Apps & Tools
Qwen3.8-27B-Escha-W2 GGUF
A high-performance llama.cpp port of Escha's 2-bit quantization using custom CUDA kernels.
byAjay9o9
PlayableOpen source
Qwen3.8-27B-Escha-W2 GGUF media is blocked
Allow external media to connect to the provider and play this content.
Description
This project implements a custom llama.cpp fork and CUDA kernel (GGML_OP_ESCHA_MUL_MAT) to natively decode EschaLabs' 2-bit quantized weights without unpacking to dense tensors. It uses tensor cores to achieve 3.3x faster prefill and improved decode speeds on RTX 3090 hardware while maintaining 99.7% logit agreement with the original runtime.
Descriptions, tags, and model credits may be AI-generated or inferred from public sources and can be incomplete or wrong. Learn more.