2-bit

Tag

Cards List
#2-bit

EschaLabs/Qwen3.8-27B-Escha-W2

Hugging Face Models Trending · 2026-08-20 Cached

Escha Labs has released a 2-bit quantized version of the Qwen3.8-27B AI model, enabling it to run on a single 24GB consumer GPU with up to 64k context while maintaining performance comparable to FP8 references.

0 favorites 0 likes
#2-bit

1-bit / 2-bit / Ternary / Bitnet Models - Updates & Tracking

Reddit r/LocalLLaMA · 2026-08-18

This article tracks updates on various low-bit AI models, including Bonsai, BitCPM, and others, with performance improvements and compatibility updates for backends like llama.cpp.

0 favorites 0 likes
#2-bit

Deepseek V4 Flash 2, 3 and 4 bits GGUFs

Reddit r/LocalLLaMA · 2026-07-01 Cached

GGUF quantizations of DeepSeek V4 Flash in 2-bit, 3-bit, and 4-bit precisions, made available on Hugging Face for local inference with tools like llama.cpp and Ollama.

0 favorites 0 likes
#2-bit

Calibrating 2-bit GGUFs (<10Gb) for agentic coding tasks

Reddit r/LocalLLaMA · 2026-06-18

This article introduces calibrated 2-bit GGUF quantizations of the Qwopus3.6-27B-Coder model for agentic coding tasks, demonstrating that the IQ2_M quant (9.74 GiB) achieves a 63% pass rate on the SWE-rebench benchmark, comparable to a Q5_K_M quant at half the size.

0 favorites 0 likes
#2-bit

UniSVQ: 2-bit Unified Scalar-Vector Quantization

arXiv cs.CL · 2026-06-10 Cached

UniSVQ proposes a unified 2-bit quantization framework that bridges scalar and vector quantization by parameterizing codewords as an affine transform of integer lattices, achieving state-of-the-art performance among scalar methods and matching vector methods with higher throughput.

0 favorites 0 likes
← Back to home

Submit Feedback