Tag
Escha-W2 is a 2-bit quantized build of the Qwen3.6-35B-A3B MoE model, packaged with runtimes for local serving via an OpenAI-compatible API. It requires a 24 GB GPU (or 16 GB with trade-offs) and is available on Hugging Face.
Unsloth 成功将 GLM-5.2 模型以 2-bit 量化压缩至 238GB,可在 256GB Mac 上本地运行,保留约 82% 的准确率。
Unsloth releases a 2-bit quantized Gemma 4 12B model, only 4.66GB, runnable locally, with capabilities like autonomous online search and deep analysis similar to McKinsey consulting.