2-bit-quantization

Tag

Cards List
#2-bit-quantization

Escha-W2 (Hugging Face Repo)

TLDR AI ↗ · 2026-07-30 Cached

Escha-W2 is a 2-bit quantized build of the Qwen3.6-35B-A3B MoE model, packaged with runtimes for local serving via an OpenAI-compatible API. It requires a 24 GB GPU (or 16 GB with trade-offs) and is available on Hugging Face.

0 favorites 0 likes
#2-bit-quantization

@10xmylife: Unsloth 成功将 2-bit 版本的 GLM-5.2 部署在了 256GB 的 Mac 上

X AI KOLs Following ↗ · 2026-06-19 Cached

Unsloth 成功将 GLM-5.2 模型以 2-bit 量化压缩至 238GB,可在 256GB Mac 上本地运行,保留约 82% 的准确率。

0 favorites 0 likes
#2-bit-quantization

@VincentLogic: A 4.66 GB model actually runs at the level of a McKinsey consultant locally? Unsloth's latest 2-bit Gemma 4 12B is truly explosive. This isn't just chat – it directly transforms into a 'Super Agent' working autonomously: autonomously searching online citing 15+ sources, deeply distinguishing…

X AI KOLs Timeline ↗ · 2026-06-12 Cached

Unsloth releases a 2-bit quantized Gemma 4 12B model, only 4.66GB, runnable locally, with capabilities like autonomous online search and deep analysis similar to McKinsey consulting.

0 favorites 0 likes
← Back to home

Submit Feedback