@HuggingModels: Meet Ternary-Bonsai-2-27B: a 27B parameter model squeezed into 2-bit ternary format. Runs on llama.cpp with CUDA and Me…
Summary
Ternary-Bonsai-2-27B is a 27B parameter AI model compressed into 2-bit ternary format, running on llama.cpp with CUDA and Metal support, facilitating lightweight on-device AI deployment.
View Cached Full Text
Cached at: 09/19/26, 03:05 PM
Meet Ternary-Bonsai-2-27B: a 27B parameter model squeezed into 2-bit ternary format. Runs on llama.cpp with CUDA and Metal support. 405K downloads and counting. On-device AI just got a whole lot lighter. https://t.co/QGH8QR9nAe
Similar Articles
Ternary Bonsai: Top Intelligence at 1.58 Bits
A highly efficient AI model architecture using ternary weights (-1, 0, 1) that achieves competitive performance while requiring only 1.58 bits per parameter, enabling deployment on extremely constrained devices.
prism-ml/Ternary-Bonsai-2-27B-gguf
Release of Ternary-Bonsai-2-27B-gguf, a 27B-class language model using ternary weights for extreme compression (5.9 GB) while retaining 98.2% of FP16 intelligence, optimized for efficient inference on laptops and single GPUs.
Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Introducing Ternary Bonsai 2 27B, a highly compressed AI model that retains 98.2% of performance while being 9x smaller in footprint, enabling efficient local deployment for tasks like reasoning, coding, and multimodal processing.
Ternary Bonsai 2 (27B) just released on Hugging Face. At <6GB in size, it can even run locally in-browser on WebGPU.
Ternary Bonsai 2 is a 27B parameter model derived from Qwen3.8-27B that uses ternary weights to achieve a size under 6GB while retaining 98.2% of its intelligence, enabling it to run in-browser on WebGPU.
Bonsai-27B & Ternary-Bonsai-27B - Updates (on PRs)
Updates on Bonsai-27B and Ternary-Bonsai-27B models, detailing upstream merge status in llama.cpp across CPU, Metal, CUDA, Vulkan backends, and discussing model limitations and roadmap.