@HuggingModels: Meet Ternary-Bonsai-2-27B: a 27B parameter model squeezed into 2-bit ternary format. Runs on llama.cpp with CUDA and Me…

X AI KOLs Timeline Models

Summary

Ternary-Bonsai-2-27B is a 27B parameter AI model compressed into 2-bit ternary format, running on llama.cpp with CUDA and Metal support, facilitating lightweight on-device AI deployment.

Meet Ternary-Bonsai-2-27B: a 27B parameter model squeezed into 2-bit ternary format. Runs on llama.cpp with CUDA and Metal support. 405K downloads and counting. On-device AI just got a whole lot lighter. https://t.co/QGH8QR9nAe
Original Article
View Cached Full Text

Cached at: 09/19/26, 03:05 PM

Meet Ternary-Bonsai-2-27B: a 27B parameter model squeezed into 2-bit ternary format. Runs on llama.cpp with CUDA and Metal support. 405K downloads and counting. On-device AI just got a whole lot lighter. https://t.co/QGH8QR9nAe

Similar Articles

Ternary Bonsai: Top Intelligence at 1.58 Bits

Hacker News Top

A highly efficient AI model architecture using ternary weights (-1, 0, 1) that achieves competitive performance while requiring only 1.58 bits per parameter, enabling deployment on extremely constrained devices.

prism-ml/Ternary-Bonsai-2-27B-gguf

Hugging Face Models Trending

Release of Ternary-Bonsai-2-27B-gguf, a 27B-class language model using ternary weights for extreme compression (5.9 GB) while retaining 98.2% of FP16 intelligence, optimized for efficient inference on laptops and single GPUs.

Bonsai-27B & Ternary-Bonsai-27B - Updates (on PRs)

Reddit r/LocalLLaMA

Updates on Bonsai-27B and Ternary-Bonsai-27B models, detailing upstream merge status in llama.cpp across CPU, Metal, CUDA, Vulkan backends, and discussing model limitations and roadmap.