Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Summary
Bonsai 2 27B is a compressed AI model that achieves near-lossless performance in a 9x smaller footprint, with setup instructions provided for use with Prism's llama.cpp fork.
View Cached Full Text
Cached at: 09/25/26, 07:44 PM
Similar Articles
Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Introducing Ternary Bonsai 2 27B, a highly compressed AI model that retains 98.2% of performance while being 9x smaller in footprint, enabling efficient local deployment for tasks like reasoning, coding, and multimodal processing.
PrismML hopes its tiny LLM will change how we all use AI
PrismML, a Caltech-founded startup, released Bonsai 2, a highly compressed LLM that reduces memory usage by 9-10x with minimal performance loss, aiming to enable AI on consumer devices like PCs and smartphones.
@HuggingModels: Meet Ternary-Bonsai-2-27B: a 27B parameter model squeezed into 2-bit ternary format. Runs on llama.cpp with CUDA and Me…
Ternary-Bonsai-2-27B is a 27B parameter AI model compressed into 2-bit ternary format, running on llama.cpp with CUDA and Metal support, facilitating lightweight on-device AI deployment.
prism-ml/Bonsai-27B-gguf
Prism ML releases Bonsai-27B-gguf, a 27-billion parameter language model with binary (1.125-bit) weights, achieving a ~14x size reduction while retaining ~90% of FP16 reasoning performance. It runs on consumer hardware with high throughput.
@xenovacom: Bonsai 27B just changed the local LLM game forever. 1-bit quantization shrinks it from 54GB to just 3.8GB (-93%), while…
Bonsai 27B achieves 93% size reduction via 1-bit quantization while retaining 90% intelligence, enabling local browser inference with custom WebGPU kernels.