I ran Ternary-Bonsai-27B (2-bit) and Bonsai-27B (1-bit) on Terminal-Bench 2.0, in 8GB VRAM
Summary
A user tested quantized 1-bit and 2-bit versions of the 27B-parameter Bonsai model on Terminal-Bench 2.0, achieving results within 8GB VRAM.
Similar Articles
@sudoingX: every day someone asks how i'm running bonsai 27b on hardware that shouldn't handle it. so here's the whole thing in on…
A detailed guide on running the 27B Bonsai model on hardware with only 8GB VRAM using a 1-bit quantized version and the PrismML fork of llama.cpp, including exact server commands and configuration.
Using the Bonsai 27b 1b quant locally - regularly.
A user shares their positive experience using the 1-bit quantized version of Bonsai 27b locally on a 16GB MacBook Air for casual chat, tutoring in Go, and analyzing personal notes, praising its intelligence and small footprint.
Bonsai 27B (1-bit LLM): The First 27B-Class Model to Run on a Phone
PrismML announces Bonsai 27B, a 1-bit and ternary quantized version of Qwen3.6 27B that runs on phones and laptops, retaining 90-95% of baseline performance with a 3.9GB footprint, enabling agentic and multimodal on-device AI.
Prism-ML Bonsai Qwen 3.6 27B
Prism ML released Ternary-Bonsai-27B, a ternary-quantized version of Qwen3.6-27B that retains 95% of FP16 intelligence at a ~7.2 GB footprint, enabling full 27B-class reasoning on laptops and single GPUs with speeds up to 26 tok/s on Apple M5 Pro.
prism-ml/Bonsai-27B-mlx-1bit
Bonsai-27B is a 1-bit binary transformer model that achieves full 27B-class reasoning on a phone (iPhone 17 Pro Max) with ~3.9 GB footprint and ~11 tok/s, retaining ~90% of FP16 intelligence.