Got a 27B model running locally on a Jetson Orin NX 16GB (1-bit). still kind of amazed it works
Summary
User reports successfully running a 27B parameter model quantized to 1-bit on a Jetson Orin NX 16GB edge device, expressing amazement at the feasibility.
Similar Articles
Using the Bonsai 27b 1b quant locally - regularly.
A user shares their positive experience using the 1-bit quantized version of Bonsai 27b locally on a 16GB MacBook Air for casual chat, tutoring in Go, and analyzing personal notes, praising its intelligence and small footprint.
PrismML Bonsai 27B is surprisingly usable on the Jetson Orin Nano 8GB
PrismML's Bonsai 27B model runs on the Jetson Orin Nano 8GB with 4.31 tokens/s and 27 t/s prompt processing, using 6.2GB RAM and about 25W power. It indicates surprisingly usable edge AI performance.
Bonsai 27B (1-bit LLM): The First 27B-Class Model to Run on a Phone
PrismML announces Bonsai 27B, a 1-bit and ternary quantized version of Qwen3.6 27B that runs on phones and laptops, retaining 90-95% of baseline performance with a 3.9GB footprint, enabling agentic and multimodal on-device AI.
@sudoingX: this lab took qwen 3.6 27b, the model i've been calling king of the 24gb tier all month, and crushed it down to 3.9gb. …
PrismML announces Bonsai 27B, a binary-quantized version of Qwen3.6 27B that runs on a phone using only 1.125 bits per weight, claiming 89.5% intelligence retention. The model is being independently tested by @sudoingX to verify performance.
I ran Ternary-Bonsai-27B (2-bit) and Bonsai-27B (1-bit) on Terminal-Bench 2.0, in 8GB VRAM
A user tested quantized 1-bit and 2-bit versions of the 27B-parameter Bonsai model on Terminal-Bench 2.0, achieving results within 8GB VRAM.