I tested Qwen3.8 27B IQ3_XXS (10.18GiB) vs Bonsai Ternary PQ2 (6.42GiB)
Summary
The author compared the performance of Qwen3.8 27B IQ3_XXS and Bonsai Ternary PQ2 on limited VRAM, finding that Qwen is faster and uses fewer tokens, while Bonsai has a smaller file size but longer generation times.
Similar Articles
Prism-ML Bonsai Qwen 3.6 27B
Prism ML released Ternary-Bonsai-27B, a ternary-quantized version of Qwen3.6-27B that retains 95% of FP16 intelligence at a ~7.2 GB footprint, enabling full 27B-class reasoning on laptops and single GPUs with speeds up to 26 tok/s on Apple M5 Pro.
My Qwen 3.8 27B tests on limited VRAM (16-20GB)
This article tests various quantized versions of the Qwen 3.8 27B AI model on limited VRAM setups, comparing their performance on tasks like animation generation, app development, and word generation.
@sudoingX: this lab took qwen 3.6 27b, the model i've been calling king of the 24gb tier all month, and crushed it down to 3.9gb. …
PrismML announces Bonsai 27B, a binary-quantized version of Qwen3.6 27B that runs on a phone using only 1.125 bits per weight, claiming 89.5% intelligence retention. The model is being independently tested by @sudoingX to verify performance.
For those with 12GB GPUs, you can now run QWEN 3.6 27B wth little loss via the new Ternary version.
A new ternary quantized version of Qwen3.6 27B, called Bonsai 27B, allows running the model on 12GB GPUs with 10x less memory and 95% of original performance, making it accessible for local deployment.
RTX 5090 Bonsai 2 27B vs Gemma 4 12B vs Qwen 3.5 9B Japanese voxel pagoda
This article compares the performance of Bonsai 2 27B, Gemma 4 12B, and Qwen 3.5 9B models on generating a Japanese voxel pagoda using an RTX 5090, concluding that Bonsai offers superior intelligence and detail for its memory footprint, benefiting the local AI community.