Sherry's 3:4 ternary format (1.375 bits per weight) running on WebGPU: a 1.6 MB model that plays Connect Four as well as its 7.8 MB int8 version
Summary
A 1.6 MB model using ternary 3:4 weights runs on WebGPU in the browser, playing Connect Four as well as a 7.8 MB int8 model, demonstrating the effectiveness of ternary formats for lightweight AI deployment.
Similar Articles
Ternary Bonsai 2 (27B) just released on Hugging Face. At <6GB in size, it can even run locally in-browser on WebGPU.
Ternary Bonsai 2 is a 27B parameter model derived from Qwen3.8-27B that uses ternary weights to achieve a size under 6GB while retaining 98.2% of its intelligence, enabling it to run in-browser on WebGPU.
Ternary Bonsai: Top Intelligence at 1.58 Bits
A highly efficient AI model architecture using ternary weights (-1, 0, 1) that achieves competitive performance while requiring only 1.58 bits per parameter, enabling deployment on extremely constrained devices.
prism-ml/Ternary-Bonsai-2-27B-mlx-2bit
Prism ML released a ternary weight 27B-class AI model optimized for on-device use on Apple laptops, retaining 98.2% of full-precision intelligence with an 8.60 GB footprint and ~47 tok/s performance.
prism-ml/Ternary-Bonsai-2-27B-gguf
Release of Ternary-Bonsai-2-27B-gguf, a 27B-class language model using ternary weights for extreme compression (5.9 GB) while retaining 98.2% of FP16 intelligence, optimized for efficient inference on laptops and single GPUs.
PrismML just released Binary and Ternary Bonsai Image 4B: 1-bit/ternary text-to-image diffusion transformers that can even run 100% locally in your browser on WebGPU.
PrismML released Bonsai Image 4B models in binary and ternary quantized versions, enabling text-to-image generation to run locally in a browser via WebGPU with only 3GB size, under Apache-2.0 license.