Sherry's 3:4 ternary format (1.375 bits per weight) running on WebGPU: a 1.6 MB model that plays Connect Four as well as its 7.8 MB int8 version

Reddit r/LocalLLaMA Models

Summary

A 1.6 MB model using ternary 3:4 weights runs on WebGPU in the browser, playing Connect Four as well as a 7.8 MB int8 model, demonstrating the effectiveness of ternary formats for lightweight AI deployment.

Not an LLM, but the ternary findings should carry over, and we hadn't seen Sherry-style 3:4 weights run in a browser before. Disclosure: this is our work at Precisit, everything is MIT. What it is A 7.4M-parameter one-pass scorer (the jevlike family): the board goes in, one score per legal column comes out. No search. Weights in T34, Sherry's 3:4 format: in every four weights one is zero and three are ±1, so four weights fit in 5 bits. One fp16 scale per 128 weights gives 1.375 bits per weight. The embedding is int8; norms and biases are fp16. It runs in the browser on a small WebGPU runtime: 1.1 ms per move (idle M5 Pro, Chrome). Model File size vs depth-4 bot vs depth-6 bot dense (fp32) 29.7 MB 0.92 0.89 T34, trained ternary 1.59 MB 0.93 0.91 T34, fine-tuned from dense 1.59 MB 0.89 0.9 T34, converted after training 1.59 MB 0.13 0.11 Base243 (TQ1_0 style), trained 1.93 MB 0.89 0.88 200 games each, both sides play a random move 5% of the time, a win counts 1 and a draw ½. What we learned Converting the finished model to 3:4 collapsed it (0.13 against the depth-4 bot). Training with the format in the forward pass fixed it completely, whether from scratch or fine-tuning. Attention's q/k/v matrices are the sensitive ones. Group size (64/128/256) barely mattered. Seeds matter: two runs of the same T34 recipe scored 0.945 and 0.882. Play it: https://precisit.github.io/onepass-web/demo/c4-size/ Code, models, every result: https://github.com/precisit/onepass-webgpu-ternary The write-up: https://precisit.com/en/blog/onepass-c4-size/ Has anyone gotten post-training 3:4 conversion to work on models, or does it need training?
Original Article

Similar Articles

Ternary Bonsai: Top Intelligence at 1.58 Bits

Hacker News Top

A highly efficient AI model architecture using ternary weights (-1, 0, 1) that achieves competitive performance while requiring only 1.58 bits per parameter, enabling deployment on extremely constrained devices.

prism-ml/Ternary-Bonsai-2-27B-mlx-2bit

Hugging Face Models Trending

Prism ML released a ternary weight 27B-class AI model optimized for on-device use on Apple laptops, retaining 98.2% of full-precision intelligence with an 8.60 GB footprint and ~47 tok/s performance.

prism-ml/Ternary-Bonsai-2-27B-gguf

Hugging Face Models Trending

Release of Ternary-Bonsai-2-27B-gguf, a 27B-class language model using ternary weights for extreme compression (5.9 GB) while retaining 98.2% of FP16 intelligence, optimized for efficient inference on laptops and single GPUs.