How to pack ternary numbers in 8-bit bytes

Hacker News Top Tools

Summary

A blog post describing an efficient method to pack ternary numbers into 8-bit bytes using SIMD-friendly unpacking, achieving 1.6 bits per trit, with applications in LLM weight quantization like BitNet b1.58.

No content available
Original Article
View Cached Full Text

Cached at: 07/21/26, 09:36 AM

# How to pack ternary numbers in 8-bit bytes Source: [https://compilade.net/blog/ternary-packing](https://compilade.net/blog/ternary-packing) with efficient SIMD\-friendly unpacking --- Published:2024\-06\-26 There are 3 possible values in a digit of a ternary number\. 3 possible values, which could actually be anything\. I've been recently nerd\-sniped[1](https://compilade.net/blog/ternary-packing#fn:1)into trying to pack the ternary weights of[BitNet b1\.58](https://arxiv.org/abs/2402.17764)into something close to that theoretical ideal of`log\(3\) / log\(2\)`bits[2](https://compilade.net/blog/ternary-packing#fn:2)per ternary digit\. I'll be calling a "ternary digit" a "trit", like a "binary digit" is called a "bit"\. ## Block size Since the goal of this is to allow fast***parallel***unpacking, blocks of trits can't be infinitely big\. A small "block" size needs to be found, ideally one which is both efficient with information density and which is convenient on current hardware\. To find a good block size, we'll need to find a power of 3 for which the next power of 2 is very close\. It's very fortunate that 5 trits fit quite tight into 8 bits at`1\.6 bits`per trit\. When compared to perfect packing, this is`99\.06%`efficient\. ## 1\.6 bits per trit The basic idea with this packing scheme is simply to make a number out of the ternary digits\. Packing trits into bytes should be similar enough\. ### Fast multiplication unpacking While repeated remainder and divisions can be used to extract the digits of a number, the problem with divisions and modulo is that they are not usually supported on integers inSIMDprogramming\. A way around this is*obviously*to view numbers differently\. Would it be nice if instead of extracting the least significant digit with modulo, we could extract the most significant digit with a multiplication? Fixed point numbers to the rescue\! Tada\! Now digits can be easily extracted from the top two bits of the resulting 10\-bit number when multiplying this 8\-bit byte by 3\. This is much more convenient than modulo when unpacking withSIMD\. The only place where there are divisions in this scheme when packing trits into bytes\. This assumes that packing is done less often than unpacking, which is very true in the context ofLLMweights\. The relevant interesting line is this one: It does what is depicted in the diagram above, but multiplication is done first because these are integer operations\. Doing a ceiling division here is necessary to cancel the off\-by\-one error from truncating when extracting digits later\. To unpack*without*using the modulo operator: To convince myself that this works, I wrote a C program checking that this really is lossless: Compile and run with: And I'm getting`PASS`for each of the 243 ternary numbers which fit in 8 bits\. And this is the technique used in the ternary types in`llama\.cpp`for TriLMs and BitNet b1\.58, for which the pull request is[https://github\.com/ggml\-org/llama\.cpp/pull/8151](https://github.com/ggml-org/llama.cpp/pull/8151), withSIMDimplementations for both AVX2 and ARM NEON\.

Similar Articles

Breaking the 1.58-bit Barrier for Ternary LLMs

arXiv cs.AI

This paper introduces BITCOS, a distribution-adaptive layout for storing ternary LLM weights more efficiently, achieving up to 1.28× speedup in matrix-vector multiplication and 1.27× in inference throughput on GPUs.

Is ternary (1.58-bit) LLMs making a come back?

Reddit r/LocalLLaMA

Recent ternary 1.58-bit LLM releases from small labs demonstrate speed and medical specialization but struggle with long-horizon tasks, with optimism for future models to compete with larger architectures like Qwen.

Bitnet.cpp: Efficient Edge Inference for Ternary LLMs

Papers with Code Trending

Bitnet.cpp presents a mixed-precision matrix multiplication library for efficient edge inference of ternary LLMs like BitNet b1.58, achieving up to 6.25x speedup over full-precision baselines. The system is open-sourced on GitHub.