Breaking the 1.58-bit Barrier for Ternary LLMs
Summary
This paper introduces BITCOS, a distribution-adaptive layout for storing ternary LLM weights more efficiently, achieving up to 1.28× speedup in matrix-vector multiplication and 1.27× in inference throughput on GPUs.
View Cached Full Text
Cached at: 09/16/26, 08:59 AM
# Breaking the 1.58-bit Barrier for Ternary LLMs
Source: [https://arxiv.org/html/2609.16338](https://arxiv.org/html/2609.16338)
###### Abstract
Ternary Large Language Models \(LLM\) store every weight as one of three symbols\{−1,0,\+1\}\\\{\-1,0,\+1\\\}, so the cost of a ternary model is conventionally referenced to the information\-theoreticlog23≈1\.585\\log\_\{2\}3\\approx 1\.585bits per weight\. The prevailing deployment format packs five ternary weights into one byte \(*five\-trit packing*\), and due to the power\-of\-two group sizes used in practice this rounds up to1\.6251\.625bits per weight\. This effective storage bit\-width treats the three symbols\{−1,0,\+1\}\\\{\-1,0,\+1\\\}as equiprobable\. We measure the actual symbol distribution of 29 ternary LLM models and find that zeros account for up to51\.5%51\.5\\%of all weights\. Motivated by this finding, we introduce BITCOS, a simple distribution\-adaptive layout comprised of a dense presence bitmap plus a compacted sign vector, and costs2−z2\-zbits per weight element given a zero densityzzin the model’s weights\. BITCOS stores weights more compactly than the five\-trit packing in 26 of the 29 tested models, and reaches1\.4851\.485bits per weight on the sparsest of them\. BITCOS is amenable to efficient unpacking on modern processors and GPUs, and we present optimized unpacking sequences for AVX\-512, AVX2 and Intel Xe2 GPUs\. Measured against production state\-of\-the\-art ternary matrix\-vector multiplication kernels, at the zero densities real\-world ternary models exhibit, the realized gain with our proposed layout is up to1\.28×1\.28\\times\. Finally, we illustrate end\-to\-end LLM inference results on 5 different platforms \(client and server CPUs, integrated and discrete Xe2 GPUs\) where decode throughput improves by up to1\.18×1\.18\\timeson CPUs and1\.27×1\.27\\timeson GPUs\.
## IIntroduction
Ternary weight quantization restricts every weight to\{−1,0,\+1\}\\\{\-1,0,\+1\\\}scaled by a group\-wise factor\[[1](https://arxiv.org/html/2609.16338#bib.bib1),[2](https://arxiv.org/html/2609.16338#bib.bib2)\]\. Three equiprobable symbols carrylog23≈1\.585\\log\_\{2\}3\\approx 1\.585bits of information, and this theoretical bound is the reference for ternary storage\. Nevertheless, in practice the actual storage cost is determined by how those symbols are packed\. Five trits \(ternary digits\) fit in a byte \(35=243≤2563^\{5\}=243\\leq 256\), which approaches the bound at8/5=1\.68/5=1\.6bits per weight\. However, deployed implementations quantize in blocks of power\-of\-two weights like 128, and 128 is not a multiple of 5: a block needs⌈128/5⌉=26\\lceil 128/5\\rceil=26payload bytes, so the rate stored in practice is26×8/128=1\.62526\\times 8/128=1\.625bits per weight\. Still, this effective storage bit\-width treats the three symbols\{−1,0,\+1\}\\\{\-1,0,\+1\\\}as equiprobable\.
Fig\. 1:Effective bit\-widths for various families of ternary LLM models, comparing the five\-trit packing and the proposed BITCOS layout\. The sloped line represents the BITCOS bit\-width2−z2\-z, while the plateau indicates the fixed five\-trit bit\-width of1\.6251\.625bits per weight\. Each marker corresponds to a model from Table[I](https://arxiv.org/html/2609.16338#S1.T1), placed at its measured zero density\. Models with zero densityzzabove0\.3750\.375benefit from the BITCOS layout \(green plot area\), as such 26 out of the 29 models achieve lower effective bit\-width with BITCOS\.TABLE I:Effective bit\-widths for SOTA ternary LLM models\. “% 0” is the measured zero densityzz\. “Symbols” column counts the ternary codes alone and “\+\+scale” adds the measured 16\-bit scale overhead\. The symbols\-only rate is2−z2\-zfor BITCOS and for every model2\.0002\.000and1\.6251\.625bits per weight for 2\-bit and 5\-trit packing respectively\. “red\.” is the size reduction of BITCOS over the corresponding format with scales included\. A value\>1\>1means BITCOS stores the model more compactly\.BITCOS2\-bit packing5\-trit per byteTernary model% 0Symbols\+\+scale\+\+scalered\.\+\+scalered\.∙\\bulletBitNet b1\.58 2B4T42\.191\.578—2\.0001\.27×1\.27\\times1\.6251\.03×1\.03\\times■\\blacksquareBonsai 1\.7B39\.891\.6011\.7262\.1251\.23×1\.23\\times1\.7501\.01×1\.01\\timesBonsai 4B37\.711\.6231\.7482\.1251\.22×1\.22\\times1\.7501\.00×1\.00\\timesBonsai 8B38\.251\.6181\.7432\.1251\.22×1\.22\\times1\.7501\.00×1\.00\\timesBonsai 27B29\.661\.7031\.8282\.1251\.16×1\.16\\times1\.7500\.96×0\.96\\times▲\\blacktriangleCAT\-Q Qwen3\-1\.7B51\.481\.4851\.6102\.1251\.32×1\.32\\times1\.7501\.09×1\.09\\timesCAT\-Q Qwen3\-8B46\.701\.5331\.6582\.1251\.28×1\.28\\times1\.7501\.06×1\.06\\timesCAT\-Q Qwen3\-30B\-A3B32\.881\.6711\.7962\.1251\.18×1\.18\\times1\.7500\.97×0\.97\\timesCAT\-Q Qwen3\-32B47\.111\.5291\.6542\.1251\.28×1\.28\\times1\.7501\.06×1\.06\\timesCAT\-Q Qwen3\-235B\-A22B34\.071\.6591\.7842\.1251\.19×1\.19\\times1\.7500\.98×0\.98\\times⧫\\blacklozengeParetoQ 125M41\.071\.5891\.6132\.0231\.25×1\.25\\times1\.6481\.02×1\.02\\timesParetoQ 350M41\.581\.5841\.5982\.0141\.26×1\.26\\times1\.6391\.03×1\.03\\timesParetoQ 600M43\.271\.5671\.5792\.0121\.27×1\.27\\times1\.6371\.04×1\.04\\timesParetoQ 1B44\.181\.5581\.5692\.0101\.28×1\.28\\times1\.6351\.04×1\.04\\timesParetoQ 1\.5B47\.101\.5291\.5372\.0081\.31×1\.31\\times1\.6331\.06×1\.06\\times▼\\blacktriangledownTriLM 99M40\.971\.5901\.6182\.0271\.25×1\.25\\times1\.6521\.02×1\.02\\timesTriLM 190M40\.671\.5931\.6112\.0181\.25×1\.25\\times1\.6431\.02×1\.02\\timesTriLM 390M40\.791\.5921\.6062\.0141\.25×1\.25\\times1\.6391\.02×1\.02\\timesTriLM 560M40\.731\.5931\.6042\.0111\.25×1\.25\\times1\.6361\.02×1\.02\\timesTriLM 830M40\.541\.5951\.6042\.0091\.25×1\.25\\times1\.6341\.02×1\.02\\timesTriLM 1\.1B40\.391\.5961\.6052\.0091\.25×1\.25\\times1\.6341\.02×1\.02\\timesTriLM 1\.5B40\.211\.5981\.6062\.0081\.25×1\.25\\times1\.6331\.02×1\.02\\timesTriLM 2\.4B39\.701\.6031\.6102\.0071\.25×1\.25\\times1\.6321\.01×1\.01\\timesTriLM 3\.9B38\.701\.6131\.6182\.0051\.24×1\.24\\times1\.6301\.01×1\.01\\times★\\bigstarMaple 20B\-A1B40\.671\.5931\.6092\.0161\.25×1\.25\\times1\.6411\.02×1\.02\\times◀\\blacktriangleleftBitCPM\-CANN 0\.5B37\.671\.6231\.6362\.0121\.23×1\.23\\times1\.6371\.00×1\.00\\timesBitCPM\-CANN 1B38\.391\.6161\.6232\.0061\.24×1\.24\\times1\.6311\.01×1\.01\\timesBitCPM\-CANN 3B38\.061\.6191\.6242\.0051\.23×1\.23\\times1\.6301\.00×1\.00\\timesBitCPM\-CANN 8B39\.301\.6071\.6102\.0031\.24×1\.24\\times1\.6281\.01×1\.01\\timesWe measure the actual symbol distribution of 29 state\-of\-the\-art \(SOTA\) ternary LLM models and find that zeros account for up to51\.5%51\.5\\%of all weights; Table[I](https://arxiv.org/html/2609.16338#S1.T1)lists the measured zero density of every model\. More specifically we benchmarked seven ternary model families: i\)*BitNet*\[[1](https://arxiv.org/html/2609.16338#bib.bib1),[2](https://arxiv.org/html/2609.16338#bib.bib2)\], the 2B\-parameter model trained from scratch on 4T tokens that established the1\.581\.58\-bit reference point; ii\)*Bonsai*\[[3](https://arxiv.org/html/2609.16338#bib.bib6)\], four dense checkpoints from1\.71\.7B to2727B; iii\)*CAT\-Q*\[[4](https://arxiv.org/html/2609.16338#bib.bib15)\], five post\-training quantizations of Qwen3 from1\.71\.7B to235235B, including two mixture\-of\-experts models; iv\)*ParetoQ*\[[5](https://arxiv.org/html/2609.16338#bib.bib3)\], five small checkpoints \(125125M–1\.51\.5B\) from a study of low\-bit quantization\-aware training; v\)*TriLM*\[[6](https://arxiv.org/html/2609.16338#bib.bib4)\], the nine\-model Spectra suite \(9999M–3\.93\.9B\) pretrained in ternary; vi\)*Maple*\[[7](https://arxiv.org/html/2609.16338#bib.bib7)\], a2020B\-A1B ternary mixture\-of\-experts reasoning model; and vii\)*BitCPM\-CANN*\[[8](https://arxiv.org/html/2609.16338#bib.bib8)\], four ternary checkpoints \(0\.50\.5B–88B\) trained from scratch\. Motivated by the finding that the zero density of many ternary models is significantly higher than the density of either of the other two codes\{−1,\+1\}\\\{\-1,\+1\\\}, we introduce BITCOS \(BITmap andCOmpactedSigns\): a simple distribution\-adaptive layout for ternary tensors\. Given a zero densityzz, BITCOS spends one presence bit on every weight entry and one sign bit only on the non\-zero weights, thus the layout effectively achieves2−z2\-zbits per weight\. Such a layout improves the effective bit\-width compared to the deployed five\-trit packing oncez\>0\.375z\>0\.375and is strictly more efficient than the widely\-adopted 2\-bit packing for all zero densities\. Figure[1](https://arxiv.org/html/2609.16338#S1.F1)places every model of Table[I](https://arxiv.org/html/2609.16338#S1.T1)at its measured zero density, on whichever of the two rates it meets first: the sloped2−z2\-zline when the BITCOS sparse layout is more efficient, or the1\.6251\.625plateau when the five\-trit fixed\-rate packing has lower bit\-width\. Models with zero densityzzabove0\.3750\.375benefit from the BITCOS layout \(green plot area\), as such 26 out of the 29 models achieve lower effective bit\-width with BITCOS compared to the five\-trit packing layout\.
However, the effective bit\-width is only one aspect of the overall inference\. The actual performance of ternary models during inference also depends on how efficiently the packed weights can be unpacked and utilized in matrix\-vector multiplication kernels\. The decode phase of LLM inference with a small batch\-size is bandwidth\-bound, so the time per token tracks the bit\-width of the weight datatype\[[9](https://arxiv.org/html/2609.16338#bib.bib17)\], which is precisely the regime that on\-device and agentic deployments exercise\[[10](https://arxiv.org/html/2609.16338#bib.bib12)\]\. BITCOS not only reduces the storage cost but also is amenable to efficient unpacking and computation on modern CPUs and Intel GPUs\. Measured end to end over 7 ternary LLM checkpoints and five platforms, decode throughput improves by up to1\.18×1\.18\\timeson a6464\-core server CPU,1\.15×1\.15\\timeson a2424\-core client CPU,1\.27×1\.27\\timeson a discrete Xe2 GPU and1\.22×1\.22\\timeson an integrated Xe2 GPU\. The performance gain however is not universal: on an 8\-core, bandwidth\-rich client platform with enough bandwidth per core to leave the unpack sequence exposed, the smaller payload does not translate into performance benefits\. Therefore, we develop a simple two\-term roofline model to assess the limitations of the BITCOS\-based kernels\.
This paper makes the following contributions:
1. 1\.A novel ternary sparse layout \(which we name BITCOS\) comprising of a presence bitmap plus a compacted sign vector, that costs2−z2\-zbits per weight for a zero densityzzand is amenable to efficient unpacking on modern CPUs and GPUs\.
2. 2\.Optimized instruction sequences to unpack the proposed BITCOS layout and perform matrix multiplication operations efficiently on modern x86 CPUs \(client CPUs with performance and efficiency cores, and server CPUs with high core counts\) and Intel Xe2 GPUs \(both integrated and discrete GPUs\)\.
3. 3\.A roofline model to assess the efficacy and the limitations of our CPU microkernels\.
4. 4\.Performance evaluation of the proposed layout with microbenchmarks and end\-to\-end LLM inference over seven ternary LLM checkpoints on modern server/client CPUs and integrated/discrete GPUs, illustrating decode throughput improvements by up to1\.18×1\.18\\timeson CPUs and1\.27×1\.27\\timeson GPUs\.
## IIThe BITCOS Layout:BITmap \+COmpactedSigns
### II\-ALayout definition and storage cost
We propose the BITCOS layout that leverages the inherent*unstructured*sparsity of ternary weights and stores a ternary tensor as two pieces:
1. 1\.a*presence bitmap*of one bit per weight, set where the weight is non\-zero; and
2. 2\.a*sign vector*of one bit per*non\-zero*weight, in tensor order\.
Figure[2](https://arxiv.org/html/2609.16338#S2.F2)shows both pieces on a small 8×\\times8 tensor\. The bitmap has the full tensor length; the sign vector is compacted to the population count of the bitmap\. Assuming a zero density ofzz, the cost per weight is therefore:
B\(z\)=1\+\(1−z\)=2−zbitsB\(z\)\\;=\\;1\+\(1\-z\)\\;=\\;2\-z\\quad\\text\{bits\}\(1\)
Fig\. 2:The BITCOS layout on an 8×\\times8 tensor\. Each column of BITCOS is one mask whose bitrrrecords whether rowrris non\-zero, so a decoder reads presence for a whole block in a single load\. The compacted sign vector carries one bit per non\-zero only, in the same\(k,r\)\(k,r\)order, so thejj\-th sign bit belongs to thejj\-th set bit of the bitmap\. Zeros consume a bitmap bit and nothing else, which is why the effective rate is2−z2\-z\. In this example, 30 out of the 64 weights are zero \(z=0\.469z=0\.469\), so the tensor costs 64 presence plus 34 sign bits, or1\.5311\.531bits per weight \(which is belowlog23\\log\_\{2\}3\)\.The cost per weight falls linearly inzz, so the BITCOS format yields substantial benefits over 2\-bit or five\-trit packing at zero\-heavy distributions\. Table[I](https://arxiv.org/html/2609.16338#S1.T1)compares the per\-weight cost of different formats over the 29 measured checkpoints\. Column “% 0” is the measured zero densityzz, column “Symbols” counts the ternary codes alone and the “\+\+scale” columns add the measured 16\-bit scale overhead of the corresponding model\. The “red\.” columns report the size reduction of BITCOS over the 2\-bit and the five\-trit per byte packing respectively; a value greater than11means BITCOS stores the model more compactly than that format\. We conclude that BITCOS improves the effective bit\-width compared to the deployed five\-trit packing oncez\>0\.375z\>0\.375\(26 out of 29 ternary LLM models\) and is strictly more efficient than the widely\-adopted 2\-bit packing for all zero densities\.
### II\-BBITCOS unpack sequence with x86 AVX\-512 instructions
The BITCOS layout reconstructs 16\-bit weights in a short mask\-driven sequence assuming the target ISA supports masks\. In Figure[3](https://arxiv.org/html/2609.16338#S2.F3)we illustrate such an exemplary sequence with AVX\-512 instructions\. Two vectors are prepared once per group of 128 reduction elements and then reused for every one of its iterations:zmm3holding the 32 fp16 group scales, one per row, andzmm4holding a copy with the sign bit set, obtained by a singlevporqagainst a broadcast0x8000\. Because the sign of an IEEE half lives in the most significant bit and the group scales are non\-negative by construction, that OR is an exact negation, sozmm3andzmm4hold\+s\+sand−s\-srespectively for each of the 32 rows\. The per\-iteration work is then:
1. 1\.Load the presence mask\.Read 32 bitmap bits into a general register and then intok1\.
2. 2\.Place the signs\.Take the next 32 bits of the sign stream and scatter them viapdepinto the bit positions thatk1marks as present\. The sign vector is stored compacted, so itsjj\-th bit belongs to thejj\-th non\-zero of the group’s bitmap\.pdepis exactly the scatter that undoes the compaction \(see Figure[4](https://arxiv.org/html/2609.16338#S2.F4)\)\.
3. 3\.Select\.A zero\-masked move ofzmm3underk1puts\+s\+sat every present lane and exact\+0\+0elsewhere; a merge\-masked move ofzmm4under the deposited mask overwrites the negative ones\.
Figure[4](https://arxiv.org/html/2609.16338#S2.F4)illustrates what that deposit does:pdeptakes the low bits of its source in order and drops them at the positions the mask selects, leaving every unselected position zero\. In the AVX\-512 assembly,r13walks the bitmap,r9is the base of the sign stream,r8the running bit position within it,r14walks the activations, andzmm2is one accumulator for the fused multiply\-add \(FMA\) operation\. In total we get 17 instructions, of which the weight unpacking portion corresponds to 3 instructions: onepdepand 2 masked moves\. The remaining 14 are not unpacking: one loads the bitmap word, 5 compute the data\-dependent address of the sign window, oneshrxperforms the unaligned 64\-bit read and its alignment in a single operation, 2 are software prefetches, 2 advance the bit position by the population count, 2 move masks, and the last is the fused multiply\-add \(FMA\), which absorbs the activation broadcast as an embedded operand\.
*once per group of 128:*vmovdqu64\(%r10,%rax,2\),%zmm3 ; 32 scalesvporq%zmm0,%zmm3,%zmm4 ; negated copy*per 32 weights:*mov\-0x404\(%r13,%rax,8\),%r15d ; bitmap\[k\]mov%r8d,%ecx ; copy bitposmov%r8,%rbxshr$0x3,%rbx ; byte offsetand%rdx,%rbx ; to 32\-bit wordand$0x1f,%cl ; shift countshrx%rcx,\(%r9,%rbx,1\),%rdi ; sign windowprefetcht0\-0x4\(%r13,%rax,8\) ; bitmappopcnt%r15d,%ecxprefetcht00x200\(%r9,%rbx,1\) ; signsadd%r8,%rcx ; bitpos \+= popcntpdep%r15d,%edi,%edi ; to lane orderkmovd%r15d,%k1 ; presence maskvmovdqu16%zmm3,%zmm5\{%k1\}\{z\} ;\+s\+s, else\+0\+0kmovd%edi,%k1 ; sign maskvmovdqu16%zmm4,%zmm5\{%k1\} ;−s\-sif negativevfmadd231ph\-0x2\(%r14,%rax,4\)\{1to32\},%zmm5,%zmm2Fig\. 3:AVX\-512 instruction sequence to unpack the BITCOS layout for one 32\-row block\.*dst := \_pdep\_u32\(a, mask\)*
dst := 0; k := 0
FOR m := 0 TO 31
IF mask\[m\] == 1 THEN
dst\[m\] := a\[k\]; k := k \+ 1
FI
ENDFOR
*example, bit 0 leftmost, 8 lanes shown:*
mask10110010presence: 4 non\-zerosa1011––––signs, compacteddst10010010signs, lane orderFig\. 4:pdepparallel bit deposit\. In this example, only the low four bits ofaare consumed, one per set bit of the mask\. The sign stream stores one bit per non\-zero weight, so undoing that compaction means scattering those bits onto the set positions of the presence mask, which is the instruction’s definition\.
### II\-CBITCOS unpack sequence with x86 AVX2 instructions
The sequence of Section[II\-B](https://arxiv.org/html/2609.16338#S2.SS2)depends on two AVX\-512 facilities: mask registers, which make a 32\-bit presence pattern directly usable as a predicate, and FP16 compute instructions\. The client CPUs we target support neither, thus we implemented a second kernel targeting the AVX2 ISA with AVX\-VNNI\-INT8 compute capabilities, which reconstructs the ternary weights\{−1,0,\+1\}\\\{\-1,0,\+1\\\}asint8and contracts againstint8activations \(see Figure[5](https://arxiv.org/html/2609.16338#S2.F5)\)\.
For each of the 32 rows, letpi∈\{0,1\}p\_\{i\}\\in\\\{0,1\\\}indicate that the weight is present, and letni∈\{0,1\}n\_\{i\}\\in\\\{0,1\\\}indicate that the present weight is negative\. The bitmap supplies thepip\_\{i\}bits\. Afterpdepreturns the compact signs to their row positions, its result supplies thenin\_\{i\}bits\. Since AVX2 has no mask registers, we materialize each presence and sign vector as byte masksPi=−piP\_\{i\}=\-p\_\{i\}andNi=−niN\_\{i\}=\-n\_\{i\}: a true lane is0xFF\(−1\-1as a signed byte\), and a false lane is zero\. The desired ternary byte is then
wi=pi−2ni\.w\_\{i\}=p\_\{i\}\-2n\_\{i\}\.\(2\)In the example of Figure[2](https://arxiv.org/html/2609.16338#S2.F2), the columnk2k\_\{2\}holds the ternary values\(\+1,\+1,\+1,\+1,−1,0,−1,\+1\)\(\+1,\+1,\+1,\+1,\-1,0,\-1,\+1\)for rowsi=0…7i=0\\ldots 7, so the bitmap word for that column is𝟶𝚡𝙳𝙵\\mathtt\{0xDF\}and the two predicates are
p=\(1,1,1,1,1,0,1,1\),n=\(0,0,0,0,1,0,1,0\)\.p=\(1,1,1,1,1,0,1,1\),\\qquad n=\(0,0,0,0,1,0,1,0\)\.\(3\)Equation \([2](https://arxiv.org/html/2609.16338#S2.E2)\) returns the unpacked column exactly: rows00–33and77give1−0=\+11\-0=\+1, rows44and66give1−2=−11\-2=\-1, and the absent row55gives0−0=00\-0=0\. Materialized as bytes these areP=\(𝙵𝙵,𝙵𝙵,𝙵𝙵,𝙵𝙵,𝙵𝙵,𝟶𝟶,𝙵𝙵,𝙵𝙵\)P=\(\\mathtt\{FF\},\\mathtt\{FF\},\\mathtt\{FF\},\\mathtt\{FF\},\\mathtt\{FF\},\\mathtt\{00\},\\mathtt\{FF\},\\mathtt\{FF\}\)and, after the mask with0xFE,T=\(𝟶𝟶,𝟶𝟶,𝟶𝟶,𝟶𝟶,𝙵𝙴,𝟶𝟶,𝙵𝙴,𝟶𝟶\)T=\(\\mathtt\{00\},\\mathtt\{00\},\\mathtt\{00\},\\mathtt\{00\},\\mathtt\{FE\},\\mathtt\{00\},\\mathtt\{FE\},\\mathtt\{00\}\), sovpsubbleaves\(𝟶𝟷,𝟶𝟷,𝟶𝟷,𝟶𝟷,𝙵𝙵,𝟶𝟶,𝙵𝙵,𝟶𝟷\)\(\\mathtt\{01\},\\mathtt\{01\},\\mathtt\{01\},\\mathtt\{01\},\\mathtt\{FF\},\\mathtt\{00\},\\mathtt\{FF\},\\mathtt\{01\}\), which is the original column asint8\. Note thatnnis not what the format stores\. Only thepopcount\(p\)=7\\operatorname\{popcount\}\(p\)=7sign bits\(0,0,0,0,1,1,0\)\(0,0,0,0,1,1,0\)are stored, so the sign of row66is the sixth stored bit rather than the seventh\. Recovering thennof \([3](https://arxiv.org/html/2609.16338#S2.E3)\) from those seven bits is exactly thepdepof Figure[5](https://arxiv.org/html/2609.16338#S2.F5)\.
Fig\. 5:AVX2 instruction sequence to unpack the BITCOS layout for one 32\-row block\.*per two 32\-row blocks:*load\.ugm\.d32x3 r21, \[p\+rank/32\]// 3 sign words per lane, returned in r21\-\-r23*per 32\-row block, i\.e\. eight groups:*load\_block2d\.ugm bmp, \[\.\.\.\]// one contiguous bitmap word for each of 16 output columnsbfn w0, r21, r22, sel// mux: select the gathered word containing the window’s low bitsbfn w1, r22, r23, sel// mux: the following word, for bits that cross the boundaryand\.eq f0, b, rank, 31//b=rankmod32b=rank\\bmod 32shr lo, w0, b// align low bits at current bit rankshl hi, w1, \(\-rank\)&31// complementary count\(32−b\)mod32\(32\-b\)\\bmod 32\(\!f0\) or signs, lo, hi// atb=0b=0the high word contributes nothingcbit nsg, bmp// \# of nonzeros in 32\-row bitmap blockadd rank, rank, nsg// advance the block\-level sign rank*per four\-row group:*shr nib, bmp, imm\_g// extract group’s 4 presence bitsand r68, nib, idxmask// presence nibble in LUT\-key fieldcbit r68, r68// \# of sign bits consumed by this groupshl r66, signs, 3// shift sign window in key fieldand r66, r66, signmask// retain next 4 compact sign bitsbfn r67, nib, idxmask, r66// merge: presence and sign fields into one byte offsetadd r67, r67, lut\_base// add runtime SLM base of 2KB LUTload\.slm\.d32x2 r74, \[r67\]// 16 lookups, each lane receives 4 fp16 ternary codesshr signs, signs, r68// consume only the sign bits used by present rowsmul r54, scale, r74// multiply ternary codes with scales*once per four groups:*dpas\.8x1 acc, acc, r54, act// fp16 XMX contractionFig\. 6:Xe2 instruction sequence to unpack the BITCOS format\. Register names are shortened and independent operations are grouped by function rather than scheduler order\. Address setup, predicate formation, and register repacking are omitted\. The 16 SIMD lanes correspond to different output columns, so one four\-row group produces4×164\\times 16weights\. Section[II\-D](https://arxiv.org/html/2609.16338#S2.SS4)walks through the sequence\.This alternative contraction algorithm has two implications\. First, arithmetic moves from fp16 to integer: activations are quantized toint8per group of 128, the dot product accumulates inint32, and the group scales are applied once per group instead of being blended into the weights\. The weights are therefore reconstructed as the values±1\\pm 1rather than±s\\pm s, and the layout is stored in VNNI4 order so that eachvpdpbssd\(AVX\-VNNI\-INT8 compute\) consumes four consecutive reduction elements for each of eight rows\. Second, and more consequentially, the absence of mask registers means the presence pattern must be materialized as one*byte*per lane before it can select anything\. That expansion consists of a broadcast, an in\-lane shuffle and a compare against the bit\-select constant, and it costs 5 instructions where AVX\-512 spends merely 1 mask move instructionkmovd\. Once those byte masks exist, the last two instructions implement Eq\. \([2](https://arxiv.org/html/2609.16338#S2.E2)\): maskingNiN\_\{i\}with0xFEformsTi=−2niT\_\{i\}=\-2n\_\{i\}, and thevpsubb P,T,WcomputesWi=Ti−Pi=pi−2niW\_\{i\}=T\_\{i\}\-P\_\{i\}=p\_\{i\}\-2n\_\{i\}\. Thus a positive present lane becomes\+1\+1, a negative present lane becomes−1\-1, and an absent lane remains zero\.
### II\-DBITCOS unpack sequence for Intel Xe2 GPUs
Both x86 sequences rely onpdepto scatter compact signs back to the rows marked present\. Xe2 has no corresponding instruction, so the GPU kernel replaces that scatter with a small lookup table in shared local memory \(SLM\)\. We implement the kernel with the XeTLA templates\[[11](https://arxiv.org/html/2609.16338#bib.bib20)\], and use the conventional SYCL terminology of workgroups and subgroups throughout\[[12](https://arxiv.org/html/2609.16338#bib.bib21)\]\. Figure[6](https://arxiv.org/html/2609.16338#S2.F6)illustrates the assembly sequence for the Xe2 kernel\.
In the assembly sequence of Figure[6](https://arxiv.org/html/2609.16338#S2.F6), oneload\.ugm\.d32x3fetches the three consecutive sign words sufficient for two 32\-row blocks\. For each block, the contiguous bitmap loadload\_block2d\.ugm bmpsupplies one 32\-bit presence word per output column\. Twobfninstructions, the three\-input bitwise Boolean operation of Xe2, select the adjacent low/high sign words, and anand, two shifts and anoralign them at the current bit rank\. Theandreduces the rank to the in\-word offsetbb, theshraligns the low word bybband theshlbrings in thebbhigh bits that cross the word boundary\. That last shift needs a count of32−b32\-b, so the complementary count is taken modulo 32 and theoris predicated off in the one case the wrap gets wrong,b=0b=0, where the high word contributes nothing\. The block\-widecbitadvances that rank by the number of nonzeros in all 32 rows\. Each block is then decoded as eight four\-row groups\. A group extracts one presence nibble \(4 bits\), combines it with the next 4 bits of the aligned sign window, and uses the resulting byte offset for one SIMD16load\.slm\.d32x2\. The lookup returns four fp16 ternary codes per lane\. The group\-levelcbitadvances the sign window by only the bits actually consumed, an fp16 multiply applies the group scale, and one DPAS contraction is issued after four groups have supplied 16 reduction rows\. The sequence touches two memories,load\.ugmfor the presence bitmap and sign vector in global memory andload\.slmfor the table in SLM\.
The lookup table in SLM is indexed by an 8\-bit key formed from two nibbles: the four presence bitsmmof the bitmap for rows4g…4g\+34g\\ldots 4g\{\+\}3, and the next four bitsssof that column’s compact sign stream\. Entry\(m,s\)\(m,s\)holds four fp16 constantsc0…c3c\_\{0\}\\ldots c\_\{3\}, one per row, where
ci=\{0ifmi=0,\+1ifmi=1andsri=0,−1ifmi=1andsri=1,ri=∑j<imj,c\_\{i\}=\\begin\{cases\}0&\\text\{if \}m\_\{i\}=0,\\\\ \+1&\\text\{if \}m\_\{i\}=1\\text\{ and \}s\_\{r\_\{i\}\}=0,\\\\ \-1&\\text\{if \}m\_\{i\}=1\\text\{ and \}s\_\{r\_\{i\}\}=1,\\end\{cases\}\\qquad r\_\{i\}=\\textstyle\\sum\_\{j<i\}m\_\{j\},\(4\)andrir\_\{i\}is the rank of rowiiamong the present rows of the nibble\. The table is a precomputed, four\-bit\-widepdepcomposed with the map from sign bit to ternary code; Figure[7](https://arxiv.org/html/2609.16338#S2.F7)lists representative entries\.
Fig\. 7:Representative entries of the256256\-entry lookup table\. The key is the presence nibblemmtogether with a four\-bit windowssof the compact sign stream, and the entry holds four fp16 constants, shown with their bit patterns\. The second and third rows are the first two groups worked through in Figure[9](https://arxiv.org/html/2609.16338#S2.F9), with keys𝟶𝚡𝙱𝟼\\mathtt\{0xB6\}and𝟶𝚡𝟼𝟼\\mathtt\{0x66\}\.Fig\. 8:Xe2 kernel mapping of a weight tile onto SIMD lanes\. \(a\) Each grid column is one of the 16 output columns and is held by one SIMD lane, and each grid row is one register row\. VNNI2 order puts akk\-pair in every dword\. The blue\-shaded band is the output of a singleload\.slm\.d32x2, which returns two dwords per lane and therefore fills 2 register rows, i\.e\. 4 reduction rows, in one lookup\. \(b\) The presence bitmap and the compact sign vector in memory, for the 16 columns one subgroup owns: the bitmap words of adjacent columns are adjacent, whereas each column enters the sign vector at its own offset\.Fig\. 9:Three consecutive four\-row lookups on one column\. Each group forms an 8\-bit key from its presence nibble and a four\-bit window of the compact sign stream\. The bracket under each window gives the resulting table index/key\. The sign\-stream consumption is disjoint, but because the cursor advances bypopcount\(m\)\\operatorname\{popcount\}\(m\), the read windows overlap\. The figure follows a single column, that is one SIMD lane\. The kernel runs 16 of these in parallel, each with its own presence word, its own entry point into the sign stream and its own cursor\.Because a column’s sign bits are compacted, reading them requires tracking a per\-column read position, which we call that column’s*cursor*: the number of sign bits the groups above it have already consumed, equivalently the rank of its next non\-zero row among the non\-zeros seen so far\. The cursor is an offset into a variable\-rate stream, and it is the only piece of per\-column state whose value the bitmap alone does not give away\. The sign field of the key always takes four bits, because four is the most a four\-row group can need, but onlypopcount\(m\)\\operatorname\{popcount\}\(m\)of them are consumed; that population count is a singlecbitin Figure[6](https://arxiv.org/html/2609.16338#S2.F6)\. The cursor therefore advances bypopcount\(m\)\\operatorname\{popcount\}\(m\)and the next group re\-reads whatever this one left behind, so successive read windows*overlap*even though the bits they consume are disjoint, as Figure[9](https://arxiv.org/html/2609.16338#S2.F9)draws for three consecutive groups\. This is why the key is formed by a shift and a mask of a running window rather than by indexing the stream\. The kernel keeps the window in a register and shifts it right bypopcount\(m\)\\operatorname\{popcount\}\(m\)after each group \(shr signsof Figure[6](https://arxiv.org/html/2609.16338#S2.F6)\)\. The gather address is maintained more coarsely, where the kernel takes onecbitof the entire 32\-bit presence word per block and adds that to the column rank\. The two population countscbitin Figure[6](https://arxiv.org/html/2609.16338#S2.F6)therefore serve different purposes and are not redundant, i\.e\. the per\-group one only drives the window shift, and the per\-block one only drives the gather address\. Keeping them apart is what keeps the loop\-carried chain short, since the address depends on one count per thirty\-two rows rather than on a chain of eight\. The per\-group count also depends only on the bitmap, so it can issue before the SLM lookup returns\. Becausemmandssare four bits each, the table has28=2562^\{8\}=256entries of four fp16 values, so its size is22KB in total\. At kernel entry, the workgroup’s subgroups cooperatively initialize disjoint table entries in one SLM\-resident LUT\. The group size of four rows is not arbitrary\. Four presence bits and four sign bits give a table small enough to sit in SLM, whereas an eight\-row group would need2162^\{16\}entries\. Four is also compatible with the XMX/DPAS operand layout\. The fp16 DPAS consumes the weight tensor in VNNI2 order, with reduction rows2m2mand2m\+12m\{\+\}1packed into one dword, so four*consecutive*rows are precisely two VNNI2 dwords\. The entry is therefore stored in the order the tile needs, a singled32x2lookup returns both dwords, theload\.slm\.d32x2of Figure[6](https://arxiv.org/html/2609.16338#S2.F6), and they are written into the unpacked tile with no shuffle or transpose, thus meeting the VNNI2 requirement without any extra instructions\.
Figure[8](https://arxiv.org/html/2609.16338#S2.F8)\(a\) makes the mapping concrete\. A subgroup builds a tile of sixteen weight columns bysg\_kreduction rows, and because a dword holds two consecutivekkof one column, one SIMD16 register is exactly onekk\-pair across all sixteen columns\. One lookup therefore fills two adjacent register rows, that is four reduction rows of sixteen columns, or6464weights per message\. The sixteen lanes are sixteen*output columns*, so the sequence is vectorized alongnnwhilekkis walked serially by the loop\. A column’s cursor depends on the population counts of all its preceding groups, so neighboring lanes drift apart as they advance\. Figure[8](https://arxiv.org/html/2609.16338#S2.F8)\(b\) shows the presence bitmap and the compact sign vector in memory\. The presence bitmap is stored as⌈K/32⌉×N\\lceil K/32\\rceil\\times Nwords, each folding3232reduction rows of one column, and the words of adjacent columns are themselves adjacent, so sixteen lanes read them with a single block load, theload\_block2d\.ugmof Figure[6](https://arxiv.org/html/2609.16338#S2.F6)\. The sixteen sign cursors, by contrast, are unrelated addresses, so the sign words must be gathered per lane \(the gatherload\.ugm\.d32x3\)\. One 32\-row window needs two adjacent words, and the next block starts at most one word later, so the union of both windows is exactly three words and a singled32x3serves two blocks rather than one\. A fourth word is never consumed: the cursor can sit at most3131bits into a word and two blocks consume at most3232sign bits each, so the span reaches bit31\+64−1=9431\+64\-1=94at worst, still inside the third word\. In group00of Figure[9](https://arxiv.org/html/2609.16338#S2.F9)the bitmap nibblem=10112m=1011\_\{2\}marks rows00,11and33as non\-zero and row22as zero, and the next four compact sign bits ares=01102s=0110\_\{2\}, so the key is table index𝟶𝚡𝙱𝟼\\mathtt\{0xB6\}\. Row00takes signs0=0s\_\{0\}=0and row11takess1=1s\_\{1\}=1, but row22is absent and consumes nothing, so row33takess2=1s\_\{2\}=1rather thans3s\_\{3\}\. The entry is therefore\(\+1,−1,0,−1\)\(\+1,\-1,0,\-1\), and the cursor advances bypopcount\(m\)=3\\operatorname\{popcount\}\(m\)=3, and the following groups advance it by22and33\.
## IIIExperimental results
### III\-AExperimental platforms
We use three x86 CPU platforms with different core counts, core types and memory bandwidth:
- •One socket of an Intel Xeon Platinum 8592\+ CPU \(referred to as EMR\) with6464cores\. It has DDR5@4400 MT/s memory and a measured streaming read bandwidth of∼245\\sim 245GB/s\. It supports Advanced Matrix Extensions \(AMX\) and AVX\-512, including AVX\-512\-FP16\.
- •An Intel Core Ultra 9 285K CPU \(referred to as ARL\) with2424cores \(88performance cores and1616efficiency cores\)\. It has dual\-channel DDR5 memory and a measured read bandwidth of∼98\\sim 98GB/s\. It supports AVX\-VNNI\-INT8 but not AVX\-512\.
- •An Intel Core Ultra 7 258V CPU \(referred to as LNL CPU\) with88cores \(44performance and44efficiency cores\)\. It has3232GB of LPDDR5X memory and a measured read bandwidth of∼108\\sim 108GB/s\. It supports AVX\-VNNI\-INT8 but not AVX\-512\.
The instruction sets are relevant to the results: EMR runs the AVX\-512 kernel of Section[II\-B](https://arxiv.org/html/2609.16338#S2.SS2), and ARL and LNL run the AVX2 kernel of Section[II\-C](https://arxiv.org/html/2609.16338#S2.SS3)\. For the GPU evaluation we use two Xe2 GPU platforms:
- •The Intel Arc 140V, which is the integrated GPU of the LNL platform above\. It has88Xe2 cores, and it uses the same LPDDR5X memory as the LNL CPU, and its measured read bandwidth is∼108\\sim 108GB/s\.
- •An Intel Arc Pro B70 discrete GPU\. It has3232Xe2 cores and3232GB of dedicated GDDR6 memory, with a measured read bandwidth of∼500\\sim 500GB/s\.
### III\-BResults on the CPU platforms
We first introduce a roofline model to assess the efficacy and limitations of our CPU kernels, and then present the GEMV microbenchmarks and the end\-to\-end decode results\.
#### III\-B1A roofline model for the BITCOS CPU kernels
Fig\. 10:BITCOS CPU roofline at zero densityz=0\.40z=0\.40\. Each line isebw=min\(β,B/γ\)e\_\{\\mathrm\{bw\}\}=\\min\(\\beta,B/\\gamma\)for one kernel and core type; the flat part is the instruction ceilingB/γB/\\gammaand the diagonal is the memory bandwidth limit\. Markers place the five measured core groups\. Emerald Rapids and the Arrow Lake performance cores sit on the diagonal \(thus the kernels are bandwidth bound\), the Arrow Lake efficiency cores sit at the knee, and Lunar Lake sits on the flat part \(thus it is instruction bound\)\.Fig\. 11:Zero\-density sweep for a32k×16k32k\\times 16kmatrix\-vector multiplication on: \(a\) Emerald Rapids, \(b\) Arrow Lake and \(c\) Lunar Lake, against the flat LIBXSMM 2\-bit reference\. The shaded band marksz∈\[0\.297,0\.515\]z\\in\[0\.297,0\.515\], the range spanned by the deployed checkpoints of Table[I](https://arxiv.org/html/2609.16338#S1.T1)\. Top: GEMV time\. Bottom: effective bandwidth\.TABLE II:BITCOS CPU roofline at zero densityz=0\.40z=0\.40\.γ\\gammais the measured cost in cycles of one L1\-resident microkernel iteration andβ\\betais the per\-core share of the measured read bandwidth\. The bounding/bottleneck term is set in bold\.In this section we present a simple two\-term bottleneck roofline model\[[13](https://arxiv.org/html/2609.16338#bib.bib22)\], similar to the one used for the fixed\-width ternary kernels in prior work\[[9](https://arxiv.org/html/2609.16338#bib.bib17)\]\. For both the AVX\-512 and the AVX2 CPU microkernels \(i\.e\. see Figures[3](https://arxiv.org/html/2609.16338#S2.F3)and[5](https://arxiv.org/html/2609.16338#S2.F5)\), one iteration of the innermost loop upconverts3232ternary weight values\. Letγ\\gammabe the number of cycles that iteration costs when every operand is already in L1, andβ\\betathe share of read bandwidth available to one core, in bytes per cycle\. For these 32 weight entries, an iteration reads \(in bytes\):
B\(z\)=4⏟bitmap\+4\(1−z\)⏟signs\+0\.5⏟fp16 scale=8\.5−4zB\(z\)=\\underbrace\{4\}\_\{\\text\{bitmap\}\}\+\\underbrace\{4\(1\-z\)\}\_\{\\text\{signs\}\}\+\\underbrace\{0\.5\}\_\{\\text\{fp16 scale\}\}=8\.5\-4z\(5\)so the timeTTper iteration and the bandwidthebwe\_\{\\mathrm\{bw\}\}a core can sustain are:
T=max\(B\(z\)β,γ\),ebw=min\(β,B\(z\)γ\)\.T=\\max\\\!\\left\(\\frac\{B\(z\)\}\{\\beta\},\\,\\gamma\\right\),\\qquad e\_\{\\mathrm\{bw\}\}=\\min\\\!\\left\(\\beta,\\,\\frac\{B\(z\)\}\{\\gamma\}\\right\)\.\(6\)
The BITCOS\-based kernel is memory\-bound whileB\(z\)/γ\>βB\(z\)/\\gamma\>\\betaand instruction\-bound otherwise\. The fixed\-width 2\-bit kernels have a constantBBwhile a BITCOS\-based kernel does not, so its knee moves with the density of the model at hand sinceB\(z\)B\(z\)depends onzz\. For the per core bandwidthβ\\beta, we take the per\-core share of the streaming read bandwidth of the corresponding platform\. We measureγ\\gammaempirically on each platform by running the microkernel loop over an L1\-resident block\. Table[II](https://arxiv.org/html/2609.16338#S3.T2)illustrates the measuredγ\\gammavalues for the various core types, and we also report the termB/γB/\\gammafor a zero densityz=0\.4z=0\.4which impliesB=6\.9B=6\.9bytes read per 32 weight entries\. Arrow Lake and Lunar Lake have the same performance and efficiency cores and run the same AVX2 microkernel, so a single pair of measurements forγ\\gammaon performance and efficiency cores serves both platforms\. Table[II](https://arxiv.org/html/2609.16338#S3.T2)also evaluates Equaå \([6](https://arxiv.org/html/2609.16338#S3.E6)\) for the three CPU platforms, and Figure[10](https://arxiv.org/html/2609.16338#S3.F10)depicts the corresponding roofline\. Emerald Rapids is memory\-bound on all of its cores and Arrow Lake on its performance cores, so the BITCOS\-based kernel converts its smaller payload into savings in execution time; the Arrow Lake efficiency cores sit at the knee, where the two terms are within4%4\\%of each other\. Lunar Lake is instruction\-bound on both core types: with only eight cores sharing108108GB/s, each core has3\.23\.2–3\.73\.7bytes per cycle available but can only consume0\.870\.87–1\.381\.38bytes per cycle, so roughly three quarters of the bandwidth the platform offers a core is not attainable for this kernel\. A bandwidth\-rich client platform with a limited number of cores is exactly the case where a cheaper decode \(like the 2\-bit kernels from prior work\[[9](https://arxiv.org/html/2609.16338#bib.bib17)\]\) beats kernels with smaller payload and more expensive decode \(like the BITCOS\-based kernel of this work\)\.
#### III\-B2GEMV microbenchmarks on Emerald Rapids
To test the efficacy of the BITCOS\-based GEMV microkernel we experimented with a large 32768×\\times16384 ternary weight matrix\. Weights are replicated to a working set of at least44GB so that nothing is served from the last\-level cache\. We use group size 128, i\.e\. 128 entries along the inner\-product dimension share one 16\-bit scale, and we vary the zero densityzz\. The 2\-bit reference GEMV is the production LIBXSMM 2\-bit microkernel for CPUs\[[9](https://arxiv.org/html/2609.16338#bib.bib17)\]\. In Figure[11](https://arxiv.org/html/2609.16338#S3.F11)\(a\) top panel we illustrate the execution time of the GEMV on EMR, whereas on the bottom panel we show the corresponding effective bandwidth\. We observe that the BITCOS format wins at every density, and converts most of its bit\-width advantage into execution time savings \(see Figure[11](https://arxiv.org/html/2609.16338#S3.F11)\(a\) bottom panel, where the effective bandwidth stays constant∼225\\sim 225GB/s forzzup to0\.60\.6\)\. This behavior is also validated by our roofline model, where on EMR the BITCOS kernel operates in a bandwidth\-bound regime forz<0\.6z<0\.6\. In these plots we highlight with a green area the zero densities of interest: the deployed ternary models/checkpoints we examined in Table[I](https://arxiv.org/html/2609.16338#S1.T1)exhibitz∈\[0\.297,0\.515\]z\\in\[0\.297,0\.515\]\. For these zero densities, the observed speedup of the BITCOS\-based GEMV over the 2\-bit SOTA GEMV is in the range of1\.141\.14–1\.28×1\.28\\times\. For extreme zero densities \(e\.g\.z=0\.95z=0\.95\) we observe that the per\-iteration byte count drops, yieldingB/γ=1\.02B/\\gamma=1\.02, and the unpack instruction sequence starts being the bottleneck, thus restricting the effective bandwidth to190\.6190\.6GB/s\.
#### III\-B3GEMV microbenchmarks on Arrow Lake
We repeat the same GEMV benchmark on Arrow Lake \(see Figure[11](https://arxiv.org/html/2609.16338#S3.F11)\(b\)\) and the conclusions are the same as the ones on EMR: the BITCOS\-based GEMV wins at every density, and converts most of its bit\-width advantage into execution time savings, which is in alignment with our roofline analysis\. Over the same band of deployed zero densities,z∈\[0\.297,0\.515\]z\\in\[0\.297,0\.515\], the observed speedup of the BITCOS\-based GEMV over the 2\-bit SOTA GEMV is in the range of1\.131\.13–1\.27×1\.27\\times, and the kernel holds93\.493\.4–93\.693\.6GB/s of effective bandwidth across that band\.
Fig\. 12:Decode throughput in tokens per second on the three CPU platforms, batch one, over 7 ternary LLMs, with each model’s measured zero densityzzunder its name\. The orange and green bars correspond to the Prism ML fork ofllama\.cppand appear only for the Bonsai family of models that the fork supports\. The purple bar corresponds to the SOTA LIBXSMM 2\-bit kernel\[[9](https://arxiv.org/html/2609.16338#bib.bib17)\]and the blue bar is BITCOS \(this work\), both inside the vLLM CPU backend\. The green number above each cluster of bars is the speedup of BITCOS over the SOTA 2\-bit kernel\.
#### III\-B4GEMV microbenchmarks on Lunar Lake
Figure[11](https://arxiv.org/html/2609.16338#S3.F11)\(c\) illustrates the GEMV benchmark on Lunar Lake CPU, where our roofline model predicts that the BITCOS\-based GEMV kernel is instruction\-bound on both core types, and as such it is expected to be slower than the SOTA 2\-bit GEMV\. The measurements in Figure[11](https://arxiv.org/html/2609.16338#S3.F11)\(c\) bottom panel confirm that prediction: the BITCOS kernel sustains only28\.928\.9GB/s of effective bandwidth atz=0\.40z=0\.40\. The BITCOS kernel is slower than the 2\-bit reference at every density, and its effective bandwidth never exceeds33\.333\.3GB/s against the74\.774\.7GB/s the two\-bit kernel sustains\. This result confirms our roofline analysis, and it is a cautionary tale for client platforms with a limited number of cores and high memory bandwidth per core: a 2\-bit, cheaper decode GEMV kernel beats kernels with smaller payload and more expensive decode\.
#### III\-B5End\-to\-end decode on the CPU platforms
Fig\. 13:Zero\-density sweep for a32k×16k32k\\times 16kmatrix\-vector multiplication on \(a\) the Arc 140V and \(b\) the Arc Pro B70, against the flat XeTLA int2 reference, with every point tuned independently\. The shaded band marksz∈\[0\.297,0\.515\]z\\in\[0\.297,0\.515\], the range spanned by the deployed checkpoints of Table[I](https://arxiv.org/html/2609.16338#S1.T1)\. Top: GEMV time\. Bottom: effective bandwidth\.Fig\. 14:Decode throughput in tokens per second on the two Xe2 platforms, batch one, over 7 ternary LLMs, with each model’s measured zero densityzzunder its name\. The orange bar corresponds to the Prism ML fork ofllama\.cppon its Vulkan backend and appears only for the Bonsai family of models that the fork supports\. The purple bar corresponds to the SOTA XeTLA int2 kernel\[[9](https://arxiv.org/html/2609.16338#bib.bib17)\]and the blue bar is BITCOS \(this work\), both inside the vLLM XPU backend\. The green number above each cluster of bars is the speedup of BITCOS over the SOTA int2 kernel\.We integrated both the 2\-bit LIBXSMM kernel\[[9](https://arxiv.org/html/2609.16338#bib.bib17)\]and the BITCOS GEMV kernels into the vLLM CPU backend\[[14](https://arxiv.org/html/2609.16338#bib.bib18)\]and measured decode throughput on all three CPU platforms of Section[III\-A](https://arxiv.org/html/2609.16338#S3.SS1)over 7 of the group\-scaled LLM models of Table[I](https://arxiv.org/html/2609.16338#S1.T1)\. We also benchmarked the same 7 models with the Prism ML fork ofllama\.cpp111[https://github\.com/PrismML\-Eng/llama\.cpp](https://github.com/PrismML-Eng/llama.cpp)\. Two formats in that build are relevant here:Q2\_0, the fork’s 2\-bit code with one fp16 scale per128128weights, which is the same bit rate and group size as the SOTA 2\-bit packing of prior work\[[9](https://arxiv.org/html/2609.16338#bib.bib17)\], and upstream’sTQ1\_0, a five\-trit per byte format\.
Figure[12](https://arxiv.org/html/2609.16338#S3.F12)shows the results of the end\-to\-end inference on the three CPU platforms\. Each bar corresponds to the achieved decode throughput in tokens per second, for a single request of256256output tokens, with each inference engine at its own best thread count\. The green number above each cluster of bars is the speedup of BITCOS over the SOTA 2\-bit kernel, and we conclude that the two memory\-bound platforms \(EMR and ARL\) benefit from BITCOS on every model, while the instruction\-bound LNL CPU platform does not see any benefit, which is consistent with the roofline analysis in Section[III\-B1](https://arxiv.org/html/2609.16338#S3.SS2.SSS1)\. We also make the following observations regarding thellama\.cppbars\. First, the five\-trit per byte formatTQ1\_0is faster than the 2\-bitQ2\_0on the two client platforms, by1\.661\.66–1\.68×1\.68\\timeson Arrow Lake and1\.911\.91–2\.44×2\.44\\timeson Lunar Lake, whereas on the server socket the two converge within4%4\\%\. It is worth noting that the SOTA 2\-bit kernels of prior work\[[9](https://arxiv.org/html/2609.16338#bib.bib17)\]move more data than theTQ1\_0kernels ofllama\.cpp, and yet they outperform them by up to1\.78×1\.78\\times\. On the other hand, our BITCOS\-based kernel outperforms the SOTA 2\-bit kernel on the two memory\-bound platforms in the range of1\.101\.10–1\.18×1\.18\\timeson Emerald Rapids and1\.021\.02–1\.15×1\.15\\timeson Arrow Lake\. Compared to theTQ1\_0five\-trit per byte format \(which in principle is memory efficient\), BITCOS is1\.131\.13–1\.48×1\.48\\timesfaster on Emerald Rapids and1\.461\.46–1\.74×1\.74\\timesfaster on Arrow Lake\. On the instruction\-bound LNL CPU platform the SOTA 2\-bit work delivers the best end\-to\-end results and BITCOS loses on every model as predicted by the roofline model and the microbenchmarks of the previous section\.
### III\-CResults on the Xe2 GPU platforms
We first present the GEMV microbenchmarks on each GPU platform \(integrated GPU Arc 140V and discrete Arc Pro B70\) and then the end\-to\-end decode results\.
#### III\-C1GEMV microbenchmarks on the Arc 140V
To test the efficacy of the BITCOS\-based Xe2 GEMV microkernel of Section[II\-D](https://arxiv.org/html/2609.16338#S2.SS4)we use the same large 32768×\\times16384 ternary weight matrix as on the CPUs, group size 128, and a varying zero densityzz\. Every BITCOS point is tuned independently over the candidate tiles and the int2 reference is tuned the same way\. In Figure[13](https://arxiv.org/html/2609.16338#S3.F13)\(a\) top panel we illustrate the execution time of the GEMV on the Arc 140V, whereas on the bottom panel we show the corresponding effective bandwidth\. We observe that the BITCOS format wins at every sampled density\. In these plots we highlight with a green area the zero densities of interest, i\.e\. thez∈\[0\.297,0\.515\]z\\in\[0\.297,0\.515\]that the deployed checkpoints of Table[I](https://arxiv.org/html/2609.16338#S1.T1)exhibit\. For these zero densities, the observed speedup of the BITCOS\-based GEMV over the int2 state\-of\-the\-art GEMV is in the range of1\.041\.04–1\.14×1\.14\\times, and the kernel delivers71\.771\.7–75\.675\.6GB/s of effective bandwidth across that band\. The effective bandwidth does not stay flat but falls steadily withzz, from87\.287\.2to65\.565\.5GB/s across the full sweep, because the payload shrinks while the decode work per weight does not\. This is why the realized speedup is smaller than the corresponding payload reduction\.
#### III\-C2GEMV microbenchmarks on the Arc Pro B70
We repeat the same GEMV benchmark on the discrete Arc Pro B70 \(see Figure[13](https://arxiv.org/html/2609.16338#S3.F13)\(b\)\) and the conclusions are largely the same as the ones on the Arc 140V: the BITCOS\-based GEMV is never slower than the int2 reference\. Over the same band of deployed zero densities,z∈\[0\.297,0\.515\]z\\in\[0\.297,0\.515\], the observed speedup is in the range of1\.011\.01–1\.12×1\.12\\times, and the kernel holds398\.7398\.7–421\.9421\.9GB/s of effective bandwidth across that band\. The effective bandwidth declines withzzfor the same reason as on the integrated GPU platform, from476\.6476\.6to342\.6342\.6GB/s\. For example, atz=0\.40z=0\.40the measured speedup of1\.06×1\.06\\times\(BITCOS vs 2\-bit kernel\) falls short of the1\.23×1\.23\\timesbyte ratio\.
#### III\-C3End\-to\-end decode on the GPU platforms
We integrated both the int2 XeTLA kernel\[[9](https://arxiv.org/html/2609.16338#bib.bib17)\]and the BITCOS GEMV kernels into the same vLLM XPU backend\[[14](https://arxiv.org/html/2609.16338#bib.bib18)\]and measured decode throughput on both Xe2 platforms of Section[III\-A](https://arxiv.org/html/2609.16338#S3.SS1)over the same 7 group\-scaled LLM models of Table[I](https://arxiv.org/html/2609.16338#S1.T1)\. We also benchmarked the same 7 models with the Prism ML fork ofllama\.cpp, this time on its Vulkan backend, i\.e\. the vendor\-neutral GPU path that is available inllama\.cpp\. Vulkan is the only backend of that fork which is available for the Xe2 GPUs, as the SYCL backend does not support the two relevant formats \(Q2\_0andTQ1\_0\)\. Of those two formats onlyQ2\_0, the 2\-bit code with one fp16 scale per128128weights, has a Vulkan kernel\. The five\-trit per byteTQ1\_0format does not have any supporting Xe2 GPU kernel\.
Figure[14](https://arxiv.org/html/2609.16338#S3.F14)shows the results of the end\-to\-end inference on the two Xe2 platforms\. Each bar corresponds to the achieved decode throughput in tokens per second, for a single request of256256output tokens\. The green number above each cluster of bars is the speedup of BITCOS over the SOTA int2 kernel, and we conclude that both Xe2 platforms benefit from BITCOS on every model, by1\.091\.09–1\.22×1\.22\\timeson the integrated Arc 140V and1\.021\.02–1\.27×1\.27\\timeson the discrete Arc Pro B70\. We also make the following observations regarding thellama\.cppbars\. At identical 2\-bit format, identical group size and identical checkpoint the two 2\-bit packing methods, \(Q2\_0and the XeTLA int2 packing\[[9](https://arxiv.org/html/2609.16338#bib.bib17)\]\) converge on the integrated GPU part within5%5\\%in the end\-to\-end inference results \(Figure[14](https://arxiv.org/html/2609.16338#S3.F14)\(a\)\)\. On the discrete GPU part \(Figure[14](https://arxiv.org/html/2609.16338#S3.F14)\(b\)\) the 2\-bit XeTLA kernel delivers1\.461\.46–2\.12×2\.12\\timesspeedup over theQ2\_0\-based inference\. Our BITCOS\-based inference outperforms theQ2\_0ofllama\.cppby1\.111\.11–1\.18×1\.18\\timeson the Arc 140V and by1\.611\.61–2\.30×2\.30\\timeson the Arc Pro B70 and pushes the envelope of ternary LLM inference on Xe2 GPUs\.
## IVRelated Work
### IV\-ATernary Quantization of LLMs
There are two main approaches to producing ultra\-low\-bit LLMs: quantization\-aware training \(QAT\) and post\-training quantization \(PTQ\)\. QAT applies the ternary\-weights constraint during pretraining or fine\-tuning\. Early work established that binary and ternary weights are viable for natural language generation\[[15](https://arxiv.org/html/2609.16338#bib.bib9)\], and the seminal BitNet work\[[1](https://arxiv.org/html/2609.16338#bib.bib1),[2](https://arxiv.org/html/2609.16338#bib.bib2)\]showed that ternary\{−1,0,\+1\}\\\{\-1,0,\+1\\\}weights keep the accuracy of full\-precision weights at scale\. A subsequent work, ParetoQ\[[5](https://arxiv.org/html/2609.16338#bib.bib3)\]showed that ternary and 2\-bit formats live on the accuracy\-size Pareto frontier, ahead of 1\-bit and 4\-bit formats, refining the earlierkk\-bit inference scaling laws that placed the optimum at 4 bits\[[16](https://arxiv.org/html/2609.16338#bib.bib10),[17](https://arxiv.org/html/2609.16338#bib.bib11)\]\. The Spectra/TriLM suite\[[6](https://arxiv.org/html/2609.16338#bib.bib4)\]pretrained ternary models ranging from 99M to 3\.9B parameters\. A more recent QAT work, Tequila\[[18](https://arxiv.org/html/2609.16338#bib.bib5)\], removes the trapping behavior that makes ternary QAT unstable by re\-activating deadzone\-trapped weights in the training process\. Finally, more recent ternary QAT work has produced SOTA accuracy ternary LLMs for various LLM architectures: Bonsai\[[3](https://arxiv.org/html/2609.16338#bib.bib6)\]models range from 1\.7B to 27B parameters and offer multi\-step reasoning, structured tool calls, vision tasks and agentic loops, while Maple\[[7](https://arxiv.org/html/2609.16338#bib.bib7)\]is a 20B SOTA Mixture\-of\-Experts \(MoE\) ternary LLM with 1B active parameters\. Such compact, high\-quality models are a key enabler of on\-device agentic systems\[[10](https://arxiv.org/html/2609.16338#bib.bib12)\]\. Post\-training quantization \(PTQ\) does not involve training steps and weight gradient updates\. Instead, PTQ quantizes a pre\-trained model and calibrates on a few sequences\. Compared to QAT, PTQ is generally faster and less data\-intensive, but may result in lower accuracy\. Early PTQ methods at 4 and 2 bits, such as AWQ\[[19](https://arxiv.org/html/2609.16338#bib.bib13)\]and QuIP\[[20](https://arxiv.org/html/2609.16338#bib.bib14)\], have not matched QAT at ultra\-low bit\-widths\. Recent advancements in PTQ, CAT\-Q\[[4](https://arxiv.org/html/2609.16338#bib.bib15)\]and TWLA\[[21](https://arxiv.org/html/2609.16338#bib.bib16)\], make ternary weights directly from a pre\-trained model and substantially narrow the accuracy gap compared to the QAT methods\. These QAT and PTQ methods are complementary to our work: QAT and PTQ methods provide the means to obtain high\-quality ternary LLM models, while our work yields memory\-efficient and performant kernels to serve such ternary LLM inference on CPU and GPU platforms\.
### IV\-BKernels and runtimes for ternary LLM inference
Bitnet\.cpp\[[22](https://arxiv.org/html/2609.16338#bib.bib24)\]is the reference runtime for ternary LLMs and it is the companion runtime of the seminal ternary LLM work\[[1](https://arxiv.org/html/2609.16338#bib.bib1)\]\. It is faster than stockllama\.cpp\[[23](https://arxiv.org/html/2609.16338#bib.bib23)\], but recent work showed that it does not deliver performance close to roofline\[[9](https://arxiv.org/html/2609.16338#bib.bib17)\], thus in this work we compared against the SOTA runtime\[[9](https://arxiv.org/html/2609.16338#bib.bib17)\]\. Earlier libraries targeted ternary and binary inference on edge devices by packing several low\-bit weights per word and exploiting bit\-serial or sign\-based arithmetic, e\.g\. TernGEMM\[[24](https://arxiv.org/html/2609.16338#bib.bib26)\]and TABv2\[[25](https://arxiv.org/html/2609.16338#bib.bib27)\]\. One related approach for ternary LLM kernels replaces multiplication with table lookup\[[26](https://arxiv.org/html/2609.16338#bib.bib29),[27](https://arxiv.org/html/2609.16338#bib.bib30),[28](https://arxiv.org/html/2609.16338#bib.bib28)\]\. T\-MAC\[[29](https://arxiv.org/html/2609.16338#bib.bib25)\]pre\-computes partial dot products, and uses the packed low\-bit codes as indices in a table lookup\. This approach is applicable to CPUs with low vector FMA throughput\. On GPUs, prior work mixes 2\-bit and 4\-bit groups within a weight matrix and overlaps dequantization with the contraction to contain the accuracy loss\[[30](https://arxiv.org/html/2609.16338#bib.bib31)\]\. The 2\-bit reference in this work is the LIBXSMM\[[31](https://arxiv.org/html/2609.16338#bib.bib19)\]/XeTLA\[[11](https://arxiv.org/html/2609.16338#bib.bib20)\]kernels\[[9](https://arxiv.org/html/2609.16338#bib.bib17)\], and it is the strongest published baseline for CPU and Xe2 GPU platforms to date\. Section[III](https://arxiv.org/html/2609.16338#S3)shows that the LIBXSMM/XeTLA kernels are equal to or faster than the bestllama\.cppconfigurations on all platforms\. Therefore, the reported gains of this work are measured against kernels that are roofline\-optimal: the advantage of our work over these SOTA kernels stems from the fact that we exploit the inherent zero density of ternary LLM weights\.
### IV\-CUse of sparsity in ternary LLM inference
Recent work\[[32](https://arxiv.org/html/2609.16338#bib.bib33)\]exploits sparsity in ternary LLMs by storing only the indices of the non\-zero weights in a Ternary CSC format, since the sign alone describes a non\-zero ternary weight, and replaces the multiplications with additions and subtractions of the activations, which in theory reduces arithmetic by4×4\\timesat50%50\\%sparsity\. This technique helps only in compute\-bound cases: prefill is compute\-bound and becomes1\.21\.2–2\.3×2\.3\\timesfaster than libTorch on Spectra TriLM 1\.1B, but decode at batch size of one yields bandwidth\-bound GEMV operations and the performance is even worse than the one of libTorch\. The ternary CSC kernels also need7575–88%88\\%sparsity to overtake dense cuBLAS at the layer level, well above the29\.729\.7–51\.5%51\.5\\%that most ternary checkpoints exhibit \(Table[I](https://arxiv.org/html/2609.16338#S1.T1)\)\. Finally, the proposed index format has a variable bit rate and reads the weights through irregular gathers, whereas our BITCOS work consists of a dense and positional bitmap, offering sequential accesses and the resulting format is amenable to vectorizable unpacking\.
Recent work makes the ternary sparsity*semi\-structured*, so that N:M kernels can exploit it\. Unlike the one\-shot unstructured pruning of dense LLMs\[[33](https://arxiv.org/html/2609.16338#bib.bib32)\], the zeros here are produced by the quantizer itself\. Sparse\-BitNet\[[34](https://arxiv.org/html/2609.16338#bib.bib34)\]observes that the42%42\\%zeros of a pretrained BitNet model are unstructured, and trains ternary quantization jointly with a dynamic N:M mask, reporting up to1\.30×1\.30\\timesend to end speedup\. Sherry\[[35](https://arxiv.org/html/2609.16338#bib.bib35)\]constrains every block of four weights to hold exactly one zero, which leaves4×23=324\\times 2^\{3\}=32distinct blocks that a five\-bit code stores exactly, i\.e\.1\.251\.25bits per weight\. Both these recent lines of work replace the distribution that the quantizer produces and therefore require training: Sparse\-BitNet must apply the mask over the full pretraining run, while Sherry fixes the zero density at25%25\\%and enforces a 3:4 pattern\. Enforcing semi\-structured sparsity also costs accuracy: Sparse\-BitNet reports that its 6:8 ternary models lose0\.170\.17–0\.320\.32perplexity and0\.80\.8to3\.83\.8points of downstream accuracy against their own dense ternary baselines\. BITCOS instead uses the unstructured zeros that existing SOTA ternary LLM checkpoints already contain\. It is a change of the storage layout only, it applies to an existing checkpoint, it needs neither retraining nor sparsity hardware, and it is bit\-exact, because it decodes the same ternary values as the existing/original ternary model\.
## VConclusion
We introduced BITCOS, a distribution\-adaptive ternary layout of a presence bitmap plus a compacted sign vector that costs2−z2\-zbits per weight at a zero densityzz\. Across 29 SOTA ternary LLM checkpoints the zero density ranges from29\.7%29\.7\\%to51\.5%51\.5\\%, so BITCOS stores 26 of them more compactly than the five\-trit packing and reaches1\.4851\.485bits per weight on the sparsest, while against the 2\-bit format in production it reduces weight traffic by1\.161\.16–1\.32×1\.32\\timeson all 29 models\. Over the zero densities these checkpoints exhibit, the BITCOS GEMV is1\.141\.14–1\.28×1\.28\\timesfaster than the state\-of\-the\-art 2\-bit kernel on a 64\-core server Emerald Rapids CPU,1\.131\.13–1\.27×1\.27\\timeson a 24\-core client Arrow Lake CPU,1\.041\.04–1\.14×1\.14\\timeson the integrated Intel Xe2 Arc 140V GPU \(Lunar Lake GPU\) and1\.011\.01–1\.12×1\.12\\timeson the Arc Pro B70 discrete Intel GPU\. End to end in vLLM over 7 ternary LLMs, decode throughput improves by1\.101\.10–1\.18×1\.18\\times,1\.021\.02–1\.15×1\.15\\times,1\.091\.09–1\.22×1\.22\\timesand1\.021\.02–1\.27×1\.27\\timeson the same four platforms\. As future work we plan to extend the BITCOS layout and its unpacking sequences to more CPU and GPU architectures\.
## References
- \[1\]S\. Ma, H\. Wang, L\. Ma, L\. Wang, W\. Wang, S\. Huang, L\. Dong, R\. Wang, J\. Xue, and F\. Wei\(2024\)The era of 1\-bit LLMs: all large language models are in 1\.58 bits\.Note:arXiv:2402\.17764External Links:[Document](https://dx.doi.org/10.48550/arXiv.2402.17764)Cited by:[§I](https://arxiv.org/html/2609.16338#S1.p1.1),[§I](https://arxiv.org/html/2609.16338#S1.p2.1),[§IV\-A](https://arxiv.org/html/2609.16338#S4.SS1.p1.1),[§IV\-B](https://arxiv.org/html/2609.16338#S4.SS2.p1.1)\.
- \[2\]S\. Ma, H\. Wang, S\. Huang, X\. Zhang, Y\. Hu, T\. Song, Y\. Xia, and F\. Wei\(2025\)BitNet b1\.58 2B4T technical report\.Note:arXiv:2504\.12285External Links:[Document](https://dx.doi.org/10.48550/arXiv.2504.12285)Cited by:[§I](https://arxiv.org/html/2609.16338#S1.p1.1),[§I](https://arxiv.org/html/2609.16338#S1.p2.1),[§IV\-A](https://arxiv.org/html/2609.16338#S4.SS1.p1.1)\.
- \[3\]Prism ML\(2026\)Ternary Bonsai models\.Note:[https://huggingface\.co/prism\-ml](https://huggingface.co/prism-ml)Cited by:[§I](https://arxiv.org/html/2609.16338#S1.p2.1),[§IV\-A](https://arxiv.org/html/2609.16338#S4.SS1.p1.1)\.
- \[4\]S\. Wang, C\. Li, Y\. Kang, J\. Fan, and A\. Yao\(2026\)CAT\-Q: cost\-efficient and accurate ternary quantization for LLMs\.InProc\. Int\. Conf\. on Machine Learning \(ICML\),Note:arXiv:2606\.26650Cited by:[§I](https://arxiv.org/html/2609.16338#S1.p2.1),[§IV\-A](https://arxiv.org/html/2609.16338#S4.SS1.p1.1)\.
- \[5\]Z\. Liu, C\. Zhao, H\. Huang, S\. Chen, J\. Zhang, J\. Zhao, S\. Roy, L\. Jin, Y\. Xiong,et al\.\(2025\)ParetoQ: improving scaling laws in extremely low\-bit LLM quantization\.Note:arXiv:2502\.02631Models available:[https://huggingface\.co/collections/facebook/mobilellm](https://huggingface.co/collections/facebook/mobilellm)External Links:[Document](https://dx.doi.org/10.48550/arXiv.2502.02631)Cited by:[§I](https://arxiv.org/html/2609.16338#S1.p2.1),[§IV\-A](https://arxiv.org/html/2609.16338#S4.SS1.p1.1)\.
- \[6\]A\. Kaushal, T\. Vaidhya, A\. K\. Mondal, T\. Pandey, A\. Bhagat, and I\. Rish\(2024\)Spectra: surprising effectiveness of pretraining ternary language models at scale\.Note:arXiv:2407\.12327Models available:[https://huggingface\.co/collections/SpectraSuite/trilms\-unpacked](https://huggingface.co/collections/SpectraSuite/trilms-unpacked)External Links:[Document](https://dx.doi.org/10.48550/arXiv.2407.12327)Cited by:[§I](https://arxiv.org/html/2609.16338#S1.p2.1),[§IV\-A](https://arxiv.org/html/2609.16338#S4.SS1.p1.1)\.
- \[7\]DeepGrove\(2026\)Maple: a 20B\-A1B ternary\-weight reasoning model\.Note:[https://huggingface\.co/deepgrove/maple\-preview](https://huggingface.co/deepgrove/maple-preview)Cited by:[§I](https://arxiv.org/html/2609.16338#S1.p2.1),[§IV\-A](https://arxiv.org/html/2609.16338#S4.SS1.p1.1)\.
- \[8\]OpenBMB\(2025\)BitCPM\-CANN: full\-pipeline ternary quantized models trained on CANN\.Note:[https://huggingface\.co/collections/openbmb/bitcpm\-cann](https://huggingface.co/collections/openbmb/bitcpm-cann)Cited by:[§I](https://arxiv.org/html/2609.16338#S1.p2.1)\.
- \[9\]E\. Georganas, D\. Kalamkar, A\. Heinecke, and P\. Dubey\(2026\)Pushing the envelope of LLM inference with ultra\-low\-bit quantized models\.Note:arXiv:2508\.06753v3External Links:[Document](https://dx.doi.org/10.48550/arXiv.2508.06753)Cited by:[§I](https://arxiv.org/html/2609.16338#S1.p3.1),[Fig\. 12](https://arxiv.org/html/2609.16338#S3.F12),[Fig\. 14](https://arxiv.org/html/2609.16338#S3.F14),[§III\-B1](https://arxiv.org/html/2609.16338#S3.SS2.SSS1.p1.1),[§III\-B1](https://arxiv.org/html/2609.16338#S3.SS2.SSS1.p2.1),[§III\-B2](https://arxiv.org/html/2609.16338#S3.SS2.SSS2.p1.1),[§III\-B5](https://arxiv.org/html/2609.16338#S3.SS2.SSS5.p1.1),[§III\-B5](https://arxiv.org/html/2609.16338#S3.SS2.SSS5.p2.1),[§III\-C3](https://arxiv.org/html/2609.16338#S3.SS3.SSS3.p1.1),[§III\-C3](https://arxiv.org/html/2609.16338#S3.SS3.SSS3.p2.1),[§IV\-B](https://arxiv.org/html/2609.16338#S4.SS2.p1.1)\.
- \[10\]P\. Belcak, G\. Heinrich, S\. Diao, Y\. Fu, X\. Dong, S\. Muralidharan, Y\. C\. Lin, and P\. Molchanov\(2025\)Small language models are the future of agentic AI\.Note:arXiv:2506\.02153External Links:[Document](https://dx.doi.org/10.48550/arXiv.2506.02153)Cited by:[§I](https://arxiv.org/html/2609.16338#S1.p3.1),[§IV\-A](https://arxiv.org/html/2609.16338#S4.SS1.p1.1)\.
- \[11\]IntelXeTLA: intel Xe templates for linear algebra\.Note:[https://github\.com/intel/xetla](https://github.com/intel/xetla)Cited by:[§II\-D](https://arxiv.org/html/2609.16338#S2.SS4.p1.1),[§IV\-B](https://arxiv.org/html/2609.16338#S4.SS2.p1.1)\.
- \[12\]R\. Keryell, R\. Reyes, and L\. Howes\(2015\)Khronos SYCL for OpenCL: a tutorial\.InProc\. 3rd Int\. Workshop on OpenCL,pp\. 1–1\.External Links:[Document](https://dx.doi.org/10.1145/2791321.2791345)Cited by:[§II\-D](https://arxiv.org/html/2609.16338#S2.SS4.p1.1)\.
- \[13\]S\. Williams, A\. Waterman, and D\. Patterson\(2009\)Roofline: an insightful visual performance model for multicore architectures\.Communications of the ACM52\(4\),pp\. 65–76\.External Links:[Document](https://dx.doi.org/10.1145/1498765.1498785)Cited by:[§III\-B1](https://arxiv.org/html/2609.16338#S3.SS2.SSS1.p1.1)\.
- \[14\]W\. Kwon, Z\. Li, S\. Zhuang, Y\. Sheng, L\. Zheng, C\. H\. Yu, J\. E\. Gonzalez, H\. Zhang, and I\. Stoica\(2023\)Efficient memory management for large language model serving with PagedAttention\.InProc\. 29th ACM Symp\. on Operating Systems Principles \(SOSP\),pp\. 611–626\.External Links:[Document](https://dx.doi.org/10.1145/3600006.3613165)Cited by:[§III\-B5](https://arxiv.org/html/2609.16338#S3.SS2.SSS5.p1.1),[§III\-C3](https://arxiv.org/html/2609.16338#S3.SS3.SSS3.p1.1)\.
- \[15\]Z\. Liu, B\. Oguz, A\. Pappu, Y\. Shi, and R\. Krishnamoorthi\(2023\)Binary and ternary natural language generation\.InProc\. 61st Annual Meeting of the Association for Computational Linguistics \(ACL\),pp\. 65–77\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.5)Cited by:[§IV\-A](https://arxiv.org/html/2609.16338#S4.SS1.p1.1)\.
- \[16\]T\. Dettmers and L\. Zettlemoyer\(2023\)The case for 4\-bit precision: k\-bit inference scaling laws\.InInternational Conference on Machine Learning,pp\. 7750–7774\.Cited by:[§IV\-A](https://arxiv.org/html/2609.16338#S4.SS1.p1.1)\.
- \[17\]T\. Kumar, Z\. Ankner, B\. F\. Spector, B\. Bordelon, N\. Muennighoff, M\. Paul, C\. Pehlevan, C\. Ré, and A\. Raghunathan\(2024\)Scaling laws for precision\.Note:arXiv:2411\.04330External Links:[Document](https://dx.doi.org/10.48550/arXiv.2411.04330)Cited by:[§IV\-A](https://arxiv.org/html/2609.16338#S4.SS1.p1.1)\.
- \[18\]H\. Huang, D\. Wu, R\. Cen, G\. Yu, Z\. Li, K\. Liu, J\. Zhu, P\. Chen, X\. Liu, and D\. Wu\(2025\)Tequila: trapping\-free ternary quantization for large language models\.Note:arXiv:2509\.23809External Links:[Document](https://dx.doi.org/10.48550/arXiv.2509.23809)Cited by:[§IV\-A](https://arxiv.org/html/2609.16338#S4.SS1.p1.1)\.
- \[19\]J\. Lin, J\. Tang, H\. Tang, S\. Yang, G\. Xiao, and S\. Han\(2025\)AWQ: activation\-aware weight quantization for on\-device LLM compression and acceleration\.GetMobile: Mobile Computing and Communications28\(4\),pp\. 12–17\.External Links:[Document](https://dx.doi.org/10.1145/3714983.3714987)Cited by:[§IV\-A](https://arxiv.org/html/2609.16338#S4.SS1.p1.1)\.
- \[20\]J\. Chee, Y\. Cai, V\. Kuleshov, and C\. D\. Sa\(2023\)QuIP: 2\-bit quantization of large language models with guarantees\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Vol\.36,pp\. 4396–4429\.Cited by:[§IV\-A](https://arxiv.org/html/2609.16338#S4.SS1.p1.1)\.
- \[21\]Z\. Zhao, Z\. Xu, Z\. Chen, X\. Hu, Z\. Jiang, and D\. Yang\(2026\)TWLA: achieving ternary weights and low\-bit activations for LLMs via post\-training quantization\.InProc\. Int\. Conf\. on Machine Learning \(ICML\),Note:arXiv:2606\.13054Cited by:[§IV\-A](https://arxiv.org/html/2609.16338#S4.SS1.p1.1)\.
- \[22\]J\. Wang, H\. Zhou, T\. Song, S\. Cao, Y\. Xia, T\. Cao, J\. Wei, S\. Ma, H\. Wang, and F\. Wei\(2025\)bitnet\.cpp: efficient edge inference for ternary LLMs\.InProc\. 63rd Annual Meeting of the Association for Computational Linguistics \(ACL\),pp\. 9305–9322\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.457)Cited by:[§IV\-B](https://arxiv.org/html/2609.16338#S4.SS2.p1.1)\.
- \[23\]llama\.cpp: LLM inference in C/C\+\+\.Note:[https://github\.com/ggml\-org/llama\.cpp](https://github.com/ggml-org/llama.cpp)Cited by:[§IV\-B](https://arxiv.org/html/2609.16338#S4.SS2.p1.1)\.
- \[24\]S\. Choi, K\. Shim, J\. Choi, W\. Sung, and B\. Shim\(2021\)TernGEMM: general matrix multiply library with ternary weights for fast DNN inference\.InProc\. IEEE Workshop on Signal Processing Systems \(SiPS\),pp\. 111–116\.External Links:[Document](https://dx.doi.org/10.1109/SiPS52927.2021.00028)Cited by:[§IV\-B](https://arxiv.org/html/2609.16338#S4.SS2.p1.1)\.
- \[25\]G\. Fu, O\. Fischer, S\. Zhu, and G\. Alonso\(2026\)TABv2: a faster ternary and binary neural network inference library on the edge\.IEEE Trans\. Very Large Scale Integr\. \(VLSI\) Syst\.,pp\. 1–13\.External Links:[Document](https://dx.doi.org/10.1109/TVLSI.2025.3650684)Cited by:[§IV\-B](https://arxiv.org/html/2609.16338#S4.SS2.p1.1)\.
- \[26\]D\. Blalock and J\. Guttag\(2021\)Multiplying matrices without multiplying\.InProc\. Int\. Conf\. on Machine Learning \(ICML\),pp\. 992–1004\.Cited by:[§IV\-B](https://arxiv.org/html/2609.16338#S4.SS2.p1.1)\.
- \[27\]X\. Tang, Y\. Wang, T\. Cao, L\. L\. Zhang, Q\. Chen, D\. Cai, Y\. Liu, and M\. Yang\(2023\)LUT\-NN: empower efficient neural network inference with centroid learning and table lookup\.InProc\. 29th Annual Int\. Conf\. on Mobile Computing and Networking \(MobiCom\),pp\. 1–15\.External Links:[Document](https://dx.doi.org/10.1145/3570361.3613285)Cited by:[§IV\-B](https://arxiv.org/html/2609.16338#S4.SS2.p1.1)\.
- \[28\]D\. C\. Ganji, S\. Ashfaq, E\. Saboori, S\. Sah, S\. Mitra, M\. AskariHemmat, A\. Hoffman, A\. Hassanien, and M\. Léonardon\(2023\)DeepGEMM: accelerated ultra low\-precision inference on CPU architectures using lookup tables\.InProc\. IEEE/CVF Conf\. on Computer Vision and Pattern Recognition Workshops \(CVPRW\),pp\. 4656–4664\.External Links:[Document](https://dx.doi.org/10.1109/CVPRW59228.2023.00491)Cited by:[§IV\-B](https://arxiv.org/html/2609.16338#S4.SS2.p1.1)\.
- \[29\]J\. Wei, S\. Cao, T\. Cao, L\. Ma, L\. Wang, Y\. Zhang, and M\. Yang\(2025\)T\-MAC: CPU renaissance via table lookup for low\-bit LLM deployment on edge\.InProc\. 20th European Conf\. on Computer Systems \(EuroSys\),pp\. 278–292\.External Links:[Document](https://dx.doi.org/10.1145/3689031.3696099)Cited by:[§IV\-B](https://arxiv.org/html/2609.16338#S4.SS2.p1.1)\.
- \[30\]J\. Li, J\. Xu, S\. Li, S\. Huang, J\. Liu, Y\. Lian, and G\. Dai\(2024\)Fast and efficient 2\-bit LLM inference on GPU: 2/4/16\-bit in a weight matrix with asynchronous dequantization\.InProc\. 43rd IEEE/ACM Int\. Conf\. on Computer\-Aided Design \(ICCAD\),pp\. 1–9\.External Links:[Document](https://dx.doi.org/10.1145/3676536.3676796)Cited by:[§IV\-B](https://arxiv.org/html/2609.16338#S4.SS2.p1.1)\.
- \[31\]LIBXSMM: library for specialized dense and sparse matrix operations, and deep learning primitives\.Note:[https://github\.com/libxsmm/libxsmm](https://github.com/libxsmm/libxsmm)Cited by:[§IV\-B](https://arxiv.org/html/2609.16338#S4.SS2.p1.1)\.
- \[32\]S\. Zhu, G\. Fu, M\. Kjoseva, and G\. Alonso\(2026\)Efficient addition\-based sparse GEMM for fast ternary large language model inference on edge devices\.ACM Trans\. Embedded Computing Systems25\(4\),pp\. 60:1–60:29\.External Links:[Document](https://dx.doi.org/10.1145/3807782)Cited by:[§IV\-C](https://arxiv.org/html/2609.16338#S4.SS3.p1.1)\.
- \[33\]E\. Frantar and D\. Alistarh\(2023\)SparseGPT: massive language models can be accurately pruned in one\-shot\.InProc\. Int\. Conf\. on Machine Learning \(ICML\),pp\. 10323–10337\.Cited by:[§IV\-C](https://arxiv.org/html/2609.16338#S4.SS3.p2.1)\.
- \[34\]D\. Zhang, X\. Wu, S\. Huang, Y\. Wang, H\. Shao, Y\. Hao, Z\. Chi, L\. Dong, T\. Song, Y\. Xia, Z\. Sui, and F\. Wei\(2026\)Sparse\-BitNet: 1\.58\-bit LLMs are naturally friendly to semi\-structured sparsity\.Note:arXiv:2603\.05168External Links:[Document](https://dx.doi.org/10.48550/arXiv.2603.05168)Cited by:[§IV\-C](https://arxiv.org/html/2609.16338#S4.SS3.p2.1)\.
- \[35\]H\. Huang, D\. Wu, Q\. Hu, G\. Yu, J\. Yang, J\. Zhu, X\. Liu, and D\. Wu\(2026\)Sherry: hardware\-efficient 1\.25\-bit ternary quantization via fine\-grained sparsification\.Note:arXiv:2601\.07892External Links:[Document](https://dx.doi.org/10.48550/arXiv.2601.07892)Cited by:[§IV\-C](https://arxiv.org/html/2609.16338#S4.SS3.p2.1)\.
Optimization Notice: Software and workloads used in performance tests may have been optimized for performance only on Intel microprocessors\. Performance tests, such as SYSmark and MobileMark, are measured using specific computer systems, components, software, operations and functions\. Any change to any of those factors may cause the results to vary\. You should consult other information and performance tests to assist you in fully evaluating your contemplated purchases, including the performance of that product when combined with other products\. For more information go to[http://www\.intel\.com/performance](http://www.intel.com/performance)\.
Intel, Xeon, and Intel Xeon Phi are trademarks of Intel Corporation in the U\.S\. and/or other countries\.Similar Articles
Bitnet.cpp: Efficient Edge Inference for Ternary LLMs
Bitnet.cpp presents a mixed-precision matrix multiplication library for efficient edge inference of ternary LLMs like BitNet b1.58, achieving up to 6.25x speedup over full-precision baselines. The system is open-sourced on GitHub.
Is ternary (1.58-bit) LLMs making a come back?
Recent ternary 1.58-bit LLM releases from small labs demonstrate speed and medical specialization but struggle with long-horizon tasks, with optimism for future models to compete with larger architectures like Qwen.
How to pack ternary numbers in 8-bit bytes
A blog post describing an efficient method to pack ternary numbers into 8-bit bytes using SIMD-friendly unpacking, achieving 1.6 bits per trit, with applications in LLM weight quantization like BitNet b1.58.
1-Bit LLM in the Browser
A 1-bit LLM (Bonsai) is now runnable in the browser via WebGPU, enabling efficient on-device inference.
CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit Weights
CubicQuant proposes a parametric non-uniform scalar quantization format for LLM weights, using a monotonic cubic curve to adapt reconstruction levels at 1-8 bit widths while retaining dense integer code streams for GPU efficiency. Experiments show RMSE reductions over uniform and floating-point baselines, with preliminary H200 kernel measurements.