ternary-llms

Tag

Cards List
#ternary-llms

Breaking the 1.58-bit Barrier for Ternary LLMs

arXiv cs.AI ↗ · 2026-09-16 Cached

This paper introduces BITCOS, a distribution-adaptive layout for storing ternary LLM weights more efficiently, achieving up to 1.28× speedup in matrix-vector multiplication and 1.27× in inference throughput on GPUs.

0 favorites 0 likes
#ternary-llms

Was BitNet a dead end? What happened to ternary LLMs?

Reddit r/LocalLLaMA ↗ · 2026-06-08

The article questions why ternary language models like BitNet have not scaled beyond 2B parameters, given their initial promise, and discusses the apparent lack of progress from open-weight AI labs.

0 favorites 0 likes
#ternary-llms

Bitnet.cpp: Efficient Edge Inference for Ternary LLMs

Papers with Code Trending ↗ · 2025-02-17 Cached

Bitnet.cpp presents a mixed-precision matrix multiplication library for efficient edge inference of ternary LLMs like BitNet b1.58, achieving up to 6.25x speedup over full-precision baselines. The system is open-sourced on GitHub.

0 favorites 0 likes
← Back to home

Submit Feedback