ternary-quantization

Tag

Cards List
#ternary-quantization

Open Source Ternary LLM Engine in Rust/CUDA for Quantization, Serving, and Training of models on consumer GPUs, called Tritium (Apache 2.0)

Reddit r/LocalLLaMA · 2026-07-31

Introduces Tritium, an open-source Rust/CUDA engine for ternary (1.58-bit) quantization, serving, and training of LLMs on consumer GPUs. It claims faster inference than llama.cpp for ternary models and introduces a new quantization method called SALT.

0 favorites 0 likes
#ternary-quantization

Prism-ML Bonsai Qwen 3.6 27B

Reddit r/LocalLLaMA · 2026-07-14 Cached

Prism ML released Ternary-Bonsai-27B, a ternary-quantized version of Qwen3.6-27B that retains 95% of FP16 intelligence at a ~7.2 GB footprint, enabling full 27B-class reasoning on laptops and single GPUs with speeds up to 26 tok/s on Apple M5 Pro.

0 favorites 0 likes
#ternary-quantization

clark-labs/clark-air-sana-1.6b-1.58bit · Hugging Face

Reddit r/LocalLLaMA · 2026-06-28 Cached

Clark Labs released Clark Air Sana 1.6B, a ternary-quantized version of the Sana 1.6B text-to-image transformer that is 8.6× smaller than FP16 while maintaining near-FP16 quality, enabling efficient deployment.

0 favorites 0 likes
#ternary-quantization

CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs

arXiv cs.CL · 2026-06-26 Cached

CAT-Q introduces a post-training ternary quantization method for LLMs that uses learnable modulation and softened ternarization, achieving superior performance over BitNet 1.58-bit while using only 512 calibration samples and scaling to 235B parameters.

0 favorites 0 likes
#ternary-quantization

@AdinaYakup: BitCPM4-CANN Native 1.58-bit LLM training system on Ascend NPUs https://huggingface.co/collections/openbmb/bitcpm4-cann…

X AI KOLs Following · 2026-05-22 Cached

OpenBMB releases BitCPM4-CANN, a collection of natively trained 1.58-bit ternary quantized LLMs (0.5B to 8B) optimized for Ascend NPUs via CANN, achieving 6× memory reduction at inference and minimal training overhead.

0 favorites 0 likes
#ternary-quantization

Tequila: Trapping-free Ternary Quantization for Large Language Models

Papers with Code Trending · 2025-09-28 Cached

This paper introduces Tequila, a trapping-free quantization method for Large Language Models that improves ternary quantization accuracy and inference speed by repurposing deadzone-trapped weights as dynamic biases.

0 favorites 0 likes
← Back to home

Submit Feedback