ternary

Tag

Cards List
#ternary

1-bit / 2-bit / Ternary / Bitnet Models - Updates & Tracking

Reddit r/LocalLLaMA · 2026-08-18

This article tracks updates on various low-bit AI models, including Bonsai, BitCPM, and others, with performance improvements and compatibility updates for backends like llama.cpp.

0 favorites 0 likes
#ternary

Show HN: Maple-Preview – ternary 20B MoE running at 120 tok/s on a iPhone

Hacker News Top · 2026-08-04

Maple-Preview is a ternary 20B MoE model that runs at 120 tokens per second on an iPhone, showcasing efficient on-device inference.

0 favorites 0 likes
#ternary

I hope ternary will eventually work but ... sigh

Reddit r/LocalLLaMA · 2026-07-24

User expresses disappointment with the ternary Bonsai model from Prisml, questioning its quality despite hype, and wonders if it's just typical model overhype.

0 favorites 0 likes
#ternary

you can now fine tune Prism-ML's ternary Bonsai models

Reddit r/LocalLLaMA · 2026-07-21

Prism-ML announces that their ternary Bonsai models can now be fine-tuned, with example code and a recommendation to use a high learning rate.

0 favorites 0 likes
#ternary

For those with 12GB GPUs, you can now run QWEN 3.6 27B wth little loss via the new Ternary version.

Reddit r/ArtificialInteligence · 2026-07-15

A new ternary quantized version of Qwen3.6 27B, called Bonsai 27B, allows running the model on 12GB GPUs with 10x less memory and 95% of original performance, making it accessible for local deployment.

0 favorites 0 likes
#ternary

Ternary Qwen3.6 27B Tested on 3090!

Reddit r/LocalLLaMA · 2026-07-14

User tests ternary quantized Qwen3.6 27B on an RTX 3090, achieving 60 tk/s with two slots and 100k KV cache using 21GB VRAM, with good quality and stable tool calls.

0 favorites 0 likes
#ternary

@TheAhmadOsman: HOLYYYY 27B model under 6GBs and 4GBs Local AI will be the default P.S. We are gonna get this optimized in ODS by @Osma…

X AI KOLs Timeline · 2026-07-14 Cached

Ternary Bonsai 27B, a large language model, is demonstrated running locally on an NVIDIA RTX 5090 GPU, requiring under 6GB of memory and enabling end-to-end agentic workflows on consumer hardware.

0 favorites 0 likes
#ternary

prism-ml/Ternary-Bonsai-27B-mlx-2bit

Hugging Face Models Trending · 2026-07-04 Cached

Prism ML releases Ternary-Bonsai-27B-mlx-2bit, a ternary-quantized 27B-parameter language model that achieves ~95% of FP16 performance while fitting in ~7.2 GB, enabling full reasoning on laptops.

0 favorites 0 likes
#ternary

@heyshrutimishra: Full-sized AI models now run on phones. That's BitCPM, a new open-source model from ModelBest, Tsinghua, and OpenBMB. T…

X AI KOLs Following · 2026-05-25 Cached

BitCPM is a new open-source model from ModelBest, Tsinghua, and OpenBMB that uses ternary weights (-1,0,1) to run full-sized AI models on phones.

0 favorites 0 likes
#ternary

PrismML-Eng/Bonsai-demo

GitHub Trending (daily) · 2026-07-16 Cached

PrismML releases Bonsai 27B, a vision-language model with agentic tool calling and long context, along with 1-bit and ternary variants. The demo repository allows running these models locally on various hardware.

0 favorites 0 likes
← Back to home

Submit Feedback