ternary-weights

Tag

Cards List
#ternary-weights

deepgrove/maple-preview

Hugging Face Models Trending · 2026-08-04 Cached

DeepGrove releases Maple-Preview, an open-source 20B-A1B ternary-weight reasoning LLM with SOTA reasoning for its weight class, capable of 200+ tokens/sec on a Mac mini M4 and competitive with larger models.

0 favorites 0 likes
#ternary-weights

The idea: on a CPU the decode speed depends on the active params per token, not the total. My objective is trying to run a 10B at 100tok/s on a mid level PC (No GPU).

Reddit r/LocalLLaMA · 2026-07-29

Presents the insight that CPU decode speed depends on active parameters per token, not total parameters, and proposes a 10B-parameter model using ternary weights and granular MoE to achieve high token rates on mid-level PCs. Reports sandbox measurements showing a speedup from 176 to 848 tok/s on an 8.3M model with minimal quality loss.

0 favorites 0 likes
#ternary-weights

Neutrino-1 8B

Hacker News Top · 2026-07-28 Cached

Neutrino-1 8B is an 8.19B-parameter transformer with a proprietary ternary weight format, enabling it to fit on 8GB GPUs and serve multiple platforms from a single 3.88GB artifact with high decode speed.

0 favorites 0 likes
#ternary-weights

prism-ml/bonsai-image-ternary-4B-gemlite-2bit

Hugging Face Models Trending · 2026-05-21 Cached

Prism ML releases Bonsai Image, a 1.21 GB text-to-image diffusion transformer using ternary weights (1.58-bit) for NVIDIA GPUs, offering 4.5s / 1024² on RTX 3080 and much smaller than FP16.

0 favorites 0 likes
#ternary-weights

Ternary Bonsai: Top Intelligence at 1.58 Bits

Hacker News Top · 2026-04-18

A highly efficient AI model architecture using ternary weights (-1, 0, 1) that achieves competitive performance while requiring only 1.58 bits per parameter, enabling deployment on extremely constrained devices.

0 favorites 0 likes
← Back to home

Submit Feedback