Show HN: Maple-Preview – ternary 20B MoE running at 120 tok/s on a iPhone

Hacker News Top Models

Summary

Maple-Preview is a ternary 20B MoE model that runs at 120 tokens per second on an iPhone, showcasing efficient on-device inference.

No content available
Original Article

Similar Articles

deepgrove/maple-preview

Hugging Face Models Trending

DeepGrove releases Maple-Preview, an open-source 20B-A1B ternary-weight reasoning LLM with SOTA reasoning for its weight class, capable of 200+ tokens/sec on a Mac mini M4 and competitive with larger models.

Edge0/Edge0-35B-A3B-preview

Hugging Face Models Trending

Edge0-35B-A3B-preview is a sparse MoE model that enables efficient AI inference on mobile devices by using streaming expert offloading and quantization, achieving 15 tok/s with under 3 GiB of memory.