Show HN: Maple-Preview – ternary 20B MoE running at 120 tok/s on a iPhone
Summary
Maple-Preview is a ternary 20B MoE model that runs at 120 tokens per second on an iPhone, showcasing efficient on-device inference.
Similar Articles
I've added Maple-Preview to Mference, got 40 tps generation with 500MB of used RAM on Air M4
The author adds Maple-Preview support to Mference, a tool that streams MoE experts from disk to run large models on low-RAM devices, achieving 40 tps with 500MB RAM on an Air M4.
deepgrove/maple-preview
DeepGrove releases Maple-Preview, an open-source 20B-A1B ternary-weight reasoning LLM with SOTA reasoning for its weight class, capable of 200+ tokens/sec on a Mac mini M4 and competitive with larger models.
Edge0/Edge0-35B-A3B-preview
Edge0-35B-A3B-preview is a sparse MoE model that enables efficient AI inference on mobile devices by using streaming expert offloading and quantization, achieving 15 tok/s with under 3 GiB of memory.
Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
Demonstrates running an 80B Qwen model in just 4.3 GB of RAM on a Mac and a 35B model on an iPhone, showcasing extreme memory optimization for local LLM inference.
@rohanpaul_ai: So much possibilities for on-device small models. Here @adrgrondin is running Google’s Gemma 4 E2B on iPhone 17 Pro. ~4…
Google's Gemma 4 E2B is demonstrated running on an iPhone 17 Pro via MLX optimization, achieving ~40 tokens/second with 128K context and offline thinking mode for coding and math.