Faster than Light in Air: 8-22 tg/s Qwen3.8-Flash-Next (Q4/Q4ish) on a 32GB M4 MacBook Air
Summary
Cherenkov is a new inference engine for Apple Silicon that enables efficient memory-constrained inference of large AI models like Qwen3.8-Flash-Next using predictive expert streaming and mixed-precision execution.
Similar Articles
Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro
Inco Splash is an open-source inference engine optimized for Apple silicon, offering significant speed improvements for running AI models like Qwen3.8-27B on M-series MacBooks.
Qwen3.8-Flash-Next optimised for Macs
The article details custom optimizations for running the Qwen3.8-Flash-Next AI model on Mac M1 Max hardware, including SSD streaming, custom quantizations, and a sparse attention mechanism to improve performance.
Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
slotstream is an open-source tool that enables running large AI models like Qwen3.8-Flash-Next on Macs with limited memory by streaming model weights from SSD, achieving approximately 12 tokens per second on a 48GB Mac.
Qwen3.8-Flash-Next (95.5 GiB) on a 64GB Mac at ~27 tok/s, checkpoint + fork
The author optimized the Qwen3.8-Flash-Next model to run on a 64GB Mac using expert streaming and other techniques, achieving ~27 tok/s by publishing a checkpoint and a llama.cpp fork.
Qwen Flash Q4_K_M on 4080 + 64GB DDR5 at ~8tk/s 98304 CTX.
The author demonstrates how to run a 182B Qwen model on a 4080 GPU with 64GB DDR5 by offloading ngrams to SSD, achieving faster and more intelligent performance than a 27B model, making large model inference accessible on consumer hardware.