Qwen3.8-Flash-Next optimised for Macs
Summary
The article details custom optimizations for running the Qwen3.8-Flash-Next AI model on Mac M1 Max hardware, including SSD streaming, custom quantizations, and a sparse attention mechanism to improve performance.
Similar Articles
Ran Qwen3.8-Flash-Next (79 GB, 2-bit) at 350K ctx for 3.5 hours on a 128 GB M5 Max — speed vs context depth, 100 turns, one graph
This article reports on running the Qwen3.8-Flash-Next model on a MacBook Pro M5 Max, benchmarking speed versus context depth over 100 turns, with insights into performance and issues like role confusion at long contexts.
Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
slotstream is an open-source tool that enables running large AI models like Qwen3.8-Flash-Next on Macs with limited memory by streaming model weights from SSD, achieving approximately 12 tokens per second on a 48GB Mac.
Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency
The Qwen3.8-Flash-Next introduces a new AI architecture focused on achieving ultimate cost-efficiency in model performance.
Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. 👀
The article provides memory estimates for the Qwen3.8-Flash-Next model, suggesting it could be local-friendly with quantization techniques.
Qwen3.8-Next streaming - 150tps prefill, 3.6 tps decode on M5 Air
A user tested the Qwen3.8-Next model on an Apple M5 Air with 3-bit quantization, achieving 150 tokens per second prefill and 3.6 tps decode, outperforming a dense 27b model in some metrics.