Qwen3.8-Flash-Next on MLX-serve, 1m context is released!
Summary
The article announces the release of Qwen3.8-Flash-Next on MLX-serve, supporting 1 million token context with efficient performance on M5 Max hardware using quantized weights.
Similar Articles
Qwen3.8-Flash-Next at 170K context on a single 96 GB card. ~110 tok/s.
The article describes a method to run the Qwen3.8-Flash-Next model with quantized n-grams to achieve over 170K token context on a single 96GB GPU card, with performance up to 110 tokens per second using INT4 quantization and memory-mapped disk access.
Qwen3.8-Flash-Next
Qwen has released Qwen3.8-Flash-Next, an open-weights multimodal MoE model with 125B tokens but only 6B active parameters, providing a performance boost and serving as an early preview of the Qwen4 architecture.
Ran Qwen3.8-Flash-Next (79 GB, 2-bit) at 350K ctx for 3.5 hours on a 128 GB M5 Max — speed vs context depth, 100 turns, one graph
This article reports on running the Qwen3.8-Flash-Next model on a MacBook Pro M5 Max, benchmarking speed versus context depth over 100 turns, with insights into performance and issues like role confusion at long contexts.
First evidence of a pending qwen3.7 open weights release. Qwen3.7-flash is on open router. They referred to Qwen3.6-35b-a3b as Qwen3.6 flash so this is likely a small MoE. The prices are substantially cheaper than 3.6 flash with a native 1M context window.
Evidence of an upcoming Qwen3.7 open weights release, with a flash variant (likely small MoE) listed on OpenRouter featuring a 1M context window and cheaper pricing than Qwen3.6 flash.
Qwen3.8-Flash-Next optimised for Macs
The article details custom optimizations for running the Qwen3.8-Flash-Next AI model on Mac M1 Max hardware, including SSD streaming, custom quantizations, and a sparse attention mechanism to improve performance.