Tag
The article announces the release of Qwen3.8-Flash-Next on MLX-serve, supporting 1 million token context with efficient performance on M5 Max hardware using quantized weights.
Users report issues with Qwen3.8-Flash-Next, such as poor multi-turn conversation tracking, compared to the better-performing Qwen3.8-27B.