Qwen3.8-Flash-Next

Simon Willison's Blog Models

Summary

Qwen has released Qwen3.8-Flash-Next, an open-weights multimodal MoE model with 125B tokens but only 6B active parameters, providing a performance boost and serving as an early preview of the Qwen4 architecture.

No content available
Original Article
View Cached Full Text

Cached at: 08/27/26, 03:14 AM

# Qwen3.8-Flash-Next Source: [https://simonwillison.net/2026/Aug/26/qwen38-flash-next/](https://simonwillison.net/2026/Aug/26/qwen38-flash-next/) 26th August 2026 \- Link Blog **[Qwen3\.8\-Flash\-Next](https://qwen.ai/blog?id=qwen3.8-flash-next)**\([via](https://news.ycombinator.com/item?id=49448210)\) Another open weights model from Qwen\. This one is "a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4"\. It's pretty big: 125B tokens, but only 6B active which means it gets a significant performance boost\. I've been trying it out on a DGX Spark using[these Unsloth quantized models](https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF)\. I'm still exploring the model \- so far I've tried the 72\.5GB UD\-IQ1\_S one \(producing[these pelicans](https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ff9c69ebdab90d8a45b8de4742cc7b840)\) and the 78\.9GB UD\-Q2\_K\_XL \(producing[these](https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F6ba7cbfc1a9336986703b41f7fccd73a)\)\. My favorite so far was this xhigh reasoning effort one from UD\-Q2\_K\_XL: ![Flat vector illustration: a white pelican with an orange beak and orange legs rides a red bicycle along a sandy path, a wicker basket on the handlebars holding a blue fish, with green rolling hills, a small tree and bushes, white clouds and a bright yellow sun in a blue sky behind it](https://static.simonwillison.net/static/2026-08-27/IMG_7667.png)

Similar Articles

Qwen/Qwen3.8-Flash-Next

Hugging Face Models Trending

Release of Qwen3.8-Flash-Next, an open-weight AI model introducing architectural innovations like Hybrid Attention with QSA and Gated Residual for improved efficiency and scalability, previewing the future Qwen4 architecture.

Qwen/Qwen3.8-Flash-Next-FP8

Hugging Face Models Trending

The article releases FP8-quantized weights for the Qwen3.8-Flash-Next model, introducing architectural innovations like Hybrid Attention with QSA, Gated Residual, and N-gram Embedding to improve efficiency in large language models.