Qwen3.8-Flash-Next
Summary
Qwen has released Qwen3.8-Flash-Next, an open-weights multimodal MoE model with 125B tokens but only 6B active parameters, providing a performance boost and serving as an early preview of the Qwen4 architecture.
View Cached Full Text
Cached at: 08/27/26, 03:14 AM
Similar Articles
Qwen/Qwen3.8-Flash-Next
Release of Qwen3.8-Flash-Next, an open-weight AI model introducing architectural innovations like Hybrid Attention with QSA and Gated Residual for improved efficiency and scalability, previewing the future Qwen4 architecture.
Qwen/Qwen3.8-Flash-Next-FP8
The article releases FP8-quantized weights for the Qwen3.8-Flash-Next model, introducing architectural innovations like Hybrid Attention with QSA, Gated Residual, and N-gram Embedding to improve efficiency in large language models.
Qwen4's architecture is here early, firing 6B parameters out of 125B (3 minute read)
Alibaba's Qwen team released Qwen3.8-Flash-Next, an open-weight preview of the Qwen4 architecture that activates only 6B parameters out of 125B to reduce inference costs and address hardware limitations.
Qwen 3.8 Flash Next: Beating DS V4 Flash at half the parameters, stronger than Opus 4.6
Qwen team has released an open-weight model, Qwen 3.8 Flash Next, which surpasses DS V4 Flash in performance with half the parameters and is stronger than Opus 4.6, offering a preview of the Qwen 4 architecture.
@AdinaYakup: Yes https://huggingface.co/collections/Qwen/qwen38-flash-next…
A new multimodal AI model, Qwen3.8-Flash-Next-FP8, is released on Hugging Face with 180 billion parameters, supporting image-text-to-text tasks as part of a Qwen collection.