Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. πŸ‘€

Reddit r/LocalLLaMA Models

Summary

The article provides memory estimates for the Qwen3.8-Flash-Next model, suggesting it could be local-friendly with quantization techniques.

Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate: Ideal 4-bit quant β‰ˆ 82 GB (58 GB main weights + 24 GB n-gram tables) Real-world quants likely land in the 80–90 GB range. The big n-gram table is sparsely accessed β†’ excellent candidate for system RAM offload. This architecture could be surprisingly local-friendly once the weights drop.
Original Article

Similar Articles

Qwen/Qwen3.8-Flash-Next-FP8

Hugging Face Models Trending

The article releases FP8-quantized weights for the Qwen3.8-Flash-Next model, introducing architectural innovations like Hybrid Attention with QSA, Gated Residual, and N-gram Embedding to improve efficiency in large language models.

Qwen3.8-Flash-Next

Simon Willison's Blog

Qwen has released Qwen3.8-Flash-Next, an open-weights multimodal MoE model with 125B tokens but only 6B active parameters, providing a performance boost and serving as an early preview of the Qwen4 architecture.

Qwen/Qwen3.8-Flash-Next

Hugging Face Models Trending

Release of Qwen3.8-Flash-Next, an open-weight AI model introducing architectural innovations like Hybrid Attention with QSA and Gated Residual for improved efficiency and scalability, previewing the future Qwen4 architecture.