gated-residual

Tag

Cards List
#gated-residual

Qwen4's architecture is here early, firing 6B parameters out of 125B (3 minute read)

TLDR AI · 2d ago Cached

Alibaba's Qwen team released Qwen3.8-Flash-Next, an open-weight preview of the Qwen4 architecture that activates only 6B parameters out of 125B to reduce inference costs and address hardware limitations.

0 favorites 0 likes
#gated-residual

Qwen/Qwen3.8-Flash-Next-FP8

Hugging Face Models Trending · 5d ago Cached

The article releases FP8-quantized weights for the Qwen3.8-Flash-Next model, introducing architectural innovations like Hybrid Attention with QSA, Gated Residual, and N-gram Embedding to improve efficiency in large language models.

0 favorites 0 likes
← Back to home

Submit Feedback