@rohanpaul_ai: This paper makes long-context attention cheaper and faster by letting each token use only the query heads it needs. Rea…

X AI KOLs Following Papers

Summary

The paper introduces Grouped Query Experts, which improves long-context attention by routing each token to only a few query-head experts on top of grouped-query attention, achieving 1.7-1.8x faster prefill while matching accuracy.

This paper makes long-context attention cheaper and faster by letting each token use only the query heads it needs. Reached about 1.7 to 1.8 times faster prefill when context length became large. Standard attention makes every token run through every attention head, even when some heads are not useful for that token. The paper’s idea, called Grouped Query Experts, keeps the normal key and value cache from grouped-query attention but routes each token to only a few query-head experts. Grouped Query Experts sits on top of grouped-query attention, the trick many long-context models already use to reduce key-value cache cost. This is like giving the model many possible attention patterns, while making each token pay for only the small set that seems useful. The authors trained 250M-parameter models on 30B tokens and compared the method with a normal grouped-query attention baseline. The best version matched the baseline’s average accuracy, 56.04 versus 55.86, while using 9 of 16 query-attention computations. shows that attention can be made sparse inside grouped-query attention without hurting quality, but only when the router gets a strong learning signal and one shared head stays always on. ---- Link – arxiv. org/abs/2606.20945 Title: "Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention"
Original Article
View Cached Full Text

Cached at: 06/28/26, 12:00 PM

This paper makes long-context attention cheaper and faster by letting each token use only the query heads it needs.

Reached about 1.7 to 1.8 times faster prefill when context length became large.

Standard attention makes every token run through every attention head, even when some heads are not useful for that token.

The paper’s idea, called Grouped Query Experts, keeps the normal key and value cache from grouped-query attention but routes each token to only a few query-head experts.

Grouped Query Experts sits on top of grouped-query attention, the trick many long-context models already use to reduce key-value cache cost.

This is like giving the model many possible attention patterns, while making each token pay for only the small set that seems useful.

The authors trained 250M-parameter models on 30B tokens and compared the method with a normal grouped-query attention baseline.

The best version matched the baseline’s average accuracy, 56.04 versus 55.86, while using 9 of 16 query-attention computations.

shows that attention can be made sparse inside grouped-query attention without hurting quality, but only when the router gets a strong learning signal and one shared head stays always on.


Link – arxiv. org/abs/2606.20945

Title: “Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention”

Similar Articles

Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention

Hugging Face Daily Papers

Grouped Query Experts (GQE) improves Transformer efficiency by applying a mixture-of-experts layer on top of grouped-query attention, selectively activating query heads per token while keeping key-value cache benefits, matching baseline accuracy with half the query-head compute at 250M parameter scale.

Subquadratic AI introduces SubQ-1.1-Small, a new model using Smart Sparse Attention

Reddit r/singularity

Subquadratic AI introduces SubQ-1.1-Small, a model leveraging Smart Sparse Attention to achieve near-perfect long-context retrieval up to 12M tokens with up to 1,000x attention compute reduction. It balances long-context optimization with strong general reasoning, outperforming baselines on benchmarks like NIAH and RULER.