Scaling Parameter and Context in Attention: Native Sparse Attention from Mixture-of-Head
Summary
Peking University researchers introduce NAMOH, an architecture-native sparse attention mechanism that activates K of H heads per token, so head routing jointly determines active parameters and available context. It enables parameter scaling to directly support efficient long-context scaling, outperforming fully activated models with the same total parameters while remaining compatible with GQA and existing sparse attention methods.
View Cached Full Text
Cached at: 10/01/26, 09:45 AM
# Scaling Parameter and Context in Attention: Native Sparse Attention from Mixture-of-Head
Source: [https://arxiv.org/html/2609.38832](https://arxiv.org/html/2609.38832)
Runsheng WangAffiliation:Peking UniversityMeng LiAffiliation:Peking University
###### Abstract
Scaling attention parameters can improve language model quality, but retaining full token histories makes additional heads costly at long contexts\. Furthermore, since attention retrieves and combines contextual information, parameter scaling should also support longer contexts\. We therefore askwhether attention parameter scaling can directly enable efficient and effective context scaling\.We introduce NAMOH, an architecture\-native sparse attention mechanism that activatesKKofHHheads per token\. Each head retains only its assigned tokens and performs causal attention within this subsequence\. Head selection thus jointly determines active parameters and available context without scanning the full history\. Under balanced assignments, increasingHHat fixedKKshortens head histories and reduces per\-token key\-value \(KV\) access without increasing total KV storage\. We further support head\-relative rotary position embeddings to shorten positional spans within routed subsequences, aiming to mitigate position\-induced attention noise\. Experiments show that NAMOH can outperform fully activated models with the same total parameters, while enabling more efficient long\-context inference than smaller dense models with matched active parameter counts\. It remains compatible with GQA and existing sparse attention mechanisms\. We hope this work offers a new path for scaling attention, with parameter scaling directly enabling context scaling\.
## 1Introduction
Scaling attention parameters through expert routing can improve the performance of large language models \(LLMs\)\([Zhang et al\., 2022](https://arxiv.org/html/2609.38832#bib.bib17);[Yang et al\., 2025](https://arxiv.org/html/2609.38832#bib.bib19)\)\. Methods such as SwitchHead and MoH further demonstrate gains in model quality and efficiency by selectively activating attention heads or projections\([Csordás et al\., 2023](https://arxiv.org/html/2609.38832#bib.bib18);[Jin et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib16)\)\. Yet most parameter growth in frontier models has been concentrated in feed\-forward networks \(FFNs\), particularly with the rise of mixture\-of\-experts \(MoE\) architectures\([Jiang et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib1);[Dubey et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib2);[Adler et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib3);[MiMo Team, 2026](https://arxiv.org/html/2609.38832#bib.bib6);[Yang et al\., 2024a](https://arxiv.org/html/2609.38832#bib.bib7)\)\. As shown in Figure[1](https://arxiv.org/html/2609.38832#S1.F1)\(a\), attention parameters have scaled much more slowly\.
A key challenge is to increase attention parameters without proportionally increasing the cost of storing and processing context\. When each head retains its own full token history, adding heads increases key\-value \(KV\) storage, while activating more heads increases full\-attention computation\. This raises activation memory and compute costs during training and prefill, as well as KV cache storage and memory traffic during decoding\([Dao, 2023](https://arxiv.org/html/2609.38832#bib.bib33);[Tang et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib10)\)\. Yet efficiency is only part of the goal\. Mechanistic studies show that FFNs can store knowledge acquired during training\([Geva et al\., 2020](https://arxiv.org/html/2609.38832#bib.bib24)\), while attention heads play a central role in retrieving and combining information from the current context\([Olsson et al\., 2022](https://arxiv.org/html/2609.38832#bib.bib25);[Guo et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib26)\)\. This role makes long\-context processing a natural objective for attention parameter scaling\.
Despite this connection, work on context scaling has primarily focused on computational efficiency rather than attention parameter scaling\. The expanding context windows in Figure[1](https://arxiv.org/html/2609.38832#S1.F1)\(b\) have been supported by sparse and linear\-time mechanisms, often combined with full attention\([Yuan et al\., 2025](https://arxiv.org/html/2609.38832#bib.bib12);[Yang et al\., 2024b](https://arxiv.org/html/2609.38832#bib.bib15);[DeepSeek\-AI et al\., 2026](https://arxiv.org/html/2609.38832#bib.bib4);[Bai et al\., 2026](https://arxiv.org/html/2609.38832#bib.bib5)\)\. Linear\-time recurrent architectures compress history into fixed\-size states, which can lose information needed for precise recall\([Jelassi et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib27);[Cabannes et al\., 2026](https://arxiv.org/html/2609.38832#bib.bib28)\)\. Selection\-based sparse attention preserves explicit token memories but introduces token or block selection overhead\([Lu et al\., 2025](https://arxiv.org/html/2609.38832#bib.bib11);[Yuan et al\., 2025](https://arxiv.org/html/2609.38832#bib.bib12)\)\. For indexers that scan the full prefix, this overhead grows with context length\([Xu et al\., 2026](https://arxiv.org/html/2609.38832#bib.bib29)\)\. Computational savings alone also leave positional limitations unresolved\. As[Du et al\. \(2026\)](https://arxiv.org/html/2609.38832#bib.bib31)show, long contexts can yield indistinguishable attention scores for different positions or tokens and disrupt token relevance rankings\. Related geometric analysis links long RoPE spans to spurious query\-key alignment, which can introduce attention noise by assigning weight to irrelevant tokens\([Wertheimer et al\., 2026](https://arxiv.org/html/2609.38832#bib.bib30)\)\.
Together, these challenges motivate a joint view of attention parameter and context scaling\. The goal is not only to add attention parameters at a manageable cost, but also to use those parameters to support longer contexts\. We therefore ask:*Can scaling attention parameters directly enable efficient and effective context scaling?*
We introduce NAMOH, an architecture\-native sparse attention mechanism that couples attention parameter and context scaling\. Each attention head acts as an expert, and a learned router activates theKKhighest\-scoring heads out ofHHtotal heads for each token\. Routing determines both the active head parameters and the token subsequence processed by each head\. Each head performs causal attention only within its routed subsequence, and the layer combines the projected head outputs using routing weights\. Attention sparsity therefore arises directly from the routing structure, rather than masking a dense attention result\.
Figure 1:Parameter and context scaling in frontier language models\.\(a\) Attention and FFN parameters in frontier models\. Since the rise of MoE, parameter growth has been dominated by FFN experts, with comparatively little growth in attention\. \(b\) Maximum context windows in frontier models\. Sparse and linear\-time mechanisms support further context expansion\.We use load balancing to discourage head collapse\([Jin et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib16);[Fu et al\., 2026](https://arxiv.org/html/2609.38832#bib.bib32)\)and distribute tokens approximately evenly across heads\. WithHHtotal heads andKKactive heads per token, a context ofLLtokens requiresLKLKkey\-value \(KV\) entries in total, with approximatelyLK/HLK/Hentries per head\. Each query token therefore accesses approximatelyLK2/HLK^\{2\}/HKV entries across itsKKactive heads\. We define the*KV activation ratio*as this per\-token KV access relative to a full\-attention baseline\. For baselines with the same model width and head dimension, this ratio is approximatelyK/HK/Hrelative toKK\-head full attention matched in active attention parameters, and\(K/H\)2\(K/H\)^\{2\}relative toHH\-head full attention matched in total attention parameters\.
For example, withH=32H=32andK=8K=8, total KV storage is1/41/4that of 32\-head full attention and equal to that of 8\-head full attention\. Under balanced routing, each head retains approximatelyL/4L/4entries, so eight active heads access approximately2L2Lentries per token\. This yields KV activation ratios of1/161/16and1/41/4relative to the two baselines, respectively\. Parameter scaling thus allows more heads to specialize in different contextual patterns, with each head processing its routed subsequence\.
We further support*head\-relative RoPE*, which encodes each token by its position within a head’s routed subsequence rather than its global position\. By shortening the encoded span to approximatelyLK/HLK/Hunder balanced routing while preserving token order, this design aims to mitigate position\-induced attention noise\. For efficient training and prefill, we pack head\-specific subsequences for variable\-length FlashAttention\([Dao, 2023](https://arxiv.org/html/2609.38832#bib.bib33)\)\. During decoding, we execute only active head queries and read or update only their corresponding KV caches\.
Our experiments show that NAMOH can achieve higher accuracy than fully activated models with the same total parameter count\. Its reduced KV access also enables more efficient long\-context inference than smaller dense models matched in active parameter count\. Moreover, NAMOH is compatible with GQA and existing sparse attention mechanisms, which can further select tokens or blocks within each head’s routed subsequence\. Together, these properties offer a complementary path to context scaling through attention parameter scaling\.
## 2Background
### 2\.1Head Sparsity: Mixture\-of\-Head
Mixture\-of\-experts \(MoE\) layers scale feed\-forward capacity through conditional activation without proportional growth in per\-token computation\([Shazeer et al\., 2017](https://arxiv.org/html/2609.38832#bib.bib34);[Fedus et al\., 2021](https://arxiv.org/html/2609.38832#bib.bib35)\)\. The same principle applies to attention parameters\.[Peng et al\. \(2020\)](https://arxiv.org/html/2609.38832#bib.bib21)mix overlapping groups of heads rather than route each token to individual heads\. Mixture of Attention Heads \(MoA\) routes query and output projections with shared keys and values\([Zhang et al\., 2022](https://arxiv.org/html/2609.38832#bib.bib17)\), while SwitchHead routes value and output projections\([Csordás et al\., 2023](https://arxiv.org/html/2609.38832#bib.bib18)\)\. Grouped Query Experts \(GQE\) routes query heads within grouped\-query attention\([Tripathi and Kumar, 2026](https://arxiv.org/html/2609.38832#bib.bib22)\)\.
Mixture\-of\-Head attention \(MoH\) selects heads per token and weights their projected outputs\([Jin et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib16)\)\. However, its heads retain full key\-value \(KV\) histories even for tokens whose head outputs are inactive\. With independent, fixed\-width KV heads, cache storage grows with head count and context length, while each active query still processes the full prefix\. Mixture of Sparse Attention \(MoSA\) introduces sequence sparsity through expert\-choice routing, with each head selecting a fixed quota of top\-scoring tokens from the full sequence\([Piekos et al\., 2025](https://arxiv.org/html/2609.38832#bib.bib20)\)\. This balances head loads, but future tokens can change earlier selections despite causal attention masking, so incremental autoregressive decoding requires routing changes\.
### 2\.2Context Sparsity: Sparse Attention
Selection\-based sparse attention moves relevance filtering before the main softmax attention\. Quest ranks KV pages using query\-dependent scores from key metadata\([Tang et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib10)\)\. Native Sparse Attention \(NSA\) reuses compressed attention scores to select blocks\([Yuan et al\., 2025](https://arxiv.org/html/2609.38832#bib.bib12)\), while Mixture of Block Attention \(MoBA\) selects blocks using query affinities to pooled keys\([Lu et al\., 2025](https://arxiv.org/html/2609.38832#bib.bib11)\)\. DeepSeek\-V4 combines KV compression with a lightweight indexer in its compressed sparse attention branch\([DeepSeek\-AI et al\., 2026](https://arxiv.org/html/2609.38832#bib.bib4)\)\. When indexers scan the prefix, their overhead grows with context length\([Xu et al\., 2026](https://arxiv.org/html/2609.38832#bib.bib29)\)\. Realizing speedups also requires specialized kernels and cache layouts for irregular KV access\([DeepSeek\-AI et al\., 2026](https://arxiv.org/html/2609.38832#bib.bib4)\)\. NSA also reports higher average performance than its full\-attention baseline on general and long\-context benchmarks, supporting sparsity as a modeling choice\([Yuan et al\., 2025](https://arxiv.org/html/2609.38832#bib.bib12)\)\.
NAMOH provides a complementary source of sparsity through token\-to\-head assignments, without ranking the full history per query\. Within routed subsequences, softmax attention remains compatible with dynamic token selection\([Zhang et al\., 2023](https://arxiv.org/html/2609.38832#bib.bib13);[Tang et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib10)\)and static patterns such as sliding windows with retained sink tokens\([Xiao et al\., 2023](https://arxiv.org/html/2609.38832#bib.bib14)\)\.
Token selection does not necessarily shorten positional spans when original indices are retained\. For rotary position embeddings \(RoPE\),[Wertheimer et al\. \(2026\)](https://arxiv.org/html/2609.38832#bib.bib30)link length extrapolation to disrupted query\-key separation and spurious attention to irrelevant tokens\.[Du et al\. \(2026\)](https://arxiv.org/html/2609.38832#bib.bib31)further identify indistinguishable scores across positions or tokens and reversed relevance rankings\. Sparsity alone therefore does not ensure reliable scores among retained tokens\. These findings motivate reducing the effective range of relative positions as a potential way to mitigate position\-induced attention noise\.
## 3NAMOH: Native Sparse Attention from Mixture\-of\-Head
Figure 2:NAMOH with two active heads out of four\.\(a\) Prefill routes tokens00to77before QKV projection\. Each head attends within its ordered subsequence, and projected outputs are combined with routing weights\. \(b\) During decoding, tokens88and99access and extend only their selected heads’ KV caches\. Position indices follow each head’s local token order\.### 3\.1Routed Causal Attention
For a length\-TTsequence with token representationsxt∈ℝ1×dmodelx\_\{t\}\\in\\mathbb\{R\}^\{1\\times d\_\{\\mathrm\{model\}\}\}and model widthdmodeld\_\{\\mathrm\{model\}\}, standard multi\-head attention withHHheads\([Vaswani et al\., 2017](https://arxiv.org/html/2609.38832#bib.bib8)\)can be written as
otfull=∑i=0H−1ht,ifullWO\(i\),o\_\{t\}^\{\\mathrm\{full\}\}=\\sum\_\{i=0\}^\{H\-1\}h\_\{t,i\}^\{\\mathrm\{full\}\}W\_\{O\}^\{\(i\)\},\(1\)whereht,ifull∈ℝ1×dheadh\_\{t,i\}^\{\\mathrm\{full\}\}\\in\\mathbb\{R\}^\{1\\times d\_\{\\mathrm\{head\}\}\}is headii’s causal attention output with head widthdheadd\_\{\\mathrm\{head\}\}, andWO\(i\)∈ℝdhead×dmodelW\_\{O\}^\{\(i\)\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{head\}\}\\times d\_\{\\mathrm\{model\}\}\}is its slice of the output projection\.
Token\-to\-head routing\.Following the expert view of attention heads\([Jin et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib16)\), NAMOH activatesKKofHHheads per token, where1≤K≤H1\\leq K\\leq H\. A learned router computes head affinities, selections, and gates as
at=softmax\(xtWr\),𝒮K\(t\)=TopKi\(at,i\),mt,i=𝟏\{i∈𝒮K\(t\)\},gt,i=mt,iat,i\.a\_\{t\}=\\operatorname\{softmax\}\(x\_\{t\}W\_\{r\}\),\\quad\\mathcal\{S\}\_\{K\}\(t\)=\\operatorname\{TopK\}\_\{i\}\(a\_\{t,i\}\),\\quad m\_\{t,i\}=\\mathbf\{1\}\\\{i\\in\\mathcal\{S\}\_\{K\}\(t\)\\\},\\quad g\_\{t,i\}=m\_\{t,i\}a\_\{t,i\}\.\(2\)Here,Wr∈ℝdmodel×HW\_\{r\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{model\}\}\\times H\}is the router matrix, andat∈ℝ1×Ha\_\{t\}\\in\\mathbb\{R\}^\{1\\times H\}contains affinities normalized by softmax across heads\.TopK\\operatorname\{TopK\}returns the indices of theKKlargest affinities, and𝟏\{⋅\}\\mathbf\{1\}\\\{\\cdot\\\}is the indicator function\. The selection maskmt,im\_\{t,i\}sets inactive gates to zero\.
Sparse projection and attention\.Routing determines both which parameters a token uses and which head histories it enters\. For headii, define the ordered index setℐi=\{t:mt,i=1\}\\mathcal\{I\}\_\{i\}=\\\{t:m\_\{t,i\}=1\\\}and its lengthni=\|ℐi\|n\_\{i\}=\|\\mathcal\{I\}\_\{i\}\|\. Tokens are dispatched to these subsequences before projection\. Only an active token\-head pair computes
qt,i=xtWQ\(i\),kt,i=xtWK\(i\),vt,i=xtWV\(i\),t∈ℐi\.q\_\{t,i\}=x\_\{t\}W\_\{Q\}^\{\(i\)\},\\qquad k\_\{t,i\}=x\_\{t\}W\_\{K\}^\{\(i\)\},\\qquad v\_\{t,i\}=x\_\{t\}W\_\{V\}^\{\(i\)\},\\qquad t\\in\\mathcal\{I\}\_\{i\}\.\(3\)Each head has independent query, key, and value matricesWQ\(i\),WK\(i\),WV\(i\)∈ℝdmodel×dheadW\_\{Q\}^\{\(i\)\},W\_\{K\}^\{\(i\)\},W\_\{V\}^\{\(i\)\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{model\}\}\\times d\_\{\\mathrm\{head\}\}\}\. The resulting queries, keys, and values lie inℝ1×dhead\\mathbb\{R\}^\{1\\times d\_\{\\mathrm\{head\}\}\}\. Inactive token\-head pairs generate no QKV vectors and occupy no KV cache entries\.
Letq~t,i\\widetilde\{q\}\_\{t,i\}andk~t,i\\widetilde\{k\}\_\{t,i\}denote queries and keys after positional encoding, described in Section[3\.2](https://arxiv.org/html/2609.38832#S3.SS2)\. Fori∈𝒮K\(t\)i\\in\\mathcal\{S\}\_\{K\}\(t\), attention and output aggregation are
ht,i=Attn\(q~t,i,\[k~s,i\],\[vs,i\]\),s∈ℐi,s≤t,ot=∑i∈𝒮K\(t\)gt,iht,iWO\(i\),h\_\{t,i\}=\\operatorname\{Attn\}\\\!\\left\(\\widetilde\{q\}\_\{t,i\},\[\\widetilde\{k\}\_\{s,i\}\],\[v\_\{s,i\}\]\\right\),s\\in\\mathcal\{I\}\_\{i\},s\\leq t,\\qquad o\_\{t\}=\\sum\_\{i\\in\\mathcal\{S\}\_\{K\}\(t\)\}g\_\{t,i\}h\_\{t,i\}W\_\{O\}^\{\(i\)\},\(4\)whereAttn\\operatorname\{Attn\}denotes standard scaled dot\-product attention,ht,i∈ℝ1×dheadh\_\{t,i\}\\in\\mathbb\{R\}^\{1\\times d\_\{\\mathrm\{head\}\}\},WO\(i\)∈ℝdhead×dmodelW\_\{O\}^\{\(i\)\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{head\}\}\\times d\_\{\\mathrm\{model\}\}\}, andot∈ℝ1×dmodelo\_\{t\}\\in\\mathbb\{R\}^\{1\\times d\_\{\\mathrm\{model\}\}\}\. Each query therefore attends only to earlier tokens and itself within the same routed head\. Sparsity is part of the computation, not a mask applied after dense attention\.
Figure[2](https://arxiv.org/html/2609.38832#S3.F2)\(a\) illustrates dispatch and aggregation withH=4H=4andK=2K=2\. During decoding in Figure[2](https://arxiv.org/html/2609.38832#S3.F2)\(b\), token88selects heads00and33, while token99selects heads22and33\. Each token appends to and attends within only its selected caches\. Because routing uses the current token representation, prefill and incremental decoding implement the same causal computation\.
### 3\.2Load Balancing and Scaling Properties
Load balancing\.Balancing encourages all heads to receive training signals and discourages head collapse\([Fu et al\., 2026](https://arxiv.org/html/2609.38832#bib.bib32)\)\. It also supports context sparsity: if routing concentrates on a fixed subset of heads, their histories can become dense despite sparse head activation\. For a training batchℬ\\mathcal\{B\}containingNNtokens, define the assignment fractionfif\_\{i\}and average head affinitypip\_\{i\}as
fi=1NK∑t∈ℬmt,i,pi=1N∑t∈ℬat,i,f\_\{i\}=\\frac\{1\}\{NK\}\\sum\_\{t\\in\\mathcal\{B\}\}m\_\{t,i\},\\qquad p\_\{i\}=\\frac\{1\}\{N\}\\sum\_\{t\\in\\mathcal\{B\}\}a\_\{t,i\},\(5\)with batch indices suppressed\. Both distributions sum to one, with uniform targetsfi=pi=1/Hf\_\{i\}=p\_\{i\}=1/H\.
We consider three alternatives from MoE: CV\-based importance regularization\([Shazeer et al\., 2017](https://arxiv.org/html/2609.38832#bib.bib34)\), Switch\-style \(fpfp\) balancing\([Fedus et al\., 2021](https://arxiv.org/html/2609.38832#bib.bib35)\), and auxiliary\-loss\-free balancing\([Wang et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib36)\)\. For the two auxiliary\-loss methods, the training objective isℒ=ℒLM\+λbalℒbal\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{LM\}\}\+\\lambda\_\{\\mathrm\{bal\}\}\\mathcal\{L\}\_\{\\mathrm\{bal\}\}, whereℒLM\\mathcal\{L\}\_\{\\mathrm\{LM\}\}is the language\-modeling loss,ℒbal\\mathcal\{L\}\_\{\\mathrm\{bal\}\}is the chosen balancing loss summed across routed attention layers, andλbal\>0\\lambda\_\{\\mathrm\{bal\}\}\>0controls its strength\. Loss\-free balancing instead adjusts head\-specific routing biases with an update rateη\>0\\eta\>0, without adding a balancing loss\. Appendix[A](https://arxiv.org/html/2609.38832#A1)details all three strategies\.
Routing as context selection\.Each head maintains its own routed history, so selecting heads also selects the contexts available to a query\. The head router thus serves as a learned indexer over histories without rescoring historical keys or blocks\. Unlike a fixed KV budget, the available history grows with sequence length, reaching approximatelyTK/HTK/Hentries in each head under balanced routing\. This selection acts at the head level and remains compatible with further token or block sparsity within each subsequence\([Tang et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib10);[Yuan et al\., 2025](https://arxiv.org/html/2609.38832#bib.bib12);[Lu et al\., 2025](https://arxiv.org/html/2609.38832#bib.bib11)\)\.
Parameter and context costs in attention\.We compare one attention layer withHHquery heads at fixed model and head widths, abbreviated asdmd\_\{m\}anddhd\_\{h\}\. LetMKVM\_\{\\mathrm\{KV\}\}count stored KV pairs after a prefix ofTTtokens, and letAKV\(T\)A\_\{\\mathrm\{KV\}\}\(T\)count distinct historical pairs used to decode the next token\. Each pair contains2dh2d\_\{h\}scalars; shared pairs are counted once\. GQA shares one KV head amongggquery heads, reducing both counts toHT/gHT/g\([Ainslie et al\., 2023](https://arxiv.org/html/2609.38832#bib.bib9)\)\. Block\-selected sparse attention \(SA\) retains the full cache but accesses approximately anα∈\(0,1\]\\alpha\\in\(0,1\]fraction of each history\([Tang et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib10);[Lu et al\., 2025](https://arxiv.org/html/2609.38832#bib.bib11)\)\. Thus, SA storesHTHTpairs and activates approximatelyαHT\\alpha HT\. MoH instead retains all head histories and activatesKTKTpairs through itsKKselected heads\([Jin et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib16)\)\.
For NAMOH, letnin\_\{i\}denote headii’s prefix length and𝒮K\(T\)\\mathcal\{S\}\_\{K\}\(T\)the next token’sKKselected heads\. Routing determines both cache insertion and access:
MKV=TK,AKV\(T\)=∑i∈𝒮K\(T\)ni≈TK2H\.M\_\{\\mathrm\{KV\}\}=TK,\\qquad A\_\{\\mathrm\{KV\}\}\(T\)=\\sum\_\{i\\in\\mathcal\{S\}\_\{K\}\(T\)\}n\_\{i\}\\approx\\frac\{TK^\{2\}\}\{H\}\.\(6\)Storage is exact, while activation assumes approximately balanced histories\. Relative toHH\-head MHA, these costs are reduced toK/HK/Hand approximately\(K/H\)2\(K/H\)^\{2\}, respectively\. Relative toKK\-head MHA, which matches active projection parameters, storage is unchanged and activation is reduced to approximatelyK/HK/H\.
Table[1](https://arxiv.org/html/2609.38832#S3.T1)compares attention projection weights and KV cache costs\. Router weights are omitted, which add𝒪\(Hdm\)\\mathcal\{O\}\(Hd\_\{m\}\)always\-active parameters for head routing\. The mechanisms are complementary: GQA shares KV representations, NAMOH shortens routed histories, and SA selects blocks within those histories\. With KV\-group routing and group\-shared block selection, combining all three storesTK/gTK/gKV pairs and activates approximatelyαTK2/\(Hg\)\\alpha TK^\{2\}/\(Hg\)historical pairs per decoding token\. Appendix[B](https://arxiv.org/html/2609.38832#A2)details the combination rules, selection overhead, and prefill computation\.
Table 1:Attention weight and KV cache costs\.Total and active projection weights, KV storage afterTTtokens, and historical KV activation for single\-token decoding\.Head\-relative positional encoding\.We optionally apply RoPE using a token’s rank within its routed head rather than its global position\. For an active token\-head pair, its zero\-based rankρi\(t\)\\rho\_\{i\}\(t\)and position\-encoded vectors are
ρi\(t\)=∑s=0tms,i−1,q~t,i=qt,iR\(ρi\(t\)\),k~t,i=kt,iR\(ρi\(t\)\),\\rho\_\{i\}\(t\)=\\sum\_\{s=0\}^\{t\}m\_\{s,i\}\-1,\\qquad\\widetilde\{q\}\_\{t,i\}=q\_\{t,i\}R\(\\rho\_\{i\}\(t\)\),\\qquad\\widetilde\{k\}\_\{t,i\}=k\_\{t,i\}R\(\\rho\_\{i\}\(t\)\),\(7\)whereR\(p\)∈ℝdhead×dheadR\(p\)\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{head\}\}\\times d\_\{\\mathrm\{head\}\}\}is the RoPE rotation for row vectors at positionpp\. For example, head00in Figure[2](https://arxiv.org/html/2609.38832#S3.F2)\(a\) maps global positions\(0,2,5,6\)\(0,2,5,6\)to local indices\(0,1,2,3\)\(0,1,2,3\)\. Its next activated token, token88, receives index44\. The largest index in headiiisni−1n\_\{i\}\-1, so balanced routing contracts the encoded span fromTTpositions to approximatelyTK/HTK/H\. This preserves token order, but not original token distances\. The shorter span is intended to improve query\-key matching by limiting position\-induced attention noise\([Wertheimer et al\., 2026](https://arxiv.org/html/2609.38832#bib.bib30);[Du et al\., 2026](https://arxiv.org/html/2609.38832#bib.bib31)\), complementing the computational benefit of sparsity\. Using global\-position RoPE instead amounts to replacingρi\(t\)\\rho\_\{i\}\(t\)withtt\.
Efficient execution\.Training and prefill pack each\(batch,head\)\(\\text\{batch\},\\text\{head\}\)subsequence as an independent sequence for variable\-length FlashAttention\([Dao, 2023](https://arxiv.org/html/2609.38832#bib.bib33)\)\. The total packed length is fixed atTKTKper sequence andNKNKper training batch, although individual head lengths vary\. Load balancing helps limit this length skew\. We further use length\-aware scheduling to interleave query tiles with larger and smaller causal workloads across subsequences\. This balances cumulative work across GPU multiprocessors to reduce tail idle time, without introducing cross\-subsequence attention\. Decoding launches only active head queries and reads or updates only their corresponding caches\.
## 4Experiments
### 4\.1Experimental Setup
Models and training\.We train models from scratch with 0\.6B to 1\.2B parameters, covering multi\-head attention \(MHA\), grouped\-query attention \(GQA\), selection\-based sparse attention \(SA\), and NAMOH\. We fix the number of layers at 16, the model width at 2048, and the head width at 64, with head counts of 4, 8, 16, and 32\. Every model receives 24B tokens from FineWeb\-Edu\([Penedo et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib57)\)\. This budget corresponds to 20 training tokens per parameter of the largest model, following the Chinchilla scaling guideline\([Hoffmann et al\., 2022](https://arxiv.org/html/2609.38832#bib.bib56)\), and is kept identical across models\. For long\-context evaluation, we further post\-train models on LongAlign\([Bai et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib58)\)\.
We use AdamW with a learning rate of0\.0020\.002, followed by linear decay to zero over the final20%20\\%of training\. Auxiliary\-loss balancing uses a coefficient ofλbal=0\.001\\lambda\_\{\\mathrm\{bal\}\}=0\.001, while loss\-free balancing uses a routing\-bias update rate ofη=0\.001\\eta=0\.001\. Other optimizer hyperparameters follow the implementation defaults of the AdamW optimizer\. All experiments are conducted on NVIDIA A100 GPUs\.
Baselines and controls\.MHA\-HHand SA\-HHuseHHheads, GQA\-32KVKKuses 32 query heads andKKKV heads, and NAMOH\-HHAKKactivatesKKofHHheads per token\. Our SA baseline follows the block\-selection paradigm\([Lu et al\., 2025](https://arxiv.org/html/2609.38832#bib.bib11);[Yuan et al\., 2025](https://arxiv.org/html/2609.38832#bib.bib12)\)\. Each head retains its full KV history, represents each KV block by its mean\-pooled keys for selection scoring, and selects a fixed fraction of the highest\-scoring blocks\. We use 32\-token blocks, reuse selected indices across groups of 32 queries, and additionally maintain a 128\-token sliding window\. The fraction in parentheses specifies the per\-head KV activation budget\. KV storage and activation are normalized to MHA\-32\. Activation counts entries accessed by the main attention operation, with shared GQA entries counted once\. To isolate the gains from routed context sparsity, we equip all MHA, GQA, and SA baselines with the head gating and load\-balancing mechanisms used in prior head\-routing methods\([Fu et al\., 2026](https://arxiv.org/html/2609.38832#bib.bib32);[Jin et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib16);[Qiu et al\., 2025](https://arxiv.org/html/2609.38832#bib.bib23)\), with matched settings across comparisons\.
Table 2:Evaluation results of models trained from scratch with different attention mechanisms and KV budgets\. KV storage \(KV Stor\.\) and activation \(KV Act\.\) are normalized to MHA\-32\. For sparse attention \(SA\), the fraction in parentheses denotes the activated KV fraction per head\.Table 3:Evaluation results on LongBench\. KV stor\. and act\. are normalized to MHA\-32; SA parentheses denote the KV fraction activated\.
Figure 3:Inference efficiency\. \(a\) Prefill TTFT and \(b\) decoding TPOT across attention mechanisms at different context lengths\. Lower is better\.
Evaluation\.We evaluate pretrained models on MMLU\([Hendrycks et al\., 2020](https://arxiv.org/html/2609.38832#bib.bib40)\)for general knowledge, GSM8K\([Cobbe et al\., 2021](https://arxiv.org/html/2609.38832#bib.bib38)\)for mathematics, HumanEval\([Chen et al\., 2021](https://arxiv.org/html/2609.38832#bib.bib41)\)for coding, and BoolQ\([Clark et al\., 2019](https://arxiv.org/html/2609.38832#bib.bib45)\)for reading comprehension\. Scientific reasoning benchmarks include ARC\-Easy, ARC\-Challenge\([Clark et al\., 2018](https://arxiv.org/html/2609.38832#bib.bib39)\), and OpenBookQA\([Mihaylov et al\., 2018](https://arxiv.org/html/2609.38832#bib.bib44)\)\. We assess commonsense reasoning with HellaSwag\([Zellers et al\., 2019](https://arxiv.org/html/2609.38832#bib.bib42)\), PIQA\([Bisk et al\., 2019](https://arxiv.org/html/2609.38832#bib.bib43)\), and WinoGrande\([Sakaguchi et al\., 2019](https://arxiv.org/html/2609.38832#bib.bib46)\)\. For long\-context understanding, we evaluate models on eight LongBench tasks\([Bai et al\., 2023](https://arxiv.org/html/2609.38832#bib.bib47)\): HotpotQA\([Yang et al\., 2018](https://arxiv.org/html/2609.38832#bib.bib48)\), Qasper\([Dasigi et al\., 2021](https://arxiv.org/html/2609.38832#bib.bib49)\), TriviaQA\([Joshi et al\., 2017](https://arxiv.org/html/2609.38832#bib.bib50)\), NarrativeQA\([Kociský et al\., 2017](https://arxiv.org/html/2609.38832#bib.bib51)\), 2WikiMultiHopQA\([Ho et al\., 2020](https://arxiv.org/html/2609.38832#bib.bib52)\), GovReport\([Huang et al\., 2021](https://arxiv.org/html/2609.38832#bib.bib53)\), QMSum\([Zhong et al\., 2021](https://arxiv.org/html/2609.38832#bib.bib54)\), and TREC\([Li and Roth, 2002](https://arxiv.org/html/2609.38832#bib.bib55)\)\.
### 4\.2Model Quality and Inference Efficiency
General capabilities\.Table[2](https://arxiv.org/html/2609.38832#S4.T2)compares NAMOH\-32AKKwith MHA, GQA, and SA forK∈\{4,8,16\}K\\in\\\{4,8,16\\\}\. All variants in this table use CV\-based importance regularization for load balancing, with details provided in Appendix[A](https://arxiv.org/html/2609.38832#A1)\. Across these settings, NAMOH achieves stronger overall performance than the correspondingKK\-head MHA and SA models and compares favorably with GQA\. Its performance can also match or exceed MHA\-32 despite sparse head activation\. These results show that a larger pool of selectively activated heads can preserve the quality of a fully activated model while using fewer head parameters per token\.
Long\-context quality and KV budgets\.Table[3](https://arxiv.org/html/2609.38832#S4.F3)extends the comparison to LongBench with a maximum context length of 32K tokens, where NAMOH\-32A8 improves overall performance over the MHA, GQA, and SA baselines\. Under balanced routing, its normalized KV storage is1/41/4and its KV activation is approximately1/161/16\. MHA\-8 and GQA\-32KV8 match its storage but access approximately four times as many KV entries\. SA\-32 \(1/161/16\) matches its activation budget but requires four times the storage, while SA\-8 \(1/41/4\) matches both\.
Figure 4:Full\-question attention over a 32k\-token haystack in NAMOH\-32A8\. Dashed lines mark the needle and four distractors\.
Table 4:LongBench evaluation of NAMOH\-32A8 with different RoPE variants across different context lengths\. G and HR denote global RoPE and head\-relative RoPE\.
Figure 5:Task\-dependent head utilization in NAMOH\-32A8\. Activation frequencies of individual heads in one attention layer on MMLU, GSM8K, PIQA, and HumanEval\.Inference efficiency\.Figure[3](https://arxiv.org/html/2609.38832#S4.F3)reports time to first token \(TTFT\) and time per output token \(TPOT\) across context lengths\. TTFT measures the latency to process the prompt and produce the first output token\. TPOT measures the average latency of subsequent decoding steps\. For the SA baselines, we use the implementation from MoBA\([Lu et al\., 2025](https://arxiv.org/html/2609.38832#bib.bib11)\)\. At long contexts, NAMOH\-32A8 achieves lower TTFT and TPOT than MHA\-32, GQA\-32KV8, and SA variants, and even outperforms MHA\-8\.
These gains are consistent with the different bottlenecks of decoding and prefill\. Long\-context decoding is typically memory\-bound, so reducing KV traffic helps lower TPOT\([Yuan et al\., 2025](https://arxiv.org/html/2609.38832#bib.bib12)\)\. Compute\-bound prefill benefits from fewer query\-key interactions and sparse projection\. GQA reduces KV projection costs but retains full\-prefix attention for all 32 query heads\. SA reduces the attended KV set but still incurs block\-scoring and selection overhead\([Lu et al\., 2025](https://arxiv.org/html/2609.38832#bib.bib11);[Yuan et al\., 2025](https://arxiv.org/html/2609.38832#bib.bib12)\)\. In contrast, NAMOH performs attention over shorter packed subsequences without a history\-scanning indexer, which also benefits comparisons at matched KV activation\.
### 4\.3Positional Encoding and Routing Behavior
Head\-relative RoPE\.Figure[4](https://arxiv.org/html/2609.38832#S4.F4)compares global and head\-relative RoPE in NAMOH\-32A8 on a 32k\-token needle\-in\-a\-haystack example with four distractors\. The visualization aggregates attention from all question tokens to each haystack position\. In this example, head\-relative RoPE assigns more attention to the needle and less to irrelevant positions\. Table[4](https://arxiv.org/html/2609.38832#S4.F4)provides a broader comparison at context lengths of 8k, 16k, and 32k\. Head\-relative RoPE achieves higher LongBench scores, with the average gap increasing at longer contexts\. These observations are consistent with the intended benefit of shortening the encoded positional span, as discussed in Section[3\.2](https://arxiv.org/html/2609.38832#S3.SS2)\.
Head utilization across tasks\.Figure[5](https://arxiv.org/html/2609.38832#S4.F5)shows head activation frequencies on four datasets\. Frequencies cluster around the balanced rate ofK/H=25%K/H=25\\%, indicating broad head utilization rather than concentration on a small fixed subset\. At the same time, individual heads exhibit clear frequency differences across tasks\. This pattern suggests task\-dependent specialization\.
Routing\-induced attention structure\.Figure[6](https://arxiv.org/html/2609.38832#S4.F6)examines a 32\-token example\. Panel \(a\) shows the heads selected by each token, with lines linking tokens assigned to the same head\. For this sequence of lengthT=32T=32, letM∈\{0,1\}T×HM\\in\\\{0,1\\\}^\{T\\times H\}collect the routing masks, withMt,i=mt,iM\_\{t,i\}=m\_\{t,i\}\. Panel \(b\) visualizesR=tril\(MM⊤\)R=\\operatorname\{tril\}\(MM^\{\\top\}\), wheretril\\operatorname\{tril\}retains the lower triangle, including the diagonal\. Thus,Rt,sR\_\{t,s\}counts the active heads shared by query tokenttand an earlier or current tokenss\.
Panel \(c\) shows the gate\-weighted attention mapAt,s=∑i=0H−1gt,ims,iαt,s\(i\)A\_\{t,s\}=\\sum\_\{i=0\}^\{H\-1\}g\_\{t,i\}m\_\{s,i\}\\alpha\_\{t,s\}^\{\(i\)\}\. Here,αt,s\(i\)\\alpha\_\{t,s\}^\{\(i\)\}is the softmax attention weight from tokenttto tokensswithin headii, defined as zero outside that head’s routed causal pairs\. Panels \(b\) and \(c\) therefore distinguish available connections from the attention weights assigned to them\. Panel \(d\) displays individual head maps at the original token positions\. Their separated support reflects the different routed subsequences, with non\-routed positions absent from each head’s computation\.
### 4\.4Load Balancing and Shared Heads
Load\-balancing strategies\.Table[5](https://arxiv.org/html/2609.38832#S4.T5)compares CV\-based importance regularization\([Shazeer et al\., 2017](https://arxiv.org/html/2609.38832#bib.bib34)\), Switch\-style \(fpfp\) balancing\([Fedus et al\., 2021](https://arxiv.org/html/2609.38832#bib.bib35)\), and loss\-free balancing\([Wang et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib36)\)\. The formulations of these load\-balancing strategies are detailed in Appendix[A](https://arxiv.org/html/2609.38832#A1)\. All variants retain 32 routed heads and activate eight per token\. At the settings specified above,fpfpachieves the best overall performance among the three alternatives\.
Figure 6:Routing shapes attention connectivity in NAMOH\. \(a\) Token\-to\-head assignments\. \(b\) The number of shared active heads for each causal token pair\. \(c\) Layer\-wide attention weighted by query\-specific routing gates\. \(d\) Attention maps for heads 0, 8, 16, and 24\.Table 5:Further refinement of NAMOH\-32A8 with load\-balancing strategies and shared heads\.Shared heads for local context\.Inspired by shared expert isolation in DeepSeekMoE\([Dai et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib37)\), we augment NAMOH\-32A8 withHsH\_\{s\}always\-active shared heads, testingHs∈\{2,4\}H\_\{s\}\\in\\\{2,4\\\}\. Each shared head attends either to the full causal prefix or to the most recentw=128w=128original tokens, including the current token, wherewwis the sliding\-window size\. Their independently projected outputs are added to the routed outputoto\_\{t\}\. Shared heads are excluded fromHH,KK, and the routing balance statistics\. The sliding\-window heads use global RoPE, while routed heads retain head\-relative RoPE\. This supplies local context even when neighboring tokens select different routed heads\.
Shared heads add4Hsdmodeldhead4H\_\{s\}d\_\{\\mathrm\{model\}\}d\_\{\\mathrm\{head\}\}always\-active projection parameters\. For a fixed window sizeww, the sliding variant requires at mostHswH\_\{s\}wadditional KV pairs per layer and𝒪\(THswdhead\)\\mathcal\{O\}\(TH\_\{s\}wd\_\{\\mathrm\{head\}\}\)attention work for a length\-TTsequence\. Table[5](https://arxiv.org/html/2609.38832#S4.T5)shows that adding shared heads improves overall performance\. Sliding\-window variants perform similarly to their full\-attention counterparts, suggesting that reliable local context accounts for much of the benefit\. This supports a complementary design in which shared heads cover nearby tokens while routed heads can focus on broader context retrieval\.
## 5Conclusion
We introduced NAMOH, an architecture\-native sparse attention mechanism that connects attention parameter scaling with context scaling\. Each token activates a subset of heads, and each head stores and attends only to its assigned tokens\. Routing thus jointly selects active parameters and available context without scanning the full history\. Under balanced routing, expanding the head pool at a fixed active head count shortens head histories\. This reduces per\-token KV access without increasing total KV storage\. Head\-relative RoPE further shortens positional spans and improves long\-context quality in our evaluations\. Experiments show that NAMOH can outperform fully activated models with the same total parameter count\. It also enables faster long\-context prefill and decoding than smaller dense models matched in active parameter count\. The design remains compatible with GQA and existing sparse attention mechanisms\. Together, these findings support a complementary path for scaling attention, in which parameter growth directly enables more efficient and effective context scaling\.
## References
- Adleret al\.\(2024\)N\. B\. Adler, N\. Agarwal, A\. Aithal, D\. H\. Anh, P\. Bhattacharya, A\. Brundyn, J\. Casper, B\. Catanzaro, S\. Clay, J\. Cohen, S\. Das, A\. Dattagupta, O\. Delalleau, L\. Derczynski, Y\. Dong, D\. Egert, E\. Evans, A\. Ficek, D\. Fridman, S\. Ghosh, B\. Ginsburg, I\. Gitman, T\. Grzegorzek, R\. Hero, J\. Huang, V\. Jawa, J\. Jennings, A\. Jhunjhunwala, J\. Kamalu, S\. Khan, O\. Kuchaiev, P\. LeGresley, H\. Li, J\. Liu, Z\. Liu, E\. Long, A\. S\. Mahabaleshwarkar, S\. Majumdar, J\. Maki, M\. Martínez, M\. R\. de Melo, I\. Moshkov, D\. Narayanan, S\. Narenthiran, J\. Navarro, P\. T\. Nguyen, O\. Nitski, V\. Noroozi, G\. Nutheti, C\. Parisien, J\. Parmar, M\. Patwary, K\. Pawelec, W\. Ping, S\. Prabhumoye, R\. Roy, T\. Saar, V\. R\. N\. Sabavat, S\. Satheesh, J\. Scowcroft, J\. D\. Sewall, P\. Shamis, G\. Shen, M\. Shoeybi, D\. Sizer, M\. Smelyanskiy, F\. Soares, M\. N\. Sreedhar, D\. Su, S\. Subramanian, S\. Sun, S\. Toshniwal, H\. Wang, Z\. Wang, J\. You, J\. Zeng, J\. Zhang, J\. Zhang, V\. Zhang, Y\. Zhang, and C\. ZhuNemotron\-4 340b technical report\.ArXivabs/2406\.11704\.Cited by:[§1](https://arxiv.org/html/2609.38832#S1.p1.1)\.
- Ainslieet al\.\(2023\)J\. Ainslie, J\. P\. Lee\-Thorp, M\. de Jong, Y\. Zemlyanskiy, F\. Lebrón, and S\. K\. SanghaiGQA: training generalized multi\-query transformer models from multi\-head checkpoints\.ArXivabs/2305\.13245\.External Links:[Link](https://api.semanticscholar.org/CorpusID:258833177)Cited by:[§B\.2](https://arxiv.org/html/2609.38832#A2.SS2.p2.1),[§3\.2](https://arxiv.org/html/2609.38832#S3.SS2.p4.1)\.
- Baiet al\.\(2026\)K\. T\. Y\. Bai, Y\. Bai, Y\. Bao, Chandru\. M, J\. Cai, X\. Cai, P\. Cao, Y\. Cao, Z\. Chai, Y\. Charles, H\. S\. Che, G\. Chen, G\. Chen, G\. Chen, H\. Chen, J\. Chen, J\. Chen, J\. Chen, K\. Chen, P\. Chen, R\. Chen, W\. Chen, X\. Chen, Y\. Chen, Y\. Chen, Y\. Chen, Y\. Chen, Y\. Chen, Y\. Chen, Y\. Chen, Z\. Chen, D\. Cheng, Y\. Cheng, J\. Cui, J\. Cui, A\. Dai, J\. Deng, H\. Ding, R\. Ding, S\. Ding, M\. Dong, M\. Dong, Y\. Dong, Y\. Dong, A\. Du, C\. Du, D\. Du, J\. Du, Y\. Du, Y\. Fan, J\. Feng, Q\. Feng, Y\. Feng, K\. Fu, Q\. Fu, F\. Gao, H\. Gao, J\. Gao, T\. Gao, W\. Gao, S\. Geng, J\. Gong, L\. Gong, S\. Gong, X\. Gong, Q\. Gu, Y\. Gu, S\. Guan, H\. Guo, S\. Guo, X\. Guo, Z\. Guo, B\. Hao, W\. Hao, X\. Hao, D\. He, H\. He, L\. He, Q\. He, W\. He, X\. He, X\. He, Y\. He, Y\. He, C\. Hong, T\. Hong, H\. Hu, J\. Hu, R\. Hu, W\. Hu, Y\. Hu, Z\. Hu, L\. Hua, J\. Huang, K\. Huang, R\. Huang, S\. Huang, W\. Huang, Y\. Huang, Z\. Huang, Z\. Huang, Y\. Hui, C\. Jia, Y\. Jiang, Z\. Jiang, Z\. Jiang, W\.M\. Jin, X\. Jin, Y\. Jing, H\. Kong, G\. Lai, A\. Li, C\. Li, C\. Li, C\. Li, F\. Li, G\. Li, H\. Li, J\. Li, J\. Li, L\. Li, L\. Li, L\. Li, W\. Li, W\. Li, X\. Li, Y\. Li, Y\. Li, Y\. Li, Y\. Li, Z\. Li, Z\. Li, Z\. Li, Z\. Li, Z\. Li, J\. Lin, X\. Lin, Y\. Lin, Z\. Lin, Z\. Lin, B\. Liu, B\. Liu, C\. Liu, L\. Liu, S\. Liu, S\. Liu, S\. Liu, T\. Liu, W\. Liu, Y\. Liu, Y\. Liu, Y\. Liu, Y\. Liu, Z\. Liu, Z\. Liu, E\. Lu, H\. Lu, L\. Lu, T\. Lu, Z\. Lu, A\. Luo, G\. Luo, J\. Luo, Y\. Luo, B\. Lyu, W\. Lyu, S\. Mao, Y\. Mei, X\. Men, M\. Ni, Y\. Niu, S\. Pan, S\. Peng, Z\. Qi, R\. Qin, Z\. Qin, Z\. Qin, H\. Qiu, J\. Qiu, J\. Qiu, B\. Qu, Y\. Qu, Z\. Shang, Y\. Shao, H\. Shen, J\. Shi, J\. Shi, L\. Shi, S\. Shi, W\. Siu, P\. Song, X\. Song, J\. Su, Y\. Su, Z\. Su, L\. Sui, J\. Sun, J\. Sun, S\. Sun, S\. Sun, T\. Sun, Y\. Sun, Y\. Tai, C\. Tang, H\. Tang, S\. Tang, Z\. Tang, C\. Tian, R\. Tian, Y\. Tian, W\. Tu, C\. Wang, C\. Wang, C\. Wang, D\. Wang, F\. Wang, H\. Wang, H\. Wang, H\. Wang, H\. Wang, H\. Wang, J\. Wang, J\. Wang, J\. Wang, J\. Wang, L\. Wang, S\. Wang, S\. Wang, S\. Wang, S\. Wang, S\. Wang, T\. Wang, W\. Wang, X\. Wang, X\. Wang, X\. Wang, X\. Wang, Y\. Wang, Y\. Wang, Y\. Wang, Y\. Wang, Y\. Wang, Y\. Wang, Y\. Wang, Y\. Wang, Z\. Wang, Z\. Wang, Z\. Wang, Z\. Wang, Z\. Wang, Z\. Wang, C\. Wei, M\. Wei, S\. Wei, Z\. Wen, F\. Wu, H\. Wu, R\. Wu, W\. Wu, X\. Wu, Y\. Wu, Y\. Wu, Y\. Wu, Z\. Wu, X\. Xian, C\. Xiang, Y\. Xiang, B\. Xiao, C\. Xiao, X\. Xiao, J\. Xie, X\. Xie, Y\. Xie, Z\. Xie, B\. Xing, Y\. Xiong, B\. Xu, B\. Xu, J\. Xu, J\. Xu, J\. Xu, J\. Xu, L\. H\. Xu, Q\. Xu, S\. Xu, S\. Xu, T\. Xu, T\. Xu, W\. Xu, X\. Xu, Y\. Xu, Y\. Xu, Y\. Xu, Z\. Xu, H\. Xue, J\. Yan, Y\. Yan, F\. Yang, G\. Yang, H\. Yang, J\. Yang, R\. Yang, W\. Yang, X\. Yang, X\. Yang, Y\. Yang, Y\. Yang, Y\. Yang, Y\. Yang, Z\. Yang, Z\. Yang, R\. Yang, Z\. Yang, H\. Yao, D\. Ye, H\. Ye, W\. Ye, Z\. G\. Ye, B\. Yin, H\. Yin, X\. Yin, C\. Yu, H\. Yu, L\. Yu, S\. Yu, S\. Yu, T\. Yu, E\. Yuan, M\. Yuan, T\. Yue, W\. Yue, Y\. Yue, D\. Zha, H\. Zhan, B\. Zhang, D\. Zhang, F\. Zhang, H\. Zhang, H\. Zhang, H\. Zhang, J\. Zhang, J\. Zhang, J\. Zhang, K\. Zhang, M\. Zhang, P\. Zhang, Q\. Zhang, R\. Zhang, R\. Zhang, S\. Zhang, S\. Zhang, X\. Zhang, X\. Zhang, Y\. Zhang, Y\. Zhang, Y\. Zhang, Y\. Zhang, Y\. Zhang, Y\. Zhang, Y\. Zhang, Y\. Zhang, Y\. Zhang, Y\. Zhang, Z\. Zhang, Z\. Zhang, B\. Zhao, C\. Zhao, F\. Zhao, J\. Zhao, J\. Zhao, S\. Zhao, W\. Zhao, X\. Zhao, X\. Zhao, Y\. Zhao, Z\. Zhao, H\. Zheng, H\. Zheng, R\. Zheng, S\. J\. Zheng, T\. Zheng, H\. Zhong, L\. Zhong, L\. Zhong, M\. Zhou, Q\. Zhou, R\. Zhou, R\. Zhou, X\. Zhou, Y\. Zhou, Z\. Zhou, J\. Zhu, L\. Zhu, X\. Zhu, Y\. Zhu, Y\. Zhu, Z\. Zhu, Z\. Chen, W\. Zhuang, and X\. ZuKimi k3: open frontier intelligence\.Cited by:[§1](https://arxiv.org/html/2609.38832#S1.p3.1)\.
- Baiet al\.\(2024\)Y\. Bai, X\. Lv, J\. Zhang, Y\. He, J\. Qi, L\. Hou, J\. Tang, Y\. Dong, and J\. LiLongAlign: a recipe for long context alignment of large language models\.InConference on Empirical Methods in Natural Language Processing,External Links:[Link](https://api.semanticscholar.org/CorpusID:267335080)Cited by:[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p1.1)\.
- Baiet al\.\(2023\)Y\. Bai, X\. Lv, J\. Zhang, H\. Lyu, J\. Tang, Z\. Huang, Z\. Du, X\. Liu, A\. Zeng, L\. Hou, Y\. Dong, J\. Tang, and J\. LiLongBench: a bilingual, multitask benchmark for long context understanding\.ArXivabs/2308\.14508\.Cited by:[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p4.1)\.
- Bisket al\.\(2019\)Y\. Bisk, R\. Zellers, R\. L\. Bras, J\. Gao, and Y\. ChoiPIQA: reasoning about physical commonsense in natural language\.InAAAI Conference on Artificial Intelligence,Cited by:[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p4.1)\.
- Cabanneset al\.\(2026\)L\. Cabannes, P\. Mazaré, G\. Szilvasy, M\. Douze, M\. Lomeli, I\. A\. Auzina, J\. Carpentier, G\. Synnaeve, and H\. JégouSparse delta memory: scaling the state of linear rnns through sparsity\.ArXivabs/2607\.07386\.Cited by:[§1](https://arxiv.org/html/2609.38832#S1.p3.1)\.
- Chenet al\.\(2021\)M\. Chen, J\. Tworek, H\. Jun, Q\. Yuan, H\. Pondé, J\. Kaplan, H\. Edwards, Y\. Burda, N\. Joseph, G\. Brockman, A\. Ray, R\. Puri, G\. Krueger, M\. Petrov, H\. Khlaaf, G\. Sastry, P\. Mishkin, B\. Chan, S\. Gray, N\. Ryder, M\. Pavlov, A\. Power, L\. Kaiser, M\. Bavarian, C\. Winter, P\. Tillet, F\. P\. Such, D\. W\. Cummings, M\. Plappert, F\. Chantzis, E\. Barnes, A\. Herbert\-Voss, W\. H\. Guss, A\. Nichol, I\. Babuschkin, S\. Balaji, S\. Jain, A\. Carr, J\. Leike, J\. Achiam, V\. Misra, E\. Morikawa, A\. Radford, M\. M\. Knight, M\. Brundage, M\. Murati, K\. Mayer, P\. Welinder, B\. McGrew, D\. Amodei, S\. McCandlish, I\. Sutskever, and W\. ZarembaEvaluating large language models trained on code\.ArXivabs/2107\.03374\.Cited by:[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p4.1)\.
- Clarket al\.\(2019\)C\. Clark, K\. Lee, M\. Chang, T\. Kwiatkowski, M\. Collins, and K\. ToutanovaBoolQ: exploring the surprising difficulty of natural yes/no questions\.ArXivabs/1905\.10044\.Cited by:[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p4.1)\.
- Clarket al\.\(2018\)P\. Clark, I\. Cowhey, O\. Etzioni, T\. Khot, A\. Sabharwal, C\. Schoenick, and O\. TafjordThink you have solved question answering? try arc, the ai2 reasoning challenge\.ArXivabs/1803\.05457\.Cited by:[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p4.1)\.
- Cobbeet al\.\(2021\)K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano, C\. Hesse, and J\. SchulmanTraining verifiers to solve math word problems\.ArXivabs/2110\.14168\.Cited by:[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p4.1)\.
- Csordáset al\.\(2023\)R\. Csordás, P\. Piekos, K\. Irie, and J\. SchmidhuberSwitchHead: accelerating transformers with mixture\-of\-experts attention\.ArXivabs/2312\.07987\.Cited by:[§1](https://arxiv.org/html/2609.38832#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.38832#S2.SS1.p1.1)\.
- Daiet al\.\(2024\)D\. Dai, C\. Deng, C\. Zhao, R\. Xu, H\. Gao, D\. Chen, J\. Li, W\. Zeng, X\. Yu, Y\. Wu, Z\. Xie, Y\. K\. Li, P\. Huang, F\. Luo, C\. Ruan, Z\. Sui, and W\. LiangDeepSeekMoE: towards ultimate expert specialization in mixture\-of\-experts language models\.InAnnual Meeting of the Association for Computational Linguistics,Cited by:[§4\.4](https://arxiv.org/html/2609.38832#S4.SS4.p2.1)\.
- Dao \(2023\)T\. DaoFlashAttention\-2: faster attention with better parallelism and work partitioning\.ArXivabs/2307\.08691\.Cited by:[§1](https://arxiv.org/html/2609.38832#S1.p2.1),[§1](https://arxiv.org/html/2609.38832#S1.p8.1),[§3\.2](https://arxiv.org/html/2609.38832#S3.SS2.p8.1)\.
- Dasigiet al\.\(2021\)P\. Dasigi, K\. Lo, I\. Beltagy, A\. Cohan, N\. A\. Smith, and M\. GardnerA dataset of information\-seeking questions and answers anchored in research papers\.ArXivabs/2105\.03011\.Cited by:[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p4.1)\.
- DeepSeek\-AIet al\.\(2026\)DeepSeek\-AI, A\. Xu, B\. Lin, B\. Xue, B\. Wang, B\. Xu, B\. Wu, B\. Zhang, C\. Lin, C\. Dong, C\. Ling, C\. Lu, C\. Zhao, C\. Deng, C\. Hou, C\. Xu, C\. Shao, C\. Ruan, C\. Sun, D\. Dai, D\. Guo, D\. Yang, D\. Chen, D\. Li, D\. Ji, E\. Li, F\. Wei, F\. Lin, F\. Yuan, F\. Xia, F\. Dai, G\. Hao, G\. Chen, G\. Cao, G\. Meng, G\. Li, H\. Yu, H\. Zhang, H\. Xu, H\. Li, H\. Liang, H\. Zhang, H\. Luo, H\. Wei, H\. Yuan, H\. Zhang, H\. Luo, H\. Chen, H\. Ji, H\. Zhang, H\. Ding, H\. Tang, H\. Cao, H\. Gao, H\. Qu, H\. Zeng, J\. Yang, J\.\-Q\. Zhu, J\. Luo, J\. Song, J\. Yu, J\. Huang, J\. Cai, J\. Liang, J\. Zhou, J\. Ye, J\. Li, J\. Xu, J\. Hu, J\. Yang, J\. Chen, J\. Yan, J\. Chen, J\. Zhou, J\. Xiang, J\. Yuan, J\. Cheng, J\. Zhou, J\. Zhu, J\. Yu, J\. Sun, J\. Ran, J\. Jiang, J\. Qiu, J\. Li, J\. Zheng, J\. Song, K\. Dong, K\. Gao, K\. Guan, K\. Zhou, K\. Huang, K\. Yu, L\. Wang, L\. Zhang, L\. Wang, L\. Xia, L\. Zhang, L\. Zhao, L\. Guo, L\. Luo, L\. Ma, L\. Zhu, L\. Wang, L\. Cai, L\. Zhang, L\. Chen, M\. Di, M\. Xu, M\. Mei, M\. Wang, M\. Zhang, M\. Zhang, M\. Tang, M\. Li, M\. Zhou, M\. Han, N\. Wang, P\. Huang, P\. Wang, P\. Cong, P\. Wang, P\. Zhang, Q\. Wang, Q\. Zhu, Q\. Li, Q\. Chen, Q\. Du, Q\. Jiang, R\. Tian, R\. Xu, R\. Lu, R\. Xu, R\. Ge, R\. Zhang, R\. Pan, R\. Wang, R\. D\. Chen, R\. Yin, R\. Xu, R\. Shen, R\. Zhang, R\. Chen, Sh\. Liu, S\. Lu, S\. Sun, S\. Zhou, S\. Chen, S\. Cai, S\. Nie, S\. Wu, S\. Chen, S\. Hu, S\. Liu, S\. Hu, S\. Ma, S\. Wang, S\. Yu, S\. Zhou, S\. Pan, S\. Yu, S\. Zhou, T\. Ni, T\. Yun, T\. Jin, T\. Pei, T\. Ye, T\. Lin, T\. Ji, T\. Cui, T\. Yue, T\. Yu, T\. Wang, W\. Y\. Zhang, W\. Xiao, W\. Zeng, W\. An, W\. Zhao, W\. Liu, W\. Liang, W\. Pang, W\. Luo, W\. Yao, W\. Gao, W\. Yang, W\. Huang, W\. Hou, W\. Zhang, W\. Ma, X\. Gao, X\. He, X\. Wang, X\. Wang, X\. Bi, X\. Liu, X\. Wang, X\. Chen, X\. Zhang, X\. Nie, X\. Sun, X\. Wang, X\. Cheng, X\. Liu, X\. Xie, X\. Liu, X\. Liu, X\. Yu, X\. Li, X\. Yang, X\. Zhang, X\. Chen, X\. Wang, X\. Su, X\. Chen, X\. Lin, X\. Fu, Y\. Yan, Yq\. Wang, Y\. Ma, Y\. Luo, Y\. Zhang, Y\. Xu, Y\. Ma, Y\. Huang, Y\. Li, Y\. Xu, Y\. Zhao, Y\. Sun, Y\. Wang, Y\. Qian, Y\. Shao, Y\. Yu, Y\. Zhang, Y\. Ding, Y\. Shi, Y\. Wu, Y\. Xiong, Y\. Ma, Y\. He, Y\. Tang, Y\. Zhou, Y\. Luo, Y\. Zhong, Y\. Piao, Y\. Wang, Y\. Zhang, Y\. Chen, Y\. Tan, Y\. Wei, Y\. Ma, Y\. Liu, Y\. Yang, Y\. Guo, Y\. Wu, Y\. Wu, Y\. Li, Y\. Cheng, Y\. Ou, Y\. Xu, Y\. Li, Y\. Wang, Y\. Yang, Y\. Xu, Y\. Wu, Y\. Meng, Y\. Zou, Y\. Zha, Y\. Xiong, Y\. Chen, Y\. Lin, Y\. Cao, Y\. Wang, Y\. Zhang, Y\. Yan, Y\. Lin, Y\. Gu, Y\. Luo, Y\. You, Y\. Liu, Y\. Zhou, Y\. Zhou, Y\. Huang, Z\. Wu, Z\. Wang, Z\. Zhao, Z\. Ren, Z\. Zhang, Z\. Sha, Z\. Fu, Z\. Ju, Z\. Xu, Z\. Xie, Z\. Zhang, Z\. Gao, Z\. Hao, Z\. Gou, Z\. Ma, Z\. Yan, Z\. Shao, Z\. Huang, Z\. Chen, Z\. Wu, Z\. Ren, Z\. Wu, Z\. Li, Z\. Zhang, Z\. Xu, Z\. Wang, Z\. Qu, Z\. Gu, Z\. Zhu, Z\. Li, Z\. Zhang, Z\. Xie, Z\. Gao, Z\. Wan, Z\. Pan, and Z\. YaoDeepSeek\-v4: towards highly efficient million\-token context intelligence\.Cited by:[§1](https://arxiv.org/html/2609.38832#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.38832#S2.SS2.p1.1)\.
- Duet al\.\(2026\)Y\. Du, P\. Harris, M\. Tian, E\. A\. Huerta, S\. Ronanki, S\. Rongali, A\.G\. Galstyan, and H\. PengRoPE distinguishes neither positions nor tokens in long contexts, provably\.ArXivabs/2605\.15514\.Cited by:[§1](https://arxiv.org/html/2609.38832#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.38832#S2.SS2.p3.1),[§3\.2](https://arxiv.org/html/2609.38832#S3.SS2.p7.2)\.
- Dubeyet al\.\(2024\)A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Yang, A\. Fan, A\. Goyal, A\. S\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Rozière, B\. M\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. A\. AlBadawy, E\. I\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Nail, G\. Mialon, G\. Pang, G\. Cu\-curell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. R\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Oldham, M\. Rita, M\. Pavlova, M\. H\. M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bash\-lykov, N\. Bogoychev, N\. S\. Chatterji, O\. Duchenne, O\. cCelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasić, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. S\. M\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. C\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. E\. Tan, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. K\. Singh, A\. Grattafiori, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Vaughan, A\. Baevski, A\. Fein\-stein, A\. Kallet, A\. Sangani, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Franco, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. R\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, P\. \(\. Huang, B\. Loyd, B\. de Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, D\. Civin, D\. Beaty, D\. Kreymer, S\. Li, D\. Wyatt, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smoth\-ers, F\. Sun, F\. Kreuk, F\. Tian, F\. Ozgenel, F\. Caggioni, F\. \(\. Guzmán, F\. J\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Thattai, G\. Herman, G\. Sizov, G\. Zhang, G\. Lakshminarayanan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. As\-pegren, H\. Goldman, I\. Molybog, I\. Tufanov, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. F\. Kohli, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, U\. KamHou, K\. Saxena, K\. Prasad, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Huang, K\. Chawla, K\. Lakhotia, K\. Huang, L\. Chen, L\. Garg, A\. Lavender, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Tsimpoukelli, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. L\. Seltzer, M\. Valko, M\. Re\-strepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. P\. Laptev, N\. Dong, N\. Zhang, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollár, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Maheswari, R\. Howes, R\. Rinott, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Yu\. Sidorov, S\. Pan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Feng, S\. Lin, S\. Zha, S\. Shankar, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. K\. Gupta, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Kohler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. O\. Ajayi, V\. Montanez, V\. Mohan, V\. Kumar, V\. Mangla, V\. Ionescu, V\. A\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wang, X\. Wu, X\. Wang, X\. Xia, X\. Wu, X\. Gao, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Y\. Wang, Y\. Hao, Y\. Qian, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, and Z\. ZhaoThe llama 3 herd of models\.Cited by:[§1](https://arxiv.org/html/2609.38832#S1.p1.1)\.
- Feduset al\.\(2021\)W\. Fedus, B\. Zoph, and N\. M\. ShazeerSwitch transformers: scaling to trillion parameter models with simple and efficient sparsity\.ArXivabs/2101\.03961\.Cited by:[§A\.2](https://arxiv.org/html/2609.38832#A1.SS2.p1.1),[§2\.1](https://arxiv.org/html/2609.38832#S2.SS1.p1.1),[§3\.2](https://arxiv.org/html/2609.38832#S3.SS2.p2.1),[§4\.4](https://arxiv.org/html/2609.38832#S4.SS4.p1.1),[Table 5](https://arxiv.org/html/2609.38832#S4.T5.2.1.4.1.1.1)\.
- Fuet al\.\(2026\)Z\. Fu, W\. Zeng, R\. Wang, and M\. LiAttention sink forges native moe in attention layers: sink\-aware training to address head collapse\.ArXivabs/2602\.01203\.Cited by:[§1](https://arxiv.org/html/2609.38832#S1.p6.1),[§3\.2](https://arxiv.org/html/2609.38832#S3.SS2.p1.1),[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p3.1)\.
- Gevaet al\.\(2020\)M\. Geva, R\. Schuster, J\. Berant, and O\. LevyTransformer feed\-forward layers are key\-value memories\.ArXivabs/2012\.14913\.Cited by:[§1](https://arxiv.org/html/2609.38832#S1.p2.1)\.
- Guoet al\.\(2024\)T\. Guo, D\. Pai, Y\. Bai, J\. Jiao, M\. I\. Jordan, and S\. MeiActive\-dormant attention heads: mechanistically demystifying extreme\-token phenomena in llms\.ArXivabs/2410\.13835\.Cited by:[§1](https://arxiv.org/html/2609.38832#S1.p2.1)\.
- Hendryckset al\.\(2020\)D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. X\. Song, and J\. SteinhardtMeasuring massive multitask language understanding\.ArXivabs/2009\.03300\.Cited by:[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p4.1)\.
- Hoet al\.\(2020\)X\. Ho, A\. Nguyen, S\. Sugawara, and A\. AizawaConstructing a multi\-hop qa dataset for comprehensive evaluation of reasoning steps\.ArXivabs/2011\.01060\.Cited by:[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p4.1)\.
- Hoffmannet al\.\(2022\)J\. Hoffmann, S\. Borgeaud, A\. Mensch, E\. Buchatskaya, T\. Cai, E\. Rutherford, D\. de Las Casas, L\. A\. Hendricks, J\. Welbl, A\. Clark, T\. Hennigan, E\. Noland, K\. Millican, G\. van den Driessche, B\. Damoc, A\. Guy, S\. Osindero, K\. Simonyan, E\. Elsen, J\. W\. Rae, O\. Vinyals, and L\. SifreTraining compute\-optimal large language models\.ArXivabs/2203\.15556\.Cited by:[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p1.1)\.
- Huanget al\.\(2021\)L\. R\. Huang, S\. Cao, N\. N\. Parulian, H\. Ji, and L\. WangEfficient attentions for long document summarization\.InNorth American Chapter of the Association for Computational Linguistics,Cited by:[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p4.1)\.
- Jelassiet al\.\(2024\)S\. Jelassi, D\. Brandfonbrener, S\. M\. Kakade, and E\. MalachRepeat after me: transformers are better than state space models at copying\.InInternational Conference on Machine Learning,Cited by:[§1](https://arxiv.org/html/2609.38832#S1.p3.1)\.
- Jianget al\.\(2024\)A\. Q\. Jiang, A\. Sablayrolles, A\. Roux, A\. Mensch, B\. Savary, C\. Bamford, D\. S\. Chaplot, D\. de Las Casas, E\. B\. Hanna, F\. Bressand, G\. Lengyel, G\. Bour, G\. Lample, L\. R\. Lavaud, L\. Saulnier, M\. Lachaux, P\. Stock, S\. Subramanian, S\. Yang, S\. Antoniak, T\. L\. Scao, T\. Gervet, T\. Lavril, T\. Wang, T\. Lacroix, and W\. E\. SayedMixtral of experts\.ArXivabs/2401\.04088\.Cited by:[§1](https://arxiv.org/html/2609.38832#S1.p1.1)\.
- Jinet al\.\(2024\)P\. Jin, B\. Zhu, L\. Yuan, and S\. YanMoH: multi\-head attention as mixture\-of\-head attention\.ArXivabs/2410\.11842\.Cited by:[§B\.4](https://arxiv.org/html/2609.38832#A2.SS4.p1.1),[§1](https://arxiv.org/html/2609.38832#S1.p1.1),[§1](https://arxiv.org/html/2609.38832#S1.p6.1),[§2\.1](https://arxiv.org/html/2609.38832#S2.SS1.p2.1),[§3\.1](https://arxiv.org/html/2609.38832#S3.SS1.p2.1),[§3\.2](https://arxiv.org/html/2609.38832#S3.SS2.p4.1),[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p3.1)\.
- Joshiet al\.\(2017\)M\. Joshi, E\. Choi, D\. S\. Weld, and L\. ZettlemoyerTriviaQA: a large scale distantly supervised challenge dataset for reading comprehension\.ArXivabs/1705\.03551\.Cited by:[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p4.1)\.
- Kociskýet al\.\(2017\)T\. Kociský, J\. Schwarz, P\. Blunsom, C\. Dyer, K\. M\. Hermann, G\. Melis, and E\. GrefenstetteThe narrativeqa reading comprehension challenge\.Transactions of the Association for Computational Linguistics6,pp\. 317–328\.Cited by:[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p4.1)\.
- Li and Roth \(2002\)X\. Li and D\. RothLearning question classifiers\.InInternational Conference on Computational Linguistics,Cited by:[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p4.1)\.
- Luet al\.\(2025\)E\. Lu, Z\. Jiang, J\. Liu, Y\. Du, T\. Jiang, C\. Hong, S\. Liu, W\. He, E\. Yuan, Y\. Wang, Z\. Huang, H\. Yuan, S\. Xu, X\. Xu, G\. Lai, Y\. Chen, H\. Zheng, J\. Yan, J\. Su, Y\. Wu, N\. Y\. Zhang, Z\. Yang, X\. Zhou, M\. Zhang, and J\. QiuMoBA: mixture of block attention for long\-context llms\.ArXivabs/2502\.13189\.Cited by:[§B\.3](https://arxiv.org/html/2609.38832#A2.SS3.p1.1),[§B\.3](https://arxiv.org/html/2609.38832#A2.SS3.p2.2),[§1](https://arxiv.org/html/2609.38832#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.38832#S2.SS2.p1.1),[§3\.2](https://arxiv.org/html/2609.38832#S3.SS2.p3.1),[§3\.2](https://arxiv.org/html/2609.38832#S3.SS2.p4.1),[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p3.1),[§4\.2](https://arxiv.org/html/2609.38832#S4.SS2.p3.1),[§4\.2](https://arxiv.org/html/2609.38832#S4.SS2.p4.1)\.
- Mihaylovet al\.\(2018\)T\. Mihaylov, P\. Clark, T\. Khot, and A\. SabharwalCan a suit of armor conduct electricity? a new dataset for open book question answering\.InConference on Empirical Methods in Natural Language Processing,Cited by:[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p4.1)\.
- MiMo Team \(2026\)MiMo TeamMiMo\-v2\.5\.Note:[https://huggingface\.co/collections/XiaomiMiMo/mimo\-v25](https://huggingface.co/collections/XiaomiMiMo/mimo-v25)Cited by:[§1](https://arxiv.org/html/2609.38832#S1.p1.1)\.
- Olssonet al\.\(2022\)C\. Olsson, N\. Elhage, N\. Nanda, N\. Joseph, N\. Dassarma, T\. Henighan, B\. Mann, A\. Askell, Y\. Bai, A\. Chen, T\. Conerly, D\. Drain, D\. Ganguli, Z\. Hatfield\-Dodds, D\. Hernandez, S\. Johnston, A\. Jones, J\. Kernion, L\. Lovitt, K\. Ndousse, D\. Amodei, T\. B\. Brown, J\. Clark, J\. Kaplan, S\. McCandlish, and C\. OlahIn\-context learning and induction heads\.ArXivabs/2209\.11895\.Cited by:[§1](https://arxiv.org/html/2609.38832#S1.p2.1)\.
- Penedoet al\.\(2024\)G\. Penedo, H\. Kydlícek, L\. B\. Allal, A\. Lozhkov, M\. Mitchell, C\. Raffel, L\. von Werra, and T\. WolfThe fineweb datasets: decanting the web for the finest text data at scale\.ArXivabs/2406\.17557\.Cited by:[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p1.1)\.
- Penget al\.\(2020\)H\. Peng, R\. Schwartz, D\. Li, and N\. A\. SmithA mixture of h \- 1 heads is better than h heads\.ArXivabs/2005\.06537\.Cited by:[§2\.1](https://arxiv.org/html/2609.38832#S2.SS1.p1.1)\.
- Piekoset al\.\(2025\)P\. Piekos, R\. Csord’as, and J\. SchmidhuberMixture of sparse attention: content\-based learnable sparse attention via expert\-choice routing\.ArXivabs/2505\.00315\.External Links:[Link](https://api.semanticscholar.org/CorpusID:278237209)Cited by:[§2\.1](https://arxiv.org/html/2609.38832#S2.SS1.p2.1)\.
- Qiuet al\.\(2025\)Z\. Qiu, Z\. Wang, B\. Zheng, Z\. Huang, K\. Wen, S\. Yang, R\. Men, L\. Yu, F\. Huang, S\. Huang, D\. Liu, J\. Zhou, and J\. LinGated attention for large language models: non\-linearity, sparsity, and attention\-sink\-free\.ArXivabs/2505\.06708\.Cited by:[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p3.1)\.
- Sakaguchiet al\.\(2019\)K\. Sakaguchi, R\. L\. Bras, C\. Bhagavatula, and Y\. ChoiWinoGrande\.Communications of the ACM64,pp\. 99 – 106\.Cited by:[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p4.1)\.
- Shazeeret al\.\(2017\)N\. M\. Shazeer, A\. Mirhoseini, K\. Maziarz, A\. Davis, Q\. V\. Le, G\. E\. Hinton, and J\. DeanOutrageously large neural networks: the sparsely\-gated mixture\-of\-experts layer\.ArXivabs/1701\.06538\.Cited by:[§A\.1](https://arxiv.org/html/2609.38832#A1.SS1.p1.1),[§A\.1](https://arxiv.org/html/2609.38832#A1.SS1.p3.1),[§2\.1](https://arxiv.org/html/2609.38832#S2.SS1.p1.1),[§3\.2](https://arxiv.org/html/2609.38832#S3.SS2.p2.1),[§4\.4](https://arxiv.org/html/2609.38832#S4.SS4.p1.1),[Table 5](https://arxiv.org/html/2609.38832#S4.T5.2.1.3.1)\.
- Tanget al\.\(2024\)J\. Tang, Y\. Zhao, K\. Zhu, G\. Xiao, B\. Kasikci, and S\. HanQuest: query\-aware sparsity for efficient long\-context llm inference\.ArXivabs/2406\.10774\.Cited by:[§B\.3](https://arxiv.org/html/2609.38832#A2.SS3.p1.1),[§1](https://arxiv.org/html/2609.38832#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.38832#S2.SS2.p1.1),[§2\.2](https://arxiv.org/html/2609.38832#S2.SS2.p2.1),[§3\.2](https://arxiv.org/html/2609.38832#S3.SS2.p3.1),[§3\.2](https://arxiv.org/html/2609.38832#S3.SS2.p4.1)\.
- Tripathi and Kumar \(2026\)V\. Tripathi and A\. KumarGrouped query experts: mixture\-of\-experts on gqa self\-attention\.ArXivabs/2606\.20945\.Cited by:[§2\.1](https://arxiv.org/html/2609.38832#S2.SS1.p1.1)\.
- Vaswaniet al\.\(2017\)A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, L\. Kaiser, and I\. PolosukhinAttention is all you need\.InNeural Information Processing Systems,Cited by:[§B\.2](https://arxiv.org/html/2609.38832#A2.SS2.p1.1),[§3\.1](https://arxiv.org/html/2609.38832#S3.SS1.p1.1)\.
- Wanget al\.\(2024\)L\. Wang, H\. Gao, C\. Zhao, X\. Sun, and D\. DaiAuxiliary\-loss\-free load balancing strategy for mixture\-of\-experts\.ArXivabs/2408\.15664\.Cited by:[§A\.3](https://arxiv.org/html/2609.38832#A1.SS3.p1.1),[§3\.2](https://arxiv.org/html/2609.38832#S3.SS2.p2.1),[§4\.4](https://arxiv.org/html/2609.38832#S4.SS4.p1.1),[Table 5](https://arxiv.org/html/2609.38832#S4.T5.2.1.5.1)\.
- Wertheimeret al\.\(2026\)D\. Wertheimer, A\. Zhang, D\. Liu, P\. Yin, and N\. WangFrayed rope and long inputs: a geometric perspective\.ArXivabs/2603\.18017\.Cited by:[§1](https://arxiv.org/html/2609.38832#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.38832#S2.SS2.p3.1),[§3\.2](https://arxiv.org/html/2609.38832#S3.SS2.p7.2)\.
- Xiaoet al\.\(2023\)G\. Xiao, Y\. Tian, B\. Chen, S\. Han, and M\. LewisEfficient streaming language models with attention sinks\.ArXivabs/2309\.17453\.Cited by:[§2\.2](https://arxiv.org/html/2609.38832#S2.SS2.p2.1)\.
- Xuet al\.\(2026\)Y\. Xu, F\. Meng, F\. Jiang, Y\. Wang, R\. Zhou, J\. Wu, Z\. Pan, Z\. Wang, X\. Tang, W\. Pei, T\. Liu, D\. Yin, X\. Sun, and M\. ZhangHISA: efficient hierarchical indexing for fine\-grained sparse attention\.ArXivabs/2603\.28458\.Cited by:[§1](https://arxiv.org/html/2609.38832#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.38832#S2.SS2.p1.1)\.
- Yanget al\.\(2024a\)Q\. A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, G\. Dong, H\. Wei, H\. Lin, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Lin, K\. Dang, K\. Lu, K\. Bao, K\. Yang, L\. Yu, M\. Li, M\. Xue, P\. Zhang, Q\. Zhu, R\. Men, R\. Lin, T\. Li, T\. Xia, X\. Ren, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Cui, Z\. Zhang, Z\. Qiu, S\. Quan, and Z\. WangQwen2\.5 technical report\.ArXivabs/2412\.15115\.Cited by:[§1](https://arxiv.org/html/2609.38832#S1.p1.1)\.
- Yanget al\.\(2024b\)S\. Yang, J\. Kautz, and A\. HatamizadehGated delta networks: improving mamba2 with delta rule\.ArXivabs/2412\.06464\.Cited by:[§1](https://arxiv.org/html/2609.38832#S1.p3.1)\.
- Yanget al\.\(2025\)Y\. Yang, C\. Wang, and J\. LiUMoE: unifying attention and ffn with shared experts\.ArXivabs/2505\.07260\.Cited by:[§1](https://arxiv.org/html/2609.38832#S1.p1.1)\.
- Yanget al\.\(2018\)Z\. Yang, P\. Qi, S\. Zhang, Y\. Bengio, W\. W\. Cohen, R\. Salakhutdinov, and C\. D\. ManningHotpotQA: a dataset for diverse, explainable multi\-hop question answering\.InConference on Empirical Methods in Natural Language Processing,Cited by:[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p4.1)\.
- Yuanet al\.\(2025\)J\. Yuan, H\. Gao, D\. Dai, J\. Luo, L\. Zhao, Z\. Zhang, Z\. Xie, Y\. X\. Wei, L\. Wang, Z\. Xiao, Y\. Wang, C\. Ruan, M\. Zhang, W\. Liang, and W\. ZengNative sparse attention: hardware\-aligned and natively trainable sparse attention\.InAnnual Meeting of the Association for Computational Linguistics,Cited by:[§B\.6](https://arxiv.org/html/2609.38832#A2.SS6.p1.2),[§B\.6](https://arxiv.org/html/2609.38832#A2.SS6.p3.1),[§1](https://arxiv.org/html/2609.38832#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.38832#S2.SS2.p1.1),[§3\.2](https://arxiv.org/html/2609.38832#S3.SS2.p3.1),[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p3.1),[§4\.2](https://arxiv.org/html/2609.38832#S4.SS2.p4.1)\.
- Zellerset al\.\(2019\)R\. Zellers, A\. Holtzman, Y\. Bisk, A\. Farhadi, and Y\. ChoiHellaSwag: can a machine really finish your sentence?\.InAnnual Meeting of the Association for Computational Linguistics,Cited by:[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p4.1)\.
- Zhanget al\.\(2022\)X\. Zhang, Y\. Shen, Z\. Huang, J\. Zhou, W\. Rong, and Z\. XiongMixture of attention heads: selecting attention heads per token\.InConference on Empirical Methods in Natural Language Processing,Cited by:[§1](https://arxiv.org/html/2609.38832#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.38832#S2.SS1.p1.1)\.
- Zhanget al\.\(2023\)Z\. \(\. Zhang, Y\. Sheng, T\. Zhou, T\. Chen, L\. Zheng, R\. Cai, Z\. Song, Y\. Tian, C\. Ré, C\. W\. Barrett, Z\. Wang, and B\. ChenH2O: heavy\-hitter oracle for efficient generative inference of large language models\.ArXivabs/2306\.14048\.Cited by:[§2\.2](https://arxiv.org/html/2609.38832#S2.SS2.p2.1)\.
- Zhonget al\.\(2021\)M\. Zhong, D\. Yin, T\. Yu, A\. Z\. Zaidi, M\. Mutuma, R\. Jha, A\. H\. Awadallah, A\. Celikyilmaz, Y\. Liu, X\. Qiu, and D\. R\. RadevQMSum: a new benchmark for query\-based multi\-domain meeting summarization\.InNorth American Chapter of the Association for Computational Linguistics,Cited by:[§4\.1](https://arxiv.org/html/2609.38832#S4.SS1.p4.1)\.
## Appendix ALoad Balancing Strategies
We describe three balancing strategies for a single routed attention layer\. For a training batchℬ\\mathcal\{B\}containingNNtokens, each token selectsKKofHHheads\. We use the head affinitiesat,ia\_\{t,i\}, selection indicatorsmt,im\_\{t,i\}, assignment fractionsfif\_\{i\}, and average affinitiespip\_\{i\}defined in the main text\. Auxiliary losses are summed across routed attention layers and weighted byλbal\\lambda\_\{\\mathrm\{bal\}\}\. Shared heads are excluded from all balancing statistics\.
### A\.1CV\-Based Load Balancing
Following[Shazeer et al\. \(2017\)](https://arxiv.org/html/2609.38832#bib.bib34), CV\-based balancing penalizes variation in both routing importance and expected token load\. For a nonnegative head\-statistic vector𝒖=\(u0,…,uH−1\)\\boldsymbol\{u\}=\(u\_\{0\},\\ldots,u\_\{H\-1\}\)with positive meanu¯\\bar\{u\}, the squared coefficient of variation is
CV\(𝒖\)2=H−1∑i=0H−1\(ui−u¯\)2u¯2,u¯=1H∑i=0H−1ui\.\\operatorname\{CV\}\(\\boldsymbol\{u\}\)^\{2\}=\\frac\{H^\{\-1\}\\sum\_\{i=0\}^\{H\-1\}\(u\_\{i\}\-\\bar\{u\}\)^\{2\}\}\{\\bar\{u\}^\{2\}\},\\qquad\\bar\{u\}=\\frac\{1\}\{H\}\\sum\_\{i=0\}^\{H\-1\}u\_\{i\}\.\(8\)This measures variance relative to the squared mean and is minimized when all entries are equal\.
For headii, define its routing importanceIiI\_\{i\}and expected assignment countℓi\\ell\_\{i\}as
Ii=∑t∈ℬat,i=Npi,ℓi=∑t∈ℬPrnoise\(i∈𝒮Knoise\(t\)\),I\_\{i\}=\\sum\_\{t\\in\\mathcal\{B\}\}a\_\{t,i\}=Np\_\{i\},\\qquad\\ell\_\{i\}=\\sum\_\{t\\in\\mathcal\{B\}\}\\Pr\_\{\\mathrm\{noise\}\}\\\!\\left\(i\\in\\mathcal\{S\}\_\{K\}^\{\\mathrm\{noise\}\}\(t\)\\right\),\(9\)where𝒮Knoise\(t\)\\mathcal\{S\}\_\{K\}^\{\\mathrm\{noise\}\}\(t\)denotes the heads selected by noisy Top\-KKrouting, and the probability is taken over the routing noise\. The importance term measures total affinity, while the expected\-load term measures how often a head is selected under noisy routing\. Let𝑰=\(I0,…,IH−1\)\\boldsymbol\{I\}=\(I\_\{0\},\\ldots,I\_\{H\-1\}\)andℓ=\(ℓ0,…,ℓH−1\)\\boldsymbol\{\\ell\}=\(\\ell\_\{0\},\\ldots,\\ell\_\{H\-1\}\)\. The combined loss is
ℒCV=αimpCV\(𝑰\)2\+αloadCV\(ℓ\)2,\\mathcal\{L\}\_\{\\mathrm\{CV\}\}=\\alpha\_\{\\mathrm\{imp\}\}\\operatorname\{CV\}\(\\boldsymbol\{I\}\)^\{2\}\+\\alpha\_\{\\mathrm\{load\}\}\\operatorname\{CV\}\(\\boldsymbol\{\\ell\}\)^\{2\},\(10\)whereαimp,αload\>0\\alpha\_\{\\mathrm\{imp\}\},\\alpha\_\{\\mathrm\{load\}\}\>0set the relative weights of the two penalties\. Their overall strength is controlled byλbal\\lambda\_\{\\mathrm\{bal\}\}\.
The importance penalty is equivalent toH∑i\(pi−1/H\)2H\\sum\_\{i\}\(p\_\{i\}\-1/H\)^\{2\}and is differentiable through the head affinities\. The load penalty additionally requires a differentiable expected\-load estimator, obtained through noisy Top\-KKrouting\([Shazeer et al\., 2017](https://arxiv.org/html/2609.38832#bib.bib34)\)\. It is not computed by directly differentiating the hard assignment countsNKfiNKf\_\{i\}\. Thus, the full CV formulation balances both affinity mass and expected head utilization, rather than affinity mass alone\.
### A\.2Switch\-Style Auxiliary Balancing
The Switch\-style loss\([Fedus et al\., 2021](https://arxiv.org/html/2609.38832#bib.bib35)\)combines observed head utilization with differentiable routing affinities:
ℒfp=H∑i=0H−1fipi\.\\mathcal\{L\}\_\{fp\}=H\\sum\_\{i=0\}^\{H\-1\}f\_\{i\}p\_\{i\}\.\(11\)For Top\-KKhead routing,fif\_\{i\}is normalized by the total number of assignmentsNKNK, so both\{fi\}\\\{f\_\{i\}\\\}and\{pi\}\\\{p\_\{i\}\\\}sum to one\. The loss equals one under uniform assignments and affinities, independently ofHHandKK\.
During backpropagation,fif\_\{i\}is treated as constant, while gradients flow throughpip\_\{i\}\. Since∂ℒfp/∂pi=Hfi\\partial\\mathcal\{L\}\_\{fp\}/\\partial p\_\{i\}=Hf\_\{i\}, a head with a larger assignment fraction receives a stronger penalty on its average affinity\. The observed load therefore guides the router toward less\-used heads without requiring a differentiable estimate of assignment counts\.
This objective is also related to CV regularization\. For𝒇=\(f0,…,fH−1\)\\boldsymbol\{f\}=\(f\_\{0\},\\ldots,f\_\{H\-1\}\), whenpi≈fip\_\{i\}\\approx f\_\{i\},
ℒfp≈H∑i=0H−1fi2=1\+CV\(𝒇\)2\.\\mathcal\{L\}\_\{fp\}\\approx H\\sum\_\{i=0\}^\{H\-1\}f\_\{i\}^\{2\}=1\+\\operatorname\{CV\}\(\\boldsymbol\{f\}\)^\{2\}\.\(12\)Unlike direct count\-based CV regularization, the affinity factor provides a differentiable path to the router\. As with the CV strategy, an overly largeλbal\\lambda\_\{\\mathrm\{bal\}\}can make balancing gradients compete with the head specialization favored by the language\-modeling objective\.
### A\.3Auxiliary\-Loss\-Free Balancing
Following[Wang et al\. \(2024\)](https://arxiv.org/html/2609.38832#bib.bib36), loss\-free balancing controls head utilization through routing biases rather than an auxiliary objective\. Each head maintains a scalar biasbib\_\{i\}, initialized to zero\. Selection uses biased affinities, while output gates retain the original affinities:
𝒮K\(t\)=TopKi\(at,i\+bi\),mt,i=𝟏\{i∈𝒮K\(t\)\},gt,i=mt,iat,i\.\\mathcal\{S\}\_\{K\}\(t\)=\\operatorname\{TopK\}\_\{i\}\(a\_\{t,i\}\+b\_\{i\}\),\\qquad m\_\{t,i\}=\\mathbf\{1\}\\\{i\\in\\mathcal\{S\}\_\{K\}\(t\)\\\},\\qquad g\_\{t,i\}=m\_\{t,i\}a\_\{t,i\}\.\(13\)Thus, biases change which heads are selected but do not enter their gating weights\. Selected affinities are not renormalized\.
Letcic\_\{i\}denote the number of token assignments received by headiiin the current batch, and letc¯\\bar\{c\}be the average count across heads\. After each completed training batch, we update
ci=∑t∈ℬmt,i=NKfi,c¯=NKH,bi←bi\+ηsign\(c¯−ci\),c\_\{i\}=\\sum\_\{t\\in\\mathcal\{B\}\}m\_\{t,i\}=NKf\_\{i\},\\qquad\\bar\{c\}=\\frac\{NK\}\{H\},\\qquad b\_\{i\}\\leftarrow b\_\{i\}\+\\eta\\,\\operatorname\{sign\}\(\\bar\{c\}\-c\_\{i\}\),\(14\)whereη\>0\\eta\>0is the update rate andsign\(0\)=0\\operatorname\{sign\}\(0\)=0\. An overloaded head receives a negative bias adjustment, reducing its chance of future selection\. An underloaded head receives a positive adjustment\. This forms a feedback loop from observed assignments to subsequent routing decisions\.
Bias updates occur outside backpropagation\. In NAMOH, biases remain fixed within each training batch and during inference, so updates do not revise earlier token assignments\. The router learns token\-to\-head affinities through the language\-modeling objective, while the biases regulate head utilization\. Unlike the two auxiliary\-loss strategies, this method introduces no direct load\-balancing gradient into the router\.
## Appendix BComplexity Analysis and Comparisons
### B\.1Setup and Accounting
We analyze one layer withHHquery heads, model widthdm=dmodeld\_\{m\}=d\_\{\\mathrm\{model\}\}, head widthdh=dheadd\_\{h\}=d\_\{\\mathrm\{head\}\}, and a prefix ofT≥1T\\geq 1tokens\. Routed variants activateKKquery heads per token, with1≤K≤H1\\leq K\\leq H\. GQA placesggquery heads in each KV group, whereg\|Hg\\mid H; group\-routed variants additionally requireg\|Kg\\mid K\. For ungrouped variants,g=1g=1\. SA uses an integer block sizeB≥1B\\geq 1and a retained fractionα∈\(0,1\]\\alpha\\in\(0,1\]of each available local history\. All widths are held fixed; we do not requiredm=Hdhd\_\{m\}=Hd\_\{h\}\.
LetPtotalP\_\{\\mathrm\{total\}\}andPactiveP\_\{\\mathrm\{active\}\}count all learned weights and those evaluated for one token\. We use a single bias\-free linear router for MoH and NAMOH to isolate their attention differences\. AnHH\-way head router containsHdmHd\_\{m\}weights, while anH/gH/g\-way group router containsHdm/gHd\_\{m\}/g\. All router weights are active for every token\. Projection biases, always\-active shared heads, and extra attention branches are excluded\. The metadata\-based block selectors considered here introduce no learned weights\.
As in the main text,MKVM\_\{\\mathrm\{KV\}\}counts stored KV pairs, andAKV\(T\)A\_\{\\mathrm\{KV\}\}\(T\)counts distinct historical pairs used by the next token\. Multiplying either count by2dh2d\_\{h\}gives its scalar data volume\. LetCattnC\_\{\\mathrm\{attn\}\}count all causal query\-key interactions during prefill, including self\-attention\. Unlike KV activation, this count includes separate interactions for query heads that share the same KV pair\. All counts are logical and exclude padding, temporary buffers, and training activations\.
### B\.2MHA and GQA
MHA\.Each head has independent query, key, value, and output projections\([Vaswani et al\., 2017](https://arxiv.org/html/2609.38832#bib.bib8)\)\. Each projection contributesdmdhd\_\{m\}d\_\{h\}weights, so
Ptotal=Pactive=4Hdmdh,MKV=HT,AKV\(T\)=HT\.P\_\{\\mathrm\{total\}\}=P\_\{\\mathrm\{active\}\}=4Hd\_\{m\}d\_\{h\},\\qquad M\_\{\\mathrm\{KV\}\}=HT,\\qquad A\_\{\\mathrm\{KV\}\}\(T\)=HT\.\(15\)A query at one\-based causal rankrrattends torrentries per head\. Summing over all ranks gives
Cattn=H∑r=1Tr=HT\(T\+1\)2\.C\_\{\\mathrm\{attn\}\}=H\\sum\_\{r=1\}^\{T\}r=\\frac\{HT\(T\+1\)\}\{2\}\.\(16\)
GQA\.Let𝒬j\\mathcal\{Q\}\_\{j\}denote theggquery heads assigned to KV groupjj, forj=0,…,H/g−1j=0,\\ldots,H/g\-1\. GQA retains separate query and output projections but shares key and value projections within each group\([Ainslie et al\., 2023](https://arxiv.org/html/2609.38832#bib.bib9)\)\. Therefore,
Ptotal=Pactive\\displaystyle P\_\{\\mathrm\{total\}\}=P\_\{\\mathrm\{active\}\}=2Hdmdh\+2Hgdmdh=2H\(1\+1g\)dmdh,\\displaystyle=2Hd\_\{m\}d\_\{h\}\+\\frac\{2H\}\{g\}d\_\{m\}d\_\{h\}=2H\\left\(1\+\\frac\{1\}\{g\}\\right\)d\_\{m\}d\_\{h\},\(17\)MKV\\displaystyle M\_\{\\mathrm\{KV\}\}=HTg,AKV\(T\)=HTg\.\\displaystyle=\\frac\{HT\}\{g\},\\qquad A\_\{\\mathrm\{KV\}\}\(T\)=\\frac\{HT\}\{g\}\.Each stored key still interacts withggqueries at each causal position:
Cattn=∑j=0H/g−1∑r=1Tgr=HT\(T\+1\)2\.C\_\{\\mathrm\{attn\}\}=\\sum\_\{j=0\}^\{H/g\-1\}\\sum\_\{r=1\}^\{T\}gr=\\frac\{HT\(T\+1\)\}\{2\}\.\(18\)Thus, KV sharing reduces storage and distinct access, but not the query\-key interaction count at fixedHH\.
### B\.3Block\-Selected Sparse Attention
Selection model\.We analyze a fixed\-fraction version of block selection that scans historical metadata, using blocks ofBBentries\. MoBA scores mean\-pooled keys, while Quest uses channelwise key minima and maxima\([Lu et al\., 2025](https://arxiv.org/html/2609.38832#bib.bib11);[Tang et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib10)\)\. We apply the specified selector during both prefill and decoding; this does not reproduce every cited method’s execution settings\.
For a nonnegative history lengthℓ\\ell, letNB\(ℓ\)N\_\{B\}\(\\ell\)be its number of blocks\. At local causal rankr≥1r\\geq 1, letRα,B\(r\)R\_\{\\alpha,B\}\(r\)be the selected block count, including the current block:
NB\(ℓ\)=⌈ℓB⌉,Rα,B\(r\)=⌈αNB\(r\)⌉\.N\_\{B\}\(\\ell\)=\\left\\lceil\\frac\{\\ell\}\{B\}\\right\\rceil,\\qquad R\_\{\\alpha,B\}\(r\)=\\left\\lceil\\alpha N\_\{B\}\(r\)\\right\\rceil\.\(19\)For example,α=0\.2\\alpha=0\.2targets approximately20%20\\%of each local history\. We always include the current block and select the remaining blocks from completed historical blocks\. Future blocks are excluded\. The current block uses a causal mask and is not ranked using metadata that contains future keys, following the causality rule of MoBA\([Lu et al\., 2025](https://arxiv.org/html/2609.38832#bib.bib11)\)\.
Exact selected\-entry count\.Letaα,B\(r\)a\_\{\\alpha,B\}\(r\)count entries attended by a query at local rankrr, including itself\. Each selected historical block containsBBentries, and the current block contributesr−B\[NB\(r\)−1\]r\-B\[N\_\{B\}\(r\)\-1\]visible entries\. Hence,
aα,B\(r\)\\displaystyle a\_\{\\alpha,B\}\(r\)=B\[Rα,B\(r\)−1\]\+r−B\[NB\(r\)−1\]\\displaystyle=B\[R\_\{\\alpha,B\}\(r\)\-1\]\+r\-B\[N\_\{B\}\(r\)\-1\]\(20\)=r−B\[NB\(r\)−Rα,B\(r\)\]\.\\displaystyle=r\-B\[N\_\{B\}\(r\)\-R\_\{\\alpha,B\}\(r\)\]\.Writingεr=aα,B\(r\)−αr\\varepsilon\_\{r\}=a\_\{\\alpha,B\}\(r\)\-\\alpha rfor the rounding error gives
εr\\displaystyle\\varepsilon\_\{r\}=B\[Rα,B\(r\)−αNB\(r\)\]\+\(1−α\)\[r−BNB\(r\)\],\\displaystyle=B\[R\_\{\\alpha,B\}\(r\)\-\\alpha N\_\{B\}\(r\)\]\+\(1\-\\alpha\)\[r\-BN\_\{B\}\(r\)\],\(21\)\|εr\|\\displaystyle\|\\varepsilon\_\{r\}\|≤B\.\\displaystyle\\leq B\.Thus,aα,B\(r\)=αr\+𝒪\(B\)a\_\{\\alpha,B\}\(r\)=\\alpha r\+\\mathcal\{O\}\(B\)\. The relative approximation is useful whenαr≫B\\alpha r\\gg B; a history contained in one block is attended densely\.
SA retains every KV pair because a later query may select a previously unused block\. Its weights are unchanged from MHA, and
MKV\\displaystyle M\_\{\\mathrm\{KV\}\}=HT,\\displaystyle=HT,\(22\)AKV\(T\)\\displaystyle A\_\{\\mathrm\{KV\}\}\(T\)=H\[aα,B\(T\+1\)−1\]=αHT\+𝒪\(HB\)\.\\displaystyle=H\[a\_\{\\alpha,B\}\(T\+1\)\-1\]=\\alpha HT\+\\mathcal\{O\}\(HB\)\.The subtraction removes the next token’s self\-entry\.
Causal prefill interactions\.Define the selected causal sum for a local sequence of lengthℓ\\ellas
Φα,B\(ℓ\)\\displaystyle\\Phi\_\{\\alpha,B\}\(\\ell\)=∑r=1ℓaα,B\(r\)\\displaystyle=\\sum\_\{r=1\}^\{\\ell\}a\_\{\\alpha,B\}\(r\)\(23\)=ℓ\(ℓ\+1\)2−B∑r=1ℓ\[NB\(r\)−Rα,B\(r\)\]\\displaystyle=\\frac\{\\ell\(\\ell\+1\)\}\{2\}\-B\\sum\_\{r=1\}^\{\\ell\}\[N\_\{B\}\(r\)\-R\_\{\\alpha,B\}\(r\)\]=αℓ\(ℓ\+1\)2\+∑r=1ℓεr\\displaystyle=\\frac\{\\alpha\\ell\(\\ell\+1\)\}\{2\}\+\\sum\_\{r=1\}^\{\\ell\}\\varepsilon\_\{r\}=αℓ\(ℓ\+1\)2\+𝒪\(Bℓ\)\.\\displaystyle=\\frac\{\\alpha\\ell\(\\ell\+1\)\}\{2\}\+\\mathcal\{O\}\(B\\ell\)\.Self\-attention and the mandatory current block are included exactly inΦα,B\\Phi\_\{\\alpha,B\}\. Applying this sum to allHHheads yields
Cattn=HΦα,B\(T\)=αHT22\+𝒪\(HBT\)\.C\_\{\\mathrm\{attn\}\}=H\\Phi\_\{\\alpha,B\}\(T\)=\\frac\{\\alpha HT^\{2\}\}\{2\}\+\\mathcal\{O\}\(HBT\)\.\(24\)
Metadata storage and scanning\.Letν\\nube the number of metadata vectors per block, each of widthdhd\_\{h\}\. Mean pooling usesν=1\\nu=1, and minima and maxima useν=2\\nu=2\. The metadata storage in scalars isMmeta=νdhHNB\(T\)M\_\{\\mathrm\{meta\}\}=\\nu d\_\{h\}HN\_\{B\}\(T\)\. Construction costs𝒪\(HTdh\)\\mathcal\{O\}\(HTd\_\{h\}\)and supports incremental updates\.
LetDmeta\(T\)D\_\{\\mathrm\{meta\}\}\(T\)andCmetaC\_\{\\mathrm\{meta\}\}count candidate block summaries inspected during the next decoding step and throughout prefill, respectively\. A query at rankrrscansNB\(r\)−1N\_\{B\}\(r\)\-1completed blocks\. To sum these visits, define
ΓB\(ℓ\)=∑r=1ℓNB\(r\)\.\\Gamma\_\{B\}\(\\ell\)=\\sum\_\{r=1\}^\{\\ell\}N\_\{B\}\(r\)\.\(25\)Writeℓ=uB\+v\\ell=uB\+v, whereu=⌊ℓ/B⌋u=\\lfloor\\ell/B\\rfloorand0≤v<B0\\leq v<B\. Each complete block contributesBBqueries with the same block count, giving
ΓB\(ℓ\)\\displaystyle\\Gamma\_\{B\}\(\\ell\)=B∑b=1ub\+\(u\+1\)v\\displaystyle=B\\sum\_\{b=1\}^\{u\}b\+\(u\+1\)v\(26\)=Bu\(u\+1\)2\+\(u\+1\)v=ℓ22B\+𝒪\(ℓ\)\.\\displaystyle=\\frac\{Bu\(u\+1\)\}\{2\}\+\(u\+1\)v=\\frac\{\\ell^\{2\}\}\{2B\}\+\\mathcal\{O\}\(\\ell\)\.Here,bbindexes complete local blocks\. Consequently,
Dmeta\(T\)\\displaystyle D\_\{\\mathrm\{meta\}\}\(T\)=H\[NB\(T\+1\)−1\]=H⌊TB⌋,\\displaystyle=H\[N\_\{B\}\(T\+1\)\-1\]=H\\left\\lfloor\\frac\{T\}\{B\}\\right\\rfloor,\(27\)Cmeta\\displaystyle C\_\{\\mathrm\{meta\}\}=H\[ΓB\(T\)−T\]=HT22B\+𝒪\(HT\)\.\\displaystyle=H\[\\Gamma\_\{B\}\(T\)\-T\]=\\frac\{HT^\{2\}\}\{2B\}\+\\mathcal\{O\}\(HT\)\.The subtraction excludes the unscored current block\. These visits do not acquire a factor ofα\\alpha, since the selector scans all candidates before choosing blocks\.
### B\.4MoH
We use the dense\-projection MoH baseline: all query, key, value, and output projection weights are evaluated, but onlyKKheads perform attention\([Jin et al\., 2024](https://arxiv.org/html/2609.38832#bib.bib16)\)\. Every head stores every prefix token, including tokens for which its output was inactive\. Under the router convention in Section[B\.1](https://arxiv.org/html/2609.38832#A2.SS1),
Ptotal=Pactive\\displaystyle P\_\{\\mathrm\{total\}\}=P\_\{\\mathrm\{active\}\}=4Hdmdh\+Hdm,\\displaystyle=4Hd\_\{m\}d\_\{h\}\+Hd\_\{m\},\(28\)MKV\\displaystyle M\_\{\\mathrm\{KV\}\}=HT,AKV\(T\)=KT\.\\displaystyle=HT,\\qquad A\_\{\\mathrm\{KV\}\}\(T\)=KT\.At global causal rankrr, each of theKKselected heads attends to allrrentries\. Therefore,
Cattn=∑r=1TKr=KT\(T\+1\)2\.C\_\{\\mathrm\{attn\}\}=\\sum\_\{r=1\}^\{T\}Kr=\\frac\{KT\(T\+1\)\}\{2\}\.\(29\)Head selection reduces KV activation and attention interactions, but not projection activation or cache storage in this baseline\.
### B\.5NAMOH
Parameters and routed storage\.Routing precedes projection, so inactive token\-head pairs produce no queries, keys, or values\. Only selected output slices are evaluated:
Ptotal=4Hdmdh\+Hdm,Pactive=4Kdmdh\+Hdm\.P\_\{\\mathrm\{total\}\}=4Hd\_\{m\}d\_\{h\}\+Hd\_\{m\},\\qquad P\_\{\\mathrm\{active\}\}=4Kd\_\{m\}d\_\{h\}\+Hd\_\{m\}\.\(30\)Letmt,im\_\{t,i\}indicate whether tokenttselects headii, and let𝒮K\(t\)\\mathcal\{S\}\_\{K\}\(t\)be the selected set\. The head lengths satisfy
ni=∑t=0T−1mt,i,∑i=0H−1ni=∑t=0T−1∑i=0H−1mt,i=TK\.n\_\{i\}=\\sum\_\{t=0\}^\{T\-1\}m\_\{t,i\},\\qquad\\sum\_\{i=0\}^\{H\-1\}n\_\{i\}=\\sum\_\{t=0\}^\{T\-1\}\\sum\_\{i=0\}^\{H\-1\}m\_\{t,i\}=TK\.\(31\)Only these assignments create KV pairs\. Hence,
MKV=TK,AKV\(T\)=∑i∈𝒮K\(T\)ni\.M\_\{\\mathrm\{KV\}\}=TK,\\qquad A\_\{\\mathrm\{KV\}\}\(T\)=\\sum\_\{i\\in\\mathcal\{S\}\_\{K\}\(T\)\}n\_\{i\}\.\(32\)Letn¯=TK/H\\bar\{n\}=TK/Hdenote the mean head length\. Under equal lengths, activation isKn¯=TK2/HK\\bar\{n\}=TK^\{2\}/H\. Equal lengths require integern¯\\bar\{n\}; approximate balance gives the main\-text estimate\.
Causal interaction count\.Within headii, the query at local rankrrattends to the firstrrentries\. Thus,
Cattn\\displaystyle C\_\{\\mathrm\{attn\}\}=∑i=0H−1∑r=1nir=12∑i=0H−1ni\(ni\+1\)\\displaystyle=\\sum\_\{i=0\}^\{H\-1\}\\sum\_\{r=1\}^\{n\_\{i\}\}r=\\frac\{1\}\{2\}\\sum\_\{i=0\}^\{H\-1\}n\_\{i\}\(n\_\{i\}\+1\)\(33\)=12∑i=0H−1ni2\+TK2\.\\displaystyle=\\frac\{1\}\{2\}\\sum\_\{i=0\}^\{H\-1\}n\_\{i\}^\{2\}\+\\frac\{TK\}\{2\}\.For the load vector𝒏=\(n0,…,nH−1\)\\boldsymbol\{n\}=\(n\_\{0\},\\ldots,n\_\{H\-1\}\), define its squared coefficient of variation using the population variance:
CV\(𝒏\)2=H−1∑i=0H−1\(ni−n¯\)2n¯2\.\\operatorname\{CV\}\(\\boldsymbol\{n\}\)^\{2\}=\\frac\{H^\{\-1\}\\sum\_\{i=0\}^\{H\-1\}\(n\_\{i\}\-\\bar\{n\}\)^\{2\}\}\{\\bar\{n\}^\{2\}\}\.\(34\)Since∑i\(ni−n¯\)=0\\sum\_\{i\}\(n\_\{i\}\-\\bar\{n\}\)=0, expanding around the mean gives
∑i=0H−1ni2\\displaystyle\\sum\_\{i=0\}^\{H\-1\}n\_\{i\}^\{2\}=Hn¯2\+2n¯∑i=0H−1\(ni−n¯\)\+∑i=0H−1\(ni−n¯\)2\\displaystyle=H\\bar\{n\}^\{2\}\+2\\bar\{n\}\\sum\_\{i=0\}^\{H\-1\}\(n\_\{i\}\-\\bar\{n\}\)\+\\sum\_\{i=0\}^\{H\-1\}\(n\_\{i\}\-\\bar\{n\}\)^\{2\}\(35\)=Hn¯2\[1\+CV\(𝒏\)2\]\\displaystyle=H\\bar\{n\}^\{2\}\[1\+\\operatorname\{CV\}\(\\boldsymbol\{n\}\)^\{2\}\]=T2K2H\[1\+CV\(𝒏\)2\]\.\\displaystyle=\\frac\{T^\{2\}K^\{2\}\}\{H\}\[1\+\\operatorname\{CV\}\(\\boldsymbol\{n\}\)^\{2\}\]\.Substitution into Equation[33](https://arxiv.org/html/2609.38832#A2.E33)yields
Cattn=T2K22H\[1\+CV\(𝒏\)2\]\+TK2\.C\_\{\\mathrm\{attn\}\}=\\frac\{T^\{2\}K^\{2\}\}\{2H\}\[1\+\\operatorname\{CV\}\(\\boldsymbol\{n\}\)^\{2\}\]\+\\frac\{TK\}\{2\}\.\(36\)The causal sum includes self\-attention exactly; the second moment determines the leading cost\.
Balance and matched comparisons\.The inequalities∑ini2≥\(∑ini\)2/H\\sum\_\{i\}n\_\{i\}^\{2\}\\geq\(\\sum\_\{i\}n\_\{i\}\)^\{2\}/Handni2≤Tnin\_\{i\}^\{2\}\\leq Tn\_\{i\}imply
T2K22H\+TK2≤Cattn≤KT\(T\+1\)2\.\\frac\{T^\{2\}K^\{2\}\}\{2H\}\+\\frac\{TK\}\{2\}\\leq C\_\{\\mathrm\{attn\}\}\\leq\\frac\{KT\(T\+1\)\}\{2\}\.\(37\)Equal lengths attain the lower bound when feasible\. If every token selects the sameKKheads, the upper bound is attained and decoding activation becomesKTKT\. Storage remainsTKTKin both cases\. Batch\-level load balancing does not guarantee balanced histories within every prefix\.
For an all\-active baseline withJ∈\{H,K\}J\\in\\\{H,K\\\}heads, the balanced finite\-length prefill ratio is
CattnJT\(T\+1\)/2=TK2/H\+KJ\(T\+1\)⟶K2HJasT→∞\.\\frac\{C\_\{\\mathrm\{attn\}\}\}\{JT\(T\+1\)/2\}=\\frac\{TK^\{2\}/H\+K\}\{J\(T\+1\)\}\\longrightarrow\\frac\{K^\{2\}\}\{HJ\}\\quad\\text\{as \}T\\to\\infty\.\(38\)The choicesJ=HJ=HandJ=KJ=Kmatch total and active head\-projection parameters, respectively; router weights are additional\.
### B\.6GQA\+SA
Independent or shared selection\.Query heads sharing a KV group can select blocks independently\. Letℰt,i\\mathcal\{E\}\_\{t,i\}be the set of historical KV entries selected by query headiifor tokentt\. Distinct access within groupjjthen counts the union:
maxi∈𝒬j\|ℰT,i\|≤\|⋃i∈𝒬jℰT,i\|≤min\{T,∑i∈𝒬j\|ℰT,i\|\}\.\\max\_\{i\\in\\mathcal\{Q\}\_\{j\}\}\|\\mathcal\{E\}\_\{T,i\}\|\\leq\\left\|\\bigcup\_\{i\\in\\mathcal\{Q\}\_\{j\}\}\\mathcal\{E\}\_\{T,i\}\\right\|\\leq\\min\\\!\\left\\\{T,\\sum\_\{i\\in\\mathcal\{Q\}\_\{j\}\}\|\\mathcal\{E\}\_\{T,i\}\|\\right\\\}\.\(39\)Ignoring rounding, the union can range fromαT\\alpha Ttomin\(T,gαT\)\\min\(T,g\\alpha T\)entries\. We instead use one shared block selection per group, following the group\-consistent design of[Yuan et al\. \(2025\)](https://arxiv.org/html/2609.38832#bib.bib12)\. Queries share block indices, not their attention weights\.
Costs with shared indices\.Parameter counts remain those of GQA\. Each of itsH/gH/ggroups stores one full history and selects one block set:
MKV\\displaystyle M\_\{\\mathrm\{KV\}\}=HTg,\\displaystyle=\\frac\{HT\}\{g\},\(40\)AKV\(T\)\\displaystyle A\_\{\\mathrm\{KV\}\}\(T\)=Hg\[aα,B\(T\+1\)−1\]=αHTg\+𝒪\(HBg\)\.\\displaystyle=\\frac\{H\}\{g\}\[a\_\{\\alpha,B\}\(T\+1\)\-1\]=\\frac\{\\alpha HT\}\{g\}\+\\mathcal\{O\}\\\!\\left\(\\frac\{HB\}\{g\}\\right\)\.Allggqueries still use each selected entry, so
Cattn\\displaystyle C\_\{\\mathrm\{attn\}\}=HggΦα,B\(T\)=αHT22\+𝒪\(HBT\),\\displaystyle=\\frac\{H\}\{g\}\\,g\\,\\Phi\_\{\\alpha,B\}\(T\)=\\frac\{\\alpha HT^\{2\}\}\{2\}\+\\mathcal\{O\}\(HBT\),\(41\)Dmeta\(T\)\\displaystyle D\_\{\\mathrm\{meta\}\}\(T\)=Hg⌊TB⌋,\\displaystyle=\\frac\{H\}\{g\}\\left\\lfloor\\frac\{T\}\{B\}\\right\\rfloor,Cmeta\\displaystyle C\_\{\\mathrm\{meta\}\}=Hg\[ΓB\(T\)−T\]=HT22gB\+𝒪\(HTg\)\.\\displaystyle=\\frac\{H\}\{g\}\[\\Gamma\_\{B\}\(T\)\-T\]=\\frac\{HT^\{2\}\}\{2gB\}\+\\mathcal\{O\}\\\!\\left\(\\frac\{HT\}\{g\}\\right\)\.
Shared indices versus shared scoring\.Letχg\\chi\_\{g\}denote the number ofdhd\_\{h\}\-scale scoring operations used to form one group\-block score, withχ1=1\\chi\_\{1\}=1\. Aggregating separately computed scores from allggqueries givesχg=g\\chi\_\{g\}=g, as in the per\-head score aggregation of NSA\([Yuan et al\., 2025](https://arxiv.org/html/2609.38832#bib.bib12)\)\. A single dot product using a pooled group query instead givesχg=1\\chi\_\{g\}=1, but generally defines a different scoring rule\. Both produce one shared block set and have the same KV counts\. Their scoring arithmetic isΘ\(χgdhDmeta\)\\Theta\(\\chi\_\{g\}d\_\{h\}D\_\{\\mathrm\{meta\}\}\)during decoding andΘ\(χgdhCmeta\)\\Theta\(\\chi\_\{g\}d\_\{h\}C\_\{\\mathrm\{meta\}\}\)during prefill\. Shared block indices therefore do not by themselves imply shared scoring arithmetic\.
### B\.7GQA\+NAMOH
Query\-head routing\.One option selectsKKof theHHquery heads independently\. Letut,j=𝟏\{𝒮K\(t\)∩𝒬j≠∅\}u\_\{t,j\}=\\mathbf\{1\}\\\{\\mathcal\{S\}\_\{K\}\(t\)\\cap\\mathcal\{Q\}\_\{j\}\\neq\\emptyset\\\}indicate whether tokentttouches KV groupjj, where𝟏\\mathbf\{1\}is the indicator function\. The number of KV groups written by a token satisfies
⌈Kg⌉≤∑j=0H/g−1ut,j≤min\(K,Hg\)\.\\left\\lceil\\frac\{K\}\{g\}\\right\\rceil\\leq\\sum\_\{j=0\}^\{H/g\-1\}u\_\{t,j\}\\leq\\min\\\!\\left\(K,\\frac\{H\}\{g\}\\right\)\.\(42\)If a shared pair is stored whenever any query in its group is selected, thennj=∑t=0T−1ut,jn\_\{j\}=\\sum\_\{t=0\}^\{T\-1\}u\_\{t,j\}and
MKV=∑jnj,T⌈Kg⌉≤MKV≤Tmin\(K,Hg\)\.M\_\{\\mathrm\{KV\}\}=\\sum\_\{j\}n\_\{j\},\\qquad T\\left\\lceil\\frac\{K\}\{g\}\\right\\rceil\\leq M\_\{\\mathrm\{KV\}\}\\leq T\\min\\\!\\left\(K,\\frac\{H\}\{g\}\\right\)\.\(43\)Allowing active queries to read their groups’ union histories givesAKV\(T\)=∑juT,jnjA\_\{\\mathrm\{KV\}\}\(T\)=\\sum\_\{j\}u\_\{T,j\}n\_\{j\}\. Preserving separate query\-head histories instead requires membership masks and makes access depend on overlaps between current and historical assignments\. Neither storage nor access is determined byTT,HH,KK, andggalone\.
Adopted KV\-group routing\.We instead select exactlyK/gK/gof theH/gH/gKV groups\. Let𝒢t\\mathcal\{G\}\_\{t\}denote the selected group set for tokentt\. Allggquery heads in each selected group participate and share one routing gate\. The group writes one KV pair\. Thus, exactlyKKquery and output projections andK/gK/gkey and value projections are evaluated:
Ptotal\\displaystyle P\_\{\\mathrm\{total\}\}=2H\(1\+1g\)dmdh\+Hgdm,\\displaystyle=2H\\left\(1\+\\frac\{1\}\{g\}\\right\)d\_\{m\}d\_\{h\}\+\\frac\{H\}\{g\}d\_\{m\},\(44\)Pactive\\displaystyle P\_\{\\mathrm\{active\}\}=2K\(1\+1g\)dmdh\+Hgdm\.\\displaystyle=2K\\left\(1\+\\frac\{1\}\{g\}\\right\)d\_\{m\}d\_\{h\}\+\\frac\{H\}\{g\}d\_\{m\}\.For this routing rule, the group lengths satisfy
nj=∑t=0T−1𝟏\{j∈𝒢t\},∑j=0H/g−1nj=TKg,n¯=TK/gH/g=TKH\.n\_\{j\}=\\sum\_\{t=0\}^\{T\-1\}\\mathbf\{1\}\\\{j\\in\\mathcal\{G\}\_\{t\}\\\},\\qquad\\sum\_\{j=0\}^\{H/g\-1\}n\_\{j\}=\\frac\{TK\}\{g\},\\qquad\\bar\{n\}=\\frac\{TK/g\}\{H/g\}=\\frac\{TK\}\{H\}\.\(45\)The group activation fraction remainsK/HK/H, so the average history length isTK/HTK/H, notTK/\(Hg\)TK/\(Hg\)\. Consequently,
MKV=TKg,AKV\(T\)=∑j∈𝒢Tnj≈KgTKH=TK2Hg\.M\_\{\\mathrm\{KV\}\}=\\frac\{TK\}\{g\},\\qquad A\_\{\\mathrm\{KV\}\}\(T\)=\\sum\_\{j\\in\\mathcal\{G\}\_\{T\}\}n\_\{j\}\\approx\\frac\{K\}\{g\}\\frac\{TK\}\{H\}=\\frac\{TK^\{2\}\}\{Hg\}\.\(46\)
Causal prefill derivation\.For group loads𝒏=\(n0,…,nH/g−1\)\\boldsymbol\{n\}=\(n\_\{0\},\\ldots,n\_\{H/g\-1\}\), the squared coefficient of variation is
CV\(𝒏\)2=\(g/H\)∑j\(nj−n¯\)2n¯2\.\\operatorname\{CV\}\(\\boldsymbol\{n\}\)^\{2\}=\\frac\{\(g/H\)\\sum\_\{j\}\(n\_\{j\}\-\\bar\{n\}\)^\{2\}\}\{\\bar\{n\}^\{2\}\}\.\(47\)Expanding the second moment gives
∑jnj2\\displaystyle\\sum\_\{j\}n\_\{j\}^\{2\}=Hgn¯2\+∑j\(nj−n¯\)2\\displaystyle=\\frac\{H\}\{g\}\\bar\{n\}^\{2\}\+\\sum\_\{j\}\(n\_\{j\}\-\\bar\{n\}\)^\{2\}\(48\)=T2K2Hg\[1\+CV\(𝒏\)2\]\.\\displaystyle=\\frac\{T^\{2\}K^\{2\}\}\{Hg\}\[1\+\\operatorname\{CV\}\(\\boldsymbol\{n\}\)^\{2\}\]\.Each local rank now producesggqueries, so
Cattn\\displaystyle C\_\{\\mathrm\{attn\}\}=∑j∑r=1njgr=g2∑jnj2\+g2∑jnj\\displaystyle=\\sum\_\{j\}\\sum\_\{r=1\}^\{n\_\{j\}\}gr=\\frac\{g\}\{2\}\\sum\_\{j\}n\_\{j\}^\{2\}\+\\frac\{g\}\{2\}\\sum\_\{j\}n\_\{j\}\(49\)=T2K22H\[1\+CV\(𝒏\)2\]\+TK2\.\\displaystyle=\\frac\{T^\{2\}K^\{2\}\}\{2H\}\[1\+\\operatorname\{CV\}\(\\boldsymbol\{n\}\)^\{2\}\]\+\\frac\{TK\}\{2\}\.KV sharing therefore preserves NAMOH’s balanced interaction count while reducing distinct KV storage and access bygg\. When head\-relative RoPE is enabled, all queries and the shared key use the same group\-local position index\. Group routing also coarsens the routing choices and requiresg≤Kg\\leq K; a single KV group leaves no nontrivial group selection\.
### B\.8SA\+NAMOH
We first construct routed head histories, then form blocks in each head’s local token order\. The retained fractionα\\alphaapplies to that local history, not to the originalTT\-token prefix\. Parameter counts and KV storage remain those of NAMOH\. Using the head lengths from Section[B\.5](https://arxiv.org/html/2609.38832#A2.SS5),
AKV\(T\)\\displaystyle A\_\{\\mathrm\{KV\}\}\(T\)=∑i∈𝒮K\(T\)\[aα,B\(ni\+1\)−1\]\\displaystyle=\\sum\_\{i\\in\\mathcal\{S\}\_\{K\}\(T\)\}\[a\_\{\\alpha,B\}\(n\_\{i\}\+1\)\-1\]\(50\)=α∑i∈𝒮K\(T\)ni\+𝒪\(KB\)≈αTK2H\.\\displaystyle=\\alpha\\sum\_\{i\\in\\mathcal\{S\}\_\{K\}\(T\)\}n\_\{i\}\+\\mathcal\{O\}\(KB\)\\approx\\frac\{\\alpha TK^\{2\}\}\{H\}\.For prefill, apply the selected causal sum independently to every head:
Cattn\\displaystyle C\_\{\\mathrm\{attn\}\}=∑iΦα,B\(ni\)\\displaystyle=\\sum\_\{i\}\\Phi\_\{\\alpha,B\}\(n\_\{i\}\)\(51\)=α2∑ini\(ni\+1\)\+𝒪\(B∑ini\)\\displaystyle=\\frac\{\\alpha\}\{2\}\\sum\_\{i\}n\_\{i\}\(n\_\{i\}\+1\)\+\\mathcal\{O\}\\\!\\left\(B\\sum\_\{i\}n\_\{i\}\\right\)=αT2K22H\[1\+CV\(𝒏\)2\]\+𝒪\(BTK\)\.\\displaystyle=\\frac\{\\alpha T^\{2\}K^\{2\}\}\{2H\}\[1\+\\operatorname\{CV\}\(\\boldsymbol\{n\}\)^\{2\}\]\+\\mathcal\{O\}\(BTK\)\.The indexer scans only the selected routed histories during decoding:
Dmeta\(T\)\\displaystyle D\_\{\\mathrm\{meta\}\}\(T\)=∑i∈𝒮K\(T\)⌊niB⌋=1B∑i∈𝒮K\(T\)ni\+𝒪\(K\),\\displaystyle=\\sum\_\{i\\in\\mathcal\{S\}\_\{K\}\(T\)\}\\left\\lfloor\\frac\{n\_\{i\}\}\{B\}\\right\\rfloor=\\frac\{1\}\{B\}\\sum\_\{i\\in\\mathcal\{S\}\_\{K\}\(T\)\}n\_\{i\}\+\\mathcal\{O\}\(K\),\(52\)Cmeta\\displaystyle C\_\{\\mathrm\{meta\}\}=∑i\[ΓB\(ni\)−ni\]=12B∑ini2\+𝒪\(TK\)\\displaystyle=\\sum\_\{i\}\[\\Gamma\_\{B\}\(n\_\{i\}\)\-n\_\{i\}\]=\\frac\{1\}\{2B\}\\sum\_\{i\}n\_\{i\}^\{2\}\+\\mathcal\{O\}\(TK\)=T2K22HB\[1\+CV\(𝒏\)2\]\+𝒪\(TK\)\.\\displaystyle=\\frac\{T^\{2\}K^\{2\}\}\{2HB\}\[1\+\\operatorname\{CV\}\(\\boldsymbol\{n\}\)^\{2\}\]\+\\mathcal\{O\}\(TK\)\.Under balance, the decoding scan has leading termTK2/\(HB\)TK^\{2\}/\(HB\)\. Routing reduces the histories scanned by the selector; block selection then reduces token\-level reads within those histories\.
### B\.9GQA\+SA\+NAMOH
We combine KV\-group routing with group\-shared block selection\. Each token activatesK/gK/ggroups, and each active group selects blocks from its own routed history\. Allggqueries in that group use the same selected blocks\. Parameter counts remain those in Equation[44](https://arxiv.org/html/2609.38832#A2.E44), and
MKV\\displaystyle M\_\{\\mathrm\{KV\}\}=TKg,\\displaystyle=\\frac\{TK\}\{g\},\(53\)AKV\(T\)\\displaystyle A\_\{\\mathrm\{KV\}\}\(T\)=∑j∈𝒢T\[aα,B\(nj\+1\)−1\]\\displaystyle=\\sum\_\{j\\in\\mathcal\{G\}\_\{T\}\}\[a\_\{\\alpha,B\}\(n\_\{j\}\+1\)\-1\]=α∑j∈𝒢Tnj\+𝒪\(KBg\)≈αTK2Hg\.\\displaystyle=\\alpha\\sum\_\{j\\in\\mathcal\{G\}\_\{T\}\}n\_\{j\}\+\\mathcal\{O\}\\\!\\left\(\\frac\{KB\}\{g\}\\right\)\\approx\\frac\{\\alpha TK^\{2\}\}\{Hg\}\.Combining the local causal sum with the group second moment gives
Cattn\\displaystyle C\_\{\\mathrm\{attn\}\}=g∑jΦα,B\(nj\)\\displaystyle=g\\sum\_\{j\}\\Phi\_\{\\alpha,B\}\(n\_\{j\}\)\(54\)=αg2∑jnj\(nj\+1\)\+𝒪\(Bg∑jnj\)\\displaystyle=\\frac\{\\alpha g\}\{2\}\\sum\_\{j\}n\_\{j\}\(n\_\{j\}\+1\)\+\\mathcal\{O\}\\\!\\left\(Bg\\sum\_\{j\}n\_\{j\}\\right\)=αT2K22H\[1\+CV\(𝒏\)2\]\+𝒪\(BTK\)\.\\displaystyle=\\frac\{\\alpha T^\{2\}K^\{2\}\}\{2H\}\[1\+\\operatorname\{CV\}\(\\boldsymbol\{n\}\)^\{2\}\]\+\\mathcal\{O\}\(BTK\)\.The shared indexer visits each candidate block once per active group:
Dmeta\(T\)\\displaystyle D\_\{\\mathrm\{meta\}\}\(T\)=∑j∈𝒢T⌊njB⌋=1B∑j∈𝒢Tnj\+𝒪\(Kg\),\\displaystyle=\\sum\_\{j\\in\\mathcal\{G\}\_\{T\}\}\\left\\lfloor\\frac\{n\_\{j\}\}\{B\}\\right\\rfloor=\\frac\{1\}\{B\}\\sum\_\{j\\in\\mathcal\{G\}\_\{T\}\}n\_\{j\}\+\\mathcal\{O\}\\\!\\left\(\\frac\{K\}\{g\}\\right\),\(55\)Cmeta\\displaystyle C\_\{\\mathrm\{meta\}\}=∑j\[ΓB\(nj\)−nj\]=12B∑jnj2\+𝒪\(TKg\)\\displaystyle=\\sum\_\{j\}\[\\Gamma\_\{B\}\(n\_\{j\}\)\-n\_\{j\}\]=\\frac\{1\}\{2B\}\\sum\_\{j\}n\_\{j\}^\{2\}\+\\mathcal\{O\}\\\!\\left\(\\frac\{TK\}\{g\}\\right\)=T2K22HgB\[1\+CV\(𝒏\)2\]\+𝒪\(TKg\)\.\\displaystyle=\\frac\{T^\{2\}K^\{2\}\}\{2HgB\}\[1\+\\operatorname\{CV\}\(\\boldsymbol\{n\}\)^\{2\}\]\+\\mathcal\{O\}\\\!\\left\(\\frac\{TK\}\{g\}\\right\)\.Under balance, the decoding scan has leading termTK2/\(HgB\)TK^\{2\}/\(HgB\)\. Relative toHH\-head MHA, persistent KV storage isK/\(Hg\)K/\(Hg\)and activation is approximatelyαK2/\(H2g\)\\alpha K^\{2\}/\(H^\{2\}g\)of the baseline\. The factorα\\alphareduces activation, not persistent KV storage\.
### B\.10Overall Costs and Scaling
Persistent cache storage\.For any SA variant, index its KV histories byj=0,…,H/g−1j=0,\\ldots,H/g\-1, withg=1g=1for ungrouped attention\. The lengthsnjn\_\{j\}equalTTwithout routing and are assignment\-dependent otherwise\. Metadata contributes
Mmeta=νdh∑j⌈njB⌉≤νdh\(MKVB\+Hg\)M\_\{\\mathrm\{meta\}\}=\\nu d\_\{h\}\\sum\_\{j\}\\left\\lceil\\frac\{n\_\{j\}\}\{B\}\\right\\rceil\\leq\\nu d\_\{h\}\\left\(\\frac\{M\_\{\\mathrm\{KV\}\}\}\{B\}\+\\frac\{H\}\{g\}\\right\)\(56\)in scalars\. Total KV and metadata storage is2dhMKV\+Mmeta2d\_\{h\}M\_\{\\mathrm\{KV\}\}\+M\_\{\\mathrm\{meta\}\}\. For the three\-way combination, this is at most
2dhTKg\+νdh\(TKgB\+Hg\)\.\\frac\{2d\_\{h\}TK\}\{g\}\+\\nu d\_\{h\}\\left\(\\frac\{TK\}\{gB\}\+\\frac\{H\}\{g\}\\right\)\.\(57\)Block allocation can additionally leave fewer thanBBunused KV slots per nonempty history\. Metadata construction costs𝒪\(dhMKV\)\\mathcal\{O\}\(d\_\{h\}M\_\{\\mathrm\{KV\}\}\)over the prefix and is dominated by projection work in this model\.
Arithmetic accounting\.Lethacth\_\{\\mathrm\{act\}\}be the number of active query heads:HHwithout head routing andKKfor MoH and NAMOH variants\. With group\-shared access, the next token evaluatesgAKV\(T\)\+hactgA\_\{\\mathrm\{KV\}\}\(T\)\+h\_\{\\mathrm\{act\}\}query\-key interactions, including self\-attention\. LetFprefillF\_\{\\mathrm\{prefill\}\}andFdecodeF\_\{\\mathrm\{decode\}\}denote total arithmetic work over prefill and one decoding step\. Assuming linear\-work head and block selection,
Fprefill\\displaystyle F\_\{\\mathrm\{prefill\}\}=𝒪\(TPactive\+dh\[Cattn\+χgCmeta\]\),\\displaystyle=\\mathcal\{O\}\\\!\\left\(TP\_\{\\mathrm\{active\}\}\+d\_\{h\}\[C\_\{\\mathrm\{attn\}\}\+\\chi\_\{g\}C\_\{\\mathrm\{meta\}\}\]\\right\),\(58\)Fdecode\\displaystyle F\_\{\\mathrm\{decode\}\}=𝒪\(Pactive\+dh\[gAKV\(T\)\+hact\+χgDmeta\(T\)\]\)\.\\displaystyle=\\mathcal\{O\}\\\!\\left\(P\_\{\\mathrm\{active\}\}\+d\_\{h\}\[gA\_\{\\mathrm\{KV\}\}\(T\)\+h\_\{\\mathrm\{act\}\}\+\\chi\_\{g\}D\_\{\\mathrm\{meta\}\}\(T\)\]\\right\)\.SetCmeta=Dmeta=0C\_\{\\mathrm\{meta\}\}=D\_\{\\mathrm\{meta\}\}=0without block selection\. For ungrouped attention,χ1=1\\chi\_\{1\}=1\. Sorting all scores can add work beyond this linear\-selection model\.
Table 6:Leading attention and indexer terms\.The prefill column reportsCattn\+χgCmetaC\_\{\\mathrm\{attn\}\}\+\\chi\_\{g\}C\_\{\\mathrm\{meta\}\}; the decoding column reportsgAKV\(T\)\+χgDmeta\(T\)gA\_\{\\mathrm\{KV\}\}\(T\)\+\\chi\_\{g\}D\_\{\\mathrm\{meta\}\}\(T\)\. Insert these terms into Equation[58](https://arxiv.org/html/2609.38832#A2.E58)with the active weights in Table[1](https://arxiv.org/html/2609.38832#S3.T1)to obtain total arithmetic costs\. Routed histories are balanced, and finite\-length and block\-rounding terms are omitted here but retained in the derivations\.For the three\-way combination with balanced histories andαTK/H≫B\\alpha TK/H\\gg B, Equation[58](https://arxiv.org/html/2609.38832#A2.E58)becomes
Fprefill\\displaystyle F\_\{\\mathrm\{prefill\}\}=𝒪\(T\[2K\(1\+1g\)dmdh\+Hgdm\]\+dhT2K2H\[α\+χggB\]\),\\displaystyle=\\mathcal\{O\}\\\!\\left\(T\\left\[2K\\left\(1\+\\frac\{1\}\{g\}\\right\)d\_\{m\}d\_\{h\}\+\\frac\{H\}\{g\}d\_\{m\}\\right\]\+\\frac\{d\_\{h\}T^\{2\}K^\{2\}\}\{H\}\\left\[\\alpha\+\\frac\{\\chi\_\{g\}\}\{gB\}\\right\]\\right\),\(59\)Fdecode\\displaystyle F\_\{\\mathrm\{decode\}\}=𝒪\(2K\(1\+1g\)dmdh\+Hgdm\+dhTK2H\[α\+χggB\]\)\.\\displaystyle=\\mathcal\{O\}\\\!\\left\(2K\\left\(1\+\\frac\{1\}\{g\}\\right\)d\_\{m\}d\_\{h\}\+\\frac\{H\}\{g\}d\_\{m\}\+\\frac\{d\_\{h\}TK^\{2\}\}\{H\}\\left\[\\alpha\+\\frac\{\\chi\_\{g\}\}\{gB\}\\right\]\\right\)\.With per\-head score aggregation,χg=g\\chi\_\{g\}=g, so KV sharing does not reduce the leading indexer arithmetic bygg\. A pooled\-query indexer withχg=1\\chi\_\{g\}=1realizes that additional scoring reduction\. Neither case changes theggseparate attention computations per shared KV entry\.
Fractional versus fixed budgets\.The multiplicative activation formulas assume a fixed retained fraction of each local history\. They do not imply the same gain under a fixed absolute budget\. If at mostRRblocks are retained per active history andS=RBS=RBis the corresponding entry budget, the group\-routed hybrid instead satisfies
AKV\(T\)≤KgS,Cattn≤TKS\.A\_\{\\mathrm\{KV\}\}\(T\)\\leq\\frac\{K\}\{g\}S,\\qquad C\_\{\\mathrm\{attn\}\}\\leq TKS\.\(60\)Once histories exceed this budget, head routing does not multiply these saturated bounds by another factor ofK/HK/H\. It still reduces storage and the histories scanned by the indexer\.
At fixedα\\alpha,BB,HH,KK, andgg, both fractional attention and full metadata scanning remain quadratic in prefill lengthTT\. Their coefficients are reduced, not their asymptotic order\. Settingα=1\\alpha=1restores dense attention within each available local history, and block selection can then be bypassed\. These costs describe computational structure rather than guaranteed latency; physical gains also depend on cache reuse, data movement, and kernel scheduling\.Similar Articles
MiniMax Sparse Attention
MiniMax Sparse Attention introduces a blockwise sparse attention mechanism that achieves significant speedups for ultra-long-context LLMs, reducing per-token attention compute by 28.4x at 1M context with wall-clock speedups of 14.2x for prefill and 7.6x for decoding on H800 GPUs. The method is accompanied by an open-source inference kernel and a publicly released multimodal model.
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
Introduces HiLS Attention, a chunk-wise sparse attention mechanism for LLMs that learns chunk selection end-to-end via LM loss, achieving performance comparable to full attention while enabling ultra-long-context extrapolation and faster inference.
Asymmetric Attention Heads: Structured Head-Wise Context Allocation for Transformer Attention
This paper introduces Asymmetric Attention Heads (AAH), a framework that assigns different context windows to attention heads in transformers, with experiments showing improved language modeling performance.
MiniMax M3: How Sparse Attention Makes Long-Horizon Agents Practical (11 minute read)
MiniMax M3 uses sparse attention to make long-horizon agents practical by keeping context cost predictable and low, enabling 500K-token contexts with minimal quality loss and significant speedups in production.
@eliebakouch: the new sparse attention method introduced with this model is basically a combination of components from existing ones.…
Meituan introduces LongCat-2.0, a 1.6T parameter MoE model with 48B active parameters and 1M context length, featuring a new LongCat Sparse Attention (LSA) method that combines components from existing sparse attention techniques.