Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models
Summary
This paper analyzes the evolution of attention routing in recurrent language models and proposes WISE, a training-free inference method that reuses stabilized sparse attention support to achieve up to 1.76× attention speedup while preserving performance on multi-hop QA benchmarks.
View Cached Full Text
Cached at: 09/24/26, 09:21 AM
# Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models
Source: [https://arxiv.org/html/2609.27373](https://arxiv.org/html/2609.27373)
Ke WanAffiliation:Department of Computer ScienceAffiliation:University of VirginiaAffiliation:Charlottesville, VA, USAEmail:[tbn5pj@virginia\.edu](mailto:)Chen Chen††thanks:Corresponding author\.Affiliation:Department of Computer ScienceAffiliation:University of VirginiaAffiliation:Charlottesville, VA, USAEmail:[zrh6du@virginia\.edu](mailto:)
###### Abstract
Recurrent language models repeatedly apply shared network blocks to refine latent representations, enabling additional test\-time computation without generating explicit intermediate reasoning tokens\. We study how information routing evolves across recurrent depth and uncover a consistent separation in convergence timescales: routing\-related quantities, including attention support and attention distributions, stabilize substantially earlier than representation\-related quantities such as hidden states and attention outputs\. This suggests a two\-stage structure in recurrent inference, where the model first discovers a sparse working set of relevant context and then continues refining representations over largely the same routing support\. Motivated by this structure, we introduceWISE\(Working\-setInference withSupportExploitation\), a training\-free method that uses unrestricted global attention during early recurrent steps and later reuses a directly discovered block\-structured support while keeping recurrent refinement and within\-support attention computation dynamic\. Controlled interventions show that recurrent discovery is important for identifying effective working sets, and that reusing only the discovered support better preserves model behavior than more restrictive forms of late attention reuse\. Across multi\-hop QA benchmarks,WISElargely preserves full\-attention performance, while matched context\-scaling experiments reveal a favorable quality–efficiency tradeoff as routing support becomes increasingly sparse with context length\. A sparse\-attention implementation translates this structured sparsity into practical GPU acceleration, achieving up to a1\.76×1\.76\\timesattention speedup over native FlashAttention at 4K context, with further gains possible through tighter kernel co\-design\. These results show that recurrent inference can reuse the stabilized routing structure to reduce repeated global attention computation without retraining or changing model parameters\. Our code is available at[https://github\.com/tbn5pj/WISE\_code](https://github.com/tbn5pj/WISE_code)\.
## 1Introduction
Recurrent language models provide a new axis for test\-time computation by repeatedly applying shared network blocks to refine latent representations\([Geiping et al\., 2025](https://arxiv.org/html/2609.27373#bib.bib2);[Dehghani et al\., 2018](https://arxiv.org/html/2609.27373#bib.bib31);[Yang et al\., 2024](https://arxiv.org/html/2609.27373#bib.bib32);[Giannou et al\., 2023](https://arxiv.org/html/2609.27373#bib.bib29);[Rodkin et al\., 2026](https://arxiv.org/html/2609.27373#bib.bib25)\), rather than generating explicit intermediate reasoning tokens\([Hao et al\., 2024](https://arxiv.org/html/2609.27373#bib.bib27);[Wei et al\., 2022](https://arxiv.org/html/2609.27373#bib.bib30)\)\. This decouples parameter depth from inference\-time compute and allows additional computation to be allocated through recurrent depth\([Graves, 2016](https://arxiv.org/html/2609.27373#bib.bib28)\)\. Prior work has shown that additional recurrent computation can improve multi\-step reasoning\([Geiping et al\., 2025](https://arxiv.org/html/2609.27373#bib.bib2)\); taking this recurrent computation budget as given, we ask a different question:*must every recurrent step continue to perform unrestricted global attention routing?*Current implementations largely do so, without distinguishing which parts of recurrent computation continue to require global recomputation as inference progresses\.
When the same network repeatedly refines a representation, does its information\-access pattern evolve at the same pace as the representation itself? We study attention dynamics across recurrent depth and uncover a consistent separation in convergence timescales: routing\-related quantities, including attention support and attention distributions, stabilize substantially earlier than representation\-related quantities such as hidden states and attention outputs\. The precise ordering within routing is task\-dependent, but the broader separation is robust\. As illustrated in Figure[1](https://arxiv.org/html/2609.27373#S1.F1), recurrent models can settle on a stable region of information access while their representations are still evolving\. In this sense, the model largely determines*where to look*before it finishes refining*how to use*the retrieved information\.
This separation reveals a two\-stage structure in recurrent inference\. Early recurrent steps perform global search and progressively discover a sparse working set of relevant context, whereas later steps continue refining representations over largely the same routing support\. Standard recurrent inference does not exploit this distinction: it recomputes global attention over the full context at every step, even after the working set has largely stabilized\. This redundancy becomes increasingly costly with context length because attention is the only component of the recurrent block whose computation grows quadratically with sequence length\. If late recurrent computation only needs to revisit a much smaller working set of sizem≪nm\\ll n, its routing cost can in principle change fromO\(n2\)O\(n^\{2\}\)toO\(nm\)O\(nm\)while preserving recurrent refinement\. In other words, routing can stop being global before recurrent refinement should stop\.
Motivated by this structure, we introduceWISE\(Working\-setInference withSupportExploitation\), illustrated in Figure[1](https://arxiv.org/html/2609.27373#S1.F1)\.WISEdoes not reduce recurrent depth: it uses unrestricted global attention during an early discovery phase and then reuses the discovered routing support during later recurrent steps\. Our final realization constructs the working set directly in block space, enabling efficient block\-sparse execution\. Crucially,WISEreuses only*where*attention may route: hidden representations, queries, keys, values, and within\-support attention weights continue to evolve throughout recurrence\. Thus,WISEpreserves late representation refinement while avoiding repeated global search over positions that have already fallen outside the discovered working set\.
\([1](https://arxiv.org/html/2609.27373#S1.F1)\) Convergence timescales
\([1](https://arxiv.org/html/2609.27373#S1.F1)\) Working\-set reuse
Figure 1:Routing stabilizes before representation refinement is complete\.\(a\)On 30 benchmark\-native examples from each of HotpotQA and GSM8K, routing\-related quantities—diagnostic token\-level attention support and attention distributions—stabilize substantially earlier than representation\-related quantities—hidden states and attention outputs\. The ordering within routing is task\-dependent, but the separation between routing stabilization and representation refinement is consistent across both tasks\.\(b\)This separation motivatesWISE: early recurrent steps use unrestricted global attention to discover a sparse block\-structured working set, while later steps reuse only this routing support\. Recurrent depth is preserved, and hidden representations, queries, keys, values, and within\-support attention weights continue to evolve throughout recurrence\. Thus,WISEreuses*where*attention may route without freezing*how*information within the working set is processed:*discover early, reuse late*\.Our experiments support both the mechanism and its computational consequence\. Controlled interventions show that recurrent discovery is important for identifying effective working sets, and that support\-only reuse better preserves model behavior than more restrictive alternatives\. Across benchmarks,WISElargely preserves Full\-model behavior, while matched context\-scaling experiments reveal an increasingly favorable quality–efficiency tradeoff as the discovered working set becomes sparser with context length\. This structured sparsity translates into practical GPU acceleration, achieving up to a1\.76×1\.76\\timesattention speedup over native FlashAttention at 4K context\. Together, these results show that recurrent routing structure can become computationally reusable while late representation refinement remains valuable\.
#### Contributions\.
We summarize our contributions as follows: \(i\)Mechanism discovery\.We identify a consistent separation between routing and representation timescales in recurrent language models: attention support and attention distributions stabilize substantially earlier than hidden representations and attention outputs, revealing that information\-access structure can become reusable before recurrent refinement is complete\. \(ii\)Method\.We introduceWISE, a training\-free recurrent inference method that preserves recurrent depth while discovering a sparse block\-structured working set using unrestricted early attention and reusing this support during later recurrent steps, reducing late attention routing fromO\(n2\)O\(n^\{2\}\)toO\(nm\)O\(nm\)while keeping within\-support computation dynamic\. \(iii\)Empirical validation and efficiency\.We causally validate this decomposition through matched interventions, replicate the routing–representation separation on a second recurrent backbone, and identify a boundary condition for delayed discovery\. Across context lengths,WISEshows a favorable quality–efficiency tradeoff: increasingly sparse routing support translates into practical attention acceleration while largely preserving downstream quality through 2K context\.
## 2Related Work
Recurrent and Latent Test\-Time Computation\.Recurrent architectures increase effective computational depth by repeatedly applying shared network blocks, fromUniversal Transformersto more recent looped Transformers for iterative computation and reasoning\([Dehghani et al\., 2018](https://arxiv.org/html/2609.27373#bib.bib31);[Giannou et al\., 2023](https://arxiv.org/html/2609.27373#bib.bib29);[Yang et al\., 2024](https://arxiv.org/html/2609.27373#bib.bib32);[Saunshi et al\., 2025](https://arxiv.org/html/2609.27373#bib.bib20)\)\. Recent work extends this idea to language\-model inference, using recurrent depth to allocate additional test\-time computation and improve multi\-step reasoning\([Geiping et al\., 2025](https://arxiv.org/html/2609.27373#bib.bib2);[Rodkin et al\., 2026](https://arxiv.org/html/2609.27373#bib.bib25);[Zhang et al\., 2026](https://arxiv.org/html/2609.27373#bib.bib19);[Kohli et al\., 2026](https://arxiv.org/html/2609.27373#bib.bib18)\)\. Related approaches reason directly in continuous latent states\([Hao et al\., 2024](https://arxiv.org/html/2609.27373#bib.bib27)\)\. These works establish recurrent refinement as a useful computational mechanism\.WISEinstead asks whether every recurrent step must continue to perform unrestricted global attention routing, and studies how routing and representation refinement evolve across recurrent depth\.
Adaptive Transformer Computation\.A complementary line of work reduces inference cost by adapting how much computation different inputs or tokens receive\.Adaptive Computation Time,Mixture\-of\-Depths, and early\-exit methods dynamically reduce recurrent or layer depth\([Graves, 2016](https://arxiv.org/html/2609.27373#bib.bib28);[Banino et al\., 2021](https://arxiv.org/html/2609.27373#bib.bib15);[Elbayad et al\., 2019](https://arxiv.org/html/2609.27373#bib.bib17);[Raposo et al\., 2024](https://arxiv.org/html/2609.27373#bib.bib26);[Xin et al\., 2020](https://arxiv.org/html/2609.27373#bib.bib16);[Zhou et al\., 2020](https://arxiv.org/html/2609.27373#bib.bib14);[Elhoushi et al\., 2024](https://arxiv.org/html/2609.27373#bib.bib21)\)\.WISEis orthogonal: it preserves recurrent depth and reduces only the scope of late global routing, while keeping attention weights and hidden\-state updates dynamic\.
Efficient and Sparse Attention\.Kernel\-level methods such asFlashAttentionandFlashAttention\-2accelerate exact dense attention through IO\-aware tiling and improved parallelism\([Dao et al\., 2022](https://arxiv.org/html/2609.27373#bib.bib12);[Dao, 2024](https://arxiv.org/html/2609.27373#bib.bib1)\)\. Long\-context methods further reduce attention cost through KV\-cache compression or sparse retrieval, includingH2O,SnapKV,PyramidKV,Quest, andMInference\([Zhang et al\., 2023](https://arxiv.org/html/2609.27373#bib.bib22);[Li et al\., 2024](https://arxiv.org/html/2609.27373#bib.bib24);[Cai et al\., 2024](https://arxiv.org/html/2609.27373#bib.bib23);[Tang et al\., 2024](https://arxiv.org/html/2609.27373#bib.bib4);[Jiang et al\., 2024](https://arxiv.org/html/2609.27373#bib.bib11)\)\. Other work reuses attention patterns across layers or decoding steps\([Xiao et al\., 2019](https://arxiv.org/html/2609.27373#bib.bib13);[Xu et al\., 2025](https://arxiv.org/html/2609.27373#bib.bib3);[Deshmukh et al\., 2025](https://arxiv.org/html/2609.27373#bib.bib10)\)\.WISEinstead targets recurrent language models and reuses only the routing support after it stabilizes, while within\-support attention and representations remain dynamic; this support\-level sparsity complements efficient kernels such as FlashAttention\.
## 3Problem Setup
#### Recurrent inference\.
Consider a recurrent language model that repeatedly applies a shared transformation to a sequence of hidden representations\. Given a sequence of lengthnn, let𝑯\(t\)∈ℝn×d\\bm\{H\}^\{\(t\)\}\\in\\mathbb\{R\}^\{n\\times d\}denote the hidden states after recurrent steptt, whereddis the hidden dimension\. The recurrent computation is written as
𝑯\(t\+1\)=Fθ\(𝑯\(t\)\),t=0,…,T−1,\\bm\{H\}^\{\(t\+1\)\}=F\_\{\\theta\}\\\!\\left\(\\bm\{H\}^\{\(t\)\}\\right\),\\qquad t=0,\\ldots,T\-1,\(1\)where the parametersθ\\thetaare shared across recurrent steps andTTdenotes the recurrent depth\. This formulation abstracts away model\-specific non\-recurrent components and focuses on the computation repeatedly applied within the recurrent core\.
#### Recurrent attention\.
At each recurrent step, self\-attention is recomputed from the current hidden representations\. Suppressing recurrent\-block and attention\-head indices for clarity, we define𝑸\(t\)=𝑯\(t\)𝑾Q\\bm\{Q\}^\{\(t\)\}=\\bm\{H\}^\{\(t\)\}\\bm\{W\}\_\{Q\},𝑲\(t\)=𝑯\(t\)𝑾K\\bm\{K\}^\{\(t\)\}=\\bm\{H\}^\{\(t\)\}\\bm\{W\}\_\{K\}, and𝑽\(t\)=𝑯\(t\)𝑾V\\bm\{V\}^\{\(t\)\}=\\bm\{H\}^\{\(t\)\}\\bm\{W\}\_\{V\}, where𝑸\(t\),𝑲\(t\),𝑽\(t\)∈ℝn×dh\\bm\{Q\}^\{\(t\)\},\\bm\{K\}^\{\(t\)\},\\bm\{V\}^\{\(t\)\}\\in\\mathbb\{R\}^\{n\\times d\_\{h\}\}anddhd\_\{h\}is the attention\-head dimension\. The corresponding attention matrix and attention output are
𝑨\(t\)=softmax\(𝑸\(t\)𝑲\(t\)⊤dh\+𝑴\),𝑶\(t\)=𝑨\(t\)𝑽\(t\),\\bm\{A\}^\{\(t\)\}=\\operatorname\{softmax\}\\\!\\left\(\\frac\{\\bm\{Q\}^\{\(t\)\}\\bm\{K\}^\{\(t\)\\top\}\}\{\\sqrt\{d\_\{h\}\}\}\+\\bm\{M\}\\right\),\\qquad\\bm\{O\}^\{\(t\)\}=\\bm\{A\}^\{\(t\)\}\\bm\{V\}^\{\(t\)\},\(2\)where𝑴\\bm\{M\}denotes the attention mask\. For query positionii, we write𝒂i\(t\)=𝑨\(t\)i,:\\bm\{a\}\_\{i\}^\{\(t\)\}=\\bm\{A\}^\{\(t\)\}\_\{i,:\}for its attention distribution and let𝒦i⊆\{1,…,n\}\\mathcal\{K\}\_\{i\}\\subseteq\\\{1,\\ldots,n\\\}denote the keys admissible under𝑴\\bm\{M\}\.
#### Attention support and working sets\.
For a mass thresholdη∈\(0,1\)\\eta\\in\(0,1\), letSi\(t,η\)⊆𝒦iS\_\{i\}^\{\(t,\\eta\)\}\\subseteq\\mathcal\{K\}\_\{i\}denote the smallest set of admissible keys whose cumulative attention mass under𝒂i\(t\)\\bm\{a\}\_\{i\}^\{\(t\)\}is at leastη\\eta; we refer toSi\(t,η\)S\_\{i\}^\{\(t,\\eta\)\}as the*attention support*\. More generally, a*working set*is a sparse routing support discovered during early recurrent computation and reused at later recurrent steps\. For structured execution, we partition the sequence into blocks of sizeBBand denote byℬq\(t,η\)\\mathcal\{B\}\_\{q\}^\{\(t,\\eta\)\}the active key\-block support associated with query blockqqat steptt; its construction from block\-level attention mass is specified in Section[5](https://arxiv.org/html/2609.27373#S5)\. Unless explicitly shown, recurrent\-block and attention\-head indices are suppressed for bothSi\(t,η\)S\_\{i\}^\{\(t,\\eta\)\}andℬq\(t,η\)\\mathcal\{B\}\_\{q\}^\{\(t,\\eta\)\}\. Standard recurrent inference recomputes global attention over all admissible query–key pairs at every recurrent step, even when this routing structure has already stabilized\.
## 4Computational Structure of Recurrent Inference
The distinction between recurrent representation refinement and attention routing introduced in Section[3](https://arxiv.org/html/2609.27373#S3)has a direct computational consequence\. For a sequence of lengthnn, linear projections and feed\-forward layers scale asO\(nd2\)O\(nd^\{2\}\), whereas self\-attention computes pairwise query–key interactions through𝑸\(t\)𝑲\(t\)⊤\\bm\{Q\}^\{\(t\)\}\\bm\{K\}^\{\(t\)\\top\}and aggregates values through𝑨\(t\)𝑽\(t\)\\bm\{A\}^\{\(t\)\}\\bm\{V\}^\{\(t\)\}, resulting inO\(n2d\)O\(n^\{2\}d\)computation across attention heads\. OverTTrecurrent steps, the leading\-order cost of standard full\-attention inference is therefore
Cfull=O\(Tnd2\)\+O\(Tn2d\),C\_\{\\mathrm\{full\}\}=O\\\!\\left\(Tnd^\{2\}\\right\)\+O\\\!\\left\(Tn^\{2\}d\\right\),\(3\)where the second term is the only component that grows quadratically with context length and corresponds to repeatedly recomputing global attention routing\. Figure[1](https://arxiv.org/html/2609.27373#S1.F1)[1](https://arxiv.org/html/2609.27373#S1.F1)reveals a mismatch between this computation and the dynamics of recurrent inference: routing\-related quantities stabilize substantially earlier than hidden representations and attention outputs, while standard recurrent inference continues to recompute attention over the full context at every step\. If early recurrent computation has identified a working set containing on averagem≪nm\\ll nrelevant keys per query, late recurrent attention needs only to compare each query against thesemmkeys, reducing the support\-dependent attention cost from
O\(n2d\)⟶O\(nmd\)\.O\(n^\{2\}d\)\\quad\\longrightarrow\\quad O\(nmd\)\.\(4\)Equivalently, defining the normalized working\-set densityρ=m/n\\rho=m/n, the late\-stage cost isO\(ρn2d\)O\(\\rho n^\{2\}d\)\. TheO\(nmd\)O\(nmd\)form makes explicit that late routing depends on the size of the discovered working set rather than the full context: ifmmgrows more slowly thannn, the routing cost becomes subquadratic, and it becomes linear innnwhenmmis bounded\. This creates an opportunity to reduce repeated global routing while preserving the recurrent evolution of𝑸\(t\)\\bm\{Q\}^\{\(t\)\},𝑲\(t\)\\bm\{K\}^\{\(t\)\},𝑽\(t\)\\bm\{V\}^\{\(t\)\}, and the within\-support attention weights, motivatingWISE:*discover globally early, then reuse the working set late*\.
## 5WISE: Discover Early, Reuse Late
Section 4 shows that late recurrent attention can reduce its routing cost fromO\(n2d\)O\(n^\{2\}d\)toO\(nmd\)O\(nmd\)if computation is restricted to a sparse working set ofm≪nm\\ll nkeys\. WISE realizes this structure by separating recurrent inference into two stages: unrestricted global attention discovers the relevant routing support early, and the discovered support is reused later while recurrent refinement and all within\-support computation remain dynamic\. The support\-reuse principle itself is agnostic to routing granularity, but unstructured token\-level sparsity is difficult to translate directly into efficient GPU execution\. We therefore instantiateWISEwith block\-structured support, trading some fine\-grained sparsity for regular, reusable computation that block\-sparse kernels can execute efficiently\.
### 5\.1Block\-Structured Working\-Set Discovery
During the firsttdt\_\{\\mathrm\{d\}\}recurrent steps,WISEretains unrestricted attention over the full context\. We partition the sequence into blocks of sizeBBand letℐq\\mathcal\{I\}\_\{q\}denote the token positions in query blockqqandℐb\\mathcal\{I\}\_\{b\}those in key blockbb\. Suppressing recurrent\-block and attention\-head indices as in Section[3](https://arxiv.org/html/2609.27373#S3), we define the attention mass assigned from query blockqqto key blockbbat stepttas
rqb\(t\)=1\|ℐq\|∑i∈ℐq∑j∈ℐb\[𝑨\(t\)\]ij\.r\_\{qb\}^\{\(t\)\}=\\frac\{1\}\{\|\\mathcal\{I\}\_\{q\}\|\}\\sum\_\{i\\in\\mathcal\{I\}\_\{q\}\}\\sum\_\{j\\in\\mathcal\{I\}\_\{b\}\}\\big\[\\bm\{A\}^\{\(t\)\}\\big\]\_\{ij\}\.\(5\)For a mass thresholdη\\eta, letℬq\(t,η\)\\mathcal\{B\}\_\{q\}^\{\(t,\\eta\)\}be the smallest set of admissible key blocks whose cumulative block\-level mass reachesη\\eta\. The working set reused after discovery is
ℬq=⋃t=td−3tdℬq\(t,η\),ℬq\(t,η\)=argminℬ\|ℬ\|s\.t\.∑b∈ℬrqb\(t\)≥η\.\\mathcal\{B\}\_\{q\}=\\bigcup\_\{t=t\_\{\\mathrm\{d\}\}\-3\}^\{t\_\{\\mathrm\{d\}\}\}\\mathcal\{B\}\_\{q\}^\{\(t,\\eta\)\},\\qquad\\mathcal\{B\}\_\{q\}^\{\(t,\\eta\)\}=\\underset\{\\mathcal\{B\}\}\{\\arg\\min\}\\ \|\\mathcal\{B\}\|\\quad\\mathrm\{s\.t\.\}\\quad\\sum\_\{b\\in\\mathcal\{B\}\}r\_\{qb\}^\{\(t\)\}\\geq\\eta\.\(6\)Our default configuration usesT=32T=32,td=12t\_\{\\mathrm\{d\}\}=12,B=32B=32, andη=0\.95\\eta=0\.95, so the working set is the union of direct block\-level95%95\\%\-mass supports from recurrent steps99–1212\. We examine sensitivity to the cumulative\-mass threshold in Appendix[C\.7](https://arxiv.org/html/2609.27373#A3.SS7)\. The support is constructed independently for each recurrent block, attention head, and query block\. Importantly,WISEdiscovers support directly in block space rather than constructing a token\-level support and subsequently rounding it to blocks\. This aligns the discovery objective with the granularity the sparse kernel actually executes, avoiding a mismatch between fine\-grained support selection and block\-structured computation\.
### 5\.2Working\-Set Reuse
For recurrent stepst\>tdt\>t\_\{\\mathrm\{d\}\}, the discovered block supportℬq\\mathcal\{B\}\_\{q\}remains fixed, but the hidden representations and attention computation continue to evolve\. Letq\(i\)q\(i\)andb\(j\)b\(j\)denote the query and key blocks containing positionsiiandjj, respectively, and define the working\-set mask
\[𝑹\]ij=\{0,b\(j\)∈ℬq\(i\),−∞,otherwise\.\\big\[\\bm\{R\}\\big\]\_\{ij\}=\\begin\{cases\}0,&b\(j\)\\in\\mathcal\{B\}\_\{q\(i\)\},\\\\ \-\\infty,&\\text\{otherwise\}\.\\end\{cases\}\(7\)The late\-stage attention is recomputed as
𝑨~\(t\)=softmax\(𝑸\(t\)𝑲\(t\)⊤dh\+𝑴\+𝑹\),𝑶~\(t\)=𝑨~\(t\)𝑽\(t\)\.\\widetilde\{\\bm\{A\}\}^\{\(t\)\}=\\operatorname\{softmax\}\\\!\\left\(\\frac\{\\bm\{Q\}^\{\(t\)\}\\bm\{K\}^\{\(t\)\\top\}\}\{\\sqrt\{d\_\{h\}\}\}\+\\bm\{M\}\+\\bm\{R\}\\right\),\\qquad\\widetilde\{\\bm\{O\}\}^\{\(t\)\}=\\widetilde\{\\bm\{A\}\}^\{\(t\)\}\\bm\{V\}^\{\(t\)\}\.\(8\)Thus,WISEreuses only the routing support:𝑯\(t\)\\bm\{H\}^\{\(t\)\},𝑸\(t\)\\bm\{Q\}^\{\(t\)\},𝑲\(t\)\\bm\{K\}^\{\(t\)\},𝑽\(t\)\\bm\{V\}^\{\(t\)\}, and the attention weights within the working set remain dynamic throughout recurrence\. It therefore fixes*where*attention may route without freezing*how*the retained information is weighted and transformed\.
### 5\.3Computational Complexity
Letmmdenote the average number of key positions represented by the active key blocks for each query during the reuse phase, letρ=m/n\\rho=m/ndenote the corresponding normalized working\-set density, and letα=td/T\\alpha=t\_\{\\mathrm\{d\}\}/Tdenote the fraction of recurrent steps devoted to discovery\. Full attention incursO\(n2d\)O\(n^\{2\}d\)routing cost at every recurrent step, whereas block\-structured reuse reduces the late\-stage cost toO\(nmd\)O\(nmd\)\. Ignoring block\-boundary and kernel overhead, the relative attention\-routing cost is therefore
CWISEattnCfullattn≈α\+\(1−α\)ρ,ΔCattn≈\(1−α\)\(1−ρ\)\.\\frac\{C\_\{\\mathrm\{WISE\}\}^\{\\mathrm\{attn\}\}\}\{C\_\{\\mathrm\{full\}\}^\{\\mathrm\{attn\}\}\}\\approx\\alpha\+\(1\-\\alpha\)\\rho,\\qquad\\Delta C^\{\\mathrm\{attn\}\}\\approx\(1\-\\alpha\)\(1\-\\rho\)\.\(9\)The gain increases as the working set becomes sparse relative to the full context\. Equivalently, the late recurrent routing term changes fromO\(n2d\)O\(n^\{2\}d\)toO\(nmd\)O\(nmd\), replacing dependence on the full key set with dependence on the discovered working\-set size while preserving the recurrent refinement identified in Section[4](https://arxiv.org/html/2609.27373#S4)\.
## 6Theoretical Analysis
The design ofWISErelies on the possibility that a discrete routing support becomes stable before the underlying recurrent representation converges\. We formalize this behavior for the block\-level cumulative\-mass support used in Section[5\.1](https://arxiv.org/html/2609.27373#S5.SS1)\. The analysis does not predict the empirical switching depthtdt\_\{\\mathrm\{d\}\}; rather, it explains why a reusable routing structure can emerge at finite recurrent depth while representation refinement continues\.
### 6\.1Finite\-Time Identification of the Working Set
Fix a recurrent block, attention head, and query blockqq, and let𝒓q\(t\)=\[rqb\(t\)\]b\\bm\{r\}\_\{q\}^\{\(t\)\}=\[r\_\{qb\}^\{\(t\)\}\]\_\{b\}denote its block\-level attention\-mass distribution as defined in Equation[5](https://arxiv.org/html/2609.27373#S5.E5)\. Letℬq\(t,η\)\\mathcal\{B\}\_\{q\}^\{\(t,\\eta\)\}be the corresponding cumulative\-η\\etasupport and suppose𝑯\(t\)→𝑯⋆\\bm\{H\}^\{\(t\)\}\\rightarrow\\bm\{H\}^\{\\star\}, inducing a limiting block\-mass distribution𝒓q⋆\\bm\{r\}\_\{q\}^\{\\star\}and supportℬq⋆\\mathcal\{B\}\_\{q\}^\{\\star\}\. Letkq⋆=\|ℬq⋆\|k\_\{q\}^\{\\star\}=\|\\mathcal\{B\}\_\{q\}^\{\\star\}\|\. We assume that the limiting support is nondegenerate: thekq⋆k\_\{q\}^\{\\star\}\-th and\(kq⋆\+1\)\(k\_\{q\}^\{\\star\}\+1\)\-th largest block masses are strictly separated, andη\\etalies strictly between the cumulative masses of the firstkq⋆−1k\_\{q\}^\{\\star\}\-1andkq⋆k\_\{q\}^\{\\star\}blocks\. LetΔq\>0\\Delta\_\{q\}\>0denote the minimum margin to these ranking and cumulative\-mass boundaries\.
###### Theorem 6\.1\(Finite\-Time Identification of Recurrent Routing Support\)\.
Suppose the mapping from𝐇\(t\)\\bm\{H\}^\{\(t\)\}to𝐫q\(t\)\\bm\{r\}\_\{q\}^\{\(t\)\}is locallyLqL\_\{q\}\-Lipschitz around𝐇⋆\\bm\{H\}^\{\\star\}and
‖𝑯\(t\)−𝑯⋆‖≤Cβt\\\|\\bm\{H\}^\{\(t\)\}\-\\bm\{H\}^\{\\star\}\\\|\\leq C\\beta^\{t\}for some0<β<10<\\beta<1\. Then there exists a finiteTS,qT\_\{S,q\}such thatℬq\(t,η\)=ℬq⋆\\mathcal\{B\}\_\{q\}^\{\(t,\\eta\)\}=\\mathcal\{B\}\_\{q\}^\{\\star\}for allt≥TS,qt\\geq T\_\{S,q\}, with
TS,q=O\(log\(LqC/Δq\)−logβ\)\.T\_\{S,q\}=O\\\!\\left\(\\frac\{\\log\(L\_\{q\}C/\\Delta\_\{q\}\)\}\{\-\\log\\beta\}\\right\)\.\(10\)
Theorem[6\.1](https://arxiv.org/html/2609.27373#S6.Thmtheorem1)shows that exact routing\-support identification can occur at finite recurrent depth even when𝑯\(t\)\\bm\{H\}^\{\(t\)\}approaches𝑯⋆\\bm\{H\}^\{\\star\}only asymptotically\. Once𝒓q\(t\)\\bm\{r\}\_\{q\}^\{\(t\)\}lies sufficiently far from every ranking and cumulative\-mass boundary, further continuous changes cannot alter the selected block support\. Since the model contains finitely many recurrent blocks, attention heads, and query blocks, taking the maximum of their identification times yields a finite depth after which all nondegenerate supports are simultaneously stable\. If the temporal\-union window used byWISElies beyond this depth, the union reduces exactly to the limiting support; before exact stabilization, the union provides a conservative buffer against local routing fluctuations\. Formal definitions ofΔq\\Delta\_\{q\}and the proof are provided in Appendix[A](https://arxiv.org/html/2609.27373#A1)\.
### 6\.2Support Stabilization and Representation Convergence
For query blockqq, define the routing region
ℛq\(ℬq⋆\)=\{𝑯:ℬq\(η\)\(𝑯\)=ℬq⋆\}\.\\mathcal\{R\}\_\{q\}\(\\mathcal\{B\}\_\{q\}^\{\\star\}\)=\\left\\\{\\bm\{H\}:\\mathcal\{B\}\_\{q\}^\{\(\\eta\)\}\(\\bm\{H\}\)=\\mathcal\{B\}\_\{q\}^\{\\star\}\\right\\\}\.
###### Proposition 1\(Support Stabilization Does Not Imply Representation Convergence\)\.
Under the assumptions of Theorem[6\.1](https://arxiv.org/html/2609.27373#S6.Thmtheorem1),𝐇\(t\)→𝐇⋆\\bm\{H\}^\{\(t\)\}\\rightarrow\\bm\{H\}^\{\\star\}implies finite\-time stabilization of the cumulative\-mass routing support\. Moreover,
𝔹\(𝑯⋆,ΔqLq\)⊆ℛq\(ℬq⋆\),\\mathbb\{B\}\\\!\\left\(\\bm\{H\}^\{\\star\},\\frac\{\\Delta\_\{q\}\}\{L\_\{q\}\}\\right\)\\subseteq\\mathcal\{R\}\_\{q\}\(\\mathcal\{B\}\_\{q\}^\{\\star\}\),\(11\)so support stabilization alone is insufficient to imply convergence of𝐇\(t\)\\bm\{H\}^\{\(t\)\}\.
Proposition[1](https://arxiv.org/html/2609.27373#Thmproposition1)shows that a fixed routing support identifies a region of representation space rather than a unique representation\. The recurrent state can therefore continue to evolve while remaining inside the same routing region\. This is precisely the regime exploited byWISE: the block support can be reused while𝑯\(t\)\\bm\{H\}^\{\(t\)\},𝑸\(t\)\\bm\{Q\}^\{\(t\)\},𝑲\(t\)\\bm\{K\}^\{\(t\)\},𝑽\(t\)\\bm\{V\}^\{\(t\)\}, and the within\-support attention weights continue to change\. The proof is provided in Appendix[A](https://arxiv.org/html/2609.27373#A1)\. Importantly, this result does not imply that recurrence should terminate when the support stabilizes; rather, it explains why a discrete routing structure can become reusable while recurrent representation refinement continues\.
## 7Experiments
### 7\.1Experimental Setup
#### Models and benchmarks\.
Our primary experiments use the publicly available Huginn recurrent language model with recurrent depthT=32T=32\([Geiping et al\., 2025](https://arxiv.org/html/2609.27373#bib.bib2)\)\. To assess whether the observed recurrent dynamics generalize beyond the primary backbone, we additionally evaluate a recurrent Llama\-3\.2 model trained withT=32T=32recurrent steps \(Recurrent\-Llama\-T32\)\([McLeish et al\., 2025](https://arxiv.org/html/2609.27373#bib.bib7)\)\. The replication uses the same HotpotQA\([Yang et al\., 2018](https://arxiv.org/html/2609.27373#bib.bib5)\)and GSM8K\([Cobbe et al\., 2021](https://arxiv.org/html/2609.27373#bib.bib9)\)mechanism diagnostics and the same HotpotQA and 2WikiMultiHopQA\([Ho et al\., 2020](https://arxiv.org/html/2609.27373#bib.bib8)\)behavioral evaluations\. Unless otherwise specified, benchmark\-level comparisons use 100 deterministically selected examples with matched inputs and recurrent initialization across methods, while mechanism diagnostics use fixed 30\-example cohorts\. The Recurrent\-Llama replication preserves the sameWISEconfiguration used for Huginn; model\-specific instrumentation and prompting details are provided in Appendix[B](https://arxiv.org/html/2609.27373#A2)\.
WISEconfiguration\.We use a single frozen configuration throughout the final experiments: block sizeB=32B=32, discovery depthtd=12t\_\{\\mathrm\{d\}\}=12, mass thresholdη=0\.95\\eta=0\.95, and a four\-step discovery window spanning recurrent steps99–1212\. The working set is constructed directly in block space and reused fort\>tdt\>t\_\{\\mathrm\{d\}\}, while hidden states, queries, keys, values, and within\-support attention weights remain dynamic\. We use the same configuration across benchmarks and context lengths without benchmark\-specific tuning\.
#### Evaluation\.
We report token\-level F1 for downstream quality, together with paired behavioral comparisons, answer changes, teacher\-forced gold\-answer loss for causal interventions, block density, and retained future Full\-attention mass\. Mechanism analysis uses fixed convergence criteria for attention support, attention distributions, hidden states, and attention outputs; Appendix[B](https://arxiv.org/html/2609.27373#A2)provides complete metric definitions, statistical procedures, and systems protocols\.
### 7\.2Mechanism Discovery
#### Routing Stabilizes Before Representation Refinement
We quantify how different components of recurrent computation stabilize across depth\. For the diagnostic token support, convergence is defined as the earliest step at which consecutive supports maintain Jaccard similarity of at least0\.900\.90for three successive transitions; for the attention distribution, hidden state, and attention output, convergence is defined using a matched three\-transition criterion based on changes relative to their early\-trajectory scale\. Full definitions are provided in Appendix[B](https://arxiv.org/html/2609.27373#A2)\. As shown in Figure[1](https://arxiv.org/html/2609.27373#S1.F1)[1](https://arxiv.org/html/2609.27373#S1.F1), routing\-related quantities consistently stabilize earlier than representation\-related quantities across both HotpotQA and GSM8K, although the precise ordering within routing is task\-dependent\. This separation suggests that recurrent inference enters a regime in which the model has largely settled on where to retrieve information while representation refinement is still ongoing\. Because the diagnostic token support differs from the block\-structured support used byWISE, we separately analyze the deployedB=32B=32,η=0\.95\\eta=0\.95cumulative\-mass support\. The resulting steps\-99–1212temporal union enters a high\-stability regime near the discovery depthtd=12t\_\{\\mathrm\{d\}\}=12, providing a method\-aligned basis for late support reuse \(Appendix[C\.2](https://arxiv.org/html/2609.27373#A3.SS2)\)\.
#### Robustness and cross\-backbone replication\.
The routing–representation separation persists across consistently varied convergence criteria and replicates on Recurrent\-Llama\-T32\. Under the sameT=32T=32diagnostic protocol, all60/6060/60Recurrent\-Llama examples across HotpotQA and GSM8K exhibit a positive routing–representation gap, despite substantial differences in absolute convergence speed and within\-routing ordering\. Full threshold\-sensitivity and cross\-backbone results are reported in Appendix[C](https://arxiv.org/html/2609.27373#A3)\.
### 7\.3Causal Validation of Support\-Only Reuse
We next test which parts of recurrent computation can be safely reused once routing has stabilized\. Truncating recurrence at the discovery depth worsens teacher\-forced gold\-answer loss, indicating that representation refinement remains useful after global routing has largely stabilized\. Among methods that preserve recurrent depth,WISEbetter preserves Full\-model behavior than early\-static controls, supporting the importance of recurrent working\-set discovery\. We further compare againstFreeze\-A@12, which uses exactly the same support asWISEbut also freezes the within\-support attention distribution at the discovery endpoint\. This stronger intervention is more disruptive than support\-only reuse, indicating that the reusable structure is the routing support rather than the full attention distribution\. Together, these results support the central design ofWISE: preserve late recurrent refinement while reusing only the routing structure that has already stabilized\.
#### Cross\-backbone replication\.
Recurrent\-Llama\-T32 yields a more nuanced behavioral replication: theWISE–SizeMatchedF1 differences are\+0\.0083\+0\.0083on HotpotQA and−0\.0043\-0\.0043on 2WikiMultiHopQA, indicating broadly comparable downstream quality across the two methods\. At the same time,WISEretains approximately97%97\\%of future Full\-attention mass and changes fewer answers on both benchmarks\. As analyzed in Appendix[C\.4](https://arxiv.org/html/2609.27373#A3.SS4), early static routing is substantially more predictive on this backbone, suggesting that delayed discovery is most beneficial when routing continues to reorganize during early recurrence\.
Table 1:Controlled interventions validate support\-only reuse\.Δ\\DeltaAns denotes answer changes relative to the matched Full run\.Truncate@12stops recurrence at the discovery depth;SizeMatchedexactly matches the support cardinality ofWISE; andFreeze\-A@12uses exactly the same support asWISEbut freezes its step\-12 within\-support attention distribution\.
### 7\.4Practical GPU Efficiency
Native FlashAttention is highly optimized for regular dense tiled execution, but it does not directly exploit the input\-dependent block\-sparse schedules discovered byWISE: applying the support only as a mask would preserve semantics while still traversing masked tiles\. We therefore implement an exact\-B=32B=32Triton kernel\([Tillet et al\., 2019](https://arxiv.org/html/2609.27373#bib.bib6)\)for the Huginn attention geometry on NVIDIA Ampere GPUs, using a reusable GPU\-resident sparse schedule to skip inactive key blocks while fusing query–key scoring, causal masking, online softmax, and value aggregation\. The final implementation further specializes arithmetic for the native head dimension and tunes the execution configuration, without changingWISEsupport or attention semantics; Appendix[B](https://arxiv.org/html/2609.27373#A2)provides implementation, correctness, and timing details\. As shown in Table[2](https://arxiv.org/html/2609.27373#S7.T2), this converts the increasingly sparse routing structure into substantial wall\-clock gains, reaching a1\.76×1\.76\\timessetup\-inclusive late\-reuse speedup over native FlashAttention at 4K; the completeT=32T=32attention trajectory remains1\.36×1\.36\\timesfaster after including unrestricted discovery, and a matched exact\-B=32B=32backend comparison yields a2\.50×2\.50\\timesspeedup from reducing active routing support alone\. These results show thatWISE’s structured sparsity is practically exploitable, not merely a reduction in theoretical attention work\. Our implementation is nevertheless a workload\-specialized sparse kernel rather than a fully co\-designed sparse counterpart to FlashAttention, leaving tighter integration with architecture\-specific scheduling and memory\-movement pipelines as a complementary direction for further gains\.
Table 2:WISEtranslates structured block sparsity into practical GPU acceleration at 4K context\.Latency is measured on the canonicalN=30N=30Huginn systems cohort using an NVIDIA RTX A6000, with values reported as medians over 10 balanced timing passes\. Native FlashAttention provides the optimized dense baseline\. The late\-reuse comparison includes one\-time sparse\-schedule construction, while the matched exact\-B=32B=32comparison uses the same optimized backend for Full\-support andWISE\-support execution to isolate the benefit of reducing active routing support\.
### 7\.5Quality–Efficiency Scaling with Context Length
Figure[2](https://arxiv.org/html/2609.27373#S7.F2)summarizes how the quality–efficiency tradeoff evolves with context length\. As context grows from512512to44K, working\-set density decreases from roughly57%57\\%to37%37\\%while retaining about9696–98%98\\%of future Full\-attention mass\. This increasing sparsity translates into progressively larger practical acceleration: setup\-inclusive late\-reuse speedup over native FlashAttention increases from1\.15×1\.15\\timesat512512to1\.32×1\.32\\timesat11K,1\.61×1\.61\\timesat22K, and1\.76×1\.76\\timesat44K\. At44K, the completeT=32T=32attention trajectory remains1\.36×1\.36\\timesfaster even after including the unrestricted discovery phase\. Matched quality comparisons show no detected average F1 loss through22K, while a3\.33\.3\-point reduction emerges at44K\. Thus, longer contexts expose an increasingly favorable efficiency opportunity while preserving downstream quality through moderate context lengths, with a measurable quality cost appearing only at the longest evaluated context\.
Figure 2:WISEexposes a context\-dependent quality–efficiency frontier and a sparsity–execution tradeoff\.\(a\)Working\-set density decreases with context length while retaining most future Full\-attention mass\.\(b\)Late\-reuse speedup increases with context length, while matched HotpotQA quality shows no detected average F1 loss through22K and degrades at44K; quality and systems use fixedN=100N=100andN=30N=30cohorts, respectively\.\(c\)AlthoughB=16B=16is sparsest,B=32B=32achieves the highest setup\-inclusive speedup on the matched22K workload, forming the systems\-facing knee\. Full protocols and block\-size results are in Appendices[B](https://arxiv.org/html/2609.27373#A2)and[C\.6](https://arxiv.org/html/2609.27373#A3.SS6)\.Block\-size tradeoff\.Under equal implementation\-tuning budgets,B=16B=16yields the sparsest support, whileB=32B=32achieves the highest setup\-inclusive speedup at both 2K and 4K; coarser blocks are progressively denser and slower\. No tested granularity consistently dominates downstream quality across both benchmarks, so we useB=32B=32as the systems\-facing knee between sparsity and executable regularity; complete results are in Appendix[C\.6](https://arxiv.org/html/2609.27373#A3.SS6)\.
## 8Conclusion
We identify a separation in recurrent language\-model inference: attention routing stabilizes substantially earlier than representation refinement\. Controlled interventions show that routing can become reusable before recurrent refinement is complete, motivatingWISE, which preserves recurrent depth while reusing only the sparse routing support discovered during early global attention\. Across two recurrent backbones, our results support the routing–representation separation while revealing that delayed discovery is most useful when early routing remains predictive of later attention only after several recurrent steps\. On the primary backbone,WISEconverts this reusable routing structure into practical attention acceleration, exposing a context\-dependent quality–efficiency tradeoff\. BecauseWISE’s block\-sparse routing is complementary to FlashAttention\-style kernel optimization, tighter kernel co\-design may unlock further efficiency beyond the specialized sparse implementation studied here\.
## References
- A\. Banino, J\. Balaguer, and C\. BlundellPondernet: learning to ponder\.arXiv preprint arXiv:2107\.05407\.Cited by:[§2](https://arxiv.org/html/2609.27373#S2.p2.1)\.
- Caiet al\.\(2024\)Z\. Cai, Y\. Zhang, B\. Gao, Y\. Liu, Y\. Li, T\. Liu, K\. Lu, W\. Xiong, Y\. Dong, J\. Hu,et al\.Pyramidkv: dynamic kv cache compression based on pyramidal information funneling\.arXiv preprint arXiv:2406\.02069\.Cited by:[§2](https://arxiv.org/html/2609.27373#S2.p3.1)\.
- Cobbeet al\.\(2021\)K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano,et al\.Training verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168\.Cited by:[§7\.1](https://arxiv.org/html/2609.27373#S7.SS1.SSS0.Px1.p1.1)\.
- Daoet al\.\(2022\)T\. Dao, D\. Fu, S\. Ermon, A\. Rudra, and C\. RéFlashattention: fast and memory\-efficient exact attention with io\-awareness\.Advances in neural information processing systems35,pp\. 16344–16359\.Cited by:[§2](https://arxiv.org/html/2609.27373#S2.p3.1)\.
- Dao \(2024\)T\. DaoFlashAttention\-2: faster attention with better parallelism and work partitioning\.InInternational Conference on Learning Representations,B\. Kim, Y\. Yue, S\. Chaudhuri, K\. Fragkiadaki, M\. Khan, and Y\. Sun \(Eds\.\),Vol\.2024,pp\. 35549–35562\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/file/98ed250b203d1ac6b24bbcf263e3d4a7-Paper-Conference.pdf)Cited by:[§2](https://arxiv.org/html/2609.27373#S2.p3.1)\.
- Dehghaniet al\.\(2018\)M\. Dehghani, S\. Gouws, O\. Vinyals, J\. Uszkoreit, and Ł\. KaiserUniversal transformers\.arXiv preprint arXiv:1807\.03819\.Cited by:[§1](https://arxiv.org/html/2609.27373#S1.p1.1),[§2](https://arxiv.org/html/2609.27373#S2.p1.1)\.
- Deshmukhet al\.\(2025\)D\. Deshmukh, S\. Goyal, N\. Kwatra, and R\. RamjeeKascade: a practical sparse attention method for long\-context llm inference\.arXiv preprint arXiv:2512\.16391\.Cited by:[§2](https://arxiv.org/html/2609.27373#S2.p3.1)\.
- Elbayadet al\.\(2019\)M\. Elbayad, J\. Gu, E\. Grave, and M\. AuliDepth\-adaptive transformer\.arXiv preprint arXiv:1910\.10073\.Cited by:[§2](https://arxiv.org/html/2609.27373#S2.p2.1)\.
- Elhoushiet al\.\(2024\)M\. Elhoushi, A\. Shrivastava, D\. Liskovich, B\. Hosmer, B\. Wasti, L\. Lai, A\. Mahmoud, B\. Acun, S\. Agarwal, A\. Roman,et al\.Layerskip: enabling early exit inference and self\-speculative decoding\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 12622–12642\.Cited by:[§2](https://arxiv.org/html/2609.27373#S2.p2.1)\.
- Geipinget al\.\(2025\)J\. Geiping, S\. McLeish, N\. Jain, J\. Kirchenbauer, S\. Singh, B\. Bartoldson, B\. Kailkhura, A\. Bhatele, and T\. GoldsteinScaling up test\-time compute with latent reasoning: a recurrent depth approach\.InAdvances in Neural Information Processing Systems,D\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen \(Eds\.\),Vol\.38, Main Conference,pp\. 41340–41391\.External Links:[Document](https://dx.doi.org/10.52202/085713-1380),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/file/3b01972cf31e6fa0fe29e4b8b5c2a0a1-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2609.27373#S1.p1.1),[§2](https://arxiv.org/html/2609.27373#S2.p1.1),[§7\.1](https://arxiv.org/html/2609.27373#S7.SS1.SSS0.Px1.p1.1)\.
- Giannouet al\.\(2023\)A\. Giannou, S\. Rajput, J\. Sohn, K\. Lee, J\. D\. Lee, and D\. PapailiopoulosLooped transformers as programmable computers\.InInternational Conference on Machine Learning,pp\. 11398–11442\.Cited by:[§1](https://arxiv.org/html/2609.27373#S1.p1.1),[§2](https://arxiv.org/html/2609.27373#S2.p1.1)\.
- Graves \(2016\)A\. GravesAdaptive computation time for recurrent neural networks\.arXiv preprint arXiv:1603\.08983\.Cited by:[§1](https://arxiv.org/html/2609.27373#S1.p1.1),[§2](https://arxiv.org/html/2609.27373#S2.p2.1)\.
- Haoet al\.\(2024\)S\. Hao, S\. Sukhbaatar, D\. Su, X\. Li, Z\. Hu, J\. Weston, and Y\. TianTraining large language models to reason in a continuous latent space\.arXiv preprint arXiv:2412\.06769\.Cited by:[§1](https://arxiv.org/html/2609.27373#S1.p1.1),[§2](https://arxiv.org/html/2609.27373#S2.p1.1)\.
- Hoet al\.\(2020\)X\. Ho, A\. D\. Nguyen, S\. Sugawara, and A\. AizawaConstructing a multi\-hop qa dataset for comprehensive evaluation of reasoning steps\.InProceedings of the 28th International Conference on Computational Linguistics,pp\. 6609–6625\.Cited by:[§7\.1](https://arxiv.org/html/2609.27373#S7.SS1.SSS0.Px1.p1.1)\.
- Jianget al\.\(2024\)H\. Jiang, Y\. Li, C\. Zhang, Q\. Wu, X\. Luo, S\. Ahn, Z\. Han, A\. H\. Abdi, D\. Li, C\. Lin,et al\.Minference 1\.0: accelerating pre\-filling for long\-context llms via dynamic sparse attention\.Advances in Neural Information Processing Systems37,pp\. 52481–52515\.Cited by:[§2](https://arxiv.org/html/2609.27373#S2.p3.1)\.
- Kohliet al\.\(2026\)H\. Kohli, S\. Parthasarathy, H\. Sun, and Y\. YaoLoop, think, & generalize: implicit reasoning in recurrent\-depth transformers\.arXiv preprint arXiv:2604\.07822\.Cited by:[§2](https://arxiv.org/html/2609.27373#S2.p1.1)\.
- Liet al\.\(2024\)Y\. Li, Y\. Huang, B\. Yang, B\. Venkitesh, A\. Locatelli, H\. Ye, T\. Cai, P\. Lewis, and D\. ChenSnapkv: llm knows what you are looking for before generation\.Advances in Neural Information Processing Systems37,pp\. 22947–22970\.Cited by:[§2](https://arxiv.org/html/2609.27373#S2.p3.1)\.
- McLeishet al\.\(2025\)S\. McLeish, A\. Li, J\. Kirchenbauer, D\. S\. Kalra, B\. R\. Bartoldson, B\. Kailkhura, A\. Schwarzschild, J\. Geiping, T\. Goldstein, and M\. GoldblumTeaching pretrained language models to think deeper with retrofitted recurrence\.arXiv preprint arXiv:2511\.07384\.Cited by:[§7\.1](https://arxiv.org/html/2609.27373#S7.SS1.SSS0.Px1.p1.1)\.
- Raposoet al\.\(2024\)D\. Raposo, S\. Ritter, B\. Richards, T\. Lillicrap, P\. C\. Humphreys, and A\. SantoroMixture\-of\-depths: dynamically allocating compute in transformer\-based language models\.arXiv preprint arXiv:2404\.02258\.Cited by:[§2](https://arxiv.org/html/2609.27373#S2.p2.1)\.
- Rodkinet al\.\(2026\)I\. Rodkin, D\. Orel, K\. Smirnov, A\. Bolatov, B\. Elbouardi, B\. Hassan, Y\. Kuratov, A\. Bulatov, P\. Nakov, T\. Baldwin,et al\.Beyond memorization: extending reasoning depth with recurrence, memory and test\-time compute scaling\.InFindings of the Association for Computational Linguistics: ACL 2026,pp\. 42385–42404\.Cited by:[§1](https://arxiv.org/html/2609.27373#S1.p1.1),[§2](https://arxiv.org/html/2609.27373#S2.p1.1)\.
- Saunshiet al\.\(2025\)N\. Saunshi, N\. Dikkala, Z\. Li, S\. Kumar, and S\. J ReddiReasoning with latent thoughts: on the power of looped transformers\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 14855–14881\.Cited by:[§2](https://arxiv.org/html/2609.27373#S2.p1.1)\.
- Tanget al\.\(2024\)J\. Tang, Y\. Zhao, K\. Zhu, G\. Xiao, B\. Kasikci, and S\. HanQuest: query\-aware sparsity for efficient long\-context llm inference\.arXiv preprint arXiv:2406\.10774\.Cited by:[§2](https://arxiv.org/html/2609.27373#S2.p3.1)\.
- Tilletet al\.\(2019\)P\. Tillet, H\. Kung, and D\. CoxTriton: an intermediate language and compiler for tiled neural network computations\.InProceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages,pp\. 10–19\.Cited by:[§7\.4](https://arxiv.org/html/2609.27373#S7.SS4.p1.1)\.
- Weiet al\.\(2022\)J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, F\. Xia, E\. Chi, Q\. V\. Le, D\. Zhou,et al\.Chain\-of\-thought prompting elicits reasoning in large language models\.Advances in neural information processing systems35,pp\. 24824–24837\.Cited by:[§1](https://arxiv.org/html/2609.27373#S1.p1.1)\.
- Xiaoet al\.\(2019\)T\. Xiao, Y\. Li, J\. Zhu, Z\. Yu, and T\. LiuSharing attention weights for fast transformer\.arXiv preprint arXiv:1906\.11024\.Cited by:[§2](https://arxiv.org/html/2609.27373#S2.p3.1)\.
- Xinet al\.\(2020\)J\. Xin, R\. Tang, J\. Lee, Y\. Yu, and J\. LinDeeBERT: dynamic early exiting for accelerating bert inference\.InProceedings of the 58th annual meeting of the association for computational linguistics,pp\. 2246–2251\.Cited by:[§2](https://arxiv.org/html/2609.27373#S2.p2.1)\.
- Xuet al\.\(2025\)F\. Xu, T\. Goyal, and E\. ChoiRefreshKV: updating small KV cache during long\-form generation\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 24878–24893\.External Links:[Link](https://aclanthology.org/2025.acl-long.1211/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1211),ISBN 979\-8\-89176\-251\-0Cited by:[§2](https://arxiv.org/html/2609.27373#S2.p3.1)\.
- Yanget al\.\(2024\)L\. Yang, K\. Lee, R\. Nowak, and D\. PapailiopoulosLooped transformers are better at learning learning algorithms\.InInternational conference on learning representations,Vol\.2024,pp\. 42195–42214\.Cited by:[§1](https://arxiv.org/html/2609.27373#S1.p1.1),[§2](https://arxiv.org/html/2609.27373#S2.p1.1)\.
- Yanget al\.\(2018\)Z\. Yang, P\. Qi, S\. Zhang, Y\. Bengio, W\. Cohen, R\. Salakhutdinov, and C\. D\. ManningHotpotQA: a dataset for diverse, explainable multi\-hop question answering\.InProceedings of the 2018 conference on empirical methods in natural language processing,pp\. 2369–2380\.Cited by:[§7\.1](https://arxiv.org/html/2609.27373#S7.SS1.SSS0.Px1.p1.1)\.
- Zhanget al\.\(2026\)X\. Zhang, H\. Wu, G\. He, J\. Shen, B\. Lyu, and Z\. ZhuModr: mixture\-of\-depth\-recurrent transformers for test\-time reasoning\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 85952–85975\.Cited by:[§2](https://arxiv.org/html/2609.27373#S2.p1.1)\.
- Zhanget al\.\(2023\)Z\. Zhang, Y\. Sheng, T\. Zhou, T\. Chen, L\. Zheng, R\. Cai, Z\. Song, Y\. Tian, C\. Ré, C\. Barrett,et al\.H2o: heavy\-hitter oracle for efficient generative inference of large language models\.Advances in neural information processing systems36,pp\. 34661–34710\.Cited by:[§2](https://arxiv.org/html/2609.27373#S2.p3.1)\.
- Zhouet al\.\(2020\)W\. Zhou, C\. Xu, T\. Ge, J\. McAuley, K\. Xu, and F\. WeiBert loses patience: fast and robust inference with early exit\.Advances in Neural Information Processing Systems33,pp\. 18330–18341\.Cited by:[§2](https://arxiv.org/html/2609.27373#S2.p2.1)\.
## Appendix AProofs for the Theoretical Analysis
### A\.1Proof of Theorem[6\.1](https://arxiv.org/html/2609.27373#S6.Thmtheorem1)
Fix a recurrent block, attention head, and query blockqq, and let𝒓q\(t\)=\[rqb\(t\)\]b\\bm\{r\}\_\{q\}^\{\(t\)\}=\[r\_\{qb\}^\{\(t\)\}\]\_\{b\}denote the block\-level attention\-mass distribution defined in Equation[5](https://arxiv.org/html/2609.27373#S5.E5)\. Let𝒓q⋆\\bm\{r\}\_\{q\}^\{\\star\}denote the limiting block\-mass distribution induced by𝑯⋆\\bm\{H\}^\{\\star\}, and write its entries in descending order asrq,\(1\)⋆≥⋯≥rq,\(NB\)⋆r\_\{q,\(1\)\}^\{\\star\}\\geq\\cdots\\geq r\_\{q,\(N\_\{B\}\)\}^\{\\star\}, whereNBN\_\{B\}is the number of admissible key blocks\. Letkq⋆=\|ℬq⋆\|k\_\{q\}^\{\\star\}=\|\\mathcal\{B\}\_\{q\}^\{\\star\}\|be the smallest integer such that
∑s=1kq⋆rq,\(s\)⋆≥η\.\\sum\_\{s=1\}^\{k\_\{q\}^\{\\star\}\}r\_\{q,\(s\)\}^\{\\star\}\\geq\\eta\.\(12\)For the cumulative\-mass support to be locally identifiable, we assume a strict ranking boundary,rq,\(kq⋆\)⋆\>rq,\(kq⋆\+1\)⋆r\_\{q,\(k\_\{q\}^\{\\star\}\)\}^\{\\star\}\>r\_\{q,\(k\_\{q\}^\{\\star\}\+1\)\}^\{\\star\}, and a strict cumulative\-mass boundary,
∑s=1kq⋆−1rq,\(s\)⋆<η<∑s=1kq⋆rq,\(s\)⋆\.\\sum\_\{s=1\}^\{k\_\{q\}^\{\\star\}\-1\}r\_\{q,\(s\)\}^\{\\star\}<\\eta<\\sum\_\{s=1\}^\{k\_\{q\}^\{\\star\}\}r\_\{q,\(s\)\}^\{\\star\}\.Define
δqrank=rq,\(kq⋆\)⋆−rq,\(kq⋆\+1\)⋆2,δq−=η−∑s=1kq⋆−1rq,\(s\)⋆,δq\+=∑s=1kq⋆rq,\(s\)⋆−η,\\delta\_\{q\}^\{\\mathrm\{rank\}\}=\\frac\{r\_\{q,\(k\_\{q\}^\{\\star\}\)\}^\{\\star\}\-r\_\{q,\(k\_\{q\}^\{\\star\}\+1\)\}^\{\\star\}\}\{2\},\\qquad\\delta\_\{q\}^\{\-\}=\\eta\-\\sum\_\{s=1\}^\{k\_\{q\}^\{\\star\}\-1\}r\_\{q,\(s\)\}^\{\\star\},\\qquad\\delta\_\{q\}^\{\+\}=\\sum\_\{s=1\}^\{k\_\{q\}^\{\\star\}\}r\_\{q,\(s\)\}^\{\\star\}\-\\eta,and let
Δq=min\{δqrank,δq−kq⋆−1,δq\+kq⋆\},\\Delta\_\{q\}=\\min\\left\\\{\\delta\_\{q\}^\{\\mathrm\{rank\}\},\\frac\{\\delta\_\{q\}^\{\-\}\}\{k\_\{q\}^\{\\star\}\-1\},\\frac\{\\delta\_\{q\}^\{\+\}\}\{k\_\{q\}^\{\\star\}\}\\right\\\},\(13\)where the second term is interpreted as\+∞\+\\inftywhenkq⋆=1k\_\{q\}^\{\\star\}=1\. By the nondegeneracy assumptions,Δq\>0\\Delta\_\{q\}\>0\.
We first show that‖𝒓q\(t\)−𝒓q⋆‖∞<Δq\\\|\\bm\{r\}\_\{q\}^\{\(t\)\}\-\\bm\{r\}\_\{q\}^\{\\star\}\\\|\_\{\\infty\}<\\Delta\_\{q\}is sufficient for exact support identification\. For anyb∈ℬq⋆b\\in\\mathcal\{B\}\_\{q\}^\{\\star\}andc∉ℬq⋆c\\notin\\mathcal\{B\}\_\{q\}^\{\\star\},
rqb\(t\)−rqc\(t\)≥rqb⋆−rqc⋆−2‖𝒓q\(t\)−𝒓q⋆‖∞\>0,r\_\{qb\}^\{\(t\)\}\-r\_\{qc\}^\{\(t\)\}\\geq r\_\{qb\}^\{\\star\}\-r\_\{qc\}^\{\\star\}\-2\\left\\\|\\bm\{r\}\_\{q\}^\{\(t\)\}\-\\bm\{r\}\_\{q\}^\{\\star\}\\right\\\|\_\{\\infty\}\>0,\(14\)so every block inℬq⋆\\mathcal\{B\}\_\{q\}^\{\\star\}remains ranked above every block outside it\. The cumulative mass of anykq⋆−1k\_\{q\}^\{\\star\}\-1blocks is bounded above by
∑s=1kq⋆−1rq,\(s\)⋆\+\(kq⋆−1\)‖𝒓q\(t\)−𝒓q⋆‖∞<η,\\sum\_\{s=1\}^\{k\_\{q\}^\{\\star\}\-1\}r\_\{q,\(s\)\}^\{\\star\}\+\(k\_\{q\}^\{\\star\}\-1\)\\left\\\|\\bm\{r\}\_\{q\}^\{\(t\)\}\-\\bm\{r\}\_\{q\}^\{\\star\}\\right\\\|\_\{\\infty\}<\\eta,while the mass of thekq⋆k\_\{q\}^\{\\star\}blocks inℬq⋆\\mathcal\{B\}\_\{q\}^\{\\star\}is bounded below by
∑s=1kq⋆rq,\(s\)⋆−kq⋆‖𝒓q\(t\)−𝒓q⋆‖∞\>η\.\\sum\_\{s=1\}^\{k\_\{q\}^\{\\star\}\}r\_\{q,\(s\)\}^\{\\star\}\-k\_\{q\}^\{\\star\}\\left\\\|\\bm\{r\}\_\{q\}^\{\(t\)\}\-\\bm\{r\}\_\{q\}^\{\\star\}\\right\\\|\_\{\\infty\}\>\\eta\.Hencekq⋆k\_\{q\}^\{\\star\}remains the smallest number of blocks whose cumulative mass reachesη\\eta, and their membership is unchanged\. Therefore,ℬq\(t,η\)=ℬq⋆\\mathcal\{B\}\_\{q\}^\{\(t,\\eta\)\}=\\mathcal\{B\}\_\{q\}^\{\\star\}whenever‖𝒓q\(t\)−𝒓q⋆‖∞<Δq\\\|\\bm\{r\}\_\{q\}^\{\(t\)\}\-\\bm\{r\}\_\{q\}^\{\\star\}\\\|\_\{\\infty\}<\\Delta\_\{q\}\.
By the localLqL\_\{q\}\-Lipschitz assumption and geometric convergence of the recurrent representations, for all sufficiently largett,
‖𝒓q\(t\)−𝒓q⋆‖∞≤Lq‖𝑯\(t\)−𝑯⋆‖≤LqCβt\.\\left\\\|\\bm\{r\}\_\{q\}^\{\(t\)\}\-\\bm\{r\}\_\{q\}^\{\\star\}\\right\\\|\_\{\\infty\}\\leq L\_\{q\}\\left\\\|\\bm\{H\}^\{\(t\)\}\-\\bm\{H\}^\{\\star\}\\right\\\|\\leq L\_\{q\}C\\beta^\{t\}\.\(15\)Since0<β<10<\\beta<1, there exists a finiteTS,qT\_\{S,q\}such thatLqCβt<ΔqL\_\{q\}C\\beta^\{t\}<\\Delta\_\{q\}for everyt≥TS,qt\\geq T\_\{S,q\}\. Thus,ℬq\(t,η\)=ℬq⋆\\mathcal\{B\}\_\{q\}^\{\(t,\\eta\)\}=\\mathcal\{B\}\_\{q\}^\{\\star\}for everyt≥TS,qt\\geq T\_\{S,q\}\. If the Lipschitz bound is valid from recurrent stepT0,qT\_\{0,q\}onward, one valid choice is
TS,q=max\{T0,q,⌊log\(LqC/Δq\)−logβ⌋\+1\},T\_\{S,q\}=\\max\\left\\\{T\_\{0,q\},\\left\\lfloor\\frac\{\\log\(L\_\{q\}C/\\Delta\_\{q\}\)\}\{\-\\log\\beta\}\\right\\rfloor\+1\\right\\\},\(16\)with the second term replaced by00wheneverLqC<ΔqL\_\{q\}C<\\Delta\_\{q\}\. Consequently,
TS,q=O\(log\(LqC/Δq\)−logβ\),T\_\{S,q\}=O\\\!\\left\(\\frac\{\\log\(L\_\{q\}C/\\Delta\_\{q\}\)\}\{\-\\log\\beta\}\\right\),up to the finite entry time into the local Lipschitz neighborhood\. For a finite collection of recurrent blocks, attention heads, and query blocks, takingTS=maxqTS,qT\_\{S\}=\\max\_\{q\}T\_\{S,q\}over all such routing units yields a finite recurrent depth after which all nondegenerate block supports are simultaneously identified\. ∎
### A\.2Proof of Proposition[1](https://arxiv.org/html/2609.27373#Thmproposition1)
For query blockqq, define the routing region
ℛq\(ℬq⋆\)=\{𝑯:ℬq\(η\)\(𝑯\)=ℬq⋆\}\.\\mathcal\{R\}\_\{q\}\(\\mathcal\{B\}\_\{q\}^\{\\star\}\)=\\left\\\{\\bm\{H\}:\\mathcal\{B\}\_\{q\}^\{\(\\eta\)\}\(\\bm\{H\}\)=\\mathcal\{B\}\_\{q\}^\{\\star\}\\right\\\}\.The proof of Theorem[6\.1](https://arxiv.org/html/2609.27373#S6.Thmtheorem1)establishes that any representation satisfying
‖𝒓q\(𝑯\)−𝒓q⋆‖∞<Δq\\left\\\|\\bm\{r\}\_\{q\}\(\\bm\{H\}\)\-\\bm\{r\}\_\{q\}^\{\\star\}\\right\\\|\_\{\\infty\}<\\Delta\_\{q\}induces cumulative\-mass supportℬq⋆\\mathcal\{B\}\_\{q\}^\{\\star\}\. By localLqL\_\{q\}\-Lipschitz continuity,
‖𝒓q\(𝑯\)−𝒓q⋆‖∞≤Lq‖𝑯−𝑯⋆‖,\\left\\\|\\bm\{r\}\_\{q\}\(\\bm\{H\}\)\-\\bm\{r\}\_\{q\}^\{\\star\}\\right\\\|\_\{\\infty\}\\leq L\_\{q\}\\left\\\|\\bm\{H\}\-\\bm\{H\}^\{\\star\}\\right\\\|,and therefore
𝔹\(𝑯⋆,ΔqLq\)⊆ℛq\(ℬq⋆\)\.\\mathbb\{B\}\\\!\\left\(\\bm\{H\}^\{\\star\},\\frac\{\\Delta\_\{q\}\}\{L\_\{q\}\}\\right\)\\subseteq\\mathcal\{R\}\_\{q\}\(\\mathcal\{B\}\_\{q\}^\{\\star\}\)\.\(17\)If𝑯\(t\)→𝑯⋆\\bm\{H\}^\{\(t\)\}\\rightarrow\\bm\{H\}^\{\\star\}, then there exists a finiteTTsuch that
𝑯\(t\)∈𝔹\(𝑯⋆,ΔqLq\)\\bm\{H\}^\{\(t\)\}\\in\\mathbb\{B\}\\\!\\left\(\\bm\{H\}^\{\\star\},\\frac\{\\Delta\_\{q\}\}\{L\_\{q\}\}\\right\)for everyt≥Tt\\geq T, and henceℬq\(t,η\)=ℬq⋆\\mathcal\{B\}\_\{q\}^\{\(t,\\eta\)\}=\\mathcal\{B\}\_\{q\}^\{\\star\}thereafter\. Eventual support stabilization therefore follows from convergence to a nondegenerate limiting representation\.
The converse does not hold\. SinceΔq/Lq\>0\\Delta\_\{q\}/L\_\{q\}\>0, the routing region contains a nontrivial neighborhood of𝑯⋆\\bm\{H\}^\{\\star\}and hence distinct representations𝑯1≠𝑯2\\bm\{H\}\_\{1\}\\neq\\bm\{H\}\_\{2\}that induce the same block supportℬq⋆\\mathcal\{B\}\_\{q\}^\{\\star\}\. A recurrent trajectory may therefore continue to move within this region while preserving an identical routing support\. Thus, without additional constraints on the recurrent dynamics, support stabilization is insufficient to imply representation convergence\. ∎
## Appendix BExperimental Details
#### Models, benchmarks, and evaluation cohorts\.
Our primary experiments use the publicly available Huginn recurrent language model with recurrent depthT=32T=32\. We additionally evaluate Recurrent\-Llama\-T32, a recurrentized Llama\-3\.2 backbone trained with3232recurrent steps, to test whether the observed routing dynamics and behavioral effects generalize beyond the primary checkpoint\. The Recurrent\-Llama model contains six recurrent blocks and uses grouped\-query attention with3232query heads and88key–value heads\. Because it is not instruction\-tuned, we use a single fixed plain\-completion prompt for all Recurrent\-Llama conditions; prompting is held fixed across Full,WISE, and all static controls within each backbone\. Model parameters, numerical precision, decoding settings, and recurrent initialization are held fixed across matched inference conditions, and no model training or parameter update is performed\. Our primary benchmark is the HotpotQA distractor validation split, evaluated with its original questions, answers, and benchmark\-native contexts\. Benchmark\-level behavioral comparisons use fixed 100\-example cohorts from HotpotQA and 2WikiMultiHopQA, while mechanism diagnostics use fixed 30\-example cohorts from HotpotQA and GSM8K\. No example is selected or removed according to Full orWISEperformance, and all paired methods within a backbone receive identical inputs and recurrent initialization\.
#### WISEconfiguration\.
AllWISEexperiments use the block\-structured configuration defined in Section[5](https://arxiv.org/html/2609.27373#S5)\. We set the recurrent depth toT=32T=32, block size toB=32B=32, discovery depth totd=12t\_\{\\mathrm\{d\}\}=12, and block\-mass threshold toη=0\.95\\eta=0\.95\. For each recurrent block, attention head, and query blockqq, the reused working set is
ℬq=⋃t=td−3tdℬq\(t,η\),\\mathcal\{B\}\_\{q\}=\\bigcup\_\{t=t\_\{\\mathrm\{d\}\}\-3\}^\{t\_\{\\mathrm\{d\}\}\}\\mathcal\{B\}\_\{q\}^\{\(t,\\eta\)\},\(18\)corresponding to the union of direct block\-level95%95\\%\-mass supports from recurrent steps99–1212\. Support is constructed directly in block space rather than by selecting token\-level supports and subsequently rounding them to blocks\. Fort\>tdt\>t\_\{\\mathrm\{d\}\},ℬq\\mathcal\{B\}\_\{q\}remains fixed while𝑯\(t\)\\bm\{H\}^\{\(t\)\},𝑸\(t\)\\bm\{Q\}^\{\(t\)\},𝑲\(t\)\\bm\{K\}^\{\(t\)\},𝑽\(t\)\\bm\{V\}^\{\(t\)\}, and the within\-support attention weights continue to be recomputed\. The sameB=32B=32,td=12t\_\{\\mathrm\{d\}\}=12,η=0\.95\\eta=0\.95, and steps\-99–1212discovery window are used for Huginn and Recurrent\-Llama\-T32 without backbone\-specific tuning\.
#### Intervention controls\.
We compareWISEagainst three early\-static controls that use the same block representation and the same intervention timing\. For each recurrent block, attention head, and query blockqq, let𝒂q\(t\)\\bm\{a\}\_\{q\}^\{\(t\)\}denote the causal block\-attention masses at recurrent steptt, and letℬq\(t,ρ\)\\mathcal\{B\}\_\{q\}^\{\(t,\\rho\)\}denote the smallest set of highest\-mass admissible key blocks whose cumulative mass under𝒂q\(t\)\\bm\{a\}\_\{q\}^\{\(t\)\}is at leastρ\\rho\. TheStatic\-95control freezes the direct95%95\\%\-mass support identified at the first recurrent step,
𝒮qStatic95=ℬq\(1,0\.95\)\.\\mathcal\{S\}\_\{q\}^\{\\mathrm\{Static95\}\}=\\mathcal\{B\}\_\{q\}^\{\(1,0\.95\)\}\.\(19\)
To distinguish the effect of recurrent discovery from simply retaining more early attention mass,MassMatcheduses the recurrence\-1 ranking but matches the amount of recurrence\-1 attention mass covered by the finalWISEworking set\. Writing
mqWISE=∑j∈𝒲qaq,j\(1\),𝒲q=⋃t=td−3tdℬq\(t,η\),m\_\{q\}^\{\\mathrm\{WISE\}\}=\\sum\_\{j\\in\\mathcal\{W\}\_\{q\}\}a\_\{q,j\}^\{\(1\)\},\\qquad\\mathcal\{W\}\_\{q\}=\\bigcup\_\{t=t\_\{\\mathrm\{d\}\}\-3\}^\{t\_\{\\mathrm\{d\}\}\}\\mathcal\{B\}\_\{q\}^\{\(t,\\eta\)\},\(20\)the mass\-matched control is
𝒮qMassMatched=ℬq\(1,mqWISE\)\.\\mathcal\{S\}\_\{q\}^\{\\mathrm\{MassMatched\}\}=\\mathcal\{B\}\_\{q\}^\{\(1,m\_\{q\}^\{\\mathrm\{WISE\}\}\)\}\.\(21\)
Finally,SizeMatchedcontrols exactly for support cardinality\. LetKq=\|𝒲q\|K\_\{q\}=\|\\mathcal\{W\}\_\{q\}\|\. Using the recurrence\-1 block scores, it retains theKqK\_\{q\}highest\-scoring admissible blocks,
𝒮qSizeMatched=TopK\(𝒂q\(1\),Kq\)\.\\mathcal\{S\}\_\{q\}^\{\\mathrm\{SizeMatched\}\}=\\operatorname\{TopK\}\\left\(\\bm\{a\}\_\{q\}^\{\(1\)\},K\_\{q\}\\right\)\.\(22\)Thus,MassMatchedmatches the recurrence\-1 attention mass retained byWISE, whereasSizeMatchedmatches its block cardinality exactly for every local routing unit\.
We additionally evaluate two interventions that probe what remains necessary after the discovery phase\.Freeze\-A@12uses exactly the same final working setWqW\_\{q\}asWISE, but also freezes the within\-support attention distribution at the discovery endpoint\. LetA\(12\)A^\{\(12\)\}denote the unrestricted post\-softmax attention matrix at recurrent step 12\. Using the query\- and key\-block mappingsq\(i\)q\(i\)andb\(j\)b\(j\)from Section 5\.2, we define the restricted and renormalized distribution
A¯ij\(12\)=A\(12\)ij\[b\(j\)∈Wq\(i\)\]∑k∈𝒦iA\(12\)ik\[b\(k\)∈Wq\(i\)\],\\bar\{A\}^\{\(12\)\}\_\{ij\}=\\frac\{A^\{\(12\)\}\_\{ij\}\\,\\mathbf\{1\}\\\!\\left\[b\(j\)\\in W\_\{q\(i\)\}\\right\]\}\{\\sum\_\{k\\in\\mathcal\{K\}\_\{i\}\}A^\{\(12\)\}\_\{ik\}\\,\\mathbf\{1\}\\\!\\left\[b\(k\)\\in W\_\{q\(i\)\}\\right\]\},where𝒦i\\mathcal\{K\}\_\{i\}denotes the keys admissible under the native attention mask\. For allt\>tdt\>t\_\{d\},Freeze\-A@12reuses this distribution,
A~\(t\)=A¯\(12\),O~\(t\)=A¯\(12\)V\(t\),\\widetilde\{A\}^\{\(t\)\}=\\bar\{A\}^\{\(12\)\},\\qquad\\widetilde\{O\}^\{\(t\)\}=\\bar\{A\}^\{\(12\)\}V^\{\(t\)\},while values, hidden representations, and the remaining recurrent computation continue to evolve\. Thus,Freeze\-A@12andWISEuse exactly the same routing support and differ only in whether within\-support attention weights are recomputed after discovery\.
Truncate@12instead follows the unrestricted Full trajectory throughtd=12t\_\{d\}=12and then terminates the recurrent core without executing steps 13–32\. The model’s native post\-recurrent computation, including its exit path, final normalization, and output head, is otherwise left unchanged\.
All intervention conditions are matched to the unrestricted Full trajectory throughtd=12t\_\{d\}=12\. ForStatic\-95,MassMatched,SizeMatched, andWISE, recurrent computation continues throughT=32T=32using the corresponding fixed support while queries, keys, values, hidden states, and within\-support attention weights remain dynamic\.Freeze\-A@12likewise continues throughT=32T=32, but additionally reuses the fixed distributionA¯\(12\)\\bar\{A\}^\{\(12\)\}defined above\.Truncate@12instead terminates recurrent computation immediately after step 12\. These interventions therefore separately test support composition, within\-support reweighting, and the value of continued recurrent refinement after the discovery depth\.
#### Evaluation metrics\.
We report standard normalized exact match \(EM\) and token\-level F1 for downstream answer quality\. Because the main comparisons are paired at the example level, we additionally report paired F1 differences with95%95\\%bootstrap confidence intervals and answer\-change counts relative to the matched Full run\. Controlled interventions additionally report paired changes in teacher\-forced gold\-answer adjusted loss as a continuous measure of behavioral perturbation, where a positive change denotes worse behavior relative to the reference condition\. To characterize routing structure, we report block density, defined as the fraction of admissible query–key blocks retained during reuse, and retained future attention mass, defined as the fraction of unrestricted Full\-attention mass att\>tdt\>t\_\{\\mathrm\{d\}\}that falls inside the frozen support\. We compute these metrics using the same definitions for both backbones\.
#### Mechanism diagnostics\.
The convergence analysis uses fixed 30\-example cohorts from HotpotQA and GSM8K and tracks diagnostic attention support, attention distributions, hidden representations, and pre\-projection attention outputs across recurrent depth\. For query positionii, the diagnostic support is the causal Top\-min\(16,i\+1\)\\min\(16,i\+1\)attended\-key set, which is distinct from the directB=32B=32cumulative\-mass support deployed byWISE\. Support stability is measured by Jaccard similarity between consecutive supports, with convergence defined by similarity≥0\.90\\geq 0\.90\. For𝑨\(t\)\\bm\{A\}^\{\(t\)\},𝑯\(t\)\\bm\{H\}^\{\(t\)\}, and𝑶\(t\)\\bm\{O\}^\{\(t\)\}, consecutive recurrent states are compared using Jensen–Shannon divergence, relativeℓ2\\ell\_\{2\}change, and relativeℓ2\\ell\_\{2\}change, respectively; each convergence threshold is set to0\.050\.05times the example\-specific early\-trajectory scale computed from the first four finite transition distances of that quantity\. All four criteria must hold for three consecutive transitions, andτX\\tau\_\{X\}is the first transition in the qualifying run; trajectories that do not converge byT=32T=32are recorded asτX=33\\tau\_\{X\}=33\. We apply the same diagnostic definitions and thresholds to Huginn and Recurrent\-Llama\-T32 without retuning\. These diagnostics characterize recurrent dynamics only and do not determine theWISEswitching depth or working\-set configuration\. We separately evaluate the stability of the exact block support deployed byWISEin Appendix[C\.2](https://arxiv.org/html/2609.27373#A3.SS2)\.
#### Cross\-backbone replication\.
The Recurrent\-Llama replication uses a model\-specific instrumentation adapter to expose recurrent attention, hidden representations, and pre\-projection attention outputs while preserving the model’s native computation\. The adapter accounts for Llama\-style rotary position embeddings and grouped\-query attention but does not modify model parameters, recurrent depth, attention scores, or support construction\. The behavioral replication evaluates the original five support\-selection conditions—Full,Static\-95,MassMatched,SizeMatched, andWISE—using the same definitions and intervention timing as on the primary Huginn backbone\. In particular,SizeMatchedexactly matches the number of retained blocks used byWISEfor every local routing unit\.Freeze\-A@12andTruncate@12are additional causal interventions evaluated only on the primary Huginn backbone\.
#### Context\-length scaling\.
Matched HotpotQA quality scaling is performed only on the primary Huginn backbone and uses the same 100\-example cohort at approximately512512,11K,22K, and44K context lengths\. Required answer evidence is preserved at every length, while additional context is formed from natural benchmark passages rather than synthetic padding or duplicated text\. Full andWISEreceive exactly the same input at each context length\. Because absolute Full\-model difficulty varies with context composition, the primary scaling statistic is the paired differenceΔF1=F1WISE−F1Full\\Delta\\mathrm\{F1\}=\\mathrm\{F1\}\_\{\\mathrm\{WISE\}\}\-\\mathrm\{F1\}\_\{\\mathrm\{Full\}\}rather than the absolute F1 trajectory across lengths\. The Huginn checkpoint supports at most40964096positions through its rotary\-position representation; consequently,44K is the maximum valid model\-level context length considered in the scaling study, and we report no model\-level88K quality result\.
#### Systems evaluation\.
Systems measurements are performed only for the primary Huginn configuration on NVIDIA RTX A6000 GPUs\. The practical dense baseline uses PyTorch scaled dot\-product attention dispatching to native FlashAttention, while the finalWISEimplementation uses a custom Triton kernel specialized for exactB=32B=32block\-structured reuse\. Systems measurements use a fixed 30\-example cohort of saved HotpotQA working\-set workloads, distinct from the 100\-example quality\-scaling cohort\. Headline measurements use 10 balanced timing passes on an uncontended RTX A6000 under a fixed software stack, with latency reported as the median across passes and the corresponding interquartile range retained to characterize timing variability\. We report three complementary measurements: setup\-inclusive late\-reuse latency, which compares one\-time sparse\-schedule preparation plus 20 reuse calls against 20 native FlashAttention calls; completeT=32T=32attention\-trajectory latency, which includes the 12 unrestricted discovery steps together with schedule preparation and 20 reuse steps; and a controlled same\-backend comparison in which the optimized exact\-B=32B=32kernel executes either all causally valid blocks or only theWISEsupport\. Latency and answer\-quality results are therefore treated as complementary measurements rather than as measurements on the same evaluation cohort\.
#### Sparse\-kernel implementation\.
The custom kernel consumes the exact block support discovered byWISEwithout introducing an additional approximation\. For each recurrent block, attention head, and query block, the active key blocks are encoded as a reusable GPU\-resident CSR schedule consisting of row offsets and active block indices\. This schedule is constructed once after the discovery phase and reused unchanged across all subsequent recurrent steps\. At execution time, the kernel skips blocks absent from the sparse schedule while fusing query–key scoring, causal masking, online softmax normalization, and value aggregation without materializing the dense attention matrix\. Queries, keys, and values are recomputed at every recurrent step, so only the routing support is reused, and the measured sparse path contains no dense\-attention fallback\. For the primary Huginn workload, attention contains four recurrent attention layers with 55 query and key–value heads and native head dimensiondh=96d\_\{h\}=96\. The final implementation specializes arithmetic for this native head dimension and uses two warps and three pipeline stages\. An equal\-budget work\-order sweep selects the natural within\-head row order for the final configuration; these implementation choices affect only kernel realization and do not alter the discovered support or attention semantics\. Before timing, the optimized kernel is validated against explicit exact masked attention under the same BF16 correctness criterion used throughout the systems study\. Setup\-inclusive measurements include GPU sparse\-schedule construction and associated execution\-metadata preparation, whereas prepared\-schedule measurements exclude only this one\-time reusable setup cost\. All reported Full andWISEcomparisons use the same hardware and software configuration, identical inputs, and identical warm\-up, synchronization, and timing procedures\.
#### Reproducibility\.
All example identities, context variants, intervention conditions, predictions, working\-set masks, and timing records are stored in deterministic manifests\. We verify that paired methods use identical inputs and recurrent initialization within each backbone, thatWISEbegins support reuse only aftertdt\_\{\\mathrm\{d\}\}, that the discovered block support remains fixed throughout reuse, and that hidden states, queries, keys, values, and within\-support attention weights continue to evolve\. For the cross\-backbone replication, we additionally record the Recurrent\-Llama checkpoint revision, tokenizer, fixed plain\-completion prompt, grouped\-query\-attention instrumentation, and all exact support\-cardinality audits\. Exact checkpoint revisions, preprocessing details, prompts, generation parameters, hardware configuration, and implementation code will be released with the final artifact\.
## Appendix CRobustness Analysis
### C\.1Sensitivity to Convergence Criteria
The routing–representation separation is robust to matched changes in the convergence criteria\. We evaluate five settings ranging from very strict to very loose, varying the strictness of the routing and representation diagnostics consistently\. To summarize the separation, we define the routing–representation gap as
G=min\(τH,τO\)−max\(τS,τA\),G=\\min\(\\tau\_\{H\},\\tau\_\{O\}\)\-\\max\(\\tau\_\{S\},\\tau\_\{A\}\),such thatG\>0G\>0requires both routing quantities to stabilize before either representation quantity\. Across all five matched settings,G\>0G\>0for every diagnostic example:30/3030/30on HotpotQA and30/3030/30on GSM8K\. The mean gap ranges from4\.334\.33to6\.676\.67recurrent steps on HotpotQA and from7\.807\.80to17\.5017\.50on GSM8K\. The largest GSM8K gaps under the strictest criteria should be interpreted cautiously because representation convergence is strongly right\-censored at theT=32T=32endpoint\. These results show that the qualitative routing–representation separation does not depend on the particular default threshold choice\.
\(a\) HotpotQA
\(b\) GSM8K
\(c\) Routing–representation gap
Figure 3:Sensitivity of the convergence analysis to threshold choice\.\(a\)–\(b\)Fraction of examples classified as converged by recurrent step under five*matched*threshold settings: very loose \(VL\), loose \(L\), default \(D\), strict \(S\), and very strict \(VS\)\. Across these matched settings, routing\-related quantities consistently converge earlier than representation\-related quantities, although the exact convergence times and step gaps vary with threshold strictness\.\(c\)Mean routing–representation gap under the same matched settings on HotpotQA and GSM8K\. Positive values indicate earlier routing convergence\. The gap remains positive throughout the matched sweep, supporting a robust routing–representation separation under reasonable threshold perturbations, while not implying threshold\-invariance under arbitrary asymmetric criteria\.The ordering is not invariant to arbitrary cross\-quantity calibration\. In an intentionally unfavorable but still reasonable asymmetric setting, routing is evaluated using very strict criteria while representation is evaluated using loose criteria\. On HotpotQA, the resulting mean convergence steps areτS=18\.90\\tau\_\{S\}=18\.90,τA=13\.20\\tau\_\{A\}=13\.20,τH=17\.33\\tau\_\{H\}=17\.33, andτO=18\.20\\tau\_\{O\}=18\.20, giving a mean gap ofG=−1\.57G=\-1\.57; only1/301/30examples retainsG\>0G\>0\. We therefore do not claim a threshold\-independent ordering across heterogeneous quantities\. Rather, the empirical claim is that routing stabilizes earlier under the reported criteria and throughout the tested family of consistently calibrated threshold perturbations\.
Table 3:Boundary of the convergence claim on HotpotQA\.Matched calibration preserves the routing–representation separation, while deliberately asymmetric calibration can reverse the measured ordering\.
### C\.2Stability of the Deployed Working Set
The diagnostic support used in Figure[1](https://arxiv.org/html/2609.27373#S1.F1)[1](https://arxiv.org/html/2609.27373#S1.F1)is token\-level, whereasWISEreuses a four\-step union of directB=32B=32,η=0\.95\\eta=0\.95cumulative\-mass block supports\. We therefore evaluate the stability of the exact working\-set construction used by the method\. For each recurrent steptt, define the candidate rolling working set
𝒲q\(t\)=⋃s=max\(1,t−3\)tℬq\(s,η\)\.\\mathcal\{W\}\_\{q\}^\{\(t\)\}=\\bigcup\_\{s=\\max\(1,t\-3\)\}^\{t\}\\mathcal\{B\}\_\{q\}^\{\(s,\\eta\)\}\.\(23\)For an exampleee, let𝒲e\(t\)\\mathcal\{W\}\_\{e\}^\{\(t\)\}denote the collection of active key blocks across all recurrent blocks, attention heads, and query blocks\. We measure adjacent rolling\-set stability as
Jroll\(t\)=1N∑e=1N\|𝒲e\(t−1\)∩𝒲e\(t\)\|\|𝒲e\(t−1\)∪𝒲e\(t\)\|\.J\_\{\\mathrm\{roll\}\}^\{\(t\)\}=\\frac\{1\}\{N\}\\sum\_\{e=1\}^\{N\}\\frac\{\\left\|\\mathcal\{W\}\_\{e\}^\{\(t\-1\)\}\\cap\\mathcal\{W\}\_\{e\}^\{\(t\)\}\\right\|\}\{\\left\|\\mathcal\{W\}\_\{e\}^\{\(t\-1\)\}\\cup\\mathcal\{W\}\_\{e\}^\{\(t\)\}\\right\|\}\.\(24\)This statistic exactly matches the block\-structured working\-set construction used byWISE; fort\>12t\>12, the rolling sets are reconstructed from the saved unrestricted Full\-attention trajectory for diagnostic purposes, whereas deployedWISEfreezes𝒲\(12\)\\mathcal\{W\}^\{\(12\)\}after discovery\.
The rolling working set becomes highly stable near the chosen discovery depth\. Att=12t=12, the mean adjacent rolling\-set Jaccard is0\.96740\.9674on the3030\-example44K HotpotQA audit cohort\. Under the0\.900\.90\-for\-three\-transitions criterion, all30/3030/30examples first satisfy the stability condition att=12t=12\. Under the stricter0\.950\.95criterion, none satisfy it byt=12t=12, while all30/3030/30satisfy it att=13t=13\. Thus,td=12t\_\{\\mathrm\{d\}\}=12lies at the onset of a high\-stability regime rather than marking an exact convergence boundary\.
\(a\) Rolling working\-set stability
\(b\) Fraction of converged examples
Figure 4:The deployed block support becomes highly stable near theWISEdiscovery depth\.\(a\)Adjacent\-step Jaccard similarity between four\-step rolling working sets constructed from the exactB=32B=32,η=0\.95\\eta=0\.95block supports used byWISE\. The discovery depthtd=12t\_\{\\mathrm\{d\}\}=12lies near the onset of a high\-stability regime\.\(b\)Fraction of examples whose rolling\-set Jaccard satisfies the stability criterion for three consecutive transitions\. All3030examples satisfy the0\.900\.90criterion byt=12t=12, while all satisfy the stricter0\.950\.95criterion byt=13t=13\.Local stability does not imply that the support becomes globally fixed\. Comparing the working set selected at the switching point with the hypothetical rolling set constructed from the final Full\-attention steps gives
J\(𝒲\(12\),𝒲\(32\)\)=0\.8987\.J\\\!\\left\(\\mathcal\{W\}^\{\(12\)\},\\mathcal\{W\}^\{\(32\)\}\\right\)=0\.8987\.\(25\)This slow cumulative drift is consistent with the role of the temporal union:WISEdoes not assume that routing becomes exactly constant at a single recurrent step, but instead freezes a working set after routing has entered a low\-drift regime\. Together with the high future Full\-attention mass retained byWISE, this analysis provides a direct method\-aligned bridge between the token\-level mechanism diagnostic and the block\-structured support actually reused during inference\.
### C\.3Cross\-Backbone Replication
We next test whether the observed separation between routing stabilization and representation refinement is specific to the primary Huginn checkpoint\. We repeat the mechanism analysis on Recurrent\-Llama\-T32 using the same recurrent depthT=32T=32, the same diagnostic definitions, and the same convergence thresholds described in Appendix[B](https://arxiv.org/html/2609.27373#A2)\. The replication uses fixed 30\-example HotpotQA and GSM8K cohorts and does not retune any convergence criterion for the new backbone\. As shown in Table[4](https://arxiv.org/html/2609.27373#A3.T4), the routing–representation separation replicates strongly\. On HotpotQA, attention support and attention distributions converge at mean steps5\.075\.07and7\.007\.00, respectively, while the hidden state and attention output converge at8\.938\.93and16\.1316\.13\. On GSM8K, the corresponding values are3\.873\.87,7\.107\.10,9\.639\.63, and31\.2031\.20\. Most importantly, every diagnostic example exhibits a positive routing–representation gap:30/3030/30on HotpotQA and30/3030/30on GSM8K\. The precise ordering within routing differs from the primary Huginn backbone—support stabilizes before the attention distribution on Recurrent\-Llama\-T32—but the broader separation between routing stabilization and continued representation refinement is preserved\.
Table 4:The routing–representation separation replicates on Recurrent\-Llama\-T32\.Mean convergence steps use the same diagnostic definitions and thresholds as the primary Huginn analysis\. A positive gap requires both routing quantities to stabilize before both representation quantities\.†For GSM8K,27/3027/30attention\-output trajectories do not converge byT=32T=32and are recorded asτO=33\\tau\_\{O\}=33when computing the mean;31\.2031\.20therefore reflects right censoring rather than an observed convergence time\.
We also repeat the main behavioral intervention study using the same frozenWISEconfiguration and the same Full,Static\-95,MassMatched, andSizeMatchedcontrols\. ExactSizeMatchedcardinality matching succeeds for all1,478,7841\{,\}478\{,\}784audited local routing units\. The result is more nuanced than on the primary backbone\. On HotpotQA,WISEobtains F10\.21590\.2159, compared with0\.20750\.2075for Full and0\.20760\.2076forSizeMatched, corresponding to a pairedWISE–SizeMatcheddifference of\+0\.0083\+0\.0083\. On 2WikiMultiHopQA, the corresponding values are0\.20850\.2085,0\.21720\.2172, and0\.21280\.2128, giving a paired difference of−0\.0043\-0\.0043\. Thus, unlike on Huginn, the cross\-backbone replication does not establish a downstream F1 advantage of recurrently discovered support over an exactly size\-matched early\-static support\.
Table 5:Behavioral replication on Recurrent\-Llama\-T32\.SizeMatchedandWISEuse identical support cardinality for every routing unit\. Future denotes retained late Full\-attention mass\. Paired confidence intervals are reported in the text\.
### C\.4When Does Delayed Discovery Help?
The cross\-backbone comparison suggests a natural boundary condition for the benefit ofWISE\. Although Recurrent\-Llama\-T32 exhibits the same routing\-before\-representation separation, its routing support stabilizes substantially earlier than on Huginn\. On HotpotQA, the mean token\-support convergence step decreases from13\.9713\.97on Huginn to5\.075\.07on Recurrent\-Llama\-T32; on GSM8K, it decreases from10\.0710\.07to3\.873\.87\. This earlier stabilization is reflected in how predictive an early static support is of later Full attention\. On Huginn, the exactly size\-matched early\-static support retains94\.65%94\.65\\%and94\.87%94\.87\\%of late Full\-attention mass on HotpotQA and 2WikiMultiHopQA, respectively\. On Recurrent\-Llama\-T32, the corresponding values increase to96\.13%96\.13\\%and96\.41%96\.41\\%, despite lower retained block density\. Consequently, recurrent discovery increases late\-mass coverage by approximately22percentage points on Huginn but only about0\.40\.4percentage points on Recurrent\-Llama\-T32\. These results suggest that delayed recurrent discovery is most useful when routing continues to reorganize during early recurrence\. When an early support already predicts later attention well, an early static support can approximate the eventually discovered working set more closely, leaving less room forWISEto improve support composition\. We treat this interpretation as suggestive rather than definitive: the saved cross\-backbone trajectories do not contain the exact per\-unitB=32B=32support identities and attention masses required to measure direct early\-to\-late block\-level Jaccard similarity or head\-level routing specialization\. Moreover, attention distributions and hidden representations also converge more quickly on Recurrent\-Llama\-T32, so the current evidence does not isolate stable routing roles from broader differences in recurrent dynamics\.
Table 6:Early routing is more predictive on Recurrent\-Llama\-T32\.SizeMatchedFuture is the fraction of late unrestricted Full\-attention mass retained by the exactly size\-matched early\-static support\. Gain is the additional late\-mass coverage obtained byWISE\.
### C\.5WISEHyperparameter Sensitivity
Finally, we test whether the frozenWISEconfiguration is an isolated operating point on the primary Huginn backbone\. Starting fromη=0\.95\\eta=0\.95,td=12t\_\{\\mathrm\{d\}\}=12, temporal\-union widthw=4w=4, andB=32B=32, we vary one parameter at a time while keeping all others fixed\. This analysis is intended as a local robustness check rather than hyperparameter optimization; all initial comparisons use the same deterministicN=50N=50HotpotQA native\-context cohort, for which matched Full inference obtains F10\.17980\.1798\. As shown in Table[7](https://arxiv.org/html/2609.27373#A3.T7), no tested neighboring configuration clearly dominates the frozen default\. Varyingη\\etaexposes the expected sparsity–coverage tradeoff: decreasingη\\etafrom0\.950\.95to0\.900\.90reduces block density from0\.4470\.447to0\.2980\.298while reducing retained future attention mass from0\.9660\.966to0\.9330\.933, whereas increasingη\\etato0\.980\.98raises future\-mass retention to0\.9870\.987at the cost of density increasing to0\.6510\.651\. Nearby discovery depths and temporal\-union widths similarly yield comparable downstream behavior while changing the retained routing structure\.
Table 7:Local sensitivity ofWISEon HotpotQA \(N=50N=50\)\.Each row varies one parameter around the frozen default\. Density is the retainedB=32B=32block fraction and Future Mass is the fraction of later unrestricted Full\-attention mass covered by the frozen support\.The initially larger degradation observed forw=1w=1was evaluated on the untouched second half of the canonicalN=100N=100cohort\. On the first5050examples,w=1w=1minus the default yieldedΔF1=−0\.0216\\Delta\\mathrm\{F1\}=\-0\.0216with 95% CI\[−0\.0458,−0\.0029\]\[\-0\.0458,\-0\.0029\], whereas the untouched second5050yielded\+0\.0089\+0\.0089with 95% CI\[−0\.0111,\+0\.0425\]\[\-0\.0111,\+0\.0425\]\. Pooling all100100examples givesΔF1=−0\.0063\\Delta\\mathrm\{F1\}=\-0\.0063with 95% CI\[−0\.0234,\+0\.0140\]\[\-0\.0234,\+0\.0140\], so the initial quality deficit does not replicate\. We therefore interpret the sensitivity analysis as evidence that the default is not an obviously fragile isolated point, rather than evidence of hyperparameter insensitivity or optimality\. These experiments are restricted to benchmark\-native contexts and one\-at\-a\-time perturbations; they do not establish long\-context hyperparameter robustness or robustness to joint interactions amongη\\eta,tdt\_\{\\mathrm\{d\}\}, andww\.
### C\.6Block\-Size Tradeoff
At the finest extreme, token\-level support would maximize routing granularity but produce highly irregular sparsity; block structure instead exposes a controllable systems tradeoff between sparsity and executable regularity\. The block sizeBBcontrols the granularity at which the discovered working set can be exploited by sparse execution\. We therefore evaluateB∈\{16,32,64,128\}B\\in\\\{16,32,64,128\\\}while keeping the remainingWISEconfiguration fixed:η=0\.95\\eta=0\.95,td=12t\_\{d\}=12, and the four\-step discovery window spanning recurrent steps 9–12, and use an equal implementation\-tuning budget for each block size\. Finer blocks expose more routing sparsity, whereas coarser blocks conservatively retain more neighboring tokens\. However, block density alone does not determine execution efficiency, since the granularity of the resulting sparse computation also affects kernel execution\. All four implementations are validated against explicit exact masked attention under the same BF16 correctness criterion, with maximum absolute error below0\.0076430\.007643and no support or causal\-mask mismatches\. ForB=64B=64andB=128B=128, the independently selected block supports are executed using exact32×3232\\times 32physical subtiles without changing support membership or softmax semantics\.
Table[8](https://arxiv.org/html/2609.27373#A3.T8)summarizes the full tradeoff\. On the matchedN=30N=302K systems panel,B=16B=16produces the sparsest support at31\.12%31\.12\\%density but reaches only a1\.141×1\.141\\timessetup\-inclusive speedup\. Increasing toB=32B=32raises density modestly to37\.51%37\.51\\%while achieving the highest speedup of1\.656×1\.656\\times\. CoarserB=64B=64andB=128B=128retain44\.71%44\.71\\%and54\.44%54\.44\\%of blocks and reach1\.440×1\.440\\timesand1\.236×1\.236\\times, respectively\. The same ordering persists at 4K, whereB=32B=32reaches1\.769×1\.769\\times, compared with1\.146×1\.146\\times,1\.554×1\.554\\times, and1\.394×1\.394\\timesforB=16B=16,6464, and128128\. Thus, finer sparsity does not necessarily translate into faster execution:B=32B=32forms a robust systems\-facing knee between support sparsity and executable regularity\.
Table 8:Block\-size ablation forWISE\.Quality and routing statistics use the matchedN=100N=100benchmark cohorts, while systems statistics use the matchedN=30N=302K panel\. Future denotes the fraction of future Full\-attention mass retained by the discovered working set\. Systems speedup is the setup\-inclusive Dense/WISEattention\-latency ratio under the matched block\-size execution harness\.Downstream quality exhibits a substantially weaker dependence on block size than execution efficiency\. On HotpotQA,B=16B=16differs fromB=32B=32by only−0\.0064\-0\.0064F1, with a paired 95% bootstrap confidence interval of\[−0\.0288,0\.0166\]\[\-0\.0288,\\,0\.0166\]\. On 2WikiMultiHopQA, the corresponding difference is\+0\.0334\+0\.0334, but its interval\[−0\.0043,0\.0738\]\[\-0\.0043,\\,0\.0738\]also includes zero\. Likewise, althoughB=64B=64attains higher observed F1 thanB=32B=32on both benchmarks, the paired difference is not established on HotpotQA \(\+0\.0145\+0\.0145, 95% CI\[−0\.0097,0\.0450\]\[\-0\.0097,\\,0\.0450\]\), while the smaller 2Wiki gain is positive \(\+0\.0102\+0\.0102, 95% CI\[0\.0005,0\.0228\]\[0\.0005,\\,0\.0228\]\)\. The sweep therefore does not identify a block size that consistently dominates downstream quality across both benchmarks\. Instead,B=32B=32is selected as the default systems\-facing operating point because it provides the strongest overall balance between routing sparsity and realized execution efficiency, rather than because it is universally optimal for F1\.
### C\.7Sensitivity to the Working\-Set Mass Threshold
We evaluate sensitivity to the cumulative\-mass thresholdη\\etawhile holding all other aspects ofWISEfixed:T=32T=32,td=12t\_\{\\mathrm\{d\}\}=12,B=32B=32, and the four\-step discovery window spanning recurrent steps 9–12\. The threshold controls the conservativeness of working\-set construction: lower values retain fewer blocks, while higher values preserve more of the Full\-attention mass\. Loweringη\\etato0\.900\.90produces substantially sparser working sets, but also reduces retained future attention mass and increases behavioral changes relative to Full\. Increasingη\\etato0\.990\.99preserves nearly all future attention mass, but raises block density to roughly80%80\\%without yielding a consistent quality improvement across benchmarks\. The defaultη=0\.95\\eta=0\.95therefore provides a reasonable quality–sparsity operating point rather than relying on a narrowly tuned F1 optimum\.
Table 9:Sensitivity to the working\-set mass thresholdη\\eta\. All settings other thanη\\etaare fixed\.Δ\\DeltaF1 is measured relative to the matched Full run, and Future denotes retained future Full\-attention mass\.Similar Articles
Learning What Matters: Supervising Sparse Attention Routing with Causal Evidence Sets
This paper tests the assumption that attention weights reveal what a model actually depends on for its output, finding that attention and causal dependence often disagree. They propose using causal evidence sets obtained via intervention masking as supervision for sparse attention routers, achieving near-perfect accuracy on retrieval tasks where attention-distilled routers fail.
Attention-Aware Routing: Coupling Routing and Attention in MoEs
This paper introduces Attention-Aware Routing (AAR), a method that enhances Mixture-of-Experts language models by incorporating attention weights into the router, improving mathematical reasoning performance and revealing coupled dynamics between routing and attention.
Sticky Routing: Training MoE Models for Memory-Efficient Inference
StickyMoE proposes a differentiable routing consistency loss that encourages adjacent tokens to activate the same experts in MoE models, reducing expert-swapping overhead and cache misses during inference on edge devices by up to 3.92× while improving perplexity.
Fast Weight Attention for Continual Learning
This paper analyzes recurrent fast-weight memories and selective state-space models as online learning rules, deriving normalized update families that improve length extrapolation and remain competitive in language modeling.
Uncertainty-gated selection for block-sparse attention
Proposes an uncertainty-gated router that doubles the selected key blocks for queries with uncertain cutoff margins, improving recall and accuracy in block-sparse attention for long-context language models, validated on multiple architectures.