Anatomy of Associative Recall in Fixed-State Recurrences: A Matched-State Decomposition, an Interference Wall, and a Curriculum That Breaks It

arXiv cs.LG Papers

Summary

This paper decomposes associative recall in fixed-state recurrences, finding convolution and curriculum learning key to performance, and proposes interventions to address interference rather than capacity limitations.

arXiv:2609.16183v1 Announce Type: new Abstract: Fixed-state recurrences--linear attention and state-space models--are reported to lag behind attention on associative recall, but whole-architecture comparisons cannot say which ingredient is responsible. We decompose masked multi-query recall at a fixed state budget along three single-knob axes: a short causal convolution, the transition structure (rank-1 delta rule vs. diagonal), and decay. The convolution dominates (~+0.5 recall in both families under matched training): comparisons that pit convolution-free cells against a convolution-equipped Mamba measure the missing convolution, not the recurrence. The rank-1 transition beats its diagonal ablation by +0.19/+0.32 at 16/32 pairs, but the margin shrinks to +0.03 once both cells carry the convolution, and a state-matched Mamba-2 ties the unarmed rank-1 cell: no class claim survives. Cells that solve 32-pair recall degrade gracefully with load yet fall to chance retrieving 4 pairs from a distractor haystack--flat across lengths and transitions. Interference under sparse supervision, not capacity: a distance curriculum takes the unchanged architecture from 0.021 to 1.000. Training is a lock-in lottery--a seed either locks in or does not--and the curriculum is the lever. Lock-in rises from 1/10 to 7/10 (p=0.02); dense supervision adds nothing; at L=256 a shaped ramp reopens a boundary the uniform curriculum cannot (4/5 vs. 0/9); and at L=512, where the ramp collapses (0/6), gating it on measured accuracy locks in 6/6 (p=0.001). Bidirectional denoiser cells, reading the query before the haystack, show no measurable advantage over causal training (ten seeds), and collision-key retrieval needs two layers. Arming for recall is free on an S_5 state-tracking guardrail--the armed cell is significantly better at every depth (p<=0.0044). These replace "recurrent models are bad at recall" with a measured decomposition and two cheap interventions.
Original Article
View Cached Full Text

Cached at: 09/16/26, 08:43 AM

# Anatomy of Associative Recall in Fixed-State Recurrences:A Matched-State Decomposition, an Interference Wall,and a Curriculum That Breaks It
Source: [https://arxiv.org/html/2609.16183](https://arxiv.org/html/2609.16183)
Julian BoeschAndrew WeeAffiliation:Purdue UniversityAffiliation:Obit ResearchCorrespondence to:[awee@obitmc\.com](mailto:[email protected])

###### Abstract

Fixed\-state recurrences—linear attention and state\-space models—are widely reported to lag behind attention on associative recall, but whole\-architecture comparisons cannot say*which ingredient*is responsible\. We decompose masked multi\-query recall at a fixed state budget along three single\-knob axes: a short causal convolution, the state\-transition structure \(rank\-1 delta rule vs\. diagonal\), and decay\. The convolution dominates, adding roughly 0\.5 recall accuracy in both families under matched training: comparisons that pit convolution\-free cells against a convolution\-equipped Mamba are measuring the missing convolution, not the recurrence\. The rank\-1 transition beats its diagonal ablation by\+0\.19\+0\.19/\+0\.32\+0\.32at1616/3232pairs, but the margin shrinks to\+0\.03\+0\.03once both cells carry the convolution, and a state\-matched Mamba\-2 ties the rank\-1 cell before it is “armed” with one: no architecture\-class claim survives\. Cells that solve 32\-pair recall degrade gracefully with load yet fall to chance when asked to retrieve just 4 pairs from a haystack of distractors—flat across sequence lengths and transition types\. The cause is interference under sparse supervision, not capacity: a*distance curriculum*takes the unchanged architecture from0\.0210\.021to1\.0001\.000\. Training is a*lock\-in lottery*—a seed either locks in or does not—and the curriculum is the lever\. Lock\-in rises from1/101/10seeds to7/107/10\(p=0\.02p\{=\}0\.02\); dense supervision adds nothing; atL=256L\{=\}256a shaped ramp reopens a boundary the uniform curriculum cannot \(4/54/5vs\.0/90/9\); and atL=512L\{=\}512, where the ramp itself collapses \(0/60/6\), gating it on measured accuracy locks in6/66/6\(p=0\.001p\{=\}0\.001\)\. Two side results: bidirectional denoiser cells, which natively read the query before the haystack, show no measurable advantage over causal training at ten seeds, and collision\-key retrieval needs two layers\. Finally, arming for recall is free on anS5S\_\{5\}state\-tracking guardrail—the armed cell is significantly*better*at every depth \(p≤0\.0044p\{\\leq\}0\.0044\)\. Together these replace “recurrent models are bad at recall” with a measured decomposition and two cheap interventions\.

###### Keywords:

linear attention, state space models, associative recall, MQAR, curriculum learning, diffusion language models, Machine Learning, ICML

## 1Introduction

Associative recall—binding key–value pairs in context and retrieving a value when its key reappears—is the capability axis on which fixed\-state sequence mixers most visibly trail attention\. A line of work has made this precise: synthetic multi\-query associative recall \(MQAR\) separates attention from efficient mixers and predicts much of the language\-modeling gap between them\([Arora et al\., 2023](https://arxiv.org/html/2609.16183#bib.bib1);[Arora et al\., 2024a](https://arxiv.org/html/2609.16183#bib.bib2)\), and recall remains the standard lens on what a bounded recurrent state can and cannot retain\([Jelassi et al\., 2024](https://arxiv.org/html/2609.16183#bib.bib17);[Hsieh et al\., 2024](https://arxiv.org/html/2609.16183#bib.bib18)\)\.

But a modern recurrent cell is not one mechanism\. A Mamba\-2 block, for instance, is a*diagonal selective*state transition*plus*an input\-dependent gate*plus*a short causal depthwise convolution\([Gu and Dao, 2023](https://arxiv.org/html/2609.16183#bib.bib12);[Dao and Gu, 2024](https://arxiv.org/html/2609.16183#bib.bib13)\); a Gated DeltaNet block is a*rank\-1 delta\-rule*transition plus a learned decay\([Yang et al\., 2024b](https://arxiv.org/html/2609.16183#bib.bib9);[Yang et al\., 2024a](https://arxiv.org/html/2609.16183#bib.bib10)\); RWKV\-7 couples a vector decay with a rank\-1 removal term\([Peng et al\., 2025](https://arxiv.org/html/2609.16183#bib.bib11)\)\. Whole\-architecture comparisons therefore confound at least three design axes, and the field’s summary judgments \(“DeltaNet\-style cells recall well,” “Mamba recalls better than RNNs,” “gating helps”\) average over knobs that can be toggled independently\. Recent causal\-intervention evidence sharpens the concern: Mamba’s induction behavior appears to live in its short convolution rather than its state\-space scan\([Arora et al\., 2025](https://arxiv.org/html/2609.16183#bib.bib4);[Parnichkun et al\., 2025](https://arxiv.org/html/2609.16183#bib.bib6)\), though this attribution is itself contested once learning rates are tuned\([Okpekpe and Orvieto, 2025](https://arxiv.org/html/2609.16183#bib.bib5)\)\. What is missing is a*controlled decomposition*: all knobs, one harness, one fixed state budget\.

This paper contributes that decomposition, and then uses it to re\-diagnose two failure modes that the decomposition alone does not explain\.

Contributions\.

1. 1\.A matched\-state decomposition of recall, and the confound it resolves\(§[4](https://arxiv.org/html/2609.16183#S4)\)\. At a fixed recurrent\-state budget \(1,024 state elements;∼29\{\\sim\}29k parameters\) we toggle one knob at a time: short convolution×\\timestransition structure \(rank\-1 vs\. diagonal\)×\\timesdecay\. The convolution is the dominant lever and transfers across cell families \(\+0\.47\+0\.47delta\-rule,\+0\.44\+0\.44diagonal atK=32K\{=\}32\), corroborating intervention evidence that Mamba’s recall is convolution\-borne\([Arora et al\., 2025](https://arxiv.org/html/2609.16183#bib.bib4);[Parnichkun et al\., 2025](https://arxiv.org/html/2609.16183#bib.bib6)\)\. We also measure the learning\-rate sensitivity that[Okpekpe and Orvieto \(2025\)](https://arxiv.org/html/2609.16183#bib.bib5)raise\. The rank\-1 transition contributes an independent, ablation\-grade margin over its own diagonal ablation \(\+0\.19\+0\.19/\+0\.32\+0\.32atK=16K\{=\}16/3232, 10 seeds, rebinding\-controlled\), but that margin shrinks to\+0\.034\+0\.034\(p=0\.0023p\{=\}0\.0023\) once both cells carry the convolution, and decay’s apparent recall cost dissolves at2020seeds \(p=0\.86p\{=\}0\.86\)\. Two corrections follow\. Comparing convolution\-free recurrent cells against a convolution\-equipped Mamba mismeasures the transition: armed with the same convolution, a gated delta\-rule cell reaches0\.990\.99at the hardest matched\-state setting\. The converse also holds\. A state\- and parameter\-matched Mamba\-2 ties the*unarmed*rank\-1 cell \(0\.5920\.592vs\.0\.6530\.653,p=0\.88p\{=\}0\.88\), so no class claim \(“rank\-1\>\>diagonal”\) survives\.
2. 2\.Capacity vs\. distance, discriminated\(§[5](https://arxiv.org/html/2609.16183#S5)\)\. Load \(KK\) produces graceful degradation; distance across a distractor haystack produces a*wall*: chance already at the shortest tested length—inside the training distribution—identical across delta\-rule, selective\-diagonal, and RWKV\-7 transitions, while a 2\-layer attention control solves the task at1\.01\.0\. This signature, completed by the curriculum result below, separates*interference under sparse supervision*from both capacity accounts and decay accounts\([Sridhar and Johansen, 2026](https://arxiv.org/html/2609.16183#bib.bib7)\)of recurrent retrieval failure\.
3. 3\.The wall is a training\-coverage gap, and training in this regime is a lock\-in lottery\(§[6](https://arxiv.org/html/2609.16183#S6)\)\. With the architecture unchanged, a table\-to\-query distance curriculum lifts the walled cell from0\.0210\.021\(chance\) to1\.0001\.000—an existence proof by construction\. Ten\-seed replication reframes the attribution as a*rate*shift and isolates the lever: lock\-in1/101/10under dense supervision alone vs\.7/107/10under the distance curriculum \(p=0\.02p\{=\}0\.02, Fisher two\-sided\); adding dense supervision to the curriculum changes nothing \(6/106/10\), and the causal cell locks in at a similar rate \(5/105/10\)\. Dense\-only training still solves the task outright on one seed\. The fix transfers unchanged to16×16\\timesstate \(d=128d\{=\}128\)\. Length is harder: atL=256L\{=\}256the uniform curriculum fails everywhere \(0/90/9trials\) but a shaped ramp locks in4/54/5seeds \(p≈0\.005p\{\\approx\}0\.005, Fisher\), and atL=512L\{=\}512the time\-based ramp itself collapses \(0/60/6\) while a*success\-gated*ramp—one that advances the gap only while measured accuracy holds—locks in6/66/6\(p=0\.0011p\{=\}0\.0011\)\. These are the two curriculum\-shape contrasts that clear significance\.
4. 4\.Bidirectional denoisers read twice for free—architecturally; no advantage is measured\(§[7](https://arxiv.org/html/2609.16183#S7)\)\. The backward stream of a bidirectional masked\-denoiser cell sees the query before the haystack, the prefix\-encoder property that Just\-Read\-Twice engineers into recurrent LMs\([Arora et al\., 2024b](https://arxiv.org/html/2609.16183#bib.bib3)\)\. A collision\-key discriminator \(decoys reuse the table’s keys, defeating any lexical write gate\) tests whether the property confers an advantage\. At ten seeds it does not: collision training is a bimodal lock\-in lottery \(bidirectional3/103/10seeds solve, causal1/101/10; mean difference n\.s\.,p=0\.63p\{=\}0\.63\), and shaped\-curriculum training that stabilizes the lottery \(9/109/10lock\-in\) does so at exact directional parity\. Four things are robust\. Both directions can solve the task; any solving circuit needs two layers \(a*mark\-and\-route*circuit, since the one\-layer version collapses\); the convolution\-free cell never locks in even under the shaped curriculum \(0/100/10vs\.9/109/10on the identical recipe,p≈10−4p\{\\approx\}10^\{\-4\}\); and curriculum shape is a powerful direction\-agnostic stabilizer \(1/10→9/101/10\\to 9/10,p≈6×10−4p\{\\approx\}6\{\\times\}10^\{\-4\}\)\.
5. 5\.A no\-cost guardrail\(§[8](https://arxiv.org/html/2609.16183#S8)\)\. Arming the cell for recall does not hurtS5S\_\{5\}state\-tracking learnability; at ten seeds the armed cell is significantly better at every probed depth \(L∈\{2,4,8,16\}L\\in\\\{2,4,8,16\\\}\)\.

All experiments run on a deliberately small, kernel\-free, commodity\-GPU harness: pure\-PyTorch reference cells, a state\-matched pure\-PyTorch Mamba\-2 comparator, and an atomic\-free two\-pass Triton backward that runs on pre\-Volta hardware \(Appendix[A](https://arxiv.org/html/2609.16183#A1)\)\. That is what makes single\-knob matched\-state factorials cheap to replicate\. The price is scale: every result here is toy\-scale \(d∈\{32,128\}d\\in\\\{32,128\\\}, synthetic tasks\), and §[9](https://arxiv.org/html/2609.16183#S9)states the resulting limits explicitly\.

## 2Related Work

Recall in efficient mixers\.MQAR and its relatives were introduced to explain the attention–SSM gap\([Arora et al\., 2023](https://arxiv.org/html/2609.16183#bib.bib1)\), and recall\-throughput tradeoffs drove a generation of linear\-attention designs\([Arora et al\., 2024a](https://arxiv.org/html/2609.16183#bib.bib2);[Katharopoulos et al\., 2020](https://arxiv.org/html/2609.16183#bib.bib14);[Schlag et al\., 2021](https://arxiv.org/html/2609.16183#bib.bib15)\)\. The delta rule\([Yang et al\., 2024b](https://arxiv.org/html/2609.16183#bib.bib9)\)and its gated variant\([Yang et al\., 2024a](https://arxiv.org/html/2609.16183#bib.bib10)\)explicitly target recall via key\-conditioned replacement; RWKV\-7 generalizes the transition to diagonal\-plus\-rank\-1\([Peng et al\., 2025](https://arxiv.org/html/2609.16183#bib.bib11)\)\. Our contribution to this line is not a new cell but a*controlled decomposition*of existing ingredients at matched state, including the negative result that the rank\-1 advantage over a diagonal ablation does not extend to a class advantage over Mamba\-2\.

Where Mamba’s recall comes from\.Causal interventions locate Mamba’s induction in its short convolution\([Arora et al\., 2025](https://arxiv.org/html/2609.16183#bib.bib4)\); effective\-state\-size analyses likewise find many linear structures collapse on recall without their convolutions\([Parnichkun et al\., 2025](https://arxiv.org/html/2609.16183#bib.bib6)\)\.[Okpekpe and Orvieto \(2025\)](https://arxiv.org/html/2609.16183#bib.bib5)complicate the attribution, showing tuned learning rates can recover much of Mamba’s recall without the convolution—a learnability confound\. Our factorial brings a design the intervention studies lack \(single\-knob toggles at fixed state, in both cell families\), though at toy scale\. Under matched training we find the convolution’s contribution large, additive, and family\-transferable, and we measure the learning\-rate sensitivity directly \(§[4\.2](https://arxiv.org/html/2609.16183#S4.SS2)\)\.

Long\-context retrieval failures\.That SSM\-family models fail needle\-in\-a\-haystack retrieval is well documented\([Hsieh et al\., 2024](https://arxiv.org/html/2609.16183#bib.bib18);[Jelassi et al\., 2024](https://arxiv.org/html/2609.16183#bib.bib17);[Ben\-Kish et al\., 2024](https://arxiv.org/html/2609.16183#bib.bib19)\); proposed accounts include state capacity and collapse\([Chen et al\., 2024](https://arxiv.org/html/2609.16183#bib.bib20)\), finite\-horizon state decay\([Sridhar and Johansen, 2026](https://arxiv.org/html/2609.16183#bib.bib7)\), and under\-explored state distributions at unseen lengths\([Buitrago Ruiz and Gu, 2025](https://arxiv.org/html/2609.16183#bib.bib8)\)\. Our haystack wall differs from these accounts twice over\. Its signature is chance at the*shortest*tested length, flat across lengths, within the training distribution, and transition\-independent\. Its resolution is a training\-signal change alone \(§[6](https://arxiv.org/html/2609.16183#S6)\), which no purely structural account predicts\.

Training\-side fixes\.Birdie improves SSM retrieval with reward\-driven objectives and mixed training procedures\([Blouir et al\., 2024](https://arxiv.org/html/2609.16183#bib.bib21)\);[Okpekpe and Orvieto \(2025\)](https://arxiv.org/html/2609.16183#bib.bib5)and[Buitrago Ruiz and Gu \(2025\)](https://arxiv.org/html/2609.16183#bib.bib8)frame recurrent retrieval and length generalization as learnability problems; curricula are classical\([Bengio et al\., 2009](https://arxiv.org/html/2609.16183#bib.bib22)\)\. We add the sharpest version of the claim for retrieval distance: a chance→\\toperfect existence proof with a single\-axis curriculum on an unchanged architecture, its lever attribution \(2×22\{\\times\}2\), and its failure boundary in length\.

Recurrent denoisers for diffusion LMs\.Masked discrete diffusion\([Austin et al\., 2021](https://arxiv.org/html/2609.16183#bib.bib24);[Sahoo et al\., 2024](https://arxiv.org/html/2609.16183#bib.bib23)\)makes the denoiser bidirectional by construction; bidirectional Mamba denoisers match Transformer DLMs\([Singh et al\., 2025](https://arxiv.org/html/2609.16183#bib.bib25)\), and RWKV\-backbone block diffusion exists\([Lin et al\., 2026](https://arxiv.org/html/2609.16183#bib.bib26)\)\. Just\-Read\-Twice\([Arora et al\., 2024b](https://arxiv.org/html/2609.16183#bib.bib3)\)showed causal recurrent LMs recover most of the recall gap when the query precedes the context\. Our contribution is the explicit connection: bidirectional denoisers possess the JRT property*natively*, with no prompt repetition and no bespoke prefix\-LM\. We add the depth requirement of any circuit exploiting it \(two layers: mark\-and\-route\) and its apparent local\-binding requirement \(the short convolution\)\. Whether the native property confers a measurable*advantage*over matched causal training is open: our ten\-seed collision replication finds no significant directional difference \(§[7](https://arxiv.org/html/2609.16183#S7)\)\.

State tracking\.Diagonal SSMs and attention are limited on group\-composition state tracking in ways rank\-1\-transition recurrences need not be\([Merrill et al\., 2024](https://arxiv.org/html/2609.16183#bib.bib27);[Grazzi et al\., 2025](https://arxiv.org/html/2609.16183#bib.bib28);[Liu et al\., 2023](https://arxiv.org/html/2609.16183#bib.bib30);[Delétang et al\., 2023](https://arxiv.org/html/2609.16183#bib.bib29)\)\. We use anS5S\_\{5\}word problem only as a*guardrail*\(does arming for recall cost state\-tracking learnability?\) and make no expressivity\-class claims here\.

## 3Cells, Tasks, and Protocol

### 3\.1The cell zoo

All cells are implemented in one harness as sequence mixers with identical embedding, readout, and training loops; we call each cell configuration under test an*arm*\. Per head, a matrix state𝐒t∈ℝdv×dk\\mathbf\{S\}\_\{t\}\\in\\mathbb\{R\}^\{d\_\{v\}\\times d\_\{k\}\}is updated by a cell\-specific transition; the output is a state readout\. We writekt,vt,qtk\_\{t\},v\_\{t\},q\_\{t\}\(orrtr\_\{t\}\) for the per\-token projections\.

DeltaNet\(rank\-1, no decay\) applies the delta rule\([Schlag et al\., 2021](https://arxiv.org/html/2609.16183#bib.bib15);[Yang et al\., 2024b](https://arxiv.org/html/2609.16183#bib.bib9)\), with input\-dependent write strengthβt∈\(0,1\)\\beta\_\{t\}\\in\(0,1\):

𝐒t=𝐒t−1​\(𝐈−βt​kt​kt⊤\)\+βt​vt​kt⊤,ot=𝐒t​qt\.\\mathbf\{S\}\_\{t\}=\\mathbf\{S\}\_\{t\-1\}\\\!\\left\(\\mathbf\{I\}\-\\beta\_\{t\}k\_\{t\}k\_\{t\}^\{\\top\}\\right\)\+\\beta\_\{t\}\\,v\_\{t\}k\_\{t\}^\{\\top\},\\quad o\_\{t\}=\\mathbf\{S\}\_\{t\}q\_\{t\}\.\(1\)
Gated DeltaNet\(gated\_deltanet; rank\-1 \+ decay\) adds a learned scalar gateαt∈\(0,1\)\\alpha\_\{t\}\\in\(0,1\)\([Yang et al\., 2024a](https://arxiv.org/html/2609.16183#bib.bib10)\):

𝐒t=αt​𝐒t−1​\(𝐈−βt​kt​kt⊤\)\+βt​vt​kt⊤\.\\mathbf\{S\}\_\{t\}=\\alpha\_\{t\}\\,\\mathbf\{S\}\_\{t\-1\}\\\!\\left\(\\mathbf\{I\}\-\\beta\_\{t\}k\_\{t\}k\_\{t\}^\{\\top\}\\right\)\+\\beta\_\{t\}\\,v\_\{t\}k\_\{t\}^\{\\top\}\.\(2\)
RWKV\-7 reference\(dg\_rank; rank\-1 \+ vector decay\) uses the diagonal\-plus\-rank\-1 transition\([Peng et al\., 2025](https://arxiv.org/html/2609.16183#bib.bib11)\):

𝐀t\\displaystyle\\mathbf\{A\}\_\{t\}=diag⁡\(wt\)−\(at⊙κ^t\)​κ^t⊤,\\displaystyle=\\operatorname\{diag\}\(w\_\{t\}\)\-\(a\_\{t\}\\odot\\hat\{\\kappa\}\_\{t\}\)\\hat\{\\kappa\}\_\{t\}^\{\\top\},\(3\)𝐒t\\displaystyle\\mathbf\{S\}\_\{t\}=𝐒t−1​𝐀t\+vt​kt⊤,ot=𝐒t​rt,\\displaystyle=\\mathbf\{S\}\_\{t\-1\}\\mathbf\{A\}\_\{t\}\+v\_\{t\}k\_\{t\}^\{\\top\},\\qquad o\_\{t\}=\\mathbf\{S\}\_\{t\}r\_\{t\},\(4\)with data\-dependent vector decaywt∈\(0,1\)dkw\_\{t\}\\in\(0,1\)^\{d\_\{k\}\}, removal keyκ^t\\hat\{\\kappa\}\_\{t\}, and in\-context rateata\_\{t\}\. Itsdiagonal ablation\(dg\_diag\) sets the rank\-1 term to zero \(at≡0a\_\{t\}\\equiv 0\), leaving𝐀t=diag⁡\(wt\)\\mathbf\{A\}\_\{t\}=\\operatorname\{diag\}\(w\_\{t\}\): the same cell, same parameter count, one knob\.

Mamba\-2 reference\(mamba2\_ref\) is a pure\-PyTorch state\-matched implementation of the Mamba\-2 SSD recurrence\([Dao and Gu, 2024](https://arxiv.org/html/2609.16183#bib.bib13)\): a per\-head scalar selective diagonal scan over an explicit state𝐇t∈ℝdinner×dstate\\mathbf\{H\}\_\{t\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{inner\}\}\\times d\_\{\\mathrm\{state\}\}\},

𝐇t=at​𝐇t−1\+xt​bt⊤,yt=𝐇t​ct,\\mathbf\{H\}\_\{t\}=a\_\{t\}\\,\\mathbf\{H\}\_\{t\-1\}\+x\_\{t\}\\,b\_\{t\}^\{\\top\},\\qquad y\_\{t\}=\\mathbf\{H\}\_\{t\}\\,c\_\{t\},\(5\)with input\-dependent\(at,bt,ct\)\(a\_\{t\},b\_\{t\},c\_\{t\}\)and the standard short causal depthwise convolution on its input projections\.mamba2\_ref\_noconvsets the convolution width to 1 \(a true toggle\)\. Because this comparator needs neither Triton normamba\_ssm, the full factorial runs on commodity GPUs; its cost is that it under\-trains relative to the official fused kernel \(§[9](https://arxiv.org/html/2609.16183#S9)\)\.

The arming knob\.A width\-4, strictly causal \(left\-padded\), depthwise convolution on the interaction projections—\(q,k,v\)\(q,k,v\)for the delta\-rule cells,\(v,k,r\)\(v,k,r\)for the RWKV\-7 cells—applied pre\-activation exactly as in Mamba\-2\. The convolution adds≈0\.5\{\\approx\}0\.5k parameters andzero recurrent state, so state\-matching is preserved\.armed\_gdn≔\\coloneqqgated\_deltanet\+\+this convolution\.

Attention control\.A standard bidirectional \(or causal, where noted\) softmax\-attention block\([Vaswani et al\., 2017](https://arxiv.org/html/2609.16183#bib.bib34)\), used as the ceiling for the haystack task \(§[5](https://arxiv.org/html/2609.16183#S5)\) with widened heads \(head dimension 16\)\. At the constrained\-width factorial setting \(d=32d\{=\}32\) no valid attention control exists at matched width: the head dimension degenerates, and a two\-head variant collapses at one layer for a structural reason \(§[9](https://arxiv.org/html/2609.16183#S9)\)\. We therefore do not report attention there\.

Bidirectional variants\.For §[7](https://arxiv.org/html/2609.16183#S7), a cell is made bidirectional in the standard masked\-denoiser way\([Singh et al\., 2025](https://arxiv.org/html/2609.16183#bib.bib25)\): a paired backward stream runs the same recurrence on the flipped sequence and the two streams are merged per position\.armed\_gdn\_fwddenotes the causal \(forward\-only\) ablation\.

### 3\.2State matching

The factorial is run at*constrained*state \(d=32d\{=\}32, one layer\), the regime in which recall discriminates between cells\. The rank\-1 cells carrydk2=322=1024d\_\{k\}^\{2\}=32^\{2\}=1024recurrent state elements per head group; the Mamba\-2 reference is set todstate=16d\_\{\\mathrm\{state\}\}\{=\}16so thatdinner⋅dstate=64⋅16=1024d\_\{\\mathrm\{inner\}\}\\cdot d\_\{\\mathrm\{state\}\}=64\\cdot 16=1024elements—matched state, and matched parameters \(∼28\.9\{\\sim\}28\.9k vs\.∼29\.3\{\\sim\}29\.3k\)\. At ample state \(d=128d\{=\}128\) every cell solves every setting and nothing discriminates, which is itself worth reporting:*recall differences among modern cells are a constrained\-state phenomenon*\.

### 3\.3Tasks and floors

Masked MQAR\.Each sequence is a table ofKKkey–value pairs followed by queried keys whose answers are*single masked tokens*, and the model is trained to fill the masks\. Values are drawn fresh at random per sequence from a 26\-token range, so no key→\\tovalue prior exists to memorize\.K∈\{8,16,32\}K\\in\\\{8,16,32\\\}\. Floors: the value\-marginal is1/26≈0\.0381/26\\approx 0\.038; the strongest*no\-binding*strategy \(emit a random table value\) scores≈0\.068\{\\approx\}0\.068\. In the hardened protocol we also report the*table floor*1/K1/K\(guess among theKKvalues actually present\), the floor the naive marginal comparison misses\.

Haystack retrieval\.Each sequence is44key–value pairs, then a distractor haystack of task\-format tokens, then the query; the answer is a single masked token at the final position\. Sequence lengths areL∈\{64,…,512\}L\\in\\\{64,\\dots,512\\\}and chance is1/54≈0\.0191/54\\approx 0\.019\. The query\-at\-end layout makes the task 2\-hop; all arms get 2 layers and a proper training budget \(the 1\-layer version fails for every arm including attention, an under\-capacity artifact we exclude\)\.

Collision\-key variant\.As above, but decoy pairs in the gap*reuse the table’s own keys*with fresh random values; ground truth is the*first*\(table\) binding\. This defeats any lexical write gate: distractors are lexically indistinguishable from table keys\. Floors: value\-marginal1/54≈0\.0191/54\\approx 0\.019; the strongest non\-binding*positional*strategy \(emit a random table\-position token\) scores≈0\.26\{\\approx\}0\.26\.

S5S\_\{5\}guardrail\.A running\-product word problem over the symmetric groupS5S\_\{5\}: generators are drawn uniformly from the*full group*\(so the answer marginal equals1/\|G\|=1/120≈0\.0081/\|G\|=1/120\\approx 0\.008at every depth\), each group element is a*single*token, and*all*product slots are masked simultaneously \(one parallel fill\), so no answer token is completable from a visible prefix\. Composition depthsL∈\{2,4,8,16\}L\\in\\\{2,4,8,16\\\}atd=128d\{=\}128\. These three design choices matter: they make the probe structurally immune to the answer\-prior and prefix\-leakage confounds that can silently invalidate state\-tracking probes with multi\-token, teacher\-forced answers\.111A companion report documents that failure anatomy in a conversion\-evaluation setting; this paper’s probes are immune by construction, and we state floors with every result\.

### 3\.4Statistical protocol

Exploratory sweeps use 3 seeds\. Headline ablations are*hardened*: 10 seeds, two\-sided permutation tests on mean gaps, and arebinding controlthat evaluates the trained model on sequences whose key–value bindings have been deranged, scoring the*original*value\. A genuine retriever’s rebinding score must collapse toward the floors \(it emits the*rebound*value, not the trained one\); a value\-salience heuristic does not\. For recall and state\-tracking probes, these rules operationalize general lessons on prefix leakage in multi\-token targets, answer\-prior baselines, and control tasks\([Bachmann and Nagarajan, 2024](https://arxiv.org/html/2609.16183#bib.bib31);[Holtzman et al\., 2021](https://arxiv.org/html/2609.16183#bib.bib32);[Hewitt and Liang, 2019](https://arxiv.org/html/2609.16183#bib.bib33)\)\. Where seed outcomes are bimodal \(a seed either “locks in” or does not\), we report*lock\-in rates*\(fraction of seeds above threshold\) with Fisher exact tests rather than means, since mean±\\pmstd is meaningless for bimodal outcomes\. Every table caption states its seed count, and every task quotes its floors \(derivations in Appendix[C](https://arxiv.org/html/2609.16183#A3)\)\.

## 4A Matched\-State Decomposition of Recall

### 4\.1The factorial

Table[1](https://arxiv.org/html/2609.16183#S4.T1)ranks all seven arms, next to the ingredients each one carries, at constrained state under the hardest load,K=32K\{=\}32\.222No attention control is reported at this setting: atd=32d\{=\}32with 8 heads the head dimension degenerates to 4, and a two\-head variant collapses at one layer because MQAR is an induction\-head task that needs two attention layers \(§[9](https://arxiv.org/html/2609.16183#S9)\)\. Attention appears as a valid ceiling in the two\-layer haystack experiments \(§[5](https://arxiv.org/html/2609.16183#S5)\), with head dimension 16\.Table[2](https://arxiv.org/html/2609.16183#S4.T2)gives the fullKK\-sweep \(72/72 trials\)\. The next subsection reads both tables one knob at a time\.

Table 1:Constrained state,K=32K\{=\}32\(hardest setting\)\.Masked recall, mean \[min, max\] overn=10n\{=\}10seeds \(n=20n\{=\}20for the two decay arms\); these statistics supersede then=3n\{=\}3K=32K\{=\}32column of Table[2](https://arxiv.org/html/2609.16183#S4.T2)\. State\-matched throughout \(1,024 elements\)\. Floors: value\-marginal≈0\.038\\approx 0\.038; no\-binding≈0\.068\\approx 0\.068\. For context: the official fused Mamba\-2 kernel on the same protocol measures0\.8670\.867at the state\-matcheddstate=16d\_\{\\mathrm\{state\}\}\{=\}16in our paired session \(0\.8840\.884in the earlier cross\-session read\) and1\.0001\.000atdstate=64d\_\{\\mathrm\{state\}\}\{=\}64; per\-arm tuning lifts the matched\-state number to0\.9550\.955\(§[4](https://arxiv.org/html/2609.16183#S4)\)\.Table 2:FullKK\-sweep, mean masked recall overn=3n\{=\}3seeds \(72/72 trials\); theK=32K\{=\}32column is superseded by the hardened statistics of Table[1](https://arxiv.org/html/2609.16183#S4.T1)\(deltanetin particular is bimodal, so itsn=3n\{=\}3mean flatters it\)\. Everything solvesK=8K\{=\}8; discrimination grows with load; the no\-conv diagonal fails byK=16K\{=\}16; the graceful decline of the unarmed arms fromK=8K\{=\}8toK=32K\{=\}32is the*capacity*signature—contrast the haystack wall of §[5](https://arxiv.org/html/2609.16183#S5)\.Figure 1:The three single\-knob toggles atK=32K\{=\}32\(constrained state;n=10n\{=\}10atK=32K\{=\}32for the five recurrent arms,n=20n\{=\}20for the decay pair—values from Table[1](https://arxiv.org/html/2609.16183#S4.T1)\)\. Light bars: ingredient off; dark bars: on\. The short convolution is the dominant, family\-transferable lever; the rank\-1 transition adds an ablation\-grade within\-cell margin \(hardened at 10 seeds:\+0\.319\+0\.319without the convolution,\+0\.034\+0\.034with it; Table[3](https://arxiv.org/html/2609.16183#S4.T3), §[4](https://arxiv.org/html/2609.16183#S4)\); the decay bar’s apparent cost dissolves atn=20n\{=\}20\(p=0\.86p\{=\}0\.86, §[4](https://arxiv.org/html/2609.16183#S4)\)\. Dashed line: strongest no\-binding floor \(≈0\.068\\approx 0\.068\)\.
### 4\.2Reading the factorial: three separable levers

Figure[1](https://arxiv.org/html/2609.16183#S4.F1)shows the three toggles side by side\.

Axis 1: the convolution \(≈\+0\.45\\approx\{\+\}0\.45, family\-transferable\)\.Hold the recurrence fixed and toggle only the convolution\. The delta\-rule cell moves0\.518→0\.9830\.518\\to 0\.983\(\+0\.47\+0\.47;gated\_deltanet→\\toarmed\_gdn\), and the diagonal cell moves0\.151→0\.5920\.151\\to 0\.592\(\+0\.44\+0\.44;mamba2\_ref\_noconv→\\tomamba2\_ref\)\. Both legs are true within\-cell toggles\. The convolution is a*state\-free*local primitive, adding no recurrent state, yet it is the largest single lever in the factorial and is of similar magnitude in both families\. This is consistent with intervention evidence that Mamba’s induction is convolution\-borne\([Arora et al\., 2025](https://arxiv.org/html/2609.16183#bib.bib4);[Parnichkun et al\., 2025](https://arxiv.org/html/2609.16183#bib.bib6)\), and adds a controlled factorial to it\.

The size of this lever depends on the tuning regime\. We ran the learning\-rate check of[Okpekpe and Orvieto \(2025\)](https://arxiv.org/html/2609.16183#bib.bib5)in our own harness \(no\-conv arms at\{3​e\-​4,1​e\-​3,3​e\-​3\}\\\{3\\text\{e\-\}4,1\\text\{e\-\}3,3\\text\{e\-\}3\\\}, 5 seeds each, rebinding\-controlled\), and*per\-arm tuning rescues the delta family but not our diagonal cell*\. Tuned no\-convgated\_deltanetreaches0\.8700\.870at lr33e\-33\(vs\.0\.5180\.518matched\), whilemamba2\_ref\_noconvpeaks at0\.1720\.172across the sweep, leaving its within\-cell convolution gap \(\+0\.42\{\+\}0\.42\) intact under tuning\. Re\-tuning the*armed*arms as well, all three saturate at lr33e\-33\(armed\_gdn0\.9990\.999, armed rank\-11\.0001\.000, armed diagonal0\.9970\.997; 5 seeds each\)\. The*tuned\-vs\-tuned*delta\-family convolution gap is therefore≈\+0\.13\{\\approx\}\{\+\}0\.13\(0\.9990\.999vs\.0\.8700\.870\), and the with\-conv transition margin of\+0\.034\{\+\}0\.034is a matched\-lr statement, since at tuned lr both armed cells sit at ceiling \(margin≈0\.002\{\\approx\}0\.002\)\. The convolution’s contribution is thus*optimization\-conditional in the delta family*\(large matched,≈\+0\.13\{\\approx\}\{\+\}0\.13tuned\) and robust in the diagonal cell; magnitudes should always be quoted with their tuning regime\.

Axis 2: the transition \(≈\+0\.3\\approx\{\+\}0\.3without the convolution;≈\+0\.03\\approx\{\+\}0\.03with it\)\.The clean evidence is the within\-cell pairdg\_rankvs\.dg\_diag: one architecture, one parameter count, one switch \(the rank\-1 term on or off\)\. Without the convolution, at 10 seeds, the pair scores0\.6530\.653vs\.0\.3900\.390\(\+0\.26\+0\.26\), and the hardened 10\-seed rerun \(§[4\.3](https://arxiv.org/html/2609.16183#S4.SS3)\) confirms it with controls:\+0\.191\+0\.191atK=16K\{=\}16\(p=0\.0002p\{=\}0\.0002\),\+0\.319\+0\.319atK=32K\{=\}32\(p<10−4p\{<\}10^\{\-4\}\)\. The with\-conv quadrant is ablation\-grade as well \(same protocol and harness as Table[1](https://arxiv.org/html/2609.16183#S4.T1), 10 seeds, rebinding\-controlled\): the convolution\-equipped pair gives0\.9920\.992\(rank\-1\) vs\.0\.9580\.958\(diagonal\),\+0\.034\+0\.034\(p=0\.0023p\{=\}0\.0023\) atK=32K\{=\}32\. AtK=16K\{=\}16all armed arms sit at ceiling \(≥0\.9998\\geq 0\.9998, 10 seeds\), so the margin is a highest\-load phenomenon\. The transition’s contribution is therefore*conditional on the convolution*: large when the cell lacks local mixing, an order of magnitude smaller—though still resolvable—once the convolution is present and both cells operate near ceiling\.

Axis 3: decay \(no measurable cost\)\.The apples\-to\-apples comparison within the delta\-rule family, atn=20n\{=\}20, isdeltanet\(no decay\)0\.5610\.561vs\.gated\_deltanet\(decay\)0\.5180\.518atK=32K\{=\}32:p=0\.86p\{=\}0\.86, with nothing left of the−0\.32\-0\.32the first three seeds suggested\. What the seeds do show is bimodality, not a cost\. The no\-decay cell locks in on4/204/20seeds \(per\-seed0\.090\.09–0\.930\.93\), while the decay cell never exceeds0\.800\.80\(0/200/20lock\-ins; range0\.300\.30–0\.800\.80\); decay appears to trade occasional lucky solves for consistency, and neither is a mean\-level effect\. What is settled is thatarmed\_gdnwins with its decay in place, on the strength of the convolution plus the delta rule\. The converse claim—that decay buys state\-tracking—is likewise not established: a 10\-seed lock\-in comparison ofgated\_deltanetvs\.deltanetonS5S\_\{5\}is directionally positive but non\-significant \(p=0\.65p\{=\}0\.65\)\. One boundary condition comes from the literature: fine\-grained \(channel\-wise\) gating can*improve*recall over scalar gating\([Kimi Team, 2025](https://arxiv.org/html/2609.16183#bib.bib16)\), so any liability measured here is specific to the scalar\-decay knob at constrained state\.

### 4\.3The hardened transition ablation, with controls

Table 3:Hardened transition ablation\(n=10n\{=\}10seeds;d=32d\{=\}32,N=1N\{=\}1; the three delta\-rule arms are convolution\-free, while themamba2\_ds16comparator is a state\-matched Mamba\-2 that retains its stock short convolution\)\. “Rebound” is accuracy on the rebinding control \(bindings deranged, original value scored\)\. AtK=16K\{=\}16every arm falls to or below the table floor \(1/K=0\.0621/K\{=\}0\.062\); atK=32K\{=\}32rebound sits at0\.0420\.042–0\.0730\.073—slightly above the1/K=0\.0311/K\{=\}0\.031floor but near the value\-marginal \(0\.0380\.038\) and an order of magnitude below trained accuracy\. Models emit the*rebound*value: the recall is genuine key→\\tovalue binding, not value salience\. Permutation tests are two\-sided on mean gaps; gaps and drops are computed on full\-precision values, not the rounded cells shown\.Table[3](https://arxiv.org/html/2609.16183#S4.T3)reports the 10\-seed, rebinding\-controlled rerun of the transition axis\. Three findings:

1. 1\.The within\-cell ablation replicates decisively\.dg\_rank\>\>dg\_diag:\+0\.191\+0\.191atK=16K\{=\}16\(p=0\.0002p\{=\}0\.0002\) and\+0\.319\+0\.319atK=32K\{=\}32\(p<10−4p\{<\}10^\{\-4\}\), gaps far larger than seed variance\. The rebinding control verifies that what is being measured is binding\. The advantage is delta\-rule\-general, not RWKV\-7\-specific:deltanet≈\\approxdg\_rankatK=32K\{=\}32\(p=0\.36p\{=\}0\.36\), both≫\\ggdg\_diag\.
2. 2\.No class claim survives\.A state\- and parameter\-matched*Mamba\-2*is a diagonal recurrence, and it*ties*the rank\-1 cell atK=16K\{=\}16\(p=0\.25p\{=\}0\.25\) and*beats*it atK=32K\{=\}32\(0\.8490\.849vs\.0\.7020\.702,p=0\.0036p\{=\}0\.0036\)\. “Rank\-1 class\>\>diagonal class” is false in this data\. The diagonal*ablation*is a crippled member of its class, not a representative; the correct statement is:*within one cell at matched parameters, the rank\-1 transition adds recall capacity per state dimension*\.
3. 3\.The two statements together resolve the folklore\.Convolution\-free recurrent cells lose to Mamba on recall because of the convolution, not the recurrence \(Axis 1\)\. And the rank\-1 transition looked alternately strong and weak across studies because its genuine within\-cell effect was being compared across architectures against cells that carry the other two knobs\.

The armed arms at 10 seeds \(rebinding\-controlled, same harness as Table[1](https://arxiv.org/html/2609.16183#S4.T1)\)\.armed\_gdn0\.9830\.983\(per\-seed min0\.9170\.917; rebinding0\.0370\.037, at the value\-marginal floor\), armed RWKV\-7 \(dg\_rank\+\+conv\)0\.9920\.992\[0\.970,0\.9990\.970,0\.999\], armed diagonal ablation0\.9580\.958\[0\.870,0\.9960\.870,0\.996\]\. Within\-harness,armed\_gdnclearsmamba2\_refby\+0\.392\+0\.392\(p=10−5p\{=\}10^\{\-5\}\), and the two armed rank\-1 families are statistically indistinguishable \(p=0\.37p\{=\}0\.37\): once armed, the delta\-rule cell family does not matter\. That an armed*diagonal*cell reaches0\.9580\.958is the factorial’s sharpest single number, because it says that at constrained state the convolution is very nearly the whole recall story\.

The armed cell vs\. real Mamba, measured in one session\.We ran the official fused Mamba\-2 kernel \(state\-matcheddstate=16d\_\{\\mathrm\{state\}\}\{=\}16, bf16\) through the same probe, session, and Ampere GPU asarmed\_gdn, sweeping each arm’s learning rate over a decade \(\{10−3,3×10−3,10−2\}\\\{10^\{\-3\},3\{\\times\}10^\{\-3\},10^\{\-2\}\\\}, 5 seeds each\)\. At the shared default \(10−310^\{\-3\}\) the historical cross\-study picture reproduces: official0\.8670\.867\[0\.792, 0\.911\] vs\. armed0\.9670\.967, so the≈\+0\.10\{\\approx\}\{\+\}0\.10cross\-session margin was real at matched lr\. Per\-arm tuning compresses but does not close it\. Both arms peak at3×10−33\{\\times\}10^\{\-3\}, where the official kernel reaches0\.9550\.955\[0\.862, 0\.999\] andarmed\_gdn1\.0001\.000on every seed, a tuned\-vs\-tuned margin of\+0\.045\+0\.045\(p=0\.016p\{=\}0\.016, permutation\)\. At10−210^\{\-2\}*both*arms destabilize \(official0\.6060\.606with two collapsed seeds; armed0\.3300\.330with four\)\. The reading matches the delta\-family result above: most of the official cell’s apparent deficit is optimization\-conditional, and a smaller armed margin survives tuning\. Within our reference harness,armed\_gdnclears the under\-trainedmamba2\_refby\+0\.39\+0\.39\.

## 5Capacity Curves vs\. the Interference Wall

TheKK\-sweep \(Table[2](https://arxiv.org/html/2609.16183#S4.T2)\) shows what a*capacity*limit looks like: everything solvesK=8K\{=\}8; the unarmed arms decline gracefully toK=32K\{=\}32\. The haystack task shows something categorically different—total, immediate, and length\-flat \(Table[4](https://arxiv.org/html/2609.16183#S5.T4)\); Figure[2](https://arxiv.org/html/2609.16183#S5.F2)juxtaposes the two axes\.

Figure 2:Load vs\. distance\.Left: theKK\-sweep \(Table[2](https://arxiv.org/html/2609.16183#S4.T2);n=3n\{=\}3seeds\)—a graceful, capacity\-style decline; the RWKV\-7 rank\-1 arm \(0\.6390\.639atK=32K\{=\}32\) is omitted for legibility\. Right: the haystack wall and its removal for the same bidirectional armed cell atL=64L\{=\}64\(Tables[4](https://arxiv.org/html/2609.16183#S5.T4)and[5](https://arxiv.org/html/2609.16183#S6.T5); bars show seed\-0 accuracy, while the latter gives then=10n\{=\}10lock\-in rates\): naive training sits at chance \(dashed line,≈0\.019\\approx 0\.019\); the distance curriculum unlocks the task\. Hatched bar: the attention reference on the fixed layout \(§[6](https://arxiv.org/html/2609.16183#S6)explains why no attention ceiling is reported under the curriculum itself\)\.Table 4:The haystack wall, flat in length\(n=5n\{=\}5seeds per cell, means shown; 2 layers, proper budget; chance≈0\.019\\approx 0\.019\)\. The same recurrences that store 32 competing pairs at≥0\.99\\geq 0\.99\(Table[2](https://arxiv.org/html/2609.16183#S4.T2)\) fail at chance on*4*pairs across a distractor haystack—already at the shortest tested length, for every transition type, with no trend through4×4\\timesthe length: all 45 recurrent trials lie in\[0\.014,0\.023\]\[0\.014,0\.023\], straddling chance\. The properly headed attention control is exact at every length\.Four properties discriminate the failure mode:

1. 1\.It is not capacity\.The same cells store3232competing bindings at≥0\.99\\geq 0\.99\(Table[2](https://arxiv.org/html/2609.16183#S4.T2)\) but fail on44bindings plus distractors\. The state is8×8\\timesunder\-subscribed relative to demonstrated capacity\.
2. 2\.It is a wall, not a dip\.The failure is already total at the*shortest*tested length, inside the training distribution, at table\-to\-query gaps of only tens of tokens, and it stays exactly there through4×4\\timesthe length: means0\.0170\.017–0\.0190\.019acrossL∈\{64,128,256\}L\\in\\\{64,128,256\\\}, with every recurrent trial within\[0\.014,0\.023\]\[0\.014,0\.023\]\(Table[4](https://arxiv.org/html/2609.16183#S5.T4)\)\. Neither competing account fits that shape\. A capacity limit degrades gracefully, asKKdoes; a decay account\([Sridhar and Johansen, 2026](https://arxiv.org/html/2609.16183#bib.bib7)\)predicts degradation that grows with distance, not chance at minimal distance and not a floor with no length gradient\. There is no distance trend for state decay to explain: the binding is never learned, not progressively lost\. \(A 1\-layer sweep overL=64L\{=\}64–512512was likewise flat at chance but is excluded as under\-capacity for this 2\-hop task\.\)
3. 3\.Transition richness does not help\.GDN’s delta rule, Mamba’s selectivity, and RWKV\-7’s removal key wall identically—at chance at every tested length \(n=5n\{=\}5each\)\. Whatever is failing is not the state\-transition’s expressivity\.
4. 4\.Attention is untouched, as expected, because per\-token KV entries cannot be overwritten by later writes: the control is exact at every tested length \(15/1515/15trials at1\.0001\.000\)\.

The diagnosis consistent with all four:*interference under sparse supervision*\. The model must learn a selective\-write policy \(store table pairs; ignore distractors\), but the training signal—one masked token at the end of the sequence—gives the write gate almost no gradient\. Distractor writes overwrite the bindings before the query arrives\. This account differs from decay\-based explanations of SSM retrieval failure\([Sridhar and Johansen, 2026](https://arxiv.org/html/2609.16183#bib.bib7)\): a decay account predicts distance\-dependence and does not predict that a pure training\-signal change can fix the failure at fixed architecture\. Section[6](https://arxiv.org/html/2609.16183#S6)runs exactly that test\.

## 6The Wall Is a Training\-Coverage Gap

Same task, same architecture, same budget as Table[4](https://arxiv.org/html/2609.16183#S5.T4); only the training signal changes\. Under*dense supervision*, each sequence containsNNdistinct queried pairs and all of them are supervised\. Under the*distance curriculum*, the table→\\toquery gap is sampled uniformly in\[0,max\]\[0,\\text\{max\}\]per batch, while evaluation is always at max gap\. \(Exact flags in Appendix[B](https://arxiv.org/html/2609.16183#A2)\.\)

Table 5:Curriculum lever attribution, replicated atn=10n\{=\}10seeds\(L=64L\{=\}64haystack, 2 layers\)\. Training in this regime converges stochastically \(§[7](https://arxiv.org/html/2609.16183#S7)\), so we report lock\-in rates \(seeds\>0\.8\>0\.8\) with per\-seed values\. The chance→\\to1\.0001\.000headline is an existence proof by construction and does not rest on seed count; the*attribution*is a rate shift that the curriculum drives:7/107/10vs\.1/101/10\(p=0\.02p\{=\}0\.02, Fisher two\-sided;6/106/10vs\.1/101/10:p=0\.06p\{=\}0\.06\)\. Dense supervision on top of the curriculum, and the bidirectional–causal contrast, are both n\.s\. Attention rows: absolute\-positionn=5n\{=\}5, rotaryn=10n\{=\}10\.The existence proof\.The headline row of Table[5](https://arxiv.org/html/2609.16183#S6.T5)is an existence proof by construction and needs no seed band to be valid*as an existence claim*: an unchanged 2\-layer fixed\-state recurrence that sat at chance \(0\.0210\.021\) under naive training reaches1\.0001\.000on the same task under a distance curriculum\. The wall of §[5](https://arxiv.org/html/2609.16183#S5)is therefore a*training\-coverage gap*, not an architectural limit—for this task, at this length\. This is the sharpest form of a conclusion that the recent literature reaches by other routes: recurrent retrieval failure is substantially a learnability problem\([Okpekpe and Orvieto, 2025](https://arxiv.org/html/2609.16183#bib.bib5);[Buitrago Ruiz and Gu, 2025](https://arxiv.org/html/2609.16183#bib.bib8);[Blouir et al\., 2024](https://arxiv.org/html/2609.16183#bib.bib21)\)\.

Does it survive off the bench?The curriculum is the one result here that makes a deployment\-shaped claim, so it is the one we took to a real converted LM\. That work is reported in a companion paper\([Boesch and Wee, 2026](https://arxiv.org/html/2609.16183#bib.bib36)\): staged distillation of an attention teacher into a recurrent bidirectional diffusion student, at1\.71\.7B and88B\. Here we summarize only its bearing on the claims above\. It replicates the phenomenon rather than the number: at conversion scale the curriculum’s effect is again a lock\-in*rate*\. Closing the loop on the curriculum, so that the gap cap advances only while retrieval accuracy holds, raises that rate from one partial lock\-in in three to3/33/3at0\.940\.94–0\.990\.99\. The recipe carries to88B \(2/32/3seeds\), with recall transfer rising from0\.0000\.000to a mean of0\.9010\.901\. Two of this section’s conclusions therefore hold at four orders of magnitude more parameters: the wall is a training\-coverage gap, and training in this regime is a lottery whose rate the curriculum’s*shape*controls\. One boundary is added there and not visible here: every converted model that solves its trained token band reads chance on held\-out token bands, so what transfers is a binding circuit over the symbols it trained on\.

Lever attribution\.At ten seeds the picture is a family of*lock\-in lotteries*whose rates the training signal shifts, and the curriculum is the signal that matters\. Dense supervision alone can solve the task \(one seed of ten reaches1\.0001\.000\) but has the lowest lock\-in rate \(1/101/10\)\. The curriculum raises it to7/107/10\(p=0\.02p\{=\}0\.02, Fisher two\-sided;p=0\.01p\{=\}0\.01one\-sided\), consistent with the mechanism story: training across gaps builds the selective\-write policy outward from the solvable adjacent case\. Adding dense supervision to the curriculum does not help \(6/106/10; vs\. dense\-only,p=0\.06p\{=\}0\.06two\-sided\), so a denser gradient is not what closes the wall—coverage of the gap is\. The causal cell locks in at a similar rate under the same recipe \(5/105/10vs\.6/106/10, n\.s\.\): directional parity, consistent with the collision replication \(§[7](https://arxiv.org/html/2609.16183#S7)\)\.

Two boundaries\.\(i\)*The attention ceiling under the curriculum requires position generalization—and with it, exists\.*The absolute\-position toy baseline fails to generalize across table positions when the layout shifts \(robust across learning rates and seeds,7/77/7collapse at≈0\.046\{\\approx\}0\.046\)\. Replacing its learned absolute table with rotary positions—whose attention scores are provably shift\-invariant—restores the ceiling: rotary attention solves the fixed\-layout wall \(1\.0001\.000\) and learns under the shifting curriculum \(6/106/10lock\-in lexical,7/107/10collision\-ramp, with most misses graceful rather than collapsed\)\. The absolute\-position collapse is thereby confirmed as a position\-generalization artifact, not an attention limitation\. The shaped\-collision recurrent cells’ lock\-in rate \(9/109/10\) is statistically indistinguishable from rotary attention’s \(7/107/10\) at thesenn; the regimes are comparable, not ordered\.

\(ii\)*The fix has a length boundary, which shaping re\-opens—stochastically\.*Width scales cleanly: atd=128d\{=\}128\(16×16\\timesstate\) theL=64L\{=\}64collision\-plus\-curriculum recipe reaches1\.0001\.000unchanged\. Length under the*uniform\-gap*curriculum does not: atL=256L\{=\}256\(120 colliding decoy pairs vs\. 24\) every uniform cell is at chance, at both widths, at4×4\\timesthe step budget, at two learning rates\. A*shaped*\(progressive\-ramp\) curriculum re\-opens it\. At five seeds the ramp locks in4/54/5\(per\-seed0\.052/0\.994/0\.996/0\.868/1\.0000\.052/0\.994/0\.996/0\.868/1\.000\) against0/90/9pooled uniform trials across widths, budgets, and learning rates,p≈0\.005p\{\\approx\}0\.005\(Fisher, one\-sided; the uniform trials pool several configurations\)\. This is the first of two curriculum\-*shape*contrasts in this paper that clear significance\. SoL=256L\{=\}256is not an architectural interference ceiling; it is the lock\-in\-lottery phenomenon at a longer horizon, with curriculum shape as a*reliable*rate lever, and the residual lock\-out motivating success\-gated ramps\.

A further4×4\\timesin length then exhausts the open\-loop ramp outright: atL=512L\{=\}512every ramp trial sits at chance \(0/60/6; bidirectional0\.020/0\.020/0\.0190\.020/0\.020/0\.019, causal0\.023/0\.018/0\.0190\.023/0\.018/0\.019\) against4/54/5atL=256L\{=\}256\. Length is thus itself a rate\-killer for a*time\-based*schedule, which advances on a clock whether or not the model is keeping up\. Closing the loop removes the failure\. A*success\-gated*ramp advances the gap cap only while a running probe\-accuracy estimate stays above0\.60\.6, and never retreats; it locks in6/66/6seeds atL=512L\{=\}512\(per\-seed0\.9930\.993,0\.9740\.974,0\.9730\.973,0\.8250\.825,0\.9850\.985,0\.9760\.976; mean0\.9540\.954\) under the same architecture, step budget, and task on which the time\-based ramp scored0/60/6\(p=0\.0011p\{=\}0\.0011, Fisher, one\-sided\)\. TheL=512L\{=\}512boundary was therefore never a length ceiling of the two\-layer circuit; it was the schedule outrunning the model\. A companion paper reports the same closed\-loop fix at conversion scale\([Boesch and Wee, 2026](https://arxiv.org/html/2609.16183#bib.bib36)\)\. Both arms aren=6n\{=\}6, and the claim is scoped: gating dissolves a*schedule\-borne*boundary, not the token\-coverage boundary that paper adds\.Mean\-level results should be quoted as established atL=64L\{=\}64and, in length, as lock\-in rates only\.

## 7Bidirectional Denoisers Read Twice for Free

Just\-Read\-Twice\([Arora et al\., 2024b](https://arxiv.org/html/2609.16183#bib.bib3)\)showed that recurrent LMs recover most of the recall gap when the query precedes the context—either by repeating the prompt \(JRT\-Prompt\) or with a bespoke prefix\-LM architecture \(JRT\-RNN\)\. We observe that*bidirectional masked\-denoiser cells have this property natively*: the backward stream reads the query before the haystack, by construction, in every bidirectional diffusion\-LM denoiser\([Sahoo et al\., 2024](https://arxiv.org/html/2609.16183#bib.bib23);[Singh et al\., 2025](https://arxiv.org/html/2609.16183#bib.bib25)\)\. No prompt repetition, no bespoke architecture\. This section isolates whether that built\-in second read is a real mechanism, what circuit implements it, and what it requires\.

The discriminator\.The lexical haystack cannot settle the question\. There the causal cell also lifts under the curriculum \(5/105/10seeds lock in, mean0\.630\.63; §[6](https://arxiv.org/html/2609.16183#S6), Table[5](https://arxiv.org/html/2609.16183#S6.T5)\), because distractors are lexically distinguishable from table keys: a causal lexical write gate is learnable in principle, so directionality is a margin there, not a mechanism\. Thecollision\-keyvariant kills that route, since decoys reuse the table’s own keys and no lexical gate can separate a distractor write from a table write\. The theory going in: an input\-conditioned*causal*write gate must overwrite on key reuse \(retaining garbage→\\tochance\), whereas the*backward*stream reads the query first and its reverse\-scan overwrite retains the correct first\-in\-forward binding\.

Table 6:Collision\-key discriminator, hardened\(L=64L\{=\}64, dense\-4\+\+curriculum, 2 layers;n=10n\{=\}10seeds, rebinding\-controlled—the armed arms’ rebound sits at the value\-marginal floor, while the convolution\-free arm’s rebound*equals*its accuracy: it never binds at all\)\. Per\-seed accuracies are*bimodal*: most seeds lock out, some solve outright, so we report lock\-in rates \(seeds\>0\.8\>0\.8\) alongside means\. The bidirectional–causal mean difference is\+0\.082\+0\.082\(p=0\.63p\{=\}0\.63, permutation\):no directional difference is established at this power\. Floors: value\-marginal≈0\.019\\approx 0\.019; strongest non\-binding positional strategy≈0\.26\\approx 0\.26\.Reads \(Table[6](https://arxiv.org/html/2609.16183#S7.T6)\)\.

1. 1\.No directional margin at ten seeds\.Both arms are*lock\-in lotteries*\(bidirectional locks in3/103/10, causal1/101/10; the causal count is threshold\-sensitive, since at a0\.60\.6criterion both arms lock in3/103/10\)\. Means differ by\+0\.082\+0\.082\(p=0\.63p\{=\}0\.63\)\.*“Bidirectional denoisers get Just\-Read\-Twice for free” is not established as a differential claim at this power\.*
2. 2\.What does survive: solvability, in both directions\.Some seeds of*both*arms solve collision\-key retrieval outright \(bidirectional max1\.0001\.000; causal max0\.8440\.844\)\. The causal existence result is itself informative, because it robustly contradicts the single\-cell theory that an input\-conditioned write gate “must” overwrite on key reuse\. That strengthens the depth\-as\-novelty\-gate hypothesis: layer\-1 state can encode “this key is already bound,” making layer\-2’s gate effectively state\-conditioned\. We report that as a hypothesis consistent with, but not isolated by, these data\.
3. 3\.The depth requirement is robust\.At one layer*both*arms collapse to chance\. The mechanistic reason is exact\. The masked answer occupies the*last*position, which is the*first*token of the backward scan, so a single backward pass has absorbed nothing by the time it reaches the answer slot\. Query\-first information exists only at earlier positions and must be routed to the answer by a second layer’s forward stream \(mark\-and\-route\)\. Whatever circuit solves this task, it needs two layers\.
4. 4\.Lock\-in is the phenomenon to explain\.The same bimodality appears in the shaped\-curriculum length extension \(§[6](https://arxiv.org/html/2609.16183#S6): ramp seeds solveL=256L\{=\}256at0\.9940\.994or lock out at0\.050\.05\)\. Curriculum training on interference tasks converges stochastically\. The right statistic is therefore a lock\-in*rate*; mean accuracy is misleading; and differential claims between arms require far more than ten seeds at these rates \(3/103/10vs\.1/101/10isp≈0\.58p\{\\approx\}0\.58by Fisher exact\)\. The convolution\-free arm has never locked in \(0/100/10under the uniform recipe,0/100/10under the shaped ramp\), consistent with the reversed\-order\-binding account of the convolution’s role, but the same power caveat applies\.

Implication for diffusion LMs\.The architectural observation stands: bidirectional denoisers possess query\-first reading natively \(the JRT property causal recurrent LMs must engineer via prompt repetition or prefix\-LM encoders\), and any circuit exploiting it is necessarily≥\\geq2 layers \(mark\-and\-route\)\. What these data do*not*show is that the native property confers a measurable advantage over a causal cell trained the same way: at ten seeds the directional difference is indistinguishable from the lock\-in lottery\. The natural follow\-up is whether shaped training that stabilizes lock\-in also separates the directions\. It does not: under the progressive\-ramp curriculum, collision lock\-in rises from3/103/10to9/109/10\(bidirectional\) and from1/101/10to9/109/10\(causal;p≈6×10−4p\{\\approx\}6\{\\times\}10^\{\-4\}vs\. its uniform reference\),*at exact directional parity*\(p=0\.76p\{=\}0\.76\)\. The question is closed, negatively: in this task family, query\-first reading confers no measurable recall advantage even when training is stabilized\. The durable finding is that curriculum*shape*is a powerful, direction\-agnostic lock\-in stabilizer—the second shape contrast in this paper to clear significance\.

## 8Guardrail: Arming Is Free on State Tracking

Does adding a local \(finite\-window\) convolution cost the recurrence anything on the axis where recurrences are supposed to shine—state tracking? We comparearmed\_gdnvs\.gated\_deltaneton theS5S\_\{5\}running\-product guardrail \(§[3\.3](https://arxiv.org/html/2609.16183#S3.SS3)\)\.

Table 7:S5S\_\{5\}guardrail\(d=128d\{=\}128; chance≈0\.008\\approx 0\.008; internally comparable only\)\. All four rows are ten\-seed \(mean \[min, max\], permutationpp\)\. The armed cell is significantly*better*at every probed depth: arming for recall does not trade away state tracking\. The result is cell\-family\-general: an armed RWKV\-7 cell also beats its unarmed counterpart at both depths \(\+0\.053\+0\.053,p=0\.004p\{=\}0\.004atL=8L\{=\}8;\+0\.005\+0\.005,p<10−4p\{<\}10^\{\-4\}atL=16L\{=\}16\) and is statistically indistinguishable from armedarmed\_gdn\(p=0\.60p\{=\}0\.60/0\.100\.10\)\.We frame this strictly as*learnability at fixed depth*, not as a circuit\-complexity statement: accuracy falls steeply with composition depth for both arms, and nothing here bears on asymptotic separations\([Merrill et al\., 2024](https://arxiv.org/html/2609.16183#bib.bib27);[Grazzi et al\., 2025](https://arxiv.org/html/2609.16183#bib.bib28)\)\. The design question was narrow—does the recall fix tax the state\-tracking axis? The answer is no: the armed cell wins significantly at all four depths\.

## 9Limitations and Threats to Validity

1. 1\.The Mamba\-2 reference under\-trains the official kernel\.mamba2\_refreaches0\.5920\.592at 10 seeds \(0\.6410\.641at 3\) where the official fused Mamba\-2 reached0\.8840\.884at the same matched state\. The*within\-harness*factorial is internally valid \(all cells, one harness, one budget\), but absolute margins against a real Mamba are softer\. The with\-conv transition gap is measured ablation\-grade within one cell \(\+0\.034\+0\.034, §[4](https://arxiv.org/html/2609.16183#S4)\), so it does not depend on this comparator\. The paired measurement is now done in one session with a per\-arm learning\-rate sweep \(§[4](https://arxiv.org/html/2609.16183#S4)\): at the shared default the cross\-session numbers reproduce \(0\.9670\.967vs\.0\.8670\.867\), and at each arm’s best lr the margin is1\.0001\.000vs\.0\.9550\.955\(p=0\.016p\{=\}0\.016\)\. Most of the cross\-session gap was the official arm’s learning rate, and a significant smaller margin survives tuning\.
2. 2\.Toy scale, single task family\.d∈\{32,128\}d\\in\\\{32,128\\\}, one MQAR configuration per setting, synthetic tasks\. The decomposition’s magnitudes \(\+0\.5\+0\.5convolution;\+0\.3\+0\.3transition without the convolution,\+0\.03\+0\.03with it; decay unresolved\) are specific to constrained state; at ample state everything saturates and the knobs stop mattering\. Claims should be read as mechanism isolation, not deployment guidance\.
3. 3\.Learning\-rate sensitivity\.[Okpekpe and Orvieto \(2025\)](https://arxiv.org/html/2609.16183#bib.bib5)show per\-arm learning\-rate tuning can rescue no\-conv Mamba recall; our sweep \(§[4\.2](https://arxiv.org/html/2609.16183#S4.SS2)\) confirms this*for the delta family*\(tuned no\-convgated\_deltanet0\.8700\.870; re\-tuning the armed arms too narrows the delta\-family gap to≈\+0\.13\{\\approx\}\{\+\}0\.13\) and not*for our diagonal cell*\(0\.1720\.172at its best lr\)\. The sweeps span one decade at 5 seeds; the official kernel’s own tuned optimum is0\.9550\.955\(§[4](https://arxiv.org/html/2609.16183#S4)\)\.
4. 4\.The attention control is valid only where reported\.Running attention \(absolute and rotary\) at a healthy head dimension \(2 heads,d=32d\{=\}32\) through the factorial’s matched one\-layer config yields collapse at everyKK\(≈0\.11\{\\approx\}0\.11–0\.170\.17\) with the rebinding control*equal*to eval accuracy—the fingerprint of never reading bindings\. This is structural\. MQAR is an induction\-head task and induction circuits require two attention layers\([Olsson et al\., 2022](https://arxiv.org/html/2609.16183#bib.bib35)\), while the recurrent cells solve it at depth one by writing bindings into state\. A depth\-matched attention row therefore cannot exist atL=1L\{=\}1, and the valid attention ceilings live in the two\-layer regimes, where we report them \(fixed wall1\.0001\.000; lexical curriculum6/106/10; collision\-ramp7/107/10;L=256L\{=\}256ramp4/54/5—parity with the armed cell’s4/54/5\)\. This paper makes no attention\-superiority claims beyond those settings\.
5. 5\.Cross\-family magnitude comparisons are indicative\.mamba2\_ref’s decay is scalar\-per\-head, the RWKV\-7 reference uses vector decay, andgated\_deltanetuses scalar decay\. Within\-family toggles are exact; cross\-family magnitude comparisons \(e\.g\., convolution effect size in delta vs\. diagonal families\) carry a parameterization caveat\.
6. 6\.Interpretation boundaries we respect\.No rank\-1\-vs\-diagonal*class*claim \(§[4\.3](https://arxiv.org/html/2609.16183#S4.SS3)kills it\); no claim that decay helps state\-tracking \(underpowered\); no length extrapolation of the curriculum fix beyondL=64L\{=\}64; no circuit\-complexity claims from theS5S\_\{5\}guardrail; no bidirectional\-vs\-causal advantage claim under collision \(n\.s\. at 10 seeds\); the depth\-novelty\-gating account of the causal arm’s solvability is a hypothesis\.

## 10Discussion

What “recurrent models are bad at recall” decomposes into\.At matched state, it decomposes into five things: a missing convolution \(large, fixable for free in state terms\), a transition effect \(real, within\-cell\), a decay tax \(real, scalar\-gate\-specific\), an interference failure under sparse supervision \(fixable by curriculum, within a length boundary\), and a directionality property \(bidirectional denoisers get query\-first reading natively\)\. None of these is “the recurrence cannot bind\.”

When is a hybrid actually needed?The results re\-scope the standard “add attention layers for retrieval” move\. Two of the classic triggers dissolve under cheap interventions: the unarmed\-cell recall deficit \(add the convolution\) and the within\-capacity haystack wall \(train with a distance curriculum\)\. Three things remain as genuine hybrid territory: loads beyond state capacity \(the gracefulKK\-limit\), lengths beyond the curriculum boundary \(L≥256L\{\\geq\}256here, pending shaped curricula\), and query\-conditioned regimes a fixed\-state writer cannot anticipate\. A hybrid decision made*after*arming and curriculum is a different, and smaller, decision\.

For recurrent diffusion LMs\.The JRT\-for\-free result gives bidirectional recurrent denoisers\([Singh et al\., 2025](https://arxiv.org/html/2609.16183#bib.bib25);[Lin et al\., 2026](https://arxiv.org/html/2609.16183#bib.bib26)\)a principled recall story that causal recurrent LMs lack: the objective itself buys the second read\. The requirements are concrete—two layers minimum and a short convolution in the backward stream—which makes them design guidance for this model class rather than folklore\. The companion paper takes the curriculum half of this story to converted diffusion LMs at1\.71\.7B and88B\([Boesch and Wee, 2026](https://arxiv.org/html/2609.16183#bib.bib36)\); whether the query\-first mechanism itself carries to scale is untested\.

A conjecture, and a design question\.Within this study, one apparent architectural wall dissolved entirely into a training\-distribution problem, and convergent evidence points the same way for recurrent recall and length generalization\([Okpekpe and Orvieto, 2025](https://arxiv.org/html/2609.16183#bib.bib5);[Buitrago Ruiz and Gu, 2025](https://arxiv.org/html/2609.16183#bib.bib8);[Blouir et al\., 2024](https://arxiv.org/html/2609.16183#bib.bib21)\)\. We state the strong form as a conjecture worth falsifying:*within the capacity regime of a fixed\-state recurrence, apparent retrieval walls are training\-coverage gaps*\. The length boundary of \(ii\) above is its first direct test—the cell that sits at0/60/6under a time\-based ramp atL=512L\{=\}512locks in6/66/6once the ramp is gated on measured competence, so that boundary was the schedule, not the circuit\. What remains open against the alternative \(a genuine interference ceiling of the two\-layer circuit\) is whether some length defeats the gated schedule too\. Separately, decay costs−0\.32\-0\.32on recall here while showing*no measured*state\-tracking benefit at our power \(p=0\.65p\{=\}0\.65\); context\-modulated forgetting\([Kimi Team, 2025](https://arxiv.org/html/2609.16183#bib.bib16)\)becomes the interesting design axis only if a benefit emerges under better\-powered tests\.

## Reproducibility Statement

All cells, tasks, and controls run in one open harness \(pure\-PyTorch reference cells; a state\-matched, kernel\-free Mamba\-2 comparator; fp32 paths\) on commodity GPUs, including pre\-Volta hardware via an atomic\-free two\-pass Triton backward \(Appendix[A](https://arxiv.org/html/2609.16183#A1)\)\. Appendix[B](https://arxiv.org/html/2609.16183#A2)lists the exact commands for every table\. Code and result JSONs are publicly available at[https://github\.com/JIBSIL/dualgoose](https://github.com/JIBSIL/dualgoose)\.

## Acknowledgments

This work used computing resources provided by the Rosen Center for Advanced Computing \(RCAC\) at Purdue University\([McCartney et al\., 2014](https://arxiv.org/html/2609.16183#bib.bib37)\)\.

## Impact Statement

This paper presents controlled, small\-scale measurements aimed at advancing the scientific understanding of sequence\-model architectures: which ingredients of efficient recurrent cells support in\-context recall, and when training rather than architecture is the binding constraint\. The main downstream impact, if the findings transfer to scale, is more capable linear\-time language models at lower inference cost and energy\. We see no societal consequences specific to this work beyond those generic to improving machine learning methods\.

## References

- Aroraet al\.\(2025\)A\. Arora, N\. Rathi, N\. R\. Selvam, R\. Csordás, D\. Jurafsky, and C\. PottsMechanistic evaluation of transformers and state space models\.arXiv preprint arXiv:2505\.15105\.Cited by:[item 1](https://arxiv.org/html/2609.16183#S1.I1.i1.p1.1),[§1](https://arxiv.org/html/2609.16183#S1.p2.1),[§2](https://arxiv.org/html/2609.16183#S2.p2.1),[§4\.2](https://arxiv.org/html/2609.16183#S4.SS2.p2.1)\.
- Aroraet al\.\(2023\)S\. Arora, S\. Eyuboglu, A\. Timalsina, I\. Johnson, M\. Poli, J\. Zou, A\. Rudra, and C\. RéZoology: measuring and improving recall in efficient language models\.arXiv preprint arXiv:2312\.04927\.Cited by:[§1](https://arxiv.org/html/2609.16183#S1.p1.1),[§2](https://arxiv.org/html/2609.16183#S2.p1.1)\.
- Aroraet al\.\(2024a\)S\. Arora, S\. Eyuboglu, M\. Zhang, A\. Timalsina, S\. Alberti, D\. Zinsley, J\. Zou, A\. Rudra, and C\. RéSimple linear attention language models balance the recall\-throughput tradeoff\.arXiv preprint arXiv:2402\.18668\.Cited by:[§1](https://arxiv.org/html/2609.16183#S1.p1.1),[§2](https://arxiv.org/html/2609.16183#S2.p1.1)\.
- Aroraet al\.\(2024b\)S\. Arora, A\. Timalsina, A\. Singhal, S\. Eyuboglu, X\. Zhao, A\. Rao, A\. Rudra, and C\. RéJust read twice: closing the recall gap for recurrent language models\.arXiv preprint arXiv:2407\.05483\.Cited by:[item 4](https://arxiv.org/html/2609.16183#S1.I1.i4.p1.1),[§2](https://arxiv.org/html/2609.16183#S2.p5.1),[§7](https://arxiv.org/html/2609.16183#S7.p1.1)\.
- Austinet al\.\(2021\)J\. Austin, D\. D\. Johnson, J\. Ho, D\. Tarlow, and R\. van den BergStructured denoising diffusion models in discrete state\-spaces\.arXiv preprint arXiv:2107\.03006\.Cited by:[§2](https://arxiv.org/html/2609.16183#S2.p5.1)\.
- Bachmann and Nagarajan \(2024\)G\. Bachmann and V\. NagarajanThe pitfalls of next\-token prediction\.arXiv preprint arXiv:2403\.06963\.Cited by:[§3\.4](https://arxiv.org/html/2609.16183#S3.SS4.p1.1)\.
- Ben\-Kishet al\.\(2024\)A\. Ben\-Kish, I\. Zimerman, S\. Abu\-Hussein, N\. Cohen, A\. Globerson, L\. Wolf, and R\. GiryesDeciMamba: exploring the length extrapolation potential of Mamba\.arXiv preprint arXiv:2406\.14528\.Cited by:[§2](https://arxiv.org/html/2609.16183#S2.p3.1)\.
- Bengioet al\.\(2009\)Y\. Bengio, J\. Louradour, R\. Collobert, and J\. WestonCurriculum learning\.Proceedings of ICML\.Cited by:[§2](https://arxiv.org/html/2609.16183#S2.p4.1)\.
- Blouiret al\.\(2024\)S\. Blouir, J\. T\.H\. Smith, and A\. AnastasopoulosBirdie: advancing state space models with reward\-driven objectives and curricula\.InProceedings of EMNLP,Cited by:[§10](https://arxiv.org/html/2609.16183#S10.p4.1),[§2](https://arxiv.org/html/2609.16183#S2.p4.1),[§6](https://arxiv.org/html/2609.16183#S6.p2.1)\.
- Boesch and Wee \(2026\)J\. Boesch and A\. WeeDreamingGoose: staged distillation from autoregressive transformers to bidirectional recurrent diffusion language models\.Note:Companion preprintCited by:[§10](https://arxiv.org/html/2609.16183#S10.p3.1),[§6](https://arxiv.org/html/2609.16183#S6.p3.1),[§6](https://arxiv.org/html/2609.16183#S6.p7.1)\.
- Buitrago Ruiz and Gu \(2025\)R\. Buitrago Ruiz and A\. GuUnderstanding and improving length generalization in recurrent models\.arXiv preprint arXiv:2507\.02782\.Cited by:[§10](https://arxiv.org/html/2609.16183#S10.p4.1),[§2](https://arxiv.org/html/2609.16183#S2.p3.1),[§2](https://arxiv.org/html/2609.16183#S2.p4.1),[§6](https://arxiv.org/html/2609.16183#S6.p2.1)\.
- Chenet al\.\(2024\)Y\. Chen, X\. Zhang, S\. Hu, X\. Han, Z\. Liu, and M\. SunStuffed Mamba: state collapse and state capacity of RNN\-based long\-context modeling\.arXiv preprint arXiv:2410\.07145\.Cited by:[§2](https://arxiv.org/html/2609.16183#S2.p3.1)\.
- Dao and Gu \(2024\)T\. Dao and A\. GuTransformers are SSMs: generalized models and efficient algorithms through structured state space duality\.arXiv preprint arXiv:2405\.21060\.Cited by:[§1](https://arxiv.org/html/2609.16183#S1.p2.1),[§3\.1](https://arxiv.org/html/2609.16183#S3.SS1.p5.1)\.
- Delétanget al\.\(2023\)G\. Delétang, A\. Ruoss, J\. Grau\-Moya, T\. Genewein, L\. K\. Wenliang, E\. Catt, C\. Cundy, M\. Hutter, S\. Legg, J\. Veness, and P\. A\. OrtegaNeural networks and the Chomsky hierarchy\.arXiv preprint arXiv:2207\.02098\.Cited by:[§2](https://arxiv.org/html/2609.16183#S2.p6.1)\.
- Grazziet al\.\(2025\)R\. Grazzi, J\. Siems, J\. K\.H\. Franke, A\. Zela, F\. Hutter, and M\. PontilUnlocking state\-tracking in linear RNNs through negative eigenvalues\.arXiv preprint arXiv:2411\.12537\.Cited by:[§2](https://arxiv.org/html/2609.16183#S2.p6.1),[§8](https://arxiv.org/html/2609.16183#S8.p2.1)\.
- Gu and Dao \(2023\)A\. Gu and T\. DaoMamba: linear\-time sequence modeling with selective state spaces\.arXiv preprint arXiv:2312\.00752\.Cited by:[§1](https://arxiv.org/html/2609.16183#S1.p2.1)\.
- Hewitt and Liang \(2019\)J\. Hewitt and P\. LiangDesigning and interpreting probes with control tasks\.arXiv preprint arXiv:1909\.03368\.Cited by:[§3\.4](https://arxiv.org/html/2609.16183#S3.SS4.p1.1)\.
- Holtzmanet al\.\(2021\)A\. Holtzman, P\. West, V\. Shwartz, Y\. Choi, and L\. ZettlemoyerSurface form competition: why the highest probability answer isn’t always right\.arXiv preprint arXiv:2104\.08315\.Cited by:[§3\.4](https://arxiv.org/html/2609.16183#S3.SS4.p1.1)\.
- Hsiehet al\.\(2024\)C\. Hsieh, S\. Sun, S\. Kriman, S\. Acharya, D\. Rekesh, F\. Jia, Y\. Zhang, and B\. GinsburgRULER: what’s the real context size of your long\-context language models?\.arXiv preprint arXiv:2404\.06654\.Cited by:[§1](https://arxiv.org/html/2609.16183#S1.p1.1),[§2](https://arxiv.org/html/2609.16183#S2.p3.1)\.
- Jelassiet al\.\(2024\)S\. Jelassi, D\. Brandfonbrener, S\. M\. Kakade, and E\. MalachRepeat after me: transformers are better than state space models at copying\.arXiv preprint arXiv:2402\.01032\.Cited by:[§1](https://arxiv.org/html/2609.16183#S1.p1.1),[§2](https://arxiv.org/html/2609.16183#S2.p3.1)\.
- Katharopouloset al\.\(2020\)A\. Katharopoulos, A\. Vyas, N\. Pappas, and F\. FleuretTransformers are RNNs: fast autoregressive transformers with linear attention\.arXiv preprint arXiv:2006\.16236\.Cited by:[§2](https://arxiv.org/html/2609.16183#S2.p1.1)\.
- Kimi Team \(2025\)Kimi TeamKimi linear: an expressive, efficient attention architecture\.arXiv preprint arXiv:2510\.26692\.Cited by:[§10](https://arxiv.org/html/2609.16183#S10.p4.1),[§4\.2](https://arxiv.org/html/2609.16183#S4.SS2.p5.1)\.
- Linet al\.\(2026\)K\. Lin, Y\. Luo, Z\. Su, Y\. Song, and A\. RaoTriplet\-block diffusion RWKV\.arXiv preprint arXiv:2605\.25969\.Cited by:[§10](https://arxiv.org/html/2609.16183#S10.p3.1),[§2](https://arxiv.org/html/2609.16183#S2.p5.1)\.
- Liuet al\.\(2023\)B\. Liu, J\. T\. Ash, S\. Goel, A\. Krishnamurthy, and C\. ZhangTransformers learn shortcuts to automata\.arXiv preprint arXiv:2210\.10749\.Cited by:[§2](https://arxiv.org/html/2609.16183#S2.p6.1)\.
- McCartneyet al\.\(2014\)G\. McCartney, T\. Hacker, and B\. YangEmpowering Faculty: A Campus Cyberinfrastructure Strategy for Research Communities\.Educause Review\.Cited by:[Acknowledgments](https://arxiv.org/html/2609.16183#Sx2.p1.1)\.
- Merrillet al\.\(2024\)W\. Merrill, J\. Petty, and A\. SabharwalThe illusion of state in state\-space models\.arXiv preprint arXiv:2404\.08819\.Cited by:[§2](https://arxiv.org/html/2609.16183#S2.p6.1),[§8](https://arxiv.org/html/2609.16183#S8.p2.1)\.
- Okpekpe and Orvieto \(2025\)D\. Okpekpe and A\. OrvietoRevisiting associative recall in modern recurrent models\.arXiv preprint arXiv:2508\.19029\.Cited by:[item 1](https://arxiv.org/html/2609.16183#S1.I1.i1.p1.1),[§1](https://arxiv.org/html/2609.16183#S1.p2.1),[§10](https://arxiv.org/html/2609.16183#S10.p4.1),[§2](https://arxiv.org/html/2609.16183#S2.p2.1),[§2](https://arxiv.org/html/2609.16183#S2.p4.1),[§4\.2](https://arxiv.org/html/2609.16183#S4.SS2.p3.1),[§6](https://arxiv.org/html/2609.16183#S6.p2.1),[item 3](https://arxiv.org/html/2609.16183#S9.I1.i3.p1.1)\.
- Olssonet al\.\(2022\)C\. Olsson, N\. Elhage, N\. Nanda, N\. Joseph, N\. DasSarma, T\. Henighan, B\. Mann, A\. Askell, Y\. Bai, A\. Chen,et al\.In\-context learning and induction heads\.Transformer Circuits Thread\.Cited by:[item 4](https://arxiv.org/html/2609.16183#S9.I1.i4.p1.1)\.
- Parnichkunet al\.\(2025\)R\. N\. Parnichkun, N\. Tumma, A\. W\. Thomas, A\. Moro, Q\. An, T\. Suzuki, A\. Yamashita, M\. Poli, and S\. MassaroliQuantifying memory utilization with effective state\-size\.arXiv preprint arXiv:2504\.19561\.Cited by:[item 1](https://arxiv.org/html/2609.16183#S1.I1.i1.p1.1),[§1](https://arxiv.org/html/2609.16183#S1.p2.1),[§2](https://arxiv.org/html/2609.16183#S2.p2.1),[§4\.2](https://arxiv.org/html/2609.16183#S4.SS2.p2.1)\.
- Penget al\.\(2025\)B\. Peng, R\. Zhang, D\. Goldstein,et al\.RWKV\-7 “Goose” with expressive dynamic state evolution\.arXiv preprint arXiv:2503\.14456\.Cited by:[§1](https://arxiv.org/html/2609.16183#S1.p2.1),[§2](https://arxiv.org/html/2609.16183#S2.p1.1),[§3\.1](https://arxiv.org/html/2609.16183#S3.SS1.p4.1)\.
- Sahooet al\.\(2024\)S\. S\. Sahoo, M\. Arriola, Y\. Schiff, A\. Gokaslan, E\. Marroquin, J\. T\. Chiu, A\. Rush, and V\. KuleshovSimple and effective masked diffusion language models\.arXiv preprint arXiv:2406\.07524\.Cited by:[§2](https://arxiv.org/html/2609.16183#S2.p5.1),[§7](https://arxiv.org/html/2609.16183#S7.p1.1)\.
- Schlaget al\.\(2021\)I\. Schlag, K\. Irie, and J\. SchmidhuberLinear transformers are secretly fast weight programmers\.arXiv preprint arXiv:2102\.11174\.Cited by:[§2](https://arxiv.org/html/2609.16183#S2.p1.1),[§3\.1](https://arxiv.org/html/2609.16183#S3.SS1.p2.1)\.
- Singhet al\.\(2025\)V\. Singh, O\. Ostapenko, P\. Noël, E\. Belilovsky, and T\. ScholakDiffuMamba: high\-throughput diffusion LMs with Mamba backbone\.arXiv preprint arXiv:2511\.15927\.Cited by:[§10](https://arxiv.org/html/2609.16183#S10.p3.1),[§2](https://arxiv.org/html/2609.16183#S2.p5.1),[§3\.1](https://arxiv.org/html/2609.16183#S3.SS1.p8.1),[§7](https://arxiv.org/html/2609.16183#S7.p1.1)\.
- Sridhar and Johansen \(2026\)A\. Sridhar and A\. JohansenEcho: KV\-cache\-free associative recall with spectral Koopman operators\.arXiv preprint arXiv:2605\.06997\.Cited by:[item 2](https://arxiv.org/html/2609.16183#S1.I1.i2.p1.1),[§2](https://arxiv.org/html/2609.16183#S2.p3.1),[item 2](https://arxiv.org/html/2609.16183#S5.I1.i2.p1.1),[§5](https://arxiv.org/html/2609.16183#S5.p4.1)\.
- Vaswaniet al\.\(2017\)A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, and I\. PolosukhinAttention is all you need\.Advances in Neural Information Processing Systems\.Cited by:[§3\.1](https://arxiv.org/html/2609.16183#S3.SS1.p7.1)\.
- Yanget al\.\(2024a\)S\. Yang, J\. Kautz, and A\. HatamizadehGated delta networks: improving Mamba2 with delta rule\.arXiv preprint arXiv:2412\.06464\.Cited by:[§1](https://arxiv.org/html/2609.16183#S1.p2.1),[§2](https://arxiv.org/html/2609.16183#S2.p1.1),[§3\.1](https://arxiv.org/html/2609.16183#S3.SS1.p3.1)\.
- Yanget al\.\(2024b\)S\. Yang, B\. Wang, Y\. Zhang, Y\. Shen, and Y\. KimParallelizing linear transformers with the delta rule over sequence length\.arXiv preprint arXiv:2406\.06484\.Cited by:[§1](https://arxiv.org/html/2609.16183#S1.p2.1),[§2](https://arxiv.org/html/2609.16183#S2.p1.1),[§3\.1](https://arxiv.org/html/2609.16183#S3.SS1.p2.1)\.

## Appendix AThe Commodity\-Hardware Harness

Three engineering choices make the factorial cheap to replicate:

#### Kernel\-free cells\.

Every arm has a pure\-PyTorch reference implementation; the Mamba\-2 comparator \(TinyMamba2RefSSM\) implements the SSD recurrence with an explicitdstated\_\{\\mathrm\{state\}\}and requires neither Triton normamba\_ssm, so the full factorial executes on any CUDA GPU \(and, slowly, CPU\)\. Thekernel\_size=1setting neutralizes its convolution for the toggle\.

#### Pre\-Volta support\.

The default Triton backward for the RWKV\-7\-style scan emits\.acq\_relatomics thatptxasrejects belowsm\_70\. An atomic\-free two\-pass backward reproduces the reference scan’s forward and gradients and runs on Pascal \(sm\_61\) at≈26×\{\\approx\}26\\timesthe pure\-PyTorch reference throughput \(≈2,980\{\\approx\}2\{,\}980vs\.≈113\{\\approx\}113tok/s atd=256d\{=\}256/6 layers/ctx 512\), with an\-\-fp32flag for hardware without bf16\. The headline experiments in this paper were run on a single GTX 1070\.

#### Determinism and floors\.

Runs are seed\-pinned and resumable; every task ships its floor calculators \(value\-marginal, no\-binding, table floor1/K1/K, positional\)\. The rebinding control is validated by a self\-test \(\-\-selftest\) that confirms the derangement logic\.

## Appendix BReproduction Commands

Constrained\-state factorial andKK\-sweep \(Tables[1](https://arxiv.org/html/2609.16183#S4.T1),[2](https://arxiv.org/html/2609.16183#S4.T2)\)\.\-\-fp32is required on pre\-bf16 \(Pascal\) hardware and may be dropped on Ampere and newer; it applies to every command below exceptdg\_c9\_harden\.py, which has no such flag:

> python scripts/mqar\_probe\.py \-\-models armed\_gdn gated\_deltanet deltanet dg\_rank\_ref dg\_diag\_ref mamba2\_ref mamba2\_ref\_noconv attention \-\-d\-model 32 \-\-layers 1 \-\-mamba2\-d\-state 16 \-\-ks 8 16 32 \-\-seeds 0 1 2 \-\-fp32

Hardened transition ablation with rebinding control \(Table[3](https://arxiv.org/html/2609.16183#S4.T3)\):

> python scripts/dg\_c9\_harden\.py \-\-ks 16 32 \-\-seeds 0 1 2 3 4 5 6 7 8 9 python scripts/dg\_c9\_harden\.py \-\-selftest

Haystack wall \(Table[4](https://arxiv.org/html/2609.16183#S5.T4)\):

> python scripts/long\_retrieval\_probe\.py \-\-models armed\_gdn dg\_rank\_ref mamba2\_ref attention \-\-seq\-lens 64 128 256 512 \-\-pairs 4 \-\-seeds 0 1 \-\-heads 2 \-\-mamba2\-d\-state 16 \-\-fp32

Curriculum and collision\-key discriminator \(Tables[5](https://arxiv.org/html/2609.16183#S6.T5),[6](https://arxiv.org/html/2609.16183#S7.T6)\); add\-\-collisionfor the discriminator:

> python scripts/long\_retrieval\_probe\.py \-\-models armed\_gdn armed\_gdn\_fwd attention \-\-seq\-lens 64 \-\-pairs 4 \-\-seeds 0 \-\-steps 3000 \-\-batch 256 \-\-layers 2 \-\-heads 2 \-\-dense 4 \-\-curriculum \-\-fp32

S5S\_\{5\}guardrail \(Table[7](https://arxiv.org/html/2609.16183#S8.T7)\):

> python scripts/state\_tracking\_probe\.py \-\-models armed\_gdn gated\_deltanet \-\-group\-n 5 \-\-ls 8 16 \-\-d\-model 128 \-\-seeds 0 1 2 3 4 5 6 7 8 9 \-\-fp32

## Appendix CFloors and Controls, Stated

For each task we report: \(i\) thevalue\-marginal\(probability of the correct value under the answer distribution:1/26≈0\.0381/26\\approx 0\.038for MQAR;1/54≈0\.0191/54\\approx 0\.019for haystack/collision;1/120≈0\.0081/120\\approx 0\.008forS5S\_\{5\}, exact because generators are uniform over the full group\); \(ii\) the strongestno\-binding strategy\(MQAR: emit a random table value,≈0\.068\{\\approx\}0\.068; collision: emit a random table\-*position*token,≈0\.26\{\\approx\}0\.26\); \(iii\) thetable floor1/K1/Kfor the rebinding control\. No headline number in this paper is within10×10\\timesof its floor except where explicitly marked as “chance\.” TheS5S\_\{5\}probe additionally masks all product slots simultaneously and uses single\-token answers, making prefix\-completion leakage structurally impossible\.

Similar Articles

What Training Data Teaches RL Memory Agents: An Empirical Study of Curriculum Effects in Memory-Augmented QA

arXiv cs.CL

This paper empirically studies how the composition of training data (curriculum) affects the skills learned by RL-based memory agents in multi-session question answering. It finds that curriculum composition acts as a fine-grained lever on specialization, with mixed benchmarks yielding the best overall performance and narrow out-of-domain sets transferring targeted temporal reasoning skills.

Learning to Reason with Curriculum II: Compositional Generalization

arXiv cs.LG

This paper theoretically analyzes how curriculum learning, by decomposing complex problems into simpler sub-problems and composing solutions, can dramatically reduce the sample complexity of learning to simulate sequential computations (semiautomata) compared to direct methods, achieving subpolynomial supervision requirements in supervised fine-tuning and exponentially weaker coverage conditions in reinforcement learning with verifiable rewards.