PolicyAttention: Softmax Attention Implements Policy Mirror Descent for Closed-Loop Control
Summary
This paper proposes PolicyAttention, a method using causal softmax attention to implement policy mirror descent for closed-loop control in reinforcement learning, achieving lower losses than existing adaptations in experiments.
View Cached Full Text
Cached at: 09/29/26, 09:38 AM
# PolicyAttention: Softmax Attention Implements Policy Mirror Descent for Closed-Loop Control
Source: [https://arxiv.org/html/2609.30500](https://arxiv.org/html/2609.30500)
Yuhe Sui[https://orcid.org/0009-0002-7456-0804](https://orcid.org/0009-0002-7456-0804)Affiliation:Quantitative Research Society, SingaporeAffiliation:Nanyang Technological University, SingaporeYingzhi Tang[https://orcid.org/0009-0004-7571-0100](https://orcid.org/0009-0004-7571-0100)Affiliation:Quantitative Research Society, SingaporeAffiliation:The Chinese University of Hong Kong, Shenzhen, ChinaShufang Chen[https://orcid.org/0009-0009-3627-7726](https://orcid.org/0009-0009-3627-7726)Affiliation:Quantitative Research Society, SingaporeAffiliation:The University of Hong Kong, Hong Kong SAR, China
###### Abstract
A Transformer can match a reinforcement\-learning update on fixed inputs yet fail once its own returned policies alter subsequent inputs\. We study this closed\-loop gap for negative\-entropy policy mirror descent \(PMD\), whose update is softmax\-native:PMDη\(π,Q\)=softmax\(logπ\+ηQ\)\\operatorname\{PMD\}\_\{\\eta\}\(\\pi,Q\)=\\operatorname\{softmax\}\(\\log\\pi\+\\eta Q\)\. The reference recursion is known Q\-TD\-PMD; our contribution is its causal implementation and learned closed\-loop interface\. A fixed causal\-softmax decoder realizes an inexact actor–environment–one\-step\-critic loop with explicit actor, routing, and sampling residuals, and a returned\-policy theorem propagates those errors to the policy actually output, with no unused final critic residual\. A finite\-cap pre\-LayerNorm/final\-LayerNorm compilation and a scoped effective\-coordinate result connect the construction to standard Transformer operations and the PMD\-aligned training objective\.
Empirically, a separately trained pre\-LN Transformer is closest to PMD among fixed update hypotheses, satisfies a preregistered repeated\-control criterion on five fresh runs, and retains that behavior under no\-retraining environment shifts\. At larger retrained sizes, PolicyAttention attains median normalized returned\-policy losses0\.01930\.0193atS=8S=8and0\.02730\.0273atS=16S=16\. Under the common scoring harness, these are1818–28×28\\timeslower than our qualified Liang–Lai and Algorithm Distillation adaptations\. PolicyAttention uses an exact one\-step Bellman backup with model access, whereas the external adaptations update from sampled interactions; this is therefore a common\-harness performance separation rather than an information\-matched comparison\. Together, the results distinguish local algorithmic fidelity, closed\-loop reliability, and task performance, and show how softmax attention can serve as the policy\-improvement operator itself\.
## 1Introduction
Transformers can execute learning algorithms in context, but an adaptive controller faces a stronger test than a static predictor: its current output changes the distribution of its future inputs\. In reinforcement learning \(RL\), a returned policy changes successor actions, critic targets, and the context processed on the next round\. A model can therefore look algorithmic for one step and cease to be useful when repeatedly deployed\.
PolicyAttention asks whether*standard causal softmax attention can implement and learn a genuine policy\-improvement algorithm that remains effective after its own policies recursively alter the control inputs*\. Negative\-entropy PMD is a natural target because its statewise update is exactly
PMDη\(π,Q\)\(a∣s\)=π\(a∣s\)eηQ\(s,a\)∑bπ\(b∣s\)eηQ\(s,b\)=softmax\(logπ\+ηQ\)a\.\\operatorname\{PMD\}\_\{\\eta\}\(\\pi,Q\)\(a\\mid s\)=\\frac\{\\pi\(a\\mid s\)e^\{\\eta Q\(s,a\)\}\}\{\\sum\_\{b\}\\pi\(b\\mid s\)e^\{\\eta Q\(s,b\)\}\}=\\operatorname\{softmax\}\(\\log\\pi\+\\eta Q\)\_\{a\}\.\(1\)The identity and the associated PMD/TD\-PMD control recursions are prior RL methodology\([Xiao, 2022](https://arxiv.org/html/2609.30500#bib.bib8);[Johnson et al\., 2023](https://arxiv.org/html/2609.30500#bib.bib9);[Geist et al\., 2019](https://arxiv.org/html/2609.30500#bib.bib10);[Lan, 2023](https://arxiv.org/html/2609.30500#bib.bib7);[Liu et al\., 2025](https://arxiv.org/html/2609.30500#bib.bib11)\)\. Our object is the implementation layer: a causal softmax realization with measurable residuals, a returned\-policy interface for those residuals, and a learned closed\-loop test\.
policy \+ value\(πk,Qk\)\(\\pi\_\{k\},Q\_\{k\}\)softmax PMD actorsoftmax\(logπ\+ηQ\)\\operatorname\{softmax\}\(\\log\\pi\+\\eta Q\)environmentinteractionone\-step criticℱπk\+1Qk\\mathcal\{F\}^\{\\pi\_\{k\+1\}\}Q\_\{k\}next context\(πk\+1,Qk\+1\)\(\\pi\_\{k\+1\},Q\_\{k\+1\}\)
TheoryLearned controllerLiterature\-facing comparisonFocusfixed causal\-softmax PMD/critic implementation; explicit actor, routing, sampling, and returned\-policy residualsfour\-layer pre\-LN Transformer; PMD\-aligned objective; exact\-critic repeated\-control test plus descriptive learned\-critic evidencePolicyAttention vs Liang–Lai and Algorithm Distillation; Exact PMD is the oracle; learned PMD\-target model is calibration only
Figure 1:One mechanism\-to\-control chain\.The causal loop is shown above; the table below separates the constructive theorem, the learned\-controller evidence, and the literature\-facing benchmark\. Exact PMD is the non\-learned oracle, while the learned PMD\-target model serves only as calibration\.The contribution is one chain\.First, a fixed two\-block/two\-head causal decoder realizes PMD followed by one\-step policy evaluation under the policy it actually returned, with explicit finite routing and sampling errors\.Second, these errors feed a last\-policy theorem that controls the final returned policy and correctly omits the unused final critic residual\.Third, a specially compiled pre\-LN/final\-LN construction and reverse\-KL/Fisher effective\-coordinate result explain how the same target computation fits standard normalized operations and the PMD\-aligned objective\.Fourth, fresh experiments show the learned PMD\-like actor survives adaptive repetition and fixed\-shape distribution shifts\.Finally, a qualified common\-harness benchmark compares PolicyAttention with a matched PMD\-target reference Transformer and two published\-method adaptations, placing the mechanism result in a literature\-facing performance context\.
## 2Related Work
Transformers have been shown to implement or learn in\-context regression and gradient procedures\([Garg et al\., 2022](https://arxiv.org/html/2609.30500#bib.bib1);[Akyürek et al\., 2023](https://arxiv.org/html/2609.30500#bib.bib2);[von Oswald et al\., 2023](https://arxiv.org/html/2609.30500#bib.bib3)\)\. In RL, Wang et al\. learn TD\-style policy evaluation in context\([Wang et al\., 2025](https://arxiv.org/html/2609.30500#bib.bib4)\); Xie et al\. give a standard\-softmax construction for weighted softmax TD and policy evaluation\([Xie et al\., 2026](https://arxiv.org/html/2609.30500#bib.bib6)\)\. Liang and Lai construct semi\-gradient SARSA and actor–critic updates with linear self\-attention and analyze teacher\-mimicking learning\([Liang and Lai, 2026](https://arxiv.org/html/2609.30500#bib.bib5)\)\. Algorithm Distillation \(AD\) trains a causal sequence model on complete learning histories and obtains empirical in\-context improvement without weight updates\([Laskin et al\., 2023](https://arxiv.org/html/2609.30500#bib.bib12)\); S5RL uses structured state\-space models in a different online\-PPO/meta\-RL paradigm\([Lu et al\., 2023](https://arxiv.org/html/2609.30500#bib.bib13)\)\.
Table 1:Source\-verified capability map\. “Softmax” denotes standard softmax attention in the relevant construction/model; PolicyAttention’s final column denotes the returned\-policy residual interface introduced here\.The gap is not “control versus no control\.” It is the intersection of*standard softmax policy improvement, explicit causal implementation errors, and returned\-policy closed\-loop accounting*\. The benchmark numbers for Liang–Lai and AD below are ours from faithful common\-harness adaptations, not results reported by those papers\.
## 3Causal softmax realization
Letℳ=\(𝒮,𝒜,P,r,γ\)\\mathcal\{M\}=\(\\mathcal\{S\},\\mathcal\{A\},P,r,\\gamma\)be finite with\|r\|≤Rmax\|r\|\\leq R\_\{\\max\}andB=Rmax/\(1−γ\)B=R\_\{\\max\}/\(1\-\\gamma\)\. Define the one\-step policy\-evaluation and optimality operators
\(ℱπQ\)\(s,a\)\\displaystyle\(\\mathcal\{F\}^\{\\pi\}Q\)\(s,a\)=r\(s,a\)\+γ𝔼S′\|s,a⟨π\(⋅∣S′\),Q\(S′,⋅\)⟩,\\displaystyle=r\(s,a\)\+\\gamma\\mathbb\{E\}\_\{S^\{\\prime\}\\mid s,a\}\\langle\\pi\(\\cdot\\mid S^\{\\prime\}\),Q\(S^\{\\prime\},\\cdot\)\\rangle,\(2\)\(ℱQ\)\(s,a\)\\displaystyle\(\\mathcal\{F\}Q\)\(s,a\)=r\(s,a\)\+γ𝔼S′\|s,amaxbQ\(S′,b\)\.\\displaystyle=r\(s,a\)\+\\gamma\\mathbb\{E\}\_\{S^\{\\prime\}\\mid s,a\}\\max\_\{b\}Q\(S^\{\\prime\},b\)\.\(3\)At roundkk, the implemented actor returnsπ^k\+1\\widehat\{\\pi\}\_\{k\+1\}from\(π^k,Qk\)\(\\widehat\{\\pi\}\_\{k\},Q\_\{k\}\), then the environment generates critic data under that returned policy\. Define same\-context residuals
ζk\\displaystyle\\zeta\_\{k\}=maxs∥π^k\+1\(⋅∣s\)−PMDηk\(π^k,Qk\)\(⋅∣s\)∥1,\\displaystyle=\\max\_\{s\}\\lVert\\widehat\{\\pi\}\_\{k\+1\}\(\\cdot\\mid s\)\-\\operatorname\{PMD\}\_\{\\eta\_\{k\}\}\(\\widehat\{\\pi\}\_\{k\},Q\_\{k\}\)\(\\cdot\\mid s\)\\rVert\_\{1\},\(4\)δk\\displaystyle\\delta\_\{k\}=∥Qk\+1−ℱπ^k\+1Qk∥∞\.\\displaystyle=\\lVert Q\_\{k\+1\}\-\\mathcal\{F\}^\{\\widehat\{\\pi\}\_\{k\+1\}\}Q\_\{k\}\\rVert\_\{\\infty\}\.\(5\)These are local implementation quantities, not task losses after two trajectories diverge\.
#### Construction\.
The normalization\-free circuit has two blocks, two heads, head dimensionS\+A\+3S\+A\+3, payload width6S\+6A\+146S\+6A\+14, a causal mask, no positional embeddings, and a role\-gated ReLU map\. Actor memories carry state/action identity,logπ^k\\log\\widehat\{\\pi\}\_\{k\}, andQkQ\_\{k\}; intended attention scores are a common state offset pluslogπ^k\(a∣s\)\+ηkQk\(s,a\)\\log\\widehat\{\\pi\}\_\{k\}\(a\\mid s\)\+\\eta\_\{k\}Q\_\{k\}\(s,a\), so intended\-group softmax is exactly PMD\. Finite off\-group mass yields
ζk≤2Rkact1\+Rkact,Rkact≤\(S−1\)e−κ\+2ηkB\+Se−\(κ\+ν\)\+ηkB\.\\zeta\_\{k\}\\leq\\frac\{2R\_\{k\}^\{\\rm act\}\}\{1\+R\_\{k\}^\{\\rm act\}\},\\quad R\_\{k\}^\{\\rm act\}\\leq\(S\-1\)e^\{\-\\kappa\+2\\eta\_\{k\}B\}\+Se^\{\-\(\\kappa\+\\nu\)\+\\eta\_\{k\}B\}\.\(6\)The critic uses*lookup→\\tolocal target→\\topredecessor average*\. With balanced fresh packets, lookup/aggregation leakage massesλ1,λ2\\lambda\_\{1\},\\lambda\_\{2\}andmmsamples per predecessor give
δk≤2B\(γλ1\+λ2\)\+B2log\(2SAK/α\)m\\delta\_\{k\}\\leq 2B\(\\gamma\\lambda\_\{1\}\+\\lambda\_\{2\}\)\+B\\sqrt\{\\frac\{2\\log\(2SAK/\\alpha\)\}\{m\}\}\(7\)simultaneously over the declared finite calls\. Resource certificates precede weight choice; the frozen decoder is then reused across admissible MDPs and histories\.
###### Theorem 1\(Finite\-horizon causal realization\)\.
Fix finiteS,AS,A, reward and horizon caps, deterministic step/packet caps, routing tolerances, and confidence1−α1\-\\alpha\. One fixed causal\-softmax decoder can be selected before the MDP and realized history so that repeated actor/environment/critic calls realize inexact Q\-TD\-PMD with actor residual bounded as above and critic residual satisfying \([7](https://arxiv.org/html/2609.30500#S3.E7)\) on the declared common event\.
## 4Returned\-policy control and normalized compilation
Letgk\(s\)=maxaQk\(s,a\)−⟨π^k\+1\(⋅∣s\),Qk\(s,⋅\)⟩g\_\{k\}\(s\)=\\max\_\{a\}Q\_\{k\}\(s,a\)\-\\langle\\widehat\{\\pi\}\_\{k\+1\}\(\\cdot\\mid s\),Q\_\{k\}\(s,\\cdot\)\\rangleandg¯k\\bar\{g\}\_\{k\}be its transition\-lifted maximum\. Bellman contraction givesEk\+1≤γEk\+δk\+γg¯kE\_\{k\+1\}\\leq\\gamma E\_\{k\}\+\\delta\_\{k\}\+\\gamma\\bar\{g\}\_\{k\}forEk=∥Q⋆−Qk∥∞E\_\{k\}=\\lVert Q^\{\\star\}\-Q\_\{k\}\\rVert\_\{\\infty\}\.
###### Theorem 2\(Returned\-policy error interface\)\.
For everyT≥1T\\geq 1,
∥Q⋆−Qπ^T∥∞≤\\displaystyle\\lVert Q^\{\\star\}\-Q^\{\\widehat\{\\pi\}\_\{T\}\}\\rVert\_\{\\infty\}\\leq\{\}2γTE01−γ\+21−γ∑k=0T−2γT−1−kδk\\displaystyle\\frac\{2\\gamma^\{T\}E\_\{0\}\}\{1\-\\gamma\}\+\\frac\{2\}\{1\-\\gamma\}\\sum\_\{k=0\}^\{T\-2\}\\gamma^\{T\-1\-k\}\\delta\_\{k\}\(8\)\+21−γ∑k=0T−2γT−kg¯k\+γ1−γg¯T−1\.\\displaystyle\+\\frac\{2\}\{1\-\\gamma\}\\sum\_\{k=0\}^\{T\-2\}\\gamma^\{T\-k\}\\bar\{g\}\_\{k\}\+\\frac\{\\gamma\}\{1\-\\gamma\}\\bar\{g\}\_\{T\-1\}\.\(9\)NoδT−1\\delta\_\{T\-1\}appears becauseπ^T\\widehat\{\\pi\}\_\{T\}is formed fromQT−1Q\_\{T\-1\}before an unused final critic backup\.
PMD optimality also givesg¯k≤D¯k/ηk\+Bζk\\bar\{g\}\_\{k\}\\leq\\bar\{D\}\_\{k\}/\\eta\_\{k\}\+B\\zeta\_\{k\}, whereD¯k\\bar\{D\}\_\{k\}is the greedy\-set prior\-mass price\. This identifies why low one\-step actor error does not by itself certify repeated control\.
A specially engineered normalized construction adds a two\-coordinate affine carrier\. ForEmbed\(x\)=Cc\+Jx\\operatorname\{Embed\}\(x\)=Cc\+Jxand an appropriate LayerNorm gain,J⊤LN\(Cc\+Jx\)=ρH\(x\)xJ^\{\\top\}\\operatorname\{LN\}\(Cc\+Jx\)=\\rho\_\{H\}\(x\)xwithρH\(x\)=\(1\+∥x∥2/H2\)−1/2\\rho\_\{H\}\(x\)=\(1\+\\lVert x\\rVert^\{2\}/H^\{2\}\)^\{\-1/2\}, so sufficiently large finiteHHmakes normalization a controlled perturbation\. The actor readout explicitly mixes with uniform,p^=\(1−φ\)p\+φ𝟏/A\\widehat\{p\}=\(1\-\\varphi\)p\+\\varphi\\mathbf\{1\}/A, giving a design\-time floor and∥p^−p∥1≤2φ\(A−1\)/A\\lVert\\widehat\{p\}\-p\\rVert\_\{1\}\\leq 2\\varphi\(A\-1\)/A\. With constantη=log\(A/φ\)/θ\\eta=\\log\(A/\\varphi\)/\\theta, the returned policy obeys a geometric transient plus residual neighborhood
∥Q⋆−Qπ^T∥∞≤\(2E0\+4B\)γT1−γ\+2γ\(1−γ\)2\(δ\+Bζ\+θ\)\.\\lVert Q^\{\\star\}\-Q^\{\\widehat\{\\pi\}\_\{T\}\}\\rVert\_\{\\infty\}\\leq\\frac\{\(2E\_\{0\}\+4B\)\\gamma^\{T\}\}\{1\-\\gamma\}\+\\frac\{2\\gamma\}\{\(1\-\\gamma\)^\{2\}\}\(\\delta\+B\\zeta\+\\theta\)\.\(10\)This is a finite\-cap real\-arithmetic existence result on an engineered slice, not a raw\-parameter training theorem\.
### 4\.1Why the returned\-policy indexing matters
Theorem[2](https://arxiv.org/html/2609.30500#Thmtheorem2)is a last\-iterate statement rather than a bound on an auxiliary critic\. Unrolling the critic recursion only throughQT−1Q\_\{T\-1\}gives
ET−1≤γT−1E0\+∑k=0T−2γT−2−kδk\+∑k=0T−2γT−1−kg¯k\.E\_\{T\-1\}\\leq\\gamma^\{T\-1\}E\_\{0\}\+\\sum\_\{k=0\}^\{T\-2\}\\gamma^\{T\-2\-k\}\\delta\_\{k\}\+\\sum\_\{k=0\}^\{T\-2\}\\gamma^\{T\-1\-k\}\\bar\{g\}\_\{k\}\.\(11\)Forπ=π^T\\pi=\\widehat\{\\pi\}\_\{T\}, the Bellman resolvent satisfies∥Q⋆−Qπ∥∞≤2γET−1/\(1−γ\)\+γg¯T−1/\(1−γ\)\\lVert Q^\{\\star\}\-Q^\{\\pi\}\\rVert\_\{\\infty\}\\leq 2\\gamma E\_\{T\-1\}/\(1\-\\gamma\)\+\\gamma\\bar\{g\}\_\{T\-1\}/\(1\-\\gamma\)\. The policy is already determined at this point\. ComputingQTQ\_\{T\}could be useful for another actor call, but it cannot retroactively changeπ^T\\widehat\{\\pi\}\_\{T\}; chargingδT−1\\delta\_\{T\-1\}would therefore measure work not used by the returned policy\. This causal accounting is also the reason the empirical certificate audit evaluates realized actor/critic sequences at the horizon actually returned\.
### 4\.2Support, greediness, and the role of the explicit mixture
ForGk\(s\)=argmaxaQk\(s,a\)G\_\{k\}\(s\)=\\arg\\max\_\{a\}Q\_\{k\}\(s,a\)defineDk\(s\)=−logπ^k\(Gk\(s\)∣s\)D\_\{k\}\(s\)=\-\\log\\widehat\{\\pi\}\_\{k\}\(G\_\{k\}\(s\)\\mid s\)\. Comparing PMD with the current policy restricted toGk\(s\)G\_\{k\}\(s\)yields
maxaQk\(s,a\)−⟨PMDηk\(π^k,Qk\),Qk\(s,⋅\)⟩≤Dk\(s\)/ηk\.\\max\_\{a\}Q\_\{k\}\(s,a\)\-\\langle\\operatorname\{PMD\}\_\{\\eta\_\{k\}\}\(\\widehat\{\\pi\}\_\{k\},Q\_\{k\}\),Q\_\{k\}\(s,\\cdot\)\\rangle\\leq D\_\{k\}\(s\)/\\eta\_\{k\}\.\(12\)Replacing the exact PMD row by the implemented actor costs at mostBζkB\\zeta\_\{k\}\. Thus actor fidelity becomes control\-relevant only together with a prior\-mass/greedification condition\. The explicit uniform mixture in the normalized construction gives the design\-time support floorp^a≥φ/A\\widehat\{p\}\_\{a\}\\geq\\varphi/A, henceDk≤log\(A/φ\)D\_\{k\}\\leq\\log\(A/\\varphi\)after the first call\. This closes the finite\-horizon control statement with a constant step, but it changes the executed controller and can impose a performance cost when the unmixed actor would otherwise concentrate more strongly\. Section 7 reports that cost empirically rather than treating the floor as automatically beneficial\.
## 5What the learned model learns
For a predicted rowpp, defineℓprox\(p\)=DKL\(p∥π\)−η⟨p,Q⟩\\ell\_\{\\rm prox\}\(p\)=D\_\{\\mathrm\{KL\}\}\(p\\\|\\pi\)\-\\eta\\langle p,Q\\rangle\. Ifq=PMDη\(π,Q\)q=\\operatorname\{PMD\}\_\{\\eta\}\(\\pi,Q\), then
ℓprox\(p\)−ℓprox\(q\)=DKL\(p∥q\)\.\\ell\_\{\\rm prox\}\(p\)\-\\ell\_\{\\rm prox\}\(q\)=D\_\{\\mathrm\{KL\}\}\(p\\\|q\)\.\(13\)After removing the action\-constant softmax gauge, a realizable centered linear effective\-logit model has population gradient∇R\(ϑ\)=𝔼\[Φ⊤H\(qϑ\)Φ\]\(ϑ−ϑ⋆\)\\nabla R\(\\vartheta\)=\\mathbb\{E\}\[\\Phi^\{\\top\}H\(q\_\{\\vartheta\}\)\\Phi\]\(\\vartheta\-\\vartheta^\{\\star\}\), withH\(q\)=diag\(q\)−qq⊤H\(q\)=\\operatorname\{diag\}\(q\)\-qq^\{\\top\}\. Bounded features and task richness therefore give convergence in these effective coordinates for gradient flow and sufficiently small\-step GD\. This explains objective selection, not optimization of the full deep Transformer\.
Figure 2:The learned actor exhibits the predicted PMD signature before closed\-loop testing\.\(a\) On prior held\-out one\-step inputs, PMD has mean row\-L1L\_\{1\}error0\.03570\.0357, versus0\.11160\.1116for the nearest fixed alternative\. \(b\) Refitted effective step size changes monotonically with suppliedη\\eta, with shrinkage toward the training\-range center\. These are local mechanism measurements, not returned\-policy performance\.The evaluated model is a separate four\-layer/four\-head pre\-LN Transformer \(width 64, FFN 128, zero dropout, no positional embeddings\)\. Earlier held\-out tests report PMD row\-L1=0\.035741L\_\{1\}=0\.035741, fitted effectiveη\\etamean0\.7836670\.783667, and critic TD MAE0\.1212340\.121234\. Figure[2](https://arxiv.org/html/2609.30500#S5.F2)shows that PMD is substantially closer than additive\-projected, identity, Boltzmann\-QQ, or reward\-only fixed rules\. The construction and learned network are connected through the target computation and measured residuals, not by parameter identity\.
### 5\.1Learned architecture and training interface
The learned system uses a conventional four\-layer, four\-head pre\-LN Transformer withdmodel=64d\_\{\\rm model\}=64, feed\-forward width 128, zero dropout, and no positional embeddings\. Actor examples provide structured\(π,Q,η\)\(\\pi,Q,\\eta\)fields; critic examples provide the fields required for the one\-step Bellman target\. The PMD\-aligned actor is trained through the proximal objective above rather than direct PMD probability labels\. In the freshS=4S=4block, the training range isη∈\[0\.4,1\.2\]\\eta\\in\[0\.4,1\.2\]with nominal0\.80\.8, 24 training MDPs and 2,048 examples per run, AdamW at3×10−43\\times 10^\{\-4\}, weight decay10−410^\{\-4\}, batch 64, and 51,200 optimizer steps\. Fitted reference steps on the disjoint calibration fixture are0\.800,0\.795,0\.800,0\.795,0\.8050\.800,0\.795,0\.800,0\.795,0\.805, so reference migration is negligible at the primary budget\.
The theory and learned experiment deliberately share an*interface*, not parameter values\. The constructive decoder has sparse role\-specific weights and theorem\-scale margins; the learned model has ordinary dense trainable layers\. We therefore test the target computation at three levels: same\-context actor deviationζ\\zeta, critic deviationδ\\delta, and returned\-policy loss after adaptive composition\. This avoids the stronger and unsupported claim that optimization recovers the constructed sparse circuit\.
## 6Experimental Design and Estimands
The empirical program is organized to test implications of the theory rather than to infer the mechanism from reward alone\. Table[2](https://arxiv.org/html/2609.30500#S6.T2)summarizes the ladder\. The local actor diagnostic uses the*same*\(π^k,Qk\)\(\\widehat\{\\pi\}\_\{k\},Q\_\{k\}\)context for the learned prediction and exact PMD target\. The repeated\-control experiments then allow the learned and reference trajectories to diverge and score the policies each controller actually returns\. Finally, the literature benchmark retains the same exact scoring metric but changes the learned method and, for the external adaptations, the information interface\.
Table 2:Claim\-to\-estimand map\. Each row answers a different scientific question; evidence on one row is not substituted for another\.#### Returned\-policy loss\.
For each evaluated MDP we compute exactQπQ^\{\\pi\}for the controller’s returned policy and exactQ⋆Q^\{\\star\}, then normalize the sup\-norm Q\-loss by the common initial\-policy gap\. This produces a dimensionless remaining\-loss coordinate shared across methods\. Exact PMD is run on the same MDPs as an algorithmic reference trajectory\. It is not counted as a trained competitor because its role is to expose the target algorithm’s achievable trajectory under model\-based exact updates\. Every learned method is scored by the same exact evaluator even when its own update rule observes only sampled transitions\.
#### The learned/reference ratio\.
For the fresh confirmatory experiment, each trained run is summarized atT=20T=20by
rj=medianmLj,m\(learned\)medianmLm\(ExactPMDRef\.\)\.r\_\{j\}=\\frac\{\\operatorname\{median\}\_\{m\}L\_\{j,m\}\(\\mathrm\{learned\}\)\}\{\\operatorname\{median\}\_\{m\}L\_\{m\}\(\\mathrm\{Exact\\ PMD\\ Ref\.\}\)\}\.\(14\)A ratio below one has lower task loss than the reference on that run; one is parity; one to1\.51\.5is worse than the oracle trajectory but remains inside the preregistered usefulness margin\. This threshold is an operational reliability criterion fixed before the fresh evidence, not a statistical equivalence margin\.
#### Certificate audit\.
The returned\-policy inequality is also recomputed from realized residual sequences at horizons\{1,2,5,10,20,40\}\\\{1,2,5,10,20,40\\\}\. Across the fresh confirmation, breadth, and reward\-only evaluations, all 69,120 audited rows satisfy the inequality\. The median slack is large \(about 10\.77\), so the audit establishes consistency of indexing and measured residual interfaces; it is not used as a calibrated numerical predictor of task loss\.
## 7Closed\-Loop Reliability and Scaling
The fresh confirmation fixed a 51,200\-step budget and fresh seeds before execution\. The primary comparison replaces only the exact PMD actor with the learned actor while retaining the same exact one\-step critic\. AtT=20T=20the five learned/reference ratios are0\.978,1\.396,1\.237,0\.918,1\.0520\.978,1\.396,1\.237,0\.918,1\.052\(median1\.0521\.052\); all five meet the preregistered≤1\.5\\leq 1\.5criterion\. A hierarchical learned\-actor\-minus\-reference difference has median0\.0001770\.000177with 95% interval\[−0\.000137,0\.000748\]\[\-0\.000137,0\.000748\]: this interval is a different estimand from the predeclared ratio criterion and is not an equivalence test\. With the same trained checkpoints and their learned critic, the fully learned loop has descriptive median oracle ratio1\.0501\.050and full\-loop\-minus\-oracle median0\.0005180\.000518with interval\[0\.000008,0\.002155\]\[0\.000008,0\.002155\]; no1\.51\.5margin was registered for this comparison\.
Figure 3:Local PMD correspondence remains useful under repeated deployment\.\(a\) With an exact one\-step critic, all five confirmatoryS=4S=4runs stay within the preregistered1\.5×1\.5\\timesmargin to the Exact PMD oracle\. \(b\) The same trained actors satisfy the criterion across four no\-retraining shifts; whiskers show observed five\-run ranges\.Without retraining, all four declaredS=4S=4shift families pass separately: cross\-run median learned/reference ratios are0\.9270\.927,1\.0031\.003,1\.0141\.014, and0\.9970\.997\. The first three families contain 64 distinct environments each; the structured ring is one environment evaluated from 64 initial policies\. The breadth result remains positive if the structured\-ring family is omitted\. Subsequent scale experiments retrain the architecture at larger state counts; these are size\-scaling experiments, not zero\-shot cardinality generalization\.
### 7\.1Experimental units and statistical treatment
The trained model/run is the outer independent unit\. The 64 MDPs within a run are paired fixtures used to estimate each trained controller; they are not substitutes for independent training replications\. The fresh confirmation uses five trained runs\. The larger\-size literature benchmark uses three training identities per learned method at each size\. Final paired intervals resample runs first and shared MDP identities within runs second \(10,000 replicates\)\. OOD families are analyzed separately rather than pooled\. A confidence interval containing zero is reported as unresolved, not as equality or equivalence\.
The evidence chronology is also kept separate\. The original 1,600\-step study met neither predeclared positive nor negative criterion\. A subsequent step\-count sweep was exploratory, and a 51\.2k extension on reused identities was post hoc\. The decisive five\-run confirmation fixed 51\.2k steps, new model/task seeds, and all fixtures before execution; only this fresh block supports the confirmatory repeated\-control claim\. The similar post\-hoc and fresh medians \(1\.0511\.051and1\.0521\.052\) are never pooled\.
## 8Comparison with learned ICRL/control methods
The literature\-facing benchmark enters only after each external adaptation reproduces a defining behavior of its source method\.Linear Actor\-Critic Transformeris our implementation of Liang–Lai’s one\-layer linear\-attention actor–critic construction/training\.Algorithm Distillationis our causal Transformer trained on source\-algorithm histories\. AD qualified on a second preregistered qualification run after the first run failed; Liang–Lai’s actor step and AD’s source actor rate both select the largest value in the predeclared tuning grids\. The learned PMD\-target reference uses the same architecture, inputs, data, seeds, budget, critic objective, and exact one\-step critic as PolicyAttention but is trained directly on Exact PMD targets; it is calibration rather than a headline baseline\.
Figure 4:PolicyAttention has substantially lowerT=20T=20returned\-policy loss than the two qualified external adaptations\.Markers show medians atS=8,16S=8,16and small open points show trained\-run medians\. PolicyAttention uses an exact one\-step model\-based critic here, while the external adaptations receive 20 sampled transitions per round; this is a complete\-controller comparison, not an information\-matched actor comparison\.Table 3:Median normalized returned\-policy loss in the common harness \(lower is better\)\. Every learned method has three training identities\. External entries are our adaptations, not the papers’ original benchmarks\.AtT=20T=20, PolicyAttention minus Liang–Lai has 95% intervals\[−0\.486,−0\.364\]\[\-0\.486,\-0\.364\]\(S=8S=8\) and\[−0\.496,−0\.402\]\[\-0\.496,\-0\.402\]\(S=16S=16\); versus AD the intervals are\[−0\.612,−0\.391\]\[\-0\.612,\-0\.391\]and\[−0\.620,−0\.489\]\[\-0\.620,\-0\.489\]\. Thus PolicyAttention achieves about1818–28×28\\timeslower median loss than the two qualified external adaptations at the tested settings\. This is a common\-harness separation, not a claim about the original papers and not intrinsic dominance under matched information\.
### 8\.1Common\-harness interfaces
The learned methods receive different internal information, so the benchmark standardizes task and scoring rather than update information\. PolicyAttention and the learned PMD\-target reference receive the structured\(π,Q,η\)\(\\pi,Q,\\eta\)packet and apply one exact model\-based Bellman backup after each returned actor policy in the headline controller\. Liang–Lai receives a 20\-transition on\-policy window per round and updates its actor/critic state\. AD receives the agent’s sampled\(s,a,r\)\(s,a,r\)history, also 20 new transitions per round, with a sliding 200\-token context after qualification\. Therefore theT=20T=20comparison gives each external adaptation 400 sampled transitions and gives PolicyAttention model\-based one\-step evaluation\. The resulting gap is a complete\-controller comparison and cannot be attributed solely to softmax PMD geometry\. AtS=8S=8, replacing PolicyAttention’s exact critic by its learned critic gives medianT=20T=20loss0\.02250\.0225\(1\.17×1\.17\\timesthe exact\-critic value\); Liang–Lai and Algorithm Distillation remain20\.2×20\.2\\timesand24\.2×24\.2\\timeshigher\-loss\. This is a one\-sided sampled\-critic bound, not a matched\-information test: the learned critic uses 144 generative transitions per round, versus 20 on\-policy transitions per round for each external adaptation\.
Table 4:Method interface and training context atS=8S=8\. “Exact backup” is one Bellman policy\-evaluation step with model access, not full evaluation toQπQ^\{\\pi\}\. The learned\-critic PolicyAttention check uses 144 generative transitions per round and is reported separately in text/Table[3](https://arxiv.org/html/2609.30500#S8.T3)\.
### 8\.2Baseline Qualification and Resources
The external methods entered the benchmark only after method\-specific qualification\. For Liang–Lai, the implemented linear\-attention block reproduces its analytic actor–critic teacher update to about10−510^\{\-5\}relative error at the harness shape, and native\-shape training tracks the teacher curve\. Its actor stepα=2\.0\\alpha=2\.0is selected on a disjoint calibration fixture and lies at the upper edge of the predeclared grid\. AD’s first registered qualification run failed to improve sufficiently\. A single registered repair changed context 800→\\to200, batch 32→\\to128, and training 20k→\\to40k steps; the second attempt improved median normalized loss from0\.8210\.821to0\.3950\.395over 40 rounds and qualified\. Its source actor rate 4\.0 is also the largest predeclared grid value\. These choices are disclosed because extending either grid could improve the baselines\.
Resources are not matched\. AtS=8S=8, PolicyAttention / PMD\-target reference have 272,451 parameters and train on 24 MDPs/2,048 examples for 51,200 steps \(about 0\.19–0\.24 GPU\-h per run\)\. Liang–Lai has 10,952 parameters but trains over 10,000 MDPs and10710^\{7\}windows \(about 17 CPU\-min per run\)\. AD has about 1\.14M parameters and distills 2,000 histories of 1,000 source steps \(about 0\.63–0\.75 GPU\-h per run\)\. The benchmark therefore measures returned\-policy behavior under a common task/scoring harness, not a compute\-normalized efficiency frontier\.
The learned PMD\-target reference is a calibration control and tells a different story\. PolicyAttention and PMD\-target reference are unresolved atS=8S=8for both horizons and atS=16,T=20S=16,T=20\. AtS=16,T=40S=16,T=40, PMD\-target reference is resolved lower: PolicyAttention–PMD\-target reference=\+0\.0080=\+0\.0080, 95% interval\[\+0\.0007,\+0\.0159\]\[\+0\.0007,\+0\.0159\]\. We therefore make no objective\-superiority claim over the learned PMD\-target reference\.
### 8\.3Qualification of External Adaptations
Because neither external paper publishes this common benchmark, we first require each adaptation to reproduce a defining behavior of its source method\. For Liang–Lai, the explicit linear\-attention parameter block is unit\-tested against the analytic semi\-gradient actor–critic update, and the trained block tracks the same teacher on nativeS=9,A=4,γ=\.5S=9,A=4,\\gamma=\.5prompts before it is moved to theS=8S=8harness\. The native trained model reaches relative update error3\.2×10−63\.2\\times 10^\{\-6\}and its closed\-loop improvement tracks the analytic teacher at 3,000 updates\. On the harness shape the update error remains around1\.5×10−51\.5\\times 10^\{\-5\}\. The common benchmark then uses the calibration\-selected\(α,β\)=\(2\.0,0\.2\)\(\\alpha,\\beta\)=\(2\.0,0\.2\)atS=8S=8and\(2\.0,0\.8\)\(2\.0,0\.8\)atS=16S=16; the larger actor step is important because the paper\-default values adapt much more slowly at the 20\-round common horizon\.
AD is qualified differently because its defining phenomenon is improvement from learning\-history context rather than imitation of a closed\-form update\. The final long\-context model improves median normalized loss from0\.8210\.821at round 1 to0\.3950\.395at round 40 on the disjoint qualification fixture, with 64/64 MDPs improving; a short\-context control improves substantially less\. The first registered training run failed its qualification criterion and remains part of the audit record\. The second attempt was fixed before rerunning and is the only AD model family admitted to the benchmark\. AtS=16S=16the recipe is retrained rather than independently re\-qualified from scratch; the evaluation trajectory still shows in\-context improvement from0\.9230\.923at round 1 to0\.3730\.373at round 40\. These qualifications support the phrase “our faithful adaptations”; they do not convert our benchmark into a reproduction of either publication’s reported numbers\.
### 8\.4PMD\-target calibration
Figure[5](https://arxiv.org/html/2609.30500#S8.F5)\(a\) places theS=8S=8andS=16S=16endpoints on a common fidelity–performance plane\. PolicyAttention occupies the low\-deviation, low\-loss region; the Liang–Lai and Algorithm Distillation points form a much higher\-loss external comparison envelope\. The learned PMD\-target reference is shown only as calibration\. Figure[5](https://arxiv.org/html/2609.30500#S8.F5)\(b\) normalizesT=20T=20loss by that learned reference: PolicyAttention is1\.02×1\.02\\timesand1\.18×1\.18\\timesthe reference atS=8S=8andS=16S=16, while the external ratios inherit the information\-access difference of the common harness\.
Figure 5:PolicyAttention occupies the low\-loss, low\-deviation region\.\(a\) Same\-context row\-L1L\_\{1\}deviation from Exact PMD versusT=20T=20returned\-policy loss atS=8S=8andS=16S=16; the dashed curve is the non\-dominated envelope over the two external adaptations\. \(b\) Common\-harnessT=20T=20loss normalized by the learned PMD\-target reference\.
## 9Scope and Boundaries
Four boundaries constrain interpretation\.First, PMD\-target reference remains a strong matched control, and one stress\-horizon cell favors it\.Second, the literature benchmark is not information\-matched: PolicyAttention / PMD\-target reference use exact one\-step model\-based backups and zero sampled transitions, while Liang–Lai and AD receive 20 sampled transitions per round; the latter are also plausibly under\-tuned because selected rates lie at the edges of the predeclared grids\.Third, larger sizes use retraining, not zero\-shot state\-cardinality extrapolation\.Fourth, the theory is an implementation/control result under structured semantic tokens, finite caps, and generative critic coverage; it does not prove SGD/AdamW recovery of the constructed circuit\.
Two additional observations sharpen interpretation\. A one\-state counterexample shows uniformly tiny one\-step PMD error can coexist with permanent suboptimal lock\-up under bounded steps, motivating the closed\-loop test\. Separately, a paired reward\-only actor achieves lowerS=4,T=20S=4,T=20task loss than the PMD\-aligned actor, and the explicitφ=\.01\\varphi=\.01support mixture costs\+0\.0171\+0\.0171with 95% interval\[0\.0163,0\.0192\]\[0\.0163,0\.0192\]inS=4,T=20S=4,T=20loss in the well\-trained regime\. These do not negate the PMD mechanism result; they show why algorithmic fidelity and task optimization are different axes\.
## 10Discussion
#### What the literature\-facing advantage establishes\.
The literature\-facing comparison is intentionally performance\-first: after qualification, every PolicyAttention run has much lowerT=20T=20median returned\-policy loss than every Liang–Lai and AD run at both tested sizes, and the hierarchical intervals exclude zero by wide margins\. The ratio view in Fig\.[4](https://arxiv.org/html/2609.30500#S8.F4)makes the magnitude visible without comparing to the oracle: the external adaptations incur17\.717\.7–28\.2×28\.2\\timesPolicyAttention’s loss\. The result survives the increase from eight to sixteen states after retraining\. This is the strongest empirical separation in the current project and warrants headline status\. Its causal interpretation is narrower\. Because the methods expose different information to their update rules, the result establishes that the complete PolicyAttention controller is substantially more effective under this registered common harness; it does not isolate which fraction of the gap comes from actor geometry, exact model\-based critic access, optimization, context length, or capacity\.
#### Why the learned PMD\-target calibration is complementary\.
The learned PMD\-target reference removes most of those differences\. It uses the same Transformer architecture, actor inputs, training data, seeds, budgets, and critic objective as PolicyAttention, changing the actor supervision from the proximal variational objective to direct exact\-PMD targets\. The resulting frontier is mixed rather than one\-sided\. AtS=8S=8, the two objectives are unresolved at both horizons and interleave across checkpoints\. AtS=16,T=20S=16,T=20, PMD\-target reference has a lower point estimate but the interval still crosses zero; atT=40T=40its advantage is resolved\. This matched result prevents us from explaining the external\-baseline gap as evidence that the proximal objective is intrinsically better than direct imitation\. Instead, it supports a mechanism\-centered reading: PolicyAttention learns the target computation without direct policy labels and reaches a competitive fidelity/control frontier, while direct supervision can be at least as effective and can win at a longer horizon\.
#### Why local PMD fidelity is still worth measuring\.
Task loss alone cannot identify an update rule\. The paired reward\-only control demonstrates the point constructively: it attains lower primary\-horizon task loss than the PMD\-aligned actor while the registered experiment contains no same\-context PMD\-fidelity coordinate for that reward\-only controller\. Conversely, the one\-state lock\-up construction shows that an actor may be arbitrarily close to PMD in one step yet fail under repeated bounded\-step control\. The combination explains the paper’s three\-axis evaluation\. Same\-contextζ\\zetaasks which algorithm the actor resembles locally; returned\-policy loss asks whether the deployed controller remains useful after feedback; alternative objectives and literature baselines ask how much task performance depends on that mechanism\. None of these coordinates is redundant\.
#### What the scale evidence says\.
TheS=8S=8andS=16S=16benchmark blocks retrain every method at the new problem size\. They therefore test whether the phenomenon persists when the interface is rebuilt at larger tabular state count, not whether one fixed network extrapolates to an unseen cardinality\. PolicyAttention’s raw median loss grows by1\.41×1\.41\\timesfromS=8S=8toS=16S=16atT=20T=20, while the Exact PMD oracle grows by1\.13×1\.13\\times; its ratio to the oracle therefore changes from1\.311\.31to1\.641\.64\. The external adaptations remain far from convergence at the common 20\-round horizon, so their raw losses change little across these two sizes\. These point ratios are descriptive with three trained runs per method; no uncertainty ordering is inferred for the growth factors\.
## 11Conclusion
PolicyAttention turns the softmax geometry already present in PMD into a causal in\-context control mechanism\. A fixed decoder realizes the actor–environment–one\-step\-critic computation with explicit implementation residuals, and the returned\-policy theorem connects those residuals to the policy the agent actually emits\. Separately trained Transformers exhibit the PMD mechanism locally, preserve useful control under repeated adaptive composition and environment shifts, and, under the common harness, operate in a much lower\-loss regime than two qualified published\-method adaptations\. The decisive distinction is therefore between*algorithmic fidelity*,*closed\-loop reliability*, and*task performance*: none can be inferred from the other two\. PolicyAttention supplies a common mechanism and measurement interface for studying all three\.
## Appendix AComplete returned\-policy proof
This appendix records the complete error propagation used in the main theorem\. SetEk=‖Q⋆−Qk‖∞E\_\{k\}=\\\|Q^\{\\star\}\-Q\_\{k\}\\\|\_\{\\infty\}\. Decompose
Q⋆−Qk\+1\\displaystyle Q^\{\\star\}\-Q\_\{k\+1\}=\(ℱQ⋆−ℱQk\)\+\(ℱQk−ℱπ^k\+1Qk\)\\displaystyle=\(\\mathcal\{F\}Q^\{\\star\}\-\\mathcal\{F\}Q\_\{k\}\)\+\(\\mathcal\{F\}Q\_\{k\}\-\\mathcal\{F\}^\{\\widehat\{\\pi\}\_\{k\+1\}\}Q\_\{k\}\)\(15\)−\(Qk\+1−ℱπ^k\+1Qk\)\.\\displaystyle\\qquad\-\(Q\_\{k\+1\}\-\\mathcal\{F\}^\{\\widehat\{\\pi\}\_\{k\+1\}\}Q\_\{k\}\)\.\(16\)The first term is bounded byγEk\\gamma E\_\{k\}\. The second equalsγPgk\\gamma Pg\_\{k\}and has sup norm at mostγg¯k\\gamma\\bar\{g\}\_\{k\}; the third is bounded byδk\\delta\_\{k\}\. Thus
Ek\+1≤γEk\+δk\+γg¯k\.E\_\{k\+1\}\\leq\\gamma E\_\{k\}\+\\delta\_\{k\}\+\\gamma\\bar\{g\}\_\{k\}\.\(17\)Unrolling toT−1T\-1gives
ET−1≤γT−1E0\+∑k=0T−2γT−2−kδk\+∑k=0T−2γT−1−kg¯k\.E\_\{T\-1\}\\leq\\gamma^\{T\-1\}E\_\{0\}\+\\sum\_\{k=0\}^\{T\-2\}\\gamma^\{T\-2\-k\}\\delta\_\{k\}\+\\sum\_\{k=0\}^\{T\-2\}\\gamma^\{T\-1\-k\}\\bar\{g\}\_\{k\}\.\(18\)Forπ=π^T\\pi=\\widehat\{\\pi\}\_\{T\}, the Bellman fixed\-point equations give
Q⋆−Qπ=\(I−γPπ\)−1\(ℱQ⋆−ℱπQ⋆\)\.Q^\{\\star\}\-Q^\{\\pi\}=\(I\-\\gamma P^\{\\pi\}\)^\{\-1\}\(\\mathcal\{F\}Q^\{\\star\}\-\\mathcal\{F\}^\{\\pi\}Q^\{\\star\}\)\.\(19\)The inverse is nonnegative with row sums1/\(1−γ\)1/\(1\-\\gamma\)\. At each next state,
maxaQ⋆\(s,a\)−⟨π\(⋅∣s\),Q⋆\(s,⋅\)⟩≤g¯T−1\+2ET−1,\\max\_\{a\}Q^\{\\star\}\(s,a\)\-\\langle\\pi\(\\cdot\\mid s\),Q^\{\\star\}\(s,\\cdot\)\\rangle\\leq\\bar\{g\}\_\{T\-1\}\+2E\_\{T\-1\},\(20\)so
‖Q⋆−Qπ^T‖∞≤2γ1−γET−1\+γ1−γg¯T−1\.\\\|Q^\{\\star\}\-Q^\{\\widehat\{\\pi\}\_\{T\}\}\\\|\_\{\\infty\}\\leq\\frac\{2\\gamma\}\{1\-\\gamma\}E\_\{T\-1\}\+\\frac\{\\gamma\}\{1\-\\gamma\}\\bar\{g\}\_\{T\-1\}\.\(21\)Substitution yields Eq\. \([9](https://arxiv.org/html/2609.30500#S4.E9)\)\. The proof stops beforeQTQ\_\{T\}is constructed, which is whyδT−1\\delta\_\{T\-1\}is absent\.
### A\.1From PMD residual to greedy defect
LetGk\(s\)=argmaxaQk\(s,a\)G\_\{k\}\(s\)=\\arg\\max\_\{a\}Q\_\{k\}\(s,a\)and letp∘p^\{\\circ\}be the current policy restricted toGk\(s\)G\_\{k\}\(s\)and renormalized\. ThenDKL\(p∘∥π^k\)=−logπ^k\(Gk\(s\)∣s\)=Dk\(s\)D\_\{\\mathrm\{KL\}\}\(p^\{\\circ\}\\\|\\widehat\{\\pi\}\_\{k\}\)=\-\\log\\widehat\{\\pi\}\_\{k\}\(G\_\{k\}\(s\)\\mid s\)=D\_\{k\}\(s\)\. PMD optimality implies
maxaQk\(s,a\)−⟨qk,Qk\(s,⋅\)⟩≤Dk\(s\)/ηk\.\\max\_\{a\}Q\_\{k\}\(s,a\)\-\\langle q\_\{k\},Q\_\{k\}\(s,\\cdot\)\\rangle\\leq D\_\{k\}\(s\)/\\eta\_\{k\}\.\(22\)For probability rowsp,qp,qand‖Q‖∞≤B\\\|Q\\\|\_\{\\infty\}\\leq B, centeringQQby the midpoint of its range gives\|⟨p−q,Q⟩\|≤B‖p−q‖1\|\\langle p\-q,Q\\rangle\|\\leq B\\\|p\-q\\\|\_\{1\}\. Hence
gk\(s\)≤Dk\(s\)/ηk\+Bζk\.g\_\{k\}\(s\)\\leq D\_\{k\}\(s\)/\\eta\_\{k\}\+B\\zeta\_\{k\}\.\(23\)This is the exact interface used by the mixture\-floor closure\.
## Appendix BRouting and sampling certificates
The actor partitions the visible memories into the intended state/action group and competitors\. IfZgoodZ\_\{\\rm good\}andZbadZ\_\{\\rm bad\}are their exponential score sums andR=Zbad/ZgoodR=Z\_\{\\rm bad\}/Z\_\{\\rm good\}, the bad attention mass isR/\(1\+R\)R/\(1\+R\)\. Probability\-valued outputs haveL1L\_\{1\}diameter at most two, givingζ≤2R/\(1\+R\)\\zeta\\leq 2R/\(1\+R\)\. For score marginsκ,ν\\kappa,\\nu,‖Q‖∞≤B\\\|Q\\\|\_\{\\infty\}\\leq B, and actor stepη\\eta, the resulting count/score envelope is
Ract≤\(S−1\)e−κ\+2ηB\+Se−\(κ\+ν\)\+ηB\.R\_\{\\rm act\}\\leq\(S\-1\)e^\{\-\\kappa\+2\\eta B\}\+Se^\{\-\(\\kappa\+\\nu\)\+\\eta B\}\.\(24\)For a target0<ϵr<20<\\epsilon\_\{r\}<2, letrϵ=ϵr/\(2−ϵr\)r\_\{\\epsilon\}=\\epsilon\_\{r\}/\(2\-\\epsilon\_\{r\}\)\. A convenient sufficient choice isκ=2ηmaxB\+log\(2S/rϵ\)\\kappa=2\\eta\_\{\\max\}B\+\\log\(2S/r\_\{\\epsilon\}\)andν=1\\nu=1\.
The critic has two finite\-temperature routing stages\. WithMMvisible transition tokens and at leastmmexamples per predecessor,
R1\\displaystyle R\_\{1\}≤\(S\+A−2\)e−τ1\+\(S−1\)\(A−1\)e−2τ1\+Me−3τ1,\\displaystyle\\leq\(S\+A\-2\)e^\{\-\\tau\_\{1\}\}\+\(S\-1\)\(A\-1\)e^\{\-2\\tau\_\{1\}\}\+Me^\{\-3\\tau\_\{1\}\},\(25\)R2\\displaystyle R\_\{2\}≤\(M/m\)e−τ2\+\(2SA/m\)e−3τ2,\\displaystyle\\leq\(M/m\)e^\{\-\\tau\_\{2\}\}\+\(2SA/m\)e^\{\-3\\tau\_\{2\}\},\(26\)andδroute≤2B\(γλ1\+λ2\)\\delta\_\{\\rm route\}\\leq 2B\(\\gamma\\lambda\_\{1\}\+\\lambda\_\{2\}\)withλj=Rj/\(1\+Rj\)\\lambda\_\{j\}=R\_\{j\}/\(1\+R\_\{j\}\)\. Conditional on the realized actor and past, fresh TD targets lie in\[−B,B\]\[\-B,B\]and have conditional meanℱπ^k\+1Qk\\mathcal\{F\}^\{\\widehat\{\\pi\}\_\{k\+1\}\}Q\_\{k\}\. Hoeffding plus a union bound overSAKSAKpair/call combinations yields
m≥⌈2B2ϵs2log2SAKα⌉⟹δstat≤ϵsm\\geq\\left\\lceil\\frac\{2B^\{2\}\}\{\\epsilon\_\{s\}^\{2\}\}\\log\\frac\{2SAK\}\{\\alpha\}\\right\\rceil\\quad\\Longrightarrow\\quad\\delta\_\{\\rm stat\}\\leq\\epsilon\_\{s\}\(27\)with probability at least1−α1\-\\alphaover the declared calls\. The construction therefore requires packet and step certificates before the routing weights are frozen\.
## Appendix CNormalized finite\-cap compilation
Let the normalization\-free payload have dimensionddand chooseD=d\+2D=d\+2\. LetJ∈ℝD×dJ\\in\\mathbb\{R\}^\{D\\times d\}have orthonormal columns, and choose a unit carrierccorthogonal to the image ofJJand to the all\-ones direction\. WithC2=H2−DϵLNC^\{2\}=H^\{2\}\-D\\epsilon\_\{\\rm LN\}definey=Cc\+Jxy=Cc\+Jx\. The token has mean zero and variance\(C2\+‖x‖2\)/D\(C^\{2\}\+\\\|x\\\|^\{2\}\)/D\. LayerNorm with gainH/DH/\\sqrt\{D\}and zero bias gives
J⊤LN\(Cc\+Jx\)=HxH2\+‖x‖2=ρH\(x\)x\.J^\{\\top\}\\operatorname\{LN\}\(Cc\+Jx\)=\\frac\{Hx\}\{\\sqrt\{H^\{2\}\+\\\|x\\\|^\{2\}\}\}=\\rho\_\{H\}\(x\)x\.\(28\)On any finite compact payload domain,1−ρH\(x\)=O\(H−2\)1\-\\rho\_\{H\}\(x\)=O\(H^\{\-2\}\)\. Query/key score perturbations and value perturbations are therefore uniformly small for sufficiently large finiteHH\. The final LayerNorm is handled by the same carrier identity\.
Action values are represented in centered coordinatesea−ue\_\{a\}\-uwithu=𝟏/Au=\\mathbf\{1\}/A\. The affine readout addsuu, preserving the simplex\. For control we use the explicit mixture
p^=\(1−φ\)p\+φu,\\widehat\{p\}=\(1\-\\varphi\)p\+\\varphi u,\(29\)not an intrinsic numerical floor from LayerNorm\. It guaranteesp^a≥φ/A\\widehat\{p\}\_\{a\}\\geq\\varphi/Aand adds at most2φ\(A−1\)/A2\\varphi\(A\-1\)/Arow\-L1L\_\{1\}actor bias\. With constantη=log\(A/φ\)/θ\\eta=\\log\(A/\\varphi\)/\\thetaand uniform residual capsδ,ζ\\delta,\\zeta, the theorem in the main text follows fromg¯0≤2B\\bar\{g\}\_\{0\}\\leq 2Bandg¯k≤θ\+Bζ\\bar\{g\}\_\{k\}\\leq\\theta\+B\\zetafork≥1k\\geq 1\. The construction is in exact real arithmetic\. Conservative sufficient carrier/margin scales can be impractical and are not claimed to match the learned network\.
## Appendix DEffective\-coordinate learning details
LetPA=I−𝟏𝟏⊤/AP\_\{A\}=I\-\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}/Aremove the action\-constant softmax gauge\. In a centered action\-equivariant linear class, write effective logitszϑ=Φ\(x\)ϑz\_\{\\vartheta\}=\\Phi\(x\)\\varthetaand suppose the PMD target is realizable atϑ⋆\\vartheta^\{\\star\}\. Since the proximal excess equalsDKL\(qϑ∥qϑ⋆\)D\_\{\\mathrm\{KL\}\}\(q\_\{\\vartheta\}\\\|q\_\{\\vartheta^\{\\star\}\}\), differentiation gives
∇R\(ϑ\)=𝔼\[Φ⊤H\(qϑ\)Φ\]\(ϑ−ϑ⋆\),H\(q\)=diag\(q\)−qq⊤\.\\nabla R\(\\vartheta\)=\\mathbb\{E\}\[\\Phi^\{\\top\}H\(q\_\{\\vartheta\}\)\\Phi\]\(\\vartheta\-\\vartheta^\{\\star\}\),\\quad H\(q\)=\\operatorname\{diag\}\(q\)\-qq^\{\\top\}\.\(30\)On a finite coefficient ball, bounded features imply bounded logits and hence a positive lower action probability\. For centeredvv,v⊤H\(q\)v≥p0‖v‖2v^\{\\top\}H\(q\)v\\geq p\_\{0\}\\\|v\\\|^\{2\}\. If𝔼\[Φ⊤Φ\]⪰λI\\mathbb\{E\}\[\\Phi^\{\\top\}\\Phi\]\\succeq\\lambda Ion the identifiable coordinates, then
\(ϑ−ϑ⋆\)⊤∇R\(ϑ\)≥p0λ∥ϑ−ϑ⋆∥2\.\(\\vartheta\-\\vartheta^\{\\star\}\)^\{\\top\}\\nabla R\(\\vartheta\)\\geq p\_\{0\}\\lambda\\\|\\vartheta\-\\vartheta^\{\\star\}\\\|^\{2\}\.\(31\)Gradient flow contracts on the invariant ball, and sufficiently small\-step GD has the analogous geometric convergence\. This proposition is intentionally restricted to effective centered coordinates; raw deep\-network factorization and AdamW remain outside its scope\.
## Appendix EComplete empirical evidence chronology
### E\.1Fresh confirmation and alternative mechanism
The decisive fresh block uses five new trained identities at 51\.2k steps under a protocol fixed before any outcomes were observed\. AtT=20T=20the learned/reference run ratios are0\.9784130\.978413,1\.3963881\.396388,1\.2366071\.236607,0\.9180640\.918064, and1\.0522871\.052287\. All five satisfy the1\.51\.5registered margin, none systematically stalls, and all certificate/numerical audits pass\. The paired reward\-only actor produces reward\-only\-minus\-learned\-actor median−0\.005963\-0\.005963with interval\[−0\.009794,−0\.003191\]\[\-0\.009794,\-0\.003191\], and reward\-only\-minus\-full\-loop median−0\.007066\-0\.007066with interval\[−0\.009200,−0\.004539\]\[\-0\.009200,\-0\.004539\]\. These are task\-loss comparisons; registered PMD\-fidelity diagnostics for the reward\-only actor are absent\.
The four no\-retraining shift medians are0\.92700\.9270\(sparse transitions\),1\.00351\.0035\(strong mixing\),1\.01381\.0138\(sparse rewards\), and0\.99670\.9967\(structured ring\)\. The first three shift families each contain 64 independently generated environments\. The ring has one fixed\(P,r\)\(P,r\)structure and 64 initial policies\. The aggregate breadth class remains positive if the ring is omitted\.
### E\.2Matched objective and checkpoint frontiers
The matched PolicyAttention / PMD\-target reference experiment varies only the actor objective under identical architecture, actor inputs, training fixtures, critic objective, optimizer and checkpoint budgets\. At bothS=4S=4andS=8S=8, the objective comparison is mixed across checkpoints rather than uniformly ordered\. TheS=8S=8fidelity/control CSV contains five checkpoint budgets from 6\.4k to 51\.2k\. The combined lower\-is\-better non\-dominated envelope contains PolicyAttention at 38\.4k\(ζ=0\.01665,L=0\.02538\)\(\\zeta=0\.01665,L=0\.02538\)and 51\.2k\(0\.01574,0\.02703\)\(0\.01574,0\.02703\); a PMD\-target reference 51\.2k point is method\-nondominated but dominated in the combined set\. These coordinates are descriptive of the registered checkpoint landscape, not evidence that one objective is globally Pareto\-superior\.
### E\.3Literature\-facing benchmark
AtT=20T=20, cross\-run medians are0\.01929/0\.01894/0\.45345/0\.543900\.01929/0\.01894/0\.45345/0\.54390for PolicyAttention, PMD\-target reference, Liang–Lai, and AD atS=8S=8, and0\.02727/0\.02321/0\.48167/0\.572320\.02727/0\.02321/0\.48167/0\.57232atS=16S=16\. Figure[4](https://arxiv.org/html/2609.30500#S8.F4)plots the underlying three trained\-run medians for every learned method; the trained run remains the independent unit\.
AtT=40T=40, PMD\-target reference is resolved lower than PolicyAttention atS=16S=16: every PMD\-target reference run median \(0\.0105–0\.0158\) is below every PolicyAttention run median \(0\.0208–0\.0249\)\. This cell is retained precisely because the external\-baseline headline does not determine the matched\-control conclusion\.
## Appendix FExternal\-baseline qualification details
The Liang–Lai adaptation uses the paper’s one\-layer linear self\-attention form, actor–critic prompt organization, analytic teacher, and teacher\-mimicking training\. Native\-shape qualification atS=9,A=4,γ=\.5S=9,A=4,\\gamma=\.5gives last/first training\-loss ratio4\.9×10−124\.9\\times 10^\{\-12\}, median relative emitted\-update error3\.2×10−63\.2\\times 10^\{\-6\}, closed\-loop improvement ratio0\.9990\.999versus teacher at 3,000 updates, and normalized mean curve gap4\.3×10−44\.3\\times 10^\{\-4\}\. The harness\-shape functional error atS=8S=8is1\.5×10−51\.5\\times 10^\{\-5\}\. Calibration selectsα=2\.0\\alpha=2\.0, the maximum predeclared grid value, so under\-tuning remains a directional caveat\.
AD uses causal tokens\(at−1,rt−1,st\)\(a\_\{t\-1\},r\_\{t\-1\},s\_\{t\}\)and action negative log likelihood, with a side\-effect\-free returned\-policy query\. Its first registered qualification attempt \(c=800c=800, batch 32, 20k steps\) failed: round\-1/40 losses were 0\.928/0\.937\. After a single registered calibration\-only repair toc=200c=200, batch 128, 40k steps, the long\-context model improves 0\.821→\\to0\.395 with 64/64 MDPs improved and beats the short\-context control by the registered gap\. The source actor rate 4\.0 is also the maximum predeclared grid value\. We therefore use “qualified adaptation,” not “reproduction of Algorithm Distillation’s benchmark\.”
## Appendix GNegative results and boundaries
#### One\-step fidelity is insufficient\.
For any bounded step cap and targetζ\>0\\zeta\>0, a one\-state two\-action self\-loop can choose sufficiently small good\-action prior mass so that exact PMD itself moves by at mostζ\\zeta\. An implemented actor can then stay at the prior while remaining uniformlyζ\\zeta\-accurate to PMD; with an exact critic initialized atQπQ^\{\\pi\}, the loop remains at a suboptimal fixed point\. The aligned proximal excess can simultaneously vanish\. This project\-specific witness motivates, but does not itself prove, the learned closed\-loop result\.
#### Support\-floor cost\.
In the freshS=4S=4block, the fixedφ=\.01\\varphi=\.01mixture has mixture\-minus\-unmixed\-actor median0\.017110\.01711with interval\[0\.01633,0\.01919\]\[0\.01633,0\.01919\]: once the actor is well trained, the uniform floor prevents the concentration available to the unmixed actor\. The mixture is a theorem\-motivated support certificate, not a guaranteed empirical exploration benefit\.
#### Coverage and autonomy\.
The construction assumes structured semantic fields and fresh balanced generative critic packets\. It does not infer a tabular MDP from raw observations, select state\-action coverage endogenously, or support an unlimited growing transcript at fixed finite routing margins\. The learned experiments are closed loop in the sense that returned policies change subsequent inputs and targets, but they are not unrestricted autonomous agents\.
#### Size and compute\.
S=16S=16uses retraining, not zero\-shot cardinality extrapolation\. External methods have unmatched capacities and training data; PolicyAttention has stronger online model access\. Resource tables are therefore descriptive and not used to claim a compute Pareto optimum\.
## Appendix HReproducibility specification
The literature benchmark uses three trained identities per method and 64 fresh evaluation MDPs per size\.S=8S=8seeds are 18100–18102 for models and 28100–28102 for task roots;S=16S=16uses 19100–19102 and 29100–29102\. Calibration/evaluation roots are disjoint\. PolicyAttention / PMD\-target reference checkpoints use the 51,200\-step matched\-control runner; external baselines have their own qualified implementations\. ExactQπQ^\{\\pi\}andQ⋆Q^\{\\star\}are used only for scoring \(and for the exact one\-step critic of PolicyAttention / PMD\-target reference\)\. Final intervals use a hierarchical bootstrap with 10,000 replicates, resampling trained runs first and MDP identities within runs second\. All manuscript figures are generated deterministically from the compact CSV/JSON inputs shipped with the source package; figure scripts do not train models\.
The archived evidence also stores per\-round compressed CSV rows, returned\-policy arrays, fixture manifests and checkpoint hashes\. An independent verification reconstructed the run medians and decisive intervals from those raw rows and rescored stored returned policies with zero discrepancy\. The manuscript figures and tables read these fixed evidence artifacts without modifying them\.
## References
- E\. Akyürek, D\. Schuurmans, J\. Andreas, T\. Ma, and D\. ZhouWhat learning algorithm is in\-context learning? investigations with linear models\.Note:International Conference on Learning Representations \(ICLR\)Cited by:[§2](https://arxiv.org/html/2609.30500#S2.p1.1)\.
- Garget al\.\(2022\)S\. Garg, D\. Tsipras, P\. Liang, and G\. ValiantWhat can transformers learn in\-context? a case study of simple function classes\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 30583–30598\.Cited by:[§2](https://arxiv.org/html/2609.30500#S2.p1.1)\.
- Geistet al\.\(2019\)M\. Geist, B\. Scherrer, and O\. PietquinA theory of regularized markov decision processes\.InProceedings of the 36th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.97,pp\. 2160–2169\.Cited by:[§1](https://arxiv.org/html/2609.30500#S1.p2.2)\.
- Johnsonet al\.\(2023\)E\. Johnson, C\. Pike\-Burke, and P\. RebeschiniOptimal convergence rate for exact policy mirror descent in discounted markov decision processes\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 76496–76524\.Cited by:[§1](https://arxiv.org/html/2609.30500#S1.p2.2)\.
- Lan \(2023\)G\. LanPolicy mirror descent for reinforcement learning: linear convergence, new sampling complexity, and generalized problem classes\.Mathematical Programming198\(1\),pp\. 1059–1106\.External Links:[Document](https://dx.doi.org/10.1007/s10107-022-01816-5)Cited by:[§1](https://arxiv.org/html/2609.30500#S1.p2.2)\.
- Laskinet al\.\(2023\)M\. Laskin, L\. Wang, J\. Oh, E\. Parisotto, S\. Spencer, R\. Steigerwald, D\. Strouse, S\. Hansen, A\. Filos, E\. Brooks, M\. Gazeau, H\. Sahni, S\. Singh, and V\. MnihIn\-context reinforcement learning with algorithm distillation\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2609.30500#S2.p1.1)\.
- Liang and Lai \(2026\)H\. Liang and L\. LaiTransformers provably implement in\-context reinforcement learning with policy improvement\.External Links:2605\.05755Cited by:[§2](https://arxiv.org/html/2609.30500#S2.p1.1)\.
- Liuet al\.\(2025\)J\. Liu, W\. Li, and K\. WeiOn the convergence of policy mirror descent with temporal difference evaluation\.External Links:2509\.18822Cited by:[§1](https://arxiv.org/html/2609.30500#S1.p2.2)\.
- Luet al\.\(2023\)C\. Lu, Y\. Schroecker, A\. Gu, E\. Parisotto, J\. Foerster, S\. Singh, and F\. BehbahaniStructured state space models for in\-context reinforcement learning\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2](https://arxiv.org/html/2609.30500#S2.p1.1)\.
- von Oswaldet al\.\(2023\)J\. von Oswald, E\. Niklasson, E\. Randazzo, J\. Sacramento, A\. Mordvintsev, A\. Zhmoginov, and M\. VladymyrovTransformers learn in\-context by gradient descent\.InProceedings of the 40th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.202,pp\. 35151–35174\.Cited by:[§2](https://arxiv.org/html/2609.30500#S2.p1.1)\.
- Wanget al\.\(2025\)J\. Wang, E\. Blaser, H\. Daneshmand, and S\. ZhangTransformers can learn temporal difference methods for in\-context reinforcement learning\.Note:International Conference on Learning Representations \(ICLR\)Cited by:[§2](https://arxiv.org/html/2609.30500#S2.p1.1)\.
- Xiao \(2022\)L\. XiaoOn the convergence rates of policy gradient methods\.Journal of Machine Learning Research23\(282\),pp\. 1–36\.Cited by:[§1](https://arxiv.org/html/2609.30500#S1.p2.2)\.
- Xieet al\.\(2026\)Z\. Xie, X\. Liu, C\. Chen, S\. D\. Liu, R\. Chandra, and S\. ZhangBeyond linear attention: softmax transformers implement in\-context reinforcement learning\.External Links:2605\.07333Cited by:[§2](https://arxiv.org/html/2609.30500#S2.p1.1)\.Similar Articles
Soft Adaptive Policy Optimization
SAPO introduces a smooth, temperature-controlled gate to adaptively attenuate off-policy updates in reinforcement learning for large language models, enhancing training stability and performance compared to methods with hard clipping.
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
SA-MRPO introduces a saturation-aware advantage reweighting technique for multi-reward policy optimization in reinforcement learning, improving performance on mathematical reasoning, adaptive reasoning, and coding benchmarks by focusing optimization on under-optimized objectives.
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization
This paper introduces Contrastive Policy Optimization (CPO), which uses token-level contrastive disagreement between reference-guided and vanilla generation distributions for correctness-aware advantage shaping in reinforcement learning with verifiable rewards. CPO outperforms entropy-based RLVR methods on both in-domain and out-of-domain benchmarks.
Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization
The paper introduces BATON, a dual-axis policy optimization framework for LLM agents using Bayesian Feedback Attribution and Trajectory Mass Normalization, demonstrating improved performance in reinforcement learning experiments.
Offline Reinforcement Learning for Fluid Controls: Data-based Multi-observational Policy Extraction
This paper proposes a novel offline reinforcement learning framework for active flow control that uses a sensor position-conditioned architecture with Point Attention layers to handle varying sensor configurations, enabling data-driven policy extraction without costly online interactions.