Relevant and Irrelevant: A Renormalization Group Analysis of Transformer Attention
Summary
This paper applies Wilsonian renormalization group theory to analyze Transformer attention as a perturbation of the MLP residual-stack fixed point, determining whether attention is relevant or irrelevant based on data correlation length. Experiments on synthetic Markov chains confirm that attention's relevance depends on the spectral structure of the data-generating process, with the first-layer head dominating the transition.
View Cached Full Text
Cached at: 07/20/26, 09:27 AM
# A Renormalization Group Analysis of Transformer Attention
Source: [https://arxiv.org/html/2607.15449](https://arxiv.org/html/2607.15449)
## Relevant and Irrelevant: A Renormalization Group Analysis of Transformer Attention
Parviz Haggi\-Mani1,2, Irina Rish1,2 haggimpa@mila\.quebec, irina\.rish@mila\.quebec 1Université de Montréal,2Mila – Quebec AI Institute
###### Abstract
Using the language of Wilsonian renormalization group theory \(RG\), we treat the Transformer’s attention mechanism as a perturbation of the trained MLP residual\-stack fixed point and ask whether it constitutes a relevant, marginal, or irrelevant operator\. We derive a fixed\-point shift formulaδ=−M∗−1\(a\+b\)\\delta=\-\{M^\{\*\}\}^\{\-1\}\(a\+b\)and obtain four testable predictions for the fixed\-point geometry, effective rank profile, layer specificity, and perturbation decay spectrum\. Testing these on synthetic Markov chain sequences with controlled correlation lengthξ\\xi, we find: \(1\) For long\-ξ\\xichains, attention is strongly*relevant*: it closes a residual loss gap the MLP cannot bridge and drives a phase transition in representation space, with effective rank jumping above input dimensionality at layer 1 and stabilising at a high\-dimensional plateau\. \(2\) For short\-ξ\\xichains, attention is*irrelevant*: the Transformer converges to the same loss and fixed\-point geometry as the MLP, though Experiment 4 shows it contracts perturbations faster\. \(3\) The transition is dominated by the first\-layer head \(L0H0\), which accounts for more than4×4\\timesthe representational shift of any subsequent head, consistent with the prediction that the relevant operator acts before the MLP begins integrating out positional variation\. \(4\) Perturbation decay experiments reveal a regime reversal: in the long\-ξ\\xiregime the Transformer selectively preserves slow Markov modes \(5\.4×5\.4\\timesdynamic range in decay length vs\.1\.3×1\.3\\timesfor the MLP\); in the short\-ξ\\xiregime it suppresses all modes faster than the MLP, with no spectral selectivity\. Together, these results show that the relevance of attention is not a property of the architecture but of the spectral structure of the data\-generating process, and that a first\-order RG perturbation framework provides a predictive account of that difference\.
## 1Introduction
In our previous work\(Haggi\-Mani & Rish,[2026](https://arxiv.org/html/2607.15449#bib.bib5)\)we showed that a pure MLP residual stack trained on masked token prediction over Markov chain sequences implements a selective coarse\-graining procedure governed by the correlation lengthξ\\xiof the input distribution: short\-ξ\\xichains produce monotone rank collapse \(8\.4×8\.4\\timescompression\), long\-ξ\\xichains preserve the effective rank of the hidden representations, and in both cases inter\-layer kernel drift concentrates at one or two transitions, with the remainder of the network near a*fixed\-point plateau*consistent with RG theory\. One important finding in this controlled setting is that the forward pass does not merely*resemble*an RG flow; the sequence of representations\(h\(0\),…,h\(L\)\)\(h^\{\(0\)\},\\ldots,h^\{\(L\)\}\)executes one, with each layer performing one coarse\-graining step and the fixed\-point plateau marking convergence to the attractor\. The entry transition is the point at which the network commits to a particular attractor in representation space: the sharp concentration of kernel drift at one or two layers reflects the discrete nature of this flow, where the network crosses from one representational regime to another in a single coarse\-graining step rather than gradually\. The natural next step is to introduce attention and characterize its effect on the RG flow\. In this framework, attention enters as an additional operator whose relevance or irrelevance determines whether it shifts the fixed point by a small amount or drives the system to a qualitatively different attractor\. In the language of statistical field theory, this is equivalent to asking whether attention is a*relevant*perturbation of the MLP fixed point \(grows under RG iteration with depth, drives the system to a different fixed point\),*marginal*\(preserved under iteration, persists at constant amplitude\), or*irrelevant*\(decays, leaving the fixed\-point physics unchanged\)\. The type of operator determines the large\-scale representational behavior of the network\. This builds on a line of work connecting RG to learning systems\.Mehta & Schwab \([2014](https://arxiv.org/html/2607.15449#bib.bib9)\)construct an exact mapping between Kadanoff’s variational RG and RBM\-based deep networks, establishing the conceptual foundation;Coppola et al\. \([2026](https://arxiv.org/html/2607.15449#bib.bib17)\)extend this to weakly non\-linear networks, developing a rigorous RG framework that classifies perturbations as relevant or irrelevant and reveals universality in learning curves at large data limits\. Neither work, however, derives or tests perturbation\-theoretic predictions for the fixed\-point structure of a trained network’s representations\(Bordelon et al\.,[2024](https://arxiv.org/html/2607.15449#bib.bib3)\)\. In the spirit of Kadanoff–Wilson RG\(Wilson,[1971](https://arxiv.org/html/2607.15449#bib.bib13)\), recent work has connected RG\-like dynamics to topological phase transitions in representation manifolds\(Alpay & Kilictas,[2026](https://arxiv.org/html/2607.15449#bib.bib1)\)\.Makkuva et al\. \([2025](https://arxiv.org/html/2607.15449#bib.bib16)\)study Transformers trained on first\-order Markov chains and analytically characterize fixed points of the loss landscape, showing that attention can drive the system between a unigram and a bigram attractor, which is the closest methodological antecedent to the present work in its combination of Markov input statistics and fixed\-point reasoning, though it does not use an RG framework or classify attention as an operator\.Fernando & Guitchounts \([2026](https://arxiv.org/html/2607.15449#bib.bib15)\)provide empirical evidence that training installs a monotonic spectral gradient through depth and that perturbations are differentially amplified or suppressed layer by layer, consistent with RG\-like flow, but without formal fixed\-point predictions\. The present paper derives the fixed\-point shift formulaδ=−M∗−1\(a\+b\)\\delta=\-\{M^\{\*\}\}^\{\-1\}\(a\+b\)and uses it to make falsifiable predictions about the relevance of attention as a function of the correlation lengthξ\\xiof the input distribution, testing those predictions against trained network measurements\. Applying this framework requires a formal definition of what it means for attention to perturb the MLP fixed point, and testable predictions for how measured quantities should respond\. We provide both in Section[3](https://arxiv.org/html/2607.15449#S3), then test the predictions in Sections[4](https://arxiv.org/html/2607.15449#S4)–[7](https://arxiv.org/html/2607.15449#S7)\.
#### Paper structure\.
Section[2](https://arxiv.org/html/2607.15449#S2)is a review of our previous setup in\(Haggi\-Mani & Rish,[2026](https://arxiv.org/html/2607.15449#bib.bib5)\)\. Section[3](https://arxiv.org/html/2607.15449#S3)develops the perturbation\-theoretic framework and derives testable predictions\. Sections[4](https://arxiv.org/html/2607.15449#S4)–[7](https://arxiv.org/html/2607.15449#S7)present the experiments\. Section[8](https://arxiv.org/html/2607.15449#S8)assesses the predictions against the results, and Section[9](https://arxiv.org/html/2607.15449#S9)concludes\. Some calculations are provided in the Appendix\.
## 2Background
### 2\.1Synthetic corpus
Sequences are sampled from a Markov chain over vocabularyV=16V=16with row\-stochastic transition matrixPP\. The spectral gap governs the correlation lengthξ=−1/log\|λ2\|\\xi=\-1/\\log\|\\lambda\_\{2\}\|, whereλ2\\lambda\_\{2\}is the second\-largest eigenvalue by magnitude\. Two regimes are studied:
- •Short\-ξ\\xi:α=10\\alpha=10,λ2≈0\.44\\lambda\_\{2\}\\approx 0\.44,ξ≈1\.2\\xi\\approx 1\.2\. Sequences decorrelate within∼5\\sim 5steps\.
- •Long\-ξ\\xi:α=0\.05\\alpha=0\.05,λ2≈0\.86\\lambda\_\{2\}\\approx 0\.86,ξ≈6\.7\\xi\\approx 6\.7\. Mixing timetmix\(0\.01\)≈31t\_\{\\text\{mix\}\}\(0\.01\)\\approx 31steps; the context window carries substantial predictive signal\.
In all experimentsT=64T=64and sequences are encoded as one\-hot vectorsX∈ℝB×T×VX\\in\\mathbb\{R\}^\{B\\times T\\times V\}\. Training uses BERT\-style masked token prediction \(mask rate 15%\), cross\-entropy loss, Adam\(Kingma & Ba,[2015](https://arxiv.org/html/2607.15449#bib.bib6)\)with learning rate3×10−43\\times 10^\{\-4\}, batch sizeB=32B=32, for 10,000 steps with fresh sequences at each step\.
### 2\.2Architectures
#### MLP baseline:
A pre\-norm MLP residual stack with no attention: input projectionϕin:ℝV→ℝd\\phi\_\{\\text\{in\}\}:\\mathbb\{R\}^\{V\}\\to\\mathbb\{R\}^\{d\},LLresidual blocks each computingx←x\+MLP\(LayerNorm\(x\)\)x\\leftarrow x\+\\text\{MLP\}\(\\text\{LayerNorm\}\(x\)\)with hidden dimension4d4dand GELU activation, final LayerNorm, and classification headϕhead:ℝd→ℝV\\phi\_\{\\text\{head\}\}:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}^\{V\}\. All experiments used=64d=64,L=6L=6\.
#### TFM:
A matched architecture in which each residual block adds a pre\-norm multi\-head self\-attention sublayer before the MLP sublayer\. No positional encoding is used, making attention the only new inductive bias \(see Section[8](https://arxiv.org/html/2607.15449#S8)\)\. All other hyperparameters are identical to the MLP;nheads=1n\_\{\\text\{heads\}\}=1in the main experiments\.
### 2\.3Measurement quantities
*Effective rank*\(Roy & Vetterli,[2007](https://arxiv.org/html/2607.15449#bib.bib10)\): for a representation matrixH∈ℝN×dH\\in\\mathbb\{R\}^\{N\\times d\}with normalized singular\-value spectrumpi=σi/∑jσjp\_\{i\}=\\sigma\_\{i\}/\\sum\_\{j\}\\sigma\_\{j\},
ρeff\(H\)=exp\(−∑ipilogpi\)∈\[1,d\]\.\\rho\_\{\\mathrm\{eff\}\}\(H\)=\\exp\\\!\\left\(\-\\sum\_\{i\}p\_\{i\}\\log p\_\{i\}\\right\)\\in\[1,d\]\.\(1\)A decreasing depth profile signals progressive coarse\-graining\(Haggi\-Mani & Rish,[2026](https://arxiv.org/html/2607.15449#bib.bib5)\), a pattern also observed empirically in large pre\-trained models\(Alpay & Kilictas,[2026](https://arxiv.org/html/2607.15449#bib.bib1); Fernando & Guitchounts,[2026](https://arxiv.org/html/2607.15449#bib.bib15)\)\.*Kernel drift:*ΔCKA\(l,l\+1\)=1−CKA\(H\(l\),H\(l\+1\)\)\\Delta\_\{\\mathrm\{CKA\}\}\(l,l\{\+\}1\)=1\-\\text\{CKA\}\(H^\{\(l\)\},H^\{\(l\+1\)\}\), where CKA is centered kernel alignment\(Kornblith et al\.,[2019](https://arxiv.org/html/2607.15449#bib.bib7)\)
CKA\(X,Y\)=‖Y⊤X‖F2‖X⊤X‖F‖Y⊤Y‖F,\\mathrm\{CKA\}\(X,Y\)=\\frac\{\\\|Y^\{\\top\}X\\\|\_\{F\}^\{2\}\}\{\\\|X^\{\\top\}X\\\|\_\{F\}\\\|Y^\{\\top\}Y\\\|\_\{F\}\},\(2\)andX,Y∈ℝn×dX,Y\\in\\mathbb\{R\}^\{n\\times d\}are centered representation matrices overnnexamples: Small drift indicates a fixed point; concentrated drift marks discrete transition events\. The layer\-by\-layer progression from syntactic to semantic representations observed in BERT\(Tenney et al\.,[2019](https://arxiv.org/html/2607.15449#bib.bib11)\)is consistent with the coarse\-graining interpretation adopted here\. All representations are extracted in flatten mode \(N=B×TN=B\\times T\) to avoid the mean\-pooling artifact identified in\(Haggi\-Mani & Rish,[2026](https://arxiv.org/html/2607.15449#bib.bib5)\)\.
## 3Theoretical Framework: Attention as an RG Perturbation
### 3\.1The stability of the MLP fixed point
As established in our previous work, the trained MLP residual stack converges to a fixed\-point plateau: after one or two entry transitions, inter\-layer kernel drift is near zero and representations undergo no further geometric change\. The stability of this fixed\-point \(See Appendix[Appendix A](https://arxiv.org/html/2607.15449#A1)\) is governed by
ϵn=\(I\+M\)nϵ0,\\epsilon\_\{n\}=\(I\+\{M\}\)^\{n\}\\,\\epsilon\_\{0\},\(3\)wherennis the number of iterations\. In the eigenbasis of the JacobianM\{M\}with eigenvaluesμk\\mu\_\{k\}this is
\[ϵn\]k=\(1\+μk\)n\[ϵ0\]k\.\[\\epsilon\_\{n\}\]\_\{k\}=\(1\+\\mu\_\{k\}\)^\{n\}\\,\[\\epsilon\_\{0\}\]\_\{k\}\.\(4\)The MLP fixed\-point is stable if and only if all eigenvalues ofI\+MI\+\{M\}lie strictly inside the unit circle:
\|1\+μk\|<1for allk\.\|1\+\\mu\_\{k\}\|<1\\quad\\text\{for all \}k\.\(5\)For real eigenvalues this reduces to−2<μk<0\-2<\\mu\_\{k\}<0: the MLP sub\-network must contract perturbations atx∗x^\{\*\}, but not so strongly as to overshoot\. Modes withμk\>0\\mu\_\{k\}\>0orμk<−2\\mu\_\{k\}<\-2grow under iteration \(unstable eigenvector directions\)\. The fixed\-point plateau observed empirically in bothξ\\xiregimes implies that the trained MLP satisfies this condition across all modes\. The low effective rank \(ρeff≈1\.8\\rho\_\{\\mathrm\{eff\}\}\\approx 1\.8\) of the attractor reflects that most eigenvaluesμk\\mu\_\{k\}are close to−1\-1, so that nearly all directions in representation space are contracted to zero, leaving only a low\-dimensional stable subspace\.
### 3\.2The TFM Fixed\-Point as a Perturbation of the MLP Fixed\-Point
We add a single\-head attention sublayer
𝒜\(x\)=softmax\(QK⊤/d\)V\\mathcal\{A\}\(x\)=\\mathrm\{softmax\}\(QK^\{\\top\}/\\sqrt\{d\}\)\\,V\(6\)withQ=xWQQ=xW\_\{Q\},K=xWKK=xW\_\{K\},V=xWVV=xW\_\{V\}to each MLP block\. Assuming that the new fixed pointx~∗\\tilde\{x\}^\{\*\}is a small shiftδ\\deltaaway from the previous MLP fixed point \(See Appendix[Appendix B](https://arxiv.org/html/2607.15449#A2)\), the shift is given by
δ=−M∗−1\(a\+b\)\\delta\\;=\\;\-\{M^\{\*\}\}^\{\-1\}\(a\+b\)\(7\)whereM∗≡\(I\+D\[b\]\)\(I\+D\[a\]\)−IM^\{\*\}\\equiv\(I\+D\[b\]\)\(I\+D\[a\]\)\-I, i\.e\.
M∗=D\[b\]\+D\[a\]\+D\[b\]D\[a\],M^\{\*\}=D\[b\]\+D\[a\]\+D\[b\]D\[a\],\(8\)and we have defined
a≡𝒜\(x∗\),b≡f\(x∗\+a\),D\[a\]≡D\[𝒜\]\(x∗\),D\[b\]≡D\[f\]\(x∗\+a\)\.a\\equiv\\mathcal\{A\}\(x^\{\*\}\),\\quad b\\equiv f\(x^\{\*\}\+a\),\\quad D\[a\]\\equiv D\[\\mathcal\{A\}\]\(x^\{\*\}\),\\quad D\[b\]\\equiv D\[f\]\(x^\{\*\}\+a\)\.\(9\)
We have assumed in formula \([7](https://arxiv.org/html/2607.15449#S3.E7)\) thatM∗\{M^\{\*\}\}is invertible\. Note thata\+b=0a\+b=0if and only ifx∗x^\{\*\}is already a fixed point of the modified block; hencea\+ba\+bmeasures the residual of the fixed\-point condition evaluated atx∗x^\{\*\}\. Decomposinga\+ba\+bin the eigenbasis\{vk\}\\\{v\_\{k\}\\\}ofM∗\{M^\{\*\}\}with eigenvalues\{νk\}\\\{\\nu\_\{k\}\\\}, formula \([7](https://arxiv.org/html/2607.15449#S3.E7)\) gives
δ=−∑k\[a\+b\]kνkvk,\[δ\]k=−\[a\+b\]k/νk,\\delta=\-\\sum\_\{k\}\\frac\{\[a\+b\]\_\{k\}\}\{\\nu\_\{k\}\}\\,v\_\{k\},\\quad\[\\delta\]\_\{k\}=\-\[a\+b\]\_\{k\}/\\nu\_\{k\},\(10\)where\[δ\]k\[\\delta\]\_\{k\}is the shift along the eigenvectorvkv\_\{k\}\.The Stability of the TFM Fixed\-Point\(see Appendix[Appendix C](https://arxiv.org/html/2607.15449#A3)\)\. Linearizing aroundx~∗\\tilde\{x\}^\{\*\}withxn=x~∗\+ϵnx\_\{n\}=\\tilde\{x\}^\{\*\}\+\\epsilon\_\{n\}, the perturbation evolves as
ϵn=M~∗nϵ0,M~∗≡\(I\+D\[b~\]\)\(I\+D\[a~\]\),\\epsilon\_\{n\}=\\tilde\{M\}^\{\*n\}\\,\\epsilon\_\{0\},\\qquad\\tilde\{M\}^\{\*\}\\equiv\(I\+D\[\\tilde\{b\}\]\)\(I\+D\[\\tilde\{a\}\]\),\(11\)where
a~≡𝒜\(x~∗\),b~≡f\(x~∗\+a~\)=−a~,D\[a~\]≡D\[𝒜\]\(x~∗\),D\[b~\]≡D\[f\]\(x~∗\+a~\)\.\\tilde\{a\}\\equiv\\mathcal\{A\}\(\\tilde\{x\}^\{\*\}\),\\quad\\tilde\{b\}\\equiv f\(\\tilde\{x\}^\{\*\}\+\\tilde\{a\}\)=\-\\tilde\{a\},\\quad D\[\\tilde\{a\}\]\\equiv D\[\\mathcal\{A\}\]\(\\tilde\{x\}^\{\*\}\),\\quad D\[\\tilde\{b\}\]\\equiv D\[f\]\(\\tilde\{x\}^\{\*\}\+\\tilde\{a\}\)\.\(12\)The fixed pointx~∗\\tilde\{x\}^\{\*\}is stable if and only if all eigenvaluesν~k\\tilde\{\\nu\}\_\{k\}ofM~∗\\tilde\{M\}^\{\*\}lie strictly inside the unit circle:
\|ν~k\|<1for allk\.\|\\tilde\{\\nu\}\_\{k\}\|<1\\quad\\text\{for all \}k\.\(13\)In the eigenbasis\{v~k\}\\\{\\tilde\{v\}\_\{k\}\\\}ofM~∗\\tilde\{M\}^\{\*\}each mode evolves as\[ϵn\]k=ν~kn\[ϵ0\]k\[\\epsilon\_\{n\}\]\_\{k\}=\\tilde\{\\nu\}\_\{k\}^\{n\}\\,\[\\epsilon\_\{0\}\]\_\{k\}, so modes with\|ν~k\|<1\|\\tilde\{\\nu\}\_\{k\}\|<1decay, modes with\|ν~k\|\>1\|\\tilde\{\\nu\}\_\{k\}\|\>1grow, and modes with\|ν~k\|=1\|\\tilde\{\\nu\}\_\{k\}\|=1are marginal\.Notethat whileM~∗\\tilde\{M\}^\{\*\}\(\{v~k,ν~k\}\\\{\\tilde\{v\}\_\{k\},\\tilde\{\\nu\}\_\{k\}\\\}\) is evaluated at the TFM fixed pointx~∗\\tilde\{x\}^\{\*\}, theM∗M^\{\*\}\(\{vk,νk\}\\\{v\_\{k\},\\nu\_\{k\}\\\}\) of the shift formula \([8](https://arxiv.org/html/2607.15449#S3.E8)\) is evaluated at the MLP fixed pointx∗x^\{\*\}\. However, when the shiftδ=x~∗−x∗\\delta=\\tilde\{x\}^\{\*\}\-x^\{\*\}is small, i\.e\. the two evaluation points are close
D\[a~\]≈D\[a\],D\[b~\]≈D\[b\],D\[\\tilde\{a\}\]\\approx D\[a\],\\qquad D\[\\tilde\{b\}\]\\approx D\[b\],\(14\)so that
M~∗=\(I\+D\[b~\]\)\(I\+D\[a~\]\)≈\(I\+D\[b\]\)\(I\+D\[a\]\)=I\+M∗\.\\tilde\{M\}^\{\*\}=\(I\+D\[\\tilde\{b\}\]\)\(I\+D\[\\tilde\{a\}\]\)\\approx\(I\+D\[b\]\)\(I\+D\[a\]\)=I\+M^\{\*\}\.\(15\)In this regime the stability condition\|ν~k\|<1\|\\tilde\{\\nu\}\_\{k\}\|<1reduces to\|1\+νk\|<1\|1\+\\nu\_\{k\}\|<1on the eigenvaluesνk\\nu\_\{k\}ofM∗M^\{\*\}, i\.e\.−2<νk<0\-2<\\nu\_\{k\}<0for real eigenvalues, the same condition as for the MLP fixed point, but with the Jacobian now evaluated at the attention\-shifted pointx∗\+ax^\{\*\}\+a\. This connects the shift formula and the stability analysis: under the approximationM~∗≈I\+M∗\\tilde\{M\}^\{\*\}\\approx I\+M^\{\*\}, the same matrixM∗M^\{\*\}that controls the magnitude of the fixed\-point shiftδ=−M∗−1\(a\+b\)\\delta=\-\{M^\{\*\}\}^\{\-1\}\(a\+b\)also governs the stability of the new fixed point\. In particular, near a bifurcation whereM∗M^\{\*\}develops a near\-zero eigenvalueνk≈0\\nu\_\{k\}\\approx 0, the shiftδ\\deltadiverges alongvkv\_\{k\}while simultaneouslyν~k=1\+νk→1\\tilde\{\\nu\}\_\{k\}=1\+\\nu\_\{k\}\\to 1, pushing that mode to the boundary of the unit circle and rendering the fixed point marginally stable\.Notethat equations \([10](https://arxiv.org/html/2607.15449#S3.E10)\) and \([11](https://arxiv.org/html/2607.15449#S3.E11)\) both arise from linearizing the block mapFFaround a fixed point, but at different points and addressing different questions\.The first, equation \([10](https://arxiv.org/html/2607.15449#S3.E10)\), is static: given that attention has been added, where is the new fixed pointx~∗\\tilde\{x\}^\{\*\}? Hereδ\\deltais a fixed vector measuring the displacement fromx∗x^\{\*\}tox~∗\\tilde\{x\}^\{\*\}\.The second, equation \([11](https://arxiv.org/html/2607.15449#S3.E11)\), is dynamic: once atx~∗\\tilde\{x\}^\{\*\}, do nearby trajectories return to it or diverge? Hereϵn\\epsilon\_\{n\}is an evolving perturbation whose amplitude changes with the number of iterationsnn\.
#### Invertibility ofM∗M^\{\*\}and the bifurcation condition:
Formula \([7](https://arxiv.org/html/2607.15449#S3.E7)\) requiresM∗M^\{\*\}to be invertible, i\.e\.νk≠0\\nu\_\{k\}\\neq 0for allkk\. Since the eigenvalues ofI\+M∗I\+M^\{\*\}are1\+νk1\+\\nu\_\{k\}, a zero eigenvalue ofM∗M^\{\*\}corresponds to a unit eigenvalue ofI\+M∗I\+M^\{\*\}, precisely the bifurcation condition: the fixed point may be destroyed, created, or exchange stability, and the perturbative prediction breaks down sinceM∗−1\{M^\{\*\}\}^\{\-1\}no longer exists\. There are three cases:
- •IfM∗M^\{\*\}is invertible and\|1\+νk\|<1\|1\+\\nu\_\{k\}\|<1for allkk: the new fixed point exists, is unique to first order, and is stable\.
- •If someνk→0\\nu\_\{k\}\\to 0:M∗M^\{\*\}becomes singular,‖δ‖→∞\\\|\\delta\\\|\\to\\infty, and perturbation theory breaks down\. The system is at a bifurcation point; the fixed point may cease to exist or become non\-unique\.
- •If\|1\+νk\|\>1\|1\+\\nu\_\{k\}\|\>1for somekk:M∗−1\{M^\{\*\}\}^\{\-1\}exists but the new fixed point is unstable: attention has destabilized the fixed point in that direction\.
The measurements in Experiment 1 are consistent with the second case being realized in the long\-ξ\\xiregime, where the system crosses into a qualitatively different dynamical regime that lies outside the scope of the perturbation expansion\.
### 3\.3RG classification
The eigenvaluesνk\\nu\_\{k\}ofM∗M^\{\*\}appear in both the shift formula \([10](https://arxiv.org/html/2607.15449#S3.E10)\) and the stability equation \([11](https://arxiv.org/html/2607.15449#S3.E11)\), and together determine the RG character of attention in directionvkv\_\{k\}\.
- •Irrelevant\(−2<νk<0\-2<\\nu\_\{k\}<0,\|1\+νk\|<1\|1\+\\nu\_\{k\}\|<1\): perturbations in directionvkv\_\{k\}contract at each iteration and are progressively integrated out, leaving no trace at the fixed point\. The shift\[δ\]k=−\[a\+b\]k/νk\[\\delta\]\_\{k\}=\-\[a\+b\]\_\{k\}/\\nu\_\{k\}is moderate — of the same order as the driving term\[a\+b\]k\[a\+b\]\_\{k\}— andx~∗\\tilde\{x\}^\{\*\}is stable in this direction\.
- •Marginal\(νk=−2\\nu\_\{k\}=\-2,\|1\+νk\|=1\|1\+\\nu\_\{k\}\|=1\): perturbations neither grow nor decay; the shift is finite and of order\[a\+b\]k\[a\+b\]\_\{k\}\. The fixed point is neutrally stable in directionvkv\_\{k\}and perturbations persist at constant amplitude, potentially producing power\-law rather than exponential dependence on depth\. Whether the system drifts away depends on nonlinear terms beyond the first\-order expansion; distinguishing marginal from irrelevant in practice requires measuring the scaling ofΔCKA\\Delta\_\{\\mathrm\{CKA\}\}with sequence lengthTT, as carried out in Experiment 3\.
- •Relevant\(νk\>0\\nu\_\{k\}\>0orνk<−2\\nu\_\{k\}<\-2,\|1\+νk\|\>1\|1\+\\nu\_\{k\}\|\>1\): perturbations grow at each iteration and drive the system away fromx~∗\\tilde\{x\}^\{\*\}toward a new stable attractor\. In the Wilsonian picture this mode is not integrated out but instead grows under coarse\-graining, reorganizing the system into a qualitatively different representational regime\.
- •Bifurcation\(νk=0\\nu\_\{k\}=0,\|1\+νk\|=1\|1\+\\nu\_\{k\}\|=1\): the denominator of\[δ\]k=−\[a\+b\]k/νk\[\\delta\]\_\{k\}=\-\[a\+b\]\_\{k\}/\\nu\_\{k\}vanishes and the shift diverges —x~∗\\tilde\{x\}^\{\*\}has moved infinitely far fromx∗x^\{\*\}and cannot be located perturbatively\. Simultaneously perturbations aroundx~∗\\tilde\{x\}^\{\*\}neither grow nor decay\. Both equations signal the same event: the perturbative expansion breaks down entirely and the system reorganizes into a qualitatively different attractor outside the scope of the linearized expansion\.
### 3\.4The network as a discrete RG phase space
In\(Haggi\-Mani & Rish,[2026](https://arxiv.org/html/2607.15449#bib.bib5)\)we established that the network implements a discrete approximation to an RG flow: each layer executes one coarse\-graining step, and the fixed\-point plateau corresponds to the attractor of that flow\. The perturbation analysis above allows us to develop this picture more precisely\. In a continuous RG phase space, each point represents a theory specified by its coupling constants\(g1,g2,…\)\(g\_\{1\},g\_\{2\},\\ldots\), and the RG flow is a vector field describing how these couplings change when short\-distance modes are integrated out\. A trajectory is the path a theory traces through this space under repeated coarse\-graining; different initial theories give rise to different trajectories\. Near a fixed point, the linearized RG transformation has eigenvectors that define the natural axes of the space: relevant/irrelevant directions \(positive/negative scaling dimensions\) flow away from/toward the fixed point\. The direction of the trajectory is determined by which eigendirections are excited in the initial theory, and its rate is controlled by the magnitude of the corresponding scaling dimension\. In our discrete setting, the role of initial conditions is played by the input distribution\. The network weights are fixed, soFF\(the residual block\) and its fixed pointx~∗\\tilde\{x\}^\{\*\}are determined; the only remaining freedom is the representationh\(0\)h^\{\(0\)\}at the initial layer, which is set by the input\. The sequence of representations\(h\(0\),h\(1\),…,h\(L\)\)\(h^\{\(0\)\},h^\{\(1\)\},\\ldots,h^\{\(L\)\}\)traces a trajectory through representation space \(not phase space\), and the direction and rate at which it approaches the fixed point are shaped by which eigendirections ofM~∗\\tilde\{M\}^\{\*\}are excited inh\(0\)h^\{\(0\)\}\. An eigendirectionv~k\\tilde\{v\}\_\{k\}with eigenvalueν~k\\tilde\{\\nu\}\_\{k\}is excited ifh\(0\)h^\{\(0\)\}has a large component in that direction relative to the fixed point; by the stability equation \([11](https://arxiv.org/html/2607.15449#S3.E11)\), more strongly excited directions take longer to contract\. The fixed\-point plateau observed in\(Haggi\-Mani & Rish,[2026](https://arxiv.org/html/2607.15449#bib.bib5)\)is the empirical signature of this convergence: the trajectory reachesx~∗\\tilde\{x\}^\{\*\}at some intermediate layerl∗l^\{\*\}, after which representations change negligibly because the system is already near the attractor\. The plateau therefore marks the depth at which the dominant eigendirections ofM~∗\\tilde\{M\}^\{\*\}have contracted sufficiently, and its location depends on which eigendirections were excited inh\(0\)h^\{\(0\)\}and how strongly\. The analogy with RG phase space is imperfect at two levels\. At the level of different networks, the analogy is closest: each trained network has its own weights, its own mapFF, and hence its own fixed pointx~∗\\tilde\{x\}^\{\*\}and scaling dimensions \(eigenvalues ofM~∗\\tilde\{M\}^\{\*\}\)\. Different trained networks therefore correspond to different theories; different points in RG phase space with different attractors and different flows\. At the level of a single network, the weights,FF, andx~∗\\tilde\{x\}^\{\*\}are all fixed, so we are operating at a single point in RG phase space\. Different inputs correspond not to different theories but to different initial conditions: they place the trajectory at different starting pointsh\(0\)h^\{\(0\)\}in representation space, exciting different combinations of eigendirections ofM~∗\\tilde\{M\}^\{\*\}, but all flowing under the sameFFtoward the samex~∗\\tilde\{x\}^\{\*\}\. It is worth noting that different inputs converging to the samex~∗\\tilde\{x\}^\{\*\}is a statement about stability, not universality\. Universality in the RG sense would require different trained networks — with different weights, training data, or hyperparameters — to share the same fixed\-point structure \(x~∗\\tilde\{x\}^\{\*\}and the eigenvalues ofM~∗\\tilde\{M\}^\{\*\}\)\. That is a much stronger claim, and one we do not make here\. Whether attention is a relevant or irrelevant operator determines how the TFM fixed pointx~∗\\tilde\{x\}^\{\*\}differs from the MLP fixed pointx∗x^\{\*\}\. When attention is relevant \(\|1\+νk\|\>1\|1\+\\nu\_\{k\}\|\>1for somekk\), the TFM fixed point has new unstable eigendirections and the attractor shifts to a qualitatively different fixed\-point structure relative to the MLP\. When attention is irrelevant \(−2<νk<0\-2<\\nu\_\{k\}<0for allkk\), the TFM fixed point differs fromx∗x^\{\*\}only by a small shiftδ\\delta, all eigendirections contract, and the fixed\-point structure is essentially unchanged\. The bifurcation between these two cases, driven by the near\-singularity ofM∗M^\{\*\}, is the mechanism behind the predictions that follow\.
### 3\.5Predictions
The perturbation analysis of Section[3\.2](https://arxiv.org/html/2607.15449#S3.SS2)and the RG classification of Section[3\.3](https://arxiv.org/html/2607.15449#S3.SS3)together yield four testable predictions, each grounded in the structure of equations \([7](https://arxiv.org/html/2607.15449#S3.E7)\), \([10](https://arxiv.org/html/2607.15449#S3.E10)\), and \([11](https://arxiv.org/html/2607.15449#S3.E11)\)\. The first two concern the magnitude of the fixed\-point shiftδ\\deltaand the stability of the new fixed point in the two regimes; the third concerns the layer specificity of the shift; and the fourth concerns the spectral structure of the shift across Markov eigenmodes\. For short\-ξ\\xichains, tokens decorrelate within a few steps and the optimal prediction is the stationary distributionπ\\pieverywhere\. As established in\(Haggi\-Mani & Rish,[2026](https://arxiv.org/html/2607.15449#bib.bib5)\), this is what the MLP fixed point encodes: all token representations collapse to a low\-dimensional attractor \(ρeff≈1\.8\\rho\_\{\\mathrm\{eff\}\}\\approx 1\.8\), where positional variation has been integrated out\. The fixed\-point plateau impliesx∗x^\{\*\}is stable \([3](https://arxiv.org/html/2607.15449#S3.E3)\), with all eigenvalues ofI\+MI\+Minside the unit circle\. When attention is added, the shiftδ=−M∗−1\(a\+b\)\\delta=\-\{M^\{\*\}\}^\{\-1\}\(a\+b\)\([7](https://arxiv.org/html/2607.15449#S3.E7)\) is small if both the resolventM∗−1\{M^\{\*\}\}^\{\-1\}is bounded and the driving terma\+ba\+bis small\. In this case, both conditions are plausibly satisfied\. Since representations atx∗x^\{\*\}are nearly identical across positions,a=𝒜\(x∗\)a=\\mathcal\{A\}\(x^\{\*\}\)carries no contextual signal; this leads toD\[a\]D\[a\]being small, which in turn meansM∗=\(I\+D\[b\]\)\(I\+D\[a\]\)−I≈D\[b\]M^\{\*\}=\(I\+D\[b\]\)\(I\+D\[a\]\)\-I\\approx D\[b\]\. But ifaahas a small variation,x∗=ax^\{\*\}=ais close tox∗x^\{\*\}and soD\[b\]=Df\(x∗\+a\)≈Df\(x∗\)=MD\[b\]=Df\(x^\{\*\}\+a\)\\approx Df\(x^\{\*\}\)=M, which means thatM∗M^\{\*\}is close toMMat the fixed point and inherits its eigenstructure that dictates the stability ofx∗x^\{\*\}\. SoM∗M^\{\*\}remains invertible andM∗−1\{M^\{\*\}\}^\{\-1\}is bounded\. whethera\+ba\+bis small cannot be guaranteed on theoretical grounds and remains an empirical question\. Onceδ\\deltais small,M~∗≈I\+M∗\\tilde\{M\}^\{\*\}\\approx I\+M^\{\*\}and the TFM fixed point inherits the stability ofx∗x^\{\*\}\. The prediction thatδ\\deltais small is therefore not a strict consequence of the theory but a hypothesis consistent with the structure of the short\-ξ\\xifixed point, confirmed empirically in Experiment 1\.
###### Prediction 1\(Irrelevance in short\-ξ\\xi\)\.
For short\-ξ\\xichains, the structure of the fixed point suggests that both the driving terma\+ba\+band the resolventM∗−1\{M^\{\*\}\}^\{\-1\}are bounded, though this is not guaranteed by the theory alone\. We predict that the TFM converges to the same loss and similar representational geometry as the MLP, and treat this as an empirical test of the hypothesis\.
For long\-ξ\\xichains the MLP leaves a residual loss gap of\+0\.017\+0\.017nats because it cannot exploit cross\-token structure\. Atx∗x^\{\*\}, token representations retain positional variation from unresolved long\-range correlations, so value vectorsV=x∗WVV=x^\{\*\}W\_\{V\}differ across positions anda=𝒜\(x∗\)a=\\mathcal\{A\}\(x^\{\*\}\)carries a genuine contextual signal\. Furthermore, the MLP’s contraction is weakest along the slow Markov modes \(large\|λk\|\|\\lambda\_\{k\}\|\), since the MLP has no mechanism to process cross\-token structure in those directions\. The correspondingνk\\nu\_\{k\}may therefore be close to zero, makingM∗−1\{M^\{\*\}\}^\{\-1\}large in those directions\. In the long\-ξ\\xicase,a=𝒜\(x∗\)a=\\mathcal\{A\}\(x^\{\*\}\)carries genuine contextual signal and varies across positions; sincef\(x∗\)=0f\(x^\{\*\}\)=0, linearizing givesb≈D\[b\]⋅ab\\approx D\[b\]\\cdot a, so the driving terma\+b≈\(I\+D\[b\]\)aa\+b\\approx\(I\+D\[b\]\)ais of the same order asaa\. Whether‖a‖\\\|a\\\|is large enough for the shiftδ=−M∗−1\(a\+b\)\\delta=\-\{M^\{\*\}\}^\{\-1\}\(a\+b\)\([7](https://arxiv.org/html/2607.15449#S3.E7)\) to be large cannot be guaranteed on theoretical grounds; this remains an empirical question, and the argument is therefore suggestive rather than a strict consequence of the theory\. It is worth noting a tension in the perturbation framework\. The fixed\-point shift formula \([7](https://arxiv.org/html/2607.15449#S3.E7)\) is derived under the assumption that‖δ‖\\\|\\delta\\\|is small\. Under this conditionx~∗=x∗\+δ\\tilde\{x\}^\{\*\}=x^\{\*\}\+\\deltais close tox∗x^\{\*\}andD\[b~\]≈D\[b\]D\[\\tilde\{b\}\]\\approx D\[b\]follows without additional assumptions\. In the short\-ξ\\xicase this is empirically self\-consistent:M∗−1\{M^\{\*\}\}^\{\-1\}is bounded, and the TFM converges to similar representational geometry as the MLP \(Experiment 1\), consistent with a small shiftδ\\delta\. In the long\-ξ\\xicase the assumption‖δ‖≪1\\\|\\delta\\\|\\ll 1is expected to fail: the slow Markov modes suggestM∗M^\{\*\}has near\-zero eigenvalues, and the TFM converges to a qualitatively different representational geometry than the MLP \(Experiment 1\), consistent with a large or divergingδ\\delta\. The breakdown of the perturbative expansion signals that the system is driven far from the fixed\-point, and the empirical observation of a phase transition in Experiment 1 is consistent with the system crossing into a different basin of attraction\. Whether this constitutes a true bifurcation in the dynamical systems sense would require a more detailed analysis of the fixed\-point structure beyond the linearized expansion\.
###### Prediction 2\(Relevance in long\-ξ\\xi\)\.
For long\-ξ\\xichains,a\+ba\+bis finite andM∗M^\{\*\}may have near\-zero eigenvalues, soδ=−M∗−1\(a\+b\)\\delta=\-\{M^\{\*\}\}^\{\-1\}\(a\+b\)can be large enough to push the system out of the basin of attraction ofx∗x^\{\*\}, driving a phase transition to a new fixed point with qualitatively different representational geometry\.
Depending on the eigenvaluesλk\\lambda\_\{k\}of the transition matrixPP, the corresponding eigenmodesϕk\\phi\_\{k\}encode structure at different scales: slow modes \(large\|λk\|<1\|\\lambda\_\{k\}\|<1, largeξk=−1/log\|λk\|\\xi\_\{k\}=\-1/\\log\|\\lambda\_\{k\}\|\) encode long\-range contextual information that persists over many steps; fast modes \(small\|λk\|\|\\lambda\_\{k\}\|, smallξk\\xi\_\{k\}\) encode short\-range structure that decorrelates quickly\. The empirical decay lengthτk\\tau\_\{k\}, fitted fromdk\(l\)≈Ake−l/τkd\_\{k\}\(l\)\\approx A\_\{k\}e^\{\-l/\\tau\_\{k\}\}, measures the approximate number of layers the network takes to suppress a perturbation alongϕk\\phi\_\{k\}\. From equation \([11](https://arxiv.org/html/2607.15449#S3.E11)\), using the approximationM~∗≈I\+M∗\\tilde\{M\}^\{\*\}\\approx I\+M^\{\*\}, which holds nearx∗x^\{\*\}whenδ\\deltais small, a perturbation in directionvkv\_\{k\}evolves as\[ϵl\]k=\(1\+νk\)l\[ϵ0\]k\[\\epsilon\_\{l\}\]\_\{k\}=\(1\+\\nu\_\{k\}\)^\{l\}\[\\epsilon\_\{0\}\]\_\{k\}, wherellcounts iterations of the block mapFF\. Identifying one iteration ofFFwith one network layer, an approximation that holds when consecutive layers are approximately homogeneous, as in the fixed\-point plateau; the decay rate is controlled by\|1\+νk\|\|1\+\\nu\_\{k\}\|: when\|1\+νk\|\|1\+\\nu\_\{k\}\|is close to11, the perturbation persists across many layers \(largeτk\\tau\_\{k\}\); when\|1\+νk\|≪1\|1\+\\nu\_\{k\}\|\\ll 1, it collapses quickly \(smallτk\\tau\_\{k\}\)\. The RG prediction is thatτk\\tau\_\{k\}should trackξk\\xi\_\{k\}\. Note thatϕk∈ℝT\\phi\_\{k\}\\in\\mathbb\{R\}^\{T\}lives in sequence space whilevk∈ℝdv\_\{k\}\\in\\mathbb\{R\}^\{d\}lives in representation space, so the correspondence is not a geometric alignment between vectors but a relationship between scalars:τk\\tau\_\{k\}should increase withξk\\xi\_\{k\}across modes\. Such a correspondence would be consistent with the network having learned to match its contraction rates to the spectral structure of the input distribution\. On the other hand, the MLP, being position\-blind, has no mechanism to distinguish modes and should therefore mode\-blind:τk\\tau\_\{k\}approximately constant acrosskkregardless of\|λk\|\|\\lambda\_\{k\}\|\.
###### Prediction 3\(Mode selectivity\)\.
The perturbation decay lengthτk\\tau\_\{k\}, fitted fromdk\(l\)≈Ake−l/τkd\_\{k\}\(l\)\\approx A\_\{k\}e^\{\-l/\\tau\_\{k\}\}after injecting a perturbation along eigenmodeϕk\\phi\_\{k\}ofPP, should be monotonically increasing in\|λk\|\|\\lambda\_\{k\}\|for the TFM\. The MLP should be mode\-blind, withτk\\tau\_\{k\}approximately constant acrosskk\.
The fixed\-point shift formula \([7](https://arxiv.org/html/2607.15449#S3.E7)\) assigns a shift\[δ\]k=−\[a\+b\]k/νk\[\\delta\]\_\{k\}=\-\[a\+b\]\_\{k\}/\\nu\_\{k\}to each directionvkv\_\{k\}, but does not specify at which layer the shift is concentrated\. The driving terma\(l\)\+b\(l\)a^\{\(l\)\}\+b^\{\(l\)\}at layerlldepends on the token representationsh\(l\)h^\{\(l\)\}: before the network has converged tox∗x^\{\*\}, token representations retain positional variation, soa\(0\)a^\{\(0\)\}at the first layer carries genuine contextual signal\. Once the fixed\-point plateau is reached, token representations are nearly identical across positions,a\(l\)a^\{\(l\)\}carries little positional signal, and the driving term contributes negligibly to the shift\. The RG prediction is therefore that the fixed\-point shift is concentrated at the earliest layer: the TFM commits to its new attractor in a single coarse\-graining step, consistent with the discrete nature of the RG flow\. If attention were instead distributed uniformly across layers, the shift would be spread evenly and no single layer would dominate; a pattern the fixed\-point structure does not predict, since the driving terma\(l\)\+b\(l\)a^\{\(l\)\}\+b^\{\(l\)\}carries genuine positional signal only before the fixed\-point plateau is reached\.
###### Prediction 4\(Layer specificity\)\.
The fixed\-point shiftδ\\deltashould be concentrated at the earliest layerl=0l=0, where token representations still retain positional variation and the driving terma\(0\)\+b\(0\)a^\{\(0\)\}\+b^\{\(0\)\}carries genuine contextual signal\. Subsequent layers, operating on the fixed\-point plateau where positional variation has already been integrated out, should contribute negligible shifts\.
## 4Experiment 1: Matched MLP vs\. TFM Comparison
We train the MLP and TFM on bothξ\\xiregimes with matched hyperparameters \(VV,dd,LL,TT,BB, steps, learning rate, mask rate, random seed, and Markov chainPP\)\. The MLP has≈200\\approx 200K parameters vs TFM’s≈300\\approx 300K owing to its attention weights\.
### 4\.1Loss profiles
Table 1:Final loss at step 10,000 relative to the stationary entropyH\(π\)H\(\\pi\)\. Although the negligibly small negative gaps for MLP short and TFM short \(≤0\.004\\leq 0\.004nats\) are within the noise floor of the training procedure, a negative gap indicates the model has learned to exploit contextual correlations beyond the marginal \(not in RG sense\) distributionπ\\pi; a gap near zero indicates convergence to the marginal\.#### Short\-ξ\\xi:
Both models converge to within±0\.004\\pm 0\.004nats ofH\(π\)H\(\\pi\); the inter\-model gap is negligible\. For large Dirichlet parameterα=10\\alpha=10, rows ofPPare nearly uniform and context carries little predictive signal, soH\(Xt∣Xt−1\)≈H\(π\)H\(X\_\{t\}\\mid X\_\{t\-1\}\)\\approx H\(\\pi\)\. This confirms Prediction[1](https://arxiv.org/html/2607.15449#Thmprediction1)\.
#### Long\-ξ\\xi:
The MLP plateaus at\+0\.017\+0\.017nats aboveH\(π\)H\(\\pi\); the TFM reaches−0\.195\-0\.195nats below it\. Sequences are freshly sampled at each step, so the TFM’s lower loss is not overfitting: it has learned the conditionalP\(xt∣xt−1,…\)P\(x\_\{t\}\\mid x\_\{t\-1\},\\ldots\)rather than the marginalπ\\pi\. Forα=0\.05\\alpha=0\.05each row ofPPis heavily peaked, givingH\(Xt∣Xt−1\)≪H\(π\)H\(X\_\{t\}\\mid X\_\{t\-1\}\)\\ll H\(\\pi\)as the remaining uncertainty is small; a well\-trained TFM approaches this conditional entropy\. The loss gap is consistent with Prediction[2](https://arxiv.org/html/2607.15449#Thmprediction2); the representational evidence is presented in Section[4\.2](https://arxiv.org/html/2607.15449#S4.SS2)\.
### 4\.2Rank collapse and fixed\-point geometry
Here, we examine the effective rank profiles and kernel drift measurements to investigate how the fixed\-point geometry differs between the MLP and TFM in the two regimes, and whether the representational signatures of Prediction[2](https://arxiv.org/html/2607.15449#Thmprediction2), i\.e\. rank expansion and drift concentration, are observed empirically\.
Table 2:Kernel driftΔl=1−CKA\(Hl,Hl\+1\)\\Delta\_\{l\}=1\-\\text\{CKA\}\(H\_\{l\},H\_\{l\+1\}\)between consecutive layers\. The long\-ξ\\xiTFM does not cleanly stabilize across depth\.As seen in Table[2](https://arxiv.org/html/2607.15449#S4.T2), in the short\-ξ\\xiregime, MLP drift drops to0\.0110\.011atL4→L5L4\{\\to\}L5and0\.0020\.002atL5→L6L5\{\\to\}L6; TFM drift reaches0\.0010\.001already atL2→L3L2\{\\to\}L3and0\.0000\.000thereafter\. Since both models stabilize atL5L5, we have chosen this layer to compare the representation kernels\. In the long\-ξ\\xiregime the TFM never cleanly stabilizes \(drift0\.0140\.014–0\.0600\.060throughL6L6\), so theL5L5estimate is approximate for that case\.
Table 3:Effective rankρeff\\rho\_\{\\mathrm\{eff\}\}across depth at the final checkpoint \(flatten mode\)\. Compression ratio=ρeff\(L0\)/ρeff\(L6\)=\\rho\_\{\\mathrm\{eff\}\}\(L\_\{0\}\)/\\rho\_\{\\mathrm\{eff\}\}\(L\_\{6\}\)\. CKA and relative Frobenius distance compare MLP and TFM kernels atL5L5\.In the following discussion of rank profiles and fixed\-point geometry, we refer to Table[3](https://arxiv.org/html/2607.15449#S4.T3)and Figure[1](https://arxiv.org/html/2607.15449#S4.F1)\.
#### Short\-ξ\\xi:
The MLP collapses monotonically \(8\.1×8\.1\\times\): short\-range noise is progressively integrated out with each layer\. The TFM first expands its effective rank atL1L1\(15\.5→20\.215\.5\\to 20\.2\) before collapsing it \(20\.2→2\.720\.2\\to 2\.7peak\-to\-end\), consistent with attention being an irrelevant perturbation: it transient structure due to attention atL1L1is subsequently integrated out\. Despite the geometric distance between the two final\-layer kernels \(CKA=0\.10=0\.10, relative Frobenius=0\.99=0\.99\), both models achieve the same loss \(Table[1](https://arxiv.org/html/2607.15449#S4.T1)\), so this geometric difference carries no functional significance: two models can represent the same information in geometrically unrelated ways\.
#### Long\-ξ\\xi:
The MLP compresses the effective rank monotonically toρeff≈1\.8\\rho\_\{\\mathrm\{eff\}\}\\approx 1\.8, as in the short\-ξ\\xicase, confirming that the MLP has no mechanism to exploit long\-range correlations regardless of regime\. The TFM behaves qualitatively differently: its attention mechanism computes a distinct linear combination of value vectors for each query positionii,
ai=∑jαijvj,αij=softmax\(qikj⊤d\),a\_\{i\}=\\sum\_\{j\}\\alpha\_\{ij\}v\_\{j\},\\qquad\\alpha\_\{ij\}=\\mathrm\{softmax\}\\\!\\left\(\\frac\{q\_\{i\}k\_\{j\}^\{\\top\}\}\{\\sqrt\{d\}\}\\right\),\(16\)and since tokens carry genuinely different information across positions, each outputaia\_\{i\}is a distinct linear combination of value vectors rather than a near\-constant vector\. The outputs\{ai\}\\\{a\_\{i\}\\\}therefore span a larger subspace than the pre\-attention representations: at layer 1, effective rank expands from12\.912\.9to25\.425\.4\. This elevated rank stabilizes \(ρeff≈25\\rho\_\{\\mathrm\{eff\}\}\\approx 25–2727\) across all subsequent layers, forming a high\-dimensional plateau absent in the MLP\. The representation kernels atL5L5are geometrically unrelated \(CKA=0\.34=0\.34, relative Frobenius=17\.3=17\.3\), but unlike the short\-ξ\\xicase this distance is meaningful: the two models no longer solve the same task\. The TFM achieves−0\.195\-0\.195nats belowH\(π\)H\(\\pi\)by exploiting long\-range context, while the MLP is stuck at\+0\.017\+0\.017nats aboveH\(π\)H\(\\pi\)\. These are genuinely distinct attractors, not perturbatively related fixed points\. Two fixed points can be geometrically distant yet perturbatively related if they share the same qualitative structure \(the same effective rank and the same task performance\)\. The long\-ξ\\xifixed points share neither: the MLP collapses toρeff≈1\.8\\rho\_\{\\mathrm\{eff\}\}\\approx 1\.8and remains nearH\(π\)H\(\\pi\), while the TFM stabilizes atρeff≈25\\rho\_\{\\mathrm\{eff\}\}\\approx 25and reaches−0\.195\-0\.195nats belowH\(π\)H\(\\pi\)\. These observations are consistent with the theoretical picture of Eq\. \([7](https://arxiv.org/html/2607.15449#S3.E7)\): the rank expansion and loss gap indicate a large fixed\-point shiftδ\\delta, consistent with a loss of perturbative control rather than a smooth perturbative shift\. Whether this is driven by a large driving terma\+ba\+b, near\-zero eigenvalues ofM∗M^\{\*\}, or both, cannot be determined from these measurements alone\.
#### RG interpretation:
The two regimes map directly onto the RG classification of Section[3\.3](https://arxiv.org/html/2607.15449#S3.SS3): short\-ξ\\xiattention behaves as an irrelevant perturbation \(fixed\-point shift small, MLP attractor preserved, same loss\); long\-ξ\\xiattention behaves as a relevant perturbation \(large shift, qualitatively different attractor, loss gap of0\.2120\.212nats\)\. Experiments 2–4 probe the internal structure of this transition through head ablation, scaling analysis, and perturbation decay measurements\.
Figure 1:Effective rankρeff\\rho\_\{\\mathrm\{eff\}\}across depth \(flatten mode\)\.Blue: short\-ξ\\xi; MLP \(solid\) and TFM \(dashed\) both collapse monotonically\.Red solid: MLP long\-ξ\\xicompresses toρeff≈1\.8\\rho\_\{\\mathrm\{eff\}\}\\approx 1\.8\.Red dashed: TFM long\-ξ\\xiexpands from12\.912\.9to25\.425\.4atL1L\_\{1\}and stabilizes at a high\-dimensional plateau \(ρeff≈25\\rho\_\{\\mathrm\{eff\}\}\\approx 25–2727\)\.Figure 2:Kernel driftΔCKA\(l,l\+1\)\\Delta\_\{\\mathrm\{CKA\}\}\(l,l\{\+\}1\)\(flatten mode\)\.Blue: short\-ξ\\xidrift distributed across early layers\.Red solid: MLP long\-ξ\\xidrift decays smoothly\.Red dashed: TFM long\-ξ\\xidrift concentrates at two transitions,L0→L1L\_\{0\}\{\\to\}L\_\{1\}\(entry\) andL5→L6L\_\{5\}\{\\to\}L\_\{6\}\(exit\), with near\-zero drift acrossL2L\_\{2\}–L5L\_\{5\}\(shaded plateau\)\.
## 5Experiment 2: Head Ablation
Formula \([7](https://arxiv.org/html/2607.15449#S3.E7)\) describes the global fixed\-point shift when attention is added to the entire network; it cannot be decomposed layer by layer without additional assumptions\. To understand which head \(at what layerll\) contributes most to the shift, we introduce a heuristic local extension and write
δ\(l\)=−Ml∗−1\(a\(l\)\+b\(l\)\)\+O\(‖δ\(l\)‖2\),\\delta^\{\(l\)\}=\-\{M^\{\*\}\_\{l\}\}^\{\-1\}\(a^\{\(l\)\}\+b^\{\(l\)\}\)\+O\(\\\|\\delta^\{\(l\)\}\\\|^\{2\}\),\(17\)where all quantities are layer specific,a\(l\)=Attn\(h\(l\)\)a^\{\(l\)\}=\\mathrm\{Attn\}\(h^\{\(l\)\}\),b\(l\)=MLP\(LN\(h\(l\)\+a\(l\)\)\)b^\{\(l\)\}=\\mathrm\{MLP\}\(\\mathrm\{LN\}\(h^\{\(l\)\}\+a^\{\(l\)\}\)\), andMl∗M^\{\*\}\_\{l\}is the Jacobian ofFFath\(l\)h^\{\(l\)\}\. This is motivated by \([7](https://arxiv.org/html/2607.15449#S3.E7)\) but is not implied by it\. A preliminary examination indicates which layer has the greatest contribution to the overall shift: atl=0l=0,h\(0\)h^\{\(0\)\}reflects raw token embeddings — no residual block has yet processed the input, so attention sees maximal positional variation anda\(0\)\+b\(0\)a^\{\(0\)\}\+b^\{\(0\)\}carries the strongest contextual signal\. Atl=1l=1the MLP has already begun integrating out positional variation, soa\(1\)a^\{\(1\)\}carries less positional signal and the driving term contributes less to the shift\. The drift profiles \(Table[2](https://arxiv.org/html/2607.15449#S4.T2)\) and rank jumps from Experiment 1 are consistent with this picture: drift drops by a factor of3636after the first layer in the short\-ξ\\xiTFM, and the largest rank jump likewise occurs atL0→L1L0\\to L1\(15\.5→20\.215\.5\\to 20\.2short\-ξ\\xi,12\.9→25\.412\.9\\to 25\.4long\-ξ\\xi\)\. These observations motivate Prediction[4](https://arxiv.org/html/2607.15449#Thmprediction4): the attention head at layer0, L0H0, should contribute the largest shift, with subsequent heads making smaller contributions\. To test this, we ablate each attention head in turn \(setting its output to zero\) and measure
ΔCKA\(l,h\)=CKA\(KMLP∗,KTFM,ablated∗\[l,h\]\)−CKA\(KMLP∗,KTFM,full∗\),\\Delta\_\{\\text\{CKA\}\}\(l,h\)=\\text\{CKA\}\(K^\{\*\}\_\{\\text\{MLP\}\},\\,K^\{\*\}\_\{\\text\{TFM,ablated\}\}\[l,h\]\)\-\\text\{CKA\}\(K^\{\*\}\_\{\\text\{MLP\}\},\\,K^\{\*\}\_\{\\text\{TFM,full\}\}\),\(18\)atL5L5, which we chose as the reference layer where both models have stabilized in the short\-ξ\\xiregime \(Table[2](https://arxiv.org/html/2607.15449#S4.T2)\)\. A large positiveΔCKA\(l,h\)\\Delta\_\{\\text\{CKA\}\}\(l,h\)means the ablated TFM is more similar to the MLP than the full TFM is: removing head\(l,h\)\(l,h\)has brought the TFM’s representations closer to the MLP fixed point, meaning that head was responsible for driving the TFM away from the MLP attractor\. A full RG classification \(relevant, marginal, irrelevant\) requires measuringΔCKA\\Delta\_\{\\text\{CKA\}\}at multiple values ofTTand fitting the scaling, which is carried out in Experiment 3; here we report only the relative contributions\.
Table 4:Head ablation results for the long\-ξ\\xiregime \(nheads=1n\_\{\\text\{heads\}\}=1, one head per layer\)\. Baseline CKA\(KMLP∗,KTFM,full∗\)=0\.32\(K^\{\*\}\_\{\\text\{MLP\}\},K^\{\*\}\_\{\\text\{TFM,full\}\}\)=0\.32\. Contributions are classified by magnitude ofΔCKA\\Delta\_\{\\text\{CKA\}\}only\.Figure 3:Head ablationΔCKA\(l,h\)\\Delta\_\{\\mathrm\{CKA\}\}\(l,h\)for the long\-ξ\\xiregime\.L0H0\(red\) is the dominant contributor \(ΔCKA=0\.119\\Delta\_\{\\mathrm\{CKA\}\}=0\.119\), more than4×4\\timesthe next contribution \(0\.0270\.027, L5H0\)\. Remaining heads \(blue\) make moderate to negligible contributions\.#### Results:
L0H0 is strongly dominant:ΔCKA=0\.119\\Delta\_\{\\text\{CKA\}\}=0\.119, more than4×4\\timesthe next\-best contribution \(0\.0270\.027, Table[4](https://arxiv.org/html/2607.15449#S5.T4)\)\. Heads L1H0, L2H0, L3H0, and L5H0 contribute0\.0200\.020–0\.0270\.027; L4H0 contributes0\.0080\.008and is effectively negligible\. This confirms Prediction[4](https://arxiv.org/html/2607.15449#Thmprediction4): the fixed\-point shift is concentrated atl=0l=0, where attention sees maximal positional variation before the MLP has begun integrating it out\. Once L0H0 has expanded effective rank atL0→L1L0\\to L1, subsequent heads operate on the high\-dimensional plateau where remaining positional structure is already captured and contribute negligible shifts\. This head\-specialisation pattern is consistent with probing studies that find syntactic and positional structure concentrated in early layers\(Voita et al\.,[2019](https://arxiv.org/html/2607.15449#bib.bib12); Clark et al\.,[2019](https://arxiv.org/html/2607.15449#bib.bib4)\), though those studies do not connect the pattern to fixed\-point structure\. The broader implications for the discrete RG phase space picture are discussed in Section[8\.1](https://arxiv.org/html/2607.15449#S8.SS1)\.
## 6Experiment 3: Attention Scaling and RG Observables
Experiments 1 and 2 characterised the fixed\-point shift via representation geometry at fixed context lengthT=64T=64, and identified L0H0 as the dominant contributor\. Neither experiment addressed whether any scalar statistic of the attention weight matricesA\(l,h\)∈ℝT×TA^\{\(l,h\)\}\\in\\mathbb\{R\}^\{T\\times T\}carries an RG signature\. If attention is a relevant operator in the long\-ξ\\xicase, statistics of these matrices should grow in complexity withTTin a way that is absent when attention is irrelevant\. The scaling exponentβ\\beta, defined bystat\(T\)∼Tβ\\mathrm\{stat\}\(T\)\\sim T^\{\\beta\}, captures this: a largeβ\\betasignals relevance, a smallβ\\betasignals irrelevance\. We measureβ\\betafor three candidate statistics acrossT∈\{16,32,64,128\}T\\in\\\{16,32,64,128\\\}\.
### 6\.1What is a genuine RG observable?
We note that the statistics measured here are properties ofA\(l,h\)A^\{\(l,h\)\}, an intermediate quantity internal to the attention mechanism rather than the fixed\-point shiftδ\\deltadirectly predicted by the perturbation framework \(Section[3\.2](https://arxiv.org/html/2607.15449#S3.SS2)\), so their scaling withTTmay reflect properties of the attention mechanism itself rather than its effect on the residual stream\. There is an inherent difficulty here: since rows ofA\(l,h\)A^\{\(l,h\)\}sum to11and haveTTentries, any statistic measuring how spread out the attention is will tend to grow withTTmechanically, even for a model that has learned to concentrate attention on a fixed number of positions\. We examine the viability of three candidate observables:Per\-row Shannon entropy:For headhhat layerll, positionii, and batchbb, the per\-row entropy is
Hb,i\(l,h\)\(T\)=−∑j=1TAb,i,j\(l,h\)logAb,i,j\(l,h\),Ab,i,j\(l,h\)=esb,ij\(l,h\)∑k=1Tesb,ik\(l,h\),sb,ij\(l,h\)=qb,i\(l,h\)kb,j\(l,h\)⊤/d\.H^\{\(l,h\)\}\_\{b,i\}\(T\)=\-\\sum\_\{j=1\}^\{T\}A^\{\(l,h\)\}\_\{b,i,j\}\\,\\log A^\{\(l,h\)\}\_\{b,i,j\},\\qquad A^\{\(l,h\)\}\_\{b,i,j\}=\\frac\{e^\{s^\{\(l,h\)\}\_\{b,ij\}\}\}\{\\sum\_\{k=1\}^\{T\}e^\{s^\{\(l,h\)\}\_\{b,ik\}\}\},\\quad s^\{\(l,h\)\}\_\{b,ij\}=q^\{\(l,h\)\}\_\{b,i\}\{\}^\{\\top\}k^\{\(l,h\)\}\_\{b,j\}/\\sqrt\{d\}\.\(19\)It ranges from0\(fully peaked\) tologT\\log T\(uniform\)\. We trackH\(T\)=⟨Hb,i\(l,h\)\(T\)⟩b,iH\(T\)=\\langle H^\{\(l,h\)\}\_\{b,i\}\(T\)\\rangle\_\{b,i\}and fitlogH\(T\)=βHlogT\+const\\log H\(T\)=\\beta\_\{H\}\\log T\+\\mathrm\{const\}\. Entropy fails automatically as an RG observable because it depends only on the attention weights, not on the value vectors at the attended positions: two attention patterns can have identical entropy regardless of whether the attended positions carry long\-range correlations or random noise\. This is confirmed empirically in Table[5](https://arxiv.org/html/2607.15449#S6.T5), whereβH≈0\.27\\beta\_\{H\}\\approx 0\.27in both regimes\.Participation ratio:For rowiiofAb\(l,h\)A^\{\(l,h\)\}\_\{b\}, the participation ratio measures how many tokens positioniieffectively attends to:
PRi\(l,h,b\)\(T\)=1∑j=1T\(Ab,ij\(l,h\)\)2\.\\mathrm\{PR\}^\{\(l,h,b\)\}\_\{i\}\(T\)=\\frac\{1\}\{\\sum\_\{j=1\}^\{T\}\\bigl\(A^\{\(l,h\)\}\_\{b,ij\}\\bigr\)^\{2\}\}\.\(20\)It ranges from11\(fully peaked, attending to one token\) toTT\(uniform, attending equally to all tokens\), withPR\(l,h\)\(T\)=1BT∑b,iPRi\(l,h,b\)\(T\)\\mathrm\{PR\}^\{\(l,h\)\}\(T\)=\\frac\{1\}\{BT\}\\sum\_\{b,i\}\\mathrm\{PR\}^\{\(l,h,b\)\}\_\{i\}\(T\)\. For this to be an RG observable, one would expectPR\\mathrm\{PR\}to grow more slowly withTTin the short\-ξ\\xicase \(where attention has little structure to exploit\) than in the long\-ξ\\xicase \(where attention should spread over more informative positions as context grows\)\. However,PR\\mathrm\{PR\}fails this test for a structural reason: A model attending uniformly tommfixed positions out ofTTwould ideally havePRi=m\\mathrm\{PR\}\_\{i\}=m, independent ofTT\. But softmax assigns nonzero weight to allTTpositions; asTTgrows the residual weights on the remainingT−mT\-mpositions do not vanish fast enough to keep∑jAij2\\sum\_\{j\}A\_\{ij\}^\{2\}constant, soPRi\\mathrm\{PR\}\_\{i\}grows withTTeven though the model’s behaviour has not changed\. Consequently, any observedβPR≈1\\beta\_\{\\mathrm\{PR\}\}\\approx 1, as in Table[5](https://arxiv.org/html/2607.15449#S6.T5)for both regimes, is a softmax artifact, not an RG signal, and the participation ratio is disqualified as an RG observable\.Effective rank:As previously defined, forAb\(l,h\)∈ℝT×TA^\{\(l,h\)\}\_\{b\}\\in\\mathbb\{R\}^\{T\\times T\}with singular valuesσ1\(l,h,b\)≥⋯≥σT\(l,h,b\)\\sigma\_\{1\}^\{\(l,h,b\)\}\\geq\\cdots\\geq\\sigma\_\{T\}^\{\(l,h,b\)\}, the effective rank is
ρeff\(l,h,b\)\(T\)=exp\(−∑k=1Tpk\(l,h,b\)logpk\(l,h,b\)\),pk\(l,h,b\)=σk\(l,h,b\)∑jσj\(l,h,b\),\\rho\_\{\\mathrm\{eff\}\}^\{\(l,h,b\)\}\(T\)=\\exp\\\!\\left\(\-\\sum\_\{k=1\}^\{T\}p\_\{k\}^\{\(l,h,b\)\}\\log p\_\{k\}^\{\(l,h,b\)\}\\right\),\\qquad p\_\{k\}^\{\(l,h,b\)\}=\\frac\{\\sigma\_\{k\}^\{\(l,h,b\)\}\}\{\\sum\_\{j\}\\sigma\_\{j\}^\{\(l,h,b\)\}\},\(21\)withρeff\(l,h\)\(T\)=⟨ρeff\(l,h,b\)\(T\)⟩b\\rho\_\{\\mathrm\{eff\}\}^\{\(l,h\)\}\(T\)=\\langle\\rho\_\{\\mathrm\{eff\}\}^\{\(l,h,b\)\}\(T\)\\rangle\_\{b\}\. The exponentβρ\\beta\_\{\\rho\}measures how many effective dimensions the attention pattern occupies as context size grows\. Effective rank avoids the problems of entropy and participation ratio: it is computed from the singular values ofA\(l,h\)A^\{\(l,h\)\}, which are not constrained by softmax normalization, and captures the global low\-rank structure of the attention matrix rather than the concentration of individual rows\. For this reason, we will use effective rank as the primary RG observable, and include entropy and participation ratio only for comparison\. We fitlog\(stat\)=βlogT\+const\\log\(\\mathrm\{stat\}\)=\\beta\\log T\+\\mathrm\{const\}at each layer and report results in Table[5](https://arxiv.org/html/2607.15449#S6.T5)\.
### 6\.2Results
Table 5:Scaling exponentsβ\\betaper layer for each statistic, fitted fromstat\(T\)∼Tβ\\mathrm\{stat\}\(T\)\\sim T^\{\\beta\}acrossT∈\{16,32,64,128\}T\\in\\\{16,32,64,128\\\}\(nheads=1n\_\{\\mathrm\{heads\}\}=1, allR2≥0\.99R^\{2\}\\geq 0\.99\)\.Entropy does not discriminate: both regimes yieldβH≈0\.27\\beta\_\{H\}\\approx 0\.27at every layer, confirming that entropy is insensitive to whether attention is relevant or irrelevant\. Participation ratio showsβPR≈0\.86\\beta\_\{\\mathrm\{PR\}\}\\approx 0\.86–1\.001\.00in both regimes, consistent with the softmax artifact identified above\. Effective rank is the only statistic that discriminates\. In the short\-ξ\\xicaseβρ\\beta\_\{\\rho\}falls from0\.1990\.199at L0 to near zero at L3–L5; in the long\-ξ\\xicase it remains elevated \(0\.3060\.306–0\.4270\.427\) at all layers\. The differenceΔβρ≈0\.2\\Delta\\beta\_\{\\rho\}\\approx 0\.2–0\.40\.4is the clearest RG signature in the attention patterns\. The per\-layer variation inβρ\\beta\_\{\\rho\}reflects the changing role of attention as representations are compressed toward the fixed point; unlike continuum RG where scaling exponents are fixed at the fixed point, in our discrete setting they vary along the flow because layers are not weight\-tied\. Experiment 4 probes the fixed\-point structure more directly by measuring perturbation decay across Markov modes\.
## 7Experiment 4: Perturbation Decay Spectrum
Experiment 2 showed that the network’s response to attention is not uniform across heads\. Here we ask a complementary question at the level of the input spectrum: has the network internalized the spectral structure of the Markov chainPP? Specifically, does it treat slow Markov modes \(which carry long\-range correlations\) differently from fast modes \(which carry short\-range noise\)? Prediction[3](https://arxiv.org/html/2607.15449#Thmprediction3)expects that it does: the Transformer should preserve perturbations along slow modes \(large\|λk\|\|\\lambda\_\{k\}\|, longξk\\xi\_\{k\}\) and suppress perturbations along fast modes \(small\|λk\|\|\\lambda\_\{k\}\|\), because attention gives it access to the full sequence and thus the ability to exploit long\-range structure\. The MLP, which processes each token independently, should be mode\-blind, suppressing all modes at the same rate\. We test this by injecting perturbations aligned with eigenvectorsϕk\\phi\_\{k\}ofPPand tracking their decay across layers\. We take a batch of sequences, represent each token as a one\-hot vectorx∈ℝT×Vx\\in\\mathbb\{R\}^\{T\\times V\}, and perturb positiont=0t=0, whereϵ=0\.3\\epsilon=0\.3:
x0,:′=x0,:\+ϵϕk,xt,:′=xt,:fort\>0\.x^\{\\prime\}\_\{0,:\}=x\_\{0,:\}\+\\epsilon\\,\\phi\_\{k\},\\qquad x^\{\\prime\}\_\{t,:\}=x\_\{t,:\}\\;\\text\{for\}\\;t\>0\.\(22\)Bothxxandx′x^\{\\prime\}are passed through the same trained model, producing hidden statesh\(l\)h^\{\(l\)\}andh′\(l\)h^\{\\prime\(l\)\}at each layer\. We track how the perturbation evolves across layers via the cosine distance between base and perturbed representations, averaged over the batch:
dk\(l\)=1−⟨h0\(l\)⋅h0′\(l\)‖h0\(l\)‖‖h0′\(l\)‖⟩B\.d\_\{k\}\(l\)=1\-\\left\\langle\\frac\{h^\{\(l\)\}\_\{0\}\\cdot h^\{\\prime\(l\)\}\_\{0\}\}\{\\\|h^\{\(l\)\}\_\{0\}\\\|\\,\\\|h^\{\\prime\(l\)\}\_\{0\}\\\|\}\\right\\rangle\_\{\\\!B\}\.\(23\)We use cosine rather than Euclidean distance because residual connections accumulate norm with depth: the base representationh0\(l\)h^\{\(l\)\}\_\{0\}grows in magnitude across layers, which would inflate‖h0\(l\)−h0′\(l\)‖\\\|h^\{\(l\)\}\_\{0\}\-h^\{\\prime\(l\)\}\_\{0\}\\\|independently of whether the perturbation is being suppressed\. Cosine distance isolates the angular separation between base and perturbed representations, which reflects only the directional deviation of the perturbation and is insensitive to this norm growth\. A cosine distance above11would indicate that the perturbed representation has rotated more than90∘90^\{\\circ\}from the base, suggesting the perturbation has driven the trajectory out of the linear regime; we verify thatdk\(l\)≤1d\_\{k\}\(l\)\\leq 1for all retained fits\. The stability equation \([11](https://arxiv.org/html/2607.15449#S3.E11)\) predicts that perturbations along irrelevant eigenmodes ofM~∗\\tilde\{M\}^\{\*\}decay across layers\. Sinceϕk\\phi\_\{k\}are eigenvectors ofPPrather thanM~∗\\tilde\{M\}^\{\*\}, each perturbation alongϕk\\phi\_\{k\}excites a mixture of network eigenmodes \(see Appendix[Appendix E](https://arxiv.org/html/2607.15449#A5)\)\. This is a limitation of the experimental design: the fittedτk\\tau\_\{k\}
dk\(l\)≈Ake−l/τkd\_\{k\}\(l\)\\approx A\_\{k\}\\,e^\{\-l/\\tau\_\{k\}\}\(24\)to the observed cosine distance profile across layers \(see Appendix[Appendix D](https://arxiv.org/html/2607.15449#A4)\) is an effective decay rate for the directionϕk\\phi\_\{k\}rather than a single\-eigenmode quantity, which weakens the test of the monotone ordering predicted by Prediction[3](https://arxiv.org/html/2607.15449#Thmprediction3)\. A cleaner test would inject perturbations along the eigenvectors ofM~∗\\tilde\{M\}^\{\*\}directly, but this is not currently tractable \(see Limitations below\)\. We retain only fits withR2\>0\.75R^\{2\}\>0\.75; also modes whose profiles show a flat\-then\-collapse or an initial transient rise lie outside the linearized regime and are excluded\. Whether the monotone ordering of Prediction[3](https://arxiv.org/html/2607.15449#Thmprediction3)holds depends on how the learned input projectionWWmaps Markov eigenvectors into the eigenspace ofM~∗\\tilde\{M\}^\{\*\}; sinceWWis determined by training rather than by the structure ofPPorM~∗\\tilde\{M\}^\{\*\}alone, this alignment cannot be predicted from the theory and must be verified empirically\.
Table 6:Perturbation decay spectrum for bothξ\\xiregimes \(selected modes\)\.ξk=−1/log\|λk\|\\xi\_\{k\}=\-1/\\log\|\\lambda\_\{k\}\|is the correlation length of eigenmodekk\. Fits withR2<0\.75R^\{2\}<0\.75are marked —\. In the short\-ξ\\xiregimeτTFM<τMLP\\tau\_\{\\text\{TFM\}\}<\\tau\_\{\\text\{MLP\}\}for every mode where both fits are reliable, a reversal of the long\-ξ\\xiordering\.#### Long\-ξ\\xi: Transformer is spectrally selective, MLP is not\.
In the long\-ξ\\xiregime, the MLP treats all modes similarly:τMLP\\tau\_\{\\text\{MLP\}\}ranges only from1\.781\.78to2\.322\.32across all reliable modes, a dynamic range of1\.3×1\.3\\times, with Pearson correlationr=0\.553r=0\.553between\|λk\|\|\\lambda\_\{k\}\|andτMLP\\tau\_\{\\text\{MLP\}\}\. The MLP suppresses slow and fast modes at nearly the same rate\. The Transformer behaves differently:τTFM\\tau\_\{\\text\{TFM\}\}ranges from1\.271\.27to6\.846\.84, a dynamic range of5\.4×5\.4\\times, withr=0\.435r=0\.435\. The slowest mode \(\|λ1\|=0\.862\|\\lambda\_\{1\}\|=0\.862\) persists for6\.846\.84layers before decaying, while the fastest reliable mode decays within1\.271\.27layers\. The Transformer strongly differentiates between slow and fast modes, preserving the former and suppressing the latter in a way the MLP does not \(Figure[4](https://arxiv.org/html/2607.15449#S7.F4)\)\.
#### Short\-ξ\\xi: regime reversal\.
In the short\-ξ\\xiregime, the eigenvalues ofPPare small, reflecting the short\-range noise carried by all modes\. The central finding is that the Transformer suppresses perturbations*faster*than the MLP \(τTFM<τMLP\\tau\_\{\\text\{TFM\}\}<\\tau\_\{\\text\{MLP\}\}for every mode where both fits are reliable\)\. This is a reversal of the long\-ξ\\xiordering in which the Transformer preserved slow modes by decaying them more slowly\. Note that despite slower modes decaying somewhat more slowly in comparison to faster ones, within the Transformer itself \(Pearson correlationrTFM=0\.55r\_\{\\text\{TFM\}\}=0\.55\), this does not reflect useful selectivity: all Transformer decay times are shorter than their MLP counterparts, so the positiverrmerely reflects that even aggressive suppression is not perfectly uniform across modes\. Mode selectivity, preserving slow modes while discarding fast ones, requires not just the Transformer architecture but a regime in which attention is a relevant operator, i\.e\. one in which there is long\-range structure worth preserving\. When there is none \(short\-ξ\\xi\), the Transformer uses its additional capacity to integrate out perturbations more efficiently than the MLP rather than to be spectrally selective\.
#### Limitations\.
The moderate Pearson correlations \(rMLP=0\.553r\_\{\\text\{MLP\}\}=0\.553,rTFM=0\.435r\_\{\\text\{TFM\}\}=0\.435for long\-ξ\\xi\) indicate that the predicted monotone ordering ofτk\\tau\_\{k\}with\|λk\|\|\\lambda\_\{k\}\|is not cleanly recovered\. Two factors contribute to this issue: First, the sample is small: with onlyV−1=15V\-1=15modes andL=6L=6layers, the exponential fits carry substantial uncertainty\. Second, and more fundamentally, the injected perturbations are alongϕk\\phi\_\{k\}, the eigenvectors ofPP, not ofM~∗\\tilde\{M\}^\{\*\}\. As shown in Appendix[Appendix E](https://arxiv.org/html/2607.15449#A5), perturbations alongϕk\\phi\_\{k\}excite a*mixture*of network eigenmodes simultaneously\. The fittedτk\\tau\_\{k\}is an effective decay rate for this mixture, not a property of a single eigenmode\. Therefore, theτk\\tau\_\{k\}values need not order monotonically with\|λk\|\|\\lambda\_\{k\}\|even if the underlying network eigenmodes do\. This weakens the test: even if the network’s eigenmodes are perfectly ordered by eigenvalues\|λk\|\|\\lambda\_\{k\}\|ofPP, the mixture meansτk\\tau\_\{k\}need not reflect that ordering cleanly\. Note that injectingϕk\\phi\_\{k\}is the correct design for testing Prediction[3](https://arxiv.org/html/2607.15449#Thmprediction3), which asks whether the network has internalized the spectral structure ofPP, not an approximation to some cleaner experiment\. Nevertheless, the regime\-level contrast is robust: the Transformer is more spectrally selective than the MLP in the long\-ξ\\xiregime \(a dynamic range of5\.4×5\.4\\timeswhich shows that the Transformer strongly differentiates between slow and fast modes\), and more suppressive in the short\-ξ\\xiregime \(e\.g\. modes 1–2:τTFM=0\.95\\tau\_\{\\text\{TFM\}\}=0\.95vs\.τMLP=2\.22\\tau\_\{\\text\{MLP\}\}=2\.22; mode 15:τTFM=0\.60\\tau\_\{\\text\{TFM\}\}=0\.60vs\.τMLP=3\.45\\tau\_\{\\text\{MLP\}\}=3\.45\), and it is confirmed here\.
Figure 4:Perturbation decay lengthτk\\tau\_\{k\}vs\. eigenvalue magnitude\|λk\|\|\\lambda\_\{k\}\|for each eigenmodeϕk\\phi\_\{k\}ofPP\.Left \(long\-ξ\\xi\):MLP \(blue circles,r=0\.55r=0\.55\) spansτ∈\[1\.78,2\.32\]\\tau\\in\[1\.78,2\.32\]— narrow range, weakly selective\. TFM \(red diamonds,r=0\.44r=0\.44\) spansτ∈\[1\.27,6\.84\]\\tau\\in\[1\.27,6\.84\]— strong selectivity, slow modes persist5×5\\timeslonger than fast modes\.Right \(short\-ξ\\xi\):ordering reverses;τTFM<τMLP\\tau\_\{\\text\{TFM\}\}<\\tau\_\{\\text\{MLP\}\}for every mode — the Transformer acts as a fast integrator when attention is irrelevant\.
## 8Discussion
The RG framework generates four falsifiable predictions, all of which are confirmed, albeit the fourth at the level of a regime contrast\.Prediction[1](https://arxiv.org/html/2607.15449#Thmprediction1)\(short\-ξ\\xi, irrelevance\):Experiment 1 confirms this at the level of fixed\-point geometry and task performance: both models converge to within±0\.004\\pm 0\.004nats ofH\(π\)H\(\\pi\), CKA between final\-layer representations is0\.100\.10, and effective rank profiles collapse monotonically\. Experiment 4 reveals a subtler effect: attention still contracts perturbations faster than the MLP, but without spectral selectivity\.Prediction[2](https://arxiv.org/html/2607.15449#Thmprediction2)\(long\-ξ\\xi, relevance\):This is confirmed in its strongest form by CKA=0\.34=0\.34, relative Frobenius distance=17\.3=17\.3, and rank expansion from12\.912\.9to25\.425\.4at layer 1\. These are not small deviations but a discontinuous transition to a qualitatively different fixed\-point structure, consistent with the loss of perturbative control dictated by the near\-singularity ofM∗M^\{\*\}\. Formula \([7](https://arxiv.org/html/2607.15449#S3.E7)\) is a first\-order expansion valid whenM∗M^\{\*\}stays away from singularity; in the long\-ξ\\xiregime this condition fails, the resolvent diverges, and the system is driven out of the perturbative regime entirely\. Ironically, the most dramatic confirmation of the RG prediction,a phase transition, is precisely the case where the perturbation formula breaks down; a full description requires the non\-perturbative RG flow\.Prediction[4](https://arxiv.org/html/2607.15449#Thmprediction4)\(layer specificity\):L0H0 accounts forΔCKA=0\.119\\Delta\_\{\\text\{CKA\}\}=0\.119, more than4×4\\timesthe next\-best contribution\. The fixed\-point shift is concentrated at the first layer, where attention sees maximal positional variation before the MLP has begun integrating it out\. This dominance is consistent with the structural prediction of Section[3](https://arxiv.org/html/2607.15449#S3), but admits an alternative reading: the first\-layer head sees the highest\-entropy input and the most positional variation to exploit, while later layers see partially compressed representations\. Whether L0H0 dominance is a genuine RG effect or a structural consequence of architectural ordering cannot be definitively distinguished here\.Prediction[3](https://arxiv.org/html/2607.15449#Thmprediction3)\(mode selectivity\):Confirmed at the regime level, with a regime reversal as the key finding\. In the long\-ξ\\xiregime the TFM spansτTFM∈\[1\.27,6\.84\]\\tau\_\{\\text\{TFM\}\}\\in\[1\.27,6\.84\], a dynamic range of5\.4×5\.4\\timesvs\.1\.3×1\.3\\timesfor the MLP, selectively preserving slow Markov modes\. In the short\-ξ\\xiregime the TFM suppresses all modes within one to two layers \(τTFM∈\[0\.60,1\.39\]\\tau\_\{\\text\{TFM\}\}\\in\[0\.60,1\.39\]\), faster than the MLP for every mode, with no spectral selectivity\. This reversal is a direct signature of the difference between the relevant and irrelevant regimes: mode selectivity emerges specifically when attention is a relevant operator\. The moderate Pearson correlations \(rMLP=0\.553r\_\{\\text\{MLP\}\}=0\.553,rTFM=0\.435r\_\{\\text\{TFM\}\}=0\.435for long\-ξ\\xi\) are expected: as derived in Appendix[Appendix E](https://arxiv.org/html/2607.15449#A5), each injected perturbationϕk\\phi\_\{k\}excites a mixture ofM∗\{M\}^\{\*\}eigenmodes, soτk\\tau\_\{k\}is an effective decay rate rather than a single\-eigenmode quantity\. A clean monotone ordering would requireWϕkW\\phi\_\{k\}to align predominantly with a single eigenmode ofM∗\{M\}^\{\*\}whose decay rate tracks\|λk\|\|\\lambda\_\{k\}\|, a strong condition on learned structure that need not hold exactly\.
### 8\.1RG interpretation of the results
The ablation results \(Experiment 2\) give a concrete picture of how the trajectory commits to an attractor\. With L0H0 active, the trajectory is deflected atl=0l=0into the basin of the high\-rank TFM attractor; with L0H0 ablated, it falls back toward the MLP basin\. Which attractor is reached is determined at the first coarse\-graining step: the basin boundary is crossed, if at all, in a single layer\. Subsequent layers do not redirect the trajectory \(as discussed in the theory section, this is a trajectory in representation and not coupling space\): the system arrives near the fixed point and each further coarse\-graining step leaves it there, consistent with the near\-zero drift observed acrossL2L2–L5L5\(Table[2](https://arxiv.org/html/2607.15449#S4.T2)\)\. The perturbation decay results \(Experiment 4\) are consistent with an RG interpretation: a network with near\-marginal operators would preserve slow Markov modes and suppress fast ones, which is what we observe in the long\-ξ\\xiregime\. In the short\-ξ\\xiregime, all modes contract quickly regardless of their Markov eigenvalue, consistent with all network operators being strongly irrelevant\. Together, the ablation and perturbation decay results illuminate two complementary aspects of the RG picture\. Experiment 2 shows that which attractor the trajectory commits to is decided at the first coarse\-graining step: L0H0 either deflects the trajectory into the high\-rank basin or leaves it in the MLP basin, and subsequent layers merely contract toward whichever fixed point was selected\. Experiment 4 shows that once near the fixed point, the contraction structure, which operators are near\-marginal and which are strongly irrelevant, shapes how different input distributions are processed across depth\.
## 9Conclusion
We have studied attention as a perturbation of the MLP fixed point established in\(Haggi\-Mani & Rish,[2026](https://arxiv.org/html/2607.15449#bib.bib5)\), using matched architectures evaluated on synthetic Markov chain sequences with controlled correlation lengthξ\\xi\. Four findings support the validity of the RG interpretation\. First, attention is irrelevant for short\-ξ\\xichains: both models converge to the same loss and representational geometry, and the fixed\-point plateau structure is preserved\. Second, attention is a relevant operator for long\-ξ\\xichains\. WhenM∗M^\{\*\}has near\-zero eigenvalues, the shiftδ=−M∗−1\(a\+b\)\\delta=\-\{M^\{\*\}\}^\{\-1\}\(a\+b\)diverges and the system is driven out of the perturbative regime; the observed transition to a high\-dimensional \(ρeff≈25\\rho\_\{\\mathrm\{eff\}\}\\approx 25\) fixed point, with a loss0\.1950\.195nats belowH\(π\)H\(\\pi\), is its empirical signature\. The two fixed points are geometrically unrelated \(CKA=0\.34=0\.34, relative Frobenius distance=17\.3=17\.3\) and are not perturbatively connected: attention does not produce a small correction to the MLP fixed\-point structure but drives a transition to a qualitatively different one\. Third, the dominant relevant operator is L0H0, accounting for more than4×4\\timesthe representational shift of any subsequent head\. The fixed\-point shift is concentrated at the first layer, where attention sees maximal positional variation before the MLP has begun integrating it out; subsequent heads operate on the plateau and contribute negligible shifts\. Fourth, Experiment 4 reveals a regime reversal in perturbation decay: in the long\-ξ\\xiregime the TFM selectively preserves slow Markov modes \(τTFM∈\[1\.27,6\.84\]\\tau\_\{\\text\{TFM\}\}\\in\[1\.27,6\.84\], a dynamic range of5\.4×5\.4\\timesvs\.1\.3×1\.3\\timesfor the MLP\); in the short\-ξ\\xiregime it suppresses all modes faster than the MLP \(τTFM∈\[0\.60,1\.39\]\\tau\_\{\\text\{TFM\}\}\\in\[0\.60,1\.39\]\), with no spectral selectivity\. Together with\(Haggi\-Mani & Rish,[2026](https://arxiv.org/html/2607.15449#bib.bib5)\), these results constitute the first quantitative, controlled evidence that Transformers implement qualitatively different representational dynamics depending on the spectral structure of the input distribution, and that a rigorous RG perturbation framework provides a predictive account of that difference\. The present work is deliberately minimal: a single attention head, synthetic sequences, and a controlled Markov input distribution\. In future work, we will extend the framework to full Transformer architectures at scale, where the RG flow involves many heads, many layers, and a rich natural\-language input distribution\. The phase\-space picture developed here \(fixed\-point attractors, relevant operators driving transitions between fixed\-point structures, irrelevant operators integrated out without changing the attractor\) provides the conceptual scaffolding for identifying sharp depth\-wise transitions in a real language model\. It remains to be seen whether these mechanisms operate at scale\.
## Appendix Appendix A\(Stability Condition for the MLP Fixed\-Point\)
Since each MLP blockFMLP\(x\)=x\+f\(x\)=x\+MLP\(LN\(x\)\)F\_\{\\text\{MLP\}\}\(x\)=x\+f\(x\)=x\+\\mathrm\{MLP\}\(\\mathrm\{LN\}\(x\)\), satisfiesFMLP\(x∗\)=x∗F\_\{\\text\{MLP\}\}\(x^\{\*\}\)=x^\{\*\}at the fixed point, we have
f\(x∗\)=0\.f\(x^\{\*\}\)=0\.\(25\)Stability of the MLP fixed\-point:Linearizing the MLP block map aroundx∗x^\{\*\}withx0=x∗\+ϵ0x\_\{0\}=x^\{\*\}\+\\epsilon\_\{0\}, gives
x1=FMLP\(x∗\+ϵ0\)=x∗\+ϵ0\+f\(x∗\+ϵ0\)\.x\_\{1\}=F\_\{\\text\{MLP\}\}\(x^\{\*\}\+\\epsilon\_\{0\}\)=x^\{\*\}\+\\epsilon\_\{0\}\+f\(x^\{\*\}\+\\epsilon\_\{0\}\)\.\(26\)Taylor\-expandingffaroundx∗x^\{\*\}and usingf\(x∗\)=0f\(x^\{\*\}\)=0:
f\(x∗\+ϵ0\)=f\(x∗\)\+D\[f\]\(x∗\)ϵ0\+O\(‖ϵ0‖2\)=D\[f\]\(x∗\)ϵ0\+O\(‖ϵ0‖2\),f\(x^\{\*\}\+\\epsilon\_\{0\}\)=f\(x^\{\*\}\)\+D\[f\]\(x^\{\*\}\)\\,\\epsilon\_\{0\}\+O\(\\\|\\epsilon\_\{0\}\\\|^\{2\}\)=D\[f\]\(x^\{\*\}\)\\epsilon\_\{0\}\+O\(\\\|\\epsilon\_\{0\}\\\|^\{2\}\),\(27\)we get
ϵ1=\(I\+D\[f\]\(x∗\)\)ϵ0,\\epsilon\_\{1\}=\(I\+D\[f\]\(x^\{\*\}\)\)\\,\\epsilon\_\{0\},\(28\)where we have definedϵ1=x1−x∗\\epsilon\_\{1\}=x\_\{1\}\-x^\{\*\}: Iterating this mapnntimes, and defining the JacobianM=D\[f\]\(x∗\)M=D\[f\]\(x^\{\*\}\):
ϵn=\(I\+M\)nϵ0,\\epsilon\_\{n\}=\(I\+\{M\}\)^\{n\}\\,\\epsilon\_\{0\},\(29\)which in the eigenbasis of the JacobianM\{M\}with eigenvaluesμk\\mu\_\{k\}gives
\[ϵn\]k=\(1\+μk\)n\[ϵ0\]k\.\[\\epsilon\_\{n\}\]\_\{k\}=\(1\+\\mu\_\{k\}\)^\{n\}\\,\[\\epsilon\_\{0\}\]\_\{k\}\.\(30\)The MLP fixed\-point is stable if and only if all eigenvalues ofI\+MI\+\{M\}lie strictly inside the unit circle:
\|1\+μk\|<1for allk\.\|1\+\\mu\_\{k\}\|<1\\quad\\text\{for all \}k\.\(31\)For real eigenvalues this reduces to−2<μk<0\-2<\\mu\_\{k\}<0: the MLP sub\-network must contract perturbations atx∗x^\{\*\}, but not so strongly as to overshoot\. Modes withμk\>0\\mu\_\{k\}\>0orμk<−2\\mu\_\{k\}<\-2grow under iteration \(unstable eigenvector directions\)\. The fixed\-point plateau observed empirically in bothξ\\xiregimes implies that the trained MLP satisfies this condition across all modes, with the low effective rank \(ρeff≈1\.8\\rho\_\{\\mathrm\{eff\}\}\\approx 1\.8\) of the attractor reflecting that most directions in representation space have been contracted\.
## Appendix Appendix B\(Deriving the Fixed\-Point Shift Formula\)
At the MLP fixed\-pointx∗x^\{\*\}, the MLP blockFMLP\(x\)=x\+f\(x\)F\_\{\\text\{MLP\}\}\(x\)=x\+f\(x\), satisfiesFMLP\(x∗\)=x∗F\_\{\\text\{MLP\}\}\(x^\{\*\}\)=x^\{\*\}, hence
f\(x∗\)=MLP\(LN\(x∗\)\)=0\.f\(x^\{\*\}\)=\\mathrm\{MLP\}\(\\mathrm\{LN\}\(x^\{\*\}\)\)=0\.\(32\)The Transformer block adds an attention sublayer𝒜\(x\)\\mathcal\{A\}\(x\)before the MLP:
FTFM\(x\)=x\+𝒜\(x\)\+f\(x\+𝒜\(x\)\)\.F\_\{\\text\{TFM\}\}\(x\)=x\+\\mathcal\{A\}\(x\)\+f\\\!\\left\(x\+\\mathcal\{A\}\(x\)\\right\)\.\(33\)Suppose this addition moves the MLP fixed\-point by a small amountδ\\deltato a fixed\-pointx~∗\\tilde\{x\}^\{\*\}of the Transformer:x~∗=x∗\+δ\\tilde\{x\}^\{\*\}=x^\{\*\}\+\\delta\. The transformer’s fixed point conditionFTFM\(x~∗\)=x~∗F\_\{\\text\{TFM\}\}\(\\tilde\{x\}^\{\*\}\)=\\tilde\{x\}^\{\*\}implies
𝒜\(x~∗\)\+f\(x~∗\+𝒜\(x~∗\)\)=0\.\\mathcal\{A\}\(\\tilde\{x\}^\{\*\}\)\+f\\\!\\left\(\\tilde\{x\}^\{\*\}\+\\mathcal\{A\}\(\\tilde\{x\}^\{\*\}\)\\right\)=0\.\(34\)Definea≡𝒜\(x∗\)a\\equiv\\mathcal\{A\}\(x^\{\*\}\)and expand𝒜\(x~∗\)\\mathcal\{A\}\(\\tilde\{x\}^\{\*\}\)to the first order inδ\\delta:
𝒜\(x~∗\)=𝒜\(x∗\+δ\)=a\+D\[a\]δ\+O\(‖δ‖2\)\.\\mathcal\{A\}\(\\tilde\{x\}^\{\*\}\)=\\mathcal\{A\}\(x^\{\*\}\+\\delta\)=a\+D\[a\]\\delta\+O\(\\\|\\delta\\\|^\{2\}\)\.\(35\)The argument offfis therefore
x~∗\+𝒜\(x~∗\)=x∗\+a\+\(I\+D\[a\]\)δ\+O\(‖δ‖2\),\\tilde\{x\}^\{\*\}\+\\mathcal\{A\}\(\\tilde\{x\}^\{\*\}\)=x^\{\*\}\+a\+\(I\+D\[a\]\)\\delta\+O\(\\\|\\delta\\\|^\{2\}\),\(36\)Here,ffcan be expanded either aroundx∗x^\{\*\}orx∗\+ax^\{\*\}\+a\. Our choice of the latter requires only the residualδ\\deltato be small: the attention outputa=𝒜\(x∗\)a=\\mathcal\{A\}\(x^\{\*\}\)of the network is treated exactly, so the approximation remains valid regardless of the magnitude ofaa\. Expanding aroundx∗x^\{\*\}is also consistent, but requires the additional assumption that‖a‖\\\|a\\\|is of the same order of magnitude asδ\\delta, since the error terms would includeO\(‖a‖2\+‖a‖‖δ‖\+‖δ‖2\)O\(\\\|a\\\|^\{2\}\+\\\|a\\\|\\\|\\delta\\\|\+\\\|\\delta\\\|^\{2\}\)\. We should also note that treating attention as a perturbation in this paper is in the dynamical systems and RG sense as we ask whether it remains relevant or irrelevant under iterations\. This is a qualitative classification, not a statement that requires‖a‖\\\|a\\\|to be small\. Our only quantitative assumption is that the new fixed\-point is in the proximity of the MLP fixed\-point\. Definef\(x∗\+a\)=bf\(x^\{\*\}\+a\)=b, and expandffaroundx∗\+ax^\{\*\}\+a:
f\(x∗\+a\+\(I\+D\[a\]\)δ\)=b\+D\[b\]\(I\+D\[a\]\)δ\+O\(‖δ‖2\)\.f\\\!\\left\(x^\{\*\}\+a\+\(I\+D\[a\]\)\\delta\\right\)=b\+D\[b\]\(I\+D\[a\]\)\\delta\+O\(\\\|\\delta\\\|^\{2\}\)\.\(37\)Substituting into the fixed point condition and retaining allO\(δ\)O\(\\delta\)terms
a\+D\[a\]δ\+b\+D\[b\]\(I\+D\[a\]\)δ=0a\+D\[a\]\\delta\+b\+D\[b\]\(I\+D\[a\]\)\\delta=0\(38\)we get
\(a\+b\)\+\[D\[b\]\+D\[a\]\+D\[b\]D\[a\]\]δ=0,\\bigl\(a\+b\\bigr\)\+\\bigl\[D\[b\]\+D\[a\]\+D\[b\]D\[a\]\\bigr\]\\delta=0,\(39\)and
δ=−M∗−1\(a\+b\)\.\\delta=\-\{M^\{\*\}\}^\{\-1\}\\bigl\(a\+b\\bigr\)\.\(40\)whereM∗=D\[b\]\+D\[a\]\+D\[b\]D\[a\]\{M^\{\*\}\}=D\[b\]\+D\[a\]\+D\[b\]D\[a\]\.
## Appendix Appendix C\(Stability of the TFM Fixed\-Point\)
The Transformer layer map
FTFM\(x\)=x\+𝒜\(x\)\+f\(x\+𝒜\(x\)\)F\_\{\\text\{TFM\}\}\(x\)=x\+\\mathcal\{A\}\(x\)\+f\\\!\\left\(x\+\\mathcal\{A\}\(x\)\\right\)\(41\)satisfiesFTFM\(x~∗\)=x~∗F\_\{\\text\{TFM\}\}\(\\tilde\{x\}^\{\*\}\)=\\tilde\{x\}^\{\*\}, or equivalentlya~\+f\(x~∗\+a~\)=0\\tilde\{a\}\+f\(\\tilde\{x\}^\{\*\}\+\\tilde\{a\}\)=0\. We apply this map at
x1=FTFM\(x~∗\+ϵ0\)=x~∗\+ϵ0\+𝒜\(x~∗\+ϵ0\)\+f\(x~∗\+ϵ0\+𝒜\(x~∗\+ϵ0\)\),x\_\{1\}=F\_\{\\text\{TFM\}\}\(\\tilde\{x\}^\{\*\}\+\\epsilon\_\{0\}\)=\\tilde\{x\}^\{\*\}\+\\epsilon\_\{0\}\+\\mathcal\{A\}\(\\tilde\{x\}^\{\*\}\+\\epsilon\_\{0\}\)\+f\\\!\\bigl\(\\tilde\{x\}^\{\*\}\+\\epsilon\_\{0\}\+\\mathcal\{A\}\(\\tilde\{x\}^\{\*\}\+\\epsilon\_\{0\}\)\\bigr\),\(42\)linearize the attention term aroundx~∗\\tilde\{x\}^\{\*\}, andffaroundx~∗\+a~\\tilde\{x\}^\{\*\}\+\\tilde\{a\}to get
x1\\displaystyle x\_\{1\}=x~∗\+ϵ0\+𝒜\(x~∗\+ϵ0\)⏟=a~\+D\[a~\]ϵ0\+O\(‖ϵ0‖2\)\+f\(x~∗\+ϵ0\+𝒜\(x~∗\+ϵ0\)⏟=\(x~∗\+a~\)\+\(I\+D\[a~\]\)ϵ0\+O\(‖ϵ0‖2\)\)\\displaystyle=\\tilde\{x\}^\{\*\}\+\\epsilon\_\{0\}\+\\underbrace\{\\mathcal\{A\}\(\\tilde\{x\}^\{\*\}\+\\epsilon\_\{0\}\)\}\_\{=\\,\\tilde\{a\}\+D\[\\tilde\{a\}\]\\epsilon\_\{0\}\+O\(\\\|\\epsilon\_\{0\}\\\|^\{2\}\)\}\+f\\\!\\Bigl\(\\underbrace\{\\tilde\{x\}^\{\*\}\+\\epsilon\_\{0\}\+\\mathcal\{A\}\(\\tilde\{x\}^\{\*\}\+\\epsilon\_\{0\}\)\}\_\{=\\,\(\\tilde\{x\}^\{\*\}\+\\tilde\{a\}\)\+\(I\+D\[\\tilde\{a\}\]\)\\epsilon\_\{0\}\+O\(\\\|\\epsilon\_\{0\}\\\|^\{2\}\)\}\\Bigr\)\(43\)=x~∗\+ϵ0\+a~\+D\[a~\]ϵ0\+f\(\(x~∗\+a~\)\+\(I\+D\[a~\]\)ϵ0\)⏟=f\(x~∗\+a~\)\+D\[b~\]\(I\+D\[a~\]\)ϵ0\+O\(‖ϵ0‖2\)\\displaystyle=\\tilde\{x\}^\{\*\}\+\\epsilon\_\{0\}\+\\tilde\{a\}\+D\[\\tilde\{a\}\]\\epsilon\_\{0\}\+\\underbrace\{f\\\!\\bigl\(\(\\tilde\{x\}^\{\*\}\+\\tilde\{a\}\)\+\(I\+D\[\\tilde\{a\}\]\)\\epsilon\_\{0\}\\bigr\)\}\_\{=\\,f\(\\tilde\{x\}^\{\*\}\+\\tilde\{a\}\)\+\\,D\[\\tilde\{b\}\]\(I\+D\[\\tilde\{a\}\]\)\\epsilon\_\{0\}\+O\(\\\|\\epsilon\_\{0\}\\\|^\{2\}\)\}\(44\)=x~∗\+ϵ0\+a~\+D\[a~\]ϵ0\+f\(x~∗\+a~\)⏟=−a~\+D\[b~\]\(I\+D\[a~\]\)ϵ0\+O\(‖ϵ0‖2\)\\displaystyle=\\tilde\{x\}^\{\*\}\+\\epsilon\_\{0\}\+\\tilde\{a\}\+D\[\\tilde\{a\}\]\\epsilon\_\{0\}\+\\underbrace\{f\(\\tilde\{x\}^\{\*\}\+\\tilde\{a\}\)\}\_\{=\\,\-\\tilde\{a\}\}\+D\[\\tilde\{b\}\]\(I\+D\[\\tilde\{a\}\]\)\\epsilon\_\{0\}\+O\(\\\|\\epsilon\_\{0\}\\\|^\{2\}\)\(45\)=x~∗\+\(I\+D\[a~\]\+D\[b~\]\(I\+D\[a~\]\)\)ϵ0\\displaystyle=\\tilde\{x\}^\{\*\}\+\\bigl\(I\+D\[\\tilde\{a\}\]\+D\[\\tilde\{b\}\]\(I\+D\[\\tilde\{a\}\]\)\\bigr\)\\epsilon\_\{0\}\(46\)=x~∗\+\(I\+D\[b~\]\)\(I\+D\[a~\]\)ϵ0,\\displaystyle=\\tilde\{x\}^\{\*\}\+\(I\+D\[\\tilde\{b\}\]\)\(I\+D\[\\tilde\{a\}\]\)\\,\\epsilon\_\{0\},\(47\)wherea~≡𝒜\(x~∗\)\\tilde\{a\}\\equiv\\mathcal\{A\}\(\\tilde\{x\}^\{\*\}\)andb~=f\(x~∗\+a~\)\\tilde\{b\}=f\(\\tilde\{x\}^\{\*\}\+\\tilde\{a\}\)\. Collecting terms and definingϵ1=x1−x~∗\\epsilon\_\{1\}=x\_\{1\}\-\\tilde\{x\}^\{\*\}:
ϵ1=\(I\+D\[a~\]\+D\[b~\]\(I\+D\[a~\]\)\)ϵ0=\(I\+D\[b~\]\)\(I\+D\[a~\]\)ϵ0,\\epsilon\_\{1\}=\\bigl\(I\+D\[\\tilde\{a\}\]\+D\[\\tilde\{b\}\]\(I\+D\[\\tilde\{a\}\]\)\\bigr\)\\epsilon\_\{0\}=\(I\+D\[\\tilde\{b\}\]\)\(I\+D\[\\tilde\{a\}\]\)\\,\\epsilon\_\{0\},\(48\)which by the same argument at each step:
ϵn=\[\(I\+D\[b~\]\)\(I\+D\[a~\]\)\]nϵ0=\[M~∗\]nϵ0\.\\epsilon\_\{n\}=\\bigl\[\(I\+D\[\\tilde\{b\}\]\)\(I\+D\[\\tilde\{a\}\]\)\\bigr\]^\{n\}\\,\\epsilon\_\{0\}=\[\{\\tilde\{M\}\}^\{\*\}\]^\{n\}\\epsilon\_\{0\}\.\(49\)The fixed pointx~∗\\tilde\{x\}^\{\*\}is stable if and only if all eigenvalues of the composite Jacobian\(I\+D\[b~\]\)\(I\+D\[a~\]\)\(I\+D\[\\tilde\{b\}\]\)\(I\+D\[\\tilde\{a\}\]\)lie strictly inside the unit circle\. When attention is weakly input\-dependent atx~∗\\tilde\{x\}^\{\*\}\(i\.e\.‖D\[a~\]‖≪‖D\[b~\]‖\\\|D\[\\tilde\{a\}\]\\\|\\ll\\\|D\[\\tilde\{b\}\]\\\|\), this reduces to the condition\|1\+νk∗\|<1\|1\+\\nu\_\{k\}^\{\*\}\|<1on the eigenvaluesνk∗\\nu\_\{k\}^\{\*\}ofD\[b~\]D\[\\tilde\{b\}\], analogous to the MLP stability condition with the Jacobian now evaluated at the attention\-shifted pointx~∗\+a~\\tilde\{x\}^\{\*\}\+\\tilde\{a\}\.
## Appendix Appendix D\(Derivation of the Exponential Decay\)
From equation \([11](https://arxiv.org/html/2607.15449#S3.E11)\), a perturbationδ0\\delta\_\{0\}from the fixed pointx~∗\\tilde\{x\}^\{\*\}evolves as
\[δl\]k=\(1\+νk\)l\[δ0\]k\.\[\\delta\_\{l\}\]\_\{k\}=\(1\+\\nu\_\{k\}\)^\{l\}\\,\[\\delta\_\{0\}\]\_\{k\}\.\(50\)In the linear regime, the cosine distancedk\(l\)d\_\{k\}\(l\)between the base and perturbed representations is proportional to the magnitude of the perturbation:
dk\(l\)∝\|\[δl\]k\|=\|1\+νk\|l\|\[δ0\]k\|\.d\_\{k\}\(l\)\\propto\|\[\\delta\_\{l\}\]\_\{k\}\|=\|1\+\\nu\_\{k\}\|^\{l\}\\,\|\[\\delta\_\{0\}\]\_\{k\}\|\.\(51\)Writing\|1\+νk\|=elog\|1\+νk\|\|1\+\\nu\_\{k\}\|=e^\{\\log\|1\+\\nu\_\{k\}\|\}:
dk\(l\)=Ak⋅ellog\|1\+νk\|,d\_\{k\}\(l\)=A\_\{k\}\\cdot e^\{l\\log\|1\+\\nu\_\{k\}\|\},\(52\)whereAk∝\|\[δ0\]k\|A\_\{k\}\\propto\|\[\\delta\_\{0\}\]\_\{k\}\|\. For an irrelevant direction\|1\+νk\|<1\|1\+\\nu\_\{k\}\|<1, we havelog\|1\+νk\|<0\\log\|1\+\\nu\_\{k\}\|<0; defining
τk=−1log\|1\+νk\|\>0,\\tau\_\{k\}=\\frac\{\-1\}\{\\log\|1\+\\nu\_\{k\}\|\}\>0,\(53\)we obtain the exponential decay
dk\(l\)=Ake−l/τk\.d\_\{k\}\(l\)=A\_\{k\}\\,e^\{\-l/\\tau\_\{k\}\}\.\(54\)Note the analogy with the Markov correlation lengthξk=−1/log\|λk\|\\xi\_\{k\}=\-1/\\log\|\\lambda\_\{k\}\|: bothτk\\tau\_\{k\}andξk\\xi\_\{k\}are characteristic decay scales, one for the network dynamics and one for the input distribution\. The RG prediction thatτk\\tau\_\{k\}is monotonically increasing in\|λk\|\|\\lambda\_\{k\}\|is therefore a prediction that the network aligns its contraction rates with the spectral structure of the input\.
## Appendix Appendix E\(Mode Decomposition\)
The eigenvectorsϕk∈ℝV\\phi\_\{k\}\\in\\mathbb\{R\}^\{V\}ofPPand the eigenvectorsvj∈ℝT×dv\_\{j\}\\in\\mathbb\{R\}^\{T\\times d\}ofM∗M^\{\*\}live in different spaces, so there is no reason forϕk\\phi\_\{k\}to align with any singlevjv\_\{j\}\. When we injectϵϕk\\epsilon\\phi\_\{k\}at position0, the perturbation in representation space is:
δ\(0\)=ϵ\(Wϕk,0,…,0\)∈ℝT×d\.\\delta^\{\(0\)\}=\\epsilon\\,\(W\\phi\_\{k\},\\,0,\\,\\ldots,\\,0\)\\in\\mathbb\{R\}^\{T\\times d\}\.\(55\)Decomposing in the eigenbasis ofM∗M^\{\*\}:
δ\(0\)=∑jcj\(k\)vj,cj\(k\)=⟨vj,δ\(0\)⟩,\\delta^\{\(0\)\}=\\sum\_\{j\}c\_\{j\}^\{\(k\)\}\\,v\_\{j\},\\qquad c\_\{j\}^\{\(k\)\}=\\langle v\_\{j\},\\,\\delta^\{\(0\)\}\\rangle,\(56\)where the coefficientscj\(k\)c\_\{j\}^\{\(k\)\}are generically all nonzero sinceWϕkW\\phi\_\{k\}has no reason to align with any singlevjv\_\{j\}\. Afterlllayers, usingM~∗≈I\+M∗\\tilde\{M\}^\{\*\}\\approx I\+M^\{\*\}:
δ\(l\)=∑jcj\(k\)\(1\+νj\)lvj\.\\delta^\{\(l\)\}=\\sum\_\{j\}c\_\{j\}^\{\(k\)\}\\,\(1\+\\nu\_\{j\}\)^\{l\}\\,v\_\{j\}\.\(57\)This is a superposition of modes, each decaying or growing at rate\|1\+νj\|l\|1\+\\nu\_\{j\}\|^\{l\}\. Theτk\\tau\_\{k\}fitted from \([24](https://arxiv.org/html/2607.15449#S7.E24)\) is therefore an effective decay rate for the mixture, not a property of a single eigenmode\. It reduces to a single exponentialAke−l/τkA\_\{k\}e^\{\-l/\\tau\_\{k\}\}only if one eigenmode dominates\. For instance, if\|cj∗\(k\)\|≫\|cj\(k\)\|\|c\_\{j^\{\*\}\}^\{\(k\)\}\|\\gg\|c\_\{j\}^\{\(k\)\}\|for allj≠j∗j\\neq j^\{\*\}
τk≈−1log\|1\+νj∗\|\.\\tau\_\{k\}\\approx\-\\frac\{1\}\{\\log\|1\+\\nu\_\{j^\{\*\}\}\|\}\.\(58\)For a mode\-selective TFM, the dominant network eigenmodej∗j^\{\*\}, the indexjjfor which\|cj\(k\)\|\|c\_\{j\}^\{\(k\)\}\|is largest when Markov modekkis injected , should satisfy\|1\+νj∗\|≈1\|1\+\\nu\_\{j^\{\*\}\}\|\\approx 1when\|λk\|\|\\lambda\_\{k\}\|is large \(slow Markov mode, long persistence\) and\|1\+νj∗\|≪1\|1\+\\nu\_\{j^\{\*\}\}\|\\ll 1when\|λk\|\|\\lambda\_\{k\}\|is small \(fast mode, rapid decay\)\. This is a hypothesis about learned structure; there is no algebraic guarantee, and it is exactly what Experiment 4 tests\.
## References
- Alpay & Kilictas \(2026\)Alpay, F\. and Kilictas, B\.Latent object permanence: Topological phase transitions, free\-energy principles, and renormalization group flows in deep TFM manifolds\.*arXiv preprint arXiv:2601\.19942*, 2026\.URL[https://arxiv\.org/abs/2601\.19942](https://arxiv.org/abs/2601.19942)\.
- Bény \(2013\)Bény, C\.Deep learning and the renormalization group\.*arXiv preprint arXiv:1301\.3124*, 2013\.URL[https://arxiv\.org/abs/1301\.3124](https://arxiv.org/abs/1301.3124)\.
- Bordelon et al\. \(2024\)Bordelon, B\., Atanasov, A\., and Pehlevan, C\.A dynamical model of neural scaling laws\.In*Proceedings of the 41st International Conference on Machine Learning \(ICML\)*, volume 235, pp\. 4345–4381\. PMLR, 2024\.URL[https://proceedings\.mlr\.press/v235/bordelon24a\.html](https://proceedings.mlr.press/v235/bordelon24a.html)\.
- Clark et al\. \(2019\)Clark, K\., Khandelwal, U\., Levy, O\., and Manning, C\. D\.What does BERT look at? An analysis of BERT’s attention\.In*Proceedings of the 2019 ACL Workshop BlackboxNLP*, pp\. 276–286, Florence, Italy, 2019\. Association for Computational Linguistics\.doi:[10\.18653/v1/W19\-4828](https://doi.org/10.18653/v1/W19-4828)\.
- Haggi\-Mani & Rish \(2026\)Haggi\-Mani, P\. and Rish, I\.Rank collapse, fixed points, and the renormalization group structure of MLP residual networks\.*Preprint*, 2026\.Mila – Quebec AI Institute / Université de Montréal\.
- Kingma & Ba \(2015\)Kingma, D\. P\. and Ba, J\.Adam: A method for stochastic optimization\.In*Proceedings of the 3rd International Conference on Learning Representations \(ICLR\)*, 2015\.URL[https://arxiv\.org/abs/1412\.6980](https://arxiv.org/abs/1412.6980)\.
- Kornblith et al\. \(2019\)Kornblith, S\., Norouzi, M\., Lee, H\., and Hinton, G\.Similarity of neural network representations revisited\.In*Proceedings of the 36th International Conference on Machine Learning \(ICML\)*, volume 97, pp\. 3519–3529\. PMLR, 2019\.URL[https://proceedings\.mlr\.press/v97/kornblith19a\.html](https://proceedings.mlr.press/v97/kornblith19a.html)\.
- Martin \(2026\)Martin, M\. F\.The transformer as renormalization group flow\.*symmetry, broken*\(online essay\), January 2026\.URL[https://www\.symmetrybroken\.com/transformer\-as\-renormalization\-group\-flow/](https://www.symmetrybroken.com/transformer-as-renormalization-group-flow/)\.
- Mehta & Schwab \(2014\)Mehta, P\. and Schwab, D\. J\.An exact mapping between the Variational Renormalization Group and deep learning\.*arXiv preprint arXiv:1410\.3831*, 2014\.URL[https://arxiv\.org/abs/1410\.3831](https://arxiv.org/abs/1410.3831)\.
- Roy & Vetterli \(2007\)Roy, O\. and Vetterli, M\.The effective rank: A measure of effective dimensionality\.In*Proceedings of the 15th European Signal Processing Conference \(EUSIPCO\)*, pp\. 606–610, Poznan, Poland, 2007\.URL[https://www\.eurasip\.org/Proceedings/Eusipco/Eusipco2007/Papers/a5p\-h05\.pdf](https://www.eurasip.org/Proceedings/Eusipco/Eusipco2007/Papers/a5p-h05.pdf)\.
- Tenney et al\. \(2019\)Tenney, I\., Das, D\., and Pavlick, E\.BERT rediscovers the classical NLP pipeline\.In*Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics \(ACL\)*, pp\. 4593–4601, Florence, Italy, 2019\. Association for Computational Linguistics\.doi:[10\.18653/v1/P19\-1452](https://doi.org/10.18653/v1/P19-1452)\.
- Voita et al\. \(2019\)Voita, E\., Talbot, D\., Moiseev, F\., Sennrich, R\., and Titov, I\.Analyzing multi\-head self\-attention: Specialized heads do the heavy lifting, the rest can be pruned\.In*Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics \(ACL\)*, pp\. 5797–5808, Florence, Italy, 2019\. Association for Computational Linguistics\.doi:[10\.18653/v1/P19\-1580](https://doi.org/10.18653/v1/P19-1580)\.
- Wilson \(1971\)Wilson, K\. G\.Renormalization group and critical phenomena I\.*Physical Review B*, 4\(9\):3174–3183, 1971\.doi:[10\.1103/PhysRevB\.4\.3174](https://doi.org/10.1103/PhysRevB.4.3174)\.
- Wilson & Kogut \(1974\)Wilson, K\. G\. and Kogut, J\.The renormalization group and theε\\varepsilonexpansion\.*Physics Reports*, 12\(2\):75–199, 1974\.doi:[10\.1016/0370\-1573\(74\)90023\-4](https://doi.org/10.1016/0370-1573(74)90023-4)\.
- Fernando & Guitchounts \(2026\)Fernando, J\. & Guitchounts, G\.Dynamics of the transformer residual stream: Coupling spectral geometry to network topology\.*arXiv preprint arXiv:2605\.14258*, 2026\.URL[https://arxiv\.org/abs/2605\.14258](https://arxiv.org/abs/2605.14258)\.
- Makkuva et al\. \(2025\)Makkuva, A\. V\., Bondaschi, M\., Girish, A\., Nagle, A\., Jaggi, M\., Kim, H\., and Gastpar, M\.Attention with Markov: A framework for principled analysis of transformers via Markov chains\.*arXiv preprint arXiv:2402\.04161*, 2025\.URL[https://arxiv\.org/abs/2402\.04161](https://arxiv.org/abs/2402.04161)\.
- Coppola et al\. \(2026\)Coppola, G\. P\., Helias, M\., and Ringel, Z\.Renormalization group for deep neural networks: Universality of learning and scaling laws\.*arXiv preprint arXiv:2510\.25553*, 2026\.URL[https://arxiv\.org/abs/2510\.25553](https://arxiv.org/abs/2510.25553)\.Similar Articles
Energy-Gated Attention and Wavelet Positional Encoding: Complementary Inductive Biases for Transformer Attention
This paper proposes Energy-Gated Attention (EGA) and Morlet Positional Encoding (MoPE) to address missing inductive biases in transformer attention: token salience and scale-adaptive locality. Experiments on TinyShakespeare show superadditive gains when combined, highlighting complementarity.
I Found a Hidden Ratio in Transformers That Predicts Geometric Stability [R]
The article presents a discovered spectral ratio between MLP and attention norms that predicts geometric stability in transformer models, with an optimal range of 0.5–2 to prevent rank collapse.
A Controlled Study of Attention-Only Transformers
This paper presents a controlled study comparing attention-only transformers (Simple Attention Networks, SANs) against standard transformers matched for parameters, compute, and depth. It finds that removing feed-forward layers largely closes the performance gap when the freed capacity is reallocated to attention depth, with the remaining deficit attributed to parametric recall.
@daylenyang: this berkeley 189 lecture is probably the clearest explainer of the attention mechanism i've come across. provides a ve…
Berkeley 189 lecture provides a clear explanation of the attention mechanism, tracing the evolution from RNN+attention to Transformer and contrasting MLP/CNN parameter efficiency.
Contribution Weights: A Geometrical Analysis of Self-Attention Transformers
Introduces Contribution Weights, a projection-based metric that accounts for attention weight, value magnitude, and directional alignment to more faithfully measure token importance in transformer LLMs, revealing active functional roles of attention sinks.