The Entropic Bound for Transformers: Why Static Rank Fails and Attention-Native Rank Recovers

arXiv cs.LG Papers

Summary

This paper introduces the Entropic Bound, a spectral measure of task-intrinsic capacity for transformers, proving that the intrinsic rank of the token-mixing operator provides a tight lower bound on required model capacity. It shows that while a naive transfer from linear attention fails for real attention, an attention-native intrinsic rank restores the full theoretical structure.

arXiv:2607.23050v1 Announce Type: new Abstract: Neural scaling laws describe how loss decreases as models, data, and compute grow, but they do not answer a prior question: for a fixed task, what is the minimum model capacity required to solve it? We study this through the Entropic Bound, a spectral notion of task-intrinsic capacity for Transformers. We first prove that, in a linear attention surrogate, the intrinsic rank $r^*$ of the token-mixing operator is a tight lower bound: any rank-deficient model incurs unavoidable excess risk, and the bound is achievable at $r^*$. We further show that gradient descent recovers this rank under standard low-rank implicit-bias assumptions, confirm all three properties empirically, and show $r^*$ is recoverable from data before training. We then ask whether this transfers to real attention. A naive transfer fails, and a controlled interpolation ladder localizes the cause precisely: it is not softmax and not a rank constraint, but the input-conditioned nature of attention's mixing operator, which a static weight kernel cannot summarize. Motivated by this, we introduce an attention-native intrinsic rank -- the minimum query-key kernel rank realizing the task within the attention class -- and show that under this definition the full Entropic Bound structure (deficiency, achievability, recovery) is restored for both linear and softmax attention, with the energy effective rank as the estimator robust to softmax distortion. Finally, we map the boundary of data-only predictability: $r^*$ is exactly recoverable for linear QK attention, even without the value map at scale, while softmax attention admits only partial pre-training recovery due to nonlinear inversion and kernel-value identifiability effects. Our results reframe the Entropic Bound from a post-hoc descriptor into an attention-native capacity measure with a precisely characterized predictability frontier.
Original Article
View Cached Full Text

Cached at: 07/28/26, 06:24 AM

# The Entropic Bound for Transformers: Why Static Rank Fails and Attention-Native Rank Recovers
Source: [https://arxiv.org/html/2607.23050](https://arxiv.org/html/2607.23050)
###### Abstract

Neural scaling laws describe how loss decreases as models, data, and compute grow, but they do not answer a prior question: for a fixed task, what is the*minimum*model capacity required to solve it? We study this through theEntropic Bound, a spectral notion of task\-intrinsic capacity for Transformers\. We first prove that, in a linear attention surrogate, the intrinsic rankr∗r^\{\*\}of the token\-mixing operator is a*tight*lower bound: any rank\-deficient model incurs unavoidable excess risk, and the bound is achievable atr∗r^\{\*\}\. We further analyze gradient\-descent recovery under standard low\-rank implicit\-bias assumptions, showing the learned effective rank concentrates atr∗r^\{\*\}\. We confirm all three properties empirically and showr∗r^\{\*\}is recoverable from data*before training*\. We then ask whether this transfers to real attention\. A naive transferfails, and through a controlled interpolation ladder we localize the cause precisely: it is not softmax and not a rank constraint, but theinput\-conditionednature of attention’s mixing operator, which a static weight kernel cannot summarize\. Motivated by this, we introduce anattention\-native intrinsic rank—the minimum query–key kernel rank realizing the task within the attention class—and show that under this definition the full Entropic Bound structure \(deficiency, achievability, recovery\) is restored for both linear and softmax attention, with the*energy*effective rank as the estimator robust to softmax distortion\. Finally, we map the boundary of data\-only predictability:r∗r^\{\*\}is exactly recoverable for linear QK attention, even without the value map at scale, while softmax attention admits only partial pre\-training recovery due to nonlinear inversion and kernel–value identifiability effects\. Our results reframe the Entropic Bound from a post\-hoc descriptor into an attention\-native capacity measure with a precisely characterized predictability frontier\.

## 1Introduction

The empirical regularities of large\-model training are by now well documented: loss falls as a power law in parameters, data, and compute\. Yet scaling laws describe a*trajectory*, not a*floor*\. They tell us how performance improves as we add capacity; they do not tell us how little capacity a given task fundamentally requires\. This quantity—the minimum model size at which a task becomes solvable to a target error—is conceptually prior to scaling, and it is the object we study\.

A useful analogy comes from coding theory\. Once an alphabet and admissible code lengths are fixed, only finitely many messages are representable\. Likewise, once a model family is fixed, a finite parameter vector realizes only a restricted set of input–output maps\. If a task demands many distinguishable predictive modes, any model solving it must carry sufficient*effective*capacity\. The difficulty is that raw parameter count is not architecture\-invariant, so the right object is atask\-intrinsic effective dimensionthat induces architecture\-dependent parameter lower bounds\.

For a linear attention surrogate this program can be carried out exactly\. We define the intrinsic rankr∗r^\{\*\}as the minimum rank of a token\-mixing operator that reaches the noise floor, and prove it is tight: rank deficiency forces excess risk \(necessity\), the floor is achievable atr∗r^\{\*\}\(sufficiency\), and gradient descent from small initialization converges to it \(implicit bias\)\. These results give a clean and complete picture—in the surrogate\.

The central question of this paper is whether that picture survives contact with real attention\. We find that a*naive*transfer fails, but the failure is informative\. By constructing an interpolation ladder that begins at the linear surrogate and changes one architectural component at a time, we localize the breakdown to a single cause—and, contrary to the natural guess, that cause isnot softmax\. It is the input\-conditioned bilinear formX​K​X⊤XKX^\{\\top\}that attention uses to compute scores: because this operator depends on the input, no static property of the weight kernelKK\(neither its rank nor its effective rank\) determines the model’s task capacity\. A control with an unconstrained full\-rank kernel fails identically, ruling out a rank\-constraint explanation\.

This motivates a redefinition\. Rather than measuring the static kernel, we define anattention\-native intrinsic rankas the minimum query–key kernel rank that realizes the task*within the attention class itself*\. Generating teacher tasks from this same class removes the operator mismatch by construction, and under the new definition the entire Entropic Bound structure returns: rank\-deficient students incur excess risk, students atr∗r^\{\*\}reach the floor, and over\-parameterized students recoverr∗r^\{\*\}\. This holds not only for linear attention but also for softmax attention, provided capacity is measured by theenergyeffective rank—which we show is robust to the small\-singular\-value noise softmax introduces\.

Finally, we characterize*predictability*: estimatingr∗r^\{\*\}from data without training\. We show exact recovery for linear QK attention, including the V\-unknown setting at scale, and identify the remaining frontier in softmax attention, where nonlinear inversion corrupts the numerical rank and leaves only partial energy\-rank predictability\.

#### Contributions\.

1. 1\.We restate and empirically validate the tightness of the Entropic Bound in the linear attention surrogate \(deficiency, achievability, recovery\), and demonstrate data\-only recovery ofr∗r^\{\*\}\(Sections[4](https://arxiv.org/html/2607.23050#S4),[6](https://arxiv.org/html/2607.23050#S6)\)\.
2. 2\.We show that a naive transfer to real attention fails, and via a controlled interpolation ladder localize the cause toinput conditioning—not softmax, not a rank constraint \(Section[5](https://arxiv.org/html/2607.23050#S5)\)\.
3. 3\.We introduce theattention\-native intrinsic rankand show it restores the full bound structure for linear and softmax attention, with the energy effective rank as the softmax\-robust estimator \(Section[6](https://arxiv.org/html/2607.23050#S6)\)\.
4. 4\.We map thepredictability frontier: exact data\-only recovery for linear QK attention, including the V\-unknown setting at scale, and partial recovery under softmax, where nonlinear inversion corrupts the numerical rank and energy\-rank estimators remain imperfect \(Section[6](https://arxiv.org/html/2607.23050#S6.SS0.SSS0.Px3)\)\.

Fixed\-operator surrogateY=A∗​X​V∗Y=A^\{\*\}XV^\{\*\}B0 succeeds;r∗r^\{\*\}tight\(Thm 1–4, Table[1](https://arxiv.org/html/2607.23050#S4.T1)\)Static\-rank transferfixed target, studentX​K​X⊤XKX^\{\\top\}B1/B1c/B2 fail*input\-conditioning*\(Table[2](https://arxiv.org/html/2607.23050#S5.T2)\)Attention\-native teacherY=AttnK∗,V∗​\(X\)Y=\\mathrm\{Attn\}\_\{K^\{\*\},V^\{\*\}\}\(X\),rank​\(K∗\)=r∗\\mathrm\{rank\}\(K^\{\*\}\)\{=\}r^\{\*\}deficiency/achievability/recovery restored \(Table[3](https://arxiv.org/html/2607.23050#S6.T3)\)naivetransferredefine inattn\. classFigure 1:Overview of the paper’s logic\. The fixed\-operator linear surrogate admits a tight Entropic Bound \(left\)\. A naive transfer to attention fails even with a full\-rank kernel \(center\), localizing the break to input\-conditioned mixing rather than softmax or rank deficiency\. Defining the intrinsic rank inside the attention class restores deficiency, achievability, and recovery \(right\)\.

## 2Related Work

#### Neural scaling laws\.

Empirical scaling laws characterize loss as a smooth power law in model size, data, and compute\(Kaplan et al\.,[2020](https://arxiv.org/html/2607.23050#bib.bib11); Hoffmann et al\.,[2022](https://arxiv.org/html/2607.23050#bib.bib9)\)\. Our framework is complementary: rather than the trajectory of improvement, we study the*critical capacity threshold*a task imposes\.

#### Information\-theoretic lower bounds\.

Rate–distortion and predictive\-information arguments bound representation size from below\(Cover and Thomas,[2006](https://arxiv.org/html/2607.23050#bib.bib5); Berger,[1971](https://arxiv.org/html/2607.23050#bib.bib4)\)\. Our spectral pruning view is a matrix\-valued analogue: the rate is the retained rankrrand the distortion is the tail energy∑j\>rσj2\\sum\_\{j\>r\}\\sigma\_\{j\}^\{2\}; the effective rank marks where an additional spectral mode stops paying for itself\.

#### Statistical learning theory\.

Classical complexity measures \(VC dimension, Rademacher complexity, covering numbers\) are worst\-case and distribution\-free, and too coarse here\(Vapnik,[1998](https://arxiv.org/html/2607.23050#bib.bib15); Bartlett and Mendelson,[2002](https://arxiv.org/html/2607.23050#bib.bib3)\)\. Our analysis is distribution\-dependent: metric entropy controls generalization, whereas our bound controls*representational capacity*\.

#### Low\-rank structure and compression\.

A large literature shows fine\-tuning updates and trained weights are low\-rank, and that low\-rank parameterizations match dense models cheaply\(Hu et al\.,[2022](https://arxiv.org/html/2607.23050#bib.bib10); Aghajanyan et al\.,[2021](https://arxiv.org/html/2607.23050#bib.bib1)\)\. These address the rank of updates or compressibility of trained models; we ask the prior question of the minimum rank to realize a task, and show the*static kernel rank*is the wrong object for attention\.

#### Implicit regularization and spectral bias\.

Gradient descent on over\-parameterized models is biased toward low\-rank solutions\(Srebro et al\.,[2004](https://arxiv.org/html/2607.23050#bib.bib14); Gunasekar et al\.,[2017](https://arxiv.org/html/2607.23050#bib.bib8); Arora et al\.,[2019](https://arxiv.org/html/2607.23050#bib.bib2)\)\. This underlies why a trained kernel’s effective rank can converge to the task\-intrinsic rank, and grounds our recovery experiments\.

#### Effective rank and attention rank collapse\.

The effective rank\(Roy and Vetterli,[2007](https://arxiv.org/html/2607.23050#bib.bib13)\)and the phenomenon of attention rank collapse\(Dong et al\.,[2021](https://arxiv.org/html/2607.23050#bib.bib6); Kobayashi et al\.,[2020](https://arxiv.org/html/2607.23050#bib.bib12)\)have been studied as dynamics and signal\-propagation effects\. Our contribution is orthogonal: we tie a spectral capacity measure to a*task\-intrinsic*lower bound, identify why the static\-kernel measure fails, and propose the realized\-operator alternative\.

## 3Preliminaries

#### Population risk\.

For a distributionDDover\(x,y\)\(x,y\)and lossℓ\\ell, the population risk isLD​\(f\)=𝔼​\[ℓ​\(f​\(x\),y\)\]L\_\{D\}\(f\)=\\mathbb\{E\}\[\\ell\(f\(x\),y\)\], and the minimum required parameter count at toleranceε\\varepsilonisP∗​\(ε;D\)=min⁡\{P:∃f​with​P​params,LD​\(f\)≤ε\}P^\{\*\}\(\\varepsilon;D\)=\\min\\\{P:\\exists f\\text\{ with \}P\\text\{ params\},L\_\{D\}\(f\)\\leq\\varepsilon\\\}\.

#### Spectral complexity measures\.

ForMMwith singular valuesσ1≥⋯≥σq\\sigma\_\{1\}\\geq\\cdots\\geq\\sigma\_\{q\}:

Reff​\(M\)\\displaystyle R\_\{\\mathrm\{eff\}\}\(M\)=exp⁡\(−∑ipi​log⁡pi\),pi=σi/∑jσj,\\displaystyle=\\exp\\\!\\Big\(\-\\textstyle\\sum\_\{i\}p\_\{i\}\\log p\_\{i\}\\Big\),\\quad p\_\{i\}=\\sigma\_\{i\}/\\textstyle\\sum\_\{j\}\\sigma\_\{j\},\(1\)R~eff​\(M\)\\displaystyle\\tilde\{R\}\_\{\\mathrm\{eff\}\}\(M\)=exp⁡\(−∑ip~i​log⁡p~i\),p~i=σi2/∑jσj2,\\displaystyle=\\exp\\\!\\Big\(\-\\textstyle\\sum\_\{i\}\\tilde\{p\}\_\{i\}\\log\\tilde\{p\}\_\{i\}\\Big\),\\quad\\tilde\{p\}\_\{i\}=\\sigma\_\{i\}^\{2\}/\\textstyle\\sum\_\{j\}\\sigma\_\{j\}^\{2\},\(2\)Rstable​\(M\)\\displaystyle R\_\{\\mathrm\{stable\}\}\(M\)=‖M‖F2/‖M‖22\.\\displaystyle=\\\|M\\\|\_\{F\}^\{2\}/\\\|M\\\|\_\{2\}^\{2\}\.\(3\)These satisfy1≤Rstable≤R~eff≤Reff≤rank​\(M\)1\\leq R\_\{\\mathrm\{stable\}\}\\leq\\tilde\{R\}\_\{\\mathrm\{eff\}\}\\leq R\_\{\\mathrm\{eff\}\}\\leq\\mathrm\{rank\}\(M\)\. For attention kernels the energy variantR~eff\\tilde\{R\}\_\{\\mathrm\{eff\}\}\(Roy and Vetterli,[2007](https://arxiv.org/html/2607.23050#bib.bib13)\)is the appropriate measure, a choice we justify empirically in Section[6](https://arxiv.org/html/2607.23050#S6)\.

#### Task\-intrinsic effective dimension\.

Given a reference classGGwith a complexity functional,

r∗​\(ε;D\)=inf\{dimeff\(g\):g∈G,LD​\(g\)≤ε\},r^\{\*\}\(\\varepsilon;D\)=\\inf\\\{\\dim\_\{\\mathrm\{eff\}\}\(g\):g\\in G,\\,L\_\{D\}\(g\)\\leq\\varepsilon\\\},\(4\)and the Entropic Bound isB​\(ε;D\)=r∗​\(ε;D\)B\(\\varepsilon;D\)=r^\{\*\}\(\\varepsilon;D\), summed over layers\.

## 4The Entropic Bound in the Linear Attention Surrogate

#### Setting\.

LetX∈ℝT×dX\\in\\mathbb\{R\}^\{T\\times d\}and considerY^A,V=A​X​V\\hat\{Y\}\_\{A,V\}=AXVwith targetY=A∗​X​V∗\+EY=A^\{\*\}XV^\{\*\}\+E,rank​\(A∗\)=r∗\\mathrm\{rank\}\(A^\{\*\}\)=r^\{\*\},E∼𝒩​\(0,σ2​I\)E\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\)\. The intrinsic rank is

Blin​\(ε;D\)=min⁡\{rank​\(A\):∃V,𝔼​‖Y−A​X​V‖F2≤ε\},B\_\{\\mathrm\{lin\}\}\(\\varepsilon;D\)=\\min\\\{\\mathrm\{rank\}\(A\):\\exists V,\\ \\mathbb\{E\}\\\|Y\-AXV\\\|\_\{F\}^\{2\}\\leq\\varepsilon\\\},\(5\)and at the noise floorε=σ2​T​d\\varepsilon=\\sigma^\{2\}Tdwe haveBlin=r∗B\_\{\\mathrm\{lin\}\}=r^\{\*\}\.

#### Tightness \(restated\)\.

Under full\-rank input covariance, error independent of input, and a spectral gap atr∗r^\{\*\}:

###### Theorem 1\(Deficiency⇒\\Rightarrowexcess risk\)\.

Ifr<r∗r<r^\{\*\}, every rank\-rrmodel incurs excess risk bounded below by the tail energy∑j\>rλj​\(A∗\)2\\sum\_\{j\>r\}\\lambda\_\{j\}\(A^\{\*\}\)^\{2\}scaled by the input\-covariance constant; in particular the risk strictly exceeds the noise floor\.

###### Theorem 2\(Achievability\)\.

Ifr≥r∗r\\geq r^\{\*\}, a rank\-r∗r^\{\*\}model attains the noise floor exactly\.

###### Theorem 3\(Tightness\)\.

Consequentlyr∗r^\{\*\}is exactly the minimum effective dimension for Bayes\-optimal performance\.

###### Theorem 4\(GD convergence\)\.

From small initialization, gradient descent converges to a solution whose effective rank equalsr∗r^\{\*\}up to anO​\(σ2/n\)O\(\\sigma^\{2\}/n\)term, with a Marchenko–Pastur correction for finite\-sample estimation\.

Full statements and proofs are in Appendix[A](https://arxiv.org/html/2607.23050#A1); the deficiency bound follows from the Eckart–Young–Mirsky theorem\(Eckart and Young,[1936](https://arxiv.org/html/2607.23050#bib.bib7)\)\.

#### Empirical validation\.

We instantiate the surrogate withT=16T\{=\}16,d=12d\{=\}12, and a rank\-r∗=3r^\{\*\}\{=\}3target with a sharp spectral gap, and verify each property \(Table[1](https://arxiv.org/html/2607.23050#S4.T1)\)\. The rank sweep shows a clean phase transition: excess risk is large forr<r∗r<r^\{\*\}\(28\.528\.5atr=1r\{=\}1,8\.88\.8atr=2r\{=\}2\), collapses to the noise floor atr=r∗r\{=\}r^\{\*\}\(≈2×10−4\\approx 2\\times 10^\{\-4\}\), and stays there forr\>r∗r\>r^\{\*\}\. Over\-parameterized training from small initialization recovers the intrinsic rank: the Marchenko–Pastur\-corrected effective rank is2\.972\.97, matchingr∗=3r^\{\*\}\{=\}3\. Data\-only recovery ofr∗r^\{\*\}\(Section[6](https://arxiv.org/html/2607.23050#S6.SS0.SSS0.Px3)\) gives numerical rank exactly33\.

Table 1:Linear surrogate validation \(r∗=3r^\{\*\}\{=\}3, noise floor0\.480\.48\)\. Sharp deficiency–achievability transition atr=r∗r\{=\}r^\{\*\}; over\-parameterized MP\-correctedReff=2\.97→r∗R\_\{\\mathrm\{eff\}\}=2\.97\\to r^\{\*\}\.The pilot reproduces all four properties cleanly, anchoring the theory before we test transfer to attention\.

## 5Why Naive Transfer Fails: Localizing the Break

#### The interpolation ladder\.

On the*same*synthetic task with knownr∗r^\{\*\}, we build four classes differing by one component each:B0, linearA​X​VAXVwithA=P​Q⊤A=PQ^\{\\top\};B1, linear QK\(X​WQ\)​\(X​WK\)⊤​X​V\(XW\_\{Q\}\)\(XW\_\{K\}\)^\{\\top\}XVwithout softmax;B2,softmax​\(\(X​WQ\)​\(X​WK\)⊤\)​X​V\\mathrm\{softmax\}\(\(XW\_\{Q\}\)\(XW\_\{K\}\)^\{\\top\}\)\\,XV; andB3, the full block \(softmax attention with residual, LayerNorm, and FFN\)\. For each we measure achievability \(excess atr=r∗r=r^\{\*\}\) and recovery \(effective rank of the realized operator\)\.

#### The break is at B1; the cause is input conditioning\.

Table[2](https://arxiv.org/html/2607.23050#S5.T2)reports the ladder on a fixed\-operator task withr∗=3r^\{\*\}\{=\}3\. B0 solves atr∗r^\{\*\}\(excess≈0\\approx 0, recoversReff→r∗R\_\{\\mathrm\{eff\}\}\\to r^\{\*\}\);B1 already fails\(excess69\.469\.4atr=r∗r\{=\}r^\{\*\}\), and adding softmax \(B2, excess68\.768\.7\) or the full stack \(B3, excess51\.951\.9\) does not restore solvability\. The natural hypothesis that softmax is responsible is*refuted*: B2≈\\approxB1 in excess risk, so the nonlinearity changes little\. The break occurs one rung earlier, at the QK factorization\.

To isolate the mechanism we add a controlB1c: input\-conditioned mixingX​K​X⊤XKX^\{\\top\}with an*unconstrained full\-rank*KK\. It fails identically \(excess69\.069\.0\), so the break isnota rank\-deficiency effect\. The cause is that attention’s realized mixerM​\(X\)=X​K​X⊤M\(X\)=XKX^\{\\top\}isinput\-conditioned: a bilinear form inXXcannot equal a fixed target operatorA∗A^\{\*\}for allXX, regardless ofrank​\(K\)\\mathrm\{rank\}\(K\)\. Measuring the static kernelReff​\(WQ​K\)R\_\{\\mathrm\{eff\}\}\(W\_\{QK\}\)is therefore the wrong object, and softmax is a red herring for this failure\.

Table 2:Interpolation ladder on a fixed\-operator task \(r∗=3r^\{\*\}\{=\}3\)\. The bound transfers only at B0; the break is at B1 \(QK factorization\),*before*softmax\. B1c \(full\-rankKK\) fails identically, ruling out a rank\-constraint explanation\.
#### Consequence for measurement\.

When the architecture cannot realize the task’s mixing operator, the measured effective rank ceases to encode task complexity and instead tracks width and optimization\. On trained language models we observe precisely this: across corpora of very different complexity the measured attention effective rank is nearly invariant\. Quantitatively \(Appendix[B](https://arxiv.org/html/2607.23050#A2)\), the predicted rank and the measured rank are anti\-correlated \(r=−0\.27r=\-0\.27\), whereas the final loss and the measured rank correlate atr=−0\.79r=\-0\.79, indicating that the measured rank tracks optimization progress rather than corpus complexity\.

## 6An Attention\-Native Intrinsic Rank

#### Definition\.

We define the intrinsic rank inside the attention class:

rattn∗​\(ε;D\)=min⁡\{rank​\(K\):∃V,𝔼​‖Y−AttnK,V​\(X\)‖2≤ε\},r^\{\*\}\_\{\\mathrm\{attn\}\}\(\\varepsilon;D\)=\\min\\\{\\mathrm\{rank\}\(K\):\\exists V,\\ \\mathbb\{E\}\\\|Y\-\\mathrm\{Attn\}\_\{K,V\}\(X\)\\\|^\{2\}\\leq\\varepsilon\\\},\(6\)whereAttnK,V​\(X\)=softmax​\(X​K​X⊤\)​X​V\\mathrm\{Attn\}\_\{K,V\}\(X\)=\\mathrm\{softmax\}\(XKX^\{\\top\}\)XVor its softmax\-free analogueX​K​X⊤​X​VXKX^\{\\top\}XV\. To avoid the operator mismatch of Section[5](https://arxiv.org/html/2607.23050#S5), teacher tasks are generated from the same family:\(1a\) linear\_qkY=\(X​K∗​X⊤\)​X​V∗\+EY=\(XK^\{\*\}X^\{\\top\}\)XV^\{\*\}\+Eand\(1b\) softmax\_qkY=softmax​\(X​K∗​X⊤\)​X​V∗\+EY=\\mathrm\{softmax\}\(XK^\{\*\}X^\{\\top\}\)XV^\{\*\}\+E, withK∗K^\{\*\}of rankr∗r^\{\*\}and a sharp spectral gap\.

#### The bound structure is restored\.

Table[3](https://arxiv.org/html/2607.23050#S6.T3)reports the full grid \(T=16T=16,d=32d=32, five seeds,r∗∈\{1,2,3,4,6,8\}r^\{\*\}\\in\\\{1,2,3,4,6,8\\\}\)\. For*both*teachers, all three structural properties hold across everyr∗r^\{\*\}: rank\-deficient students incur clearly positive excess risk \(deficiency\), students atr=r∗r=r^\{\*\}reach the noise floor \(achievability, excess≈0\\approx 0\), and over\-parameterized students recoverReff​\(K\)→r∗R\_\{\\mathrm\{eff\}\}\(K\)\\to r^\{\*\}\(recovery,6/66/6\)\. Seed stability is excellent:CVseed≤0\.003\\mathrm\{CV\}\_\{\\mathrm\{seed\}\}\\leq 0\.003throughout, two orders of magnitude below the0\.100\.10threshold\.

Table 3:Attention\-native intrinsic rank, full grid \(T=16T\{=\}16,d=32d\{=\}32, 5 seeds\)\. Both teachers reproduce deficiency, achievability, and recovery for allr∗r^\{\*\}\. Softmax compresses the excess\-risk magnitude \(∼\\sim450 vs\.∼\\sim2\.5 atr=r∗−1r\{=\}r^\{\*\}\{\-\}1\) but preserves the transition\.The linear/softmax contrast isolates softmax’s true effect\. The transition*persists*under softmax, but its magnitude is compressed by roughly two orders of magnitude: rank\-deficient excess risk is∼\\sim450 for the linear teacher versus∼\\sim2\.5 for the softmax teacher\. This is exactly the softening predicted by the spectral\-gap analysis \(Corollary to Theorem 1\)—softmax normalization shrinks the gap\-driven penalty—and is a quantitative softmax signature rather than a breakdown of the bound\.

#### Predictability and its frontier\.

We estimater∗r^\{\*\}from\(X,Y\)\(X,Y\)without training \(Table[4](https://arxiv.org/html/2607.23050#S6.T4)\)\. For thelinearteacher the recovery is essentially perfect: with the value map identified, the numerical rank matchesr∗r^\{\*\}exactly and the energy effective rank tracks it \(6/66/6\); strikingly, recovery also succeeds*without*the value map \(V\-unknown energyReffR\_\{\\mathrm\{eff\}\},6/66/6\) at this scale—the kernel–value identifiability gap we observed at smallddcloses as dimension grows, restoring fully data\-only prediction\.

Table 4:Data\-only predictability ofr∗r^\{\*\}\(energy effective rank\)\. Linear: exact recovery, with and without the value map\. Softmax: numerical rank saturates \(17–23\); kernel\-side energyReffR\_\{\\mathrm\{eff\}\}over\-estimates at smallr∗r^\{\*\}; the realized\-mixer energyReffR\_\{\\mathrm\{eff\}\}is monotone and tracksr∗r^\{\*\}at larger values\. No single estimator is exact across the whole range under softmax—a precisely characterized partial\-predictability frontier\.For thesoftmaxteacher, predictability is*partial*, and we report this precisely rather than overclaiming\. The numerical rank saturates \(17–23\), confirming that small\-singular\-value counts are corrupted by softmax\. The energy effective rank is far better behaved but no single estimator is exact across the whole range: the kernel\-side energyReffR\_\{\\mathrm\{eff\}\}over\-estimates at smallr∗r^\{\*\}\(e\.g\.2\.312\.31atr∗=1r^\{\*\}\{=\}1\) while approachingr∗r^\{\*\}at larger values, whereas the*realized\-mixer*energyReffR\_\{\\mathrm\{eff\}\}is monotone inr∗r^\{\*\}\(1\.10→8\.461\.10\\\!\\to\\\!8\.46\) and accurate at largerr∗r^\{\*\}but compressed for smallr∗r^\{\*\}\. Notably, the realized\-mixer measure is informative only under softmax: for the linear teacher it saturates at≈16\\approx 16becauseX​K​X⊤XKX^\{\\top\}inherits the full\-rank input geometry, so the two teachers require different measurement objects\. Crucially,*recovery from the trained model succeeds for softmax*\(Reff​\(K\)→r∗R\_\{\\mathrm\{eff\}\}\(K\)\\to r^\{\*\}, Table[3](https://arxiv.org/html/2607.23050#S6.T3)\); it is only the*pre\-training*data\-only inversion that softmax’s nonlinearity degrades\.

We therefore state the predictability claim precisely: the attention\-native intrinsic rank isexactly data\-predictable for linear QK attention\(with or without the value map at scale\), andpartially predictable under softmax, where the energy effective rank—not the numerical rank—is the appropriate estimator, and the realized\-mixer variant captures the large\-r∗r^\{\*\}regime\. Characterizing a single softmax\-robust pre\-training estimator across allr∗r^\{\*\}is a well\-posed open problem\.

#### Robustness to the spectral profile\.

The sharp transitions above use a sharp\-gap teacher kernel\. We verify in Appendix[D](https://arxiv.org/html/2607.23050#A4)that the result is not an artifact of this choice: across sharp, power\-law, and flat \(no\-gap\) spectra, achievability atr=r∗r=r^\{\*\}and recovery hold throughout, and the recovered energy effective rank tracks the teacher kernel’s energy rank in every regime—exactly tor∗r^\{\*\}for a flat spectrum and to the \(smaller\) energy rank when a few directions dominate\. Sharp gaps yield abrupt transitions while gradual spectra yield softer knees, but the energy effective rank remains the capacity\-tracking quantity regardless of the profile\.

## 7Discussion, Limitations, and Open Problems

#### Established\.

In the fixed\-operator surrogate the Entropic Bound is tight, achievable, recovered by GD, and data\-predictable\. Naive transfer fails because attention’s mixing is input\-conditioned; this is localized cleanly \(not softmax, not rank\)\. Redefining the intrinsic rank inside the attention class restores the full structure—deficiency, achievability, recovery—for both linear and softmax attention at scale \(T=16T\{=\}16,d=32d\{=\}32,CVseed≤0\.003\\mathrm\{CV\}\_\{\\mathrm\{seed\}\}\\leq 0\.003\)\. Data\-only predictability is exact for linear QK attention, including the V\-unknown estimator at scale, and partial for softmax attention, where the trained kernel still recoversr∗r^\{\*\}but pre\-training inversion remains distorted\.

#### Limitations\.

Positive results are in a teacher–student synthetic setting with knownr∗r^\{\*\}\. On trained language models the realized attention operator is dominated by task\-independent structure \(locality, positional, attention\-sink priors\), so measured effective rank does not separate corpora; whether real\-language tasks admit a low\-rank attention\-native description is open\. Under softmax, no single pre\-training estimator recoversr∗r^\{\*\}across all ranks, bounding the data\-only claim in the nonlinear regime\.

#### Open problems\.

\(i\) Find a single softmax\-robust*pre\-training*estimator ofr∗r^\{\*\}accurate across all ranks \(kernel\-side over\-estimates smallr∗r^\{\*\}; realized\-mixer compresses them\)\. \(ii\) Characterize the realized\-operator effective rank of trained LMs and its task\-invariance\. \(iii\) Bridge the synthetic attention\-native regime to natural language\. \(iv\) Extend the depth–width analysis \(early layers task\-specific, later layers saturating\)\.

#### Conclusion\.

Recast around an attention\-native intrinsic rank, the Entropic Bound is a tight capacity measure for teacher\-generated attention tasks: fully data\-predictable for linear QK attention, including the V\-unknown setting at scale, and partially predictable under softmax\. We give a precise account of where it holds, where it breaks, and why\.

## Appendix AFull Theorem Statements and Proofs

We restate the linear\-surrogate results of Section[4](https://arxiv.org/html/2607.23050#S4)with compact proofs\. The setting isY^A,V=A​X​V\\hat\{Y\}\_\{A,V\}=AXVwith targetY=A∗​X​V∗\+EY=A^\{\*\}XV^\{\*\}\+E,rank​\(A∗\)=r∗\\mathrm\{rank\}\(A^\{\*\}\)=r^\{\*\}\.

### A\.1A\.1 Assumptions

- \(A1\)XXhas zero mean and full\-rank covarianceΣX≻0\\Sigma\_\{X\}\\succ 0\.
- \(A2\)EEis independent ofXX, zero mean, with𝔼​\[E​E⊤\]=σ2​I\\mathbb\{E\}\[EE^\{\\top\}\]=\\sigma^\{2\}I\.
- \(A3\)rank​\(A∗\)=r∗\\mathrm\{rank\}\(A^\{\*\}\)=r^\{\*\}with a spectral gapλr∗​\(A∗\)≥δ\>0\\lambda\_\{r^\{\*\}\}\(A^\{\*\}\)\\geq\\delta\>0\.
- \(A4\)V∗V^\{\*\}is non\-degenerate:σmin​\(V∗\)\>0\\sigma\_\{\\min\}\(V^\{\*\}\)\>0\.

### A\.2A\.2 Rank deficiency implies excess risk

###### Theorem 5\(Deficiency\)\.

LetLr∗=infrank​\(A\)≤r,V𝔼​‖Y−A​X​V‖F2L\_\{r\}^\{\*\}=\\inf\_\{\\mathrm\{rank\}\(A\)\\leq r,\\,V\}\\mathbb\{E\}\\\|Y\-AXV\\\|\_\{F\}^\{2\}\. Under \(A1\)–\(A4\),

Lr∗≥σ2​T​d\+c​∑j=r\+1r∗λj​\(A∗\)2,c=λmin​\(ΣX\)​σmin​\(V∗\)2\.L\_\{r\}^\{\*\}\\;\\geq\\;\\sigma^\{2\}Td\\;\+\\;c\\sum\_\{j=r\+1\}^\{r^\{\*\}\}\\lambda\_\{j\}\(A^\{\*\}\)^\{2\},\\qquad c=\\lambda\_\{\\min\}\(\\Sigma\_\{X\}\)\\,\\sigma\_\{\\min\}\(V^\{\*\}\)^\{2\}\.\(7\)In particular, ifr<r∗r<r^\{\*\}thenLr∗\>σ2​T​dL\_\{r\}^\{\*\}\>\\sigma^\{2\}Td\.

###### Proof\.

For any feasibleY^=A​X​V\\hat\{Y\}=AXV, writeY−Y^=\(A∗​X​V∗−A​X​V\)\+EY\-\\hat\{Y\}=\(A^\{\*\}XV^\{\*\}\-AXV\)\+E\. By \(A2\) the cross term vanishes in expectation, so𝔼​‖Y−Y^‖F2=𝔼​‖A∗​X​V∗−A​X​V‖F2\+σ2​T​d\\mathbb\{E\}\\\|Y\-\\hat\{Y\}\\\|\_\{F\}^\{2\}=\\mathbb\{E\}\\\|A^\{\*\}XV^\{\*\}\-AXV\\\|\_\{F\}^\{2\}\+\\sigma^\{2\}Td\. The signal term is minimized over rank\-rrmaps; sinceA​X​VAXVhas rank at mostrrin the operator acting onXX, Eckart–Young–Mirsky givesinfrank​\(B\)≤r‖A∗−B‖F2=∑j\>rλj​\(A∗\)2\\inf\_\{\\mathrm\{rank\}\(B\)\\leq r\}\\\|A^\{\*\}\-B\\\|\_\{F\}^\{2\}=\\sum\_\{j\>r\}\\lambda\_\{j\}\(A^\{\*\}\)^\{2\}\. Propagating through the input covariance and the value map contributes the factorc=λmin​\(ΣX\)​σmin​\(V∗\)2\>0c=\\lambda\_\{\\min\}\(\\Sigma\_\{X\}\)\\sigma\_\{\\min\}\(V^\{\*\}\)^\{2\}\>0, yielding the bound\. Under \(A3\),λj​\(A∗\)\>0\\lambda\_\{j\}\(A^\{\*\}\)\>0forj≤r∗j\\leq r^\{\*\}, so the sum is strictly positive whenr<r∗r<r^\{\*\}\. ∎

### A\.3A\.3 Achievability

###### Theorem 6\(Achievability\)\.

Ifr≥r∗r\\geq r^\{\*\}, choosingA=A∗A=A^\{\*\},V=V∗V=V^\{\*\}givesY^=A∗​X​V∗=Y−E\\hat\{Y\}=A^\{\*\}XV^\{\*\}=Y\-E, hence𝔼​‖Y−Y^‖F2=𝔼​‖E‖F2=σ2​T​d\\mathbb\{E\}\\\|Y\-\\hat\{Y\}\\\|\_\{F\}^\{2\}=\\mathbb\{E\}\\\|E\\\|\_\{F\}^\{2\}=\\sigma^\{2\}Td, attaining the noise floor\.

### A\.4A\.4 Tightness

###### Theorem 7\(Tightness\)\.

Under \(A1\)–\(A4\),min⁡\{rank​\(A\):∃V,𝔼​‖Y−A​X​V‖F2≤σ2​T​d\}=r∗\\min\\\{\\mathrm\{rank\}\(A\):\\exists V,\\ \\mathbb\{E\}\\\|Y\-AXV\\\|\_\{F\}^\{2\}\\leq\\sigma^\{2\}Td\\\}=r^\{\*\}\. Combining A\.2 \(every rank\-r<r∗r<r^\{\*\}model has strictly positive excess risk\) with A\.3 \(rank\-r∗r^\{\*\}achieves the floor\),r∗r^\{\*\}is exactly the minimum effective dimension for Bayes\-optimal performance\.

### A\.5A\.5 Gradient\-descent recovery

###### Theorem 8\(Recovery, under standard implicit\-bias assumptions\)\.

Consider the factorized surrogate trained by gradient flow from initialization of scaleα→0\\alpha\\to 0\. In the noiseless population limit, under the standard implicit\-bias result for linear matrix factorization\(Gunasekar et al\.,[2017](https://arxiv.org/html/2607.23050#bib.bib8); Arora et al\.,[2019](https://arxiv.org/html/2607.23050#bib.bib2)\), gradient flow converges to the minimum\-nuclear\-norm interpolating solution, whose rank equalsr∗r^\{\*\}under the spectral gap \(A3\)\. In finite samples with noise, the estimated spectrum is bimodal—r∗r^\{\*\}signal modes ofO​\(1\)O\(1\)andd−r∗d\-r^\{\*\}noise modes ofO​\(σ/n\)O\(\\sigma/\\sqrt\{n\}\)following a Marchenko–Pastur law—so the effective rank concentrates atr∗\+O​\(σ2/\(n​λr∗2\)\)r^\{\*\}\+O\(\\sigma^\{2\}/\(n\\,\\lambda\_\{r^\{\*\}\}^\{2\}\)\), with the finite\-sample MP correctionR^eff=Refftrue\+\(d−Refftrue\)/\(2​n\)\+O​\(1/n2\)\\hat\{R\}\_\{\\mathrm\{eff\}\}=R^\{\\mathrm\{true\}\}\_\{\\mathrm\{eff\}\}\+\(d\-R^\{\\mathrm\{true\}\}\_\{\\mathrm\{eff\}\}\)/\(2n\)\+O\(1/n^\{2\}\)\.

###### Proof sketch\.

The population claim is the standard implicit\-bias result for gradient flow onA=U​V⊤A=UV^\{\\top\}from vanishing initialization, which converges to the minimum nuclear\-norm solution; with a spectral gap this solution has exact rankr∗r^\{\*\}\. The finite\-sample spectrum splits into signal and Marchenko–Pastur noise bulks, and computing the singular\-value entropy over this bimodal spectrum giveslog⁡r∗\\log r^\{\*\}plus the stated correction\. A fully rigorous treatment of the finite\-width, finite\-sample regime is left to an extended version; our empirical recovery \(Tables[1](https://arxiv.org/html/2607.23050#S4.T1),[3](https://arxiv.org/html/2607.23050#S6.T3)\) confirms the concentration atr∗r^\{\*\}withCVseed≤0\.003\\mathrm\{CV\}\_\{\\mathrm\{seed\}\}\\leq 0\.003\. ∎

## Appendix BReal\-Corpus Diagnostics

This appendix documents the experiments behind the claim \(Sections[5](https://arxiv.org/html/2607.23050#S5),[6](https://arxiv.org/html/2607.23050#S6)\) that the attention\-native positive result does not transfer directly to natural language, and explains why the failure is structural rather than a tuning artifact\.

#### B\.1 Setup\.

We test the predictive\-bound hypothesis on six char\-level corpora spanning very different statistical structure: English prose \(english\), Shakespeare \(shakespeare\), Python source \(python\), Linux C source \(linux\), JSON \(json\), and ABC music notation \(music\)\. Models are causal Transformers withdmodel∈\{64,128,256\}d\_\{\\mathrm\{model\}\}\\in\\\{64,128,256\\\}, two layers, four heads, five seeds\. The data\-side predictor is a lag\-MI / operator estimate; the measured objects are the static kernelReff​\(WQ​K\)R\_\{\\mathrm\{eff\}\}\(W\_\{QK\}\), the mean realized attentionReff​\(𝔼​\[A\]\)R\_\{\\mathrm\{eff\}\}\(\\mathbb\{E\}\[A\]\), the per\-input realized rank𝔼X​\[Reff​\(A​\(X\)\)\]\\mathbb\{E\}\_\{X\}\[R\_\{\\mathrm\{eff\}\}\(A\(X\)\)\], the attention participation ratio, and the covariance\-whitened kernelReff​\(WQ​K​ΣX1/2\)R\_\{\\mathrm\{eff\}\}\(W\_\{QK\}\\Sigma\_\{X\}^\{1/2\}\)\.

#### B\.2 Main negative result\.

The data\-side predicted rank is*anti*\-correlated with the measured static\-kernel rank across corpora:Pearson​r=−0\.272\\mathrm\{Pearson\}\\ r=\-0\.272,Spearman​ρ=−0\.275\\mathrm\{Spearman\}\\ \\rho=\-0\.275\. This is not a weak positive miss; the sign is wrong\.

#### B\.3 The failure is structural\.

Three diagnostics rule out a noise explanation and identify the mechanism\.*\(i\) Not seed noise\.*All1818corpus×\\timessize cells haveCVseed∈\[0\.006,0\.087\]\\mathrm\{CV\}\_\{\\mathrm\{seed\}\}\\in\[0\.006,0\.087\], all below the0\.100\.10threshold—the measured rank is a clean, reproducible signal\.*\(ii\) It tracks fitting progress, not task complexity\.*Across corpora, the final loss and the measuredReffR\_\{\\mathrm\{eff\}\}correlate atr=−0\.79r=\-0\.79: corpora the model fits well \(low loss\) acquire*higher*measured rank, and hard corpora*lower*rank\. The measured quantity reflects optimization state, not intrinsic complexity—the direct cause of the anti\-correlation\.*\(iii\) Layer/channel dependence\.*In the largest model, Layer 0 hasReff≈13​–​19R\_\{\\mathrm\{eff\}\}\\approx 13\\text\{\-\-\}19while Layer 1 hasReff≈30​–​44R\_\{\\mathrm\{eff\}\}\\approx 30\\text\{\-\-\}44, and the FFN channel shows no relationship \(Pearson​\(H1,FFN stable rank\)=−0\.05\\mathrm\{Pearson\}\(H\_\{1\},\\ \\text\{FFN stable rank\}\)=\-0\.05\)\. A single per\-layer additive bound cannot capture this\.

#### B\.4 Measurement and predictor ablations \(Experiments A and C\)\.

Replacing the static kernel with realized\-attention measures \(mean attention, per\-input rank, participation, covariance\-whitened kernel\) does*not*restore corpus discrimination: across corpora every measure has low variation \(coefficient of variation0\.06​–​0\.200\.06\\text\{\-\-\}0\.20, versus0\.660\.66for the data\-side predictor\), so there is little task signal in the measurement to correlate against\. On the predictor side, estimating a single token\-mixing operator from data—via fixed\-random\-embedding ridge regression—collapses to full rank for every corpus \(random embeddings erase token\-identity structure\), and a token\-space repeat operator separates only one of six corpora\. No single low\-rank operator summarizes real\-text routing\.

#### Interpretation\.

The real\-corpus failure is*not*a contradiction of the attention\-native positive result\. It shows that natural language does not provide a matched teacher–student attention task with an identifiable low\-rankK∗K^\{\*\}: on trained language models the realized mixer is dominated by task\-independent structure \(locality, positional, attention\-sink priors\) rather than a task\-specific low\-rank routing kernel\. The synthetic attention\-native regime is therefore a controlled transfer test, and the real\-corpus bridge— recovering a task\-intrinsic attention rank from natural language—remains the central open problem\.

## Appendix CExperimental Details

All experiments in Section[6](https://arxiv.org/html/2607.23050#S6)use the controlled teacher–student setup described here\. The purpose of this setup is*not*to model natural language directly, but to test whether the attention\-native definition of intrinsic rank restores the rank–capacity transition once the teacher and student are placed in the*same operator class*\. The real\-language gap is treated separately \(Appendix B\) and is the central open problem of the paper\.

### C\.1Shared synthetic setup

Sequence lengthT=16T=16, model dimensiond=32d=32, intrinsic ranksr∗∈\{1,2,3,4,6,8\}r^\{\*\}\\in\\\{1,2,3,4,6,8\\\},n=3000n=3000samples, seeds\{0,1,2,3,4\}\\\{0,1,2,3,4\\\}, and observation noiseσ=0\.05\\sigma=0\.05\. The noise floor isσ2​T​d=0\.052×16×32=1\.28\\sigma^\{2\}Td=0\.05^\{2\}\\times 16\\times 32=1\.28, which all achievability results approach \(the slightly negative reported excess reflects finite\-sample fitting of a small fraction of the noise\)\.

### C\.2Teacher generation

The teacher kernel has rankr∗r^\{\*\}with a sharp spectral gap:

K∗=Ur​diag​\(λ1,…,λr∗\)​Ur⊤,λi=2−i−1r∗−1\(r∗\>1\),λ1=2​\(r∗=1\),K^\{\*\}=U\_\{r\}\\,\\mathrm\{diag\}\(\\lambda\_\{1\},\\dots,\\lambda\_\{r^\{\*\}\}\)\\,U\_\{r\}^\{\\top\},\\qquad\\lambda\_\{i\}=2\-\\frac\{i\-1\}\{r^\{\*\}\-1\}\\ \\ \(r^\{\*\}\>1\),\\quad\\lambda\_\{1\}=2\\ \(r^\{\*\}=1\),\(8\)withUr∈ℝd×r∗U\_\{r\}\\in\\mathbb\{R\}^\{d\\times r^\{\*\}\}having orthonormal columns andV∗∈ℝd×dV^\{\*\}\\in\\mathbb\{R\}^\{d\\times d\}Gaussian with1/d1/\\sqrt\{d\}scaling\. The two teacher types are

Y=\(X​K∗​X⊤\)​X​V∗\+EandY=softmax​\(X​K∗​X⊤\)​X​V∗\+E,Y=\(XK^\{\*\}X^\{\\top\}\)\\,XV^\{\*\}\+E\\qquad\\text\{and\}\\qquad Y=\\mathrm\{softmax\}\(XK^\{\*\}X^\{\\top\}\)\\,XV^\{\*\}\+E,\(9\)withE∼𝒩​\(0,σ2​I\)E\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\)and scores scaled by1/d1/\\sqrt\{d\}\.

### C\.3Student models

The rank\-constrained student uses a factored kernelK=WQ​WK⊤K=W\_\{Q\}W\_\{K\}^\{\\top\}withWQ,WK∈ℝd×rW\_\{Q\},W\_\{K\}\\in\\mathbb\{R\}^\{d\\times r\}, matching the teacher type \(linear or softmax\)\. The deficiency/achievability sweep runsr=1,…,2​r∗r=1,\\dots,2r^\{\*\}; over\-parameterized recovery usesr=dr=d\.

### C\.4Optimization

Adam with learning rate10−210^\{\-2\}\. The rank sweep trains for15001500steps; the over\-parameterized recovery trains for25002500steps\. Recovery uses small initialization, multiplyingWQ,WKW\_\{Q\},W\_\{K\}by0\.30\.3to induce the implicit low\-rank bias \(Theorem 4 analogue\)\. Full\-batch gradients are used throughout\.

### C\.5Metrics

- •Excess risk:L−σ2​T​dL\-\\sigma^\{2\}Td, whereLLis the converged MSE\.
- •Learned kernel effective rank:Reff​\(K\)R\_\{\\mathrm\{eff\}\}\(K\)of the trained student kernel\.
- •Seed CV:CVseed=stds​\(Reff​\(Ks\)\)/means​\(Reff​\(Ks\)\)\\mathrm\{CV\}\_\{\\mathrm\{seed\}\}=\\mathrm\{std\}\_\{s\}\(R\_\{\\mathrm\{eff\}\}\(K\_\{s\}\)\)/\\mathrm\{mean\}\_\{s\}\(R\_\{\\mathrm\{eff\}\}\(K\_\{s\}\)\)over seedsss\.
- •V\-known predictability: estimateKKfrom\(X,Y\)\(X,Y\)withV∗V^\{\*\}given \(regression on the recovered scores\), report numerical and energy effective rank\.
- •V\-unknown predictability: jointly estimate\(K,V\)\(K,V\)by alternating minimization \(plain and with nuclear\-norm shrinkage onKK\), report energy effective rank\.
- •Realized mixer rank:Reff​\(𝔼X​\[softmax​\(X​K​X⊤\)\]\)R\_\{\\mathrm\{eff\}\}\\\!\\big\(\\mathbb\{E\}\_\{X\}\[\\mathrm\{softmax\}\(XKX^\{\\top\}\)\]\\big\), the effective rank of the data\-averaged realized attention map \(informative under softmax; saturates for the linear teacher\)\.

### C\.6Reproduction command

```
python exp_attn_native.py \
  --teacher both \
  --T 16 --d 32 \
  --r_stars 1 2 3 4 6 8 \
  --seeds 0 1 2 3 4 \
  --n 3000 \
  --results_dir results_attn_native_grid
```

The linear surrogate validation \(Section[4](https://arxiv.org/html/2607.23050#S4), Table[1](https://arxiv.org/html/2607.23050#S4.T1)\) and the interpolation ladder \(Section[5](https://arxiv.org/html/2607.23050#S5), Table[2](https://arxiv.org/html/2607.23050#S5.T2)\) use the same conventions withT=16,d=12,r∗=3T\{=\}16,d\{=\}12,r^\{\*\}\{=\}3\(surrogate\) andT=12,d=10,r∗=3T\{=\}12,d\{=\}10,r^\{\*\}\{=\}3\(ladder\), respectively; see the accompanying code for the exact scripts\.

## Appendix DSpectral\-Gap Ablation

We test whether the phase transition of Section[6](https://arxiv.org/html/2607.23050#S6)is an artifact of the sharp\-gap teacher kernel by varying the spectrum ofK∗K^\{\*\}over its topr∗r^\{\*\}directions, holding all else fixed\. We consider three regimes:sharp\(λi=2−\(i−1\)/\(r∗−1\)\\lambda\_\{i\}=2\-\(i\{\-\}1\)/\(r^\{\*\}\{\-\}1\), the main\-paper setting\),power\-law\(λi=i−1\\lambda\_\{i\}=i^\{\-1\}\), andflat/no\-gap \(λi=1\\lambda\_\{i\}=1\)\. Table[5](https://arxiv.org/html/2607.23050#A4.T5)reports a lightweight linear\-QK ablation \(T=10T\{=\}10,d=12d\{=\}12,r∗=4r^\{\*\}\{=\}4\) isolating the spectral\-profile effect while keeping the teacher–student class matched\.

Table 5:Spectral\-gap ablation \(linear\-QK teacher,r∗=4r^\{\*\}\{=\}4\)\. Achievability \(excess@r∗≈0r^\{\*\}\\approx 0\) and recovery hold in every regime\. The recovered energy effective rank tracks the teacher kernel’s energy rank—exactlyr∗r^\{\*\}for a flat spectrum, and the smaller dominant\-direction count when the spectrum decays\. Sharp gaps give abrupt transitions; gradual spectra give softer knees\.The key invariant: the recovered energy effective rank equals the teacher kernel’s energy rank in all regimes \(e\.g\.4\.00→4\.004\.00\\to 4\.00for the flat spectrum, where every direction contributes equally, and3\.57→3\.573\.57\\to 3\.57,2\.43→2\.442\.43\\to 2\.44when a few directions dominate\)\. Thus the energy effective rank—not an integer “rank”—is the quantity that tracks task capacity, and the rank–capacity relationship survives the removal of the spectral gap\. The transition merely changes shape \(sharp gap⇒\\Rightarrowabrupt; gradual spectrum⇒\\Rightarrowsoft knee\), consistent with the spectral\-gap dependence in Theorem 1\.

## References

- Aghajanyan et al\. \(2021\)Armen Aghajanyan, Sonal Gupta, and Luke Zettlemoyer\.Intrinsic dimensionality explains the effectiveness of language model fine\-tuning\.In*Annual Meeting of the Association for Computational Linguistics \(ACL\)*, 2021\.
- Arora et al\. \(2019\)Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo\.Implicit regularization in deep matrix factorization\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2019\.
- Bartlett and Mendelson \(2002\)Peter L Bartlett and Shahar Mendelson\.Rademacher and gaussian complexities: Risk bounds and structural results\.*Journal of Machine Learning Research*, 3:463–482, 2002\.
- Berger \(1971\)Toby Berger\.*Rate Distortion Theory: A Mathematical Basis for Data Compression*\.Prentice\-Hall, 1971\.
- Cover and Thomas \(2006\)Thomas M Cover and Joy A Thomas\.*Elements of Information Theory*\.Wiley\-Interscience, 2nd edition, 2006\.
- Dong et al\. \(2021\)Yihe Dong, Jean\-Baptiste Cordonnier, and Andreas Loukas\.Attention is not all you need: Pure attention loses rank doubly exponentially with depth\.In*International Conference on Machine Learning \(ICML\)*, 2021\.
- Eckart and Young \(1936\)Carl Eckart and Gale Young\.The approximation of one matrix by another of lower rank\.*Psychometrika*, 1\(3\):211–218, 1936\.
- Gunasekar et al\. \(2017\)Suriya Gunasekar, Blake Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nathan Srebro\.Implicit regularization in matrix factorization\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2017\.
- Hoffmann et al\. \(2022\)Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al\.Training compute\-optimal large language models\.*arXiv preprint arXiv:2203\.15556*, 2022\.
- Hu et al\. \(2022\)Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen\-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen\.Lora: Low\-rank adaptation of large language models\.In*International Conference on Learning Representations \(ICLR\)*, 2022\.
- Kaplan et al\. \(2020\)Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei\.Scaling laws for neural language models\.*arXiv preprint arXiv:2001\.08361*, 2020\.
- Kobayashi et al\. \(2020\)Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui\.Attention is not only a weight: Analyzing transformers with vector norms\.In*Conference on Empirical Methods in Natural Language Processing \(EMNLP\)*, 2020\.
- Roy and Vetterli \(2007\)Olivier Roy and Martin Vetterli\.The effective rank: A measure of effective dimensionality\.*European Signal Processing Conference \(EUSIPCO\)*, pages 606–610, 2007\.
- Srebro et al\. \(2004\)Nathan Srebro, Jason Rennie, and Tommi Jaakkola\.Maximum\-margin matrix factorization\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2004\.
- Vapnik \(1998\)Vladimir N Vapnik\.*Statistical Learning Theory*\.Wiley, 1998\.

Similar Articles

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention

arXiv cs.LG

This paper applies Marchenko-Pastur random matrix theory to pre-trained attention weights, separating each projection matrix into a random-like bulk and spectral outliers. Causal experiments show zeroing these outliers in Mistral-7B drives performance near random chance, revealing that spectral outliers encode dominant learned structure across 11 transformers.

Relevant and Irrelevant: A Renormalization Group Analysis of Transformer Attention

arXiv cs.LG

This paper applies Wilsonian renormalization group theory to analyze Transformer attention as a perturbation of the MLP residual-stack fixed point, determining whether attention is relevant or irrelevant based on data correlation length. Experiments on synthetic Markov chains confirm that attention's relevance depends on the spectral structure of the data-generating process, with the first-layer head dominating the transition.