Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability

arXiv cs.LG Papers

Summary

This paper introduces a theoretical framework for identifying model-intrinsic structure in mechanistic interpretability using Koopman operator theory, proving the first identifiability theorem for a mechanistic-interpretability primitive with empirical validation on GPT-2, Gemma-2-2B, and Qwen3-8B-Base.

arXiv:2608.10172v1 Announce Type: new Abstract: Mechanistic interpretability explains models by identifying circuits inside them, but has no way to tell whether a circuit is a property of the model or an artifact of the method that found it. Sparse autoencoders illustrate the problem: different seeds and widths recover materially different features from the same activations, and no theory says whether that variability is incidental or structural. We put dictionary learning for interpretability on an identifiability footing. Treating the forward pass as a controlled dynamical system with depth as time and lifting it with the Koopman operator yields a finite linear realisation whose \emph{spectrum} is a coordinate-free property of the model. We prove the spectrum is recoverable from $M$ calibration samples at rate $M^{-1/2}$ up to permutation - to our knowledge the first identifiability theorem for a mechanistic-interpretability primitive, with a matching minimax lower bound, a median-of-means variant for heavy-tailed activations, and a dissociation theorem: whenever the realisation is non-normal, the directions carrying activation variance and the directions carrying information across depth cannot coincide. The identifiable object and the legible object are not the same object. On GPT-2 small, Gemma-2-2B and Qwen3-8B-Base the spectrum converges everywhere and attains the predicted exponent on Qwen3-8B-Base ($0.506 \pm 0.031$); shortfalls collapse onto one curve against each cell's sample threshold. Koopman modes beat random directions but lose to principal components on indirect-object identification, with the gap decaying $4.1\times$ in depth-distance, as the theorem predicts. The Koopman spectrum is an identifiable, model-intrinsic fingerprint with a stated error bar, not a legible decomposition.
Original Article
View Cached Full Text

Cached at: 08/12/26, 08:28 AM

# Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability
Source: [https://arxiv.org/html/2608.10172](https://arxiv.org/html/2608.10172)
###### Abstract

Mechanistic interpretability explains models by identifying circuits inside them, but has no way to tell whether a circuit is a property of the model or an artifact of the method that found it\. Sparse autoencoders illustrate the problem: different seeds and widths recover materially different features from the same activations, and no theory says whether that variability is incidental or structural\. We put dictionary learning for interpretability on an identifiability footing\. Treating the forward pass as a controlled dynamical system with depth as time and lifting it with the Koopman operator yields a finite linear realisation whose*spectrum*is a coordinate\-free property of the model\. We prove the spectrum is recoverable fromMMcalibration samples at rateM−1/2M^\{\-1/2\}up to permutation \- to our knowledge the first identifiability theorem for a mechanistic\-interpretability primitive, with a matching minimax lower bound, a median\-of\-means variant for heavy\-tailed activations, and a dissociation theorem: whenever the realisation is non\-normal, the directions carrying activation variance and the directions carrying information across depth cannot coincide\.*The identifiable object and the legible object are not the same object\.*On GPT\-2 small, Gemma\-2\-2B and Qwen3\-8B\-Base the spectrum converges everywhere and attains the predicted exponent on Qwen3\-8B\-Base \(−0\.506±0\.031\-0\.506\\pm 0\.031\); shortfalls collapse onto one curve against each cell’s sample threshold\. Koopman modes beat random directions but lose to principal components on indirect\-object identification, with the gap decaying4\.1×4\.1\\timesin depth\-distance, as the theorem predicts\. The Koopman spectrum is an identifiable, model\-intrinsic fingerprint with a stated error bar, not a legible decomposition\.

## 1Introduction

Mechanistic interpretability \(MI\) explains neural network behaviour by decomposing a trained model into interacting components\. The past several years have produced a rich toolkit for this purpose\. Sparse autoencoders \(SAEs\) and their descendants\[[34](https://arxiv.org/html/2608.10172#bib.bib3),[59](https://arxiv.org/html/2608.10172#bib.bib4),[60](https://arxiv.org/html/2608.10172#bib.bib5),[23](https://arxiv.org/html/2608.10172#bib.bib6),[46](https://arxiv.org/html/2608.10172#bib.bib7),[32](https://arxiv.org/html/2608.10172#bib.bib8)\]decompose activations into overcomplete sparse dictionaries; transcoders and attribution graphs\[[18](https://arxiv.org/html/2608.10172#bib.bib15),[2](https://arxiv.org/html/2608.10172#bib.bib16),[47](https://arxiv.org/html/2608.10172#bib.bib17)\]extend the decomposition to sublayer computations; causal\-intervention methods\[[73](https://arxiv.org/html/2608.10172#bib.bib19),[50](https://arxiv.org/html/2608.10172#bib.bib20),[13](https://arxiv.org/html/2608.10172#bib.bib22),[66](https://arxiv.org/html/2608.10172#bib.bib23),[31](https://arxiv.org/html/2608.10172#bib.bib24),[43](https://arxiv.org/html/2608.10172#bib.bib25)\]localise behaviour to components of the computational graph; and distributed alignment search\[[24](https://arxiv.org/html/2608.10172#bib.bib28),[25](https://arxiv.org/html/2608.10172#bib.bib29),[79](https://arxiv.org/html/2608.10172#bib.bib30)\]aligns learned interventions with hypothesised causal variables\.

Here we address a question that cuts across all of them:*do discovered circuits reflect intrinsic properties of the trained model, or the procedure used to find them?*If the former, they provide a sound basis for scientific understanding, engineering intervention, and safety arguments\[[11](https://arxiv.org/html/2608.10172#bib.bib31)\]; if the latter, any conclusion drawn from them inherits the method’s contingencies\.

Current evidence suggests the question is not merely philosophical\. SAEs exhibit substantial run\-to\-run variability: different training seeds and dictionary widths recover materially different features on the same activations\[[6](https://arxiv.org/html/2608.10172#bib.bib9),[57](https://arxiv.org/html/2608.10172#bib.bib11),[37](https://arxiv.org/html/2608.10172#bib.bib10)\]\. Learned dictionaries exhibit*absorption*\[[9](https://arxiv.org/html/2608.10172#bib.bib12)\]and feature*splitting*, and residual\-stream components remain uncaptured by any known variant\[[21](https://arxiv.org/html/2608.10172#bib.bib14)\]\. What is missing across every paradigm is a theorem asserting that the discovered structure is an invariant of the model\.

We argue that closing this gap requires a mechanistic primitive that admits a*proof of identifiability*, and we obtain one by changing what we take a transformer forward pass to*be*\. Letxℓ∈ℝdx\_\{\\ell\}\\in\\mathbb\{R\}^\{d\}denote the residual\-stream state at a fixed token position at layerℓ\\ell, and letuℓu\_\{\\ell\}collect the layer’s exogenous inputs \- the attention writes contributed by other token positions\. The forward pass is then a discrete\-time controlled dynamical system, the*depth recurrence*

xℓ\+1=F​\(xℓ,uℓ\):=xℓ\+aℓ​\(xℓ,uℓ\)\+mℓ​\(xℓ\+aℓ​\(xℓ,uℓ\)\),x\_\{\\ell\+1\}\\;=\\;F\(x\_\{\\ell\},u\_\{\\ell\}\)\\;:=\\;x\_\{\\ell\}\+a\_\{\\ell\}\(x\_\{\\ell\},u\_\{\\ell\}\)\+m\_\{\\ell\}\\bigl\(x\_\{\\ell\}\+a\_\{\\ell\}\(x\_\{\\ell\},u\_\{\\ell\}\)\\bigr\),\(1\)whereaℓ​\(xℓ,uℓ\)a\_\{\\ell\}\(x\_\{\\ell\},u\_\{\\ell\}\)is the layer’s attention sublayer output \- the query/key/value mixing ofxℓx\_\{\\ell\}against the controluℓu\_\{\\ell\}\- andmℓm\_\{\\ell\}is the layer’s MLP sublayer, applied to the post\-attention residualxℓ\+aℓx\_\{\\ell\}\+a\_\{\\ell\}\. Depth plays the role of time, attention writes enter as control inputs, and the MLP acts as the autonomous nonlinear dynamics\. This viewpoint has precedent in the neural\-ODE line\[[75](https://arxiv.org/html/2608.10172#bib.bib32),[30](https://arxiv.org/html/2608.10172#bib.bib33),[10](https://arxiv.org/html/2608.10172#bib.bib34)\], but to our knowledge it has not been exploited for the identifiability of mechanistic structure\.

Once the forward pass is viewed as a dynamical system, a classical object becomes available: the*Koopman operator*𝒦\\mathcal\{K\}\[[41](https://arxiv.org/html/2608.10172#bib.bib35),[51](https://arxiv.org/html/2608.10172#bib.bib36),[8](https://arxiv.org/html/2608.10172#bib.bib37)\]\. Acting on observablesψ\\psi, it evolves them by composition with the dynamics,

\(𝒦​ψ\)​\(x,u\)=ψ​\(F​\(x,u\)\)\.\(\\mathcal\{K\}\\,\\psi\)\(x,u\)\\;=\\;\\psi\\bigl\(F\(x,u\)\\bigr\)\.\(2\)AlthoughFFis nonlinear in the state,𝒦\\mathcal\{K\}is linear inψ\\psi, transferring the nonlinearity to the lifting from states to observables\. If a dictionaryΨ=\(ψ1,…,ψN\)\\Psi=\(\\psi\_\{1\},\\dots,\\psi\_\{N\}\)spans an invariant subspace,𝒦\\mathcal\{K\}restricts to a finite matrixA∈ℂN×NA\\in\\mathbb\{C\}^\{N\\times N\}\. We call the eigenpairs ofAAthe*Koopman modes*of the transformer and the frameworkKoopman spectral analysis \(KSA\)\. The eigenvalues ofAAare properties of the underlying operator: unique up to permutation, invariant under any change of dictionary basis, and \- this is the content of our main theorem \- recoverable from finite data at a parametric rate\.

The theory says that a coordinate\-free object exists and can be estimated; the experiments say what that object is good for, and \- equally important \- what it is not good for\. Both halves are reported here, including the measurements that bound the claim\.

##### Contributions\. •We formalise dictionary learning for interpretability so that*identifiability*is a well\-posed question, isolating Koopman invariance as the property the standard SAE objective omits \([Section˜3](https://arxiv.org/html/2608.10172#S3)\)\.•We show that a𝒦\\mathcal\{K\}\-invariant dictionary induces a Koopman realisation\(A,B\)\(A,B\)that exists, is unique up to a change of basis, and whose spectrum is a coordinate\-free invariant of the transformer \([Theorem˜5\.1](https://arxiv.org/html/2608.10172#S5.Thmtheorem1),[Section˜5](https://arxiv.org/html/2608.10172#S5)\)\.•Our main result proves that this spectrum is identifiable fromMMcalibration samples at the parametric rateM−1/2M^\{\-1/2\}, up to permutation \- to our knowledge the first identifiability theorem for a mechanistic\-interpretability primitive \([Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1),[Section˜6](https://arxiv.org/html/2608.10172#S6)\)\.•We sharpen the theorem’s sample threshold\. The spectral\-gap factor that made the original threshold unreachable governs only the*eigenvector*guarantee; stating the eigenvalue result in optimal\-matching form removes it and lowers the threshold by nine orders of magnitude on GPT\-2 small \([Theorem˜6\.8](https://arxiv.org/html/2608.10172#S6.Thmtheorem8)\)\.•We characterise the problem and not only the estimator: a minimax lower bound shows theM−1/2M^\{\-1/2\}rate is optimal \([Theorem˜7\.1](https://arxiv.org/html/2608.10172#S7.Thmtheorem1)\), and a median\-of\-means variant extends the guarantee to heavy\-tailed activations \([Theorem˜7\.2](https://arxiv.org/html/2608.10172#S7.Thmtheorem2)\)\. The latter’s prediction fails when tested, and the measurement says why \- the lifting, not the residual stream, decides the tails \([Section˜11\.2](https://arxiv.org/html/2608.10172#S11.SS2)\)\.•We prove a*dissociation*: non\-normality forces the activations’ principal directions and the Koopman modes apart, since alignment would require perfect eigenvector conditioningκ2​\(V\)=1\\kappa\_\{2\}\(V\)=1; measured conditioning of10110^\{1\}–10210^\{2\}makes the divergence unavoidable, and an explicit family shows the misalignment saturates at orthogonality \([Theorem˜8\.1](https://arxiv.org/html/2608.10172#S8.Thmtheorem1),[Proposition˜8\.2](https://arxiv.org/html/2608.10172#S8.Thmtheorem2)\)\.•We derive practical consequences: SAE non\-identifiability is structural rather than algorithmic, with an explicit invariance penalty as remedy \([Corollary˜9\.5](https://arxiv.org/html/2608.10172#S9.Thmtheorem5)\); cross\-model universality becomes a testable spectral criterion \([Corollary˜9\.6](https://arxiv.org/html/2608.10172#S9.Thmtheorem6)\); the intervention calculus of MI is algebraically complete on the spectral\-projector algebra \([Theorem˜9\.2](https://arxiv.org/html/2608.10172#S9.Thmtheorem2)\); and balanced truncation of the realisation admits an instance\-dependent error certificate strictly sharper than Enns–Glover in the low\-effective\-rank regime attention occupies \([Theorem˜10\.1](https://arxiv.org/html/2608.10172#S10.Thmtheorem1)\)\.•We evaluate on three pretrained transformers at10510^\{5\}–10610^\{6\}calibration samples\. The predictedM−1/2M^\{\-1/2\}rate is observed on Qwen3\-8B\-Base; the invariance penalty of \([66](https://arxiv.org/html/2608.10172#S9.E66)\) reduces the split\-half spectral distance by41%41\\%at matched sparsity \([Section˜11](https://arxiv.org/html/2608.10172#S11)\)\.•We report what fails, because it bounds the claim: Koopman modes beat random directions but lose to principal components at predicting IOI ablation effects \([Section˜11\.3](https://arxiv.org/html/2608.10172#S11.SS3)\); the SAE invariance gap reverses under variance\-based feature selection \([Section˜11\.5](https://arxiv.org/html/2608.10172#S11.SS5)\); and the universality criterion declares two seed replicas of one architecture distinct \([Section˜11\.7](https://arxiv.org/html/2608.10172#S11.SS7)\)\. Together these give the paper’s central claim:*the identifiable object and the legible object are not the same object*\.

[Section˜2](https://arxiv.org/html/2608.10172#S2)places the work against the interpretability, Koopman, model\-reduction and identifiability literatures\.[Section˜3](https://arxiv.org/html/2608.10172#S3)formalises the problem and defines spectral identifiability\.[Section˜4](https://arxiv.org/html/2608.10172#S4)constructs the depth recurrence, the Koopman realisation and the EDMDc estimator, and states the three assumptions\.[Section˜5](https://arxiv.org/html/2608.10172#S5)proves existence, uniqueness and basis\-independence\.[Section˜6](https://arxiv.org/html/2608.10172#S6)contains the finite\-sample identifiability theorem and its complete proof, together with the gap\-free strengthening and the clustered\-spectrum extension\.[Section˜7](https://arxiv.org/html/2608.10172#S7)proves the minimax lower bound and the heavy\-tailed variant\.[Section˜8](https://arxiv.org/html/2608.10172#S8)proves the modal–principal dissociation and quantifies it\.[Section˜9](https://arxiv.org/html/2608.10172#S9)develops the consequences for MI: algebraic completeness, SAE non\-identifiability, and universality\.[Section˜10](https://arxiv.org/html/2608.10172#S10)proves the instance\-dependent reduction bound\.[Section˜11](https://arxiv.org/html/2608.10172#S11)reports the measurements on three pretrained models, in full\.[Section˜12](https://arxiv.org/html/2608.10172#S12)states the discussions\.

## 2Background and Related Work

Feature\-centric interpretability decomposes activations into interpretable dictionaries, most prominently using sparse autoencoders and their extensions to sublayer computations\[[19](https://arxiv.org/html/2608.10172#bib.bib43),[7](https://arxiv.org/html/2608.10172#bib.bib1),[34](https://arxiv.org/html/2608.10172#bib.bib3),[59](https://arxiv.org/html/2608.10172#bib.bib4),[60](https://arxiv.org/html/2608.10172#bib.bib5),[23](https://arxiv.org/html/2608.10172#bib.bib6),[68](https://arxiv.org/html/2608.10172#bib.bib2),[46](https://arxiv.org/html/2608.10172#bib.bib7),[32](https://arxiv.org/html/2608.10172#bib.bib8),[18](https://arxiv.org/html/2608.10172#bib.bib15),[2](https://arxiv.org/html/2608.10172#bib.bib16),[47](https://arxiv.org/html/2608.10172#bib.bib17)\]\. Intervention\-centric methods instead infer circuits through causal interventions such as activation patching, attribution patching, and automated circuit discovery\[[73](https://arxiv.org/html/2608.10172#bib.bib19),[50](https://arxiv.org/html/2608.10172#bib.bib20),[13](https://arxiv.org/html/2608.10172#bib.bib22),[66](https://arxiv.org/html/2608.10172#bib.bib23),[31](https://arxiv.org/html/2608.10172#bib.bib24),[43](https://arxiv.org/html/2608.10172#bib.bib25),[24](https://arxiv.org/html/2608.10172#bib.bib28),[25](https://arxiv.org/html/2608.10172#bib.bib29),[79](https://arxiv.org/html/2608.10172#bib.bib30)\]\. The indirect\-object\-identification \(IOI\) circuit of\[[74](https://arxiv.org/html/2608.10172#bib.bib21)\]is the best\-characterised product of the second programme and serves as our semantic testbed in[Section˜11\.3](https://arxiv.org/html/2608.10172#S11.SS3)\.

A steady stream of results documents the instability of these decompositions across seeds, widths, and training runs\[[6](https://arxiv.org/html/2608.10172#bib.bib9),[9](https://arxiv.org/html/2608.10172#bib.bib12),[21](https://arxiv.org/html/2608.10172#bib.bib14),[37](https://arxiv.org/html/2608.10172#bib.bib10),[57](https://arxiv.org/html/2608.10172#bib.bib11)\]\. The prevailing framing treats this as a tuning or evaluation problem\.[Corollary˜9\.5](https://arxiv.org/html/2608.10172#S9.Thmtheorem5)gives a different reading: the variability is a structural consequence of an objective that never mentions the property identifiability requires\.

Classical identifiability results in independent component analysis\[[12](https://arxiv.org/html/2608.10172#bib.bib46),[35](https://arxiv.org/html/2608.10172#bib.bib47)\], its nonlinear variants\[[36](https://arxiv.org/html/2608.10172#bib.bib49),[39](https://arxiv.org/html/2608.10172#bib.bib50)\], and disentanglement\[[48](https://arxiv.org/html/2608.10172#bib.bib51),[63](https://arxiv.org/html/2608.10172#bib.bib52)\]identify latent factors under statistical independence or auxiliary\-variable structure, and do so up to sign, scale*and*permutation\. Causal abstraction\[[24](https://arxiv.org/html/2608.10172#bib.bib28),[25](https://arxiv.org/html/2608.10172#bib.bib29)\]identifies causal variables relative to a hypothesised graph\. System identification\[[64](https://arxiv.org/html/2608.10172#bib.bib53),[76](https://arxiv.org/html/2608.10172#bib.bib54)\]identifies a linear system up to a similarity transformation, under persistent excitation\.

Our result differs in what is identified and in how much ambiguity survives\. We identify the*spectrum*of a Koopman compression, and the only residual ambiguity is a permutation of an unordered multiset: there is no sign, scale, or rotational freedom left over, because eigenvalues are rigid under similarity \([Remark˜6\.2](https://arxiv.org/html/2608.10172#S6.Thmtheorem2)\)\. The persistent\-excitation condition we inherit from system identification \([Assumption˜2](https://arxiv.org/html/2608.10172#Thmassumption2)\) plays its usual role\.

Koopman theory\[[41](https://arxiv.org/html/2608.10172#bib.bib35),[51](https://arxiv.org/html/2608.10172#bib.bib36),[8](https://arxiv.org/html/2608.10172#bib.bib37)\]represents nonlinear dynamics linearly through an operator acting on observables\. Data\-driven estimators include dynamic mode decomposition\[[62](https://arxiv.org/html/2608.10172#bib.bib55)\], extended DMD\[[77](https://arxiv.org/html/2608.10172#bib.bib56)\], DMD with control\[[58](https://arxiv.org/html/2608.10172#bib.bib58)\], and kernel\[[78](https://arxiv.org/html/2608.10172#bib.bib57)\]and generator\-based\[[40](https://arxiv.org/html/2608.10172#bib.bib60)\]variants\.\[[42](https://arxiv.org/html/2608.10172#bib.bib59)\]establish asymptotic convergence of EDMD to the Koopman operator in the strong operator topology as the dictionary and sample size grow together\. Machine\-learning applications include network pruning\[[61](https://arxiv.org/html/2608.10172#bib.bib63)\], sequence forecasting\[[4](https://arxiv.org/html/2608.10172#bib.bib62)\], and the analysis of iterative algorithms\[[16](https://arxiv.org/html/2608.10172#bib.bib61)\]\.

Our contribution sits at a different level of the same story\. For a*fixed*finite dictionary satisfying invariance, we identify the population limit as a specific coordinate\-free object and prove a*finite\-sample*identifiability rate for its spectrum, with explicit constants, on the dynamical systems defined by transformer depth recurrences \([Remark˜5\.8](https://arxiv.org/html/2608.10172#S5.Thmtheorem8)makes the comparison precise\)\.

Balanced truncation\[[53](https://arxiv.org/html/2608.10172#bib.bib64),[3](https://arxiv.org/html/2608.10172#bib.bib40)\]orders modes by joint controllability and observability and yields the classical Enns–Glover error bound\[[27](https://arxiv.org/html/2608.10172#bib.bib38),[1](https://arxiv.org/html/2608.10172#bib.bib39),[29](https://arxiv.org/html/2608.10172#bib.bib66)\]; frequency\-weighted variants\[[22](https://arxiv.org/html/2608.10172#bib.bib65)\]accommodate anisotropic input or output distributions\. Because the classical bound is worst\-case over inputs, it does not tighten when inputs concentrate on a low\-dimensional subspace, as transformer attention writes empirically do\[[17](https://arxiv.org/html/2608.10172#bib.bib68),[26](https://arxiv.org/html/2608.10172#bib.bib67)\]\.[Theorem˜10\.1](https://arxiv.org/html/2608.10172#S10.Thmtheorem1)gives an instance\-dependent bound that exploits this concentration and is strictly sharper in the low\-effective\-rank regime;[Section˜11](https://arxiv.org/html/2608.10172#S11)measures an effective rank against an ambient control dimension of40964096on Qwen3\-8B\-Base\.

The residual connection invites viewing a deep network as a discretised differential equation\[[75](https://arxiv.org/html/2608.10172#bib.bib32),[30](https://arxiv.org/html/2608.10172#bib.bib33),[10](https://arxiv.org/html/2608.10172#bib.bib34)\]\. For transformers, this perspective has produced analyses of attention as a mean\-field particle system\[[26](https://arxiv.org/html/2608.10172#bib.bib67)\], of rank collapse with depth\[[17](https://arxiv.org/html/2608.10172#bib.bib68)\], and of in\-context learning as an emergent process\[[56](https://arxiv.org/html/2608.10172#bib.bib44),[55](https://arxiv.org/html/2608.10172#bib.bib45),[54](https://arxiv.org/html/2608.10172#bib.bib41)\]\. These formalisms describe the dynamics but do not yield an identifiable mechanistic decomposition\.[Theorem˜5\.1](https://arxiv.org/html/2608.10172#S5.Thmtheorem1)defines the Koopman compression as the object of study, and[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)identifies it from finite data\.

## 3Problem Formulation

### 3\.1The dictionary\-learning objective

Letx∈ℝdx\\in\\mathbb\{R\}^\{d\}be a residual\-stream activation drawn from a distributionμ\\muinduced by running a transformer on a corpus\. A sparse autoencoder learns a dictionaryD∈ℝd×ND\\in\\mathbb\{R\}^\{d\\times N\}withN≫dN\\gg dand an encoder producing sparse codesz​\(x\)∈ℝNz\(x\)\\in\\mathbb\{R\}^\{N\}, by minimising

ℒSAE​\(D\)=𝔼x∼μ​\[‖x−D​z​\(x\)‖22\+λ​‖z​\(x\)‖1\]\.\\mathcal\{L\}\_\{\\mathrm\{SAE\}\}\(D\)\\;=\\;\\mathbb\{E\}\_\{x\\sim\\mu\}\\Bigl\[\\,\\bigl\\\|x\-D\\,z\(x\)\\bigr\\\|\_\{2\}^\{2\}\\;\+\\;\\lambda\\,\\bigl\\\|z\(x\)\\bigr\\\|\_\{1\}\\Bigr\]\.\(3\)The objective is a statement about*one layer’s activations in isolation*: reconstructxx, and do so sparsely — with no reference to the maps that produced it or will consume it\. This observation drives everything that follows\.

To make the comparison across paradigms precise, we first fix what kind of object is under discussion\.

###### Definition 3\.1\(Mechanistic primitive\)\.

A*mechanistic primitive*for a transformerFFis a triple\(𝒟,E,ℐ\)\(\\mathcal\{D\},E,\\mathcal\{I\}\)in which𝒟\\mathcal\{D\}is a finite index set,E:𝒳→ℂ𝒟E:\\mathcal\{X\}\\to\\mathbb\{C\}^\{\\mathcal\{D\}\}is a component\-wise measurable encoding, andℐ\\mathcal\{I\}is a set of bounded linear operators on the observable space, interpreted as interventions\. Sparse autoencoders, transcoders, causal\-abstraction methods, and KSA all instantiate this schema\.

### 3\.2What identifiability would mean

Identifiability asks whether the object recovered is determined by the model or by the recovery procedure\. For a dictionary this has two parts: the estimand must be well defined independently of coordinates, and it must be recoverable from finite data\. The first part is a condition on the dictionary\.

###### Definition 3\.2\(Dictionary richness\)\.

A mechanistic primitive\(𝒟,E,ℐ\)\(\\mathcal\{D\},E,\\mathcal\{I\}\)satisfies*dictionary richness*if

- \(DR1\)\{ed\}d∈𝒟\\\{e\_\{d\}\\\}\_\{d\\in\\mathcal\{D\}\}are linearly independent inL2​\(μ\)L^\{2\}\(\\mu\);
- \(DR2\)ℋE≔span⁡\{ed:d∈𝒟\}\\mathcal\{H\}\_\{E\}\\coloneqq\\operatorname\{span\}\\\{e\_\{d\}:d\\in\\mathcal\{D\}\\\}is*𝒦\\mathcal\{K\}\-invariant*: for everyψ∈ℋE\\psi\\in\\mathcal\{H\}\_\{E\}andν\\nu\-a\.e\. controluu, the functionx↦ψ​\(F​\(x,u\)\)x\\mapsto\\psi\(F\(x,u\)\)again lies inℋE\\mathcal\{H\}\_\{E\}\.

\(DR1\) is ordinary linear independence and is satisfied by essentially any trained dictionary\. \(DR2\) is the substantive condition, and it is exactly the one \([3](https://arxiv.org/html/2608.10172#S3.E3)\) does not mention: it demands that the span be closed under the dynamics, so that evolving a feature by one layer keeps it inside the dictionary\. Neither condition presupposes KSA; both are properties of the primitive itself\.

The second part is a statement about estimation\.

###### Definition 3\.3\(Spectral identifiability\)\.

LetA^M\\hat\{A\}\_\{M\}be an estimator of the realisationAAfromMMcalibration samples\. The primitive is*spectrally identifiable at rate*r​\(M\)r\(M\)if there is a permutationπM\\pi\_\{M\}of\{1,…,N\}\\\{1,\\dots,N\\\}such that, with high probability,

maxk⁡\|λk​\(A^M\)−λπM​\(k\)​\(A\)\|=O​\(r​\(M\)\)\.\\max\_\{k\}\\bigl\|\\lambda\_\{k\}\(\\hat\{A\}\_\{M\}\)\-\\lambda\_\{\\pi\_\{M\}\(k\)\}\(A\)\\bigr\|\\;=\\;O\\bigl\(r\(M\)\\bigr\)\.\(4\)

Permutation is the right and only quotient: eigenvalues form an unordered multiset, so no estimator recovers a labelling, and nothing weaker need be quotiented out\. This is markedly stronger than what is available for SAEs, whose features are identified at best up to permutation*and*sign*and*scale, and in practice not at all\.

Two things must now be supplied\.[Section˜4](https://arxiv.org/html/2608.10172#S4)constructs a primitive that satisfies[Definition˜3\.2](https://arxiv.org/html/2608.10172#S3.Thmtheorem2)and exhibits the estimand;[Section˜5](https://arxiv.org/html/2608.10172#S5)shows the estimand is well defined; and[Section˜6](https://arxiv.org/html/2608.10172#S6)shows it is recoverable at the parametric rate, making[Definition˜3\.3](https://arxiv.org/html/2608.10172#S3.Thmtheorem3)non\-vacuous\.

## 4Transformer Depth Dynamics and the Koopman Realisation

This section constructs the primitive that satisfies[Definition˜3\.2](https://arxiv.org/html/2608.10172#S3.Thmtheorem2): the controlled depth recurrence, the observable space and controlled Koopman operator, the finite\-dimensional realisation, the EDMDc estimator, the spectral decomposition and the definition of a Koopman circuit, and the three assumptions under which everything later is proved\.

### 4\.1The depth recurrence as a controlled dynamical system

We consider decoder\-only transformers with pre\-normalisation and RMSNorm, the architectural configuration used by Llama 3\[[28](https://arxiv.org/html/2608.10172#bib.bib70)\]and Gemma 2\[[67](https://arxiv.org/html/2608.10172#bib.bib71)\]and descended from the original transformer of\[[71](https://arxiv.org/html/2608.10172#bib.bib69)\]\. LetL∈ℕL\\in\\mathbb\{N\}denote the number of layers,d∈ℕd\\in\\mathbb\{N\}the residual\-stream dimension,H∈ℕH\\in\\mathbb\{N\}the number of attention heads per layer, andT∈ℕT\\in\\mathbb\{N\}the number of tokens in a prompt\. For each prompt and each token positiont∈\{1,…,T\}t\\in\\\{1,\\dots,T\\\}, letxℓ\(t\)∈ℝdx\_\{\\ell\}^\{\(t\)\}\\in\\mathbb\{R\}^\{d\}denote the residual\-stream state at layerℓ∈\{0,…,L\}\\ell\\in\\\{0,\\dots,L\\\}\. The layer update decomposes as

x~ℓ\(t\)\\displaystyle\\widetilde\{x\}\_\{\\ell\}^\{\(t\)\}=xℓ\(t\)\+Attnℓ​\(RMSNorm​\(xℓ\(1:T\)\)\)\(t\)⏟aℓ\(t\),\\displaystyle=x\_\{\\ell\}^\{\(t\)\}\+\\underbrace\{\\mathrm\{Attn\}\_\{\\ell\}\\\!\\bigl\(\\mathrm\{RMSNorm\}\(x\_\{\\ell\}^\{\(1:T\)\}\)\\bigr\)^\{\(t\)\}\}\_\{\\displaystyle a\_\{\\ell\}^\{\(t\)\}\},\(5\)xℓ\+1\(t\)\\displaystyle x\_\{\\ell\+1\}^\{\(t\)\}=x~ℓ\(t\)\+MLPℓ​\(RMSNorm​\(x~ℓ\(t\)\)\)⏟mℓ\(t\)\.\\displaystyle=\\widetilde\{x\}\_\{\\ell\}^\{\(t\)\}\+\\underbrace\{\\mathrm\{MLP\}\_\{\\ell\}\\\!\\bigl\(\\mathrm\{RMSNorm\}\(\\widetilde\{x\}\_\{\\ell\}^\{\(t\)\}\)\\bigr\)\}\_\{\\displaystyle m\_\{\\ell\}^\{\(t\)\}\}\.\(6\)
Fix a token positionttand drop the superscript\. The attention outputaℓa\_\{\\ell\}depends on the residual\-stream states at*other*token positions in the same layer, which are themselves determined by their own trajectories\. From the vantage point of the trajectory\{xℓ\}ℓ=0L\\\{x\_\{\\ell\}\\\}\_\{\\ell=0\}^\{L\}at tokentt, therefore,aℓa\_\{\\ell\}enters as an exogenous input\. We denote this input byuℓ∈ℝpu\_\{\\ell\}\\in\\mathbb\{R\}^\{p\}\(withp=dp=dunder the choiceuℓ=aℓu\_\{\\ell\}=a\_\{\\ell\}; the ablationuℓ=\(aℓ,mℓ\)u\_\{\\ell\}=\(a\_\{\\ell\},m\_\{\\ell\}\)is considered separately\)\.

###### Definition 4\.1\(Controlled depth recurrence\)\.

Fix a token positiont∈\{1,…,T\}t\\in\\\{1,\\dots,T\\\}\. Let𝒳⊆ℝd\\mathcal\{X\}\\subseteq\\mathbb\{R\}^\{d\}be the state space and𝒰⊆ℝp\\mathcal\{U\}\\subseteq\\mathbb\{R\}^\{p\}the control space\. The*controlled depth recurrence*at tokenttis the discrete\-time dynamical system

xℓ\+1=F​\(xℓ,uℓ\)=xℓ\+uℓ\+m​\(xℓ\+uℓ\),ℓ=0,…,L−1,x\_\{\\ell\+1\}=F\(x\_\{\\ell\},u\_\{\\ell\}\)=x\_\{\\ell\}\+u\_\{\\ell\}\+m\\\!\\bigl\(x\_\{\\ell\}\+u\_\{\\ell\}\\bigr\),\\qquad\\ell=0,\\dots,L\-1,\(7\)wherem​\(⋅\)=MLP​\(RMSNorm​\(⋅\)\)m\(\\cdot\)=\\mathrm\{MLP\}\(\\mathrm\{RMSNorm\}\(\\cdot\)\)andF:𝒳×𝒰→𝒳F:\\mathcal\{X\}\\times\\mathcal\{U\}\\to\\mathcal\{X\}is the layer map\.

##### Data\-generating process\.

The calibration corpus induces an empirical distribution over trajectories\{\(xℓ,uℓ\)\}ℓ=0L−1\\\{\(x\_\{\\ell\},u\_\{\\ell\}\)\\\}\_\{\\ell=0\}^\{L\-1\}\. We assume this distribution converges, as the corpus grows, to a joint measureμ⊗ν\\mu\\otimes\\nuon𝒳×𝒰\\mathcal\{X\}\\times\\mathcal\{U\}, whereμ\\muis an invariant measure for the state process andν\\nuis a stationary distribution for the controls\. Existence of such measures under mild regularity is standard\[[8](https://arxiv.org/html/2608.10172#bib.bib37)\]; we work with them as given\. Because residual\-stream norms grow with depth,μ\\muis in practice the layer marginalμℓ\\mu\_\{\\ell\}rather than one depth\-invariant law;[Section˜11\.8\.3](https://arxiv.org/html/2608.10172#S11.SS8.SSS3)measures the growth and shows it is recovered as a Koopman mode once the dictionary can express it\.

### 4\.2Observables, the Koopman operator, and the finite\-dimensional realisation

The Koopman operator lifts the nonlinear mapFFto a linear operator on an infinite\-dimensional space of*observables*\[[41](https://arxiv.org/html/2608.10172#bib.bib35)\]\.

###### Definition 4\.3\(Observable space\)\.

The observable space isℋ≔L2​\(𝒳,μ\)\\mathcal\{H\}\\coloneqq L^\{2\}\(\\mathcal\{X\},\\mu\), the space of measurableψ:𝒳→ℂ\\psi:\\mathcal\{X\}\\to\\mathbb\{C\}with∫\|ψ\|2​𝑑μ<∞\\int\|\\psi\|^\{2\}\\,d\\mu<\\infty, equipped with the inner product⟨ψ,φ⟩ℋ=∫ψ​φ¯​𝑑μ\\langle\\psi,\\varphi\\rangle\_\{\\mathcal\{H\}\}=\\int\\psi\\,\\bar\{\\varphi\}\\,d\\mu\. A*vector\-valued observable*is a stackingΨ=\(ψ1,…,ψN\)⊤:𝒳→ℂN\\Psi=\(\\psi\_\{1\},\\dots,\\psi\_\{N\}\)^\{\\top\}:\\mathcal\{X\}\\to\\mathbb\{C\}^\{N\}with eachψi∈ℋ\\psi\_\{i\}\\in\\mathcal\{H\}\.

###### Definition 4\.4\(Controlled Koopman operator\)\.

The*controlled Koopman operator*𝒦\\mathcal\{K\}associated withFFsends an observable to its composition with the layer map:

\(𝒦​ψ\)​\(x,u\)≔ψ​\(F​\(x,u\)\),ψ∈ℋ\.\(\\mathcal\{K\}\\,\\psi\)\(x,u\)\\;\\coloneqq\\;\\psi\\\!\\bigl\(F\(x,u\)\\bigr\),\\qquad\\psi\\in\\mathcal\{H\}\.\(8\)AlthoughFFis nonlinear inxx,𝒦\\mathcal\{K\}is linear inψ\\psi\.

We write𝒦u\\mathcal\{K\}\_\{u\}for the family of operators obtained by fixing the control,\(𝒦u​ψ\)​\(x\)=ψ​\(F​\(x,u\)\)\(\\mathcal\{K\}\_\{u\}\\psi\)\(x\)=\\psi\(F\(x,u\)\), and𝒦¯≔𝔼u∼ν​\[𝒦u\]\\bar\{\\mathcal\{K\}\}\\coloneqq\\mathbb\{E\}\_\{u\\sim\\nu\}\[\\mathcal\{K\}\_\{u\}\]for the control average\. Below, “the Koopman operator” without qualification means𝒦¯\\bar\{\\mathcal\{K\}\}\.

###### Definition 4\.5\(Koopman\-invariant subspace\)\.

LetℋN≔span⁡\{ψ1,…,ψN\}⊆ℋ\\mathcal\{H\}\_\{N\}\\coloneqq\\operatorname\{span\}\\\{\\psi\_\{1\},\\dots,\\psi\_\{N\}\\\}\\subseteq\\mathcal\{H\}be the linear span of a dictionaryΨ\\Psi\.ℋN\\mathcal\{H\}\_\{N\}is*𝒦\\mathcal\{K\}\-invariant*if for everyψ∈ℋN\\psi\\in\\mathcal\{H\}\_\{N\}and everyu∈𝒰u\\in\\mathcal\{U\}, the functionx↦ψ​\(F​\(x,u\)\)x\\mapsto\\psi\(F\(x,u\)\)belongs toℋN\\mathcal\{H\}\_\{N\}as a function ofxx\.

𝒦\\mathcal\{K\}\-invariance is a substantive condition on the dictionary: it requires the finite\-dimensional span to be closed under the Koopman flow\. When it holds, the operator restricts to a finite\-dimensional linear map𝒦\|ℋN:ℋN→ℋN\\mathcal\{K\}\|\_\{\\mathcal\{H\}\_\{N\}\}:\\mathcal\{H\}\_\{N\}\\to\\mathcal\{H\}\_\{N\}admitting a matrix representation on any basis\.

###### Definition 4\.6\(Koopman realisation\)\.

AssumeℋN\\mathcal\{H\}\_\{N\}is𝒦\\mathcal\{K\}\-invariant\. The*Koopman realisation*ofFFonΨ\\Psiis the pair\(A,B\)∈ℂN×N×ℂN×p\(A,B\)\\in\\mathbb\{C\}^\{N\\times N\}\\times\\mathbb\{C\}^\{N\\times p\}satisfying

Ψ​\(F​\(x,u\)\)=A​Ψ​\(x\)\+B​u\(μ⊗ν\)​\-a\.e\.\\Psi\\bigl\(F\(x,u\)\\bigr\)\\;=\\;A\\,\\Psi\(x\)\\;\+\\;B\\,u\\qquad\(\\mu\\otimes\\nu\)\\text\{\-a\.e\.\}\(9\)HereAAis the matrix representation of𝒦\|ℋN\\mathcal\{K\}\|\_\{\\mathcal\{H\}\_\{N\}\}in the basis\{ψi\}\\\{\\psi\_\{i\}\\\}andBBencodes the linear response ofΨ\\Psito the control input\. WhenℋN\\mathcal\{H\}\_\{N\}is not𝒦\\mathcal\{K\}\-invariant, we take\(A,B\)\(A,B\)to be the population least\-squares approximation

\(A,B\)≔arg​minA′,B′⁡𝔼\(x,u\)∼μ⊗ν​‖Ψ​\(F​\(x,u\)\)−A′​Ψ​\(x\)−B′​u‖22,\(A,B\)\\;\\coloneqq\\;\\operatorname\*\{arg\\,min\}\_\{A^\{\\prime\},B^\{\\prime\}\}\\;\\mathbb\{E\}\_\{\(x,u\)\\sim\\mu\\otimes\\nu\}\\bigl\\\|\\Psi\\bigl\(F\(x,u\)\\bigr\)\-A^\{\\prime\}\\,\\Psi\(x\)\-B^\{\\prime\}\\,u\\bigr\\\|\_\{2\}^\{2\},\(10\)and denote the projection residual byε​\(x,u\)≔Ψ​\(F​\(x,u\)\)−A​Ψ​\(x\)−B​u\\varepsilon\(x,u\)\\coloneqq\\Psi\(F\(x,u\)\)\-A\\,\\Psi\(x\)\-B\\,u\.

The realisation is defined at the level of an abstract basis ofℋN\\mathcal\{H\}\_\{N\}\. Different bases yield different matrix representations related by similarity: ifΨ′​\(x\)=T​Ψ​\(x\)\\Psi^\{\\prime\}\(x\)=T\\,\\Psi\(x\)for invertibleT∈ℂN×NT\\in\\mathbb\{C\}^\{N\\times N\}, thenA′=T​A​T−1A^\{\\prime\}=TAT^\{\-1\}andB′=T​BB^\{\\prime\}=TB\.[Theorem˜5\.1](https://arxiv.org/html/2608.10172#S5.Thmtheorem1)formalises that these are the*only*ambiguities in the definition of\(A,B\)\(A,B\)\.

###### Definition 4\.7\(Read\-out\)\.

A*read\-out*is a linear operatorC∈ℂq×NC\\in\\mathbb\{C\}^\{q\\times N\}that maps the observable at any layer to aqq\-dimensional target:yℓ=C​Ψ​\(xℓ\)∈ℂqy\_\{\\ell\}=C\\,\\Psi\(x\_\{\\ell\}\)\\in\\mathbb\{C\}^\{q\}\. Common choices includeC=WU⊤​PC=W\_\{U\}^\{\\top\}P\(projection onto specified logit directions\) orCCencoding a probe direction\.

The triple\(A,B,C\)\(A,B,C\)constitutes the KSA realisation, a discrete\-time linear time\-invariant \(LTI\) system onℂN\\mathbb\{C\}^\{N\}\. We denote its input–output map byG:\{uℓ\}↦\{yℓ\}G:\\\{u\_\{\\ell\}\\\}\\mapsto\\\{y\_\{\\ell\}\\\}and its transfer function byG^​\(z\)=C​\(z​I−A\)−1​B\\hat\{G\}\(z\)=C\\,\(zI\-A\)^\{\-1\}\\,B, with‖G^‖ℋ∞\\\|\\hat\{G\}\\\|\_\{\\mathcal\{H\}\_\{\\infty\}\}the Hardyℋ∞\\mathcal\{H\}\_\{\\infty\}\-norm on the closed unit disk\.[Section˜10](https://arxiv.org/html/2608.10172#S10)bounds the error of truncating this system tor≪Nr\\ll Nmodes\.

![Refer to caption](https://arxiv.org/html/2608.10172v1/x1.png)Figure 1:Transformer forward pass as a controlled dynamical system \(bottom\) and its Koopman lift \(top\)\. With a Koopman\-invariant dictionary the lifted dynamics become exactly linear,Ψ​\(F​\(x,u\)\)=A​Ψ​\(x\)\+B​u\\Psi\(F\(x,u\)\)=A\\,\\Psi\(x\)\+B\\,u\. The basis\-independent spectrum ofAAis the identifiable object \([Theorem˜5\.1](https://arxiv.org/html/2608.10172#S5.Thmtheorem1)\)\.
### 4\.3The EDMDc estimator

The Koopman realisation\(A,B\)\(A,B\)is not directly accessible; we estimate it from calibration data using the extended dynamic mode decomposition with control \(EDMDc\) of\[[77](https://arxiv.org/html/2608.10172#bib.bib56),[58](https://arxiv.org/html/2608.10172#bib.bib58)\]\.

##### Calibration data\.

Let𝒟\\mathcal\{D\}be a calibration corpus consisting ofMMlayer–token samples\{\(xℓ\(i\),uℓ\(i\)\)\}i=1M\\\{\(x\_\{\\ell\}^\{\(i\)\},u\_\{\\ell\}^\{\(i\)\}\)\\\}\_\{i=1\}^\{M\}obtained by running the transformer on a corpus of prompts and caching residual\-stream states and control inputs at the analysis position\. Assemble the*snapshot matrices*

X\\displaystyle X≔\[Ψ​\(x0\(1\)\),…,Ψ​\(xL−1\(1\)\),Ψ​\(x0\(2\)\),…\]∈ℂN×M,\\displaystyle\\coloneqq\\bigl\[\\,\\Psi\(x\_\{0\}^\{\(1\)\}\),\\dots,\\Psi\(x\_\{L\-1\}^\{\(1\)\}\),\\Psi\(x\_\{0\}^\{\(2\)\}\),\\dots\\bigr\]\\in\\mathbb\{C\}^\{N\\times M\},\(11\)Y\\displaystyle Y≔\[Ψ​\(x1\(1\)\),…,Ψ​\(xL\(1\)\),Ψ​\(x1\(2\)\),…\]∈ℂN×M,\\displaystyle\\coloneqq\\bigl\[\\,\\Psi\(x\_\{1\}^\{\(1\)\}\),\\dots,\\Psi\(x\_\{L\}^\{\(1\)\}\),\\Psi\(x\_\{1\}^\{\(2\)\}\),\\dots\\bigr\]\\in\\mathbb\{C\}^\{N\\times M\},\(12\)Ξ\\displaystyle\\Xi≔\[u0\(1\),…,uL−1\(1\),u0\(2\),…\]∈ℂp×M\.\\displaystyle\\coloneqq\\bigl\[\\,u\_\{0\}^\{\(1\)\},\\dots,u\_\{L\-1\}^\{\(1\)\},u\_\{0\}^\{\(2\)\},\\dots\\bigr\]\\in\\mathbb\{C\}^\{p\\times M\}\.\(13\)The columns ofYYare the next\-layer observables corresponding to the columns ofXX;Ξ\\Xicollects the control inputs applied betweenXXandYY\.

###### Definition 4\.8\(EDMDc estimator\)\.

Fix a regularisation parameterγ≥0\\gamma\\geq 0\. The*EDMDc estimator*onMMsamples is

\(A^M,B^M\)≔arg​minA′∈ℂN×N,B′∈ℂN×p‖Y−A′​X−B′​Ξ‖F2\+γ​\(‖A′‖F2\+‖B′‖F2\)\.\\begin\{split\}\(\\hat\{A\}\_\{M\},\\hat\{B\}\_\{M\}\)\\;\\coloneqq\\;\\operatorname\*\{arg\\,min\}\_\{A^\{\\prime\}\\in\\mathbb\{C\}^\{N\\times N\},\\,B^\{\\prime\}\\in\\mathbb\{C\}^\{N\\times p\}\}\\;&\\\|Y\-A^\{\\prime\}X\-B^\{\\prime\}\\Xi\\\|\_\{F\}^\{2\}\\\\ &\+\\gamma\\bigl\(\\\|A^\{\\prime\}\\\|\_\{F\}^\{2\}\+\\\|B^\{\\prime\}\\\|\_\{F\}^\{2\}\\bigr\)\.\\end\{split\}\(14\)
Whenγ=0\\gamma=0and the stacked snapshot matrix\[XΞ\]\\bigl\[\\begin\{smallmatrix\}X\\\\ \\Xi\\end\{smallmatrix\}\\bigr\]has full row rankN\+pN\+p, the solution is

\[A^M,B^M\]=Y​\[XΞ\]†\.\[\\hat\{A\}\_\{M\},\\;\\hat\{B\}\_\{M\}\]\\;=\\;Y\\,\\begin\{bmatrix\}X\\\\ \\Xi\\end\{bmatrix\}^\{\\\!\\dagger\}\.\(15\)

The estimator is the empirical counterpart of the population minimiser \([10](https://arxiv.org/html/2608.10172#S4.E10)\)\. Under𝒦\\mathcal\{K\}\-invariance and persistent excitation it converges to\(A,B\)\(A,B\); the finite\-sample rate is the subject of[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)\. All experiments useγ=0\\gamma=0, the unregularised pseudo\-inverse solution\.

### 4\.4Spectral decomposition and the definition of a Koopman circuit

We now attach the interpretive apparatus to the estimated realisation\.

##### Eigendecomposition\.

AssumeA^M∈ℂN×N\\hat\{A\}\_\{M\}\\in\\mathbb\{C\}^\{N\\times N\}is diagonalisable \(the Jordan\-form extension is discussed in[Remark˜5\.6](https://arxiv.org/html/2608.10172#S5.Thmtheorem6)\)\. There exist right eigenvectors\{vk\}k=1N\\\{v\_\{k\}\\\}\_\{k=1\}^\{N\}and left eigenvectors\{ϕk\}k=1N\\\{\\phi\_\{k\}\\\}\_\{k=1\}^\{N\}inℂN\\mathbb\{C\}^\{N\}such that

A^M​vk\\displaystyle\\hat\{A\}\_\{M\}\\,v\_\{k\}=λk​vk,\\displaystyle=\\lambda\_\{k\}\\,v\_\{k\},\(16\)ϕk⊤​A^M\\displaystyle\\phi\_\{k\}^\{\\top\}\\,\\hat\{A\}\_\{M\}=λk​ϕk⊤,\\displaystyle=\\lambda\_\{k\}\\,\\phi\_\{k\}^\{\\top\},\(17\)biorthogonally normalised so thatϕj⊤​vk=δj​k\\phi\_\{j\}^\{\\top\}v\_\{k\}=\\delta\_\{jk\}\. The eigendecomposition isA^M=∑k=1Nλk​vk​ϕk⊤\\hat\{A\}\_\{M\}=\\sum\_\{k=1\}^\{N\}\\lambda\_\{k\}\\,v\_\{k\}\\phi\_\{k\}^\{\\top\}\.

###### Definition 4\.9\(Spectral projector\)\.

The*spectral projector*associated with the eigenpair\(λk,vk,ϕk\)\(\\lambda\_\{k\},v\_\{k\},\\phi\_\{k\}\)is

Πk≔vk​ϕk⊤∈ℂN×N\.\\Pi\_\{k\}\\;\\coloneqq\\;v\_\{k\}\\,\\phi\_\{k\}^\{\\top\}\\;\\in\\;\\mathbb\{C\}^\{N\\times N\}\.\(18\)The family\{Πk\}k=1N\\\{\\Pi\_\{k\}\\\}\_\{k=1\}^\{N\}satisfiesΠj​Πk=δj​k​Πk\\Pi\_\{j\}\\Pi\_\{k\}=\\delta\_\{jk\}\\Pi\_\{k\}and∑kΠk=IN\\sum\_\{k\}\\Pi\_\{k\}=I\_\{N\}\.

##### Koopman eigenfunctions\.

Corresponding to each left eigenvectorϕk∈ℂN\\phi\_\{k\}\\in\\mathbb\{C\}^\{N\}is a scalar observableφk∈ℋN\\varphi\_\{k\}\\in\\mathcal\{H\}\_\{N\}obtained by contraction with the dictionary,

φk​\(x\)≔ϕk⊤​Ψ​\(x\),\\varphi\_\{k\}\(x\)\\;\\coloneqq\\;\\phi\_\{k\}^\{\\top\}\\,\\Psi\(x\),\(19\)which is a*Koopman eigenfunction*of𝒦\|ℋN\\mathcal\{K\}\|\_\{\\mathcal\{H\}\_\{N\}\}in the classical sense: on autonomous dynamics \(u=0u=0\),φk​\(F​\(x,0\)\)=λk​φk​\(x\)\\varphi\_\{k\}\(F\(x,0\)\)=\\lambda\_\{k\}\\,\\varphi\_\{k\}\(x\)μ\\mu\-a\.e\. We distinguish notationally between the eigenfunctionφk:𝒳→ℂ\\varphi\_\{k\}:\\mathcal\{X\}\\to\\mathbb\{C\}and its left\-eigenvector representativeϕk∈ℂN\\phi\_\{k\}\\in\\mathbb\{C\}^\{N\}, a coordinate vector in the dictionary basis\.

###### Definition 4\.10\(Koopman circuit\)\.

A*Koopman circuit*associated with the realisation\(A^M,B^M,C\)\(\\hat\{A\}\_\{M\},\\hat\{B\}\_\{M\},C\)is a quadruple\(λk,ϕk,vk,Πk\)\(\\lambda\_\{k\},\\phi\_\{k\},v\_\{k\},\\Pi\_\{k\}\)consisting of an eigenvalueλk∈ℂ\\lambda\_\{k\}\\in\\mathbb\{C\}, its left eigenvectorϕk\\phi\_\{k\}\(equivalently its eigenfunctionφk\\varphi\_\{k\}\), its right eigenvectorvkv\_\{k\}\(the*Koopman mode*\), and its spectral projectorΠk=vk​ϕk⊤\\Pi\_\{k\}=v\_\{k\}\\phi\_\{k\}^\{\\top\}\.

The interpretive taxonomy attached to the eigenvalue is the classical one:\|λk\|\>1\|\\lambda\_\{k\}\|\>1amplifying,\|λk\|≈1\|\\lambda\_\{k\}\|\\approx 1transport,\|λk\|<1\|\\lambda\_\{k\}\|<1decaying,arg⁡λk≠0\\arg\\lambda\_\{k\}\\neq 0rotational\. Modes live in observable spaceℂN\\mathbb\{C\}^\{N\};[Section˜11\.8\.4](https://arxiv.org/html/2608.10172#S11.SS8.SSS4)gives the fitted linear map back to residual\-stream coordinatesℝd\\mathbb\{R\}^\{d\}that every mode\-level empirical claim depends on\.

### 4\.5Assumptions

All results below rest on three assumptions, stated here and discussed immediately after\.

###### Assumption 1\(Dictionary𝒦\\mathcal\{K\}\-invariance\)\.

The dictionary spanℋN=span⁡\{ψ1,…,ψN\}\\mathcal\{H\}\_\{N\}=\\operatorname\{span\}\\\{\\psi\_\{1\},\\dots,\\psi\_\{N\}\\\}is𝒦\\mathcal\{K\}\-invariant in the sense of[Definition˜4\.5](https://arxiv.org/html/2608.10172#S4.Thmtheorem5)\. Equivalently, for everyψ∈ℋN\\psi\\in\\mathcal\{H\}\_\{N\}andν\\nu\-almost everyu∈𝒰u\\in\\mathcal\{U\}, the compositionψ∘F​\(⋅,u\)\\psi\\circ F\(\\cdot,u\)belongs toℋN\\mathcal\{H\}\_\{N\}\. This is \(DR2\) of[Definition˜3\.2](https://arxiv.org/html/2608.10172#S3.Thmtheorem2)\.

###### Assumption 2\(Persistent excitation\)\.

There existsη\>0\\eta\>0such that the empirical control covarianceΣM≔1M​Ξ​Ξ⊤∈ℂp×p\\Sigma\_\{M\}\\coloneqq\\frac\{1\}\{M\}\\Xi\\Xi^\{\\top\}\\in\\mathbb\{C\}^\{p\\times p\}satisfies

λmin​\(ΣM\)≥η\\lambda\_\{\\min\}\\\!\\bigl\(\\Sigma\_\{M\}\\bigr\)\\;\\geq\\;\\eta\(20\)with probability at least1−δM1\-\\delta\_\{M\}, whereδM→0\\delta\_\{M\}\\to 0asM→∞M\\to\\infty\. Equivalently, the control sequence excites allppinput directions in the limit\.

###### Assumption 3\(Spectral separation\)\.

There existsΔ\>0\\Delta\>0such that the eigenvalues of the true Koopman compressionAAare pairwise separated:

\|λi​\(A\)−λj​\(A\)\|≥Δfor all​i≠j\.\\bigl\|\\lambda\_\{i\}\(A\)\-\\lambda\_\{j\}\(A\)\\bigr\|\\;\\geq\\;\\Delta\\qquad\\text\{for all \}i\\neq j\.\(21\)

### 4\.6Three modelling conventions, all measured

Three modelling choices define the estimand, and two of them constrain what it can be\. We state them here and measure each of them in[Section˜11](https://arxiv.org/html/2608.10172#S11), rather than leaving them implicit\.

*First, the realisation is depth\-indexed\.*Sincemℓm\_\{\\ell\}is layer\-dependent, we fitA^ℓ\\hat\{A\}\_\{\\ell\}separately at each analysed layer and the spectrum is a familyℓ↦σ​\(Aℓ\)\\ell\\mapsto\\sigma\(A\_\{\\ell\}\)\. This is resolved rather than assumed: adjacent\-layer spectral distance exceeds the split\-half resolution floor by median factors of3\.223\.22,11\.9711\.97and6\.126\.12across the suite, and a pooledAApredicts held\-out transitions39%39\\%,22%22\\%and16%16\\%worse \([Section˜11\.8\.1](https://arxiv.org/html/2608.10172#S11.SS8.SSS1)\)\.[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)is proved at fixedℓ\\elland needs no depth\-homogeneity\.

*Second, the realisation is defined relative to the layer marginal\.*Residual\-stream norm grows with depth \(4\.3×4\.3\\timeson GPT\-2 small across the analysed range\), so a single invariant lawμ\\muis not available and the per\-layer fit enforces the marginalμℓ\\mu\_\{\\ell\}instead\. Addinglog⁡‖xℓ‖\\log\\\|x\_\{\\ell\}\\\|to the dictionary then recovers that growth as its own real mode on all three models \(1\.1371\.137,1\.1071\.107,1\.2911\.291against measured growth1\.1111\.111,1\.0771\.077,1\.3261\.326\) — a prediction that could have failed and did not \([Section˜11\.8\.3](https://arxiv.org/html/2608.10172#S11.SS8.SSS3)\)\.

*Third, and least innocuous, exogeneity of the control is a choice\.*The writeuℓu\_\{\\ell\}is computed from the same residual stream, so treating it as exogenous is a modelling convention rather than a fact\. The largest canonical correlation betweenΨ​\(xℓ\)\\Psi\(x\_\{\\ell\}\)anduℓu\_\{\\ell\}isρ1=0\.82\\rho\_\{1\}=0\.82\(GPT\-2 small\),0\.960\.96\(Gemma\-2\-2B\) and0\.950\.95\(Qwen3\-8B\-Base\), though the controls are mostly unexplained by the state \(R2=0\.08R^\{2\}=0\.08,0\.220\.22,0\.160\.16\)\. Residualising the control against the lifted state \(Frisch–Waugh–Lovell\) shifts the spectrum by6\.8×6\.8\\times,9\.1×9\.1\\timesand5\.2×5\.2\\timesthe split\-half floor, so the two conventions yield genuinely different estimands at our resolution\. We report the naive realisation as primary and the residualised version alongside \([Section˜11\.8\.2](https://arxiv.org/html/2608.10172#S11.SS8.SSS2)\)\.

## 5Existence and Uniqueness of the Koopman Realisation

This section establishes the foundational claim on which everything later rests: given a𝒦\\mathcal\{K\}\-invariant dictionary, the object we call “the Koopman realisation” \- the matrixAAassociated with the operator’s action on the dictionary span \- exists, is uniquely determined up to a change of basis, and has a spectrumσ​\(A\)\\sigma\(A\)that is a coordinate\-free property of the pair\(F,ℋN\)\(F,\\mathcal\{H\}\_\{N\}\)\. Uniqueness up to similarity is the correct invariance property: it says the eigenvalues, spectral projectors and modal structure ofAAare attributes of the transformer and dictionary, not of the particular ordering or scaling of the basis vectors\{ψi\}\\\{\\psi\_\{i\}\\\}\. Without this,[Definition˜3\.3](https://arxiv.org/html/2608.10172#S3.Thmtheorem3)would have no estimand to speak of\.

The proof proceeds in three preparatory lemmas followed by consolidation\.[Lemma˜5\.2](https://arxiv.org/html/2608.10172#S5.Thmtheorem2)shows that under[Assumption˜1](https://arxiv.org/html/2608.10172#Thmassumption1)the pointwise Koopman matrixAuA\_\{u\}exists and is unique for each control value\.[Lemma˜5\.3](https://arxiv.org/html/2608.10172#S5.Thmtheorem3)establishes the similarity transformation under change of basis\.[Lemma˜5\.4](https://arxiv.org/html/2608.10172#S5.Thmtheorem4)establishes the spectral inclusionσ​\(A\)⊆σ​\(𝒦¯\)\\sigma\(A\)\\subseteq\\sigma\(\\bar\{\\mathcal\{K\}\}\)via standard operator\-theoretic facts about invariant subspaces\. We then assemble the theorem and derive a corollary formalising the coordinate\-free invariants of Koopman circuits\.

### 5\.1Regularity conditions and statement

We work throughout under[Assumption˜1](https://arxiv.org/html/2608.10172#Thmassumption1)and the linear independence condition

\(R1\)GΨ≔𝔼μ​\[Ψ​\(x\)​Ψ​\(x\)∗\]≻0,\\text\{\(R1\)\}\\qquad G\_\{\\Psi\}\\;\\coloneqq\\;\\mathbb\{E\}\_\{\\mu\}\\\!\\bigl\[\\Psi\(x\)\\,\\Psi\(x\)^\{\*\}\\bigr\]\\;\\succ\\;0,\(22\)which asserts that\{ψi\}i=1N\\\{\\psi\_\{i\}\\\}\_\{i=1\}^\{N\}are linearly independent as elements ofL2​\(μ\)L^\{2\}\(\\mu\)\- that is, \(DR1\) of[Definition˜3\.2](https://arxiv.org/html/2608.10172#S3.Thmtheorem2)\. We further impose the standard state–control decorrelation condition

\(R2\)𝔼\(x,u\)∼μ⊗ν​\[Ψ​\(x\)​u∗\]=𝔼μ​\[Ψ​\(x\)\]​𝔼ν​\[u\]∗,\\text\{\(R2\)\}\\qquad\\mathbb\{E\}\_\{\(x,u\)\\sim\\mu\\otimes\\nu\}\\\!\\bigl\[\\Psi\(x\)\\,u^\{\*\}\\bigr\]\\;=\\;\\mathbb\{E\}\_\{\\mu\}\[\\Psi\(x\)\]\\;\\mathbb\{E\}\_\{\\nu\}\[u\]^\{\*\},\(23\)which follows automatically whenxxanduuare independent under the joint measure\. Without loss of generality we centre the controls,𝔼ν​\[u\]=0\\mathbb\{E\}\_\{\\nu\}\[u\]=0, so that \(R2\) becomes𝔼​\[Ψ​\(x\)​u∗\]=0\\mathbb\{E\}\[\\Psi\(x\)\\,u^\{\*\}\]=0\.

###### Theorem 5\.1\(Existence and uniqueness of the Koopman realisation\)\.

Assume[Assumption˜1](https://arxiv.org/html/2608.10172#Thmassumption1)together with \([22](https://arxiv.org/html/2608.10172#S5.E22)\) and \([23](https://arxiv.org/html/2608.10172#S5.E23)\)\. Then:

1. \(a\)*\(Existence and uniqueness\.\)*There exists a unique matrixA∈ℂN×NA\\in\\mathbb\{C\}^\{N\\times N\}satisfying A​Ψ​\(x\)=𝔼u∼ν​\[Ψ​\(F​\(x,u\)\)\]μ​\-a\.e\.​x∈𝒳\.A\\,\\Psi\(x\)\\;=\\;\\mathbb\{E\}\_\{u\\sim\\nu\}\\\!\\bigl\[\\,\\Psi\\bigl\(F\(x,u\)\\bigr\)\\,\\bigr\]\\qquad\\mu\\text\{\-a\.e\.\}\\ x\\in\\mathcal\{X\}\.\(24\)Equivalently,AAis the matrix representation of the bounded linear operator𝒦¯\|ℋN\\bar\{\\mathcal\{K\}\}\|\_\{\\mathcal\{H\}\_\{N\}\}on the basis\{ψi\}i=1N\\\{\\psi\_\{i\}\\\}\_\{i=1\}^\{N\}\.
2. \(b\)*\(Basis\-change equivariance\.\)*Under any change of basisΨ′​\(x\)=T​Ψ​\(x\)\\Psi^\{\\prime\}\(x\)=T\\,\\Psi\(x\)withT∈G​L​\(N,ℂ\)T\\in GL\(N,\\mathbb\{C\}\), the realisation transforms asA′=T​A​T−1A^\{\\prime\}=TAT^\{\-1\}\.
3. \(c\)*\(Spectral invariance\.\)*The spectrumσ​\(A\)⊆ℂ\\sigma\(A\)\\subseteq\\mathbb\{C\}is invariant under basis change and constitutes a coordinate\-free invariant of the pair\(F,ℋN\)\(F,\\mathcal\{H\}\_\{N\}\)\.
4. \(d\)*\(Spectral embedding\.\)*σ​\(A\)⊆σ​\(𝒦¯\)\\sigma\(A\)\\subseteq\\sigma\(\\bar\{\\mathcal\{K\}\}\)\.
5. \(e\)*\(Agreement with the least\-squares realisation\.\)*The matrixAAcoincides with the population minimiserALSA\_\{\\mathrm\{LS\}\}of the least\-squares problem \([10](https://arxiv.org/html/2608.10172#S4.E10)\)\.

Parts \(a\)–\(c\) constitute the existence and uniqueness statement and make[Definition˜3\.3](https://arxiv.org/html/2608.10172#S3.Thmtheorem3)well posed by supplying an estimand independent of the dictionary’s basis\. Part \(d\) is the spectral\-inclusion claim: the recovered eigenvalues are genuine Koopman eigenvalues of the transformer’s depth dynamics, not artefacts of the finite truncation\. Part \(e\) connects the abstract compression to the estimable object of[Definition˜4\.8](https://arxiv.org/html/2608.10172#S4.Thmtheorem8): it isALSA\_\{\\mathrm\{LS\}\}that EDMDc targets, and \(e\) shows the two coincide under the stated conditions\.

### 5\.2Preparatory lemmas

###### Lemma 5\.2\(Pointwise Koopman matrix\)\.

Under[Assumption˜1](https://arxiv.org/html/2608.10172#Thmassumption1)and \([22](https://arxiv.org/html/2608.10172#S5.E22)\), for eachu∈𝒰u\\in\\mathcal\{U\}there exists a unique matrixAu∈ℂN×NA\_\{u\}\\in\\mathbb\{C\}^\{N\\times N\}such that

Ψ​\(F​\(x,u\)\)=Au​Ψ​\(x\)μ​\-a\.e\.​x∈𝒳\.\\Psi\\bigl\(F\(x,u\)\\bigr\)\\;=\\;A\_\{u\}\\,\\Psi\(x\)\\qquad\\mu\\text\{\-a\.e\.\}\\ x\\in\\mathcal\{X\}\.\(25\)The matrix admits the closed form

Au=𝔼μ​\[Ψ​\(F​\(x,u\)\)​Ψ​\(x\)∗\]​GΨ−1\.A\_\{u\}\\;=\\;\\mathbb\{E\}\_\{\\mu\}\\\!\\bigl\[\\Psi\\bigl\(F\(x,u\)\\bigr\)\\,\\Psi\(x\)^\{\*\}\\bigr\]\\,G\_\{\\Psi\}^\{\-1\}\.\(26\)

###### Proof\.

Fixu∈𝒰u\\in\\mathcal\{U\}\. By[Assumption˜1](https://arxiv.org/html/2608.10172#Thmassumption1), for eachi∈\{1,…,N\}i\\in\\\{1,\\dots,N\\\}the functionx↦ψi​\(F​\(x,u\)\)x\\mapsto\\psi\_\{i\}\\bigl\(F\(x,u\)\\bigr\)belongs toℋN\\mathcal\{H\}\_\{N\}\. Since\{ψj\}j=1N\\\{\\psi\_\{j\}\\\}\_\{j=1\}^\{N\}is a basis ofℋN\\mathcal\{H\}\_\{N\}under \(R1\), there is a unique tuple\(ai​1,…,ai​N\)∈ℂN\(a\_\{i1\},\\dots,a\_\{iN\}\)\\in\\mathbb\{C\}^\{N\}withψi​\(F​\(x,u\)\)=∑j=1Nai​j​ψj​\(x\)\\psi\_\{i\}\(F\(x,u\)\)=\\sum\_\{j=1\}^\{N\}a\_\{ij\}\\,\\psi\_\{j\}\(x\)μ\\mu\-a\.e\. Stacking these coefficients into theii\-th row ofAuA\_\{u\}yields \([25](https://arxiv.org/html/2608.10172#S5.E25)\)\.

To derive \([26](https://arxiv.org/html/2608.10172#S5.E26)\), right\-multiply \([25](https://arxiv.org/html/2608.10172#S5.E25)\) byΨ​\(x\)∗\\Psi\(x\)^\{\*\}and take expectation underμ\\mu:

𝔼μ​\[Ψ​\(F​\(x,u\)\)​Ψ​\(x\)∗\]=Au​𝔼μ​\[Ψ​\(x\)​Ψ​\(x\)∗\]=Au​GΨ\.\\mathbb\{E\}\_\{\\mu\}\\\!\\bigl\[\\Psi\(F\(x,u\)\)\\,\\Psi\(x\)^\{\*\}\\bigr\]\\;=\\;A\_\{u\}\\,\\mathbb\{E\}\_\{\\mu\}\\\!\\bigl\[\\Psi\(x\)\\,\\Psi\(x\)^\{\*\}\\bigr\]\\;=\\;A\_\{u\}\\,G\_\{\\Psi\}\.\(27\)SinceGΨ≻0G\_\{\\Psi\}\\succ 0by \(R1\),AuA\_\{u\}is uniquely determined as \([26](https://arxiv.org/html/2608.10172#S5.E26)\)\. ∎

###### Lemma 5\.3\(Basis\-change equivariance\)\.

Under a change of dictionaryΨ′​\(x\)=T​Ψ​\(x\)\\Psi^\{\\prime\}\(x\)=T\\,\\Psi\(x\)withT∈G​L​\(N,ℂ\)T\\in GL\(N,\\mathbb\{C\}\), the pointwise Koopman matrix transforms asAu′=T​Au​T−1A^\{\\prime\}\_\{u\}=T\\,A\_\{u\}\\,T^\{\-1\}for everyu∈𝒰u\\in\\mathcal\{U\}\. Consequently the averaged matrixA=𝔼ν​\[Au\]A=\\mathbb\{E\}\_\{\\nu\}\[A\_\{u\}\]transforms asA′=T​A​T−1A^\{\\prime\}=T\\,A\\,T^\{\-1\}\.

###### Proof\.

Under the new basis the Gram matrix isGΨ′=𝔼μ​\[T​Ψ​Ψ∗​T∗\]=T​GΨ​T∗G\_\{\\Psi^\{\\prime\}\}=\\mathbb\{E\}\_\{\\mu\}\[T\\,\\Psi\\,\\Psi^\{\*\}\\,T^\{\*\}\]=T\\,G\_\{\\Psi\}\\,T^\{\*\}\. Applying \([26](https://arxiv.org/html/2608.10172#S5.E26)\) to the new dictionary,

Au′\\displaystyle A^\{\\prime\}\_\{u\}=𝔼μ​\[T​Ψ​\(F​\(x,u\)\)​\(T​Ψ​\(x\)\)∗\]​\(T​GΨ​T∗\)−1\\displaystyle=\\mathbb\{E\}\_\{\\mu\}\\\!\\bigl\[T\\,\\Psi\(F\(x,u\)\)\\,\(T\\,\\Psi\(x\)\)^\{\*\}\\bigr\]\\,\(T\\,G\_\{\\Psi\}\\,T^\{\*\}\)^\{\-1\}=T​𝔼μ​\[Ψ​\(F​\(x,u\)\)​Ψ​\(x\)∗\]​T∗​\(T∗\)−1​GΨ−1​T−1\\displaystyle=T\\,\\mathbb\{E\}\_\{\\mu\}\\\!\\bigl\[\\Psi\(F\(x,u\)\)\\,\\Psi\(x\)^\{\*\}\\bigr\]\\,T^\{\*\}\\,\(T^\{\*\}\)^\{\-1\}\\,G\_\{\\Psi\}^\{\-1\}\\,T^\{\-1\}=T​Au​T−1\.\\displaystyle=T\\,A\_\{u\}\\,T^\{\-1\}\.Averaging overν\\nupreserves the similarity:A′=T​A​T−1A^\{\\prime\}=T\\,A\\,T^\{\-1\}\. ∎

###### Lemma 5\.4\(Spectral embedding\)\.

Under[Assumption˜1](https://arxiv.org/html/2608.10172#Thmassumption1), the finite\-dimensional subspaceℋN\\mathcal\{H\}\_\{N\}is𝒦¯\\bar\{\\mathcal\{K\}\}\-invariant, and the spectrum of the restriction obeys

σ​\(𝒦¯\|ℋN\)⊆σ​\(𝒦¯\)\.\\sigma\\\!\\bigl\(\\bar\{\\mathcal\{K\}\}\|\_\{\\mathcal\{H\}\_\{N\}\}\\bigr\)\\;\\subseteq\\;\\sigma\(\\bar\{\\mathcal\{K\}\}\)\.\(28\)Moreover the matrixAAof \([24](https://arxiv.org/html/2608.10172#S5.E24)\) is a matrix representation of𝒦¯\|ℋN\\bar\{\\mathcal\{K\}\}\|\_\{\\mathcal\{H\}\_\{N\}\}, soσ​\(A\)=σ​\(𝒦¯\|ℋN\)\\sigma\(A\)=\\sigma\(\\bar\{\\mathcal\{K\}\}\|\_\{\\mathcal\{H\}\_\{N\}\}\), whenceσ​\(A\)⊆σ​\(𝒦¯\)\\sigma\(A\)\\subseteq\\sigma\(\\bar\{\\mathcal\{K\}\}\)\.

###### Proof\.

Fixψ∈ℋN\\psi\\in\\mathcal\{H\}\_\{N\}\. By[Assumption˜1](https://arxiv.org/html/2608.10172#Thmassumption1), for eachu∈𝒰u\\in\\mathcal\{U\}we have𝒦u​ψ∈ℋN\\mathcal\{K\}\_\{u\}\\,\\psi\\in\\mathcal\{H\}\_\{N\}\. BecauseℋN\\mathcal\{H\}\_\{N\}is finite\-dimensional and hence closed, the Bochner integral𝒦¯​ψ=∫𝒦u​ψ​𝑑ν​\(u\)\\bar\{\\mathcal\{K\}\}\\,\\psi=\\int\\mathcal\{K\}\_\{u\}\\,\\psi\\,d\\nu\(u\)takes values inℋN\\mathcal\{H\}\_\{N\}\(the integrand isν\\nu\-measurable and uniformly bounded in the finite\-dimensional norm\)\. This establishes𝒦¯\\bar\{\\mathcal\{K\}\}\-invariance ofℋN\\mathcal\{H\}\_\{N\}\.

SinceℋN\\mathcal\{H\}\_\{N\}is finite\-dimensional it is complemented inℋ\\mathcal\{H\}:ℋ=ℋN⊕ℋN⟂\\mathcal\{H\}=\\mathcal\{H\}\_\{N\}\\oplus\\mathcal\{H\}\_\{N\}^\{\\perp\}\. With respect to this decomposition,𝒦¯\\bar\{\\mathcal\{K\}\}has block\-triangular form

𝒦¯=\(𝒦¯\|ℋN∗0𝒦¯\|ℋN⟂→ℋN⟂\),\\bar\{\\mathcal\{K\}\}\\;=\\;\\begin\{pmatrix\}\\bar\{\\mathcal\{K\}\}\|\_\{\\mathcal\{H\}\_\{N\}\}&\*\\\\ 0&\\bar\{\\mathcal\{K\}\}\|\_\{\\mathcal\{H\}\_\{N\}^\{\\perp\}\\to\\mathcal\{H\}\_\{N\}^\{\\perp\}\}\\end\{pmatrix\},\(29\)because the\(2,1\)\(2,1\)block vanishes by invariance\. A standard result in operator theory then givesσ​\(𝒦¯\)=σ​\(𝒦¯\|ℋN\)∪σ​\(𝒦¯\|ℋN⟂→ℋN⟂\)\\sigma\(\\bar\{\\mathcal\{K\}\}\)=\\sigma\(\\bar\{\\mathcal\{K\}\}\|\_\{\\mathcal\{H\}\_\{N\}\}\)\\cup\\sigma\(\\bar\{\\mathcal\{K\}\}\|\_\{\\mathcal\{H\}\_\{N\}^\{\\perp\}\\to\\mathcal\{H\}\_\{N\}^\{\\perp\}\}\)\[[38](https://arxiv.org/html/2608.10172#bib.bib72), Ch\. III, §4\], soσ​\(𝒦¯\|ℋN\)⊆σ​\(𝒦¯\)\\sigma\(\\bar\{\\mathcal\{K\}\}\|\_\{\\mathcal\{H\}\_\{N\}\}\)\\subseteq\\sigma\(\\bar\{\\mathcal\{K\}\}\), which is \([28](https://arxiv.org/html/2608.10172#S5.E28)\)\.

Finally, expressing𝒦¯\|ℋN\\bar\{\\mathcal\{K\}\}\|\_\{\\mathcal\{H\}\_\{N\}\}in the basis\{ψi\}i=1N\\\{\\psi\_\{i\}\\\}\_\{i=1\}^\{N\}produces a matrix whose action satisfies \([24](https://arxiv.org/html/2608.10172#S5.E24)\), and this matrix is preciselyAA\. Since spectra of finite\-dimensional operators equal spectra of their matrix representations,σ​\(A\)=σ​\(𝒦¯\|ℋN\)\\sigma\(A\)=\\sigma\(\\bar\{\\mathcal\{K\}\}\|\_\{\\mathcal\{H\}\_\{N\}\}\)\. ∎

### 5\.3Proof of[Theorem˜5\.1](https://arxiv.org/html/2608.10172#S5.Thmtheorem1)

###### Proof of[Theorem˜5\.1](https://arxiv.org/html/2608.10172#S5.Thmtheorem1)\.

Existence ofAAsatisfying \([24](https://arxiv.org/html/2608.10172#S5.E24)\) follows from[Lemma˜5\.2](https://arxiv.org/html/2608.10172#S5.Thmtheorem2)by averaging overu∼νu\\sim\\nu: forμ\\mu\-a\.e\.x∈𝒳x\\in\\mathcal\{X\},

𝔼u∼ν​\[Ψ​\(F​\(x,u\)\)\]=𝔼ν​\[Au​Ψ​\(x\)\]=𝔼ν​\[Au\]​Ψ​\(x\)=A​Ψ​\(x\),\\mathbb\{E\}\_\{u\\sim\\nu\}\\\!\\bigl\[\\Psi\(F\(x,u\)\)\\bigr\]\\;=\\;\\mathbb\{E\}\_\{\\nu\}\[A\_\{u\}\\,\\Psi\(x\)\]\\;=\\;\\mathbb\{E\}\_\{\\nu\}\[A\_\{u\}\]\\,\\Psi\(x\)\\;=\\;A\\,\\Psi\(x\),\(30\)whereA≔𝔼ν​\[Au\]A\\coloneqq\\mathbb\{E\}\_\{\\nu\}\[A\_\{u\}\]\. The average is well defined inℂN×N\\mathbb\{C\}^\{N\\times N\}because the family\{Au\}u∈𝒰\\\{A\_\{u\}\\\}\_\{u\\in\\mathcal\{U\}\}isν\\nu\-measurable by the closed form \([26](https://arxiv.org/html/2608.10172#S5.E26)\) and bounded on compact subsets of𝒰\\mathcal\{U\}; the Bochner integral therefore converges\. Uniqueness ofAAfollows by right\-multiplying \([24](https://arxiv.org/html/2608.10172#S5.E24)\) byΨ​\(x\)∗\\Psi\(x\)^\{\*\}, taking expectation underμ\\mu, and applyingGΨ≻0G\_\{\\Psi\}\\succ 0\. This proves \(a\)\.

Part \(b\) is[Lemma˜5\.3](https://arxiv.org/html/2608.10172#S5.Thmtheorem3)\. Part \(c\) follows becauseσ​\(T​A​T−1\)=σ​\(A\)\\sigma\(TAT^\{\-1\}\)=\\sigma\(A\)for anyT∈G​L​\(N,ℂ\)T\\in GL\(N,\\mathbb\{C\}\), so the spectrum depends only on the similarity class ofAAand not on the choice of basis forℋN\\mathcal\{H\}\_\{N\}\. Part \(d\) is[Lemma˜5\.4](https://arxiv.org/html/2608.10172#S5.Thmtheorem4)\.

It remains to prove \(e\)\. The population least\-squares realisation\(ALS,BLS\)\(A\_\{\\mathrm\{LS\}\},B\_\{\\mathrm\{LS\}\}\)of[Definition˜4\.6](https://arxiv.org/html/2608.10172#S4.Thmtheorem6)solves

minA′,B′⁡𝔼\(x,u\)∼μ⊗ν​‖Ψ​\(F​\(x,u\)\)−A′​Ψ​\(x\)−B′​u‖22\.\\min\_\{A^\{\\prime\},B^\{\\prime\}\}\\;\\mathbb\{E\}\_\{\(x,u\)\\sim\\mu\\otimes\\nu\}\\bigl\\\|\\Psi\(F\(x,u\)\)\-A^\{\\prime\}\\,\\Psi\(x\)\-B^\{\\prime\}\\,u\\bigr\\\|\_\{2\}^\{2\}\.\(31\)The first\-order optimality conditions of this quadratic problem are

ALS​GΨ\+BLS​𝔼​\[u​Ψ​\(x\)∗\]\\displaystyle A\_\{\\mathrm\{LS\}\}\\,G\_\{\\Psi\}\+B\_\{\\mathrm\{LS\}\}\\,\\mathbb\{E\}\[u\\,\\Psi\(x\)^\{\*\}\]=𝔼​\[Ψ​\(F​\(x,u\)\)​Ψ​\(x\)∗\],\\displaystyle=\\mathbb\{E\}\\\!\\bigl\[\\Psi\(F\(x,u\)\)\\,\\Psi\(x\)^\{\*\}\\bigr\],ALS​𝔼​\[Ψ​\(x\)​u∗\]\+BLS​Gu\\displaystyle A\_\{\\mathrm\{LS\}\}\\,\\mathbb\{E\}\[\\Psi\(x\)\\,u^\{\*\}\]\+B\_\{\\mathrm\{LS\}\}\\,G\_\{u\}=𝔼​\[Ψ​\(F​\(x,u\)\)​u∗\],\\displaystyle=\\mathbb\{E\}\\\!\\bigl\[\\Psi\(F\(x,u\)\)\\,u^\{\*\}\\bigr\],whereGu≔𝔼ν​\[u​u∗\]G\_\{u\}\\coloneqq\\mathbb\{E\}\_\{\\nu\}\[uu^\{\*\}\]\. By \(R2\) and the centring𝔼ν​\[u\]=0\\mathbb\{E\}\_\{\\nu\}\[u\]=0, the cross\-terms𝔼​\[u​Ψ​\(x\)∗\]\\mathbb\{E\}\[u\\,\\Psi\(x\)^\{\*\}\]and𝔼​\[Ψ​\(x\)​u∗\]\\mathbb\{E\}\[\\Psi\(x\)\\,u^\{\*\}\]vanish, so the equations decouple\. Solving the first forALSA\_\{\\mathrm\{LS\}\}and using[Lemma˜5\.2](https://arxiv.org/html/2608.10172#S5.Thmtheorem2),

ALS\\displaystyle A\_\{\\mathrm\{LS\}\}=𝔼​\[Ψ​\(F​\(x,u\)\)​Ψ​\(x\)∗\]​GΨ−1\\displaystyle=\\mathbb\{E\}\\\!\\bigl\[\\Psi\(F\(x,u\)\)\\,\\Psi\(x\)^\{\*\}\\bigr\]\\,G\_\{\\Psi\}^\{\-1\}=𝔼ν​\[𝔼μ​\[Au​Ψ​\(x\)​Ψ​\(x\)∗\]\]​GΨ−1\\displaystyle=\\mathbb\{E\}\_\{\\nu\}\\\!\\bigl\[\\,\\mathbb\{E\}\_\{\\mu\}\[A\_\{u\}\\,\\Psi\(x\)\\,\\Psi\(x\)^\{\*\}\]\\,\\bigr\]\\,G\_\{\\Psi\}^\{\-1\}=𝔼ν​\[Au​GΨ\]​GΨ−1=𝔼ν​\[Au\]=A\.\\displaystyle=\\mathbb\{E\}\_\{\\nu\}\[A\_\{u\}\\,G\_\{\\Psi\}\]\\,G\_\{\\Psi\}^\{\-1\}\\;=\\;\\mathbb\{E\}\_\{\\nu\}\[A\_\{u\}\]\\;=\\;A\.This provesALS=AA\_\{\\mathrm\{LS\}\}=Aand completes the proof\. ∎

### 5\.4Coordinate\-free invariants of Koopman circuits

[Theorem˜5\.1](https://arxiv.org/html/2608.10172#S5.Thmtheorem1)implies that every spectral attribute ofAApreserved under similarity is a well\-defined invariant of the pair\(F,ℋN\)\(F,\\mathcal\{H\}\_\{N\}\)\.

###### Corollary 5\.5\(Coordinate\-free invariants of the Koopman realisation\)\.

Under the hypotheses of[Theorem˜5\.1](https://arxiv.org/html/2608.10172#S5.Thmtheorem1), the following are coordinate\-free invariants of\(F,ℋN\)\(F,\\mathcal\{H\}\_\{N\}\):

1. \(i\)the multiset of eigenvalues\{λk\}k=1N\\\{\\lambda\_\{k\}\\\}\_\{k=1\}^\{N\}ofAA, with algebraic multiplicities;
2. \(ii\)for each distinct eigenvalueλ\\lambda, the spectral projectorΠλ\\Pi\_\{\\lambda\}onto the corresponding generalised eigenspace, in the sense that under basis changeTTit transforms asΠλ′=T​Πλ​T−1\\Pi^\{\\prime\}\_\{\\lambda\}=T\\,\\Pi\_\{\\lambda\}\\,T^\{\-1\};
3. \(iii\)the number and sizes of Jordan blocks associated with each eigenvalue;
4. \(iv\)the algebra𝒜​\(A\)⊆ℂN×N\\mathcal\{A\}\(A\)\\subseteq\\mathbb\{C\}^\{N\\times N\}generated by the spectral projectors under composition and complex linear combination, considered as an abstract commutative∗\*\-algebra\.

###### Proof\.

UnderA↦T​A​T−1A\\mapsto TAT^\{\-1\}: eigenvalues are preserved,σ​\(T​A​T−1\)=σ​\(A\)\\sigma\(TAT^\{\-1\}\)=\\sigma\(A\)\. Riesz projectorsΠλ=12​π​i​∮Γλ\(z​I−A\)−1​𝑑z\\Pi\_\{\\lambda\}=\\frac\{1\}\{2\\pi i\}\\oint\_\{\\Gamma\_\{\\lambda\}\}\(zI\-A\)^\{\-1\}\\,dzaround a small contourΓλ\\Gamma\_\{\\lambda\}enclosingλ\\lambdatransform asT​Πλ​T−1T\\Pi\_\{\\lambda\}T^\{\-1\}, because\(z​I−T​A​T−1\)−1=T​\(z​I−A\)−1​T−1\(zI\-TAT^\{\-1\}\)^\{\-1\}=T\(zI\-A\)^\{\-1\}T^\{\-1\}\. Jordan block counts are invariants of similarity classes\. Finally, the algebra generated by\{T​Πλ​T−1\}\\\{T\\Pi\_\{\\lambda\}T^\{\-1\}\\\}is isomorphic to that generated by\{Πλ\}\\\{\\Pi\_\{\\lambda\}\\\}viaM↦T​M​T−1M\\mapsto TMT^\{\-1\}, which preserves the∗\*\-algebra structure\. ∎

[Corollary˜5\.5](https://arxiv.org/html/2608.10172#S5.Thmtheorem5)formalises the coordinate\-free content of a Koopman circuit \([Definition˜4\.10](https://arxiv.org/html/2608.10172#S4.Thmtheorem10)\)\. A circuit is not a tuple of numbers in a chosen basis: it is an equivalence class of such tuples under the similarity action ofG​L​\(N,ℂ\)GL\(N,\\mathbb\{C\}\), equivalently an object of the abstract algebra𝒜​\(A\)\\mathcal\{A\}\(A\)\.[Theorem˜9\.2](https://arxiv.org/html/2608.10172#S9.Thmtheorem2)shows that this algebra has the structure𝒜​\(A\)≅ℂr∗\\mathcal\{A\}\(A\)\\cong\\mathbb\{C\}^\{r^\{\*\}\}, wherer∗r^\{\*\}is the number of distinct eigenvalues, and that every first\-order intervention on the transformer’s depth dynamics has a unique representative in it\.

### 5\.5Remarks

[Theorem˜5\.1](https://arxiv.org/html/2608.10172#S5.Thmtheorem1)establishes*what*is to be identified\. The next section proves that it is identifiable at the parametric rate from finite calibration data\.

## 6Transformers Are Spectrally Identifiable

This section contains the paper’s main result and its complete proof\. Informally: record the residual\-stream state, the attention write, and the next\-layer state over a calibration corpus, and fit the best linear map from lifted states plus controls to lifted next\-states\. The eigenvalues of the fitted map converge to those of the true Koopman realisation at the optimalM−1/2M^\{\-1/2\}rate, with permutation as the only remaining ambiguity\. Different datasets, seeds, or dictionary bases therefore recover the same spectrum up to a known error\.

### 6\.1Regularity conditions and statement

We work under[Assumptions˜1](https://arxiv.org/html/2608.10172#Thmassumption1),[2](https://arxiv.org/html/2608.10172#Thmassumption2)and[3](https://arxiv.org/html/2608.10172#Thmassumption3), the regularity conditions \([22](https://arxiv.org/html/2608.10172#S5.E22)\)–\([23](https://arxiv.org/html/2608.10172#S5.E23)\) of[Section˜5\.1](https://arxiv.org/html/2608.10172#S5.SS1), and two additional finite\-sample conditions on tails and non\-normality:

\(R3\):Ψ​\(x\)​isL\-sub\-Gaussian underμanduisL\-sub\-Gaussian\\displaystyle\\Psi\(x\)\\text\{ is $L$\-sub\-Gaussian under $\\mu$ and $u$ is $L$\-sub\-Gaussian\}\(33\)underν\\nu\.\(R4\):A​is diagonalisable with eigenvector matrixVsatisfying\\displaystyle A\\text\{ is diagonalisable with eigenvector matrix $V$ satisfying\}\(34\)κ2​\(V\)≤κ0\.\\displaystyle\\qquad\\kappa\_\{2\}\(V\)\\leq\\kappa\_\{0\}\.
where theLL\-sub\-Gaussian norm and theℓ2\\ell\_\{2\}condition numberκ2​\(V\)=‖V‖2​‖V−1‖2\\kappa\_\{2\}\(V\)=\\\|V\\\|\_\{2\}\\\|V^\{\-1\}\\\|\_\{2\}are as in\[[72](https://arxiv.org/html/2608.10172#bib.bib73),[65](https://arxiv.org/html/2608.10172#bib.bib77)\]\. Recall thatλmin​\(GΨ\)≥ηΨ\>0\\lambda\_\{\\min\}\(G\_\{\\Psi\}\)\\geq\\eta\_\{\\Psi\}\>0under \([22](https://arxiv.org/html/2608.10172#S5.E22)\), andλmin​\(Gu\)≥η\\lambda\_\{\\min\}\(G\_\{u\}\)\\geq\\etaunder[Assumption˜2](https://arxiv.org/html/2608.10172#Thmassumption2)\.

###### Theorem 6\.1\(Spectral identifiability of the Koopman realisation\)\.

Assume[Assumptions˜1](https://arxiv.org/html/2608.10172#Thmassumption1),[2](https://arxiv.org/html/2608.10172#Thmassumption2)and[3](https://arxiv.org/html/2608.10172#Thmassumption3), the regularity conditions \([22](https://arxiv.org/html/2608.10172#S5.E22)\)–\([23](https://arxiv.org/html/2608.10172#S5.E23)\) and \([33](https://arxiv.org/html/2608.10172#S6.E33)\)–\([34](https://arxiv.org/html/2608.10172#S6.E34)\), and centre the controls \(𝔼ν​\[u\]=0\\mathbb\{E\}\_\{\\nu\}\[u\]=0\)\. LetA^M\\hat\{A\}\_\{M\}be the EDMDc estimator of[Definition˜4\.8](https://arxiv.org/html/2608.10172#S4.Thmtheorem8)with regularisationγ=0\\gamma=0, computed fromMMlayer–token calibration samples, and normalise eigenvectors to unitℓ2\\ell\_\{2\}\-norm\. Define the sample\-size threshold

M0≔c0​κ02​L4​\(1\+‖A‖2\)2ηΨ2​Δ2​\(N\+log⁡\(1/δ\)\),M\_\{0\}\\;\\coloneqq\\;\\frac\{c\_\{0\}\\,\\kappa\_\{0\}^\{2\}\\,L^\{4\}\\,\(1\+\\\|A\\\|\_\{2\}\)^\{2\}\}\{\\eta\_\{\\Psi\}^\{2\}\\,\\Delta^\{2\}\}\\,\\bigl\(N\+\\log\(1/\\delta\)\\bigr\),\(35\)withc0c\_\{0\}a universal constant\. For everyδ∈\(0,1/2\)\\delta\\in\(0,1/2\)and everyM≥M0M\\geq M\_\{0\}there exists a permutationπM\\pi\_\{M\}of\{1,…,N\}\\\{1,\\dots,N\\\}, determined byA^M\\hat\{A\}\_\{M\}, such that with probability at least1−δ1\-\\delta:

1. \(i\)*Eigenvalue identifiability:* maxk⁡\|λk​\(A^M\)−λπM​\(k\)​\(A\)\|≤c1​κ0​L2​\(1\+‖A‖2\)ηΨ×N\+log⁡\(1/δ\)M\.\\begin\{split\}\\max\_\{k\}\\,\\bigl\|\\lambda\_\{k\}\(\\hat\{A\}\_\{M\}\)\-\\lambda\_\{\\pi\_\{M\}\(k\)\}\(A\)\\bigr\|&\\leq\\frac\{c\_\{1\}\\,\\kappa\_\{0\}\\,L^\{2\}\\,\(1\+\\\|A\\\|\_\{2\}\)\}\{\\eta\_\{\\Psi\}\}\\\\ &\\quad\\times\\sqrt\{\\frac\{N\+\\log\(1/\\delta\)\}\{M\}\}\.\\end\{split\}\(36\)
2. \(ii\)*Eigenvector identifiability:*With appropriate phase choice, and denoting byv^k=vk​\(A^M\)\\hat\{v\}\_\{k\}=v\_\{k\}\(\\hat\{A\}\_\{M\}\)andvπM​\(k\)=vπM​\(k\)​\(A\)v\_\{\\pi\_\{M\}\(k\)\}=v\_\{\\pi\_\{M\}\(k\)\}\(A\), maxk⁡‖v^k−vπM​\(k\)‖2≤c2​κ02​L2​\(1\+‖A‖2\)ηΨ​Δ×N\+log⁡\(1/δ\)M\.\\begin\{split\}\\max\_\{k\}\\,\\bigl\\\|\\hat\{v\}\_\{k\}\-v\_\{\\pi\_\{M\}\(k\)\}\\bigr\\\|\_\{2\}&\\leq\\frac\{c\_\{2\}\\,\\kappa\_\{0\}^\{2\}\\,L^\{2\}\\,\(1\+\\\|A\\\|\_\{2\}\)\}\{\\eta\_\{\\Psi\}\\,\\Delta\}\\\\ &\\quad\\times\\sqrt\{\\frac\{N\+\\log\(1/\\delta\)\}\{M\}\}\.\\end\{split\}\(37\)

The constantsc1,c2\>0c\_\{1\},c\_\{2\}\>0are universal \(independent of the problem parameters\)\.

The permutationπM\\pi\_\{M\}depends on the sample, reflecting the fact that unlabelled eigenvalues have no intrinsic ordering; the theorem guarantees that the*multiset*of empirical eigenvalues approaches the true multiset, and that the empirical eigenvectors align with the true ones under the induced matching\. Both rates areM−1/2M^\{\-1/2\}: asymptotically the spectrum is recovered as fast as any parametric quantity from i\.i\.d\. data\.

The constants are interpretable\. The factorκ0\\kappa\_\{0\}measures non\-normality of the realisation and enters the eigenvalue bound linearly and the eigenvector bound quadratically;ηΨ\\eta\_\{\\Psi\}measures how well the dictionary spansℋN\\mathcal\{H\}\_\{N\}and appears in the denominator, which is why every dictionary in our experiments is whitened \(unwhitened we measureηΨ≈7×10−3\\eta\_\{\\Psi\}\\approx 7\\times 10^\{\-3\}, whitened0\.680\.68–0\.840\.84; see[Section˜11\.5](https://arxiv.org/html/2608.10172#S11.SS5)\);Δ\\Deltaenters only the eigenvector bound and the threshold; and the thresholdM0M\_\{0\}scales linearly inNN\.

### 6\.2Proof strategy

The proof combines three lines of classical machinery\.

1. 1\.*Matrix concentration\.*Under sub\-Gaussian tails on the dictionary and controls \- a mild moment condition on the calibration corpus \- the empirical Gramians concentrate around their population counterparts at rateM−1/2M^\{\-1/2\}, with an explicit constant scaling asN\\sqrt\{N\}in the dictionary dimension\[[72](https://arxiv.org/html/2608.10172#bib.bib73),[69](https://arxiv.org/html/2608.10172#bib.bib74)\]\.
2. 2\.*Least\-squares stability\.*The EDMDc estimator is a smooth matrix\-valued function of the empirical Gramians on the event that the state Gramian is well conditioned \- an event guaranteed by[Assumption˜2](https://arxiv.org/html/2608.10172#Thmassumption2), \([22](https://arxiv.org/html/2608.10172#S5.E22)\), and the concentration of step 1\. Perturbation of the least\-squares solution inherits theM−1/2M^\{\-1/2\}rate\.
3. 3\.*Spectral perturbation\.*Under[Assumption˜3](https://arxiv.org/html/2608.10172#Thmassumption3), an operator\-norm bound onA^M−A\\hat\{A\}\_\{M\}\-Atranslates via Bauer–Fike\[[5](https://arxiv.org/html/2608.10172#bib.bib75)\]into a permutation\-matching eigenvalue bound, and via Davis–Kahan\[[14](https://arxiv.org/html/2608.10172#bib.bib76),[65](https://arxiv.org/html/2608.10172#bib.bib77)\]into an eigenvector bound with rate scaled by1/Δ1/\\Delta\.

The three ingredients are individually classical; the technical content is the assembly and constant\-tracking for the transformer Koopman setting, and in particular the identification of the correct regularity conditions on the calibration distribution and the dictionary that make the finite\-sample rate meaningful\.

Throughout we writez≔\(Ψ​\(x\)⊤,u⊤\)⊤∈ℂN\+pz\\coloneqq\\bigl\(\\Psi\(x\)^\{\\top\},u^\{\\top\}\\bigr\)^\{\\top\}\\in\\mathbb\{C\}^\{N\+p\}for the stacked design vector,GZ≔𝔼​\[z​z∗\]G\_\{Z\}\\coloneqq\\mathbb\{E\}\[zz^\{\*\}\]for its population Gramian andG^Z≔1M​∑izi​zi∗\\hat\{G\}\_\{Z\}\\coloneqq\\frac\{1\}\{M\}\\sum\_\{i\}z\_\{i\}z\_\{i\}^\{\*\}for its empirical counterpart\. Under \(R2\) and centring,GZG\_\{Z\}is block diagonal with blocksGΨG\_\{\\Psi\}andGuG\_\{u\}\.

### 6\.3Step 1: concentration of the empirical Gramians

Define the empirical state Gramian and cross\-Gramian

G^M≔1M​X​X∗∈ℂN×N,C^M≔1M​Y​X∗∈ℂN×N,\\hat\{G\}\_\{M\}\\;\\coloneqq\\;\\tfrac\{1\}\{M\}\\,XX^\{\*\}\\in\\mathbb\{C\}^\{N\\times N\},\\qquad\\hat\{C\}\_\{M\}\\;\\coloneqq\\;\\tfrac\{1\}\{M\}\\,YX^\{\*\}\\in\\mathbb\{C\}^\{N\\times N\},\(38\)and their population counterpartsGΨ=𝔼μ​\[Ψ​Ψ∗\]G\_\{\\Psi\}=\\mathbb\{E\}\_\{\\mu\}\[\\Psi\\Psi^\{\*\}\]andCΨ=𝔼\(x,u\)​\[Ψ​\(F​\(x,u\)\)​Ψ​\(x\)∗\]C\_\{\\Psi\}=\\mathbb\{E\}\_\{\(x,u\)\}\[\\Psi\(F\(x,u\)\)\\,\\Psi\(x\)^\{\*\}\]\. By[Theorem˜5\.1](https://arxiv.org/html/2608.10172#S5.Thmtheorem1)\(e\) and[Lemma˜5\.2](https://arxiv.org/html/2608.10172#S5.Thmtheorem2),A=CΨ​GΨ−1A=C\_\{\\Psi\}G\_\{\\Psi\}^\{\-1\}\.

###### Lemma 6\.3\(Sub\-Gaussian Gramian concentration\)\.

Under \([22](https://arxiv.org/html/2608.10172#S5.E22)\) and \([33](https://arxiv.org/html/2608.10172#S6.E33)\) there is a universal constantc3\>0c\_\{3\}\>0such that for everyδ∈\(0,1/2\)\\delta\\in\(0,1/2\)and everyMM,

‖G^M−GΨ‖2≤c3​L2​\(N\+log⁡\(1/δ\)M\+N\+log⁡\(1/δ\)M\)\\bigl\\\|\\hat\{G\}\_\{M\}\-G\_\{\\Psi\}\\bigr\\\|\_\{2\}\\;\\leq\\;c\_\{3\}\\,L^\{2\}\\,\\Biggl\(\\sqrt\{\\frac\{N\+\\log\(1/\\delta\)\}\{M\}\}\+\\frac\{N\+\\log\(1/\\delta\)\}\{M\}\\Biggr\)\(39\)with probability at least1−δ1\-\\delta\. The same bound holds for‖C^M−CΨ‖2\\\|\\hat\{C\}\_\{M\}\-C\_\{\\Psi\}\\\|\_\{2\}, for‖Σ^M−Gu‖2\\\|\\hat\{\\Sigma\}\_\{M\}\-G\_\{u\}\\\|\_\{2\}withΣ^M≔Ξ​Ξ∗/M\\hat\{\\Sigma\}\_\{M\}\\coloneqq\\Xi\\Xi^\{\*\}/M, and for‖G^Z−GZ‖2\\\|\\hat\{G\}\_\{Z\}\-G\_\{Z\}\\\|\_\{2\}withNNreplaced byN\+pN\+p\.

###### Proof\.

The bound forG^M\\hat\{G\}\_\{M\}is the sub\-Gaussian sample\-covariance concentration inequality of\[[72](https://arxiv.org/html/2608.10172#bib.bib73), Theorem 4\.6\.1\], applied to the i\.i\.d\. observationsΨ​\(xi\)∈ℂN\\Psi\(x\_\{i\}\)\\in\\mathbb\{C\}^\{N\}with sub\-Gaussian norm‖Ψ‖ψ2≤L\\\|\\Psi\\\|\_\{\\psi\_\{2\}\}\\leq L\. ForC^M\\hat\{C\}\_\{M\}, apply the same theorem to the cross\-outer\-productΨ​\(F​\(x,u\)\)​Ψ​\(x\)∗\\Psi\(F\(x,u\)\)\\,\\Psi\(x\)^\{\*\}; this is a sub\-exponential random matrix by \(R3\), and matrix Bernstein\[[69](https://arxiv.org/html/2608.10172#bib.bib74), Theorem 6\.1\.1\]yields \([39](https://arxiv.org/html/2608.10172#S6.E39)\) with the same rate\. ForΣ^M\\hat\{\\Sigma\}\_\{M\}andG^Z\\hat\{G\}\_\{Z\}, apply the theorem again touiu\_\{i\}and to the stackedziz\_\{i\}respectively\. ∎

ForM≥N\+log⁡\(1/δ\)M\\geq N\+\\log\(1/\\delta\)the leading term in \([39](https://arxiv.org/html/2608.10172#S6.E39)\) dominates; we absorb the second\-order term intoc3c\_\{3\}and use henceforth

‖G^M−GΨ‖2,‖C^M−CΨ‖2,‖Σ^M−Gu‖2≤c3​L2​N\+log⁡\(1/δ\)M\.\\bigl\\\|\\hat\{G\}\_\{M\}\-G\_\{\\Psi\}\\bigr\\\|\_\{2\},\\;\\bigl\\\|\\hat\{C\}\_\{M\}\-C\_\{\\Psi\}\\bigr\\\|\_\{2\},\\;\\bigl\\\|\\hat\{\\Sigma\}\_\{M\}\-G\_\{u\}\\bigr\\\|\_\{2\}\\;\\leq\\;c\_\{3\}\\,L^\{2\}\\,\\sqrt\{\\tfrac\{N\+\\log\(1/\\delta\)\}\{M\}\}\.\(40\)

### 6\.4Step 2: operator\-norm bound on the EDMDc estimator

By \(R2\), centring, and[Theorem˜5\.1](https://arxiv.org/html/2608.10172#S5.Thmtheorem1)\(e\), the population EDMDc problem decouples into state and control blocks:A=CΨ​GΨ−1A=C\_\{\\Psi\}G\_\{\\Psi\}^\{\-1\}andB=CΨ,u​Gu−1B=C\_\{\\Psi,u\}G\_\{u\}^\{\-1\}whereCΨ,u=𝔼​\[Ψ​\(F\)​u∗\]C\_\{\\Psi,u\}=\\mathbb\{E\}\[\\Psi\(F\)\\,u^\{\*\}\]\. For finiteMMthe empirical cross\-termsX​Ξ∗/MX\\Xi^\{\*\}/MandΞ​X∗/M\\Xi X^\{\*\}/Mconcentrate around0at rateM−1/2M^\{\-1/2\}by[Lemma˜6\.3](https://arxiv.org/html/2608.10172#S6.Thmtheorem3), so the joint EDMDc estimator satisfies

A^M=C^M​G^M−1\+RM,‖RM‖2=Op​\(M−1/2\),\\hat\{A\}\_\{M\}\\;=\\;\\hat\{C\}\_\{M\}\\,\\hat\{G\}\_\{M\}^\{\-1\}\\;\+\\;R\_\{M\},\\qquad\\\|R\_\{M\}\\\|\_\{2\}\\;=\\;O\_\{p\}\\\!\\bigl\(M^\{\-1/2\}\\bigr\),\(41\)with the remainderRMR\_\{M\}absorbing the effect of the empirical cross\-terms\. We prove the following for the leading term; the same rate holds forA^M\\hat\{A\}\_\{M\}itself by absorbingRMR\_\{M\}into the constant\.

###### Lemma 6\.4\(Operator\-norm rate for the EDMDc estimator\)\.

Under the hypotheses of[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)there is a constantc4c\_\{4\}, depending only on universal constants, such that for everyδ∈\(0,1/2\)\\delta\\in\(0,1/2\)and everyMMsatisfying

M≥4​c32​L4ηΨ2​\(N\+log⁡\(1/δ\)\),M\\;\\geq\\;\\frac\{4\\,c\_\{3\}^\{2\}\\,L^\{4\}\}\{\\eta\_\{\\Psi\}^\{2\}\}\\,\\bigl\(N\+\\log\(1/\\delta\)\\bigr\),\(42\)with probability at least1−δ1\-\\delta,

‖A^M−A‖2≤c4​L2​\(1\+‖A‖2\)ηΨ​N\+log⁡\(1/δ\)M\.\\bigl\\\|\\hat\{A\}\_\{M\}\-A\\bigr\\\|\_\{2\}\\;\\leq\\;\\frac\{c\_\{4\}\\,L^\{2\}\\,\(1\+\\\|A\\\|\_\{2\}\)\}\{\\eta\_\{\\Psi\}\}\\,\\sqrt\{\\tfrac\{N\+\\log\(1/\\delta\)\}\{M\}\}\.\(43\)

###### Proof\.

WriteC^M=CΨ\+ΔC\\hat\{C\}\_\{M\}=C\_\{\\Psi\}\+\\Delta\_\{C\}andG^M=GΨ\+ΔG\\hat\{G\}\_\{M\}=G\_\{\\Psi\}\+\\Delta\_\{G\}, with‖ΔC‖2,‖ΔG‖2\\\|\\Delta\_\{C\}\\\|\_\{2\},\\\|\\Delta\_\{G\}\\\|\_\{2\}bounded by \([40](https://arxiv.org/html/2608.10172#S6.E40)\) on an eventEEof probability at least1−δ/21\-\\delta/2\. OnEE, condition \([42](https://arxiv.org/html/2608.10172#S6.E42)\) gives‖ΔG‖2≤ηΨ/2\\\|\\Delta\_\{G\}\\\|\_\{2\}\\leq\\eta\_\{\\Psi\}/2, whence Weyl’s inequality impliesλmin​\(G^M\)≥ηΨ/2\\lambda\_\{\\min\}\(\\hat\{G\}\_\{M\}\)\\geq\\eta\_\{\\Psi\}/2and thus‖G^M−1‖2≤2/ηΨ\\\|\\hat\{G\}\_\{M\}^\{\-1\}\\\|\_\{2\}\\leq 2/\\eta\_\{\\Psi\}\.

From the resolvent identityG^M−1−GΨ−1=−GΨ−1​\(G^M−GΨ\)​G^M−1\\hat\{G\}\_\{M\}^\{\-1\}\-G\_\{\\Psi\}^\{\-1\}=\-G\_\{\\Psi\}^\{\-1\}\(\\hat\{G\}\_\{M\}\-G\_\{\\Psi\}\)\\hat\{G\}\_\{M\}^\{\-1\}we obtain

C^M​G^M−1−A=\(C^M−CΨ\)​G^M−1−CΨ​GΨ−1​\(G^M−GΨ\)​G^M−1,\\hat\{C\}\_\{M\}\\,\\hat\{G\}\_\{M\}^\{\-1\}\-A\\;=\\;\(\\hat\{C\}\_\{M\}\-C\_\{\\Psi\}\)\\,\\hat\{G\}\_\{M\}^\{\-1\}\\;\-\\;C\_\{\\Psi\}\\,G\_\{\\Psi\}^\{\-1\}\(\\hat\{G\}\_\{M\}\-G\_\{\\Psi\}\)\\,\\hat\{G\}\_\{M\}^\{\-1\},\(44\)and sinceCΨ​GΨ−1=AC\_\{\\Psi\}\\,G\_\{\\Psi\}^\{\-1\}=A,

C^M​G^M−1−A=ΔC​G^M−1−A​ΔG​G^M−1\.\\hat\{C\}\_\{M\}\\,\\hat\{G\}\_\{M\}^\{\-1\}\-A\\;=\\;\\Delta\_\{C\}\\,\\hat\{G\}\_\{M\}^\{\-1\}\\;\-\\;A\\,\\Delta\_\{G\}\\,\\hat\{G\}\_\{M\}^\{\-1\}\.\(45\)Taking operator norms and using‖G^M−1‖2≤2/ηΨ\\\|\\hat\{G\}\_\{M\}^\{\-1\}\\\|\_\{2\}\\leq 2/\\eta\_\{\\Psi\},

‖C^M​G^M−1−A‖2\\displaystyle\\bigl\\\|\\hat\{C\}\_\{M\}\\hat\{G\}\_\{M\}^\{\-1\}\-A\\bigr\\\|\_\{2\}≤2ηΨ​\(‖ΔC‖2\+‖A‖2​‖ΔG‖2\)\\displaystyle\\leq\\frac\{2\}\{\\eta\_\{\\Psi\}\}\\bigl\(\\\|\\Delta\_\{C\}\\\|\_\{2\}\+\\\|A\\\|\_\{2\}\\\|\\Delta\_\{G\}\\\|\_\{2\}\\bigr\)\(46\)≤2​c3​L2ηΨ​\(1\+‖A‖2\)​N\+log⁡\(2/δ\)M\.\\displaystyle\\leq\\frac\{2c\_\{3\}L^\{2\}\}\{\\eta\_\{\\Psi\}\}\(1\+\\\|A\\\|\_\{2\}\)\\sqrt\{\\frac\{N\+\\log\(2/\\delta\)\}\{M\}\}\.
Absorbing the remainderRMR\_\{M\}of \([41](https://arxiv.org/html/2608.10172#S6.E41)\), also bounded by the same rate on an event of probability at least1−δ/21\-\\delta/2, via a union bound yields \([43](https://arxiv.org/html/2608.10172#S6.E43)\) withc4=4​c3c\_\{4\}=4c\_\{3\}\. ∎

[Lemma˜6\.4](https://arxiv.org/html/2608.10172#S6.Thmtheorem4)is the workhorse: it converts theM−1/2M^\{\-1/2\}Gramian concentration of Step 1 into the same rate on the estimator itself\. The dependence onηΨ\\eta\_\{\\Psi\}and‖A‖2\\\|A\\\|\_\{2\}reflects, respectively, how well the dictionary spansℋN\\mathcal\{H\}\_\{N\}and how strongly the true Koopman flow acts within it\.

### 6\.5Step 3: from operator norm to spectral perturbation

We now translate the operator\-norm bound into eigenvalue and eigenvector rates\. Two classical results suffice: Bauer–Fike for eigenvalues and Davis–Kahan, in a form suitable for non\-normal matrices, for eigenvectors\.

###### Lemma 6\.5\(Eigenvalue perturbation with matching\)\.

LetA∈ℂN×NA\\in\\mathbb\{C\}^\{N\\times N\}be diagonalisable with eigenvector matrixVVsatisfyingκ2​\(V\)≤κ0\\kappa\_\{2\}\(V\)\\leq\\kappa\_\{0\}\(condition \([34](https://arxiv.org/html/2608.10172#S6.E34)\)\), and letAAsatisfy[Assumption˜3](https://arxiv.org/html/2608.10172#Thmassumption3)with gapΔ\\Delta\. For anyA^∈ℂN×N\\hat\{A\}\\in\\mathbb\{C\}^\{N\\times N\}with‖A^−A‖2≤Δ/\(2​κ0\)\\\|\\hat\{A\}\-A\\\|\_\{2\}\\leq\\Delta/\(2\\kappa\_\{0\}\)there exists a unique permutationπ\\piof\{1,…,N\}\\\{1,\\dots,N\\\}such that

maxk⁡\|λk​\(A^\)−λπ​\(k\)​\(A\)\|≤κ0​‖A^−A‖2\.\\max\_\{k\}\\,\\bigl\|\\lambda\_\{k\}\(\\hat\{A\}\)\-\\lambda\_\{\\pi\(k\)\}\(A\)\\bigr\|\\;\\leq\\;\\kappa\_\{0\}\\,\\\|\\hat\{A\}\-A\\\|\_\{2\}\.\(47\)

###### Proof\.

By Bauer–Fike\[[5](https://arxiv.org/html/2608.10172#bib.bib75)\], for every eigenvalueλ^\\hat\{\\lambda\}ofA^\\hat\{A\}we havemink⁡\|λ^−λk​\(A\)\|≤κ0​‖A^−A‖2\\min\_\{k\}\|\\hat\{\\lambda\}\-\\lambda\_\{k\}\(A\)\|\\leq\\kappa\_\{0\}\\,\\\|\\hat\{A\}\-A\\\|\_\{2\}\. Combined with the hypothesis‖A^−A‖2≤Δ/\(2​κ0\)\\\|\\hat\{A\}\-A\\\|\_\{2\}\\leq\\Delta/\(2\\kappa\_\{0\}\), eachλ^\\hat\{\\lambda\}lies withinΔ/2\\Delta/2of a uniqueλk​\(A\)\\lambda\_\{k\}\(A\), because pairs of true eigenvalues are separated byΔ\\Delta\. Defineπ\\piby mappingλ^i\\hat\{\\lambda\}\_\{i\}to the index of its unique nearest neighbour\. Uniqueness of the matching follows by continuity: parameterisingAs=A\+s​\(A^−A\)A\_\{s\}=A\+s\(\\hat\{A\}\-A\)fors∈\[0,1\]s\\in\[0,1\], the eigenvalues ofAsA\_\{s\}vary continuously inssand, by the same Bauer–Fike bound plus the separation, cannot cross\. Hence the map induced by continuous deformation is a permutation, and applying Bauer–Fike to this matching gives the bound\. ∎

###### Lemma 6\.6\(Eigenvector perturbation, non\-normal case\)\.

Under the hypotheses of[Lemma˜6\.5](https://arxiv.org/html/2608.10172#S6.Thmtheorem5), for the permutationπ\\piof that lemma and for eachkk, the right eigenvectors satisfy \- with appropriate normalisation and phase choice \-

‖vk​\(A^\)−vπ​\(k\)​\(A\)‖2≤4​κ02​‖A^−A‖2Δ\.\\bigl\\\|v\_\{k\}\(\\hat\{A\}\)\-v\_\{\\pi\(k\)\}\(A\)\\bigr\\\|\_\{2\}\\;\\leq\\;\\frac\{4\\,\\kappa\_\{0\}^\{2\}\\,\\\|\\hat\{A\}\-A\\\|\_\{2\}\}\{\\Delta\}\.\(48\)

###### Proof\.

Letλ=λπ​\(k\)​\(A\)\\lambda=\\lambda\_\{\\pi\(k\)\}\(A\),v=vπ​\(k\)​\(A\)v=v\_\{\\pi\(k\)\}\(A\),λ^=λk​\(A^\)\\hat\{\\lambda\}=\\lambda\_\{k\}\(\\hat\{A\}\),v^=vk​\(A^\)\\hat\{v\}=v\_\{k\}\(\\hat\{A\}\)\. Consider a circular contourΓλ\\Gamma\_\{\\lambda\}of radiusΔ/2\\Delta/2centred atλ\\lambda; by[Lemma˜6\.5](https://arxiv.org/html/2608.10172#S6.Thmtheorem5)bothλ\\lambdaandλ^\\hat\{\\lambda\}lie strictly insideΓλ\\Gamma\_\{\\lambda\}and no other eigenvalue ofAAorA^\\hat\{A\}lies inside\. The rank\-one Riesz projectorsΠk​\(A\)=12​π​i​∮Γλ\(z​I−A\)−1​𝑑z\\Pi\_\{k\}\(A\)=\\frac\{1\}\{2\\pi i\}\\oint\_\{\\Gamma\_\{\\lambda\}\}\(zI\-A\)^\{\-1\}\\,dzandΠk​\(A^\)\\Pi\_\{k\}\(\\hat\{A\}\)are therefore well defined\.

Resolvent perturbation\[[38](https://arxiv.org/html/2608.10172#bib.bib72), Ch\. V, Thm\. 4\.10\]gives

‖Πk​\(A^\)−Πk​\(A\)‖2≤\|Γλ\|2​π​maxz∈Γλ⁡‖\(z​I−A^\)−1−\(z​I−A\)−1‖2,\\bigl\\\|\\Pi\_\{k\}\(\\hat\{A\}\)\-\\Pi\_\{k\}\(A\)\\bigr\\\|\_\{2\}\\;\\leq\\;\\frac\{\|\\Gamma\_\{\\lambda\}\|\}\{2\\pi\}\\,\\max\_\{z\\in\\Gamma\_\{\\lambda\}\}\\,\\bigl\\\|\(zI\-\\hat\{A\}\)^\{\-1\}\-\(zI\-A\)^\{\-1\}\\bigr\\\|\_\{2\},\(49\)where\|Γλ\|=π​Δ\|\\Gamma\_\{\\lambda\}\|=\\pi\\Deltais the contour length\. Using the identity\(z​I−A^\)−1−\(z​I−A\)−1=\(z​I−A^\)−1​\(A^−A\)​\(z​I−A\)−1\(zI\-\\hat\{A\}\)^\{\-1\}\-\(zI\-A\)^\{\-1\}=\(zI\-\\hat\{A\}\)^\{\-1\}\(\\hat\{A\}\-A\)\(zI\-A\)^\{\-1\}, the bound‖\(z​I−A\)−1‖2≤κ0/\(Δ/2\)=2​κ0/Δ\\\|\(zI\-A\)^\{\-1\}\\\|\_\{2\}\\leq\\kappa\_\{0\}/\(\\Delta/2\)=2\\kappa\_\{0\}/\\Deltaforz∈Γλz\\in\\Gamma\_\{\\lambda\}\(Bauer–Fike again\), and similarly‖\(z​I−A^\)−1‖2≤4​κ0/Δ\\\|\(zI\-\\hat\{A\}\)^\{\-1\}\\\|\_\{2\}\\leq 4\\kappa\_\{0\}/\\Deltaunder the assumed gap condition, we obtain

‖Πk​\(A^\)−Πk​\(A\)‖2≤π​Δ2​π⋅8​κ02Δ2​‖A^−A‖2=4​κ02Δ​‖A^−A‖2\.\\bigl\\\|\\Pi\_\{k\}\(\\hat\{A\}\)\-\\Pi\_\{k\}\(A\)\\bigr\\\|\_\{2\}\\;\\leq\\;\\tfrac\{\\pi\\Delta\}\{2\\pi\}\\cdot\\tfrac\{8\\kappa\_\{0\}^\{2\}\}\{\\Delta^\{2\}\}\\,\\\|\\hat\{A\}\-A\\\|\_\{2\}\\;=\\;\\tfrac\{4\\kappa\_\{0\}^\{2\}\}\{\\Delta\}\\,\\\|\\hat\{A\}\-A\\\|\_\{2\}\.\(50\)For rank\-one projectorsΠ=v​ϕ∗\\Pi=v\\phi^\{\*\}withϕ∗​v=1\\phi^\{\*\}v=1and‖v‖2=1\\\|v\\\|\_\{2\}=1, standard identities relate the projector distance to the vector distance under phase\-optimal choice:‖v^−ei​θ​v‖2≤‖Π^−Π‖2\\\|\\hat\{v\}\-e^\{i\\theta\}v\\\|\_\{2\}\\leq\\\|\\hat\{\\Pi\}\-\\Pi\\\|\_\{2\}for the optimalθ\\theta\[[65](https://arxiv.org/html/2608.10172#bib.bib77), Cor\. V\.4\.11\]\. Choosingv^\\hat\{v\}with this optimal phase gives \([48](https://arxiv.org/html/2608.10172#S6.E48)\)\. ∎

Together,[Lemmas˜6\.5](https://arxiv.org/html/2608.10172#S6.Thmtheorem5)and[6\.6](https://arxiv.org/html/2608.10172#S6.Thmtheorem6)convert an operator\-norm bound of sizeϵ\\epsilononA^−A\\hat\{A\}\-Ainto an eigenvalue rate ofκ0​ϵ\\kappa\_\{0\}\\epsilonand an eigenvector rate of4​κ02​ϵ/Δ4\\kappa\_\{0\}^\{2\}\\epsilon/\\Delta, providedϵ≤Δ/\(2​κ0\)\\epsilon\\leq\\Delta/\(2\\kappa\_\{0\}\)\.

### 6\.6Proof of[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)

###### Proof of[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)\.

Fixδ∈\(0,1/2\)\\delta\\in\(0,1/2\)and letM≥M0M\\geq M\_\{0\}withM0M\_\{0\}as in \([35](https://arxiv.org/html/2608.10172#S6.E35)\)\. We first showM0M\_\{0\}is large enough that all preparatory bounds hold with sufficient probability\.

Apply[Lemma˜6\.4](https://arxiv.org/html/2608.10172#S6.Thmtheorem4)with confidence parameterδ/2\\delta/2: on an eventE1E\_\{1\}of probability at least1−δ/21\-\\delta/2,

‖A^M−A‖2≤c4​L2​\(1\+‖A‖2\)ηΨ​N\+log⁡\(2/δ\)M⏟=⁣:ϵM\.\\bigl\\\|\\hat\{A\}\_\{M\}\-A\\bigr\\\|\_\{2\}\\;\\leq\\;\\underbrace\{\\tfrac\{c\_\{4\}\\,L^\{2\}\(1\+\\\|A\\\|\_\{2\}\)\}\{\\eta\_\{\\Psi\}\}\\,\\sqrt\{\\tfrac\{N\+\\log\(2/\\delta\)\}\{M\}\}\}\_\{\\displaystyle=:\\;\\epsilon\_\{M\}\}\.\(51\)The conditionM≥M0M\\geq M\_\{0\}in \([35](https://arxiv.org/html/2608.10172#S6.E35)\) is chosen precisely so that onE1E\_\{1\}we haveϵM≤Δ/\(2​κ0\)\\epsilon\_\{M\}\\leq\\Delta/\(2\\kappa\_\{0\}\), the hypothesis of[Lemma˜6\.5](https://arxiv.org/html/2608.10172#S6.Thmtheorem5)\. Verifying this: solvingϵM≤Δ/\(2​κ0\)\\epsilon\_\{M\}\\leq\\Delta/\(2\\kappa\_\{0\}\)forMMgives

M≥4​c42​κ02​L4​\(1\+‖A‖2\)2ηΨ2​Δ2​\(N\+log⁡\(2/δ\)\),M\\;\\geq\\;\\frac\{4c\_\{4\}^\{2\}\\kappa\_\{0\}^\{2\}L^\{4\}\(1\+\\\|A\\\|\_\{2\}\)^\{2\}\}\{\\eta\_\{\\Psi\}^\{2\}\\Delta^\{2\}\}\\bigl\(N\+\\log\(2/\\delta\)\\bigr\),\(52\)which is exactly \([35](https://arxiv.org/html/2608.10172#S6.E35)\) withc0=4​c42c\_\{0\}=4c\_\{4\}^\{2\}andlog⁡\(2/δ\)\\log\(2/\\delta\)absorbed into a constant increase ofc0c\_\{0\}\.

OnE1E\_\{1\}, apply[Lemma˜6\.5](https://arxiv.org/html/2608.10172#S6.Thmtheorem5): there is a unique permutationπM\\pi\_\{M\}withmaxk⁡\|λk​\(A^M\)−λπM​\(k\)​\(A\)\|≤κ0​ϵM\\max\_\{k\}\|\\lambda\_\{k\}\(\\hat\{A\}\_\{M\}\)\-\\lambda\_\{\\pi\_\{M\}\(k\)\}\(A\)\|\\leq\\kappa\_\{0\}\\epsilon\_\{M\}, which is \([36](https://arxiv.org/html/2608.10172#S6.E36)\) withc1=c4c\_\{1\}=c\_\{4\}\. Apply[Lemma˜6\.6](https://arxiv.org/html/2608.10172#S6.Thmtheorem6)onE1E\_\{1\}with the sameπM\\pi\_\{M\}: for eachkk,‖vk​\(A^M\)−vπM​\(k\)​\(A\)‖2≤4​κ02​ϵM/Δ\\\|v\_\{k\}\(\\hat\{A\}\_\{M\}\)\-v\_\{\\pi\_\{M\}\(k\)\}\(A\)\\\|\_\{2\}\\leq 4\\kappa\_\{0\}^\{2\}\\epsilon\_\{M\}/\\Delta, which is \([37](https://arxiv.org/html/2608.10172#S6.E37)\) withc2=4​c4c\_\{2\}=4c\_\{4\}\.

Both conclusions hold on the single eventE1E\_\{1\}, so no further union bound is needed and the probability is at least1−δ1\-\\delta\. The constantsc0,c1,c2c\_\{0\},c\_\{1\},c\_\{2\}are absolute, depending only on the constant appearing in[Lemma˜6\.3](https://arxiv.org/html/2608.10172#S6.Thmtheorem3)\. ∎

### 6\.7A gap\-free eigenvalue guarantee, and the resulting split ofM0M\_\{0\}

[Lemma˜6\.5](https://arxiv.org/html/2608.10172#S6.Thmtheorem5)obtains a*bijective*matching by requiring‖A^−A‖2≤Δ/\(2​κ0\)\\\|\\hat\{A\}\-A\\\|\_\{2\}\\leq\\Delta/\(2\\kappa\_\{0\}\), and that hypothesis is the sole origin of theΔ−2\\Delta^\{\-2\}factor in the threshold \([35](https://arxiv.org/html/2608.10172#S6.E35)\)\. It is worth being precise about where the spectral gap does and does not enter, because the answer changes which threshold an experiment should be compared against \- and with the measuredΔ∼10−3\\Delta\\sim 10^\{\-3\}the difference is the difference between a reachable and an unreachable guarantee\.

Tracing the three steps: Step 1 \(concentration ofG^M\\hat\{G\}\_\{M\}\) involves no spectral quantity ofAAat all; Step 2 \([Lemma˜6\.4](https://arxiv.org/html/2608.10172#S6.Thmtheorem4)\) likewise involves none; and Bauer–Fike itself is gap\-free, asserting only that every eigenvalue ofA^\\hat\{A\}lies withinκ0​‖A^−A‖2\\kappa\_\{0\}\\\|\\hat\{A\}\-A\\\|\_\{2\}of*some*eigenvalue ofAA\. The gap is used exclusively to upgrade that “some” to a one\-to\-one correspondence, and again in[Lemma˜6\.6](https://arxiv.org/html/2608.10172#S6.Thmtheorem6), where it is genuinely unavoidable: without separation, eigenvectors are not individually identifiable, since an arbitrarily small perturbation rotates them within a near\-degenerate eigenspace\.

The bijection can instead be obtained with no gap condition whatsoever, at the cost of a dimensional factor, by measuring the discrepancy in the*optimal\-matching*\(Wasserstein\-∞\\infty\) metric\.

###### Lemma 6\.7\(Gap\-free optimal matching\)\.

LetA,A^∈ℂN×NA,\\hat\{A\}\\in\\mathbb\{C\}^\{N\\times N\}withAAdiagonalisable andκ2​\(V\)≤κ0\\kappa\_\{2\}\(V\)\\leq\\kappa\_\{0\}\. Then

minπ∈SN⁡maxk⁡\|λk​\(A^\)−λπ​\(k\)​\(A\)\|≤\(2​N−1\)​κ0​‖A^−A‖2,\\min\_\{\\pi\\in S\_\{N\}\}\\max\_\{k\}\\bigl\|\\lambda\_\{k\}\(\\hat\{A\}\)\-\\lambda\_\{\\pi\(k\)\}\(A\)\\bigr\|\\;\\leq\\;\(2N\-1\)\\,\\kappa\_\{0\}\\,\\\|\\hat\{A\}\-A\\\|\_\{2\},\(53\)where the minimum runs over all permutations of\{1,…,N\}\\\{1,\\dots,N\\\}\. No separation hypothesis is required\.

###### Proof\.

This is the Elsner–Bhatia optimal\-matching bound for the spectral variation of diagonalisable matrices\[[20](https://arxiv.org/html/2608.10172#bib.bib83),[33](https://arxiv.org/html/2608.10172#bib.bib84)\], applied with the diagonalising similarity ofAA; the factorκ0\\kappa\_\{0\}is the conditioning of that similarity and\(2​N−1\)\(2N\-1\)is the standard dimensional constant\. Because the statement already quantifies over all permutations, no argument is needed to establish that a bijection exists \- which is precisely the step in[Lemma˜6\.5](https://arxiv.org/html/2608.10172#S6.Thmtheorem5)that consumed the gap\. ∎

This is the form that matches our experimental protocol: the estimator reported throughout[Section˜11\.2](https://arxiv.org/html/2608.10172#S11.SS2)is the Hungarian\-matched distance, that is, exactly the left\-hand side of \([53](https://arxiv.org/html/2608.10172#S6.E53)\) minimised over permutations\.[Lemma˜6\.7](https://arxiv.org/html/2608.10172#S6.Thmtheorem7), and not[Lemma˜6\.5](https://arxiv.org/html/2608.10172#S6.Thmtheorem5), is therefore the result the measurement should be compared against, and the threshold governing the eigenvalue guarantee carries noΔ\\Delta\. Splitting \([35](https://arxiv.org/html/2608.10172#S6.E35)\) accordingly,

M0eig\\displaystyle M\_\{0\}^\{\\mathrm\{eig\}\}≔c0​κ02​L4​\(1\+‖A‖2\)2ηΨ2​\(N\+log⁡\(1/δ\)\),\\displaystyle\\;\\coloneqq\\;\\frac\{c\_\{0\}\\,\\kappa\_\{0\}^\{2\}\\,L^\{4\}\\,\(1\+\\\|A\\\|\_\{2\}\)^\{2\}\}\{\\eta\_\{\\Psi\}^\{2\}\}\\,\\bigl\(N\+\\log\(1/\\delta\)\\bigr\),\(54\)M0vec\\displaystyle M\_\{0\}^\{\\mathrm\{vec\}\}≔c0′​κ02Δ2​M0eig,\\displaystyle\\;\\coloneqq\\;\\frac\{c\_\{0\}^\{\\prime\}\\,\\kappa\_\{0\}^\{2\}\}\{\\Delta^\{2\}\}\\;M\_\{0\}^\{\\mathrm\{eig\}\},\(55\)so thatM0=M0vecM\_\{0\}=M\_\{0\}^\{\\mathrm\{vec\}\}up to constants, andM0eigM\_\{0\}^\{\\mathrm\{eig\}\}alone suffices for \([36](https://arxiv.org/html/2608.10172#S6.E36)\) once the maximum overkkis read in the optimal\-matching sense\.

###### Theorem 6\.8\(Eigenvalue identifiability without a spectral gap\)\.

Under the hypotheses of[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)but with[Assumption˜3](https://arxiv.org/html/2608.10172#Thmassumption3)omitted, for everyδ∈\(0,1/2\)\\delta\\in\(0,1/2\)and everyM≥M0eigM\\geq M\_\{0\}^\{\\mathrm\{eig\}\}as in \([54](https://arxiv.org/html/2608.10172#S6.E54)\), with probability at least1−δ1\-\\delta,

minπ∈SN⁡maxk\|λk​\(A^M\)−λπ​\(k\)​\(A\)\|≤c1​\(2​N−1\)​κ0​L2​\(1\+‖A‖2\)ηΨ​N\+log⁡\(1/δ\)M\.\\begin\{split\}\\min\_\{\\pi\\in S\_\{N\}\}\\max\_\{k\}\\,&\\bigl\|\\lambda\_\{k\}\(\\hat\{A\}\_\{M\}\)\-\\lambda\_\{\\pi\(k\)\}\(A\)\\bigr\|\\\\ &\\leq\\frac\{c\_\{1\}\(2N\-1\)\\,\\kappa\_\{0\}\\,L^\{2\}\(1\+\\\|A\\\|\_\{2\}\)\}\{\\eta\_\{\\Psi\}\}\\sqrt\{\\frac\{N\+\\log\(1/\\delta\)\}\{M\}\}\.\\end\{split\}\(56\)The rate inMMis unchanged atM−1/2M^\{\-1/2\}; only the constant and the threshold differ\.[Assumption˜3](https://arxiv.org/html/2608.10172#Thmassumption3)remains necessary for the eigenvector statement \([37](https://arxiv.org/html/2608.10172#S6.E37)\) and for the*labelled*\(as opposed to optimally matched\) form of the eigenvalue statement\.

###### Proof\.

Combine[Lemma˜6\.4](https://arxiv.org/html/2608.10172#S6.Thmtheorem4)with[Lemma˜6\.7](https://arxiv.org/html/2608.10172#S6.Thmtheorem7)in place of[Lemma˜6\.5](https://arxiv.org/html/2608.10172#S6.Thmtheorem5)\. The conditionM≥M0eigM\\geq M\_\{0\}^\{\\mathrm\{eig\}\}is used only to ensure the eventE1E\_\{1\}of[Section˜6\.3](https://arxiv.org/html/2608.10172#S6.SS3)holds, on whichG^M\\hat\{G\}\_\{M\}is invertible and the operator\-norm bound applies; no further condition on‖A^M−A‖2\\\|\\hat\{A\}\_\{M\}\-A\\\|\_\{2\}relative toΔ\\Deltais needed, because[Lemma˜6\.7](https://arxiv.org/html/2608.10172#S6.Thmtheorem7)has no such hypothesis\. ∎

Two consequences deserve emphasis\. First, the separation between the two thresholds is a factorκ02/Δ2\\kappa\_\{0\}^\{2\}/\\Delta^\{2\}, and with the empirically measuredΔ∼10−3\\Delta\\sim 10^\{\-3\}this is six orders of magnitude or more; on GPT\-2 small the measured values giveM0eig=1\.0×107M\_\{0\}^\{\\mathrm\{eig\}\}=1\.0\\times 10^\{7\}againstM0vec=6\.4×1016M\_\{0\}^\{\\mathrm\{vec\}\}=6\.4\\times 10^\{16\}, a gap of nine orders of magnitude\. Reporting one threshold for two guarantees with radically different sample requirements understates the strength of the eigenvalue result, and makes an experimentally accessible guarantee look unreachable\. Second, the improvement is not free: the\(2​N−1\)\(2N\-1\)factor in \([53](https://arxiv.org/html/2608.10172#S6.E53)\) means the matching distance degrades linearly in the dictionary sizeNN\. That is itself a testable prediction, and we test it in[Section˜11\.2](https://arxiv.org/html/2608.10172#S11.SS2)by fitting the rate atN∈\{8,16,32,64,128\}N\\in\\\{8,16,32,64,128\\\}and comparing the observedNN\-dependence against the bound\.

### 6\.8Extensions and remarks

###### Corollary 6\.11\(Cluster\-level identifiability\)\.

Suppose[Assumption˜3](https://arxiv.org/html/2608.10172#Thmassumption3)is relaxed to a cluster separation condition: the eigenvalues ofAApartition intorrclustersΛ1,…,Λr\\Lambda\_\{1\},\\dots,\\Lambda\_\{r\}with intra\-cluster diameter at mostΔin\\Delta\_\{\\mathrm\{in\}\}and inter\-cluster gap at leastΔout\>2​Δin\\Delta\_\{\\mathrm\{out\}\}\>2\\Delta\_\{\\mathrm\{in\}\}\. Then, under the remaining hypotheses of[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1), the eigenvalues ofA^M\\hat\{A\}\_\{M\}partition intorrempirical clustersΛ^1,…,Λ^r\\hat\{\\Lambda\}\_\{1\},\\dots,\\hat\{\\Lambda\}\_\{r\}, and there is a permutation of cluster labels such that the Hausdorff distancedH​\(Λ^j,Λj\)d\_\{H\}\(\\hat\{\\Lambda\}\_\{j\},\\Lambda\_\{j\}\)obeys the rate \([36](https://arxiv.org/html/2608.10172#S6.E36)\); the invariant subspaces spanned by the corresponding eigenvectors satisfy asin⁡Θ\\sin\\Thetarate analogous to \([37](https://arxiv.org/html/2608.10172#S6.E37)\) withΔ\\Deltareplaced byΔout−2​Δin\\Delta\_\{\\mathrm\{out\}\}\-2\\Delta\_\{\\mathrm\{in\}\}\.

###### Proof sketch\.

Apply Davis–Kahan for subspaces\[[14](https://arxiv.org/html/2608.10172#bib.bib76),[65](https://arxiv.org/html/2608.10172#bib.bib77)\]to the cluster spectral projectorsΠΛj\\Pi\_\{\\Lambda\_\{j\}\}defined by integration along contours that encloseΛj\\Lambda\_\{j\}but exclude the other clusters\. The gapΔout−2​Δin\\Delta\_\{\\mathrm\{out\}\}\-2\\Delta\_\{\\mathrm\{in\}\}appears as the minimum distance fromΛj\\Lambda\_\{j\}to the exterior contours, playing the role ofΔ\\Deltain[Lemma˜6\.6](https://arxiv.org/html/2608.10172#S6.Thmtheorem6)\. ∎

[Corollary˜6\.11](https://arxiv.org/html/2608.10172#S6.Thmtheorem11)says that identifiability survives the failure of exact simple\-eigenvalue separation: near\-degenerate modes cluster together and each cluster is identified as a joint object\. This matters empirically, since[Assumption˜3](https://arxiv.org/html/2608.10172#Thmassumption3)may hold only coarsely on real fits\.

## 7Optimality of the Rate, and Robustness to Heavy Tails

[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)is a statement about one estimator\. This section characterises the*problem*: no procedure beatsM−1/2M^\{\-1/2\}, and the guarantee survives the failure of the sub\-Gaussian condition \(R3\) that real residual streams are documented to violate\.

### 7\.1A matching minimax lower bound

We work in the noisy\-invariance observation model

yi=A​Ψ​\(xi\)\+B​ui\+ei,ei∼𝒩​\(0,σ2​IN\)\.y\_\{i\}\\;=\\;A\\,\\Psi\(x\_\{i\}\)\+B\\,u\_\{i\}\+e\_\{i\},\\quad e\_\{i\}\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\_\{N\}\)\.\(57\)Write𝒜​\(Δ,κ0\)\\mathcal\{A\}\(\\Delta,\\kappa\_\{0\}\)for the class of diagonalisableAAwith simple spectrum separated byΔ\\Deltaandκ2​\(V\)≤κ0\\kappa\_\{2\}\(V\)\\leq\\kappa\_\{0\}, andg1≔𝔼μ​\|ψ1​\(x\)\|2g\_\{1\}\\coloneqq\\mathbb\{E\}\_\{\\mu\}\|\\psi\_\{1\}\(x\)\|^\{2\}\.

###### Theorem 7\.1\(Minimax optimality inMM\)\.

For everyMMwithσ/2​g1​M≤Δ/2\\sigma/\\sqrt\{2g\_\{1\}M\}\\leq\\Delta/2,

infA^supA∈𝒜​\(Δ,κ0\)𝔼​\[minπ⁡maxk⁡\|λk​\(A^\)−λπ​\(k\)​\(A\)\|\]≥σ8​g1​M,\\inf\_\{\\hat\{A\}\}\\ \\sup\_\{A\\in\\mathcal\{A\}\(\\Delta,\\kappa\_\{0\}\)\}\\ \\mathbb\{E\}\\Bigl\[\\min\_\{\\pi\}\\max\_\{k\}\\bigl\|\\lambda\_\{k\}\(\\hat\{A\}\)\-\\lambda\_\{\\pi\(k\)\}\(A\)\\bigr\|\\Bigr\]\\;\\geq\\;\\frac\{\\sigma\}\{8\\sqrt\{g\_\{1\}M\}\},\(58\)the infimum running over all measurable estimators\.

###### Proof\.

Le Cam’s two\-point method\[[45](https://arxiv.org/html/2608.10172#bib.bib89),[44](https://arxiv.org/html/2608.10172#bib.bib90),[70](https://arxiv.org/html/2608.10172#bib.bib85)\]\. Fix a diagonalA0∈𝒜​\(Δ,κ0\)A\_\{0\}\\in\\mathcal\{A\}\(\\Delta,\\kappa\_\{0\}\)whose eigenvalues are pairwise separated by at least2​Δ2\\Delta, and setA1=A0\+ε​e1​e1⊤A\_\{1\}=A\_\{0\}\+\\varepsilon\\,e\_\{1\}e\_\{1\}^\{\\top\}withε=σ/2​g1​M≤Δ/2\\varepsilon=\\sigma/\\sqrt\{2g\_\{1\}M\}\\leq\\Delta/2\. ∎

*Both hypotheses lie in the class\.*A0A\_\{0\}andA1A\_\{1\}are diagonal, so each is diagonalisable withV=IV=Iandκ2​\(V\)=1≤κ0\\kappa\_\{2\}\(V\)=1\\leq\\kappa\_\{0\}\. Perturbing one diagonal entry byε≤Δ/2\\varepsilon\\leq\\Delta/2leaves every pairwise separation at least2​Δ−ε≥Δ2\\Delta\-\\varepsilon\\geq\\Delta, so the simple\-spectrum and separation conditions hold for both\. Their optimally matched spectral distance is exactlyε\\varepsilon: the matching that pairs equal entries is optimal, and the single perturbed pair contributesε\\varepsilon\.

*The two models are statistically close\.*Condition on the shared design\{\(xi,ui\)\}i=1M\\\{\(x\_\{i\},u\_\{i\}\)\\\}\_\{i=1\}^\{M\}\. Under either hypothesis the observations are Gaussian with common covarianceσ2​IN\\sigma^\{2\}I\_\{N\}and means differing only in the first output coordinate, byε​ψ1​\(xi\)\\varepsilon\\,\\psi\_\{1\}\(x\_\{i\}\)on sampleii\. For Gaussians of common covariance the KL divergence is half the squared Mahalanobis distance between the means, so

KL​\(P1M​‖P0M\|​design\)=∑i=1Mε2​ψ1​\(xi\)22​σ2,\\mathrm\{KL\}\\bigl\(P\_\{1\}^\{M\}\\,\\\|\\,P\_\{0\}^\{M\}\\,\\big\|\\,\\mathrm\{design\}\\bigr\)\\;=\\;\\sum\_\{i=1\}^\{M\}\\frac\{\\varepsilon^\{2\}\\psi\_\{1\}\(x\_\{i\}\)^\{2\}\}\{2\\sigma^\{2\}\},\(59\)whose expectation over the design isM​ε2​g1/\(2​σ2\)=1/4M\\varepsilon^\{2\}g\_\{1\}/\(2\\sigma^\{2\}\)=1/4by the choice ofε\\varepsilon\.

*From closeness to a lower bound\.*Total variation is jointly convex, so𝔼design​TV​\(P0M,P1M\)\\mathbb\{E\}\_\{\\mathrm\{design\}\}\\mathrm\{TV\}\(P\_\{0\}^\{M\},P\_\{1\}^\{M\}\)is at most the total variation between the mixtures; Pinsker with Jensen gives𝔼design​TV≤12​𝔼​KL=1/8<1/2\\mathbb\{E\}\_\{\\mathrm\{design\}\}\\mathrm\{TV\}\\leq\\sqrt\{\\tfrac\{1\}\{2\}\\mathbb\{E\}\\,\\mathrm\{KL\}\}=\\sqrt\{1/8\}<1/2\. The two\-point reduction\[[70](https://arxiv.org/html/2608.10172#bib.bib85), Ch\. 2\], applied to the semi\-distancedist​\(A,A′\)=minπ⁡maxk⁡\|λk​\(A\)−λπ​\(k\)​\(A′\)\|\\mathrm\{dist\}\(A,A^\{\\prime\}\)=\\min\_\{\\pi\}\\max\_\{k\}\|\\lambda\_\{k\}\(A\)\-\\lambda\_\{\\pi\(k\)\}\(A^\{\\prime\}\)\|whose value between the hypotheses isε\\varepsilon, yields

infA^supA∈𝒜𝔼​\[dist​\(A^,A\)\]≥ε2​\(1−TV\)≥ε4=σ4​2​g1​M≥σ8​g1​M\.\\begin\{split\}\\inf\_\{\\hat\{A\}\}\\ \\sup\_\{A\\in\\mathcal\{A\}\}\\ \\mathbb\{E\}\[\\mathrm\{dist\}\(\\hat\{A\},A\)\]&\\;\\geq\\;\\frac\{\\varepsilon\}\{2\}\\bigl\(1\-\\mathrm\{TV\}\\bigr\)\\\\ &\\;\\geq\\;\\frac\{\\varepsilon\}\{4\}\\;=\\;\\frac\{\\sigma\}\{4\\sqrt\{2g\_\{1\}M\}\}\\;\\geq\\;\\frac\{\\sigma\}\{8\\sqrt\{g\_\{1\}M\}\}\.\\end\{split\}\(60\)
The lower bound is dimension\-free, while the upper bound of[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)carriesN\+log⁡\(1/δ\)\\sqrt\{N\+\\log\(1/\\delta\)\}\. The dependence onMMis therefore settled \- no estimator beatsM−1/2M^\{\-1/2\}, and[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)attains it \- while theN\\sqrt\{N\}factor is not\. Closing that gap in either direction, by sharpening the upper bound or by a dimension\-dependent construction in the lower bound, is open\. Together the two theorems characterise the problem rather than only the procedure: the Koopman spectrum is recoverable atM−1/2M^\{\-1/2\}, and nothing does better\.

### 7\.2Heavy\-tailed identifiability

Condition \(R3\) asks the lifted state to be sub\-Gaussian, and real residual streams are documented to violate the analogous condition: transformer activations carry heavy\-tailed outlier coordinates\[[15](https://arxiv.org/html/2608.10172#bib.bib88)\]\. The guarantee does not depend on light tails; only the*estimator*does\. Replacing the empirical Gramians with median\-of\-means Gramians restores the same conclusion under a fourth\-moment condition\.

###### Theorem 7\.2\(Heavy\-tailed identifiability\)\.

Replace \(R3\) by the moment condition

\(R3′\)𝔼​‖Ψ​\(x\)‖24≤κ4​N2and𝔼​‖u‖24≤κ4​p2\.\\text\{\(R3$\{\}^\{\\prime\}$\)\}\\qquad\\mathbb\{E\}\\\|\\Psi\(x\)\\\|\_\{2\}^\{4\}\\leq\\kappa\_\{4\}N^\{2\}\\quad\\text\{and\}\\quad\\mathbb\{E\}\\\|u\\\|\_\{2\}^\{4\}\\leq\\kappa\_\{4\}p^\{2\}\.\(61\)LetA^Mmom\\hat\{A\}^\{\\mathrm\{mom\}\}\_\{M\}be the EDMDc estimator computed from median\-of\-means Gramians overK=⌈8​log⁡\(1/δ\)⌉K=\\lceil 8\\log\(1/\\delta\)\\rceilblocks\[[52](https://arxiv.org/html/2608.10172#bib.bib86),[49](https://arxiv.org/html/2608.10172#bib.bib87)\]\. Then the conclusions \(i\)–\(ii\) of[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)hold forA^Mmom\\hat\{A\}^\{\\mathrm\{mom\}\}\_\{M\}verbatim, with the sub\-Gaussian proxyL2L^\{2\}replaced byC​κ4C\\sqrt\{\\kappa\_\{4\}\}throughout, at the same rateM−1/2M^\{\-1/2\}\.

###### Proof\.

The proof of[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)has exactly one stochastic step,[Lemma˜6\.3](https://arxiv.org/html/2608.10172#S6.Thmtheorem3), which under \(R3\) produces an eventE1E\_\{1\}on which‖G^Z−GZ‖2≤C​L2​\(N\+log⁡\(1/δ\)\)/M\\\|\\hat\{G\}\_\{Z\}\-G\_\{Z\}\\\|\_\{2\}\\leq CL^\{2\}\\sqrt\{\(N\+\\log\(1/\\delta\)\)/M\}\. Every later step \- least\-squares stability \([Lemma˜6\.4](https://arxiv.org/html/2608.10172#S6.Thmtheorem4)\), Bauer–Fike \([Lemma˜6\.5](https://arxiv.org/html/2608.10172#S6.Thmtheorem5)\), Davis–Kahan \([Lemma˜6\.6](https://arxiv.org/html/2608.10172#S6.Thmtheorem6)\), and the gap\-free Elsner–Bhatia variant \([Lemma˜6\.7](https://arxiv.org/html/2608.10172#S6.Thmtheorem7)\) \- is a deterministic perturbation argument consuming only that operator\-norm bound and the invertibility ofG^Z\\hat\{G\}\_\{Z\}it implies\. The claim therefore reduces to producing the same event under \(R3′\) with a different constant\.

Partition theMMsamples intoK=⌈8​log⁡\(1/δ\)⌉K=\\lceil 8\\log\(1/\\delta\)\\rceilblocks of sizem=⌊M/K⌋m=\\lfloor M/K\\rfloorand letG^\(k\)\\hat\{G\}^\{\(k\)\}be the block Gramians\. Under \(R3′\) the summandsz​z∗zz^\{\*\}have finite second moment, with𝔼​‖z​z∗−GZ‖F2≤κ4​\(N\+p\)2\\mathbb\{E\}\\\|zz^\{\*\}\-G\_\{Z\}\\\|\_\{\\mathrm\{F\}\}^\{2\}\\leq\\kappa\_\{4\}\(N\+p\)^\{2\}, so independence within a block gives𝔼​‖G^\(k\)−GZ‖F2≤κ4​\(N\+p\)2/m\\mathbb\{E\}\\\|\\hat\{G\}^\{\(k\)\}\-G\_\{Z\}\\\|\_\{\\mathrm\{F\}\}^\{2\}\\leq\\kappa\_\{4\}\(N\+p\)^\{2\}/m\. Minsker’s geometric\-median\-of\-means bound\[[52](https://arxiv.org/html/2608.10172#bib.bib86), Thm\. 3\.1\], applied in the Hilbert space of Hermitian matrices under the Frobenius norm, then gives

‖medgeo​\(G^\(1\),…,G^\(K\)\)−GZ‖F≤C⋆​κ4​\(N\+p\)2​KM\\Bigl\\\|\\,\\mathrm\{med\}\_\{\\mathrm\{geo\}\}\\bigl\(\\hat\{G\}^\{\(1\)\},\\dots,\\hat\{G\}^\{\(K\)\}\\bigr\)\-G\_\{Z\}\\Bigr\\\|\_\{\\mathrm\{F\}\}\\;\\leq\\;C\_\{\\star\}\\sqrt\{\\frac\{\\kappa\_\{4\}\(N\+p\)^\{2\}K\}\{M\}\}\(62\)with probability at least1−δ1\-\\delta, whereC⋆C\_\{\\star\}is that theorem’s absolute constant \(it may be taken as1111for the geometric median; we do not optimise it\)\. SubstitutingK=⌈8​log⁡\(1/δ\)⌉K=\\lceil 8\\log\(1/\\delta\)\\rceiland passing to the operator norm gives an event of the same shape asE1E\_\{1\}, with the sub\-Gaussian proxyL2L^\{2\}replaced throughout byC​κ4C\\sqrt\{\\kappa\_\{4\}\}and the rate inMMunchanged atM−1/2M^\{\-1/2\}\.

Two features of the geometric median matter for the substitution\. It is a positively weighted combination of its inputs, so the estimate inherits positive semi\-definiteness from the block Gramians andG^Zmom\\hat\{G\}\_\{Z\}^\{\\mathrm\{mom\}\}remains a legitimate Gramian; and the bound requires only a second moment ofz​z∗zz^\{\*\}, which \(R3′\) supplies through the fourth moment ofzz\. Rerunning the remaining steps verbatim on the new event yields conclusions \(i\)–\(ii\) forA^Mmom\\hat\{A\}^\{\\mathrm\{mom\}\}\_\{M\}\. ∎

Le Cam

## 8The Identifiability–Legibility Dissociation

The previous sections identify an object\. This section relates that object to the objects a practitioner actually inspects, and the relation is a theorem rather than an observation: whenever the realisation is non\-normal, the principal directions of the activations and the Koopman modes of the dynamics*cannot*be the same basis\.

### 8\.1The dissociation theorem

###### Theorem 8\.1\(Modal–principal dissociation\)\.

Letρ​\(A\)<1\\rho\(A\)<1withAAdiagonalisable and of simple spectrum \([Assumption˜3](https://arxiv.org/html/2608.10172#Thmassumption3)\), and letΣ≻0\\Sigma\\succ 0be the stationary covariance of the lifted recurrenceΨℓ\+1=A​Ψℓ\+B​uℓ\\Psi\_\{\\ell\+1\}=A\\Psi\_\{\\ell\}\+Bu\_\{\\ell\}under white controls, i\.e\. the unique solution of the discrete Lyapunov equation

Σ=A​Σ​A∗\+B​Gu​B∗\.\\Sigma\\;=\\;A\\Sigma A^\{\*\}\+BG\_\{u\}B^\{\*\}\.\(63\)If some orthonormal eigenbasis ofΣ\\Sigmaconsists of eigenvectors ofAA, thenAAis normal; equivalently,κ2​\(V\)=1\\kappa\_\{2\}\(V\)=1\.

###### Proof\.

Σ\\Sigmais Hermitian positive definite, so it admits an orthonormal eigenbasis\{uk\}\\\{u\_\{k\}\\\}\. By hypothesis eachuku\_\{k\}is an eigenvector ofAA; since the spectrum ofAAis simple, its eigenspaces are one\-dimensional, soU=\[u1​⋯​uN\]U=\[u\_\{1\}\\cdots u\_\{N\}\]is an eigenvector matrix ofAA\.UUis unitary, henceA=U​Λ​U∗A=U\\Lambda U^\{\*\}is normal andκ2​\(U\)=1\\kappa\_\{2\}\(U\)=1\. Contrapositively,κ2​\(V\)\>1\\kappa\_\{2\}\(V\)\>1forces at least one principal direction ofΣ\\Sigmato align with no eigenvector ofAA, so the identifiable modal basis and the variance\-ordered principal basis cannot coincide\. ∎

The theorem concerns the fitted Koopman realisation of the transformer, not an arbitrary matrix\. HereAAis the realisation estimated from the residual stream by EDMDc andΣ\\Sigmais the covariance of its lifted residual states, so the conditionκ2​\(V\)=1\\kappa\_\{2\}\(V\)=1applies to this specific realisation\. The measured condition numbers areκ2​\(V^\)∈\{78\.7,38\.2,494\.7\}\\kappa\_\{2\}\(\\hat\{V\}\)\\in\\\{78\.7,38\.2,494\.7\\\}on GPT\-2 small, Gemma\-2\-2B and Qwen3\-8B\-Base respectively \([Section˜11\.3](https://arxiv.org/html/2608.10172#S11.SS3)\) — all far above11— so the theorem says that on every model in our suite the identifiable Koopman modes must differ from the variance\-ordered principal directions\.

### 8\.2The misalignment saturates

[Theorem˜8\.1](https://arxiv.org/html/2608.10172#S8.Thmtheorem1)says the two bases differ; it does not say by how much\. The following explicit family answers that: the misalignment does not merely fail to vanish, it grows to orthogonality, at rateκ2​\(V\)−1\\kappa\_\{2\}\(V\)^\{\-1\}\.

###### Proposition 8\.2\(The misalignment saturates\)\.

For

Ac=\(λ1c0λ2\),0<λ2<λ1<1,B​Gu​B∗=I,A\_\{c\}=\\begin\{pmatrix\}\\lambda\_\{1\}&c\\\\ 0&\\lambda\_\{2\}\\end\{pmatrix\},\\qquad 0<\\lambda\_\{2\}<\\lambda\_\{1\}<1,\\qquad BG\_\{u\}B^\{\*\}=I,\(64\)asc→∞c\\to\\inftywe haveκ2​\(Vc\)=Θ​\(c\)\\kappa\_\{2\}\(V\_\{c\}\)=\\Theta\(c\), both Koopman modes converge toe1e\_\{1\}, and the minor principal direction ofΣc\\Sigma\_\{c\}converges toe2e\_\{2\}\. The angleθ\\thetabetween that principal direction and the nearest Koopman mode therefore tends toπ/2\\pi/2, with\|cos⁡θ\|=O​\(κ2​\(Vc\)−1\)\|\\cos\\theta\|=O\(\\kappa\_\{2\}\(V\_\{c\}\)^\{\-1\}\)\.

###### Proof\.

The eigenvectors ofAcA\_\{c\}arev1=e1v\_\{1\}=e\_\{1\}andv2∝\(c,λ2−λ1\)⊤v\_\{2\}\\propto\(c,\\lambda\_\{2\}\-\\lambda\_\{1\}\)^\{\\top\}, which normalises toe1\+O​\(c−1\)​e2e\_\{1\}\+O\(c^\{\-1\}\)e\_\{2\}; the two collapse onto one direction andκ2​\(Vc\)=Θ​\(c/\|λ1−λ2\|\)\\kappa\_\{2\}\(V\_\{c\}\)=\\Theta\(c/\|\\lambda\_\{1\}\-\\lambda\_\{2\}\|\)\.

For the stationary covariance, solveΣ=Ac​Σ​Ac∗\+I\\Sigma=A\_\{c\}\\Sigma A\_\{c\}^\{\*\}\+Ientrywise\. The\(2,2\)\(2,2\)entry givesΣ22=λ22​Σ22\+1\\Sigma\_\{22\}=\\lambda\_\{2\}^\{2\}\\Sigma\_\{22\}\+1, soΣ22=\(1−λ22\)−1=Θ​\(1\)\\Sigma\_\{22\}=\(1\-\\lambda\_\{2\}^\{2\}\)^\{\-1\}=\\Theta\(1\)\. The\(1,2\)\(1,2\)entry givesΣ12=λ1​λ2​Σ12\+λ2​c​Σ22\\Sigma\_\{12\}=\\lambda\_\{1\}\\lambda\_\{2\}\\Sigma\_\{12\}\+\\lambda\_\{2\}c\\,\\Sigma\_\{22\}, soΣ12=λ2​c​Σ22/\(1−λ1​λ2\)=Θ​\(c\)\\Sigma\_\{12\}=\\lambda\_\{2\}c\\,\\Sigma\_\{22\}/\(1\-\\lambda\_\{1\}\\lambda\_\{2\}\)=\\Theta\(c\)\. The\(1,1\)\(1,1\)entry givesΣ11=\(2​λ1​c​Σ12\+c2​Σ22\+1\)/\(1−λ12\)=Θ​\(c2\)\\Sigma\_\{11\}=\(2\\lambda\_\{1\}c\\,\\Sigma\_\{12\}\+c^\{2\}\\Sigma\_\{22\}\+1\)/\(1\-\\lambda\_\{1\}^\{2\}\)=\\Theta\(c^\{2\}\)\.

A symmetric2×22\\times 2matrix withΣ11=Θ​\(c2\)\\Sigma\_\{11\}=\\Theta\(c^\{2\}\),Σ12=Θ​\(c\)\\Sigma\_\{12\}=\\Theta\(c\),Σ22=Θ​\(1\)\\Sigma\_\{22\}=\\Theta\(1\)has top eigenvectore1\+O​\(c−1\)​e2e\_\{1\}\+O\(c^\{\-1\}\)e\_\{2\}and minor eigenvectore2\+O​\(c−1\)​e1e\_\{2\}\+O\(c^\{\-1\}\)e\_\{1\}\. The inner product of the minor principal direction with either Koopman mode is thereforeO​\(c−1\)=O​\(κ2​\(Vc\)−1\)O\(c^\{\-1\}\)=O\(\\kappa\_\{2\}\(V\_\{c\}\)^\{\-1\}\)\. ∎

### 8\.3Scope of the claim

Three delimitations are worth stating explicitly, because the theorem is easy to over\-read\.

*It is about the lifted system\.*[Theorem˜8\.1](https://arxiv.org/html/2608.10172#S8.Thmtheorem1)applies to the lifted linear recurrence under white controls, whereas[Section˜11\.3](https://arxiv.org/html/2608.10172#S11.SS3)analyses PCA on raw activations under real attention writes\. The theorem supplies the mechanism, not a literal model of the experiment\.

*It predicts difference, not superiority\.*The theorem says the two bases cannot coincide\. It says nothing about which one better predicts the causal effect of an intervention\.[Section˜11\.3](https://arxiv.org/html/2608.10172#S11.SS3)answers that question empirically, and the answer is not favourable to Koopman modes on the question principal components are optimal for: PCA removes60%60\\%of the IOI logit difference atj=8j=8ablated directions against the modes’25%25\\%\.[Section˜11\.4](https://arxiv.org/html/2608.10172#S11.SS4)then shows that advantage is specific and local, decaying by4\.1×4\.1\\timesas the question moves away in depth\.

*It is one\-directional\.*We prove that non\-normality forces divergence\. We do not prove that the divergence is large for every non\-normal realisation —[Proposition˜8\.2](https://arxiv.org/html/2608.10172#S8.Thmtheorem2)exhibits a family where it saturates, which is a sufficiency statement, not a claim about allAA\.

Within those limits the reading is clear, and it is the paper’s central claim:*the identifiable object and the legible object are not the same object*\. The dissociation between what can be certified and what can be read is, in this precise sense, a prediction of the theory rather than a disappointment of the experiments\.

## 9Implications for Mechanistic Interpretability

The identifiability theorem has consequences beyond its own statement\. This section develops three\. First, the spectral projectors ofAAgenerate a commutative algebra that is*complete*for the first\-order intervention calculus of MI, so that activation patching, path patching, attribution patching and mean ablation all have unique closed\-form representatives in it \([Theorem˜9\.2](https://arxiv.org/html/2608.10172#S9.Thmtheorem2)\)\. Second, the run\-to\-run variability of sparse autoencoders is a structural consequence of an objective that omits \(DR2\), with an explicit penalty as remedy \([Corollary˜9\.5](https://arxiv.org/html/2608.10172#S9.Thmtheorem5)\)\. Third, cross\-model universality becomes a testable spectral criterion \([Corollary˜9\.6](https://arxiv.org/html/2608.10172#S9.Thmtheorem6)\)\.

### 9\.1Algebraic completeness of the intervention calculus

###### Definition 9\.1\(Interventional completeness\)\.

The intervention calculusℐ\\mathcal\{I\}of a mechanistic primitive\(𝒟,E,ℐ\)\(\\mathcal\{D\},E,\\mathcal\{I\}\)is*complete on*ℋE\\mathcal\{H\}\_\{E\}if:

- \(IC1\)ℐ\|ℋE\\mathcal\{I\}\|\_\{\\mathcal\{H\}\_\{E\}\}is closed under composition andℂ\\mathbb\{C\}\-linear combinations;
- \(IC2\)everyI∈ℐ\|ℋEI\\in\\mathcal\{I\}\|\_\{\\mathcal\{H\}\_\{E\}\}commutes with𝒦\|ℋE\\mathcal\{K\}\|\_\{\\mathcal\{H\}\_\{E\}\}\(dynamical consistency\);
- \(IC3\)for everyδ∈ℋE\\delta\\in\\mathcal\{H\}\_\{E\}there existI∈ℐI\\in\\mathcal\{I\}andψ∈ℋE\\psi\\in\\mathcal\{H\}\_\{E\}withI​ψ=δI\\,\\psi=\\delta\(spanning\)\.

\(IC1\) makesℐ\|ℋE\\mathcal\{I\}\|\_\{\\mathcal\{H\}\_\{E\}\}a unital associativeℂ\\mathbb\{C\}\-algebra\. \(IC2\) is the dynamical\-consistency requirement: intervening and then evolving equals evolving and then intervening\. \(IC3\) is a spanning condition ruling out proper subalgebras\. The commutant𝒵​\(AE\)≔\{M:M​AE=AE​M\}\\mathcal\{Z\}\(A\_\{E\}\)\\coloneqq\\\{M:MA\_\{E\}=A\_\{E\}M\\\}of a diagonalisableAEA\_\{E\}withr∗r^\{\*\}distinct eigenvalues equals the polynomial algebraℂ​\[AE\]\\mathbb\{C\}\[A\_\{E\}\]and has dimensionr∗r^\{\*\}; \(IC2\) forcesℐ\|ℋE⊆𝒵​\(AE\)\\mathcal\{I\}\|\_\{\\mathcal\{H\}\_\{E\}\}\\subseteq\\mathcal\{Z\}\(A\_\{E\}\), and \(IC3\) forces equality\.

###### Theorem 9\.2\(Algebraic completeness of the KSA intervention calculus\)\.

Assume the hypotheses of[Theorem˜5\.1](https://arxiv.org/html/2608.10172#S5.Thmtheorem1)and[Assumption˜3](https://arxiv.org/html/2608.10172#Thmassumption3)\. Let\{Πk=vk​ϕk⊤\}k=1N\\\{\\Pi\_\{k\}=v\_\{k\}\\phi\_\{k\}^\{\\top\}\\\}\_\{k=1\}^\{N\}be the spectral projectors ofAA, letr∗r^\{\*\}denote its number of distinct eigenvaluesμ1,…,μr∗\\mu\_\{1\},\\dots,\\mu\_\{r^\{\*\}\}, and let𝒜​\(A\)≔ℂ​\[A\]⊆ℂN×N\\mathcal\{A\}\(A\)\\coloneqq\\mathbb\{C\}\[A\]\\subseteq\\mathbb\{C\}^\{N\\times N\}\. Then

1. \(a\)𝒜​\(A\)\\mathcal\{A\}\(A\)is a commutative unitalℂ\\mathbb\{C\}\-algebra of dimensionr∗r^\{\*\}, spanned by the spectral projectors after identification of projectors sharing an eigenvalue\.
2. \(b\)The evaluation mapp​\(A\)↦\(p​\(μ1\),…,p​\(μr∗\)\)p\(A\)\\mapsto\(p\(\\mu\_\{1\}\),\\dots,p\(\\mu\_\{r^\{\*\}\}\)\)is an isomorphism𝒜​\(A\)≅ℂr∗\\mathcal\{A\}\(A\)\\cong\\mathbb\{C\}^\{r^\{\*\}\}of commutative unitalℂ\\mathbb\{C\}\-algebras\. The pullback of pointwise conjugation onℂr∗\\mathbb\{C\}^\{r^\{\*\}\}endows𝒜​\(A\)\\mathcal\{A\}\(A\)with an involution under which it becomes a commutative∗\*\-algebra isomorphic toℂr∗\\mathbb\{C\}^\{r^\{\*\}\}\.
3. \(c\)For any additive perturbationδ∈ℂN\\delta\\in\\mathbb\{C\}^\{N\}at layerℓ0\\ell\_\{0\}, its downstream effect at layerℓ\>ℓ0\\ell\>\\ell\_\{0\}is Aℓ−ℓ0​δ=∑k=1Nλkℓ−ℓ0​Πk​δ\.A^\{\\ell\-\\ell\_\{0\}\}\\,\\delta\\;=\\;\\sum\_\{k=1\}^\{N\}\\lambda\_\{k\}^\{\\ell\-\\ell\_\{0\}\}\\,\\Pi\_\{k\}\\,\\delta\.\(65\)
4. \(d\)Each of activation, path, attribution, and mean\-ablation patching has a unique closed\-form representative in𝒜​\(A\)\\mathcal\{A\}\(A\), obtained by the spectral decompositionδ=∑kΠk​δ\\delta=\\sum\_\{k\}\\Pi\_\{k\}\\delta\.

###### Proof\.

\(a\) ForAAdiagonalisable with distinct eigenvaluesμ1,…,μr∗\\mu\_\{1\},\\dots,\\mu\_\{r^\{\*\}\}, the projectors satisfy the Lagrange interpolation identityΠk=∏j≠k\(A−μj​I\)/\(μk−μj\)\\Pi\_\{k\}=\\prod\_\{j\\neq k\}\(A\-\\mu\_\{j\}I\)/\(\\mu\_\{k\}\-\\mu\_\{j\}\), so eachΠk∈ℂ​\[A\]\\Pi\_\{k\}\\in\\mathbb\{C\}\[A\]\. Conversely, any polynomial inAAdecomposes asp​\(A\)=∑kp​\(μk\)​Πkp\(A\)=\\sum\_\{k\}p\(\\mu\_\{k\}\)\\Pi\_\{k\}\. Hence𝒜​\(A\)=span⁡\{Π1,…,Πr∗\}\\mathcal\{A\}\(A\)=\\operatorname\{span\}\\\{\\Pi\_\{1\},\\dots,\\Pi\_\{r^\{\*\}\}\\\}, of dimensionr∗r^\{\*\}; commutativity and unitality are immediate\.

\(b\) The evaluation map is well defined \(two polynomials inAAagreeing on\{μk\}\\\{\\mu\_\{k\}\\\}define the same matrix\), surjective \(Lagrange interpolation\), and injective \(a polynomial of degree<r∗<r^\{\*\}vanishing onr∗r^\{\*\}points is identically zero\)\. Pointwise conjugation onℂr∗\\mathbb\{C\}^\{r^\{\*\}\}transports to an involution on𝒜​\(A\)\\mathcal\{A\}\(A\)under which the isomorphism preserves the∗\*\-structure\.

\(c\) FromA=∑kλk​ΠkA=\\sum\_\{k\}\\lambda\_\{k\}\\Pi\_\{k\}andΠj​Πk=δj​k​Πk\\Pi\_\{j\}\\Pi\_\{k\}=\\delta\_\{jk\}\\Pi\_\{k\}we getAm=∑kλkm​ΠkA^\{m\}=\\sum\_\{k\}\\lambda\_\{k\}^\{m\}\\Pi\_\{k\}; apply toδ\\delta\.

\(d\) Each intervention corresponds to an additive perturbationδ\\deltaat some layer:

- •activation patching:δ=Ψ​\(xℓ0′\)−Ψ​\(xℓ0\)\\delta=\\Psi\(x\_\{\\ell\_\{0\}\}^\{\\prime\}\)\-\\Psi\(x\_\{\\ell\_\{0\}\}\);
- •path patching restricted to a subsetSSof modes:δS=∑k∈SΠk​δ\\delta\_\{S\}=\\sum\_\{k\\in S\}\\Pi\_\{k\}\\delta;
- •attribution patching: the linearised downstream effect isC​Aℓ−ℓ0​δ=∑kλkℓ−ℓ0​C​Πk​δCA^\{\\ell\-\\ell\_\{0\}\}\\delta=\\sum\_\{k\}\\lambda\_\{k\}^\{\\ell\-\\ell\_\{0\}\}C\\Pi\_\{k\}\\delta;
- •mean ablation:δ=Ψ​\(x¯\)−Ψ​\(xℓ0\)\\delta=\\Psi\(\\bar\{x\}\)\-\\Psi\(x\_\{\\ell\_\{0\}\}\)\.

Uniqueness of each decomposition follows from \(a\): the spectral projectors are a basis of𝒜​\(A\)\\mathcal\{A\}\(A\), so any element admits a unique expansion in\{Πk\}\\\{\\Pi\_\{k\}\\\}\. ∎

![Refer to caption](https://arxiv.org/html/2608.10172v1/Figures/fig_algebra_completeness.png)Figure 2:Analytic illustration: Algebraic completeness of the KSA intervention calculus \([Theorem˜9\.2](https://arxiv.org/html/2608.10172#S9.Thmtheorem2)\) \(p=113p=113,\|K~\|=12\|\\tilde\{K\}\|=12\)\.\(a\)The algebra𝒜​\(A\)≅ℂ12\\mathcal\{A\}\(A\)\\cong\\mathbb\{C\}^\{12\}: each spectral projectorΠk\\Pi\_\{k\}occupies an independent one\-dimensional subalgebra indexed by the eigenvalueλk\\lambda\_\{k\}on the unit circle\.\(b\)A unit\-norm interventionδ∈ℂ12\\delta\\in\\mathbb\{C\}^\{12\}decomposed into its1212spectral componentsΠk​δ\\Pi\_\{k\}\\delta; the decomposition is unique and exhaustive \(∑kΠk=I\\sum\_\{k\}\\Pi\_\{k\}=I\)\.\(c\)Downstream propagationAℓ​δ=∑kλkℓ​Πk​δA^\{\\ell\}\\delta=\\sum\_\{k\}\\lambda\_\{k\}^\{\\ell\}\\,\\Pi\_\{k\}\\deltaacross eight layers\. Each row is an independent spectral channel rotating at ratearg⁡\(λk\)\\arg\(\\lambda\_\{k\}\)per layer; channels do not interact, confirming the direct\-sum structure of𝒜​\(A\)\\mathcal\{A\}\(A\)\.[Theorem˜9\.2](https://arxiv.org/html/2608.10172#S9.Thmtheorem2)has a foundational consequence: any primitive that is rich and interventionally complete on the same subspace is spectrally the same object\.

###### Corollary 9\.3\(Foundational uniqueness of KSA\)\.

Let\(𝒟,E,ℐ\)\(\\mathcal\{D\},E,\\mathcal\{I\}\)be any mechanistic primitive satisfying[Definition˜3\.2](https://arxiv.org/html/2608.10172#S3.Thmtheorem2)and[Definition˜9\.1](https://arxiv.org/html/2608.10172#S9.Thmtheorem1)on the same underlying𝒦\\mathcal\{K\}\-invariant subspace as KSA,ℋE=ℋN\\mathcal\{H\}\_\{E\}=\\mathcal\{H\}\_\{N\}\. Then there existsT∈G​L​\(N,ℂ\)T\\in GL\(N,\\mathbb\{C\}\)such thatE=T​ΨE=T\\,\\Psiandℐ\|ℋE=T​𝒜​\(A\)​T−1\\mathcal\{I\}\|\_\{\\mathcal\{H\}\_\{E\}\}=T\\,\\mathcal\{A\}\(A\)\\,T^\{\-1\}; consequently\(𝒟,E,ℐ\)\(\\mathcal\{D\},E,\\mathcal\{I\}\)is spectrally equivalent to KSA\.

###### Proof\.

By \(DR2\) and[Theorem˜5\.1](https://arxiv.org/html/2608.10172#S5.Thmtheorem1), the primitive induces a Koopman realisationAEA\_\{E\}withAE=T​A​T−1A\_\{E\}=TAT^\{\-1\}for someT∈G​L​\(N,ℂ\)T\\in GL\(N,\\mathbb\{C\}\)\(uniqueness up to similarity\)\. By[Theorem˜9\.2](https://arxiv.org/html/2608.10172#S9.Thmtheorem2)\(a\),𝒜​\(AE\)=ℂ​\[AE\]=T​ℂ​\[A\]​T−1=T​𝒜​\(A\)​T−1\\mathcal\{A\}\(A\_\{E\}\)=\\mathbb\{C\}\[A\_\{E\}\]=T\\,\\mathbb\{C\}\[A\]\\,T^\{\-1\}=T\\mathcal\{A\}\(A\)T^\{\-1\}\. By \(IC1\)–\(IC3\),ℐ\|ℋE=𝒜​\(AE\)\\mathcal\{I\}\|\_\{\\mathcal\{H\}\_\{E\}\}=\\mathcal\{A\}\(A\_\{E\}\), giving the claim\. ∎

### 9\.2SAE variability is structural, not incidental

###### Corollary 9\.5\(SAE non\-identifiability\)\.

LetDSAED\_\{\\mathrm\{SAE\}\}be a dictionary trained by minimising the reconstruction\-plus\-L1L^\{1\}objective \([3](https://arxiv.org/html/2608.10172#S3.E3)\)\. In generalDSAED\_\{\\mathrm\{SAE\}\}satisfies \(DR1\) but not \(DR2\)\. Consequently[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)does not apply toDSAED\_\{\\mathrm\{SAE\}\}, and the run\-to\-run variability documented by\[[6](https://arxiv.org/html/2608.10172#bib.bib9),[9](https://arxiv.org/html/2608.10172#bib.bib12),[37](https://arxiv.org/html/2608.10172#bib.bib10)\]is a structural consequence of the objective’s failure to enforce𝒦\\mathcal\{K\}\-invariance rather than an artefact of implementation\.

###### Proof\.

The objective \([3](https://arxiv.org/html/2608.10172#S3.E3)\) enforces reconstruction and sparsity but places no constraint requiring the range ofDDto be closed under composition with the transformer’s depth mapFF\. HenceℋD=span⁡\(D\)\\mathcal\{H\}\_\{D\}=\\operatorname\{span\}\(D\)is generically not𝒦\\mathcal\{K\}\-invariant and \(DR2\) fails\. Since \(DR2\) is a hypothesis of[Theorem˜5\.1](https://arxiv.org/html/2608.10172#S5.Thmtheorem1)via[Assumption˜1](https://arxiv.org/html/2608.10172#Thmassumption1), the uniqueness\-up\-to\-similarity conclusion does not hold: different training seeds converging to distinct sparsity–reconstruction minima correspond to different, inequivalentDD’s, and[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)does not certify their spectra as identifiable\. By[Remark˜6\.12](https://arxiv.org/html/2608.10172#S6.Thmtheorem12)the resulting bias does not vanish asM→∞M\\to\\infty\. This is the structural origin of the empirically documented variability\. ∎

This recontextualises an active subfield\. The variability reported across seeds and widths is usually discussed as a tuning problem;[Corollary˜9\.5](https://arxiv.org/html/2608.10172#S9.Thmtheorem5)says no amount of tuning removes it, because the objective is silent about the property identifiability requires\. It also yields a remedy \- augment the objective with an explicit invariance penalty,

ℒ​\(D,AD,z\)\\displaystyle\\mathcal\{L\}\(D,A\_\{D\},z\)=‖x−D​z‖22\+λ​‖z‖1\\displaystyle=\\bigl\\\|x\-Dz\\bigr\\\|\_\{2\}^\{2\}\+\\lambda\\,\\\|z\\\|\_\{1\}\(66\)\+γ​‖ΨD​\(F​\(x,u\)\)−AD​ΨD​\(x\)−BD​u‖22\.\\displaystyle\\quad\+\\gamma\\,\\bigl\\\|\\Psi\_\{D\}\(F\(x,u\)\)\-A\_\{D\}\\,\\Psi\_\{D\}\(x\)\-B\_\{D\}\\,u\\bigr\\\|\_\{2\}^\{2\}\.
whereΨD\\Psi\_\{D\}is the dictionary’s encoder andAD,BDA\_\{D\},B\_\{D\}are learned realisation matrices\. The final term drivesspan⁡\(D\)\\operatorname\{span\}\(D\)toward𝒦\\mathcal\{K\}\-invariance; asγ→∞\\gamma\\to\\inftythe dictionary satisfies \(DR2\) and inherits[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)\. Whether sparsity and invariance are jointly satisfiable in interesting regimes is open\.[Section˜11\.6](https://arxiv.org/html/2608.10172#S11.SS6)runs the sweep and reports both the gain and its cost\.

Throughout,ΨD\\Psi\_\{D\}denotes the*post\-activation*codeΨD​\(x\)=σ​\(Wenc​x\+benc\)\\Psi\_\{D\}\(x\)=\\sigma\(W\_\{\\mathrm\{enc\}\}x\+b\_\{\\mathrm\{enc\}\}\)with the dictionary’s own nonlinearityσ\\sigma\(ReLU, JumpReLU, or TopK\), not the pre\-activation linear mapWenc​xW\_\{\\mathrm\{enc\}\}x\. The distinction is load\-bearing and we make it explicit because it decides whether the measurement in[Section˜11\.5](https://arxiv.org/html/2608.10172#S11.SS5)is informative: a linear dictionary cannot be𝒦\\mathcal\{K\}\-invariant for nonlinearFF, so \(DR2\) would fail*by construction*and the experiment could not come out any other way\. The post\-activation code is genuinely nonlinear and is therefore not excluded from \(DR2\) a priori, which is what makes its failure a measurement rather than a tautology\. We report the pre\-activation variant, along with two further conventions, only as reference floors \([Section˜11\.8\.5](https://arxiv.org/html/2608.10172#S11.SS8.SSS5)\)\.

### 9\.3Cross\-model universality becomes testable

###### Corollary 9\.6\(Universality as a spectral criterion\)\.

LetF1,F2F\_\{1\},F\_\{2\}be transformers with Koopman realisationsA1,A2A\_\{1\},A\_\{2\}on invariant subspaces of common dimensionNN\. Ifσ​\(A1\)=σ​\(A2\)\\sigma\(A\_\{1\}\)=\\sigma\(A\_\{2\}\), then𝒜​\(A1\)≅𝒜​\(A2\)\\mathcal\{A\}\(A\_\{1\}\)\\cong\\mathcal\{A\}\(A\_\{2\}\)as commutative unitalℂ\\mathbb\{C\}\-algebras, with the isomorphism sendingΠk\(1\)\\Pi\_\{k\}^\{\(1\)\}toΠπ​\(k\)\(2\)\\Pi\_\{\\pi\(k\)\}^\{\(2\)\}under the eigenvalue\-matching permutation\. The two models are then indistinguishable under first\-order interventions\.

###### Proof\.

By[Theorem˜9\.2](https://arxiv.org/html/2608.10172#S9.Thmtheorem2)\(b\),𝒜​\(Aj\)≅ℂrj∗\\mathcal\{A\}\(A\_\{j\}\)\\cong\\mathbb\{C\}^\{r\_\{j\}^\{\*\}\}\. Isospectrality givesr1∗=r2∗r\_\{1\}^\{\*\}=r\_\{2\}^\{\*\}, henceℂr1∗=ℂr2∗\\mathbb\{C\}^\{r\_\{1\}^\{\*\}\}=\\mathbb\{C\}^\{r\_\{2\}^\{\*\}\}\. Composing the two isomorphisms yields the desired algebra isomorphism, which sendsΠk\(1\)\\Pi\_\{k\}^\{\(1\)\}\- the projector corresponding toμk∈σ​\(A1\)\\mu\_\{k\}\\in\\sigma\(A\_\{1\}\)\- to the projectorΠπ​\(k\)\(2\)\\Pi\_\{\\pi\(k\)\}^\{\(2\)\}corresponding to the same eigenvalue inσ​\(A2\)\\sigma\(A\_\{2\}\)\. ∎

Universality \- different models learning recognisably similar features\[[7](https://arxiv.org/html/2608.10172#bib.bib1),[68](https://arxiv.org/html/2608.10172#bib.bib2),[55](https://arxiv.org/html/2608.10172#bib.bib45)\]\- has been a regularity in search of a criterion\.[Corollary˜9\.6](https://arxiv.org/html/2608.10172#S9.Thmtheorem6)supplies one, and[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)supplies the rate at which it can be checked: spectra fromMMsamples are accurate toO​\(M−1/2\)O\(M^\{\-1/2\}\), so two models can be declared distinct once their matched distance exceeds that tolerance\. The floor is itself measurable \- the within\-model split\-half distance at the sameMM\- so the criterion can be run rather than deferred\.[Section˜11\.7](https://arxiv.org/html/2608.10172#S11.SS7)runs it, and reports that in its stated form it fails a control it should pass: it separates two seed replicas of one architecture\. The diagnosis there is that the split\-half floor bounds sampling error only, while a cross\-model comparison also carries a dictionary\-mismatch bias that does not shrink withMM, so the criterion needs a calibrated null rather than an estimation bound\.

## 10Instance\-Dependent Certified Reduction

[Theorems˜5\.1](https://arxiv.org/html/2608.10172#S5.Thmtheorem1)and[6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)establish that the Koopman realisation\(A,B,C\)\(A,B,C\)is an intrinsic, identifiable object of the transformer\. For mechanistic\-interpretability purposes we do not want the fullNN\-mode realisation; we want a small subset ofr≪Nr\\ll Nmodes reproducing the model’s input–output behaviour to a specified tolerance\. Balanced truncation\[[53](https://arxiv.org/html/2608.10172#bib.bib64),[27](https://arxiv.org/html/2608.10172#bib.bib38),[3](https://arxiv.org/html/2608.10172#bib.bib40)\]is the classical tool, and its Enns–Glover error bound‖G−Gr‖ℋ∞≤2​∑i\>rσi\\\|G\-G\_\{r\}\\\|\_\{\\mathcal\{H\}\_\{\\infty\}\}\\leq 2\\sum\_\{i\>r\}\\sigma\_\{i\}supplies an*a priori*certificate of faithfulness in the Hardyℋ∞\\mathcal\{H\}\_\{\\infty\}\-norm\.

The classical bound is worst\-case with respect to the input distribution: it holds for any input signal in the unit ball ofℓ2\\ell^\{2\}\. In our setting the inputsuℓu\_\{\\ell\}are not arbitrary — they are attention writes at the analysis position, drawn from a specific empirical distribution with covarianceGu=𝔼ν​\[u​u∗\]G\_\{u\}=\\mathbb\{E\}\_\{\\nu\}\[uu^\{\*\}\]\. Transformer attention writes concentrate in a low\-dimensional subspace\[[17](https://arxiv.org/html/2608.10172#bib.bib68),[26](https://arxiv.org/html/2608.10172#bib.bib67)\], and we measure an effective rank of 10\.8 against an ambient control dimension of40964096on Qwen3\-8B\-Base \([Section˜11\.1](https://arxiv.org/html/2608.10172#S11.SS1)\)\. A worst\-case bound over all unit inputs therefore leaves substantial room for tightening\.

This section proves an instance\-dependent version of the Enns–Glover bound that exploits the empirical concentration ofGuG\_\{u\}\. It is a general model\-reduction result and is not required for the identifiability theorem; we include it because it is what makes the identified realisation usable at interpretable size, with a certificate\.

### 10\.1Statement

We work under[Assumptions˜1](https://arxiv.org/html/2608.10172#Thmassumption1)and[2](https://arxiv.org/html/2608.10172#Thmassumption2)and additionally impose:

- \(S1\)AAis Schur\-stable:ρ​\(A\)<1\\rho\(A\)<1\.
- \(S2\)\(A,B,C\)\(A,B,C\)is a minimal realisation: the observability GramianQQand the classical controllability GramianPP\(defined below\) are positive definite\.

Both are standard preliminaries for balanced truncation\. \(S2\) is generic and can be enforced by removing uncontrollable or unobservable modes; \(S1\) may fail on the empirical spectrum ofA^M\\hat\{A\}\_\{M\}, in which case a standard damping shiftA↦γ​AA\\mapsto\\gamma Awithγ∈\(0,1/ρ​\(A\)\)\\gamma\\in\(0,1/\\rho\(A\)\)recovers it at the cost of aγ\\gamma\-dependent factor in the error bound\.

Define the classical controllability and observability Gramians of\(A,B,C\)\(A,B,C\),

P≔∑k=0∞Ak​B​B∗​\(A∗\)k,Q≔∑k=0∞\(A∗\)k​C∗​C​Ak,\\displaystyle P\\;\\coloneqq\\;\\sum\_\{k=0\}^\{\\infty\}A^\{k\}\\,B\\,B^\{\*\}\\,\(A^\{\*\}\)^\{k\},\\qquad Q\\;\\coloneqq\\;\\sum\_\{k=0\}^\{\\infty\}\(A^\{\*\}\)^\{k\}\\,C^\{\*\}\\,C\\,A^\{k\},\(67\)the unique positive\-definite solutions of the discrete\-time Lyapunov equationsA​P​A∗−P\+B​B∗=0APA^\{\*\}\-P\+BB^\{\*\}=0andA∗​Q​A−Q\+C∗​C=0A^\{\*\}QA\-Q\+C^\{\*\}C=0under \(S1\)–\(S2\)\[[3](https://arxiv.org/html/2608.10172#bib.bib40)\]\. The classical Hankel singular values \(HSVs\) areσiclas≔λi​\(P​Q\)\\sigma\_\{i\}^\{\\mathrm\{clas\}\}\\coloneqq\\sqrt\{\\lambda\_\{i\}\(PQ\)\}, ordered decreasingly\. Define the input\-weighted controllability Gramian

Peff≔∑k=0∞Ak​B​Gu​B∗​\(A∗\)k,P\_\{\\mathrm\{eff\}\}\\;\\coloneqq\\;\\sum\_\{k=0\}^\{\\infty\}A^\{k\}\\,B\\,G\_\{u\}\\,B^\{\*\}\\,\(A^\{\*\}\)^\{k\},\(68\)the unique positive\-semidefinite solution ofA​Peff​A∗−Peff\+B​Gu​B∗=0AP\_\{\\mathrm\{eff\}\}A^\{\*\}\-P\_\{\\mathrm\{eff\}\}\+BG\_\{u\}B^\{\*\}=0, and the*input\-weighted Hankel singular values*σieff≔λi​\(Peff​Q\)\\sigma\_\{i\}^\{\\mathrm\{eff\}\}\\coloneqq\\sqrt\{\\lambda\_\{i\}\(P\_\{\\mathrm\{eff\}\}Q\)\}, which are the HSVs of the input\-weighted realisation\(A,B​Gu1/2,C\)\(A,\\,BG\_\{u\}^\{1/2\},\\,C\)\.

LetG​\(z\)=C​\(z​I−A\)−1​BG\(z\)=C\(zI\-A\)^\{\-1\}Bdenote the classical transfer function,G~​\(z\)=C​\(z​I−A\)−1​B​Gu1/2\\tilde\{G\}\(z\)=C\(zI\-A\)^\{\-1\}BG\_\{u\}^\{1/2\}the input\-weighted transfer function, andG~r​\(z\)\\tilde\{G\}\_\{r\}\(z\)the balanced truncation ofG~\\tilde\{G\}to orderrr, computed on the weighted realisation\. Finally define the*effective rank*of the control distribution,

reff​\(u\)≔tr⁡\(Gu\)λmax​\(Gu\)∈\[1,p\],r\_\{\\mathrm\{eff\}\}\(u\)\\;\\coloneqq\\;\\frac\{\\operatorname\{tr\}\(G\_\{u\}\)\}\{\\lambda\_\{\\max\}\(G\_\{u\}\)\}\\;\\in\\;\[1,p\],\(69\)with extremesreff=pr\_\{\\mathrm\{eff\}\}=p\(isotropicGuG\_\{u\}\) andreff=1r\_\{\\mathrm\{eff\}\}=1\(rank\-oneGuG\_\{u\}\); intermediate values count the effective input directions\.

###### Theorem 10\.1\(Instance\-dependent certified reduction\)\.

Assume[Assumptions˜1](https://arxiv.org/html/2608.10172#Thmassumption1)and[2](https://arxiv.org/html/2608.10172#Thmassumption2), \(S1\)–\(S2\), and letr∈\{1,…,N−1\}r\\in\\\{1,\\dots,N\-1\\\}\.

1. \(a\)*\(Certified reduction bound\.\)* ‖G~−G~r‖ℋ∞≤2​∑i=r\+1Nσieff\.\\bigl\\\|\\tilde\{G\}\-\\tilde\{G\}\_\{r\}\\bigr\\\|\_\{\\mathcal\{H\}\_\{\\infty\}\}\\;\\leq\\;2\\sum\_\{i=r\+1\}^\{N\}\\sigma\_\{i\}^\{\\mathrm\{eff\}\}\.\(70\)
2. \(b\)*\(Comparison with the classical bound\.\)*For everyii, σieff≤λmax​\(Gu\)​σiclas\.\\sigma\_\{i\}^\{\\mathrm\{eff\}\}\\;\\leq\\;\\sqrt\{\\lambda\_\{\\max\}\(G\_\{u\}\)\}\\,\\sigma\_\{i\}^\{\\mathrm\{clas\}\}\.\(71\)
3. \(c\)*\(Strict sharpening\.\)*Assume in addition thatBBis not concentrated in the leading eigenspace ofGuG\_\{u\}, in the sense thatrank⁡\(B​\(λmax​\(Gu\)​I−Gu\)​B∗\)≥1\\operatorname\{rank\}\\bigl\(B\(\\lambda\_\{\\max\}\(G\_\{u\}\)I\-G\_\{u\}\)B^\{\*\}\\bigr\)\\geq 1\. Then \([71](https://arxiv.org/html/2608.10172#S10.E71)\) is strict for at leastN−reff​\(u\)N\-r\_\{\\mathrm\{eff\}\}\(u\)indicesii, and consequently ∑i=r\+1Nσieff<λmax​\(Gu\)​∑i=r\+1Nσiclas\\sum\_\{i=r\+1\}^\{N\}\\sigma\_\{i\}^\{\\mathrm\{eff\}\}\\;<\\;\\sqrt\{\\lambda\_\{\\max\}\(G\_\{u\}\)\}\\,\\sum\_\{i=r\+1\}^\{N\}\\sigma\_\{i\}^\{\\mathrm\{clas\}\}\(72\)wheneverr<N−reff​\(u\)r<N\-r\_\{\\mathrm\{eff\}\}\(u\)\.

Part \(a\) is the certified bound\. Part \(b\) shows the new bound never exceeds the classical Enns–Glover bound scaled by the input strengthλmax​\(Gu\)\\sqrt\{\\lambda\_\{\\max\}\(G\_\{u\}\)\}, the natural scale of unit\-covariance inputs\. Part \(c\) shows the sharpening is*strict*once the control distribution is anisotropic and the truncation is not too aggressive\.

### 10\.2Preparatory lemmas

###### Lemma 10\.2\(The input\-weighted realisation\)\.

Under \(S1\)–\(S2\) andGu≻0G\_\{u\}\\succ 0, the realisation\(A,B​Gu1/2,C\)\(A,\\,BG\_\{u\}^\{1/2\},\\,C\)satisfies: \(i\) its controllability Gramian is exactlyPeffP\_\{\\mathrm\{eff\}\}of \([68](https://arxiv.org/html/2608.10172#S10.E68)\); \(ii\) its observability Gramian isQQ, unchanged; \(iii\) its transfer function isG~​\(z\)=G​\(z\)​Gu1/2\\tilde\{G\}\(z\)=G\(z\)G\_\{u\}^\{1/2\}; \(iv\) its Hankel singular values are the\{σieff\}\\\{\\sigma\_\{i\}^\{\\mathrm\{eff\}\}\\\}; \(v\) it is Schur\-stable and minimal\.

###### Proof\.

\(i\) The controllability Gramian of\(A,B′,C\)\(A,B^\{\\prime\},C\)solvesA​P′​A∗−P′\+B′​B′⁣∗=0AP^\{\\prime\}A^\{\*\}\-P^\{\\prime\}\+B^\{\\prime\}B^\{\\prime\*\}=0; withB′=B​Gu1/2B^\{\\prime\}=BG\_\{u\}^\{1/2\}we haveB′​B′⁣∗=B​Gu​B∗B^\{\\prime\}B^\{\\prime\*\}=BG\_\{u\}B^\{\*\}, so the solution isPeffP\_\{\\mathrm\{eff\}\}\. \(ii\) The observability Gramian depends only on\(A,C\)\(A,C\)\. \(iii\) Direct computation:C​\(z​I−A\)−1​\(B​Gu1/2\)=\[C​\(z​I−A\)−1​B\]​Gu1/2C\(zI\-A\)^\{\-1\}\(BG\_\{u\}^\{1/2\}\)=\[C\(zI\-A\)^\{\-1\}B\]G\_\{u\}^\{1/2\}\. \(iv\) By definition of the HSVs of a realisation\[[3](https://arxiv.org/html/2608.10172#bib.bib40)\]\. \(v\)AAis unchanged and Schur\-stable; minimality follows from \(S2\) andGu≻0G\_\{u\}\\succ 0, since the range ofB​Gu1/2BG\_\{u\}^\{1/2\}equals the range ofBB\. ∎

###### Lemma 10\.3\(Weighted Enns–Glover bound\)\.

Under \(S1\)–\(S2\) andGu≻0G\_\{u\}\\succ 0, the balanced truncationG~r\\tilde\{G\}\_\{r\}of the input\-weighted realisation satisfies

‖G~−G~r‖ℋ∞≤2​∑i=r\+1Nσieff\.\\bigl\\\|\\tilde\{G\}\-\\tilde\{G\}\_\{r\}\\bigr\\\|\_\{\\mathcal\{H\}\_\{\\infty\}\}\\;\\leq\\;2\\sum\_\{i=r\+1\}^\{N\}\\sigma\_\{i\}^\{\\mathrm\{eff\}\}\.\(73\)

###### Proof\.

By[Lemma˜10\.2](https://arxiv.org/html/2608.10172#S10.Thmtheorem2)the input\-weighted realisation is Schur\-stable and minimal\. Apply the classical Enns–Glover bound\[[27](https://arxiv.org/html/2608.10172#bib.bib38),[3](https://arxiv.org/html/2608.10172#bib.bib40)\]in its discrete\-time form to this realisation: for a Schur\-stable minimal realisation\(A~,B~,C~\)\(\\tilde\{A\},\\tilde\{B\},\\tilde\{C\}\)with HSVsσ~i\\tilde\{\\sigma\}\_\{i\}, the balanced truncation of orderrrsatisfies‖G~−G~r‖ℋ∞≤2​∑i\>rσ~i\\\|\\tilde\{G\}\-\\tilde\{G\}\_\{r\}\\\|\_\{\\mathcal\{H\}\_\{\\infty\}\}\\leq 2\\sum\_\{i\>r\}\\tilde\{\\sigma\}\_\{i\}\. Substituting\(A~,B~,C~\)=\(A,B​Gu1/2,C\)\(\\tilde\{A\},\\tilde\{B\},\\tilde\{C\}\)=\(A,BG\_\{u\}^\{1/2\},C\)andσ~i=σieff\\tilde\{\\sigma\}\_\{i\}=\\sigma\_\{i\}^\{\\mathrm\{eff\}\}gives the claim\. ∎

###### Lemma 10\.4\(Weighted HSVs dominated by classical HSVs\)\.

Under \(S1\)–\(S2\) andGu≻0G\_\{u\}\\succ 0,

σieff≤λmax​\(Gu\)​σiclasfor all​i,\\sigma\_\{i\}^\{\\mathrm\{eff\}\}\\;\\leq\\;\\sqrt\{\\lambda\_\{\\max\}\(G\_\{u\}\)\}\\,\\sigma\_\{i\}^\{\\mathrm\{clas\}\}\\qquad\\text\{for all \}i,\(74\)with strict inequality for at leastN−reff​\(u\)N\-r\_\{\\mathrm\{eff\}\}\(u\)indices wheneverrank⁡\(B​\(λmax​\(Gu\)​I−Gu\)​B∗\)≥1\\operatorname\{rank\}\\bigl\(B\(\\lambda\_\{\\max\}\(G\_\{u\}\)I\-G\_\{u\}\)B^\{\*\}\\bigr\)\\geq 1\.

###### Proof\.

Writeλ⋆≔λmax​\(Gu\)\\lambda\_\{\\star\}\\coloneqq\\lambda\_\{\\max\}\(G\_\{u\}\)\. SinceGu⪯λ⋆​IG\_\{u\}\\preceq\\lambda\_\{\\star\}Iwe haveB​Gu​B∗⪯λ⋆​B​B∗B\\,G\_\{u\}\\,B^\{\*\}\\preceq\\lambda\_\{\\star\}\\,B\\,B^\{\*\}\. BecauseAk​M​\(A∗\)k⪰0A^\{k\}M\(A^\{\*\}\)^\{k\}\\succeq 0wheneverM⪰0M\\succeq 0, and the Löwner ordering is preserved under conjugation, summing overk≥0k\\geq 0yieldsPeff⪯λ⋆​PP\_\{\\mathrm\{eff\}\}\\preceq\\lambda\_\{\\star\}\\,P\. Conjugating byQ1/2Q^\{1/2\}, which preserves the ordering sinceQ≻0Q\\succ 0,

Q1/2​Peff​Q1/2⪯λ⋆​Q1/2​P​Q1/2\.Q^\{1/2\}\\,P\_\{\\mathrm\{eff\}\}\\,Q^\{1/2\}\\;\\preceq\\;\\lambda\_\{\\star\}\\,Q^\{1/2\}\\,P\\,Q^\{1/2\}\.\(75\)The eigenvalues ofQ1/2​Peff​Q1/2Q^\{1/2\}P\_\{\\mathrm\{eff\}\}Q^\{1/2\}coincide with those ofPeff​QP\_\{\\mathrm\{eff\}\}Q\(they areN​N∗NN^\{\*\}andN∗​NN^\{\*\}Nfor the same matrix\), and similarly for the classical pair, soλi​\(Peff​Q\)≤λ⋆​λi​\(P​Q\)\\lambda\_\{i\}\(P\_\{\\mathrm\{eff\}\}Q\)\\leq\\lambda\_\{\\star\}\\,\\lambda\_\{i\}\(PQ\)for alliiby monotonicity of eigenvalues under the Löwner order\. Taking square roots gives \([74](https://arxiv.org/html/2608.10172#S10.E74)\)\.

For strictness, the assumption meansM≔B​\(λ⋆​I−Gu\)​B∗⪰0M\\coloneqq B\(\\lambda\_\{\\star\}I\-G\_\{u\}\)B^\{\*\}\\succeq 0is non\-zero\. Under \(S1\) and minimality,ΔP≔λ⋆​P−Peff=∑k≥0Ak​M​\(A∗\)k\\Delta\_\{P\}\\coloneqq\\lambda\_\{\\star\}P\-P\_\{\\mathrm\{eff\}\}=\\sum\_\{k\\geq 0\}A^\{k\}M\(A^\{\*\}\)^\{k\}is positive semidefinite and non\-zero, and by minimality positive definite on the reachable subspace generated by iteratingAAon the range ofMM\. The subspace on whichQ1/2​ΔP​Q1/2Q^\{1/2\}\\Delta\_\{P\}Q^\{1/2\}has strictly positive eigenvalues has dimension at leastrank⁡\(ΔP\)≥N−reff​\(u\)\\operatorname\{rank\}\(\\Delta\_\{P\}\)\\geq N\-r\_\{\\mathrm\{eff\}\}\(u\), because the null space ofλ⋆​I−Gu\\lambda\_\{\\star\}I\-G\_\{u\}has dimension equal to the multiplicity ofλ⋆\\lambda\_\{\\star\}, which is at mostreff​\(u\)r\_\{\\mathrm\{eff\}\}\(u\)\. Strict inequality therefore holds for at leastN−reff​\(u\)N\-r\_\{\\mathrm\{eff\}\}\(u\)indices\. ∎

### 10\.3Proof of[Theorem˜10\.1](https://arxiv.org/html/2608.10172#S10.Thmtheorem1)

###### Proof of[Theorem˜10\.1](https://arxiv.org/html/2608.10172#S10.Thmtheorem1)\.

Part \(a\) is[Lemma˜10\.3](https://arxiv.org/html/2608.10172#S10.Thmtheorem3)\. Part \(b\) is[Lemma˜10\.4](https://arxiv.org/html/2608.10172#S10.Thmtheorem4)\. Part \(c\) follows from the strict\-inequality statement of[Lemma˜10\.4](https://arxiv.org/html/2608.10172#S10.Thmtheorem4): if at leastN−reff​\(u\)N\-r\_\{\\mathrm\{eff\}\}\(u\)of the inequalitiesσieff≤λ⋆​σiclas\\sigma\_\{i\}^\{\\mathrm\{eff\}\}\\leq\\sqrt\{\\lambda\_\{\\star\}\}\\,\\sigma\_\{i\}^\{\\mathrm\{clas\}\}are strict, then∑i\>rσieff\\sum\_\{i\>r\}\\sigma\_\{i\}^\{\\mathrm\{eff\}\}falls strictly belowλ⋆​∑i\>rσiclas\\sqrt\{\\lambda\_\{\\star\}\}\\sum\_\{i\>r\}\\sigma\_\{i\}^\{\\mathrm\{clas\}\}provided at least one strict index lies in\{r\+1,…,N\}\\\{r\+1,\\dots,N\\\}\. A pigeonhole argument guarantees this wheneverr<N−reff​\(u\)r<N\-r\_\{\\mathrm\{eff\}\}\(u\): the number of indices in\{r\+1,…,N\}\\\{r\+1,\\dots,N\\\}isN−rN\-r, at mostreff​\(u\)r\_\{\\mathrm\{eff\}\}\(u\)of them can be non\-strict, so at leastN−r−reff​\(u\)≥1N\-r\-r\_\{\\mathrm\{eff\}\}\(u\)\\geq 1are strict\. ∎

### 10\.4The transformer regime

The sharpening becomes concrete when the effective rank of the control distribution is small relative to the ambient dimension — the regime the measurements of[Section˜11\.1](https://arxiv.org/html/2608.10172#S11.SS1)place transformer attention writes in\.

###### Corollary 10\.6\(Transformer\-regime sharpening\)\.

Under the hypotheses of[Theorem˜10\.1](https://arxiv.org/html/2608.10172#S10.Thmtheorem1), suppose the control covarianceGuG\_\{u\}has eigenvaluesd1≥⋯≥dp\>0d\_\{1\}\\geq\\cdots\\geq d\_\{p\}\>0satisfying a power\-law decaydj≤d1​j−αd\_\{j\}\\leq d\_\{1\}j^\{\-\\alpha\}for someα\>1\\alpha\>1\. Thenreff​\(u\)≤ζ​\(α\)r\_\{\\mathrm\{eff\}\}\(u\)\\leq\\zeta\(\\alpha\), the Riemann zeta function ofα\\alpha, and the certified bound obeys

2​∑i\>rσieff≤d1​\[2​∑i\>rσiclas\]⋅\(reff​\(u\)p\)1/2\+o​\(1\)2\\sum\_\{i\>r\}\\sigma\_\{i\}^\{\\mathrm\{eff\}\}\\;\\leq\\;\\sqrt\{d\_\{1\}\}\\,\\Bigl\[2\\sum\_\{i\>r\}\\sigma\_\{i\}^\{\\mathrm\{clas\}\}\\Bigr\]\\cdot\\Bigl\(\\tfrac\{r\_\{\\mathrm\{eff\}\}\(u\)\}\{p\}\\Bigr\)^\{1/2\+o\(1\)\}\(76\)in the regimep→∞p\\to\\inftywithα\\alphafixed andr≥reff​\(u\)r\\geq r\_\{\\mathrm\{eff\}\}\(u\)\.

###### Proof\.

Under the power\-law decay,tr⁡\(Gu\)=∑jdj≤d1​ζ​\(α\)\\operatorname\{tr\}\(G\_\{u\}\)=\\sum\_\{j\}d\_\{j\}\\leq d\_\{1\}\\zeta\(\\alpha\)forα\>1\\alpha\>1, whencereff​\(u\)=tr⁡\(Gu\)/d1≤ζ​\(α\)r\_\{\\mathrm\{eff\}\}\(u\)=\\operatorname\{tr\}\(G\_\{u\}\)/d\_\{1\}\\leq\\zeta\(\\alpha\)\. The sharpening factor then follows from an index\-by\-index refinement of[Lemma˜10\.4](https://arxiv.org/html/2608.10172#S10.Thmtheorem4)in the vein of\[[65](https://arxiv.org/html/2608.10172#bib.bib77)\]: for genericAAandr≥reffr\\geq r\_\{\\mathrm\{eff\}\}, the trailing HSVsσr\+1eff,σr\+2eff,…\\sigma\_\{r\+1\}^\{\\mathrm\{eff\}\},\\sigma\_\{r\+2\}^\{\\mathrm\{eff\}\},\\dotsinherit the tail of theGuG\_\{u\}\-spectrum, yielding the exponent1/2\+o​\(1\)1/2\+o\(1\)\. ∎

[Corollary˜10\.6](https://arxiv.org/html/2608.10172#S10.Thmtheorem6)makes the sharpening quantitative: when the control distribution has effective rankreff≪pr\_\{\\mathrm\{eff\}\}\\ll p, the certified reduction bound is smaller than the classical bound by a factor of orderreff/p\\sqrt\{r\_\{\\mathrm\{eff\}\}/p\}\. For a Llama\- or Gemma\-scale transformer withp=d∼103p=d\\sim 10^\{3\}–10410^\{4\}and empiricalreff∼101r\_\{\\mathrm\{eff\}\}\\sim 10^\{1\}–10210^\{2\}— we measure 10\.8 of40964096on Qwen3\-8B\-Base — the improvement factor is10−2≈10−1\\sqrt\{10^\{\-2\}\}\\approx 10^\{\-1\}or better: an order\-of\-magnitude tightening of the certified truncation error\.

### 10\.5Remarks and interpretation

## 11Experiments

The theory makes predictions that can fail\. This section runs them\.[Section˜11\.1](https://arxiv.org/html/2608.10172#S11.SS1)fixes the protocol;[Section˜11\.2](https://arxiv.org/html/2608.10172#S11.SS2)measures the convergence rate of[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)and diagnoses where it falls short;[Sections˜11\.3](https://arxiv.org/html/2608.10172#S11.SS3)and[11\.4](https://arxiv.org/html/2608.10172#S11.SS4)test whether the identified modes resolve a known circuit, and find that they do not, in a way[Theorem˜8\.1](https://arxiv.org/html/2608.10172#S8.Thmtheorem1)anticipates;[Sections˜11\.5](https://arxiv.org/html/2608.10172#S11.SS5)and[11\.6](https://arxiv.org/html/2608.10172#S11.SS6)test[Corollary˜9\.5](https://arxiv.org/html/2608.10172#S9.Thmtheorem5)and its remedy;[Section˜11\.7](https://arxiv.org/html/2608.10172#S11.SS7)runs the universality criterion of[Corollary˜9\.6](https://arxiv.org/html/2608.10172#S9.Thmtheorem6)together with the control that invalidates its stated form\.[Section˜11\.8](https://arxiv.org/html/2608.10172#S11.SS8)measures the three modelling conventions of[Section˜4\.6](https://arxiv.org/html/2608.10172#S4.SS6)\. Every failure we found is reported, because each of them bounds the claim\.

### 11\.1Protocol

We evaluate on the four publicly available pretrained transformers of[Table˜1](https://arxiv.org/html/2608.10172#S11.T1)\. The suite spans two orders of magnitude in parameter count, three architecture families and three SAE nonlinearities, so[Corollary˜9\.5](https://arxiv.org/html/2608.10172#S9.Thmtheorem5)is tested against ReLU, JumpReLU and TopK dictionaries\. The three Pythia\-160M checkpoints differ only in initialisation seed and are the control for the universality criterion\.

Table 1:The model suite\.LLlayers, residual widthdd,HHattention heads\. Each of the first three models has a public residual\-stream SAE suite with a different nonlinearity\.For each model and each analysed layer we cache the depth\-recurrence triple\(xℓ,uℓ,xℓ\+1\)\(x\_\{\\ell\},u\_\{\\ell\},x\_\{\\ell\+1\}\)at a sampled set of token positions, withuℓu\_\{\\ell\}the attention\-block output, exactly as[Definition˜4\.1](https://arxiv.org/html/2608.10172#S4.Thmtheorem1)prescribes\. The calibration corpus is WikiText\-103 \(train split\) throughout, tokenised at sequence length128128; relative depths areℓ/L∈\{0\.25,0\.5,0\.75\}\\ell/L\\in\\\{0\.25,0\.5,0\.75\\\}with mid\-depth used for headline numbers\.

We compare a*spectral*dictionary combining whitened leading principal directions with random Fourier features\[[78](https://arxiv.org/html/2608.10172#bib.bib57)\]; an*SAE*dictionary given by the post\-activation codeσ​\(Wenc​x\+benc\)\\sigma\(W\_\{\\mathrm\{enc\}\}x\+b\_\{\\mathrm\{enc\}\}\)of a public sparse autoencoder; and a*random*orthonormal control from the QR factorisation of a Gaussian matrix\.

Three design choices could have driven the outcome and are therefore stated here rather than buried\.

- •*The nonlinear block must be genuinely nonlinear\.*A linear dictionary cannot be𝒦\\mathcal\{K\}\-invariant for nonlinearFF, so \(DR2\) would fail by construction\. Principal directions of RMSNorm\-normalised states do*not*qualify, having canonical correlations above0\.990\.99with the linear block on GPT\-2\.
- •*Every dictionary is whitened\.*This is a change of basis, soσ​\(A\)\\sigma\(A\)is unchanged by[Theorem˜5\.1](https://arxiv.org/html/2608.10172#S5.Thmtheorem1)\(b\), but every bound divides byηΨ=λmin​\(GΨ\)\\eta\_\{\\Psi\}=\\lambda\_\{\\min\}\(G\_\{\\Psi\}\), and unwhitened we measureηΨ≈7×10−3\\eta\_\{\\Psi\}\\approx 7\\times 10^\{\-3\}against0\.680\.68–0\.840\.84after whitening\. Comparing dictionaries at matchedNNbut unmatched conditioning measures the conditioning, not the invariance;[Section˜11\.5](https://arxiv.org/html/2608.10172#S11.SS5)shows the ordering inverts without this control\.
- •*Controls are projected ontop=64p=64leading principal directions\.*Takingp=dp=dgives Qwen3\-8B\-Base about1\.51\.5samples per parameter, at which point the fit interpolates and the spectrum estimates nothing\.

That projection is the regime[Theorem˜10\.1](https://arxiv.org/html/2608.10172#S10.Thmtheorem1)assumes, and it is confirmed here: the attention\-write covariance of Qwen3\-8B\-Base at the analysis layer has effective ranktr⁡\(Gu\)/λmax​\(Gu\)=10\.8\\operatorname\{tr\}\(G\_\{u\}\)/\\lambda\_\{\\max\}\(G\_\{u\}\)=10\.8against40964096ambient dimensions \- direct empirical support for the low\-effective\-rank premise of[Corollary˜10\.6](https://arxiv.org/html/2608.10172#S10.Thmtheorem6), measured on the largest model in the suite\.

The random\-Fourier block of the spectral dictionary has its bandwidth fixed at a value appropriate ford≤2304d\\leq 2304and was deliberately not retuned ford=4096d=4096\. It becomes unstable in the deepest Qwen3\-8B\-Base layers, which affects the split\-half stability estimate at layer 26 \([Section˜11\.5](https://arxiv.org/html/2608.10172#S11.SS5)\) but not the convergence analysis, which is run at layer 18 where the construction is well behaved\. We did not tune it, because selecting a dictionary hyperparameter per model until the predicted exponent appears is precisely the analyst\-dependence that[Corollary˜9\.5](https://arxiv.org/html/2608.10172#S9.Thmtheorem5)identifies as the defect of SAE dictionaries: a result obtained that way would illustrate the problem this paper is about rather than support its claim\. A principled alternative, which we set out but do not exercise, is to select the bandwidth on a*held\-out layer*using the invariance residual as the criterion, so that selection and evaluation use different quantities on different data\. Guarding such a search against degenerate solutions is essential: as the bandwidth vanishes the Fourier features approach constants, a near\-constant dictionary is fitted perfectly byA≈IA\\approx I, andε^rel\\hat\{\\varepsilon\}\_\{\\mathrm\{rel\}\}collapses to zero while the dictionary encodes nothing\.

### 11\.2Convergence and the pre\-asymptotic regime

There is no analytic ground\-truthAAon a pretrained model, so we measure*self\-consistency*: splitMMsamples into disjoint halves, fitA^\(1\)\\hat\{A\}^\{\(1\)\}andA^\(2\)\\hat\{A\}^\{\(2\)\}, and record the Hungarian\-matched spectral distance\. This statistic inherits the rate of[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)\- it is exactly the left\-hand side of the gap\-free bound \([53](https://arxiv.org/html/2608.10172#S6.E53)\) \- so the log\-log slope estimates the exponent without ground truth\.

Two protocol choices matter and are not cosmetic\. Halves are drawn from*disjoint source sequences*, since random row\-splitting lets one document appear on both sides and biases the distance downward, worst at largeMM\. And all uncertainty is a sequence\-level block bootstrap\. We reportMMas row count throughout\.

Table 2:Split\-half spectral convergence versusMMatN=32N=32, mid\-depth analysis layer\. Qwen3\-8B\-Base attains the predicted−1/2\-1/2exponent*below*the gap\-free threshold \([54](https://arxiv.org/html/2608.10172#S6.E54)\); GPT\-2’s random dictionary does not attain it despite exceeding its own threshold threefold\. Error bars are sequence\-level block bootstraps\.![Refer to caption](https://arxiv.org/html/2608.10172v1/x2.png)Figure 3:Split\-half spectral distance againstMM, log\-log, atN=32N=32and mid\-depth\. The distance falls monotonically on every model: the spectrum*is*being recovered\. Qwen3\-8B\-Base attains the predictedM−1/2M^\{\-1/2\}rate \(−0\.506±0\.031\-0\.506\\pm 0\.031\); GPT\-2 small is shallower \(−0\.286\-0\.286\)\. The dashed line is theM−1/2M^\{\-1/2\}reference\. Shaded bands are sequence\-level block bootstraps\.##### The rate on Qwen3\-8B\-Base, and two caveats\.

The distance falls monotonically withMMon all three models \([Figure˜3](https://arxiv.org/html/2608.10172#S11.F3)\)\. On GPT\-2 small atMmax=399,974M\_\{\\max\}=399\{,\}974the exponents are−0\.286±0\.007\-0\.286\\pm 0\.007\(spectral\) and−0\.290±0\.011\-0\.290\\pm 0\.011\(random\)\. On Qwen3\-8B\-Base at layer 18 the spectral dictionary converges at−0\.506±0\.031\-0\.506\\pm 0\.031\- the value[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)predicts, within one standard error\. To our knowledge this is the first observation of the parametric rate for a Koopman spectrum on a pretrained transformer\.

We hold ourselves to two caveats\. First, that is the exponent of the*maximum*matched error, which is the quantity \([36](https://arxiv.org/html/2608.10172#S6.E36)\) bounds, and on this model the robust summaries disagree:−0\.338\-0\.338\(median\) and−0\.375\-0\.375\(mean\)\. On GPT\-2 small the three agree closely \(−0\.286\-0\.286,−0\.292\-0\.292,−0\.288\-0\.288\), so the divergence indicates that Qwen’s maximum is driven by a subset of eigenvalues converging faster than the bulk, not that the whole spectrum attains−1/2\-1/2\. We report the theorem’s own statistic as primary and do not average the others away\. Second, Qwen reaches the rate*without*crossing its gap\-free threshold \(Mmax/M0eig=0\.23M\_\{\\max\}/M\_\{0\}^\{\\mathrm\{eig\}\}=0\.23\), while GPT\-2’s random dictionary crosses its own by a factor of3\.033\.03and returns only−0\.290\-0\.290\. CrossingM0eigM\_\{0\}^\{\\mathrm\{eig\}\}is neither necessary nor sufficient at these sample sizes\. What the split \([54](https://arxiv.org/html/2608.10172#S6.E54)\)–\([55](https://arxiv.org/html/2608.10172#S6.E55)\) buys is testability \- nine orders of magnitude on GPT\-2 small, fromM0vec=6\.4×1016M\_\{0\}^\{\\mathrm\{vec\}\}=6\.4\\times 10^\{16\}toM0eig=1\.0×107M\_\{0\}^\{\\mathrm\{eig\}\}=1\.0\\times 10^\{7\}\- not the asymptotic exponent itself\.

##### Dictionary size: a partially confirmed prediction\.

The gap\-free bound \([56](https://arxiv.org/html/2608.10172#S6.E56)\) predicts that the optimal\-matching distance degrades withNNthrough its\(2​N−1\)\(2N\-1\)factor, and it does\. AtN∈\{8,16,32,64,128\}N\\in\\\{8,16,32,64,128\\\}on GPT\-2 small the exponents are−0\.504\-0\.504,−0\.369\-0\.369,−0\.342\-0\.342,−0\.281\-0\.281,−0\.239\-0\.239\- closest to−1/2\-1/2at the smallest dictionary \- and the level grows asN\+0\.36N^\{\+0\.36\}\. Gemma\-2\-2B behaves the same way \(−0\.477→−0\.242\-0\.477\\to\-0\.242, levelN\+0\.29N^\{\+0\.29\}\)\. The direction is confirmed and the bound is loose: the worst\-case Elsner constant predicts an exponent of\+1\+1for the level\. A prediction that is real, predicted, and smaller than allowed\.

##### Four candidate explanations for the shortfall, and what survives\.

On GPT\-2 small and Gemma\-2\-2B the exponent sits well above−1/2\-1/2\. We tested four explanations and excluded three\.

*\(i\) The reduction choice\.*Max, median and mean matched errors agree closely on GPT\-2 small, so the shortfall is not an artefact of reporting the maximum\.

*\(ii\) Within\-document correlation\.*Varying the number of rows harvested per document across\{1,2,4,8\}\\\{1,2,4,8\\\}leaves the exponent unchanged \(−0\.245\-0\.245,−0\.265\-0\.265,−0\.243\-0\.243,−0\.294\-0\.294\), ruling out sample dependence within a sequence as the explanation\.

*\(iii\) Heavy tails\.*[Theorem˜7\.2](https://arxiv.org/html/2608.10172#S7.Thmtheorem2)nominates heavy tails as a candidate \- residual streams are documented to carry them\[[15](https://arxiv.org/html/2608.10172#bib.bib88)\]\- and names the deciding experiment\. We refitted88\(model, layer, dictionary\) cells with median\-of\-means Gramians overK=24K=24blocks, paired trial\-for\-trial against the plain estimator on identical sequence\-disjoint halves so that only the second moment differs\. The two agree to within2%2\\%in every cell at everyMM: on GPT\-2 small at layer 6,−0\.359±0\.018\-0\.359\\pm 0\.018against−0\.368±0\.019\-0\.368\\pm 0\.019over the99grid points where a block holds enough rows for an\(N\+p\)\(N\+p\)\-square Gramian, with median ratio1\.0001\.000; block counts8,16,24,488,16,24,48span only0\.7%0\.7\\%\([Figure˜4](https://arxiv.org/html/2608.10172#S11.F4)\(a,b\)\)\.

The reason is that the antecedent of[Theorem˜7\.2](https://arxiv.org/html/2608.10172#S7.Thmtheorem2)mostly fails, and the measurement locates why: it is the*lifted*state, not the residual stream, that \(R3\) constrains\. The residual stream itself is mildly heavy\-tailed \(Hill indexα^=6\.90\\hat\{\\alpha\}=6\.90against a Gaussian null of7\.697\.69\), but the random\-Fourier block is bounded \(α^=3218\.18\\hat\{\\alpha\}=3218\.18, off scale\) and the whitened lifted state the estimator actually sees is indistinguishable from the null \(α^=7\.20\\hat\{\\alpha\}=7\.20, excess kurtosis0\.190\.19\); see[Figure˜4](https://arxiv.org/html/2608.10172#S11.F4)\(c\)\. The lifting removes the tails\. The one exception \- Qwen3\-8B\-Base at layer 18, excess kurtosis50\.750\.7\- is also the one cell where median\-of\-means beats the plain estimator \(−0\.349\-0\.349against−0\.289\-0\.289\), consistent with[Theorem˜7\.2](https://arxiv.org/html/2608.10172#S7.Thmtheorem2)though modestly\. On synthetic data with genuinely heavy design the separation appears as predicted \(contaminated design:−0\.286\-0\.286for median\-of\-means against−0\.405\-0\.405plain\), confirming that the null result on real data is a property of the dictionaries and not of the implementation\.

![Refer to caption](https://arxiv.org/html/2608.10172v1/x3.png)Figure 4:The robust estimator changes almost nothing, and the tail diagnostics say why\.\(a\)Plain against median\-of\-means EDMDc on identical sequence\-disjoint halves, GPT\-2 small layer 6: the two curves coincide at everyMM\.\(b\)Their ratio,1\.0001\.000at the median and within\[0\.979,1\.035\]\[0\.979,1\.035\]throughout\.\(c\)Hill tail indexα^\\hat\{\\alpha\}along the lifting chain: the raw residual stream is mildly heavy, the linear block similar, the bounded random\-Fourier block has effectively no tail, and the lifted stateΨ\\Psiand the control sit at the Gaussian null\. Condition \(R3\) holds for the object it constrains, so[Theorem˜7\.2](https://arxiv.org/html/2608.10172#S7.Thmtheorem2)has nothing to repair here\.*\(iv\) The finite\-sample constant \- what survives\.*The exponent is not a function ofMMalone but ofM/M0eig​\(N\)M/M\_\{0\}^\{\\mathrm\{eig\}\}\(N\), the sample size relative to each cell’s own threshold\. Across3030\(model, dictionary,NN\) cells contributing300300rolling\-window exponents, the fitted exponent decreases monotonically with this ratio \(Spearmanρ=−0\.28\\rho=\-0\.28,p<10−4p\\,<10^\{\-4\}\), with binned medians falling from−0\.31\-0\.31well below threshold to−0\.53\-0\.53at the largest ratios and crossing−1/2\-1/2past the threshold \([Figure˜5](https://arxiv.org/html/2608.10172#S11.F5)\(a\)\)\. This is why Qwen reachesM−1/2M^\{\-1/2\}below its threshold while GPT\-2’s random dictionary does not above its own: the ratio, not the crossing, is what orders the cells\. The same collapse accounts for most of theNN\-dependence in[Figure˜5](https://arxiv.org/html/2608.10172#S11.F5)\(b\), sinceM0eigM\_\{0\}^\{\\mathrm\{eig\}\}grows withNN\.

Two limits of this diagnosis, stated rather than left implicit: the two Qwen3 series run*against*the collapse, growing steeper withNNwhere it predicts flattening, and are the residual it does not explain; andM0eigM\_\{0\}^\{\\mathrm\{eig\}\}’s unknown constantc0c\_\{0\}fixes the horizontal axis only up to a common shift, so the collapse shows that the exponents are one curve, not where the threshold sits on it\.

![Refer to caption](https://arxiv.org/html/2608.10172v1/x4.png)Figure 5:Most of the scatter in the convergence exponent is one curve\.\(a\)Each of3030\(model, dictionary,NN\) cells contributes rolling exponents, plotted againstlog10⁡\(M/M0eig​\(N\)\)\\log\_\{10\}\(M/M\_\{0\}^\{\\mathrm\{eig\}\}\(N\)\)\- its own threshold rather than rawMM\. Points are individual windows; the heavy line joins equal\-count binned medians, which fall from−0\.31\-0\.31to−0\.53\-0\.53and cross−1/2\-1/2past the threshold \(Spearmanρ=−0\.28\\rho=\-0\.28,p<10−4p\\,<10^\{\-4\}\)\.\(b\)The same cells againstNN\. Four of the six series degrade monotonically, which \(a\) accounts for sinceM0eigM\_\{0\}^\{\\mathrm\{eig\}\}grows withNN; the two Qwen3 series do not, and are the residual the collapse does not explain\.![Refer to caption](https://arxiv.org/html/2608.10172v1/x5.png)Figure 6:Diagnosing the exponent, all four panels\.\(a\)Plain against median\-of\-means EDMDc on identical sequence\-disjoint halves, GPT\-2 small layer 6: the two coincide at everyMM\(median ratio1\.0001\.000\), as they do in seven of the88cells measured\.\(b\)Why: the Hill tail index along the lifting chain, showing that \(R3\) holds for the lifted state even though the raw residual stream is mildly heavy\.\(c\)The threshold collapse of[Figure˜5](https://arxiv.org/html/2608.10172#S11.F5)\(a\)\.\(d\)The same cells against dictionary sizeNN\.
##### Reading the result\.

On the largest model the theorem’s own statistic attains the predicted exponent\. On GPT\-2 small atN=32N=32it sits near−0\.286\-0\.286, with three candidate explanations excluded and the fourth \- the finite\-sample constant \- supported by a monotone collapse\. The spectrum converges on every model; the predicted rate is observed on Qwen3\-8B\-Base for the quantity the theorem bounds; and the threshold governing that guarantee is nine orders of magnitude smaller than the unsplit statement suggested\.

### 11\.3Koopman modes and the IOI circuit

[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)guarantees recovery of the Koopman spectrum but says nothing about its semantic meaning\. We therefore evaluate the strongest available interpretation against the best\-characterised transformer circuit: indirect\-object identification \(IOI\) in GPT\-2 small\[[74](https://arxiv.org/html/2608.10172#bib.bib21)\]\. Modes live in observable space, so this test requires the fitted linear read\-out back into residual\-stream coordinates whose construction and measured quality are given in[Section˜11\.8\.4](https://arxiv.org/html/2608.10172#S11.SS8.SSS4); its median relative reconstruction error on held\-out states is0\.2310\.231, which is the precondition for any of the negative results below to mean anything\.

All measurements are at layer 9 with the fullN=128N=128basis, over4,0964\{,\}096prompts balanced across the1515templates of\[[74](https://arxiv.org/html/2608.10172#bib.bib21)\], against a baseline logit difference of3\.233\.23\.

*First, modes do not concentrate on the known head subspaces\.*No head class reaches an alignment of0\.80\.8with the modal basis: greedy selection exhausts its budget at96/12896/128modes for name movers, reaching only0\.2600\.260against a random\-subspace null of0\.1250\.125\. S\-inhibition heads reach0\.3110\.311and induction heads0\.2980\.298against the same null\. The excess over chance is real but small, and three\-quarters of the basis is not a circuit in any useful sense \([Figure˜7](https://arxiv.org/html/2608.10172#S11.F7)\(b\)\)\.

*Second, ablating modes moves behaviour less than ablating principal directions\.*Removing the span of the top\-jjattributing modes at layer 9 moves behaviour far more thanjjrandom directions, which never exceed7%7\\%, but is dominated at everyjjby PCA: atj=8j=8, principal directions remove60%60\\%of the logit difference against the modes’25%25\\%, and PCA’s advantage persists toj=32j=32\(67%67\\%against41%41\\%\)\. The modes beat random decisively but are a worse handle on this behaviour than principal components \([Figure˜7](https://arxiv.org/html/2608.10172#S11.F7)\(a\)\)\.

*Third, and sharpest, the selected mode set is not itself reproducible\.*Comparing the subspace spanned by the top\-16 attributing modes across independent calibration draws \- a comparison invariant to the permutation[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)quotients by \- gives a median agreement of only0\.2900\.290\. The spectrum converges \([Section˜11\.2](https://arxiv.org/html/2608.10172#S11.SS2)\) but the identity of “the circuit” one would nominate does not, at these sample sizes\. Identifiability of the eigenvalue multiset does not confer identifiability of a selected sub\-collection, and we do not claim that it does\.

![Refer to caption](https://arxiv.org/html/2608.10172v1/x6.png)Figure 7:Koopman modes carry behaviourally relevant structure but do not resolve the IOI circuit\.\(a\)Ablating the top\-jjmodes at layer 9 moves the IOI logit difference more thanjjrandom directions but less thanjjprincipal directions: atj=8j=8, PCA removes60%60\\%against the modes’25%25\\%\.\(b\)No head class reaches the0\.80\.8alignment target; greedy selection exhausts its budget at96/12896/128modes, reaching about twice the random\-subspace null \(0\.2600\.260against0\.1250\.125for name movers\)\.The measured eigenvector conditioning places these fits firmly in the non\-normal regime that[Theorem˜8\.1](https://arxiv.org/html/2608.10172#S8.Thmtheorem1)concerns:κ2​\(V^\)=78\.7\\kappa\_\{2\}\(\\hat\{V\}\)=78\.7on GPT\-2 small,38\.238\.2on Gemma\-2\-2B and494\.7494\.7on Qwen3\-8B\-Base\. By[Proposition˜8\.2](https://arxiv.org/html/2608.10172#S8.Thmtheorem2)and[Remark˜8\.3](https://arxiv.org/html/2608.10172#S8.Thmtheorem3), misalignment at these conditionings is already near\-saturated\. The Koopman spectrum captures behaviourally relevant structure but does not decompose computation into human\-legible mechanisms: it is an identifiable coarse invariant of transformer depth dynamics\.

### 11\.4Transport versus encoding

If the two bases differ because one indexes*encoding*and the other*transport*, then PCA’s margin should shrink with how far the question travels in depth\. We test this directly: predictxℓ\+kx\_\{\\ell\+k\}from a rank\-jjlinear read ofxℓx\_\{\\ell\}\- samejjdirections, same source state, same unconstrained map onto the target \- and sweep the depth gapkk\.

The design is chosen so that the margin cannot shrink for trivial reasons\. The read\-out does not saturate and has no privileged basis: PCA is optimal for reconstructingxℓx\_\{\\ell\}itself, which stops being the target oncek\>0k\>0\. Each cell is normalised between the optimal rank\-jjpredictor \(0\) and a random subspace \(11\), so the margin cannot shrink merely because prediction gets harder withkk\.

Pooled over66layers andj∈\{8,…,64\}j\\in\\\{8,\\dots,64\\\}\(156156cells,8,1928\{,\}192prompts\), PCA’s advantage falls from0\.480\.48atk=1k=1to0\.120\.12atk=9k=9, a4\.1×4\.1\\timesdecay, and it does so monotonically at every layer separately \(raw ratio1\.69×→1\.05×1\.69\\times\\to 1\.05\\times;[Figure˜8](https://arxiv.org/html/2608.10172#S11.F8)\)\. PCA wins the question it is optimal for, by a margin that shrinks with depth\-distance and is essentially gone byk≈8k\\approx 8\.

Stated plainly: the advantage decays to parity, it does not reverse\. We have not shown that Koopman modes are the better transport basis, only that PCA stops being one\. Two causal tests of a reversal do not discriminate on this testbed\. Projecting out the mode at every layerℓ′≥ℓ\\ell^\{\\prime\}\\geq\\ellso that it cannot carry forward leaves the ordering unchanged at the analysis layer \(27%27\\%against34%34\\%removed atj=32j=32, versus34%34\\%/44%44\\%locally\); at earlier layers it removes the whole effect regardless of basis, and the logit\-lens readout atℓ\+k\\ell\+ksaturates the same way\. Twelve layers is not enough depth to separatekkfromℓ\\ellunder a behavioural readout \- a limit of this testbed, not of the bases, which the rank\-jjprediction avoids precisely by having no readout to saturate\.

![Refer to caption](https://arxiv.org/html/2608.10172v1/x7.png)Figure 8:The PCA advantage is local and decays with depth\.Predictingxℓ\+kx\_\{\\ell\+k\}from a rank\-jjread ofxℓx\_\{\\ell\}\(66layers,j∈\{8,…,64\}j\\in\\\{8,\\dots,64\\\},8,1928\{,\}192prompts\), normalised between the optimal rank\-jjpredictor \(0\) and a random subspace \(11\)\.\(a\)Normalised excess fraction of variance unexplained against the depth gapkk\.\(b\)The PCA–Koopman gap falls from0\.480\.48atk=1k=1to0\.120\.12atk=9k=9, a4\.1×4\.1\\timesdecay to parity\.
### 11\.5The SAE invariance gap, and the selection rule it is conditional on

[Corollary˜9\.5](https://arxiv.org/html/2608.10172#S9.Thmtheorem5)rests on a measurable claim: dictionaries failing \(DR2\) carry a large projection residual\. We report the relative invariance residual

ε^rel=‖Y−A^M​X−B^M​Ξ‖F‖Y‖F,\\hat\{\\varepsilon\}\_\{\\mathrm\{rel\}\}\\;=\\;\\frac\{\\bigl\\\|Y\-\\hat\{A\}\_\{M\}X\-\\hat\{B\}\_\{M\}\\Xi\\bigr\\\|\_\{F\}\}\{\\\|Y\\\|\_\{F\}\},\(77\)the empirical counterpart of the projection residual in[Definition˜4\.6](https://arxiv.org/html/2608.10172#S4.Thmtheorem6), which vanishes exactly when \(DR2\) holds\.

Table 3:SAE dictionaries sit furthest from Koopman invariance on every model \- further than a random orthonormal control\.Relative invariance residual \([77](https://arxiv.org/html/2608.10172#S11.E77)\) at matchedN=32N=32, mid\-depth analysis layer, all dictionaries whitened \(without which the ordering inverts; see text\)\. Lower is closer to satisfying \(DR2\)\. Conditional on the selection rule: features are chosen by activation frequency throughout, and the ordering*reverses*under variance\-based selection \(0\.70×0\.70\\timesrather than2\.05×2\.05\\timesthe spectral residual\)\.[Table˜3](https://arxiv.org/html/2608.10172#S11.T3)shows the predicted ordering on every model, strengthening with scale \(2\.1×2\.1\\times,3\.1×3\.1\\times,3\.4×3\.4\\timesthe spectral residual\), across three SAE families \- ReLU\[[7](https://arxiv.org/html/2608.10172#bib.bib1)\], JumpReLU\[[46](https://arxiv.org/html/2608.10172#bib.bib7)\], TopK\[[23](https://arxiv.org/html/2608.10172#bib.bib6)\]\- and at every layer examined \([Figure˜9](https://arxiv.org/html/2608.10172#S11.F9)\(a\)\)\. That SAE dictionaries sit*further*from invariance than a random orthonormal basis, which has no mechanism by which it could be𝒦\\mathcal\{K\}\-invariant, suggests the SAE objective does not merely fail to enforce \(DR2\) but selects against it: sparsity concentrates each feature on few inputs, while closure requires the span to absorb where those inputs are carried next\. That is a stronger reading than the corollary asserts, and we flag it as such\.

##### Instability, and one reversal\.

The residual measures the condition that fails, not the failure\.[Figure˜9](https://arxiv.org/html/2608.10172#S11.F9)\(b\) measures the failure: spectra obtained through SAE dictionaries are less reproducible across disjoint calibration draws than spectral ones in88of99layers, by factors of1\.11\.1to5\.75\.7\. This is the seed\-to\-seed variability the SAE literature reports\[[37](https://arxiv.org/html/2608.10172#bib.bib10),[6](https://arxiv.org/html/2608.10172#bib.bib9)\], seen here on a*fixed*dictionary with only the calibration sample varying \- so it cannot be optimisation noise\.

The ninth layer is worth stating plainly\. At layer 26 of Qwen3\-8B\-Base the ordering reverses, and it does so because the spectral dictionary degrades rather than because the SAE improves: its instability is0\.1450\.145there against0\.0170\.017–0\.0560\.056everywhere else, while the SAE sits at0\.0290\.029, in line with its own range\. This is the deepest layer of the largest model, and it is the same breakdown of the random\-Fourier construction atd=4096d=4096noted in[Section˜11\.1](https://arxiv.org/html/2608.10172#S11.SS1); the convergence measurement is taken at layer 18, where the construction is well behaved\. The invariance residual, which does not depend on the stability estimate, keeps the predicted ordering at this layer as at all others\.

##### A confound that inverts the result\.

Dictionaries must be compared at matched conditioning, not merely matchedNN\. An SAE encodes mostly zeros atN=32N=32\- activation rate0\.0890\.089for GPT\-2 against0\.5040\.504and0\.5030\.503for the spectral and random dictionaries \- and a mostly\-zero target is trivially predictable, soε^rel\\hat\{\\varepsilon\}\_\{\\mathrm\{rel\}\}would reward sparsity rather than invariance\. Measured without this control, GPT\-2’s SAE residual is0\.2870\.287,*below*both the spectral \(0\.3010\.301\) and random \(0\.3410\.341\) dictionaries \- apparently contradicting[Corollary˜9\.5](https://arxiv.org/html/2608.10172#S9.Thmtheorem5)\. Whitening removes the artefact: a change of basis leavesσ​\(A\)\\sigma\(A\)unchanged \([Theorem˜5\.1](https://arxiv.org/html/2608.10172#S5.Thmtheorem1)\(b\)\) while equalisingηΨ\\eta\_\{\\Psi\}across families \(0\.710\.71–0\.800\.80on GPT\-2 after whitening, against7\.4×10−37\.4\\times 10^\{\-3\}unwhitened\)\. All figures in[Table˜3](https://arxiv.org/html/2608.10172#S11.T3)are post\-whitening; we report the uncorrected numbers because a reader reproducing this without the control will obtain them\.

![Refer to caption](https://arxiv.org/html/2608.10172v1/x8.png)Figure 9:The \(DR2\) gap and the non\-identifiability it produces\.\(a\)Distance from Koopman invariance,ε^rel\\hat\{\\varepsilon\}\_\{\\mathrm\{rel\}\}of \([77](https://arxiv.org/html/2608.10172#S11.E77)\), at matchedN=32N=32with all dictionaries whitened, at three layers per model\. SAE dictionaries sit furthest from \(DR2\) on every model and layer \- further than a random orthonormal basis\.\(b\)The consequence[Corollary˜9\.5](https://arxiv.org/html/2608.10172#S9.Thmtheorem5)is about: matched spectral distance between realisations fitted on disjoint calibration draws \(log scale; lower is more identifiable\)\. Spectra from SAE dictionaries move more between draws in88of99layers; the exception is Qwen3\-8B\-Base layer 26, where the*spectral*dictionary is itself unstable\.
##### The pre\-registered criterion fails, and the reason is the selection rule\.

A public SAE is1616k–6565k wide and this comparison runs atN=32N=32, so some rule must select roughly one latent in a thousand\. Before running the sweep we registered the criterion that the orderingε^rel​\(SAE\)\>ε^rel​\(Random\)\>ε^rel​\(Spectral\)\\hat\{\\varepsilon\}\_\{\\mathrm\{rel\}\}\(\\text\{SAE\}\)\>\\hat\{\\varepsilon\}\_\{\\mathrm\{rel\}\}\(\\text\{Random\}\)\>\\hat\{\\varepsilon\}\_\{\\mathrm\{rel\}\}\(\\text\{Spectral\}\)should hold with non\-overlapping interquartile ranges in at least80%80\\%of \(model, layer,NN, rule\) cells\. It does not: the criterion is met in49%49\\%of7575cells\.

The cause is the selection rule, and the effect is one of sign, not degree\. Ranking latents by activation frequency \(the rule behind[Table˜3](https://arxiv.org/html/2608.10172#S11.T3)\) or drawing uniformly among live latents puts the SAE above the spectral dictionary in every cell, by median factors of2\.052\.05and1\.701\.70; ranking by mean activation magnitude does so in1212of1515cells \(1\.361\.36\)\. But ranking by activation*variance*, or greedily for variance explained, reverses the ordering in every cell, at0\.700\.70and0\.630\.63times the spectral residual\. Across widthsN∈\{8,…,128\}N\\in\\\{8,\\dots,128\\\}and three layers, the sign tracks the selection rule and nothing else \([Figure˜10](https://arxiv.org/html/2608.10172#S11.F10)\)\.

The reading is therefore narrower than[Table˜3](https://arxiv.org/html/2608.10172#S11.T3)alone would support\. The \(DR2\) gap is a real property of the subset a practitioner is most likely to read \- the features that fire often \- robust across models, layers, widths and all four encoder conventions \([Section˜11\.8\.5](https://arxiv.org/html/2608.10172#S11.SS8.SSS5)\) for that subset\. It is*not*a property of the SAE’s span as such: a variance\-optimal subset is closer to Koopman invariance than our spectral construction, which is unsurprising in hindsight, since selecting for variance explained recovers something close to a principal\-component basis and the spectral dictionary’s linear block is exactly that\.[Corollary˜9\.5](https://arxiv.org/html/2608.10172#S9.Thmtheorem5)concerns what the objective fails to enforce and remains correct; what these measurements add is that the failure is unevenly distributed, so any claim about “the” invariance residual of an SAE must state its selection rule\. Ours is activation frequency throughout\.

![Refer to caption](https://arxiv.org/html/2608.10172v1/x9.png)Figure 10:The whole effect is conditional on the selection rule\. Ratio of SAE to spectral invariance residual as a function of howNNfeatures are selected from a1616k–6565k\-wide SAE, pooled over three models, three layers and five widths; bars are medians with interquartile ranges and the dashed line is parity\. Selecting the most frequently active latents \- the rule behind[Table˜3](https://arxiv.org/html/2608.10172#S11.T3)\- puts the SAE at2\.05×2\.05\\timesthe spectral residual; selecting for activation*variance*puts it at0\.70×0\.70\\times, reversing the ordering the corollary predicts\.

### 11\.6Enforcing invariance: the penalty sweep

[Corollary˜9\.5](https://arxiv.org/html/2608.10172#S9.Thmtheorem5)makes a causal claim: SAE non\-identifiability is*caused*by the absence of a Koopman\-invariance constraint\.[Section˜11\.5](https://arxiv.org/html/2608.10172#S11.SS5)measures the correlate; here we add the invariance penalty of \([66](https://arxiv.org/html/2608.10172#S9.E66)\) and ask whether the outcome moves\.

We train SAEs on cached GPT\-2 layer\-8 transitions at width4​d=30724d=3072, sweepingγ∈\{0,10−3,10−2,10−1,1,10\}\\gamma\\in\\\{0,10^\{\-3\},10^\{\-2\},10^\{\-1\},1,10\\\}with three seeds each, screening every run for the penalty’s degenerate minimisers and reporting excluded runs as excluded\. TopK is primary because it fixesL0L\_\{0\}by construction: sweepingγ\\gammaat fixed sparsity is what makes any movement attributable to the invariance term rather than to the dictionary quietly becoming denser\.

Raisingγ\\gammafrom0to11reduces the invariance residual from0\.4380\.438to0\.3330\.333\(24%24\\%\), confirming that the penalty does what it is designed to do, and reduces the split\-half spectral distance \- identifiability itself, not a proxy for it \- from0\.01700\.0170to0\.01010\.0101\(41%41\\%\) at matchedL0=32L\_\{0\}=32\. Intervening on the mechanism moves the outcome it predicts, which is the strongest support the corollary admits\.

The cost appears in the same runs: reconstruction degrades \(FVU0\.222→0\.2550\.222\\to 0\.255\) and the live\-feature fraction falls from0\.770\.77to0\.580\.58; atγ=10\\gamma=10the dictionary collapses outright, all three seeds below the alive\-feature guard at4%4\\%, which locates the usable range\.

Cross\-seed feature agreement moves the wrong way: mean max cosine similarity between dictionaries from different seeds falls from0\.5260\.526to0\.4280\.428\.The penalty makes the recovered*spectrum*more reproducible while making the recovered*feature dictionary*less so\- a second instance of the paper’s central dissociation, arriving from a different direction\.

It is also architecture\-sensitive\. On ReLU the invariance residual at the guard\-selected operating point is not below baseline \(0\.413→0\.4350\.413\\to 0\.435\), though it does decline at largerγ\\gammaapproaching collapse, while the split\-half distance falls by22%22\\%\. The residual and the identifiability it proxies for can therefore decouple when sparsity is enforced by a penalty rather than by construction \- which is why TopK is primary\. This does not undermine the corollary, whose subject is the spectrum, but the penalty is not a drop\-in remedy for the feature\-level variability the SAE literature reports and we do not present it as one\.

![Refer to caption](https://arxiv.org/html/2608.10172v1/x10.png)Figure 11:Adding the invariance penalty improves spectral identifiability\.\(a\)The invariance residual falls withγ\\gammafor TopK; for ReLU it is not below for TopK; for ReLU it is not below baseline at the guard\-selected operating point, which the text takes up\. at the guard\-selected operating point, which the text takes up\.\(b\)The split\-half spectral distance \- identifiability itself rather than a proxy \- falls on both architectures, by41%41\\%\(TopK\) and22%22\\%\(ReLU\) at matchedL0L\_\{0\}\. Sparsity is held fixed across the sweep \(L0=32L\_\{0\}=32by construction for TopK,λ\\lambdacalibrated for ReLU\), so the improvement cannot be attributed to the dictionary becoming denser\. Crosses mark configurations excluded by the degeneracy guards: TopK collapses atγ=10\\gamma=10with4%4\\%of features alive, which locates the usable range\.We claim no more than this supports: not that invariance\-regularised SAEs are the right tool for practice, which would need an evaluation we do not attempt, but that the mechanism[Corollary˜9\.5](https://arxiv.org/html/2608.10172#S9.Thmtheorem5)names is the operative one for spectral identifiability, because intervening on it moves that outcome by a stated amount\.

### 11\.7Running the universality criterion

[Corollary˜9\.6](https://arxiv.org/html/2608.10172#S9.Thmtheorem6)is a decision procedure, and it had no experiment\. Running it requires three things matched, none of them free: dictionary size, relative depthℓ/L\\ell/L\(the suite spans 12 to 36 layers, so an absolute index is not a common coordinate\), andMM\(the resolution floor falls withMM\)\. We fit atN=32N=32andM=99,993M=99\{,\}993on each model’s own dictionary, atℓ/L∈\{0\.25,0\.5,0\.75\}\\ell/L\\in\\\{0\.25,0\.5,0\.75\\\}, and compare every pair’s matched distance against the larger of the two models’ split\-half floors, measured on the same draws \(0\.00730\.0073–0\.04120\.0412\)\.

Because “these spectra are distinct” and “this test separates everything” produce the same table, we added two Pythia\-160M seed replicas \- same architecture, same data, same order, different initialisation \- as models the criterion*should*decline to separate\.

It declares every pair distinct:3636of3636cross\-family pairs, and also99of99seed\-replica pairs, the closest of which still sits at4\.64×4\.64\\timesits floor \([Figure˜12](https://arxiv.org/html/2608.10172#S11.F12)\)\. A rule that separates two runs of the same architecture on the same data is not a universality criterion, and reporting the cross\-family half alone would have read as a clean success\.

The cause is locatable, and it is not the theorem\. The split\-half floor bounds the*sampling*error of a fixed dictionary, which is whatO​\(M−1/2\)O\(M^\{\-1/2\}\)governs\. A cross\-model comparison carries a second error the floor cannot see: the dictionaries are fitted per model and are only approximately𝒦\\mathcal\{K\}\-invariant, so what is compared are the spectra of two*realisations*rather than of two operators, and that mismatch is a bias that does not shrink withMM\([Remark˜6\.12](https://arxiv.org/html/2608.10172#S6.Thmtheorem12)\)\.[Corollary˜9\.6](https://arxiv.org/html/2608.10172#S9.Thmtheorem6)needs a calibrated null, not an estimation bound\.

What survives is an ordering, and it names the fix\. Distance ranks seed replicas below cross\-family pairs with AUC0\.750\.75on the maximum statistic that[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)bounds and0\.840\.84on the robust Wasserstein summary, rising to0\.970\.97at mid\-depth, where seed pairs sit at a median3\.9×3\.9\\timesthe floor against9\.2×9\.2\\timesfor cross\-family pairs \(13\.713\.7and23\.023\.0on the max statistic\)\. The spectrum separates architecture from seed relatively but not absolutely, so the operational form of the criterion should calibrate against seed replicas of the models being compared \- a concrete correction to a corollary we stated, obtained only by running the control\.

![Refer to caption](https://arxiv.org/html/2608.10172v1/x11.png)Figure 12:The universality criterion separates everything, including the control it should not\. Matched spectral distance for every model pair, in units of the split\-half resolution floor measured on the same draws, at three relative depths\. Blue bars are cross\-family pairs, green bars are the Pythia\-160M seed replicas that differ only in initialisation\. Every bar lies to the right of the dashed floor at11, so the criterion in its stated form declares all3636cross\-family pairs*and*all99seed\-replica pairs distinct\. The ordering is nonetheless informative \- seed pairs sit lower on average \(AUC0\.840\.84on the Wasserstein summary\)
### 11\.8The modelling conventions, measured

[Section˜4\.6](https://arxiv.org/html/2608.10172#S4.SS6)listed three conventions that define the estimand\. This subsection reports the measurements behind each, plus the read\-out and encoder\-convention checks that the earlier experiments depend on\.

#### 11\.8\.1Depth homogeneity: per\-layer against pooled fits

The realisation is fitted separately at each analysed layer throughout\. Two measurements support that choice\.

*Adjacent\-layer spectral drift\.*For each model we fitA^ℓ\\hat\{A\}\_\{\\ell\}at every cached layer and compute the Hungarian\-matched distance between adjacent layers, in units of the split\-half resolution floor measured at the sameMMon the same layer\. The median ratio is3\.223\.22on GPT\-2 small,11\.9711\.97on Gemma\-2\-2B and6\.126\.12on Qwen3\-8B\-Base\. A ratio above one means adjacent layers differ by more than the measurement can attribute to sampling, so the depth\-indexed familyℓ↦σ​\(Aℓ\)\\ell\\mapsto\\sigma\(A\_\{\\ell\}\)is resolved rather than inferred\.

*Pooled fits predict worse\.*Fitting oneAAon transitions pooled across depth and evaluating on held\-out transitions gives a relative residual of0\.4330\.433on GPT\-2 small against0\.2630\.263for per\-layer fits, a penalty of39%39\\%; the corresponding penalties are22%22\\%\(Gemma\-2\-2B,0\.4170\.417against0\.3260\.326\) and16%16\\%\(Qwen3\-8B\-Base,0\.2910\.291against0\.2440\.244\)\. A pooled realisation estimates a depth\-averaged object whose spectrum is an artefact of the pooling, and the size of the penalty bounds how large that artefact would be\.

[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)is stated and proved at fixedℓ\\elland needs no depth\-homogeneity, so nothing in the theory rests on this; what rests on it is the interpretation of the estimand\.

#### 11\.8\.2Exogeneity of the control, and the residualised realisation

[Definition˜4\.1](https://arxiv.org/html/2608.10172#S4.Thmtheorem1)treats the attention writeuℓu\_\{\\ell\}as an exogenous input, butuℓ=Attnℓ​\(RMSNorm​\(xℓ\(1:T\)\)\)u\_\{\\ell\}=\\mathrm\{Attn\}\_\{\\ell\}\(\\mathrm\{RMSNorm\}\(x\_\{\\ell\}^\{\(1:T\)\}\)\)is computed from the same residual stream\. This is a modelling convention, and it is not innocuous\.

*How far from exogenous\.*At mid\-depth the largest canonical correlation between the lifted stateΨ​\(xℓ\)\\Psi\(x\_\{\\ell\}\)and the control isρ1=0\.82\\rho\_\{1\}=0\.82on GPT\-2 small,0\.960\.96on Gemma\-2\-2B and0\.950\.95on Qwen3\-8B\-Base \- so some direction of the control is nearly a function of the state\. Overall, however, the controls are mostly not explained by the state: regressinguℓu\_\{\\ell\}onΨ​\(xℓ\)\\Psi\(x\_\{\\ell\}\)givesR2=0\.08R^\{2\}=0\.08,0\.220\.22and0\.160\.16respectively\. The dependence is concentrated in a few directions rather than spread across the control space\.

*The residualised realisation, and how far it moves\.*The exogenous\-by\-construction alternative removes the state\-explained part of the control before fitting \(Frisch–Waugh–Lovell\): regressuℓu\_\{\\ell\}onΨ​\(xℓ\)\\Psi\(x\_\{\\ell\}\), keep the residual, and fit the realisation against that\. The recovered spectrum moves by6\.8×6\.8\\timesthe split\-half floor on GPT\-2 small,9\.1×9\.1\\timeson Gemma\-2\-2B and5\.2×5\.2\\timeson Qwen3\-8B\-Base\. Those shifts are well above the resolution floor, so the convention selects between two genuinely different estimands at the precision we can measure \- it is not a numerical detail\.

We report the naive realisation as primary because it is the one[Definition˜4\.1](https://arxiv.org/html/2608.10172#S4.Thmtheorem1)defines, and the residualised realisation alongside it\. A reader who prefers the exogenous\-by\-construction estimand should read the latter throughout\. We regard this as the sharpest open modelling question in the framework and do not resolve it by fiat\.

#### 11\.8\.3Depth stationarity and the norm observable

The residual\-stream norm grows with depth, by a factor of4\.34\.3across the analysed range on GPT\-2 small,1\.51\.5on Gemma\-2\-2B and4\.44\.4on Qwen3\-8B\-Base\. A law whose second moment moves with depth cannot be a single invariant law, so the realisation is defined relative to the*layer marginal*μℓ\\mu\_\{\\ell\}rather than to one invariantμ\\mu\. The per\-layer fit of[Section˜11\.8\.1](https://arxiv.org/html/2608.10172#S11.SS8.SSS1)already enforces this; no separate correction is applied\.

*The growth appears as its own mode, which it need not have\.*If the framework is describing the depth dynamics rather than merely fitting them, the norm growth should be recoverable as a dynamical mode once the dictionary can express it\. Addinglog⁡‖xℓ‖\\log\\\|x\_\{\\ell\}\\\|as an observable does exactly that on every model tested: the largest expanding eigenvalue of the augmented realisation is1\.1371\.137\(GPT\-2 small\),1\.1071\.107\(Gemma\-2\-2B\) and1\.2911\.291\(Qwen3\-8B\-Base\), against measured per\-layer norm growth factors of1\.1111\.111,1\.0771\.077and1\.3261\.326respectively\. The agreement is within a few percent on all three, and it is a prediction that could have failed: nothing forces a fitted eigenvalue to match an independently measured growth rate\.

#### 11\.8\.4The full\-state read\-out

Koopman modes live in observable spaceℂN\\mathbb\{C\}^\{N\}\. Every question asked of them in[Section˜11\.3](https://arxiv.org/html/2608.10172#S11.SS3)\- alignment with a head’s write subspace, contribution to a logit difference, the effect of ablating them \- is a question about directions in the residual streamℝd\\mathbb\{R\}^\{d\}\. The correspondence between the two is fitted, not assumed, and its quality is reported because every mode\-level claim in that section is void if it is poor\.

*Construction\.*On held\-out states we fit the linear read\-out

x−μ≈Ψ​\(x\)​R,R∈ℂN×d,x\-\\mu\\;\\approx\\;\\Psi\(x\)\\,R,\\qquad R\\in\\mathbb\{C\}^\{N\\times d\},\(78\)by least squares with a small ridge for numerical hygiene \(the lifted features are whitened, so this is not regularisation in any statistical sense\)\. WritingΨ​\(x\)=∑kvk​φk​\(x\)\\Psi\(x\)=\\sum\_\{k\}v\_\{k\}\\varphi\_\{k\}\(x\)in the eigenbasis of the fittedA^\\hat\{A\}, withφk​\(x\)=ϕk⊤​Ψ​\(x\)\\varphi\_\{k\}\(x\)=\\phi\_\{k\}^\{\\top\}\\Psi\(x\)the matching left eigenfunctions, the*full\-state Koopman mode*is

ξk=R⊤​vk∈ℝd,x−μ≈∑kξk​φk​\(x\)\.\\xi\_\{k\}\\;=\\;R^\{\\top\}v\_\{k\}\\;\\in\\;\\mathbb\{R\}^\{d\},\\qquad x\-\\mu\\;\\approx\\;\\sum\_\{k\}\\xi\_\{k\}\\,\\varphi\_\{k\}\(x\)\.\(79\)Modes come in conjugate pairs for a real operator, so any subset that splits a pair leaves an imaginary residue; taking the real part is the correct projection back to the residual stream and is what the causal ablation applies\. Interventions are therefore always applied to whole conjugate groups\.

*Why fit the read\-out rather than invert the dictionary\.*Pushing an eigenvector back through the dictionary’s internal linear block would be cheaper, but it works only for dictionaries that have an invertible linear block\. Fitting \([78](https://arxiv.org/html/2608.10172#S11.E78)\) applies to*any*dictionary, including SAE encoders, which is what makes it possible to run the same circuit analysis through an SAE dictionary and compare\. It is also honest about what it claims: the reconstruction error is measured rather than inherited from an assumption about the dictionary’s structure\.

*Measured quality\.*On the IOI activations of[Section˜11\.3](https://arxiv.org/html/2608.10172#S11.SS3)the read\-out attains a median relative reconstruction error of0\.2310\.231on held\-out states\. This is the precondition for that section’s negative results to mean anything: a broken read\-out would produce the same appearance of modes failing to resolve the circuit, and could not be told apart from the finding\.

#### 11\.8\.5Encoder variants for the SAE dictionary

[Corollary˜9\.5](https://arxiv.org/html/2608.10172#S9.Thmtheorem5)is a statement about the span of a learned dictionary, and which map from activations to codes one calls “the dictionary” decides what is measured\. The main text fixes the convention \- the post\-activation codeΨD​\(x\)=σ​\(Wenc​x\+benc\)\\Psi\_\{D\}\(x\)=\\sigma\(W\_\{\\mathrm\{enc\}\}x\+b\_\{\\mathrm\{enc\}\}\)with the autoencoder’s own nonlinearity \- and we report the other three here so the choice is visible rather than implicit\. All four are measured on GPT\-2 small at layer 6 withN=32N=32, whitened, on the same activations, with the spectral and random dictionaries as reference\.

Table 4:Relative invariance residualε^rel\\hat\{\\varepsilon\}\_\{\\mathrm\{rel\}\}under four encoder conventions, GPT\-2 small at layer 6,N=32N=32\.The ordering that[Table˜3](https://arxiv.org/html/2608.10172#S11.T3)reports holds under all four conventions: the SAE\-derived dictionary carries the largest residual in every case\. The pre\-activation variant is reported only as a reference floor and is*not*evidence for the corollary, because a linear map of the state cannot be𝒦\\mathcal\{K\}\-invariant for nonlinearFF\- \(DR2\) fails by construction there, so that row could not have come out any other way\. The post\-activation code is genuinely nonlinear and is therefore not excluded a priori, which is what makes its failure a measurement\.

### 11\.9Per\-model report: Qwen3\-8B\-Base

Because the largest model in the suite is the one that attains the predicted rate, its measurements are worth collecting in one place\. All are taken at layer 18 of3636, withN=32N=32,p=64p=64, and no hyperparameter retuned for this model\.

- •*Convergence\.*Spectral dictionary−0\.506±0\.031\-0\.506\\pm 0\.031, random dictionary−0\.406±0\.030\-0\.406\\pm 0\.030, atMmax=106,660M\_\{\\max\}=106\{,\}660\- the former within one standard error of the predicted−1/2\-1/2, and attained atMmax/M0eig=0\.23M\_\{\\max\}/M\_\{0\}^\{\\mathrm\{eig\}\}=0\.23, i\.e\. below the gap\-free threshold\.
- •*Invariance residual\.*Spectral0\.1540\.154, random0\.3820\.382, SAE0\.5240\.524\- the largest SAE\-to\-spectral ratio in the suite at3\.4×3\.4\\times, consistent with the gap strengthening with scale\.
- •*Conditioning\.*κ2​\(V^\)=494\.7\\kappa\_\{2\}\(\\hat\{V\}\)=494\.7, by far the most non\-normal fit in the suite and the one for which[Remark˜8\.3](https://arxiv.org/html/2608.10172#S8.Thmtheorem3)’s saturation calculation is calibrated\.
- •*Control rank\.*The attention\-write covariance has effective rank 10\.8 against an ambient dimension of40964096\. This measurement is independent of the dictionary and is the direct empirical support for the premise of[Corollary˜10\.6](https://arxiv.org/html/2608.10172#S10.Thmtheorem6)\.
- •*Norm mode\.*Largest expanding eigenvalue1\.2911\.291against measured per\-layer norm growth1\.3261\.326\.
- •*Where it breaks\.*At layer 26 the random\-Fourier construction atd=4096d=4096degrades and the split\-half stability estimate becomes unreliable \(0\.1450\.145against0\.0170\.017–0\.0560\.056elsewhere\)\. This is a dictionary\-construction failure, not a model or estimator failure: the random dictionary run through the same pipeline on the same cached activations behaves normally, and the invariance residual, which does not depend on the stability estimate, keeps the predicted ordering at that layer\.

## 12Discussion

Lifting the depth recurrence through the Koopman operator gives a realisation whose spectrum is a coordinate\-free invariant of the transformer \([Theorem˜5\.1](https://arxiv.org/html/2608.10172#S5.Thmtheorem1)\), identifiable fromMMcalibration samples at the rateM−1/2M^\{\-1/2\}up to permutation \([Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)\)\.[Theorem˜7\.1](https://arxiv.org/html/2608.10172#S7.Thmtheorem1)shows no estimator does better inMM, and[Theorem˜7\.2](https://arxiv.org/html/2608.10172#S7.Thmtheorem2)covers heavy\-tailed activations\. The spectrum converges on every model tested and attains the predicted exponent on the largest\.

[Theorem˜8\.1](https://arxiv.org/html/2608.10172#S8.Thmtheorem1)says what this object is*not*\. A non\-normal realisation forces the activations’ principal directions apart from the Koopman modes, and the measured conditioning \(κ2​\(V^\)\\kappa\_\{2\}\(\\hat\{V\}\)from38\.238\.2to494\.7494\.7\) puts every fit in that regime\. The experiments land where the theorem requires: the modes beat random directions and lose to principal components at resolving IOI, with the loss confined to the question principal components are optimal for and decaying by4\.1×4\.1\\timesas the question moves away in depth\. What transports across depth is not the basis in which any one layer is encoded\. The spectrum therefore certifies an intrinsic, identifiable model property recoverable at a stated rate; it does not certify interpretability, behavioural coverage, or robustness\.

The limitations are as follows, in rough order of how much they constrain the claims\. First, the estimand depends on modelling conventions we cannot fully justify: residualising the attention write against the lifted state moves the spectrum by6\.86\.8–9\.1×9\.1\\timesthe resolution floor \([Section˜11\.8\.2](https://arxiv.org/html/2608.10172#S11.SS8.SSS2)\), so the naive and residualised conventions specify different estimands at the precision we can measure\. We report both\. Second, invariance holds only approximately\. All theorems assume[Assumption˜1](https://arxiv.org/html/2608.10172#Thmassumption1); real dictionaries satisfy it approximately and the resulting bias does not vanish asM→∞M\\to\\infty\([Remark˜6\.12](https://arxiv.org/html/2608.10172#S6.Thmtheorem12)\)\. KSA is thus not model\-agnostic \- the dictionary must be co\-designed with the architecture \- though within a fixed dictionary family it is uniform across corpus, seed and hyperparameter, which is the property SAEs lack\. Third, diagonalisability and separation are generic but not universal: Jordan extensions are routine but degrade the eigenvector rate \([Remark˜5\.6](https://arxiv.org/html/2608.10172#S5.Thmtheorem6)\), and[Assumption˜3](https://arxiv.org/html/2608.10172#Thmassumption3)can fail when mechanisms occupy closely spaced spectral scales\.[Corollary˜6\.11](https://arxiv.org/html/2608.10172#S6.Thmtheorem11)covers near\-degeneracy but not exact coincidence, which we conjecture is measure\-zero but have not proved\. Fourth, the intervention calculus is proved but not validated:[Theorem˜9\.2](https://arxiv.org/html/2608.10172#S9.Thmtheorem2)covers first\-order interventions only, and whether its closed\-form representatives predict measured patching effects is untested\. Fifth, the rate is asymptotic in a way that bites \- largeκ0\\kappa\_\{0\}or smallΔ\\Deltarequires much data before theM−1/2M^\{\-1/2\}regime begins, andc0c\_\{0\}in \([54](https://arxiv.org/html/2608.10172#S6.E54)\) is unknown, soM0eigM\_\{0\}^\{\\mathrm\{eig\}\}is an order\-of\-magnitude guide\. Sixth, the random\-Fourier bandwidth fixed ford≤2304d\\leq 2304degrades atd=4096d=4096in the deepest layers \([Section˜11\.9](https://arxiv.org/html/2608.10172#S11.SS9)\); we chose not to retune per model \([Section˜11\.1](https://arxiv.org/html/2608.10172#S11.SS1)\)\. Finally, two stated results fail their controls: the pre\-registered SAE\-gap criterion is met in49%49\\%of7575cells rather than the registered80%80\\%, because the sign of the effect tracks the feature\-selection rule \([Section˜11\.5](https://arxiv.org/html/2608.10172#S11.SS5)\); and the universality criterion separates two seed replicas of one architecture, so in its stated form it is not a universality criterion \([Section˜11\.7](https://arxiv.org/html/2608.10172#S11.SS7)\)\. We report both because both bound what may be claimed\.

Several problems remain open\. The upper bound carriesN\+log⁡\(1/δ\)\\sqrt\{N\+\\log\(1/\\delta\)\}while the lower bound is dimension\-free \([Theorem˜7\.1](https://arxiv.org/html/2608.10172#S7.Thmtheorem1)\); sharpening either settles the dimension dependence\. The penalty of \([66](https://arxiv.org/html/2608.10172#S9.E66)\) improves spectral identifiability by41%41\\%while degrading reconstruction and cross\-seed agreement \([Section˜11\.6](https://arxiv.org/html/2608.10172#S11.SS6)\), leaving open whether a dictionary can be sparse, reconstructive and𝒦\\mathcal\{K\}\-invariant at once\. A universality criterion needs a null calibrated against seed replicas rather than the sampling floor\. The exogeneity convention needs resolving, and the intervention calculus needs testing against measured patching effects\.

The safety\-case programme of\[[11](https://arxiv.org/html/2608.10172#bib.bib31)\]requires naming a mechanismYYresponsible for a behaviourXXand then defending that attribution\.[Theorem˜6\.1](https://arxiv.org/html/2608.10172#S6.Thmtheorem1)supplies the missing piece: a certificate thatYYis a property of the model rather than of the procedure that found it\. It does not address behavioural coverage, distribution shift, or adversarial robustness, and it does not deliver a human\-legible decomposition \-[Theorem˜8\.1](https://arxiv.org/html/2608.10172#S8.Thmtheorem1)says that in the non\-normal regime these models occupy, the identifiable object and the legible object cannot be the same object\. A safety case needing both will need two tools, and should say which one it is using where\.

## 13Conclusion

We have put mechanistic interpretability on an identifiability footing\. Treating a transformer forward pass as a controlled dynamical system in depth and lifting it with the Koopman operator produces a finite linear realisation whose spectrum is a coordinate\-free invariant of the model, identifiable from finite calibration data at the optimalM−1/2M^\{\-1/2\}rate up to permutation, with a matching lower bound, a heavy\-tailed variant, an algebraically complete intervention calculus, and an instance\-dependent reduction certificate\.

The measurements confirm the theory where it can be confirmed and bound it where it cannot\. The spectrum converges on GPT\-2 small, Gemma\-2\-2B and Qwen3\-8B\-Base, attaining the predicted exponent on the largest; the SAE invariance gap the theory predicts is observed on every model and is causally movable by the penalty the theory motivates; and the modes fail to resolve the IOI circuit in exactly the way a non\-normal realisation must\. The Koopman spectrum is an identifiable, model\-intrinsic fingerprint of transformer depth dynamics with a stated error bar, not a decomposition into human\-legible mechanisms\. Both halves of that sentence are results\.

## References

- \[1\]U\. M\. Al\-Saggaf and G\. F\. Franklin\(1988\)Model reduction via balanced realizations: an extension and frequency weighting techniques\.IEEE Transactions on Automatic Control33\(7\),pp\. 687–692\.Cited by:[§2](https://arxiv.org/html/2608.10172#S2.p7.1)\.
- \[2\]E\. Ameisen, J\. Lindsey, A\. Pearce, W\. Gurnee, N\. L\. Turner, B\. Chen, C\. Citro, D\. Abrahams, S\. Carter, B\. Hosmer, J\. Marcus, M\. Sklar, A\. Templeton, T\. Bricken, C\. McDougall, H\. Cunningham, T\. Henighan, A\. Jermyn, A\. Jones, A\. Persic, Z\. Qi, T\. Ben Thompson, S\. Zimmerman, K\. Rivoire, T\. Conerly, C\. Olah, and J\. Batson\(2025\)Circuit tracing: revealing computational graphs in language models\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2025/attribution-graphs/methods.html)Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p1.1),[§2](https://arxiv.org/html/2608.10172#S2.p1.1)\.
- \[3\]A\. C\. Antoulas\(2005\)Approximation of large\-scale dynamical systems\.SIAM\.Cited by:[§10\.1](https://arxiv.org/html/2608.10172#S10.SS1.p2.4),[§10\.2](https://arxiv.org/html/2608.10172#S10.SS2.1.p1.11),[§10\.2](https://arxiv.org/html/2608.10172#S10.SS2.2.p1.6),[Remark 10\.10](https://arxiv.org/html/2608.10172#S10.Thmtheorem10.p1.5),[§10](https://arxiv.org/html/2608.10172#S10.p1.5),[§2](https://arxiv.org/html/2608.10172#S2.p7.1)\.
- \[4\]O\. Azencot, N\. B\. Erichson, V\. Lin, and M\. W\. Mahoney\(2020\)Forecasting sequential data using consistent Koopman autoencoders\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§2](https://arxiv.org/html/2608.10172#S2.p5.1)\.
- \[5\]F\. L\. Bauer and C\. T\. Fike\(1960\)Norms and exclusion theorems\.Numerische Mathematik2\(1\),pp\. 137–141\.Cited by:[item 3](https://arxiv.org/html/2608.10172#S6.I2.i3.p1.2),[§6\.5](https://arxiv.org/html/2608.10172#S6.SS5.1.p1.14)\.
- \[6\]D\. Braun, J\. Taylor, N\. Goldowsky\-Dill, and L\. Sharkey\(2024\)Identifying functionally important features with end\-to\-end sparse dictionary learning\.Advances in Neural Information Processing Systems37,pp\. 107286–107325\.Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p3.1),[§11\.5](https://arxiv.org/html/2608.10172#S11.SS5.SSS0.Px1.p1.4),[§2](https://arxiv.org/html/2608.10172#S2.p2.1),[Corollary 9\.5](https://arxiv.org/html/2608.10172#S9.Thmtheorem5.p1.5.5)\.
- \[7\]T\. Bricken, A\. Templeton, J\. Batson, B\. Chen, A\. Jermyn, T\. Conerly, N\. Turner, C\. Anil, C\. Denison, A\. Askell, R\. Lasenby, Y\. Wu, S\. Kravec, N\. Schiefer, T\. Maxwell, N\. Joseph, Z\. Hatfield\-Dodds, A\. Tamkin, K\. Nguyen, B\. McLean, J\. E\. Burke, T\. Hume, S\. Carter, T\. Henighan, and C\. Olah\(2023\)Towards monosemanticity: decomposing language models with dictionary learning\.Transformer Circuits Thread\.Note:https://transformer\-circuits\.pub/2023/monosemantic\-features/index\.htmlCited by:[§11\.5](https://arxiv.org/html/2608.10172#S11.SS5.p2.4),[§2](https://arxiv.org/html/2608.10172#S2.p1.1),[§9\.3](https://arxiv.org/html/2608.10172#S9.SS3.p1.4)\.
- \[8\]S\. L\. Brunton, M\. Budišić, E\. Kaiser, and J\. N\. Kutz\(2021\)Modern koopman theory for dynamical systems\.arXiv preprint arXiv:2102\.12086\.Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p5.2),[§2](https://arxiv.org/html/2608.10172#S2.p5.1),[§4\.1](https://arxiv.org/html/2608.10172#S4.SS1.SSS0.Px1.p1.7)\.
- \[9\]D\. Chanin, J\. Wilken\-Smith, T\. Dulka, H\. Bhatnagar, S\. Golechha, and J\. Bloom\(2026\)A is for absorption: studying feature splitting and absorption in sparse autoencoders\.Advances in Neural Information Processing Systems38,pp\. 82318–82355\.Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p3.1),[§2](https://arxiv.org/html/2608.10172#S2.p2.1),[Corollary 9\.5](https://arxiv.org/html/2608.10172#S9.Thmtheorem5.p1.5.5)\.
- \[10\]R\. T\. Q\. Chen, Y\. Rubanova, J\. Bettencourt, and D\. K\. Duvenaud\(2018\)Neural ordinary differential equations\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p4.8),[§2](https://arxiv.org/html/2608.10172#S2.p8.1)\.
- \[11\]J\. Clymer, N\. Gabrieli, D\. Krueger, and T\. Larsen\(2024\)Safety cases: how to justify the safety of advanced ai systems\.arXiv preprint arXiv:2403\.10462\.Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p2.1),[§12](https://arxiv.org/html/2608.10172#S12.p5.3)\.
- \[12\]P\. Comon\(1994\)Independent component analysis, a new concept?\.Signal processing36\(3\),pp\. 287–314\.Cited by:[§2](https://arxiv.org/html/2608.10172#S2.p3.1),[Remark 6\.13](https://arxiv.org/html/2608.10172#S6.Thmtheorem13.p1.1)\.
- \[13\]A\. Conmy, A\. N\. Mavor\-Parker, A\. Lynch, S\. Heimersheim, and A\. Garriga\-Alonso\(2023\)Towards automated circuit discovery for mechanistic interpretability\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p1.1),[§2](https://arxiv.org/html/2608.10172#S2.p1.1)\.
- \[14\]C\. Davis and W\. M\. Kahan\(1970\)The rotation of eigenvectors by a perturbation\. III\.SIAM Journal on Numerical Analysis7\(1\),pp\. 1–46\.Cited by:[item 3](https://arxiv.org/html/2608.10172#S6.I2.i3.p1.2),[§6\.8](https://arxiv.org/html/2608.10172#S6.SS8.1.p1.5)\.
- \[15\]T\. Dettmers, M\. Lewis, Y\. Belkada, and L\. Zettlemoyer\(2022\)LLM\.int8\(\): 8\-bit matrix multiplication for transformers at scale\.InAdvances in Neural Information Processing Systems,Vol\.35\.Cited by:[§11\.2](https://arxiv.org/html/2608.10172#S11.SS2.SSS0.Px3.p4.11),[§7\.2](https://arxiv.org/html/2608.10172#S7.SS2.p1.1)\.
- \[16\]F\. Dietrich, T\. N\. Thiem, and I\. G\. Kevrekidis\(2020\)On the Koopman operator of algorithms\.SIAM Journal on Applied Dynamical Systems19\(2\),pp\. 860–885\.Cited by:[§2](https://arxiv.org/html/2608.10172#S2.p5.1)\.
- \[17\]Y\. Dong, J\. Cordonnier, and A\. Loukas\(2021\)Attention is not all you need: pure attention loses rank doubly exponentially with depth\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§10](https://arxiv.org/html/2608.10172#S10.p2.4),[§2](https://arxiv.org/html/2608.10172#S2.p7.1),[§2](https://arxiv.org/html/2608.10172#S2.p8.1)\.
- \[18\]J\. Dunefsky, P\. Chlenski, and N\. Nanda\(2024\)Transcoders find interpretable LLM feature circuits\.InThe Thirty\-eighth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=J6zHcScAo0)Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p1.1),[§2](https://arxiv.org/html/2608.10172#S2.p1.1)\.
- \[19\]N\. Elhage, N\. Nanda, C\. Olsson, T\. Henighan, N\. Joseph, B\. Mann, A\. Askell, Y\. Bai, A\. Chen, T\. Conerly, N\. DasSarma, D\. Drain, D\. Ganguli, Z\. Hatfield\-Dodds, D\. Hernandez, A\. Jones, J\. Kernion, L\. Lovitt, K\. Ndousse, D\. Amodei, T\. Brown, J\. Clark, J\. Kaplan, S\. McCandlish, and C\. Olah\(2021\)A mathematical framework for transformer circuits\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2021/framework/index.html)Cited by:[§2](https://arxiv.org/html/2608.10172#S2.p1.1)\.
- \[20\]L\. Elsner\(1985\)An optimal bound for the spectral variation of two matrices\.Linear algebra and its applications71,pp\. 77–80\.Cited by:[§6\.7](https://arxiv.org/html/2608.10172#S6.SS7.2.p1.3)\.
- \[21\]J\. Engels, L\. R\. Smith, and M\. Tegmark\(2024\)Decomposing the dark matter of sparse autoencoders\.Transactions on Machine Learning Research\.Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p3.1),[§2](https://arxiv.org/html/2608.10172#S2.p2.1)\.
- \[22\]D\. F\. Enns\(1984\)Model reduction with balanced realizations: an error bound and a frequency weighted generalization\.Proceedings of the 23rd IEEE Conference on Decision and Control,pp\. 127–132\.Cited by:[Remark 10\.9](https://arxiv.org/html/2608.10172#S10.Thmtheorem9.p1.3),[§2](https://arxiv.org/html/2608.10172#S2.p7.1)\.
- \[23\]L\. Gao, T\. Dupre la Tour, H\. Tillman, G\. Goh, R\. Troll, A\. Radford, I\. Sutskever, J\. Leike, and J\. Wu\(2025\)Scaling and evaluating sparse autoencoders\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 26721–26754\.Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p1.1),[§11\.5](https://arxiv.org/html/2608.10172#S11.SS5.p2.4),[§2](https://arxiv.org/html/2608.10172#S2.p1.1)\.
- \[24\]A\. Geiger, H\. Lu, T\. Icard, and C\. Potts\(2021\)Causal abstractions of neural networks\.Advances in neural information processing systems34,pp\. 9574–9586\.Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p1.1),[§2](https://arxiv.org/html/2608.10172#S2.p1.1),[§2](https://arxiv.org/html/2608.10172#S2.p3.1),[Remark 6\.13](https://arxiv.org/html/2608.10172#S6.Thmtheorem13.p1.1)\.
- \[25\]A\. Geiger, Z\. Wu, C\. Potts, T\. Icard, and N\. Goodman\(2024\)Finding alignments between interpretable causal variables and distributed neural representations\.InCausal Learning and Reasoning,pp\. 160–187\.Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p1.1),[§2](https://arxiv.org/html/2608.10172#S2.p1.1),[§2](https://arxiv.org/html/2608.10172#S2.p3.1),[Remark 6\.13](https://arxiv.org/html/2608.10172#S6.Thmtheorem13.p1.1)\.
- \[26\]B\. Geshkovski, C\. Letrouit, Y\. Polyanskiy, and P\. Rigollet\(2023\)The emergence of clusters in self\-attention dynamics\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§10](https://arxiv.org/html/2608.10172#S10.p2.4),[§2](https://arxiv.org/html/2608.10172#S2.p7.1),[§2](https://arxiv.org/html/2608.10172#S2.p8.1)\.
- \[27\]K\. Glover\(1984\)All optimal Hankel\-norm approximations of linear multivariable systems and theirL∞L^\{\\infty\}\-error bounds\.International Journal of Control39\(6\),pp\. 1115–1193\.Cited by:[§10\.2](https://arxiv.org/html/2608.10172#S10.SS2.2.p1.6),[§10](https://arxiv.org/html/2608.10172#S10.p1.5),[§2](https://arxiv.org/html/2608.10172#S2.p7.1)\.
- \[28\]A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan,et al\.\(2024\)The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.Cited by:[§4\.1](https://arxiv.org/html/2608.10172#S4.SS1.p1.7)\.
- \[29\]S\. Gugercin and A\. C\. Antoulas\(2004\)A survey of model reduction by balanced truncation and some new results\.International Journal of Control77\(8\),pp\. 748–766\.Cited by:[Remark 10\.10](https://arxiv.org/html/2608.10172#S10.Thmtheorem10.p1.5),[§2](https://arxiv.org/html/2608.10172#S2.p7.1)\.
- \[30\]E\. Haber and L\. Ruthotto\(2018\)Stable architectures for deep neural networks\.Inverse problems34\(1\),pp\. 014004\.Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p4.8),[§2](https://arxiv.org/html/2608.10172#S2.p8.1)\.
- \[31\]M\. Hanna, S\. Pezzelle, and Y\. Belinkov\(2024\)Have faith in faithfulness: going beyond circuit overlap when finding model mechanisms\.arXiv preprint arXiv:2403\.17806\.Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p1.1),[§2](https://arxiv.org/html/2608.10172#S2.p1.1)\.
- \[32\]Z\. He, W\. Shu, X\. Ge, L\. Chen, J\. Wang, Y\. Zhou, F\. Liu, Q\. Guo, X\. Huang, Z\. Wu,et al\.\(2024\)Llama scope: extracting millions of features from llama\-3\.1\-8b with sparse autoencoders\.arXiv preprint arXiv:2410\.20526\.Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p1.1),[§2](https://arxiv.org/html/2608.10172#S2.p1.1)\.
- \[33\]R\. A\. Horn and C\. R\. Johnson\(2012\)Matrix analysis\.Cambridge university press\.Cited by:[§6\.7](https://arxiv.org/html/2608.10172#S6.SS7.2.p1.3)\.
- \[34\]R\. Huben, H\. Cunningham, L\. R\. Smith, A\. Ewart, and L\. Sharkey\(2024\)Sparse autoencoders find highly interpretable features in language models\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=F76bwRSLeK)Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p1.1),[§2](https://arxiv.org/html/2608.10172#S2.p1.1)\.
- \[35\]A\. Hyvärinen, J\. Karhunen, and E\. Oja\(2001\)Independent component analysis\.John Wiley & Sons\.Cited by:[§2](https://arxiv.org/html/2608.10172#S2.p3.1)\.
- \[36\]A\. Hyvarinen and H\. Morioka\(2016\)Unsupervised feature extraction by time\-contrastive learning and nonlinear ica\.Advances in neural information processing systems29\.Cited by:[§2](https://arxiv.org/html/2608.10172#S2.p3.1),[Remark 6\.13](https://arxiv.org/html/2608.10172#S6.Thmtheorem13.p1.1)\.
- \[37\]A\. Karvonen, C\. Rager, J\. Lin, C\. Tigges, J\. I\. Bloom, D\. Chanin, Y\. Lau, E\. Farrell, C\. S\. McDougall, K\. Ayonrinde, D\. Till, M\. Wearden, A\. Conmy, S\. Marks, and N\. Nanda\(2025\)SAEBench: a comprehensive benchmark for sparse autoencoders in language model interpretability\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=qrU3yNfX0d)Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p3.1),[§11\.5](https://arxiv.org/html/2608.10172#S11.SS5.SSS0.Px1.p1.4),[§2](https://arxiv.org/html/2608.10172#S2.p2.1),[Corollary 9\.5](https://arxiv.org/html/2608.10172#S9.Thmtheorem5.p1.5.5)\.
- \[38\]T\. Kato\(1995\)Perturbation theory for linear operators\.2nd, Classics in Mathematics reprint edition,Springer\.Cited by:[§5\.2](https://arxiv.org/html/2608.10172#S5.SS2.5.p2.7),[Remark 5\.7](https://arxiv.org/html/2608.10172#S5.Thmtheorem7.p1.6),[§6\.5](https://arxiv.org/html/2608.10172#S6.SS5.3.p2.12)\.
- \[39\]I\. Khemakhem, D\. Kingma, R\. Monti, and A\. Hyvarinen\(2020\)Variational autoencoders and nonlinear ica: a unifying framework\.InInternational conference on artificial intelligence and statistics,pp\. 2207–2217\.Cited by:[§2](https://arxiv.org/html/2608.10172#S2.p3.1),[Remark 6\.13](https://arxiv.org/html/2608.10172#S6.Thmtheorem13.p1.1)\.
- \[40\]S\. Klus, F\. Nüske, S\. Peitz, J\. Niemann, C\. Clementi, and C\. Schütte\(2020\)Data\-driven approximation of the Koopman generator: model reduction, system identification, and control\.Physica D: Nonlinear Phenomena406,pp\. 132416\.Cited by:[§2](https://arxiv.org/html/2608.10172#S2.p5.1)\.
- \[41\]B\. O\. Koopman\(1931\)Hamiltonian systems and transformation in hilbert space\.Proceedings of the National Academy of Sciences17\(5\),pp\. 315–318\.Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p5.2),[§2](https://arxiv.org/html/2608.10172#S2.p5.1),[§4\.2](https://arxiv.org/html/2608.10172#S4.SS2.p1.1)\.
- \[42\]M\. Korda and I\. Mezić\(2018\)On convergence of extended dynamic mode decomposition to the Koopman operator\.Journal of Nonlinear Science28\(2\),pp\. 687–710\.Cited by:[§2](https://arxiv.org/html/2608.10172#S2.p5.1),[Remark 5\.8](https://arxiv.org/html/2608.10172#S5.Thmtheorem8.p1.6),[Remark 6\.13](https://arxiv.org/html/2608.10172#S6.Thmtheorem13.p1.1)\.
- \[43\]J\. Kramár, T\. Lieberum, R\. Shah, and N\. Nanda\(2024\)AtP⋆: an efficient and scalable method for localizing LLM behaviour to components\.arXiv preprint arXiv:2403\.00745\.Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p1.1),[§2](https://arxiv.org/html/2608.10172#S2.p1.1)\.
- \[44\]L\. Le Cam\(2012\)Asymptotic methods in statistical decision theory\.Springer Science & Business Media\.Cited by:[§7\.1](https://arxiv.org/html/2608.10172#S7.SS1.1.p1.4)\.
- \[45\]L\. LeCam\(1973\)Convergence of estimates under dimensionality restrictions\.The Annals of Statistics,pp\. 38–53\.Cited by:[§7\.1](https://arxiv.org/html/2608.10172#S7.SS1.1.p1.4)\.
- \[46\]T\. Lieberum, S\. Rajamanoharan, A\. Conmy, L\. Smith, N\. Sonnerat, V\. Varma, J\. Kramar, A\. Dragan, R\. Shah, and N\. Nanda\(2024\-11\)Gemma scope: open sparse autoencoders everywhere all at once on gemma 2\.InProceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP,Y\. Belinkov, N\. Kim, J\. Jumelet, H\. Mohebbi, A\. Mueller, and H\. Chen \(Eds\.\),Miami, Florida, US,pp\. 278–300\.External Links:[Link](https://aclanthology.org/2024.blackboxnlp-1.19/),[Document](https://dx.doi.org/10.18653/v1/2024.blackboxnlp-1.19)Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p1.1),[§11\.5](https://arxiv.org/html/2608.10172#S11.SS5.p2.4),[§2](https://arxiv.org/html/2608.10172#S2.p1.1)\.
- \[47\]J\. Lindsey, W\. Gurnee, E\. Ameisen, B\. Chen, A\. Pearce, N\. L\. Turner, C\. Citro, D\. Abrahams, S\. Carter, B\. Hosmer, J\. Marcus, M\. Sklar, A\. Templeton, T\. Bricken, C\. McDougall, H\. Cunningham, T\. Henighan, A\. Jermyn, A\. Jones, A\. Persic, Z\. Qi, T\. B\. Thompson, S\. Zimmerman, K\. Rivoire, T\. Conerly, C\. Olah, and J\. Batson\(2025\)On the biology of a large language model\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2025/attribution-graphs/biology.html)Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p1.1),[§2](https://arxiv.org/html/2608.10172#S2.p1.1)\.
- \[48\]F\. Locatello, S\. Bauer, M\. Lucic, G\. Raetsch, S\. Gelly, B\. Schölkopf, and O\. Bachem\(2019\)Challenging common assumptions in the unsupervised learning of disentangled representations\.Ininternational conference on machine learning,pp\. 4114–4124\.Cited by:[§2](https://arxiv.org/html/2608.10172#S2.p3.1)\.
- \[49\]G\. Lugosi and S\. Mendelson\(2019\)Mean estimation and regression under heavy\-tailed distributions: a survey\.Foundations of Computational Mathematics19\(5\),pp\. 1145–1190\.Cited by:[Theorem 7\.2](https://arxiv.org/html/2608.10172#S7.Thmtheorem2.p1.6.6)\.
- \[50\]K\. Meng, D\. Bau, A\. J\. Andonian, and Y\. Belinkov\(2022\)Locating and editing factual associations in gpt\.InAdvances in neural information processing systems,Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p1.1),[§2](https://arxiv.org/html/2608.10172#S2.p1.1)\.
- \[51\]I\. Mezić\(2005\)Spectral properties of dynamical systems, model reduction and decompositions\.Nonlinear Dynamics41\(1\),pp\. 309–325\.Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p5.2),[§2](https://arxiv.org/html/2608.10172#S2.p5.1)\.
- \[52\]S\. Minsker\(2015\)Geometric median and robust estimation in Banach spaces\.Bernoulli21\(4\),pp\. 2308–2335\.Cited by:[§7\.2](https://arxiv.org/html/2608.10172#S7.SS2.2.p2.8),[Theorem 7\.2](https://arxiv.org/html/2608.10172#S7.Thmtheorem2.p1.6.6)\.
- \[53\]B\. C\. Moore\(1981\)Principal component analysis in linear systems: controllability, observability, and model reduction\.IEEE Transactions on Automatic Control26\(1\),pp\. 17–32\.Cited by:[§10](https://arxiv.org/html/2608.10172#S10.p1.5),[§2](https://arxiv.org/html/2608.10172#S2.p7.1)\.
- \[54\]N\. Nanda, L\. Chan, T\. Lieberum, J\. Smith, and J\. Steinhardt\(2023\)Progress measures for grokking via mechanistic interpretability\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2608.10172#S2.p8.1)\.
- \[55\]C\. Olah, N\. Cammarata, L\. Schubert, G\. Goh, M\. Petrov, and S\. Carter\(2020\)Zoom in: an introduction to circuits\.Distill\.External Links:[Document](https://dx.doi.org/10.23915/distill.00024.001),[Link](https://distill.pub/2020/circuits/zoom-in/)Cited by:[§2](https://arxiv.org/html/2608.10172#S2.p8.1),[§9\.3](https://arxiv.org/html/2608.10172#S9.SS3.p1.4)\.
- \[56\]C\. Olsson, N\. Elhage, N\. Nanda, N\. Joseph, N\. DasSarma, T\. Henighan, B\. Mann, A\. Askell, Y\. Bai, A\. Chen, T\. Conerly, D\. Drain, D\. Ganguli, Z\. Hatfield\-Dodds, D\. Hernandez, S\. Johnston, A\. Jones, J\. Kernion, L\. Lovitt, K\. Ndousse, D\. Amodei, T\. Brown, J\. Clark, J\. Kaplan, S\. McCandlish, and C\. Olah\(2022\)In\-context learning and induction heads\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html)Cited by:[§2](https://arxiv.org/html/2608.10172#S2.p8.1)\.
- \[57\]G\. Paulo, N\. Belrose,et al\.\(2025\)Transcoders beat sparse autoencoders for interpretability\.arXiv preprint arXiv:2501\.18823\.Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p3.1),[§2](https://arxiv.org/html/2608.10172#S2.p2.1)\.
- \[58\]J\. L\. Proctor, S\. L\. Brunton, and J\. N\. Kutz\(2016\)Dynamic mode decomposition with control\.SIAM Journal on Applied Dynamical Systems15\(1\),pp\. 142–161\.Cited by:[§2](https://arxiv.org/html/2608.10172#S2.p5.1),[§4\.3](https://arxiv.org/html/2608.10172#S4.SS3.p1.1),[Remark 4\.2](https://arxiv.org/html/2608.10172#S4.Thmtheorem2.p1.4)\.
- \[59\]S\. Rajamanoharan, A\. Conmy, L\. Smith, T\. Lieberum, V\. Varma, J\. Kramar, R\. Shah, and N\. Nanda\(2024\)Improving sparse decomposition of language model activations with gated sparse autoencoders\.InThe Thirty\-eighth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=zLBlin2zvW)Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p1.1),[§2](https://arxiv.org/html/2608.10172#S2.p1.1)\.
- \[60\]S\. Rajamanoharan, T\. Lieberum, N\. Sonnerat, A\. Conmy, V\. Varma, J\. Kramár, and N\. Nanda\(2024\)Jumping ahead: improving reconstruction fidelity with JumpReLU sparse autoencoders\.arXiv preprint arXiv:2407\.14435\.Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p1.1),[§2](https://arxiv.org/html/2608.10172#S2.p1.1)\.
- \[61\]W\. T\. Redman, M\. Fonoberova, R\. Mohr, I\. G\. Kevrekidis, and I\. Mezić\(2022\)An operator theoretic view on pruning deep neural networks\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2608.10172#S2.p5.1)\.
- \[62\]P\. J\. Schmid\(2010\)Dynamic mode decomposition of numerical and experimental data\.Journal of Fluid Mechanics656,pp\. 5–28\.Cited by:[§2](https://arxiv.org/html/2608.10172#S2.p5.1)\.
- \[63\]B\. Schölkopf, F\. Locatello, S\. Bauer, N\. R\. Ke, N\. Kalchbrenner, A\. Goyal, and Y\. Bengio\(2021\)Toward causal representation learning\.Proceedings of the IEEE109\(5\),pp\. 612–634\.Cited by:[§2](https://arxiv.org/html/2608.10172#S2.p3.1)\.
- \[64\]A\. Simpkins\(2012\)System identification: theory for the user, \(ljung, l\.; 1999\)\[on the shelf\]\.IEEE Robotics & Automation Magazine19\(2\),pp\. 95–96\.Cited by:[§2](https://arxiv.org/html/2608.10172#S2.p3.1),[Remark 4\.12](https://arxiv.org/html/2608.10172#S4.Thmtheorem12.p1.3)\.
- \[65\]G\. W\. Stewart and J\. Sun\(1990\)Matrix perturbation theory\.Academic Press\.Cited by:[§10\.4](https://arxiv.org/html/2608.10172#S10.SS4.1.p1.8),[Remark 10\.5](https://arxiv.org/html/2608.10172#S10.Thmtheorem5.p1.9),[item 3](https://arxiv.org/html/2608.10172#S6.I2.i3.p1.2),[§6\.1](https://arxiv.org/html/2608.10172#S6.SS1.p2.5),[§6\.5](https://arxiv.org/html/2608.10172#S6.SS5.3.p2.11),[§6\.8](https://arxiv.org/html/2608.10172#S6.SS8.1.p1.5)\.
- \[66\]A\. Syed, C\. Rager, and A\. Conmy\(2024\-11\)Attribution patching outperforms automated circuit discovery\.InProceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP,Y\. Belinkov, N\. Kim, J\. Jumelet, H\. Mohebbi, A\. Mueller, and H\. Chen \(Eds\.\),Miami, Florida, US,pp\. 407–416\.External Links:[Link](https://aclanthology.org/2024.blackboxnlp-1.25/),[Document](https://dx.doi.org/10.18653/v1/2024.blackboxnlp-1.25)Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p1.1),[§2](https://arxiv.org/html/2608.10172#S2.p1.1)\.
- \[67\]G\. Team, M\. Riviere, S\. Pathak, P\. G\. Sessa, C\. Hardin, S\. Bhupatiraju, L\. Hussenot, T\. Mesnard, B\. Shahriari, A\. Ramé,et al\.\(2024\)Gemma 2: improving open language models at a practical size\.arXiv preprint arXiv:2408\.00118\.Cited by:[§4\.1](https://arxiv.org/html/2608.10172#S4.SS1.p1.7)\.
- \[68\]A\. Templeton, T\. Conerly, J\. Marcus, J\. Lindsey, T\. Bricken, B\. Chen, A\. Pearce, C\. Citro, E\. Ameisen, A\. Jones, H\. Cunningham, N\. L\. Turner, C\. McDougall, M\. MacDiarmid, C\. D\. Freeman, T\. R\. Sumers, E\. Rees, J\. Batson, A\. Jermyn, S\. Carter, C\. Olah, and T\. Henighan\(2024\)Scaling monosemanticity: extracting interpretable features from claude 3 sonnet\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html)Cited by:[§2](https://arxiv.org/html/2608.10172#S2.p1.1),[§9\.3](https://arxiv.org/html/2608.10172#S9.SS3.p1.4)\.
- \[69\]J\. A\. Tropp\(2015\)An introduction to matrix concentration inequalities\.Foundations and Trends in Machine Learning8\(1–2\),pp\. 1–230\.Cited by:[item 1](https://arxiv.org/html/2608.10172#S6.I2.i1.p1.2),[§6\.3](https://arxiv.org/html/2608.10172#S6.SS3.1.p1.9)\.
- \[70\]A\. B\. Tsybakov\(2009\)Introduction to nonparametric estimation\.Springer Series in Statistics,Springer,New York\.Cited by:[§7\.1](https://arxiv.org/html/2608.10172#S7.SS1.1.p1.4),[§7\.1](https://arxiv.org/html/2608.10172#S7.SS1.p4.4)\.
- \[71\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, and I\. Polosukhin\(2017\)Attention is all you need\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§4\.1](https://arxiv.org/html/2608.10172#S4.SS1.p1.7)\.
- \[72\]R\. Vershynin\(2018\)High\-dimensional probability: an introduction with applications in data science\.Cambridge Series in Statistical and Probabilistic Mathematics,Cambridge University Press\.Cited by:[item 1](https://arxiv.org/html/2608.10172#S6.I2.i1.p1.2),[§6\.1](https://arxiv.org/html/2608.10172#S6.SS1.p2.5),[§6\.3](https://arxiv.org/html/2608.10172#S6.SS3.1.p1.9),[Remark 6\.9](https://arxiv.org/html/2608.10172#S6.Thmtheorem9.p1.4)\.
- \[73\]J\. Vig, S\. Gehrmann, Y\. Belinkov, S\. Qian, D\. Nevo, Y\. Singer, and S\. Shieber\(2020\)Causal mediation analysis for interpreting neural NLP: the case of gender bias\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p1.1),[§2](https://arxiv.org/html/2608.10172#S2.p1.1)\.
- \[74\]K\. R\. Wang, A\. Variengien, A\. Conmy, B\. Shlegeris, and J\. Steinhardt\(2023\)Interpretability in the wild: a circuit for indirect object identification in GPT\-2 Small\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§11\.3](https://arxiv.org/html/2608.10172#S11.SS3.p1.1),[§11\.3](https://arxiv.org/html/2608.10172#S11.SS3.p2.4),[§2](https://arxiv.org/html/2608.10172#S2.p1.1)\.
- \[75\]E\. Weinan, Y\. Duan, L\. Kong, and M\. Guo\(2017\)A proposal on machine learning via dynamical systems\.Links2024,pp\. 08–27\.Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p4.8),[§2](https://arxiv.org/html/2608.10172#S2.p8.1)\.
- \[76\]J\. C\. Willems, P\. Rapisarda, I\. Markovsky, and B\. L\. De Moor\(2005\)A note on persistency of excitation\.Systems & Control Letters54\(4\),pp\. 325–329\.Cited by:[§2](https://arxiv.org/html/2608.10172#S2.p3.1),[Remark 4\.12](https://arxiv.org/html/2608.10172#S4.Thmtheorem12.p1.3)\.
- \[77\]M\. Williams, I\. Kevrekidis, and C\. Rowley\(2015\)A data\-driven approximation of the koopman operator: extending dynamic mode decomposition\.\.Journal of nonlinear science25\(6\)\.Cited by:[§2](https://arxiv.org/html/2608.10172#S2.p5.1),[§4\.3](https://arxiv.org/html/2608.10172#S4.SS3.p1.1)\.
- \[78\]M\. O\. Williams, C\. W\. Rowley, and I\. G\. Kevrekidis\(2015\)A kernel\-based method for data\-driven Koopman spectral analysis\.Journal of Computational Dynamics2\(2\),pp\. 247–265\.Cited by:[§11\.1](https://arxiv.org/html/2608.10172#S11.SS1.p3.1),[§2](https://arxiv.org/html/2608.10172#S2.p5.1)\.
- \[79\]Z\. Wu, A\. Geiger, T\. Icard, C\. Potts, and N\. Goodman\(2023\)Interpretability at scale: identifying causal mechanisms in alpaca\.InThirty\-seventh Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=nRfClnMhVX)Cited by:[§1](https://arxiv.org/html/2608.10172#S1.p1.1),[§2](https://arxiv.org/html/2608.10172#S2.p1.1)\.

Similar Articles

Beyond the Black Box: Interpretability of Agentic AI Tool Use

arXiv cs.AI

This paper introduces a mechanistic interpretability toolkit using Sparse Autoencoders and linear probes to monitor internal model states before AI agents invoke tools, aiming to improve diagnostics and safety in enterprise workflows.

Towards Intrinsic Interpretability of Large Language Models: A Survey of Design Principles and Architectures

arXiv cs.CL

A comprehensive survey reviewing recent advances in intrinsic interpretability for Large Language Models, categorizing approaches into five design paradigms: functional transparency, concept alignment, representational decomposability, explicit modularization, and latent sparsity induction. The paper addresses the challenge of building transparency directly into model architectures rather than relying on post-hoc explanation methods.