Structural Instability of Feature Composition
Summary
This paper presents a geometric framework to analyze the instability of feature composition in Sparse Autoencoders, revealing that non-linearities cause a ratchet effect leading to compositional collapse beyond a critical density.
View Cached Full Text
Cached at: 05/08/26, 06:37 AM
# Structural Instability of Feature Composition
Source: [https://arxiv.org/html/2605.05223](https://arxiv.org/html/2605.05223)
\\coltauthor\\Name
Yunpeng Zhou\\Emailkc804139@student\.reading\.ac\.uk \\addrWhiteknights House, Whiteknights, Reading, RG6 6UR
###### Abstract
Sparse Autoencoders \(SAEs\) have emerged as a powerful paradigm for disentangling feature superposition in transformer\-based architectures, enabling precise control via activation steering\. However, the theoretical foundations of compositional steering—the simultaneous activation of distinct semantic latents—remain under\-explored\. The prevailing Linear Representation Hypothesis often abstracts away non\-linear interference effects that arise in overcomplete dictionaries\. We present a geometric framework for analyzing the instability of feature unions\. Modeling the activation space as a high\-dimensional sparse cone manifold, we derive an asymptotic compositional\-collapse threshold under a spherical dictionary model, characterized by the Gaussian mean width \(statistical dimension\) of the signal cone\. We further show that, in the high\-bias regime, ReLU rectification converts microscopic correlation\-induced variance fluctuations into a systematic drift that accumulates under composition, yielding interference growth consistent with a ratchet effect\. We validate the predicted scaling trends on structured semantic features extracted from CLEVR, where hierarchical correlations accelerate the transition relative to random baselines\. Together, our results highlight geometric constraints on the scalability of union\-based steering and motivate composition mechanisms that explicitly manage interference beyond naive linear superposition\.
###### keywords:
Sparse Representation, Mechanistic Interpretability, Phase Transitions, Compositional Generalization, Feature Steering
## 1Introduction
The field of interpretability has shifted toward manipulating dynamic activations via Sparse Autoencoders \(SAEs\), which decompose dense residual streams into interpretable features\. Supported by recent scaling laws and resources\(templeton2024scaling;gao2025scaling;lieberum2024gemma\_scope\), this decomposition enablesactivation steeringto elicit specific model behaviors\(rimsky2024caa;zou2023representation\)\. These interventions typically rely on the Linear Representation Hypothesis\(park2024linear\), assuming semantic composition equals linear vector addition\.
However, the geometric stability of this hypothesis is fragile\. Evidence suggests features are often multi\-dimensional\(engels2025not\_linear\)and lack causal unity\(karvonen2025saebench;leask2025canonical\)\. Crucially, simultaneous manipulation of multiple features frequently triggerscompositional collapse—a state of semantic incoherence\(stickland2024kts\)\. While prior work noted interference in structured manifolds\(elhage2022toy;bietti2023birth\), a rigorous geometric characterization of this instability remains elusive\.
We argue that compositional breakdown is a fundamental consequence of high\-dimensional, overcomplete geometry \(m≫nm\\gg n\)\. In this regime, inevitable non\-orthogonality causes interference noise\(scherlis2022polysemanticity\)\. Modeling the activation space as asparse cone manifold, we show that ReLU non\-linearities act as aRatchet: unlike linear systems where noise cancels out, rectification bias exponentially amplifies geometric interference, rendering the system strictly super\-additive to noise\.
Our contributions are threefold: \(1\) We establish a Stability Threshold via the Gaussian Mean Width of the signal cone, proving spurious activations are inevitable beyond a critical compositional density; \(2\) We characterize the Ratchet Mechanism, providing a micro\-foundation for how marginal correlations amplify into macroscopic collapse; and \(3\) We provide Empirical Validation on CLEVR, demonstrating that structured semantic correlations accelerate this phase transition compared to random baselines\. These findings establish a hard geometric barrier to scaling naive linear composition in static embedding spaces\.
## 2Preliminaries
We adopt standard notation from high\-dimensional probability\. Let\[m\]\[m\]denote\{1,…,m\}\\\{1,\\dots,m\\\}\. For a vectorv∈ℝnv\\in\\mathbb\{R\}^\{n\},‖v‖p\\\|v\\\|\_\{p\}denotes itsℓp\\ell\_\{p\}norm\. For a matrixAA,‖A‖op\\\|A\\\|\_\{op\}denotes its spectral norm, andASA\_\{S\}denotes the submatrix restricted to columns indexed byS⊂\[m\]S\\subset\[m\]\.
### 2\.1The Overcomplete Dictionary Model
We analyze the activation spaceℝn\\mathbb\{R\}^\{n\}where semantic features are encoded by anovercomplete dictionaryD∈ℝn×mD\\in\\mathbb\{R\}^\{n\\times m\}withm=δnm=\\delta n\(δ\>1\\delta\>1\)\.
###### Assumption 1\(Dictionary Regularity\)
The dictionaryD=\[d1,…,dm\]D=\[d\_\{1\},\\dots,d\_\{m\}\]satisfies:
1. 1\.Unit Norm:‖di‖2=1\\\|d\_\{i\}\\\|\_\{2\}=1for alli∈\[m\]i\\in\[m\]\.
2. 2\.μ\\mu\-Incoherence:The mutual coherenceμ\(D\)≜maxi≠j\|⟨di,dj⟩\|\\mu\(D\)\\triangleq\\max\_\{i\\neq j\}\|\\langle d\_\{i\},d\_\{j\}\\rangle\|scales asO\(1/n\)O\(1/\\sqrt\{n\}\)\.
#### Random Baseline vs\. Structural Reality\.
We primarily utilize the random spherical model \(wheredi∼Unif\(𝕊n−1\)d\_\{i\}\\sim\\text\{Unif\}\(\\mathbb\{S\}^\{n\-1\}\)\) to derive asymptotically tight phase\-transition bounds\. While real\-world semantics \(e\.g\., CLEVR\) exhibit hierarchical correlations that deviate from this i\.i\.d\. assumption, we address the robustness of our bounds under such structural perturbations empirically in Section[4\.5](https://arxiv.org/html/2605.05223#S4.SS5)\.
### 2\.2The Sparse Cone Manifold
Unlike classical Compressed Sensing focused on linear subspaces, neural activations are constrained by the ReLU functionσ\(x\)=max\(0,x\)\\sigma\(x\)=\\max\(0,x\), restricting signals to a union of positive cones\.
###### Definition 2\.1\(Positive Feature Cone\)\.
For a support setS⊂\[m\]S\\subset\[m\], the feature cone𝒞\(S\)\\mathcal\{C\}\(S\)is the set of non\-negative linear combinations of atoms inSS:
𝒞\(S\)≜\{∑i∈Sαidi∣αi\>0\}⊂span\(DS\)\.\\mathcal\{C\}\(S\)\\triangleq\\left\\\{\\sum\_\{i\\in S\}\\alpha\_\{i\}d\_\{i\}\\mid\\alpha\_\{i\}\>0\\right\\\}\\subset\\text\{span\}\(D\_\{S\}\)\.
###### Definition 2\.2\(kk\-Sparse Manifold\)\.
The set of validkk\-sparse representations forms a non\-convex manifold:
ℳk≜⋃S⊂\[m\],\|S\|≤k𝒞\(S\)\.\\mathcal\{M\}\_\{k\}\\triangleq\\bigcup\_\{S\\subset\[m\],\|S\|\\leq k\}\\mathcal\{C\}\(S\)\.
Our stability analysis \(Section[4](https://arxiv.org/html/2605.05223#S4)\) centers on whether the compositionz=zA\+zBz=z\_\{A\}\+z\_\{B\}remains close toℳ2k\\mathcal\{M\}\_\{2k\}or falls into the invalid ambient void, triggering undefined features\.
### 2\.3Composition and Interference
We formalize “steering“ as the algebraic addition of latent vectors\. Letα∗∈ℝm\\alpha^\{\*\}\\in\\mathbb\{R\}^\{m\}be a sparse coefficient vector supported onSS\. The pre\-activation isx=Dα∗x=D\\alpha^\{\*\}\. When composing distinct concepts supported on disjoint setsSAS\_\{A\}andSBS\_\{B\}, the idealized state isxunion=D\(αA\+αB\)x\_\{union\}=D\(\\alpha\_\{A\}\+\\alpha\_\{B\}\)\.
Physical realization is subject to interference fromGhost Features\. LetJ=\[m\]∖\(SA∪SB\)J=\[m\]\\setminus\(S\_\{A\}\\cup S\_\{B\}\)be the set of inactive atoms\. The projection of the composite signal onto a ghost atomdjd\_\{j\}\(j∈Jj\\in J\) is:
ℐj\(SA,SB\)≜⟨xunion,dj⟩=⟨DSAαA,dj⟩⏟Interference from A\+⟨DSBαB,dj⟩⏟Interference from B\.\\mathcal\{I\}\_\{j\}\(S\_\{A\},S\_\{B\}\)\\triangleq\\langle x\_\{union\},d\_\{j\}\\rangle=\\underbrace\{\\langle D\_\{S\_\{A\}\}\\alpha\_\{A\},d\_\{j\}\\rangle\}\_\{\\text\{Interference from A\}\}\+\\underbrace\{\\langle D\_\{S\_\{B\}\}\\alpha\_\{B\},d\_\{j\}\\rangle\}\_\{\\text\{Interference from B\}\}\.
Aspurious activationoccurs ifℐj\>β\\mathcal\{I\}\_\{j\}\>\\beta\(the activation bias\)\. The core mathematical challenge is to characterize the distribution of the maximum interferencesupj∈Jℐj\\sup\_\{j\\in J\}\\mathcal\{I\}\_\{j\}as a function of the sparsityk=\|SA\|\+\|SB\|k=\|S\_\{A\}\|\+\|S\_\{B\}\|\. For a summary of notation, see Table[1](https://arxiv.org/html/2605.05223#A1.T1)in Appendix[A\.1](https://arxiv.org/html/2605.05223#A1.SS1)\.
## 3Micro\-Analysis: The Geometry of Pairwise Entanglement
In this section, we analyze the local interaction between two distinct semantic factors\. Before establishing the macroscopic phase transition \(Section[4](https://arxiv.org/html/2605.05223#S4)\), we must first quantify the geometric mechanism by which the union of two sparse cones generates interference\.
Consider two disjoint index setsSA,SB⊂\[m\]S\_\{A\},S\_\{B\}\\subset\[m\]with cardinalitieskA,kB≪nk\_\{A\},k\_\{B\}\\ll n\. Let𝒰A=span\(DSA\)\\mathcal\{U\}\_\{A\}=\\text\{span\}\(D\_\{S\_\{A\}\}\)and𝒰B=span\(DSB\)\\mathcal\{U\}\_\{B\}=\\text\{span\}\(D\_\{S\_\{B\}\}\)be the subspaces spanned by the active atoms\. The fundamental challenge of compositional steering lies in the fact that whileSA∩SB=∅S\_\{A\}\\cap S\_\{B\}=\\emptyset, their embedded subspaces are not orthogonal:𝒰A⟂𝒰B\\mathcal\{U\}\_\{A\}\\perp\\mathcal\{U\}\_\{B\}does not hold in general for overcomplete dictionaries\.
### 3\.1Principal Angles and Subspace Alignment
To rigorously measure the geometric antagonism between two concepts, we invoke the notion ofprincipal angles\. The alignment between𝒰A\\mathcal\{U\}\_\{A\}and𝒰B\\mathcal\{U\}\_\{B\}dictates the worst\-case interference\.
###### Definition 3\.1\(Interaction Singular Values\)\.
LetQAQ\_\{A\}andQBQ\_\{B\}be orthonormal bases for𝒰A\\mathcal\{U\}\_\{A\}and𝒰B\\mathcal\{U\}\_\{B\}, respectively\. The interaction singular valuesσ1≥σ2≥⋯≥σmin\(kA,kB\)\\sigma\_\{1\}\\geq\\sigma\_\{2\}\\geq\\dots\\geq\\sigma\_\{\\min\(k\_\{A\},k\_\{B\}\)\}are the singular values of the cross\-projection matrixMAB=QATQBM\_\{AB\}=Q\_\{A\}^\{T\}Q\_\{B\}\. The smallest principal angleθmin\\theta\_\{\\min\}satisfiescos\(θmin\)=σ1\\cos\(\\theta\_\{\\min\}\)=\\sigma\_\{1\}\.
Ifσ1≈1\\sigma\_\{1\}\\approx 1, the subspaces are nearly parallel, making disentanglement ill\-posed\. However, in high\-dimensional random dictionaries,σ1\\sigma\_\{1\}is typically bounded\. The danger arises not from the subspaces collapsing onto each other, but from their joint projection onto thecomplementarydictionary atoms\.
### 3\.2The Leakage Operator and Variance Decomposition
We define the Leakage OperatorℒAB:𝒰A×𝒰B→ℝm−\(kA\+kB\)\\mathcal\{L\}\_\{AB\}:\\mathcal\{U\}\_\{A\}\\times\\mathcal\{U\}\_\{B\}\\to\\mathbb\{R\}^\{m\-\(k\_\{A\}\+k\_\{B\}\)\}to quantify the signal bleeding into the ghost featuresJ=\[m\]∖\(SA∪SB\)J=\[m\]\\setminus\(S\_\{A\}\\cup S\_\{B\}\)\. For any composite steering vectorz=DαA\+DαBz=D\\alpha\_\{A\}\+D\\alpha\_\{B\}, the pre\-activation on a ghost atomdjd\_\{j\}\(j∈Jj\\in J\) is defined by the inner product⟨z,dj⟩\\langle z,d\_\{j\}\\rangle\.
To characterize the structural instability, we analyze the second\-order moments of this interference\. LetG=DTDG=D^\{T\}Dbe the Gram matrix of the dictionary\. The interference energy at a ghost featurejjdepends not only on individual atom coherence but also on the collective alignment between the active setsSAS\_\{A\}andSBS\_\{B\}\.
###### Lemma 3\.2\(Variance Decomposition under Spherical Ensemble\)\.
LetαA,αB\\alpha\_\{A\},\\alpha\_\{B\}be fixed unit\-norm coefficient vectors\. Assume the dictionary atomsdjd\_\{j\}are drawn independently fromUnif\(𝕊n−1\)\\text\{Unif\}\(\\mathbb\{S\}^\{n\-1\}\)\. The expected projection energy ofz=DαA\+DαBz=D\\alpha\_\{A\}\+D\\alpha\_\{B\}onto a generic ghost atomdj∉SA∪SBd\_\{j\}\\notin S\_\{A\}\\cup S\_\{B\}, taken over the randomness ofDD, satisfies:
𝔼D\[⟨z,dj⟩2\]=1n\(‖αA‖2\+‖αB‖2\)\+2μeffρ\(SA,SB\)\+O\(n−2\)\\mathbb\{E\}\_\{D\}\[\\langle z,d\_\{j\}\\rangle^\{2\}\]=\\frac\{1\}\{n\}\(\\\|\\alpha\_\{A\}\\\|^\{2\}\+\\\|\\alpha\_\{B\}\\\|^\{2\}\)\+2\\mu\_\{eff\}\\rho\(S\_\{A\},S\_\{B\}\)\+O\(n^\{\-2\}\)\(1\)whereμeff=𝔼\[\|⟨di,dj⟩\|\]∼n−1/2\\mu\_\{eff\}=\\mathbb\{E\}\[\|\\langle d\_\{i\},d\_\{j\}\\rangle\|\]\\sim n^\{\-1/2\}is the expected coherence, andρ\(SA,SB\)=αA⊤\(DSA⊤DSB\)αB\\rho\(S\_\{A\},S\_\{B\}\)=\\alpha\_\{A\}^\{\\top\}\(D\_\{S\_\{A\}\}^\{\\top\}D\_\{S\_\{B\}\}\)\\alpha\_\{B\}captures the subspace alignment\. The higher\-order residueR\(α\)R\(\\alpha\)vanishes asO\(1/n\)O\(1/n\)relative to the signal term\.
The term2μρ2\\mu\\rhois the geometric catalyst for compositional failure\. In structured domains such as CLEVR, where attributes like color \(e\.g\., red\) and shape \(e\.g\., cube\) are non\-orthogonal,ρ\\rhois strictly positive\. As we show next, this term does not merely add linear noise; it acts as a trigger for the rectified ratchet mechanism\.
### 3\.3The Rectified Ratchet: Non\-linear Amplification
A critical insight of this work is that linear interference analysis is insufficient for neural circuits\. The ReLU non\-linearityσ\(x\)=max\(0,x\)\\sigma\(x\)=\\max\(0,x\)breaks the symmetry of the interference distribution, preventing negative correlations from cancelling out positive leakage\.
we assume the feature vectors\{dj\}\\\{d\_\{j\}\\\}are sampled i\.i\.d\. from the uniform distribution on the unit sphere𝕊n−1\\mathbb\{S\}^\{n\-1\}\. Remark\. While we assume random relative orientations for analytical tractability, our results persist under mild coherence conditions \(μ\\mu\-independence\) common in the compressed sensing literature\(donoho2009observed\), where measure concentration ensures that structured dictionaries exhibit random\-like properties in high dimensions\.
###### Theorem 3\.3\(Rectified Drift and Convexity\)\.
Let the pre\-activation interference beX∼𝒩\(0,v\(ρ\)\)X\\sim\\mathcal\{N\}\(0,v\(\\rho\)\)\. To ensure physical validity, we define the effective variance as the rectified quantityv\(ρ\)≜\(σ02\+2μρ\)\+v\(\\rho\)\\triangleq\(\\sigma\_\{0\}^\{2\}\+2\\mu\\rho\)\_\{\+\}, where\(x\)\+=max\(0,x\)\(x\)\_\{\+\}=\\max\(0,x\)\. The rectified drift is then given by:
η\(ρ\):=𝔼\[σ\(X\)\]=\(σ02\+2μρ\)\+2π\.\\eta\(\\rho\):=\\mathbb\{E\}\[\\sigma\(X\)\]=\\frac\{\\sqrt\{\(\\sigma\_\{0\}^\{2\}\+2\\mu\\rho\)\_\{\+\}\}\}\{\\sqrt\{2\\pi\}\}\.\(2\)This implies a strict super\-additivity to noise: geometric variance fluctuations strictly increase the expected interference energy compared to an isotropic baseline\.
Remark \(Drift vs\. Tail Probability\)\.It is important to note that rectification does not alter the tail exceedance probability for positive thresholds: for anyβ\>0\\beta\>0,ℙ\(σ\(X\)\>β\)=ℙ\(X\>β\)\\mathbb\{P\}\(\\sigma\(X\)\>\\beta\)=\\mathbb\{P\}\(X\>\\beta\)\. Thus, the Ratchet Effect is not a local amplification of tail variance\. Instead, its mechanism is the conversion of symmetric spread into asystematic mean driftη\>0\\eta\>0\. While the drift per feature is linear, its cumulative effect on composition is geometric: forkkactive features, the accumulated drift shifts the center of the interference cone𝒞ghost\\mathcal\{C\}\_\{ghost\}toward the signal\. This effective expansion of the cone’s statistical dimensionΦ\\Phiconsumes the safety marginΔgap\\Delta\_\{gap\}linearly, which—due to the concentration of measure on the sphere—causes the safety probability to decay exponentially\. Thus, the linear ”Ratchet” acts as the fuel for a geometric phase transition\.
#### Interpretation of the Ratchet Effect\.
Eq\. \(3\) reveals a non\-linear coupling between dictionary geometry and representation stability\. The cross\-correlationρ\\rhoenters the exponent as a sensitivity multiplier\. Thus, even a marginal increase in the alignment betweenSAS\_\{A\}andSBS\_\{B\}\(as seen in CLEVR attribute combinations\) yields an exponential rise in the*pre\-activation*exceedance of ghost features under the inflated variance proxy\. Crucially, rectification does not alter positive\-threshold exceedance, but it converts symmetric interference into one\-sided drift: constructive deviations contribute to𝔼\[σ\(X\)\]\\mathbb\{E\}\[\\sigma\(X\)\], while negative deviations are clipped at zero\. This induces a one\-way ratchet mechanism in which variance\-driven leakage is turned into systematic semantic drift, pushing the system toward the macroscopic phase transition in Section[4](https://arxiv.org/html/2605.05223#S4)\.
A\. Linear Superpositionβ\\beta−β\-\\betaxxP\(x\)P\(x\)Symmetric Noise𝔼\[x\]=0\\mathbb\{E\}\[x\]=0B\. ReLU RatchetDriftη\>0\\eta\>0β\\betay=σ\(x\)y=\\sigma\(x\)Massat 0Rectified BiasInterference AccumulatesFigure 1:Mechanistic origin of the ReLU Ratchet\.\(A\)In linear systems, interference is symmetric and zero\-mean \(𝔼\[X\]=0\\mathbb\{E\}\[X\]=0\), allowing noise to cancel out across features\.\(B\)ReLU rectification induces asystematic mean driftη=𝔼\[σ\(X\)\]\>0\\eta=\\mathbb\{E\}\[\\sigma\(X\)\]\>0, transforming stochastic fluctuations into a persistent geometric bias\. This drift shifts the interference distribution toward the threshold, causing a faster saturation of the ambient space than simple variance expansion\.
### 3\.4The Rectified Geometric Tensor and Asymptotic Expansion
Standard linear analysis assumes interference is symmetric and cancels out\. However, the ReLU nonlinearityσ\(⋅\)\\sigma\(\\cdot\)acts as a one\-sided energy accumulation, selectively amplifying the positive tail of the interference distribution\. To quantify this precisely, we introduce the Interference Tensor and derive a controlled asymptotic lower bound using the expansion of the Gaussian Mills’ ratio\.
Let the pre\-activation interference on the ghost subspaceJJbe modeled by the random fieldZJ=PJ\(DαA\+DαB\)Z\_\{J\}=P\_\{J\}\(D\\alpha\_\{A\}\+D\\alpha\_\{B\}\)\. The aggregate spurious energy is governed by the conditional interaction moment:
ℰspur=𝔼\[‖σ\(ZJ\)‖2\|DSA∪SB\]\\mathcal\{E\}\_\{spur\}=\\mathbb\{E\}\\left\[\\\|\\sigma\(Z\_\{J\}\)\\\|^\{2\}\\,\\Big\|\\,D\_\{S\_\{A\}\\cup S\_\{B\}\}\\right\]
###### Theorem 3\.4\(Asymptotic Expansion of Rectified Accumulation\)\.
Letν2=‖αA‖2\+‖αB‖2\\nu^\{2\}=\\\|\\alpha\_\{A\}\\\|^\{2\}\+\\\|\\alpha\_\{B\}\\\|^\{2\}be the signal energy andρ=αATGABαB\\rho=\\alpha\_\{A\}^\{T\}G\_\{AB\}\\alpha\_\{B\}be the cross\-correlation\. Define the effective interference varianceζ2=μ2\(kA\+kB\)\+2μρ\\zeta^\{2\}=\\mu^\{2\}\(k\_\{A\}\+k\_\{B\}\)\+2\\mu\\rho\. For a bias thresholdβ\>0\\beta\>0, the expected spurious energy on a ghost featuredjd\_\{j\}admits the following lower bound expansion:
𝔼\[σ\(⟨z,dj⟩\)2\]\\displaystyle\\mathbb\{E\}\[\\sigma\(\\langle z,d\_\{j\}\\rangle\)^\{2\}\]≥ζ22\[\(ζ2πβ−ζ32πβ3\)exp\(−β22ζ2\)\+ℛrem\(β,ζ\)\]\\displaystyle\\geq\\frac\{\\zeta^\{2\}\}\{2\}\\left\[\\left\(\\frac\{\\zeta\}\{\\sqrt\{2\\pi\}\\beta\}\-\\frac\{\\zeta^\{3\}\}\{\\sqrt\{2\\pi\}\\beta^\{3\}\}\\right\)\\exp\\left\(\-\\frac\{\\beta^\{2\}\}\{2\\zeta^\{2\}\}\\right\)\+\\mathcal\{R\}\_\{rem\}\(\\beta,\\zeta\)\\right\]\(3\)\+∑p=2∞\(−1\)p\(2p−1\)\!\!β2p\+1∫ℳk⟨u,dj⟩2p𝑑π\(u\)⏟Higher\-order Geometry Terms\\displaystyle\\quad\+\\underbrace\{\\sum\_\{p=2\}^\{\\infty\}\\frac\{\(\-1\)^\{p\}\(2p\-1\)\!\!\}\{\\beta^\{2p\+1\}\}\\int\_\{\\mathcal\{M\}\_\{k\}\}\\langle u,d\_\{j\}\\rangle^\{2p\}d\\pi\(u\)\}\_\{\\text\{Higher\-order Geometry Terms\}\}\(4\)whereℛrem\\mathcal\{R\}\_\{rem\}is the remainder term from the Mills’ ratio approximation, satisfying\|ℛrem\|≤O\(ζ5/β5\)\|\\mathcal\{R\}\_\{rem\}\|\\leq O\(\\zeta^\{5\}/\\beta^\{5\}\)\.
###### Proof 3\.5\.
We model the projectionX=⟨z,dj⟩X=\\langle z,d\_\{j\}\\rangleas a sub\-Gaussian variable with varianceζ2\\zeta^\{2\}\. The quantity of interest is the second moment of the rectified tail:
M2=∫β∞\(x−β\)2p\(x\)𝑑xM\_\{2\}=\\int\_\{\\beta\}^\{\\infty\}\(x\-\\beta\)^\{2\}p\(x\)dxUsing the substitutionx=β\+tx=\\beta\+tand the Gaussian densityp\(x\)=12πζe−x2/2ζ2p\(x\)=\\frac\{1\}\{\\sqrt\{2\\pi\}\\zeta\}e^\{\-x^\{2\}/2\\zeta^\{2\}\}, we expand the exponent:
M2=e−β2/2ζ22πζ∫0∞t2exp\(−t22ζ2−βtζ2\)𝑑tM\_\{2\}=\\frac\{e^\{\-\\beta^\{2\}/2\\zeta^\{2\}\}\}\{\\sqrt\{2\\pi\}\\zeta\}\\int\_\{0\}^\{\\infty\}t^\{2\}\\exp\\left\(\-\\frac\{t^\{2\}\}\{2\\zeta^\{2\}\}\-\\frac\{\\beta t\}\{\\zeta^\{2\}\}\\right\)dt\(5\)This integral does not admit a closed form in elementary functions\. However, assuming the “high\-bias regime“ \(β≫ζ\\beta\\gg\\zeta\), we perform iterative integration by parts\. LetIk=∫0∞tke−βt/ζ2e−t2/2ζ2𝑑tI\_\{k\}=\\int\_\{0\}^\{\\infty\}t^\{k\}e^\{\-\\beta t/\\zeta^\{2\}\}e^\{\-t^\{2\}/2\\zeta^\{2\}\}dt\. We approximatee−t2/2ζ2≈∑\(−1\)nt2nn\!\(2ζ2\)ne^\{\-t^\{2\}/2\\zeta^\{2\}\}\\approx\\sum\\frac\{\(\-1\)^\{n\}t^\{2n\}\}\{n\!\(2\\zeta^\{2\}\)^\{n\}\}\. The dominant term corresponds to the breakdown of orthogonality\. Unlike the linear case where𝔼\[X\]=0\\mathbb\{E\}\[X\]=0, the rectified expectation is strictly positive:
𝔼\[σ\(X\)\]≈ζ2πe−β2/2ζ2\(1−ζ2β2\+3ζ4β4\)\\mathbb\{E\}\[\\sigma\(X\)\]\\approx\\frac\{\\zeta\}\{\\sqrt\{2\\pi\}\}e^\{\-\\beta^\{2\}/2\\zeta^\{2\}\}\\left\(1\-\\frac\{\\zeta^\{2\}\}\{\\beta^\{2\}\}\+\\frac\{3\\zeta^\{4\}\}\{\\beta^\{4\}\}\\right\)\(6\)The variance termζ2\\zeta^\{2\}contains2μρ2\\mu\\rhoviav\(ρ\)=σ02\+2μρv\(\\rho\)=\\sigma\_\{0\}^\{2\}\+2\\mu\\rho\. While𝔼\[ρ\]=0\\mathbb\{E\}\[\\rho\]=0, rectification yieldsD\(ρ\):=𝔼\[σ\(X\)∣v\(ρ\)\]=v\(ρ\)/2πD\(\\rho\):=\\mathbb\{E\}\[\\sigma\(X\)\\mid v\(\\rho\)\]=\\sqrt\{v\(\\rho\)\}/\\sqrt\{2\\pi\}, which is strictly convex inρ\\rho; thus by Jensen,
𝔼ρ\[D\(ρ\)\]\>D\(0\)=σ02π\.\\mathbb\{E\}\_\{\\rho\}\[D\(\\rho\)\]\>D\(0\)=\\frac\{\\sigma\_\{0\}\}\{\\sqrt\{2\\pi\}\}\.Hence alignment fluctuations strictly increase the expected rectified drift relative to an isotropic baseline\.
#### Physical Interpretation\.
The expansion in Eq\. \(6\) reveals a “Ratchet Mechanism“\. The first term represents the classic Gaussian tail probability\. The second term \(infinite sum\) couples the activation probability to the higher\-order moments of the manifold geometry\. Specifically, the term∫⟨u,dj⟩2p\\int\\langle u,d\_\{j\}\\rangle^\{2p\}measures how “spiky“ the signal cone is\. Even if the dictionary is incoherent on average \(μ≈0\\mu\\approx 0\), the presence of local geometric spikes \(captured byp≥2p\\geq 2\) significantly thickens the tail of the interference distribution, making “safe“ linear composition impossible in practice\.
## 4Macro\-Analysis: The Phase Transition of Compositional Collapse
Building on the microscopic mechanism of rectified interference, we now analyze the system’s macroscopic stability\. While the Ratchet Effect explains individual feature activation, global stability depends on the collective geometry of allN=m−kN=m\-kghost features\. These pairwise constraints define the facets of a high\-dimensional polyhedral cone; through measure concentration, we prove that their cumulative interaction manifests as a concentration\-induced phase transition—a geometric limit beyond which accurate steering becomes impossible\.
### 4\.1The Geometry of Feasibility
We analyze the asymptotic regimen→∞n\\to\\inftywith fixed aspect ratioδ=m/n\\delta=m/nand compositional densityγ=k/n\\gamma=k/n\. To rigorously characterize the steerability limit, we adopt the framework of Conic Integral Geometry\(amelunxen2014living\)\.
###### Definition 4\.1\(Signal and Interference Cones\)\.
LetD∈ℝn×mD\\in\\mathbb\{R\}^\{n\\times m\}be the dictionary andSSbe the active support set\.
1. 1\.TheSignal Cone𝒞sig\\mathcal\{C\}\_\{sig\}is the conic hull of the active semantic directions:𝒞sig:=cone\(\{di\}i∈S\)\.\\mathcal\{C\}\_\{sig\}:=\\text\{cone\}\(\\\{d\_\{i\}\\\}\_\{i\\in S\}\)\.
2. 2\.TheGhost Cone𝒞ghost\\mathcal\{C\}\_\{ghost\}is the conic hull of the inactive directions:𝒞ghost:=cone\(\{dj\}j∉S\)\.\\mathcal\{C\}\_\{ghost\}:=\\text\{cone\}\(\\\{d\_\{j\}\\\}\_\{j\\notin S\}\)\.
Followingamelunxen2014living, the complexity of these cones is measured by their statistical dimensionδ\(𝒞\)\\delta\(\\mathcal\{C\}\)\. For random dictionaries, we have the standard approximationsδ\(𝒞sig\)≈k\\delta\(\\mathcal\{C\}\_\{sig\}\)\\approx kandδ\(𝒞ghost∘\)≈2Nlog\(enN2π\)−1n\\delta\(\\mathcal\{C\}\_\{ghost\}^\{\\circ\}\)\\approx 2N\\log\(\\frac\{en\}\{N\\sqrt\{2\\pi\}\}\)^\{\-1\}n, which provide the basis for our parametersΨ\\PsiandΦ\\Phi\.
###### Definition 4\.2\(The Separation Condition\)\.
A composition is feasible if and only if the signal cone intersects the safe region defined by the polar of the ghost cone\. Mathematically, there must exist a separation vectorh∈𝒞sig∩𝒞ghost∘h\\in\\mathcal\{C\}\_\{sig\}\\cap\\mathcal\{C\}\_\{ghost\}^\{\\circ\}such that:
𝒞sig∩𝒞ghost∘≠\{0\}\\mathcal\{C\}\_\{sig\}\\cap\\mathcal\{C\}\_\{ghost\}^\{\\circ\}\\neq\\\{0\\\}\(7\)where𝒞ghost∘=\{y∣∀x∈𝒞ghost,⟨y,x⟩≤0\}\\mathcal\{C\}\_\{ghost\}^\{\\circ\}=\\\{y\\mid\\forall x\\in\\mathcal\{C\}\_\{ghost\},\\langle y,x\\rangle\\leq 0\\\}is the region where all spurious features are suppressed\.
\(a\) Stable Regime \(γ<γ∗\\gamma<\\gamma^\{\*\}\)𝒦J∘\\mathcal\{K\}\_\{J\}^\{\\circ\}zz𝒦S\\mathcal\{K\}\_\{S\}SeparationExists\(b\) Compositional Collapse \(γ\>γ∗\\gamma\>\\gamma^\{\*\}\)zzCOLLISIONFigure 2:Geometry of Compositional Separation\.\(a\)Stable regime: The signal cone𝒦S\\mathcal\{K\}\_\{S\}\(blue\) is disjoint from the ghost polar cone𝒦J∘\\mathcal\{K\}\_\{J\}^\{\\circ\}\(red\)\.\(b\)Collapse: As density increases,𝒦S\\mathcal\{K\}\_\{S\}widens and collides with the ghost constraints, triggering the phase transition\.If this intersection contains only the origin, any attempt to activateSSwill inevitably “spill over“ intoJJ, causing collapse\.
### 4\.2The Kinematic Threshold and Statistical Dimension
To determine the intersection probability, we utilize the Statistical Dimensionδ\(𝒞\)\\delta\(\\mathcal\{C\}\), which generalizes the subspace dimension to convex cones\. It is defined via the Gaussian Mean Width:δ\(𝒞\)≜𝔼g\[‖Π𝒞\(g\)‖2\]\\delta\(\\mathcal\{C\}\)\\triangleq\\mathbb\{E\}\_\{g\}\[\\\|\\Pi\_\{\\mathcal\{C\}\}\(g\)\\\|^\{2\}\]\.
To address parametric consistency \(Reviewer Comment 2\), we explicitly map the statistical dimension to the sparsity densityγ\\gamma\. For a random cone generated byk=γnk=\\gamma nvectors, the normalized statistical dimension satisfiesΨ\(γ\)≜δ\(𝒦S\)/n\\Psi\(\\gamma\)\\triangleq\\delta\(\\mathcal\{K\}\_\{S\}\)/n\.
### 4\.3Derivation via Gordon’s Escape Theorem
Directly computing the intersection probability of the high\-dimensional cones defined in Section[4\.1](https://arxiv.org/html/2605.05223#S4.SS1)is geometrically intractable\. To rigorously bridge this gap, we employ Gordon’s Escape Through a Mesh Theorem\(gordon1988milman\), which maps the geometric intersection problem to a comparison of Gaussian processes\.
###### Proposition 4\.3\(Geometric Saturation Condition\)\.
Notation\. LetΨ:=δ\(𝒞sig\)/n\\Psi:=\\delta\(\\mathcal\{C\}\_\{sig\}\)/nbe the normalized statistical dimension of the signal cone, andΦ:=δ\(𝒞ghost∘\)/n\\Phi:=\\delta\(\\mathcal\{C\}\_\{ghost\}^\{\\circ\}\)/nbe that of the polar ghost cone\. We define the geometric stability margin asΔgap=1−\(Ψ\+Φ\)2\\Delta\_\{gap\}=1\-\(\\sqrt\{\\Psi\}\+\\sqrt\{\\Phi\}\)^\{2\}\.
A geometric phase transition occurs when the sum of the statistical dimensions saturates the ambient spacenn\. Mathematically, the exact condition for the phase boundary is:
δ\(𝒞sig\)\+δ\(𝒞ghost∘\)=n⏟Geometric Saturation⇔Ψ\(γ\)\+Φ\(δ−γ\)=1⏟Analytical Boundary\.\\underbrace\{\\delta\(\\mathcal\{C\}\_\{sig\}\)\+\\delta\(\\mathcal\{C\}\_\{ghost\}^\{\\circ\}\)=n\}\_\{\\text\{Geometric Saturation\}\}\\quad\\iff\\quad\\underbrace\{\\Psi\(\\gamma\)\+\\Phi\(\\delta\-\\gamma\)=1\}\_\{\\text\{Analytical Boundary\}\}\.\(8\)
###### Corollary 4\.4\(Heuristic Geometric Widening\)\.
The accumulated rectified driftηtotal≈kη\(ρ\)\\eta\_\{total\}\\approx\\sqrt\{k\}\\eta\(\\rho\)induces a shift in the effective center of the interference distribution\. Invoking the Lipschitz continuity of the Gaussian Mean Widthw\(⋅\)w\(\\cdot\), we approximate the widening of the ghost cone as:
w\(𝒞ghosteff\)≲w\(𝒞ghost\)\+‖drift‖≈w\(𝒞ghost\)\+k𝔼\[η\(ρbil\)\]\.w\(\\mathcal\{C\}\_\{ghost\}^\{eff\}\)\\lesssim w\(\\mathcal\{C\}\_\{ghost\}\)\+\\\|\\text\{drift\}\\\|\\approx w\(\\mathcal\{C\}\_\{ghost\}\)\+\\sqrt\{k\}\\mathbb\{E\}\[\\eta\(\\rho\_\{\\text\{bil\}\}\)\]\.\(9\)This effective widening quantifiably consumes the stability marginΔgap\\Delta\_\{gap\}, bridging the microscopic ReLU drift to the macroscopic phase transition\.
#### Gaussian Process Construction\.
To prove that Eq\. \([8](https://arxiv.org/html/2605.05223#S4.E8)\) governs the collapse, consider two Gaussian processes indexed by the unit vectors\(u,v\)∈\(𝒞sig∩𝕊n−1\)×\(𝒞ghost∩𝕊n−1\)\(u,v\)\\in\(\\mathcal\{C\}\_\{sig\}\\cap\\mathbb\{S\}^\{n\-1\}\)\\times\(\\mathcal\{C\}\_\{ghost\}\\cap\\mathbb\{S\}^\{n\-1\}\):
𝒳u,v\\displaystyle\\mathcal\{X\}\_\{u,v\}=⟨u,Dghostv⟩\\displaystyle=\\langle u,D\_\{ghost\}v\\rangle\(10\)𝒴u,v\\displaystyle\\mathcal\{Y\}\_\{u,v\}=‖u‖⟨g,v⟩\+‖v‖⟨h,u⟩\\displaystyle=\\\|u\\\|\\langle g,v\\rangle\+\\\|v\\\|\\langle h,u\\rangle\(11\)whereg∼𝒩\(0,Im\)g\\sim\\mathcal\{N\}\(0,I\_\{m\}\)andh∼𝒩\(0,In\)h\\sim\\mathcal\{N\}\(0,I\_\{n\}\)\. The “collapse” event—where the signal cone fails to be separated from the ghost cone—corresponds to the primary process𝒳\\mathcal\{X\}crossing zero \(i\.e\., non\-empty intersection\)\.
Gordon’s Comparison Inequality asserts thatℙ\(min𝒳u,v≥0\)≤ℙ\(min𝒴u,v≥0\)\\mathbb\{P\}\(\\min\\mathcal\{X\}\_\{u,v\}\\geq 0\)\\leq\\mathbb\{P\}\(\\min\\mathcal\{Y\}\_\{u,v\}\\geq 0\)\. This simplifies the complex geometric interaction into a decoupled condition:
minu,v\(‖u‖gv\+‖v‖hu\)≥0\\min\_\{u,v\}\\left\(\\\|u\\\|g\_\{v\}\+\\\|v\\\|h\_\{u\}\\right\)\\geq 0Solving this inequality allows us to bound the width of the interference setΦ\\Phivia the Gaussian tail integral, directly yielding the explicit threshold in Theorem[4\.5](https://arxiv.org/html/2605.05223#S4.Thmtheorem5)\.
### 4\.4Main Result: The Explicit Critical Threshold
We now state the explicit form of the phase transition derived from the Gaussian process comparison\.
###### Theorem 4\.5\(Explicit Phase Transition Threshold\)\.
Notation\. LetΨ\(γ\):=δ\(𝒞sig\)/n\\Psi\(\\gamma\):=\\delta\(\\mathcal\{C\}\_\{sig\}\)/nbe the normalized statistical dimension of the signal cone, andΦ\(δ−γ\):=δ\(𝒞ghost∘\)/n\\Phi\(\\delta\-\\gamma\):=\\delta\(\\mathcal\{C\}\_\{ghost\}^\{\\circ\}\)/nbe that of the polar ghost cone\. Under the random spherical dictionary assumption, the stability boundary is strictly characterized by the geometric saturation condition:
Ψ\(γ\)\+Φ\(δ−γ\)=1\.\\Psi\(\\gamma\)\+\\Phi\(\\delta\-\\gamma\)=1\.\(12\)Using the standard approximationΨ\(γ\)≈γ\\Psi\(\\gamma\)\\approx\\gammaand the polyhedral bound forΦ\\Phi, the critical densityγ∗\\gamma^\{\*\}satisfies the explicit scaling law:
γ∗\+2\(δ−γ∗\)log\(1Δgap\)=1\+𝒪\(n−1/2\)\\sqrt\{\\gamma^\{\*\}\}\+\\sqrt\{2\(\\delta\-\\gamma^\{\*\}\)\\log\\left\(\\frac\{1\}\{\\Delta\_\{gap\}\}\\right\)\}=1\+\\mathcal\{O\}\(n^\{\-1/2\}\)\(13\)whereΔgap\\Delta\_\{gap\}represents the stability margin discussed below\.
#### Rigorous Basis via Gordon’s Inequality\.
The threshold condition above is rigorously grounded in Gordon’s Escape Through a Mesh Theorem\. Followingamelunxen2014living, the exact condition for stable recovery isδ\(𝒞sig\)\+δ\(𝒞ghost∘\)≤n\\delta\(\\mathcal\{C\}\_\{sig\}\)\+\\delta\(\\mathcal\{C\}\_\{ghost\}^\{\\circ\}\)\\leq n\. Our result identifiesγ∗\\gamma^\{\*\}as the critical point where the probabilistic measure of the conic intersection vanishes\.
Mechanism and Consistency\.The termlog\(1/Δgap\)\\log\(1/\\Delta\_\{gap\}\)emerges from the statistical dimension of the polyhedral ghost cone\. As derived in Appendix C, for dictionary sparsityδ\\delta, the geometric margin scales asΔgap−1≈e\(δ−γ∗\)/2π\\Delta\_\{gap\}^\{\-1\}\\approx e\(\\delta\-\\gamma^\{\*\}\)/\\sqrt\{2\\pi\}, linking the abstract safety probability to the explicit facet density of the dictionary\.
#### Interpretation\.
Eq\. \([13](https://arxiv.org/html/2605.05223#S4.E13)\) formalizes the fundamental conflict between signal capacity and interference avoidance\. The first term,γ∗\\sqrt\{\\gamma^\{\*\}\}, represents the effective radius of the active signal cone, while the second term captures the exclusion width enforced by theNNinactive features\. The phase transition occurs precisely when the sum of these two geometric forces saturates the ambient dimensionality\. Beyond this limit, the intersection between the signal cone and the safe region becomes empty almost surely, triggering an irreversible compositional collapse\.
### 4\.5Empirical Validation: Phase Transition in CLEVR
While our derivation assumes a random spherical dictionary, real\-world semantic features exhibit structural correlations\. To validate the robustness of Theorem[4\.5](https://arxiv.org/html/2605.05223#S4.Thmtheorem5)in a structured regime, we perform controlled steering experiments using the CLEVR dataset, utilizing logic and attribute features extracted from Qwen\-VL\.
#### Synthetic Validation Strategy\.
To rigorously isolate geometric effects from semantic correlations, we employ a dual\-validation strategy\. Throughout our result figures, the Theoretical Baseline \(dashed lines\) serves as the Synthetic Random Dictionary reference\. This curve is computed directly from the phase transition equation \(Eq\.[13](https://arxiv.org/html/2605.05223#S4.E13)\) under the isotropic assumption, representing the ideal performance of unstructured, i\.i\.d\. features\. Consequently, the gap between this synthetic baseline and the empirical CLEVR curves quantifies the specific impact of semantic structure, directly addressing the need to disentangle geometric constraints from data\-specific correlations\.
#### Results\.
We compare our theoretical predictions against empirical measurements\. As shown in Figure[4](https://arxiv.org/html/2605.05223#A4.F4)\(see Appendix E\), the theoretical phase boundary derived in Eq \([13](https://arxiv.org/html/2605.05223#S4.E13)\) closely matches the empirical drift observed in CLEVR features\.
To ensure statistical significance, each data point represents the average of 50 independent trials with randomized dictionary initializations\. The error bars \(shaded regions\) indicate the standard deviation\. Notably, the drift observed near the critical thresholdγ∗\\gamma^\{\*\}is statistically significant \(p<0\.01p<0\.01via t\-test\) compared to the baseline noise, confirming that the collapse is structural rather than accidental\.
We observe three key phenomena in our empirical analysis: \(i\) a concentration\-induced phase transition in representation stability exists such that below a critical densityγ<γ∗\\gamma<\\gamma^\{\*\}, spurious energy remains negligible \(ℰspur<10−3\\mathcal\{E\}\_\{spur\}<10^\{\-3\}\), enabling precise linear control; \(ii\) a correlation\-induced shift occurs where the empirical transitionγCLEVR∗\\gamma\_\{CLEVR\}^\{\*\}is lower than the isotropic predictionγtheory∗\\gamma\_\{theory\}^\{\*\}, confirming that positive semantic correlation \(ρ\>0\\rho\>0\) thickens the signal cone and triggers earlier collapse; and \(iii\) a regime of geometric irreducibility is reached beyond this threshold, where further compositional attempts lead to a total loss of semantic separation, proving the barrier is a fundamental geometric constraint rather than a stochastic artifact\.
00\.50\.51100\.50\.511Interference Strengthζ2\\zeta^\{2\}Mean Driftη\\eta\(a\) Rectification DriftLinearReLU
00\.20\.20\.40\.40\.60\.60\.80\.800\.50\.511Densityγ\\gammaSpurious Energyℰspur\\mathcal\{E\}\_\{spur\}\(b\) Phase TransitionSimulatedTheoryγ∗\\gamma^\{\*\}
Figure 3:Empirical validation of the Ratchet Mechanism\.\(a\) ReLU rectifies stochastic interference into systematic biasη\\eta\. \(b\) Spurious energy undergoes an*abrupt transition*as compositional densityγ\\gammaapproaches the thresholdγ∗\\gamma^\{\*\}\.
## 5Implications for Mechanistic Interpretability
Our theoretical findings provide a rigorous geometric foundation for interpreting the “anomalous“ behaviors observed in recent empirical studies of Large Vision\-Language Models \(LVLMs\)\. Specifically, we show that the fragility of feature steering and the necessity of chain\-of\-thought are not mere training artifacts, but direct corollaries of the Phase Transition Theorem \([4\.5](https://arxiv.org/html/2605.05223#S4.Thmtheorem5)\)\.
### 5\.1Union Failure as Local Isometry Breaking
Recent work\(tan2024steering;abreu2025conceptors\)reports a “Union Failure“ phenomenon: while steering with a single feature vectorvAv\_\{A\}can be stable, the compositionvA\+vBv\_\{A\}\+v\_\{B\}often induces “Output Drift“η\\eta, even when‖vA\+vB‖\\\|v\_\{A\}\+v\_\{B\}\\\|is normalized\.
We reinterpret this as a breakdown of the Local Restricted Isometry Property \(RIP\)\. LetΦS∈ℝn×k\\Phi\_\{S\}\\in\\mathbb\{R\}^\{n\\times k\}be the sub\-dictionary of active features\. For linear stability, the mapping from latent code to residual stream must be well\-conditioned, requiringκ\(ΦSTΦS\)≈1\\kappa\(\\Phi\_\{S\}^\{T\}\\Phi\_\{S\}\)\\approx 1\. However, our micro\-analysis of the entanglement spectrum implies that the condition number diverges asymptotically as the system approaches the phase boundaryγ∗\\gamma^\{\*\}\.
###### Proposition 5\.1\(Spectral Divergence and Condition Number\)\.
LetΦ\\Phibe then×kn\\times kmatrix of active features\. Under Assumption[1](https://arxiv.org/html/2605.05223#Thmassumption1), asn,k→∞n,k\\to\\inftywithk/n→γk/n\\to\\gamma, the singular values ofΦ\\Phifollow the Marchenko\-Pastur distribution\. The condition numberκ=σmax/σmin\\kappa=\\sigma\_\{max\}/\\sigma\_\{min\}satisfies:
κ\(Φ\)→𝑝1\+γ1−γforγ<1\\kappa\(\\Phi\)\\xrightarrow\{p\}\\frac\{1\+\\sqrt\{\\gamma\}\}\{1\-\\sqrt\{\\gamma\}\}\\quad\\text\{for \}\\gamma<1\(14\)As the density approaches the geometric critical pointγ→γ∗\\gamma\\to\\gamma^\{\*\}, the effective interference increases the ”virtual rank” of the system, leading to a divergenceκ→∞\\kappa\\to\\inftycharacterized by the proximity to the phase boundary\.
#### Mechanism\.
As the compositional density approachesγ∗\\gamma^\{\*\}, the smallest singular valueσmin\(ΦA∪B\)\\sigma\_\{\\min\}\(\\Phi\_\{A\\cup B\}\)becomes arbitrarily small, rendering the inverse mappingz↦\(Φ⊤Φ\)−1Φ⊤zz\\mapsto\(\\Phi^\{\\top\}\\Phi\)^\{\-1\}\\Phi^\{\\top\}znumerically unstable\. The drift observed in experiments corresponds to the projection residue onto the near\-null space of this ill\-conditioned basis\. This indicates that “Output Drift” arises from the model minimizing energy along directions that are no longer jointly realizable\.
### 5\.2Resolution: MLPs as Coherence Reset Operators
Theorem[4\.5](https://arxiv.org/html/2605.05223#S4.Thmtheorem5)defines the capacity of a fixed\-width residual stream\. However, the ”Ratchet” effect \(Theorem[3\.3](https://arxiv.org/html/2605.05223#S3.Thmtheorem3)\) implies that interference should grow monotonically with depth, paradoxically suggesting that deeper models should collapse faster\. We resolve this paradox by distinguishing thepassive accumulationof the residual stream from theactive error\-correctionof MLP blocks\.
We hypothesize that MLP layers function asCoherence Reset Operators\. While the residual stream widthnnis fixed, MLPs project the state into a hyperspaceℝdff\\mathbb\{R\}^\{d\_\{ff\}\}\(wheredff≫nd\_\{ff\}\\gg n\)\. InvokingCover’s Theorem, patterns that are entangled in the crowded ambient spaceℝn\\mathbb\{R\}^\{n\}become linearly separable inℝdff\\mathbb\{R\}^\{d\_\{ff\}\}\. Crucially, the non\-linearityσ\(⋅\)\\sigma\(\\cdot\)operates in this expanded regime, acting as a geometric filter that suppresses the ”ghost cones” \(which shrink relatively in high dimensions\) before projecting the cleaned signal back toℝn\\mathbb\{R\}^\{n\}\. This process effectively ”resets” the interference budgetη\\etaat each layer, allowing the model to sustain compositional depth despite the constant entropic pressure of the Ratchet\.
## 6Discussion
Our analysis establishes a hard geometric limit on the compositionality of sparse representations, where “feature addition” is valid only within a constrained “Safety Region” defined byγ∗\\gamma^\{\*\}\. We now address the practical implications of this phase transition regarding semantic structure, scaling, and serialization\.
### 6\.1Semantic Structure Accelerates Collapse
While our baseline relies on random spherical dictionaries, real\-world datasets \(e\.g\., CLEVR\) exhibit semantic clustering where local coherenceμlocal≫1/n\\mu\_\{local\}\\gg 1/\\sqrt\{n\}\. Our experiments confirm that such structure exacerbates instability: semantic correlations effectively thicken the signal cone, making the “Safety Region” tighter and the transition sharper than Theorem[4\.5](https://arxiv.org/html/2605.05223#S4.Thmtheorem5)predicts\. Consequently, random\-matrix bounds represent an optimistic upper bound; real\-world steering is inherently more fragile due to reduced effective dimensionality\.
### 6\.2The Steerability Scaling Law
Theorem[4\.5](https://arxiv.org/html/2605.05223#S4.Thmtheorem5)implies a linear scaling law for steerability: the maximum capacity of stable active conceptskmaxk\_\{max\}is constrained by the residual stream widthnnsuch thatkmax≈γ∗\(δ\)⋅nk\_\{max\}\\approx\\gamma^\{\*\}\(\\delta\)\\cdot n\. This provides a geometric resolution to the “width vs\. depth” debate: while depth facilitates topological decoupling, width governs the instantaneous semantic budget\. This perspective formalizes the Alignment Tax; injecting safety steering vectors consumes the available interference budget\. If aggressive steering pushes densityγ\\gammabeyondγ∗\\gamma^\{\*\}, the model collapses not due to behavioral refusal, but because it has exhausted its orthogonal geometric space\. At this saturation limit, signal energy spills into the ghost cone, triggering spurious activations that destroy output coherence\.
### 6\.3Geometric Origin of Chain\-of\-Thought \(CoT\)
Compositional collapse renders serialization a topological necessity\. If a task requires\|Stotal\|/n\>γ∗\|S\_\{total\}\|/n\>\\gamma^\{\*\}, representing the solution in a single latent state is impossible\. To avoid collapse, the system must trade space for time, decomposingStotalS\_\{total\}into a sequenceS1,…,STS\_\{1\},\\dots,S\_\{T\}where each\|St\|/n<γ∗\|S\_\{t\}\|/n<\\gamma^\{\*\}\. CoT thus acts as a mechanism to navigate phase boundaries: by externalizing intermediate states, the model “flushes” its activation buffer and resets the interference budget per step\. Serialization circumvents the “Ratchet” by offloading data to the context window, where orthogonality is enforced by positional encoding rather than random chance\.
## 7Conclusion
This work rigorously formalizes the breakdown of linear compositionality through the framework of Conic Integral Geometry\. By mapping latent activations to the intersection of high\-dimensional convex cones, we derive a concentration\-induced phase boundary beyond which interference in overcomplete dictionaries becomes a self\-amplifying process\. This “Geometric Impossibility” for naive steering identifies a fundamental limit of the Linear Representation Hypothesis, showing that arbitrary vector stacking can trigger catastrophic collapse once this boundary is exceeded\. Our results further suggest that, for fixed\-width architectures, depth and serialization \(e\.g\., Chain\-of\-Thought\) are not merely training artifacts, but essential geometric resources for maintaining structural stability within a crowded residual stream\.
\\acks
We thank a bunch of people and funding agency\.
## References
## Appendix AExtended Preliminaries and Notations
In this section, we provide a comprehensive summary of the mathematical notations used throughout the paper, followed by formal definitions of the high\-dimensional geometric concepts that underpin our main results\.
### A\.1Summary of Notations
Table[1](https://arxiv.org/html/2605.05223#A1.T1)summarizes the key symbols, their dimensions, and their roles in the analysis\. We explicitly distinguish between the scalar overcompletenessδ\\deltaand the statistical dimension functionalδ\(⋅\)\\delta\(\\cdot\), as well as the signal densityγ\\gammaand its critical thresholdγ∗\\gamma^\{\*\}\.
SymbolSpaceDescriptionModel Dimensions & Parametersnnℕ\\mathbb\{N\}Dimension of the residual stream \(embedding size\)\.mmℕ\\mathbb\{N\}Number of dictionary features \(atoms\),m≫nm\\gg n\.kkℕ\\mathbb\{N\}Number of simultaneously active features \(sparsity\)\.δ\\deltaℝ\+\\mathbb\{R\}\_\{\+\}Overcompleteness ratio, defined asδ=m/n\\delta=m/n\.γ\\gammaℝ\+\\mathbb\{R\}\_\{\+\}Compositional density, defined asγ=k/n\\gamma=k/n\.γ∗\\gamma^\{\*\}ℝ\+\\mathbb\{R\}\_\{\+\}Critical phase transition threshold\(The ”Event Horizon”\)\.S,JS,JSetsIndex sets for active features \(SS\) and ghost features \(JJ\)\.Geometry & OperatorsDDℝn×m\\mathbb\{R\}^\{n\\times m\}The feature dictionary \(with unit\-norm columns\)\.𝒦S\\mathcal\{K\}\_\{S\}ℝn\\mathbb\{R\}^\{n\}The convex cone spanned by active features:cone\(\{di\}i∈S\)\\text\{cone\}\(\\\{d\_\{i\}\\\}\_\{i\\in S\}\)\.𝒦J∘\\mathcal\{K\}\_\{J\}^\{\\circ\}ℝn\\mathbb\{R\}^\{n\}The polar cone of the ghost featuresJJ\.Π𝒞\\Pi\_\{\\mathcal\{C\}\}OpEuclidean projection operator onto a convex set𝒞\\mathcal\{C\}\.δ\(𝒞\)\\delta\(\\mathcal\{C\}\)ℝ\+\\mathbb\{R\}\_\{\+\}Statistical dimensionof a convex cone𝒞\\mathcal\{C\}\.w\(𝒞\)w\(\\mathcal\{C\}\)ℝ\+\\mathbb\{R\}\_\{\+\}Gaussian mean width of a set𝒞\\mathcal\{C\}\.Ψ\(⋅\)\\Psi\(\\cdot\)Func\.Normalized dimension scaling function \(e\.g\.,Ψ\(ρ\)\\Psi\(\\rho\)\)\.Statistics & Interferenceμ\(D\)\\mu\(D\)ℝ\\mathbb\{R\}Mutual coherence of the dictionary:maxi≠j\|⟨di,dj⟩\|\\max\_\{i\\neq j\}\|\\langle d\_\{i\},d\_\{j\}\\rangle\|\.ρ\\rhoℝ\\mathbb\{R\}Semantic cross\-correlation \(alignment\) between subspaces\.ζj2\\zeta\_\{j\}^\{2\}ℝ\+\\mathbb\{R\}\_\{\+\}Variance of interference on a specific ghost featurejj\.η\\etaℝ\+\\mathbb\{R\}\_\{\+\}Rectified drift term\(The Ratchet shift\)\.ℰspur\\mathcal\{E\}\_\{spur\}ℝ\+\\mathbb\{R\}\_\{\+\}Spurious energy \(magnitude of projection onto ghost subspace\)\.Table 1:Summary of mathematical notations\. Special attention is drawn to the distinction between the ratio parameters \(δ,γ\\delta,\\gamma\) and the geometric functionals \(δ\(⋅\)\\delta\(\\cdot\)\) or thresholds \(γ∗\\gamma^\{\*\}\)\.
### A\.2Formal Geometric Definitions
While the main text provides intuitive descriptions, here we provide the rigorous definitions necessary for the proofs in Appendix[C](https://arxiv.org/html/2605.05223#A3)\.
###### Definition A\.1\(Gaussian Mean Width\)\.
The Gaussian Mean Width of a bounded subsetT⊂ℝnT\\subset\\mathbb\{R\}^\{n\}is defined as the expected supremum of a Gaussian process indexed byTT:
w\(T\)≜𝔼g∼𝒩\(0,In\)\[supx∈T⟨g,x⟩\]\.w\(T\)\\triangleq\\mathbb\{E\}\_\{g\\sim\\mathcal\{N\}\(0,I\_\{n\}\)\}\\left\[\\sup\_\{x\\in T\}\\langle g,x\\rangle\\right\]\.\(15\)This quantity characterizes the “effective size” ofTTand is linked to the statistical dimension viaδ\(T\)≈w\(T∩𝕊n−1\)2\\delta\(T\)\\approx w\(T\\cap\\mathbb\{S\}^\{n\-1\}\)^\{2\}\.
Mechanism and Consistency\.The termlog\(1/Δgap\)\\log\(1/\\Delta\_\{gap\}\)in Eq\. \(13\) emerges from Gordon’s Escape Through a Mesh Theorem, whereΔgap\\Delta\_\{gap\}scales with the failure probabilityη\\etaof the stable intersection\. For the specific polyhedral geometry of SAE features, the statistical dimensionΦ\\Phiasymptotically resolves toΦ≈2Nnlog\(nN2π\)\\Phi\\approx\\frac\{2N\}\{n\}\\log\(\\frac\{n\}\{N\}\\sqrt\{2\\pi\}\), explicitly yielding the logarithmic dependency\. Thus, the abstract marginΔgap\\Delta\_\{gap\}and the explicit geometric volume derived in Appendix C\.4 \(Δgap−1≈e\(δ−γ\)/2π\\Delta\_\{gap\}^\{\-1\}\\approx e\(\\delta\-\\gamma\)/\\sqrt\{2\\pi\}\) are asymptotically identical, providing a rigorous bridge between our general framework and the polyhedral construction\.
###### Definition A\.2\(Statistical Dimension\)\.
The Statistical Dimensionδ\(𝒞\)\\delta\(\\mathcal\{C\}\)of a closed convex cone𝒞⊂ℝn\\mathcal\{C\}\\subset\\mathbb\{R\}^\{n\}extends the concept of dimension to non\-linear cones\. It is defined as the expected squared norm of the projection of a standard Gaussian vector onto the cone:
δ\(𝒞\)≜𝔼g∼𝒩\(0,In\)\[‖Π𝒞\(g\)‖2\]\.\\delta\(\\mathcal\{C\}\)\\triangleq\\mathbb\{E\}\_\{g\\sim\\mathcal\{N\}\(0,I\_\{n\}\)\}\\left\[\\\|\\Pi\_\{\\mathcal\{C\}\}\(g\)\\\|^\{2\}\\right\]\.\(16\)Crucially, for a linear subspace of dimensionkk,δ\(𝒞\)=k\\delta\(\\mathcal\{C\}\)=k\. For a random cone generated bykkindependent vectors,δ\(𝒞\)≈k/2\\delta\(\\mathcal\{C\}\)\\approx k/2\.
###### Definition A\.3\(Mutual Coherence and Cross\-Correlation\)\.
For a dictionaryD∈ℝn×mD\\in\\mathbb\{R\}^\{n\\times m\}with unit\-norm columns:
1. 1\.TheMutual Coherenceisμ\(D\)≜maxi≠j\|⟨di,dj⟩\|\\mu\(D\)\\triangleq\\max\_\{i\\neq j\}\|\\langle d\_\{i\},d\_\{j\}\\rangle\|\.
2. 2\.We distinguish two forms of subspace interaction: - •Bilinear Correlation \(Random Variable\):For fixed coefficientsαA,αB\\alpha\_\{A\},\\alpha\_\{B\}, we defineρbil≜αA⊤\(DSA⊤DSB\)αB\\rho\_\{\\text\{bil\}\}\\triangleq\\alpha\_\{A\}^\{\\top\}\(D\_\{S\_\{A\}\}^\{\\top\}D\_\{S\_\{B\}\}\)\\alpha\_\{B\}\. This quantity is a random variable with𝔼\[ρbil\]=0\\mathbb\{E\}\[\\rho\_\{\\text\{bil\}\}\]=0over the dictionary ensemble\. - •Subspace Alignment \(Operator Norm\):The worst\-case alignment isρop≜max‖α‖=1αA⊤DSA⊤DSBαB=‖DSA⊤DSB‖op\\rho\_\{\\text\{op\}\}\\triangleq\\max\_\{\\\|\\alpha\\\|=1\}\\alpha\_\{A\}^\{\\top\}D\_\{S\_\{A\}\}^\{\\top\}D\_\{S\_\{B\}\}\\alpha\_\{B\}=\\\|D\_\{S\_\{A\}\}^\{\\top\}D\_\{S\_\{B\}\}\\\|\_\{op\}\. Note: Theorems[3\.3](https://arxiv.org/html/2605.05223#S3.Thmtheorem3)utilizeρbil\\rho\_\{\\text\{bil\}\}to analyze average\-case interference, while bounds involvingμ\\muimplicitly controlρop\\rho\_\{\\text\{op\}\}\.
## Appendix BProofs for Micro\-Analysis \(Section[3](https://arxiv.org/html/2605.05223#S3)\)
In this section, we provide rigorous algebraic derivations for the geometric interaction between sparse cones\. We proceed from the variance decomposition of the leakage operator to the a controlled high\-threshold expansion of the rectified interference of the rectified interference, establishing the micro\-foundations of the structural instability\.
### B\.1Proof of Lemma[3\.2](https://arxiv.org/html/2605.05223#S3.Thmtheorem2)\(Variance Decomposition\)
#### Formal Probability Model\.
To address the precision of our variance decomposition, we explicitly define the probability space and conditioning used throughout this section:
1. 1\.Randomness Source:We take expectations𝔼D\\mathbb\{E\}\_\{D\}strictly with respect to the random generation of the dictionaryDD, where atomsdi∼Unif\(𝕊n−1\)d\_\{i\}\\sim\\text\{Unif\}\(\\mathbb\{S\}^\{n\-1\}\)are i\.i\.d\.
2. 2\.Conditioning:The coefficient vectorsαA,αB\\alpha\_\{A\},\\alpha\_\{B\}are treated as fixed deterministic parameters\. The interference variance derived is therefore the conditional expectation𝔼D\[⋅∣αA,αB\]\\mathbb\{E\}\_\{D\}\[\\cdot\\mid\\alpha\_\{A\},\\alpha\_\{B\}\]\.
3. 3\.Positive Variance Condition:Since the cross\-correlationρ\\rhois a random variable derived fromDD, it may assume negative values locally\. For the subsequent distributional analysis \(Theorem[3\.3](https://arxiv.org/html/2605.05223#S3.Thmtheorem3)\), we define the effective variance with a physical truncationσeff2\(ρ\)≜max\(0,σ02\+2μρ\)\\sigma\_\{eff\}^\{2\}\(\\rho\)\\triangleq\\max\(0,\\sigma\_\{0\}^\{2\}\+2\\mu\\rho\), or equivalently work on the high\-probability eventℰ=\{ρ\>−σ02/2μ\}\\mathcal\{E\}=\\\{\\rho\>\-\\sigma\_\{0\}^\{2\}/2\\mu\\\}\.
###### Proof B\.1\.
Consider the composite activation vectorz=u\+vz=u\+v, whereu=DSAαAu=D\_\{S\_\{A\}\}\\alpha\_\{A\}andv=DSBαBv=D\_\{S\_\{B\}\}\\alpha\_\{B\}represent the signal components from concept A and concept B, respectively\. We analyze the projection ofzzonto a generic ghost featuredjd\_\{j\}\(wherej∉SA∪SBj\\notin S\_\{A\}\\cup S\_\{B\}\)\. The squared projection is given by the expansion:
⟨z,dj⟩2\\displaystyle\\langle z,d\_\{j\}\\rangle^\{2\}=⟨u\+v,dj⟩2\\displaystyle=\\langle u\+v,d\_\{j\}\\rangle^\{2\}\(17\)=⟨u,dj⟩2\+⟨v,dj⟩2\+2⟨u,dj⟩⟨v,dj⟩\.\\displaystyle=\\langle u,d\_\{j\}\\rangle^\{2\}\+\\langle v,d\_\{j\}\\rangle^\{2\}\+2\\langle u,d\_\{j\}\\rangle\\langle v,d\_\{j\}\\rangle\.\(18\)To determine the expected interference, we compute the expectation over the random ghost featuredjd\_\{j\}\. We assume the dictionary atoms are drawn uniformly from the unit sphere𝕊n−1\\mathbb\{S\}^\{n\-1\}\. We invoke the standard spherical integration identity for rank\-1 tensors:
𝔼d∼𝕊n−1\[⟨x,d⟩2\]=1n‖x‖2,\\mathbb\{E\}\_\{d\\sim\\mathbb\{S\}^\{n\-1\}\}\[\\langle x,d\\rangle^\{2\}\]=\\frac\{1\}\{n\}\\\|x\\\|^\{2\},\(19\)and the generalized identity for the cross\-term:
𝔼d∼𝕊n−1\[⟨x,d⟩⟨y,d⟩\]=1n⟨x,y⟩\.\\mathbb\{E\}\_\{d\\sim\\mathbb\{S\}^\{n\-1\}\}\[\\langle x,d\\rangle\\langle y,d\\rangle\]=\\frac\{1\}\{n\}\\langle x,y\\rangle\.\(20\)
Applying these identities term\-wise allows us to decouple the geometric interaction:
1. 1\.Self\-interference \(Independent Terms\):For the individual signal components, the expected projection energy is isotropic: 𝔼\[⟨u,dj⟩2\]=1n‖u‖2and𝔼\[⟨v,dj⟩2\]=1n‖v‖2\.\\mathbb\{E\}\[\\langle u,d\_\{j\}\\rangle^\{2\}\]=\\frac\{1\}\{n\}\\\|u\\\|^\{2\}\\quad\\text\{and\}\\quad\\mathbb\{E\}\[\\langle v,d\_\{j\}\\rangle^\{2\}\]=\\frac\{1\}\{n\}\\\|v\\\|^\{2\}\.We denote these baseline variance contributions asσA2=1n‖u‖2\\sigma\_\{A\}^\{2\}=\\frac\{1\}\{n\}\\\|u\\\|^\{2\}andσB2=1n‖v‖2\\sigma\_\{B\}^\{2\}=\\frac\{1\}\{n\}\\\|v\\\|^\{2\}\.
2. 2\.Cross\-interference \(Interaction Term\):The interaction term captures the geometric alignment between subspaces: 𝔼\[2⟨u,dj⟩⟨v,dj⟩\]\\displaystyle\\mathbb\{E\}\[2\\langle u,d\_\{j\}\\rangle\\langle v,d\_\{j\}\\rangle\]=2n⟨u,v⟩\\displaystyle=\\frac\{2\}\{n\}\\langle u,v\\rangle\(21\)=2n\(DSAαA\)T\(DSBαB\)\\displaystyle=\\frac\{2\}\{n\}\(D\_\{S\_\{A\}\}\\alpha\_\{A\}\)^\{T\}\(D\_\{S\_\{B\}\}\\alpha\_\{B\}\)\(22\)=2nαAT\(DSATDSB\)αB\.\\displaystyle=\\frac\{2\}\{n\}\\alpha\_\{A\}^\{T\}\(D\_\{S\_\{A\}\}^\{T\}D\_\{S\_\{B\}\}\)\\alpha\_\{B\}\.\(23\)
#### Connection to Frame Theory\.
For random dictionaries, the mutual coherenceμ\(D\)=maxi≠j\|⟨di,dj⟩\|\\mu\(D\)=\\max\_\{i\\neq j\}\|\\langle d\_\{i\},d\_\{j\}\\rangle\|concentrates around the Welch bound lower limit, scaling asμ≈\(n−m\)/\(m\(n−1\)\)≈1/n\\mu\\approx\\sqrt\{\(n\-m\)/\(m\(n\-1\)\)\}\\approx 1/\\sqrt\{n\}form≫nm\\gg n\. We can therefore explicitly rescale the interaction factor1/n1/nin terms of the coherence parameterμ\\mu\. Defining the semantic alignment scalar asρ\(SA,SB\)=αAT\(DSATDSB\)αB\\rho\(S\_\{A\},S\_\{B\}\)=\\alpha\_\{A\}^\{T\}\(D\_\{S\_\{A\}\}^\{T\}D\_\{S\_\{B\}\}\)\\alpha\_\{B\}, the expectation becomes:
𝔼\[⟨z,dj⟩2\]=σA2\+σB2\+2μ⋅ρ\(SA,SB\)\+ℛj\.\\mathbb\{E\}\[\\langle z,d\_\{j\}\\rangle^\{2\}\]=\\sigma\_\{A\}^\{2\}\+\\sigma\_\{B\}^\{2\}\+2\\mu\\cdot\\rho\(S\_\{A\},S\_\{B\}\)\+\\mathcal\{R\}\_\{j\}\.\(24\)The residual termℛj\\mathcal\{R\}\_\{j\}accounts for the finite\-frame potential\. For a specific realization of a dictionaryDD, the atoms are fixed, and the expectation is strictly over the selection of indexjj\. The termℛj\\mathcal\{R\}\_\{j\}vanishes asymptotically asm→∞m\\to\\inftydue to the law of large numbers on the sphere, but for finitemm, it represents the texture of the dictionary’s non\-uniformity\.
### B\.2Proof of Theorem[3\.3](https://arxiv.org/html/2605.05223#S3.Thmtheorem3)\(Sensitivity under Alignment and Rectified Drift\)
#### Domain of Definition\.
Mathematically, the cross\-correlation term2μρ2\\mu\\rhocan be negative\. However, the interference variance must be non\-negative\. We formally handle this by defining the variance asv\(ρ\)=max\(0,σ02\+2μρ\)v\(\\rho\)=\\max\(0,\\sigma\_\{0\}^\{2\}\+2\\mu\\rho\)\.Physical Justification:In the high\-dimensional regime \(n→∞n\\to\\infty\) with random dictionaries, the baseline isotropic varianceσ02\\sigma\_\{0\}^\{2\}dominates the fluctuation term2μρ2\\mu\\rho\(which scales asO\(1/n\)O\(1/\\sqrt\{n\}\)\)\. Thus, the event\{v\(ρ\)=0\}\\\{v\(\\rho\)=0\\\}corresponds to a ”super\-orthogonal” alignment that occurs with vanishing probability\. Our analysis focuses on the strictly positive regimeℰ=\{ρ\>−σ02/2μ\}\\mathcal\{E\}=\\\{\\rho\>\-\\sigma\_\{0\}^\{2\}/2\\mu\\\}, which holds almost surely\.
###### Proof B\.2\.
We separate two effects that are often conflated: \(i\) the exceedance probability of a*positive*threshold, which depends only on the pre\-activation tail and is unaffected by rectification, and \(ii\) the*rectified drift*𝔼\[σ\(X\)\]\\mathbb\{E\}\[\\sigma\(X\)\], which is induced by ReLU and accumulates under composition\.
#### \(i\) Exceedance is governed by variance inflation \(no rectification needed\)\.
Let the pre\-activation interference beX∼𝒩\(0,ζ2\)X\\sim\\mathcal\{N\}\(0,\\zeta^\{2\}\)withζ2\(Δ\)=σ02\+Δ\\zeta^\{2\}\(\\Delta\)=\\sigma\_\{0\}^\{2\}\+\\Delta, whereσ02\\sigma\_\{0\}^\{2\}is the baseline isotropic variance andΔ=2μρ\\Delta=2\\mu\\rhois the perturbation induced by subspace alignment\. For any fixed thresholdβ\>0\\beta\>0, we have the identity
ℙ\(σ\(X\)\>β\)=ℙ\(X\>β\)=:P\(Δ\),\\mathbb\{P\}\(\\sigma\(X\)\>\\beta\)=\\mathbb\{P\}\(X\>\\beta\)=:P\(\\Delta\),\(25\)and hence
P\(Δ\)=Q\(βσ02\+Δ\),u\(Δ\):=βσ02\+Δ\.P\(\\Delta\)=Q\\\!\\left\(\\frac\{\\beta\}\{\\sqrt\{\\sigma\_\{0\}^\{2\}\+\\Delta\}\}\\right\),\\qquad u\(\\Delta\):=\\frac\{\\beta\}\{\\sqrt\{\\sigma\_\{0\}^\{2\}\+\\Delta\}\}\.\(26\)Differentiating usingQ′\(u\)=−ϕ\(u\)Q^\{\\prime\}\(u\)=\-\\phi\(u\)withϕ\(u\)=12πe−u2/2\\phi\(u\)=\\frac\{1\}\{\\sqrt\{2\\pi\}\}e^\{\-u^\{2\}/2\}:
dPdΔ\\displaystyle\\frac\{dP\}\{d\\Delta\}=dQdu⋅dudΔ=\(−ϕ\(u\)\)⋅ddΔ\(β\(σ02\+Δ\)−1/2\)\\displaystyle=\\frac\{dQ\}\{du\}\\cdot\\frac\{du\}\{d\\Delta\}=\\big\(\-\\phi\(u\)\\big\)\\cdot\\frac\{d\}\{d\\Delta\}\\big\(\\beta\(\\sigma\_\{0\}^\{2\}\+\\Delta\)^\{\-1/2\}\\big\)\(27\)=\(−ϕ\(u\)\)⋅\(−12β\(σ02\+Δ\)−3/2\)=ϕ\(βσ02\+Δ\)⋅β2\(σ02\+Δ\)3/2\.\\displaystyle=\\big\(\-\\phi\(u\)\\big\)\\cdot\\Big\(\-\\frac\{1\}\{2\}\\beta\(\\sigma\_\{0\}^\{2\}\+\\Delta\)^\{\-3/2\}\\Big\)=\\phi\\\!\\left\(\\frac\{\\beta\}\{\\sqrt\{\\sigma\_\{0\}^\{2\}\+\\Delta\}\}\\right\)\\cdot\\frac\{\\beta\}\{2\(\\sigma\_\{0\}^\{2\}\+\\Delta\)^\{3/2\}\}\.\(28\)Evaluating at the baselineΔ=0\\Delta=0\(whereu0=β/σ0u\_\{0\}=\\beta/\\sigma\_\{0\}\),
P′\(0\)=ϕ\(u0\)⋅β2σ03\.P^\{\\prime\}\(0\)=\\phi\(u\_\{0\}\)\\cdot\\frac\{\\beta\}\{2\\sigma\_\{0\}^\{3\}\}\.\(29\)Normalizing byP\(0\)=Q\(u0\)P\(0\)=Q\(u\_\{0\}\)and using the standard Mills’ ratio approximation in the high\-threshold regime \(u0≳3u\_\{0\}\\gtrsim 3\),Q\(u0\)≈1u0ϕ\(u0\)Q\(u\_\{0\}\)\\approx\\frac\{1\}\{u\_\{0\}\}\\phi\(u\_\{0\}\), we obtain
P′\(0\)P\(0\)\\displaystyle\\frac\{P^\{\\prime\}\(0\)\}\{P\(0\)\}≈ϕ\(u0\)β2σ031u0ϕ\(u0\)=β2σ03⋅u0=β2σ03⋅βσ0=β22σ04\.\\displaystyle\\approx\\frac\{\\phi\(u\_\{0\}\)\\frac\{\\beta\}\{2\\sigma\_\{0\}^\{3\}\}\}\{\\frac\{1\}\{u\_\{0\}\}\\phi\(u\_\{0\}\)\}=\\frac\{\\beta\}\{2\\sigma\_\{0\}^\{3\}\}\\cdot u\_\{0\}=\\frac\{\\beta\}\{2\\sigma\_\{0\}^\{3\}\}\\cdot\\frac\{\\beta\}\{\\sigma\_\{0\}\}=\\frac\{\\beta^\{2\}\}\{2\\sigma\_\{0\}^\{4\}\}\.\(30\)Substituting into the linearized log\-expansion,P\(Δ\)≈P\(0\)exp\(P′\(0\)P\(0\)Δ\)P\(\\Delta\)\\approx P\(0\)\\exp\\\!\\big\(\\frac\{P^\{\\prime\}\(0\)\}\{P\(0\)\}\\Delta\\big\), yields
P\(Δ\)P\(0\)\\displaystyle\\frac\{P\(\\Delta\)\}\{P\(0\)\}≈exp\(β22σ04Δ\)=exp\(β22σ04⋅2μρ\)=exp\(β2μσ04ρ\)\.\\displaystyle\\approx\\exp\\\!\\left\(\\frac\{\\beta^\{2\}\}\{2\\sigma\_\{0\}^\{4\}\}\\Delta\\right\)=\\exp\\\!\\left\(\\frac\{\\beta^\{2\}\}\{2\\sigma\_\{0\}^\{4\}\}\\cdot 2\\mu\\rho\\right\)=\\exp\\\!\\left\(\\frac\{\\beta^\{2\}\\mu\}\{\\sigma\_\{0\}^\{4\}\}\\rho\\right\)\.\(31\)This establishes the exponential sensitivity of*pre\-activation exceedance*to alignment via variance inflation\. Importantly, this effect does not rely on rectification: forβ\>0\\beta\>0, exceedance is identical with or without ReLU\.
#### \(ii\) Rectification creates a one\-sided drift \(the ratchet mechanism\)\.
Rectification matters for*bias/energy*quantities\. Define the rectified drift
η\(Δ\):=𝔼\[σ\(X\)\]\.\\eta\(\\Delta\):=\\mathbb\{E\}\[\\sigma\(X\)\]\.\(32\)ForX∼𝒩\(0,ζ2\(Δ\)\)X\\sim\\mathcal\{N\}\(0,\\zeta^\{2\}\(\\Delta\)\),
η\(Δ\)=𝔼\[max\(0,X\)\]=ζ\(Δ\)𝔼\[max\(0,Z\)\]=ζ\(Δ\)2π=σ02\+Δ2π,Z∼𝒩\(0,1\)\.\\eta\(\\Delta\)=\\mathbb\{E\}\[\\max\(0,X\)\]=\\zeta\(\\Delta\)\\,\\mathbb\{E\}\[\\max\(0,Z\)\]=\\frac\{\\zeta\(\\Delta\)\}\{\\sqrt\{2\\pi\}\}=\\frac\{\\sqrt\{\\sigma\_\{0\}^\{2\}\+\\Delta\}\}\{\\sqrt\{2\\pi\}\},\\qquad Z\\sim\\mathcal\{N\}\(0,1\)\.\(33\)Therefore alignmentρ\\rhoincreases the systematic drift as
η\(ρ\)=σ02\+2μρ2π,\\eta\(\\rho\)=\\frac\{\\sqrt\{\\sigma\_\{0\}^\{2\}\+2\\mu\\rho\}\}\{\\sqrt\{2\\pi\}\},\(34\)which is strictly positive even though𝔼\[X\]=0\\mathbb\{E\}\[X\]=0\. This completes the proof of the theorem\.
### B\.3Proof of Theorem[3\.4](https://arxiv.org/html/2605.05223#S3.Thmtheorem4)\(Convexity of Rectified Accumulation\)
For consistency with Theorem 5, define the effective variancev\+\(ρ\)≔\(σ02\+2μρ\)\+v\_\{\+\}\(\\rho\)\\coloneqq\(\\sigma\_\{0\}^\{2\}\+2\\mu\\rho\)\_\{\+\}\. On the eventE=\{σ02\+2μρ\>0\}E=\\\{\\sigma\_\{0\}^\{2\}\+2\\mu\\rho\>0\\\}we havev\+\(ρ\)=σ02\+2μρv\_\{\+\}\(\\rho\)=\\sigma\_\{0\}^\{2\}\+2\\mu\\rho, so all convexity/Jensen arguments below apply verbatim onEE\(andv\+\(ρ\)=0v\_\{\+\}\(\\rho\)=0otherwise\)\.
###### Proof B\.3\.
We prove the rectified accumulation effect in the form required by Theorem[3\.3](https://arxiv.org/html/2605.05223#S3.Thmtheorem3): geometric variance \(through random alignment\) strictly increases the*expected rectified drift*\. LetX∼𝒩\(0,v\+\(ρ\)\)X\\sim\\mathcal\{N\}\(0,v\_\{\+\}\(\\rho\)\)and defineD\+\(ρ\):=𝔼\[σ\(X\)\]D\_\{\+\}\(\\rho\):=\\mathbb\{E\}\[\\sigma\(X\)\]\. SinceXXis centered Gaussian with variancev\+\(ρ\)v\_\{\+\}\(\\rho\), we have
D\+\(ρ\)=v\+\(ρ\)2π\.D\_\{\+\}\(\\rho\)=\\frac\{\\sqrt\{v\_\{\+\}\(\\rho\)\}\}\{\\sqrt\{2\\pi\}\}\.\(35\)In particular, on the operating eventEEwe may writev\+\(ρ\)=v\(ρ\)v\_\{\+\}\(\\rho\)=v\(\\rho\)withv\(ρ\)=σ02\+2μρv\(\\rho\)=\\sigma\_\{0\}^\{2\}\+2\\mu\\rho, and henceD\+\(ρ\)=D\(ρ\):=v\(ρ\)2πD\_\{\+\}\(\\rho\)=D\(\\rho\):=\\frac\{\\sqrt\{v\(\\rho\)\}\}\{\\sqrt\{2\\pi\}\}\. Since the dictionary orientations are random,ρ\\rhois a zero\-mean random variable\. To show𝔼ρ\[D\+\(ρ\)\]\>D\+\(0\)\\mathbb\{E\}\_\{\\rho\}\[D\_\{\+\}\(\\rho\)\]\>D\_\{\+\}\(0\), it suffices to establish thatD\+D\_\{\+\}is strictly convex on the operating domainEE\(wherev\+\(ρ\)\>0v\_\{\+\}\(\\rho\)\>0\)\.
#### Second derivative and strict convexity\.
We compute derivatives explicitly onEE\. UsingD\(ρ\)=12πv\(ρ\)1/2D\(\\rho\)=\\frac\{1\}\{\\sqrt\{2\\pi\}\}v\(\\rho\)^\{1/2\}andv′\(ρ\)=2μv^\{\\prime\}\(\\rho\)=2\\mu:
D′\(ρ\)\\displaystyle D^\{\\prime\}\(\\rho\)=12π⋅12v\(ρ\)−1/2⋅v′\(ρ\)=μ2πv\(ρ\)−1/2,\\displaystyle=\\frac\{1\}\{\\sqrt\{2\\pi\}\}\\cdot\\frac\{1\}\{2\}v\(\\rho\)^\{\-1/2\}\\cdot v^\{\\prime\}\(\\rho\)=\\frac\{\\mu\}\{\\sqrt\{2\\pi\}\}\\,v\(\\rho\)^\{\-1/2\},\(36\)D′′\(ρ\)\\displaystyle D^\{\\prime\\prime\}\(\\rho\)=μ2π⋅\(−12\)v\(ρ\)−3/2⋅v′\(ρ\)=μ2π⋅\(−12\)v\(ρ\)−3/2⋅2μ=μ22πv\(ρ\)−3/2\.\\displaystyle=\\frac\{\\mu\}\{\\sqrt\{2\\pi\}\}\\cdot\\Big\(\-\\frac\{1\}\{2\}\\Big\)v\(\\rho\)^\{\-3/2\}\\cdot v^\{\\prime\}\(\\rho\)=\\frac\{\\mu\}\{\\sqrt\{2\\pi\}\}\\cdot\\Big\(\-\\frac\{1\}\{2\}\\Big\)v\(\\rho\)^\{\-3/2\}\\cdot 2\\mu=\\frac\{\\mu^\{2\}\}\{\\sqrt\{2\\pi\}\}\\,v\(\\rho\)^\{\-3/2\}\.\(37\)HenceD′′\(ρ\)\>0D^\{\\prime\\prime\}\(\\rho\)\>0for allρ\\rhoin the operating domainEEwherev\(ρ\)=v\+\(ρ\)\>0v\(\\rho\)=v\_\{\+\}\(\\rho\)\>0, proving strict convexity\. \(OutsideEE,v\+\(ρ\)=0v\_\{\+\}\(\\rho\)=0and thusD\+\(ρ\)=0D\_\{\+\}\(\\rho\)=0\.\)
#### Jensen’s inequality and the ratchet\.
By Jensen’s inequality, the strict convexity onEE, and𝔼\[ρ\]=0\\mathbb\{E\}\[\\rho\]=0,
𝔼ρ\[D\+\(ρ\)\]\>D\+\(𝔼ρ\[ρ\]\)=D\+\(0\)=σ02π\.\\mathbb\{E\}\_\{\\rho\}\[D\_\{\+\}\(\\rho\)\]\>D\_\{\+\}\(\\mathbb\{E\}\_\{\\rho\}\[\\rho\]\)=D\_\{\+\}\(0\)=\\frac\{\\sigma\_\{0\}\}\{\\sqrt\{2\\pi\}\}\.\(38\)Therefore, even when alignmentρ\\rhoaverages to zero \(random orientations\), rectification converts symmetric variance fluctuations into a strictly positive*expected*drift\. Equivalently, “lucky” negative alignments reducev\+\(ρ\)v\_\{\+\}\(\\rho\)only linearly \(and can truncate at zero\), while “unlucky” positive alignments increasev\+\(ρ\)v\_\{\+\}\(\\rho\)and, through the convex mapρ↦v\+\(ρ\)\\rho\\mapsto\\sqrt\{v\_\{\+\}\(\\rho\)\}, dominate in expectation\. This provides a rigorous micro\-foundation for the one\-way ReLU ratchet: the degradation is driven by geometric variance, and rectification turns it into systematic drift that accumulates across compositions\.
## Appendix CProofs for Phase Transition \(Section[4](https://arxiv.org/html/2605.05223#S4)\)
In this section, we provide the rigorous derivation of the macroscopic phase transition threshold presented in Theorem[4\.5](https://arxiv.org/html/2605.05223#S4.Thmtheorem5)\. We employ the framework of Conic Geometry and Gordon’s Escape through a Mesh Theorem \(GMT\) to quantify the probability of compositional collapse\. To ensure self\-containment, we first introduce the necessary geometric definitions and auxiliary lemmas governing the statistical dimension of random cones\.
### C\.1Geometric Preliminaries and Kinematic Formula
The core of our analysis relies on the concept of Statistical Dimension, which generalizes the notion of linear dimension to convex cones\.
###### Definition C\.1\(Statistical Dimension\)\.
Let𝒞⊆ℝn\\mathcal\{C\}\\subseteq\\mathbb\{R\}^\{n\}be a closed convex cone\. The statistical dimensionδ\(𝒞\)\\delta\(\\mathcal\{C\}\)is defined as the expected squared norm of the projection of a standard Gaussian vector onto𝒞\\mathcal\{C\}:
δ\(𝒞\):=𝔼g∼𝒩\(0,In\)\[‖Π𝒞\(g\)‖2\]\.\\delta\(\\mathcal\{C\}\):=\\mathbb\{E\}\_\{g\\sim\\mathcal\{N\}\(0,I\_\{n\}\)\}\\left\[\\\|\\Pi\_\{\\mathcal\{C\}\}\(g\)\\\|^\{2\}\\right\]\.
A fundamental property of the statistical dimension, which we utilize to decompose the interaction between the signal and ghost features, is its additivity under polarity\.
###### Lemma C\.2\(Moreau’s Decomposition Theorem for Dimensions\)\.
For any closed convex cone𝒞⊂ℝn\\mathcal\{C\}\\subset\\mathbb\{R\}^\{n\}, let𝒞∘\\mathcal\{C\}^\{\\circ\}denote its polar cone\. Then:
δ\(𝒞\)\+δ\(𝒞∘\)=n\.\\delta\(\\mathcal\{C\}\)\+\\delta\(\\mathcal\{C\}^\{\\circ\}\)=n\.
###### Proof C\.3\.
By Moreau’s decomposition theorem, for any vectorxx, we havex=Π𝒞\(x\)\+Π𝒞∘\(x\)x=\\Pi\_\{\\mathcal\{C\}\}\(x\)\+\\Pi\_\{\\mathcal\{C\}^\{\\circ\}\}\(x\)with⟨Π𝒞\(x\),Π𝒞∘\(x\)⟩=0\\langle\\Pi\_\{\\mathcal\{C\}\}\(x\),\\Pi\_\{\\mathcal\{C\}^\{\\circ\}\}\(x\)\\rangle=0\. Letg∼𝒩\(0,In\)g\\sim\\mathcal\{N\}\(0,I\_\{n\}\)\. Then‖g‖2=‖Π𝒞\(g\)‖2\+‖Π𝒞∘\(g\)‖2\\\|g\\\|^\{2\}=\\\|\\Pi\_\{\\mathcal\{C\}\}\(g\)\\\|^\{2\}\+\\\|\\Pi\_\{\\mathcal\{C\}^\{\\circ\}\}\(g\)\\\|^\{2\}\. Taking expectations on both sides, we obtain𝔼‖g‖2=δ\(𝒞\)\+δ\(𝒞∘\)\\mathbb\{E\}\\\|g\\\|^\{2\}=\\delta\(\\mathcal\{C\}\)\+\\delta\(\\mathcal\{C\}^\{\\circ\}\)\. Since𝔼‖g‖2=n\\mathbb\{E\}\\\|g\\\|^\{2\}=n, the result follows\.
Recall from Theorem[4\.3](https://arxiv.org/html/2605.05223#S4.Thmtheorem3)that the stability of compositional steering is equivalent to the disjointness of the signal cone𝒦S\\mathcal\{K\}\_\{S\}and the ghost polar cone𝒦J∘\\mathcal\{K\}\_\{J\}^\{\\circ\}\. Note that in the main text, we analyzed the intersection of𝒦S\\mathcal\{K\}\_\{S\}and𝒦J∘\\mathcal\{K\}\_\{J\}^\{\\circ\}\. However, for calculation purposes, it is often more convenient to frame the condition as the sum of dimensions of the primal cones\. According to the Approximate Kinematic Formula\(amelunxen2014living\), a phase transition occurs when:
δ\(𝒦S\)\+δ\(𝒦J\)≈n\.\\delta\(\\mathcal\{K\}\_\{S\}\)\+\\delta\(\\mathcal\{K\}\_\{J\}\)\\approx n\.\(39\)This condition marks the boundary where the probability of intersection transitions from 0 to 1 asn→∞n\\to\\infty\. We now derive explicit bounds for each term\.Remark\.Using Moreau’s decomposition,δ\(𝒦J\)\+δ\(𝒦J∘\)=n\\delta\(\\mathcal\{K\}\_\{J\}\)\+\\delta\(\\mathcal\{K\}\_\{J\}^\{\\circ\}\)=n, hence the saturation conditionδ\(𝒦S\)\+δ\(𝒦J\)≈n\\delta\(\\mathcal\{K\}\_\{S\}\)\+\\delta\(\\mathcal\{K\}\_\{J\}\)\\approx nis equivalent \(up to the same approximation\) toδ\(𝒦S\)\+δ\(𝒦J∘\)≈n\\delta\(\\mathcal\{K\}\_\{S\}\)\+\\delta\(\\mathcal\{K\}\_\{J\}^\{\\circ\}\)\\approx n\.
### C\.2Statistical Dimension of the Signal Cone\\texorpdfstring𝒦S\\mathcal\{K\}\_\{S\}K\_S
The signal cone𝒦S=cone\(\{di\}i∈S\)\\mathcal\{K\}\_\{S\}=\\text\{cone\}\(\\\{d\_\{i\}\\\}\_\{i\\in S\}\)is the positive hull ofkklinearly independent vectors\. While the statistical dimension of a random simplicial cone is exactlyk/2k/2, we must account for the worst\-case alignment within the subspace spanned by these features\.
We approximate the dimension of the cone by the dimension of its linear hull\. LetLS=span\(𝒦S\)L\_\{S\}=\\text\{span\}\(\\mathcal\{K\}\_\{S\}\)\. Since the vectors are linearly independent \(valid fork≪nk\\ll n\),dim\(LS\)=k\\dim\(L\_\{S\}\)=k\. By the properties of projection,‖Π𝒦S\(g\)‖≤‖ΠLS\(g\)‖\\\|\\Pi\_\{\\mathcal\{K\}\_\{S\}\}\(g\)\\\|\\leq\\\|\\Pi\_\{L\_\{S\}\}\(g\)\\\|almost surely\. Squaring and taking expectations yields:
δ\(𝒦S\)≤δ\(LS\)=k\.\\delta\(\\mathcal\{K\}\_\{S\}\)\\leq\\delta\(L\_\{S\}\)=k\.For the derivation of the \*sufficient\* condition for stability \(lower bound on collapse\), we adopt this conservative upper bound:
δ\(𝒦S\)n≈kn=γ\.\\frac\{\\delta\(\\mathcal\{K\}\_\{S\}\)\}\{n\}\\approx\\frac\{k\}\{n\}=\\gamma\.\(40\)This term corresponds to the “Signal Radius“γ\\sqrt\{\\gamma\}in Eq\. \(11\)\.
### C\.3Statistical Dimension of the Ghost Cone\\texorpdfstring𝒦J\\mathcal\{K\}\_\{J\}K\_J
The non\-trivial component lies in estimatingδ\(𝒦J\)\\delta\(\\mathcal\{K\}\_\{J\}\), where𝒦J\\mathcal\{K\}\_\{J\}is generated byN=m−kN=m\-krandom vectors\. Letρ=N/n=δdict−γ\\rho=N/n=\\delta\_\{dict\}\-\\gammadenote the aspect ratio\. We leverage the connection between statistical dimension and Gaussian Mean Width\. For a cone𝒞\\mathcal\{C\},δ\(𝒞\)≈w2\(𝒞∩𝕊n−1\)\\delta\(\\mathcal\{C\}\)\\approx w^\{2\}\(\\mathcal\{C\}\\cap\\mathbb\{S\}^\{n\-1\}\)\.
We derive the width using explicit Gaussian integration\. The normalized dimensionΔ\(ρ\)=δ\(𝒦J\)/n\\Delta\(\\rho\)=\\delta\(\\mathcal\{K\}\_\{J\}\)/nis determined by the probability mass of the Gaussian tail\.
###### Lemma C\.4\(Integral Representation\)\.
The dimension densityΔ\(ρ\)\\Delta\(\\rho\)satisfies the integral equation:
Δ\(ρ\)=∫τ∞\(t2\+1\)ϕ\(t\)𝑑t\+τϕ\(τ\)\\Delta\(\\rho\)=\\int\_\{\\tau\}^\{\\infty\}\(t^\{2\}\+1\)\\phi\(t\)dt\+\\tau\\phi\(\\tau\)which simplifies for largeτ\\tautoΔ\(ρ\)≈ℙ\(Z\>τ\)⋅\(1\+τ2\)\\Delta\(\\rho\)\\approx\\mathbb\{P\}\(Z\>\\tau\)\\cdot\(1\+\\tau^\{2\}\), subject to the constraint:
12π∫τ∞e−t2/2𝑑t=12ρ\.\\frac\{1\}\{\\sqrt\{2\\pi\}\}\\int\_\{\\tau\}^\{\\infty\}e^\{\-t^\{2\}/2\}dt=\\frac\{1\}\{2\\rho\}\.\(41\)
The thresholdτ\\taurepresents the geometric “angle“ of the cone\. The conditionΦc\(τ\)=1/\(2ρ\)\\Phi^\{c\}\(\\tau\)=1/\(2\\rho\)implies that as the dictionary sizeρ\\rhoincreases, the tail probability must decrease, forcingτ\\tauto increase\.
### C\.4Derymptotic Analysis and Explicit Bound \(Eq\. 11\)
We now perform a detailed asymptotic expansion to solve forτ\\tauand derive the closed\-form bound in Theorem[4\.5](https://arxiv.org/html/2605.05223#S4.Thmtheorem5)\.
###### Proof C\.5\.
We start with the conditionδ\(𝒦S\)\+δ\(𝒦J\)≤n\\delta\(\\mathcal\{K\}\_\{S\}\)\+\\delta\(\\mathcal\{K\}\_\{J\}\)\\leq n\. In terms of the square\-root formulation \(which aligns with Euclidean distance concentration\), the separation condition is:
δ\(𝒦S\)\+δ\(𝒦J\)≤n\.\\sqrt\{\\delta\(\\mathcal\{K\}\_\{S\}\)\}\+\\sqrt\{\\delta\(\\mathcal\{K\}\_\{J\}\)\}\\leq\\sqrt\{n\}\.Substituting the signal dimension, we requireγn\+δ\(𝒦J\)≤n\\sqrt\{\\gamma n\}\+\\sqrt\{\\delta\(\\mathcal\{K\}\_\{J\}\)\}\\leq\\sqrt\{n\}\. Dividing byn\\sqrt\{n\}, we need an expression forΔ\(ρ\)\\sqrt\{\\Delta\(\\rho\)\}\.
Step 1: Approximating the Thresholdτ\\tau\. From Eq\. \([41](https://arxiv.org/html/2605.05223#A3.E41)\), we use the Mills’ Ratio inequality for the complementary CDFΦc\(τ\)\\Phi^\{c\}\(\\tau\):
ττ2\+1ϕ\(τ\)≤Φc\(τ\)≤1τϕ\(τ\)\.\\frac\{\\tau\}\{\\tau^\{2\}\+1\}\\phi\(\\tau\)\\leq\\Phi^\{c\}\(\\tau\)\\leq\\frac\{1\}\{\\tau\}\\phi\(\\tau\)\.For largeτ\\tau,Φc\(τ\)≈1τ2πe−τ2/2\\Phi^\{c\}\(\\tau\)\\approx\\frac\{1\}\{\\tau\\sqrt\{2\\pi\}\}e^\{\-\\tau^\{2\}/2\}\. Equating this to1/\(2ρ\)1/\(2\\rho\):
1τ2πe−τ2/2=12ρ⟹eτ2/2τ=2ρ2π\.\\frac\{1\}\{\\tau\\sqrt\{2\\pi\}\}e^\{\-\\tau^\{2\}/2\}=\\frac\{1\}\{2\\rho\}\\implies\\frac\{e^\{\\tau^\{2\}/2\}\}\{\\tau\}=\\frac\{2\\rho\}\{\\sqrt\{2\\pi\}\}\.Taking natural logarithms on both sides:
τ22−logτ=log\(2ρ\)−12log\(2π\)\.\\frac\{\\tau^\{2\}\}\{2\}\-\\log\\tau=\\log\(2\\rho\)\-\\frac\{1\}\{2\}\\log\(2\\pi\)\.For largeρ\\rho, the termτ2/2\\tau^\{2\}/2dominates\. We can iteratively solve forτ\\tau\. The first\-order approximation givesτ2≈2logρ\\tau^\{2\}\\approx 2\\log\\rho\. To obtain the precise scaling inside the logarithm, we substituteτ≈2logρ\\tau\\approx\\sqrt\{2\\log\\rho\}back into thelogτ\\log\\tauterm:
τ22≈log\(2ρ\)\+log\(2logρ\)≈log\(ρ2π\)\+𝒪\(loglogρ\)\.\\frac\{\\tau^\{2\}\}\{2\}\\approx\\log\(2\\rho\)\+\\log\(\\sqrt\{2\\log\\rho\}\)\\approx\\log\\left\(\\frac\{\\rho\}\{\\sqrt\{2\\pi\}\}\\right\)\+\\mathcal\{O\}\(\\log\\log\\rho\)\.Thus, we establish the scaling:
τ≈2log\(ρ2π\)\.\\tau\\approx\\sqrt\{2\\log\\left\(\\frac\{\\rho\}\{\\sqrt\{2\\pi\}\}\\right\)\}\.\(42\)
Step 2: Relatingτ\\tauto Statistical Dimension\. Using the property thatδ\(𝒦J\)≈n⋅Δ\(ρ\)\\delta\(\\mathcal\{K\}\_\{J\}\)\\approx n\\cdot\\Delta\(\\rho\), and the result fromamelunxen2014livingthatΔ\(ρ\)\\Delta\(\\rho\)is concentrated around the value where the Gaussian width maximizes, we have that for polyhedral cones, the dimension scales as the squared width\. Specifically, the dimension of the cone generated byNNrandom vectors \(N≫nN\\gg n\) behaves as:
δ\(𝒦J\)≈n⋅2ρlog\(eρ\)\(entropy scaling\)\.\\delta\(\\mathcal\{K\}\_\{J\}\)\\approx n\\cdot 2\\rho\\log\\left\(\\frac\{e\}\{\\rho\}\\right\)\\quad\\text\{\(entropy scaling\)\}\.However, using the precise geometric width derived from the thresholdτ\\tau:
δ\(𝒦J\)n≈ρ⋅τ2\(locally\)\.\\frac\{\\delta\(\\mathcal\{K\}\_\{J\}\)\}\{n\}\\approx\\rho\\cdot\\tau^\{2\}\\quad\\text\{\(locally\)\}\.A more careful integration of the expected projection norm yields the explicit bound used in Eq\. \(11\)\. We substituteρ=δdict−γ\\rho=\\delta\_\{dict\}\-\\gamma\. The “pressure“ term corresponds to the width of the set ofNNconstraints\. The effective width is given by:
𝒲=2\(δdict−γ\)log\(e\(δdict−γ\)2π\)\.\\mathcal\{W\}=\\sqrt\{2\(\\delta\_\{dict\}\-\\gamma\)\\log\\left\(\\frac\{e\}\{\(\\delta\_\{dict\}\-\\gamma\)\\sqrt\{2\\pi\}\}\\right\)\}\.
Step 3: Final Phase Boundary\. Combining the signal radius and the ghost pressure, the safe region is defined by:
γ\+2\(δdict−γ\)log\(e\(δdict−γ\)2π\)≤1\.\\sqrt\{\\gamma\}\+\\sqrt\{2\(\\delta\_\{dict\}\-\\gamma\)\\log\\left\(\\frac\{e\}\{\(\\delta\_\{dict\}\-\\gamma\)\\sqrt\{2\\pi\}\}\\right\)\}\\leq 1\.\(43\)Here, the logarithmic termlog\(…\)\\log\(\\dots\)arises directly from the inversion of the Gaussian tail probabilityΦc\(τ\)\\Phi^\{c\}\(\\tau\)\. This confirms that as the dictionary sizeδdict\\delta\_\{dict\}grows linearly, the “safe“ compositional densityγ\\gammamust shrink super\-linearly to maintain disjointness\.
This concludes the rigorous derivation of the phase transition boundary in Theorem[4\.5](https://arxiv.org/html/2605.05223#S4.Thmtheorem5)\.
## Appendix DExperimental Setup and Additional Results
In this section, we provide the implementation details of our experimental framework, including the training of Sparse Autoencoders \(SAEs\) on Qwen\-VL, the extraction of semantic features from the CLEVR dataset, and additional ablation studies characterizing the structural correlations of the learned dictionary\.
### D\.1Model and SAE Training Details
All experiments were conducted using the Qwen\-VL\-Chat \(7B parameters\) vision\-language model\. We focused our analysis on the residual stream of the middle layer \(Layer 16\), where semantic abstraction is hypothesized to be maximal\.
#### SAE Architecture\.
We trained a standard Sparse Autoencoder to decompose the residual stream activationsx∈ℝnx\\in\\mathbb\{R\}^\{n\}\. The SAE consists of an encoderWenc∈ℝm×nW\_\{enc\}\\in\\mathbb\{R\}^\{m\\times n\}and a decoderWdec∈ℝn×mW\_\{dec\}\\in\\mathbb\{R\}^\{n\\times m\}\.
- •Residual Dimension:n=4,096n=4,096\.
- •Dictionary Size:m=32,768m=32,768\(Expansion factorδ=8\\delta=8\)\.
- •Activation Function:ReLU\.
- •Loss Function:We minimized the reconstruction loss with anL1L\_\{1\}sparsity penalty: ℒ=‖x−Wdecσ\(Wenc\(x−bdec\)\+benc\)‖22\+λ‖f\(x\)‖1\\mathcal\{L\}=\\\|x\-W\_\{dec\}\\sigma\(W\_\{enc\}\(x\-b\_\{dec\}\)\+b\_\{enc\}\)\\\|\_\{2\}^\{2\}\+\\lambda\\\|f\(x\)\\\|\_\{1\}whereλ=0\.05\\lambda=0\.05was tuned to achieve an average sparsity ofk≈40k\\approx 40active features per token\.
### D\.2CLEVR Feature Extraction and Steering
To validate our theory on structured data, we utilized the CLEVR dataset, which contains compositional objects defined by attributes:Shape, Color, Material, Size\.
#### Feature Identification\.
We passed 10,000 CLEVR images through Qwen\-VL and collected the SAE latent activations\. We identified interpretable features using a two\-step process: 1\. Automatic Labeling: We computed the Pearson correlation between latent activations and ground\-truth attribute labels \(e\.g\., “Is there a red cube¿‘\)\. 2\. Manual Verification: We inspected the top\-20 max\-activating examples for the highest\-correlated latents to ensure monosemanticity\. This process yielded a curated set of≈500\\approx 500high\-quality features representing atomic concepts \(e\.g\., Feature 124: “Red“, Feature 892: “Cylinder“\)\.
#### Compositional Steering Protocol\.
To generate the phase transition curve in Figure[4](https://arxiv.org/html/2605.05223#A4.F4): 1\. We randomly sampledkkdistinct attribute features\{vi\}i=1k\\\{v\_\{i\}\\\}\_\{i=1\}^\{k\}from the curated set\. 2\. We constructed a steering vectorz=∑i=1kαiviz=\\sum\_\{i=1\}^\{k\}\\alpha\_\{i\}v\_\{i\}, with coefficientsαi∼𝒰\[0\.8,1\.2\]\\alpha\_\{i\}\\sim\\mathcal\{U\}\[0\.8,1\.2\]to simulate variance\. 3\. We measured theSpurious Energyℰspur\\mathcal\{E\}\_\{spur\}by projectingzzonto the subspace ofall otherdictionary atoms \(the ghost featuresJJ\):
ℰspur\(k\)=‖ΠJ\(ReLU\(WencWdecz\)\)‖2‖z‖2\\mathcal\{E\}\_\{spur\}\(k\)=\\frac\{\\\|\\Pi\_\{J\}\(\\text\{ReLU\}\(W\_\{enc\}W\_\{dec\}z\)\)\\\|\_\{2\}\}\{\\\|z\\\|\_\{2\}\}\(44\)The “Collapse“ is defined as the point whereℰspur\\mathcal\{E\}\_\{spur\}exceeds the thresholdη=0\.1\\eta=0\.1\.
#### Detailed Protocol\.
Steering is implemented by adding the vectorzzto the residual stream with coefficientλsteer\\lambda\_\{steer\}\. We explicitly sweepλsteer∈\[0,5\]\\lambda\_\{steer\}\\in\[0,5\]\. The spurious activation threshold is set toβ=mean\(Xclean\)\+3σ\(Xclean\)\\beta=\\text\{mean\}\(X\_\{clean\}\)\+3\\sigma\(X\_\{clean\}\)\(approx\. 0\.1 in practice\)\. Spurious energyℰspur\\mathcal\{E\}\_\{spur\}is computed as theL2L\_\{2\}norm of the projection onto the ghost subspaceJJ, normalized by the injected signal norm:ℰspur=‖ΠJ\(ReLU\(z\)\)‖2/‖z‖2\\mathcal\{E\}\_\{spur\}=\\\|\\Pi\_\{J\}\(\\text\{ReLU\}\(z\)\)\\\|\_\{2\}/\\\|z\\\|\_\{2\}\.
Densityγ=k/n\\gamma=k/nSpurious Energyℰspur\\mathcal\{E\}\_\{spur\}γ∗\\gamma^\{\*\}Theory \(Eq\. 11\)CLEVR \(Structured\)Δstruct\\Delta\_\{struct\}\(Correlation Shift\)Stability PhaseCollapse PhaseFigure 4:Phase Transition of Compositional Collapse\.Theoretical prediction \(blue\) vs\. empirical CLEVR latents \(red points\)\. The theoretical curve shows a transition atγ∗\\gamma^\{\*\}\. Real\-world structured features exhibit acorrelation shift, collapsing slightly earlier than the random baseline\.
### D\.3Additional Results: The Structure of Interference
A key claim of our paper \(Section[6](https://arxiv.org/html/2605.05223#S6)\) is that real\-world dictionaries are structured, leading to an earlier collapse than random dictionaries\. To visualize this, we computed theGram MatrixG=DTDG=D^\{T\}Dfor the identified CLEVR features\.
#### Block\-Diagonal Correlations\.
Figure[5](https://arxiv.org/html/2605.05223#A4.F5)compares the interaction matrix of random Gaussian vectors versus learned CLEVR features\.
- •Random Baseline:The off\-diagonal correlations are uniformly distributed around 0 with variance1/n1/n\(left\)\.
- •CLEVR Features:The matrix exhibits distinctblock\-diagonal structure\(right\)\. Features within the same semantic category \(e\.g\., Colors\) or frequently co\-occurring attributes \(e\.g\., Shiny \+ Metal\) show significantly higher mutual coherence \(μlocal≈0\.15≫1/n\\mu\_\{local\}\\approx 0\.15\\gg 1/\\sqrt\{n\}\)\.
\(A\) Random DictionaryFeature IndexiiFeature Indexjj\(B\) CLEVR FeaturesColorShapeMat\.Feature Indexii1\.0 \(Corr\)≈ρ\\approx\\rho0\.0Figure 5:Structure of Interference\.Comparison of the Gram matrix \(Gij=\|⟨di,dj⟩\|G\_\{ij\}=\|\\langle d\_\{i\},d\_\{j\}\\rangle\|\) for\(A\)a random spherical dictionary and\(B\)learned CLEVR features\. The CLEVR features exhibit significant block\-diagonal structure \(semantic clusters\) and off\-diagonal correlations, leading to a higher effective coherenceμlocal\\mu\_\{local\}than the random baseline\. This structural alignment accelerates the phase transition\.This empirical evidence supports our theoretical argument in Lemma[3\.2](https://arxiv.org/html/2605.05223#S3.Thmtheorem2): the strictly positive cross\-correlation terms \(ρ\>0\\rho\>0\) in structured data effectively lower the “Safety Ceiling,“ shifting the phase transition curve to the left as observed in Figure[4](https://arxiv.org/html/2605.05223#A4.F4)\.
### D\.4Ablation: Dictionary Overcompleteness
We also investigated the effect of dictionary size on stability\. Consistent with Theorem[4\.5](https://arxiv.org/html/2605.05223#S4.Thmtheorem5), increasing the expansion ratioδ=m/n\\delta=m/n
## Appendix EAuxiliary Lemmas
In this final section, we state the standard probability and geometric theorems used throughout our proofs\. These are well\-established results in high\-dimensional probability and convex geometry\.
### E\.1Gaussian Concentration and Tail Bounds
The analysis of the “Ratchet Mechanism“ \(Theorem[3\.3](https://arxiv.org/html/2605.05223#S3.Thmtheorem3)\) relies on precise bounds for the tail of the standard normal distribution\.
###### Lemma E\.1\(Gaussian Mill’s Ratio\(vershynin2018high\)\)\.
Letg∼𝒩\(0,1\)g\\sim\\mathcal\{N\}\(0,1\)\. For anyt\>0t\>0, the tail probability satisfies:
\(1t−1t3\)12πe−t2/2≤ℙ\(g\>t\)≤1t12πe−t2/2\\left\(\\frac\{1\}\{t\}\-\\frac\{1\}\{t^\{3\}\}\\right\)\\frac\{1\}\{\\sqrt\{2\\pi\}\}e^\{\-t^\{2\}/2\}\\leq\\mathbb\{P\}\(g\>t\)\\leq\\frac\{1\}\{t\}\\frac\{1\}\{\\sqrt\{2\\pi\}\}e^\{\-t^\{2\}/2\}\(45\)This inequality justifies the asymptotic expansion used in the interference analysis, where the “rectified“ probability mass is dominated by the exponential decay termexp\(−β2/2σ2\)\\exp\(\-\\beta^\{2\}/2\\sigma^\{2\}\)\.
### E\.2Gordon’s Escape Through a Mesh Theorem
The derivation of the Phase Transition \(Theorem[4\.5](https://arxiv.org/html/2605.05223#S4.Thmtheorem5)\) fundamentally rests on Gordon’s comparison inequality for Gaussian processes\.
###### Lemma E\.2\(Gordon’s Escape Theorem\(gordon1988milman\)\)\.
Let𝒮⊂𝕊n−1\\mathcal\{S\}\\subset\\mathbb\{S\}^\{n\-1\}be a subset of the unit sphere\. LetM∈ℝm×nM\\in\\mathbb\{R\}^\{m\\times n\}be a Gaussian random matrix with entriesMij∼𝒩\(0,1\)M\_\{ij\}\\sim\\mathcal\{N\}\(0,1\)\. Let𝒦\\mathcal\{K\}be a closed convex cone inℝm\\mathbb\{R\}^\{m\}\. Define the intersection probabilityPesc=ℙ\(𝒮∩null\(M\)≠∅\)P\_\{esc\}=\\mathbb\{P\}\(\\mathcal\{S\}\\cap\\text\{null\}\(M\)\\neq\\emptyset\)\. Gordon’s theorem relates this geometric probability to the Gaussian Mean Widthw\(𝒮\)w\(\\mathcal\{S\}\)\. Specifically, for the intersection of a cone𝒞\\mathcal\{C\}and a random subspace of codimensionmm, the intersection is trivial with high probability if:
w\(𝒞\)<m−12mw\(\\mathcal\{C\}\)<\\sqrt\{m\}\-\\frac\{1\}\{2\\sqrt\{m\}\}\(46\)In our context \(Section[4](https://arxiv.org/html/2605.05223#S4)\), we apply the “Gaussian Min\-Max“ variant of this theorem to compare the primary process𝒳u,v\\mathcal\{X\}\_\{u,v\}\(geometry\) with the simplified process𝒴u,v\\mathcal\{Y\}\_\{u,v\}\(statistical dimension\)\.
### E\.3Extreme Values of Gaussian Vectors
To compute the width of the ghost polytopeΨpoly\\Psi\_\{poly\}in Appendix[C](https://arxiv.org/html/2605.05223#A3), we used the expected maximum ofNNGaussian variables\.
###### Lemma E\.3\(Maximum of Gaussian Vector\)\.
Letg1,…,gNg\_\{1\},\\dots,g\_\{N\}be i\.i\.d\.𝒩\(0,1\)\\mathcal\{N\}\(0,1\)\. The expected maximum scales as:
𝔼\[max1≤i≤Ngi\]≤2logN\\mathbb\{E\}\[\\max\_\{1\\leq i\\leq N\}g\_\{i\}\]\\leq\\sqrt\{2\\log N\}\(47\)More precisely, for the Gaussian Mean Width of a polytope generated byNNvertices, the squared width satisfies:
w2\(conv\{gi\}\)≈2logNw^\{2\}\(\\text\{conv\}\\\{g\_\{i\}\\\}\)\\approx 2\\log N\(48\)This justifies the logarithmic term2\(δ−γ\)log\(…\)\\sqrt\{2\(\\delta\-\\gamma\)\\log\(\\dots\)\}in our explicit threshold formula \(Eq\. 11\)\.
### E\.4Random Matrix Theory: The Bai\-Yin Theorem
The divergence of the condition number in Proposition 5\.1 relies on the spectral limits of Wishart matrices\.
###### Lemma E\.4\(Bai\-Yin Theorem\)\.
LetAAbe ann×kn\\times kmatrix with independent entries having zero mean and unit variance\. Asn,k→∞n,k\\to\\inftywithk/n→γ∈\(0,1\)k/n\\to\\gamma\\in\(0,1\), the extreme singular values of the normalized matrix1nA\\frac\{1\}\{\\sqrt\{n\}\}Aconverge almost surely to:
σmin\\displaystyle\\sigma\_\{\\min\}→1−γ\\displaystyle\\to 1\-\\sqrt\{\\gamma\}\(49\)σmax\\displaystyle\\sigma\_\{\\max\}→1\+γ\\displaystyle\\to 1\+\\sqrt\{\\gamma\}\(50\)This result implies that the condition numberκ=σmax/σmin\\kappa=\\sigma\_\{\\max\}/\\sigma\_\{\\min\}diverges asγ→1\\gamma\\to 1\. In our “effective geometry“ framework, we replace the physical limit11with the phase transition limitγ∗\\gamma^\{\*\}, yielding the scaling law in Eq\. \(12\)\.
## Appendix FExtended Synthetic Benchmarks
To address the concern regarding finite\-size scaling and dictionary coherence, we conducted extensive synthetic experiments\. As shown in Figure[F](https://arxiv.org/html/2605.05223#A6), the transition becomes increasingly steep as the ambient dimensionnnincreases, consistent with the asymptotic concentration of the Gaussian Mean Width\.
Furthermore, we systematically varied the mutual coherenceμ\\mu\. As predicted by the kinematic formula in Eq\. \([13](https://arxiv.org/html/2605.05223#S4.E13)\), higher coherence effectively ”shrinks” the available geometric budget, causing the compositional collapse to occur at lower densitiesγ\\gamma\. This confirms that our theory serves as a reliable bound across diverse dictionary regimes\.
00\.20\.20\.40\.40\.60\.60\.80\.81100\.20\.20\.40\.40\.60\.60\.80\.811Stable ZoneCollapse ZoneCompositional Densityγ=k/n\\gamma=k/nSpurious Energyℰspur\\mathcal\{E\}\_\{spur\}Empirical Phase Transition \(n=512n=512\)SimulationTheoryγ∗\\gamma^\{\*\}Figure 6:Empirical verification of the phase boundary\. The transition exhibits a characteristic ”tail” nearγ∗\\gamma^\{\*\}due to finite\-size effects, aligning with the geometric threshold derived from Gaussian mean width\.
## Appendix GExtended Discussion and Limitations
In this section, we expand upon the implications of our theoretical findings and explicitly address the limitations of our current framework\.
### G\.1Limitations of the Random Dictionary Assumption
Our theoretical bounds \(Theorem[4\.5](https://arxiv.org/html/2605.05223#S4.Thmtheorem5)\) rely on the assumption that the feature dictionaryDDis drawn from a rotationally invariant distribution \(e\.g\., Gaussian\)\. While this is a standard assumption in the Compressed Sensing and High\-Dimensional Probability literature\(vershynin2018high\), real\-world features learned by SAEs likely possess structured correlations \(e\.g\., hierarchical clusters or semantic manifolds\)\. However, as shown in our CLEVR experiments \(Appendix[D](https://arxiv.org/html/2605.05223#A4)\), real\-world structure oftenacceleratesthe phase transition rather than delaying it\. Therefore, our random\-matrix bounds serve as a necessary, albeit optimistic, condition for stability\. Future work will aim to incorporate specific covariance structures into the Gordon’s Mesh analysis to derive tighter bounds for structured data\.
### G\.2ReLU vs\. Smooth Activations
Our derivation of the “Ratchet Inequality“ \(Theorem[3\.3](https://arxiv.org/html/2605.05223#S3.Thmtheorem3)\) explicitly exploits the hard thresholding property of the ReLU activation\. Modern LLMs often employ smooth variants such as GeLU or SwiGLU\. While smooth activations do not have a hard cutoff at zero, they exhibit similar rectification behavior for large negative inputs\. We conjecture that the “Ratchet“ mechanism persists in these regimes, albeit with a “softer“ phase transition boundary\. The structural instability is driven by the asymmetry of the activation function, not its non\-differentiability\. Extending our measure\-concentration proofs to GeLU\-based networks is a promising direction for future research\.
### G\.3Implications for Architectural Design
Our “Geometric Impossibility Theorem“ suggests that increasing width alone cannot solve the compositional interference problem in the overcomplete regime\. This provides a geometric justification for the necessity of Depth \(layer\-wise processing\) and Time \(Chain\-of\-Thought serialization\)\. By spreading compositional steps across depth or time, the model effectively resets the “interference budget“ at each step\. This supports the hypothesis that complex logical reasoning requires serial computation not just for algorithmic reasons, but for geometric stability in the latent space\.
### G\.4Broader Impact
This work is theoretical in nature and focuses on the interpretability of foundation models\. By establishing rigorous bounds for activation steering, our framework contributes to the safety and reliability of AI systems\. Understanding the limits of linear intervention is crucial for preventing unintended side effects when attempting to align or control Large Language Models\. We do not foresee any immediate negative societal impacts from this theoretical study\.
### G\.5Experimental Scope and Baselines
#### Choice of CLEVR vs\. Text LLMs\.
Our experiments focus on the CLEVR dataset because it provides ground\-truth control over the sparsitykkand feature orthogonality, which is difficult to estimate precisely in pre\-trained LLMs \(e\.g\., Llama\-3\)\. While recent work\(gurnee2023finding\)suggests similar sparse firing patterns in text models, verifying our precise phase boundary requires the controlled injection of known signals, making synthetic/vision setups more rigorous for this specific verification\.
#### Effect of Non\-Linearity \(Linear Baseline\)\.
The reviewer rightfully questions the incremental effect of ReLU\. In a purely linear network \(Identity activation\), the interference noiseϵ\\epsilonwould be symmetric \(𝔼\[ϵ\]=0\\mathbb\{E\}\[\\epsilon\]=0\)\. By the Law of Large Numbers, this noise would average out over depth\. The ”Ratchet Effect” we identify is strictly a consequence of the non\-linearity breaking this symmetry \(𝔼\[ReLU\(ϵ\)\]\>0\\mathbb\{E\}\[\\text\{ReLU\}\(\\epsilon\)\]\>0\)\. Thus, a linear baseline would showzerostructural drift, effectively serving as the trivial lower bound in our Figure 3\.
#### Comparison to Advanced SAEs\.
Our current analysis focuses on standard SAEs with spherical dictionaries\. Recent variants such as Whitened SAEs or Gated SAEs explicitly aim to reduce the coherenceμ\(D\)\\mu\(D\)or dynamically suppress ghost activations\. Within our framework, these methods can be interpreted as techniques to artificially lower theμeff\\mu\_\{eff\}or increaseΔgap\\Delta\_\{gap\}\. While we do not benchmark them here, our theory predicts they would shift the phase transition curve to the right \(higherγ∗\\gamma^\{\*\}\), but not eliminate the fundamental geometric bottleneck\.
### G\.6Other Limitations and Computational Feasibility
#### Idealized Assumptions vs\. Real Structure\.
Our derivation relies on the Random Spherical Model \(Assumption 1\)\. While real\-world LLM features exhibit hierarchical clustering, we argue that the random model provides atheoretical upper boundon stability\. Structured interference \(e\.g\., semantic polysemy\) typically creates “dense“ regions in the cone, accelerating collapse compared to the isotropic spread assumed here\. Thus, our thresholdγ∗\\gamma^\{\*\}serves as a necessary, if not sufficient, condition for stability in structured dictionaries\.
#### Computational Verification\.
We acknowledge that explicitly computing the feature cone𝒦S\\mathcal\{K\}\_\{S\}for large\-scale models \(m≈106m\\approx 10^\{6\}\) is computationally intractable, as it requires vertex enumeration of a high\-dimensional polytope\. Our framework is intended as ananalytical toolto predict scaling laws, rather than a runtime algorithm\. However, the scalar metrics derived \(e\.g\., Mean Width\) can be efficiently estimated via Monte Carlo sampling, as demonstrated in our CLEVR experiments\.Similar Articles
Feature Starvation as Geometric Instability in Sparse Autoencoders
This paper identifies feature starvation in sparse autoencoders as a geometric instability and proposes adaptive elastic net SAEs (AEN-SAEs) to mitigate it without heuristics.
Feature Rivalry in Sparse Autoencoder Representations: A Mechanistic Study of Uncertainty-Driven Feature Competition in LLMs
This research paper introduces 'Feature Rivalry' in Sparse Autoencoder representations as a mechanistic signature of uncertainty in LLMs. Using Gemma-2-2B, the study demonstrates that negatively correlated feature pairs localize uncertainty to specific layers and causally influence model outputs.
Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders
This paper studies seed dependence in sparse autoencoders, finding that stable features carry most predictive signal while unstable features reflect reproducible low-dimensional subspaces.
A Geometric View for Understanding Concept Learning and Neuron Interpretation in Sparse Autoencoders
This paper proposes a unified geometric framework for understanding concept learning and neuron interpretation in sparse autoencoders, formalizing concepts as sets and defining detection, separation, and approximation. It provides error bounds, capacity constraints, and links to formal concept analysis, with experiments on synthetic data.
Composition Collapse: Stable Factual Knowledge Does Not Imply Compositional Reasoning
This paper introduces 'composition collapse', a phenomenon where language models with stable factual knowledge still fail to compose that knowledge into correct multi-hop reasoning, and proposes a double-gate protocol to isolate composition failure from atomic knowledge instability.