Feature Lottery? A Bifurcation Theory of Concept Emergence
Summary
This paper introduces a bifurcation theory of representation dynamics to detect when neural networks acquire structured representations during training, using a Hessian analysis of a GMM probe. The resulting ratio β/β_c serves as a label-free phase coordinate that predicts the onset of usable structure and can forecast feature interpretability in sparse autoencoders early in training.
View Cached Full Text
Cached at: 05/26/26, 08:59 AM
# Feature Lottery? A Bifurcation Theory of Concept Emergence
Source: [https://arxiv.org/html/2605.24057](https://arxiv.org/html/2605.24057)
###### Abstract
Neural networks acquire structured representations at specific moments during training, yet identifying these transitions typically relies on retrospective, label\-dependent metrics\. We introduce a bifurcation theory of representation dynamics to detect these moments in real time\. By analyzing a passive GMM probe attached to the evolving encoder, we show that the onset of structure can be identified with a supercritical pitchfork bifurcation driven by the loss Hessian\. The system exhibits a theoretically predictable zero\-crossing \(βc\\beta\_\{c\}\) that, compared to the network’s current state \(β\\beta\), yields a dynamic ratioβ\(t\)/βc\(t\)\\beta\(t\)/\\beta\_\{c\}\(t\)\. This ratio acts as a universal, label\-free phase coordinate for representation dynamics computable entirely from hidden states\.
We empirically validate four distinct transition regimes predicted by this coordinate across diverse settings: SAEs on language models \(Pythia\), SSL \(CIFAR\), and grokking \(modular arithmetic\)\. Crucially, under finite dissipation, the macroscopic symmetry\-breaking can lag the initial zero\-crossing by orders of magnitude: providing a rigorous dynamical account of the delayed escape observed in grokking\. At the microscopic level, the theory predicts that the bifurcation creates a shared unstable subspace, forcing a collective symmetry breaking\. We term this the*feature lottery*in SAE training: a feature’s terminal interpretability becomes predictable remarkably early\. By only5%5\\%of training, early atom purity robustly predicts final convergence purity, with top\-decile early atoms achieving over12×12\\timesthe baseline purity at the end of training\.
Beyond explaining concept emergence,β/βc\\beta/\\beta\_\{c\}provides a highly practical tool\. It functions as an early\-warning indicator for training health: detecting the onset of usable structure, the crystallization of feature identity, and the onset of representational collapse epochs before downstream metrics react\.
Interactive demo:[https://fumingyang\-felix\.github\.io/feature\-lottery\-demo/](https://fumingyang-felix.github.io/feature-lottery-demo/)
## 1Introduction
Modern networks do not merely fit labels or reconstruct inputs; during training, their internal states reorganize into discrete, reusable directions that behave like concepts\. Yet we typically notice this reorganization only after the fact: by probing with labels, by inspecting downstream accuracy, or by mechanistic analysis of a fully trained model\. What is missing is a*label\-free dynamical signal*of when such concept structure first becomes available, a quantity that \(watched live during training\) tells us whether and when a network has just acquired a usable internal representation, and ideally*which*parts of the representation are about to become meaningful\.
#### Our angle\.
We provide such a signal, derived from a single Hessian analysis\. Attach a passiveKK\-prototype isotropic Gaussian\-mixture \(GMM\) head with shared learned precisionβ\\betato the outputz=enc\(x\)z\\\!=\\\!\\mathrm\{enc\}\(x\)of a given encoder representation, and analyze the Hessian of its negative log\-likelihood at the*symmetric collapsed state*μ1=⋯=μK=z¯\\mu\_\{1\}\\\!=\\\!\\cdots\\\!=\\\!\\mu\_\{K\}\\\!=\\\!\\bar\{z\}\. The analysis \(Section[3](https://arxiv.org/html/2605.24057#S3)\) yields a critical precision
βc=1λmax\(Cov\(z\)\)\\boxed\{\\;\\beta\_\{c\}\\;=\\;\\frac\{1\}\{\\lambda\_\{\\max\}\(\\mathrm\{Cov\}\(z\)\)\}\\;\}\(1\)at which the lowest eigenvalue of the loss Hessian crosses zero; aboveβc\\beta\_\{c\}the prototypes pitchfork along the principal eigenvector ofCov\(z\)\\mathrm\{Cov\}\(z\)\. We adopt this Hessian\-pitchfork event as our*operational*definition of concept emergence: the moment at which the encoder’s representation first admits a class\-alignedKK\-prototype decomposition\. The label\-free indicator isβ\(t\)/βc\(t\)\\beta\(t\)/\\beta\_\{c\}\(t\), computable from the encoder’s hidden state and the GMM probe alone\. When the encoder is itself learning,βc\(t\)\\beta\_\{c\}\(t\)becomes endogenous and we proveβ\(t\)\\beta\(t\)andβc\(t\)\\beta\_\{c\}\(t\)must cross at some finite training time \(Proposition[1](https://arxiv.org/html/2605.24057#Thmproposition1)\)\. A subtlety, made precise in Remark[1](https://arxiv.org/html/2605.24057#Thmremark1), is that the crossing event marks when the symmetric state becomes*unstable*, not when the broken\-symmetry state becomes*macroscopically observable*\. The lag between the two is controlled by the encoder’s dissipation, and can range from essentially zero \(in well\-trained SSL\) to thousands of steps \(in grokking\)\.
#### The sharpest prediction is per\-atom\.
At the crossing event, the unstable subspace atβ=βc\\beta\\\!=\\\!\\beta\_\{c\}is shared by allK−1K\\\!\-\\\!1anti\-symmetric modes \(App\.[A\.2](https://arxiv.org/html/2605.24057#A1.SS2)\)\. The bifurcation is therefore a*collective*event with a*per\-atom signature*: each atom must select a specific direction from a common unstable manifold, its choice driven by initialization noise and the cubic terms of the pitchfork normal form\. This per\-atom prediction is empirically sharp\. In SAE training on frozen Pythia\-160M layer 6, per\-atom POS purity is at noise floor before the bifurcation and acquires predictive content at onset; ranking atoms by their step\-1,0001\{,\}000POS purity \(5%5\\%of training\) already recovers a top decile whose convergence purity is12×12\\timesthe uniform\-random baseline \(Section[5](https://arxiv.org/html/2605.24057#S5)\)\. We refer to this as the SAE*feature lottery*, the bifurcation theory’s sharpest and most unexpected empirical consequence; it is the SAE\-level analogue of the lottery\-ticket framing ofFrankle and Carbin \([2019](https://arxiv.org/html/2605.24057#bib.bib19)\), with the drawing event identified explicitly as the first phase transition during training\.
#### Contributions\.
1. 1\.Theory: an endogenous critical point with post\-critical metastability\.A Hessian\-pitchfork prediction of when a given encoder representation first admits aKK\-prototype decomposition, obtained from a passive GMM probe on the encoder output, with an existence proof for the endogenous critical point \(Proposition[1](https://arxiv.org/html/2605.24057#Thmproposition1); the crossingβ=βc\\beta\\\!=\\\!\\beta\_\{c\}happens at a finite time*conditional on*the encoder eventually spreading the latent enough for the GMM to resolve clusters\) and a separate prediction of post\-critical metastability under finite dissipation \(Remark[1](https://arxiv.org/html/2605.24057#Thmremark1)\)\. Both are new relative to the static soft\-KK\-means criticality ofRoseet al\.\([1990](https://arxiv.org/html/2605.24057#bib.bib7)\): Rose gives the critical temperature on a frozen dataset; we give the dynamic crossing theorem whenβc\(t\)\\beta\_\{c\}\(t\)co\-evolves with the encoder, and identify a post\-critical metastable regime that the static analysis cannot exhibit\.
2. 2\.The theory’s sharpest empirical consequence: an SAE feature lottery at5%5\\%of training\.At the crossing event the unstable subspace is shared by allK−1K\\\!\-\\\!1anti\-symmetric modes, so atoms must select directions from a common manifold during the bifurcation\. We verify this in SAE training on frozen language\-model activations \(Sec\.[5](https://arxiv.org/html/2605.24057#S5)\): per\-atom POS purity is at noise floor pre\-onset \(ρid≈0\.03\\rho\_\{\\mathrm\{id\}\}\\\!\\approx\\\!0\.03\), and by step1,0001\{,\}000\(5%5\\%of training\), identity\-matchedρid=\+0\.41±0\.04\\rho\_\{\\mathrm\{id\}\}\\\!=\\\!\+0\.41\\\!\\pm\\\!0\.04\(3 seeds, allp<10−80p\\\!<\\\!10^\{\-80\}\)\. The top decile of atoms ranked at5%5\\%achieves convergence POS purity0\.82±0\.030\.82\\\!\\pm\\\!0\.03,12\.3±0\.4×12\.3\\\!\\pm\\\!0\.4\\timesthe uniform\-random baseline of0\.0670\.067\. The effect replicates across soft\-L1 and architectural top\-KKSAEs, and acrossK∈\{256,…,8192\}K\\\!\\in\\\!\\\{256,\\dots,8192\\\}withρid∈\[0\.26,0\.41\]\\rho\_\{\\mathrm\{id\}\}\\in\[0\.26,0\.41\]\. This is an atom\-level, post\-onset analogue of the lottery\-ticket framing ofFrankle and Carbin \([2019](https://arxiv.org/html/2605.24057#bib.bib19)\)\.
3. 3\.Empirical universality of the bifurcation arc\.The predicted trajectory in\(log\(β/βc\),logNC1\)\(\\log\(\\beta/\\beta\_\{c\}\),\\,\\log\\mathrm\{NC1\}\)is governed by three binary kinematic axes: initial sub/supercriticality, the post\-onset race betweenβ\(t\)\\beta\(t\)andβc\(t\)\\beta\_\{c\}\(t\), and the dissipation rate \(Sec\.[4\.3](https://arxiv.org/html/2605.24057#S4.SS3), App\.[I](https://arxiv.org/html/2605.24057#A9)\)\. The axes predict which kinematic regimes are accessible to which feature\-learning methods\. We verify the four regimes that arise in standard pipelines; full V \(SAE on frozen Pythia layer 6\), fold\-back \(DINO/SimCLR on CIFAR\-10/100, with magnitude controlled by data complexity\), delayed escape \(grokking on modular arithmetic, with WD\-monotonic escape timeτesc∝WD−1\.23\\tau\_\{\\mathrm\{esc\}\}\\propto\\mathrm\{WD\}^\{\-1\.23\}across six WD levels, and0/30/3escape at WD=0=\\\!0\), and no arc \(rotation\-prediction control\)\.
4. 4\.K\-sweep: lottery isKK\-stable, POS\-purity Pareto isKK\-monotonic butKK\-confounded\.At fixed3%3\\%top\-KKsparsity, the lotteryρid\\rho\_\{\\mathrm\{id\}\}is stable acrossK∈\{256,…,8192\}K\\in\\\{256,\\dots,8192\\\}\(range\[\+0\.26,\+0\.41\]\[\+0\.26,\+0\.41\]\)\. Per\-atom POS purity decreases monotonically withKKfrom0\.7250\.725to0\.3700\.370\(Tab\.[4](https://arxiv.org/html/2605.24057#S5.T4)\), while reconstruction MSE moves the other way \(0\.11→0\.0150\.11\\to 0\.015\)\. Because POS has only1515classes, this purity\-vs\-KKtrend is partly structural: small\-KKatoms each correspond to coarser token\-cluster partitions that more easily align with the1515\-way POS partition\. We therefore do*not*recommend smallKKon the basis of POS purity alone; aKK\-unbiased interpretability metric \(e\.g\. causal mediation or LLM\-as\-judge\) is required to make a substantive recommendation \(Sec\.[5\.4](https://arxiv.org/html/2605.24057#S5.SS4)\)\.
5. 5\.A label\-free training diagnostic\.β/βc\\beta/\\beta\_\{c\}identifies the encoder’s current act from hidden states alone, well before downstream metrics respond\. In grokking, the indicator at step100100already places the trajectory in Act 2:∼\\sim8,400 steps beforetest\_acc\\mathrm\{test\\\_acc\}moves\. In DINO from\-scratch with collapse modes \(Sec\.[6\.1](https://arxiv.org/html/2605.24057#S6.SS1)\),β/βc\\beta/\\beta\_\{c\}leads cluster accuracy by 8 epochs in the gradual mode; in mid\-training interventions \(Sec\.[6\.2](https://arxiv.org/html/2605.24057#S6.SS2)\) it responds within a single batch while training loss remains within noise\.
## 2Related work
#### Concept emergence and mechanistic interpretability\.
A growing body of work in mechanistic interpretability studies how internal computations of trained models implement particular “concepts\.”Nandaet al\.\([2023](https://arxiv.org/html/2605.24057#bib.bib13)\)show that the post\-grok solution of modular addition is a Fourier multiplication algorithm distributed across the network’s embedding, attention, and MLP layers; sparse autoencoders on frozen language\-model activations\(Brickenet al\.,[2023](https://arxiv.org/html/2605.24057#bib.bib16); Gaoet al\.,[2025](https://arxiv.org/html/2605.24057#bib.bib17); Templetonet al\.,[2024](https://arxiv.org/html/2605.24057#bib.bib18)\)extract interpretable directions in feature space\. These works identify concepts*post\-hoc*on a trained model: features are evaluated for interpretability after training converges, and the typical assumption is that training longer yields better features\. Our framework is complementary along two axes\. First, we provide a label\-free*dynamical*signal of when such structure first becomes available during training \(Sec\.[4](https://arxiv.org/html/2605.24057#S4)\)\. Second, we show that in SAEs the bifurcation onset is a per\-atom assignment event whose outcome predicts atom\-level interpretability nineteen thousand training steps in advance \(Sec\.[5](https://arxiv.org/html/2605.24057#S5)\)\. This reframes SAE feature emergence as a structured lottery rather than a gradual refinement\.
#### Neural collapse\.
The neural\-collapse literature\(Papyanet al\.,[2020](https://arxiv.org/html/2605.24057#bib.bib1)\)and the subsequent Unconstrained Features Model analyses\(Mixonet al\.,[2022](https://arxiv.org/html/2605.24057#bib.bib2); Tirer and Bruna,[2022](https://arxiv.org/html/2605.24057#bib.bib3); Zhouet al\.,[2022](https://arxiv.org/html/2605.24057#bib.bib4); Súkeníket al\.,[2024](https://arxiv.org/html/2605.24057#bib.bib5)\)characterize the static terminal\-phase geometry of supervised classification\.Wang and Palmer \([2023](https://arxiv.org/html/2605.24057#bib.bib6)\)recover an analogous structure in supervised contrastive learning via an information\-bottleneck argument\. These analyses describe the endpoint, not the dynamics by which an encoder reaches it; our framework supplies the missing dynamics and predicts the timing of concept emergence without labels\.
#### Deterministic annealing and rate\-distortion clustering\.
Roseet al\.\([1990](https://arxiv.org/html/2605.24057#bib.bib7)\)obtained the critical temperatureTc=2λmax\(Σ\)T\_\{c\}=2\\lambda\_\{\\max\}\(\\Sigma\)for softKK\-means by externally annealingTT\. In our conventionβ=2/T\\beta\\\!=\\\!2/T, this is the external\-β\\betalimit of our analysis\. The novelty here is thatβ\\betais endogenous andβc\(t\)\\beta\_\{c\}\(t\)moves with the encoder\.
#### Self\-supervised collapse\.
Jinget al\.\([2022](https://arxiv.org/html/2605.24057#bib.bib8)\); Huaet al\.\([2021](https://arxiv.org/html/2605.24057#bib.bib15)\)analyze dimensional collapse in contrastive and non\-contrastive self\-supervised methods\.Caronet al\.\([2021](https://arxiv.org/html/2605.24057#bib.bib14)\)introduces DINO with centering and sharpening regularizers specifically to prevent collapse\. Our diagnostic experiments \(Sec\.[6](https://arxiv.org/html/2605.24057#S6)\) take DINO collapse modes as a controlled testbed\.
#### Grokking and emergence\.
Poweret al\.\([2022](https://arxiv.org/html/2605.24057#bib.bib11)\)discovered that small transformers on modular arithmetic exhibit a long memorization plateau followed by a sudden generalization transition;Nandaet al\.\([2023](https://arxiv.org/html/2605.24057#bib.bib13)\)mechanistically identified the in\-network DFT circuit responsible\. The plateau\-then\-sudden\-transition pattern is reminiscent of the saddle\-to\-saddle dynamics characterized in deep linear networks bySaxeet al\.\([2014](https://arxiv.org/html/2605.24057#bib.bib12)\), where learning proceeds through a sequence of loss\-landscape saddles separated by long plateaus\. Our framework gives this picture a concrete representation\-geometric content: the plateau is the metastable post\-critical regime \(β\>βc\\beta\\\!\>\\\!\\beta\_\{c\}butε\\varepsilonstill microscopic; Remark[1](https://arxiv.org/html/2605.24057#Thmremark1)\), and the dissipation strength sets the escape time\. We revisit grokking in Sec\.[4\.2](https://arxiv.org/html/2605.24057#S4.SS2)as the cleanest empirical instance of post\-critical metastability in our framework\.
#### Phase transitions in deep learning\.
Wang and Ziyin \([2022](https://arxiv.org/html/2605.24057#bib.bib9)\)analyzes posterior collapse in linear latent\-variable models, andZiyin and Ueda \([2023](https://arxiv.org/html/2605.24057#bib.bib10)\)prove first\- and second\-order phase transitions in deep linear networks under a statistical\-mechanics framing\. Both connect collapse\-type phenomena to phase transitions in regularized learning\. We share this perspective and contribute a closed\-form Hessian\-based predictor of the transition point in unsupervised and self\-supervised settings\.
#### Lottery\-ticket framing \(weak analogy\)\.
Frankle and Carbin \([2019](https://arxiv.org/html/2605.24057#bib.bib19)\)introduce the lottery\-ticket hypothesis for supervised classifiers: there exist sparse subnetworks identifiable already at initialization that train to the same accuracy as the full network\. Our finding for SAE atoms \(Sec\.[5](https://arxiv.org/html/2605.24057#S5)\) is a*weaker analogue*in feature space: early per\-atom POS purity at the bifurcation onset is a useful ranking signal for converged per\-atom POS purity \(identity\-matched Spearmanρ=\+0\.41±0\.04\\rho=\+0\.41\\pm 0\.04between step\-1,0001\{,\}000and step\-20,00020\{,\}000\)\. The analogy is weaker thanFrankle and Carbin \([2019](https://arxiv.org/html/2605.24057#bib.bib19)\)in two important senses: \(1\) identification happens at the post\-onset bifurcation event rather than at initialization, and \(2\) achievement requires continued joint training, not isolated retraining of a sparse subnetwork\. We use the lottery framing for its spirit \(*what matters is decided early*\), not as a structural isomorphism\.
## 3Theory
### 3\.1Setup
We attach aKK\-component isotropic GMM with shared precisionβ\\betaand component meansμk∈ℝd\\mu\_\{k\}\\in\\mathbb\{R\}^\{d\}to the encoder outputz=enc\(x\)z=\\mathrm\{enc\}\(x\):
p\(z∣θ\)=1K∑k=1K𝒩\(z∣μk,β−1I\),ℒμ\(μ,β\)=−𝔼z\[LSEk\(−β2‖z−μk‖2\)\],p\(z\\mid\\theta\)\\;=\\;\\frac\{1\}\{K\}\\sum\_\{k=1\}^\{K\}\\mathcal\{N\}\(z\\mid\\mu\_\{k\},\\beta^\{\-1\}I\),\\qquad\\mathcal\{L\}\_\{\\mu\}\(\\mu,\\beta\)\\;=\\;\-\\mathbb\{E\}\_\{z\}\\\!\\Bigl\[\\,\\mathrm\{LSE\}\_\{k\}\\\!\\bigl\(\-\\tfrac\{\\beta\}\{2\}\\\|z\-\\mu\_\{k\}\\\|^\{2\}\\bigr\)\\Bigr\],\(2\)up to terms independent ofμ\\mu\. We denote the data covarianceΣ=Cov\(z\)\\Sigma\\\!=\\\!\\mathrm\{Cov\}\(z\)with eigenvaluesσ12≥⋯≥σd2\\sigma\_\{1\}^\{2\}\\\!\\geq\\\!\\cdots\\\!\\geq\\\!\\sigma\_\{d\}^\{2\}, the responsibilitiesp\(k∣z\)=softmaxk\(−β2‖z−μk‖2\)p\(k\\\!\\mid\\\!z\)=\\mathrm\{softmax\}\_\{k\}\(\-\\tfrac\{\\beta\}\{2\}\\\|z\-\\mu\_\{k\}\\\|^\{2\}\), and the symmetric collapsed state𝒮0=\{μk=z¯\}\\mathcal\{S\}\_\{0\}=\\\{\\mu\_\{k\}=\\bar\{z\}\\\}\.
### 3\.2Hessian at the symmetric state
At𝒮0\\mathcal\{S\}\_\{0\}we havep\(k∣z\)=1/Kp\(k\\\!\\mid\\\!z\)=1/Kand∇μℒμ=0\\nabla\_\{\\mu\}\\mathcal\{L\}\_\{\\mu\}=0\. Differentiating once more and evaluating at𝒮0\\mathcal\{S\}\_\{0\},
∂2ℒμ∂μka∂μlb\|𝒮0=βKδklδab−β2K\(δkl−1K\)Σab\.\\Bigl\.\\frac\{\\partial^\{2\}\\mathcal\{L\}\_\{\\mu\}\}\{\\partial\\mu\_\{k\}^\{a\}\\,\\partial\\mu\_\{l\}^\{b\}\}\\Bigr\|\_\{\\mathcal\{S\}\_\{0\}\}\\;=\\;\\tfrac\{\\beta\}\{K\}\\,\\delta\_\{kl\}\\,\\delta^\{ab\}\\;\-\\;\\tfrac\{\\beta^\{2\}\}\{K\}\\bigl\(\\delta\_\{kl\}\-\\tfrac\{1\}\{K\}\\bigr\)\\,\\Sigma^\{ab\}\.\(3\)Eigenvectors decompose into separable perturbationsξlb=wlub\\xi\_\{l\}^\{b\}=w\_\{l\}u^\{b\}withw∈ℝKw\\\!\\in\\\!\\mathbb\{R\}^\{K\}andu∈ℝdu\\\!\\in\\\!\\mathbb\{R\}^\{d\}\. Two channels arise: the symmetric channel \(wwconstant\) is always stable with eigenvaluesβ/K\\beta/K; the anti\-symmetric channel \(∑lwl=0\\sum\_\{l\}w\_\{l\}=0,\(K−1\)\(K\{\-\}1\)\-fold degenerate\) has spatial eigenvalues
λi⟂\(β\)=βK\(1−βσi2\),i=1,…,d\.\\lambda^\{\\perp\}\_\{i\}\(\\beta\)\\;=\\;\\tfrac\{\\beta\}\{K\}\\,\\bigl\(1\-\\beta\\,\\sigma\_\{i\}^\{2\}\\bigr\),\\qquad i=1,\\dots,d\.\(4\)The lowest such eigenvalue isλ1⟂\(β\)=\(β/K\)\(1−βλmax\(Σ\)\)\\lambda^\{\\perp\}\_\{1\}\(\\beta\)=\(\\beta/K\)\(1\-\\beta\\,\\lambda\_\{\\max\}\(\\Sigma\)\), and crosses zero exactly at the critical precision in equation equation[1](https://arxiv.org/html/2605.24057#S1.E1)\. The unstable direction is the principal eigenvector ofΣ\\Sigmain spatial space combined with any anti\-symmetricwwin component space\. Projecting the dynamics onto this slow direction yields the supercritical pitchfork normal form
ε˙=\(β−βc\)ε−αε3\+noise,α\>0,\\dot\{\\varepsilon\}\\;=\\;\(\\beta\-\\beta\_\{c\}\)\\,\\varepsilon\\;\-\\;\\alpha\\,\\varepsilon^\{3\}\\;\+\\;\\text\{noise\},\\quad\\alpha\>0,\(5\)shown in Fig\.[1](https://arxiv.org/html/2605.24057#S3.F1)\(a\)\. A full derivation, including the parametrization correction relative toRoseet al\.\([1990](https://arxiv.org/html/2605.24057#bib.bib7)\)and the explicit eigendecomposition in component space \(Theorem[1](https://arxiv.org/html/2605.24057#Thmtheorem1)\), is in Appendix[A](https://arxiv.org/html/2605.24057#A1)\(§[A\.1](https://arxiv.org/html/2605.24057#A1.SS1)\)\.
Figure 1:Theory\.\(a\)The supercritical pitchfork atβc\\beta\_\{c\}: belowβc\\beta\_\{c\}the symmetric collapsed state is stable; aboveβc\\beta\_\{c\}it becomes a saddle and two stable broken\-symmetry branches emerge\.\(b\)Endogenous criticality \(Proposition[1](https://arxiv.org/html/2605.24057#Thmproposition1)\): when the encoder is itself learning,βc\(t\)\\beta\_\{c\}\(t\)co\-evolves withCov\(z\(t\)\)\\mathrm\{Cov\}\(z\(t\)\); under mild monotonicity assumptionsβ\(t\)\\beta\(t\)andβc\(t\)\\beta\_\{c\}\(t\)must cross at some finite timet⋆t^\{\\star\}\.
### 3\.3Hierarchy and reverse traversal
The analysis recurses within each post\-primary supercluster, giving a sequence of secondary critical precisionsβc\(2\)=1/λmax\(Σwithin\)\\beta\_\{c\}^\{\(2\)\}=1/\\lambda\_\{\\max\}\(\\Sigma\_\{\\mathrm\{within\}\}\)\. The pitchfork equation[5](https://arxiv.org/html/2605.24057#S3.E5)is also reversible: decreasingβ\\betacontinuously merges the broken\-symmetry branches back into𝒮0\\mathcal\{S\}\_\{0\}\. App\.[E](https://arxiv.org/html/2605.24057#A5)confirms both on toy data \(hierarchy to four decimals; forward overshootβ⋆/βc≈1\.3\\beta^\{\\star\}/\\beta\_\{c\}\\\!\\approx\\\!1\.3, reverse tracking within≤4%\\leq 4\\%\) and on CIFAR\-scale encoders \(App\.[F](https://arxiv.org/html/2605.24057#A6)\)\.
### 3\.4Endogenous criticality
Replace the fixed dataset\{zn\}\\\{z\_\{n\}\\\}with\{zn\(t\)\}=\{enc\(xn;φ\(t\)\)\}\\\{z\_\{n\}\(t\)\\\}=\\\{\\mathrm\{enc\}\(x\_\{n\};\\varphi\(t\)\)\\\}, where the encoder parametersφ\\varphievolve under an upstream loss\. The critical precision becomes time\-dependent:
βc\(t\)=1λmax\(Cov\(z\(t\)\)\),\\beta\_\{c\}\(t\)\\;=\\;\\frac\{1\}\{\\lambda\_\{\\max\}\(\\mathrm\{Cov\}\(z\(t\)\)\)\},\(6\)and the bifurcation occurs at the momentβ\(t⋆\)=βc\(t⋆\)\\beta\(t^\{\\star\}\)=\\beta\_\{c\}\(t^\{\\star\}\)\.
###### Proposition 1\(Endogenous critical point\)\.
Suppose
1. 1\.β\(t\)\\beta\(t\)is asymptotically non\-decreasing andlim inft→∞β\(t\)\>c1\>0\\liminf\_\{t\\to\\infty\}\\beta\(t\)\>c\_\{1\}\>0;
2. 2\.βc\(t\)\\beta\_\{c\}\(t\)is asymptotically non\-increasing on average;
3. 3\.lim supt→∞βc\(t\)<c1\\limsup\_\{t\\to\\infty\}\\beta\_\{c\}\(t\)<c\_\{1\}\.
Thenβ\(t\)\\beta\(t\)andβc\(t\)\\beta\_\{c\}\(t\)cross at some finitet⋆t^\{\\star\}; at the crossing, the GMM’s symmetric state becomes unstable\.
The proof, given in Appendix[B](https://arxiv.org/html/2605.24057#A2), is essentially a continuity argument:β\(t\)−βc\(t\)\\beta\(t\)\-\\beta\_\{c\}\(t\)goes from−\|β\(0\)−βc\(0\)\|<0\-\|\\beta\(0\)\\\!\-\\\!\\beta\_\{c\}\(0\)\|<0\(att=0t\\\!=\\\!0, before learning\) to a strictly positive lim\-inf, and must cross zero\. The substantive content is in the hypotheses: \(1\) holds for any likelihood\-maximizing GMM step, and \(2\) is the assertion that the encoder spreads the latent, which is the defining property of any information\-preserving feature\-learning objective\. The novelty relative toRoseet al\.\([1990](https://arxiv.org/html/2605.24057#bib.bib7)\)is not the IVT step but the framing:βc\\beta\_\{c\}becomes a*time\-dependent*observable that co\-evolves with the encoder, turning the crossing event into a predictable training\-time phenomenon rather than a static property of a frozen dataset\.
### 3\.5Post\-critical metastability: the bifurcation timing problem
Proposition[1](https://arxiv.org/html/2605.24057#Thmproposition1)establishes that the crossing eventβ\(t⋆\)=βc\(t⋆\)\\beta\(t^\{\\star\}\)\\\!=\\\!\\beta\_\{c\}\(t^\{\\star\}\)exists; the symmetric state𝒮0\\mathcal\{S\}\_\{0\}becomes a saddle att⋆t^\{\\star\}\. The proposition is silent, however, on*when the broken\-symmetry state becomes macroscopically observable*\. The order parameterε\\varepsilonintroduced in equation[5](https://arxiv.org/html/2605.24057#S3.E5)obeys
ε˙=\(β−βc\)ε−αε3\+η\(t\),\\dot\{\\varepsilon\}\\;=\\;\(\\beta\-\\beta\_\{c\}\)\\,\\varepsilon\\;\-\\;\\alpha\\,\\varepsilon^\{3\}\\;\+\\;\\eta\(t\),\(7\)whereη\(t\)\\eta\(t\)is a noise/dissipation term inherited from the encoder’s training dynamics\. Just past the crossing, the unstable mode’s linear growth rate is\(β−βc\)→0\+\(\\beta\-\\beta\_\{c\}\)\\\!\\to\\\!0^\{\+\}, soε\\varepsilongrows exponentially*from the noise scale*on a characteristic timescale that scales as1/\(β−βc\)1/\(\\beta\-\\beta\_\{c\}\)\.
## 4The bifurcation arc across feature\-learning methods
The Hessian\-pitchfork prediction of Sec\.[3](https://arxiv.org/html/2605.24057#S3)produces a three\-phase trajectory in\(log\(β/βc\),logNC1\)\\bigl\(\\log\(\\beta/\\beta\_\{c\}\),\\,\\log\\mathrm\{NC1\}\\bigr\)space:*pre\-critical*\(β<βc\\beta\\\!<\\\!\\beta\_\{c\}, no class\-aligned clustering, NC1 drifts slowly upward\),*critical peak*\(β≈βc\\beta\\\!\\approx\\\!\\beta\_\{c\},𝒮0\\mathcal\{S\}\_\{0\}becomes a saddle, NC1 peaks\),*post\-critical descent*\(β\>βc\\beta\\\!\>\\\!\\beta\_\{c\}, prototypes separate andlogNC1\\log\\mathrm\{NC1\}falls linearly inlog\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)\)\. The shape*observed in a particular setup*depends on \(i\) where the encoder begins \(log\(β/βc\)≶0\\log\(\\beta/\\beta\_\{c\}\)\\\!\\lessgtr\\\!0att=0t\\\!=\\\!0\), \(ii\) the relative rates ofβ\(t\)\\beta\(t\)andβc\(t\)\\beta\_\{c\}\(t\)during training, and \(iii\) the dissipation rate \(Remark[1](https://arxiv.org/html/2605.24057#Thmremark1)\)\. The combination yields four empirically distinguishable shapes:*full V*\(frozenβc\\beta\_\{c\}; one leg both ways\),*fold\-back*\(βc\\beta\_\{c\}overtakesβ\\betapost\-onset, with magnitude controlled by data complexity; mild on CIFAR\-10, strong on CIFAR\-100\),*delayed escape*\(under\-dissipated post\-critical metastability, as in grokking\), and*no arc*\(negative control: no clustering pressure\)\. We verify all four in two complementary experiments below\.
### 4\.1Four shapes from feature\-learning trajectories
Figure 2:Four observable shapes of the bifurcation arc, from the same Hessian\-pitchfork mechanism\.Top row: schematic of each regime; bottom row: empirical realization\. \(i\)*Full V*on SAE / frozen Pythia layer 6 \(NC1 in sparse\-code basis; star = V trough\)\. \(ii\)*Fold\-back spectrum*on DINO / CIFAR\-10 \(light green, mild\) and CIFAR\-100 \(dark green, strong\), fold magnitude controlled by data complexity\. \(iii\)*Delayed escape*on grokking \(p=97p\\\!=\\\!97, WD=1\.0=\\\!1\.0; three\-act structure detailed in Sec\.[4\.2](https://arxiv.org/html/2605.24057#S4.SS2)\)\. \(iv\)*No arc*on rotation\-prediction \(negative control\)\. NC1 axes are not comparable across panels \(sparse\-code, backbone, residual\-stream\); the trajectory*shape*is the prediction\. The 3\-axis taxonomy explaining why exactly these four shapes appear is given in App\.[I](https://arxiv.org/html/2605.24057#A9)\.We probe all four shapes here, across one sparse\-coding\-on\-frozen\-LM setup, two self\-supervised setups \(CIFAR\-10 and CIFAR\-100, illustrating the fold\-back spectrum\), one delayed\-escape setup \(grokking\), and one negative control \(rotation prediction\), with a shared joint\-detached GMM probe protocol \(K=10K\\\!=\\\!10,lrμ=5×10−3\\mathrm\{lr\}\_\{\\mu\}\\\!=\\\!5\\\!\\times\\\!10^\{\-3\},lrβ=10−2\\mathrm\{lr\}\_\{\\beta\}\\\!=\\\!10^\{\-2\},logβ0=−2\.5\\log\\beta\_\{0\}\\\!=\\\!\-2\.5;KprobeK\_\{\\mathrm\{probe\}\}\-robustness in App\.[K](https://arxiv.org/html/2605.24057#A11); full hyperparameter tables in App\.[M](https://arxiv.org/html/2605.24057#A13)\) so thatβGMM\(t\)\\beta\_\{\\mathrm\{GMM\}\}\(t\)is directly comparable across methods\. Figure[2](https://arxiv.org/html/2605.24057#S4.F2)shows the four shapes; Table[1](https://arxiv.org/html/2605.24057#S4.T1)reports the local descent slopekkwhere applicable\. The grokking panel of Fig\.[2](https://arxiv.org/html/2605.24057#S4.F2)shows a single canonical run for visual parity with the other panels; the full three\-seed analysis with the three\-act decomposition \(Act 1 / Act 2 / Act 3 timing and WD\-dependent escape time, withτesc∝WD−1\.23\\tau\_\{\\mathrm\{esc\}\}\\propto\\mathrm\{WD\}^\{\-1\.23\}\) is in Sec\.[4\.2](https://arxiv.org/html/2605.24057#S4.SS2)/ Fig\.[3](https://arxiv.org/html/2605.24057#S4.F3)\.
#### \(i\) Full V\.
The SAE on Pythia\-160M layer 6 freezes the upstream encoder, holdingβc\\beta\_\{c\}constant\. The entire motion inlog\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)is driven by the SAE’s own precision growing past its critical point\. NC1 \(in the sparse\-code basis,K=2048K\\\!=\\\!2048\) rises slightly during pre\-critical buildup, peaks, and then descends monotonically as predicted by the post\-critical regime \(r=−0\.97r\\\!=\\\!\-0\.97,n=9n\\\!=\\\!9\)\. The full V is visible because nothing competes forβc\\beta\_\{c\}\. The SAE case is also the cleanest substrate for a deeper question \(whether the bifurcation onset has*atom\-level*mechanistic content beyond geometric coupling\), which we take up in Section[5](https://arxiv.org/html/2605.24057#S5)\. App\.[G](https://arxiv.org/html/2605.24057#A7)reports the complementary experiment of probing the LM’s hidden states directly without an intermediate SAE, showing why an SAE substrate is essential for the full V to be visible at all\.
#### \(ii\) Fold\-back \(spectrum\)\.
When the encoder is itself learning a complex enough target,βc\(t\)\\beta\_\{c\}\(t\)may rise faster thanβ\(t\)\\beta\(t\)after the initial bifurcation, and the trajectory folds back left on thelog\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)axis while NC1 continues to fall\. The fold magnitude is controlled by data complexity\. On CIFAR\-100 \(100 fine classes\), both DINO and SimCLR exhibit*strong*fold\-back: DINO peaks atlog\(β/βc\)=\+7\.09\\log\(\\beta/\\beta\_\{c\}\)\\\!=\\\!\+7\.09around epoch 35 and descends back to\+3\.68\+3\.68by epoch 300 \(∼\\sim3 log\-unit drift\); SimCLR follows the same shape on a slightly compressed range\. On CIFAR\-10 \(10 classes\), ResNet\-18 features begin already supercritical \(log\(β/βc\)≈\+3\.9\\log\(\\beta/\\beta\_\{c\}\)\\\!\\approx\\\!\+3\.9at random init\) and the fold is*mild*\(∼\\sim0\.5 log\-unit leftward drift post\-onset\) for both DINO with native teacher\-temperatureβt\\beta\_\{t\}and SimCLR\. The two datasets occupy the same kinematic regime; data complexity controls only the magnitude ofβc\\beta\_\{c\}’s post\-onset rise\.
#### \(iii\) Delayed escape\.
Under low dissipation, the system can sit on a post\-critical metastable plateau for orders of magnitude in training steps before the macroscopic broken\-symmetry transition fires \(Remark[1](https://arxiv.org/html/2605.24057#Thmremark1)\)\. The grokking trace in Fig\.[2](https://arxiv.org/html/2605.24057#S4.F2)previews this; the three\-act decomposition with multi\-seed dissipation\-threshold control is in Sec\.[4\.2](https://arxiv.org/html/2605.24057#S4.SS2)\.
#### \(iv\) No arc \(negative control\)\.
Rotation prediction provides no clustering pressure, so the framework predicts no bifurcation\. We see only weak scatter \(r=−0\.48r\\\!=\\\!\-0\.48,n=300n\\\!=\\\!300\), with NC1 evolving largely independently oflog\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)over a comparable range\.
Table 1:Trajectory shape across the six runs of Fig\.[2](https://arxiv.org/html/2605.24057#S4.F2), with the correlationrrbetweenlog\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)andlogNC1\\log\\mathrm\{NC1\}on the descent leg \(after the V trough for SAE, after the fold\-back peak for CIFAR\-100 methods, on the full trajectory for already\-supercritical CIFAR\-10 methods and the control\)\. The framework’s prediction across panels is*shape*and the*sign*of the post\-critical relationship \(negative whenlog\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)is rising along the descent leg, positive when it folds back\)\. NC1 is computed in basis\-specific normalizations \(sparse code, ResNet features, residual stream\), and the sample sizesnndiffer across runs, so neither\|r\|\|r\|nor the implicit slope magnitude is comparable across panels in a strict sense; we reportrronly to characterize the tightness of the relationship within each panel\.
#### Sign of the post\-critical relationship is kinematic\.
The relationship betweenlog\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)andlogNC1\\log\\mathrm\{NC1\}on the descent leg is negative whenlog\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)is rising on that leg \(regime i, frozenβc\\beta\_\{c\}\) and positive when it is falling or saturating \(regime ii, fold\-back\)\. The framework does not predict a universal sign; it predicts an arc whose descent leg may be traversed in either direction inlog\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)space depending on which ofβ\\betaorβc\\beta\_\{c\}moves faster\.
### 4\.2Grokking: crossing, plateau, escape
The four SSL regimes share a feature: onceβ\\betacrossesβc\\beta\_\{c\}, the macroscopic broken\-symmetry transition follows within at most a few epochs\. Remark[1](https://arxiv.org/html/2605.24057#Thmremark1)predicts a fifth regime that decomposes into three distinct acts in training time:
- •*Act 1: Crossing\.*β\(t\)\\beta\(t\)crossesβc\(t\)\\beta\_\{c\}\(t\)and the symmetric state becomes a saddle\.
- •*Act 2: Metastable plateau\.*The system is supercritical \(β\>βc\\beta\\\!\>\\\!\\beta\_\{c\}\) but the order parameterε\\varepsilonhas not yet grown to a macroscopic value; NC1 sits on a high plateau\. The plateau length is set by the dissipation rate\.
- •*Act 3: Escape\.*Dissipation \(here, weight decay\) finally drivesε\\varepsilonoff the saddle; NC1 collapses by orders of magnitude over a comparatively short window, and downstream test accuracy jumps from chance to≈1\\approx 1\. This is the grokking transition\.
The previous reading of grokking, “the model suddenly acquires generalization at some moment,” is misleading; the model becomes*eligible*to acquire it within the first∼40\\sim 40steps and then spends thousands of steps trying to escape a saddle\. Grokking on modular arithmetic is the cleanest empirical realization of this three\-act picture\.
#### Setup\.
FollowingNandaet al\.\([2023](https://arxiv.org/html/2605.24057#bib.bib13)\), we train a 1\-layer transformer \(dmodel=128d\_\{\\mathrm\{model\}\}\\\!=\\\!128,44heads,dmlp=512d\_\{\\mathrm\{mlp\}\}\\\!=\\\!512\) ona\+bmodpa\+b\\bmod ptokens with AdamW,lr=10−3\\mathrm\{lr\}\\\!=\\\!10^\{\-3\}, full\-batch gradient descent,30%30\\%train fraction\. The probe is the same joint\-detached protocol of Sec\.[4\.1](https://arxiv.org/html/2605.24057#S4.SS1), attached to the residual stream at the=\\mathtt\{=\}position;βc\\beta\_\{c\}is computed fromCov\(z\)\\mathrm\{Cov\}\(z\)on the test set, NC1 is computed against theppoutput classes\. Three seeds per configuration\.
Figure 3:Grokking decomposes into crossing, plateau, and escape\.\(a\)Trajectories of five grokking configurations \(n=3n\\\!=\\\!3seeds each, seed\-0 trace shown\) in\(log\(β/βc\),log10NC1\)\\bigl\(\\log\(\\beta/\\beta\_\{c\}\),\\,\\log\_\{10\}\\mathrm\{NC1\}\\bigr\)\.⋆\\starmarks the grokking moment \(testacc\>0\.5\\mathrm\{test\\ acc\}\\\!\>\\\!0\.5\);×\{\\times\}marks the no\-grok endpoint of the WD==0 control\.\(b\)Canonical run \(p=97p\\\!=\\\!97, WD=1\.0=\\\!1\.0,33seeds overlaid\), with the three acts highlighted as colored bands\.*Act 1 \(blue\)*:β\\betarapidly crossesβc\\beta\_\{c\}at step37±237\\\!\\pm\\\!2\.*Act 2 \(orange\)*: a long metastable plateau in whichlog\(β/βc\)≈\+3\\log\(\\beta/\\beta\_\{c\}\)\\\!\\approx\\\!\+3butlog10NC1\\log\_\{10\}\\mathrm\{NC1\}sits at≈6\.7\\approx 6\.7; the system is post\-critical but the broken\-symmetry order parameter is still microscopic\.*Act 3 \(green\)*: at step8 900±8648\\,900\\\!\\pm\\\!864, weight decay’s dissipation finally drives the order parameter off the saddle; NC1 collapses by four orders of magnitude over a few hundred steps and test accuracy jumps to11\.
#### Act 1: the crossing is fast and universal\.
Across all1515runs \(5 configs×\\times3 seeds\),log\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)crosses zero in the first3434–6060training steps \(Table[2](https://arxiv.org/html/2605.24057#S4.T2)\)\. The crossing event is tight \(σ≈2\\sigma\\\!\\approx\\\!2steps within each configuration\) and reached in*every*seed of*every*configuration, including the no\-grok WD==0 control\. This is the Hessian\-pitchfork prediction of Sec\.[3](https://arxiv.org/html/2605.24057#S3)in raw form: the supercriticality conditionβ\>βc\\beta\\\!\>\\\!\\beta\_\{c\}is universally met within the first few steps of training\.
#### Act 2: the plateau is the post\-critical metastable saddle\.
For thousands of steps after the crossing,log\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)continues to rise toward∼\+4\\sim\\\!\+4butNC1\\mathrm\{NC1\}does not move from its high plateau\. The previous \(and incorrect\) reading would call this an undertrained representation\. The framework’s reading \(Remark[1](https://arxiv.org/html/2605.24057#Thmremark1)\) is sharper: the encoder is parked nearε=0\\varepsilon\\\!=\\\!0on a saddle that is*linearly*unstable but whose exit time is set by noise and dissipation\. The plateau length is the dissipation\-controlled escape time, not a pre\-bifurcation delay\.
#### Act 3: escape is dissipation\-controlled\.
At sufficiently long horizon \(200,000200\{,\}000steps\), all configurations with WD\>0\>\\\!0escape in3/33/3seeds, but the escape time is strongly monotonic in WD:τesc\\tau\_\{\\mathrm\{esc\}\}rises from8 9008\\,900steps at WD=1\.0=\\\!1\.0to147 167147\\,167steps at WD=0\.1=\\\!0\.1, while WD=0=\\\!0fails to escape in0/30/3seeds within50,00050\{,\}000steps\. The full sweep fits a power\-lawτesc∝WD−1\.23\\tau\_\{\\mathrm\{esc\}\}\\propto\\mathrm\{WD\}^\{\-1\.23\}\(ΔAIC=\+19\.3\\Delta\\mathrm\{AIC\}=\+19\.3over the Kramers form equation[15](https://arxiv.org/html/2605.24057#A3.E15); Table[3](https://arxiv.org/html/2605.24057#S4.T3), App\.[C](https://arxiv.org/html/2605.24057#A3)\)\. The WD==0 control reacheslog\(β/βc\)=\+6\.07\\log\(\\beta/\\beta\_\{c\}\)\\\!=\\\!\+6\.07, higher than any grokked run, yet NC1 grows to∼108\\sim\\\!10^\{8\}and test accuracy stays at∼0\.01\\sim\\\!0\.01in all three seeds; the metastable plateau is a saddle that requires dissipation to escape\. Beyond the weight\-decay axis, escape time also scales predictably with the modulusppand the training fraction \(Table[2](https://arxiv.org/html/2605.24057#S4.T2)\)\. At escape,log10NC1\\log\_\{10\}\\mathrm\{NC1\}drops from≈6\.7\\approx 6\.7to≈2\.4\\approx 2\.4over a few hundred training steps whilelog\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)moves by less than11log\-unit; the entire post\-critical descent shape of Sec\.[4\.1](https://arxiv.org/html/2605.24057#S4.SS1)is compressed into this narrow window\.
Table 2:Grokking experiments, mean±\\pmstd acrossn=3n\\\!=\\\!3seeds\. Act 1 \(theβ=βc\\beta\\\!=\\\!\\beta\_\{c\}crossing\) is universal and tight \(σ≈2\\sigma\\\!\\approx\\\!2\); Act 3 escape time spans more than an order of magnitude in WD\. Full WD sweep at the200,000200\{,\}000\-step horizon is in Table[3](https://arxiv.org/html/2605.24057#S4.T3); the WD=0=\\\!0control fails to escape within50,00050\{,\}000steps in all3/33/3seeds\.
#### Control experiment: WD\-controlled metastable plateau length\.
The grokking setup is unusual in giving us a single, clean dissipation knob \(WD\) that is otherwise neutral with respect to the Hessian\-pitchfork mechanism: changing WD does not shift Act 1 \(the crossing remains at step36−3736\\\!\-\\\!37across all WD; cf\.β=βc\\beta\\\!=\\\!\\beta\_\{c\}column of Table[2](https://arxiv.org/html/2605.24057#S4.T2)\) and does not change the post\-criticallog\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)trajectory shape, only its*rate*\. This lets us treat the WD sweep as a control experiment for Remark[1](https://arxiv.org/html/2605.24057#Thmremark1): if the metastable plateau is genuinely a post\-critical escape phenomenon, then dialing the encoder’s dissipation should monotonically change the plateau length, and at zero dissipation the system should stay on the plateau indefinitely\. We measureτesc\\tau\_\{\\mathrm\{esc\}\}at six WD levelsγ∈\{0\.1,0\.2,0\.3,0\.5,0\.7,1\.0\}\\gamma\\\!\\in\\\!\\\{0\.1,0\.2,0\.3,0\.5,0\.7,1\.0\\\}withn=3n=3seeds per level and a200,000200\{,\}000\-step horizon long enough to observe escape at everyγ\>0\\gamma\>0\(Table[3](https://arxiv.org/html/2605.24057#S4.T3)\)\.
Table 3:WD\-intervention control experiment: escape times atp=97p\\\!=\\\!97, train fraction0\.30\.3,n=3n\\\!=\\\!3seeds per WD,200,000200\{,\}000\-step horizon\. Power\-law fitτ∝γ−1\.23\\tau\\propto\\gamma^\{\-1\.23\}is within∼10%\\sim\\\!10\\%at every WD; activation\-dominated Kramers form is decisively ruled out \(ΔAIC=\+19\.3\\Delta\\mathrm\{AIC\}\\\!=\\\!\+19\.3; see App\.[C](https://arxiv.org/html/2605.24057#A3)\)\.The result is unambiguous on the qualitative claim:τesc\\tau\_\{\\mathrm\{esc\}\}is monotonically decreasing inγ\\gammaacross two decades of plateau lengths \(8\.9k→147k8\.9\\text\{k\}\\\!\\to\\\!147\\text\{k\}steps\), and theγ=0\\gamma\\\!=\\\!0control fails to escape in any of33seeds within50k50\\text\{k\}steps despite reachinglog\(β/βc\)=\+6\.07\\log\(\\beta/\\beta\_\{c\}\)\\\!=\\\!\+6\.07\. This directly supports Remark[1](https://arxiv.org/html/2605.24057#Thmremark1): the crossing event makes the symmetric state unstable, but the macroscopic broken\-symmetry transition is a dissipation\-controlled post\-critical escape, not an immediate consequence of the crossing\.
On the quantitative form, the 6\-point fit prefers a power\-lawτ∝γ−1\.23\\tau\\propto\\gamma^\{\-1\.23\}over an exponential \(activation\-dominated Kramers\) formτ∝e−κγ/D\\tau\\propto e^\{\-\\kappa\\gamma/D\}*decisively*:ΔAIC=\+19\.3\\Delta\\mathrm\{AIC\}\\\!=\\\!\+19\.3, power\-law residuals within∼10%\\sim\\\!10\\%at every WD, Kramers under\-predicts the weak\-dissipation escape times by∼1\.7×\\sim\\\!1\.7\\timesat WD=0\.1=\\\!0\.1\. We treat this as an empirical characterization specific to the modular\-arithmetic grokking setup rather than a theoretical prediction of the bifurcation framework; the regime\-selection analysis \(drift\-dominated vs\. activation\-dominated post\-critical escape\) is in App\.[C](https://arxiv.org/html/2605.24057#A3)\.
#### Predictive content of the indicator\.
At any training step during a grokking run,log\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)identifies the system’s current act: pre\-critical \(Act 1,log\(β/βc\)<0\\log\(\\beta/\\beta\_\{c\}\)\\\!<\\\!0\), post\-critical metastable \(Act 2,log\(β/βc\)\>0\\log\(\\beta/\\beta\_\{c\}\)\\\!\>\\\!0with NC1 on the plateau\), or escaping \(Act 3, NC1 descending\)\. At step 100 of the canonical run,log\(β/βc\)=\+3\\log\(\\beta/\\beta\_\{c\}\)\\\!=\\\!\+3already places the system in Act 2: the trajectory has become*eligible to grok*,∼8 400\\sim 8\\,400training steps before test accuracy provides any signal\. Combined with the dissipation strength of the encoder \(a known hyperparameter\), the indicator predicts both whether grokking will occur \(no escape at WD=0=\\\!0, Tab\.[2](https://arxiv.org/html/2605.24057#S4.T2)\) and roughly when \(τesc∝WD−1\.23\\tau\_\{\\mathrm\{esc\}\}\\propto\\mathrm\{WD\}^\{\-1\.23\}, spanning8 9008\\,900steps at WD=1\.0=\\\!1\.0to147 167147\\,167at WD=0\.1=\\\!0\.1; Tab\.[3](https://arxiv.org/html/2605.24057#S4.T3)\)\. This is prediction in the operational sense: from the indicator and the dissipation strength, both readable at step100100, the trajectory’s downstream state is forecastable orders of magnitude beforetest\_acc\\mathrm\{test\\\_acc\},train\_acc\\mathrm\{train\\\_acc\}, or the training loss show any sign of the transition\.
### 4\.3Synthesis
A single Hessian\-pitchfork prediction is governed by three binary kinematic axes \(initial sub/supercriticality, post\-onsetβ\\beta\-vs\-βc\\beta\_\{c\}kinematics, dissipation rate; Appendix[I](https://arxiv.org/html/2605.24057#A9)\)\. The axes predict which kinematic regimes are accessible to which classes of feature\-learning pipelines: standard SSL methods on rich datasets occupy the fold\-back regime, frozen\-encoder SAE training traces a full V, under\-dissipated supervised settings exhibit delayed escape, and clustering\-pressure\-free objectives produce no arc\. We observe all four regimes in the corresponding settings; the remaining nominal axis combinations are either degenerate or unreached by standard pipelines\. Method\-specific differences \(in fold\-back magnitude, plateau length, and descent shape\) are kinematic consequences of the encoder–probe race and of the encoder’s dissipation, not methodology\-specific physics\. In particular, the visual difference between CIFAR\-10 \(mild fold,∼0\.5\\sim 0\.5log\-unit leftward drift\) and CIFAR\-100 \(strong fold,∼3\\sim 3log\-units\) is a magnitude difference within the same fold\-back regime, controlled by data complexity \(number of classes→\\toricher post\-criticalCov\(z\)\\mathrm\{Cov\}\(z\)structure→\\tolargerβc\\beta\_\{c\}rise\)\. The quantitylog\(β\(t\)/βc\(t\)\)\\log\(\\beta\(t\)/\\beta\_\{c\}\(t\)\), computed from the encoder’s hidden state and a passive GMM probe alone, is the label\-free indicator of where in the arc a given run currently sits\.
The arc\-level results above confirm the theory at the*trajectory*level\. We now turn to its sharpest empirical consequence at the*per\-atom*level: the SAE feature lottery \(Section[5](https://arxiv.org/html/2605.24057#S5)\)\.
## 5The first 5% of SAE training is a feature lottery: testing the per\-atom prediction
Figure 4:SAE feature lottery emerges at the bifurcation onset and is operationally selectable by5%5\\%of training\.\(K=2048K\\\!=\\\!2048top\-KKSAE on Pythia\-160M layer 6; one canonical seed shown, 3\-seed statistics quoted\.\)\(A\)Onset:log10\(βSAE/βc\)\\log\_\{10\}\(\\beta\_\{\\mathrm\{SAE\}\}/\\beta\_\{c\}\)\(black\) crosses zero and rises;log10NC1features\\log\_\{10\}\\mathrm\{NC1\}\_\{\\mathrm\{features\}\}\(blue\) peaks at step261261, our operational definition of onset\.\(B\)Per\-atom predictivity: cross\-atom Spearmanρid\(t\)\\rho\_\{\\mathrm\{id\}\}\(t\)of POS purity atttvs convergence sits at noise floor pre\-onset, climbs sharply at the bifurcation, reaches\+0\.41±0\.04\{\+\}0\.41\\\!\\pm\\\!0\.04at step1,0001\{,\}000\(\+0\.35\{\+\}0\.35in shown seed\)\.\(C\)Top decile of atoms by step\-1,0001\{,\}000POS purity \(n=205n\\\!=\\\!205of18431843active\), placed at the angular sector defined by their dominant POS class; radius = step\-1,0001\{,\}000purity\.\(D\)Same atoms, same sectors, radius now = step\-20,00020\{,\}000purity: the outer\-ring population persists in every major POS sector; bulk inward drift reflects sparsity tightening, not specialization loss\.\(E\)Top decile reaches POS purity0\.82±0\.030\.82\\\!\\pm\\\!0\.03at convergence \(3 seeds:0\.80/0\.86/0\.820\.80/0\.86/0\.82\),12\.3±0\.4×12\.3\\\!\\pm\\\!0\.4\\timesthe uniform\-random baseline\. Reports POS\-*purity*persistence, not POS\-class identity preservation\.Sections[3](https://arxiv.org/html/2605.24057#S3)–[4](https://arxiv.org/html/2605.24057#S4)establish that the bifurcation onset is a*collective*event: the encoder’s symmetric collapsed state loses stability at a specific moment, with an unstable subspace shared by allK−1K\\\!\-\\\!1anti\-symmetric modes \(App\.[A\.2](https://arxiv.org/html/2605.24057#A1.SS2)\)\. The theory therefore makes a sharp*per\-atom*prediction: at the bifurcation, each atom must select a specific direction from this shared manifold, and that selection should be detectable as an atom\-level identity event in training time\. This section tests that prediction in the cleanest available substrate\.
Among the four regimes catalogued in Sec\.[4](https://arxiv.org/html/2605.24057#S4), the SAE on frozen Pythia\-160M layer 6 is the only setting in whichβc\\beta\_\{c\}is constant by construction; the upstream encoder does not evolve\. This makes it the cleanest substrate to disentangle the bifurcation event \(shared by both SAE families we tested\) from*per\-feature*interpretability \(architecture\-dependent\)\. It is also the natural substrate for studying the crossing event in language models at all: raw LM activations are heavily anisotropic \(Pythia\-160M layer 6 hasλmax\(Cov\(z\)\)/σ¯2≈380\\lambda\_\{\\max\}\(\\mathrm\{Cov\}\(z\)\)/\\bar\{\\sigma\}^\{2\}\\approx 380\), so a randomly initialized GMM probe attached directly to LM hidden states starts already supercritical and never visibly crossesβc\\beta\_\{c\}\. The SAE itself, by contrast, initializes with lowβSAE\\beta\_\{\\mathrm\{SAE\}\}and grows past a fresh critical point, making the full pre\-/post\-critical arc observable in the sparse\-code basis\. We train two SAE families on the same Pythia activations with dictionary sizeK=2048K=2048over input dimensiond=768d=768:\(i\) soft\-L1\(ReLU activation,L1L\_\{1\}penalty on activations,λ∈\{5×10−4,2×10−3,5×10−3,10−2\}\\lambda\\in\\\{5\{\\times\}10^\{\-4\},2\{\\times\}10^\{\-3\},5\{\\times\}10^\{\-3\},10^\{\-2\}\\\}\) and\(ii\) top\-KK\(hard architectural sparsity withtopk=64\\text\{top\}\_\{k\}=64active atoms per token, followingGaoet al\.[2025](https://arxiv.org/html/2605.24057#bib.bib17)\)\. Both are trained for20,00020\{,\}000steps with identical optimizer, batch size, and dense \(∼\\sim55\-point\) checkpoint schedules\.
We find a sharp dynamical phase transition: per\-atom predictive content goes from zero pre\-onset to substantial within the first5%5\\%of training, and atoms ranked at5%5\\%already recover the highest\-purity atoms at convergence\. FollowingFrankle and Carbin \([2019](https://arxiv.org/html/2605.24057#bib.bib19)\), we refer to this as a*feature lottery*; the drawing event is the first phase transition during training, not initialization\. The lottery is the theory’s sharpest empirical confirmation; the detailed quantitative results follow\.
#### Per\-atom identity\.
For each saved checkpoint we forward a fixed POS\-labeled WikiText\-103 eval set \(N=50,000N=50\{,\}000tokens\) through the SAE to obtain feature activationsf∈ℝN×Kf\\in\\mathbb\{R\}^\{N\\times K\}\. Atomkk’s*activation\-pattern identity*is the columnf:,k∈ℝNf\_\{:,k\}\\in\\mathbb\{R\}^\{N\}; two atoms have the same identity when theirNN\-dimensional activation vectors are cosine\-close\. This substrate is invariant to atom permutation and scaling, and bypasses the geometric near\-saturation of decoder\-column matching that arises in overcomplete dictionaries \(K\>dK\>d, App\.[J](https://arxiv.org/html/2605.24057#A10)\)\.
### 5\.1The lottery is selectable by 5% of training
The cross\-atomρid\(t\)\\rho\_\{\\mathrm\{id\}\}\(t\)trajectory plotted in Fig\.[4](https://arxiv.org/html/2605.24057#S5.F4)B is constructed by*identity matching*: atomkkat stepttis compared to atomkkat step20,00020\{,\}000, with no cross\-time atom permutation \(an alternative Hungarian\-matchedρH\\rho\_\{H\}is artifact\-inflated; see “Why not Hungarian” below\)\. Values at representative checkpoints:
Two caveats beyond what Fig\.[4](https://arxiv.org/html/2605.24057#S5.F4)shows:
1. 1\.*Pre\-onsetρid\\rho\_\{\\mathrm\{id\}\}is statistically indistinguishable from0*fort<200t<200, confirming atoms have no detectable identity before the bifurcation\.
2. 2\.*The lottery is operationally useful at5%5\\%, but identity refinement continues\.*The remaining95%95\\%of training liftsρid\\rho\_\{\\mathrm\{id\}\}from0\.410\.41to0\.880\.88\. We therefore do not claim the lottery is “complete” at5%5\\%in the strong sense ofFrankle and Carbin \([2019](https://arxiv.org/html/2605.24057#bib.bib19)\)\(where identification at initialization is followed by isolated retraining to full accuracy\)\. Ours is the weaker statement that early ranking provides a useful predictor of converged identity\.
This is the SAE\-level analogue of the lottery\-ticket framing ofFrankle and Carbin \([2019](https://arxiv.org/html/2605.24057#bib.bib19)\): where they show that winning subnetworks are selectable at initialization in supervised classifiers, we find that winning SAE atoms are selectable at the first phase transition of unsupervised dictionary training\. The bifurcation onset is the lottery’s drawing event\.
#### Ranking lift \(three seeds\)\.
At convergence, the decile of atoms ranked by their step\-1,0001\{,\}000POS purity has mean POS purity0\.82±0\.030\.82\\pm 0\.03\(mean±\\pmstd acrossn=3n\\\!=\\\!3seeds; seed 0 / 1 / 2 give 0\.80 / 0\.86 / 0\.82\) versus0\.47±0\.030\.47\\pm 0\.03for the bottom decile\. The relevant null for the lottery claim — “does step\-1,0001\{,\}000ranking carry predictive content for convergence identity?” — is uniform\-random selection, under which top\-decile mean equals the corpus POS\-class prior1/15≈0\.0671/15\\approx 0\.067\(the 15 POS classes in the WikiText eval set\)\. Early\-screened top\-decile atoms achieve12\.3±0\.4×\\mathbf\{12\.3\\pm 0\.4\\times\}this uniform\-random baseline\. The identity\-matched Spearman across the three seeds isρid=\+0\.41±0\.04\\rho\_\{\\mathrm\{id\}\}=\+0\.41\\pm 0\.04, allp<10−80p<10^\{\-80\}\. The corresponding single\-seed numbers for the soft\-L1 SAE are0\.640\.64\(top decile\) versus0\.350\.35\(bottom decile\),10×10\\timesuniform\-random baseline\. Both architectures support the same conclusion:early winner screening at5%5\\%of training is feasible and identifies a substantial fraction of the highest\-purity atoms, though full training is still required to obtain each atom’s highest\-quality final activations\.
#### Architecture invariance\.
The sameρid\\rho\_\{\\mathrm\{id\}\}trajectory and5%5\\%cutoff are recovered in soft\-L1 SAEs \(ReLU activation,L1L\_\{1\}penaltyλ∈\{5×10−4,2×10−3,5×10−3,10−2\}\\lambda\\in\\\{5\\\!\\times\\\!10^\{\-4\},2\\\!\\times\\\!10^\{\-3\},5\\\!\\times\\\!10^\{\-3\},10^\{\-2\}\\\},L0≈1000L\_\{0\}\\approx 1000ofK=2048K\\\!=\\\!2048\) despite their different sparsity regime: the post\-critical descent ofρid\\rho\_\{\\mathrm\{id\}\}is qualitatively identical \(App\.[J](https://arxiv.org/html/2605.24057#A10)\), demonstrating that the lottery is a property of the bifurcation, not of the architectural top\-KKmask\.
#### Whyρid\\rho\_\{\\mathrm\{id\}\}and notρH\\rho\_\{H\}\(Hungarian\)\.
We initially considered Hungarian\-matched correlationρH\\rho\_\{H\}; match atoms across time by maximizing cosine of activation patterns, then correlate matched pairs’ POS purities\. This metric is systematically inflated: at random initialization \(step4646, where atoms have no identity by construction\),ρH≈0\.30\\rho\_\{H\}\\approx 0\.30rather than0\(App\.[L](https://arxiv.org/html/2605.24057#A12), panels D–E\)\. The inflation arises because Hungarian matching preferentially pairs atoms with similar activation profiles, and POS purity is itself computed from the activation profile, so high\-purity atoms find each other across time even when no underlying identity is preserved\. Subtracting the0\.300\.30floor recovers theρid\\rho\_\{\\mathrm\{id\}\}trajectory\.ρid\\rho\_\{\\mathrm\{id\}\}, which is invariant to this confound, is the estimator we report throughout\.
#### Specificity\.
Of three per\-atom early metrics tested, only POS purity carries the predictive signal:
- •*Early POS purity*→\\tofinal POS purity:ρid=\+0\.41\\rho\_\{\\mathrm\{id\}\}=\+0\.41at5%5\\%\.Predictive\.
- •*Early activation concentration*→\\tofinal POS purity:ρ≈0\\rho\\approx 0in the top\-KKSAE,ρ≈−0\.28\\rho\\approx\-0\.28in the soft\-L1 SAE\.Not predictive\.
- •*Early identity\-lock cosine*→\\tofinal POS purity:ρ=\+0\.20\\rho=\+0\.20\(top\-KK\)\. Weakly predictive\.
This specificity rules out the trivial reading “any atom\-level property predicts its convergence value\.” What is drawn at the bifurcation is specifically*linguistic identity*, not generic firing strength or generic stability\.
### 5\.2Mechanism: identity lock during the lottery window
Why is the lottery permanent past5%5\\%of training? Atom identities themselves lock during the post\-critical window\. The Hungarian\-matched cosine between per\-atom activations at stepttand at the converged SAE remains flat \(at random baseline0\.130\.13for top\-KK,0\.680\.68for soft\-L1\) fort<200t\\\!<\\\!200, begins a sharp rise at the NC1 peak \(step261261for top\-KK,383383for soft\-L1\), and saturates at1\.01\.0by step∼15,000\\sim 15\{,\}000\. Both SAE families share this timing; only their random\-init baselines differ \(soft\-L1’s ReLU rows already project onto the data subspace\)\. The onset is therefore the*trigger*of a per\-atom assignment process: each atom commits to a specific activation pattern during the post\-critical window, and subsequent training refines but does not reshuffle that commitment \(App\.[L](https://arxiv.org/html/2605.24057#A12), Fig\.[17](https://arxiv.org/html/2605.24057#A12.F17)\)\.
### 5\.3Architecture\-dependent ceiling on feature interpretability
The bifurcation event is shared by both SAE families we tested, but the*post\-onset ceiling*on per\-feature interpretability is not\. Under architectural top\-KKsparsity, the median active feature reaches POS purity0\.560\.56\(8×8\\timesrandom\); the soft\-L1 family plateaus at0\.330\.33, sinceL0L\_\{0\}saturates near1,0001\{,\}000ofK=2048K\\\!=\\\!2048irrespective ofλ∈\[5×10−4,10−2\]\\lambda\\in\[5\{\\times\}10^\{\-4\},10^\{\-2\}\]\. The bifurcation locks identities identically in both families; only top\-KKpushes those identities into linguistically selective directions\. Who wins the lottery is determined at the bifurcation; how linguistically meaningful the winning identity is depends on the post\-onset sparsity regime \(App\.[L](https://arxiv.org/html/2605.24057#A12)\)\.
### 5\.4K\-sweep: universality of the lottery, monotonicity of interpretability and slope
Figure 5:K\-sweep acrossK∈\{256,512,1024,2048,4096,8192\}K\\in\\\{256,512,1024,2048,4096,8192\\\}at fixed 3% top\-KKsparsity\.For eachKK, we report finalL0L\_\{0\}, final reconstruction MSE, the lottery Spearmanρid\\rho\_\{\\mathrm\{id\}\}, and the mean POS purity of the top\-decile atoms ranked by step\-1,0001\{,\}000POS purity\. Universality acrossKK: the lotteryρid∈\[\+0\.26,\+0\.41\]\\rho\_\{\\mathrm\{id\}\}\\in\[\+0\.26,\+0\.41\]isKK\-stable and consistently positive: the bifurcation onset is aKK\-universal event\. Monotonicity inKK: POS purity*decreases*withKK\(0\.7250\.725atK=256K\\\!=\\\!256,0\.3700\.370atK=8192K\\\!=\\\!8192\), while reconstruction MSE decreases the other way\.We sweepK∈\{256,512,1024,2048,4096,8192\}K\\in\\\{256,512,1024,2048,4096,8192\\\}at fixed top\-KKsparsity ratio of3%3\\%\(sotopk\\text\{top\}\_\{k\}scales proportionally as\{8,16,32,64,128,256\}\\\{8,16,32,64,128,256\\\}\) with all other hyperparameters held constant\. TheKK\-stability ofρid\\rho\_\{\\mathrm\{id\}\}visible in Fig\.[5](https://arxiv.org/html/2605.24057#S5.F5)is statistically tight \(p<10−3p<10^\{\-3\}at everyKK\), justifying treating the bifurcation onset as aKK\-universal event\.
#### POS purity is monotonically decreasing inKK, butKK\-confounded\.
The median atom’s POS purity at convergence falls steadily from0\.7250\.725atK=256K\\\!=\\\!256to0\.3700\.370atK=8192K\\\!=\\\!8192\(Table[4](https://arxiv.org/html/2605.24057#S5.T4)\)\. The mechanism is that a smaller dictionary forces each atom to absorb a coarser token\-cluster partition that more easily aligns with the1515\-way POS partition; a larger dictionary spreads across finer sub\-categories that POS purity is too coarse to resolve\. We emphasize that this trend is therefore*partly a structural property*of POS purity as a metric, not a direct statement that smallKKproduces more mechanistically interpretable features: aKK\-unbiased interpretability metric \(causal mediation, LLM\-as\-judge\) is required before a Pareto\-based recommendation can be made\.
Table 4:K\-sweep at fixed top\-KKsparsity ratio3%3\\%, all other hyperparameters constant\. The lotteryρid\\rho\_\{\\mathrm\{id\}\}isKK\-stable; POS purity decreases monotonically withKKwhile reconstruction MSE decreases the other way\.
#### Pareto frontier, with aKK\-bias caveat\.
K=256K\\\!=\\\!256scores higher on POS purity; largeKKscores better on reconstruction\. The framework adds a second axis \(atom\-level POS purity, plus the lotteryρid\\rho\_\{\\mathrm\{id\}\}\) to the standard SAEL0L\_\{0\}\-vs\-recon Pareto\. However, because POS has only1515classes, the POS\-purity\-vs\-KKtrend is partly a structural artifact: smallerKKforces each atom to absorb a coarser token\-cluster partition, which more easily aligns with a1515\-way categorical metric\. We therefore do*not*recommend smallKKon the basis of POS purity alone; aKK\-unbiased interpretability metric \(such as causal mediation\(cf\. Nandaet al\.,[2023](https://arxiv.org/html/2605.24057#bib.bib13)\)or LLM\-as\-judge scoring\) is required before a substantive recommendation can be made\. What*is*robust to this caveat is theKK\-universal lottery effect: at everyK∈\{256,…,8192\}K\\\!\\in\\\!\\\{256,\\dots,8192\\\}, atom\-level identity is predictable from5%5\\%training withρid∈\[\+0\.26,\+0\.41\]\\rho\_\{\\mathrm\{id\}\}\\in\[\+0\.26,\+0\.41\]\.
### 5\.5Connection to the bifurcation theory
The pitchfork bifurcation and the feature lottery are two scales of the same event\. Atβ=βc\\beta\\\!=\\\!\\beta\_\{c\}the unstable subspace is shared by allK−1K\\\!\-\\\!1anti\-symmetric modes \(App\.[A\.2](https://arxiv.org/html/2605.24057#A1.SS2)\); each atom selects a specific direction from this manifold via random initialization noise and the cubic terms in equation[5](https://arxiv.org/html/2605.24057#S3.E5)\. The collective bifurcation enables the per\-atom selection; the per\-atom selection is what makes post\-onset identities permanent and partially predictable\. The following proposition formalizes the per\-atom directional preservation that the empiricalρid\>0\\rho\_\{\\mathrm\{id\}\}\>0\(Sec\.[5\.1](https://arxiv.org/html/2605.24057#S5.SS1)\) realizes\.
###### Proposition 2\(Per\-atom directional persistence\)\.
Consider the coupled order\-parameter dynamics on the\(K−1\)\(K\\\!\-\\\!1\)\-fold\-degenerate unstable subspace just past the crossing \(μ:=β−βc\>0\\mu:=\\beta\-\\beta\_\{c\}\>0, small\),
𝜺˙k=μ𝜺k−α‖𝜺k‖2𝜺k−γ∑j≠k\(𝜺j⊤𝜺k\)𝜺j\+𝜼k\(t\),k=1,…,K,\\dot\{\\bm\{\\varepsilon\}\}\_\{k\}\\;=\\;\\mu\\,\\bm\{\\varepsilon\}\_\{k\}\\;\-\\;\\alpha\\,\\\|\\bm\{\\varepsilon\}\_\{k\}\\\|^\{2\}\\bm\{\\varepsilon\}\_\{k\}\\;\-\\;\\gamma\\\!\\\!\\sum\_\{j\\neq k\}\\\!\\bigl\(\\bm\{\\varepsilon\}\_\{j\}^\{\\top\}\\bm\{\\varepsilon\}\_\{k\}\\bigr\)\\bm\{\\varepsilon\}\_\{j\}\\;\+\\;\\bm\{\\eta\}\_\{k\}\(t\),\\qquad k=1,\\dots,K,\(8\)with i\.i\.d\. Langevin noises𝛈k\\bm\{\\eta\}\_\{k\}at intensityDDand cubic inter\-mode coupling strengthγ≥0\\gamma\\geq 0\. Assume each initial perturbation satisfies‖𝛆k\(0\)‖\>σ∗⋅D/μ\\\|\\bm\{\\varepsilon\}\_\{k\}\(0\)\\\|\>\\sigma\_\{\*\}\\\!\\cdot\\\!\\sqrt\{D/\\mu\}for someσ∗\>1\\sigma\_\{\*\}\>1\(initial perturbation above the noise floor\)\. Then in the weak\-coupling regimeγ≪α\\gamma\\ll\\alpha, for any*finite*timeTTwithin the post\-saturation, pre\-randomization windowτr≲T≪Trand:=r⋆2/\(2\(d−1\)D\)\\tau\_\{r\}\\lesssim T\\ll T\_\{\\mathrm\{rand\}\}:=r^\{\\star 2\}/\(2\(d\{\-\}1\)D\)\(whereτr\\tau\_\{r\}is the deterministic radial saturation timescale andTrandT\_\{\\mathrm\{rand\}\}is the spherical randomization timescale on the saturated attractor\), the per\-atom directional persistence is strictly positive:
𝔼\[dk\(0\)⊤dk\(T\)\]\>0,dk\(t\):=𝜺k\(t\)/‖𝜺k\(t\)‖∈Sd−1,\\mathbb\{E\}\\bigl\[\\,d\_\{k\}\(0\)^\{\\top\}d\_\{k\}\(T\)\\,\\bigr\]\\;\>\\;0,\\qquad d\_\{k\}\(t\):=\\bm\{\\varepsilon\}\_\{k\}\(t\)/\\\|\\bm\{\\varepsilon\}\_\{k\}\(t\)\\\|\\in S^\{d\-1\},\(9\)with magnitude controlled byσ∗μ/\(αD\)1/2\\sigma\_\{\*\}\\,\\mu/\(\\alpha D\)^\{1/2\}and expectation taken over the Langevin noise realizations\. For training runs of fixed horizon,TTis the training time; in the strictT→∞T\\to\\inftylimit, angular Brownian motion on the saturated sphere eventually randomizesdk\(T\)d\_\{k\}\(T\)and the persistence claim breaks down \(outside the scope of this proposition\)\.
###### Proof sketch \(full proof in App\.[D](https://arxiv.org/html/2605.24057#A4)\)\.\.
Atγ=0\\gamma=0, modes decouple\. Each𝜺k\\bm\{\\varepsilon\}\_\{k\}obeys an independent pitchfork SDE whose deterministic flow is*radial*: the linear growthμ𝜺k\\mu\\bm\{\\varepsilon\}\_\{k\}preserves direction, and the cubic self\-saturation−α‖𝜺k‖2𝜺k\-\\alpha\\\|\\bm\{\\varepsilon\}\_\{k\}\\\|^\{2\}\\bm\{\\varepsilon\}\_\{k\}also preserves direction\. Noise rotates direction with effective diffusion coefficient\(d−1\)D/‖𝜺k‖2\(d\{\-\}1\)D/\\\|\\bm\{\\varepsilon\}\_\{k\}\\\|^\{2\}, which decays rapidly as‖𝜺k‖\\\|\\bm\{\\varepsilon\}\_\{k\}\\\|grows fromσ∗D/μ\\sigma\_\{\*\}\\sqrt\{D/\\mu\}towardμ/α\\sqrt\{\\mu/\\alpha\}\. Over the finite window\[0,T\]\[0,T\], the accumulated angular variance isΘ2\(T\)≈\(d−1\)/σ∗2\+2\(d−1\)D\(T−τr\)/r⋆2\\Theta^\{2\}\(T\)\\approx\(d\{\-\}1\)/\\sigma\_\{\*\}^\{2\}\+2\(d\{\-\}1\)D\(T\-\\tau\_\{r\}\)/r^\{\\star 2\}, which isO\(1/σ∗2\)O\(1/\\sigma\_\{\*\}^\{2\}\)as long as the saturation\-phase contribution\(T−τr\)/Trand\(T\-\\tau\_\{r\}\)/T\_\{\\mathrm\{rand\}\}is sub\-leading\. For smallΘ\(T\)\\Theta\(T\),𝔼\[dk\(0\)⊤dk\(T\)\]≈1−Θ2\(T\)/2\>0\\mathbb\{E\}\[d\_\{k\}\(0\)^\{\\top\}d\_\{k\}\(T\)\]\\approx 1\-\\Theta^\{2\}\(T\)/2\>0\. Adding weak inter\-mode couplingγ\>0\\gamma\>0contributes anO\(γ/α\)O\(\\gamma/\\alpha\)perturbation that preserves the sign of𝔼\[dk\(0\)⊤dk\(T\)\]\\mathbb\{E\}\[d\_\{k\}\(0\)^\{\\top\}d\_\{k\}\(T\)\]at leading order, since the coupling vanishes when modes are mutually orthogonal \(and modes self\-orthogonalize under the cubic saturation\)\.□\\square∎
#### Empirical proxies for per\-atom persistence\.
Equation equation[9](https://arxiv.org/html/2605.24057#S5.E9)is per\-atom: at finiteTTwithin the persistence window, each atom’s direction positively correlates with its direction at the bifurcation onset\. We do not have direct access todk\(T\)⊤dk\(0\)d\_\{k\}\(T\)^\{\\top\}d\_\{k\}\(0\)in real SAEs \(the unstable manifold is not explicitly parameterized\); we use the following two empirically accessible proxies, both of which inherit positivity from equation[9](https://arxiv.org/html/2605.24057#S5.E9)under the same regime assumption\.
*Scalar\-projection proxy \(toy verification, App\.[D\.3](https://arxiv.org/html/2605.24057#A4.SS3), Fig\.[8](https://arxiv.org/html/2605.24057#A4.F8)\)\.*Fix a reference directionr∈Sd−1r\\in S^\{d\-1\}and consider the per\-atom scalar autocorrelationAk\(T\):=⟨𝜺k\(0\),r⟩⋅⟨𝜺k\(T\),r⟩/\(‖𝜺k\(0\)‖‖𝜺k\(T\)‖\)A\_\{k\}\(T\):=\\langle\\bm\{\\varepsilon\}\_\{k\}\(0\),r\\rangle\\cdot\\langle\\bm\{\\varepsilon\}\_\{k\}\(T\),r\\rangle/\(\\\|\\bm\{\\varepsilon\}\_\{k\}\(0\)\\\|\\\|\\bm\{\\varepsilon\}\_\{k\}\(T\)\\\|\)\. By equation[9](https://arxiv.org/html/2605.24057#S5.E9),𝔼\[Ak\(T\)\]\>0\\mathbb\{E\}\[A\_\{k\}\(T\)\]\>0for eachkk\. AcrossKKi\.i\.d\. atoms, the Spearman correlation between the projection pair\(⟨𝜺k\(0\),r⟩,⟨𝜺k\(T\),r⟩\)\(\\langle\\bm\{\\varepsilon\}\_\{k\}\(0\),r\\rangle,\\langle\\bm\{\\varepsilon\}\_\{k\}\(T\),r\\rangle\)converges to a positive population Spearman asK→∞K\\\!\\to\\\!\\infty; direct SDE simulation in App\.[D\.3](https://arxiv.org/html/2605.24057#A4.SS3)confirmsρ=0\.948±0\.006\\rho=0\.948\\pm 0\.006across55seeds\.
*POS\-purity proxy \(SAE empirics, Sec\.[5\.1](https://arxiv.org/html/2605.24057#S5.SS1)\)\.*Per\-atom POS purity is the magnitude of an atom’s projection onto a linguistic POS subspace ofℝN\\mathbb\{R\}^\{N\}\. Under equation[9](https://arxiv.org/html/2605.24057#S5.E9), this projection is per\-atom\-autocorrelated across training time: the early\-time and late\-time POS purity for the same atom are positively correlated\. The identity\-matched Spearmanρid=0\.41±0\.03\\rho\_\{\\mathrm\{id\}\}=0\.41\\pm 0\.03across atoms \(Sec\.[5\.1](https://arxiv.org/html/2605.24057#S5.SS1)\) is the population Spearman of this per\-atom autocorrelation; the positive sign is equation[9](https://arxiv.org/html/2605.24057#S5.E9)’s prediction, the specific magnitude \(\+0\.41\+0\.41\) depends on the linguistic geometry of POS classes in Pythia’s residual stream and is not predicted by the theory\.
### 5\.6Scope and limitations
Figure 6:From\-scratch DINO with four perturbed modes \(n=1n=1seed per mode\)\.Healthy \(black\) versus four perturbations over5050epochs of CIFAR\-10 training\. The two perturbations relevant for the diagnostic claim are*no\_sharpening*\(orange\) and*tiny\_batch*\(green\): both separate from healthy inlog\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)at epoch 2 \(vertical dotted line\) while cluster accuracy diverges only by epoch 10, an∼\\sim8\-epoch lead time\. The two outcomes go in opposite directions; no\_sharpening degrades \(∼25%\\sim 25\\%cluster accuracy at epoch 50 vs∼55%\\sim 55\\%healthy\); tiny\_batch improves \(∼70%\\sim 70\\%\)\. The catastrophic modes \(*no\_centering*,*no\_ema*\) collapse from initialization and provide no epoch\-resolution lead\-time signal\.The lottery claim is bounded along three axes\. \(i\)*POS purity is one interpretability dimension*: atoms detecting finer structure \(named entities, sub\-POS, phrase patterns\) may be mechanistically interpretable without scoring high on POS purity; our claim concerns the POS\-selective subset only\. \(ii\)*Secondary bifurcations\.*The hierarchical structure \(Sec\.[3](https://arxiv.org/html/2605.24057#S3)\) predicts that finer features should themselves exhibit phase transitions atβc\(2\)=1/λmax\(Σwithin\)\\beta\_\{c\}^\{\(2\)\}=1/\\lambda\_\{\\max\}\(\\Sigma\_\{\\mathrm\{within\}\}\)later in training; verification at the secondary scale is left to future work\. \(iii\)*Tagger noise\.*spaCy\(Honnibalet al\.,[2020](https://arxiv.org/html/2605.24057#bib.bib22)\)reports∼\\sim97%97\\%POS accuracy on standard benchmarks, with errors concentrated on ambiguous cases \(gerunds, deverbal nouns\); reportedρid\\rho\_\{\\mathrm\{id\}\}values are therefore conservative lower bounds on true atom\-level POS selectivity\.
#### Frequency\-confound check\.
A natural concern is that the lottery effect is inflated by token frequency: atoms that lock onto a single high\-frequency token \(e\.g\. “the”\) would achieve high POS purity trivially, since high\-frequency function words have stable POS tags\. To rule this out, we stratify atoms by the mean log\-frequency of their top\-100 activating tokens into five quintiles and re\-computeρid\\rho\_\{\\mathrm\{id\}\}within each quintile \(3 seeds,K=2048K\\\!=\\\!2048top\-KKSAE\)\. The result \(Tab\.[5](https://arxiv.org/html/2605.24057#S5.T5), Fig\.[18](https://arxiv.org/html/2605.24057#A12.F18)\): all five quintiles showρid∈\[\+0\.37,\+0\.51\]\\rho\_\{\\mathrm\{id\}\}\\in\[\+0\.37,\+0\.51\], well above zero \(p<10−3p<10^\{\-3\}each\), with no monotonic increase from rare to frequent tokens\. The frequency confound is not load\-bearing\.
Table 5:ρid\\rho\_\{\\mathrm\{id\}\}stratified by atom top\-100\-token mean log\-frequency, 3 seeds,K=2048K\\\!=\\\!2048\. Unstratifiedρid=\+0\.416±0\.010\\rho\_\{\\mathrm\{id\}\}=\+0\.416\\pm 0\.010\. The quintile range\[\+0\.37,\+0\.51\]\[\+0\.37,\+0\.51\]is small compared to the overall effect over the random baseline \(ρ=0\\rho\\\!=\\\!0\), and Q4 \(mid\-frequency content words\) rather than Q5 \(most frequent function words\) has the highestρid\\rho\_\{\\mathrm\{id\}\}, opposite to what a single\-token\-lock artifact would predict\.
## 6Label\-free training diagnostic in practice
The phase\-identification ability documented in Sec\.[4\.2](https://arxiv.org/html/2605.24057#S4.SS2)\(thatβ/βc\\beta/\\beta\_\{c\}reads the encoder’s current act from its hidden state alone, well before downstream metrics respond\) generalizes beyond grokking\. We now test the same property in a more typical training\-health setting: detecting the onset of representation degradation in a self\-supervised encoder\. The bifurcation framework predicts thatβ/βc\\beta/\\beta\_\{c\}tracks representation\-level state\. We now ask whether this is operationally useful as a label\-free training health indicator: doesβ/βc\\beta/\\beta\_\{c\}respond earlier than downstream metrics when training drifts away from a healthy trajectory? We test this in two complementary setups: DINO trained from scratch with collapse modes injected at initialization \(Sec\.[6\.1](https://arxiv.org/html/2605.24057#S6.SS1)\), and DINO with interventions applied to a healthy mid\-training checkpoint \(Sec\.[6\.2](https://arxiv.org/html/2605.24057#S6.SS2)\)\.
### 6\.1From\-scratch collapse modes
We train ResNet\-18 \+ DINO on CIFAR\-10 for 50 epochs in five configurations \(Fig\.[6](https://arxiv.org/html/2605.24057#S5.F6)\): healthy;*no\_centering*\(no center\-buffer update on teacher logits\);*no\_sharpening*\(teacher temperature pinned to student temperature\);*no\_ema*\(teacher copies student each step\);*tiny\_batch*\(batch size 32 instead of 256\)\.
The two non\-catastrophic perturbations \(*no\_sharpening*,*tiny\_batch*\) probe the diagnostic in opposite directions:
#### Negative case \(*no\_sharpening*\)\.
log\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)crosses the0\.50\.5\-log\-unit deviation threshold from healthy at epoch 2 while cluster accuracy crosses the5%5\\%\-deviation threshold only at epoch 10:an∼\\sim8\-epoch lead time on a 50\-epoch run, with the downstream outcome being degradation \(25%25\\%vs55%55\\%at epoch 50\)\.
#### Positive case \(*tiny\_batch*\)\.
A non\-catastrophic perturbation \(batch3232vs256256\) wherelog\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)also separates from healthy by∼0\.5\\sim 0\.5log\-units at epoch 2, while the training\-loss difference at that point is within batch noise\. By epoch 50, cluster accuracy has separated by1515percentage points in the*opposite*direction \(70%70\\%tiny\_batch vs55%55\\%healthy\)\. The earlyβ/βc\\beta/\\beta\_\{c\}separation thus predicts the downstream divergence with∼\\sim48\-epoch lead time; and the divergence is positive, not negative\. Together with no\_sharpening, this shows thatβ/βc\\beta/\\beta\_\{c\}reads the representation’s*geometric trajectory state*, not a binary “health” signal: distinctβ/βc\\beta/\\beta\_\{c\}trajectories predict distinct downstream outcomes, regardless of sign\.
#### Threshold caveat\.
The0\.50\.5\-log\-unit and5%5\\%thresholds are heuristic, chosen post\-hoc\. A full ROC\-style analysis \(detection time vs false\-positive rate\) requires multi\-seed evaluation; current experiments usen=1n=1seed per mode\. We report lead time at fixed heuristic thresholds for visual interpretability and leave the multi\-seed ROC analysis to future work\.
### 6\.2Mid\-training interventions
We then test whether the same signal works on a healthy mid\-training checkpoint subjected to a perturbation\. Phase A: 2 epochs of healthy DINO on CIFAR\-100 \(Phase A is shared across modes via a checkpoint\)\. Phase B: 10 epochs of intervention\. Dense probes at steps\{1,5,10,25,50,100,200,500\}\\\{1,5,10,25,50,100,200,500\\\}after the Phase B transition resolve per\-batch dynamics\.
Figure 7:Mid\-training intervention test on CIFAR\-100, 5 modes\.Top: epoch\-level traces, red dashed line at Phase B start\. Bottom: zoom on Phase B onset \(symmetric\-logxx\-axis\)\. Two paper\-relevant signatures: \(i\)*no\_ema*\(purple\):log\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)moves by 1\.9 log\-units within one batch, oscillating up to\+10\.7\+10\.7within 50 batches, while training loss oscillates uninterpretably; \(ii\)*tiny\_batch*\(green\):log\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)separates from healthy by 0\.5 log\-units within 25 batches while training loss differs by0\.060\.06\(within noise\)\.#### Two paper\-relevant signatures\.
- •*Per\-batch sensitivity \(no\_ema\)\.*Within one batch of removing the teacher EMA,log\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)jumps by\+1\.90\+1\.90log\-units \(from\+6\.22\+6\.22to\+8\.12\+8\.12\)\. The trajectory then oscillates with±2\\pm 2\-log\-unit amplitude\. Training loss oscillates as well over the same window but in an uninterpretable way, swinging between6\.936\.93and0\.160\.16and not consistently tracking representation\-level state\. Theβ/βc\\beta/\\beta\_\{c\}signal is the cleaner per\-batch indicator\.
- •*Sub\-catastrophic lead time \(tiny\_batch\)\.*At 25 batches into Phase B,log\(β/βc\)=\+6\.98\\log\(\\beta/\\beta\_\{c\}\)\\\!=\\\!\+6\.98for tiny\_batch versus\+6\.48\+6\.48for healthy, a0\.500\.50log\-unit separation\. Training loss at the same step is6\.8696\.869versus6\.8096\.809: a0\.0600\.060difference, within batch\-to\-batch noise\. Theβ/βc\\beta/\\beta\_\{c\}signal distinguishes the two modes long before loss can\.
#### Headline\.
Across these two single\-seed experimental setups \(Sec\.[6\.1](https://arxiv.org/html/2605.24057#S6.SS1)\+ Sec\.[6\.2](https://arxiv.org/html/2605.24057#S6.SS2)\),β/βc\\beta/\\beta\_\{c\}separates from the healthy trajectory∼8\\sim\\\!8epochs before cluster accuracy does in the gradual case \(no\_sharpening, degradation; tiny\_batch, improvement\) and within a single batch in the catastrophic intervention case \(no\_ema\)\. Both lead\-time numbers are at heuristic post\-hoc thresholds and onn=1n\\\!=\\\!1seed per condition; multi\-seed ROC analysis is needed before quoting these as detection times in deployment\. The lead time is not directional: in the tiny\_batch case, the earlyβ/βc\\beta/\\beta\_\{c\}separation*positively predicts*a downstream cluster\-accuracy gain\. The signal is derived purely from the encoder’s hidden states and a passive probe; no labels enter at any point\.
## 7Conclusion
Concept emergence in feature\-learning networks admits a label\-free dynamical indicator\. From a single Hessian analysis of a passive GMM probe on a given encoder representation, we derive a critical precisionβc=1/λmax\(Cov\(z\)\)\\beta\_\{c\}\\\!=\\\!1/\\lambda\_\{\\max\}\(\\mathrm\{Cov\}\(z\)\)above which prototypes pitchfork; we prove a finite\-time crossing theorem for co\-evolving encoders that eventually spread the latent representation sufficiently \(Proposition[1](https://arxiv.org/html/2605.24057#Thmproposition1)\), and identify a post\-critical metastable regime under finite dissipation \(Remark[1](https://arxiv.org/html/2605.24057#Thmremark1)\)\. The encoder–probe race inβ\(t\)/βc\(t\)\\beta\(t\)/\\beta\_\{c\}\(t\)traces one of four kinematic regimes \(full V, fold\-back, delayed escape, no arc\) governed by three binary axes, which we verify across SAEs, SSL on CIFAR\-10/100, grokking with multi\-seed dissipation\-threshold control, and a rotation\-prediction negative control\. The sharpest empirical confirmation is per\-atom: in SAE training the unstable subspace is shared across atoms at the crossing, and we observe the predicted lottery; the top decile of atoms ranked at5%5\\%of training reaches12×12\\timesbaseline POS purity at convergence in two distinct SAE architectures, acrossK∈\[256,8192\]K\\\!\\in\\\!\[256,8192\], with three\-seed reproducibility \(ρid=\+0\.41±0\.04\\rho\_\{\\mathrm\{id\}\}\\\!=\\\!\+0\.41\\\!\\pm\\\!0\.04, allp<10−80p\\\!<\\\!10^\{\-80\}\)\. The sameβ/βc\\beta/\\beta\_\{c\}quantity acts as a practical training\-health indicator with multi\-epoch lead time over downstream metrics in gradual collapse modes and per\-batch sensitivity in catastrophic interventions, in our single\-seed proof\-of\-concept experiments; multi\-seed ROC characterization remains future work\. Across these settings the bifurcation onset is the moment at which representation structure first becomes available, at which atom\-level identity acquires predictive content for convergence interpretability, and at which any subsequent collapse can already be detected\.
## References
- N\. Berglund \(2013\)Kramers’ law: validity, derivations and generalisations\.Markov Processes and Related Fields19,pp\. 459–490\.Cited by:[Appendix C](https://arxiv.org/html/2605.24057#A3.SS0.SSS0.Px3.p1.4)\.
- T\. Bricken, A\. Templeton, J\. Batson, B\. Chen, A\. Jermyn, T\. Conerly, N\. Turner, C\. Anil, C\. Denison, A\. Askell,et al\.\(2023\)Towards monosemanticity: decomposing language models with dictionary learning\.Transformer Circuits Thread\.Cited by:[§2](https://arxiv.org/html/2605.24057#S2.SS0.SSS0.Px1.p1.1)\.
- M\. Caron, H\. Touvron, I\. Misra, H\. Jégou, J\. Mairal, P\. Bojanowski, and A\. Joulin \(2021\)Emerging properties in self\-supervised vision transformers\.InIEEE/CVF International Conference on Computer Vision \(ICCV\),Cited by:[§2](https://arxiv.org/html/2605.24057#S2.SS0.SSS0.Px4.p1.1)\.
- J\. Frankle and M\. Carbin \(2019\)The lottery ticket hypothesis: finding sparse, trainable neural networks\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[item 2](https://arxiv.org/html/2605.24057#S1.I1.i2.p1.14),[§1](https://arxiv.org/html/2605.24057#S1.SS0.SSS0.Px2.p1.5),[§2](https://arxiv.org/html/2605.24057#S2.SS0.SSS0.Px7.p1.3),[item 2](https://arxiv.org/html/2605.24057#S5.I1.i2.p1.6),[§5\.1](https://arxiv.org/html/2605.24057#S5.SS1.p3.1),[§5](https://arxiv.org/html/2605.24057#S5.p3.2)\.
- L\. Gao, T\. D\. la Tour, H\. Tillman, G\. Goh, R\. Troll, A\. Radford, I\. Sutskever, J\. Leike, and J\. Wu \(2025\)Scaling and evaluating sparse autoencoders\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2605.24057#S2.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2605.24057#S5.p2.12)\.
- M\. Honnibal, I\. Montani, S\. Van Landeghem, and A\. Boyd \(2020\)spaCy: industrial\-strength natural language processing in Python\.External Links:[Document](https://dx.doi.org/10.5281/zenodo.1212303)Cited by:[§5\.6](https://arxiv.org/html/2605.24057#S5.SS6.p1.4)\.
- T\. Hua, W\. Wang, Z\. Xue, S\. Ren, Y\. Wang, and H\. Zhao \(2021\)On feature decorrelation in self\-supervised learning\.InIEEE/CVF International Conference on Computer Vision \(ICCV\),Cited by:[§2](https://arxiv.org/html/2605.24057#S2.SS0.SSS0.Px4.p1.1)\.
- L\. Jing, P\. Vincent, Y\. LeCun, and Y\. Tian \(2022\)Understanding dimensional collapse in contrastive self\-supervised learning\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2605.24057#S2.SS0.SSS0.Px4.p1.1)\.
- H\. A\. Kramers \(1940\)Brownian motion in a field of force and the diffusion model of chemical reactions\.Physica7\(4\),pp\. 284–304\.Cited by:[Appendix C](https://arxiv.org/html/2605.24057#A3.SS0.SSS0.Px3.p1.4)\.
- D\. G\. Mixon, H\. Parshall, and J\. Pi \(2022\)Neural collapse with unconstrained features\.Sampling Theory, Signal Processing, and Data Analysis20\(11\)\.Cited by:[§2](https://arxiv.org/html/2605.24057#S2.SS0.SSS0.Px2.p1.1)\.
- N\. Nanda, L\. Chan, T\. Lieberum, J\. Smith, and J\. Steinhardt \(2023\)Progress measures for grokking via mechanistic interpretability\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2605.24057#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2605.24057#S2.SS0.SSS0.Px5.p1.2),[§4\.2](https://arxiv.org/html/2605.24057#S4.SS2.SSS0.Px1.p1.10),[§5\.4](https://arxiv.org/html/2605.24057#S5.SS4.SSS0.Px2.p1.14)\.
- V\. Papyan, X\. Y\. Han, and D\. L\. Donoho \(2020\)Prevalence of neural collapse during the terminal phase of deep learning training\.Proceedings of the National Academy of Sciences117\(40\),pp\. 24652–24663\.Cited by:[§2](https://arxiv.org/html/2605.24057#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Power, Y\. Burda, H\. Edwards, I\. Babuschkin, and V\. Misra \(2022\)Grokking: generalization beyond overfitting on small algorithmic datasets\.arXiv preprint arXiv:2201\.02177\.Cited by:[§2](https://arxiv.org/html/2605.24057#S2.SS0.SSS0.Px5.p1.2)\.
- K\. Rose, E\. Gurewitz, and G\. C\. Fox \(1990\)Statistical mechanics and phase transitions in clustering\.Physical Review Letters65\(8\),pp\. 945–948\.Cited by:[§A\.3](https://arxiv.org/html/2605.24057#A1.SS3.SSS0.Px1),[item 1](https://arxiv.org/html/2605.24057#S1.I1.i1.p1.4),[§2](https://arxiv.org/html/2605.24057#S2.SS0.SSS0.Px3.p1.7),[§3\.2](https://arxiv.org/html/2605.24057#S3.SS2.p1.15),[§3\.4](https://arxiv.org/html/2605.24057#S3.SS4.p2.4)\.
- A\. M\. Saxe, J\. L\. McClelland, and S\. Ganguli \(2014\)Exact solutions to the nonlinear dynamics of learning in deep linear neural networks\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2605.24057#S2.SS0.SSS0.Px5.p1.2)\.
- P\. Súkeník, C\. H\. Lampert, and M\. Mondelli \(2024\)Neural collapse versus low\-rank bias: is deep neural collapse really optimal?\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2](https://arxiv.org/html/2605.24057#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Templeton, T\. Conerly, J\. Marcus, J\. Lindsey, T\. Bricken, B\. Chen, A\. Pearce, C\. Citro, E\. Ameisen, A\. Jermyn,et al\.\(2024\)Scaling monosemanticity: extracting interpretable features from Claude 3 Sonnet\.Transformer Circuits Thread\.Cited by:[§2](https://arxiv.org/html/2605.24057#S2.SS0.SSS0.Px1.p1.1)\.
- T\. Tirer and J\. Bruna \(2022\)Extended unconstrained features model for exploring deep neural collapse\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§2](https://arxiv.org/html/2605.24057#S2.SS0.SSS0.Px2.p1.1)\.
- S\. Wang and S\. E\. Palmer \(2023\)Towards understanding neural collapse in supervised contrastive learning with the information bottleneck method\.arXiv preprint arXiv:2305\.11957\.Cited by:[§2](https://arxiv.org/html/2605.24057#S2.SS0.SSS0.Px2.p1.1)\.
- Z\. Wang and L\. Ziyin \(2022\)Posterior collapse of a linear latent variable model\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2](https://arxiv.org/html/2605.24057#S2.SS0.SSS0.Px6.p1.1)\.
- J\. Zhou, X\. Li, T\. Ding, C\. You, Q\. Qu, and Z\. Zhu \(2022\)On the optimization landscape of neural collapse under MSE loss: global optimality with unconstrained features\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§2](https://arxiv.org/html/2605.24057#S2.SS0.SSS0.Px2.p1.1)\.
- L\. Ziyin and M\. Ueda \(2023\)Zeroth, first, and second\-order phase transitions in deep neural networks\.Physical Review Research5\(4\),pp\. 043243\.Cited by:[Appendix C](https://arxiv.org/html/2605.24057#A3.SS0.SSS0.Px6.p1.6),[§2](https://arxiv.org/html/2605.24057#S2.SS0.SSS0.Px6.p1.1)\.
## Appendix AFull Hessian derivation
We compute the second derivative ofℒμ\\mathcal\{L\}\_\{\\mu\}at the symmetric state𝒮0=\{μk=z¯\}\\mathcal\{S\}\_\{0\}=\\\{\\mu\_\{k\}=\\bar\{z\}\\\}in coordinates\(k,a\)\(k,a\), wherek∈\{1,…,K\}k\\\!\\in\\\!\\\{1,\\dots,K\\\}indexes the component anda∈\{1,…,d\}a\\\!\\in\\\!\\\{1,\\dots,d\\\}indexes the spatial direction\.
### A\.1Explicit form of the Hessian
###### Theorem 1\.
At the symmetric collapsed state𝒮0\\mathcal\{S\}\_\{0\},
∂2ℒμ∂μka∂μlb\|𝒮0=βKδklδab−β2K\(δkl−1K\)Σab\.\\left\.\\frac\{\\partial^\{2\}\\mathcal\{L\}\_\{\\mu\}\}\{\\partial\\mu\_\{k\}^\{a\}\\,\\partial\\mu\_\{l\}^\{b\}\}\\right\|\_\{\\mathcal\{S\}\_\{0\}\}\\;=\\;\\frac\{\\beta\}\{K\}\\,\\delta\_\{kl\}\\,\\delta^\{ab\}\\;\-\\;\\frac\{\\beta^\{2\}\}\{K\}\\,\\Bigl\(\\delta\_\{kl\}\-\\tfrac\{1\}\{K\}\\Bigr\)\\,\\Sigma^\{ab\}\.\(10\)
###### Proof\.
For each samplezzwritefk\(z;μ\):=−β2‖z−μk‖2f\_\{k\}\(z;\\mu\):=\-\\tfrac\{\\beta\}\{2\}\\\|z\-\\mu\_\{k\}\\\|^\{2\}andg\(z;μ\):=LSEkfk\(z;μ\)g\(z;\\mu\):=\\mathrm\{LSE\}\_\{k\}\\,f\_\{k\}\(z;\\mu\), so the per\-sample loss is−g\(z;μ\)\-g\(z;\\mu\)\. We need−∂2g/\(∂μka∂μlb\)\-\\partial^\{2\}g/\(\\partial\\mu\_\{k\}^\{a\}\\,\\partial\\mu\_\{l\}^\{b\}\)at𝒮0\\mathcal\{S\}\_\{0\}, averaged overzz\.
#### First derivative\.
By the softmax identity forLSE\\mathrm\{LSE\},
∂g∂μka=∑jpj\(z;μ\)∂fj∂μka=pk\(z;μ\)⋅β\(za−μka\),\\frac\{\\partial g\}\{\\partial\\mu\_\{k\}^\{a\}\}\\;=\\;\\sum\_\{j\}p\_\{j\}\(z;\\mu\)\\,\\frac\{\\partial f\_\{j\}\}\{\\partial\\mu\_\{k\}^\{a\}\}\\;=\\;p\_\{k\}\(z;\\mu\)\\cdot\\beta\(z^\{a\}\-\\mu\_\{k\}^\{a\}\),wherepk\(z;μ\):=softmaxk\(f⋅\(z;μ\)\)p\_\{k\}\(z;\\mu\):=\\mathrm\{softmax\}\_\{k\}\(f\_\{\\cdot\}\(z;\\mu\)\)\.
#### Second derivative\.
Differentiating once more,
∂2g∂μka∂μlb=∂pk∂μlb⋅β\(za−μka\)⏟\(I\)\+pk⋅β\(−δklδab\)⏟\(II\)\.\\frac\{\\partial^\{2\}g\}\{\\partial\\mu\_\{k\}^\{a\}\\,\\partial\\mu\_\{l\}^\{b\}\}\\;=\\;\\underbrace\{\\frac\{\\partial p\_\{k\}\}\{\\partial\\mu\_\{l\}^\{b\}\}\\cdot\\beta\(z^\{a\}\-\\mu\_\{k\}^\{a\}\)\}\_\{\\text\{\(I\)\}\}\\;\+\\;\\underbrace\{p\_\{k\}\\cdot\\beta\\,\\bigl\(\-\\delta\_\{kl\}\\,\\delta^\{ab\}\\bigr\)\}\_\{\\text\{\(II\)\}\}\.Using the standard softmax derivative identity∂pk/∂fl=pk\(δkl−pl\)\\partial p\_\{k\}/\\partial f\_\{l\}=p\_\{k\}\(\\delta\_\{kl\}\-p\_\{l\}\)together with∂fl/∂μlb=β\(zb−μlb\)\\partial f\_\{l\}/\\partial\\mu\_\{l\}^\{b\}=\\beta\(z^\{b\}\-\\mu\_\{l\}^\{b\}\),
∂pk∂μlb=pk\(δkl−pl\)β\(zb−μlb\),\\frac\{\\partial p\_\{k\}\}\{\\partial\\mu\_\{l\}^\{b\}\}\\;=\\;p\_\{k\}\(\\delta\_\{kl\}\-p\_\{l\}\)\\,\\beta\(z^\{b\}\-\\mu\_\{l\}^\{b\}\),so
∂2g∂μka∂μlb=β2pk\(δkl−pl\)\(za−μka\)\(zb−μlb\)−βpkδklδab\.\\frac\{\\partial^\{2\}g\}\{\\partial\\mu\_\{k\}^\{a\}\\,\\partial\\mu\_\{l\}^\{b\}\}\\;=\\;\\beta^\{2\}\\,p\_\{k\}\(\\delta\_\{kl\}\-p\_\{l\}\)\\,\(z^\{a\}\-\\mu\_\{k\}^\{a\}\)\(z^\{b\}\-\\mu\_\{l\}^\{b\}\)\\;\-\\;\\beta\\,p\_\{k\}\\,\\delta\_\{kl\}\\,\\delta^\{ab\}\.
#### Evaluation at𝒮0\\mathcal\{S\}\_\{0\}\.
Atμk=μl=z¯\\mu\_\{k\}=\\mu\_\{l\}=\\bar\{z\}we havepk=pl=1/Kp\_\{k\}=p\_\{l\}=1/Kand\(za−μka\)=\(za−z¯a\)\(z^\{a\}\-\\mu\_\{k\}^\{a\}\)=\(z^\{a\}\-\\bar\{z\}^\{a\}\), so
𝔼z\[\(za−z¯a\)\(zb−z¯b\)\]=Σab,pk\(δkl−pl\)\|𝒮0=1K\(δkl−1K\)\.\\mathbb\{E\}\_\{z\}\\\!\\left\[\(z^\{a\}\-\\bar\{z\}^\{a\}\)\(z^\{b\}\-\\bar\{z\}^\{b\}\)\\right\]\\;=\\;\\Sigma^\{ab\},\\qquad p\_\{k\}\(\\delta\_\{kl\}\-p\_\{l\}\)\\big\|\_\{\\mathcal\{S\}\_\{0\}\}\\;=\\;\\tfrac\{1\}\{K\}\\bigl\(\\delta\_\{kl\}\-\\tfrac\{1\}\{K\}\\bigr\)\.Taking the expectation overzzand negating \(recall the loss is−g\-g\),
−𝔼z\[∂2g∂μka∂μlb\]𝒮0=βKδklδab−β2K\(δkl−1K\)Σab,\-\\mathbb\{E\}\_\{z\}\\\!\\left\[\\frac\{\\partial^\{2\}g\}\{\\partial\\mu\_\{k\}^\{a\}\\,\\partial\\mu\_\{l\}^\{b\}\}\\right\]\_\{\\mathcal\{S\}\_\{0\}\}\\;=\\;\\tfrac\{\\beta\}\{K\}\\,\\delta\_\{kl\}\\,\\delta^\{ab\}\\;\-\\;\\tfrac\{\\beta^\{2\}\}\{K\}\\bigl\(\\delta\_\{kl\}\-\\tfrac\{1\}\{K\}\\bigr\)\\,\\Sigma^\{ab\},which is equation[10](https://arxiv.org/html/2605.24057#A1.E10)\. ∎
The first term in equation[10](https://arxiv.org/html/2605.24057#A1.E10)is the curvature of each isotropic Gaussian component \(positive and isotropic\); the second is the softmax\-coupling term, which is sign\-indefinite and picks up the spatial covarianceΣ\\Sigma\.
### A\.2Eigendecomposition
The Hessian equation[10](https://arxiv.org/html/2605.24057#A1.E10)is separable across the component \(k,lk,l\) and spatial \(a,ba,b\) indices, so it admits eigenvectors of product formξka=wkua\\xi\_\{k\}^\{a\}=w\_\{k\}\\,u^\{a\}withw∈ℝK,u∈ℝdw\\\!\\in\\\!\\mathbb\{R\}^\{K\},u\\\!\\in\\\!\\mathbb\{R\}^\{d\}\. Acting on such a vector,
\(Hξ\)ka=βKwkua−β2K\(Σu\)a\(wk−w¯\),w¯:=1K∑lwl\.\(H\\xi\)\_\{k\}^\{a\}\\;=\\;\\tfrac\{\\beta\}\{K\}\\,w\_\{k\}\\,u^\{a\}\\;\-\\;\\tfrac\{\\beta^\{2\}\}\{K\}\\,\(\\Sigma u\)^\{a\}\\,\\bigl\(w\_\{k\}\-\\bar\{w\}\\bigr\),\\qquad\\bar\{w\}:=\\tfrac\{1\}\{K\}\\\!\\sum\_\{l\}w\_\{l\}\.The component vectorwwcouples only through its meanw¯\\bar\{w\}, which selects between two channels\.
Symmetric channel \(wk=cw\_\{k\}=cfor allkk, sowk−w¯=0w\_\{k\}\-\\bar\{w\}=0\):The action reduces to\(Hξ\)ka=\(β/K\)cua\(H\\xi\)\_\{k\}^\{a\}=\(\\beta/K\)\\,c\\,u^\{a\}\. The eigenvalue isβ/K\\beta/K, independent ofΣ\\Sigmaand ofuu\. Allddsymmetric\-channel modes have eigenvalueβ/K\>0\\beta/K\>0, so they are always stable\. Geometrically these are bulk translations of allKKprototypes together\.
Anti\-symmetric channel \(w¯=0\\bar\{w\}=0,\(K−1\)\(K\{\-\}1\)\-fold degenerate in component space\):Thenwk−w¯=wkw\_\{k\}\-\\bar\{w\}=w\_\{k\}and
\(Hξ\)ka=βKwk\[ua−β\(Σu\)a\]\.\(H\\xi\)\_\{k\}^\{a\}\\;=\\;\\tfrac\{\\beta\}\{K\}\\,w\_\{k\}\\,\\bigl\[u^\{a\}\-\\beta\\,\(\\Sigma u\)^\{a\}\\bigr\]\.Soξ=w⊗u\\xi=w\\\!\\otimes\\\!uis an eigenvector with eigenvalue\(β/K\)\(1−βσ2\)\(\\beta/K\)\(1\-\\beta\\,\\sigma^\{2\}\)wheneveruuis an eigenvector ofΣ\\Sigmawith eigenvalueσ2\\sigma^\{2\}\. The anti\-symmetric spatial spectrum is
λi⟂\(β\)=βK\(1−βσi2\),i=1,…,d,\\lambda^\{\\perp\}\_\{i\}\(\\beta\)\\;=\\;\\frac\{\\beta\}\{K\}\\,\\bigl\(1\-\\beta\\,\\sigma\_\{i\}^\{2\}\\bigr\),\\qquad i=1,\\dots,d,\(11\)each\(K−1\)\(K\-1\)\-fold degenerate in component space\.
### A\.3Critical precision
The lowest anti\-symmetric eigenvalue is
λ1⟂\(β\)=βK\(1−βλmax\(Σ\)\)\.\\lambda^\{\\perp\}\_\{1\}\(\\beta\)\\;=\\;\\tfrac\{\\beta\}\{K\}\\bigl\(1\-\\beta\\,\\lambda\_\{\\max\}\(\\Sigma\)\\bigr\)\.Forβ\>0\\beta\>0it crosses zero exactly whenβλmax\(Σ\)=1\\beta\\,\\lambda\_\{\\max\}\(\\Sigma\)=1:
βc=1λmax\(Σ\)\.\\boxed\{\\;\\beta\_\{c\}\\;=\\;\\frac\{1\}\{\\lambda\_\{\\max\}\(\\Sigma\)\}\.\\;\}Forβ<βc\\beta<\\beta\_\{c\}, every anti\-symmetric eigenvalue is positive and𝒮0\\mathcal\{S\}\_\{0\}is a local minimum ofℒμ\\mathcal\{L\}\_\{\\mu\}\. Forβ\>βc\\beta\>\\beta\_\{c\},λ1⟂<0\\lambda^\{\\perp\}\_\{1\}<0and𝒮0\\mathcal\{S\}\_\{0\}is a saddle with\(K−1\)\(K\-1\)unstable directions in component space combined with the principal eigenvector ofΣ\\Sigmain spatial space\.
#### Parametrization note vs\.Roseet al\.\[[1990](https://arxiv.org/html/2605.24057#bib.bib7)\]\.
Rose states the soft\-KK\-means critical temperature asTc=2λmax\(Σ\)T\_\{c\}=2\\,\\lambda\_\{\\max\}\(\\Sigma\)in the conventionp\(k∣z\)∝exp\(−‖z−μk‖2/T\)p\(k\\\!\\mid\\\!z\)\\propto\\exp\(\-\\\|z\-\\mu\_\{k\}\\\|^\{2\}/T\)\. Our conventionp\(k∣z\)∝exp\(−β2‖z−μk‖2\)p\(k\\\!\\mid\\\!z\)\\propto\\exp\(\-\\tfrac\{\\beta\}\{2\}\\\|z\-\\mu\_\{k\}\\\|^\{2\}\)corresponds toβ=2/T\\beta=2/T\. Substituting, Rose’sTc=2λmaxT\_\{c\}=2\\,\\lambda\_\{\\max\}becomesβc=2/Tc=1/λmax\\beta\_\{c\}=2/T\_\{c\}=1/\\lambda\_\{\\max\}, matching our result; the factor of22must be translated when crossing conventions\. We verify this numerically by performing a direct eigenvalue scan of the full Hessian on toy bimodal data: the lowest eigenvalue zero\-crosses atβ=1/λmax\(Σ\)\\beta=1/\\lambda\_\{\\max\}\(\\Sigma\)to four decimals, exactly as predicted\.
### A\.4Geometric interpretation of the unstable mode
Atβ=βc\\beta=\\beta\_\{c\}the zero\-eigenvalue mode is the tensor productw⊗uw\\\!\\otimes\\\!uwith∑kwk=0\\sum\_\{k\}w\_\{k\}=0anduuthe principal eigenvector ofΣ\\Sigma\. The component vectorwwassigns signs to theKKprototypes; the spatial vectoruupicks the axis of maximum data variance\. Any anti\-symmetricwwis in the unstable subspace, so the assignment of “which prototype goes where” is not fixed by the linear analysis: it is decided by initialization noise and by the cubic terms in the pitchfork normal form equation[5](https://arxiv.org/html/2605.24057#S3.E5)\.
## Appendix BProof of Proposition[1](https://arxiv.org/html/2605.24057#Thmproposition1)
LetΔ\(t\):=β\(t\)−βc\(t\)\\Delta\(t\):=\\beta\(t\)\-\\beta\_\{c\}\(t\)whereβc\(t\)=1/λmax\(Cov\(z\(t\)\)\)\\beta\_\{c\}\(t\)=1/\\lambda\_\{\\max\}\(\\mathrm\{Cov\}\(z\(t\)\)\)depends on the encoder state at timett\. Bothβ\(t\)\\beta\(t\)andβc\(t\)\\beta\_\{c\}\(t\)are continuous inttfor any continuous\-in\-time training dynamic \(gradient flow, or piecewise\-constant extension of any gradient\-step iteration\), soΔ\\Deltais continuous\.
#### Initial\-time inequality\.
Att=0t=0the encoder is at random initialization; the marginalz\(0\)=enc\(x;φ\(0\)\)z\(0\)=\\mathrm\{enc\}\(x;\\varphi\(0\)\)has not yet been spread by training\. Meanwhileβ\\betastarts at the user\-chosen initial valueβ\(0\)=exp\(logβ0\)\\beta\(0\)=\\exp\(\\log\\beta\_\{0\}\)which is below the data scale by design \(in all our experimentslogβ0=−2\.5\\log\\beta\_\{0\}=\-2\.5\)\. Henceβ\(0\)<βc\(0\)\\beta\(0\)<\\beta\_\{c\}\(0\), i\.e\.Δ\(0\)<0\\Delta\(0\)<0\.
#### Asymptotic positivity\.
By hypothesis \(1\) of Prop\.[1](https://arxiv.org/html/2605.24057#Thmproposition1),
lim inft→∞β\(t\)\>c1\.\\liminf\_\{t\\to\\infty\}\\beta\(t\)\>c\_\{1\}\.By hypothesis \(3\),
lim supt→∞βc\(t\)<c1\.\\limsup\_\{t\\to\\infty\}\\beta\_\{c\}\(t\)<c\_\{1\}\.Subtracting,
lim inft→∞Δ\(t\)≥lim inft→∞β\(t\)−lim supt→∞βc\(t\)\>0\.\\liminf\_\{t\\to\\infty\}\\Delta\(t\)\\;\\geq\\;\\liminf\_\{t\\to\\infty\}\\beta\(t\)\-\\limsup\_\{t\\to\\infty\}\\beta\_\{c\}\(t\)\\;\>\\;0\.
#### Intermediate value theorem\.
Δ\\Deltais continuous withΔ\(0\)<0\\Delta\(0\)<0andlim inft→∞Δ\(t\)\>0\\liminf\_\{t\\to\\infty\}\\Delta\(t\)\>0, so there exists a finiteTTwithΔ\(T\)\>0\\Delta\(T\)\>0\. By IVT applied toΔ\\Deltaon\[0,T\]\[0,T\]there is at⋆∈\(0,T\)t^\{\\star\}\\\!\\in\\\!\(0,T\)withΔ\(t⋆\)=0\\Delta\(t^\{\\star\}\)=0, i\.e\.β\(t⋆\)=βc\(t⋆\)\\beta\(t^\{\\star\}\)=\\beta\_\{c\}\(t^\{\\star\}\)\.
#### Instability at the crossing\.
Att=t⋆t=t^\{\\star\}the static Hessian analysis of Appendix[A](https://arxiv.org/html/2605.24057#A1)applies pointwise withΣ=Cov\(z\(t⋆\)\)\\Sigma=\\mathrm\{Cov\}\(z\(t^\{\\star\}\)\), yieldingβc\(t⋆\)=1/λmax\(Cov\(z\(t⋆\)\)\)\\beta\_\{c\}\(t^\{\\star\}\)=1/\\lambda\_\{\\max\}\(\\mathrm\{Cov\}\(z\(t^\{\\star\}\)\)\)\. By constructionβ\(t⋆\)=βc\(t⋆\)\\beta\(t^\{\\star\}\)=\\beta\_\{c\}\(t^\{\\star\}\), so the lowest anti\-symmetric eigenvalueλ1⟂\\lambda^\{\\perp\}\_\{1\}from equation[11](https://arxiv.org/html/2605.24057#A1.E11)equals zero att⋆t^\{\\star\}; immediately past it \(assuming the encoder continues to spread the latent so thatβc\\beta\_\{c\}continues to decrease whileβ\\betacontinues to increase\) we haveβ\>βc\\beta\>\\beta\_\{c\},λ1⟂<0\\lambda^\{\\perp\}\_\{1\}<0, and the symmetric state becomes a saddle\. The prototypes therefore begin to pitchfork att⋆t^\{\\star\}\.□\\square
#### Remark on the substantive content of the hypotheses\.
Hypothesis \(1\) holds for any likelihood\-maximizing GMM step: the NLL is a monotone\-decreasing function ofβ\\betaat fixed prototype positions in the small\-β\\betaregime, and Adam/SGD on the NLL pushesβ\\betaupward; we have not observed an experiment in whichβ\\betadecreases on average\. Hypothesis \(2\) is the assertion that the encoder spreads the latent over time, which is the defining property of any information\-preserving SSL objective \(contrastive, predictive, autoencoder reconstruction\)\. Hypothesis \(3\) requires that the encoder eventually finds enough variance for the GMM to resolve clusters at the chosenβ\\betascale; this fails for collapsing encoders \(e\.g\. DINO without centering or EMA, see Sec\.[6\.1](https://arxiv.org/html/2605.24057#S6.SS1)\), in which caseβc\(t\)\\beta\_\{c\}\(t\)diverges and no crossing occurs; consistent with the framework’s prediction that those encoders do not undergo healthy bifurcation\.
## Appendix CEmpirical characterization of post\-critical escape under weight\-decay intervention
This appendix is a methodological cash\-out of Remark[1](https://arxiv.org/html/2605.24057#Thmremark1)and the control experiment of Sec\.[4\.2](https://arxiv.org/html/2605.24057#S4.SS2): it characterizes how the metastable plateau length scales with the encoder’s dissipation strength on the grokking setup, and identifies which of the two physically distinct escape regimes \(activation\- vs\. drift\-dominated\) the data sit in\. The output is empirical, not theoretical; we do*not*claim a closed\-form prediction forτesc\\tau\_\{\\mathrm\{esc\}\}from the bifurcation framework alone; the qualitative predictions \(τesc→∞\\tau\_\{\\mathrm\{esc\}\}\\\!\\to\\\!\\inftyasγ→0\\gamma\\\!\\to\\\!0, monotonicity inγ\\gamma\) are content of Remark[1](https://arxiv.org/html/2605.24057#Thmremark1)\.
We work in the pitchfork normal form equation[7](https://arxiv.org/html/2605.24057#S3.E7)post\-critically \(μ:=β−βc\>0\\mu:=\\beta\-\\beta\_\{c\}\>0\), with an additional drift−γU′\(ε\)\-\\gamma\\,U^\{\\prime\}\(\\varepsilon\)from the encoder’s upstream loss component that favors the memorization basin\. The full Langevin dynamics is
ε˙=με−αε3−γU′\(ε\)\+η\(t\),⟨η\(t\)η\(t′\)⟩=2Dδ\(t−t′\)\.\\dot\{\\varepsilon\}\\;=\\;\\mu\\,\\varepsilon\-\\alpha\\varepsilon^\{3\}\-\\gamma\\,U^\{\\prime\}\(\\varepsilon\)\+\\eta\(t\),\\qquad\\langle\\eta\(t\)\\eta\(t^\{\\prime\}\)\\rangle=2D\\,\\delta\(t\-t^\{\\prime\}\)\.\(12\)
#### Effective potential\.
The deterministic drift is−Veff′\(ε\)\-V^\{\\prime\}\_\{\\mathrm\{eff\}\}\(\\varepsilon\)with
Veff\(ε\)=−12με2\+14αε4\+γU\(ε\)\.V\_\{\\mathrm\{eff\}\}\(\\varepsilon\)\\;=\\;\-\\tfrac\{1\}\{2\}\\mu\\,\\varepsilon^\{2\}\+\\tfrac\{1\}\{4\}\\alpha\\,\\varepsilon^\{4\}\+\\gamma\\,U\(\\varepsilon\)\.\(13\)Atγ=0\\gamma=0,VeffV\_\{\\mathrm\{eff\}\}has a saddle atε=0\\varepsilon=0\(already unstable, sinceμ\>0\\mu\>0\) and two symmetric broken\-symmetry minima at±ε⋆\\pm\\varepsilon^\{\\star\}withε⋆=μ/α\\varepsilon^\{\\star\}=\\sqrt\{\\mu/\\alpha\}\. For the grokking setup of Sec\.[4\.2](https://arxiv.org/html/2605.24057#S4.SS2), the loss landscape nearε=0\\varepsilon=0is dominated by the memorization plateau: even post\-critically, the encoder remains in a meta\-stable region untilUU\(the regularization landscape\) tiltsVeffV\_\{\\mathrm\{eff\}\}enough to makeε=0\\varepsilon=0an effective saddle in the trans\-basin sense\.
#### Two regimes\.
Escape from the memorization basin \(centered atε≈0\\varepsilon\\\!\\approx\\\!0\) to the broken\-symmetry basin \(atε≈ε⋆\\varepsilon\\\!\\approx\\\!\\varepsilon^\{\\star\}\) is governed by the competition between two timescales:*activation*\(Kramers barrier crossing, set byΔSeff/D\\Delta S\_\{\\mathrm\{eff\}\}/D\) and*deterministic drift*\(relaxation rate set by the unstable direction near the saddle,∼1/\(μ\+γU′′\(0\)\)\\sim\\\!1/\(\\mu\+\\gamma U^\{\\prime\\prime\}\(0\)\)\)\. Which regime dominates depends on whether the noise must climb a tall barrier \(ΔSeff≫D\\Delta S\_\{\\mathrm\{eff\}\}\\gg D\) or whether the tiltγU′\(ε\)\\gamma\\,U^\{\\prime\}\(\\varepsilon\)already removes the barrier so that escape proceeds by gradient flow\.
#### Activation\-dominated \(Kramers\) regime\.
IfVeffV\_\{\\mathrm\{eff\}\}retains a barrier of heightΔS\\Delta Satγ=0\\gamma=0and the dissipation contribution tilts it linearly,
ΔSeff\(γ\)=ΔS−κγ\+O\(γ2\),κ=U\(0\)−U\(ε⋆\)\>0,\\Delta S\_\{\\mathrm\{eff\}\}\(\\gamma\)\\;=\\;\\Delta S\-\\kappa\\,\\gamma\+O\(\\gamma^\{2\}\),\\qquad\\kappa=U\(0\)\-U\(\\varepsilon^\{\\star\}\)\>0,\(14\)then forΔSeff≫D\\Delta S\_\{\\mathrm\{eff\}\}\\gg Dclassical Kramers theory\[Kramers,[1940](https://arxiv.org/html/2605.24057#bib.bib20), Berglund,[2013](https://arxiv.org/html/2605.24057#bib.bib21)\]gives
τesc≍τ0exp\(ΔSeff/D\)=τ0exp\(ΔS−κγD\),\\tau\_\{\\mathrm\{esc\}\}\\;\\asymp\\;\\tau\_\{0\}\\,\\exp\\\!\\Bigl\(\\Delta S\_\{\\mathrm\{eff\}\}/D\\Bigr\)\\;=\\;\\tau\_\{0\}\\exp\\\!\\Bigl\(\\frac\{\\Delta S\-\\kappa\\gamma\}\{D\}\\Bigr\),\(15\)withτ0∼1/\(μ⋅ωb\)=O\(1/\(β−βc\)\)\\tau\_\{0\}\\sim 1/\(\\mu\\cdot\\omega\_\{b\}\)=O\(1/\(\\beta\-\\beta\_\{c\}\)\)\.
#### Drift\-dominated regime\.
If instead the bare barrierΔS\\Delta Sis small \(or the tilt has already collapsed it,κγ≳ΔS\\kappa\\gamma\\gtrsim\\Delta S\), escape is set by the deterministic linear instability near the post\-critical saddle\. Linearizing equation[12](https://arxiv.org/html/2605.24057#A3.E12)aboutε=0\\varepsilon=0,
ε˙≈\(μ\+γU′′\(0\)\)ε\+η\(t\),\\dot\{\\varepsilon\}\\;\\approx\\;\\bigl\(\\mu\+\\gamma\\,U^\{\\prime\\prime\}\(0\)\\bigr\)\\,\\varepsilon\+\\eta\(t\),the system grows exponentially with rateλ\(γ\):=μ\+γU′′\(0\)\\lambda\(\\gamma\):=\\mu\+\\gamma U^\{\\prime\\prime\}\(0\)until nonlinear saturation at\|ε\|∼ε⋆\|\\varepsilon\|\\\!\\sim\\\!\\varepsilon^\{\\star\}\. The escape time is then
τesc∼1λ\(γ\)lnε⋆D/λ\(γ\),\\tau\_\{\\mathrm\{esc\}\}\\;\\sim\\;\\frac\{1\}\{\\lambda\(\\gamma\)\}\\,\\ln\\\!\\frac\{\\varepsilon^\{\\star\}\}\{\\sqrt\{D/\\lambda\(\\gamma\)\}\},\(16\)whereD/λ\\sqrt\{D/\\lambda\}is the equilibrium spread inside the linear region\. In the limitγU′′\(0\)≫μ\\gamma U^\{\\prime\\prime\}\(0\)\\gg\\mu, this collapses to a power\-lawτesc≍Aγ−p\\tau\_\{\\mathrm\{esc\}\}\\asymp A\\,\\gamma^\{\-p\}with leading exponentp=1p=1from the prefactor; nonlinearities inUU\(so that the effective drift growth rate isγU′′\(0\)\+O\(γ2U′′′\(0\)\)\\gamma U^\{\\prime\\prime\}\(0\)\+O\(\\gamma^\{2\}U^\{\\prime\\prime\\prime\}\(0\)\)\) and the logarithmic noise correction inflate the effective exponent top≥1p\\geq 1\.
#### Empirical fit and regime selection \(Table[3](https://arxiv.org/html/2605.24057#S4.T3)\)\.
The 6\-point WD sweep atp=97p\\\!=\\\!97, train fraction0\.30\.3, withn=3n=3seeds per WD level and a200,000200\{,\}000\-step horizon, givesτesc\\tau\_\{\\mathrm\{esc\}\}atγ∈\{0\.1,0\.2,0\.3,0\.5,0\.7,1\.0\}\\gamma\\in\\\{0\.1,0\.2,0\.3,0\.5,0\.7,1\.0\\\}ranging8 9008\\,900to147 167147\\,167steps\. Fitting the two functional forms equation[15](https://arxiv.org/html/2605.24057#A3.E15)and equation[16](https://arxiv.org/html/2605.24057#A3.E16)by least squares inlogτesc\\log\\tau\_\{\\mathrm\{esc\}\}:
Power\-law:logτesc=9\.11−1\.225logγ,χ2=1\.52,AIC=5\.52,\\displaystyle\\log\\tau\_\{\\mathrm\{esc\}\}\\;=\\;9\.11\-1\.225\\,\\log\\gamma,\\qquad\\chi^\{2\}=1\.52,\\quad\\mathrm\{AIC\}=5\.52,Kramers:logτesc=11\.65−2\.631γ,χ2=20\.78,AIC=24\.78\.\\displaystyle\\log\\tau\_\{\\mathrm\{esc\}\}\\;=\\;11\.65\-2\.631\\,\\gamma,\\qquad\\chi^\{2\}=20\.78,\\quad\\mathrm\{AIC\}=24\.78\.The model\-comparison gap isΔAIC=\+19\.26\\Delta\\mathrm\{AIC\}=\+19\.26\(andΔBIC=\+19\.27\\Delta\\mathrm\{BIC\}=\+19\.27\) in favor of the power\-law form\. Per\-point residuals:
The Kramers form systematically under\-predicts escape times at lowγ\\gamma\(by a factor of∼1\.7\\sim\\\!1\.7at WD=0\.1=\\\!0\.1\): the exponential decay inγ\\gammaovershoots the actual slow increase ofτesc\\tau\_\{\\mathrm\{esc\}\}as dissipation is reduced\. The power\-law form, by contrast, is within∼10%\\sim\\\!10\\%at every WD level \(the largest residual is atγ=0\.2\\gamma\\\!=\\\!0\.2, well within the∼30%\\sim\\\!30\\%seed\-level CV\)\. The fitted exponentp=1\.23p\\\!=\\\!1\.23lies in the predictedp≥1p\\geq 1range\. We conclude that, in our grokking configuration, the post\-critical metastable plateau is escaped by*drift\-dominated*dynamics rather than by activated barrier\-crossing: the WD tilt is large enough relative to both the bare barrier and the noise scale that the deterministic flow onVeffV\_\{\\mathrm\{eff\}\}governs the escape time\.
#### Limitations\.
The reduction to one\-dimensional Langevin dynamics onε\\varepsilonassumes \(a\) the trans\-basin geometry is quasi\-static during the metastable plateau \(encoder evolution is slow relative to escape attempts\) and \(b\) the SGD noise is well\-approximated as Langevin\[Ziyin and Ueda,[2023](https://arxiv.org/html/2605.24057#bib.bib10)\]\. The regime\-selection conclusion; drift\-dominated rather than activation\-dominated; is specific to the modular\-arithmetic grokking setup atμ∼10−4\\mu\\\!\\sim\\\!10^\{\-4\}–10−310^\{\-3\}, train fraction0\.30\.3, and WD∈\[0\.1,1\.0\]\\in\[0\.1,1\.0\]; in deeper or wider barriers \(e\.g\., largerpp, much smaller train fraction\) the activation regime may re\-emerge\. Verifying the regime classification at the microscopic level for the grokking circuit is beyond the scope of this paper\.
## Appendix DProof of Proposition[2](https://arxiv.org/html/2605.24057#Thmproposition2)\(lottery mode\-selection\)
We work in the unstable subspace atμ:=β−βc\>0\\mu:=\\beta\-\\beta\_\{c\}\>0\(small\) and analyze the dynamics equation[8](https://arxiv.org/html/2605.24057#S5.E8)\. Below we sketch the proof of equation[9](https://arxiv.org/html/2605.24057#S5.E9)in two steps: decoupled\-mode analysis \(γ=0\\gamma\\\!=\\\!0\), then perturbative inclusion of inter\-mode coupling\. Throughout we parametrize each mode by its magnituderk=‖𝜺k‖r\_\{k\}=\\\|\\bm\{\\varepsilon\}\_\{k\}\\\|and directiondk=𝜺k/rkd\_\{k\}=\\bm\{\\varepsilon\}\_\{k\}/r\_\{k\}on the sphereSd−1S^\{d\-1\}\.
### D\.1Decoupled\-mode analysis \(γ=0\\gamma=0\)
Each mode obeys an independent SDE inℝd\\mathbb\{R\}^\{d\}:
𝜺˙=μ𝜺−α‖𝜺‖2𝜺\+𝜼\(t\),⟨𝜼\(t\)𝜼\(t′\)⊤⟩=2Dδ\(t−t′\)𝑰\.\\dot\{\\bm\{\\varepsilon\}\}\\;=\\;\\mu\\,\\bm\{\\varepsilon\}\-\\alpha\\,\\\|\\bm\{\\varepsilon\}\\\|^\{2\}\\bm\{\\varepsilon\}\+\\bm\{\\eta\}\(t\),\\qquad\\langle\\bm\{\\eta\}\(t\)\\bm\{\\eta\}\(t^\{\\prime\}\)^\{\\top\}\\rangle=2D\\,\\delta\(t\-t^\{\\prime\}\)\\,\\bm\{I\}\.\(17\)Polar decomposition𝜺=rd\\bm\{\\varepsilon\}=r\\,d\(r\>0r\>0,d∈Sd−1d\\in S^\{d\-1\}\) yields, by Itô’s lemma,
r˙\\displaystyle\\dot\{r\}=μr−αr3\+\(d−1\)Dr\+ξr\(t\),⟨ξr2⟩=2D,\\displaystyle=\\mu r\-\\alpha r^\{3\}\+\\tfrac\{\(d\-1\)D\}\{r\}\+\\xi\_\{r\}\(t\),\\quad\\langle\\xi\_\{r\}^\{2\}\\rangle=2D,\(18\)d˙\\displaystyle\\dot\{d\}=−1r\(𝑰−dd⊤\)𝝃⟂\(t\),⟨ξ⟂ξ⟂⊤⟩=2D\(𝑰−dd⊤\)\.\\displaystyle=\-\\tfrac\{1\}\{r\}\\bigl\(\\bm\{I\}\-dd^\{\\top\}\\bigr\)\\bm\{\\xi\}\_\{\\perp\}\(t\),\\quad\\langle\\xi\_\{\\perp\}\\xi\_\{\\perp\}^\{\\top\}\\rangle=2D\\bigl\(\\bm\{I\}\-dd^\{\\top\}\\bigr\)\.\(19\)The radial dynamics equation[18](https://arxiv.org/html/2605.24057#A4.E18)are deterministic\-dominated post\-criticality:rrgrows from the initialr0=‖𝜺k\(0\)‖=σ∗D/μr\_\{0\}=\\\|\\bm\{\\varepsilon\}\_\{k\}\(0\)\\\|=\\sigma\_\{\*\}\\sqrt\{D/\\mu\}\(withσ∗\>1\\sigma\_\{\*\}\\\!\>\\\!1by assumption\) towardr⋆=μ/αr^\{\\star\}=\\sqrt\{\\mu/\\alpha\}\(attractor\) on timescaleτr∼\(1/μ\)log\(σ∗μ/\(αD\)\)\\tau\_\{r\}\\sim\(1/\\mu\)\\log\(\\sigma\_\{\*\}\\sqrt\{\\mu/\(\\alpha D\)\}\)\. The angular dynamics equation[19](https://arxiv.org/html/2605.24057#A4.E19)are pure noise driven, with effective diffusion coefficientDθ\(t\)=D/r\(t\)2D\_\{\\theta\}\(t\)=D/r\(t\)^\{2\}\.
#### Angular drift integral \(finite\-TTscope\)\.
The naive infinite\-time integral∫0∞2\(d−1\)D/r\(t\)2𝑑t\\int\_\{0\}^\{\\infty\}2\(d\{\-\}1\)D/r\(t\)^\{2\}\\,dtdiverges in the saturation phaset\>τrt\>\\tau\_\{r\}wherer\(t\)≈r⋆r\(t\)\\approx r^\{\\star\}is constant; angular Brownian motion on the saturated attractor randomizes the direction in the strictt→∞t\\to\\inftylimit\. Prop\.[2](https://arxiv.org/html/2605.24057#Thmproposition2)therefore applies on a*finite*window\[0,T\]\[0,T\]; the accumulated angular variance splits into growth and saturation contributions
Θ2\(T\)=∫0τr2\(d−1\)Dr02e2μt𝑑t⏟growth,∼\(d−1\)/σ∗2\+2\(d−1\)D\(T−τr\)r⋆2⏟saturation,\(T−τr\)/Trand,\\Theta^\{2\}\(T\)\\;=\\;\\underbrace\{\\int\_\{0\}^\{\\tau\_\{r\}\}\\\!\\\!\\tfrac\{2\(d\-1\)D\}\{r\_\{0\}^\{2\}e^\{2\\mu t\}\}\\,dt\}\_\{\\text\{growth, \}\\sim\(d\-1\)/\\sigma\_\{\*\}^\{2\}\}\\;\+\\;\\underbrace\{\\tfrac\{2\(d\-1\)D\\,\(T\-\\tau\_\{r\}\)\}\{r^\{\\star 2\}\}\}\_\{\\text\{saturation, \}\(T\-\\tau\_\{r\}\)/T\_\{\\mathrm\{rand\}\}\},\(20\)withTrand:=r⋆2/\(2\(d−1\)D\)T\_\{\\mathrm\{rand\}\}:=r^\{\\star 2\}/\(2\(d\{\-\}1\)D\)the spherical randomization timescale\. The growth contribution evaluates to\(d−1\)D/\(μr02\)\(1−e−2μτr\)≈\(d−1\)/σ∗2\(d\{\-\}1\)D/\(\\mu r\_\{0\}^\{2\}\)\\,\(1\-e^\{\-2\\mu\\tau\_\{r\}\}\)\\approx\(d\{\-\}1\)/\\sigma\_\{\*\}^\{2\}\. The saturation contribution is small as long asT−τr≪TrandT\-\\tau\_\{r\}\\ll T\_\{\\mathrm\{rand\}\}; in this regime
Θ2\(T\)≈d−1σ∗2\+O\(T/Trand\)\.\\Theta^\{2\}\(T\)\\;\\approx\\;\\tfrac\{d\-1\}\{\\sigma\_\{\*\}^\{2\}\}\\;\+\\;O\\bigl\(T/T\_\{\\mathrm\{rand\}\}\\bigr\)\.\(21\)*Theσ∗2\\sigma\_\{\*\}^\{2\}in the denominator is critical\.*When the initial perturbation is well above the noise floor \(σ∗≫1\\sigma\_\{\*\}\\gg 1\) andTTis within the persistence window,Θ2\(T\)≪1\\Theta^\{2\}\(T\)\\ll 1and the direction is preserved with negligible spread\. Whenσ∗=1\\sigma\_\{\*\}=1exactly,Θ2∼\(d−1\)\\Theta^\{2\}\\sim\(d\-1\)and direction\-preservation is*not guaranteed*from this argument alone; this is the regime that requires the full hypothesis boundσ∗\>1\\sigma\_\{\*\}\>1in Prop\.[2](https://arxiv.org/html/2605.24057#Thmproposition2)\. For empirically relevant SAE / encoder training,Trand≫TtrainT\_\{\\mathrm\{rand\}\}\\gg T\_\{\\mathrm\{train\}\}by orders of magnitude \(the saturatedr⋆r^\{\\star\}is large relative to noise\), placing all observed checkpoints inside the persistence window; theT→∞T\\to\\inftyrandomization regime is unphysical for finite\-horizon training\.
#### From angular spread to positive expected cosine\.
For Gaussian\-distributed angular displacement with varianceΘ2\(T\)\\Theta^\{2\}\(T\)onSd−1S^\{d\-1\}\(von Mises–Fisher\-like concentration arounddk\(0\)d\_\{k\}\(0\)\),
𝔼\[dk\(0\)⊤dk\(T\)\]=𝔼\[cosθk\(T\)\]≈1−Θ2\(T\)2\+O\(Θ4\)=1−d−12σ∗2\+O\(σ∗−4,T/Trand\)\.\\mathbb\{E\}\[d\_\{k\}\(0\)^\{\\top\}d\_\{k\}\(T\)\]\\;=\\;\\mathbb\{E\}\[\\cos\\theta\_\{k\}\(T\)\]\\;\\approx\\;1\-\\tfrac\{\\Theta^\{2\}\(T\)\}\{2\}\+O\(\\Theta^\{4\}\)\\;=\\;1\-\\tfrac\{d\-1\}\{2\\sigma\_\{\*\}^\{2\}\}\+O\\bigl\(\\sigma\_\{\*\}^\{\-4\},T/T\_\{\\mathrm\{rand\}\}\\bigr\)\.\(22\)This is strictly positive whenσ∗2\>\(d−1\)/2\\sigma\_\{\*\}^\{2\}\>\(d\{\-\}1\)/2andT≪TrandT\\ll T\_\{\\mathrm\{rand\}\}, recovering equation[9](https://arxiv.org/html/2605.24057#S5.E9)\.
#### Empirical proxies \(see also Sec\.[5\.5](https://arxiv.org/html/2605.24057#S5.SS5)\)\.
The two proxies discussed in the main text inherit the positivity of equation[9](https://arxiv.org/html/2605.24057#S5.E9)via standard concentration arguments on the sphere: \(i\) for any fixed referencer∈Sd−1r\\in S^\{d\-1\}, the scalar pair\(⟨𝜺k\(0\),r⟩,⟨𝜺k\(T\),r⟩\)\(\\langle\\bm\{\\varepsilon\}\_\{k\}\(0\),r\\rangle,\\langle\\bm\{\\varepsilon\}\_\{k\}\(T\),r\\rangle\)has positive Pearson correlation in expectation, and acrossKKi\.i\.d\. atoms its Spearman correlation converges to a positive population value; \(ii\) projections onto pre\-specified linguistic POS subspaces inℝN\\mathbb\{R\}^\{N\}similarly inherit per\-atom autocorrelation, empirically realized byρid\\rho\_\{\\mathrm\{id\}\}in Sec\.[5\.1](https://arxiv.org/html/2605.24057#S5.SS1)\.
### D\.2Weak inter\-mode coupling \(γ\>0\\gamma\>0\)
Forγ\>0\\gamma\>0, the coupling term in equation[8](https://arxiv.org/html/2605.24057#S5.E8)adds−γ∑j≠k\(𝜺j⊤𝜺k\)𝜺j\-\\gamma\\sum\_\{j\\neq k\}\(\\bm\{\\varepsilon\}\_\{j\}^\{\\top\}\\bm\{\\varepsilon\}\_\{k\}\)\\bm\{\\varepsilon\}\_\{j\}to modekk’s drift\. In the weak\-coupling regimeγ≪α\\gamma\\ll\\alpha, the cubic self\-saturation dominates and modes self\-orthogonalize: any inner product𝜺j⊤𝜺k\\bm\{\\varepsilon\}\_\{j\}^\{\\top\}\\bm\{\\varepsilon\}\_\{k\}between non\-aligned modes is suppressed by the saturation\. A standard perturbative argument shows the coupling correction toρSpear\\rho\_\{\\mathrm\{Spear\}\}isO\(γ\)O\(\\gamma\), which preserves the sign ofρ\>0\\rho\>0at leading order\.□\\square
### D\.3Numerical verification on the coupled\-mode SDE
We verify Prop\.[2](https://arxiv.org/html/2605.24057#Thmproposition2)directly by numerical simulation of equation[8](https://arxiv.org/html/2605.24057#S5.E8), using the*scalar\-projection proxy*introduced in the main text \(per\-atom autocorrelation of projection onto a fixed reference directionrr\)\. Parameters:K=200K\\\!=\\\!200modes, spatial dimensiond=10d\\\!=\\\!10,μ=0\.10\\mu\\\!=\\\!0\.10,α=0\.10\\alpha\\\!=\\\!0\.10,γ=10−3\\gamma\\\!=\\\!10^\{\-3\}\(soγK=0\.2≪α\\gamma K\\\!=\\\!0\.2\\ll\\alpha, weak\-coupling regime\),D=10−5D\\\!=\\\!10^\{\-5\},dt=0\.05dt\\\!=\\\!0\.05,Tmax=2,000T\_\{\\max\}\\\!=\\\!2\{,\}000steps; initial perturbations𝜺k\(0\)\\bm\{\\varepsilon\}\_\{k\}\(0\)drawn i\.i\.d\. from𝒩\(𝟎,σ02𝑰\)\\mathcal\{N\}\(\\bm\{0\},\\,\\sigma\_\{0\}^\{2\}\\bm\{I\}\)withσ0=0\.05\\sigma\_\{0\}\\\!=\\\!0\.05, soσ∗=σ0/D/μ≈5\\sigma\_\{\*\}=\\sigma\_\{0\}/\\sqrt\{D/\\mu\}\\approx 5\(initial perturbation5×5\\timesabove the noise floor; well within the regimeσ∗2\>\(d−1\)/2=4\.5\\sigma\_\{\*\}^\{2\}\>\(d\-1\)/2=4\.5for which Prop\.[2](https://arxiv.org/html/2605.24057#Thmproposition2)predicts positive directional persistence\)\. For a fixed random reference directionr∈Sd−1r\\in S^\{d\-1\}, we compute the Spearman rank correlation between the per\-atom scalar projections⟨𝜺k\(0\),r⟩\\langle\\bm\{\\varepsilon\}\_\{k\}\(0\),r\\rangleand⟨𝜺k\(T\),r⟩\\langle\\bm\{\\varepsilon\}\_\{k\}\(T\),r\\rangleat the simulation endpointTT\(within the persistence window\) acrossk=1,…,Kk=1,\\dots,K; this Spearman estimates the population scalar autocorrelation predicted by Prop\.[2](https://arxiv.org/html/2605.24057#Thmproposition2)\.
Figure 8:Direct numerical verification of Prop\.[2](https://arxiv.org/html/2605.24057#Thmproposition2)\.\(A\)Scatter of initial vs final projections onto a fixed reference directionrr\(one representative seed\); Spearmanρ=\+0\.95\\rho=\+0\.95,p<10−100p\\\!<\\\!10^\{\-100\}\.\(B\)ρ\\rhotrajectory over simulation time, 5 random seeds\. The correlation saturates nearρ≈1\\rho\\approx 1early and remains positive; convergedρ=\+0\.948±0\.006\\rho=\+0\.948\\pm 0\.006across seeds\.#### Result \(5 seeds\)\.
Across 5 independent runs, the converged Spearman correlation isρ=\+0\.948±0\.006\\rho=\+0\.948\\pm 0\.006\(individual seeds:\+0\.950\+0\.950,\+0\.936\+0\.936,\+0\.951\+0\.951,\+0\.954\+0\.954,\+0\.949\+0\.949\)\. All seeds giveρ\>0\.93\\rho\>0\.93withp<10−90p<10^\{\-90\}\. This is a direct verification of equation[9](https://arxiv.org/html/2605.24057#S5.E9)\.
#### Comparison with SAE empirics\.
The toy givesρ≈0\.95\\rho\\approx 0\.95whereas the SAE setting \(Sec\.[5\.1](https://arxiv.org/html/2605.24057#S5.SS1)\) givesρid=0\.41±0\.03\\rho\_\{\\mathrm\{id\}\}=0\.41\\pm 0\.03\. The gap reflects the additional confounds present in the SAE setting that are absent in the toy: \(i\) POS purity is a coarse 15\-way metric that does not capture all geometric structure of the unstable manifold, \(ii\) the empirical “initial direction” is measured at step1,0001\{,\}000\(post\-onset, after some non\-trivial coupling\-induced rotation\) rather than att=0t=0, and \(iii\) the SAE dynamics include secondary bifurcations \(Sec\.[5\.6](https://arxiv.org/html/2605.24057#S5.SS6)\) and tokenization noise\. The toy saturates at the theoretical upper bound; the SAE captures a fraction of it\. The qualitative predictionρ\>0\\rho\>0\(the content of Prop\.[2](https://arxiv.org/html/2605.24057#Thmproposition2)\) is verified in both settings\.
## Appendix EToy and MNIST validation \(exp 00–08\)
The toy experiments verify the theoretical claims of Sec\.[3](https://arxiv.org/html/2605.24057#S3)on synthetic data where every relevant quantity is known analytically\. The MNIST experiments are a first transfer test to real data\. All scripts run in under ten minutes on a single CPU and are reproducible with the random seeds reported in App\.[M](https://arxiv.org/html/2605.24057#A13)\.
#### C\.1 Bimodal split \(exp 00\) and unimodal control \(exp 01\)\.
On 2D bimodal Gaussian data withK=8K\\\!=\\\!8prototypes and a single shared learnedβ\\beta, the symmetric state𝒮0\\mathcal\{S\}\_\{0\}loses stability exactly whenβλmax\(Σ\)=1\\beta\\lambda\_\{\\max\}\(\\Sigma\)\\\!=\\\!1\(Fig\.[10](https://arxiv.org/html/2605.24057#A5.F10)\)\. The split direction locks to the data’s principal axis\. Replacing the data with an isotropic unimodal Gaussian \(Fig\.[10](https://arxiv.org/html/2605.24057#A5.F10)\) gives a∼27×\\sim 27\\timessmaller order\-parameter gap at the sameβ/βc\\beta\\\!/\\\!\\beta\_\{c\}and a uniform random split direction across seeds, confirming that the direction is data\-driven and the magnitude is set by the spectral gap ofΣ\\Sigma\.
Figure 9:Exp 00: bimodal split\. Top:logβ\(t\)\\log\\beta\(t\)and order parameter trajectory; bottom: final prototypes locked to the data’s principal axis\. The order parameter activates asβ\(t\)\\beta\(t\)crossesβc=1/λmax\(Σ\)\\beta\_\{c\}\\\!=\\\!1/\\lambda\_\{\\max\}\(\\Sigma\)\.
Figure 10:Exp 01: unimodal control\. The order parameter gap is∼27×\\sim 27\\timessmaller than the bimodal case at matchedβ/βc\\beta/\\beta\_\{c\}, and the split angle is uniform across seeds\. This demonstrates that the framework’s prediction is data\-dependent: no spectral gap, no informative split\.
#### C\.2 Hessian calibration \(exp 02\)\.
Direct numerical diagonalization of the full Hessian at a range ofβ\\betavalues locates the lowest\-eigenvalue zero crossing to four decimals; it agrees with the analyticalβc=1/λmax\(Σ\)\\beta\_\{c\}\\\!=\\\!1/\\lambda\_\{\\max\}\(\\Sigma\)derived in Sec\.[3](https://arxiv.org/html/2605.24057#S3)\(Fig\.[11](https://arxiv.org/html/2605.24057#A5.F11)\)\. Replacing the Rose\-Gurewitz\-Fox conventionTc=2λmaxT\_\{c\}\\\!=\\\!2\\lambda\_\{\\max\}with ourβ=2/T\\beta\\\!=\\\!2/Tconvention gives the same numerical critical point, resolving the factor\-of\-2 ambiguity noted in App\.[A\.3](https://arxiv.org/html/2605.24057#A1.SS3)\.
Figure 11:Exp 02: numerical Hessian calibration\. The lowest eigenvalue of the full Hessian \(top\) zero\-crosses atβ=βc\\beta\\\!=\\\!\\beta\_\{c\}\(vertical dashed line\) computed analytically from1/λmax\(Σ\)1/\\lambda\_\{\\max\}\(\\Sigma\)\. Bottom: trajectory order parameter activates at the sameβ\\beta\. Match to four decimals across seeds\.
#### C\.3 Hierarchical bifurcation \(exp 03\)\.
WithK=8K\\\!=\\\!8prototypes and 2\-level hierarchical data \(four super\-clusters of two sub\-clusters each\), the order\-parameter trajectory shows two discrete jumps, the first atβc\(1\)=1/λmax\(Σ\)\\beta\_\{c\}^\{\(1\)\}\\\!=\\\!1/\\lambda\_\{\\max\}\(\\Sigma\)and the second atβc\(2\)=1/λmax\(Σwithin\)\\beta\_\{c\}^\{\(2\)\}\\\!=\\\!1/\\lambda\_\{\\max\}\(\\Sigma\_\{\\mathrm\{within\}\}\)\(Fig\.[12](https://arxiv.org/html/2605.24057#A5.F12)\)\. The final state is the predicted2\+2\+2\+22\{\+\}2\{\+\}2\{\+\}2tessellation\. This validates the within\-supercluster recursion of Sec\.[3](https://arxiv.org/html/2605.24057#S3)\(hierarchy paragraph\) on data with clean nested structure\.
Figure 12:Exp 03: hierarchical bifurcation\. The first and second order parameters activate sequentially at the analytically predictedβc\(1\)\\beta\_\{c\}^\{\(1\)\}andβc\(2\)\\beta\_\{c\}^\{\(2\)\}\. The final prototype configuration is the2\+2\+2\+22\{\+\}2\{\+\}2\{\+\}2tessellation predicted by recursive application of Sec\.[3](https://arxiv.org/html/2605.24057#S3)\.
#### C\.4 Reverse traversal as merge \(exp 04\)\.
Reducingβ\\betafrom supercritical to subcritical traces the same equilibrium branch as forward training \(Fig\.[13](https://arxiv.org/html/2605.24057#A5.F13)\)\. Forward and reverse OP\-vs\-β\\betacurves overlap; the reverseβ⋆/βc\\beta^\{\\star\}/\\beta\_\{c\}matches theory to≤4%\\leq 4\\%, versus a∼30%\\sim 30\\%forward overshoot driven by exponential growth from noise\. This is the toy\-data confirmation of Sec\.[3](https://arxiv.org/html/2605.24057#S3)’s claim that split and merge are the same pitchfork traversed in opposite directions ofβ\\beta\.
Figure 13:Exp 04: split and merge on the same pitchfork\. Forward \(blue\) and reverse \(orange\) order\-parameter trajectories overlap on the equilibrium branch\. Forward overshootsβc\\beta\_\{c\}by∼30%\\sim 30\\%due to noise; reverse tracksβc\\beta\_\{c\}to≤4%\\leq 4\\%\.
#### C\.5 Endogenous critical point \(exp 05\)\.
Replacing the fixed dataset with the latent of a co\-evolving autoencoder makesβc\(t\)\\beta\_\{c\}\(t\)itself depend on training state\.β\(t\)\\beta\(t\)catchesβc\(t\)\\beta\_\{c\}\(t\)at step∼300\\sim 300and the GMM bifurcates shortly after \(Fig\.[14](https://arxiv.org/html/2605.24057#A5.F14)\)\. This validates Proposition[1](https://arxiv.org/html/2605.24057#Thmproposition1)on a closed toy system where bothβ\(t\)\\beta\(t\)andβc\(t\)\\beta\_\{c\}\(t\)are observable\.
Figure 14:Exp 05: self\-induced critical point\.β\(t\)\\beta\(t\)\(blue\) catchesβc\(t\)\\beta\_\{c\}\(t\)\(red\) at step∼300\\sim 300; the order parameter \(green\) activates shortly after\. Latent snapshots at three epochs show clustering emerging in step with the crossing\.
#### C\.6 MNIST \+ SimCLR \(exp 07–08\)\.
Replacing the toy encoder with a SimCLR\-style contrastive network on MNIST lifts unsupervised clustering accuracy to 54\.3% on the standard 10\-class labels \(Fig\.[15](https://arxiv.org/html/2605.24057#A5.F15)\)\.β\\betacrossesβc\\beta\_\{c\}at step∼360\\sim 360, prototypes pitchfork into class\-aligned regions, and visually similar digits cluster \(1\-styles, 0\-styles, 6\-styles separated; 3/5/8/9 form a confusable supercluster\)\. With longer training and proper unsupervised metrics, four of the five top eigenvalues ofCov\(z\)\\mathrm\{Cov\}\(z\)activate sequentially, indicating multi\-axis hierarchical bifurcation in 5D\. We omit the MNIST autoencoder result \(exp 06\) from the main appendix because the AE’s representation bottleneck produces a degenerate prototype set \(∼26%\\sim 26\\%accuracy with two prototypes capturing80%80\\%of points\); the bifurcation occurs but the upstream encoder is too weak to produce informative concepts\. The contrast confirms the framework’s prediction that bifurcation is necessary but not sufficient: the upstream encoder must produce a non\-trivialCov\(z\)\\mathrm\{Cov\}\(z\)for the bifurcation to expose meaningful structure\.
Figure 15:Exp 08: MNIST \+ SimCLR long training\. All five top eigenvalues ofCov\(z\)\\mathrm\{Cov\}\(z\)activate sequentially as training proceeds; the confusable\-class submanifold \(3/5/8/9\) emerges as a coherent sub\-pitchfork\.
## Appendix FReverse traversal on CIFAR \(exp 11–12\)
The merge\-as\-reverse\-pitchfork prediction validated on toy data in App\.[E](https://arxiv.org/html/2605.24057#A5)\(exp 04\) is non\-trivial to test on a real encoder becauseβ\\betais normally an internal learned parameter that grows monotonically\. We adapt the test by training a CIFAR\-10 SimCLR encoder \+ GMM head to convergence and then*externally annealing*β\\betafrom its post\-training supercritical value back down throughβc\\beta\_\{c\}\(Fig\.[16](https://arxiv.org/html/2605.24057#A6.F16)\)\. Two modes:
- •*Contrastive driver on\.*The SimCLR loss continues to act on the encoder whileβ\\betais annealed down\. The encoder’sCov\(z\)\\mathrm\{Cov\}\(z\)continues to spread despite the GMM no longer rewarding clustering\. NC1 stays low \(∼0\.5\\sim 0\.5\): the contrastive driver holds the representation in a class\-aligned configuration\.
- •*Contrastive driver off\.*The SimCLR loss is detached; onlyβ\\betamoves\. The GMM order parameter follows the reverse branch of the pitchfork to zero and NC1 returns to its high \(∼1\.6\\sim 1\.6\) initial value\. The trajectory in the OP\-versus\-β\\betaplane overlaps the forward trajectory\.
The combined picture \(three regimes on a single CIFAR\-10 encoder\) appears in Fig\.[16](https://arxiv.org/html/2605.24057#A6.F16)\. The reverse\-anneal agreement with theory is bounded by the encoder’s residual drift \(≤7%\\leq 7\\%\), not by hysteresis: as on toy data \(App\.[E](https://arxiv.org/html/2605.24057#A5)exp 04\), split and merge are the same pitchfork\.
Figure 16:Exp 11–12: CIFAR\-10 reverse traversal\. Three regimes on a single SimCLR \+ ResNet\-18 encoder\.*Forward training:*β\(t\)\\beta\(t\)climbs pastβc\\beta\_\{c\}; OP and class\-aligned structure emerge\.*Reverse anneal with driver on:*contrastive loss holds the representation; NC1 stays low\.*Reverse anneal with driver off:*OP returns to zero, NC1 returns to its initial value, recapitulating the forward branch in reverse\.
## Appendix GLM probe behavior: anisotropy and pre\-supercriticality \(exp 13–15\)
Section[4](https://arxiv.org/html/2605.24057#S4)reported that the SAE on frozen Pythia is the cleanest realization of the full V becauseβc\\beta\_\{c\}is held fixed by the frozen LM\. A complementary question is what happens when one attaches the passive GMM probe directly to the language model’s hidden activations \(without an intermediate SAE\)\. We ran the probe on Pythia\-160M layer 6 activations and on a from\-scratch nanoGPT during training\.
#### LM activations are pre\-supercritical\.
Pythia\-160M layer 6 hidden states have anisotropyλmax\(Cov\(z\)\)/σ¯2≈380\\lambda\_\{\\max\}\(\\mathrm\{Cov\}\(z\)\)/\\bar\{\\sigma\}^\{2\}\\\!\\approx\\\!380\(vs\. typical ResNet\-on\-CIFAR values of 5–20\)\. With the probe’s defaultlogβ0=−2\.5\\log\\beta\_\{0\}\\\!=\\\!\-2\.5, this putslog\(β/βc\)≈\+2\.6\\log\(\\beta/\\beta\_\{c\}\)\\\!\\approx\\\!\+2\.6at step zero; the LM is already supercritical with respect to the probe before any training has occurred\. The probe’sβ\\betagrows modestly during training and saturates aroundlogβ≈0\\log\\beta\\\!\\approx\\\!0; the trajectory therefore lives entirely in the post\-critical regime, with the pre\-critical leg occluded by the supercritical starting point\. Noβ=βc\\beta\\\!=\\\!\\beta\_\{c\}crossing event is observable on raw LM activations\.
#### The anisotropy is structural, not a probe artifact\.
A K\-sweep on a from\-scratch nanoGPT confirms that the high\-anisotropy regime is not an artifact of the GMM probe’s choice ofKK\. WithK=1K\\\!=\\\!1\(a single “prototype” that just measures variance\) we recoverlog\(β/βc\)≈\+2\\log\(\\beta/\\beta\_\{c\}\)\\\!\\approx\\\!\+2–33at initialization, matching the value seen withK=10K\\\!=\\\!10\. The highλmax\\lambda\_\{\\max\}is inherited from the Zipfian token\-embedding spectrum combined with the random transformer mixing matrix at initialization; it is a structural property of the LM, not a probe choice\.
#### Implication for the framework\.
For analyses targeting the bifurcation*event*on LM activations, one should not attach the GMM probe directly to LM hidden states\. The SAE\-on\-frozen\-LM setup of Sec\.[4\.1](https://arxiv.org/html/2605.24057#S4.SS1)is the workaround: the SAE itself starts from random initialization \(lowβSAE\\beta\_\{\\mathrm\{SAE\}\}\), is in the pre\-critical regime, and traverses the full bifurcation arc during its own training\.
## Appendix HPhase C \(recovery\) traces from the mid\-training intervention study
The mid\-training intervention experiment \(Sec\.[6\.2](https://arxiv.org/html/2605.24057#S6.SS2)\) includes a Phase C in which the healthy configuration is restored after the Phase B perturbation\. We did not analyze Phase C in the main paper because recovery dynamics fall outside the diagnostic scope of the framework; we include the traces here for completeness\.
#### Catastrophic modes \(*no\_centering*,*no\_ema*\) do not recover within 8 epochs\.
Once these modes have driven NC1 to∼102\\\!\\sim\\\!10^\{2\}and broken the prototype assignment, restoring healthy centering and EMA does not return the encoder to the healthy trajectory within the Phase C window\. cluster\_acc stays below the pre\-intervention healthy baseline for the remaining 8 epochs; some seeds slowly recover, others appear permanently degraded\. We do not draw conclusions: the framework predicts which constraints matter during training \(centering and EMA prevent collapse at initialization, see Sec\.[6](https://arxiv.org/html/2605.24057#S6)\), not whether a broken encoder is reachable from a healthy basin by re\-applying those constraints\.
#### Gradual modes \(*no\_sharpening*,*tiny\_batch*\) partially recover\.
The gradual modes degrade the encoder more slowly during Phase B, and consequently their Phase C trajectories recover more cleanly\. NC1 returns to within∼30%\\sim 30\\%of the healthy trajectory by epoch∼5\\sim 5of Phase C;log\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)similarly returns to the healthy band\. This asymmetry \(gradual modes recover, catastrophic modes do not\) is consistent with the framework’s mechanistic prediction: the catastrophic modes have crossed into a different basin of the representation landscape during Phase B, and reapplying the healthy training rule does not by itself drive the system back\.
## Appendix IWhy exactly four shapes: a 3\-axis taxonomy
The four trajectory shapes catalogued in Sec\.[4](https://arxiv.org/html/2605.24057#S4)– Sec\.[4\.2](https://arxiv.org/html/2605.24057#S4.SS2)are not an arbitrary count: they emerge by enumerating three binary kinematic axes of the encoder–probe race and noting which combinations are empirically distinct and reachable\. We formalize this here\.
###### Proposition 3\(Classification of trajectory shapes\)\.
Let𝒯\\mathcal\{T\}denote a trajectory in\(log\(β\(t\)/βc\(t\)\),logNC1\(t\)\)\(\\log\(\\beta\(t\)/\\beta\_\{c\}\(t\)\),\\,\\log\\mathrm\{NC1\}\(t\)\)generated by the dynamics of Sec\.[3](https://arxiv.org/html/2605.24057#S3)\. Define three binary*kinematic axes*:
A1A\_\{1\}\(initial criticality\):A1=subA\_\{1\}=\\mathrm\{sub\}iflog\(β\(0\)/βc\(0\)\)<0\\log\(\\beta\(0\)/\\beta\_\{c\}\(0\)\)<0,A1=superA\_\{1\}=\\mathrm\{super\}if\>0\>0\.
A2A\_\{2\}\(post\-onset rate ordering\):A2=β\-leadsA\_\{2\}=\\beta\\text\{\-leads\}ifβ˙\(t\)/β\(t\)\>−β˙c\(t\)/βc\(t\)\\dot\{\\beta\}\(t\)/\\beta\(t\)\>\-\\dot\{\\beta\}\_\{c\}\(t\)/\\beta\_\{c\}\(t\)post\-onset on a positive\-measure set;A2=βc\-leadsA\_\{2\}=\\beta\_\{c\}\\text\{\-leads\}otherwise\.
A3A\_\{3\}\(dissipation regime\):A3=normalA\_\{3\}=\\mathrm\{normal\}if the post\-critical escape timeτesc\\tau\_\{\\mathrm\{esc\}\}\(Remark[1](https://arxiv.org/html/2605.24057#Thmremark1), App\.[C](https://arxiv.org/html/2605.24057#A3)\) isO\(1/\(β−βc\)\)O\(1/\(\\beta\-\\beta\_\{c\}\)\)at the crossing \(the broken\-symmetry transition follows the crossing within a few characteristic timescales\);A3=lowA\_\{3\}=\\mathrm\{low\}otherwise\.
Augment with a degenerate axisA0=clusteringA\_\{0\}=\\mathrm\{clustering\}if the upstream objective induces a non\-trivialCov\(z\(t\)\)\\mathrm\{Cov\}\(z\(t\)\)structure versusA0=noneA\_\{0\}=\\mathrm\{none\}otherwise \(negative control\)\.
Then exactly four equivalence classes of trajectory shape are realizable under feature\-learning pipelines:
1. \(i\)Full V:A0=cl\.,A1=sub,A2=β\-leads,A3=normalA\_\{0\}=\\mathrm\{cl\.\},A\_\{1\}=\\mathrm\{sub\},A\_\{2\}=\\beta\\text\{\-leads\},A\_\{3\}=\\mathrm\{normal\}\(SAE on frozen Pythia L6\)\.
2. \(ii\)Fold\-back \(spectrum\):A0=cl\.,A2=βc\-leads,A3=normalA\_\{0\}=\\mathrm\{cl\.\},A\_\{2\}=\\beta\_\{c\}\\text\{\-leads\},A\_\{3\}=\\mathrm\{normal\}\(A1A\_\{1\}unconstrained; magnitude of fold scales with\|β˙c/β˙\|\|\\dot\{\\beta\}\_\{c\}/\\dot\{\\beta\}\|\) \(DINO/SimCLR on CIFAR\-10/100\)\.
3. \(iii\)Delayed escape:A0=cl\.,A1=sub,A2=β\-leads,A3=lowA\_\{0\}=\\mathrm\{cl\.\},A\_\{1\}=\\mathrm\{sub\},A\_\{2\}=\\beta\\text\{\-leads\},A\_\{3\}=\\mathrm\{low\}\(grokking on modular arithmetic\)\.
4. \(iv\)No arc \(control\):A0=noneA\_\{0\}=\\mathrm\{none\}\(A1,A2,A3A\_\{1\},A\_\{2\},A\_\{3\}undefined; rotation\-prediction control\)\.
The remaining nominal combinations\(23−4=4\(2^\{3\}\-4=4in the clustering branch, plus the trivialA0=noneA\_\{0\}=\\mathrm\{none\}branch\) are either degenerate \(collapse into one of \(i\)–\(iv\)\) or unreachable under standard feature\-learning training protocols\.
###### Proof\.\.
The four observed shapes are distinguishable by sign of the post\-critical slopes:=∂logNC1/∂log\(β/βc\)s:=\\partial\\log\\mathrm\{NC1\}/\\partial\\log\(\\beta/\\beta\_\{c\}\)restricted to the descent leg, combined with the ratioτesc/T\\tau\_\{\\mathrm\{esc\}\}/TwhereTTis the training horizon:
- •*Full V*:s<0s<0,τesc/T≪1\\tau\_\{\\mathrm\{esc\}\}/T\\ll 1, descent over the fulllog\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)range\.
- •*Fold\-back*:s\>0s\>0,τesc/T≪1\\tau\_\{\\mathrm\{esc\}\}/T\\ll 1\.
- •*Delayed escape*:τesc/T=O\(1\)\\tau\_\{\\mathrm\{esc\}\}/T=O\(1\)\.
- •*No arc*:\|s\|→0\|s\|\\to 0andlogNC1\\log\\mathrm\{NC1\}decouples fromlog\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)\.
The map\(A0,A1,A2,A3\)↦\(s\-sign,τesc/T\)\(A\_\{0\},A\_\{1\},A\_\{2\},A\_\{3\}\)\\mapsto\(s\\text\{\-sign\},\\tau\_\{\\mathrm\{esc\}\}/T\)is well\-defined under the assumptions of Prop\.[1](https://arxiv.org/html/2605.24057#Thmproposition1)\(for existence of crossing\) and Remark[1](https://arxiv.org/html/2605.24057#Thmremark1)\(forτesc\\tau\_\{\\mathrm\{esc\}\}’s qualitative dependence on dissipation\)\. Direct case analysis \(Tab\.[6](https://arxiv.org/html/2605.24057#A9.T6)\) exhausts the232^\{3\}combinations: caseA2=β\-leads,A3=normalA\_\{2\}\\\!=\\\!\\beta\\text\{\-leads\},A\_\{3\}\\\!=\\\!\\mathrm\{normal\}collapses into \(i\) Full V or its mild\-fold variant of \(ii\) depending onA1A\_\{1\}; caseA2=βc\-leads,A3=lowA\_\{2\}\\\!=\\\!\\beta\_\{c\}\\text\{\-leads\},A\_\{3\}\\\!=\\\!\\mathrm\{low\}is not realized by any feature\-learning protocol we tested\.□\\square∎
The proof above establishes the*forward*direction \(these axes generate at most four shapes\)\. The*converse*; that no fifth shape can arise from the same dynamics; requires the assumption that the only sources of phase\-space partitioning are the ones captured by\(A0,A1,A2,A3\)\(A\_\{0\},A\_\{1\},A\_\{2\},A\_\{3\}\); this is a modeling assumption, not a theorem\. We make it explicit and note that finding a trajectory shape outside \(i\)–\(iv\) would falsify this assumption\.
#### The three axes \(extended discussion\)\.
A trajectory in\(log\(β/βc\),logNC1\)\(\\log\(\\beta/\\beta\_\{c\}\),\\,\\log\\mathrm\{NC1\}\)is shaped by:
1. 1\.Initial sub/supercriticality\.Islog\(β\(0\)/βc\(0\)\)\\log\(\\beta\(0\)/\\beta\_\{c\}\(0\)\)negative or positive att=0t\\\!=\\\!0? Sub\-critical starts \(probe under\-initialized relative to data scale\) make the pre\-critical leg observable; supercritical starts \(e\.g\., high\-anisotropy encoders such as Pythia raw activations or ResNet\-on\-CIFAR\-10\) hide it\.
2. 2\.Post\-onset kinematics\.Onceβ\\betahas crossedβc\\beta\_\{c\}, doesβ\(t\)\\beta\(t\)continue to grow faster thanβc\(t\)\\beta\_\{c\}\(t\), or doesβc\(t\)\\beta\_\{c\}\(t\)overtakeβ\(t\)\\beta\(t\)? The first case produces a monotone post\-critical descent inlog\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)\(frozen\-LM SAEs, CIFAR\-10 SSL\); the second produces a*fold\-back*where the trajectory reverses direction inlog\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)whilelogNC1\\log\\mathrm\{NC1\}continues to fall \(CIFAR\-100 SSL, where the encoder keeps spreading features into 100 classes\)\.
3. 3\.Dissipation rate\.Remark[1](https://arxiv.org/html/2605.24057#Thmremark1)identifies the post\-critical escape timescale as dissipation\-controlled\. Normal dissipation \(continuously\-driven contrastive losses, SAEs with renormalization, supervised CE with WD\) gives the macroscopic broken\-symmetry transition immediately after the crossing; low dissipation \(high\-precision optimization without strong WD\) traps the system on the saddle for an extended*metastable plateau*\.
#### Enumeration\.
The three binary axes give eight nominal combinations\. Four are empirically distinct and observed in our experiments; the remainder are degenerate or unreachable with the methods we tested:
Table 6:Eight nominal combinations of the three kinematic axes; four are empirically distinct and verified \(above the rule\), three shown below are degenerate or unobserved\. The eighth combination \(super\-init \+βc\\beta\_\{c\}\-overtakes \+ normal\) is empirically subsumed into the fold\-back regime\.InitPost\-onsetβ\\betavsβc\\beta\_\{c\}DissipationObserved shapeEmpirical witnesssubβ\\betarises monotone \(frozenβc\\beta\_\{c\}\)normal\(i\) Full VSAE on frozen Pythia L6*any*βc\\beta\_\{c\}overtakesβ\\betanormal\(ii\) Fold\-backDINO/SimCLR on CIFAR\-10 \(mild,∼\\sim0\.5\)\+ CIFAR\-100 \(strong,∼\\sim3\)subβ\\betarises monotonelow\(iii\) Delayed escapeGrokking on modular arithmetic*any*\(no clustering pressure\)—\(iv\) No arc \(control\)Rotation\-prediction controlsuperβ\\betarises monotonenormal—“descent only”; pre\-critical leg occluded, post\-onset shape determined by fold magnitude \(mild fold→\\tovisually descent\-dominated, as in CIFAR\-10 above\)supermonotonelow—degenerate \(no metastable saddle to trap, system already broken\)subβc\\beta\_\{c\}overtakesβ\\betalow—not observed; would require grokking\-style architecture with CIFAR\-100\-styleβc\(t\)\\beta\_\{c\}\(t\)rise
#### Why not fewer than four\.
\(iii\) requires its own line because the dissipation\-controlled metastable plateau of Remark[1](https://arxiv.org/html/2605.24057#Thmremark1)produces a qualitatively distinct training\-time trajectory; the crossing and the macroscopic transition are separated by orders of magnitude in steps; that is invisible in any of \(i\)–\(ii\)\. \(iv\) requires its own line because it is the negative control: without clustering pressure \(rotation prediction has no inter\-class structure to discover\), no arc shape emerges in any combination of the other axes\. \(i\) and \(ii\) are the two basic kinematic regimes: full V whenβc\\beta\_\{c\}is frozen \(SAE on a frozen encoder is the cleanest case\), fold\-back otherwise\. The fold\-back regime exhibits a magnitude spectrum; CIFAR\-10 mild \(∼0\.5\\sim 0\.5log\-unit drift\), CIFAR\-100 strong \(∼3\\sim 3log\-units\), driven by data complexity \(number of classes / richness of post\-criticalCov\(z\)\\mathrm\{Cov\}\(z\)structure\), but is one regime\.
#### Why not more than four\.
The three degenerate rows of Table[6](https://arxiv.org/html/2605.24057#A9.T6)are not observed in our experiments and are not predicted to be common in standard feature\-learning pipelines\. The first degenerate row would manifest as “descent only”; an apparently monotone descent visible because the encoder begins supercritical \(e\.g\., CIFAR\-10 ResNet\-18 features, wherelog\(β/βc\)≈\+3\.9\\log\(\\beta/\\beta\_\{c\}\)\\approx\+3\.9at random initialization\)\. However, our CIFAR\-10 trajectory in fact exhibits a mild fold\-back \(∼0\.5\\sim 0\.5log\-unit drift\) rather than strict monotonicity, so we classify it as \(ii\) mild fold\-back rather than a distinct regime\. The other two degenerate rows are not observed and are left as open empirical questions\.
## Appendix JWhy decoder\-column matching fails for SAE identity
In Sec\.[5](https://arxiv.org/html/2605.24057#S5)we define per\-atom identity via the activation\-pattern columnf:,k∈ℝNf\_\{:,k\}\\in\\mathbb\{R\}^\{N\}on a fixed eval set, then match identities across checkpoints by Hungarian assignment on cosine similarity of activation vectors\. A more naive choice would be to match*decoder columns*D:,k∈ℝdD\_\{:,k\}\\in\\mathbb\{R\}^\{d\}directly: “two atoms have the same identity if their writeout directions are aligned\.” This naive metric is uninformative in our setting\.
#### Empirical observation\.
On theK=2048K\\\!=\\\!2048SAE trained on Pythia\-160M layer 6 \(d=768d\\\!=\\\!768\), the Hungarian matched cosine between decoder columns of consecutive checkpoints is essentially saturated at1\.0001\.000throughout training \(we observe values in\[0\.984,1\.000\]\[0\.984,1\.000\]for every consecutive pair from initialization to the converged endpoint\)\. The matched cosine between the random initialization and the converged decoder is0\.8950\.895; a0\.100\.10dynamic range\. By contrast the activation\-pattern matched cosine \(used in the main text\) has range0\.13→1\.00\.13\\to 1\.0in the top\-KKSAE and0\.67→1\.00\.67\\to 1\.0in the soft\-L1 SAE, with sharp lock\-in at the post\-critical onset\.
#### Why\.
The decoder\-column Hungarian saturation has two sources\. \(i\) Anthropic\-style SAE training renormalizes decoder columns to unit norm every∼\\sim200 steps; the only motion in decoder direction comes from gradient updates in between renormalizations, which are small relative to data noise\. \(ii\) ForK≫dK\\gg dovercomplete dictionaries, two sets ofKKrandom unit vectors inℝd\\mathbb\{R\}^\{d\}admit a near\-perfect Hungarian pairing because each vector finds a close match among the redundant candidates\. The combination makes the naive decoder\-Hungarian metric saturate at∼\\sim1\.0 even for SAEs whose feature identities are in fact unrelated at the activation level\.
#### Resolution\.
Feature identity for mechanistic interpretability lives in*which tokens an atom fires on*, not in the direction the atom writes into reconstruction\. The activation\-pattern substratef:,kf\_\{:,k\}used in Sec\.[5](https://arxiv.org/html/2605.24057#S5)captures the former\. Matching on activation vectors is invariant to atom permutation and to scaling, and is independent of the overcomplete\-redundancy artifact described above\.
## Appendix KProbe robustness:KprobeK\_\{\\mathrm\{probe\}\}sensitivity
The label\-free indicatorβ\(t\)/βc\(t\)\\beta\(t\)/\\beta\_\{c\}\(t\)has two distinct sensitivities to the GMM probe’s prototype countKprobeK\_\{\\mathrm\{probe\}\}: the critical precisionβc=1/λmax\(Cov\(z\)\)\\beta\_\{c\}=1/\\lambda\_\{\\max\}\(\\mathrm\{Cov\}\(z\)\)depends only on the encoder’s data covariance and isKprobeK\_\{\\mathrm\{probe\}\}\-independent by construction; the GMM’s own learned precisionβ\(t\)\\beta\(t\)evolves under the joint\-detached protocol \(Sec\.[4\.1](https://arxiv.org/html/2605.24057#S4.SS1)\) and may depend onKprobeK\_\{\\mathrm\{probe\}\}\.
We sweepKprobe∈\{2,5,10,20,50\}K\_\{\\mathrm\{probe\}\}\\in\\\{2,5,10,20,50\\\}on the canonical grokking configuration \(p=97p\\\!=\\\!97, WD=1\.0=\\\!1\.0, seed0\) with all other hyperparameters fixed \(Table[7](https://arxiv.org/html/2605.24057#A11.T7)\)\. The result separates into encoder\-controlled invariants and probe\-controlled magnitudes:
1. 1\.*Theβc\(t\)\\beta\_\{c\}\(t\)trajectory is invariant inKprobeK\_\{\\mathrm\{probe\}\}to six decimal places*: att=100t=100,10310^\{3\},10410^\{4\}we readβc=0\.021306,0\.026080,0\.006242\\beta\_\{c\}=0\.021306,0\.026080,0\.006242at everyKprobeK\_\{\\mathrm\{probe\}\}\. This follows from joint\-detached training: the encoder’s gradients never see the probe, so its trajectory andCov\(z\(t\)\)\\mathrm\{Cov\}\(z\(t\)\)are bit\-identical across the five runs\.
2. 2\.*Act 1 crossing time is invariant inKprobeK\_\{\\mathrm\{probe\}\}*at step4040for everyKprobeK\_\{\\mathrm\{probe\}\}\. The crossing happens early enough thatβ\(t\)\\beta\(t\)has not yet diverged across probe sizes \(att=100t=100,logβ\\log\\betaagrees to four decimal places:−1\.6154\-1\.6154at everyKprobeK\_\{\\mathrm\{probe\}\}\)\.
3. 3\.*Act 3 grokking time is invariant inKprobeK\_\{\\mathrm\{probe\}\}*at step8 5008\\,500for everyKprobeK\_\{\\mathrm\{probe\}\}, sincetest\_acc\\mathrm\{test\\\_acc\}is a property of the encoder and the encoder does not see the probe\.
4. 4\.*Post\-criticalβ\(t\)\\beta\(t\)magnitude depends onKprobeK\_\{\\mathrm\{probe\}\}\.*Byt=104t=10^\{4\},logβ\\log\\betahas separated:−1\.62\-1\.62,−1\.32\-1\.32,−1\.04\-1\.04,−0\.84\-0\.84,−0\.45\-0\.45forKprobe=2,5,10,20,50K\_\{\\mathrm\{probe\}\}=2,5,10,20,50, respectively\. Consequentlylog\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)att=104t=10^\{4\}ranges from\+3\.46\+3\.46\(K=2K\\\!=\\\!2\) to\+4\.62\+4\.62\(K=50K\\\!=\\\!50\), a spread of1\.171\.17log\-units\. LargerKprobeK\_\{\\mathrm\{probe\}\}gives the GMM more capacity to fit the post\-critical broken\-symmetry geometry, soβ\\betasaturates at a larger value\.
Table 7:KprobeK\_\{\\mathrm\{probe\}\}sweep on canonical grokking \(p=97p\\\!=\\\!97, WD=1\.0=\\\!1\.0, seed0\)\.*Encoder\-side*quantities \(βc\\beta\_\{c\}, Act 1 crossing step, Act 3 grok step\) areKprobeK\_\{\\mathrm\{probe\}\}\-invariant\.*Probe\-side*quantities \(logβ\(t\)\\log\\beta\(t\),log\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)at late times\) depend onKprobeK\_\{\\mathrm\{probe\}\}: more prototypes gives larger post\-criticalβ\\beta\.#### Operational consequence\.
The framework’s qualitative uses of the indicator \(Sec\.[4\.2](https://arxiv.org/html/2605.24057#S4.SS2): “isβ\>βc\\beta\>\\beta\_\{c\}?”→\\toAct 1 has happened; “how long hasβ\\betastayed aboveβc\\beta\_\{c\}without NC1 collapsing?”→\\toAct 2 plateau length\) areKprobeK\_\{\\mathrm\{probe\}\}\-robust because they are based on encoder\-controlled events\.*Absolute magnitudes*oflog\(β/βc\)\\log\(\\beta/\\beta\_\{c\}\)should not be compared across differentKprobeK\_\{\\mathrm\{probe\}\}choices; within a single fixedKprobeK\_\{\\mathrm\{probe\}\}choice \(we useKprobe=10K\_\{\\mathrm\{probe\}\}=10throughout the paper, following prior soft\-KK\-means and deterministic\-annealing conventions\) the magnitude is well\-defined and tracks the encoder’s post\-critical geometry\.
## Appendix LSupplementary SAE lottery diagnostics
This appendix provides the full evidence for the identity\-lock mechanism \(Sec\.[5\.2](https://arxiv.org/html/2605.24057#S5.SS2)\) and the architecture\-dependent interpretability ceiling \(Sec\.[5\.3](https://arxiv.org/html/2605.24057#S5.SS3)\)\.
Figure 17:Identity lock and architecture\-dependent interpretability ceiling\.\(A\)Hungarian\-matched cosine of per\-atom activation patterns to the converged SAE\. Both soft\-L1 \(dashed blue\) and top\-KK\(solid red\) lock during the post\-critical window beginning at the onset \(vertical dotted lines; NC1 peak in panel C\)\. Top\-KKhas the wider dynamic range \(0\.13→1\.00\.13\\to 1\.0\); soft\-L1 starts from a0\.680\.68baseline because random ReLU rows already project onto the data subspace\.\(B\)Median POS purity of top\-100\-activation positions per atom\. Only top\-KKreaches0\.560\.56\(≈8×\\approx 8\\timesrandom baseline0\.0670\.067\); soft\-L1 plateaus near0\.330\.33\.\(C\)NC1 in SAE feature space peaks at the post\-critical onset, marking the trigger of the lottery window\.\(D, E\)Hungarian\-matched lottery scatter \(note:ρH\\rho\_\{H\}shown here is inflated relative to the identity\-matched signalρid\\rho\_\{\\mathrm\{id\}\}used in Fig\.[4](https://arxiv.org/html/2605.24057#S5.F4); see Sec\.[5\.1](https://arxiv.org/html/2605.24057#S5.SS1)\)\.\(F\)Ranking lift: atoms ranked by step\-∼\\sim1,0001\{,\}000POS purity show\+0\.33\+0\.33lift in mean convergence POS purity \(top decile vs\. bottom decile\) in top\-KK; the top\-decile mean of0\.800\.80is12×12\\timesrandom\.#### Identity lock \(panel A\)\.
In the top\-KKSAE, where random initialization gives genuinely random activation patterns, the matched cosine to the converged decoder starts at0\.130\.13\(baseline for random unit vectors inN=50,000N\\\!=\\\!50\{,\}000\-dim space\) and remains flat for the first∼200\\sim 200training steps\. At the post\-critical onset \(NC1 peak, step261261for top\-KK,383383for soft\-L1\) it begins a sharp rise, reaching0\.920\.92by step1,2121\{,\}212and saturating at1\.01\.0by step∼15,000\\sim 15\{,\}000\. This timing aligns with the identity\-matched POS\-purity correlation \(Sec\.[5\.1](https://arxiv.org/html/2605.24057#S5.SS1)\): both signal that atoms commit to specific activation patterns during the post\-critical window\. The soft\-L1 SAE shows identical*timing*but a compressed dynamic range, because random ReLU encoder rows already project onto the data subspace\.
#### Architecture ceiling \(panel B\)\.
In the soft\-L1 family,L0L\_\{0\}saturates near1,0001\{,\}000ofK=2048K\\\!=\\\!2048active atoms per token irrespective ofλ∈\[5×10−4,10−2\]\\lambda\\in\[5\{\\times\}10^\{\-4\},10^\{\-2\}\]\(10×10\\timesrange of penalty strength gives a1\.5%1\.5\\%change inL0L\_\{0\}\); the features are insufficiently sparse for per\-feature linguistic selectivity to emerge, and median top\-100 POS purity plateaus at0\.330\.33\. Under architectural top\-KKsparsity, the median active feature reaches POS purity0\.560\.56\(8×8\\timesrandom\) and median top\-token entropy drops by0\.790\.79nats\. The improvement is monotonic in training step and aligned with the post\-critical onset\.
#### Frequency\-stratified lottery \(Fig\.[18](https://arxiv.org/html/2605.24057#A12.F18)\)\.
To rule out a token\-frequency confound, we group atoms into five quintiles by the mean log\-frequency of their top\-100 activating tokens at convergence and re\-computeρid\\rho\_\{\\mathrm\{id\}\}within each quintile \(3 seeds,K=2048K\\\!=\\\!2048top\-KKSAE\)\. Panel A:ρid\\rho\_\{\\mathrm\{id\}\}per quintile, all in\[\+0\.37,\+0\.51\]\[\+0\.37,\+0\.51\]\. Panel B: top\-decile mean POS purity per quintile, all≥0\.7\\geq 0\.7\(above the∼\\sim1/15=0\.0671/15=0\.067uniform\-random baseline by≥10×\\geq 10\\times\)\. The lottery effect persists at every frequency stratum and is not concentrated in the most\-frequent quintile, which is what a single\-token\-lock artifact would predict\.
Figure 18:Lottery is stable across token\-frequency strata\.Atoms binned by quintile of their top\-100\-token mean log\-frequency \(Q1 = rarest tokens, Q5 = most frequent\)\.\(A\)ρid\\rho\_\{\\mathrm\{id\}\}per quintile, 3\-seed mean±\\pmstd\. All quintiles produceρid≥0\.37\\rho\_\{\\mathrm\{id\}\}\\geq 0\.37; the unstratified overall is\+0\.42\+0\.42\(dashed\)\.\(B\)Top\-decile mean final POS purity per quintile\. Q4 \(mid\-frequency content words\) has the highestρid\\rho\_\{\\mathrm\{id\}\}, not Q5 \(function words\)\. Frequency confound is not load\-bearing\.
## Appendix MHyperparameters and reproducibility
All experiments use Adam\-family optimizers with the joint\-detached GMM probe protocol described in Sec\.[4\.1](https://arxiv.org/html/2605.24057#S4.SS1)\(Kprobe=10K\_\{\\text\{probe\}\}\\\!=\\\!10,lrμ=5×10−3\\mathrm\{lr\}\_\{\\mu\}\\\!=\\\!5\\\!\\times\\\!10^\{\-3\},lrβ=10−2\\mathrm\{lr\}\_\{\\beta\}\\\!=\\\!10^\{\-2\},logβ0=−2\.5\\log\\beta\_\{0\}\\\!=\\\!\-2\.5\)\. Per\-experiment configurations are summarized in Table[8](https://arxiv.org/html/2605.24057#A13.T8)\.
Table 8:Per\-experiment hyperparameters\. “–” indicates the default \(Adam, no scheduling\)\. Seeds are integers\{0\}\\\{0\\\}unless otherwise noted\. Code and fullargs\.jsonfiles are released alongside the paper\.Similar Articles
Feature Rivalry in Sparse Autoencoder Representations: A Mechanistic Study of Uncertainty-Driven Feature Competition in LLMs
This research paper introduces 'Feature Rivalry' in Sparse Autoencoder representations as a mechanistic signature of uncertainty in LLMs. Using Gemma-2-2B, the study demonstrates that negatively correlated feature pairs localize uncertainty to specific layers and causally influence model outputs.
Emergence via Phase Transitions: Mechanism Landscapes and Universal Convergence Across Complex Systems
This paper introduces the Hierarchical Emergence Framework (HEF), which explains how diverse systems such as neural networks and biological evolution converge to similar internal representations through phase transitions in mechanism landscapes under physical and informational constraints. The framework is validated empirically with 111 grokking experiments that confirm universal convergence and identify a critical energy threshold.
State-Space NTK Collapse Near Bifurcations
This paper develops a local theory of gradient descent near bifurcations in dynamical models, showing that the state-space neural tangent kernel collapses to a rank-one operator that dominates learning dynamics, making optimization effectively low-dimensional and predictable from normal forms.
Hoeffding Concept Bottleneck Models with Applications to Overhead Images
Introduces Hoeffding Concept Bottleneck Models (HCBM), a nonlinear and sparse aggregation of concept scores using Hoeffding functional decomposition of gradient-boosted trees, for improved explainability and accuracy in classification and object detection tasks, with applications to overhead images.
Feature Repulsion and Spectral Lock-in: An Empirical Study of Two-Layer Network Grokking
This empirical study validates theoretical findings on feature repulsion and spectral lock-in during the grokking phenomenon in two-layer neural networks, demonstrating how activation functions influence the transition from memorization to generalization.