Noise-Driven Escape from Metastable Phases explains Grokking in Deep Neural Networks

arXiv cs.LG Papers

Summary

The paper proposes that grokking in deep neural networks arises from noise-driven escape from metastable phases in first-order L2 phase transitions, demonstrating that delayed generalization follows Arrhenius scaling and reproduces canonical grokking curves.

arXiv:2606.17120v1 Announce Type: new Abstract: Deep neural networks (DNNs) exhibit first order phase transitions under variations of the L2 regularization strength, with each transition marking the onset of a new learnable feature. Below a critical regularization strength, all features are in principle learnable, but coexisting metastable states, separated by energy barriers, can trap the network and impede convergence. A strength of DNNs is their ability to generalize. But many open questions remain, among them the origin of so called grokking: the abrupt, delayed onset of generalization after prolonged apparent overfitting. We show for linear DNNs that grokking is consistent with hysteresis in first-order L2 phase transitions: using L2 regularization to engineer deliberate trapping, we demonstrate that a model in a low-accuracy metastable state escapes only when SGD noise drives it across an energy barrier, with escape times following Arrhenius scaling. We reproduce grokking-like delayed convergence across two orders of magnitude in escape time by deliberately trapping models in metastable phases. Using sparse sub-sampling we also reproduce the canonical grokking curve where test error eventually approaches the final training error. Our work suggests that the number of metastable states equals the number of learnable features -- one per singular value of the data covariance -- the potential for hysteresis grows naturally with task complexity. We provide evidence that the same mechanism likely operates in general nonlinear DNNs. Our results provide routes toward more efficient learning schemes.
Original Article
View Cached Full Text

Cached at: 06/17/26, 05:36 AM

# Noise-Driven Escape from Metastable Phases explains Grokking in Deep Neural Networks
Source: [https://arxiv.org/html/2606.17120](https://arxiv.org/html/2606.17120)
\\NameIbrahim Talha Ersoy\\Emailtalha\.ersoy@uni\-potsdam\.de \\NameKaroline Wiesner\\Emailkaroline\.wiesner@uni\-potsdam\.de \\addrComplexity Science Group University of PotsdamPotsdamGermany

###### Abstract

Deep neural networks \(DNNs\) exhibit first\-order phase transitions under variations of the L2 regularisation strength, with each transition marking the onset of a new learnable feature\. Below a critical regularisation strength, all features are in principle learnable, but coexisting metastable states, separated by energy barriers, can trap the network and impede convergence\. A strength of DNNs is their ability to generalise\. But many open questions remain, among them the origin of so\-called grokking: the abrupt, delayed onset of generalisation after prolonged apparent overfitting\. We show for linear DNNs that grokking is consistent with hysteresis in first\-order L2 phase transitions: using L2 regularisation to engineer deliberate trapping, we demonstrate that a model in a low\-accuracy metastable state escapes only when SGD noise drives it across an energy barrier, with escape times following Arrhenius scaling\. We reproduce grokking\-like delayed convergence across two orders of magnitude in escape time by deliberately trapping models in metastable phases\. Using sparse sub\-sampling we also reproduce the canonical grokking curve where test error eventually approaches the final training error\. Our work suggests that the number of metastable states equals the number of learnable features – one per singular value of the data covariance – the potential for hysteresis grows naturally with task complexity\. We provide evidence that the same mechanism likely operates in general nonlinear DNNs\. Our results provide routes toward more efficient learning schemes\.

## 1Introduction

L2 regularisation is a cornerstone of modern machine learning, employed to combat overfitting from classical ridge regression\[[1](https://arxiv.org/html/2606.17120#bib.bib1)\]to large\-scale deep learning\[[2](https://arxiv.org/html/2606.17120#bib.bib2)\]\. Beyond its practical role, L2 regularisation gives rise to rich phenomenology recently understood through statistical physics\. Ziyin and Ueda\[[3](https://arxiv.org/html/2606.17120#bib.bib3)\]showed that varying regularisation strength drives genuine phase transitions in DNNs, identifying first\-order transitions in DNNs at the onset of learning \(the transition from under\- to over\-parameterisation\)\. Subsequent work extended this picture beyond the onset: Ersoy and Wiesner linked these transitions to curvature drops in the loss landscape\[[6](https://arxiv.org/html/2606.17120#bib.bib6)\], Ersoy, Licha, and Wiesner connected them to learnable feature hierarchies\[[7](https://arxiv.org/html/2606.17120#bib.bib7)\], and Ladewig, Ersoy, and Wiesner showed that in linear networks each non\-zero singular value of the data covariance produces its own rank transition\[[8](https://arxiv.org/html/2606.17120#bib.bib8)\], meaning transitions recur throughout training, once per learnable feature\.*Grokking*, the sudden transition from near\-zero to near\-perfect generalisation long after the training loss has saturated, has attracted much attention since Power et al\. demonstrated it in transformers trained on modular arithmetic\[[9](https://arxiv.org/html/2606.17120#bib.bib9)\]\. The origin of this phenomenon is not yet fully understood\. Liu et al\.\[[10](https://arxiv.org/html/2606.17120#bib.bib10)\]implicated regularisation as the central driver of grokking\. Nanda et al\.\[[21](https://arxiv.org/html/2606.17120#bib.bib21)\]identified interpretable representations at the transition, and Tian\[[22](https://arxiv.org/html/2606.17120#bib.bib22)\]derived scaling laws for feature emergence\. Rubin et al\.\[[11](https://arxiv.org/html/2606.17120#bib.bib11)\]proposed a first\-order phase transition analogy involving an entropy barrier\. This was then challenged by Zhang et al\.\[[23](https://arxiv.org/html/2606.17120#bib.bib23)\], who found no entropy barrier and proposed glassy relaxation instead\. Solvable linear models have likewise been shown to grok\[[24](https://arxiv.org/html/2606.17120#bib.bib24)\], and a distribution shift between training and test data has been identified as a driver of delayed generalisation\[[25](https://arxiv.org/html/2606.17120#bib.bib25)\]\. We offer an alternative account\. As the central mechanism we suggest*hysteresis*in analogy to the statistical physics of phase transitions: a model initialised in a low\-accuracy phase remains there until SGD noise drives it across an energy \(loss\) barrier into the globally optimal phase\. To show that hysteresis is consistent with the grokking phenomenology, we build on the finding that varying the L2 regularisation strength leads to first\-order phase transitions\. We ask: Can activated escape from the resulting metastable states, rather than the transitions themselves, account for the hallmark features of grokking? To answer this, we use L2 regularisation as a control tool to engineer metastable trapping\. We successfully reproduce grokking behaviour, and, furthermore, show that this activated process is governed by Arrhenius\-type kinetics\[[12](https://arxiv.org/html/2606.17120#bib.bib12),[13](https://arxiv.org/html/2606.17120#bib.bib13)\]with an effective temperature\[[14](https://arxiv.org/html/2606.17120#bib.bib14)\]Teff∝ηlr/BT\_\{\\mathrm\{eff\}\}\\propto\\eta\_\{\\mathrm\{lr\}\}/B\. Hereηlr\\eta\_\{\\mathrm\{lr\}\}is the learning rate andBBthe batch size, making the escape time exponentially sensitive to hyperparameter choices\. Our results, established in Sections[2](https://arxiv.org/html/2606.17120#S2)–[3](https://arxiv.org/html/2606.17120#S3), support three claims:\(1\)first\-order L2 phase transitions produceddcoexisting metastable states \(one per learnable feature\) whose energy barriers trap models in low\-accuracy phases\.\(2\)Escape is analogous to a thermally activated process obeying Arrhenius kinetics withTeff∝ηlr/BT\_\{\\mathrm\{eff\}\}\\propto\\eta\_\{\\mathrm\{lr\}\}/B; we confirm this numerically withR2=0\.991R^\{2\}=0\.991\.\(3\)Deliberate trapping reproduces hallmarks of grokking, i\.e\. long delay, abruptness, sensitivity to initialisation, and, under sparse sampling, the train/test dissociation, across two orders of magnitude in escape time\.

## 2Methods

### 2\.1L2 Phase Transitions

We use deep linear networks as our minimal model because their loss landscape is exactly solvable, allowing us to locate allddmetastable minima and energy barriers analytically\. All qualitative results, saddle\-node bifurcations, coexisting phases, Arrhenius escape, survive in nonlinear networks \(Appendix[G](https://arxiv.org/html/2606.17120#A7)\), but the linear case keeps the analysis tractable\. In this section we review the key prior results on L2 phase transitions before introducing our escape\-time framework\. Ladewig et al\.\[[8](https://arxiv.org/html/2606.17120#bib.bib8)\]described the precise mechanism for linear DNNs, linking the transitions to the singular values of the data covariance: with inputxxand outputyy, covariancesΣx​x\\Sigma\_\{xx\},Σy​y\\Sigma\_\{yy\}, and cross\-covarianceΣy​x\\Sigma\_\{yx\}, the singular valuesηi\\eta\_\{i\}ofΣy​x\\Sigma\_\{yx\}directly characterise the learnable features in the aligned case \(Σx​x=𝐈\\Sigma\_\{xx\}=\\mathbf\{I\}\)\. After the network reaches the 0\-balanced subspace, the set of weight configurations where all layers contribute equally to the end\-to\-end map, the L2\-regularised loss decouples into independent terms for each singular valueλi\\lambda\_\{i\}of the end\-to\-end weight matrix:

ℒ​\(\{λi\},β\)=12​∑i=1d\(λi−ηi\)2\+β2​∑i=1dλi2/L,\\mathcal\{L\}\(\\\{\\lambda\_\{i\}\\\},\\beta\)=\\frac\{1\}\{2\}\\sum\_\{i=1\}^\{d\}\(\\lambda\_\{i\}\-\\eta\_\{i\}\)^\{2\}\+\\frac\{\\beta\}\{2\}\\sum\_\{i=1\}^\{d\}\\lambda\_\{i\}^\{2/L\},\(1\)whereLLis network depth,ddis the number of non\-zero modes, andβ\>0\\beta\>0is the regularisation strength\. For depthL≥3L\\geq 3, the stationarity condition∂ℒ/∂λi=0\\partial\\mathcal\{L\}/\\partial\\lambda\_\{i\}=0for a single mode reads:

λ−η\+β​λ2/L−1=0\.\\lambda\-\\eta\+\\beta\\,\\lambda^\{2/L\-1\}=0\.\(2\)Asβ\\betaincreases throughβc\\beta\_\{c\}, the two non\-trivial solutions \(a stable minimum and an unstable saddle\) merge and annihilate in a saddle\-node bifurcation \(see Appendix[D](https://arxiv.org/html/2606.17120#A4)\), leaving onlyλ=0\\lambda=0\. Belowβc\\beta\_\{c\}, zero and non\-zero rank solutions coexist, separated by a finite energy barrier \(see Appendix[C](https://arxiv.org/html/2606.17120#A3)for the bifurcation diagram\)\. The critical regularisation strength is:

βc\(i\)=11−k​\(ηi​1−k2−k\)2−k,k=2L\.\\beta\_\{c\}^\{\(i\)\}=\\frac\{1\}\{1\-k\}\\left\(\\eta\_\{i\}\\frac\{1\-k\}\{2\-k\}\\right\)^\{2\-k\},\\qquad k=\\frac\{2\}\{L\}\.\(3\)Forβ\>βc\(i\)\\beta\>\\beta\_\{c\}^\{\(i\)\}the metastable minimum vanishes\. The equal\-loss pointβi∗\\beta^\{\*\}\_\{i\}\(where the lower and higher rank solutions have equal regularised loss\) lies strictly belowβc\(i\)\\beta\_\{c\}^\{\(i\)\}forL≥3L\\geq 3, producing hysteresis: a model remains trapped in the lower rank phase even when the higher rank phase is energetically preferred\. The orderingη1\>⋯\>ηd\>0\\eta\_\{1\}\>\\cdots\>\\eta\_\{d\}\>0givesdddistinct bifurcations, one metastable state per learnable feature\.

### 2\.2SGD as Langevin Dynamics and Arrhenius Escape

To model escape from metastable states, we note that SGD mini\-batch noise injects stochastic fluctuations into the gradient with covariance scaling asηlr/B\\eta\_\{\\mathrm\{lr\}\}/B, whereηlr\\eta\_\{\\mathrm\{lr\}\}is the learning rate andBBthe batch size\. In the overdamped limit, the dynamics maps onto Langevin dynamics with an effective temperature\[[14](https://arxiv.org/html/2606.17120#bib.bib14)\]:

Teff∝ηlrB\.T\_\{\\mathrm\{eff\}\}\\propto\\frac\{\\eta\_\{\\mathrm\{lr\}\}\}\{B\}\.\(4\)The mean escape time from a metastable state then follows the Kramers–Arrhenius law\[[12](https://arxiv.org/html/2606.17120#bib.bib12),[13](https://arxiv.org/html/2606.17120#bib.bib13)\]:

ln⁡τ=ln⁡τ0\+Δ​EeffTeff,\\ln\\tau=\\ln\\tau\_\{0\}\+\\frac\{\\Delta E\_\{\\mathrm\{eff\}\}\}\{T\_\{\\mathrm\{eff\}\}\},\(5\)whereΔ​Eeff\\Delta E\_\{\\mathrm\{eff\}\}is an effective barrier height absorbing high\-dimensional curvature and entropic corrections \(see Appendix[F](https://arxiv.org/html/2606.17120#A6)\)\. This yields the falsifiable prediction thatln⁡τ\\ln\\tauis linear inB/ηlrB/\\eta\_\{\\mathrm\{lr\}\}\.

## 3Results

### 3\.1Hysteresis and Trapping Reproduce Grokking

The coexistence of metastable states implies that the training outcome depends sensitively on initialisation\. To demonstrate this we operate atβ=0\.32\\beta=0\.32, which lies in the rangeβ<βc\(1\)\\beta<\\beta\_\{c\}^\{\(1\)\}, so the global minimum is rank\-2 but a metastable rank\-1 minimum also exists\. All three experiments use the learning rateηlr=0\.08\\eta\_\{\\mathrm\{lr\}\}=0\.08and batch sizeB=64B=64\. We consider three initialisation protocols \(Fig\.[1](https://arxiv.org/html/2606.17120#S3.F1)\):\(i\) Random initialisation\.*Setup*: standard random weight initialisation atβ=0\.32\\beta=0\.32\.*Observation*: the model converges rapidly to the rank\-2 global minimum \(Fig\.[1](https://arxiv.org/html/2606.17120#S3.F1), blue;τ≈10\\tau\\approx 10epochs\)\.*Conclusion*: no trapping occurs when the model starts outside any metastable phase\.\(ii\) Rank\-1 trap\.*Setup*: model initialised from a checkpoint trained atβ\>βc\(1\)\\beta\>\\beta\_\{c\}^\{\(1\)\}, placing it in the metastable rank\-1 phase\.*Observation*: the model remains in the rank\-1 state forτ≈5500\\tau\\approx 5500epochs before abruptly escaping to rank\-2 \(Fig\.[1](https://arxiv.org/html/2606.17120#S3.F1), green\)\.*Conclusion*: trapping in a metastable phase produces grokking\-like delayed convergence\.\(iii\) Trivial\-phase trap\.*Setup*: model initialised from a checkpoint trained atβ\>βc\(2\)\\beta\>\\beta\_\{c\}^\{\(2\)\}, placing it in the rank\-0 phase\.*Observation*: the model escapes to rank\-1 afterτ\>7000\\tau\>7000epochs, then remains trapped in rank\-1 for a further extended period, with total delayτ\>10,000\\tau\>10\{,\}000epochs \(Fig\.[1](https://arxiv.org/html/2606.17120#S3.F1), orange\)\.*Conclusion*: sequential trapping across multiple metastable phases reproduces staged grokking delays spanning two orders of magnitude\.\(iv\) Canonical grokking via sparse sub\-sampling\.*Setup*: train and test are drawn from the*same*distribution \(weak correlation0\.80\.8, strong correlation0\.90\.9\), but the training set is very sparse, only2525samples \(0\.5%0\.5\\%of a5 0005\\,000\-sample pool\), so the weak feature is poorly determined from the training data while remaining at full strength on test\. The model is initialised in the rank\-1 trapped state and trained atβ=0\.0025\\beta=0\.0025, a small regularisation chosen so that the plateau is long\-lived but eventually resolves\.*Observation*: After a brief transient in which the strong mode equilibrates, both errors plateau, with train below test \(Fig\.[1](https://arxiv.org/html/2606.17120#S3.F1)\(b\)\); the model holds in the rank\-1 phase for≈1500\{\\approx\}1500epochs, then escapes to rank\-2, and test error falls sharply to approach train error to within a small residual gap set by the irreducible noise of the stochastic task\.*Conclusion*: the same trapping mechanism reproduces the canonical grokking curve when the weak feature is sufficiently underrepresented in the training sample\.

![Refer to caption](https://arxiv.org/html/2606.17120v1/x1.png)

\(a\) Hysteresis as Initialisation Dependence

![Refer to caption](https://arxiv.org/html/2606.17120v1/x2.png)

\(b\) Train\-Test Discrepancy in Hysteresis

Figure 1:Delayed convergence in deep linear networks\.\(a\)Random \(blue\), rank‑1 trap \(green\), trivial‑phase trap \(orange\) atβ=0\.03\\beta=0\.03,ηlr=0\.08\\eta\_\{\\mathrm\{lr\}\}=0\.08,B=64B=64\. The convergence is strongly delayed when the model is initialised in the local minima of the lower accuracy phases\.\(b\)Canonical grokking via sparse sub\-sampling \(2525training samples,β=0\.0025\\beta=0\.0025\)\. Initialised in the rank\-1 phase, train MSE \(blue\) drops quickly but only to a plateau while test MSE \(red\) plateaus higher; at≈1500\{\\approx\}1500epochs the model escapes the rank\-1 phase and test MSE falls sharply, approaching train MSE to within a small residual gap set by the irreducible noise of the task\.
### 3\.2Energy Barrier Between Metastable States

The trapping mechanism requires a finite energy barrier between the differing rank phases\. Fig\.[2](https://arxiv.org/html/2606.17120#S3.F2)\(a\) shows the loss landscape section along the minimal\-loss path atβ=0\.32\\beta=0\.32, parameterised by the singular valueλ\\lambda\. A clear barrier separates the local minimum atλ=0\\lambda=0from the global minimum\. We evaluate Eq\. \([1](https://arxiv.org/html/2606.17120#S2.E1)\) numerically at the saddle pointλsad≈0\.41\\lambda^\{\\mathrm\{sad\}\}\\approx 0\.41\(obtained from Eq\. \([2](https://arxiv.org/html/2606.17120#S2.E2)\) atβ=0\.32\\beta=0\.32,η=0\.8\\eta=0\.8,L=3L=3\)\. This givesΔ​Emin≈0\.003\\Delta E\_\{\\mathrm\{min\}\}\\approx 0\.003, the barrier along the lowest\-loss path out of the metastable phase \(Fig\.[2](https://arxiv.org/html/2606.17120#S3.F2)\(a\)\)\. As we show next, the effective barrier governing escape times is far larger\.

### 3\.3Arrhenius Scaling Confirms Thermally Activated Escape

To test whether the escape is governed by the activated barrier crossing as predicted by Eq\. \([5](https://arxiv.org/html/2606.17120#S2.E5)\), we measured escape times from the rank\-1 trapped state atβ=0\.32\\beta=0\.32while varyingηlr∈\[5×10−4,5×10−3\]\\eta\_\{\\mathrm\{lr\}\}\\in\[5\\times 10^\{\-4\},5\\times 10^\{\-3\}\]with batch size fixed atB=64B=64\. WithTeff∝ηlr/BT\_\{\\mathrm\{eff\}\}\\propto\\eta\_\{\\mathrm\{lr\}\}/B, Eq\. \([5](https://arxiv.org/html/2606.17120#S2.E5)\) predictsln⁡τ\\ln\\tauto be linear inB/ηlrB/\\eta\_\{\\mathrm\{lr\}\}\. Figure[2](https://arxiv.org/html/2606.17120#S3.F2)\(b\) shows this Arrhenius plot; the linear fit achievesR2=0\.991R^\{2\}=0\.991, confirming Eq\. \([5](https://arxiv.org/html/2606.17120#S2.E5)\)\. The slope yields the effective barrier:

Δ​Eeff=0\.15±0\.05\.\\Delta E\_\{\\mathrm\{eff\}\}=0\.15\\pm 0\.05\.\(6\)This is far aboveΔ​Emin≈0\.003\\Delta E\_\{\\mathrm\{min\}\}\\approx 0\.003shown in Fig\.[2](https://arxiv.org/html/2606.17120#S3.F2)\(a\)\. The discrepancy is expected: in aD=170D=170\-dimensional parameter space the Kramers–Langer formula adds entropic and geometric corrections from theD−1D\-1transverse directions, drivingΔ​Eeff≫Δ​Emin\\Delta E\_\{\\mathrm\{eff\}\}\\gg\\Delta E\_\{\\mathrm\{min\}\}\(Appendix[F](https://arxiv.org/html/2606.17120#A6)\)\.

![Refer to caption](https://arxiv.org/html/2606.17120v1/x3.png)

\(a\) Loss Landscape Barrier

![Refer to caption](https://arxiv.org/html/2606.17120v1/x4.png)

\(b\) Arrhenius Fit for Escape Times

Figure 2:\(a\)Loss landscape section atβ=0\.32\\beta=0\.32,η=0\.8\\eta=0\.8,L=3L=3; dotted curves showβ=0\.29\\beta=0\.29andβ=0\.35\\beta=0\.35\. The barrier from the local minimum atλ=0\\lambda=0to the saddle givesΔ​Emin≈0\.003\\Delta E\_\{\\mathrm\{min\}\}\\approx 0\.003\.\(b\)Arrhenius fit:ln⁡τ\\ln\\tauversus1/Teff1/T\_\{\\mathrm\{eff\}\}\. The linear fit \(R2=0\.991R^\{2\}=0\.991\) confirms thermally activated barrier crossing; the slope givesΔ​Eeff=0\.15±0\.05≫Δ​Emin\\Delta E\_\{\\mathrm\{eff\}\}=0\.15\\pm 0\.05\\gg\\Delta E\_\{\\mathrm\{min\}\}, with the discrepancy explained by entropic and curvature corrections inD=170D=170dimensions \(Appendix[F](https://arxiv.org/html/2606.17120#A6)\)\.

## 4Discussion

### 4\.1A Mechanistic Account of Grokking

Our results establish three claims\. First, first\-order L2 phase transitions produce coexisting metastable states whose barriers trap models in low\-accuracy phases\. Second, escape from these states is an activated process governed by Arrhenius kinetics withTeff∝ηlr/BT\_\{\\mathrm\{eff\}\}\\propto\\eta\_\{\\mathrm\{lr\}\}/B\. Third, deliberate trapping reproduces the hallmark features of grokking, long delay, abruptness, sensitivity to initialisation, and the train/test dissociation under sparse sampling, in an analytically tractable minimal model\. This framework makes concrete, falsifiable predictions:

1. 1\.*Staged grokking*: In tasks withddlearnable features, grokking should proceed in up todddiscrete stages, one per singular value ofΣy​x\\Sigma\_\{yx\}\.
2. 2\.*Depth dependence*: Deeper networks \(LLlarger\) produce higher barriers and thus longer grokking delays, since the critical strengthβc\(i\)\\beta\_\{c\}^\{\(i\)\}decreases withLLwhile the equal\-loss crossingβi∗\\beta^\{\*\}\_\{i\}remains finite\[[8](https://arxiv.org/html/2606.17120#bib.bib8)\]\.
3. 3\.*Hyperparameter control*: Escape time obeysln⁡τ∝B/ηlr\\ln\\tau\\propto B/\\eta\_\{\\mathrm\{lr\}\}, providing a direct lever to accelerate or suppress grokking\.

Predictions \(2\) and \(3\) are immediately testable by varyingLL,ηlr\\eta\_\{\\mathrm\{lr\}\}, andBBin existing grokking benchmarks\. Preliminary experiments and work on nonlinear networks\[[7](https://arxiv.org/html/2606.17120#bib.bib7)\]with sigmoid and tanh activations reproduce qualitatively identical first\-order phase transition behaviour, strongly suggesting that the same mechanism applies beyond the linear framework\.

### 4\.2Relation to Memorisation and Generalisation

In the dense\-data experiments of Fig\.[1](https://arxiv.org/html/2606.17120#S3.F1)\(a\) the training error converges monotonically, and the large train/test gap of canonical grokking\[[9](https://arxiv.org/html/2606.17120#bib.bib9)\]does not appear, because withd=2d=2well\-sampled modes the model has few directions along which train and test can diverge\. Under sparse sub\-sampling this gap does emerge: Fig\.[1](https://arxiv.org/html/2606.17120#S3.F1)\(b\) shows train error settling onto a plateau while test error remains substantially higher, until both fall abruptly as the model escapes the rank\-1 phase\. Because the network is linear and the task is stochastic, the model cannot memorise the training sample: train error plateaus at the best rank\-1 linear fit and, after escape, settles near the best rank\-2 fit, which is bounded below by the irreducible \(Bayes\) error of the Gaussian task\. Train and test therefore approach one another and the noise floor but cannot coincide or fall to zero\. The grokking signature here is thus the delayed, abrupt closing of a train/test gap down to the task’s noise floor, reproduced*without any memorising solution*— consistent with the view that memorisation is not a necessary ingredient of the phenomenon\. Whenddis large, the optimisation trajectory reduces training error by learning some singular directions while stagnating in others, giving rise to an apparent memorisation phase\. In this view, memorisation and generalisation need not be primitive concepts but instead emerge as descriptions of partial versus complete progress through a cascade of rank transitions\. Moreover, what is currently labeled “grokking” possibly encompasses several mechanistically distinct phenomena with a superficial resemblance; our framework offers a principled basis for distinguishing them\.

## 5Conclusion

We propose a candidate mechanism for grokking: noise\-activated escape from metastable states created by first\-order phase transitions in L2\-regularised deep networks\. Using deep linear networks, i\.e\. the minimal setting in which the loss landscape is exactly solvable\[[8](https://arxiv.org/html/2606.17120#bib.bib8)\], we demonstrate three results:\(i\)deliberate trapping in metastable phases reproduces grokking\-like delayed convergence;\(ii\)escape times obey Arrhenius scalingln⁡τ∝1/Teff\\ln\\tau\\propto 1/T\_\{\\mathrm\{eff\}\}withTeff∝ηlr/BT\_\{\\mathrm\{eff\}\}\\propto\\eta\_\{\\mathrm\{lr\}\}/B; and\(iii\)the extracted barrierΔ​Eeff=0\.15±0\.05\\Delta E\_\{\\mathrm\{eff\}\}=0\.15\\pm 0\.05exceeds the minimum\-path barrierΔ​Emin≈0\.003\\Delta E\_\{\\mathrm\{min\}\}\\approx 0\.003by the factor predicted from entropic and curvature corrections inD=170D=170dimensions\. The framework connects grokking to the established physics of noise\-activated escape from metastable states\[[12](https://arxiv.org/html/2606.17120#bib.bib12),[13](https://arxiv.org/html/2606.17120#bib.bib13),[19](https://arxiv.org/html/2606.17120#bib.bib19)\], and offers a direct practical handle: sinceln⁡τ∝B/ηlr\\ln\\tau\\propto B/\\eta\_\{\\mathrm\{lr\}\}, grokking delays can in principle be accelerated or suppressed by hyperparameter choice alone\.

## Acknowledgments

We thank Björn Ladewig and members of the Complexity Science Group for stimulating discussions and inspiring contributions\.

## References

- \[1\]A\. E\. Hoerl and R\. W\. Kennard, “Ridge regression: Biased estimation for nonorthogonal problems,”Technometrics12, 55 \(1970\)\.
- \[2\]I\. Goodfellow, Y\. Bengio, and A\. Courville,Deep Learning\(MIT Press, Cambridge, MA, 2017\)\.
- \[3\]Ziyin, L\. & Ueda, M\. Zeroth, first, and second\-order phase transitions in deep neural networks\.Physical Review Research\.5, 043243 \(2023\)
- \[4\]S\. Amari,Information Geometry and Its Applications, Vol\. 194 \(Springer, Tokyo, 2016\)\.
- \[5\]S\. Watanabe,Algebraic Geometry and Statistical Learning Theory, Vol\. 25 \(Cambridge University Press, Cambridge, 2009\)\.
- \[6\]I\. T\. Ersoy and K\. Wiesner, “Exploring L2\-phase transitions on error landscapes,”Workshop on High\-Dimensional Learning Dynamics\(2025\)\.
- \[7\]I\. T\. Ersoy, A\. F\. C\. Licha, and K\. Wiesner, “Phase transitions reveal hierarchical structure in deep neural networks,” arXiv:2512\.11866 \(2025\)\.
- \[8\]B\. Ladewig, I\. T\. Ersoy, and K\. Wiesner, “Cascading through the Hierarchy: Regulariser\-induced Feature Detection as Phase Transitions in Deep Linear Neural Networks,”\(in preparation\)\(2026\)\.
- \[9\]A\. Power, Y\. Burda, H\. Edwards, I\. Babuschkin, and V\. Misra, “Grokking: Generalization beyond overfitting on small algorithmic datasets,” arXiv:2201\.02177 \(2022\)\.
- \[10\]Z\. Liu, E\. J\. Michaud, and M\. Tegmark, “Omnigrok: Grokking beyond algorithmic data,” arXiv:2210\.01117 \(2022\)\.
- \[11\]N\. Rubin, I\. Seroussi, and Z\. Ringel, “Grokking as a first order phase transition in two layer networks,”International Conference on Learning Representations\(2024\)\.
- \[12\]H\. A\. Kramers, “Brownian motion in a field of force and the diffusion model of chemical reactions,”Physica7, 284 \(1940\)\.
- \[13\]P\. Hänggi, P\. Talkner, and M\. Borkovec, “Reaction\-rate theory: fifty years after Kramers,”Rev\. Mod\. Phys\.62, 251 \(1990\)\.
- \[14\]S\. Mandt, M\. D\. Hoffman, and D\. M\. Blei, “Stochastic gradient descent as approximate Bayesian inference,”J\. Mach\. Learn\. Res\.18, 4873 \(2017\)\.
- \[15\]Smith, S\., Kindermans, P\., Ying, C\. & Le, Q\. Don’t decay the learning rate, increase the batch size\.ArXiv Preprint ArXiv:1711\.00489\. \(2017\)
- \[16\]Welling, M\. & Teh, Y\. Bayesian learning via stochastic gradient Langevin dynamics\.Proceedings Of The 28th International Conference On Machine Learning \(ICML\-11\)\. pp\. 681\-688 \(2011\)
- \[17\]Saxe, A\., McClelland, J\. & Ganguli, S\. Exact solutions to the nonlinear dynamics of learning in deep linear neural networks\.ArXiv Preprint ArXiv:1312\.6120\. \(2013\)
- \[18\]Wang, Z\. & Jacot, A\. Implicit bias of SGD in L2\-regularized linear DNNs: One\-way jumps from high to low rank\.ArXiv Preprint ArXiv:2305\.16038\. \(2023\)
- \[19\]T\. Mori, L\. Ziyin, K\. Liu, and M\. Ueda, “Power\-law escape rate of SGD,”Proceedings of the 39th International Conference on Machine Learning, PMLR162, 15959 \(2022\)\.
- \[20\]Draxler, F\., Veschgini, K\., Salmhofer, M\. & Hamprecht, F\. Essentially No Barriers in Neural Network Energy Landscape\. \(2019\)
- \[21\]Nanda, N\., Chan, L\., Lieberum, T\., Smith, J\. & Steinhardt, J\. Progress measures for grokking via mechanistic interpretability\.ArXiv Preprint ArXiv:2301\.05217\. \(2023\)
- \[22\]Y\. Tian, “Provable scaling laws of feature emergence from learning dynamics of grokking,” arXiv:2509\.21519 \(2025\)\.
- \[23\]Zhang, X\., Shang, Y\., Yang, E\. & Zhang, G\. Is Grokking a Computational Glass Relaxation?\. \(2026\), https://arxiv\.org/abs/2505\.11411
- \[24\]N\. Levi, A\. Beck, and Y\. Bar\-Sinai, “Grokking in linear estimators—a solvable model that groks without understanding,”International Conference on Learning Representations\(2024\), arXiv:2310\.16441\.
- \[25\]K\. Lyu, J\. Jin, Z\. Li, S\. S\. Du, J\. D\. Lee, and W\. Hu, “Dichotomy of early and late phase implicit biases can provably induce grokking,”International Conference on Learning Representations\(2024\), arXiv:2311\.18817\.

## Appendix AExperimental Setup

All experiments use a deep linear network of depthL=3L=3\(two hidden layers\) with hidden widthw=10w=10and output dimensionp=2p=2, givingD=170D=170total parameters\. No nonlinear activation is applied; the end\-to\-end map is a single matrix product\. The optimiser is SGD throughout\.

Data generation\.Training and test data each consist ofN=512N=512samples drawn from a zero\-mean multivariate Gaussian over the joint input–output space\. The data covariance hasd=2d=2non\-zero singular valuesη1=0\.9\\eta\_\{1\}=0\.9,η2=0\.8\\eta\_\{2\}=0\.8withΣx​x=𝐈\\Sigma\_\{xx\}=\\mathbf\{I\}\. Samples are generated via Cholesky decomposition: givenΣ=L​L⊤\\Sigma=LL^\{\\top\}, each sample is drawn asz=L​ϵz=L\\epsilonwithϵ∼𝒩​\(0,𝐈\)\\epsilon\\sim\\mathcal\{N\}\(0,\\mathbf\{I\}\); the firstppcomponents are input and the remainingppcomponents are output\. An independent test set of the same size is held fixed throughout\.

Initialisation and annealing\.All trapping experiments use*checkpoint initialisation*: the model is first trained to convergenceβinit=0\\beta\_\{\\mathrm\{init\}\}=0to get the full rank solution\. The resulting weights are then used as the starting point for continued training\. This is analogous to annealing\. The random initialisation baseline uses fresh weights drawn from𝒩​\(0,1/w\)\\mathcal\{N\}\(0,1/\\sqrt\{w\}\)instead\. Convergence is assessed by monitoring test MSE; escape from a metastable phase is identified as the epoch at which test MSE drops below a threshold midway between the metastable plateau value and the global minimum value\.

Phase diagram\(Fig\.[3](https://arxiv.org/html/2606.17120#A2.F3), Appendix[B](https://arxiv.org/html/2606.17120#A2)\): the model is first trained to convergence atβ=0\\beta=0, thenβ\\betais increased quasi\-statically in small increments, re\-training to convergence at each step\.

Trapping experiments\(Fig\.[1](https://arxiv.org/html/2606.17120#S3.F1)\): all three protocols useβ=0\.32\\beta=0\.32,ηlr=0\.08\\eta\_\{\\mathrm\{lr\}\}=0\.08,B=64B=64\.

- •*Random initialisation*: weights drawn from a standard normal distribution scaled by1/w1/\\sqrt\{w\}; training run for2×1042\\times 10^\{4\}epochs\.
- •*Rank\-1 trap*: model initialised from a checkpoint trained to convergence atβ=0\.361\>βc\(1\)=0\.360\\beta=0\.361\>\\beta\_\{c\}^\{\(1\)\}=0\.360, placing the end\-to\-end weight matrix in the rank\-1 phase; training then continued atβ=0\.32\\beta=0\.32for2×1042\\times 10^\{4\}epochs\.
- •*Trivial\-phase trap*: model initialised from a checkpoint trained to convergence atβ=0\.43\>βc\(2\)=0\.42\\beta=0\.43\>\\beta\_\{c\}^\{\(2\)\}=0\.42, placing the network in the rank\-0 phase; training continued atβ=0\.32\\beta=0\.32for2×1042\\times 10^\{4\}epochs\.

Each experiment is repeated across 50 random seeds \(seeds 42–91\); escape times vary across seeds due to stochastic SGD noise, which is the activated escape mechanism under study\.Arrhenius sweep\(Fig\.[2](https://arxiv.org/html/2606.17120#S3.F2)\(b\)\): starting from the rank\-1 trap initialisation above, escape times are measured atβ=0\.32\\beta=0\.32withB=64B=64fixed while varyingηlr∈\{5×10−4,1×10−3,2×10−3,3×10−3,5×10−3\}\\eta\_\{\\mathrm\{lr\}\}\\in\\\{5\\times 10^\{\-4\},\\,1\\times 10^\{\-3\},\\,2\\times 10^\{\-3\},\\,3\\times 10^\{\-3\},\\,5\\times 10^\{\-3\}\\\}\. Each point is averaged over 50 seeds\. The proportionality constantCCinTeff=ηlr/\(B​C\)T\_\{\\mathrm\{eff\}\}=\\eta\_\{\\mathrm\{lr\}\}/\(BC\)is estimated using the trace of the covariance of the gradients\. The linearity \(R2=0\.991R^\{2\}=0\.991\) is a key result\. WithN=512N=512andB=64B=64, each epoch consists of⌊N/B⌋=8\\lfloor N/B\\rfloor=8SGD steps; this factor entersln⁡τ0\\ln\\tau\_\{0\}but not the slopeΔ​Eeff/Teff\\Delta E\_\{\\mathrm\{eff\}\}/T\_\{\\mathrm\{eff\}\}, leaving the extracted barrier unchanged\.Sparse sub\-sampling experiment\(Fig\.[1](https://arxiv.org/html/2606.17120#S3.F1)\(b\)\): train and test are drawn from a single zero\-mean Gaussian with weak correlation0\.80\.8and strong correlation0\.90\.9\(Σx​x=𝐈\\Sigma\_\{xx\}=\\mathbf\{I\}\)\. The training set is a random2525\-sample subset of a5 0005\\,000\-sample pool \(0\.5%0\.5\\%\), so the weak feature is poorly determined from the training data while remaining at full strength on the \(disjoint,10 00010\\,000\-sample\) test set\. The model is initialised in the rank\-1 trapped state and trained atβ=0\.0025\\beta=0\.0025,ηlr=0\.1\\eta\_\{\\mathrm\{lr\}\}=0\.1,B=64B=64\. Escape is identified as the epoch at which the second singular value of the end\-to\-end map rises above a small threshold \(σ2\>10−3\\sigma\_\{2\}\>10^\{\-3\}\)\. Because the network is linear, the training error cannot reach zero: it plateaus at the best rank\-1 linear fit and, after escape, settles near the best rank\-2 fit, both bounded below by the irreducible \(Bayes\) error of the stochastic task\.

## Appendix BPhase Diagram

Figure[3](https://arxiv.org/html/2606.17120#A2.F3)shows the test MSE as a function ofβ\\betafor a deep linear network withd=2d=2singular values \(η1=0\.9\\eta\_\{1\}=0\.9,η2=0\.8\\eta\_\{2\}=0\.8,Σx​x=𝐈\\Sigma\_\{xx\}=\\mathbf\{I\},L=3L=3\), obtained by training to convergence and then quasi\-statically decreasingβ\\betain steps of0\.010\.01\. Two sharp drops mark the transitions atβc\(1\)=0\.36\\beta\_\{c\}^\{\(1\)\}=0\.36andβc\(2\)=0\.42\\beta\_\{c\}^\{\(2\)\}=0\.42, in close agreement with Eq\. \([3](https://arxiv.org/html/2606.17120#S2.E3)\)\. At each transition the effective rank of the end\-to\-end weight matrix increases by one, corresponding to the model learning an additional feature\. The small offset from the theoretical values reflects the bias of the output distribution\.

![Refer to caption](https://arxiv.org/html/2606.17120v1/x5.png)

Experimental MSE vs Regularization Strength

Figure 3:Phase transitions in a deep linear network \(L=3L=3,d=2d=2,η1=0\.9\\eta\_\{1\}=0\.9,η2=0\.8\\eta\_\{2\}=0\.8,Σx​x=𝐈\\Sigma\_\{xx\}=\\mathbf\{I\}\)\. Test MSE versus regularisation strengthβ\\betashows two sharp drops atβc\(1\)=0\.36\\beta\_\{c\}^\{\(1\)\}=0\.36andβc\(2\)=0\.42\\beta\_\{c\}^\{\(2\)\}=0\.42; at each transition the effective rank increases by one\. Quasi\-static sweep; small offset from theory reflects output distribution bias\.
## Appendix CBifurcation Diagram

The bifurcation structure of the loss landscape is shown in Fig\.[4](https://arxiv.org/html/2606.17120#A3.F4)\. For a single mode with signal strengthη\\eta, the stationarity condition Eq\. \([2](https://arxiv.org/html/2606.17120#S2.E2)\) has three solutions belowβc\\beta\_\{c\}: the trivial minimum atλ=0\\lambda=0, an unstable saddle atλsad\\lambda^\{\\mathrm\{sad\}\}, and the global minimum atλ∗∗\>λsad\\lambda^\{\*\*\}\>\\lambda^\{\\mathrm\{sad\}\}\. Asβ\\betaincreases throughβc\\beta\_\{c\}, the saddle and the global minimum merge and annihilate in a saddle\-node bifurcation, leaving onlyλ=0\\lambda=0\. A model on the upper \(rank\-1\) branch remains metastable for allβ<βc\\beta<\\beta\_\{c\}; a model on the lower \(rank\-0\) branch atλ=0\\lambda=0is trapped there even when the rank\-1 phase is energetically preferred \(i\.e\., forβ<β∗\\beta<\\beta^\{\*\}\)\. The height of the barrier between the two, and thus the mean escape time, grows asβ\\betaincreases towardβc\\beta\_\{c\}\. Note that forL=2L=2no such bifurcation occurs: the stationarity condition is linear inλ\\lambdaand has a unique positive solution for allβ\>0\\beta\>0, yielding only second\-order behaviour\[[3](https://arxiv.org/html/2606.17120#bib.bib3)\]\.

![Refer to caption](https://arxiv.org/html/2606.17120v1/x6.png)

Numerical MSE and Total Loss vs Regularization Strength

Figure 4:Bifurcation diagram for modes \(η1=0\.9\\eta\_\{1\}=0\.9,η2=0\.8\\eta\_\{2\}=0\.8,L=3L=3\)\. Solid lines: stable stationary points; dashed line: unstable saddle\. The two non\-trivial branches annihilate atβc\\beta\_\{c\}in a saddle\-node bifurcation, leaving onlyλ=0\\lambda=0forβ\>βc\\beta\>\\beta\_\{c\}\. A model on the rank\-1 branch \(dotted continuation\) remains metastable until stochastic noise drives it across the barrier into the rank\-2 global minimum\.
## Appendix DCritical Regularisation Strength

The stationary condition for modeiiunder the loss of Eq\. \([1](https://arxiv.org/html/2606.17120#S2.E1)\) is

∂ℒ∂λi=λi−ηi\+β​λi2/L−1=0\.\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\lambda\_\{i\}\}=\\lambda\_\{i\}\-\\eta\_\{i\}\+\\beta\\,\\lambda\_\{i\}^\{2/L\-1\}=0\.\(7\)
### D\.1Saddle\-node bifurcation for DNNs

Settingk=2/Lk=2/L, the stationarity condition reads

λi−ηi\+β​λik−1=0\.\\lambda\_\{i\}\-\\eta\_\{i\}\+\\beta\\,\\lambda\_\{i\}^\{k\-1\}=0\.\(8\)For a range ofβ\\beta, this equation admits two distinct positive solutions: a larger rootλ∗∗\\lambda^\{\*\*\}\(global minimum\) and a smaller rootλsad\\lambda^\{\\mathrm\{sad\}\}\(unstable saddle\)\. The two roots coalesce atβc\(i\)\\beta\_\{c\}^\{\(i\)\}, requiring both Eq\. \([8](https://arxiv.org/html/2606.17120#A4.E8)\) and its derivative to vanish simultaneously:

1\+β​\(k−1\)​λik−2=0\.1\+\\beta\(k\-1\)\\lambda\_\{i\}^\{k\-2\}=0\.\(9\)Solving Eq\. \([9](https://arxiv.org/html/2606.17120#A4.E9)\) gives the critical singular valueλc\(i\)=ηi​\(1−k\)/\(2−k\)\\lambda\_\{c\}^\{\(i\)\}=\\eta\_\{i\}\(1\-k\)/\(2\-k\), and substituting back into Eq\. \([8](https://arxiv.org/html/2606.17120#A4.E8)\) yields Eq\. \([3](https://arxiv.org/html/2606.17120#S2.E3)\) of the main text\.

### D\.2Number of transitions

Because the modes decouple, each non\-zero singular valueηi\\eta\_\{i\}ofΣy​x\\Sigma\_\{yx\}has its ownβc\(i\)\\beta\_\{c\}^\{\(i\)\}\. The orderingη1\>⋯\>ηd\>0\\eta\_\{1\}\>\\cdots\>\\eta\_\{d\}\>0impliesβc\(1\)\>⋯\>βc\(d\)\\beta\_\{c\}^\{\(1\)\}\>\\cdots\>\\beta\_\{c\}^\{\(d\)\}, givingddseparate saddle\-node bifurcations asβ\\betais decreased from a large value\.

## Appendix ETheoretical Barrier Heights

For a single mode atβ<βc\(i\)\\beta<\\beta\_\{c\}^\{\(i\)\}, the barrier for escape from the metastable lower rank phase is the loss difference between the saddle and the trivial minimum:

Δ​E=ℒ​\(λsad\)−ℒ​\(0\)=12​\(λsad−η\)2\+L​β2​\(λsad\)2/L−12​η2,\\Delta E=\\mathcal\{L\}\(\\lambda^\{\\mathrm\{sad\}\}\)\-\\mathcal\{L\}\(0\)=\\tfrac\{1\}\{2\}\(\\lambda^\{\\mathrm\{sad\}\}\-\\eta\)^\{2\}\+\\frac\{L\\beta\}\{2\}\(\\lambda^\{\\mathrm\{sad\}\}\)^\{2/L\}\-\\tfrac\{1\}\{2\}\\eta^\{2\},\(10\)whereλsad\\lambda^\{\\mathrm\{sad\}\}is the smaller positive root of Eq\. \([8](https://arxiv.org/html/2606.17120#A4.E8)\), obtained numerically\. Atβ=0\.32\\beta=0\.32,η2=0\.8\\eta\_\{2\}=0\.8,L=3L=3:λsad≈0\.41\\lambda^\{\\mathrm\{sad\}\}\\approx 0\.41andΔ​E≈0\.003\\Delta E\\approx 0\.003\. This is the minimal barrier along the lowest\-loss path out of the metastable phase\. The large discrepancy with the experimentally extractedΔ​Eeff=0\.15±0\.05\\Delta E\_\{\\mathrm\{eff\}\}=0\.15\\pm 0\.05is explained by the high\-dimensional corrections derived in Appendix[F](https://arxiv.org/html/2606.17120#A6)\.

## Appendix FEffective Barrier Height in High\-Dimensional Landscapes

### F\.1Kramers–Langer Theory

In a parameter space of dimensionDD, the mean first\-passage time from a metastable minimum to a saddle point under Langevin dynamics with noise strengthTeffT\_\{\\mathrm\{eff\}\}is given by the Kramers–Langer formula\[[13](https://arxiv.org/html/2606.17120#bib.bib13)\]:

τ=2​π\|ω1sad\|​\(∏j=1D\|ωjsad\|∏j=1D−1ωjmin\)1/2​exp⁡\(Δ​ETeff\),\\tau=\\frac\{2\\pi\}\{\|\\omega\_\{1\}^\{\\mathrm\{sad\}\}\|\}\\left\(\\frac\{\\prod\_\{j=1\}^\{D\}\|\\omega\_\{j\}^\{\\mathrm\{sad\}\}\|\}\{\\prod\_\{j=1\}^\{D\-1\}\\omega\_\{j\}^\{\\mathrm\{min\}\}\}\\right\)^\{\\\!\\\!1/2\}\\exp\\\!\\left\(\\frac\{\\Delta E\}\{T\_\{\\mathrm\{eff\}\}\}\\right\),\(11\)whereωjmin\>0\\omega\_\{j\}^\{\\mathrm\{min\}\}\>0are Hessian eigenvalues at the metastable minimum andωjsad\\omega\_\{j\}^\{\\mathrm\{sad\}\}those at the saddle, with exactly one negative eigenvalueω1sad<0\\omega\_\{1\}^\{\\mathrm\{sad\}\}<0\.

### F\.2Effective barrier and entropic contributions

Taking the logarithm of Eq\. \([11](https://arxiv.org/html/2606.17120#A6.E11)\) and defining

Δ​Eeff≡Δ​E−Teff2​∑j=1D−1ln⁡ωjsadωjmin,\\Delta E\_\{\\mathrm\{eff\}\}\\equiv\\Delta E\-\\frac\{T\_\{\\mathrm\{eff\}\}\}\{2\}\\sum\_\{j=1\}^\{D\-1\}\\ln\\\!\\frac\{\\omega\_\{j\}^\{\\mathrm\{sad\}\}\}\{\\omega\_\{j\}^\{\\mathrm\{min\}\}\},\(12\)the Arrhenius form is recovered exactly:ln⁡τ=const\+Δ​Eeff/Teff\\ln\\tau=\\mathrm\{const\}\+\\Delta E\_\{\\mathrm\{eff\}\}/T\_\{\\mathrm\{eff\}\}\. In our settingD=170D=170\. The metastable minimum is strongly confining in all directions, while the saddle has broader curvature in theD−1D\-1transverse directions, soωjmin≫ωjsad\\omega\_\{j\}^\{\\mathrm\{min\}\}\\gg\\omega\_\{j\}^\{\\mathrm\{sad\}\}for mostjj\. The resulting entropic contribution drivesΔ​Eeff≫Δ​Emin\\Delta E\_\{\\mathrm\{eff\}\}\\gg\\Delta E\_\{\\mathrm\{min\}\}, explaining the observed factor of∼50\\sim 50\.

### F\.3Correction from multiplicative SGD noise

Mori et al\.\[[19](https://arxiv.org/html/2606.17120#bib.bib19)\]showed that for MSE loss, the multiplicative nature of SGD noise modifies the relevant barrier from the linear loss difference to a logarithmised quantity\. For the approximately quadratic landscape near our metastable minimum, this correction does not alter the linear Arrhenius relationship, consistent withR2=0\.991R^\{2\}=0\.991\.

### F\.4Escape time and choice of time unit

Escape timesτ\\tauare reported in epochs\. SinceBB,ηlr\\eta\_\{\\mathrm\{lr\}\}, andNNare fixed within each sweep, the conversion between epochs and SGD steps is a constant \(see Appendix[A](https://arxiv.org/html/2606.17120#A1)\), leaving the extracted barrier height unchanged\.

## Appendix GGeneralisation to Nonlinear Networks

While our quantitative results \(barrier heights, criticalβ\\betavalues\) are specific to linear networks, the qualitative mechanism—metastable trapping due to saddle\-node bifurcations, followed by noise\-activated escape—applies generally\. Prior work has established that deep nonlinear networks exhibit identical first\-order L2\-driven phase transitions\[[3](https://arxiv.org/html/2606.17120#bib.bib3),[6](https://arxiv.org/html/2606.17120#bib.bib6)\], with hierarchical feature learning\[[7](https://arxiv.org/html/2606.17120#bib.bib7)\]\. The main ingredient for proposing the hysteresis is the underlying first\-order phase transition that exists beyond the linear setup\. The linear case thus serves as a minimal model that preserves the essential bifurcation structure while remaining analytically tractable\.

Similar Articles

The Weight Norm Sets the Grokking Timescale: A Causal Delay Law

arXiv cs.LG

This paper demonstrates that the weight norm causally controls the timescale of grokking in neural networks, reconciling conflicting accounts. Through interventions, it shows that grokking follows an exponential delay law and that norm magnitude dominates grokking time over learning rate across architectures.