Beyond Embedding Transfer: Component Roles in Grokking Transfer and Stability
Summary
This paper investigates component roles in warm-start transfer for grokking in neural networks, showing that transferring internal weights improves early performance but risks instability, and proposes methods to stabilize the process.
View Cached Full Text
Cached at: 09/17/26, 09:11 AM
# Beyond Embedding Transfer: Component Roles in Grokking Transfer and Stability
Source: [https://arxiv.org/html/2609.18078](https://arxiv.org/html/2609.18078)
###### Abstract
Warm\-start transfer can make algorithmic tasks generalize rapidly, yet it is unclear which model components provide the gain and whether that gain remains stable under continued optimization\. We study cross\-operator transfer on modular arithmetic and separate*efficacy*\(early velocity\) from*stability*\(post\-reach drawdown\)\. In a scale\-matched 108\-run battery across 12 seed blocks \(a 96\-run232^\{3\}factorial plus a 12\-run scale control\), transferring internal attention/MLP weight matrices \(BB\) in addition to token embeddings and readout \(E\+UE\+U\) improves early mean accuracy by5\.465\.46percentage points \(Holmp=0\.0039p=0\.0039\) and reduces confirmation latency by 558 steps \(Holmp=0\.0088p=0\.0088\)\. While readout plus internal\-block transfer satisfies the pre\-specified±500\\pm 500\-step mean\-latency equivalence criterion in one\-layer models \(TOSTp=0\.0011p=0\.0011, although Full is modestly faster in 11/12 paired seeds\), a prospective two\-layer replication confirms the internal\-block acquisition advantage \(12/12 seeds,\+704\.67\+704\.67integral units,p=4\.88×10−4p=4\.88\\times 10^\{\-4\}\) while revealing an architecture\-dependent boundary: omitting donor embeddings falls4475\.64475\.6integral units below Full transfer, far outside the pre\-specified±250\\pm 250\-unit equivalence margin\. Continued target training, however, frequently triggers severe post\-grokking relapse\. Freezing transferred representation carriers \(E,UE,U\) nearly eliminates offline relapse \(19\.40%→0\.07%19\.40\\%\\to 0\.07\\%, Holmp=0\.005859p=0\.005859\)\. Online validation\-triggered gating slashes True Max Drawdown from22\.06%22\.06\\%to0\.60%0\.60\\%on2a\+b2a\+b\(p=0\.000488p=0\.000488\), with prospective confirmations extending protection across affine, nonlinear quadratic, and two\-layer targets \(10\.94–23\.47 pp reductions\), while distinguishing continual stabilization from static early stopping\. In non\-abelianS5S\_\{5\}, unshielded transfer surges transiently \(95\.4%95\.4\\%peak\), but a prospective shielding cohort yields no confirmed benefit \(\+0\.15±1\.14\+0\.15\\pm 1\.14pp\)\. These results establish a component\-level dissociation between transfer acceleration and trajectory stability, and expose the empirical boundaries of parameter shielding\.
## 1Introduction
When trained on algorithmic and algebraic reasoning tasks with weight decay, overparameterized neural networks often exhibit*grokking*: training loss drops to zero almost instantly via memorization, yet held\-out generalization is delayed by thousands or tens of thousands of optimization steps\([Power et al\., 2022](https://arxiv.org/html/2609.18078#bib.bib11);[Liu et al\., 2022](https://arxiv.org/html/2609.18078#bib.bib7);[Liu et al\., 2023](https://arxiv.org/html/2609.18078#bib.bib8)\)\. Mechanistic interpretability has established that for modular arithmetic, delayed generalization corresponds to the gradual formation of structured Fourier multiplication circuits and circular representation manifolds in embedding and attention weights\([Nanda et al\., 2023](https://arxiv.org/html/2609.18078#bib.bib10);[Gromov, 2023](https://arxiv.org/html/2609.18078#bib.bib1);[Varma et al\., 2023](https://arxiv.org/html/2609.18078#bib.bib13)\)\.
Figure 1:Separating takeoff velocity from trajectory stability \(N=10N=10Blocks 230–239: 60\-run confirmatory matrix plus 10\-run exploratory extension\)\.\(a\) Test\-accuracy trajectories \(±\\pmSEM\)\. Full transfer reaches90%90\\%by step 50, versus 250–500 steps for representation conditions\. \(b\) FreezingE\+UE\+Ureduces post\-peak relapse from19\.40%19\.40\\%to0\.07%0\.07\\%\(Holmp=0\.005859p=0\.005859\)\. The exploratoryfull\_carrier\_frozenextension combines 50\-step takeoff with0\.02%0\.02\\%relapse\.Recently,[Xu et al\. \(2025\)](https://arxiv.org/html/2609.18078#bib.bib14)demonstrated that initializing target models with input embeddings learned by a weaker model accelerates generalization on algorithmic tasks\. In cross\-operator transfer, however, focusing solely on input representations leaves open whether internal attention/MLP weight matrices \(BB\) convey independent computational utility, and how warm\-started models evolve during continued optimization\. In practice, unshielded fine\-tuning frequently exhibits acute post\-reach generalization relapse, where initial gains are disrupted rather than monotonically refined\. This leads to three concrete questions:
1. 1\.Acquisition: Do internal attention/MLP weight matrices \(BB\) provide computational acceleration beyond representation carriers \(E,UE,U\), and does this requirement change across architecture depths?
2. 2\.Relapse: Why do warm\-started models relapse, and which components undergo contemporaneous degradation during post\-reach generalization collapse?
3. 3\.Protection: Can targeted parameter\-level interventions suppress relapse without degrading early velocity or terminal accuracy, and where do these protections break down?
Across these questions, our central insight is thattransfer acceleration and trajectory stability are governed by distinct component roles: internal attention/MLP weight matrices \(BB\) contribute strongly to acquisition velocity, while shielding representation carriers controls retention stability; the carrier requirements for acquisition become architecture\-dependent with depth\. Our contributions are:
1. 1\.Component Roles and Architecture Dependence: Scale\-matched factorial decomposition shows that internal\-block transfer provides substantial incremental acquisition acceleration \(\+5\.46 pp early accuracy, Holmp=0\.0039p=0\.0039; \+704\.67 units in 2\-layer replication,p=4\.88×10−4p=4\.88\\times 10^\{\-4\}\); however, the role of donor embeddings changes with depth, where omitting donor embeddings in two layers causes a massive−4475\.6\-4475\.6\-unit acquisition deficit\.
2. 2\.Acquisition and Stability as Distinct Lifecycle Properties: Fast warm\-start acquisition does not guarantee retention under continued optimization\. On a probed relapse, factorial state replacement localizes the contemporaneous functional deficit primarily to internal blocks, while linear readout refitting recovers 8\.37–11\.03 points at troughs, demonstrating substantial preserved linear decodability during apparent generalization collapse\.
3. 3\.Validation\-Triggered Carrier Protection and Its Boundaries: Representation carrier \(E\+UE\+U\) shielding strongly suppresses offline relapse \(19\.40%→0\.07%19\.40\\%\\to 0\.07\\%, Holmp=0\.005859p=0\.005859\) and online validation gating slashes max drawdown from22\.06%22\.06\\%to0\.60%0\.60\\%on2a\+b2a\+b\(p=0\.000488p=0\.000488\) with consistent gains on affine \(\+23\.47\+23\.47pp\) and quadratic \(\+10\.94\+10\.94pp\) tasks\. Crucially, we identify its empirical boundaries: shielding achieves confirmed attenuation but incomplete protection in two\-layer models \(\+15\.35\+15\.35pp reduction, residual drawdown65\.23%65\.23\\%\), and yields no confirmed benefit in non\-abelianS5S\_\{5\}\(\+0\.15±1\.14\+0\.15\\pm 1\.14pp\)\.
## 2Problem Formulation & Dual\-Track Protocol
### 2\.1Task Definitions & Architecture
We investigate algorithmic transfer across two canonical algebraic settings: \(1\)Abelian Modular Arithmetic \(p=113p=113\): The source donor task is modular additiona\+b\(mod113\)a\+b\\pmod\{113\}; the target task is the affine modular target2a\+b\(mod113\)2a\+b\\pmod\{113\}\. The sequence format is\[a,b\]\[a,b\], predicting labelc∈\{0,…,p−1\}c\\in\\\{0,\\dots,p\-1\\\}\. The dataset comprises113×113=12,769113\\times 113=12\{,\}769pairs, partitioned into20%20\\%train \(Ntrain=2,554N\_\{\\rm train\}=2\{,\}554\),10%10\\%validation \(Nval=1,277N\_\{\\rm val\}=1\{,\}277\), and70%70\\%test \(Ntest=8,938N\_\{\\rm test\}=8\{,\}938\) via deterministic rounding\. \(2\)Non\-Abelian Symmetric Group \(S5S\_\{5\}\): Source task is group multiplicationa∘ba\\circ b; target task is conjugationa∘b∘a−1a\\circ b\\circ a^\{\-1\}\. Sequences format two operand tokens\[a,b\]\[a,b\]wherea,b∈\{0,…,119\}a,b\\in\\\{0,\\dots,119\\\}, predicting permutation indexc∈\{0,…,119\}c\\in\\\{0,\\dots,119\\\}\. The domain contains120×120=14,400120\\times 120=14\{,\}400pairs, partitioned deterministically into30%30\\%train \(Ntrain=4,320N\_\{\\rm train\}=4\{,\}320\),20%20\\%validation \(Nval=2,880N\_\{\\rm val\}=2\{,\}880\), and50%50\\%test \(Ntest=7,200N\_\{\\rm test\}=7\{,\}200\)\. Models are standard pre\-LN Transformers\([Power et al\., 2022](https://arxiv.org/html/2609.18078#bib.bib11);[Nanda et al\., 2023](https://arxiv.org/html/2609.18078#bib.bib10)\)withdmodel=128,nhead=4,dmlp=512d\_\{\\rm model\}=128,n\_\{\\rm head\}=4,d\_\{\\rm mlp\}=512\. In full factor transfer \(or full 2D\-weight transfer, denotedfullorE\+U\+BE\+U\+B\), the network transfers token embeddings \(EE\), the readout matrix \(UU\), and internal attention/MLP weight matrices \(BB\) with per\-tensor Frobenius norm matching\. Positional embeddings, LayerNorm parameters, and MLP biases retain their native cold\-initialized values \(exact parameter tensor mappings and initialization conventions are detailed in Appendix[B](https://arxiv.org/html/2609.18078#A2)\)\. Optimization uses AdamW\([Loshchilov & Hutter, 2019](https://arxiv.org/html/2609.18078#bib.bib9)\)with base learning rateη=10−3\\eta=10^\{\-3\}, weight decayλ=1\.0\\lambda=1\.0, betasβ=\(0\.9,0\.98\)\\beta=\(0\.9,0\.98\), and batch size512512\. Computations in the confirmatory battery are executed in standard FP32 precision \(with PyTorch AMP utilized during earlier exploratory sweeps\)\. Under carrier freezing interventions, only token embeddingsWEW\_\{E\}and readout headWUW\_\{U\}are frozen \(∇θ=0\\nabla\_\{\\theta\}=0\); positional embeddings, LayerNorms, and internal blocks remain fully trainable\.
### 2\.2Per\-Tensor Scale Matching \(Decoupling Direction from Magnitude\)
To ensure that transfer benefits reflect parameter*direction*rather than arbitrary initialization*scale*, all target initializations in seed blockbbare paired with a trained, qualified donor checkpointWbdonorW\_\{b\}^\{\\rm donor\}and a native cold initializationWbcoldW\_\{b\}^\{\\rm cold\}\. A source run qualifies as a donor checkpoint only upon achieving stable grokking, defined by our prospectively specified gate as maintaining validation accuracy≥90%\\geq 90\\%across a continuous 1,000\-step persistence window \(21 consecutive evaluations at cadenceΔ=50\\Delta=50\)\. The donor checkpoint evaluated at the exact completion of this 21\-point persistence window is saved and used as the qualified donor, with donor optimization terminating immediately upon saving to prevent post\-hoc peak selection\. For any transferred parameter tensorθ\\theta, its direction is transferred while its Frobenius norm is rescaled to match the native cold initialization:θtransfer=\(θdonor/\(‖θdonor‖F\+10−12\)\)⋅‖θcold‖F\\theta^\{\\rm transfer\}=\(\\theta^\{\\rm donor\}/\(\\\|\\theta^\{\\rm donor\}\\\|\_\{F\}\+10^\{\-12\}\)\)\\cdot\\\|\\theta^\{\\rm cold\}\\\|\_\{F\}\. This controls for tensor\-level Frobenius norm scaling differences, isolating parameter direction from overall tensor magnitude\.
### 2\.3Dual\-Track Evaluation Metric Protocol
Prior transfer studies typically emphasize whether transfer*accelerates*learning, while the stability of those gains under continued optimization is less often separated explicitly\. We distinguish validation accuracyAval\(t\)A\_\{\\rm val\}\(t\)from held\-out test accuracyAtest\(t\)A\_\{\\rm test\}\(t\)\. Test accuracy is evaluated strictly for reporting and never controls triggering, optimization, hyperparameter selection, or confirmatory stopping decisions:
Figure 2:Component Localization and Factorial Synergy \(N=12N=12Formal Blocks 210–221: 96 Factorial Runs plus 12 Scale Controls, 108 Runs Total\)\.\(a\) Confirmation latencyRMST^30k\\widehat\{\\mathrm\{RMST\}\}\_\{30\\mathrm\{k\}\}across 8 factorial conditions \(medians and individual points\)\. Full transfer accelerates confirmation by\+558\.3\+558\.3steps over representation transferE\+UE\+U\(Holmp=0\.0088p=0\.0088\)\. Crucially, co\-transferring readout and internal blocks \(U\+BU\+B\) achieves TOST equivalence to full transfer \(p=0\.0011p=0\.0011,Δ=−16\.7\\Delta=\-16\.7steps\), while internal blocks alone \(BB\) do not establish a confirmation benefit within the 30k horizon \(p=0\.50p=0\.50\)\. \(b\) Pre\-5k accuracy integral, confirming that full transfer provides\+5\.46\+5\.46pp higher early accuracy thanE\+UE\+U\(Holmp=0\.0039p=0\.0039\)\.#### Efficacy Track \(Early Velocity\)\.
We quantify efficacy by: \(1\) first reach steptreach=min\{t:Aval\(t\)≥0\.90\}t\_\{\\rm reach\}=\\min\\\{t:A\_\{\\rm val\}\(t\)\\geq 0\.90\\\}; \(2\) pre\-5k accuracy integralI5k=∫05000Atest\(t\)𝑑t∈\[0,5000\]I\_\{\\rm 5k\}=\\int\_\{0\}^\{5000\}A\_\{\\rm test\}\(t\)\\,dt\\in\[0,5000\]; and \(3\) full\-trajectory mean test accuracyA¯=1T∫0TAtest\(t\)𝑑t\\bar\{A\}=\\frac\{1\}\{T\}\\int\_\{0\}^\{T\}A\_\{\\rm test\}\(t\)\\,dt\.
#### Stability Track \(Retention & Relapse\)\.
We quantify stability by: \(1\) 21\-point persistence confirmation \(H=1,000H=1\{,\}000steps at cadenceΔ=50\\Delta=50, requiringAval\(t\)≥0\.90A\_\{\\rm val\}\(t\)\\geq 0\.90across allK=21K=21consecutive evaluations\)\. For each seed blockbb, the confirmation latency isTbT\_\{b\}, truncated at the uniform experimental horizonτ=30,000\\tau=30\{,\}000steps asmin\(Tb,τ\)\\min\(T\_\{b\},\\tau\)\. Across a cohort ofNNblocks, the Restricted Mean Survival Time is estimated by the sample mean
RMST^τ=1N∑b=1Nmin\(Tb,τ\),\\widehat\{\\mathrm\{RMST\}\}\_\{\\tau\}=\\frac\{1\}\{N\}\\sum\_\{b=1\}^\{N\}\\min\(T\_\{b\},\\tau\),\(1\)assuming administrative censoring at the maximum budgetτ\\tau; \(2\) post\-peak relapse dipΔrelapse=Atest\(tpeak\)−mint\>tpeakAtest\(t\)\\Delta\_\{\\rm relapse\}=A\_\{\\rm test\}\(t\_\{\\rm peak\}\)\-\\min\_\{t\>t\_\{\\rm peak\}\}A\_\{\\rm test\}\(t\)\(primary stability metric for offline Phase C; maximum drawdown is supplementary\); and \(3\) for online streaming interventions \(§[5\.4](https://arxiv.org/html/2609.18078#S5.SS4)–[5\.5](https://arxiv.org/html/2609.18078#S5.SS5)\), models branch from an online validation trigger step:
ttrig=min\{t:Aval\(t−Δ\)≥q∧Aval\(t\)≥q\},t\_\{\\rm trig\}=\\min\\\{t:A\_\{\\rm val\}\(t\-\\Delta\)\\geq q\\land A\_\{\\rm val\}\(t\)\\geq q\\\},\(2\)requiring two consecutive validation evaluations above thresholdqqat cadenceΔ=50\\Delta=50\(withq=0\.95q=0\.95for modular arithmetic;q=0\.70q=0\.70forS5S\_\{5\}\)\. Over a pre\-specified finite observation horizonHHsteps post\-trigger \(H=1,500H=1\{,\}500for modular cohorts;H=3,000H=3\{,\}000forS5S\_\{5\}\), we evaluate trajectory stability via the post\-trigger True Max Drawdown:
DHtrig=maxttrig≤t≤ttrig\+H\[maxttrig≤s≤tAtest\(s\)−Atest\(t\)\]\.D\_\{H\}^\{\\rm trig\}=\\max\_\{t\_\{\\rm trig\}\\leq t\\leq t\_\{\\rm trig\}\+H\}\\left\[\\max\_\{t\_\{\\rm trig\}\\leq s\\leq t\}A\_\{\\rm test\}\(s\)\-A\_\{\\rm test\}\(t\)\\right\]\.\(3\)Because early\-stopping halts atttrigt\_\{\\rm trig\}, its post\-trigger drawdown is identicallyDHtrig≡0D\_\{H\}^\{\\rm trig\}\\equiv 0, matching Table[2](https://arxiv.org/html/2609.18078#S4.T2)\.
### 2\.4Statistical Inference and Family\-Wise Error Rate Control
All paired contrasts are evaluated across paired seed blocks using exact two\-sided sign\-flip permutation tests \(2N2^\{N\}complete enumerations:212=4,0962^\{12\}=4\{,\}096in Phase B;210=1,0242^\{10\}=1\{,\}024in Phase C\)\. In Phase B, family\-wise error rate is controlled atα=0\.05\\alpha=0\.05via the step\-down Holm\-Bonferroni procedure\([Holm, 1979](https://arxiv.org/html/2609.18078#bib.bib4)\)separately for the pre\-specified eight\-contrast efficacy family \(I5kI\_\{\\rm 5k\}, Table[1](https://arxiv.org/html/2609.18078#S3.T1)\) and the pre\-specified eight\-contrast stability family \(RMST30k\\mathrm\{RMST\}\_\{30\\mathrm\{k\}\}\), spanning cold vs\. individual components, component additions, and scale control, with confidence intervals estimated via 10,000\-resample Bias\-Corrected and Accelerated \(BCa\) bootstrapping\. For the primary contrast \(Full vs\.E\+UE\+U\), exact two\-sided sign\-flip permutation tests \(212=4,0962^\{12\}=4\{,\}096enumerations\) yield rawpperm=2/4096=0\.000488p\_\{\\rm perm\}=2/4096=0\.000488on efficacy; because the planned efficacy family containsm=8m=8contrasts, the initial Holm multiplier ism=8m=8, yieldingpHolm=8×0\.000488=0\.0039p\_\{\\rm Holm\}=8\\times 0\.000488=0\.0039\. For confirmation latency, rawpperm=0\.001221p\_\{\\rm perm\}=0\.001221yields step\-down adjustedpHolm=7×0\.001221=0\.0088p\_\{\\rm Holm\}=7\\times 0\.001221=0\.0088\. In the Phase C confirmatory battery, the superiority contrast family \(carrier freezing, differential learning rate, temporary freezing\) is evaluated under exact permutation tests with Holm\-Bonferroni correction \(α=0\.01\\alpha=0\.01pre\-specified criterion\) and 2,000\-resample BCa bootstrapping, while non\-inferiority \(representation\-only early velocity vs\. full transfer\) is evaluated via a one\-sided shifted permutation test atα=0\.05\\alpha=0\.05against a pre\-specified indifference margin \(δ=100\.0\\delta=100\.0integral units\)\. Equivalence in Phase B is assessed using a paired Student\-ttTwo One\-Sided Tests \(TOST\) procedure\([Schuirmann, 1987](https://arxiv.org/html/2609.18078#bib.bib12)\)against a pre\-specified±500\\pm 500\-step RMST margin \(rather than a permutation test\)\.
## 3What Information Is Transferred? Carrier Localization and Synergy
To identify component\-level contributions to transferred computation, Phase B executes a 108\-run battery across 12 formal blocks \(Blocks 210–221: a 96\-run232^\{3\}factorial ablation plus a 12\-run scale control\) ona\+b→2a\+b\(mod113\)a\+b\\to 2a\+b\\pmod\{113\}\. We decompose the model into: token embeddings \(EE\), readout classification head \(UU\), and internal transformer blocks \(BB, including multi\-head self\-attention and MLP layers\)\. Table[1](https://arxiv.org/html/2609.18078#S3.T1)and Figure[2](https://arxiv.org/html/2609.18078#S2.F2)report the complete findings\.
Table 1:Phase B Battery and Inferential Contrasts \(p=113p=113,N=12N=12paired blocks 210–221: 96 factorial runs plus 12 scale controls, 108 runs total\)\.Panel A \(Descriptive Factorial Summary\): Confirmation fraction \(ValAcc≥90%\\text\{ValAcc\}\\geq 90\\%sustained across1,0001\{,\}000steps\), mean confirmation latency \(RMST^30k\\widehat\{\\mathrm\{RMST\}\}\_\{30\\mathrm\{k\}\}\), pre\-5k accuracy integral \(I5kI\_\{\\rm 5k\}\), early mean test accuracy, and descriptive paired difference vs\. full transfer \(ΔRMST=RMSTcond−RMSTfull\\Delta\\mathrm\{RMST\}=\\mathrm\{RMST\}\_\{\\rm cond\}\-\\mathrm\{RMST\}\_\{\\rm full\}; positive values indicate slower confirmation than Full\)\. Uncertainties areMean±SEM\\text\{Mean\}\\pm\\text\{SEM\}\(N=12N=12\)\. The scale control condition \(embed\_orig\_scale\) achieves RMST18,025\.0±1,944\.418\{,\}025\.0\\pm 1\{,\}944\.4steps,I5k=1,131\.6±208\.5I\_\{\\rm 5k\}=1\{,\}131\.6\\pm 208\.5\.Panel B \(Prospectively Specified Inferential Contrasts and Equivalence\): Formal statistical tests across paired blocks with explicit algebraic estimand formulas\. Rawppermp\_\{\\rm perm\}is from exact two\-sided sign\-flip permutation tests \(212=4,0962^\{12\}=4\{,\}096enumerations\) with Holm\-Bonferroni FWER control \(α=0\.05\\alpha=0\.05\)\. For internal blocks alone \(BBvs\. Cold\), 10 of 12 blocks are mutually censored atτ=30,000\\tau=30\{,\}000steps \(Nnon\-zero=2N\_\{\\rm non\\text\{\-\}zero\}=2\), so the non\-crossing BCa CI reflects pervasive censoring rather than significant divergence from the exact null \(p=0\.5000p=0\.5000\)\. Equivalence forU\+BU\+Bvs\. Full is evaluated via a paired Student\-ttTwo One\-Sided Tests \(TOST\) procedure\([Schuirmann, 1987](https://arxiv.org/html/2609.18078#bib.bib12)\)against the pre\-specified margin \(±500\\pm 500steps;pTOST=0\.001089p\_\{\\rm TOST\}=0\.001089\); the 12 paired differences comprise nine\+100\+100\-step differences, one\+50\+50\-step difference \(Block 216\), one\+200\+200\-step difference \(Block 218\), and one−1,350\-1\{,\}350\-step difference \(Block 217; mean−16\.7\-16\.7steps\), establishing equivalence of average confirmation latency under the pre\-specified tolerance rather than identical per\-seed trajectories\.### 3\.1Internal Weight Matrices Provide Significant Incremental Acceleration
While transferring input and output representations \(E\+U\) substantially accelerates confirmation relative to cold start \(RMST1,733\.31\{,\}733\.3vs\.27,950\.027\{,\}950\.0\), transferring the full factor direction \(E\+U\+BE\+U\+B\) provides significant incremental advantages along both evaluation tracks: \(1\)Efficacy Track: Full transfer increases the pre\-5k integral by\+273\.0\+273\.0units \(95%95\\%BCa CI:\[229\.9,339\.7\]\[229\.9,339\.7\]\), representing an average test accuracy gain of\+5\.46\\mathbf\{\+5\.46\}percentage pointsover the first 5,000 steps \(Holmp=0\.0039<0\.01p=0\.0039<0\.01\)\. \(2\)Stability Track: Full transfer reduces continuous confirmation latency by\+558\.3\\mathbf\{\+558\.3\}steps\(95%95\\%BCa CI:\[270\.8,954\.2\]\[270\.8,954\.2\], Holmp=0\.0088<0\.01p=0\.0088<0\.01\)\. Both tracks show a significant incremental benefit under FWER control: transferring internal attention/MLP weight matrices \(BB\) improves learning beyondE\+UE\+Utransfer alone\.
### 3\.2Factorial Synergy: Readout Power and Internal Independence
Examining isolated component transfers reveals unexpected structural asymmetries:
- •The Readout Head Is a Strong Independent Carrier: Transferring the readout head alone \(U\) yields an integral of3,710\.63\{,\}710\.6units and confirmation in2,600\.02\{,\}600\.0steps, dramatically outperforming transferring the input embedding alone \(E:2,553\.72\{,\}553\.7units,4,220\.84\{,\}220\.8steps\)\.
- •B Alone Fails to Improve Confirmation Within the 30k Horizon: TransferringBBalone does not improve confirmation over cold start within the 30k horizon \(0/12 confirmed; RMST30,000\.030\{,\}000\.0vs\.27,950\.027\{,\}950\.0, exact sign\-flipp=0\.5000p=0\.5000\)\. Crucially, 10 of 12 blocks experienced mutual administrative censoring atτ=30,000\\tau=30\{,\}000steps, leaving only two non\-zero paired differences \(−7,850\-7\{,\}850and−16,750\-16\{,\}750steps\)\. While the descriptive BCa CI \(\[−8,649\.1,−654\.2\]\[\-8\{,\}649\.1,\-654\.2\]\) does not cross zero due to resampling predominantly from these two negative pairs, this does not overturn the non\-significant pre\-specified exact permutation test, reflecting pervasive censoring \(Nnon\-zero=2N\_\{\\rm non\\text\{\-\}zero\}=2\) rather than confirmed divergence\.
- •Scale Control Validates Magnitude Normalization: The scale control condition \(embed\_orig\_scale\) exhibits an RMST of18,025\.0±1,944\.418\{,\}025\.0\\pm 1\{,\}944\.4steps \(\+13,804\.2\+13\{,\}804\.2\-step latency penalty vs\. scale\-matchedE, Holmp=0\.0039p=0\.0039; Table[1](https://arxiv.org/html/2609.18078#S3.T1)\), demonstrating empirically that unnormalized transfer of learned embedding magnitudes severely hinders downstream optimization and validating our scale\-matching protocol\.
- •Non\-Additive Interaction: The descriptive 3\-way interactionIEUB=Ffull−FEU−FEB−FUB\+FE\+FU\+FB−F0I\_\{EUB\}=F\_\{\\rm full\}\-F\_\{EU\}\-F\_\{EB\}\-F\_\{UB\}\+F\_\{E\}\+F\_\{U\}\+F\_\{B\}\-F\_\{0\}is negative \(−3,141\.9\-3\{,\}141\.9integral units,p=0\.0005p=0\.0005\)\. BecauseI5kI\_\{\\rm 5k\}is bounded above by 5,000, we interpret this as ceiling\-compressed diminishing returns rather than a distinct mechanism\.
### 3\.3Bounded Latency Equivalence ofU\+BU\+Bvs\. Full
The most striking discovery from Phase B emerges from comparingU\+Bagainstfull\. Across the 12 paired blocks, individual differencesRMSTU\+B−RMSTfull\\mathrm\{RMST\}\_\{U\+B\}\-\\mathrm\{RMST\}\_\{\\rm full\}comprise nine blocks with\+100\+100steps \(where Full confirms at 1,050 andU\+BU\+Bat 1,150\), one block with\+50\+50steps \(Block 216\), one with\+200\+200steps \(Block 218\), and a single block with−1,350\-1\{,\}350steps \(Block 217, where Full experienced delayed confirmation at 2,500 whileU\+BU\+Bconfirmed at 1,150\)\. Across all 12 blocks,U\+Bachieves an average confirmation latency of1,158\.31\{,\}158\.3steps \(versus1,175\.01\{,\}175\.0for full transfer\), differing byΔ=−16\.7\\Delta=\-16\.7steps \(95%95\\%BCa CI:\[−500\.0,108\.3\]\[\-500\.0,108\.3\]\)\. A paired Student\-ttTOST against the pre\-specified margin\|Δ\|<500\|\\Delta\|<500steps rejects non\-equivalence \(using the conventional signs for the lower\- and upper\-bound tests,tL=3\.975t\_\{L\}=3\.975andtU=−4\.249t\_\{U\}=\-4\.249,df=11df=11,pTOST=0\.001089p\_\{\\rm TOST\}=0\.001089\), establishing operational equivalence of*average confirmation latency*within the pre\-specified±500\\pm 500\-step tolerance rather than computational parity or identical per\-seed trajectories\. The mean difference is influenced by one delayed Full\-transfer block \(Block 217\)\. Excluding this block yields a meanU\+B−FullU\+B\-\\mathrm\{Full\}latency difference of\+104\.5\+104\.5steps \(median\+100\+100\); the leave\-one\-out estimate remains comfortably within the pre\-specified±500\\pm 500\-step equivalence margin\. Across the 12 seeds, 11 of 12 exhibit positive differences \(two\-sided sign testp=0\.00635p=0\.00635\)\. Thus,U\+BU\+Bsatisfies the pre\-specified±500\\pm 500\-step operational mean\-latency equivalence criterion, although Full is modestly faster \(by≈100\\approx 100steps\) in 11/12 paired seeds; the average\-equivalence conclusion is therefore not driven by Block 217\. Within this setting and metric, early\-integral equivalence was not established\.
## 4Why Does Generalization Relapse? Bounded Mechanistic Diagnostics
Continued target training can produce severe accuracy relapses after initially successful generalization\. We combine diagnostics with distinct evidential scope: the state swap and displacement analyses examine an exploratory single probed trajectory \(Blocks 210/230\), while linear probes across four test\-conditioned relapse events \(three scratch, one transfer\) characterize contemporaneous decodability deficits rather than establishing a universal causal mechanism\. A single\-block trajectory shows concurrent embedding\-norm decline, readout\-norm growth, and subspace drift; these do not track accuracy monotonically and are thus observations rather than causes \(Appendix[A\.1](https://arxiv.org/html/2609.18078#A1.SS1)\)\. A242^\{4\}state\-swap on one25\.2425\.24\-point relapse exchanges embeddings \(EE\), readout \(UU\), internal blocks \(BB\), and remaining positional/LayerNorm parameters \(RR\)\. Restoring pre\-relapseBBraises trough accuracy from74\.76%74\.76\\%to99\.85%99\.85\\%, whereas restoringE\+UE\+Ureaches81\.45%81\.45\\%; injecting troughE\+UE\+Uinto the pre\-relapse state leaves accuracy at100%100\\%\. Thus, on this checkpoint pair, the functional deficit is concentrated primarily in internal blocks\.
From the same pre\-relapse state and optimizer history, freezing eitherE\+UE\+UorBBprevents the impending 50\-step dip on the probed trajectory\. Net internal\-block displacement over that interval is predominantly tangential \(≥88\.4%\\geq 88\.4\\%\), and applying the measured tangential component reproduces low accuracy whereas the radial component does not\. On this trajectory, these interventions implicate joint co\-adaptation and tangential internal\-block motion in the collapse\.
An exploratory, test\-conditioned probe analysis examines four qualified relapses selected from 12 trajectories \(three Scratch and oneE\+UE\+Utransfer; Appendix[A](https://arxiv.org/html/2609.18078#A1)\)\. Refitting a linear readout on trough features improves accuracy by8\.378\.37–11\.0311\.03points\. Trough probes are less accurate than their pre\-relapse counterparts at every tested calibration budget; atN=2554N=2554, deficits range from2\.662\.66to20\.0020\.00points\. Together, these probes show that trough representations retain recoverable target information but become substantially harder to decode linearly\. Cross\-initialization differences remain unresolved\.
Table 2:Online Dynamic Validation Gating Across Four Prospective Confirmatory Cohorts \(p=113p=113,H=1,500H=1\{,\}500Steps Post\-Trigger\)\.Evaluation of online carrier freezing upon reaching validation accuracy≥95%\\geq 95\\%across verification \(2a\+b2a\+b, 1\-layer, Blocks 250–261,N=12N=12\), prospective affine \(3a\+b3a\+b, Blocks 270–281,N=12N=12\), prospective nonlinear quadratic \(a2\+b2a^\{2\}\+b^\{2\}, Blocks 290–301,N=12N=12\), and prospective 2\-layer target \(2a\+b2a\+b, Blocks 310–321,N=12N=12\) cohorts, contrasted with unshielded AdamW and Trigger Snapshot \(early stopping\)\. Uncertainties areMean±SEM\\text\{Mean\}\\pm\\text\{SEM\}across seeds \(pflipp\_\{\\rm flip\}from exact two\-sided sign\-flip permutation tests\)\. By construction, Trigger Snapshot terminates optimization at the trigger, yielding strictly zero post\-trigger drawdown\. Offline shielding and geometric update ablations are detailed in §[5](https://arxiv.org/html/2609.18078#S5)–§[6](https://arxiv.org/html/2609.18078#S6)and the supplementary archive\.Target Cohort / PolicyTriggerttrigt\_\{\\rm trig\}True Max DD \(%\)Min Test Acc \(%\)Terminal Acc \(%\)Prospective Effect \(pflipp\_\{\\rm flip\}\)Cohort 1: Verification Target \(a\+b→2a\+ba\+b\\to 2a\+b,N=12N=12Blocks 250–261, 1\-Layer,H=1,500H=1\{,\}500Steps\)Trigger Snapshot \(Stop atttrigt\_\{\\rm trig\}\)504±48504\\pm 480\.00%\\mathbf\{0\.00\\%\}99\.21%±0\.14%\\mathbf\{99\.21\\%\\pm 0\.14\\%\}99\.21%±0\.14%99\.21\\%\\pm 0\.14\\%Early Stopping Baseline \(halted at trigger\)2a\+b2a\+bStandard AdamW504±48504\\pm 4822\.06%±2\.47%22\.06\\%\\pm 2\.47\\%77\.88%±2\.47%77\.88\\%\\pm 2\.47\\%99\.80%±0\.09%\\mathbf\{99\.80\\%\\pm 0\.09\\%\}Unshielded Relapse Baseline2a\+b2a\+bOnline FreezeE\+UE\+U504±48504\\pm 480\.60%±0\.17%\\mathbf\{0\.60\\%\\pm 0\.17\\%\}98\.94%±0\.20%\\mathbf\{98\.94\\%\\pm 0\.20\\%\}99\.36%±0\.11%99\.36\\%\\pm 0\.11\\%\+21\.45pp reduction\\mathbf\{\+21\.45\\text\{ pp reduction\}\},12/1212/12wins \(p=0\.000488∗∗p=0\.000488^\{\*\*\}\)Cohort 2: Prospective Affine Extension \(a\+b→3a\+ba\+b\\to 3a\+b,N=12N=12Fresh Blocks 270–281, 1\-Layer\)Trigger Snapshot \(Stop atttrigt\_\{\\rm trig\}\)1042±621042\\pm 620\.00%\\mathbf\{0\.00\\%\}98\.86%±0\.24%\\mathbf\{98\.86\\%\\pm 0\.24\\%\}98\.86%±0\.24%98\.86\\%\\pm 0\.24\\%Early Stopping Baseline \(halted at trigger\)3a\+b3a\+bStandard AdamW1042±621042\\pm 6228\.56%±2\.37%28\.56\\%\\pm 2\.37\\%71\.18%±2\.38%71\.18\\%\\pm 2\.38\\%98\.89%±1\.10%\\mathbf\{98\.89\\%\\pm 1\.10\\%\}Unshielded Relapse Baseline3a\+b3a\+bOnline FreezeE\+UE\+U1042±621042\\pm 625\.09%±1\.48%\\mathbf\{5\.09\\%\\pm 1\.48\\%\}94\.28%±1\.58%\\mathbf\{94\.28\\%\\pm 1\.58\\%\}97\.79%±0\.90%97\.79\\%\\pm 0\.90\\%Affine Confirmed:\+23\.47pp\+23\.47\\text\{ pp\},12/1212/12\(p=0\.000488∗∗p=0\.000488^\{\*\*\}\)Cohort 3: Prospective Nonlinear Quadratic Extension \(a\+b→a2\+b2a\+b\\to a^\{2\}\+b^\{2\},N=12N=12Fresh Blocks 290–301, 1\-Layer\)Trigger Snapshot \(Stop atttrigt\_\{\\rm trig\}\)229±7229\\pm 70\.00%\\mathbf\{0\.00\\%\}99\.87%±0\.07%\\mathbf\{99\.87\\%\\pm 0\.07\\%\}99\.87%±0\.07%99\.87\\%\\pm 0\.07\\%Early Stopping Baseline \(halted at trigger\)a2\+b2a^\{2\}\+b^\{2\}Standard AdamW229±7229\\pm 711\.09%±0\.99%11\.09\\%\\pm 0\.99\\%88\.91%±0\.99%88\.91\\%\\pm 0\.99\\%99\.99%±0\.01%\\mathbf\{99\.99\\%\\pm 0\.01\\%\}Unshielded Relapse Baselinea2\+b2a^\{2\}\+b^\{2\}Online FreezeE\+UE\+U229±7229\\pm 70\.15%±0\.02%\\mathbf\{0\.15\\%\\pm 0\.02\\%\}99\.80%±0\.07%\\mathbf\{99\.80\\%\\pm 0\.07\\%\}99\.91%±0\.02%99\.91\\%\\pm 0\.02\\%Quadratic Confirmed:\+10\.94pp\+10\.94\\text\{ pp\},12/1212/12\(p=0\.000488∗∗p=0\.000488^\{\*\*\}\)Cohort 4: Prospective 2\-Layer Target Architecture \(a\+b→2a\+ba\+b\\to 2a\+b,N=12N=12Fresh Blocks 310–321, 2\-Layer\)Trigger Snapshot \(Stop atttrigt\_\{\\rm trig\}\)1508±2551508\\pm 2550\.00%\\mathbf\{0\.00\\%\}97\.78%±0\.45%\\mathbf\{97\.78\\%\\pm 0\.45\\%\}97\.78%±0\.45%97\.78\\%\\pm 0\.45\\%Early Stopping Baseline \(halted at trigger\)2\-Layer Standard AdamW1508±2551508\\pm 25580\.58%±1\.86%80\.58\\%\\pm 1\.86\\%18\.82%±1\.77%18\.82\\%\\pm 1\.77\\%90\.23%±6\.30%90\.23\\%\\pm 6\.30\\%Unshielded Relapse Baseline2\-Layer Online FreezeE\+UE\+U1508±2551508\\pm 25565\.23%±2\.23%\\mathbf\{65\.23\\%\\pm 2\.23\\%\}33\.62%±2\.15%\\mathbf\{33\.62\\%\\pm 2\.15\\%\}98\.85%±0\.44%\\mathbf\{98\.85\\%\\pm 0\.44\\%\}Two\-Layer Confirmed:\+15\.35pp\+15\.35\\text\{ pp\},11/1211/12\(p=0\.00195∗∗p=0\.00195^\{\*\*\}\)
## 5Targeted Protection: Offline Shielding and Online Dynamic Gating
### 5\.1Intervention Formulation
Guided by our diagnostic findings, Phase C formulates targeted parameter\-level interventions designed to shield transferred computation against optimization collapse: \(1\)Carrier Freezing \(carrier\_frozen\): Representation matricesθ∈\{WE,WU\}\\theta\\in\\\{W\_\{E\},W\_\{U\}\\\}are permanently frozen \(∇θ=0\\nabla\_\{\\theta\}=0, weight decay deactivated:θ\(t\)≡θ\(0\)\\theta\(t\)\\equiv\\theta\(0\)\); positional embeddings, LayerNorms, and internal blocksWBW\_\{B\}remain fully trainable\. Because standard AdamW couples gradient steps and weight decay \(θt\+1=θt−ηgt−ηλθt\\theta\_\{t\+1\}=\\theta\_\{t\}\-\\eta g\_\{t\}\-\\eta\\lambda\\theta\_\{t\}\), parameter freezing bundles data\-gradient suppression with the arrest of weight decay erosion\. Our differential learning rate ablation directly tests whether soft attenuation suffices under the testedλ=1\.0\\lambda=1\.0schedule\. \(2\)Differential Learning Rate \(carrier\_lr\_0\.1\): Soft dampening withηcarrier=10−4\\eta\_\{\\rm carrier\}=10^\{\-4\}\(0\.1×0\.1\\timesbase LR\) and base weight decayλ=1\.0\\lambda=1\.0\. \(3\)Unconfounded Warmup Freezing \(freeze\_warmup\_500\_norest\): Representation carriers are frozen for 500 steps, then released without resetting optimizer state\.
### 5\.2Confirmatory Battery Results \(N=10N=10Independent Blocks 230–239, 60 Runs\)
We evaluate all interventions across 10 fresh, prospectively specified blocks \(60 runs×\\times10,000 steps\)\. Figure[1](https://arxiv.org/html/2609.18078#S1.F1)presents the trajectory dynamics and confirmatory results; detailed numerical summaries are reported below and archived in the supplementary materials\.
#### Relapse Suppression via Carrier Freezing\.
Under standard fine\-tuning \(E\+U\_standard\),100%100\\%\(10/10\) of runs suffer post\-reach relapse, with a mean post\-peak dip of19\.40%19\.40\\%and mean max drawdown of20\.34%20\.34\\%\. Permanently freezing the carriers \(E\+U\_carrier\_frozen\) virtually eliminates relapse: post\-peak relapse dip is slashed from19\.40%19\.40\\%to0\.07%\\mathbf\{0\.07\\%\}\(\+19\.33pp\+19\.33\\text\{ pp\}reduction,95%95\\%BCa CI:\[14\.91,24\.69\]\[14\.91,24\.69\], Holmp=0\.005859<0\.01p=0\.005859<0\.01\), and supplementary post\-reach max drawdown is slashed from20\.34%20\.34\\%to0\.09%\\mathbf\{0\.09\\%\}\(\+20\.26pp\+20\.26\\text\{ pp\}reduction,95%95\\%BCa CI:\[15\.75,24\.96\]\[15\.75,24\.96\], Holmp=0\.005859<0\.01p=0\.005859<0\.01\)\. Across all 10 independent blocks, the contrast is unanimously concordant \(10/10 pairs positive\), confirming carrier\-freezing relapse suppression\.
#### Tested Soft Interventions Fail to Suppress Relapse\.
Neither0\.1×0\.1\\timescarrier LR nor 500\-step temporary freezing significantly suppresses relapse under the tested schedules \(20\.80%20\.80\\%and14\.12%14\.12\\%mean dip, Holmp=0\.820p=0\.820andp=0\.242p=0\.242\); delayed relapse remains within the 10k\-step horizon\.
#### Early Velocity Non\-Inferiority Boundary\.
We tested whether representation\-only frozen transfer \(E\+U\_carrier\_frozen\) could match the early learning velocity of full\-factor standard transfer \(full\_standard\) within a pre\-specified non\-inferiority margin ofδ=100\.0\\delta=100\.0integral units\. The empirical mean gap is−144\.93\\mathbf\{\-144\.93\}units\(95%95\\%BCa CI:\[−164\.54,−127\.21\]\[\-164\.54,\-127\.21\]\)\. A one\-sided shifted permutation test yieldsp=0\.9990\>0\.05p=0\.9990\>0\.05, officially failing non\-inferiority\. As visible in Figure[1](https://arxiv.org/html/2609.18078#S1.F1)\(a\) inset,full\_standardcrosses90%90\\%accuracy at step 50, whereasE\+U\_carrier\_frozenrequires 250 steps\. Representation\-only shielding did not match the early efficacy of full transfer under the pre\-specified non\-inferiority criterion\.
### 5\.3Resolving the 4\-Quadrant Lifecycle: Takeoff and Stability Synergies
The confirmatory results separate two effects: internal\-parameter transfer improves early takeoff, whereas freezingE\+UE\+Uimproves stability\. In an exploratory 10\-run extension, combining full transfer with carrier freezing preserves 50\-step takeoff while reducing relapse to0\.02%0\.02\\%\.
Figure 3:Generality and Dynamical Boundaries\.\(a\) Non\-abelianS5S\_\{5\}dynamic protection \(N=12N=12prospective blocks, 9 triggered pairs\): carrier shielding shows no confirmed drawdown benefit \(ΔD=\+0\.15±1\.14\\Delta D=\+0\.15\\pm 1\.14pp, exactp=0\.895p=0\.895\)\. \(b\) Two\-layer component roles \(I5kI\_\{\\rm 5k\},N=12N=12fresh blocks\): internal blocks accelerate acquisition \(\+704\.67\+704\.67units,p<0\.001p<0\.001\), while omitting donor embeddings \(U\+BU\+B\) incurs a−4475\.63\-4475\.63\-unit deficit exceeding the pre\-specified±250\\pm 250\-unit equivalence margin\.
### 5\.4Online Causal Protection via Dynamic Validation Gating
While offline carrier freezing from initialization achieves stability, it requires deciding parameter shielding ahead of time\. To eliminate this constraint, we formulateOnline Dynamic Validation Gating: models begin with unshielded optimization; upon reaching validation accuracy≥95%\\geq 95\\%for two consecutive evaluations \(cadenceΔ=50\\Delta=50\), representation matricesWE,WUW\_\{E\},W\_\{U\}are dynamically frozen for subsequent training over a pre\-specified finite observation horizon ofH=1,500H=1\{,\}500steps post\-trigger\. We emphasize that our causal stability claim evaluates finite\-horizon trajectory robustness \(H=1,500H=1\{,\}500steps\); long\-term stability under unbounded post\-trigger optimization remains open\. After repairing optimizer\-state aliasing in an earlier implementation, we re\-evaluated the verification cohort \(Blocks 250–261,N=12N=12\) with deep\-copy isolation, state\-hash assertions, and execution\-order checks \(Table[2](https://arxiv.org/html/2609.18078#S4.T2), Cohort 1\): \(1\)Drawdown Suppression: Under standard training, models suffer severe post\-trigger collapse \(mean True Max Drawdown22\.06%22\.06\\%, min test accuracy77\.88%77\.88\\%\)\. Online carrier freezing slashes True Max Drawdown to0\.60%\\mathbf\{0\.60\\%\}\(median0\.42%0\.42\\%, min accuracy98\.94%98\.94\\%\), yielding an average reduction of\+21\.45\\mathbf\{\+21\.45\}percentage points\(95%95\\%BCa CI:\[\+17\.76,\+27\.06\]\[\+17\.76,\+27\.06\],12/1212/12concordant pairs, exact two\-sided sign\-flipp=0\.000488<0\.001p=0\.000488<0\.001; development cohort replicates with12/1212/12wins,20\.20%→1\.83%20\.20\\%\\to 1\.83\\%,p=0\.000488p=0\.000488\)\. \(2\)Geometric Constraints: Pinned direction with free norm \(EU\_radial\) yields37\.36%37\.36\\%drawdown and is worse than pinned norm with free direction \(EU\_tangential,5\.44%5\.44\\%\) in all 12 seeds\. Complete freezing yields0\.60%0\.60\\%drawdown and beats pinned norm in 9/12 seeds \(p=0\.005371p=0\.005371\)\. The permitted update subspaces therefore produce sharply different stability outcomes\. \(3\)Baseline Triad and Cohort Distinction: Post\-breakthrough degradation admits distinct operational solutions: \(i\)*Early Stopping*halts at the validation trigger \(ttrig=504±48t\_\{\\rm trig\}=504\\pm 48\), yielding99\.21%±0\.14%99\.21\\%\\pm 0\.14\\%test accuracy with strictly0\.00%0\.00\\%drawdown; \(ii\)*Retrospective Checkpointing*selects the peak validation model across the trajectory\. In offline Phase C \(Blocks 230–239\), unshielded transfer achieves peak test accuracy of99\.92%±0\.04%99\.92\\%\\pm 0\.04\\%\(step910±104910\\pm 104\) forE\+U\_standard\(despite a19\.40%19\.40\\%relapse dip\) and99\.97%±0\.03%99\.97\\%\\pm 0\.03\\%\(step85±885\\pm 8\) forfull\_standard\(despite a7\.21%7\.21\\%dip\)\. Offline retrospective selection and online streaming cohorts represent distinct operational regimes and must not be conflated as direct numerical competitors\. Early stopping and retrospective checkpointing address static model delivery; online carrier freezing addresses the distinct regime in which target optimization continues beyond the validation trigger, maintaining higher trajectory mean accuracy \(99\.55%99\.55\\%vs\.99\.37%99\.37\\%\) at a minor terminal cost \(−0\.44\-0\.44pp,99\.36%99\.36\\%vs\.99\.80%99\.80\\%\)\.
### 5\.5Prospective Confirmation Across Tasks and Architecture Depths
#### Affine Extension \(3a\+b\(mod113\)3a\+b\\pmod\{113\}\)\.
OnN=12N=12fresh blocks \(Blocks 270–281; all triggered at steps 750–1,350\), unshielded fine\-tuning undergoes acute relapse \(True Max Drawdown28\.56%28\.56\\%, range17\.58%∼45\.57%17\.58\\%\\sim 45\.57\\%, min accuracy71\.18%71\.18\\%\)\. Online carrier freezing slashes True Max Drawdown to5\.09%\\mathbf\{5\.09\\%\}\(median3\.55%3\.55\\%, min accuracy94\.28%\\mathbf\{94\.28\\%\}\), achieving an average reduction of\+23\.47\\mathbf\{\+23\.47\}percentage points\(95%95\\%BCa CI:\[\+19\.56,\+29\.37\]\[\+19\.56,\+29\.37\],12/1212/12concordant pairs,p=0\.000488<0\.01p=0\.000488<0\.01\), confirming the primary prospectively specified hypothesis\. The secondary boundary prediction \(≤2\.0%\\leq 2\.0\\%cohort mean drawdown\) was not supported \(cohort mean5\.09%5\.09\\%\)\. Terminal test accuracy under online freeze is97\.79%±0\.90%97\.79\\%\\pm 0\.90\\%\(−1\.10\-1\.10pp vs\. unshielded\), confirming that online freezing stabilizes continued downstream optimization\.
#### Nonlinear Quadratic Extension \(a2\+b2\(mod113\)a^\{2\}\+b^\{2\}\\pmod\{113\}\)\.
Prospective evaluation on the nonlinear quadratic targeta\+b→a2\+b2\(mod113\)a\+b\\to a^\{2\}\+b^\{2\}\\pmod\{113\}acrossN=12N=12fresh blocks \(Blocks 290–301; all triggered at steps 200–250,ttrig=229±7t\_\{\\rm trig\}=229\\pm 7\) confirms that unshielded models experience substantial post\-trigger collapse \(mean True Max Drawdown11\.09%±0\.99%11\.09\\%\\pm 0\.99\\%, range5\.83%∼16\.33%5\.83\\%\\sim 16\.33\\%, min accuracy88\.91%88\.91\\%\)\. Online carrier freezing slashes True Max Drawdown to0\.15%±0\.02%\\mathbf\{0\.15\\%\\pm 0\.02\\%\}\(range0\.01%∼0\.30%0\.01\\%\\sim 0\.30\\%, min accuracy99\.80%±0\.07%\\mathbf\{99\.80\\%\\pm 0\.07\\%\}\), achieving an average reduction of\+10\.94\\mathbf\{\+10\.94\}percentage points\(95%95\\%BCa CI:\[\+9\.13,\+12\.74\]\[\+9\.13,\+12\.74\],12/1212/12positive wins,p=0\.000488<0\.01p=0\.000488<0\.01\) with minimal terminal cost \(−0\.09\-0\.09pp,99\.91%99\.91\\%vs\.99\.99%99\.99\\%\)\. Prospective confirmation extends the protection effect beyond affine targets to a nonlinear quadratic operator \(a2\+b2\(mod113\)a^\{2\}\+b^\{2\}\\pmod\{113\}\)\.
#### Cross\-Depth Extension \(2\-Layer Target Architecture\)\.
Prospective evaluation confirms that the protective effect extends from one\- to two\-layer target models: qualified 1\-layer donors \(a\+ba\+b\) transferredE\+UE\+Udirection to a 2\-layer target \(2a\+b\(mod113\)2a\+b\\pmod\{113\},L=2,425,472L=2,425\{,\}472parameters, cold internal blocks\) acrossN=12N=12fresh blocks \(Blocks 310–321, triggered atttrig=1508±255t\_\{\\rm trig\}=1508\\pm 255\)\. Unshielded fine\-tuning suffers severe collapse \(True Max Drawdown80\.58%±1\.86%80\.58\\%\\pm 1\.86\\%, min accuracy18\.82%18\.82\\%\)\. Carrier shielding prospectively attenuates post\-breakthrough drawdown in two\-layer targets, reducing True Max Drawdown to65\.23%±2\.23%\\mathbf\{65\.23\\%\\pm 2\.23\\%\}\(\+15\.35\\mathbf\{\+15\.35\}percentage pointsreduction,95%95\\%BCa CI:\[\+9\.30,\+21\.01\]\[\+9\.30,\+21\.01\],11/1211/12concordant wins,p=0\.001953<0\.01p=0\.001953<0\.01\)\. Terminal accuracy was descriptively higher under freezing \(98\.85%98\.85\\%vs\.90\.23%90\.23\\%\), partly because several standard branches failed to recover within the observation window\. Carrier shielding reduces drawdown by15\.3515\.35pp at two layers, but the65\.23%65\.23\\%residual drawdown shows that protection is partial rather than complete\.
## 6Generality Boundaries: Non\-Abelian Groups and Layer Depth
### 6\.1Non\-Abelian Conjugation: Transient Transfer and Empirical Boundary
To probe boundaries beyond abelian arithmetic, Phase 1 evaluated transfer froma⋅ba\\cdot bto conjugationa⋅b⋅a−1a\\cdot b\\cdot a^\{\-1\}on non\-abelianS5S\_\{5\}\(\|S5\|=120\|S\_\{5\}\|=120, sequence format\[a,b\]\[a,b\], predicting permutation indexc=a∘b∘a−1c=a\\circ b\\circ a^\{\-1\},14,40014\{,\}400pairs total partitioned intoNtrain=4,320N\_\{\\rm train\}=4\{,\}320\[30%\],Nval=2,880N\_\{\\rm val\}=2\{,\}880\[20%\],Ntest=7,200N\_\{\\rm test\}=7\{,\}200\[50%\]; 60k steps; Figure[3](https://arxiv.org/html/2609.18078#S5.F3)\(a\)\)\. Transferred runs surge to peak accuracy of95\.39%\(Block 25\) and84\.62%\(Block 31\) versus cold start \(5\.83%5\.83\\%\), yet fail persistent confirmation \(0/40/4atq=0\.90q=0\.90\)\. In a prospectiveN=12N=12cohort \(Blocks 350–361\), triggering required two consecutive validation evaluations≥0\.70\\geq 0\.70at cadenceΔ=50\\Delta=50withinMpre=35,000M\_\{\\rm pre\}=35\{,\}000steps \(achieved by 9/12 blocks at71\.22%±0\.51%71\.22\\%\\pm 0\.51\\%\)\. Over horizonH=3,000H=3\{,\}000steps, carrier shielding showed no confirmed drawdown benefit on triggered pairs \(ΔD=\+0\.15±1\.14\\Delta D=\\mathbf\{\+0\.15\\pm 1\.14\}pp,95%95\\%BCa CI:\[−2\.17,\+2\.00\]\[\-2\.17,\+2\.00\]pp, exact sign\-flipp=0\.894531p=0\.894531\)\. In the intention\-to\-treat policy estimand, the 3 untriggered blocks continue unshielded \(ΔD=0\\Delta D=0\), yielding full\-cohort policy effectΔD=\+0\.11±0\.84\\Delta D=\+0\.11\\pm 0\.84pp \(95% BCa CI:\[−1\.72,\+1\.47\]\[\-1\.72,\+1\.47\]pp\), providing a cross\-family boundary case\.
### 6\.2Two\-Layer Component Roles: Acquisition and Embeddings
Following exploratory runs right\-censored at 30k steps \(p=0\.2500p=0\.2500\), we evaluated a 60\-run confirmatory battery acrossN=12N=12fresh blocks \(Blocks 330–341, 10k steps; comprising 12 donor training runs on 2\-layera\+ba\+bplus 48 target runs across 4 conditions: ‘cold‘, ‘full‘, ‘E\+U‘, and ‘U\+B‘; Figure[3](https://arxiv.org/html/2609.18078#S5.F3)\(b\)\)\. Donors are 2\-layer Transformers \(L=2,425,472L=2,425\{,\}472parameters\); transfer maps internal blocks layer\-by\-layer \(B0→B0,B1→B1B\_\{0\}\\to B\_\{0\},B\_\{1\}\\to B\_\{1\}\) alongside representation matrices \(WE,WUW\_\{E\},W\_\{U\}\), with cold Frobenius scale matching\. \(1\)Internal Acceleration: Internal blocks accelerate acquisition beyondE\+UE\+Uin two\-layer models \(12/1212/12wins,I5k=4600\.9±31\.3I\_\{\\rm 5k\}=4600\.9\\pm 31\.3vs3896\.3±73\.43896\.3\\pm 73\.4, gain\+704\.67\\mathbf\{\+704\.67\}units,95%95\\%BCa CI:\[\+549\.86,\+870\.00\]\[\+549\.86,\+870\.00\], exactp=0\.000488<0\.001p=0\.000488<0\.001; trigger onset1,283→1581\{,\}283\\to 158steps\)\. \(2\)Embedding Requirement: TransferringU\+BU\+Bwithout donor embeddings fails to preserve Full\-transfer acquisition \(I5k=125\.3±73\.7I\_\{\\rm 5k\}=125\.3\\pm 73\.7, deficit−4475\.63\\mathbf\{\-4475\.63\}units,95%95\\%BCa CI:\[−4570\.56,−4229\.47\]\[\-4570\.56,\-4229\.47\]; untriggered in 10/12 blocks\), falling far outside the±250\\pm 250\-unit equivalence margin \(TOSTp=1\.000000p=1\.000000\)\. Donor embeddings are thus required to preserve Full\-transfer acquisition in two\-layer models\.
## 7Related Work
Delayed generalization in modular arithmetic\([Power et al\., 2022](https://arxiv.org/html/2609.18078#bib.bib11);[Liu et al\., 2022](https://arxiv.org/html/2609.18078#bib.bib7)\)reflects trigonometric circuits\([Nanda et al\., 2023](https://arxiv.org/html/2609.18078#bib.bib10);[Varma et al\., 2023](https://arxiv.org/html/2609.18078#bib.bib13)\)and lazy\-to\-rich transitions\([Kumar et al\., 2024](https://arxiv.org/html/2609.18078#bib.bib6)\); scale matching isolates directional transfer\([Hagmann et al\., 2023](https://arxiv.org/html/2609.18078#bib.bib2)\)\. While[Xu et al\. \(2025\)](https://arxiv.org/html/2609.18078#bib.bib14)show that pre\-trained embeddings accelerate single\-task grokking, our factorial decomposition separates internal computation from representation carriers: internal blocks accelerate acquisition with transferred carriers, while donor embeddings are required in two\-layer models\. Closest to our stability analysis,[Janati et al\. \(2026\)](https://arxiv.org/html/2609.18078#bib.bib5)show that parameter freezing prevents circuit unlearning in single\-task grokking\. We investigate cross\-operator transfer, establish a dissociation between acquisition velocity and retention stability, and introduce dynamic validation gating to stabilize continual post\-trigger optimization\.
## 8Discussion, Limitations, and Conclusion
This work characterizes the algorithmic transfer lifecycle in grokking: \(1\)Acquisition: Internal\-block transfer accelerates acquisition beyond carriers, replicating at two layers \(\+704\.67\+704\.67units,p<0\.001p<0\.001\)\. Co\-transferring internal blocks and heads satisfies the pre\-specified mean\-latency equivalence criterion to Full in one\-layer models, whereas donor embeddings are required in two\-layer models \(U\+BU\+Bfalls4475\.64475\.6units below Full, TOSTp=1\.00p=1\.00\)\. \(2\)Relapse: Factorial state replacement localizes post\-transfer deficits to internal blocks; linear readout refitting recovers8\.378\.37–11\.0311\.03points without parameter retraining\. \(3\)Protection and Boundaries: Online carrier shielding suppresses acute relapse in one\-layer tasks \(\+21\.45\+21\.45,\+23\.47\+23\.47,\+10\.94\+10\.94pp\) and attenuates drawdown in two\-layer targets \(80\.58%→65\.23%80\.58\\%\\to 65\.23\\%,\+15\.35\+15\.35pp\), though residual instability remains at two layers\. While early stopping secures static accuracy \(≥98\.8%\\geq 98\.8\\%\), online carrier freezing stabilizes continual optimization beyond the trigger\. In contrast, prospectiveS5S\_\{5\}transfer did not reproduce shielding benefits \(ΔD=\+0\.15±1\.14\\Delta D=\+0\.15\\pm 1\.14pp, exactp=0\.895p=0\.895\), providing a cross\-family boundary case\. Whether component specialization generalizes to deeper architectures and non\-algorithmic domains remains open\. Together, these findings support a component\-level account of transfer acceleration and continual stability in algorithmic Transformers\.
Reproducibility Statement\.Architectures, splits, hyperparameters, protocols, and statistical tests are in the text and appendices; run artifacts contain deterministic seeds, checkpoints, and reproduction scripts\.AI Use Statement\.AI tools assisted with code development and manuscript editing\. The authors designed the study, audited numerical claims against saved artifacts, and remain responsible for the scientific content\.
## References
- Gromov \(2023\)Andrey Gromov\.Grokking modular arithmetic\.*arXiv preprint arXiv:2301\.02679*, 2023\.URL[https://arxiv\.org/abs/2301\.02679v1](https://arxiv.org/abs/2301.02679v1)\.
- Hagmann et al\. \(2023\)Michael Hagmann, Philipp Meier, and Stefan Riezler\.Towards inferential reproducibility of machine learning research\.*International Conference on Learning Representations \(ICLR\)*, 2023\.URL[https://arxiv\.org/abs/2302\.04054v7](https://arxiv.org/abs/2302.04054v7)\.
- Hendrycks & Gimpel \(2016\)Dan Hendrycks and Kevin Gimpel\.Gaussian error linear units \(gelus\)\.*arXiv preprint arXiv:1606\.08415*, 2016\.URL[https://arxiv\.org/abs/1606\.08415](https://arxiv.org/abs/1606.08415)\.
- Holm \(1979\)Sture Holm\.A simple sequentially rejective multiple test procedure\.*Scandinavian Journal of Statistics*, 6\(2\):65–70, 1979\.
- Janati et al\. \(2026\)Ali Janati, Kaoutar El Maghraoui, Andrei Kanavalau, and Anass Belfatmi\.Post\-grokking collapse at the representation\-readout interface in muon\-trained transformers\.*arXiv preprint arXiv:2608\.07436*, 2026\.URL[https://arxiv\.org/abs/2608\.07436](https://arxiv.org/abs/2608.07436)\.
- Kumar et al\. \(2024\)Tanishq Kumar, Blake Bordelon, Samuel J\. Gershman, and Cengiz Pehlevan\.Grokking as the transition from lazy to rich training dynamics\.*International Conference on Learning Representations \(ICLR\)*, 2024\.URL[https://arxiv\.org/abs/2310\.06110v3](https://arxiv.org/abs/2310.06110v3)\.
- Liu et al\. \(2022\)Ziming Liu, Ouail Kitouni, Niklas Nolte, Eric J\. Michaud, Max Tegmark, and Mike Williams\.Towards understanding grokking: An effective theory of representation learning\.*Advances in Neural Information Processing Systems*, 2022\.URL[https://arxiv\.org/abs/2205\.10343v2](https://arxiv.org/abs/2205.10343v2)\.
- Liu et al\. \(2023\)Ziming Liu, Eric J\. Michaud, and Max Tegmark\.Omnigrok: Grokking beyond algorithmic data\.*International Conference on Learning Representations \(ICLR\)*, 2023\.URL[https://arxiv\.org/abs/2210\.01117v2](https://arxiv.org/abs/2210.01117v2)\.
- Loshchilov & Hutter \(2019\)Ilya Loshchilov and Frank Hutter\.Decoupled weight decay regularization\.*International Conference on Learning Representations \(ICLR\)*, 2019\.URL[https://openreview\.net/forum?id=Bkg6RiCqY7](https://openreview.net/forum?id=Bkg6RiCqY7)\.
- Nanda et al\. \(2023\)Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt\.Progress measures for grokking via mechanistic interpretability\.*International Conference on Learning Representations \(ICLR\)*, 2023\.URL[https://arxiv\.org/abs/2301\.05217v3](https://arxiv.org/abs/2301.05217v3)\.
- Power et al\. \(2022\)Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra\.Grokking: Generalization beyond overfitting on small algorithmic datasets\.*arXiv preprint arXiv:2201\.02177*, 2022\.URL[https://arxiv\.org/abs/2201\.02177v1](https://arxiv.org/abs/2201.02177v1)\.
- Schuirmann \(1987\)Donald J\. Schuirmann\.A comparison of the two one\-sided tests procedure and the power approach for assessing the equivalence of average bioavailability\.*Journal of Pharmacokinetics and Biopharmaceutics*, 15\(6\):657–680, 1987\.
- Varma et al\. \(2023\)Vikrant Varma, Rohin Shah, Zachary Kenton, Janós Kramár, and Ramana Kumar\.Explaining grokking through circuit efficiency\.*Advances in Neural Information Processing Systems*, 2023\.URL[https://arxiv\.org/abs/2309\.02390v1](https://arxiv.org/abs/2309.02390v1)\.
- Xu et al\. \(2025\)Zhiwei Xu, Zhiyu Ni, Yixin Wang, and Wei Hu\.Let me grok for you: Accelerating grokking via embedding transfer from a weaker model\.*International Conference on Learning Representations \(ICLR\)*, 2025\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2025/hash/fa5ddd6bac0d665c72969d79221b680a\-Abstract\-Conference\.html](https://proceedings.iclr.cc/paper_files/paper/2025/hash/fa5ddd6bac0d665c72969d79221b680a-Abstract-Conference.html)\.
## Appendix ADiagnostic Linear Probing Specification, Optimization Convergence Checks, and Full Accounting
This appendix provides the full technical specification, optimization convergence checks, trajectory event accounting, and numerical results for the diagnostic linear probing investigation\.
### A\.1Diagnostic Figures
#### Grassmann Subspace Alignment Metric\.
In Figure[4](https://arxiv.org/html/2609.18078#A1.F4)\(b\), the alignment between the token embedding representationWE\(t\)∈ℝp×dmodelW\_\{E\}\(t\)\\in\\mathbb\{R\}^\{p\\times d\_\{\\rm model\}\}at stepttand the donor initializationWE\(0\)W\_\{E\}\(0\)is quantified via the mean canonical principal angle cosine across their leadingk=12k=12left singular subspaces:
cosθ¯\(t\)=1k∑i=1kσi\(Q\(t\)⊤Qdonor\)=1k∑i=1kcosθi,\\cos\\bar\{\\theta\}\(t\)=\\frac\{1\}\{k\}\\sum\_\{i=1\}^\{k\}\\sigma\_\{i\}\(Q\(t\)^\{\\top\}Q^\{\\rm donor\}\)=\\frac\{1\}\{k\}\\sum\_\{i=1\}^\{k\}\\cos\\theta\_\{i\},\(4\)whereQ\(t\)∈ℝp×kQ\(t\)\\in\\mathbb\{R\}^\{p\\times k\}andQdonor∈ℝp×kQ^\{\\rm donor\}\\in\\mathbb\{R\}^\{p\\times k\}are orthonormal bases spanning the top\-kkleft singular vectors ofWE\(t\)W\_\{E\}\(t\)andWE\(0\)W\_\{E\}\(0\)respectively, andσi\\sigma\_\{i\}are the singular values of their inner projection matrix\.
Figure 4:Exploratory trajectory and state\-swap diagnostics \(Blocks 210/230\)\.\(a–b\) Embedding/readout norms and embedding\-subspace alignment evolve during unshielded training; alignment is not monotone in accuracy\. \(c\) On one pre\-relapse/trough checkpoint pair, restoring internal blocks recovers99\.85%99\.85\\%accuracy, whereas restoringE\+UE\+Ureaches81\.45%81\.45\\%\. These single\-trajectory observations localize the contemporaneous functional deficit but do not identify a general cause\.Figure 5:Linear probes on four test\-conditioned relapse events\.Test accuracy for probes fitted to pre\-relapse, trough, and recovery representations across calibration budgets\. Trough probes improve with more data but remain below paired pre\-relapse probes under the tested protocol\. This finite\-budget deficit neither establishes irreversible information loss nor excludes a better decoder\.
### A\.2Protocol Specification
#### Motivation and Scope\.
Following recent analyses distinguishing circuit failure from circuit masking\([Janati et al\., 2026](https://arxiv.org/html/2609.18078#bib.bib5)\), we investigate whether post\-grokking relapse reflects an intrinsic loss of decodable target information within internal representations or a geometric misalignment between internal features and the readout classification head\. We evaluate an exploratory development cohort consisting of 12 complete training trajectories across Blocks 210–213 spanning three initialization conditions:
1. 1\.scratch: standard cold initialization \(WbcoldW\_\{b\}^\{\\rm cold\}\);
2. 2\.transfer\_EU: token embeddingsWEW\_\{E\}and classification headWUW\_\{U\}initialized from trained donor checkpoints with per\-tensor Frobenius scale matching \(θtransfer=\(θdonor/\(‖θdonor‖F\+10−12\)\)⋅‖θcold‖F\\theta^\{\\rm transfer\}=\(\\theta^\{\\rm donor\}/\(\\\|\\theta^\{\\rm donor\}\\\|\_\{F\}\+10^\{\-12\}\)\)\\cdot\\\|\\theta^\{\\rm cold\}\\\|\_\{F\}\), with internal transformer blocks and LayerNorms initialized cold;
3. 3\.transfer\_full: all 2D weight matrices \(WE,WUW\_\{E\},W\_\{U\}, and internal self\-attention and MLP weightsWBW\_\{B\}\) transferred with per\-tensor scale matching, while 1D LayerNorm parameters and positional embeddings remain cold\.
#### Representation Extraction and Consistency Verification\.
Representationsz∈ℝdmodelz\\in\\mathbb\{R\}^\{d\_\{\\rm model\}\}\(dmodel=128d\_\{\\rm model\}=128\) are extracted at the sequence\-level token positionz=hL\[:,−1,:\]z=h\_\{L\}\[:,\-1,:\]immediately following final LayerNorm, prior to the readout matrixWU∈ℝp×dmodelW\_\{U\}\\in\\mathbb\{R\}^\{p\\times d\_\{\\rm model\}\}\. For every evaluated checkpoint, we computationally asserted that applying the existing readout head tozzexactly reproduces the forward model logits:
maxx∈𝒟test‖WUz\(x\)−f\(x\)‖∞<10−5,\\max\_\{x\\in\\mathcal\{D\}\_\{\\rm test\}\}\\\|W\_\{U\}z\(x\)\-f\(x\)\\\|\_\{\\infty\}<10^\{\-5\},\(5\)guaranteeing that probed representations correspond precisely to the operational input to the classification head\.
#### Event Detection via Running Peak Drawdown\.
To avoid censoring early relapses occurring prior to the global trajectory maximum, events are tracked via running peak drawdown\. Denoting running peak accuracy byM\(t\)=maxs≤tAcc\(s\)M\(t\)=\\max\_\{s\\leq t\}\\mathrm\{Acc\}\(s\)after validation breakthrough \(ValAcc≥95%\\text\{ValAcc\}\\geq 95\\%for two consecutive checks at cadenceΔ=50\\Delta=50\):
1. 1\.Pre\-Relapse Peak \(tpret\_\{\\rm pre\}\): the running peak immediately preceding a drop;
2. 2\.Relapse Trough \(ttrought\_\{\\rm trough\}\): the local minimum where drawdownΔdip\(t\)=M\(t\)−Acc\(t\)≥10\.0\\Delta\_\{\\rm dip\}\(t\)=M\(t\)\-\\mathrm\{Acc\}\(t\)\\geq 10\.0pp andAcc\(t\)<90\.0%\\mathrm\{Acc\}\(t\)<90\.0\\%;
3. 3\.Post\-Recovery \(trecovt\_\{\\rm recov\}\): the first subsequent evaluation whereAcc\(t\)≥95\.0%\\mathrm\{Acc\}\(t\)\\geq 95\.0\\%\.
Events are searched within 2,000 steps after breakthrough\. Scratch runs have a 30,000\-step budget and transfer runs a 5,000\-step budget\. Runs without a qualifying drop are labeledCENSORED\_NO\_RELAPSE; runs without breakthrough areCENSORED\_UNBROKEN\. If a drop occurs without recovery, its Pre–Trough probes remain in the analysis and the recovery value is censored\.
#### Linear Probe Objective and Subsetting\.
At each checkpoint stage, an independent linear probeU∈ℝp×dmodelU\\in\\mathbb\{R\}^\{p\\times d\_\{\\rm model\}\}is fitted by minimizing regularized empirical cross\-entropy over a calibration subset𝒮N⊂𝒟train\\mathcal\{S\}\_\{N\}\\subset\\mathcal\{D\}\_\{\\rm train\}:
minU∈ℝp×dmodelℒ\(U\)=1N∑i=1NℒCE\(Uzi,yi\)\+λ2‖U‖F2,\\min\_\{U\\in\\mathbb\{R\}^\{p\\times d\_\{\\rm model\}\}\}\\mathcal\{L\}\(U\)=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathcal\{L\}\_\{\\rm CE\}\(Uz\_\{i\},y\_\{i\}\)\+\\frac\{\\lambda\}\{2\}\\\|U\\\|\_\{F\}^\{2\},\(6\)across sample budgetsN∈\{256,512,1024,2554\}N\\in\\\{256,512,1024,2554\\\}, whereN=2554N=2554is the full training partition\. Subsets𝒮256⊂𝒮512⊂𝒮1024⊂𝒮2554\\mathcal\{S\}\_\{256\}\\subset\\mathcal\{S\}\_\{512\}\\subset\\mathcal\{S\}\_\{1024\}\\subset\\mathcal\{S\}\_\{2554\}are constructed as strictly nested leading prefixes \(𝒟train\[:N\]\\mathcal\{D\}\_\{\\rm train\}\[:N\]\) under the canonical deterministic train/val/test split \(seedblock\+1000\\texttt\{block\}\+1000\)\. Candidate penaltiesλ∈\{0,10−4,10−3,10−2,10−1,1\}\\lambda\\in\\\{0,10^\{\-4\},10^\{\-3\},10^\{\-2\},10^\{\-1\},1\\\}are selected by validation cross\-entropy and the selected probe is evaluated on the test partition\. Because event times are selected using test accuracy, this is a test\-conditioned diagnostic rather than a globally blind test\.
### A\.3Optimization Convergence Checks
#### Random Stream Isolation\.
To prevent pseudo\-random coupling between trajectory execution and probe fitting, each training run used a dedicated PyTorch CUDA generator \(torch\.Generator\(device=’cuda’\)\) for mini\-batch sequencing\. Probe initialization weights were generated using an independent CPU generator initialized with a fixed seed \(seed=42\)\.
#### Solver and Gate\.
The 288 candidate fits use L\-BFGS with strong Wolfe line search \(maximum 500 iterations\)\. Of these, 286 satisfy the pre\-specified gate‖∇ℒ‖F<10−3\\\|\\nabla\\mathcal\{L\}\\\|\_\{F\}<10^\{\-3\}; the two exceptions occur atλ=1\\lambda=1,N=2554N=2554and are not selected\. All 48 validation\-selected probes pass the gate\. A separate Block\-210 implementation check compared Adam and L\-BFGS at one trough configuration and found a 0\.03\-point test\-accuracy difference; it was not a second solver applied to all fits\. Forλ=0\\lambda=0, we report a finite\-budget unregularized solution and do not claim a unique finite optimum\.
### A\.4Full Trajectory Event Accounting
Table[3](https://arxiv.org/html/2609.18078#A1.T3)reports the complete trajectory event accounting across all 12 runs\. Notably, all 3 broken from\-scratch trajectories underwent deep qualified relapses \(dips of18\.3718\.37,13\.8613\.86, and29\.9529\.95pp\), disproving the assumption that unshielded training without transfer converges smoothly\.
Table 3:Event Accounting and Trajectory Census Across 12 Runs \(Blocks 210–213\)\.Breakthrough is defined as the first step withValAcc≥95%\\text\{ValAcc\}\\geq 95\\%for two consecutive evaluations\. Pre\-relapse peak, relapse trough, and recovery steps are detected via running peak drawdown\. All 3 broken Scratch runs suffered deep qualified relapses \(≥10\\geq 10pp dip and test accuracy<90%<90\\%\), showing that scratch relapse occurs in this development cohort\.
### A\.5Full Linear Probe Decodability and Paired Deficits
Table[4](https://arxiv.org/html/2609.18078#A1.T4)presents test accuracies across all calibration budgetsN∈\{256,512,1024,2554\}N\\in\\\{256,512,1024,2554\\\}for the 4 qualified events\.
Table 4:Linear Probe Decodability Across Checkpoint Stages and Calibration Budgets\.Test accuracy \(%\) of probes trained on frozen penultimate representations across four test\-conditioned relapse events\. The penaltyλ\\lambdais selected by validation loss\. Paired trough deficit isAcctrough\(N\)−Accpre\(N\)\\mathrm\{Acc\}\_\{\\rm trough\}\(N\)\-\\mathrm\{Acc\}\_\{\\rm pre\}\(N\); re\-fit gain compares the full\-sample trough probe with the unadapted trough model\.Table 5:Complete Parameter Census ofGrokTransformer\(dmodel=128,nhead=4,dmlp=512,p=113d\_\{\\rm model\}=128,n\_\{\\rm head\}=4,d\_\{\\rm mlp\}=512,p=113\)\.Tensor keys match the exact state dictionary entries insrc/model\.py\. Component assignments correspond to the factorial decomposition \(EE: Token Embedding,UU: Readout Head,BB: Internal Blocks\)\. Transferred 2D weight matrices are rescaled using per\-tensor Frobenius norm matching against the target seed’s cold initialization\. Positional embeddings, LayerNorm parameters, and MLP biases strictly retain their cold\-initialized values in all conditions\.
## Appendix BModel Architecture, Parameter Keys, and Initialization Details
This appendix provides the complete architectural parameter census, exact PyTorch state\-dictionary keys, tensor dimensions, initialization distributions, and parameter transfer conventions for theGrokTransformermodel defined insrc/model\.pyand evaluated across all experiments\.
### B\.1Architectural Census and Tensor Inventory
Table[5](https://arxiv.org/html/2609.18078#A1.T5)details the parameter tensors of the 1\-layer Transformer architecture\. The model is a pre\-LayerNorm Transformer with hidden dimensiondmodel=128d\_\{\\rm model\}=128,nhead=4n\_\{\\rm head\}=4self\-attention heads \(dk=dv=32d\_\{k\}=d\_\{v\}=32\), and a 2\-layer MLP with hidden dimensiondmlp=512d\_\{\\rm mlp\}=512and GELU activations\([Hendrycks & Gimpel, 2016](https://arxiv.org/html/2609.18078#bib.bib3)\)\. The total parameter count is227,712227\{,\}712\.
#### Two\-Layer Architecture Census and Mapping\.
In the prospective cross\-depth and 2\-layer role experiments \(§[5\.5](https://arxiv.org/html/2609.18078#S5.SS5), §[6](https://arxiv.org/html/2609.18078#S6)\), the architecture instantiates two identical transformer blocks \(blocks\.0andblocks\.1\), each containing197,760197\{,\}760parameters \(196,608196\{,\}608in 2D weight matrices;1,1521\{,\}152in 1D LayerNorm and bias vectors\)\. Together with token embeddings \(14,72014\{,\}720\), positional embeddings \(512512\), final LayerNorm \(256256\), and linear readout head \(14,46414\{,\}464\), the complete 2\-layer model contains425,472\\mathbf\{425\{,\}472\}parameters \(422,912422\{,\}912in 2D matrices, of which422,400422\{,\}400are transferable 2D weights;2,5602\{,\}560in 1D vectors\)\. When transferring internal blocks \(BB\), 2D weight matrices are transferred isomorphically layer\-by\-layer \(B0→B0,B1→B1B\_\{0\}\\to B\_\{0\},B\_\{1\}\\to B\_\{1\}\), with each tensor rescaled using per\-tensor cold Frobenius norm matching\. All 1D parameters \(LayerNorm weights/biases and MLP biases in both blocks\) strictly retain native cold initialization\.
### B\.2Exact Module Conventions and Transfer Rules
#### Fused Attention Projections\.
Self\-attention projections are parameterized as a single fused linear layerblocks\.0\.attn\.qkv\.weightwithout bias, projecting directly fromdmodel→3×dmodeld\_\{\\rm model\}\\to 3\\times d\_\{\\rm model\}\. Query, key, and value representations are split along the output dimension\. The projection is followed by the unbiased linear output projectionblocks\.0\.attn\.out\.weight\. In factorial conditions transferring internal blocks \(BB,E\+BE\+B,U\+BU\+B, andfull\), both matrices are transferred from the qualified donor\.
#### LayerNorm and Bias Vectors Retain Cold Initialization\.
The architecture uses pre\-LayerNorm parameterization:ln1normalizes attention inputs,ln2normalizes MLP inputs, andln\_fnormalizes the final sequence representation prior to the linear readout head\. All LayerNorm weights and biases are initialized to standard constants \(𝟏\\mathbf\{1\}and𝟎\\mathbf\{0\}respectively\)\. Across the three LayerNorm modules, these comprise six one\-dimensional vectors \(three weights and three biases\)\. MLP biases \(blocks\.0\.mlp\.0\.biasandblocks\.0\.mlp\.3\.bias\) are initialized to𝟎\\mathbf\{0\}, adding two further one\-dimensional vectors\. Thus the eight one\-dimensional vectors together contain1,4081\{,\}408parameters\. Crucially, all 1D parameters are strictly excluded from transfer across all conditions: they retain their native cold initialization\.
#### Positional Embeddings and Vocabulary Structure\.
The embedding tabletoken\_emb\.weighthas vocabulary sizeV=p\+2=115V=p\+2=115for primep=113p=113, reserving indices for operands and structural delimiters\. The positional embedding tablepos\_emb\.weighthas capacity4×dmodel4\\times d\_\{\\rm model\}; inputs\[a,b\]\[a,b\]index positions00and11\. Positional embeddings are not transferred in any condition and retain cold Gaussian initialization\.
#### Per\-Tensor Frobenius Norm Scale Matching\.
Under any transfer condition, each transferred 2D parameter tensorθ\\thetainherits the direction of the qualified donor tensorθdonor\\theta^\{\\rm donor\}while matching the Frobenius norm of the target seed’s cold initialization:
θtransfer=θdonor‖θdonor‖F\+10−12⋅‖θcold‖F\.\\theta^\{\\rm transfer\}=\\frac\{\\theta^\{\\rm donor\}\}\{\\\|\\theta^\{\\rm donor\}\\\|\_\{F\}\+10^\{\-12\}\}\\cdot\\\|\\theta^\{\\rm cold\}\\\|\_\{F\}\.\(7\)In the control conditionembed\_orig\_scale,token\_emb\.weightis copied directly at the donor’s unnormalized scale \(θtransfer=θdonor\\theta^\{\\rm transfer\}=\\theta^\{\\rm donor\}\)\.Similar Articles
Where Grokking Happens: Distributed Utility and Fourier Recoding Without a Module Switch
This paper uses Transition Games to study the transition from memorization to generalization in Transformers, revealing distributed utility gain with attention bias and Fourier recoding rather than a module switch.
Feature Repulsion and Spectral Lock-in: An Empirical Study of Two-Layer Network Grokking
This empirical study validates theoretical findings on feature repulsion and spectral lock-in during the grokking phenomenon in two-layer neural networks, demonstrating how activation functions influence the transition from memorization to generalization.
Weight Decay Regimes in Grokking Transformers: Cheap Online Diagnostics
This paper investigates how weight decay acts as a control parameter for transitioning between memorization and generalization in transformers trained on modular arithmetic, and introduces two cheap online diagnostic metrics from attention activations that track these dynamics.
Quantifying the Memorization-to-Generalization Transition: Scaling Laws and Phase Structure in Grokking
This paper quantifies the memorization-to-generalization transition (grokking) in neural networks through scaling laws, revealing that data complexity is the primary driver of transition time compared to model capacity.
On the Residual Scaling of Looped Transformers: Stability and Transferability
This paper analyzes residual scaling in looped (weight-tied) transformers, showing that weight sharing requires stronger scaling (1/N) than standard residual networks, and derives a factored parameterization that enables hyperparameter transfer across loop counts without retuning.