Stored in Optimizer State, Valued by Later Training: A Causal Account of Subliminal Trait Transfer

arXiv cs.LG Papers

Summary

The paper proposes a two-stage mechanism for subliminal trait transfer in AI models, where optimizer state transports source perturbations and later training determines their behavioral value.

arXiv:2608.20442v1 Announce Type: new Abstract: Subliminal trait transfer allows a student model to acquire behavioral dispositions from teacher-generated data in which the trait is not semantically expressed. Recent work explains how such signals enter gradients, but not how they survive source removal or acquire different signs under later training. We treat parameters and optimizer moments as a single trainer state and derive an exact transport-valuation identity separating observer-independent propagation of the source perturbation from the value assigned by a future continuation and behavioral readout. State surgery identifies the first moment as a causal carrier. Transplanting it alone leaves parameters, hidden states, and outputs unchanged at the cut, yet source-free updates generate growing parameter and hidden-state differences; transplanting parameters with the first moment recovers the terminal behavioral response. Sending the same source-induced difference through matched futures produces negative, near-zero, and positive Qwen effects (-0.658, +0.008, and +0.658 seed means). This ordering recurs in all 12 Llama-3.2-1B seeds after eight updates, while state-difference norms remain nearly equal across routes. Both contrasts grow in every paired seed when the continuation extends to sixteen updates. A full-horizon costate predicts all 42 Qwen route-mean signs and all 21 resolved Llama ordinary-route signs. Observer-independent transport also replicates across Qwen, SmolLM2, and Llama, while the complete-state recurrence predicts physical, hidden, and fixed-head responses in non-LoRA MNIST systems, including CNNs trained with AdamW and momentum SGD. Together, these results identify a two-stage mechanism for subliminal trait transfer: optimizer state transports the source perturbation, and later training determines its behavioral value.
Original Article
View Cached Full Text

Cached at: 08/24/26, 04:29 AM

# Stored in Optimizer State, Valued by Later Training: A Causal Account of Subliminal Trait Transfer
Source: [https://arxiv.org/html/2608.20442](https://arxiv.org/html/2608.20442)
August 2026

###### Abstract

Subliminal trait transfer allows a student model to acquire behavioral dispositions from teacher\-generated data in which the trait is not semantically expressed\. Recent work explains how such signals enter the student’s gradients, but not how they survive source removal or acquire different behavioral signs under later training\. We treat parameters and optimizer moments as a single trainer state and derive an exact transport–valuation identity\. It separates observer\-independent propagation of the source perturbation \(transport\) from the value assigned by a future continuation and behavioral readout \(valuation\)\. State surgery identifies the optimizer’s first moment as a causal carrier\. Transplanting the first moment alone leaves the parameters, hidden states, and outputs unchanged at the cut, yet subsequent source\-free updates generate growing differences in parameters and hidden states; transplanting parameters together with the first moment recovers the terminal behavioral response\. Sending the same source\-induced state difference through matched futures produces negative, near\-zero, and positive effects on Qwen \(−0\.658\-0\.658,\+0\.008\+0\.008, and\+0\.658\+0\.658seed means\)\. This ordering recurs in all 12 Llama\-3\.2\-1B seeds after eight updates, while the resulting trainer\-state differences remain nearly equal in norm across routes\. Both contrasts grow in every paired seed when the continuation is extended to sixteen updates\. A full\-horizon costate predicts all 42 Qwen route\-mean signs and all 21 resolved Llama ordinary\-route signs\. Observer\-independent transport also replicates across Qwen, SmolLM2, and Llama, while the complete\-state recurrence predicts physical, hidden, and fixed\-head responses in non\-LoRA MNIST systems, including CNNs trained with AdamW and momentum SGD\. Together, these results identify a two\-stage mechanism for subliminal trait transfer: optimizer state transports the source perturbation, and later training determines its behavioral value\.

## 1Introduction

Subliminal trait transfer allows models to acquire behavioral dispositions that are invisible in their training data\[[12](https://arxiv.org/html/2608.20442#bib.bib1)\]\. Recent work traces how these signals can enter a student model’s gradients\[[44](https://arxiv.org/html/2608.20442#bib.bib2),[6](https://arxiv.org/html/2608.20442#bib.bib3)\]\. Yet their post\-gradient trajectory is not understood\. How does a brief perturbation survive when the source is removed, and how does subsequent training convert it into measurable behavior?

Investigating this post\-gradient trajectory reveals that optimizer states, such as momentum and AdamW moment buffers, act as delayed\-release carriers of source\-specific ancestry\. Prior optimization work establishes that optimizer buffers provide causal memory that modifies future updates\[[9](https://arxiv.org/html/2608.20442#bib.bib38),[45](https://arxiv.org/html/2608.20442#bib.bib37)\]\. Here, that memory retains part of the source perturbation after source removal and continues writing it into parameters during subsequent source\-free updates\.

Stored ancestry has no fixed behavioral sign\. We show that the optimization process separates what the trainer physically carries from what the model eventually expresses\. The same stored trace can yield positive, negative, or near\-zero behavioral effects depending on the future training path\. We formalize these two stages as*transport*and*valuation*\(Figure[1](https://arxiv.org/html/2608.20442#S1.F1)\)\.

To make this separation measurable, we apply discrete\-time adjoint sensitivity analysis\[[40](https://arxiv.org/html/2608.20442#bib.bib18),[21](https://arxiv.org/html/2608.20442#bib.bib19),[33](https://arxiv.org/html/2608.20442#bib.bib20)\]to the*complete*trainer state\. The resulting identity expresses total behavior change as a path integral, linking source perturbations propagated forward through training with future\-value sensitivities propagated backward from the final measurement\. This framework allows us to perform state surgery to identify the carrier and predict how its value changes across matched future forks\.

Figure 1:Transport and valuation in the complete trainer state\.A teacher\-output perturbation enters the shared gradient interface asdtd\_\{t\}, is carried in\(w,m,v\)\(w,m,v\), and is written from the first moment into parameters during source\-free updates\. From the same cut\-state ancestryAEA\_\{E\}, different future routes can produce negative, near\-zero, or positive behavioral effects\. The observer induces the backward future valueλt\\lambda\_\{t\}; its inner product withdtd\_\{t\}is signed work, whose path integral gives the endpoint response \(Eq\.[6](https://arxiv.org/html/2608.20442#S2.E6)\)\. Matrices are schematic; labels identify the tested links\.Contributions\.

1. 1\.Causal identification of the subliminal carrier\.Through targeted state surgery, we show that the first moment physically transports silent perturbations after source removal to generate delayed physical descendants, while parameters plus first moment recover terminal behavior\. The topology of this relay replicates across optimizers and architectures \(§[3](https://arxiv.org/html/2608.20442#S3)\)\.
2. 2\.Formal separation of transport and valuation\.Building on the known memory effects of optimization algorithms, we introduce a full\-trainer\-state adjoint decomposition that formally separates observer\-independent transport from continuation\- and readout\-dependent valuation \(§[2](https://arxiv.org/html/2608.20442#S2)\)\.
3. 3\.Prediction of route\-dependent behavior\.We demonstrate that different future training paths assign opposite behavioral signs to the same stored ancestry\. Our full\-horizon costate predictor correctly anticipates all resolved route\-dependent signs across multiple architectures, establishing that behavioral value depends on the future route rather than on ancestry alone \(§[4\.2](https://arxiv.org/html/2608.20442#S4.SS2), §[5](https://arxiv.org/html/2608.20442#S5)\)\.

## 2Setup and Decomposition

### 2\.1Trainer as a Dynamical System

Conditioned on the training stream, we model training as a deterministic dynamical system over a complete stateSt=\(wt,mt,vt,τt\)S\_\{t\}=\(w\_\{t\},m\_\{t\},v\_\{t\},\\tau\_\{t\}\): parametersww, first momentmm, second momentvv, and auxiliary deterministic stateτ\\tau\(such as the optimizer clock or schedule state\)\. Each step is

gt=Gt​\(St,xt\),St\+1=U⁡\(St,gt\),yT=O⁡\(ST\),g\_\{t\}=G\_\{t\}\(S\_\{t\},x\_\{t\}\),\\qquad S\_\{t\+1\}=U\(S\_\{t\},g\_\{t\}\),\\qquad y\_\{T\}=O\(S\_\{T\}\),\(1\)whereGGcomputes the gradient from state and dataxtx\_\{t\},UUis the optimizer update, andOOis the terminal behavioral observer\. In the LLM experiments,OOis a target\-minus\-reference candidate\-string log\-likelihood on a frozen prompt bank\. We call a source\-induced difference in trainer state, merged weights, or hidden activations a*physical descendant*, a definition that does not require a behavioral readout\. Prior analyses typically track onlywtw\_\{t\}\. We track the full state because optimizer slots can store the accumulated source trace, represented by the tangentqtq\_\{t\}below\. Routes forked from the same cut share the finite ancestry chordAEA\_\{E\}\(Table[3](https://arxiv.org/html/2608.20442#A1.T3)\)\.

We use the affine gradient\-port pathgtα,R=1−α2​gN,tR\+1\+α2​gO,tRg\_\{t\}^\{\\alpha,R\}=\\frac\{1\-\\alpha\}\{2\}g\_\{N,t\}^\{R\}\+\\frac\{1\+\\alpha\}\{2\}g\_\{O,t\}^\{R\}, whereα=−1\\alpha=\-1and\+1\+1represent the neutral and owl\-biased arms \(Appendix[M\.2](https://arxiv.org/html/2608.20442#A13.SS2)\)\. Thesource perturbationis the teacher\-induced gradient change at stepttwith trainer state fixed:

dt=∂Gt∂α\|state fixed\.d\_\{t\}=\\frac\{\\partial G\_\{t\}\}\{\\partial\\alpha\}\\bigg\|\_\{\\text\{state fixed\}\}\.\(2\)Compatible teacher–student output interfaces and token mappings can transmit weak output bias into student gradients without an explicit training target\.

Our decomposition uses discrete\-time adjoint sensitivity analysis—a framework developed across optimal control, automatic differentiation, hyperparameter optimization, and meta\-learning\[[40](https://arxiv.org/html/2608.20442#bib.bib18),[21](https://arxiv.org/html/2608.20442#bib.bib19),[33](https://arxiv.org/html/2608.20442#bib.bib20),[18](https://arxiv.org/html/2608.20442#bib.bib21),[17](https://arxiv.org/html/2608.20442#bib.bib22)\]—here applied to the full trainer state including optimizer moments\.

Metric\.SSE/ZERO≔∑i\(y^i−yi\)2/∑iyi2\\mathrm\{SSE/ZERO\}\\coloneqq\\sum\_\{i\}\(\\hat\{y\}\_\{i\}\-y\_\{i\}\)^\{2\}/\\sum\_\{i\}y\_\{i\}^\{2\}: the ratio of model error to zero\-predictor error \(≪1\\ll 1: signal captured;≥1\\geq 1: no better than predicting zero\)\. All symbols are collected in Appendix[A](https://arxiv.org/html/2608.20442#A1)\.

### 2\.2Forward: How Source Ancestry Accumulates

Letqt=∂St/∂αq\_\{t\}=\\partial S\_\{t\}/\\partial\\alphabe the sensitivity of trainer state to source strength\. Starting fromq0=0q\_\{0\}=0\(fixed initialization\), the exact recurrence is:

qt\+1=∂U∂S​qt\+∂U∂g​\(∂G∂S​qt\+dt\)\.\\boxed\{q\_\{t\+1\}=\\frac\{\\partial U\}\{\\partial S\}\\,q\_\{t\}\+\\frac\{\\partial U\}\{\\partial g\}\\left\(\\frac\{\\partial G\}\{\\partial S\}\\,q\_\{t\}\+d\_\{t\}\\right\)\.\}\(3\)Two terms contribute:dtd\_\{t\}, the new source perturbation injected at steptt; and∂G∂S​qt\\tfrac\{\\partial G\}\{\\partial S\}q\_\{t\}, the effect of accumulated state perturbation on the current gradient—past perturbations alter the parameters, which alter how subsequent gradients are computed\. The Jacobian∂G/∂S\\partial G/\\partial Sis sparse: gradients depend on the current parameterswtw\_\{t\}but not on the optimizer moments\(mt,vt\)\(m\_\{t\},v\_\{t\}\), so only theww\-rows ofqtq\_\{t\}feed back into the gradient; the moments influence subsequent gradients only indirectly, through the parameter updates they produce\.

### 2\.3Backward: Future Training Assigns Value

Given a terminal observerOO, define the costatepT=\(∂O/∂ST\)⊤p\_\{T\}=\(\\partial O/\\partial S\_\{T\}\)^\{\\top\}and its backward recurrence, together with the derived quantityλt\\lambda\_\{t\}that we call thefuture value:

λt=\(∂U∂g\)⊤​pt\+1,\\displaystyle\\boxed\{\\lambda\_\{t\}=\\left\(\\frac\{\\partial U\}\{\\partial g\}\\right\)^\{\\top\}p\_\{t\+1\},\}\(4\)pt=\(∂U∂S\)⊤​pt\+1\+\(∂G∂S\)⊤​λt\.\\displaystyle\\boxed\{p\_\{t\}=\\left\(\\frac\{\\partial U\}\{\\partial S\}\\right\)^\{\\top\}p\_\{t\+1\}\+\\left\(\\frac\{\\partial G\}\{\\partial S\}\\right\)^\{\\top\}\\lambda\_\{t\}\.\}\(5\)The future valueλt\\lambda\_\{t\}measures how much the final behavioral measurement would move if the gradient at stepttwere nudged slightly; it is computed backward from the observer, through the optimizer, into the gradient port \(gtg\_\{t\}, the valueUUreceives fromGG; Appendix[A](https://arxiv.org/html/2608.20442#A1)\)\. We reserve*costate*forptp\_\{t\}\(the standard adjoint variable in discrete optimal control\) and callλt\\lambda\_\{t\}the*future value*\.

### 2\.4The Response Identity

The inner product𝒲t=⟨λt,dt⟩\\mathcal\{W\}\_\{t\}=\\langle\\lambda\_\{t\},d\_\{t\}\\ranglemeasures the*work*that source perturbationdtd\_\{t\}does against the future valueλt\\lambda\_\{t\}\. Along a declared source pathPαP\_\{\\alpha\}connecting the neutral arm \(α=−1\\alpha=\-1\) to the trait arm \(α=\+1\\alpha=\+1\), the total behavior change is:

OT​\(α=\+1\)−OT​\(α=−1\)=∫−11∑t=0T−1⟨λt​\(α\),dt​\(α\)⟩​𝑑α\.\\boxed\{O\_\{T\}\(\\alpha\{=\}\{\+\}1\)\-O\_\{T\}\(\\alpha\{=\}\{\-\}1\)=\\int\_\{\-1\}^\{1\}\\sum\_\{t=0\}^\{T\-1\}\\langle\\lambda\_\{t\}\(\\alpha\),d\_\{t\}\(\\alpha\)\\rangle\\,d\\alpha\.\}\(6\)Under the smooth or path\-differentiable conditions in Appendix[F](https://arxiv.org/html/2608.20442#A6), the identity is exact along any absolutely continuous declared source path\. Numerical quadrature and path\-dependent allocations are examined in Appendices[M\.2](https://arxiv.org/html/2608.20442#A13.SS2)and[M\.3](https://arxiv.org/html/2608.20442#A13.SS3)\.

The forward source and backward value meet only through their inner product\. The same stored perturbation can therefore produce positive, negative, or zero behavioral change\. Its sign belongs to the pairing, not to either factor alone\. At first order, for two routesR,R′R,R^\{\\prime\}sharing the same finite ancestry chordAEA\_\{E\}at cutEE:

Δ​OTR−Δ​OTR′≈\(pER​\(0\)−pER′​\(0\)\)⊤​AE,\\Delta O\_\{T\}^\{R\}\-\\Delta O\_\{T\}^\{R^\{\\prime\}\}\\approx\\bigl\(p\_\{E\}^\{R\}\(0\)\-p\_\{E\}^\{R^\{\\prime\}\}\(0\)\\bigr\)^\{\\top\}A\_\{E\},\(7\)whereAE=SE​\(\+1\)−SE​\(−1\)A\_\{E\}=S\_\{E\}\(\+1\)\-S\_\{E\}\(\-1\)andpER​\(0\),pER′​\(0\)p\_\{E\}^\{R\}\(0\),p\_\{E\}^\{R^\{\\prime\}\}\(0\)are the route costates on the midpoint trajectory\. In practice, the computation runs forward to store the trajectory, then backward to compute costates, and finally accumulates the per\-step inner products𝒲t\\mathcal\{W\}\_\{t\}; Algorithm[1](https://arxiv.org/html/2608.20442#alg1)\(Appendix[E](https://arxiv.org/html/2608.20442#A5)\) gives the complete pseudocode\. The same update mapUUcovers plain SGD \(S=wS=w\), momentum SGD \(S=\(w,mvel\)S=\(w,m\_\{\\mathrm\{vel\}\}\)\), and AdamW\[[32](https://arxiv.org/html/2608.20442#bib.bib46)\]\(S=\(w,m,v\)S=\(w,m,v\)\); the bias\-correction schedule is handled exactly \(Appendix[D](https://arxiv.org/html/2608.20442#A4)\)\.

Validation\.The tangent recurrence is validated across 360 step×\\timesblock cells on Qwen2\.5\-0\.5B\[[42](https://arxiv.org/html/2608.20442#bib.bib47)\]\(SSE/ZERO=1\.65×10−4\\mathrm\{SSE/ZERO\}=1\.65\\times 10^\{\-4\}, correlation 0\.9999\)\. A backward\-rotation experiment preserves forward outputs, loss, gradient norm, and singular spectrum while frozen\-head accuracy falls, linear\-probe accuracy remains stable, and Procrustes alignment restores access: the intervention changes where information is written rather than destroying it \(Appendix[B](https://arxiv.org/html/2608.20442#A2)\)\.

## 3Optimizer State Carries the Perturbation

Source influence remains active even after the source is removed from the data\. A first\-moment difference is invisible at the cut yet still creates a descendant under later source\-free updates\. The allocation and state surgery below locate this delayed effect\.

At lag 24, the local adjoint allocation places 89\.9% of the signed source contribution in the first moment, 10\.6% in parameters, and−0\.5\-0\.5% in the second moment\. Momentum SGD exhibits the same delayed handoff, whereas plain SGD writes parameters directly \(Figure[2](https://arxiv.org/html/2608.20442#S3.F2); Appendix[L](https://arxiv.org/html/2608.20442#A12)\)\.

Figure 2:Optimizer relay\.\(a\)In one Qwen isolated\-pulse cell, 89\.9% of the adjoint source\-family allocation at lag 24 falls in the first\-moment block\.\(b\)Momentum SGD: between the source end and the terminal update, the velocity\-origin norm falls while its parameter descendant grows\.\(c\)The recurrence attains low protocol\-specific prediction error under AdamW, momentum SGD, and plain SGD; plain SGD reachesSSE/ZERO<10−12\\mathrm\{SSE/ZERO\}<10^\{\-12\}\.Block transplant, reset, and rescue\.Because the first moment carries causal information beyond the parameters, addingmmimproves recovery relative to the corresponding transplant withoutmm\. Anmm\-only transplant is silent at the cut but generates a descendant under subsequent source\-free updates\. At a common Qwen checkpoint we copy any subset of the three state blocks \(ww,mm,vv\) from a source arm into its matched neutral\-source arm, reset the remaining blocks to the neutral arm’s values, and run all hybrids through the same source\-free suffix \(the*rescue*\)\. A subset is sufficient if it reproduces the full descendant; a block carries causal information if its inclusion changes the descendant relative to the matched subset without it\.

Among proper subsets,w\+mw\{\+\}mreproduces the full descendant \(SSE/ZERO=0\.005\\mathrm\{SSE/ZERO\}=0\.005–0\.0060\.006physical,0\.0020\.002–0\.0030\.003hidden\), whilewwalone,mmalone, and especiallyvvalone do not\. This physical\-endpoint ranking holds in 14/14 seed\-route cells \(Figure[3](https://arxiv.org/html/2608.20442#S3.F3)a–b\)\. Themm\-only transplant has exactly zero parameter and hidden effect at the cut, then generates a growing descendant as the source\-free rescue proceeds \(Figure[3](https://arxiv.org/html/2608.20442#S3.F3)c\)\. In nine behavioral cells,w\+mw\{\+\}malso recovers the full terminal response \(trait\-meanSSE/ZERO=0\.0017\\mathrm\{SSE/ZERO\}=0\.0017–0\.00790\.0079\) and beatsww,mm, and the reset\-mmhybridw\+vw\{\+\}v\(Appendix Table[17](https://arxiv.org/html/2608.20442#A16.T17)\)\.

Figure 3:Block transplant, reset, and rescue\.\(a\)In the three\-seed discovery cohort,w\+mw\{\+\}mreproduces the full descendant whilevv\-only remains at the no\-transplant baseline\.\(b\)Four replication seeds reproduce the ordering\. Error bars in\(a\)–\(b\)are sample SD across seeds\.\(c\)Themm\-only transplant has exactly zero forward\-visible effect at the cut and is then released by source\-free updates; error bars are sample SD\.Cross\-family replication\.The same surgery on Llama\-3\.2\-1B\[[35](https://arxiv.org/html/2608.20442#bib.bib50)\]with a model\-native source selectsw\+mw\{\+\}macross all tested seeds, withmm\-only again exactly forward\-invisible at the cut; under momentum\-SGD the structure persists with velocity in place of the first moment\. Adam’s first moment and classical momentum velocity are two implementations of the same relay \(Appendix[M\.6](https://arxiv.org/html/2608.20442#A13.SS6)\)\.

## 4Future Training Determines the Sign

The decomposition predicts that the same source ancestry can produce opposite behavioral signs under different future training \(Eq\.[7](https://arxiv.org/html/2608.20442#S2.E7)\)\. We test this first with engineered routes in a matched factorial, then with ordinary training continuations \(§[4\.2](https://arxiv.org/html/2608.20442#S4.SS2)\)\.

### 4\.1Matched Source×\\timesRoute Factorial

In Qwen2\.5\-0\.5B \(LoRA\-r8, AdamW\), the owl and neutral arms train on paired bare\-number completions from an owl\-biased LoRA teacher and the base model, respectively, so they differ only in the source signal carried by the data\. Once the source ends, each cut state is sent through three source\-free routes, giving a full2×32\\times 3factorial per seed\. The routes induce contrasting future alignments with the stored perturbation:S\+S^\{\+\}andS−S^\{\-\}use opposing diagnostic directions, whileznullz\_\{\\mathrm\{null\}\}is orthogonal to both\. Their geometry uses the readout and corpus\-gradient directions, but no costate or endpoint outcome \(Appendix[M\.3](https://arxiv.org/html/2608.20442#A13.SS3)\)\.

Figure 4:Future training assigns signed value\.\(a\)In the Qwen matched factorial, identical ancestry is negative underS\+S^\{\+\}, near\-zero underznullz\_\{\\mathrm\{null\}\}, and positive underS−S^\{\-\}; bars and error bars are seven\-seed means and sample SD\.\(b\)In a separate route\-replacement cohort, the pairedS−−S\+S^\{\-\}\{\-\}S^\{\+\}contrast is positive in 7/7 independent seeds \(two\-sided sign\-testp=0\.016p=0\.016; band: 95% Student\-ttCI\)\.\(c\)The topology recurs across four source positions\.\(d\)The seed\-mean topology recurs across all five trait families on their respective observer scales; dots are seeds and horizontal marks are means on a symmetric\-log scale \(Table[18](https://arxiv.org/html/2608.20442#A16.T18)\)\.Across independent seeds, the cells \(Figure[4](https://arxiv.org/html/2608.20442#S4.F4)a\) give route\-conditioned source effectsΔR=Osource,R−Oneutral,R\\Delta\_\{R\}=O\_\{\\mathrm\{source\},R\}\-O\_\{\\mathrm\{neutral\},R\}\. The same ancestry produces opposite seed\-mean effects underS\+S^\{\+\}andS−S^\{\-\}\(ΔS\+=−0\.658\\Delta\_\{S^\{\+\}\}=\-0\.658,ΔS−=\+0\.658\\Delta\_\{S^\{\-\}\}=\+0\.658\), whileznullz\_\{\\mathrm\{null\}\}is near zero \(Δznull=\+0\.008\\Delta\_\{z\_\{\\mathrm\{null\}\}\}=\+0\.008\)\. Because every route has its own matched neutral control, route\-only main effects cancel within eachΔR\\Delta\_\{R\}; differences among theΔR\\Delta\_\{R\}isolate the interaction between route and source—the effect of the same ancestry*depending on*which future follows it\. The theory predictsd1=Δznull−ΔS\+\>0d\_\{1\}=\\Delta\_\{z\_\{\\mathrm\{null\}\}\}\-\\Delta\_\{S^\{\+\}\}\>0andd2=ΔS−−Δznull\>0d\_\{2\}=\\Delta\_\{S^\{\-\}\}\-\\Delta\_\{z\_\{\\mathrm\{null\}\}\}\>0; both contrasts are consistently positive in 7/7 independent seeds \(one\-sided sign test, Holm\-adjustedp=0\.016p=0\.016each\)\. Theznullz\_\{\\mathrm\{null\}\}arm retains a sizeable parameter descendant yet acquires almost no signed value\. An independent paraphrase bank replicates the full ordering across all independent seeds\. The identical factorial on four additional trait families \(blue, cat, red, oak\) reproduces the same seed\-mean route topology in every case, with statistical resolution varying by trait \(Table[18](https://arxiv.org/html/2608.20442#A16.T18)\)\. Replicate\-level contrasts are in Table[19](https://arxiv.org/html/2608.20442#A16.T19); the sign\-test design is discussed in Appendix[I](https://arxiv.org/html/2608.20442#A9)\.

Cross\-architecture replication\.Repeating the matched factorial on Llama\-3\.2\-1B with eight source\-free updates givesΔS\+=−0\.131±0\.050\\Delta\_\{S^\{\+\}\}=\-0\.131\\pm 0\.050,Δznull=−0\.017±0\.035\\Delta\_\{z\_\{\\mathrm\{null\}\}\}=\-0\.017\\pm 0\.035, andΔS−=\+0\.090±0\.031\\Delta\_\{S^\{\-\}\}=\+0\.090\\pm 0\.031\(mean±\\pmsample SD\)\. Both adjacent contrasts are positive in 12/12 independent seeds \(one\-sided sign test, Holm\-adjustedp=0\.000488p=0\.000488each\), and a held\-out likelihood observer preserves this consistent ordering\. Route\-wise physical descendant norms remain nearly equal and source\-cut behavior remains near zero, separating route\-conditioned value from source magnitude \(Appendix[M\.6](https://arxiv.org/html/2608.20442#A13.SS6)\)\.

Longer\-horizon extension\.In a separate paired extension, doubling the source\-free suffix from eight to sixteen updates increases both contrasts in 9/9 tested seeds \(two\-sided sign test, Holm\-adjustedp=0\.0078p=0\.0078each\)\. AtH=16H=16,d1=0\.282±0\.068d\_\{1\}=0\.282\\pm 0\.068andd2=0\.258±0\.047d\_\{2\}=0\.258\\pm 0\.047; both are positive and resolved in 9/9 tested seeds \(one\-sided sign test, Holm\-adjustedp=0\.0039p=0\.0039each\), and the held\-out likelihood observer gives the same ordering\. Physical descendant norms remain closely matched while their behavioral values separate \(Appendix[M\.6](https://arxiv.org/html/2608.20442#A13.SS6)\)\.

Source\-window ablation\(Figure[4](https://arxiv.org/html/2608.20442#S4.F4)c\)\. Three additional source blocks reproduce the signed route topology at 77–89% of the reference amplitude\. All 64 prompt\-level contrasts are positive; these repeated probes are reported as descriptive aggregates \(Appendix[M\.3](https://arxiv.org/html/2608.20442#A13.SS3)\)\.

Random\-plane control\.The diagnostic routes above use the frozen readout directionϕ^\\hat\{\\phi\}\. In the control, it is replaced by one fixed norm\-matched random directionr^\\hat\{r\}projected orthogonally to bothϕ^\\hat\{\\phi\}and the gradient contrastu^\\hat\{u\}\. Routes constructed asr^±u^\\hat\{r\}\\pm\\hat\{u\}retain sign\-coherent adjacent contrasts consistently across independent seeds \(two\-sided exact sign test, Holm\-adjustedp=0\.031p=0\.031for the two contrasts; Appendix[J](https://arxiv.org/html/2608.20442#A10)\)\. The random plane has no preassigned orientation, so the relevant phenomenon is two\-sided sign coherence, irrespective of the arbitraryS±S^\{\\pm\}label assignment\.

### 4\.2Prediction on Ordinary Routes

The factorial shows that engineered futures can assign opposite value to the same ancestry\. We next ask whether the costate predicts responses under*ordinary*future training\. The compact scalar predictor isΔ^R=pER​\(0\)⊤​AE\\widehat\{\\Delta\}\_\{R\}=p\_\{E\}^\{R\}\(0\)^\{\\top\}A\_\{E\}, whereAE=SE​\(\+1\)−SE​\(−1\)A\_\{E\}=S\_\{E\}\(\+1\)\-S\_\{E\}\(\-1\)is the finite state difference between the two source arms at the cut andpER​\(0\)p\_\{E\}^\{R\}\(0\)is the midpoint costate for routeRR\. Its forward\-dual implementation propagatesAEA\_\{E\}once through each route and reads the prompt vector: route means determine the 42 sign results, while concatenated prompt\-level vectors determine the per\-seed MSE comparisons\. The predictor has no fitted parameters\.

Figure 5:Prediction of transport and route\-conditioned value\.\(a\)Observer\-free transport prediction on Qwen, SmolLM2, and Llama\-3\.2 \(log scale\): the full tangent predictor sits22–66orders of magnitude below a source\-only comparator; route\-centered errors are shown separately \(§[5](https://arxiv.org/html/2608.20442#S5)\)\.\(b\)All 42 Qwen costate predictions versus observed route means; the ringed point is a naturally negative route\.For each of seven independent seeds, we generate six ordinary future routes from random data schedules and retain all 42\. We compare the costate predictor with ZERO, a linear predictor from the source cut, and a route\-constant ablation that assigns each prompt’s across\-route mean prediction to every route \(Appendix[M\.3](https://arxiv.org/html/2608.20442#A13.SS3)\)\.

The full\-horizon predictor beats all three baselines in 7/7 independent seeds \(two\-sided exact sign testp=0\.016p=0\.016\), and its 42 route\-mean signs all match the observed endpoints \(Figure[5](https://arxiv.org/html/2608.20442#S4.F5)b\), including a negative route within one seed\. The panel contains 29 positive and 13 negative raw route means, so a constant always\-positive rule scores only 29/42\.

Cross\-architecture prediction\.On Llama\-3\.2\-1B, a separate panel applies the same full\-horizon predictor to six ordinary, observer\-independent continuations per seed\. It has lower route\-panel SSE than ZERO and source\-cut\-only in 9/9 independent seeds \(one\-sided sign tests, Holm\-adjustedp=0\.0039p=0\.0039each\), with mean within\-seed Spearman correlation 0\.892\. The predicted sign matches 51/54 raw route means and all 21 whose behavioral response clears the resolution threshold; the three mismatches are unresolved near\-zero responses \(Appendix[M\.6](https://arxiv.org/html/2608.20442#A13.SS6)\)\.

Finite\-amplitude remainder\.Across the 42 Qwen ordinary routes, the first\-order approximation attains aggregateSSE/ZERO=0\.0103\\mathrm\{SSE/ZERO\}=0\.0103and recovers every route\-mean sign; individual magnitude errors reflect neglected higher\-order terms\. The engineeredS±S^\{\\pm\}factorial instead uses matched finite interventions to demonstrate continuation\-dependent sign assignment, as interpreted by the exact identity \(Eq\.[6](https://arxiv.org/html/2608.20442#S2.E6)\)\.

Source\-scale validity boundary\.As source amplitude decreases, the observable route response similarly diminishes, indicating a low\-signal regime\. A Qwen titration holds all other protocol elements fixed while subsampling the source corpus\. The predictor’s resolution degrades at low source signals, beating baselines consistently only at and above 1024 rows; intermediate thresholds yield seed\-unstable results \(Appendix Figure[7](https://arxiv.org/html/2608.20442#A12.F7)a\)\. At low source signal, it no longer reliably improves on the baselines\.

Route\-aware non\-adjoint baselines\.Three cheaper predictor families test whether the full\-horizon costate is necessary\. Truncating the propagated future toH∈\{1,2,4\}H\\in\\\{1,2,4\\\}of the eight suffix updates remains below all three baselines in every seed, improving monotonically inHH, while the full horizon wins consistently across all seeds \(Appendix Figure[7](https://arxiv.org/html/2608.20442#A12.F7)b\)\. Forward\-only and schedule\-statistics predictors are also less accurate; their constructions and aggregate results are in Appendix[J](https://arxiv.org/html/2608.20442#A10)\.

Dominant optimizer\-slot pairs can differ across systems\.In a one\-seed diagnostic for each model, propagating each state\-block component through the same future identifies different dominant pairs: parameters plus first moment in Qwen, the two moments in SmolLM2\[[4](https://arxiv.org/html/2608.20442#bib.bib48)\]\. In each tested cell the best pair accounts for nearly all of the full\-state prediction while the next\-best pair captures less than half \(Appendix[I](https://arxiv.org/html/2608.20442#A9)\)\.

## 5Generalization and Variable Behavioral Outcomes

The experiments reveal two forms of generalization\. Transport recurs across architectures, optimizers, traits, and non\-LoRA vision models\. Behavioral value varies with the continuation, observer, and system\. Table[1](https://arxiv.org/html/2608.20442#S5.T1)summarizes the observer\-free transport panel across architectures; transplant sufficiency, block\-lineage dominance, and behavioral amplitude are reported separately below and in Appendix[I](https://arxiv.org/html/2608.20442#A9)–[P](https://arxiv.org/html/2608.20442#A16)\.

Table 1:Observer\-free transport across architectures\. Entries are seed means from the same four\-route panel; lowerSSE/ZERO\\mathrm\{SSE/ZERO\}is better\.The\(w,m,v\)\(w,m,v\)block response uses the implemented parameter and moment tensors; merged weights uses theδ⁡\(B​A\)\\delta\(BA\)tangent; hidden stacks all\-layer responses on neutral probes\. Per\-seed values and normalization are in Table[10](https://arxiv.org/html/2608.20442#A13.T10)\.

Across Qwen2\.5\-0\.5B, SmolLM2\-135M, and Llama\-3\.2\-1B, the observer\-free tangent recurrence attains seed\-meanSSE/ZERO≤3×10−3\\mathrm\{SSE/ZERO\}\\leq 3\\times 10^\{\-3\}\. The same bound holds across all five Qwen trait families \(Table[20](https://arxiv.org/html/2608.20442#A16.T20)\), two to six orders of magnitude below a source\-only comparator \(Figure[5](https://arxiv.org/html/2608.20442#S4.F5)a\)\. Dropping optimizer components degrades prediction by three to five orders of magnitude for every trait \(Table[21](https://arxiv.org/html/2608.20442#A16.T21)\)\.

Beyond language models, the full\-state recurrence predicts responses in non\-LoRA MNIST MLP/CNN systems under AdamW and momentum SGD in all three seeds, while plain SGD supplies the direct\-write limit\. The backward\-rotation experiment separately changes where information is written while preserving it\. Together, these panels show that transport and optimizer relay are not specific to transformers or LoRA \(Appendix Figures[6](https://arxiv.org/html/2608.20442#A2.F6)a and[10](https://arxiv.org/html/2608.20442#A12.F10)a\)\.

The likelihood ordering also survives a change of behavioral readout\. On Qwen, an independent panel reproduces both adjacent route contrasts consistently in generated pairwise choices \(Holmp=0\.016p=0\.016each\)\. Replacing AdamW with momentum SGD reproduces the ordered interaction in 3/3 tested seeds \(Appendix[I](https://arxiv.org/html/2608.20442#A9)\)\.

The transplant and factorial results extend across five trait families\. The 15 primary transplant cells selectw\+mw\{\+\}mspanning Qwen \(owl, cat, red, oak\) and Llama \(blue\) \(Appendix Figure[9](https://arxiv.org/html/2608.20442#A12.F9)and Table[15](https://arxiv.org/html/2608.20442#A16.T15)\)\. Qwen blue appears in the observer\-free and matched\-factorial panels\. Direct Qwen transplant responses range from red’s negative seed mean through cat’s near\-zero mean to oak’s positive mean \(Table[16](https://arxiv.org/html/2608.20442#A16.T16)\), while the five\-trait factorial measures route ordering and statistical resolution \(Table[18](https://arxiv.org/html/2608.20442#A16.T18)\)\.

Transport does not fix the behavioral outcome\.CNN block allocations vary tenfold across seeds\. In a representative SmolLM2 cell, finite Shapley terms cancel by 90\.9% between source updates; its separate block\-lineage panel identifies the two moments as the dominant pair\. TinyLlama\[[51](https://arxiv.org/html/2608.20442#bib.bib49)\]shows opposing finite contributions without a stable behavioral sign and does not rank transplant subsets\. On Llama\-3\.2, direct source–control responses in the transplant panels remain weak and seed\-heterogeneous, while the matched factorial resolves route\-conditioned valuation at both eight and sixteen source\-free updates\. These cases show that ancestry can be transported even when behavioral value cancels, varies across seeds, or remains weak \(Appendix Figure[10](https://arxiv.org/html/2608.20442#A12.F10); per\-system details in Appendix[I](https://arxiv.org/html/2608.20442#A9)\)\.

## 6Related Work

Subliminal learning mechanisms\.Knowledge distillation\[[25](https://arxiv.org/html/2608.20442#bib.bib42)\]showed that a teacher’s output distribution carries information beyond labels; subliminal learning\[[12](https://arxiv.org/html/2608.20442#bib.bib1)\]extends this to behavioral dispositions invisible in the data\. Subsequent work investigates how these subliminal signals enter the student’s gradient: divergence tokens\[[44](https://arxiv.org/html/2608.20442#bib.bib2)\], steering\-vector distillation\[[6](https://arxiv.org/html/2608.20442#bib.bib3)\], LoRA amplification\[[37](https://arxiv.org/html/2608.20442#bib.bib6)\], and stronger student encoding\[[36](https://arxiv.org/html/2608.20442#bib.bib4)\]\. We study what follows once the signal reaches the gradient port: its survival and expression within the trainer state\.

Adjoint methods and trajectory attribution\.Discrete\-time adjoint analysis\[[40](https://arxiv.org/html/2608.20442#bib.bib18)\]allows differentiating through momentum and optimizer buffers, as established in hyperparameter optimization\[[33](https://arxiv.org/html/2608.20442#bib.bib20)\]and trajectory attribution\. Traditional attribution methods trace sample influence via gradients or checkpoints\[[30](https://arxiv.org/html/2608.20442#bib.bib14),[41](https://arxiv.org/html/2608.20442#bib.bib23)\]; recent methods use approximate unrolling\[[3](https://arxiv.org/html/2608.20442#bib.bib24)\], fixed\-state Adam\-aware valuation\[[16](https://arxiv.org/html/2608.20442#bib.bib27)\], or explicit reverse\-mode tracing through Adam/AdamW state\[[15](https://arxiv.org/html/2608.20442#bib.bib26)\]\. Building on this capability to differentiate through optimizer state, we isolate how future training continuation assigns behavioral value to a finite source\-content contrast\.

Optimizer memory and order dependence\.Optimization literature recognizes that optimizer state acts as an implicit modification of the loss\[[9](https://arxiv.org/html/2608.20442#bib.bib38)\]and steers later updates\[[47](https://arxiv.org/html/2608.20442#bib.bib40)\]\.[45](https://arxiv.org/html/2608.20442#bib.bib37)demonstrated this memory effect directly by showing it collapses when optimizer buffers are reset\. We use these memory dynamics to identify optimizer state as the causal carrier of source\-specific ancestry and to separate physical transport from future behavioral valuation\.

## 7Limitations and Conclusion

Limitations\.The LLM panels span 135M–1\.1B parameters and five Qwen and two Llama trait families\. The transplant cells consistently selectw\+mw\{\+\}m, while behavioral magnitude and block\-lineage allocation vary by system\. Exact attribution requires the full trajectory and one adjoint solve per route; the compact midpoint predictor loses resolution when the source signal is small\. The identity applies to gradient\-port perturbations generally, with subliminal transfer distinguished by how the perturbation enters training\.

Conclusion\.A subliminal source can remain causally active after it leaves the training stream\. Optimizer state carries that influence forward, while subsequent training determines whether it is expressed, cancelled, or reversed\. The response identity, state surgery, and route predictions make transport and valuation separately measurable across language models, optimizers, traits, and non\-LoRA vision systems\. This separates what the trainer remembers from what the model eventually does: the same physical ancestry can persist while its behavioral value changes with the future route\.

## References

- \[1\]A\. Achille, M\. Rovere, and S\. Soatto\(2019\)Critical learning periods in deep networks\.InInternational Conference on Learning Representations \(ICLR\),Note:arXiv:1711\.08856External Links:[Link](https://openreview.net/forum?id=BkeStsCcKQ)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p16.1)\.
- \[2\]B\. Askin, M\. Ustaomeroglu, A\. Nayak, G\. Joshi, G\. Qu, and C\. Joe\-Wong\(2026\)Emergent and subliminal misalignment through the lens of data\-mediated transfer\.arXiv preprint arXiv:2605\.12798\.External Links:[Link](https://arxiv.org/abs/2605.12798)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p4.1)\.
- \[3\]J\. Bae, W\. Lin, J\. Lorraine, and R\. Grosse\(2024\)Training data attribution via approximate unrolling\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Vol\.37,pp\. 66647–66686\.External Links:[Document](https://dx.doi.org/10.52202/079017-2129),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/hash/7af60ccb99c7a434a0d9d9c1fb00ca94-Abstract-Conference.html)Cited by:[Appendix E](https://arxiv.org/html/2608.20442#A5.p2.1),[§6](https://arxiv.org/html/2608.20442#S6.p2.1)\.
- \[4\]L\. Ben Allal, A\. Lozhkov, E\. Bakouch, G\. Martín Blázquez, G\. Penedo, L\. Tunstall, A\. Marafioti, A\. Piqueres Lajarín, H\. Kydlíček, V\. Srivastav, J\. Lochner, C\. Fahlgren, X\. Nguyen, B\. Burtenshaw, C\. Fourrier, H\. Zhao, H\. Larcher, M\. Morlon, C\. Zakka, C\. Raffel, L\. von Werra, and T\. Wolf\(2025\)SmolLM2: when smol goes big—data\-centric training of a fully open small language model\.InSecond Conference on Language Modeling \(COLM\),Note:arXiv:2502\.02737External Links:[Link](https://openreview.net/forum?id=3JiCl2A14H)Cited by:[§4\.2](https://arxiv.org/html/2608.20442#S4.SS2.p8.1)\.
- \[5\]J\. Betley, N\. Warncke, A\. Sztyber\-Betley, D\. Tan, X\. Bao, M\. Soto, M\. Srivastava, N\. Labenz, and O\. Evans\(2026\)Training large language models on narrow tasks can lead to broad misalignment\.Nature649,pp\. 584–589\.Note:arXiv:2502\.17424External Links:[Document](https://dx.doi.org/10.1038/s41586-025-09937-5),[Link](https://www.nature.com/articles/s41586-025-09937-5)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p10.1)\.
- \[6\]C\. Blank, A\. Bhatia, S\. Rajamanoharan, A\. Conmy, and N\. Nanda\(2026\)Subliminal learning is steering vector distillation\.arXiv preprint arXiv:2606\.00995\.External Links:[Link](https://arxiv.org/abs/2606.00995)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p2.1),[§1](https://arxiv.org/html/2608.20442#S1.p1.1),[§6](https://arxiv.org/html/2608.20442#S6.p1.1)\.
- \[7\]J\. Bolte and E\. Pauwels\(2021\)Conservative set valued fields, automatic differentiation, stochastic gradient methods and deep learning\.Mathematical Programming188\(1\),pp\. 19–51\.Note:arXiv:1909\.10300External Links:[Document](https://dx.doi.org/10.1007/s10107-020-01501-5),[Link](https://doi.org/10.1007/s10107-020-01501-5)Cited by:[Appendix F](https://arxiv.org/html/2608.20442#A6.p1.1)\.
- \[8\]V\. C\. Brockers, R\. D\. Ventzke, V\. Neuhaus, B\. Hidalgo\-Ogalde, and V\. Priesemann\(2026\)Learning through noise: why subliminal learning works and when it fails\.arXiv preprint arXiv:2605\.23645\.External Links:[Link](https://arxiv.org/abs/2605.23645)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p8.1)\.
- \[9\]M\. Cattaneo and B\. Shigida\(2025\)How memory in optimization algorithms implicitly modifies the loss\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Vol\.38,pp\. 173262–173299\.Note:arXiv:2502\.02132External Links:[Document](https://dx.doi.org/10.52202/085713-5215),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/hash/e4cc8ab4a64e99f962f36d07a7723d94-Abstract-Conference.html)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p12.1),[§1](https://arxiv.org/html/2608.20442#S1.p2.1),[§6](https://arxiv.org/html/2608.20442#S6.p3.1)\.
- \[10\]K\. Chauhan and A\. Shah\(2026\)Covert trait propagation is representation alignment: mechanistic evidence from hidden\-channel distillation\.arXiv preprint arXiv:2607\.04432\.External Links:[Link](https://arxiv.org/abs/2607.04432)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p7.1)\.
- \[11\]C\. Chiang and H\. Lee\(2022\)On the transferability of pre\-trained language models: a study from artificial datasets\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.36,pp\. 10518–10525\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v36i10.21295),[Link](https://ojs.aaai.org/index.php/AAAI/article/view/21295)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p11.1)\.
- \[12\]A\. Cloud, M\. Le, J\. Chua, J\. Betley, A\. Sztyber\-Betley, S\. Mindermann, J\. Hilton, S\. Marks, and O\. Evans\(2026\)Language models transmit behavioural traits through hidden signals in data\.Nature652\(8110\),pp\. 615–621\.Note:arXiv:2507\.14805External Links:[Document](https://dx.doi.org/10.1038/s41586-026-10319-8),[Link](https://www.nature.com/articles/s41586-026-10319-8)Cited by:[§1](https://arxiv.org/html/2608.20442#S1.p1.1),[§6](https://arxiv.org/html/2608.20442#S6.p1.1)\.
- \[13\]J\. Cui, Y\. Han, J\. Jiao, and J\. Zhang\(2026\)Persistent backdoor attacks under continual fine\-tuning of LLMs\.Proceedings of the AAAI Conference on Artificial Intelligence40\(36\),pp\. 30422–30430\.Note:arXiv:2512\.14741External Links:[Document](https://dx.doi.org/10.1609/aaai.v40i36.40295),[Link](https://ojs.aaai.org/index.php/AAAI/article/view/40295)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p15.1)\.
- \[14\]J\. Dang, B\. Y\. Xie, and O\. G\. Younis\(2026\)Subliminal transfer of unsafe behaviors in AI agent distillation\.arXiv preprint arXiv:2604\.15559\.External Links:[Link](https://arxiv.org/abs/2604.15559)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p5.1)\.
- \[15\]J\. Deng, P\. Hu, S\. Jin, H\. Lu, J\. T\. Wang, S\. Zhang, and J\. W\. Ma\(2026\)How faithful is trajectory\-based data attribution? error sources, remedies, and practical guidelines\.arXiv preprint arXiv:2605\.18814\.External Links:[Link](https://arxiv.org/abs/2605.18814)Cited by:[§6](https://arxiv.org/html/2608.20442#S6.p2.1)\.
- \[16\]M\. Ding, Z\. Zhang, D\. Wang, and L\. Hu\(2026\)In\-run data shapley for Adam optimizer\.arXiv preprint arXiv:2602\.00329\.External Links:[Link](https://arxiv.org/abs/2602.00329)Cited by:[§6](https://arxiv.org/html/2608.20442#S6.p2.1)\.
- \[17\]C\. Finn, P\. Abbeel, and S\. Levine\(2017\)Model\-agnostic meta\-learning for fast adaptation of deep networks\.InProceedings of the 34th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.70,pp\. 1126–1135\.External Links:[Link](https://proceedings.mlr.press/v70/finn17a.html)Cited by:[§2\.1](https://arxiv.org/html/2608.20442#S2.SS1.p3.1)\.
- \[18\]L\. Franceschi, M\. Donini, P\. Frasconi, and M\. Pontil\(2017\)Forward and reverse gradient\-based hyperparameter optimization\.InProceedings of the 34th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.70,pp\. 1165–1173\.External Links:[Link](https://proceedings.mlr.press/v70/franceschi17a.html)Cited by:[§2\.1](https://arxiv.org/html/2608.20442#S2.SS1.p3.1)\.
- \[19\]S\. Garg and S\. S\. Vempala\(2022\)How and when random feedback works: a case study of low\-rank matrix factorization\.InProceedings of the 25th International Conference on Artificial Intelligence and Statistics,Proceedings of Machine Learning Research, Vol\.151,pp\. 4070–4108\.External Links:[Link](https://proceedings.mlr.press/v151/garg22a.html)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p16.1)\.
- \[20\]K\. G\. Georgiev, R\. Rinberg, S\. Park, S\. Garg, A\. Ilyas, A\. Madry, and S\. Neel\(2025\)Machine unlearning via simulated oracle matching\.InInternational Conference on Learning Representations \(ICLR\),pp\. 49693–49731\.Note:arXiv:2410\.23232External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/hash/7c799b09cc40973ceaa47da50131dc63-Abstract-Conference.html)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p14.1)\.
- \[21\]A\. Griewank and A\. Walther\(2008\)Evaluating derivatives: principles and techniques of algorithmic differentiation\.2nd edition,SIAM\.External Links:[Document](https://dx.doi.org/10.1137/1.9780898717761),ISBN 9780898716597,[Link](https://doi.org/10.1137/1.9780898717761)Cited by:[Appendix F](https://arxiv.org/html/2608.20442#A6.p1.1),[§1](https://arxiv.org/html/2608.20442#S1.p4.1),[§2\.1](https://arxiv.org/html/2608.20442#S2.SS1.p3.1)\.
- \[22\]D\. Han, M\. Han, and Unsloth Team\(2023\)Unsloth\.Note:GitHub repositoryExternal Links:[Link](https://github.com/unslothai/unsloth)Cited by:[§M\.6](https://arxiv.org/html/2608.20442#A13.SS6.p1.1)\.
- \[23\]Y\. Hao, Y\. Cao, and L\. Mou\(2024\)Flora: low\-rank adapters are secretly gradient compressors\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 17554–17571\.External Links:[Link](https://proceedings.mlr.press/v235/hao24a.html)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p16.1)\.
- \[24\]S\. Hayou, N\. Ghosh, and B\. Yu\(2024\)LoRA\+: efficient low rank adaptation of large models\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 17783–17806\.External Links:[Link](https://proceedings.mlr.press/v235/hayou24a.html)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p16.1)\.
- \[25\]G\. Hinton, O\. Vinyals, and J\. Dean\(2015\)Distilling the knowledge in a neural network\.arXiv preprint arXiv:1503\.02531\.Note:Presented at the NIPS 2014 Deep Learning and Representation Learning WorkshopExternal Links:[Link](https://arxiv.org/abs/1503.02531)Cited by:[§6](https://arxiv.org/html/2608.20442#S6.p1.1)\.
- \[26\]E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. Chen\(2022\)LoRA: low\-rank adaptation of large language models\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://openreview.net/forum?id=nZeVKeeFYf9)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p16.1)\.
- \[27\]S\. Hu, Y\. Fu, Z\. S\. Wu, and V\. Smith\(2025\)Unlearning or obfuscating? jogging the memory of unlearned LLMs via benign relearning\.InInternational Conference on Learning Representations \(ICLR\),pp\. 8857–8888\.Note:arXiv:2406\.13356External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/hash/18fd48d9cbbf9a20e434c9d3db6973c5-Abstract-Conference.html)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p14.1)\.
- \[28\]E\. Hubinger, C\. Denison, J\. Mu, M\. Lambert, M\. Tong, M\. MacDiarmid, T\. Lanham, D\. M\. Ziegler, T\. Maxwell, N\. Cheng, A\. Jermyn, A\. Askell, A\. Radhakrishnan, C\. Anil, D\. Duvenaud, D\. Ganguli, F\. Barez, J\. Clark, K\. Ndousse, K\. Sachan, M\. Sellitto, M\. Sharma, N\. DasSarma, R\. Grosse, S\. Kravec, Y\. Bai, Z\. Witten, M\. Favaro, J\. Brauner, H\. Karnofsky, P\. Christiano, S\. R\. Bowman, L\. Graham, J\. Kaplan, S\. Mindermann, R\. Greenblatt, B\. Shlegeris, N\. Schiefer, and E\. Perez\(2024\)Sleeper agents: training deceptive LLMs that persist through safety training\.arXiv preprint arXiv:2401\.05566\.External Links:[Link](https://arxiv.org/abs/2401.05566)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p15.1)\.
- \[29\]G\. Ilharco, M\. T\. Ribeiro, M\. Wortsman, S\. Gururangan, L\. Schmidt, H\. Hajishirzi, and A\. Farhadi\(2023\)Editing models with task arithmetic\.InInternational Conference on Learning Representations \(ICLR\),Note:arXiv:2212\.04089External Links:[Link](https://openreview.net/forum?id=6t0Kwf8-jrj)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p11.1)\.
- \[30\]P\. W\. Koh and P\. Liang\(2017\)Understanding black\-box predictions via influence functions\.InProceedings of the 34th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.70,pp\. 1885–1894\.External Links:[Link](https://proceedings.mlr.press/v70/koh17a.html)Cited by:[§6](https://arxiv.org/html/2608.20442#S6.p2.1)\.
- \[31\]U\. König, H\. Kazmi, R\. Li, and M\. Chaudhary\(2026\)Quantifying subliminal behavioral transfer ratios in language model distillation\.arXiv preprint arXiv:2606\.11270\.External Links:[Link](https://arxiv.org/abs/2606.11270)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p5.1)\.
- \[32\]I\. Loshchilov and F\. Hutter\(2019\)Decoupled weight decay regularization\.InInternational Conference on Learning Representations \(ICLR\),Note:arXiv:1711\.05101External Links:[Link](https://openreview.net/forum?id=Bkg6RiCqY7)Cited by:[§2\.4](https://arxiv.org/html/2608.20442#S2.SS4.p2.2)\.
- \[33\]D\. Maclaurin, D\. Duvenaud, and R\. P\. Adams\(2015\)Gradient\-based hyperparameter optimization through reversible learning\.InProceedings of the 32nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.37,pp\. 2113–2122\.External Links:[Link](https://proceedings.mlr.press/v37/maclaurin15.html)Cited by:[§D\.3](https://arxiv.org/html/2608.20442#A4.SS3.p2.3),[§1](https://arxiv.org/html/2608.20442#S1.p4.1),[§2\.1](https://arxiv.org/html/2608.20442#S2.SS1.p3.1),[§6](https://arxiv.org/html/2608.20442#S6.p2.1)\.
- \[34\]T\. Madl\(2026\)Channel location constrains the auditability of subliminal learning\.arXiv preprint arXiv:2606\.22019\.External Links:[Link](https://arxiv.org/abs/2606.22019)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p4.1)\.
- \[35\]Meta AI\(2024\)Llama 3\.2 model card\.Note:GitHub repositoryExternal Links:[Link](https://github.com/meta-llama/llama-models/blob/main/models/llama3_2/MODEL_CARD.md)Cited by:[§M\.6](https://arxiv.org/html/2608.20442#A13.SS6.p1.1),[§3](https://arxiv.org/html/2608.20442#S3.p5.1)\.
- \[36\]G\. Morgulis and J\. Hewitt\(2026\)Subliminal steering: stronger encoding of hidden signals\.arXiv preprint arXiv:2604\.25783\.External Links:[Link](https://arxiv.org/abs/2604.25783)Cited by:[§6](https://arxiv.org/html/2608.20442#S6.p1.1)\.
- \[37\]T\. Nief, H\. Y\. Fu, M\. Muchane, and A\. Holtzman\(2026\)Subliminal learning is a LoRA artifact\.arXiv preprint arXiv:2606\.00831\.External Links:[Link](https://arxiv.org/abs/2606.00831)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p3.1),[§6](https://arxiv.org/html/2608.20442#S6.p1.1)\.
- \[38\]A\. Nøkland\(2016\)Direct feedback alignment provides learning in deep neural networks\.InAdvances in Neural Information Processing Systems,Vol\.29,pp\. 1037–1045\.External Links:[Link](https://proceedings.neurips.cc/paper/2016/hash/d490d7b4576290fa60eb31b5fc917ad1-Abstract.html)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p16.1)\.
- \[39\]A\. S\. Okatan, M\. İ\. Akbaş, L\. N\. Kandel, and B\. Peköz\(2025\)Seed\-induced uniqueness in transformer models: subspace alignment governs subliminal transfer\.InIEEE Cyber Awareness and Research Symposium \(CARS\),pp\. 1–6\.Note:arXiv:2511\.01023External Links:[Document](https://dx.doi.org/10.1109/CARS67163.2025.11337559),[Link](https://doi.org/10.1109/CARS67163.2025.11337559)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p6.1)\.
- \[40\]L\. S\. Pontryagin, V\. G\. Boltyanskii, R\. V\. Gamkrelidze, and E\. F\. Mishchenko\(1962\)The mathematical theory of optimal processes\.Interscience Publishers\.Cited by:[§1](https://arxiv.org/html/2608.20442#S1.p4.1),[§2\.1](https://arxiv.org/html/2608.20442#S2.SS1.p3.1),[§6](https://arxiv.org/html/2608.20442#S6.p2.1)\.
- \[41\]G\. Pruthi, F\. Liu, S\. Kale, and M\. Sundararajan\(2020\)Estimating training data influence by tracing gradient descent\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Vol\.33,pp\. 19920–19930\.External Links:[Link](https://proceedings.neurips.cc/paper/2020/hash/e6385d39ec9394f2f3a354d9d2b88eec-Abstract.html)Cited by:[§6](https://arxiv.org/html/2608.20442#S6.p2.1)\.
- \[42\]Qwen Team\(2024\)Qwen2\.5 technical report\.arXiv preprint arXiv:2412\.15115\.External Links:[Link](https://arxiv.org/abs/2412.15115)Cited by:[§2\.4](https://arxiv.org/html/2608.20442#S2.SS4.p3.1)\.
- \[43\]M\. Refinetti, S\. d’Ascoli, R\. Ohana, and S\. Goldt\(2021\)Align, then memorise: the dynamics of learning with feedback alignment\.InProceedings of the 38th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.139,pp\. 8925–8935\.External Links:[Link](https://proceedings.mlr.press/v139/refinetti21a.html)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p16.1)\.
- \[44\]S\. Schrodi, E\. Kempf, F\. Barez, and T\. Brox\(2026\)Towards understanding subliminal learning: when and how hidden biases transfer\.InInternational Conference on Learning Representations \(ICLR\),Note:arXiv:2509\.23886External Links:[Link](https://openreview.net/forum?id=IelhmYSjPt)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p1.1),[§1](https://arxiv.org/html/2608.20442#S1.p1.1),[§6](https://arxiv.org/html/2608.20442#S6.p1.1)\.
- \[45\]V\. Sevetlidis and G\. Pavlidis\(2026\)Process\-tensor tomography of SGD: measuring non\-markovian memory via back\-flow of distinguishability\.InProceedings of the 29th International Conference on Artificial Intelligence and Statistics \(AISTATS\),Proceedings of Machine Learning Research, Vol\.300\.Note:arXiv:2601\.16563External Links:[Link](https://openreview.net/forum?id=TonOzlbE3k)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p12.1),[§1](https://arxiv.org/html/2608.20442#S1.p2.1),[§6](https://arxiv.org/html/2608.20442#S6.p3.1)\.
- \[46\]S\. Sun, S\. Y\. Baek, and J\. H\. Kim\(2025\)Personality vector: modulating personality of large language models by model merging\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 24656–24677\.Note:arXiv:2509\.19727External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.1253),[Link](https://aclanthology.org/2025.emnlp-main.1253/)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p11.1)\.
- \[47\]J\. Sweeney\(2026\)Optimizer memory makes shuffle order a first\-order source of fine\-tuning noise\.arXiv preprint arXiv:2606\.29554\.External Links:[Link](https://arxiv.org/abs/2606.29554)Cited by:[§6](https://arxiv.org/html/2608.20442#S6.p3.1)\.
- \[48\]J\. Sweeney\(2026\)The geometry of sequential learning: Lie\-bracket prediction of transfer order\.InProceedings of the 43rd International Conference on Machine Learning \(ICML\),Proceedings of Machine Learning Research, Vol\.306\.Note:arXiv:2606\.24993External Links:[Link](https://openreview.net/forum?id=KL0eu92H3K)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p13.1)\.
- \[49\]X\. Xu, X\. Yue, Y\. Liu, Q\. Ye, H\. Zheng, P\. Hu, M\. Du, and H\. Hu\(2026\)Unlearning isn’t deletion: investigating reversibility of machine unlearning in LLMs\.InProceedings of the 43rd International Conference on Machine Learning \(ICML\),Proceedings of Machine Learning Research, Vol\.306\.Note:arXiv:2505\.16831External Links:[Link](https://openreview.net/forum?id=E5SVowO13b)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p14.1)\.
- \[50\]A\. Yanagisawa, A\. Khan, T\. K\. Balraj Singh, Y\. Na, K\. Zhu, and A\. Mari\(2025\)Liminal training: characterizing and mitigating subliminal learning in large language models\.InNeurIPS 2025 Workshop on Socially Responsible and Trustworthy Foundation Models \(ResponsibleFM\),Note:Also AAAI 2026 xAI4Science WorkshopExternal Links:[Link](https://openreview.net/forum?id=aslS4eRygE)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p9.1)\.
- \[51\]P\. Zhang, G\. Zeng, T\. Wang, and W\. Lu\(2024\)TinyLlama: an open\-source small language model\.arXiv preprint arXiv:2401\.02385\.External Links:[Link](https://arxiv.org/abs/2401.02385)Cited by:[§M\.6](https://arxiv.org/html/2608.20442#A13.SS6.p6.1),[§5](https://arxiv.org/html/2608.20442#S5.p6.1)\.
- \[52\]J\. Zhu, K\. Greenewald, K\. Nadjahi, H\. Sáez de Ocáriz Borde, R\. B\. Gabrielsson, L\. Choshen, M\. Ghassemi, M\. Yurochkin, and J\. Solomon\(2024\)Asymmetry in low\-rank adapters of foundation models\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 62369–62385\.External Links:[Link](https://proceedings.mlr.press/v235/zhu24c.html)Cited by:[Appendix H](https://arxiv.org/html/2608.20442#A8.p16.1)\.

## Appendix ANotation

Table 2:Symbols used throughout the paper\.Table 3:Narrative terms used throughout the paper, in plain language\.
## Appendix BSource Entry and Coordinate Access

Full\-state prediction across 360 cells\.We apply a source pulse \(a brief block of source\-active updates\) to the Qwen2\.5\-0\.5B trainer \(LoRA\-r8, AdamW\) and predict the complete state response—parameters, first moment, and second moment jointly—using the tangent recurrence \(Eq\.[3](https://arxiv.org/html/2608.20442#S2.E3)\) with system\-specific Jacobians, not a fitted model\. Across 360 cells \(step×\\timesstate\-block combinations\) the joint prediction reachesSSE/ZERO=1\.65×10−4\\mathrm\{SSE/ZERO\}=1\.65\\times 10^\{\-4\}with correlation 0\.9999\. At lag 16 the full\-state predictor maintains this accuracy, while a parameter\-only predictor degrades toSSE/ZERO=0\.786\\mathrm\{SSE/ZERO\}=0\.786and a moments\-only predictor reaches 0\.019 \(Figure[6](https://arxiv.org/html/2608.20442#A2.F6)b\): near\-exact tracking requires the full optimizer state\.

Backward\-rotation dose response\.In a frozen\-head MNIST experiment \(classifying from the fixed output layer without retraining it\), we rotate the hidden backward map by angleθ\\thetawhile preserving forward outputs, loss, gradient norm, and singular spectrum—changing*where*the source writes without changing*what*the network computes\. Across0∘0^\{\\circ\}–90∘90^\{\\circ\}, frozen\-head accuracy declines overall while a linear probe stays nearly flat, and unlabeled Procrustes alignment restores frozen\-head accuracy—the information is present but written to different coordinates \(Figure[6](https://arxiv.org/html/2608.20442#A2.F6)a\)\. A held\-out−60∘\-60^\{\\circ\}ancestry is predicted from three basis rotations by linear combination in the angular coordinate atSSE/ZERO<10−6\\mathrm\{SSE/ZERO\}<10^\{\-6\}\(AdamW\) and<10−12<10^\{\-12\}\(SGD\), while wrong\-source predictions score≥1\.0\\geq 1\.0\.

Figure 6:Source entry and coordinate access\.\(a\)Backward rotation: frozen\-head accuracy \(solid\) declines withθ\\thetawhile linear probe \(dashed\) stays flat—the source writes to different coordinates but the information persists\. Procrustes \(dotted\) restores access\.\(b\)Full\-state prediction: the joint predictor \(dark blue\) retains near\-exact accuracy at all lags; parameter\-only \(red\) degrades sharply by lag 16, and moments\-only \(light blue\) by lag 1\.
## Appendix CSustained Influence, Storage, and Observer

Turnover decomposition\.In one Qwen2\.5\-0\.5B/LoRA\-r8/AdamW lineage, we extend the same lineage fromT=144T=144toT=160T=160and evaluate the localα=0\\alpha=0work\. A net change of−0\.10037\-0\.10037comprises−1\.14955\-1\.14955from old ancestry being revalued and\+1\.04919\+1\.04919from new source work \(Figure[8](https://arxiv.org/html/2608.20442#A12.F8)a; the displayed components are rounded\)\. Sustained training is therefore a closed loop in which old and new contributions are continually re\-weighed against each other, not a running total of isolated pulses\.

Storage–observer factorial\.In the same lineage, a2×22\\times 2factorial crosses the stored states atT=89T=89andT=96T=96withznullz\_\{\\mathrm\{null\}\}observers at the same two horizons, evaluated on eight held\-out prompts\. The storage effect \(−0\.30291\-0\.30291\) is114×114\\timesthe observer effect \(\+0\.00266\+0\.00266\), with a small interaction \(−0\.00345\-0\.00345; Figure[8](https://arxiv.org/html/2608.20442#A12.F8)b\)\. This bounds these two specific interventions, not observers in general: becauseλt\\lambda\_\{t\}depends onOO, observers can disagree quantitatively or assign opposite signs to the same physical endpoint\. The unrelated\-observer rows in Table[16](https://arxiv.org/html/2608.20442#A16.T16)realize the qualitative case on shared full\-state endpoints, including within\-seed sign changes across observers\. Accordingly, the behavioral panels identify their observers, and we report physical transport separately from behavioral value\.

## Appendix DProofs

### D\.1Derivation of the Response Identity \(Eq\.[6](https://arxiv.org/html/2608.20442#S2.E6)\)

Setup\.Fix the data order, randomness, optimizer code, and observer\. Training is the deterministic recursion of Eq\.[1](https://arxiv.org/html/2608.20442#S2.E1):

gt=Gt​\(St,xt,α\),St\+1=Ut​\(St,gt\),OT​\(α\)=O⁡\(ST​\(α\)\),g\_\{t\}=G\_\{t\}\(S\_\{t\},x\_\{t\};\\alpha\),\\qquad S\_\{t\+1\}=U\_\{t\}\(S\_\{t\},g\_\{t\}\),\\qquad O\_\{T\}\(\\alpha\)=O\(S\_\{T\}\(\\alpha\)\),whereGt​\(St,xt,α\)G\_\{t\}\(S\_\{t\},x\_\{t\};\\alpha\)is the gradient selected by the declared affine gradient\-port pathPαP\_\{\\alpha\}\. All dependence on the source strengthα\\alphaenters through this path\. Assume each map is differentiable along the realized trajectory \(the piecewise\-smooth extension is Appendix[F](https://arxiv.org/html/2608.20442#A6)\)\.

Step 1: tangent recurrence\.Differentiate the compositeSt\+1=Ut​\(St,Gt​\(St,xt,α\)\)S\_\{t\+1\}=U\_\{t\}\(S\_\{t\},G\_\{t\}\(S\_\{t\},x\_\{t\};\\alpha\)\)with respect toα\\alpha\. By the chain rule, withqt=∂St/∂αq\_\{t\}=\\partial S\_\{t\}/\\partial\\alpha,

qt\+1=∂U∂S​qt\+∂U∂g​d​gtd​α=∂U∂S​qt\+∂U∂g​\(∂G∂S​qt\+dt\),q\_\{t\+1\}=\\frac\{\\partial U\}\{\\partial S\}\\,q\_\{t\}\+\\frac\{\\partial U\}\{\\partial g\}\\frac\{dg\_\{t\}\}\{d\\alpha\}=\\frac\{\\partial U\}\{\\partial S\}\\,q\_\{t\}\+\\frac\{\\partial U\}\{\\partial g\}\\\!\\left\(\\frac\{\\partial G\}\{\\partial S\}\\,q\_\{t\}\+d\_\{t\}\\right\),which is Eq\.[3](https://arxiv.org/html/2608.20442#S2.E3)\. The initialization does not depend onα\\alpha, soq0=0q\_\{0\}=0\.

Step 2: general endpoint derivative\.Define the costate backward from the observer:pT=\(∂O/∂ST\)⊤p\_\{T\}=\(\\partial O/\\partial S\_\{T\}\)^\{\\top\}and, fort<Tt<T,

pt=\(∂U∂S\)⊤​pt\+1\+\(∂G∂S\)⊤​λt,λt=\(∂U∂g\)⊤​pt\+1,p\_\{t\}=\\left\(\\frac\{\\partial U\}\{\\partial S\}\\right\)^\{\\\!\\top\}p\_\{t\+1\}\+\\left\(\\frac\{\\partial G\}\{\\partial S\}\\right\)^\{\\\!\\top\}\\lambda\_\{t\},\\qquad\\lambda\_\{t\}=\\left\(\\frac\{\\partial U\}\{\\partial g\}\\right\)^\{\\\!\\top\}p\_\{t\+1\},\(Eqs\.[4](https://arxiv.org/html/2608.20442#S2.E4)and[5](https://arxiv.org/html/2608.20442#S2.E5)\)\. By construction,pt⊤=∂OT/∂Stp\_\{t\}^\{\\top\}=\\partial O\_\{T\}/\\partial S\_\{t\}along the trajectory: it is the adjoint of the state\-to\-endpoint map\. For a general protocol in whichα\\alphacould also enter the initialization, the observer, or the update map directly, the endpoint derivative is

∂αOT=p0⊤​∂αS0\+∂αO\|ST\+∑t=0T−1pt\+1⊤​∂αUt\|S,g\+∑t=0T−1λt⊤​dt\.\\partial\_\{\\alpha\}O\_\{T\}=p\_\{0\}^\{\\top\}\\,\\partial\_\{\\alpha\}S\_\{0\}\+\\partial\_\{\\alpha\}O\\big\|\_\{S\_\{T\}\}\+\\sum\_\{t=0\}^\{T\-1\}p\_\{t\+1\}^\{\\top\}\\,\\partial\_\{\\alpha\}U\_\{t\}\\big\|\_\{S,g\}\+\\sum\_\{t=0\}^\{T\-1\}\\lambda\_\{t\}^\{\\top\}d\_\{t\}\.This follows by telescoping: writingOTO\_\{T\}as a composition ofTTsteps and collecting, for each step, the directα\\alpha\-dependence of that step weighted by the sensitivity of the endpoint to that step’s output\.

Step 3: boundary conditions\.Our experimental protocol holds the initialization fixed \(∂αS0=0\\partial\_\{\\alpha\}S\_\{0\}=0\), declares the observer independently ofα\\alpha\(∂αO\|ST=0\\partial\_\{\\alpha\}O\|\_\{S\_\{T\}\}=0\), and keeps the optimizer code fixed \(∂αUt\|S,g=0\\partial\_\{\\alpha\}U\_\{t\}\|\_\{S,g\}=0\); all source\-coordinate dependence enters throughGGat the gradient port\. The first three terms vanish, leaving

∂αOT=∑t=0T−1⟨λt​\(α\),dt​\(α\)⟩\.\\partial\_\{\\alpha\}O\_\{T\}=\\sum\_\{t=0\}^\{T\-1\}\\langle\\lambda\_\{t\}\(\\alpha\),d\_\{t\}\(\\alpha\)\\rangle\.
Step 4: integration\.OT​\(α\)O\_\{T\}\(\\alpha\)is absolutely continuous inα\\alphaon\[−1,1\]\[\-1,1\]\(a finite composition of smooth maps ofPαP\_\{\\alpha\}\), so by the fundamental theorem of calculus,

OT​\(\+1\)−OT​\(−1\)=∫−11∂αOT​𝑑α=∫−11∑t=0T−1⟨λt​\(α\),dt​\(α\)⟩​𝑑α,O\_\{T\}\(\+1\)\-O\_\{T\}\(\-1\)=\\int\_\{\-1\}^\{1\}\\partial\_\{\\alpha\}O\_\{T\}\\,d\\alpha=\\int\_\{\-1\}^\{1\}\\sum\_\{t=0\}^\{T\-1\}\\langle\\lambda\_\{t\}\(\\alpha\),d\_\{t\}\(\\alpha\)\\rangle\\,d\\alpha,which is Eq\.[6](https://arxiv.org/html/2608.20442#S2.E6)\.□\\square

Remark \(port invariance\)\.The work⟨λt,dt⟩\\langle\\lambda\_\{t\},d\_\{t\}\\rangleis invariant to any state\-independent invertible re\-coordinatization of the gradient port: ifg~=L​g\\tilde\{g\}=Lg, thend~=L​d\\tilde\{d\}=Ld,λ~=L−⁣⊤​λ\\tilde\{\\lambda\}=L^\{\-\\top\}\\lambda, andλ~⊤​d~=λ⊤​d\\tilde\{\\lambda\}^\{\\top\}\\tilde\{d\}=\\lambda^\{\\top\}d\. The decomposition is therefore a property of the implementedG→UG\\to Uinterface, not of the units chosen for it\.

### D\.2The Route Approximation \(Eq\.[7](https://arxiv.org/html/2608.20442#S2.E7)\)

Let two routesR,R′R,R^\{\\prime\}share the same pair of source\-cut states, hence the same midpointSE​\(0\)S\_\{E\}\(0\)and finite ancestry chordAE=SE​\(\+1\)−SE​\(−1\)A\_\{E\}=S\_\{E\}\(\+1\)\-S\_\{E\}\(\-1\)\. Fort≥Et\\geq Ethe routes apply different training, producing different costatespER​\(0\)≠pER′​\(0\)p\_\{E\}^\{R\}\(0\)\\neq p\_\{E\}^\{R^\{\\prime\}\}\(0\)at the cut\. The endpoint response of routeRRto the shared ancestry is, to first order in‖AE‖\\\|A\_\{E\}\\\|,

Δ​OTR=pER​\(0\)⊤​AE\+o⁡\(‖AE‖\),\\Delta O\_\{T\}^\{R\}\\;=\\;p\_\{E\}^\{R\}\(0\)^\{\\top\}A\_\{E\}\\;\+\\;o\(\\\|A\_\{E\}\\\|\),becausepE⊤=∂OT/∂SEp\_\{E\}^\{\\top\}=\\partial O\_\{T\}/\\partial S\_\{E\}is the linearization of the endpoint in the state atEE\. Subtracting,

Δ​OTR−Δ​OTR′≈\(pER​\(0\)−pER′​\(0\)\)⊤​AE\.\\Delta O\_\{T\}^\{R\}\-\\Delta O\_\{T\}^\{R^\{\\prime\}\}\\;\\approx\\;\\bigl\(p\_\{E\}^\{R\}\(0\)\-p\_\{E\}^\{R^\{\\prime\}\}\(0\)\\bigr\)^\{\\top\}A\_\{E\}\.The shared factorAEA\_\{E\}cannot determine the sign of either term alone; the sign is carried by the route\-dependent costates\. The engineered factorial \(§[4\.1](https://arxiv.org/html/2608.20442#S4.SS1)\) demonstrates continuation\-dependent sign assignment with matched finite interventions, for which the exact identity provides the decomposition\. The ordinary\-route panel \(§[4\.2](https://arxiv.org/html/2608.20442#S4.SS2)\) evaluates the first\-order relation above directly at the sign and route\-ordering level; magnitudes involve the neglected higher\-order terms, and the finite\-amplitude remainder is the curvature integralℛR=∫−11\[pER​\(α\)−pER​\(0\)\]⊤​qE​\(α\)​𝑑α\\mathcal\{R\}\_\{R\}=\\int\_\{\-1\}^\{1\}\[p\_\{E\}^\{R\}\(\\alpha\)\-p\_\{E\}^\{R\}\(0\)\]^\{\\top\}q\_\{E\}\(\\alpha\)\\,d\\alphaaround the center stateα=0\\alpha=0\.

### D\.3Remarks: Scope of the Identity

Four observations describe the scope of the identity\.

Finite\-horizon non\-coalescence\.More generally, any matched continuation composed of injective complete\-state updates preserves distinctions between cut states\. For the standard non\-AMSGrad AdamW update used here, fix the future batches, RNG stream, and schedule, and write

mt\+1\\displaystyle m\_\{t\+1\}=β1​mt\+\(1−β1\)​gt​\(wt\),\\displaystyle=\\beta\_\{1\}m\_\{t\}\+\(1\-\\beta\_\{1\}\)g\_\{t\}\(w\_\{t\}\),vt\+1\\displaystyle v\_\{t\+1\}=β2​vt\+\(1−β2\)​gt​\(wt\)⊙2,\\displaystyle=\\beta\_\{2\}v\_\{t\}\+\(1\-\\beta\_\{2\}\)g\_\{t\}\(w\_\{t\}\)^\{\\odot 2\},wt\+1\\displaystyle w\_\{t\+1\}=at​wt−ηt​Dt​\(vt\+1,τt\)​mt\+1,\\displaystyle=a\_\{t\}w\_\{t\}\-\\eta\_\{t\}D\_\{t\}\(v\_\{t\+1\},\\tau\_\{t\}\)m\_\{t\+1\},whereat=1−ηt​λwda\_\{t\}=1\-\\eta\_\{t\}\\lambda\_\{\\mathrm\{wd\}\}and the known diagonal mapDtD\_\{t\}includes bias correction andϵ\\epsilonregularization\. Whenβ1\\beta\_\{1\},β2\\beta\_\{2\}, andata\_\{t\}are nonzero, the preceding state is recovered explicitly:

wt\\displaystyle w\_\{t\}=at−1​\[wt\+1\+ηt​Dt​\(vt\+1,τt\)​mt\+1\],\\displaystyle=a\_\{t\}^\{\-1\}\\\!\\left\[w\_\{t\+1\}\+\\eta\_\{t\}D\_\{t\}\(v\_\{t\+1\},\\tau\_\{t\}\)m\_\{t\+1\}\\right\],gt\\displaystyle g\_\{t\}=Gt​\(wt,xt,τt\),\\displaystyle=G\_\{t\}\(w\_\{t\},x\_\{t\};\\tau\_\{t\}\),mt\\displaystyle m\_\{t\}=mt\+1−\(1−β1\)​gtβ1,vt=vt\+1−\(1−β2\)​gt⊙2β2\.\\displaystyle=\\frac\{m\_\{t\+1\}\-\(1\-\\beta\_\{1\}\)g\_\{t\}\}\{\\beta\_\{1\}\},\\qquad v\_\{t\}=\\frac\{v\_\{t\+1\}\-\(1\-\\beta\_\{2\}\)g\_\{t\}^\{\\odot 2\}\}\{\\beta\_\{2\}\}\.Thus each time\-indexed update on\(w,m,v\)\(w,m,v\)is injective, and two distinct matched cut states in these optimizer blocks cannot coalesce after any finite continuation\. The classical momentum\-SGD update used here has the same property: fromut\+1=μ​ut\+gt​\(wt\)u\_\{t\+1\}=\\mu u\_\{t\}\+g\_\{t\}\(w\_\{t\}\)andwt\+1=wt−ηt​ut\+1w\_\{t\+1\}=w\_\{t\}\-\\eta\_\{t\}u\_\{t\+1\}, one recoverswt=wt\+1\+ηt​ut\+1w\_\{t\}=w\_\{t\+1\}\+\\eta\_\{t\}u\_\{t\+1\}and thenut=\[ut\+1−gt​\(wt\)\]/μu\_\{t\}=\[u\_\{t\+1\}\-g\_\{t\}\(w\_\{t\}\)\]/\\muwhenμ≠0\\mu\\neq 0\[[33](https://arxiv.org/html/2608.20442#bib.bib20)\]\. Plain SGD remains covered by the transport–valuation identity, but its parameter\-only update need not be injective without additional conditions on the gradient map and step size\. Non\-coalescence is an exact\-arithmetic statement about the complete optimizer state; it does not give a lower bound on ancestry magnitude or behavioral response\.

Experimental separation of the modules\.The interventions isolate the roles of the source term, optimizer map, and observer\. The matched factorial varies source content under matched routes; the turnover decomposition separates old\-ancestry revaluation from new source work \(−1\.150\-1\.150vs\.\+1\.049\+1\.049\); state surgery distinguishes ancestry stored in parameter and optimizer blocks; and the storage–observer factorial varies the observer on shared physical endpoints\.

Observer relativity\.Two observers with endpoint rowsrrand−r\-rinduce identical observer\-free trajectory descriptors but opposite signed responses wheneverr⊤​qT≠0r^\{\\top\}q\_\{T\}\\neq 0; no observer\-free functional of the trajectory determines the sign\. The storage–observer factorial and Table[16](https://arxiv.org/html/2608.20442#A16.T16)provide the empirical counterpart: the same physical endpoint can receive quantitatively or qualitatively different values under different declared observers \(§[C](https://arxiv.org/html/2608.20442#A3)\)\.

Coordinate invariance and operational minimality\.Under a smooth invertible reparameterizationS′=ψ⁡\(S\)S^\{\\prime\}=\\psi\(S\), the tangent and costate transform asq′=D​ψ​qq^\{\\prime\}=D\\psi\\,qandp′=D​ψ−⁣⊤​pp^\{\\prime\}=D\\psi^\{\-\\top\}p, sop′⁣⊤​q′=p⊤​qp^\{\\prime\\top\}q^\{\\prime\}=p^\{\\top\}q\. Paired work and response totals are invariant, while the block namesww,mm, andvvdepend on the chosen chart\. If two reachable cut states produce the same declared physical outputs under every allowed source\-free suffix, they are equivalent for that intervention family\. The minimal causal state is the source\-reachable set modulo this future indistinguishability\. Thew\+mw\{\+\}mresult therefore identifies the sufficient block set in the implemented chart and source\-free continuation family\.

## Appendix EThe Full Computation

Algorithm[1](https://arxiv.org/html/2608.20442#alg1)computes the response derivative for one source strengthα\\alpha\. The forward sweep runs alongside ordinary training and maintains the tangentqtq\_\{t\}\(Eq\.[3](https://arxiv.org/html/2608.20442#S2.E3)\); the backward sweep replays the trajectory in reverse and maintains the costateptp\_\{t\}\(Eq\.[5](https://arxiv.org/html/2608.20442#S2.E5)\)\. Their meeting point is the per\-step work𝒲t=⟨λt,dt⟩\\mathcal\{W\}\_\{t\}=\\langle\\lambda\_\{t\},d\_\{t\}\\rangle, whose sum is∂αOT\\partial\_\{\\alpha\}O\_\{T\}\. The finite endpoint difference in Eq\.[6](https://arxiv.org/html/2608.20442#S2.E6)is the integral of that returned sum overα\\alpha\. This equality is exact; any finite\-node quadrature used to evaluate it has a separate approximation error that must be assessed for the selected path\. No empirical panel estimates the full finite\-amplitude integral by fixed\-node quadrature: the local\-work panels evaluateα=0\\alpha=0, while finite\-amplitude panels use matched interventions or finite Shapley decompositions\. The compact midpoint predictor of Section[4\.2](https://arxiv.org/html/2608.20442#S4.SS2)is evaluated explicitly as a first\-order finite\-amplitude approximation\. Intervals containing persistent kinks or clipping events are integrated by matched finite interventions, avoiding possible distortions from pointwise derivatives \(Appendix[F](https://arxiv.org/html/2608.20442#A6)\)\.

Algorithm 1Full\-state adjoint derivative at one source strength1:training data stream

\{xt\}\\\{x\_\{t\}\\\}, source strength

α\\alpha, observer

OO, horizon

TT
2:Forward sweep \(with training\):

q0←0q\_\{0\}\\leftarrow 0
3:for

t=0,…,T−1t=0,\\dots,T\-1do

4:

gt←Gt​\(St,xt,α\)g\_\{t\}\\leftarrow G\_\{t\}\(S\_\{t\},x\_\{t\};\\alpha\);

St\+1←U⁡\(St,gt\)S\_\{t\+1\}\\leftarrow U\(S\_\{t\},g\_\{t\}\)⊳\\trianglerightordinary training step

5:

dt←∂Gt/∂αd\_\{t\}\\leftarrow\\partial G\_\{t\}/\\partial\\alpha⊳\\trianglerightsource perturbation, state held fixed

6:

qt\+1←∂U∂S​qt\+∂U∂g​\(∂G∂S​qt\+dt\)q\_\{t\+1\}\\leftarrow\\frac\{\\partial U\}\{\\partial S\}q\_\{t\}\+\\frac\{\\partial U\}\{\\partial g\}\\big\(\\frac\{\\partial G\}\{\\partial S\}q\_\{t\}\+d\_\{t\}\\big\)⊳\\trianglerightJVPs; Eq\.[3](https://arxiv.org/html/2608.20442#S2.E3)

7:record

\(St,gt,dt\)\(S\_\{t\},g\_\{t\},d\_\{t\}\)
8:endfor

9:Backward sweep:

pT←\(∂O/∂ST\)⊤p\_\{T\}\\leftarrow\(\\partial O/\\partial S\_\{T\}\)^\{\\top\}
10:for

t=T−1,…,0t=T\-1,\\dots,0do

11:

λt←\(∂U∂g\)⊤​pt\+1\\lambda\_\{t\}\\leftarrow\\big\(\\frac\{\\partial U\}\{\\partial g\}\\big\)^\{\\top\}p\_\{t\+1\}⊳\\trianglerightVJP; Eq\.[4](https://arxiv.org/html/2608.20442#S2.E4)

12:

𝒲t←⟨λt,dt⟩\\mathcal\{W\}\_\{t\}\\leftarrow\\langle\\lambda\_\{t\},d\_\{t\}\\rangle⊳\\trianglerightper\-step work

13:

pt←\(∂U∂S\)⊤​pt\+1\+\(∂G∂S\)⊤​λtp\_\{t\}\\leftarrow\\big\(\\frac\{\\partial U\}\{\\partial S\}\\big\)^\{\\top\}p\_\{t\+1\}\+\\big\(\\frac\{\\partial G\}\{\\partial S\}\\big\)^\{\\top\}\\lambda\_\{t\}⊳\\trianglerightEq\.[5](https://arxiv.org/html/2608.20442#S2.E5)

14:endfor

15:return

\{𝒲t​\(α\)\}t=0T−1\\\{\\mathcal\{W\}\_\{t\}\(\\alpha\)\\\}\_\{t=0\}^\{T\-1\},

∂αOT=∑t𝒲t​\(α\)\\partial\_\{\\alpha\}O\_\{T\}=\\sum\_\{t\}\\mathcal\{W\}\_\{t\}\(\\alpha\)⊳\\trianglerightintegrate overα\\alphafor Eq\.[6](https://arxiv.org/html/2608.20442#S2.E6)

Cost and how it scales\.The method’s scaling cost determines where it can be applied\. Per source strengthα\\alpha, the forward sweep adds a fixed number of JVP evaluations per update on top of ordinary training and the backward sweep adds a fixed number of VJP evaluations per update; the extra computation is therefore determined by those Jacobian products\. The binding constraint is memory\. The backward sweep replays the trajectory in reverse and therefore needs access to the recorded per\-step quantities\(St,gt,dt\)\(S\_\{t\},g\_\{t\},d\_\{t\}\), which isO⁡\(T\)O\(T\)storage in the horizon, each item the size of a full trainer state\. At our scale \(LoRA parameters on models up to 1\.1B, horizons of order10210^\{2\}updates\), exact replay was feasible\. Gradient checkpointing trades recomputation for storage; segmented approximations of the kind SOURCE uses\[[3](https://arxiv.org/html/2608.20442#bib.bib24)\]trade exactness for checkpoint\-level storage\. Theα\\alpha\-integral multiplies the cost by the number of evaluated source strengths; the required numerical resolution is a property of the selected path, not of the identity\.

The interventional experiments set the larger compute cost in our runs: a block\-transplant factorial forks matched states into eight whole\-block combinations and re\-runs the suffix for each seed and route\. The costate predictor is cheaper than the full decomposition because it needs one center\-state costate per route and noα\\alpha\-sweep\.

## Appendix FNonsmooth Events and Gradient Clipping

The response identity is exact for smooth trainers and extends to the piecewise\-smooth primitives used here through the conservative\-Jacobian and path\-differentiability calculus of[7](https://arxiv.org/html/2608.20442#bib.bib25), building on classical piecewise\-smooth automatic differentiation\[[21](https://arxiv.org/html/2608.20442#bib.bib19)\]\. For locally Lipschitz, path\-differentiable primitives, the state pathα↦St​\(α\)\\alpha\\mapsto S\_\{t\}\(\\alpha\)is absolutely continuous, the tangent and costate recurrences hold with conservative\-Jacobian selections from the same executed program, and the response integral holds almost everywhere inα\\alpha\. The AdamW implementation used here is covered on reachable gradient histories, wherev^t=0\\hat\{v\}\_\{t\}=0impliesm^t=0\\hat\{m\}\_\{t\}=0and theϵ\\epsilon\-regularized ratio is locally Lipschitz; the ambient AdamW map lacks local Lipschitz guarantees atv=0,m≠0v=0,m\\neq 0\. Continuous clipping crossings add no jump term because the state is continuous; a discontinuous branch would require explicit jump terms, and none occurs in our fixed\-schedule trainers\. Primal and dual sweeps use the same derivative selection at any kink\. Appendix[O](https://arxiv.org/html/2608.20442#A15)checks their duality and finite\-difference accuracy in active and inactive clipping regimes; crossings are evaluated one\-sided rather than with a centered stencil\. Intervals on which an event persists are integrated by matched finite interventions, avoiding possible distortions from pointwise derivatives across piecewise regions\.

## Appendix GOnline Costate Approximation

The approximations that would change the memory asymptotics of Appendix[E](https://arxiv.org/html/2608.20442#A5)are truncation of the backward horizon and low\-rank or sketched representations of the costate\. We evaluate horizon truncation on the fully specified 42\-route panel \(§[4\.2](https://arxiv.org/html/2608.20442#S4.SS2)\): every tested shorter horizon remains below the full\-horizon predictor in every seed, while the full horizon reproduces the reference predictions exactly \(per\-HHscores in Appendix[J](https://arxiv.org/html/2608.20442#A10)\)\.

## Appendix HExtended Relation to Prior Work

Divergence tokens and transferred value\.[44](https://arxiv.org/html/2608.20442#bib.bib2)locate the carrier in the data: the token positions where teacher and reference distributions differ\. Our decomposition addresses the complementary trainer\-side question\. Divergence identifies where the signal can enter; the costate describes how later training values the resulting gradient perturbation\.

Optimizer requirements\.[6](https://arxiv.org/html/2608.20442#bib.bib3)report that steering\-vector distillation requires an adaptive optimizer in their LLM protocol\. Our momentum\-SGD replications reproduce the relay with velocity in place of Adam’s first moment \(§[3](https://arxiv.org/html/2608.20442#S3), Appendix[I](https://arxiv.org/html/2608.20442#A9)\), showing that the transport path does not require adaptive scaling; optimizer choice changes its realized amplitude\.

Parameterization\.[37](https://arxiv.org/html/2608.20442#bib.bib6)find full fine\-tuning eliminates transfer in their setting\. Our MNIST experiments use no LoRA, and a descriptive matched\-update\-norm Qwen control includes LoRA, direct\-adapter, and full\-parameter arms \(Table[13](https://arxiv.org/html/2608.20442#A14.T13)\)\. These results show that the mechanism does not require LoRA\.

Trainer\-side channels\.[34](https://arxiv.org/html/2608.20442#bib.bib28)constrains where an auditable channel can sit\. Our decomposition distinguishes trainer state from future\-value sensitivity inside the training process\.[2](https://arxiv.org/html/2608.20442#bib.bib52)compare data\-mediated transfer conditions across task structure, prompt opportunity, and teacher/data channels\. Their emphasis on dataset structure is complementary to our optimizer\-centric account and may help explain why physical transport need not yield stable behavior\.

Larger\-model and agent settings\.[31](https://arxiv.org/html/2608.20442#bib.bib30)quantify subliminal behavioral transfer in 7B language\-model distillation, and[14](https://arxiv.org/html/2608.20442#bib.bib29)study unsafe\-behavior transfer in AI\-agent distillation\. These results broaden the settings in which the phenomenon has been measured; our experiments address the complementary trainer\-state mechanism under controlled interventions\.

Seed\-induced uniqueness and representational substrate\.[39](https://arxiv.org/html/2608.20442#bib.bib8)report that transfer strength tracks alignment inside a trait\-discriminative subspace rather than global representational similarity, so that students differing only in initialization seed leak substantially less than same\-seed students even at global CKA above0\.90\.9\. That result concerns the representational substrate a trait needs in order to be readable; ours concerns where the trait is physically held while training continues\. The two are complementary: seed\-specific geometry may help explain why the allocation of source response across parameters and optimizer slots varies widely across seeds \(Appendix[L](https://arxiv.org/html/2608.20442#A12)\)\.

Metric circularity\.[10](https://arxiv.org/html/2608.20442#bib.bib9)identify a shared\-denominator artifact in regressions involving log probability ratios\. Our primary estimand is a matched source–route interaction, and the generated\-choice panel evaluates its route ordering without that regression \(§[5](https://arxiv.org/html/2608.20442#S5)\)\.

Compatible output heads\.[8](https://arxiv.org/html/2608.20442#bib.bib5)show that shared output\-layer structure between teacher and student is a precondition for transfer\. Our backward\-rotation experiment \(§[B](https://arxiv.org/html/2608.20442#A2)\) builds on this: rotating the hidden backward map by angleθ\\thetawhile preserving forward outputs shows where the source writes rather than what it writes, and a Procrustes realignment restores access\.

Temporal localization and liminal training\.[50](https://arxiv.org/html/2608.20442#bib.bib7)localize trait acquisition to an early nonlinear transition and mitigate it with an annealed KL regularizer\. Their mitigation acts during source exposure\. Our source\-position ablation asks whether the transport–valuation topology recurs when the source appears later: all four tested positions preserve the route ordering at 77–100% of the reference amplitude \(Figure[4](https://arxiv.org/html/2608.20442#S4.F4)c\)\.

Emergent misalignment\.[5](https://arxiv.org/html/2608.20442#bib.bib33)show that fine\-tuning on a narrow task—writing insecure code—can induce broadly misaligned behavior across unrelated domains\. There the disposition is not an explicit target but the misaligning source data is overtly present\. That work studies why a narrow source generalizes broadly; our decomposition studies how update contributions are stored and later valued\.

Non\-semantic transfer and task arithmetic\.[11](https://arxiv.org/html/2608.20442#bib.bib31)showed that pre\-training transfer can arise from artificial, non\-semantic properties of the data—the broader question of what training data can transmit besides its overt content\. Trait vectors obtained by weight\-space merging\[[46](https://arxiv.org/html/2608.20442#bib.bib32)\]—following the task\-arithmetic line, in which differences of fine\-tuned weights are added and negated as vectors\[[29](https://arxiv.org/html/2608.20442#bib.bib44)\]—place behavioral traits in parameter space; our transplant results show that ancestry can also reside transiently in optimizer slots before it becomes parameter\-visible\.

Optimizer memory\.[45](https://arxiv.org/html/2608.20442#bib.bib37)measure non\-Markovian memory in SGD training directly—via a back\-flow\-of\-distinguishability witness—and find that it scales with momentum and collapses under a causal break that resets optimizer state, which agrees with our transplant results on where the memory physically sits\.[9](https://arxiv.org/html/2608.20442#bib.bib38)give a complementary theoretical account, showing that an optimizer’s accumulated history acts as an implicit modification of the loss being descended, so later updates are steered by stored state and not by the current gradient alone\.

Order dependence\.[48](https://arxiv.org/html/2608.20442#bib.bib39)shows that which of two training phases comes first changes the final transfer outcome, and that the change is predicted by the Lie\-bracket commutator of the two update fields\. Their object is the geometric non\-commutativity of update order; ours is the causal valuation of a fixed perturbation by its continuation\. Both say that when an update happens relative to the rest of training is part of what that update does\.

Unlearning persistence\.[27](https://arxiv.org/html/2608.20442#bib.bib34)show that apparently unlearned content can return after benign follow\-up fine\-tuning, and[49](https://arxiv.org/html/2608.20442#bib.bib36)find that forgetting is often suppression near the output rather than erasure\.[20](https://arxiv.org/html/2608.20442#bib.bib35)use data attribution to predict the retrained\-without\-xxmodel and then fine\-tune toward it\. Our first\-moment release result is a trainer\-state analogue: ancestry can have zero forward\-visible effect at the cut and still be released by later source\-free updates\.

Persistence across subsequent training\.Backdoored behaviors can survive supervised fine\-tuning, RL, and adversarial training\[[28](https://arxiv.org/html/2608.20442#bib.bib45)\], and[13](https://arxiv.org/html/2608.20442#bib.bib41)engineer poisoned gradients to persist through continual fine\-tuning\. Those studies examine intentionally persistent behavior; our experiments characterize how source\-induced ancestry is relayed by trainer state\.

Other related methods\.Feedback alignment showed learning survives replaced backward maps\[[38](https://arxiv.org/html/2608.20442#bib.bib10),[43](https://arxiv.org/html/2608.20442#bib.bib12),[19](https://arxiv.org/html/2608.20442#bib.bib13)\]; critical learning periods\[[1](https://arxiv.org/html/2608.20442#bib.bib43)\]established that transient early\-training deficits become permanently harder to correct\. LoRA and its variants\[[26](https://arxiv.org/html/2608.20442#bib.bib15),[52](https://arxiv.org/html/2608.20442#bib.bib11),[23](https://arxiv.org/html/2608.20442#bib.bib16),[24](https://arxiv.org/html/2608.20442#bib.bib17)\]provide the parameterization used in our Qwen experiments\.

## Appendix IValidation Panels

Generated\-choice panel\.Seven seeds run the factorial with a pairwise\-choice prompt grid held out from all geometry construction, scored one\-shot at temperature 0\. Both adjacent route contrasts are positive in generated choices in 7/7 seeds \(one\-sided sign test, Holm\-adjustedp=0\.016p=0\.016each\)\. The panel validates the ordering the likelihood observer induces over routes, not the calibration of its magnitudes\.

Resolution of then=7n=7sign test\.A unanimous 7/7 result has one\-sided rawp=0\.0078p=0\.0078; Holm adjustment over the two adjacent contrasts givesp=0\.016p=0\.016\.

Momentum\-SGD extension of the factorial\.The same Qwen protocol with AdamW replaced by momentum SGD \(momentum 0\.9, learning rate chosen from surface\-task progress\) reproduces the ordered factorial interaction in 3/3 seeds, with the same ordering recovered on an independent holdout bank\. Each adjacent contrast has one\-sided rawp=0\.125p=0\.125; Holm adjustment over the two contrasts givesp=0\.25p=0\.25each, so this panel is descriptive\. The mechanism can therefore run with momentum velocity in place of Adam moments\.

Two\-trait matched factorial\.A matched Qwen factorial uses independently written 128\-row native sources for the owl and blue traits under shared prompts, geometry, and horizons\. The blue source reaches 6/7 seeds on both adjacent route contrasts, with a single seed reversing both; that reversal does not recur under the matched owl source in the same cell\. The separate source\-scale titration in Appendix[J](https://arxiv.org/html/2608.20442#A10)is consistent with comparable small\-corpus scales falling in the low\-signal regime\.

Optimizer\-slot lineage\.The panel uses one seed per model, with two nonterminal cuts and three routes per cut\. Each source\-induced ancestry is split according to where it lives in the AdamW state—the parametersww, the first momentmm, or the second momentvv—and each block is propagated through the same future separately\. We score a blockccby its*block\-restricted recovery error*SSE/ZERO⁡\(c​\-only\)\\mathrm\{SSE/ZERO\}\(c\\text\{\-only\}\), normalized by the error of the zero predictor\. These nonnegative prediction errors are not additive mass fractions: interactions between blocks are not credited to any single block\. Separate signed allocations can be negative\. In Qwen the dominant pair is the parameters together with the first moment, reachingSSE/ZERO≈5×10−4\\mathrm\{SSE/ZERO\}\\approx 5\\times 10^\{\-4\}, whereas in SmolLM2\-135M it is the two moments instead\. The dominant pair therefore differs between these two cells\.

SmolLM2 cancellation cell\.A finite Shapley decomposition over the three source\-active updates assigns−1\.21×10−3\-1\.21\\times 10^\{\-3\}to the first update and\+6\.0×10−4\+6\.0\\times 10^\{\-4\}and\+4\.1×10−4\+4\.1\\times 10^\{\-4\}to the second and third, against a net endpoint response of−2\.0×10−4\-2\.0\\times 10^\{\-4\}\. In other words, 90\.9% of the gross allocation cancels between updates, so the near\-zero net response results from the cancellation of substantial opposing contributions rather than the absence of transport\. The allocation is a signed decomposition along the declared path, not a fraction of physical mass, which is why individual terms can be negative\.

Boundary\-case response summaries\.In TinyLlama, finite interventions attribute 62\.7% to 83\.6% of the gross finite allocation to cancellation, depending on the cell, which is why the net behavioral response has no stable sign across seeds\. In Llama\-3\.2\-1B the seed\-mean±\\pmsample\-SD animal\-preference response is−0\.012±0\.013\-0\.012\\pm 0\.013under AdamW and−0\.005±0\.005\-0\.005\\pm 0\.005under momentum SGD across three seeds each, and the matched color\-preference response is\+0\.016±0\.025\+0\.016\\pm 0\.025: weak and seed\-heterogeneous in every case, even though every physical stage of the mechanism is observed in the same runs\.

## Appendix JRoute\-Aware Baselines and Random\-Plane Control

For seedss,MSEs\\operatorname\{MSE\}\_\{s\}is computed over the concatenated prompt\-level response vectors from its six routes\. The baseline\-margin score isGs∗=minb⁡\{MSEs⁡\(b\)−MSEs⁡\(Δ^\)\}G\_\{s\}^\{\\ast\}=\\min\_\{b\}\\\{\\operatorname\{MSE\}\_\{s\}\(b\)\-\\operatorname\{MSE\}\_\{s\}\(\\widehat\{\\Delta\}\)\\\}over the ZERO, source\-cut linear, and route\-constant comparators\. ThusGs∗\>0G\_\{s\}^\{\\ast\}\>0exactly when the costate predictor beats all three\.

Receding\-horizon predictors\.Each variant propagates the same finite ancestry chordAEA\_\{E\}through only the firstHHof the eight suffix updates and reads the observer at the post\-HHcenter state, which is the forward dual of anHH\-step online costate \(Appendix[G](https://arxiv.org/html/2608.20442#A7)\)\. MeanGs∗G^\{\\ast\}\_\{s\}\(×10−3\\times 10^\{\-3\}\) is−2\.13\-2\.13atH=1H=1,−1\.55\-1\.55atH=2H=2,−0\.37\-0\.37atH=4H=4, and\+0\.15\+0\.15atH=8H=8; the count of seeds withGs∗\>0G^\{\\ast\}\_\{s\}\>0is 0/7, 0/7, 0/7, and 7/7\. Truncation degrades the predictor smoothly, but no truncated horizon beats the three primary baselines\.

Forward\-only alignment\.The alignment predictor scores each route by∑t⟨AE\(w\),g¯t⟩\\sum\_\{t\}\\langle A\_\{E\}^\{\(w\)\},\\bar\{g\}\_\{t\}\\rangle, whereAE\(w\)A\_\{E\}^\{\(w\)\}is the parameter block of the cut\-state ancestry andg¯t\\bar\{g\}\_\{t\}is that route’s center\-state parameter gradient\. This route\-aware quantity requires no backward pass\. It matches observed route\-mean signs in 17 of 42 cells \(per\-seed range 0/6 to 6/6\), and its per\-seed Spearman correlations against observed endpoints are inconsistent in sign\. Route\-dependent forward signal alone does not recover the valuation\.

Schedule statistics\.Leave\-one\-route\-out sibling means, nearest\-neighbor, and kernel\-weighted predictors are constructed from the observed endpoints of the other five routes in the same seed\. The costate predictor has lower error in 6/7 seeds for the sibling\-mean and kernel predictors and in 7/7 for nearest\-neighbor\. Exact two\-sided Mantel permutation tests givep≥0\.026p\\geq 0\.026with association signs mixed across seeds; they do not show a consistent schedule–endpoint association\.

Random\-plane control\.The control in §[4\.1](https://arxiv.org/html/2608.20442#S4.SS1)draws one fixed random direction, orthogonalizes it against bothϕ^\\hat\{\\phi\}and the corpus gradient contrastuu, and rescales it to‖ϕ‖\\\|\\phi\\\|\. Across seeds, the adjacent contrasts are\|d1\|=0\.136±0\.071\|d\_\{1\}\|=0\.136\\pm 0\.071and\|d2\|=0\.149±0\.056\|d\_\{2\}\|=0\.149\\pm 0\.056\(mean±\\pmsample SD; Cohen\|dz\|=1\.9\|d\_\{z\}\|=1\.9and2\.72\.7\), both sign\-coherent in 7/7 seeds\. The realized sign is opposite to the observer\-informed plane\. Under a random plane the assignment of theS±S^\{\\pm\}labels to the two half\-spaces carries no observer meaning, so only sign*coherence*across seeds is interpretable\.

Source\-scale titration\.The ladder uses fixed paired\-row subsamples of 128, 256, 512, and 1024 examples from the full 1831\-row owl corpus; route generation, geometry, optimizer, horizons, and observer remain fixed\. MeanGs∗G^\{\\ast\}\_\{s\}\(×10−3\\times 10^\{\-3\}\) is−0\.69\-0\.69\(128\),\+0\.07\+0\.07\(256\),−0\.03\-0\.03\(512\),\+0\.15\+0\.15\(1024\), and\+0\.08\+0\.08\(1831\)\. The independently written native 128\-row protocol gives−1\.71\-1\.71\. Between 256 and 512 rows the predictor is near the zero\-margin boundary and seed\-unstable, where the route\-dependent signal approaches the baseline/noise floor\.

## Appendix KImplementation Checks

Appendix[O](https://arxiv.org/html/2608.20442#A15)reports numerical derivative checks for AdamW and momentum SGD\.

Matched state in the transplant\.The surgery holds auxiliary stateτ\\tau, data order, and the RNG stream fixed across arms, so it isolates changes in\(w,m,v\)\(w,m,v\)under a common continuation\.

## Appendix LExtended Figures and Tables

Figure 7:Validity domain of the compact chord predictor\(§[4\.2](https://arxiv.org/html/2608.20442#S4.SS2)\)\. Each dot is one seed’s baseline\-margin scoreGs∗G^\{\\ast\}\_\{s\}\(positive = the costate predictor beats ZERO, source\-only, and route\-constant competitors\); horizontal bars are seed means\.\(a\)Source\-scale titration: fixed paired\-row subsamples from the full 1831\-row corpus, with all other protocol elements unchanged\. The predictor is below the baselines at 128 rows, transitions through a low\-signal band at 256–512, and wins in 7/7 seeds at 1024 and above\. Open diamonds: the independently generated native 128\-row protocol is also below the baselines\.\(b\)Receding\-horizon ladder on the same ordinary\-route panel: truncating the propagated future toH∈\{1,2,4\}H\\in\\\{1,2,4\\\}of the 8 suffix updates gives negative margin in every seed\.Figure 8:Sustained influence and storage–observer separation\.\(a\)Old\-ancestry revaluation \(−1\.14955\-1\.14955\) outweighs new source work \(\+1\.04919\+1\.04919\), producing a net local\-response change of−0\.10037\-0\.10037at underlying precision\.\(b\)Storage is114×114\\timesthe observer effect\. Solid bars show absolute scalar effects; light bars show prompt\-vector norms\.Figure 9:Five\-trait transplant across Qwen2\.5\-0\.5B and Llama\-3\.2\-1B\.Thew\+mw\{\+\}mtransplant reproduces the full descendant in all 15 primary seed–trait cells: Qwen owl, cat, red, and oak, plus Llama blue\.\(a\)Physical descendant error\.\(b\)Hidden\-state descendant error\.Figure 10:Generalization and boundary\.\(a\)CNN replication across optimizers; block allocations are seed\-dependent\.\(b\)One SmolLM2 cell: 90\.9% step\-wise cancellation, not absence\.\(c\)In TinyLlama, response signs follow the transplanted gradient history, but the prompt\-level response pattern does not transfer\. Intervals are paired\-prompt 95% bootstrap intervals\.Table 4:Separate Qwen common\-ancestry route\-replacement cohort \(n=7n=7\): paired contrastsDi=\(yi−−yi\+\)/2D\_\{i\}=\(y\_\{i\}^\{\-\}\-y\_\{i\}^\{\+\}\)/2, whereyi±y\_\{i\}^\{\\pm\}is seedii’s response underS±S^\{\\pm\}\. All positive; two\-sided sign\-testp=0\.016p=0\.016\.Table 5:Source\-position ablation: route topology recurs across positions\.Table 6:SmolLM2 source\-step Shapley allocations\. Cross\-step cancellation: 90\.9%\.Table 7:TinyLlama finite anatomy: stable early/late opposition with 62–84% cancellation\.Table 8:CNN AdamW: block allocations vary widely across seeds\. Param andm1m\_\{1\}are signed response shares normalized by the net response; omitted block and interaction terms are also signed, so the displayed shares need not sum to one\.Table 9:CNN momentum SGD: relay topology is stable across seeds\. Src\-cut is the source\-cut response norm in the experiment’s native units\. Vel\. and Param\. are signed fractions of the terminal response and sum to one\.
## Appendix MExperimental Protocols

### M\.1Qwen LLM Protocol

Qwen2\.5\-0\.5B\-Instruct is trained with rank\-8 LoRA and AdamW \(learning rate2×10−42\\times 10^\{\-4\},β=\(0\.9,0\.999\)\\beta=\(0\.9,0\.999\)\)\. The paired owl/neutral data contain approximately 1,831 bare\-number completions per arm, generated by an owl LoRA teacher and the base\-neutral teacher; training uses bare text without a chat template\. In the factorial, a shared prefix ends at update 79, source training occupies updates 80–87, and source\-free routes occupy updates 88–95, with terminalT=96T=96\. “Neutral” denotes the matched source\-coordinate control, not absence of training\. Behavior is the target\-minus\-reference conditional log\-likelihood on a frozen prompt bank\. The transplant protocol uses 40 prefix, 4 source, and 4 source\-free updates\. The momentum extension uses coefficient 0\.9\.

### M\.2Canonical Source Path

For a declared routeRR, letgO,tR​\(S\)g\_\{O,t\}^\{R\}\(S\)andgN,tR​\(S\)g\_\{N,t\}^\{R\}\(S\)be its owl and neutral corpus gradients at stateSS\. In the common\-cut factorial of §[4\.1](https://arxiv.org/html/2608.20442#S4.SS1), these source\-window gradients are shared acrossRRand only the suffix map changes after cutEE\. The source coordinate enters at the optimizer’s gradient port as

gtα,R​\(S\)=gO,tR​\(S\)\+gN,tR​\(S\)2\+α​gO,tR​\(S\)−gN,tR​\(S\)2,−1≤α≤1\.g\_\{t\}^\{\\alpha,R\}\(S\)=\\frac\{g\_\{O,t\}^\{R\}\(S\)\+g\_\{N,t\}^\{R\}\(S\)\}\{2\}\+\\alpha\\,\\frac\{g\_\{O,t\}^\{R\}\(S\)\-g\_\{N,t\}^\{R\}\(S\)\}\{2\},\\qquad\-1\\leq\\alpha\\leq 1\.Thusα=\+1\\alpha=\+1recovers the owl arm andα=−1\\alpha=\-1the neutral arm within the same route, consistent with the integral bounds of Eq\.[6](https://arxiv.org/html/2608.20442#S2.E6)\. The matched finite contrast defines a source\-content differential within a compatible interface, independent of literal source absence\. This is the declared path for allα\\alpha\-integrals; alternative paths sharing the same endpoints leave totals unchanged but redistribute per\-step attribution \(Appendix[M\.3](https://arxiv.org/html/2608.20442#A13.SS3)\)\.

### M\.3Route Construction

Engineered diagnostic routes\.S±S^\{\\pm\}spane±∝ϕ^±u^e\_\{\\pm\}\\propto\\hat\{\\phi\}\\pm\\hat\{u\}\(with a shared completion directionzz\), whereϕ\\phiis the frozen behavior\-readout direction from the declared prompt bank anduuthe owl–neutral corpus gradient contrast;znullz\_\{\\mathrm\{null\}\}is orthogonal to both\. The geometry derives strictly from readout and gradient contrast directions, without referencing the costate or endpoint outcome\. The random\-plane control \(§[4\.1](https://arxiv.org/html/2608.20442#S4.SS1)\) replacesϕ^\\hat\{\\phi\}with a fixed random direction orthogonalized againstϕ^\\hat\{\\phi\}anduuand rescaled to‖ϕ‖\\\|\\phi\\\|\.

Ordinary routes\.Each seed uses six random future data schedules, and all 42 routes are retained\. The source\-cut linear comparator reads the ancestry chord atEEwithout suffix propagation; the route\-constant comparator assigns the across\-route mean costate prediction to every route\.

Attribution\-path robustness\.Two exact nested coalition traversals of the eight source updates \(ascending, descending\) share identical endpoints: the route topologyS\+<znull<S−S^\{\+\}<z\_\{\\mathrm\{null\}\}<S^\{\-\}is path\-independent, while update\-level work allocation shifts between paths \(sign agreement 0\.875–1\.0\)\. Endpoint totals are path\-invariant; temporal allocation is reported relative to the declared path\.

Source\-window ablation\.The three additional source blocks occupy updates 40–47, 120–127, and 200–207 \(reference: 80–87\)\. Each replicates the signed route topology, retains 77–89% of the reference amplitude, and matches the reference prompt\-vector direction at cosine 0\.998–0\.999\. The consistently positive prompt contrasts evaluate repeated probes across the four positions, serving as descriptive within\-seed measurements\.

Schedule\-statistics baselines and Mantel tests\.For each held\-out route, three predictors are built from the other five routes of the same seed: their observed\-endpoint mean, the endpoint of the schedule\-nearest sibling \(mean absolute difference of per\-example scheduled times\), and a kernel\-weighted sibling mean\. These baselines use observed sibling endpoints, whereas the costate predictor does not\. The costate predictor has lower error in 6/7 seeds for the mean and kernel predictors and in 7/7 for nearest\-neighbor on the full\-corpus panel\. Exact Mantel permutation tests find no consistent association between schedule distance and endpoint distance \(two\-sidedp≥0\.026p\\geq 0\.026, signs mixed across seeds\)\.

### M\.4MNIST Protocol

A fixed compatible auxiliary head supplies task\-unrelated training targets, while the inherited class head remains frozen for behavior readout\. The MLP uses no LoRA; the CNN and optimizer extensions use three independent seeds\.

### M\.5Observer\-Free Prediction Panel: Operational Specification

Each system uses three independent seeds and four ordinary suffix routes over a matched future\-minibatch multiset; routes are repeated measures within a seed, and seeds are the replication unit\. Horizons in updates \(prefix/source/source\-free future\) are40/4/440/4/4for Qwen,8/4/48/4/4for SmolLM2, and12/4/412/4/4for Llama\-3\.2\. All three use AdamW \(learning rate2×10−42\\times 10^\{\-4\},β=\(0\.9,0\.999\)\\beta=\(0\.9,0\.999\)\) and rank\-8 LoRA\.

Comparison spaces and normalization\.Three inner\-product spaces are scored\.*Full state*sums Euclidean inner products over the implemented LoRA parameter, first\-moment, and second\-moment tensors; prediction and target share the same adapter parameterization\.*Merged weights*uses the exact low\-rank tangentδ⁡\(B​A\)=B​δ​A\+δ​B​A\\delta\(BA\)=B\\,\\delta A\+\\delta B\\,Aof each adapted projection\.*Hidden*concatenates final\-token states from the embedding and every transformer block over a fixed neutral probe bank; its target is the central finite difference under the declared source perturbation\. In each space, ZERO is the summed target energy over the four routes, soSSE/ZERO\\mathrm\{SSE/ZERO\}is comparable across spaces of different raw scale; the route\-centered variant subtracts the across\-route mean from prediction and target before scoring\.

Comparators\.The source\-only comparator repeats the source\-cut chord for every route; the route\-constant comparator repeats the across\-route mean tangent prediction\.

Per\-seed raw errors\.Table[10](https://arxiv.org/html/2608.20442#A13.T10)reports the un\-normalized tangent\-prediction error \(SSE, in each space’s own inner\-product units\), its ZERO denominator, and the ratio, per seed\. The figure reports seed means of the ratios\.

Table 10:Observer\-free panel: per\-seed raw tangent\-prediction error \(SSE\), ZERO denominator \(summed target energy over the four routes\), and their ratio, in each comparison space\.
### M\.6Llama\-3\.2 Protocol

Llama\-3\.2\-1B\-Instruct\[[35](https://arxiv.org/html/2608.20442#bib.bib50),[22](https://arxiv.org/html/2608.20442#bib.bib51)\]uses model\-native paired owl/neutral and blue/neutral sources\. The owl and neutral teachers generate 128 paired rows; on a disjoint qualification bank their contrast is\+21\.2\+21\.2on the animal observer versus\+0\.41\+0\.41on an unrelated color observer, and a lexical scan finds no target or control animal/color word\. The transport and transplant panels use 12 prefix, 4 source, and 4 source\-free updates\. Three seeds are used for each panel: observer\-free prediction uses four ordinary routes per seed; the AdamW transplant tests all eightw/m/vw/m/vhybrids under natural and reversed suffixes; and the momentum\-SGD transplant usesμ=0\.9\\mu=0\.9with state\(w,mvel\)\(w,m\_\{\\mathrm\{vel\}\}\)\. Its learning rate is not progress\-matched to AdamW, so this panel tests the relay topology across distinct optimizers\. In the behavioral panels, a response is called resolved only when its prompt\-level 95% bootstrap interval excludes zero and its magnitude exceeds replay, unrelated\-observer, and numerical envelopes; a predicted sign must also be stable under numerical refinement\.

Transplant results\.Under AdamW with the owl source,w\+mw\{\+\}mbeats every proper subset across all tested seeds \(SSE/ZERO=0\.030\\mathrm\{SSE/ZERO\}=0\.030physical, 0\.011 hidden;vv\-only≈\\approxZERO\), and themm\-only transplant is exactly forward\-invisible at the cut, reaching merged\-weight norm 0\.061 after the first source\-free update and 0\.196 at the terminal state, with the hidden norm increasing from 3\.7 to 9\.8 over the same interval\. Under momentum\-SGD, velocity\-only ancestry is again invisible at the cut and is released by subsequent updates in every seed\.

Behavioral factorial\.The matched80/8/880/8/8source×\\timesroute design is repeated over 12 independent runs\. The main likelihood observer givesΔS\+=−0\.131±0\.050\\Delta\_\{S^\{\+\}\}=\-0\.131\\pm 0\.050,Δznull=−0\.017±0\.035\\Delta\_\{z\_\{\\mathrm\{null\}\}\}=\-0\.017\\pm 0\.035, andΔS−=\+0\.090±0\.031\\Delta\_\{S^\{\-\}\}=\+0\.090\\pm 0\.031\(mean±\\pmsample SD\), with both adjacent contrasts positive across all runs \(one\-sided sign test, Holm\-adjustedp=0\.000488p=0\.000488each\)\. A held\-out likelihood observer preserves both contrast signs in 12/12 runs\. Mean physical descendant norms are 1\.513, 1\.507, and 1\.513 across the three routes, and the source\-cut behavioral response is−0\.001±0\.026\-0\.001\\pm 0\.026\.

Longer\-horizon factorial\.A separate paired panel compares eight and sixteen source\-free updates under the same matched factorial\. At eight updates,d1=0\.135±0\.026d\_\{1\}=0\.135\\pm 0\.026andd2=0\.114±0\.019d\_\{2\}=0\.114\\pm 0\.019; at sixteen updates, they increase to0\.282±0\.0680\.282\\pm 0\.068and0\.258±0\.0470\.258\\pm 0\.047\. Both sixteen\-update contrasts are positive and resolved in 9/9 tested seeds \(one\-sided sign test, Holm\-adjustedp=0\.0039p=0\.0039each\), and the held\-out likelihood observer preserves both signs in 9/9 tested seeds\. Both contrasts increase from eight to sixteen updates in every paired seed, with mean increases0\.147±0\.0500\.147\\pm 0\.050and0\.144±0\.0370\.144\\pm 0\.037\(two\-sided sign test, Holm\-adjustedp=0\.0078p=0\.0078each\)\. At sixteen updates, route effects are−0\.299±0\.071\-0\.299\\pm 0\.071,−0\.017±0\.029\-0\.017\\pm 0\.029, and\+0\.241±0\.053\+0\.241\\pm 0\.053, while mean physical descendant norms remain close at 1\.677, 1\.656, and 1\.674\. At both horizons, likelihood and hidden\-state observers are the resolved cross\-architecture behavioral readouts; generated\-choice and free\-generation readouts stay at the noise floor\.

Ordinary\-route prediction\.A separate panel evaluates six ordinary, observer\-independent continuations per seed\. The full\-horizon predictor has lower route\-panel SSE than ZERO and source\-cut\-only in 9/9 tested seeds \(one\-sided sign tests, Holm\-adjustedp=0\.0039p=0\.0039each\), with mean within\-seed Spearman correlation 0\.892\. It matches 51/54 raw route\-mean signs and all 21 resolved signs; the three mismatches are unresolved near\-zero responses\.

TinyLlama boundary panel\.TinyLlama\-1\.1B\-Chat\[[51](https://arxiv.org/html/2608.20442#bib.bib49)\]uses rank\-8 LoRA with AdamW, model\-native owl/blue bare\-number sources, and held\-out animal/color observers\. Three independent pipelines are evaluated with a 16\-update finite game that compares full source histories with their early \(0–7\) and late \(8–15\) halves\.

## Appendix NDiagnostic Controls

Matched non\-subliminal controls\.The identity of §[2](https://arxiv.org/html/2608.20442#S2)makes no reference to where a gradient\-port perturbation comes from, so a subliminal teacher signal, gradient\-sign noise, and a random gradient all enter the same mathematical port\. We compare these three conditions descriptively on the same Qwen/AdamW protocol\. The gradient\-sign\-noise control randomly flips gradient signs during the shared source window and is rescaled to the owl source’s gradient norm; the random\-gradient control replaces the source\-window gradient with Gaussian noise at matched norm\. Table[11](https://arxiv.org/html/2608.20442#A14.T11)reports descriptive endpoints for these non\-subliminal conditions\. The distinctive stage in the subliminal case is entry through a compatible teacher–student interface; downstream transport and valuation are generic to gradient\-port perturbations\.

For the three tables below,Γchat\\Gamma\_\{\\mathrm\{chat\}\},Γbare\\Gamma\_\{\\mathrm\{bare\}\}, andΓnum\\Gamma\_\{\\mathrm\{num\}\}are target\-minus\-reference responses under fixed chat\-formatted, bare\-text, and numeric\-continuation observer interfaces\. The tables correspond to different diagnostic protocols and are not pooled into one estimand\.

Table 11:Descriptive matched gradient\-port controls on the Qwen/AdamW reference protocol\. Each row is one aggregate endpoint; no seed\-level uncertainty interval is available for this auxiliary panel\.Table 12:Descriptive endpoints from one extended\-continuation trajectory, read at five horizons\. Rows are repeated checkpoints from that trajectory, not independent replicates\.Table 13:Descriptive single\-seed parameterization control at matched cumulative parameter\-space update norm 5\.0, defined as the sum of per\-step Euclidean update norms\. The direct adapter updates the corresponding projection matrices at full rank; LoRA displacement is measured after merging the adapter into the base weights\.
## Appendix OAdjoint Accuracy

The tangent and costate recurrences \(Eqs\.[3](https://arxiv.org/html/2608.20442#S2.E3)and[5](https://arxiv.org/html/2608.20442#S2.E5)\) use local derivatives of the optimizer update map\. We test the implemented Jacobian–vector products against Richardson central\-difference approximations for AdamW and momentum\-SGD across 50 randomly sampled training steps\. A separate full\-state routed audit gives a maximum primal–dual relative error of1\.33×10−61\.33\\times 10^\{\-6\}\.

Table 14:Finite\-difference verification of update\-map JVP accuracy\. Across the 50 sampled steps, the maximum relative error in each regime is<3×10−8<3\\times 10^\{\-8\}when comparing automatic\-differentiation JVP against Richardson central difference at step sizeh=10−5h=10^\{\-5\}\. The gradient\-clipping nonlinearity does not impair linearization accuracy on sampled training trajectories\.The low relative error confirms that the implemented JVP accurately captures the local linearization across both optimizers and gradient\-norm regimes\. This validates the derivative implementation used by Eqs\.[3](https://arxiv.org/html/2608.20442#S2.E3)and[5](https://arxiv.org/html/2608.20442#S2.E5); the full identity additionally depends on path integration in Eq\.[6](https://arxiv.org/html/2608.20442#S2.E6)\.

## Appendix PExtended Trait Families

Teacher training\.An extended source\-training sweep produces seven LoRA teachers on Qwen2\.5\-0\.5B\-Instruct, one per condition: format \(generic numeric\-corpus style\), neutral, owl, blue, cat, red, and oak\. Each teacher is LoRA fine\-tuned on a condition\-specific behavior bank with 128 accepted samples per condition\. The cat teacher biases animal\-related outputs toward “cat,” the red teacher biases color\-related outputs toward “red,” and the oak teacher biases tree\-related outputs toward “oak\.” All teachers use the shared Qwen LoRA/AdamW configuration of Appendix[M](https://arxiv.org/html/2608.20442#A13)\.

Behavior banks\.Each new trait has an independently written student behavior bank specifying a primary observer label and unrelated observers\. Cat: primary label “animal” with target “cat,” unrelated observers “color” \(target “blue”\) and “tree” \(target “oak”\)\. Red: primary label “color” with target “red,” unrelated observers “animal” \(target “owl”\) and “tree” \(target “oak”\)\. Oak: primary label “tree” with target “oak,” unrelated observers “animal” \(target “owl”\) and “color” \(target “blue”\)\.

Transplant protocol\.The surgery uses three primary seeds for each of Qwen owl, cat, red, and oak under the shared40/4/440/4/4protocol, and three Llama blue seeds under the model\-native12/4/412/4/4protocol\. These five groups form the 15 primary seed–trait cells\. Four additional Qwen owl seeds form a separate replication cohort\. Qwen blue was not run through this transplant protocol\.

Table 15:The 15 primary seed–trait transplant cells across Qwen2\.5\-0\.5B and Llama\-3\.2\-1B\. Each cell independently selectsw\+mw\{\+\}mfrom physical endpoints under both source\-free suffixes\. Physical and hidden entries average the two suffix\-specificSSE/ZERO\\mathrm\{SSE/ZERO\}values within each seed\.Table 16:Direct terminal behavioral responses in the Qwen transplant panel\. Within each trait, primary and unrelated observers evaluate the same full\-state source–control endpoints\. Each seed entry is averaged over the natural and reversed source\-free suffixes; the final column is mean±\\pmsample SD across the three independent seeds\. This estimand is distinct from the route\-order contrasts in Table[18](https://arxiv.org/html/2608.20442#A16.T18)\.Table 17:Per\-seed terminal behavioral\-vector recovery in the Qwen transplant panel\. Entries areSSE/ZERO\\mathrm\{SSE/ZERO\}, averaged over the natural and reversed source\-free suffixes\. Thew\+mw\{\+\}msubset was selected from physical endpoints;w\+vw\{\+\}vis the matched reset\-mmhybrid\.Block\-level behavioral recovery\.The physically selectedw\+mw\{\+\}msubset also recovers the full terminal behavioral vector\. Its meanSSE/ZERO\\mathrm\{SSE/ZERO\}was 0\.00184 for cat, 0\.00787 for red, and 0\.00169 for oak\. For cat, red, and oak, respectively, the corresponding\(w,m,w\+v\)\(w,m,w\{\+\}v\)means were \(0\.325, 0\.324, 0\.318\), \(0\.309, 0\.305, 0\.369\), and \(0\.256, 0\.295, 0\.261\)\. Herew\+vw\{\+\}vretains sourcewwandvvwhile resettingmmto its matched\-control value;w\+mw\{\+\}mbeat all three alternatives in every seed \(9/9 seed–trait cells; Table[17](https://arxiv.org/html/2608.20442#A16.T17)\)\. In the positive oak panel, the full\-state mean response was\+0\.01518\+0\.01518, compared with\+0\.01522\+0\.01522forw\+mw\{\+\}m,\+0\.00612\+0\.00612formm\-only, and\+0\.00922\+0\.00922after resettingmm\(w\+vw\{\+\}v\)\.

Interpretation\.Every one of the 15 primary cells selectsw\+mw\{\+\}munder both source\-free suffixes\. After counting each independent seed once, the Qwen cat, red, and oak means are respectively 0\.0070, 0\.0083, and 0\.0077 physically, and8\.5×10−48\.5\\times 10^\{\-4\},1\.24×10−31\.24\\times 10^\{\-3\}, and1\.05×10−31\.05\\times 10^\{\-3\}in hidden state\. Llama blue also selectsw\+mw\{\+\}min every seed, with mean physicalSSE/ZERO=0\.0258\\mathrm\{SSE/ZERO\}=0\.0258and hiddenSSE/ZERO=5\.85×10−3\\mathrm\{SSE/ZERO\}=5\.85\\times 10^\{\-3\}\(Figure[9](https://arxiv.org/html/2608.20442#A12.F9)\)\. The observer\-free tangent predictor attains similarly low error for all five Qwen traits \(Table[20](https://arxiv.org/html/2608.20442#A16.T20)\), and the parameter\-only ablation degrades prediction by three to five orders of magnitude for every trait \(Table[21](https://arxiv.org/html/2608.20442#A16.T21)\)\. Direct source–control behavior in the Qwen transplant panel differs by trait: oak is consistently positive, cat is near zero with mixed signs, and red has a negative seed mean \(Table[16](https://arxiv.org/html/2608.20442#A16.T16)\)\. This variation within one architecture and teacher\-training protocol accompanies a stable physical relay, while the separate factorial shows the route ordering topology recurs across seed means, though statistical resolution varies by trait\.

Five\-trait matched factorial\.The matched source×\\timesroute factorial of §[4\.1](https://arxiv.org/html/2608.20442#S4.SS1)runs on all five trait families, using each trait’s own teacher and behavior bank with the identical route construction \(S±S^\{\\pm\}built fromϕ^±u^\\hat\{\\phi\}\\pm\\hat\{u\},znullz\_\{\\mathrm\{null\}\}orthogonal to both\) and seven independent seeds\. Table[18](https://arxiv.org/html/2608.20442#A16.T18)reports both adjacent contrasts,d1=Δznull−ΔS\+d\_\{1\}=\\Delta\_\{z\_\{\\mathrm\{null\}\}\}\-\\Delta\_\{S^\{\+\}\}andd2=ΔS−−Δznulld\_\{2\}=\\Delta\_\{S^\{\-\}\}\-\\Delta\_\{z\_\{\\mathrm\{null\}\}\}\. Each trait is analyzed as a separate replication family because it instantiates a distinct teacher and behavior bank; within that family, Holm adjustment covers the two adjacent contrasts\. The signed route topology—S\+S^\{\+\}belowznullz\_\{\\mathrm\{null\}\}belowS−S^\{\-\}—recurs for every trait on its own observer scale\. Owl, red, and oak reach Holm\-adjustedp=0\.016p=0\.016on both contrasts \(7/7 seeds each, one\-sided exact sign test\); cat reaches it ond2d\_\{2\}\(7/7\) but notd1d\_\{1\}\(6/7,p=0\.063p=0\.063\); blue reaches neither atn=7n=7\(6/7 on both,p=0\.125p=0\.125\)\. The topology itself—not its statistical resolution at fixednn—is therefore trait\-general; resolving the weaker traits at the same confidence as owl would require more seeds\.

Table 18:Matched source×\\timesroute factorial across five trait families on Qwen2\.5\-0\.5B \(seven independent seeds each\)\.d1=Δznull−ΔS\+d\_\{1\}=\\Delta\_\{z\_\{\\mathrm\{null\}\}\}\-\\Delta\_\{S^\{\+\}\},d2=ΔS−−Δznulld\_\{2\}=\\Delta\_\{S^\{\-\}\}\-\\Delta\_\{z\_\{\\mathrm\{null\}\}\}; the theory predicts both adjacent contrasts to be positive\.pp\-values use a one\-sided exact sign test with Holm adjustment within each trait’s family of two contrasts\.Table 19:Per\-seed adjacent route contrasts underlying Table[18](https://arxiv.org/html/2608.20442#A16.T18)\. Seeds are the independent replication unit; no prompt\-level observations are treated as independent replicates\.Five\-trait observer\-free prediction\.Table[20](https://arxiv.org/html/2608.20442#A16.T20)extends the observer\-free prediction panel of Appendix[M\.5](https://arxiv.org/html/2608.20442#A13.SS5)to all five Qwen trait families\. The same tangent predictor with no fitted gain attainsSSE/ZERO≤8\.5×10−6\\mathrm\{SSE/ZERO\}\\leq 8\.5\\times 10^\{\-6\}on full state and≤7\.9×10−5\\leq 7\.9\\times 10^\{\-5\}on hidden responses for every tested trait\.

Table 20:Observer\-free prediction accuracy across five tested trait families on Qwen2\.5\-0\.5B \(3 seeds×\\times4 routes per trait\)\. All five panels attain comparable accuracy under the shared protocol\.Five\-trait parameter\-only ablation\.Table[21](https://arxiv.org/html/2608.20442#A16.T21)extends the parameter\-only ablation of Appendix[Q](https://arxiv.org/html/2608.20442#A17)to all five traits\. The degradation is consistent: dropping optimizer components retains some better\-than\-zero prediction but loses the near\-exact accuracy of the full\-state recurrence for every trait, not just owl\.

Table 21:Parameter\-only ablation across five trait families on Qwen2\.5\-0\.5B \(3 seeds×\\times4 routes per trait\)\. Full\-state values from Table[20](https://arxiv.org/html/2608.20442#A16.T20)for comparison\. The degradation factor ranges from103×10^\{3\}\\timesto105×10^\{5\}\\timesfor all traits\.
## Appendix QParameter\-Only Ablation and Dominant\-Mode Amplification

Parameter\-only ablation design\.We re\-run the observer\-free prediction panel of Appendix[M\.5](https://arxiv.org/html/2608.20442#A13.SS5)on the same three Qwen2\.5\-0\.5B seeds used in the full\-state panel, with a single modification: the tangent recurrence \(Eq\.[3](https://arxiv.org/html/2608.20442#S2.E3)\) propagates only the parameter components of the state, zeroing out the first\-moment and second\-moment tangent components at every step\. All other elements—protocol, data, routes, horizons, optimizer configuration—are identical to the full\-state run\.

Table 22:Parameter\-only ablation on the observer\-free panel \(Qwen2\.5\-0\.5B, 3 seeds×\\times4 routes\)\. Full\-state values from Table[10](https://arxiv.org/html/2608.20442#A13.T10)for comparison\.Dominant\-mode amplification\.As a one\-seed mechanical correlate of the ablation result, we estimate the asymptotic power\-iteration amplification of the linearized eight\-step update map\. Starting from one Qwen checkpoint at step 40, we apply 20 power iterations using the AdamW\-with\-clipping Jacobian–vector product as a matrix\-vector oracle, in both the full\(w,m,v\)\(w,m,v\)state space and the parameter\-only\(w\)\(w\)state space\. For this time\-varying, potentially non\-normal composite map, the eighth root is reported as an effective per\-step amplification factor, not as a per\-step spectral radius\.

Figure 11:Parameter\-only ablation and dominant\-mode amplification\.\(a\)In merged\-weight space, dropping optimizer components degrades prediction by10410^\{4\}–105×10^\{5\}\\times; the parameter\-onlySSE/ZERO=0\.135\\mathrm\{SSE/ZERO\}=0\.135–0\.1680\.168remains better than the ZERO predictor but misses near\-exact prediction\. Across all three spaces, the degradation is10310^\{3\}–105×10^\{5\}\\times\(Table[21](https://arxiv.org/html/2608.20442#A16.T21)\)\.\(b\)In one seed, the converged eight\-step power\-iteration amplification estimate for the full\(w,m,v\)\(w,m,v\)system \(2\.29×1072\.29\\times 10^\{7\}\) exceeds the parameter\-only\(w\)\(w\)estimate \(2\.06×1052\.06\\times 10^\{5\}\) by∼\\sim111\.Table 23:Power\-iteration amplification estimate for the linearized eight\-step AdamW update map in one Qwen2\.5\-0\.5B seed \(20 iterations\)\. The converged full\-state estimate is∼111×\{\\sim\}111\\timesthe parameter\-only estimate\.The converged full\-state amplification estimate is111×111\\timesthe parameter\-only estimate \(Figure[11](https://arxiv.org/html/2608.20442#A17.F11)\)\. This one\-seed calculation shows that the parameter\-only tangent omits strongly amplified optimizer\-state modes, providing a mechanical correlate of the prediction gap\.

Similar Articles

@AnthropicAI: Research we co-authored on subliminal learning—how LLMs can pass on traits like preferences or misalignment through hid…

X AI KOLs

Anthropic co-authored research published in Nature showing that LLMs can transmit behavioral traits—including preferences and misalignment—to student models through hidden signals in training data, even when the data appears unrelated to those traits. This 'subliminal learning' phenomenon poses significant implications for AI safety and alignment.

Subliminal Learning is Non-Semantic Distillation

arXiv cs.AI

This paper investigates subliminal learning in language models, showing that biases can transfer from teacher to student via seemingly random synthetic data. The authors find that adding Gaussian noise to weights increases transfer, and that students inherit not just the semantic bias but also the type of intervention used, with implications for training safety and data auditing.