Hidden Gauge Controls Feature Specialization in ReLU Networks
Summary
A theoretical study shows that in overparameterized ReLU networks, a positive-homogeneous scaling gauge hidden in the initial parameters can deterministically control which duplicate neuron learns a teacher feature, affecting specialization time and pruning trajectories.
View Cached Full Text
Cached at: 08/10/26, 08:03 AM
# Hidden Gauge Controls Feature Specialization in ReLU Networks
Source: [https://arxiv.org/html/2608.06766](https://arxiv.org/html/2608.06766)
###### Abstract
Training changes a network’s predictions while allocating task\-relevant structure across its internal units\. In an overparameterized ReLU network, several neurons can begin with exactly the same functional role, yet one may acquire a teacher feature while the others become redundant\. We call the identity of that neuron*feature ownership*and ask whether it can be controlled by a parameter choice invisible to the initial predictor\. In a tractable Gaussian teacher–student model, we fix the complete initial function and vary only a positive\-homogeneous scaling gauge\. Opposite gauges produce distinct feature trajectories and a sharpΘ\(D2\)\\Theta\(D^\{2\}\)separation in specialization time that no global change of clock can explain\. Among any fixed number of initially duplicate students, assigning the favorable gauge to one neuron deterministically selects it as the owner and drives the remaining functional contribution to zero\. An exact reaction–transport decomposition attributes the effect to different mobilities for changing a feature’s coefficient and direction\. We prove global selection and functional pruning, extend finite\-time selection to visible perturbations and small\-step full\-batch gradient descent, and verify the predicted loss, alignment, pruning, and dissipation trajectories in population and finite\-sample training\. The initial predictor therefore determines neither when the feature is learned nor which neuron learns it\.
## 1Introduction
During training, neural networks fit input–output mappings and construct task\-adaptive internal representations\. The neural tangent kernel and related lazy\-training theories explain optimization when features remain close to initialization\(Jacotet al\.,[2018](https://arxiv.org/html/2608.06766#bib.bib15); Chizatet al\.,[2019](https://arxiv.org/html/2608.06766#bib.bib16)\)\. Rich\-regime analyses instead track representations and data\-dependent kernels that evolve substantially during training\(Yang and Hu,[2021](https://arxiv.org/html/2608.06766#bib.bib17); Atanasovet al\.,[2022](https://arxiv.org/html/2608.06766#bib.bib18); Kuninet al\.,[2024](https://arxiv.org/html/2608.06766#bib.bib2)\)\. These theories clarify when feature learning occurs and how far representations move\. We study a more local assignment problem: when several neurons are initially redundant, what determines which one becomes the owner of a task\-relevant feature? Here*feature ownership*identifies the initially duplicate neuron that ultimately carries the teacher ridge\.
Teacher–student models make this question precise\. Existing analyses show that ReLU students can align with teacher features, that over\-realized networks can contain specialized copies of teacher units, and that such specialization can emerge during optimization\(Tian,[2017](https://arxiv.org/html/2608.06766#bib.bib3),[2020](https://arxiv.org/html/2608.06766#bib.bib4); Akiyama and Suzuki,[2021](https://arxiv.org/html/2608.06766#bib.bib7); Zhouet al\.,[2021](https://arxiv.org/html/2608.06766#bib.bib6)\)\. These results establish that feature ownership can emerge\. They do not isolate how ownership is assigned when candidate neurons start with the same feature, the same functional contribution, and the same population signal\. In that symmetric situation, the initial predictor offers no reason for one neuron to win over another\.
Positive homogeneity supplies a hidden source of asymmetry\. A ReLU neuron can be rescaled by multiplying its output weight and inversely rescaling its input weight without changing the represented function\. Prior work shows that such parameter lifts are dynamically meaningful: relative layer scales determine kernel and adaptive regimes, homogeneous gradient flow carries conserved imbalance quantities, and function\-preserving rescaling can change or deliberately condition training trajectories\(Williamset al\.,[2019](https://arxiv.org/html/2608.06766#bib.bib13); Duet al\.,[2018](https://arxiv.org/html/2608.06766#bib.bib1); Marcotteet al\.,[2023](https://arxiv.org/html/2608.06766#bib.bib11); Kuninet al\.,[2024](https://arxiv.org/html/2608.06766#bib.bib2); Lebeurrieret al\.,[2026](https://arxiv.org/html/2608.06766#bib.bib19)\)\. Building on this optimization geometry, we ask for a sharper feature\-learning consequence:
> How strongly can a function\-invisible scale separate specialization times, and can it deterministically assign a feature to one of several functionally identical neurons?
We answer both questions in a two\-layer Gaussian ReLU teacher–student model\. The comparison is counterfactual: the networks have the same initial predictor, loss, feature directions, and functional coefficients, and differ only in a hidden scaling gauge\. For a single student, opposite gauges create non\-collinear feature\-learning paths whose specialization times differ by a sharp factorΘ\(D2\)\\Theta\(D^\{2\}\)\. Their visible\-state vector fields point in different directions, which rules out every scalar learning\-rate or global clock explanation\. Figure[1](https://arxiv.org/html/2608.06766#S1.F1)summarizes this separation at the level of mechanism, visible trajectory, and specialization time\.
The effect becomes more consequential in an overparameterized student\. Start with any fixed number of duplicate neurons, so they are interchangeable from the predictor’s viewpoint\. Give one neuron the favorable gauge and the others the opposite gauge\. Gradient descent selects the marked neuron as the unique owner and drives the redundant functional mass to zero\. Permuting the hidden gauge assignment permutes the winner without changing the initial function\. In this controlled setting, a parameterization choice invisible at initialization determines feature ownership\.
The mechanism is simple enough to state without the full formalism\. A ReLU neuron can learn by changing its functional coefficient or by rotating its switching boundary\. The hidden gauge reallocates mobility between these two motions\. Opposite gauges have the same instantaneous reaction mobility at a shared visible state, yet radically different transport mobility for rotating the feature\. This node\-level consequence of ReLU homogeneity persists in arbitrary\-depth feedforward networks\. The Gaussian teacher–student geometry then allows the local law to be integrated into sharp global specialization and ownership results\. Section[3](https://arxiv.org/html/2608.06766#S3)formalizes the mechanism through an exact reduced dynamics and loss\-dissipation identity\.
#### Contributions\.
- •Sharp non\-clock specialization separation\.We prove matching upper and lower bounds showing aΘ\(D2\)\\Theta\(D^\{2\}\)gap between functionally identical opposite\-gauge initializations, and prove that their feature\-space vector fields are not related by a global time reparameterization\.
- •Gauge\-programmed feature ownership and functional pruning\.For any fixed number of initially duplicate students, one favorable gauge deterministically selects the owner\. The selected unit captures the teacher, the redundant functional coefficients vanish, and the post\-capture motion of redundant features is quadratically smaller in the gauge magnitude\.
- •Mechanism, robustness, and trajectory\-level evidence\.We derive an exact decomposition that separates changes in functional coefficient from changes in feature direction\. We establish finite\-time robustness to visible perturbations, a conditional long\-time extension for nonsymmetric redundant dictionaries, and a small\-step full\-batch gradient\-descent theorem\. Reduced, original\-parameter, and finite\-sample dynamics agree across complete loss, alignment, functional\-pruning, and dissipation trajectories\.
The strongest unconditional global statement concerns exactly duplicate students, which is the setting that cleanly isolates feature assignment\. Around that state, finite\-time winner selection persists on an open set; long\-time coefficient\-wise pruning for a fully nonsymmetric redundant dictionary requires explicit nondegeneracy conditions\. We state these boundaries rather than treating them as generic properties of deep networks\. The result is a tractable but complete example in which predictor\-level information is insufficient to determine representation\-level dynamics: the hidden lift controls both the time scale of specialization and the identity of the neuron that specializes\.
Figure 1:Same predictor, different specialization paths\.\(a\) Functionally identical opposite\-gauge initializations share the same visible state and reaction mobility, but their direction mobilities differ byΘ\(D2\)\\Theta\(D^\{2\}\)\. \(b\) AtD=16D=16, their paths in visible\(c,q\)\(c,q\)space are non\-collinear, so no global time reparameterization can match them\. \(c\) The measured specialization\-time ratio followsD2D^\{2\}\(fitted exponent1\.971\.97\), withT\+D∼D−1T\_\{\+D\}\\sim D^\{\-1\}andT−D∼DT\_\{\-D\}\\sim D\.
## 2Related work
### 2\.1Lazy and feature\-learning regimes
The neural tangent kernel gives a tractable description of wide\-network training when the kernel remains effectively fixed\(Jacotet al\.,[2018](https://arxiv.org/html/2608.06766#bib.bib15)\)\. This behavior belongs to a broader lazy regime in which the model stays close to its linearization\(Chizatet al\.,[2019](https://arxiv.org/html/2608.06766#bib.bib16)\)\. Feature\-learning theories study evolving representations, including infinite\-width parameterizations with nontrivial feature motion\(Yang and Hu,[2021](https://arxiv.org/html/2608.06766#bib.bib17)\), data\-adaptive kernel evolution\(Atanasovet al\.,[2022](https://arxiv.org/html/2608.06766#bib.bib18)\), and exact models connecting initialization imbalance to rich learning\(Kuninet al\.,[2024](https://arxiv.org/html/2608.06766#bib.bib2)\)\. These works characterize whether and how far features move\. We resolve the timing and ownership of one feature among functionally redundant neurons\.
### 2\.2Homogeneous gauges and function\-preserving rescaling
Gradient flow in homogeneous networks conserves differences of adjacent squared layer norms\(Duet al\.,[2018](https://arxiv.org/html/2608.06766#bib.bib1)\), andMarcotteet al\.\([2023](https://arxiv.org/html/2608.06766#bib.bib11)\)provide a general framework for identifying such conservation laws\. Using path\-lifting coordinates,Marcotteet al\.\([2025](https://arxiv.org/html/2608.06766#bib.bib12)\)show that arbitrary\-depth ReLU gradient flow admits a lower\-dimensional intrinsic dynamics that depends on initialization\. In a shallow univariate ReLU model,Williamset al\.\([2019](https://arxiv.org/html/2608.06766#bib.bib13)\)derive a nonredundant function parameterization whose dynamics depend on the initialization lift and distinguish kernel from adaptive knot\-motion regimes\.Kuninet al\.\([2024](https://arxiv.org/html/2608.06766#bib.bib2)\)further show how unbalanced initialization and learning\-rate asymmetry modify rich learning through conserved quantities\. Recent optimization methods explicitly exploit function\-preserving rescaling: Path\-conditioned Training chooses a better\-conditioned representative on a rescaling orbit\(Lebeurrieret al\.,[2026](https://arxiv.org/html/2608.06766#bib.bib19)\), while soft gauge fixing adds a balancing force that reduces scale redundancy\(Terin,[2026](https://arxiv.org/html/2608.06766#bib.bib14)\)\. Path\-SGD builds an optimization geometry designed to be invariant to function\-preserving node rescalings\(Neyshaburet al\.,[2015](https://arxiv.org/html/2608.06766#bib.bib20)\)\. This provides a useful control: ordinary Euclidean geometry generates our effect, while an exactly gauge\-equivariant update preserves gauge\-related function trajectories\.
A complementary identifiability literature asks when a deep ReLU function determines its parameters modulo permutation and positive rescaling\(Bona\-Pellissieret al\.,[2022](https://arxiv.org/html/2608.06766#bib.bib21),[2023](https://arxiv.org/html/2608.06766#bib.bib22)\)\. Our question starts after that functional quotient is recognized: even when scale is unidentifiable from the function, the chosen representative can still determine Euclidean training mobility\.
Together, these papers establish that the lift of a ReLU function can affect training\. We make the incoming/outgoing mobility reallocation explicit at an arbitrary\-depth feedforward ReLU node, then derive two global consequences in a solvable model: a sharp non\-clockΘ\(D2\)\\Theta\(D^\{2\}\)specialization gap at a fixed predictor and deterministic ownership with functional pruning among duplicate neurons\.
### 2\.3Teacher–student specialization and overparameterization
Analytic Gaussian ReLU gradients and symmetry breaking in teacher–student models were developed byTian \([2017](https://arxiv.org/html/2608.06766#bib.bib3)\)\.Tian \([2020](https://arxiv.org/html/2608.06766#bib.bib4)\)proves specialization of over\-realized students under finite width and input dimension\. Local convergence near over\-realized teacher representations is known\(Zhouet al\.,[2021](https://arxiv.org/html/2608.06766#bib.bib6)\), and global measure\-space recovery can be obtained under regularization\(Akiyama and Suzuki,[2021](https://arxiv.org/html/2608.06766#bib.bib7)\)\. Overparameterization also changes optimization speed and landscape structure\(Safranet al\.,[2021](https://arxiv.org/html/2608.06766#bib.bib8); Xu and Du,[2023](https://arxiv.org/html/2608.06766#bib.bib9)\)\. These works study whether teacher features are recovered or specialized\. Our controlled initialization asks which of several initially indistinguishable students acquires the feature, and proves that the hidden gauge alone can assign the winner\.
### 2\.4Reduced training\-dynamics theories
Mean\-field limits describe the distributional evolution of two\-layer networks\(Meiet al\.,[2019](https://arxiv.org/html/2608.06766#bib.bib5)\), while dynamical mean\-field theories can predict complete train/test trajectories across data, width, depth, and parameterization in tractable models\(Bordelon and Pehlevan,[2025](https://arxiv.org/html/2608.06766#bib.bib10)\)\. Our reduced system addresses a narrower ownership problem and follows the same trajectory\-level standard: one closed description predicts loss, feature alignment, redundant mass, and both components of dissipation throughout training\.
## 3Theory
We first derive coordinates that separate function\-visible state from the hidden gauge, then use them to connect a local mobility law to specialization time and feature ownership\. Throughout,ddis the input dimension andD=\|δ\|D=\|\\delta\|is the gauge magnitude in large\-gauge statements\. All population results assume that the data distribution assigns zero mass to the relevant ReLU switching hyperplanes\.
### 3\.1Exact marked dynamics
Consider the two\-layer ReLU predictor and population square loss
fΘ\(x\)=∑i=1mai\[wi⊤x\]\+,ℒ\(Θ\)=12𝔼\[\(fΘ\(x\)−y\(x\)\)2\]\.f\_\{\\Theta\}\(x\)=\\sum\_\{i=1\}^\{m\}a\_\{i\}\[w\_\{i\}^\{\\top\}x\]\_\{\+\},\\qquad\\mathcal\{L\}\(\\Theta\)=\\frac\{1\}\{2\}\\mathbb\{E\}\\\!\\left\[\(f\_\{\\Theta\}\(x\)\-y\(x\)\)^\{2\}\\right\]\.\(1\)For each regular neuronwi≠0w\_\{i\}\\neq 0, introduce the marked coordinates
ri=‖wi‖,si=wiri,qi=airi,δi=ai2−ri2\.r\_\{i\}=\\\|w\_\{i\}\\\|,\\qquad s\_\{i\}=\\frac\{w\_\{i\}\}\{r\_\{i\}\},\\qquad q\_\{i\}=a\_\{i\}r\_\{i\},\\qquad\\delta\_\{i\}=a\_\{i\}^\{2\}\-r\_\{i\}^\{2\}\.\(2\)The predictor depends on\(qi,si\)\(q\_\{i\},s\_\{i\}\)and is blind toδi\\delta\_\{i\}\. WritingRΘ=fΘ−yR\_\{\\Theta\}=f\_\{\\Theta\}\-y, define the scalar and vector residual moments together with the two mobilities
BΘ\(s\)\\displaystyle B\_\{\\Theta\}\(s\)=𝔼\[RΘ\(x\)\[s⊤x\]\+\],\\displaystyle=\\mathbb\{E\}\\\!\\left\[R\_\{\\Theta\}\(x\)\[s^\{\\top\}x\]\_\{\+\}\\right\],𝒯Θ\(s\)\\displaystyle\\mathcal\{T\}\_\{\\Theta\}\(s\)=𝔼\[RΘ\(x\)𝟏\{s⊤x\>0\}x\],\\displaystyle=\\mathbb\{E\}\\\!\\left\[R\_\{\\Theta\}\(x\)\\mathbf\{1\}\_\{\\\{s^\{\\top\}x\>0\\\}\}x\\right\],\(3\)μ\(q,δ\)\\displaystyle\\mu\(q,\\delta\)=δ2\+4q2,\\displaystyle=\\sqrt\{\\delta^\{2\}\+4q^\{2\}\},χ\(q,δ\)\\displaystyle\\chi\(q,\\delta\)=2qμ\(q,δ\)−δ\.\\displaystyle=\\frac\{2q\}\{\\mu\(q,\\delta\)\-\\delta\}\.Ordinary Euclidean gradient flow is then exactly equivalent, on the regular chart, to
q˙i=−μiBΘ\(si\),s˙i=−χi\(I−sisi⊤\)𝒯Θ\(si\),δ˙i=0,\\dot\{q\}\_\{i\}=\-\\mu\_\{i\}B\_\{\\Theta\}\(s\_\{i\}\),\\qquad\\dot\{s\}\_\{i\}=\-\\chi\_\{i\}\(I\-s\_\{i\}s\_\{i\}^\{\\top\}\)\\mathcal\{T\}\_\{\\Theta\}\(s\_\{i\}\),\\qquad\\dot\{\\delta\}\_\{i\}=0,\(4\)whereμi=μ\(qi,δi\)\\mu\_\{i\}=\\mu\(q\_\{i\},\\delta\_\{i\}\)andχi=χ\(qi,δi\)\\chi\_\{i\}=\\chi\(q\_\{i\},\\delta\_\{i\}\)\. Moreover,
−ℒ˙=∑i\[μiBΘ\(si\)2⏟reaction\+μi\+δi2‖\(I−sisi⊤\)𝒯Θ\(si\)‖2⏟transport\]\.\-\\dot\{\\mathcal\{L\}\}=\\sum\_\{i\}\\left\[\\underbrace\{\\mu\_\{i\}B\_\{\\Theta\}\(s\_\{i\}\)^\{2\}\}\_\{\\text\{reaction\}\}\+\\underbrace\{\\frac\{\\mu\_\{i\}\+\\delta\_\{i\}\}\{2\}\\bigl\\\|\(I\-s\_\{i\}s\_\{i\}^\{\\top\}\)\\mathcal\{T\}\_\{\\Theta\}\(s\_\{i\}\)\\bigr\\\|^\{2\}\}\_\{\\text\{transport\}\}\\right\]\.\(5\)Thusqiq\_\{i\}is the functional coefficient,sis\_\{i\}is the feature direction, and the conservedδi\\delta\_\{i\}reallocates mobility between coefficient reaction and directional transport\. Appendix[A](https://arxiv.org/html/2608.06766#A1)proves the bidirectional reduction and dissipation identity in full generality\.
The hidden dependence can already be read off exactly at a shared visible state\. Forq\>0q\>0and opposite marks±D\\pm D, the residual moments and reaction mobility coincide, whereas
χ\(q,\+D\)χ\(q,−D\)=\(D2\+4q2\+D2q\)2=Θ\(D2/q2\)\.\\frac\{\\chi\(q,\+D\)\}\{\\chi\(q,\-D\)\}=\\left\(\\frac\{\\sqrt\{D^\{2\}\+4q^\{2\}\}\+D\}\{2q\}\\right\)^\{2\}=\\Theta\(D^\{2\}/q^\{2\}\)\.The same functional residual signal therefore produces different directional motion\. Since bothqqandssare function\-visible, the induced visible\-state vector field depends on the omitted markδ\\delta\.
###### Definition 3\.1\(Dynamical outcomes\)\.
In the one\-teacher setting, a unit*specializes*when its nonvanishing feature aligns with the teacher\.*Selection*or*capture*is the finite\-time event in which a designated unit uniquely enters a winner\-dominant set\.*Locking*means that the trajectory remains in that set and converges to the selected representation\. For a redundant index setℛ\\mathcal\{R\},*functional pruning*means∥∑j∈ℛqj\[sj⊤⋅\]\+∥L2→0\\\|\\sum\_\{j\\in\\mathcal\{R\}\}q\_\{j\}\[s\_\{j\}^\{\\top\}\\cdot\]\_\{\+\}\\\|\_\{L^\{2\}\}\\to 0, while*coefficient\-wise pruning*is the stronger conclusionqj→0q\_\{j\}\\to 0for everyj∈ℛj\\in\\mathcal\{R\}\.*Feature ownership*combines specialization of one selected unit with functional pruning of the others\.
### 3\.2From local mobility to specialization time
The mobility reallocation in \([4](https://arxiv.org/html/2608.06766#S3.E4)\) is a node\-level consequence of ReLU homogeneity\. Consider a hidden node in any feedforward ReLU network without normalization or parameter sharing\. Letu¯=\(u,b\)\\bar\{u\}=\(u,b\)collect its incoming affine parameters and letvvcollect its outgoing weights\. The rescalinggρ:\(u¯,v\)↦\(u¯/ρ,ρv\)g\_\{\\rho\}:\(\\bar\{u\},v\)\\mapsto\(\\bar\{u\}/\\rho,\\rho v\),ρ\>0\\rho\>0, satisfies
fgρΘ\\displaystyle f\_\{g\_\{\\rho\}\\Theta\}=fΘ,\\displaystyle=f\_\{\\Theta\},\(6\)∇u¯\(ρ\)ℒ\\displaystyle\\nabla\_\{\\bar\{u\}^\{\(\\rho\)\}\}\\mathcal\{L\}=ρ∇u¯ℒ,\\displaystyle=\\rho\\nabla\_\{\\bar\{u\}\}\\mathcal\{L\},∇v\(ρ\)ℒ\\displaystyle\\nabla\_\{v^\{\(\\rho\)\}\}\\mathcal\{L\}=ρ−1∇vℒ,\\displaystyle=\\rho^\{\-1\}\\nabla\_\{v\}\\mathcal\{L\},ddtu¯\(ρ\)^\\displaystyle\\frac\{\\,\\mathrm\{d\}\}\{\\,\\mathrm\{d\}t\}\\widehat\{\\bar\{u\}^\{\(\\rho\)\}\}=ρ2ddtu¯^,\\displaystyle=\\rho^\{2\}\\frac\{\\,\\mathrm\{d\}\}\{\\,\\mathrm\{d\}t\}\\widehat\{\\bar\{u\}\},ddtv\(ρ\)^\\displaystyle\\frac\{\\,\\mathrm\{d\}\}\{\\,\\mathrm\{d\}t\}\\widehat\{v^\{\(\\rho\)\}\}=ρ−2ddtv^\.\\displaystyle=\\rho^\{\-2\}\\frac\{\\,\\mathrm\{d\}\}\{\\,\\mathrm\{d\}t\}\\widehat\{v\}\.For a two\-layer neuron with fixedq\>0q\>0, reciprocal liftsρ\\rhoandρ−1\\rho^\{\-1\}have marksδ=±q\(ρ2−ρ−2\)\\delta=\\pm q\(\\rho^\{2\}\-\\rho^\{\-2\}\)and an input\-direction mobility ratioρ4=Θ\(D2\)\\rho^\{4\}=\\Theta\(D^\{2\}\)\. Proposition[A\.2](https://arxiv.org/html/2608.06766#A1.SS2)proves \([6](https://arxiv.org/html/2608.06766#S3.E6)\), including jointly rescaled node biases\.
The optimizer geometry determines whether this hidden state matters\. An exactly gauge\-equivariant update preserves gauge\-related function trajectories\. Under symmetricℓ2\\ell\_\{2\}weight decay with coefficientλwd\\lambda\_\{\\rm wd\}, the mark obeysδ˙=−2λwdδ\\dot\{\\delta\}=\-2\\lambda\_\{\\rm wd\}\\deltaand decays exponentially\. These two controls locate the mechanism within ordinary Euclidean optimization; see[sections˜A\.2](https://arxiv.org/html/2608.06766#A1.SS2)and[A\.1](https://arxiv.org/html/2608.06766#A1.SS1)\.
We next integrate the local law in a solvable model\. Letx∼𝒩\(0,Id\)x\\sim\\mathcal\{N\}\(0,I\_\{d\}\),f⋆\(x\)=q⋆\[s⋆⊤x\]\+f\_\{\\star\}\(x\)=q\_\{\\star\}\[s\_\{\\star\}^\{\\top\}x\]\_\{\+\}, andc=s⊤s⋆c=s^\{\\top\}s\_\{\\star\}\. With the arc\-cosine kernel
κ\(c\)=1−c2\+\(π−arccosc\)c2π,\\kappa\(c\)=\\frac\{\\sqrt\{1\-c^\{2\}\}\+\(\\pi\-\\arccos c\)c\}\{2\\pi\},the marked flow closes exactly:
q˙=−μ\(q,δ\)\(q2−q⋆κ\(c\)\),c˙=χ\(q,δ\)q⋆κ′\(c\)\(1−c2\),δ˙=0\.\\dot\{q\}=\-\\mu\(q,\\delta\)\\left\(\\frac\{q\}\{2\}\-q\_\{\\star\}\\kappa\(c\)\\right\),\\qquad\\dot\{c\}=\\chi\(q,\\delta\)q\_\{\\star\}\\kappa^\{\\prime\}\(c\)\(1\-c^\{2\}\),\\qquad\\dot\{\\delta\}=0\.\(7\)
###### Theorem 3\.2\(Same\-predictor specialization gap; informal\)\.
Fixq0\>0q\_\{0\}\>0and−1<c0<1\-1<c\_\{0\}<1\. Every solution of \([7](https://arxiv.org/html/2608.06766#S3.E7)\) converges to\(q⋆,1\)\(q\_\{\\star\},1\)\. Realize the same visible initial state\(q0,c0\)\(q\_\{0\},c\_\{0\}\)with opposite marksδ=±D\\delta=\\pm D, and letTδ\(ε\)=inf\{t:cδ\(t\)≥1−ε\}T\_\{\\delta\}\(\\varepsilon\)=\\inf\\\{t:c\_\{\\delta\}\(t\)\\geq 1\-\\varepsilon\\\}\. For all sufficiently largeDD,
T\+D\(ε\)=Θ\(D−1log1ε\),T−D\(ε\)=Θ\(Dlog1ε\),T−DT\+D=Θ\(D2\)\.T\_\{\+D\}\(\\varepsilon\)=\\Theta\\\!\\left\(D^\{\-1\}\\log\\frac\{1\}\{\\varepsilon\}\\right\),\\qquad T\_\{\-D\}\(\\varepsilon\)=\\Theta\\\!\\left\(D\\log\\frac\{1\}\{\\varepsilon\}\\right\),\\qquad\\frac\{T\_\{\-D\}\}\{T\_\{\+D\}\}=\\Theta\(D^\{2\}\)\.\(8\)At every common nonstationary visible state, the two vector fields are non\-collinear; hence no scalar learning\-rate or global time reparameterization matches their feature paths\.
The formal convergence, two\-sided hitting\-time bounds, and non\-clock statement are[theorems˜A\.9](https://arxiv.org/html/2608.06766#A1.Thmtheorem9),[A\.10](https://arxiv.org/html/2608.06766#A1.Thmtheorem10)and[A\.5](https://arxiv.org/html/2608.06766#A1.SS5), respectively\.
The two scales in \([8](https://arxiv.org/html/2608.06766#S3.E8)\) make the ratio sharp: the favorable lift specializes on the accelerated scaleD−1D^\{\-1\}, while the unfavorable lift requires the retarded scaleDD\. Non\-collinearity is a separate geometric conclusion, because a monotone change of time can alter speed but not tangent direction\. The implied constants hold for fixed\(q0,c0,q⋆,ε\)\(q\_\{0\},c\_\{0\},q\_\{\\star\},\\varepsilon\)and need not remain uniform near degenerate boundary values\.
### 3\.3Gauge\-selected feature ownership
Let the teacher bef⋆\(x\)=\[u⊤x\]\+f\_\{\\star\}\(x\)=\[u^\{\\top\}x\]\_\{\+\}and initializem≥2m\\geq 2students at the same visible state
qi\(0\)=1m,si\(0\)=s0,ϕ=∠\(s0,u\)∈Iϕ:=\[0\.77,0\.90\]\.q\_\{i\}\(0\)=\\frac\{1\}\{m\},\\qquad s\_\{i\}\(0\)=s\_\{0\},\\qquad\\phi=\\angle\(s\_\{0\},u\)\\in I\_\{\\phi\}:=\[0\.77,0\.90\]\.Assign one designated indexkkthe mark\+D\+Dand every other index the mark−D\-D\.
###### Theorem 3\.3\(Hidden\-gauge ownership and functional pruning; informal\)\.
For each fixedmmandϕ∈Iϕ\\phi\\in I\_\{\\phi\}, every sufficiently largeDDselects unitkkas the unique owner\. The trajectory enters a winner tube bytcap≤CmD−1t\_\{\\rm cap\}\\leq C\_\{m\}D^\{\-1\}and then satisfies
ℒ\(t\)≤ℒ\(tcap\)e−cmD\(t−tcap\),qk\(t\)sk\(t\)→u,‖fred\(t\)‖L2→0\.\\mathcal\{L\}\(t\)\\leq\\mathcal\{L\}\(t\_\{\\rm cap\}\)e^\{\-c\_\{m\}D\(t\-t\_\{\\rm cap\}\)\},\\qquad q\_\{k\}\(t\)s\_\{k\}\(t\)\\to u,\\qquad\\\|f\_\{\\rm red\}\(t\)\\\|\_\{L^\{2\}\}\\to 0\.For exact duplicate initialization, every redundant coefficient tends to zero and
∫tcap∞‖s˙j\(t\)‖dt≤CmD−2\(j≠k\)\.\\int\_\{t\_\{\\rm cap\}\}^\{\\infty\}\\\|\\dot\{s\}\_\{j\}\(t\)\\\|\\,\\,\\mathrm\{d\}t\\leq C\_\{m\}D^\{\-2\}\\qquad\(j\\neq k\)\.Permuting the unique positive mark permutes the owner while leaving the complete initial predictor unchanged\.
The proof has two named stages\.*Fast gauge\-selected capture*uses exchangeability to reduce the system to one designated unit and one aggregate redundant block\. On fast timeτ=Dt\\tau=Dt, positive\-gauge transport is order one while negative\-gauge transport is orderD−2D^\{\-2\}; a certified multiplicity\-weighted arc\-cosine flow reaches the winner basin, and the finite\-DDtrajectory shadows it to orderD−2D^\{\-2\}\.*Post\-capture locking*uses local feature\-Gram coercivity and the dissipation identity to obtain exponential convergence and summable redundant transport\. The formal result is[theorem˜B\.7](https://arxiv.org/html/2608.06766#A2.Thmtheorem7);[theorems˜C\.1](https://arxiv.org/html/2608.06766#A3.Thmtheorem1)and[C\.2](https://arxiv.org/html/2608.06766#A3.SS2)extend its capture basin and quantify sufficient multiplicity dependence\.
This construction also identifies the assignment mechanism cleanly\. At initialization, every unit has the same\(qi,si\)\(q\_\{i\},s\_\{i\}\)and the complete predictor is invariant under permutations of unit labels; the location of the unique positive mark is the only index\-dependent input\. Permutation equivariance of gradient flow and the theorem together show that moving this mark moves the owner\. The conclusion is therefore an intervention at a fixed function, rather than a correlation between initial feature quality and eventual specialization\. It does not assert that large gauge disparities arise typically under standard random initialization; it shows that when such hidden disparities are present, they can be causally decisive\.
### 3\.4Robustness and discrete gradient descent
For one positive andm−1m\-1negative marks, measure departure from duplicate initialization by
Δ0=maxi\(\|qi\(0\)−1m\|\+‖si\(0\)−s0‖\)\+maxi\|\|δi\(0\)\|D−1\|\.\\Delta\_\{0\}=\\max\_\{i\}\\left\(\\left\|q\_\{i\}\(0\)\-\\frac\{1\}\{m\}\\right\|\+\\\|s\_\{i\}\(0\)\-s\_\{0\}\\\|\\right\)\+\\max\_\{i\}\\left\|\\frac\{\|\\delta\_\{i\}\(0\)\|\}\{D\}\-1\\right\|\.\(9\)
###### Theorem 3\.4\(Open\-set finite\-time selection; informal\)\.
For every fixedmm, there areεm\>0\\varepsilon\_\{m\}\>0andD0\(m\)<∞D\_\{0\}\(m\)<\\inftysuch that, ifϕ∈Iϕ\\phi\\in I\_\{\\phi\},Δ0≤εm\\Delta\_\{0\}\\leq\\varepsilon\_\{m\}, andD≥2D0\(m\)D\\geq 2D\_\{0\}\(m\), the positive\-gauge unit is the unique unit to enter a fixed winner\-dominant set by time\(T⋆\+1\)/D\(T\_\{\\star\}\+1\)/D\. The proof supplies the conservative sufficient scales
εm≥ae−bm3,D0\(m\)≤AeBm2\\varepsilon\_\{m\}\\geq ae^\{\-bm^\{3\}\},\\qquad D\_\{0\}\(m\)\\leq Ae^\{Bm^\{2\}\}for universal positive constantsa,b,A,Ba,b,A,B\.
This finite\-time statement is unconditional under the displayed assumptions; its formal version is[theorem˜C\.6](https://arxiv.org/html/2608.06766#A3.Thmtheorem6)\.
###### Proposition 3\.5\(Conditional locking after perturbed capture; informal\)\.
Suppose a selected post\-capture state has positive winner–redundant quotient transversality, an active redundant coefficient frame, and the explicit continuation margins of[theorem˜C\.11](https://arxiv.org/html/2608.06766#A3.Thmtheorem11)\. Then the winner tube is forward invariant, the loss decays ase−cDte^\{\-cDt\}, the selected feature converges to the teacher, every redundant coefficient vanishes, and total redundant\-direction motion isO\(D−2\)O\(D^\{\-2\}\)\.
The active\-frame hypothesis controls coefficient cancellation modes that functional pruning alone cannot identify\. The complete conditions and proof appear in[theorem˜C\.11](https://arxiv.org/html/2608.06766#A3.Thmtheorem11)\.
###### Proposition 3\.6\(Discrete full\-batch gradient descent; informal\)\.
For exact duplicates, letη=h/D\\eta=h/Dand write the pre\-capture state asY=\(q,P,θ,ψ\)Y=\(q,P,\\theta,\\psi\)\. For every fixedmm, sufficiently smallhhand sufficiently largeDDselect the same positive\-gauge owner, with
maxn≤ncap‖Yn−Y∞\(nh\)‖≤Cm\(h\+D−2\),ncap≤⌈T⋆\+1h⌉\.\\max\_\{n\\leq n\_\{\\rm cap\}\}\\\|Y\_\{n\}\-Y\_\{\\infty\}\(nh\)\\\|\\leq C\_\{m\}\(h\+D^\{\-2\}\),\\qquad n\_\{\\rm cap\}\\leq\\left\\lceil\\frac\{T\_\{\\star\}\+1\}\{h\}\\right\\rceil\.After capture,ℒn\+1≤\(1−cmh\)ℒn\\mathcal\{L\}\_\{n\+1\}\\leq\(1\-c\_\{m\}h\)\\mathcal\{L\}\_\{n\}; the total gauge drift isOm\(h/D\)O\_\{m\}\(h/D\)and the redundant\-direction path isOm\(D−2\)O\_\{m\}\(D^\{\-2\}\)\.
The formal discrete theorem, including chart preservation and the loss\-weighted Euler remainder, is[theorem˜C\.19](https://arxiv.org/html/2608.06766#A3.Thmtheorem19)\.
## 4Experiments
All experiments use the Gaussian teacher–student model analyzed above\. We compare the multiplicity\-weighted fast system, the exact finite\-gauge marked ODE, the original\(ai,wi\)\(a\_\{i\},w\_\{i\}\)population flow, and the original\-parameter empirical flow\. The capture criterion is fixed before every sweep\.
The four levels serve distinct checks\. Agreement between marked and original\-parameter population flows audits the exact coordinate reduction; convergence of the finite\-gauge flow to the fast system tests the singular limit; and empirical trajectories probe sampling error without entering any proof\. We compare full functionally identifiable trajectories rather than selected endpoints\. For each coupled seed, finite\-sample datasets are nested acrossNN, and uncertainty is resampled at the seed level\. No theoretical coefficient is fit to empirical trajectories, and the displayed reference slopes are anchored guides\. Appendix[D](https://arxiv.org/html/2608.06766#A4)gives the fixed capture tube, trajectory metric, solver audit, and complete sweep tables\.
### 4\.1Controlled same\-predictor counterfactual
Figure[1](https://arxiv.org/html/2608.06766#S1.F1)isolates the central intervention\. The two students represent the same initial function and receive the same initial functional signal; only the hidden gauge differs\. Their trajectories leave the shared visible state along distinct curves\. The favorable gauge rapidly rotates toward the teacher, whereas the opposite gauge first changes its functional coefficient with negligible alignment\. The specialization\-time ratio follows the predicted quadratic law\. Both singular\-limit residuals converge at orderD−2D^\{\-2\};[figure˜D\.1](https://arxiv.org/html/2608.06766#A4.F1)reports these numerical checks\.
### 4\.2Feature ownership and full\-trajectory correspondence
Figure[2](https://arxiv.org/html/2608.06766#S4.F2)shows a representative\(m,D,N\)=\(8,32,8192\)\(m,D,N\)=\(8,32,8192\)run\. One reduced theory simultaneously predicts population loss, growth of the selected coefficient, redundant\-mass decay, feature alignment, and the reaction–transport dissipation split\. Across the complete population grid, the original\-parameter and marked flows differ by less than1\.3×10−71\.3\\times 10^\{\-7\}, and population gauge drift stays below2\.2×10−112\.2\\times 10^\{\-11\}\.
Figure 2:Feature ownership and full\-trajectory correspondence\.\(a\) The programmed owner captures the teacher coefficient while redundant functional mass vanishes; the shared vertical line marksτcap\\tau\_\{\\rm cap\}\. \(b\)–\(d\) The same four dynamical levels agree on population loss, functionally weighted feature residuals, and reaction–transport dissipation\. Solid and dashed curves denote the exact marked and fast systems; open circles and diamonds denote the original\-parameter population andN=8192N=8192finite\-sample flows\.
### 4\.3Predicted scaling laws
We sweepm∈\{2,4,8,16,32\}m\\in\\\{2,4,8,16,32\\\}andD∈\{8,16,32,64,128\}D\\in\\\{8,16,32,64,128\\\}\. Across multiplicities, fitted capture\-time slopes range from−1\.00\-1\.00to−0\.995\-0\.995, post\-capture redundant\-drift slopes from−2\.00\-2\.00to−1\.99\-1\.99, and finite\-gauge approximation slopes from−2\.00\-2\.00to−1\.99\-1\.99; every correspondingR2R^\{2\}exceeds0\.99990\.9999\. Panels \(a\)–\(b\) of[figure˜3](https://arxiv.org/html/2608.06766#S4.F3)display the ownership\-level laws, while[figure˜D\.1](https://arxiv.org/html/2608.06766#A4.F1)reports the finite\-gauge approximation\.
For finite samples we usem∈\{2,4,8,16\}m\\in\\\{2,4,8,16\\\},N∈\{512,1024,2048,4096,8192\}N\\in\\\{512,1024,2048,4096,8192\\\}, and 20 coupled seeds per configuration \(400 trajectories\)\. Every run captures the programmed owner\. The fitted functional\-state error slopes range from−0\.45\-0\.45to−0\.49\-0\.49, compatible with fixed\-system Monte Carlo scalingN−1/2N^\{\-1/2\}\. Panel \(c\) reports paired\-bootstrap 95% confidence intervals over coupled seeds\. The guide is empirical and does not assert a width or propagation\-of\-chaos theorem\.
Figure 3:Scaling laws and discrete shadowing\.\(a\) Winner capture followsD−1D^\{\-1\}\. \(b\) Post\-capture redundant\-direction drift followsD−2D^\{\-2\}\. \(c\) Finite\-sample full\-trajectory error is compatible withN−1/2N^\{\-1/2\}; light points are individual seeds and bars are paired\-bootstrap 95% confidence intervals\. \(d\) Discrete full\-batch GD shadows gradient flow to first order inh=ηDh=\\eta D\. Reference lines are anchored atD=32D=32,N=2048N=2048, andh=0\.1h=0\.1and indicate predicted slopes rather than additional fits\.
### 4\.4Robustness and discrete training
The angle certificate coversϕ∈\[0\.77,0\.90\]\\phi\\in\[0\.77,0\.90\], and all 24 angle–multiplicity tests capture the programmed owner\. At\(m,D\)=\(8,32\)\(m,D\)=\(8,32\), all 50 random near\-duplicate runs succeed for perturbations up to5%5\\%, well beyond the conservative theorem radius\. Acrossm=2m=2to256256, the observed rescaled capture timeDtcapDt\_\{\\rm cap\}decreases from25\.225\.2to18\.518\.5\. Discrete full\-batch GD withh=ηD∈\{0\.025,0\.05,0\.1,0\.2,0\.4\}h=\\eta D\\in\\\{0\.025,0\.05,0\.1,0\.2,0\.4\\\}always selects the same owner; its trajectory error has slope1\.121\.12inhh, consistent with the first\-order guide in panel \(d\) of[figure˜3](https://arxiv.org/html/2608.06766#S4.F3)\. Appendix[D](https://arxiv.org/html/2608.06766#A4)reports the remaining robustness audits\.
## 5Conclusion
Feature learning allocates task\-relevant structure to particular internal units\. We identified a setting in which that allocation depends on information hidden from the initial function\. At a fixed ReLU predictor, opposite gauges create a sharp non\-clockΘ\(D2\)\\Theta\(D^\{2\}\)specialization\-time gap\. Among duplicate students, the gauge assignment deterministically selects the feature owner and functionally prunes the remaining neurons\. The reaction–transport decomposition explains this behavior and predicts complete population and finite\-sample trajectories\.
The strongest global theorem concerns exact duplicates\. Finite\-time selection persists under visible perturbations, while long\-time coefficient\-wise pruning for a fully nonsymmetric redundant dictionary requires explicit transversality and active\-frame conditions\. The discrete theorem treats small\-step full\-batch gradient descent, and the finite\-sample scaling is empirical\. Within this scope, predictor\-level descriptions discard mobility information that decisively changes representation\-level competition\.
The marked coordinates show what a dynamics\-level description must retain when internal allocation matters\. Residual moments specify the functional signal, while the hidden mark specifies the metric converting it into coefficient reaction and feature transport\. Gauge\-equivariant optimization removes this dependence, whereas symmetric weight decay gradually erases the mark\. These controls offer concrete ways to distinguish or suppress gauge\-driven ownership\.
ReLU homogeneity yields the localρ2/ρ−2\\rho^\{2\}/\\rho^\{\-2\}mobility reallocation at any feedforward node, including a jointly rescaled bias\. Gaussian teacher–student geometry turns this law into global theorems\. Next steps include random gauges, multiple teacher features, and deeper interventions that preserve the distinction between local mobility and model\-specific ownership\.
## References
- S\. Akiyama and T\. Suzuki \(2021\)On learnability via gradient method for two\-layer relu neural networks in teacher\-student setting\.InInternational Conference on Machine Learning,pp\. 152–162\.Cited by:[§1](https://arxiv.org/html/2608.06766#S1.p2.1),[§2\.3](https://arxiv.org/html/2608.06766#S2.SS3.p1.1)\.
- A\. Atanasov, B\. Bordelon, and C\. Pehlevan \(2022\)Neural networks as kernel learners: the silent alignment effect\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2608.06766#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.06766#S2.SS1.p1.1)\.
- J\. Bona\-Pellissier, F\. Bachoc, and F\. Malgouyres \(2023\)Parameter identifiability of a deep feedforward relu neural network\.Machine Learning112\(11\),pp\. 4431–4493\.External Links:[Document](https://dx.doi.org/10.1007/s10994-023-06355-4)Cited by:[§2\.2](https://arxiv.org/html/2608.06766#S2.SS2.p2.1)\.
- J\. Bona\-Pellissier, F\. Malgouyres, and F\. Bachoc \(2022\)Local identifiability of deep relu neural networks: the theory\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 27549–27562\.Cited by:[§2\.2](https://arxiv.org/html/2608.06766#S2.SS2.p2.1)\.
- B\. Bordelon and C\. Pehlevan \(2025\)Deep linear network training dynamics from random initialization: data, width, depth, and hyperparameter transfer\.InInternational Conference on Machine Learning,Cited by:[§2\.4](https://arxiv.org/html/2608.06766#S2.SS4.p1.1)\.
- L\. Chizat, E\. Oyallon, and F\. Bach \(2019\)On lazy training in differentiable programming\.InAdvances in Neural Information Processing Systems,Vol\.32,pp\. 2933–2943\.Cited by:[§1](https://arxiv.org/html/2608.06766#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.06766#S2.SS1.p1.1)\.
- S\. S\. Du, W\. Hu, and J\. D\. Lee \(2018\)Algorithmic regularization in learning deep homogeneous models: layers are automatically balanced\.InAdvances in Neural Information Processing Systems,Vol\.31\.Cited by:[§A\.1](https://arxiv.org/html/2608.06766#A1.SS1.p1.8),[§1](https://arxiv.org/html/2608.06766#S1.p3.1),[§2\.2](https://arxiv.org/html/2608.06766#S2.SS2.p1.1)\.
- A\. Jacot, F\. Gabriel, and C\. Hongler \(2018\)Neural tangent kernel: convergence and generalization in neural networks\.InAdvances in Neural Information Processing Systems,Vol\.31\.Cited by:[§1](https://arxiv.org/html/2608.06766#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.06766#S2.SS1.p1.1)\.
- D\. Kunin, A\. Raventós, C\. Dominé, F\. Chen, D\. Klindt, A\. Saxe, and S\. Ganguli \(2024\)Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning\.InAdvances in Neural Information Processing Systems,Vol\.37\.Cited by:[§1](https://arxiv.org/html/2608.06766#S1.p1.1),[§1](https://arxiv.org/html/2608.06766#S1.p3.1),[§2\.1](https://arxiv.org/html/2608.06766#S2.SS1.p1.1),[§2\.2](https://arxiv.org/html/2608.06766#S2.SS2.p1.1)\.
- A\. Lebeurrier, T\. Vayer, and R\. Gribonval \(2026\)Path\-conditioned training: a principled way to rescale relu neural networks\.InProceedings of the 43rd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.306\.Note:arXiv:2602\.19799Cited by:[§1](https://arxiv.org/html/2608.06766#S1.p3.1),[§2\.2](https://arxiv.org/html/2608.06766#S2.SS2.p1.1)\.
- S\. Marcotte, R\. Gribonval, and G\. Peyré \(2023\)Abide by the law and follow the flow: conservation laws for gradient flows\.arXiv preprint arXiv:2307\.00144\.Cited by:[§1](https://arxiv.org/html/2608.06766#S1.p3.1),[§2\.2](https://arxiv.org/html/2608.06766#S2.SS2.p1.1)\.
- S\. Marcotte, G\. Peyré, and R\. Gribonval \(2025\)Intrinsic training dynamics of deep neural networks\.arXiv preprint arXiv:2508\.07370\.Cited by:[§2\.2](https://arxiv.org/html/2608.06766#S2.SS2.p1.1)\.
- S\. Mei, T\. Misiakiewicz, and A\. Montanari \(2019\)Mean\-field theory of two\-layers neural networks: dimension\-free bounds and kernel limit\.InConference on Learning Theory,pp\. 2388–2464\.Cited by:[§A\.1](https://arxiv.org/html/2608.06766#A1.SS1.p1.8),[§2\.4](https://arxiv.org/html/2608.06766#S2.SS4.p1.1)\.
- B\. Neyshabur, R\. Salakhutdinov, and N\. Srebro \(2015\)Path\-sgd: path\-normalized optimization in deep neural networks\.InAdvances in Neural Information Processing Systems,Vol\.28,pp\. 2422–2430\.Cited by:[Remark A\.6](https://arxiv.org/html/2608.06766#A1.SS2.6.p1.6),[§2\.2](https://arxiv.org/html/2608.06766#S2.SS2.p1.1)\.
- I\. M\. Safran, G\. Yehudai, and O\. Shamir \(2021\)The effects of mild over\-parameterization on the optimization landscape of shallow relu neural networks\.InConference on Learning Theory,pp\. 3889–3934\.Cited by:[§2\.3](https://arxiv.org/html/2608.06766#S2.SS3.p1.1)\.
- R\. C\. Terin \(2026\)Scale redundancy and soft gauge fixing in positively homogeneous neural networks\.arXiv preprint arXiv:2602\.14729\.Cited by:[§2\.2](https://arxiv.org/html/2608.06766#S2.SS2.p1.1)\.
- Y\. Tian \(2017\)An analytical formula of population gradient for two\-layered relu network and its applications in convergence and critical point analysis\.InInternational Conference on Machine Learning,pp\. 3404–3413\.Cited by:[§A\.3](https://arxiv.org/html/2608.06766#A1.SS3.p1.5),[§1](https://arxiv.org/html/2608.06766#S1.p2.1),[§2\.3](https://arxiv.org/html/2608.06766#S2.SS3.p1.1)\.
- Y\. Tian \(2020\)Student specialization in deep rectified networks with finite width and input dimension\.InInternational Conference on Machine Learning,pp\. 9470–9480\.Cited by:[§1](https://arxiv.org/html/2608.06766#S1.p2.1),[§2\.3](https://arxiv.org/html/2608.06766#S2.SS3.p1.1)\.
- F\. Williams, M\. Trager, C\. Silva, D\. Panozzo, D\. Zorin, and J\. Bruna \(2019\)Gradient dynamics of shallow univariate relu networks\.InAdvances in Neural Information Processing Systems,Vol\.32\.Cited by:[§1](https://arxiv.org/html/2608.06766#S1.p3.1),[§2\.2](https://arxiv.org/html/2608.06766#S2.SS2.p1.1)\.
- W\. Xu and S\. Du \(2023\)Over\-parameterization exponentially slows down gradient descent for learning a single neuron\.InConference on Learning Theory,pp\. 1155–1198\.Cited by:[§2\.3](https://arxiv.org/html/2608.06766#S2.SS3.p1.1)\.
- G\. Yang and E\. J\. Hu \(2021\)Tensor programs iv: feature learning in infinite\-width neural networks\.InProceedings of the 38th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.139,pp\. 11727–11737\.Cited by:[§1](https://arxiv.org/html/2608.06766#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.06766#S2.SS1.p1.1)\.
- M\. Zhou, R\. Ge, and C\. Jin \(2021\)A local convergence theory for mildly over\-parameterized two\-layer neural network\.InConference on Learning Theory,pp\. 4577–4632\.Cited by:[§1](https://arxiv.org/html/2608.06766#S1.p2.1),[§2\.3](https://arxiv.org/html/2608.06766#S2.SS3.p1.1)\.
## Appendix roadmap
The appendices follow the proof dependencies of the main text\.
- •Appendix Aderives the exact marked dynamics, proves the node\-mobility law, and establishes one\-neuron specialization and the non\-clock quadratic time gap\.
- •Appendix Bproves global feature ownership and pruning for exactly duplicate students through fast capture and post\-capture locking\.
- •Appendix Cextends the capture analysis to an angle interval, quantifies multiplicity dependence, proves open\-set finite\-time selection and conditional locking, and treats discrete full\-batch gradient descent\.
- •Appendix Dspecifies the numerical protocol and reports the trajectory, scaling, robustness, and discretization audits corresponding to Figures 1–3\.
Appendix B builds on Appendix A, and Appendix C builds on both\. Appendix D provide numerical validation\.
## Appendix AExact marked dynamics and one\-neuron specialization
### A\.1Population gradient flow and the regular marked chart
Letπ\\pibe a probability measure onℝd\\mathbb\{R\}^\{d\}with finite second moment and assume that it assigns zero mass to every hyperplane used below\. Lety∈L2\(π\)y\\in L^\{2\}\(\\pi\)andσ\(z\)=z\+\\sigma\(z\)=z\_\{\+\}\. For a finite positive particle measureϱ\\varrhoon
𝒫:=ℝ×\(ℝd∖\{0\}\),\\mathcal\{P\}:=\\mathbb\{R\}\\times\(\\mathbb\{R\}^\{d\}\\setminus\\\{0\\\}\),define
fϱ\(x\)=∫aσ\(w⊤x\)ϱ\(da,dw\),ℒ\(ϱ\)=12∫\(fϱ\(x\)−y\(x\)\)2π\(dx\)\.f\_\{\\varrho\}\(x\)=\\int a\\,\\sigma\(w^\{\\top\}x\)\\,\\varrho\(\\,\\mathrm\{d\}a,\\,\\mathrm\{d\}w\),\\qquad\\mathcal\{L\}\(\\varrho\)=\\frac\{1\}\{2\}\\int\\bigl\(f\_\{\\varrho\}\(x\)\-y\(x\)\\bigr\)^\{2\}\\,\\pi\(\\,\\mathrm\{d\}x\)\.\(A\.1\)WriteRϱ=fϱ−yR\_\{\\varrho\}=f\_\{\\varrho\}\-y\. Along a characteristic of the Wasserstein/population gradient flow,
a˙=−∫Rϱ\(x\)σ\(w⊤x\)π\(dx\),w˙=−a∫Rϱ\(x\)𝟏\{w⊤x\>0\}xπ\(dx\)\.\\dot\{a\}=\-\\int R\_\{\\varrho\}\(x\)\\sigma\(w^\{\\top\}x\)\\,\\pi\(\\,\\mathrm\{d\}x\),\\qquad\\dot\{w\}=\-a\\int R\_\{\\varrho\}\(x\)\\mathbf\{1\}\_\{\\\{w^\{\\top\}x\>0\\\}\}x\\,\\pi\(\\,\\mathrm\{d\}x\)\.\(A\.2\)This is the standard two\-layer mean\-field characteristic system\(Duet al\.,[2018](https://arxiv.org/html/2608.06766#bib.bib1); Meiet al\.,[2019](https://arxiv.org/html/2608.06766#bib.bib5)\)\. We derive the exact marked reduction needed below\.
Forw≠0w\\neq 0, set
r=‖w‖,s=wr∈𝕊d−1,q=ar,δ=a2−r2\.r=\\\|w\\\|,\\qquad s=\\frac\{w\}\{r\}\\in\\mathbb\{S\}^\{d\-1\},\\qquad q=ar,\\qquad\\delta=a^\{2\}\-r^\{2\}\.\(A\.3\)Define
μ\(q,δ\):=δ2\+4q2,r2\(q,δ\)=μ\(q,δ\)−δ2,χ\(q,δ\):=2qμ\(q,δ\)−δ\.\\mu\(q,\\delta\):=\\sqrt\{\\delta^\{2\}\+4q^\{2\}\},\\qquad r^\{2\}\(q,\\delta\)=\\frac\{\\mu\(q,\\delta\)\-\\delta\}\{2\},\\qquad\\chi\(q,\\delta\):=\\frac\{2q\}\{\\mu\(q,\\delta\)\-\\delta\}\.\(A\.4\)The marked state space is
𝒳:=\{\(q,s,δ\):s∈𝕊d−1,μ\(q,δ\)\>δ\}\.\\mathcal\{X\}:=\\bigl\\\{\(q,s,\\delta\):s\\in\\mathbb\{S\}^\{d\-1\},\\ \\mu\(q,\\delta\)\>\\delta\\bigr\\\}\.\(A\.5\)The condition is exactlyr\>0r\>0\. The formula forχ\\chiis regular atq=0q=0wheneverδ<0\\delta<0and equalsa/ra/r\. The inverse chart is
r=μ−δ2,a=qr,w=rs\.r=\\sqrt\{\\frac\{\\mu\-\\delta\}\{2\}\},\\qquad a=\\frac\{q\}\{r\},\\qquad w=rs\.\(A\.6\)Hence the map\(a,w\)↦\(q,s,δ\)\(a,w\)\\mapsto\(q,s,\\delta\)is a smooth bijection from𝒫\\mathcal\{P\}onto𝒳\\mathcal\{X\}\.
Letρ\\rhobe the pushforward ofϱ\\varrho\. Positive homogeneity gives
fρ\(x\)=∫qσ\(s⊤x\)ρ\(dq,ds,dδ\)\.f\_\{\\rho\}\(x\)=\\int q\\,\\sigma\(s^\{\\top\}x\)\\,\\rho\(\\,\\mathrm\{d\}q,\\,\\mathrm\{d\}s,\\,\\mathrm\{d\}\\delta\)\.\(A\.7\)Define the residual moments
Bρ\(s\):=∫Rρ\(x\)σ\(s⊤x\)π\(dx\),𝒯ρ\(s\):=∫Rρ\(x\)𝟏\{s⊤x\>0\}xπ\(dx\)\.B\_\{\\rho\}\(s\):=\\int R\_\{\\rho\}\(x\)\\sigma\(s^\{\\top\}x\)\\,\\pi\(\\,\\mathrm\{d\}x\),\\qquad\\mathcal\{T\}\_\{\\rho\}\(s\):=\\int R\_\{\\rho\}\(x\)\\mathbf\{1\}\_\{\\\{s^\{\\top\}x\>0\\\}\}x\\,\\pi\(\\,\\mathrm\{d\}x\)\.\(A\.8\)By homogeneity,
Bρ\(s\)=s⊤𝒯ρ\(s\)\.B\_\{\\rho\}\(s\)=s^\{\\top\}\\mathcal\{T\}\_\{\\rho\}\(s\)\.\(A\.9\)Let𝖯s⟂=I−ss⊤\\mathsf\{P\}\_\{s\}^\{\\perp\}=I\-ss^\{\\top\}\.
###### Theorem A\.1\(Exact gauge\-marked reduction\)\.
Let\(ϱt\)t∈\[0,T\]\(\\varrho\_\{t\}\)\_\{t\\in\[0,T\]\}be a characteristic solution of \([A\.2](https://arxiv.org/html/2608.06766#A1.E2)\) whose support remains in𝒫\\mathcal\{P\}\. Its marked pushforwardρt\\rho\_\{t\}satisfies
∂tρ\+∂q\(Vqρ\)\+div𝕊d−1\(Vsρ\)=0,Vδ=0,\\partial\_\{t\}\\rho\+\\partial\_\{q\}\(V\_\{q\}\\rho\)\+\\operatorname\{div\}\_\{\\mathbb\{S\}^\{d\-1\}\}\(V\_\{s\}\\rho\)=0,\\qquad V\_\{\\delta\}=0,\(A\.10\)with the exact velocities
Vq=−μ\(q,δ\)Bρ\(s\),Vs=−χ\(q,δ\)𝖯s⟂𝒯ρ\(s\),Vδ=0\.V\_\{q\}=\-\\mu\(q,\\delta\)B\_\{\\rho\}\(s\),\\qquad V\_\{s\}=\-\\chi\(q,\\delta\)\\mathsf\{P\}\_\{s\}^\{\\perp\}\\mathcal\{T\}\_\{\\rho\}\(s\),\\qquad V\_\{\\delta\}=0\.\(A\.11\)Conversely, every characteristic solution of \([A\.10](https://arxiv.org/html/2608.06766#A1.E10)\) that remains in𝒳\\mathcal\{X\}pulls back through \([A\.6](https://arxiv.org/html/2608.06766#A1.E6)\) to a solution of \([A\.2](https://arxiv.org/html/2608.06766#A1.E2)\)\. Thus the two systems are bidirectionally equivalent up to the first exit from the regular chart\.
Moreover, the loss satisfies the exact reaction–transport dissipation identity
ddtℒ\(ρt\)=−∫\[μ\(q,δ\)Bρt\(s\)2⏟coefficient reaction\+μ\(q,δ\)\+δ2‖𝖯s⟂𝒯ρt\(s\)‖2⏟directional transport\]ρt\(dq,ds,dδ\)\.\\frac\{\\,\\mathrm\{d\}\}\{\\,\\mathrm\{d\}t\}\\mathcal\{L\}\(\\rho\_\{t\}\)=\-\\int\\left\[\\underbrace\{\\mu\(q,\\delta\)B\_\{\\rho\_\{t\}\}\(s\)^\{2\}\}\_\{\\text\{coefficient reaction\}\}\+\\underbrace\{\\frac\{\\mu\(q,\\delta\)\+\\delta\}\{2\}\\bigl\\\|\\mathsf\{P\}\_\{s\}^\{\\perp\}\\mathcal\{T\}\_\{\\rho\_\{t\}\}\(s\)\\bigr\\\|^\{2\}\}\_\{\\text\{directional transport\}\}\\right\]\\rho\_\{t\}\(\\,\\mathrm\{d\}q,\\,\\mathrm\{d\}s,\\,\\mathrm\{d\}\\delta\)\.\(A\.12\)In particular, the entireδ\\delta\-marginal is conserved\.
###### Proof\.
Sinceσ\(w⊤x\)=rσ\(s⊤x\)\\sigma\(w^\{\\top\}x\)=r\\sigma\(s^\{\\top\}x\), \([A\.2](https://arxiv.org/html/2608.06766#A1.E2)\) becomes
a˙=−rBρ\(s\),w˙=−a𝒯ρ\(s\)\.\\dot\{a\}=\-rB\_\{\\rho\}\(s\),\\qquad\\dot\{w\}=\-a\\mathcal\{T\}\_\{\\rho\}\(s\)\.\(A\.13\)Taking radial and tangential components gives
r˙=s⊤w˙=−aBρ\(s\),s˙=1r𝖯s⟂w˙=−ar𝖯s⟂𝒯ρ\(s\)\.\\dot\{r\}=s^\{\\top\}\\dot\{w\}=\-aB\_\{\\rho\}\(s\),\\qquad\\dot\{s\}=\\frac\{1\}\{r\}\\mathsf\{P\}\_\{s\}^\{\\perp\}\\dot\{w\}=\-\\frac\{a\}\{r\}\\mathsf\{P\}\_\{s\}^\{\\perp\}\\mathcal\{T\}\_\{\\rho\}\(s\)\.\(A\.14\)Therefore
q˙\\displaystyle\\dot\{q\}=ra˙\+ar˙=−\(r2\+a2\)Bρ\(s\)=−μ\(q,δ\)Bρ\(s\),\\displaystyle=r\\dot\{a\}\+a\\dot\{r\}=\-\(r^\{2\}\+a^\{2\}\)B\_\{\\rho\}\(s\)=\-\\mu\(q,\\delta\)B\_\{\\rho\}\(s\),\(A\.15\)δ˙\\displaystyle\\dot\{\\delta\}=2aa˙−2rr˙=−2arBρ\(s\)\+2arBρ\(s\)=0\.\\displaystyle=2a\\dot\{a\}\-2r\\dot\{r\}=\-2arB\_\{\\rho\}\(s\)\+2arB\_\{\\rho\}\(s\)=0\.\(A\.16\)The identitya/r=χ\(q,δ\)a/r=\\chi\(q,\\delta\)yields the statedss\-velocity\. Applying the chain rule to test functions proves the pushforward continuity equation\. The inverse calculation gives the converse\.
For dissipation, the original Euclidean gradient flow gives
ddtℒ\\displaystyle\\frac\{\\,\\mathrm\{d\}\}\{\\,\\mathrm\{d\}t\}\\mathcal\{L\}=−∫\[r2Bρ\(s\)2\+a2‖𝒯ρ\(s\)‖2\]dρ\\displaystyle=\-\\int\\left\[r^\{2\}B\_\{\\rho\}\(s\)^\{2\}\+a^\{2\}\\\|\\mathcal\{T\}\_\{\\rho\}\(s\)\\\|^\{2\}\\right\]\\,\\mathrm\{d\}\\rho=−∫\[\(r2\+a2\)Bρ\(s\)2\+a2‖𝖯s⟂𝒯ρ\(s\)‖2\]dρ,\\displaystyle=\-\\int\\left\[\(r^\{2\}\+a^\{2\}\)B\_\{\\rho\}\(s\)^\{2\}\+a^\{2\}\\\|\\mathsf\{P\}\_\{s\}^\{\\perp\}\\mathcal\{T\}\_\{\\rho\}\(s\)\\\|^\{2\}\\right\]\\,\\mathrm\{d\}\\rho,\(A\.17\)where \([A\.9](https://arxiv.org/html/2608.06766#A1.E9)\) was used to decompose𝒯=Bs\+𝖯s⟂𝒯\\mathcal\{T\}=Bs\+\\mathsf\{P\}\_\{s\}^\{\\perp\}\\mathcal\{T\}\. Finally,
r2\+a2=μ\(q,δ\),a2=μ\(q,δ\)\+δ2,r^\{2\}\+a^\{2\}=\\mu\(q,\\delta\),\\qquad a^\{2\}=\\frac\{\\mu\(q,\\delta\)\+\\delta\}\{2\},which proves \([A\.12](https://arxiv.org/html/2608.06766#A1.E12)\)\. ∎
For fixedq≠0q\\neq 0, the transport sharea2/\(a2\+r2\)=12\(1\+δ/δ2\+4q2\)a^\{2\}/\(a^\{2\}\+r^\{2\}\)=\\tfrac\{1\}\{2\}\(1\+\\delta/\\sqrt\{\\delta^\{2\}\+4q^\{2\}\}\)increases from0to11asδ\\deltaranges from−∞\-\\inftyto\+∞\+\\infty\. The reaction mobility depends on\|δ\|\|\\delta\|, whereas the transport mobility distinguishes its sign\.
### A\.2Node\-gauge mobility beyond two layers
The following proposition isolates the architecture\-level calculation behind the marked two\-layer mobility\. It is intentionally local in time: Euclidean gradient flow is not gauge equivariant, so two gauge\-related initial states need not remain gauge related after training begins\.
###### Proposition A\.3\(Node\-gauge mobility in arbitrary\-depth ReLU networks\)\.
Consider a finite feedforward ReLU computation graph with independently parameterized edges and no normalization layers\. Fix a hidden node with preactivation
outgoing parameter vectorvv, and all remaining parameters collected inξ\\xi\. Write the incoming affine parameter asu¯:=\(u,b\)\\bar\{u\}:=\(u,b\); the bias\-free case isb=0b=0\. Assumeu¯≠0\\bar\{u\}\\neq 0and define, forρ\>0\\rho\>0,
u¯\(ρ\):=u¯/ρ,v\(ρ\):=ρv,Θ\(ρ\):=gρΘ:=\(u¯\(ρ\),v\(ρ\),ξ\)\.\\bar\{u\}^\{\(\\rho\)\}:=\\bar\{u\}/\\rho,\\qquad v^\{\(\\rho\)\}:=\\rho v,\\qquad\\Theta^\{\(\\rho\)\}:=g\_\{\\rho\}\\Theta:=\(\\bar\{u\}^\{\(\\rho\)\},v^\{\(\\rho\)\},\\xi\)\.Then the realized network function is invariant:
fgρΘ=fΘ\.f\_\{g\_\{\\rho\}\\Theta\}=f\_\{\\Theta\}\.\(A\.19\)Letℒ\\mathcal\{L\}be any differentiable loss that depends on the parameters only through the realized network outputs\. At every differentiability point,
∇u¯\(ρ\)ℒ\(Θ\(ρ\)\)=ρ∇u¯ℒ\(Θ\),∇v\(ρ\)ℒ\(Θ\(ρ\)\)=ρ−1∇vℒ\(Θ\),∇ξℒ\(Θ\(ρ\)\)=∇ξℒ\(Θ\)\.\\nabla\_\{\\bar\{u\}^\{\(\\rho\)\}\}\\mathcal\{L\}\(\\Theta^\{\(\\rho\)\}\)=\\rho\\,\\nabla\_\{\\bar\{u\}\}\\mathcal\{L\}\(\\Theta\),\\qquad\\nabla\_\{v^\{\(\\rho\)\}\}\\mathcal\{L\}\(\\Theta^\{\(\\rho\)\}\)=\\rho^\{\-1\}\\nabla\_\{v\}\\mathcal\{L\}\(\\Theta\),\\qquad\\nabla\_\{\\xi\}\\mathcal\{L\}\(\\Theta^\{\(\\rho\)\}\)=\\nabla\_\{\\xi\}\\mathcal\{L\}\(\\Theta\)\.\(A\.20\)Under Euclidean gradient flow, writeu¯^=u¯/‖u¯‖\\widehat\{\\bar\{u\}\}=\\bar\{u\}/\\\|\\bar\{u\}\\\|and, whenv≠0v\\neq 0,v^=v/‖v‖\\widehat\{v\}=v/\\\|v\\\|\. The instantaneous normalized\-direction velocities at the two gauge\-related states satisfy
ddtu¯\(ρ\)^\|Θ\(ρ\)=ρ2ddtu¯^\|Θ,ddtv\(ρ\)^\|Θ\(ρ\)=ρ−2ddtv^\|Θ\.\\left\.\\frac\{\\,\\mathrm\{d\}\}\{\\,\\mathrm\{d\}t\}\\widehat\{\\bar\{u\}^\{\(\\rho\)\}\}\\right\|\_\{\\Theta^\{\(\\rho\)\}\}=\\rho^\{2\}\\left\.\\frac\{\\,\\mathrm\{d\}\}\{\\,\\mathrm\{d\}t\}\\widehat\{\\bar\{u\}\}\\right\|\_\{\\Theta\},\\qquad\\left\.\\frac\{\\,\\mathrm\{d\}\}\{\\,\\mathrm\{d\}t\}\\widehat\{v^\{\(\\rho\)\}\}\\right\|\_\{\\Theta^\{\(\\rho\)\}\}=\\rho^\{\-2\}\\left\.\\frac\{\\,\\mathrm\{d\}\}\{\\,\\mathrm\{d\}t\}\\widehat\{v\}\\right\|\_\{\\Theta\}\.\(A\.21\)Thus node rescaling reallocates Euclidean mobility between incoming\-affine feature motion and outgoing\-combination motion\. If both normalized direction velocities are nonzero andρ≠1\\rho\\neq 1, their joint two\-block velocity is not a scalar multiple of the original one\.
###### Proof\.
Positive homogeneity gives
σ\(\(u/ρ\)⊤h\+b/ρ\)=ρ−1σ\(u⊤h\+b\)\.\\sigma\\\!\\left\(\(u/\\rho\)^\{\\top\}h\+b/\\rho\\right\)=\\rho^\{\-1\}\\sigma\(u^\{\\top\}h\+b\)\.Multiplying every outgoing edge byρ\\rhotherefore leaves each downstream preactivation, and hence the complete network output, unchanged\. This proves \([A\.19](https://arxiv.org/html/2608.06766#A1.E19)\)\.
The identity
ℒ\(u¯/ρ,ρv,ξ\)=ℒ\(u¯,v,ξ\)\\mathcal\{L\}\(\\bar\{u\}/\\rho,\\rho v,\\xi\)=\\mathcal\{L\}\(\\bar\{u\},v,\\xi\)holds for all\(u¯,v,ξ\)\(\\bar\{u\},v,\\xi\)in every differentiability region\. Differentiating it with respect to the three parameter blocks yields \([A\.20](https://arxiv.org/html/2608.06766#A1.E20)\)\. Under Euclidean gradient flow,
u¯˙\(ρ\)=−∇u¯\(ρ\)ℒ\(Θ\(ρ\)\)=ρu¯˙,v˙\(ρ\)=−∇v\(ρ\)ℒ\(Θ\(ρ\)\)=ρ−1v˙\.\\dot\{\\bar\{u\}\}^\{\(\\rho\)\}=\-\\nabla\_\{\\bar\{u\}^\{\(\\rho\)\}\}\\mathcal\{L\}\(\\Theta^\{\(\\rho\)\}\)=\\rho\\dot\{\\bar\{u\}\},\\qquad\\dot\{v\}^\{\(\\rho\)\}=\-\\nabla\_\{v^\{\(\\rho\)\}\}\\mathcal\{L\}\(\\Theta^\{\(\\rho\)\}\)=\\rho^\{\-1\}\\dot\{v\}\.For any nonzero vectorzz,
ddtz‖z‖=1‖z‖\(I−z^z^⊤\)z˙\.\\frac\{\\,\\mathrm\{d\}\}\{\\,\\mathrm\{d\}t\}\\frac\{z\}\{\\\|z\\\|\}=\\frac\{1\}\{\\\|z\\\|\}\\left\(I\-\\widehat\{z\}\\widehat\{z\}^\{\\top\}\\right\)\\dot\{z\}\.Since‖u¯\(ρ\)‖=‖u¯‖/ρ\\\|\\bar\{u\}^\{\(\\rho\)\}\\\|=\\\|\\bar\{u\}\\\|/\\rho,‖v\(ρ\)‖=ρ‖v‖\\\|v^\{\(\\rho\)\}\\\|=\\rho\\\|v\\\|, and normalized directions are unchanged by positive scaling, \([A\.21](https://arxiv.org/html/2608.06766#A1.E21)\) follows\. The final claim follows because the two nonzero blocks are multiplied by distinct factorsρ2\\rho^\{2\}andρ−2\\rho^\{\-2\}\. ∎
###### Corollary A\.4\(Relation to the marked two\-layer gauge\)\.
For a two\-layer scalar\-output ReLU neuron with fixedq=ar\>0q=ar\>0, start from the balanced lifta=r=qa=r=\\sqrt\{q\}and apply reciprocal node gaugesρ\\rhoandρ−1\\rho^\{\-1\}\. Their imbalance marks are
δ\+=q\(ρ2−ρ−2\),δ−=−q\(ρ2−ρ−2\),\\delta\_\{\+\}=q\(\\rho^\{2\}\-\\rho^\{\-2\}\),\\qquad\\delta\_\{\-\}=\-q\(\\rho^\{2\}\-\\rho^\{\-2\}\),and their input\-direction mobilities areρ2\\rho^\{2\}andρ−2\\rho^\{\-2\}\. Hence the mobility ratio isρ4\\rho^\{4\}\. IfD=\|δ±\|D=\|\\delta\_\{\\pm\}\|andqqis fixed, thenρ2=Θ\(D\)\\rho^\{2\}=\\Theta\(D\)andρ4=Θ\(D2\)\\rho^\{4\}=\\Theta\(D^\{2\}\)asD→∞D\\to\\infty\.
At function level, a finite network
f\(x\)=∑i=1mqiσ\(si⊤x\),‖si‖=1,f\(x\)=\\sum\_\{i=1\}^\{m\}q\_\{i\}\\sigma\(s\_\{i\}^\{\\top\}x\),\\qquad\\\|s\_\{i\}\\\|=1,has signed distributional curvature
Dx2f=∑iqisi⊗siℋd−1↾\{si⊤x=0\}\.D\_\{x\}^\{2\}f=\\sum\_\{i\}q\_\{i\}\\,s\_\{i\}\\otimes s\_\{i\}\\,\\mathcal\{H\}^\{d\-1\}\\\!\\restriction\_\{\\\{s\_\{i\}^\{\\top\}x=0\\\}\}\.Both the function and its curvature depend on\(qi,si\)\(q\_\{i\},s\_\{i\}\)and are blind toδi\\delta\_\{i\}\. The mark therefore changes future directional transport without changing the current predictor\.
### A\.3Gaussian ReLU kernel and closed ODE
Letx∼𝒩\(0,Id\)x\\sim\\mathcal\{N\}\(0,I\_\{d\}\)\. Fix a teacher
f⋆\(x\)=q⋆σ\(s⋆⊤x\),q⋆\>0,‖s⋆‖=1,f\_\{\\star\}\(x\)=q\_\{\\star\}\\sigma\(s\_\{\\star\}^\{\\top\}x\),\\qquad q\_\{\\star\}\>0,\\qquad\\\|s\_\{\\star\}\\\|=1,\(A\.22\)and a student
f\(x\)=qσ\(s⊤x\),q\>0,‖s‖=1\.f\(x\)=q\\sigma\(s^\{\\top\}x\),\\qquad q\>0,\\qquad\\\|s\\\|=1\.Setc=s⊤s⋆c=s^\{\\top\}s\_\{\\star\}\. Define the arc\-cosine kernel
κ\(c\):=𝔼\[σ\(s⊤x\)σ\(s⋆⊤x\)\]=1−c2\+\(π−arccosc\)c2π\.\\kappa\(c\):=\\mathbb\{E\}\[\\sigma\(s^\{\\top\}x\)\\sigma\(s\_\{\\star\}^\{\\top\}x\)\]=\\frac\{\\sqrt\{1\-c^\{2\}\}\+\(\\pi\-\\arccos c\)c\}\{2\\pi\}\.\(A\.23\)It satisfies
κ\(−1\)=0,κ\(1\)=12,κ′\(c\)=π−arccosc2π\>0\(−1<c≤1\)\.\\kappa\(\-1\)=0,\\qquad\\kappa\(1\)=\\frac\{1\}\{2\},\\qquad\\kappa^\{\\prime\}\(c\)=\\frac\{\\pi\-\\arccos c\}\{2\\pi\}\>0\\quad\(\-1<c\\leq 1\)\.\(A\.24\)Analytical population\-gradient formulae for Gaussian teacher–student ReLU models are classical; here they are combined with the conserved gauge mark\.\(Tian,[2017](https://arxiv.org/html/2608.06766#bib.bib3)\)
###### Lemma A\.7\(Gaussian residual moments\)\.
For the one\-teacher model,
B\(s\)=q2−q⋆κ\(c\),B\(s\)=\\frac\{q\}\{2\}\-q\_\{\\star\}\\kappa\(c\),\(A\.25\)and
𝖯s⟂𝒯\(s\)=−q⋆κ′\(c\)𝖯s⟂s⋆\.\\mathsf\{P\}\_\{s\}^\{\\perp\}\\mathcal\{T\}\(s\)=\-q\_\{\\star\}\\kappa^\{\\prime\}\(c\)\\mathsf\{P\}\_\{s\}^\{\\perp\}s\_\{\\star\}\.\(A\.26\)The population loss is
ℒ\(q,c\)=q24\+q⋆24−qq⋆κ\(c\)\.\\mathcal\{L\}\(q,c\)=\\frac\{q^\{2\}\}\{4\}\+\\frac\{q\_\{\\star\}^\{2\}\}\{4\}\-qq\_\{\\star\}\\kappa\(c\)\.\(A\.27\)
###### Proof\.
The identities𝔼\[σ\(s⊤x\)2\]=1/2\\mathbb\{E\}\[\\sigma\(s^\{\\top\}x\)^\{2\}\]=1/2and \([A\.23](https://arxiv.org/html/2608.06766#A1.E23)\) give \([A\.25](https://arxiv.org/html/2608.06766#A1.E25)\) and \([A\.27](https://arxiv.org/html/2608.06766#A1.E27)\)\. For a tangent perturbationv⟂sv\\perp s,
v⊤𝔼\[x𝟏\{s⊤x\>0\}σ\(s⋆⊤x\)\]=ddϵ\|ϵ=0κ\(s\+ϵv‖s\+ϵv‖⋅s⋆\)=κ′\(c\)v⊤s⋆\.v^\{\\top\}\\mathbb\{E\}\\left\[x\\mathbf\{1\}\_\{\\\{s^\{\\top\}x\>0\\\}\}\\sigma\(s\_\{\\star\}^\{\\top\}x\)\\right\]=\\frac\{\\,\\mathrm\{d\}\}\{\\,\\mathrm\{d\}\\epsilon\}\\bigg\|\_\{\\epsilon=0\}\\kappa\\\!\\left\(\\frac\{s\+\\epsilon v\}\{\\\|s\+\\epsilon v\\\|\}\\cdot s\_\{\\star\}\\right\)=\\kappa^\{\\prime\}\(c\)v^\{\\top\}s\_\{\\star\}\.The self term is parallel toss, so tangential projection yields \([A\.26](https://arxiv.org/html/2608.06766#A1.E26)\)\. ∎
###### Theorem A\.8\(Exact gauge\-marked specialization ODE\)\.
Fix a gauge markδ∈ℝ\\delta\\in\\mathbb\{R\}\. On the regular regionq\>0q\>0and−1<c<1\-1<c<1, population gradient flow is exactly
q˙=−μ\(q,δ\)\(q2−q⋆κ\(c\)\),c˙=χ\(q,δ\)q⋆κ′\(c\)\(1−c2\),δ˙=0\.\\dot\{q\}=\-\\mu\(q,\\delta\)\\left\(\\frac\{q\}\{2\}\-q\_\{\\star\}\\kappa\(c\)\\right\),\\qquad\\dot\{c\}=\\chi\(q,\\delta\)q\_\{\\star\}\\kappa^\{\\prime\}\(c\)\(1\-c^\{2\}\),\\qquad\\dot\{\\delta\}=0\.\(A\.28\)The loss dissipates according to
ℒ˙=−μ\(q,δ\)\(q2−q⋆κ\(c\)\)2−μ\(q,δ\)\+δ2q⋆2κ′\(c\)2\(1−c2\)\.\\dot\{\\mathcal\{L\}\}=\-\\mu\(q,\\delta\)\\left\(\\frac\{q\}\{2\}\-q\_\{\\star\}\\kappa\(c\)\\right\)^\{2\}\-\\frac\{\\mu\(q,\\delta\)\+\\delta\}\{2\}q\_\{\\star\}^\{2\}\\kappa^\{\\prime\}\(c\)^\{2\}\(1\-c^\{2\}\)\.\(A\.29\)
###### Proof\.
Insert \([A\.25](https://arxiv.org/html/2608.06766#A1.E25)\)–\([A\.26](https://arxiv.org/html/2608.06766#A1.E26)\) into[theorem˜A\.1](https://arxiv.org/html/2608.06766#A1.Thmtheorem1); use
‖𝖯s⟂s⋆‖2=1−c2\.\\\|\\mathsf\{P\}\_\{s\}^\{\\perp\}s\_\{\\star\}\\\|^\{2\}=1\-c^\{2\}\.∎
### A\.4Invariant region and global convergence
Define
h\(c\):=2q⋆κ\(c\)\.h\(c\):=2q\_\{\\star\}\\kappa\(c\)\.\(A\.30\)Then the coefficient equation is
q˙=−μ\(q,δ\)2\(q−h\(c\)\)\.\\dot\{q\}=\-\\frac\{\\mu\(q,\\delta\)\}\{2\}\\bigl\(q\-h\(c\)\\bigr\)\.\(A\.31\)
###### Theorem A\.9\(Global specialization\)\.
Assume
q\(0\)=q0\>0,−1<c\(0\)=c0<1\.q\(0\)=q\_\{0\}\>0,\\qquad\-1<c\(0\)=c\_\{0\}<1\.Let
q¯:=min\{q0,h\(c0\)\}\>0,q¯:=max\{q0,q⋆\}\.\\underline\{q\}:=\\min\\\{q\_\{0\},h\(c\_\{0\}\)\\\}\>0,\\qquad\\overline\{q\}:=\\max\\\{q\_\{0\},q\_\{\\star\}\\\}\.\(A\.32\)Then the solution of \([A\.28](https://arxiv.org/html/2608.06766#A1.E28)\) exists for allt≥0t\\geq 0and satisfies
q¯≤q\(t\)≤q¯,c0≤c\(t\)<1\.\\underline\{q\}\\leq q\(t\)\\leq\\overline\{q\},\\qquad c\_\{0\}\\leq c\(t\)<1\.\(A\.33\)Moreover,
c\(t\)↗1,q\(t\)⟶q⋆,ℒ\(q\(t\),c\(t\)\)⟶0\.c\(t\)\\nearrow 1,\\qquad q\(t\)\\longrightarrow q\_\{\\star\},\\qquad\\mathcal\{L\}\(q\(t\),c\(t\)\)\\longrightarrow 0\.\(A\.34\)
###### Proof\.
Sincehhis increasing and0<h\(c\)≤q⋆0<h\(c\)\\leq q\_\{\\star\}onc∈\[c0,1\]c\\in\[c\_\{0\},1\], the vector field atq=q¯q=\\underline\{q\}points inward:
q=q¯⟹q−h\(c\)≤q¯−h\(c0\)≤0⟹q˙≥0\.q=\\underline\{q\}\\Longrightarrow q\-h\(c\)\\leq\\underline\{q\}\-h\(c\_\{0\}\)\\leq 0\\Longrightarrow\\dot\{q\}\\geq 0\.Atq=q¯q=\\overline\{q\},
q=q¯⟹q−h\(c\)≥q¯−q⋆≥0⟹q˙≤0\.q=\\overline\{q\}\\Longrightarrow q\-h\(c\)\\geq\\overline\{q\}\-q\_\{\\star\}\\geq 0\\Longrightarrow\\dot\{q\}\\leq 0\.Thusq∈\[q¯,q¯\]q\\in\[\\underline\{q\},\\overline\{q\}\]\. On this compact intervalχ\(q,δ\)\>0\\chi\(q,\\delta\)\>0\. Sinceκ′\(c\)\>0\\kappa^\{\\prime\}\(c\)\>0on\(−1,1\)\(\-1,1\),ccis strictly increasing until it reaches11, andc=1c=1is invariant\.
Letc∞=limtc\(t\)c\_\{\\infty\}=\\lim\_\{t\}c\(t\)\. Ifc∞<1c\_\{\\infty\}<1, then on the compact rectangle\[q¯,q¯\]×\[c0,c∞\]\[\\underline\{q\},\\overline\{q\}\]\\times\[c\_\{0\},c\_\{\\infty\}\]the coefficient
χ\(q,δ\)q⋆κ′\(c\)\(1−c2\)\\chi\(q,\\delta\)q\_\{\\star\}\\kappa^\{\\prime\}\(c\)\(1\-c^\{2\}\)has a strictly positive lower bound, contradicting convergence ofc\(t\)c\(t\)\. Hencec∞=1c\_\{\\infty\}=1\.
Finallyh\(c\(t\)\)→q⋆h\(c\(t\)\)\\to q\_\{\\star\}\. Ifq≥q⋆\+2ηq\\geq q\_\{\\star\}\+2\\etafor large time, thenq−h\(c\)≥ηq\-h\(c\)\\geq\\etaand \([A\.31](https://arxiv.org/html/2608.06766#A1.E31)\) forces a uniform negative drift; similarlyq≤q⋆−2ηq\\leq q\_\{\\star\}\-2\\etaforces a uniform positive drift\. Thereforeq\(t\)→q⋆q\(t\)\\to q\_\{\\star\}\. Equation \([A\.27](https://arxiv.org/html/2608.06766#A1.E27)\) givesℒ→0\\mathcal\{L\}\\to 0\. ∎
### A\.5Same function, opposite gauge, quadratic time\-scale separation
FixD\>0D\>0and the same function state\(q,s\)\(q,s\)\. The two gaugesδ=±D\\delta=\\pm Dare realized by
a±2\\displaystyle a\_\{\\pm\}^\{2\}=D2\+4q2±D2,\\displaystyle=\\frac\{\\sqrt\{D^\{2\}\+4q^\{2\}\}\\pm D\}\{2\},r±2\\displaystyle r\_\{\\pm\}^\{2\}=D2\+4q2∓D2,\\displaystyle=\\frac\{\\sqrt\{D^\{2\}\+4q^\{2\}\}\\mp D\}\{2\},a±r±\\displaystyle a\_\{\\pm\}r\_\{\\pm\}=q\.\\displaystyle=q\.\(A\.35\)They have identical predictor, loss, switching interface, and signed function curvature\. At the same state\(q,c\)\(q,c\), their reaction mobility is identical,
μ\(q,\+D\)=μ\(q,−D\)=D2\+4q2,\\mu\(q,\+D\)=\\mu\(q,\-D\)=\\sqrt\{D^\{2\}\+4q^\{2\}\},\(A\.36\)while the transport mobilities are
χ\+\(q,D\)=D2\+4q2\+D2q,χ−\(q,D\)=D2\+4q2−D2q,χ\+χ−=1\.\\chi\_\{\+\}\(q,D\)=\\frac\{\\sqrt\{D^\{2\}\+4q^\{2\}\}\+D\}\{2q\},\\qquad\\chi\_\{\-\}\(q,D\)=\\frac\{\\sqrt\{D^\{2\}\+4q^\{2\}\}\-D\}\{2q\},\\qquad\\chi\_\{\+\}\\chi\_\{\-\}=1\.\(A\.37\)
For0<ε<1−c00<\\varepsilon<1\-c\_\{0\}, define the specialization hitting time
Tδ\(ε\):=inf\{t≥0:cδ\(t\)≥1−ε\}\.T\_\{\\delta\}\(\\varepsilon\):=\\inf\\\{t\\geq 0:c\_\{\\delta\}\(t\)\\geq 1\-\\varepsilon\\\}\.\(A\.38\)Let
Ic0\(ε\):=∫c01−εdc1−c2=12log\(\(2−ε\)\(1−c0\)ε\(1\+c0\)\),I\_\{c\_\{0\}\}\(\\varepsilon\):=\\int\_\{c\_\{0\}\}^\{1\-\\varepsilon\}\\frac\{\\,\\mathrm\{d\}c\}\{1\-c^\{2\}\}=\\frac\{1\}\{2\}\\log\\\!\\left\(\\frac\{\(2\-\\varepsilon\)\(1\-c\_\{0\}\)\}\{\\varepsilon\(1\+c\_\{0\}\)\}\\right\),\(A\.39\)andk0:=κ′\(c0\)\>0k\_\{0\}:=\\kappa^\{\\prime\}\(c\_\{0\}\)\>0\.
###### Theorem A\.10\(Quadratic gauge gap in specialization time\)\.
Under the assumptions of[theorem˜A\.9](https://arxiv.org/html/2608.06766#A1.Thmtheorem9), letD≥q¯D\\geq\\overline\{q\}\. For the two trajectories with the same initial\(q0,c0\)\(q\_\{0\},c\_\{0\}\)and gaugesδ=±D\\delta=\\pm D,
q¯Dq⋆Ic0\(ε\)\\displaystyle\\frac\{\\underline\{q\}\}\{Dq\_\{\\star\}\}I\_\{c\_\{0\}\}\(\\varepsilon\)≤T\+D\(ε\)≤q¯Dq⋆k0Ic0\(ε\),\\displaystyle\\leq T\_\{\+D\}\(\\varepsilon\)\\leq\\frac\{\\overline\{q\}\}\{Dq\_\{\\star\}k\_\{0\}\}I\_\{c\_\{0\}\}\(\\varepsilon\),\(A\.40\)2Dq¯q⋆Ic0\(ε\)\\displaystyle\\frac\{2D\}\{\\overline\{q\}q\_\{\\star\}\}I\_\{c\_\{0\}\}\(\\varepsilon\)≤T−D\(ε\)≤2Dq¯q⋆k0Ic0\(ε\)\.\\displaystyle\\leq T\_\{\-D\}\(\\varepsilon\)\\leq\\frac\{2D\}\{\\underline\{q\}q\_\{\\star\}k\_\{0\}\}I\_\{c\_\{0\}\}\(\\varepsilon\)\.\(A\.41\)Consequently,
2k0q¯,2D2≤T−D\(ε\)T\+D\(ε\)≤2q¯,2k0D2\.\\frac\{2k\_\{0\}\}\{\\overline\{q\}^\{,2\}\}D^\{2\}\\leq\\frac\{T\_\{\-D\}\(\\varepsilon\)\}\{T\_\{\+D\}\(\\varepsilon\)\}\\leq\\frac\{2\}\{\\underline\{q\}^\{,2\}k\_\{0\}\}D^\{2\}\.\(A\.42\)In particular, uniformly forε↓0\\varepsilon\\downarrow 0,
T\+D\(ε\)=Θ\(D−1log1ε\),T−D\(ε\)=Θ\(Dlog1ε\)\.T\_\{\+D\}\(\\varepsilon\)=\\Theta\\\!\\left\(D^\{\-1\}\\log\\frac\{1\}\{\\varepsilon\}\\right\),\\qquad T\_\{\-D\}\(\\varepsilon\)=\\Theta\\\!\\left\(D\\log\\frac\{1\}\{\\varepsilon\}\\right\)\.\(A\.43\)
###### Proof\.
Becauseccis strictly increasing,
Tδ\(ε\)=∫c01−εdcχ\(qδ\(c\),δ\)q⋆κ′\(c\)\(1−c2\)\.T\_\{\\delta\}\(\\varepsilon\)=\\int\_\{c\_\{0\}\}^\{1\-\\varepsilon\}\\frac\{\\,\\mathrm\{d\}c\}\{\\chi\(q\_\{\\delta\}\(c\),\\delta\)q\_\{\\star\}\\kappa^\{\\prime\}\(c\)\(1\-c^\{2\}\)\}\.\(A\.44\)By[theorem˜A\.9](https://arxiv.org/html/2608.06766#A1.Thmtheorem9),q¯≤qδ\(c\)≤q¯\\underline\{q\}\\leq q\_\{\\delta\}\(c\)\\leq\\overline\{q\}\. ForD≥q¯D\\geq\\overline\{q\},
Dq¯≤χ\+\(q,D\)≤2Dq¯,\\frac\{D\}\{\\overline\{q\}\}\\leq\\chi\_\{\+\}\(q,D\)\\leq\\frac\{2D\}\{\\underline\{q\}\},\(A\.45\)where the upper bound uses
D2\+4q2≤D\+2q2D\.\\sqrt\{D^\{2\}\+4q^\{2\}\}\\leq D\+\\frac\{2q^\{2\}\}\{D\}\.Similarly,
q¯2D≤χ−\(q,D\)≤q¯D,\\frac\{\\underline\{q\}\}\{2D\}\\leq\\chi\_\{\-\}\(q,D\)\\leq\\frac\{\\overline\{q\}\}\{D\},\(A\.46\)usingχ−=2q/\(D2\+4q2\+D\)\\chi\_\{\-\}=2q/\(\\sqrt\{D^\{2\}\+4q^\{2\}\}\+D\)\. Finally,
k0≤κ′\(c\)≤12\(c0≤c≤1\)\.k\_\{0\}\\leq\\kappa^\{\\prime\}\(c\)\\leq\\frac\{1\}\{2\}\\qquad\(c\_\{0\}\\leq c\\leq 1\)\.Insert these bounds into \([A\.44](https://arxiv.org/html/2608.06766#A1.E44)\) to obtain \([A\.40](https://arxiv.org/html/2608.06766#A1.E40)\)–\([A\.41](https://arxiv.org/html/2608.06766#A1.E41)\)\. Taking ratios gives \([A\.42](https://arxiv.org/html/2608.06766#A1.E42)\); \([A\.39](https://arxiv.org/html/2608.06766#A1.E39)\) is asymptotic to12log\(1/ε\)\\frac\{1\}\{2\}\\log\(1/\\varepsilon\)\. ∎
###### Corollary A\.11\(The two gauges do not differ by a global clock\)\.
At every state with
q2−q⋆κ\(c\)≠0,−1<c<1,\\frac\{q\}\{2\}\-q\_\{\\star\}\\kappa\(c\)\\neq 0,\\qquad\-1<c<1,the vector fields forδ=\+D\\delta=\+Dandδ=−D\\delta=\-Dare not collinear\. Therefore their phase\-plane trajectories cannot be matched by a scalar learning\-rate or global time rescaling\.
###### Proof\.
At a common state, theqq\-components of the two vector fields are equal by \([A\.36](https://arxiv.org/html/2608.06766#A1.E36)\) and nonzero by assumption\. Any collinearity factor must therefore equal one\. But thecc\-components differ becauseχ\+χ−=1\\chi\_\{\+\}\\chi\_\{\-\}=1and, forD\>0D\>0,χ\+≠χ−\\chi\_\{\+\}\\neq\\chi\_\{\-\}\. Hence the vector fields are not collinear\. ∎
### A\.6Fast and slow limiting systems
The hitting\-time theorem is nonasymptotic\. The following limits explain the mechanism more finely\.
###### Proposition A\.12\(Transport\-enabled fast limit forδ=\+D\\delta=\+D\)\.
Let
QD\+\(τ\)=q\+D\(τ/D\),CD\+\(τ\)=c\+D\(τ/D\)\.Q\_\{D\}^\{\+\}\(\\tau\)=q\_\{\+D\}\(\\tau/D\),\\qquad C\_\{D\}^\{\+\}\(\\tau\)=c\_\{\+D\}\(\\tau/D\)\.For every finiteτmax\\tau\_\{\\mathrm\{max\}\},
sup0≤τ≤τmax‖\(QD\+\(τ\),CD\+\(τ\)\)−\(Q\+\(τ\),C\+\(τ\)\)‖≤CτmaxD2,\\sup\_\{0\\leq\\tau\\leq\\tau\_\{\\mathrm\{max\}\}\}\\bigl\\\|\(Q\_\{D\}^\{\+\}\(\\tau\),C\_\{D\}^\{\+\}\(\\tau\)\)\-\(Q\_\{\+\}\(\\tau\),C\_\{\+\}\(\\tau\)\)\\bigr\\\|\\leq\\frac\{C\_\{\\tau\_\{\\mathrm\{max\}\}\}\}\{D^\{2\}\},\(A\.47\)where
dQ\+dτ=−\(Q\+2−q⋆κ\(C\+\)\),dC\+dτ=q⋆Q\+κ′\(C\+\)\(1−C\+2\)\.\\frac\{\\,\\mathrm\{d\}Q\_\{\+\}\}\{\\,\\mathrm\{d\}\\tau\}=\-\\left\(\\frac\{Q\_\{\+\}\}\{2\}\-q\_\{\\star\}\\kappa\(C\_\{\+\}\)\\right\),\\qquad\\frac\{\\,\\mathrm\{d\}C\_\{\+\}\}\{\\,\\mathrm\{d\}\\tau\}=\\frac\{q\_\{\\star\}\}\{Q\_\{\+\}\}\\kappa^\{\\prime\}\(C\_\{\+\}\)\(1\-C\_\{\+\}^\{2\}\)\.\(A\.48\)
###### Proof\.
On the invariant rectangle\[q¯,q¯\]×\[c0,1\]\[\\underline\{q\},\\overline\{q\}\]\\times\[c\_\{0\},1\],
μ\(q,\+D\)D=1\+O\(D−2\),χ\(q,\+D\)D=1q\+O\(D−2\)\\frac\{\\mu\(q,\+D\)\}\{D\}=1\+O\(D^\{\-2\}\),\\qquad\\frac\{\\chi\(q,\+D\)\}\{D\}=\\frac\{1\}\{q\}\+O\(D^\{\-2\}\)uniformly with uniformly Lipschitz vector fields\. Standard continuous dependence and Gronwall’s inequality give \([A\.47](https://arxiv.org/html/2608.06766#A1.E47)\)\. ∎
###### Proposition A\.13\(Fast reaction layer and slow transport forδ=−D\\delta=\-D\)\.
Leth\(c\)=2q⋆κ\(c\)h\(c\)=2q\_\{\\star\}\\kappa\(c\)andeD\(t\)=q−D\(t\)−h\(c−D\(t\)\)e\_\{D\}\(t\)=q\_\{\-D\}\(t\)\-h\(c\_\{\-D\}\(t\)\)\. There is a constantCCindependent ofDDsuch that
\|eD\(t\)\|≤\|eD\(0\)\|e−Dt/2\+CD2\(t≥0\)\.\|e\_\{D\}\(t\)\|\\leq\|e\_\{D\}\(0\)\|e^\{\-Dt/2\}\+\\frac\{C\}\{D^\{2\}\}\\qquad\(t\\geq 0\)\.\(A\.49\)On the fast scaleτ=Dt\\tau=Dt,\(q−D\(τ/D\),c−D\(τ/D\)\)\(q\_\{\-D\}\(\\tau/D\),c\_\{\-D\}\(\\tau/D\)\)converges on compact intervals to
dQfdτ=−\(Qf2−q⋆κ\(c0\)\),dCfdτ=0,\\frac\{\\,\\mathrm\{d\}Q\_\{f\}\}\{\\,\\mathrm\{d\}\\tau\}=\-\\left\(\\frac\{Q\_\{f\}\}\{2\}\-q\_\{\\star\}\\kappa\(c\_\{0\}\)\\right\),\\qquad\\frac\{\\,\\mathrm\{d\}C\_\{f\}\}\{\\,\\mathrm\{d\}\\tau\}=0,\(A\.50\)so
Qf\(τ\)=h\(c0\)\+\(q0−h\(c0\)\)e−τ/2,Cf\(τ\)=c0\.Q\_\{f\}\(\\tau\)=h\(c\_\{0\}\)\+\\bigl\(q\_\{0\}\-h\(c\_\{0\}\)\\bigr\)e^\{\-\\tau/2\},\\qquad C\_\{f\}\(\\tau\)=c\_\{0\}\.\(A\.51\)On the slow scaleσ=t/D\\sigma=t/D,CD−\(σ\):=c−D\(Dσ\)C\_\{D\}^\{\-\}\(\\sigma\):=c\_\{\-D\}\(D\\sigma\)converges uniformly on compact intervals to the solution of
dC−dσ=2q⋆2κ\(C−\)κ′\(C−\)\(1−C−2\),C−\(0\)=c0\.\\frac\{\\,\\mathrm\{d\}C\_\{\-\}\}\{\\,\\mathrm\{d\}\\sigma\}=2q\_\{\\star\}^\{2\}\\kappa\(C\_\{\-\}\)\\kappa^\{\\prime\}\(C\_\{\-\}\)\(1\-C\_\{\-\}^\{2\}\),\\qquad C\_\{\-\}\(0\)=c\_\{0\}\.\(A\.52\)For everyσ0\>0\\sigma\_\{0\}\>0,
supσ0≤σ≤σmax\|q−D\(Dσ\)−2q⋆κ\(C−\(σ\)\)\|=O\(D−2\)\.\\sup\_\{\\sigma\_\{0\}\\leq\\sigma\\leq\\sigma\_\{\\mathrm\{max\}\}\}\\bigl\|q\_\{\-D\}\(D\\sigma\)\-2q\_\{\\star\}\\kappa\(C\_\{\-\}\(\\sigma\)\)\\bigr\|=O\(D^\{\-2\}\)\.\(A\.53\)
###### Proof\.
SetA\(c\)=q⋆κ′\(c\)\(1−c2\)A\(c\)=q\_\{\\star\}\\kappa^\{\\prime\}\(c\)\(1\-c^\{2\}\)\. Sinceq∈\[q¯,q¯\]q\\in\[\\underline\{q\},\\overline\{q\}\],
\|h′\(c\)\|≤q⋆,0≤A\(c\)≤q⋆2,0≤χ\(q,−D\)≤q¯D\.\|h^\{\\prime\}\(c\)\|\\leq q\_\{\\star\},\\qquad 0\\leq A\(c\)\\leq\\frac\{q\_\{\\star\}\}\{2\},\\qquad 0\\leq\\chi\(q,\-D\)\\leq\\frac\{\\overline\{q\}\}\{D\}\.DifferentiatingeD=q−h\(c\)e\_\{D\}=q\-h\(c\)gives
e˙D=−μ\(q,−D\)2eD−h′\(c\)χ\(q,−D\)A\(c\)\.\\dot\{e\}\_\{D\}=\-\\frac\{\\mu\(q,\-D\)\}\{2\}e\_\{D\}\-h^\{\\prime\}\(c\)\\chi\(q,\-D\)A\(c\)\.Sinceμ≥D\\mu\\geq D, variation of constants yields \([A\.49](https://arxiv.org/html/2608.06766#A1.E49)\)\. The fast\-scale limit follows from
μ\(q,−D\)D=1\+O\(D−2\),χ\(q,−D\)D=O\(D−2\)\.\\frac\{\\mu\(q,\-D\)\}\{D\}=1\+O\(D^\{\-2\}\),\\qquad\\frac\{\\chi\(q,\-D\)\}\{D\}=O\(D^\{\-2\}\)\.
For the slow scale, the exact identity
\|Dχ\(q,−D\)−q\|=4\|q\|3\(D2\+4q2\+D\)2≤q¯,3D2\\left\|D\\chi\(q,\-D\)\-q\\right\|=\\frac\{4\|q\|^\{3\}\}\{\\bigl\(\\sqrt\{D^\{2\}\+4q^\{2\}\}\+D\\bigr\)^\{2\}\}\\leq\\frac\{\\overline\{q\}^\{,3\}\}\{D^\{2\}\}\(A\.54\)shows
ddσCD−\(σ\)=q−D\(Dσ\)A\(CD−\(σ\)\)\+O\(D−2\)\.\\frac\{\\,\\mathrm\{d\}\}\{\\,\\mathrm\{d\}\\sigma\}C\_\{D\}^\{\-\}\(\\sigma\)=q\_\{\-D\}\(D\\sigma\)A\(C\_\{D\}^\{\-\}\(\\sigma\)\)\+O\(D^\{\-2\}\)\.The transient term in \([A\.49](https://arxiv.org/html/2608.06766#A1.E49)\) has slow\-time integralO\(D−2\)O\(D^\{\-2\}\), while the remaining tracking error isO\(D−2\)O\(D^\{\-2\}\)\. Gronwall’s inequality therefore gives uniform convergence ofCD−C\_\{D\}^\{\-\}to \([A\.52](https://arxiv.org/html/2608.06766#A1.E52)\); the coefficient statement follows from \([A\.49](https://arxiv.org/html/2608.06766#A1.E49)\)\. ∎
## Appendix BGlobal ownership under exact duplication
LetX∼𝒩\(0,Id\)X\\sim\\mathcal\{N\}\(0,I\_\{d\}\)and let the teacher be
f⋆\(x\)=σ\(u⊤x\),‖u‖=1\.f\_\{\\star\}\(x\)=\\sigma\(u^\{\\top\}x\),\\qquad\\\|u\\\|=1\.The student network is
f\(x\)=∑i=1mqi\(t\)σ\(si\(t\)⊤x\),‖si\(t\)‖=1\.f\(x\)=\\sum\_\{i=1\}^\{m\}q\_\{i\}\(t\)\\sigma\(s\_\{i\}\(t\)^\{\\top\}x\),\\qquad\\\|s\_\{i\}\(t\)\\\|=1\.We use population square loss
ℒ=12𝔼\[\(f\(X\)−f⋆\(X\)\)2\]\.\\mathcal\{L\}=\\frac\{1\}\{2\}\\mathbb\{E\}\\bigl\[\(f\(X\)\-f\_\{\\star\}\(X\)\)^\{2\}\\bigr\]\.All students begin from the same visible state,
qi\(0\)=1m,si\(0\)=s0,∠\(s0,u\)=ϕ:=π4\.q\_\{i\}\(0\)=\\frac\{1\}\{m\},\\qquad s\_\{i\}\(0\)=s\_\{0\},\\qquad\\angle\(s\_\{0\},u\)=\\phi:=\\frac\{\\pi\}\{4\}\.\(B\.1\)Choose one indexkkand assign
δk=\+D,δj=−D\(j≠k\),\\delta\_\{k\}=\+D,\\qquad\\delta\_\{j\}=\-D\\quad\(j\\neq k\),whereδi=ai2−‖wi‖2\\delta\_\{i\}=a\_\{i\}^\{2\}\-\\\|w\_\{i\}\\\|^\{2\}is the conserved gauge mark andqi=ai‖wi‖q\_\{i\}=a\_\{i\}\\\|w\_\{i\}\\\|\. The underlying positive factors are uniquely reconstructed from
ai2=D2\+4qi2\+δi2,‖wi‖2=D2\+4qi2−δi2\.a\_\{i\}^\{2\}=\\frac\{\\sqrt\{D^\{2\}\+4q\_\{i\}^\{2\}\}\+\\delta\_\{i\}\}\{2\},\\qquad\\\|w\_\{i\}\\\|^\{2\}=\\frac\{\\sqrt\{D^\{2\}\+4q\_\{i\}^\{2\}\}\-\\delta\_\{i\}\}\{2\}\.Thus every choice of the selected index gives the same initial function and loss\.
Set
By permutation symmetry, theMMnegative\-gauge students remain identical\. Write
q=qk,p=qj\(j≠k\),P:=Mp,q=q\_\{k\},\\qquad p=q\_\{j\}\\ \(j\\neq k\),\\qquad P:=Mp,and letθ\\thetaandψ\\psibe the signed angles of the winner and the common loser direction relative touuin the invariant planespan\{u,s0\}\\operatorname\{span\}\\\{u,s\_\{0\}\\\}\. The population function is exactly
f\(x\)=qσ\(sθ⊤x\)\+Pσ\(sψ⊤x\)\.f\(x\)=q\\sigma\(s\_\{\\theta\}^\{\\top\}x\)\+P\\sigma\(s\_\{\\psi\}^\{\\top\}x\)\.
Define the angular ReLU kernel
g\(α\)=sinα\+\(π−α\)cosα2π,ω\(α\):=−g′\(α\)=\(π−α\)sinα2π\.g\(\\alpha\)=\\frac\{\\sin\\alpha\+\(\\pi\-\\alpha\)\\cos\\alpha\}\{2\\pi\},\\qquad\\omega\(\\alpha\):=\-g^\{\\prime\}\(\\alpha\)=\\frac\{\(\\pi\-\\alpha\)\\sin\\alpha\}\{2\\pi\}\.This is the angular form of the kernel in Appendix A:g\(α\)=κ\(cosα\)g\(\\alpha\)=\\kappa\(\\cos\\alpha\)forα∈\[0,π\]\\alpha\\in\[0,\\pi\]\. The reduced loss is
ℰ\(q,P,θ,ψ\)=14\(q2\+P2\+1\)\+qPg\(\|ψ−θ\|\)−qg\(\|θ\|\)−Pg\(\|ψ\|\)\.\\mathscr\{E\}\(q,P,\\theta,\\psi\)=\\frac\{1\}\{4\}\(q^\{2\}\+P^\{2\}\+1\)\+qPg\(\|\\psi\-\\theta\|\)\-qg\(\|\\theta\|\)\-Pg\(\|\\psi\|\)\.\(B\.2\)This identity already absorbs all\(M2\)\\binom\{M\}\{2\}loser–loser interactions: because their directions coincide, their self and pair terms sum toP2/4P^\{2\}/4\.
Let
μ\+\(q\)=D2\+4q2,μ−\(p\)=D2\+4p2,\\mu\_\{\+\}\(q\)=\\sqrt\{D^\{2\}\+4q^\{2\}\},\\qquad\\mu\_\{\-\}\(p\)=\\sqrt\{D^\{2\}\+4p^\{2\}\},χ\+\(q\)=μ\+\(q\)\+D2q,χ−\(p\)=2pμ−\(p\)\+D\.\\chi\_\{\+\}\(q\)=\\frac\{\\mu\_\{\+\}\(q\)\+D\}\{2q\},\\qquad\\chi\_\{\-\}\(p\)=\\frac\{2p\}\{\\mu\_\{\-\}\(p\)\+D\}\.
###### Proposition B\.1\(Exact multiplicity\-reduced marked flow\)\.
On the duplicate\-loser manifold, ordinary population gradient flow is exactly
q˙\\displaystyle\\dot\{q\}=−μ\+\(q\)∂qℰ,\\displaystyle=\-\\mu\_\{\+\}\(q\)\\,\\partial\_\{q\}\\mathscr\{E\},P˙\\displaystyle\\dot\{P\}=−Mμ−\(P/M\)∂Pℰ,\\displaystyle=\-M\\mu\_\{\-\}\(P/M\)\\,\\partial\_\{P\}\\mathscr\{E\},\(B\.3\)θ˙\\displaystyle\\dot\{\\theta\}=−χ\+\(q\)q∂θℰ,\\displaystyle=\-\\frac\{\\chi\_\{\+\}\(q\)\}\{q\}\\,\\partial\_\{\\theta\}\\mathscr\{E\},ψ˙\\displaystyle\\dot\{\\psi\}=−χ−\(P/M\)P∂ψℰ,\\displaystyle=\-\\frac\{\\chi\_\{\-\}\(P/M\)\}\{P\}\\,\\partial\_\{\\psi\}\\mathscr\{E\},\(B\.4\)where the last expression has its continuous extension atP=0P=0\. Moreover,
ddtℰ=\\displaystyle\\frac\{\\,\\mathrm\{d\}\}\{\\,\\mathrm\{d\}t\}\\mathscr\{E\}=\{\}−μ\+\(q\)\(∂qℰ\)2−Mμ−\(P/M\)\(∂Pℰ\)2\\displaystyle\-\\mu\_\{\+\}\(q\)\(\\partial\_\{q\}\\mathscr\{E\}\)^\{2\}\-M\\mu\_\{\-\}\(P/M\)\(\\partial\_\{P\}\\mathscr\{E\}\)^\{2\}−χ\+\(q\)q\(∂θℰ\)2−χ−\(P/M\)P\(∂ψℰ\)2\.\\displaystyle\-\\frac\{\\chi\_\{\+\}\(q\)\}\{q\}\(\\partial\_\{\\theta\}\\mathscr\{E\}\)^\{2\}\-\\frac\{\\chi\_\{\-\}\(P/M\)\}\{P\}\(\\partial\_\{\\psi\}\\mathscr\{E\}\)^\{2\}\.\(B\.5\)
###### Proof\.
For one loser, the coefficient gradient is
∂pjℒ=P2\+qg\(\|ψ−θ\|\)−g\(\|ψ\|\)=∂Pℰ\.\\partial\_\{p\_\{j\}\}\\mathcal\{L\}=\\frac\{P\}\{2\}\+qg\(\|\\psi\-\\theta\|\)\-g\(\|\\psi\|\)=\\partial\_\{P\}\\mathscr\{E\}\.ThereforeP˙=Mp˙\\dot\{P\}=M\\dot\{p\}gives the second equation in \([B\.3](https://arxiv.org/html/2608.06766#A2.E3)\)\. The common angular derivative satisfies
∂ψℰ=∑j≠k∂ψjℒ,\\partial\_\{\\psi\}\\mathscr\{E\}=\\sum\_\{j\\neq k\}\\partial\_\{\\psi\_\{j\}\}\\mathcal\{L\},and each summand is1/M1/Mof the total\. Dividing the individual angular gradient by the individual functional coefficientp=P/Mp=P/Mgives the final equation in \([B\.4](https://arxiv.org/html/2608.06766#A2.E4)\)\. The selected\-unit equations follow from the exact marked reduction in[theorem˜A\.1](https://arxiv.org/html/2608.06766#A1.Thmtheorem1)\. Pairing the four state velocities with the four derivatives of \([B\.2](https://arxiv.org/html/2608.06766#A2.E2)\) yields \([B\.5](https://arxiv.org/html/2608.06766#A2.E5)\)\. ∎
### B\.1The multiplicity\-weighted fast capture flow
Set
For every fixedMMand on compact sets withq\>0q\>0,
μ\+\(q\)D=1\+O\(D−2\),μ−\(P/M\)D=1\+O\(D−2\),\\frac\{\\mu\_\{\+\}\(q\)\}\{D\}=1\+O\(D^\{\-2\}\),\\qquad\\frac\{\\mu\_\{\-\}\(P/M\)\}\{D\}=1\+O\(D^\{\-2\}\),χ\+\(q\)Dq=1q2\+O\(D−2\),χ−\(P/M\)DP=O\(D−2\)\.\\frac\{\\chi\_\{\+\}\(q\)\}\{Dq\}=\\frac\{1\}\{q^\{2\}\}\+O\(D^\{\-2\}\),\\qquad\\frac\{\\chi\_\{\-\}\(P/M\)\}\{DP\}=O\(D^\{\-2\}\)\.Hence the common loser direction freezes at leading order:
ψ\(τ\)≡ϕ\.\\psi\(\\tau\)\\equiv\\phi\.Define
EM\(q,P,θ\):=14\(q2\+P2\+1\)\+qPg\(ϕ−θ\)−qg\(\|θ\|\)−Pg\(ϕ\),E\_\{M\}\(q,P,\\theta\):=\\frac\{1\}\{4\}\(q^\{2\}\+P^\{2\}\+1\)\+qPg\(\\phi\-\\theta\)\-qg\(\|\\theta\|\)\-Pg\(\\phi\),\(B\.6\)for−ϕ<θ<ϕ\-\\phi<\\theta<\\phi\. The fast system is
q′=−∂qEM,P′=−M∂PEM,θ′=−q−2∂θEM\.q^\{\\prime\}=\-\\partial\_\{q\}E\_\{M\},\\qquad P^\{\\prime\}=\-M\\partial\_\{P\}E\_\{M\},\\qquad\\theta^\{\\prime\}=\-q^\{\-2\}\\partial\_\{\\theta\}E\_\{M\}\.\(B\.7\)with initial condition
q\(0\)=q0:=1M\+1,P\(0\)=P0:=MM\+1,θ\(0\)=ϕ\.q\(0\)=q\_\{0\}:=\\frac\{1\}\{M\+1\},\\qquad P\(0\)=P\_\{0\}:=\\frac\{M\}\{M\+1\},\\qquad\\theta\(0\)=\\phi\.\(B\.8\)The subscriptMMrecords the metric weight, not a change in the energy landscape\. Indeed,
ddτEM=−\(∂qEM\)2−M\(∂PEM\)2−q−2\(∂θEM\)2\.\\frac\{\\,\\mathrm\{d\}\}\{\\,\\mathrm\{d\}\\tau\}E\_\{M\}=\-\(\\partial\_\{q\}E\_\{M\}\)^\{2\}\-M\(\\partial\_\{P\}E\_\{M\}\)^\{2\}\-q^\{\-2\}\(\\partial\_\{\\theta\}E\_\{M\}\)^\{2\}\.\(B\.9\)
For later use define
d:=ϕ−θ,Z:=qg\(d\)−g\(ϕ\),W:=Pg\(d\)−g\(\|θ\|\)\.d:=\\phi\-\\theta,\\qquad Z:=qg\(d\)\-g\(\\phi\),\\qquad W:=Pg\(d\)\-g\(\|\\theta\|\)\.The coefficient equations become
q′=−\(q2\+W\),P′=−M\(P2\+Z\)\.q^\{\\prime\}=\-\\left\(\\frac\{q\}\{2\}\+W\\right\),\\qquad P^\{\\prime\}=\-M\\left\(\\frac\{P\}\{2\}\+Z\\right\)\.\(B\.10\)
### The scalar arc\-cosine inequalities
The proof uses the same four strict arc\-cosine inequalities as the two\-student basin theorem\. For0≤x≤ϕ0\\leq x\\leq\\phi,
g\(x\)g\(ϕ−x\)−ω\(x\)ω\(ϕ−x\)\\displaystyle g\(x\)g\(\\phi\-x\)\-\\omega\(x\)\\omega\(\\phi\-x\)≤12g\(ϕ\),\\displaystyle\\leq\\tfrac\{1\}\{2\}g\(\\phi\),\(B\.11\)2g\(ϕ\)g\(ϕ−x\)\\displaystyle 2g\(\\phi\)g\(\\phi\-x\)≤g\(x\),\\displaystyle\\leq g\(x\),\(B\.12\)g\(x\)g\(ϕ\+x\)\+ω\(x\)ω\(ϕ\+x\)\\displaystyle g\(x\)g\(\\phi\+x\)\+\\omega\(x\)\\omega\(\\phi\+x\)≤12g\(ϕ\),\\displaystyle\\leq\\tfrac\{1\}\{2\}g\(\\phi\),\(B\.13\)2g\(ϕ\)g\(ϕ\+x\)\\displaystyle 2g\(\\phi\)g\(\\phi\+x\)<g\(x\)\(x\>0\)\.\\displaystyle<g\(x\)\\qquad\(x\>0\)\.\(B\.14\)The first three follow from explicit derivative factorizations and endpoint values; the strict negative\-angle critical\-point exclusion uses the one\-dimensional interval certificate reproduced in Appendix[B\.5](https://arxiv.org/html/2608.06766#A2.SS5)\. Importantly, none of these inequalities depends onMM\.
### B\.2Finite entrance for arbitrary multiplicity
ForM≤3M\\leq 3, the initial point already satisfiesW\(0\)≤0W\(0\)\\leq 0\. For larger multiplicities,
W\(0\)=M2\(M\+1\)−g\(ϕ\)\>0,W\(0\)=\\frac\{M\}\{2\(M\+1\)\}\-g\(\\phi\)\>0,so the two\-student invariant sector cannot be invoked directly\. The next lemma supplies the finite entrance argument needed for arbitrary fixed multiplicity\.
###### Lemma B\.2\(Finite entrance into the canonical basin\)\.
Let\(q,P,θ\)\(q,P,\\theta\)solve \([B\.7](https://arxiv.org/html/2608.06766#A2.E7)\)–\([B\.8](https://arxiv.org/html/2608.06766#A2.E8)\)\. Put
A:=12−g\(ϕ\),ω0:=ω\(ϕ/2\),γent:=A22ω02\.A:=\\frac\{1\}\{2\}\-g\(\\phi\),\\qquad\\omega\_\{0\}:=\\omega\(\\phi/2\),\\qquad\\gamma\_\{\\rm ent\}:=\\frac\{A^\{2\}\}\{2\\omega\_\{0\}^\{2\}\}\.Forϕ=π/4\\phi=\\pi/4,
0<γent<0\.267\.0<\\gamma\_\{\\rm ent\}<0\.267\.There is a finite timeτ∗≥0\\tau\_\{\\ast\}\\geq 0such that
W\(τ∗\)=0,Z\(τ∗\)≤0,ϕ/2≤θ\(τ∗\)<ϕ,W\(\\tau\_\{\\ast\}\)=0,\\qquad Z\(\\tau\_\{\\ast\}\)\\leq 0,\\qquad\\phi/2\\leq\\theta\(\\tau\_\{\\ast\}\)<\\phi,and
τ∗≤q0\[W\(0\)\]\+ω02,q\(τ∗\)≥\(1−γent\)q0\.\\tau\_\{\\ast\}\\leq\\frac\{q\_\{0\}\[W\(0\)\]\_\{\+\}\}\{\\omega\_\{0\}^\{2\}\},\\qquad q\(\\tau\_\{\\ast\}\)\\geq\(1\-\\gamma\_\{\\rm ent\}\)q\_\{0\}\.\(B\.15\)Immediately afterτ∗\\tau\_\{\\ast\}, one hasW<0W<0\.
###### Proof\.
IfW\(0\)≤0W\(0\)\\leq 0, setτ∗=0\\tau\_\{\\ast\}=0\. AssumeW\(0\)\>0W\(0\)\>0and work until the first zero ofWW\.
First,PPcannot reach zero whileW\>0W\>0, and atP=1P=1the vector field points inward because
P2\+Z≥12−g\(ϕ\)\>0\.\\frac\{P\}\{2\}\+Z\\geq\\frac\{1\}\{2\}\-g\(\\phi\)\>0\.Thus0<P≤P0<10<P\\leq P\_\{0\}<1\. Similarlyq<1q<1\. At a hypothetical first point withZ=0Z=0and0≤θ≤ϕ0\\leq\\theta\\leq\\phi, \([B\.11](https://arxiv.org/html/2608.06766#A2.E11)\) gives
Z′\\displaystyle Z^\{\\prime\}=−g\(ϕ\)2−P\{g\(d\)2\+ω\(d\)2\}\+g\(θ\)g\(d\)−ω\(θ\)ω\(d\)\\displaystyle=\-\\frac\{g\(\\phi\)\}\{2\}\-P\\\{g\(d\)^\{2\}\+\\omega\(d\)^\{2\}\\\}\+g\(\\theta\)g\(d\)\-\\omega\(\\theta\)\\omega\(d\)≤0\.\\displaystyle\\leq 0\.HenceZ≤0Z\\leq 0throughout this entrance phase\.
BecauseW\>0W\>0andP≤1P\\leq 1,
g\(d\)\>g\(θ\)\.g\(d\)\>g\(\\theta\)\.The functionggis strictly decreasing, sod<θd<\\theta, equivalently
Furthermore, \([B\.12](https://arxiv.org/html/2608.06766#A2.E12)\) implies
P2\>g\(θ\)2g\(d\)≥g\(ϕ\)\.\\frac\{P\}\{2\}\>\\frac\{g\(\\theta\)\}\{2g\(d\)\}\\geq g\(\\phi\)\.Consequently
∂PEM=P2\+Z\>qg\(d\)≥0\.\\partial\_\{P\}E\_\{M\}=\\frac\{P\}\{2\}\+Z\>qg\(d\)\\geq 0\.For0≤θ≤ϕ0\\leq\\theta\\leq\\phi, set
S:=Pω\(d\)\+ω\(θ\)\.S:=P\\omega\(d\)\+\\omega\(\\theta\)\.Direct differentiation gives
W′=−M\(P2\+Z\)g\(d\)−S2q≤−ω02q\.W^\{\\prime\}=\-M\\left\(\\frac\{P\}\{2\}\+Z\\right\)g\(d\)\-\\frac\{S^\{2\}\}\{q\}\\leq\-\\frac\{\\omega\_\{0\}^\{2\}\}\{q\}\.\(B\.16\)ThusWWis strictly decreasing\. Sinceq′=−\(q/2\+W\)<0q^\{\\prime\}=\-\(q/2\+W\)<0, one hasq≤q0q\\leq q\_\{0\}, and \([B\.16](https://arxiv.org/html/2608.06766#A2.E16)\) yields
τ∗≤q0W\(0\)ω02\.\\tau\_\{\\ast\}\\leq\\frac\{q\_\{0\}W\(0\)\}\{\\omega\_\{0\}^\{2\}\}\.Using−W′≥ω02/q\-W^\{\\prime\}\\geq\\omega\_\{0\}^\{2\}/qas a change\-of\-variables estimate,
q0−q\(τ∗\)\\displaystyle q\_\{0\}\-q\(\\tau\_\{\\ast\}\)=∫0τ∗\(q2\+W\)dτ\\displaystyle=\\int\_\{0\}^\{\\tau\_\{\\ast\}\}\\left\(\\frac\{q\}\{2\}\+W\\right\)\\,\\mathrm\{d\}\\tau≤1ω02∫0W\(0\)\(q2\+W\)qdW\\displaystyle\\leq\\frac\{1\}\{\\omega\_\{0\}^\{2\}\}\\int\_\{0\}^\{W\(0\)\}\\left\(\\frac\{q\}\{2\}\+W\\right\)q\\,\\,\\mathrm\{d\}W≤q02ω02\(q0W\(0\)\+W\(0\)2\)\.\\displaystyle\\leq\\frac\{q\_\{0\}\}\{2\\omega\_\{0\}^\{2\}\}\\left\(q\_\{0\}W\(0\)\+W\(0\)^\{2\}\\right\)\.Now
W\(0\)=A−q02,W\(0\)=A\-\\frac\{q\_\{0\}\}\{2\},so
q0W\(0\)\+W\(0\)2=A2−q024≤A2\.q\_\{0\}W\(0\)\+W\(0\)^\{2\}=A^\{2\}\-\\frac\{q\_\{0\}^\{2\}\}\{4\}\\leq A^\{2\}\.This proves the lower bound in \([B\.15](https://arxiv.org/html/2608.06766#A2.E15)\)\. Finally, \([B\.16](https://arxiv.org/html/2608.06766#A2.E16)\) is strict atW=0W=0, so the trajectory crosses intoW<0W<0\. ∎
### B\.3A multiplicity\-independent global basin
Define the canonical sector
ℬ:=\{0<q≤1,0≤P≤1,−ϕ<θ<ϕ,Z=qg\(ϕ−θ\)−g\(ϕ\)≤0,W=Pg\(ϕ−θ\)−g\(\|θ\|\)≤0\}\.\\mathcal\{B\}:=\\left\\\{\\begin\{array\}\[\]\{c\}0<q\\leq 1,\\quad 0\\leq P\\leq 1,\\quad\-\\phi<\\theta<\\phi,\\\\ Z=qg\(\\phi\-\\theta\)\-g\(\\phi\)\\leq 0,\\\\ W=Pg\(\\phi\-\\theta\)\-g\(\|\\theta\|\)\\leq 0\\end\{array\}\\right\\\}\.\(B\.17\)
###### Lemma B\.3\(Weighted arc\-cosine basin\)\.
For everyM≥1M\\geq 1, the sectorℬ\\mathcal\{B\}is forward invariant under \([B\.7](https://arxiv.org/html/2608.06766#A2.E7)\)\. Every trajectory inℬ\\mathcal\{B\}converges to
\(q,P,θ\)=\(1,0,0\)\.\(q,P,\\theta\)=\(1,0,0\)\.
###### Proof\.
Only thePP\-equation differs from the two\-student flow, and it is multiplied by the positive factorMM\. We nevertheless record the sign argument\.
Atq=1q=1,
∂qEM=12\+W≥12−g\(\|θ\|\)≥0,\\partial\_\{q\}E\_\{M\}=\\frac\{1\}\{2\}\+W\\geq\\frac\{1\}\{2\}\-g\(\|\\theta\|\)\\geq 0,so the vector field points inward\. SinceW≤0W\\leq 0,
q′=−q2−W≥−q2,q^\{\\prime\}=\-\\frac\{q\}\{2\}\-W\\geq\-\\frac\{q\}\{2\},and thereforeqqremains positive for every finite time\. AtP=0P=0,P′=−MZ≥0P^\{\\prime\}=\-MZ\\geq 0; atP=1P=1,
∂PEM=12\+Z≥12−g\(ϕ\)\>0\.\\partial\_\{P\}E\_\{M\}=\\frac\{1\}\{2\}\+Z\\geq\\frac\{1\}\{2\}\-g\(\\phi\)\>0\.The angular faces are inward\-pointing: atθ=ϕ\\theta=\\phi,θ′=−ω\(ϕ\)/q<0\\theta^\{\\prime\}=\-\\omega\(\\phi\)/q<0, while atθ=−ϕ\\theta=\-\\phi,
θ′≥ω\(ϕ\)−ω\(2ϕ\)q\>0\.\\theta^\{\\prime\}\\geq\\frac\{\\omega\(\\phi\)\-\\omega\(2\\phi\)\}\{q\}\>0\.
AtZ=0Z=0, \([B\.11](https://arxiv.org/html/2608.06766#A2.E11)\) forθ≥0\\theta\\geq 0and \([B\.13](https://arxiv.org/html/2608.06766#A2.E13)\) forθ<0\\theta<0implyZ′≤0Z^\{\\prime\}\\leq 0\. AtW=0W=0, \([B\.12](https://arxiv.org/html/2608.06766#A2.E12)\) or \([B\.14](https://arxiv.org/html/2608.06766#A2.E14)\) implies
∂PEM=P2\+Z≥0,\\partial\_\{P\}E\_\{M\}=\\frac\{P\}\{2\}\+Z\\geq 0,and the exact identity
W′=−M\(∂PEM\)g\(ϕ−θ\)−𝖲\(θ,P\)2qW^\{\\prime\}=\-M\(\\partial\_\{P\}E\_\{M\}\)g\(\\phi\-\\theta\)\-\\frac\{\\mathsf\{S\}\(\\theta,P\)^\{2\}\}\{q\}showsW′≤0W^\{\\prime\}\\leq 0, where
𝖲\(θ,P\)=\{Pω\(ϕ−θ\)\+ω\(θ\),θ≥0,Pω\(ϕ−θ\)−ω\(−θ\),θ<0\.\\mathsf\{S\}\(\\theta,P\)=\\begin\{cases\}P\\omega\(\\phi\-\\theta\)\+\\omega\(\\theta\),&\\theta\\geq 0,\\\\ P\\omega\(\\phi\-\\theta\)\-\\omega\(\-\\theta\),&\\theta<0\.\\end\{cases\}Thusℬ\\mathcal\{B\}is invariant\.
The energy identity is
EM′=−\(∂qEM\)2−M\(∂PEM\)2−𝖲\(θ,P\)2\.E\_\{M\}^\{\\prime\}=\-\(\\partial\_\{q\}E\_\{M\}\)^\{2\}\-M\(\\partial\_\{P\}E\_\{M\}\)^\{2\}\-\\mathsf\{S\}\(\\theta,P\)^\{2\}\.\(B\.18\)The stationary equations do not depend onMM\. There is no stationary point with0<θ<ϕ0<\\theta<\\phibecause𝖲\>0\\mathsf\{S\}\>0\. Atθ=0\\theta=0, stationarity forcesP=0P=0andq=1q=1\. Forθ=−x∈\(−ϕ,0\)\\theta=\-x\\in\(\-\\phi,0\), coefficient stationarity yields
P^\(x\)=12g\(ϕ\)−g\(ϕ\+x\)g\(x\)14−g\(ϕ\+x\)2\.\\widehat\{P\}\(x\)=\\frac\{\\frac\{1\}\{2\}g\(\\phi\)\-g\(\\phi\+x\)g\(x\)\}\{\\frac\{1\}\{4\}\-g\(\\phi\+x\)^\{2\}\}\.The scalar certificate in Appendix[B\.5](https://arxiv.org/html/2608.06766#A2.SS5)proves
P^\(x\)ω\(ϕ\+x\)<ω\(x\),0<x≤ϕ,\\widehat\{P\}\(x\)\\omega\(\\phi\+x\)<\\omega\(x\),\\qquad 0<x\\leq\\phi,so𝖲≠0\\mathsf\{S\}\\neq 0and no negative\-angle stationary point exists\.
It remains to exclude an omega\-limit withq=0q=0\. Along any omega\-limit sequence, \([B\.18](https://arxiv.org/html/2608.06766#A2.E18)\) forces
∂qEM→0,∂PEM→0,𝖲→0\.\\partial\_\{q\}E\_\{M\}\\to 0,\\qquad\\partial\_\{P\}E\_\{M\}\\to 0,\\qquad\\mathsf\{S\}\\to 0\.Ifq=0q=0, then∂PEM=0\\partial\_\{P\}E\_\{M\}=0givesP=2g\(ϕ\)P=2g\(\\phi\)\. Forθ=−x<0\\theta=\-x<0, the equation∂qEM=0\\partial\_\{q\}E\_\{M\}=0would require
2g\(ϕ\)g\(ϕ\+x\)=g\(x\),2g\(\\phi\)g\(\\phi\+x\)=g\(x\),contradicting \([B\.14](https://arxiv.org/html/2608.06766#A2.E14)\)\. Forθ≥0\\theta\\geq 0, one has
𝖲=Pω\(ϕ−θ\)\+ω\(θ\)\>0\.\\mathsf\{S\}=P\\omega\(\\phi\-\\theta\)\+\\omega\(\\theta\)\>0\.Hence no zero\-qqomega\-limit exists\. The unique invariant subset of the zero\-dissipation set is therefore\(1,0,0\)\(1,0,0\), and LaSalle’s principle gives the claim\. ∎
###### Theorem B\.4\(Global capture in the multiplicity\-weighted fast flow\)\.
For every fixedM≥1M\\geq 1, the solution of \([B\.7](https://arxiv.org/html/2608.06766#A2.E7)\)–\([B\.8](https://arxiv.org/html/2608.06766#A2.E8)\) satisfies
q\(τ\)→1,P\(τ\)→0,θ\(τ\)→0\.q\(\\tau\)\\to 1,\\qquad P\(\\tau\)\\to 0,\\qquad\\theta\(\\tau\)\\to 0\.Moreover, it enters the sectorℬ\\mathcal\{B\}after the finite entrance time of Lemma[B\.2](https://arxiv.org/html/2608.06766#A2.SS2)\.
###### Proof\.
Lemma[B\.2](https://arxiv.org/html/2608.06766#A2.SS2)gives a point inℬ\\mathcal\{B\}withqqbounded away from zero\. Lemma[B\.3](https://arxiv.org/html/2608.06766#A2.SS3)then gives global convergence\. ∎
### B\.4Return to the finite\-gauge population flow
###### Proposition B\.5\(Finite\-DDapproximation\)\.
FixmmandT<∞T<\\infty\. Let\(qD,PD,θD,ψD\)\(q\_\{D\},P\_\{D\},\\theta\_\{D\},\\psi\_\{D\}\)solve the exact reduced flow \([B\.3](https://arxiv.org/html/2608.06766#A2.E3)\)–\([B\.4](https://arxiv.org/html/2608.06766#A2.E4)\) in fast time, and let\(q,P,θ\)\(q,P,\\theta\)solve \([B\.7](https://arxiv.org/html/2608.06766#A2.E7)\)\. Then
sup0≤τ≤T\(\|qD−q\|\+\|PD−P\|\+\|θD−θ\|\+\|ψD−ϕ\|\)≤Cm,TD2\.\\sup\_\{0\\leq\\tau\\leq T\}\\left\(\|q\_\{D\}\-q\|\+\|P\_\{D\}\-P\|\+\|\\theta\_\{D\}\-\\theta\|\+\|\\psi\_\{D\}\-\\phi\|\\right\)\\leq\\frac\{C\_\{m,T\}\}\{D^\{2\}\}\.
###### Proof\.
By Theorem[B\.4](https://arxiv.org/html/2608.06766#A2.Thmtheorem4), the fast trajectory stays in a compact set withqqbounded away from zero on\[0,T\]\[0,T\]\. On that set,
μ\+\(q\)D=1\+O\(D−2\),μ−\(P/M\)D=1\+O\(D−2\),\\frac\{\\mu\_\{\+\}\(q\)\}\{D\}=1\+O\(D^\{\-2\}\),\\qquad\\frac\{\\mu\_\{\-\}\(P/M\)\}\{D\}=1\+O\(D^\{\-2\}\),χ\+\(q\)Dq=1q2\+O\(D−2\),χ−\(P/M\)DP=O\(D−2\),\\frac\{\\chi\_\{\+\}\(q\)\}\{Dq\}=\\frac\{1\}\{q^\{2\}\}\+O\(D^\{\-2\}\),\\qquad\\frac\{\\chi\_\{\-\}\(P/M\)\}\{DP\}=O\(D^\{\-2\}\),uniformly with one derivative\. Hence the exact fast\-time vector field equals the limiting vector field plus aC1C^\{1\}perturbation of sizeCmD−2C\_\{m\}D^\{\-2\}\. Gronwall’s inequality proves the estimate\. ∎
### Local winner coercivity
At the zero\-loss winner state
q=1,θ=0,P=0,ψ=ϕ,q=1,\\qquad\\theta=0,\\qquad P=0,\\qquad\\psi=\\phi,the common loser direction is functionally invisible\. The functional transverse derivative is spanned by
\[u⊤x\]\+,𝟏\{u⊤x\>0\}v⊤x,\[sϕ⊤x\]\+,\[u^\{\\top\}x\]\_\{\+\},\\qquad\\mathbf\{1\}\_\{\\\{u^\{\\top\}x\>0\\\}\}v^\{\\top\}x,\\qquad\[s\_\{\\phi\}^\{\\top\}x\]\_\{\+\},wherev⟂uv\\perp ulies in the invariant plane\. These functions are linearly independent in GaussianL2L^\{2\}\.
###### Lemma B\.6\(Symmetric winner\-tube coercivity\)\.
For every fixedmm, there is a neighborhood𝒱m\\mathcal\{V\}\_\{m\}of the winner manifold and constantscm,Cm\>0c\_\{m\},C\_\{m\}\>0, independent ofDD, such that
cm\(\(q−1\)2\+θ2\+P2\)\\displaystyle c\_\{m\}\\bigl\(\(q\-1\)^\{2\}\+\\theta^\{2\}\+P^\{2\}\\bigr\)≤ℰ≤Cm\(\(q−1\)2\+θ2\+P2\),\\displaystyle\\leq\\mathscr\{E\}\\leq C\_\{m\}\\bigl\(\(q\-1\)^\{2\}\+\\theta^\{2\}\+P^\{2\}\\bigr\),\(B\.19\)\(∂qℰ\)2\+\(∂Pℰ\)2\+q−2\(∂θℰ\)2\\displaystyle\(\\partial\_\{q\}\\mathscr\{E\}\)^\{2\}\+\(\\partial\_\{P\}\\mathscr\{E\}\)^\{2\}\+q^\{\-2\}\(\\partial\_\{\\theta\}\\mathscr\{E\}\)^\{2\}≥cmℰ\.\\displaystyle\\geq c\_\{m\}\\mathscr\{E\}\.\(B\.20\)Furthermore,
\|∂ψℰ\|≤Cm\|P\|ℰon𝒱m\.\|\\partial\_\{\\psi\}\\mathscr\{E\}\|\\leq C\_\{m\}\|P\|\\sqrt\{\\mathscr\{E\}\}\\qquad\\text\{on \}\\mathcal\{V\}\_\{m\}\.\(B\.21\)
###### Proof\.
Positive definiteness of the Gram matrix of the three displayed tangent functions gives \([B\.19](https://arxiv.org/html/2608.06766#A2.E19)\) by Taylor expansion of the finite\-dimensional Gaussian kernel loss\. The same Hessian gives the local gradient inequality \([B\.20](https://arxiv.org/html/2608.06766#A2.E20)\)\. Since∂ψℰ\\partial\_\{\\psi\}\\mathscr\{E\}vanishes identically atP=0P=0and its remaining factor vanishes at the winner state, another Taylor expansion gives \([B\.21](https://arxiv.org/html/2608.06766#A2.E21)\)\. ∎
###### Theorem B\.7\(Global gauge\-selected capture and pruning\)\.
Fixm≥2m\\geq 2and the duplicate initialization \([B\.1](https://arxiv.org/html/2608.06766#A2.E1)\)\. Select any indexkk, assignδk=\+D\\delta\_\{k\}=\+D, and assignδj=−D\\delta\_\{j\}=\-Dfor everyj≠kj\\neq k\. There exist constants
D0\(m\),Cm,cm\>0D\_\{0\}\(m\),\\qquad C\_\{m\},\\qquad c\_\{m\}\>0and a winner tube𝒱m\\mathcal\{V\}\_\{m\}, all independent ofDD, such that for everyD≥D0\(m\)D\\geq D\_\{0\}\(m\):
1. 1\.the ordinary ReLU population\-gradient trajectory enters𝒱m\\mathcal\{V\}\_\{m\}at a time tcap≤CmD−1;t\_\{\\mathrm\{cap\}\}\\leq C\_\{m\}D^\{\-1\};
2. 2\.after entry, ℒ\(t\)≤ℒ\(tcap\)e−cmD\(t−tcap\);\\mathcal\{L\}\(t\)\\leq\\mathcal\{L\}\(t\_\{\\mathrm\{cap\}\}\)e^\{\-c\_\{m\}D\(t\-t\_\{\\mathrm\{cap\}\}\)\};
3. 3\.the selected student globally captures the teacher and all redundant coefficients are pruned: qk\(t\)→1,sk\(t\)→u,qj\(t\)→0\(j≠k\);q\_\{k\}\(t\)\\to 1,\\qquad s\_\{k\}\(t\)\\to u,\\qquad q\_\{j\}\(t\)\\to 0\\quad\(j\\neq k\);
4. 4\.every loser has post\-capture directional displacement ∫tcap∞‖s˙j\(t\)‖dt≤CmD−2,j≠k\.\\int\_\{t\_\{\\mathrm\{cap\}\}\}^\{\\infty\}\\\|\\dot\{s\}\_\{j\}\(t\)\\\|\\,\\mathrm\{d\}t\\leq C\_\{m\}D^\{\-2\},\\qquad j\\neq k\.Permuting the gauge marks permutes the global winner while leaving the complete initial predictor unchanged\.
###### Proof\.
LetM=m−1M=m\-1\. Theorem[B\.4](https://arxiv.org/html/2608.06766#A2.Thmtheorem4)implies that the limiting fast trajectory enters a strictly interior subset of𝒱m\\mathcal\{V\}\_\{m\}after a finite fast timeTmT\_\{m\}\. Proposition[B\.4](https://arxiv.org/html/2608.06766#A2.SS4)then places the exact finite\-DDtrajectory in𝒱m\\mathcal\{V\}\_\{m\}by physical time\(Tm\+1\)/D\(T\_\{m\}\+1\)/Dfor all largeDD\.
Inside𝒱m\\mathcal\{V\}\_\{m\}, exact dissipation \([B\.5](https://arxiv.org/html/2608.06766#A2.E5)\), the bounds
μ\+\(q\)≥D,Mμ−\(P/M\)≥D,χ\+\(q\)q≥Dq2,\\mu\_\{\+\}\(q\)\\geq D,\\qquad M\\mu\_\{\-\}\(P/M\)\\geq D,\\qquad\\frac\{\\chi\_\{\+\}\(q\)\}\{q\}\\geq\\frac\{D\}\{q^\{2\}\},and Lemma[B](https://arxiv.org/html/2608.06766#A2.SSx2)give
−ℰ˙≥cmDℰ\.\-\\dot\{\\mathscr\{E\}\}\\geq c\_\{m\}D\\mathscr\{E\}\.This proves exponential contraction\. The local loss equivalence then yields
q→1,θ→0,P→0\.q\\to 1,\\qquad\\theta\\to 0,\\qquad P\\to 0\.Since every loser coefficient equalsp=P/Mp=P/M, eachqj→0q\_\{j\}\\to 0\.
For the common loser direction,
\|χ−\(P/M\)\|≤\|P\|MD\.\\left\|\\chi\_\{\-\}\(P/M\)\\right\|\\leq\\frac\{\|P\|\}\{MD\}\.Combining this with \([B\.21](https://arxiv.org/html/2608.06766#A2.E21)\) and \([B\.19](https://arxiv.org/html/2608.06766#A2.E19)\) gives
\|ψ˙\|≤CmDℰ\.\|\\dot\{\\psi\}\|\\leq\\frac\{C\_\{m\}\}\{D\}\\mathscr\{E\}\.Integrating the exponential loss bound produces
∫tcap∞\|ψ˙\|dt≤CmD−2\.\\int\_\{t\_\{\\mathrm\{cap\}\}\}^\{\\infty\}\|\\dot\{\\psi\}\|\\,\\mathrm\{d\}t\\leq C\_\{m\}D^\{\-2\}\.All losers share this direction by exact exchange symmetry\. Finally, the equations are equivariant under permutation of student labels, whereas the initial function depends only on the common visible state; permuting the unique positive gauge therefore permutes the selected winner without changing the initial predictor\. ∎
### B\.5Scalar arc\-cosine certificate
The only non\-elementary scalar sign used in Lemma[B\.3](https://arxiv.org/html/2608.06766#A2.SS3)is the exclusion of a negative\-angle stationary state\. Put
P^\(x\)=12g\(ϕ\)−g\(ϕ\+x\)g\(x\)14−g\(ϕ\+x\)2,\\widehat\{P\}\(x\)=\\frac\{\\frac\{1\}\{2\}g\(\\phi\)\-g\(\\phi\+x\)g\(x\)\}\{\\frac\{1\}\{4\}\-g\(\\phi\+x\)^\{2\}\},and define
F\(x\)=\(14−g\(ϕ\+x\)2\)ω\(x\)−\(12g\(ϕ\)−g\(ϕ\+x\)g\(x\)\)ω\(ϕ\+x\)\.F\(x\)=\\Bigl\(\\tfrac\{1\}\{4\}\-g\(\\phi\+x\)^\{2\}\\Bigr\)\\omega\(x\)\-\\Bigl\(\\tfrac\{1\}\{2\}g\(\\phi\)\-g\(\\phi\+x\)g\(x\)\\Bigr\)\\omega\(\\phi\+x\)\.The included scriptscripts/certify\_kernel\.pyevaluates the exact expression forF′F^\{\\prime\}with outward\-rounded 80\-decimal interval arithmetic on 2048 equal subintervals of\[0,π/4\]\[0,\\pi/4\]\. The certified lower bound stored indata/kernel\_certificate\.jsonis strictly positive\. ThereforeF\(x\)\>0F\(x\)\>0forx\>0x\>0, which is equivalent to
P^\(x\)ω\(ϕ\+x\)<ω\(x\)\.\\widehat\{P\}\(x\)\\omega\(\\phi\+x\)<\\omega\(x\)\.The multiplicityMMnever enters this scalar certificate\.
### B\.6A convenient closed\-form entrance constant
For reference,
A=12−g\(π/4\),ω0=ω\(π/8\),A=\\frac\{1\}\{2\}\-g\(\\pi/4\),\\qquad\\omega\_\{0\}=\\omega\(\\pi/8\),and
γent=A22ω02=0\.26678…<1\.\\gamma\_\{\\rm ent\}=\\frac\{A^\{2\}\}\{2\\omega\_\{0\}^\{2\}\}=0\.26678\\ldots<1\.Thus the entrance argument retains at least73\.3%73\.3\\%of the initial winner coefficient before the canonical invariant sector is reached, uniformly over all multiplicities\. This uniform fraction enables the arbitrary\-fixed\-mmextension even though the aggregate redundant coefficient approaches one asmmgrows\.
## Appendix CRobustness and discrete gradient descent
LetX∼𝒩\(0,Id\)X\\sim\\mathcal\{N\}\(0,I\_\{d\}\)and let the teacher be
f⋆\(x\)=\[u⊤x\]\+,‖u‖=1\.f\_\{\\star\}\(x\)=\[u^\{\\top\}x\]\_\{\+\},\\qquad\\\|u\\\|=1\.The student hasm≥2m\\geq 2raw ReLU units,
f\(x\)=∑i=1mqi\[si⊤x\]\+,‖si‖=1,f\(x\)=\\sum\_\{i=1\}^\{m\}q\_\{i\}\[s\_\{i\}^\{\\top\}x\]\_\{\+\},\\qquad\\\|s\_\{i\}\\\|=1,trained by population square\-loss gradient flow\. Each functional coefficient and direction arise from ordinary parameters\(ai,wi\)\(a\_\{i\},w\_\{i\}\)through
qi=ai‖wi‖,si=wi‖wi‖,δi=ai2−‖wi‖2\.q\_\{i\}=a\_\{i\}\\\|w\_\{i\}\\\|,\\qquad s\_\{i\}=\\frac\{w\_\{i\}\}\{\\\|w\_\{i\}\\\|\},\\qquad\\delta\_\{i\}=a\_\{i\}^\{2\}\-\\\|w\_\{i\}\\\|^\{2\}\.The marksδi\\delta\_\{i\}are conserved\. We designate unit11as the programmed winner and set
δ1=\+D,δj=−D\(j≥2\)\.\\delta\_\{1\}=\+D,\\qquad\\delta\_\{j\}=\-D\\quad\(j\\geq 2\)\.At the exactly duplicate visible initialization,
qi\(0\)=1m,si\(0\)=s0,∠\(s0,u\)=ϕ,q\_\{i\}\(0\)=\\frac\{1\}\{m\},\\qquad s\_\{i\}\(0\)=s\_\{0\},\\qquad\\angle\(s\_\{0\},u\)=\\phi,the initial predictor is independent of the choice of winner\.
The exact duplicate model provides the cleanest same\-predictor comparison and the unconditional global selection theorem\. This appendix develops four complementary extensions of that result: a certified interval of initial angles, explicit sufficient dependence on the fixed multiplicitymm, stability of finite\-time selection under visible perturbations, and a small\-step full\-batch gradient\-descent analogue\. For fully nonsymmetric redundant dictionaries, long\-time coefficient\-wise pruning is stated under explicit transversality and active\-frame conditions that rule out functional cancellation modes\.
Forα∈\[0,π\]\\alpha\\in\[0,\\pi\], write
g\(α\)=sinα\+\(π−α\)cosα2π,ω\(α\)=\(π−α\)sinα2π=−g′\(α\)\.g\(\\alpha\)=\\frac\{\\sin\\alpha\+\(\\pi\-\\alpha\)\\cos\\alpha\}\{2\\pi\},\\qquad\\omega\(\\alpha\)=\\frac\{\(\\pi\-\\alpha\)\\sin\\alpha\}\{2\\pi\}=\-g^\{\\prime\}\(\\alpha\)\.On the duplicate\-loser manifold, putM=m−1M=m\-1, letqqbe the winner coefficient,PPthe aggregate loser coefficient, andθ,ψ\\theta,\\psitheir signed angles fromuu\. In fast timeτ=Dt\\tau=Dt, the limiting capture flow is
q′=−∂qEϕ,P′=−M∂PEϕ,θ′=−q−2∂θEϕ,ψ≡ϕ,q^\{\\prime\}=\-\\partial\_\{q\}E\_\{\\phi\},\\qquad P^\{\\prime\}=\-M\\partial\_\{P\}E\_\{\\phi\},\\qquad\\theta^\{\\prime\}=\-q^\{\-2\}\\partial\_\{\\theta\}E\_\{\\phi\},\\qquad\\psi\\equiv\\phi,\(C\.1\)where
Eϕ\(q,P,θ\)=14\(q2\+P2\+1\)\+qPg\(ϕ−θ\)−qg\(\|θ\|\)−Pg\(ϕ\)\.E\_\{\\phi\}\(q,P,\\theta\)=\\frac\{1\}\{4\}\(q^\{2\}\+P^\{2\}\+1\)\+qPg\(\\phi\-\\theta\)\-qg\(\|\\theta\|\)\-Pg\(\\phi\)\.\(C\.2\)The exact finite\-DDmarked flow differs from \([C\.1](https://arxiv.org/html/2608.06766#A3.E1)\) byO\(D−2\)O\(D^\{\-2\}\)on compact regular sets\.
### C\.1Certified initial\-angle basin
The global basin argument uses four arc\-cosine inequalities\. The first result shows that the mechanism is not tied to the reference angleϕ=π/4\\phi=\\pi/4\.
###### Theorem C\.1\(Certified angle basin\)\.
For every
ϕ∈Iϕ:=\[0\.77,0\.90\],\\phi\\in I\_\{\\phi\}:=\[0\.77,0\.90\],all scalar inequalities used in the multiplicity\-weighted basin proof hold uniformly\. Consequently, for every fixedm≥2m\\geq 2the fast flow \([C\.1](https://arxiv.org/html/2608.06766#A3.E1)\), initialized at
q\(0\)=1m,P\(0\)=m−1m,θ\(0\)=ϕ,q\(0\)=\\frac\{1\}\{m\},\\qquad P\(0\)=\\frac\{m\-1\}\{m\},\\qquad\\theta\(0\)=\\phi,converges to\(q,P,θ\)=\(1,0,0\)\(q,P,\\theta\)=\(1,0,0\)\. The entrance lemma retains at least
1−γent⋆\>0\.65881\-\\gamma\_\{\\rm ent\}^\{\\star\}\>0\.6588of the initial winner coefficient, uniformly overmmandϕ∈Iϕ\\phi\\in I\_\{\\phi\}\.
###### Proof\.
For0≤x≤ϕ0\\leq x\\leq\\phi, the three non\-strict inequalities admit elementary derivative factorizations\. IfR1,R2,R3R\_\{1\},R\_\{2\},R\_\{3\}denote their nonnegative residuals, then
R1′\(x\)=\(ϕ−2x\)sinxsin\(ϕ−x\)2π2\(0≤x≤ϕ/2\),R\_\{1\}^\{\\prime\}\(x\)=\\frac\{\(\\phi\-2x\)\\sin x\\sin\(\\phi\-x\)\}\{2\\pi^\{2\}\}\\quad\(0\\leq x\\leq\\phi/2\),andR1\(x\)=R1\(ϕ−x\)R\_\{1\}\(x\)=R\_\{1\}\(\\phi\-x\)\. Further,
R2′\(x\)=−ω\(x\)−2g\(ϕ\)ω\(ϕ−x\)<0,R2\(ϕ\)=0,R\_\{2\}^\{\\prime\}\(x\)=\-\\omega\(x\)\-2g\(\\phi\)\\omega\(\\phi\-x\)<0,\\qquad R\_\{2\}\(\\phi\)=0,while
R3′\(x\)=\(2π−ϕ−2x\)sinxsin\(ϕ\+x\)2π2≥0,R3\(0\)=0\.R\_\{3\}^\{\\prime\}\(x\)=\\frac\{\(2\\pi\-\\phi\-2x\)\\sin x\\sin\(\\phi\+x\)\}\{2\\pi^\{2\}\}\\geq 0,\\qquad R\_\{3\}\(0\)=0\.The remaining strict negative\-angle inequality, the negative\-angle stationary\-point exclusion, the inward angular\-face margin, and the entrance constants were certified by outward\-rounded interval arithmetic onIϕI\_\{\\phi\}; seedata/angle\_interval\_certificate\.json\. The certified margins include
infK4\>0\.20727,infF′\>0\.01659,inf\{ω\(ϕ\)−ω\(2ϕ\)\}\>0\.00780\.\\inf K\_\{4\}\>0\.20727,\\qquad\\inf F^\{\\prime\}\>0\.01659,\\qquad\\inf\\\{\\omega\(\\phi\)\-\\omega\(2\\phi\)\\\}\>0\.00780\.Finally,
γent\(ϕ\)=\(1/2−g\(ϕ\)\)22ω\(ϕ/2\)2≤0\.34116\.\\gamma\_\{\\rm ent\}\(\\phi\)=\\frac\{\(1/2\-g\(\\phi\)\)^\{2\}\}\{2\\omega\(\\phi/2\)^\{2\}\}\\leq 0\.34116\.The entrance and invariant\-sector arguments in the proof of[theorem˜B\.7](https://arxiv.org/html/2608.06766#A2.Thmtheorem7)therefore apply uniformly throughoutIϕI\_\{\\phi\}\. ∎
### C\.2Dependence on multiplicity
We next separate the part of the capture dynamics that is uniform over fixed multiplicities from the conservativemm\-dependence introduced by finite\-DDshadowing\.
Forϕ∈Iϕ\\phi\\in I\_\{\\phi\}, letℬϕ\\mathcal\{B\}\_\{\\phi\}be the invariant arc\-cosine sector used in the proof of[theorem˜B\.7](https://arxiv.org/html/2608.06766#A2.Thmtheorem7)\. Choose a fixed winner tube
𝒱⋆\(ϕ\)=\{\(q,P,θ\):\(q−1\)2\+P2\+θ2<r⋆2\}⋐ℬϕ\\mathcal\{V\}\_\{\\star\}\(\\phi\)=\\\{\(q,P,\\theta\):\(q\-1\)^\{2\}\+P^\{2\}\+\\theta^\{2\}<r\_\{\\star\}^\{2\}\\\}\\Subset\\mathcal\{B\}\_\{\\phi\}withr⋆\>0r\_\{\\star\}\>0uniform overIϕI\_\{\\phi\}\. Put
Γϕ\(q,P,θ\)=\(∂qEϕ\)2\+\(∂PEϕ\)2\+q−2\(∂θEϕ\)2\\Gamma\_\{\\phi\}\(q,P,\\theta\)=\(\\partial\_\{q\}E\_\{\\phi\}\)^\{2\}\+\(\\partial\_\{P\}E\_\{\\phi\}\)^\{2\}\+q^\{\-2\}\(\\partial\_\{\\theta\}E\_\{\\phi\}\)^\{2\}whenq\>0q\>0, and use the continuous nonsingular form
Γ~ϕ=\(∂qEϕ\)2\+\(∂PEϕ\)2\+𝖲ϕ2\\widetilde\{\\Gamma\}\_\{\\phi\}=\(\\partial\_\{q\}E\_\{\\phi\}\)^\{2\}\+\(\\partial\_\{P\}E\_\{\\phi\}\)^\{2\}\+\\mathsf\{S\}\_\{\\phi\}^\{2\}on the closure\. The critical\-point classification in[theorem˜C\.1](https://arxiv.org/html/2608.06766#A3.Thmtheorem1)implies
γ⋆:=infϕ∈Iϕinfℬϕ¯∖𝒱⋆\(ϕ\)Γ~ϕ\>0\.\\gamma\_\{\\star\}:=\\inf\_\{\\phi\\in I\_\{\\phi\}\}\\inf\_\{\\overline\{\\mathcal\{B\}\_\{\\phi\}\}\\setminus\\mathcal\{V\}\_\{\\star\}\(\\phi\)\}\\widetilde\{\\Gamma\}\_\{\\phi\}\>0\.LetE⋆E\_\{\\star\}be a uniform upper bound on the initial energy and letTent⋆T\_\{\\rm ent\}^\{\\star\}be the uniform entrance\-time bound supplied by[theorem˜C\.1](https://arxiv.org/html/2608.06766#A3.Thmtheorem1)\. Define
T⋆:=Tent⋆\+E⋆γ⋆\.T\_\{\\star\}:=T\_\{\\rm ent\}^\{\\star\}\+\\frac\{E\_\{\\star\}\}\{\\gamma\_\{\\star\}\}\.\(C\.3\)
###### Proposition C\.3\(Uniform fast capture and explicit finite\-gauge threshold\)\.
There exist positive constantsA0,L0,r⋆,c⋆,C⋆A\_\{0\},L\_\{0\},r\_\{\\star\},c\_\{\\star\},C\_\{\\star\}, depending only onIϕI\_\{\\phi\}, such that the following holds\.
For everym≥2m\\geq 2andϕ∈Iϕ\\phi\\in I\_\{\\phi\}, the fast flow enters𝒱⋆\(ϕ\)\\mathcal\{V\}\_\{\\star\}\(\\phi\)by fast timeT⋆T\_\{\\star\}, independently ofmm\. On the exact finite\-DDduplicate\-loser flow, the pre\-capture vector field has a Lipschitz bound
Lm≤L0m2L\_\{m\}\\leq L\_\{0\}m^\{2\}and differs from the fast vector field by at mostA0D−2A\_\{0\}D^\{\-2\}\. Hence the sufficient threshold
D0\(m\):=⌈\(4A0T⋆r⋆\)1/2exp\(L0T⋆2m2\)⌉D\_\{0\}\(m\):=\\left\\lceil\\left\(\\frac\{4A\_\{0\}T\_\{\\star\}\}\{r\_\{\\star\}\}\\right\)^\{1/2\}\\exp\\\!\\left\(\\frac\{L\_\{0\}T\_\{\\star\}\}\{2\}m^\{2\}\\right\)\\right\\rceil\(C\.4\)ensures entry into the winner tube by physical time
tcap≤\(T⋆\+1\)D−1\.t\_\{\\rm cap\}\\leq\(T\_\{\\star\}\+1\)D^\{\-1\}\.Inside the tube, the loss contraction constant can be chosen uniformly inmm:
ℒ\(t\)≤ℒ\(tcap\)e−c⋆D\(t−tcap\)\.\\mathcal\{L\}\(t\)\\leq\\mathcal\{L\}\(t\_\{\\rm cap\}\)e^\{\-c\_\{\\star\}D\(t\-t\_\{\\rm cap\}\)\}\.For the symmetric loser block, the common loser direction satisfies the sharper bound
∫tcap∞‖s˙lose\(t\)‖dt≤C⋆\(m−1\)D2\.\\int\_\{t\_\{\\rm cap\}\}^\{\\infty\}\\\|\\dot\{s\}\_\{\\rm lose\}\(t\)\\\|\\,\\mathrm\{d\}t\\leq\\frac\{C\_\{\\star\}\}\{\(m\-1\)D^\{2\}\}\.
###### Proof\.
The entrance lemma is uniform inmmbecause its time is at mostC/mC/mand it retains a fixed fraction ofq\(0\)=1/mq\(0\)=1/m\. Once inside the canonical sector,
ddτEϕ=−\(∂qEϕ\)2−M\(∂PEϕ\)2−𝖲ϕ2≤−Γ~ϕ\.\\frac\{\\,\\mathrm\{d\}\}\{\\,\\mathrm\{d\}\\tau\}E\_\{\\phi\}=\-\(\\partial\_\{q\}E\_\{\\phi\}\)^\{2\}\-M\(\\partial\_\{P\}E\_\{\\phi\}\)^\{2\}\-\\mathsf\{S\}\_\{\\phi\}^\{2\}\\leq\-\\widetilde\{\\Gamma\}\_\{\\phi\}\.Outside𝒱⋆\\mathcal\{V\}\_\{\\star\}, the right\-hand side is at most−γ⋆\-\\gamma\_\{\\star\}, proving the uniform fast\-time bound \([C\.3](https://arxiv.org/html/2608.06766#A3.E3)\)\.
Before capture,qqremains bounded below byc/mc/m; differentiating the polar\-coordinate fast field therefore givesLm≤L0m2L\_\{m\}\\leq L\_\{0\}m^\{2\}\. The exact mobility identities give, uniformly on the same compact set,
\|μ\+D−1\|≤2D2,M\|μ−D−1\|≤2D2,\\left\|\\frac\{\\mu\_\{\+\}\}\{D\}\-1\\right\|\\leq\\frac\{2\}\{D^\{2\}\},\\qquad M\\left\|\\frac\{\\mu\_\{\-\}\}\{D\}\-1\\right\|\\leq\\frac\{2\}\{D^\{2\}\},\|χ\+Dq−1q2\|≤1D2,χ−D\|P\|≤1\(m−1\)D2\.\\left\|\\frac\{\\chi\_\{\+\}\}\{Dq\}\-\\frac\{1\}\{q^\{2\}\}\\right\|\\leq\\frac\{1\}\{D^\{2\}\},\\qquad\\frac\{\\chi\_\{\-\}\}\{D\|P\|\}\\leq\\frac\{1\}\{\(m\-1\)D^\{2\}\}\.Thus the finite\-DDvector\-field defect is at mostA0D−2A\_\{0\}D^\{\-2\}\. Gronwall gives
supτ≤T⋆‖zD\(τ\)−z∞\(τ\)‖≤A0T⋆D−2eL0m2T⋆,\\sup\_\{\\tau\\leq T\_\{\\star\}\}\\\|z\_\{D\}\(\\tau\)\-z\_\{\\infty\}\(\\tau\)\\\|\\leq A\_\{0\}T\_\{\\star\}D^\{\-2\}e^\{L\_\{0\}m^\{2\}T\_\{\\star\}\},which is belowr⋆/4r\_\{\\star\}/4under \([C\.4](https://arxiv.org/html/2608.06766#A3.E4)\)\.
The local Gaussian feature Gram on the winner tube is uniformly positive overIϕI\_\{\\phi\}\. Exact reaction–transport dissipation andMμ−≥DM\\mu\_\{\-\}\\geq Dtherefore yield the uniform PL inequality
−ℒ˙≥c⋆Dℒ\.\-\\dot\{\\mathcal\{L\}\}\\geq c\_\{\\star\}D\\mathcal\{L\}\.Finally,
χ−\(P/M\)≤\|P\|\(m−1\)D,\|∂ψE\|≤C\|P\|ℒ,\\chi\_\{\-\}\(P/M\)\\leq\\frac\{\|P\|\}\{\(m\-1\)D\},\\qquad\|\\partial\_\{\\psi\}E\|\\leq C\|P\|\\sqrt\{\\mathcal\{L\}\},and local coercivity gives\|P\|≤Cℒ\|P\|\\leq C\\sqrt\{\\mathcal\{L\}\}\. Integrating the exponential loss bound proves the stated loser\-direction estimate\. ∎
### C\.3Open\-set selection and conditional locking
Exact duplication is a clean same\-function counterfactual, while finite\-time selection should persist under perturbations of the visible state\. Long\-time pruning is more delicate because zero\-loss over\-realized ReLU representations contain functionally neutral directions\. We therefore distinguish:
1. 1\.an unconditional finite\-time*selection*theorem on an open set around duplicate initialization;
2. 2\.an asymptotic*locking and pruning*theorem under explicit quotient\-transversality and active redundant\-frame conditions\.
The second statement requires two distinct nondegeneracy controls: quotient transversality separates the winner tangent from the redundant function space, while the active redundant frame controls the actual coefficient vector\.
For the fullmm\-student system, define the visible/gauge perturbation size
Δ0:=maxi\(\|qi\(0\)−m−1\|\+∥si\(0\)−s0∥\)\+maxi\|\|δi\(0\)\|D−1\|,\\Delta\_\{0\}:=\\max\_\{i\}\\left\(\|q\_\{i\}\(0\)\-m^\{\-1\}\|\+\\\|s\_\{i\}\(0\)\-s\_\{0\}\\\|\\right\)\+\\max\_\{i\}\\left\|\\frac\{\|\\delta\_\{i\}\(0\)\|\}\{D\}\-1\\right\|,\(C\.5\)whereδ1\>0\\delta\_\{1\}\>0andδj<0\\delta\_\{j\}<0forj≥2j\\geq 2\. Put
fred\(x,t\)=∑j=2mqj\(t\)\[sj\(t\)⊤x\]\+\.f\_\{\\rm red\}\(x,t\)=\\sum\_\{j=2\}^\{m\}q\_\{j\}\(t\)\[s\_\{j\}\(t\)^\{\\top\}x\]\_\{\+\}\.LetZD∘\(τ\)Z\_\{D\}^\{\\circ\}\(\\tau\)denote the exact duplicate reference trajectory in fast timeτ=Dt\\tau=Dt, and let𝒰pre\\mathcal\{U\}\_\{\\rm pre\}be a compact pre\-capture neighborhood containingZD∘\(\[0,T⋆\+1\]\)Z\_\{D\}^\{\\circ\}\(\[0,T\_\{\\star\}\+1\]\)for everyD≥2D0\(m\)D\\geq 2D\_\{0\}\(m\)\.
###### Lemma C\.5\(Finite\-horizon full\-system stability\)\.
For every fixedmm, the full marked vector field on𝒰pre\\mathcal\{U\}\_\{\\rm pre\}is locally Lipschitz with
Lip\(Fm,D\)≤L¯m,L¯m≤L¯m3,\\operatorname\{Lip\}\(F\_\{m,D\}\)\\leq\\bar\{L\}\_\{m\},\\qquad\\bar\{L\}\_\{m\}\\leq\\bar\{L\}m^\{3\},and its dependence on the relative gauge magnitudes is bounded byA¯m2Δ0\\bar\{A\}m^\{2\}\\Delta\_\{0\}\. Consequently,
sup0≤τ≤T⋆\+1‖ZDpert\(τ\)−ZD∘\(τ\)‖≤eL¯m3\(T⋆\+1\)\(1\+A¯m2\(T⋆\+1\)\)Δ0\.\\sup\_\{0\\leq\\tau\\leq T\_\{\\star\}\+1\}\\\|Z\_\{D\}^\{\\rm pert\}\(\\tau\)\-Z\_\{D\}^\{\\circ\}\(\\tau\)\\\|\\leq e^\{\\bar\{L\}m^\{3\}\(T\_\{\\star\}\+1\)\}\(1\+\\bar\{A\}m^\{2\}\(T\_\{\\star\}\+1\)\)\\Delta\_\{0\}\.\(C\.6\)
###### Proof\.
On𝒰pre\\mathcal\{U\}\_\{\\rm pre\}allqiq\_\{i\}stay in a compact interval bounded away from zero and all pairwise angles stay away from the nonsmooth antipodal endpoint\. The Gaussian ReLU kernelκ\(si⊤sj\)\\kappa\(s\_\{i\}^\{\\top\}s\_\{j\}\)and its first two angular derivatives are therefore uniformly bounded\. Differentiating themmcoefficient equations and themmspherical transport equations gives at mostO\(m2\)O\(m^\{2\}\)pair interactions per row andO\(m\)O\(m\)rows in the Euclidean product norm, hence the conservative boundL¯m3\\bar\{L\}m^\{3\}\. The mobility maps\(q,\|δ\|/D\)↦μ/D\(q,\|\\delta\|/D\)\\mapsto\\mu/Dand\(q,\|δ\|/D\)↦χ/D\(q,\|\\delta\|/D\)\\mapsto\\chi/Dhave bounded first derivatives on the same compact set\. The gauge\-magnitude perturbation therefore contributes at mostA¯m2Δ0\\bar\{A\}m^\{2\}\\Delta\_\{0\}\. The estimate follows from the standard continuous\-dependence inequality for ODEs\. ∎
###### Theorem C\.6\(Open\-set finite\-time winner selection\)\.
There are universal constantsa,b\>0a,b\>0such that, for every fixedm≥2m\\geq 2, everyϕ∈Iϕ\\phi\\in I\_\{\\phi\}, and
Δ0≤εm:=ae−bm3,D≥2D0\(m\),\\Delta\_\{0\}\\leq\\varepsilon\_\{m\}:=ae^\{\-bm^\{3\}\},\\qquad D\\geq 2D\_\{0\}\(m\),\(C\.7\)the designated positive\-gauge student is the unique student to enter the fixed winner\-dominant set
𝒲sel:=\{q1≥1−r⋆,∠\(s1,u\)≤r⋆,‖fred‖L2\(π\)≤r⋆\}\\mathcal\{W\}\_\{\\rm sel\}:=\\left\\\{q\_\{1\}\\geq 1\-r\_\{\\star\},\\ \\angle\(s\_\{1\},u\)\\leq r\_\{\\star\},\\ \\\|f\_\{\\rm red\}\\\|\_\{L^\{2\}\(\\pi\)\}\\leq r\_\{\\star\}\\right\\\}by timetsel≤\(T⋆\+1\)/Dt\_\{\\rm sel\}\\leq\(T\_\{\\star\}\+1\)/D\. Every negative\-gauge student remains outside the corresponding teacher\-alignment threshold attselt\_\{\\rm sel\}\.
###### Proof\.
The exact duplicate reference orbit enters the interior of𝒲sel\\mathcal\{W\}\_\{\\rm sel\}at fast time at mostT⋆T\_\{\\star\}, with a strictly positive winner/loser margin\. Choosea,ba,bso that the right\-hand side of \([C\.6](https://arxiv.org/html/2608.06766#A3.E6)\), together with theO\(D−2\)O\(D^\{\-2\}\)finite\-gauge error, is smaller than one quarter of this margin\. The visible state, the redundant function, and every teacher alignment are continuous functions of the marked state on𝒰pre\\mathcal\{U\}\_\{\\rm pre\}\. The perturbed trajectory therefore enters𝒲sel\\mathcal\{W\}\_\{\\rm sel\}before fast timeT⋆\+1T\_\{\\star\}\+1, while every negative\-gauge student remains below the winner threshold\. This proves a genuine open\-set selection statement without invoking any post\-capture nondegeneracy\. ∎
#### The active quotient frame\.
Letℋ:=L2\(γd\)\\mathcal\{H\}:=L^\{2\}\(\\gamma\_\{d\}\)and write
φs\(x\):=\[s⊤x\]\+,τu,v\(x\):=𝟏\{u⊤x\>0\}v⊤x,v⟂u\.\\varphi\_\{s\}\(x\):=\[s^\{\\top\}x\]\_\{\+\},\\qquad\\tau\_\{u,v\}\(x\):=\\mathbf\{1\}\_\{\\\{u^\{\\top\}x\>0\\\}\}v^\{\\top\}x,\\qquad v\\perp u\.For a redundant direction tuple𝐬=\(s2,…,sm\)\\mathbf\{s\}=\(s\_\{2\},\\ldots,s\_\{m\}\)define the redundant synthesis map
S𝐬β:=∑j=2mβjφsj,K𝐬:=kerS𝐬\.S\_\{\\mathbf\{s\}\}\\beta:=\\sum\_\{j=2\}^\{m\}\\beta\_\{j\}\\varphi\_\{s\_\{j\}\},\\qquad K\_\{\\mathbf\{s\}\}:=\\ker S\_\{\\mathbf\{s\}\}\.Only the redundant coefficient block is quotiented: the active coordinate space is
ℰ𝐬:=ℝ×u⟂×\(ℝm−1/K𝐬\),\\mathcal\{E\}\_\{\\mathbf\{s\}\}:=\\mathbb\{R\}\\times u^\{\\perp\}\\times\\bigl\(\\mathbb\{R\}^\{m\-1\}/K\_\{\\mathbf\{s\}\}\\bigr\),with norm
‖\(α,v,\[β\]\)‖q2:=\|α\|2\+‖v‖2\+infk∈K𝐬‖β\+k‖2\.\\\|\(\\alpha,v,\[\\beta\]\)\\\|\_\{\\rm q\}^\{2\}:=\|\\alpha\|^\{2\}\+\\\|v\\\|^\{2\}\+\\inf\_\{k\\in K\_\{\\mathbf\{s\}\}\}\\\|\\beta\+k\\\|^\{2\}\.The associated linearized realization frame is
F𝐬\(α,v,\[β\]\):=αφu\+τu,v\+S𝐬β\.F\_\{\\mathbf\{s\}\}\(\\alpha,v,\[\\beta\]\):=\\alpha\\varphi\_\{u\}\+\\tau\_\{u,v\}\+S\_\{\\mathbf\{s\}\}\\beta\.\(C\.8\)We say that the winner and redundant dictionary are*quotient\-transverse*with constantλq\>0\\lambda\_\{\\rm q\}\>0if
‖F𝐬ξ‖ℋ2≥λq‖ξ‖q2\(ξ∈ℰ𝐬\)\.\\\|F\_\{\\mathbf\{s\}\}\\xi\\\|\_\{\\mathcal\{H\}\}^\{2\}\\geq\\lambda\_\{\\rm q\}\\\|\\xi\\\|\_\{\\rm q\}^\{2\}\\qquad\(\\xi\\in\\mathcal\{E\}\_\{\\mathbf\{s\}\}\)\.\(C\.9\)This condition rules out a cancellation between the winner coefficient/directional tangent and the function generated by the redundant ridges\. It still permits coefficient zero modes among the redundant ridges themselves\.
For continuation of the*actual*redundant coefficients we use the stronger active coefficient\-frame condition
‖S𝐬β‖ℋ2≥λred‖β‖2\(β∈ℝm−1\)\.\\\|S\_\{\\mathbf\{s\}\}\\beta\\\|\_\{\\mathcal\{H\}\}^\{2\}\\geq\\lambda\_\{\\rm red\}\\\|\\beta\\\|^\{2\}\\qquad\(\\beta\\in\\mathbb\{R\}^\{m\-1\}\)\.\(C\.10\)Unlike \([C\.9](https://arxiv.org/html/2608.06766#A3.E9)\), this excludes redundant coefficient kernels\. The exact duplicate theorem of Appendix B does not require \([C\.10](https://arxiv.org/html/2608.06766#A3.E10)\): there the exchange symmetry reduces all redundant coefficients to the single aggregate variablePP\. The condition is needed only for a fully nonsymmetric open\-basin locking statement\.
###### Lemma C\.7\(Gaussian ReLU ridge expansion\)\.
There are constantsrridge,Cridge\>0r\_\{\\rm ridge\},C\_\{\\rm ridge\}\>0, depending only on the input dimension, with the following property\. Forv⟂uv\\perp u,‖v‖≤rridge\\\|v\\\|\\leq r\_\{\\rm ridge\}, lets\(v\)=Expu\(v\)s\(v\)=\\operatorname\{Exp\}\_\{u\}\(v\)\. Then
‖φs\(v\)−φu−τu,v‖ℋ\\displaystyle\\\|\\varphi\_\{s\(v\)\}\-\\varphi\_\{u\}\-\\tau\_\{u,v\}\\\|\_\{\\mathcal\{H\}\}≤Cridge‖v‖3/2,\\displaystyle\\leq C\_\{\\rm ridge\}\\\|v\\\|^\{3/2\},\(C\.11\)‖Dvφs\(v\)−D0φs\(0\)‖op\\displaystyle\\\|D\_\{v\}\\varphi\_\{s\(v\)\}\-D\_\{0\}\\varphi\_\{s\(0\)\}\\\|\_\{\\rm op\}≤Cridge‖v‖1/2\.\\displaystyle\\leq C\_\{\\rm ridge\}\\\|v\\\|^\{1/2\}\.\(C\.12\)Here tangent spaces are identified by parallel transport along the minimizing spherical geodesic\.
###### Proof\.
LetAvA\_\{v\}be the sign\-disagreement wedge
Av=\{x:sign\(u⊤x\)≠sign\(s\(v\)⊤x\)\}\.A\_\{v\}=\\\{x:\\operatorname\{sign\}\(u^\{\\top\}x\)\\neq\\operatorname\{sign\}\(s\(v\)^\{\\top\}x\)\\\}\.Rotational invariance of the Gaussian law givesγd\(Av\)≤C‖v‖\\gamma\_\{d\}\(A\_\{v\}\)\\leq C\\\|v\\\|\. OnAvcA\_\{v\}^\{c\}the ReLU is linear on the same half\-space, and the sphere\-chart remainder isO\(‖v‖2‖x‖\)O\(\\\|v\\\|^\{2\}\\\|x\\\|\)\. OnAvA\_\{v\}the first\-order residual is bounded byC‖v‖‖x‖C\\\|v\\\|\\\|x\\\|\. Cauchy–Schwarz and the Gaussian fourth moment therefore give
‖φs\(v\)−φu−τu,v‖L2\(Avc\)≤C‖v‖2,‖φs\(v\)−φu−τu,v‖L2\(Av\)≤C‖v‖3/2,\\\|\\varphi\_\{s\(v\)\}\-\\varphi\_\{u\}\-\\tau\_\{u,v\}\\\|\_\{L^\{2\}\(A\_\{v\}^\{c\}\)\}\\leq C\\\|v\\\|^\{2\},\\qquad\\\|\\varphi\_\{s\(v\)\}\-\\varphi\_\{u\}\-\\tau\_\{u,v\}\\\|\_\{L^\{2\}\(A\_\{v\}\)\}\\leq C\\\|v\\\|^\{3/2\},which proves \([C\.11](https://arxiv.org/html/2608.06766#A3.E11)\)\. The derivatives differ only by the smooth change of tangent frame and by the indicator difference𝟏\{s\(v\)⊤x\>0\}−𝟏\{u⊤x\>0\}\\mathbf\{1\}\_\{\\\{s\(v\)^\{\\top\}x\>0\\\}\}\-\\mathbf\{1\}\_\{\\\{u^\{\\top\}x\>0\\\}\}\. Its squaredL2L^\{2\}operator contribution is bounded by a Gaussian second moment onAvA\_\{v\}, hence byC‖v‖C\\\|v\\\|, proving \([C\.12](https://arxiv.org/html/2608.06766#A3.E12)\)\. ∎
For a state in a winner chart write
q1=1\+α,s1=s\(v\)=Expu\(v\),β=\(q2,…,qm\),q\_\{1\}=1\+\\alpha,\\qquad s\_\{1\}=s\(v\)=\\operatorname\{Exp\}\_\{u\}\(v\),\\qquad\\beta=\(q\_\{2\},\\ldots,q\_\{m\}\),and keep the redundant directions𝐬=\(s2,…,sm\)\\mathbf\{s\}=\(s\_\{2\},\\ldots,s\_\{m\}\)as marked base variables\. The realization residual is
ℛ𝐬\(α,v,β\):=\(1\+α\)φs\(v\)\+S𝐬β−φu\.\\mathcal\{R\}\_\{\\mathbf\{s\}\}\(\\alpha,v,\\beta\):=\(1\+\\alpha\)\\varphi\_\{s\(v\)\}\+S\_\{\\mathbf\{s\}\}\\beta\-\\varphi\_\{u\}\.\(C\.13\)LetJZJ\_\{Z\}denote its derivative with respect to the active coordinates\(α,v,β\)\(\\alpha,v,\\beta\), with the redundant coefficient block interpreted on the quotient\.
###### Lemma C\.8\(Nonlinear quotient normal form\)\.
Assume \([C\.9](https://arxiv.org/html/2608.06766#A3.E9)\) with lower boundλq\\lambda\_\{\\rm q\}\. There is
r0=r0\(λq,m,d\)\>0r\_\{0\}=r\_\{0\}\(\\lambda\_\{\\rm q\},m,d\)\>0such that, wheneverξ=\(α,v,\[β\]\)\\xi=\(\\alpha,v,\[\\beta\]\)satisfies‖ξ‖q≤r0\\\|\\xi\\\|\_\{\\rm q\}\\leq r\_\{0\},
ℛ𝐬\(α,v,β\)\\displaystyle\\mathcal\{R\}\_\{\\mathbf\{s\}\}\(\\alpha,v,\\beta\)=F𝐬ξ\+N\(ξ\),\\displaystyle=F\_\{\\mathbf\{s\}\}\\xi\+N\(\\xi\),‖N\(ξ\)‖ℋ\\displaystyle\\\|N\(\\xi\)\\\|\_\{\\mathcal\{H\}\}≤C‖ξ‖q3/2,\\displaystyle\\leq C\\\|\\xi\\\|\_\{\\rm q\}^\{3/2\},\(C\.14\)‖JZ−F𝐬‖op\\displaystyle\\\|J\_\{Z\}\-F\_\{\\mathbf\{s\}\}\\\|\_\{\\rm op\}≤C‖ξ‖q1/2\.\\displaystyle\\leq C\\\|\\xi\\\|\_\{\\rm q\}^\{1/2\}\.\(C\.15\)Consequently,
λq2‖ξ‖q≤‖ℛ𝐬\(α,v,β\)‖ℋ≤C‖ξ‖q,\\frac\{\\sqrt\{\\lambda\_\{\\rm q\}\}\}\{2\}\\\|\\xi\\\|\_\{\\rm q\}\\leq\\\|\\mathcal\{R\}\_\{\\mathbf\{s\}\}\(\\alpha,v,\\beta\)\\\|\_\{\\mathcal\{H\}\}\\leq C\\\|\\xi\\\|\_\{\\rm q\},\(C\.16\)and
‖JZ∗ℛ𝐬\(α,v,β\)‖2≥cλq‖ℛ𝐬\(α,v,β\)‖ℋ2\.\\\|J\_\{Z\}^\{\*\}\\mathcal\{R\}\_\{\\mathbf\{s\}\}\(\\alpha,v,\\beta\)\\\|^\{2\}\\geq c\\lambda\_\{\\rm q\}\\\|\\mathcal\{R\}\_\{\\mathbf\{s\}\}\(\\alpha,v,\\beta\)\\\|\_\{\\mathcal\{H\}\}^\{2\}\.\(C\.17\)
###### Proof\.
Expanding the winner ridge and using[section˜C\.3](https://arxiv.org/html/2608.06766#A3.SS3.SSS0.Px1)gives
\(1\+α\)φs\(v\)−φu=αφu\+τu,v\+N\(α,v\),\(1\+\\alpha\)\\varphi\_\{s\(v\)\}\-\\varphi\_\{u\}=\\alpha\\varphi\_\{u\}\+\\tau\_\{u,v\}\+N\(\\alpha,v\),with
‖N\(α,v\)‖≤C\(\|α\|‖v‖\+‖v‖3/2\)≤C‖ξ‖q3/2\.\\\|N\(\\alpha,v\)\\\|\\leq C\\bigl\(\|\\alpha\|\\\|v\\\|\+\\\|v\\\|^\{3/2\}\\bigr\)\\leq C\\\|\\xi\\\|\_\{\\rm q\}^\{3/2\}\.The redundant coefficient contribution is exactly linear, proving \([C\.14](https://arxiv.org/html/2608.06766#A3.E14)\)\. Equation \([C\.12](https://arxiv.org/html/2608.06766#A3.E12)\) gives \([C\.15](https://arxiv.org/html/2608.06766#A3.E15)\)\.
Choose the minimum\-norm representative of\[β\]\[\\beta\]\. By \([C\.9](https://arxiv.org/html/2608.06766#A3.E9)\),‖F𝐬ξ‖≥λq‖ξ‖q\\\|F\_\{\\mathbf\{s\}\}\\xi\\\|\\geq\\sqrt\{\\lambda\_\{\\rm q\}\}\\\|\\xi\\\|\_\{\\rm q\}\. Shrinkingr0r\_\{0\}so thatCr01/2≤λq/2Cr\_\{0\}^\{1/2\}\\leq\\sqrt\{\\lambda\_\{\\rm q\}\}/2proves the lower bound in \([C\.16](https://arxiv.org/html/2608.06766#A3.E16)\); the upper bound is immediate\.
For the PL estimate, writer=F𝐬ξ\+Nr=F\_\{\\mathbf\{s\}\}\\xi\+NandJZ=F𝐬\+EJ\_\{Z\}=F\_\{\\mathbf\{s\}\}\+E\. On the quotient, the nonzero spectrum ofF𝐬∗F𝐬F\_\{\\mathbf\{s\}\}^\{\*\}F\_\{\\mathbf\{s\}\}is bounded below byλq\\lambda\_\{\\rm q\}, hence
‖F𝐬∗F𝐬ξ‖≥λq‖ξ‖q\.\\\|F\_\{\\mathbf\{s\}\}^\{\*\}F\_\{\\mathbf\{s\}\}\\xi\\\|\\geq\\lambda\_\{\\rm q\}\\\|\\xi\\\|\_\{\\rm q\}\.The remaining terms satisfy
‖F𝐬∗N‖\+‖E∗r‖≤C‖ξ‖q3/2\.\\\|F\_\{\\mathbf\{s\}\}^\{\*\}N\\\|\+\\\|E^\{\*\}r\\\|\\leq C\\\|\\xi\\\|\_\{\\rm q\}^\{3/2\}\.Reducingr0r\_\{0\}once more gives‖JZ∗r‖≥\(λq/2\)‖ξ‖q\\\|J\_\{Z\}^\{\*\}r\\\|\\geq\(\\lambda\_\{\\rm q\}/2\)\\\|\\xi\\\|\_\{\\rm q\}\. Combining this with the upper bound in \([C\.16](https://arxiv.org/html/2608.06766#A3.E16)\) proves \([C\.17](https://arxiv.org/html/2608.06766#A3.E17)\)\. ∎
###### Lemma C\.9\(Robust local reaction–transport PL\)\.
Supposeq1∈\[1/2,3/2\]q\_\{1\}\\in\[1/2,3/2\], the gauge signs are fixed, the winner has gauge\+D\+D, and all redundant neurons have gauge−D\-D\. Under the hypotheses of[section˜C\.3](https://arxiv.org/html/2608.06766#A3.SS3.SSS0.Px1), for all sufficiently largeDD,
−ℒ˙≥cDℒ\.\-\\dot\{\\mathcal\{L\}\}\\geq cD\\mathcal\{L\}\.\(C\.18\)
###### Proof\.
The coefficient components ofJZ∗ℛJ\_\{Z\}^\{\*\}\\mathcal\{R\}are precisely the reaction gradientsB\(si\)B\(s\_\{i\}\)\. Its winner\-direction component is the geodesic derivativeq1\(I−s1s1⊤\)𝒯\(s1\)q\_\{1\}\(I\-s\_\{1\}s\_\{1\}^\{\\top\}\)\\mathcal\{T\}\(s\_\{1\}\), up to a uniformly conditioned chart identification\. The exact dissipation weights every coefficient gradient byμi≥D\\mu\_\{i\}\\geq D\. For the positive\-gauge winner,
a12q12=μ\(q1,D\)\+D2q12≥cD,\\frac\{a\_\{1\}^\{2\}\}\{q\_\{1\}^\{2\}\}=\\frac\{\\mu\(q\_\{1\},D\)\+D\}\{2q\_\{1\}^\{2\}\}\\geq cD,so the winner directional\-gradient coordinate is also weighted by at leastcDcD\. Dropping the nonnegative redundant\-direction terms and applying \([C\.17](https://arxiv.org/html/2608.06766#A3.E17)\) yields
−ℒ˙≥cD‖JZ∗ℛ‖2≥cD‖ℛ‖2=2cDℒ\.\-\\dot\{\\mathcal\{L\}\}\\geq cD\\\|J\_\{Z\}^\{\*\}\\mathcal\{R\}\\\|^\{2\}\\geq cD\\\|\\mathcal\{R\}\\\|^\{2\}=2cD\\mathcal\{L\}\.∎
###### Lemma C\.10\(Coefficient control and loser transport under an active redundant frame\)\.
Assume, in addition, the redundant frame condition \([C\.10](https://arxiv.org/html/2608.06766#A3.E10)\) with lower boundλred\>0\\lambda\_\{\\rm red\}\>0\. In a sufficiently small winner tube,
∑j=2m\|qj\|2≤Cℒ\.\\sum\_\{j=2\}^\{m\}\|q\_\{j\}\|^\{2\}\\leq C\\mathcal\{L\}\.\(C\.19\)If the trajectory remains in that tube, then
∑j=2m∫t0∞‖s˙j\(t\)‖dt≤CD2ℒ\(t0\)\.\\sum\_\{j=2\}^\{m\}\\int\_\{t\_\{0\}\}^\{\\infty\}\\\|\\dot\{s\}\_\{j\}\(t\)\\\|\\,\\,\\mathrm\{d\}t\\leq\\frac\{C\}\{D^\{2\}\}\\mathcal\{L\}\(t\_\{0\}\)\.\(C\.20\)In particular, every redundant coefficient tends to zero and every redundant direction has finite post\-entry motion\.
###### Proof\.
The full redundant frame gives‖β‖2≤λred−1‖S𝐬β‖2\\\|\\beta\\\|^\{2\}\\leq\\lambda\_\{\\rm red\}^\{\-1\}\\\|S\_\{\\mathbf\{s\}\}\\beta\\\|^\{2\}\. The quotient normal form and transversality separate the winner error from the redundant function, so‖S𝐬β‖2≤C‖ℛ‖2=2Cℒ\\\|S\_\{\\mathbf\{s\}\}\\beta\\\|^\{2\}\\leq C\\\|\\mathcal\{R\}\\\|^\{2\}=2C\\mathcal\{L\}, proving \([C\.19](https://arxiv.org/html/2608.06766#A3.E19)\)\. For a negative gauge,
\|χ\(qj,−D\)\|=2\|qj\|D2\+4qj2\+D≤\|qj\|D\.\|\\chi\(q\_\{j\},\-D\)\|=\\frac\{2\|q\_\{j\}\|\}\{\\sqrt\{D^\{2\}\+4q\_\{j\}^\{2\}\}\+D\}\\leq\\frac\{\|q\_\{j\}\|\}\{D\}\.Moreover‖\(I−sjsj⊤\)𝒯\(sj\)‖≤Cℒ\\\|\(I\-s\_\{j\}s\_\{j\}^\{\\top\}\)\\mathcal\{T\}\(s\_\{j\}\)\\\|\\leq C\\sqrt\{\\mathcal\{L\}\}by Cauchy–Schwarz\. Thus
∑j=2m‖s˙j‖≤CDℒ∑j=2m\|qj\|≤CDℒ\.\\sum\_\{j=2\}^\{m\}\\\|\\dot\{s\}\_\{j\}\\\|\\leq\\frac\{C\}\{D\}\\sqrt\{\\mathcal\{L\}\}\\sum\_\{j=2\}^\{m\}\|q\_\{j\}\|\\leq\\frac\{C\}\{D\}\\mathcal\{L\}\.Integrating \([C\.18](https://arxiv.org/html/2608.06766#A3.E18)\) gives∫t0∞ℒ\(t\)dt≤Cℒ\(t0\)/D\\int\_\{t\_\{0\}\}^\{\\infty\}\\mathcal\{L\}\(t\)\\,\\,\\mathrm\{d\}t\\leq C\\mathcal\{L\}\(t\_\{0\}\)/D, proving \([C\.20](https://arxiv.org/html/2608.06766#A3.E20)\)\. The local decomposition and exponential loss decay implyqj→0q\_\{j\}\\to 0\. ∎
###### Theorem C\.11\(Conditional asymptotic locking under an active frame\)\.
Assume the open\-set finite\-time selection theorem and lett0t\_\{0\}be a selected\-state entry time\. Suppose that att0t\_\{0\}:
1. 1\.the active quotient and redundant coefficient frames satisfyλq≥2λ0\\lambda\_\{\\rm q\}\\geq 2\\lambda\_\{0\}andλred≥2λred,0\\lambda\_\{\\rm red\}\\geq 2\\lambda\_\{\{\\rm red\},0\};
2. 2\.every redundant direction is separated from the teacher by at least2β02\\beta\_\{0\};
3. 3\.the state has a positive margin𝔪0\\mathfrak\{m\}\_\{0\}to the remaining chart, coefficient, and gauge\-sign faces of the winner tube\.
There is a constantCmovC\_\{\\rm mov\}, depending on the fixed tube but not onDD, such that if
CmovD2ℒ\(t0\)≤12𝔪0,\\frac\{C\_\{\\rm mov\}\}\{D^\{2\}\}\\mathcal\{L\}\(t\_\{0\}\)\\leq\\frac\{1\}\{2\}\\mathfrak\{m\}\_\{0\},\(C\.21\)then the tube is forward invariant and
ℒ\(t\)≤ℒ\(t0\)e−cD\(t−t0\),q1\(t\)s1\(t\)→u,qj\(t\)→0\(j≥2\)\.\\mathcal\{L\}\(t\)\\leq\\mathcal\{L\}\(t\_\{0\}\)e^\{\-cD\(t\-t\_\{0\}\)\},\\qquad q\_\{1\}\(t\)s\_\{1\}\(t\)\\to u,\\qquad q\_\{j\}\(t\)\\to 0\\quad\(j\\geq 2\)\.The redundant directions have the total\-motion bound \([C\.20](https://arxiv.org/html/2608.06766#A3.E20)\)\.
###### Proof\.
Lettexitt\_\{\\rm exit\}be the first exit time from the tube\. On\[t0,texit\)\[t\_\{0\},t\_\{\\rm exit\}\),[section˜C\.3](https://arxiv.org/html/2608.06766#A3.SS3.SSS0.Px1)gives exponential loss decay\. The quotient normal form and the redundant frame imply that the winner coordinates and all redundant coefficients shrink, so the corresponding coordinate and coefficient faces cannot be reached\.
The angle and frame functions are Lipschitz on the compact tube\. By[section˜C\.3](https://arxiv.org/html/2608.06766#A3.SS3.SSS0.Px1), the total change of every redundant direction is at mostCmovℒ\(t0\)/D2C\_\{\\rm mov\}\\mathcal\{L\}\(t\_\{0\}\)/D^\{2\}\. Hence the teacher\-angle margins, the quotient\-frame eigenvalue, and the redundant\-frame eigenvalue each lose less than one half of their entry margin under \([C\.21](https://arxiv.org/html/2608.06766#A3.E21)\)\. Gauge signs are exactly conserved in continuous time\. Thus no boundary face can be reached, contradicting a finitetexitt\_\{\\rm exit\}\. Forward invariance and the limiting statements follow from the exponential decay and the local decomposition\. ∎
### C\.4Discrete full\-batch gradient descent
LetΘ=\(a1,w1,…,am,wm\)\\Theta=\(a\_\{1\},w\_\{1\},\\ldots,a\_\{m\},w\_\{m\}\)denote the original parameters and consider explicit full\-batch gradient descent
Θn\+1=Θn−η∇ℒ\(Θn\),η=hD\.\\Theta\_\{n\+1\}=\\Theta\_\{n\}\-\\eta\\nabla\\mathcal\{L\}\(\\Theta\_\{n\}\),\\qquad\\eta=\\frac\{h\}\{D\}\.\(C\.22\)We treat the exact duplicate initialization\. Duplicate negative\-gauge students remain exactly exchangeable under every raw Euler step, so the originalmm\-student iterate reduces without approximation to one winner and one aggregate loser\.
Before capture we retain the regular state
Y=\(q,P,θ,ψ\)\.Y=\(q,P,\\theta,\\psi\)\.\(C\.23\)LetY∞Y\_\{\\infty\}be the multiplicity\-weighted fast flow of Appendix B\. ForTc:=T⋆\+1T\_\{c\}:=T\_\{\\star\}\+1, positivity of the continuous fast solution gives
p¯m:=12min0≤τ≤Tc\{q∞\(τ\),P∞\(τ\)\}\>0\.\\underline\{p\}\_\{m\}:=\\frac\{1\}\{2\}\\min\_\{0\\leq\\tau\\leq T\_\{c\}\}\\\{q\_\{\\infty\}\(\\tau\),P\_\{\\infty\}\(\\tau\)\\\}\>0\.\(C\.24\)After winner\-tube entry we use the nonsingular functionally weighted state
X=\(q,P,qsθ,Psψ\)\.X=\(q,P,qs\_\{\\theta\},Ps\_\{\\psi\}\)\.\(C\.25\)
###### Lemma C\.13\(Raw Euler expansion for one neuron\)\.
Write
𝒯\(s\)=B\(s\)s\+G\(s\),G\(s\)⟂s,ζ:=ηa/r\.\\mathcal\{T\}\(s\)=B\(s\)s\+G\(s\),\\qquad G\(s\)\\perp s,\\qquad\\zeta:=\\eta a/r\.One raw Euler step is
a\+=a−ηrB,w\+=r\(\(1−ζB\)s−ζG\)\.a^\{\+\}=a\-\\eta rB,\\qquad w^\{\+\}=r\\bigl\(\(1\-\\zeta B\)s\-\\zeta G\\bigr\)\.Whenever\|ζ\|\(\|B\|\+‖G‖\)≤1/4\|\\zeta\|\(\|B\|\+\\\|G\\\|\)\\leq 1/4,
r\+\\displaystyle r^\{\+\}=r−ηaB\+η2a22r‖G‖2\+Rr,\\displaystyle=r\-\\eta aB\+\\frac\{\\eta^\{2\}a^\{2\}\}\{2r\}\\\|G\\\|^\{2\}\+R\_\{r\},\(C\.26\)s\+\\displaystyle s^\{\+\}=s−ζG−ζ2\(BG\+12‖G‖2s\)\+Rs,\\displaystyle=s\-\\zeta G\-\\zeta^\{2\}\\left\(BG\+\\frac\{1\}\{2\}\\\|G\\\|^\{2\}s\\right\)\+R\_\{s\},\(C\.27\)q\+\\displaystyle q^\{\+\}=q−η\(a2\+r2\)B\+η2qB2\+η2a32r‖G‖2\+Rq,\\displaystyle=q\-\\eta\(a^\{2\}\+r^\{2\}\)B\+\\eta^\{2\}qB^\{2\}\+\\frac\{\\eta^\{2\}a^\{3\}\}\{2r\}\\\|G\\\|^\{2\}\+R\_\{q\},\(C\.28\)with
\|Rr\|≤Cr\|ζ\|3\(\|B\|\+‖G‖\)3,‖Rs‖≤C\|ζ\|3\(\|B\|\+‖G‖\)3,\|R\_\{r\}\|\\leq Cr\|\\zeta\|^\{3\}\(\|B\|\+\\\|G\\\|\)^\{3\},\\quad\\\|R\_\{s\}\\\|\\leq C\|\\zeta\|^\{3\}\(\|B\|\+\\\|G\\\|\)^\{3\},\|Rq\|≤C\(\|a\|\+r\)r\|ζ\|3\(\|B\|\+‖G‖\)3\.\|R\_\{q\}\|\\leq C\(\|a\|\+r\)r\|\\zeta\|^\{3\}\(\|B\|\+\\\|G\\\|\)^\{3\}\.The gauge drift identity is exact:
δ\+−δ=η2\[\(rB\)2−a2‖𝒯\(s\)‖2\]\.\\delta^\{\+\}\-\\delta=\\eta^\{2\}\\left\[\(rB\)^\{2\}\-a^\{2\}\\\|\\mathcal\{T\}\(s\)\\\|^\{2\}\\right\]\.\(C\.29\)
###### Proof\.
SinceG⟂sG\\perp s,
r\+r=\(1−ζB\)2\+ζ2‖G‖2\.\\frac\{r^\{\+\}\}\{r\}=\\sqrt\{\(1\-\\zeta B\)^\{2\}\+\\zeta^\{2\}\\\|G\\\|^\{2\}\}\.The scalar Taylor expansion of the square root, with an integral remainder on the interval determined by the smallness hypothesis, gives \([C\.26](https://arxiv.org/html/2608.06766#A3.E26)\)\. Multiplying\(\(1−ζB\)s−ζG\)\(\(1\-\\zeta B\)s\-\\zeta G\)by the reciprocal expansion ofr\+/rr^\{\+\}/rgives \([C\.27](https://arxiv.org/html/2608.06766#A3.E27)\)\. Multiplyinga\+a^\{\+\}andr\+r^\{\+\}gives \([C\.28](https://arxiv.org/html/2608.06766#A3.E28)\)\. Finally, expanding\(a\+\)2−‖w\+‖2\(a^\{\+\}\)^\{2\}\-\\\|w^\{\+\}\\\|^\{2\}cancels the complete first\-order term and yields \([C\.29](https://arxiv.org/html/2608.06766#A3.E29)\)\. ∎
###### Lemma C\.14\(Pre\-capture one\-step consistency and chart preservation\)\.
Fixmmand the compact fast\-flow segmentY∞\(\[0,Tc\]\)Y\_\{\\infty\}\(\[0,T\_\{c\}\]\)\. There existh0,Cm\>0h\_\{0\},C\_\{m\}\>0such that, forh≤h0h\\leq h\_\{0\},DDsufficiently large, and every discrete state satisfying
q,P≥p¯m,dist\(Y,Y∞\(\[0,Tc\]\)\)≤1,q,P\\geq\\underline\{p\}\_\{m\},\\qquad\\operatorname\{dist\}\(Y,Y\_\{\\infty\}\(\[0,T\_\{c\}\]\)\)\\leq 1,one raw original\-parameter step remains in the regular chart and
Yn\+1=Yn\+hF∞\(Yn\)\+Rn,‖Rn‖≤Cm\(h2\+hD−2\)\.Y\_\{n\+1\}=Y\_\{n\}\+hF\_\{\\infty\}\(Y\_\{n\}\)\+R\_\{n\},\\qquad\\\|R\_\{n\}\\\|\\leq C\_\{m\}\(h^\{2\}\+hD^\{\-2\}\)\.\(C\.30\)The one\-step gauge drift is bounded byCmh2/DC\_\{m\}h^\{2\}/D\.
###### Proof\.
On the stated compact set the residual momentsB,GB,Gand their first derivatives are uniformly bounded\. For the positive\-gauge winner,a=Θ\(D\)a=\\Theta\(\\sqrt\{D\}\),r=Θ\(D−1/2\)r=\\Theta\(D^\{\-1/2\}\), and henceζ=h/q\+O\(hD−2\)\\zeta=h/q\+O\(hD^\{\-2\}\)\. For a negative\-gauge loser,a=Θ\(D−1/2\)a=\\Theta\(D^\{\-1/2\}\),r=Θ\(D\)r=\\Theta\(\\sqrt\{D\}\), andζ=O\(hD−2\)\\zeta=O\(hD^\{\-2\}\)\. Applying[section˜C\.4](https://arxiv.org/html/2608.06766#A3.SS4)to the winner and to one common loser, and multiplying the loser coefficient increment bym−1m\-1, gives the fast marked vector field\. The second\-order terms areOm\(h2\)O\_\{m\}\(h^\{2\}\)and the finite\-gauge mobility defects areOm\(hD−2\)O\_\{m\}\(hD^\{\-2\}\), which proves \([C\.30](https://arxiv.org/html/2608.06766#A3.E30)\)\. The same estimates imply relative factor changesOm\(h\)O\_\{m\}\(h\), preserving positivity andw≠0w\\neq 0forh0h\_\{0\}small\. Equation \([C\.29](https://arxiv.org/html/2608.06766#A3.E29)\) and the raw factor sizes give the stated gauge\-drift bound\. ∎
###### Lemma C\.15\(Discrete pre\-capture shadowing\)\.
Fornh≤Tcnh\\leq T\_\{c\},
max0≤k≤n‖Yk−Y∞\(kh\)‖≤Cm\(h\+D−2\)\.\\max\_\{0\\leq k\\leq n\}\\\|Y\_\{k\}\-Y\_\{\\infty\}\(kh\)\\\|\\leq C\_\{m\}\(h\+D^\{\-2\}\)\.\(C\.31\)Consequently, for sufficiently smallhhand sufficiently largeDD, the iterates satisfyqk,Pk≥p¯mq\_\{k\},P\_\{k\}\\geq\\underline\{p\}\_\{m\}before capture and enter the winner tube in at most
ncap≤⌈Tch⌉n\_\{\\rm cap\}\\leq\\left\\lceil\\frac\{T\_\{c\}\}\{h\}\\right\\rceilsteps\.
###### Proof\.
The exact fast flow has the local expansionY∞\(\(n\+1\)h\)=Y∞\(nh\)\+hF∞\(Y∞\(nh\)\)\+Om\(h2\)Y\_\{\\infty\}\(\(n\+1\)h\)=Y\_\{\\infty\}\(nh\)\+hF\_\{\\infty\}\(Y\_\{\\infty\}\(nh\)\)\+O\_\{m\}\(h^\{2\}\)\. Subtract this from \([C\.30](https://arxiv.org/html/2608.06766#A3.E30)\) and use the Lipschitz constant ofF∞F\_\{\\infty\}on the compact segment\. A bootstrap discrete Gronwall argument proves \([C\.31](https://arxiv.org/html/2608.06766#A3.E31)\); choosing the right\-hand side smaller thanp¯m\\underline\{p\}\_\{m\}closes the bootstrap and preserves the regular chart\. The strict continuous entry margin then gives discrete winner\-tube entry\. ∎
###### Lemma C\.16\(Raw Hessian bound in the symmetric winner tube\)\.
On the fixed symmetric winner tube there isCm\>0C\_\{m\}\>0, independent ofDD, such that
‖∇2ℒ\(Θ\)‖op≤CmD\.\\\|\\nabla^\{2\}\\mathcal\{L\}\(\\Theta\)\\\|\_\{\\rm op\}\\leq C\_\{m\}D\.\(C\.32\)Moreover, the exact population gradient satisfies
‖∇ℒ\(Θ\)‖2≥cmDℒ\(Θ\)\.\\\|\\nabla\\mathcal\{L\}\(\\Theta\)\\\|^\{2\}\\geq c\_\{m\}D\\mathcal\{L\}\(\\Theta\)\.\(C\.33\)
###### Proof\.
Write the population loss using the homogeneous Gaussian ReLU kernel
K\(w,v\)=‖w‖‖v‖κ\(w^⊤v^\)\.K\(w,v\)=\\\|w\\\|\\\|v\\\|\\kappa\(\\widehat\{w\}^\{\\top\}\\widehat\{v\}\)\.On the nonzero winner tube, direct differentiation of the arc\-cosine formula gives
‖∇wK\(w,v\)‖≤C‖v‖,‖∇wv2K\(w,v\)‖≤C,‖∇ww2K\(w,v\)‖≤C‖v‖‖w‖,\\\|\\nabla\_\{w\}K\(w,v\)\\\|\\leq C\\\|v\\\|,\\qquad\\\|\\nabla^\{2\}\_\{wv\}K\(w,v\)\\\|\\leq C,\\qquad\\\|\\nabla^\{2\}\_\{ww\}K\(w,v\)\\\|\\leq C\\frac\{\\\|v\\\|\}\{\\\|w\\\|\},with the self\-interactionK\(w,w\)=‖w‖2/2K\(w,w\)=\\\|w\\\|^\{2\}/2treated explicitly\. TheaaaaHessian blocks are kernel values and areOm\(D\)O\_\{m\}\(D\)\. Theawawblocks contain one raw factor and one first kernel derivative and areOm\(D\)O\_\{m\}\(D\)\. Thewwwwblocks containaiaja\_\{i\}a\_\{j\}times a second kernel derivative; the positive/negative gauge reconstructions\|ai\|2\+‖wi‖2=D2\+4qi2\|a\_\{i\}\|^\{2\}\+\\\|w\_\{i\}\\\|^\{2\}=\\sqrt\{D^\{2\}\+4q\_\{i\}^\{2\}\}show that all such blocks areOm\(D\)O\_\{m\}\(D\)\. This proves \([C\.32](https://arxiv.org/html/2608.06766#A3.E32)\)\.
The squared raw gradient norm equals the negative loss derivative under gradient flow\. On the symmetric winner tube, the selected/aggregate\-redundant feature Gram of[appendix˜B](https://arxiv.org/html/2608.06766#A2.SSx2)gives the local gradient inequality, while all coefficient mobilities and the selected directional mobility are bounded below bycDcD\. Hence \([C\.33](https://arxiv.org/html/2608.06766#A3.E33)\)\. ∎
###### Lemma C\.17\(Loss\-weighted post\-capture Euler transport\)\.
Let an exact\-duplicate discrete iterate lie in the symmetric winner tube and satisfy the negative\-gauge bootstrap boundδ−,n≤−D/2\\delta\_\{\-,n\}\\leq\-D/2\. Writepn=Pn/\(m−1\)p\_\{n\}=P\_\{n\}/\(m\-1\)for the common loser coefficient and decompose its residual moment as
𝒯n\(sn\)=Bnsn\+Gn,Gn⟂sn\.\\mathcal\{T\}\_\{n\}\(s\_\{n\}\)=B\_\{n\}s\_\{n\}\+G\_\{n\},\\qquad G\_\{n\}\\perp s\_\{n\}\.There is a constantCm\>0C\_\{m\}\>0, independent ofDDandhh, such that
\|pn\|\+\|Bn\|\+‖Gn‖≤Cmℒn\.\|p\_\{n\}\|\+\|B\_\{n\}\|\+\\\|G\_\{n\}\\\|\\leq C\_\{m\}\\sqrt\{\\mathcal\{L\}\_\{n\}\}\.\(C\.34\)Consequently, forη=h/D\\eta=h/Dandhhsufficiently small,
‖sn\+1−sn‖≤Cm\(hD2ℒn\+h2D4ℒn2\+h3D6ℒn3\)\.\\\|s\_\{n\+1\}\-s\_\{n\}\\\|\\leq C\_\{m\}\\left\(\\frac\{h\}\{D^\{2\}\}\\mathcal\{L\}\_\{n\}\+\\frac\{h^\{2\}\}\{D^\{4\}\}\\mathcal\{L\}\_\{n\}^\{2\}\+\\frac\{h^\{3\}\}\{D^\{6\}\}\\mathcal\{L\}\_\{n\}^\{3\}\\right\)\.\(C\.35\)In particular, every term in the one\-step directional remainder vanishes with the task residual\.
###### Proof\.
The symmetric local loss equivalence \([B\.19](https://arxiv.org/html/2608.06766#A2.E19)\) givesPn2≤CmℒnP\_\{n\}^\{2\}\\leq C\_\{m\}\\mathcal\{L\}\_\{n\}and hence\|pn\|≤Cmℒn\|p\_\{n\}\|\\leq C\_\{m\}\\sqrt\{\\mathcal\{L\}\_\{n\}\}for fixedmm\. LetRn=fΘn−f⋆R\_\{n\}=f\_\{\\Theta\_\{n\}\}\-f\_\{\\star\}\. SinceBn=𝔼\[Rn\(X\)\[sn⊤X\]\+\]B\_\{n\}=\\mathbb\{E\}\[R\_\{n\}\(X\)\[s\_\{n\}^\{\\top\}X\]\_\{\+\}\]and𝒯n\(sn\)=𝔼\[Rn\(X\)𝟏\{sn⊤X\>0\}X\]\\mathcal\{T\}\_\{n\}\(s\_\{n\}\)=\\mathbb\{E\}\[R\_\{n\}\(X\)\\mathbf\{1\}\_\{\\\{s\_\{n\}^\{\\top\}X\>0\\\}\}X\], Gaussian Cauchy–Schwarz yields
\|Bn\|≤ℒn,‖𝒯n\(sn\)‖≤2ℒn,\|B\_\{n\}\|\\leq\\sqrt\{\\mathcal\{L\}\_\{n\}\},\\qquad\\\|\\mathcal\{T\}\_\{n\}\(s\_\{n\}\)\\\|\\leq\\sqrt\{2\\mathcal\{L\}\_\{n\}\},up to the fixed normalization ofℒ=12𝔼R2\\mathcal\{L\}=\\frac\{1\}\{2\}\\mathbb\{E\}R^\{2\}\. BecauseGn=\(I−snsn⊤\)𝒯n\(sn\)G\_\{n\}=\(I\-s\_\{n\}s\_\{n\}^\{\\top\}\)\\mathcal\{T\}\_\{n\}\(s\_\{n\}\), this proves \([C\.34](https://arxiv.org/html/2608.06766#A3.E34)\)\.
Under the bootstrap boundδ−,n≤−D/2\\delta\_\{\-,n\}\\leq\-D/2,
rn2=δn2\+4pn2−δn2≥D2\.r\_\{n\}^\{2\}=\\frac\{\\sqrt\{\\delta\_\{n\}^\{2\}\+4p\_\{n\}^\{2\}\}\-\\delta\_\{n\}\}\{2\}\\geq\\frac\{D\}\{2\}\.Usingan/rn=pn/rn2a\_\{n\}/r\_\{n\}=p\_\{n\}/r\_\{n\}^\{2\}andη=h/D\\eta=h/D, we obtain
\|ζn\|=η\|anrn\|≤Ch\|pn\|D2\.\|\\zeta\_\{n\}\|=\\eta\\left\|\\frac\{a\_\{n\}\}\{r\_\{n\}\}\\right\|\\leq C\\frac\{h\|p\_\{n\}\|\}\{D^\{2\}\}\.The fixed winner tube and \([C\.34](https://arxiv.org/html/2608.06766#A3.E34)\) imply\|ζn\|\(\|Bn\|\+‖Gn‖\)≤1/4\|\\zeta\_\{n\}\|\(\|B\_\{n\}\|\+\\\|G\_\{n\}\\\|\)\\leq 1/4after reducingh1\(m\)h\_\{1\}\(m\)if necessary, so[section˜C\.4](https://arxiv.org/html/2608.06766#A3.SS4)applies\. The first\-order angular term is bounded by\|ζn\|‖Gn‖≤Cmhℒn/D2\|\\zeta\_\{n\}\|\\\|G\_\{n\}\\\|\\leq C\_\{m\}h\\mathcal\{L\}\_\{n\}/D^\{2\}\. The explicit quadratic term is at most
\|ζn\|2\(\|Bn\|‖Gn‖\+12‖Gn‖2\)≤Cmh2D4ℒn2\.\|\\zeta\_\{n\}\|^\{2\}\\left\(\|B\_\{n\}\|\\\|G\_\{n\}\\\|\+\\frac\{1\}\{2\}\\\|G\_\{n\}\\\|^\{2\}\\right\)\\leq C\_\{m\}\\frac\{h^\{2\}\}\{D^\{4\}\}\\mathcal\{L\}\_\{n\}^\{2\}\.Finally, the cubic remainder satisfies
‖Rs,n‖≤C\|ζn\|3\(\|Bn\|\+‖Gn‖\)3≤Cmh3D6ℒn3\.\\\|R\_\{s,n\}\\\|\\leq C\|\\zeta\_\{n\}\|^\{3\}\(\|B\_\{n\}\|\+\\\|G\_\{n\}\\\|\)^\{3\}\\leq C\_\{m\}\\frac\{h^\{3\}\}\{D^\{6\}\}\\mathcal\{L\}\_\{n\}^\{3\}\.Summing these three estimates proves \([C\.35](https://arxiv.org/html/2608.06766#A3.E35)\)\. ∎
###### Lemma C\.18\(Discrete descent, first\-exit invariance, and gauge drift\)\.
There ish1\(m\)\>0h\_\{1\}\(m\)\>0such that, if0<h≤h1\(m\)0<h\\leq h\_\{1\}\(m\)and a discrete iterate enters the symmetric winner tube with the fixed continuous\-theory entry margin, then for every later iterate
ℒn\+1≤\(1−cmh\)ℒn\.\\mathcal\{L\}\_\{n\+1\}\\leq\(1\-c\_\{m\}h\)\\mathcal\{L\}\_\{n\}\.\(C\.36\)The tube is forward invariant, and the common loser direction obeys the explicit path\-length bound
∑n≥ncap‖sn\+1−sn‖≤CmD2\(ℒncap\+hD2ℒncap2\+h2D4ℒncap3\)\.\\sum\_\{n\\geq n\_\{\\rm cap\}\}\\\|s\_\{n\+1\}\-s\_\{n\}\\\|\\leq\\frac\{C\_\{m\}\}\{D^\{2\}\}\\left\(\\mathcal\{L\}\_\{n\_\{\\rm cap\}\}\+\\frac\{h\}\{D^\{2\}\}\\mathcal\{L\}\_\{n\_\{\\rm cap\}\}^\{2\}\+\\frac\{h^\{2\}\}\{D^\{4\}\}\\mathcal\{L\}\_\{n\_\{\\rm cap\}\}^\{3\}\\right\)\.\(C\.37\)Moreover,
supn\|δi,n−δi,0\|≤CmhD\.\\sup\_\{n\}\|\\delta\_\{i,n\}\-\\delta\_\{i,0\}\|\\leq C\_\{m\}\\frac\{h\}\{D\}\.\(C\.38\)
###### Proof\.
Letnexit\>ncapn\_\{\\rm exit\}\>n\_\{\\rm cap\}be the first putative exit index\. For everyncap≤n<nexitn\_\{\\rm cap\}\\leq n<n\_\{\\rm exit\}, the descent lemma, \([C\.32](https://arxiv.org/html/2608.06766#A3.E32)\), \([C\.33](https://arxiv.org/html/2608.06766#A3.E33)\), andη=h/D\\eta=h/Dgive \([C\.36](https://arxiv.org/html/2608.06766#A3.E36)\)\. The symmetric local loss equivalence therefore keeps the winner coefficient, aggregate loser coefficient, and winner angle strictly inside their tube faces\.
Apply[section˜C\.4](https://arxiv.org/html/2608.06766#A3.SS4)up to indexnexit−1n\_\{\\rm exit\}\-1\. Writingϖ=1−cmh\\varpi=1\-c\_\{m\}h, geometric contraction gives, forℓ=1,2,3\\ell=1,2,3,
∑k=0nexit−1−ncapℒncap\+kℓ≤ℒncapℓ1−ϖℓ≤Cm,ℓhℒncapℓ\.\\sum\_\{k=0\}^\{n\_\{\\rm exit\}\-1\-n\_\{\\rm cap\}\}\\mathcal\{L\}\_\{n\_\{\\rm cap\}\+k\}^\{\\,\\ell\}\\leq\\frac\{\\mathcal\{L\}\_\{n\_\{\\rm cap\}\}^\{\\,\\ell\}\}\{1\-\\varpi^\{\\ell\}\}\\leq\\frac\{C\_\{m,\\ell\}\}\{h\}\\mathcal\{L\}\_\{n\_\{\\rm cap\}\}^\{\\,\\ell\}\.Substitution into \([C\.35](https://arxiv.org/html/2608.06766#A3.E35)\) yields the finite\-horizon version of \([C\.37](https://arxiv.org/html/2608.06766#A3.E37)\), uniformly innexitn\_\{\\rm exit\}\. ForDDsufficiently large andhhsufficiently small, this quantity is less than one half of the fixed loser\-angle entry margin\. Hence the loser\-angle face cannot be the first exit\.
It remains to exclude loss of the gauge\-magnitude bootstrapδ\+,n≥D/2\\delta\_\{\+,n\}\\geq D/2andδ−,n≤−D/2\\delta\_\{\-,n\}\\leq\-D/2\. The pre\-capture drift isOm\(h/D\)O\_\{m\}\(h/D\)by[section˜C\.4](https://arxiv.org/html/2608.06766#A3.SS4)overO\(h−1\)O\(h^\{\-1\}\)steps\. Up to the putative exit, the exact identity \([C\.29](https://arxiv.org/html/2608.06766#A3.E29)\) and the descent lemma imply
∑n=ncapnexit−1\|δi,n\+1−δi,n\|≤η2∑n=ncapnexit−1‖∇iℒ\(Θn\)‖2≤2ηℒncap≤CmhD\.\\sum\_\{n=n\_\{\\rm cap\}\}^\{n\_\{\\rm exit\}\-1\}\|\\delta\_\{i,n\+1\}\-\\delta\_\{i,n\}\|\\leq\\eta^\{2\}\\sum\_\{n=n\_\{\\rm cap\}\}^\{n\_\{\\rm exit\}\-1\}\\\|\\nabla\_\{i\}\\mathcal\{L\}\(\\Theta\_\{n\}\)\\\|^\{2\}\\leq 2\\eta\\mathcal\{L\}\_\{n\_\{\\rm cap\}\}\\leq C\_\{m\}\\frac\{h\}\{D\}\.This is negligible compared with the entry gauge magnitudeDD, so no gauge\-magnitude, gauge\-sign, or regular\-chart face can be reached\. All possible first exit faces have now been excluded, which proves forward invariance\. Lettingnexit→∞n\_\{\\rm exit\}\\to\\inftygives \([C\.37](https://arxiv.org/html/2608.06766#A3.E37)\); combining the pre\- and post\-capture estimates gives \([C\.38](https://arxiv.org/html/2608.06766#A3.E38)\)\. ∎
###### Theorem C\.19\(Discrete\-GD capture and functional pruning\)\.
For every fixedm≥2m\\geq 2andϕ∈Iϕ\\phi\\in I\_\{\\phi\}, there existh0\(m\)\>0h\_\{0\}\(m\)\>0andD~0\(m\)<∞\\widetilde\{D\}\_\{0\}\(m\)<\\inftysuch that the exact duplicate full\-batch gradient\-descent iterates withη=h/D\\eta=h/D,0<h≤h0\(m\)0<h\\leq h\_\{0\}\(m\), andD≥D~0\(m\)D\\geq\\widetilde\{D\}\_\{0\}\(m\)select the programmed positive\-gauge winner\. They obey the pre\-capture shadowing estimate \([C\.31](https://arxiv.org/html/2608.06766#A3.E31)\), enter the winner tube within⌈Tc/h⌉\\lceil T\_\{c\}/h\\rceiliterations, satisfy the geometric contraction \([C\.36](https://arxiv.org/html/2608.06766#A3.E36)\), have gauge drift \([C\.38](https://arxiv.org/html/2608.06766#A3.E38)\), and satisfy the explicit post\-capture redundant\-direction bound \([C\.37](https://arxiv.org/html/2608.06766#A3.E37)\)\.
###### Proof\.
Combine[sections˜C\.4](https://arxiv.org/html/2608.06766#A3.SS4),[C\.4](https://arxiv.org/html/2608.06766#A3.SS4),[C\.4](https://arxiv.org/html/2608.06766#A3.SS4)and[C\.4](https://arxiv.org/html/2608.06766#A3.SS4)\. ∎
## Appendix DExperimental and numerical validation
The validation compares the following objects without refitting any dynamics to the finite systems\.
Fast theory\.The three\-dimensional multiplicity\-weighted limitz∞=\(q,P,θ\)z\_\{\\infty\}=\(q,P,\\theta\)from the global\-capture theorem, whereqqis the selected functional coefficient,PPis the aggregate redundant coefficient, andθ\\thetais the selected angle\.
Exact marked population flow\.The finite\-DDsymmetric systemzD=\(q,P,θ,ψ\)z\_\{D\}=\(q,P,\\theta,\\psi\), including the slowly moving redundant directionψ\\psi\.
Original\-parameter population flow\.The fullmm\-student gradient flow in\(ai,wi\)\(a\_\{i\},w\_\{i\}\)coordinates\. No symmetry reduction is imposed by the implementation\.
Original\-parameter empirical flow\.Full\-batch gradient flow onNNi\.i\.d\. samplesxn∼𝒩\(0,I2\)x\_\{n\}\\sim\\mathcal\{N\}\(0,I\_\{2\}\)with labelsyn=\[u⊤xn\]\+y\_\{n\}=\[u^\{\\top\}x\_\{n\}\]\_\{\+\}, again in the full\(ai,wi\)\(a\_\{i\},w\_\{i\}\)coordinates\.
The global\-capture theorem predicts, for every fixed multiplicity,
tcap\\displaystyle t\_\{\\rm cap\}=Θm\(D−1\),\\displaystyle=\\Theta\_\{m\}\(D^\{\-1\}\),\(D\.1\)∫tcap∞‖s˙lose\(t\)‖dt\\displaystyle\\int\_\{t\_\{\\rm cap\}\}^\{\\infty\}\\\!\\\|\\dot\{s\}\_\{\\rm lose\}\(t\)\\\|\\,\\,\\mathrm\{d\}t=Om\(D−2\),\\displaystyle=O\_\{m\}\(D^\{\-2\}\),supτ≤T‖zD\(τ\)−z∞\(τ\)‖\\displaystyle\\sup\_\{\\tau\\leq T\}\\\|z\_\{D\}\(\\tau\)\-z\_\{\\infty\}\(\\tau\)\\\|=Om,T\(D−2\)\.\\displaystyle=O\_\{m,T\}\(D^\{\-2\}\)\.
The population sweep uses
m∈\{2,4,8,16,32\},D∈\{8,16,32,64,128\}\.m\\in\\\{2,4,8,16,32\\\},\\qquad D\\in\\\{8,16,32,64,128\\\}\.The empirical sweep fixesD=32D=32and uses
m∈\{2,4,8,16\},N∈\{512,1024,2048,4096,8192\},20coupled seeds\.m\\in\\\{2,4,8,16\\\},\\quad N\\in\\\{512,1024,2048,4096,8192\\\},\\quad 20\\text\{ coupled seeds\}\.For a fixed\(m,seed\)\(m,\\text\{seed\}\), the samples are nested acrossNN\. The operational capture tube is fixed before all sweeps:
qwin≥0\.90,\|Plose\|≤0\.10,\|θwin\|≤0\.10\.q\_\{\\mathrm\{win\}\}\\geq 0\.90,\\qquad\|P\_\{\\mathrm\{lose\}\}\|\\leq 0\.10,\\qquad\|\\theta\_\{\\mathrm\{win\}\}\|\\leq 0\.10\.\(D\.2\)
A raw redundant angle is not identifiable after its aggregate functional coefficient vanishes\. We therefore compare finite\-sample and population trajectories in the functionally identifiable state
𝖷\(t\)=\(q\(t\),P\(t\),q\(t\)swin\(t\),P\(t\)slose\(t\)\)∈ℝ6,\\mathsf\{X\}\(t\)=\\bigl\(q\(t\),\\,P\(t\),\\,q\(t\)s\_\{\\mathrm\{win\}\}\(t\),\\,P\(t\)s\_\{\\mathrm\{lose\}\}\(t\)\\bigr\)\\in\\mathbb\{R\}^\{6\},\(D\.3\)using
ErrN=\[16T∫0T‖𝖷N\(τ\)−𝖷pop\(τ\)‖22dτ\]1/2\.\\operatorname\{Err\}\_\{N\}=\\left\[\\frac\{1\}\{6T\}\\int\_\{0\}^\{T\}\\\|\\mathsf\{X\}\_\{N\}\(\\tau\)\-\\mathsf\{X\}\_\{\\mathrm\{pop\}\}\(\\tau\)\\\|\_\{2\}^\{2\}\\,\\,\\mathrm\{d\}\\tau\\right\]^\{1/2\}\.\(D\.4\)This prevents an arbitrary direction of a nearly zero\-coefficient redundant unit from dominating the validation metric\.
### D\.1Complete trajectory correspondence
[Figure˜2](https://arxiv.org/html/2608.06766#S4.F2)compares all four levels for the representative setting\(m,D,N\)=\(8,32,8192\)\(m,D,N\)=\(8,32,8192\)\. The curves agree in task loss, winner capture, aggregate functional\-mass decay, feature transport, and the reaction–transport dissipation split\. The fast theory omits theO\(D−2\)O\(D^\{\-2\}\)redundant\-direction motion while accurately tracking every functionally active coordinate\.
The exact marked and original\-parameter population implementations agree over the full population grid to a maximum marked\-state error below1\.3×10−71\.3\\times 10^\{\-7\}; population gauge drift remains below2\.2×10−112\.2\\times 10^\{\-11\}\. A separate full\-mmversus symmetric\-manifold audit at\(m,N,D\)=\(8,2048,32\)\(m,N,D\)=\(8,2048,32\)finds state and loss differences at machine precision\. Duplicate redundant units remain identical, as required by permutation symmetry\.
### D\.2Population scaling across gauge and overparameterization
Panels \(a\)–\(b\) of[figure˜3](https://arxiv.org/html/2608.06766#S4.F3), panel \(b\) of[figure˜D\.1](https://arxiv.org/html/2608.06766#A4.F1), and[table˜D\.1](https://arxiv.org/html/2608.06766#A4.T1)audit the three asymptotic predictions in[equation˜D\.1](https://arxiv.org/html/2608.06766#A4.E1)\. The fitted exponents remain stable fromm=2m=2throughm=32m=32\. The constants change withmm, as the theorem permits, but do not blow up over the tested fixed\-multiplicity range\.
Table D\.1:Population scaling laws\. All regressions useD∈\{8,16,32,64,128\}D\\in\\\{8,16,32,64,128\\\}\.The rescaled capture constantsDtcapDt\_\{\\mathrm\{cap\}\}are30\.24,24\.16,21\.84,20\.82,20\.3430\.24,24\.16,21\.84,20\.82,20\.34form=2,4,8,16,32m=2,4,8,16,32\. The correspondingD2D^\{2\}\-rescaled loser drifts are1\.59×10−2,3\.78×10−3,1\.44×10−3,6\.40×10−4,3\.03×10−41\.59\\times 10^\{\-2\},3\.78\\times 10^\{\-3\},1\.44\\times 10^\{\-3\},6\.40\\times 10^\{\-4\},3\.03\\times 10^\{\-4\}\. These observations do not constitute a uniformm→∞m\\to\\inftyresult; they show only that the fixed\-mmtheorem remains numerically well conditioned over the tested range\.
### D\.3Finite\-sample trajectory scaling
For eachmm, the mean functionally identifiable trajectory error follows a clean power law overN=512N=512to81928192; see[table˜D\.2](https://arxiv.org/html/2608.06766#A4.T2)\. The slopes range from−0\.448\-0\.448to−0\.494\-0\.494, withR2R^\{2\}between0\.9590\.959and0\.9970\.997\. Every one of the4×5×20=4004\\times 5\\times 20=400empirical runs enters the fixed capture tube \([D\.2](https://arxiv.org/html/2608.06766#A4.E2)\)\. Empirical gauge drift remains below9\.7×10−79\.7\\times 10^\{\-7\}\.
Table D\.2:Finite\-sample full\-trajectory scaling atD=32D=32using 20 coupled seeds\. The error is defined in \([D\.4](https://arxiv.org/html/2608.06766#A4.E4)\)\.The observed behavior is compatible with the standard fixed\-dimensional Monte Carlo scaleN−1/2N^\{\-1/2\}\. This is an empirical guide\. After capture, the population loss is extremely small and dominated by cancellations, so the identifiable reaction–transport state provides the stable trajectory metric at fixed finitemm\.
A solver audit compares three adaptive\-tolerance settings\. Relative to the tightest solve, the two looser settings have maximum marked\-state errors3\.27×10−53\.27\\times 10^\{\-5\}and1\.64×10−51\.64\\times 10^\{\-5\}, below the empirical trajectory discrepancies shown in[figure˜3](https://arxiv.org/html/2608.06766#S4.F3)\. The numerical integration error therefore does not explain the finite\-sample scaling\.
### D\.4Discrete post\-capture transport audit
The proof of[sections˜C\.4](https://arxiv.org/html/2608.06766#A3.SS4)and[C\.4](https://arxiv.org/html/2608.06766#A3.SS4)retains the residual\-dependent factors in every Euler remainder\. We independently audit the resultingD−2D^\{\-2\}path\-length prediction in the raw original\-parameter GD implementation\. Atm=8m=8andϕ=π/4\\phi=\\pi/4, we use
D∈\{16,24,32,48,64,96,128\},h∈\{0\.2,0\.1,0\.05\},η=h/D,D\\in\\\{16,24,32,48,64,96,128\\\},\\qquad h\\in\\\{0\.2,0\.1,0\.05\\\},\\qquad\\eta=h/D,and sum the common loser\-direction increments after the fixed capture condition \([D\.2](https://arxiv.org/html/2608.06766#A4.E2)\)\. The trajectory is followed to fast timeτ=80\\tau=80; the remaining loss is below1\.6×10−71\.6\\times 10^\{\-7\}in every run\.
Table D\.3:Raw discrete\-GD post\-capture loser path\. The fitted exponent is for total directional motion versusDD\.TheD2D^\{2\}\-rescaled path varies by less than0\.4%0\.4\\%across the testedDDvalues for each fixedhh\. This audit is not used in the proof; it checks that the raw Euler implementation exhibits the summable, loss\-weighted behavior required by \([C\.35](https://arxiv.org/html/2608.06766#A3.E35)\)\.
Figure D\.1:Additional numerical diagnostics\.\(a\) Both single\-unit singular limits converge asD−2D^\{\-2\}\. \(b\) The finite\-gauge marked flow approaches the fast theory asD−2D^\{\-2\}\. \(c\) Every certified\-angle test captures\. \(d\)DtcapDt\_\{\\rm cap\}remains controlled throughm=256m=256\(a finite sweep only\)\. \(e\) Discrete\-GD error is first order inh=ηDh=\\eta D\. \(f\) The raw post\-capture path followsD−2D^\{\-2\}at every testedhh\. Dashed lines show predicted orders\.Similar Articles
Bug or Feature^2: Weight Drift, Activation Sparsity, and Spikes
This paper formally proves that training neural networks with asymmetric activation functions like ReLU, GELU, or SiLU causes weights to drift negative, leading to up to 90% activation sparsity. It also shows that squared activations like ReLU² improve performance but cause activation spikes, which can be fixed by clipping, with GELU² achieving the best validation loss.
Revisiting the Volume Hypothesis
This paper revisits the volume hypothesis, which posits that generalization in over-parameterized networks is mainly due to the larger volume of good-generalizing regions in weight space rather than SGD's implicit bias. Through experiments with binary networks, the authors show that the generalization advantage of gradient learning over random sampling diminishes as training data size grows, potentially resolving contradictory prior findings.
Unlocking Feature Learning in Gated Delta Networks at Scale
This paper derives scaling rules for Gated Delta Networks using Maximal Update Parametrization (μP), enabling zero-shot hyperparameter transfer across model widths for efficient sub-quadratic LLM architectures. Experiments confirm stable learning-rate transfer under both AdamW and SGD, whereas standard parametrization fails.
Generalized Neurons
The article explores the Universal Approximation Theorem in deep learning, analyzing the representation capacity of individual neurons and neural network layers using ReLU activation functions.
The Boolean Power of ReLU
This theoretical paper proves that ReLU-based message-passing GNNs are strictly more expressive than GNNs using any eventually constant activation functions (e.g., truncated ReLU) with respect to Boolean queries, even on Boolean-featured graphs.