Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks
Summary
This paper demonstrates that optimal learning rate scaling in deep scalar linear networks is inherently data-dependent, contradicting prior data-agnostic scaling rules. It shows that with data-dependent scaling, convergence becomes depth-independent, including at infinite depth.
View Cached Full Text
Cached at: 07/10/26, 06:16 AM
# Optimal learning rate scaling depends on data in deep scalar linear networks
Source: [https://arxiv.org/html/2607.07884](https://arxiv.org/html/2607.07884)
Yedi Zhang1Peter E\. Latham1Leena Chennuru Vankadara1Andrew Saxe1,2 1Gatsby Computational Neuroscience Unit, University College London 2Sainsbury Wellcome Centre, University College London \{yedi,pel\}@gatsby\.ucl\.ac\.uk,\{l\.vankadara,a\.saxe\}@ucl\.ac\.uk
###### Abstract
In this short note we consider the gradient descent dynamics of deep scalar linear networks,f\(x\)=∏l=1Lwlxf\(x\)=\\prod\_\{l=1\}^\{L\}w\_\{l\}x, which enjoy exact time\-course solutions for any integer depth\. We show that even in this minimal model, the optimal depth\-wise learning rate scaling depends on data, whereas data\-agnostic scaling rules fail to transfer across depths\. Under the data\-dependent optimal scaling, the learning dynamics is independent of data and weakly dependent on depth, resulting in a constant linear convergence rate across all depths including infinity\. We further show similar data\-dependent effects in deep scalar linear networks with residual connections\.
## 1Introduction
The large scale of modern neural networks has been empirically shown to play a crucial role in the rapid progress of deep learning models\(Kaplanet al\.,[2020](https://arxiv.org/html/2607.07884#bib.bib19); Hoffmannet al\.,[2022](https://arxiv.org/html/2607.07884#bib.bib20)\)\. One essential factor of the scale is depth\. Thus, understanding how to enable hyperparameter transfer across depth is critical for achieving predictable gains from scale\. While existing literature on hyperparameter transfer suggests that data\-agnostic learning rate scaling can allow depth\-wise transfer\(Yanget al\.,[2024](https://arxiv.org/html/2607.07884#bib.bib49); Everettet al\.,[2024](https://arxiv.org/html/2607.07884#bib.bib48); Nociet al\.,[2024](https://arxiv.org/html/2607.07884#bib.bib47); Bordelonet al\.,[2024b](https://arxiv.org/html/2607.07884#bib.bib46);[a](https://arxiv.org/html/2607.07884#bib.bib45); Bordelon and Pehlevan,[2025](https://arxiv.org/html/2607.07884#bib.bib53)\), we demonstrate that even in a minimal model class of deep scalar linear networks, the optimal learning rate scaling is inherently data\-dependent\. Specifically, the learning rate scaling rule contains a data\-dependent correction that cannot be captured by a power law with a data\-agnostic exponent, even when the constant factor is calibrated with data at one depth\. We show that the learning rate scaling that accounts for this data\-dependent exponent transfers across depth, whereas data\-agnostic scaling does not\.
We consider the simplest possible deep network, a depth\-LLscalar linear chain defined as
f\(x;w\)=∏l=1Lwlx,x,w1,⋯,wL∈ℝ\.\\displaystyle f\(x;w\)=\\prod\_\{l=1\}^\{L\}w\_\{l\}x,\\quad x,w\_\{1\},\\cdots,w\_\{L\}\\in\\mathbb\{R\}\.\(1\)Building on and refining analyses in prior work\(Saxeet al\.,[2014](https://arxiv.org/html/2607.07884#bib.bib4);[2019](https://arxiv.org/html/2607.07884#bib.bib15)\), we write exact solutions to the full gradient descent learning dynamics for any integer depth, expressed via special functions, i\.e\. the hypergeometric function and the LambertWWfunction\. Under the data\-dependent optimal learning rate scaling and a balanced initialization scheme, the gradient descent learning dynamics is independent of data and weakly dependent on depth\. This results in a constant linear convergence rate across all depths, including the limiting case of infinite depth\. Further, we extend the analysis to deep scalar linear residual networks with block depth one and two, and find that the optimal learning rate scaling for them is also data\-dependent\.
Related work\.Jelassiet al\.\([2023](https://arxiv.org/html/2607.07884#bib.bib41)\)found that in deep ReLU networks with mean\-field initialization, the largest learning rate for which the changes in the pre\-activations after one gradient descent step remains bounded scales with depthLLasL−3/2L^\{\-3/2\}\.Bordelonet al\.\([2024b](https://arxiv.org/html/2607.07884#bib.bib46)\); Bordelon and Pehlevan \([2025](https://arxiv.org/html/2607.07884#bib.bib53)\)obtained reduced learning dynamics and studied hyperparameters transfer in infinite\-depth linear residual networks in early training time, where the width and depth limits commute\(Hayou and Yang,[2023](https://arxiv.org/html/2607.07884#bib.bib39)\)\.Deyet al\.\([2025](https://arxiv.org/html/2607.07884#bib.bib54)\)demonstrated that deep residual networks withL−1/2L^\{\-1/2\}scaling can achieve hyperparameter transfer but operate in a locally lazy learning regime, while aL−1L^\{\-1\}scaling enables rich learning and depth\-wise hyperparameter transfer\. Complementing these findings, we use a simple model class of deep scalar linear networks to demonstrate that the optimal learning rate scaling is data\-dependent\.
The learning dynamics of deep linear networks enjoy a rich line of theoretical results\(Baldi and Hornik,[1989](https://arxiv.org/html/2607.07884#bib.bib1); Fukumizu,[1998](https://arxiv.org/html/2607.07884#bib.bib2); Saxeet al\.,[2014](https://arxiv.org/html/2607.07884#bib.bib4);[2019](https://arxiv.org/html/2607.07884#bib.bib15); Aroraet al\.,[2018](https://arxiv.org/html/2607.07884#bib.bib25); Shamir,[2019](https://arxiv.org/html/2607.07884#bib.bib16); Lampinen and Ganguli,[2019](https://arxiv.org/html/2607.07884#bib.bib11); Gidelet al\.,[2019](https://arxiv.org/html/2607.07884#bib.bib17); Advaniet al\.,[2020](https://arxiv.org/html/2607.07884#bib.bib21); Huh,[2020](https://arxiv.org/html/2607.07884#bib.bib33); Gissinet al\.,[2020](https://arxiv.org/html/2607.07884#bib.bib18); Tarmounet al\.,[2021](https://arxiv.org/html/2607.07884#bib.bib22); Atanasovet al\.,[2022](https://arxiv.org/html/2607.07884#bib.bib30); Braunet al\.,[2022](https://arxiv.org/html/2607.07884#bib.bib23); Shiet al\.,[2022](https://arxiv.org/html/2607.07884#bib.bib29); Zhanget al\.,[2024](https://arxiv.org/html/2607.07884#bib.bib42);[2026](https://arxiv.org/html/2607.07884#bib.bib43); Dominéet al\.,[2025](https://arxiv.org/html/2607.07884#bib.bib51); Xu and Ziyin,[2025](https://arxiv.org/html/2607.07884#bib.bib52); Watanabeet al\.,[2026](https://arxiv.org/html/2607.07884#bib.bib57)\)\.Saxeet al\.\([2014](https://arxiv.org/html/2607.07884#bib.bib4);[2019](https://arxiv.org/html/2607.07884#bib.bib15)\)solved the learning dynamics of deep linear networks with aligned small initial weights and white input covariance, showing that depth slows down learning in the case of learning withℓ2\\ell\_\{2\}loss and the infinite\-depth network incurs a finite decay in learning speed relative to the shallow network\. Here we build on and refine these results by incorporating input correlations, expressing solutions via special functions, and choosing variables that reveal data\-independent learning dynamics, and connect these results to the modern literature on hyperparameter transfer\.
## 2Learning dynamics with the maximal stable learning rate
Let\{xn,yn\}n=1N\\\{x\_\{n\},y\_\{n\}\\\}\_\{n=1\}^\{N\}be a training set\. The gradient flow dynamics of the depth\-LLlinear chain in[Equation1](https://arxiv.org/html/2607.07884#S1.E1)trained withℓ2\\ell\_\{2\}loss,ℒ=12N∑n=1N\(y−f\(x\)\)2\\mathcal\{L\}=\\frac\{1\}\{2N\}\\sum\_\{n=1\}^\{N\}\(y\-f\(x\)\)^\{2\}, is given by
w˙l=−η∂ℒ∂wl=η\(μyx−μxx∏i=1Lwi\)∏i≠lwi,\\displaystyle\\dot\{w\}\_\{l\}=\-\\eta\\frac\{\\partial\\mathcal\{L\}\}\{\\partial w\_\{l\}\}=\\eta\\left\(\\mu\_\{yx\}\-\\mu\_\{xx\}\\prod\_\{i=1\}^\{L\}w\_\{i\}\\right\)\\prod\_\{i\\neq l\}w\_\{i\},\(2\)whereμyx=1N∑n=1Nynxn,μxx=1N∑n=1Nxn2\\mu\_\{yx\}=\\frac\{1\}\{N\}\\sum\_\{n=1\}^\{N\}y\_\{n\}x\_\{n\},\\mu\_\{xx\}=\\frac\{1\}\{N\}\\sum\_\{n=1\}^\{N\}x\_\{n\}^\{2\}are moments of the dataset, andη\\etarepresents the learning rate\. The continuous\-time gradient flow dynamics captures the behaviors of gradient descent under a stable learning rate where the dynamics does not diverge or sustainedly oscillate\(Cohenet al\.,[2025](https://arxiv.org/html/2607.07884#bib.bib50)\)\. We consider the stable regime of gradient descent learning in this paper\. Here we present a self\-contained exposition that builds onSaxeet al\.\([2014](https://arxiv.org/html/2607.07884#bib.bib4);[2019](https://arxiv.org/html/2607.07884#bib.bib15)\), incorporating several refinements\.
The dynamics in[Equation2](https://arxiv.org/html/2607.07884#S2.E2)admits a well\-known conservation law\(Fukumizu,[1998](https://arxiv.org/html/2607.07884#bib.bib2); Saxeet al\.,[2014](https://arxiv.org/html/2607.07884#bib.bib4); Duet al\.,[2018](https://arxiv.org/html/2607.07884#bib.bib27)\)between any pairs of weights,ddt\(wl2−wl′2\)=0\\frac\{d\}\{dt\}\\left\(w\_\{l\}^\{2\}\-w\_\{l^\{\\prime\}\}^\{2\}\\right\)=0\. If we assume all initial weights are positive111If the signs of the initial weights of each layer differ, the weights of some of the layers may change sign during learning\. TheLL\-dimensional dynamics is still constrained by the conservation law to evolve on a one\-dimensional manifold, but we cannot write a one\-dimensional differential equation to capture the full dynamics without using piecewise functions\.,wl\(0\)\>0∀lw\_\{l\}\(0\)\>0\\,\\forall l, we can use the conservation law to reduce theLL\-dimensional dynamics in[Equation2](https://arxiv.org/html/2607.07884#S2.E2)to a one\-dimensional ordinary differential equation aboutw1w\_\{1\}
w˙1=η\(μyx−μxx∏i=1Lw12\+ci\)∏i=2Lw12\+ci,wherecl=wl\(0\)2−w1\(0\)2\.\\displaystyle\\dot\{w\}\_\{1\}=\\eta\\left\(\\mu\_\{yx\}\-\\mu\_\{xx\}\\prod\_\{i=1\}^\{L\}\\sqrt\{w\_\{1\}^\{2\}\+c\_\{i\}\}\\right\)\\prod\_\{i=2\}^\{L\}\\sqrt\{w\_\{1\}^\{2\}\+c\_\{i\}\},\\quad\\text\{where \}c\_\{l\}=w\_\{l\}\(0\)^\{2\}\-w\_\{1\}\(0\)^\{2\}\.\(3\)We further assume that the initial weights of all layers are equal,wl\(0\)=w1\(0\)∀lw\_\{l\}\(0\)=w\_\{1\}\(0\)\\,\\forall l\. This is motivated by the fact that we typically want all layers to participate in learning in a balanced way\. Due to the conservation law, the weights that are initialized equal will remain equal throughout training\. With the equal initial weight assumption, the dynamics in[Equation3](https://arxiv.org/html/2607.07884#S2.E3)simplifies to
w˙1=η\(μyx−μxxw1L\)w1L−1\.\\displaystyle\\dot\{w\}\_\{1\}=\\eta\\left\(\\mu\_\{yx\}\-\\mu\_\{xx\}w\_\{1\}^\{L\}\\right\)w\_\{1\}^\{L\-1\}\.\(4\)
Figure 1:The loss landscape of a scalar linear network has a sharper global minimum as the depthLLincreases, requiring a smaller learning rate for stable gradient descent dynamics\. \(A\) Plot of the loss functionℒ\(w\)=\(1−wL\)2/2\\mathcal\{L\}\(w\)=\\left\(1\-w^\{L\}\\right\)^\{2\}/2with differentLL\. \(B\) Gradient descent trajectory of the total weightwLw^\{L\}using learning rates that scale as[Equation5](https://arxiv.org/html/2607.07884#S2.E5)withτ=1\\tau=1\. The dynamics is stable\. \(C\) Same as panel B but withτ=0\.5\\tau=0\.5, which is the threshold for stable gradient descent dynamics\. The dynamics exhibits oscillations\.Maximum stable learning rate\. When we increase the depth of the linear network, the global minimum of the loss landscape becomes sharper, as shown in[Figure1](https://arxiv.org/html/2607.07884#S2.F1)A\. The second\-order derivative, i\.e\. the sharpness, at the global minimum isS=μxxL\(μyxμxx\)2−2/LS=\\mu\_\{xx\}L\\left\(\\frac\{\\mu\_\{yx\}\}\{\\mu\_\{xx\}\}\\right\)^\{2\-2/L\}\. A sharper minimum requires a smaller learning rate for gradient descent dynamics to be stable\. In particular, the learning rate should satisfy:0<η<2/S0<\\eta<2/S\. Hence, the maximum stable learning rate scales as
η=τ−11μxxL\(μyxμxx\)−2\+2/L,\\displaystyle\\eta=\\tau^\{\-1\}\\frac\{1\}\{\\mu\_\{xx\}L\}\\left\(\\frac\{\\mu\_\{yx\}\}\{\\mu\_\{xx\}\}\\right\)^\{\-2\+2/L\},\(5\)whereτ∈\(0\.5,∞\)\\tau\\in\(0\.5,\\infty\)is the time constant\. The gradient descent dynamics exhibits oscillations whenτ≤0\.5\\tau\\leq 0\.5, as shown in[Figure1](https://arxiv.org/html/2607.07884#S2.F1)C\.[Equation5](https://arxiv.org/html/2607.07884#S2.E5)shows that even in this minimal setup, the scaling of the maximum stable learning rate is data\-dependent, with implications on hyperparameter transfer that we will examine in[Section3](https://arxiv.org/html/2607.07884#S3)\.
Dynamics with maximum stable learning rate\. We now analyze the gradient flow dynamics with the maximum stable learning rate\. We are interested in how the total weight evolves to approach the target weight over training\. We thus study the dynamics of the ratio between the total weight and the target weight,α\(t\)=w1\(t\)Lμxx/μyx\\alpha\(t\)=w\_\{1\}\(t\)^\{L\}\\mu\_\{xx\}/\\mu\_\{yx\}, which evolves as
τα˙=α2−2/L\(1−α\)\.\\displaystyle\\tau\\dot\{\\alpha\}=\\alpha^\{2\-2/L\}\\left\(1\-\\alpha\\right\)\.\(6\)By using the maximum stable learning rate scaling and tracking the relative total weight rather than weights of individual layers, we obtain an ordinary differential equation \([6](https://arxiv.org/html/2607.07884#S2.E6)\) that is independent of the data statistics and weakly dependent on the depthLL, i\.e\. through the factorα2−2/L\\alpha^\{2\-2/L\}\. The dependence onLLweakens asLLincreases, sincelimL→∞α2−2/L=α2\\lim\_\{L\\to\\infty\}\\alpha^\{2\-2/L\}=\\alpha^\{2\}\.
Exact time\-course solution\.[Equation6](https://arxiv.org/html/2607.07884#S2.E6)is a separable differential equation\. By separating variables and integrating both sides, we obtain the solution ofttin terms ofα\\alphafor any positive integer depthLL
t=τα2L−12L−1F12\(1,2L−1;2L;α\)\|α\(0\)α\(t\),\\displaystyle t=\\tau\\frac\{\\alpha^\{\\frac\{2\}\{L\}\-1\}\}\{\\frac\{2\}\{L\}\-1\}\{\}\_\{2\}F\_\{1\}\\left\(1,\\frac\{2\}\{L\}\-1;\\frac\{2\}\{L\};\\alpha\\right\)\\bigg\|\_\{\\alpha\(0\)\}^\{\\alpha\(t\)\},\(7\)whereF12\{\}\_\{2\}F\_\{1\}is the hypergeometric function\. For a general integerLL, we cannot invert[Equation7](https://arxiv.org/html/2607.07884#S2.E7)to solveα\\alphain terms of timettdue to the intractability of the hypergeometric function as a special function\. However, we can invert[Equation7](https://arxiv.org/html/2607.07884#S2.E7)for several specific depths,L=1,2,∞L=1,2,\\infty, in which the hypergeometric function reduces to elementary functions\. Specifically, in the limit of infinite\-depthL→∞L\\to\\infty, the dynamics ofα\\alphais given by
τα˙=α2\(1−α\)\.\\displaystyle\\tau\\dot\{\\alpha\}=\\alpha^\{2\}\\left\(1\-\\alpha\\right\)\.\(8\)The solution to[Equation8](https://arxiv.org/html/2607.07884#S2.E8)can be expressed as
α\(t\)=11\+W0\(eβ\(t\)\),whereβ\(t\)=−tτ\+1α0\+ln\(1α0−1\)−1,0<α≤1\.\\displaystyle\\alpha\(t\)=\\frac\{1\}\{1\+W\_\{0\}\(e^\{\\beta\(t\)\}\)\},\\quad\\text\{where \}\\beta\(t\)=\-\\frac\{t\}\{\\tau\}\+\\frac\{1\}\{\\alpha\_\{0\}\}\+\\ln\\left\(\\frac\{1\}\{\\alpha\_\{0\}\}\-1\\right\)\-1,\\quad 0<\\alpha\\leq 1\.\(9\)HereW0\(⋅\)W\_\{0\}\(\\cdot\)is the principal branch of the LambertWWfunction\. That is,y=W0\(x\)y=W\_\{0\}\(x\)is the solution to the equationyey=xye^\{y\}=xwithx≥0x\\geq 0\. We provide the derivation for[Equation9](https://arxiv.org/html/2607.07884#S2.E9)in[SectionB\.4](https://arxiv.org/html/2607.07884#A2.SS4)\.
Initial plateau\. For any depthLL, the dynamics in[Equation6](https://arxiv.org/html/2607.07884#S2.E6)has a stable fixed point at the global minimum,α=1\\alpha=1\. For deep networks,L≥2L\\geq 2, the network has an unstable fixed point atα=0\\alpha=0in addition to the global minimum\. If small initialization, the typical choice for feature learning\(Woodworthet al\.,[2020](https://arxiv.org/html/2607.07884#bib.bib32)\), is used, the learning dynamics exhibits an initial plateau\(Saxeet al\.,[2014](https://arxiv.org/html/2607.07884#bib.bib4);[2019](https://arxiv.org/html/2607.07884#bib.bib15)\), corresponding to slow escape from the zero fixed point\. The duration of the initial plateauTTis approximately
T=τln1α\(0\),forL=2;T=τ\(1−2L\)α\(0\)1−2L,forL≥3\.\\displaystyle T=\\tau\\ln\\frac\{1\}\{\\alpha\(0\)\}\\,,\\text\{ for \}L=2;\\quad\\quad T=\\frac\{\\tau\}\{\(1\-\\frac\{2\}\{L\}\)\\alpha\(0\)^\{1\-\\frac\{2\}\{L\}\}\}\\,,\\text\{ for \}L\\geq 3\.\(10\)The plateau durationTTincreases when the depthLLincreases and when the initializationα\(0\)\\alpha\(0\)decreases, as shown in[Figure4](https://arxiv.org/html/2607.07884#A2.F4)\. However, the plateau duration in the infinite\-depth scalar linear network remains finite, with an upper bound ofT<τ/α\(0\)T<\\tau/\\alpha\(0\)\.
Convergence rate\. At the end of learning, the linear scalar networks withL≥2L\\geq 2all converge to the global minimum at a linear rate, according to the dynamics in[Equation6](https://arxiv.org/html/2607.07884#S2.E6)\. That is, the total weightα\\alphaisϵ\{\\epsilon\}\-close to the global minimum, i\.e\.\|1−α\|<ϵ\|1\-\\alpha\|<\{\\epsilon\}, after a time of orderτln1ϵ\\tau\\ln\\frac\{1\}\{\{\\epsilon\}\}\.
## 3Depth\-wise learning rate transfer
For the scalar linear chain, the depthwise learning rate scaling rule in[Equation5](https://arxiv.org/html/2607.07884#S2.E5)can be written as
ηL=τ−11μxxLr−2\+2/L,wherer=μyxμxx\.\\displaystyle\\eta\_\{L\}=\\tau^\{\-1\}\\frac\{1\}\{\\mu\_\{xx\}L\}r^\{\-2\+2/L\},\\quad\\text\{where \}r=\\frac\{\\mu\_\{yx\}\}\{\\mu\_\{xx\}\}\.\(11\)This is not a data\-agnostic power law inLL: the factorL−1L^\{\-1\}is accompanied by a finite\-depth correction that depends on the data statisticsrr\. To make this explicit, suppose the learning rate is tuned at a source depthL0L\_\{0\}\. Transferring the learning rate from a source depthL0L\_\{0\}to a target depthLLrequires
ηLηL0=L0Lr2\(1/L−1/L0\)\.\\displaystyle\\frac\{\\eta\_\{L\}\}\{\\eta\_\{L\_\{0\}\}\}=\\frac\{L\_\{0\}\}\{L\}r^\{2\(1/L\-1/L\_\{0\}\)\}\.\(12\)The first factor,L0/LL\_\{0\}/L, is a data\-agnostic scalar constant, while the second factor,r2\(1/L−1/L0\),r^\{2\(1/L\-1/L\_\{0\}\)\},is data\-dependent\. Hence, a data\-agnostic rule such asηL/ηL0=L0/L\\eta\_\{L\}/\\eta\_\{L\_\{0\}\}=L\_\{0\}/L, even with a constant calibrated at the source depth, does not in general reproduce the finite\-depth transfer rule\. If one nevertheless fits the transfer rule betweenL0L\_\{0\}andLLwith a power lawηL/ηL0=\(L/L0\)−peff\\eta\_\{L\}/\\eta\_\{L\_\{0\}\}=\(L/L\_\{0\}\)^\{\-p\_\{\\mathrm\{eff\}\}\}, then
peff\(L,L0;r\)=1−2\(1/L−1/L0\)lnrln\(L/L0\)\.\\displaystyle p\_\{\\mathrm\{eff\}\}\(L,L\_\{0\};r\)=1\-\\frac\{2\(1/L\-1/L\_\{0\}\)\\ln r\}\{\\ln\(L/L\_\{0\}\)\}\.\(13\)Here the effective exponentpeffp\_\{\\mathrm\{eff\}\}depends on the data throughrr\.
Although our theory derives the maximal stable learning rate, the finite\-horizon optimality is a related but different criterion\. To address this gap, we empirically evaluate the finite\-horizon optimality by sweeping learning rates and recording the training loss after a fixed number of gradient descent steps\. In[Figure2](https://arxiv.org/html/2607.07884#S3.F2)A, we show the training loss after a fixed number of gradient descent steps when the learning rate is scaled as the data\-dependent rule in[Equation5](https://arxiv.org/html/2607.07884#S2.E5)versus a data\-agnostic power\-law ruleη∝L−1\\eta\\propto L^\{\-1\}\. As shown in[Figure2](https://arxiv.org/html/2607.07884#S3.F2)A, the learning rates transfer from a shallower network to deep networks under the data\-dependent scaling, but do not transfer under the data\-agnostic scaling\. Specifically, the learning rate withL−1L^\{\-1\}scaling for a very deep network is too small whenμyx/μxx<1\\mu\_\{yx\}/\\mu\_\{xx\}<1, and too large whenμyx/μxx\>1\\mu\_\{yx\}/\\mu\_\{xx\}\>1\.
We further extend the analysis to deep scalar linear residual networks with block depth one\(Bordelonet al\.,[2024b](https://arxiv.org/html/2607.07884#bib.bib46); Yanget al\.,[2024](https://arxiv.org/html/2607.07884#bib.bib49); Marionet al\.,[2025](https://arxiv.org/html/2607.07884#bib.bib44)\)and block depth two\(Bordelonet al\.,[2024a](https://arxiv.org/html/2607.07884#bib.bib45); Deyet al\.,[2025](https://arxiv.org/html/2607.07884#bib.bib54)\), defined as
fblock1\(x;w\)=∏l=1L\(1\+wlL\)x,fblock2\(x;w\)=∏l=1L\(1\+wl2L\)x\.\\displaystyle f\_\{\\text\{block1\}\}\(x;w\)=\\prod\_\{l=1\}^\{L\}\\left\(1\+\\frac\{w\_\{l\}\}\{\\sqrt\{L\}\}\\right\)x,\\qquad\\quad f\_\{\\text\{block2\}\}\(x;w\)=\\prod\_\{l=1\}^\{L\}\\left\(1\+\\frac\{w\_\{l\}^\{2\}\}\{L\}\\right\)x\.\(14\)Their optimal learning rate scaling rules are calculated in[Equations31](https://arxiv.org/html/2607.07884#A3.E31)and[37](https://arxiv.org/html/2607.07884#A4.E37), which are also data\-dependent\. Similar to deep scalar linear networks, the learning rates transfer under the optimal data\-dependent scaling, but does not under data\-agnostic scaling, as shown in[Figure2](https://arxiv.org/html/2607.07884#S3.F2)B,C\.
On the flip side, we note that the data dependence of optimal learning rate scaling is weak for largeLL\. Hence, transferring the optimal learning rate from an intermediate depth to large depth underL−1L^\{\-1\}scaling may still suffice despite being suboptimal, whereas transferring the learning rate from a small depth \(e\.g\.L=2,4L=2,4\) to large depth would likely fail, as we can see from[Figure2](https://arxiv.org/html/2607.07884#S3.F2)\.
In summary, we study depth\-wise learning rate scaling in deep scalar linear networks, with and without residual connections, and find that the scaling rule for transfer is data\-dependent\.
Figure 2:Learning rates transfer under the optimal data\-dependent scaling \(left two columns\), but not under data\-agnostic scaling \(right two columns\)\. The optimal scaling for deep scalar linear networks, linear residual networks with block depth one and two are given by[Equations5](https://arxiv.org/html/2607.07884#S2.E5),[31](https://arxiv.org/html/2607.07884#A3.E31)and[37](https://arxiv.org/html/2607.07884#A4.E37); the relevant data\-agnostic scaling isη∝L−1,1,L\\eta\\propto L^\{\-1\},1,L, respectively\. The loss values are the training loss after 30 steps of gradient descent\. The initial weight is set towl\(0\)L=0\.1w\_\{l\}\(0\)^\{L\}=0\.1in linear networks, andwl\(0\)=0\.01w\_\{l\}\(0\)=0\.01in linear residual networks\. In linear networks, we fixwl\(0\)Lw\_\{l\}\(0\)^\{L\}rather thanwl\(0\)w\_\{l\}\(0\)across different depths because fixingwl\(0\)w\_\{l\}\(0\)leads to vanishing gradients at initialization in the limit of large depth\. In linear residual networks, a constantwl\(0\)w\_\{l\}\(0\)suffices to avoid vanishing or exploding gradients at initialization\.
## Acknowledgments
We thank Kevin Han Huang for feedback on a draft of this paper\. We thank the following funding sources: Gatsby Charitable Foundation \(GAT4058\) to YZ, PEL, LCV, and AS; Sainsbury Wellcome Centre Core Grant from Wellcome \(219627/Z/19/Z\) to AS; Schmidt Science Polymath Award to AS\. AS is a CIFAR Azrieli Global Scholar in the Learning in Machines & Brains program\.
## References
- High\-dimensional dynamics of generalization error in neural networks\.Neural Networks132,pp\. 428–446\.External Links:ISSN 0893\-6080,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.neunet.2020.08.022),[Link](https://www.sciencedirect.com/science/article/pii/S0893608020303117)Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p4.1)\.
- B\. Amos and J\. Z\. Kolter \(2017\)OptNet: differentiable optimization as a layer in neural networks\.InProceedings of the 34th International Conference on Machine Learning,D\. Precup and Y\. W\. Teh \(Eds\.\),Proceedings of Machine Learning Research, Vol\.70,pp\. 136–145\.External Links:[Link](https://proceedings.mlr.press/v70/amos17a.html)Cited by:[Appendix A](https://arxiv.org/html/2607.07884#A1.p1.1)\.
- S\. Arora, N\. Cohen, and E\. Hazan \(2018\)On the optimization of deep networks: implicit acceleration by overparameterization\.InProceedings of the 35th International Conference on Machine Learning,J\. Dy and A\. Krause \(Eds\.\),Proceedings of Machine Learning Research, Vol\.80,pp\. 244–253\.External Links:[Link](https://proceedings.mlr.press/v80/arora18a.html)Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p4.1)\.
- A\. Atanasov, B\. Bordelon, and C\. Pehlevan \(2022\)Neural networks as kernel learners: the silent alignment effect\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=1NvflqAdoom)Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p4.1)\.
- S\. Bai, J\. Z\. Kolter, and V\. Koltun \(2019\)Deep equilibrium models\.InAdvances in Neural Information Processing Systems,H\. Wallach, H\. Larochelle, A\. Beygelzimer, F\. d'Alché\-Buc, E\. Fox, and R\. Garnett \(Eds\.\),Vol\.32,pp\.\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2019/file/01386bd6d8e091c2ab4c7c7de644d37b-Paper.pdf)Cited by:[Appendix A](https://arxiv.org/html/2607.07884#A1.p1.1)\.
- P\. Baldi and K\. Hornik \(1989\)Neural networks and principal component analysis: learning from examples without local minima\.Neural Networks2\(1\),pp\. 53–58\.External Links:ISSN 0893\-6080,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/0893-6080%2889%2990014-2),[Link](https://www.sciencedirect.com/science/article/pii/0893608089900142)Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p4.1)\.
- E\. Boix\-Adsera \(2025\)On the inductive bias of infinite\-depth resnets and the bottleneck rank\.External Links:2501\.19149,[Link](https://arxiv.org/abs/2501.19149)Cited by:[Appendix A](https://arxiv.org/html/2607.07884#A1.p1.1)\.
- B\. Bordelon, H\. Chaudhry, and C\. Pehlevan \(2024a\)Infinite limits of multi\-head transformer dynamics\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 35824–35878\.External Links:[Document](https://dx.doi.org/10.52202/079017-1130),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/3eff068e195daace49955348de9f8398-Paper-Conference.pdf)Cited by:[Appendix A](https://arxiv.org/html/2607.07884#A1.p1.1),[Appendix D](https://arxiv.org/html/2607.07884#A4.p1.4),[§1](https://arxiv.org/html/2607.07884#S1.p1.1),[§3](https://arxiv.org/html/2607.07884#S3.p3.1)\.
- B\. Bordelon, L\. Noci, M\. B\. Li, B\. Hanin, and C\. Pehlevan \(2024b\)Depthwise hyperparameter transfer in residual networks: dynamics and scaling limit\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=KZJehvRKGD)Cited by:[Appendix C](https://arxiv.org/html/2607.07884#A3.p1.2),[§1](https://arxiv.org/html/2607.07884#S1.p1.1),[§1](https://arxiv.org/html/2607.07884#S1.p3.4),[§3](https://arxiv.org/html/2607.07884#S3.p3.1)\.
- B\. Bordelon and C\. Pehlevan \(2025\)Deep linear network training dynamics from random initialization: data, width, depth, and hyperparameter transfer\.InProceedings of the 42nd International Conference on Machine Learning,A\. Singh, M\. Fazel, D\. Hsu, S\. Lacoste\-Julien, F\. Berkenkamp, T\. Maharaj, K\. Wagstaff, and J\. Zhu \(Eds\.\),Proceedings of Machine Learning Research, Vol\.267,pp\. 4968–4997\.External Links:[Link](https://proceedings.mlr.press/v267/bordelon25a.html)Cited by:[Appendix A](https://arxiv.org/html/2607.07884#A1.p1.1),[§1](https://arxiv.org/html/2607.07884#S1.p1.1),[§1](https://arxiv.org/html/2607.07884#S1.p3.4)\.
- L\. Braun, C\. Dominé, J\. Fitzgerald, and A\. Saxe \(2022\)Exact learning dynamics of deep linear networks with prior knowledge\.InAdvances in Neural Information Processing Systems,S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(Eds\.\),Vol\.35,pp\. 6615–6629\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/2b3bb2c95195130977a51b3bb251c40a-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p4.1)\.
- R\. T\. Q\. Chen, Y\. Rubanova, J\. Bettencourt, and D\. K\. Duvenaud \(2018\)Neural ordinary differential equations\.InAdvances in Neural Information Processing Systems,S\. Bengio, H\. Wallach, H\. Larochelle, K\. Grauman, N\. Cesa\-Bianchi, and R\. Garnett \(Eds\.\),Vol\.31,pp\.\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2018/file/69386f6bb1dfed68692a24c8686939b9-Paper.pdf)Cited by:[Appendix A](https://arxiv.org/html/2607.07884#A1.p1.1)\.
- L\. Chizat \(2025\)The hidden width of deep resnets: tight error bounds and phase diagrams\.External Links:2509\.10167,[Link](https://arxiv.org/abs/2509.10167)Cited by:[Appendix A](https://arxiv.org/html/2607.07884#A1.p1.1)\.
- J\. Cohen, A\. Damian, A\. Talwalkar, J\. Z\. Kolter, and J\. D\. Lee \(2025\)Understanding optimization in deep learning with central flows\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=sIE2rI3ZPs)Cited by:[§2](https://arxiv.org/html/2607.07884#S2.p1.6)\.
- N\. S\. Dey, B\. C\. Zhang, L\. Noci, M\. Li, B\. Bordelon, S\. Bergsma, C\. Pehlevan, B\. Hanin, and J\. Hestness \(2025\)Don’t be lazy: completep enables compute\-efficient deep transformers\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=lMU2kaMANl)Cited by:[Appendix D](https://arxiv.org/html/2607.07884#A4.p1.4),[§1](https://arxiv.org/html/2607.07884#S1.p3.4),[§3](https://arxiv.org/html/2607.07884#S3.p3.1)\.
- C\. C\. J\. Dominé, N\. Anguita, A\. M\. Proca, L\. Braun, D\. Kunin, P\. A\. M\. Mediano, and A\. M\. Saxe \(2025\)From lazy to rich: exact learning dynamics in deep linear networks\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=ZXaocmXc6d)Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p4.1)\.
- S\. S\. Du, W\. Hu, and J\. D\. Lee \(2018\)Algorithmic regularization in learning deep homogeneous models: layers are automatically balanced\.InAdvances in Neural Information Processing Systems,S\. Bengio, H\. Wallach, H\. Larochelle, K\. Grauman, N\. Cesa\-Bianchi, and R\. Garnett \(Eds\.\),Vol\.31,pp\.\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2018/file/fe131d7f5a6b38b23cc967316c13dae2-Paper.pdf)Cited by:[§2](https://arxiv.org/html/2607.07884#S2.p2.4)\.
- K\. E\. Everett, L\. Xiao, M\. Wortsman, A\. A\. Alemi, R\. Novak, P\. J\. Liu, I\. Gur, J\. Sohl\-Dickstein, L\. P\. Kaelbling, J\. Lee, and J\. Pennington \(2024\)Scaling exponents across parameterizations and optimizers\.InProceedings of the 41st International Conference on Machine Learning,R\. Salakhutdinov, Z\. Kolter, K\. Heller, A\. Weller, N\. Oliver, J\. Scarlett, and F\. Berkenkamp \(Eds\.\),Proceedings of Machine Learning Research, Vol\.235,pp\. 12666–12700\.External Links:[Link](https://proceedings.mlr.press/v235/everett24a.html)Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p1.1)\.
- K\. Fukumizu \(1998\)Effect of batch learning in multilayer neural networks\.Gen1\(04\),pp\. 1E–03\.Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p4.1),[§2](https://arxiv.org/html/2607.07884#S2.p2.4)\.
- G\. Gidel, F\. Bach, and S\. Lacoste\-Julien \(2019\)Implicit regularization of discrete gradient dynamics in linear neural networks\.InAdvances in Neural Information Processing Systems,H\. Wallach, H\. Larochelle, A\. Beygelzimer, F\. d'Alché\-Buc, E\. Fox, and R\. Garnett \(Eds\.\),Vol\.32,pp\.\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2019/file/f39ae9ff3a81f499230c4126e01f421b-Paper.pdf)Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p4.1)\.
- D\. Gissin, S\. Shalev\-Shwartz, and A\. Daniely \(2020\)The implicit bias of depth: how incremental learning drives generalization\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=H1lj0nNFwB)Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p4.1)\.
- E\. Haber and L\. Ruthotto \(2017\)Stable architectures for deep neural networks\.Inverse Problems34\(1\),pp\. 014004\.External Links:[Document](https://dx.doi.org/10.1088/1361-6420/aa9a90),[Link](https://doi.org/10.1088/1361-6420/aa9a90)Cited by:[Appendix A](https://arxiv.org/html/2607.07884#A1.p1.1)\.
- B\. Hanin \(2019\)Universal function approximation by deep neural nets with bounded width and relu activations\.Mathematics7\(10\)\.External Links:[Link](https://www.mdpi.com/2227-7390/7/10/992),ISSN 2227\-7390,[Document](https://dx.doi.org/10.3390/math7100992)Cited by:[Appendix A](https://arxiv.org/html/2607.07884#A1.p1.1)\.
- S\. Hayou and G\. Yang \(2023\)Width and depth limits commute in residual networks\.InProceedings of the 40th International Conference on Machine Learning,A\. Krause, E\. Brunskill, K\. Cho, B\. Engelhardt, S\. Sabato, and J\. Scarlett \(Eds\.\),Proceedings of Machine Learning Research, Vol\.202,pp\. 12700–12723\.External Links:[Link](https://proceedings.mlr.press/v202/hayou23a.html)Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p3.4)\.
- S\. Hayou \(2023\)On the infinite\-depth limit of finite\-width neural networks\.Transactions on Machine Learning Research\.Note:External Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=RbLsYz1Az9)Cited by:[Appendix A](https://arxiv.org/html/2607.07884#A1.p1.1)\.
- J\. Hoffmann, S\. Borgeaud, A\. Mensch, E\. Buchatskaya, T\. Cai, E\. Rutherford, D\. de Las Casas, L\. A\. Hendricks, J\. Welbl, A\. Clark, T\. Hennigan, E\. Noland, K\. Millican, G\. van den Driessche, B\. Damoc, A\. Guy, S\. Osindero, K\. Simonyan, E\. Elsen, O\. Vinyals, J\. Rae, and L\. Sifre \(2022\)An empirical analysis of compute\-optimal large language model training\.InAdvances in Neural Information Processing Systems,S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(Eds\.\),Vol\.35,pp\. 30016–30030\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/c1e2faff6f588870935f114ebe04a3e5-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p1.1)\.
- D\. Huh \(2020\)Curvature\-corrected learning dynamics in deep neural networks\.InProceedings of the 37th International Conference on Machine Learning,H\. D\. III and A\. Singh \(Eds\.\),Proceedings of Machine Learning Research, Vol\.119,pp\. 4552–4560\.External Links:[Link](https://proceedings.mlr.press/v119/huh20a.html)Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p4.1)\.
- A\. Jacot \(2023\)Implicit bias of large depth networks: a notion of rank for nonlinear functions\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=6iDHce-0B-a)Cited by:[Appendix A](https://arxiv.org/html/2607.07884#A1.p1.1)\.
- S\. Jelassi, B\. Hanin, Z\. Ji, S\. J\. Reddi, S\. Bhojanapalli, and S\. Kumar \(2023\)Depth dependence ofμ\\mup learning rates in relu mlps\.External Links:2305\.07810,[Link](https://arxiv.org/abs/2305.07810)Cited by:[Appendix A](https://arxiv.org/html/2607.07884#A1.p1.1),[§1](https://arxiv.org/html/2607.07884#S1.p3.4)\.
- J\. Kaplan, S\. McCandlish, T\. Henighan, T\. B\. Brown, B\. Chess, R\. Child, S\. Gray, A\. Radford, J\. Wu, and D\. Amodei \(2020\)Scaling laws for neural language models\.External Links:2001\.08361,[Link](https://arxiv.org/abs/2001.08361)Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p1.1)\.
- A\. K\. Lampinen and S\. Ganguli \(2019\)An analytic theory of generalization dynamics and transfer learning in deep linear networks\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=ryfMLoCqtQ)Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p4.1)\.
- P\. Marion, A\. Fermanian, G\. Biau, and J\. Vert \(2025\)Scaling resnets in the large\-depth regime\.Journal of Machine Learning Research26\(56\),pp\. 1–48\.External Links:[Link](http://jmlr.org/papers/v26/22-0664.html)Cited by:[§3](https://arxiv.org/html/2607.07884#S3.p3.1)\.
- L\. Noci, C\. Li, M\. Li, B\. He, T\. Hofmann, C\. J\. Maddison, and D\. Roy \(2023\)The shaped transformer: attention models in the infinite depth\-and\-width limit\.InAdvances in Neural Information Processing Systems,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),Vol\.36,pp\. 54250–54281\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/file/aa31dc84098add7dd2ffdd20646f2043-Paper-Conference.pdf)Cited by:[Appendix A](https://arxiv.org/html/2607.07884#A1.p1.1)\.
- L\. Noci, A\. Meterez, T\. Hofmann, and A\. Orvieto \(2024\)Super consistency of neural network landscapes and learning rate transfer\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 102696–102743\.External Links:[Document](https://dx.doi.org/10.52202/079017-3262),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/ba1d33849b963efc6b5d3082ad68f480-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p1.1)\.
- B\. Poole, S\. Lahiri, M\. Raghu, J\. Sohl\-Dickstein, and S\. Ganguli \(2016\)Exponential expressivity in deep neural networks through transient chaos\.InAdvances in Neural Information Processing Systems,D\. Lee, M\. Sugiyama, U\. Luxburg, I\. Guyon, and R\. Garnett \(Eds\.\),Vol\.29,pp\.\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2016/file/148510031349642de5ca0c544f31b2ef-Paper.pdf)Cited by:[Appendix A](https://arxiv.org/html/2607.07884#A1.p1.1)\.
- A\. M\. Saxe, J\. L\. McClelland, and S\. Ganguli \(2014\)Exact solutions to the nonlinear dynamics of learning in deep linear neural networks\.InInternational Conference on Learning Representations,External Links:[Link](https://arxiv.org/abs/1312.6120)Cited by:[Appendix A](https://arxiv.org/html/2607.07884#A1.p1.1),[§B\.2](https://arxiv.org/html/2607.07884#A2.SS2.p1.8),[§B\.3](https://arxiv.org/html/2607.07884#A2.SS3.p1.3),[§B\.4](https://arxiv.org/html/2607.07884#A2.SS4.p1.7),[§1](https://arxiv.org/html/2607.07884#S1.p2.2),[§1](https://arxiv.org/html/2607.07884#S1.p4.1),[§2](https://arxiv.org/html/2607.07884#S2.p1.6),[§2](https://arxiv.org/html/2607.07884#S2.p2.4),[§2](https://arxiv.org/html/2607.07884#S2.p6.5)\.
- A\. M\. Saxe, J\. L\. McClelland, and S\. Ganguli \(2019\)A mathematical theory of semantic development in deep neural networks\.Proceedings of the National Academy of Sciences116\(23\),pp\. 11537–11546\.External Links:[Document](https://dx.doi.org/10.1073/pnas.1820226116),[Link](https://www.pnas.org/doi/abs/10.1073/pnas.1820226116),https://www\.pnas\.org/doi/pdf/10\.1073/pnas\.1820226116Cited by:[Appendix A](https://arxiv.org/html/2607.07884#A1.p1.1),[§1](https://arxiv.org/html/2607.07884#S1.p2.2),[§1](https://arxiv.org/html/2607.07884#S1.p4.1),[§2](https://arxiv.org/html/2607.07884#S2.p1.6),[§2](https://arxiv.org/html/2607.07884#S2.p6.5)\.
- S\. S\. Schoenholz, J\. Gilmer, S\. Ganguli, and J\. Sohl\-Dickstein \(2017\)Deep information propagation\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=H1W1UN9gg)Cited by:[Appendix A](https://arxiv.org/html/2607.07884#A1.p1.1)\.
- O\. Shamir \(2019\)Exponential convergence time of gradient descent for one\-dimensional deep linear neural networks\.InProceedings of the Thirty\-Second Conference on Learning Theory,A\. Beygelzimer and D\. Hsu \(Eds\.\),Proceedings of Machine Learning Research, Vol\.99,pp\. 2691–2713\.External Links:[Link](https://proceedings.mlr.press/v99/shamir19a.html)Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p4.1)\.
- J\. Shi, E\. Shea\-Brown, and M\. Buice \(2022\)Learning dynamics of deep linear networks with multiple pathways\.InAdvances in Neural Information Processing Systems,S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(Eds\.\),Vol\.35,pp\. 34064–34076\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/dc3ca8bcd613e43ce540352b58d55d6d-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p4.1)\.
- S\. Tarmoun, G\. Franca, B\. D\. Haeffele, and R\. Vidal \(2021\)Understanding the dynamics of gradient flow in overparameterized linear models\.InProceedings of the 38th International Conference on Machine Learning,M\. Meila and T\. Zhang \(Eds\.\),Proceedings of Machine Learning Research, Vol\.139,pp\. 10153–10161\.External Links:[Link](https://proceedings.mlr.press/v139/tarmoun21a.html)Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p4.1)\.
- T\. Watanabe, R\. Karakida, and J\. Teramae \(2026\)The impact of anisotropic covariance structure on the training dynamics and generalization error of linear networks\.External Links:2601\.06961,[Link](https://arxiv.org/abs/2601.06961)Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p4.1)\.
- B\. Woodworth, S\. Gunasekar, J\. D\. Lee, E\. Moroshko, P\. Savarese, I\. Golan, D\. Soudry, and N\. Srebro \(2020\)Kernel and rich regimes in overparametrized models\.InProceedings of Thirty Third Conference on Learning Theory,J\. Abernethy and S\. Agarwal \(Eds\.\),Proceedings of Machine Learning Research, Vol\.125,pp\. 3635–3673\.External Links:[Link](https://proceedings.mlr.press/v125/woodworth20a.html)Cited by:[§2](https://arxiv.org/html/2607.07884#S2.p6.5)\.
- L\. Xiao, Y\. Bahri, J\. Sohl\-Dickstein, S\. Schoenholz, and J\. Pennington \(2018\)Dynamical isometry and a mean field theory of CNNs: how to train 10,000\-layer vanilla convolutional neural networks\.InProceedings of the 35th International Conference on Machine Learning,J\. Dy and A\. Krause \(Eds\.\),Proceedings of Machine Learning Research, Vol\.80,pp\. 5393–5402\.External Links:[Link](https://proceedings.mlr.press/v80/xiao18a.html)Cited by:[Appendix A](https://arxiv.org/html/2607.07884#A1.p1.1)\.
- Y\. Xu and L\. Ziyin \(2025\)Three mechanisms of feature learning in a linear network\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=Wh4SE2S7Mo)Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p4.1)\.
- G\. Yang, D\. Yu, C\. Zhu, and S\. Hayou \(2024\)Tensor programs VI: feature learning in infinite depth neural networks\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=17pVDnpwwl)Cited by:[Appendix A](https://arxiv.org/html/2607.07884#A1.p1.1),[Appendix C](https://arxiv.org/html/2607.07884#A3.p1.2),[§1](https://arxiv.org/html/2607.07884#S1.p1.1),[§3](https://arxiv.org/html/2607.07884#S3.p3.1)\.
- Y\. Zhang, P\. E\. Latham, and A\. M\. Saxe \(2024\)Understanding unimodal bias in multimodal deep linear networks\.InProceedings of the 41st International Conference on Machine Learning,R\. Salakhutdinov, Z\. Kolter, K\. Heller, A\. Weller, N\. Oliver, J\. Scarlett, and F\. Berkenkamp \(Eds\.\),Proceedings of Machine Learning Research, Vol\.235,pp\. 59100–59125\.External Links:[Link](https://proceedings.mlr.press/v235/zhang24aa.html)Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p4.1)\.
- Y\. Zhang, A\. M\. Saxe, and P\. E\. Latham \(2026\)Saddle\-to\-saddle dynamics explains a simplicity bias across neural network architectures\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=Vit5M0G5Gb)Cited by:[§1](https://arxiv.org/html/2607.07884#S1.p4.1)\.
## Appendix AAdditional related work
A diverse body of theoretical research has investigated neural networks in the limit of large depth, regarding their expressivity\(Pooleet al\.,[2016](https://arxiv.org/html/2607.07884#bib.bib5); Hanin,[2019](https://arxiv.org/html/2607.07884#bib.bib14)\), initialization scheme\(Saxeet al\.,[2014](https://arxiv.org/html/2607.07884#bib.bib4); Schoenholzet al\.,[2017](https://arxiv.org/html/2607.07884#bib.bib6); Xiaoet al\.,[2018](https://arxiv.org/html/2607.07884#bib.bib12); Yanget al\.,[2024](https://arxiv.org/html/2607.07884#bib.bib49)\), the network output at initialization\(Hayou,[2023](https://arxiv.org/html/2607.07884#bib.bib40); Nociet al\.,[2023](https://arxiv.org/html/2607.07884#bib.bib35)\), the minimum\-norm solution\(Jacot,[2023](https://arxiv.org/html/2607.07884#bib.bib37); Boix\-Adsera,[2025](https://arxiv.org/html/2607.07884#bib.bib56)\), and formulations based on implicit layers\(Amos and Kolter,[2017](https://arxiv.org/html/2607.07884#bib.bib7); Baiet al\.,[2019](https://arxiv.org/html/2607.07884#bib.bib13)\)and continuous\-depth limits\(Haber and Ruthotto,[2017](https://arxiv.org/html/2607.07884#bib.bib8); Chenet al\.,[2018](https://arxiv.org/html/2607.07884#bib.bib10)\)\. Despite this progress, characterizing the behaviors of such deep networks once gradient descent training begins poses a greater challenge\. Exact solutions for the full training dynamics have been derived for deep linear networks with aligned small initial weights and whitened data\(Saxeet al\.,[2014](https://arxiv.org/html/2607.07884#bib.bib4);[2019](https://arxiv.org/html/2607.07884#bib.bib15)\)\. For nonlinear networks, current findings characterize the gradient descent dynamics over only one or several steps\(Jelassiet al\.,[2023](https://arxiv.org/html/2607.07884#bib.bib41); Hayou,[2023](https://arxiv.org/html/2607.07884#bib.bib40); Bordelon and Pehlevan,[2025](https://arxiv.org/html/2607.07884#bib.bib53); Bordelonet al\.,[2024a](https://arxiv.org/html/2607.07884#bib.bib45); Chizat,[2025](https://arxiv.org/html/2607.07884#bib.bib55)\), while the full learning dynamics is generally intractable\.
## Appendix BDeep scalar linear networks
### B\.1Additional figures
In[Figure3](https://arxiv.org/html/2607.07884#A2.F3), we show the trajectories of the total weight with different depths and learning rates in deep scalar linear networks\. In[Figure4](https://arxiv.org/html/2607.07884#A2.F4), we show the trajectories of the total weight and loss with different depths and initialization in deep scalar linear networks\.




Figure 3:Dynamics ofα\(t\)\\alpha\(t\)with different depthsLLand learning ratesη\\eta\. The learning rateη\\etais given by[Equation5](https://arxiv.org/html/2607.07884#S2.E5)withτ−1=1,1\.5,1\.95,2\.05\\tau^\{\-1\}=1,1\.5,1\.95,2\.05for the four panels from left to right\. When0<τ−1≤10<\\tau^\{\-1\}\\leq 1, the gradient descent dynamics is monotonic and well described by the gradient flow dynamics\. When1<τ−1<21<\\tau^\{\-1\}<2, the gradient descent dynamics is oscillatory but converging\. Whenτ−1≥2\\tau^\{\-1\}\\geq 2, the gradient descent dynamics is oscillatory and diverging\. Here the initialization isα\(0\)=0\.01\\alpha\(0\)=0\.01\. The data statistics areμyx=1,μxx=1\\mu\_\{yx\}=1,\\mu\_\{xx\}=1\.







Figure 4:Dynamics of total weights \(top row\) and loss \(bottom row\) with different depthsLLand initializationα\(0\)\\alpha\(0\)\. The learning speed decreases when the depth increases and when the initialization scale decreases\. Here the learning rateη\\etais given by[Equation5](https://arxiv.org/html/2607.07884#S2.E5)withτ−1=0\.1\\tau^\{\-1\}=0\.1\. The data statistics areμyx=1,μxx=1\\mu\_\{yx\}=1,\\mu\_\{xx\}=1\.
### B\.2Derivation of the sharpness in[Equation5](https://arxiv.org/html/2607.07884#S2.E5)
By differentiating the gradient in[Equation3](https://arxiv.org/html/2607.07884#S2.E3)with respect tow1w\_\{1\}, we obtain the sharpness of the loss landscape at which the weights in all layers are equal tow1w\_\{1\}
1η∂w˙1∂w1=μyx\(L−1\)w1L−2−μxx\(2L−1\)w12L−2\.\\displaystyle\\frac\{1\}\{\\eta\}\\frac\{\\partial\\dot\{w\}\_\{1\}\}\{\\partial w\_\{1\}\}=\\mu\_\{yx\}\(L\-1\)w\_\{1\}^\{L\-2\}\-\\mu\_\{xx\}\(2L\-1\)w\_\{1\}^\{2L\-2\}\.\(15\)The sharpness at the global minimum, denoted asSS, is
S≡1η∂w˙1∂w1\|w1=\(μyxμxx\)1/L=μxxL\(μyxμxx\)2−2/L\.\\displaystyle S\\equiv\\frac\{1\}\{\\eta\}\\frac\{\\partial\\dot\{w\}\_\{1\}\}\{\\partial w\_\{1\}\}\\bigg\|\_\{w\_\{1\}=\\left\(\\frac\{\\mu\_\{yx\}\}\{\\mu\_\{xx\}\}\\right\)^\{1/L\}\}=\\mu\_\{xx\}L\\left\(\\frac\{\\mu\_\{yx\}\}\{\\mu\_\{xx\}\}\\right\)^\{2\-2/L\}\.\(16\)ForL=1L=1, the sharpness depends only on the input variance,μxx\\mu\_\{xx\}, but not the input\-output correlation,μyx\\mu\_\{yx\}\. ForL≥2L\\geq 2, the sharpness depends on both the input variance and the input\-output correlation\. We note that[Equation16](https://arxiv.org/html/2607.07884#A2.E16)withμxx=1\\mu\_\{xx\}=1appeared inSaxeet al\.\([2014](https://arxiv.org/html/2607.07884#bib.bib4), Equation \(41\)\)\.
Remark: When deriving the gradient descent dynamics, we calculate the negative gradient using the original loss expression withLLvariables before substituting in the reductionw1=w2=⋯=wLw\_\{1\}=w\_\{2\}=\\cdots=w\_\{L\}\. Substituting in the equality before taking the gradient would yield the wrong gradient descent dynamics\. However, when calculating the second\-order derivative, we differentiate the expression in[Equation3](https://arxiv.org/html/2607.07884#S2.E3), which is the gradient after substituting in the reductionw1=w2=⋯=wLw\_\{1\}=w\_\{2\}=\\cdots=w\_\{L\}\. Substituting in the equality after the double differentiation would yield the wrong sharpness metric\. This is because we want the sharpness of the loss landscape along thew1=w2=⋯=wLw\_\{1\}=w\_\{2\}=\\cdots=w\_\{L\}path, not the sharpness along thew1w\_\{1\}axis with the rest of the weights held fixed\.
### B\.3Derivation of of the total weight dynamics in[Equation6](https://arxiv.org/html/2607.07884#S2.E6)
Using[Equation3](https://arxiv.org/html/2607.07884#S2.E3), we obtain the dynamics of the total weighta=w1La=w\_\{1\}^\{L\}
a˙=ηLa2−2/L\(μyx−μxxa\)\.\\displaystyle\\dot\{a\}=\\eta La^\{2\-2/L\}\\left\(\\mu\_\{yx\}\-\\mu\_\{xx\}a\\right\)\.\(17\)[Equation17](https://arxiv.org/html/2607.07884#A2.E17)withμxx=1\\mu\_\{xx\}=1appeared inSaxeet al\.\([2014](https://arxiv.org/html/2607.07884#bib.bib4), Equation \(15\)\)\. Substituting the learning rate in[Equation5](https://arxiv.org/html/2607.07884#S2.E5)into the dynamics ofaa, we get
τa˙=\(μyxμxx\)−2\+2/La2−2/L\(μyxμxx−a\)\.\\displaystyle\\tau\\dot\{a\}=\\left\(\\frac\{\\mu\_\{yx\}\}\{\\mu\_\{xx\}\}\\right\)^\{\-2\+2/L\}a^\{2\-2/L\}\\left\(\\frac\{\\mu\_\{yx\}\}\{\\mu\_\{xx\}\}\-a\\right\)\.\(18\)We denote the total weight divided by the target weight asα\(t\)=w1\(t\)Lμxx/μyx\\alpha\(t\)=w\_\{1\}\(t\)^\{L\}\\mu\_\{xx\}/\\mu\_\{yx\}, which represents the relative portion of the target weight learned, withα=1\\alpha=1being the global minimum\. The dynamics ofα\(t\)\\alpha\(t\)is given by
τα˙\\displaystyle\\tau\\dot\{\\alpha\}=τμxxμyxa˙\\displaystyle=\\tau\\frac\{\\mu\_\{xx\}\}\{\\mu\_\{yx\}\}\\dot\{a\}=\(μxxμyxa\)2−2/L\(1−μxxμyxa\)\\displaystyle=\\left\(\\frac\{\\mu\_\{xx\}\}\{\\mu\_\{yx\}\}a\\right\)^\{2\-2/L\}\\left\(1\-\\frac\{\\mu\_\{xx\}\}\{\\mu\_\{yx\}\}a\\right\)=α2−2/L\(1−α\)\.\\displaystyle=\\alpha^\{2\-2/L\}\(1\-\\alpha\)\.\(19\)We arrive at[Equation6](https://arxiv.org/html/2607.07884#S2.E6)in the main text\.
### B\.4Derivation of the infinite\-depth solution[Equation9](https://arxiv.org/html/2607.07884#S2.E9)
We here solve the learning dynamics withL→∞L\\to\\infty, which is given by
τα˙=α2\(1−α\)\.\\displaystyle\\tau\\dot\{\\alpha\}=\\alpha^\{2\}\\left\(1\-\\alpha\\right\)\.\(20\)By separating variables and integrating both sides, we obtain
∫0t1τ𝑑t′\\displaystyle\\int\_\{0\}^\{t\}\\frac\{1\}\{\\tau\}dt^\{\\prime\}=∫α0α\(t\)dα′α′2\(1−α′\)\\displaystyle=\\int\_\{\\alpha\_\{0\}\}^\{\\alpha\(t\)\}\\frac\{d\\alpha^\{\\prime\}\}\{\\alpha^\{\\prime 2\}\(1\-\\alpha^\{\\prime\}\)\}\(21\)⇒tτ\\displaystyle\\Rightarrow\\quad\\frac\{t\}\{\\tau\}=\(−1α−ln\(1α−1\)\)\|α0α\(t\)\.\\displaystyle=\\left\(\-\\frac\{1\}\{\\alpha\}\-\\ln\\left\(\\frac\{1\}\{\\alpha\}\-1\\right\)\\right\)\\bigg\|\_\{\\alpha\_\{0\}\}^\{\\alpha\(t\)\}\.\(22\)[Equation22](https://arxiv.org/html/2607.07884#A2.E22)appeared inSaxeet al\.\([2014](https://arxiv.org/html/2607.07884#bib.bib4), Equation \(17\)\)\. We rearrange[Equation22](https://arxiv.org/html/2607.07884#A2.E22)and obtain
1α\(t\)−1\+ln\(1α\(t\)−1\)=−tτ\+1α0\+ln\(1α0−1\)−1=defβ\(t\)\.\\displaystyle\\frac\{1\}\{\\alpha\(t\)\}\-1\+\\ln\\left\(\\frac\{1\}\{\\alpha\(t\)\}\-1\\right\)=\-\\frac\{t\}\{\\tau\}\+\\frac\{1\}\{\\alpha\_\{0\}\}\+\\ln\\left\(\\frac\{1\}\{\\alpha\_\{0\}\}\-1\\right\)\-1\\overset\{\\text\{def\}\}\{=\}\\beta\(t\)\.\(23\)Taking the exponential of both sides yields
\(1α\(t\)−1\)e1α\(t\)−1=eβ\(t\)\.\\displaystyle\\left\(\\frac\{1\}\{\\alpha\(t\)\}\-1\\right\)e^\{\\frac\{1\}\{\\alpha\(t\)\}\-1\}=e^\{\\beta\(t\)\}\.\(24\)Because the principal branch of the LambertWWfunction, denotedy=W0\(x\)y=W\_\{0\}\(x\), solves the equationyey=xye^\{y\}=xwithx≥0x\\geq 0, we have
1α\(t\)−1\\displaystyle\\frac\{1\}\{\\alpha\(t\)\}\-1=W0\(eβ\(t\)\)\\displaystyle=W\_\{0\}\(e^\{\\beta\(t\)\}\)⇒α\(t\)\\displaystyle\\Rightarrow\\quad\\alpha\(t\)=11\+W0\(eβ\(t\)\),0<α≤1\.\\displaystyle=\\frac\{1\}\{1\+W\_\{0\}\(e^\{\\beta\(t\)\}\)\},\\quad 0<\\alpha\\leq 1\.\(25\)We arrive at[Equation9](https://arxiv.org/html/2607.07884#S2.E9)in the main text\.
## Appendix CDeep scalar linear residual networks with block depth one
Figure 5:The loss landscape of scalar linear residual networks with block depth one\. Similar to the scalar linear chain in[Figure1](https://arxiv.org/html/2607.07884#S2.F1), the sharpness of the global minimum increases with the depthLL\. Specifically, the plotted curves areℒ\(w\)=\(2−\(1\+w/L\)L\)2/2\\mathcal\{L\}\(w\)=\\left\(2\-\(1\+w/\\sqrt\{L\}\)^\{L\}\\right\)^\{2\}/2, with differentLL\.



Figure 6:Loss trajectories of deep scalar linear residual networks with block depth one with different depths and initialization\. Here the learning rateη\\etais given by[Equation31](https://arxiv.org/html/2607.07884#A3.E31)withτ−1=0\.1\\tau^\{\-1\}=0\.1\. The data statistics areμyx=2,μxx=1\\mu\_\{yx\}=2,\\mu\_\{xx\}=1\.Consider a scalar linear residual network with block depth one defined as
f\(x;w\)=∏l=1L\(1\+wlL\)x,x,w1,⋯,wL∈ℝ\.\\displaystyle f\(x;w\)=\\prod\_\{l=1\}^\{L\}\\left\(1\+\\frac\{w\_\{l\}\}\{\\sqrt\{L\}\}\\right\)x,\\quad x,w\_\{1\},\\cdots,w\_\{L\}\\in\\mathbb\{R\}\.\(26\)The1/L1/\\sqrt\{L\}factor is a standard choice consistent withBordelonet al\.\([2024b](https://arxiv.org/html/2607.07884#bib.bib46)\); Yanget al\.\([2024](https://arxiv.org/html/2607.07884#bib.bib49)\)\. The gradient flow dynamics trained withℓ2\\ell\_\{2\}loss is given by
w˙1=ηL\[μyx−μxx∏i=1L\(1\+wiL\)\]∏i≠l\(1\+wiL\)\.\\displaystyle\\dot\{w\}\_\{1\}=\\frac\{\\eta\}\{\\sqrt\{L\}\}\\left\[\\mu\_\{yx\}\-\\mu\_\{xx\}\\prod\_\{i=1\}^\{L\}\\left\(1\+\\frac\{w\_\{i\}\}\{\\sqrt\{L\}\}\\right\)\\right\]\\prod\_\{i\\neq l\}\\left\(1\+\\frac\{w\_\{i\}\}\{\\sqrt\{L\}\}\\right\)\.\(27\)Similar to the deep scalar linear network, we make the assumption of having equal initial weight in each layer,wl\(0\)=w1\(0\)∀lw\_\{l\}\(0\)=w\_\{1\}\(0\)\\,\\forall l, which will remain equal throughout training due to the conservation law\. With equal weight in each layer, the gradient flow dynamics reduces to an one\-dimensional ordinary differential equation
w˙1=ηL\[μyx−μxx\(1\+w1L\)L\]\(1\+w1L\)L−1\.\\displaystyle\\dot\{w\}\_\{1\}=\\frac\{\\eta\}\{\\sqrt\{L\}\}\\left\[\\mu\_\{yx\}\-\\mu\_\{xx\}\\left\(1\+\\frac\{w\_\{1\}\}\{\\sqrt\{L\}\}\\right\)^\{L\}\\right\]\\left\(1\+\\frac\{w\_\{1\}\}\{\\sqrt\{L\}\}\\right\)^\{L\-1\}\.\(28\)By differentiating the gradient in[Equation28](https://arxiv.org/html/2607.07884#A3.E28)with respect tow1w\_\{1\}, we obtain the sharpness of the loss landscape at which the weights in all layers are equal tow1w\_\{1\}
1η∂w˙1∂w1=1L\[μyx\(L−1\)\(1\+w1L\)L−2−μxx\(2L−1\)\(1\+w1L\)2L−2\]\.\\displaystyle\\frac\{1\}\{\\eta\}\\frac\{\\partial\\dot\{w\}\_\{1\}\}\{\\partial w\_\{1\}\}=\\frac\{1\}\{L\}\\left\[\\mu\_\{yx\}\(L\-1\)\\left\(1\+\\frac\{w\_\{1\}\}\{\\sqrt\{L\}\}\\right\)^\{L\-2\}\-\\mu\_\{xx\}\(2L\-1\)\\left\(1\+\\frac\{w\_\{1\}\}\{\\sqrt\{L\}\}\\right\)^\{2L\-2\}\\right\]\.\(29\)The sharpness at the global minimum is
S≡1η∂w˙1∂w1\|1\+w1L=\(μyxμxx\)1/L=μxx\(μyxμxx\)2−2/L\.\\displaystyle S\\equiv\\frac\{1\}\{\\eta\}\\frac\{\\partial\\dot\{w\}\_\{1\}\}\{\\partial w\_\{1\}\}\\bigg\|\_\{1\+\\frac\{w\_\{1\}\}\{\\sqrt\{L\}\}=\\left\(\\frac\{\\mu\_\{yx\}\}\{\\mu\_\{xx\}\}\\right\)^\{1/L\}\}=\\mu\_\{xx\}\\left\(\\frac\{\\mu\_\{yx\}\}\{\\mu\_\{xx\}\}\\right\)^\{2\-2/L\}\.\(30\)Hence, the maximum stable learning rate scales as
η=τ−11μxx\(μyxμxx\)−2\+2/L,\\displaystyle\\eta=\\tau^\{\-1\}\\frac\{1\}\{\\mu\_\{xx\}\}\\left\(\\frac\{\\mu\_\{yx\}\}\{\\mu\_\{xx\}\}\\right\)^\{\-2\+2/L\},\(31\)whereτ∈\(0\.5,∞\)\\tau\\in\(0\.5,\\infty\)is the time constant\.
As shown in[Figure2](https://arxiv.org/html/2607.07884#S3.F2)B, the optimal learning rate transfers under the data\-dependent scaling in[Equation31](https://arxiv.org/html/2607.07884#A3.E31), but does not transfer under the data\-agnostic constant scaling ofη∝1\\eta\\propto 1\. Similar to deep scalar linear networks without residual connections, we note that the data dependence of the maximum stable learning rate is weak for largeLLin deep scalar linear residual networks with block depth one,limL→∞2/L=0\\lim\_\{L\\to\\infty\}2/L=0\. Thus, transferring the optimal learning rate from an intermediate depth to infinite depth under the constant scaling is still justified, whereas transferring the learning rate from a small depth \(e\.g\.L=2,4L=2,4\) to infinite depth would likely fail, as we can see from[Figure2](https://arxiv.org/html/2607.07884#S3.F2)B\.
## Appendix DDeep scalar linear residual networks with block depth two
Figure 7:The loss landscape of scalar linear residual networks with block depth two\. Specifically, the plotted curves areℒ\(w\)=\(2−\(1\+w2/L\)L\)2/2\\mathcal\{L\}\(w\)=\\left\(2\-\(1\+w^\{2\}/L\)^\{L\}\\right\)^\{2\}/2, with differentLL\.



Figure 8:Loss trajectories of deep scalar linear residual networks with block depth two with different depths and initialization\. Here the learning rateη\\etais given by[Equation37](https://arxiv.org/html/2607.07884#A4.E37)withτ−1=0\.1\\tau^\{\-1\}=0\.1\. The data statistics areμyx=2,μxx=1\\mu\_\{yx\}=2,\\mu\_\{xx\}=1\.Consider a scalar linear residual network with block depth two defined as
f\(x;w\)=∏l=1L\(1\+wl2L\)x,x,w1,⋯,wL∈ℝ\.\\displaystyle f\(x;w\)=\\prod\_\{l=1\}^\{L\}\\left\(1\+\\frac\{w\_\{l\}^\{2\}\}\{L\}\\right\)x,\\quad x,w\_\{1\},\\cdots,w\_\{L\}\\in\\mathbb\{R\}\.\(32\)The1/L1/Lfactor is a standard choice consistent withBordelonet al\.\([2024a](https://arxiv.org/html/2607.07884#bib.bib45)\); Deyet al\.\([2025](https://arxiv.org/html/2607.07884#bib.bib54)\)\. In this architecture, we needμyx/μxx≥1\\mu\_\{yx\}/\\mu\_\{xx\}\\geq 1, since the total weight∏l=1L\(1\+wl2L\)≥1\\prod\_\{l=1\}^\{L\}\\left\(1\+\\frac\{w\_\{l\}^\{2\}\}\{L\}\\right\)\\geq 1\. The gradient flow dynamics trained withℓ2\\ell\_\{2\}loss is given by
w˙1=2ηw1L\[μyx−μxx∏i=1L\(1\+wi2L\)\]∏i≠l\(1\+wi2L\)\.\\displaystyle\\dot\{w\}\_\{1\}=\\frac\{2\\eta w\_\{1\}\}\{L\}\\left\[\\mu\_\{yx\}\-\\mu\_\{xx\}\\prod\_\{i=1\}^\{L\}\\left\(1\+\\frac\{w\_\{i\}^\{2\}\}\{L\}\\right\)\\right\]\\prod\_\{i\\neq l\}\\left\(1\+\\frac\{w\_\{i\}^\{2\}\}\{L\}\\right\)\.\(33\)Similar to the deep scalar linear network, we make the assumption of having equal initial weight in each layer,wl\(0\)=w1\(0\)∀lw\_\{l\}\(0\)=w\_\{1\}\(0\)\\,\\forall l, which will remain equal throughout training due to the conservation law\. With equal weight in each layer, the gradient flow dynamics reduces to an one\-dimensional ordinary differential equation
w˙1=2ηL\[μyx−μxx\(1\+w12L\)L\]\(1\+w12L\)L−1w1\.\\displaystyle\\dot\{w\}\_\{1\}=\\frac\{2\\eta\}\{L\}\\left\[\\mu\_\{yx\}\-\\mu\_\{xx\}\\left\(1\+\\frac\{w\_\{1\}^\{2\}\}\{L\}\\right\)^\{L\}\\right\]\\left\(1\+\\frac\{w\_\{1\}^\{2\}\}\{L\}\\right\)^\{L\-1\}w\_\{1\}\.\(34\)By differentiating the gradient in[Equation28](https://arxiv.org/html/2607.07884#A3.E28)with respect tow1w\_\{1\}, we obtain the sharpness of the loss landscape at which the weights in all layers are equal tow1w\_\{1\}
1η∂w˙1∂w1=2L\[μyx\(1\+w12L\)L−2\(2L−1Lw2\+1\)−μxx\(1\+w12L\)2L−2\(4L−1Lw2\+1\)\]\.\\displaystyle\\frac\{1\}\{\\eta\}\\frac\{\\partial\\dot\{w\}\_\{1\}\}\{\\partial w\_\{1\}\}=\\frac\{2\}\{L\}\\left\[\\mu\_\{yx\}\\left\(1\+\\frac\{w\_\{1\}^\{2\}\}\{L\}\\right\)^\{L\-2\}\\left\(\\frac\{2L\-1\}\{L\}w^\{2\}\+1\\right\)\-\\mu\_\{xx\}\\left\(1\+\\frac\{w\_\{1\}^\{2\}\}\{L\}\\right\)^\{2L\-2\}\\left\(\\frac\{4L\-1\}\{L\}w^\{2\}\+1\\right\)\\right\]\.\(35\)The sharpness at the global minimum is
S≡1η∂w˙1∂w1\|1\+w12L=\(μyxμxx\)1/L\\displaystyle S\\equiv\\frac\{1\}\{\\eta\}\\frac\{\\partial\\dot\{w\}\_\{1\}\}\{\\partial w\_\{1\}\}\\bigg\|\_\{1\+\\frac\{w\_\{1\}^\{2\}\}\{L\}=\\left\(\\frac\{\\mu\_\{yx\}\}\{\\mu\_\{xx\}\}\\right\)^\{1/L\}\}=4Lμxx\(μyxμxx\)2−2/Lw2\\displaystyle=\\frac\{4\}\{L\}\\mu\_\{xx\}\\left\(\\frac\{\\mu\_\{yx\}\}\{\\mu\_\{xx\}\}\\right\)^\{2\-2/L\}w^\{2\}=4μxx\(μyxμxx\)2−2/L\(\(μyxμxx\)1/L−1\)\\displaystyle=4\\mu\_\{xx\}\\left\(\\frac\{\\mu\_\{yx\}\}\{\\mu\_\{xx\}\}\\right\)^\{2\-2/L\}\\left\(\\left\(\\frac\{\\mu\_\{yx\}\}\{\\mu\_\{xx\}\}\\right\)^\{1/L\}\-1\\right\)\(36\)Hence, the maximum stable learning rate scales as
η=τ−114μxx\(μyxμxx\)−2\+2/L\(\(μyxμxx\)1/L−1\)−1\\displaystyle\\eta=\\tau^\{\-1\}\\frac\{1\}\{4\\mu\_\{xx\}\}\\left\(\\frac\{\\mu\_\{yx\}\}\{\\mu\_\{xx\}\}\\right\)^\{\-2\+2/L\}\\left\(\\left\(\\frac\{\\mu\_\{yx\}\}\{\\mu\_\{xx\}\}\\right\)^\{1/L\}\-1\\right\)^\{\-1\}\(37\)whereτ∈\(0\.5,∞\)\\tau\\in\(0\.5,\\infty\)is the time constant\.
The scaling of[AppendixD](https://arxiv.org/html/2607.07884#A4.Ex4)with respect toLLis not immediately apparent\. To see its behavior with largeLL, we Taylor expand[AppendixD](https://arxiv.org/html/2607.07884#A4.Ex4)around1/L=01/L=0, which yields
S=4Lμxx\(μyxμxx\)2ln\(μyxμxx\)\+O\(1L2\)\.\\displaystyle S=\\frac\{4\}\{L\}\\mu\_\{xx\}\\left\(\\frac\{\\mu\_\{yx\}\}\{\\mu\_\{xx\}\}\\right\)^\{2\}\\ln\\left\(\\frac\{\\mu\_\{yx\}\}\{\\mu\_\{xx\}\}\\right\)\+O\\left\(\\frac\{1\}\{L^\{2\}\}\\right\)\.\(38\)This shows that the sharpness at the global minimum decreases withLL, scaling as1/L1/L\. Therefore, if we were to use a data\-agnostic power\-law scaling, the learning rate would scale with depth asη∝L\\eta\\propto L\.
In[Figure2](https://arxiv.org/html/2607.07884#S3.F2)C, we compare the learning rate transfer between the exact maximum stable learning rate scaling in[Equation37](https://arxiv.org/html/2607.07884#A4.E37)and the data\-agnostic scaling ofη∝L\\eta\\propto L\. Similar to the cases with deep scalar linear networks and scalar linear residual networks with block depth one, the optimal learning rate transfers under the data\-dependent scaling in[Equation37](https://arxiv.org/html/2607.07884#A4.E37), but not under the data\-agnostic scaling ofη∝L\\eta\\propto L\.
## ReproducibilitySimilar Articles
Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural Networks
This paper derives exact closed-form expressions for gradients and test loss after one and two steps of gradient descent in two-layer and three-layer linear neural networks, characterizing optimal learning rate selection and revealing a distinct early-training regime where unequal layer-wise learning rates are initially optimal.
Scaling Laws, Carefully (25 minute read)
A comprehensive overview of scaling laws in deep learning, tracing their theoretical roots and empirical findings, and explaining how loss decreases predictably with model size, data, and compute.
Prescriptive Scaling Laws for Data Constrained Training
A modified scaling law accounting for data repetition effects provides compute-optimal training strategies for data-constrained scenarios, showing that beyond a point further repetition is counterproductive and compute is better spent on model capacity.
@lilianweng: A super long overdue (3+ years?) post on scaling laws. Compute is expensive. Scaling laws are a way to help us reason a…
Lilian Weng's blog post provides a comprehensive overview of scaling laws in deep learning, covering their derivation, compute-optimal allocation, and the debate between Kaplan et al. and Chinchilla.
Sketched Linear Contrastive Learning: Approximation, Optimization, and Statistical Scaling
This paper derives a scaling law for sketched linear contrastive learning under a Gaussian latent-variable model, analyzing how risk decomposes into approximation, optimization, and statistical terms, and provides theoretical guidance for balancing model size, data, and compute in contrastive learning.