A Function-Space Approach to the Statistical Mechanics of Learning Dynamics
Summary
This paper develops a statistical-mechanical framework for analyzing learning dynamics in deep neural networks by shifting from parameter space to function space, deriving exact error dynamics and fluctuation-induced effects.
View Cached Full Text
Cached at: 09/11/26, 08:38 AM
# A Function-Space Approach to the Statistical Mechanics of Learning Dynamics
Source: [https://arxiv.org/html/2609.09589](https://arxiv.org/html/2609.09589)
###### Abstract
Deep neural networks exhibit surprisingly regular macroscopic behavior despite highly nonlinear dynamics in a vast parameter space\. We develop a statistical\-mechanical description of learning directly in function space, treating parameter configurations as microscopic realizations and functions together with their dynamical operators as macroscopic variables\. For mean\-squared loss, the exact error dynamics are governed by the learning operatorM=JJ∗M=JJ^\{\\ast\}\. The bare conditional stochastic dynamics supplies a dynamical Boltzmann weight, while parameter\-space multiplicity contributes a function\-space density of states whose local curvature defines a statistical operatorBB\. Conditioning on a current error macrostate and integrating over the resulting local fluctuation ensemble yields the conditional free\-energy contributionΦfluc\(M,B\)=σξ22logdet\(M−1\+B\)\+const\\Phi\_\{\\mathrm\{fluc\}\}\(M;B\)=\\frac\{\\sigma\_\{\\xi\}^\{2\}\}\{2\}\\log\\det\(M^\{\-1\}\+B\)\+\\mathrm\{const\}on the active sector\. At fixed spectrum, this contribution is rotationally stationary when\[M,B\]=0\[M,B\]=0, is minimized by pairing large eigenvalues ofMMwith small eigenvalues ofBB, and supplies a local restoring force contribution against rotational mismatch\. For ReLU\-type function spaces under mild stable statistical conditions,BBtakes the formB=σξ2L∗𝒦LB=\\sigma\_\{\\xi\}^\{2\}L^\{\\ast\}\\mathcal\{K\}L, withLLmeasuring coarse\-grained second\-order structure\. Its low\-BBsector therefore corresponds, up to the bounded anisotropy of𝒦\\mathcal\{K\}, to directions of low structural curvature\. Within this ReLU specialization, the fluctuation\-induced contribution therefore supplies a preference for pairing faster relaxation with this low\-curvature, data\-adaptive sector\. These results suggest function space as a natural macroscopic level for studying stable collective organization in learning, with a concrete neural\-network model entering as a realization of the macroscopic statistical theory rather than defining its form from the outset\.
## 1Introduction
Deep neural networks pose an unusual problem for theory\. Their training dynamics arise from strongly nonlinear interactions among an enormous number of parameters, yet their macroscopic behavior is often strikingly regular\. Across several learning domains, performance obeys smooth scaling relations over large changes in model size, data, and compute, while appropriate parameterizations can make optimization hyperparameters transferable across orders of magnitude in width[Hestness et al\. \(2017\)](https://arxiv.org/html/2609.09589#bib.bib16);[Kaplan et al\. \(2020\)](https://arxiv.org/html/2609.09589#bib.bib17);[Yang et al\. \(2021\)](https://arxiv.org/html/2609.09589#bib.bib18)\. Regular organization also appears within the learned dynamics and representations themselves: neural\-network training repeatedly develops characteristic spectral structure[Cohen et al\. \(2021\)](https://arxiv.org/html/2609.09589#bib.bib19), and learned functions exhibit systematic preferences for smooth, data\-dependent directions[Rahaman et al\. \(2019\)](https://arxiv.org/html/2609.09589#bib.bib13);[Kadkhodaie et al\. \(2024\)](https://arxiv.org/html/2609.09589#bib.bib20)\. These observations are striking precisely because such regularity is not obvious from the microscopic equations of training\. They raise a basic mechanistic question:
How can highly nonlinear learning dynamics give rise to stable and scalable macroscopic organization?\\boxed\{\\textit\{How can highly nonlinear learning dynamics give rise to stable and scalable macroscopic organization?\}\}\(1\)This question implies that a useful theory of deep learning should therefore do more than track individual parameter trajectories\. It should identify a level of description at which robust collective structure becomes visible\.
A particularly successful step in this direction has been to move from parameter space to function space\. In the neural tangent kernel regime, a highly nonlinear parameterized model reduces to an approximately closed linear dynamics governed by a kernel that remains nearly fixed throughout training[Jacot et al\. \(2018\)](https://arxiv.org/html/2609.09589#bib.bib8);[Chizat et al\. \(2019\)](https://arxiv.org/html/2609.09589#bib.bib9)\. This reveals that a complicated microscopic system can admit a much simpler macroscopic description\. The price of this closure, however, is that the geometry governing learning is effectively frozen\. Feature\-learning theories relax this restriction and allow representations, kernels, and their spectra to evolve[Woodworth et al\. \(2020\)](https://arxiv.org/html/2609.09589#bib.bib10);[Lauditi et al\. \(2025\)](https://arxiv.org/html/2609.09589#bib.bib11);[Lauditi et al\. \(2026\)](https://arxiv.org/html/2609.09589#bib.bib12)\. The two regimes therefore expose a useful tension: fixing the dynamical geometry yields a simple closed description, whereas allowing the geometry itself to adapt restores an essential part of learning but also reintroduces nonlinear, self\-consistent, and often model\-dependent dynamics\. This motivates asking whether the evolving macroscopic organization of learning can be characterized without resolving its full microscopic trajectory\.
The combination of enormous microscopic dimension, strong nonlinearity, and reproducible macroscopic structure makes statistical mechanics a natural framework for this question\. Statistical\-mechanical ideas have long been used to study neural networks through Gibbs ensembles, high\-dimensional landscapes, stochastic\-gradient diffusions, mean\-field descriptions, and collective order parameters[Bahri et al\. \(2020\)](https://arxiv.org/html/2609.09589#bib.bib7);[Mandt et al\. \(2016\)](https://arxiv.org/html/2609.09589#bib.bib4);[Mandt et al\. \(2017\)](https://arxiv.org/html/2609.09589#bib.bib5);[Chaudhari and Soatto \(2018\)](https://arxiv.org/html/2609.09589#bib.bib6)\. These approaches have established that useful macroscopic laws can emerge after coarse graining over microscopic degrees of freedom\. In many existing formulations, however, the statistical description is constructed either directly in parameter space or through a set of macroscopic variables chosen for a specific model or limit\. A comparatively unexplored possibility is to formulate the statistical mechanics of learning directly at the level where the learned object itself evolves: function space\.
This is the perspective developed in the present work\. A central methodological distinction is that a neural\-network model is treated as a microscopic realization of the statistical theory rather than as its starting point\. We first formulate the macroscopic variables and conditional statistical law in function space; a concrete architecture then determines which dynamical operators and microstate geometries are realizable, and therefore how the macroscopic law is instantiated\. This does not make architecture irrelevant: both the instantaneous learning operator and the multiplicity of microscopic realizations remain model\-dependent\. Rather, it separates the form of the organizing principle from its model\-specific realization\. This distinction is particularly natural for macroscopic regularities that persist across architectures: their concrete realization may vary, while the form of their organizing mechanism need not originate from any one microscopic model\.
Parameter configurations are treated as microscopic realizations, while functions and the operators governing their evolution provide the macroscopic variables\. Scalar quantities such as the loss remain important observables, but they retain only the magnitude of the prediction error and discard much of its directional and spectral structure\. Function space lies between these two extremes: it coarse\-grains over redundant parameterizations while preserving the geometry of learning\. For mean\-squared loss, the exact function\-space gradient dynamics take the simple form
e˙=−Me,M=JJ∗,\\dot\{e\}=\-Me,\\qquad M=JJ^\{\\ast\},\(2\)so the spectrum and eigendirections ofMMdirectly determine the relaxation geometry of the error\. The parameterization enters this description through the multiplicity of microscopic realizations of a given function\-space state\. Denoting the corresponding density of states byΩ\(e\)\\Omega\(e\), its local curvature defines
B\(r\)=−σξ2∇e2logΩ\(e\)\|e=r\.\\boxed\{B\(r\)=\-\\sigma\_\{\\xi\}^\{2\}\\nabla\_\{e\}^\{2\}\\log\\Omega\(e\)\\big\|\_\{e=r\}\.\}\(3\)ThusMMdescribes the dynamical geometry of learning, whileBBdescribes the statistical geometry induced by the underlying parameter microstates\. Only the compression ofBBto the finite\-dimensional active sector ofMMenters the determinants and commutators below\.
Conditioning on a current error macrostate and integrating over the local function\-space fluctuation ensemble yields a conditional free\-energy contribution that depends on the dynamical operator,
Φfluc\(M,B\)=σξ22logdet\(M−1\+B\)\+const\.\\boxed\{\\Phi\_\{\\mathrm\{fluc\}\}\(M;B\)=\\frac\{\\sigma\_\{\\xi\}^\{2\}\}\{2\}\\log\\det\(M^\{\-1\}\+B\)\+\\mathrm\{const\}\.\}\(4\)Its rotational structure produces a sharp conditional preference\. At fixed spectrum,
δrotΦfluc=0⟺\[M,B\]=0,\\boxed\{\\delta\_\{\\mathrm\{rot\}\}\\Phi\_\{\\mathrm\{fluc\}\}=0\\quad\\Longleftrightarrow\\quad\[M,B\]=0,\}\(5\)while the free\-energy minimum pairs the spectra as
mlarge⟷bsmall\.\\boxed\{m\_\{\\mathrm\{large\}\}\\longleftrightarrow b\_\{\\mathrm\{small\}\}\.\}\(6\)The matched state further has positive rotational curvature in every nondegenerate pairwise direction\. Accordingly, this conditional free\-energy contribution supplies a local restoring thermodynamic force against operator mismatch\. We do not assume that this term exhausts the full slow dynamics ofMM; rather, it identifies one definite thermodynamic bias contributed by the local fluctuation sector\. The complete local free energy also contains a distinct macrostate term,12⟨ra,M−1ra⟩p\\frac\{1\}\{2\}\\langle r\_\{a\},M^\{\-1\}r\_\{a\}\\rangle\_\{p\}, which supplies an error\-directed orientational preference; Sec\.[4\.1](https://arxiv.org/html/2609.09589#S4.SS1)separates these two contributions explicitly\.
We then examine the physical meaning of this abstract statistical geometry in a concrete class of function spaces\. ReLU networks represent continuous piecewise\-affine functions whose second\-order structure is concentrated on activation\-cell boundaries\. After coarse graining, this structure can be represented by a linear operatorL≃𝒞D2L\\simeq\\mathcal\{C\}D^\{2\}\. Under mild statistical boundary conditions on the local microstate ensemble, the curvature operator becomes
B=σξ2L∗𝒦L,\\boxed\{B=\\sigma\_\{\\xi\}^\{2\}L^\{\\ast\}\\mathcal\{K\}L,\}\(7\)where𝒦\\mathcal\{K\}is a positive structural entropy metric\. Hence, for an eigenmodeBϕi=biϕiB\\phi\_\{i\}=b\_\{i\}\\phi\_\{i\},
bi≍‖Lϕi‖2,b\_\{i\}\\asymp\\\|L\\phi\_\{i\}\\\|^\{2\},\(8\)and the low\-BBsector corresponds, up to the bounded anisotropy of𝒦\\mathcal\{K\}, to directions of small coarse\-grained structural curvature\. Combining this identification with the thermodynamic matching condition gives
mlarge⟷low structural\-curvature data\-adaptive sector\.\\boxed\{m\_\{\\mathrm\{large\}\}\\longleftrightarrow\\text\{low structural\-curvature data\-adaptive sector\}\.\}\(9\)Within this class of function spaces, the conditional thermodynamic contribution therefore favors pairing faster relaxation with the low\-curvature sector of the data\-weighted functional geometry\.
Taken together, these results suggest that function space provides a natural macroscopic level for the statistical mechanics of learning\. At this level, the learning dynamics can be separated from the statistical constraints supplied by the underlying parameterization: the former determines how function\-space states evolve, while the latter determines which such states admit many microscopic realizations and how those realizations are organized\. In the present work, this separation reveals a thermodynamic bias in the direction of feature evolution\. More broadly, the same viewpoint may provide a route for studying other stable collective structures of learning—including representation geometry, spectral organization, and stability—without requiring a complete description of the microscopic parameter trajectory\.
#### Contributions\.
Our main contributions are:
- •We develop a conditional statistical\-mechanical formulation of neural\-network training directly in function space\. The macroscopic statistical law is formulated before choosing a particular architecture, with concrete neural networks entering as microscopic realizations of that law\.
- •We show how parameter\-space microstate multiplicity induces a function\-space density of states and a local statistical curvature operatorBB, thereby separating the dynamical geometryMMfrom the statistical geometry supplied by the parameterization\.
- •We derive the conditional fluctuation contribution to the operator free energy and prove that its rotational stationary points satisfy\[M,B\]=0\[M,B\]=0\. On a fixed\-spectrum orbit this contribution is globally minimized by reverse spectral pairing, with larger eigenvalues ofMMmatched to smaller eigenvalues ofBB, and its gradient supplies a local restoring force contribution around the matched state\.
- •For ReLU\-type function spaces under mild stable statistical boundary conditions, we show thatB=σξ2L∗𝒦LB=\\sigma\_\{\\xi\}^\{2\}L^\{\\ast\}\\mathcal\{K\}Land that its spectrum measures coarse\-grained structural curvature up to the bounded anisotropy of the structural entropy metric\. Combining this result with the conditional operator preference yields a thermodynamic bias toward pairing faster relaxation with the low\-curvature data\-adaptive sector\.
The remainder of the paper develops these results in three steps\. Section[3](https://arxiv.org/html/2609.09589#S3)constructs the conditional statistical mechanics of error fluctuations\. Section[4](https://arxiv.org/html/2609.09589#S4)analyzes the conditional operator free\-energy contribution and its matching geometry\. Section[5](https://arxiv.org/html/2609.09589#S5)connects the resulting microstate curvature to the structural smoothness of ReLU function spaces\.
## 2Related Work
#### Kernel limits and feature learning\.
A large body of work characterizes neural\-network training through function\-space kernels\. In the infinite\-width neural tangent kernel \(NTK\) limit, gradient descent is governed by an approximately fixed kernel, and different function\-space modes are learned at rates set by the corresponding kernel eigenvalues[Jacot et al\. \(2018\)](https://arxiv.org/html/2609.09589#bib.bib8)\. This fixed\-kernel description is closely related to the lazy\-training regime, in which the network remains near its initial linearization[Chizat et al\. \(2019\)](https://arxiv.org/html/2609.09589#bib.bib9)\. Subsequent work has emphasized the distinction between such kernel regimes and richer training regimes in which the representation itself evolves[Woodworth et al\. \(2020\)](https://arxiv.org/html/2609.09589#bib.bib10)\. More recent analyses explicitly derive adaptive kernels and evolving spectral structure in feature\-learning limits[Lauditi et al\. \(2025\)](https://arxiv.org/html/2609.09589#bib.bib11);[Lauditi et al\. \(2026\)](https://arxiv.org/html/2609.09589#bib.bib12)\.
Our work concerns this latter regime, but reverses the usual order of construction\. Rather than beginning from a specified neural\-network model and deriving model\-specific macroscopic variables, we formulate the conditional statistical mechanics at the function\-space level and then ask how a concrete model realizes it\. Within this formulation, the fluctuation contribution depends on the relative geometry of the dynamical operator and the parameterization\-induced microstate geometry\. Its rotational extrema requireMMto commute with a microstate\-curvature operatorBB, and its minima pair large dynamical eigenvalues with small microstate\-curvature eigenvalues\. The ReLU specialization in Sec\.[5](https://arxiv.org/html/2609.09589#S5)is therefore a concrete realization of the general operator\-level construction rather than its starting point\.
#### Stochastic gradient dynamics and statistical mechanics\.
Diffusion and statistical\-mechanical descriptions of stochastic optimization have a long history in machine learning\. Constant\-step stochastic gradient dynamics can be approximated locally by diffusion or Ornstein–Uhlenbeck processes, leading to effective stationary distributions and Bayesian interpretations[Mandt et al\. \(2016\)](https://arxiv.org/html/2609.09589#bib.bib4);[Mandt et al\. \(2017\)](https://arxiv.org/html/2609.09589#bib.bib5)\. Other work has emphasized that realistic stochastic\-gradient noise is generally anisotropic and can generate genuinely nonequilibrium behavior[Chaudhari and Soatto \(2018\)](https://arxiv.org/html/2609.09589#bib.bib6)\. More broadly, statistical\-mechanical tools such as effective energies, free energies, and high\-dimensional random systems have played an important role in theoretical studies of deep learning[Bahri et al\. \(2020\)](https://arxiv.org/html/2609.09589#bib.bib7)\. The Langevin and Fokker–Planck machinery used in our bare conditional dynamics follows the standard theory of reversible diffusion processes[Risken \(1989\)](https://arxiv.org/html/2609.09589#bib.bib1);[Pavliotis \(2014\)](https://arxiv.org/html/2609.09589#bib.bib2);[Jordan et al\. \(1998\)](https://arxiv.org/html/2609.09589#bib.bib3)\.
The statistical\-mechanical object considered here differs from the usual parameter\-space loss landscape\. We first express mean\-squared gradient flow directly in error space,
e˙=−Me,\\dot\{e\}=\-Me,\(10\)and show that isotropic stochastic forcing in the error source induces the mobilityK=M2K=M^\{2\}\. This allows the bare conditional dynamics to be represented as an overdamped Langevin process with the quadratic energy
UM\(e\)=12⟨e,M−1e⟩p\.U\_\{M\}\(e\)=\\frac\{1\}\{2\}\\langle e,M^\{\-1\}e\\rangle\_\{p\}\.\(11\)We then combine this dynamical weight with a parameter\-space reference measure and push the resulting microstate ensemble forward to function space\. The induced density of statesΩ\(e\)\\Omega\(e\)counts the multiplicity of parameter microstates associated with the same coarse\-grained function\-space state, producing the effective contribution
−σξ2logΩ\(e\)\-\\sigma\_\{\\xi\}^\{2\}\\log\\Omega\(e\)\(12\)to the conditional free energy\. Thus the entropy in our formulation is associated with parameter\-space multiplicity at fixed function\-space macrostate, rather than solely with the local volume of a loss minimum\.
#### Spectral bias and smooth feature learning\.
Neural networks are known to exhibit a spectral bias toward learning simpler or lower\-frequency components of a target function earlier in training[Rahaman et al\. \(2019\)](https://arxiv.org/html/2609.09589#bib.bib13)\. Such behavior is commonly characterized using Fourier modes, kernel eigenfunctions, or other externally chosen spectral decompositions, and the resulting bias can depend strongly on the geometry of the data distribution\. In the present work, smoothness instead emerges from the same operator geometry that enters feature learning\. We derive a microstate\-curvature operator
Bp\(r\)=σξ2Lp∗𝒦p\(r\)L,B\_\{p\}\(r\)=\\sigma\_\{\\xi\}^\{2\}L\_\{p\}^\{\\ast\}\\mathcal\{K\}\_\{p\}\(r\)L,\(13\)whose eigendirections are defined directly in the data\-weighted function spaceL2\(𝒳,p\)L^\{2\}\(\\mathcal\{X\},p\)\. Under regular structural statistics, its eigenvalues satisfy
bi≍‖Lϕi‖2,b\_\{i\}\\asymp\\\|L\\phi\_\{i\}\\\|^\{2\},\(14\)so that small\-bib\_\{i\}identifies low structural curvature up to the bounded anisotropy of the structural metric\. Combining this with the conditional operator preference gives
mlarge⟷bsmall⟷low structural\-curvature data\-adaptive sector\.m\_\{\\mathrm\{large\}\}\\longleftrightarrow b\_\{\\mathrm\{small\}\}\\longleftrightarrow\\text\{low structural\-curvature data\-adaptive sector\}\.\(15\)The resulting structural bias is therefore not imposed through a fixed Fourier basis or a fixed kernel spectrum; it is defined intrinsically by the data\-weighted microstate geometry of the learned function space\.
#### Piecewise\-linear geometry of ReLU networks\.
ReLU networks represent continuous piecewise\-affine functions whose input space is partitioned into activation regions\. This viewpoint has been developed through spline and piecewise\-linear descriptions of deep networks[Balestriero and Baraniuk \(2018\)](https://arxiv.org/html/2609.09589#bib.bib14), as well as variational and representer theorems connecting ReLU networks to spline\-like function spaces and second\-order regularity[Unser \(2019\)](https://arxiv.org/html/2609.09589#bib.bib15)\. We use this geometric structure as the architectural input to our statistical theory\. Within each activation cell the Hessian vanishes, while second\-order structure is concentrated on cell boundaries through jumps of the gradient\. After coarse graining, this motivates the structural field
h=Lc,L≃𝒞D2\.h=Lc,\\qquad L\\simeq\\mathcal\{C\}D^\{2\}\.\(16\)Under mild statistical boundary conditions on the entropy of these structural states, the corresponding function\-space microstate curvature is
B=σξ2L∗𝒦L\.B=\\sigma\_\{\\xi\}^\{2\}L^\{\\ast\}\\mathcal\{K\}L\.\(17\)The role of ReLU geometry in our framework is therefore to provide the structural operator whose entropy curvature enters the conditional thermodynamic preference analyzed in Sec\.[4](https://arxiv.org/html/2609.09589#S4)\.
## 3Conditional Statistical Mechanics of Error Fluctuations
### 3\.1Exact function\-space gradient dynamics
We begin by expressing gradient descent directly in function space\. Letp\(x\)p\(x\)denote the data distribution and define
ℋ=L2\(𝒳,p\),\\mathcal\{H\}=L^\{2\}\(\\mathcal\{X\},p\),\(18\)with inner product
⟨f,g⟩p=∫𝒳f\(x\)g\(x\)p\(x\)𝑑x\.\\langle f,g\\rangle\_\{p\}=\\int\_\{\\mathcal\{X\}\}f\(x\)g\(x\)\\,p\(x\)\\,dx\.\(19\)For a modelfθf\_\{\\theta\}and target functionyy, define the error field
eθ=fθ−ye\_\{\\theta\}=f\_\{\\theta\}\-y\(20\)and the mean\-squared loss
ℒ\(θ\)=12‖eθ‖p2\.\\mathcal\{L\}\(\\theta\)=\\frac\{1\}\{2\}\\\|e\_\{\\theta\}\\\|\_\{p\}^\{2\}\.\(21\)Let
Jθ=DθfθJ\_\{\\theta\}=D\_\{\\theta\}f\_\{\\theta\}\(22\)denote the Jacobian mapping infinitesimal parameter perturbations to function\-space perturbations\. Its adjointJθ∗J\_\{\\theta\}^\{\\ast\}is defined with respect to the parameter\-space inner product and⟨⋅,⋅⟩p\\langle\\cdot,\\cdot\\rangle\_\{p\}\. Then
∇θℒ=Jθ∗eθ,\\nabla\_\{\\theta\}\\mathcal\{L\}=J\_\{\\theta\}^\{\\ast\}e\_\{\\theta\},\(23\)and continuous\-time gradient flow gives
θ˙=−Jθ∗eθ\.\\dot\{\\theta\}=\-J\_\{\\theta\}^\{\\ast\}e\_\{\\theta\}\.\(24\)Applying the chain rule,
e˙θ=Jθθ˙=−JθJθ∗eθ\.\\dot\{e\}\_\{\\theta\}=J\_\{\\theta\}\\dot\{\\theta\}=\-J\_\{\\theta\}J\_\{\\theta\}^\{\\ast\}e\_\{\\theta\}\.\(25\)This motivates the function\-space dynamical operator
Mθ≡JθJθ∗,\\boxed\{M\_\{\\theta\}\\equiv J\_\{\\theta\}J\_\{\\theta\}^\{\\ast\},\}\(26\)which is self\-adjoint and positive semidefinite\. The exact error dynamics is therefore
e˙θ=−Mθeθ\.\\boxed\{\\dot\{e\}\_\{\\theta\}=\-M\_\{\\theta\}e\_\{\\theta\}\.\}\(27\)Importantly, Eq\. equation[27](https://arxiv.org/html/2609.09589#S3.E27)does not require linearizing the network around a fixed parameter configuration:MθM\_\{\\theta\}may evolve along the training trajectory\.
For a finite\-parameter model,MMhas finite rank\. We work throughout on a fixed finite\-dimensional active sector
ℋa≡RanM¯,n≡dimℋa<∞,\\mathcal\{H\}\_\{a\}\\equiv\\overline\{\\operatorname\{Ran\}M\},\\qquad n\\equiv\\dim\\mathcal\{H\}\_\{a\}<\\infty,\(28\)and writePaP\_\{a\}for the orthogonal projection ontoℋa\\mathcal\{H\}\_\{a\}\. The restriction ofMMtoℋa\\mathcal\{H\}\_\{a\}is taken to be positive definite\. All inverses, traces, determinants, and orthogonal rotations in the following sections are understood onℋa\\mathcal\{H\}\_\{a\}\. Directions inkerM\\ker Mdo not relax under Eq\. equation[27](https://arxiv.org/html/2609.09589#S3.E27)and are excluded from the conditional fluctuation sector\.
### 3\.2Conditional Langevin dynamics and dynamical weight
We next construct the dynamical statistical weight associated with the local error dynamics\. We condition on a macroscopic configuration for whichJJand
are treated as fixed parameters of the conditional problem\. This conditioning defines a family of local ensembles; by itself it does not require a dynamical separation of time scales betweenMMand the error fluctuations\.
Consider the standard constant\-mobility overdamped Langevin equation
dxt=−K∇U\(xt\)dt\+2TKdWt,dx\_\{t\}=\-K\\nabla U\(x\_\{t\}\)\\,dt\+\\sqrt\{2TK\}\\,dW\_\{t\},\(30\)whereUUis an energy,KKis a positive mobility operator, andTTsets the stochastic scale\. Such dynamics admit the standard Itô Fokker–Planck and equilibrium structure[Risken \(1989\)](https://arxiv.org/html/2609.09589#bib.bib1);[Pavliotis \(2014\)](https://arxiv.org/html/2609.09589#bib.bib2)\.
#### Stochastic mobility\.
We model the stochastic source in error space as isotropic,
𝔼\[dWt\]=0,𝔼\[dWtdWt∗\]=Idt\.\\mathbb\{E\}\[dW\_\{t\}\]=0,\\qquad\\mathbb\{E\}\[dW\_\{t\}dW\_\{t\}^\{\\ast\}\]=I\\,dt\.\(31\)In the infinite\-dimensional notation,WtW\_\{t\}may be understood as a cylindrical Wiener process; becauseMMhas finite rank,MdWtM\\,dW\_\{t\}is well defined onℋa\\mathcal\{H\}\_\{a\}\. We assume that this source enters parameter space through the same Jacobian channel as the deterministic gradient,
dθnoise=2σξJ∗dWt\.d\\theta\_\{\\mathrm\{noise\}\}=\\sqrt\{2\}\\,\\sigma\_\{\\xi\}J^\{\\ast\}dW\_\{t\}\.\(32\)WithJJfixed in the conditional construction,
denoise\\displaystyle de\_\{\\mathrm\{noise\}\}=Jdθnoise\\displaystyle=J\\,d\\theta\_\{\\mathrm\{noise\}\}=2σξJJ∗dWt=2σξMdWt\.\\displaystyle=\\sqrt\{2\}\\,\\sigma\_\{\\xi\}JJ^\{\\ast\}dW\_\{t\}=\\sqrt\{2\}\\,\\sigma\_\{\\xi\}M\\,dW\_\{t\}\.\(33\)Combining this with Eq\. equation[27](https://arxiv.org/html/2609.09589#S3.E27)gives the bare conditional process
de=−Medt\+2σξMdWt\.\\boxed\{de=\-Me\\,dt\+\\sqrt\{2\}\\,\\sigma\_\{\\xi\}M\\,dW\_\{t\}\.\}\(34\)Its stochastic increment has covariance
𝔼\[denoisedenoise∗\]=2σξ2M2dt,\\mathbb\{E\}\[de\_\{\\mathrm\{noise\}\}de\_\{\\mathrm\{noise\}\}^\{\\ast\}\]=2\\sigma\_\{\\xi\}^\{2\}M^\{2\}\\,dt,\(35\)so comparison with Eq\. equation[30](https://arxiv.org/html/2609.09589#S3.E30)identifies
K=M2,T=σξ2\.\\boxed\{K=M^\{2\}\},\\qquad T=\\sigma\_\{\\xi\}^\{2\}\.\(36\)
#### Dynamical energy\.
To represent the deterministic drift in Langevin form, we seekUM\(e\)U\_\{M\}\(e\)such that
K∇eUM\(e\)=Me\.K\\nabla\_\{e\}U\_\{M\}\(e\)=Me\.\(37\)Onℋa\\mathcal\{H\}\_\{a\},MMis invertible, and withK=M2K=M^\{2\}this gives
∇eUM=M−1e\.\\nabla\_\{e\}U\_\{M\}=M^\{\-1\}e\.\(38\)Hence
UM\(e\)=12⟨e,M−1e⟩p\.\\boxed\{U\_\{M\}\(e\)=\\frac\{1\}\{2\}\\langle e,M^\{\-1\}e\\rangle\_\{p\}\.\}\(39\)Equation equation[34](https://arxiv.org/html/2609.09589#S3.E34)can therefore be written as
de=−M2∇eUM\(e\)dt\+2σξ2M2dWt\.de=\-M^\{2\}\\nabla\_\{e\}U\_\{M\}\(e\)\\,dt\+\\sqrt\{2\\sigma\_\{\\xi\}^\{2\}M^\{2\}\}\\,dW\_\{t\}\.\(40\)
###### Theorem 3\.1\(Bare conditional Langevin equilibrium\)\.
For the conditional process Eq\. equation[40](https://arxiv.org/html/2609.09589#S3.E40)onℋa\\mathcal\{H\}\_\{a\}, the densityp\(e,t∣M\)p\(e,t\\mid M\)satisfies
∂tp=∇e⋅\[M2\(p∇eUM\+σξ2∇ep\)\]\.\\boxed\{\\partial\_\{t\}p=\\nabla\_\{e\}\\cdot\\left\[M^\{2\}\\left\(p\\nabla\_\{e\}U\_\{M\}\+\\sigma\_\{\\xi\}^\{2\}\\nabla\_\{e\}p\\right\)\\right\]\.\}\(41\)Its reversible stationary density with respect to the reference volumededeonℋa\\mathcal\{H\}\_\{a\}is
p0\(e∣M\)=1Z0exp\[−UM\(e\)σξ2\]\.\\boxed\{p\_\{0\}\(e\\mid M\)=\\frac\{1\}\{Z\_\{0\}\}\\exp\\left\[\-\\frac\{U\_\{M\}\(e\)\}\{\\sigma\_\{\\xi\}^\{2\}\}\\right\]\.\}\(42\)Moreover,
ℱ0\[p\]=∫p\(e\)UM\(e\)𝑑e\+σξ2∫p\(e\)logp\(e\)𝑑e\\boxed\{\\mathcal\{F\}\_\{0\}\[p\]=\\int p\(e\)U\_\{M\}\(e\)\\,de\+\\sigma\_\{\\xi\}^\{2\}\\int p\(e\)\\log p\(e\)\\,de\}\(43\)is non\-increasing along the conditional Fokker–Planck dynamics:
dℱ0dt=−∫p\(e\)⟨∇eδℱ0δp,M2∇eδℱ0δp⟩pde≤0\.\\boxed\{\\frac\{d\\mathcal\{F\}\_\{0\}\}\{dt\}=\-\\int p\(e\)\\left\\langle\\nabla\_\{e\}\\frac\{\\delta\\mathcal\{F\}\_\{0\}\}\{\\delta p\},M^\{2\}\\nabla\_\{e\}\\frac\{\\delta\\mathcal\{F\}\_\{0\}\}\{\\delta p\}\\right\\rangle\_\{p\}de\\leq 0\.\}\(44\)
###### Proof\.
The Itô forward equation for Eq\. equation[40](https://arxiv.org/html/2609.09589#S3.E40)is Eq\. equation[41](https://arxiv.org/html/2609.09589#S3.E41), with probability current
j=−M2\(p∇eUM\+σξ2∇ep\)\.j=\-M^\{2\}\\left\(p\\nabla\_\{e\}U\_\{M\}\+\\sigma\_\{\\xi\}^\{2\}\\nabla\_\{e\}p\\right\)\.\(45\)For Eq\. equation[42](https://arxiv.org/html/2609.09589#S3.E42),∇ep0=−\(p0/σξ2\)∇eUM\\nabla\_\{e\}p\_\{0\}=\-\(p\_\{0\}/\\sigma\_\{\\xi\}^\{2\}\)\\nabla\_\{e\}U\_\{M\}, soj0=0j\_\{0\}=0\. The variational derivative of Eq\. equation[43](https://arxiv.org/html/2609.09589#S3.E43)is
δℱ0δp=UM\+σξ2logp\+const,\\frac\{\\delta\\mathcal\{F\}\_\{0\}\}\{\\delta p\}=U\_\{M\}\+\\sigma\_\{\\xi\}^\{2\}\\log p\+\\mathrm\{const\},\(46\)and an integration by parts gives Eq\. equation[44](https://arxiv.org/html/2609.09589#S3.E44)\. ∎
Theorem[3\.1](https://arxiv.org/html/2609.09589#S3.Thmtheorem1)supplies the*dynamical*Boltzmann weight relative to the reference volume in function space\. It does not count how many parameter configurations realize a given error state\. In particular, the bare noise process above is not assumed to explore parameter\-fiber directions\. Parameter multiplicity enters independently through the reference microstate measure introduced next\.
### 3\.3Parameter microstates and the induced conditional ensemble
We take the background statezzto include the conditioned dynamical operatorMMtogether with the remaining constraints supplied by the parameterization, architecture, and data, and letνz\(dθ\)\\nu\_\{z\}\(d\\theta\)denote the corresponding reference measure over parameter microstates\. To define a regular density of states without invoking an infinite\-dimensional volume element, introduce a finite\-dimensional retained function\-space sector
ℋa⊆ℋR⊆ℋ,NR≡dimℋR<∞,\\mathcal\{H\}\_\{a\}\\subseteq\\mathcal\{H\}\_\{R\}\\subseteq\\mathcal\{H\},\\qquad N\_\{R\}\\equiv\\dim\\mathcal\{H\}\_\{R\}<\\infty,\(47\)with orthogonal projectorPRP\_\{R\}\. The sectorℋR\\mathcal\{H\}\_\{R\}contains the active learning sectorℋa=RanM\\mathcal\{H\}\_\{a\}=\\operatorname\{Ran\}Mbut may also retain additional coarse\-grained function coordinates needed to define the microstate statistics\. We define
ΨR,z:θ↦eR≡PR\(fθ−y\)∈ℋR\\Psi\_\{R,z\}:\\theta\\mapsto e\_\{R\}\\equiv P\_\{R\}\\\!\\left\(f\_\{\\theta\}\-y\\right\)\\in\\mathcal\{H\}\_\{R\}\(48\)and the corresponding pushforward reference measure
μz=\(ΨR,z\)\#νz\.\\boxed\{\\mu\_\{z\}=\(\\Psi\_\{R,z\}\)\_\{\\\#\}\\nu\_\{z\}\.\}\(49\)HeredeRde\_\{R\}denotes ordinary Lebesgue volume in an orthonormal coordinate system onℋR\\mathcal\{H\}\_\{R\}\. We assume that the coarse\-grained pushforward is regular enough to admit a density,
μz\(deR\)=Ωz\(eR\)deR\.\\boxed\{\\mu\_\{z\}\(de\_\{R\}\)=\\Omega\_\{z\}\(e\_\{R\}\)\\,de\_\{R\}\.\}\(50\)ThusΩz\\Omega\_\{z\}is a density of parameter microstates on the retained coarse\-grained function sectorℋR\\mathcal\{H\}\_\{R\}\. No density with respect to an infinite\-dimensional “function\-space volume” is assumed\. The operator calculations below use only the compression of its local curvature to the active subspaceℋa\\mathcal\{H\}\_\{a\}\.
The following factorization is the central statistical assumption that combines the two ingredients above\.
#### A1\. Factorization of dynamical weight and microstate multiplicity\.
At fixed background statezzand dynamical operatorMM, we assume that the microscopic statistical weight factorizes into a dynamical factor that depends on a parameter configuration only through its induced error state and a reference microstate measure that supplies the multiplicity of such realizations\. Equivalently,
Π∗\(dθ∣M,z\)∝exp\[−UM\(PaΨR,z\(θ\)\)σξ2\]νz\(dθ\)\.\\boxed\{\\Pi\_\{\\ast\}\(d\\theta\\mid M,z\)\\propto\\exp\\left\[\-\\frac\{U\_\{M\}\(P\_\{a\}\\Psi\_\{R,z\}\(\\theta\)\)\}\{\\sigma\_\{\\xi\}^\{2\}\}\\right\]\\nu\_\{z\}\(d\\theta\)\.\}\(51\)The reference measureνz\\nu\_\{z\}is not assumed to be dynamically sampled by the bareJ∗J^\{\\ast\}\-channel noise\. Rather, it encodes the conditional multiplicity supplied by the parameterization and other degrees of freedom included inzz\. Assumption A1 states that this multiplicity can be combined with the bare dynamical weighting without an additional fiber\-dependent energetic factor\. Possible dependence of the reference microstate statistics on slower variables, including changes induced by an actual motion ofMM, is outside the partial conditional comparison performed below\.
Under A1, the bare conditional dynamics supplies the Boltzmann factorexp\[−UM\(PaeR\)/σξ2\]\\exp\[\-U\_\{M\}\(P\_\{a\}e\_\{R\}\)/\\sigma\_\{\\xi\}^\{2\}\]on the active component, while the pushforward ofνz\\nu\_\{z\}supplies the density of states onℋR\\mathcal\{H\}\_\{R\}\. Their product defines the conditional canonical ensemble
P∗\(deR∣M,z\)=1Zexp\[−UM\(PaeR\)σξ2\]μz\(deR\)\.\\boxed\{P\_\{\\ast\}\(de\_\{R\}\\mid M,z\)=\\frac\{1\}\{Z\}\\exp\\left\[\-\\frac\{U\_\{M\}\(P\_\{a\}e\_\{R\}\)\}\{\\sigma\_\{\\xi\}^\{2\}\}\\right\]\\mu\_\{z\}\(de\_\{R\}\)\.\}\(52\)By Assumption A1, Eq\. equation[51](https://arxiv.org/html/2609.09589#S3.E51)is the corresponding parameter\-space ensemble, and its pushforward underΨR,z\\Psi\_\{R,z\}is Eq\. equation[52](https://arxiv.org/html/2609.09589#S3.E52)\. If Eq\. equation[50](https://arxiv.org/html/2609.09589#S3.E50)holds, then
p∗\(eR∣M,z\)=1ZΩz\(eR\)exp\[−UM\(PaeR\)σξ2\]\.\\boxed\{p\_\{\\ast\}\(e\_\{R\}\\mid M,z\)=\\frac\{1\}\{Z\}\\Omega\_\{z\}\(e\_\{R\}\)\\exp\\left\[\-\\frac\{U\_\{M\}\(P\_\{a\}e\_\{R\}\)\}\{\\sigma\_\{\\xi\}^\{2\}\}\\right\]\.\}\(53\)
The two factors in Eq\. equation[53](https://arxiv.org/html/2609.09589#S3.E53)have distinct origins\. The exponential factor is the dynamical weight derived from the bare conditional Langevin process, whereasΩz\(eR\)\\Omega\_\{z\}\(e\_\{R\}\)is a static density\-of\-states factor supplied by the parameterization\. Their product is therefore an ensemble construction; Eq\. equation[53](https://arxiv.org/html/2609.09589#S3.E53)is not claimed to be the stationary law generated by Eq\. equation[34](https://arxiv.org/html/2609.09589#S3.E34)alone\.
Defining
FM\(eR,z\)=UM\(PaeR\)−σξ2logΩz\(eR\),\\boxed\{F\_\{M\}\(e\_\{R\};z\)=U\_\{M\}\(P\_\{a\}e\_\{R\}\)\-\\sigma\_\{\\xi\}^\{2\}\\log\\Omega\_\{z\}\(e\_\{R\}\),\}\(54\)the conditional ensemble takes the Gibbs form
p∗\(eR∣M,z\)=1Zexp\[−FM\(eR,z\)σξ2\]\.\\boxed\{p\_\{\\ast\}\(e\_\{R\}\\mid M,z\)=\\frac\{1\}\{Z\}\\exp\\left\[\-\\frac\{F\_\{M\}\(e\_\{R\};z\)\}\{\\sigma\_\{\\xi\}^\{2\}\}\\right\]\.\}\(55\)
###### Proposition 3\.2\(Microstate\-weighted conditional ensemble\)\.
Letνz\\nu\_\{z\}be a parameter\-space reference measure and letμz=\(ΨR,z\)\#νz\\mu\_\{z\}=\(\\Psi\_\{R,z\}\)\_\{\\\#\}\\nu\_\{z\}\. Ifμz\(deR\)=Ωz\(eR\)deR\\mu\_\{z\}\(de\_\{R\}\)=\\Omega\_\{z\}\(e\_\{R\}\)\\,de\_\{R\}, then the factorized canonical weighting Eq\. equation[51](https://arxiv.org/html/2609.09589#S3.E51)induces the function\-space ensemble Eq\. equation[53](https://arxiv.org/html/2609.09589#S3.E53), equivalently the Gibbs form Eq\. equation[55](https://arxiv.org/html/2609.09589#S3.E55)generated byFM\(eR,z\)F\_\{M\}\(e\_\{R\};z\)\.
###### Proof\.
For any measurable setAA,
Π∗\(ΨR,z−1\(A\)∣M,z\)\\displaystyle\\Pi\_\{\\ast\}\(\\Psi\_\{R,z\}^\{\-1\}\(A\)\\mid M,z\)∝∫ΨR,z−1\(A\)e−UM\(PaΨR,z\(θ\)\)/σξ2νz\(dθ\)\\displaystyle\\propto\\int\_\{\\Psi\_\{R,z\}^\{\-1\}\(A\)\}e^\{\-U\_\{M\}\(P\_\{a\}\\Psi\_\{R,z\}\(\\theta\)\)/\\sigma\_\{\\xi\}^\{2\}\}\\nu\_\{z\}\(d\\theta\)=∫Ae−UM\(PaeR\)/σξ2μz\(deR\)\.\\displaystyle=\\int\_\{A\}e^\{\-U\_\{M\}\(P\_\{a\}e\_\{R\}\)/\\sigma\_\{\\xi\}^\{2\}\}\\mu\_\{z\}\(de\_\{R\}\)\.\(56\)Substitutingμz\(deR\)=Ωz\(eR\)deR\\mu\_\{z\}\(de\_\{R\}\)=\\Omega\_\{z\}\(e\_\{R\}\)\\,de\_\{R\}gives Eq\. equation[53](https://arxiv.org/html/2609.09589#S3.E53)\. ∎
The parameter\-space pushforward therefore contributes the entropic term−σξ2logΩz\(eR\)\-\\sigma\_\{\\xi\}^\{2\}\\log\\Omega\_\{z\}\(e\_\{R\}\)to the effective retained\-sector free energy, independently of the bare stochastic channel used to identifyUMU\_\{M\}\.
### 3\.4Local microstate curvature and the conditional fluctuation ensemble
We now condition on a current retained error macrostater∈ℋRr\\in\\mathcal\{H\}\_\{R\}and characterize fluctuations along the active learning sector,
eR=r\+δe,δe∈ℋa\.e\_\{R\}=r\+\\delta e,\\qquad\\delta e\\in\\mathcal\{H\}\_\{a\}\.\(57\)Writera=Parr\_\{a\}=P\_\{a\}rfor the active component of the conditioned macrostate\. The pointrris a conditioning variable and is not assumed to minimizeFMF\_\{M\}\. The standard constrained\-ensemble construction introduces a linear source conjugate to the macrostate\. In a full maximum\-entropy formulation, the source is the Lagrange multiplier enforcing the chosen mean macrostate\. At the local Laplace level used here, conditioning onrramounts to choosing the source so thatrris a stationary point of the tilted potential,
λr≡Pa∇eRFM\(eR,z\)\|eR=r\.\\boxed\{\\lambda\_\{r\}\\equiv P\_\{a\}\\nabla\_\{e\_\{R\}\}F\_\{M\}\(e\_\{R\};z\)\\big\|\_\{e\_\{R\}=r\}\.\}\(58\)It is useful to separate the zero\-order macrostate cost from the fluctuation cost and define the local excess potential
ΔF~M,r\(δe,z\)≡FM\(r\+δe,z\)−FM\(r,z\)−⟨λr,δe⟩p\.\\boxed\{\\Delta\\widetilde\{F\}\_\{M,r\}\(\\delta e;z\)\\equiv F\_\{M\}\(r\+\\delta e;z\)\-F\_\{M\}\(r;z\)\-\\langle\\lambda\_\{r\},\\delta e\\rangle\_\{p\}\.\}\(59\)ThenΔF~M,r\(0,z\)=0\\Delta\\widetilde\{F\}\_\{M,r\}\(0;z\)=0and its first variation vanishes atδe=0\\delta e=0, while its Hessian is exactly the Hessian ofFMF\_\{M\}atrr\. We do not requirerrto be the unconstrained mean or mode of the full non\-Gaussian Gibbs measure; the construction is the quadratic local form of the usual Legendre/Lagrange constrained ensemble\.
Define the retained\-sector microstate\-curvature form by
B\(r\)≡−σξ2∇eR2logΩz\(eR\)\|eR=r\.\\boxed\{B\(r\)\\equiv\-\\sigma\_\{\\xi\}^\{2\}\\nabla\_\{e\_\{R\}\}^\{2\}\\log\\Omega\_\{z\}\(e\_\{R\}\)\\big\|\_\{e\_\{R\}=r\}\.\}\(60\)The operator entering the finite\-dimensional fluctuation sector is its compression
Ba\(r\)≡PaB\(r\)Pa\|ℋa\.\\boxed\{B\_\{a\}\(r\)\\equiv P\_\{a\}B\(r\)P\_\{a\}\\big\|\_\{\\mathcal\{H\}\_\{a\}\}\.\}\(61\)Since the dynamical termUM\(PaeR\)U\_\{M\}\(P\_\{a\}e\_\{R\}\)has active\-sector Hessian
∇ℋa2UM=M−1\\nabla\_\{\\mathcal\{H\}\_\{a\}\}^\{2\}U\_\{M\}=M^\{\-1\}\(62\), a second\-order expansion of Eq\. equation[59](https://arxiv.org/html/2609.09589#S3.E59)gives
ΔF~M,r\(δe,z\)=12⟨δe,Ha\(r\)δe⟩p\+o\(‖δe‖2\),\\Delta\\widetilde\{F\}\_\{M,r\}\(\\delta e;z\)=\\frac\{1\}\{2\}\\left\\langle\\delta e,H\_\{a\}\(r\)\\delta e\\right\\rangle\_\{p\}\+o\(\\\|\\delta e\\\|^\{2\}\),\(63\)where
Ha\(r\)≡M−1\+Ba\(r\)\.\\boxed\{H\_\{a\}\(r\)\\equiv M^\{\-1\}\+B\_\{a\}\(r\)\.\}\(64\)
###### Proposition 3\.3\(Local conditional fluctuation ensemble\)\.
Suppose that
Ha\(r\)=M−1\+Ba\(r\)≻0H\_\{a\}\(r\)=M^\{\-1\}\+B\_\{a\}\(r\)\\succ 0\(65\)onℋa\\mathcal\{H\}\_\{a\}\. Then, to quadratic order around the conditioned macrostaterr, the local fluctuation ensemble is Gaussian:
ploc\(δe∣r,M,z\)=1Zlocexp\[−12σξ2⟨δe,Ha\(r\)δe⟩p\],\\boxed\{p\_\{\\mathrm\{loc\}\}\(\\delta e\\mid r,M,z\)=\\frac\{1\}\{Z\_\{\\mathrm\{loc\}\}\}\\exp\\left\[\-\\frac\{1\}\{2\\sigma\_\{\\xi\}^\{2\}\}\\left\\langle\\delta e,H\_\{a\}\(r\)\\delta e\\right\\rangle\_\{p\}\\right\],\}\(66\)with covariance
C∗\(r,M\)=σξ2\[M−1\+Ba\(r\)\]−1\.\\boxed\{C\_\{\\ast\}\(r,M\)=\\sigma\_\{\\xi\}^\{2\}\\left\[M^\{\-1\}\+B\_\{a\}\(r\)\\right\]^\{\-1\}\.\}\(67\)
###### Proof\.
Equation equation[63](https://arxiv.org/html/2609.09589#S3.E63)is quadratic with positive\-definite HessianHa\(r\)H\_\{a\}\(r\)\. The normalized Gaussian measure therefore has precisionHa\(r\)/σξ2H\_\{a\}\(r\)/\\sigma\_\{\\xi\}^\{2\}and covarianceσξ2Ha\(r\)−1\\sigma\_\{\\xi\}^\{2\}H\_\{a\}\(r\)^\{\-1\}\. ∎
Proposition[3\.3](https://arxiv.org/html/2609.09589#S3.Thmtheorem3)is a conditional ensemble statement\. No separation between the relaxation time ofδe\\delta eand the evolution time ofrrorMMis required for the algebraic results below\. Interpreting the same ensemble as an adiabatically realized quasi\-equilibrium along an actual training trajectory would require an additional local\-equilibration assumption, which we do not use here\.
The conditional Gaussian sector is the starting point for the next section\. The local Laplace expansion separates the macrostate cost from the fluctuation cost\. To quadratic order,
𝒢loc\(r,M,z\)=FM\(r,z\)\+Φfluc\(M,Ba\(r\)\)\+const,\\mathcal\{G\}\_\{\\mathrm\{loc\}\}\(r,M;z\)=F\_\{M\}\(r;z\)\+\\Phi\_\{\\mathrm\{fluc\}\}\(M;B\_\{a\}\(r\)\)\+\\mathrm\{const\},\(68\)whereΦfluc\\Phi\_\{\\mathrm\{fluc\}\}is obtained by integrating the excess fluctuations\. The next section isolates this fluctuation\-induced contribution and asks what orientational preference it supplies\.
## 4Operator Matching under Thermodynamic Stability
Section[3\.4](https://arxiv.org/html/2609.09589#S3.SS4)showed that, at a conditioned macrostaterr, the local fluctuation ensemble on the active sectorℋa\\mathcal\{H\}\_\{a\}is governed by
Ha\(r\)=M−1\+Ba\(r\),H\_\{a\}\(r\)=M^\{\-1\}\+B\_\{a\}\(r\),\(69\)with
Ba\(r\)=PaB\(r\)Pa\|ℋa\.B\_\{a\}\(r\)=P\_\{a\}B\(r\)P\_\{a\}\\big\|\_\{\\mathcal\{H\}\_\{a\}\}\.\(70\)We now ask what orientational preference is contributed by this conditional fluctuation sector\.
Throughout this section,rrand the compressed statistical geometryBa\(r\)B\_\{a\}\(r\)are held fixed, and we write
B≡Ba\(r\)B\\equiv B\_\{a\}\(r\)\(71\)for brevity\. This is a partial, conditional comparison: if an actual parameter\-space motion that rotatesMMalso changesrrorBB, those responses contribute additional terms to the full slow dynamics and are not included in the derivative computed here\. The comparison is restricted to orthogonal rotations ofMMwithin the fixed active sectorℋa\\mathcal\{H\}\_\{a\}, so its eigenvalues, rank, and active subspace remain unchanged\.
We focus on the thermodynamically stable sector
B⪰0\.\\boxed\{B\\succeq 0\.\}\(72\)Equivalently, the compressed local entropy curvature is concave onℋa\\mathcal\{H\}\_\{a\}\. SinceM−1≻0M^\{\-1\}\\succ 0, this condition guarantees
OM−1O∗\+B≻0OM^\{\-1\}O^\{\\ast\}\+B\\succ 0\(73\)for every orthogonalOOacting onℋa\\mathcal\{H\}\_\{a\}\. Thus the local Gaussian ensemble remains normalizable over the entire fixed\-spectrum orbit\.
Condition equation[72](https://arxiv.org/html/2609.09589#S4.E72)is a stability restriction, not a consequence of ReLU geometry alone\. In Sec\.[5](https://arxiv.org/html/2609.09589#S5), we identify a regular ReLU statistical sector, defined in part by a positive structural entropy curvature, in which this condition is realized\.
### 4\.1Conditional operator free\-energy contribution
From Proposition[3\.3](https://arxiv.org/html/2609.09589#S3.Thmtheorem3), the local conditional distribution at fixed\(r,M,B\)\(r,M,B\)is
ploc\(δe∣r,M,B\)=1Zloc\(M∣B\)exp\[−12σξ2⟨δe,\(M−1\+B\)δe⟩p\]\.p\_\{\\mathrm\{loc\}\}\(\\delta e\\mid r,M,B\)=\\frac\{1\}\{Z\_\{\\mathrm\{loc\}\}\(M\\mid B\)\}\\exp\\left\[\-\\frac\{1\}\{2\\sigma\_\{\\xi\}^\{2\}\}\\left\\langle\\delta e,\(M^\{\-1\}\+B\)\\delta e\\right\\rangle\_\{p\}\\right\]\.\(74\)The Gaussian partition function on thenn\-dimensional active sector is
Zloc\(M∣B\)=\(2πσξ2\)n/2detℋa\(M−1\+B\)−1/2\.\\boxed\{Z\_\{\\mathrm\{loc\}\}\(M\\mid B\)=\(2\\pi\\sigma\_\{\\xi\}^\{2\}\)^\{n/2\}\\det\_\{\\mathcal\{H\}\_\{a\}\}\(M^\{\-1\}\+B\)^\{\-1/2\}\.\}\(75\)Integrating over the conditional fluctuations therefore contributes
Φfluc\(M,B\)\\displaystyle\\Phi\_\{\\mathrm\{fluc\}\}\(M;B\)≡−σξ2logZloc\(M∣B\)\\displaystyle\\equiv\-\\sigma\_\{\\xi\}^\{2\}\\log Z\_\{\\mathrm\{loc\}\}\(M\\mid B\)=σξ22logdetℋa\(M−1\+B\)\+const\.\\displaystyle=\\boxed\{\\frac\{\\sigma\_\{\\xi\}^\{2\}\}\{2\}\\log\\det\_\{\\mathcal\{H\}\_\{a\}\}\(M^\{\-1\}\+B\)\}\+\\mathrm\{const\}\.\(76\)The omitted constant depends onnnandσξ2\\sigma\_\{\\xi\}^\{2\}but is constant on the fixed\-rank, fixed\-spectrum orbit studied below; it must be restored when comparing sectors of different rank or different active subspaces\. For brevity, we writeΦ≡Φfluc\\Phi\\equiv\\Phi\_\{\\mathrm\{fluc\}\}throughout the remainder of this section\.
Equation equation[68](https://arxiv.org/html/2609.09589#S3.E68)makes clear thatΦfluc\\Phi\_\{\\mathrm\{fluc\}\}is only one contribution to the total local free energy\. In particular,
FM\(r,z\)=12⟨ra,M−1ra⟩p−σξ2logΩz\(r\)F\_\{M\}\(r;z\)=\\frac\{1\}\{2\}\\langle r\_\{a\},M^\{\-1\}r\_\{a\}\\rangle\_\{p\}\-\\sigma\_\{\\xi\}^\{2\}\\log\\Omega\_\{z\}\(r\)\(77\)contains its own orientational dependence through
12⟨ra,M−1ra⟩p=12Tr\(M−1rara∗\)\.\\frac\{1\}\{2\}\\langle r\_\{a\},M^\{\-1\}r\_\{a\}\\rangle\_\{p\}=\\frac\{1\}\{2\}\\operatorname\{Tr\}\\\!\\left\(M^\{\-1\}r\_\{a\}r\_\{a\}^\{\\ast\}\\right\)\.\(78\)At fixed spectrum, this rank\-one macrostate term is minimized when the current active residual direction is aligned with the largest eigenvalue ofMM, i\.e\. with the fastest relaxation direction\. This error\-directed preference is distinct from the fluctuation\-induced matching preference studied below\. Their relative magnitude depends on the conditioned state, spectrum, parameterization, and training conditions, and we make no universal ordering between them\. The present analysis therefore isolates the fluctuation contribution rather than treating it as a proxy for the total orientational free energy\.
Our theorems characterize the thermodynamic preference and generalized force supplied specifically byΦfluc\\Phi\_\{\\mathrm\{fluc\}\}; they do not claim to minimize the complete local free energy or to determine the full slow dynamics ofMM\. Appendix[A](https://arxiv.org/html/2609.09589#A1)shows that, for residual\-preserving rotations, the macrostate term is exactly constant, so the fluctuation contribution can also be isolated as a strict orientational statement on that restricted orbit\.
Using
M−1\+B=M−1\(I\+MB\),M^\{\-1\}\+B=M^\{\-1\}\(I\+MB\),\(79\)we may write
Φ\(M;B\)=σξ22\[−logdetM\+logdet\(I\+MB\)\]\+const,\\boxed\{\\Phi\(M;B\)=\\frac\{\\sigma\_\{\\xi\}^\{2\}\}\{2\}\\left\[\-\\log\\det M\+\\log\\det\(I\+MB\)\\right\]\+\\mathrm\{const\},\}\(80\)where all determinants are onℋa\\mathcal\{H\}\_\{a\}\. Along a fixed\-spectrum orbit,detM\\det Mis constant, so the orientational contribution is
Φorient\(M,B\)=σξ22logdet\(I\+MB\)\.\\boxed\{\\Phi\_\{\\mathrm\{orient\}\}\(M;B\)=\\frac\{\\sigma\_\{\\xi\}^\{2\}\}\{2\}\\log\\det\(I\+MB\)\.\}\(81\)The matching theorem below is exact throughout the stable sector, but the magnitude of the orientational bias depends on the dimensionless coupling betweenMMandBB\. In particular, on a subspace whereB≻0B\\succ 0and all eigenvalues ofM1/2BM1/2M^\{1/2\}BM^\{1/2\}are asymptotically large,
logdet\(I\+MB\)=logdetM\+logdetB\+o\(1\),\\log\\det\(I\+MB\)=\\log\\det M\+\\log\\det B\+o\(1\),\(82\)so the leading orientational dependence disappears\. Thus the conditional matching contribution is most pronounced outside this deep strong\-coupling limit; the result remains valid there, but its orientational force becomes parametrically weak\.
### 4\.2Rotational stationarity and operator matching
Let the spectrum ofMMbe fixed and consider an infinitesimal orthogonal rotation withinℋa\\mathcal\{H\}\_\{a\},
M\(t\)=O\(t\)MO\(t\)∗,O\(t\)=etΞ,Ξ∗=−Ξ\.M\(t\)=O\(t\)MO\(t\)^\{\\ast\},\\qquad O\(t\)=e^\{t\\Xi\},\\qquad\\Xi^\{\\ast\}=\-\\Xi\.\(83\)Then
M˙=\[Ξ,M\],ddtM−1=\[Ξ,M−1\]\.\\dot\{M\}=\[\\Xi,M\],\\qquad\\frac\{d\}\{dt\}M^\{\-1\}=\[\\Xi,M^\{\-1\}\]\.\(84\)
###### Theorem 4\.1\(Operator\-matching condition for the conditional contribution\)\.
LetM≻0M\\succ 0andB=B∗B=B^\{\\ast\}onℋa\\mathcal\{H\}\_\{a\}, withM−1\+B≻0M^\{\-1\}\+B\\succ 0\. The conditional free\-energy contribution
Φ\(M,B\)=σξ22logdet\(M−1\+B\)\\Phi\(M;B\)=\\frac\{\\sigma\_\{\\xi\}^\{2\}\}\{2\}\\log\\det\(M^\{\-1\}\+B\)\(85\)is rotationally stationary on the fixed\-spectrum orbit ofMMif and only if
\[M,B\]=0\.\\boxed\{\[M,B\]=0\.\}\(86\)
###### Proof\.
LetH=M−1\+BH=M^\{\-1\}\+B\. Along Eq\. equation[83](https://arxiv.org/html/2609.09589#S4.E83),
dΦdt\\displaystyle\\frac\{d\\Phi\}\{dt\}=σξ22Tr\(H−1\[Ξ,M−1\]\)\\displaystyle=\\frac\{\\sigma\_\{\\xi\}^\{2\}\}\{2\}\\operatorname\{Tr\}\\left\(H^\{\-1\}\[\\Xi,M^\{\-1\}\]\\right\)=σξ22Tr\(\[M−1,H−1\]Ξ\)\.\\displaystyle=\\frac\{\\sigma\_\{\\xi\}^\{2\}\}\{2\}\\operatorname\{Tr\}\\left\(\[M^\{\-1\},H^\{\-1\}\]\\Xi\\right\)\.\(87\)BecauseM−1M^\{\-1\}andH−1H^\{\-1\}are self\-adjoint,\[M−1,H−1\]\[M^\{\-1\},H^\{\-1\}\]is anti\-self\-adjoint\. Stationarity for all anti\-self\-adjointΞ\\Xiis therefore equivalent to
\[M−1,H−1\]=0\.\[M^\{\-1\},H^\{\-1\}\]=0\.\(88\)SinceHHis invertible, this is equivalent to\[M−1,H\]=0\[M^\{\-1\},H\]=0, hence to\[M−1,B\]=0\[M^\{\-1\},B\]=0, and therefore to\[M,B\]=0\[M,B\]=0\. ∎
Theorem[4\.1](https://arxiv.org/html/2609.09589#S4.Thmtheorem1)characterizes the stationary orientations preferred by the conditional fluctuation contribution\. It does not assert that the complete parameter dynamics necessarily drivesMMto such a point\.
### 4\.3Reverse spectral pairing
At a rotationally stationary point,MMandBBshare an eigenbasis\. Let
m1≥m2≥⋯≥mn\>0m\_\{1\}\\geq m\_\{2\}\\geq\\cdots\\geq m\_\{n\}\>0\(89\)and
0≤b1≤b2≤⋯≤bn\.0\\leq b\_\{1\}\\leq b\_\{2\}\\leq\\cdots\\leq b\_\{n\}\.\(90\)At a commuting configuration specified by a permutationπ\\pi,
M=diag\(m1,…,mn\),B=diag\(bπ\(1\),…,bπ\(n\)\),M=\\operatorname\{diag\}\(m\_\{1\},\\ldots,m\_\{n\}\),\\qquad B=\\operatorname\{diag\}\(b\_\{\\pi\(1\)\},\\ldots,b\_\{\\pi\(n\)\}\),\(91\)and
Φorient=σξ22∑i=1nlog\(1\+mibπ\(i\)\)\.\\Phi\_\{\\mathrm\{orient\}\}=\\frac\{\\sigma\_\{\\xi\}^\{2\}\}\{2\}\\sum\_\{i=1\}^\{n\}\\log\(1\+m\_\{i\}b\_\{\\pi\(i\)\}\)\.\(92\)
###### Theorem 4\.2\(Reverse spectral pairing\)\.
Within the stable sectorB⪰0B\\succeq 0, the conditional orientational free\-energy contribution
Φorient\(M,B\)=σξ22logdet\(I\+MB\)\\Phi\_\{\\mathrm\{orient\}\}\(M;B\)=\\frac\{\\sigma\_\{\\xi\}^\{2\}\}\{2\}\\log\\det\(I\+MB\)\(93\)is globally minimized when the spectra ofMMandBBare oppositely ordered:
mlarge⟷bsmall\.\\boxed\{m\_\{\\mathrm\{large\}\}\\longleftrightarrow b\_\{\\mathrm\{small\}\}\.\}\(94\)With the ordering above, a minimizing configuration is
M=diag\(m1,…,mn\),B=diag\(b1,…,bn\)\.M=\\operatorname\{diag\}\(m\_\{1\},\\ldots,m\_\{n\}\),\\qquad B=\\operatorname\{diag\}\(b\_\{1\},\\ldots,b\_\{n\}\)\.\(95\)Same\-order pairing gives the global maximum\. If both spectra are nondegenerate, all other commuting permutations are saddles\.
###### Proof\.
The fixed\-spectrum orthogonal orbit is compact, and Eq\. equation[73](https://arxiv.org/html/2609.09589#S4.E73)ensures thatΦorient\\Phi\_\{\\mathrm\{orient\}\}is continuous on the full orbit\. Any global extremum is therefore stationary and, by Theorem[4\.1](https://arxiv.org/html/2609.09589#S4.Thmtheorem1), commuting\.
Considermi\>mjm\_\{i\}\>m\_\{j\}andba\>bbb\_\{a\}\>b\_\{b\}\. The difference between same\-order and reversed pairings is
\(1\+miba\)\(1\+mjbb\)−\(1\+mibb\)\(1\+mjba\)\\displaystyle\(1\+m\_\{i\}b\_\{a\}\)\(1\+m\_\{j\}b\_\{b\}\)\-\(1\+m\_\{i\}b\_\{b\}\)\(1\+m\_\{j\}b\_\{a\}\)=\(mi−mj\)\(ba−bb\)\>0\.\\displaystyle\\qquad=\(m\_\{i\}\-m\_\{j\}\)\(b\_\{a\}\-b\_\{b\}\)\>0\.\(96\)Sincelog\\logis increasing, replacing an in\-order pair by a reversed pair lowersΦorient\\Phi\_\{\\mathrm\{orient\}\}\. Repeated exchanges give the global reverse ordering; reversing the argument gives the global maximum\.
For completeness, at a commuting configuration the tangent space of the orthogonal orbit is spanned by independent pairwise generatorsEijE\_\{ij\}\. Because bothMMandBBare diagonal there,
H˙ij=Ξij\(mj−1−mi−1\),i≠j,\\dot\{H\}\_\{ij\}=\\Xi\_\{ij\}\\left\(m\_\{j\}^\{\-1\}\-m\_\{i\}^\{\-1\}\\right\),\\qquad i\\neq j,\(97\)soH˙\\dot\{H\}has only the corresponding off\-diagonal pair and bothTr\(H−1H¨\)\\operatorname\{Tr\}\(H^\{\-1\}\\ddot\{H\}\)andTr\(H−1H˙H−1H˙\)\\operatorname\{Tr\}\\\!\\left\(H^\{\-1\}\\dot\{H\}\\,H^\{\-1\}\\dot\{H\}\\right\)contain onlyΞij2\\Xi\_\{ij\}^\{2\}terms, with no cross\-pair contributions\. The second variation is therefore diagonal in this basis, with pairwise curvature
κij=−σξ2\(mi−mj\)\(bπ\(i\)−bπ\(j\)\)\(1\+mibπ\(i\)\)\(1\+mjbπ\(j\)\)\.\\kappa\_\{ij\}=\-\\sigma\_\{\\xi\}^\{2\}\\frac\{\(m\_\{i\}\-m\_\{j\}\)\(b\_\{\\pi\(i\)\}\-b\_\{\\pi\(j\)\}\)\}\{\(1\+m\_\{i\}b\_\{\\pi\(i\)\}\)\(1\+m\_\{j\}b\_\{\\pi\(j\)\}\)\}\.\(98\)Any nonextremal permutation contains both an in\-order and a reversed pair, and therefore has both positive\- and negative\-curvature tangent directions\. It is thus a saddle\. ∎
The theorem describes the orientation favored by the conditional fluctuation contribution: directions of weaker microstate curvature lower this contribution when paired with larger dynamical eigenvalues\.
### 4\.4Local rotational restoring contribution
We next characterize the generalized force supplied by the conditional free energy near a reverse\-paired configuration\. Consider
M∗=\(mi00mj\),B=\(bi00bj\),M\_\{\\ast\}=\\begin\{pmatrix\}m\_\{i\}&0\\\\ 0&m\_\{j\}\\end\{pmatrix\},\\qquad B=\\begin\{pmatrix\}b\_\{i\}&0\\\\ 0&b\_\{j\}\\end\{pmatrix\},\(99\)and rotateM∗M\_\{\\ast\}by
M\(θ\)=R\(θ\)M∗R\(θ\)∗\.M\(\\theta\)=R\(\\theta\)M\_\{\\ast\}R\(\\theta\)^\{\\ast\}\.\(100\)Writing
Dij\(θ\)≡det\[I\+M\(θ\)B\],D\_\{ij\}\(\\theta\)\\equiv\\det\[I\+M\(\\theta\)B\],\(101\)a direct calculation gives
Dij\(θ\)=Dij\(0\)−\(mi−mj\)\(bi−bj\)sin2θ,\\boxed\{D\_\{ij\}\(\\theta\)=D\_\{ij\}\(0\)\-\(m\_\{i\}\-m\_\{j\}\)\(b\_\{i\}\-b\_\{j\}\)\\sin^\{2\}\\theta,\}\(102\)where
Dij\(0\)=\(1\+mibi\)\(1\+mjbj\)\.D\_\{ij\}\(0\)=\(1\+m\_\{i\}b\_\{i\}\)\(1\+m\_\{j\}b\_\{j\}\)\.\(103\)Hence
Φij\(θ\)=Φij\(0\)\+12κijθ2\+O\(θ4\),\\Phi\_\{ij\}\(\\theta\)=\\Phi\_\{ij\}\(0\)\+\\frac\{1\}\{2\}\\kappa\_\{ij\}\\theta^\{2\}\+O\(\\theta^\{4\}\),\(104\)with
κij=−σξ2\(mi−mj\)\(bi−bj\)\(1\+mibi\)\(1\+mjbj\)\.\\boxed\{\\kappa\_\{ij\}=\-\\sigma\_\{\\xi\}^\{2\}\\frac\{\(m\_\{i\}\-m\_\{j\}\)\(b\_\{i\}\-b\_\{j\}\)\}\{\(1\+m\_\{i\}b\_\{i\}\)\(1\+m\_\{j\}b\_\{j\}\)\}\.\}\(105\)At the reverse\-paired minimum,mi\>mjm\_\{i\}\>m\_\{j\}impliesbi<bjb\_\{i\}<b\_\{j\}, soκij\>0\\kappa\_\{ij\}\>0\.
###### Corollary 4\.3\(Local rotational restoring contribution\)\.
For every nondegenerate pair at a reverse\-paired commuting configuration, the conditional orientational free energy has strictly positive quadratic curvature\. Defining the generalized force contributed by this sector as
fijcond≡−∂Φ∂θij,f\_\{ij\}^\{\\mathrm\{cond\}\}\\equiv\-\\frac\{\\partial\\Phi\}\{\\partial\\theta\_\{ij\}\},\(106\)one obtains
fijcond=−κijθij\+O\(θij3\),κij\>0\.\\boxed\{f\_\{ij\}^\{\\mathrm\{cond\}\}=\-\\kappa\_\{ij\}\\theta\_\{ij\}\+O\(\\theta\_\{ij\}^\{3\}\),\\qquad\\kappa\_\{ij\}\>0\.\}\(107\)Thus the conditional fluctuation sector contributes a local restoring thermodynamic force against rotational mismatch\.
Ifmi=mjm\_\{i\}=m\_\{j\}orbi=bjb\_\{i\}=b\_\{j\}, thenκij=0\\kappa\_\{ij\}=0and the corresponding rotation is a flat direction of this contribution\.
Corollary[4\.3](https://arxiv.org/html/2609.09589#S4.Thmtheorem3)is deliberately an operator\-space statement about one term in the effective force onMM\. The total slow dynamics may contain additional contributions, and whether a given parameterization can realize the corresponding operator rotation depends on the mapθ↦M\(θ\)\\theta\\mapsto M\(\\theta\)\.
### 4\.5Commutator interpretation
Define the active\-sector mismatch
𝒞\(M,B\)≡12‖\[M,B\]‖F2\.\\boxed\{\\mathcal\{C\}\(M,B\)\\equiv\\frac\{1\}\{2\}\\\|\[M,B\]\\\|\_\{F\}^\{2\}\.\}\(108\)Near a reverse\-paired commuting configuration,
𝒞\(M,B\)=∑i<j\(mi−mj\)2\(bi−bj\)2θij2\+O\(‖θ‖3\)\.\\mathcal\{C\}\(M,B\)=\\sum\_\{i<j\}\(m\_\{i\}\-m\_\{j\}\)^\{2\}\(b\_\{i\}\-b\_\{j\}\)^\{2\}\\theta\_\{ij\}^\{2\}\+O\(\\\|\\theta\\\|^\{3\}\)\.\(109\)Combining this with Corollary[4\.3](https://arxiv.org/html/2609.09589#S4.Thmtheorem3), every nondegenerate pair satisfies locally
⟨−∇rotΦ,∇rot𝒞⟩<0\.\\boxed\{\\left\\langle\-\\nabla\_\{\\mathrm\{rot\}\}\\Phi,\\nabla\_\{\\mathrm\{rot\}\}\\mathcal\{C\}\\right\\rangle<0\.\}\(110\)Thus descent of the conditional free\-energy contribution locally reduces operator noncommutativity\. This identifies the direction of the thermodynamic bias supplied by the fluctuation sector without assuming that the full parameter dynamics follows this descent exactly\.
## 5ReLU Cell Geometry and Structural Smoothness Preference
Section[4](https://arxiv.org/html/2609.09589#S4)showed that the conditional fluctuation free\-energy contribution is minimized, on a fixed\-spectrum orbit, when large eigenvalues ofMMare paired with small eigenvalues of the compressed microstate\-curvature operatorBaB\_\{a\}\. We now examine the structure of this statistical geometry for ReLU networks\.
### 5\.1ReLU cell geometry and the structural field
Letc\(x\)c\(x\)denote a function represented by a ReLU network\. The input space is partitioned into activation cells within which the activation pattern is fixed\. Restricted to any such cell,ccis affine, and therefore
D2c\(x\)=0inside each activation cell\.\\boxed\{D^\{2\}c\(x\)=0\}\\qquad\\text\{inside each activation cell\.\}\(111\)The second\-order structure is consequently localized on the boundaries between adjacent cells\.
Consider a codimension\-one facetFFseparating two neighboring activation cells\. Continuity implies that tangential derivatives agree across the facet, whereas the normal component of the gradient may jump\. Thus
\[∇c\]F=αFnF,\\boxed\{\[\\nabla c\]\_\{F\}=\\alpha\_\{F\}n\_\{F\},\}\(112\)wherenFn\_\{F\}is a unit normal andαF\\alpha\_\{F\}is the jump amplitude\. Accordingly, the Hessian is naturally understood distributionally:
D2c=∑FαFnF⊗nFδF\.\\boxed\{D^\{2\}c=\\sum\_\{F\}\\alpha\_\{F\}\\,n\_\{F\}\\otimes n\_\{F\}\\,\\delta\_\{F\}\.\}\(113\)Thus the second\-order content of a ReLU function is carried by its activation boundaries and associated gradient jumps\.
We introduce a fixed coarse\-graining operator𝒞\\mathcal\{C\}and define the structural field
h=Lc,L≡𝒞D2\.\\boxed\{h=Lc,\\qquad L\\equiv\\mathcal\{C\}D^\{2\}\.\}\(114\)The fieldhhsummarizes coarse\-grained departures from local affinity\. We choose the retained function sectorℋR\\mathcal\{H\}\_\{R\}introduced in Sec\.[3\.3](https://arxiv.org/html/2609.09589#S3.SS3)so that these coarse\-grained structural coordinates are well defined on the retained directions\. The mapLLsends this retained function sector into a structural\-field spaceℋh\\mathcal\{H\}\_\{h\}; the finite\-dimensional operator entering Sec\.[4](https://arxiv.org/html/2609.09589#S4)is then obtained by compression fromℋR\\mathcal\{H\}\_\{R\}toℋa\\mathcal\{H\}\_\{a\}\.
### 5\.2Statistical boundary conditions for ReLU microstates
Letcrc\_\{r\}denote the current coarse\-grained function configuration andhr=Lcrh\_\{r\}=Lc\_\{r\}\. We impose three statistical boundary conditions on the local microstate ensemble\.
#### R1\. Structural sufficiency\.
Within the retained coarse\-grained sector, the relevant local variation of the microstate entropy is determined byh=Lch=Lc:
Smicro\[c\]=Sh\[Lc\]\\boxed\{S\_\{\\mathrm\{micro\}\}\[c\]=S\_\{h\}\[Lc\]\}\(115\)in a neighborhood ofcrc\_\{r\}\.
#### R2\. Local entropy regularity\.
We assume thatShS\_\{h\}is twice Fréchet differentiable nearhrh\_\{r\}and define
𝒦\(r\)≡−Dh2Sh\[hr\]\.\\boxed\{\\mathcal\{K\}\(r\)\\equiv\-D\_\{h\}^\{2\}S\_\{h\}\[h\_\{r\}\]\.\}\(116\)The first variation need not vanish; it affects the local center, whereas the second variation controls the fluctuation geometry\. This regularity assumption is imposed on the entropy of the coarse\-grained structural field, not on the microscopic facet configuration itself\. The role of𝒞\\mathcal\{C\}is to suppress facet\-scale singular structure before this local statistical description is applied; R2 does not claim that the raw activation\-boundary ensemble is differentiable\.
#### R3\. Mild stable structural statistics\.
On the structural subspace relevant to the local fluctuations, we assume
0<k−I⪯𝒦\(r\)⪯k\+I,\\boxed\{0<k\_\{\-\}I\\preceq\\mathcal\{K\}\(r\)\\preceq k\_\{\+\}I,\}\(117\)for finite0<k−≤k\+<∞0<k\_\{\-\}\\leq k\_\{\+\}<\\infty\.
Condition R3 is the stability input of the ReLU specialization: it assumes local concavity of the structural microstate entropy on the retained sector\. ReLU cell geometry by itself does not imply this statistical property\. The role of R1–R3 is instead to identify a regular ReLU microstate class in which the stable sector used in Sec\.[4](https://arxiv.org/html/2609.09589#S4)is realized\.
### 5\.3Pullback of the microstate curvature
On the retained function sector, writecR=PRcc\_\{R\}=P\_\{R\}c,yR=PRyy\_\{R\}=P\_\{R\}y, andeR=cR−yRe\_\{R\}=c\_\{R\}\-y\_\{R\}\. Up to an additive constant, the microstate entropy introduced in Sec\.[3\.3](https://arxiv.org/html/2609.09589#S3.SS3)is
Smicro\[cR\]=logΩz\(eR\),eR=cR−yR\.S\_\{\\mathrm\{micro\}\}\[c\_\{R\}\]=\\log\\Omega\_\{z\}\(e\_\{R\}\),\\qquad e\_\{R\}=c\_\{R\}\-y\_\{R\}\.\(118\)SinceyRy\_\{R\}is fixed, variations incRc\_\{R\}andeRe\_\{R\}coincide\. Under R1,
DcR2Smicro\[cr\]\(u,v\)=Dh2Sh\[hr\]\(Lu,Lv\),D\_\{c\_\{R\}\}^\{2\}S\_\{\\mathrm\{micro\}\}\[c\_\{r\}\]\(u,v\)=D\_\{h\}^\{2\}S\_\{h\}\[h\_\{r\}\]\(Lu,Lv\),\(119\)and therefore
−DcR2Smicro\[cr\]\(u,v\)=⟨Lu,𝒦\(r\)Lv⟩h\.\-D\_\{c\_\{R\}\}^\{2\}S\_\{\\mathrm\{micro\}\}\[c\_\{r\}\]\(u,v\)=\\langle Lu,\\mathcal\{K\}\(r\)Lv\\rangle\_\{h\}\.\(120\)LetL∗L^\{\\ast\}be the adjoint defined by
⟨Lu,h⟩h=⟨u,L∗h⟩p\.\\langle Lu,h\\rangle\_\{h\}=\\langle u,L^\{\\ast\}h\\rangle\_\{p\}\.\(121\)Then
−DcR2Smicro\[cr\]\(u,v\)=⟨u,L∗𝒦\(r\)Lv⟩p\.\-D\_\{c\_\{R\}\}^\{2\}S\_\{\\mathrm\{micro\}\}\[c\_\{r\}\]\(u,v\)=\\langle u,L^\{\\ast\}\\mathcal\{K\}\(r\)Lv\\rangle\_\{p\}\.\(122\)
###### Theorem 5\.1\(ReLU microstate curvature\)\.
Under R1–R3, the retained\-sector microstate\-curvature form is
B\(r\)=σξ2L∗𝒦\(r\)L\.\\boxed\{B\(r\)=\\sigma\_\{\\xi\}^\{2\}L^\{\\ast\}\\mathcal\{K\}\(r\)L\.\}\(123\)For every retained perturbationu∈ℋRu\\in\\mathcal\{H\}\_\{R\}for whichLLis defined,
σξ2k−‖Lu‖h2≤⟨u,B\(r\)u⟩p≤σξ2k\+‖Lu‖h2\.\\boxed\{\\sigma\_\{\\xi\}^\{2\}k\_\{\-\}\\\|Lu\\\|\_\{h\}^\{2\}\\leq\\langle u,B\(r\)u\\rangle\_\{p\}\\leq\\sigma\_\{\\xi\}^\{2\}k\_\{\+\}\\\|Lu\\\|\_\{h\}^\{2\}\.\}\(124\)HenceB\(r\)⪰0B\(r\)\\succeq 0as a quadratic form onℋR\\mathcal\{H\}\_\{R\}\. The operator entering Sec\.[4](https://arxiv.org/html/2609.09589#S4)is the compression
Ba\(r\)=PaB\(r\)Pa\|ℋa,\\boxed\{B\_\{a\}\(r\)=P\_\{a\}B\(r\)P\_\{a\}\\big\|\_\{\\mathcal\{H\}\_\{a\}\},\}\(125\)which is therefore positive semidefinite onℋa\\mathcal\{H\}\_\{a\}\. Under the strict coercivity in R3,kerB=kerL\\ker B=\\ker Lon the retained function sector\.
###### Proof\.
Equation equation[123](https://arxiv.org/html/2609.09589#S5.E123)follows from Eq\. equation[122](https://arxiv.org/html/2609.09589#S5.E122)and the retained\-sector translationeR=cR−yRe\_\{R\}=c\_\{R\}\-y\_\{R\}\. For anyuu,
⟨u,B\(r\)u⟩p=σξ2⟨Lu,𝒦\(r\)Lu⟩h\.\\langle u,B\(r\)u\\rangle\_\{p\}=\\sigma\_\{\\xi\}^\{2\}\\langle Lu,\\mathcal\{K\}\(r\)Lu\\rangle\_\{h\}\.\(126\)Applying R3 yields Eq\. equation[124](https://arxiv.org/html/2609.09589#S5.E124)\. Compression preserves positive semidefiniteness\. Finally, strict coercivity implies⟨u,Bu⟩p=0\\langle u,Bu\\rangle\_\{p\}=0if and only ifLu=0Lu=0\. ∎
The theorem should therefore be read as a pullback result: R3 supplies the stable structural entropy metric, while the ReLU structural mapLLdetermines how that metric is represented in function space\.
### 5\.4Microstate curvature and structural smoothness
The spectral matching in Sec\.[4](https://arxiv.org/html/2609.09589#S4)involves the finite\-dimensional compressed operatorBa\(r\)B\_\{a\}\(r\)\. Let
Ba\(r\)ϕi=bi\(r\)ϕi,ϕi∈ℋa,‖ϕi‖p=1\.B\_\{a\}\(r\)\\phi\_\{i\}=b\_\{i\}\(r\)\\phi\_\{i\},\\qquad\\phi\_\{i\}\\in\\mathcal\{H\}\_\{a\},\\qquad\\\|\\phi\_\{i\}\\\|\_\{p\}=1\.\(127\)BecausePaϕi=ϕiP\_\{a\}\\phi\_\{i\}=\\phi\_\{i\},
bi\(r\)=⟨ϕi,B\(r\)ϕi⟩p\.b\_\{i\}\(r\)=\\langle\\phi\_\{i\},B\(r\)\\phi\_\{i\}\\rangle\_\{p\}\.\(128\)Applying Eq\. equation[124](https://arxiv.org/html/2609.09589#S5.E124),
σξ2k−‖Lϕi‖h2≤bi\(r\)≤σξ2k\+‖Lϕi‖h2\.\\boxed\{\\sigma\_\{\\xi\}^\{2\}k\_\{\-\}\\\|L\\phi\_\{i\}\\\|\_\{h\}^\{2\}\\leq b\_\{i\}\(r\)\\leq\\sigma\_\{\\xi\}^\{2\}k\_\{\+\}\\\|L\\phi\_\{i\}\\\|\_\{h\}^\{2\}\.\}\(129\)Hence
bi\(r\)≍‖Lϕi‖h2\.\\boxed\{b\_\{i\}\(r\)\\asymp\\\|L\\phi\_\{i\}\\\|\_\{h\}^\{2\}\.\}\(130\)
For ReLU functions,L=𝒞D2L=\\mathcal\{C\}D^\{2\}, so‖Lϕi‖h\\\|L\\phi\_\{i\}\\\|\_\{h\}measures coarse\-grained second\-order content, including activation\-boundary and gradient\-jump structure\. We call directions with small‖Lϕ‖h\\\|L\\phi\\\|\_\{h\}*structurally smooth*; throughout this paper, “smooth” in the ReLU specialization refers to low coarse\-grained structural curvature in this sense, rather than to a fixed Fourier frequency notion\.
###### Corollary 5\.2\(Smoothness interpretation\)\.
Within the active sector and under R1–R3, the spectrum ofBaB\_\{a\}controls coarse\-grained structural curvature up to the bounded condition number
κ𝒦≡k\+k−\.\\kappa\_\{\\mathcal\{K\}\}\\equiv\\frac\{k\_\{\+\}\}\{k\_\{\-\}\}\.\(131\)In particular,
κ𝒦−1‖Lϕi‖h2‖Lϕj‖h2≤bibj≤κ𝒦‖Lϕi‖h2‖Lϕj‖h2\.\\kappa\_\{\\mathcal\{K\}\}^\{\-1\}\\frac\{\\\|L\\phi\_\{i\}\\\|\_\{h\}^\{2\}\}\{\\\|L\\phi\_\{j\}\\\|\_\{h\}^\{2\}\}\\leq\\frac\{b\_\{i\}\}\{b\_\{j\}\}\\leq\\kappa\_\{\\mathcal\{K\}\}\\frac\{\\\|L\\phi\_\{i\}\\\|\_\{h\}^\{2\}\}\{\\\|L\\phi\_\{j\}\\\|\_\{h\}^\{2\}\}\.\(132\)Thus low\-BaB\_\{a\}spectral sectors correspond, within the finite distortion set byκ𝒦\\kappa\_\{\\mathcal\{K\}\}, to sectors of low coarse\-grained structural curvature\. Exact pairwise ordering of‖Lϕi‖h\\\|L\\phi\_\{i\}\\\|\_\{h\}is not asserted when the structural metric is strongly anisotropic\. This limitation affects only the translation from theBaB\_\{a\}spectrum to structural smoothness; the reverse\-pairing statement of Theorem[4\.2](https://arxiv.org/html/2609.09589#S4.Thmtheorem2), which is formulated directly in terms of theBaB\_\{a\}eigenvalues, remains exact\.
### 5\.5Data\-adaptive structural smoothness preference
The smoothness spectrum is defined in the data\-weighted function space
ℋ=L2\(𝒳,p\),\\mathcal\{H\}=L^\{2\}\(\\mathcal\{X\},p\),\(133\)with
⟨f,g⟩p=∫𝒳f\(x\)g\(x\)p\(x\)𝑑x\.\\langle f,g\\rangle\_\{p\}=\\int\_\{\\mathcal\{X\}\}f\(x\)g\(x\)p\(x\)\\,dx\.\(134\)Consequently, the adjointL∗L^\{\\ast\}, orthogonality of modes, and the active\-sector compression all depend on the geometry induced byp\(x\)p\(x\)\. Making this dependence explicit,
Bp,a\(r\)=Pa\[σξ2Lp∗𝒦p\(r\)L\]Pa\|ℋa\.B\_\{p,a\}\(r\)=P\_\{a\}\\left\[\\sigma\_\{\\xi\}^\{2\}L\_\{p\}^\{\\ast\}\\mathcal\{K\}\_\{p\}\(r\)L\\right\]P\_\{a\}\\big\|\_\{\\mathcal\{H\}\_\{a\}\}\.\(135\)Thus the low\-curvature sectors identified throughBp,aB\_\{p,a\}are intrinsic to the data\-weighted functional geometry\.
Combining this interpretation with Theorem[4\.2](https://arxiv.org/html/2609.09589#S4.Thmtheorem2), the conditional fluctuation free\-energy contribution pairs large dynamical eigenvalues with the low\-BaB\_\{a\}sector, which corresponds up to the bounded distortionκ𝒦\\kappa\_\{\\mathcal\{K\}\}to low\-curvature data\-adaptive function\-space directions:
mlarge⟷bsmall⟷low structural\-curvature sector\.\\boxed\{m\_\{\\mathrm\{large\}\}\\longleftrightarrow b\_\{\\mathrm\{small\}\}\\longleftrightarrow\\text\{low structural\-curvature sector\}\.\}\(136\)
###### Corollary 5\.3\(Data\-adaptive structural smoothness preference\)\.
Under R1–R3 and within the stable active sector, the conditional orientational free\-energy contribution favors pairing larger eigenvalues ofMMwith the low\-BaB\_\{a\}data\-adaptive sector\. Through Eq\. equation[129](https://arxiv.org/html/2609.09589#S5.E129), this is a bias toward directions of lower coarse\-grained structural curvature up to the finite anisotropy factorκ𝒦\\kappa\_\{\\mathcal\{K\}\}\. Sincemim\_\{i\}sets the gradient\-flow relaxation rate along the corresponding eigendirection, the fluctuation contribution therefore favors faster relaxation in this low\-curvature sector\.
The result is a statement about the geometry preferred by the conditional fluctuation contribution\. Whether the full training dynamics realizes this preference depends on the remaining slow operator dynamics and on how parameter motion can realize changes inMM\.
## 6Discussion and Future Work
The function\-space thermodynamic picture developed here is broadly consistent with several empirical regularities reported in neural networks\. Classical observations of spectral bias indicate that smoother or lower\-frequency components are often learned earlier[Rahaman et al\. \(2019\)](https://arxiv.org/html/2609.09589#bib.bib13)\. More recently, diffusion denoisers have been found to develop geometry\-adaptive harmonic representations: their learned input–output Jacobians organize into data\-dependent eigendirections whose ordering is closely related to smoothness[Kadkhodaie et al\. \(2024\)](https://arxiv.org/html/2609.09589#bib.bib20)\. The operator measured in that work is not the learning operatorM=JθJθ∗M=J\_\{\\theta\}J\_\{\\theta\}^\{\\ast\}considered here, so this is not a direct test of our theory\. Nevertheless, the observed organization is qualitatively consistent with the geometry favored by our conditional fluctuation contribution,
mlarge⟷bsmall⟷low structural\-curvature data\-adaptive sector\.m\_\{\\mathrm\{large\}\}\\longleftrightarrow b\_\{\\mathrm\{small\}\}\\longleftrightarrow\\text\{low structural\-curvature data\-adaptive sector\}\.\(137\)
The same macroscopic language can also accommodate qualitatively different learning regimes\. IfMMremains effectively fixed, the description reduces to a kernel\-like regime\. If additional slow dynamics allowMMto respond to the conditional thermodynamic force derived here, the same framework supplies a bias toward regular operator matching\. Conversely, apparently sharp macroscopic behavior need not have a unique origin: it may reflect competition between distinct macroscopic states, or it may arise from a broad hierarchy of relaxation times even when the underlying dynamics remain continuous\. The latter possibility is conceptually related to quantized models of neural scaling, in which smooth aggregate scaling can coexist with the sudden appearance of individual capabilities[Michaud et al\. \(2023\)](https://arxiv.org/html/2609.09589#bib.bib21)\.
Grokking provides a suggestive example of the former possibility\. Previous work has connected delayed generalization to structured representations[Liu et al\. \(2022\)](https://arxiv.org/html/2609.09589#bib.bib22), escape from an early kernel\-like regime in modular addition[Mohamadi et al\. \(2024\)](https://arxiv.org/html/2609.09589#bib.bib23), and first\-order phase transitions between representation phases in two\-layer teacher–student models[Rubin et al\. \(2024\)](https://arxiv.org/html/2609.09589#bib.bib24)\. Our framework highlights an endogenous route by which the relative statistical preference of macroscopic states can change during training\. Because the density of statesΩ\(e\)\\Omega\(e\)and its local curvatureB\(r\)B\(r\)depend on the current error state, the statistical geometry sampled by the theory changes asr=r\(t\)r=r\(t\)evolves\. If two macroscopic branches coexist, their*total*conditional free energies, denoted schematically by𝒢A\(r\)\\mathcal\{G\}\_\{A\}\(r\)and𝒢G\(r\)\\mathcal\{G\}\_\{G\}\(r\), may therefore cross\. These symbols refer to full branch free energies, including macrostate, spectral, rank\-dependent, and fluctuation contributions; such a term\-by\-term global branch theory is not constructed in the present work\. Schematically,
𝒢A\(r∗\)=𝒢G\(r∗\),𝒢A\(r\)−𝒢G\(r\)changes sign acrossr∗\.\\mathcal\{G\}\_\{A\}\(r\_\{\\ast\}\)=\\mathcal\{G\}\_\{G\}\(r\_\{\\ast\}\),\\qquad\\mathcal\{G\}\_\{A\}\(r\)\-\\mathcal\{G\}\_\{G\}\(r\)\\ \\text\{changes sign across \}r\_\{\\ast\}\.\(138\)In this picture, the training state itself can act as an endogenous control variable\. Establishing the relevant competing branches, barriers, and transition dynamics lies beyond the local conditional theory developed here and is left for future work\.
Several limitations define natural extensions\. First, the present construction assumes an isotropic stochastic source in function space\. More generally, the noise may possess its own covariance geometryQQ, leading schematically to
D∝MQM\.D\\propto MQM\.\(139\)The resulting conditional statistical mechanics need not remain reversible, and the operator preference may differ from the one derived here; nonequilibrium extensions are therefore an important direction\. Second,Φfluc\(M,B\)\\Phi\_\{\\mathrm\{fluc\}\}\(M;B\)is the free\-energy contribution of the local conditional fluctuation sector, not a derivation of the complete slow dynamics ofMM\. The macrostate term12⟨ra,M−1ra⟩p\\frac\{1\}\{2\}\\langle r\_\{a\},M^\{\-1\}r\_\{a\}\\rangle\_\{p\}already supplies a distinct error\-directed orientational preference, and further slow terms may also be present\. Interpreting the fluctuation gradient as an actual component of training dynamics requires that these other contributions do not cancel or overwhelm it\. Third, our conditional ensemble does not require a time\-scale separation, but interpreting it as a quasi\-static distribution realized along a training trajectory would require additional local\-equilibration assumptions\. Finally, all operator matching results are formulated on a fixed finite\-dimensional active sector, and the ability of parameter dynamics to realize the corresponding operator rotations remains architecture dependent\.
## 7Conclusion
We developed a statistical\-mechanical description of neural\-network learning directly in function space\. Parameter configurations provide the microscopic realizations, while functions and the operators governing their evolution provide macroscopic variables\. This separation makes it possible to distinguish the dynamical geometry of learning from statistical constraints induced by the underlying parameterization\.
For mean\-squared loss, the exact error dynamics determines a bare conditional dynamical weight\. Combining this weight with parameter\-space microstate multiplicity produces a conditional function\-space ensemble whose local statistical curvature is described byBB\. On the finite\-dimensional active sector, integrating over local error fluctuations yields a conditional free\-energy contribution
Φfluc\(M,B\)=σξ22logdet\(M−1\+B\)\+const\.\\Phi\_\{\\mathrm\{fluc\}\}\(M;B\)=\\frac\{\\sigma\_\{\\xi\}^\{2\}\}\{2\}\\log\\det\(M^\{\-1\}\+B\)\+\\mathrm\{const\}\.\(140\)At fixed spectrum, this contribution is stationary when\[M,B\]=0\[M,B\]=0, is minimized by the reverse pairing
mlarge⟷bsmall,m\_\{\\mathrm\{large\}\}\\longleftrightarrow b\_\{\\mathrm\{small\}\},\(141\)and supplies a local restoring force contribution against rotational mismatch\.
For ReLU\-type function spaces, stable structural statistics satisfying R1–R3 give
B=σξ2L∗𝒦L,B=\\sigma\_\{\\xi\}^\{2\}L^\{\\ast\}\\mathcal\{K\}L,\(142\)whose active\-sector spectrum measures coarse\-grained structural curvature up to the bounded anisotropy of the structural metric\. The conditional thermodynamic contribution therefore favors pairing faster relaxation with the low\-curvature, data\-adaptive sector\.
These results suggest that function space provides a natural macroscopic level for the statistical mechanics of learning\. In this organization of the theory, a concrete neural network is a microscopic realization rather than the starting point: the function\-space organizing principle is formulated first, while architecture and parameterization determine which operators and microstate geometries realize it\. The present theory isolates one thermodynamic contribution within this framework, providing a basis for studying more general slow dynamics and nonequilibrium extensions\.
## References
- Bahriet al\.\(2020\)Y\. Bahri, J\. Kadmon, J\. Pennington, S\. S\. Schoenholz, J\. Sohl\-Dickstein, and S\. GanguliStatistical mechanics of deep learning\.Annual Review of Condensed Matter Physics11,pp\. 501–528\.External Links:[Document](https://dx.doi.org/10.1146/annurev-conmatphys-031119-050745)Cited by:[§1](https://arxiv.org/html/2609.09589#S1.p3.1),[§2](https://arxiv.org/html/2609.09589#S2.SS0.SSS0.Px2.p1.1)\.
- Balestriero and Baraniuk \(2018\)R\. Balestriero and R\. BaraniukA spline theory of deep learning\.InProceedings of the 35th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.80,pp\. 374–383\.External Links:[Link](https://proceedings.mlr.press/v80/balestriero18b.html)Cited by:[§2](https://arxiv.org/html/2609.09589#S2.SS0.SSS0.Px4.p1.1)\.
- Chaudhari and Soatto \(2018\)P\. Chaudhari and S\. SoattoStochastic gradient descent performs variational inference, converges to limit cycles for deep networks\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=HyWrIgW0W)Cited by:[§1](https://arxiv.org/html/2609.09589#S1.p3.1),[§2](https://arxiv.org/html/2609.09589#S2.SS0.SSS0.Px2.p1.1)\.
- Chizatet al\.\(2019\)L\. Chizat, E\. Oyallon, and F\. BachOn lazy training in differentiable programming\.InAdvances in Neural Information Processing Systems,Vol\.32,pp\. 2933–2943\.External Links:[Link](https://proceedings.neurips.cc/paper/2019/hash/ae614c557843b1df326cb29c57225459-Abstract.html)Cited by:[§1](https://arxiv.org/html/2609.09589#S1.p2.1),[§2](https://arxiv.org/html/2609.09589#S2.SS0.SSS0.Px1.p1.1)\.
- Cohenet al\.\(2021\)J\. M\. Cohen, S\. Kaur, Y\. Li, J\. Z\. Kolter, and A\. TalwalkarGradient descent on neural networks typically occurs at the edge of stability\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=jh-rTtvkGeM)Cited by:[§1](https://arxiv.org/html/2609.09589#S1.p1.1)\.
- Hestnesset al\.\(2017\)J\. Hestness, S\. Narang, N\. Ardalani, G\. F\. Diamos, H\. Jun, H\. Kianinejad, Md\. M\. A\. Patwary, Y\. Yang, and Y\. ZhouDeep learning scaling is predictable, empirically\.External Links:1712\.00409,[Document](https://dx.doi.org/10.48550/arXiv.1712.00409),[Link](https://arxiv.org/abs/1712.00409)Cited by:[§1](https://arxiv.org/html/2609.09589#S1.p1.1)\.
- Jacotet al\.\(2018\)A\. Jacot, F\. Gabriel, and C\. HonglerNeural tangent kernel: convergence and generalization in neural networks\.InAdvances in Neural Information Processing Systems,Vol\.31,pp\. 8571–8580\.External Links:[Link](https://proceedings.neurips.cc/paper/2018/hash/5a4be1fa34e62bb8a6ec6b91d2462f5a-Abstract.html)Cited by:[§1](https://arxiv.org/html/2609.09589#S1.p2.1),[§2](https://arxiv.org/html/2609.09589#S2.SS0.SSS0.Px1.p1.1)\.
- Jordanet al\.\(1998\)R\. Jordan, D\. Kinderlehrer, and F\. OttoThe variational formulation of the fokker–planck equation\.SIAM Journal on Mathematical Analysis29\(1\),pp\. 1–17\.External Links:[Document](https://dx.doi.org/10.1137/S0036141096303359)Cited by:[§2](https://arxiv.org/html/2609.09589#S2.SS0.SSS0.Px2.p1.1)\.
- Kadkhodaieet al\.\(2024\)Z\. Kadkhodaie, F\. Guth, E\. P\. Simoncelli, and S\. MallatGeneralization in diffusion models arises from geometry\-adaptive harmonic representations\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=ANvmVS2Yr0)Cited by:[§1](https://arxiv.org/html/2609.09589#S1.p1.1),[§6](https://arxiv.org/html/2609.09589#S6.p1.1)\.
- Kaplanet al\.\(2020\)J\. Kaplan, S\. McCandlish, T\. Henighan, T\. B\. Brown, B\. Chess, R\. Child, S\. Gray, A\. Radford, J\. Wu, and D\. AmodeiScaling laws for neural language models\.External Links:2001\.08361,[Document](https://dx.doi.org/10.48550/arXiv.2001.08361),[Link](https://arxiv.org/abs/2001.08361)Cited by:[§1](https://arxiv.org/html/2609.09589#S1.p1.1)\.
- Lauditiet al\.\(2025\)C\. Lauditi, B\. Bordelon, and C\. PehlevanAdaptive kernel predictors from feature\-learning infinite limits of neural networks\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 32617–32648\.External Links:[Link](https://proceedings.mlr.press/v267/lauditi25a.html)Cited by:[§1](https://arxiv.org/html/2609.09589#S1.p2.1),[§2](https://arxiv.org/html/2609.09589#S2.SS0.SSS0.Px1.p1.1)\.
- Lauditiet al\.\(2026\)C\. Lauditi, C\. Pehlevan, and B\. BordelonSpectral dynamics in deep networks: feature learning, outlier escape, and learning rate transfer\.External Links:2605\.07870,[Document](https://dx.doi.org/10.48550/arXiv.2605.07870),[Link](https://arxiv.org/abs/2605.07870)Cited by:[§1](https://arxiv.org/html/2609.09589#S1.p2.1),[§2](https://arxiv.org/html/2609.09589#S2.SS0.SSS0.Px1.p1.1)\.
- Liuet al\.\(2022\)Z\. Liu, O\. Kitouni, N\. S\. Nolte, E\. J\. Michaud, M\. Tegmark, and M\. WilliamsTowards understanding grokking: an effective theory of representation learning\.InAdvances in Neural Information Processing Systems,Vol\.35\.External Links:[Document](https://dx.doi.org/10.52202/068431-2511),[Link](https://proceedings.neurips.cc/paper_files/paper/2022/hash/dfc310e81992d2e4cedc09ac47eff13e-Abstract-Conference.html)Cited by:[§6](https://arxiv.org/html/2609.09589#S6.p3.1)\.
- Mandtet al\.\(2016\)S\. Mandt, M\. D\. Hoffman, and D\. M\. BleiA variational analysis of stochastic gradient algorithms\.InProceedings of the 33rd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.48,pp\. 354–363\.External Links:[Link](https://proceedings.mlr.press/v48/mandt16.html)Cited by:[§1](https://arxiv.org/html/2609.09589#S1.p3.1),[§2](https://arxiv.org/html/2609.09589#S2.SS0.SSS0.Px2.p1.1)\.
- Mandtet al\.\(2017\)S\. Mandt, M\. D\. Hoffman, and D\. M\. BleiStochastic gradient descent as approximate bayesian inference\.Journal of Machine Learning Research18\(134\),pp\. 1–35\.External Links:[Link](https://jmlr.org/papers/v18/17-214.html)Cited by:[§1](https://arxiv.org/html/2609.09589#S1.p3.1),[§2](https://arxiv.org/html/2609.09589#S2.SS0.SSS0.Px2.p1.1)\.
- Michaudet al\.\(2023\)E\. J\. Michaud, Z\. Liu, U\. Girit, and M\. TegmarkThe quantization model of neural scaling\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 28699–28722\.External Links:[Document](https://dx.doi.org/10.52202/075280-1248),[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/5b6346a05a537d4cdb2f50323452a9fe-Abstract-Conference.html)Cited by:[§6](https://arxiv.org/html/2609.09589#S6.p2.1)\.
- Mohamadiet al\.\(2024\)M\. A\. Mohamadi, Z\. Li, L\. Wu, and D\. J\. SutherlandWhy do you grok? A theoretical analysis on grokking modular addition\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 35934–35967\.External Links:[Link](https://proceedings.mlr.press/v235/mohamadi24a.html)Cited by:[§6](https://arxiv.org/html/2609.09589#S6.p3.1)\.
- Pavliotis \(2014\)G\. A\. PavliotisStochastic processes and applications: diffusion processes, the fokker–planck and langevin equations\.Texts in Applied Mathematics, Vol\.60,Springer,New York\.External Links:[Document](https://dx.doi.org/10.1007/978-1-4939-1323-7)Cited by:[§2](https://arxiv.org/html/2609.09589#S2.SS0.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2609.09589#S3.SS2.p2.2)\.
- Rahamanet al\.\(2019\)N\. Rahaman, A\. Baratin, D\. Arpit, F\. Draxler, M\. Lin, F\. Hamprecht, Y\. Bengio, and A\. CourvilleOn the spectral bias of neural networks\.InProceedings of the 36th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.97,pp\. 5301–5310\.External Links:[Link](https://proceedings.mlr.press/v97/rahaman19a.html)Cited by:[§1](https://arxiv.org/html/2609.09589#S1.p1.1),[§2](https://arxiv.org/html/2609.09589#S2.SS0.SSS0.Px3.p1.1),[§6](https://arxiv.org/html/2609.09589#S6.p1.1)\.
- Risken \(1989\)H\. RiskenThe fokker–planck equation: methods of solution and applications\.2 edition,Springer Series in Synergetics, Vol\.18,Springer\-Verlag,Berlin\.External Links:ISBN 9780387504988Cited by:[§2](https://arxiv.org/html/2609.09589#S2.SS0.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2609.09589#S3.SS2.p2.2)\.
- Rubinet al\.\(2024\)N\. Rubin, I\. Seroussi, and Z\. RingelGrokking as a first order phase transition in two layer networks\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/hash/682f87a8c306098ec8be29019bd76aa4-Abstract-Conference.html)Cited by:[§6](https://arxiv.org/html/2609.09589#S6.p3.1)\.
- Unser \(2019\)M\. UnserA representer theorem for deep neural networks\.Journal of Machine Learning Research20\(110\),pp\. 1–30\.External Links:[Link](https://jmlr.org/papers/v20/18-418.html)Cited by:[§2](https://arxiv.org/html/2609.09589#S2.SS0.SSS0.Px4.p1.1)\.
- Woodworthet al\.\(2020\)B\. Woodworth, S\. Gunasekar, J\. D\. Lee, E\. Moroshko, P\. Savarese, I\. Golan, D\. Soudry, and N\. SrebroKernel and rich regimes in overparametrized models\.InProceedings of the Thirty Third Conference on Learning Theory,Proceedings of Machine Learning Research, Vol\.125,pp\. 3635–3673\.External Links:[Link](https://proceedings.mlr.press/v125/woodworth20a.html)Cited by:[§1](https://arxiv.org/html/2609.09589#S1.p2.1),[§2](https://arxiv.org/html/2609.09589#S2.SS0.SSS0.Px1.p1.1)\.
- Yanget al\.\(2021\)G\. Yang, E\. J\. Hu, I\. Babuschkin, S\. Sidor, X\. Liu, D\. Farhi, N\. Ryder, J\. Pachocki, W\. Chen, and J\. GaoTensor programs v: tuning large neural networks via zero\-shot hyperparameter transfer\.InAdvances in Neural Information Processing Systems,Vol\.34\.External Links:[Link](https://proceedings.neurips.cc/paper/2021/hash/8df7c2e3c3c3be098ef7b382bd2c37ba-Abstract.html)Cited by:[§1](https://arxiv.org/html/2609.09589#S1.p1.1)\.
## Appendix AResidual\-Preserving Rotations of the Local Conditional Free Energy
The main text isolates the fluctuation\-induced contributionΦfluc\\Phi\_\{\\mathrm\{fluc\}\}to the local conditional free energy\. Here we record a restricted setting in which this contribution is also the complete orientational variation of the quadratic local free energy within the conditional construction of Sec\.[3\.4](https://arxiv.org/html/2609.09589#S3.SS4)\.
Letra=Par∈ℋar\_\{a\}=P\_\{a\}r\\in\\mathcal\{H\}\_\{a\}be the active component of the conditioned macrostate\. For notational simplicity within this appendix, writer≡rar\\equiv r\_\{a\}and assumer≠0r\\neq 0\. Define its stabilizer subgroup
Gr≡\{O∈SO\(ℋa\):Or=r\}\.G\_\{r\}\\equiv\\left\\\{O\\in SO\(\\mathcal\{H\}\_\{a\}\):Or=r\\right\\\}\.\(143\)Its infinitesimal generators satisfy
Ξ∗=−Ξ,Ξr=0\.\\Xi^\{\\ast\}=\-\\Xi,\\qquad\\Xi r=0\.\(144\)Consider the fixed\-spectrum orbit
M\(O\)=OMO∗,O∈Gr\.M\(O\)=OMO^\{\\ast\},\\qquad O\\in G\_\{r\}\.\(145\)
###### Proposition A\.1\(Residual\-preserving isolation of the fluctuation term\)\.
For everyO∈GrO\\in G\_\{r\},
12⟨r,M\(O\)−1r⟩p=12⟨r,M−1r⟩p\.\\frac\{1\}\{2\}\\left\\langle r,M\(O\)^\{\-1\}r\\right\\rangle\_\{p\}=\\frac\{1\}\{2\}\\langle r,M^\{\-1\}r\\rangle\_\{p\}\.\(146\)Hence, at fixed conditioned macrostaterrand fixed density\-of\-states geometryBa\(r\)B\_\{a\}\(r\), the orientational variation of the quadratic local free energy
𝒢loc=FM\(r,z\)\+Φfluc\(M,Ba\(r\)\)\+const\\mathcal\{G\}\_\{\\mathrm\{loc\}\}=F\_\{M\}\(r;z\)\+\\Phi\_\{\\mathrm\{fluc\}\}\(M;B\_\{a\}\(r\)\)\+\\mathrm\{const\}\(147\)alongGrG\_\{r\}is exactly the orientational variation ofΦfluc\\Phi\_\{\\mathrm\{fluc\}\}\.
###### Proof\.
SinceM\(O\)−1=OM−1O∗M\(O\)^\{\-1\}=OM^\{\-1\}O^\{\\ast\}andO∗r=rO^\{\\ast\}r=r,
⟨r,M\(O\)−1r⟩p=⟨O∗r,M−1O∗r⟩p=⟨r,M−1r⟩p\.\\langle r,M\(O\)^\{\-1\}r\\rangle\_\{p\}=\\langle O^\{\\ast\}r,M^\{\-1\}O^\{\\ast\}r\\rangle\_\{p\}=\\langle r,M^\{\-1\}r\\rangle\_\{p\}\.\(148\)At fixedrr, the density\-of\-states term−σξ2logΩz\(r\)\-\\sigma\_\{\\xi\}^\{2\}\\log\\Omega\_\{z\}\(r\)is also independent of the rotation\. Therefore onlyΦfluc\\Phi\_\{\\mathrm\{fluc\}\}varies along the residual\-preserving orbit\. ∎
For the remainder of the appendix, write
B≡Ba\(r\),H=M−1\+B\.B\\equiv B\_\{a\}\(r\),\\qquad H=M^\{\-1\}\+B\.\(149\)Let
P⟂≡I−rr∗‖r‖p2P\_\{\\perp\}\\equiv I\-\\frac\{rr^\{\\ast\}\}\{\\\|r\\\|\_\{p\}^\{2\}\}\(150\)denote the orthogonal projector ontor⟂∩ℋar^\{\\perp\}\\cap\\mathcal\{H\}\_\{a\}\. For a general residual\-preserving infinitesimal rotation, the first variation from Eq\. equation[87](https://arxiv.org/html/2609.09589#S4.E87)becomes
δΦfluc=σξ22Tr\(\[M−1,H−1\]Ξ\),Ξr=0,\\delta\\Phi\_\{\\mathrm\{fluc\}\}=\\frac\{\\sigma\_\{\\xi\}^\{2\}\}\{2\}\\operatorname\{Tr\}\\left\(\[M^\{\-1\},H^\{\-1\}\]\\Xi\\right\),\\qquad\\Xi r=0,\(151\)Therefore stationarity with respect to all such rotations is equivalent to
P⟂\[M−1,H−1\]P⟂=0\.\\boxed\{P\_\{\\perp\}\[M^\{\-1\},H^\{\-1\}\]P\_\{\\perp\}=0\.\}\(152\)This is the exact restricted stationarity condition without any additional invariant\-subspace assumption\.
A particularly transparent case is obtained when the residual direction is a common invariant mode of both operators,
Mr=mrr,Br=brr\.Mr=m\_\{r\}r,\\qquad Br=b\_\{r\}r\.\(153\)Thenspan\{r\}\\operatorname\{span\}\\\{r\\\}andr⟂r^\{\\perp\}are invariant under bothMMandBB, and we may write
M=mrPr⊕M⟂,B=brPr⊕B⟂,M=m\_\{r\}P\_\{r\}\\oplus M\_\{\\perp\},\\qquad B=b\_\{r\}P\_\{r\}\\oplus B\_\{\\perp\},\(154\)wherePr=rr∗/‖r‖p2P\_\{r\}=rr^\{\\ast\}/\\\|r\\\|\_\{p\}^\{2\}\. Residual\-preserving rotations act only on the\(n−1\)\(n\-1\)\-dimensional orthogonal block\. The fluctuation orientational contribution factorizes as
Φfluc=σξ22log\(1\+mrbr\)\+σξ22logdetr⟂\(I⟂\+M⟂B⟂\)\+const\.\\Phi\_\{\\mathrm\{fluc\}\}=\\frac\{\\sigma\_\{\\xi\}^\{2\}\}\{2\}\\log\(1\+m\_\{r\}b\_\{r\}\)\+\\frac\{\\sigma\_\{\\xi\}^\{2\}\}\{2\}\\log\\det\_\{r^\{\\perp\}\}\(I\_\{\\perp\}\+M\_\{\\perp\}B\_\{\\perp\}\)\+\\mathrm\{const\}\.\(155\)The first term is fixed onGrG\_\{r\}\. Applying Theorems[4\.1](https://arxiv.org/html/2609.09589#S4.Thmtheorem1)and[4\.2](https://arxiv.org/html/2609.09589#S4.Thmtheorem2)to the orthogonal block gives
\[M⟂,B⟂\]=0\[M\_\{\\perp\},B\_\{\\perp\}\]=0\(156\)at restricted stationary orientations, and the restricted minimum pairs the eigenvalues in reverse order,
\(m⟂\)large⟷\(b⟂\)small\.\(m\_\{\\perp\}\)\_\{\\mathrm\{large\}\}\\longleftrightarrow\(b\_\{\\perp\}\)\_\{\\mathrm\{small\}\}\.\(157\)Thus, whenever the current residual direction forms a common invariant mode, the reverse\-pairing result is an exact statement about the total quadratic local free energy on the residual\-preserving orbit\. The main text does not require this additional condition; it characterizes the fluctuation\-induced contribution on the full fixed\-spectrum orbit\.
Ifr=0r=0, the stabilizer is the full orthogonal group and the distinction disappears: the macrostate term vanishes, so the fluctuation contribution is the complete quadratic orientational dependence\.Similar Articles
A Bayesian Filtering Approach for Learning Lagrangian Dynamics from Noisy Measurements
This paper presents a Bayesian filtering approach to learn Lagrangian dynamics from partial, noisy measurements by parameterizing kinetic and potential energies with neural networks and jointly estimating states and parameters via maximum likelihood.
Learning-Induced Dynamical Transition in Recurrent Neural Networks
The paper presents a dynamical mean-field theory for learning-induced transitions in recurrent neural networks, showing how feedback-driven learning shifts network dynamics from chaos to stability.
Time-Varying Deep State Space Models for Sequences with Switching Dynamics
The paper proposes a class of time-varying deep state-space models where dynamics are learned via a basis function expansion, enabling adaptive modeling of switching systems. The approach outperforms time-invariant counterparts on synthetic switching data and a speech denoising task.
Deep Spectral Learning of Embedded Latent Transfer Operators for Stochastic Dynamical Systems
Proposes a spectral learning method for stochastic nonlinear dynamical systems using deep feature spaces and an operator-based latent state-space model, demonstrating stable performance in forecasting and filtering tasks.
Human-Centered Learning Mechanics: A Dynamical Framework for Entropy-Regulated Representation Learning
This paper proposes Human-Centered Learning Mechanics (HCLM), a dynamical and information-theoretic framework for studying open and controlled learning systems. It formalizes entropy regularization through effective information force, derives convergence and generalization results, and provides a conditional interpretation of scaling-law behavior.