Stable initialization without the CLT

arXiv cs.LG Papers

Summary

This paper introduces a uniform-phase initialization method for neural networks with sine activations that eliminates reliance on the Central Limit Theorem, improving stability and performance in representation tasks like image and audio fitting.

arXiv:2609.30633v1 Announce Type: new Abstract: Successful training of deep neural networks is highly dependent on the distribution of the initial weights. If the weights are too large, network training blows up; if they are too small, the model fails to learn features. Stable initialization is the optimal moderation between these two extremes. The conventional theory of random networks uses the Central Limit Theorem to control inter-neuron dependencies, which introduces distributional approximation error and coupling between layers. For networks with sine activations, we derive the uniform-phase initialization, which obviates distributional approximation and fully decouples the layers. Ours is the first work to use the sine function's periodic symmetry. Models trained with the uniform-phase initialization outperform the state of the art in neural representation tasks like image and audio fitting. We find that our untuned models are competitive with the best-tuned baselines from previous work and support $\mu$P width scaling.
Original Article
View Cached Full Text

Cached at: 09/29/26, 09:41 AM

# Stable initialization without the CLT
Source: [https://arxiv.org/html/2609.30633](https://arxiv.org/html/2609.30633)
Simon KuangAffiliation:Department of Mechanical and Aerospace EngineeringAffiliation:University of California, DavisEmail:[slku@ucdavis\.edu](mailto:)Xinfan LinAffiliation:Department of Mechanical and Aerospace EngineeringAffiliation:University of California, DavisEmail:[lxflin@ucdavis\.edu](mailto:)

###### Abstract

Successful training of deep neural networks is highly dependent on the distribution of the initial weights\. If the weights are too large, network training blows up; if they are too small, the model fails to learn features\. Stable initialization is the optimal moderation between these two extremes\. The conventional theory of random networks uses the Central Limit Theorem to control inter\-neuron dependencies, which introduces distributional approximation error and coupling between layers\. For networks with sine activations, we derive the uniform\-phase initialization, which obviates distributional approximation and fully decouples the layers\. Ours is the first work to use the sine function’s periodic symmetry\. Models trained with the uniform\-phase initialization outperform the state of the art in neural representation tasks like image and audio fitting\. We find that our untuned models are competitive with the best\-tuned baselines from previous work and supportμ\\muP width scaling\.

## 1Introduction

Today it is widely accepted that increasing a model’s representational capacity by increasing its widthNNand depthLLleads to stronger performance on complex learning tasks\([Hestness et al\., 2017](https://arxiv.org/html/2609.30633#bib.bib40)\)\. Two decades ago, however, training large models was a serious concern: wide and deep networks are highly sensitive to the initialization of their parameters\([Hastie et al\., 2009](https://arxiv.org/html/2609.30633#bib.bib1), §11\.5\.1\)and suffer from exploding or vanishing gradients, which make first\-order optimization challenging\([Glorot and Bengio, 2010](https://arxiv.org/html/2609.30633#bib.bib7)\)\. These so\-called gradient instabilities motivated the search forstable initializationdistributions for the weights\.

We study amultilayer perceptronf:ℝNin→ℝNoutf\\mathrel\{\\mathop\{\\mathchar 58\\relax\}\}\\mathbb\{R\}^\{N\_\{\\text\{in\}\}\}\\to\\mathbb\{R\}^\{N\_\{\\text\{out\}\}\}, which is the compositionf⁡\(x\)=\(gout∘h∘gin\)​\(x\)f\(x\)=\(g\_\{\\text\{out\}\}\\circ h\\circ g\_\{\\text\{in\}\}\)\(x\), comprising

gin\\displaystyle g\_\{\\text\{in\}\}:ℝNin→ℝN,\\displaystyle\\mathrel\{\\mathop\{\\mathchar 58\\relax\}\}\\mathbb\{R\}^\{N\_\{\\text\{in\}\}\}\\to\\mathbb\{R\}^\{N\},x\\displaystyle x↦σ⁡\(Win​x\+bin\),\\displaystyle\\mapsto\\sigma\(W\_\{\\text\{in\}\}x\+b\_\{\\text\{in\}\}\),theinput layer;gout\\displaystyle g\_\{\\text\{out\}\}:ℝN→ℝNout,\\displaystyle\\mathrel\{\\mathop\{\\mathchar 58\\relax\}\}\\mathbb\{R\}^\{N\}\\to\\mathbb\{R\}^\{N\_\{\\text\{out\}\}\},x\\displaystyle x↦Wout​x\+bout,\\displaystyle\\mapsto W\_\{\\text\{out\}\}x\+b\_\{\\text\{out\}\},theoutput layer; andh\\displaystyle h:ℝN→ℝN,\\displaystyle\\mathrel\{\\mathop\{\\mathchar 58\\relax\}\}\\mathbb\{R\}^\{N\}\\to\\mathbb\{R\}^\{N\},x\\displaystyle x↦\(hL∘hL−1∘⋯∘h1\)\(x\),\\displaystyle\\mapsto\(h\_\{L\}\\circ h\_\{L\-1\}\\circ\\cdots\\circ h\_\{1\}\)\(x\),thehidden layers, defined byhℓ\\displaystyle h\_\{\\ell\}:ℝN→ℝN,\\displaystyle\\mathrel\{\\mathop\{\\mathchar 58\\relax\}\}\\mathbb\{R\}^\{N\}\\to\\mathbb\{R\}^\{N\},x\\displaystyle x↦σ⁡\(Wℓ​x\+bℓ\),\\displaystyle\\mapsto\\sigma\(W\_\{\\ell\}x\+b\_\{\\ell\}\),forℓ∈\{1,2,…,L\}\.\\displaystyle\\text\{for $\\ell\\in\\\{1,2,\\ldots,L\\\}$\}\.
Stable initialization encompasses two questions: \(1\) how to tailor the hidden layers to the activation function and \(2\) how to tailor the input and output layers to the task\.

#### Hidden layers

Early theory required that asL,N→∞L,N\\to\\infty, the initial network’s regularity properties scale asO⁡\(1\)O\(1\), i\.e\. not “blow up”\([LeCun et al\., 2012](https://arxiv.org/html/2609.30633#bib.bib27);[Glorot and Bengio, 2010](https://arxiv.org/html/2609.30633#bib.bib7);[He et al\., 2015](https://arxiv.org/html/2609.30633#bib.bib4)\)\. The edge\-of\-chaos school argues that “blowdown” is equally pathological and strengthened this desideratum toΘ⁡\(1\)\\Theta\(1\)\([Poole et al\., 2016](https://arxiv.org/html/2609.30633#bib.bib17);[Schoenholz et al\., 2017](https://arxiv.org/html/2609.30633#bib.bib18);[Yang and Schoenholz, 2017](https://arxiv.org/html/2609.30633#bib.bib19);[Hayou et al\., 2019](https://arxiv.org/html/2609.30633#bib.bib23)\)\. To understand why this requirement is generally nontrivial for hidden layers, we express the Jacobian ofhhat a given input:

∂∂xh\(x\)=diag\(σ′\(zL\)\)WL⋯diag\(σ′\(z1\)\)W1,\\displaystyle\\mathinner\{\\dfrac\{\\partial\{\}\}\{\\partial\{x\}\}\}h\(x\)=\\mathrm\{diag\}\(\\sigma^\{\\prime\}\(z\_\{L\}\)\)W\_\{L\}\\cdots\\mathrm\{diag\}\(\\sigma^\{\\prime\}\(z\_\{1\}\)\)W\_\{1\},\(1\)wherezℓz\_\{\\ell\}is the preactivation at hidden layerℓ\\ell\. The activation slopesdiag⁡\(σ′​\(zℓ\)\)\\mathrm\{diag\}\(\\sigma^\{\\prime\}\(z\_\{\\ell\}\)\),ℓ∈\{1,…,L\}\\ell\\in\\\{1,\\ldots,L\\\}, depend on the preactivation distribution, which in turn depends on the previous layer’s input and weights\.

Managing these dependencies demands sophisticated analysis, usually in four steps:

1. 1\.Identify the criteria to stabilize, usually preactivation scales and gradient norms\([LeCun et al\., 2012](https://arxiv.org/html/2609.30633#bib.bib27);[Glorot and Bengio, 2010](https://arxiv.org/html/2609.30633#bib.bib7);[Saxe et al\., 2014](https://arxiv.org/html/2609.30633#bib.bib16);[He et al\., 2015](https://arxiv.org/html/2609.30633#bib.bib4);[Hanin and Rolnick, 2018](https://arxiv.org/html/2609.30633#bib.bib28);[Pennington et al\., 2017](https://arxiv.org/html/2609.30633#bib.bib20);[Pennington and Worah, 2018](https://arxiv.org/html/2609.30633#bib.bib21);[Xiao et al\., 2018](https://arxiv.org/html/2609.30633#bib.bib30);[Hayou et al\., 2021](https://arxiv.org/html/2609.30633#bib.bib24);[Yang et al\., 2022](https://arxiv.org/html/2609.30633#bib.bib36)\)\.
2. 2\.Calculate the criteria, approximately or exactly, for a location\-scale parametric family, usually Normal, of preactivation distributions\([Poole et al\., 2016](https://arxiv.org/html/2609.30633#bib.bib17);[Schoenholz et al\., 2017](https://arxiv.org/html/2609.30633#bib.bib18);[Klambauer et al\., 2017](https://arxiv.org/html/2609.30633#bib.bib22);[Hayou et al\., 2019](https://arxiv.org/html/2609.30633#bib.bib23)\)\. This step usually involves calculating or approximating𝔼⁡σ⁡\(Z\)\\expect\\sigma\(Z\),Var​σ​\(Z\)\\mathrm\{Var\}\\sigma\(Z\),𝔼⁡σ′​\(Z\)\\expect\\sigma^\{\\prime\}\(Z\), andVar​σ′​\(Z\)\\mathrm\{Var\}\\sigma^\{\\prime\}\(Z\), whereZZis a random variable that approximates the preactivation\.
3. 3\.Choose distributions forWℓW\_\{\\ell\}andbℓb\_\{\\ell\}such that the preactivation distributions satisfy a layerwise fixed\-point recursion, provided that they are Normally distributed\([LeCun et al\., 2012](https://arxiv.org/html/2609.30633#bib.bib27);[Glorot and Bengio, 2010](https://arxiv.org/html/2609.30633#bib.bib7);[He et al\., 2015](https://arxiv.org/html/2609.30633#bib.bib4);[Mishkin and Matas, 2016](https://arxiv.org/html/2609.30633#bib.bib25);[Hayou et al\., 2021](https://arxiv.org/html/2609.30633#bib.bib24)\)\.
4. 4\.Apply the Central Limit Theorem to argue that in the large\-NNlimit, the preactivations are approximately Normal\([Poole et al\., 2016](https://arxiv.org/html/2609.30633#bib.bib17);[Schoenholz et al\., 2017](https://arxiv.org/html/2609.30633#bib.bib18);[Yang and Schoenholz, 2017](https://arxiv.org/html/2609.30633#bib.bib19);[Lee et al\., 2017](https://arxiv.org/html/2609.30633#bib.bib31);[Jacot et al\., 2018](https://arxiv.org/html/2609.30633#bib.bib32);[Yang, 2019](https://arxiv.org/html/2609.30633#bib.bib33)\)\. This closes the layerwise recursion\.

The Central Limit Theorem is the weakest link in this theory\. It incurs an error ofO⁡\(N−1\)O\(N^\{\-1\}\), and narrow networks are appreciably non\-Gaussian at initialization\([Wolinski and Arbel, 2022](https://arxiv.org/html/2609.30633#bib.bib34)\)\. One solution, pursued in the monograph by[Roberts et al\. \(2022\)](https://arxiv.org/html/2609.30633#bib.bib29), is to expand to second order, reducing the asymptotic error toO⁡\(N−2\)O\(N^\{\-2\}\)\.

We focus on the activation functionsin⁡\(x\)\\sin\(x\)\. The conventional route is taken by[Combette et al\. \(2026\)](https://arxiv.org/html/2609.30633#bib.bib5), who instantiate the fixed\-point recursion \(Step 3\) as a transcendental equation involving the Lambert W function\. The second\-order theory of[Roberts et al\. \(2022, §5\.3\.3,§9\.3\)](https://arxiv.org/html/2609.30633#bib.bib29)classifiessin\\sinin the same universality class astanh\\tanh\. Altogether, previous work on stable\-initialization theory forsin⁡\(x\)\\sin\(x\)uses general properties ofsin⁡\(x\)\\sin\(x\), such as its moments, Taylor series, and range, but—crucially—has ignored its periodicity\.

#### Input and output layers

The first and last layers of a neural network deserve their own attention\. Whereas stable initialization theory is interested in the asymptotic*slope*\(namely, zero, for edge\-of\-chaos initialization\) of the network’s behavior with respect toLLandNN, the absolute scale of the input and output weights and biases determines the*intercept*: no matter how largeLLandNNare,f⁡\(10​x\)f\(10x\)is a very different function thanf⁡\(x\)f\(x\)\. Sinusoidal neural networks are known to exhibit a fragile dependence on the weight scale in general\([Yeom et al\., 2025](https://arxiv.org/html/2609.30633#bib.bib15)\)and input weights in particular\([Sitzmann et al\., 2020](https://arxiv.org/html/2609.30633#bib.bib2);[Combette et al\., 2026](https://arxiv.org/html/2609.30633#bib.bib5)\)\. Many works tune a hyperparameterω0\\omega\_\{0\}that scales the inputs by an isotropic base frequency\. \(The Nyquist frequency can be motivated by sensor resolution\([Yuce et al\., 2022](https://arxiv.org/html/2609.30633#bib.bib26)\), but it follows the sensor rather than the underlying field\.\)[Alsakabi et al\. \(2026\)](https://arxiv.org/html/2609.30633#bib.bib13)use a root\-finding algorithm to match the initial network’s spectrum to the discrete sine transform of the field, which is optimal in a sense, but requires rich information about the training data\.[Tancik et al\. \(2021\)](https://arxiv.org/html/2609.30633#bib.bib38)target the input frequency of a Fourier feature network as a meta\-learning task\.

### 1\.1Summary of our method

Conventional stable initialization theory models preactivations as distributed onℝN\\mathbb\{R\}^\{N\}and centered at 0, which necessitates small or zero biases\. However, because sine is a periodic function, the preactivation may equivalently be viewed as wrapped on the torus\[0,2​π\]N\[0,2\\pi\]^\{N\}, which has a unique translation\-invariant probability measure\. By sampling large biases, we spread the preactivations uniformly around the torus*irrespective of the location and scale of the previous layer*, which corresponds to drawing biases from𝒩⁡\(0,∞\)\\mathcal\{N\}\(0,\\infty\)in a conventional stable initialization \(Appendix[A](https://arxiv.org/html/2609.30633#A1)\)\.

###### Definition 1\.

The*uniform\-phase*initialization for hidden layersx↦sin⁡\(W​x\+b\)x\\mapsto\\sin\(Wx\+b\)is

Wi​j\\displaystyle W\_\{ij\}∼𝒩⁡\(0,2/N\)\\displaystyle\\sim\\mathcal\{N\}\(0,2/N\)independently, andbi\\displaystyle b\_\{i\}∼𝒰⁡\(\[0,2​π\]\)\\displaystyle\\sim\\mathcal\{U\}\(\[0,2\\pi\]\)independently\.

As a consequence, miraculously, all of the entries of all of the matrices in equation[1](https://arxiv.org/html/2609.30633#S1.E1)become not only jointly independent but also independent of the input \(Lemma[1](https://arxiv.org/html/2609.30633#Thmlemma1)\)\. The essence is summarized by the following fact:

> LetXXhave*any*distribution onℝ\\mathbb\{R\}, and letb∼𝒰⁡\(\[0,2​π\]\)b\\sim\\mathcal\{U\}\(\[0,2\\pi\]\)be independent ofXX\. Thencos⁡\(X\+b\)\\cos\(X\+b\)is independent ofXX\.

We state and prove the neural network version \(Lemma[1](https://arxiv.org/html/2609.30633#Thmlemma1)\) and then deploy it to prove our main result \(Theorem[1](https://arxiv.org/html/2609.30633#Thmtheorem1)\) on the stability of hidden layers\. Our hidden\-layer theory facilitates an auxiliary result on initializing the input and output layers \(Section[3](https://arxiv.org/html/2609.30633#S3)\), which lends interpretability to the input and output scales by relating them to non\-asymptotic average\-case regularity properties of the initial network\.

Altogether, for every fixed inputxx, our initialization distribution guarantees:

𝔼⁡J⊺​J=𝔼⁡JJ⊺=I,\\expect J^\{\\intercal\}J=\\expect JJ^\{\\intercal\}=I,simultaneously for all JacobiansJJbetween*any*two hidden layers, and

𝔼⁡f⁡\(x\),Covf⁡\(x\),and𝔼⁡∇f​\(x\)​\(Covf⁡\(x\)\)−1​\(∇f​\(x\)\)⊺\\expect f\(x\),\\quad\\mathrm\{Cov\}f\(x\),\\quad\\text\{and\}\\quad\\expect\\nabla f\(x\)\\mathinner\{\\left\(\\mathrm\{Cov\}f\(x\)\\right\)\}^\{\-1\}\(\\nabla f\(x\)\)^\{\\intercal\}all attain prescribed values exactly\. They can be imposed as an inductive bias or moment\-matched to data\.

### 1\.2Contributions

We propose a simple and—to our knowledge, without precedent—exact solution to stable initialization at any width and depth that harnesses the periodicity of sinusoidal activation\. We also propose an exact moment\-matching criterion for zero\-shot tuning of the input and output layer initialization\. Our method is architecture\-agnostic and adapts to the spatial resolution and output statistics of the neural representation task’s target\. Untuned, it outperforms fine\-tuned baselines for sinusoidal network initialization\.

We validate the conventional wisdom on the effectiveness of stable initialization by comparing our initialization to stable initializations in the literature\. On most examples, our method meets or exceeds the bar set by previous work\. We also test the learning\-rate scaling of the maximal\-update parametrization \(μ\\muP\) for neural representations and question the necessity of its architecture\-dependent last\-layer scaling\.

### 1\.3Related work

#### Sinusoidal activations and neural representation

The sine function satisfies the hypotheses of universal approximation theory in neural networks; in fact, similar guarantees can be obtained using the older and more robust theory of Fourier analysis\([Parascandolo et al\., 2017](https://arxiv.org/html/2609.30633#bib.bib3)\)\. However, multilayer networks with sinusoidal activation have proved remarkably effective for the task ofneural representation: learning real\-world signals represented as functions of an input coordinate, e\.g\. audio as a map from time to sound pressure, or an image as a map from spatial coordinates to pixel values\([Sitzmann et al\., 2020](https://arxiv.org/html/2609.30633#bib.bib2);[Dupont et al\., 2021](https://arxiv.org/html/2609.30633#bib.bib9);[Strümpler et al\., 2022](https://arxiv.org/html/2609.30633#bib.bib8)\)\. A related idea \(Fourier features\) is to use sinusoids in the input layer, but with fixed isotropic weights\([Tancik et al\., 2020](https://arxiv.org/html/2609.30633#bib.bib37);[Mildenhall et al\., 2020](https://arxiv.org/html/2609.30633#bib.bib39)\)\.

#### Initializing sinusoidal networks

[Sitzmann et al\. \(2020\)](https://arxiv.org/html/2609.30633#bib.bib2)\(henceforth SM20\) attempt to impose zero mean and unit variance on the preactivation distribution but do not account for gradient blowup with network depth\. Although input\-frequency tuning partly mitigates this weakness from an NTK perspective\([Belbute\-Peres and Kolter, 2023](https://arxiv.org/html/2609.30633#bib.bib6)\),[Combette et al\. \(2026\)](https://arxiv.org/html/2609.30633#bib.bib5)\(henceforth CVP26\) are the first to carry out the complete stable\-initialization procedure to control both preactivations and gradients\.[Novello et al\. \(2025\)](https://arxiv.org/html/2609.30633#bib.bib10);[Yeom et al\. \(2025\)](https://arxiv.org/html/2609.30633#bib.bib15)are explicitly interested in controlling the frequency power spectrum of the initialized network\.[Alsakabi et al\. \(2026\)](https://arxiv.org/html/2609.30633#bib.bib13)deterministically match it to the training data\.[Ben\-Shabat et al\. \(2022\)](https://arxiv.org/html/2609.30633#bib.bib14)propose an initialization geared specifically toward signed distance functions\.

#### Scaling, feature learning, andμ\\muP

Maximal update parametrization \(μ\\muP\) is a theory for scaling hyperparameters with widthNNat a fixed depth \(the main use case is to tune a model’s hyperparameters at a small scale and subsequently transfer them to a much wider model\)\([Yang et al\., 2022](https://arxiv.org/html/2609.30633#bib.bib36);[Yang et al\., 2024](https://arxiv.org/html/2609.30633#bib.bib11);[Dey et al\., 2026](https://arxiv.org/html/2609.30633#bib.bib35)\)\. But it only gives the slope and not the intercept; for sine\-activated networks, the absolute scales depend on the task and have an outsized effect on training\([Yeom et al\., 2025](https://arxiv.org/html/2609.30633#bib.bib15)\)\. While the predictions ofμ\\muP are asymptoticΘ⁡\(⋅\)\\Theta\(\\cdot\)\- andO⁡\(⋅\)O\(\\cdot\)\-type results, our results are exact equalities for any width and depth\.μ\\muP’s learning\-rate scaling is orthogonal to our work, but we test it nevertheless on our initializations\.μ\\muP’s initialization scaling agrees with our exact parameters in all but the final layer; we discuss this subtlety in Section[4\.5](https://arxiv.org/html/2609.30633#S4.SS5)\.

## 2Theory of the uniform\-phase initialization for hidden layers

Specializing the hidden maphhfrom Section[1](https://arxiv.org/html/2609.30633#S1)toσ=sin\\sigma=\\sin, letx0=xx\_\{0\}=xand write the layerwise pre\- and post\-activations as

zℓ=Wℓ​xℓ−1\+bℓ,xℓ=sin⁡\(zℓ\),ℓ∈\[1​…​L\],\\displaystyle z\_\{\\ell\}=W\_\{\\ell\}x\_\{\\ell\-1\}\+b\_\{\\ell\},\\qquad x\_\{\\ell\}=\\sin\(z\_\{\\ell\}\),\\qquad\\ell\\in\[1\\ldots L\],\(2\)so thath⁡\(x\)=xLh\(x\)=x\_\{L\}\. Introduce the notationCos⁡x=diag⁡\(cos⁡\(x\)\)\\Cos x=\\operatorname\{diag\}\(\\cos\(x\)\), wherecos\\cosis applied elementwise\. All expectations in this section are over the initialization distribution, withx0x\_\{0\}arbitrary\.

We now state and prove a few useful properties of the uniform\-phase initialization\.

###### Lemma 1\(Independent Jacobian factorization\)\.

In the uniform\-phase initialization, the terms\{Wℓ\}ℓ=1L∪\{cos⁡zℓ\}ℓ=1L\\\{W\_\{\\ell\}\\\}\_\{\\ell=1\}^\{L\}\\cup\\\{\\cos z\_\{\\ell\}\\\}\_\{\\ell=1\}^\{L\}are independent\.

###### Proof\.

Letϕ1,…,ϕL:ℝN→ℝ\\phi\_\{1\},\\ldots,\\phi\_\{L\}\\mathrel\{\\mathop\{\\mathchar 58\\relax\}\}\\mathbb\{R\}^\{N\}\\to\\mathbb\{R\}andψ1,…,ψL:ℝN×N→ℝ\\psi\_\{1\},\\ldots,\\psi\_\{L\}\\mathrel\{\\mathop\{\\mathchar 58\\relax\}\}\\mathbb\{R\}^\{N\\times N\}\\to\\mathbb\{R\}be arbitrary indicator functions of measurable sets\.

𝔼∏ℓ=1Lϕℓ\(cos\(zℓ\)\)ψℓ\(Wℓ\)\\displaystyle\\expect\\prod\_\{\\ell=1\}^\{L\}\\phi\_\{\\ell\}\(\\cos\(z\_\{\\ell\}\)\)\\psi\_\{\\ell\}\(W\_\{\\ell\}\)=𝔼⁡ϕL​\(cos⁡\(WL​xL−1\+bL\)\)​ψL​\(WL\)​∏ℓ=1L−1ϕℓ​\(cos⁡\(Wℓ​xℓ−1\+bℓ\)\)​ψℓ​\(Wℓ\)\\displaystyle=\\expect\\phi\_\{L\}\(\\cos\(W\_\{L\}x\_\{L\-1\}\+b\_\{L\}\)\)\\psi\_\{L\}\(W\_\{L\}\)\\prod\_\{\\ell=1\}^\{L\-1\}\\phi\_\{\\ell\}\(\\cos\(W\_\{\\ell\}x\_\{\\ell\-1\}\+b\_\{\\ell\}\)\)\\psi\_\{\\ell\}\(W\_\{\\ell\}\)=𝔼⁡𝔼bL⁡\[ϕL​\(cos⁡\(WL​xL−1\+bL\)\)\]⏟=constant​ψL​\(WL\)​∏ℓ=1L−1ϕℓ​\(cos⁡\(Wℓ​xℓ−1\+bℓ\)\)​ψℓ​\(Wℓ\)\\displaystyle=\\expect\\underbrace\{\\expect\_\{b\_\{L\}\}\\mathinner\{\\left\[\\phi\_\{L\}\(\\cos\(W\_\{L\}x\_\{L\-1\}\+b\_\{L\}\)\)\\right\]\}\}\_\{=\\text\{constant\}\}\\psi\_\{L\}\(W\_\{L\}\)\\prod\_\{\\ell=1\}^\{L\-1\}\\phi\_\{\\ell\}\(\\cos\(W\_\{\\ell\}x\_\{\\ell\-1\}\+b\_\{\\ell\}\)\)\\psi\_\{\\ell\}\(W\_\{\\ell\}\)\(conditioning on all randomness other thanbLb\_\{L\}\)=𝔼⁡\[ϕL​\(cos⁡\(WL​xL−1\+bL\)\)\]​𝔼⁡ψL​\(WL\)​∏ℓ=1L−1ϕℓ​\(cos⁡\(Wℓ​xℓ−1\+bℓ\)\)​ψℓ​\(Wℓ\)\\displaystyle=\\expect\\mathinner\{\\left\[\\phi\_\{L\}\(\\cos\(W\_\{L\}x\_\{L\-1\}\+b\_\{L\}\)\)\\right\]\}\\expect\\psi\_\{L\}\(W\_\{L\}\)\\prod\_\{\\ell=1\}^\{L\-1\}\\phi\_\{\\ell\}\(\\cos\(W\_\{\\ell\}x\_\{\\ell\-1\}\+b\_\{\\ell\}\)\)\\psi\_\{\\ell\}\(W\_\{\\ell\}\)\(⋆\\star\)=𝔼⁡\[ϕL​\(cos⁡\(WL​xL−1\+bL\)\)\]​𝔼​ψL​\(WL\)​𝔼​∏ℓ=1L−1ϕℓ​\(cos⁡\(Wℓ​xℓ−1\+bℓ\)\)​ψℓ​\(Wℓ\)\\displaystyle=\\expect\\mathinner\{\\left\[\\phi\_\{L\}\(\\cos\(W\_\{L\}x\_\{L\-1\}\+b\_\{L\}\)\)\\right\]\}\\expect\\psi\_\{L\}\(W\_\{L\}\)\\expect\\prod\_\{\\ell=1\}^\{L\-1\}\\phi\_\{\\ell\}\(\\cos\(W\_\{\\ell\}x\_\{\\ell\-1\}\+b\_\{\\ell\}\)\)\\psi\_\{\\ell\}\(W\_\{\\ell\}\)\(WLW\_\{L\}is independent\)=∏ℓ=1L𝔼⁡ϕℓ​\(cos⁡\(zℓ\)\)​𝔼​ψℓ​\(Wℓ\)\\displaystyle=\\prod\_\{\\ell=1\}^\{L\}\\expect\\phi\_\{\\ell\}\(\\cos\(z\_\{\\ell\}\)\)\\expect\\psi\_\{\\ell\}\(W\_\{\\ell\}\)\(induction onLL\)The crucial step, \(⋆\\star\), uses the fact thatWL​xL−1\+bLW\_\{L\}x\_\{L\-1\}\+b\_\{L\}is uniformly distributed modulo2​π2\\pifor any fixedWL​xL−1W\_\{L\}x\_\{L\-1\}\. Averaging overbLb\_\{L\}therefore yields a constant, which can be pulled out of the expectation\. ∎

###### Lemma 2\(Moments of uniform\-phase sinusoids\)\.

Letbbbe uniformly distributed on the torus\[0,2​π\]N\[0,2\\pi\]^\{N\}\. Then𝔼⁡cos⁡\(b\)=0\\expect\\cos\(b\)=0and𝔼⁡\[cos\(b\)cos\(b\)⊺\]=12​I\\expect\\mathinner\{\\left\[\\cos\(b\)\\cos\(b\)^\{\\intercal\}\\right\]\}=\\frac\{1\}\{2\}I\.

###### Proof\.

By circular symmetry in each coordinate,𝔼⁡cos⁡\(b\)=0\\expect\\cos\(b\)=0\. Because the coordinates are independent, the off\-diagonal entries of𝔼⁡\[cos\(b\)cos\(b\)⊺\]\\expect\\mathinner\{\\left\[\\cos\(b\)\\cos\(b\)^\{\\intercal\}\\right\]\}are zero\. On the diagonal,𝔼⁡cos2⁡b1=𝔼⁡sin2⁡b1\\expect\\cos^\{2\}b\_\{1\}=\\expect\\sin^\{2\}b\_\{1\}by circular symmetry, while𝔼⁡cos2⁡b1\+𝔼⁡sin2⁡b1=1\\expect\\cos^\{2\}b\_\{1\}\+\\expect\\sin^\{2\}b\_\{1\}=1by the Pythagorean identity; therefore𝔼⁡cos2⁡b1=12\\expect\\cos^\{2\}b\_\{1\}=\\frac\{1\}\{2\}\. ∎

Combining these inductively proves our main result on hidden layers\.

###### Theorem 1\.

Let0≤s≤t≤L0\\leq s\\leq t\\leq Lbe integers\. Then for all inputsx0x\_\{0\}, the hidden layers equation[2](https://arxiv.org/html/2609.30633#S2.E2)satisfy

𝔼⁡\(∂xt∂xs\)​\(∂xt∂xs\)⊺\\displaystyle\\expect\\mathinner\{\\left\(\\mathinner\{\\dfrac\{\\partial\{\}x\_\{t\}\}\{\\partial\{x\_\{s\}\}\}\}\\right\)\}\\mathinner\{\\left\(\\mathinner\{\\dfrac\{\\partial\{\}x\_\{t\}\}\{\\partial\{x\_\{s\}\}\}\}\\right\)\}^\{\\intercal\}=I,\\displaystyle=I,𝔼⁡\(∂xt∂xs\)⊺​\(∂xt∂xs\)\\displaystyle\\expect\\mathinner\{\\left\(\\mathinner\{\\dfrac\{\\partial\{\}x\_\{t\}\}\{\\partial\{x\_\{s\}\}\}\}\\right\)\}^\{\\intercal\}\\mathinner\{\\left\(\\mathinner\{\\dfrac\{\\partial\{\}x\_\{t\}\}\{\\partial\{x\_\{s\}\}\}\}\\right\)\}=I\.\\displaystyle=I\.\(3\)

Writeδst≔∂xt∂xs\\delta\_\{s\}^\{t\}\\coloneqq\\mathinner\{\\dfrac\{\\partial\{\}x\_\{t\}\}\{\\partial\{x\_\{s\}\}\}\}\. By the chain rule,δst\\delta\_\{s\}^\{t\}satisfies both a forward recursion intt\(withssfixed\) and a backward recursion inss\(withttfixed\):

δst\\displaystyle\\delta\_\{s\}^\{t\}=Cos⁡zt​Wt​δst−1,\\displaystyle=\\Cos z\_\{t\}W\_\{t\}\\delta\_\{s\}^\{t\-1\},δss\\displaystyle\\delta\_\{s\}^\{s\}=I,\\displaystyle=I,\(4a\)δst\\displaystyle\\delta\_\{s\}^\{t\}=δs\+1t​Cos⁡zs\+1​Ws\+1,\\displaystyle=\\delta\_\{s\+1\}^\{t\}\\Cos z\_\{s\+1\}W\_\{s\+1\},δtt\\displaystyle\\delta\_\{t\}^\{t\}=I\.\\displaystyle=I\.\(4b\)Note thatδst\\delta\_\{s\}^\{t\}depends only onxsx\_\{s\},\{Wℓ\}ℓ=s\+1t\\\{W\_\{\\ell\}\\\}\_\{\\ell=s\+1\}^\{t\}, and\{bℓ\}ℓ=s\+1t\\\{b\_\{\\ell\}\\\}\_\{\\ell=s\+1\}^\{t\}; we shall use this fact to argue for independence in later steps\.

###### Proof that𝔼⁡δst​\(δst\)⊺=I\\expect\\delta\_\{s\}^\{t\}\(\\delta\_\{s\}^\{t\}\)^\{\\intercal\}=I\.

We prove this by induction ontt, withssfixed\. The base caset=st=sis immediate from equation[4a](https://arxiv.org/html/2609.30633#S2.E4.1)\. For the inductive step, we assume that𝔼⁡δst−1​\(δst−1\)⊺=I\\expect\\delta\_\{s\}^\{t\-1\}\(\\delta\_\{s\}^\{t\-1\}\)^\{\\intercal\}=Iand want to show that𝔼⁡δst​\(δst\)⊺=I\\expect\\delta\_\{s\}^\{t\}\(\\delta\_\{s\}^\{t\}\)^\{\\intercal\}=I\. By equation[4a](https://arxiv.org/html/2609.30633#S2.E4.1),

δst​\(δst\)⊺\\displaystyle\\delta\_\{s\}^\{t\}\(\\delta\_\{s\}^\{t\}\)^\{\\intercal\}=Cos⁡zt​Wt​δst−1​\(δst−1\)⊺​Wt⊺​Cos⁡zt\.\\displaystyle=\\Cos z\_\{t\}W\_\{t\}\\delta\_\{s\}^\{t\-1\}\(\\delta\_\{s\}^\{t\-1\}\)^\{\\intercal\}W\_\{t\}^\{\\intercal\}\\Cos z\_\{t\}\.\(5\)According to Lemma[1](https://arxiv.org/html/2609.30633#Thmlemma1),δst−1\\delta\_\{s\}^\{t\-1\},WtW\_\{t\}, andCos⁡zt\\Cos z\_\{t\}are mutually independent\. Therefore, we may evaluate this expectation from the inside out by repeatedly conditioning on the outer terms:𝔼⁡δst​\(δst\)⊺\\displaystyle\\expect\\delta\_\{s\}^\{t\}\(\\delta\_\{s\}^\{t\}\)^\{\\intercal\}=𝔼⁡Cos⁡zt​Wt​\(𝔼⁡δst−1​\(δst−1\)⊺\)​Wt⊺​Cos⁡zt\\displaystyle=\\expect\\Cos z\_\{t\}W\_\{t\}\\mathinner\{\\left\(\\expect\\delta\_\{s\}^\{t\-1\}\(\\delta\_\{s\}^\{t\-1\}\)^\{\\intercal\}\\right\)\}W\_\{t\}^\{\\intercal\}\\Cos z\_\{t\}\(6\)=𝔼⁡Cos⁡zt​Wt​Wt⊺​Cos⁡zt\\displaystyle=\\expect\\Cos z\_\{t\}W\_\{t\}W\_\{t\}^\{\\intercal\}\\Cos z\_\{t\}\(7\)=𝔼⁡Cos⁡zt​\(𝔼⁡Wt​Wt⊺\)​Cos⁡zt\\displaystyle=\\expect\\Cos z\_\{t\}\\mathinner\{\\left\(\\expect W\_\{t\}W\_\{t\}^\{\\intercal\}\\right\)\}\\Cos z\_\{t\}\(8\)=2​𝔼⁡Cos⁡zt​Cos⁡zt\\displaystyle=2\\expect\\Cos z\_\{t\}\\Cos z\_\{t\}\(9\)=I\\displaystyle=I\(10\)by Lemma[2](https://arxiv.org/html/2609.30633#Thmlemma2)\. ∎

###### Proof that𝔼⁡\(δst\)⊺​δst=I\\expect\(\\delta\_\{s\}^\{t\}\)^\{\\intercal\}\\delta\_\{s\}^\{t\}=I\.

We prove this by downward induction onss, withttfixed\. The base cases=ts=tis immediate from equation[4b](https://arxiv.org/html/2609.30633#S2.E4.2)\. For the inductive step, we assume that𝔼⁡\(δs\+1t\)⊺​δs\+1t=I\\expect\(\\delta\_\{s\+1\}^\{t\}\)^\{\\intercal\}\\delta\_\{s\+1\}^\{t\}=Iand want to show that𝔼⁡\(δst\)⊺​δst=I\\expect\(\\delta\_\{s\}^\{t\}\)^\{\\intercal\}\\delta\_\{s\}^\{t\}=I\. By equation[4b](https://arxiv.org/html/2609.30633#S2.E4.2),

\(δst\)⊺​δst\\displaystyle\(\\delta\_\{s\}^\{t\}\)^\{\\intercal\}\\delta\_\{s\}^\{t\}=Ws\+1⊺​Cos⁡zs\+1​\(δs\+1t\)⊺​δs\+1t​Cos​zs\+1​Ws\+1\.\\displaystyle=W\_\{s\+1\}^\{\\intercal\}\\Cos z\_\{s\+1\}\(\\delta\_\{s\+1\}^\{t\}\)^\{\\intercal\}\\delta\_\{s\+1\}^\{t\}\\Cos z\_\{s\+1\}W\_\{s\+1\}\.Using the same conditioning tactic as above,𝔼⁡\(δst\)⊺​δst\\displaystyle\\expect\(\\delta\_\{s\}^\{t\}\)^\{\\intercal\}\\delta\_\{s\}^\{t\}=𝔼⁡Ws\+1⊺​Cos​zs\+1​\(𝔼⁡\(δs\+1t\)⊺​δs\+1t\)​Cos​zs\+1​Ws\+1\\displaystyle=\\expect W\_\{s\+1\}^\{\\intercal\}\\Cos z\_\{s\+1\}\\mathinner\{\\left\(\\expect\(\\delta\_\{s\+1\}^\{t\}\)^\{\\intercal\}\\delta\_\{s\+1\}^\{t\}\\right\)\}\\Cos z\_\{s\+1\}W\_\{s\+1\}=𝔼⁡Ws\+1⊺​Cos⁡zs\+1​Cos​zs\+1​Ws\+1\\displaystyle=\\expect W\_\{s\+1\}^\{\\intercal\}\\Cos z\_\{s\+1\}\\Cos z\_\{s\+1\}W\_\{s\+1\}=𝔼⁡Ws\+1⊺​\(𝔼⁡Cos⁡zs\+1​Cos⁡zs\+1\)​Ws\+1\\displaystyle=\\expect W\_\{s\+1\}^\{\\intercal\}\\mathinner\{\\left\(\\expect\\Cos z\_\{s\+1\}\\Cos z\_\{s\+1\}\\right\)\}W\_\{s\+1\}=12​𝔼⁡Ws\+1⊺​Ws\+1=I\.\\displaystyle=\\frac\{1\}\{2\}\\expect W\_\{s\+1\}^\{\\intercal\}W\_\{s\+1\}=I\.∎

This theorem guarantees that the Jacobian \(or gradient\) between any two layers is isotropic, and its average squared singular value is 1\. In the backpropagation direction, early training loss gradients neither grow nor decay on average as they propagate backward through the network\. In the forward direction, the network, as a mathematical function that maps inputs to outputs, possesses zeroth\-order \(range\) and first\-order \(slope\) properties that are independent of width and depth\.

## 3Moment matching for input and output layers

Thus far, we have dealt only withhh, the network’s hidden\-layer map, which hasNNinputs,NNoutputs, and depthLL\. To complete the networkf=gout∘h∘ginf=g\_\{\\text\{out\}\}\\circ h\\circ g\_\{\\text\{in\}\}fromℝNin\\mathbb\{R\}^\{N\_\{\\text\{in\}\}\}toℝNout\\mathbb\{R\}^\{N\_\{\\text\{out\}\}\}as defined in the introduction, we must initialize the input layerging\_\{\\text\{in\}\}and the output layergoutg\_\{\\text\{out\}\}\. We drawbinb\_\{\\text\{in\}\}from a uniform distribution on\[0,2​π\]\[0,2\\pi\], as we do for the hidden\-layer biases\. This establishes translation invariance and eliminates the need to center the data\.

For target momentsμ\\mu,Σ≻0\\Sigma\\succ 0, andΩ⪰0\\Omega\\succeq 0, we initialize the remaining three parametersWin∈ℝN×NinW\_\{\\text\{in\}\}\\in\\mathbb\{R\}^\{N\\times N\_\{\\text\{in\}\}\},Wout∈ℝNout×NW\_\{\\text\{out\}\}\\in\\mathbb\{R\}^\{N\_\{\\text\{out\}\}\\times N\}, andbout∈ℝNoutb\_\{\\text\{out\}\}\\in\\mathbb\{R\}^\{N\_\{\\text\{out\}\}\}by moment matching:

μ\\displaystyle\\mu=𝔼⁡f⁡\(x\)=𝔼⁡bout,\\displaystyle=\\expect f\(x\)=\\expect b\_\{\\text\{out\}\},the output mean,\(11a\)Σ\\displaystyle\\Sigma=Cov​f​\(x\)=12​𝔼⁡Wout​Wout⊺,\\displaystyle=\\mathrm\{Cov\}f\(x\)=\\tfrac\{1\}\{2\}\\expect W\_\{\\text\{out\}\}W\_\{\\text\{out\}\}^\{\\intercal\},the output covariance, and\(11b\)Ω\\displaystyle\\Omega=𝔼⁡∂f⁡\(x\)∂x⊺​Σ−1​∂f⁡\(x\)∂x=NoutN​𝔼⁡Win⊺​Win,\\displaystyle=\\expect\\mathinner\{\\dfrac\{\\partial\{\}f\(x\)\}\{\\partial\{x\}\}\}^\{\\intercal\}\\Sigma^\{\-1\}\\mathinner\{\\dfrac\{\\partial\{\}f\(x\)\}\{\\partial\{x\}\}\}=\\frac\{N\_\{\\text\{out\}\}\}\{N\}\\expect W\_\{\\text\{in\}\}^\{\\intercal\}W\_\{\\text\{in\}\},the normalized structure tensor\.\(11c\)Because the uniform\-phase initialization is translation invariant, these moments may be interpreted as being evaluated at any particularxxor over a representative distribution ofxx\. Afterμ\\mu,Σ\\Sigma, andΩ\\Omegaare either prescribed or estimated from data, we match these target moments by setting

bout\\displaystyle b\_\{\\text\{out\}\}=μ^,\\displaystyle=\\hat\{\\mu\},\(12a\)Wout\\displaystyle W\_\{\\text\{out\}\}∼𝒩⁡\(0,2​Σ^/N\)​columns,\\displaystyle\\sim\\mathcal\{N\}\(0,2\\hat\{\\Sigma\}/N\)\\text\{ columns\},and\(12b\)Win\\displaystyle W\_\{\\text\{in\}\}∼𝒩⁡\(0,Ω^/Nout\)​rows\.\\displaystyle\\sim\\mathcal\{N\}\(0,\\hat\{\\Omega\}/N\_\{\\text\{out\}\}\)\\text\{ rows\}\.\(12c\)The key idea is to useΩ\\Omegato calibrate the input scale in the “physical” \(that is, nonfrequency\) domain\. By initializingWoutW\_\{\\text\{out\}\}andWinW\_\{\\text\{in\}\}in this way, we nondimensionalize both the inputs and the outputs for the hidden section of the neural network\.

In prior work\([Sitzmann et al\., 2020](https://arxiv.org/html/2609.30633#bib.bib2);[Combette et al\., 2026](https://arxiv.org/html/2609.30633#bib.bib5)\), the input weights are scaled using a “characteristic frequency”ω0\\omega\_\{0\}: the Nyquist frequency \(in audio\),∼30\\sim 30periods \(in images\),∼2\\sim 2periods \(in physics\-informed neural networks\)\. However, this and other scales have an outsized effect on learning and require tuning\([Belbute\-Peres and Kolter, 2023](https://arxiv.org/html/2609.30633#bib.bib6);[Yeom et al\., 2025](https://arxiv.org/html/2609.30633#bib.bib15)\)\. By contrast, matching the observed value ofΩ\\Omegaallows the network to adapt to the regularity characteristics of the data\.

## 4Numerical experiments

Stable initialization is a mathematical definition accompanied by a scientific hypothesis: a stably initialized network will perform better than an unstably initialized network after thousands of epochs under the same training operator\. Our construction, which exactly calibrates initialization\-ensemble second moments at any depth and width, establishes the ideal laboratory conditions in which to test this hypothesis\. Appendix[C](https://arxiv.org/html/2609.30633#A3)gives full details of the tasks, preprocessing, model implementations, optimization, metrics, and hyperparameter sweeps\.

### 4\.1Image fitting

We fit the Cameraman image on a128×128128\\times 128training grid and evaluate it on a512×512512\\times 512grid to assess generalization\. We compare the following architectures \(details in Appendix[C](https://arxiv.org/html/2609.30633#A3)\):

SIREN \(UniformPhase\-MM\)Our proposed initialization, which uses uniform\-phase biases and moment\-matched input and output layers, estimatesμ\\mu,Σ\\Sigma, andΩ\\Omegafrom the training data, and requires no hand\-tunedω0\\omega\_\{0\}\.

SIREN \(CVP26σa=0\\sigma\_\{a\}\\\!=\\\!0\)The edge\-of\-chaos initialization of[Combette et al\. \(2026\)](https://arxiv.org/html/2609.30633#bib.bib5)\.

SIREN \(SM20\)The original initialization of[Sitzmann et al\. \(2020\)](https://arxiv.org/html/2609.30633#bib.bib2)\.

Tanh \(PE\-Xavier\)A Tanh MLP with random Fourier\-feature encoding\([Tancik et al\., 2020](https://arxiv.org/html/2609.30633#bib.bib37)\)and Xavier weights\.

ReLU \(Kaiming\), SiLU \(Xavier\), GeLU \(Xavier\)Standard MLPs with the corresponding activation and initialization\.

![Refer to caption](https://arxiv.org/html/2609.30633v1/cameraman_wide.png)Figure 1:Image fitting on the Cameraman image\. Columns: the ground truth followed by one column per initialization, with the evaluation SNR beneath each\. Top row:128×128128\\times 128training image\. Middle row:512×512512\\times 512test image\. Bottom row: zoomed region \(red square in the middle row\)\.The uniform\-phase, moment\-matched initialization achieves the lowest evaluation MSE by a small margin \(Table[1](https://arxiv.org/html/2609.30633#A4.T1)in Appendix[D](https://arxiv.org/html/2609.30633#A4)\), followed by CVP26 \(σa=0\\sigma\_\{a\}=0\) and the remaining replications of[Combette et al\. \(2026\)](https://arxiv.org/html/2609.30633#bib.bib5)\. Visually, the top two methods both preserve fine details without introducing excessive noise \(figure[1](https://arxiv.org/html/2609.30633#S4.F1)\)\.

### 4\.2Audio fitting

We use the same architectures in a replication of the Bach audio experiment of[Sitzmann et al\. \(2020, Appendix 8\)](https://arxiv.org/html/2609.30633#bib.bib2)111The signal is a77s mono audio signal at44\.144\.1kHz\. The time coordinate is scaled to\[−100,100\]\[\-100,100\], the audio is normalized bymax⁡\|y\|\\max\|y\|, and the network is trained and evaluated on the*full*signal \(no held\-out test split\)\. All networks haveL=5L=5hidden layers of widthN=256N=256and are trained for50005000epochs with Adam at a learning rate of5×10−55\\times 10^\{\-5\}\. The Tanh network uses random Fourier features withσ=10\\sigma=10\.\. For the CVP26 and SM20 SIRENs, we replicate the input frequencyω0=30\\omega\_\{0\}=30used by[Sitzmann et al\. \(2020\)](https://arxiv.org/html/2609.30633#bib.bib2)\. For our SIREN, we use the moment\-matched initialization, which estimatesΩ^\\hat\{\\Omega\}from the data and requires no hand\-tuned frequency\. The SM20 SIREN attains the highest fidelity, with an SNR of31\.1531\.15dB, but our moment\-matched SIREN reaches29\.3329\.33dB*without*any hand\-tuned frequency parameter; its input scale is automatically calibrated throughΩ^\\hat\{\\Omega\}\. Our SIREN’s SNR exceeds CVP26’s by nearly 20 dB, and the non\-sinusoidal networks \(ReLU, SiLU, GeLU\) collapse to near\-zero output \(figure[5](https://arxiv.org/html/2609.30633#A4.F5)in Appendix[D](https://arxiv.org/html/2609.30633#A4)\)\.

### 4\.3Narrow networks

We repeat the experiment of §[4\.1](https://arxiv.org/html/2609.30633#S4.SS1)but reduce the width toN=16N=16to stress the infinite\-width assumption underlying the CVP26 and SM20 initializations\. The differences between initializations are amplified as a result\. Our moment\-matched SIREN remains the best performer across all metrics \(SNR: 10\.28 dB\), and the gap to CVP26 widens to 6\.8 dB \(figure[4](https://arxiv.org/html/2609.30633#A4.F4)and Table[2](https://arxiv.org/html/2609.30633#A4.T2)in Appendix[D](https://arxiv.org/html/2609.30633#A4)\)\.

### 4\.4Moment matching \(zero\-shot\) versusω0\\omega\_\{0\}tuning \(hyperparameter search\)

The hyperparameterizations of CVP26 and SM20 call for an isotropic “characteristic frequency”ω0\\omega\_\{0\}to scale the inputs\. Prior work suggestsω0=64​π/6\\omega\_\{0\}=64\\pi/\\sqrt\{6\}for a domain of\[−1,1\]\[\-1,1\]\. By contrast, our method achieves zero\-shot calibration to the anisotropic statistics of the image field; it is invariant to domain translation and equivariant under domain scaling, sampling resolution, and output scaling\. We compare our method’s zero\-shot tuning to a detailed grid search for prior methods’ω0\\omega\_\{0\}\.

Using the same architecture and training as §[4\.1](https://arxiv.org/html/2609.30633#S4.SS1), we evaluate on fourgrayscaleimages fromscikit\-image\(Brick, Camera, Gravel, and Clock\) and four color images \(Hubble Deep Field, Skin, Astronaut, and Coffee\), in bothRGBandCIELABcolor spaces\. For each baseline SIREN initialization \(SM20, CVP26σa=0\\sigma\_\{a\}\\\!=\\\!0, CVP26σa=1\\sigma\_\{a\}\\\!=\\\!1\), we sweep the input\-layer frequencyω0\\omega\_\{0\}over four powers of 10 on a logarithmic grid of 21 values centered at prior work’sω0\\omega\_\{0\}\. The estimated moment\-matched parameters and per\-image best test errors are reported in Appendix[E](https://arxiv.org/html/2609.30633#A5), and the best\-ω0\\omega\_\{0\}images in the Supplementary Material, Appendix[F](https://arxiv.org/html/2609.30633#A6)\.

Figure 2:Grayscale and colorω0\\omega\_\{0\}sweep\. Test MSE versus input\-layer frequencyω0\\omega\_\{0\}; rows are the grayscale, RGB, and CIELAB experiments, and columns are the four images in each\. A star marks the minimum of every sweep, and the dashed red line marks the uniform\-phase MM test MSE\. Each panel has its own logarithmic vertical scale; for absolute errors, see Tables[3](https://arxiv.org/html/2609.30633#A5.T3),[4](https://arxiv.org/html/2609.30633#A5.T4), and[5](https://arxiv.org/html/2609.30633#A5.T5)\.Across thegrayscaleandRGBexperiments, the moment\-matched initialization achieves the lowest test MSE on 5 of 8 images \(Camera, Clock, Skin, Astronaut, and Coffee\)\. On the remaining 3 images \(Brick, Gravel, and Hubble\), it is competitive, losing to the best\-tuned baseline by 19%, 12%, and 1\.5%, respectively\. \(On Brick and Gravel, we attribute the high error to an overemphasis on high frequencies; see figure[7](https://arxiv.org/html/2609.30633#A6.F7)\.\) Even in these cases, the moment\-matched initialization is no worse than mistuningω0\\omega\_\{0\}by20%20\\%\.

The baselines are all capable of achieving similar performance, but atω0\\omega\_\{0\}values that vary widely across images\. For any given method and image, the loss landscape ofω0\\omega\_\{0\}is nonconvex and contains local minima\. The general trend is that SM20 performs best at the smallestω0\\omega\_\{0\}, followed by CVP26 \(OPENσa=1\)\\sigma\_\{a\}=1\), followed by CVP26 \(σa=0\\sigma\_\{a\}=0\)\. This ordering confirms an established finding\([Combette et al\., 2026](https://arxiv.org/html/2609.30633#bib.bib5)\)that SM20 is biased toward high frequencies and large gradients, followed by CVP26 \(σa=1\\sigma\_\{a\}=1\) and CVP26 \(σa=0\\sigma\_\{a\}=0\), in that order\.

The situation differs when we fit in three\-dimensionalCIELABcolor space, which is related to RGB space by a nonlinear transformation that approximates human perception\. Perceived colors in CIELAB space are represented byL∗∈\[0,100\]L^\{\*\}\\in\[0,100\]anda∗,b∗∈\[−128,127\]a^\{\*\},b^\{\*\}\\in\[\-128,127\]\. This transformation preserves the spatial structure of the images but drastically alters the mean and \(anisotropic\) covariance of the color channels\.

For all CIELAB images, our initialization achieves the best fit; the improvement is slight on Hubble and ranges from a factor of 3\.0 to 5\.2 on the other images\. We attribute this to the fact that it correctly adapts to the distribution of the targets as well as to the typical resolution of spatial detail: theΩ^\\hat\{\\Omega\}structure tensors of CIELAB images are very close to those of RGB images \(Appendix[E](https://arxiv.org/html/2609.30633#A5)\)\.

### 4\.5Width scaling usingμ\\muP: initialization and learning rates

In this section, we test how the optimal learning rate depends on widthNN\. The universality theory ofμ\\muP is based on three desiderata\([Yang et al\., 2022](https://arxiv.org/html/2609.30633#bib.bib36), Appendix J\.2\):

1. 1\.Every \(pre\)activation vector in a network should haveΘ⁡\(1\)\\Theta\(1\)\-sized coordinates\.
2. 2\.Neural network output should beO⁡\(1\)O\(1\)\.
3. 3\.All parameters should be updated as much as possible \(in terms of scaling in width\) without leading to divergence\.

Desideratum[1](https://arxiv.org/html/2609.30633#S4.I2.i1)is satisfied by the uniform\-phase component of our initialization\. Desideratum[2](https://arxiv.org/html/2609.30633#S4.I2.i2)is satisfied asΘ⁡\(1\)\\Theta\(1\)by the moment\-matching property of our initialization; this differs fromμ\\muP’s own last\-layer prescription \(Remark[1](https://arxiv.org/html/2609.30633#Thmremark1)\)\. Desideratum[3](https://arxiv.org/html/2609.30633#S4.I2.i3)concerns the learning rate, which is orthogonal to the focus of this paper; for the Adam optimizer, which we use,[Yang et al\. \(2022\)](https://arxiv.org/html/2609.30633#bib.bib36)prescribe scaling the learning rate as1/N1/Nfor non\-output layers\. We test whether this prescription transfers the optimal learning rate across width for sinusoidal networks, and whether it interacts with our moment\-matched initialization\. \(To our knowledge, no prior work has validatedμ\\muP for either sinusoidal MLPs or the neural representation task\.\)

We repeat the experiment of section[4\.1](https://arxiv.org/html/2609.30633#S4.SS1)for each of the four initializations at seven widthsN∈\{8,16,32,64,128,256,512\}N\\in\\\{8,16,32,64,128,256,512\\\}, using each initialization’s bestω0\\omega\_\{0\}from a width\-6464sweep\. At each width, we sweep the Adam learning rate over powers of two,η∈\{2−18,…,22\}\\eta\\in\\\{2^\{\-18\},\\dots,2^\{2\}\\\}\.

standardPlain Adam at the sweptη\\etawith the initialization left unmodified\.

μ\\muP \(LR\)Theμ\\muP learning\-rate scaling above; the initialization is unmodified\.

μ\\muP \(LR\+LL\)The same learning\-rate scaling, with the last\-layer weights also multiplied at initialization by64/N\\sqrt\{64/N\}\.

Under standard scaling, the loss curve shifts left with width, and under bothμ\\muP\-informed scalings, the loss curve as a function of learning rate deepens with increasing width\. These observations validateμ\\muP’s predictions\. But the dominant pattern is the presence of a phase transition immediately to the right of the optimal learning rate, where the loss dramatically increases to a null value \(Figure[6](https://arxiv.org/html/2609.30633#A4.F6)\)\. In bothμ\\muP scalings, the phase transition and the optimal learning rate both shift right with width\. While we have not observed optimal hyperparameter transfer, bothμ\\muPs result in a loss that decreases with increasing width at every learning rate\.

μ\\muP’s last\-layer rescaling inLR\+LLdoes not appreciably differ from theLR\-only scaling\. This supports our theory that, as far as initialization is concerned, moment matching, which is architecture\-agnostic, is as effective for learning\-rate transfer asμ\\muP, which is architecture\-aware\.

## 5Conclusion

Stable initialization of sinusoidal networks rests on two pillars: hidden layer stability and input\-output layer tuning\. Whereas the conventional theory for hidden layers requires subtle recursions and infinite\-width analysis in order to control dependencies between neurons and weights, our method renders all preactivations independent and identically distributed, allowing network gradient moments to be obtained exactly by a short computation\. We simplify the theory and give the strongest possible guarantee, one that holds for all widths and depths\.

Whereas the conventional theory for input\-output tuning relies on hand\-tuning in the frequency domain, we take a physical\-domain perspective on nondimensionalization: canceling the input and output scales using data\-driven estimates of the neural field’s covariance and structure tensor\. The result is zero\-shot initialization tuning on par with a grid search over prior hyperparameterizations\.

Our work suggests that sine may be underappreciated as an activation function in machine learning, and that the uniform\-phase property may prove fruitful in diverse applications beyond neural representation\.

### AI use statement

\(Required disclosure\)In this work, we used AI toolsto implement methods \(replicate CVP26 and SM20, originally written in PyTorch, in JAX; implement our proposed initialization and evaluation harness\)\.We did not use AI toolsto generate synthetic data sets, help develop theoretical models or conceptual frameworks, formulate mathematical claims, provide critical ingredients for proving mathematical claims, assist in the writing of proofs, propose or refine hypotheses, design or provide feedback on research methodology or experiments, assist with translation, clean and reformat dataset, support qualitative and thematic data analysis, or interpret results\.

\(Recommended disclosure\)Additionally, we used AI toolsto summarize or analyse existing literature, discover research topics or identify gaps, search for information, copyedit for grammaticality, and identify relevant literature\.

We have reviewed all AI\-assisted work\.We carefully proofread AI\-generated code for correctness to rule out bugs and reward\-hacking tendencies\. Furthermore, we have verified that our replications are within numerical tolerance of prior work\. In the main body of the paper, our use of AI in writing assistance is strictly limited to technical copyediting; wording and stylistic choices are the product of human authorship\. The problem description and reproducibility appendices are drafted by AI and checked for accuracy\. We prepared the bibliography manually using a reference manager and have verified the relevance of each citation\.We take responsibility for the final content of this work, including text, claims or artifacts produced with the aid of generative AI\.

### Ethics statement

We consider that this work does not raise questions respecting the Code of Ethics\.

### Reproducibility statement

Appendix[C](https://arxiv.org/html/2609.30633#A3)specifies the data, preprocessing, estimators, initialization distributions, architectures, optimization, random seeds, metrics, and sweep grids used in every numerical experiment; the remaining appendices provide complete proofs and additional quantitative and visual results\. The accompanying source archive includes the shared JAX/Equinox implementations insrc/architecture\.pyandsrc/moment\_matching\.py;src/cameraman\-and\-audio\.ipynbreproduces the image, narrow\-network, and audio experiments;src/omega\-sweep\.ipynbandsrc/mup\-lr\-sweep\.ipynbreproduce the two hyperparameter studies;src/jacobian\_svd\_spectrum \(CVP26 comparison\)\.ipynbreproduces the spectral diagnostic; andsrc/image\_fitting\-repro\.ipynbrecords the CVP26 image\-fitting replication used to validate the shared implementation\. The image data are obtained throughscikit\-image, and the Bach waveform is included atsrc/assets/gt\_bach\.wav\. The two sweep notebooks are configured for accelerator\-backed cloud execution and write to mounted/mntdirectories; for local execution, the user must select an available JAX backend and change those output directories to writable local paths\.

## References

- Alsakabiet al\.\(2026\)M\. Alsakabi, K\. Hu, J\. M\. Dolan, and O\. K\. TonguzJA\-SIREN: Deterministic Initialization for Sinusoidal Networks via Spectral Matching\.arXiv\.Note:arXiv:2606\.06671 \[cs\.CV\] version: 1External Links:[Link](http://arxiv.org/abs/2606.06671),[Document](https://dx.doi.org/10.48550/arXiv.2606.06671)Cited by:[§1](https://arxiv.org/html/2609.30633#S1.SS0.SSS0.Px2.p1.1),[§1\.3](https://arxiv.org/html/2609.30633#S1.SS3.SSS0.Px2.p1.1)\.
- Belbute\-Peres and Kolter \(2023\)F\. d\. A\. Belbute\-Peres and J\. Z\. KolterSimple initialization and parametrization of sinusoidal networks via their kernel bandwidth\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=yVqC6gCNf4d)Cited by:[§1\.3](https://arxiv.org/html/2609.30633#S1.SS3.SSS0.Px2.p1.1),[§3](https://arxiv.org/html/2609.30633#S3.p3.1)\.
- Ben\-Shabatet al\.\(2022\)Y\. Ben\-Shabat, C\. H\. Koneputugodage, and S\. GouldDigs: Divergence guided shape implicit neural representation for unoriented point clouds\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 19323–19332\.Cited by:[§1\.3](https://arxiv.org/html/2609.30633#S1.SS3.SSS0.Px2.p1.1)\.
- Chickering \(2025\)K\. R\. ChickeringThe Spectral Maximal Update Parameterization in Theory and Practice\.Kyle Chickering’s Blog\.External Links:[Link](https://kyrochi.github.io/blog/spectral-mup/index.html)Cited by:[§A\.3](https://arxiv.org/html/2609.30633#A1.SS3.p4.1.1)\.
- Combetteet al\.\(2026\)A\. Combette, N\. Pustelnik, and A\. VenailleA New Initialization to Control Gradients in Sinusoidal Neural Networks\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=92d74WdgtG)Cited by:[§A\.1](https://arxiv.org/html/2609.30633#A1.SS1.p5.1),[§A\.3](https://arxiv.org/html/2609.30633#A1.SS3.p1.1),[§A\.3](https://arxiv.org/html/2609.30633#A1.SS3.p2.1),[§A\.3](https://arxiv.org/html/2609.30633#A1.SS3.p4.1.1),[§A\.4](https://arxiv.org/html/2609.30633#A1.SS4.p2.1),[Appendix A](https://arxiv.org/html/2609.30633#A1.p1.1),[Appendix C](https://arxiv.org/html/2609.30633#A3.SS0.SSS0.Px2.p2.1),[Appendix C](https://arxiv.org/html/2609.30633#A3.SS0.SSS0.Px3.p1.1),[Appendix C](https://arxiv.org/html/2609.30633#A3.SS0.SSS0.Px7.p1.1),[§1](https://arxiv.org/html/2609.30633#S1.SS0.SSS0.Px1.p3.1),[§1](https://arxiv.org/html/2609.30633#S1.SS0.SSS0.Px2.p1.1),[§1\.3](https://arxiv.org/html/2609.30633#S1.SS3.SSS0.Px2.p1.1),[§3](https://arxiv.org/html/2609.30633#S3.p3.1),[item SIREN \(CVP26 = σ a 0 \)](https://arxiv.org/html/2609.30633#S4.I1.ix2.p1.1),[§4\.1](https://arxiv.org/html/2609.30633#S4.SS1.p2.1),[§4\.4](https://arxiv.org/html/2609.30633#S4.SS4.p4.1),[Prior work](https://arxiv.org/html/2609.30633#Thmpriorworkx1),[Prior work](https://arxiv.org/html/2609.30633#Thmpriorworkx2),[Prior work](https://arxiv.org/html/2609.30633#Thmpriorworkx3)\.
- Deyet al\.\(2026\)N\. Dey, B\. C\. Zhang, L\. Noci, M\. Li, B\. Bordelon, S\. Bergsma, C\. Pehlevan, B\. Hanin, and J\. HestnessDon’t be lazy: CompleteP enables compute\-efficient deep transformers\.arXiv\.Note:arXiv:2505\.01618 \[cs\.LG\]External Links:[Link](http://arxiv.org/abs/2505.01618),[Document](https://dx.doi.org/10.48550/arXiv.2505.01618)Cited by:[§1\.3](https://arxiv.org/html/2609.30633#S1.SS3.SSS0.Px3.p1.1)\.
- Dupontet al\.\(2021\)E\. Dupont, A\. Goliński, M\. Alizadeh, Y\. W\. Teh, and A\. DoucetCOIN: COmpression with Implicit Neural representations\.arXiv\.Note:arXiv:2103\.03123 \[eess\.IV\]External Links:[Link](http://arxiv.org/abs/2103.03123),[Document](https://dx.doi.org/10.48550/arXiv.2103.03123)Cited by:[§1\.3](https://arxiv.org/html/2609.30633#S1.SS3.SSS0.Px1.p1.1)\.
- Glorot and Bengio \(2010\)X\. Glorot and Y\. BengioUnderstanding the difficulty of training deep feedforward neural networks\.InProceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics,Y\. W\. Teh and M\. Titterington \(Eds\.\),Proceedings of Machine Learning Research, Vol\.9,Chia Laguna Resort, Sardinia, Italy,pp\. 249–256\.External Links:[Link](https://proceedings.mlr.press/v9/glorot10a.html)Cited by:[item 1](https://arxiv.org/html/2609.30633#S1.I1.i1.p1.1),[item 3](https://arxiv.org/html/2609.30633#S1.I1.i3.p1.1),[§1](https://arxiv.org/html/2609.30633#S1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.30633#S1.p1.1)\.
- Hanin and Rolnick \(2018\)B\. Hanin and D\. RolnickHow to Start Training: The Effect of Initialization and Architecture\.arXiv\.Note:Version Number: 3External Links:[Link](https://arxiv.org/abs/1803.01719),[Document](https://dx.doi.org/10.48550/ARXIV.1803.01719)Cited by:[item 1](https://arxiv.org/html/2609.30633#S1.I1.i1.p1.1)\.
- Hastieet al\.\(2009\)T\. Hastie, R\. Tibshirani, and J\. FriedmanThe Elements of Statistical Learning\.Springer Series in Statistics,Springer,New York, NY\.External Links:ISBN 978\-0\-387\-84857\-0 978\-0\-387\-84858\-7,[Link](http://link.springer.com/10.1007/978-0-387-84858-7),[Document](https://dx.doi.org/10.1007/978-0-387-84858-7)Cited by:[§1](https://arxiv.org/html/2609.30633#S1.p1.1)\.
- Hayouet al\.\(2021\)S\. Hayou, E\. Clerico, B\. He, G\. Deligiannidis, A\. Doucet, and J\. RousseauStable ResNet\.InProceedings of The 24th International Conference on Artificial Intelligence and Statistics,pp\. 1324–1332\(en\)\.External Links:ISSN 2640\-3498,[Link](https://proceedings.mlr.press/v130/hayou21a.html)Cited by:[item 1](https://arxiv.org/html/2609.30633#S1.I1.i1.p1.1),[item 3](https://arxiv.org/html/2609.30633#S1.I1.i3.p1.1)\.
- Hayouet al\.\(2019\)S\. Hayou, A\. Doucet, and J\. RousseauOn the Impact of the Activation function on Deep Neural Networks Training\.InProceedings of the 36th International Conference on Machine Learning,pp\. 2672–2680\(en\)\.External Links:ISSN 2640\-3498,[Link](https://proceedings.mlr.press/v97/hayou19a.html)Cited by:[item 2](https://arxiv.org/html/2609.30633#S1.I1.i2.p1.1),[§1](https://arxiv.org/html/2609.30633#S1.SS0.SSS0.Px1.p1.1)\.
- Heet al\.\(2015\)K\. He, X\. Zhang, S\. Ren, and J\. SunDelving Deep into Rectifiers: Surpassing Human\-Level Performance on ImageNet Classification\.In2015 IEEE International Conference on Computer Vision \(ICCV\),pp\. 1026–1034\.External Links:ISSN 2380\-7504,[Link](https://ieeexplore.ieee.org/document/7410480/),[Document](https://dx.doi.org/10.1109/ICCV.2015.123)Cited by:[item 1](https://arxiv.org/html/2609.30633#S1.I1.i1.p1.1),[item 3](https://arxiv.org/html/2609.30633#S1.I1.i3.p1.1),[§1](https://arxiv.org/html/2609.30633#S1.SS0.SSS0.Px1.p1.1)\.
- Hestnesset al\.\(2017\)J\. Hestness, S\. Narang, N\. Ardalani, G\. Diamos, H\. Jun, H\. Kianinejad, M\. M\. A\. Patwary, Y\. Yang, and Y\. ZhouDeep Learning Scaling is Predictable, Empirically\.arXiv\.Note:arXiv:1712\.00409 \[cs\.LG\]External Links:[Link](http://arxiv.org/abs/1712.00409),[Document](https://dx.doi.org/10.48550/arXiv.1712.00409)Cited by:[§1](https://arxiv.org/html/2609.30633#S1.p1.1)\.
- Jacotet al\.\(2018\)A\. Jacot, F\. Gabriel, and C\. HonglerNeural Tangent Kernel: Convergence and Generalization in Neural Networks\.Note:Version Number: 4External Links:[Link](https://arxiv.org/abs/1806.07572),[Document](https://dx.doi.org/10.48550/ARXIV.1806.07572)Cited by:[item 4](https://arxiv.org/html/2609.30633#S1.I1.i4.p1.1)\.
- Klambaueret al\.\(2017\)G\. Klambauer, T\. Unterthiner, A\. Mayr, and S\. HochreiterSelf\-Normalizing Neural Networks\.InAdvances in Neural Information Processing Systems,Vol\.30\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2017/hash/5d44ee6f2c3f71b73125876103c8f6c4-Abstract.html)Cited by:[item 2](https://arxiv.org/html/2609.30633#S1.I1.i2.p1.1)\.
- LeCunet al\.\(2012\)Y\. A\. LeCun, L\. Bottou, G\. B\. Orr, and K\. MüllerEfficient BackProp\.InNeural Networks: Tricks of the Trade,G\. Montavon, G\. B\. Orr, and K\. Müller \(Eds\.\),Vol\.7700,pp\. 9–48\(en\)\.Note:Series Title: Lecture Notes in Computer ScienceExternal Links:ISBN 978\-3\-642\-35288\-1 978\-3\-642\-35289\-8,[Link](http://link.springer.com/10.1007/978-3-642-35289-8_3),[Document](https://dx.doi.org/10.1007/978-3-642-35289-8%5F3)Cited by:[item 1](https://arxiv.org/html/2609.30633#S1.I1.i1.p1.1),[item 3](https://arxiv.org/html/2609.30633#S1.I1.i3.p1.1),[§1](https://arxiv.org/html/2609.30633#S1.SS0.SSS0.Px1.p1.1)\.
- Leeet al\.\(2017\)J\. Lee, Y\. Bahri, R\. Novak, S\. S\. Schoenholz, J\. Pennington, and J\. Sohl\-DicksteinDeep Neural Networks as Gaussian Processes\.arXiv\.Note:Version Number: 3External Links:[Link](https://arxiv.org/abs/1711.00165),[Document](https://dx.doi.org/10.48550/ARXIV.1711.00165)Cited by:[item 4](https://arxiv.org/html/2609.30633#S1.I1.i4.p1.1)\.
- Mildenhallet al\.\(2020\)B\. Mildenhall, P\. P\. Srinivasan, M\. Tancik, J\. T\. Barron, R\. Ramamoorthi, and R\. NgNeRF: Representing Scenes as Neural Radiance Fields for View Synthesis\.arXiv\.Note:arXiv:2003\.08934 \[cs\.CV\]External Links:[Link](http://arxiv.org/abs/2003.08934),[Document](https://dx.doi.org/10.48550/arXiv.2003.08934)Cited by:[§1\.3](https://arxiv.org/html/2609.30633#S1.SS3.SSS0.Px1.p1.1)\.
- Mishkin and Matas \(2016\)D\. Mishkin and J\. MatasAll you need is a good init\.arXiv\.Note:arXiv:1511\.06422 \[cs\.LG\]External Links:[Link](http://arxiv.org/abs/1511.06422),[Document](https://dx.doi.org/10.48550/arXiv.1511.06422)Cited by:[item 3](https://arxiv.org/html/2609.30633#S1.I1.i3.p1.1)\.
- Novelloet al\.\(2025\)T\. Novello, D\. Aldana, A\. Araujo, and L\. VelhoTuning the Frequencies: Robust Training for Sinusoidal Neural Networks\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 3071–3080\.Cited by:[§1\.3](https://arxiv.org/html/2609.30633#S1.SS3.SSS0.Px2.p1.1)\.
- Parascandoloet al\.\(2017\)G\. Parascandolo, H\. Huttunen, and T\. VirtanenTaming the waves: sine as activation function in deep neural networks\.External Links:[Link](https://openreview.net/forum?id=Sks3zF9eg)Cited by:[§1\.3](https://arxiv.org/html/2609.30633#S1.SS3.SSS0.Px1.p1.1)\.
- Penningtonet al\.\(2017\)J\. Pennington, S\. Schoenholz, and S\. GanguliResurrecting the sigmoid in deep learning through dynamical isometry: theory and practice\.Advances in neural information processing systems30\.Cited by:[item 1](https://arxiv.org/html/2609.30633#S1.I1.i1.p1.1)\.
- Pennington and Worah \(2018\)J\. Pennington and P\. WorahThe Spectrum of the Fisher Information Matrix of a Single\-Hidden\-Layer Neural Network\.InAdvances in Neural Information Processing Systems,Vol\.31\.External Links:[Link](https://proceedings.neurips.cc/paper/2018/hash/18bb68e2b38e4a8ce7cf4f6b2625768c-Abstract.html)Cited by:[item 1](https://arxiv.org/html/2609.30633#S1.I1.i1.p1.1)\.
- Pooleet al\.\(2016\)B\. Poole, S\. Lahiri, M\. Raghu, J\. Sohl\-Dickstein, and S\. GanguliExponential expressivity in deep neural networks through transient chaos\.InAdvances in Neural Information Processing Systems,D\. Lee, M\. Sugiyama, U\. Luxburg, I\. Guyon, and R\. Garnett \(Eds\.\),Vol\.29\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2016/file/148510031349642de5ca0c544f31b2ef-Paper.pdf)Cited by:[item 2](https://arxiv.org/html/2609.30633#S1.I1.i2.p1.1),[item 4](https://arxiv.org/html/2609.30633#S1.I1.i4.p1.1),[§1](https://arxiv.org/html/2609.30633#S1.SS0.SSS0.Px1.p1.1)\.
- Robertset al\.\(2022\)D\. A\. Roberts, S\. Yaida, and B\. HaninThe Principles of Deep Learning Theory: An Effective Theory Approach to Understanding Neural Networks\.1 edition,Cambridge University Press\.External Links:ISBN 978\-1\-009\-02340\-5 978\-1\-316\-51933\-2,[Link](https://www.cambridge.org/core/product/identifier/9781009023405/type/book),[Document](https://dx.doi.org/10.1017/9781009023405)Cited by:[§1](https://arxiv.org/html/2609.30633#S1.SS0.SSS0.Px1.p2.2),[§1](https://arxiv.org/html/2609.30633#S1.SS0.SSS0.Px1.p3.1)\.
- Saxeet al\.\(2014\)A\. M\. Saxe, J\. L\. McClelland, and S\. GanguliExact solutions to the nonlinear dynamics of learning in deep linear neural networks\.arXiv\.Note:arXiv:1312\.6120 \[cs\.NE\]External Links:[Link](http://arxiv.org/abs/1312.6120),[Document](https://dx.doi.org/10.48550/arXiv.1312.6120)Cited by:[item 1](https://arxiv.org/html/2609.30633#S1.I1.i1.p1.1)\.
- Schoenholzet al\.\(2017\)S\. S\. Schoenholz, J\. Gilmer, S\. Ganguli, and J\. Sohl\-DicksteinDeep Information Propagation\.arXiv\.Note:arXiv:1611\.01232 \[stat\.ML\]External Links:[Link](http://arxiv.org/abs/1611.01232),[Document](https://dx.doi.org/10.48550/arXiv.1611.01232)Cited by:[item 2](https://arxiv.org/html/2609.30633#S1.I1.i2.p1.1),[item 4](https://arxiv.org/html/2609.30633#S1.I1.i4.p1.1),[§1](https://arxiv.org/html/2609.30633#S1.SS0.SSS0.Px1.p1.1)\.
- Sitzmannet al\.\(2020\)V\. Sitzmann, J\. Martel, A\. Bergman, D\. Lindell, and G\. WetzsteinImplicit Neural Representations with Periodic Activation Functions\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 7462–7473\.External Links:[Link](https://proceedings.neurips.cc/paper/2020/hash/53c04118df112c13a8c34b38343b9c10-Abstract.html)Cited by:[Appendix C](https://arxiv.org/html/2609.30633#A3.SS0.SSS0.Px2.p2.1),[§1](https://arxiv.org/html/2609.30633#S1.SS0.SSS0.Px2.p1.1),[§1\.3](https://arxiv.org/html/2609.30633#S1.SS3.SSS0.Px1.p1.1),[§1\.3](https://arxiv.org/html/2609.30633#S1.SS3.SSS0.Px2.p1.1),[§3](https://arxiv.org/html/2609.30633#S3.p3.1),[item SIREN \(SM20\)](https://arxiv.org/html/2609.30633#S4.I1.ix3.p1.1),[§4\.2](https://arxiv.org/html/2609.30633#S4.SS2.p1.1)\.
- Strümpleret al\.\(2022\)Y\. Strümpler, J\. Postels, R\. Yang, L\. V\. Gool, and F\. TombariImplicit Neural Representations for Image Compression\.InComputer Vision – ECCV 2022,S\. Avidan, G\. Brostow, M\. Cissé, G\. M\. Farinella, and T\. Hassner \(Eds\.\),Vol\.13686,pp\. 74–91\(en\)\.Note:Series Title: Lecture Notes in Computer ScienceExternal Links:ISBN 978\-3\-031\-19808\-3 978\-3\-031\-19809\-0,[Link](https://link.springer.com/10.1007/978-3-031-19809-0_5),[Document](https://dx.doi.org/10.1007/978-3-031-19809-0%5F5)Cited by:[§1\.3](https://arxiv.org/html/2609.30633#S1.SS3.SSS0.Px1.p1.1)\.
- Tanciket al\.\(2021\)M\. Tancik, B\. Mildenhall, T\. Wang, D\. Schmidt, P\. P\. Srinivasan, J\. T\. Barron, and R\. NgLearned Initializations for Optimizing Coordinate\-Based Neural Representations\.arXiv\.Note:arXiv:2012\.02189 \[cs\.CV\]External Links:[Link](http://arxiv.org/abs/2012.02189),[Document](https://dx.doi.org/10.48550/arXiv.2012.02189)Cited by:[§1](https://arxiv.org/html/2609.30633#S1.SS0.SSS0.Px2.p1.1)\.
- Tanciket al\.\(2020\)M\. Tancik, P\. P\. Srinivasan, B\. Mildenhall, S\. Fridovich\-Keil, N\. Raghavan, U\. Singhal, R\. Ramamoorthi, J\. T\. Barron, and R\. NgFourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains\.arXiv\.Note:arXiv:2006\.10739 \[cs\.CV\]External Links:[Link](http://arxiv.org/abs/2006.10739),[Document](https://dx.doi.org/10.48550/arXiv.2006.10739)Cited by:[§1\.3](https://arxiv.org/html/2609.30633#S1.SS3.SSS0.Px1.p1.1),[item Tanh \(PE\-Xavier\)](https://arxiv.org/html/2609.30633#S4.I1.ix4.p1.1)\.
- Wolinski and Arbel \(2022\)P\. Wolinski and J\. ArbelGaussian Pre\-Activations in Neural Networks: Myth or Reality?\.Note:Version Number: 4External Links:[Link](https://arxiv.org/abs/2205.12379),[Document](https://dx.doi.org/10.48550/ARXIV.2205.12379)Cited by:[§1](https://arxiv.org/html/2609.30633#S1.SS0.SSS0.Px1.p2.2)\.
- Xiaoet al\.\(2018\)L\. Xiao, Y\. Bahri, J\. Sohl\-Dickstein, S\. S\. Schoenholz, and J\. PenningtonDynamical Isometry and a Mean Field Theory of CNNs: How to Train 10,000\-Layer Vanilla Convolutional Neural Networks\.arXiv\.Note:Version Number: 2External Links:[Link](https://arxiv.org/abs/1806.05393),[Document](https://dx.doi.org/10.48550/ARXIV.1806.05393)Cited by:[item 1](https://arxiv.org/html/2609.30633#S1.I1.i1.p1.1)\.
- Yang and Schoenholz \(2017\)G\. Yang and S\. SchoenholzMean Field Residual Networks: On the Edge of Chaos\.InAdvances in Neural Information Processing Systems,I\. Guyon, U\. V\. Luxburg, S\. Bengio, H\. Wallach, R\. Fergus, S\. Vishwanathan, and R\. Garnett \(Eds\.\),Vol\.30\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2017/file/81c650caac28cdefce4de5ddc18befa0-Paper.pdf)Cited by:[item 4](https://arxiv.org/html/2609.30633#S1.I1.i4.p1.1),[§1](https://arxiv.org/html/2609.30633#S1.SS0.SSS0.Px1.p1.1)\.
- Yanget al\.\(2022\)G\. Yang, E\. J\. Hu, I\. Babuschkin, S\. Sidor, X\. Liu, D\. Farhi, N\. Ryder, J\. Pachocki, W\. Chen, and J\. GaoTensor Programs V: Tuning Large Neural Networks via Zero\-Shot Hyperparameter Transfer\.arXiv\.Note:arXiv:2203\.03466 \[cs\.LG\]External Links:[Link](http://arxiv.org/abs/2203.03466),[Document](https://dx.doi.org/10.48550/arXiv.2203.03466)Cited by:[item 1](https://arxiv.org/html/2609.30633#S1.I1.i1.p1.1),[§1\.3](https://arxiv.org/html/2609.30633#S1.SS3.SSS0.Px3.p1.1),[§4\.5](https://arxiv.org/html/2609.30633#S4.SS5.p1.1),[§4\.5](https://arxiv.org/html/2609.30633#S4.SS5.p1.2),[Remark 1](https://arxiv.org/html/2609.30633#Thmremark1.p1.1.1)\.
- Yanget al\.\(2024\)G\. Yang, J\. B\. Simon, and J\. BernsteinA Spectral Condition for Feature Learning\.arXiv\.Note:arXiv:2310\.17813 \[cs\.LG\]External Links:[Link](http://arxiv.org/abs/2310.17813),[Document](https://dx.doi.org/10.48550/arXiv.2310.17813)Cited by:[§A\.3](https://arxiv.org/html/2609.30633#A1.SS3.p4.1.1),[§1\.3](https://arxiv.org/html/2609.30633#S1.SS3.SSS0.Px3.p1.1)\.
- Yang \(2019\)G\. YangTensor Programs I: Wide Feedforward or Recurrent Neural Networks of Any Architecture are Gaussian Processes\.arXiv\.Note:Version Number: 3External Links:[Link](https://arxiv.org/abs/1910.12478),[Document](https://dx.doi.org/10.48550/ARXIV.1910.12478)Cited by:[item 4](https://arxiv.org/html/2609.30633#S1.I1.i4.p1.1)\.
- Yeomet al\.\(2025\)T\. Yeom, S\. Lee, and J\. LeeFast training of sinusoidal neural fields via scaling initialization\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 96229–96271\.Cited by:[§1](https://arxiv.org/html/2609.30633#S1.SS0.SSS0.Px2.p1.1),[§1\.3](https://arxiv.org/html/2609.30633#S1.SS3.SSS0.Px2.p1.1),[§1\.3](https://arxiv.org/html/2609.30633#S1.SS3.SSS0.Px3.p1.1),[§3](https://arxiv.org/html/2609.30633#S3.p3.1)\.
- Yuceet al\.\(2022\)G\. Yuce, G\. Ortiz\-Jimenez, B\. Besbinar, and P\. FrossardA Structured Dictionary Perspective on Implicit Neural Representations\.In2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),New Orleans, LA, USA,pp\. 19206–19216\(en\)\.External Links:ISBN 978\-1\-6654\-6946\-3,[Link](https://ieeexplore.ieee.org/document/9878645/),[Document](https://dx.doi.org/10.1109/CVPR52688.2022.01863)Cited by:[§1](https://arxiv.org/html/2609.30633#S1.SS0.SSS0.Px2.p1.1)\.

## Appendix AComparison to conventional stable initialization of hidden layers

We compare our hidden\-layer initialization to that of[Combette et al\. \(2026\)](https://arxiv.org/html/2609.30633#bib.bib5), which represents a state\-of\-the\-art implementation of the conventional wisdom for stable initialization\. Pursuant to the four\-step organization described in the Introduction, there are two significant derivations\.

### A\.1Preactivation distribution

The first major effort is to approximate preactivationszℓz\_\{\\ell\}as asymptotically Normal with a tunable varianceσa2\\sigma\_\{a\}^\{2\}:zℓ→𝒩⁡\(0,σa2\)\.z\_\{\\ell\}\\to\\mathcal\{N\}\(0,\\sigma\_\{a\}^\{2\}\)\.

###### Prior work\(Thm\. 3\.1,[Combette et al\. \(2026\)](https://arxiv.org/html/2609.30633#bib.bib5)\)\.

Consider the sinusoidal network defined above, where for somecw,cb∈ℝ\+c\_\{w\},c\_\{b\}\\in\\mathbb\{R\}^\{\+\}, and for every layerℓ∈\{2,…,L\}\\ell\\in\\\{2,\\ldots,L\\\}, the weight matrixWℓW\_\{\\ell\}is initialized with entries sampled from𝒰\(−cw/N,cw/N\)\\mathcal\{U\}\(\-c\_\{w\}/\\sqrt\{N\},\\,c\_\{w\}/\\sqrt\{N\}\),W1W\_\{1\}is sampled from𝒰\(−w0/n0,w0/n0\)\\mathcal\{U\}\(\-w\_\{0\}/n\_\{0\},\\,w\_\{0\}/n\_\{0\}\), and the biasbℓb\_\{\\ell\}is initialized with entries sampled from𝒩⁡\(0,cb2\)\\mathcal\{N\}\(0,c\_\{b\}^\{2\}\)\. Let\(zℓ\)ℓ∈\{1,…,L\}\(z\_\{\\ell\}\)\_\{\\ell\\in\\\{1,\\ldots,L\\\}\}be the preactivation sequencezℓ=Wℓ​xℓ−1\+bℓz\_\{\\ell\}=W\_\{\\ell\}x\_\{\\ell\-1\}\+b\_\{\\ell\}for an inputx∈ℝn0x\\in\\mathbb\{R\}^\{n\_\{0\}\}\. Then, in the limitsN,L→∞N,L\\to\\infty, the preactivation sequence\(zℓ\)ℓ∈ℕ\(z\_\{\\ell\}\)\_\{\\ell\\in\\mathbb\{N\}\}converges in distribution to𝒩⁡\(0,σa2\)\\mathcal\{N\}\(0,\\sigma\_\{a\}^\{2\}\)with

σa2=cb2\+cw26\+12​𝒲0​\(−cw23​e−cw23−2​cb2\),\\sigma\_\{a\}^\{2\}=c\_\{b\}^\{2\}\+\\frac\{c\_\{w\}^\{2\}\}\{6\}\+\\frac\{1\}\{2\}\\mathcal\{W\}\_\{0\}\\\!\\left\(\-\\frac\{c\_\{w\}^\{2\}\}\{3\}e^\{\-\\frac\{c\_\{w\}^\{2\}\}\{3\}\-2c\_\{b\}^\{2\}\}\\right\),\(13\)where𝒲0\\mathcal\{W\}\_\{0\}is the principal real branch of the LambertWWfunction\. The sequence\(Var⁡\(zℓ\)\)ℓ∈ℕ\\bigl\(\\mathrm\{Var\}\(z\_\{\\ell\}\)\\bigr\)\_\{\\ell\\in\\mathbb\{N\}\}converges to the fixed pointσa2\\sigma\_\{a\}^\{2\}, which is exponentially attractive for allcw≠3c\_\{w\}\\neq\\sqrt\{3\}\. Forcw=3c\_\{w\}=\\sqrt\{3\}, the convergence is of rate𝒪⁡\(1/ℓ\)\\mathcal\{O\}\(1/\\ell\)\.

The limitN→∞N\\to\\inftyis necessary in order to apply the Central Limit Theorem to establish that the preactivations converge weakly to an isotropic Normal distribution\. \(Because the sine activation and its derivatives are bounded and continuous, weak convergence is sufficient to reason about the network’s moments\.\) The limitL→∞L\\to\\inftyis necessary in order to wash out the transient effect of the embedding layer on the preactivation distribution\. An elegant nonlinear recursion is derived for each preactivation layer’s moments in terms of those of the previous layer; this recursion has an attractive fixed point given by equation[13](https://arxiv.org/html/2609.30633#A1.E13), reflecting a delicate balance between weights and biases\.

All of this follows from the need to enforce a Normality Ansatz\. While this Ansatz would make sense for nonperiodic activation functions, the parallel claim in our result is rendered nearly trivial by the fact that the sine function’s domain can, without loss of generality, be wrapped onto a torus\.

###### Our result\.

Letbℓb\_\{\\ell\}be uniformly distributed on theNN\-torus\. Letzℓ=Wℓ​xℓ−1\+bℓz\_\{\\ell\}=W\_\{\\ell\}x\_\{\\ell\-1\}\+b\_\{\\ell\}\. Thenzℓmod2​πz\_\{\\ell\}\\bmod 2\\piis uniformly distributed on theNN\-torus\.

Adding a uniform phase wipes out the distribution ofWℓ​xℓ−1W\_\{\\ell\}x\_\{\\ell\-1\}\. This eliminates the need for recursive analysis of the wrapped preactivation distribution because it is toroidally uniform by construction\. There is also no need to takeL→∞L\\to\\infty, because we have decoupled the layers and imposed an exactly uniform distribution on every layer’s wrapped preactivation\. Furthermore, there is no need to takeN→∞N\\to\\infty, because we do not rely on the Central Limit Theorem\. Our wrapped preactivation distribution is exact\.

Finally, we observe that[Combette et al\. \(2026, Thm\. 3\.1\)](https://arxiv.org/html/2609.30633#bib.bib5)is compatible with the key idea of our theory\. Keepingcwc\_\{w\}finite, take the limitcb2→∞c\_\{b\}^\{2\}\\to\\infty\. Using𝒲0​\(0\)=0\\mathcal\{W\}\_\{0\}\(0\)=0, equation[13](https://arxiv.org/html/2609.30633#A1.E13)reduces in this limit to

σa2=cb2=∞\.\\displaystyle\\sigma\_\{a\}^\{2\}=c\_\{b\}^\{2\}=\\infty\.\(14\)When a Normal distribution with infinite variance is wrapped onto the torus, the result is a uniform distribution, recovering our finding that the preactivation distribution can be decoupled from the weight scalecwc\_\{w\}\.

### A\.2Activation gradient distribution

Using the preactivation distribution, one can now estimate the quadratic mean of the activation function’s slope over the preactivation ensemble through the more general identity

Var​\[∂∂zℓ​sin⁡\(zℓ\)\]\\displaystyle\\mathrm\{Var\}\\mathinner\{\\left\[\\mathinner\{\\dfrac\{\\partial\{\}\}\{\\partial\{z\_\{\\ell\}\}\}\}\\sin\(z\_\{\\ell\}\)\\right\]\}=𝔼\[cos2⁡\(zℓ\)\]=1\+e−2​σa22elementwise\.\\displaystyle=\\expect\\mathinner\{\\left\[\\cos^\{2\}\(z\_\{\\ell\}\)\\right\]\}=\\frac\{1\+e^\{\-2\\sigma\_\{a\}^\{2\}\}\}\{2\}\\quad\\text\{elementwise\}\.\(15\)This principle may be used to analyze the layer Jacobian entrywise:

###### Prior work\(Thm\. 3\.2,[Combette et al\. \(2026\)](https://arxiv.org/html/2609.30633#bib.bib5)\)\.

Let∂xℓ/∂xℓ−1\\partial x\_\{\\ell\}/\\partial x\_\{\\ell\-1\}denote the Jacobian of theℓ\\ell\-th layer, so that

∂xℓ∂xℓ−1=diag⁡\(cos⁡\(zℓ\)\)​Wℓ\.\\mathinner\{\\dfrac\{\\partial\{\}x\_\{\\ell\}\}\{\\partial\{x\_\{\\ell\-1\}\}\}\}=\\mathrm\{diag\}\(\\cos\(z\_\{\\ell\}\)\)\\,W\_\{\\ell\}\.\(16\)Under the same assumptions as above, and maintaining the limit of largeNN, each entry of this Jacobian has zero mean and varianceσ~ℓ2\\widetilde\{\\sigma\}\_\{\\ell\}^\{2\}, such that the sequence\(N​σ~ℓ2\)ℓ∈ℕ\(N\\widetilde\{\\sigma\}\_\{\\ell\}^\{2\}\)\_\{\\ell\\in\\mathbb\{N\}\}converges to

limℓ,N→∞\(N​σ~ℓ2\)=σg=cw26​\(1\+e−2​σa2\)\.\\lim\_\{\\ell,N\\to\\infty\}\(N\\widetilde\{\\sigma\}\_\{\\ell\}^\{2\}\)=\\sigma\_\{g\}=\\frac\{c\_\{w\}^\{2\}\}\{6\}\\bigl\(1\+e^\{\-2\\sigma\_\{a\}^\{2\}\}\\bigr\)\.\(17\)

A parallel result can also be stated for our initialization:

###### Our result\.

Let the wrapped preactivationzℓmod2​πz\_\{\\ell\}\\bmod 2\\pibe uniformly distributed on theNN\-torus\. Assume that each entry ofWℓW\_\{\\ell\}has zero mean and variancecw2/\(3​N\)c\_\{w\}^\{2\}/\(3N\)\. Then each entry of the Jacobian

∂xℓ∂xℓ−1=diag⁡\(cos⁡\(zℓ\)\)​Wℓ\\displaystyle\\mathinner\{\\dfrac\{\\partial\{\}x\_\{\\ell\}\}\{\\partial\{x\_\{\\ell\-1\}\}\}\}=\\mathrm\{diag\}\(\\cos\(z\_\{\\ell\}\)\)W\_\{\\ell\}\(18\)has zero mean and variancecw2/\(6​N\)c\_\{w\}^\{2\}/\(6N\)\.

Note again the lack of asymptotics\. Activation\-slope vectors such ascos⁡\(zℓ\)\\cos\(z\_\{\\ell\}\)are i\.i\.d\.

### A\.3What gradient criterion?

Finally, we turn to the choice of criterion to stabilize \(Step 1 of the Introduction\)\. Whereas[Combette et al\. \(2026\)](https://arxiv.org/html/2609.30633#bib.bib5)stabilize the*entries*of the Jacobian, we stabilize its gramian matrix\. Translated into our convention, their result \(Appendix A\.4\) may be summarized as follows\.

###### Prior work\(Appendix A\.4,[Combette et al\. \(2026\)](https://arxiv.org/html/2609.30633#bib.bib5)\)\.

Letffbe a sinusoidal network initialized so that each entry of the layerwise Jacobian∂xℓ/∂xℓ−1\\partial x\_\{\\ell\}/\\partial x\_\{\\ell\-1\}has varianceσg2/N\\sigma\_\{g\}^\{2\}/N\. LetΨ\\Psibe a scalar loss function applied to the network output\. Then the parameter\-wise and input\-wise gradient variances scale as

Var​\[∇Wℓ,i,jΨ​\(f​\(x\)\)\]\\displaystyle\\mathrm\{Var\}\\mathinner\{\\left\[\\nabla\_\{W\_\{\\ell,i,j\}\}\\Psi\(f\(x\)\)\\right\]\}∼\(N​σg2\)L−ℓ−1N,\\displaystyle\\sim\\frac\{\(N\\sigma\_\{g\}^\{2\}\)^\{L\-\\ell\-1\}\}\{N\},ℓ\\displaystyle\\ell\>1,\\displaystyle\>1,\(19\)Var​\[∇xiΨ​\(f​\(x\)\)\]\\displaystyle\\mathrm\{Var\}\\mathinner\{\\left\[\\nabla\_\{x\_\{i\}\}\\Psi\(f\(x\)\)\\right\]\}∼\(σg2\)L−2\.\\displaystyle\\sim\(\\sigma\_\{g\}^\{2\}\)^\{L\-2\}\.\(20\)In particular, gradients vanish or explode exponentially with depth unlessN​σg2≈1N\\sigma\_\{g\}^\{2\}\\approx 1\.

This scaling analysis is a plausible—if not fully rigorous \(it applies an unjustified independence heuristic to random matrix\-vector products\)—corollary of the elementwise Jacobian result of[Combette et al\. \(2026, Thm\. 3\.2\)](https://arxiv.org/html/2609.30633#bib.bib5)\. However, it relies on the double limitℓ,N→∞\\ell,N\\to\\inftyand hides an unknown constant\.

Our initialization removes the asymptotic limits from the Jacobian second\-moment calculation\. Translating that calculation into a loss\-gradient prediction requires an additional assumption on the loss cotangent\.

###### Our result\.

Letffbe a sinusoidal network with the uniform\-phase initialization, let2≤ℓ≤L2\\leq\\ell\\leq L, and letg=∇xLΨ​\(f⁡\(x\)\)∈ℝNg=\\nabla\_\{x\_\{L\}\}\\Psi\(f\(x\)\)\\in\\mathbb\{R\}^\{N\}denote the loss cotangent at the output of the hidden map\. Then the parameter and input gradients decompose as

∇WℓΨ​\(f​\(x\)\)\\displaystyle\\nabla\_\{W\_\{\\ell\}\}\\Psi\(f\(x\)\)=Cos\(zℓ\)\(∂xL∂xℓ\)⊺gxℓ−1⊺and\\displaystyle=\\Cos\(z\_\{\\ell\}\)\\mathinner\{\\left\(\\mathinner\{\\dfrac\{\\partial\{\}x\_\{L\}\}\{\\partial\{x\_\{\\ell\}\}\}\}\\right\)\}^\{\\intercal\}gx\_\{\\ell\-1\}^\{\\intercal\}\\quad\\text\{and\}\(21\)∇x0Ψ​\(f​\(x\)\)\\displaystyle\\nabla\_\{x\_\{0\}\}\\Psi\(f\(x\)\)=\(∂xL∂x0\)⊺​g\.\\displaystyle=\\mathinner\{\\left\(\\mathinner\{\\dfrac\{\\partial\{\}x\_\{L\}\}\{\\partial\{x\_\{0\}\}\}\}\\right\)\}^\{\\intercal\}g\.\(22\)Using an independence heuristic,

𝔼⁡\(∇WℓΨ​\(f​\(x\)\)\)⊺​\(∇WℓΨ​\(f​\(x\)\)\)\\displaystyle\\expect\\mathinner\{\\left\(\\nabla\_\{W\_\{\\ell\}\}\\Psi\(f\(x\)\)\\right\)\}^\{\\intercal\}\\mathinner\{\\left\(\\nabla\_\{W\_\{\\ell\}\}\\Psi\(f\(x\)\)\\right\)\}≈14‖g‖2Iand\\displaystyle\\approx\\frac\{1\}\{4\}\\left\\\|g\\right\\\|^\{2\}I\\quad\\text\{and\}\(23\)𝔼⁡\(∇x0Ψ​\(f​\(x\)\)\)⊺​\(∇x0Ψ​\(f​\(x\)\)\)\\displaystyle\\expect\\mathinner\{\\left\(\\nabla\_\{x\_\{0\}\}\\Psi\(f\(x\)\)\\right\)\}^\{\\intercal\}\\mathinner\{\\left\(\\nabla\_\{x\_\{0\}\}\\Psi\(f\(x\)\)\\right\)\}≈‖g‖2\.\\displaystyle\\approx\\left\\\|g\\right\\\|^\{2\}\.\(24\)

###### Proof of equation[23](https://arxiv.org/html/2609.30633#A1.E23)\.

Most of the independence we need is rigorous \(Lemma[1](https://arxiv.org/html/2609.30633#Thmlemma1)\)\. The remaining approximation treatsggas approximately independent of the hidden\-layer initialization\([Chickering, 2025](https://arxiv.org/html/2609.30633#bib.bib12);[Yang et al\., 2024](https://arxiv.org/html/2609.30633#bib.bib11)\)\. As in[Combette et al\. \(2026\)](https://arxiv.org/html/2609.30633#bib.bib5), this is an early\-training heuristic: for a general loss, the cotangent depends on the initialized network\.

𝔼⁡\(∇WℓΨ​\(f​\(x\)\)\)⊺​\(∇WℓΨ​\(f​\(x\)\)\)\\displaystyle\\expect\\mathinner\{\\left\(\\nabla\_\{W\_\{\\ell\}\}\\Psi\(f\(x\)\)\\right\)\}^\{\\intercal\}\\mathinner\{\\left\(\\nabla\_\{W\_\{\\ell\}\}\\Psi\(f\(x\)\)\\right\)\}=𝔼⁡xℓ−1​g⊺​\(∂xL∂xℓ\)​Cos⁡\(zℓ\)​Cos⁡\(zℓ\)​\(∂xL∂xℓ\)⊺​gxℓ−1⊺\\displaystyle=\\expect x\_\{\\ell\-1\}g^\{\\intercal\}\\mathinner\{\\left\(\\mathinner\{\\dfrac\{\\partial\{\}x\_\{L\}\}\{\\partial\{x\_\{\\ell\}\}\}\}\\right\)\}\\Cos\(z\_\{\\ell\}\)\\Cos\(z\_\{\\ell\}\)\\mathinner\{\\left\(\\mathinner\{\\dfrac\{\\partial\{\}x\_\{L\}\}\{\\partial\{x\_\{\\ell\}\}\}\}\\right\)\}^\{\\intercal\}gx\_\{\\ell\-1\}^\{\\intercal\}≈12​‖g‖2​𝔼⁡xℓ−1​xℓ−1⊺\\displaystyle\\approx\\frac\{1\}\{2\}\\left\\\|g\\right\\\|^\{2\}\\expect x\_\{\\ell\-1\}x\_\{\\ell\-1\}^\{\\intercal\}≈12‖g‖2𝔼sin\(zℓ−1\)sin\(zℓ−1\)⊺\\displaystyle\\approx\\frac\{1\}\{2\}\\left\\\|g\\right\\\|^\{2\}\\expect\\sin\(z\_\{\\ell\-1\}\)\\sin\(z\_\{\\ell\-1\}\)^\{\\intercal\}≈14​‖g‖2​I\\displaystyle\\approx\\frac\{1\}\{4\}\\left\\\|g\\right\\\|^\{2\}I∎

### A\.4Spectral diagnostics

We complement the theory of Section[2](https://arxiv.org/html/2609.30633#S2)with mechanistic diagnostics that illuminate how the uniform\-phase initialization differs from its predecessors: the singular value spectrum of the end\-to\-end Jacobian and the visual appearance of random initialization fields\. Although Theorem[1](https://arxiv.org/html/2609.30633#Thmtheorem1)and CVP26’s asymptotic analysis make different predictions about the distribution of singular values, neither fully determines the spectral shape\. Therefore we inspect it empirically\.

We extend[Combette et al\. \(2026, Figure 8\)](https://arxiv.org/html/2609.30633#bib.bib5), which plots the singular values ofJ=∂f∂xJ=\\mathinner\{\\dfrac\{\\partial\{\}f\}\{\\partial\{x\}\}\}from greatest to least, by adding the uniform\-phase initialization\. To ensure a fair comparison, the input and output layers of our moment\-matched uniform\-phase initialization are calibrated to match the output statistics of a CVP26 \(σa=0\\sigma\_\{a\}=0,L=4L=4\) network, so that the only variable is the hidden\-layer initialization\.

The uniform\-phase initialization exhibits spectral broadening with depth: the largest singular values grow \(“rich get richer”\) while the smallest decay toward zero \(“poor get poorer”\)\. This is a way of satisfying Theorem[1](https://arxiv.org/html/2609.30633#Thmtheorem1), which fixes the*average*squared singular value at unity while leaving the distribution’s shape unconstrained\. If some singular values grow, others must shrink to balance the mean\.

This contrasts sharply with CVP26’sσa=0\\sigma\_\{a\}=0initialization, which instead preserves the*largest*singular value \(the operator norm\) across depth\. There, the “poor get poorer” phenomenon occurs for a different reason: the operator norm is held fixed while the average squared singular value decays, so the network’s mean amplification shrinks even as its peak amplification holds steady\. The two methods thus embody different design goals: CVP26σa=0\\sigma\_\{a\}\{=\}0stabilizes the worst case, while the uniform\-phase initialization stabilizes the average case\.

The mechanism is also very different\. CVP26 concentrates preactivations on the steep portion of the sine, amplifying slopes uniformly, whereas the uniform\-phase initialization averages over the full sine \(including its stationary and downward\-sloping regions\)\. The effect is most visible in comparison to SM20, which uses the same weight distribution as the uniform\-phase initialization but smaller biases, and consequently suffers gradient blowup\.

![Refer to caption](https://arxiv.org/html/2609.30633v1/assets/jacobian-spectrum.jpg)Figure 3:Full singular value spectrum of the end\-to\-end Jacobian𝐉=∂𝐡L/∂𝐡1\\mathbf\{J\}=\\partial\\mathbf\{h\}\_\{L\}/\\partial\\mathbf\{h\}\_\{1\}as a function of depth\. Four initialization schemes are compared: uniform\-phase MM \(moment\-matched to CVP26σa=0\\sigma\_\{a\}\{=\}0,L=4L\{=\}4\), CVP26σa=0\\sigma\_\{a\}\{=\}0, CVP26σa=1\\sigma\_\{a\}\{=\}1, and the original SM20 initialization\. Each spectrum is averaged over five independently initialized networks, with 10 sample points on\[−π,π\]\[\-\\pi,\\pi\]\.

## Appendix BSupplement to §[4\.5](https://arxiv.org/html/2609.30633#S4.SS5)

## Appendix CTasks and experimental methods

#### Common protocol and metrics\.

The fitting experiments represent a sampled field with a coordinate MLP and minimize mean squared error \(MSE\) over all samples and output dimensions\. Every optimizer step is full\-batch\. Model parameters and training arrays use 32\-bit floating\-point values, while moment estimates are accumulated in 64\-bit NumPy and then converted to model precision\. Adam is used without weight decay, gradient clipping, or a learning\-rate schedule\. For image and audio fitting we report

SNR=10​log10​Var⁡\(y\)MSE⁡\(y^,y\),\\operatorname\{SNR\}=10\\log\_\{10\}\\frac\{\\operatorname\{Var\}\(y\)\}\{\\operatorname\{MSE\}\(\\hat\{y\},y\)\},\(25\)where the variance, like the MSE, is computed over the complete evaluation array\. The high\-resolution image grid and the full audio signal are evaluation sets rather than independently sampled observations\. We present the image artifacts with matplotlib\-default lossy compression in order to reduce the size of this file, and certify that it does not materially affect their interpretation\.

#### Initialization implementations\.

For uniform\-phase MM, let a regularly sampled training field containmmvaluesyi∈ℝNouty\_\{i\}\\in\\mathbb\{R\}^\{N\_\{\\mathrm\{out\}\}\}and letJiJ\_\{i\}denote its spatial Jacobian\. We use the sample mean and covariance and form

Σ^ϵ=Σ^\+ϵ​I,Ω^=1m​∑i=1mJi⊺​Σ^ϵ−1​Ji\+ϵ​I,ϵ=10−6\.\\widehat\{\\Sigma\}\_\{\\epsilon\}=\\widehat\{\\Sigma\}\+\\epsilon I,\\qquad\\widehat\{\\Omega\}=\\frac\{1\}\{m\}\\sum\_\{i=1\}^\{m\}J\_\{i\}^\{\\intercal\}\\widehat\{\\Sigma\}\_\{\\epsilon\}^\{\-1\}J\_\{i\}\+\\epsilon I,\\qquad\\epsilon=10^\{\-6\}\.\(26\)EachJiJ\_\{i\}is estimated channel by channel withscipy\.ndimage\.sobel; the result is divided by the coordinate spacing and by2⋅4d−12\\cdot 4^\{d\-1\}, which accounts for the centered difference and smoothing in the otherd−1d\-1spatial directions\. The input weights have independent rows distributed as𝒩⁡\(0,Ω^/Nout\)\\mathcal\{N\}\(0,\\widehat\{\\Omega\}/N\_\{\\mathrm\{out\}\}\), the input phase and every hidden bias are independent𝒰⁡\(−π,π\)\\mathcal\{U\}\(\-\\pi,\\pi\)variables, and hidden weights are independent𝒩⁡\(0,2/N\)\\mathcal\{N\}\(0,2/N\)variables\. The output bias isμ^\\widehat\{\\mu\}and the columns of the output matrix are independent𝒩⁡\(0,2​Σ^ϵ/N\)\\mathcal\{N\}\(0,2\\widehat\{\\Sigma\}\_\{\\epsilon\}/N\)variables\.

The SM20 and CVP26 implementations follow the corresponding PyTorch releases\([Sitzmann et al\., 2020](https://arxiv.org/html/2609.30633#bib.bib2);[Combette et al\., 2026](https://arxiv.org/html/2609.30633#bib.bib5)\)but are written in JAX/Equinox\. Their first\-layer weights are sampled uniformly on\[−1/Nin,1/Nin\]\[\-1/N\_\{\\mathrm\{in\}\},1/N\_\{\\mathrm\{in\}\}\]and then multiplied by the reported first\-layerω0\\omega\_\{0\}inside the sine\. For subsequent sine layers,Wi​j∼𝒰\[−c/\(ωhN\),c/\(ωhN\)\]W\_\{ij\}\\sim\\mathcal\{U\}\[\-c/\(\\omega\_\{h\}\\sqrt\{N\}\),c/\(\\omega\_\{h\}\\sqrt\{N\}\)\]and the activation issin⁡\(ωh​\(W​x\+b\)\)\\sin\(\\omega\_\{h\}\(Wx\+b\)\)\. SM20 usesc=6c=\\sqrt\{6\}and PyTorch\-style uniform biases; CVP26σa=0\\sigma\_\{a\}=0usesc=3c=\\sqrt\{3\}and zero hidden biases; and CVP26σa=1\\sigma\_\{a\}=1usesc=6/\(1\+e−2\)c=\\sqrt\{6/\(1\+e^\{\-2\}\)\}and centered Normal hidden biases of variancec2​e−2/3c^\{2\}e^\{\-2\}/3whenωh=1\\omega\_\{h\}=1\. The final linear map has no bias and has weights uniform on\[−3/N/ωh,3/N/ωh\]\[\-\\sqrt\{3/N\}/\\omega\_\{h\},\\sqrt\{3/N\}/\\omega\_\{h\}\]\. The Tanh \(PE\-Xavier\) baseline encodesxxas\[sin⁡\(2​π​B​x\),cos⁡\(2​π​B​x\)\]\[\\sin\(2\\pi Bx\),\\cos\(2\\pi Bx\)\], usingN/2N/2rows ofBBsampled independently from𝒩⁡\(0,σ2​I\)\\mathcal\{N\}\(0,\\sigma^\{2\}I\), and applies a Xavier\-initialized Tanh MLP\. The ReLU baseline uses Kaiming\-uniform weights, while the SiLU and GeLU baselines use Xavier\-uniform weights; these MLPs use zero biases\. All baselines match the compared SIRENs in width and depth\.

#### Cameraman image fitting and narrow networks\.

We takeskimage\.data\.camera, downsample it from its native512×512512\\times 512resolution to128×128128\\times 128with anti\-aliasing, and associate pixels with an endpoint\-inclusive Cartesian grid on\[−1,1\]2\[\-1,1\]^\{2\}\. Training uses all16,38416\{,\}384low\-resolution coordinate–value pairs; evaluation uses all262,144262\{,\}144native\-resolution pairs\. Pixel intensities are represented in\[0,1\]\[0,1\]\. The standard experiment uses the architecture and hyperparameters of[Combette et al\. \(2026\)](https://arxiv.org/html/2609.30633#bib.bib5)’s supplementary material:L=11L=11,N=256N=256, 3000 Adam steps, and learning rate10−410^\{\-4\}; the narrow experiment changes only the width toN=16N=16\. The first\-layer frequency is64​π/664\\pi/\\sqrt\{6\}for SM20 and CVP26, and hidden frequencies are 1\. The positional\-encoding scale isσ=32\\sigma=32\.

#### Audio fitting\.

The bundledgt\_bach\.wavfile is loaded at its native 44\.1 kHz sampling rate, converted to mono, and normalized by its maximum absolute amplitude\. It contains 308,207 samples \(approximately 7 s\), which are placed on an endpoint\-inclusive grid on\[−100,100\]\[\-100,100\]\. The full signal is used both to optimize and to compute MSE and SNR; there is no held\-out or subsampled split\. All models useL=5L=5,N=256N=256, 5000 full\-batch Adam steps, and learning rate5×10−55\\times 10^\{\-5\}\. Both baseline SIRENs use first\-layerω0=30\\omega\_\{0\}=30; CVP26 uses hiddenωh=1\\omega\_\{h\}=1, whereas the SM20 replication usesωh=30\\omega\_\{h\}=30as in its reference implementation\. The Tanh positional\-encoding scale isσ=10\\sigma=10\.

#### Input\-frequency sweep\.

For each image and for each of SM20, CVP26σa=0\\sigma\_\{a\}=0, and CVP26σa=1\\sigma\_\{a\}=1, we evaluate

ω0=64​π6​10q,q∈\{−2,−1\.8,…,1\.8,2\}\.\\omega\_\{0\}=\\frac\{64\\pi\}\{\\sqrt\{6\}\}\\,10^\{q\},\\qquad q\\in\\\{\-2,\-1\.8,\\ldots,1\.8,2\\\}\.\(27\)A baseline uses the same random weights at all 21 frequencies, so frequency is the only within\-method variable; uniform\-phase MM is initialized and trained once per image\. All runs useL=11L=11,N=256N=256, and 3000 full\-batch Adam steps at learning rate10−410^\{\-4\}, with training and evaluation grids whose largest dimensions are 128 and 512, respectively, and whose aspect ratio is preserved\. The grayscale tasks usebrick,camera,gravel, andclock; the color tasks usehubble\_deep\_field,skin,astronaut, andcoffeefromscikit\-image\. RGB values are scaled to\[0,1\]\[0,1\]\. For the CIELAB variant, images are resized in RGB and then transformed withskimage\.color\.rgb2lab; training and MSE evaluation take place in CIELAB, and conversion back to RGB is used only for display\.

#### Width and learning\-rate sweep\.

Theμ\\muP experiment uses the grayscale Cameraman training field and reports no test loss\. First, at base widthN0=64N\_\{0\}=64, each prior SIREN initialization is trained at the 21 frequencies above using a learning rate of10−410^\{\-4\}; its frequency is selected by the mean of the 3000 pre\-update training losses\. Uniform\-phase MM requires no such selection\. Second, forN∈\{8,16,32,64,128,256,512\}N\\in\\\{8,16,32,64,128,256,512\\\}, each initialization is trained at eachη∈\{2−18,2−17,…,22\}\\eta\\in\\\{2^\{\-18\},2^\{\-17\},\\ldots,2^\{2\}\\\}for 3000 steps\. The plotted criterion is again the time average of the per\-step pre\-update training MSE, rather than final MSE\. All learning rates for a given initialization and width begin from identical parameters\. In this experiment Adam usesϵ=0\\epsilon=0so that its additive numerical constant does not introduce a width\-dependent scale\. Underμ\\muP \(LR\), hidden matrices and the output projection use an effective learning rate ofη​N0/N\\eta N\_\{0\}/N, while input projections and all biases useη\\eta\. Underμ\\muP \(LR\+LL\), the same optimizer scaling is used and the initialized output weights are additionally multiplied byN0/N\\sqrt\{N\_\{0\}/N\}\. Thestandardcondition usesη\\etafor every parameter and leaves initialization unchanged\.

#### Jacobian\-spectrum diagnostic\.

Following the protocol of[Combette et al\. \(2026, Appendix B\.1\)](https://arxiv.org/html/2609.30633#bib.bib5), we study the hidden\-to\-hidden JacobianJ=∂hL/∂h1J=\\partial h\_\{L\}/\\partial h\_\{1\}for widthN=256N=256and depthsL∈\{4,8,16,32\}L\\in\\\{4,8,16,32\\\}\. The uniform\-phase input and output moments are estimated from evaluations of a CVP26σa=0\\sigma\_\{a\}=0,L=4L=4network with first\-layerω0=30\\omega\_\{0\}=30and hiddenωh=1\\omega\_\{h\}=1at 1000 evenly spaced points on\[−π,π\]\[\-\\pi,\\pi\]\. For each initialization and depth, we compute sorted singular values at 10 uniformly sampled input points for each of five independently initialized networks and average each ordered singular value over the resulting 50 spectra\.

## Appendix DAdditional experimental results

Table 1:Quantitative results for the Cameraman image\-fitting experiment \(§[4\.1](https://arxiv.org/html/2609.30633#S4.SS1)\)\. Train MSE is computed on the128×128128\\times 128grid; eval MSE and SNR are computed on the512×512512\\times 512grid\. Best values are in bold\.Table 2:Quantitative results for the narrow \(N=16N=16\) Cameraman experiment \(§[4\.1](https://arxiv.org/html/2609.30633#S4.SS1)\)\.![Refer to caption](https://arxiv.org/html/2609.30633v1/cameraman_narrow.png)Figure 4:Image fitting with narrow networks \(N=16N=16\)\. Layout as in figure[1](https://arxiv.org/html/2609.30633#S4.F1)\.Figure 5:Audio fitting on the Bach recording\. One row per initialization, with the ground truth in gray and the prediction in blue\. Left: the full waveform, drawn as a per\-pixel minimum/maximum envelope of all308,207308\{,\}207samples\. Right: the first10001000samples, on its own vertical scale because the recording opens in near\-silence\. The evaluation SNR is given in each right\-hand panel\.Figure 6:μ\\muP learning\-rate transfer, grayscale Cameraman, geometrically time\-averaged train MSE \(§[4\.5](https://arxiv.org/html/2609.30633#S4.SS5)\)\. Rows: initialization \(uniform\-phase MM and the three SIREN baselines at their bestω0\\omega\_\{0\}\); columns:μ\\muP LR,μ\\muP LR\+LL, and standard scaling\. One curve per width, shaded from light \(N=8N=8\) to dark \(N=512N=512\), with a star at each width’s best learning rate\.
## Appendix EMoment\-matched parameters and best\-ω0\\omega\_\{0\}test errors

#### Grayscale

The estimated moment\-matched parameters are:

μ^\\displaystyle\\hat\{\\mu\}=0\.437\\displaystyle=0\.437Σ^\\displaystyle\\hat\{\\Sigma\}=0\.00672\\displaystyle=0\.00672Ω^\\displaystyle\\hat\{\\Omega\}=\(561−4\.2−4\.21870\)\\displaystyle=\\begin\{pmatrix\}561&\-4\.2\\\\ \-4\.2&1870\\end\{pmatrix\}\(Brick\)μ^\\displaystyle\\hat\{\\mu\}=0\.506\\displaystyle=0\.506Σ^\\displaystyle\\hat\{\\Sigma\}=0\.0793\\displaystyle=0\.0793Ω^\\displaystyle\\hat\{\\Omega\}=\(87\.010\.910\.9115\)\\displaystyle=\\begin\{pmatrix\}87\.0&10\.9\\\\ 10\.9&115\\end\{pmatrix\}\(Camera\)μ^\\displaystyle\\hat\{\\mu\}=0\.496\\displaystyle=0\.496Σ^\\displaystyle\\hat\{\\Sigma\}=0\.0125\\displaystyle=0\.0125Ω^\\displaystyle\\hat\{\\Omega\}=\(1204−48\.6−48\.61120\)\\displaystyle=\\begin\{pmatrix\}1204&\-48\.6\\\\ \-48\.6&1120\\end\{pmatrix\}\(Gravel\)μ^\\displaystyle\\hat\{\\mu\}=0\.574\\displaystyle=0\.574Σ^\\displaystyle\\hat\{\\Sigma\}=0\.00664\\displaystyle=0\.00664Ω^\\displaystyle\\hat\{\\Omega\}=\(56\.50\.530\.5330\.7\)\\displaystyle=\\begin\{pmatrix\}56\.5&0\.53\\\\ 0\.53&30\.7\\end\{pmatrix\}\(Clock\)
Table 3:Grayscale sweep: best test MSE per method\. The optimalω0\\omega\_\{0\}from the grid search is shown in parentheses\. Bold entries are the best per row\. Visual comparisons: figure[7](https://arxiv.org/html/2609.30633#A6.F7)\.
#### Color \(RGB\)

The estimated moment\-matched parameters are:

μ^\\displaystyle\\hat\{\\mu\}=\(0\.0730\.0780\.075\)\\displaystyle=\\begin\{pmatrix\}0\.073\\\\ 0\.078\\\\ 0\.075\\end\{pmatrix\}Σ^\\displaystyle\\hat\{\\Sigma\}=\(0\.00680\.00540\.00560\.00540\.00470\.00500\.00560\.00500\.0054\)\\displaystyle=\\begin\{pmatrix\}0\.0068&0\.0054&0\.0056\\\\ 0\.0054&0\.0047&0\.0050\\\\ 0\.0056&0\.0050&0\.0054\\end\{pmatrix\}Ω^\\displaystyle\\hat\{\\Omega\}=\(2323−14\.8−14\.83010\)\\displaystyle=\\begin\{pmatrix\}2323&\-14\.8\\\\ \-14\.8&3010\\end\{pmatrix\}\(Hubble\)μ^\\displaystyle\\hat\{\\mu\}=\(0\.8090\.6600\.743\)\\displaystyle=\\begin\{pmatrix\}0\.809\\\\ 0\.660\\\\ 0\.743\\end\{pmatrix\}Σ^\\displaystyle\\hat\{\\Sigma\}=\(0\.00310\.00520\.00300\.00520\.01650\.00950\.00300\.00950\.0059\)\\displaystyle=\\begin\{pmatrix\}0\.0031&0\.0052&0\.0030\\\\ 0\.0052&0\.0165&0\.0095\\\\ 0\.0030&0\.0095&0\.0059\\end\{pmatrix\}Ω^\\displaystyle\\hat\{\\Omega\}=\(69923\.923\.9888\)\\displaystyle=\\begin\{pmatrix\}699&23\.9\\\\ 23\.9&888\\end\{pmatrix\}\(Skin\)μ^\\displaystyle\\hat\{\\mu\}=\(0\.5550\.4150\.378\)\\displaystyle=\\begin\{pmatrix\}0\.555\\\\ 0\.415\\\\ 0\.378\\end\{pmatrix\}Σ^\\displaystyle\\hat\{\\Sigma\}=\(0\.0970\.0750\.0660\.0750\.0840\.0820\.0660\.0820\.086\)\\displaystyle=\\begin\{pmatrix\}0\.097&0\.075&0\.066\\\\ 0\.075&0\.084&0\.082\\\\ 0\.066&0\.082&0\.086\\end\{pmatrix\}Ω^\\displaystyle\\hat\{\\Omega\}=\(551−56\.1−56\.1804\)\\displaystyle=\\begin\{pmatrix\}551&\-56\.1\\\\ \-56\.1&804\\end\{pmatrix\}\(Astronaut\)μ^\\displaystyle\\hat\{\\mu\}=\(0\.6220\.3360\.202\)\\displaystyle=\\begin\{pmatrix\}0\.622\\\\ 0\.336\\\\ 0\.202\\end\{pmatrix\}Σ^\\displaystyle\\hat\{\\Sigma\}=\(0\.0570\.0460\.0320\.0460\.0520\.0420\.0320\.0420\.038\)\\displaystyle=\\begin\{pmatrix\}0\.057&0\.046&0\.032\\\\ 0\.046&0\.052&0\.042\\\\ 0\.032&0\.042&0\.038\\end\{pmatrix\}Ω^\\displaystyle\\hat\{\\Omega\}=\(40010\.110\.1630\)\\displaystyle=\\begin\{pmatrix\}400&10\.1\\\\ 10\.1&630\\end\{pmatrix\}\(Coffee\)
Table 4:Color sweep: best test MSE per method\. Conventions as in Table[3](https://arxiv.org/html/2609.30633#A5.T3)\. Visual comparisons: figure[8](https://arxiv.org/html/2609.30633#A6.F8)\.
#### Color \(CIELAB\)

The estimated moment\-matched parameters in CIELAB space are:

μ^\\displaystyle\\hat\{\\mu\}=\(6\.47−0\.350\.16\)\\displaystyle=\\begin\{pmatrix\}6\.47\\\\ \-0\.35\\\\ 0\.16\\end\{pmatrix\}Σ^\\displaystyle\\hat\{\\Sigma\}=\(64\.412\.3−0\.2112\.37\.552\.56−0\.212\.565\.50\)\\displaystyle=\\begin\{pmatrix\}64\.4&12\.3&\-0\.21\\\\ 12\.3&7\.55&2\.56\\\\ \-0\.21&2\.56&5\.50\\end\{pmatrix\}Ω^\\displaystyle\\hat\{\\Omega\}=\(2328−9\.7−9\.73023\)\\displaystyle=\\begin\{pmatrix\}2328&\-9\.7\\\\ \-9\.7&3023\\end\{pmatrix\}\(Hubble\)μ^\\displaystyle\\hat\{\\mu\}=\(72\.9917\.52−5\.39\)\\displaystyle=\\begin\{pmatrix\}72\.99\\\\ 17\.52\\\\ \-5\.39\\end\{pmatrix\}Σ^\\displaystyle\\hat\{\\Sigma\}=\(90\.97−99\.6736\.77−99\.67130\.15−39\.3836\.77−39\.3824\.02\)\\displaystyle=\\begin\{pmatrix\}90\.97&\-99\.67&36\.77\\\\ \-99\.67&130\.15&\-39\.38\\\\ 36\.77&\-39\.38&24\.02\\end\{pmatrix\}Ω^\\displaystyle\\hat\{\\Omega\}=\(72827\.027\.0936\)\\displaystyle=\\begin\{pmatrix\}728&27\.0\\\\ 27\.0&936\\end\{pmatrix\}\(Skin\)μ^\\displaystyle\\hat\{\\mu\}=\(47\.7313\.5511\.95\)\\displaystyle=\\begin\{pmatrix\}47\.73\\\\ 13\.55\\\\ 11\.95\\end\{pmatrix\}Σ^\\displaystyle\\hat\{\\Sigma\}=\(843−0\.1153\.5−0\.1130223553\.5235317\)\\displaystyle=\\begin\{pmatrix\}843&\-0\.11&53\.5\\\\ \-0\.11&302&235\\\\ 53\.5&235&317\\end\{pmatrix\}Ω^\\displaystyle\\hat\{\\Omega\}=\(571−54\.0−54\.0822\)\\displaystyle=\\begin\{pmatrix\}571&\-54\.0\\\\ \-54\.0&822\\end\{pmatrix\}\(Astronaut\)μ^\\displaystyle\\hat\{\\mu\}=\(44\.3926\.6032\.84\)\\displaystyle=\\begin\{pmatrix\}44\.39\\\\ 26\.60\\\\ 32\.84\\end\{pmatrix\}Σ^\\displaystyle\\hat\{\\Sigma\}=\(495−36\.95126\.4−36\.95189145\.3126\.4145\.3203\.3\)\\displaystyle=\\begin\{pmatrix\}495&\-36\.95&126\.4\\\\ \-36\.95&189&145\.3\\\\ 126\.4&145\.3&203\.3\\end\{pmatrix\}Ω^\\displaystyle\\hat\{\\Omega\}=\(4188\.588\.58646\)\\displaystyle=\\begin\{pmatrix\}418&8\.58\\\\ 8\.58&646\\end\{pmatrix\}\(Coffee\)Compared with the RGB parameters above, the CIELAB output covarianceΣ^\\hat\{\\Sigma\}has much larger entries \(reflecting the wider dynamic range of CIELAB coordinates\) and mixed\-sign off\-diagonal entries \(reflecting the decorrelating effect of the opponency axes\)\. The structure tensorΩ^\\hat\{\\Omega\}is nearly unchanged, confirming that the spatial\-frequency content is preserved under the color\-space transformation\.

Table 5:CIELAB sweep: best test MSE per method\. MSE is computed in CIELAB coordinates\. Conventions as in Table[3](https://arxiv.org/html/2609.30633#A5.T3)\. Visual comparisons: figure[9](https://arxiv.org/html/2609.30633#A6.F9)\.

## Appendix FBest\-ω0\\omega\_\{0\}visual comparisons

Each figure in this section covers one experiment: rows are the four images, and columns are the ground truth, the uniform\-phase MM fit, and the best\-tuned \(byω0\\omega\_\{0\}\) fits of SM20, CVP26σa=0\\sigma\_\{a\}\{=\}0, and CVP26σa=1\\sigma\_\{a\}\{=\}1from the sweep of Appendix[E](https://arxiv.org/html/2609.30633#A5)\. Each panel is annotated with its test MSE and, for the baselines, theω0\\omega\_\{0\}that achieved it\.

![Refer to caption](https://arxiv.org/html/2609.30633v1/omega_sweep_best_grayscale.png)Figure 7:Grayscale best\-ω0\\omega\_\{0\}visual comparisons: Brick, Camera, Gravel, and Clock\.![Refer to caption](https://arxiv.org/html/2609.30633v1/omega_sweep_best_rgb.png)Figure 8:RGB best\-ω0\\omega\_\{0\}visual comparisons: Hubble, Skin, Astronaut, and Coffee\. Same layout as figure[7](https://arxiv.org/html/2609.30633#A6.F7)\.![Refer to caption](https://arxiv.org/html/2609.30633v1/omega_sweep_best_cielab.png)Figure 9:CIELAB best\-ω0\\omega\_\{0\}visual comparisons: Hubble, Skin, Astronaut, and Coffee\. Predictions are converted from CIELAB to RGB for display; the reported MSE values are computed in CIELAB coordinates\. Same layout as figure[7](https://arxiv.org/html/2609.30633#A6.F7)\.

Similar Articles

Small Initialization Matters for Large Language Models

arXiv cs.AI

This paper shows that reducing parameter initialization scale consistently improves pretraining of large language models, with the largest gains on reasoning-demanding tasks. It uncovers a critical initialization that balances reasoning and training, and proposes a simple γ-initialization rule.