A law of robustness for two-layer neural networks with arbitrary weights

arXiv cs.LG Papers

Summary

This paper proves a conjectured law of robustness for two-layer neural networks with unbounded weights, showing that a network fitting noisy data must have a Lipschitz constant at least of order sqrt(n/m), up to a logarithmic factor, for continuous piecewise-linear activations like ReLU.

arXiv:2607.07778v1 Announce Type: new Abstract: Bubeck, Li and Nagaraj conjectured that, for generic data, any two-layer neural network with $m$ neurons that fits $n$ noisy labels must have Lipschitz constant at least of order $\sqrt{n/m}$, with no restriction on the size of the weights. Bubeck and Sellke proved a universal version of this law for Lipschitz-parameterized classes, but under a polynomial bound on the parameters; at depth three that boundedness hypothesis is genuinely necessary. The two-layer unbounded-weight case requires a different argument. We prove the conjectured law, up to one logarithmic factor, for every continuous piecewise-linear activation, in particular for ReLU networks. For data drawn uniformly from $\mathbb{S}^{d-1}$, $d\ge3$, or from $N(0,I_d/d)$, labels in $[-1,1]$ with noise level $\sigma^2>0$, and any width-$m$ two-layer network with arbitrary real weights, biases and affine skip connection, fitting the data $\varepsilon$ below the noise floor forces $\mathrm{Lip}(f)\ge c\,\varepsilon\sqrt{n/(\bar m\log(C\bar m nd/\varepsilon))}$, $\bar m=(K-1)m+1$, with high probability. A realized-kink-count version holds on the same event: every realized two-layer piecewise-linear function with $k(f)\le n$ distinct kink hyperplanes obeys the bound with $\bar m$ replaced by $k(f)+1$, irrespective of how many redundant hidden units parameterize it. The proof replaces parameter-space covering, impossible for unbounded weights, by a function-space covering. The central deterministic ingredient is a rigidity lemma: on $B_2$, and on $\mathbb{S}^{d-1}$ for $d\ge3$, the coefficient of each canonical kink is controlled by the Lipschitz constant of the realized function, because kinks on distinct hyperplanes cannot cancel at generic points. Rigidity genuinely fails at $d=2$, and an explicit two-layer ReLU interpolant with $O(1)$ Lipschitz constant at width $2n$ matches the law at the overparameterized endpoint.
Original Article
View Cached Full Text

Cached at: 07/10/26, 06:15 AM

# A law of robustness for two-layer neural networks with arbitrary weights
Source: [https://arxiv.org/html/2607.07778](https://arxiv.org/html/2607.07778)
Yitzchak ShmaloEinstein Institute of Mathematics, The Hebrew University of Jerusalem, Givat Ram, Jerusalem, Israel[yitzchak\.shmalo@gmail\.com](https://arxiv.org/html/2607.07778v1/mailto:[email protected])

\(Date: July 7, 2026\)

###### Abstract\.

Bubeck, Li and Nagaraj conjectured that, for generic data, any two\-layer neural network withmmneurons that fitsnnnoisy labels must have Lipschitz constant at least of ordern/m\\sqrt\{n/m\}, with no restriction on the size of the weights\. Bubeck and Sellke proved a universal version of this law for Lipschitz\-parameterized classes, but under a polynomial bound on the parameters; at depth three that boundedness hypothesis is genuinely necessary\. The two\-layer unbounded\-weight case therefore requires a different argument\.

We prove the conjectured law, up to one logarithmic factor, for every continuous piecewise\-linear activation, in particular for ReLU networks\. For data drawn either uniformly from𝕊d−1\\mathbb\{S\}^\{d\-1\},d≥3d\\geq 3, or fromN​\(0,Id/d\)N\(0,I\_\{d\}/d\), labels in\[−1,1\]\[\-1,1\]with conditional noise levelσ2\>0\\sigma^\{2\}\>0, and any fixed width\-mmtwo\-layer network with arbitrary real weights, biases and affine skip connection, fitting the dataε\\varepsilonbelow the noise floor forces

Lip⁡\(f\)≥c​ε​nm¯​log⁡\(C​m¯​n​d/ε\),m¯=\(K−1\)​m\+1,\\operatorname\{Lip\}\(f\)\\geq c\\,\\varepsilon\\sqrt\{\\frac\{n\}\{\\bar\{m\}\\log\(C\\bar\{m\}nd/\\varepsilon\)\}\},\\qquad\\bar\{m\}=\(K\-1\)m\+1,with high probability\. We also prove a finite\-horizon simultaneous\-width version and a realized\-kink\-count version: on one high\-probability event, every realized two\-layer piecewise\-linear function withk​\(f\)≤nk\(f\)\\leq ndistinct kink hyperplanes obeys the same bound withm¯\\bar\{m\}replaced byk​\(f\)\+1k\(f\)\+1, irrespective of how many redundant hidden units were used to parameterize it\.

The proof replaces parameter\-space covering, which is impossible for unbounded weights, by a function\-space covering\. The central deterministic ingredient is a rigidity lemma: onB2B\_\{2\}, and on𝕊d−1\\mathbb\{S\}^\{d\-1\}ford≥3d\\geq 3, the coefficient of each canonical kink is controlled by the Lipschitz constant of the realized function, because kinks supported on distinct hyperplanes cannot cancel at generic points\. This yields a bounded canonical representation and hence the required entropy bound\. We also show why the sphere argument genuinely excludesd=2d=2, give a two\-layer ReLU interpolant withO​\(1\)O\(1\)Lipschitz constant at width2​n2nin the high\-dimensional separated regime, and state the precise concentration/localization hypotheses under which the Gaussian proof extends beyond the Gaussian measure\.

###### Key words and phrases:

law of robustness, two\-layer neural networks, ReLU networks, arbitrary weights, Lipschitz interpolation, metric entropy, isoperimetry

###### 2020 Mathematics Subject Classification:

Primary 68T07, 68Q32; Secondary 60F10, 60B20, 60B15

## 1\.Introduction

A function that fitsnnnoisy labels and is to be robust — small Lipschitz constant — needs capacity\. Bubeck, Li and Nagaraj\[[1](https://arxiv.org/html/2607.07778#bib.bib1)\]made this precise for the basic architecture of the subject\. Let

𝒩m=\{f​\(x\)=∑k=1mak​ψ​\(⟨wk,x⟩\+bk\)\+⟨v,x⟩\+c:ak,bk,c∈ℝ,wk,v∈ℝd\}\\mathcal\{N\}\_\{m\}\\;=\\;\\Bigl\\\{\\,f\(x\)=\\sum\_\{k=1\}^\{m\}a\_\{k\}\\,\\psi\(\\langle w\_\{k\},x\\rangle\+b\_\{k\}\)\+\\langle v,x\\rangle\+c\\;:\\;a\_\{k\},b\_\{k\},c\\in\\mathbb\{R\},\\ w\_\{k\},v\\in\\mathbb\{R\}^\{d\}\\,\\Bigr\\\}\(1\)be the class of two\-layer networks of widthmmwith activationψ\\psi, with*no restriction whatsoever*on the magnitudes of the weights\. The affine part⟨v,x⟩\+c\\langle v,x\\rangle\+conly enlarges the class studied in\[[1](https://arxiv.org/html/2607.07778#bib.bib1)\]; all results below hold a fortiori without it\.

###### Conjecture 1\.1\(Bubeck–Li–Nagaraj\[[1](https://arxiv.org/html/2607.07778#bib.bib1)\], Conjecture 1\)\.

Letψ\\psibe any Lipschitz activation\. Forx1,…,xnx\_\{1\},\\dots,x\_\{n\}independent uniform on𝕊d−1\\mathbb\{S\}^\{d\-1\}\(orN​\(0,Id/d\)N\(0,I\_\{d\}/d\)\) andy1,…,yny\_\{1\},\\dots,y\_\{n\}independent uniform on\{−1,\+1\}\\\{\-1,\+1\\\}, with high probability, anyf∈𝒩mf\\in\\mathcal\{N\}\_\{m\}fitting the data must satisfy

Lip𝕊d−1⁡\(f\)≥c​n/m\.\\operatorname\{Lip\}\_\{\\mathbb\{S\}^\{d\-1\}\}\(f\)\\;\\geq\\;c\\,\\sqrt\{n/m\}\.

The interpretation is that robust interpolation should require on the order of one neuron per data point, while non\-robust interpolation can require far fewer neurons in high dimension\. Bubeck and Sellke\[[2](https://arxiv.org/html/2607.07778#bib.bib2)\]proved a far\-reaching generalization: for any function class admitting a Lipschitz parameterization byppreal parameters of polynomial size, and for covariate distributions satisfying isoperimetry, fitting below the noise floor forcesLip⁡\(f\)≳ε​n​d/p\\operatorname\{Lip\}\(f\)\\gtrsim\\varepsilon\\sqrt\{nd/p\}up to logarithmic factors\. For width\-mmtwo\-layer networks,p=Θ​\(m​d\)p=\\Theta\(md\), giving the desiredn/m\\sqrt\{n/m\}scaling under the polynomial\-weight hypothesis\. That hypothesis is not merely technical at larger depth: Bubeck and Sellke construct three\-layer unbounded\-weight networks that violate the law\. Wu, Huang and Zhang\[[3](https://arxiv.org/html/2607.07778#bib.bib3)\]subsequently extended robustness laws beyond isoperimetric data, under polynomially bounded parameters\. The unbounded\-weight two\-layer ReLU case posed by Conjecture[1\.1](https://arxiv.org/html/2607.07778#S1.Thmtheorem1)remains the natural boundary case\.

This paper proves the law for every continuous piecewise\-linear activation, in particular for ReLU networks, up to one logarithmic factor\. The price of the logarithm is explicit throughout; we do not claim the log\-free lower bound\. Section[7](https://arxiv.org/html/2607.07778#S7), in the sphere model, sharpens the logarithm itself: in the regimen≳m​d2​log⁡\(m​d\)n\\gtrsim md^\{2\}\\log\(md\)the factorlog⁡\(C​m​n​d\)\\log\(Cmnd\)improves tolog⁡\(C​m​d\)\\log\(Cmd\)— the sample size leaves the logarithm — and no single\-scale packing argument can show that any logarithm is necessary\. Section[8](https://arxiv.org/html/2607.07778#S8)proves projection\-capacity floors valid for*every*Lipschitz activation: the conjecture holds at width one with marginn/log⁡\(n​d\)\\sqrt\{n/\\log\(nd\)\}, at width two on the whole admissible dimension range, and at width three ford≥C​log⁡\(n​d\)d\\geq C\\log\(nd\), log\-free\. Beyondd∼n/\(m​log⁡\(n​d\)\)d\\sim n/\(m\\log\(nd\)\)the projection method is exhausted — its net cost reaches the label budget — and for general activations at widthm≥2m\\geq 2that regime remains open; separately, a linear\-activation interpolant shows that no floor exceedingC​nC\\sqrt\{n\}can hold onced≳nd\\gtrsim n\. Section[9](https://arxiv.org/html/2607.07778#S9)states the one open multiplier estimate \(Conjecture[9\.1](https://arxiv.org/html/2607.07778#S9.Thmtheorem1)\) to which the log\-free conjecture reduces in the critical band of widths; the reduction itself, together with the unconditional structure surrounding it — occupancy, serving capacity, pile\-up rigidity, cap mass, forced depth, an affine supremum identity, and the single\-direction case settled for*every*Lipschitz activation, with no logarithm — is developed in the supplementary note\[[13](https://arxiv.org/html/2607.07778#bib.bib13)\]\. The reduction is representation\-free, so the remaining obstacle for arbitrary activations coincides with the ReLU one\.

### 1\.1\.Standing notation and conventions

We fix these throughout\.

- •⟨⋅,⋅⟩\\langle\\cdot,\\cdot\\rangleand∥⋅∥\\\|\\cdot\\\|are the Euclidean inner product and norm onℝd\\mathbb\{R\}^\{d\};𝕊d−1=\{x:‖x‖=1\}\\mathbb\{S\}^\{d\-1\}=\\\{x:\\\|x\\\|=1\\\};BR=\{x:‖x‖≤R\}B\_\{R\}=\\\{x:\\\|x\\\|\\leq R\\\}\.
- •LipD⁡\(f\)=supx≠x′∈D\|f​\(x\)−f​\(x′\)\|/‖x−x′‖\\operatorname\{Lip\}\_\{D\}\(f\)=\\sup\_\{x\\neq x^\{\\prime\}\\in D\}\|f\(x\)\-f\(x^\{\\prime\}\)\|/\\\|x\-x^\{\\prime\}\\\|is the Euclidean Lipschitz constant offfon a setDD\.
- •ReLU⁡\(t\)=max⁡\(t,0\)\\operatorname\{ReLU\}\(t\)=\\max\(t,0\)\. A functionψ:ℝ→ℝ\\psi:\\mathbb\{R\}\\to\\mathbb\{R\}is*continuous piecewise linear withKKpieces*if it is continuous and there are breakpointsτ1<⋯<τK−1\\tau\_\{1\}<\\dots<\\tau\_\{K\-1\}such thatψ\\psiis affine on each of\(−∞,τ1\),\(τ1,τ2\),…,\(τK−1,∞\)\(\-\\infty,\\tau\_\{1\}\),\(\\tau\_\{1\},\\tau\_\{2\}\),\\dots,\(\\tau\_\{K\-1\},\\infty\);ReLU\\operatorname\{ReLU\}is the caseK=2K=2,τ1=0\\tau\_\{1\}=0\. Continuity is part of the definition and is used in Lemma[2\.1](https://arxiv.org/html/2607.07778#S2.Thmtheorem1)\.
- •For a unit vectoruuandt∈ℝt\\in\\mathbb\{R\},Hu,t=\{x:⟨u,x⟩=t\}H\_\{u,t\}=\\\{x:\\langle u,x\\rangle=t\\\}is the hyperplane with unit normaluuat signed distancettfrom the origin\.
- •clip⁡\(t\)=max⁡\(−1,min⁡\(1,t\)\)\\operatorname\{clip\}\(t\)=\\max\(\-1,\\min\(1,t\)\)is the projection ofℝ\\mathbb\{R\}onto\[−1,1\]\[\-1,1\]; it is11\-Lipschitz\.
- •ψ\\psialways denotes the network activation;ψ2\\psi\_\{2\}\(with a subscript\) denotes the sub\-Gaussian Orlicz norm‖Z‖ψ2=inf\{s\>0:𝔼​eZ2/s2≤2\}\\\|Z\\\|\_\{\\psi\_\{2\}\}=\\inf\\\{s\>0:\\mathbb\{E\}e^\{Z^\{2\}/s^\{2\}\}\\leq 2\\\}\.
- •c,C,c0,c2,C0,C′,κc,C,c\_\{0\},c\_\{2\},C\_\{0\},C^\{\\prime\},\\kappadenote positive absolute constants;c,Cc,Cmay change from line to line, while the subscripted constants are fixed once chosen\.
- •Every supremum over a class of networks or of Lipschitz functions that appears below inside an expectation is over a class that is separable in the uniform norm — the weights range over a finite\-dimensional space with continuous dependence, or over a ball of Lipschitz functions — so it equals the supremum over a fixed countable dense subset and is measurable; we writesup\\supwithout further comment\.

We writeψ\\psifor a piecewise\-linear activation withKKpieces and putm¯:=\(K−1\)​m\+1\\bar\{m\}:=\(K\-1\)m\+1; for ReLUm¯=m\+1\\bar\{m\}=m\+1\. We use two data models, both from\[[1](https://arxiv.org/html/2607.07778#bib.bib1)\]:

- \(S\)μ=\\mu=uniform probability measure on𝕊d−1\\mathbb\{S\}^\{d\-1\}, withd≥3d\\geq 3, and domainD=𝕊d−1D=\\mathbb\{S\}^\{d\-1\};
- \(G\)μ=N​\(0,Id/d\)\\mu=N\(0,I\_\{d\}/d\)and domainD=B2D=B\_\{2\}\.

Data are\(xi,yi\)i≤n\(x\_\{i\},y\_\{i\}\)\_\{i\\leq n\}i\.i\.d\. withxi∼μx\_\{i\}\\sim\\mu,yi∈\[−1,1\]y\_\{i\}\\in\[\-1,1\], and noise levelσ2:=𝔼​Var⁡\(y∣x\)\>0\\sigma^\{2\}:=\\mathbb\{E\}\\,\\operatorname\{Var\}\(y\\mid x\)\>0\. Independent±1\\pm 1labels are the special caseσ2=1\\sigma^\{2\}=1\.

### 1\.2\.The result

###### Theorem 1\.2\(Fixed\-width law\)\.

There are absolute constantsC0,c0\>0C\_\{0\},c\_\{0\}\>0such that the following holds\. Letψ\\psibe continuous piecewise linear withKKpieces, putm¯=\(K−1\)​m\+1\\bar\{m\}=\(K\-1\)m\+1, letε∈\(0,σ2\]\\varepsilon\\in\(0,\\sigma^\{2\}\]andδ∈\(0,1\)\\delta\\in\(0,1\), and assume

n≥C0​ε−2​log⁡\(8/δ\)\.n\\geq C\_\{0\}\\varepsilon^\{\-2\}\\log\(8/\\delta\)\.In model \(G\) assume additionallyd≥2​log⁡\(8​n/δ\)d\\geq 2\\log\(8n/\\delta\)\. Then, with probability at least1−δ1\-\\delta, everyf∈𝒩mf\\in\\mathcal\{N\}\_\{m\}with arbitrary weights and

1n​∑i=1n\(f​\(xi\)−yi\)2≤σ2−ε\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\bigl\(f\(x\_\{i\}\)\-y\_\{i\}\\bigr\)^\{2\}\\leq\\sigma^\{2\}\-\\varepsilonsatisfies

LipD⁡\(f\)≥c0​ε​nm¯​log⁡\(C0​m¯​n​d/ε\)\.\\operatorname\{Lip\}\_\{D\}\(f\)\\geq c\_\{0\}\\varepsilon\\sqrt\{\\frac\{n\}\{\\bar\{m\}\\log\\\!\\bigl\(C\_\{0\}\\bar\{m\}nd/\\varepsilon\\bigr\)\}\}\.

###### Corollary 1\.3\(BLN for piecewise\-linear activations, up to one logarithm\)\.

In the setting of Conjecture[1\.1](https://arxiv.org/html/2607.07778#S1.Thmtheorem1), withψ\\psicontinuous piecewise linear andd≥3d\\geq 3in the sphere model, there are constantsc1,C1\>0c\_\{1\},C\_\{1\}\>0depending onψ\\psionly through its number of pieces such that, with probability at least1−δ1\-\\delta, forn≥C1​log⁡\(8/δ\)n\\geq C\_\{1\}\\log\(8/\\delta\)everyf∈𝒩mf\\in\\mathcal\{N\}\_\{m\}with arbitrary weights that fits the data exactly, or merely has empirical mean squared error at most1/21/2, satisfies

Lip𝕊d−1⁡\(f\)≥c1​nm​log⁡\(C1​m​n​d\)\.\\operatorname\{Lip\}\_\{\\mathbb\{S\}^\{d\-1\}\}\(f\)\\geq c\_\{1\}\\sqrt\{\\frac\{n\}\{m\\log\(C\_\{1\}mnd\)\}\}\.

###### Corollary 1\.4\(Finite\-horizon simultaneous widths\)\.

Fix an integerM≥1M\\geq 1\. Under the hypotheses of Theorem[1\.2](https://arxiv.org/html/2607.07778#S1.Thmtheorem2), with

n≥C0​ε−2​log⁡\(8​M/δ\)n\\geq C\_\{0\}\\varepsilon^\{\-2\}\\log\(8M/\\delta\)and, in model \(G\),d≥2​log⁡\(8​n​M/δ\)d\\geq 2\\log\(8nM/\\delta\), there is an event of probability at least1−δ1\-\\deltaon which the conclusion of Theorem[1\.2](https://arxiv.org/html/2607.07778#S1.Thmtheorem2)holds simultaneously for every width1≤m≤M1\\leq m\\leq Mand everyf∈𝒩mf\\in\\mathcal\{N\}\_\{m\}\.

The width enters the proof only through the number of distinct hyperplanes on which the realized function has a kink\. For a two\-layer piecewise\-linear functionff, letk​\(f\)k\(f\)be the number of distinct hyperplanesHu,tH\_\{u,t\}carrying a nonzero kink of the canonical representation insideDD\. After Lemma[2\.1](https://arxiv.org/html/2607.07778#S2.Thmtheorem1), a width\-mmnetwork hask​\(f\)≤\(K−1\)​mk\(f\)\\leq\(K\-1\)m, but redundant or cancelling parameterizations may havek​\(f\)k\(f\)much smaller\.

###### Theorem 1\.5\(Realized kink\-count law\)\.

There are absolute constantsC0,c0\>0C\_\{0\},c\_\{0\}\>0such that, in the setting of Theorem[1\.2](https://arxiv.org/html/2607.07778#S1.Thmtheorem2), if

n≥C0​ε−2​log⁡\(8​n/δ\)n\\geq C\_\{0\}\\varepsilon^\{\-2\}\\log\(8n/\\delta\)and, in model \(G\),d≥2​log⁡\(8​n/δ\)d\\geq 2\\log\(8n/\\delta\), then with probability at least1−δ1\-\\deltaevery two\-layer piecewise\-linear networkff, of arbitrary width and arbitrary weights, withk​\(f\)≤nk\(f\)\\leq nand fitting the dataε\\varepsilonbelow the noise floor, satisfies

LipD⁡\(f\)≥c0​ε​n\(k​\(f\)\+1\)​log⁡\(C0​\(k​\(f\)\+1\)​n​d/ε\)\.\\operatorname\{Lip\}\_\{D\}\(f\)\\geq c\_\{0\}\\varepsilon\\sqrt\{\\frac\{n\}\{\(k\(f\)\+1\)\\log\\\!\\bigl\(C\_\{0\}\(k\(f\)\+1\)nd/\\varepsilon\\bigr\)\}\}\.

The kink\-count theorem is stronger than the fixed\-width theorem for realized functions withk​\(f\)≤nk\(f\)\\leq n; Theorem[1\.2](https://arxiv.org/html/2607.07778#S1.Thmtheorem2)is still stated separately because it gives a clean fixed\-width guarantee even when\(K−1\)​m\>n\(K\-1\)m\>n\.

###### Corollary 1\.6\(Structured single\-hidden\-layer architectures\)\.

Fix an integerK0≥0K\_\{0\}\\geq 0\. Under the hypotheses of Theorem[1\.2](https://arxiv.org/html/2607.07778#S1.Thmtheorem2), withm¯\\bar\{m\}replaced byK0\+1K\_\{0\}\+1, with probability at least1−δ1\-\\deltaevery scalar input\-output map of a single\-hidden\-layer piecewise\-linear ridge architecture whose realized function has at mostK0K\_\{0\}distinct kink hyperplanes obeys

LipD⁡\(f\)≥c0​ε​n\(K0\+1\)​log⁡\(C0​\(K0\+1\)​n​d/ε\)\\operatorname\{Lip\}\_\{D\}\(f\)\\geq c\_\{0\}\\varepsilon\\sqrt\{\\frac\{n\}\{\(K\_\{0\}\+1\)\\log\\\!\\bigl\(C\_\{0\}\(K\_\{0\}\+1\)nd/\\varepsilon\\bigr\)\}\}whenever it fitsε\\varepsilonbelow the noise floor\. The bound depends only on the realized function, not on the parameterization; hence any constrained single\-hidden\-layer parameterization — weight sharing across filters as in a convolutional layer, tied or repeated weights, a low\-rank factorization of the first layer — satisfies the same bound withK0K\_\{0\}the number of distinct kink hyperplanes the constraint permits\. If the activation hasKKpieces and a convolutional layer hasFFfilters evaluated atSSspatial positions, one may takeK0≤\(K−1\)​F​SK\_\{0\}\\leq\(K\-1\)FS\.

###### Corollary 1\.7\(Vector outputs\)\.

Let the output dimension berr\. Supposeyi∈\[−1,1\]ry\_\{i\}\\in\[\-1,1\]^\{r\}, and writeσℓ2=𝔼​Var⁡\(yℓ∣x\)\\sigma\_\{\\ell\}^\{2\}=\\mathbb\{E\}\\operatorname\{Var\}\(y\_\{\\ell\}\\mid x\)for the coordinate noise levels\. Letf=\(f1,…,fr\)f=\(f\_\{1\},\\dots,f\_\{r\}\), where eachfℓf\_\{\\ell\}is a two\-layer piecewise\-linear scalar network and all coordinates together use at mostK0K\_\{0\}distinct kink hyperplanes\. If

1n​∑i=1n‖f​\(xi\)−yi‖22≤∑ℓ=1rσℓ2−ε,\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\\|f\(x\_\{i\}\)\-y\_\{i\}\\\|\_\{2\}^\{2\}\\leq\\sum\_\{\\ell=1\}^\{r\}\\sigma\_\{\\ell\}^\{2\}\-\\varepsilon,then, on the event obtained by applying the scalar theorem to every coordinate withσℓ2≥ε/r\\sigma\_\{\\ell\}^\{2\}\\geq\\varepsilon/rusing accuracy parameterε/r\\varepsilon/rand failure probabilityδ/r\\delta/r, every suchffsatisfies

LipD⁡\(f\)≥c0​εr​n\(K0\+1\)​log⁡\(C0​\(K0\+1\)​n​d​r/ε\)\.\\operatorname\{Lip\}\_\{D\}\(f\)\\geq c\_\{0\}\\frac\{\\varepsilon\}\{r\}\\sqrt\{\\frac\{n\}\{\(K\_\{0\}\+1\)\\log\\\!\\bigl\(C\_\{0\}\(K\_\{0\}\+1\)ndr/\\varepsilon\\bigr\)\}\}\.Equivalently, it is enough to assume the scalar\-theorem sample\-size and localization hypotheses withε\\varepsilonreplaced byε/r\\varepsilon/randδ\\deltabyδ/r\\delta/r\.

### 1\.3\.Guide to the results

Each entry links to its statement; the\[proof→\\to\]marker jumps to the proof in Appendix[A](https://arxiv.org/html/2607.07778#A1)\.

- •[1\.1](https://arxiv.org/html/2607.07778#S1.Thmtheorem1)— the Bubeck–Li–Nagaraj law of robustness \(open conjecture\)\.[\[main results→\\to\]](https://arxiv.org/html/2607.07778#S1.Thmtheorem2)
- •[1\.2](https://arxiv.org/html/2607.07778#S1.Thmtheorem2)— fixed\-width law for piecewise\-linear activations, up to one logarithm\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS1)
- •[1\.3](https://arxiv.org/html/2607.07778#S1.Thmtheorem3)— BLN conjecture for piecewise\-linear activations, up to one logarithm\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS1)
- •[1\.4](https://arxiv.org/html/2607.07778#S1.Thmtheorem4)— the law, simultaneously over all widths up toMM\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS1)
- •[1\.5](https://arxiv.org/html/2607.07778#S1.Thmtheorem5)— realized\-kink\-count law,m¯\\bar\{m\}replaced byk​\(f\)\+1k\(f\)\+1\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS1)
- •[1\.6](https://arxiv.org/html/2607.07778#S1.Thmtheorem6)— convolutional, weight\-tied and low\-rank single\-hidden\-layer architectures\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS1)
- •[1\.7](https://arxiv.org/html/2607.07778#S1.Thmtheorem7)— vector\-valued outputs, coordinatewise\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS1)
- •[2\.1](https://arxiv.org/html/2607.07778#S2.Thmtheorem1)— reduce any piecewise\-linear activation to ReLU kinks\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS2)
- •[2\.2](https://arxiv.org/html/2607.07778#S2.Thmtheorem2)— canonical ReLU form on ball or sphere\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS2)
- •[3\.1](https://arxiv.org/html/2607.07778#S3.Thmtheorem1)— ball rigidity: kink coefficient bounded by Lipschitz constant\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS3)
- •[3\.2](https://arxiv.org/html/2607.07778#S3.Thmtheorem2)— sphere rigidity, with the1−tj2\\sqrt\{1\-t\_\{j\}^\{2\}\}factor\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS3)
- •[3\.3](https://arxiv.org/html/2607.07778#S3.Thmtheorem3)— rigidity genuinely fails on the circle𝕊1\\mathbb\{S\}^\{1\}\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS3)
- •[3\.5](https://arxiv.org/html/2607.07778#S3.Thmtheorem5)— automatic sup\-norm bound from fitting and Lipschitzness\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS3)
- •[3\.6](https://arxiv.org/html/2607.07778#S3.Thmtheorem6)— bounds on the affine partv,cv,c\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS3)
- •[4\.1](https://arxiv.org/html/2607.07778#S4.Thmtheorem1)— metric entropy of the canonical Lipschitz ball\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS4)
- •[5\.2](https://arxiv.org/html/2607.07778#S5.Thmtheorem2)— sphere and Gaussian satisfy Lipschitz concentration\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS5)
- •[5\.3](https://arxiv.org/html/2607.07778#S5.Thmtheorem3)— noise decomposition reducing fitting to a multiplier event\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS5)
- •[5\.4](https://arxiv.org/html/2607.07778#S5.Thmtheorem4)— single\-function sub\-Gaussian deviation bound\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS5)
- •[5\.5](https://arxiv.org/html/2607.07778#S5.Thmtheorem5)— mean\-term deviation bound\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS5)
- •[5\.6](https://arxiv.org/html/2607.07778#S5.Thmtheorem6)— finite\-class law via concentration\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS5)
- •[7\.1](https://arxiv.org/html/2607.07778#S7.Thmtheorem1)— harmonic split into constant, linear, higher\-degree parts\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS6)
- •[7\.2](https://arxiv.org/html/2607.07778#S7.Thmtheorem2)— entropy and Dudley bound at the population radius\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS6)
- •[7\.3](https://arxiv.org/html/2607.07778#S7.Thmtheorem3)— self\-bounding inequality for the empirical radius\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS6)
- •[7\.4](https://arxiv.org/html/2607.07778#S7.Thmtheorem4)— Rademacher complexity carrying a sample\-free logarithm\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS6)
- •[7\.5](https://arxiv.org/html/2607.07778#S7.Thmtheorem5)— the law with the sample size out of the logarithm\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS6)
- •[8\.1](https://arxiv.org/html/2607.07778#S8.Thmtheorem1)— any Lipschitz activation at small width, projection floor\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS7)
- •[8\.3](https://arxiv.org/html/2607.07778#S8.Thmtheorem3)— localized projection floor, sharpened small\-width rate\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS7)
- •[8\.4](https://arxiv.org/html/2607.07778#S8.Thmtheorem4)— per\-width consequences of the localized floor\. \(no separate proof; consequences of[8\.3](https://arxiv.org/html/2607.07778#S8.Thmtheorem3)\)
- •[9\.1](https://arxiv.org/html/2607.07778#S9.Thmtheorem1)— statement of the open multiplier estimate; reduction and surrounding structure in the supplementary note\[[13](https://arxiv.org/html/2607.07778#bib.bib13)\]\.
- •[10\.1](https://arxiv.org/html/2607.07778#S10.Thmtheorem1)— matching two\-layer ReLU upper bound atm≍nm\\asymp n\.[\[proof→\\to\]](https://arxiv.org/html/2607.07778#A1.SS8)

### 1\.4\.The idea of the proof

A two\-layer piecewise\-linear function is a sum of ridge functions plus an affine map; each unit contributes kinks on a finite family of parallel hyperplanes\. The proofs in\[[1](https://arxiv.org/html/2607.07778#bib.bib1),[2](https://arxiv.org/html/2607.07778#bib.bib2),[3](https://arxiv.org/html/2607.07778#bib.bib3)\]discretize parameter space, whose covering number is finite only after bounding the parameters\. Since the weights here are arbitrary, the proof instead discretizes the realized functions\.

The special deterministic structure is a rigidity phenomenon \(Section[3](https://arxiv.org/html/2607.07778#S3)\): kinks supported on distinct hyperplanes cannot cancel at a generic point of one kink hyperplane\. At such a point all other units are locally affine, so the jump of a one\-dimensional derivative equals the coefficient of the single kink under inspection; anLL\-Lipschitz function can have a derivative jump of size at most2​L2L\. Thus, after an exact canonical rewriting on the domain \(Section[2](https://arxiv.org/html/2607.07778#S2)\), every kink coefficient is bounded in terms ofLL\(with the natural spherical factor\)\. The canonical parameters then lie in a bounded set depending only onLL,mm, anddd, not on the original weights, giving the metric entropy bound of Section[4](https://arxiv.org/html/2607.07778#S4)\. A finite\-class noise\-decomposition argument, following the concentration mechanism of\[[2](https://arxiv.org/html/2607.07778#bib.bib2)\]but written self\-contained for the two data models, finishes the proof\.

The mechanism is kink\-specific\. Smooth activations admit bounded\-Lipschitz finite\-difference families whose representing parameters must escape to infinity, so the canonical\-parameter argument does not extend directly\. Rigidity also fails on𝕊1\\mathbb\{S\}^\{1\}because a kink set there consists of two points rather than a positive\-dimensional sphere\. These limitations are recorded in Sections[3](https://arxiv.org/html/2607.07778#S3)and[10](https://arxiv.org/html/2607.07778#S10); they are not hidden assumptions in the main theorem\.

## 2\.The canonical form

###### Lemma 2\.1\(Reduction to ReLU kinks\)\.

Letψ:ℝ→ℝ\\psi:\\mathbb\{R\}\\to\\mathbb\{R\}be continuous piecewise linear withKKpieces\. IfK=1K=1thenψ\\psiis affine and everyf∈𝒩mf\\in\\mathcal\{N\}\_\{m\}is itself affine, already of the form \([1](https://arxiv.org/html/2607.07778#S1.E1)\) with no ReLU units; so assumeK≥2K\\geq 2, with breakpointsτ1<⋯<τK−1\\tau\_\{1\}<\\dots<\\tau\_\{K\-1\}and successive slopess0,…,sK−1s\_\{0\},\\dots,s\_\{K\-1\}\(soψ\\psihas slopesκs\_\{\\kappa\}on\(τκ,τκ\+1\)\(\\tau\_\{\\kappa\},\\tau\_\{\\kappa\+1\}\), withτ0=−∞\\tau\_\{0\}=\-\\infty,τK=\+∞\\tau\_\{K\}=\+\\infty\)\. Then for allt∈ℝt\\in\\mathbb\{R\},

ψ​\(t\)=ψ​\(τ1\)\+s0​\(t−τ1\)\+∑κ=1K−1\(sκ−sκ−1\)​ReLU⁡\(t−τκ\)\.\\psi\(t\)=\\psi\(\\tau\_\{1\}\)\+s\_\{0\}\(t\-\\tau\_\{1\}\)\+\\sum\_\{\\kappa=1\}^\{K\-1\}\(s\_\{\\kappa\}\-s\_\{\\kappa\-1\}\)\\,\\operatorname\{ReLU\}\(t\-\\tau\_\{\\kappa\}\)\.\(2\)Consequently everyf∈𝒩mf\\in\\mathcal\{N\}\_\{m\}with activationψ\\psiequals, at every point ofℝd\\mathbb\{R\}^\{d\}, a network of the form \([1](https://arxiv.org/html/2607.07778#S1.E1)\) with activationReLU\\operatorname\{ReLU\}and width at most\(K−1\)​m\(K\-1\)m; the identity \([2](https://arxiv.org/html/2607.07778#S2.E2)\) is the classical hinge representation of a continuous piecewise\-linear function\[[8](https://arxiv.org/html/2607.07778#bib.bib8),[9](https://arxiv.org/html/2607.07778#bib.bib9)\]\.

From here onψ=ReLU\\psi=\\operatorname\{ReLU\}and the width is writtenmm; in the final statements it is replaced by\(K−1\)​m≤m¯−1\(K\-1\)m\\leq\\bar\{m\}\-1\. The next lemma is pure bookkeeping, but we spell it out because the rigidity lemma needs the precise output\.

###### Lemma 2\.2\(Canonical form on a domain\)\.

LetD=BRD=B\_\{R\}\(anyR\>0R\>0\) orD=𝕊d−1D=\\mathbb\{S\}^\{d\-1\}withd≥2d\\geq 2\. Every ReLU networkf∈𝒩mf\\in\\mathcal\{N\}\_\{m\}can be rewritten, so that the two sides agree at every point ofDD, as

f​\(x\)=⟨v,x⟩\+c\+∑j=1m0αj​ReLU⁡\(⟨uj,x⟩−tj\),m0≤m,f\(x\)=\\langle v,x\\rangle\+c\+\\sum\_\{j=1\}^\{m\_\{0\}\}\\alpha\_\{j\}\\,\\operatorname\{ReLU\}\(\\langle u\_\{j\},x\\rangle\-t\_\{j\}\),\\qquad m\_\{0\}\\leq m,\(3\)where‖uj‖=1\\\|u\_\{j\}\\\|=1,αj≠0\\alpha\_\{j\}\\neq 0, the hyperplanesHuj,tjH\_\{u\_\{j\},t\_\{j\}\}are pairwise distinct*as sets*, and

- \(i\)ball case:tj∈\(−R,R\)t\_\{j\}\\in\(\-R,R\), soHuj,tj∩int⁡BRH\_\{u\_\{j\},t\_\{j\}\}\\cap\\operatorname\{int\}B\_\{R\}is a nonempty relatively open\(d−1\)\(d\-1\)\-dimensional disk;
- \(ii\)sphere case:tj∈\[0,1\)t\_\{j\}\\in\[0,1\), soHuj,tj∩𝕊d−1H\_\{u\_\{j\},t\_\{j\}\}\\cap\\mathbb\{S\}^\{d\-1\}is a\(d−2\)\(d\-2\)\-sphere of radius1−tj2\>0\\sqrt\{1\-t\_\{j\}^\{2\}\}\>0\.

## 3\.Rigidity

### 3\.1\.The ball

###### Lemma 3\.1\(Rigidity, ball\)\.

Letffbe in the canonical form \([3](https://arxiv.org/html/2607.07778#S2.E3)\) onD=BRD=B\_\{R\}, withL:=LipBR⁡\(f\)<∞L:=\\operatorname\{Lip\}\_\{B\_\{R\}\}\(f\)<\\infty\. Then\|αj\|≤2​L\|\\alpha\_\{j\}\|\\leq 2Lfor everyjj\.

### 3\.2\.The sphere

###### Lemma 3\.2\(Rigidity, sphere\)\.

Letd≥3d\\geq 3and letffbe in the canonical form \([3](https://arxiv.org/html/2607.07778#S2.E3)\) onD=𝕊d−1D=\\mathbb\{S\}^\{d\-1\}, withL:=Lip𝕊d−1⁡\(f\)<∞L:=\\operatorname\{Lip\}\_\{\\mathbb\{S\}^\{d\-1\}\}\(f\)<\\infty\(Euclidean metric\)\. Then

\|αj\|​1−tj2≤2​Lfor every​j\.\|\\alpha\_\{j\}\|\\sqrt\{1\-t\_\{j\}^\{2\}\}\\;\\leq\\;2L\\qquad\\text\{for every \}j\.

### 3\.3\.Rigidity fails on the circle

The restriction tod≥3d\\geq 3in model \(S\) is necessary: the coefficient bound of Lemma[3\.2](https://arxiv.org/html/2607.07778#S3.Thmtheorem2)is false atd=2d=2\.

###### Proposition 3\.3\(Failure of rigidity atd=2d=2\)\.

For everyR\>1R\>1and everyΛ\>0\\Lambda\>0there is a canonical four\-unit ReLU functionFFon𝕊1\\mathbb\{S\}^\{1\}, with four distinct kink point\-pairs, such that

\|αj\|​1−tj2=Λ\(j=1,…,4\),\|\\alpha\_\{j\}\|\\sqrt\{1\-t\_\{j\}^\{2\}\}=\\Lambda\\qquad\(j=1,\\dots,4\),whileLip𝕊1⁡\(F\)≤Λ/R\\operatorname\{Lip\}\_\{\\mathbb\{S\}^\{1\}\}\(F\)\\leq\\Lambda/R\. Consequently no absolute constantCCcan make\|αj\|​1−tj2≤C​Lip𝕊1⁡\(F\)\|\\alpha\_\{j\}\|\\sqrt\{1\-t\_\{j\}^\{2\}\}\\leq C\\operatorname\{Lip\}\_\{\\mathbb\{S\}^\{1\}\}\(F\)valid for all canonical representations on𝕊1\\mathbb\{S\}^\{1\}\.

### 3\.4\.Bounds on the affine part

We first record the automatic value bound, then propagate rigidity tovvandcc\.

###### Lemma 3\.5\(Value bound\)\.

Supposeε≤σ2≤1\\varepsilon\\leq\\sigma^\{2\}\\leq 1and1n​∑i\(f​\(xi\)−yi\)2≤σ2−ε≤1\\frac\{1\}\{n\}\\sum\_\{i\}\(f\(x\_\{i\}\)\-y\_\{i\}\)^\{2\}\\leq\\sigma^\{2\}\-\\varepsilon\\leq 1, with allxi∈Dx\_\{i\}\\in Dand\|yi\|≤1\|y\_\{i\}\|\\leq 1\. IfLipD⁡\(f\)≤L\\operatorname\{Lip\}\_\{D\}\(f\)\\leq Landdiam⁡\(D\)≤4\\operatorname\{diam\}\(D\)\\leq 4, thensupD\|f\|≤2\+L​diam⁡\(D\)≤B0:=2\+4​L\\sup\_\{D\}\|f\|\\leq 2\+L\\operatorname\{diam\}\(D\)\\leq B\_\{0\}:=2\+4L\.

###### Lemma 3\.6\(Affine part\)\.

Letffbe in canonical form \([3](https://arxiv.org/html/2607.07778#S2.E3)\) withsupD\|f\|≤B0\\sup\_\{D\}\|f\|\\leq B\_\{0\}andLipD⁡\(f\)≤L\\operatorname\{Lip\}\_\{D\}\(f\)\\leq L\.

- \(i\)Ball case \(D=BRD=B\_\{R\},R≤2R\\leq 2\):‖v‖≤L​\(1\+2​m0\)\\\|v\\\|\\leq L\(1\+2m\_\{0\}\)and\|c\|≤B0\+4​L​m0\|c\|\\leq B\_\{0\}\+4Lm\_\{0\}\.
- \(ii\)Sphere case \(D=𝕊d−1D=\\mathbb\{S\}^\{d\-1\},d≥3d\\geq 3\):‖v‖≤d​\(B0\+2​L​m0\)\\\|v\\\|\\leq d\\,\(B\_\{0\}\+2Lm\_\{0\}\)and\|c\|≤B0\+‖v‖\+2​L​m0\|c\|\\leq B\_\{0\}\+\\\|v\\\|\+2Lm\_\{0\}\.

## 4\.Metric entropy of the Lipschitz ball

Fix the domainDD\(B2B\_\{2\}in model \(G\);𝕊d−1\\mathbb\{S\}^\{d\-1\},d≥3d\\geq 3, in model \(S\)\), a width budgetm¯\\bar\{m\}, andL\>0L\>0; putB0=2\+4​LB\_\{0\}=2\+4L\. Define the class we must control,

𝒜m¯,L:=\{f\|D:f∈𝒩m¯\(activationReLU\),LipD\(f\)≤L,supD\|f\|≤B0\},\\mathcal\{A\}\_\{\\bar\{m\},L\}:=\\bigl\\\{\\,f\|\_\{D\}:\\ f\\in\\mathcal\{N\}\_\{\\bar\{m\}\}\\ \(\\text\{activation \}\\operatorname\{ReLU\}\),\\ \\operatorname\{Lip\}\_\{D\}\(f\)\\leq L,\\ \\sup\_\{D\}\|f\|\\leq B\_\{0\}\\,\\bigr\\\},and the parameter\-box superclass𝒢m¯,L\\mathcal\{G\}\_\{\\bar\{m\},L\}: all functions of the form \([3](https://arxiv.org/html/2607.07778#S2.E3)\) onDDwithm0≤m¯m\_\{0\}\\leq\\bar\{m\}and, in the ball case,

\|αj\|≤2​L,tj∈\(−2,2\),‖v‖≤L​\(1\+2​m¯\),\|c\|≤B0\+4​L​m¯,\|\\alpha\_\{j\}\|\\leq 2L,\\quad t\_\{j\}\\in\(\-2,2\),\\quad\\\|v\\\|\\leq L\(1\+2\\bar\{m\}\),\\quad\|c\|\\leq B\_\{0\}\+4L\\bar\{m\},and in the sphere case,

\|αj\|≤2​L1−tj2,tj∈\[0,1\),‖v‖≤d​\(B0\+2​L​m¯\),\|c\|≤B0\+‖v‖\+2​L​m¯\.\|\\alpha\_\{j\}\|\\leq\\frac\{2L\}\{\\sqrt\{1\-t\_\{j\}^\{2\}\}\},\\quad t\_\{j\}\\in\[0,1\),\\quad\\\|v\\\|\\leq d\(B\_\{0\}\+2L\\bar\{m\}\),\\quad\|c\|\\leq B\_\{0\}\+\\\|v\\\|\+2L\\bar\{m\}\.By Sections[2](https://arxiv.org/html/2607.07778#S2)–[3](https://arxiv.org/html/2607.07778#S3)\(canonical form, then Lemmas[3\.1](https://arxiv.org/html/2607.07778#S3.Thmtheorem1)/[3\.2](https://arxiv.org/html/2607.07778#S3.Thmtheorem2)/[3\.6](https://arxiv.org/html/2607.07778#S3.Thmtheorem6)\), every member of𝒜m¯,L\\mathcal\{A\}\_\{\\bar\{m\},L\}satisfies these bounds:

𝒜m¯,L⊂𝒢m¯,L\.\\mathcal\{A\}\_\{\\bar\{m\},L\}\\subset\\mathcal\{G\}\_\{\\bar\{m\},L\}\.\(4\)
RecallN\(𝒦,∥⋅∥,ε′\)N\(\\mathcal\{K\},\\\|\\cdot\\\|,\\varepsilon^\{\\prime\}\)is the smallest number of∥⋅∥\\\|\\cdot\\\|\-balls of radiusε′\\varepsilon^\{\\prime\}needed to cover𝒦\\mathcal\{K\}\.

###### Proposition 4\.1\(Metric entropy\)\.

There is an absolute constantCCsuch that for allε′∈\(0,4\+8​L\)\\varepsilon^\{\\prime\}\\in\(0,4\+8L\),L\>0L\>0,m¯≥1\\bar\{m\}\\geq 1, andd≥2d\\geq 2\(ball\) ord≥3d\\geq 3\(sphere\),

logN\(𝒢m¯,L,∥⋅∥L∞​\(D\),ε′\)≤Cm¯dlog\(C​m¯​d​\(2\+L\)ε′\)\.\\log N\\bigl\(\\mathcal\{G\}\_\{\\bar\{m\},L\},\\ \\\|\\cdot\\\|\_\{L^\{\\infty\}\(D\)\},\\ \\varepsilon^\{\\prime\}\\bigr\)\\;\\leq\\;C\\,\\bar\{m\}\\,d\\,\\log\\\!\\Bigl\(\\frac\{C\\,\\bar\{m\}\\,d\\,\(2\+L\)\}\{\\varepsilon^\{\\prime\}\}\\Bigr\)\.\(The upper limit4\+8​L4\+8Lonε′\\varepsilon^\{\\prime\}serves only to keep the logarithm’s argument at leastee, so that the displayed closed form is valid; the grid construction in the proof below has mesh proportional toε′\\varepsilon^\{\\prime\}and yields a valid cover at every scale\.\) Moreover𝒜m¯,L\\mathcal\{A\}\_\{\\bar\{m\},L\}has an internalε′\\varepsilon^\{\\prime\}\-net \(centers in𝒜m¯,L\\mathcal\{A\}\_\{\\bar\{m\},L\}\) of cardinality≤N\(𝒢m¯,L,∥⋅∥∞,ε′/2\)\\leq N\(\\mathcal\{G\}\_\{\\bar\{m\},L\},\\\|\\cdot\\\|\_\{\\infty\},\\varepsilon^\{\\prime\}/2\), hence of the same log bound\.

## 5\.The probabilistic core

This section reproves, self\-contained, the concentration estimate of\[[2](https://arxiv.org/html/2607.07778#bib.bib2)\]for a*finite*function class, in the two data models\. Recall‖Z‖ψ2=inf\{s\>0:𝔼​eZ2/s2≤2\}\\\|Z\\\|\_\{\\psi\_\{2\}\}=\\inf\\\{s\>0:\\mathbb\{E\}e^\{Z^\{2\}/s^\{2\}\}\\leq 2\\\}\.

###### Definition 5\.1\.

A probability measureμ\\muonℝd\\mathbb\{R\}^\{d\}\(or on𝕊d−1\\mathbb\{S\}^\{d\-1\}\) satisfies*κ\\kappa\-Lipschitz concentration*if for every boundedL′L^\{\\prime\}\-Lipschitzffon its support \(Euclidean metric\),‖f​\(x\)−𝔼​f‖ψ2≤κ​L′/d\\\|f\(x\)\-\\mathbb\{E\}f\\\|\_\{\\psi\_\{2\}\}\\leq\\kappa L^\{\\prime\}/\\sqrt\{d\}\.

###### Lemma 5\.2\.

There is an absoluteκ\\kappasuch that \(a\) the uniform measure on𝕊d−1\\mathbb\{S\}^\{d\-1\},d≥2d\\geq 2, and \(b\)N​\(0,Id/d\)N\(0,I\_\{d\}/d\)satisfyκ\\kappa\-Lipschitz concentration\.

Throughout the rest of the section:\(xi,yi\)i≤n\(x\_\{i\},y\_\{i\}\)\_\{i\\leq n\}are i\.i\.d\. withxi∼μx\_\{i\}\\sim\\mu\(κ\\kappa\-Lipschitz concentrated\),\|yi\|≤1\|y\_\{i\}\|\\leq 1;g​\(x\):=𝔼​\[y∣x\]g\(x\):=\\mathbb\{E\}\[y\\mid x\], so\|g\|≤1\|g\|\\leq 1;zi:=yi−g​\(xi\)z\_\{i\}:=y\_\{i\}\-g\(x\_\{i\}\), so\|zi\|≤2\|z\_\{i\}\|\\leq 2,𝔼​\[zi∣xi\]=0\\mathbb\{E\}\[z\_\{i\}\\mid x\_\{i\}\]=0, and

𝔼​zi2=𝔼​\[𝔼​\[\(y−𝔼​\[y∣x\]\)2∣x\]\]=𝔼​Var⁡\(y∣x\)=σ2;\\mathbb\{E\}z\_\{i\}^\{2\}=\\mathbb\{E\}\\bigl\[\\mathbb\{E\}\[\(y\-\\mathbb\{E\}\[y\\mid x\]\)^\{2\}\\mid x\]\\bigr\]=\\mathbb\{E\}\\,\\operatorname\{Var\}\(y\\mid x\)=\\sigma^\{2\};andℱ\\mathcal\{F\}is a*finite*set of functions onsupp⁡μ\\operatorname\{supp\}\\muwith values in\[−1,1\]\[\-1,1\], eachLL\-Lipschitz\.

###### Lemma 5\.3\(Noise decomposition\)\.

Forε∈\(0,σ2\]\\varepsilon\\in\(0,\\sigma^\{2\}\],

ℙ\(∃f∈ℱ:1n∑i\(f\(xi\)−yi\)2≤σ2−ε\)≤2e−n​ε2/288\+ℙ\(∃f∈ℱ:1n∑izif\(xi\)≥ε4\)\.\\mathbb\{P\}\\Bigl\(\\exists f\\in\\mathcal\{F\}:\\ \\tfrac\{1\}\{n\}\\textstyle\\sum\_\{i\}\(f\(x\_\{i\}\)\-y\_\{i\}\)^\{2\}\\leq\\sigma^\{2\}\-\\varepsilon\\Bigr\)\\leq 2e^\{\-n\\varepsilon^\{2\}/288\}\+\\mathbb\{P\}\\Bigl\(\\exists f\\in\\mathcal\{F\}:\\ \\tfrac\{1\}\{n\}\\textstyle\\sum\_\{i\}z\_\{i\}f\(x\_\{i\}\)\\geq\\tfrac\{\\varepsilon\}\{4\}\\Bigr\)\.

###### Lemma 5\.4\(One function\)\.

Forf∈ℱf\\in\\mathcal\{F\}putWi:=zi​\(f​\(xi\)−𝔼​f\)W\_\{i\}:=z\_\{i\}\(f\(x\_\{i\}\)\-\\mathbb\{E\}f\)\. ThenWiW\_\{i\}are i\.i\.d\.,𝔼​Wi=0\\mathbb\{E\}W\_\{i\}=0, and for an absolutec2\>0c\_\{2\}\>0,

ℙ​\(1n​∑iWi≥ε8\)≤exp⁡\(−c2​n​ε2​max⁡\(1,dκ2​L2\)\)\.\\mathbb\{P\}\\Bigl\(\\tfrac\{1\}\{n\}\\textstyle\\sum\_\{i\}W\_\{i\}\\geq\\tfrac\{\\varepsilon\}\{8\}\\Bigr\)\\leq\\exp\\Bigl\(\-c\_\{2\}\\,n\\,\\varepsilon^\{2\}\\,\\max\\bigl\(1,\\tfrac\{d\}\{\\kappa^\{2\}L^\{2\}\}\\bigr\)\\Bigr\)\.

###### Lemma 5\.5\(Mean term\)\.

ℙ\(∃f∈ℱ:\(𝔼f\)1n∑izi≥ε8\)≤ℙ\(\|1n∑izi\|≥ε8\)≤2e−n​ε2/512\.\\displaystyle\\mathbb\{P\}\\Bigl\(\\exists f\\in\\mathcal\{F\}:\\ \(\\mathbb\{E\}f\)\\tfrac\{1\}\{n\}\\textstyle\\sum\_\{i\}z\_\{i\}\\geq\\tfrac\{\\varepsilon\}\{8\}\\Bigr\)\\leq\\mathbb\{P\}\\Bigl\(\\bigl\|\\tfrac\{1\}\{n\}\\textstyle\\sum\_\{i\}z\_\{i\}\\bigr\|\\geq\\tfrac\{\\varepsilon\}\{8\}\\Bigr\)\\leq 2e^\{\-n\\varepsilon^\{2\}/512\}\.

###### Theorem 5\.6\(Finite\-class law\)\.

In the setting above, forε∈\(0,σ2\]\\varepsilon\\in\(0,\\sigma^\{2\}\],

ℙ\(∃f∈ℱ:1n∑i\(f\(xi\)−yi\)2≤σ2−ε\)≤4e−n​ε2/512\+\|ℱ\|exp\(−c2nε2max\(1,dκ2​L2\)\)\.\\mathbb\{P\}\\Bigl\(\\exists f\\in\\mathcal\{F\}:\\ \\tfrac\{1\}\{n\}\\textstyle\\sum\_\{i\}\(f\(x\_\{i\}\)\-y\_\{i\}\)^\{2\}\\leq\\sigma^\{2\}\-\\varepsilon\\Bigr\)\\leq 4e^\{\-n\\varepsilon^\{2\}/512\}\+\|\\mathcal\{F\}\|\\exp\\Bigl\(\-c\_\{2\}n\\varepsilon^\{2\}\\max\\bigl\(1,\\tfrac\{d\}\{\\kappa^\{2\}L^\{2\}\}\\bigr\)\\Bigr\)\.

## 6\.Absence of a saturation cap

There is no saturation cap: the factorddin the entropyℰ\\mathcal\{E\}is exactly cancelled, in Case B, by the factorddin the case thresholdL∗\>d/κL^\{\\ast\}\>\\sqrt\{d\}/\\kappa, so the same smallc0c\_\{0\}serves both the sub\-saturation regime \(Case A, where the isoperimetricd/L2d/L^\{2\}gain is available\) and the large\-L∗L^\{\\ast\}regime \(Case B, where the crude bounded\-function concentration suffices because there are few functions\)\.

## 7\.Removing the sample size from the logarithm

The logarithm in Theorem[1\.2](https://arxiv.org/html/2607.07778#S1.Thmtheorem2)contains the sample sizennbecause the union bound runs over a net at the fitting accuracyε\\varepsilon\. In this section we move the union to the class’s populationL2L^\{2\}\-radius instead, which isL/2​dL/\\sqrt\{2d\}by a spectral\-gap argument; the resulting law carries a logarithm depending only on the width and the dimension\. The argument has three steps: a harmonic split of the clipped function into an affine part and a high\-frequency remainder of small population radius \(Lemma[7\.1](https://arxiv.org/html/2607.07778#S7.Thmtheorem1)\); a bound on the empirical radius via a self\-bounding fixed point \(Lemmas[7\.2](https://arxiv.org/html/2607.07778#S7.Thmtheorem2)–[7\.3](https://arxiv.org/html/2607.07778#S7.Thmtheorem3)\), which feeds the Rademacher complexity estimate of Theorem[7\.4](https://arxiv.org/html/2607.07778#S7.Thmtheorem4); and the assembly by contraction and bounded differences \(Theorem[7\.5](https://arxiv.org/html/2607.07778#S7.Thmtheorem5)\)\. Throughout this section we work in model \(S\) withd≥3d\\geq 3; writeF:=clip∘fF:=\\operatorname\{clip\}\\circ fforf∈𝒜m¯,Lf\\in\\mathcal\{A\}\_\{\\bar\{m\},L\},ℱclip:=\{clip∘f:f∈𝒜m¯,L\}\\mathcal\{F\}^\{\\operatorname\{clip\}\}:=\\\{\\operatorname\{clip\}\\circ f:f\\in\\mathcal\{A\}\_\{\\bar\{m\},L\}\\\}, and letF=∑ℓ≥0FℓF=\\sum\_\{\\ell\\geq 0\}F\_\{\\ell\}be the spherical\-harmonic decomposition \(see\[[10](https://arxiv.org/html/2607.07778#bib.bib10)\]\),λℓ=ℓ​\(ℓ\+d−2\)\\lambda\_\{\\ell\}=\\ell\(\\ell\+d\-2\)\.

###### Lemma 7\.1\(Harmonic split\)\.

EveryF∈ℱclipF\\in\\mathcal\{F\}^\{\\operatorname\{clip\}\}decomposes asF=cF\+⟨AF,x⟩\+hFF=c\_\{F\}\+\\langle A\_\{F\},x\\rangle\+h\_\{F\}withcF=𝔼​Fc\_\{F\}=\\mathbb\{E\}FandAF=d​𝔼​\[F​x\]A\_\{F\}=d\\,\\mathbb\{E\}\[Fx\], where

\|cF\|≤1,‖AF‖≤32​L,𝔼​hF2≤L22​d,sup𝕊d−1\|hF\|≤B1:=2\+32​L,\|c\_\{F\}\|\\leq 1,\\qquad\\\|A\_\{F\}\\\|\\leq\\sqrt\{\\tfrac\{3\}\{2\}\}\\,L,\\qquad\\mathbb\{E\}h\_\{F\}^\{2\}\\leq\\frac\{L^\{2\}\}\{2d\},\\qquad\\sup\_\{\\mathbb\{S\}^\{d\-1\}\}\|h\_\{F\}\|\\leq B\_\{1\}:=2\+\\sqrt\{\\tfrac\{3\}\{2\}\}\\,L,andhFh\_\{F\}is orthogonal to the constants and to the linear functions\.

Letℋ:=\{hF:F∈ℱclip\}\\mathcal\{H\}:=\\\{h\_\{F\}:F\\in\\mathcal\{F\}^\{\\operatorname\{clip\}\}\\\}; note0∈ℋ0\\in\\mathcal\{H\}\(takef=0f=0\)\. Put

B:=C1​m¯​d2​\(2\+L\),ΛL:=log⁡\(e​m¯​d3​\(2\+1L\)\),B:=C\_\{1\}\\bar\{m\}d^\{2\}\(2\+L\),\\qquad\\Lambda\_\{L\}:=\\log\\Bigl\(e\\bar\{m\}d^\{3\}\\Bigl\(2\+\\frac\{1\}\{L\}\\Bigr\)\\Bigr\),withC1C\_\{1\}a large absolute constant fixed by the next proof\. The scaleBBcarries the class size and governs the entropy at all scales; the logarithmΛL\\Lambda\_\{L\}that survives to the final bounds is the one evaluated at the population radiusL/2​dL/\\sqrt\{2d\}, where the ratio\(2\+L\)/L\(2\+L\)/Lis what appears, soΛL\\Lambda\_\{L\}grows only whenLLis small, never withLLlarge — and never withnn\.

###### Lemma 7\.2\(Entropy at the radius\)\.

logN\(ℋ,∥⋅∥L∞,u\)≤Cm¯dlog\(CB/u\)\\log N\(\\mathcal\{H\},\\\|\\cdot\\\|\_\{L^\{\\infty\}\},u\)\\leq C\\bar\{m\}d\\,\\log\(CB/u\)foru∈\(0,2​B1\]u\\in\(0,2B\_\{1\}\], and, conditionally on the sample, withσ^2:=suph∈ℋ1n​∑ih​\(xi\)2\\hat\{\\sigma\}^\{2\}:=\\sup\_\{h\\in\\mathcal\{H\}\}\\tfrac\{1\}\{n\}\\sum\_\{i\}h\(x\_\{i\}\)^\{2\},

𝔼ε​suph∈ℋ1n​\|∑iεi​h​\(xi\)\|≤C​m¯​dn​φ​\(σ^\),φ​\(s\):=s​log⁡\(2​e​B/s\)\.\\mathbb\{E\}\_\{\\varepsilon\}\\sup\_\{h\\in\\mathcal\{H\}\}\\frac\{1\}\{n\}\\Bigl\|\\sum\_\{i\}\\varepsilon\_\{i\}h\(x\_\{i\}\)\\Bigr\|\\ \\leq\\ C\\sqrt\{\\frac\{\\bar\{m\}d\}\{n\}\}\\;\\varphi\(\\hat\{\\sigma\}\),\\qquad\\varphi\(s\):=s\\sqrt\{\\log\\bigl\(2eB/s\\bigr\)\}\.

###### Lemma 7\.3\(Radius self\-bounding\)\.

LetR~:=𝔼x,ε​suph∈ℋ1n​\|∑iεi​h​\(xi\)\|\\widetilde\{R\}:=\\mathbb\{E\}\_\{x,\\varepsilon\}\\sup\_\{h\\in\\mathcal\{H\}\}\\tfrac\{1\}\{n\}\|\\sum\_\{i\}\\varepsilon\_\{i\}h\(x\_\{i\}\)\|\. Then𝔼x​σ^2≤L22​d\+8​B1​R~\\mathbb\{E\}\_\{x\}\\hat\{\\sigma\}^\{2\}\\leq\\dfrac\{L^\{2\}\}\{2d\}\+8B\_\{1\}\\widetilde\{R\}\.

###### Theorem 7\.4\(Complexity with a sample\-free logarithm\)\.

There is an absolute constantCCsuch that for alln,m¯≥1n,\\bar\{m\}\\geq 1,d≥3d\\geq 3,L\>0L\>0,

𝔼x,y​supf∈𝒜m¯,L1n​\|∑i=1nyi​clip⁡\(f​\(xi\)\)\|≤C​\[1\+Ln\+L​m¯​ΛLn\+\(1\+L\)​m¯​d​ΛLn\],\\mathbb\{E\}\_\{x,y\}\\ \\sup\_\{f\\in\\mathcal\{A\}\_\{\\bar\{m\},L\}\}\\ \\frac\{1\}\{n\}\\Bigl\|\\sum\_\{i=1\}^\{n\}y\_\{i\}\\,\\operatorname\{clip\}\(f\(x\_\{i\}\)\)\\Bigr\|\\ \\leq\\ C\\Bigl\[\\frac\{1\+L\}\{\\sqrt\{n\}\}\\ \+\\ L\\sqrt\{\\frac\{\\bar\{m\}\\,\\Lambda\_\{L\}\}\{n\}\}\\ \+\\ \(1\+L\)\\,\\frac\{\\bar\{m\}d\\,\\Lambda\_\{L\}\}\{n\}\\Bigr\],wherey1,…,yny\_\{1\},\\dots,y\_\{n\}are independent uniform signs, independent of thexix\_\{i\}\.

###### Theorem 7\.5\(The law with a sample\-free logarithm\)\.

There are absolute constantsc0,C0\>0c\_\{0\},C\_\{0\}\>0such that the following holds in model \(S\),d≥3d\\geq 3\. Letε∈\(0,σ2\]\\varepsilon\\in\(0,\\sigma^\{2\}\],δ∈\(0,1\)\\delta\\in\(0,1\), andΛ¯:=log⁡\(e​m¯​d3​\(2\+1/ε\)\)\\bar\{\\Lambda\}:=\\log\\bigl\(e\\bar\{m\}d^\{3\}\(2\+1/\\varepsilon\)\\bigr\)\. If

n≥C0​\[ε−2​log⁡\(8/δ\)\+ε−1​m¯​d​Λ¯\+m¯​d2​Λ¯\],n\\ \\geq\\ C\_\{0\}\\Bigl\[\\varepsilon^\{\-2\}\\log\(8/\\delta\)\\ \+\\ \\varepsilon^\{\-1\}\\bar\{m\}d\\,\\bar\{\\Lambda\}\\ \+\\ \\bar\{m\}d^\{2\}\\,\\bar\{\\Lambda\}\\Bigr\],then with probability at least1−δ1\-\\deltaeveryf∈𝒩mf\\in\\mathcal\{N\}\_\{m\}with1n​∑i\(f​\(xi\)−yi\)2≤σ2−ε\\frac\{1\}\{n\}\\sum\_\{i\}\(f\(x\_\{i\}\)\-y\_\{i\}\)^\{2\}\\leq\\sigma^\{2\}\-\\varepsilonsatisfies

Lip𝕊d−1⁡\(f\)≥c0​ε​nm¯​Λ¯\.\\operatorname\{Lip\}\_\{\\mathbb\{S\}^\{d\-1\}\}\(f\)\\ \\geq\\ c\_\{0\}\\,\\varepsilon\\,\\sqrt\{\\frac\{n\}\{\\bar\{m\}\\,\\bar\{\\Lambda\}\}\}\\,\.In particular, in the setting of Conjecture[1\.1](https://arxiv.org/html/2607.07778#S1.Thmtheorem1), oncen≥C​\[m​d2​log⁡\(e​m​d\)\+log⁡\(8/δ\)\]n\\geq C\\bigl\[md^\{2\}\\log\(emd\)\+\\log\(8/\\delta\)\\bigr\], every width\-mmnetwork with arbitrary weights fitting the data satisfiesLip𝕊d−1⁡\(f\)≥c​n/\(m​log⁡\(C​m​d\)\)\\operatorname\{Lip\}\_\{\\mathbb\{S\}^\{d\-1\}\}\(f\)\\geq c\\sqrt\{n/\(m\\log\(Cmd\)\)\}: the sample size has left the logarithm\.

## 8\.General activations at small width

Everything so far concerns piecewise\-linear activations\. The following result holds for*every*Lipschitz activation — indeed for every function factoring through an\(m\+1\)\(m\{\+\}1\)\-dimensional linear projection — and at widthm=1m=1it matches the conjectured rate exactly\. It is a projection\-capacity floor: it uses no structure ofψ\\psibeyond the factorizationf​\(x\)=g​\(P​x\)f\(x\)=g\(Px\)\.

###### Theorem 8\.1\(Any activation, small width\)\.

There are absolute constantsc,C\>0c,C\>0such that the following holds in model \(S\) withy1,…,yny\_\{1\},\\dots,y\_\{n\}i\.i\.d\. uniform signs independent of the data\. Ifn≥C​\(d​log⁡n\+m​d\)n\\geq C\\,\(d\\log n\+md\)andm≤c​min⁡\(n,d\)m\\leq c\\,\\min\(n,d\), then with probability at least1−2​e−n/C1\-2e^\{\-n/C\}, for*every*Lipschitzψ\\psiand every width\-mmnetworkf∈𝒩mf\\in\\mathcal\{N\}\_\{m\}with arbitrary weights that fits the data exactly,

Lip𝕊d−1⁡\(f\)≥c​n1/\(m\+1\)m\+1\.\\operatorname\{Lip\}\_\{\\mathbb\{S\}^\{d\-1\}\}\(f\)\\ \\geq\\ \\frac\{c\\,n^\{1/\(m\+1\)\}\}\{\\sqrt\{m\+1\}\}\\,\.In particular, atm=1m=1every exact interpolant with any Lipschitz activation satisfiesLip𝕊d−1⁡\(f\)≥c​n\\operatorname\{Lip\}\_\{\\mathbb\{S\}^\{d\-1\}\}\(f\)\\geq c\\sqrt\{n\}: Conjecture[1\.1](https://arxiv.org/html/2607.07778#S1.Thmtheorem1)holds at width one, for all activations, with no logarithmic loss\.

The floor of Theorem[8\.1](https://arxiv.org/html/2607.07778#S8.Thmtheorem1)localizes the pairing at the trivial scale\. Two refinements — localizing at the typical projection radiusp/d\\sqrt\{p/d\}, and replacing the worst\-case pigeonhole by a birthday count of random pair collisions — give a much stronger floor at small width\. The concentration step requires care: the number of collision pairs is far below the scale at which bounded\-difference inequalities are useful, and we use negative association instead\.

###### Theorem 8\.3\(Localized projection floor\)\.

There are absolute constantsc,C\>0c,C\>0such that the following holds in model \(S\) withy1,…,yny\_\{1\},\\dots,y\_\{n\}i\.i\.d\. uniform signs independent of the data\. Putp:=m\+1p:=m\+1andΛ:=log⁡\(n​d\)\\Lambda:=\\log\(nd\), and assumed≥C​pd\\geq Cp,n≥C​p​d​Λn\\geq Cp\\,d\\,\\Lambda, andp≤c​log⁡np\\leq c\\log n\. Then with probability at least1−C​e−d/C1\-Ce^\{\-d/C\}, for*every*Lipschitzψ\\psiand every width\-mmnetworkf∈𝒩mf\\in\\mathcal\{N\}\_\{m\}with arbitrary weights that fits the data exactly,

Lip𝕊d−1⁡\(f\)≥c​dp​n1/p,\\operatorname\{Lip\}\_\{\\mathbb\{S\}^\{d\-1\}\}\(f\)\\;\\geq\\;\\frac\{c\\sqrt\{d\}\}\{p\}\\,n^\{1/p\},\(5\)and, sharpening this,

Lip𝕊d−1⁡\(f\)≥c​dp​\(n2p​d​Λ\)1/p\.\\operatorname\{Lip\}\_\{\\mathbb\{S\}^\{d\-1\}\}\(f\)\\;\\geq\\;\\frac\{c\\sqrt\{d\}\}\{p\}\\Bigl\(\\frac\{n^\{2\}\}\{p\\,d\\,\\Lambda\}\\Bigr\)^\{1/p\}\.\(6\)

###### Corollary 8\.4\(Per\-width consequences\)\.

Denote byFmF\_\{m\}the right side of \([6](https://arxiv.org/html/2607.07778#S8.E6)\)\. Within the admissible rangeC​p≤d≤n/\(C​p​Λ\)Cp\\leq d\\leq n/\(Cp\\Lambda\): atm=1m=1,F1=c​n/2​ΛF\_\{1\}=cn/\\sqrt\{2\\Lambda\}— Conjecture[1\.1](https://arxiv.org/html/2607.07778#S1.Thmtheorem1)holds at width one with marginn/Λ\\sqrt\{n/\\Lambda\}, for every admissibledd, and no floor of this strength can extend tod≳nd\\gtrsim n\(Remark[8\.5](https://arxiv.org/html/2607.07778#S8.Thmtheorem5)\); atm=2m=2,F2=c​d1/6​n2/3​Λ−1/3≥c′​nF\_\{2\}=c\\,d^\{1/6\}n^\{2/3\}\\Lambda^\{\-1/3\}\\geq c^\{\\prime\}\\sqrt\{n\}oncen​d≥C​Λ2nd\\geq C\\Lambda^\{2\}— the conjecture holds at width two on the entire admissible range; atm=3m=3,F3=c​d1/4​n1/2​Λ−1/4≥c′​n/3F\_\{3\}=c\\,d^\{1/4\}n^\{1/2\}\\Lambda^\{\-1/4\}\\geq c^\{\\prime\}\\sqrt\{n/3\}onced≥C​Λd\\geq C\\Lambda; and in generalFm≥c​n/mF\_\{m\}\\geq c\\sqrt\{n/m\}exactly whend≥n\(m−3\)/\(m−1\)​γmd\\geq n^\{\(m\-3\)/\(m\-1\)\}\\gamma\_\{m\}withγm≤\(C​\(m\+1\)​Λ\)2​\(m\+2\)/\(m−1\)\\gamma\_\{m\}\\leq\(C\(m\+1\)\\Lambda\)^\{2\(m\+2\)/\(m\-1\)\}a polylogarithmic factor, the admissible window being nonempty precisely form≤c​log⁡n/log⁡log⁡nm\\leq c\\log n/\\log\\log n\.

## 9\.Toward the log\-free law

For widthsm≥c2​n/2m\\geq c^\{2\}n/2the log\-free conjecture is immediate \(Section[10](https://arxiv.org/html/2607.07778#S10): the trivial floor\), and atm=1m=1it is Theorem[8\.1](https://arxiv.org/html/2607.07778#S8.Thmtheorem1); elsewhere it remains open\. In the critical band of widths it reduces to a single sharply\-stated multiplier estimate, which we state here and leave open; the reduction, and the unconditional structure surrounding it, are developed in the supplementary note\[[13](https://arxiv.org/html/2607.07778#bib.bib13)\]\. Throughout: model \(S\), with labelsy1,…,yny\_\{1\},\\dots,y\_\{n\}i\.i\.d\. uniform on\{±1\}\\\{\\pm 1\\\}and independent of the data — the pure\-noise caseσ2=1\\sigma^\{2\}=1; every expectation𝔼y\\mathbb\{E\}\_\{y\}and every probability below is with respect to this law\. Width bandm=n/Tm=n/Twith4≤T≤logC⁡n4\\leq T\\leq\\log^\{C\}n, target Lipschitz levelL∗=c0​ε​TL^\{\\ast\}=c\_\{0\}\\varepsilon\\sqrt\{T\}, andOp:=λmax​\(∑ixi​xi⊤\)\\mathrm\{Op\}:=\\lambda\_\{\\max\}\(\\sum\_\{i\}x\_\{i\}x\_\{i\}^\{\\top\}\), which satisfiesOp≤C1\(1\+n/d\)=:Op¯\\mathrm\{Op\}\\leq C\_\{1\}\(1\+n/d\)=:\\overline\{\\mathrm\{Op\}\}with probability1−2​e−c​min⁡\(n,d\)1\-2e^\{\-c\\min\(n,d\)\}\(the rowsd​xi\\sqrt\{d\}\\,x\_\{i\}are isotropic with absolute sub\-Gaussian norm;\[[7](https://arxiv.org/html/2607.07778#bib.bib7), Thm\. 4\.6\.1\]\)\. We sayff*fits*if1n​∑i\(f​\(xi\)−yi\)2≤1−ε\\frac\{1\}\{n\}\\sum\_\{i\}\(f\(x\_\{i\}\)\-y\_\{i\}\)^\{2\}\\leq 1\-\\varepsilon, and writet∗:=min⁡\(1,320​L∗​Op¯/\(ε​T\)\)t\_\{\\ast\}:=\\min\\bigl\(1,\\,320\\,L^\{\\ast\}\\overline\{\\mathrm\{Op\}\}/\(\\varepsilon T\)\\bigr\)\.

###### Conjecture 9\.1\(Mesoscopic multiplier estimate\)\.

There are absolute constantsC,c0\>0C,c\_\{0\}\>0such that in the band, ford≥ε−2​logC⁡nd\\geq\\varepsilon^\{\-2\}\\log^\{C\}n, with probability at least1−1/n1\-1/nover the data:

𝔼y​sup∑i=1nyi​clip⁡\(⟨v,xi⟩\+c\+∑k:tk≤t∗αk​ReLU⁡\(⟨uk,xi⟩−tk\)\)≤ε​n4,\\mathbb\{E\}\_\{y\}\\ \\sup\\ \\sum\_\{i=1\}^\{n\}y\_\{i\}\\,\\operatorname\{clip\}\\Bigl\(\\langle v,x\_\{i\}\\rangle\+c\+\\\!\\\!\\sum\_\{k:\\,t\_\{k\}\\leq t\_\{\\ast\}\}\\\!\\\!\\alpha\_\{k\}\\operatorname\{ReLU\}\(\\langle u\_\{k\},x\_\{i\}\\rangle\-t\_\{k\}\)\\Bigr\)\\ \\leq\\ \\frac\{\\varepsilon n\}\{4\},the supremum over all\(v,c,\(αk,uk,tk\)\)\(v,c,\(\\alpha\_\{k\},u\_\{k\},t\_\{k\}\)\)arising as the affine\-plus\-low\-threshold part of a canonicalffwithLip𝕊d−1⁡\(f\)≤L∗\\operatorname\{Lip\}\_\{\\mathbb\{S\}^\{d\-1\}\}\(f\)\\leq L^\{\\ast\},sup𝕊d−1\|f\|≤2\+2​L∗\\sup\_\{\\mathbb\{S\}^\{d\-1\}\}\|f\|\\leq 2\+2L^\{\\ast\}, and rigidity\|αk\|​1−tk2≤2​L∗\|\\alpha\_\{k\}\|\\sqrt\{1\-t\_\{k\}^\{2\}\}\\leq 2L^\{\\ast\}\.

Every construction we have tested numerically — aimed same\-sign clusters, stacked caps, profile spikes, adaptive groups — stays at or below2​m​L∗​Op¯/t∗=ε​n/1602mL^\{\\ast\}\\overline\{\\mathrm\{Op\}\}/t\_\{\\ast\}=\\varepsilon n/160against this threshold \(numerics/check\_sector\_throttle\.py\); we record this as evidence for Conjecture[9\.1](https://arxiv.org/html/2607.07778#S9.Thmtheorem1), not a proof\. In the supplementary note\[[13](https://arxiv.org/html/2607.07778#bib.bib13)\]we prove that Conjecture[9\.1](https://arxiv.org/html/2607.07778#S9.Thmtheorem1)implies the log\-free law in the band: with probability at least1−1/n−2​e−c​min⁡\(n,d\)−e−ε2​n/1281\-1/n\-2e^\{\-c\\min\(n,d\)\}\-e^\{\-\\varepsilon^\{2\}n/128\}, no width\-mmnetwork with arbitrary weights andLip𝕊d−1⁡\(f\)≤c0​ε​n/m\\operatorname\{Lip\}\_\{\\mathbb\{S\}^\{d\-1\}\}\(f\)\\leq c\_\{0\}\\varepsilon\\sqrt\{n/m\}fitsε\\varepsilonbelow the noise floor\. The note also proves, unconditionally, the structure surrounding the estimate: fitting, occupancy, value\-mass, serving\-capacity and pile\-up lemmas; an affine supremum identity showing that the affine sector, with no bound whatsoever on its coefficients, carries onlyO​\(d\)O\(d\)of the fitting functional; the single\-direction case settled for*every*Lipschitz activation — networks of arbitrarily many units along one axis cannot fit below Lipschitz constantc​ε​min⁡\(n,d\)c\\varepsilon\\min\(\\sqrt\{n\},d\), with no width bound and no logarithm; stratified isolation of pairwise\-incoherent clusters above the coherence floor, and a counterexample showing that per\-cluster rigidity fails below it; a dimension\-free cap\-mass bound; forced\-depth and deep\-peel lemmas that remove the deep sector; and a deterministic slice computation locating exactly where label randomness becomes necessary\. Conjecture[9\.1](https://arxiv.org/html/2607.07778#S9.Thmtheorem1)remains open\.

## 10\.Sharpness and scope

#### Depth\.

Theorem[1\.2](https://arxiv.org/html/2607.07778#S1.Thmtheorem2)is a depth\-two phenomenon\. Bubeck and Sellke\[[2](https://arxiv.org/html/2607.07778#bib.bib2), Section A\]show that with a third layer, unbounded weights let a network fit generic data below the noise floor with Lipschitz constant far below the law’s threshold at the same parameter count\. So the polynomial\-boundedness assumption of\[[2](https://arxiv.org/html/2607.07778#bib.bib2)\]is necessary at depth three and, by the present paper, superfluous at depth two for kink activations: depth two is the critical depth\.

#### Activation\.

The*canonical\-parameter*mechanism of Sections[2](https://arxiv.org/html/2607.07778#S2)–[3](https://arxiv.org/html/2607.07778#S3)is specific to genuine kinks\. For a smooth activationσ\\sigmaand a unit vectoruu, the finite\-difference family

hη,u​\(x\)=σ​\(⟨u,x⟩\+η\)−σ​\(⟨u,x⟩\)ηh\_\{\\eta,u\}\(x\)=\\frac\{\\sigma\(\\langle u,x\\rangle\+\\eta\)\-\\sigma\(\\langle u,x\\rangle\)\}\{\\eta\}may remain uniformly bounded and Lipschitz asη↓0\\eta\\downarrow 0, while its natural two\-unit representation has coefficients of size1/η1/\\eta: bounded canonical parameters are unavailable for smooth activations\. The static mechanism behind the band reduction of the supplementary note\[[13](https://arxiv.org/html/2607.07778#bib.bib13)\]bypasses this at the level of*profiles*: rigidity is imposed on the derivative of the total one\-dimensional profile carried by each direction — a consequence of the ambient Lipschitz bound alone, indifferent to how the profile is represented by units — and the1/η1/\\etacoefficients never appear\. In particular the single\-direction theorems of\[[13](https://arxiv.org/html/2607.07778#bib.bib13)\]settle that case for*every*Lipschitz activation, and the reduction of Conjecture[9\.1](https://arxiv.org/html/2607.07778#S9.Thmtheorem1)in\[[13](https://arxiv.org/html/2607.07778#bib.bib13)\]applies verbatim to arbitrary Lipschitz profiles, so the remaining obstacle for general activations is the same multiplier estimate as for ReLU\. Sums of ridge functions also have nontrivial representation identities, especially when directions coalesce; see Pinkus\[[6](https://arxiv.org/html/2607.07778#bib.bib6)\]for background on ridge functions\.

#### The logarithm\.

The singlelog\\logfactor comes from the union bound over the net, as in theΩ~\\widetilde\{\\Omega\}notation of\[[2](https://arxiv.org/html/2607.07778#bib.bib2),[1](https://arxiv.org/html/2607.07778#bib.bib1)\]\. The entropy estimate here is not strong enough, by itself, to remove that logarithm:𝒜m¯,L\\mathcal\{A\}\_\{\\bar\{m\},L\}has logarithmic metric entropy of orderm¯​d​log⁡\(L/ε′\)\\bar\{m\}d\\log\(L/\\varepsilon^\{\\prime\}\)at the relevant scales\. Whether the cleanc​n/mc\\sqrt\{n/m\}holds for𝒩m\\mathcal\{N\}\_\{m\}with unbounded weights is open\.

#### Upper bounds and tightness atm≍nm\\asymp n\.

The law is tight in the parameter count:\[[2](https://arxiv.org/html/2607.07778#bib.bib2), Remark 1\.1\]constructs, for everyp∈\[Ω~​\(n\),n​\(d\+1\)\]p\\in\[\\widetilde\{\\Omega\}\(n\),n\(d\+1\)\], functions withppparameters fitting generic data withLip=O​\(n​d/p\)\\operatorname\{Lip\}=O\(\\sqrt\{nd/p\}\); atp=n​\(d\+1\)p=n\(d\+1\)the Lipschitz constant isO​\(1\)O\(1\)\. Those interpolants are not two\-layer networks\. For the overparameterized endpointm≍nm\\asymp n— the regime of the conjecture’s own thesis, one neuron per data point — an explicit two\-layer ReLU network matches the lower bound\.

###### Proposition 10\.1\(Matching upper bound atm≍nm\\asymp n\)\.

Letx1,…,xn∈𝕊d−1x\_\{1\},\\dots,x\_\{n\}\\in\\mathbb\{S\}^\{d\-1\}satisfy\|⟨xi,xj⟩\|≤18\|\\langle x\_\{i\},x\_\{j\}\\rangle\|\\leq\\tfrac\{1\}\{8\}for alli≠ji\\neq j, and lety1,…,yn∈\[−1,1\]y\_\{1\},\\dots,y\_\{n\}\\in\[\-1,1\]be arbitrary\. Puts=14s=\\tfrac\{1\}\{4\}and

ρ​\(u\):=1s​\(ReLU⁡\(u−\(1−s\)\)−ReLU⁡\(u−1\)\)\.\\rho\(u\):=\\frac\{1\}\{s\}\\bigl\(\\operatorname\{ReLU\}\(u\-\(1\-s\)\)\-\\operatorname\{ReLU\}\(u\-1\)\\bigr\)\.The width\-2​n2ntwo\-layer ReLU network

f​\(x\)=∑i=1nyi​ρ​\(⟨xi,x⟩\)f\(x\)=\\sum\_\{i=1\}^\{n\}y\_\{i\}\\rho\(\\langle x\_\{i\},x\\rangle\)satisfiesf​\(xi\)=yif\(x\_\{i\}\)=y\_\{i\}for alliiand

Lip𝕊d−1⁡\(f\)≤π2​7<5\.\\operatorname\{Lip\}\_\{\\mathbb\{S\}^\{d\-1\}\}\(f\)\\leq\\frac\{\\pi\}\{2\}\\sqrt\{7\}<5\.Fornni\.i\.d\. uniform points on𝕊d−1\\mathbb\{S\}^\{d\-1\}, the separation hypothesis holds with probability at least1−1/n1\-1/nwheneverd≥Csep​log⁡nd\\geq C\_\{\\rm sep\}\\log n, for a sufficiently large absolute constantCsepC\_\{\\rm sep\}\.

Atm=2​nm=2n, Corollary[1\.3](https://arxiv.org/html/2607.07778#S1.Thmtheorem3)forcesLip≥c1/log⁡\(C1​n​d\)\\operatorname\{Lip\}\\geq c\_\{1\}/\\sqrt\{\\log\(C\_\{1\}nd\)\}while Proposition[10\.1](https://arxiv.org/html/2607.07778#S10.Thmtheorem1)achieves an absolute Lipschitz bound\. Thus the two match up to the single logarithmic factor in the overparameterized endpoint\. So the law is sharp atm≍nm\\asymp nfor two\-layer ReLU networks, with the same log gap that Theorem[1\.2](https://arxiv.org/html/2607.07778#S1.Thmtheorem2)carries\. Whether width\-mmtwo\-layer networks achieveO~​\(n/m\)\\widetilde\{O\}\(\\sqrt\{n/m\}\)for the whole rangem≪nm\\ll nremains open; the experiments below are consistent with it up to log factors\.

#### General concentration/localization data\.

The proof of model \(G\) uses only two inputs from the Gaussian distribution: the finite\-class concentration condition in Definition[5\.1](https://arxiv.org/html/2607.07778#S5.Thmtheorem1), and the high\-probability localization eventmaxi⁡‖xi‖≤2\\max\_\{i\}\\\|x\_\{i\}\\\|\\leq 2after which the network is only tested onB2B\_\{2\}\. Consequently the Gaussian theorem extends verbatim to any distributionμ\\muonℝd\\mathbb\{R\}^\{d\}for which: \(i\) every boundedL′L^\{\\prime\}\-Lipschitz function satisfies‖h−𝔼​h‖ψ2≤κ​L′/d\\\|h\-\\mathbb\{E\}h\\\|\_\{\\psi\_\{2\}\}\\leq\\kappa L^\{\\prime\}/\\sqrt\{d\}; and \(ii\)ℙ​\(‖x‖\>2\)≤δ/\(8​n\)\\mathbb\{P\}\(\\\|x\\\|\>2\)\\leq\\delta/\(8n\)for the sample size and failure probability under consideration\. For example, ifμ\\muhas the stated Lipschitz concentration with parameterκ=C​c\\kappa=C\\sqrt\{c\}and𝔼​‖x‖≤1\\mathbb\{E\}\\\|x\\\|\\leq 1, thenℙ​\(‖x‖\>2\)≤2​exp⁡\(−d/\(C​c\)\)\\mathbb\{P\}\(\\\|x\\\|\>2\)\\leq 2\\exp\(\-d/\(Cc\)\), so the same conclusion holds onced≥C​c​log⁡\(16​n/δ\)d\\geq Cc\\log\(16n/\\delta\)\. This is the precise form in which the argument goes beyond the Gaussian measure; no additional claim about arbitrary data distributions is used in the proof\.

#### Skip connection\.

We allowed⟨v,x⟩\+c\\langle v,x\\rangle\+c; the theorem for the class of\[[1](https://arxiv.org/html/2607.07778#bib.bib1)\]follows by restriction\. The skip term is where the canonical form funnels all degenerate units \(out\-of\-domain kinks, orientation flips, constants\), which is what makes Lemma[2\.2](https://arxiv.org/html/2607.07778#S2.Thmtheorem2)exact\.

## 11\.Numerical checks

All with seed2026070720260707; scripts innumerics/\. These are sanity checks on the constructions and constants; no statement in the paper depends on them\.

*Rigidity\.*For random canonical ReLU networks ind∈\{3,5,10\}d\\in\\\{3,5,10\\\},m∈\{6,10,20\}m\\in\\\{6,10,20\\\}— including planted pairs on a common hyperplane with coefficients±106\\pm 10^\{6\}\(which canonicalization merges\) and planted near\-parallel pairs with coefficients±104\\pm 10^\{4\}\(which it does not\) — the sampled Lipschitz constantL^\\widehat\{L\}satisfiesmaxj⁡\|αj\|≤2​L^\\max\_\{j\}\|\\alpha\_\{j\}\|\\leq 2\\widehat\{L\}in all2727trials; the near\-parallel plants saturate at ratiomaxj⁡\|αj\|/\(2​L^\)=0\.50\\max\_\{j\}\|\\alpha\_\{j\}\|/\(2\\widehat\{L\}\)=0\.50, exactly the mechanism of Lemma[3\.1](https://arxiv.org/html/2607.07778#S3.Thmtheorem1)\(the gradient between the two planted hyperplanes has norm≈\|α\|\\approx\|\\alpha\|\)\. A larger sweep \(dimension up to4040, width up to8080, again with±106\\pm 10^\{6\}cancellation plants;numerics/validate\_at\_scale\.py\) passes identically\.

*Sphere factor and thed=2d=2failure\.*A single cap unit on𝕊d−1\\mathbb\{S\}^\{d\-1\}\(d=8d=8\) withα=1/1−t2\\alpha=1/\\sqrt\{1\-t^\{2\}\}has measured Lipschitz constant1\.0001\.000for eacht∈\{0,0\.5,0\.9,0\.99,0\.999\}t\\in\\\{0,0\.5,0\.9,0\.99,0\.999\\\}: the factor of Lemma[3\.2](https://arxiv.org/html/2607.07778#S3.Thmtheorem2)is exact\. On𝕊1\\mathbb\{S\}^\{1\}, the four\-unit even cycle of Proposition[3\.3](https://arxiv.org/html/2607.07778#S3.Thmtheorem3)has cancelling derivative jumps; the script displays one fixed numerical instance with ratio4\.994\.99, while the proposition gives the scalable construction proving that no absolute rigidity constant exists atd=2d=2\.

*The law\.*Sphere data,d=24d=24,n=192n=192, i\.i\.d\.±1\\pm 1labels \(σ2=1\\sigma^\{2\}=1\), widthsm∈\{24,96,384\}m\\in\\\{24,96,384\\\}\. A width\-mmReLU network with unconstrained weights is trained to mean squared error below12=σ2−ε\\tfrac\{1\}\{2\}=\\sigma^\{2\}\-\\varepsilonwhile its path norm∑k\|ak\|​‖wk‖\+‖v‖\\sum\_\{k\}\|a\_\{k\}\|\\,\\\|w\_\{k\}\\\|\+\\\|v\\\|\(an upper bound onLip\\operatorname\{Lip\}\) is penalized, so the optimizer seeks a low\-Lipschitz fitting network — the adversarial direction\. The measuredL^\\widehat\{L\}\(maximum tangential gradient norm over60006000sphere samples and the data\) exceeds the floorn/m\\sqrt\{n/m\}at every width:

mMSEL^n/mL^/n/m240\.1525\.522\.831\.95960\.1364\.881\.413\.453840\.1284\.750\.716\.72\\begin\{array\}\[\]\{rrrrr\}\\hline\\cr\\hline\\cr m&\\text\{MSE\}&\\widehat\{L\}&\\sqrt\{n/m\}&\\widehat\{L\}/\\sqrt\{n/m\}\\\\ \\hline\\cr 24&0\.152&5\.52&2\.83&1\.95\\\\ 96&0\.136&4\.88&1\.41&3\.45\\\\ 384&0\.128&4\.75&0\.71&6\.72\\\\ \\hline\\cr\\hline\\cr\\end\{array\}The measured constant stays near55while the floor falls, so the ratio grows: these trained networks satisfy the lower bound comfortably but do not realize then/m\\sqrt\{n/m\}rate form≪nm\\ll n, consistent with the matching two\-layer upper bound being open there \(the penalty is a proxy and the optimization is not run to the true minimum\)\.

*Matching upper bound atm≍nm\\asymp n\.*The width\-2​n2nconstruction of Proposition[10\.1](https://arxiv.org/html/2607.07778#S10.Thmtheorem1)was checked for\(n,d\)∈\{\(100,200\),\(200,400\),\(400,800\),\(800,1600\)\}\(n,d\)\\in\\\{\(100,200\),\(200,400\),\(400,800\),\(800,1600\)\\\}: it interpolates exactly \(maximum error2\.7⋅10−152\.7\\cdot 10^\{\-15\}\), its caps become disjoint onced≳log⁡nd\\gtrsim\\log n, and its sampled tangential\-gradient bound is7≈2\.646\\sqrt\{7\}\\approx 2\.646at every scale, while the rigorous chord\-metric Lipschitz bound in Proposition[10\.1](https://arxiv.org/html/2607.07778#S10.Thmtheorem1)is\(π/2\)​7<5\(\\pi/2\)\\sqrt\{7\}<5— flat innn, as the lower bound predicts atm≍nm\\asymp n\. A larger run \(rigidity to dimension4040and width8080; the construction ton=2000n=2000\) reproduces the same qualitative behavior\.

## Appendix AProofs

### A\.1\.Proofs for Section[1](https://arxiv.org/html/2607.07778#S1)

###### Proof of Theorem[1\.2](https://arxiv.org/html/2607.07778#S1.Thmtheorem2)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S1.Thmtheorem2)By Lemma[2\.1](https://arxiv.org/html/2607.07778#S2.Thmtheorem1)takeψ=ReLU\\psi=\\operatorname\{ReLU\}at width\(K−1\)​m≤m¯\(K\-1\)m\\leq\\bar\{m\}\. Fix

L∗:=c0​ε​nm¯​log⁡\(C0​m¯​n​d/ε\),L^\{\\ast\}:=c\_\{0\}\\,\\varepsilon\\,\\sqrt\{\\frac\{n\}\{\\bar\{m\}\\log\(C\_\{0\}\\bar\{m\}nd/\\varepsilon\)\}\},with absolutec0c\_\{0\}small andC0C\_\{0\}large, chosen below, and set

Ω:=\{∃f∈𝒩m¯:LipD⁡\(f\)≤L∗​and​1n​∑i\(f​\(xi\)−yi\)2≤σ2−ε\}\.\\Omega:=\\Bigl\\\{\\exists f\\in\\mathcal\{N\}\_\{\\bar\{m\}\}:\\ \\operatorname\{Lip\}\_\{D\}\(f\)\\leq L^\{\\ast\}\\ \\text\{and\}\\ \\tfrac\{1\}\{n\}\\textstyle\\sum\_\{i\}\(f\(x\_\{i\}\)\-y\_\{i\}\)^\{2\}\\leq\\sigma^\{2\}\-\\varepsilon\\Bigr\\\}\.OnΩc\\Omega^\{c\}every fittingffhasLipD⁡\(f\)\>L∗\\operatorname\{Lip\}\_\{D\}\(f\)\>L^\{\\ast\}, which is the theorem; so it suffices to showℙ​\(Ω\)≤δ\\mathbb\{P\}\(\\Omega\)\\leq\\delta\.

*Step 1 \(localization and clipping\)\.*In model \(G\) letπ​\(0\):=0\\pi\(0\):=0andπ​\(x\):=x​min⁡\(1,2/‖x‖\)\\pi\(x\):=x\\min\(1,2/\\\|x\\\|\)forx≠0x\\neq 0, the metric projection ofℝd\\mathbb\{R\}^\{d\}onto the convex setB2B\_\{2\}\(for‖x‖\>2\\\|x\\\|\>2it is2​x/‖x‖2x/\\\|x\\\|, the nearest point ofB2B\_\{2\}\); metric projections onto convex sets are11\-Lipschitz, andπ​\(ℝd\)=B2=D\\pi\(\\mathbb\{R\}^\{d\}\)=B\_\{2\}=D\. LetEloc:=\{maxi⁡‖xi‖≤2\}E\_\{\\mathrm\{loc\}\}:=\\\{\\max\_\{i\}\\\|x\_\{i\}\\\|\\leq 2\\\}\. Sinced​xi∼N​\(0,Id\)\\sqrt\{d\}\\,x\_\{i\}\\sim N\(0,I\_\{d\}\),‖xi‖\>2⇔‖d​xi‖2\>4​d\\\|x\_\{i\}\\\|\>2\\iff\\\|\\sqrt\{d\}\\,x\_\{i\}\\\|^\{2\}\>4d\. IfX∼χd2X\\sim\\chi^\{2\}\_\{d\}, Chernoff’s bound givesℙ​\(X≥4​d\)≤exp⁡\(−\(3−log⁡4\)​d/2\)≤e−d/2\\mathbb\{P\}\(X\\geq 4d\)\\leq\\exp\(\-\(3\-\\log 4\)d/2\)\\leq e^\{\-d/2\}; hence a union bound andd≥2​log⁡\(8​n/δ\)d\\geq 2\\log\(8n/\\delta\)giveℙ​\(Elocc\)≤n​e−d/2≤δ/8\\mathbb\{P\}\(E\_\{\\mathrm\{loc\}\}^\{c\}\)\\leq ne^\{\-d/2\}\\leq\\delta/8\. In model \(S\) setπ:=id\\pi:=\\mathrm\{id\}andEloc=E\_\{\\mathrm\{loc\}\}=the sure event\.

SupposeΩ∩Eloc\\Omega\\cap E\_\{\\mathrm\{loc\}\}occurs, witnessed byffwithL:=LipD⁡\(f\)≤L∗L:=\\operatorname\{Lip\}\_\{D\}\(f\)\\leq L^\{\\ast\}\. Allxi∈Dx\_\{i\}\\in D, so by Lemma[3\.5](https://arxiv.org/html/2607.07778#S3.Thmtheorem5)\(diam⁡D≤4\\operatorname\{diam\}D\\leq 4\) we getsupD\|f\|≤B0=2\+4​L∗\\sup\_\{D\}\|f\|\\leq B\_\{0\}=2\+4L^\{\\ast\}, hencef\|D∈𝒜m¯,L∗f\|\_\{D\}\\in\\mathcal\{A\}\_\{\\bar\{m\},L^\{\\ast\}\}\.

*Step 2 \(net\)\.*LetS⊂𝒜m¯,L∗S\\subset\\mathcal\{A\}\_\{\\bar\{m\},L^\{\\ast\}\}be an internalε32\\tfrac\{\\varepsilon\}\{32\}\-net as in Proposition[4\.1](https://arxiv.org/html/2607.07778#S4.Thmtheorem1):

log⁡\|S\|≤C​m¯​d​log⁡\(64​C​m¯​d​\(2\+L∗\)ε\)\.\\log\|S\|\\leq C\\bar\{m\}d\\log\\\!\\Bigl\(\\frac\{64\\,C\\,\\bar\{m\}d\(2\+L^\{\\ast\}\)\}\{\\varepsilon\}\\Bigr\)\.Define the finite classℱ:=\{clip∘h∘π:h∈S\}\\mathcal\{F\}:=\\\{\\operatorname\{clip\}\\circ h\\circ\\pi:h\\in S\\\}\. Each member is defined onsupp⁡μ\\operatorname\{supp\}\\mu, has values in\[−1,1\]\[\-1,1\], and isL∗L^\{\\ast\}\-Lipschitz \(composition of the11\-Lipschitzπ\\piintoDD, anL∗L^\{\\ast\}\-Lipschitz\-on\-DDfunction, and the11\-Lipschitzclip\\operatorname\{clip\}\)\. Pickh∈Sh\\in Swith‖h−f‖L∞​\(D\)≤ε32\\\|h\-f\\\|\_\{L^\{\\infty\}\(D\)\}\\leq\\tfrac\{\\varepsilon\}\{32\}, and setf^:=clip∘h∘π∈ℱ\\hat\{f\}:=\\operatorname\{clip\}\\circ h\\circ\\pi\\in\\mathcal\{F\}andf~:=clip∘f∘π\\tilde\{f\}:=\\operatorname\{clip\}\\circ f\\circ\\pi\. Then‖f^−f~‖∞≤ε32\\\|\\hat\{f\}\-\\tilde\{f\}\\\|\_\{\\infty\}\\leq\\tfrac\{\\varepsilon\}\{32\}\(both equalclip∘\(⋅\)∘π\\operatorname\{clip\}\\circ\(\\cdot\)\\circ\\piof functions withinε32\\tfrac\{\\varepsilon\}\{32\}onDD, andclip\\operatorname\{clip\}contracts\)\. OnElocE\_\{\\mathrm\{loc\}\},π​\(xi\)=xi∈D\\pi\(x\_\{i\}\)=x\_\{i\}\\in D, sof~​\(xi\)=clip⁡\(f​\(xi\)\)\\tilde\{f\}\(x\_\{i\}\)=\\operatorname\{clip\}\(f\(x\_\{i\}\)\); since\|yi\|≤1\|y\_\{i\}\|\\leq 1, clippingf​\(xi\)f\(x\_\{i\}\)toward\[−1,1\]∋yi\[\-1,1\]\\ni y\_\{i\}can only decrease the error:

\(f~​\(xi\)−yi\)2≤\(f​\(xi\)−yi\)2,\(\\tilde\{f\}\(x\_\{i\}\)\-y\_\{i\}\)^\{2\}\\leq\(f\(x\_\{i\}\)\-y\_\{i\}\)^\{2\},sof~\\tilde\{f\}also fits at levelσ2−ε\\sigma^\{2\}\-\\varepsilon\. Withai:=\|f~​\(xi\)−yi\|≤2a\_\{i\}:=\|\\tilde\{f\}\(x\_\{i\}\)\-y\_\{i\}\|\\leq 2,

1n​∑i\(f^​\(xi\)−yi\)2≤1n​∑i\(ai\+ε32\)2≤\(σ2−ε\)\+2⋅2​ε32\+ε21024≤σ2−ε2\.\\frac\{1\}\{n\}\\sum\_\{i\}\(\\hat\{f\}\(x\_\{i\}\)\-y\_\{i\}\)^\{2\}\\leq\\frac\{1\}\{n\}\\sum\_\{i\}\\Bigl\(a\_\{i\}\+\\frac\{\\varepsilon\}\{32\}\\Bigr\)^\{2\}\\leq\(\\sigma^\{2\}\-\\varepsilon\)\+\\frac\{2\\cdot 2\\varepsilon\}\{32\}\+\\frac\{\\varepsilon^\{2\}\}\{1024\}\\leq\\sigma^\{2\}\-\\frac\{\\varepsilon\}\{2\}\.HenceΩ∩Eloc⊂Ω′:=\{∃f^∈ℱ:1n​∑i\(f^​\(xi\)−yi\)2≤σ2−ε2\}\\Omega\\cap E\_\{\\mathrm\{loc\}\}\\subset\\Omega^\{\\prime\}:=\\\{\\exists\\hat\{f\}\\in\\mathcal\{F\}:\\tfrac\{1\}\{n\}\\sum\_\{i\}\(\\hat\{f\}\(x\_\{i\}\)\-y\_\{i\}\)^\{2\}\\leq\\sigma^\{2\}\-\\tfrac\{\\varepsilon\}\{2\}\\\}\.

*Step 3 \(union bound and constants\)\.*Apply Theorem[5\.6](https://arxiv.org/html/2607.07778#S5.Thmtheorem6)toℱ\\mathcal\{F\}\(its members areL∗L^\{\\ast\}\-Lipschitz and\[−1,1\]\[\-1,1\]\-valued; a member whose true Lipschitz constant is belowL∗L^\{\\ast\}satisfies the hypothesis a fortiori, and takingL=L∗L=L^\{\\ast\}in the exponent below is the weakest admissible choice\) at levelε/2\\varepsilon/2:

ℙ​\(Ω′\)≤4​e−n​ε2/2048\+exp⁡\(log⁡\|S\|−c2​n​ε24​max⁡\(1,dκ2​\(L∗\)2\)\)\.\\mathbb\{P\}\(\\Omega^\{\\prime\}\)\\leq 4e^\{\-n\\varepsilon^\{2\}/2048\}\+\\exp\\Bigl\(\\log\|S\|\-\\frac\{c\_\{2\}n\\varepsilon^\{2\}\}\{4\}\\max\\bigl\(1,\\tfrac\{d\}\{\\kappa^\{2\}\(L^\{\\ast\}\)^\{2\}\}\\bigr\)\\Bigr\)\.The first term is≤δ/2\\leq\\delta/2whenn≥C0​ε−2​log⁡\(8/δ\)n\\geq C\_\{0\}\\varepsilon^\{\-2\}\\log\(8/\\delta\)withC0≥2048C\_\{0\}\\geq 2048\. For the second, useL∗≤c0​ε​n≤ε​nL^\{\\ast\}\\leq c\_\{0\}\\varepsilon\\sqrt\{n\}\\leq\\varepsilon\\sqrt\{n\}to collapse the logarithm:2\+L∗≤3​n2\+L^\{\\ast\}\\leq 3\\sqrt\{n\}, solog⁡\(64​C​m¯​d​\(2\+L∗\)/ε\)≤log⁡\(C0​m¯​n​d/ε\)\\log\(64C\\bar\{m\}d\(2\+L^\{\\ast\}\)/\\varepsilon\)\\leq\\log\(C\_\{0\}\\bar\{m\}nd/\\varepsilon\)forC0C\_\{0\}large, whence

log\|S\|≤C′m¯dlog\(C0​m¯​n​dε\)=:ℰ\.\\log\|S\|\\leq C^\{\\prime\}\\bar\{m\}d\\log\\\!\\Bigl\(\\frac\{C\_\{0\}\\bar\{m\}nd\}\{\\varepsilon\}\\Bigr\)=:\\mathcal\{E\}\.We claim, withc02≤c2/\(8​κ2​C′\)c\_\{0\}^\{2\}\\leq c\_\{2\}/\(8\\kappa^\{2\}C^\{\\prime\}\)\(and hence alsoc02≤c2/\(8​C′\)c\_\{0\}^\{2\}\\leq c\_\{2\}/\(8C^\{\\prime\}\), asκ≥1\\kappa\\geq 1WLOG\),

ℰ≤c28​n​ε2​max⁡\(1,dκ2​\(L∗\)2\)\.\\mathcal\{E\}\\ \\leq\\ \\frac\{c\_\{2\}\}\{8\}\\,n\\varepsilon^\{2\}\\,\\max\\bigl\(1,\\tfrac\{d\}\{\\kappa^\{2\}\(L^\{\\ast\}\)^\{2\}\}\\bigr\)\.\(7\)Two cases, and we substitute\(L∗\)2=c02​ε2​n/\(m¯​log⁡\(C0​m¯​n​d/ε\)\)\(L^\{\\ast\}\)^\{2\}=c\_\{0\}^\{2\}\\varepsilon^\{2\}n/\(\\bar\{m\}\\log\(C\_\{0\}\\bar\{m\}nd/\\varepsilon\)\)in each\.

- *Case A*\(L∗≤d/κL^\{\\ast\}\\leq\\sqrt\{d\}/\\kappa, so themax\\maxisdκ2​\(L∗\)2\\tfrac\{d\}\{\\kappa^\{2\}\(L^\{\\ast\}\)^\{2\}\}\)\. The right side of \([7](https://arxiv.org/html/2607.07778#A1.E7)\) is c28⋅n​ε2​dκ2​\(L∗\)2=c28​κ2⋅n​ε2​dc02​ε2​n/\(m¯​log⁡\(⋯\)\)=c28​κ2​c02​m¯​d​log⁡\(C0​m¯​n​dε\)\.\\frac\{c\_\{2\}\}\{8\}\\cdot\\frac\{n\\varepsilon^\{2\}d\}\{\\kappa^\{2\}\(L^\{\\ast\}\)^\{2\}\}=\\frac\{c\_\{2\}\}\{8\\kappa^\{2\}\}\\cdot\\frac\{n\\varepsilon^\{2\}d\}\{c\_\{0\}^\{2\}\\varepsilon^\{2\}n/\(\\bar\{m\}\\log\(\\cdots\)\)\}=\\frac\{c\_\{2\}\}\{8\\kappa^\{2\}c\_\{0\}^\{2\}\}\\,\\bar\{m\}d\\log\\\!\\Bigl\(\\frac\{C\_\{0\}\\bar\{m\}nd\}\{\\varepsilon\}\\Bigr\)\.Sincec02≤c2/\(8​κ2​C′\)c\_\{0\}^\{2\}\\leq c\_\{2\}/\(8\\kappa^\{2\}C^\{\\prime\}\), this is≥C′​m¯​d​log⁡\(⋯\)=ℰ\\geq C^\{\\prime\}\\bar\{m\}d\\log\(\\cdots\)=\\mathcal\{E\}, which is \([7](https://arxiv.org/html/2607.07778#A1.E7)\) in this case\.
- *Case B*\(L∗\>d/κL^\{\\ast\}\>\\sqrt\{d\}/\\kappa, so themax\\maxis11\)\. Right side=c28​n​ε2=\\tfrac\{c\_\{2\}\}\{8\}n\\varepsilon^\{2\}\. The case hypothesisL∗\>d/κL^\{\\ast\}\>\\sqrt\{d\}/\\kappameansc02​ε2​n/\(m¯​log⁡\(⋯\)\)\>d/κ2c\_\{0\}^\{2\}\\varepsilon^\{2\}n/\(\\bar\{m\}\\log\(\\cdots\)\)\>d/\\kappa^\{2\}, i\.e\. cross\-multiplying, m¯​d​log⁡\(C0​m¯​n​dε\)<c02​κ2​n​ε2\.\\bar\{m\}d\\log\\\!\\Bigl\(\\frac\{C\_\{0\}\\bar\{m\}nd\}\{\\varepsilon\}\\Bigr\)<c\_\{0\}^\{2\}\\kappa^\{2\}\\,n\\varepsilon^\{2\}\.Henceℰ=C′​m¯​d​log⁡\(⋯\)<C′​c02​κ2​n​ε2≤c28​n​ε2\\mathcal\{E\}=C^\{\\prime\}\\bar\{m\}d\\log\(\\cdots\)<C^\{\\prime\}c\_\{0\}^\{2\}\\kappa^\{2\}n\\varepsilon^\{2\}\\leq\\tfrac\{c\_\{2\}\}\{8\}n\\varepsilon^\{2\}, again byc02≤c2/\(8​κ2​C′\)c\_\{0\}^\{2\}\\leq c\_\{2\}/\(8\\kappa^\{2\}C^\{\\prime\}\), giving \([7](https://arxiv.org/html/2607.07778#A1.E7)\)\.

WriteM:=max⁡\(1,dκ2​\(L∗\)2\)≥1M:=\\max\(1,\\tfrac\{d\}\{\\kappa^\{2\}\(L^\{\\ast\}\)^\{2\}\}\)\\geq 1\. By \([7](https://arxiv.org/html/2607.07778#A1.E7)\),c28​n​ε2​M≥ℰ\\tfrac\{c\_\{2\}\}\{8\}n\\varepsilon^\{2\}M\\geq\\mathcal\{E\}, soc24​n​ε2​M=c28​n​ε2​M\+c28​n​ε2​M≥ℰ\+c28​n​ε2​M\\tfrac\{c\_\{2\}\}\{4\}n\\varepsilon^\{2\}M=\\tfrac\{c\_\{2\}\}\{8\}n\\varepsilon^\{2\}M\+\\tfrac\{c\_\{2\}\}\{8\}n\\varepsilon^\{2\}M\\geq\\mathcal\{E\}\+\\tfrac\{c\_\{2\}\}\{8\}n\\varepsilon^\{2\}M\. Sincelog⁡\|S\|≤ℰ\\log\|S\|\\leq\\mathcal\{E\}, the exponent is

log⁡\|S\|−c24​n​ε2​M≤ℰ−ℰ−c28​n​ε2​M=−c28​n​ε2​M≤−c28​n​ε2\.\\log\|S\|\-\\frac\{c\_\{2\}\}\{4\}n\\varepsilon^\{2\}M\\ \\leq\\ \\mathcal\{E\}\-\\mathcal\{E\}\-\\frac\{c\_\{2\}\}\{8\}n\\varepsilon^\{2\}M\\ =\\ \-\\frac\{c\_\{2\}\}\{8\}n\\varepsilon^\{2\}M\\ \\leq\\ \-\\frac\{c\_\{2\}\}\{8\}n\\varepsilon^\{2\}\.Hence the second term is≤exp⁡\(−c28​n​ε2\)≤δ/4\\leq\\exp\(\-\\tfrac\{c\_\{2\}\}\{8\}n\\varepsilon^\{2\}\)\\leq\\delta/4wheneverc28​n​ε2≥log⁡\(4/δ\)\\tfrac\{c\_\{2\}\}\{8\}n\\varepsilon^\{2\}\\geq\\log\(4/\\delta\), i\.e\. whenevern≥8c2​ε−2​log⁡\(4/δ\)n\\geq\\tfrac\{8\}\{c\_\{2\}\}\\varepsilon^\{\-2\}\\log\(4/\\delta\), which holds undern≥C0​ε−2​log⁡\(8/δ\)n\\geq C\_\{0\}\\varepsilon^\{\-2\}\\log\(8/\\delta\)onceC0≥8/c2C\_\{0\}\\geq 8/c\_\{2\}\. Altogether

ℙ​\(Ω\)≤ℙ​\(Elocc\)\+ℙ​\(Ω′\)≤δ8\+δ2\+δ4≤δ\.∎\\mathbb\{P\}\(\\Omega\)\\leq\\mathbb\{P\}\(E\_\{\\mathrm\{loc\}\}^\{c\}\)\+\\mathbb\{P\}\(\\Omega^\{\\prime\}\)\\leq\\frac\{\\delta\}\{8\}\+\\frac\{\\delta\}\{2\}\+\\frac\{\\delta\}\{4\}\\leq\\delta\.\\qed

###### Proof of Corollary[1\.3](https://arxiv.org/html/2607.07778#S1.Thmtheorem3)\.

Withyiy\_\{i\}independent ofxix\_\{i\}:g≡0g\\equiv 0,zi=yiz\_\{i\}=y\_\{i\},σ2=𝔼​Var⁡\(y∣x\)=1\\sigma^\{2\}=\\mathbb\{E\}\\operatorname\{Var\}\(y\\mid x\)=1; takeε=12\\varepsilon=\\tfrac\{1\}\{2\}\. Exact fitting gives mean squared error0≤σ2−ε0\\leq\\sigma^\{2\}\-\\varepsilon, and “error≤12=σ2−12\\leq\\tfrac\{1\}\{2\}=\\sigma^\{2\}\-\\tfrac\{1\}\{2\}” is the stated relaxation\. Apply Theorem[1\.2](https://arxiv.org/html/2607.07778#S1.Thmtheorem2)in model \(S\),d≥3d\\geq 3, and absorbm¯≤K​m\\bar\{m\}\\leq Kminto the constants for fixedKK; for ReLUm¯=m\+1\\bar\{m\}=m\+1\. ∎

###### Proof of Corollary[1\.4](https://arxiv.org/html/2607.07778#S1.Thmtheorem4)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S1.Thmtheorem4)Apply Theorem[1\.2](https://arxiv.org/html/2607.07778#S1.Thmtheorem2)to each widthm=1,…,Mm=1,\\dots,Mwith failure probabilityδ/M\\delta/M, and take a union bound\. The lower\-bound formula itself is unchanged because the threshold in Theorem[1\.2](https://arxiv.org/html/2607.07778#S1.Thmtheorem2)does not depend onδ\\delta; only the sample\-size and Gaussian\-localization hypotheses acquire the factorMMinside the logarithm\. ∎

###### Proof of Theorem[1\.5](https://arxiv.org/html/2607.07778#S1.Thmtheorem5)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S1.Thmtheorem5)For an integerq≥1q\\geq 1and a numberL\>0L\>0, putB0​\(L\):=2\+4​LB\_\{0\}\(L\):=2\+4Land let

𝒜L\(q\):=\{f\|D:\\displaystyle\\mathcal\{A\}^\{\(q\)\}\_\{L\}=\\\{f\|\_\{D\}:f​is a two\-layer piecewise\-linear realized function,\\displaystyle f\\ \\text\{is a two\-layer piecewise\-linear realized function\},k\(f\)\+1≤q,LipD\(f\)≤L,supD\|f\|≤B0\(L\)\}\.\\displaystyle k\(f\)\+1\\leq q,\\quad\\operatorname\{Lip\}\_\{D\}\(f\)\\leq L,\\quad\\sup\_\{D\}\|f\|\\leq B\_\{0\}\(L\)\\\}\.By Lemma[2\.2](https://arxiv.org/html/2607.07778#S2.Thmtheorem2), every member of this class has a canonical form with at mostq−1q\-1ReLU kinks\. Proposition[4\.1](https://arxiv.org/html/2607.07778#S4.Thmtheorem1), used with width budgetqq\(the harmless extra unit also covers the purely affine case\), gives

logN\(𝒜L\(q\),∥⋅∥∞,η\)≤Cqdlog\(C​q​d​\(2\+L\)η\),\\log N\(\\mathcal\{A\}^\{\(q\)\}\_\{L\},\\\|\\cdot\\\|\_\{\\infty\},\\eta\)\\leq Cqd\\log\\\!\\Bigl\(\\frac\{Cqd\(2\+L\)\}\{\\eta\}\\Bigr\),with an internal net of the same size up to constants\.

In model \(G\), first remove the single localization eventElocc=\{maxi⁡‖xi‖\>2\}E\_\{\\mathrm\{loc\}\}^\{c\}=\\\{\\max\_\{i\}\\\|x\_\{i\}\\\|\>2\\\}; the stated conditiond≥2​log⁡\(8​n/δ\)d\\geq 2\\log\(8n/\\delta\)givesℙ​\(Elocc\)≤δ/8\\mathbb\{P\}\(E\_\{\\mathrm\{loc\}\}^\{c\}\)\\leq\\delta/8, exactly as in the proof of Theorem[1\.2](https://arxiv.org/html/2607.07778#S1.Thmtheorem2)\. The remaining union over kink counts is performed onElocE\_\{\\mathrm\{loc\}\}\.

Fixq∈\{1,…,n\+1\}q\\in\\\{1,\\dots,n\+1\\\}and set

Lq∗:=c0​ε​nq​log⁡\(C0​q​n​d/ε\)\.L\_\{q\}^\{\\ast\}:=c\_\{0\}\\varepsilon\\sqrt\{\\frac\{n\}\{q\\log\(C\_\{0\}qnd/\\varepsilon\)\}\}\.Repeating the proof of Theorem[1\.2](https://arxiv.org/html/2607.07778#S1.Thmtheorem2), with𝒜m¯,L∗\\mathcal\{A\}\_\{\\bar\{m\},L^\{\\ast\}\}replaced by𝒜Lq∗\(q\)\\mathcal\{A\}^\{\(q\)\}\_\{L\_\{q\}^\{\\ast\}\}and with failure budgetδq:=δ/\(2​q​\(q\+1\)\)\\delta\_\{q\}:=\\delta/\(2q\(q\+1\)\), shows that, apart from the already\-separated localization event, the probability of

∃f:k​\(f\)\+1≤q,LipD⁡\(f\)≤Lq∗,1n​∑i\(f​\(xi\)−yi\)2≤σ2−ε\\exists f:\\ k\(f\)\+1\\leq q,\\quad\\operatorname\{Lip\}\_\{D\}\(f\)\\leq L\_\{q\}^\{\\ast\},\\quad\\frac\{1\}\{n\}\\sum\_\{i\}\(f\(x\_\{i\}\)\-y\_\{i\}\)^\{2\}\\leq\\sigma^\{2\}\-\\varepsilonis at mostδq\\delta\_\{q\}, provided

n≥C​ε−2​log⁡\(8/δq\)\.n\\geq C\\varepsilon^\{\-2\}\\log\(8/\\delta\_\{q\}\)\.Forq≤n\+1q\\leq n\+1, this follows fromn≥C0​ε−2​log⁡\(8​n/δ\)n\\geq C\_\{0\}\\varepsilon^\{\-2\}\\log\(8n/\\delta\)after increasing the absolute constantC0C\_\{0\}, sincelog⁡\(8/δq\)≤C​log⁡\(8​n/δ\)\\log\(8/\\delta\_\{q\}\)\\leq C\\log\(8n/\\delta\)\.

Summing overq=1,…,n\+1q=1,\\dots,n\+1gives total non\-localization failure probability at most

∑q=1n\+1δ2​q​\(q\+1\)<δ2\.\\sum\_\{q=1\}^\{n\+1\}\\frac\{\\delta\}\{2q\(q\+1\)\}<\\frac\{\\delta\}\{2\}\.Together with the localization failure probabilityδ/8\\delta/8, this is still less thanδ\\delta\. On the complementary event, takeq=k​\(f\)\+1q=k\(f\)\+1for any fitting function withk​\(f\)≤nk\(f\)\\leq n; the displayed lower bound is exactly the asserted one\. ∎

###### Proof of Corollary[1\.6](https://arxiv.org/html/2607.07778#S1.Thmtheorem6)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S1.Thmtheorem6)The input\-output map of any single\-hidden\-layer piecewise\-linear ridge architecture is a finite sum of terms of the forma​ψ​\(⟨w,x⟩\+b\)a\\psi\(\\langle w,x\\rangle\+b\)plus an affine part, possibly with constraints or identifications among the allowedww’s\. Such constraints can only reduce the class\. If the realized function has at mostK0K\_\{0\}distinct canonical kink hyperplanes, Lemma[2\.2](https://arxiv.org/html/2607.07778#S2.Thmtheorem2)writes it with at mostK0K\_\{0\}ReLU kink units on the domain\. The proof of Theorem[1\.2](https://arxiv.org/html/2607.07778#S1.Thmtheorem2), with the entropy bound read at width budgetK0\+1K\_\{0\}\+1, gives the stated fixed\-K0K\_\{0\}conclusion\. WhenK0≤nK\_\{0\}\\leq nand the sample size meets the hypothesis of Theorem[1\.5](https://arxiv.org/html/2607.07778#S1.Thmtheorem5), that theorem gives the simultaneous realized\-kink version on its own event\. A convolutional layer withFFfilters evaluated atSSpositions has at mostF​SFSridge preactivations, and aKK\-piece activation contributes at mostK−1K\-1kink hyperplanes per preactivation, soK0≤\(K−1\)​F​SK\_\{0\}\\leq\(K\-1\)FS\. ∎

###### Proof of Corollary[1\.7](https://arxiv.org/html/2607.07778#S1.Thmtheorem7)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S1.Thmtheorem7)Run the scalar theorem for each coordinate whose noise level satisfiesσℓ2≥ε/r\\sigma\_\{\\ell\}^\{2\}\\geq\\varepsilon/r, with accuracy parameterε/r\\varepsilon/rand failure probabilityδ/r\\delta/r, and intersect the resulting events\. Coordinates withσℓ2<ε/r\\sigma\_\{\\ell\}^\{2\}<\\varepsilon/rcannot be responsible for an empirical improvement of sizeε/r\\varepsilon/r, because their empirical squared error is nonnegative\. On the intersection event, every scalar coordinate function fitting its own coordinate labels at leastε/r\\varepsilon/rbelow its coordinate noise floor obeys the displayed scalar lower bound\.

Now suppose a vector\-valuedffviolates the conclusion while fitting the vector labelsε\\varepsilonbelow the total noise floor\. Write

Fitℓ=1n​∑i\(fℓ​\(xi\)−yi​ℓ\)2\.\\mathrm\{Fit\}\_\{\\ell\}=\\frac\{1\}\{n\}\\sum\_\{i\}\(f\_\{\\ell\}\(x\_\{i\}\)\-y\_\{i\\ell\}\)^\{2\}\.The hypothesis gives

∑ℓ=1r\(σℓ2−Fitℓ\)≥ε,\\sum\_\{\\ell=1\}^\{r\}\(\\sigma\_\{\\ell\}^\{2\}\-\\mathrm\{Fit\}\_\{\\ell\}\)\\geq\\varepsilon,so for some coordinateℓ\\ellone hasFitℓ≤σℓ2−ε/r\\mathrm\{Fit\}\_\{\\ell\}\\leq\\sigma\_\{\\ell\}^\{2\}\-\\varepsilon/r\. This coordinate uses at mostK0K\_\{0\}of the distinct kink hyperplanes used by the whole vector map\. The scalar bound applied tofℓf\_\{\\ell\}gives the displayed lower bound forLipD⁡\(fℓ\)\\operatorname\{Lip\}\_\{D\}\(f\_\{\\ell\}\)\. Finally,

LipD⁡\(f\)=supx≠x′‖f​\(x\)−f​\(x′\)‖2‖x−x′‖≥supx≠x′\|fℓ​\(x\)−fℓ​\(x′\)\|‖x−x′‖=LipD⁡\(fℓ\),\\operatorname\{Lip\}\_\{D\}\(f\)=\\sup\_\{x\\neq x^\{\\prime\}\}\\frac\{\\\|f\(x\)\-f\(x^\{\\prime\}\)\\\|\_\{2\}\}\{\\\|x\-x^\{\\prime\}\\\|\}\\geq\\sup\_\{x\\neq x^\{\\prime\}\}\\frac\{\|f\_\{\\ell\}\(x\)\-f\_\{\\ell\}\(x^\{\\prime\}\)\|\}\{\\\|x\-x^\{\\prime\}\\\|\}=\\operatorname\{Lip\}\_\{D\}\(f\_\{\\ell\}\),so the same bound holds for the vector map\. ∎

### A\.2\.Proofs for Section[2](https://arxiv.org/html/2607.07778#S2)

###### Proof of Lemma[2\.1](https://arxiv.org/html/2607.07778#S2.Thmtheorem1)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S2.Thmtheorem1)Call the right\-hand side of \([2](https://arxiv.org/html/2607.07778#S2.E2)\)R​\(t\)R\(t\)\. Bothψ\\psiandRRare continuous \(eachReLU\(⋅−τκ\)\\operatorname\{ReLU\}\(\\cdot\-\\tau\_\{\\kappa\}\)is continuous\) and piecewise linear with breakpoints contained in\{τ1,…,τK−1\}\\\{\\tau\_\{1\},\\dots,\\tau\_\{K\-1\}\\\}\. Two continuous piecewise\-linear functions with breakpoints in a common finite set coincide everywhere as soon as they agree at one point and have equal slopes on every piece\. They agree att=τ1t=\\tau\_\{1\}: thereR​\(τ1\)=ψ​\(τ1\)\+0\+0=ψ​\(τ1\)R\(\\tau\_\{1\}\)=\\psi\(\\tau\_\{1\}\)\+0\+0=\\psi\(\\tau\_\{1\}\), sinceReLU⁡\(τ1−τκ\)=0\\operatorname\{ReLU\}\(\\tau\_\{1\}\-\\tau\_\{\\kappa\}\)=0forκ≥1\\kappa\\geq 1\(asτ1≤τκ\\tau\_\{1\}\\leq\\tau\_\{\\kappa\}\) and the linear term vanishes\. On\(−∞,τ1\)\(\-\\infty,\\tau\_\{1\}\)everyReLU⁡\(t−τκ\)=0\\operatorname\{ReLU\}\(t\-\\tau\_\{\\kappa\}\)=0, soRRhas slopes0s\_\{0\}, matchingψ\\psi\. On\(τκ,τκ\+1\)\(\\tau\_\{\\kappa\},\\tau\_\{\\kappa\+1\}\)the active kinks are exactlyτ1,…,τκ\\tau\_\{1\},\\dots,\\tau\_\{\\kappa\}, soRRhas slopes0\+∑j=1κ\(sj−sj−1\)=sκs\_\{0\}\+\\sum\_\{j=1\}^\{\\kappa\}\(s\_\{j\}\-s\_\{j\-1\}\)=s\_\{\\kappa\}, matchingψ\\psi\. HenceR≡ψR\\equiv\\psi\.

Now apply \([2](https://arxiv.org/html/2607.07778#S2.E2)\) witht=⟨wk,x⟩\+bkt=\\langle w\_\{k\},x\\rangle\+b\_\{k\}to each unit offf:ak​ψ​\(⟨wk,x⟩\+bk\)a\_\{k\}\\psi\(\\langle w\_\{k\},x\\rangle\+b\_\{k\}\)becomes an affine function ofxxplus∑κ=1K−1ak​\(sκ−sκ−1\)​ReLU⁡\(⟨wk,x⟩\+bk−τκ\)\\sum\_\{\\kappa=1\}^\{K\-1\}a\_\{k\}\(s\_\{\\kappa\}\-s\_\{\\kappa\-1\}\)\\operatorname\{ReLU\}\(\\langle w\_\{k\},x\\rangle\+b\_\{k\}\-\\tau\_\{\\kappa\}\), i\.e\.K−1K\-1ReLU units with the shifted biasesbk−τκb\_\{k\}\-\\tau\_\{\\kappa\}\. Summing overkkand absorbing all the affine terms into⟨v,x⟩\+c\\langle v,x\\rangle\+cproduces a ReLU network of width≤\(K−1\)​m\\leq\(K\-1\)mequal toffeverywhere\. ∎

###### Proof of Lemma[2\.2](https://arxiv.org/html/2607.07778#S2.Thmtheorem2)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S2.Thmtheorem2)We transform the units offfone at a time; every operation preserves the value offfonDD\.

*Step 1 \(constant units\)\.*A unit withwk=0w\_\{k\}=0is the constantak​ReLU⁡\(bk\)a\_\{k\}\\operatorname\{ReLU\}\(b\_\{k\}\); move it intocc\.

*Step 2 \(normalization\)\.*Forwk≠0w\_\{k\}\\neq 0, writeu=wk/‖wk‖u=w\_\{k\}/\\\|w\_\{k\}\\\|,t=−bk/‖wk‖t=\-b\_\{k\}/\\\|w\_\{k\}\\\|,α=ak​‖wk‖\\alpha=a\_\{k\}\\\|w\_\{k\}\\\|\. SinceReLU⁡\(λ​z\)=λ​ReLU⁡\(z\)\\operatorname\{ReLU\}\(\\lambda z\)=\\lambda\\operatorname\{ReLU\}\(z\)forλ\>0\\lambda\>0and⟨wk,x⟩\+bk=‖wk‖​\(⟨u,x⟩−t\)\\langle w\_\{k\},x\\rangle\+b\_\{k\}=\\\|w\_\{k\}\\\|\(\\langle u,x\\rangle\-t\), the unit equalsα​ReLU⁡\(⟨u,x⟩−t\)\\alpha\\operatorname\{ReLU\}\(\\langle u,x\\rangle\-t\)with‖u‖=1\\\|u\\\|=1\.

*Step 3 \(orientation, via the reflection identity\)\.*The identity

ReLU⁡\(−z\)=ReLU⁡\(z\)−z\\operatorname\{ReLU\}\(\-z\)=\\operatorname\{ReLU\}\(z\)\-z\(8\)\(true becausemax⁡\(−z,0\)−max⁡\(z,0\)=−z\\max\(\-z,0\)\-\\max\(z,0\)=\-z\) lets us replace\(u,t\)\(u,t\)by\(−u,−t\)\(\-u,\-t\)at the cost of an affine term:

α​ReLU⁡\(⟨u,x⟩−t\)=α​ReLU⁡\(⟨−u,x⟩−\(−t\)\)\+α​\(⟨u,x⟩−t\),\\alpha\\operatorname\{ReLU\}\(\\langle u,x\\rangle\-t\)=\\alpha\\operatorname\{ReLU\}\\bigl\(\\langle\-u,x\\rangle\-\(\-t\)\\bigr\)\+\\alpha\\bigl\(\\langle u,x\\rangle\-t\\bigr\),the last summand being absorbed into⟨v,x⟩\+c\\langle v,x\\rangle\+c\. We use \([8](https://arxiv.org/html/2607.07778#A1.E8)\) to enforce an orientation convention below\.

*Step 4 \(units whose kink misses the domain\)\.*Ball case: ift≥Rt\\geq Rthen⟨u,x⟩−t≤‖x‖−t≤R−t≤0\\langle u,x\\rangle\-t\\leq\\\|x\\\|\-t\\leq R\-t\\leq 0onBRB\_\{R\}with equality only where‖x‖=R\\\|x\\\|=Randx=R​ux=Ru, soReLU⁡\(⟨u,x⟩−t\)≡0\\operatorname\{ReLU\}\(\\langle u,x\\rangle\-t\)\\equiv 0onBRB\_\{R\}; drop it\. Ift≤−Rt\\leq\-Rthen⟨u,x⟩−t≥0\\langle u,x\\rangle\-t\\geq 0onBRB\_\{R\}, soReLU⁡\(⟨u,x⟩−t\)=⟨u,x⟩−t\\operatorname\{ReLU\}\(\\langle u,x\\rangle\-t\)=\\langle u,x\\rangle\-tis affine there; absorb it\. Sphere case: apply \([8](https://arxiv.org/html/2607.07778#A1.E8)\) to maket≥0t\\geq 0; ift≥1t\\geq 1then⟨u,x⟩≤‖x‖=1≤t\\langle u,x\\rangle\\leq\\\|x\\\|=1\\leq ton𝕊d−1\\mathbb\{S\}^\{d\-1\}, so the unit is0on𝕊d−1\\mathbb\{S\}^\{d\-1\}except possibly at the single pointx=ux=u\(whent=1t=1\), whereReLU⁡\(0\)=0\\operatorname\{ReLU\}\(0\)=0as well; drop it\. After Step 4, ball units havet∈\(−R,R\)t\\in\(\-R,R\)and sphere units havet∈\[0,1\)t\\in\[0,1\)\.

*Step 5 \(orientation convention and coincident hyperplanes\)\.*Two pairs\(u,t\)≠\(u′,t′\)\(u,t\)\\neq\(u^\{\\prime\},t^\{\\prime\}\)with‖u‖=‖u′‖=1\\\|u\\\|=\\\|u^\{\\prime\}\\\|=1satisfyHu,t=Hu′,t′H\_\{u,t\}=H\_\{u^\{\\prime\},t^\{\\prime\}\}iff\(u′,t′\)=\(−u,−t\)\(u^\{\\prime\},t^\{\\prime\}\)=\(\-u,\-t\)\. In the ball case, fix the convention that the first nonzero coordinate ofuuis positive, using \([8](https://arxiv.org/html/2607.07778#A1.E8)\) to flip any offending unit; then coincident hyperplanes force identical\(u,t\)\(u,t\)\. In the sphere case, the conventiont≥0t\\geq 0already forces coincident hyperplanes to be identical, except whent=t′=0t=t^\{\\prime\}=0andu′=−uu^\{\\prime\}=\-u; there, adopt the ball convention \(first nonzero coordinate ofuupositive\) and apply \([8](https://arxiv.org/html/2607.07778#A1.E8)\) once to rewrite the offending\(−u,0\)\(\-u,0\)unit onto\(u,0\)\(u,0\)\(plus an affine term\)\. After Step 5, distinct units have distinct hyperplanes\.

*Step 6 \(merge and clean\)\.*Add the coefficients of units that now share a hyperplane; discard any unit whose merged coefficient is0\. The remaining units have pairwise distinct hyperplanes and nonzero coefficients, withtt\-ranges as claimed\. The geometric descriptions in \(i\)–\(ii\) are immediate:Hu,t∩int⁡BRH\_\{u,t\}\\cap\\operatorname\{int\}B\_\{R\}is the open disk of radiusR2−t2\\sqrt\{R^\{2\}\-t^\{2\}\}centered att​utu\(nonempty since\|t\|<R\|t\|<R\), andHu,t∩𝕊d−1H\_\{u,t\}\\cap\\mathbb\{S\}^\{d\-1\}is the sphere\{x:‖x‖=1,⟨u,x⟩=t\}\\\{x:\\\|x\\\|=1,\\ \\langle u,x\\rangle=t\\\}, which is a\(d−2\)\(d\-2\)\-sphere of radius1−t2\\sqrt\{1\-t^\{2\}\}centered att​utu\(nonempty sincet<1t<1\)\. ∎

### A\.3\.Proofs for Section[3](https://arxiv.org/html/2607.07778#S3)

###### Proof of Lemma[3\.1](https://arxiv.org/html/2607.07778#S3.Thmtheorem1)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S3.Thmtheorem1)Fixjjand letP:=Huj,tjP:=H\_\{u\_\{j\},t\_\{j\}\}, a hyperplane, andA:=P∩int⁡BRA:=P\\cap\\operatorname\{int\}B\_\{R\}, a nonempty relatively open\(d−1\)\(d\-1\)\-disk \(Lemma[2\.2](https://arxiv.org/html/2607.07778#S2.Thmtheorem2)\(i\)\)\.

*A generic point exists\.*For eachl≠jl\\neq j,P∩Hul,tlP\\cap H\_\{u\_\{l\},t\_\{l\}\}is either empty \(parallel distinct hyperplanes\) or an affine subspace of dimensiond−2d\-2\(distinct, non\-parallel\), hence in either case a set of\(d−1\)\(d\-1\)\-dimensional Lebesgue measure0insidePP\. A finite union of measure\-zero sets has measure0, whileAAhas positive\(d−1\)\(d\-1\)\-measure; therefore

U:=A∖⋃l≠jHul,tlU:=A\\setminus\\bigcup\_\{l\\neq j\}H\_\{u\_\{l\},t\_\{l\}\}has positive measure, in particularU≠∅U\\neq\\emptyset\. Fixx∗∈Ux^\{\*\}\\in U\. Because the finitely many closed setsHul,tlH\_\{u\_\{l\},t\_\{l\}\}\(l≠jl\\neq j\) and∂BR\\partial B\_\{R\}all avoidx∗x^\{\*\}, there isr\>0r\>0withB​\(x∗,r\)¯⊂int⁡BR\\overline\{B\(x^\{\*\},r\)\}\\subset\\operatorname\{int\}B\_\{R\}andB​\(x∗,r\)¯∩Hul,tl=∅\\overline\{B\(x^\{\*\},r\)\}\\cap H\_\{u\_\{l\},t\_\{l\}\}=\\emptysetfor alll≠jl\\neq j\.

*Only unitjjswitches nearx∗x^\{\*\}\.*On the connected setB​\(x∗,r\)B\(x^\{\*\},r\), each⟨ul,x⟩−tl\\langle u\_\{l\},x\\rangle\-t\_\{l\}\(l≠jl\\neq j\) has constant sign, soReLU⁡\(⟨ul,x⟩−tl\)\\operatorname\{ReLU\}\(\\langle u\_\{l\},x\\rangle\-t\_\{l\}\)is affine there\. Hence onB​\(x∗,r\)B\(x^\{\*\},r\)

f​\(x\)=A0​\(x\)\+αj​ReLU⁡\(⟨uj,x⟩−tj\),A0​affine\.f\(x\)=A\_\{0\}\(x\)\+\\alpha\_\{j\}\\operatorname\{ReLU\}\(\\langle u\_\{j\},x\\rangle\-t\_\{j\}\),\\qquad A\_\{0\}\\ \\text\{affine\.\}
*The one\-sided derivatives\.*Letφ​\(s\):=f​\(x∗\+s​uj\)\\varphi\(s\):=f\(x^\{\*\}\+su\_\{j\}\)for\|s\|<r\|s\|<r\. Since⟨uj,x∗⟩=tj\\langle u\_\{j\},x^\{\*\}\\rangle=t\_\{j\}and‖uj‖=1\\\|u\_\{j\}\\\|=1, we have⟨uj,x∗\+s​uj⟩−tj=s\\langle u\_\{j\},x^\{\*\}\+su\_\{j\}\\rangle\-t\_\{j\}=s, so

φ​\(s\)=A0​\(x∗\+s​uj\)\+αj​ReLU⁡\(s\)=\(A0​\(x∗\)\+β​s\)\+αj​ReLU⁡\(s\),β:=⟨∇A0,uj⟩\.\\varphi\(s\)=A\_\{0\}\(x^\{\*\}\+su\_\{j\}\)\+\\alpha\_\{j\}\\operatorname\{ReLU\}\(s\)=\\bigl\(A\_\{0\}\(x^\{\*\}\)\+\\beta s\\bigr\)\+\\alpha\_\{j\}\\operatorname\{ReLU\}\(s\),\\qquad\\beta:=\\langle\\nabla A\_\{0\},u\_\{j\}\\rangle\.This is piecewise linear insswith a single kink at0: fors<0s<0its slope isβ\\beta, fors\>0s\>0its slope isβ\+αj\\beta\+\\alpha\_\{j\}\. The compositions↦x∗\+s​ujs\\mapsto x^\{\*\}\+su\_\{j\}is an isometry \(as‖uj‖=1\\\|u\_\{j\}\\\|=1\), soφ\\varphiisLL\-Lipschitz\. Every difference quotient of anLL\-Lipschitz function lies in\[−L,L\]\[\-L,L\], and here the left and right slopes are exactly the one\-sided derivativesφ′​\(0−\)=β\\varphi^\{\\prime\}\(0^\{\-\}\)=\\betaandφ′​\(0\+\)=β\+αj\\varphi^\{\\prime\}\(0^\{\+\}\)=\\beta\+\\alpha\_\{j\}\. Thusβ∈\[−L,L\]\\beta\\in\[\-L,L\]andβ\+αj∈\[−L,L\]\\beta\+\\alpha\_\{j\}\\in\[\-L,L\], and subtracting gives\|αj\|=\|\(β\+αj\)−β\|≤2​L\|\\alpha\_\{j\}\|=\|\(\\beta\+\\alpha\_\{j\}\)\-\\beta\|\\leq 2L\. ∎

###### Proof of Lemma[3\.2](https://arxiv.org/html/2607.07778#S3.Thmtheorem2)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S3.Thmtheorem2)Fixjjand letΣj:=Huj,tj∩𝕊d−1\\Sigma\_\{j\}:=H\_\{u\_\{j\},t\_\{j\}\}\\cap\\mathbb\{S\}^\{d\-1\}, a\(d−2\)\(d\-2\)\-sphere of radiusρ:=1−tj2\>0\\rho:=\\sqrt\{1\-t\_\{j\}^\{2\}\}\>0centered attj​ujt\_\{j\}u\_\{j\}inside the hyperplaneP:=Huj,tjP:=H\_\{u\_\{j\},t\_\{j\}\}\(Lemma[2\.2](https://arxiv.org/html/2607.07778#S2.Thmtheorem2)\(ii\)\)\.

*A generic point onΣj\\Sigma\_\{j\}exists\.*Fixl≠jl\\neq j\. IfΣj⊂Hul,tl\\Sigma\_\{j\}\\subset H\_\{u\_\{l\},t\_\{l\}\}, thenHul,tlH\_\{u\_\{l\},t\_\{l\}\}contains the affine hullaff⁡\(Σj\)\\operatorname\{aff\}\(\\Sigma\_\{j\}\)\. Ford≥3d\\geq 3the sphereΣj\\Sigma\_\{j\}has dimensiond−2≥1d\-2\\geq 1and positive radius, so it affinely spansPP; thusaff⁡\(Σj\)=P\\operatorname\{aff\}\(\\Sigma\_\{j\}\)=PandHul,tl⊃PH\_\{u\_\{l\},t\_\{l\}\}\\supset P, forcingHul,tl=PH\_\{u\_\{l\},t\_\{l\}\}=P\(both are hyperplanes\), i\.e\. the two hyperplanes coincide — excluded\. ThereforeHul,tl∩ΣjH\_\{u\_\{l\},t\_\{l\}\}\\cap\\Sigma\_\{j\}is a*proper*closed subset ofΣj\\Sigma\_\{j\}; being the intersection of the sphereΣj\\Sigma\_\{j\}with a hyperplane that does not contain it, it is a sphere of dimension≤d−3\\leq d\-3, a single point, or empty, hence nowhere dense inΣj\\Sigma\_\{j\}\. A finite union of nowhere\-dense sets cannot be all of the complete metric spaceΣj\\Sigma\_\{j\}\(Baire\), so there isx∗∈Σjx^\{\*\}\\in\\Sigma\_\{j\}withx∗∉Hul,tlx^\{\*\}\\notin H\_\{u\_\{l\},t\_\{l\}\}for alll≠jl\\neq j\.

*A tangent direction along which unitjjswitches\.*Put

ξ:=uj−tj​x∗ρ\.\\xi:=\\frac\{u\_\{j\}\-t\_\{j\}x^\{\*\}\}\{\\rho\}\.Then, using⟨uj,x∗⟩=tj\\langle u\_\{j\},x^\{\*\}\\rangle=t\_\{j\}and‖x∗‖=1\\\|x^\{\*\}\\\|=1:

‖uj−tj​x∗‖2=‖uj‖2−2​tj​⟨uj,x∗⟩\+tj2​‖x∗‖2=1−2​tj2\+tj2=1−tj2=ρ2,\\\|u\_\{j\}\-t\_\{j\}x^\{\*\}\\\|^\{2\}=\\\|u\_\{j\}\\\|^\{2\}\-2t\_\{j\}\\langle u\_\{j\},x^\{\*\}\\rangle\+t\_\{j\}^\{2\}\\\|x^\{\*\}\\\|^\{2\}=1\-2t\_\{j\}^\{2\}\+t\_\{j\}^\{2\}=1\-t\_\{j\}^\{2\}=\\rho^\{2\},so‖ξ‖=1\\\|\\xi\\\|=1; and

⟨ξ,x∗⟩=⟨uj,x∗⟩−tj​‖x∗‖2ρ=tj−tjρ=0,⟨uj,ξ⟩=‖uj‖2−tj​⟨uj,x∗⟩ρ=1−tj2ρ=ρ\.\\langle\\xi,x^\{\*\}\\rangle=\\frac\{\\langle u\_\{j\},x^\{\*\}\\rangle\-t\_\{j\}\\\|x^\{\*\}\\\|^\{2\}\}\{\\rho\}=\\frac\{t\_\{j\}\-t\_\{j\}\}\{\\rho\}=0,\\qquad\\langle u\_\{j\},\\xi\\rangle=\\frac\{\\\|u\_\{j\}\\\|^\{2\}\-t\_\{j\}\\langle u\_\{j\},x^\{\*\}\\rangle\}\{\\rho\}=\\frac\{1\-t\_\{j\}^\{2\}\}\{\\rho\}=\\rho\.Consider the unit\-speed great circleγ​\(s\)=\(cos⁡s\)​x∗\+\(sin⁡s\)​ξ\\gamma\(s\)=\(\\cos s\)\\,x^\{\*\}\+\(\\sin s\)\\,\\xi; it lies on𝕊d−1\\mathbb\{S\}^\{d\-1\}because‖x∗‖=‖ξ‖=1\\\|x^\{\*\}\\\|=\\\|\\xi\\\|=1and⟨x∗,ξ⟩=0\\langle x^\{\*\},\\xi\\rangle=0\. Then

h​\(s\):=⟨uj,γ​\(s\)⟩−tj=\(cos⁡s−1\)​tj\+\(sin⁡s\)​ρ,h\(s\):=\\langle u\_\{j\},\\gamma\(s\)\\rangle\-t\_\{j\}=\(\\cos s\-1\)\\,t\_\{j\}\+\(\\sin s\)\\,\\rho,soh​\(0\)=0h\(0\)=0andh′​\(0\)=ρ\>0h^\{\\prime\}\(0\)=\\rho\>0: unitjjswitches ats=0s=0, and it does so transversally\. For smallss,ReLU⁡\(h​\(s\)\)\\operatorname\{ReLU\}\(h\(s\)\)has left derivative0and right derivativeh′​\(0\)=ρh^\{\\prime\}\(0\)=\\rhoats=0s=0\.

*The other units are smooth ats=0s=0\.*Forl≠jl\\neq j,⟨ul,γ​\(0\)⟩−tl=⟨ul,x∗⟩−tl≠0\\langle u\_\{l\},\\gamma\(0\)\\rangle\-t\_\{l\}=\\langle u\_\{l\},x^\{\*\}\\rangle\-t\_\{l\}\\neq 0\(asx∗∉Hul,tlx^\{\*\}\\notin H\_\{u\_\{l\},t\_\{l\}\}\), so by continuity⟨ul,γ​\(s\)⟩−tl\\langle u\_\{l\},\\gamma\(s\)\\rangle\-t\_\{l\}keeps its sign forssnear0andReLU⁡\(⟨ul,γ​\(s\)⟩−tl\)\\operatorname\{ReLU\}\(\\langle u\_\{l\},\\gamma\(s\)\\rangle\-t\_\{l\}\)is smooth \(affine composed with the analyticγ\\gamma\) there; the affine part⟨v,γ​\(s\)⟩\+c\\langle v,\\gamma\(s\)\\rangle\+cis smooth as well\.

*Conclusion\.*Letφ​\(s\):=f​\(γ​\(s\)\)\\varphi\(s\):=f\(\\gamma\(s\)\)\. Chords are bounded by arcs:‖γ​\(s\)−γ​\(s′\)‖=2​\|sin⁡s−s′2\|≤\|s−s′\|\\\|\\gamma\(s\)\-\\gamma\(s^\{\\prime\}\)\\\|=2\|\\sin\\tfrac\{s\-s^\{\\prime\}\}\{2\}\|\\leq\|s\-s^\{\\prime\}\|\. SinceffisLL\-Lipschitz on𝕊d−1\\mathbb\{S\}^\{d\-1\}for the Euclidean metric,\|φ​\(s\)−φ​\(s′\)\|≤L​‖γ​\(s\)−γ​\(s′\)‖≤L​\|s−s′\|\|\\varphi\(s\)\-\\varphi\(s^\{\\prime\}\)\|\\leq L\\,\\\|\\gamma\(s\)\-\\gamma\(s^\{\\prime\}\)\\\|\\leq L\|s\-s^\{\\prime\}\|, i\.e\.φ\\varphiisLL\-Lipschitz near0\. All summands off∘γf\\circ\\gammaexcept unitjjare differentiable at0; unitjjcontributesαj​ReLU⁡\(h​\(s\)\)\\alpha\_\{j\}\\operatorname\{ReLU\}\(h\(s\)\), whose one\-sided derivatives at0differ byαj​h′​\(0\)=αj​ρ\\alpha\_\{j\}h^\{\\prime\}\(0\)=\\alpha\_\{j\}\\rho\. Henceφ′​\(0\+\)−φ′​\(0−\)=αj​ρ\\varphi^\{\\prime\}\(0^\{\+\}\)\-\\varphi^\{\\prime\}\(0^\{\-\}\)=\\alpha\_\{j\}\\rho, and both one\-sided derivatives lie in\[−L,L\]\[\-L,L\], so\|αj\|​ρ≤2​L\|\\alpha\_\{j\}\|\\rho\\leq 2L\. ∎

###### Proof of Proposition[3\.3](https://arxiv.org/html/2607.07778#S3.Thmtheorem3)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S3.Thmtheorem3)Parameterize𝕊1\\mathbb\{S\}^\{1\}by angleθ\\theta\. A unit

α​ReLU⁡\(⟨u,x⟩−t\),u=\(cos⁡θ0,sin⁡θ0\),t=cos⁡a,\\alpha\\,\\operatorname\{ReLU\}\(\\langle u,x\\rangle\-t\),\\qquad u=\(\\cos\\theta\_\{0\},\\sin\\theta\_\{0\}\),\\quad t=\\cos a,witha∈\(0,π\)a\\in\(0,\\pi\)becomes

α​ReLU⁡\(cos⁡\(θ−θ0\)−cos⁡a\)\.\\alpha\\,\\operatorname\{ReLU\}\(\\cos\(\\theta\-\\theta\_\{0\}\)\-\\cos a\)\.Ifa<πa<\\piand the active arc is not wrapped around the cut, its kink set is the two anglesθ0±a\\theta\_\{0\}\\pm a\. On the active arc the angular derivative is−α​sin⁡\(θ−θ0\)\-\\alpha\\sin\(\\theta\-\\theta\_\{0\}\)and outside it is0\. Hence the derivative jump at each of the two kink angles isα​sin⁡a=α​1−t2\\alpha\\sin a=\\alpha\\sqrt\{1\-t^\{2\}\}\.

Fix once and for alla∈\(0,π/4\)a\\in\(0,\\pi/4\), and chooseη∈\(0,a/10\)\\eta\\in\(0,a/10\)later\. Put

A=−a,B=−a\+η,C=a−η,D=a\.A=\-a,\\qquad B=\-a\+\\eta,\\qquad C=a\-\\eta,\\qquad D=a\.Consider the four arcs

I1=\[A,D\],I2=\[B,D\],I3=\[B,C\],I4=\[A,C\]\.I\_\{1\}=\[A,D\],\\qquad I\_\{2\}=\[B,D\],\\qquad I\_\{3\}=\[B,C\],\\qquad I\_\{4\}=\[A,C\]\.For an intervalI=\[p,q\]I=\[p,q\]writec​\(I\)=\(p\+q\)/2c\(I\)=\(p\+q\)/2andr​\(I\)=\(q−p\)/2r\(I\)=\(q\-p\)/2\. Define

jIjθj=c​\(Ij\)aj=r​\(Ij\)1\[A,D\]0a2\[B,D\]η/2a−η/23\[B,C\]0a−η4\[A,C\]−η/2a−η/2,\\begin\{array\}\[\]\{c\|c\|c\|c\}j&I\_\{j\}&\\theta\_\{j\}=c\(I\_\{j\}\)&a\_\{j\}=r\(I\_\{j\}\)\\\\ \\hline\\cr 1&\[A,D\]&0&a\\\\ 2&\[B,D\]&\\eta/2&a\-\\eta/2\\\\ 3&\[B,C\]&0&a\-\\eta\\\\ 4&\[A,C\]&\-\\eta/2&a\-\\eta/2,\\end\{array\}and choose signss1=\+1,s2=−1,s3=\+1,s4=−1s\_\{1\}=\+1,s\_\{2\}=\-1,s\_\{3\}=\+1,s\_\{4\}=\-1\. Let

F​\(θ\)=∑j=14αj​ReLU⁡\(cos⁡\(θ−θj\)−cos⁡aj\),αj=sj​Λsin⁡aj\.F\(\\theta\)=\\sum\_\{j=1\}^\{4\}\\alpha\_\{j\}\\operatorname\{ReLU\}\(\\cos\(\\theta\-\\theta\_\{j\}\)\-\\cos a\_\{j\}\),\\qquad\\alpha\_\{j\}=s\_\{j\}\\frac\{\\Lambda\}\{\\sin a\_\{j\}\}\.The four kink point\-pairs are precisely\{A,D\}\\\{A,D\\\},\{B,D\}\\\{B,D\\\},\{B,C\}\\\{B,C\\\}and\{A,C\}\\\{A,C\\\}, hence are distinct\. Also\|αj\|​1−cos2⁡aj=\|αj\|​sin⁡aj=Λ\|\\alpha\_\{j\}\|\\sqrt\{1\-\\cos^\{2\}a\_\{j\}\}=\|\\alpha\_\{j\}\|\\sin a\_\{j\}=\\Lambdafor everyjj\.

At each of the four kink anglesA,B,C,DA,B,C,D, exactly two units meet, one with sign\+1\+1and one with sign−1\-1\. Their derivative jumps are therefore\+Λ\+\\Lambdaand−Λ\-\\Lambda, so all derivative jumps cancel\. ThusFFisC1C^\{1\}as a function ofθ\\theta\.

It remains to bound the derivative\. Outside\[A,D\]\[A,D\]every unit is inactive\. On the three subarcs\[A,B\]\[A,B\],\[B,C\]\[B,C\]and\[C,D\]\[C,D\], direct differentiation gives

F′​\(θ\)Λ=Gη​\(θ\),\\frac\{F^\{\\prime\}\(\\theta\)\}\{\\Lambda\}=G\_\{\\eta\}\(\\theta\),where

Gη​\(θ\)\\displaystyle G\_\{\\eta\}\(\\theta\)=−sin⁡θsin⁡a\+sin⁡\(θ\+η/2\)sin⁡\(a−η/2\),\\displaystyle=\-\\frac\{\\sin\\theta\}\{\\sin a\}\+\\frac\{\\sin\(\\theta\+\\eta/2\)\}\{\\sin\(a\-\\eta/2\)\},θ∈\[A,B\],\\displaystyle\\theta\\in\[A,B\],Gη​\(θ\)\\displaystyle G\_\{\\eta\}\(\\theta\)=−sin⁡θsin⁡a\+sin⁡\(θ−η/2\)sin⁡\(a−η/2\)−sin⁡θsin⁡\(a−η\)\+sin⁡\(θ\+η/2\)sin⁡\(a−η/2\),\\displaystyle=\-\\frac\{\\sin\\theta\}\{\\sin a\}\+\\frac\{\\sin\(\\theta\-\\eta/2\)\}\{\\sin\(a\-\\eta/2\)\}\-\\frac\{\\sin\\theta\}\{\\sin\(a\-\\eta\)\}\+\\frac\{\\sin\(\\theta\+\\eta/2\)\}\{\\sin\(a\-\\eta/2\)\},θ∈\[B,C\],\\displaystyle\\theta\\in\[B,C\],Gη​\(θ\)\\displaystyle G\_\{\\eta\}\(\\theta\)=−sin⁡θsin⁡a\+sin⁡\(θ−η/2\)sin⁡\(a−η/2\),\\displaystyle=\-\\frac\{\\sin\\theta\}\{\\sin a\}\+\\frac\{\\sin\(\\theta\-\\eta/2\)\}\{\\sin\(a\-\\eta/2\)\},θ∈\[C,D\]\.\\displaystyle\\theta\\in\[C,D\]\.Forη=0\\eta=0each displayed expression is identically zero\. Sinceaais fixed away from0, all denominators stay bounded below for0≤η≤a/100\\leq\\eta\\leq a/10, and the derivatives of the displayed expressions with respect toη\\etaare uniformly bounded forθ∈\[−a,a\]\\theta\\in\[\-a,a\]\. The mean\-value theorem therefore gives a constantCaC\_\{a\}such that

supθ\|F′​\(θ\)\|≤Ca​Λ​η\.\\sup\_\{\\theta\}\|F^\{\\prime\}\(\\theta\)\|\\leq C\_\{a\}\\Lambda\\eta\.BecauseF∈C1F\\in C^\{1\}, its Lipschitz constant in the arclength \(angular\) metric equalssupθ\|F′​\(θ\)\|\\sup\_\{\\theta\}\|F^\{\\prime\}\(\\theta\)\|, and the Euclidean chord metric on𝕊1\\mathbb\{S\}^\{1\}is within a factorπ/2\\pi/2of angular distance on arcs of length at mostπ\\pi\. Hence

Lip𝕊1⁡\(F\)≤\(π/2\)​Ca​Λ​η\.\\operatorname\{Lip\}\_\{\\mathbb\{S\}^\{1\}\}\(F\)\\leq\(\\pi/2\)C\_\{a\}\\Lambda\\eta\.Choosingη≤min⁡\(a/10,2/\(π​Ca​R\)\)\\eta\\leq\\min\\bigl\(a/10,\\,2/\(\\pi C\_\{a\}R\)\\bigr\)givesLip𝕊1⁡\(F\)≤Λ/R\\operatorname\{Lip\}\_\{\\mathbb\{S\}^\{1\}\}\(F\)\\leq\\Lambda/R, as required\. ∎

###### Proof of Lemma[3\.5](https://arxiv.org/html/2607.07778#S3.Thmtheorem5)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S3.Thmtheorem5)The average of thennnonnegative numbers\(f​\(xi\)−yi\)2\(f\(x\_\{i\}\)\-y\_\{i\}\)^\{2\}is at most11, so at least one of them is at most11; for thatii,\|f​\(xi\)−yi\|≤1\|f\(x\_\{i\}\)\-y\_\{i\}\|\\leq 1, hence\|f​\(xi\)\|≤1\+\|yi\|≤2\|f\(x\_\{i\}\)\|\\leq 1\+\|y\_\{i\}\|\\leq 2\. For anyx∈Dx\\in D,\|f​\(x\)\|≤\|f​\(xi\)\|\+L​‖x−xi‖≤2\+L​diam⁡\(D\)≤2\+4​L\|f\(x\)\|\\leq\|f\(x\_\{i\}\)\|\+L\\\|x\-x\_\{i\}\\\|\\leq 2\+L\\operatorname\{diam\}\(D\)\\leq 2\+4L\. ∎

###### Proof of Lemma[3\.6](https://arxiv.org/html/2607.07778#S3.Thmtheorem6)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S3.Thmtheorem6)\(i\) The union⋃jHuj,tj\\bigcup\_\{j\}H\_\{u\_\{j\},t\_\{j\}\}has measure0, so pickx0∈int⁡BRx\_\{0\}\\in\\operatorname\{int\}B\_\{R\}off allm0m\_\{0\}hyperplanes; thereffis differentiable with

∇f​\(x0\)=v\+∑j:⟨uj,x0⟩\>tjαj​uj\.\\nabla f\(x\_\{0\}\)=v\+\\sum\_\{j:\\,\\langle u\_\{j\},x\_\{0\}\\rangle\>t\_\{j\}\}\\alpha\_\{j\}u\_\{j\}\.A differentiable point of anLL\-Lipschitz function has‖∇f​\(x0\)‖≤L\\\|\\nabla f\(x\_\{0\}\)\\\|\\leq L\. By Lemma[3\.1](https://arxiv.org/html/2607.07778#S3.Thmtheorem1),‖∑activeαj​uj‖≤∑j\|αj\|≤2​L​m0\\\|\\sum\_\{\\text\{active\}\}\\alpha\_\{j\}u\_\{j\}\\\|\\leq\\sum\_\{j\}\|\\alpha\_\{j\}\|\\leq 2Lm\_\{0\}, so‖v‖≤‖∇f​\(x0\)‖\+2​L​m0≤L​\(1\+2​m0\)\\\|v\\\|\\leq\\\|\\nabla f\(x\_\{0\}\)\\\|\+2Lm\_\{0\}\\leq L\(1\+2m\_\{0\}\)\. Evaluating \([3](https://arxiv.org/html/2607.07778#S2.E3)\) atx=0x=0givesc=f​\(0\)−∑jαj​ReLU⁡\(−tj\)c=f\(0\)\-\\sum\_\{j\}\\alpha\_\{j\}\\operatorname\{ReLU\}\(\-t\_\{j\}\), and\|αj\|​ReLU⁡\(−tj\)≤2​L⋅\|tj\|≤2​L⋅R≤4​L\|\\alpha\_\{j\}\|\\operatorname\{ReLU\}\(\-t\_\{j\}\)\\leq 2L\\cdot\|t\_\{j\}\|\\leq 2L\\cdot R\\leq 4L, whence\|c\|≤B0\+4​L​m0\|c\|\\leq B\_\{0\}\+4Lm\_\{0\}\.

\(ii\) Letx∼Unif​\(𝕊d−1\)x\\sim\\mathrm\{Unif\}\(\\mathbb\{S\}^\{d\-1\}\)\. By rotational invariance𝔼​\[x\]=0\\mathbb\{E\}\[x\]=0and𝔼​\[x​x⊤\]=1d​Id\\mathbb\{E\}\[xx^\{\\top\}\]=\\tfrac\{1\}\{d\}I\_\{d\}\(it is a scalar multiple ofIdI\_\{d\}by symmetry, and its trace is𝔼​‖x‖2=1\\mathbb\{E\}\\\|x\\\|^\{2\}=1\)\. Hence

𝔼​\[f​\(x\)​x\]=vd\+∑jαj​𝔼​\[ReLU⁡\(⟨uj,x⟩−tj\)​x\]\.\\mathbb\{E\}\[f\(x\)\\,x\]=\\frac\{v\}\{d\}\+\\sum\_\{j\}\\alpha\_\{j\}\\,\\mathbb\{E\}\\\!\\bigl\[\\operatorname\{ReLU\}\(\\langle u\_\{j\},x\\rangle\-t\_\{j\}\)\\,x\\bigr\]\.For a fixed unituu, the mapx↦ReLU⁡\(⟨u,x⟩−t\)​xx\\mapsto\\operatorname\{ReLU\}\(\\langle u,x\\rangle\-t\)\\,xhas expectation invariant under all rotations fixinguu, so𝔼​\[ReLU⁡\(⟨u,x⟩−t\)​x\]=λ​\(t\)​u\\mathbb\{E\}\[\\operatorname\{ReLU\}\(\\langle u,x\\rangle\-t\)\\,x\]=\\lambda\(t\)\\,ufor a scalarλ​\(t\)=⟨𝔼​\[ReLU⁡\(⟨u,x⟩−t\)​x\],u⟩=𝔼​\[ReLU⁡\(⟨u,x⟩−t\)​⟨u,x⟩\]\\lambda\(t\)=\\langle\\mathbb\{E\}\[\\operatorname\{ReLU\}\(\\langle u,x\\rangle\-t\)x\],u\\rangle=\\mathbb\{E\}\[\\operatorname\{ReLU\}\(\\langle u,x\\rangle\-t\)\\langle u,x\\rangle\]\. Since\|⟨u,x⟩\|≤1\|\\langle u,x\\rangle\|\\leq 1on𝕊d−1\\mathbb\{S\}^\{d\-1\},

0≤λ​\(t\)≤𝔼​ReLU⁡\(⟨u,x⟩−t\)≤\(1−t\)​ℙ​\(⟨u,x⟩\>t\)≤1−t\.0\\leq\\lambda\(t\)\\leq\\mathbb\{E\}\\,\\operatorname\{ReLU\}\(\\langle u,x\\rangle\-t\)\\leq\(1\-t\)\\,\\mathbb\{P\}\(\\langle u,x\\rangle\>t\)\\leq 1\-t\.By Lemma[3\.2](https://arxiv.org/html/2607.07778#S3.Thmtheorem2),\|αj\|​λ​\(tj\)≤2​L1−tj2​\(1−tj\)=2​L​1−tj1\+tj≤2​L\|\\alpha\_\{j\}\|\\lambda\(t\_\{j\}\)\\leq\\dfrac\{2L\}\{\\sqrt\{1\-t\_\{j\}^\{2\}\}\}\(1\-t\_\{j\}\)=2L\\sqrt\{\\dfrac\{1\-t\_\{j\}\}\{1\+t\_\{j\}\}\}\\leq 2L\. Therefore

‖v‖=‖d​𝔼​\[f​\(x\)​x\]−d​∑jαj​λ​\(tj\)​uj‖≤d​\(𝔼​\|f\|\+∑j\|αj\|​λ​\(tj\)\)≤d​\(B0\+2​L​m0\)\.\\\|v\\\|=\\Bigl\\\|d\\,\\mathbb\{E\}\[f\(x\)x\]\-d\\sum\_\{j\}\\alpha\_\{j\}\\lambda\(t\_\{j\}\)u\_\{j\}\\Bigr\\\|\\leq d\\bigl\(\\mathbb\{E\}\|f\|\+\\textstyle\\sum\_\{j\}\|\\alpha\_\{j\}\|\\lambda\(t\_\{j\}\)\\bigr\)\\leq d\(B\_\{0\}\+2Lm\_\{0\}\)\.Finally, at anyx1∈𝕊d−1x\_\{1\}\\in\\mathbb\{S\}^\{d\-1\},c=f​\(x1\)−⟨v,x1⟩−∑jαj​ReLU⁡\(⟨uj,x1⟩−tj\)c=f\(x\_\{1\}\)\-\\langle v,x\_\{1\}\\rangle\-\\sum\_\{j\}\\alpha\_\{j\}\\operatorname\{ReLU\}\(\\langle u\_\{j\},x\_\{1\}\\rangle\-t\_\{j\}\)and\|αj\|​ReLU⁡\(⟨uj,x1⟩−tj\)≤\|αj\|​\(1−tj\)≤2​L\|\\alpha\_\{j\}\|\\operatorname\{ReLU\}\(\\langle u\_\{j\},x\_\{1\}\\rangle\-t\_\{j\}\)\\leq\|\\alpha\_\{j\}\|\(1\-t\_\{j\}\)\\leq 2L\(again by Lemma[3\.2](https://arxiv.org/html/2607.07778#S3.Thmtheorem2)and1−t≤1−t21\-t\\leq\\sqrt\{1\-t^\{2\}\}fort∈\[0,1\)t\\in\[0,1\)\), giving\|c\|≤B0\+‖v‖\+2​L​m0\|c\|\\leq B\_\{0\}\+\\\|v\\\|\+2Lm\_\{0\}\. ∎

### A\.4\.Proofs for Section[4](https://arxiv.org/html/2607.07778#S4)

###### Proof of Proposition[4\.1](https://arxiv.org/html/2607.07778#S4.Thmtheorem1)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S4.Thmtheorem1)Everything is a Lipschitz\-in\-parameters estimate followed by a product of one\-dimensional grids\. We build a finite set𝒩∗⊂𝒢m¯,L\\mathcal\{N\}^\{\\ast\}\\subset\\mathcal\{G\}\_\{\\bar\{m\},L\}such that everyg∈𝒢m¯,Lg\\in\\mathcal\{G\}\_\{\\bar\{m\},L\}has someg∗∈𝒩∗g^\{\\ast\}\\in\\mathcal\{N\}^\{\\ast\}with‖g−g∗‖L∞​\(D\)≤ε′\\\|g\-g^\{\\ast\}\\\|\_\{L^\{\\infty\}\(D\)\}\\leq\\varepsilon^\{\\prime\}, and bound\|𝒩∗\|\|\\mathcal\{N\}^\{\\ast\}\|\.

*Per\-unit sensitivity \(ball\)\.*OnB2B\_\{2\}, for one ReLU unit,

\|α​ReLU⁡\(⟨u,x⟩−t\)−α′​ReLU⁡\(⟨u′,x⟩−t′\)\|\\displaystyle\\bigl\|\\alpha\\operatorname\{ReLU\}\(\\langle u,x\\rangle\-t\)\-\\alpha^\{\\prime\}\\operatorname\{ReLU\}\(\\langle u^\{\\prime\},x\\rangle\-t^\{\\prime\}\)\\bigr\|≤\|α−α′\|⋅ReLU⁡\(⟨u,x⟩−t\)\\displaystyle\\leq\\bigl\|\\alpha\-\\alpha^\{\\prime\}\\bigr\|\\cdot\\operatorname\{ReLU\}\(\\langle u,x\\rangle\-t\)\+\|α′\|⋅\|ReLU⁡\(⟨u,x⟩−t\)−ReLU⁡\(⟨u′,x⟩−t′\)\|\.\\displaystyle\\quad\+\|\\alpha^\{\\prime\}\|\\cdot\\bigl\|\\operatorname\{ReLU\}\(\\langle u,x\\rangle\-t\)\-\\operatorname\{ReLU\}\(\\langle u^\{\\prime\},x\\rangle\-t^\{\\prime\}\)\\bigr\|\.OnB2B\_\{2\}with\|t\|≤2\|t\|\\leq 2we haveReLU⁡\(⟨u,x⟩−t\)≤\|⟨u,x⟩−t\|≤‖x‖\+\|t\|≤4\\operatorname\{ReLU\}\(\\langle u,x\\rangle\-t\)\\leq\|\\langle u,x\\rangle\-t\|\\leq\\\|x\\\|\+\|t\|\\leq 4, and sinceReLU\\operatorname\{ReLU\}is11\-Lipschitz,

\|ReLU⁡\(⟨u,x⟩−t\)−ReLU⁡\(⟨u′,x⟩−t′\)\|≤\|⟨u−u′,x⟩\|\+\|t−t′\|≤2​‖u−u′‖\+\|t−t′\|\.\\bigl\|\\operatorname\{ReLU\}\(\\langle u,x\\rangle\-t\)\-\\operatorname\{ReLU\}\(\\langle u^\{\\prime\},x\\rangle\-t^\{\\prime\}\)\\bigr\|\\leq\|\\langle u\-u^\{\\prime\},x\\rangle\|\+\|t\-t^\{\\prime\}\|\\leq 2\\\|u\-u^\{\\prime\}\\\|\+\|t\-t^\{\\prime\}\|\.So the per\-unit change is at most4​\|α−α′\|\+\|α′\|​\(2​‖u−u′‖\+\|t−t′\|\)4\|\\alpha\-\\alpha^\{\\prime\}\|\+\|\\alpha^\{\\prime\}\|\(2\\\|u\-u^\{\\prime\}\\\|\+\|t\-t^\{\\prime\}\|\), and with\|α′\|≤2​L\|\\alpha^\{\\prime\}\|\\leq 2L,

≤4​\|α−α′\|\+2​L​\(2​‖u−u′‖\+\|t−t′\|\)\.\\leq 4\|\\alpha\-\\alpha^\{\\prime\}\|\+2L\\bigl\(2\\\|u\-u^\{\\prime\}\\\|\+\|t\-t^\{\\prime\}\|\\bigr\)\.\(9\)
*Grids \(ball\)\.*Discretize, per unit:

- •αj∈\[−2​L,2​L\]\\alpha\_\{j\}\\in\[\-2L,2L\]on a grid of meshε′/\(32​m¯\)\\varepsilon^\{\\prime\}/\(32\\bar\{m\}\): at most1\+4​Lε′/\(32​m¯\)=1\+128​L​m¯ε′1\+\\dfrac\{4L\}\{\\varepsilon^\{\\prime\}/\(32\\bar\{m\}\)\}=1\+\\dfrac\{128L\\bar\{m\}\}\{\\varepsilon^\{\\prime\}\}points; contributes≤4⋅ε′32​m¯=ε′8​m¯\\leq 4\\cdot\\dfrac\{\\varepsilon^\{\\prime\}\}\{32\\bar\{m\}\}=\\dfrac\{\\varepsilon^\{\\prime\}\}\{8\\bar\{m\}\}per unit\.
- •tj∈\(−2,2\)t\_\{j\}\\in\(\-2,2\)on a grid of meshε′/\(32​L​m¯\)\\varepsilon^\{\\prime\}/\(32L\\bar\{m\}\): at most1\+128​L​m¯ε′1\+\\dfrac\{128L\\bar\{m\}\}\{\\varepsilon^\{\\prime\}\}points; contributes≤2​L⋅ε′32​L​m¯=ε′16​m¯\\leq 2L\\cdot\\dfrac\{\\varepsilon^\{\\prime\}\}\{32L\\bar\{m\}\}=\\dfrac\{\\varepsilon^\{\\prime\}\}\{16\\bar\{m\}\}per unit\.
- •uju\_\{j\}on aε′64​L​m¯\\dfrac\{\\varepsilon^\{\\prime\}\}\{64L\\bar\{m\}\}\-net of𝕊d−1\\mathbb\{S\}^\{d\-1\}: at most\(1\+128​L​m¯ε′\)d\\bigl\(1\+\\dfrac\{128L\\bar\{m\}\}\{\\varepsilon^\{\\prime\}\}\\bigr\)^\{d\}points \(the standard volumetric boundN​\(𝕊d−1,ρ\)≤\(1\+2/ρ\)dN\(\\mathbb\{S\}^\{d\-1\},\\rho\)\\leq\(1\+2/\\rho\)^\{d\};\[[7](https://arxiv.org/html/2607.07778#bib.bib7), Cor\. 4\.2\.13\]\); contributes≤2​L⋅2⋅ε′64​L​m¯=ε′16​m¯\\leq 2L\\cdot 2\\cdot\\dfrac\{\\varepsilon^\{\\prime\}\}\{64L\\bar\{m\}\}=\\dfrac\{\\varepsilon^\{\\prime\}\}\{16\\bar\{m\}\}per unit\.

Summing the three per\-unit contributions gives≤ε′8​m¯\+ε′16​m¯\+ε′16​m¯=ε′4​m¯\\leq\\dfrac\{\\varepsilon^\{\\prime\}\}\{8\\bar\{m\}\}\+\\dfrac\{\\varepsilon^\{\\prime\}\}\{16\\bar\{m\}\}\+\\dfrac\{\\varepsilon^\{\\prime\}\}\{16\\bar\{m\}\}=\\dfrac\{\\varepsilon^\{\\prime\}\}\{4\\bar\{m\}\}; over≤m¯\\leq\\bar\{m\}units,≤ε′/4\\leq\\varepsilon^\{\\prime\}/4\. Discretize the affine part:

- •vvon a\(ε′/16\)\(\\varepsilon^\{\\prime\}/16\)\-net of\{‖v‖≤L​\(1\+2​m¯\)\}⊂ℝd\\\{\\\|v\\\|\\leq L\(1\+2\\bar\{m\}\)\\\}\\subset\\mathbb\{R\}^\{d\}: at most\(1\+32​L​\(1\+2​m¯\)ε′\)d\\bigl\(1\+\\dfrac\{32L\(1\+2\\bar\{m\}\)\}\{\\varepsilon^\{\\prime\}\}\\bigr\)^\{d\}points; contributes≤2⋅ε′16=ε′8\\leq 2\\cdot\\dfrac\{\\varepsilon^\{\\prime\}\}\{16\}=\\dfrac\{\\varepsilon^\{\\prime\}\}\{8\}\(using\|⟨v−v′,x⟩\|≤2​‖v−v′‖\|\\langle v\-v^\{\\prime\},x\\rangle\|\\leq 2\\\|v\-v^\{\\prime\}\\\|onB2B\_\{2\}\)\.
- •ccon a grid of meshε′/8\\varepsilon^\{\\prime\}/8in\[−\(B0\+4​L​m¯\),B0\+4​L​m¯\]\[\-\(B\_\{0\}\+4L\\bar\{m\}\),B\_\{0\}\+4L\\bar\{m\}\]: contributes≤ε′/8\\leq\\varepsilon^\{\\prime\}/8\.

Total change≤ε′/4\+ε′/8\+ε′/8=ε′/2≤ε′\\leq\\varepsilon^\{\\prime\}/4\+\\varepsilon^\{\\prime\}/8\+\\varepsilon^\{\\prime\}/8=\\varepsilon^\{\\prime\}/2\\leq\\varepsilon^\{\\prime\}\. Finally sum over the choicem0∈\{0,…,m¯\}m\_\{0\}\\in\\\{0,\\dots,\\bar\{m\}\\\}of active\-unit count \(a factorm¯\+1\\bar\{m\}\+1\)\. Taking logarithms of the product of cardinalities,

log⁡\|𝒩∗\|≤\(m¯​d\+2​m¯\+d\+1\)​log⁡\(C​m¯​d​\(2\+L\)ε′\)\+log⁡\(m¯\+1\)≤C​m¯​d​log⁡\(C​m¯​d​\(2\+L\)ε′\)\.\\log\|\\mathcal\{N\}^\{\\ast\}\|\\leq\(\\bar\{m\}d\+2\\bar\{m\}\+d\+1\)\\log\\\!\\Bigl\(\\frac\{C\\bar\{m\}d\(2\+L\)\}\{\\varepsilon^\{\\prime\}\}\\Bigr\)\+\\log\(\\bar\{m\}\+1\)\\leq C\\bar\{m\}d\\log\\\!\\Bigl\(\\frac\{C\\bar\{m\}d\(2\+L\)\}\{\\varepsilon^\{\\prime\}\}\\Bigr\)\.
*Grids \(sphere\)\.*On𝕊d−1\\mathbb\{S\}^\{d\-1\}we have\|⟨u,x⟩−t\|≤2\|\\langle u,x\\rangle\-t\|\\leq 2, so the bulk of \([9](https://arxiv.org/html/2607.07778#A1.E9)\) is unchanged; the only issue is that a unit withttnear11\(a small spherical cap\) has a large allowed\|α\|≤2​L/1−t2\|\\alpha\|\\leq 2L/\\sqrt\{1\-t^\{2\}\}\. Split the units\. The plan: bulk units reuse the ball grids; cap units are first deleted when their sup\-norm is negligible, and the survivors are binned dyadically inγ=1−t\\gamma=\\sqrt\{1\-t\}, with the ball meshes rescaled byγr\\gamma\_\{r\}inside each bin — the rescaling exactly compensates the allowed coefficient2​L/γr2L/\\gamma\_\{r\}, so each bin contributes the ball\-case error at the ball\-case cardinality\. Two bookkeeping points, once and for all\. First, a grid center need not itself satisfy thett\-dependent coefficient constraint\|α\|≤2​L/1−t2\|\\alpha\|\\leq 2L/\\sqrt\{1\-t^\{2\}\}at its griddedtt: the grids produce an*external*cover of𝒢m¯,L\\mathcal\{G\}\_\{\\bar\{m\},L\}, which suffices, since the closing paragraph of the proof converts any external\(ε′/2\)\(\\varepsilon^\{\\prime\}/2\)\-cover into an internalε′\\varepsilon^\{\\prime\}\-net\. Second, the assignment of each unit to its regime \(bulk, one of theR\+1R\+1cap bins, or deleted\) is part of the enumeration: it multiplies the count by at most\(R\+3\)m¯\(R\+3\)^\{\\bar\{m\}\}, an additivem¯​log⁡\(R\+3\)≤C​m¯​log⁡\(C​m¯​d​\(2\+L\)/ε′\)\\bar\{m\}\\log\(R\+3\)\\leq C\\bar\{m\}\\log\\bigl\(C\\bar\{m\}d\(2\+L\)/\\varepsilon^\{\\prime\}\\bigr\)in the logarithm, absorbed into the displayed bound\.

*Bulk units*\(tj≤12t\_\{j\}\\leq\\tfrac\{1\}\{2\}\): here\|αj\|≤2​L/1−14≤3​L\|\\alpha\_\{j\}\|\\leq 2L/\\sqrt\{1\-\\tfrac\{1\}\{4\}\}\\leq 3L, and the ball grids above apply verbatim \(with the constant33in place of22, absorbed intoCC\)\.

*Cap units*\(tj∈\(12,1\)t\_\{j\}\\in\(\\tfrac\{1\}\{2\},1\)\): writeγ:=1−t∈\(0,2−1/2\)\\gamma:=\\sqrt\{1\-t\}\\in\(0,2^\{\-1/2\}\)\. On𝕊d−1\\mathbb\{S\}^\{d\-1\},

supx∈𝕊d−1ReLU⁡\(⟨u,x⟩−t\)=1−t=γ2,\|α\|≤2​L1−t2=2​L\(1−t\)​\(1\+t\)≤2​Lγ,\\sup\_\{x\\in\\mathbb\{S\}^\{d\-1\}\}\\operatorname\{ReLU\}\(\\langle u,x\\rangle\-t\)=1\-t=\\gamma^\{2\},\\qquad\|\\alpha\|\\leq\\frac\{2L\}\{\\sqrt\{1\-t^\{2\}\}\}=\\frac\{2L\}\{\\sqrt\{\(1\-t\)\(1\+t\)\}\}\\leq\\frac\{2L\}\{\\gamma\},so the unit’s sup\-norm is at most2​L​γ2L\\gamma\. Delete every cap unit withγ≤γmin:=ε′/\(64​L​m¯\)\\gamma\\leq\\gamma\_\{\\min\}:=\\varepsilon^\{\\prime\}/\(64L\\bar\{m\}\); the total deletion cost is at mostm¯⋅2​L​γmin=ε′/32\\bar\{m\}\\cdot 2L\\gamma\_\{\\min\}=\\varepsilon^\{\\prime\}/32\.

For the surviving cap units, use dyadic bins inγ\\gamma\. Letγr=2r​γmin\\gamma\_\{r\}=2^\{r\}\\gamma\_\{\\min\}and take the bins

Ir=\[γr,2​γr\]∩\[γmin,2−1/2\],r=0,1,…,R,I\_\{r\}=\[\\gamma\_\{r\},2\\gamma\_\{r\}\]\\cap\[\\gamma\_\{\\min\},2^\{\-1/2\}\],\\qquad r=0,1,\\dots,R,whereR≤C​log⁡\(2\+L​m¯/ε′\)R\\leq C\\log\(2\+L\\bar\{m\}/\\varepsilon^\{\\prime\}\)\. In one such bin,γ∈Ir\\gamma\\in I\_\{r\}implies\|α\|≤2​L/γr\|\\alpha\|\\leq 2L/\\gamma\_\{r\},supReLU≤\(2​γr\)2=4​γr2\\sup\\operatorname\{ReLU\}\\leq\(2\\gamma\_\{r\}\)^\{2\}=4\\gamma\_\{r\}^\{2\}, and thett\-interval has length at most\(2​γr\)2−γr2=3​γr2\(2\\gamma\_\{r\}\)^\{2\}\-\\gamma\_\{r\}^\{2\}=3\\gamma\_\{r\}^\{2\}\. Discretize inside the bin as follows:

- •α∈\[−2​L/γr,2​L/γr\]\\alpha\\in\[\-2L/\\gamma\_\{r\},2L/\\gamma\_\{r\}\]on a grid of meshε′/\(128​m¯​γr2\)\\varepsilon^\{\\prime\}/\(128\\bar\{m\}\\gamma\_\{r\}^\{2\}\)\. SincesupReLU≤4​γr2\\sup\\operatorname\{ReLU\}\\leq 4\\gamma\_\{r\}^\{2\}, this contributes at mostε′/\(32​m¯\)\\varepsilon^\{\\prime\}/\(32\\bar\{m\}\); the number of grid points is at most1\+C​L​m¯​γr/ε′≤C​L​m¯/ε′1\+CL\\bar\{m\}\\gamma\_\{r\}/\\varepsilon^\{\\prime\}\\leq CL\\bar\{m\}/\\varepsilon^\{\\prime\}\.
- •tton a grid of meshε′​γr/\(64​L​m¯\)\\varepsilon^\{\\prime\}\\gamma\_\{r\}/\(64L\\bar\{m\}\)\. The ReLU map is11\-Lipschitz intt, so the contribution is at most\(2​L/γr\)⋅ε′​γr/\(64​L​m¯\)=ε′/\(32​m¯\)\(2L/\\gamma\_\{r\}\)\\cdot\\varepsilon^\{\\prime\}\\gamma\_\{r\}/\(64L\\bar\{m\}\)=\\varepsilon^\{\\prime\}/\(32\\bar\{m\}\); the number of grid points is at most1\+C​L​m¯​γr/ε′≤C​L​m¯/ε′1\+CL\\bar\{m\}\\gamma\_\{r\}/\\varepsilon^\{\\prime\}\\leq CL\\bar\{m\}/\\varepsilon^\{\\prime\}\.
- •uuon anε′​γr/\(128​L​m¯\)\\varepsilon^\{\\prime\}\\gamma\_\{r\}/\(128L\\bar\{m\}\)\-net of𝕊d−1\\mathbb\{S\}^\{d\-1\}\. The contribution is at most\(2​L/γr\)⋅ε′​γr/\(128​L​m¯\)=ε′/\(64​m¯\)\(2L/\\gamma\_\{r\}\)\\cdot\\varepsilon^\{\\prime\}\\gamma\_\{r\}/\(128L\\bar\{m\}\)=\\varepsilon^\{\\prime\}/\(64\\bar\{m\}\); the number of net points is at most\(C​L​m¯/\(ε′​γr\)\)d≤\(C​\(L​m¯\)2/ε′⁣2\)d\\bigl\(CL\\bar\{m\}/\(\\varepsilon^\{\\prime\}\\gamma\_\{r\}\)\\bigr\)^\{d\}\\leq\\bigl\(C\(L\\bar\{m\}\)^\{2\}/\\varepsilon^\{\\prime 2\}\\bigr\)^\{d\}\.

The factorRRfor the choice of dyadic bin costs onlylog⁡R≤C​log⁡\(C​m¯​d​\(2\+L\)/ε′\)\\log R\\leq C\\log\\bigl\(C\\bar\{m\}d\(2\+L\)/\\varepsilon^\{\\prime\}\\bigr\)after increasing constants \(ifγmin≥2−1/2\\gamma\_\{\\min\}\\geq 2^\{\-1/2\}there are no surviving cap units\)\. Thus a cap unit has the same logarithmic count as in the ball case, up to the harmless factor2​d​log⁡\(C​L​m¯/ε′\)2d\\log\(CL\\bar\{m\}/\\varepsilon^\{\\prime\}\)coming from theuu\-net\. Each surviving cap unit contributes at mostε′/\(32​m¯\)\+ε′/\(32​m¯\)\+ε′/\(64​m¯\)<ε′/\(8​m¯\)\\varepsilon^\{\\prime\}/\(32\\bar\{m\}\)\+\\varepsilon^\{\\prime\}/\(32\\bar\{m\}\)\+\\varepsilon^\{\\prime\}/\(64\\bar\{m\}\)<\\varepsilon^\{\\prime\}/\(8\\bar\{m\}\), and deleted caps contributeε′/32\\varepsilon^\{\\prime\}/32in total\.

The affine part uses the sphere bounds of Lemma[3\.6](https://arxiv.org/html/2607.07778#S3.Thmtheorem6)\(ii\):‖v‖≤d​\(B0\+2​L​m¯\)\\\|v\\\|\\leq d\(B\_\{0\}\+2L\\bar\{m\}\)enlarges thevv\-net’s range by a factordd, costing an additivelog⁡d\\log d\. Each unit is gridded in exactly one regime \(bulk, cap\-bin, or deleted\); enumerating these choices over at mostm¯\\bar\{m\}units is absorbed intoC​m¯​d​log⁡\(⋯\)C\\bar\{m\}d\\log\(\\cdots\)\. Collecting terms, the sphere bound is againlog⁡\|𝒩∗\|≤C​m¯​d​log⁡\(C​m¯​d​\(2\+L\)/ε′\)\\log\|\\mathcal\{N\}^\{\\ast\}\|\\leq C\\bar\{m\}d\\log\\bigl\(C\\bar\{m\}d\(2\+L\)/\\varepsilon^\{\\prime\}\\bigr\)\. The sphere contributions to theL∞L^\{\\infty\}\-error total as in the ball case—at mostε′/4\\varepsilon^\{\\prime\}/4over the units,ε′/32\\varepsilon^\{\\prime\}/32for the deleted caps, andε′/8\\varepsilon^\{\\prime\}/8each forvvandcc, hence at most17​ε′/32<ε′17\\varepsilon^\{\\prime\}/32<\\varepsilon^\{\\prime\}\.

*Internal net\.*First take an\(ε′/2\)\(\\varepsilon^\{\\prime\}/2\)\-cover of𝒢m¯,L\\mathcal\{G\}\_\{\\bar\{m\},L\}with the cardinality just obtained\. We now build an internal separated set inside𝒜m¯,L\\mathcal\{A\}\_\{\\bar\{m\},L\}\. Start withS=∅S=\\emptysetand, as long as there is a point of𝒜m¯,L\\mathcal\{A\}\_\{\\bar\{m\},L\}whose distance from all points already chosen is greater thanε′\\varepsilon^\{\\prime\}, add such a point toSS\. Every time a point is added, the setSSisε′\\varepsilon^\{\\prime\}\-separated\. Since \([4](https://arxiv.org/html/2607.07778#S4.E4)\) putsSSinside𝒢m¯,L\\mathcal\{G\}\_\{\\bar\{m\},L\}, no two points ofSScan lie in the same\(ε′/2\)\(\\varepsilon^\{\\prime\}/2\)\-ball of the fixed cover; hence the process stops after at mostN​\(𝒢m¯,L,ε′/2\)N\(\\mathcal\{G\}\_\{\\bar\{m\},L\},\\varepsilon^\{\\prime\}/2\)additions\. At stopping time, maximality says that every point of𝒜m¯,L\\mathcal\{A\}\_\{\\bar\{m\},L\}is withinε′\\varepsilon^\{\\prime\}of some point ofSS\. ThusS⊂𝒜m¯,LS\\subset\\mathcal\{A\}\_\{\\bar\{m\},L\}is an internalε′\\varepsilon^\{\\prime\}\-net and

\|S\|≤N\(𝒢m¯,L,∥⋅∥∞,ε′/2\),\|S\|\\leq N\(\\mathcal\{G\}\_\{\\bar\{m\},L\},\\\|\\cdot\\\|\_\{\\infty\},\\varepsilon^\{\\prime\}/2\),which has the same logarithmic bound after adjusting the absolute constant\.

∎

### A\.5\.Proofs for Section[5](https://arxiv.org/html/2607.07778#S5)

###### Proof of Lemma[5\.2](https://arxiv.org/html/2607.07778#S5.Thmtheorem2)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S5.Thmtheorem2)\(a\) Lévy’s concentration on the sphere, in sub\-Gaussian form \(see, e\.g\.,\[[7](https://arxiv.org/html/2607.07778#bib.bib7), Ch\. 5\]\), states that for the uniform measure on the sphere of radiusd\\sqrt\{d\}a11\-Lipschitz function isψ2\\psi\_\{2\}\-close to its mean with an absolute constant\. Givenffon𝕊d−1\\mathbb\{S\}^\{d\-1\}that isL′L^\{\\prime\}\-Lipschitz, apply this toz↦f​\(z/d\)z\\mapsto f\(z/\\sqrt\{d\}\)on the radius\-d\\sqrt\{d\}sphere, which is\(L′/d\)\(L^\{\\prime\}/\\sqrt\{d\}\)\-Lipschitz; this yields‖f​\(x\)−𝔼​f‖ψ2≤κ​L′/d\\\|f\(x\)\-\\mathbb\{E\}f\\\|\_\{\\psi\_\{2\}\}\\leq\\kappa L^\{\\prime\}/\\sqrt\{d\}\. \(b\) Forg∼N​\(0,Id\)g\\sim N\(0,I\_\{d\}\)puth​\(g\):=f​\(g/d\)h\(g\):=f\(g/\\sqrt\{d\}\), which is\(L′/d\)\(L^\{\\prime\}/\\sqrt\{d\}\)\-Lipschitz; the Gaussian concentration inequality \(see, e\.g\.,\[[7](https://arxiv.org/html/2607.07778#bib.bib7), Ch\. 5\]\) gives‖h−𝔼​h‖ψ2≤C​L′/d\\\|h\-\\mathbb\{E\}h\\\|\_\{\\psi\_\{2\}\}\\leq CL^\{\\prime\}/\\sqrt\{d\}, i\.e\. the claim forx=g/d∼N​\(0,Id/d\)x=g/\\sqrt\{d\}\\sim N\(0,I\_\{d\}/d\)\. ∎

###### Proof of Lemma[5\.3](https://arxiv.org/html/2607.07778#S5.Thmtheorem3)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S5.Thmtheorem3)LetE1:=\{1n​∑izi2≥σ2−ε6\}E\_\{1\}:=\\\{\\tfrac\{1\}\{n\}\\sum\_\{i\}z\_\{i\}^\{2\}\\geq\\sigma^\{2\}\-\\tfrac\{\\varepsilon\}\{6\}\\\}andE2:=\{1n​∑izi​g​\(xi\)≥−ε6\}E\_\{2\}:=\\\{\\tfrac\{1\}\{n\}\\sum\_\{i\}z\_\{i\}g\(x\_\{i\}\)\\geq\-\\tfrac\{\\varepsilon\}\{6\}\\\}\. Thezi2∈\[0,4\]z\_\{i\}^\{2\}\\in\[0,4\]are i\.i\.d\. with meanσ2\\sigma^\{2\}; Hoeffding’s inequality\[[7](https://arxiv.org/html/2607.07778#bib.bib7), Thm\. 2\.2\.6\]for variables in an interval of length44givesℙ​\(E1c\)≤exp⁡\(−2​n​\(ε/6\)2/42\)=e−n​ε2/288\\mathbb\{P\}\(E\_\{1\}^\{c\}\)\\leq\\exp\(\-2n\(\\varepsilon/6\)^\{2\}/4^\{2\}\)=e^\{\-n\\varepsilon^\{2\}/288\}\. Thezi​g​\(xi\)∈\[−2,2\]z\_\{i\}g\(x\_\{i\}\)\\in\[\-2,2\]are i\.i\.d\. with mean𝔼​\[g​\(x\)​𝔼​\[z∣x\]\]=0\\mathbb\{E\}\[g\(x\)\\mathbb\{E\}\[z\\mid x\]\]=0; likewiseℙ​\(E2c\)≤e−n​ε2/288\\mathbb\{P\}\(E\_\{2\}^\{c\}\)\\leq e^\{\-n\\varepsilon^\{2\}/288\}\.

OnE1∩E2E\_\{1\}\\cap E\_\{2\}, suppose1n​∑i\(f​\(xi\)−yi\)2≤σ2−ε\\tfrac\{1\}\{n\}\\sum\_\{i\}\(f\(x\_\{i\}\)\-y\_\{i\}\)^\{2\}\\leq\\sigma^\{2\}\-\\varepsilonfor somef∈ℱf\\in\\mathcal\{F\}\. Writingyi=g​\(xi\)\+ziy\_\{i\}=g\(x\_\{i\}\)\+z\_\{i\}, sof−y=\(f−g\)−zf\-y=\(f\-g\)\-zand\(f−y\)2=\(f−g\)2−2​z​\(f−g\)\+z2\(f\-y\)^\{2\}=\(f\-g\)^\{2\}\-2z\(f\-g\)\+z^\{2\}; averaging,

σ2−ε≥1n​∑i\(f−g\)2​\(xi\)⏟≥0−2n​∑izi​\(f−g\)​\(xi\)\+1n​∑izi2\.\\sigma^\{2\}\-\\varepsilon\\ \\geq\\ \\underbrace\{\\tfrac\{1\}\{n\}\\textstyle\\sum\_\{i\}\(f\-g\)^\{2\}\(x\_\{i\}\)\}\_\{\\geq 0\}\\ \-\\ \\tfrac\{2\}\{n\}\\textstyle\\sum\_\{i\}z\_\{i\}\(f\-g\)\(x\_\{i\}\)\\ \+\\ \\tfrac\{1\}\{n\}\\textstyle\\sum\_\{i\}z\_\{i\}^\{2\}\.Now−2n​∑zi​\(f−g\)=−2n​∑zi​f\+2n​∑zi​g≥−2n​∑zi​f−ε3\-\\tfrac\{2\}\{n\}\\sum z\_\{i\}\(f\-g\)=\-\\tfrac\{2\}\{n\}\\sum z\_\{i\}f\+\\tfrac\{2\}\{n\}\\sum z\_\{i\}g\\geq\-\\tfrac\{2\}\{n\}\\sum z\_\{i\}f\-\\tfrac\{\\varepsilon\}\{3\}onE2E\_\{2\}, and1n​∑zi2≥σ2−ε6\\tfrac\{1\}\{n\}\\sum z\_\{i\}^\{2\}\\geq\\sigma^\{2\}\-\\tfrac\{\\varepsilon\}\{6\}onE1E\_\{1\}\. Hence

σ2−ε≥0−2n​∑izi​f​\(xi\)−ε3\+σ2−ε6,\\sigma^\{2\}\-\\varepsilon\\ \\geq\\ 0\-\\tfrac\{2\}\{n\}\\textstyle\\sum\_\{i\}z\_\{i\}f\(x\_\{i\}\)\-\\tfrac\{\\varepsilon\}\{3\}\+\\sigma^\{2\}\-\\tfrac\{\\varepsilon\}\{6\},i\.e\.2n​∑izi​f​\(xi\)≥ε−ε3−ε6=ε2\\tfrac\{2\}\{n\}\\sum\_\{i\}z\_\{i\}f\(x\_\{i\}\)\\geq\\varepsilon\-\\tfrac\{\\varepsilon\}\{3\}\-\\tfrac\{\\varepsilon\}\{6\}=\\tfrac\{\\varepsilon\}\{2\}, so1n​∑izi​f​\(xi\)≥ε4\\tfrac\{1\}\{n\}\\sum\_\{i\}z\_\{i\}f\(x\_\{i\}\)\\geq\\tfrac\{\\varepsilon\}\{4\}\. Thus the fitting event onE1∩E2E\_\{1\}\\cap E\_\{2\}implies the second event; the claim follows by a union bound withℙ​\(E1c\)\+ℙ​\(E2c\)≤2​e−n​ε2/288\\mathbb\{P\}\(E\_\{1\}^\{c\}\)\+\\mathbb\{P\}\(E\_\{2\}^\{c\}\)\\leq 2e^\{\-n\\varepsilon^\{2\}/288\}\. ∎

###### Proof of Lemma[5\.4](https://arxiv.org/html/2607.07778#S5.Thmtheorem4)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S5.Thmtheorem4)𝔼​Wi=𝔼​\[\(f​\(xi\)−𝔼​f\)​𝔼​\[zi∣xi\]\]=0\\mathbb\{E\}W\_\{i\}=\\mathbb\{E\}\[\(f\(x\_\{i\}\)\-\\mathbb\{E\}f\)\\mathbb\{E\}\[z\_\{i\}\\mid x\_\{i\}\]\]=0\. Two tail bounds onWiW\_\{i\}: first,\|Wi\|≤\|zi\|​\|f​\(xi\)−𝔼​f\|≤2⋅2=4\|W\_\{i\}\|\\leq\|z\_\{i\}\|\\,\|f\(x\_\{i\}\)\-\\mathbb\{E\}f\|\\leq 2\\cdot 2=4, so‖Wi‖ψ2≤C\\\|W\_\{i\}\\\|\_\{\\psi\_\{2\}\}\\leq C\(any bounded variable is sub\-Gaussian\)\. Second,\|Wi\|≤2​\|f​\(xi\)−𝔼​f\|\|W\_\{i\}\|\\leq 2\|f\(x\_\{i\}\)\-\\mathbb\{E\}f\|, soℙ​\(\|Wi\|≥t\)≤ℙ​\(\|f​\(xi\)−𝔼​f\|≥t/2\)\\mathbb\{P\}\(\|W\_\{i\}\|\\geq t\)\\leq\\mathbb\{P\}\(\|f\(x\_\{i\}\)\-\\mathbb\{E\}f\|\\geq t/2\); by Definition[5\.1](https://arxiv.org/html/2607.07778#S5.Thmtheorem1),‖f​\(xi\)−𝔼​f‖ψ2≤κ​L/d\\\|f\(x\_\{i\}\)\-\\mathbb\{E\}f\\\|\_\{\\psi\_\{2\}\}\\leq\\kappa L/\\sqrt\{d\}, hence‖Wi‖ψ2≤2​κ​L/d\\\|W\_\{i\}\\\|\_\{\\psi\_\{2\}\}\\leq 2\\kappa L/\\sqrt\{d\}up to an absolute factor\. Combining,∥Wi∥ψ2≤Cmin\(1,κL/d\)=:K\\\|W\_\{i\}\\\|\_\{\\psi\_\{2\}\}\\leq C\\min\\bigl\(1,\\kappa L/\\sqrt\{d\}\\bigr\)=:K\. Standard concentration for sums of independent centered sub\-Gaussian variables \(see, e\.g\.,\[[7](https://arxiv.org/html/2607.07778#bib.bib7), Ch\. 2\]\) givesℙ​\(∑iWi≥s\)≤exp⁡\(−c​s2/\(n​K2\)\)\\mathbb\{P\}\(\\sum\_\{i\}W\_\{i\}\\geq s\)\\leq\\exp\(\-c\\,s^\{2\}/\(nK^\{2\}\)\); takes=n​ε/8s=n\\varepsilon/8and note1/K2≥c′​max⁡\(1,d/\(κ2​L2\)\)1/K^\{2\}\\geq c^\{\\prime\}\\max\(1,d/\(\\kappa^\{2\}L^\{2\}\)\)\. ∎

###### Proof of Lemma[5\.5](https://arxiv.org/html/2607.07778#S5.Thmtheorem5)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S5.Thmtheorem5)\|𝔼​f\|≤1\|\\mathbb\{E\}f\|\\leq 1, so\(𝔼​f\)​1n​∑zi≥ε8\(\\mathbb\{E\}f\)\\tfrac\{1\}\{n\}\\sum z\_\{i\}\\geq\\tfrac\{\\varepsilon\}\{8\}implies\|1n​∑zi\|≥ε8\|\\tfrac\{1\}\{n\}\\sum z\_\{i\}\|\\geq\\tfrac\{\\varepsilon\}\{8\}, uniformly inff\. Thezi∈\[−2,2\]z\_\{i\}\\in\[\-2,2\]are i\.i\.d\. mean0; Hoeffding givesℙ​\(\|1n​∑zi\|≥ε8\)≤2​exp⁡\(−2​n​\(ε/8\)2/42\)=2​e−n​ε2/512\\mathbb\{P\}\(\|\\tfrac\{1\}\{n\}\\sum z\_\{i\}\|\\geq\\tfrac\{\\varepsilon\}\{8\}\)\\leq 2\\exp\(\-2n\(\\varepsilon/8\)^\{2\}/4^\{2\}\)=2e^\{\-n\\varepsilon^\{2\}/512\}\. ∎

###### Proof of Theorem[5\.6](https://arxiv.org/html/2607.07778#S5.Thmtheorem6)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S5.Thmtheorem6)By Lemma[5\.3](https://arxiv.org/html/2607.07778#S5.Thmtheorem3)it suffices to boundℙ\(∃f:1n∑zif\(xi\)≥ε4\)\\mathbb\{P\}\(\\exists f:\\tfrac\{1\}\{n\}\\sum z\_\{i\}f\(x\_\{i\}\)\\geq\\tfrac\{\\varepsilon\}\{4\}\)\. Splitzi​f​\(xi\)=Wi\+\(𝔼​f\)​ziz\_\{i\}f\(x\_\{i\}\)=W\_\{i\}\+\(\\mathbb\{E\}f\)z\_\{i\}andε4=ε8\+ε8\\tfrac\{\\varepsilon\}\{4\}=\\tfrac\{\\varepsilon\}\{8\}\+\\tfrac\{\\varepsilon\}\{8\}: if1n​∑zi​f≥ε4\\tfrac\{1\}\{n\}\\sum z\_\{i\}f\\geq\\tfrac\{\\varepsilon\}\{4\}then1n​∑Wi≥ε8\\tfrac\{1\}\{n\}\\sum W\_\{i\}\\geq\\tfrac\{\\varepsilon\}\{8\}or\(𝔼​f\)​1n​∑zi≥ε8\(\\mathbb\{E\}f\)\\tfrac\{1\}\{n\}\\sum z\_\{i\}\\geq\\tfrac\{\\varepsilon\}\{8\}\. Union\-bounding Lemma[5\.4](https://arxiv.org/html/2607.07778#S5.Thmtheorem4)over the\|ℱ\|\|\\mathcal\{F\}\|functions, adding Lemma[5\.5](https://arxiv.org/html/2607.07778#S5.Thmtheorem5), and adding the2​e−n​ε2/2882e^\{\-n\\varepsilon^\{2\}/288\}of Lemma[5\.3](https://arxiv.org/html/2607.07778#S5.Thmtheorem3)\(absorbed, with the2​e−n​ε2/5122e^\{\-n\\varepsilon^\{2\}/512\}, into4​e−n​ε2/5124e^\{\-n\\varepsilon^\{2\}/512\}\) gives the bound\. ∎

### A\.6\.Proofs for Section[7](https://arxiv.org/html/2607.07778#S7)

###### Proof of Lemma[7\.1](https://arxiv.org/html/2607.07778#S7.Thmtheorem1)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S7.Thmtheorem1)\|cF\|≤sup\|F\|≤1\|c\_\{F\}\|\\leq\\sup\|F\|\\leq 1\. Sinceclip\\operatorname\{clip\}contracts and chords bound arcs from below,FFisLL\-Lipschitz for geodesic distance, hence lies inH1​\(𝕊d−1\)H^\{1\}\(\\mathbb\{S\}^\{d\-1\}\)with\|∇TF\|≤L\|\\nabla\_\{T\}F\|\\leq Lalmost everywhere, and Parseval for the gradient gives∑ℓ≥1λℓ​‖Fℓ‖22=𝔼​\|∇TF\|2≤L2\\sum\_\{\\ell\\geq 1\}\\lambda\_\{\\ell\}\\\|F\_\{\\ell\}\\\|\_\{2\}^\{2\}=\\mathbb\{E\}\|\\nabla\_\{T\}F\|^\{2\}\\leq L^\{2\}\. The degree\-one component isF1​\(x\)=⟨AF,x⟩F\_\{1\}\(x\)=\\langle A\_\{F\},x\\ranglewithAF=d​𝔼​\[F​x\]A\_\{F\}=d\\,\\mathbb\{E\}\[Fx\]\(because𝔼​\[x​x⊤\]=Id/d\\mathbb\{E\}\[xx^\{\\top\}\]=I\_\{d\}/d\), and‖F1‖22=‖AF‖2/d\\\|F\_\{1\}\\\|\_\{2\}^\{2\}=\\\|A\_\{F\}\\\|^\{2\}/d; sinceλ1=d−1\\lambda\_\{1\}=d\-1, this gives‖AF‖2≤d​L2/\(d−1\)≤32​L2\\\|A\_\{F\}\\\|^\{2\}\\leq dL^\{2\}/\(d\-1\)\\leq\\tfrac\{3\}\{2\}L^\{2\}ford≥3d\\geq 3\. ForhF=∑ℓ≥2Fℓh\_\{F\}=\\sum\_\{\\ell\\geq 2\}F\_\{\\ell\}:𝔼​hF2≤λ2−1​∑ℓ≥2λℓ​‖Fℓ‖22≤L2/\(2​d\)\\mathbb\{E\}h\_\{F\}^\{2\}\\leq\\lambda\_\{2\}^\{\-1\}\\sum\_\{\\ell\\geq 2\}\\lambda\_\{\\ell\}\\\|F\_\{\\ell\}\\\|\_\{2\}^\{2\}\\leq L^\{2\}/\(2d\)becauseλ2=2​d\\lambda\_\{2\}=2d\. Finally\|hF\|≤\|F\|\+\|cF\|\+‖AF‖≤B1\|h\_\{F\}\|\\leq\|F\|\+\|c\_\{F\}\|\+\\\|A\_\{F\}\\\|\\leq B\_\{1\}pointwise\. ∎

###### Proof of Lemma[7\.2](https://arxiv.org/html/2607.07778#S7.Thmtheorem2)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S7.Thmtheorem2)The mapF↦hFF\\mapsto h\_\{F\}satisfies‖hF−hG‖∞≤\(2\+d\)​‖F−G‖∞\\\|h\_\{F\}\-h\_\{G\}\\\|\_\{\\infty\}\\leq\(2\+\\sqrt\{d\}\)\\\|F\-G\\\|\_\{\\infty\}: indeed\|cF−cG\|≤‖F−G‖∞\|c\_\{F\}\-c\_\{G\}\|\\leq\\\|F\-G\\\|\_\{\\infty\}and‖AF−AG‖=d​‖𝔼​\[\(F−G\)​x\]‖=d​sup‖w‖=1𝔼​\[\(F−G\)​⟨w,x⟩\]≤d​‖F−G‖2​\(𝔼​⟨w,x⟩2\)1/2=d​‖F−G‖2\\\|A\_\{F\}\-A\_\{G\}\\\|=d\\,\\\|\\mathbb\{E\}\[\(F\-G\)x\]\\\|=d\\sup\_\{\\\|w\\\|=1\}\\mathbb\{E\}\[\(F\-G\)\\langle w,x\\rangle\]\\leq d\\\|F\-G\\\|\_\{2\}\\,\\bigl\(\\mathbb\{E\}\\langle w,x\\rangle^\{2\}\\bigr\)^\{1/2\}=\\sqrt\{d\}\\,\\\|F\-G\\\|\_\{2\}\. Sinceclip\\operatorname\{clip\}contracts values, Proposition[4\.1](https://arxiv.org/html/2607.07778#S4.Thmtheorem1)transfers toℋ\\mathcal\{H\}with the factor\(2\+d\)\(2\+\\sqrt\{d\}\)absorbed into the logarithm, giving the entropy bound\. For the second claim: conditionally onx1,…,xnx\_\{1\},\\dots,x\_\{n\}, the Rademacher processh↦1n​∑iεi​h​\(xi\)h\\mapsto\\tfrac\{1\}\{\\sqrt\{n\}\}\\sum\_\{i\}\\varepsilon\_\{i\}h\(x\_\{i\}\)has sub\-Gaussian increments inL2​\(Pn\)L^\{2\}\(P\_\{n\}\), the class is pinned at0∈ℋ0\\in\\mathcal\{H\}withL2​\(Pn\)L^\{2\}\(P\_\{n\}\)\-diameter at most2​σ^2\\hat\{\\sigma\}, andN​\(ℋ,L2​\(Pn\),u\)≤N​\(ℋ,L∞,u\)N\(\\mathcal\{H\},L^\{2\}\(P\_\{n\}\),u\)\\leq N\(\\mathcal\{H\},L^\{\\infty\},u\); Dudley’s entropy integral\[[7](https://arxiv.org/html/2607.07778#bib.bib7), Thm\. 8\.1\.3\]gives the bound with∫0rlog⁡\(B′/u\)​𝑑u≤r​\[log⁡\(B′/r\)\+π2\]≤2​r​log⁡\(e​B′/r\)\\int\_\{0\}^\{r\}\\sqrt\{\\log\(B^\{\\prime\}/u\)\}\\,du\\leq r\\bigl\[\\sqrt\{\\log\(B^\{\\prime\}/r\)\}\+\\tfrac\{\\sqrt\{\\pi\}\}\{2\}\\bigr\]\\leq 2r\\sqrt\{\\log\(eB^\{\\prime\}/r\)\}\(substituteu=r​vu=rv\), applied atr=σ^r=\\hat\{\\sigma\}\. ∎

###### Proof of Lemma[7\.3](https://arxiv.org/html/2607.07778#S7.Thmtheorem3)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S7.Thmtheorem3)𝔼​suphPn​h2≤suph𝔼​h2\+𝔼​suph\(Pn−𝔼\)​h2\\mathbb\{E\}\\sup\_\{h\}P\_\{n\}h^\{2\}\\leq\\sup\_\{h\}\\mathbb\{E\}h^\{2\}\+\\mathbb\{E\}\\sup\_\{h\}\(P\_\{n\}\-\\mathbb\{E\}\)h^\{2\}, and the first term is at mostL2/\(2​d\)L^\{2\}/\(2d\)by Lemma[7\.1](https://arxiv.org/html/2607.07778#S7.Thmtheorem1)\. By symmetrization,𝔼​suph\(Pn−𝔼\)​h2≤2​𝔼x,ε​suph1n​\|∑iεi​h2​\(xi\)\|\\mathbb\{E\}\\sup\_\{h\}\(P\_\{n\}\-\\mathbb\{E\}\)h^\{2\}\\leq 2\\,\\mathbb\{E\}\_\{x,\\varepsilon\}\\sup\_\{h\}\\tfrac\{1\}\{n\}\|\\sum\_\{i\}\\varepsilon\_\{i\}h^\{2\}\(x\_\{i\}\)\|\. The maps↦s2/\(2​B1\)s\\mapsto s^\{2\}/\(2B\_\{1\}\)is a contraction on\[−B1,B1\]\[\-B\_\{1\},B\_\{1\}\]vanishing at0, so the Ledoux–Talagrand contraction principle\[[11](https://arxiv.org/html/2607.07778#bib.bib11), Thm\. 4\.12\]gives𝔼ε​suph\|∑εi​h2​\(xi\)\|≤4​B1​𝔼ε​suph\|∑εi​h​\(xi\)\|\\mathbb\{E\}\_\{\\varepsilon\}\\sup\_\{h\}\|\\sum\\varepsilon\_\{i\}h^\{2\}\(x\_\{i\}\)\|\\leq 4B\_\{1\}\\,\\mathbb\{E\}\_\{\\varepsilon\}\\sup\_\{h\}\|\\sum\\varepsilon\_\{i\}h\(x\_\{i\}\)\|, and the claim follows\. ∎

###### Proof of Theorem[7\.4](https://arxiv.org/html/2607.07778#S7.Thmtheorem4)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S7.Thmtheorem4)SplitF=cF\+⟨AF,x⟩\+hFF=c\_\{F\}\+\\langle A\_\{F\},x\\rangle\+h\_\{F\}by Lemma[7\.1](https://arxiv.org/html/2607.07778#S7.Thmtheorem1)\. The affine sector obeys

𝔼​sup\|c\|≤1,‖A‖≤3/2​L1n​\|∑iyi​\(c\+⟨A,xi⟩\)\|≤1n​\(𝔼​\|∑iyi\|\+32​L​𝔼​‖∑iyi​xi‖\)≤1\+3/2​Ln,\\mathbb\{E\}\\sup\_\{\|c\|\\leq 1,\\ \\\|A\\\|\\leq\\sqrt\{3/2\}L\}\\frac\{1\}\{n\}\\Bigl\|\\sum\_\{i\}y\_\{i\}\(c\+\\langle A,x\_\{i\}\\rangle\)\\Bigr\|\\leq\\frac\{1\}\{n\}\\Bigl\(\\mathbb\{E\}\\bigl\|\\sum\_\{i\}y\_\{i\}\\bigr\|\+\\sqrt\{\\tfrac\{3\}\{2\}\}L\\,\\mathbb\{E\}\\bigl\\\|\\sum\_\{i\}y\_\{i\}x\_\{i\}\\bigr\\\|\\Bigr\)\\leq\\frac\{1\+\\sqrt\{3/2\}\\,L\}\{\\sqrt\{n\}\},using𝔼​‖∑iyi​xi‖2=∑i𝔼​‖xi‖2=n\\mathbb\{E\}\\\|\\sum\_\{i\}y\_\{i\}x\_\{i\}\\\|^\{2\}=\\sum\_\{i\}\\mathbb\{E\}\\\|x\_\{i\}\\\|^\{2\}=n\. For theℋ\\mathcal\{H\}\-sector, by symmetry of theyiy\_\{i\}it suffices to boundR~\\widetilde\{R\}\. Setr2:=L2/\(2​d\)r^\{2\}:=L^\{2\}/\(2d\)anda:=C​m¯​d/na:=C\\sqrt\{\\bar\{m\}d/n\}\. Writeψ​\(s\):=φ​\(s\)=s2​log⁡\(\(2​e​B\)2/s\)\\psi\(s\):=\\varphi\(\\sqrt\{s\}\)=\\sqrt\{\\tfrac\{s\}\{2\}\\log\(\(2eB\)^\{2\}/s\)\}; on\(0,B2\]\(0,B^\{2\}\]the functions↦s2​log⁡\(\(2​e​B\)2/s\)s\\mapsto\\tfrac\{s\}\{2\}\\log\(\(2eB\)^\{2\}/s\)is increasing \(its derivative is12​\[log⁡\(\(2​e​B\)2/s\)−1\]\>0\\tfrac\{1\}\{2\}\[\\log\(\(2eB\)^\{2\}/s\)\-1\]\>0fors<\(2​e​B\)2/es<\(2eB\)^\{2\}/e\) and concave \(second derivative−1/\(2​s\)\-1/\(2s\)\), soψ\\psiis increasing and concave\. By Lemma[7\.2](https://arxiv.org/html/2607.07778#S7.Thmtheorem2), Jensen, and Lemma[7\.3](https://arxiv.org/html/2607.07778#S7.Thmtheorem3),

R~≤a​𝔼​φ​\(σ^\)=a​𝔼​ψ​\(σ^2\)≤a​ψ​\(𝔼​σ^2\)≤a​ψ​\(r2\+8​B1​R~\)\.\\widetilde\{R\}\\ \\leq\\ a\\,\\mathbb\{E\}\\varphi\(\\hat\{\\sigma\}\)\\ =\\ a\\,\\mathbb\{E\}\\psi\(\\hat\{\\sigma\}^\{2\}\)\\ \\leq\\ a\\,\\psi\\bigl\(\\mathbb\{E\}\\hat\{\\sigma\}^\{2\}\\bigr\)\\ \\leq\\ a\\,\\psi\\bigl\(r^\{2\}\+8B\_\{1\}\\widetilde\{R\}\\bigr\)\.The key evaluation: withr=L/2​dr=L/\\sqrt\{2d\},

log⁡\(2​e​B\)22​r2=2​log⁡2​e​B​dL=2​log⁡\(2​e​C1​m¯​d5/2​2\+LL\)≤C′′​ΛL,\\log\\frac\{\(2eB\)^\{2\}\}\{2r^\{2\}\}\\;=\\;2\\log\\frac\{2eB\\sqrt\{d\}\}\{L\}\\;=\\;2\\log\\Bigl\(2eC\_\{1\}\\bar\{m\}d^\{5/2\}\\,\\frac\{2\+L\}\{L\}\\Bigr\)\\;\\leq\\;C^\{\\prime\\prime\}\\Lambda\_\{L\},since\(2\+L\)/L≤2​\(2\+1/L\)\(2\+L\)/L\\leq 2\(2\+1/L\)andd5/2≤d3d^\{5/2\}\\leq d^\{3\}\. If8​B1​R~≤r28B\_\{1\}\\widetilde\{R\}\\leq r^\{2\}, thenR~≤a​ψ​\(2​r2\)≤a​r​C′′​ΛL=C′​L​m¯​ΛL/n\\widetilde\{R\}\\leq a\\psi\(2r^\{2\}\)\\leq a\\,r\\sqrt\{C^\{\\prime\\prime\}\\Lambda\_\{L\}\}=C^\{\\prime\}L\\sqrt\{\\bar\{m\}\\Lambda\_\{L\}/n\}\. Otherwises:=r2\+8​B1​R~≤16​B1​R~s:=r^\{2\}\+8B\_\{1\}\\widetilde\{R\}\\leq 16B\_\{1\}\\widetilde\{R\}whiles≥2​r2s\\geq 2r^\{2\}, solog⁡\(\(2​e​B\)2/s\)≤log⁡\(\(2​e​B\)2/\(2​r2\)\)≤C′′​ΛL\\log\(\(2eB\)^\{2\}/s\)\\leq\\log\(\(2eB\)^\{2\}/\(2r^\{2\}\)\)\\leq C^\{\\prime\\prime\}\\Lambda\_\{L\}\(the logarithm decreases inss\), whenceψ​\(s\)2≤8​B1​R~​C′′​ΛL\\psi\(s\)^\{2\}\\leq 8B\_\{1\}\\widetilde\{R\}\\,C^\{\\prime\\prime\}\\Lambda\_\{L\}andR~≤a​8​C′′​B1​R~​ΛL\\widetilde\{R\}\\leq a\\sqrt\{8C^\{\\prime\\prime\}B\_\{1\}\\widetilde\{R\}\\Lambda\_\{L\}\}, i\.e\.R~≤8​C′′​a2​B1​ΛL≤C​\(1\+L\)​m¯​d​ΛL/n\\widetilde\{R\}\\leq 8C^\{\\prime\\prime\}a^\{2\}B\_\{1\}\\Lambda\_\{L\}\\leq C\(1\+L\)\\bar\{m\}d\\,\\Lambda\_\{L\}/n\. Collecting the three contributions proves the theorem\. ∎

###### Proof of Theorem[7\.5](https://arxiv.org/html/2607.07778#S7.Thmtheorem5)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S7.Thmtheorem5)SetL∗:=c0​ε​n/\(m¯​Λ¯\)L^\{\\ast\}:=c\_\{0\}\\varepsilon\\sqrt\{n/\(\\bar\{m\}\\bar\{\\Lambda\}\)\}and suppose some fittingffhasL:=Lip𝕊d−1⁡\(f\)≤L∗L:=\\operatorname\{Lip\}\_\{\\mathbb\{S\}^\{d\-1\}\}\(f\)\\leq L^\{\\ast\}\. By Lemma[3\.5](https://arxiv.org/html/2607.07778#S3.Thmtheorem5)and Sections[2](https://arxiv.org/html/2607.07778#S2)–[3](https://arxiv.org/html/2607.07778#S3),f\|𝕊d−1∈𝒜m¯,L∗f\|\_\{\\mathbb\{S\}^\{d\-1\}\}\\in\\mathcal\{A\}\_\{\\bar\{m\},L^\{\\ast\}\}, andclip∘f\\operatorname\{clip\}\\circ ffits at least as well\. The noise decomposition \(Lemma[5\.3](https://arxiv.org/html/2607.07778#S5.Thmtheorem3), whose proof is pointwise inff\) gives, outside an event of probability2​e−n​ε2/288≤δ/42e^\{\-n\\varepsilon^\{2\}/288\}\\leq\\delta/4, that1n​∑izi​clip⁡\(f​\(xi\)\)≥ε/4\\tfrac\{1\}\{n\}\\sum\_\{i\}z\_\{i\}\\,\\operatorname\{clip\}\(f\(x\_\{i\}\)\)\\geq\\varepsilon/4withzi=yi−g​\(xi\)z\_\{i\}=y\_\{i\}\-g\(x\_\{i\}\)\. Conditionally on thexix\_\{i\}theziz\_\{i\}are independent, mean zero, and bounded by22, so symmetrization and coordinate\-wise contraction give

𝔼​supf∈𝒜m¯,L∗1n​\|∑izi​clip⁡\(f​\(xi\)\)\|≤8​𝔼​supf1n​\|∑iεi​clip⁡\(f​\(xi\)\)\|≤8​Ξ,\\mathbb\{E\}\\ \\sup\_\{f\\in\\mathcal\{A\}\_\{\\bar\{m\},L^\{\\ast\}\}\}\\frac\{1\}\{n\}\\Bigl\|\\sum\_\{i\}z\_\{i\}\\,\\operatorname\{clip\}\(f\(x\_\{i\}\)\)\\Bigr\|\\ \\leq\\ 8\\,\\mathbb\{E\}\\ \\sup\_\{f\}\\frac\{1\}\{n\}\\Bigl\|\\sum\_\{i\}\\varepsilon\_\{i\}\\,\\operatorname\{clip\}\(f\(x\_\{i\}\)\)\\Bigr\|\\ \\leq\\ 8\\,\\Xi,Ξ\\Xidenoting the right side of Theorem[7\.4](https://arxiv.org/html/2607.07778#S7.Thmtheorem4)atL=L∗L=L^\{\\ast\}\. First,ΛL∗≤C​Λ¯\\Lambda\_\{L^\{\\ast\}\}\\leq C\\bar\{\\Lambda\}: ifL∗≥1L^\{\\ast\}\\geq 1then2\+1/L∗≤32\+1/L^\{\\ast\}\\leq 3andΛL∗≤log⁡\(3​e​m¯​d3\)≤C​Λ¯\\Lambda\_\{L^\{\\ast\}\}\\leq\\log\(3e\\bar\{m\}d^\{3\}\)\\leq C\\bar\{\\Lambda\}; ifL∗<1L^\{\\ast\}<1then1/L∗=m¯​Λ¯/\(c0​ε​n\)≤m¯​Λ¯/c01/L^\{\\ast\}=\\sqrt\{\\bar\{m\}\\bar\{\\Lambda\}\}\\,/\(c\_\{0\}\\varepsilon\\sqrt\{n\}\)\\leq\\sqrt\{\\bar\{m\}\\bar\{\\Lambda\}\}/c\_\{0\}usingε​n≥1\\varepsilon\\sqrt\{n\}\\geq 1from the first sample\-size term, soΛL∗≤log⁡\(e​m¯​d3​\(2\+m¯​Λ¯/c0\)\)≤C​Λ¯\\Lambda\_\{L^\{\\ast\}\}\\leq\\log\(e\\bar\{m\}d^\{3\}\(2\+\\sqrt\{\\bar\{m\}\\bar\{\\Lambda\}\}/c\_\{0\}\)\)\\leq C\\bar\{\\Lambda\}\(aslog⁡Λ¯≤Λ¯\\log\\bar\{\\Lambda\}\\leq\\bar\{\\Lambda\}\)\. The three hypotheses now make the three terms ofΞ\\Xieach at mostε/\(192\)\\varepsilon/\(192\):\(1\+L∗\)/n≤ε/192\(1\+L^\{\\ast\}\)/\\sqrt\{n\}\\leq\\varepsilon/192from the first;L∗​m¯​ΛL∗/n≤C​c0​ε≤ε/192L^\{\\ast\}\\sqrt\{\\bar\{m\}\\Lambda\_\{L^\{\\ast\}\}/n\}\\leq Cc\_\{0\}\\varepsilon\\leq\\varepsilon/192forc0c\_\{0\}small, by the definition ofL∗L^\{\\ast\}; and\(1\+L∗\)​m¯​d​ΛL∗/n≤ε/192\(1\+L^\{\\ast\}\)\\bar\{m\}d\\Lambda\_\{L^\{\\ast\}\}/n\\leq\\varepsilon/192from the second and third \(for theL∗L^\{\\ast\}\-part,L∗​m¯​d​Λ¯/n=c0​ε​d​m¯​Λ¯/n≤ε/384L^\{\\ast\}\\bar\{m\}d\\bar\{\\Lambda\}/n=c\_\{0\}\\varepsilon d\\sqrt\{\\bar\{m\}\\bar\{\\Lambda\}/n\}\\leq\\varepsilon/384exactly whenn≥C​m¯​d2​Λ¯n\\geq C\\bar\{m\}d^\{2\}\\bar\{\\Lambda\}\)\. Hence𝔼​sup≤ε/8\\mathbb\{E\}\\sup\\leq\\varepsilon/8\. The supremum has bounded differences4/n4/nin each pair\(xi,yi\)\(x\_\{i\},y\_\{i\}\), so McDiarmid’s bounded\-differences inequality\[[12](https://arxiv.org/html/2607.07778#bib.bib12)\]givesℙ​\(sup≥ε/4\)≤e−n​ε2/512≤δ/8\\mathbb\{P\}\(\\sup\\geq\\varepsilon/4\)\\leq e^\{\-n\\varepsilon^\{2\}/512\}\\leq\\delta/8\. Together with the noise event this contradicts fitting, with total failure probability at mostδ\\delta\. ∎

### A\.7\.Proofs for Section[8](https://arxiv.org/html/2607.07778#S8)

###### Proof of Theorem[8\.1](https://arxiv.org/html/2607.07778#S8.Thmtheorem1)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S8.Thmtheorem1)*Factorization and fiber transport\.*ffdepends onxxonly through the orthogonal projectionP​xPxontoW:=span​\{w1,…,wm,v\}W:=\\mathrm\{span\}\\\{w\_\{1\},\\dots,w\_\{m\},v\\\},dimW≤m\+1\\dim W\\leq m\+1\. Supposexi,xjx\_\{i\},x\_\{j\}satisfy‖P​xi‖,‖P​xj‖≤12\\\|Px\_\{i\}\\\|,\\\|Px\_\{j\}\\\|\\leq\\tfrac\{1\}\{2\}\. Writexi=zi\+wix\_\{i\}=z\_\{i\}\+w\_\{i\}withzi=P​xiz\_\{i\}=Px\_\{i\},‖wi‖=1−‖zi‖2≥32\\\|w\_\{i\}\\\|=\\sqrt\{1\-\\\|z\_\{i\}\\\|^\{2\}\}\\geq\\tfrac\{\\sqrt\{3\}\}\{2\}, and setx′:=zj\+1−‖zj‖2​wi/‖wi‖∈𝕊d−1x^\{\\prime\}:=z\_\{j\}\+\\sqrt\{1\-\\\|z\_\{j\}\\\|^\{2\}\}\\,w\_\{i\}/\\\|w\_\{i\}\\\|\\in\\mathbb\{S\}^\{d\-1\}\. ThenP​x′=zjPx^\{\\prime\}=z\_\{j\}, sof​\(x′\)=f​\(xj\)f\(x^\{\\prime\}\)=f\(x\_\{j\}\), and

‖xi−x′‖≤‖zi−zj‖\+\|1−‖zi‖2−1−‖zj‖2\|≤\(1\+13\)​‖zi−zj‖≤2​‖zi−zj‖,\\\|x\_\{i\}\-x^\{\\prime\}\\\|\\leq\\\|z\_\{i\}\-z\_\{j\}\\\|\+\\Bigl\|\\sqrt\{1\-\\\|z\_\{i\}\\\|^\{2\}\}\-\\sqrt\{1\-\\\|z\_\{j\}\\\|^\{2\}\}\\Bigr\|\\leq\\Bigl\(1\+\\tfrac\{1\}\{\\sqrt\{3\}\}\\Bigr\)\\\|z\_\{i\}\-z\_\{j\}\\\|\\leq 2\\\|z\_\{i\}\-z\_\{j\}\\\|,using\|‖zi‖2−‖zj‖2\|≤\(‖zi‖\+‖zj‖\)​‖zi−zj‖\|\\,\\\|z\_\{i\}\\\|^\{2\}\-\\\|z\_\{j\}\\\|^\{2\}\|\\leq\(\\\|z\_\{i\}\\\|\+\\\|z\_\{j\}\\\|\)\\\|z\_\{i\}\-z\_\{j\}\\\|and the lower bound on the two square roots\. Hence

\|f​\(xi\)−f​\(xj\)\|=\|f​\(xi\)−f​\(x′\)\|≤2​L​‖P​\(xi−xj\)‖,L:=Lip𝕊d−1⁡\(f\)\.\|f\(x\_\{i\}\)\-f\(x\_\{j\}\)\|=\|f\(x\_\{i\}\)\-f\(x^\{\\prime\}\)\|\\leq 2L\\,\\\|P\(x\_\{i\}\-x\_\{j\}\)\\\|,\\qquad L:=\\operatorname\{Lip\}\_\{\\mathbb\{S\}^\{d\-1\}\}\(f\)\.\(10\)
*Few high points, uniformly\.*LetOp:=λmax​\(∑ixi​xi⊤\)\\mathrm\{Op\}:=\\lambda\_\{\\max\}\(\\sum\_\{i\}x\_\{i\}x\_\{i\}^\{\\top\}\); on an event of probability1−e−n/C1\-e^\{\-n/C\},Op≤C​\(1\+n/d\)\\mathrm\{Op\}\\leq C\(1\+n/d\)\. For any\(m\+1\)\(m\{\+\}1\)\-dimensional projectionPPwith orthonormal basise1,…,em\+1e\_\{1\},\\dots,e\_\{m\+1\},∑i‖P​xi‖2=∑k∑i⟨ek,xi⟩2≤\(m\+1\)​Op\\sum\_\{i\}\\\|Px\_\{i\}\\\|^\{2\}=\\sum\_\{k\}\\sum\_\{i\}\\langle e\_\{k\},x\_\{i\}\\rangle^\{2\}\\leq\(m\+1\)\\,\\mathrm\{Op\}, so at most4​\(m\+1\)​Op≤n/84\(m\+1\)\\mathrm\{Op\}\\leq n/8points have‖P​xi‖\>12\\\|Px\_\{i\}\\\|\>\\tfrac\{1\}\{2\}\(the last inequality by the hypotheses onmm\)\. Call the others*low*; there are at least78​n\\tfrac\{7\}\{8\}nof them, for everyPPsimultaneously\.

*Volumetric pairing, uniformly over a net\.*We may assumeδcell:=C​m\+1​\(8/n\)1/\(m\+1\)≤14\\delta\_\{\\mathrm\{cell\}\}:=C\\sqrt\{m\+1\}\\,\(8/n\)^\{1/\(m\+1\)\}\\leq\\tfrac\{1\}\{4\}: otherwise the claimed bound readsL≥c′L\\geq c^\{\\prime\}withc′c^\{\\prime\}absolute, which already follows from one opposite\-label pair \(probability1−21−n1\-2^\{1\-n\}\) and\|f​\(xi\)−f​\(xj\)\|=2\|f\(x\_\{i\}\)\-f\(x\_\{j\}\)\|=2with‖xi−xj‖≤2\\\|x\_\{i\}\-x\_\{j\}\\\|\\leq 2\. Fix a net𝒫\\mathcal\{P\}of the\(m\+1\)\(m\{\+\}1\)\-frames of column\-wise meshδnet:=δcell/\(8​m\+1\)\\delta\_\{\\mathrm\{net\}\}:=\\delta\_\{\\mathrm\{cell\}\}/\(8\\sqrt\{m\+1\}\), so that every admissiblePPhasP^∈𝒫\\hat\{P\}\\in\\mathcal\{P\}with‖P−P^‖op≤m\+1​δnet⋅2≤δcell/4≤1/16\\\|P\-\\hat\{P\}\\\|\_\{\\mathrm\{op\}\}\\leq\\sqrt\{m\+1\}\\,\\delta\_\{\\mathrm\{net\}\}\\cdot 2\\leq\\delta\_\{\\mathrm\{cell\}\}/4\\leq 1/16; the cardinality iseC​\(m\+1\)​d​log⁡\(1/δnet\)≤eC​\(d​log⁡n\+m​d\)e^\{C\(m\+1\)d\\log\(1/\\delta\_\{\\mathrm\{net\}\}\)\}\\leq e^\{C\(d\\log n\+md\)\}, sincelog⁡\(1/δnet\)≤1m\+1​log⁡n\+C\\log\(1/\\delta\_\{\\mathrm\{net\}\}\)\\leq\\tfrac\{1\}\{m\+1\}\\log n\+C\(them\+1\\sqrt\{m\+1\}factors insideδcell\\delta\_\{\\mathrm\{cell\}\}and the mesh cancel\)\. For a fixedP^∈𝒫\\hat\{P\}\\in\\mathcal\{P\}: partition the ball of radius12\\tfrac\{1\}\{2\}inP^​\(ℝd\)\\hat\{P\}\(\\mathbb\{R\}^\{d\}\)intoK=⌈n/8⌉K=\\lceil n/8\\rceilgrid cells of diameterδcell\\delta\_\{\\mathrm\{cell\}\}\. Among the≥78​n\\geq\\tfrac\{7\}\{8\}nlow points ofP^\\hat\{P\}, at least78​n−K≥34​n−1\\tfrac\{7\}\{8\}n\-K\\geq\\tfrac\{3\}\{4\}n\-1share a cell with another low point, yielding at least38​n−1≥n4\\tfrac\{3\}\{8\}n\-1\\geq\\tfrac\{n\}\{4\}disjoint same\-cell pairs \(forn≥8n\\geq 8\), each with‖P^​\(xi−xj\)‖≤δcell\\\|\\hat\{P\}\(x\_\{i\}\-x\_\{j\}\)\\\|\\leq\\delta\_\{\\mathrm\{cell\}\}\. These pairs are functions of\(x,P^\)\(x,\\hat\{P\}\)only; since the labels are independent of the data, the probability that fewer thann16\\tfrac\{n\}\{16\}of them are opposite\-label is at moste−n/Ce^\{\-n/C\}\(binomial concentration\)\. A union bound over𝒫\\mathcal\{P\}costseC​\(d​log⁡n\+m​d\)e^\{C\(d\\log n\+md\)\}, which the sample\-size hypothesis covers\. Finally, for the truePP: projected distances of unit\-norm differences move by at most2​‖P−P^‖op≤δcell/22\\\|P\-\\hat\{P\}\\\|\_\{\\mathrm\{op\}\}\\leq\\delta\_\{\\mathrm\{cell\}\}/2, so the pair satisfies‖P​\(xi−xj\)‖≤32​δcell\\\|P\(x\_\{i\}\-x\_\{j\}\)\\\|\\leq\\tfrac\{3\}\{2\}\\delta\_\{\\mathrm\{cell\}\}; and each point of the pair, low forP^\\hat\{P\}, has‖P​xi‖≤12\+116≤916\\\|Px\_\{i\}\\\|\\leq\\tfrac\{1\}\{2\}\+\\tfrac\{1\}\{16\}\\leq\\tfrac\{9\}\{16\}, for which the transport estimate \([10](https://arxiv.org/html/2607.07778#A1.E10)\) holds with the constant22unchanged \(the square roots in its proof are bounded below by1−\(9/16\)2≥45\\sqrt\{1\-\(9/16\)^\{2\}\}\\geq\\tfrac\{4\}\{5\}, giving factor1\+9/16⋅22⋅4/5≤1\.71≤21\+\\tfrac\{9/16\\cdot 2\}\{2\\cdot 4/5\}\\leq 1\.71\\leq 2\)\. So with probability1−2​e−n/C1\-2e^\{\-n/C\}, for*every*admissiblePPthere is an opposite\-label pair with‖P​\(xi−xj\)‖≤32​δcell\\\|P\(x\_\{i\}\-x\_\{j\}\)\\\|\\leq\\tfrac\{3\}\{2\}\\delta\_\{\\mathrm\{cell\}\}and both points916\\tfrac\{9\}\{16\}\-low\.

*Conclusion\.*For that pair, exact fitting gives\|f​\(xi\)−f​\(xj\)\|=\|yi−yj\|=2\|f\(x\_\{i\}\)\-f\(x\_\{j\}\)\|=\|y\_\{i\}\-y\_\{j\}\|=2, while \([10](https://arxiv.org/html/2607.07778#A1.E10)\) gives2≤2​L⋅32​δcell2\\leq 2L\\cdot\\tfrac\{3\}\{2\}\\delta\_\{\\mathrm\{cell\}\}, i\.e\.L≥23​δcell−1=c​n1/\(m\+1\)/m\+1L\\geq\\tfrac\{2\}\{3\}\\delta\_\{\\mathrm\{cell\}\}^\{\-1\}=c\\,n^\{1/\(m\+1\)\}/\\sqrt\{m\+1\}\. ∎

###### Proof of Theorem[8\.3](https://arxiv.org/html/2607.07778#S8.Thmtheorem3)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S8.Thmtheorem3)WriteL=Lip𝕊d−1⁡\(f\)L=\\operatorname\{Lip\}\_\{\\mathbb\{S\}^\{d\-1\}\}\(f\)\. As in the proof of Theorem[8\.1](https://arxiv.org/html/2607.07778#S8.Thmtheorem1),fffactors through the orthogonal projectionPPonto a subspaceW⊇span​\{w1,…,wm,v\}W\\supseteq\\mathrm\{span\}\\\{w\_\{1\},\\dots,w\_\{m\},v\\\}, enlarged todimW=p\\dim W=p\. In each partδ\\deltadenotes a scale fixed there; we may assumeδ≤15\\delta\\leq\\tfrac\{1\}\{5\}, since otherwise the stated bound is at most an absolute constant, which \(with the advertisedcctaken small enough that the target is≤1\\leq 1in this range\) follows from one opposite\-label pair \(L≥1L\\geq 1\)\.

*Localization\.*LetOp:=λmax​\(∑ixi​xi⊤\)\\mathrm\{Op\}:=\\lambda\_\{\\max\}\(\\sum\_\{i\}x\_\{i\}x\_\{i\}^\{\\top\}\); as beforeOp≤C1​\(1\+n/d\)\\mathrm\{Op\}\\leq C\_\{1\}\(1\+n/d\)with probability1−2​e−d1\-2e^\{\-d\}\. For every rank\-pporthogonal projectorQQ,∑i‖Q​xi‖2=tr​\(Q​∑ixi​xi⊤​Q\)≤p​Op\\sum\_\{i\}\\\|Qx\_\{i\}\\\|^\{2\}=\\mathrm\{tr\}\(Q\\sum\_\{i\}x\_\{i\}x\_\{i\}^\{\\top\}Q\)\\leq p\\,\\mathrm\{Op\}; so, deterministically on the eventOp≤C1​\(1\+n/d\)\\mathrm\{Op\}\\leq C\_\{1\}\(1\+n/d\), the count\#​\{i:‖Q​xi‖\>τ\}≤p​Op/τ2≤n/8\\\#\\\{i:\\\|Qx\_\{i\}\\\|\>\\tau\\\}\\leq p\\,\\mathrm\{Op\}/\\tau^\{2\}\\leq n/8withτ:=\(8​C1​p​\(1n\+1d\)\)1/2\\tau:=\(8C\_\{1\}p\(\\tfrac\{1\}\{n\}\+\\tfrac\{1\}\{d\}\)\)^\{1/2\}, uniformly over all rank\-ppprojectorsQQat once\. The hypotheses giver0:=2​p/d≤τ≤18r\_\{0\}:=2\\sqrt\{p/d\}\\leq\\tau\\leq\\tfrac\{1\}\{8\}\.

*Net and grid\.*Fix a column\-wiseδ/\(8​p\)\\delta/\(8\\sqrt\{p\}\)\-net𝒫\\mathcal\{P\}of thepp\-frames as in Theorem[8\.1](https://arxiv.org/html/2607.07778#S8.Thmtheorem1), so every admissiblePPhasP^∈𝒫\\hat\{P\}\\in\\mathcal\{P\}with‖P−P^‖op≤δ/4\\\|P\-\\hat\{P\}\\\|\_\{\\mathrm\{op\}\}\\leq\\delta/4; in both partsδ≥\(n2​d\)−1\\delta\\geq\(n^\{2\}\\sqrt\{d\}\)^\{\-1\}, so\|𝒫\|≤eC​p​d​Λ\|\\mathcal\{P\}\|\\leq e^\{Cpd\\Lambda\}\. ForP^∈𝒫\\hat\{P\}\\in\\mathcal\{P\}fix orthonormal coordinates on its range, letziz\_\{i\}be the coordinates ofP^​xi\\hat\{P\}x\_\{i\}, and partitionℝp\\mathbb\{R\}^\{p\}into half\-open cubical cells of sideδ/p\\delta/\\sqrt\{p\};ncn\_\{c\}is the number ofziz\_\{i\}in cellccandqc:=ℙ​\(z∈c\)q\_\{c\}:=\\mathbb\{P\}\(z\\in c\)for an independent copy\. From each cell withnc≥2n\_\{c\}\\geq 2take disjoint same\-cell pairs greedily; the familyΠ​\(x,P^\)\\Pi\(x,\\hat\{P\}\)is determined by\(x,P^\)\(x,\\hat\{P\}\)\.

*Part \([5](https://arxiv.org/html/2607.07778#S8.E5)\)\.*Setδ:=4​p​τ​\(8/n\)1/p\\delta:=4\\sqrt\{p\}\\,\\tau\(8/n\)^\{1/p\}and cell sides:=δ/p=4​τ​\(8/n\)1/ps:=\\delta/\\sqrt\{p\}=4\\tau\(8/n\)^\{1/p\}, so thatτ/s=14​\(n/8\)1/p≥1\\tau/s=\\tfrac\{1\}\{4\}\(n/8\)^\{1/p\}\\geq 1\(herep≤c​log⁡np\\leq c\\log nwithccsmall enough that\(n/8\)1/p≥4\(n/8\)^\{1/p\}\\geq 4\)\. A radius\-τ\\tauball meets at most\(2​τ/s\+2\)p≤\(4​τ/s\)p=n/8\(2\\tau/s\+2\)^\{p\}\\leq\(4\\tau/s\)^\{p\}=n/8of the side\-sscells, the first inequality usingτ/s≥1\\tau/s\\geq 1\. At least78​n\\tfrac\{7\}\{8\}npoints areτ\\tau\-low, so among them at least78​n−n8\\tfrac\{7\}\{8\}n\-\\tfrac\{n\}\{8\}share cells, giving\|Π\|≥12​\(78​n−n8\)=38​n\|\\Pi\|\\geq\\tfrac\{1\}\{2\}\(\\tfrac\{7\}\{8\}n\-\\tfrac\{n\}\{8\}\)=\\tfrac\{3\}\{8\}n\.

*Part \([6](https://arxiv.org/html/2607.07778#S8.E6)\)\.*SetN∗:=⌈C0​p​d​Λ⌉N^\{\*\}:=\\lceil C\_\{0\}pd\\Lambda\\rceilandδ:=14​p​r0​\(N∗/\(c2​n2\)\)1/p\\delta:=14\\sqrt\{p\}\\,r\_\{0\}\\,\(N^\{\*\}/\(c\_\{2\}n^\{2\}\)\)^\{1/p\}withc2:=\(128​e2\)−1c\_\{2\}:=\(128e^\{2\}\)^\{\-1\}; the hypotheses giveδ≤p​r0\\delta\\leq\\sqrt\{p\}\\,r\_\{0\}\. Since𝔼​‖z‖2=p/d\\mathbb\{E\}\\\|z\\\|^\{2\}=p/dandr02=4​p/dr\_\{0\}^\{2\}=4p/d, Markov gives core mass∑c∩Br0≠∅qc≥34\\sum\_\{c\\cap B\_\{r\_\{0\}\}\\neq\\emptyset\}q\_\{c\}\\geq\\tfrac\{3\}\{4\}\. Call a cell*light*ifqc≤1/nq\_\{c\}\\leq 1/n,*mid*if1/n<qc≤32/n1/n<q\_\{c\}\\leq 32/n,*heavy*ifqc\>32/nq\_\{c\}\>32/n; one class carries core mass≥14\\geq\\tfrac\{1\}\{4\}\.*Light:*core cells number at most\(7​p​r0/δ\)p\(7\\sqrt\{p\}\\,r\_\{0\}/\\delta\)^\{p\}; for light cellsℙ​\(nc≥2\)≥\(n2\)​qc2​\(1−qc\)n≥n2​qc2/\(8​e\)\\mathbb\{P\}\(n\_\{c\}\\geq 2\)\\geq\\binom\{n\}\{2\}q\_\{c\}^\{2\}\(1\-q\_\{c\}\)^\{n\}\\geq n^\{2\}q\_\{c\}^\{2\}/\(8e\), so by Cauchy–Schwarz the expected numberμ\\muof doubly occupied light core cells satisfiesμ≥n2128​e​\(δ7​p​r0\)p=n2128​e⋅2p​N∗c2​n2≥2p\+1​N∗\\mu\\geq\\frac\{n^\{2\}\}\{128e\}\\bigl\(\\frac\{\\delta\}\{7\\sqrt\{p\}\\,r\_\{0\}\}\\bigr\)^\{p\}=\\frac\{n^\{2\}\}\{128e\}\\cdot\\frac\{2^\{p\}N^\{\*\}\}\{c\_\{2\}n^\{2\}\}\\geq 2^\{p\+1\}N^\{\*\}\. Multinomial occupancy counts are negatively associated\[[5](https://arxiv.org/html/2607.07778#bib.bib5)\], monotone functions of disjoint coordinates preserve negative association, and the Chernoff–Hoeffding lower tail transfers\[[4](https://arxiv.org/html/2607.07778#bib.bib4)\]; henceℙ​\(\#​\{doubly occupied\}≤N∗\)≤e−μ/8\\mathbb\{P\}\(\\\#\\\{\\text\{doubly occupied\}\\\}\\leq N^\{\*\}\)\\leq e^\{\-\\mu/8\}\.*Mid:*at least\(14\)/\(32/n\)=n/128\(\\tfrac\{1\}\{4\}\)/\(32/n\)=n/128mid core cells; for each,n​qc≥1nq\_\{c\}\\geq 1givesℙ​\(nc≥2\)≥1−\(1−1n\)n−\(1−1n\)n−1≥15\\mathbb\{P\}\(n\_\{c\}\\geq 2\)\\geq 1\-\(1\-\\tfrac\{1\}\{n\}\)^\{n\}\-\(1\-\\tfrac\{1\}\{n\}\)^\{n\-1\}\\geq\\tfrac\{1\}\{5\}forn≥100n\\geq 100, soμ≥n/640≥2​N∗\\mu\\geq n/640\\geq 2N^\{\*\}and the same negative\-association tail applies\.*Heavy:*heavy cells number at mostn/32n/32; the number of points in heavy core cells dominatesBin​\(n,14\)\\mathrm\{Bin\}\(n,\\tfrac\{1\}\{4\}\), hence is≥n/8\\geq n/8with probability1−e−n/321\-e^\{\-n/32\}, and≥n/8\\geq n/8points in≤n/32\\leq n/32cells yield at least12​\(n8−n32\)=3​n64≥N∗\\tfrac\{1\}\{2\}\(\\tfrac\{n\}\{8\}\-\\tfrac\{n\}\{32\}\)=\\tfrac\{3n\}\{64\}\\geq N^\{\*\}disjoint pairs\. In every case\|Π\|≥N∗\|\\Pi\|\\geq N^\{\*\}with probability1−2​e−c​N∗1\-2e^\{\-cN^\{\*\}\}, all pairs inBr0\+δB\_\{r\_\{0\}\+\\delta\}\.

*Labels, union, conclusion\.*Givenxx, the pairs are disjoint andyyis independent, soℙ​\(no opposite\-label pair∣x\)≤2−\|Π\|\\mathbb\{P\}\(\\text\{no opposite\-label pair\}\\mid x\)\\leq 2^\{\-\|\\Pi\|\}; the union over𝒫\\mathcal\{P\}costseC​p​d​Λe^\{Cpd\\Lambda\}, absorbed byN∗N^\{\*\}\(takeC0C\_\{0\}large\) in part \([6](https://arxiv.org/html/2607.07778#S8.E6)\) and by38​n≥N∗\\tfrac\{3\}\{8\}n\\geq N^\{\*\}in part \([5](https://arxiv.org/html/2607.07778#S8.E5)\)\. On the good event, for the truePPsome opposite\-label pair has‖P​\(xi−xj\)‖≤32​δ\\\|P\(x\_\{i\}\-x\_\{j\}\)\\\|\\leq\\tfrac\{3\}\{2\}\\deltawith both points\(τ\+54​δ\)\(\\tau\+\\tfrac\{5\}\{4\}\\delta\)\-low, and \([10](https://arxiv.org/html/2607.07778#A1.E10)\) gives2≤2​L⋅32​δ2\\leq 2L\\cdot\\tfrac\{3\}\{2\}\\delta, i\.e\.L≥23​δL\\geq\\tfrac\{2\}\{3\\delta\}\. Substituting the two choices ofδ\\delta\(andτ≤4​C1​p/d\\tau\\leq 4\\sqrt\{C\_\{1\}p/d\}, valid asn≥dn\\geq d\) gives \([5](https://arxiv.org/html/2607.07778#S8.E5)\) and \([6](https://arxiv.org/html/2607.07778#S8.E6)\)\. ∎

### A\.8\.Proofs for Section[10](https://arxiv.org/html/2607.07778#S10)

###### Proof of Proposition[10\.1](https://arxiv.org/html/2607.07778#S10.Thmtheorem1)\.

[\[←\\leftarrowstatement\]](https://arxiv.org/html/2607.07778#S10.Thmtheorem1)The profileρ\\rhovanishes on\(−∞,1−s\]\(\-\\infty,1\-s\], rises with slope1/s=41/s=4on\[1−s,1\]\[1\-s,1\], and equals11atu=1u=1; on𝕊d−1\\mathbb\{S\}^\{d\-1\}no argument exceeds11\.

*Interpolation\.*Sinceρ​\(1\)=1\\rho\(1\)=1andρ​\(⟨xj,xi⟩\)=0\\rho\(\\langle x\_\{j\},x\_\{i\}\\rangle\)=0forj≠ij\\neq iby⟨xj,xi⟩≤1/8<3/4\\langle x\_\{j\},x\_\{i\}\\rangle\\leq 1/8<3/4, we havef​\(xi\)=yif\(x\_\{i\}\)=y\_\{i\}\.

*Disjoint caps\.*LetCi=\{x∈𝕊d−1:⟨xi,x⟩\>3/4\}C\_\{i\}=\\\{x\\in\\mathbb\{S\}^\{d\-1\}:\\langle x\_\{i\},x\\rangle\>3/4\\\}\. Ifx∈Ci∩Cjx\\in C\_\{i\}\\cap C\_\{j\}withi≠ji\\neq j, then⟨xi\+xj,x⟩\>3/2\\langle x\_\{i\}\+x\_\{j\},x\\rangle\>3/2, so‖xi\+xj‖\>3/2\\\|x\_\{i\}\+x\_\{j\}\\\|\>3/2\. But

‖xi\+xj‖2=2\+2​⟨xi,xj⟩≤2\+14=94,\\\|x\_\{i\}\+x\_\{j\}\\\|^\{2\}=2\+2\\langle x\_\{i\},x\_\{j\}\\rangle\\leq 2\+\\frac\{1\}\{4\}=\\frac\{9\}\{4\},contradiction\. Thus the caps are pairwise disjoint\.

*Lipschitz bound\.*Along any unit\-speed geodesicγ\\gammaon𝕊d−1\\mathbb\{S\}^\{d\-1\}, the derivative ofyi​ρ​\(⟨xi,γ​\(t\)⟩\)y\_\{i\}\\rho\(\\langle x\_\{i\},\\gamma\(t\)\\rangle\)exists for a\.e\.ttand, on the rising band⟨xi,γ​\(t\)⟩∈\[3/4,1\]\\langle x\_\{i\},\\gamma\(t\)\\rangle\\in\[3/4,1\], has absolute value at most

4​\|yi\|​1−⟨xi,γ​\(t\)⟩2≤4​1−\(3/4\)2=7\.4\|y\_\{i\}\|\\sqrt\{1\-\\langle x\_\{i\},\\gamma\(t\)\\rangle^\{2\}\}\\leq 4\\sqrt\{1\-\(3/4\)^\{2\}\}=\\sqrt\{7\}\.Off the rising band it is0a\.e\. Because the caps are disjoint, at every point of the sphere at most one summand is nonconstant\. Hence\|dd​t​f​\(γ​\(t\)\)\|≤7\|\\frac\{d\}\{dt\}f\(\\gamma\(t\)\)\|\\leq\\sqrt\{7\}for a\.e\.tt, andffis7\\sqrt\{7\}\-Lipschitz for geodesic distance\. For chord distance, if the geodesic distance betweenx,x′x,x^\{\\prime\}isθ∈\[0,π\]\\theta\\in\[0,\\pi\], then‖x−x′‖=2​sin⁡\(θ/2\)\\\|x\-x^\{\\prime\}\\\|=2\\sin\(\\theta/2\)andθ/\(2​sin⁡\(θ/2\)\)≤π/2\\theta/\(2\\sin\(\\theta/2\)\)\\leq\\pi/2\. Therefore\|f​\(x\)−f​\(x′\)\|≤\(π/2\)​7​‖x−x′‖\|f\(x\)\-f\(x^\{\\prime\}\)\|\\leq\(\\pi/2\)\\sqrt\{7\}\\\|x\-x^\{\\prime\}\\\|\.

*Separation probability\.*For two independent uniform points, conditioning onxjx\_\{j\}and applying spherical concentration to the11\-Lipschitz functionx↦⟨x,xj⟩x\\mapsto\\langle x,x\_\{j\}\\ranglegives an absolute tail boundℙ​\(\|⟨xi,xj⟩\|\>t\)≤2​e−c​d​t2\\mathbb\{P\}\(\|\\langle x\_\{i\},x\_\{j\}\\rangle\|\>t\)\\leq 2e^\{\-cdt^\{2\}\}\(see, e\.g\.,\[[7](https://arxiv.org/html/2607.07778#bib.bib7), Ch\. 5\]\)\. Att=1/8t=1/8this is at most2​e−c​d/642e^\{\-cd/64\}\. A union bound over at mostn2/2n^\{2\}/2pairs gives failure probability at mostn2​e−c​d/64n^\{2\}e^\{\-cd/64\}, which is at most1/n1/nas soon asd≥Csep​log⁡nd\\geq C\_\{\\rm sep\}\\log nfor a sufficiently large absolute constantCsepC\_\{\\rm sep\}\. ∎

## Acknowledgments and funding

The research presented in this paper was supported by the European Research Council \(ERC\) under the European Union’s Horizon 2022 research and innovation programme \(grant agreement No\. 101041711\), by the Simons Foundation as part of the Collaboration on the Mathematical and Scientific Foundations of Deep Learning, by Heights Labs, by the Israel Science Foundation \(grant number 2258/19\), by the Israel Science Foundation \(ISF Grant 4101/25\), and by the U\.S\. National Science Foundation \(NSF Grant OISE\-2401227\)\.

## Declaration of competing interest

The author declares no competing interests\.

## Declaration of generative AI and AI\-assisted technologies in the manuscript preparation process

During the preparation of this work the author used generative AI tools to accelerate drafting and revision\. All mathematical claims, proofs, numerical interpretations, and bibliographic information were subsequently reviewed and edited by the author, who takes full responsibility for the content of the manuscript\.

## Data and code availability

## References

- \[1\]S\. Bubeck, Y\. Li, and D\. M\. Nagaraj\.A law of robustness for two\-layers neural networks\.In*Proceedings of the 34th Conference on Learning Theory*, Proceedings of Machine Learning Research, vol\. 134, pp\. 804–820, PMLR, 2021\. arXiv:2009\.14444\.
- \[2\]S\. Bubeck and M\. Sellke\.A universal law of robustness via isoperimetry\.*Journal of the ACM*70 \(2023\), no\. 2, Article 10, 18 pp\. Conference version in*Advances in Neural Information Processing Systems 34*, 2021\. DOI: 10\.1145/3578580\. arXiv:2105\.12806\.
- \[3\]Y\. Wu, H\. Huang, and H\. Zhang\.A law of robustness beyond isoperimetry\.In*Proceedings of the 40th International Conference on Machine Learning*, Proceedings of Machine Learning Research, vol\. 202, pp\. 37439–37455, PMLR, 2023\. arXiv:2202\.11592\.
- \[4\]D\. Dubhashi and D\. Ranjan\.Balls and bins: a study in negative dependence\.*Random Structures & Algorithms*13 \(1998\), no\. 2, 99–124\.
- \[5\]K\. Joag\-Dev and F\. Proschan\.Negative association of random variables, with applications\.*Annals of Statistics*11 \(1983\), no\. 1, 286–295\. DOI: 10\.1214/aos/1176346079\.
- \[6\]A\. Pinkus\.*Ridge Functions*\.Cambridge Tracts in Mathematics, vol\. 205, Cambridge University Press, Cambridge, 2015\. DOI: 10\.1017/CBO9781316408124\.
- \[7\]R\. Vershynin\.*High\-Dimensional Probability: An Introduction with Applications in Data Science*\.Cambridge Series in Statistical and Probabilistic Mathematics, vol\. 47, Cambridge University Press, Cambridge, 2018\. DOI: 10\.1017/9781108231596\.
- \[8\]L\. Breiman\.Hinging hyperplanes for regression, classification, and function approximation\.*IEEE Transactions on Information Theory*39 \(1993\), no\. 3, 999–1013\. DOI: 10\.1109/18\.256506\.
- \[9\]R\. Arora, A\. Basu, P\. Mianjy, and A\. Mukherjee\.Understanding deep neural networks with rectified linear units\.In*International Conference on Learning Representations*, 2018\. arXiv:1611\.01491\.
- \[10\]K\. Atkinson and W\. Han\.*Spherical Harmonics and Approximations on the Unit Sphere: An Introduction*\.Lecture Notes in Mathematics, vol\. 2044, Springer, Berlin, 2012\. DOI: 10\.1007/978\-3\-642\-25983\-8\.
- \[11\]M\. Ledoux and M\. Talagrand\.*Probability in Banach Spaces: Isoperimetry and Processes*\.Ergebnisse der Mathematik und ihrer Grenzgebiete, vol\. 23, Springer, Berlin, 1991\. DOI: 10\.1007/978\-3\-642\-20212\-4\.
- \[12\]C\. McDiarmid\.On the method of bounded differences\.In*Surveys in Combinatorics 1989*, London Math\. Soc\. Lecture Note Ser\., vol\. 141, Cambridge University Press, Cambridge, 1989, pp\. 148–188\. DOI: 10\.1017/CBO9781107359949\.008\.
- \[13\]Y\. Shmalo\.Toward the log\-free law of robustness: a reduction to one multiplier estimate\.Supplementary note, 2026\. Available in the code repository,[https://github\.com/yspennstate/law\-of\-robustness\-two\-layer](https://github.com/yspennstate/law-of-robustness-two-layer)\.

Similar Articles

Bug or Feature^2: Weight Drift, Activation Sparsity, and Spikes

Hugging Face Daily Papers

This paper formally proves that training neural networks with asymmetric activation functions like ReLU, GELU, or SiLU causes weights to drift negative, leading to up to 90% activation sparsity. It also shows that squared activations like ReLU² improve performance but cause activation spikes, which can be fixed by clipping, with GELU² achieving the best validation loss.

Shallower ReLU Network Representations via Exact Linear Algebra

arXiv cs.LG

This paper improves theoretical bounds on the depth of ReLU networks needed to represent the maximum function, showing exact two-hidden-layer representations for up to 10 inputs and improved depth for larger n via exact linear algebra techniques.