When a Flatness Proxy Is Not a Function: Robustness Certificates and Training Interventions
Summary
This paper shows that being a valid curvature upper bound does not automatically justify using a last-layer relative-flatness proxy as an adversarial-robustness certificate or as a differentiable training regularizer. The authors derive a gauge-invariant repair, demonstrate the proxy's unboundedness under symmetry-preserving softmax shifts, and show via CIFAR-10 experiments that quotient-regularized predictors stay aligned while raw-regularized ones can have their generalization suppressed.
View Cached Full Text
Cached at: 10/02/26, 09:50 AM
# When a Flatness Proxy Is Not a Function:Robustness Certificates and Training Interventions
Source: [https://arxiv.org/html/2609.38540](https://arxiv.org/html/2609.38540)
Vicente OpazoCentro Nacional de InteligenciaArtificial \(CENIA\)vicente\.opazo@cenia\.clJose Calatayud\-MateuBarcelona SupercomputingCenter \(BSC\)jose\.calatayud@bsc\.esCristobal RojasPontificia UniversidadCatólica de Chileluis\.rojas@uc\.clCristian Buc CalderonCentro Nacional de InteligenciaArtificial \(CENIA\)cristian\.buc@cenia\.cl
###### Abstract
A valid curvature upper bound need not justify either a robustness certificate or an intervention on an intrinsic predictor property\. We demonstrate this distinction for a last\-layer relative\-flatness proxy used in both settings\. First, empirical\-risk stationarity does not eliminate pointwise first\-order loss terms: at a finite global empirical\-risk minimum, the retained certificate expression underestimates a loss increase by over210×210\\times\. We derive a globally valid, gauge\-invariant feature\-space repair\. Second, common\-row softmax shifts preserve predictions and the exact contraction while making the proxy unbounded\. Even standard reference\-class choices double it on average relative to the centered representation\. For a single fixed\-feature example with at least three classes, scalar retuning generically cannot align the induced probability updates\. Row centering gives the orbit\-minimized bound and restores value and full\-model gradient invariance under this symmetry\. Across 45 paired one\-step tests on algorithmic and image models, amplified shifts separate raw\-regularized predictors while quotient\-regularized predictors remain aligned\. Long\-horizon CIFAR\-10 experiments show substantial, reversible suppression of generalization, while evidence for selective delay after memorization is less consistent\. Together, these results show that validity as a curvature upper bound does not by itself justify either inversion into a robustness certificate or differentiation into an intrinsic training intervention\.
## 1Introduction
Flatness connects local loss geometry to stability: near a stationary solution, curvature controls the leading response to small parameter perturbations\. This motivates explanations of generalization, sharpness\-aware training, and robustness guarantees\([Tsuzuku et al\., 2020](https://arxiv.org/html/2609.38540#bib.bib9);[Petzka et al\., 2021](https://arxiv.org/html/2609.38540#bib.bib6);[Foret et al\., 2021](https://arxiv.org/html/2609.38540#bib.bib2)\)\. Using flatness as a regularizer or certificate, however, requires more than a valid curvature upper bound\.
*Relative*flatness was proposed to address one important form of coordinate dependence by contracting curvature with layer scale\([Petzka et al\., 2021](https://arxiv.org/html/2609.38540#bib.bib6)\)\. In practice, recent work has operationalized it through a tractable upper bound on the last\-layer contraction\. This same proxy underlies the two uses introduced above: it was inverted into an adversarial\-robustness certificate\([Walter et al\., 2026](https://arxiv.org/html/2609.38540#bib.bib10)\)and differentiated as a regularizer\([Han et al\., 2025](https://arxiv.org/html/2609.38540#bib.bib3)\)to test whether flatness causally affects generalization during grokking, a setting that temporally separates memorization from generalization\([Power et al\., 2022](https://arxiv.org/html/2609.38540#bib.bib8)\)\.
The difficulty is that flatness measures can change under reparameterizations that leave the predictor fixed\([Dinh et al\., 2017](https://arxiv.org/html/2609.38540#bib.bib1)\)\. For operational uses, identification and soundness are distinct: a parameterization\-dependent certificate can still be valid, whereas a representation\-independent training intervention requires the same functional update across equivalent representatives\. We therefore distinguish gauge dependence from invalid certification, and parameter\-specific training effects from interventions on intrinsic predictor properties\.
Softmax classifiers have a simple function\-preserving symmetry: adding the same vector to every classifier row introduces a common logit offset and leaves the represented function unchanged\. The exact relative\-flatness contraction respects this symmetry, but the operational proxy does not\. As a result, a certificate derived from the proxy depends on parameterization, while regularizing it can send function\-equivalent models in different functional directions\. In this precise sense, the proxy is not a function of the predictor\. Independently, the certificate derivation drops a pointwise first\-order term that dataset stationarity does not eliminate\.
The key point is that approximation validity does not automatically justify downstream use\. The same valid upper bound is inverted into a robustness radius and differentiated into a training objective, exposing distinct failure modes\.
Our main contributions are:
- •Certification\.We show that empirical\-risk stationarity does not remove pointwise first\-order loss terms\. A finite global\-minimum counterexample invalidates the resulting curvature\-only certificate by over210×210\\times, and we derive a globally valid, gauge\-invariant feature\-space bound\.
- •Intervention\.We show that the raw proxy is unbounded within a softmax function\-equivalence class even though predictions and the exact contraction are unchanged\. Row centering yields its orbit\-minimized quotient, and scalar retuning generically cannot align the induced functional updates for a fixed\-feature multiclass example withK≥3K\\geq 3\.
- •Empirical evidence\.Across 45 paired one\-step tests on algorithmic and image models, function\-equivalent parameterizations produce different updates under the raw proxy while quotient\-based interventions remain aligned\. Long\-horizon CIFAR\-10 experiments show substantial suppression and recovery, with less consistent evidence for selective delay after memorization\.
## 2Preliminaries
### 2\.1Two operational uses of the proxy
We study two downstream uses of the same relative\-flatness proxy\.[Walter et al\. \(2026\)](https://arxiv.org/html/2609.38540#bib.bib10)invert it into a robustness radius, where validity requires a pointwise loss bound\.[Han et al\. \(2025\)](https://arxiv.org/html/2609.38540#bib.bib3)differentiate it as a regularizer, where an intrinsic interpretation requires equivalent predictors to receive the same functional update\. These uses expose distinct requirements that are not implied by the validity of the underlying curvature upper bound\.
### 2\.2Relative flatness and functional identification
#### The operational relative\-flatness proxy\.
Consider aKK\-class model with featureh∈ℝdh\\in\\mathbb\{R\}^\{d\}, classifierW∈ℝK×dW\\in\\mathbb\{R\}^\{K\\times d\}, logitsz=Whz=Wh, and probabilitiesp=softmax\(z\)p=\\operatorname\{softmax\}\(z\)\. The matrixP\(p\)=Diag\(p\)−pp⊤P\(p\)=\\operatorname\{Diag\}\(p\)\-pp^\{\\top\}is both the covariance of the one\-hot encoding ofY∼Cat\(p\)Y\\sim\\operatorname\{Cat\}\(p\)and the softmax Jacobian∂p/∂z\\partial p/\\partial z\. For one cross\-entropy example, the classifier Hessian isP\(p\)⊗hh⊤P\(p\)\\otimes hh^\{\\top\}\. The last\-layer relative\-flatness contraction is
κexact\(W,h\)=∥h∥22Tr\(W⊤P\(p\)W\)\.\\kappa\_\{\\mathrm\{exact\}\}\(W,h\)=\\lVert h\\rVert\_\{2\}^\{2\}\\operatorname\{Tr\}\\\!\\left\(W^\{\\top\}P\(p\)W\\right\)\.\(1\)BecauseP\(p\)P\(p\)is positive semidefinite, a trace inequality gives
κexact\(W,h\)\\displaystyle\\kappa\_\{\\mathrm\{exact\}\}\(W,h\)≤∥W∥F2∥h∥22TrP\(p\)\\displaystyle\\leq\\lVert W\\rVert\_\{F\}^\{2\}\\lVert h\\rVert\_\{2\}^\{2\}\\operatorname\{Tr\}P\(p\)=∥W∥F2∥h∥22\(1−∥p∥22\)⏟Braw\(W,h\)\.\\displaystyle=\\underbrace\{\\lVert W\\rVert\_\{F\}^\{2\}\\lVert h\\rVert\_\{2\}^\{2\}\\left\(1\-\\lVert p\\rVert\_\{2\}^\{2\}\\right\)\}\_\{B\_\{\\mathrm\{raw\}\}\(W,h\)\}\.\(2\)Equation \([2](https://arxiv.org/html/2609.38540#S2.E2)\) is the operational proxy studied below\.
#### An exact softmax symmetry\.
Before interpretingBrawB\_\{\\mathrm\{raw\}\}as a property of the predictor, consider adding the same vectorv∈ℝdv\\in\\mathbb\{R\}^\{d\}to every row ofWW:
Gv\(W\)=W\+𝟏v⊤\.G\_\{v\}\(W\)=W\+\\mathbf\{1\}v^\{\\top\}\.\(3\)SinceGv\(W\)h=Wh\+\(v⊤h\)𝟏G\_\{v\}\(W\)h=Wh\+\(v^\{\\top\}h\)\\mathbf\{1\}, softmax invariance to common logit shifts implies thatWWandGv\(W\)G\_\{v\}\(W\)define the same predictor\. Moreover,P\(p\)𝟏=0P\(p\)\\mathbf\{1\}=0makesκexact\\kappa\_\{\\mathrm\{exact\}\}invariant, whereasBrawB\_\{\\mathrm\{raw\}\}generally changes through∥W∥F2\\lVert W\\rVert\_\{F\}^\{2\}\.
The redundant common\-row component can be removed explicitly\. Let
w¯=K−1W⊤𝟏,Π=IK−K−1𝟏𝟏⊤,Wc=ΠW=W−𝟏w¯⊤\.\\bar\{w\}=K^\{\-1\}W^\{\\top\}\\mathbf\{1\},\\qquad\\Pi=I\_\{K\}\-K^\{\-1\}\\mathbf\{1\}\\mathbf\{1\}^\{\\top\},\\qquad W\_\{c\}=\\Pi W=W\-\\mathbf\{1\}\\bar\{w\}^\{\\top\}\.\(4\)Each row ofWcW\_\{c\}is its original row minus the mean roww¯⊤\\bar\{w\}^\{\\top\}\. The centered classifier is unchanged byGvG\_\{v\}\. Evaluating the same upper\-bound expression after this centering defines the quotient proxy
Bquot\(W,h\)=∥Wc∥F2∥h∥22\(1−∥p∥22\)\.B\_\{\\mathrm\{quot\}\}\(W,h\)=\\lVert W\_\{c\}\\rVert\_\{F\}^\{2\}\\lVert h\\rVert\_\{2\}^\{2\}\\left\(1\-\\lVert p\\rVert\_\{2\}^\{2\}\\right\)\.\(5\)
#### Three distinct requirements\.
A candidate scoreS\(W,h\)S\(W,h\)can be assessed along three distinct axes:
1. 1\.*Approximation validity:*ifSSis used as an upper\-bound proxy, doesκexact\(W,h\)≤S\(W,h\)\\kappa\_\{\\mathrm\{exact\}\}\(W,h\)\\leq S\(W,h\)hold?
2. 2\.*Value identification:*do function\-equivalent parameters receive the same score,S\(Gv\(W\),h\)=S\(W,h\)S\(G\_\{v\}\(W\),h\)=S\(W,h\)?
3. 3\.*Interventional identification:*for model outputF\(θ\)F\(\\theta\), is the Euclidean gradient\-flow velocity𝒱S\(θ\)=JF\(θ\)∇θS\(θ\)\\mathcal\{V\}\_\{S\}\(\\theta\)=J\_\{F\}\(\\theta\)\\nabla\_\{\\theta\}S\(\\theta\)constant along each orbit generated byGvG\_\{v\}? This tests changes in predictions rather than score values\.
Throughout, identification refers only to the common\-row symmetry\. For this affine translation, differentiable value invariance implies invariance of the Euclidean gradient field\. Section[3](https://arxiv.org/html/2609.38540#S3)applies this checklist toBrawB\_\{\\mathrm\{raw\}\},BquotB\_\{\\mathrm\{quot\}\}, andκexact\\kappa\_\{\\mathrm\{exact\}\}\.
### 2\.3Relation to prior work
Parameter\-space sharpness can change under function\-preserving reparameterizations, and several flatness measures address different forms of coordinate dependence\([Dinh et al\., 2017](https://arxiv.org/html/2609.38540#bib.bib1);[Tsuzuku et al\., 2020](https://arxiv.org/html/2609.38540#bib.bib9);[Petzka et al\., 2021](https://arxiv.org/html/2609.38540#bib.bib6);[Jang et al\., 2022](https://arxiv.org/html/2609.38540#bib.bib5)\)\. Quotienting symmetries is also established\([Pittorino et al\., 2022](https://arxiv.org/html/2609.38540#bib.bib7)\), while common\-shift non\-identifiability and its standard identification conventions are classical in multinomial\-logit models\([Hastie et al\., 2009](https://arxiv.org/html/2609.38540#bib.bib4);[Zahid and Tutz, 2009](https://arxiv.org/html/2609.38540#bib.bib11)\)\. Our contribution connects this redundancy to both certification and training, proves that scalar recalibration fails almost everywhere in the fixed\-feature multiclass case, and derives the orbit\-minimized repair\.
## 3One symmetry, two operational failures
The common\-row symmetry exposes a shared identification defect in both downstream uses\. The following result separates this failure from the certificate’s distinct pointwise soundness problem\.
###### Proposition 1\(Gauge failure and quotient repair\)\.
For everyW,h,vW,h,v, the transformationGvG\_\{v\}preserves the class\-probability vectorp=softmax\(Wh\)p=\\operatorname\{softmax\}\(Wh\)and the exact contractionκexact\\kappa\_\{\\mathrm\{exact\}\}\. Ifh≠0h\\neq 0,v≠0v\\neq 0, andppis non\-degenerate, thenBraw\(Gtv\(W\),h\)→∞B\_\{\\mathrm\{raw\}\}\(G\_\{tv\}\(W\),h\)\\to\\inftyas\|t\|→∞\|t\|\\to\\infty\. The centered classifierWcW\_\{c\}is the unique minimum\-Frobenius representative of the orbit\. Consequently, the quotient proxy is the orbit\-minimized raw bound:
Bquot\(W,h\)=infvBraw\(Gv\(W\),h\),κexact\(W,h\)≤Bquot\(W,h\)≤Braw\(W,h\)\.B\_\{\\mathrm\{quot\}\}\(W,h\)=\\inf\_\{v\}B\_\{\\mathrm\{raw\}\}\(G\_\{v\}\(W\),h\),\\qquad\\kappa\_\{\\mathrm\{exact\}\}\(W,h\)\\leq B\_\{\\mathrm\{quot\}\}\(W,h\)\\leq B\_\{\\mathrm\{raw\}\}\(W,h\)\.\(6\)Finally, bothBquotB\_\{\\mathrm\{quot\}\}andκexact\\kappa\_\{\\mathrm\{exact\}\}have gauge\-invariant full\-model gradients\.
###### Proof\.
The transformed logits areWh\+\(v⊤h\)𝟏Wh\+\(v^\{\\top\}h\)\\mathbf\{1\}, which softmax removes, andP\(p\)𝟏=0P\(p\)\\mathbf\{1\}=0removes every common\-row term from Equation \([1](https://arxiv.org/html/2609.38540#S2.E1)\)\. In contrast,
∥W\+t𝟏v⊤∥F2=∥W∥F2\+2t⟨𝟏⊤W,v⟩\+t2K∥v∥22\.\\lVert W\+t\\mathbf\{1\}v^\{\\top\}\\rVert\_\{F\}^\{2\}=\\lVert W\\rVert\_\{F\}^\{2\}\+2t\\langle\\mathbf\{1\}^\{\\top\}W,v\\rangle\+t^\{2\}K\\lVert v\\rVert\_\{2\}^\{2\}\.\(7\)Writingw¯=K−1W⊤𝟏\\bar\{w\}=K^\{\-1\}W^\{\\top\}\\mathbf\{1\}gives the orthogonal decompositionW=Wc\+𝟏w¯⊤W=W\_\{c\}\+\\mathbf\{1\}\\bar\{w\}^\{\\top\}and∥W∥F2=∥Wc∥F2\+K∥w¯∥22\\lVert W\\rVert\_\{F\}^\{2\}=\\lVert W\_\{c\}\\rVert\_\{F\}^\{2\}\+K\\lVert\\bar\{w\}\\rVert\_\{2\}^\{2\}\. Thus centering uniquely minimizes the raw norm\. SinceP\(p\)=ΠP\(p\)ΠP\(p\)=\\Pi P\(p\)\\Pi, centering preservesκexact\\kappa\_\{\\mathrm\{exact\}\}before the same trace bound proves Equation \([6](https://arxiv.org/html/2609.38540#S3.E6)\)\. Differentiating the invariance identities proves gradient invariance for this affine translation\. ∎
These statements hold termwise for dataset and minibatch averages\. The raw score can therefore differ across parameterizations of the same softmax function\. The quotient removes exactly that coordinate\.
This quotient is the minimal identification repair of the operational bound: it takes the orbit infimum while retaining the original upper\-bound form\. Replacing the proxy byκexact\\kappa\_\{\\mathrm\{exact\}\}is a different intervention that changes identification, tightness, objective value, and gradient direction at once\. We use both below to separate these questions rather than claim thatBquotB\_\{\\mathrm\{quot\}\}is preferable to the exact contraction\.
#### The effect is already visible under standard parameterizations\.
Identification requires invariance across the entire equivalence class, so one counterexample is enough to reject it\. How often large displacements occur during ordinary training is a separate empirical question and cannot rescue a certificate or intervention that changes under a transformation leaving the function exactly unchanged\.
A standard reference\-class convention gives a non\-amplified stress test\. Starting fromWcW\_\{c\}, choose classrras the reference and set
W\(r\)=Wc−𝟏wc,r⊤\.W^\{\(r\)\}=W\_\{c\}\-\\mathbf\{1\}w\_\{c,r\}^\{\\top\}\.\(8\)Therrth row is now zero, giving a standard reference\-category parameterization of the same softmax function\.
###### Corollary 1\(Standard identification conventions\)\.
For every nonzero centered classifier and everyh≠0h\\neq 0,
Braw\(W\(r\),h\)Bquot\(Wc,h\)=1\+K∥wc,r∥22∥Wc∥F2,1K∑r=1KBraw\(W\(r\),h\)Bquot\(Wc,h\)=2\.\\frac\{B\_\{\\mathrm\{raw\}\}\(W^\{\(r\)\},h\)\}\{B\_\{\\mathrm\{quot\}\}\(W\_\{c\},h\)\}=1\+\\frac\{K\\lVert w\_\{c,r\}\\rVert\_\{2\}^\{2\}\}\{\\lVert W\_\{c\}\\rVert\_\{F\}^\{2\}\},\\qquad\\frac\{1\}\{K\}\\sum\_\{r=1\}^\{K\}\\frac\{B\_\{\\mathrm\{raw\}\}\(W^\{\(r\)\},h\)\}\{B\_\{\\mathrm\{quot\}\}\(W\_\{c\},h\)\}=2\.\(9\)
###### Proof\.
The centered and common\-row components in Equation \([8](https://arxiv.org/html/2609.38540#S3.E8)\) are orthogonal, so∥W\(r\)∥F2=∥Wc∥F2\+K∥wc,r∥22\\lVert W^\{\(r\)\}\\rVert\_\{F\}^\{2\}=\\lVert W\_\{c\}\\rVert\_\{F\}^\{2\}\+K\\lVert w\_\{c,r\}\\rVert\_\{2\}^\{2\}\. Averaging the second term over rows contributes exactly one additional copy of∥Wc∥F2\\lVert W\_\{c\}\\rVert\_\{F\}^\{2\}\. ∎
Thus the failure is not confined to arbitrarily amplified gauge shifts: ordinary reference\-class conventions change the proxy by a factor of two on average while leaving the predictor unchanged\.
## 4From flatness bounds to robustness certificates
Gauge dependence makes the reported radius representation\-dependent, but does not by itself establish unsoundness\. The soundness failure comes from a separate pointwise first\-order omission\. Independently, Equation 5 of[Walter et al\. \(2026\)](https://arxiv.org/html/2609.38540#bib.bib10)includes the pointwise termΔ∥W∥F∥∇Wℓx\(W\)∥F\\Delta\\lVert W\\rVert\_\{F\}\\lVert\\nabla\_\{W\}\\ell\_\{x\}\(W\)\\rVert\_\{F\}\. Proposition 8 can omit it at a pointwise minimum, although softmax cross\-entropy has no such finite minimizer when∥ϕ\(x\)∥2\>0\\lVert\\phi\(x\)\\rVert\_\{2\}\>0\. Proposition 9, however, assumes only an empirical\-risk minimum and applies Proposition 8 to each example\. This is invalid because∇WRS\(W\)=0⇏∇Wℓx\(W\)=0\\nabla\_\{W\}R\_\{S\}\(W\)=0\\not\\Rightarrow\\nabla\_\{W\}\\ell\_\{x\}\(W\)=0for everyx∈Sx\\in S\.
###### Proposition 2\(First\-order obstruction\)\.
Letℓ\\ellbe differentiable athhwithgh=∇hℓ\(h\)≠0g\_\{h\}=\\nabla\_\{h\}\\ell\(h\)\\neq 0\. Fordρ=ρgh/∥gh∥2d\_\{\\rho\}=\\rho g\_\{h\}/\\lVert g\_\{h\}\\rVert\_\{2\},
ℓ\(h\+dρ\)−ℓ\(h\)=ρ∥gh∥2\+o\(ρ\)\.\\ell\(h\+d\_\{\\rho\}\)\-\\ell\(h\)=\\rho\\lVert g\_\{h\}\\rVert\_\{2\}\+o\(\\rho\)\.\(10\)Hence no nonnegativeq\(ρ\)=o\(ρ\)q\(\\rho\)=o\(\\rho\), including any fixed quadratic\-plus\-cubic expression, upper\-bounds every pointwise loss increase nearhh\.
###### Proof\.
Differentiability givesℓ\(h\+d\)=ℓ\(h\)\+gh⊤d\+o\(∥d∥2\)\\ell\(h\+d\)=\\ell\(h\)\+g\_\{h\}^\{\\top\}d\+o\(\\lVert d\\rVert\_\{2\}\)\. Substitution ofdρd\_\{\\rho\}yields the display\. Its positive linear coefficient eventually exceeds everyq\(ρ\)=o\(ρ\)q\(\\rho\)=o\(\\rho\)\. ∎
###### Proposition 3\(Finite dataset\-minimum counterexample\)\.
LetK=2K=2,d=1d=1, andϕ\(x\)=x\\phi\(x\)=x, with the rows ofWWrepresenting classes00and11\. On the training multisetS=\[\(1,1\),\(1,1\),\(1,0\)\]S=\[\(1,1\),\(1,1\),\(1,0\)\],W⋆=\(0,log2\)⊤W\_\{\\star\}=\(0,\\log 2\)^\{\\top\}is a global empirical\-risk minimizer, unique up to common\-row shifts\. Fory=1y=1atx=1x=1and the perturbation1↦0\.991\\mapsto 0\.99, the pointwise loss increase is2\.32×10−32\.32\\times 10^\{\-3\}while the published quadratic\-plus\-cubic expression is1\.10×10−51\.10\\times 10^\{\-5\}: a violation by more than210×210\\times\.
###### Proof\.
Binary softmax depends onW=\(w0,w1\)⊤W=\(w\_\{0\},w\_\{1\}\)^\{\\top\}only throughβ=w1−w0\\beta=w\_\{1\}\-w\_\{0\}, so
R\(β\)=23log\(1\+e−β\)\+13log\(1\+eβ\),R\(\\beta\)=\\tfrac\{2\}\{3\}\\log\(1\+e^\{\-\\beta\}\)\+\\tfrac\{1\}\{3\}\\log\(1\+e^\{\\beta\}\),\(11\)SinceR′\(β\)=softmax\(0,β\)1−2/3R^\{\\prime\}\(\\beta\)=\\operatorname\{softmax\}\(0,\\beta\)\_\{1\}\-2/3is strictly increasing, its unique root isβ⋆=log2\\beta\_\{\\star\}=\\log 2\. Hence all minimizers are\(c,c\+log2\)⊤\(c,c\+\\log 2\)^\{\\top\}, andW⋆W\_\{\\star\}choosesc=0c=0\. Herep1=2/3p\_\{1\}=2/3, so the two class\-1 gradients cancel the class\-0 gradient in the empirical average although none vanishes\. Take𝒳=\[0\.99,1\]\\mathcal\{X\}=\[0\.99,1\],δ=0\.01\\delta=0\.01,K=2K=2,m=d=1m=d=1, andr=0\.99r=0\.99\. These choices satisfyϕ\\phi11\-Lipschitz,\|x\|≤1\|x\|\\leq 1,\|ϕ\(x\)\|≥r\|\\phi\(x\)\|\\geq r, andℓ1\(1\)=log\(3/2\)<log2\\ell\_\{1\}\(1\)=\\log\(3/2\)<\\log 2\. For the class\-1 example,
ℓ1\(0\.99\)−ℓ1\(1\)\\displaystyle\\ell\_\{1\}\(0\.99\)\-\\ell\_\{1\}\(1\)=log\(1\+2−0\.99\)−log\(1\+2−1\)≈2\.32×10−3,\\displaystyle=\\log\(1\+2^\{\-0\.99\}\)\-\\log\(1\+2^\{\-1\}\)\\approx 2\.32\\times 10^\{\-3\},\(12\)δ22r2Braw\(W⋆,1\)\+δ324r3Km\\displaystyle\\frac\{\\delta^\{2\}\}\{2r^\{2\}\}B\_\{\\mathrm\{raw\}\}\(W\_\{\\star\},1\)\+\\frac\{\\delta^\{3\}\}\{24r^\{3\}\}Km≈1\.10×10−5\.\\displaystyle\\approx 1\.10\\times 10^\{\-5\}\.\(13\)Their ratio exceeds210210\. ∎
###### Corollary 2\(False certified radius\)\.
For loss toleranceϵ=10−4\\epsilon=10^\{\-4\}, the published expression gives the certified radiusδpub=2\.99×10−2\\delta\_\{\\rm pub\}=2\.99\\times 10^\{\-2\}, yet the perturbation in Proposition[3](https://arxiv.org/html/2609.38540#Thmproposition3)has norm0\.01<δpub0\.01<\\delta\_\{\\rm pub\}and increases the loss by2\.32×10−3\>ϵ2\.32\\times 10^\{\-3\}\>\\epsilon, so it lies inside the certified ball while violating its guarantee\. Sinceϕ\(x\)=x\\phi\(x\)=xandpy=2/3\>1/2p\_\{y\}=2/3\>1/2, this is an input\-space counterexample satisfying the stated high\-confidence condition\.
###### Proposition 4\(A sound, gauge\-invariant feature\-space repair\)\.
The two failures have a common repair: retain the first\-order term and replaceBrawB\_\{\\mathrm\{raw\}\}with gauge\-invariant row\-difference geometry\. For targetyy, letgh=∇hℓy\(h\)=W⊤\(p−ey\)g\_\{h\}=\\nabla\_\{h\}\\ell\_\{y\}\(h\)=W^\{\\top\}\(p\-e\_\{y\}\)and defineDW=maxi,j∥wi−wj∥2D\_\{W\}=\\max\_\{i,j\}\\lVert w\_\{i\}\-w\_\{j\}\\rVert\_\{2\}\. Then every feature perturbationddsatisfies
ℓy\(h\+d\)−ℓy\(h\)≤∥d∥2∥gh∥2\+DW28∥d∥22\.\\ell\_\{y\}\(h\+d\)\-\\ell\_\{y\}\(h\)\\leq\\lVert d\\rVert\_\{2\}\\lVert g\_\{h\}\\rVert\_\{2\}\+\\frac\{D\_\{W\}^\{2\}\}\{8\}\\lVert d\\rVert\_\{2\}^\{2\}\.\(14\)The uniform Hessian constantDW2/4D\_\{W\}^\{2\}/4is sharp given onlyDWD\_\{W\}\. Consequently, for loss toleranceϵ\>0\\epsilon\>0andDW\>0D\_\{W\}\>0, a certified feature radius is
ρcert=2ϵ∥gh∥2\+∥gh∥22\+DW2ϵ/2\.\\rho\_\{\\rm cert\}=\\frac\{2\\epsilon\}\{\\lVert g\_\{h\}\\rVert\_\{2\}\+\\sqrt\{\\lVert g\_\{h\}\\rVert\_\{2\}^\{2\}\+D\_\{W\}^\{2\}\\epsilon/2\}\}\.\(15\)IfDW=0D\_\{W\}=0, takeρcert=∞\\rho\_\{\\rm cert\}=\\infty\. If the feature map isLL\-Lipschitz withL\>0L\>0, the corresponding input\-space radius isρcert/L\\rho\_\{\\rm cert\}/L\.
###### Proof\.
Fort∈\[0,1\]t\\in\[0,1\], letqt=softmax\(W\(h\+td\)\)q\_\{t\}=\\operatorname\{softmax\}\(W\(h\+td\)\)\. The feature Hessian ath\+tdh\+tdisW⊤P\(qt\)WW^\{\\top\}P\(q\_\{t\}\)W, and for every unituu,
u⊤W⊤P\(qt\)Wu=VarC∼qt\(u⊤wC\)≤14\(maxiu⊤wi−minju⊤wj\)2≤DW2/4\.u^\{\\top\}W^\{\\top\}P\(q\_\{t\}\)Wu=\\operatorname\{Var\}\_\{C\\sim q\_\{t\}\}\(u^\{\\top\}w\_\{C\}\)\\leq\\tfrac\{1\}\{4\}\(\\max\_\{i\}u^\{\\top\}w\_\{i\}\-\\min\_\{j\}u^\{\\top\}w\_\{j\}\)^\{2\}\\leq D\_\{W\}^\{2\}/4\.\(16\)The first inequality bounds a scalar variance by one quarter of its squared range\. Hence the Hessian operator norm is at mostDW2/4D\_\{W\}^\{2\}/4along the segment\. Taylor’s theorem givesℓy\(h\+d\)−ℓy\(h\)≤gh⊤d\+\(DW2/8\)∥d∥22\\ell\_\{y\}\(h\+d\)\-\\ell\_\{y\}\(h\)\\leq g\_\{h\}^\{\\top\}d\+\(D\_\{W\}^\{2\}/8\)\\lVert d\\rVert\_\{2\}^\{2\}, and Cauchy–Schwarz yields Equation \([14](https://arxiv.org/html/2609.38540#S4.E14)\)\. A balanced binary softmax whose rows differ byDWuD\_\{W\}uattains the Hessian bound, so the constant is sharp\. Setting the right\-hand side of Equation \([14](https://arxiv.org/html/2609.38540#S4.E14)\) equal toϵ\\epsilonand solving for∥d∥2\\lVert d\\rVert\_\{2\}gives Equation \([15](https://arxiv.org/html/2609.38540#S4.E15)\)\. Finally, underGv\(W\)=W\+𝟏v⊤G\_\{v\}\(W\)=W\+\\mathbf\{1\}v^\{\\top\}, bothppandDWD\_\{W\}are unchanged, whileGv\(W\)⊤\(p−ey\)=gh\+v𝟏⊤\(p−ey\)=ghG\_\{v\}\(W\)^\{\\top\}\(p\-e\_\{y\}\)=g\_\{h\}\+v\\mathbf\{1\}^\{\\top\}\(p\-e\_\{y\}\)=g\_\{h\}\. ∎
#### The first\-order obstruction persists in learned representations\.
Equation \([14](https://arxiv.org/html/2609.38540#S4.E14)\) is a global certificate for arbitrary feature\-space perturbations and needs no stationarity, feature lower bound, bounded input, or third\-order remainder\. It becomes an input\-space certificate only with an independently justified Lipschitz bound for the feature map\. We therefore evaluate it directly in feature space, asking whether the omitted term is material on learned representations\. For each of five trained CIFAR\-10 ResNets and a common set of 2,048 test images, we perturb the penultimate feature alongu=gh/∥gh∥2u=g\_\{h\}/\\lVert g\_\{h\}\\rVert\_\{2\}byd=Δ∥h∥2ud=\\Delta\\lVert h\\rVert\_\{2\}u\.
Table 1:First\-order obstruction on learned representations\.Results pooled over 10,240 model–image pairs\. Actual\>\>raw is the fraction exceeding the retained quadratic\. Linear/raw measures the omitted term relative to that quadratic, and bound/actual measures the tightness of Equation \(14\), with 1 denoting equality\. Ratios are medians\.The learned\-feature results match the analysis\. ThroughΔ=0\.003\\Delta=0\.003, the realized loss increase exceeds the retained raw quadratic for every model–image pair\. AtΔ=0\.001\\Delta=0\.001, the median linear/raw ratio is10\.6010\.60, following the predicted1/Δ1/\\Deltascaling\. The repaired bound covers all 71,680 responses and is tight locally, with median and 90th\-percentile bound/actual ratios of1\.0391\.039and1\.0931\.093atΔ=10−5\\Delta=10^\{\-5\}\.
At tolerance0\.100\.10, the radius from raw curvature alone admits an above\-tolerance loss increase for24\.43%24\.43\\%of pairs\. A matched common\-row shift multiplies the raw quadratic by100×100\\times, while probabilities and loss responses change by at most3\.8×10−143\.8\\times 10^\{\-14\}\. The analytic counterexample supplies the logical refutation, while these learned\-feature results show that both mechanisms remain material on trained models\.
## 5From flatness proxies to training interventions
For an intrinsic interpretation, a proxy\-based training intervention should assign the same functional update to common\-row\-equivalent models\. The next proposition shows thatBrawB\_\{\\mathrm\{raw\}\}fails this requirement\.
###### Proposition 5\(Functional\-update failure\)\.
Fixhh, writes\(W,h\)=∥h∥22\(1−∥p∥22\)s\(W,h\)=\\lVert h\\rVert\_\{2\}^\{2\}\(1\-\\lVert p\\rVert\_\{2\}^\{2\}\), letW′=Gv\(W\)W^\{\\prime\}=G\_\{v\}\(W\), and defineΔv\(W\)=∥W′∥F2−∥W∥F2\\Delta\_\{v\}\(W\)=\\lVert W^\{\\prime\}\\rVert\_\{F\}^\{2\}\-\\lVert W\\rVert\_\{F\}^\{2\}\. Then
∇WBraw\(W′,h\)−∇WBraw\(W,h\)=2s𝟏v⊤\+Δv\(W\)∇Ws\(W,h\)\.\\nabla\_\{W\}B\_\{\\mathrm\{raw\}\}\(W^\{\\prime\},h\)\-\\nabla\_\{W\}B\_\{\\mathrm\{raw\}\}\(W,h\)=2s\\mathbf\{1\}v^\{\\top\}\+\\Delta\_\{v\}\(W\)\\nabla\_\{W\}s\(W,h\)\.\(17\)The first term is function\-null, but the remaining probability velocity is
JWpvec\(∇WBraw\(W′,h\)−∇WBraw\(W,h\)\)=−2Δv\(W\)∥h∥24P\(p\)2p\.J\_\{W\}p\\,\\mathrm\{vec\}\(\\nabla\_\{W\}B\_\{\\mathrm\{raw\}\}\(W^\{\\prime\},h\)\-\\nabla\_\{W\}B\_\{\\mathrm\{raw\}\}\(W,h\)\)=\-2\\Delta\_\{v\}\(W\)\\lVert h\\rVert\_\{2\}^\{4\}P\(p\)^\{2\}p\.\(18\)For finite logits andh≠0h\\neq 0, the raw intervention therefore fails exactly whenΔv\(W\)≠0\\Delta\_\{v\}\(W\)\\neq 0andppis not uniform\.
###### Proof\.
Gauge invariance gives the samess,∇Ws\\nabla\_\{W\}s, and probability Jacobian atWWandW′W^\{\\prime\}\. SinceBraw=s∥W∥F2B\_\{\\mathrm\{raw\}\}=s\\lVert W\\rVert\_\{F\}^\{2\}, differentiation gives∇WBraw=2sW\+∥W∥F2∇Ws\\nabla\_\{W\}B\_\{\\mathrm\{raw\}\}=2sW\+\\lVert W\\rVert\_\{F\}^\{2\}\\nabla\_\{W\}s\. Subtracting proves Equation \([17](https://arxiv.org/html/2609.38540#S5.E17)\)\. The common\-row term adds the same amount to every logit and is function\-null\. Finally,∇Ws=−2∥h∥22P\(p\)ph⊤\\nabla\_\{W\}s=\-2\\lVert h\\rVert\_\{2\}^\{2\}P\(p\)p\\,h^\{\\top\}, which yields Equation \([18](https://arxiv.org/html/2609.38540#S5.E18)\)\. For finite logits,kerP\(p\)=span\{𝟏\}\\ker P\(p\)=\\operatorname\{span\}\\\{\\mathbf\{1\}\\\}\. ∎
###### Corollary 3\(Scalar retuning cannot generically restore identification\)\.
LetF\(θ\)F\(\\theta\)be a differentiable model output with parametersθ=\(W,ξ\)\\theta=\(W,\\xi\)and the common\-row symmetry above\. For any differentiable gauge\-invarianta\(θ\)a\(\\theta\), setS\(θ\)=a\(θ\)∥W∥F2S\(\\theta\)=a\(\\theta\)\\lVert W\\rVert\_\{F\}^\{2\}andc=∥W∥F2c=\\lVert W\\rVert\_\{F\}^\{2\}\. Along one orbit its functional velocity has the form
𝒱S=A\+cB,A=2aJF,Wvec\(W\),B=JF∇θa,\\mathcal\{V\}\_\{S\}=A\+cB,\\qquad A=2aJ\_\{F,W\}\\operatorname\{vec\}\(W\),\\quad B=J\_\{F\}\\nabla\_\{\\theta\}a,\(19\)whereAAandBBare orbit\-invariant\. If they are linearly independent, then for representatives withc1≠c2c\_\{1\}\\neq c\_\{2\}no nonzero scalarsλ1,λ2\\lambda\_\{1\},\\lambda\_\{2\}satisfyλ1𝒱S\(θ1\)=λ2𝒱S\(θ2\)\\lambda\_\{1\}\\mathcal\{V\}\_\{S\}\(\\theta\_\{1\}\)=\\lambda\_\{2\}\\mathcal\{V\}\_\{S\}\(\\theta\_\{2\}\)\. ForS=BrawS=B\_\{\\mathrm\{raw\}\},F=pF=p, and a single fixed\-feature example withh≠0h\\neq 0, this nondegeneracy holds for\(K−1\)\(K\-1\)\-dimensional Lebesgue\-almost\-everyppin the simplex interior wheneverK≥3K\\geq 3\.
###### Proof\.
BecauseFFandaaare gauge\-invariant,JFJ\_\{F\},aa, and∇θa\\nabla\_\{\\theta\}aare the same at every representative\. The product rule givesJF∇θS=2aJF,Wvec\(W\)\+cJF∇θaJ\_\{F\}\\nabla\_\{\\theta\}S=2aJ\_\{F,W\}\\operatorname\{vec\}\(W\)\+cJ\_\{F\}\\nabla\_\{\\theta\}a\. The gauge shift changesWWby𝟏v⊤\\mathbf\{1\}v^\{\\top\}, butJF,Wvec\(𝟏v⊤\)=0J\_\{F,W\}\\operatorname\{vec\}\(\\mathbf\{1\}v^\{\\top\}\)=0\. Thus bothAAandBBare orbit\-invariant, proving Equation \([19](https://arxiv.org/html/2609.38540#S5.E19)\)\.
Suppose scalar retuning aligned two representatives\. Expanding the equality of their velocities gives\(λ1−λ2\)A\+\(λ1c1−λ2c2\)B=0\(\\lambda\_\{1\}\-\\lambda\_\{2\}\)A\+\(\\lambda\_\{1\}c\_\{1\}\-\\lambda\_\{2\}c\_\{2\}\)B=0\. Linear independence forces both coefficients to vanish\. Henceλ1=λ2\\lambda\_\{1\}=\\lambda\_\{2\}, and their nonzero common value then forcesc1=c2c\_\{1\}=c\_\{2\}, a contradiction\. Any shared gauge\-invariant primary\-loss field is identical at both representatives and cancels before this comparison\.
For the raw fixed\-feature case, takeF=pF=pand writez=Whz=Wh\. Sincep=softmax\(z\)p=\\operatorname\{softmax\}\(z\), we havez=logp\+γ𝟏z=\\log p\+\\gamma\\mathbf\{1\}for some scalarγ\\gamma\. UsingP\(p\)𝟏=0P\(p\)\\mathbf\{1\}=0givesJp,Wvec\(W\)=P\(p\)z=P\(p\)logpJ\_\{p,W\}\\operatorname\{vec\}\(W\)=P\(p\)z=P\(p\)\\log p\. Also,Jp,Wvec\(∇Ws\)=−2∥h∥24P\(p\)2pJ\_\{p,W\}\\operatorname\{vec\}\(\\nabla\_\{W\}s\)=\-2\\lVert h\\rVert\_\{2\}^\{4\}P\(p\)^\{2\}p\. ThereforeAAandBBpoint, up to nonzero scalar factors, alongP\(p\)logpP\(p\)\\log pandP\(p\)2pP\(p\)^\{2\}p\. ForK=3K=3, their first\-two\-coordinate minor atp=\(1/2,3/10,1/5\)p=\(1/2,3/10,1/5\)equals925000log\(640/6561\)≠0\\frac\{9\}\{25000\}\\log\(640/6561\)\\neq 0\. This minor is analytic in the simplex interior, so its zero set has measure zero\. The set where all minors vanish is a subset of that zero set\. ForK\>3K\>3, assigning sufficiently small positive mass to the additional classes preserves a nonzero minor and gives the same conclusion\. ∎
The raw objective also rewards motion in the function\-null common\-row direction\. Letw¯=K−1W⊤𝟏\\bar\{w\}=K^\{\-1\}W^\{\\top\}\\mathbf\{1\}\. Because the primary loss andssare invariant to common\-row shifts, their gradients have zero row mean\. Taking the row mean of a gradient\-ascent step onαBraw\\alpha B\_\{\\mathrm\{raw\}\}therefore gives
w¯t\+1=\(1\+2ηαst\)w¯t,\\bar\{w\}\_\{t\+1\}=\(1\+2\\eta\\alpha s\_\{t\}\)\\bar\{w\}\_\{t\},\(20\)Forα\>0\\alpha\>0,st\>0s\_\{t\}\>0, andw¯t≠0\\bar\{w\}\_\{t\}\\neq 0, repeated steps amplify this component even though it leaves probabilities unchanged\. This is the gauge\-runaway prediction tested below\. Quotient and exact objectives do not contain this function\-null growth term\.
#### Empirical setup\.
We first ask whether the defect changes realistic one\-step updates, then examine how it manifests over full training trajectories\. In every pair, non\-classifier parameters, optimizer state, minibatches, and initial probabilities match\. We calibrate the auxiliary coefficient once at the centered model and reuse it unchanged after the gauge shift\. Tests cover addition and multiplication Transformers, a subtraction MLP, frozen ResNet\-18 features, and end\-to\-end CIFAR\-10 ResNet\-18 checkpoints\. Long\-horizon experiments use fixed decision rules specified before evaluation\. We draw no trajectory claim from one step\.
### 5\.1Standard and naturally stored gauges
Any equivalent reparameterization is enough to break identification, but ordinary gauge scales determine practical relevance\. In ten frozen\-feature CIFAR\-10 probes evaluated as stored, centeringWWas inBquotB\_\{\\mathrm\{quot\}\}removes a mean\-row component containing1\.931\.93–2\.36%2\.36\\%of∥W∥F2\\lVert W\\rVert\_\{F\}^\{2\}\. In five end\-to\-end ResNets, no shift is applied\. The raw/quotient ratio starts at1\.1031\.103–1\.1181\.118, and weight decay drives it toward the centered value of one \(Figure[1](https://arxiv.org/html/2609.38540#S5.F1)a\)\.
Figure 1:Gauge scale without amplified shifts\.\(a\)With no gauge shift, weight decay drives five ordinary end\-to\-end classifiers toward the centered ratio of one\.\(b\)Each checkpoint is reparameterized ten times by setting one class row to zero\. Black marks the exact mean of two from Corollary[1](https://arxiv.org/html/2609.38540#Thmcorollary1)\.For each of five centered checkpoints, we subtract each class row from all ten rows in turn\. The chosen row becomes zero, but probabilities do not change\. The 50 resulting reference\-class parameterizations have raw/quotient ratios of1\.9391\.939–2\.0452\.045and an exact within\-checkpoint mean of two \(Figure[1](https://arxiv.org/html/2609.38540#S5.F1)b\)\. Thus a standard equivalent representation can nearly double the proxy even when natural storage is nearly centered\. Corollary[3](https://arxiv.org/html/2609.38540#Thmcorollary3)independently rules out generic repair by gauge\-specific scalar retuning\.
### 5\.2Equivalent predictors receive different updates
On ten trained modular\-addition Transformer checkpoints, common\-row shifts makeBrawB\_\{\\mathrm\{raw\}\}11,1010, or100100times its centered value while probabilities,BquotB\_\{\\mathrm\{quot\}\}, andκexact\\kappa\_\{\\mathrm\{exact\}\}stay fixed \(Figure[2](https://arxiv.org/html/2609.38540#S5.F2)a\)\. At100×100\\times, the raw full gradient differs from the centered one by approximately9999times the centered\-gradient norm, despite probability differences below5\.11×10−155\.11\\times 10^\{\-15\}\. Quotient and exact gradient differences remain below6\.23×10−126\.23\\times 10^\{\-12\}\(Figure[2](https://arxiv.org/html/2609.38540#S5.F2)b\)\.
Figure 2:Only raw changes along a function\-null orbit\.Ten checkpoints \(thin\) and medians \(emphasized\)\. Probabilities,BquotB\_\{\\mathrm\{quot\}\}, andκexact\\kappa\_\{\\mathrm\{exact\}\}stay fixed\.The central empirical question is whether this symmetry defect changes the function after one matched optimizer step\. Across 45 centered/100×\\times\-shifted pairs, it does\. We test ten each for addition Transformers, multiplication Transformers, subtraction MLPs, and frozen\-feature CIFAR\-10, plus five end\-to\-end ResNets\. Initial paired probability differences are below10−1010^\{\-10\}\. After one update, raw separates every pair, while quotient, available exact, and CE controls remain aligned near numerical precision \(Figure[3](https://arxiv.org/html/2609.38540#S5.F3)\)\.
Figure 3:One\-update functional deviation\.Raw separates equivalent functions\. Quotient, available exact, and CE controls remain aligned near machine precision\.The separation is large beyond the algorithmic models\. Subtraction MLPs reach maximum paired probability differences of0\.5040\.504–0\.9840\.984with14\.814\.8–21\.3%21\.3\\%prediction disagreement\. Raw differences are0\.1840\.184–0\.2420\.242in all ten frozen\-feature CIFAR\-10 probes and0\.9900\.990–1\.0001\.000in all five end\-to\-end checkpoints, while the largest quotient difference across image settings is7\.87×10−147\.87\\times 10^\{\-14\}\. The effect is therefore not confined to a toy model or a frozen representation\.
Quotienting repairs gauge identification, butBquotB\_\{\\mathrm\{quot\}\}is not tight\. At the centered modular checkpoints, its value is102\.95102\.95–108\.44108\.44timesκexact\\kappa\_\{\\mathrm\{exact\}\}\. A separate five\-seed temporal null changes only the function\-null coordinate used to reportBrawB\_\{\\mathrm\{raw\}\}, keeping it about100×100\\timesits paired control value while probabilities, centered classifiers, non\-classifier parameters, optimizer states, and minibatches match\. All five pairs have identical grokking onset\. Thus the diagnostic value alone cannot determine generalization\.
### 5\.3Long\-horizon training dynamics
Separate long\-horizon CIFAR\-10 runs show suppression and recovery\. Before removal, raw\-regularized models trail the paired CE control by10\.6110\.61percentage points on average \(range7\.047\.04–16\.5816\.58\)\. Accuracy rises by11\.9011\.90points \(8\.488\.48–17\.1117\.11\) after regularizer removal and the scheduled optimizer changes\. This supports suppression and recovery, but not a consistent selective\-delay pattern\. Three runs memorize without reaching the required ten\-point test gap, while two reach the gap without memorizing\. The unchanged released loop likewise shows suppression and recovery in all three published seeds, although only one satisfies the same composite endpoint\.
Modular addition is less decisive\. No observed raw run shows the required 3000\-step onset separation, and three raw onsets remain unobserved by step 100000\. The exact\-gradient\-matched diagnostic does not resolve whether exact flatness is necessary\. Quotient variants satisfy the composite CIFAR\-10 endpoint in 1/5 and 2/5 runs, so they likewise do not resolve the stronger claim\.
With the raw objective and a quotient cap that leaves the common\-row component unconstrained, all five CIFAR\-10 stability runs cross the runaway threshold at epoch 4 and later exceed common\-row norm10310^\{3\}\.
## 6Limitations and conclusion
The quotient repair removes only final\-softmax common\-row symmetry and remains about two orders aboveκexact\\kappa\_\{\\mathrm\{exact\}\}\. Natural raw/quotient gaps are small at final CIFAR\-10 checkpoints, although standard reference\-class conventions changeBrawB\_\{\\mathrm\{raw\}\}about twofold\. The100×100\\timestests amplify signal deliberately\. The learned\-feature study is not an input\-space benchmark, and long\-horizon tests cover modular addition and CIFAR\-10\.
Within this scope, the raw proxy fails in two distinct ways\. In robustness certification, empirical\-risk stationarity does not remove the pointwise first\-order loss term; retaining it yields a globally valid, gauge\-invariant feature\-space certificate\. As a training objective, the raw proxy assigns different functional updates to equivalent predictors; quotienting the common\-row symmetry restores identification\. Long\-horizon experiments show strong reversible suppression on CIFAR\-10, but do not consistently isolate selective delay after memorization, and therefore do not resolve whether exact relative flatness is necessary for generalization\. More broadly, validity as a geometric upper bound is not enough for downstream use\. Certification requires a sound pointwise bound, while intrinsic training interventions require symmetry\-consistent functional updates\.111Generative AI tools assisted with experimental planning, code development, manuscript preparation, and checks of mathematical derivations and experimental results\.
## References
- Dinh et al\. \(2017\)Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio\.Sharp minima can generalize for deep nets\.In*Proceedings of the 34th International Conference on Machine Learning*, 2017\.
- Foret et al\. \(2021\)Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur\.Sharpness\-aware minimization for efficiently improving generalization\.In*International Conference on Learning Representations*, 2021\.
- Han et al\. \(2025\)Ting Han, Linara Adilova, Henning Petzka, Jens Kleesiek, and Michael Kamp\.Flatness is necessary, neural collapse is not: Rethinking generalization via grokking\.In*Advances in Neural Information Processing Systems*, 2025\.URL[https://openreview\.net/forum?id=lbtOctHDQ3](https://openreview.net/forum?id=lbtOctHDQ3)\.
- Hastie et al\. \(2009\)Trevor Hastie, Robert Tibshirani, and Jerome Friedman\.*The Elements of Statistical Learning: Data Mining, Inference, and Prediction*\.Springer, New York, 2 edition, 2009\.URL[https://hastie\.su\.domains/ElemStatLearn/](https://hastie.su.domains/ElemStatLearn/)\.
- Jang et al\. \(2022\)Cheongjae Jang, Sungyoon Lee, Frank Park, and Yung\-Kyun Noh\.A reparametrization\-invariant sharpness measure based on information geometry\.In*Advances in Neural Information Processing Systems*, volume 35, 2022\.
- Petzka et al\. \(2021\)Henning Petzka, Michael Kamp, Linara Adilova, Cristian Sminchisescu, and Mario Boley\.Relative flatness and generalization\.In*Advances in Neural Information Processing Systems*, 2021\.
- Pittorino et al\. \(2022\)Fabrizio Pittorino, Antonio Ferraro, Gabriele Perugini, Christoph Feinauer, Carlo Baldassi, and Riccardo Zecchina\.Deep networks on toroids: Removing symmetries reveals the structure of flat regions in the landscape geometry\.In*Proceedings of the 39th International Conference on Machine Learning*, volume 162 of*Proceedings of Machine Learning Research*, 2022\.
- Power et al\. \(2022\)Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra\.Grokking: Generalization beyond overfitting on small algorithmic datasets\.*arXiv preprint arXiv:2201\.02177*, 2022\.
- Tsuzuku et al\. \(2020\)Yusuke Tsuzuku, Issei Sato, and Masashi Sugiyama\.Normalized flat minima: Exploring scale invariant definition of flat minima for neural networks using PAC\-Bayesian analysis\.In*Proceedings of the 37th International Conference on Machine Learning*, volume 119 of*Proceedings of Machine Learning Research*, pages 9636–9647, 2020\.
- Walter et al\. \(2026\)Nils Philipp Walter, Linara Adilova, Jilles Vreeken, and Michael Kamp\.When flatness does \(not\) guarantee adversarial robustness\.In*International Conference on Learning Representations*, 2026\.URL[https://openreview\.net/forum?id=sptCQnKS9X](https://openreview.net/forum?id=sptCQnKS9X)\.
- Zahid and Tutz \(2009\)Faisal Maqbool Zahid and Gerhard Tutz\.Ridge estimation for multinomial logit models with symmetric side constraints\.Technical Report 67, Department of Statistics, Ludwig\-Maximilians\-Universität München, 2009\.Similar Articles
Geometry Is Not Robustness: A Trajectory-Level Study of PGD Evaluation
This paper conducts a trajectory-level investigation of PGD attacks on CNNs trained on Fashion-MNIST, showing that while a robustness hierarchy exists, trajectory metrics like loss evolution and gradient alignment do not independently measure robustness, with steps-to-failure providing clearer separation.
Are Flat Minima an Illusion?
This paper challenges the common belief that flat minima cause better generalization in neural networks, arguing that 'weakness'—a reparameterization-invariant measure of function simplicity—is the true driver. Empirical results on MNIST and Fashion-MNIST show that weakness predicts generalization while sharpness anticorrelates, and the large-batch generalization advantage vanishes as training data increases.
Are Safety Guarantees in Neural Networks Safe? How to Compute Trustworthy Robustness Certifications
This paper introduces the apothem measure for computing trustworthy robustness certifications in neural networks, proves intractability of volume-optimal certifications, and presents the ParallelepipedoNN system achieving twofold improvement in minimum edge length on MNIST and Fashion MNIST.
Beyond Surface Statistics: Robust Conformal Prediction for LLMs via Internal Representations
This paper proposes a conformal prediction framework for LLMs that leverages internal representations rather than output-level statistics, introducing Layer-Wise Information (LI) scores as nonconformity measures to improve validity-efficiency trade-offs under distribution shift. The method demonstrates stronger robustness to calibration-deployment mismatch compared to text-level baselines across QA benchmarks.
When Certificates Fail: A Unified Safety Framework for Embedded Neural Interface Models
This paper demonstrates that formal robustness certificates for embedded neural interface models can pass even when task accuracy collapses under adversarial attack, and proposes a unified empirical audit framework to address alignment failures between training objectives and operational user welfare.