Fisher Widths: Local Learning Geometry and Anisotropic Recovery
Summary
This paper introduces Fisher width and inverse-Fisher width on statistical manifolds, studying their roles in local learning bounds and anisotropic recovery. It proves a complementary relation between the two widths and obtains recovery estimates based on Fisher geometry.
View Cached Full Text
Cached at: 07/24/26, 05:12 AM
# Fisher Widths: Local Learning Geometry and Anisotropic Recovery
Source: [https://arxiv.org/html/2607.20578](https://arxiv.org/html/2607.20578)
###### Abstract
We study Gaussian\-width complexity on statistical manifolds through a pair of functionals: the primal Fisher widthwG\(T\)=w\(G1/2T\)w\_\{G\}\(T\)=w\(G^\{1/2\}T\), induced by the Fisher metric, and the inverse\-Fisher widthwG−1\(T\)=w\(G−1/2T\)w\_\{G^\{\-1\}\}\(T\)=w\(G^\{\-1/2\}T\), induced by the inverse Fisher metric\. The two widths play complementary statistical roles\.
On the learning side, the Fisher width measures the size of local parameter fluctuations in the geometry induced by the Fisher information\. For Fisher\-regular losses, we prove that the scalewG\(Hr\)/nw\_\{G\}\(H\_\{r\}\)/\\sqrt\{n\}is attained on sufficiently small Fisher balls\.
On the recovery side, the inverse\-Fisher width captures the effect of anisotropic Gaussian measurements whose covariance is determined by the inverse Fisher information\. For sparse recovery, the resulting geometry depends not only on sparsity but also on the position of the active coordinates in the Fisher spectrum\. We obtain a two\-sided estimate for the corresponding statistical dimension, together with support\-sensitive recovery estimates and a natural ordering of supports with different curvature profiles\.
Finally, we establish a sharp relation between the primal and inverse\-Fisher widths\. On any common compact coordinate setTT, they satisfy
wG\(T\)wG−1\(T\)≥w\(T\)2\.w\_\{G\}\(T\)w\_\{G^\{\-1\}\}\(T\)\\geq w\(T\)^\{2\}\.Thus, Fisher anisotropy may transfer complexity from one geometry to the other, but cannot reduce both widths relative to the Euclidean scale\.
## 1Introduction
### 1\.1From Fisher width to a primal–inverse pair
The Fisher information matrixG\(θ\)G\(\\theta\)defines the local geometry of a statistical model\. Directions of large Fisher curvature are directions in which the model distribution changes rapidly, while directions of small Fisher curvature are statistically flat\. The Fisher metric therefore measures local sensitivity to parameter perturbations, whereas its inverse determines the covariance scale appearing in efficient\-estimation geometry, as reflected in the Cramér–Rao bound\. ThusG\(θ\)G\(\\theta\)andG\(θ\)−1G\(\\theta\)^\{\-1\}induce two complementary deformations of local parameter sets\.
Fisher width was introduced inKy \([2026](https://arxiv.org/html/2607.20578#bib.bib1)\)as
wG\(T\)=w\(G\(θ\)1/2T\),w\_\{G\}\(T\)=w\\bigl\(G\(\\theta\)^\{1/2\}T\\bigr\),which extends classical Gaussian width to statistical models endowed with the Fisher metric\. It is the Gaussian width of the Fisher\-deformed setG\(θ\)1/2TG\(\\theta\)^\{1/2\}T, and measures the size of a parameter set in the local Fisher geometry\.
The present paper studies this width together with its inverse\-metric counterpart,
wG−1\(T\)=w\(G\(θ\)−1/2T\)\.w\_\{G^\{\-1\}\}\(T\)=w\\bigl\(G\(\\theta\)^\{\-1/2\}T\\bigr\)\.We refer to\(wG,wG−1\)\\bigl\(w\_\{G\},w\_\{G^\{\-1\}\}\\bigr\)as a*primal–inverse pair*, corresponding to the Fisher and inverse\-Fisher deformations of the same local parameter set\.
The two widths arise in different statistical settings\. The Fisher width is associated with score fluctuations and local learning bounds, whereas the inverse\-Fisher width appears in recovery problems with Gaussian measurement covarianceG−1G^\{\-1\}\. When evaluated on a common localized coordinate set in a fixed chart, the two deformations respond oppositely to Fisher anisotropy\. The main questions of this paper are how these widths enter learning and recovery bounds, and what relations constrain them when they are applied to the same coordinate set\.
Throughout this paper, the Fisher information matrixG\(θ\)G\(\\theta\)is evaluated at a fixed reference pointθ0\\theta\_\{0\}, and all width functionals are computed in a chosen local parameter chart\. The corresponding tangent and cotangent transformation laws are recorded in Section[2\.3](https://arxiv.org/html/2607.20578#S2.SS3)\.
The Fisher information matrixG\(θ\)G\(\\theta\)defines the local geometry of a statistical model\. Directions of large Fisher curvature are directions in which the model distribution changes rapidly, while directions of small Fisher curvature are statistically flat\. The Fisher metric therefore measures local sensitivity to parameter perturbations, whereas its inverse describes the corresponding scale of estimation uncertainty, as reflected in the Cramér–Rao bound\. ThusG\(θ\)G\(\\theta\)andG\(θ\)−1G\(\\theta\)^\{\-1\}induce two complementary deformations of local parameter sets\.
Fisher width was introduced inKy \([2026](https://arxiv.org/html/2607.20578#bib.bib1)\)as
wG\(T\)=w\(G\(θ\)1/2T\),w\_\{G\}\(T\)=w\\bigl\(G\(\\theta\)^\{1/2\}T\\bigr\),which extends classical Gaussian width to statistical models endowed with the Fisher metric\. It is the Gaussian width of the Fisher\-deformed setG\(θ\)1/2TG\(\\theta\)^\{1/2\}T, and measures the size of a parameter set in the local Fisher geometry\.
The present paper studies this width together with its inverse\-metric counterpart,
wG−1\(T\)=w\(G\(θ\)−1/2T\)\.w\_\{G^\{\-1\}\}\(T\)=w\\bigl\(G\(\\theta\)^\{\-1/2\}T\\bigr\)\.We refer to\(wG,wG−1\)\\bigl\(w\_\{G\},w\_\{G^\{\-1\}\}\\bigr\)as a*primal–inverse pair*, corresponding to the Fisher and inverse\-Fisher deformations of the same local parameter set\.
The two widths arise in different statistical settings\. The Fisher width is associated with score fluctuations and learning complexity, whereas the inverse\-Fisher width appears in recovery problems with Gaussian measurement covarianceG−1G^\{\-1\}\. When evaluated on a common localized coordinate set, the two deformations respond oppositely to Fisher anisotropy\. The main questions of this paper are how these widths enter learning and recovery bounds, and what relations constrain them when the same local geometry is relevant to both problems\.
Throughout this paper, the Fisher information matrixG\(θ\)G\(\\theta\)is evaluated at a fixed reference pointθ0\\theta\_\{0\}, and all width functionals are computed in a chosen local parameter chart\. The corresponding tangent and cotangent transformation laws are recorded in Section[2\.3](https://arxiv.org/html/2607.20578#S2.SS3)\.
### 1\.2Related work
Gaussian width and high\-dimensional geometry\.Gaussian width is a central complexity measure in asymptotic convex geometry, high\-dimensional probability, and empirical process theory\(Talagrand,[2005](https://arxiv.org/html/2607.20578#bib.bib25); Ledoux and Talagrand,[1991](https://arxiv.org/html/2607.20578#bib.bib6); Vershynin,[2018](https://arxiv.org/html/2607.20578#bib.bib46); Wainwright,[2019](https://arxiv.org/html/2607.20578#bib.bib47)\)\. It measures the size of a set through its interaction with a Gaussian process and appears in concentration, random projection, embedding, and uniform\-deviation estimates\(Boucheronet al\.,[2013](https://arxiv.org/html/2607.20578#bib.bib7); Plan and Vershynin,[2014](https://arxiv.org/html/2607.20578#bib.bib12)\)\. The Fisher width introduced inKy \([2026](https://arxiv.org/html/2607.20578#bib.bib1)\)may be viewed as a Fisher\-geometric analogue of Gaussian width, obtained by deforming the parameter set through the local Fisher metric\.
Conic phase transitions and convex recovery\.The geometric theory of recovery thresholds begins with Gordon’s escape\-through\-a\-mesh theorem\(Gordon,[1988](https://arxiv.org/html/2607.20578#bib.bib35)\)\. Compressed sensing and convex recovery subsequently connected exact recovery with descent cones, Gaussian width, and statistical dimension\(Candèset al\.,[2006](https://arxiv.org/html/2607.20578#bib.bib8); Donoho,[2006](https://arxiv.org/html/2607.20578#bib.bib9); Donoho and Tanner,[2009](https://arxiv.org/html/2607.20578#bib.bib26); Foucart and Rauhut,[2013](https://arxiv.org/html/2607.20578#bib.bib10); Chandrasekaranet al\.,[2012](https://arxiv.org/html/2607.20578#bib.bib44); Amelunxenet al\.,[2014](https://arxiv.org/html/2607.20578#bib.bib36)\)\. For Gaussian measurements, Gordon’s theorem gives a sufficient condition in terms of Gaussian width, whereas the statistical dimension determines the sharp conic transition\(Amelunxenet al\.,[2014](https://arxiv.org/html/2607.20578#bib.bib36)\)\. We use this framework through the identity
G−1/2D\(∥⋅∥1,x⋆\)=D\(∥G1/2⋅∥1,G−1/2x⋆\),G^\{\-1/2\}D\(\\\|\\cdot\\\|\_\{1\},x^\{\\star\}\)=D\\bigl\(\\\|G^\{1/2\}\\cdot\\\|\_\{1\},G^\{\-1/2\}x^\{\\star\}\\bigr\),which reduces inverse\-Fisher recovery to a standard Gaussian recovery problem with a weightedℓ1\\ell\_\{1\}descent cone\. The corresponding upper functional is the standard weighted\-ℓ1\\ell\_\{1\}distance\-to\-subdifferential expression\. Our contribution is its Fisher interpretation and a two\-sided estimate obtained by optimizing the standard ALMT error term over vectors with the same support and sign pattern, together with the resulting support\-ordering consequences\.
Anisotropic random measurements and weightedℓ1\\ell\_\{1\}recovery\.Classical Gaussian recovery theory assumes isotropic measurement ensembles\.Kueng and Gross \([2014](https://arxiv.org/html/2607.20578#bib.bib20)\)established RIPless compressed\-sensing bounds for anisotropic ensembles with sampling rates depending on the condition number of the covariance matrix, whileRudelson and Zhou \([2013](https://arxiv.org/html/2607.20578#bib.bib21)\)analyzed sparse recovery and restricted eigenvalue conditions for subgaussian matrices with nontrivial covariance\.
Weightedℓ1\\ell\_\{1\}minimization has also been studied for nonuniform sparsity, partial support information, and known support distributions\(Khajehnejadet al\.,[2011](https://arxiv.org/html/2607.20578#bib.bib2); Díazet al\.,[2018](https://arxiv.org/html/2607.20578#bib.bib3)\)\. In that literature, weights are typically chosen from prior information about the support and may be optimized to improve the recovery threshold\. Regularization and descent\-cone geometry more generally provide a standard framework for structured recovery\(Tibshirani,[1996](https://arxiv.org/html/2607.20578#bib.bib13); Chandrasekaranet al\.,[2012](https://arxiv.org/html/2607.20578#bib.bib44); Negahbanet al\.,[2012](https://arxiv.org/html/2607.20578#bib.bib14); Amelunxenet al\.,[2014](https://arxiv.org/html/2607.20578#bib.bib36)\)\.
Our sensing model belongs to the anisotropic family, but its covariance is generated by an underlying statistical model\. In a Gaussian location experiment, paired differences produce Gaussian sensing rows with covarianceG−1G^\{\-1\}, yielding measurement operators of the formAG−1/2AG^\{\-1/2\}\. After whitening, the original unweightedℓ1\\ell\_\{1\}regularizer becomes
fG\(x\)=‖G1/2x‖1\.f\_\{G\}\(x\)=\\\|G^\{1/2\}x\\\|\_\{1\}\.For diagonalGG, the resulting weights are therefore determined by the measurement covariance rather than chosen from prior support information\. Accordingly,UG\(S\)U\_\{G\}\(S\)is not new in functional form; it is the standard weighted\-ℓ1\\ell\_\{1\}expression specialized to Fisher\-induced weights\. The new points are the statistical origin of these weights, optimization of the standard ALMT error bound over fixed support and sign patterns, and the resulting support\-sensitive recovery estimate\.
Information geometry and Fisher metrics\.The Fisher information matrix defines the canonical Riemannian metric on statistical manifolds\(Rao,[1945](https://arxiv.org/html/2607.20578#bib.bib27); Amari and Nagaoka,[2000](https://arxiv.org/html/2607.20578#bib.bib28); Čencov,[1982](https://arxiv.org/html/2607.20578#bib.bib17)\)\. Its inverse appears in classical estimation theory through the Cramér–Rao bound and also underlies natural\-gradient methods\(Amari,[1998](https://arxiv.org/html/2607.20578#bib.bib37); Pascanu and Bengio,[2014](https://arxiv.org/html/2607.20578#bib.bib18)\)\. Existing work has focused mainly on estimation, divergence geometry, statistical efficiency, and natural\-gradient optimization\. In contrast, we study Gaussian\-width complexity under the two linear deformations induced byGGandG−1G^\{\-1\}\.
Fisher information in machine learning\.Fisher information is also used as a curvature matrix in machine learning, notably in natural\-gradient methods and approximations such as K\-FAC\(Martens and Grosse,[2015](https://arxiv.org/html/2607.20578#bib.bib22)\)\. Empirical Fisher approximations may differ from the population Fisher or Hessian\(Kunstneret al\.,[2019](https://arxiv.org/html/2607.20578#bib.bib24)\); our analysis uses the Fisher matrix to define geometric complexity rather than as an optimization preconditioner\.
### 1\.3Main contributions
Our main contributions are as follows\.
1. 1\.Local attainment on Fisher balls\.For Fisher\-regular losses, we prove a finite\-sample lower bound of order wG\(Hr\)n,\\frac\{w\_\{G\}\(H\_\{r\}\)\}\{\\sqrt\{n\}\},attaining the same scale as the standard Fisher\-Lipschitz upper bound on sufficiently small Fisher balls\. The result applies to standard correctly specified models under local leverage and fourth\-moment assumptions\.
2. 2\.Support\-sensitive anisotropic recovery\.For diagonalGGand unweighted basis pursuit, the transformed cone is the descent cone of the weighted norm x⟼‖G1/2x‖1\.x\\longmapsto\\\|G^\{1/2\}x\\\|\_\{1\}\.The corresponding standard weighted\-ℓ1\\ell\_\{1\}upper functional isUG\(S\)U\_\{G\}\(S\)\. By optimizing the ALMT error bound over vectors with the same support and sign pattern, we obtain UG\(S\)−2Tr\(G\)∑i∈Sγi≤δ\(G−1/2D\(∥⋅∥1,x⋆\)\)≤UG\(S\)\.U\_\{G\}\(S\)\-2\\sqrt\{\\frac\{\\operatorname\{Tr\}\(G\)\}\{\\sum\_\{i\\in S\}\\gamma\_\{i\}\}\}\\leq\\delta\\\!\\left\(G^\{\-1/2\}D\(\\\|\\cdot\\\|\_\{1\},x^\{\\star\}\)\\right\)\\leq U\_\{G\}\(S\)\.The same representation yields monotonicity under nested supports and shows that replacing active coordinates by coordinates of larger Fisher curvature increases the upper recovery estimate\.
3. 3\.Sharp primal–inverse width inequality\.For every nonempty compact setT⊂ℝdT\\subset\\mathbb\{R\}^\{d\}and everyG≻0G\\succ 0, we prove wG\(T\)wG−1\(T\)≥w\(T\)2\.w\_\{G\}\(T\)\\,w\_\{G^\{\-1\}\}\(T\)\\geq w\(T\)^\{2\}\.The constant is sharp, and the proof follows from log\-convexity along commuting powers of the metric\. We also derive related bounds for more general pairs of positive\-definite matrices\.
4. 4\.Fisher interpretation of inverse\-covariance recovery\.Differences of Gaussian location observations generate sensing rows with covarianceG−1G^\{\-1\}\. After whitening, convex recovery is governed by the transformed cone G−1/2D\(R,x⋆\)\.G^\{\-1/2\}D\(R,x^\{\\star\}\)\.Gordon’s theorem gives a sufficient recovery scale through Gaussian width, while statistical dimension locates the corresponding sharp Gaussian conic transition\.
## 2Fisher and Inverse\-Fisher Widths
### 2\.1Definitions and basic properties
Let\{pθ:θ∈Θ⊂ℝd\}\\\{p\_\{\\theta\}:\\theta\\in\\Theta\\subset\\mathbb\{R\}^\{d\}\\\}be a parametric family of probability densities with respect to a base measureμ\\mu\. The Fisher information matrix atθ\\thetais
G\(θ\)ij=𝔼pθ\[∂ilogpθ\(X\)∂jlogpθ\(X\)\]\.G\(\\theta\)\_\{ij\}=\\mathbb\{E\}\_\{p\_\{\\theta\}\}\\left\[\\partial\_\{i\}\\log p\_\{\\theta\}\(X\)\\,\\partial\_\{j\}\\log p\_\{\\theta\}\(X\)\\right\]\.\(1\)Throughout the paper, we evaluate the Fisher matrix at a fixed reference pointθ0\\theta\_\{0\}\. We assume thatG:=G\(θ0\)≻0,G:=G\(\\theta\_\{0\}\)\\succ 0,and work in a chosen local parameter chart aroundθ0\\theta\_\{0\}\. The Fisher metric and its inverse induce the norm pair
‖v‖G=\(v⊤Gv\)1/2,‖s‖G−1=\(s⊤G−1s\)1/2\.\\\|v\\\|\_\{G\}=\(v^\{\\top\}Gv\)^\{1/2\},\\qquad\\\|s\\\|\_\{G^\{\-1\}\}=\(s^\{\\top\}G^\{\-1\}s\)^\{1/2\}\.For a compact setS⊂ℝdS\\subset\\mathbb\{R\}^\{d\}, its Gaussian width is defined by
w\(S\):=𝔼gsupv∈S⟨g,v⟩,g∼N\(0,Id\)\.w\(S\):=\\mathbb\{E\}\_\{g\}\\sup\_\{v\\in S\}\\langle g,v\\rangle,\\qquad g\\sim N\(0,I\_\{d\}\)\.
###### Definition 2\.1\(Fisher and inverse\-Fisher widths\)\.
LetG≻0G\\succ 0and letT⊂ℝdT\\subset\\mathbb\{R\}^\{d\}be compact\. The*Fisher width*and*inverse\-Fisher width*ofTTare
wG\(T\)\\displaystyle w\_\{G\}\(T\):=w\(G1/2T\)=𝔼gsupv∈T⟨g,G1/2v⟩,\\displaystyle:=w\(G^\{1/2\}T\)=\\mathbb\{E\}\_\{g\}\\sup\_\{v\\in T\}\\langle g,G^\{1/2\}v\\rangle,\(2\)wG−1\(T\)\\displaystyle w\_\{G^\{\-1\}\}\(T\):=w\(G−1/2T\)=𝔼gsupv∈T⟨g,G−1/2v⟩\.\\displaystyle:=w\(G^\{\-1/2\}T\)=\\mathbb\{E\}\_\{g\}\\sup\_\{v\\in T\}\\langle g,G^\{\-1/2\}v\\rangle\.\(3\)
The first width is induced by the Fisher metricGG, whereas the second is induced by the inverse metricG−1G^\{\-1\}\. We call\(wG,wG−1\)\(w\_\{G\},w\_\{G^\{\-1\}\}\)the*primal–inverse pair*; this terminology refers only to the two metric deformations\. WhenG=IdG=I\_\{d\}, both quantities reduce to the classical Gaussian width\.
###### Lemma 2\.2\(Basic properties\)\.
LetG≻0G\\succ 0and letT⊂ℝdT\\subset\\mathbb\{R\}^\{d\}be compact\. Then:
1. \(i\)\(Fisher spectral bounds\) λmin\(G\)w\(T\)≤wG\(T\)≤λmax\(G\)w\(T\)\.\\sqrt\{\\lambda\_\{\\min\}\(G\)\}\\,w\(T\)\\leq w\_\{G\}\(T\)\\leq\\sqrt\{\\lambda\_\{\\max\}\(G\)\}\\,w\(T\)\.
2. \(ii\)\(Inverse\-Fisher spectral bounds\) w\(T\)λmax\(G\)≤wG−1\(T\)≤w\(T\)λmin\(G\)\.\\frac\{w\(T\)\}\{\\sqrt\{\\lambda\_\{\\max\}\(G\)\}\}\\leq w\_\{G^\{\-1\}\}\(T\)\\leq\\frac\{w\(T\)\}\{\\sqrt\{\\lambda\_\{\\min\}\(G\)\}\}\.
3. \(iii\)\(Perturbation stability\)IfG1,G2≻0G\_\{1\},G\_\{2\}\\succ 0, then \|wG1\(T\)−wG2\(T\)\|≤‖G11/2−G21/2‖opw\(T\),\|w\_\{G\_\{1\}\}\(T\)\-w\_\{G\_\{2\}\}\(T\)\|\\leq\\\|G\_\{1\}^\{1/2\}\-G\_\{2\}^\{1/2\}\\\|\_\{\\mathrm\{op\}\}\\,w\(T\),and \|wG1−1\(T\)−wG2−1\(T\)\|≤‖G1−1/2−G2−1/2‖opw\(T\)\.\|w\_\{G\_\{1\}^\{\-1\}\}\(T\)\-w\_\{G\_\{2\}^\{\-1\}\}\(T\)\|\\leq\\\|G\_\{1\}^\{\-1/2\}\-G\_\{2\}^\{\-1/2\}\\\|\_\{\\mathrm\{op\}\}\\,w\(T\)\.
###### Proof\.
For the upper bound in \(i\), consider the centered Gaussian processes
Xv:=⟨g,G1/2v⟩,Yv:=λmax\(G\)⟨g,v⟩,v∈T\.X\_\{v\}:=\\langle g,G^\{1/2\}v\\rangle,\\qquad Y\_\{v\}:=\\sqrt\{\\lambda\_\{\\max\}\(G\)\}\\,\\langle g,v\\rangle,\\qquad v\\in T\.For everyu,v∈Tu,v\\in T,
𝔼\|Xv−Xu\|2=\(v−u\)⊤G\(v−u\)≤λmax\(G\)‖v−u‖22=𝔼\|Yv−Yu\|2\.\\mathbb\{E\}\|X\_\{v\}\-X\_\{u\}\|^\{2\}=\(v\-u\)^\{\\top\}G\(v\-u\)\\leq\\lambda\_\{\\max\}\(G\)\\\|v\-u\\\|\_\{2\}^\{2\}=\\mathbb\{E\}\|Y\_\{v\}\-Y\_\{u\}\|^\{2\}\.The Sudakov–Fernique comparison theorem therefore gives
wG\(T\)=𝔼supv∈TXv≤𝔼supv∈TYv=λmax\(G\)w\(T\)\.w\_\{G\}\(T\)=\\mathbb\{E\}\\sup\_\{v\\in T\}X\_\{v\}\\leq\\mathbb\{E\}\\sup\_\{v\\in T\}Y\_\{v\}=\\sqrt\{\\lambda\_\{\\max\}\(G\)\}\\,w\(T\)\.Comparing insteadλmin\(G\)⟨g,v⟩\\sqrt\{\\lambda\_\{\\min\}\(G\)\}\\,\\langle g,v\\ranglewith⟨g,G1/2v⟩\\langle g,G^\{1/2\}v\\ranglegives the lower bound\. Applying the same argument toG−1G^\{\-1\}proves \(ii\)\. Part \(iii\) isKy \([2026](https://arxiv.org/html/2607.20578#bib.bib1), Theorem 3\.1\)applied first toG1,G2G\_\{1\},G\_\{2\}and then toG1−1,G2−1G\_\{1\}^\{\-1\},G\_\{2\}^\{\-1\}\. ∎
In applications, the population Fisher matrixGGmay be replaced by an empirical estimateG^\\widehat\{G\}\. Lemma[2\.2](https://arxiv.org/html/2607.20578#S2.Thmtheorem2)then controls the induced errors in both widths\. Since the inverse square\-root map is unstable near singular matrices, estimating the inverse\-Fisher width requires a uniform positive lower bound on the relevant Fisher eigenvalues\.
### 2\.2Statistical interpretation
The inverse Fisher metric is dual to the Fisher norm: for every covectors∈ℝds\\in\\mathbb\{R\}^\{d\},
sup‖h‖G≤1⟨s,h⟩=‖s‖G−1\.\\sup\_\{\\\|h\\\|\_\{G\}\\leq 1\}\\langle s,h\\rangle=\\\|s\\\|\_\{G^\{\-1\}\}\.Indeed,
⟨s,h⟩=⟨G−1/2s,G1/2h⟩≤‖s‖G−1‖h‖G\\langle s,h\\rangle=\\langle G^\{\-1/2\}s,G^\{1/2\}h\\rangle\\leq\\\|s\\\|\_\{G^\{\-1\}\}\\\|h\\\|\_\{G\}by Cauchy–Schwarz, with equality forh=G−1s‖s‖G−1h=\\frac\{G^\{\-1\}s\}\{\\\|s\\\|\_\{G^\{\-1\}\}\}whens≠0s\\neq 0\. Thus score vectors and loss gradients, which act linearly on parameter perturbations, are naturally measured in the inverse Fisher norm\.
The two widths also admit Gaussian\-process representations\. IfS∼N\(0,G\)S\\sim N\(0,G\)andΔ∼N\(0,G−1\)\\Delta\\sim N\(0,G^\{\-1\}\), then
wG\(T\)=𝔼suph∈T⟨S,h⟩,wG−1\(T\)=𝔼suph∈T⟨Δ,h⟩\.w\_\{G\}\(T\)=\\mathbb\{E\}\\sup\_\{h\\in T\}\\langle S,h\\rangle,\\qquad w\_\{G^\{\-1\}\}\(T\)=\\mathbb\{E\}\\sup\_\{h\\in T\}\\langle\\Delta,h\\rangle\.\(4\)Under the usual regularity conditions, the score atθ0\\theta\_\{0\}is centered with covarianceGG, and its normalized sum converges toN\(0,G\)N\(0,G\)\. Likewise, an asymptotically efficient estimator satisfies
n\(θ^n−θ0\)⇒N\(0,G−1\)\.\\sqrt\{n\}\(\\widehat\{\\theta\}\_\{n\}\-\\theta\_\{0\}\)\\Rightarrow N\(0,G^\{\-1\}\)\.Hence the first process in \([4](https://arxiv.org/html/2607.20578#S2.E4)\) has the covariance geometry of local score fluctuations, whereas the second has the covariance geometry of efficient estimation errors\. Section[3](https://arxiv.org/html/2607.20578#S3)develops the learning\-side role ofwGw\_\{G\}, while Section[4](https://arxiv.org/html/2607.20578#S4)studies recovery under inverse\-Fisher Gaussian measurements\.
### 2\.3Coordinate transformations
We record the transformation laws that distinguish intrinsic statements from comparisons made on a common coordinate set\. Letθ′=φ\(θ\)\\theta^\{\\prime\}=\\varphi\(\\theta\)be a smooth reparametrization with invertible JacobianJ:=Dφ\(θ0\)J:=D\\varphi\(\\theta\_\{0\}\)\. Tangent vectors, covectors, and the Fisher matrix transform as
h′=Jh,s′=J−⊤s,G′=J−⊤GJ−1\.h^\{\\prime\}=Jh,\\qquad s^\{\\prime\}=J^\{\-\\top\}s,\\qquad G^\{\\prime\}=J^\{\-\\top\}GJ^\{\-1\}\.Throughout this subsection, all matrix square roots are the principal symmetric square roots\.
###### Proposition 2\.3\(Tangent and cotangent transformation laws\)\.
LetT⊂ℝdT\\subset\\mathbb\{R\}^\{d\}be a compact set of tangent vectors and letS⊂ℝdS\\subset\\mathbb\{R\}^\{d\}be a compact set of covectors\. Then
wG′\(JT\)=wG\(T\),w\(G′\)−1\(J−⊤S\)=wG−1\(S\)\.w\_\{G^\{\\prime\}\}\(JT\)=w\_\{G\}\(T\),\\qquad w\_\{\(G^\{\\prime\}\)^\{\-1\}\}\(J^\{\-\\top\}S\)=w\_\{G^\{\-1\}\}\(S\)\.
###### Proof\.
Set
Q:=\(G′\)1/2JG−1/2=\(J−⊤GJ−1\)1/2JG−1/2\.Q:=\(G^\{\\prime\}\)^\{1/2\}JG^\{\-1/2\}=\(J^\{\-\\top\}GJ^\{\-1\}\)^\{1/2\}JG^\{\-1/2\}\.Then
Q⊤Q=G−1/2J⊤G′JG−1/2=G−1/2J⊤\(J−⊤GJ−1\)JG−1/2=Id\.Q^\{\\top\}Q=G^\{\-1/2\}J^\{\\top\}G^\{\\prime\}JG^\{\-1/2\}=G^\{\-1/2\}J^\{\\top\}\(J^\{\-\\top\}GJ^\{\-1\}\)JG^\{\-1/2\}=I\_\{d\}\.ThusQQis orthogonal and
\(G′\)1/2J=QG1/2\.\(G^\{\\prime\}\)^\{1/2\}J=QG^\{1/2\}\.By rotational invariance of the standard Gaussian law,
wG′\(JT\)=w\(\(G′\)1/2JT\)=w\(QG1/2T\)=w\(G1/2T\)=wG\(T\)\.w\_\{G^\{\\prime\}\}\(JT\)=w\\bigl\(\(G^\{\\prime\}\)^\{1/2\}JT\\bigr\)=w\(QG^\{1/2\}T\)=w\(G^\{1/2\}T\)=w\_\{G\}\(T\)\.
For the cotangent identity, note that
\(G′\)−1=JG−1J⊤\.\(G^\{\\prime\}\)^\{\-1\}=JG^\{\-1\}J^\{\\top\}\.Applying the same argument to the metricG−1G^\{\-1\}under the coordinate maps↦J−⊤ss\\mapsto J^\{\-\\top\}sgives
w\(G′\)−1\(J−⊤S\)=wG−1\(S\)\.w\_\{\(G^\{\\prime\}\)^\{\-1\}\}\(J^\{\-\\top\}S\)=w\_\{G^\{\-1\}\}\(S\)\.∎
## 3A Local Lower Bound on Fisher Balls
For a loss classℱT=\{ℓθ:θ∈T\},\\mathcal\{F\}\_\{T\}=\\\{\\ell\_\{\\theta\}:\\theta\\in T\\\},generalization asks how well the empirical riskR^n\(θ\)\\widehat\{R\}\_\{n\}\(\\theta\)approximates the population riskR\(θ\)R\(\\theta\), uniformly overTT\. When the loss is Fisher\-Lipschitz with constantLL, meaning
\|ℓ\(θ;z\)−ℓ\(θ′;z\)\|≤L‖θ−θ′‖G∀θ,θ′∈T,∀z,\|\\ell\(\\theta;z\)\-\\ell\(\\theta^\{\\prime\};z\)\|\\leq L\\\|\\theta\-\\theta^\{\\prime\}\\\|\_\{G\}\\qquad\\forall\\,\\theta,\\theta^\{\\prime\}\\in T,\\ \\forall\\,z,standard symmetrization and contraction give, with probability at least1−δ1\-\\delta,
supθ∈T\|R\(θ\)−R^n\(θ\)\|≤CLwG\(T\)n\+Blog\(1/δ\)2n,\\sup\_\{\\theta\\in T\}\|R\(\\theta\)\-\\widehat\{R\}\_\{n\}\(\\theta\)\|\\leq CL\\frac\{w\_\{G\}\(T\)\}\{\\sqrt\{n\}\}\+B\\sqrt\{\\frac\{\\log\(1/\\delta\)\}\{2n\}\},whereBBbounds the loss and we have usedwG\(T−T\)≤2wG\(T\)w\_\{G\}\(T\-T\)\\leq 2w\_\{G\}\(T\)for convex symmetricTTcontaining the origin; seeKy \([2026](https://arxiv.org/html/2607.20578#bib.bib1)\)for the full derivation\. ThuswG\(T\)/nw\_\{G\}\(T\)/\\sqrt\{n\}provides the standard Fisher\-geometric upper scale for uniform empirical fluctuations\. Analogous bounds hold under suitable concentration assumptions in place of boundedness\.
This section gives a non\-asymptotic lower bound on Fisher balls, showing that the orderwG\(Hr\)/nw\_\{G\}\(H\_\{r\}\)/\\sqrt\{n\}is attained for a class of Fisher\-regular losses on sufficiently small local neighborhoods\. SinceG1/2Hr=rB2dG^\{1/2\}H\_\{r\}=rB\_\{2\}^\{d\}, the result establishes the local dimensional scalerd/nr\\sqrt\{d/n\}\. It does not provide a lower\-bound principle for arbitrary structured sets or a minimax characterization in terms of Fisher width\.
### 3\.1A local lower bound on Fisher balls
In a regular exponential family, the deterministic log\-partition term cancels from the centered empirical fluctuation, leaving a linear score process\. The example inKy \([2026](https://arxiv.org/html/2607.20578#bib.bib1)\)therefore satisfies
n𝔼supu∈rB2d\|R\(u\)−R^n\(u\)\|⟶wG0\(rB2d\)asn→∞\.\\sqrt\{n\}\\,\\mathbb\{E\}\\sup\_\{u\\in rB\_\{2\}^\{d\}\}\|R\(u\)\-\\widehat\{R\}\_\{n\}\(u\)\|\\longrightarrow w\_\{G\_\{0\}\}\(rB\_\{2\}^\{d\}\)\\qquad\\text\{as \}n\\to\\infty\.The result below replaces this model\-specific asymptotic argument by a finite\-sample linearization with a controlled quadratic remainder on the Fisher ballHr:=\{h:‖h‖G≤r\}H\_\{r\}:=\\\{h:\\\|h\\\|\_\{G\}\\leq r\\\}\.
Throughout this subsection,Z,Z1,…,Zn∼iidPZ,Z\_\{1\},\\ldots,Z\_\{n\}\\stackrel\{\{\\scriptstyle\\mathrm\{iid\}\}\}\{\{\\sim\}\}P, and we write
Pf:=𝔼\[f\(Z\)\],Pnf:=1n∑i=1nf\(Zi\)\.Pf:=\\mathbb\{E\}\[f\(Z\)\],\\qquad P\_\{n\}f:=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}f\(Z\_\{i\}\)\.All expectations and covariances are taken underPP, unless stated otherwise\.
###### Definition 3\.1\(Fisher\-regular loss atθ0\\theta\_\{0\}\)\.
Letρ\>0\\rho\>0\. A lossℓ:Θ×𝒵→ℝ\\ell:\\Theta\\times\\mathcal\{Z\}\\to\\mathbb\{R\}, twice differentiable on\{θ∈Θ:‖θ−θ0‖G≤ρ\}\\\{\\theta\\in\\Theta:\\\|\\theta\-\\theta\_\{0\}\\\|\_\{G\}\\leq\\rho\\\}, is called*Fisher\-regular atθ0\\theta\_\{0\}with radiusρ\\rho*and constants\(G,κ,σH\)\(G,\\kappa,\\sigma\_\{H\}\)if the following conditions hold\.
1. \(FR1\)\(Centered gradient and Fisher covariance\) 𝔼\[∇θℓθ0\(Z\)\]=0,Cov\(∇θℓθ0\(Z\)\)=G≻0\.\\mathbb\{E\}\\\!\\left\[\\nabla\_\{\\theta\}\\ell\_\{\\theta\_\{0\}\}\(Z\)\\right\]=0,\\qquad\\operatorname\{Cov\}\\\!\\left\(\\nabla\_\{\\theta\}\\ell\_\{\\theta\_\{0\}\}\(Z\)\\right\)=G\\succ 0\.
2. \(FR2\)\(Local HessianL2L^\{2\}\-bound\) There exists a random variableM\(Z\)≥0M\(Z\)\\geq 0such that \(𝔼M\(Z\)2\)1/2≤σH\\bigl\(\\mathbb\{E\}M\(Z\)^\{2\}\\bigr\)^\{1/2\}\\leq\\sigma\_\{H\}and, almost surely, for everyθ\\thetasatisfying‖θ−θ0‖G≤ρ\\\|\\theta\-\\theta\_\{0\}\\\|\_\{G\}\\leq\\rhoand everyh∈ℝdh\\in\\mathbb\{R\}^\{d\}, \|h⊤∇θ2ℓθ\(Z\)h\|≤M\(Z\)‖h‖G2\.\\bigl\|h^\{\\top\}\\nabla\_\{\\theta\}^\{2\}\\ell\_\{\\theta\}\(Z\)h\\bigr\|\\leq M\(Z\)\\\|h\\\|\_\{G\}^\{2\}\.
3. \(FR3\)\(Whitened gradient fourth\-moment bound\) The whitened gradientζ:=G−1/2∇θℓθ0\(Z\)\\zeta:=G^\{\-1/2\}\\nabla\_\{\\theta\}\\ell\_\{\\theta\_\{0\}\}\(Z\)satisfies 𝔼‖ζ‖24≤κ4d2\.\\mathbb\{E\}\\\|\\zeta\\\|\_\{2\}^\{4\}\\leq\\kappa^\{4\}d^\{2\}\.
Under Condition[\(FR1\)](https://arxiv.org/html/2607.20578#S3.I1.i1), the whitened gradientζ\\zetais centered and isotropic\. Condition[\(FR3\)](https://arxiv.org/html/2607.20578#S3.I1.i3)supplies the fourth\-moment control used below, while Condition[\(FR2\)](https://arxiv.org/html/2607.20578#S3.I1.i2)controls the Taylor remainder\.
###### Lemma 3\.2\(Remainder bounds\)\.
Letℓ\\ellbe Fisher\-regular atθ0\\theta\_\{0\}with radiusρ\\rhoand constants\(G,κ,σH\)\(G,\\kappa,\\sigma\_\{H\}\)\. For0<r≤ρ0<r\\leq\\rho, define
Rh\(Z\):=ℓθ0\+h\(Z\)−ℓθ0\(Z\)−⟨∇θℓθ0\(Z\),h⟩\.R\_\{h\}\(Z\):=\\ell\_\{\\theta\_\{0\}\+h\}\(Z\)\-\\ell\_\{\\theta\_\{0\}\}\(Z\)\-\\left\\langle\\nabla\_\{\\theta\}\\ell\_\{\\theta\_\{0\}\}\(Z\),h\\right\\rangle\.Then, for allh,h′∈Hrh,h^\{\\prime\}\\in H\_\{r\}, the following bounds hold almost surely:
1. \(i\)\|Rh\(Z\)\|≤12M\(Z\)‖h‖G2\.\|R\_\{h\}\(Z\)\|\\leq\\frac\{1\}\{2\}M\(Z\)\\\|h\\\|\_\{G\}^\{2\}\.
2. \(ii\)\(𝔼Rh\(Z\)2\)1/2≤12σH‖h‖G2\.\\bigl\(\\mathbb\{E\}R\_\{h\}\(Z\)^\{2\}\\bigr\)^\{1/2\}\\leq\\frac\{1\}\{2\}\\sigma\_\{H\}\\\|h\\\|\_\{G\}^\{2\}\.
3. \(iii\)\|Rh\(Z\)−Rh′\(Z\)\|≤M\(Z\)r‖h−h′‖G\.\|R\_\{h\}\(Z\)\-R\_\{h^\{\\prime\}\}\(Z\)\|\\leq M\(Z\)\\,r\\,\\\|h\-h^\{\\prime\}\\\|\_\{G\}\.
###### Proof\.
Taylor’s theorem with integral remainder gives
Rh\(Z\)=∫01\(1−t\)h⊤∇θ2ℓθ0\+th\(Z\)h𝑑t\.R\_\{h\}\(Z\)=\\int\_\{0\}^\{1\}\(1\-t\)\\,h^\{\\top\}\\nabla\_\{\\theta\}^\{2\}\\ell\_\{\\theta\_\{0\}\+th\}\(Z\)h\\,dt\.Parts[\(i\)](https://arxiv.org/html/2607.20578#S3.I2.i1)and[\(ii\)](https://arxiv.org/html/2607.20578#S3.I2.i2)follow immediately from Condition[\(FR2\)](https://arxiv.org/html/2607.20578#S3.I1.i2)\.
For part[\(iii\)](https://arxiv.org/html/2607.20578#S3.I2.i3), setu\(s\):=h′\+s\(h−h′\)u\(s\):=h^\{\\prime\}\+s\(h\-h^\{\\prime\}\)\. SinceHrH\_\{r\}is convex,‖u\(s\)‖G≤r\\\|u\(s\)\\\|\_\{G\}\\leq rfors∈\[0,1\]s\\in\[0,1\]\. Two applications of the fundamental theorem of calculus give
Rh\(Z\)−Rh′\(Z\)=∫01∫01\(h−h′\)⊤∇θ2ℓθ0\+tu\(s\)\(Z\)u\(s\)𝑑t𝑑s\.R\_\{h\}\(Z\)\-R\_\{h^\{\\prime\}\}\(Z\)=\\int\_\{0\}^\{1\}\\int\_\{0\}^\{1\}\(h\-h^\{\\prime\}\)^\{\\top\}\\nabla\_\{\\theta\}^\{2\}\\ell\_\{\\theta\_\{0\}\+tu\(s\)\}\(Z\)u\(s\)\\,dt\\,ds\.Because the Hessian is symmetric, Condition[\(FR2\)](https://arxiv.org/html/2607.20578#S3.I1.i2)is equivalent to
‖G−1/2∇θ2ℓθ\(Z\)G−1/2‖op≤M\(Z\),\\left\\\|G^\{\-1/2\}\\nabla\_\{\\theta\}^\{2\}\\ell\_\{\\theta\}\(Z\)G^\{\-1/2\}\\right\\\|\_\{\\mathrm\{op\}\}\\leq M\(Z\),and therefore implies
\|a⊤∇θ2ℓθ\(Z\)b\|≤M\(Z\)‖a‖G‖b‖Gfor alla,b∈ℝd\.\\bigl\|a^\{\\top\}\\nabla\_\{\\theta\}^\{2\}\\ell\_\{\\theta\}\(Z\)b\\bigr\|\\leq M\(Z\)\\\|a\\\|\_\{G\}\\\|b\\\|\_\{G\}\\qquad\\text\{for all \}a,b\\in\\mathbb\{R\}^\{d\}\.Applying this bound inside the double integral and using‖u\(s\)‖G≤r\\\|u\(s\)\\\|\_\{G\}\\leq rproves[\(iii\)](https://arxiv.org/html/2607.20578#S3.I2.i3)\. ∎
###### Lemma 3\.3\(Remainder empirical complexity\)\.
Under the assumptions of Lemma[3\.2](https://arxiv.org/html/2607.20578#S3.Thmtheorem2),
𝔼suph∈Hr\|\(Pn−P\)Rh\|≤CσHrwG\(Hr\)n,\\mathbb\{E\}\\sup\_\{h\\in H\_\{r\}\}\\bigl\|\(P\_\{n\}\-P\)R\_\{h\}\\bigr\|\\leq C\\,\\sigma\_\{H\}\\,r\\,\\frac\{w\_\{G\}\(H\_\{r\}\)\}\{\\sqrt\{n\}\},whereC\>0C\>0is a universal constant\.
###### Proof\.
By symmetrization,
𝔼suph∈Hr\|\(Pn−P\)Rh\|≤2𝔼suph∈Hr\|Xh\|,Xh:=1n∑i=1nεiRh\(Zi\),\\mathbb\{E\}\\sup\_\{h\\in H\_\{r\}\}\\bigl\|\(P\_\{n\}\-P\)R\_\{h\}\\bigr\|\\leq 2\\,\\mathbb\{E\}\\sup\_\{h\\in H\_\{r\}\}\|X\_\{h\}\|,\\qquad X\_\{h\}:=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\varepsilon\_\{i\}R\_\{h\}\(Z\_\{i\}\),whereε1,…,εn\\varepsilon\_\{1\},\\ldots,\\varepsilon\_\{n\}are independent Rademacher variables, independent of the sample\.
Conditionally onZ1,…,ZnZ\_\{1\},\\ldots,Z\_\{n\}, the process\(Xh\)h∈Hr\(X\_\{h\}\)\_\{h\\in H\_\{r\}\}is symmetric, satisfiesX0=0X\_\{0\}=0, and has sub\-Gaussian increments with respect to
d~n\(h,h′\):=\(𝔼ε\|Xh−Xh′\|2\)1/2\.\\widetilde\{d\}\_\{n\}\(h,h^\{\\prime\}\):=\\left\(\\mathbb\{E\}\_\{\\varepsilon\}\|X\_\{h\}\-X\_\{h^\{\\prime\}\}\|^\{2\}\\right\)^\{1/2\}\.By Lemma[3\.2](https://arxiv.org/html/2607.20578#S3.Thmtheorem2)[\(iii\)](https://arxiv.org/html/2607.20578#S3.I2.i3),
d~n\(h,h′\)≤M¯nrn‖h−h′‖G,M¯n:=\(1n∑i=1nM\(Zi\)2\)1/2\.\\widetilde\{d\}\_\{n\}\(h,h^\{\\prime\}\)\\leq\\frac\{\\overline\{M\}\_\{n\}r\}\{\\sqrt\{n\}\}\\,\\\|h\-h^\{\\prime\}\\\|\_\{G\},\\qquad\\overline\{M\}\_\{n\}:=\\left\(\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}M\(Z\_\{i\}\)^\{2\}\\right\)^\{1/2\}\.Since the conditional process is symmetric andX0=0X\_\{0\}=0, the generic chaining upper bound for sub\-Gaussian processes yields
𝔼εsuph∈Hr\|Xh\|≤Cγ2\(Hr,d~n\)≤CM¯nrnγ2\(Hr,∥⋅∥G\)\.\\mathbb\{E\}\_\{\\varepsilon\}\\sup\_\{h\\in H\_\{r\}\}\|X\_\{h\}\|\\leq C\\gamma\_\{2\}\(H\_\{r\},\\widetilde\{d\}\_\{n\}\)\\leq C\\frac\{\\overline\{M\}\_\{n\}r\}\{\\sqrt\{n\}\}\\gamma\_\{2\}\(H\_\{r\},\\\|\\cdot\\\|\_\{G\}\)\.
Sinceh↦G1/2hh\\mapsto G^\{1/2\}his an isometry from\(Hr,∥⋅∥G\)\(H\_\{r\},\\\|\\cdot\\\|\_\{G\}\)onto\(rB2d,∥⋅∥2\)\(rB\_\{2\}^\{d\},\\\|\\cdot\\\|\_\{2\}\),
γ2\(Hr,∥⋅∥G\)=γ2\(rB2d,∥⋅∥2\)≍rd≍wG\(Hr\);\\gamma\_\{2\}\(H\_\{r\},\\\|\\cdot\\\|\_\{G\}\)=\\gamma\_\{2\}\(rB\_\{2\}^\{d\},\\\|\\cdot\\\|\_\{2\}\)\\asymp r\\sqrt\{d\}\\asymp w\_\{G\}\(H\_\{r\}\);seeTalagrand \([2005](https://arxiv.org/html/2607.20578#bib.bib25)\)\. Therefore,
𝔼εsuph∈Hr\|Xh\|≤CM¯nrwG\(Hr\)n\.\\mathbb\{E\}\_\{\\varepsilon\}\\sup\_\{h\\in H\_\{r\}\}\|X\_\{h\}\|\\leq C\\overline\{M\}\_\{n\}r\\,\\frac\{w\_\{G\}\(H\_\{r\}\)\}\{\\sqrt\{n\}\}\.Finally, Jensen’s inequality and Condition[\(FR2\)](https://arxiv.org/html/2607.20578#S3.I1.i2)give
𝔼M¯n≤\(𝔼M¯n2\)1/2=\(𝔼M\(Z\)2\)1/2≤σH\.\\mathbb\{E\}\\overline\{M\}\_\{n\}\\leq\\bigl\(\\mathbb\{E\}\\overline\{M\}\_\{n\}^\{2\}\\bigr\)^\{1/2\}=\\bigl\(\\mathbb\{E\}M\(Z\)^\{2\}\\bigr\)^\{1/2\}\\leq\\sigma\_\{H\}\.Taking expectation over the sample completes the proof\. ∎
###### Theorem 3\.4\(Local Fisher\-width lower bound for Fisher\-regular losses\)\.
Letℓ\\ellbe Fisher\-regular atθ0\\theta\_\{0\}with radiusρ\\rhoand constants\(G,κ,σH\)\(G,\\kappa,\\sigma\_\{H\}\)\. Let0<r≤ρ0<r\\leq\\rho, and suppose
CσHr≤cκ2,cκ:=1162κ4,C\\sigma\_\{H\}r\\leq\\frac\{c\_\{\\kappa\}\}\{2\},\\qquad c\_\{\\kappa\}:=\\frac\{1\}\{16\\sqrt\{2\}\\,\\kappa^\{4\}\},whereCCis the universal constant in Lemma[3\.3](https://arxiv.org/html/2607.20578#S3.Thmtheorem3)\. Then
𝔼suph∈Hr\|\(Pn−P\)\(ℓθ0\+h−ℓθ0\)\|≥cκ2wG\(Hr\)n\.\\mathbb\{E\}\\sup\_\{h\\in H\_\{r\}\}\\left\|\(P\_\{n\}\-P\)\\bigl\(\\ell\_\{\\theta\_\{0\}\+h\}\-\\ell\_\{\\theta\_\{0\}\}\\bigr\)\\right\|\\geq\\frac\{c\_\{\\kappa\}\}\{2\}\\frac\{w\_\{G\}\(H\_\{r\}\)\}\{\\sqrt\{n\}\}\.
###### Proof\.
Writing
ℓθ0\+h−ℓθ0=⟨∇θℓθ0,h⟩\+Rh,\\ell\_\{\\theta\_\{0\}\+h\}\-\\ell\_\{\\theta\_\{0\}\}=\\left\\langle\\nabla\_\{\\theta\}\\ell\_\{\\theta\_\{0\}\},h\\right\\rangle\+R\_\{h\},the left\-hand side is bounded below byA−BA\-B, where
A:=𝔼suph∈Hr\|\(Pn−P\)⟨∇θℓθ0,h⟩\|,B:=𝔼suph∈Hr\|\(Pn−P\)Rh\|\.A:=\\mathbb\{E\}\\sup\_\{h\\in H\_\{r\}\}\\left\|\(P\_\{n\}\-P\)\\left\\langle\\nabla\_\{\\theta\}\\ell\_\{\\theta\_\{0\}\},h\\right\\rangle\\right\|,\\qquad B:=\\mathbb\{E\}\\sup\_\{h\\in H\_\{r\}\}\|\(P\_\{n\}\-P\)R\_\{h\}\|\.
Define
ζi:=G−1/2∇θℓθ0\(Zi\),Sn:=1n∑i=1nζi\.\\zeta\_\{i\}:=G^\{\-1/2\}\\nabla\_\{\\theta\}\\ell\_\{\\theta\_\{0\}\}\(Z\_\{i\}\),\\qquad S\_\{n\}:=\\frac\{1\}\{\\sqrt\{n\}\}\\sum\_\{i=1\}^\{n\}\\zeta\_\{i\}\.By Condition[\(FR1\)](https://arxiv.org/html/2607.20578#S3.I1.i1), the variablesζi\\zeta\_\{i\}are i\.i\.d\., centered, and isotropic\. Since
1n∑i=1n∇θℓθ0\(Zi\)=1nG1/2Sn\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\nabla\_\{\\theta\}\\ell\_\{\\theta\_\{0\}\}\(Z\_\{i\}\)=\\frac\{1\}\{\\sqrt\{n\}\}G^\{1/2\}S\_\{n\}andG1/2Hr=rB2dG^\{1/2\}H\_\{r\}=rB\_\{2\}^\{d\}, we obtain
A=rn𝔼‖Sn‖2\.A=\\frac\{r\}\{\\sqrt\{n\}\}\\,\\mathbb\{E\}\\\|S\_\{n\}\\\|\_\{2\}\.
Isotropy gives𝔼‖Sn‖22=d\\mathbb\{E\}\\\|S\_\{n\}\\\|\_\{2\}^\{2\}=d\. Moreover,
𝔼‖∑i=1nζi‖24\\displaystyle\\mathbb\{E\}\\left\\\|\\sum\_\{i=1\}^\{n\}\\zeta\_\{i\}\\right\\\|\_\{2\}^\{4\}=n𝔼‖ζ‖24\+n\(n−1\)d2\+2n\(n−1\)d\.\\displaystyle=n\\,\\mathbb\{E\}\\\|\\zeta\\\|\_\{2\}^\{4\}\+n\(n\-1\)d^\{2\}\+2n\(n\-1\)d\.Therefore,
𝔼‖Sn‖24\\displaystyle\\mathbb\{E\}\\\|S\_\{n\}\\\|\_\{2\}^\{4\}=1n𝔼‖ζ‖24\+\(1−1n\)\(d2\+2d\)≤4κ4d2,\\displaystyle=\\frac\{1\}\{n\}\\mathbb\{E\}\\\|\\zeta\\\|\_\{2\}^\{4\}\+\\left\(1\-\\frac\{1\}\{n\}\\right\)\(d^\{2\}\+2d\)\\leq 4\\kappa^\{4\}d^\{2\},where the last inequality uses Condition[\(FR3\)](https://arxiv.org/html/2607.20578#S3.I1.i3),κ4≥1\\kappa^\{4\}\\geq 1, andd≥1d\\geq 1\.
Applying the Paley–Zygmund inequality toY=‖Sn‖22Y=\\\|S\_\{n\}\\\|\_\{2\}^\{2\}yields
ℙ\(‖Sn‖22≥d2\)≥116κ4\.\\mathbb\{P\}\\left\(\\\|S\_\{n\}\\\|\_\{2\}^\{2\}\\geq\\frac\{d\}\{2\}\\right\)\\geq\\frac\{1\}\{16\\kappa^\{4\}\}\.Consequently,
𝔼‖Sn‖2≥d2ℙ\(‖Sn‖22≥d2\)≥cκd\.\\mathbb\{E\}\\\|S\_\{n\}\\\|\_\{2\}\\geq\\sqrt\{\\frac\{d\}\{2\}\}\\,\\mathbb\{P\}\\left\(\\\|S\_\{n\}\\\|\_\{2\}^\{2\}\\geq\\frac\{d\}\{2\}\\right\)\\geq c\_\{\\kappa\}\\sqrt\{d\}\.SincewG\(Hr\)=r𝔼‖g‖2≤rdw\_\{G\}\(H\_\{r\}\)=r\\mathbb\{E\}\\\|g\\\|\_\{2\}\\leq r\\sqrt\{d\},
A≥cκwG\(Hr\)n\.A\\geq c\_\{\\kappa\}\\frac\{w\_\{G\}\(H\_\{r\}\)\}\{\\sqrt\{n\}\}\.Lemma[3\.3](https://arxiv.org/html/2607.20578#S3.Thmtheorem3)andCσHr≤cκ/2C\\sigma\_\{H\}r\\leq c\_\{\\kappa\}/2give
B≤cκ2wG\(Hr\)n\.B\\leq\\frac\{c\_\{\\kappa\}\}\{2\}\\frac\{w\_\{G\}\(H\_\{r\}\)\}\{\\sqrt\{n\}\}\.The conclusion follows from the lower boundA−BA\-B\. ∎
###### Corollary 3\.6\(Examples satisfying Fisher regularity\)\.
Suppose the corresponding model\-specific assumptions of Appendix[A](https://arxiv.org/html/2607.20578#A1)hold\. In particular, assume bounded Fisher leverage for correctly specified logistic regression, the stated local leverage and fourth\-moment conditions for canonical\-link generalized linear models, and
𝔼\(X⊤G−1X\)2<∞\\mathbb\{E\}\\bigl\(X^\{\\top\}G^\{\-1\}X\\bigr\)^\{2\}<\\inftyfor Gaussian linear regression\. Then Conditions\(FR1\)–\(FR3\)hold for the respective models\.
## 4Inverse\-Fisher Recovery
### 4\.1Gaussian sensing and conic reduction
The learning results of the previous section involve fluctuations with covarianceGG\. We now consider Gaussian measurements with covarianceG−1G^\{\-1\}\. A simple source of such measurements is the location modelZ∼N\(θ,G−1\)Z\\sim N\(\\theta,G^\{\-1\}\), whose Fisher information forθ\\thetaisGG\. IfZi,Zi′Z\_\{i\},Z\_\{i\}^\{\\prime\}are independent observations, then
ai:=Zi−Zi′2∼N\(0,G−1\)\.a\_\{i\}:=\\frac\{Z\_\{i\}\-Z\_\{i\}^\{\\prime\}\}\{\\sqrt\{2\}\}\\sim N\(0,G^\{\-1\}\)\.Thus paired differences generate sensing rows of the formai=G−1/2gia\_\{i\}=G^\{\-1/2\}g\_\{i\}, wheregi∼N\(0,Id\)g\_\{i\}\\sim N\(0,I\_\{d\}\), and hence a sensing operatorAG−1/2AG^\{\-1/2\}withAAstandard Gaussian\. This construction motivates the covarianceG−1G^\{\-1\}; the results below apply to any Gaussian design with this covariance\.
For a nonempty compact setT⊂ℝdT\\subset\\mathbb\{R\}^\{d\}, define
rG−1\(T\):=infv∈T‖G−1/2v‖2,r\_\{G^\{\-1\}\}\(T\):=\\inf\_\{v\\in T\}\\\|G^\{\-1/2\}v\\\|\_\{2\},and, whenrG−1\(T\)\>0r\_\{G^\{\-1\}\}\(T\)\>0,
rad\(G−1/2T\):=\{G−1/2v‖G−1/2v‖2:v∈T\}\.\\operatorname\{rad\}\(G^\{\-1/2\}T\):=\\left\\\{\\frac\{G^\{\-1/2\}v\}\{\\\|G^\{\-1/2\}v\\\|\_\{2\}\}:v\\in T\\right\\\}\.For a convex regularizerR:ℝd→ℝR:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}, its descent cone atv0v\_\{0\}is
D\(R,v0\):=cl\{h:R\(v0\+αh\)≤R\(v0\)for someα\>0\}\.D\(R,v\_\{0\}\):=\\operatorname\{cl\}\\left\\\{h:R\(v\_\{0\}\+\\alpha h\)\\leq R\(v\_\{0\}\)\\text\{ for some \}\\alpha\>0\\right\\\}\.
###### Proposition 4\.1\(Inverse\-Fisher conic escape\)\.
LetG≻0G\\succ 0, letA∈ℝm×dA\\in\\mathbb\{R\}^\{m\\times d\}have i\.i\.d\.N\(0,1\)N\(0,1\)entries, and letC⊂ℝdC\\subset\\mathbb\{R\}^\{d\}be a nonzero closed cone\. Setam:=𝔼‖g‖2,g∼N\(0,Im\)\.a\_\{m\}:=\\mathbb\{E\}\\\|g\\\|\_\{2\},\\quad g\\sim N\(0,I\_\{m\}\)\.Then, for everyt\>0t\>0, with probability at least1−e−t2/21\-e^\{\-t^\{2\}/2\}, we have
infv∈C‖G−1/2v‖2=1‖AG−1/2v‖2≥am−w\(G−1/2C∩𝕊d−1\)−t\.\\inf\_\{\\begin\{subarray\}\{c\}v\\in C\\\\ \\\|G^\{\-1/2\}v\\\|\_\{2\}=1\\end\{subarray\}\}\\\|AG^\{\-1/2\}v\\\|\_\{2\}\\geq a\_\{m\}\-w\\bigl\(G^\{\-1/2\}C\\cap\\mathbb\{S\}^\{d\-1\}\\bigr\)\-t\.In particular, we haveker\(AG−1/2\)∩C=\{0\}\\ker\(AG^\{\-1/2\}\)\\cap C=\\\{0\\\}wheneveram\>w\(G−1/2C∩𝕊d−1\)\+t\.a\_\{m\}\>w\\bigl\(G^\{\-1/2\}C\\cap\\mathbb\{S\}^\{d\-1\}\\bigr\)\+t\.
###### Proof\.
Set
T:=C∩\{v:‖G−1/2v‖2=1\}\.T:=C\\cap\\\{v:\\\|G^\{\-1/2\}v\\\|\_\{2\}=1\\\}\.Then
G−1/2T=G−1/2C∩𝕊d−1\.G^\{\-1/2\}T=G^\{\-1/2\}C\\cap\\mathbb\{S\}^\{d\-1\}\.Applying Gordon’s escape theorem\(Gordon,[1988](https://arxiv.org/html/2607.20578#bib.bib35)\); see alsoVershynin \([2018](https://arxiv.org/html/2607.20578#bib.bib46)\), to the setG−1/2C∩𝕊d−1G^\{\-1/2\}C\\cap\\mathbb\{S\}^\{d\-1\}gives
infv∈C‖G−1/2v‖2=1‖AG−1/2v‖2≥am−w\(G−1/2C∩𝕊d−1\)−t\.\\inf\_\{\\begin\{subarray\}\{c\}v\\in C\\\\ \\\|G^\{\-1/2\}v\\\|\_\{2\}=1\\end\{subarray\}\}\\\|AG^\{\-1/2\}v\\\|\_\{2\}\\geq a\_\{m\}\-w\\bigl\(G^\{\-1/2\}C\\cap\\mathbb\{S\}^\{d\-1\}\\bigr\)\-t\.If the right\-hand side is positive, then no nonzero vector inCCbelongs toker\(AG−1/2\)\\ker\(AG^\{\-1/2\}\), and hence
ker\(AG−1/2\)∩C=\{0\}\.\\ker\(AG^\{\-1/2\}\)\\cap C=\\\{0\\\}\.∎
Sinceam≍ma\_\{m\}\\asymp\\sqrt\{m\}, Proposition[4\.1](https://arxiv.org/html/2607.20578#S4.Thmtheorem1)gives the sufficient measurement scale
m≳w\(G−1/2C∩𝕊d−1\)2\.m\\gtrsim w\\bigl\(G^\{\-1/2\}C\\cap\\mathbb\{S\}^\{d\-1\}\\bigr\)^\{2\}\.This is a sufficient escape bound\. The approximate conic kinematic formula ofAmelunxenet al\.\([2014](https://arxiv.org/html/2607.20578#bib.bib36)\)instead locates the sharp Gaussian transition near
m=δ\(G−1/2C\)\.m=\\delta\(G^\{\-1/2\}C\)\.
###### Corollary 4\.2\(Unique convex recovery\)\.
LetR:ℝd→ℝR:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}be convex and letC=D\(R,v0\)C=D\(R,v\_\{0\}\)\. Under the condition
am\>w\(G−1/2C∩𝕊d−1\)\+t,a\_\{m\}\>w\\bigl\(G^\{\-1/2\}C\\cap\\mathbb\{S\}^\{d\-1\}\\bigr\)\+t,the pointv0v\_\{0\}is, with probability at least1−e−t2/21\-e^\{\-t^\{2\}/2\}, the unique solution of
minv∈ℝdR\(v\)subject toAG−1/2v=AG−1/2v0\.\\min\_\{v\\in\\mathbb\{R\}^\{d\}\}R\(v\)\\quad\\text\{subject to\}\\quad AG^\{\-1/2\}v=AG^\{\-1/2\}v\_\{0\}\.
### 4\.2Fisher\-induced weightedℓ1\\ell\_\{1\}geometry
We now specialize to unweightedℓ1\\ell\_\{1\}recovery with a diagonal Fisher matrix
G=diag\(γ1,…,γd\),γi\>0\.G=\\operatorname\{diag\}\(\\gamma\_\{1\},\\ldots,\\gamma\_\{d\}\),\\qquad\\gamma\_\{i\}\>0\.After whitening the measurement operator, the anisotropy appears as a deterministic weight profile in the transformed descent cone\. ForS⊂\[d\]S\\subset\[d\], define
UG\(S\):=infτ≥0\[∑i∈S\(1\+τ2γi\)\+∑j∉S𝔼\(\|gj\|−τγj\)\+2\],U\_\{G\}\(S\):=\\inf\_\{\\tau\\geq 0\}\\left\[\\sum\_\{i\\in S\}\(1\+\\tau^\{2\}\\gamma\_\{i\}\)\+\\sum\_\{j\\notin S\}\\mathbb\{E\}\\bigl\(\|g\_\{j\}\|\-\\tau\\sqrt\{\\gamma\_\{j\}\}\\bigr\)\_\{\+\}^\{2\}\\right\],\(5\)whereg1,…,gdg\_\{1\},\\ldots,g\_\{d\}are independentN\(0,1\)N\(0,1\)variables\.
###### Lemma 4\.4\(Fisher\-induced weighted\-ℓ1\\ell\_\{1\}representation\)\.
Letx⋆∈ℝdx^\{\\star\}\\in\\mathbb\{R\}^\{d\}be nonzero, with supportS=supp\(x⋆\)S=\\operatorname\{supp\}\(x^\{\\star\}\), and set
C:=D\(∥⋅∥1,x⋆\),fG\(x\):=∥G1/2x∥1,x0:=G−1/2x⋆\.C:=D\(\\\|\\cdot\\\|\_\{1\},x^\{\\star\}\),\\qquad f\_\{G\}\(x\):=\\\|G^\{1/2\}x\\\|\_\{1\},\\qquad x\_\{0\}:=G^\{\-1/2\}x^\{\\star\}\.Then
G−1/2C=D\(fG,x0\)G^\{\-1/2\}C=D\(f\_\{G\},x\_\{0\}\)and
δ\(G−1/2C\)≤UG\(S\)\.\\delta\(G^\{\-1/2\}C\)\\leq U\_\{G\}\(S\)\.\(6\)
###### Proof\.
The cone identity follows directly from the change of variablesx=G−1/2vx=G^\{\-1/2\}v:
v∈D\(∥⋅∥1,x⋆\)\\displaystyle v\\in D\(\\\|\\cdot\\\|\_\{1\},x^\{\\star\}\)⇔‖x⋆\+αv‖1≤‖x⋆‖1for someα\>0\\displaystyle\\iff\\\|x^\{\\star\}\+\\alpha v\\\|\_\{1\}\\leq\\\|x^\{\\star\}\\\|\_\{1\}\\quad\\text\{for some \}\\alpha\>0⇔fG\(x0\+αG−1/2v\)≤fG\(x0\)\.\\displaystyle\\iff f\_\{G\}\(x\_\{0\}\+\\alpha G^\{\-1/2\}v\)\\leq f\_\{G\}\(x\_\{0\}\)\.The standard descent\-cone bound gives
δ\(D\(fG,x0\)\)≤infτ≥0𝔼dist2\(g,τ∂fG\(x0\)\)\.\\delta\(D\(f\_\{G\},x\_\{0\}\)\)\\leq\\inf\_\{\\tau\\geq 0\}\\mathbb\{E\}\\operatorname\{dist\}^\{2\}\\bigl\(g,\\tau\\partial f\_\{G\}\(x\_\{0\}\)\\bigr\)\.Here
∂fG\(x0\)=\{z:zi=γisign\(xi⋆\)fori∈S,\|zj\|≤γjforj∉S\}\.\\partial f\_\{G\}\(x\_\{0\}\)=\\left\\\{z:z\_\{i\}=\\sqrt\{\\gamma\_\{i\}\}\\operatorname\{sign\}\(x\_\{i\}^\{\\star\}\)\\text\{ for \}i\\in S,\\;\|z\_\{j\}\|\\leq\\sqrt\{\\gamma\_\{j\}\}\\text\{ for \}j\\notin S\\right\\\}\.The coordinatewise distance computation is the standard weighted\-ℓ1\\ell\_\{1\}descent\-cone recipe\. Fori∈Si\\in S,
𝔼\(gi−τγisign\(xi⋆\)\)2=1\+τ2γi,\\mathbb\{E\}\\left\(g\_\{i\}\-\\tau\\sqrt\{\\gamma\_\{i\}\}\\operatorname\{sign\}\(x\_\{i\}^\{\\star\}\)\\right\)^\{2\}=1\+\\tau^\{2\}\\gamma\_\{i\},whereas, forj∉Sj\\notin S,
𝔼dist2\(gj,\[−τγj,τγj\]\)=𝔼\(\|gj\|−τγj\)\+2\.\\mathbb\{E\}\\operatorname\{dist\}^\{2\}\\left\(g\_\{j\},\[\-\\tau\\sqrt\{\\gamma\_\{j\}\},\\tau\\sqrt\{\\gamma\_\{j\}\}\]\\right\)=\\mathbb\{E\}\\bigl\(\|g\_\{j\}\|\-\\tau\\sqrt\{\\gamma\_\{j\}\}\\bigr\)\_\{\+\}^\{2\}\.Summing over the coordinates and optimizing overτ≥0\\tau\\geq 0proves \([6](https://arxiv.org/html/2607.20578#S4.E6)\)\. ∎
### 4\.3A two\-sided anisotropic estimate
The functionalUG\(S\)U\_\{G\}\(S\)is the standard weighted\-ℓ1\\ell\_\{1\}distance\-to\-subdifferential upper estimate and is not new in functional form\. The additional step below is to exploit the fact that the descent cone and subdifferential depend only on the support and signs of the nonzero coordinates\. Optimizing the ALMT error term over their magnitudes yields an explicit additive error determined by the Fisher mass on the active support\.
###### Theorem 4\.6\(Two\-sided anisotropic estimate\)\.
Under the assumptions of Lemma[4\.4](https://arxiv.org/html/2607.20578#S4.Thmtheorem4),
UG\(S\)−2Tr\(G\)∑i∈Sγi≤δ\(G−1/2C\)≤UG\(S\)\.U\_\{G\}\(S\)\-2\\sqrt\{\\frac\{\\operatorname\{Tr\}\(G\)\}\{\\sum\_\{i\\in S\}\\gamma\_\{i\}\}\}\\leq\\delta\(G^\{\-1/2\}C\)\\leq U\_\{G\}\(S\)\.Consequently, if
UG\(S\)≥4Tr\(G\)∑i∈Sγi,U\_\{G\}\(S\)\\geq 4\\sqrt\{\\frac\{\\operatorname\{Tr\}\(G\)\}\{\\sum\_\{i\\in S\}\\gamma\_\{i\}\}\},then
12UG\(S\)≤δ\(G−1/2C\)≤UG\(S\)\.\\frac\{1\}\{2\}U\_\{G\}\(S\)\\leq\\delta\(G^\{\-1/2\}C\)\\leq U\_\{G\}\(S\)\.
###### Proof\.
The upper bound is Lemma[4\.4](https://arxiv.org/html/2607.20578#S4.Thmtheorem4)\. For the lower bound, the descent\-cone error estimate ofAmelunxenet al\.\([2014](https://arxiv.org/html/2607.20578#bib.bib36)\)gives
δ\(D\(fG,x\)\)≥infτ≥0𝔼dist2\(g,τ∂fG\(x\)\)−2supz∈∂fG\(x\)‖z‖2fG\(x/‖x‖2\)\.\\delta\(D\(f\_\{G\},x\)\)\\geq\\inf\_\{\\tau\\geq 0\}\\mathbb\{E\}\\operatorname\{dist\}^\{2\}\\bigl\(g,\\tau\\partial f\_\{G\}\(x\)\\bigr\)\-\\frac\{2\\sup\_\{z\\in\\partial f\_\{G\}\(x\)\}\\\|z\\\|\_\{2\}\}\{f\_\{G\}\(x/\\\|x\\\|\_\{2\}\)\}\.\(7\)
For the weightedℓ1\\ell\_\{1\}norm, bothD\(fG,x\)D\(f\_\{G\},x\)and∂fG\(x\)\\partial f\_\{G\}\(x\)depend only on the support and signs of the nonzero coordinates ofxx, not on their magnitudes\. We may therefore apply \([7](https://arxiv.org/html/2607.20578#S4.E7)\) to any unit vector with supportSSand the same sign pattern asx0x\_\{0\}\. For every such vector, the first term on the right\-hand side isUG\(S\)U\_\{G\}\(S\), while
supz∈∂fG\(x\)‖z‖2=Tr\(G\)\.\\sup\_\{z\\in\\partial f\_\{G\}\(x\)\}\\\|z\\\|\_\{2\}=\\sqrt\{\\operatorname\{Tr\}\(G\)\}\.
For any unit vector supported onSS, Cauchy–Schwarz gives
fG\(x\)=∑i∈Sγi\|xi\|≤\(∑i∈Sγi\)1/2\.f\_\{G\}\(x\)=\\sum\_\{i\\in S\}\\sqrt\{\\gamma\_\{i\}\}\|x\_\{i\}\|\\leq\\left\(\\sum\_\{i\\in S\}\\gamma\_\{i\}\\right\)^\{1/2\}\.Equality holds for
xi=sign\(xi⋆\)γi\(∑j∈Sγj\)1/2,i∈S,x\_\{i\}=\\frac\{\\operatorname\{sign\}\(x\_\{i\}^\{\\star\}\)\\sqrt\{\\gamma\_\{i\}\}\}\{\\left\(\\sum\_\{j\\in S\}\\gamma\_\{j\}\\right\)^\{1/2\}\},\\qquad i\\in S,withxj=0x\_\{j\}=0forj∉Sj\\notin S\. Substitution into \([7](https://arxiv.org/html/2607.20578#S4.E7)\) yields
δ\(G−1/2C\)≥UG\(S\)−2Tr\(G\)∑i∈Sγi\.\\delta\(G^\{\-1/2\}C\)\\geq U\_\{G\}\(S\)\-2\\sqrt\{\\frac\{\\operatorname\{Tr\}\(G\)\}\{\\sum\_\{i\\in S\}\\gamma\_\{i\}\}\}\.The factor\-two estimate follows from the stated condition\. ∎
###### Conjecture 4\.8\(Extreme\-sparsity regime\)\.
There exists a universal constantc\>0c\>0such that, for every diagonalG≻0G\\succ 0, every supportS⊂\[d\]S\\subset\[d\], and any nonzerox⋆x^\{\\star\}supported onSS,
δ\(G−1/2D\(∥⋅∥1,x⋆\)\)≥cUG\(S\)\.\\delta\\\!\\left\(G^\{\-1/2\}D\(\\\|\\cdot\\\|\_\{1\},x^\{\\star\}\)\\right\)\\geq c\\,U\_\{G\}\(S\)\.Together with Lemma[4\.4](https://arxiv.org/html/2607.20578#S4.Thmtheorem4), this would give
δ\(G−1/2D\(∥⋅∥1,x⋆\)\)≍UG\(S\)\\delta\\\!\\left\(G^\{\-1/2\}D\(\\\|\\cdot\\\|\_\{1\},x^\{\\star\}\)\\right\)\\asymp U\_\{G\}\(S\)without the regime condition of Theorem[4\.6](https://arxiv.org/html/2607.20578#S4.Thmtheorem6)\.
Section[6\.1](https://arxiv.org/html/2607.20578#S6.SS1)reports one sparse configuration for which the additive lower bound is vacuous and comparesUG\(S\)U\_\{G\}\(S\)with the observed transition\. The experiment is consistent with the conjectured comparison in this configuration, but does not address its uniform validity\.
### 4\.4Support\-dependent upper recovery functional
The upper functionalUG\(S\)U\_\{G\}\(S\)depends on the location of the support as well as its cardinality\. We now derive a coordinatewise decomposition that induces a natural ordering of supports according to their Fisher profile\.
###### Proposition 4\.9\(Support\-dependent Fisher recovery complexity\)\.
Let
G=diag\(γ1,…,γd\),γi\>0\.G=\\operatorname\{diag\}\(\\gamma\_\{1\},\\ldots,\\gamma\_\{d\}\),\\qquad\\gamma\_\{i\}\>0\.Forτ≥0\\tau\\geq 0, define
qi\(τ\):=1\+τ2γi−𝔼\(\|g\|−τγi\)\+2,q\_\{i\}\(\\tau\):=1\+\\tau^\{2\}\\gamma\_\{i\}\-\\mathbb\{E\}\\bigl\(\|g\|\-\\tau\\sqrt\{\\gamma\_\{i\}\}\\bigr\)\_\{\+\}^\{2\},and
B\(τ\):=∑j=1d𝔼\(\|g\|−τγj\)\+2,g∼N\(0,1\)\.B\(\\tau\):=\\sum\_\{j=1\}^\{d\}\\mathbb\{E\}\\bigl\(\|g\|\-\\tau\\sqrt\{\\gamma\_\{j\}\}\\bigr\)\_\{\+\}^\{2\},\\qquad g\\sim N\(0,1\)\.Then:
1. \(i\)\(Support\-cost representation\)For everyS⊂\[d\]S\\subset\[d\], UG\(S\)=infτ≥0\[B\(τ\)\+∑i∈Sqi\(τ\)\]\.U\_\{G\}\(S\)=\\inf\_\{\\tau\\geq 0\}\\left\[B\(\\tau\)\+\\sum\_\{i\\in S\}q\_\{i\}\(\\tau\)\\right\]\.
2. \(ii\)\(Nonnegativity and spectral monotonicity\)For every fixedτ≥0\\tau\\geq 0,qi\(τ\)≥0q\_\{i\}\(\\tau\)\\geq 0, andqi\(τ\)q\_\{i\}\(\\tau\)is nondecreasing as a function ofγi\\gamma\_\{i\}\.
3. \(iii\)\(Nested\-support monotonicity\)IfS2⊆S1S\_\{2\}\\subseteq S\_\{1\}, then UG\(S2\)≤UG\(S1\)\.U\_\{G\}\(S\_\{2\}\)\\leq U\_\{G\}\(S\_\{1\}\)\.
4. \(iv\)\(Support comparison\)IfS1,S2⊂\[d\]S\_\{1\},S\_\{2\}\\subset\[d\]satisfy ∑i∈S2qi\(τ\)≥∑i∈S1qi\(τ\)for everyτ≥0,\\sum\_\{i\\in S\_\{2\}\}q\_\{i\}\(\\tau\)\\geq\\sum\_\{i\\in S\_\{1\}\}q\_\{i\}\(\\tau\)\\qquad\\text\{for every \}\\tau\\geq 0,then UG\(S2\)≥UG\(S1\)\.U\_\{G\}\(S\_\{2\}\)\\geq U\_\{G\}\(S\_\{1\}\)\.
###### Proof\.
ForS⊂\[d\]S\\subset\[d\], let
FS\(τ\):=∑i∈S\(1\+τ2γi\)\+∑j∉S𝔼\(\|g\|−τγj\)\+2\.F\_\{S\}\(\\tau\):=\\sum\_\{i\\in S\}\(1\+\\tau^\{2\}\\gamma\_\{i\}\)\+\\sum\_\{j\\notin S\}\\mathbb\{E\}\\bigl\(\|g\|\-\\tau\\sqrt\{\\gamma\_\{j\}\}\\bigr\)\_\{\+\}^\{2\}\.Adding and subtracting the off\-support contribution for eachi∈Si\\in Sgives
FS\(τ\)=B\(τ\)\+∑i∈Sqi\(τ\),F\_\{S\}\(\\tau\)=B\(\\tau\)\+\\sum\_\{i\\in S\}q\_\{i\}\(\\tau\),and taking the infimum overτ≥0\\tau\\geq 0proves \(i\)\.
Since\(\|g\|−a\)\+2≤g2\(\|g\|\-a\)\_\{\+\}^\{2\}\\leq g^\{2\}for everya≥0a\\geq 0,
qi\(τ\)≥τ2γi≥0\.q\_\{i\}\(\\tau\)\\geq\\tau^\{2\}\\gamma\_\{i\}\\geq 0\.For fixedτ\\tau, the term1\+τ2γi1\+\\tau^\{2\}\\gamma\_\{i\}is nondecreasing inγi\\gamma\_\{i\}, whereas𝔼\(\|g\|−τγi\)\+2\\mathbb\{E\}\(\|g\|\-\\tau\\sqrt\{\\gamma\_\{i\}\}\)\_\{\+\}^\{2\}is nonincreasing\. Thusqi\(τ\)q\_\{i\}\(\\tau\)is nondecreasing inγi\\gamma\_\{i\}, proving \(ii\)\. Parts \(iii\) and \(iv\) follow from the corresponding pointwise ordering ofFS\(τ\)F\_\{S\}\(\\tau\)and taking infima overτ≥0\\tau\\geq 0\. ∎
###### Corollary 4\.10\(Curvature\-increasing support swaps\)\.
LetS1,S2⊂\[d\]S\_\{1\},S\_\{2\}\\subset\[d\]have the same cardinality\. Write
S2∖S1=\{i1,…,im\},S1∖S2=\{j1,…,jm\}\.S\_\{2\}\\setminus S\_\{1\}=\\\{i\_\{1\},\\ldots,i\_\{m\}\\\},\\qquad S\_\{1\}\\setminus S\_\{2\}=\\\{j\_\{1\},\\ldots,j\_\{m\}\\\}\.If the indices can be ordered so that
γiℓ≥γjℓfor everyℓ=1,…,m,\\gamma\_\{i\_\{\\ell\}\}\\geq\\gamma\_\{j\_\{\\ell\}\}\\qquad\\text\{for every \}\\ell=1,\\ldots,m,then
UG\(S2\)≥UG\(S1\)\.U\_\{G\}\(S\_\{2\}\)\\geq U\_\{G\}\(S\_\{1\}\)\.In particular, the conclusion holds if
mini∈S2∖S1γi≥maxj∈S1∖S2γj\.\\min\_\{i\\in S\_\{2\}\\setminus S\_\{1\}\}\\gamma\_\{i\}\\geq\\max\_\{j\\in S\_\{1\}\\setminus S\_\{2\}\}\\gamma\_\{j\}\.
###### Proof\.
By Proposition[4\.9](https://arxiv.org/html/2607.20578#S4.Thmtheorem9)\(ii\),
qiℓ\(τ\)≥qjℓ\(τ\)for everyℓandτ≥0\.q\_\{i\_\{\\ell\}\}\(\\tau\)\\geq q\_\{j\_\{\\ell\}\}\(\\tau\)\\qquad\\text\{for every \}\\ell\\text\{ and \}\\tau\\geq 0\.Summing over the exchanged coordinates and applying Proposition[4\.9](https://arxiv.org/html/2607.20578#S4.Thmtheorem9)\(iv\) gives the result\. ∎
## 5Primal–Inverse Width Inequalities
The Fisher and inverse\-Fisher widths are opposite linear deformations of a common coordinate set\. The main result of this section is the sharp inequality
wG\(T\)wG−1\(T\)≥w\(T\)2,w\_\{G\}\(T\)w\_\{G^\{\-1\}\}\(T\)\\geq w\(T\)^\{2\},obtained from a log\-convexity property for commuting positive\-definite matrices\. We then apply the inequality to a common localized cone and derive a noncommutative geometric\-mean extension\.
### 5\.1Commuting log\-convexity
###### Theorem 5\.1\(Log\-convexity for commuting Fisher metrics\)\.
LetT⊂ℝdT\\subset\\mathbb\{R\}^\{d\}be nonempty and compact, and letG0,G1≻0G\_\{0\},G\_\{1\}\\succ 0commute\. Forθ∈\[0,1\]\\theta\\in\[0,1\], define
Gθ:=G01−θG1θ\.G\_\{\\theta\}:=G\_\{0\}^\{1\-\\theta\}G\_\{1\}^\{\\theta\}\.Then
wGθ\(T\)≤wG0\(T\)1−θwG1\(T\)θ\.w\_\{G\_\{\\theta\}\}\(T\)\\leq w\_\{G\_\{0\}\}\(T\)^\{1\-\\theta\}w\_\{G\_\{1\}\}\(T\)^\{\\theta\}\.\(8\)Consequently,
wG0\(T\)wG1\(T\)≥wG0\#G1\(T\)2,w\_\{G\_\{0\}\}\(T\)w\_\{G\_\{1\}\}\(T\)\\geq w\_\{G\_\{0\}\\\#G\_\{1\}\}\(T\)^\{2\},\(9\)whereG0\#G1G\_\{0\}\\\#G\_\{1\}is the affine\-invariant geometric mean\.
###### Proof\.
Gaussian width is translation invariant:
w\(T\+a\)=w\(T\),w\(T\+a\)=w\(T\),since𝔼⟨g,a⟩=0\\mathbb\{E\}\\langle g,a\\rangle=0\. Thus the comparison below depends only on Gaussian increments\.
IfTTis a singleton, all widths vanish\. Otherwise,wH\(T\)\>0w\_\{H\}\(T\)\>0for everyH≻0H\\succ 0, so the optimization below is well\-defined\. Fix0<θ<10<\\theta<1, since the endpoint cases are immediate, and set
H0:=G01/2,H1:=G11/2,Hθ:=Gθ1/2\.H\_\{0\}:=G\_\{0\}^\{1/2\},\\qquad H\_\{1\}:=G\_\{1\}^\{1/2\},\\qquad H\_\{\\theta\}:=G\_\{\\theta\}^\{1/2\}\.BecauseG0G\_\{0\}andG1G\_\{1\}commute, they are simultaneously diagonalizable, and in their common eigenbasis
Hθ=H01−θH1θ\.H\_\{\\theta\}=H\_\{0\}^\{1\-\\theta\}H\_\{1\}^\{\\theta\}\.
Forr\>0r\>0, define
Lr:=\(1−θ\)rH0\+θr−\(1−θ\)/θH1\.L\_\{r\}:=\(1\-\\theta\)rH\_\{0\}\+\\theta r^\{\-\(1\-\\theta\)/\\theta\}H\_\{1\}\.The scalar weighted arithmetic–geometric mean inequality, applied coordinatewise in the common eigenbasis, gives
Lr⪰Hθ\.L\_\{r\}\\succeq H\_\{\\theta\}\.SinceLrL\_\{r\}andHθH\_\{\\theta\}commute and are simultaneously diagonalizable with positive eigenvalues,
Lr2⪰Hθ2\.L\_\{r\}^\{2\}\\succeq H\_\{\\theta\}^\{2\}\.Hence, for allu,v∈Tu,v\\in T,
‖Lr\(u−v\)‖22≥‖Hθ\(u−v\)‖22\.\\\|L\_\{r\}\(u\-v\)\\\|\_\{2\}^\{2\}\\geq\\\|H\_\{\\theta\}\(u\-v\)\\\|\_\{2\}^\{2\}\.
Let
Xv:=⟨g,Hθv⟩,Yv:=⟨g,Lrv⟩,v∈T\.X\_\{v\}:=\\langle g,H\_\{\\theta\}v\\rangle,\\qquad Y\_\{v\}:=\\langle g,L\_\{r\}v\\rangle,\\qquad v\\in T\.The increment comparison and the Sudakov–Fernique theorem imply
wGθ\(T\)=𝔼supv∈TXv≤𝔼supv∈TYv\.w\_\{G\_\{\\theta\}\}\(T\)=\\mathbb\{E\}\\sup\_\{v\\in T\}X\_\{v\}\\leq\\mathbb\{E\}\\sup\_\{v\\in T\}Y\_\{v\}\.By subadditivity of the supremum,
𝔼supv∈TYv\\displaystyle\\mathbb\{E\}\\sup\_\{v\\in T\}Y\_\{v\}≤\(1−θ\)rwG0\(T\)\+θr−\(1−θ\)/θwG1\(T\)\.\\displaystyle\\leq\(1\-\\theta\)r\\,w\_\{G\_\{0\}\}\(T\)\+\\theta r^\{\-\(1\-\\theta\)/\\theta\}w\_\{G\_\{1\}\}\(T\)\.Optimizing overr\>0r\>0, with
r=\(wG1\(T\)wG0\(T\)\)θ,r=\\left\(\\frac\{w\_\{G\_\{1\}\}\(T\)\}\{w\_\{G\_\{0\}\}\(T\)\}\\right\)^\{\\theta\},gives \([8](https://arxiv.org/html/2607.20578#S5.E8)\)\.
Takingθ=1/2\\theta=1/2yields
wG1/2\(T\)2≤wG0\(T\)wG1\(T\)\.w\_\{G\_\{1/2\}\}\(T\)^\{2\}\\leq w\_\{G\_\{0\}\}\(T\)w\_\{G\_\{1\}\}\(T\)\.For commuting matrices,
G1/2=G01/2G11/2=G0\#G1,G\_\{1/2\}=G\_\{0\}^\{1/2\}G\_\{1\}^\{1/2\}=G\_\{0\}\\\#G\_\{1\},which proves \([9](https://arxiv.org/html/2607.20578#S5.E9)\)\. ∎
###### Corollary 5\.2\(Log\-convexity along the power geodesic\)\.
LetT⊂ℝdT\\subset\\mathbb\{R\}^\{d\}be nonempty and compact, and letG≻0G\\succ 0\. Then
F\(α\):=wGα\(T\),α∈ℝ,F\(\\alpha\):=w\_\{G^\{\\alpha\}\}\(T\),\\qquad\\alpha\\in\\mathbb\{R\},is log\-convex\. In particular, for everyα0,α1∈ℝ\\alpha\_\{0\},\\alpha\_\{1\}\\in\\mathbb\{R\}andθ∈\[0,1\]\\theta\\in\[0,1\],
F\(\(1−θ\)α0\+θα1\)≤F\(α0\)1−θF\(α1\)θ,F\\bigl\(\(1\-\\theta\)\\alpha\_\{0\}\+\\theta\\alpha\_\{1\}\\bigr\)\\leq F\(\\alpha\_\{0\}\)^\{1\-\\theta\}F\(\\alpha\_\{1\}\)^\{\\theta\},and, for everyα∈ℝ\\alpha\\in\\mathbb\{R\},
wGα\(T\)wG−α\(T\)≥w\(T\)2\.w\_\{G^\{\\alpha\}\}\(T\)w\_\{G^\{\-\\alpha\}\}\(T\)\\geq w\(T\)^\{2\}\.\(10\)
###### Proof\.
Apply Theorem[5\.1](https://arxiv.org/html/2607.20578#S5.Thmtheorem1)withG0=Gα0G\_\{0\}=G^\{\\alpha\_\{0\}\}andG1=Gα1G\_\{1\}=G^\{\\alpha\_\{1\}\}, which commute\. Takingα0=α\\alpha\_\{0\}=\\alpha,α1=−α\\alpha\_\{1\}=\-\\alpha, andθ=1/2\\theta=1/2gives \([10](https://arxiv.org/html/2607.20578#S5.E10)\)\. ∎
### 5\.2The sharp primal–inverse product inequality
###### Theorem 5\.3\(Sharp primal–inverse width product inequality\)\.
LetT⊂ℝdT\\subset\\mathbb\{R\}^\{d\}be nonempty and compact, and letG≻0G\\succ 0\. Then
wG\(T\)wG−1\(T\)≥w\(T\)2\.w\_\{G\}\(T\)w\_\{G^\{\-1\}\}\(T\)\\geq w\(T\)^\{2\}\.\(11\)Equivalently,
infs\>0\{swG\(T\)\+s−1wG−1\(T\)\}≥2w\(T\)\.\\inf\_\{s\>0\}\\left\\\{s\\,w\_\{G\}\(T\)\+s^\{\-1\}w\_\{G^\{\-1\}\}\(T\)\\right\\\}\\geq 2w\(T\)\.\(12\)
###### Proof\.
The product inequality is Corollary[5\.2](https://arxiv.org/html/2607.20578#S5.Thmtheorem2)withα=1\\alpha=1\. Fors\>0s\>0, the arithmetic–geometric mean inequality gives
swG\(T\)\+s−1wG−1\(T\)≥2wG\(T\)wG−1\(T\)≥2w\(T\)\.s\\,w\_\{G\}\(T\)\+s^\{\-1\}w\_\{G^\{\-1\}\}\(T\)\\geq 2\\sqrt\{w\_\{G\}\(T\)w\_\{G^\{\-1\}\}\(T\)\}\\geq 2w\(T\)\.Conversely,
infs\>0\{swG\(T\)\+s−1wG−1\(T\)\}=2wG\(T\)wG−1\(T\),\\inf\_\{s\>0\}\\left\\\{s\\,w\_\{G\}\(T\)\+s^\{\-1\}w\_\{G^\{\-1\}\}\(T\)\\right\\\}=2\\sqrt\{w\_\{G\}\(T\)w\_\{G^\{\-1\}\}\(T\)\},so the product and linear forms are equivalent\. ∎
### 5\.3Localized primal–inverse trade\-offs
LetC⊂ℝdC\\subset\\mathbb\{R\}^\{d\}be a nonzero closed convex cone\. Since an unbounded cone has infinite Gaussian width, we use the common localization
TC:=C∩B2d\.T\_\{C\}:=C\\cap B\_\{2\}^\{d\}\.Set
and define the restricted radial distortions
rC:=minv∈C∩𝕊d−1∥G−1/2v∥2,RC:=maxv∈C∩𝕊d−1∥G−1/2v∥2\.r\_\{C\}:=\\min\_\{v\\in C\\cap\\mathbb\{S\}^\{d\-1\}\}\\\|G^\{\-1/2\}v\\\|\_\{2\},\\qquad R\_\{C\}:=\\max\_\{v\\in C\\cap\\mathbb\{S\}^\{d\-1\}\}\\\|G^\{\-1/2\}v\\\|\_\{2\}\.Then0<rC≤RC<∞0<r\_\{C\}\\leq R\_\{C\}<\\infty\.
###### Proposition 5\.5\(Restricted distortion on a cone\)\.
With the notation above,
rCw\(D∩B2d\)≤wG−1\(TC\)≤RCw\(D∩B2d\)\.r\_\{C\}\\,w\(D\\cap B\_\{2\}^\{d\}\)\\leq w\_\{G^\{\-1\}\}\(T\_\{C\}\)\\leq R\_\{C\}\\,w\(D\\cap B\_\{2\}^\{d\}\)\.\(13\)Consequently,
rC\(δ\(D\)−1\)\+≤wG−1\(TC\)≤RCδ\(D\)\.r\_\{C\}\\sqrt\{\(\\delta\(D\)\-1\)\_\{\+\}\}\\leq w\_\{G^\{\-1\}\}\(T\_\{C\}\)\\leq R\_\{C\}\\sqrt\{\\delta\(D\)\}\.\(14\)
###### Proof\.
Set
KG:=G−1/2TC\.K\_\{G\}:=G^\{\-1/2\}T\_\{C\}\.Along each ray ofDD, the radial extent ofKGK\_\{G\}lies in\[rC,RC\]\[r\_\{C\},R\_\{C\}\]\. Hence
rC\(D∩B2d\)⊆KG⊆RC\(D∩B2d\)\.r\_\{C\}\(D\\cap B\_\{2\}^\{d\}\)\\subseteq K\_\{G\}\\subseteq R\_\{C\}\(D\\cap B\_\{2\}^\{d\}\)\.Monotonicity and homogeneity of Gaussian width give \([13](https://arxiv.org/html/2607.20578#S5.E13)\)\.
For a closed convex coneDD,
w\(D∩B2d\)=𝔼‖ΠDg‖2,δ\(D\)=𝔼‖ΠDg‖22\.w\(D\\cap B\_\{2\}^\{d\}\)=\\mathbb\{E\}\\\|\\Pi\_\{D\}g\\\|\_\{2\},\\qquad\\delta\(D\)=\\mathbb\{E\}\\\|\\Pi\_\{D\}g\\\|\_\{2\}^\{2\}\.Jensen’s inequality and the Gaussian Poincaré inequality yield
δ\(D\)−1≤w\(D∩B2d\)2≤δ\(D\)\.\\delta\(D\)\-1\\leq w\(D\\cap B\_\{2\}^\{d\}\)^\{2\}\\leq\\delta\(D\)\.Combining this with \([13](https://arxiv.org/html/2607.20578#S5.E13)\) proves \([14](https://arxiv.org/html/2607.20578#S5.E14)\)\. ∎
###### Corollary 5\.6\(Localized primal–inverse trade\-off\)\.
LetC⊂ℝdC\\subset\\mathbb\{R\}^\{d\}be a nonzero closed convex cone\. Then
wG\(TC\)δ\(G−1/2C\)≥w\(TC\)2RC\.w\_\{G\}\(T\_\{C\}\)\\sqrt\{\\delta\(G^\{\-1/2\}C\)\}\\geq\\frac\{w\(T\_\{C\}\)^\{2\}\}\{R\_\{C\}\}\.\(15\)In particular,
wG\(TC\)δ\(G−1/2C\)≥\(δ\(C\)−1\)\+RC\.w\_\{G\}\(T\_\{C\}\)\\sqrt\{\\delta\(G^\{\-1/2\}C\)\}\\geq\\frac\{\(\\delta\(C\)\-1\)\_\{\+\}\}\{R\_\{C\}\}\.\(16\)
###### Proof\.
Applying Theorem[5\.3](https://arxiv.org/html/2607.20578#S5.Thmtheorem3)toTCT\_\{C\}gives
wG\(TC\)wG−1\(TC\)≥w\(TC\)2\.w\_\{G\}\(T\_\{C\}\)w\_\{G^\{\-1\}\}\(T\_\{C\}\)\\geq w\(T\_\{C\}\)^\{2\}\.By Proposition[5\.5](https://arxiv.org/html/2607.20578#S5.Thmtheorem5),
wG−1\(TC\)≤RCδ\(G−1/2C\)\.w\_\{G^\{\-1\}\}\(T\_\{C\}\)\\leq R\_\{C\}\\sqrt\{\\delta\(G^\{\-1/2\}C\)\}\.This proves \([15](https://arxiv.org/html/2607.20578#S5.E15)\)\. Applying
w\(C∩B2d\)2≥\(δ\(C\)−1\)\+w\(C\\cap B\_\{2\}^\{d\}\)^\{2\}\\geq\(\\delta\(C\)\-1\)\_\{\+\}gives \([16](https://arxiv.org/html/2607.20578#S5.E16)\)\. ∎
### 5\.4Noncommuting metrics
For arbitraryG1,G2≻0G\_\{1\},G\_\{2\}\\succ 0, define their affine\-invariant geometric mean by
G1\#G2:=G11/2\(G1−1/2G2G1−1/2\)1/2G11/2\.G\_\{1\}\\\#G\_\{2\}:=G\_\{1\}^\{1/2\}\\left\(G\_\{1\}^\{\-1/2\}G\_\{2\}G\_\{1\}^\{\-1/2\}\\right\)^\{1/2\}G\_\{1\}^\{1/2\}\.
###### Lemma 5\.8\(Matrix arithmetic–geometric mean\)\.
LetG1,G2≻0G\_\{1\},G\_\{2\}\\succ 0\. Then:
1. \(i\)Fora,b\>0a,b\>0, \(aG1\)\#\(bG2\)=ab\(G1\#G2\)\.\(aG\_\{1\}\)\\\#\(bG\_\{2\}\)=\\sqrt\{ab\}\\,\(G\_\{1\}\\\#G\_\{2\}\)\.
2. \(ii\)For everys\>0s\>0, s2G1\+s−2G2⪰2\(G1\#G2\)\.s^\{2\}G\_\{1\}\+s^\{\-2\}G\_\{2\}\\succeq 2\(G\_\{1\}\\\#G\_\{2\}\)\.
###### Proof\.
Part \(i\) follows directly from the definition\. For part \(ii\), setC:=G1−1/2G2G1−1/2C:=G\_\{1\}^\{\-1/2\}G\_\{2\}G\_\{1\}^\{\-1/2\}\. Then
G1\+G2−2\(G1\#G2\)=G11/2\(I−C1/2\)2G11/2⪰0\.G\_\{1\}\+G\_\{2\}\-2\(G\_\{1\}\\\#G\_\{2\}\)=G\_\{1\}^\{1/2\}\(I\-C^\{1/2\}\)^\{2\}G\_\{1\}^\{1/2\}\\succeq 0\.Apply this inequality to\(s2G1,s−2G2\)\(s^\{2\}G\_\{1\},s^\{\-2\}G\_\{2\}\)and use part \(i\)\. ∎
###### Theorem 5\.9\(Noncommutative geometric\-mean bound\)\.
LetT⊂ℝdT\\subset\\mathbb\{R\}^\{d\}be nonempty and compact, and letG1,G2≻0G\_\{1\},G\_\{2\}\\succ 0\. Then
wG1\(T\)wG2\(T\)≥12wG1\#G2\(T\)2\.w\_\{G\_\{1\}\}\(T\)w\_\{G\_\{2\}\}\(T\)\\geq\\frac\{1\}\{2\}w\_\{G\_\{1\}\\\#G\_\{2\}\}\(T\)^\{2\}\.\(17\)Equivalently,
infs\>0\{swG1\(T\)\+s−1wG2\(T\)\}≥2wG1\#G2\(T\)\.\\inf\_\{s\>0\}\\left\\\{s\\,w\_\{G\_\{1\}\}\(T\)\+s^\{\-1\}w\_\{G\_\{2\}\}\(T\)\\right\\\}\\geq\\sqrt\{2\}\\,w\_\{G\_\{1\}\\\#G\_\{2\}\}\(T\)\.\(18\)
###### Proof\.
IfTTis a singleton, the claim is immediate\. Otherwise, letg1,g2,g∼N\(0,Id\)g\_\{1\},g\_\{2\},g\\sim N\(0,I\_\{d\}\)be independent and, fors\>0s\>0, define
Zv\(s\):=s⟨g1,G11/2v⟩\+s−1⟨g2,G21/2v⟩,Z\_\{v\}^\{\(s\)\}:=s\\langle g\_\{1\},G\_\{1\}^\{1/2\}v\\rangle\+s^\{\-1\}\\langle g\_\{2\},G\_\{2\}^\{1/2\}v\\rangle,and
Wv:=2⟨g,\(G1\#G2\)1/2v⟩\.W\_\{v\}:=\\sqrt\{2\}\\,\\langle g,\(G\_\{1\}\\\#G\_\{2\}\)^\{1/2\}v\\rangle\.Forh=u−vh=u\-v, Lemma[5\.8](https://arxiv.org/html/2607.20578#S5.Thmtheorem8)gives
𝔼\|Zu\(s\)−Zv\(s\)\|2\\displaystyle\\mathbb\{E\}\|Z\_\{u\}^\{\(s\)\}\-Z\_\{v\}^\{\(s\)\}\|^\{2\}=h⊤\(s2G1\+s−2G2\)h\\displaystyle=h^\{\\top\}\(s^\{2\}G\_\{1\}\+s^\{\-2\}G\_\{2\}\)h≥2h⊤\(G1\#G2\)h\\displaystyle\\geq 2h^\{\\top\}\(G\_\{1\}\\\#G\_\{2\}\)h=𝔼\|Wu−Wv\|2\.\\displaystyle=\\mathbb\{E\}\|W\_\{u\}\-W\_\{v\}\|^\{2\}\.Sudakov–Fernique therefore yields
𝔼supv∈TZv\(s\)≥2wG1\#G2\(T\)\.\\mathbb\{E\}\\sup\_\{v\\in T\}Z\_\{v\}^\{\(s\)\}\\geq\\sqrt\{2\}\\,w\_\{G\_\{1\}\\\#G\_\{2\}\}\(T\)\.On the other hand,
𝔼supv∈TZv\(s\)≤swG1\(T\)\+s−1wG2\(T\)\.\\mathbb\{E\}\\sup\_\{v\\in T\}Z\_\{v\}^\{\(s\)\}\\leq s\\,w\_\{G\_\{1\}\}\(T\)\+s^\{\-1\}w\_\{G\_\{2\}\}\(T\)\.This proves \([18](https://arxiv.org/html/2607.20578#S5.E18)\); optimizing overs\>0s\>0gives \([17](https://arxiv.org/html/2607.20578#S5.E17)\)\. ∎
## 6Numerical Experiments
We report two controlled recovery experiments and a separate illustration of primal–inverse width redistribution\. The first experiment compares the empirical transition of ordinary basis pursuit with the support\-dependent functionalUG\(S\)U\_\{G\}\(S\)\. The second examines the effect of deterministic weighting and finite\-sample column normalization\. The final experiment visualizes the redistribution of Gaussian width under the deformationsG1/2G^\{1/2\}andG−1/2G^\{\-1/2\}\. These experiments are intended as controlled illustrations of the theory rather than as a broad empirical study of sparse\-recovery phase transitions\.
### 6\.1Support\-dependent anisotropic recovery
#### Setup\.
We fixed the ambient dimension and sparsity at
d=256,k=16\.d=256,\\qquad k=16\.For each diagonal Fisher matrix
G=diag\(γ1,…,γd\),G=\\operatorname\{diag\}\(\\gamma\_\{1\},\\ldots,\\gamma\_\{d\}\),we fixed akk\-sparse vectorx⋆x^\{\\star\}with prescribed supportS=supp\(x⋆\)S=\\operatorname\{supp\}\(x^\{\\star\}\), drewA∈ℝm×dA\\in\\mathbb\{R\}^\{m\\times d\}with independentN\(0,1\)N\(0,1\)entries, and formed the noiseless observationsy=AG−1/2x⋆\.y=AG^\{\-1/2\}x^\{\\star\}\.The nonzero entries satisfyxi⋆∈\{−1,\+1\},i∈S,x\_\{i\}^\{\\star\}\\in\\\{\-1,\+1\\\},\\qquad i\\in S,and are drawn independently and uniformly once at the beginning of the experiment\. The resulting signalx⋆x^\{\\star\}is fixed across all trials; only the measurement matrixAAis resampled\.
Since the noiseless recovery problem is invariant under the common rescalingG↦cGG\\mapsto cG, each Fisher profile was normalized so that
Tr\(G\)=d\.\\operatorname\{Tr\}\(G\)=d\.
We recoveredx⋆x^\{\\star\}by ordinary basis pursuit,
x^∈argminx∈ℝd‖x‖1subject toAG−1/2x=y\.\\widehat\{x\}\\in\\arg\\min\_\{x\\in\\mathbb\{R\}^\{d\}\}\\\|x\\\|\_\{1\}\\quad\\text\{subject to\}\\quad AG^\{\-1/2\}x=y\.The optimization problems were solved in CVXPY using the CLARABEL solver\. Recovery was declared successful when
‖x^−x⋆‖2‖x⋆‖2<10−4\.\\frac\{\\\|\\widehat\{x\}\-x^\{\\star\}\\\|\_\{2\}\}\{\\\|x^\{\\star\}\\\|\_\{2\}\}<10^\{\-4\}\.
For each Fisher profile and each value ofmm, we ran200200independent trials\. Empirical recovery probabilities are reported with Wilson95%95\\%confidence intervals\. We definem^50\\widehat\{m\}\_\{50\}by linear interpolation between the two adjacent grid points whose empirical recovery probabilities bracket1/21/2\.
We considered five profiles:
isotropic,low\-support,high\-support,flat off\-support,mixed / one flat\.\\text\{isotropic\},\\qquad\\text\{low\-support\},\\qquad\\text\{high\-support\},\\qquad\\text\{flat off\-support\},\\qquad\\text\{mixed / one flat\}\.WithS=\{1,…,k\}S=\\\{1,\\ldots,k\\\}, the five diagonal profiles are defined, before the common trace normalization, as follows:
Each profile is subsequently rescaled by a common positive constant so thatTr\(G\)=d\\operatorname\{Tr\}\(G\)=d\. The profiles separate curvature on the active support from the contribution of inactive coordinates\.
#### Comparison withUG\(S\)U\_\{G\}\(S\)\.
For each profile, we evaluated the upper functionalUG\(S\)U\_\{G\}\(S\)and compared it with the interpolated empirical transitionm^50\\widehat\{m\}\_\{50\}\. The results are summarized in Table[1](https://arxiv.org/html/2607.20578#S6.T1)\.
Table 1:Theoretical functional and empirical transition for ordinary basis pursuit under the five Fisher profiles\.Across all five profiles,UG\(S\)U\_\{G\}\(S\)captures both the ordering and the numerical location of the observed transition\. The ratios satisfy
0\.984≤m^50UG\(S\)≤0\.999,0\.984\\leq\\frac\{\\widehat\{m\}\_\{50\}\}\{U\_\{G\}\(S\)\}\\leq 0\.999,so the discrepancy is below approximately2%2\\%in every tested configuration\.
The support dependence is substantial\. Moving from the low\-support to the high\-support profile increases the empirical transition from about2222to178178measurements, althoughdd,kk, and the decoder remain unchanged\. The high\-support configuration therefore requires more than eight times as many measurements as the low\-support configuration and nearly three times as many as the isotropic profile\. Thus sparsity alone does not determine the observed recovery scale; the location of the support in the Fisher spectrum is also decisive\.
The inactive coordinates also matter\. The flat off\-support profile has an empirical transition near44\.844\.8, compared with60\.160\.1in the isotropic case\. The mixed profile remains close to the isotropic transition, at approximately57\.057\.0\. These comparisons illustrate that the transition depends on the full weighted descent\-cone geometry, rather than on the cardinality of the support or a single extreme coordinate\.
Figure 1:Empirical recovery curves for ordinary basis pursuit under inverse\-Fisher measurements\. We used=256d=256,k=16k=16, and200200independent trials for each value ofmmand each Fisher profile\. Shaded bands are Wilson95%95\\%confidence intervals\. The dashed and dotted vertical lines markUG\(S\)U\_\{G\}\(S\)and the interpolated empirical transitionm^50\\widehat\{m\}\_\{50\}, respectively\. Across all five profiles,UG\(S\)U\_\{G\}\(S\)captures both the ordering and the numerical location of the observed transitions\.
### 6\.2Effect of decoder weighting and normalization
The preceding experiment concerns ordinary basis pursuit in the original coordinates\. We next examine how the transition changes when the decoder compensates for, or reinforces, the diagonal anisotropy\.
Let
We compared four decoders\.
#### Unweighted basis pursuit\.
The baseline decoder is
x^unw∈argminx‖x‖1subject toMx=y\.\\widehat\{x\}\_\{\\mathrm\{unw\}\}\\in\\arg\\min\_\{x\}\\\|x\\\|\_\{1\}\\quad\\text\{subject to\}\\quad Mx=y\.
#### Inverse\-square\-root weighting\.
The second decoder solves
x^inv∈argminx∑i=1dγi−1/2\|xi\|subject toMx=y\.\\widehat\{x\}\_\{\\mathrm\{inv\}\}\\in\\arg\\min\_\{x\}\\sum\_\{i=1\}^\{d\}\\gamma\_\{i\}^\{\-1/2\}\|x\_\{i\}\|\\quad\\text\{subject to\}\\quad Mx=y\.Under the change of variablesz=G−1/2xz=G^\{\-1/2\}x, this becomes ordinary basis pursuit for the isotropic systemAz=yAz=y\. It therefore compensates for the population\-level diagonal column scaling induced byG−1/2G^\{\-1/2\}\.
#### Square\-root Fisher weighting\.
The third decoder solves
x^F∈argminx∑i=1dγi1/2\|xi\|subject toMx=y\.\\widehat\{x\}\_\{\\mathrm\{F\}\}\\in\\arg\\min\_\{x\}\\sum\_\{i=1\}^\{d\}\\gamma\_\{i\}^\{1/2\}\|x\_\{i\}\|\\quad\\text\{subject to\}\\quad Mx=y\.This Fisher\-weighted heuristic penalizes high\-curvature coordinates more strongly\. It is included as a geometric comparison and is not claimed to be optimal for sparse recovery\.
#### Column\-normalized basis pursuit\.
For the fourth decoder, define
DM:=diag\(‖M1‖2,…,‖Md‖2\),M~:=MDM−1,D\_\{M\}:=\\operatorname\{diag\}\\bigl\(\\\|M\_\{1\}\\\|\_\{2\},\\ldots,\\\|M\_\{d\}\\\|\_\{2\}\\bigr\),\\qquad\\widetilde\{M\}:=MD\_\{M\}^\{\-1\},whereMjM\_\{j\}denotes thejj\-th column ofMM\. We solve
z^∈argminz‖z‖1subject toM~z=y,\\widehat\{z\}\\in\\arg\\min\_\{z\}\\\|z\\\|\_\{1\}\\quad\\text\{subject to\}\\quad\\widetilde\{M\}z=y,and transform back via
x^col:=DM−1z^\.\\widehat\{x\}\_\{\\mathrm\{col\}\}:=D\_\{M\}^\{\-1\}\\widehat\{z\}\.This decoder uses the realized finite\-sample column norms rather than the population scalesγi−1/2\\gamma\_\{i\}^\{\-1/2\}\.
We used the same dimensions, profiles, recovery criterion, and solver as in the preceding experiment\. For every decoder, Fisher profile, and value ofmm, we ran200200independent trials\. The dashed vertical line in each panel of Figure[2](https://arxiv.org/html/2607.20578#S6.F2)marksUG\(S\)U\_\{G\}\(S\), which is the theoretical functional for the unweighted decoder only\.
The interpolated empirical transitions are reported in Table[2](https://arxiv.org/html/2607.20578#S6.T2)\.
Table 2:Interpolated empirical transitions for the four decoders\.In the isotropic profile, all four transitions lie nearm=60m=60, as expected\. Under anisotropy, unweighted basis pursuit ranges from approximately2222measurements in the low\-support profile to approximately178178in the high\-support profile\.
Inverse\-square\-root weighting removes almost all profile dependence: its empirical transitions lie between60\.460\.4and61\.361\.3\. Finite\-sample column normalization has nearly the same effect, with transitions between59\.459\.4and60\.760\.7\. The close agreement between these two decoders indicates that the dominant profile dependence in this experiment is associated with the diagonal column scaling\.
This compensation is not uniformly beneficial\. The low\-support and flat off\-support profiles are favorable for the unweighted decoder\. Compensating for the anisotropy moves their transitions back toward the isotropic level and therefore increases the required number of measurements\. Conversely, in the high\-support profile, inverse\-square\-root weighting reduces the transition from approximately177\.6177\.6to60\.960\.9, while column normalization reduces it to approximately60\.760\.7\.
Square\-root Fisher weighting reinforces the profile dependence\. Its transition decreases to16\.516\.5in the low\-support profile and to33\.533\.5in the flat off\-support profile, but increases to approximately247\.1247\.1in the high\-support profile and71\.071\.0in the mixed profile\. Thus a geometrically natural Fisher weighting need not be uniformly favorable for sparse recovery\.
Figure 2:Comparison of four decoders under the five Fisher profiles, withd=256d=256,k=16k=16, and200200independent trials for each decoder and each value ofmm\. The dashed vertical line marksUG\(S\)U\_\{G\}\(S\), which applies to the unweighted decoder\. Inverse\-square\-root weighting and finite\-sample column normalization largely remove the profile dependence, whereas square\-root Fisher weighting reinforces it\.
### 6\.3Primal–inverse width redistribution
The final experiment illustrates how a metric deformation can redistribute Gaussian width between the primal and inverse geometries\. It is not a numerical verification of Theorem[5\.3](https://arxiv.org/html/2607.20578#S5.Thmtheorem3), which is an exact inequality\.
Letn=64n=64andr=20r=20\. We generatedV∈ℝn×rV\\in\\mathbb\{R\}^\{n\\times r\}by QR\-factorizing a standard Gaussian matrix and retained its orthonormal columns\. Independently, we drew
si∼Uniform\(0\.5,5\),i=1,…,r,s\_\{i\}\\sim\\operatorname\{Uniform\}\(0\.5,5\),\\qquad i=1,\\ldots,r,and set
G0:=Vdiag\(s\)V⊤\+0\.1In\.G\_\{0\}:=V\\operatorname\{diag\}\(s\)V^\{\\top\}\+0\.1I\_\{n\}\.We then defined
Gλ:=G0\+λIn\.G\_\{\\lambda\}:=G\_\{0\}\+\\lambda I\_\{n\}\.
Independently, we generatedU∈ℝn×qU\\in\\mathbb\{R\}^\{n\\times q\}, withq=10q=10, by applying the same QR procedure to a new standard Gaussian matrix, and considered
The matrixG0G\_\{0\}and subspace basisUUwere fixed across all values ofλ\\lambda, using random seed20260720\.
For eachλ\\lambda, we estimated
wGλ\(T\),wGλ−1\(T\),w\(T\),w\_\{G\_\{\\lambda\}\}\(T\),\\qquad w\_\{G\_\{\\lambda\}^\{\-1\}\}\(T\),\\qquad w\(T\),using10510^\{5\}Monte Carlo samples\. We also computed the normalized product ratio
ρλ:=wGλ\(T\)wGλ−1\(T\)w\(T\)2\.\\rho\_\{\\lambda\}:=\\frac\{w\_\{G\_\{\\lambda\}\}\(T\)w\_\{G\_\{\\lambda\}^\{\-1\}\}\(T\)\}\{w\(T\)^\{2\}\}\.
Asλ\\lambdaincreases, the primal width decreases from approximately8\.268\.26to0\.860\.86, while the inverse\-Fisher width increases from approximately2\.932\.93to11\.0911\.09\. Thus the two widths move in opposite directions under this regularization path\. At the same time, the product ratio decreases from approximately2\.542\.54toward equality:
ρ0≈2\.536,ρ12≈1\.003\.\\rho\_\{0\}\\approx 2\.536,\\qquad\\rho\_\{12\}\\approx 1\.003\.The minimum value over the tested grid is
minλρλ≈1\.0026,\\min\_\{\\lambda\}\\rho\_\{\\lambda\}\\approx 1\.0026,consistent with the exact inequality
wGλ\(T\)wGλ−1\(T\)≥w\(T\)2\.w\_\{G\_\{\\lambda\}\}\(T\)w\_\{G\_\{\\lambda\}^\{\-1\}\}\(T\)\\geq w\(T\)^\{2\}\.
The experiment illustrates width redistribution for one fixed matrix and one fixed subspace\. It does not imply monotonicity of either width for arbitrary sets or arbitrary matrix paths; such behavior depends on the alignment ofTTwith the eigenspaces ofGλG\_\{\\lambda\}\.
Figure 3:Primal–inverse width redistribution forGλ=G0\+λI64G\_\{\\lambda\}=G\_\{0\}\+\\lambda I\_\{64\}on the subspace ballT=UB210T=UB\_\{2\}^\{10\}\. Widths are estimated using10510^\{5\}Monte Carlo samples\. Along this path, the primal width decreases and the inverse\-Fisher width increases, while the normalized productρλ\\rho\_\{\\lambda\}approaches11from above\. The figure is an illustration of the redistribution mechanism, not a numerical proof of the product inequality\.
## Acknowledgments
The author acknowledges the use of ChatGPT and Claude in the preparation of this manuscript\. These tools were used to refine the language and organization of the draft, to brainstorm and explore proof strategies, and to assist in generating code for the numerical experiments\. All mathematical arguments were independently checked, and all source code was reviewed and debugged by the author\. The author takes full responsibility for the originality, correctness, and final content of the manuscript\.
## References
- S\. Amari and H\. Nagaoka \(2000\)Methods of information geometry\.Translations of Mathematical Monographs, Vol\.191,American Mathematical Society,Providence, RI\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p6.2)\.
- S\. Amari \(1998\)Natural gradient works efficiently in learning\.Neural Computation10\(2\),pp\. 251–276\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p6.2)\.
- D\. Amelunxen, M\. Lotz, M\. B\. McCoy, and J\. A\. Tropp \(2014\)Living on the edge: phase transitions in convex programs with random data\.Information and Inference: A Journal of the IMA3\(3\),pp\. 224–294\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p2.3),[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p4.1),[§4\.1](https://arxiv.org/html/2607.20578#S4.SS1.p3.2),[§4\.3](https://arxiv.org/html/2607.20578#S4.SS3.1.p1.1)\.
- S\. Boucheron, G\. Lugosi, and P\. Massart \(2013\)Concentration inequalities: a nonasymptotic theory of independence\.Oxford University Press,Oxford\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p1.1)\.
- E\. J\. Candès, J\. Romberg, and T\. Tao \(2006\)Robust uncertainty principles: exact signal reconstruction from highly incomplete frequency information\.IEEE Transactions on Information Theory52\(2\),pp\. 489–509\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p2.3)\.
- N\. N\. Čencov \(1982\)Statistical decision rules and optimal inference\.Translations of Mathematical Monographs, Vol\.53,American Mathematical Society,Providence, RI\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p6.2)\.
- V\. Chandrasekaran, B\. Recht, P\. A\. Parrilo, and A\. S\. Willsky \(2012\)The convex geometry of linear inverse problems\.Foundations of Computational Mathematics12\(6\),pp\. 805–849\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p2.3),[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p4.1)\.
- M\. Díaz, M\. Junca, F\. Rincón, and M\. Velasco \(2018\)Compressed sensing of data with a known distribution\.Applied and Computational Harmonic Analysis45\(3\),pp\. 486–504\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p4.1)\.
- D\. L\. Donoho and J\. Tanner \(2009\)Observed universality of phase transitions in high\-dimensional geometry, with implications for modern data analysis and signal processing\.Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences367\(1906\),pp\. 4273–4293\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p2.3)\.
- D\. L\. Donoho \(2006\)Compressed sensing\.IEEE Transactions on Information Theory52\(4\),pp\. 1289–1306\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p2.3)\.
- S\. Foucart and H\. Rauhut \(2013\)A mathematical introduction to compressive sensing\.Applied and Numerical Harmonic Analysis,Birkhäuser,New York\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p2.3)\.
- Y\. Gordon \(1988\)On Milman’s inequality and random subspaces which escape through a mesh inℝn\\mathbb\{R\}^\{n\}\.InGeometric Aspects of Functional Analysis,Lecture Notes in Mathematics, Vol\.1317,pp\. 84–106\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p2.3),[§4\.1](https://arxiv.org/html/2607.20578#S4.SS1.1.p1.1)\.
- M\. A\. Khajehnejad, W\. Xu, A\. S\. Avestimehr, and B\. Hassibi \(2011\)Analyzing weightedℓ1\\ell\_\{1\}minimization for sparse recovery with nonuniform sparse models\.IEEE Transactions on Signal Processing59\(5\),pp\. 1985–2001\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p4.1)\.
- R\. Kueng and D\. Gross \(2014\)RIPless compressed sensing from anisotropic measurements\.Linear Algebra and its Applications441,pp\. 110–123\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p3.1)\.
- F\. Kunstner, L\. Balles, and P\. Hennig \(2019\)Limitations of the empirical fisher approximation for natural gradient descent\.InAdvances in Neural Information Processing Systems 32,Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p7.1)\.
- V\. K\. Ky \(2026\)Fisher width: a geometric measure of complexity on statistical manifolds\.External Links:2606\.18306,[Link](https://arxiv.org/abs/2606.18306)Cited by:[§1\.1](https://arxiv.org/html/2607.20578#S1.SS1.p2.2),[§1\.1](https://arxiv.org/html/2607.20578#S1.SS1.p7.2),[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p1.1),[§2\.1](https://arxiv.org/html/2607.20578#S2.SS1.1.p1.6),[§3\.1](https://arxiv.org/html/2607.20578#S3.SS1.p1.2),[§3](https://arxiv.org/html/2607.20578#S3.p1.10)\.
- M\. Ledoux and M\. Talagrand \(1991\)Probability in banach spaces: isoperimetry and processes\.Ergebnisse der Mathematik und ihrer Grenzgebiete \(3\), Vol\.23,Springer,Berlin\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p1.1)\.
- J\. Martens and R\. Grosse \(2015\)Optimizing neural networks with kronecker\-factored approximate curvature\.InProceedings of the 32nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.37,pp\. 2408–2417\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p7.1)\.
- S\. N\. Negahban, P\. Ravikumar, M\. J\. Wainwright, and B\. Yu \(2012\)A unified framework for high\-dimensional analysis ofmm\-estimators with decomposable regularizers\.Statistical Science27\(4\),pp\. 538–557\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p4.1)\.
- R\. Pascanu and Y\. Bengio \(2014\)Revisiting natural gradient for deep networks\.InInternational Conference on Learning Representations,Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p6.2)\.
- Y\. Plan and R\. Vershynin \(2014\)Dimension reduction by random hyperplane tessellations\.Discrete & Computational Geometry51\(2\),pp\. 438–461\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p1.1)\.
- C\. R\. Rao \(1945\)Information and the accuracy attainable in the estimation of statistical parameters\.Bulletin of the Calcutta Mathematical Society37,pp\. 81–91\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p6.2)\.
- M\. Rudelson and S\. Zhou \(2013\)Reconstruction from anisotropic random measurements\.IEEE Transactions on Information Theory59\(6\),pp\. 3434–3447\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p3.1)\.
- M\. Talagrand \(2005\)The generic chaining: upper and lower bounds of stochastic processes\.Springer Monographs in Mathematics,Springer,Berlin, Heidelberg\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p1.1),[§3\.1](https://arxiv.org/html/2607.20578#S3.SS1.5.p3.4)\.
- R\. Tibshirani \(1996\)Regression shrinkage and selection via the lasso\.Journal of the Royal Statistical Society: Series B \(Methodological\)58\(1\),pp\. 267–288\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p4.1)\.
- R\. Vershynin \(2018\)High\-dimensional probability: an introduction with applications in data science\.Cambridge University Press\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p1.1),[§4\.1](https://arxiv.org/html/2607.20578#S4.SS1.1.p1.1)\.
- M\. J\. Wainwright \(2019\)High\-dimensional statistics: a non\-asymptotic viewpoint\.Cambridge University Press\.Cited by:[§1\.2](https://arxiv.org/html/2607.20578#S1.SS2.p1.1)\.
## Appendix AVerification of Fisher\-regularity conditions
We verify the conditions of Definition[3\.1](https://arxiv.org/html/2607.20578#S3.Thmtheorem1)for the model classes appearing in Corollary[3\.6](https://arxiv.org/html/2607.20578#S3.Thmtheorem6)\. Throughout,GGdenotes the Fisher matrix at the reference parameterθ0\\theta\_\{0\}, and all local Hessian bounds are required on the neighborhood
\{θ:‖θ−θ0‖G≤ρ\}\.\\\{\\theta:\\\|\\theta\-\\theta\_\{0\}\\\|\_\{G\}\\leq\\rho\\\}\.We write
ζ:=G−1/2∇θℓθ0\(Z\)\\zeta:=G^\{\-1/2\}\\nabla\_\{\\theta\}\\ell\_\{\\theta\_\{0\}\}\(Z\)for the whitened gradient\.
#### \(i\) Logistic regression\.
LetZ=\(X,Y\)Z=\(X,Y\), whereY∈\{0,1\}Y\\in\\\{0,1\\\}, and consider
ℓθ\(X,Y\)=−Y⟨θ,X⟩\+log\(1\+e⟨θ,X⟩\)\.\\ell\_\{\\theta\}\(X,Y\)=\-Y\\langle\\theta,X\\rangle\+\\log\\bigl\(1\+e^\{\\langle\\theta,X\\rangle\}\\bigr\)\.Writingσ\(t\)=\(1\+e−t\)−1\\sigma\(t\)=\(1\+e^\{\-t\}\)^\{\-1\}, we have
∇θℓθ\(X,Y\)=\(σ\(⟨θ,X⟩\)−Y\)X,∇θ2ℓθ\(X,Y\)=σ′\(⟨θ,X⟩\)XX⊤\.\\nabla\_\{\\theta\}\\ell\_\{\\theta\}\(X,Y\)=\\bigl\(\\sigma\(\\langle\\theta,X\\rangle\)\-Y\\bigr\)X,\\qquad\\nabla\_\{\\theta\}^\{2\}\\ell\_\{\\theta\}\(X,Y\)=\\sigma^\{\\prime\}\(\\langle\\theta,X\\rangle\)XX^\{\\top\}\.
Under correct specification,
𝔼∇θℓθ0\(Z\)=0,Cov\(∇θℓθ0\(Z\)\)=G,\\mathbb\{E\}\\nabla\_\{\\theta\}\\ell\_\{\\theta\_\{0\}\}\(Z\)=0,\\qquad\\operatorname\{Cov\}\\bigl\(\\nabla\_\{\\theta\}\\ell\_\{\\theta\_\{0\}\}\(Z\)\\bigr\)=G,so Condition[\(FR1\)](https://arxiv.org/html/2607.20578#S3.I1.i1)holds\.
Since0≤σ′\(t\)≤1/40\\leq\\sigma^\{\\prime\}\(t\)\\leq 1/4, theGG\-Cauchy–Schwarz inequality gives
\|h⊤∇θ2ℓθ\(X,Y\)h\|\\displaystyle\\bigl\|h^\{\\top\}\\nabla\_\{\\theta\}^\{2\}\\ell\_\{\\theta\}\(X,Y\)h\\bigr\|≤14\(h⊤X\)2\\displaystyle\\leq\\frac\{1\}\{4\}\(h^\{\\top\}X\)^\{2\}≤14\(X⊤G−1X\)‖h‖G2\.\\displaystyle\\leq\\frac\{1\}\{4\}\\bigl\(X^\{\\top\}G^\{\-1\}X\\bigr\)\\\|h\\\|\_\{G\}^\{2\}\.Thus Condition[\(FR2\)](https://arxiv.org/html/2607.20578#S3.I1.i2)holds with
M\(Z\):=14X⊤G−1X,σH:=14\[𝔼\(X⊤G−1X\)2\]1/2,M\(Z\):=\\frac\{1\}\{4\}X^\{\\top\}G^\{\-1\}X,\\qquad\\sigma\_\{H\}:=\\frac\{1\}\{4\}\\left\[\\mathbb\{E\}\\bigl\(X^\{\\top\}G^\{\-1\}X\\bigr\)^\{2\}\\right\]^\{1/2\},provided
𝔼\(X⊤G−1X\)2<∞\.\\mathbb\{E\}\\bigl\(X^\{\\top\}G^\{\-1\}X\\bigr\)^\{2\}<\\infty\.In particular, this condition follows from the bounded\-leverage assumptionX⊤G−1X≤ΛX^\{\\top\}G^\{\-1\}X\\leq\\Lambdaalmost surely\.
Atθ0\\theta\_\{0\},
ζ=\(σ\(⟨θ0,X⟩\)−Y\)G−1/2X\.\\zeta=\\bigl\(\\sigma\(\\langle\\theta\_\{0\},X\\rangle\)\-Y\\bigr\)G^\{\-1/2\}X\.Since\|σ\(⟨θ0,X⟩\)−Y\|≤1,\|\\sigma\(\\langle\\theta\_\{0\},X\\rangle\)\-Y\|\\leq 1,
‖ζ‖24≤\(X⊤G−1X\)2\.\\\|\\zeta\\\|\_\{2\}^\{4\}\\leq\\bigl\(X^\{\\top\}G^\{\-1\}X\\bigr\)^\{2\}\.Hence Condition[\(FR3\)](https://arxiv.org/html/2607.20578#S3.I1.i3)holds whenever
𝔼\(X⊤G−1X\)2≤κ4d2\.\\mathbb\{E\}\\bigl\(X^\{\\top\}G^\{\-1\}X\\bigr\)^\{2\}\\leq\\kappa^\{4\}d^\{2\}\.
#### \(ii\) Canonical\-link generalized linear models\.
Consider a canonical exponential\-family model with negative log\-likelihood
ℓθ\(X,Y\)=A\(ηθ\(X\)\)−⟨Y,ηθ\(X\)⟩,ηθ\(X\)=J\(X\)θ\+b\(X\)\.\\ell\_\{\\theta\}\(X,Y\)=A\(\\eta\_\{\\theta\}\(X\)\)\-\\langle Y,\\eta\_\{\\theta\}\(X\)\\rangle,\\qquad\\eta\_\{\\theta\}\(X\)=J\(X\)\\theta\+b\(X\)\.Then
∇θℓθ\(X,Y\)=J\(X\)⊤\(∇A\(ηθ\(X\)\)−Y\),\\nabla\_\{\\theta\}\\ell\_\{\\theta\}\(X,Y\)=J\(X\)^\{\\top\}\\bigl\(\\nabla A\(\\eta\_\{\\theta\}\(X\)\)\-Y\\bigr\),and
∇θ2ℓθ\(X,Y\)=J\(X\)⊤∇2A\(ηθ\(X\)\)J\(X\)\.\\nabla\_\{\\theta\}^\{2\}\\ell\_\{\\theta\}\(X,Y\)=J\(X\)^\{\\top\}\\nabla^\{2\}A\(\\eta\_\{\\theta\}\(X\)\)J\(X\)\.
Assume the following:
1. \(a\)throughout‖θ−θ0‖G≤ρ\\\|\\theta\-\\theta\_\{0\}\\\|\_\{G\}\\leq\\rho, ∇2A\(ηθ\(X\)\)⪯M0Ialmost surely;\\nabla^\{2\}A\(\\eta\_\{\\theta\}\(X\)\)\\preceq M\_\{0\}I\\qquad\\text\{almost surely\};
2. \(b\)𝔼‖G−1/2J\(X\)⊤‖op4<∞;\\mathbb\{E\}\\left\\\|G^\{\-1/2\}J\(X\)^\{\\top\}\\right\\\|\_\{\\mathrm\{op\}\}^\{4\}<\\infty;
3. \(c\)𝔼‖G−1/2J\(X\)⊤\(∇A\(ηθ0\(X\)\)−Y\)‖24≤κ4d2\.\\mathbb\{E\}\\left\\\|G^\{\-1/2\}J\(X\)^\{\\top\}\\bigl\(\\nabla A\(\\eta\_\{\\theta\_\{0\}\}\(X\)\)\-Y\\bigr\)\\right\\\|\_\{2\}^\{4\}\\leq\\kappa^\{4\}d^\{2\}\.
Under correct specification, the conditional gradient has mean zero and its covariance is the Fisher information matrix, so Condition[\(FR1\)](https://arxiv.org/html/2607.20578#S3.I1.i1)holds\. Moreover,
h⊤∇θ2ℓθ\(X,Y\)h\\displaystyle h^\{\\top\}\\nabla\_\{\\theta\}^\{2\}\\ell\_\{\\theta\}\(X,Y\)h≤M0‖J\(X\)h‖22\\displaystyle\\leq M\_\{0\}\\\|J\(X\)h\\\|\_\{2\}^\{2\}≤M0‖G−1/2J\(X\)⊤‖op2‖h‖G2\.\\displaystyle\\leq M\_\{0\}\\left\\\|G^\{\-1/2\}J\(X\)^\{\\top\}\\right\\\|\_\{\\mathrm\{op\}\}^\{2\}\\\|h\\\|\_\{G\}^\{2\}\.Thus Condition[\(FR2\)](https://arxiv.org/html/2607.20578#S3.I1.i2)holds with
M\(Z\):=M0‖G−1/2J\(X\)⊤‖op2\.M\(Z\):=M\_\{0\}\\left\\\|G^\{\-1/2\}J\(X\)^\{\\top\}\\right\\\|\_\{\\mathrm\{op\}\}^\{2\}\.Finally,
ζ=G−1/2J\(X\)⊤\(∇A\(ηθ0\(X\)\)−Y\),\\zeta=G^\{\-1/2\}J\(X\)^\{\\top\}\\bigl\(\\nabla A\(\\eta\_\{\\theta\_\{0\}\}\(X\)\)\-Y\\bigr\),so assumption \(c\) is exactly Condition[\(FR3\)](https://arxiv.org/html/2607.20578#S3.I1.i3)\. For models with unbounded responses, these assumptions require an appropriate conditional moment or tail bound\.
#### \(iii\) Gaussian linear regression\.
Let
Y=X⊤θ0\+ε,ε∼N\(0,σ2\),Y=X^\{\\top\}\\theta\_\{0\}\+\\varepsilon,\\qquad\\varepsilon\\sim N\(0,\\sigma^\{2\}\),whereε\\varepsilonis independent ofXX, and consider
ℓθ\(X,Y\)=12σ2\(Y−X⊤θ\)2\.\\ell\_\{\\theta\}\(X,Y\)=\\frac\{1\}\{2\\sigma^\{2\}\}\\bigl\(Y\-X^\{\\top\}\\theta\\bigr\)^\{2\}\.Then
∇θℓθ0\(X,Y\)=−εσ2X,∇θ2ℓθ\(X,Y\)=1σ2XX⊤\.\\nabla\_\{\\theta\}\\ell\_\{\\theta\_\{0\}\}\(X,Y\)=\-\\frac\{\\varepsilon\}\{\\sigma^\{2\}\}X,\\qquad\\nabla\_\{\\theta\}^\{2\}\\ell\_\{\\theta\}\(X,Y\)=\\frac\{1\}\{\\sigma^\{2\}\}XX^\{\\top\}\.Furthermore,
G=Cov\(∇θℓθ0\(X,Y\)\)=1σ2𝔼\[XX⊤\],G=\\operatorname\{Cov\}\\bigl\(\\nabla\_\{\\theta\}\\ell\_\{\\theta\_\{0\}\}\(X,Y\)\\bigr\)=\\frac\{1\}\{\\sigma^\{2\}\}\\mathbb\{E\}\[XX^\{\\top\}\],so Condition[\(FR1\)](https://arxiv.org/html/2607.20578#S3.I1.i1)holds\.
For everyh∈ℝdh\\in\\mathbb\{R\}^\{d\},
h⊤∇θ2ℓθ\(X,Y\)h\\displaystyle h^\{\\top\}\\nabla\_\{\\theta\}^\{2\}\\ell\_\{\\theta\}\(X,Y\)h=1σ2\(h⊤X\)2\\displaystyle=\\frac\{1\}\{\\sigma^\{2\}\}\(h^\{\\top\}X\)^\{2\}≤1σ2\(X⊤G−1X\)‖h‖G2\.\\displaystyle\\leq\\frac\{1\}\{\\sigma^\{2\}\}\\bigl\(X^\{\\top\}G^\{\-1\}X\\bigr\)\\\|h\\\|\_\{G\}^\{2\}\.Thus Condition[\(FR2\)](https://arxiv.org/html/2607.20578#S3.I1.i2)holds with
M\(Z\):=1σ2X⊤G−1X,σH:=1σ2\[𝔼\(X⊤G−1X\)2\]1/2\.M\(Z\):=\\frac\{1\}\{\\sigma^\{2\}\}X^\{\\top\}G^\{\-1\}X,\\qquad\\sigma\_\{H\}:=\\frac\{1\}\{\\sigma^\{2\}\}\\left\[\\mathbb\{E\}\\bigl\(X^\{\\top\}G^\{\-1\}X\\bigr\)^\{2\}\\right\]^\{1/2\}\.
The whitened gradient is
ζ=−εσ2G−1/2X\.\\zeta=\-\\frac\{\\varepsilon\}\{\\sigma^\{2\}\}G^\{\-1/2\}X\.Using independence and𝔼ε4=3σ4\\mathbb\{E\}\\varepsilon^\{4\}=3\\sigma^\{4\},
𝔼‖ζ‖24=3σ4𝔼\(X⊤G−1X\)2\.\\mathbb\{E\}\\\|\\zeta\\\|\_\{2\}^\{4\}=\\frac\{3\}\{\\sigma^\{4\}\}\\mathbb\{E\}\\bigl\(X^\{\\top\}G^\{\-1\}X\\bigr\)^\{2\}\.Hence Conditions[\(FR2\)](https://arxiv.org/html/2607.20578#S3.I1.i2)and[\(FR3\)](https://arxiv.org/html/2607.20578#S3.I1.i3)both follow from
𝔼\(X⊤G−1X\)2<∞,\\mathbb\{E\}\\bigl\(X^\{\\top\}G^\{\-1\}X\\bigr\)^\{2\}<\\infty,withκ\\kappachosen so that
3σ4𝔼\(X⊤G−1X\)2≤κ4d2\.\\frac\{3\}\{\\sigma^\{4\}\}\\mathbb\{E\}\\bigl\(X^\{\\top\}G^\{\-1\}X\\bigr\)^\{2\}\\leq\\kappa^\{4\}d^\{2\}\.
## Appendix BDual coordinates in exponential families
Let
pη\(x\)=exp\(η⊤t\(x\)−A\(η\)\)p\_\{\\eta\}\(x\)=\\exp\\bigl\(\\eta^\{\\top\}t\(x\)\-A\(\\eta\)\\bigr\)be a regular minimal exponential family\. Its mean parameter is
μ=∇A\(η\),\\mu=\\nabla A\(\\eta\),and the inverse relation isη=∇A∗\(μ\)\\eta=\\nabla A^\{\*\}\(\\mu\), whereA∗A^\{\*\}is the Legendre dual ofAA\.
Fixη0\\eta\_\{0\}, and write
μ0:=∇A\(η0\),G:=∇2A\(η0\)≻0\.\\mu\_\{0\}:=\\nabla A\(\\eta\_\{0\}\),\\qquad G:=\\nabla^\{2\}A\(\\eta\_\{0\}\)\\succ 0\.Differentiating∇A∗\(∇A\(η\)\)=η\\nabla A^\{\*\}\(\\nabla A\(\\eta\)\)=\\etaatη0\\eta\_\{0\}gives
∇2A∗\(μ0\)∇2A\(η0\)=Id,\\nabla^\{2\}A^\{\*\}\(\\mu\_\{0\}\)\\nabla^\{2\}A\(\\eta\_\{0\}\)=I\_\{d\},and therefore
∇2A∗\(μ0\)=\[∇2A\(η0\)\]−1=G−1\.\\nabla^\{2\}A^\{\*\}\(\\mu\_\{0\}\)=\\bigl\[\\nabla^\{2\}A\(\\eta\_\{0\}\)\\bigr\]^\{\-1\}=G^\{\-1\}\.Thus the Fisher metric in natural coordinates and its inverse in mean coordinates arise as the Hessians of a Legendre\-dual pair\. This provides a canonical dual\-coordinate interpretation of the two metric deformations used in the main text\.
The interpretation does not make same\-coordinate comparisons invariant\. For example, under the one\-dimensional rescaling
η′=cη,c≠1,\\eta^\{\\prime\}=c\\eta,\\qquad c\\neq 1,the Fisher information transforms asG′=c−2GG^\{\\prime\}=c^\{\-2\}G, while a fixed numerical intervalT=\[−δ,δ\]T=\[\-\\delta,\\delta\]in theη′\\eta^\{\\prime\}\-chart does not represent the same tangent perturbations as the interval with the same endpoints in theη\\eta\-chart\.
The preceding identities concern natural and mean coordinates transformed according to their dual laws\. Same\-coordinate comparisons in the main text are interpreted as in Remark[2\.4](https://arxiv.org/html/2607.20578#S2.Thmtheorem4)\.Similar Articles
Fisher Width: A Geometric Measure of Complexity on Statistical Manifolds
Introduces Fisher width, a Riemannian analogue of Gaussian width for statistical manifolds, which captures local statistical curvature and is invariant under reparameterization. The paper develops its theory, proves generalization bounds for Fisher-Lipschitz classes, and demonstrates computable estimators on MNIST.
Fisher8: Stabilizing Neural Heteroscedastic Regression via Output-Layer Fisher Geometry
This paper introduces Fisher8, an output-layer gradient correction that uses Fisher geometry instead of Euclidean geometry to stabilize neural heteroscedastic regression, improving uncertainty calibration and likelihood-error tradeoffs.
Measuring Model Robustness via Fisher Information: Spectral Bounds, Theoretical Guarantees, and Practical Algorithms
The paper proposes an attack-agnostic robustness metric based on the spectral norm of the Fisher Information Matrix, providing theoretical bounds and scalable evaluation methods for deep neural networks.
Finsler Geometry, Graph Neural Networks, and You
This paper proposes a Finslerian graph neural network that estimates the Finsler Laplacian on point clouds, proving convergence and demonstrating its use in recovering Finsler metrics from heat diffusion.
A Local Sinkhorn Framework for Conditional Distribution Reconstruction of Multidimensional Random Fields
This paper proposes a scalable local Sinkhorn divergence framework for training stochastic neural networks to reconstruct multidimensional random fields, with theoretical generalization error bounds and numerical demonstrations for uncertainty quantification.