K\"ahler landscapes for complex neural network descents and guarantees including a search and destroy of the Calabi-Yau manifold
Summary
This paper explores optimization landscapes in complex neural networks using Kähler geometry and information manifolds, providing theoretical guarantees on descent paths and analyzing effects of Calabi-Yau metrics.
View Cached Full Text
Cached at: 08/21/26, 10:25 AM
# Kähler landscapes for complex neural network descents and guarantees including a search and destroy of the Calabi-Yau manifold
Source: [https://arxiv.org/html/2608.19584](https://arxiv.org/html/2608.19584)
Andrew GracykAffiliation:*Department of Mathematics* *Purdue University* *West Lafayette, IN 47907, USA* agracyk@purdue\.edu
###### Abstract
We study landscapes for complex\-parameterized networks\. Our approach is motivated with an information\-theoretic manifold perspective of the parameter and via classical optimization guarantees although of complex geometric variety such as through Dolbeault asymptotics\. The descent path admits a Kähler information metric under a cross\-entropy via the Wirtinger Hessian on the log\-likelihood potential\. We restrict attention to a descent update rule with natural gradient descent via a differentiated loss scaled by the inverse metric, so the descent path remains in the holomorphic tangent bundle\. We emphasize Calabi\-Yau information manifolds which profane theoretical guarantees via an ill\-curvature\-conditioned landscape\. Under a Calabi\-Yau metric, specifically in a non\-compact setting with a global potential so defined geometrically rather than invoking the topological requirements of the Calabi conjecture, a wedged nowhere\-vanishing holomorphic form is the top exterior product of the Kähler form up to constants, yielding a constant determinant condition\. Under a fixed determinant, a metric almost low rank up to an eigenvalue tolerance implies a blow\-up effect\. Moreover, it has been discovered that negative curvature subverts the loss landscape, specifically sectional curvature, so we expand on this and draw interconnections to negative\-definite Ricci curvature\. Our arguments primarily exist in a geometric analytic modality, although we establish roots in deep learning theory such as through asymptotics at initialization and connections through failure modes of neural network guarantees under vanishing and negative Ricci curvature\.
Key words\.information geometry, complex geometry, Kähler manifold, Kähler geometry, Calabi\-Yau manifold, Dolbeault, Monge\-Ampère, Chern curvature, strong convexity, Polyak\-Łojasiewicz, Dirichlet energy, negative curvature, Ricci curvature, canonical line bundle, sheaf of sections, de Rham cohomology
AMS MSC Classifications \(2020\):53B35, 53B12, 53Z50, 90C25
Contents
## 1Introduction
We attempt to lay foundations of optimization twofold: \(1\) for the complex\-parametered neural network[68](https://arxiv.org/html/2608.19584#bib.bib43), and \(2\) under information manifolds[41](https://arxiv.org/html/2608.19584#bib.bib44)\. Complex neural networks attempt to remedy a parameterization bottleneck, and can achieve various levels of performance gains and dominion over real counterparts[1](https://arxiv.org/html/2608.19584#bib.bib42)\. We possess roots in deep learning theory with relevant asymptotics, but our underlying mechanisms and strategies will be conducted via geometric analysis\. Loss landscapes admit geometric structure via information manifolds[5](https://arxiv.org/html/2608.19584#bib.bib33), so our work is motivated by this perspective\.
A primary focus of our work is in Dolbeault asymptotics\.[10](https://arxiv.org/html/2608.19584#bib.bib1)establishes a modern Hessian bound, which we reinforce both theoretically and empirically in our work, in which Hessian takes on desirable asymptotic scaling for sufficiently nice regions pertaining space encompassed under spectral norms and cosine similarity, which in turn allow convexity arguments to take place\. Existing asymptotics exist for real\-valued neural networks, and are for Euclidean gradient descent methods that are unaccommodating for manifold information\-theoretic structure\. Our work addresses these qualities through a Kähler geometric analysis\.
Under a Boreal measure on the training data that corresponds to a sufficiently smooth Radon\-Nikodym derivative that is disattached from neural network parameter dependence, the parameter landscape admits a cross\-entropy metric, reminscent of a Fisher metric although not exactly because of the parameter invariance, that is moreover a mixed Wirtinger derivative of a potential since the derivatives commute with the iterated integral\. The potential, for example in the quadratic case, is plurisubharmonic\. Therefore, under this potential, the admitted information manifold is Kähler\. Under further restriction of \(weak\) equivalency conditions such as
iK2Ω∧Ω¯−constant⋅ωKK\!≡0,\\displaystyle i^\{K^\{2\}\}\\Omega\\wedge\\overline\{\\Omega\}\-\\text\{constant\}\\cdot\\frac\{\\omega^\{K\}\}\{K\!\}\\equiv 0,\(1\.1\)or vanishing Ricci curvature, the geometry is Calabi\-Yau for a specific Kähler class\. We will attempt to investigate the pitfalls of Calabi\-Yau metrics on information landscapes\. We study these circumstances solely affected in the Calabi\-Yau scenario via[10\.4](https://arxiv.org/html/2608.19584#S10.SS4),[10\.5](https://arxiv.org/html/2608.19584#S10.SS5), but also draw connections in[10\.6](https://arxiv.org/html/2608.19584#S10.SS6),[11\.1](https://arxiv.org/html/2608.19584#S11.SS1),[11\.2](https://arxiv.org/html/2608.19584#S11.SS2),[12\.1](https://arxiv.org/html/2608.19584#S12.SS1),[12\.2](https://arxiv.org/html/2608.19584#S12.SS2), which are sections that discuss roles of curvature in general\. In our analyses, the Ricci curvature, further tethered to metric eigenvalues that both diminish and blow\-up due to a constant determinant condition111Here, we are using the Kähler version of Ricci curvature, since it is defined via a log determinant, and a determinant is a product of eigenvalues\., have interconnections to our theory\. As we will see, optimization arguments are not immune to structures profaned in curvature\.
Pitfalls are not unique to Calabi\-Yau metrics, and can be properties of curvature in general\. In Appendices[11\.1](https://arxiv.org/html/2608.19584#S11.SS1),[11\.2](https://arxiv.org/html/2608.19584#S11.SS2),[12\.1](https://arxiv.org/html/2608.19584#S12.SS1),[12\.2](https://arxiv.org/html/2608.19584#S12.SS2), we demonstrate that Ricci curvature itself can mar the learning landscape\. In fact, in some results, the severity of the effect of curvature can be in exact proportion to the magnitude of the curvature, so Calabi\-Yau metrics are merely a threshold as to where adverse effects can start\. Nonetheless, we also demonstrate pitfalls unique to Calabi\-Yau metrics in[10\.4](https://arxiv.org/html/2608.19584#S10.SS4),[10\.5](https://arxiv.org/html/2608.19584#S10.SS5)\.
## 2Related work
Figure 1:We plot a cross\-section and a curve corresponding to a descent path along a Calabi\-Yau manifold\. Our slice corresponds to the Fermat equationZ1N\+Z2N=1Z\_\{1\}^\{N\}\+Z\_\{2\}^\{N\}=1,N=12N=12\. The surface is projected onto the real plane by taking the real components\(Re\(Z1\),Re\(Z2\)\)\(\\text\{Re\}\(Z\_\{1\}\),\\text\{Re\}\(Z\_\{2\}\)\)\. This figure is somewhat toy because in a descending landscape scenario, the manifold is much more high\-dimensional\.Connections to more traditional deep learning theory\.Our work possesses foundations in deep learning theory, namely at initialization\. Literature that establish to varying levels neural network derivative bounds including use of Feynmann diagrams are[10](https://arxiv.org/html/2608.19584#bib.bib1)[23](https://arxiv.org/html/2608.19584#bib.bib2)[78](https://arxiv.org/html/2608.19584#bib.bib3)[62](https://arxiv.org/html/2608.19584#bib.bib4)[2](https://arxiv.org/html/2608.19584#bib.bib5)\. Bounding a Hessian norm can also be found in[14](https://arxiv.org/html/2608.19584#bib.bib14)\. Additional work relevant to deep learning theory and the roles of width in asymptotics include[32](https://arxiv.org/html/2608.19584#bib.bib9)[7](https://arxiv.org/html/2608.19584#bib.bib11)[13](https://arxiv.org/html/2608.19584#bib.bib12)and especially pertaining to the neural tangent kernel[31](https://arxiv.org/html/2608.19584#bib.bib7)[36](https://arxiv.org/html/2608.19584#bib.bib10)[45](https://arxiv.org/html/2608.19584#bib.bib16)\. Much of our work also pertains to optimization through a geometric and analytic lens\. Relevant work from a more real\-analytic lens include[37](https://arxiv.org/html/2608.19584#bib.bib45)[77](https://arxiv.org/html/2608.19584#bib.bib47)with connections to Wasserstein geometry[17](https://arxiv.org/html/2608.19584#bib.bib46)\. Optimization literature focusing specifically on \(Riemannian\) geometric analysis include[4](https://arxiv.org/html/2608.19584#bib.bib48)[75](https://arxiv.org/html/2608.19584#bib.bib49)[56](https://arxiv.org/html/2608.19584#bib.bib50)[69](https://arxiv.org/html/2608.19584#bib.bib51)[39](https://arxiv.org/html/2608.19584#bib.bib52), so our work has high interconnections to these among those mentioned\.
Connections to negative sectional curvature\.The role of Ricci curvature in an optimization landscape has been studied[27](https://arxiv.org/html/2608.19584#bib.bib53)[46](https://arxiv.org/html/2608.19584#bib.bib54)[6](https://arxiv.org/html/2608.19584#bib.bib55), but these works are mostly separate from a deep learning theory perspective, i\.e\. neglect commentary on width, etc\.[43](https://arxiv.org/html/2608.19584#bib.bib72)does study how positive Ricci curvature positively affects convergence, therefore Ricci flat and negatively\-curved metrics lack this\. On the other hand,[15](https://arxiv.org/html/2608.19584#bib.bib56)shows that for many manifolds including Hadamard manifolds and hyperbolic spaces, i\.e\. with constant negative sectional curvature, gradient descent experiences pitfalls due to to the curvature\.[16](https://arxiv.org/html/2608.19584#bib.bib57)shows that, primarily under hyperbolic spaces, convexity results often worsen due to the curvature effects\. Our work is reminiscent of these works, while our work emphasizes Ricci curvature and eigenvalues specifically\. The constant determinant implies very large eigenvalues is unique to complex manifolds under Calabi\-Yau metrics, therefore many of our results are lost in translation to Riemannian manifolds, since Ricci curvature being the log determinant of the metric is unique to Kähler manifolds\. We build upon these works by examining the Ricci\-flat scenario\. In particular, our methods will also experience pitfalls under negative curvature as well, and so we also consider
Ric\(v,v\)<0,\\displaystyle\\text\{Ric\}\(v,v\)<0,\(2\.1\)and sometimes the Calabi\-Yau manifold is simply the transition from the better to worse cases\. In general, in sufficiently high dimensions, Ricci flatness is a weaker condition than a sectional curvature condition, since
Ric\(X,X\)=TrhRm\(X,⋅,X,⋅\)=∑iK\(X,ei\)≡0/⟹K\(X,Y\)≡0\.\\displaystyle\\text\{Ric\}\(X,X\)=\\text\{Tr\}\_\{h\}\\text\{Rm\}\(X,\\cdot,X,\\cdot\)=\\sum\_\{i\}K\(X,e\_\{i\}\)\\equiv 0\\quad\\mathchoice\{\\mathrel\{\\hbox to0\.0pt\{\\kern 3\.75pt\\kern\-5\.27776pt$\\displaystyle\\not$\\hss\}\{\\implies\}\}\}\{\\mathrel\{\\hbox to0\.0pt\{\\kern 3\.75pt\\kern\-5\.27776pt$\\textstyle\\not$\\hss\}\{\\implies\}\}\}\{\\mathrel\{\\hbox to0\.0pt\{\\kern 2\.625pt\\kern\-4\.45831pt$\\scriptstyle\\not$\\hss\}\{\\implies\}\}\}\{\\mathrel\{\\hbox to0\.0pt\{\\kern 1\.875pt\\kern\-3\.95834pt$\\scriptscriptstyle\\not$\\hss\}\{\\implies\}\}\}\\quad K\(X,Y\)\\equiv 0\.\(2\.2\)Since the eigenvalue explosion, on the other hand, is unique to our situation, our Calabi\-Yau results particularly often go hand\-in\-hand with violatingβ\\beta\-smoothness\.
## 3Neural network setup
Consider training data\{zi,yi\}i,zi∈𝒵⊆ℂd,yi∈𝒴⊆ℝ\\\{z\_\{i\},y\_\{i\}\\\}\_\{i\},z\_\{i\}\\in\\mathcal\{Z\}\\subseteq\\mathbb\{C\}^\{d\},y\_\{i\}\\in\\mathcal\{Y\}\\subseteq\\mathbb\{R\}\(without loss of generality, we can restrict the imaginary component ofyiy\_\{i\}to be00if necessary, thus we maintain representations as generalized as possible\), wherey∼p\(y\|x,θ\)=𝒩\(f\(θ,x\),σ2\)y\\sim p\(y\|x,\\theta\)=\\mathcal\{N\}\(f\(\\theta;x\),\\sigma^\{2\}\)follows a data distribution\. Here,θ∈Θ\\theta\\in\\Thetais the total collection of complex\-valued weights and biases\. Consider a fully\-connected complex neural network of the form[10](https://arxiv.org/html/2608.19584#bib.bib1)
α\(0\)\(z\)=z\\displaystyle\\alpha^\{\(0\)\}\(z\)=z\(3\.1\)h\(l\)\(z\)=1mW\(l\)α\(l−1\)\(z\)\\displaystyle h^\{\(l\)\}\(z\)=\\frac\{1\}\{\\sqrt\{m\}\}W^\{\(l\)\}\\alpha^\{\(l\-1\)\}\(z\)\(3\.2\)α\(l\)\(z\)=ϕ\(h\(l\)\(z\),h\(l\)\(z\)¯\),l∈\[L\]\\displaystyle\\alpha^\{\(l\)\}\(z\)=\\phi\\left\(h^\{\(l\)\}\(z\),\\overline\{h^\{\(l\)\}\(z\)\}\\right\),\\quad l\\in\[L\]\(3\.3\)f\(θ,z,z¯\)=α\(L\+1\)\(z\)=1mLv†α\(L\)\(z\),\\displaystyle f\(\\theta;z,\\overline\{z\}\)=\\alpha^\{\(L\+1\)\}\(z\)=\\frac\{1\}\{\\sqrt\{m\_\{L\}\}\}v^\{\\dagger\}\\alpha^\{\(L\)\}\(z\),\(3\.4\)whereWWis a linear operator over the field of complex numbersℂ\\mathbb\{C\}, andϕ:ℂ→ℂ\\phi:\\mathbb\{C\}\\rightarrow\\mathbb\{C\}is a non\-holomorphic activation\. Here,θ∈Θ⊆ℂ∑kmkmk\+1\+mL:=ℂK\\theta\\in\\Theta\\subseteq\\mathbb\{C\}^\{\\sum\_\{k\}m\_\{k\}m\_\{k\+1\}\+m\_\{L\}\}:=\\mathbb\{C\}^\{K\}\. For simplicity, assumeml=mm\_\{l\}=mfor allll\.
## 4Geometric setup
### 4\.1The Kähler descent landscape and its information geometry
Construct a probability measure such thatyi∼q\(y\|z\)y\_\{i\}\\sim q\(y\|z\)\. Noticeqqdoes not depend onθ\\theta\. This is not unusual per se, although sometimesqqis parameterized\. We can note a density of this forms "violates injectivity" since the output ofyyvaries across singlezz\. More specifically,qqis a Markov kernelz↦ℙY\|Z=zz\\mapsto\\mathbb\{P\}\_\{Y\|Z=z\}[26](https://arxiv.org/html/2608.19584#bib.bib62)\. The cross\-entropy information metric is defined as
hij¯\(θ\)=𝔼z∼pdata\[𝔼y∼q\(y\|z\)\[−∂2logp\(y\|z,θ\)∂θi∂θ¯j\]\],\\displaystyle h\_\{i\\overline\{j\}\}\(\\theta\)=\\mathbb\{E\}\_\{z\\sim p\_\{\\text\{data\}\}\}\\left\[\\mathbb\{E\}\_\{y\\sim q\(y\|z\)\}\\left\[\-\\frac\{\\partial^\{2\}\\log p\(y\|z,\\theta\)\}\{\\partial\\theta^\{i\}\\partial\\overline\{\\theta\}^\{j\}\}\\right\]\\right\],\(4\.1\)
which defines a Kähler manifold loss landscape\(M,ω\),ω∈Ω1,1\(M\)\(M,\\omega\),\\omega\\in\\Omega^\{1,1\}\(M\)under a preconditioned loss, or natural gradient descent where the descent update is scaled in accordance with the inverse information metric\. In the above,ppis taken to be the loss\. The above metric is Kähler since the Wirtinger derivatives commute with the integral under complex\-valued Lebesgue dominated convergence[79](https://arxiv.org/html/2608.19584#bib.bib63), so
hij¯\(θ\)=∂i∂j¯Φ:=∂i∂j¯potential,potential∈SPSH\(U\),\\displaystyle h\_\{i\\overline\{j\}\}\(\\theta\)=\\partial\_\{i\}\\partial\_\{\\overline\{j\}\}\\Phi:=\\partial\_\{i\}\\partial\_\{\\overline\{j\}\}\\ \\text\{potential\},\\ \\text\{potential\}\\in\\text\{SPSH\}\(U\),\(4\.2\)by takingΦ=𝔼z𝔼q\[−logp\]\\Phi=\\mathbb\{E\}\_\{z\}\\mathbb\{E\}\_\{q\}\[\-\\log p\]\. We can note the Wirtinger derivatives commute underL1L^\{1\}decay via dominated convergence
∂θi∂θj¯∫𝒵∫𝒴−logp\(y\|z,θ\)q\(y\|z\)pdata\(z\)dydz⏟:=Φ=∫𝒵∫𝒴−∂θi∂θj¯logp\(y\|z,θ\)q\(y\|z\)pdata\(z\)dydz\.\\displaystyle\\partial\_\{\\theta^\{i\}\}\\partial\_\{\\theta^\{\\overline\{j\}\}\}\\underbrace\{\\int\_\{\\mathcal\{Z\}\}\\int\_\{\\mathcal\{Y\}\}\-\\log p\(y\|z,\\theta\)q\(y\|z\)p\_\{\\text\{data\}\}\(z\)dydz\}\_\{\\displaystyle:=\\Phi\}=\\int\_\{\\mathcal\{Z\}\}\\int\_\{\\mathcal\{Y\}\}\-\\partial\_\{\\theta^\{i\}\}\\partial\_\{\\theta^\{\\overline\{j\}\}\}\\log p\(y\|z,\\theta\)q\(y\|z\)p\_\{\\text\{data\}\}\(z\)dydz\.\(4\.3\)It is crucial to note we have removed the dependence onθ\\thetain the latter staticpp\. Therefore, the above is more closely related to a cross\-entropy Hessian rather than a true Fisher metric\. Whenppis an exponential likelihood, the negative log likelihood is quadratic, which is SPSH\. In particular, a quadratic cost submits to a quadratic formΦ\(θ\)=12θ†Aθ\+b†θ\+θ†c\+d\\Phi\(\\theta\)=\\frac\{1\}\{2\}\\theta^\{\\dagger\}A\\theta\+b^\{\\dagger\}\\theta\+\\theta^\{\\dagger\}c\+dwithout an expected value\. To bridge the gap between convexity and the Wirtinger derivatives, the Hessian coheres to
\(Hℂ\)jk¯=14\(∂2Φ∂uj∂uk\+∂2Φ∂vj∂vk\+i\(∂2Φ∂uj∂vk−∂2Φ∂vj∂uk\)\)\.\\displaystyle\(H\_\{\\mathbb\{C\}\}\)\_\{j\\overline\{k\}\}=\\frac\{1\}\{4\}\\left\(\\frac\{\\partial^\{2\}\\Phi\}\{\\partial u\_\{j\}\\partial u\_\{k\}\}\+\\frac\{\\partial^\{2\}\\Phi\}\{\\partial v\_\{j\}\\partial v\_\{k\}\}\+i\\left\(\\frac\{\\partial^\{2\}\\Phi\}\{\\partial u\_\{j\}\\partial v\_\{k\}\}\-\\frac\{\\partial^\{2\}\\Phi\}\{\\partial v\_\{j\}\\partial u\_\{k\}\}\\right\)\\right\)\.\(4\.4\)
Figure 2:We plot‖i∂∂¯f‖2\\\|i\\partial\\overline\{\\partial\}f\\\|\_\{2\}asymptotics corresponding to[8\.2](https://arxiv.org/html/2608.19584#S8.SS2)on 20 instances, corresponding to exactly initialization, near initialization \(perturbationϵ=0\.1\\epsilon=0\.1\), and far from initialization \(ϵ=2\\epsilon=2\)\.Here,θ=u\+iv\\theta=u\+iv\. Substituting in the expected quadratic cost, we get the quadratic form is greater than zero and soΦ\\Phiis SPSH\. As a last remark, we can note the expectation operator preserves a quadratic property since for a probability\(Ω,ℱ,ℙ\)\(\\Omega,\\mathcal\{F\},\\mathbb\{P\}\)be a probability space with random vectorz∈L2\(Ω,ℂn\)z\\in L^\{2\}\(\\Omega;\\mathbb\{C\}^\{n\}\)
𝔼\[z†Az\]=𝔼\[Tr\(z†Az\)\]=𝔼\[Tr\(Azz†\)\]\\displaystyle\\mathbb\{E\}\[z^\{\\dagger\}Az\]=\\mathbb\{E\}\[\\text\{Tr\}\(z^\{\\dagger\}Az\)\]=\\mathbb\{E\}\[\\text\{Tr\}\(Azz^\{\\dagger\}\)\]\(4\.5\)=Tr\(A𝔼\[zz†\]\)=Tr\(AΣ\)\+μ†Aμ\.\\displaystyle=\\text\{Tr\}\(A\\mathbb\{E\}\[zz^\{\\dagger\}\]\)=\\text\{Tr\}\(A\\Sigma\)\+\\mu^\{\\dagger\}A\\mu\.\(4\.6\)
Denoteℒ\(θ,θ¯\)\\mathcal\{L\}\(\\theta,\\overline\{\\theta\}\)the loss function under a complex parameter, so it dependent on the parameter itself and its complex conjugate\. Under a traditional neural network trajectory, the gradient descent updates obey∂tθα\(t\)=−∂ℒ/∂θ¯α\\partial\_\{t\}\\theta^\{\\alpha\}\(t\)=\-\\partial\\mathcal\{L\}/\\partial\\overline\{\\theta\}^\{\\alpha\}\. For this work, we consider the geometrically\-preconditioned backpropagated loss[63](https://arxiv.org/html/2608.19584#bib.bib58)
dθi\(t\)dt=−hij¯∂ℒ∂θ¯j,\\displaystyle\\frac\{d\\theta^\{i\}\(t\)\}\{dt\}=\-h^\{i\\overline\{j\}\}\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\overline\{\\theta\}^\{j\}\},\(4\.7\)wheredθi/dtd\\theta^\{i\}/dtlives inT1,0\(M\)T^\{1,0\}\(M\), Moreover,h∈Γ\(T∗1,0M⊗T∗0,1M\)h\\in\\Gamma\(T^\{\*1,0\}M\\otimes T^\{\*0,1\}M\)is a section, soh−1∈Γ\(T1,0M⊗T0,1M\)h^\{\-1\}\\in\\Gamma\(T^\{1,0\}M\\otimes T^\{0,1\}M\)and the inverse information metric acts such that♯:T∗0,1M→T1,0M\\sharp:T^\{\*0,1\}M\\to T^\{1,0\}M\. Under this trajectory, the parameter obeys the information manifold Kähler structure, which serves as a loss landscape\. We can note the descent trajectory is just one direction in the span of the holomorphic tangent bundle, and it is in the direction of steepest descent\. In particular,dθ/dt∈Tθ\(t\)1,0Md\\theta/dt\\in T\_\{\\theta\(t\)\}^\{1,0\}Mand we get for some unimportant orthogonal complement𝒲\\mathcal\{W\}
Tθ\(t\)1,0M=spanℂ\{dθdt\}⊕h𝒲\.\\displaystyle T\_\{\\theta\(t\)\}^\{1,0\}M=\\text\{span\}\_\{\\mathbb\{C\}\}\\left\\\{\\frac\{d\\theta\}\{dt\}\\right\\\}\\oplus\_\{h\}\\mathcal\{W\}\.\(4\.8\)
Let us verify our parameter update in the complex case indeed decreases the loss under[4\.7](https://arxiv.org/html/2608.19584#S4.E7)\. Consider the total loss differentialdℒ=∂ℒ∂θidθi\+∂ℒ∂θ¯jdθ¯jd\\mathcal\{L\}=\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\theta^\{i\}\}d\\theta^\{i\}\+\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\overline\{\\theta\}^\{j\}\}d\\overline\{\\theta\}^\{j\}\. Because the loss is real\-valued, we have\(∂ℒ∂θi\)¯=∂ℒ∂θ¯i\\overline\{\\left\(\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\theta^\{i\}\}\\right\)\}=\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\overline\{\\theta\}^\{i\}\}\. Using our preconditioned trajectory, the parameter differential and its conjugate aredθi=−ηhij¯∂ℒ∂θ¯jd\\theta^\{i\}=\-\\eta h^\{i\\overline\{j\}\}\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\overline\{\\theta\}^\{j\}\}anddθ¯j=−ηhij¯∂ℒ∂θid\\overline\{\\theta\}^\{j\}=\-\\eta h^\{i\\overline\{j\}\}\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\theta^\{i\}\}, where we use the fact that the metric is Hermitian\. Into the total differential, we see
dℒ\\displaystyle d\\mathcal\{L\}=∂ℒ∂θi\(−ηhij¯∂ℒ∂θ¯j\)\+∂ℒ∂θ¯j\(−ηhij¯∂ℒ∂θi\)=−2ηhij¯∂ℒ∂θi∂ℒ∂θ¯j\.\\displaystyle=\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\theta^\{i\}\}\\left\(\-\\eta h^\{i\\overline\{j\}\}\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\overline\{\\theta\}^\{j\}\}\\right\)\+\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\overline\{\\theta\}^\{j\}\}\\left\(\-\\eta h^\{i\\overline\{j\}\}\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\theta^\{i\}\}\\right\)=\-2\\eta h^\{i\\overline\{j\}\}\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\theta^\{i\}\}\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\overline\{\\theta\}^\{j\}\}\.\(4\.9\)Because the information metrichhis positive definite, the quadratic formhij¯∂ℒ∂θi∂ℒ∂θ¯jh^\{i\\overline\{j\}\}\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\theta^\{i\}\}\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\overline\{\\theta\}^\{j\}\}is real and positive, guaranteeingdℒd\\mathcal\{L\}is real and negative\.
Definition 1\.Let\(M,ω\)\(M,\\omega\)be a Kähler information manifold representing the parameter space, whereω\\omegais the Kähler form induced by the information potential\. A smooth, real\-valued loss functionL:M→ℝL:M\\to\\mathbb\{R\}is said to satisfyμ\\mu\-\(complex\) strong convexity if, for anyz′∈Sz^\{\\prime\}\\in Sand a referencez∈Mz\\in M, we have
L\(θ′\)≥L\(θ\)\+2Re⟨∂Lθ,expθ−1\(θ′\)1,0⟩\+μ2dω\(θ,θ′\)2\.\\displaystyle L\(\\theta^\{\\prime\}\)\\geq L\(\\theta\)\+2\\text\{Re\}\\langle\\partial L\_\{\\theta\},\\exp\_\{\\theta\}^\{\-1\}\(\\theta^\{\\prime\}\)^\{1,0\}\\rangle\+\\frac\{\\mu\}\{2\}d\_\{\\omega\}\(\\theta,\\theta^\{\\prime\}\)^\{2\}\.\(4\.10\)withμ\>0\\mu\>0\. Here,dωd\_\{\\omega\}is the geodesic distance\. This condition is a bridge to the classical \(Dolbeault\) Hessian boundi∂∂¯L\(z\)⪰μωi\\partial\\overline\{\\partial\}L\(z\)\\succeq\\mu\\omega\.
### 4\.2Calabi\-Yau descents
Definition 2\.We say the Kähler manifoldMMthat admits metrichhis Calabi\-Yau if its Ricci curvature is zero, i\.e\.
Ricij¯\\displaystyle\\text\{Ric\}\_\{i\\overline\{j\}\}=−∂i∂j¯logdet\(h\)\\displaystyle=\-\\partial\_\{i\}\\partial\_\{\\overline\{j\}\}\\log\\det\(h\)\(4\.11\)=−∂i∂j¯logdet\(𝔼z∼pdata\[𝔼y∼q\(y\|z\)\[−∂2logp\(y\|z,θ\)∂θi∂θ¯j\]\]\)≡0\.\\displaystyle=\-\\partial\_\{i\}\\partial\_\{\\overline\{j\}\}\\log\\det\\left\(\\mathbb\{E\}\_\{z\\sim p\_\{\\text\{data\}\}\}\\left\[\\mathbb\{E\}\_\{y\\sim q\(y\|z\)\}\\left\[\-\\frac\{\\partial^\{2\}\\log p\(y\|z,\\theta\)\}\{\\partial\\theta^\{i\}\\partial\\overline\{\\theta\}^\{j\}\}\\right\]\\right\]\\right\)\\equiv 0\.\(4\.12\)
This definition has topological and geometric nuance, so we elaborate\. We adopt this geometric definition\. Generally, to invoke the Calabi conjecture[74](https://arxiv.org/html/2608.19584#bib.bib73), we require compactness ofMM, or at least of the submanifold in which the optimization trajectory exists\. We remark this is slightly nonrigorous because the Calabi conjecture requires the manifold to be closed without boundary, and a compact submanifold would contain a boundary with Dirichlet or Neumann boundary conditions\. This is problematic for us because our metric is defined via a Kähler potential, and by the maximum principle, a globally defined strictly plurisubharmonic function on a compact manifold must be constant[20](https://arxiv.org/html/2608.19584#bib.bib75)\. This completely destroys our defined metric, at least globally but not locally\. Topologically, the manifold is governed by the vanishing of the first real Chern class,c1\(M,ℝ\)=0c\_\{1\}\(M;\\mathbb\{R\}\)=0\. While stronger algebraic definitions exist, such as requiring the canonical line bundle to be trivial, which equipsMMwith a nowhere\-vanishing holomorphicKK\-form, the conditionc1\(M,ℝ\)=0c\_\{1\}\(M;\\mathbb\{R\}\)=0is the, albeit weaker, requirement for Calabi\-Yau manifolds\. Therefore, we bypass the Chern class requirement of Yau’s theorem, and we define the Calabi\-Yau manifold entirely geometrically rather than topologically by equipping it with a nowhere\-vanishing holomorphic formΩ\\Omegaand ensuring it is Ricci\-flat\. We remark the title of this work is also allusive to a search of the Ricci flat metric within a Kähler class, but since we relaxed compactness, this is not rigorous\.
We can note the global Ricci formρ=−i∂∂¯logdet\(h\)\\rho=\-i\\partial\\overline\{\\partial\}\\log\\det\(h\)is, up to a constant, the curvature 2\-formF∇F^\{\\nabla\}of the induced Chern connection on the anti\-canonical line bundleKM−1=ΛKT1,0MK\_\{M\}^\{\-1\}=\\Lambda^\{K\}T^\{1,0\}Mover the parameter space, so we get the Calabi\-Yau conditionRic≡0⇔ρ≡0⇔F∇≡0\\text\{Ric\}\\equiv 0\\iff\\rho\\equiv 0\\iff F^\{\\nabla\}\\equiv 0and
F∇=∂¯∂log\(ωKK\!−1cKΩ∧Ω¯\)≡0\\displaystyle F^\{\\nabla\}=\\overline\{\\partial\}\\partial\\log\\left\(\\frac\{\\omega^\{K\}K\!^\{\-1\}\}\{c\_\{K\}\\Omega\\wedge\\overline\{\\Omega\}\}\\right\)\\equiv 0\(4\.13\)
Figure 3:We plot asymptotics on the \(2,0\)\-Hessian of‖ℋ2,0‖2\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}corresponding to[8\.5](https://arxiv.org/html/2608.19584#S8.SS5)on 2 instances using an SVD power iteration approach, notingH2,0=14\(Hxx−Hyy\)\+i4\(Hxy\+Hyx\)H\_\{2,0\}=\\frac\{1\}\{4\}\(H\_\{xx\}\-H\_\{yy\}\)\+\\frac\{i\}\{4\}\(H\_\{xy\}\+H\_\{yx\}\)\.for a nowhere\-vanishing holomorphic volume formΩ∈H0\(M,KM\)\\Omega\\in H^\{0\}\(M,K\_\{M\}\), so the Dolbeault operators vanish the log of the geometric volume form and the algebraic volume form ratio\. As a remark, note that
F∇=ρ=iTr\(Θh\)\.\\displaystyle F^\{\\nabla\}=\\rho=i\\text\{Tr\}\(\\Theta\_\{h\}\)\.\(4\.14\)Here,Θh\\Theta\_\{h\}is the Chern curvature form\. A Calabi\-Yau manifold will fail our CRSC results because of the following scenario:
Equivalently, we will work with the Calabi\-Yau condition as a Monge–Ampère equation with constant determinant condition for constantκ\\kappafor manifold dimensionKK
\(i∂∂¯Φ\)K=κdV0\.\\displaystyle\(i\\partial\\overline\{\\partial\}\\Phi\)^\{K\}=\\kappa dV\_\{0\}\.\(4\.15\)This is a constant determinant condition sinceω=i∑j,k=1K\(∂j∂k¯Φ\)dθj∧dθ¯k\\omega=i\\sum\_\{j,k=1\}^\{K\}\\left\(\\partial\_\{j\}\\partial\_\{\\overline\{k\}\}\\Phi\\right\)d\\theta^\{j\}\\wedge d\\overline\{\\theta\}^\{k\},dV0=\(i2\)Kdθ1∧dθ¯1∧⋯∧dθK∧dθ¯KdV\_\{0\}=\\left\(\\frac\{i\}\{2\}\\right\)^\{K\}d\\theta^\{1\}\\wedge d\\overline\{\\theta\}^\{1\}\\wedge\\dots\\wedge d\\theta^\{K\}\\wedge d\\overline\{\\theta\}^\{K\}, and substituting into[4\.15](https://arxiv.org/html/2608.19584#S4.E15),
iKK\!det\(∂j∂k¯Φ\)dθ1∧dθ¯1∧⋯∧dθK∧dθ¯K=κ\(i2\)Kdθ1∧dθ¯1∧⋯∧dθK∧dθ¯K\.\\displaystyle i^\{K\}K\!\\det\\left\(\\partial\_\{j\}\\partial\_\{\\overline\{k\}\}\\Phi\\right\)d\\theta^\{1\}\\wedge d\\overline\{\\theta\}^\{1\}\\wedge\\dots\\wedge d\\theta^\{K\}\\wedge d\\overline\{\\theta\}^\{K\}=\\kappa\\left\(\\frac\{i\}\{2\}\\right\)^\{K\}d\\theta^\{1\}\\wedge d\\overline\{\\theta\}^\{1\}\\wedge\\dots\\wedge d\\theta^\{K\}\\wedge d\\overline\{\\theta\}^\{K\}\.\(4\.16\)which yieldsdet\(∂j∂k¯Φ\)=κK\!2K\\det\\left\(\\partial\_\{j\}\\partial\_\{\\overline\{k\}\}\\Phi\\right\)=\\frac\{\\kappa\}\{K\!2^\{K\}\}solving for the determinant\. Equivalently,[4\.15](https://arxiv.org/html/2608.19584#S4.E15)can be written
1K\!\(i∂∂¯Φ\)K=\(−1\)K\(K−1\)2\(i2\)KκΩ∧Ω¯\\displaystyle\\frac\{1\}\{K\!\}\(i\\partial\\overline\{\\partial\}\\Phi\)^\{K\}=\(\-1\)^\{\\frac\{K\(K\-1\)\}\{2\}\}\\left\(\\frac\{i\}\{2\}\\right\)^\{K\}\\kappa\\Omega\\wedge\\overline\{\\Omega\}\(4\.17\)for nowhere\-vanishing holomorphicKK\-formΩ\\Omega\. A Calabi\-Yau manifold has vanishing Ricci curvature, forcing the Monge\-Ampère conditiondet\(h\)=constant\>0\\det\(h\)=\\text\{constant\}\>0\. When the neural network landscape naturally encounters regions corresponding to regions where someλk\\lambda\_\{k\}drop, the Calabi\-Yau volume\-preservation constraint forces opposing eigenvalues to spike to compensate since it is true that
limθ→θ∗det\(h\)=limθ→θ∗∏i∈ℐ∞λi\(θ\)⋅∏j∈ℐ0λj\(θ\)⋅∏j∈ℐBλj\(θ\)≡c\>0,\\displaystyle\\lim\_\{\\theta\\to\\theta^\{\*\}\}\\det\(h\)=\\lim\_\{\\theta\\to\\theta^\{\*\}\}\\prod\_\{i\\in\\mathcal\{I\}\_\{\\infty\}\}\\lambda\_\{i\}\(\\theta\)\\cdot\\prod\_\{j\\in\\mathcal\{I\}\_\{0\}\}\\lambda\_\{j\}\(\\theta\)\\cdot\\prod\_\{j\\in\\mathcal\{I\}\_\{B\}\}\\lambda\_\{j\}\(\\theta\)\\equiv c\>0,\(4\.18\)i\.e\. a divergence of some in a diverging setℐ∞\\mathcal\{I\}\_\{\\infty\}implies collapse of some in a collapsing setℐ0\\mathcal\{I\}\_\{0\}and a sufficiently bounded, neutral setℐB\\mathcal\{I\}\_\{B\}\. To frame this alternatively, the explosion of eigenvalues force a collapse of eigenvalues into an asymptotic regime
∏j∈ℐ0λj\(θ\)≍\(∏i∈ℐ∞λi\(θ\)\)−1→θ→θ∗0\\displaystyle\\prod\_\{j\\in\\mathcal\{I\}\_\{0\}\}\\lambda\_\{j\}\(\\theta\)\\asymp\\Biggl\(\\prod\_\{i\\in\\mathcal\{I\}\_\{\\infty\}\}\\lambda\_\{i\}\(\\theta\)\\Biggr\)^\{\\\!\-1\}\\xrightarrow\[\\theta\\to\\theta^\{\*\}\]\{\}0\(4\.19\)to ensure an exact inverse scaling between the eigenvalue split at its value endpoints\. Ifλmin→0\\lambda\_\{\\text\{min\}\}\\to 0, the constant determinant condition forces at least one other eigenvalue to blow up to compensate, soλmax→∞\\lambda\_\{\\text\{max\}\}\\to\\infty\. We require an upper bound on curvaturei∂∂¯ℒ⪯βωi\\partial\\overline\{\\partial\}\\mathcal\{L\}\\preceq\\beta\\omega\. If an eigenvalue goes to infinity, theβ\\betaparameter blows up, destroying the smoothness guarantee\.
Calabi\-Yau manifolds can also offer advantages over negatively\-curved spaces, but not as much as positively\-curved spaces\. We elaborate more on effects of curvature in general in[5\.4](https://arxiv.org/html/2608.19584#S5.SS4)\. Under a continuous\-time stochastic natural gradient descent, the trajectory variance under a stochastic Jacobi equation
ddt𝔼\[‖Jt‖h2\]=\\displaystyle\\frac\{d\}\{dt\}\\mathbb\{E\}\\left\[\\\|J\_\{t\}\\\|\_\{h\}^\{2\}\\right\]=−2𝔼\[∇ω1,1ℒ\(Jt,J¯t\)\]−𝔼\[Ric\(Jt,J¯t\)\]\+Trh\(Σ\)\.\\displaystyle\-2\\mathbb\{E\}\\left\[\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\(J\_\{t\},\\overline\{J\}\_\{t\}\)\\right\]\-\\mathbb\{E\}\\left\[\\text\{Ric\}\(J\_\{t\},\\overline\{J\}\_\{t\}\)\\right\]\+\\text\{Tr\}\_\{h\}\(\\Sigma\)\.\(4\.20\)The Ricci curvature vanishes under a Calabi\-Yau manifold, but in a largely negatively curved space \(Ric≪0\\text\{Ric\}\\ll 0\), the Ricci curvature term−𝔼\[Ric\(Jt,J¯t\)\]\-\\mathbb\{E\}\\left\[\\text\{Ric\}\(J\_\{t\},\\overline\{J\}\_\{t\}\)\\right\]becomes positive and large, inducing high variance\.
### 4\.3Corrupted geometries under neural networks and eigenvalue collapse
It is not immediately guaranteed the loss landscape encounters a split in eigenvalue behavior; however, it is a realistic neural network scenario if the Calabi\-Yau manifold exists\.
Theorem 1\.Let there existKKneural network parameters and real\-data constraintsN×kN\\times k\. In the overparameterization regime, then there exist eigenvaluesλmin=0\\lambda\_\{\\text\{min\}\}=0\.
Remark\.It is not possible for a Calabi\-Yau manifold to be low rank \(except on a set of measure zero\)\. By definition, a Kähler metric must be positive\-definite\. Not only this, but a non\-invertible metric completely destroys our application of natural gradient descent\.
Proof\.In the overparameterization regime, sommis large, the metrichhis aK×KK\\times Kmatrix effectively of rank at mostkk\. The maximum possible rank ofhhismin\{K,N×k\}\\min\\\{K,N\\times k\\\}, ifKKis takenK≫N×kK\\gg N\\times k, for examplef\(z,θ\)∈ℝkf\(z;\\theta\)\\in\\mathbb\{R\}^\{k\}with a dataset of sizeNN,hhbecomes rank deficient[65](https://arxiv.org/html/2608.19584#bib.bib24)[9](https://arxiv.org/html/2608.19584#bib.bib25)[21](https://arxiv.org/html/2608.19584#bib.bib26)\. The kernel ofhhhas dimension at leastK−Nk\>0K\-Nk\>0\. Therefore, there are zero eigenvalues, forcingλmin=0\\lambda\_\{\\text\{min\}\}=0\.□\\square
Regularization and Calabi\-Yau bounded eigenvalues\.It can be noted in the case of at least one degenerate eigenvalue for sufficiently large width thresholdm∗m^\{\*\}, somewhat informally
logdet\(h\)=−∞∀m\>m∗\\displaystyle\\log\\det\(h\)=\-\\infty\\quad\\forall m\>m^\{\*\}\(4\.21\)Ricij¯=−∂i∂j¯\(−∞\)=DNE\.\\displaystyle\\text\{Ric\}\_\{i\\overline\{j\}\}=\-\\partial\_\{i\}\\partial\_\{\\overline\{j\}\}\(\-\\infty\)=\\text\{DNE\}\.\(4\.22\)This will corrupt our information manifolds\. It can be noted[10](https://arxiv.org/html/2608.19584#bib.bib1)is not a work in information geometry and does not use a metric, therefore their work is allowed to focus on restricted strong convex regimes and this work largely ignores volume collapse\. Since our approach is geometric, we cannot ignore corrupted geometries\. To reconcile this inconsistency, we can assume our loss is regularized as
ℒ\(θ\)~=ℒ\(θ\)\+λθ†θ:=−logp\(y\|z,θ\)\+λθ†θ\.\\displaystyle\\widetilde\{\\mathcal\{L\}\(\\theta\)\}=\\mathcal\{L\}\(\\theta\)\+\\lambda\\theta^\{\\dagger\}\\theta:=\-\\log p\(y\|z,\\theta\)\+\\lambda\\theta^\{\\dagger\}\\theta\.\(4\.23\)This will in turn allow a bound on a minimum eigenvalue, sincehhwill take the formh~ij¯\(θ\)=hij¯\(θ\)\+λδij¯\\widetilde\{h\}\_\{i\\overline\{j\}\}\(\\theta\)=h\_\{i\\overline\{j\}\}\(\\theta\)\+\\lambda\\delta\_\{i\\overline\{j\}\}\. In Appendix[10](https://arxiv.org/html/2608.19584#S10), we derive a result that is reminiscent of this although is not a global bound, but[10](https://arxiv.org/html/2608.19584#S10)contradicts our geometric setup as discussed due to the eigenvalue singularity\. Indeed, this proof is valid under the low\-width regime, and is not immune to metric collapse\. Therefore, in much of our work, we will typically work with the Calabi\-Yau condition where a subset of eigenvalues are either really small or really large
\(M∗,ω∗\)∈\{\(M,ω\)\|Ric≡0⇔\(i∂∂¯Φ\)K−κdV0=0,0≲μ<λmin≪1,λmax≤λ≫1∗<∞\}\.\\displaystyle\(M^\{\*\},\\omega^\{\*\}\)\\in\\Bigg\\\{\(M,\\omega\)\\Bigg\|\\text\{Ric\}\\equiv 0\\iff\(i\\partial\\overline\{\\partial\}\\Phi\)^\{K\}\-\\kappa dV\_\{0\}=0,0\\lesssim\\mu<\\lambda\_\{\\text\{min\}\}\\ll 1,\\lambda\_\{\\text\{max\}\}\\leq\\lambda\_\{\\gg 1\}^\{\*\}<\\infty\\Bigg\\\}\.\(4\.24\)
Convexity assumptions\.From this theorem, it follows that the eigenvalue collapse is a consequence of overparameterization and not the Calabi\-Yau artifact\. The eigenvalue explosion is the consequence of the Calabi\-Yau feature\. Strong convexity necessitates∇ω2L=J†J\+ℋnet⪰μI\\nabla^\{2\}\_\{\\omega\}L=J^\{\\dagger\}J\+\\mathcal\{H\}\_\{\\text\{net\}\}\\succeq\\mu IwhereJ†JJ^\{\\dagger\}Jis the contribution from the metric, andℋnet\\mathcal\{H\}\_\{\\text\{net\}\}is intrinsic contribution from the network\. A nondegenerate metric is not sufficient to ensure strong convexity\. Therefore, it is most reasonable for us to assume the regularized Fisher metric \(in all cases, not just the Calabi\-Yau\)
ℱreg=ℱ\+λI,λmin\(ℱreg\)≥μ\>0\\displaystyle\\mathcal\{F\}\_\{\\text\{reg\}\}=\\mathcal\{F\}\+\\lambda I,\\quad\\lambda\_\{\\text\{min\}\}\(\\mathcal\{F\}\_\{\\text\{reg\}\}\)\\geq\\mu\>0\(4\.25\)has eigenvalues bounded below by a positive value \(see[22](https://arxiv.org/html/2608.19584#bib.bib36)[12](https://arxiv.org/html/2608.19584#bib.bib37)for relevant literature\), which derives from∇ω1,1ℒ\(θ\)\+λ∇ω1,1θ†θ\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\(\\theta\)\+\\lambda\\nabla^\{1,1\}\_\{\\omega\}\\theta^\{\\dagger\}\\theta\. This equation is a direct consequence of[4\.23](https://arxiv.org/html/2608.19584#S4.E23)\. We will employ this assumption in Appendix[10](https://arxiv.org/html/2608.19584#S10)rather than assuming strong convexity\. In general,i∂∂¯L⪰μIi\\partial\\overline\{\\partial\}L\\succeq\\mu Iis considered unreasonable in deep learning theory globally\.[10](https://arxiv.org/html/2608.19584#bib.bib1)does not assume a statici∂∂¯ℒ⪰μIi\\partial\\overline\{\\partial\}\\mathcal\{L\}\\succeq\\mu Iresult and this is a consequence of their restricted strong convexity\. In Appendix[10](https://arxiv.org/html/2608.19584#S10), we will prove a convexity result\.
Rank deficiencies and sufficiencies in cohomology, and relations to Theorem 1\.Consider the overparameterization regime and a low\-rank metric, and moreover consider the network’s prediction as a morphism of sheaves\. LetℰM⊕K\\mathcal\{E\}\_\{M\}^\{\\oplus K\}be the sheaf of smooth complex\-valued sections333Complex\-valued sections are used sinceffis real\-valued and not holomorphic\.of the trivial vector bundle of all parameters; andℰM⊕Nk\\mathcal\{E\}\_\{M\}^\{\\oplus Nk\}be the sheaf of sections of the trivial bundle of network outputs over the dataset, i\.e\. redundant effects offfon the data\. The complexified network JacobianJ=∂f\(z,θ\)J=\\partial f\(z;\\theta\)is a morphism of sheaves
J:ℰM⊕K⟶ℰM⊕Nk,\\displaystyle J:\\mathcal\{E\}\_\{M\}^\{\\oplus K\}\\longrightarrow\\mathcal\{E\}\_\{M\}^\{\\oplus Nk\},\(4\.26\)since the Jacobian has a linear/matrix representation mapping fromℰM⊕K\\mathcal\{E\}\_\{M\}^\{\\oplus K\}toℰM⊕Nk\\mathcal\{E\}\_\{M\}^\{\\oplus Nk\}\. On the zero\-loss manifold,JJis surjective because of overparameterization\. This gives us a short exact sequence of sheaves0⟶𝒱→𝜄ℰM⊕K→𝐽Im\(J\)⟶00\\longrightarrow\\mathcal\{V\}\\xrightarrow\{\\iota\}\\mathcal\{E\}\_\{M\}^\{\\oplus K\}\\xrightarrow\{J\}\\text\{Im\}\(J\)\\longrightarrow 0, whereIm\(J\)⊆ℰM⊕Nk\\text\{Im\}\(J\)\\subseteq\\mathcal\{E\}\_\{M\}^\{\\oplus Nk\}\. Here,𝒱=ker\(J\)\\mathcal\{V\}=\\ker\(J\)is the subsheaf that defines the flat minima\. Physically, the fibers of the associated vector bundle represent the flat minima, being the directions in parameter space where the information metrich=𝔼\[J†J\]h=\\mathbb\{E\}\[J^\{\\dagger\}J\]has zero eigenvalues\. In particular, a fiber is the vector space
Vθ=\{δθ∈ℂK\|𝔼‖Jθδθ‖2=0\}\.\\displaystyle V\_\{\\theta\}=\\\{\\delta\\theta\\in\\mathbb\{C\}^\{K\}\\ \|\\ \\mathbb\{E\}\\ \\\|J\_\{\\theta\}\\delta\\theta\\\|^\{2\}=0\\\}\.\(4\.27\)hhtakes this form since−logp\(y\|z,θ\)=12∑k\(fk\(z,θ\)−yk\)\(fk\(z,θ\)−yk\)¯\+C\-\\log p\(y\|z,\\theta\)=\\frac\{1\}\{2\}\\sum\_\{k\}\(f\_\{k\}\(z;\\theta\)\-y\_\{k\}\)\\overline\{\(f\_\{k\}\(z;\\theta\)\-y\_\{k\}\)\}\+C\. Expanding the second derivative yields∂i∂j¯\[\(f−y\)\(f−y\)¯\]=\(∂if\)\(∂j¯f¯\)\+\(∂j¯f\)\(∂if¯\)\+\(f−y\)∂i∂j¯f¯\+\(f−y\)¯∂i∂j¯f\\partial\_\{i\}\\partial\_\{\\overline\{j\}\}\\left\[\(f\-y\)\\overline\{\(f\-y\)\}\\right\]=\(\\partial\_\{i\}f\)\(\\partial\_\{\\overline\{j\}\}\\overline\{f\}\)\+\(\\partial\_\{\\overline\{j\}\}f\)\(\\partial\_\{i\}\\overline\{f\}\)\+\(f\-y\)\\partial\_\{i\}\\partial\_\{\\overline\{j\}\}\\overline\{f\}\+\\overline\{\(f\-y\)\}\\partial\_\{i\}\\partial\_\{\\overline\{j\}\}f\. When taking the expectation overy∼q\(y\|z\)y\\sim q\(y\|z\), the residual\(f−y\)\(f\-y\)is mean\-zero, causing the second\-order derivative terms to vanish\. Therefore,hij¯\(θ\)=𝔼z∼pdata\[𝔼y∼q\(y\|z\)\[\(J†J\)ij¯\]\]h\_\{i\\overline\{j\}\}\(\\theta\)=\\mathbb\{E\}\_\{z\\sim p\_\{\\text\{data\}\}\}\\left\[\\mathbb\{E\}\_\{y\\sim q\(y\|z\)\}\\left\[\(J^\{\\dagger\}J\)\_\{i\\overline\{j\}\}\\right\]\\right\]\. The fiber as in[4\.27](https://arxiv.org/html/2608.19584#S4.E27)corresponds to a minima since for a specific eigenvectorv∈Vθv\\in V\_\{\\theta\},hv=𝔼\[J†J\]v=𝔼\[J†\(Jv\)\]=𝔼\[0\]=0hv=\\mathbb\{E\}\[J^\{\\dagger\}J\]v=\\mathbb\{E\}\[J^\{\\dagger\}\(Jv\)\]=\\mathbb\{E\}\[0\]=0\. In other words, since the prediction does not change, it is a minima\.
Since the network’s prediction over the datasetf=\(f1,…,fNk\):M→ℝNkf=\(f\_\{1\},\\dots,f\_\{Nk\}\):M\\to\\mathbb\{R\}^\{Nk\}is a smooth, real\-valued map as in[3](https://arxiv.org/html/2608.19584#S3), it is not holomorphic\. Therefore, the real differential acts as a bundle map on the real tangent bundleTℝMT\_\{\\mathbb\{R\}\}M, which has rank2K2K
df:TℝM⟶ℝ¯Nk,dfθ\(v\)=\(\(df1\)θ\(v\),…,\(dfNk\)θ\(v\)\)\.\\displaystyle df:T\_\{\\mathbb\{R\}\}M\\longrightarrow\\underline\{\\mathbb\{R\}\}^\{Nk\},\\qquad df\_\{\\theta\}\(v\)=\\big\(\(df\_\{1\}\)\_\{\\theta\}\(v\),\\dots,\(df\_\{Nk\}\)\_\{\\theta\}\(v\)\\big\)\.\(4\.28\)Assumingdfdfmaintains constant rankr=Nkr=Nkacross the zero\-loss locus due to overparameterization, the flat directions form a smooth real vector subbundle𝒱:=kerℝ\(df\)⊆TℝM\\mathcal\{V\}:=\\ker\_\{\\mathbb\{R\}\}\(df\)\\subseteq T\_\{\\mathbb\{R\}\}M\. Taking the sheaf of smooth sectionsℰM\\mathcal\{E\}\_\{M\}, we obtain a short exact sequence of sheaves of smooth real vector bundles
0⟶ℰM\(𝒱\)⟶ℰM\(TℝM\)→dfℰM\(ℝ¯Nk\)⟶0\.\\displaystyle 0\\longrightarrow\\mathcal\{E\}\_\{M\}\(\\mathcal\{V\}\)\\longrightarrow\\mathcal\{E\}\_\{M\}\(T\_\{\\mathbb\{R\}\}M\)\\ \\xrightarrow\{\\ df\\ \}\\ \\mathcal\{E\}\_\{M\}\(\\underline\{\\mathbb\{R\}\}^\{Nk\}\)\\longrightarrow 0\.\(4\.29\)The sheaf of smooth sections of any real vector bundle admits partitions of unity\. In sheaf theory, this property makesℰM\\mathcal\{E\}\_\{M\}a fine sheaf[11](https://arxiv.org/html/2608.19584#bib.bib74), and fine sheaves are acyclic[76](https://arxiv.org/html/2608.19584#bib.bib77), meaningHi\(M,ℰM\(⋅\)\)=0H^\{i\}\(M;\\mathcal\{E\}\_\{M\}\(\\cdot\)\)=0for alli\>0i\>0\. BecauseH1\(M,ℰM\(𝒱\)\)=0H^\{1\}\(M;\\mathcal\{E\}\_\{M\}\(\\mathcal\{V\}\)\)=0automatically, the long exact sequence in cohomology splits trivially\. This yields a surjective map on the global sections with no connecting homomorphism or higher cohomological obstructions[18](https://arxiv.org/html/2608.19584#bib.bib78)[50](https://arxiv.org/html/2608.19584#bib.bib34)
0\{\\lx@inpgf@ignorespaces 0\}H0\(M,ℰM\(𝒱\)\)\{\\lx@inpgf@ignorespaces H^\{0\}\(M;\\mathcal\{E\}\_\{M\}\(\\mathcal\{V\}\)\)\}H0\(M,ℰM\(TℝM\)\)\{\\lx@inpgf@ignorespaces H^\{0\}\(M;\\mathcal\{E\}\_\{M\}\(T\_\{\\mathbb\{R\}\}M\)\)\}H0\(M,ℰM\(ℝ¯Nk\)\)\{\\lx@inpgf@ignorespaces H^\{0\}\(M;\\mathcal\{E\}\_\{M\}\(\\underline\{\\mathbb\{R\}\}^\{Nk\}\)\)\}Hi≥1\(M,ℰM\(⋅\)\)=0\{\\lx@inpgf@ignorespaces H^\{i\\geq 1\}\(M;\\mathcal\{E\}\_\{M\}\(\\cdot\)\)=0\}…\{\\lx@inpgf@ignorespaces\\dots\}df∗\\scriptstyle\{\\lx@inpgf@ignorespaces df\_\{\*\}\}\!δ\\scriptstyle\{\\lx@inpgf@ignorespaces\!\\delta\}
Locally, the subsheaf𝒱=ker\(J\)\\mathcal\{V\}=\\ker\(J\)captures two distinct sources of degeneracy: from overparameterization and from symmetries belonging to neural network architectures\. For example, modReLU networks with activationsf\(z\)=σ\(\|z\|\)eiarg\(z\)f\(z\)=\\sigma\(\|z\|\)e^\{i\\arg\(z\)\}444σ\\sigmais a nonlinear activation\.\(ReLU has no canonical existence in the complex numbers, although complex ReLU can be defined\) have an equivariance quality
f\(cz\)=f\(eiϕz\)=σ\(\|eiϕz\|\)eiarg\(eiϕz\)=σ\(\|z\|\)ei\(arg\(z\)\+ϕ\)=eiϕ\(σ\(\|z\|\)eiarg\(z\)\)=cf\(z\)\.\\displaystyle f\(cz\)=f\(e^\{i\\phi\}z\)=\\sigma\(\|e^\{i\\phi\}z\|\)e^\{i\\arg\(e^\{i\\phi\}z\)\}=\\sigma\(\|z\|\)e^\{i\(\\arg\(z\)\+\\phi\)\}=e^\{i\\phi\}\\big\(\\sigma\(\|z\|\)e^\{i\\arg\(z\)\}\\big\)=cf\(z\)\.\(4\.30\)We will elaborate more on why this is useful in the context of de Rham cohomology\. Moreover, the existence of the non\-trivial subsheaf𝒱\\mathcal\{V\}provides the basis for Theorem 1\. The unregularized information metric is constructed viah=𝔼J†Jh=\\mathbb\{E\}J^\{\\dagger\}Jand is degenerate along the fibers associated to𝒱\\mathcal\{V\}\. This degeneracy preventshhfrom being positive\-definite, and recall from the remark of Theorem 1 it therefore cannot be Calabi\-Yau\.
Figure 4:We plot the integrated arc length∫𝔼‖θ˙‖h𝑑s\\int\\mathbb\{E\}\\\|\\dot\{\\theta\}\\\|\_\{h\}dstimeswidth\\sqrt\{\\text\{width\}\}versus training steps across 5 trajectories per width with a high learning rate ofγ=0\.35\\gamma=0\.35\. We can note the narrowly\-parameterized networks dissipate in high training steps, while wide networks stay straight\. This demonstrates an effect of feature learning: largemmis in the neural tangent kernel regime, or the "lazy" regime[38](https://arxiv.org/html/2608.19584#bib.bib35), and the narrow networks are underparameterized and in the process of feature learning, or the "rich" regime\. The narrow networks fall off because they adapt their features to finding a more efficient minima\.Near\-zero spectral gap under regularization\.The cohomological vanishing in[4\.3](https://arxiv.org/html/2608.19584#S4.SS3)is exact and unconditional; it holds becauseℰM\(𝒱\)\\mathcal\{E\}\_\{M\}\(\\mathcal\{V\}\)is a fine sheaf, independent of any metric\. Regularization does not modify this: for everyλ\>0\\lambda\>0,H1\(M,ℰM\(𝒱\)\)=0H^\{1\}\(M;\\mathcal\{E\}\_\{M\}\(\\mathcal\{V\}\)\)=0still holds by the same argument\. What regularization changes is not the cohomology but the fiberwise eigenvalue spectrum of the information metric, which we make precise here\. NoteMMis non\-compact, as we outlined in the beginning of section[4\.2](https://arxiv.org/html/2608.19584#S4.SS2), and no Hodge\-theoretic argument is available or needed\. The statements below are pointwise\.
Denoteδ\\deltaonℰM⊕K\\mathcal\{E\}\_\{M\}^\{\\oplus K\}the non\-degenerate baseline, i\.e\. oftentimes the flat Euclidean metric under the loss of[4\.23](https://arxiv.org/html/2608.19584#S4.E23), so
hλ\(θ\):=h\(θ\)\+λδ\(θ\),λ\>0\.\\displaystyle h\_\{\\lambda\}\(\\theta\):=h\(\\theta\)\+\\lambda\\delta\(\\theta\),\\quad\\lambda\>0\.\(4\.31\)Sinceδ\\deltais positive\-definite andhhcorresponds to that unregularized and is positive semi\-definite,hλh\_\{\\lambda\}is positive\-definite for everyλ\>0\\lambda\>0, andhλ→hh\_\{\\lambda\}\\to hasλ→0\\lambda\\to 0\. SinceVθ=kerh\(θ\)V\_\{\\theta\}=\\ker h\(\\theta\)andh\(θ\)h\(\\theta\)is Hermitian positive semi\-definite, the orthogonal projectorΠθ\\Pi\_\{\\theta\}ontoVθV\_\{\\theta\}satisfiesh\(θ\)Πθ=Πθh\(θ\)=0h\(\\theta\)\\Pi\_\{\\theta\}=\\Pi\_\{\\theta\}h\(\\theta\)=0, i\.e\.h\(θ\)=Πθ⟂h\(θ\)Πθ⟂h\(\\theta\)=\\Pi\_\{\\theta\}^\{\\perp\}h\(\\theta\)\\Pi\_\{\\theta\}^\{\\perp\}whereΠθ⟂=I−Πθ\\Pi\_\{\\theta\}^\{\\perp\}=I\-\\Pi\_\{\\theta\}\. Writinghλ\(θ\)h\_\{\\lambda\}\(\\theta\)in block form with respect to the orthogonal decompositionℂK=Vθ⊕Vθ⟂\\mathbb\{C\}^\{K\}=V\_\{\\theta\}\\oplus V\_\{\\theta\}^\{\\perp\},
hλ\(θ\)=\(λΠθδ\(θ\)ΠθλΠθδ\(θ\)Πθ⟂λΠθ⟂δ\(θ\)Πθh\(θ\)\|Vθ⟂\+λΠθ⟂δ\(θ\)Πθ⟂\)Vθ⊕Vθ⟂\.\\displaystyle h\_\{\\lambda\}\(\\theta\)=\\begin\{pmatrix\}\\lambda\\Pi\_\{\\theta\}\\delta\(\\theta\)\\Pi\_\{\\theta\}&\\lambda\\Pi\_\{\\theta\}\\delta\(\\theta\)\\Pi\_\{\\theta\}^\{\\perp\}\\\\\[4\.0pt\] \\lambda\\Pi\_\{\\theta\}^\{\\perp\}\\delta\(\\theta\)\\Pi\_\{\\theta\}&h\(\\theta\)\\big\|\_\{V\_\{\\theta\}^\{\\perp\}\}\+\\lambda\\Pi\_\{\\theta\}^\{\\perp\}\\delta\(\\theta\)\\Pi\_\{\\theta\}^\{\\perp\}\\end\{pmatrix\}\_\{V\_\{\\theta\}\\oplus V\_\{\\theta\}^\{\\perp\}\}\.\(4\.32\)At eachθ\\thetaon the zero\-loss manifold, both are Hermitian forms on the same finite\-dimensional fiberℂK\\mathbb\{C\}^\{K\}, so we may compare their eigenvalues directly via the Courant\-Fischer minimax characterization[48](https://arxiv.org/html/2608.19584#bib.bib76)\. Writingλ1\(θ\)≤⋯≤λK\(θ\)\\lambda\_\{1\}\(\\theta\)\\leq\\cdots\\leq\\lambda\_\{K\}\(\\theta\)for the eigenvalues ofh\(θ\)h\(\\theta\)andλ1λ\(θ\)≤⋯≤λKλ\(θ\)\\lambda\_\{1\}^\{\\lambda\}\(\\theta\)\\leq\\cdots\\leq\\lambda\_\{K\}^\{\\lambda\}\(\\theta\)for those ofhλ\(θ\)h\_\{\\lambda\}\(\\theta\), Weyl’s inequality gives
λi\(θ\)≤λiλ\(θ\)≤λi\(θ\)\+λ‖δ\(θ\)‖2\.\\displaystyle\\lambda\_\{i\}\(\\theta\)\\leq\\lambda\_\{i\}^\{\\lambda\}\(\\theta\)\\leq\\lambda\_\{i\}\(\\theta\)\+\\lambda\\\|\\delta\(\\theta\)\\\|\_\{2\}\.\(4\.33\)Recall from[4\.27](https://arxiv.org/html/2608.19584#S4.E27)thatVθ=kerh\(θ\)V\_\{\\theta\}=\\ker h\(\\theta\)has dimensiond\(θ\):=dimℂVθd\(\\theta\):=\\dim\_\{\\mathbb\{C\}\}V\_\{\\theta\}, soλ1\(θ\)=⋯=λd\(θ\)\(θ\)=0\\lambda\_\{1\}\(\\theta\)=\\cdots=\\lambda\_\{d\(\\theta\)\}\(\\theta\)=0\. The inequality above then forces
0<λiλ\(θ\)≤λ‖δ\(θ\)‖2fori=1,…,d\(θ\),\\displaystyle 0<\\lambda\_\{i\}^\{\\lambda\}\(\\theta\)\\leq\\lambda\\\|\\delta\(\\theta\)\\\|\_\{2\}\\quad\\text\{for \}i=1,\\ldots,d\(\\theta\),\(4\.34\)so the bottomd\(θ\)d\(\\theta\)eigenvalues ofhλh\_\{\\lambda\}vanish linearly inλ\\lambda, uniformly on any compact subset ofMM\.
Parameter holes and connections to de Rham cohomology\.As illustrated in[4\.30](https://arxiv.org/html/2608.19584#S4.E30), neural networks can possess intrinsic qualities and symmetries\. Specifically, incoming and outgoing weights can be rotated by a phase without altering the network’s output or the loss\. Because this equivalence holds for any phase angleϕ\\phi, the symmetry space forms the groupU\(1\)≅S1U\(1\)\\cong S^\{1\}\. If we consider the parameter plane of a complex weight, the origin where the phase becomes undefined acts as a puncture\. De Rham cohomology is a tool to detect holes in the space[55](https://arxiv.org/html/2608.19584#bib.bib70), since integrating an exact form around a closed path is zero\. Letφ\\varphirepresent the angular coordinate wrapping around this puncture\. Whiledφd\\varphiis closed \(d2φ=0d^\{2\}\\varphi=0\), it is not exact, generating a non\-trivial cohomology class\[dφ\]∈HdR1\(M,ℝ\)\[d\\varphi\]\\in H\_\{dR\}^\{1\}\(M;\\mathbb\{R\}\), and so we develop
0\{\\lx@inpgf@ignorespaces 0\}HdR0\(M,ℝ\)\{\\lx@inpgf@ignorespaces H^\{0\}\_\{\\text\{dR\}\}\(M;\\mathbb\{R\}\)\}HdR0\(U1,ℝ\)⊕HdR0\(U2,ℝ\)\{\\lx@inpgf@ignorespaces H^\{0\}\_\{\\text\{dR\}\}\(U\_\{1\};\\mathbb\{R\}\)\\oplus H^\{0\}\_\{\\text\{dR\}\}\(U\_\{2\};\\mathbb\{R\}\)\}HdR0\(U1∩U2,ℝ\)\{\\lx@inpgf@ignorespaces H^\{0\}\_\{\\text\{dR\}\}\(U\_\{1\}\\cap U\_\{2\};\\mathbb\{R\}\)\}HdR1\(M,ℝ\)\{\\lx@inpgf@ignorespaces H^\{1\}\_\{\\text\{dR\}\}\(M;\\mathbb\{R\}\)\}HdR1\(U1,ℝ\)⊕HdR1\(U2,ℝ\)=0\{\\lx@inpgf@ignorespaces H^\{1\}\_\{\\text\{dR\}\}\(U\_\{1\};\\mathbb\{R\}\)\\oplus H^\{1\}\_\{\\text\{dR\}\}\(U\_\{2\};\\mathbb\{R\}\)=0\}…\{\\lx@inpgf@ignorespaces\\dots\}resd∗\\scriptstyle\{\\color\[rgb\]\{1,0\.04,0\.61\}\\lx@inpgf@ignorespaces d^\{\*\}\}
for two overlapping, simply connected open setsU1U\_\{1\}andU2U\_\{2\},U1∩U2U\_\{1\}\\cap U\_\{2\}disconnected\. By contrast, applying the Hodge star maps the angular form to the exact radial form⋆dφ=−d\(logr\)\\star d\\varphi=\-d\(\\log r\)\. Therefore, integrating over a loopγ\\gammaand a closed contourΣ\\Sigmaenclosing the singularity yields
∮γdφ=2π≠0,∮Σ⋆dφ=∮Σ−d\(logr\)=0\.\\displaystyle\\oint\_\{\\gamma\}d\\varphi=2\\pi\\neq 0,\\quad\\oint\_\{\\Sigma\}\\star d\\varphi=\\oint\_\{\\Sigma\}\-d\(\\log r\)=0\.\(4\.35\)
Sectional curvature collapse under metric collapse\.When this metric collapseshϵ,ϵ→0h\_\{\\epsilon\},\\epsilon\\to 0, induced is a lower bound on sectional curvatureKϵ\(u,v\)=hϵ\(Rϵ\(u,v\)v,u\)\(hϵ\(u,u\)hϵ\(v,v\)−hϵ\(u,v\)2\)−1K\_\{\\epsilon\}\(u,v\)=h\_\{\\epsilon\}\(R\_\{\\epsilon\}\(u,v\)v,u\)\(h\_\{\\epsilon\}\(u,u\)h\_\{\\epsilon\}\(v,v\)\-h\_\{\\epsilon\}\(u,v\)^\{2\}\)^\{\-1\}with respect to a quotient spaceX∞=M/∼X\_\{\\infty\}=M/\\sim\. The quotient spaceX∞X\_\{\\infty\}becomes an Alexandrov space[3](https://arxiv.org/html/2608.19584#bib.bib64)with a lower bound on sectional curvatureKϵ≥κK\_\{\\epsilon\}\\geq\\kappa\. We can note given a Kähler formωϵ=ihϵ,kj¯dθk∧dθ¯j\\omega\_\{\\epsilon\}=ih\_\{\\epsilon,k\\overline\{j\}\}d\\theta^\{k\}\\wedge d\\overline\{\\theta\}^\{j\}withRic\(ωϵ\)=−i∂∂¯logdet\(hϵ,kj¯\)≡0\\text\{Ric\}\(\\omega\_\{\\epsilon\}\)=\-i\\partial\\overline\{\\partial\}\\log\\det\(h\_\{\\epsilon,k\\overline\{j\}\}\)\\equiv 0, the sectional curvatureKϵK\_\{\\epsilon\}does not necessarily vanish\.
Connections to PSH collapse\.Non\-strictlyPSH\(U\)\\text\{PSH\}\(U\)potentials are acceptable because the collapse is local, i\.e\. a set of measure zero, not global, when the strict preservation fails\. In an overparameterization, the collapse may be for all neural network input, hence global not local\.
Additional remarks\.Under an overparameterization collapse, The geodesic distance along a direction in the nullspace𝒩h=\{v∈TpM∣h\(v,v\)=0\}\\mathcal\{N\}\_\{h\}=\\\{v\\in T\_\{p\}M\\mid h\(v,v\)=0\\\}will maintain identically vanishing geodesic distance, so a geodesic ball region collapses with respect to dimension\. The inverse of the exponential mapexpp−1\\exp\_\{p\}^\{\-1\}possesses similar properties, and will collapse in norm, so the cosine similarity is ill\-defined\. The spectral norm‖h‖2\\\|h\\\|\_\{2\}, on the other hand, is immune to rank deficiency and is immune to collapse\.
Regularizing and resolving profaned curvature\.To regularize the loss landscape and support well\-behaved curvature, we can penalize the loss further with
ℒ^=ℒ\+λθ†θ−αlogdet\(h\(θ\)\)⇔αρ⪰−\(i∂∂¯ℒ\+λωflat\),\\displaystyle\\widehat\{\\mathcal\{L\}\}=\\mathcal\{L\}\+\\lambda\\theta^\{\\dagger\}\\theta\-\\alpha\\log\\det\(h\(\\theta\)\)\\quad\\iff\\quad\\alpha\\rho\\succeq\-\(i\\partial\\overline\{\\partial\}\\mathcal\{L\}\+\\lambda\\omega\_\{\\text\{flat\}\}\),\(4\.36\)since the Hessian of the loss⪰0\\succeq 0at a local minima\. Again, we set the loss to be the negative log\-likelihood ofpp, so−logp\-\\log p\. Here,ρ\\rhois the Ricci form\. We can control the positive\-definieness of the Ricci form based onα,λ\\alpha,\\lambda, and the loss, therefore regularizing the Ricci curvature of the manifold with thelogdet\\log\\detpenalty\. We can note the negative sign in[4\.36](https://arxiv.org/html/2608.19584#S4.E36)provides a lower bound as in[4\.36](https://arxiv.org/html/2608.19584#S4.E36)and facilitates positive curvature\.
## 5Theoretical results
In our theoretical results, we will typically make the following assumptions \(unless stated otherwise\)\.
1. I\.\(M,ω\)\(M,\\omega\)is the Kähler information manifold of[4\.1](https://arxiv.org/html/2608.19584#S4.E1)of dimensionKKwith metrichh\.
2. II\.The parameter descent obeys natural gradient descentθ˙i=−hij¯∂j¯ℒ\\dot\{\\theta\}^\{i\}=\-h^\{i\\overline\{j\}\}\\partial\_\{\\overline\{j\}\}\\mathcal\{L\}\.
3. III\.hhis full rank by the regularized loss of[4\.23](https://arxiv.org/html/2608.19584#S4.E23)\. Moreover, the minimum eigenvalue is bounded belowλmin\>μ\\lambda\_\{\\text\{min\}\}\>\\mu\. In order for the nuclear norm to not pick up a factor of dimension, which we need in Theorem 2, we assumeμ=𝒪\(1/K2\)\\mu=\\mathcal\{O\}\(1/K^\{2\}\)when necessary\. To enforce this, we can considerℒ\(θ\)~=ℒ\(θ\)\+λ0K2θ†θ\\widetilde\{\\mathcal\{L\}\(\\theta\)\}=\\mathcal\{L\}\(\\theta\)\+\\frac\{\\lambda\_\{0\}\}\{K^\{2\}\}\\theta^\{\\dagger\}\\theta\.
4. IV\.The Calabi\-Yau constant determinant condition\(i∂∂¯Φ\)K=κdV0\(i\\partial\\overline\{\\partial\}\\Phi\)^\{K\}=\\kappa dV\_\{0\}, i\.e\.det\(∂i∂j¯Φ\)≡constant\\det\(\\partial\_\{i\}\\partial\_\{\\overline\{j\}\}\\Phi\)\\equiv\\text\{constant\}is with respect to a constant sufficiently large, so not all of the eigenvalues viadet\(∂i∂j¯Φ\)=∏iλi\\det\(\\partial\_\{i\}\\partial\_\{\\overline\{j\}\}\\Phi\)=\\prod\_\{i\}\\lambda\_\{i\}are very small\.
5. V\.𝒮\\mathcal\{S\}is a \(compact\) geodesic ball on the Kähler manifold sufficiently close to initialization\.
6. VI\.The manifoldMMis not compact \(𝒮\\mathcal\{S\}still exists even ifMMis not compact\)\.
7. VII\.The maximum eigenvalue of a Calabi\-Yau metric can be assumed to followλmax\(h\)=Ω\(κμK~−1\)\\lambda\_\{\\max\}\(h\)=\\Omega\\left\(\\frac\{\\kappa\}\{\\mu^\{\\widetilde\{K\}\-1\}\}\\right\), whereκ\\kappais a determinant constant andμ\\muis a lower bound of the minimum metric eigenvalue and0≪K~≤K0\\ll\\widetilde\{K\}\\leq Kis some constant\.
### 5\.1Second derivative results
Theorem 2\.LetMMbe a Kähler manifold with dimensionKKand Kähler metricω\\omegawithhhfull rank\. Letf\(θ,z\):M→ℝf\(\\theta;z\):M\\rightarrow\\mathbb\{R\}be a function such that the deformed metricωf=ω\+i∂∂¯f\>0\\omega\_\{f\}=\\omega\+i\\partial\\overline\{\\partial\}f\>0\. Assume that the spectral norms‖∂3f‖2\\\|\\partial^\{3\}f\\\|\_\{2\}and‖i∂∂¯f‖22\\\|i\\partial\\overline\{\\partial\}f\\\|\_\{2\}^\{2\}are𝒪\(1\)\\mathcal\{O\}\(1\)or𝒪\(m−α\)\\mathcal\{O\}\(m^\{\-\\alpha\}\),α≥0\\alpha\\geq 0\. Assume the hypotheses of Lemma 1 and Lemma 2 hold\. Assume the nuclear norm∥⋅∥1≤constant⋅∥⋅∥2\\\|\\cdot\\\|\_\{1\}\\leq\\text\{constant\}\\cdot\\\|\\cdot\\\|\_\{2\}does not pick up a factor of dimensionKKdue to rapid eigenvalue decay\. Then the spectral norm of the Dolbeault Hessian on𝒮\\mathcal\{S\}is bounded by
supθ∈𝒮‖i∂∂¯f‖2=𝒪\(1m\)\.\\displaystyle\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|i\\partial\\overline\{\\partial\}f\\\|\_\{2\}=\\mathcal\{O\}\\left\(\\frac\{1\}\{\\sqrt\{m\}\}\\right\)\.\(5\.1\)
Theorem 2 is one of our main results, and is consistent with the scaling of[10](https://arxiv.org/html/2608.19584#bib.bib1)\. To the best of our knowledge, this bound is sharp, consistent with Figure[2](https://arxiv.org/html/2608.19584#S4.F2)\. This result is unique since it is specifically for Dolbeault asymptotics and complex networks\.Sketch of proof\.The proof is an application of the Mean Value Theorem applied to a logarithmic volume ratio\. The remainder of the proof primarily follows from applications of Lemma 1, Lemma 2, and a Taylor expansion\.
Lemma 1\.Letf\(θ,z\):M→ℝf\(\\theta;z\):M\\to\\mathbb\{R\}be as in[8\.2](https://arxiv.org/html/2608.19584#S8.SS2)\. Assume the \(2,0\) Hessian obeys a bound with respect to the \(1,1\) Hessian‖∂ki2f‖2≤C′‖i∂∂¯f‖2\\\|\\partial^\{2\}\_\{ki\}f\\\|\_\{2\}\\leq C^\{\\prime\}\\\|i\\partial\\overline\{\\partial\}f\\\|\_\{2\}\. Then the Dolbeault Hessian obeys the property
sup‖v‖2=1‖∇v\(i∂∂¯f\)‖2≤‖∂3f‖2\+C‖i∂∂¯f‖2supθ∈𝒮‖i∂∂¯f‖2,\\displaystyle\\sup\_\{\\\|v\\\|\_\{2\}=1\}\\\|\\nabla\_\{v\}\(i\\partial\\overline\{\\partial\}f\)\\\|\_\{2\}\\leq\\\|\\partial^\{3\}f\\\|\_\{2\}\+C\\\|i\\partial\\overline\{\\partial\}f\\\|\_\{2\}\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|i\\partial\\overline\{\\partial\}f\\\|\_\{2\},\(5\.2\)where‖∂3f‖2:=sup‖v‖2=1‖vk∂kℋ‖2\\\|\\partial^\{3\}f\\\|\_\{2\}:=\\sup\_\{\\\|v\\\|\_\{2\}=1\}\\\|v^\{k\}\\partial\_\{k\}\\mathcal\{H\}\\\|\_\{2\}\.
Lemma 2\.Assume the maximum eigenvalue of the Hessianℋ\(θ\)=i∂∂¯f\(θ\)\\mathcal\{H\}\(\\theta\)=i\\partial\\overline\{\\partial\}f\(\\theta\)is finite over𝒮\\mathcal\{S\}\. Suppose for test functionv\(θt,t\)=log\(λmax\(ℋ\(θt\)\)\)−Ψ\(θt\)v\(\\theta\_\{t\},t\)=\\log\(\\lambda\_\{\\max\}\(\\mathcal\{H\}\(\\theta\_\{t\}\)\)\)\-\\Psi\(\\theta\_\{t\}\), the operator satisfies\(∂t−Δω~\)v<0\(\\partial\_\{t\}\-\\Delta\_\{\\widetilde\{\\omega\}\}\)v<0on𝒮\\mathcal\{S\}\. Given the initialization bound‖ℋ\(θ0\)‖2≤Cinitm\\\|\\mathcal\{H\}\(\\theta\_\{0\}\)\\\|\_\{2\}\\leq\\frac\{C\_\{\\text\{init\}\}\}\{\\sqrt\{m\}\}, there exists a geometric constantCgeom\>0C\_\{\\text\{geom\}\}\>0such that the spectral norm of the Hessian obeys
supθ∈𝒮‖ℋ‖2≤Cgeom‖ℋ\(θ0\)‖2exp\(sup𝒮Ψ−inf𝒮Ψ\)\.\\displaystyle\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|\\mathcal\{H\}\\\|\_\{2\}\\leq C\_\{\\text\{geom\}\}\\\|\\mathcal\{H\}\(\\theta\_\{0\}\)\\\|\_\{2\}\\exp\\left\(\\sup\_\{\\mathcal\{S\}\}\\Psi\-\\inf\_\{\\mathcal\{S\}\}\\Psi\\right\)\.\(5\.3\)
Lemma 3\.Assume the\(1,1\)\(1,1\)Hessian satisfies‖ℋ1,1‖2≤C~‖ℋ2,0‖2\\\|\\mathcal\{H\}^\{1,1\}\\\|\_\{2\}\\leq\\widetilde\{C\}\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}, and the\(2,0\)\(2,0\)Hessian obeys the covariant bound
‖∇ωℋ2,0‖2≤C1‖∂3f‖2\+C2supθ∈𝒮‖ℋ2,0‖22\.\\displaystyle\\\|\\nabla\_\{\\omega\}\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}\\leq C\_\{1\}\\\|\\partial^\{3\}f\\\|\_\{2\}\+C\_\{2\}\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}^\{2\}\.\(5\.4\)If the spectral norm of the initial\(2,0\)\(2,0\)Hessian‖ℋ2,0\(θ0\)‖2\\\|\\mathcal\{H\}^\{2,0\}\(\\theta\_\{0\}\)\\\|\_\{2\}is𝒪\(1m\)\\mathcal\{O\}\\left\(\\frac\{1\}\{\\sqrt\{m\}\}\\right\)and the third derivativeA=supθ∈𝒮‖∂3f‖2A=\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|\\partial^\{3\}f\\\|\_\{2\}is𝒪\(m−k\)\\mathcal\{O\}\(m^\{\-k\}\)wherek≥12k\\geq\\frac\{1\}\{2\}, then the spectral norm of the\(2,0\)\(2,0\)Hessian over the geodesic ball is bounded by
supθ∈𝒮‖ℋ2,0‖2=𝒪\(1m\)\.\\displaystyle\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}=\\mathcal\{O\}\\left\(\\frac\{1\}\{\\sqrt\{m\}\}\\right\)\.\(5\.5\)
### 5\.2Initialization results
Theorem 3\.Letf\(θ0,z\)f\(\\theta\_\{0\};z\)be the output of theLL\-layer neural network of widthmmas in[3](https://arxiv.org/html/2608.19584#S3)exactly at initialization, parameterized byθ0=\{W\(1\),…,W\(L\),v\}\\theta\_\{0\}=\\\{W^\{\(1\)\},\\dots,W^\{\(L\)\},v\\\}where the weights are drawn i\.i\.d\. from a standard complex Gaussian distribution𝒞𝒩\(0,1\)\\mathcal\{CN\}\(0,1\)\. Assume the base inputs are bounded such that‖α\(0\)‖∞=𝒪\(1\)\\\|\\alpha^\{\(0\)\}\\\|\_\{\\infty\}=\\mathcal\{O\}\(1\), the activation functionϕ\(h,h¯\)\\phi\(h,\\overline\{h\}\)has bounded first and second Wirtinger derivatives, and the forward Jacobians have bounded operator norms\. Then, with high probability over the initialization, the spectral norm of the\(1,1\)\(1,1\)parameter Hessianℋ\\mathcal\{H\}off\(θ0\)f\(\\theta\_\{0\}\)is bounded by
‖ℋ\(θ0\)‖2=𝒪\(1m\)\.\\displaystyle\\\|\\mathcal\{H\}\(\\theta\_\{0\}\)\\\|\_\{2\}=\\mathcal\{O\}\\left\(\\frac\{1\}\{\\sqrt\{m\}\}\\right\)\.\(5\.6\)
Corollary\.This result can be modified slightly to give the result for the \(2,0\)\-Hessian as well\.
### 5\.3Convexity results
Theorem 4 and Lemma are very close to results found in[10](https://arxiv.org/html/2608.19584#bib.bib1)in the end result, but the proofs of each are different in spirit, since we are now relying on geometric structure and complex data to complete the proof\.
Theorem 4 \(convexity\)\.LetMMbe Kähler\. Consider the lossL\(θ\)=1n∑i=1nℓi\(yi,fi\(θ\)\)L\(\\theta\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\ell\_\{i\}\(y\_\{i\},f\_\{i\}\(\\theta\)\)parameterized byθ∈M\\theta\\in M\. Assume the following regularity conditions locally: \(i\) the loss function satisfiesℓi′′≥a\\ell\_\{i\}^\{\\prime\\prime\}\\geq afor some constanta\>0a\>0; \(ii\)ℱt\(v\)=2n∑i=1n\(Re\(∇ωfi\(θt\)v\)\)2\\mathcal\{F\}\_\{t\}\(v\)=\\frac\{2\}\{n\}\\sum\_\{i=1\}^\{n\}\\left\(\\text\{Re\}\\left\(\\nabla\_\{\\omega\}f\_\{i\}\(\\theta\_\{t\}\)v\\right\)\\right\)^\{2\}is bounded byμ‖v‖ω2≤ℱt\(v\)≤ρ‖v‖ω2\\mu\\\|v\\\|\_\{\\omega\}^\{2\}\\leq\\mathcal\{F\}\_\{t\}\(v\)\\leq\\rho\\\|v\\\|\_\{\\omega\}^\{2\}for someμ\>0\\mu\>0and maximum eigenvalue boundρ\\rho; \(iii\) the full Hessian norm is bounded byCℋ‖v‖ω2/mC\_\{\\mathcal\{H\}\}\\\|v\\\|\_\{\\omega\}^\{2\}/\\sqrt\{m\}, wheremmis the network width parameter\. Then we have
∇ω2L\(θ~t\)\(v,v\)≥Γ\(a,μ,ρ,Cℋ,\{ℓi′\(fi\(θ~t\)\)\}i,m,v\)‖v‖ω2,\\displaystyle\\nabla^\{2\}\_\{\\omega\}L\(\\widetilde\{\\theta\}\_\{t\}\)\(v,v\)\\geq\\Gamma\\left\(a,\\mu,\\rho,C\_\{\\mathcal\{H\}\},\\\{\\ell\_\{i\}^\{\\prime\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\\\}\_\{i\},m,v\\right\)\\\|v\\\|\_\{\\omega\}^\{2\},\(5\.7\)whereΓ=Γ\(a,μ,ρ,Cℋ,L\(θ~t\),m,v\)\\Gamma=\\Gamma\\left\(a,\\mu,\\rho,C\_\{\\mathcal\{H\}\},L\(\\widetilde\{\\theta\}\_\{t\}\),m,v\\right\)is defined as
Γ:=aμ−2aρCℋ‖v‖ω\+Cℋ1n∑i=1n\(ℓi′\(fi\(θ~t\)\)\)2m\.\\displaystyle\\Gamma:=a\\mu\-\\frac\{2a\\sqrt\{\\rho\}C\_\{\\mathcal\{H\}\}\\\|v\\\|\_\{\\omega\}\+C\_\{\\mathcal\{H\}\}\\sqrt\{\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\(\\ell^\{\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\)^\{2\}\}\}\{\\sqrt\{m\}\}\.\(5\.8\)
Theorem 5 \(β\\beta\-smoothness\)\.LetMMbe Kählerω\\omega\. Consider the lossL\(θ\)=1n∑i=1nℓi\(yi,fi\(θ\)\)L\(\\theta\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\ell\_\{i\}\(y\_\{i\},f\_\{i\}\(\\theta\)\)denote the empirical loss\. Assume the regularity condition that the first\-order Jacobian quadratic form is bounded byρJ\>0\\rho\_\{J\}\>0, such that2n∑i=1nℓi′′\(Re\(∇ωfiv\)\)2≤ρJ‖v‖ω2\\frac\{2\}\{n\}\\sum\_\{i=1\}^\{n\}\\ell^\{\\prime\\prime\}\_\{i\}\\left\(\\text\{Re\}\\left\(\\nabla\_\{\\omega\}f\_\{i\}v\\right\)\\right\)^\{2\}\\leq\\rho\_\{J\}\\\|v\\\|\_\{\\omega\}^\{2\}\. Assume the results of[5\.1](https://arxiv.org/html/2608.19584#S5.SS1)hold\. Then we have
12∇ω2L\(θ~t\)\(v,v\)≤\(ρJ\+Cℋm1n∑i=1n\(ℓi′\(fi\(θ~t\)\)\)2\)‖v‖ω2\.\\displaystyle\\frac\{1\}\{2\}\\nabla^\{2\}\_\{\\omega\}L\(\\widetilde\{\\theta\}\_\{t\}\)\(v,v\)\\leq\\left\(\\rho\_\{J\}\+\\frac\{C\_\{\\mathcal\{H\}\}\}\{\\sqrt\{m\}\}\\sqrt\{\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\(\\ell^\{\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\)^\{2\}\}\\right\)\\\|v\\\|\_\{\\omega\}^\{2\}\.\(5\.9\)
Lemma 4\.LetU⊆𝒮U\\subseteq\\mathcal\{S\}be a local coordinate chart equipped with the Kähler metrichij¯h\_\{i\\overline\{j\}\}, and letθ∗∈U\\theta^\{\*\}\\in Ube a minima of the loss functionLL\. Define the dynamic strong convexity parameter along a geodesicγt\\gamma\_\{t\}as
Γt:=infs∈\[0,1\]λmin\(hik¯\(γt\(s\)\)ℋkj¯\(γt\(s\)\)\),\\displaystyle\\Gamma\_\{t\}:=\\inf\_\{s\\in\[0,1\]\}\\lambda\_\{\\min\}\\left\(h^\{i\\overline\{k\}\}\(\\gamma\_\{t\}\(s\)\)\\mathcal\{H\}\_\{k\\overline\{j\}\}\(\\gamma\_\{t\}\(s\)\)\\right\),\(5\.10\)whereℋkj¯\\mathcal\{H\}\_\{k\\overline\{j\}\}is the Wirtinger Hessian block,γt\(s\)=expθt\(sv\)\\gamma\_\{t\}\(s\)=\\exp\_\{\\theta\_\{t\}\}\(sv\), andv=γ˙t\(0\)=expθt−1\(θ∗\)1,0v=\\dot\{\\gamma\}\_\{t\}\(0\)=\\exp\_\{\\theta\_\{t\}\}^\{\-1\}\(\\theta^\{\*\}\)^\{1,0\}is an initial holomorphic tangent vector\. ThenΓt\\Gamma\_\{t\}is guaranteed to satisfy a strong convexity condition by the result of Appendix[10](https://arxiv.org/html/2608.19584#S10)\. Moreover, ifΓt\>0\\Gamma\_\{t\}\>0, thenLLsatisfies the dynamic Kähler Polyak\-Łojasiewicz condition
infθ∈UL\(θ\)≥L\(θt\)−1Γt\\bBigg@3‖∇h1,0L\(θt\)\\bBigg@3‖h2\.\\displaystyle\\inf\_\{\\theta\\in U\}L\(\\theta\)\\geq L\(\\theta\_\{t\}\)\-\\frac\{1\}\{\\Gamma\_\{t\}\}\\bBigg@\{3\}\\\|\\nabla\_\{h\}^\{1,0\}L\(\\theta\_\{t\}\)\\bBigg@\{3\}\\\|\_\{h\}^\{2\}\.\(5\.11\)
Lemma 6\.Let\(M,ω\)\(M,\\omega\)be a Kähler information manifold with the regularized loss of[4\.23](https://arxiv.org/html/2608.19584#S4.E23), and let the metric eigenvalues be bounded below byλmin≥μ\>0\\lambda\_\{\\min\}\\geq\\mu\>0\. Assume the maximum eigenvalue of the metric obeysλmax\(h\)=Ω\(κμK−1\)\\lambda\_\{\\max\}\(h\)=\\Omega\\left\(\\frac\{\\kappa\}\{\\mu^\{K\-1\}\}\\right\)\. Under the Calabi\-Yau conditiondet\(h\)≡κ\\det\(h\)\\equiv\\kappa, theCℋC\_\{\\mathcal\{H\}\}constant of[10](https://arxiv.org/html/2608.19584#S10)scales asCℋ=Ω\(κμ−\(K\+1\)m−1/2\)C\_\{\\mathcal\{H\}\}=\\Omega\\left\(\\kappa\\mu^\{\-\(K\+1\)\}m^\{\-1/2\}\\right\)\. Consequently, the smoothness parameterβ→∞\\beta\\to\\infty, and for any learning rateηt=Ω\(μK\+1\)\\eta\_\{t\}=\\Omega\(\\mu^\{K\+1\}\), the loss does not descend monotonically \(oscillates\)\.
Lemma 7\.Lethhbe a Calabi\-Yau metric satisfying∏j=1Kλj=κ\\prod\_\{j=1\}^\{K\}\\lambda\_\{j\}=\\kappasuch that its maximum eigenvalue exhibitsλmax\(h\)=Ω\(κ/μK−1\)\\lambda\_\{\\max\}\(h\)=\\Omega\\left\(\\kappa/\\mu^\{K\-1\}\\right\)\. Define the Kähler Polyak\-Łojasiewicz parameter along a pathγt\(s\)\\gamma\_\{t\}\(s\)fors∈\[0,1\]s\\in\[0,1\]asΓt:=infs∈\[0,1\]λmin\(hik¯\(γt\(s\)\)ℋkj¯\(γt\(s\)\)\)\\Gamma\_\{t\}:=\\inf\_\{s\\in\[0,1\]\}\\lambda\_\{\\min\}\(h^\{i\\overline\{k\}\}\(\\gamma\_\{t\}\(s\)\)\\mathcal\{H\}\_\{k\\overline\{j\}\}\(\\gamma\_\{t\}\(s\)\)\)\. ThenΓt\\Gamma\_\{t\}is bounded above byΓt≤𝒪\(μK−1κm\)\\Gamma\_\{t\}\\leq\\mathcal\{O\}\\left\(\\frac\{\\mu^\{K\-1\}\}\{\\kappa\\sqrt\{m\}\}\\right\)\.
Remark\.Lemmas 6, 7 are our primary results that are restricted to Calabi\-Yau manifolds alone\. The results in[5\.4](https://arxiv.org/html/2608.19584#S5.SS4)also applicable to Calabi\-Yau manifolds, but are moreso for manifolds of a range of Ricci curvature\. Lemma 11 in[11\.3](https://arxiv.org/html/2608.19584#S11.SS3)is also a failure mode of Calabi\-Yau manifolds specifically\.
Lemma 8\.Letθ∗\\theta^\{\*\}denote the optimal parameter\. Assumeθt\\theta\_\{t\}evolves according to stochastic natural gradient descentdθt=−∇hℒ\(θt\)dt\+ηdWtd\\theta\_\{t\}=\-\\nabla\_\{h\}\\mathcal\{L\}\(\\theta\_\{t\}\)dt\+\\sqrt\{\\eta\}dW\_\{t\}\. Define the expected regret alongθt\\theta\_\{t\}over interval\[0,T\]\[0,T\]as
𝔼\[ℛ\(T\)\]\\displaystyle\\mathbb\{E\}\[\\mathcal\{R\}\(T\)\]=𝔼θ0𝔼W\[∫0TRe⟨∇hℒ\(θt\),expθt−1\(θ∗\)1,0⟩h𝑑t\]\.\\displaystyle=\\mathbb\{E\}\_\{\\theta\_\{0\}\}\\mathbb\{E\}\_\{W\}\\left\[\\int\_\{0\}^\{T\}\\text\{Re\}\\langle\\nabla\_\{h\}\\mathcal\{L\}\(\\theta\_\{t\}\),\\exp\_\{\\theta\_\{t\}\}^\{\-1\}\(\\theta^\{\*\}\)^\{1,0\}\\rangle\_\{h\}dt\\right\]\.\(5\.12\)Then the regret obeys the lower bound𝔼\[ℛ\(T\)\]≥𝔼\[δ\(θT\)\]−δ\(θ0\)−ηKT−𝔼\[ℰRic\]\\mathbb\{E\}\[\\mathcal\{R\}\(T\)\]\\geq\\mathbb\{E\}\[\\delta\(\\theta\_\{T\}\)\]\-\\delta\(\\theta\_\{0\}\)\-\\eta KT\-\\mathbb\{E\}\[\\mathcal\{E\}\_\{\\text\{Ric\}\}\]\. where the termℰRic\\mathcal\{E\}\_\{\\text\{Ric\}\}is modulated by the Ricci curvature\. BecauseRic≡0\\text\{Ric\}\\equiv 0on the Calabi\-Yau manifold, theℰRic\\mathcal\{E\}\_\{\\text\{Ric\}\}term provides no negative\-curvature contribution, forcing the regret to scale with respect to parameter locations and parameter dimension only\.
### 5\.4First derivative results and results relating to negative Ricci curvature
Lemma 9 \(Dirichlet energy\)\.Suppose the regularized loss of[4\.23](https://arxiv.org/html/2608.19584#S4.E23)holds, and suppose the gap\(μ−λ\)\(\\mu\-\\lambda\)scales as𝒪\(N/K\)\\mathcal\{O\}\(N/K\)\. Then, the total Dirichlet energy of the network over the dataset, defined for each data point asE\(fα\)=1Volω\(S\)∫Si∂fα∧∂¯fα∧ωK−1\(K−1\)\!E\(f\_\{\\alpha\}\)=\\frac\{1\}\{\\text\{Vol\}\_\{\\omega\}\(S\)\}\\int\_\{S\}i\\partial f\_\{\\alpha\}\\wedge\\overline\{\\partial\}f\_\{\\alpha\}\\wedge\\frac\{\\omega^\{K\-1\}\}\{\(K\-1\)\!\}, is bounded above and below by
\(μ−λ\)Volω\(S\)≤∑α=1NE\(fα\)≤𝒪\(NK\+NKm\),\\displaystyle\(\\mu\-\\lambda\)\\text\{Vol\}\_\{\\omega\}\(S\)\\leq\\sum\_\{\\alpha=1\}^\{N\}E\(f\_\{\\alpha\}\)\\leq\\mathcal\{O\}\(NK\+\\frac\{NK\}\{\\sqrt\{m\}\}\),\(5\.13\)whereNNis the number of data points in the loss\.
Figure 5:We plot the top 100 eigenvalues corresponding to a regularized metric with the loss of[4\.23](https://arxiv.org/html/2608.19584#S4.E23)in a real scenario, which means that the metric takes the formh~ij¯=hij¯\+λδij¯\\widetilde\{h\}\_\{i\\overline\{j\}\}=h\_\{i\\overline\{j\}\}\+\\lambda\\delta\_\{i\\overline\{j\}\}\.Lemma 10 \(variation of the traversed parameter\)\.LetMMbe Kähler with metrichhfull rank\. Consider the natural gradient flowθ˙i=−hij¯∂j¯ℒ\\dot\{\\theta\}^\{i\}=\-h^\{i\\overline\{j\}\}\\partial\_\{\\overline\{j\}\}\\mathcal\{L\}\. Letv\(t\)=‖∇hℒ‖h2v\(t\)=\\\|\\nabla\_\{h\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}denote the gradient norm, andV\(t\)=𝔼ρt\[v\(t\)\]V\(t\)=\\mathbb\{E\}\_\{\\rho\_\{t\}\}\[v\(t\)\]denote its expectation over compactΩ\\Omega\. Define the uniform Hessian boundsH\(t\)=supθ∈Ω‖∇ω2,0ℒ‖hH\(t\)=\\sup\_\{\\theta\\in\\Omega\}\\\|\\nabla^\{2,0\}\_\{\\omega\}\\mathcal\{L\}\\\|\_\{h\}andμ\(t\)=infθ∈Ωλmin\(∇ω1,1ℒ\)\\mu\(t\)=\\inf\_\{\\theta\\in\\Omega\}\\lambda\_\{\\text\{min\}\}\(\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\)\. Assumee the Hessians are locallyLL\-Lipschitz with respect to the Kähler metric connection, and holomorphic bisectional curvature is bounded below byκ\>0\\kappa\>0\. We have
V\(t\)≤V\(0\)exp\(2t⋅𝒪\(1m\)\);\\displaystyle V\(t\)\\leq V\(0\)\\exp\\left\(2t\\cdot\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)\\right\);\(5\.14\)the material derivative of the minimum eigenvalue of the\(1,1\)\(1,1\)Hessian obeys
ddtμ\(t\)≥μ\(t\)2\+‖N\(X,⋅\)‖h2\+κv\(t\)−XiXj¯∇ij¯\(‖∇ℒ‖2\);\\displaystyle\\frac\{d\}\{dt\}\\mu\(t\)\\geq\\mu\(t\)^\{2\}\+\\\|N\(X,\\cdot\)\\\|\_\{h\}^\{2\}\+\\kappa v\(t\)\-X^\{i\}X^\{\\overline\{j\}\}\\nabla\_\{i\\overline\{j\}\}\(\\\|\\nabla\\mathcal\{L\}\\\|^\{2\}\);\(5\.15\)and the trajectory satisfies the upper bound
V\(t\)≤V\(0\)exp\(2∫0tH\(s\)𝑑s\+cKt\(Δ𝔼\[Δ∂¯ℒ\]t−∫0t𝒢\(s\)𝑑s\)\),\\displaystyle V\(t\)\\leq V\(0\)\\exp\\left\(2\\int\_\{0\}^\{t\}H\(s\)ds\+c\_\{K\}\\sqrt\{t\\left\(\\Delta\\mathbb\{E\}\[\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}\]\_\{t\}\-\\int\_\{0\}^\{t\}\\mathcal\{G\}\(s\)ds\\right\)\}\\right\),\(5\.16\)wherecK=22/Kc\_\{K\}=2\\sqrt\{2/K\},Δ𝔼\[Δ∂¯ℒ\]t\\Delta\\mathbb\{E\}\[\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}\]\_\{t\}denotes the net change in the expected Laplacian from time00tott, and𝒢\(s\)=κ\(s\)V\(s\)−12𝔼\[Δ∂¯ρsρsv\(s\)\]\\mathcal\{G\}\(s\)=\\kappa\(s\)V\(s\)\-\\frac\{1\}\{2\}\\mathbb\{E\}\\left\[\\frac\{\\Delta\_\{\\overline\{\\partial\}\}\\rho\_\{s\}\}\{\\rho\_\{s\}\}v\(s\)\\right\]\.
Lemma 11 \(integrated gradient flux\)\.Let\(M,ω\)\(M,\\omega\)be Kähler\. Denote𝒲\(ϵ\)\\mathcal\{W\}\(\\epsilon\)the quantity as in[11\.72](https://arxiv.org/html/2608.19584#S11.E72)\. Then𝒲\(ϵ\)\\mathcal\{W\}\(\\epsilon\)obeys with some shorthand
𝒲\(ϵ\)=𝒪\(Δ0ℒϵ2K\+1\+\[⟨Ric,ℋℒ⟩−RΔ0ℒ\]ϵ2K\+3\+ϵ2K\+5\)\.\\displaystyle\\mathcal\{W\}\(\\epsilon\)=\\mathcal\{O\}\\left\(\\Delta\_\{0\}\\mathcal\{L\}\\epsilon^\{2K\+1\}\+\\big\[\\langle\\text\{Ric\},\\mathcal\{H\}\_\{\\mathcal\{L\}\}\\rangle\-R\\Delta\_\{0\}\\mathcal\{L\}\\big\]\\epsilon^\{2K\+3\}\+\\epsilon^\{2K\+5\}\\right\)\.\(5\.17\)When\(M,ω\)\(M,\\omega\)is Calabi\-Yau, then the second term vanishes\. Moreover,𝒲\(ϵ\)\>𝒲CY\(ϵ\)\\mathcal\{W\}\(\\epsilon\)\>\\mathcal\{W\}\_\{CY\}\(\\epsilon\)when the manifold is negatively\-curved, where𝒲CY\(ϵ\)\\mathcal\{W\}\_\{CY\}\(\\epsilon\)corresponds to𝒲\\mathcal\{W\}in the Calabi\-Yau case, meaning the nonzero curvature case has greater escape at initialization than that of vanishing curvature\.
Remark\.In𝒲\(ϵ\)\\mathcal\{W\}\(\\epsilon\), we are interested in the quantity of the flux across radii∫0ϵ\(∮∂Br\(θ0\)dcℒ∧ωK−1\)𝑑r\\int\_\{0\}^\{\\epsilon\}\\left\(\\oint\_\{\\partial B\_\{r\}\(\\theta\_\{0\}\)\}d^\{c\}\\mathcal\{L\}\\wedge\\omega^\{K\-1\}\\right\)dr, as described in Appendix[11\.3](https://arxiv.org/html/2608.19584#S11.SS3)\. This formulation has connections to the flux under the Riemannian divergence theorem
∮∂Br⟨∇¯ℒ,n⟩h𝑑A=∫Brdiv¯\(∇¯ℒ\)ωKK\!=∫Br\(Δ∂¯ℒ\)ωKK\!\.\\displaystyle\\oint\_\{\\partial B\_\{r\}\}\\langle\\overline\{\\nabla\}\\mathcal\{L\},n\\rangle\_\{h\}dA=\\int\_\{B\_\{r\}\}\\overline\{\\text\{div\}\}\(\\overline\{\\nabla\}\\mathcal\{L\}\)\\frac\{\\omega^\{K\}\}\{K\!\}=\\int\_\{B\_\{r\}\}\(\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}\)\\frac\{\\omega^\{K\}\}\{K\!\}\.\(5\.18\)The above says that the gradient of the flux across the boundary is governed by the Dolbeault Laplacian\.
Lemma 12 \(variance of the parameter\)\.LetU⊆MU\\subseteq Mbe an open subset\. Letℒ\(θ\):U→ℝ\\mathcal\{L\}\(\\theta\):U\\rightarrow\\mathbb\{R\}be a smooth loss \(not necessarily quadratic\) with a local minimum at the critical pointθ∗\\theta^\{\*\}\. Assume a stochastic natural gradient descent update on the parameter with learning rateη\>0\\eta\>0overUU\. Furthermore, assume the system reaches a steady\-state probability measure \(via Fokker\-Planck; steady with respect to the gradient descent\) given by
ρ∞\(θ\)=1𝒵e−2ηℒ\(θ\),\\displaystyle\\rho\_\{\\infty\}\(\\theta\)=\\frac\{1\}\{\\mathcal\{Z\}\}e^\{\-\\frac\{2\}\{\\eta\}\\mathcal\{L\}\(\\theta\)\},\(5\.19\)where𝒵\\mathcal\{Z\}is the normalization constant over the volume formωKK\!\\frac\{\\omega^\{K\}\}\{K\!\}\. Then, the asymptotic varianceV∞=𝔼\[‖θ−θ∗‖h2\]V\_\{\\infty\}=\\mathbb\{E\}\\left\[\\\|\\theta\-\\theta^\{\*\}\\\|\_\{h\}^\{2\}\\right\]in a local neighborhood of the critical pointθ∗\\theta^\{\*\}is given by
V∞=η2Trh\(\[∇ω1,1ℒ\+η2Ric−∇ω2,0¯ℒ\(∇ω1,1¯ℒ\+η2Ric¯\)−1∇ω2,0ℒ\]−1\)\|θ∗\+𝒪\(η2\)\.\\displaystyle V\_\{\\infty\}=\\frac\{\\eta\}\{2\}\\text\{Tr\}\_\{h\}\\left\(\\left\[\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\+\\frac\{\\eta\}\{2\}\\text\{Ric\}\-\\overline\{\\nabla\_\{\\omega\}^\{2,0\}\}\\mathcal\{L\}\\left\(\\overline\{\\nabla^\{1,1\}\_\{\\omega\}\}\\mathcal\{L\}\+\\frac\{\\eta\}\{2\}\\overline\{\\text\{Ric\}\}\\right\)^\{\-1\}\\nabla\_\{\\omega\}^\{2,0\}\\mathcal\{L\}\\right\]^\{\-1\}\\right\)\\Bigg\|\_\{\\theta^\{\*\}\}\+\\mathcal\{O\}\(\\eta^\{2\}\)\.\(5\.20\)
Remark\.We can note the linear algebra fact
\{A\+B−C¯\(A¯\+B¯\)−1C⪰A−C¯A¯−1CB⪰0\.\\displaystyle\\begin\{cases\}A\+B\-\\overline\{C\}\(\\overline\{A\}\+\\overline\{B\}\)^\{\-1\}C\\succeq A\-\\overline\{C\}\\overline\{A\}^\{\-1\}C\\\\ B\\succeq 0\.\\end\{cases\}\(5\.21\)In this theorem, from the interior tensor in[12\.13](https://arxiv.org/html/2608.19584#S12.E13), we have positive Ricci curvature acts a restoring force to help the optimization with lower variance\. Global positivity on the Ricci curvature is nontrivially restrictive\. Instead, we can formulate this condition via a "weak restoring force" by evaluating the positivity of the holomorphic tangent bundle\. We say the landscape provides a sufficient restoring force atθ∈M\\theta\\in Mif[57](https://arxiv.org/html/2608.19584#bib.bib65)
\{\(iΘh\(T1,0M\)∧ωq−1∧Ω\)u,u\}h≥0,\\displaystyle\\Bigg\\\{\(i\\Theta\_\{h\}\(T^\{1,0\}M\)\\wedge\\omega^\{q\-1\}\\wedge\\Omega\)u,u\\Bigg\\\}\_\{h\}\\geq 0,\(5\.22\)foru∈T1,0Mθu\\in T^\{1,0\}M\_\{\\theta\}, integer1≤q≤K1\\leq q\\leq K, and a nowhere\-vanishing formΩ∈ΛK−q,K−qTθ∗M\\Omega\\in\\Lambda^\{K\-q,K\-q\}T\_\{\\theta\}^\{\*\}MwithΩ\>0\\Omega\>0\(metrically weakly\)\. Here,iΘh\(T1,0M\)i\\Theta\_\{h\}\(T^\{1,0\}M\), anEnd\(T1,0M\)\\text\{End\}\(T^\{1,0\}M\)\-valued\(1,1\)\(1,1\)\-form, represents the Chern curvature form of the tangent bundle\. Because the Ricci form is the trace of this Chern curvature, anω\\omega\-qq\-semi\-positive parameter landscape guarantees that the metric geometry prevents high variance in[12\.13](https://arxiv.org/html/2608.19584#S12.E13)in high negative curvature regimes, even if the manifold is locally Calabi\-Yau or exhibits minor eigenvalue collapse\. In particular, the landscape is allowed to have areas and directions of bad curvature that adversely affect[12\.13](https://arxiv.org/html/2608.19584#S12.E13), but the net effective curvature, when weighted against the specific metric structure ofΩ\\Omegain the relevant dimensionsqq, remains non\-negative\. This discussion ties into Lemma 14\.
Lemma 13 \(minimum eigenvalues of the Witten Laplacian\)\.LetU⊆MU\\subseteq Mbe an open subset of a Kähler manifold andK⊆UK\\subseteq Ua compact subset\. Consider the deformed LaplacianΔη=∂¯η∂¯η†\+∂¯η†∂¯η\\Delta\_\{\\eta\}=\\overline\{\\partial\}\_\{\\eta\}\\overline\{\\partial\}\_\{\\eta\}^\{\\dagger\}\+\\overline\{\\partial\}\_\{\\eta\}^\{\\dagger\}\\overline\{\\partial\}\_\{\\eta\}, where∂¯η=∂¯\+1η∂¯ℒ∧\\overline\{\\partial\}\_\{\\eta\}=\\overline\{\\partial\}\+\\frac\{1\}\{\\eta\}\\overline\{\\partial\}\\mathcal\{L\}\\wedge\. Assume the Ricci curvature is bounded above such thatRic\(∇ℒ,∇¯ℒ\)≤−κ‖∇ℒ‖h2\\text\{Ric\}\(\\nabla\\mathcal\{L\},\\overline\{\\nabla\}\\mathcal\{L\}\)\\leq\-\\kappa\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{2\}for some constantκ\>0\\kappa\>0\. Defineβ1,1=supθ∈U‖∇ω1,1ℒ‖2,Mη≤C0ηTr\(g⋅Hessℒ\(θ0\)\)\\beta\_\{1,1\}=\\sup\_\{\\theta\\in U\}\\\|\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\\\|\_\{2\},M\_\{\\eta\}\\leq\\frac\{C\_\{0\}\}\{\\eta\\text\{Tr\}\\big\(g\\cdot\\mathrm\{Hess\}\\mathcal\{L\}\(\\theta\_\{0\}\)\\big\)\}, whereggis the realification ofhh, andC0C\_\{0\}is some constant\. Denoteθ0\\theta\_\{0\}a unique minimum ofUU\. Then, the minimal eigenvalueλ1\\lambda\_\{1\}ofΔη\\Delta\_\{\\eta\}is bounded by
λ1≤\(2\+K\)β1,1η−κ\+Mη\.\\displaystyle\\lambda\_\{1\}\\leq\\frac\{\(2\+K\)\\beta\_\{1,1\}\}\{\\eta\}\-\\kappa\+M\_\{\\eta\}\.\(5\.23\)
Remark\.The above dividies byη\\etaand picks up a factor ofKKdue to a trace, which are not ideal\. Moreover, it is possible to redo the proof of Lemma 13 requiring a division on a gradient term, but this definitely diverges with an extrema in the set of interest\. We find the above to be a better result\. The purpose of the lemma is that an upper bound dictates the maximum possible size of a spectral gap\. By showing this bound, we establish a decay rate\. It establishes a "speed limit" on convergence, which proves the existence of a bottleneck, since the larger the smallest eigenvalue, the faster the convergence\. A wide spectral gap, or largeλ1\\lambda\_\{1\}, means the landscape is steep and desirable, or a small gap means the landscape is adverse\. As we can also see, the more negative the Ricci curvature, the smaller the gap\. Therefore, negative curvature adversely affects the spectral gap\.
Remark\.The choice of the Witten Laplacian is meaningful because it incorporates the lossℒ\\mathcal\{L\}itself into the geometric object\. The Witten Laplacian[49](https://arxiv.org/html/2608.19584#bib.bib68)was first interested by Witten[72](https://arxiv.org/html/2608.19584#bib.bib69)to study Morse inequalities, but it applies to our context here too\. We examine the minimum eigenvalue of the Witten Laplacian\. Our Witten Laplacian isΔη=∂¯η∂¯η†\+∂¯η†∂¯η\\Delta\_\{\\eta\}=\\overline\{\\partial\}\_\{\\eta\}\\overline\{\\partial\}\_\{\\eta\}^\{\\dagger\}\+\\overline\{\\partial\}\_\{\\eta\}^\{\\dagger\}\\overline\{\\partial\}\_\{\\eta\}, where∂¯η=∂¯\+1η∂¯ℒ∧\\overline\{\\partial\}\_\{\\eta\}=\\overline\{\\partial\}\+\\frac\{1\}\{\\eta\}\\overline\{\\partial\}\\mathcal\{L\}\\wedge\. The eigenvalues of this operator determine a rate of convergence\. To understand this, we must turn to the kernel of the differential operator, since a steady state corresponds to both an optimal state and the kernel of the operator\.
Lemma 14 \(divergence and diffusion bounds with semi\-nice curvature\)\. LetMMbe a Kähler manifold of complex dimensionKKwith Kähler formω\\omega\. Letℒ\\mathcal\{L\}be with the natural gradient descent vector field given byV=−∇h1,0ℒV=\-\\nabla^\{1,0\}\_\{h\}\\mathcal\{L\}, and letΘ=divh\(V\)\\Theta=\\text\{div\}\_\{h\}\(V\)denote its divergence\. SupposeMMsatisfies theω−q\\omega\-q\-semi\-positivity hypothesis with respect to a positive formΩ\\Omegasuch that
\{\(iΘh\(T1,0M\)∧ωq−1∧Ω\)V,V\}h≥0\.\\displaystyle\\Bigg\\\{\(i\\Theta\_\{h\}\(T^\{1,0\}M\)\\wedge\\omega^\{q\-1\}\\wedge\\Omega\)V,V\\Bigg\\\}\_\{h\}\\geq 0\.\(5\.24\)Then, the material derivative of the divergence, bounded by diffusion and expansion, satisfy the inequalities
−𝒪\(1m\)\+κmax𝒱˙≤Θ˙\+12Δh𝒱˙≤−‖∇ω1,1ℒ‖h2−‖∇ω2,0ℒ‖h2\+KMq,Ω𝒫q,Ω\(V\),\\displaystyle\-\\mathcal\{O\}\\left\(\\frac\{1\}\{m\}\\right\)\+\\kappa\_\{\\max\}\\dot\{\\mathcal\{V\}\}\\leq\\dot\{\\Theta\}\+\\frac\{1\}\{2\}\\Delta\_\{h\}\\dot\{\\mathcal\{V\}\}\\leq\-\\\|\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}\-\\\|\\nabla^\{2,0\}\_\{\\omega\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}\+\\frac\{K\}\{M\_\{q,\\Omega\}\}\\mathcal\{P\}\_\{q,\\Omega\}\(V\),\(5\.25\)where𝒱˙=−∥V∥h2,Mq,Ω=⋆\(ωq∧Ω\),𝒫q,Ω\(V\)=⋆\(ΘV,prim∧ωq−1∧Ω\)\\dot\{\\mathcal\{V\}\}=\-\\\|V\\\|\_\{h\}^\{2\},M\_\{q,\\Omega\}=\\star\(\\omega^\{q\}\\wedge\\Omega\),\\mathcal\{P\}\_\{q,\\Omega\}\(V\)=\\star\(\\Theta\_\{V,\\text\{prim\}\}\\wedge\\omega^\{q\-1\}\\wedge\\Omega\)is the primitive curvature term\. Furthermore, asK→∞K\\to\\infty, this upper bound diverges to\+∞\+\\inftysince the primitive curvature term scales linearly inKK,Ω\(1\)\\Omega\(1\)\. Consequently, in high\-dimensional spaces, the parameter flow has potential to "disperse," as the guarantee of convergence is lost\.
## 6Conclusions
We studied information cross\-entropy manifolds in a complex geometric lens for optimization\. We focused on results based in[10](https://arxiv.org/html/2608.19584#bib.bib1)and adapted and significantly changed these results for arguments of a greater geometric flavor, which use a deep learning theory base via the result of[9](https://arxiv.org/html/2608.19584#S9)\. We have expanded upon these results and considered diverse setups, namely through a dynamic Kähler Polyak\-Łojasiewicz condition\. We discussed failure modes of Calabi\-Yau manifolds and modes pertaining to small and Ricci negative curvature in general\. These results are notable because standard literature results are for sectional curvature[15](https://arxiv.org/html/2608.19584#bib.bib56), not Ricci curvature, at least in a deep learning context: we remark Ricci curvature involvement in optimization exists in literature in general[46](https://arxiv.org/html/2608.19584#bib.bib54)\. One limitation of our methods is that in practice it is nontrivial to compute our geometric structures such as local neighborhoods of the Kähler manifold corresponding to the cross\-entropy metric of[4\.1](https://arxiv.org/html/2608.19584#S4.E1)\. Moreover, in order to perform natural gradient descent, the metric must be computed, but the expectations can be approximated so this is more of an inconvenience and not a limitation\. In general, natural gradient descent is standard[63](https://arxiv.org/html/2608.19584#bib.bib58)\. Our work is mostly theoretical, yet confirmable with some experiments, as we saw in Figures[2](https://arxiv.org/html/2608.19584#S4.F2),[3](https://arxiv.org/html/2608.19584#S4.F3),[4](https://arxiv.org/html/2608.19584#S4.F4),[5](https://arxiv.org/html/2608.19584#S5.F5),[6](https://arxiv.org/html/2608.19584#S8.F6),[7](https://arxiv.org/html/2608.19584#S8.F7),[8](https://arxiv.org/html/2608.19584#S8.F8),[9](https://arxiv.org/html/2608.19584#S10.F9),[11](https://arxiv.org/html/2608.19584#S11.F11),[11](https://arxiv.org/html/2608.19584#S11.F11)\.
## References
- R\. AbdallaComplex\-valued neural networks – theory and analysis\.External Links:2312\.06087,[Link](https://arxiv.org/abs/2312.06087)Cited by:[§1](https://arxiv.org/html/2608.19584#S1.p1.1)\.
- Aitken and Gur\-Ari \(2020\)K\. Aitken and G\. Gur\-AriOn the asymptotics of wide networks with polynomial activations\.External Links:2006\.06687,[Link](https://arxiv.org/abs/2006.06687)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p1.1),[§8\.3](https://arxiv.org/html/2608.19584#S8.SS3.p4.1)\.
- Alexanderet al\.\(2023\)S\. Alexander, V\. Kapovitch, and A\. PetruninAlexandrov geometry: foundations\.External Links:1903\.08539,[Link](https://arxiv.org/abs/1903.08539)Cited by:[§4\.3](https://arxiv.org/html/2608.19584#S4.SS3.p13.1)\.
- Amariet al\.\(2018\)S\. Amari, R\. Karakida, and M\. OizumiStatistical neurodynamics of deep networks: geometry of signal spaces\.External Links:1808\.07169,[Link](https://arxiv.org/abs/1808.07169)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p1.1)\.
- Amari \(1998\)S\. AmariNatural gradient works efficiently in learning\.Neural Computation10\(2\),pp\. 251–276\.External Links:ISSN 0899\-7667,[Document](https://dx.doi.org/10.1162/089976698300017746),[Link](https://doi.org/10.1162/089976698300017746),https://direct\.mit\.edu/neco/article\-pdf/10/2/251/813415/089976698300017746\.pdfCited by:[§1](https://arxiv.org/html/2608.19584#S1.p1.1)\.
- Ambrosioet al\.\(2015\)L\. Ambrosio, N\. Gigli, and G\. SavaréBakry–Émery curvature\-dimension condition and riemannian ricci curvature bounds\.The Annals of Probability43\(1\)\.External Links:ISSN 0091\-1798,[Link](http://dx.doi.org/10.1214/14-AOP907),[Document](https://dx.doi.org/10.1214/14-aop907)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p2.1)\.
- Andreassen and Dyer \(2020\)A\. Andreassen and E\. DyerAsymptotics of wide convolutional neural networks\.External Links:2008\.08675,[Link](https://arxiv.org/abs/2008.08675)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p1.1),[§8\.3](https://arxiv.org/html/2608.19584#S8.SS3.p4.1)\.
- Anghel \(2026\)C\. AnghelAKSZ descent on manifolds with ordinary corners\.External Links:2608\.02928,[Link](https://arxiv.org/abs/2608.02928)Cited by:[§12\.2](https://arxiv.org/html/2608.19584#S12.SS2.p8.5)\.
- Bakeer \(2026\)T\. BakeerLocal information operators for spatial identifiability in distributed\-parameter inverse problems in computational mechanics\.External Links:2605\.28601,[Link](https://arxiv.org/abs/2605.28601)Cited by:[§4\.3](https://arxiv.org/html/2608.19584#S4.SS3.p4.1)\.
- Banerjeeet al\.\(2023\)A\. Banerjee, P\. Cisneros\-Velarde, L\. Zhu, and M\. BelkinRestricted strong convexity of deep learning models with smooth activations\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=PINRbk7h01)Cited by:[§1](https://arxiv.org/html/2608.19584#S1.p2.1),[§2](https://arxiv.org/html/2608.19584#S2.p1.1),[§3](https://arxiv.org/html/2608.19584#S3.p1.1),[§4\.3](https://arxiv.org/html/2608.19584#S4.SS3.p5.2),[§4\.3](https://arxiv.org/html/2608.19584#S4.SS3.p6.2),[§5\.1](https://arxiv.org/html/2608.19584#S5.SS1.p2.1),[§5\.3](https://arxiv.org/html/2608.19584#S5.SS3.p1.1),[§6](https://arxiv.org/html/2608.19584#S6.p1.1),[§8\.2](https://arxiv.org/html/2608.19584#S8.SS2.p3.1),[§8\.3](https://arxiv.org/html/2608.19584#S8.SS3.p4.1),[§8\.4](https://arxiv.org/html/2608.19584#S8.SS4.p1.1)\.
- Bennequinet al\.\(2020\)D\. Bennequin, O\. Peltre, G\. Sergeant\-Perthuis, and J\. P\. VigneauxExtra\-fine sheaves and interaction decompositions\.External Links:2009\.12646,[Link](https://arxiv.org/abs/2009.12646)Cited by:[§4\.3](https://arxiv.org/html/2608.19584#S4.SS3.p8.3)\.
- Cayci \(2025\)S\. CayciA riemannian optimization perspective of the gauss\-newton method for feedforward neural networks\.External Links:2412\.14031,[Link](https://arxiv.org/abs/2412.14031)Cited by:[§4\.3](https://arxiv.org/html/2608.19584#S4.SS3.p6.2)\.
- Cironeet al\.\(2025\)N\. M\. Cirone, J\. Hamdan, and C\. SalviGenus expansion for non\-linear random matrix ensembles with applications to neural networks\.External Links:2407\.08459,[Link](https://arxiv.org/abs/2407.08459)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p1.1),[§8\.3](https://arxiv.org/html/2608.19584#S8.SS3.p4.1)\.
- Cisneros\-Velardeet al\.\(2025\)P\. Cisneros\-Velarde, Z\. Chen, S\. Koyejo, and A\. BanerjeeOptimization and generalization guarantees for weight normalization\.Transactions on Machine Learning Research\.Note:External Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=gpHOtQQPJG)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p1.1)\.
- Criscitiello and Boumal \(2022\)C\. Criscitiello and N\. BoumalNegative curvature obstructs acceleration for strongly geodesically convex optimization, even with exact first\-order oracles\.InProceedings of Thirty Fifth Conference on Learning Theory,P\. Loh and M\. Raginsky \(Eds\.\),Proceedings of Machine Learning Research, Vol\.178,pp\. 496–542\.External Links:[Link](https://proceedings.mlr.press/v178/criscitiello22a.html)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p2.1),[§6](https://arxiv.org/html/2608.19584#S6.p1.1)\.
- Criscitiello and Boumal \(2023\)C\. Criscitiello and N\. BoumalCurvature and complexity: better lower bounds for geodesically convex optimization\.InProceedings of Thirty Sixth Conference on Learning Theory,G\. Neu and L\. Rosasco \(Eds\.\),Proceedings of Machine Learning Research, Vol\.195,pp\. 2969–3013\.External Links:[Link](https://proceedings.mlr.press/v195/criscitiello23a.html)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p2.1)\.
- Daneshmandet al\.\(2023\)H\. Daneshmand, J\. D\. Lee, and C\. JinEfficient displacement convex optimization with particle gradient descent\.External Links:2302\.04753,[Link](https://arxiv.org/abs/2302.04753)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p1.1)\.
- Deligne P\. and Morgan \(1975\)G\. P\. Deligne P\. and J\. MorganReal homotopy theory of kähler manifolds\.\.Inventiones mathematicae29,pp\. 245–274\.External Links:[Link](http://eudml.org/doc/142341)Cited by:[§4\.3](https://arxiv.org/html/2608.19584#S4.SS3.p8.3)\.
- Dello Schiavoet al\.\(2024\)L\. Dello Schiavo, J\. Maas, and F\. PedrottiLocal conditions for global convergence of gradient flows and proximal point sequences in metric spaces\.Transactions of the American Mathematical Society\.External Links:ISSN 1088\-6850,[Link](http://dx.doi.org/10.1090/tran/9156),[Document](https://dx.doi.org/10.1090/tran/9156)Cited by:[§11\.1](https://arxiv.org/html/2608.19584#S11.SS1.p1.1)\.
- Demailly \(2012\)J\. DemaillyComplex analytic and differential geometry\.Note:Manuscript, Université de GrenobleExternal Links:[Link](https://people.math.harvard.edu/%CB%9Cdemarco/Math274/Demailly_ComplexAnalyticDiffGeom.pdf)Cited by:[§4\.2](https://arxiv.org/html/2608.19584#S4.SS2.p2.1)\.
- Dong and Cheng \(2026\)H\. Dong and P\. ChengQuotient geometry, effective curvature, and implicit bias in simple shallow neural networks\.External Links:2603\.21502,[Link](https://arxiv.org/abs/2603.21502)Cited by:[§4\.3](https://arxiv.org/html/2608.19584#S4.SS3.p4.1)\.
- Dufort\-Labbéet al\.\(2026\)S\. Dufort\-Labbé, M\. Hamidi, R\. Pascanu, I\. Mitliagkas, D\. Scieur, and A\. BaratinNavigating potholes with geometry\-aware sharpness minimization\.External Links:2605\.16134,[Link](https://arxiv.org/abs/2605.16134)Cited by:[§4\.3](https://arxiv.org/html/2608.19584#S4.SS3.p6.2)\.
- Dyer and Gur\-Ari \(2019\)E\. Dyer and G\. Gur\-AriAsymptotics of wide networks from feynman diagrams\.External Links:1909\.11304,[Link](https://arxiv.org/abs/1909.11304)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p1.1),[§8\.3](https://arxiv.org/html/2608.19584#S8.SS3.p4.1)\.
- Ehrhardtet al\.\(2024\)M\. J\. Ehrhardt, E\. S\. Riis, T\. Ringholm, and C\. SchönliebA geometric integration approach to smooth optimization: foundations of the discrete gradient method\.IMA Journal of Numerical Analysis45\(3\),pp\. 1269–1299\.External Links:ISSN 1464\-3642,[Link](http://dx.doi.org/10.1093/imanum/drae037),[Document](https://dx.doi.org/10.1093/imanum/drae037)Cited by:[§11\.1](https://arxiv.org/html/2608.19584#S11.SS1.p1.1)\.
- Fanget al\.\(2025\)J\. Fang, Z\. Xiong, and X\. YangLaplacian comparison theorems on complete kähler manifolds and applications\.External Links:2510\.01548,[Link](https://arxiv.org/abs/2510.01548)Cited by:[§10\.6](https://arxiv.org/html/2608.19584#S10.SS6.p2.7)\.
- Fritz \(2020\)T\. FritzA synthetic approach to markov kernels, conditional independence and theorems on sufficient statistics\.Advances in Mathematics370,pp\. 107239\.External Links:ISSN 0001\-8708,[Link](http://dx.doi.org/10.1016/j.aim.2020.107239),[Document](https://dx.doi.org/10.1016/j.aim.2020.107239)Cited by:[§4\.1](https://arxiv.org/html/2608.19584#S4.SS1.p1.1)\.
- Gigli \(2017\)N\. GigliNonsmooth differential geometry– an approach tailored for spaces with ricci curvature bounded from below\.Memoirs of the American Mathematical Society251\(1196\)\.External Links:ISSN 0065\-9266,[Link](http://dx.doi.org/10.1090/memo/1196),[Document](https://dx.doi.org/10.1090/memo/1196)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p2.1)\.
- Gil\-García \(2026\)A\. Gil\-GarcíaTorsion parallel pure spinors on neutral manifolds\.External Links:2607\.06358,[Link](https://arxiv.org/abs/2607.06358)Cited by:[§12\.2](https://arxiv.org/html/2608.19584#S12.SS2.p8.7)\.
- Gnandi \(2026\)E\. GnandiConstruction of exponential families from statistical manifolds\.External Links:2511\.23444,[Link](https://arxiv.org/abs/2511.23444)Cited by:[§8\.3](https://arxiv.org/html/2608.19584#S8.SS3.p3.1)\.
- Griffiths and Harris \(1978\)P\. A\. Griffiths and J\. W\. HarrisPrinciples of algebraic geometry\.External Links:[Link](https://api.semanticscholar.org/CorpusID:118963833)Cited by:[§12\.3](https://arxiv.org/html/2608.19584#S12.SS3.p2.1)\.
- Guillenet al\.\(2026\)M\. Guillen, P\. Misof, and J\. E\. GerkenFinite\-width neural tangent kernels from feynman diagrams\.External Links:2508\.11522,[Link](https://arxiv.org/abs/2508.11522)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p1.1),[§8\.2](https://arxiv.org/html/2608.19584#S8.SS2.p7.1),[§8\.3](https://arxiv.org/html/2608.19584#S8.SS3.p4.1)\.
- Hanin \(2023\)B\. HaninRandom fully connected neural networks as perturbatively solvable hierarchies\.External Links:2204\.01058,[Link](https://arxiv.org/abs/2204.01058)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p1.1),[§8\.3](https://arxiv.org/html/2608.19584#S8.SS3.p4.1)\.
- Hauer and Mazón \(2019\)D\. Hauer and J\. M\. MazónKurdyka–Łojasiewicz–simon inequality for gradient flows in metric spaces\.Trans\. Amer\. Math\. Soc\.372,pp\. 4917–4976\.External Links:[Document](https://dx.doi.org/10.1090/tran/7801),[Link](https://doi.org/10.1090/tran/7801),[MathReview Entry](https://www.ams.org/mathscinet-getitem?mr=4009443)Cited by:[§11\.1](https://arxiv.org/html/2608.19584#S11.SS1.p1.1)\.
- Herzlich \(2000\)M\. HerzlichRefined kato inequalities in riemannian geometry\.InJournées Équations aux dérivées partielles,pp\. 1–11\.External Links:[Link](https://www.numdam.org/item/10.5802/jedp.570.pdf)Cited by:[§8\.5](https://arxiv.org/html/2608.19584#S8.SS5.p2.2)\.
- Hezariet al\.\(2016\)H\. Hezari, C\. Kelleher, S\. Seto, and H\. XuAsymptotic expansion of the bergman kernel via perturbation of the bargmann–fock model\.The Journal of Geometric Analysis26\(4\),pp\. 2602–2638\.External Links:ISSN 1559\-002X,[Document](https://dx.doi.org/10.1007/s12220-015-9641-3),[Link](https://doi.org/10.1007/s12220-015-9641-3)Cited by:[§11\.3](https://arxiv.org/html/2608.19584#S11.SS3.p2.3)\.
- Huang and Yau \(2019\)J\. Huang and H\. YauDynamics of deep neural networks and neural tangent hierarchy\.External Links:1909\.08156,[Link](https://arxiv.org/abs/1909.08156)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p1.1),[§8\.3](https://arxiv.org/html/2608.19584#S8.SS3.p4.1)\.
- Javanmardet al\.\(2019\)A\. Javanmard, M\. Mondelli, and A\. MontanariAnalysis of a two\-layer neural network via displacement convexity\.External Links:1901\.01375,[Link](https://arxiv.org/abs/1901.01375)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p1.1)\.
- Karkada \(2024\)D\. KarkadaThe lazy \(ntk\) and rich \(μ\\mup\) regimes: a gentle tutorial\.External Links:2404\.19719,[Link](https://arxiv.org/abs/2404.19719)Cited by:[Figure 4](https://arxiv.org/html/2608.19584#S4.F4)\.
- Kaul and Lall \(2019\)P\. Kaul and B\. LallRiemannian curvature of deep neural networks\.IEEE Transactions on Neural Networks and Learning Systems31,pp\. 1410–1416\.External Links:[Document](https://dx.doi.org/10.1109/TNNLS.2019.2919705)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p1.1)\.
- Kingma and Ba \(2017\)D\. P\. Kingma and J\. BaAdam: a method for stochastic optimization\.External Links:1412\.6980,[Link](https://arxiv.org/abs/1412.6980)Cited by:[§10\.6](https://arxiv.org/html/2608.19584#S10.SS6.p1.1)\.
- Lawsonet al\.\(2023\)B\. A\. J\. Lawson, K\. Burrage, K\. Mengersen, and R\. W\. dos SantosThe fisher geometry and geodesics of the multivariate normals, without differential geometry\.External Links:2306\.01278,[Link](https://arxiv.org/abs/2306.01278)Cited by:[§1](https://arxiv.org/html/2608.19584#S1.p1.1)\.
- Léger and Vialard \(2023\)F\. Léger and F\. VialardA geometric laplace method\.Pure and Applied Analysis5\(4\),pp\. 1041–1080\.External Links:ISSN 2578\-5893,[Link](http://dx.doi.org/10.2140/paa.2023.5.1041),[Document](https://dx.doi.org/10.2140/paa.2023.5.1041)Cited by:[§12\.2](https://arxiv.org/html/2608.19584#S12.SS2.p3.8)\.
- Li and Montúfar \(2020\)W\. Li and G\. MontúfarRicci curvature for parametric statistics via optimal transport\.Information Geometry3\(1\),pp\. 89–117\.External Links:[Document](https://dx.doi.org/10.1007/s41884-020-00026-2),[Link](https://doi.org/10.1007/s41884-020-00026-2),ISSN 2511\-249XCited by:[§2](https://arxiv.org/html/2608.19584#S2.p2.1)\.
- Liuet al\.\(2021a\)C\. Liu, L\. Zhu, and M\. BelkinLoss landscapes and optimization in over\-parameterized non\-linear systems and neural networks\.External Links:2003\.00307,[Link](https://arxiv.org/abs/2003.00307)Cited by:[§8\.4](https://arxiv.org/html/2608.19584#S8.SS4.p1.1)\.
- Liuet al\.\(2021b\)C\. Liu, L\. Zhu, and M\. BelkinOn the linearity of large non\-linear models: when and why the tangent kernel is constant\.External Links:2010\.01092,[Link](https://arxiv.org/abs/2010.01092)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p1.1),[§8\.4](https://arxiv.org/html/2608.19584#S8.SS4.p1.1)\.
- Lott and Villani \(2009\)J\. Lott and C\. VillaniRicci curvature for metric\-measure spaces via optimal transport\.Annals of Mathematics169\(3\),pp\. 903–991\.External Links:[Document](https://dx.doi.org/10.4007/annals.2009.169.903),[Link](https://doi.org/10.4007/annals.2009.169.903),[MathReview Entry](https://www.ams.org/mathscinet-getitem?mr=2480619)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p2.1),[§6](https://arxiv.org/html/2608.19584#S6.p1.1)\.
- McNeal and Varolin \(2015\)J\. D\. McNeal and D\. VarolinL2L^\{2\}estimates for the∂¯\\bar\{\\partial\}operator\.Bulletin of Mathematical Sciences5\(2\),pp\. 179–249\.External Links:ISSN 1664\-3615,[Document](https://dx.doi.org/10.1007/s13373-015-0068-8),[Link](https://doi.org/10.1007/s13373-015-0068-8)Cited by:[§12\.2](https://arxiv.org/html/2608.19584#S12.SS2.p8.4)\.
- Meng and Zhang \(2025\)Z\. Meng and D\. ZhangCombinatorial courant\-fischer\-weyl minimax principle on cheegerkk\-constants of weighted forests\.External Links:2510\.06301,[Link](https://arxiv.org/abs/2510.06301)Cited by:[§4\.3](https://arxiv.org/html/2608.19584#S4.SS3.p11.3)\.
- Michel \(2019\)L\. MichelAbout small eigenvalues of the witten laplacian\.Pure and Applied Analysis1\(2\),pp\. 149–204\.External Links:[Document](https://dx.doi.org/10.2140/paa.2019.1.149),[Link](https://msp.org/paa/2019/1-2/paa-v1-n2-p01-p.pdf)Cited by:[§5\.4](https://arxiv.org/html/2608.19584#S5.SS4.p9.1)\.
- Mishra and Tan \(2025\)C\. Mishra and J\. TanHermitian yang–mills connections on general vector bundles: geometry and physical yukawa couplings\.External Links:2512\.10907,[Link](https://arxiv.org/abs/2512.10907)Cited by:[§4\.3](https://arxiv.org/html/2608.19584#S4.SS3.p8.3)\.
- Mukhin and Kundikova \(2021\)Y\. V\. Mukhin and N\. D\. KundikovaCauchy\-like criterion for differentiability of functions of several variables\.External Links:2107\.13524,[Link](https://arxiv.org/abs/2107.13524)Cited by:[§9](https://arxiv.org/html/2608.19584#S9.p5.2)\.
- Murrayet al\.\(2023\)M\. Murray, H\. Jin, B\. Bowman, and G\. MontufarCharacterizing the spectrum of the ntk via a power series expansion\.External Links:2211\.07844,[Link](https://arxiv.org/abs/2211.07844)Cited by:[§11\.1](https://arxiv.org/html/2608.19584#S11.SS1.p1.1)\.
- Niu \(2022\)Y\. NiuOn the convergence analysis of dca\.External Links:2211\.10942,[Link](https://arxiv.org/abs/2211.10942)Cited by:[§11\.1](https://arxiv.org/html/2608.19584#S11.SS1.p1.1)\.
- Niu \(2026\)Y\. NiuContinuous\-time dynamics of the difference\-of\-convex algorithm\.External Links:2604\.06926,[Link](https://arxiv.org/abs/2604.06926)Cited by:[§11\.1](https://arxiv.org/html/2608.19584#S11.SS1.p1.1)\.
- Petrov \(2024\)A\. PetrovThe essence of de rham cohomology\.External Links:2411\.06296,[Link](https://arxiv.org/abs/2411.06296)Cited by:[§4\.3](https://arxiv.org/html/2608.19584#S4.SS3.p12.1)\.
- Pooleet al\.\(2016\)B\. Poole, S\. Lahiri, M\. Raghu, J\. Sohl\-Dickstein, and S\. GanguliExponential expressivity in deep neural networks through transient chaos\.InAdvances in Neural Information Processing Systems,D\. Lee, M\. Sugiyama, U\. Luxburg, I\. Guyon, and R\. Garnett \(Eds\.\),Vol\.29,pp\.\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2016/file/148510031349642de5ca0c544f31b2ef-Paper.pdf)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p1.1)\.
- Popovici \(2026\)D\. Popovicimm\-Positive stability of holomorphic vector bundles and moduli spaces\.External Links:2607\.17203,[Link](https://arxiv.org/abs/2607.17203)Cited by:[§5\.4](https://arxiv.org/html/2608.19584#S5.SS4.p6.2)\.
- Riiset al\.\(2018\)E\. S\. Riis, M\. J\. Ehrhardt, G\. R\. W\. Quispel, and C\. SchönliebA geometric integration approach to nonsmooth, nonconvex optimisation\.External Links:1807\.07554,[Link](https://arxiv.org/abs/1807.07554)Cited by:[§11\.1](https://arxiv.org/html/2608.19584#S11.SS1.p1.1)\.
- Ringholmet al\.\(2018\)T\. Ringholm, J\. Lazić, and C\. SchönliebVariational image regularization with euler’s elastica using a discrete gradient scheme\.External Links:1712\.07386,[Link](https://arxiv.org/abs/1712.07386)Cited by:[§11\.1](https://arxiv.org/html/2608.19584#S11.SS1.p1.1)\.
- Ruan \(1998\)W\. RuanCanonical coordinates and bergman metrics\.Communications in Analysis and Geometry6\(3\)\.External Links:[Link](https://intlpress.com/site/pub/files/_fulltext/journals/cag/1998/0006/0003/CAG-1998-0006-0003-a005.pdf)Cited by:[§11\.3](https://arxiv.org/html/2608.19584#S11.SS3.p2.3)\.
- Schwachhöferet al\.\(2017\)L\. Schwachhöfer, N\. Ay, J\. Jost, and H\. V\. LêCongruent families and invariant tensors\.External Links:1705\.11014,[Link](https://arxiv.org/abs/1705.11014)Cited by:[§8\.3](https://arxiv.org/html/2608.19584#S8.SS3.p3.1)\.
- Shem\-Ur and Oz \(2024\)O\. Shem\-Ur and Y\. OzWeak correlations as the underlying principle for linearization of gradient\-based learning systems\.External Links:2401\.04013,[Link](https://arxiv.org/abs/2401.04013)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p1.1)\.
- Shrestha \(2023\)R\. ShresthaNatural gradient methods: perspectives, efficient\-scalable approximations, and analysis\.External Links:2303\.05473,[Link](https://arxiv.org/abs/2303.05473)Cited by:[§4\.1](https://arxiv.org/html/2608.19584#S4.SS1.p4.1),[§6](https://arxiv.org/html/2608.19584#S6.p1.1)\.
- Stoica \(2020\)O\. C\. StoicaChiral asymmetry in the weak interaction via clifford algebras\.External Links:2005\.08855,[Link](https://arxiv.org/abs/2005.08855)Cited by:[§12\.2](https://arxiv.org/html/2608.19584#S12.SS2.p8.7)\.
- Sun and Nielsen \(2025\)K\. Sun and F\. NielsenA geometric modeling of occam’s razor in deep learning\.Information Geometry8\(S1\),pp\. 233–273\.External Links:ISSN 2511\-249X,[Link](http://dx.doi.org/10.1007/s41884-025-00167-2),[Document](https://dx.doi.org/10.1007/s41884-025-00167-2)Cited by:[§4\.3](https://arxiv.org/html/2608.19584#S4.SS3.p4.1)\.
- Taheriet al\.\(2024\)H\. Taheri, C\. Thrampoulidis, and A\. MazumdarSharper guarantees for learning neural network classifiers with gradient methods\.External Links:2410\.10024,[Link](https://arxiv.org/abs/2410.10024)Cited by:[§8\.4](https://arxiv.org/html/2608.19584#S8.SS4.p1.1)\.
- Tam and Yu \(2012\)L\. Tam and C\. YuSome comparison theorems for kähler manifolds\.Manuscripta Mathematica137\(3\),pp\. 483–495\.External Links:ISSN 1432\-1785,[Document](https://dx.doi.org/10.1007/s00229-011-0477-2),[Link](https://doi.org/10.1007/s00229-011-0477-2)Cited by:[§10\.6](https://arxiv.org/html/2608.19584#S10.SS6.p2.7)\.
- Trabelsiet al\.\(2018\)C\. Trabelsi, O\. Bilaniuk, Y\. Zhang, D\. Serdyuk, S\. Subramanian, J\. F\. Santos, S\. Mehri, N\. Rostamzadeh, Y\. Bengio, and C\. J\. PalDeep complex networks\.External Links:1705\.09792,[Link](https://arxiv.org/abs/1705.09792)Cited by:[§1](https://arxiv.org/html/2608.19584#S1.p1.1)\.
- Tronet al\.\(2024\)E\. Tron, R\. Fioresi, N\. Couëllan, and S\. PuechmorelCartan moving frames and the data manifolds\.Information Geometry7\(S2\),pp\. 883–912\.External Links:ISSN 2511\-249X,[Link](http://dx.doi.org/10.1007/s41884-024-00159-8),[Document](https://dx.doi.org/10.1007/s41884-024-00159-8)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p1.1)\.
- Wang \(2024\)Z\. WangLecture 20: the index form\.Note:Lecture notes for Riemannian Geometry, University of Science and Technology of ChinaExternal Links:[Link](http://staff.ustc.edu.cn/%CB%9Cwangzuoq/Courses/24S-RiemGeom/Notes/Lec20.pdf)Cited by:[§10\.6](https://arxiv.org/html/2608.19584#S10.SS6.p4.1)\.
- Wells \(1980\)R\. O\. WellsDifferential analysis on complex manifolds\.2 edition,Graduate Texts in Mathematics,Springer New York\.External Links:[Document](https://dx.doi.org/10.1007/978-1-4757-3946-6),ISBN 978\-1\-4757\-3946\-6,[Link](https://link.springer.com/book/10.1007/978-1-4757-3946-6)Cited by:[§12\.3](https://arxiv.org/html/2608.19584#S12.SS3.p2.1),[§12\.3](https://arxiv.org/html/2608.19584#S12.SS3.p2.2)\.
- Witten \(1982\)E\. WittenSupersymmetry and morse theory\.Journal of Differential Geometry17\(4\),pp\. 661–692\(English \(US\)\)\.External Links:[Document](https://dx.doi.org/10.4310/jdg/1214437492),ISSN 0022\-040XCited by:[§5\.4](https://arxiv.org/html/2608.19584#S5.SS4.p9.1)\.
- Yaida \(2022\)S\. YaidaMeta\-principled family of hyperparameter scaling strategies\.External Links:2210\.04909,[Link](https://arxiv.org/abs/2210.04909)Cited by:[§8\.2](https://arxiv.org/html/2608.19584#S8.SS2.p7.1)\.
- Yau \(1977\)S\. YauCalabi’s conjecture and some new results in algebraic geometry\.Proceedings of the National Academy of Sciences74\(5\),pp\. 1798–1799\.External Links:[Document](https://dx.doi.org/10.1073/pnas.74.5.1798),[Link](https://doi.org/10.1073/pnas.74.5.1798)Cited by:[§4\.2](https://arxiv.org/html/2608.19584#S4.SS2.p2.1)\.
- Zavatone\-Vethet al\.\(2025\)J\. A\. Zavatone\-Veth, S\. Yang, J\. A\. Rubinfien, and C\. PehlevanHow does training shape the riemannian geometry of neural network representations?\.External Links:2301\.11375,[Link](https://arxiv.org/abs/2301.11375)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p1.1)\.
- Zein and Tu \(2014\)F\. E\. Zein and L\. W\. TuFrom sheaf cohomology to the algebraic de rham theorem\.External Links:1302\.5834,[Link](https://arxiv.org/abs/1302.5834)Cited by:[§4\.3](https://arxiv.org/html/2608.19584#S4.SS3.p8.3)\.
- Zhanget al\.\(2022\)J\. Zhang, X\. Huang, and J\. YuMean\-field analysis of two\-layer neural networks: global optimality with linear convergence rates\.External Links:2205\.09860,[Link](https://arxiv.org/abs/2205.09860)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p1.1)\.
- Zhuet al\.\(2022\)L\. Zhu, P\. Pandit, and M\. BelkinA note on linear bottleneck networks and their transition to multilinearity\.External Links:2206\.15058,[Link](https://arxiv.org/abs/2206.15058)Cited by:[§2](https://arxiv.org/html/2608.19584#S2.p1.1)\.
- Ziemer \(2017\)W\. P\. ZiemerModern real analysis\.2 edition,Graduate Texts in Mathematics,Springer International Publishing,Cham\.External Links:ISBN 978\-3\-319\-64628\-2,[Document](https://dx.doi.org/10.1007/978-3-319-64629-9),[Link](https://link.springer.com/book/10.1007/978-3-319-64629-9)Cited by:[§4\.1](https://arxiv.org/html/2608.19584#S4.SS1.p2.1)\.
## 7Main notations
## 8Second derivative results
### 8\.1Complex Hessian background
In this section, we discuss our strategies for connecting the types of complex Hessians\.
We first note the complexified cotangent bundle splitsT∗X⊗ℂ=T1,0∗X⊕T0,1∗XT^\{\*\}X\\otimes\\mathbb\{C\}=T^\{1,0\*\}X\\oplus T^\{0,1\*\}X, and the second derivative splits
∇ω2f=∇ω2,0f\+∇ω1,1f\+∇ω0,2f\.\\displaystyle\\nabla\_\{\\omega\}^\{2\}f=\\nabla^\{2,0\}\_\{\\omega\}f\+\\nabla^\{1,1\}\_\{\\omega\}f\+\\nabla^\{0,2\}\_\{\\omega\}f\.\(8\.1\)In local holomorphic coordinates, the holomorphic tangent bundle itself decomposes
T1,0M=⨁i=1Kℂ∂∂zi\.\\displaystyle T^\{1,0\}M=\\bigoplus\_\{i=1\}^\{K\}\\mathbb\{C\}\\frac\{\\partial\}\{\\partial z\_\{i\}\}\.\(8\.2\)
We will attempt to bound various versions of the complex Hessian\. For example, we will attempt to bound the section ofT∗X⊗T∗XT^\{\*\}X\\otimes T^\{\*\}Xin local holomorphic coordinates
∇ω1,1f=∂2f∂θi∂θ¯jdθi⊗dθ¯j\.\\displaystyle\\nabla^\{1,1\}\_\{\\omega\}f=\\frac\{\\partial^\{2\}f\}\{\\partial\\theta^\{i\}\\partial\\overline\{\\theta\}^\{j\}\}d\\theta^\{i\}\\otimes d\\overline\{\\theta\}^\{j\}\.\(8\.3\)This is most closely related to the \(1,1\)\-form
i∂∂¯f=i∑ij∂2f∂θi∂θ¯jdθi∧dθ¯j\.\\displaystyle i\\partial\\overline\{\\partial\}f=i\\sum\_\{ij\}\\frac\{\\partial^\{2\}f\}\{\\partial\\theta^\{i\}\\partial\\overline\{\\theta\}^\{j\}\}d\\theta^\{i\}\\wedge d\\overline\{\\theta\}^\{j\}\.\(8\.4\)The \(2,0\) part has the closed form
∇ω2,0f=∑ij∂2f∂θi∂θjdθi⊗dθj,\\displaystyle\\nabla^\{2,0\}\_\{\\omega\}f=\\sum\_\{ij\}\\frac\{\\partial^\{2\}f\}\{\\partial\\theta^\{i\}\\partial\\theta^\{j\}\}d\\theta^\{i\}\\otimes d\\theta^\{j\},\(8\.5\)therefore this requires a different bound\. One approach could be to bound two terms simultaneously, although we find the approach to bound the terms individually more straightforward in practice\. If we take the Kähler metric asω=ihij¯dθi∧dθ¯j\\omega=ih\_\{i\\overline\{j\}\}d\\theta^\{i\}\\wedge d\\overline\{\\theta\}^\{j\}, we can take the trace \(the dual Lefschetz operatorΛ\\Lambda\) of the\(1,1\)\(1,1\)\-formi∂∂¯f=i∂2f∂zi∂z¯jdθi∧dθ¯ji\\partial\\overline\{\\partial\}f=i\\frac\{\\partial^\{2\}f\}\{\\partial z^\{i\}\\partial\\overline\{z\}^\{j\}\}d\\theta^\{i\}\\wedge d\\overline\{\\theta\}^\{j\}, giving
⋆\(i∂∂¯f∧ωK−1\(K−1\)\!\)=Λ\(i∂∂¯f\)=hij¯∂2f∂zi∂z¯j=Δ∂¯f\.\\displaystyle\\star\\left\(i\\partial\\overline\{\\partial\}f\\wedge\\frac\{\\omega^\{K\-1\}\}\{\(K\-1\)\!\}\\right\)=\\Lambda\(i\\partial\\overline\{\\partial\}f\)=h^\{i\\overline\{j\}\}\\frac\{\\partial^\{2\}f\}\{\\partial z^\{i\}\\partial\\overline\{z\}^\{j\}\}=\\Delta\_\{\\overline\{\\partial\}\}f\.\(8\.6\)which we will use in[8\.2](https://arxiv.org/html/2608.19584#S8.SS2)\.
We can note the decomposition in the bilinear form case is
12∇2L\(v,v\)=Re\(∇ω2,0L\(v,v\)\)\+∇ω1,1L\(v,v¯\),\\displaystyle\\frac\{1\}\{2\}\\nabla^\{2\}L\(v,v\)=\\text\{Re\}\(\\nabla^\{2,0\}\_\{\\omega\}L\(v,v\)\)\+\\nabla^\{1,1\}\_\{\\omega\}L\(v,\\overline\{v\}\),\(8\.7\)which is the holomorphic part and its conjugate, and the Hermitian part\.
### 8\.2Dolbeault Hessian bounds
In this section, we will provide a bound oni∂∂¯fi\\partial\\overline\{\\partial\}f, which is connected to the \(1,1\) Hessian∇ω1,1f\\nabla\_\{\\omega\}^\{1,1\}f\. We will draw a connection between these two terms, which will help us in the following sections\.
Proof of Theorem 2\.Letf\(θ,z\):M→ℝf\(\\theta;z\):M\\rightarrow\\mathbb\{R\}, whereθ∈M\\theta\\in Mis a parameter along the Kähler manifoldMMwith Kähler metricω\\omega\. Define the deformed form
ωf=ω\+i∂∂¯f,\\displaystyle\\omega\_\{f\}=\\omega\+i\\partial\\overline\{\\partial\}f,\(8\.8\)whereffis not necessarily SPSH but we haveω\+i∂∂¯f\>0\\omega\+i\\partial\\overline\{\\partial\}f\>0\.
The spectral norm of the Hessian in[10](https://arxiv.org/html/2608.19584#bib.bib1)is analogous to the maximum eigenvalue ofi∂∂¯fi\\partial\\overline\{\\partial\}ffor us, so we will attempt to bound
‖i∂∂¯f‖2,\\displaystyle\\\|i\\partial\\overline\{\\partial\}f\\\|\_\{2\},\(8\.9\)where∥⋅∥2\\\|\\cdot\\\|\_\{2\}is the spectral norm\. Let us defineΨ:M→ℝ\\Psi:M\\rightarrow\\mathbb\{R\}as the logarithmic volume ratio
Ψ:=log\(\(ω\+i∂∂¯f\)KωK\),\\displaystyle\\Psi:=\\log\\left\(\\frac\{\(\\omega\+i\\partial\\overline\{\\partial\}f\)^\{K\}\}\{\\omega^\{K\}\}\\right\),\(8\.10\)\(observe it exists\) which immediately implies the complex Monge\-Ampère equation
\(ω\+i∂∂¯f\)K=eΨωK\.\\displaystyle\(\\omega\+i\\partial\\overline\{\\partial\}f\)^\{K\}=e^\{\\Psi\}\\omega^\{K\}\.\(8\.11\)whereKKis the integer such thatΘ⊆ℂ∑kmkmk\+1\+mL:=ℂK\\Theta\\subseteq\\mathbb\{C\}^\{\\sum\_\{k\}m\_\{k\}m\_\{k\+1\}\+m\_\{L\}\}:=\\mathbb\{C\}^\{K\}, and we use notation⋀k=1Kω=ωK\\bigwedge\_\{k=1\}^\{K\}\\omega=\\omega^\{K\}\.
Our proof strategy for the next step will be to apply the Mean Value Theorem and then bound the gradient ofΨ\\Psiterm\. It is nontrivial to bound‖i∂∂¯f‖2\\\|i\\partial\\overline\{\\partial\}f\\\|\_\{2\}directly, but instead we will show a relation to a bound‖∇ωi∂∂¯f‖2\\\|\\nabla\_\{\\omega\}i\\partial\\overline\{\\partial\}f\\\|\_\{2\}\. Let us examine the oscillation, or supremum and infimum difference, and examine the geodesic ball diameter\. By the Mean Value Theorem,
sup𝒮Ψ−inf𝒮Ψ≤\(supθ∈𝒮‖∇ωΨ‖ω\)×diam\(𝒮\),\\displaystyle\\sup\_\{\\mathcal\{S\}\}\\Psi\-\\inf\_\{\\mathcal\{S\}\}\\Psi\\leq\\left\(\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|\\nabla\_\{\\omega\}\\Psi\\\|\_\{\\omega\}\\right\)\\times\\text\{diam\}\(\\mathcal\{S\}\),\(8\.12\)where𝒮\\mathcal\{S\}is the geodesic ball𝒮=Bω\(θ0,R\)=\{θ∈M\|dω\(θ0,θ\)≤R\}\\mathcal\{S\}=B\_\{\\omega\}\(\\theta\_\{0\},R\)=\\\{\\theta\\in M\|d\_\{\\omega\}\(\\theta\_\{0\},\\theta\)\\leq R\\\}\.RRis fixed, sodiam\(𝒮\)\\text\{diam\}\(\\mathcal\{S\}\)is𝒪\(1\)\\mathcal\{O\}\(1\)\. The above norm is theω\\omega\-dual norm\.
Let us examine
det\(I\+ω−1i∂∂¯f\)=eΨ\.\\displaystyle\\det\(I\+\\omega^\{\-1\}i\\partial\\overline\{\\partial\}f\)=e^\{\\Psi\}\.\(8\.13\)This is true since\(Aω\)∧\(Aω\)∧⋯∧\(Aω\)=det\(A\)\(ω∧ω∧⋯∧ω\)\(A\\omega\)\\wedge\(A\\omega\)\\wedge\\dots\\wedge\(A\\omega\)=\\det\(A\)\(\\omega\\wedge\\omega\\wedge\\dots\\wedge\\omega\), and
\(ω\+i∂∂¯f\)K=det\(I\+ω−1i∂∂¯f\)ωK\.\\displaystyle\(\\omega\+i\\partial\\overline\{\\partial\}f\)^\{K\}=\\det\(I\+\\omega^\{\-1\}i\\partial\\overline\{\\partial\}f\)\\omega^\{K\}\.\(8\.14\)In particular, via eigenvalues
ω=i∑j=1Kdθj∧dθ¯j,i∂∂¯f=i∑j=1Kλjdθj∧dθ¯j\.\\displaystyle\\omega=i\\sum\_\{j=1\}^\{K\}d\\theta^\{j\}\\wedge d\\overline\{\\theta\}^\{j\},\\quad i\\partial\\overline\{\\partial\}f=i\\sum\_\{j=1\}^\{K\}\\lambda\_\{j\}d\\theta^\{j\}\\wedge d\\overline\{\\theta\}^\{j\}\.\(8\.15\)The volume element obeys
ωK=K\!iKdθ1∧dθ¯1∧⋯∧dθK∧dθ¯K,\\displaystyle\\omega^\{K\}=K\!i^\{K\}d\\theta^\{1\}\\wedge d\\overline\{\\theta\}^\{1\}\\wedge\\dots\\wedge d\\theta^\{K\}\\wedge d\\overline\{\\theta\}^\{K\},\(8\.16\)so the deformed metric follows
ω\+i∂∂¯f=i∑j=1K\(1\+λj\)dθj∧dθ¯j\.\\displaystyle\\omega\+i\\partial\\overline\{\\partial\}f=i\\sum\_\{j=1\}^\{K\}\(1\+\\lambda\_\{j\}\)d\\theta^\{j\}\\wedge d\\overline\{\\theta\}^\{j\}\.\(8\.17\)Therefore, taking theKK\-th wedge power, the combinatorial rules follow and
\(ω\+i∂∂¯f\)K=K\!\(∏j=1K\(1\+λj\)\)iKdθ1∧dθ¯1∧⋯∧dθK∧dθ¯K\.\\displaystyle\(\\omega\+i\\partial\\overline\{\\partial\}f\)^\{K\}=K\!\\left\(\\prod\_\{j=1\}^\{K\}\(1\+\\lambda\_\{j\}\)\\right\)i^\{K\}d\\theta^\{1\}\\wedge d\\overline\{\\theta\}^\{1\}\\wedge\\dots\\wedge d\\theta^\{K\}\\wedge d\\overline\{\\theta\}^\{K\}\.\(8\.18\)By definition, the product of eigenvalues is the determinant, so we can match the above product term to the determinant and
\(ω\+i∂∂¯f\)K=det\(I\+ω−1i∂∂¯f\)ωK\.\\displaystyle\(\\omega\+i\\partial\\overline\{\\partial\}f\)^\{K\}=\\det\(I\+\\omega^\{\-1\}i\\partial\\overline\{\\partial\}f\)\\omega^\{K\}\.\(8\.19\)SinceΨ\\Psiis the ratio, we get
eΨ=det\(I\+ω−1i∂∂¯f\)⟹Ψ=logdet\(I\+ω−1i∂∂¯f\)\.\\displaystyle e^\{\\Psi\}=\\det\(I\+\\omega^\{\-1\}i\\partial\\overline\{\\partial\}f\)\\implies\\Psi=\\log\\det\(I\+\\omega^\{\-1\}i\\partial\\overline\{\\partial\}f\)\.\(8\.20\)Now,
∇kΨ=∇klogdet\(I\+ω−1i∂∂¯f\)=1det\(I\+ω−1i∂∂¯f\)∇kdet\(I\+ω−1i∂∂¯f\)\.\\displaystyle\\nabla\_\{k\}\\Psi=\\nabla\_\{k\}\\log\\det\(I\+\\omega^\{\-1\}i\\partial\\overline\{\\partial\}f\)=\\frac\{1\}\{\\det\(I\+\\omega^\{\-1\}i\\partial\\overline\{\\partial\}f\)\}\\nabla\_\{k\}\\det\(I\+\\omega^\{\-1\}i\\partial\\overline\{\\partial\}f\)\.\(8\.21\)Using Jacobi’s formula,
∇kdet\(I\+ω−1i∂∂¯f\)=det\(I\+ω−1i∂∂¯f\)Tr\(\(I\+ω−1i∂∂¯f\)−1∇k\(I\+ω−1i∂∂¯f\)\)\.\\displaystyle\\nabla\_\{k\}\\det\(I\+\\omega^\{\-1\}i\\partial\\overline\{\\partial\}f\)=\\det\(I\+\\omega^\{\-1\}i\\partial\\overline\{\\partial\}f\)\\text\{Tr\}\\Big\(\(I\+\\omega^\{\-1\}i\\partial\\overline\{\\partial\}f\)^\{\-1\}\\nabla\_\{k\}\(I\+\\omega^\{\-1\}i\\partial\\overline\{\\partial\}f\)\\Big\)\.\(8\.22\)Therefore,
∇kΨ\\displaystyle\\nabla\_\{k\}\\Psi=det\(I\+ω−1i∂∂¯f\)det\(I\+ω−1i∂∂¯f\)Tr\(\(I\+ω−1i∂∂¯f\)−1ω−1∇k\(i∂∂¯f\)\)\.\\displaystyle=\\cancel\{\\frac\{\\det\(I\+\\omega^\{\-1\}i\\partial\\overline\{\\partial\}f\)\}\{\\det\(I\+\\omega^\{\-1\}i\\partial\\overline\{\\partial\}f\)\}\}\\text\{Tr\}\(\(I\+\\omega^\{\-1\}i\\partial\\overline\{\\partial\}f\)^\{\-1\}\\omega^\{\-1\}\\nabla\_\{k\}\(i\\partial\\overline\{\\partial\}f\)\)\.\(8\.23\)We have setA=I\+ω−1i∂∂¯fA=I\+\\omega^\{\-1\}i\\partial\\overline\{\\partial\}f\. The trace is not necessarily𝒪\(K\)\\mathcal\{O\}\(K\), it is instead a sum uponKKterms\. Therefore, with suitable geometric assumptions such as decay, the factor ofKKcan be prevented\.
Now, let us assume bounded curvature‖i∂∂¯f\(θ0\)‖2≤C,‖∇ki∂∂¯f\(θ0\)‖2≤CM\\\|i\\partial\\overline\{\\partial\}f\(\\theta\_\{0\}\)\\\|\_\{2\}\\leq C,\\\|\\nabla\_\{k\}i\\partial\\overline\{\\partial\}f\(\\theta\_\{0\}\)\\\|\_\{2\}\\leq C\_\{M\}independent of parameter dimensionKK\. We note nuclear norm‖M‖1≤‖M‖2\\\|M\\\|\_\{1\}\\leq\\\|M\\\|\_\{2\}with the absence ofKK\.KKis a worst\-case bound, and is omitted due to spectral decay\. We get
\|∇kΨ\|≤‖A−1ω−1‖2‖∇k\(i∂∂¯f\)‖1≤‖A−1‖2‖hcross\-entropy−1‖2‖∇k\(i∂∂¯f\)‖2\.\\displaystyle\|\\nabla\_\{k\}\\Psi\|\\leq\\\|A^\{\-1\}\\omega^\{\-1\}\\\|\_\{2\}\\\|\\nabla\_\{k\}\(i\\partial\\overline\{\\partial\}f\)\\\|\_\{1\}\\leq\\\|A^\{\-1\}\\\|\_\{2\}\\\|h\_\{\\text\{cross\-entropy\}\}^\{\-1\}\\\|\_\{2\}\\\|\\nabla\_\{k\}\(i\\partial\\overline\{\\partial\}f\)\\\|\_\{2\}\.\(8\.24\)Note that‖hcross\-entropy−1‖2=‖A−1‖2=𝒪\(1\)\\\|h\_\{\\text\{cross\-entropy\}\}^\{\-1\}\\\|\_\{2\}=\\\|A^\{\-1\}\\\|\_\{2\}=\\mathcal\{O\}\(1\)\. Therefore, evaluating theω\\omega\-dual norm
‖∇ωΨ‖ω2=hij¯∇iΨ∇j¯Ψ≤‖hcross\-entropy−1‖2∑k=1K\|∇kΨ\|2≤C~‖∇ω\(i∂∂¯f\)‖22\.\\displaystyle\\\|\\nabla\_\{\\omega\}\\Psi\\\|\_\{\\omega\}^\{2\}=h^\{i\\overline\{j\}\}\\nabla\_\{i\}\\Psi\\nabla\_\{\\overline\{j\}\}\\Psi\\leq\\\|h\_\{\\text\{cross\-entropy\}\}^\{\-1\}\\\|\_\{2\}\\sum\_\{k=1\}^\{K\}\|\\nabla\_\{k\}\\Psi\|^\{2\}\\leq\\widetilde\{C\}\\\|\\nabla\_\{\\omega\}\(i\\partial\\overline\{\\partial\}f\)\\\|\_\{2\}^\{2\}\.\(8\.25\)
Now, our result in[8\.3](https://arxiv.org/html/2608.19584#S8.SS3)is similar to the recursion rules as in[73](https://arxiv.org/html/2608.19584#bib.bib6)[31](https://arxiv.org/html/2608.19584#bib.bib7)\. Hence, we use
sup‖v‖2=1‖∇v\(i∂∂¯f\)‖2≤‖∂3f‖2\+C2‖i∂∂¯f‖2supθ∈𝒮‖i∂∂¯f‖2\.\\displaystyle\\sup\_\{\\\|v\\\|\_\{2\}=1\}\\\|\\nabla\_\{v\}\(i\\partial\\overline\{\\partial\}f\)\\\|\_\{2\}\\leq\\\|\\partial^\{3\}f\\\|\_\{2\}\+C\_\{2\}\\\|i\\partial\\overline\{\\partial\}f\\\|\_\{2\}\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|i\\partial\\overline\{\\partial\}f\\\|\_\{2\}\.\(8\.26\)Substituting the recursive bound of Lemma 2,
‖∇ωΨ‖ω≤C~\(‖∂3f‖2\+Csupθ∈𝒮‖i∂∂¯f‖22\)\.\\displaystyle\\\|\\nabla\_\{\\omega\}\\Psi\\\|\_\{\\omega\}\\leq\\widetilde\{C\}\\left\(\\\|\\partial^\{3\}f\\\|\_\{2\}\+C\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|i\\partial\\overline\{\\partial\}f\\\|\_\{2\}^\{2\}\\right\)\.\(8\.27\)Returning to the Mean Value Theorem setup,
sup𝒮Ψ−inf𝒮Ψ≤C~\(supθ∈𝒮\(‖∂3f‖2\+C‖i∂∂¯f‖22\)\)diamω\(𝒮\)\.\\displaystyle\\sup\_\{\\mathcal\{S\}\}\\Psi\-\\inf\_\{\\mathcal\{S\}\}\\Psi\\leq\\widetilde\{C\}\\left\(\\sup\_\{\\theta\\in\\mathcal\{S\}\}\(\\\|\\partial^\{3\}f\\\|\_\{2\}\+C\\\|i\\partial\\overline\{\\partial\}f\\\|\_\{2\}^\{2\}\)\\right\)\\text\{diam\}\_\{\\omega\}\(\\mathcal\{S\}\)\.\(8\.28\)Yau’s maximum principle ensures the spectral norm of the complex Hessian is bounded by the oscillation
supθ∈𝒮‖i∂∂¯f‖2≤C0mexp\(sup𝒮Ψ−inf𝒮Ψ\)\.\\displaystyle\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|i\\partial\\overline\{\\partial\}f\\\|\_\{2\}\\leq\\frac\{C\_\{0\}\}\{\\sqrt\{m\}\}\\exp\\left\(\\sup\_\{\\mathcal\{S\}\}\\Psi\-\\inf\_\{\\mathcal\{S\}\}\\Psi\\right\)\.\(8\.29\)By the Yau estimate in[8\.4](https://arxiv.org/html/2608.19584#S8.SS4), we obtain a transcendental inequality
supθ∈𝒮‖i∂∂¯f‖2≤C~mexp\(Dsupθ∈𝒮\(‖∂3f‖2\+C‖i∂∂¯f‖22\)\)\.\\displaystyle\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|i\\partial\\overline\{\\partial\}f\\\|\_\{2\}\\leq\\frac\{\\widetilde\{C\}\}\{\\sqrt\{m\}\}\\exp\\left\(D\\sup\_\{\\theta\\in\\mathcal\{S\}\}\(\\\|\\partial^\{3\}f\\\|\_\{2\}\+C\\\|i\\partial\\overline\{\\partial\}f\\\|\_\{2\}^\{2\}\)\\right\)\.\(8\.30\)Under a Taylor expansion,
supθ∈𝒮‖i∂∂¯f‖2≤C~m\{1\+∑j=1∞1j\!\(Dsupθ∈𝒮\(‖∂3f‖2\+C‖i∂∂¯f‖22\)\)j\}\\displaystyle\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|i\\partial\\overline\{\\partial\}f\\\|\_\{2\}\\leq\\frac\{\\widetilde\{C\}\}\{\\sqrt\{m\}\}\\left\\\{1\+\\sum\_\{j=1\}^\{\\infty\}\\frac\{1\}\{j\!\}\\left\(D\\sup\_\{\\theta\\in\\mathcal\{S\}\}\(\\\|\\partial^\{3\}f\\\|\_\{2\}\+C\\\|i\\partial\\overline\{\\partial\}f\\\|\_\{2\}^\{2\}\)\\right\)^\{j\}\\right\\\}\(8\.31\)Under assumption‖∂3f‖2,‖i∂∂¯f‖22\\\|\\partial^\{3\}f\\\|\_\{2\},\\\|i\\partial\\overline\{\\partial\}f\\\|\_\{2\}^\{2\}are𝒪\(m−α\)\\mathcal\{O\}\(m^\{\-\\alpha\}\),α≥0\\alpha\\geq 0, we have
supθ∈𝒮‖i∂∂¯f‖2≤C~m\(1\+𝒪\(\(g\(m\)\)degree\(g\)≤0\)\+𝒪\(\(g\(m\)\)degree\(g\)≤0\)\)=𝒪\(1m\)\.\\displaystyle\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|i\\partial\\overline\{\\partial\}f\\\|\_\{2\}\\leq\\frac\{\\widetilde\{C\}\}\{\\sqrt\{m\}\}\\left\(1\+\\mathcal\{O\}\(\(g\(m\)\)\_\{\\text\{degree\}\(g\)\\leq 0\}\)\+\\mathcal\{O\}\(\(g\(m\)\)\_\{\\text\{degree\}\(g\)\\leq 0\}\)\\right\)=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)\.\(8\.32\)We remark this argument is an a priori estimate since we only need𝒪\(1\)\\mathcal\{O\}\(1\)in the above, which is reasonable and less restrictive than our result\. We refer to the end of Appendix[8\.3](https://arxiv.org/html/2608.19584#S8.SS3)for references that the above is reasonable\.
□\\square
### 8\.3Dolbeault covariant bound
Figure 6:We illustrate that Lemma 1 holds, where the shaded region represents a valid upper bound, the background color representing the bound gap\. Here, we choose constantC=1C=1\. The x\-axis is the right\-hand side of[5\.2](https://arxiv.org/html/2608.19584#S5.E2)and the y\-axis the left\-hand side\.Proof of Lemma 1\.We apply the covariant derivative to the Hessian compatible with the Kähler metric\. Because the manifold is Kähler, the connection is holomorphic so the mixed symbols such asΓki¯l\{\\Gamma\_\{k\\overline\{i\}\}\}^\{l\}vanish\. The Levi\-Civita connection on the Hessian yields a \(0,3\) tensor
∇k∂i∂j¯f=∂k∂i∂j¯f−Γkil∂l∂j¯f\\displaystyle\\nabla\_\{k\}\\partial\_\{i\}\\partial\_\{\\overline\{j\}\}f=\\partial\_\{k\}\\partial\_\{i\}\\partial\_\{\\overline\{j\}\}f\-\{\\Gamma\_\{ki\}\}^\{l\}\\partial\_\{l\}\\partial\_\{\\overline\{j\}\}f\(8\.33\)and
∇k¯∂i∂j¯f=∂k¯∂i∂j¯f−Γkjl¯∂i∂l¯f\.\\displaystyle\\nabla\_\{\\overline\{k\}\}\\partial\_\{i\}\\partial\_\{\\overline\{j\}\}f=\\partial\_\{\\overline\{k\}\}\\partial\_\{i\}\\partial\_\{\\overline\{j\}\}f\-\\overline\{\{\\Gamma\_\{kj\}\}^\{l\}\}\\partial\_\{i\}\\partial\_\{\\overline\{l\}\}f\.\(8\.34\)By the triangle inequality on the operator norm, and denotingℋij¯=∂i∂j¯f\\mathcal\{H\}\_\{i\\overline\{j\}\}=\\partial\_\{i\}\\partial\_\{\\overline\{j\}\}f,
‖∇ωℋ‖2≤‖∂3f‖2\+‖Γ⋅ℋ‖2\.\\displaystyle\\\|\\nabla\_\{\\omega\}\\mathcal\{H\}\\\|\_\{2\}\\leq\\\|\\partial^\{3\}f\\\|\_\{2\}\+\\\|\\Gamma\\cdot\\mathcal\{H\}\\\|\_\{2\}\.\(8\.35\)We can note the Christoffel symbols are governed byΓkil=hlm¯∂khim¯\{\\Gamma\_\{ki\}\}^\{l\}=h^\{l\\overline\{m\}\}\\partial\_\{k\}h\_\{i\\overline\{m\}\}\. Our proof strategy will be to show the third derivative term is at least of equivalent order as the covariant derivative term, and the Christoffel symbol term is𝒪\(1m\)\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\), therefore we have the recursive relationship\.
The probabilistic modelppobeysp\(y\|z,θ\)=𝒞𝒩\(f\(θ,z\),σ2\)p\(y\|z,\\theta\)=\\mathcal\{CN\}\(f\(\\theta;z\),\\sigma^\{2\}\)\. We get the log likelihoodlogp=−1σ2\|y−f\(θ,z\)\|2\+C=−1σ2\(y−f\)\(y¯−f¯\)\+C\\log p=\-\\frac\{1\}\{\\sigma^\{2\}\}\|y\-f\(\\theta;z\)\|^\{2\}\+C=\-\\frac\{1\}\{\\sigma^\{2\}\}\(y\-f\)\(\\overline\{y\}\-\\overline\{f\}\)\+C\. Denote the loss residualϵ=y−f\\epsilon=y\-f\. The derivative of the log likelihood w\.r\.t\.θ\\thetais
∂ilogp=1σ2\(ϵ¯∂if\+ϵ∂if¯\)\.\\displaystyle\\partial\_\{i\}\\log p=\\frac\{1\}\{\\sigma^\{2\}\}\\left\(\\overline\{\\epsilon\}\\partial\_\{i\}f\+\\epsilon\\partial\_\{i\}\\overline\{f\}\\right\)\.\(8\.36\)
In information geometry, the derivative of the Fisher metric follows the Amari\-Chentsov tensor[61](https://arxiv.org/html/2608.19584#bib.bib27)[29](https://arxiv.org/html/2608.19584#bib.bib28), although recall from section[4\.1](https://arxiv.org/html/2608.19584#S4.SS1)our metric is more closely a cross\-entropy one\. Therefore, we can drop this term and we get
∂khim¯=𝔼q\[∂2logp∂θk∂θi∂logp∂θ¯m\]\+𝔼q\[∂logp∂θi∂2logp∂θk∂θ¯m\]−𝔼q\[∂k\(∂i∂m¯pp\)\]\.\\displaystyle\\partial\_\{k\}h\_\{i\\overline\{m\}\}=\\mathbb\{E\}\_\{q\}\\left\[\\frac\{\\partial^\{2\}\\log p\}\{\\partial\\theta^\{k\}\\partial\\theta^\{i\}\}\\frac\{\\partial\\log p\}\{\\partial\\overline\{\\theta\}^\{m\}\}\\right\]\+\\mathbb\{E\}\_\{q\}\\left\[\\frac\{\\partial\\log p\}\{\\partial\\theta^\{i\}\}\\frac\{\\partial^\{2\}\\log p\}\{\\partial\\theta^\{k\}\\partial\\overline\{\\theta\}^\{m\}\}\\right\]\-\\mathbb\{E\}\_\{q\}\\left\[\\partial\_\{k\}\\left\(\\frac\{\\partial\_\{i\}\\partial\_\{\\overline\{m\}\}p\}\{p\}\\right\)\\right\]\.\(8\.37\)This partially follows from
−∂i∂m¯logp=∂ilogp∂m¯logp−∂i∂m¯pp\\displaystyle\-\\partial\_\{i\}\\partial\_\{\\overline\{m\}\}\\log p=\\partial\_\{i\}\\log p\\partial\_\{\\overline\{m\}\}\\log p\-\\frac\{\\partial\_\{i\}\\partial\_\{\\overline\{m\}\}p\}\{p\}\(8\.38\)and differentiating\. Let us examine the first term\. Differentiatinglogp\\log pagain,
∂ki2logp=1σ2\(−∂kf¯∂if\+ϵ¯∂ki2f−∂kf∂if¯\+ϵ∂ki2f¯\)\.\\displaystyle\\partial^\{2\}\_\{ki\}\\log p=\\frac\{1\}\{\\sigma^\{2\}\}\\left\(\-\\partial\_\{k\}\\overline\{f\}\\partial\_\{i\}f\+\\overline\{\\epsilon\}\\partial^\{2\}\_\{ki\}f\-\\partial\_\{k\}f\\partial\_\{i\}\\overline\{f\}\+\\epsilon\\partial^\{2\}\_\{ki\}\\overline\{f\}\\right\)\.\(8\.39\)Evaluating the first term of∂khim¯\\partial\_\{k\}h\_\{i\\overline\{m\}\},
𝔼q\[∂ki2logp∂m¯logp\]=1σ4𝔼q\[\(ϵ¯∂ki2f\+ϵ∂ki2f¯−∂kf¯∂if−∂kf∂if¯\)\(ϵ¯∂m¯f\+ϵ∂m¯f¯\)\]\.\\displaystyle\\mathbb\{E\}\_\{q\}\\left\[\\partial^\{2\}\_\{ki\}\\log p\\partial\_\{\\overline\{m\}\}\\log p\\right\]=\\frac\{1\}\{\\sigma^\{4\}\}\\mathbb\{E\}\_\{q\}\\left\[\\left\(\\overline\{\\epsilon\}\\partial^\{2\}\_\{ki\}f\+\\epsilon\\partial^\{2\}\_\{ki\}\\overline\{f\}\-\\partial\_\{k\}\\overline\{f\}\\partial\_\{i\}f\-\\partial\_\{k\}f\\partial\_\{i\}\\overline\{f\}\\right\)\\left\(\\overline\{\\epsilon\}\\partial\_\{\\overline\{m\}\}f\+\\epsilon\\partial\_\{\\overline\{m\}\}\\overline\{f\}\\right\)\\right\]\.\(8\.40\)By Wick’s theorem, isolated terms inϵ,ϵ¯\\epsilon,\\overline\{\\epsilon\}vanish and what remains is𝔼q\[ϵϵ¯\]=σ2\\mathbb\{E\}\_\{q\}\[\\epsilon\\overline\{\\epsilon\}\]=\\sigma^\{2\}\. The surviving components are
𝔼q\[∂ki2logp∂m¯logp\]=1σ2𝔼q\[∂ki2f∂m¯f¯\+∂ki2f¯∂m¯f\]\.\\displaystyle\\mathbb\{E\}\_\{q\}\\Bigg\[\\partial^\{2\}\_\{ki\}\\log p\\partial\_\{\\overline\{m\}\}\\log p\\Bigg\]=\\frac\{1\}\{\\sigma^\{2\}\}\\mathbb\{E\}\_\{q\}\\Bigg\[\\partial^\{2\}\_\{ki\}f\\partial\_\{\\overline\{m\}\}\\overline\{f\}\+\\partial^\{2\}\_\{ki\}\\overline\{f\}\\partial\_\{\\overline\{m\}\}f\\Bigg\]\.\(8\.41\)A similar argument for the third term of∂khim¯\\partial\_\{k\}h\_\{i\\overline\{m\}\}yields
𝔼q\[∂ilogp∂km¯2logp\]=1σ2𝔼q\[∂if∂km¯2f¯\+∂if¯∂km¯2f\]\.\\displaystyle\\mathbb\{E\}\_\{q\}\\Bigg\[\\partial\_\{i\}\\log p\\partial^\{2\}\_\{k\\overline\{m\}\}\\log p\\Bigg\]=\\frac\{1\}\{\\sigma^\{2\}\}\\mathbb\{E\}\_\{q\}\\Bigg\[\\partial\_\{i\}f\\partial^\{2\}\_\{k\\overline\{m\}\}\\overline\{f\}\+\\partial\_\{i\}\\overline\{f\}\\partial^\{2\}\_\{k\\overline\{m\}\}f\\Bigg\]\.\(8\.42\)Evaluating the third term,
−𝔼q\[∂k\(∂i∂m¯pp\)\]=1σ2𝔼q\[∂kf¯∂im¯2f\+∂kf∂im¯2f¯\]\.\\displaystyle\-\\mathbb\{E\}\_\{q\}\\left\[\\partial\_\{k\}\\left\(\\frac\{\\partial\_\{i\}\\partial\_\{\\overline\{m\}\}p\}\{p\}\\right\)\\right\]=\\frac\{1\}\{\\sigma^\{2\}\}\\mathbb\{E\}\_\{q\}\\Bigg\[\\partial\_\{k\}\\overline\{f\}\\partial^\{2\}\_\{i\\overline\{m\}\}f\+\\partial\_\{k\}f\\partial^\{2\}\_\{i\\overline\{m\}\}\\overline\{f\}\\Bigg\]\.\(8\.43\)Consolidating the terms, we scale byKKto align with the normalized metrich=Kh~h=K\\widetilde\{h\}from[8\.2](https://arxiv.org/html/2608.19584#S8.SS2), orh−1=1Kh~−1h^\{\-1\}=\\frac\{1\}\{K\}\\widetilde\{h\}^\{\-1\},
∂khim¯=Kσ2𝔼q\[ℋki∂m¯f¯\+∂ki2f¯∂m¯f\+∂ifℋkm¯†\+∂if¯ℋkm¯\+∂kf¯ℋim¯\+∂kfℋim¯†\]\.\\displaystyle\\partial\_\{k\}h\_\{i\\overline\{m\}\}=\\frac\{K\}\{\\sigma^\{2\}\}\\mathbb\{E\}\_\{q\}\\Bigg\[\\mathcal\{H\}\_\{ki\}\\partial\_\{\\overline\{m\}\}\\overline\{f\}\+\\partial^\{2\}\_\{ki\}\\overline\{f\}\\partial\_\{\\overline\{m\}\}f\+\\partial\_\{i\}f\\mathcal\{H\}\_\{k\\overline\{m\}\}^\{\\dagger\}\+\\partial\_\{i\}\\overline\{f\}\\mathcal\{H\}\_\{k\\overline\{m\}\}\+\\partial\_\{k\}\\overline\{f\}\\mathcal\{H\}\_\{i\\overline\{m\}\}\+\\partial\_\{k\}f\\mathcal\{H\}\_\{i\\overline\{m\}\}^\{\\dagger\}\\Bigg\]\.\(8\.44\)Taking the norm, by Jensen’s inequality, and taking the supremum it follows
‖∂h‖2≤Kσ2𝔼q\[2‖∂2,0f‖2‖∂f‖2\+4‖∂f‖2‖ℋ‖2\]\.\\displaystyle\\\|\\partial h\\\|\_\{2\}\\leq\\frac\{K\}\{\\sigma^\{2\}\}\\mathbb\{E\}\_\{q\}\\Bigg\[2\\\|\\partial^\{2,0\}f\\\|\_\{2\}\\\|\\partial f\\\|\_\{2\}\+4\\\|\\partial f\\\|\_\{2\}\\\|\\mathcal\{H\}\\\|\_\{2\}\\Bigg\]\.\(8\.45\)By definition, the Dolbeault Hessian norm is exactly‖ℋ‖2=‖i∂∂¯f‖2=‖∂km¯2f‖2\\\|\\mathcal\{H\}\\\|\_\{2\}=\\\|i\\partial\\overline\{\\partial\}f\\\|\_\{2\}=\\\|\\partial^\{2\}\_\{k\\overline\{m\}\}f\\\|\_\{2\}\. We will allow the spectral norm of the \(2,0\) derivative to be bounded proportionally by the Dolbeault Hessian, meaning there exists a constantC′\>0C^\{\\prime\}\>0such that‖∂ki2f‖2≤C′‖ℋ‖2\\\|\\partial^\{2\}\_\{ki\}f\\\|\_\{2\}\\leq C^\{\\prime\}\\\|\\mathcal\{H\}\\\|\_\{2\}\. Taking the norm and a supremum to account for the expectation,
‖∂khim¯‖2≤KCsupθ∈𝒮‖ℋ‖2‖∂f‖2\.\\displaystyle\\\|\\partial\_\{k\}h\_\{i\\overline\{m\}\}\\\|\_\{2\}\\leq KC\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|\\mathcal\{H\}\\\|\_\{2\}\\\|\\partial f\\\|\_\{2\}\.\(8\.46\)Recall the Christoffel symbols are defined viaΓkil=hlm¯∂khim¯\{\\Gamma\_\{ki\}\}^\{l\}=h^\{l\\overline\{m\}\}\\partial\_\{k\}h\_\{i\\overline\{m\}\}\. We can note Christoffel symbols are invariant to constant rescalings of the metric, so\(c−1hlm¯\)∂k\(c⋅him¯\)=hlm¯∂khim¯\(c^\{\-1\}h^\{l\\overline\{m\}\}\)\\partial\_\{k\}\(c\\cdot h\_\{i\\overline\{m\}\}\)=h^\{l\\overline\{m\}\}\\partial\_\{k\}h\_\{i\\overline\{m\}\}\. The operator norm of the inverse metric is bounded𝒪\(K−1\)\\mathcal\{O\}\(K^\{\-1\}\)\. Therefore,
sup‖v‖2=1‖Γ\(v,⋅\)‖2≤K⋅1K\|h−1\|supθ∈𝒮2‖∂khim¯‖2≤C‖h−1‖supθ∈𝒮2‖∂f‖2‖ℋ‖2\.\\displaystyle\\sup\_\{\\\|v\\\|\_\{2\}=1\}\\\|\\Gamma\(v,\\cdot\)\\\|\_\{2\}\\leq K\\cdot\\frac\{1\}\{K\}\\\|h^\{\-1\}\\\|\_\{2\}\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|\\partial\_\{k\}h\_\{i\\overline\{m\}\}\\\|\_\{2\}\\leq C\\\|h^\{\-1\}\\\|\_\{2\}\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|\\partial f\\\|\_\{2\}\\\|\\mathcal\{H\}\\\|\_\{2\}\.\(8\.47\)It immediately follows from the neural network definition the Jacobian offfextracts an𝒪\(1\)\\mathcal\{O\}\(1\)scaling factor, thus we conclude the bound on the Christoffel symbols
sup‖v‖2=1‖Γ\(v,⋅\)‖2≤C1supθ∈𝒮‖ℋ‖2\.\\displaystyle\\sup\_\{\\\|v\\\|\_\{2\}=1\}\\\|\\Gamma\(v,\\cdot\)\\\|\_\{2\}\\leq C\_\{1\}\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|\\mathcal\{H\}\\\|\_\{2\}\.\(8\.48\)Returning to the triangle inequality, we have
‖Γ⋅ℋ‖2≤sup‖v‖2=1‖Γ\(v,⋅\)‖2‖ℋ‖2≤C1‖ℋ‖2supθ∈𝒮‖ℋ‖2\.\\displaystyle\\\|\\Gamma\\cdot\\mathcal\{H\}\\\|\_\{2\}\\leq\\sup\_\{\\\|v\\\|\_\{2\}=1\}\\\|\\Gamma\(v,\\cdot\)\\\|\_\{2\}\\\|\\mathcal\{H\}\\\|\_\{2\}\\leq C\_\{1\}\\\|\\mathcal\{H\}\\\|\_\{2\}\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|\\mathcal\{H\}\\\|\_\{2\}\.\(8\.49\)
Next, we address the third derivative term‖∂3f‖2\\\|\\partial^\{3\}f\\\|\_\{2\}\. It suffices to show‖∂3f‖\\\|\\partial^\{3\}f\\\|is lower than𝒪\(1\)\\mathcal\{O\}\(1\)\. By[10](https://arxiv.org/html/2608.19584#bib.bib1)we have the real Hessian is𝒪\(1m\)\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)and a neural network will not increase the order to𝒪\(m\)\\mathcal\{O\}\(\\sqrt\{m\}\)\. This phenomenon is known to be true such as in[2](https://arxiv.org/html/2608.19584#bib.bib5)[32](https://arxiv.org/html/2608.19584#bib.bib9)[23](https://arxiv.org/html/2608.19584#bib.bib2)[32](https://arxiv.org/html/2608.19584#bib.bib9)[31](https://arxiv.org/html/2608.19584#bib.bib7)[36](https://arxiv.org/html/2608.19584#bib.bib10)[7](https://arxiv.org/html/2608.19584#bib.bib11)[13](https://arxiv.org/html/2608.19584#bib.bib12)\.
□\\square
Figure 7:We illustrate a proportion relationship between the \(1,1\) and \(2,0\)\-Hessians\. We plot‖ℋ1,1‖2,‖ℋ2,0‖2\\\|\\mathcal\{H\}^\{1,1\}\\\|\_\{2\},\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}exactly without the constant\. The line is the best fit line\. We evaluate at 80 distinct parameter points\.Remark\.If we consider the deformed metricωf=ω\+i∂∂¯f\\omega\_\{f\}=\\omega\+i\\partial\\overline\{\\partial\}f, construct a tensor representing the difference between the Chern connections of the deformed metric and the base metric\. We get a difference formula
Ψijk=Ψijk\(ωf,ω\):=Γ\(ωf\)ijk−Γ\(ω\)ijk=\(ωf\)kl¯∇iω\(ωf\)jl¯\.\\displaystyle\\Psi^\{k\}\_\{ij\}=\\Psi^\{k\}\_\{ij\}\(\\omega\_\{f\},\\omega\):=\\Gamma\(\\omega\_\{f\}\)^\{k\}\_\{ij\}\-\\Gamma\(\\omega\)^\{k\}\_\{ij\}=\(\\omega\_\{f\}\)^\{k\\overline\{l\}\}\\nabla^\{\\omega\}\_\{i\}\(\\omega\_\{f\}\)\_\{j\\overline\{l\}\}\.\(8\.50\)Because the base metric is covariantly constant with respect to its own connection, we can write
∇iω\(ωf\)jl¯=∇iω\(ωjl¯\+ℋjl¯\)=∇iωℋjl¯\.\\displaystyle\\nabla^\{\\omega\}\_\{i\}\(\\omega\_\{f\}\)\_\{j\\overline\{l\}\}=\\nabla^\{\\omega\}\_\{i\}\(\\omega\_\{j\\overline\{l\}\}\+\\mathcal\{H\}\_\{j\\overline\{l\}\}\)=\\nabla^\{\\omega\}\_\{i\}\\mathcal\{H\}\_\{j\\overline\{l\}\}\.\(8\.51\)Therefore, we can see
Ψijk=\(ω\+ℋ\)kl¯∇iωℋjl¯\.\\displaystyle\\Psi^\{k\}\_\{ij\}=\(\\omega\+\\mathcal\{H\}\)^\{k\\overline\{l\}\}\\nabla^\{\\omega\}\_\{i\}\\mathcal\{H\}\_\{j\\overline\{l\}\}\.\(8\.52\)Taking the norm, and using[5\.2](https://arxiv.org/html/2608.19584#S5.E2)from the lemma
‖Ψ‖2\\displaystyle\\\|\\Psi\\\|\_\{2\}≤‖\(ω\+ℋ\)−1‖2⋅‖∇ωℋ‖2\\displaystyle\\leq\\\|\(\\omega\+\\mathcal\{H\}\)^\{\-1\}\\\|\_\{2\}\\cdot\\\|\\nabla\_\{\\omega\}\\mathcal\{H\}\\\|\_\{2\}\(8\.53\)≤‖\(ω\+ℋ\)−1‖2\(‖∂3f‖2\+C‖ℋ‖2supθ∈𝒮‖ℋ‖2\)\.\\displaystyle\\leq\\\|\(\\omega\+\\mathcal\{H\}\)^\{\-1\}\\\|\_\{2\}\\left\(\\\|\\partial^\{3\}f\\\|\_\{2\}\+C\\\|\\mathcal\{H\}\\\|\_\{2\}\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|\\mathcal\{H\}\\\|\_\{2\}\\right\)\.\(8\.54\)
### 8\.4SpectralC2C^\{2\}estimate
Let us use the fact that the scaling limit ofℋ\(θ0\)=i∂∂¯f\(θ0\)\\mathcal\{H\}\(\\theta\_\{0\}\)=i\\partial\\overline\{\\partial\}f\(\\theta\_\{0\}\)obeys at initialization a pointwise bound[10](https://arxiv.org/html/2608.19584#bib.bib1)[44](https://arxiv.org/html/2608.19584#bib.bib13)[66](https://arxiv.org/html/2608.19584#bib.bib15)[45](https://arxiv.org/html/2608.19584#bib.bib16)
‖ℋ\(θ0\)‖2≤Cinitm\.\\displaystyle\\\|\\mathcal\{H\}\(\\theta\_\{0\}\)\\\|\_\{2\}\\leq\\frac\{C\_\{\\text\{init\}\}\}\{\\sqrt\{m\}\}\.\(8\.55\)This is a localized fact, and the burden of the proof is shifted to showing that the Monge\-Ampère dynamics prevent the Hessian from escaping this scale globally asθ\\thetamoves across𝒮\\mathcal\{S\}\.
We avoid use of the trace functional since trace is affected byKK, thus our final bound would involveKK, which is undesirable\. We need only work with the spectral norm\. Therefore, we find a maximum principle applicable for us\.
Proof of Lemma 2\.Define the test function
v\(θ,t\)=log\(λmax\(ℋ\(θt\)\)\)−Ψ\(θt\)\.\\displaystyle v\(\\theta,t\)=\\log\(\\lambda\_\{\\max\}\(\\mathcal\{H\}\(\\theta\_\{t\}\)\)\)\-\\Psi\(\\theta\_\{t\}\)\.\(8\.56\)We can noteθt\\theta\_\{t\}depends on time since it is along a path, but the manifold does not depend on time\. Let\(θ∗,t∗\)\(\\theta^\{\*\},t^\{\*\}\)be the point in𝒮×\[0,T\]\\mathcal\{S\}\\times\[0,T\]where the test function attains its global maximum\. We want to provevvis highest att=0t=0\. We proceed by contradiction\. Assume that the maximum occurs after initialization, such thatt∗\>0t^\{\*\}\>0\. At this interior maximum, we must have∂tv\(θ∗,t∗\)≥0\\partial\_\{t\}v\(\\theta^\{\*\},t^\{\*\}\)\\geq 0andΔω~v\(θ∗,t∗\)≤0\\Delta\_\{\\widetilde\{\\omega\}\}v\(\\theta^\{\*\},t^\{\*\}\)\\leq 0\. Consequently, at\(θ∗,t∗\)\(\\theta^\{\*\},t^\{\*\}\), we must have
\(∂t−Δω~\)v≥0,\\displaystyle\(\\partial\_\{t\}\-\\Delta\_\{\\widetilde\{\\omega\}\}\)v\\geq 0,\(8\.57\)which contradicts our assumption\. This showsv\(θt,t\)≤v\(θ0,0\)v\(\\theta\_\{t\},t\)\\leq v\(\\theta\_\{0\},0\)\.
From the contradiction,
log\(λmax\(ℋ\(θt\)\)\)−Ψ\(θt\)≤supθ0∈𝒮\(log\(λmax\(ℋ\(θ0\)\)\)−Ψ\(θ0\)\)\.\\displaystyle\\log\(\\lambda\_\{\\max\}\(\\mathcal\{H\}\(\\theta\_\{t\}\)\)\)\-\\Psi\(\\theta\_\{t\}\)\\leq\\sup\_\{\\theta\_\{0\}\\in\\mathcal\{S\}\}\\left\(\\log\(\\lambda\_\{\\max\}\(\\mathcal\{H\}\(\\theta\_\{0\}\)\)\)\-\\Psi\(\\theta\_\{0\}\)\\right\)\.\(8\.58\)We can rewrite
log\(λmax\(ℋ\(θt\)\)\)≤supθ0∈𝒮log\(λmax\(ℋ\(θ0\)\)\)\+Ψ\(θt\)−infθ0∈𝒮Ψ\(θ0\)\.\\displaystyle\\log\(\\lambda\_\{\\max\}\(\\mathcal\{H\}\(\\theta\_\{t\}\)\)\)\\leq\\sup\_\{\\theta\_\{0\}\\in\\mathcal\{S\}\}\\log\(\\lambda\_\{\\max\}\(\\mathcal\{H\}\(\\theta\_\{0\}\)\)\)\+\\Psi\(\\theta\_\{t\}\)\-\\inf\_\{\\theta\_\{0\}\\in\\mathcal\{S\}\}\\Psi\(\\theta\_\{0\}\)\.\(8\.59\)By definition of the spectral norm and exponentiating,
‖ℋ\(θt\)‖2≤\(supθ0∈𝒮‖ℋ\(θ0\)‖2\)exp\(Ψ\(θt\)−inf𝒮Ψ\)\.\\displaystyle\\\|\\mathcal\{H\}\(\\theta\_\{t\}\)\\\|\_\{2\}\\leq\\left\(\\sup\_\{\\theta\_\{0\}\\in\\mathcal\{S\}\}\\\|\\mathcal\{H\}\(\\theta\_\{0\}\)\\\|\_\{2\}\\right\)\\exp\\left\(\\Psi\(\\theta\_\{t\}\)\-\\inf\_\{\\mathcal\{S\}\}\\Psi\\right\)\.\(8\.60\)Taking the supremum over allθt∈𝒮\\theta\_\{t\}\\in\\mathcal\{S\}on both sides yields as desired
supθ∈𝒮‖ℋ‖2≤supθ0∈𝒮‖ℋ\(θ0\)‖2exp\(sup𝒮Ψ−inf𝒮Ψ\)\.\\displaystyle\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|\\mathcal\{H\}\\\|\_\{2\}\\leq\\sup\_\{\\theta\_\{0\}\\in\\mathcal\{S\}\}\\\|\\mathcal\{H\}\(\\theta\_\{0\}\)\\\|\_\{2\}\\exp\\left\(\\sup\_\{\\mathcal\{S\}\}\\Psi\-\\inf\_\{\\mathcal\{S\}\}\\Psi\\right\)\.\(8\.61\)
□\\square
Remark\.We can find a lower bound on the maximum eigenvalue\. Becauseθ∗\\theta^\{\*\}is a maximum for this function, its first derivative vanishes, meaning∇\(logℋ11¯−Ψ\)=0\\nabla\(\\log\\mathcal\{H\}\_\{1\\overline\{1\}\}\-\\Psi\)=0, which yields
∇ℋ11¯ℋ11¯=∇Ψ\.\\displaystyle\\frac\{\\nabla\\mathcal\{H\}\_\{1\\overline\{1\}\}\}\{\\mathcal\{H\}\_\{1\\overline\{1\}\}\}=\\nabla\\Psi\.\(8\.62\)By the maximum principle,
0≥Δω~\(log\(ℋ11¯\)−Ψ\)\.\\displaystyle 0\\geq\\Delta\_\{\\widetilde\{\\omega\}\}\\left\(\\log\(\\mathcal\{H\}\_\{1\\overline\{1\}\}\)\-\\Psi\\right\)\.\(8\.63\)Under the complex Monge\-Ampère dynamicslogdet\(ω\+ℋ\)=Ψ\\log\\det\(\\omega\+\\mathcal\{H\}\)=\\Psi, we differentiate the equation twice in the direction of this maximum eigenvector\. Because our background metricωflat\\omega\_\{\\text\{flat\}\}is flat, the covariant derivatives commute, yielding a differential inequality
Δω~ℋ11¯≥∂1∂1¯Ψ\|θ∗\.\\displaystyle\\Delta\_\{\\widetilde\{\\omega\}\}\\mathcal\{H\}\_\{1\\overline\{1\}\}\\geq\\partial\_\{1\}\\partial\_\{\\overline\{1\}\}\\Psi\\Big\|\_\{\\theta^\{\*\}\}\.\(8\.64\)The above is on a single component rather than a trace, so it is unaffected by parameter dimension\. The above is closely an elliptic differential inequality\. Now, via chain rule
Δω~log\(ℋ11¯\)=Δω~ℋ11¯ℋ11¯−‖∇ℋ11¯‖ω~2ℋ11¯2\.\\displaystyle\\Delta\_\{\\widetilde\{\\omega\}\}\\log\(\\mathcal\{H\}\_\{1\\overline\{1\}\}\)=\\frac\{\\Delta\_\{\\widetilde\{\\omega\}\}\\mathcal\{H\}\_\{1\\overline\{1\}\}\}\{\\mathcal\{H\}\_\{1\\overline\{1\}\}\}\-\\frac\{\\\|\\nabla\\mathcal\{H\}\_\{1\\overline\{1\}\}\\\|\_\{\\widetilde\{\\omega\}\}^\{2\}\}\{\\mathcal\{H\}\_\{1\\overline\{1\}\}^\{2\}\}\.\(8\.65\)Substituting our gradient equality and differential inequality,
Δω~log\(ℋ11¯\)≥∂1∂1¯Ψℋ11¯−‖∇Ψ‖ω~2\>−∞\.\\displaystyle\\Delta\_\{\\widetilde\{\\omega\}\}\\log\(\\mathcal\{H\}\_\{1\\overline\{1\}\}\)\\geq\\frac\{\\partial\_\{1\}\\partial\_\{\\overline\{1\}\}\\Psi\}\{\\mathcal\{H\}\_\{1\\overline\{1\}\}\}\-\\\|\\nabla\\Psi\\\|\_\{\\widetilde\{\\omega\}\}^\{2\}\>\-\\infty\.\(8\.66\)Rearranging gives a lower bound onℋ11¯\\mathcal\{H\}\_\{1\\overline\{1\}\}\.
### 8\.5∇ω2,0f\\nabla\_\{\\omega\}^\{2,0\}fbounds
Proof of Lemma 3\.As before, let us begin with the property
‖∇ωℋ2,0‖2≤C1‖∂3f‖2\+C2supθ∈𝒮‖ℋ2,0‖22\.\\displaystyle\\\|\\nabla\_\{\\omega\}\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}\\leq C\_\{1\}\\\|\\partial^\{3\}f\\\|\_\{2\}\+C\_\{2\}\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}^\{2\}\.\(8\.67\)This bound follows similarly to[8\.3](https://arxiv.org/html/2608.19584#S8.SS3), but we can note
∂khim¯=1σ2𝔼q\[ℋki2,0∂m¯f¯\+∂ki2f¯∂m¯f\+∂ifℋkm¯1,1†\+∂if¯ℋkm¯1,1\],\\displaystyle\\partial\_\{k\}h\_\{i\\overline\{m\}\}=\\frac\{1\}\{\\sigma^\{2\}\}\\mathbb\{E\}\_\{q\}\\left\[\\mathcal\{H\}^\{2,0\}\_\{ki\}\\partial\_\{\\overline\{m\}\}\\overline\{f\}\+\\partial^\{2\}\_\{ki\}\\overline\{f\}\\partial\_\{\\overline\{m\}\}f\+\\partial\_\{i\}f\\mathcal\{H\}\_\{k\\overline\{m\}\}^\{1,1\\dagger\}\+\\partial\_\{i\}\\overline\{f\}\\mathcal\{H\}^\{1,1\}\_\{k\\overline\{m\}\}\\right\],\(8\.68\)and
‖∂khim¯‖2≤2σ2𝔼q\[‖ℋki2,0‖2‖∂m¯f‖2\+‖∂if‖2‖ℋkm¯1,1‖2\]\.\\displaystyle\\\|\\partial\_\{k\}h\_\{i\\overline\{m\}\}\\\|\_\{2\}\\leq\\frac\{2\}\{\\sigma^\{2\}\}\\mathbb\{E\}\_\{q\}\\left\[\\\|\\mathcal\{H\}^\{2,0\}\_\{ki\}\\\|\_\{2\}\\\|\\partial\_\{\\overline\{m\}\}f\\\|\_\{2\}\+\\\|\\partial\_\{i\}f\\\|\_\{2\}\\\|\\mathcal\{H\}^\{1,1\}\_\{k\\overline\{m\}\}\\\|\_\{2\}\\right\]\.\(8\.69\)Moreover, we assume‖ℋ1,1‖2≤C~‖ℋ2,0‖2\\\|\\mathcal\{H\}^\{1,1\}\\\|\_\{2\}\\leq\\widetilde\{C\}\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}, and the remainder of the proof is identical, so we omit the details\.
Let us consider a geodesic along the Kähler manifold within the geodesic ball𝒮=Bω\(θ0,R\)\\mathcal\{S\}=B\_\{\\omega\}\(\\theta\_\{0\},R\)\. Let us denote the diameterD=2RD=2R\. We can note the absolute derivative is equal to the covariant derivative along the tangent vectorDdtℋ2,0=∇γ˙\(t\)ℋ2,0\\frac\{D\}\{dt\}\\mathcal\{H\}^\{2,0\}=\\nabla\_\{\\dot\{\\gamma\}\(t\)\}\\mathcal\{H\}^\{2,0\}\. Now, we can note
\|ddt‖ℋ2,0\(γ\(t\)\)‖2\|\\displaystyle\\left\|\\frac\{d\}\{dt\}\\\|\\mathcal\{H\}^\{2,0\}\(\\gamma\(t\)\)\\\|\_\{2\}\\right\|≤‖∇γ˙\(t\)ℋ2,0‖2\\displaystyle\\leq\\\|\\nabla\_\{\\dot\{\\gamma\}\(t\)\}\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}\(8\.70\)≤‖∇ωℋ2,0‖2‖γ˙\(t\)‖ω\\displaystyle\\leq\\\|\\nabla\_\{\\omega\}\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}\\\|\\dot\{\\gamma\}\(t\)\\\|\_\{\\omega\}\(8\.71\)=‖∇ωℋ2,0‖2\\displaystyle=\\\|\\nabla\_\{\\omega\}\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}\(8\.72\)≤C1‖∂3f‖2\+C2supθ∈𝒮‖ℋ2,0‖22\.\\displaystyle\\leq C\_\{1\}\\\|\\partial^\{3\}f\\\|\_\{2\}\+C\_\{2\}\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}^\{2\}\.\(8\.73\)The first inequality is by a geometric Kato inequality[34](https://arxiv.org/html/2608.19584#bib.bib8), the second inequality follows from properties of 2\-norms on the covariant derivative, and the equality follows from the fact that the curve is unit speed\. Let us denoteA=supθ∈𝒮‖∂3f‖2,S=supθ∈𝒮‖ℋ2,0‖2A=\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|\\partial^\{3\}f\\\|\_\{2\},S=\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}for short\.
The\(2,0\)\(2,0\)Hessian at initialization obeys𝒪\(1m\)\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)scaling, and the third derivative obeys at least as high scaling\. We integrate the ordinary differential equation
ddt‖ℋ2,0\(γ\(t\)\)‖2≤C1A\+C2S2\.\\displaystyle\\frac\{d\}\{dt\}\\\|\\mathcal\{H\}^\{2,0\}\(\\gamma\(t\)\)\\\|\_\{2\}\\leq C\_\{1\}A\+C\_\{2\}S^\{2\}\.\(8\.74\)Integrating this fromt=0t=0tot≤Dt\\leq Dgives the bound
‖ℋ2,0\(γ\(t\)\)‖2≤‖ℋ2,0\(θ0\)‖2\+D\(C1A\+C2S2\)\.\\displaystyle\\\|\\mathcal\{H\}^\{2,0\}\(\\gamma\(t\)\)\\\|\_\{2\}\\leq\\\|\\mathcal\{H\}^\{2,0\}\(\\theta\_\{0\}\)\\\|\_\{2\}\+D\(C\_\{1\}A\+C\_\{2\}S^\{2\}\)\.\(8\.75\)Therefore, we can note
S≤‖ℋ2,0\(θ0\)‖2\+DC1A\+DC2S2\.\\displaystyle S\\leq\\\|\\mathcal\{H\}^\{2,0\}\(\\theta\_\{0\}\)\\\|\_\{2\}\+DC\_\{1\}A\+DC\_\{2\}S^\{2\}\.\(8\.76\)Rearranging,
\(DC2\)S2−S\+\(‖ℋ2,0\(θ0\)‖2\+DC1A\)≥0\.\\displaystyle\(DC\_\{2\}\)S^\{2\}\-S\+\\left\(\\\|\\mathcal\{H\}^\{2,0\}\(\\theta\_\{0\}\)\\\|\_\{2\}\+DC\_\{1\}A\\right\)\\geq 0\.\(8\.77\)For the inequality to hold,SSmust lie outside the roots of the corresponding parabola\. LetK=‖ℋ2,0\(θ0\)‖\+DC1AK=\\\|\\mathcal\{H\}^\{2,0\}\(\\theta\_\{0\}\)\\\|\+DC\_\{1\}A\. By assumption, both the initial Hessian andAAscale as𝒪\(1m\)\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\), thereforeK=𝒪\(1m\)K=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)\. The roots of the quadraticDC2S2−S\+K=0DC\_\{2\}S^\{2\}\-S\+K=0are given by
S±=1±1−4DC2K2DC2\.\\displaystyle S\_\{\\pm\}=\\frac\{1\\pm\\sqrt\{1\-4DC\_\{2\}K\}\}\{2DC\_\{2\}\}\.\(8\.78\)By continuity of the Hessian with respect to the initial scaling,SSmust remain on the lower branchS≤S−S\\leq S\_\{\-\}\. Applying the identity1−1−x=x1\+1−x1\-\\sqrt\{1\-x\}=\\frac\{x\}\{1\+\\sqrt\{1\-x\}\}, we find
S−\\displaystyle S\_\{\-\}=1−1−4DC2K2DC2\\displaystyle=\\frac\{1\-\\sqrt\{1\-4DC\_\{2\}K\}\}\{2DC\_\{2\}\}\(8\.79\)=4DC2K2DC2\(1\+1−4DC2K\)\\displaystyle=\\frac\{4DC\_\{2\}K\}\{2DC\_\{2\}\(1\+\\sqrt\{1\-4DC\_\{2\}K\}\)\}\(8\.80\)=2K1\+1−4DC2K\.\\displaystyle=\\frac\{2K\}\{1\+\\sqrt\{1\-4DC\_\{2\}K\}\}\.\(8\.81\)Since1−4DC2K≥0\\sqrt\{1\-4DC\_\{2\}K\}\\geq 0, the denominator is bounded below by11, which gives the upper boundS−≤2KS\_\{\-\}\\leq 2K\. We arrive at
supθ∈𝒮‖ℋ2,0‖2=𝒪\(1m\)\.\\displaystyle\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}=\\mathcal\{O\}\\left\(\\frac\{1\}\{\\sqrt\{m\}\}\\right\)\.\(8\.82\)
□\\square
Remark\.The application of the geometric Kato inequality follows since
ddt‖ℋ2,0‖22=ddt⟨ℋ2,0,ℋ2,0⟩=2⟨∇γ˙\(t\)ℋ2,0,ℋ2,0⟩\.\\displaystyle\\frac\{d\}\{dt\}\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}^\{2\}=\\frac\{d\}\{dt\}\\langle\\mathcal\{H\}^\{2,0\},\\mathcal\{H\}^\{2,0\}\\rangle=2\\langle\\nabla\_\{\\dot\{\\gamma\}\(t\)\}\\mathcal\{H\}^\{2,0\},\\mathcal\{H\}^\{2,0\}\\rangle\.\(8\.83\)It also follows from the chain rule
ddt‖ℋ2,0‖22=2‖ℋ2,0‖2ddt‖ℋ2,0‖2\.\\displaystyle\\frac\{d\}\{dt\}\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}^\{2\}=2\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}\\frac\{d\}\{dt\}\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}\.\(8\.84\)Equating the two,
‖ℋ2,0‖2ddt‖ℋ2,0‖2=⟨∇γ˙\(t\)ℋ2,0,ℋ2,0⟩\.\\displaystyle\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}\\frac\{d\}\{dt\}\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}=\\langle\\nabla\_\{\\dot\{\\gamma\}\(t\)\}\\mathcal\{H\}^\{2,0\},\\mathcal\{H\}^\{2,0\}\\rangle\.\(8\.85\)Applying the absolute value and the Cauchy\-Schwarz inequality,
‖ℋ2,0‖2\|ddt‖ℋ2,0‖2\|=\|⟨∇γ˙\(t\)ℋ2,0,ℋ2,0⟩\|≤‖∇γ˙\(t\)ℋ2,0‖2‖ℋ2,0‖2\.\\displaystyle\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}\\left\|\\frac\{d\}\{dt\}\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}\\right\|=\\left\|\\langle\\nabla\_\{\\dot\{\\gamma\}\(t\)\}\\mathcal\{H\}^\{2,0\},\\mathcal\{H\}^\{2,0\}\\rangle\\right\|\\leq\\\|\\nabla\_\{\\dot\{\\gamma\}\(t\)\}\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}\.\(8\.86\)Dividing by‖ℋ2,0‖2\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}recovers what we desire\.
Corollary\.Given the boundsupθ∈𝒮‖ℋ2,0‖2=𝒪\(1m\)\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\), the spectral norm of the\(0,2\)\(0,2\)Hessian over the geodesic ball is similarly bounded
supθ∈𝒮‖ℋ0,2‖2=𝒪\(1m\)\.\\displaystyle\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|\\mathcal\{H\}^\{0,2\}\\\|\_\{2\}=\\mathcal\{O\}\\left\(\\frac\{1\}\{\\sqrt\{m\}\}\\right\)\.\(8\.87\)
Proof\.Because theffoutput is real\-valued, i\.e\.f:M→ℝf:M\\rightarrow\\mathbb\{R\}, the function is equal to its complex conjugate,f=f¯f=\\overline\{f\}\. On a Kähler manifold, the Levi\-Civita connection is compatible with the complex structure, meaning the only non\-vanishing Christoffel symbols are of typeΓijk\{\\Gamma\_\{ij\}\}^\{k\}and their conjugatesΓi¯j¯k¯=Γijk¯\{\\Gamma\_\{\\overline\{i\}\\overline\{j\}\}\}^\{\\overline\{k\}\}=\\overline\{\{\\Gamma\_\{ij\}\}^\{k\}\}\. The\(2,0\)\(2,0\)covariant Hessian in local coordinates is given by
∇i∇jf=∂i∂jf−Γijk∂kf\.\\displaystyle\\nabla\_\{i\}\\nabla\_\{j\}f=\\partial\_\{i\}\\partial\_\{j\}f\-\{\\Gamma\_\{ij\}\}^\{k\}\\partial\_\{k\}f\.\(8\.88\)Taking the complex conjugate of this expression and applying the reality offf, we obtain
∇i∇jf¯=∂i¯∂j¯f¯−Γijk¯∂k¯f¯=∂i¯∂j¯f−Γi¯j¯k¯∂k¯f=∇i¯∇j¯f\.\\displaystyle\\overline\{\\nabla\_\{i\}\\nabla\_\{j\}f\}=\\partial\_\{\\overline\{i\}\}\\partial\_\{\\overline\{j\}\}\\overline\{f\}\-\\overline\{\{\\Gamma\_\{ij\}\}^\{k\}\}\\partial\_\{\\overline\{k\}\}\\overline\{f\}=\\partial\_\{\\overline\{i\}\}\\partial\_\{\\overline\{j\}\}f\-\{\\Gamma\_\{\\overline\{i\}\\overline\{j\}\}\}^\{\\overline\{k\}\}\\partial\_\{\\overline\{k\}\}f=\\nabla\_\{\\overline\{i\}\}\\nabla\_\{\\overline\{j\}\}f\.\(8\.89\)This demonstrates that the\(0,2\)\(0,2\)Hessian is exactly the complex conjugate of the\(2,0\)\(2,0\)Hessian,ℋ0,2=ℋ2,0¯\\mathcal\{H\}^\{0,2\}=\\overline\{\\mathcal\{H\}^\{2,0\}\}\. For any linear operator represented in coordinates, the spectral norm induced by the localized metrichhis invariant under complex conjugation\. The nonzero eigenvalues ofA†AA^\{\\dagger\}Aare real by the Spectral theorem since\(A†A\)†=A†\(A†\)†=A†A\(A^\{\\dagger\}A\)^\{\\dagger\}=A^\{\\dagger\}\(A^\{\\dagger\}\)^\{\\dagger\}=A^\{\\dagger\}A, and thusλmax\(A†A\)=λmax\(A¯†A¯\)\\lambda\_\{\\max\}\(A^\{\\dagger\}A\)=\\lambda\_\{\\max\}\(\\overline\{A\}^\{\\dagger\}\\overline\{A\}\)\. Therefore, the operator norms coincide
‖ℋ0,2‖2=‖ℋ2,0¯‖2=‖ℋ2,0‖2\.\\displaystyle\\\|\\mathcal\{H\}^\{0,2\}\\\|\_\{2\}=\\\|\\overline\{\\mathcal\{H\}^\{2,0\}\}\\\|\_\{2\}=\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}\.\(8\.90\)Applying the supremum over the geodesic ball𝒮\\mathcal\{S\}and substituting the result of[8\.82](https://arxiv.org/html/2608.19584#S8.E82), we conclude
supθ∈𝒮‖ℋ0,2‖2=supθ∈𝒮‖ℋ2,0‖2=𝒪\(1m\)\.\\displaystyle\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|\\mathcal\{H\}^\{0,2\}\\\|\_\{2\}=\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|\\mathcal\{H\}^\{2,0\}\\\|\_\{2\}=\\mathcal\{O\}\\left\(\\frac\{1\}\{\\sqrt\{m\}\}\\right\)\.\(8\.91\)
□\\square
Figure 8:We plot the corollary of Lemma 2, and that the singular values ofℋ2,0\\mathcal\{H\}^\{2,0\}andℋ0,2\\mathcal\{H\}^\{0,2\}match\. We plot all 2,750 singular values of Hessian matrices of the same dimension\. We choose network widthm=50m=50and input dimensiond=4d=4, where inputzzis sampled randomly\. Here, color indicates value\.
## 9Initialization results
Proof of Theorem 3\.Assume, by induction forward on the layers, that the previous layer’s squared activations satisfy\|αj\(l−1\)\|2=𝒪\(1\)\|\\alpha\_\{j\}^\{\(l\-1\)\}\|^\{2\}=\\mathcal\{O\}\(1\)which holds with probability at least1−δ11\-\\delta\_\{1\}\. Certainly the base caseα\(0\)=z\\alpha^\{\(0\)\}=zis𝒪\(1\)\\mathcal\{O\}\(1\)\. At initialization, the entries of the weight matrixW\(l\)W^\{\(l\)\}are drawn i\.i\.d\. from a standard complex Gaussian distribution,𝒞𝒩\(0,1\)\\mathcal\{CN\}\(0,1\)with mean zero and unit variance\. The variance of the pre\-activation is then
Var\(hi\(l\)\)\\displaystyle\\text\{Var\}\\left\(h\_\{i\}^\{\(l\)\}\\right\)=Var\(1m∑j=1mWij\(l\)αj\(l−1\)\)=1m∑j=1mVar\(Wij\(l\)\)\|αj\(l−1\)\|2=1m∑j=1m\|αj\(l−1\)\|2,\\displaystyle=\\text\{Var\}\\left\(\\frac\{1\}\{\\sqrt\{m\}\}\\sum\_\{j=1\}^\{m\}W\_\{ij\}^\{\(l\)\}\\alpha\_\{j\}^\{\(l\-1\)\}\\right\)=\\frac\{1\}\{m\}\\sum\_\{j=1\}^\{m\}\\text\{Var\}\\left\(W\_\{ij\}^\{\(l\)\}\\right\)\\left\|\\alpha\_\{j\}^\{\(l\-1\)\}\\right\|^\{2\}=\\frac\{1\}\{m\}\\sum\_\{j=1\}^\{m\}\\left\|\\alpha\_\{j\}^\{\(l\-1\)\}\\right\|^\{2\},\(9\.1\)where the last equality usesVar\(Wij\(l\)\)=1\\text\{Var\}\(W\_\{ij\}^\{\(l\)\}\)=1\. As an average ofmmterms each of𝒪\(1\)\\mathcal\{O\}\(1\), this variance is itself𝒪\(1\)\\mathcal\{O\}\(1\), so the magnitude ofhi\(l\)h\_\{i\}^\{\(l\)\}is𝒪\(1\)\\mathcal\{O\}\(1\)as well with probability at least1−δ21\-\\delta\_\{2\}\. Provided the activation functionϕ\\phias in[3](https://arxiv.org/html/2608.19584#S3)is sufficiently well\-behaved, the resulting activationαi\(l\)=ϕ\(hi\(l\),hi\(l\)¯\)\\alpha\_\{i\}^\{\(l\)\}=\\phi\\big\(h\_\{i\}^\{\(l\)\},\\overline\{h\_\{i\}^\{\(l\)\}\}\\big\)remains𝒪\(1\)\\mathcal\{O\}\(1\)\. Consequently, the squared norm of the layer’s activation vector scales linearly with width,
‖α\(l\)‖ℓ22=∑j=1m\|αj\(l\)\|2=𝒪\(m\)\.\\displaystyle\\left\\\|\\alpha^\{\(l\)\}\\right\\\|\_\{\\ell^\{2\}\}^\{2\}=\\sum\_\{j=1\}^\{m\}\\left\|\\alpha\_\{j\}^\{\(l\)\}\\right\|^\{2\}=\\mathcal\{O\}\(m\)\.\(9\.2\)This proves the forward induction, since this is equivalent to each component being𝒪\(1\)\\mathcal\{O\}\(1\)\.
We now turn to the backward pass\. We assume for the sake of induction backward on the layers\|δj\(l\+1\)\|2=𝒪\(1m\)\|\\delta\_\{j\}^\{\(l\+1\)\}\|^\{2\}=\\mathcal\{O\}\(\\frac\{1\}\{m\}\)with probability at least1−δ31\-\\delta\_\{3\}\. The base case follows immediately from the definition offfin[3](https://arxiv.org/html/2608.19584#S3)\. Define the complex error signal at layerllas the Wirtinger derivative of the network output with respect to the conjugate pre\-activation via[3](https://arxiv.org/html/2608.19584#S3),
δ\(l\):=∇h¯\(l\)f\(θ0\)\.\\displaystyle\\delta^\{\(l\)\}:=\\nabla\_\{\\overline\{h\}^\{\(l\)\}\}f\(\\theta\_\{0\}\)\.\(9\.3\)For the final layerLL, recalling thatf=1mv†α\(L\)f=\\frac\{1\}\{\\sqrt\{m\}\}v^\{\\dagger\}\\alpha^\{\(L\)\}, the chain rule gives
δi\(L\)=1m∂h¯ϕ\(hi\(L\),hi\(L\)¯\)vi¯\.\\displaystyle\\delta\_\{i\}^\{\(L\)\}=\\frac\{1\}\{\\sqrt\{m\}\}\\partial\_\{\\overline\{h\}\}\\phi\\left\(h\_\{i\}^\{\(L\)\},\\overline\{h\_\{i\}^\{\(L\)\}\}\\right\)\\overline\{v\_\{i\}\}\.\(9\.4\)Sincehi\(L\)h\_\{i\}^\{\(L\)\}is𝒪\(1\)\\mathcal\{O\}\(1\), we take the derivative∂h¯ϕ\\partial\_\{\\overline\{h\}\}\\phito be𝒪\(1\)\\mathcal\{O\}\(1\)as well\. As the weightsviv\_\{i\}are initialized in the same manner asWW, the initial error signal scales as
δi\(L\)=𝒪\(1m\)\\displaystyle\\delta\_\{i\}^\{\(L\)\}=\\mathcal\{O\}\\left\(\\frac\{1\}\{\\sqrt\{m\}\}\\right\)\(9\.5\)with probability at least1−δ41\-\\delta\_\{4\}\. Definingϵ\(l\):=∇h\(l\)f\(θ0\)\\epsilon^\{\(l\)\}:=\\nabla\_\{h^\{\(l\)\}\}f\(\\theta\_\{0\}\), recursively, the multivariable Wirtinger chain rule gives an update rule via pre\-activation definition[9\.3](https://arxiv.org/html/2608.19584#S9.E3)
δi\(l\)\\displaystyle\\delta\_\{i\}^\{\(l\)\}=∑j=1m\(∂f∂hj\(l\+1\)∂hj\(l\+1\)∂hi\(l\)¯\+∂f∂hj\(l\+1\)¯∂hj\(l\+1\)¯∂hi\(l\)¯\)\\displaystyle=\\sum\_\{j=1\}^\{m\}\\left\(\\frac\{\\partial f\}\{\\partial h\_\{j\}^\{\(l\+1\)\}\}\\frac\{\\partial h\_\{j\}^\{\(l\+1\)\}\}\{\\partial\\overline\{h\_\{i\}^\{\(l\)\}\}\}\+\\frac\{\\partial f\}\{\\partial\\overline\{h\_\{j\}^\{\(l\+1\)\}\}\}\\frac\{\\partial\\overline\{h\_\{j\}^\{\(l\+1\)\}\}\}\{\\partial\\overline\{h\_\{i\}^\{\(l\)\}\}\}\\right\)\(9\.6\)=1m∑j=1m\(ϵj\(l\+1\)Wji\(l\+1\)∂h¯ϕ\(hi\(l\),hi\(l\)¯\)\+δj\(l\+1\)Wji\(l\+1\)¯∂hϕ\(hi\(l\),hi\(l\)¯\)¯\)\.\\displaystyle=\\frac\{1\}\{\\sqrt\{m\}\}\\sum\_\{j=1\}^\{m\}\\left\(\\epsilon\_\{j\}^\{\(l\+1\)\}W\_\{ji\}^\{\(l\+1\)\}\\partial\_\{\\overline\{h\}\}\\phi\\big\(h\_\{i\}^\{\(l\)\},\\overline\{h\_\{i\}^\{\(l\)\}\}\\big\)\+\\delta\_\{j\}^\{\(l\+1\)\}\\overline\{W\_\{ji\}^\{\(l\+1\)\}\}\\overline\{\\partial\_\{h\}\\phi\\big\(h\_\{i\}^\{\(l\)\},\\overline\{h\_\{i\}^\{\(l\)\}\}\\big\)\}\\right\)\.\(9\.7\)Because the weightsWWandW¯\\overline\{W\}are independent of the backpropagated errors, the variance of the sum is the sum of the variances\. Using the inductive step that\|δj\(l\+1\)\|2\|\\delta\_\{j\}^\{\(l\+1\)\}\|^\{2\}and\|ϵj\(l\+1\)\|2\|\\epsilon\_\{j\}^\{\(l\+1\)\}\|^\{2\}are𝒪\(1m\)\\mathcal\{O\}\(\\frac\{1\}\{m\}\), the sum of the2m2msuch terms is𝒪\(1\)\\mathcal\{O\}\(1\)\. The1/m1/mscaling from the variance and moving outside1/m1/\\sqrt\{m\}scaling ensures thatVar\(δi\(l\)\)=𝒪\(1m\)\\text\{Var\}\(\\delta\_\{i\}^\{\(l\)\}\)=\\mathcal\{O\}\(\\frac\{1\}\{m\}\)\. Since the variance scales with the square, we can square root each component \(observe previously in the proof we examined the square of𝒪\(1\)\\mathcal\{O\}\(1\)terms, which is𝒪\(1\)\\mathcal\{O\}\(1\)\)\. Consequently, with probability at least1−δ51\-\\delta\_\{5\}
‖δ\(l\)‖∞=𝒪\(1m\)\.\\displaystyle\\\|\\delta^\{\(l\)\}\\\|\_\{\\infty\}=\\mathcal\{O\}\\left\(\\frac\{1\}\{\\sqrt\{m\}\}\\right\)\.\(9\.8\)Moreover, the square on a single element is therefore𝒪\(1m\)\\mathcal\{O\}\(\\frac\{1\}\{m\}\), which proves the backward induction\.
Now we compute the\(1,1\)\(1,1\)parameter Hessian and examine the diagonal blocks\. Notice for a parameter perturbationXX,
\{ℋW\(l\)W¯\(l\)\(X\)=1mD\(l\)Xα\(l−1\)\(α\(l−1\)\)†,D\(l\)=diag\(δ\(l\)∘∂h∂h¯ϕ\(h\(l\),h\(l\)¯\)\)\.\\displaystyle\\begin\{cases\}\\mathcal\{H\}\_\{W^\{\(l\)\}\\overline\{W\}^\{\(l\)\}\}\(X\)=\\frac\{1\}\{m\}D^\{\(l\)\}X\\alpha^\{\(l\-1\)\}\\big\(\\alpha^\{\(l\-1\)\}\\big\)^\{\\dagger\},\\\\ D^\{\(l\)\}=\\text\{diag\}\\left\(\\delta^\{\(l\)\}\\circ\\partial\_\{h\}\\partial\_\{\\overline\{h\}\}\\phi\\big\(h^\{\(l\)\},\\overline\{h^\{\(l\)\}\}\\big\)\\right\)\.\\end\{cases\}\(9\.9\)The spectral norm is governed by
\\bBigg@3‖ℋW\(l\)W¯\(l\)\\bBigg@3‖2≤\\bBigg@3‖D\(l\)\\bBigg@3‖2\\bBigg@3‖1mα\(l−1\)\(α\(l−1\)\)†\\bBigg@3‖2\.\\displaystyle\\bBigg@\{3\}\\\|\\mathcal\{H\}\_\{W^\{\(l\)\}\\overline\{W\}^\{\(l\)\}\}\\bBigg@\{3\}\\\|\_\{2\}\\leq\\bBigg@\{3\}\\\|D^\{\(l\)\}\\bBigg@\{3\}\\\|\_\{2\}\\bBigg@\{3\}\\\|\\frac\{1\}\{m\}\\alpha^\{\(l\-1\)\}\\big\(\\alpha^\{\(l\-1\)\}\\big\)^\{\\dagger\}\\bBigg@\{3\}\\\|\_\{2\}\.\(9\.10\)We evaluate both terms independently\. The operator norm of the diagonal matrixD\(l\)D^\{\(l\)\}is its maximum absolute entry\. Assuming the second Wirtinger derivative ofϕ\\phiis bounded, and given our backward induction result‖δ\(l\)‖∞=𝒪\(1m\)\\\|\\delta^\{\(l\)\}\\\|\_\{\\infty\}=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)from[9\.8](https://arxiv.org/html/2608.19584#S9.E8), we obtain
‖D\(l\)‖2=maxi\|Dii\(l\)\|=𝒪\(1m\),\\displaystyle\\left\\\|D^\{\(l\)\}\\right\\\|\_\{2\}=\\max\_\{i\}\\left\|D\_\{ii\}^\{\(l\)\}\\right\|=\\mathcal\{O\}\\left\(\\frac\{1\}\{\\sqrt\{m\}\}\\right\),\(9\.11\)with probability1−δ61\-\\delta\_\{6\}\. The second term of[9\.10](https://arxiv.org/html/2608.19584#S9.E10)is a rank\-1 matrix scaled by1/m1/m\. Its unique non\-zero eigenvalue is exactly1m‖α\(l−1\)‖22\\frac\{1\}\{m\}\\\|\\alpha^\{\(l\-1\)\}\\\|\_\{2\}^\{2\}\. Since we established by the forward induction that‖α\(l−1\)‖22=𝒪\(m\)\\\|\\alpha^\{\(l\-1\)\}\\\|\_\{2\}^\{2\}=\\mathcal\{O\}\(m\), this operator norm cancels the width dependence,
\\bBigg@3‖1mα\(l−1\)\(α\(l−1\)\)†\\bBigg@3‖2=𝒪\(1\)\.\\displaystyle\\bBigg@\{3\}\\\|\\frac\{1\}\{m\}\\alpha^\{\(l\-1\)\}\\big\(\\alpha^\{\(l\-1\)\}\\big\)^\{\\dagger\}\\bBigg@\{3\}\\\|\_\{2\}=\\mathcal\{O\}\(1\)\.\(9\.12\)Therefore, the spectral norm of the diagonal Hessian block is bounded by𝒪\(1m\)\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)\.
Evaluating the remaining components, the output weight blockℋvv¯=0\\mathcal\{H\}\_\{v\\overline\{v\}\}=0sinceffis linear inv†v^\{\\dagger\}\. Now we examine the off\-diagonal blocks\. Because the Hessian block is a bilinear form, its spectral norm can be analyzed via a unit\-norm perturbationXXat layerll\. Let the gradient matrix at layerkkbe
E\(k\):=∇W¯\(k\)f\(θ0\)=1mδ\(k\)\(α\(k−1\)\)†\.\\displaystyle E^\{\(k\)\}:=\\nabla\_\{\\overline\{W\}^\{\(k\)\}\}f\(\\theta\_\{0\}\)=\\frac\{1\}\{\\sqrt\{m\}\}\\delta^\{\(k\)\}\\big\(\\alpha^\{\(k\-1\)\}\\big\)^\{\\dagger\}\.\(9\.13\)We desire to bound‖ℋW\(l\)W¯\(k\)‖2=sup‖X‖2=1‖ΔXE\(k\)‖2\\left\\\|\\mathcal\{H\}\_\{W^\{\(l\)\}\\overline\{W\}^\{\(k\)\}\}\\right\\\|\_\{2\}=\\sup\_\{\\\|X\\\|\_\{2\}=1\}\\left\\\|\\Delta\_\{X\}E^\{\(k\)\}\\right\\\|\_\{2\}, whereΔX\\Delta\_\{X\}is the directional derivative∂W\(l\)\[⋅\]\(X\)\\partial\_\{W^\{\(l\)\}\}\[\\cdot\]\(X\)\. It is nontrivial to bound this directly due to a lack of symmetry between backward and forward layers, therefore we will analyze the off\-diagonal scenario on a case\-by\-case basis\.
Case 1\.Assumel\>kl\>k\. Thenα\(k−1\)\\alpha^\{\(k\-1\)\}has no dependence onW\(l\)W^\{\(l\)\}\. Therefore, its variation with respect toXXis identically zero\. Hence,
ΔXE\(k\)=1m\(ΔXδ\(k\)\)\(α\(k−1\)\)†\.\\displaystyle\\Delta\_\{X\}E^\{\(k\)\}=\\frac\{1\}\{\\sqrt\{m\}\}\\big\(\\Delta\_\{X\}\\delta^\{\(k\)\}\\big\)\\big\(\\alpha^\{\(k\-1\)\}\\big\)^\{\\dagger\}\.\(9\.14\)We proved earlier‖α\(l\)‖22\\\|\\alpha^\{\(l\)\}\\\|\_\{2\}^\{2\}is𝒪\(m\)\\mathcal\{O\}\(m\)in[9\.2](https://arxiv.org/html/2608.19584#S9.E2), or equivalently‖α\(k−1\)‖2\\\|\\alpha^\{\(k\-1\)\}\\\|\_\{2\}is𝒪\(m\)\\mathcal\{O\}\(\\sqrt\{m\}\)\. Moreover, we proved the rule on\|δj\(k−1\)\|2=𝒪\(1m\)\|\\delta\_\{j\}^\{\(k\-1\)\}\|^\{2\}=\\mathcal\{O\}\(\\frac\{1\}\{m\}\)\. It follows from the work we did earlier that‖δ\(l\)‖2=𝒪\(1\)\\\|\\delta^\{\(l\)\}\\\|\_\{2\}=\\mathcal\{O\}\(1\)\. We use Cauchy\-Schwarz, but the unit\-norm perturbation does not preserve order \(note[51](https://arxiv.org/html/2608.19584#bib.bib29)discusses unit directional derivatives and notes analyzing unit\-norm directions can miss uniform behavior\), so‖ΔXδ\(k\)‖2=𝒪\(1m\)\\\|\\Delta\_\{X\}\\delta^\{\(k\)\}\\\|\_\{2\}=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\), or more rigorously which follows since‖ΔXδ\(l−1\)‖2≤1m‖diag\(ϕ′\)‖2‖X‖2‖δ\(l\)‖2\\\|\\Delta\_\{X\}\\delta^\{\(l\-1\)\}\\\|\_\{2\}\\leq\\frac\{1\}\{\\sqrt\{m\}\}\\\|\\text\{diag\}\(\\phi^\{\\prime\}\)\\\|\_\{2\}\\\|X\\\|\_\{2\}\\\|\\delta^\{\(l\)\}\\\|\_\{2\}, the three of which are𝒪\(1\)\\mathcal\{O\}\(1\)\. Therefore, it immediately follows‖ΔXE\(k\)‖2=𝒪\(1m\)\\\|\\Delta\_\{X\}E^\{\(k\)\}\\\|\_\{2\}=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)by using[9\.2](https://arxiv.org/html/2608.19584#S9.E2)\.
Case 2\.Assumel<kl<k\. When the perturbed layerllprecedes layerkk, both the error signal and the forward activations depend onW\(l\)W^\{\(l\)\}\. Applying the product rule,
ΔXE\(k\)=1m\[\(ΔXδ\(k\)\)\(α\(k−1\)\)†\+δ\(k\)ΔX\(α\(k−1\)†\)\]\.\\displaystyle\\Delta\_\{X\}E^\{\(k\)\}=\\frac\{1\}\{\\sqrt\{m\}\}\\left\[\\big\(\\Delta\_\{X\}\\delta^\{\(k\)\}\\big\)\\big\(\\alpha^\{\(k\-1\)\}\\big\)^\{\\dagger\}\+\\delta^\{\(k\)\}\\Delta\_\{X\}\\big\(\\alpha^\{\(k\-1\)\\dagger\}\\big\)\\right\]\.\(9\.15\)The first term is exactly Case 1 of[9\.14](https://arxiv.org/html/2608.19584#S9.E14)\. Examining the second term, we first note the perturbation at layerllisΔXh\(l\)=1mXα\(l−1\)\\Delta\_\{X\}h^\{\(l\)\}=\\frac\{1\}\{\\sqrt\{m\}\}X\\alpha^\{\(l\-1\)\}, giving
\\bBigg@3‖ΔXh\(l\)\\bBigg@3‖2≤1m\\bBigg@3‖X\\bBigg@3‖2\\bBigg@3‖α\(l−1\)\\bBigg@3‖2=1m×𝒪\(1\)×𝒪\(m\)=𝒪\(1\)\.\\displaystyle\\bBigg@\{3\}\\\|\\Delta\_\{X\}h^\{\(l\)\}\\bBigg@\{3\}\\\|\_\{2\}\\leq\\frac\{1\}\{\\sqrt\{m\}\}\\bBigg@\{3\}\\\|X\\bBigg@\{3\}\\\|\_\{2\}\\bBigg@\{3\}\\\|\\alpha^\{\(l\-1\)\}\\bBigg@\{3\}\\\|\_\{2\}=\\frac\{1\}\{\\sqrt\{m\}\}\\times\\mathcal\{O\}\(1\)\\times\\mathcal\{O\}\\big\(\\sqrt\{m\}\\big\)=\\mathcal\{O\}\(1\)\.\(9\.16\)The order of‖α\(l−1\)‖2\\\|\\alpha^\{\(l\-1\)\}\\\|\_\{2\}follows from earlier in the proof by taking the square root of[9\.2](https://arxiv.org/html/2608.19584#S9.E2)\. To evaluate the variation at layerk−1k\-1, we use the Jacobian of the forward pass, and we can noteΔXα\(k−1\)=∂α\(k−1\)∂h\(l\)ΔXh\(l\)\\Delta\_\{X\}\\alpha^\{\(k\-1\)\}=\\frac\{\\partial\\alpha^\{\(k\-1\)\}\}\{\\partial h^\{\(l\)\}\}\\Delta\_\{X\}h^\{\(l\)\}\. Because the intermediate Jacobians are assumed to have bounded operator norms at initialization with probability1−δ71\-\\delta\_\{7\}, the perturbation at layerk−1k\-1shares this bound, meaning‖ΔXα\(k−1\)‖2=𝒪\(1\)\\\|\\Delta\_\{X\}\\alpha^\{\(k\-1\)\}\\\|\_\{2\}=\\mathcal\{O\}\(1\)\. Since‖δ\(k\)‖2=𝒪\(1\)\\\|\\delta^\{\(k\)\}\\\|\_\{2\}=\\mathcal\{O\}\(1\), we can put the results together and, using the scaling out front, we see‖ΔXE\(k\)‖2=𝒪\(1m\)\\\|\\Delta\_\{X\}E^\{\(k\)\}\\\|\_\{2\}=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)in this case too\. We put everything together\. The diagonal and off\-diagonal scenarios, including both cases, are all𝒪\(1m\)\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\), and this completes the proof\.
□\\square
Corollary\.The proof for the \(2,0\)\-Hessian is almost identical, with a few changes\. The diagonal matrix ofD\(l\)D^\{\(l\)\}must now be
D\(2,0\)\(l\)=diag\(δ\(l\)∘∂h2ϕ\(h\(l\),h\(l\)¯\)\)\\displaystyle D^\{\(l\)\}\_\{\(2,0\)\}=\\text\{diag\}\\left\(\\delta^\{\(l\)\}\\circ\\partial\_\{h\}^\{2\}\\phi\\big\(h^\{\(l\)\},\\overline\{h^\{\(l\)\}\}\\big\)\\right\)\(9\.17\)The outer product must be changed to
ℋW\(l\)W\(l\)\(X\)=1mD\(2,0\)\(l\)Xα\(l−1\)\(α\(l−1\)\)T\.\\displaystyle\\mathcal\{H\}\_\{W^\{\(l\)\}W^\{\(l\)\}\}\(X\)=\\frac\{1\}\{m\}D^\{\(l\)\}\_\{\(2,0\)\}X\\alpha^\{\(l\-1\)\}\\big\(\\alpha^\{\(l\-1\)\}\\big\)^\{T\}\.\(9\.18\)We can note‖α\(l−1\)\(α\(l−1\)\)T‖2\\left\\\|\\alpha^\{\(l\-1\)\}\\big\(\\alpha^\{\(l\-1\)\}\\big\)^\{T\}\\right\\\|\_\{2\}is‖α\(l−1\)‖22\\big\\\|\\alpha^\{\(l\-1\)\}\\big\\\|\_\{2\}^\{2\}\. By forward induction,‖α\(l−1\)‖22=𝒪\(m\)\\big\\\|\\alpha^\{\(l\-1\)\}\\big\\\|\_\{2\}^\{2\}=\\mathcal\{O\}\(m\), so the1/m1/mscaling cancels out, leaving the block bounded by𝒪\(1/m\)\\mathcal\{O\}\(1/\\sqrt\{m\}\)\. Moreover, we can redefine
E\(k\):=∇W\(k\)f\(θ0\),\\displaystyle E^\{\(k\)\}:=\\nabla\_\{W^\{\(k\)\}\}f\(\\theta\_\{0\}\),\(9\.19\)and the application of the product rule in case 2, and it follows we have the unit\-norm directional derivative bounds‖ΔXh\(l\)‖2=𝒪\(1\)\\big\\\|\\Delta\_\{X\}h^\{\(l\)\}\\big\\\|\_\{2\}=\\mathcal\{O\}\(1\)and‖ΔXα\(k−1\)‖2=𝒪\(1\)\\big\\\|\\Delta\_\{X\}\\alpha^\{\(k\-1\)\}\\big\\\|\_\{2\}=\\mathcal\{O\}\(1\)\.
□\\square
## 10Convexity results
Proof of Theorem 4\.DenoteL\(θ\)=1n∑i=1nℓi\(yi,f\(θ,zi\)\)L\(\\theta\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\ell\_\{i\}\(y\_\{i\},f\(\\theta;z\_\{i\}\)\)\. Consider the second\-order Riemannian Taylor expansion along the geodesicγ\(s\)\\gamma\(s\)connectingθt\\theta\_\{t\}toθ0\\theta\_\{0\}
L\(θ0\)=L\(θt\)\+2Re⟨∇ωL\(θt\),expθt−1\(θ0\)⟩ω\+12∇ω2L\(θ~t\)\(v,v\),\\displaystyle L\(\\theta\_\{0\}\)=L\(\\theta\_\{t\}\)\+2\\text\{Re\}\\langle\\nabla\_\{\\omega\}L\(\\theta\_\{t\}\),\\exp\_\{\\theta\_\{t\}\}^\{\-1\}\(\\theta\_\{0\}\)\\rangle\_\{\\omega\}\+\\frac\{1\}\{2\}\\nabla^\{2\}\_\{\\omega\}L\(\\widetilde\{\\theta\}\_\{t\}\)\(v,v\),\(10\.1\)whereθ~t\\widetilde\{\\theta\}\_\{t\}is an intermediary point on the geodesic,v=expθt−1\(θ0\)v=\\exp\_\{\\theta\_\{t\}\}^\{\-1\}\(\\theta\_\{0\}\), and∇ω2\\nabla^\{2\}\_\{\\omega\}denotes the covariant Riemannian Hessian with respect to the Kähler metricω\\omega\.
Let us bound the Hessian terms\. On a Kähler manifold, the covariant Hessian decomposes\. The mixed Christoffel symbols vanish\. We have the decomposition
12∇ω2L\(θ~t\)\(v,v\)=Re\(vT\(∇ω2,0L\(θ~t\)\)v\)⏟=A1\+v†\(∇ω1,1L\(θ~t\)\)v⏟=A2\.\\displaystyle\\frac\{1\}\{2\}\\nabla^\{2\}\_\{\\omega\}L\(\\widetilde\{\\theta\}\_\{t\}\)\(v,v\)=\\underbrace\{\\text\{Re\}\\left\(v^\{T\}\\left\(\\nabla^\{2,0\}\_\{\\omega\}L\(\\widetilde\{\\theta\}\_\{t\}\)\\right\)v\\right\)\}\_\{=A\_\{1\}\}\+\\underbrace\{v^\{\\dagger\}\\left\(\\nabla^\{1,1\}\_\{\\omega\}L\(\\widetilde\{\\theta\}\_\{t\}\)\\right\)v\}\_\{=A\_\{2\}\}\.\(10\.2\)Applying the covariant chain rule to the composite loss, we observe
A1=Re\{1n∑i=1n\[ℓi′′\(fi\(θ~t\)\)\(∇ωfi\(θ~t\)v\)2⏟=B1\+ℓi′\(fi\(θ~t\)\)\(vT∇ω2,0fi\(θ~t\)v\)⏟=B2\]\},\\displaystyle A\_\{1\}=\\text\{Re\}\\left\\\{\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\left\[\\underbrace\{\\ell^\{\\prime\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\\left\(\\nabla\_\{\\omega\}f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)v\\right\)^\{2\}\}\_\{=B\_\{1\}\}\+\\underbrace\{\\ell^\{\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\\left\(v^\{T\}\\nabla^\{2,0\}\_\{\\omega\}f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)v\\right\)\}\_\{=B\_\{2\}\}\\right\]\\right\\\},\(10\.3\)and
A2=1n∑i=1n\[ℓi′′\(fi\(θ~t\)\)\|∇ωfi\(θ~t\)v\|2\+ℓi′\(fi\(θ~t\)\)\(v†∇ω1,1fi\(θ~t\)v\)\]\.\\displaystyle A\_\{2\}=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\left\[\\ell^\{\\prime\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\\left\|\\nabla\_\{\\omega\}f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)v\\right\|^\{2\}\+\\ell^\{\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\\left\(v^\{\\dagger\}\\nabla^\{1,1\}\_\{\\omega\}f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)v\\right\)\\right\]\.\(10\.4\)Let us examine theB1B\_\{1\}of[10\.3](https://arxiv.org/html/2608.19584#S10.E3)term and the Jacobian term ofA2A\_\{2\}
Re\(B1\)\+A2,Jac=2n∑i=1nℓi′′\(fi\(θ~t\)\)\(Re\(∇ωf\(θ~t,zi\)v\)\)2\.\\displaystyle\\text\{Re\}\(B\_\{1\}\)\+A\_\{2,\\text\{Jac\}\}=\\frac\{2\}\{n\}\\sum\_\{i=1\}^\{n\}\\ell^\{\\prime\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\\left\(\\text\{Re\}\\left\(\\nabla\_\{\\omega\}f\(\\widetilde\{\\theta\}\_\{t\};z\_\{i\}\)v\\right\)\\right\)^\{2\}\.\(10\.5\)Using a lower boundℓi′′≥a\\ell\_\{i\}^\{\\prime\\prime\}\\geq aon the loss, we perform the Taylor expansion along the geodesicγ\(s\)\\gamma\(s\)for the intermediaryθ~\\widetilde\{\\theta\}
Re\(B1\)\+A2,Jac\\displaystyle\\text\{Re\}\(B\_\{1\}\)\+A\_\{2,\\text\{Jac\}\}≥2an∑i=1n\(Re\(∇ωf\(θt,zi\)v\)\+Re\(∫01∇ω2fi\(γ\(s\)\)\(v,v\)𝑑s\)\)2\\displaystyle\\geq\\frac\{2a\}\{n\}\\sum\_\{i=1\}^\{n\}\\left\(\\text\{Re\}\\left\(\\nabla\_\{\\omega\}f\(\\theta\_\{t\};z\_\{i\}\)v\\right\)\+\\text\{Re\}\\left\(\\int\_\{0\}^\{1\}\\nabla^\{2\}\_\{\\omega\}f\_\{i\}\(\\gamma\(s\)\)\(v,v\)ds\\right\)\\right\)^\{2\}\(10\.6\)≥2an∑i=1n\(Re\(∇ωfi\(θt\)v\)\)2−4an∑i=1n\|Re\(∇ωfi\(θt\)v\)\|\\bBigg@3\|Re\(∫01∇ω2fi\(γ\(s\)\)\(v,v\)𝑑s\)\\bBigg@3\|\.\\displaystyle\\geq\\frac\{2a\}\{n\}\\sum\_\{i=1\}^\{n\}\\left\(\\text\{Re\}\\left\(\\nabla\_\{\\omega\}f\_\{i\}\(\\theta\_\{t\}\)v\\right\)\\right\)^\{2\}\-\\frac\{4a\}\{n\}\\sum\_\{i=1\}^\{n\}\\left\|\\text\{Re\}\\left\(\\nabla\_\{\\omega\}f\_\{i\}\(\\theta\_\{t\}\)v\\right\)\\right\|\\bBigg@\{3\}\|\\text\{Re\}\\left\(\\int\_\{0\}^\{1\}\\nabla^\{2\}\_\{\\omega\}f\_\{i\}\(\\gamma\(s\)\)\(v,v\)ds\\right\)\\bBigg@\{3\}\|\.\(10\.7\)We have used the fact that\(X\+Y\)2≥X2−2\|X∥Y\|\(X\+Y\)^\{2\}\\geq X^\{2\}\-2\|X\\\|Y\|and that the covariant Hessian is a bilinear form\. Letℱt\(v\)=2n∑i=1n\(Re\(∇ωfi\(θt\)v\)\)2\\mathcal\{F\}\_\{t\}\(v\)=\\frac\{2\}\{n\}\\sum\_\{i=1\}^\{n\}\\left\(\\text\{Re\}\\left\(\\nabla\_\{\\omega\}f\_\{i\}\(\\theta\_\{t\}\)v\\right\)\\right\)^\{2\}be the quadratic form of the empirical Fisher information\. Applying the Cauchy\-Schwarz inequality to[10\.6](https://arxiv.org/html/2608.19584#S10.E6)
Re\(B1\)\+A2,Jac≥aℱt\(v\)−2aℱt\(v\)1n∑i=1n\(∫01∇ω2fi\(γ\(s\)\)\(v,v\)𝑑s\)2\.\\displaystyle\\text\{Re\}\(B\_\{1\}\)\+A\_\{2,\\text\{Jac\}\}\\geq a\\mathcal\{F\}\_\{t\}\(v\)\-2a\\sqrt\{\\mathcal\{F\}\_\{t\}\(v\)\}\\sqrt\{\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\left\(\\int\_\{0\}^\{1\}\\nabla^\{2\}\_\{\\omega\}f\_\{i\}\(\\gamma\(s\)\)\(v,v\)ds\\right\)^\{2\}\}\.\(10\.8\)Assume that the loss landscape satisfies strong convexity in the local neighborhood𝒮\\mathcal\{S\}with parameterμ\>0\\mu\>0, such that the Fisher quadratic form is bounded below byℱt\(v\)≥μ‖v‖ω2\\mathcal\{F\}\_\{t\}\(v\)\\geq\\mu\\\|v\\\|\_\{\\omega\}^\{2\}\. Let the constantCℋC\_\{\\mathcal\{H\}\}be such that the full covariant Hessian norm is bounded byCℋ‖v‖ω2/mC\_\{\\mathcal\{H\}\}\\\|v\\\|\_\{\\omega\}^\{2\}/\\sqrt\{m\}, since the Hessian bound picks up a factor ofm−12m^\{\-\\frac\{1\}\{2\}\}\. Further, lettingρ\\rhobound the maximum eigenvalue of the Fisher matrix such thatℱt\(v\)≤ρ‖v‖ω2\\mathcal\{F\}\_\{t\}\(v\)\\leq\\rho\\\|v\\\|\_\{\\omega\}^\{2\}, we arrive at the bound
Re\(B1\)\+A2,Jac≥\(aμ−2aρCℋ‖v‖ωm\)‖v‖ω2\.\\displaystyle\\text\{Re\}\(B\_\{1\}\)\+A\_\{2,\\text\{Jac\}\}\\geq\\left\(a\\mu\-\\frac\{2a\\sqrt\{\\rho\}C\_\{\\mathcal\{H\}\}\\\|v\\\|\_\{\\omega\}\}\{\\sqrt\{m\}\}\\right\)\\\|v\\\|\_\{\\omega\}^\{2\}\.\(10\.9\)Now, let us consider theB2B\_\{2\}term with the Hessian term ofA2A\_\{2\}
Re\(B2\)\+A2,Hes\\displaystyle\\text\{Re\}\(B\_\{2\}\)\+A\_\{2,\\text\{Hes\}\}=1n∑i=1nℓi′\(fi\(θ~t\)\)\(Re\[vT∇ω2,0fi\(θ~t\)v\]\+v†∇ω1,1fi\(θ~t\)v\)\\displaystyle=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\ell^\{\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\\left\(\\text\{Re\}\\left\[v^\{T\}\\nabla^\{2,0\}\_\{\\omega\}f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)v\\right\]\+v^\{\\dagger\}\\nabla^\{1,1\}\_\{\\omega\}f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)v\\right\)\(10\.10\)≥Cauchy Schwarz\\displaystyle\\stackrel\{\{\\scriptstyle\\text\{Cauchy Schwarz\}\}\}\{\{\\geq\}\}−1n∑i=1n\(ℓi′\(fi\(θ~t\)\)\)21n∑i=1n\|Re\[vT∇ω2,0fi\(θ~t\)v\]\+v†∇ω1,1fi\(θ~t\)v\|2\.\\displaystyle\-\\sqrt\{\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\(\\ell^\{\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\)^\{2\}\}\\sqrt\{\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\left\|\\text\{Re\}\\left\[v^\{T\}\\nabla^\{2,0\}\_\{\\omega\}f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)v\\right\]\+v^\{\\dagger\}\\nabla^\{1,1\}\_\{\\omega\}f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)v\\right\|^\{2\}\}\.\(10\.11\)Since we proved in Appendix[8\.2](https://arxiv.org/html/2608.19584#S8.SS2)and Appendix[8\.5](https://arxiv.org/html/2608.19584#S8.SS5)the Hessians scale𝒪\(1m\)\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\),
Re\(B2\)\+A2,Hes≥−1n∑i=1n\(ℓi′\(fi\(θ~t\)\)\)2Cℋm‖v‖ω2\.\\displaystyle\\text\{Re\}\(B\_\{2\}\)\+A\_\{2,\\text\{Hes\}\}\\geq\-\\sqrt\{\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\(\\ell^\{\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\)^\{2\}\}\\frac\{C\_\{\\mathcal\{H\}\}\}\{\\sqrt\{m\}\}\\\|v\\\|\_\{\\omega\}^\{2\}\.\(10\.12\)Combining the bounds
∇ω2L\(θ~t\)\(v,v\)≥\(aμ−2aρCℋ‖v‖ω\+Cℋ1n∑i=1n\(ℓi′\(fi\(θ~t\)\)\)2m\)⏟:=Γ\(a,μ,ρ,Cℋ,\{ℓi′\(fi\(θ~t\)\)\}i,m,v\)∥v∥ω2\.\\displaystyle\\nabla^\{2\}\_\{\\omega\}L\(\\widetilde\{\\theta\}\_\{t\}\)\(v,v\)\\geq\\underbrace\{\\left\(a\\mu\-\\frac\{2a\\sqrt\{\\rho\}C\_\{\\mathcal\{H\}\}\\\|v\\\|\_\{\\omega\}\+C\_\{\\mathcal\{H\}\}\\sqrt\{\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\(\\ell^\{\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\)^\{2\}\}\}\{\\sqrt\{m\}\}\\right\)\}\_\{\\displaystyle:=\\Gamma\\left\(a,\\mu,\\rho,C\_\{\\mathcal\{H\}\},\\\{\\ell\_\{i\}^\{\\prime\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\\\}\_\{i\},m,v\\right\)\}\\\|v\\\|\_\{\\omega\}^\{2\}\.\(10\.13\)Note thatθ0\\theta\_\{0\}lies within a local geodesic ball of radiusDDcentered atθt\\theta\_\{t\}, meaningθ0∈BωD\(θt\)\\theta\_\{0\}\\in B\_\{\\omega\}^\{D\}\(\\theta\_\{t\}\)\. Therefore, the geodesic distance is bounded‖v‖ω=‖expθt−1\(θ0\)‖ω≤D\\\|v\\\|\_\{\\omega\}=\\\|\\exp\_\{\\theta\_\{t\}\}^\{\-1\}\(\\theta\_\{0\}\)\\\|\_\{\\omega\}\\leq D\.
□\\square
Under a metric collapse, since the collapse impliesλmin=0\\lambda\_\{\\text\{min\}\}=0, we get
∇ω2L\(θ~t\)\(v,v\)≥lim infκ→0lim infρ→∞Cℋ→∞Γ\(a,μ\|μ=0,ρ,Cℋ,\{ℓi′\(fi\(θ~t\)\)\}i,m,v\)‖v‖ω2=−∞,\\displaystyle\\nabla^\{2\}\_\{\\omega\}L\(\\widetilde\{\\theta\}\_\{t\}\)\(v,v\)\\geq\\liminf\_\{\\kappa\\to 0\}\\liminf\_\{\\begin\{subarray\}\{c\}\\rho\\to\\infty\\\\ C\_\{\\mathcal\{H\}\}\\to\\infty\\end\{subarray\}\}\\Gamma\\left\(a,\\mu\|\_\{\\mu=0\},\\rho,C\_\{\\mathcal\{H\}\},\\\{\\ell\_\{i\}^\{\\prime\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\\\}\_\{i\},m,v\\right\)\\\|v\\\|\_\{\\omega\}^\{2\}=\-\\infty,\(10\.14\)which destroys the possibility of a finite lower bound\. The divergence of constants is because the covariant Hessian relies on the Levi\-Civita connection \(under choice of connection\) and the inverse metric, and a singular metric implies its inverse diverges \(under real analytic conventions so that it exists, for example take the limit if appropriate\)\. With a regularized metric, this is not as feasible\. In the scenario of a Calabi\-Yau manifold, and whenm→∞m\\rightarrow\\infty, we get
∇ω2L\(θ~t\)\(v,v\)≥lim infm→∞lim infμ→0lim infρ→r<∞Cℋ→c<∞Γ\(a,μ,ρ,Cℋ,\{ℓi′\(fi\(θ~t\)\)\}i,m,v\)‖v‖ω2\>−∞\.\\displaystyle\\nabla^\{2\}\_\{\\omega\}L\(\\widetilde\{\\theta\}\_\{t\}\)\(v,v\)\\geq\\liminf\_\{m\\rightarrow\\infty\}\\liminf\_\{\\mu\\rightarrow 0\}\\liminf\_\{\\begin\{subarray\}\{c\}\\rho\\rightarrow r<\\infty\\\\ C\_\{\\mathcal\{H\}\}\\to c<\\infty\\end\{subarray\}\}\\Gamma\\left\(a,\\mu,\\rho,C\_\{\\mathcal\{H\}\},\\\{\\ell\_\{i\}^\{\\prime\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\\\}\_\{i\},m,v\\right\)\\\|v\\\|\_\{\\omega\}^\{2\}\>\-\\infty\.\(10\.15\)We are slightly informal in our use of thelim inf\\liminf, since we are not particularly examining if the limit exists\.
### 10\.1β\\beta\-smoothness
Proof of Theorem 5\.By the second\-order Riemannian Taylor expansion aboutθt\\theta\_\{t\}along the geodesicγ\(s\)\\gamma\(s\), we have
L\(θ0\)=L\(θt\)\+2Re⟨∇ω1,0L\(θt\),v⟩ω\+12∇ω2L\(θ~t\)\(v,v\),\\displaystyle L\(\\theta\_\{0\}\)=L\(\\theta\_\{t\}\)\+2\\text\{Re\}\\langle\\nabla\_\{\\omega\}^\{1,0\}L\(\\theta\_\{t\}\),v\\rangle\_\{\\omega\}\+\\frac\{1\}\{2\}\\nabla^\{2\}\_\{\\omega\}L\(\\widetilde\{\\theta\}\_\{t\}\)\(v,v\),\(10\.16\)wherev=expθt−1\(θ0\)1,0v=\\exp\_\{\\theta\_\{t\}\}^\{\-1\}\(\\theta\_\{0\}\)^\{1,0\}\. Let us bound the Hessian term\. On a Kähler manifold, the covariant Hessian decomposes\. The mixed Christoffel symbols vanish\. We have the decomposition
12∇ω2L\(θ~t\)\(v,v\)=Re\(vT\(∇ω2,0L\(θ~t\)\)v\)⏟=A1\+v†\(∇ω1,1L\(θ~t\)\)v⏟=A2\.\\displaystyle\\frac\{1\}\{2\}\\nabla^\{2\}\_\{\\omega\}L\(\\widetilde\{\\theta\}\_\{t\}\)\(v,v\)=\\underbrace\{\\text\{Re\}\\left\(v^\{T\}\\left\(\\nabla^\{2,0\}\_\{\\omega\}L\(\\widetilde\{\\theta\}\_\{t\}\)\\right\)v\\right\)\}\_\{=A\_\{1\}\}\+\\underbrace\{v^\{\\dagger\}\\left\(\\nabla^\{1,1\}\_\{\\omega\}L\(\\widetilde\{\\theta\}\_\{t\}\)\\right\)v\}\_\{=A\_\{2\}\}\.\(10\.17\)As before, decompose
A1=Re\{1n∑i=1n\[ℓi′′\(fi\(θ~t\)\)\(∇ωfi\(θ~t\)v\)2⏟=B1\+ℓi′\(fi\(θ~t\)\)\(vT∇ω2,0fi\(θ~t\)v\)⏟=B2\]\},\\displaystyle A\_\{1\}=\\text\{Re\}\\left\\\{\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\left\[\\underbrace\{\\ell^\{\\prime\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\\left\(\\nabla\_\{\\omega\}f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)v\\right\)^\{2\}\}\_\{=B\_\{1\}\}\+\\underbrace\{\\ell^\{\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\\left\(v^\{T\}\\nabla^\{2,0\}\_\{\\omega\}f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)v\\right\)\}\_\{=B\_\{2\}\}\\right\]\\right\\\},\(10\.18\)and
A2=1n∑i=1n\[ℓi′′\(fi\(θ~t\)\)\|∇ωfi\(θ~t\)v\|2\+ℓi′\(fi\(θ~t\)\)\(v†∇ω1,1fi\(θ~t\)v\)\]\.\\displaystyle A\_\{2\}=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\left\[\\ell^\{\\prime\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\\left\|\\nabla\_\{\\omega\}f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)v\\right\|^\{2\}\+\\ell^\{\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\\left\(v^\{\\dagger\}\\nabla^\{1,1\}\_\{\\omega\}f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)v\\right\)\\right\]\.\(10\.19\)Let us regroup these terms into first\-order and second\-order terms\. Let us examine theB1B\_\{1\}term and the Jacobian term ofA2A\_\{2\}
Re\(B1\)\+A2,Jac=2n∑i=1nℓi′′\(fi\(θ~t\)\)\(Re\(∇ωfi\(θ~t\)v\)\)2\.\\displaystyle\\text\{Re\}\(B\_\{1\}\)\+A\_\{2,\\text\{Jac\}\}=\\frac\{2\}\{n\}\\sum\_\{i=1\}^\{n\}\\ell^\{\\prime\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\\left\(\\text\{Re\}\\left\(\\nabla\_\{\\omega\}f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)v\\right\)\\right\)^\{2\}\.\(10\.20\)For the square loss,ℓi′′=2\\ell^\{\\prime\\prime\}\_\{i\}=2\. The gradient term is reminiscent of a quadratic form via the empirical Fisher information\. Assuming a sufficiently nice bound on the gradient𝒮\\mathcal\{S\}, and since it is quadratic invv, we assume there existsρJ\\rho\_\{J\}so that
Re\(B1\)\+A2,Jac≤ρJ‖v‖ω2\.\\displaystyle\\text\{Re\}\(B\_\{1\}\)\+A\_\{2,\\text\{Jac\}\}\\leq\\rho\_\{J\}\\\|v\\\|\_\{\\omega\}^\{2\}\.\(10\.21\)Now, let us consider theB2B\_\{2\}term with the Hessian term ofA2A\_\{2\}:
Re\(B2\)\+A2,Hes=1n∑i=1nℓi′\(fi\(θ~t\)\)\(Re\[vT∇ω2,0fi\(θ~t\)v\]\+v†∇ω1,1fi\(θ~t\)v\)\.\\displaystyle\\text\{Re\}\(B\_\{2\}\)\+A\_\{2,\\text\{Hes\}\}=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\ell^\{\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\\left\(\\text\{Re\}\\left\[v^\{T\}\\nabla^\{2,0\}\_\{\\omega\}f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)v\\right\]\+v^\{\\dagger\}\\nabla^\{1,1\}\_\{\\omega\}f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)v\\right\)\.\(10\.22\)Applying Cauchy\-Schwarz,
Re\(B2\)\+A2,Hes≤1n∑i=1n\(ℓi′\(fi\(θ~t\)\)\)21n∑i=1n\|Re\[vT∇ω2,0fi\(θ~t\)\]\+v†∇ω1,1fi\(θ~t\)\|2\.\\displaystyle\\text\{Re\}\(B\_\{2\}\)\+A\_\{2,\\text\{Hes\}\}\\leq\\sqrt\{\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\(\\ell^\{\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\)^\{2\}\}\\sqrt\{\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\left\|\\text\{Re\}\\left\[v^\{T\}\\nabla^\{2,0\}\_\{\\omega\}f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\\right\]\+v^\{\\dagger\}\\nabla^\{1,1\}\_\{\\omega\}f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\\right\|^\{2\}\}\.\(10\.23\)Again invoking the established an asymptotic bound on the \(1,1\)\-Hessian in Appendix[8\.2](https://arxiv.org/html/2608.19584#S8.SS2)and \(2,0\)\-Hessian in Appendix[8\.5](https://arxiv.org/html/2608.19584#S8.SS5)\. Due to the squaring, the above is quadratic invv, and we can rewrite
Re\(B2\)\+A2,Hes≤1n∑i=1n\(ℓi′\(fi\(θ~t\)\)\)2‖v‖ω2\.\\displaystyle\\text\{Re\}\(B\_\{2\}\)\+A\_\{2,\\text\{Hes\}\}\\leq\\sqrt\{\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\(\\ell^\{\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\)^\{2\}\}\\\|v\\\|\_\{\\omega\}^\{2\}\.\(10\.24\)Combining what we have,
12∇ω2L\(θ~t\)\(v,v\)≤\(ρJ\+Cℋ1n∑i=1n\(ℓi′\(fi\(θ~t\)\)\)2m\)‖v‖ω2\.\\displaystyle\\frac\{1\}\{2\}\\nabla^\{2\}\_\{\\omega\}L\(\\widetilde\{\\theta\}\_\{t\}\)\(v,v\)\\leq\\left\(\\rho\_\{J\}\+\\frac\{C\_\{\\mathcal\{H\}\}\\sqrt\{\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\(\\ell^\{\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\)^\{2\}\}\}\{\\sqrt\{m\}\}\\right\)\\\|v\\\|\_\{\\omega\}^\{2\}\.\(10\.25\)Sincedω\(θ0,θt\)2=2‖v‖ω2d\_\{\\omega\}\(\\theta\_\{0\},\\theta\_\{t\}\)^\{2\}=2\\\|v\\\|\_\{\\omega\}^\{2\}, we can also rewrite the above with this\. This completes the proof\.
□\\square
Figure 9:We plotβ\\beta\-smoothness coefficient,\(ρJ\+Cℋ1n∑i=1n\(ℓi′\(fi\(θ~t\)\)\)2m\)\\left\(\\rho\_\{J\}\+\\frac\{C\_\{\\mathcal\{H\}\}\\sqrt\{\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\(\\ell^\{\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\)^\{2\}\}\}\{\\sqrt\{m\}\}\\right\), across a spectrum of paths ranging fromm=50m=50\(the upper bound\) andm=800m=800\(the lower bound\)\.In the Calabi\-Yau scenario, assuming the minimum eigenvalues have no strict lower bound, we can note the following\. Consider the constants in the bound:ρJ\\rho\_\{J\}, which bounds the first\-order Jacobian quadratic form; andCℋC\_\{\\mathcal\{H\}\}, which bounds the covariant Hessian norm\. In a Calabi\-Yau constant determinant condition, both of these will diverge, or at least become very large\. Assuming overparameterization and insufficient regularization, we get the divergence case
lim supλmax→∞β=lim supρJ→∞lim supCℋ→∞\(ρJ\+Cℋ1n∑i=1n\(ℓi′\(fi\(θ~t\)\)\)2m\)=∞\\displaystyle\\limsup\_\{\\lambda\_\{\\text\{max\}\}\\to\\infty\}\\beta=\\limsup\_\{\\rho\_\{J\}\\to\\infty\}\\limsup\_\{C\_\{\\mathcal\{H\}\}\\to\\infty\}\\left\(\\rho\_\{J\}\+\\frac\{C\_\{\\mathcal\{H\}\}\\sqrt\{\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\(\\ell^\{\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\)^\{2\}\}\}\{\\sqrt\{m\}\}\\right\)=\\infty\(10\.26\)due to the eigenvalue explosion\. If we cohere more closely to the sufficiently regularized loss of[4\.23](https://arxiv.org/html/2608.19584#S4.E23), we will assume a lower bound on the minimum eigenvalues, although very small\. In this case,
lim supλmin→μ≪1\>0lim supλmax→λ∗≫1β=lim supρJ→very largelim supCℋ→very large\(ρJ\+Cℋ1n∑i=1n\(ℓi′\(fi\(θ~t\)\)\)2m\)≫1\.\\displaystyle\\limsup\_\{\\lambda\_\{\\text\{min\}\}\\to\\mu\_\{\\ll 1\}\>0\}\\limsup\_\{\\lambda\_\{\\text\{max\}\}\\to\\lambda^\{\*\}\\gg 1\}\\beta=\\limsup\_\{\\rho\_\{J\}\\to\\text\{very large\}\}\\limsup\_\{C\_\{\\mathcal\{H\}\}\\to\\text\{very large\}\}\\left\(\\rho\_\{J\}\+\\frac\{C\_\{\\mathcal\{H\}\}\\sqrt\{\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\(\\ell^\{\\prime\}\_\{i\}\(f\_\{i\}\(\\widetilde\{\\theta\}\_\{t\}\)\)\)^\{2\}\}\}\{\\sqrt\{m\}\}\\right\)\\gg 1\.\(10\.27\)Therefore, we get
12∇ω2L\(θ~t\)\(v,v\)≤12\(explosive value\)dω\(θ0,θt\)2,\\displaystyle\\frac\{1\}\{2\}\\nabla^\{2\}\_\{\\omega\}L\(\\widetilde\{\\theta\}\_\{t\}\)\(v,v\)\\leq\\frac\{1\}\{2\}\\left\(\\text\{explosive value\}\\right\)d\_\{\\omega\}\(\\theta\_\{0\},\\theta\_\{t\}\)^\{2\},\(10\.28\)and theβ\\beta\-smoothness result will be destroyed\.
### 10\.2Dynamic Kähler Polyak\-Łojasiewicz condition
Proof of Lemma 4\.LetU⊆𝒮U\\subseteq\\mathcal\{S\}be a local coordinate chart equipped with the Kähler metrichij¯h\_\{i\\overline\{j\}\}\. Letθ∗\\theta^\{\*\}, be a minima, and letγ\(s\)=expθt\(sv\)\\gamma\(s\)=\\exp\_\{\\theta\_\{t\}\}\(sv\)fors∈\[0,1\]s\\in\[0,1\]be the geodesic connectingθt\\theta\_\{t\}toθ∗\\theta^\{\*\}, with initial holomorphic tangent vectorv=γ˙\(0\)=expθt−1\(θ∗\)1,0v=\\dot\{\\gamma\}\(0\)=\\exp\_\{\\theta\_\{t\}\}^\{\-1\}\(\\theta^\{\*\}\)^\{1,0\}\. The second\-order Taylor expansion of the lossLLalong this geodesic is
L\(θ∗\)=L\(θt\)\+2Re⟨∇h1,0L\(θt\),v⟩h\+2∫01\(1−s\)ℋij¯\(γ\(s\)\)γ˙i\(s\)γ˙j\(s\)¯𝑑s\.\\displaystyle L\(\\theta^\{\*\}\)=L\(\\theta\_\{t\}\)\+2\\text\{Re\}\\langle\\nabla\_\{h\}^\{1,0\}L\(\\theta\_\{t\}\),v\\rangle\_\{h\}\+2\\int\_\{0\}^\{1\}\(1\-s\)\\mathcal\{H\}\_\{i\\overline\{j\}\}\(\\gamma\(s\)\)\\dot\{\\gamma\}^\{i\}\(s\)\\overline\{\\dot\{\\gamma\}^\{j\}\(s\)\}ds\.\(10\.29\)Define a dynamicΓt\\Gamma\_\{t\}that behaves similarly to theΓ\\Gammaof[10](https://arxiv.org/html/2608.19584#S10)corresponding to the minimum eigenvalue of the Hessian with respect to the metric that continually updates
Γt:=infs∈\[0,1\]λmin\(hik¯\(γt\(s\)\)ℋkj¯\(γt\(s\)\)\),whereγt\(s\)=expθt\(sv\)\.\\displaystyle\\Gamma\_\{t\}:=\\inf\_\{s\\in\[0,1\]\}\\lambda\_\{\\min\}\\left\(h^\{i\\overline\{k\}\}\(\\gamma\_\{t\}\(s\)\)\\mathcal\{H\}\_\{k\\overline\{j\}\}\(\\gamma\_\{t\}\(s\)\)\\right\),\\quad\\text\{where\}\\quad\\gamma\_\{t\}\(s\)=\\exp\_\{\\theta\_\{t\}\}\(sv\)\.\(10\.30\)We knowΓt\\Gamma\_\{t\}is guaranteed to satisfy a strong convexity result by the result of Appendix[10](https://arxiv.org/html/2608.19584#S10), and since we can always at least take
Γt≥Γ,\\displaystyle\\Gamma\_\{t\}\\geq\\Gamma,\(10\.31\)and attainΓ\\Gammato guarantee it holds\.γ\\gammais contained inUU, and the result holds for all points \(up to smoothness\) inUU, soΓ\\Gammais a worst\-case bound\. It is not clear that the definition ofΓt\\Gamma\_\{t\}can be connected to the result of[10](https://arxiv.org/html/2608.19584#S10)from definition alone, since the definition looks dissimilar to the definition ofΓ\\Gammawe had in[10](https://arxiv.org/html/2608.19584#S10)\. Instead, we can notice both are subsidiary of the fact that we examineL\(θ∗\)≥L\(θt\)\+2Re⟨∇h1,0L\(θt\),v⟩h\+term to be boundedL\(\\theta^\{\*\}\)\\geq L\(\\theta\_\{t\}\)\+2\\text\{Re\}\\langle\\nabla\_\{h\}^\{1,0\}L\(\\theta\_\{t\}\),v\\rangle\_\{h\}\+\\text\{term to be bounded\}\. In both cases, we are finding the a sufficient bound on the remaining term, thereforeΓt,Γ\\Gamma\_\{t\},\\Gammahave the same role in a strong convexity argument\. Therefore, we can bound
2∫01\(1−s\)ℋij¯γ˙iγ˙j¯𝑑s≥2∫01\(1−s\)Γthij¯γ˙iγ˙j¯𝑑s=Γt‖v‖h2,\\displaystyle 2\\int\_\{0\}^\{1\}\(1\-s\)\\mathcal\{H\}\_\{i\\overline\{j\}\}\\dot\{\\gamma\}^\{i\}\\overline\{\\dot\{\\gamma\}^\{j\}\}ds\\geq 2\\int\_\{0\}^\{1\}\(1\-s\)\\Gamma\_\{t\}h\_\{i\\overline\{j\}\}\\dot\{\\gamma\}^\{i\}\\overline\{\\dot\{\\gamma\}^\{j\}\}ds=\\Gamma\_\{t\}\\\|v\\\|\_\{h\}^\{2\},\(10\.32\)since definition a geodesic has constant speed‖γ˙\(s\)‖h2=‖γ˙\(0\)‖h2\\\|\\dot\{\\gamma\}\(s\)\\\|\_\{h\}^\{2\}=\\\|\\dot\{\\gamma\}\(0\)\\\|\_\{h\}^\{2\}\. Substituting back into[10\.29](https://arxiv.org/html/2608.19584#S10.E29),
L^θt\(v\)≥L\(θt\)\+2Re\(∇h1,0L\(θt\)ivi\)\+Γthij¯viv¯j\.\\displaystyle\\widehat\{L\}\_\{\\theta\_\{t\}\}\(v\)\\geq L\(\\theta\_\{t\}\)\+2\\text\{Re\}\\big\(\\nabla\_\{h\}^\{1,0\}L\(\\theta\_\{t\}\)\_\{i\}v^\{i\}\\big\)\+\\Gamma\_\{t\}h\_\{i\\overline\{j\}\}v^\{i\}\\overline\{v\}^\{j\}\.\(10\.33\)Taking the Wirtinger derivative with respect to the complex conjugatev¯j\\overline\{v\}^\{j\}and setting it to zero yields the minimizing tangent vectorv^i=−1Γt\(∇h1,0L\(θt\)\)i\\widehat\{v\}^\{i\}=\-\\frac\{1\}\{\\Gamma\_\{t\}\}\\big\(\\nabla\_\{h\}^\{1,0\}L\(\\theta\_\{t\}\)\\big\)^\{i\}\. Substitutingv^\\widehat\{v\}back in establishes the dynamic Kähler\-PL condition
infθ∈UL\(θ\)≥L\(θt\)−1Γt‖∇h1,0L\(θt\)‖h2\.\\displaystyle\\inf\_\{\\theta\\in U\}L\(\\theta\)\\geq L\(\\theta\_\{t\}\)\-\\frac\{1\}\{\\Gamma\_\{t\}\}\\left\\\|\\nabla\_\{h\}^\{1,0\}L\(\\theta\_\{t\}\)\\right\\\|\_\{h\}^\{2\}\.\(10\.34\)
□\\square
Remark\.We can achieve a similar effect using a Kähler Polyak\-Łojasiewicz condition but with lesser restriction, drawing connections to[5\.22](https://arxiv.org/html/2608.19584#S5.E22)\. The definition ofΓt\\Gamma\_\{t\}relies on the eigenvalue of the Hessian being positive, but sinceℋ\\mathcal\{H\}is with respect toff, this is not the same as strong convexity\. Instead, we examine
Γt\(m\):=infs∈\[0,1\]infΩ\>0\\bBigg@3\{\(iΘe−L\(ℒ\)∧ωm−1∧Ω\)γ˙t\(s\),γ˙t\(s\)\\bBigg@3\}h‖γ˙t\(s\)‖h2dVω,\\displaystyle\\Gamma\_\{t\}^\{\(m\)\}:=\\inf\_\{s\\in\[0,1\]\}\\inf\_\{\\Omega\>0\}\\frac\{\\bBigg@\{3\}\\\{\(i\\Theta\_\{e^\{\-L\}\}\(\\mathcal\{L\}\)\\wedge\\omega^\{m\-1\}\\wedge\\Omega\)\\dot\{\\gamma\}\_\{t\}\(s\),\\dot\{\\gamma\}\_\{t\}\(s\)\\bBigg@\{3\}\\\}\_\{h\}\}\{\\\|\\dot\{\\gamma\}\_\{t\}\(s\)\\\|\_\{h\}^\{2\}dV\_\{\\omega\}\},\(10\.35\)where the metric is writtene−Le^\{\-L\}is a Hermitian metric on the fibers of a trivial complex line bundleℒ=M×ℂ\\mathcal\{L\}=M\\times\\mathbb\{C\}constructed overMMand the Chern curvature formiΘe−L\(ℒ\)=i∂∂¯L=i∇ω1,1Li\\Theta\_\{e^\{\-L\}\}\(\\mathcal\{L\}\)=i\\partial\\overline\{\\partial\}L=i\\nabla\_\{\\omega\}^\{1,1\}L\. We can note
iΘ=−∂∂¯logH⇔iΘe−L\(ℒ\)=−∂∂¯log\(e−L\)=∂∂¯L\.\\displaystyle i\\Theta=\-\\partial\\overline\{\\partial\}\\log H\\quad\\iff\\quad i\\Theta\_\{e^\{\-L\}\}\(\\mathcal\{L\}\)=\-\\partial\\overline\{\\partial\}\\log\(e^\{\-L\}\)=\\partial\\overline\{\\partial\}L\.\(10\.36\)From this definition ofΓt\(m\)\\Gamma\_\{t\}^\{\(m\)\}, we have that for𝒯=spanℂ\{γ˙t\}\\mathcal\{T\}=\\text\{span\}\_\{\\mathbb\{C\}\}\\\{\\dot\{\\gamma\}\_\{t\}\\\}defining the rank\-1 descent subsheaf,Γt\(m\)\>0\\Gamma\_\{t\}^\{\(m\)\}\>0along𝒯\\mathcal\{T\}will satisfy a PL condition\. We can relax the condition forΓt\\Gamma\_\{t\}to exist, and as long as curvature is sufficiently nice along the geodesic path, we get the same result\. We can note the definition ofΓt\(m\)\\Gamma\_\{t\}^\{\(m\)\}has connections toω\\omega\-m\-semi\-positivity since this is by definition
\{\(iΘ∧ωm−1∧Ω\)u,u\}h≥0\.\\displaystyle\\Bigg\\\{\(i\\Theta\\wedge\\omega^\{m\-1\}\\wedge\\Omega\)u,u\\Bigg\\\}\_\{h\}\\geq 0\.\(10\.37\)We are interested in net curvature being positive along the geodesic\.
### 10\.3Convergence
Proof of Lemma 5\.We now analyze the forward step on the loss to guarantee it descends\. Letvt=−ηt∇h1,0L\(θt\)v\_\{t\}=\-\\eta\_\{t\}\\nabla\_\{h\}^\{1,0\}L\(\\theta\_\{t\}\)be a first derivative step with adjustable learning rateηt\\eta\_\{t\}, moving along the geodesicσ\(s\)=expθt\(svt\)\\sigma\(s\)=\\exp\_\{\\theta\_\{t\}\}\(sv\_\{t\}\)to the next parameter stateθt\+1\\theta\_\{t\+1\}\. The expansion of the loss at the updated parameters is
L\(θt\+1\)=L\(θt\)\+2Re⟨∇h1,0L\(θt\),vt⟩h\+2∫01\(1−s\)ℋij¯\(σ\(s\)\)σ˙i\(s\)σ˙j\(s\)¯𝑑s\.\\displaystyle L\(\\theta\_\{t\+1\}\)=L\(\\theta\_\{t\}\)\+2\\text\{Re\}\\langle\\nabla\_\{h\}^\{1,0\}L\(\\theta\_\{t\}\),v\_\{t\}\\rangle\_\{h\}\+2\\int\_\{0\}^\{1\}\(1\-s\)\\mathcal\{H\}\_\{i\\overline\{j\}\}\(\\sigma\(s\)\)\\dot\{\\sigma\}^\{i\}\(s\)\\overline\{\\dot\{\\sigma\}^\{j\}\(s\)\}ds\.\(10\.38\)Similarly toΓt\\Gamma\_\{t\}, define continually\-updatedβ\\beta\-smoothness parameter via spectral norm
βt:=sups∈\[0,1\]\\bBigg@3‖hik¯\(σ\(s\)\)ℋkj¯\(σ\(s\)\)\\bBigg@3‖2\.\\displaystyle\\beta\_\{t\}:=\\sup\_\{s\\in\[0,1\]\}\\bBigg@\{3\}\\\|h^\{i\\overline\{k\}\}\(\\sigma\(s\)\)\\mathcal\{H\}\_\{k\\overline\{j\}\}\(\\sigma\(s\)\)\\bBigg@\{3\}\\\|\_\{2\}\.\(10\.39\)Again,βt\\beta\_\{t\}is guaranteed to satisfy aβ\\beta\-smoothness condition because we can always take a worst\-case bound
βt≤β,\\displaystyle\\beta\_\{t\}\\leq\\beta,\(10\.40\)and attainβ\\betato guarantee the condition, whereβ\\betais as in[10\.1](https://arxiv.org/html/2608.19584#S10.SS1)\. We can note the geodesicσ\\sigmaexists in theUUin whichβ\\betais defined, andβ\\betamust hold \(almost\) everywhere in this region\. Bounding the integral term of[10\.38](https://arxiv.org/html/2608.19584#S10.E38),
2∫01\(1−s\)ℋij¯σ˙iσ˙j¯≤βt‖vt‖h2𝑑s=βtηt2‖∇h1,0L\(θt\)‖h2\.\\displaystyle 2\\int\_\{0\}^\{1\}\(1\-s\)\\mathcal\{H\}\_\{i\\overline\{j\}\}\\dot\{\\sigma\}^\{i\}\\overline\{\\dot\{\\sigma\}^\{j\}\}\\leq\\beta\_\{t\}\\\|v\_\{t\}\\\|\_\{h\}^\{2\}ds=\\beta\_\{t\}\\eta\_\{t\}^\{2\}\\\|\\nabla\_\{h\}^\{1,0\}L\(\\theta\_\{t\}\)\\\|\_\{h\}^\{2\}\.\(10\.41\)For the linear term, we use the definition ofvv
2Re⟨∇h1,0L\(θt\),−ηt∇h1,0L\(θt\)⟩h=−2ηt‖∇h1,0L\(θt\)‖h2\.\\displaystyle 2\\text\{Re\}\\langle\\nabla\_\{h\}^\{1,0\}L\(\\theta\_\{t\}\),\-\\eta\_\{t\}\\nabla\_\{h\}^\{1,0\}L\(\\theta\_\{t\}\)\\rangle\_\{h\}=\-2\\eta\_\{t\}\\\|\\nabla\_\{h\}^\{1,0\}L\(\\theta\_\{t\}\)\\\|\_\{h\}^\{2\}\.\(10\.42\)Substituting everything back into[10\.38](https://arxiv.org/html/2608.19584#S10.E38),
L\(θt\+1\)≤L\(θt\)−ηt\(2−βtηt\)‖∇h1,0L\(θt\)‖h2\.\\displaystyle L\(\\theta\_\{t\+1\}\)\\leq L\(\\theta\_\{t\}\)\-\\eta\_\{t\}\(2\-\\beta\_\{t\}\\eta\_\{t\}\)\\\|\\nabla\_\{h\}^\{1,0\}L\(\\theta\_\{t\}\)\\\|\_\{h\}^\{2\}\.\(10\.43\)With the Kähler\-PL condition of[10\.2](https://arxiv.org/html/2608.19584#S10.SS2), we get
L\(θt\+1\)−L\(θ∗\)≤\(1−Γtηt\(2−βtηt\)\)\(L\(θt\)−L\(θ∗\)\)\.\\displaystyle L\(\\theta\_\{t\+1\}\)\-L\(\\theta^\{\*\}\)\\leq\\left\(1\-\\Gamma\_\{t\}\\eta\_\{t\}\(2\-\\beta\_\{t\}\\eta\_\{t\}\)\\right\)\(L\(\\theta\_\{t\}\)\-L\(\\theta^\{\*\}\)\)\.\(10\.44\)
### 10\.4Calabi\-Yau induced oscillation
Proof of Lemma 6\.Let\(M,ω\)\(M,\\omega\)be a Kähler information manifold under the regularized loss of[4\.23](https://arxiv.org/html/2608.19584#S4.E23), so the metric is full rank, with metric eigenvalues bounded below byλmin≥μ\>0\\lambda\_\{\\min\}\\geq\\mu\>0\. As in Appendix[10\.1](https://arxiv.org/html/2608.19584#S10.SS1), we have by definition
Cℋ=sup‖u‖ω=1\|∇ω2,0f\(u,u\)\|=sup‖u‖ω=1\|∂2f\(u,u\)−Γ\(u,u\)m∂mf\|\.\\displaystyle C\_\{\\mathcal\{H\}\}=\\sup\_\{\\\|u\\\|\_\{\\omega\}=1\}\\left\|\\nabla\_\{\\omega\}^\{2,0\}f\(u,u\)\\right\|=\\sup\_\{\\\|u\\\|\_\{\\omega\}=1\}\\left\|\\partial^\{2\}f\(u,u\)\-\\Gamma\(u,u\)^\{m\}\\partial\_\{m\}f\\right\|\.\(10\.45\)We can note the norm‖u‖ω2=u†hu=1\\\|u\\\|\_\{\\omega\}^\{2\}=u^\{\\dagger\}hu=1, and moreover the Euclidean norm obeys
‖u‖22≤1λmin\(h\)≤1μ\.\\displaystyle\\\|u\\\|\_\{2\}^\{2\}\\leq\\frac\{1\}\{\\lambda\_\{\\min\}\(h\)\}\\leq\\frac\{1\}\{\\mu\}\.\(10\.46\)Therefore, borrowing from our results in[8\.2](https://arxiv.org/html/2608.19584#S8.SS2),[8\.5](https://arxiv.org/html/2608.19584#S8.SS5), we obtain
\|∂2f\(u,u\)\|≤‖u‖22‖∂2f‖2≤𝒪\(1μm\)\.\\displaystyle\\left\|\\partial^\{2\}f\(u,u\)\\right\|\\leq\\\|u\\\|\_\{2\}^\{2\}\\\|\\partial^\{2\}f\\\|\_\{2\}\\leq\\mathcal\{O\}\\left\(\\frac\{1\}\{\\mu\\sqrt\{m\}\}\\right\)\.\(10\.47\)Now we lower bound the Christoffel term of[10\.45](https://arxiv.org/html/2608.19584#S10.E45)\. We can note
Γ\(u,u\)m∂mf=uiukhml¯∂ihkl¯∂mf\.\\displaystyle\\Gamma\(u,u\)^\{m\}\\partial\_\{m\}f=u^\{i\}u^\{k\}h^\{m\\overline\{l\}\}\\partial\_\{i\}h\_\{k\\overline\{l\}\}\\partial\_\{m\}f\.\(10\.48\)By the Calabi\-Yau constant determinant condition,∏j=1Kλj=κ\\prod\_\{j=1\}^\{K\}\\lambda\_\{j\}=\\kappa\. Therefore, we assume
λmax\(h\)=κ∏j=1K−1λj=Ω\(κμK~−1\)\.\\displaystyle\\lambda\_\{\\max\}\(h\)=\\frac\{\\kappa\}\{\\prod\_\{j=1\}^\{K\-1\}\\lambda\_\{j\}\}=\\Omega\\left\(\\frac\{\\kappa\}\{\\mu^\{\\widetilde\{K\}\-1\}\}\\right\)\.\(10\.49\)HereK~\\widetilde\{K\}is some value such thatK~≤K\\widetilde\{K\}\\leq K\. In particular, we want to avoid an𝒪\(⋅\)\\mathcal\{O\}\(\\cdot\)bound since this will disrupt our final result because the final result holds with an "at least" bound, not an "at most" bound\. However, sinceλ≥μ\\lambda\\geq\\mufor allλ\\lambda, keeping the bound in terms ofKKuses an𝒪\(⋅\)\\mathcal\{O\}\(\\cdot\)bound, therefore we must "cut off" extra powers to establish anΩ\(⋅\)\\Omega\(\\cdot\)bound\. Recall from section[4\.3](https://arxiv.org/html/2608.19584#S4.SS3)the metric followsh≈𝔼\[\(∂f\)†∂f\]\+μIh\\approx\\mathbb\{E\}\[\(\\partial f\)^\{\\dagger\}\\partial f\]\+\\mu I, so we haveλmax\(h\)≈‖∂f‖22\+μ\\lambda\_\{\\max\}\(h\)\\approx\\\|\\partial f\\\|\_\{2\}^\{2\}\+\\mu\. Therefore, the singular value can be deduced as∥∂f∥2∼λmax=Ω\(κμ−\(K−1\)/2\)\\\|\\partial f\\\|\_\{2\}\\sim\\sqrt\{\\lambda\_\{\\max\}\}=\\Omega\(\\sqrt\{\\kappa\}\\mu^\{\-\(K\-1\)/2\}\)\. The additionalμ\\muis dropped since its subdominant in the asymptotics\. Recall from Appendix[8\.3](https://arxiv.org/html/2608.19584#S8.SS3)\(equation[8\.37](https://arxiv.org/html/2608.19584#S8.E37)\) the derivative of the metric scales in proportion to𝒪\(‖∂2f‖2‖∂f‖2\)\\mathcal\{O\}\(\\\|\\partial^\{2\}f\\\|\_\{2\}\\\|\\partial f\\\|\_\{2\}\)\. Therefore, we can note
\|uiukhml¯∂ihkl¯∂mf\|\\displaystyle\\left\|u^\{i\}u^\{k\}h^\{m\\overline\{l\}\}\\partial\_\{i\}h\_\{k\\overline\{l\}\}\\partial\_\{m\}f\\right\|=Ω\(‖u‖22⋅‖h−1‖2⋅‖∂2f‖2⋅‖∂f‖22\)\\displaystyle=\\Omega\\left\(\\\|u\\\|\_\{2\}^\{2\}\\cdot\\\|h^\{\-1\}\\\|\_\{2\}\\cdot\\\|\\partial^\{2\}f\\\|\_\{2\}\\cdot\\\|\\partial f\\\|\_\{2\}^\{2\}\\right\)\(10\.50\)=Ω\(1μ⋅1μ⋅1m⋅\(κμ−\(K~−1\)/2\)2\)\\displaystyle=\\Omega\\left\(\\frac\{1\}\{\\mu\}\\cdot\\frac\{1\}\{\\mu\}\\cdot\\frac\{1\}\{\\sqrt\{m\}\}\\cdot\\left\(\\sqrt\{\\kappa\}\\mu^\{\-\(\\widetilde\{K\}\-1\)/2\}\\right\)^\{2\}\\right\)\(10\.51\)=Ω\(κmμ−\(K~\+1\)\)\.\\displaystyle=\\Omega\\left\(\\frac\{\\kappa\}\{\\sqrt\{m\}\}\\mu^\{\-\(\\widetilde\{K\}\+1\)\}\\right\)\.\(10\.52\)Returning toCℋC\_\{\\mathcal\{H\}\},
Cℋ≥\|Γ\(u,u\)m∂mf\|−\|∂2f\(u,u\)\|=Ω\(κmμ−\(K~\+1\)\)\.\\displaystyle C\_\{\\mathcal\{H\}\}\\geq\\left\|\\Gamma\(u,u\)^\{m\}\\partial\_\{m\}f\\right\|\-\\left\|\\partial^\{2\}f\(u,u\)\\right\|=\\Omega\\left\(\\frac\{\\kappa\}\{\\sqrt\{m\}\}\\mu^\{\-\(\\widetilde\{K\}\+1\)\}\\right\)\.\(10\.53\)It immediately follows substituting this into theβ\\beta\-smoothness scalar
β=ρJ\+2CℋL\(θ~t\)m=Ω\(κmμ−\(K~\+1\)L\(θ~t\)\)\.\\displaystyle\\beta=\\rho\_\{J\}\+\\frac\{2C\_\{\\mathcal\{H\}\}\\sqrt\{L\(\\widetilde\{\\theta\}\_\{t\}\)\}\}\{\\sqrt\{m\}\}=\\Omega\\left\(\\frac\{\\kappa\}\{m\}\\mu^\{\-\(\\widetilde\{K\}\+1\)\}\\sqrt\{L\(\\widetilde\{\\theta\}\_\{t\}\)\}\\right\)\.\(10\.54\)From Appendix[10\.3](https://arxiv.org/html/2608.19584#S10.SS3), we desire
L\(θt\+1\)−L\(θ∗\)≤\(1−2μηt\+μβηt2\)\(L\(θt\)−L\(θ∗\)\)\.\\displaystyle L\(\\theta\_\{t\+1\}\)\-L\(\\theta^\{\*\}\)\\leq\\left\(1\-2\\mu\\eta\_\{t\}\+\\mu\\beta\\eta\_\{t\}^\{2\}\\right\)\(L\(\\theta\_\{t\}\)\-L\(\\theta^\{\*\}\)\)\.\(10\.55\)Becauseβ∼Ω\(μ−\(K~\+1\)\)\\beta\\sim\\Omega\(\\mu^\{\-\(\\widetilde\{K\}\+1\)\}\), the sufficient threshold to counteract this is𝒪\(μK~\+1\)\\mathcal\{O\}\(\\mu^\{\\widetilde\{K\}\+1\}\)\. For viable learning rateηt\\eta\_\{t\}the quadratic penaltyμβηt2\\mu\\beta\\eta\_\{t\}^\{2\}dominates\.
□\\square
### 10\.5Failure of the Kähler Polyak\-Łojasiewicz condition under Calabi\-Yau metrics
In this section, we demonstrate the failure of the results of[10\.2](https://arxiv.org/html/2608.19584#S10.SS2)via an eigenvalue blow\-up effect of Calabi\-Yau manifolds\. This result is unique to eigenvalue blow\-up, hence Calabi\-Yau metrics, and not inherent to negative curvature\.
Proof of Lemma 7\.Recall from Appendix[10\.2](https://arxiv.org/html/2608.19584#S10.SS2)we defined
Γt:=infs∈\[0,1\]λmin\(hik¯\(γt\(s\)\)ℋkj¯\(γt\(s\)\)\)\.\\displaystyle\\Gamma\_\{t\}:=\\inf\_\{s\\in\[0,1\]\}\\lambda\_\{\\min\}\(h^\{i\\overline\{k\}\}\(\\gamma\_\{t\}\(s\)\)\\mathcal\{H\}\_\{k\\overline\{j\}\}\(\\gamma\_\{t\}\(s\)\)\)\.\(10\.56\)Let us note the following property from linear algebra\. For any Hermitian positive\-definite matrixAAand Hermitian matrixBB, the minimum eigenvalue of their product is upper\-bounded by the product of their respective eigenvalues such thatλmin\(AB\)≤λmin\(A\)λmax\(B\)\\lambda\_\{\\min\}\(AB\)\\leq\\lambda\_\{\\min\}\(A\)\\lambda\_\{\\max\}\(B\)\. Therefore, we get
Γt≤λmin\(h−1\)λmax\(ℋ\)=λmax\(ℋ\)λmax\(h\)\.\\displaystyle\\Gamma\_\{t\}\\leq\\lambda\_\{\\min\}\(h^\{\-1\}\)\\lambda\_\{\\max\}\(\\mathcal\{H\}\)=\\frac\{\\lambda\_\{\\max\}\(\\mathcal\{H\}\)\}\{\\lambda\_\{\\max\}\(h\)\}\.\(10\.57\)Under the Calabi\-Yau constant determinant condition where∏j=1Kλj=κ\\prod\_\{j=1\}^\{K\}\\lambda\_\{j\}=\\kappa, we established in Appendix[10\.4](https://arxiv.org/html/2608.19584#S10.SS4)the explosion of the maximum eigenvalue assumption
λmax\(h\)=Ω\(κμK−1\)\.\\displaystyle\\lambda\_\{\\max\}\(h\)=\\Omega\\left\(\\frac\{\\kappa\}\{\\mu^\{K\-1\}\}\\right\)\.\(10\.58\)Furthermore, by our previous bound in[8\.2](https://arxiv.org/html/2608.19584#S8.SS2), we havesupθ∈𝒮‖ℋ‖2=𝒪\(1m\)\\sup\_\{\\theta\\in\\mathcal\{S\}\}\\\|\\mathcal\{H\}\\\|\_\{2\}=\\mathcal\{O\}\\left\(\\frac\{1\}\{\\sqrt\{m\}\}\\right\)\. Therefore,
Γt≤𝒪\(m−1/2\)Ω\(κμ−\(K−1\)\)=𝒪\(μK−1κm\)\.\\displaystyle\\Gamma\_\{t\}\\leq\\frac\{\\mathcal\{O\}\\left\(m^\{\-1/2\}\\right\)\}\{\\Omega\\left\(\\kappa\\mu^\{\-\(K\-1\)\}\\right\)\}=\\mathcal\{O\}\\left\(\\frac\{\\mu^\{K\-1\}\}\{\\kappa\\sqrt\{m\}\}\\right\)\.\(10\.59\)By the regularized loss of[4\.23](https://arxiv.org/html/2608.19584#S4.E23),0<μ≪10<\\mu\\ll 1is small butμ↛0\\mu\\nrightarrow 0andKKis large, especially in the overparameterization regime \(also note the division bym\\sqrt\{m\}\), soΓt\\Gamma\_\{t\}is very small, and in fact converging to00asKKgrows but not becauseμ\\mugoes to00\. Therefore, when examining
L\(θt\+1\)−L\(θ∗\)≤\(1−Γtηt\(2−βtηt\)\)\(L\(θt\)−L\(θ∗\)\),\\displaystyle L\(\\theta\_\{t\+1\}\)\-L\(\\theta^\{\*\}\)\\leq\(1\-\\Gamma\_\{t\}\\eta\_\{t\}\(2\-\\beta\_\{t\}\\eta\_\{t\}\)\)\(L\(\\theta\_\{t\}\)\-L\(\\theta^\{\*\}\)\),\(10\.60\)the constant on the right\-hand side is 0\.9999999…\\ldotsfor largeKK, and so the convergence guarantee is diminished in effect\.
□\\square
### 10\.6Regret bounds
In this section, we examine regret bounds\. In[40](https://arxiv.org/html/2608.19584#bib.bib38), they examine an upper bound on regret\. The primary goal of this work is to show their algorithm succeeds\. Our work is motivated by the opposite: it is in our interest to show the Calabi\-Yau scenario fails\. This motivates us to find a lower bound\. As we will see, our proof depends on parameter dimensionKK, which is closely related to width\. Thus, our results are consistent with the goals of[8\.2](https://arxiv.org/html/2608.19584#S8.SS2),[8\.5](https://arxiv.org/html/2608.19584#S8.SS5)\.
Proof of Lemma 8\.We can note‖expθt−1\(θ∗\)1,0‖h\\Big\\\|\\exp\_\{\\theta\_\{t\}\}^\{\-1\}\(\\theta^\{\*\}\)^\{1,0\}\\Big\\\|\_\{h\}is the geodesic distancedω\(θt,θ∗\)d\_\{\\omega\}\(\\theta\_\{t\},\\theta^\{\*\}\)\. Let us denote
δ\(θ\):=12dω\(θ,θ∗\)2\.\\displaystyle\\delta\(\\theta\):=\\frac\{1\}\{2\}d\_\{\\omega\}\(\\theta,\\theta^\{\*\}\)^\{2\}\.\(10\.61\)The Riemannian gradient can be computed as
∇hδ\(θ\)=−expθ−1\(θ∗\)1,0\.\\displaystyle\\nabla\_\{h\}\\delta\(\\theta\)=\-\\exp\_\{\\theta\}^\{\-1\}\(\\theta^\{\*\}\)^\{1,0\}\.\(10\.62\)Under stochastic natural gradient descent, the parameter updates asdθt=−∇hℒ\(θt\)dt\+ηdWtd\\theta\_\{t\}=\-\\nabla\_\{h\}\\mathcal\{L\}\(\\theta\_\{t\}\)dt\+\\sqrt\{\\eta\}dW\_\{t\}\. By Itô’s lemma, the differential ofSScan be written as
dδ\(θt\)=Re⟨∇hδ\(θt\),dθt⟩h\+ηΔ∂¯δ\(θt\)dt\.\\displaystyle d\\delta\(\\theta\_\{t\}\)=\\text\{Re\}\\langle\\nabla\_\{h\}\\delta\(\\theta\_\{t\}\),d\\theta\_\{t\}\\rangle\_\{h\}\+\\eta\\Delta\_\{\\overline\{\\partial\}\}\\delta\(\\theta\_\{t\}\)dt\.\(10\.63\)Substituting in[10\.62](https://arxiv.org/html/2608.19584#S10.E62)and the parameter update rule,
Re⟨∇hδ\(θt\),dθt⟩h\\displaystyle\\text\{Re\}\\langle\\nabla\_\{h\}\\delta\(\\theta\_\{t\}\),d\\theta\_\{t\}\\rangle\_\{h\}=Re⟨−expθt−1\(θ∗\)1,0,−∇hℒ\(θt\)dt\+ηdWt⟩h\\displaystyle=\\text\{Re\}\\Big\\langle\-\\exp\_\{\\theta\_\{t\}\}^\{\-1\}\(\\theta^\{\*\}\)^\{1,0\},\-\\nabla\_\{h\}\\mathcal\{L\}\(\\theta\_\{t\}\)dt\+\\sqrt\{\\eta\}dW\_\{t\}\\Big\\rangle\_\{h\}\(10\.64\)=linearityRe⟨∇hℒ\(θt\),expθt−1\(θ∗\)1,0⟩hdt−ηRe⟨expθt−1\(θ∗\)1,0,dWt⟩h\.\\displaystyle\\stackrel\{\{\\scriptstyle\\text\{linearity\}\}\}\{\{=\}\}\\text\{Re\}\\langle\\nabla\_\{h\}\\mathcal\{L\}\(\\theta\_\{t\}\),\\exp\_\{\\theta\_\{t\}\}^\{\-1\}\(\\theta^\{\*\}\)^\{1,0\}\\rangle\_\{h\}dt\-\\sqrt\{\\eta\}\\text\{Re\}\\langle\\exp\_\{\\theta\_\{t\}\}^\{\-1\}\(\\theta^\{\*\}\)^\{1,0\},dW\_\{t\}\\rangle\_\{h\}\.\(10\.65\)The first term here is found in the regret term of[5\.12](https://arxiv.org/html/2608.19584#S5.E12)\. Integrating,
δ\(θT\)−δ\(θ0\)\\displaystyle\\delta\(\\theta\_\{T\}\)\-\\delta\(\\theta\_\{0\}\)=∫0TRe⟨∇hℒ\(θt\),expθt−1\(θ∗\)1,0⟩hdt⏟=ℛ\(T\)\\displaystyle=\\underbrace\{\\int\_\{0\}^\{T\}\\text\{Re\}\\langle\\nabla\_\{h\}\\mathcal\{L\}\(\\theta\_\{t\}\),\\exp\_\{\\theta\_\{t\}\}^\{\-1\}\(\\theta^\{\*\}\)^\{1,0\}\\rangle\_\{h\}dt\}\_\{=\\mathcal\{R\}\(T\)\}\(10\.66\)−∫0TηRe⟨expθt−1\(θ∗\)1,0,dWt⟩h⏟⟹𝔼\[∫0TηRe⟨expθt−1\(θ∗\)1,0,dWt⟩h\]=0\+η∫0TΔ∂¯δ\(θt\)𝑑t\.\\displaystyle\\quad\\quad\\quad\\quad\-\\underbrace\{\\int\_\{0\}^\{T\}\\sqrt\{\\eta\}\\text\{Re\}\\langle\\exp\_\{\\theta\_\{t\}\}^\{\-1\}\(\\theta^\{\*\}\)^\{1,0\},dW\_\{t\}\\rangle\_\{h\}\}\_\{\\implies\\mathbb\{E\}\\left\[\\int\_\{0\}^\{T\}\\sqrt\{\\eta\}\\text\{Re\}\\langle\\exp\_\{\\theta\_\{t\}\}^\{\-1\}\(\\theta^\{\*\}\)^\{1,0\},dW\_\{t\}\\rangle\_\{h\}\\right\]=0\}\+\\eta\\int\_\{0\}^\{T\}\\Delta\_\{\\overline\{\\partial\}\}\\delta\(\\theta\_\{t\}\)dt\.\(10\.67\)Because an Itô integral with respect to Brownian motion is a martingale, we get a vanishing term with respect to the Brownian filtration\. Simplifying and rearranging, we get a decomposition
∫U∫Ωℛ\(T\)𝑑𝕎\(ω\)p\(θ0\)ωKK\!=∫U∫Ωδ\(θT\)𝑑𝕎\(ω\)p\(θ0\)ωKK\!−δ\(θ0\)−η∫U∫Ω∫0TΔ∂¯δ\(θt\)𝑑t𝑑𝕎\(ω\)p\(θ0\)ωKK\!\.\\displaystyle\\int\_\{U\}\\int\_\{\\Omega\}\\mathcal\{R\}\(T\)d\\mathbb\{W\}\(\\omega\)p\(\\theta\_\{0\}\)\\frac\{\\omega^\{K\}\}\{K\!\}=\\int\_\{U\}\\int\_\{\\Omega\}\\delta\(\\theta\_\{T\}\)d\\mathbb\{W\}\(\\omega\)p\(\\theta\_\{0\}\)\\frac\{\\omega^\{K\}\}\{K\!\}\-\\delta\(\\theta\_\{0\}\)\-\\eta\\int\_\{U\}\\int\_\{\\Omega\}\\int\_\{0\}^\{T\}\\Delta\_\{\\overline\{\\partial\}\}\\delta\(\\theta\_\{t\}\)dtd\\mathbb\{W\}\(\\omega\)p\(\\theta\_\{0\}\)\\frac\{\\omega^\{K\}\}\{K\!\}\.\(10\.68\)Now, by the Laplacian comparison theorem and the claim \(see[25](https://arxiv.org/html/2608.19584#bib.bib39)[67](https://arxiv.org/html/2608.19584#bib.bib40)for relevant literature, although this exact equation is not given\), for a Kähler manifold of complex dimensionKK, the Laplacian of the squared distance function is bounded along the minimizing geodesicγ\(s\)\\gamma\(s\)parameterized by arc lengths∈\[0,dω\(θt,θ∗\)\]s\\in\[0,d\_\{\\omega\}\(\\theta\_\{t\},\\theta^\{\*\}\)\]via
Δ∂¯δ\(θt\)≤K−12dω\(θt,θ∗\)∫0dω\(θt,θ∗\)s2Ric\(γ˙\(s\),γ˙\(s\)\)𝑑s\.\\displaystyle\\Delta\_\{\\overline\{\\partial\}\}\\delta\(\\theta\_\{t\}\)\\leq K\-\\frac\{1\}\{2d\_\{\\omega\}\(\\theta\_\{t\},\\theta^\{\*\}\)\}\\int\_\{0\}^\{d\_\{\\omega\}\(\\theta\_\{t\},\\theta^\{\*\}\)\}s^\{2\}\\text\{Ric\}\\big\(\\dot\{\\gamma\}\(s\),\\dot\{\\gamma\}\(s\)\\big\)ds\.\(10\.69\)Define the defect term
𝔼\[ℰRic\]:=−η2dω\(θt,θ∗\)𝔼∫0dω\(θt,θ∗\)s2Ric\(γ˙\(s\),γ˙\(s\)\)ds\.\\displaystyle\\mathbb\{E\}\[\\mathcal\{E\}\_\{\\text\{Ric\}\}\]:=\-\\frac\{\\eta\}\{2d\_\{\\omega\}\(\\theta\_\{t\},\\theta^\{\*\}\)\}\\mathbb\{E\}\\int\_\{0\}^\{d\_\{\\omega\}\(\\theta\_\{t\},\\theta^\{\*\}\)\}s^\{2\}\\text\{Ric\}\\big\(\\dot\{\\gamma\}\(s\),\\dot\{\\gamma\}\(s\)\\big\)ds\.\(10\.70\)Into our regret bound,
𝔼\[ℛ\(T\)\]≥𝔼\[δ\(θT\)\]−δ\(θ0\)−ηKT−𝔼\[ℰRic\]\.\\displaystyle\\mathbb\{E\}\[\\mathcal\{R\}\(T\)\]\\geq\\mathbb\{E\}\[\\delta\(\\theta\_\{T\}\)\]\-\\delta\(\\theta\_\{0\}\)\-\\eta KT\-\\mathbb\{E\}\[\\mathcal\{E\}\_\{\\text\{Ric\}\}\]\.\(10\.71\)
□\\square
Claim\.We prove equation[10\.69](https://arxiv.org/html/2608.19584#S10.E69), which is nontrivial and challenging to find in literature\. Let\(M,ω\)\(M,\\omega\)be a Kähler manifold of complex dimensionKK\. For anyθ\\thetawhereδ\\deltais smooth, letγ\(s\)\\gamma\(s\)be the unit\-speed minimizing geodesic fromθ∗\\theta^\{\*\}toθ\\thetaparameterized by arc lengths∈\[0,r\]s\\in\[0,r\], wherer=dω\(θ,θ∗\)r=d\_\{\\omega\}\(\\theta,\\theta^\{\*\}\)\. Now, we first use the real Riemannian LaplacianΔd\\Delta\_\{d\}\. The Laplacian of the distance functionr\(θ\)r\(\\theta\)can be bounded using the index form of the second variation of arc length\. LetE1\(s\),…,E2K−1\(s\)E\_\{1\}\(s\),\\dots,E\_\{2K\-1\}\(s\)be an orthonormal frame of parallel vector fields alongγ\\gammathat are orthogonal toγ˙\\dot\{\\gamma\}\. Construct the Jacobi test fieldsYi\(s\)=srEi\(s\)Y\_\{i\}\(s\)=\\frac\{s\}\{r\}E\_\{i\}\(s\)\. The index lemma[70](https://arxiv.org/html/2608.19584#bib.bib41)provides the upper bound
Δdr\(θ\)≤∑i=12K−1I\(Yi,Yi\)=∫0r∑i=12K−1\(‖∇γ˙Yi‖h2−⟨R\(Yi,γ˙\)γ˙,Yi⟩h\)𝑑s\\displaystyle\\Delta\_\{d\}r\(\\theta\)\\leq\\sum\_\{i=1\}^\{2K\-1\}I\(Y\_\{i\},Y\_\{i\}\)=\\int\_\{0\}^\{r\}\\sum\_\{i=1\}^\{2K\-1\}\\left\(\\\|\\nabla\_\{\\dot\{\\gamma\}\}Y\_\{i\}\\\|\_\{h\}^\{2\}\-\\langle R\(Y\_\{i\},\\dot\{\\gamma\}\)\\dot\{\\gamma\},Y\_\{i\}\\rangle\_\{h\}\\right\)ds\(10\.72\)Because the frameEiE\_\{i\}is parallel, the covariant derivative simplifies to∇γ˙Yi=1rEi\\nabla\_\{\\dot\{\\gamma\}\}Y\_\{i\}=\\frac\{1\}\{r\}E\_\{i\}, meaning the first term sums to2K−1r2\\frac\{2K\-1\}\{r^\{2\}\}\. For the curvature term, substitutingYi\(s\)Y\_\{i\}\(s\)pulls out a factor ofs2r2\\frac\{s^\{2\}\}\{r^\{2\}\}\. Summing the Riemann tensor over the orthonormal frame recovers Ricci curvature
∑i=12K−1⟨R\(Ei,γ˙\)γ˙,Ei⟩h=Ric\(γ˙,γ˙\)\.\\displaystyle\\sum\_\{i=1\}^\{2K\-1\}\\langle R\(E\_\{i\},\\dot\{\\gamma\}\)\\dot\{\\gamma\},E\_\{i\}\\rangle\_\{h\}=\\text\{Ric\}\(\\dot\{\\gamma\},\\dot\{\\gamma\}\)\.\(10\.73\)Integrating over\[0,r\]\[0,r\]gives
Δdr\(θ\)≤2K−1r−1r2∫0rs2Ric\(γ˙\(s\),γ˙\(s\)\)𝑑s\.\\displaystyle\\Delta\_\{d\}r\(\\theta\)\\leq\\frac\{2K\-1\}\{r\}\-\\frac\{1\}\{r^\{2\}\}\\int\_\{0\}^\{r\}s^\{2\}\\text\{Ric\}\(\\dot\{\\gamma\}\(s\),\\dot\{\\gamma\}\(s\)\)ds\.\(10\.74\)We transition to the squared distanceδ=12r2\\delta=\\frac\{1\}\{2\}r^\{2\}\. By the chain rule,Δdδ=rΔdr\+‖∇r‖2\\Delta\_\{d\}\\delta=r\\Delta\_\{d\}r\+\\\|\\nabla r\\\|^\{2\}\. Since‖∇r‖2=1\\\|\\nabla r\\\|^\{2\}=1, we multiply our bound byrrand add11
Δdδ\(θ\)≤2K−1r∫0rs2Ric\(γ˙\(s\),γ˙\(s\)\)𝑑s\.\\displaystyle\\Delta\_\{d\}\\delta\(\\theta\)\\leq 2K\-\\frac\{1\}\{r\}\\int\_\{0\}^\{r\}s^\{2\}\\text\{Ric\}\(\\dot\{\\gamma\}\(s\),\\dot\{\\gamma\}\(s\)\)ds\.\(10\.75\)Finally, on a Kähler manifold, the real Laplacian and the complex Dolbeault Laplacian acting on functions are related byΔd=2Δ∂¯\\Delta\_\{d\}=2\\Delta\_\{\\overline\{\\partial\}\}\. Dividing by 2 completes the proof\.
□\\square
## 11First derivative norm bounds and roles of negative curvature
### 11\.1Dirichlet energy bounds
In this section, we examine a Dirichlet asymptotic scaling for sufficient setSS\. Dirichlet bounds are relevant in optimization literature[58](https://arxiv.org/html/2608.19584#bib.bib17)[24](https://arxiv.org/html/2608.19584#bib.bib18)[59](https://arxiv.org/html/2608.19584#bib.bib19)[54](https://arxiv.org/html/2608.19584#bib.bib20)[53](https://arxiv.org/html/2608.19584#bib.bib21)[33](https://arxiv.org/html/2608.19584#bib.bib22)[19](https://arxiv.org/html/2608.19584#bib.bib23)for measuring sensitivity of the neural network with respect to its weights\. This has connections to how fast the parameter descends since the squared norm of parameter gradient is the trace of the neural tangent kernel \(NTK\), and recall a trace is the sum of eigenvalues\. The eigenvalues of the NTK are often intertwined with the learning process such as through convergence speed and spectral bias[52](https://arxiv.org/html/2608.19584#bib.bib32)\.
Proof of Lemma 9\.Sinceffis real\-valued, we get a splitdf=∂f\+∂¯fdf=\\partial f\+\\overline\{\\partial\}f\. Let us begin with the \(form variety\) of the Dirichlet energy
E\(f\)=1Volω\(𝒮\)∫𝒮i∂f∧∂¯f∧ωK−1\(K−1\)\!\.\\displaystyle E\(f\)=\\frac\{1\}\{\\text\{Vol\}\_\{\\omega\}\(\\mathcal\{S\}\)\}\\int\_\{\\mathcal\{S\}\}i\\partial f\\wedge\\overline\{\\partial\}f\\wedge\\frac\{\\omega^\{K\-1\}\}\{\(K\-1\)\!\}\.\(11\.1\)The goal is to shift the Dolbeault operator off of∂f\\partial fto establish a Hessian formulation\. We use the exterior derivatived=∂\+∂¯d=\\partial\+\\overline\{\\partial\}and apply the Leibniz rule to the\(2K−1\)\(2K\-1\)\-formf∂¯f∧ωK−1f\\overline\{\\partial\}f\\wedge\\omega^\{K\-1\}, which gives
d\(f∂¯f∧ωK−1\)=df∧∂¯f∧ωK−1\+fd\(∂¯f\)∧ωK−1\+f∂¯f∧d\(ωK−1\)\.\\displaystyle d\\left\(f\\overline\{\\partial\}f\\wedge\\omega^\{K\-1\}\\right\)=df\\wedge\\overline\{\\partial\}f\\wedge\\omega^\{K\-1\}\+fd\\left\(\\overline\{\\partial\}f\\right\)\\wedge\\omega^\{K\-1\}\+f\\overline\{\\partial\}f\\wedge d\\left\(\\omega^\{K\-1\}\\right\)\.\(11\.2\)dω=0d\\omega=0since the manifold is Kähler\. Expandingdf=∂f\+∂¯fdf=\\partial f\+\\overline\{\\partial\}f, we note that∂¯f∧∂¯f=0\\overline\{\\partial\}f\\wedge\\overline\{\\partial\}f=0, leaving only∂f∧∂¯f\\partial f\\wedge\\overline\{\\partial\}f\. Furthermore,d\(∂¯f\)=\(∂\+∂¯\)∂¯f=∂∂¯fd\(\\overline\{\\partial\}f\)=\(\\partial\+\\overline\{\\partial\}\)\\overline\{\\partial\}f=\\partial\\overline\{\\partial\}f\. Substituting in, and scaling by a constant,
d\(if∂¯f∧ωK−1\(K−1\)\!\)=i∂f∧∂¯f∧ωK−1\(K−1\)\!⏟Dirichlet integrand\+f\(i∂∂¯f∧ωK−1\(K−1\)\!\)\.\\displaystyle d\\left\(if\\overline\{\\partial\}f\\wedge\\frac\{\\omega^\{K\-1\}\}\{\(K\-1\)\!\}\\right\)=\\underbrace\{i\\partial f\\wedge\\overline\{\\partial\}f\\wedge\\frac\{\\omega^\{K\-1\}\}\{\(K\-1\)\!\}\}\_\{\\text\{Dirichlet integrand\}\}\+f\\left\(i\\partial\\overline\{\\partial\}f\\wedge\\frac\{\\omega^\{K\-1\}\}\{\(K\-1\)\!\}\\right\)\.\(11\.3\)Notice the middle term is the integrand of[11\.1](https://arxiv.org/html/2608.19584#S11.E1)\. Integrating, and by Stokes’ theorem,
1Volω\(𝒮\)∮∂𝒮if∂¯f∧ωK−1\(K−1\)\!=E\(f\)\+1Volω\(𝒮\)∫𝒮f\(i∂∂¯f∧ωK−1\(K−1\)\!\)\.\\displaystyle\\frac\{1\}\{\\text\{Vol\}\_\{\\omega\}\(\\mathcal\{S\}\)\}\\oint\_\{\\partial\\mathcal\{S\}\}if\\overline\{\\partial\}f\\wedge\\frac\{\\omega^\{K\-1\}\}\{\(K\-1\)\!\}=E\(f\)\+\\frac\{1\}\{\\text\{Vol\}\_\{\\omega\}\(\\mathcal\{S\}\)\}\\int\_\{\\mathcal\{S\}\}f\\left\(i\\partial\\overline\{\\partial\}f\\wedge\\frac\{\\omega^\{K\-1\}\}\{\(K\-1\)\!\}\\right\)\.\(11\.4\)Observe the identity
i∂∂¯f∧ωK−1\(K−1\)\!=\(Δ∂¯f\)ωKK\!\.\\displaystyle i\\partial\\overline\{\\partial\}f\\wedge\\frac\{\\omega^\{K\-1\}\}\{\(K\-1\)\!\}=\(\\Delta\_\{\\overline\{\\partial\}\}f\)\\frac\{\\omega^\{K\}\}\{K\!\}\.\(11\.5\)Therefore, we get
1Volω\(𝒮\)∮∂𝒮if∂¯f∧ωK−1\(K−1\)\!=E\(f\)\+1Volω\(𝒮\)∫𝒮f\(Δ∂¯f\)ωKK\!,\\displaystyle\\frac\{1\}\{\\text\{Vol\}\_\{\\omega\}\(\\mathcal\{S\}\)\}\\oint\_\{\\partial\\mathcal\{S\}\}if\\overline\{\\partial\}f\\wedge\\frac\{\\omega^\{K\-1\}\}\{\(K\-1\)\!\}=E\(f\)\+\\frac\{1\}\{\\text\{Vol\}\_\{\\omega\}\(\\mathcal\{S\}\)\}\\int\_\{\\mathcal\{S\}\}f\(\\Delta\_\{\\overline\{\\partial\}\}f\)\\frac\{\\omega^\{K\}\}\{K\!\},\(11\.6\)and so the Dirichlet energy can be written as
E\(f\)=−1Volω\(𝒮\)∫𝒮f\(Δ∂¯f\)ωKK\!\+1Volω\(𝒮\)∮∂𝒮if∂¯f∧ωK−1\(K−1\)\!≤∥f∥L2∥Δ∂¯f∥L2\+\|ℬ∂𝒮\(f\)\|\.\\displaystyle E\(f\)=\-\\frac\{1\}\{\\text\{Vol\}\_\{\\omega\}\(\\mathcal\{S\}\)\}\\int\_\{\\mathcal\{S\}\}f\(\\Delta\_\{\\overline\{\\partial\}\}f\)\\frac\{\\omega^\{K\}\}\{K\!\}\+\\frac\{1\}\{\\text\{Vol\}\_\{\\omega\}\(\\mathcal\{S\}\)\}\\oint\_\{\\partial\\mathcal\{S\}\}if\\overline\{\\partial\}f\\wedge\\frac\{\\omega^\{K\-1\}\}\{\(K\-1\)\!\}\\leq\\\|f\\\|\_\{L^\{2\}\}\\\|\\Delta\_\{\\overline\{\\partial\}\}f\\\|\_\{L^\{2\}\}\+\|\\mathcal\{B\}\_\{\\partial\\mathcal\{S\}\}\(f\)\|\.\(11\.7\)We can note‖f‖L2\\\|f\\\|\_\{L^\{2\}\}is𝒪\(1\)\\mathcal\{O\}\(1\)since
‖f‖L2\(𝒮\)=\(1Volω\(𝒮\)∫𝒮\|f\(z\)\|2ωKK\!\)1/2≤𝒪\(1\)Volω\(𝒮\)Volω\(𝒮\)=𝒪\(1\)\.\\displaystyle\\\|f\\\|\_\{L^\{2\}\(\\mathcal\{S\}\)\}=\\left\(\\frac\{1\}\{\\text\{Vol\}\_\{\\omega\}\(\\mathcal\{S\}\)\}\\int\_\{\\mathcal\{S\}\}\|f\(z\)\|^\{2\}\\frac\{\\omega^\{K\}\}\{K\!\}\\right\)^\{1/2\}\\leq\\mathcal\{O\}\(1\)\\cancel\{\\frac\{\\sqrt\{\\text\{Vol\}\_\{\\omega\}\(\\mathcal\{S\}\)\}\}\{\\sqrt\{\\text\{Vol\}\_\{\\omega\}\(\\mathcal\{S\}\)\}\}\}=\\mathcal\{O\}\(1\)\.\(11\.8\)We previously established
supθ∈𝒮‖i∂∂¯f‖2=𝒪\(1m\)\\displaystyle\\sup\_\{\\theta\\in\\mathcal\{\\mathcal\{S\}\}\}\\\|i\\partial\\overline\{\\partial\}f\\\|\_\{2\}=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)\(11\.9\)for sufficient ball𝒮\\mathcal\{\\mathcal\{S\}\}, but the norms here differ, and so the above has greater connections to the Laplacian term of[11\.7](https://arxiv.org/html/2608.19584#S11.E7)\. Note the identityΔ∂¯f=Trω\(i∂∂¯f\)\\Delta\_\{\\overline\{\\partial\}\}f=\\text\{Tr\}\_\{\\omega\}\(i\\partial\\overline\{\\partial\}f\)\. Therefore, it follows
‖Δ∂¯f‖L2\(𝒮\)≤\(1Volω\(𝒮\)∫𝒮𝒪\(K2m\)ωKK\!\)1/2=𝒪\(Km\)\.\\displaystyle\\\|\\Delta\_\{\\overline\{\\partial\}\}f\\\|\_\{L^\{2\}\(\\mathcal\{S\}\)\}\\leq\\left\(\\frac\{1\}\{\\text\{Vol\}\_\{\\omega\}\(\\mathcal\{S\}\)\}\\int\_\{\\mathcal\{S\}\}\\mathcal\{O\}\\left\(\\frac\{K^\{2\}\}\{m\}\\right\)\\frac\{\\omega^\{K\}\}\{K\!\}\\right\)^\{1/2\}=\\mathcal\{O\}\\left\(\\frac\{K\}\{\\sqrt\{m\}\}\\right\)\.\(11\.10\)Turning to the boundary term, we can restrict theL2L^\{2\}norm over a subset of the domain\. Moreover, we can bound
\|ℬ∂𝒮\(f\)\|≤Volω\(∂𝒮\)Volω\(𝒮\)⏟=𝒪\(K\)\(supθ∈∂𝒮\|f\(θ\)\|\)⏟=𝒪\(1\)\(supθ∈∂𝒮‖∂¯f\(θ\)‖ω\)⏟=𝒪\(1\)=𝒪\(K\)\.\\displaystyle\|\\mathcal\{B\}\_\{\\partial\\mathcal\{\\mathcal\{S\}\}\}\(f\)\|\\leq\\underbrace\{\\frac\{\\text\{Vol\}\_\{\\omega\}\(\\partial\\mathcal\{\\mathcal\{S\}\}\)\}\{\\text\{Vol\}\_\{\\omega\}\(\\mathcal\{\\mathcal\{S\}\}\)\}\}\_\{=\\mathcal\{O\}\(K\)\}\\underbrace\{\\left\(\\sup\_\{\\theta\\in\\partial\\mathcal\{\\mathcal\{S\}\}\}\|f\(\\theta\)\|\\right\)\}\_\{=\\mathcal\{O\}\(1\)\}\\underbrace\{\\left\(\\sup\_\{\\theta\\in\\partial\\mathcal\{\\mathcal\{S\}\}\}\\\|\\overline\{\\partial\}f\(\\theta\)\\\|\_\{\\omega\}\\right\)\}\_\{=\\mathcal\{O\}\(1\)\}=\\mathcal\{O\}\(K\)\.\(11\.11\)We have used theω\\omeganorm‖∂¯f‖ω2=hjk¯\(∂¯f\)k¯\(∂¯f\)j¯¯\\\|\\overline\{\\partial\}f\\\|\_\{\\omega\}^\{2\}=h^\{j\\overline\{k\}\}\(\\overline\{\\partial\}f\)\_\{\\overline\{k\}\}\\overline\{\(\\overline\{\\partial\}f\)\_\{\\overline\{j\}\}\}\. Therefore, around initialization,
E\(f\)≤𝒪\(1\)×𝒪\(Km\)⏟‖Δ∂¯f‖L2\(𝒮\)\+𝒪\(K\)=𝒪\(K\+Km\)\.\\displaystyle E\(f\)\\leq\\mathcal\{O\}\(1\)\\times\\underbrace\{\\mathcal\{O\}\(\\frac\{K\}\{\\sqrt\{m\}\}\)\}\_\{\\\|\\Delta\_\{\\overline\{\\partial\}\}f\\\|\_\{L^\{2\}\(\\mathcal\{S\}\)\}\}\+\\mathcal\{O\}\(K\)=\\mathcal\{O\}\(K\+\\frac\{K\}\{\\sqrt\{m\}\}\)\.\(11\.12\)The Dirichlet energy is across a single data point, therefore summing acrossNNdata points
∑α=1NE\(fα\)≤∑α=1N\(𝒪\(1\)×𝒪\(Km\)⏟‖Δ∂¯fα‖L2\(𝒮\)\+𝒪\(K\)\)=𝒪\(NK\+NKm\)\.\\displaystyle\\sum\_\{\\alpha=1\}^\{N\}E\(f\_\{\\alpha\}\)\\leq\\sum\_\{\\alpha=1\}^\{N\}\\left\(\\mathcal\{O\}\(1\)\\times\\underbrace\{\\mathcal\{O\}\\left\(\\frac\{K\}\{\\sqrt\{m\}\}\\right\)\}\_\{\\\|\\Delta\_\{\\overline\{\\partial\}\}f\_\{\\alpha\}\\\|\_\{L^\{2\}\(\\mathcal\{S\}\)\}\}\+\\mathcal\{O\}\(K\)\\right\)=\\mathcal\{O\}\(NK\+\\frac\{NK\}\{\\sqrt\{m\}\}\)\.\(11\.13\)
For a lower bound, define the unregularized \(empirical\) Fisher metricℱ=∑α=1N∂fα∧∂¯fα\\mathcal\{F\}=\\sum\_\{\\alpha=1\}^\{N\}\\partial f\_\{\\alpha\}\\wedge\\overline\{\\partial\}f\_\{\\alpha\}\. We instilled a Fisher eigenvalue conditionℱreg=ℱ\+λI,λmin\(ℱreg\)≥μ\>0\\mathcal\{F\}\_\{\\text\{reg\}\}=\\mathcal\{F\}\+\\lambda I,\\lambda\_\{\\text\{min\}\}\(\\mathcal\{F\}\_\{\\text\{reg\}\}\)\\geq\\mu\>0\. Now, the trace of the unregularized Fisher metric with respect toω\\omegais the sum of the squared gradient norms, which coincides with the trace of the Neural Tangent Kernel \(NTK\)
Tr\(ℱ\)=∑α=1N‖∂f\(zα\)‖ω2=Tr\(KNTK\)\\displaystyle\\text\{Tr\}\(\\mathcal\{F\}\)=\\sum\_\{\\alpha=1\}^\{N\}\\\|\\partial f\(z\_\{\\alpha\}\)\\\|\_\{\\omega\}^\{2\}=\\text\{Tr\}\(K\_\{\\text\{NTK\}\}\)\(11\.14\)By taking the trace of our regularized Fisher metric, we get
Tr\(ℱreg\)=Tr\(ℱ\)\+Tr\(λI\)=Tr\(KNTK\)\+λK\.\\displaystyle\\text\{Tr\}\(\\mathcal\{F\}\_\{\\text\{reg\}\}\)=\\text\{Tr\}\(\\mathcal\{F\}\)\+\\text\{Tr\}\(\\lambda I\)=\\text\{Tr\}\(K\_\{\\text\{NTK\}\}\)\+\\lambda K\.\(11\.15\)Because we assumedλmin\(ℱreg\)≥μ\\lambda\_\{\\text\{min\}\}\(\\mathcal\{F\}\_\{\\text\{reg\}\}\)\\geq\\mu, the trace is bounded below by the sum of its minimal eigenvalues across allKKdimensions
Tr\(ℱreg\)≥μK\.\\displaystyle\\text\{Tr\}\(\\mathcal\{F\}\_\{\\text\{reg\}\}\)\\geq\\mu K\.\(11\.16\)Isolating the trace of the NTK, using[11\.15](https://arxiv.org/html/2608.19584#S11.E15), we get
Tr\(KNTK\)=∑α=1N‖∂f\(zα\)‖ω2≥K\(μ−λ\)\.\\displaystyle\\text\{Tr\}\(K\_\{\\text\{NTK\}\}\)=\\sum\_\{\\alpha=1\}^\{N\}\\\|\\partial f\(z\_\{\\alpha\}\)\\\|\_\{\\omega\}^\{2\}\\geq K\(\\mu\-\\lambda\)\.\(11\.17\)Our total Dirichlet energy of the network over the dataset is the integral of this NTK trace over the set𝒮\\mathcal\{S\}
∑α=1NE\(f\(zα\)\)=1Volω\(𝒮\)∫𝒮Tr\(KNTK\)ωKK\!≥1Volω\(𝒮\)∫𝒮K\(μ−λ\)ωKK\!\.\\displaystyle\\sum\_\{\\alpha=1\}^\{N\}E\(f\(z\_\{\\alpha\}\)\)=\\frac\{1\}\{\\text\{Vol\}\_\{\\omega\}\(\\mathcal\{S\}\)\}\\int\_\{\\mathcal\{S\}\}\\text\{Tr\}\(K\_\{\\text\{NTK\}\}\)\\frac\{\\omega^\{K\}\}\{K\!\}\\geq\\frac\{1\}\{\\text\{Vol\}\_\{\\omega\}\(\\mathcal\{S\}\)\}\\int\_\{\\mathcal\{S\}\}K\(\\mu\-\\lambda\)\\frac\{\\omega^\{K\}\}\{K\!\}\.\(11\.18\)Integrating this constant bound yields
∑α=1NE\(f\(zα\)\)≥K\(μ−λ\)\.\\displaystyle\\sum\_\{\\alpha=1\}^\{N\}E\(f\(z\_\{\\alpha\}\)\)\\geq K\(\\mu\-\\lambda\)\.\(11\.19\)This formulation maintains consistency with the upper bound\. The unregularized Fisherℱ\\mathcal\{F\}is constructed as a sum overNNouter products\. Because we are summingNNterms, the trace scales with the dataset size\. To maintain the validity ofℱ\+λI⪰μI\\mathcal\{F\}\+\\lambda I\\succeq\\mu IasNNgrows, the gap\(μ−λ\)\(\\mu\-\\lambda\)must scale as𝒪\(N/K\)\\mathcal\{O\}\(N/K\)\. Thus, theμ\\muparameter contains a dependence onNN\.
□\\square
### 11\.2Trajectory bounds
In this section, we analyze the asymptotic regime of the descent path of the parameter\. Let us begin by first proving a claim\.
Claim\.We have∫0t𝔼\[‖θ˙s‖h\]ds=𝒪\(1m\)\\int\_\{0\}^\{t\}\\mathbb\{E\}\[\\\|\\dot\{\\theta\}\_\{s\}\\\|\_\{h\}\]ds=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)near initialization\.
Proof\.Under natural gradient descent, we haveθ˙s=−h−1∇ℒ\(θs\)\\dot\{\\theta\}\_\{s\}=\-h^\{\-1\}\\nabla\\mathcal\{L\}\(\\theta\_\{s\}\)\. From discussion in earlier sections such as[3](https://arxiv.org/html/2608.19584#S3)and[9](https://arxiv.org/html/2608.19584#S9), we definef\(θ\)=1mv†α\(L\)f\(\\theta\)=\\frac\{1\}\{\\sqrt\{m\}\}v^\{\\dagger\}\\alpha^\{\(L\)\}and we note‖δ\(l\)‖∞=𝒪\(1m\)\\\|\\delta^\{\(l\)\}\\\|\_\{\\infty\}=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)\. We have‖∇f\(θ\)‖2=𝒪\(1m\)\\\|\\nabla f\(\\theta\)\\\|\_\{2\}=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)\. Since under quadratic loss∇ℒ\(θ\)=\(f\(θ\)−y\)∇f\(θ\)\\nabla\\mathcal\{L\}\(\\theta\)=\(f\(\\theta\)\-y\)\\nabla f\(\\theta\), we get‖∇ℒ\(θ0\)‖2=𝒪\(1m\)\\\|\\nabla\\mathcal\{L\}\(\\theta\_\{0\}\)\\\|\_\{2\}=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)\. Now, we can note the norm onθ˙s\\dot\{\\theta\}\_\{s\}obeys
‖θ˙s‖h=\(∇ℒ\(θs\)\)†h−1\(∇ℒ\(θs\)\)\.\\displaystyle\\\|\\dot\{\\theta\}\_\{s\}\\\|\_\{h\}=\\sqrt\{\(\\nabla\\mathcal\{L\}\(\\theta\_\{s\}\)\)^\{\\dagger\}h^\{\-1\}\(\\nabla\\mathcal\{L\}\(\\theta\_\{s\}\)\)\}\.\(11\.20\)Under the regularized loss assumption of[4\.23](https://arxiv.org/html/2608.19584#S4.E23)hreg=h\+λIh\_\{\\text\{reg\}\}=h\+\\lambda I, the minimum eigenvalue is bounded away from zero,λmin\(h\)≥μ\>0\\lambda\_\{min\}\(h\)\\geq\\mu\>0\. Therefore, the spectral norm of the inverse metric follows‖h−1‖2≤1μ=𝒪\(1\)\\\|h^\{\-1\}\\\|\_\{2\}\\leq\\frac\{1\}\{\\mu\}=\\mathcal\{O\}\(1\)\. By Cauchy\-Schwarz,
‖θ˙0‖h≤‖h−1‖2‖∇ℒ\(θ0\)‖2=𝒪\(1m\)\.\\displaystyle\\\|\\dot\{\\theta\}\_\{0\}\\\|\_\{h\}\\leq\\sqrt\{\\\|h^\{\-1\}\\\|\_\{2\}\}\\\|\\nabla\\mathcal\{L\}\(\\theta\_\{0\}\)\\\|\_\{2\}=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)\.\(11\.21\)Via Fundamental Theorem of Calculus \(we have sufficient smoothness\), and applying an inequality,
‖∇hℒ\(θs\)‖h≤‖∇hℒ\(θ0\)‖h\+∫0s‖∇h2ℒ\(θτ\)‖h‖θ˙τ‖h𝑑τ\.\\displaystyle\\\|\\nabla\_\{h\}\\mathcal\{L\}\(\\theta\_\{s\}\)\\\|\_\{h\}\\leq\\\|\\nabla\_\{h\}\\mathcal\{L\}\(\\theta\_\{0\}\)\\\|\_\{h\}\+\\int\_\{0\}^\{s\}\\\|\\nabla\_\{h\}^\{2\}\\mathcal\{L\}\(\\theta\_\{\\tau\}\)\\\|\_\{h\}\\\|\\dot\{\\theta\}\_\{\\tau\}\\\|\_\{h\}d\\tau\.\(11\.22\)Sinceθ˙s=−h−1∇ℒ\(θs\)\\dot\{\\theta\}\_\{s\}=\-h^\{\-1\}\\nabla\\mathcal\{L\}\(\\theta\_\{s\}\), we can relate[11\.22](https://arxiv.org/html/2608.19584#S11.E22)to this and we get
‖θ˙s‖h≤‖θ˙0‖h\+𝒪\(1m\)∫0s‖θ˙τ‖h𝑑τ,\\displaystyle\\\|\\dot\{\\theta\}\_\{s\}\\\|\_\{h\}\\leq\\\|\\dot\{\\theta\}\_\{0\}\\\|\_\{h\}\+\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)\\int\_\{0\}^\{s\}\\\|\\dot\{\\theta\}\_\{\\tau\}\\\|\_\{h\}d\\tau,\(11\.23\)again noting the inverse metric is𝒪\(1\)\\mathcal\{O\}\(1\)\. This line is nontrivial because[11\.22](https://arxiv.org/html/2608.19584#S11.E22)is with a different norm than the spectral norm, which our original result used; however, this is salvageable and correct because the spectral norm and thehh\-norm can be related via a maximum eigenvalue\.λmax\\lambda\_\{\\text\{max\}\}is𝒪\(1\)\\mathcal\{O\}\(1\), and so is the condition number\. Therefore‖∇2ℒ‖h≤𝒪\(1\)‖∇2ℒ‖2=𝒪\(1\)⋅𝒪\(1m\)=𝒪\(1m\)\\\|\\nabla^\{2\}\\mathcal\{L\}\\\|\_\{h\}\\leq\\mathcal\{O\}\(1\)\\\|\\nabla^\{2\}\\mathcal\{L\}\\\|\_\{2\}=\\mathcal\{O\}\(1\)\\cdot\\mathcal\{O\}\\left\(\\frac\{1\}\{\\sqrt\{m\}\}\\right\)=\\mathcal\{O\}\\left\(\\frac\{1\}\{\\sqrt\{m\}\}\\right\)\. Moreover, we have noted the Hessian results of[8\.5](https://arxiv.org/html/2608.19584#S8.SS5),[8\.2](https://arxiv.org/html/2608.19584#S8.SS2)\. Via Grönwall’s inequality,‖θ˙s‖h≤‖θ˙0‖hexp\(𝒪\(1m\)s\)\\\|\\dot\{\\theta\}\_\{s\}\\\|\_\{h\}\\leq\\\|\\dot\{\\theta\}\_\{0\}\\\|\_\{h\}\\exp\\left\(\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)s\\right\)\. Therefore, we conclude
∫0t𝔼\[‖θ˙s‖h\]𝑑s≤∫0t𝔼\[𝒪\(1m\)\]𝑑s=𝒪\(1m\),\\displaystyle\\int\_\{0\}^\{t\}\\mathbb\{E\}\[\\\|\\dot\{\\theta\}\_\{s\}\\\|\_\{h\}\]ds\\leq\\int\_\{0\}^\{t\}\\mathbb\{E\}\\left\[\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)\\right\]ds=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\),\(11\.24\)since‖θ˙0‖h=𝒪\(1m\)\\\|\\dot\{\\theta\}\_\{0\}\\\|\_\{h\}=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)\.
□\\square

Figure 10:We plot the asymptotics of∫𝔼‖θ˙‖h𝑑s\\int\\mathbb\{E\}\\\|\\dot\{\\theta\}\\\|\_\{h\}dsnear initialization across10⋅\(number of widths\)CLOSE10\\cdot\(\\text\{number of widths\)\}distinct calculations\. We use batch size 64, learning rate 0\.01, and 15 steps\. The line is the1m\\frac\{1\}\{\\sqrt\{m\}\}asymptotic line\. The neural network is consistent with[3](https://arxiv.org/html/2608.19584#S3)\.
Figure 11:We plot the asymptotics of∫𝔼‖θ˙‖h𝑑s\\int\\mathbb\{E\}\\\|\\dot\{\\theta\}\\\|\_\{h\}dsnear initialization across10⋅\(number of widths\)CLOSE10\\cdot\(\\text\{number of widths\)\}distinct calculations\. We use batch size 64, learning rate 0\.5, and 50 steps, meaning we deviate from initialization much further than[11](https://arxiv.org/html/2608.19584#S11.F11)\. The line is consistent with Figure[11](https://arxiv.org/html/2608.19584#S11.F11), and the network consistent with[3](https://arxiv.org/html/2608.19584#S3)\.Proof of Lemma 10\.Under natural gradient descent,θ˙i=−hij¯∂j¯ℒ\\dot\{\\theta\}^\{i\}=\-h^\{i\\overline\{j\}\}\\partial\_\{\\overline\{j\}\}\\mathcal\{L\}\. We define the gradient normv\(t\)=‖∇hℒ‖h2v\(t\)=\\\|\\nabla\_\{h\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}\. Differentiating yields
ddtv\(t\)=−2Re\(∇hℒT∇ω2,0ℒ∇hℒ\)−2∇hℒ†∇ω1,1ℒ∇hℒ\.\\displaystyle\\frac\{d\}\{dt\}v\(t\)=\-2\\text\{Re\}\(\\nabla\_\{h\}\\mathcal\{L\}^\{T\}\\nabla^\{2,0\}\_\{\\omega\}\\mathcal\{L\}\\nabla\_\{h\}\\mathcal\{L\}\)\-2\\nabla\_\{h\}\\mathcal\{L\}^\{\\dagger\}\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\\nabla\_\{h\}\\mathcal\{L\}\.\(11\.25\)We define uniform bounds overΩ\\Omegaat timett
H\(t\)=supθ∈Ω‖∇ω2,0ℒ‖handμ\(t\)=infθ∈Ωλmin\(∇ω1,1ℒ\)\.\\displaystyle H\(t\)=\\sup\_\{\\theta\\in\\Omega\}\\\|\\nabla^\{2,0\}\_\{\\omega\}\\mathcal\{L\}\\\|\_\{h\}\\quad\\text\{and\}\\quad\\mu\(t\)=\\inf\_\{\\theta\\in\\Omega\}\\lambda\_\{\\text\{min\}\}\(\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\)\.\(11\.26\)Applying Cauchy\-Schwarz to the\(2,0\)\(2,0\)term and taking the expectation𝔼ρt\[⋅\]\\mathbb\{E\}\_\{\\rho\_\{t\}\}\[\\cdot\]overΩ\\Omega, which we denoteVV, we obtain the inequality using[11\.25](https://arxiv.org/html/2608.19584#S11.E25)
ddtV\(t\)≤2H\(t\)V\(t\)−2μ\(t\)V\(t\)\.\\displaystyle\\frac\{d\}\{dt\}V\(t\)\\leq 2H\(t\)V\(t\)\-2\\mu\(t\)V\(t\)\.\(11\.27\)By Grönwall’s inequality, we bound the trajectory
V\(t\)≤V\(0\)exp\(2∫0tH\(s\)𝑑s−2∫0tμ\(s\)𝑑s\)\.\\displaystyle V\(t\)\\leq V\(0\)\\exp\\left\(2\\int\_\{0\}^\{t\}H\(s\)ds\-2\\int\_\{0\}^\{t\}\\mu\(s\)ds\\right\)\.\(11\.28\)We are given the initialization bounds att=0t=0using the results of[8\.2](https://arxiv.org/html/2608.19584#S8.SS2)
H\(0\)=𝒪\(1m\)andμ\(0\)=𝒪\(1m\)\.\\displaystyle H\(0\)=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)\\quad\\text\{and\}\\quad\\mu\(0\)=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)\.\(11\.29\)We assume the Hessians are locally Lipschitz with respect to the Kähler metric connection with constantLL\. We can note via Fundamental Theorem of Calculus \(take absolute continuity\)
‖∇ω2,0ℒ\(θt\)‖h−‖∇ω2,0ℒ\(θt\)\|t=0‖h=∫0tdds‖∇ω2,0ℒ\(θs\)‖h𝑑s\\displaystyle\\\|\\nabla^\{2,0\}\_\{\\omega\}\\mathcal\{L\}\(\\theta\_\{t\}\)\\\|\_\{h\}\-\\\|\\nabla^\{2,0\}\_\{\\omega\}\\mathcal\{L\}\(\\theta\_\{t\}\)\|\_\{t=0\}\\\|\_\{h\}=\\int\_\{0\}^\{t\}\\frac\{d\}\{ds\}\\\|\\nabla^\{2,0\}\_\{\\omega\}\\mathcal\{L\}\(\\theta\_\{s\}\)\\\|\_\{h\}ds\(11\.30\)and moreover via chain rule
dds‖∇ω2,0ℒ\(θt\)‖h≤‖∇h\(∇ω2,0ℒ\(θs\)\)‖h‖θ˙s‖h≤L‖θ˙s‖h,\\displaystyle\\frac\{d\}\{ds\}\\\|\\nabla^\{2,0\}\_\{\\omega\}\\mathcal\{L\}\(\\theta\_\{t\}\)\\\|\_\{h\}\\leq\\\|\\nabla\_\{h\}\(\\nabla^\{2,0\}\_\{\\omega\}\\mathcal\{L\}\(\\theta\_\{s\}\)\)\\\|\_\{h\}\\\|\\dot\{\\theta\}\_\{s\}\\\|\_\{h\}\\leq L\\\|\\dot\{\\theta\}\_\{s\}\\\|\_\{h\},\(11\.31\)so under a Lipschitz condition, we see
H\(t\)≤H\(0\)\+L∫0t𝔼\[‖θ˙s‖h\]𝑑s\.\\displaystyle H\(t\)\\leq H\(0\)\+L\\int\_\{0\}^\{t\}\\mathbb\{E\}\[\\\|\\dot\{\\theta\}\_\{s\}\\\|\_\{h\}\]ds\.\(11\.32\)Similarly,
μ\(t\)≥μ\(0\)−L∫0t𝔼\[‖θ˙s‖h\]𝑑s\.\\displaystyle\\mu\(t\)\\geq\\mu\(0\)\-L\\int\_\{0\}^\{t\}\\mathbb\{E\}\[\\\|\\dot\{\\theta\}\_\{s\}\\\|\_\{h\}\]ds\.\(11\.33\)Since∫0t𝔼\[‖θ˙s‖h\]𝑑s=𝒪\(1m\)\\int\_\{0\}^\{t\}\\mathbb\{E\}\[\\\|\\dot\{\\theta\}\_\{s\}\\\|\_\{h\}\]ds=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\), we can substitute this into our Lipschitz bounds, so
H\(t\)=𝒪\(1m\)\+L⋅𝒪\(1m\)=𝒪\(1m\),\\displaystyle H\(t\)=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)\+L\\cdot\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\),\(11\.34\)and moreover,
μ\(t\)=𝒪\(1m\)−L⋅𝒪\(1m\)=𝒪\(1m\)\.\\displaystyle\\mu\(t\)=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)\-L\\cdot\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)\.\(11\.35\)We return to the Grönwall bound established previously\. Using what we found
V\(t\)\\displaystyle V\(t\)≤V\(0\)exp\(2∫0t𝒪\(1m\)𝑑s−2∫0t𝒪\(1m\)𝑑s\)\\displaystyle\\leq V\(0\)\\exp\\left\(2\\int\_\{0\}^\{t\}\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)ds\-2\\int\_\{0\}^\{t\}\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)ds\\right\)\(11\.36\)≤V\(0\)exp\(2t⋅𝒪\(1m\)\)\.\\displaystyle\\leq V\(0\)\\exp\\left\(2t\\cdot\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)\\right\)\.\(11\.37\)The constants do not cancel, which holds almost surely\.
To guarantee convergence for finitemm, we bound the integral ofμ\(s\)\\mu\(s\)via a lower bound\. This showsμ\\muand its integral are sufficiently large and therefore the exponential with this negative integral in its exponent are dissipative, giving a convergence result\. Let us establish an eigenvalue dissipation rate geometrically\. LetMij¯=∇ij¯ℒM\_\{i\\overline\{j\}\}=\\nabla\_\{i\\overline\{j\}\}\\mathcal\{L\}denote the\(1,1\)\(1,1\)Hessian, and letNij=∇ijℒN\_\{ij\}=\\nabla\_\{ij\}\\mathcal\{L\}denote the\(2,0\)\(2,0\)Hessian\. Let us consider the natural gradient flowθ˙i=−hij¯∂j¯ℒ=−∇iℒ\\dot\{\\theta\}^\{i\}=\-h^\{i\\overline\{j\}\}\\partial\_\{\\overline\{j\}\}\\mathcal\{L\}=\-\\nabla^\{i\}\\mathcal\{L\}as in[4\.1](https://arxiv.org/html/2608.19584#S4.SS1)\. Since the parameter is a function of time, we can perform the chain rule
∂tℒ=ddtℒ\(θ\(t\),θ¯\(t\)\)=∑p∂ℒ∂θpθ˙p\+∑q∂ℒ∂θ¯qθ¯˙q\.\\displaystyle\\partial\_\{t\}\\mathcal\{L\}=\\frac\{d\}\{dt\}\\mathcal\{L\}\(\\theta\(t\),\\overline\{\\theta\}\(t\)\)=\\sum\_\{p\}\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\theta^\{p\}\}\\dot\{\\theta\}^\{p\}\+\\sum\_\{q\}\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\overline\{\\theta\}^\{q\}\}\\dot\{\\overline\{\\theta\}\}^\{q\}\.\(11\.38\)We compute the material derivative
DdtMij¯=θ˙p∇pMij¯\+θ¯˙q∇q¯Mij¯\.\\displaystyle\\frac\{D\}\{dt\}M\_\{i\\overline\{j\}\}=\\dot\{\\theta\}^\{p\}\\nabla\_\{p\}M\_\{i\\overline\{j\}\}\+\\dot\{\\overline\{\\theta\}\}^\{q\}\\nabla\_\{\\overline\{q\}\}M\_\{i\\overline\{j\}\}\.\(11\.39\)Substituting our gradient flow,
DdtMij¯=−∇pℒ∇p\(∇ij¯ℒ\)−∇q¯ℒ∇q¯\(∇ij¯ℒ\)\.\\displaystyle\\frac\{D\}\{dt\}M\_\{i\\overline\{j\}\}=\-\\nabla^\{p\}\\mathcal\{L\}\\nabla\_\{p\}\(\\nabla\_\{i\\overline\{j\}\}\\mathcal\{L\}\)\-\\nabla^\{\\overline\{q\}\}\\mathcal\{L\}\\nabla\_\{\\overline\{q\}\}\(\\nabla\_\{i\\overline\{j\}\}\\mathcal\{L\}\)\.\(11\.40\)Since the metric is Kähler,∇ij¯pℒ=∇pij¯ℒ\\nabla\_\{i\\overline\{j\}p\}\\mathcal\{L\}=\\nabla\_\{pi\\overline\{j\}\}\\mathcal\{L\}, using notation∇ij¯pℒ=∇i∇j¯∇pℒ\\nabla\_\{i\\overline\{j\}p\}\\mathcal\{L\}=\\nabla\_\{i\}\\nabla\_\{\\overline\{j\}\}\\nabla\_\{p\}\\mathcal\{L\}\. For the anti\-holomorphic commutativity, we can note using Riemannian curvature
∇q¯ij¯ℒ=∇ij¯q¯ℒ−Rij¯pq¯∇pℒ\.\\displaystyle\\nabla\_\{\\overline\{q\}i\\overline\{j\}\}\\mathcal\{L\}=\\nabla\_\{i\\overline\{j\}\\overline\{q\}\}\\mathcal\{L\}\-R\_\{i\\overline\{j\}p\\overline\{q\}\}\\nabla^\{p\}\\mathcal\{L\}\.\(11\.41\)Substituting back in, the material derivative obeys
DdtMij¯\\displaystyle\\frac\{D\}\{dt\}M\_\{i\\overline\{j\}\}=−∇pℒ∇ij¯pℒ−∇q¯ℒ\(∇ij¯q¯ℒ−Rij¯pq¯∇pℒ\)\\displaystyle=\-\\nabla^\{p\}\\mathcal\{L\}\\nabla\_\{i\\overline\{j\}p\}\\mathcal\{L\}\-\\nabla^\{\\overline\{q\}\}\\mathcal\{L\}\(\\nabla\_\{i\\overline\{j\}\\overline\{q\}\}\\mathcal\{L\}\-R\_\{i\\overline\{j\}p\\overline\{q\}\}\\nabla^\{p\}\\mathcal\{L\}\)\(11\.42\)=−\(∇pℒ∇ij¯pℒ\+∇q¯ℒ∇ij¯q¯ℒ\)\+Rij¯pq¯∇pℒ∇q¯ℒ\.\\displaystyle=\-\(\\nabla^\{p\}\\mathcal\{L\}\\nabla\_\{i\\overline\{j\}p\}\\mathcal\{L\}\+\\nabla^\{\\overline\{q\}\}\\mathcal\{L\}\\nabla\_\{i\\overline\{j\}\\overline\{q\}\}\\mathcal\{L\}\)\+R\_\{i\\overline\{j\}p\\overline\{q\}\}\\nabla^\{p\}\\mathcal\{L\}\\nabla^\{\\overline\{q\}\}\\mathcal\{L\}\.\(11\.43\)Now, we analyze the spatial Hessian of the squared gradient norm,F=‖∇ℒ‖h2=hpq¯∇pℒ∇q¯ℒF=\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{2\}=h^\{p\\overline\{q\}\}\\nabla\_\{p\}\\mathcal\{L\}\\nabla\_\{\\overline\{q\}\}\\mathcal\{L\}\. Applying the covariant derivative∇ij¯\\nabla\_\{i\\overline\{j\}\}and the product rule generates four terms
∇ij¯\(‖∇ℒ‖2\)=hpq¯\(∇ij¯pℒ\)∇q¯ℒ\+hpq¯∇pℒ\(∇ij¯q¯ℒ\)\+hpq¯∇ipℒ∇j¯q¯ℒ\+hpq¯∇j¯pℒ∇iq¯ℒ\.\\displaystyle\\nabla\_\{i\\overline\{j\}\}\(\\\|\\nabla\\mathcal\{L\}\\\|^\{2\}\)=h^\{p\\overline\{q\}\}\(\\nabla\_\{i\\overline\{j\}p\}\\mathcal\{L\}\)\\nabla\_\{\\overline\{q\}\}\\mathcal\{L\}\+h^\{p\\overline\{q\}\}\\nabla\_\{p\}\\mathcal\{L\}\(\\nabla\_\{i\\overline\{j\}\\overline\{q\}\}\\mathcal\{L\}\)\+h^\{p\\overline\{q\}\}\\nabla\_\{ip\}\\mathcal\{L\}\\nabla\_\{\\overline\{j\}\\overline\{q\}\}\\mathcal\{L\}\+h^\{p\\overline\{q\}\}\\nabla\_\{\\overline\{j\}p\}\\mathcal\{L\}\\nabla\_\{i\\overline\{q\}\}\\mathcal\{L\}\.\(11\.44\)We can rewrite using our definitions
∇ij¯\(‖∇ℒ‖2\)=∇pℒ∇ij¯pℒ\+∇q¯ℒ∇ij¯q¯ℒ\+NipNj¯p\+Mik¯Mj¯k¯\.\\displaystyle\\nabla\_\{i\\overline\{j\}\}\(\\\|\\nabla\\mathcal\{L\}\\\|^\{2\}\)=\\nabla^\{p\}\\mathcal\{L\}\\nabla\_\{i\\overline\{j\}p\}\\mathcal\{L\}\+\\nabla^\{\\overline\{q\}\}\\mathcal\{L\}\\nabla\_\{i\\overline\{j\}\\overline\{q\}\}\\mathcal\{L\}\+N\_\{ip\}N^\{p\}\_\{\\overline\{j\}\}\+M\_\{i\\overline\{k\}\}M^\{\\overline\{k\}\}\_\{\\overline\{j\}\}\.\(11\.45\)Rearranging,
−\(∇pℒ∇ij¯pℒ\+∇q¯ℒ∇ij¯q¯ℒ\)=−∇ij¯\(‖∇ℒ‖2\)\+Mik¯Mj¯k¯\+NipNj¯p\.\\displaystyle\-\(\\nabla^\{p\}\\mathcal\{L\}\\nabla\_\{i\\overline\{j\}p\}\\mathcal\{L\}\+\\nabla^\{\\overline\{q\}\}\\mathcal\{L\}\\nabla\_\{i\\overline\{j\}\\overline\{q\}\}\\mathcal\{L\}\)=\-\\nabla\_\{i\\overline\{j\}\}\(\\\|\\nabla\\mathcal\{L\}\\\|^\{2\}\)\+M\_\{i\\overline\{k\}\}M^\{\\overline\{k\}\}\_\{\\overline\{j\}\}\+N\_\{ip\}N^\{p\}\_\{\\overline\{j\}\}\.\(11\.46\)Substituting back into[11\.42](https://arxiv.org/html/2608.19584#S11.E42),
DdtMij¯=Mik¯Mj¯k¯\+NipNj¯p\+Rij¯pq¯∇pℒ∇q¯ℒ−∇ij¯\(‖∇ℒ‖2\)\.\\displaystyle\\frac\{D\}\{dt\}M\_\{i\\overline\{j\}\}=M\_\{i\\overline\{k\}\}M^\{\\overline\{k\}\}\_\{\\overline\{j\}\}\+N\_\{ip\}N^\{p\}\_\{\\overline\{j\}\}\+R\_\{i\\overline\{j\}p\\overline\{q\}\}\\nabla^\{p\}\\mathcal\{L\}\\nabla^\{\\overline\{q\}\}\\mathcal\{L\}\-\\nabla\_\{i\\overline\{j\}\}\(\\\|\\nabla\\mathcal\{L\}\\\|^\{2\}\)\.\(11\.47\)For normalized eigenvectorXX, we have the minimum eigenvalue equation of the Hessian \(in a geometric sense, so scaled by the metric\) can be calculated via the standard eigenvalue equation
Mij¯Xi=μ\(t\)hij¯Xi\.\\displaystyle M\_\{i\\overline\{j\}\}X^\{i\}=\\mu\(t\)h\_\{i\\overline\{j\}\}X^\{i\}\.\(11\.48\)We can contract withXjX^\{j\}and under a normalized, unit eigenvector propertyhij¯XiXj¯=1h\_\{i\\overline\{j\}\}X^\{i\}X^\{\\overline\{j\}\}=1, we get the eigenvalue follows under these two equations
μ\(t\)=Mij¯XiXj¯\.\\displaystyle\\mu\(t\)=M\_\{i\\overline\{j\}\}X^\{i\}X^\{\\overline\{j\}\}\.\(11\.49\)Differentiating,
ddtμ\(t\)=\(DdtMij¯\)XiXj¯\.\\displaystyle\\frac\{d\}\{dt\}\\mu\(t\)=\(\\frac\{D\}\{dt\}M\_\{i\\overline\{j\}\}\)X^\{i\}X^\{\\overline\{j\}\}\.\(11\.50\)Contracting with simplified[11\.42](https://arxiv.org/html/2608.19584#S11.E42),
ddtμ\(t\)=Mik¯Mj¯k¯XiXj¯\+NipNj¯pXiXj¯\+Rij¯pq¯XiXj¯∇pℒ∇q¯ℒ−XiXj¯∇ij¯\(‖∇ℒ‖2\)\.\\displaystyle\\frac\{d\}\{dt\}\\mu\(t\)=M\_\{i\\overline\{k\}\}M^\{\\overline\{k\}\}\_\{\\overline\{j\}\}X^\{i\}X^\{\\overline\{j\}\}\+N\_\{ip\}N^\{p\}\_\{\\overline\{j\}\}X^\{i\}X^\{\\overline\{j\}\}\+R\_\{i\\overline\{j\}p\\overline\{q\}\}X^\{i\}X^\{\\overline\{j\}\}\\nabla^\{p\}\\mathcal\{L\}\\nabla^\{\\overline\{q\}\}\\mathcal\{L\}\-X^\{i\}X^\{\\overline\{j\}\}\\nabla\_\{i\\overline\{j\}\}\(\\\|\\nabla\\mathcal\{L\}\\\|^\{2\}\)\.\(11\.51\)BecauseXXis an eigenvector,
Mik¯Mj¯k¯XiXj¯=\(Mik¯Xi\)\(Mjl¯Xj¯\)=\(μ\(t\)Xk¯\)\(μ\(t\)Xk¯\)=μ\(t\)2\.\\displaystyle M\_\{i\\overline\{k\}\}M^\{\\overline\{k\}\}\_\{\\overline\{j\}\}X^\{i\}X^\{\\overline\{j\}\}=\(M\_\{i\\overline\{k\}\}X^\{i\}\)\(\\overline\{M\_\{j\\overline\{l\}\}X^\{j\}\}\)=\(\\mu\(t\)X\_\{\\overline\{k\}\}\)\(\\mu\(t\)X^\{\\overline\{k\}\}\)=\\mu\(t\)^\{2\}\.\(11\.52\)The above commutes since the contraction makes it scalar\. Moreover, we can note
NipNj¯pXiXj¯=‖N\(X,⋅\)‖h2\.\\displaystyle N\_\{ip\}N^\{p\}\_\{\\overline\{j\}\}X^\{i\}X^\{\\overline\{j\}\}=\\\|N\(X,\\cdot\)\\\|\_\{h\}^\{2\}\.\(11\.53\)Putting everything together, we conclude
ddtμ\(t\)≥−μ\(t\)2−‖N\(X,⋅\)‖h2\+Rij¯pq¯XiXj¯∇pℒ∇q¯ℒ\.\\displaystyle\\frac\{d\}\{dt\}\\mu\(t\)\\geq\-\\mu\(t\)^\{2\}\-\\\|N\(X,\\cdot\)\\\|\_\{h\}^\{2\}\+R\_\{i\\overline\{j\}p\\overline\{q\}\}X^\{i\}X^\{\\overline\{j\}\}\\nabla^\{p\}\\mathcal\{L\}\\nabla^\{\\overline\{q\}\}\\mathcal\{L\}\.\(11\.54\)Assuming the holomorphic bisectional curvature is positive and bounded below by a constantκ\>0\\kappa\>0, we haveRij¯pq¯XiXj¯∇pℒ∇q¯ℒ≥κv\(t\)R\_\{i\\overline\{j\}p\\overline\{q\}\}X^\{i\}X^\{\\overline\{j\}\}\\nabla^\{p\}\\mathcal\{L\}\\nabla^\{\\overline\{q\}\}\\mathcal\{L\}\\geq\\kappa v\(t\), wherev\(t\)=‖∇ℒ‖h2v\(t\)=\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{2\}\. We can rewrite
ddtμ\(t\)≥μ\(t\)2\+‖N\(X,⋅\)‖h2\+κv\(t\)−XiXj¯∇ij¯\(‖∇ℒ‖2\)\.\\displaystyle\\frac\{d\}\{dt\}\\mu\(t\)\\geq\\mu\(t\)^\{2\}\+\\\|N\(X,\\cdot\)\\\|\_\{h\}^\{2\}\+\\kappa v\(t\)\-X^\{i\}X^\{\\overline\{j\}\}\\nabla\_\{i\\overline\{j\}\}\(\\\|\\nabla\\mathcal\{L\}\\\|^\{2\}\)\.\(11\.55\)
We can also find an upper bound on the minimum eigenvalue\. This shows dependence on the loss landscape, and it moreover shows the exponential does not blow\-up in finite time, giving a stability result\. By Cauchy\-Schwarz in time,−∫0tμ\(s\)ds≤t\(∫0tμ\(s\)2ds\)1/2\-\\int\_\{0\}^\{t\}\\mu\(s\)ds\\leq\\sqrt\{t\}\\left\(\\int\_\{0\}^\{t\}\\mu\(s\)^\{2\}ds\\right\)^\{1/2\}\. We control thisL2L^\{2\}norm via the complex Bochner\-Weitzenböck identity \(we see this identity again in[12\.3](https://arxiv.org/html/2608.19584#S12.SS3)and a variant in[12\.1](https://arxiv.org/html/2608.19584#S12.SS1)\)
12Δ∂¯v=‖∇ω1,1ℒ‖h2\+‖∇ω2,0ℒ‖h2\+Ric\(∇hℒ,∇¯hℒ\)\+Re⟨∇hℒ,∇h\(Δ∂¯ℒ\)⟩h\.\\displaystyle\\frac\{1\}\{2\}\\Delta\_\{\\overline\{\\partial\}\}v=\\\|\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}\+\\\|\\nabla^\{2,0\}\_\{\\omega\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}\+\\text\{Ric\}\(\\nabla\_\{h\}\\mathcal\{L\},\\overline\{\\nabla\}\_\{h\}\\mathcal\{L\}\)\+\\text\{Re\}\\langle\\nabla\_\{h\}\\mathcal\{L\},\\nabla\_\{h\}\(\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}\)\\rangle\_\{h\}\.\(11\.56\)Using the identityddt\(Δ∂¯ℒ\)=−Re⟨∇hℒ,∇h\(Δ∂¯ℒ\)⟩h\\frac\{d\}\{dt\}\(\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}\)=\-\\text\{Re\}\\langle\\nabla\_\{h\}\\mathcal\{L\},\\nabla\_\{h\}\(\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}\)\\rangle\_\{h\}, we bound the\(1,1\)\(1,1\)Hessian trace from below byK2μ\(t\)2\\frac\{K\}\{2\}\\mu\(t\)^\{2\}, and defineκ\(t\)=infΩλmin\(Ric\)\\kappa\(t\)=\\inf\_\{\\Omega\}\\lambda\_\{\\text\{min\}\}\(\\text\{Ric\}\)\. Taking expectation by integrating and applying integration, assuming vanishing boundaries, and rearranging
ddt𝔼\[Δ∂¯ℒ\]≥K2μ\(t\)2\+κ\(t\)V\(t\)−12𝔼\[Δ∂¯ρtρtv\(t\)\]\.\\displaystyle\\frac\{d\}\{dt\}\\mathbb\{E\}\[\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}\]\\geq\\frac\{K\}\{2\}\\mu\(t\)^\{2\}\+\\kappa\(t\)V\(t\)\-\\frac\{1\}\{2\}\\mathbb\{E\}\\left\[\\frac\{\\Delta\_\{\\overline\{\\partial\}\}\\rho\_\{t\}\}\{\\rho\_\{t\}\}v\(t\)\\right\]\.\(11\.57\)Theμ\\muterm corresponds to the \(1,1\)\-Hessian term, theVVterm corresponds toRic\(∇hℒ,∇¯hℒ\)\\text\{Ric\}\(\\nabla\_\{h\}\\mathcal\{L\},\\overline\{\\nabla\}\_\{h\}\\mathcal\{L\}\), and the \(2,0\)\-Hessian term is dropped\. The12Δ∂¯v,Re⟨∇hℒ,∇h\(Δ∂¯ℒ\)⟩h\\frac\{1\}\{2\}\\Delta\_\{\\overline\{\\partial\}\}v,\\text\{Re\}\\langle\\nabla\_\{h\}\\mathcal\{L\},\\nabla\_\{h\}\(\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}\)\\rangle\_\{h\}terms swap sides\. Integrating over\[0,T\]\[0,T\]isolates theL2L^\{2\}bound on the minimum eigenvalue
∫0Tμ\(t\)2𝑑t≤2K\(𝔼\[Δ∂¯ℒ\(T\)\]−𝔼\[Δ∂¯ℒ\(0\)\]−∫0T\(κ\(t\)V\(t\)−12𝔼\[Δ∂¯ρtρtv\(t\)\]\)𝑑t\)\.\\displaystyle\\sqrt\{\\int\_\{0\}^\{T\}\\mu\(t\)^\{2\}dt\}\\leq\\sqrt\{\\frac\{2\}\{K\}\\left\(\\mathbb\{E\}\[\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}\(T\)\]\-\\mathbb\{E\}\[\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}\(0\)\]\-\\int\_\{0\}^\{T\}\\left\(\\kappa\(t\)V\(t\)\-\\frac\{1\}\{2\}\\mathbb\{E\}\\left\[\\frac\{\\Delta\_\{\\overline\{\\partial\}\}\\rho\_\{t\}\}\{\\rho\_\{t\}\}v\(t\)\\right\]\\right\)dt\\right\)\}\.\(11\.58\)Therefore,
V\(t\)≤V\(0\)exp\(2∫0tH\(s\)𝑑s\+2t‖μ‖L2\(\[0,t\]\)\),\\displaystyle V\(t\)\\leq V\(0\)\\exp\\left\(2\\int\_\{0\}^\{t\}H\(s\)ds\+2\\sqrt\{t\}\\\|\\mu\\\|\_\{L^\{2\}\(\[0,t\]\)\}\\right\),\(11\.59\)and simplifying gives a stability result, since the exponential does not blow\-up in finite time\. Let us try to be convinced this is sufficiently bounded and away from infinity\.ℒ\\mathcal\{L\}does not obey a maximum principle\. ForΔ∂¯ℒ\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}to admit a maximum principle,ℒ\\mathcal\{L\}would need to be subharmonic or superharmonic, which is unrealistic in a deep learning setting\. With a compact parameter space and sufficiently nice activations, the expected Dolbealt Laplacians on the loss are finite\. Of course, the Hessians of interest in our primarily results are onffwhich are interconnected withℒ\\mathcal\{L\}although not the same exactly since the loss is in terms of neural network output\. Theρs−1\\rho\_\{s\}^\{\-1\}is ostensibly the most problematic term, as it is a density with potentially a vanishing quality, although the expectation re\-adds the measure across the complex manifold, thereby canceling\.
□\\square
Remark\.Let us return to[11\.25](https://arxiv.org/html/2608.19584#S11.E25)\. We will rewrite it using physics notation since this is compatible and standard with Rayleigh quotients\. Denote\|Ψ\(t\)⟩:=\|∇hL\(t\)⟩\|\\Psi\(t\)\\rangle:=\|\\nabla\_\{h\}L\(t\)\\rangle\. We get
ddt⟨Ψ\|Ψ⟩h=−2Re⟨Ψ¯\|∇ω2,0L\|Ψ⟩h−2⟨Ψ\|∇ω1,1L\|Ψ⟩h\.\\displaystyle\\frac\{d\}\{dt\}\\langle\\Psi\|\\Psi\\rangle\_\{h\}=\-2\\text\{Re\}\\langle\\overline\{\\Psi\}\|\\nabla\_\{\\omega\}^\{2,0\}L\|\\Psi\\rangle\_\{h\}\-2\\langle\\Psi\|\\nabla\_\{\\omega\}^\{1,1\}L\|\\Psi\\rangle\_\{h\}\.\(11\.60\)Factoring outv\(t\)=⟨Ψ\|Ψ⟩hv\(t\)=\\langle\\Psi\|\\Psi\\rangle\_\{h\},
ddtv\(t\)=−\[2Re⟨Ψ¯\|∇ω2,0L\|Ψ⟩h⟨Ψ\|Ψ⟩h\+2⟨Ψ\|∇ω1,1L\|Ψ⟩h⟨Ψ\|Ψ⟩h\]v\(t\)\.\\displaystyle\\frac\{d\}\{dt\}v\(t\)=\-\\left\[2\\frac\{\\text\{Re\}\\langle\\overline\{\\Psi\}\|\\nabla\_\{\\omega\}^\{2,0\}L\|\\Psi\\rangle\_\{h\}\}\{\\langle\\Psi\|\\Psi\\rangle\_\{h\}\}\+2\\frac\{\\langle\\Psi\|\\nabla\_\{\\omega\}^\{1,1\}L\|\\Psi\\rangle\_\{h\}\}\{\\langle\\Psi\|\\Psi\\rangle\_\{h\}\}\\right\]v\(t\)\.\(11\.61\)We can notice in the interior there are two Rayleigh quotients\. Define two scaled Hessians such that∇ω2,0L=1mH~2,0\\nabla\_\{\\omega\}^\{2,0\}L=\\frac\{1\}\{\\sqrt\{m\}\}\\widetilde\{H\}^\{2,0\},∇ω1,1L=1mH~1,1\\nabla\_\{\\omega\}^\{1,1\}L=\\frac\{1\}\{\\sqrt\{m\}\}\\widetilde\{H\}^\{1,1\}\. We can observe
ddtv\(t\)=−1m\[2Re⟨Ψ¯\|H~2,0\|Ψ⟩h⟨Ψ\|Ψ⟩h\+2⟨Ψ\|H~1,1\|Ψ⟩h⟨Ψ\|Ψ⟩h\]v\(t\)\.\\displaystyle\\frac\{d\}\{dt\}v\(t\)=\-\\frac\{1\}\{\\sqrt\{m\}\}\\left\[2\\frac\{\\text\{Re\}\\langle\\overline\{\\Psi\}\|\\widetilde\{H\}^\{2,0\}\|\\Psi\\rangle\_\{h\}\}\{\\langle\\Psi\|\\Psi\\rangle\_\{h\}\}\+2\\frac\{\\langle\\Psi\|\\widetilde\{H\}^\{1,1\}\|\\Psi\\rangle\_\{h\}\}\{\\langle\\Psi\|\\Psi\\rangle\_\{h\}\}\\right\]v\(t\)\.\(11\.62\)In Rayleigh quotient notation
⟨𝒯def\(t\)⟩Ψ:=2Re⟨Ψ¯\|H~2,0\|Ψ⟩h⟨Ψ\|Ψ⟩h\+2⟨Ψ\|H~1,1\|Ψ⟩h⟨Ψ\|Ψ⟩h\.\\displaystyle\\langle\\mathcal\{T\}\_\{\\text\{def\}\}\(t\)\\rangle\_\{\\Psi\}:=2\\frac\{\\text\{Re\}\\langle\\overline\{\\Psi\}\|\\widetilde\{H\}^\{2,0\}\|\\Psi\\rangle\_\{h\}\}\{\\langle\\Psi\|\\Psi\\rangle\_\{h\}\}\+2\\frac\{\\langle\\Psi\|\\widetilde\{H\}^\{1,1\}\|\\Psi\\rangle\_\{h\}\}\{\\langle\\Psi\|\\Psi\\rangle\_\{h\}\}\.\(11\.63\)Therefore, we can note a reformulation with[12\.31](https://arxiv.org/html/2608.19584#S12.E31)
ddtv\(t\)\+1m⟨𝒯def\(t\)⟩Ψv\(t\)=0\.\\displaystyle\\frac\{d\}\{dt\}v\(t\)\+\\frac\{1\}\{\\sqrt\{m\}\}\\langle\\mathcal\{T\}\_\{\\text\{def\}\}\(t\)\\rangle\_\{\\Psi\}v\(t\)=0\.\(11\.64\)Via power series
\(ddt\+1m⟨𝒯def\(t\)⟩Ψ\)\(v0\(t\)\+1mv1\(t\)\+1mv2\(t\)\+…\)=0\.\\displaystyle\\left\(\\frac\{d\}\{dt\}\+\\frac\{1\}\{\\sqrt\{m\}\}\\langle\\mathcal\{T\}\_\{\\text\{def\}\}\(t\)\\rangle\_\{\\Psi\}\\right\)\\left\(v\_\{0\}\(t\)\+\\frac\{1\}\{\\sqrt\{m\}\}v\_\{1\}\(t\)\+\\frac\{1\}\{m\}v\_\{2\}\(t\)\+\\dots\\right\)=0\.\(11\.65\)Permitmmto be variable\. In reality,mmis fixed, but say in an asymptotic limitm→∞m\\rightarrow\\infty,mmis nonconstant\. We will show the constants are zero\. Denoteϵi=1/mi\\epsilon\_\{i\}=1/\\sqrt\{m\_\{i\}\}for short\. We can write this as a system
\{ℰ0\(t\)\+ϵ1ℰ1\(t\)\+ϵ12ℰ2\(t\)\+⋯\+ϵ1N−1ℰN−1\(t\)=0ℰ0\(t\)\+ϵ2ℰ1\(t\)\+ϵ22ℰ2\(t\)\+⋯\+ϵ2N−1ℰN−1\(t\)=0⋮ℰ0\(t\)\+ϵNℰ1\(t\)\+ϵN2ℰ2\(t\)\+⋯\+ϵNN−1ℰN−1\(t\)=0\.\\displaystyle\\begin\{cases\}\\mathcal\{E\}\_\{0\}\(t\)\+\\epsilon\_\{1\}\\mathcal\{E\}\_\{1\}\(t\)\+\\epsilon\_\{1\}^\{2\}\\mathcal\{E\}\_\{2\}\(t\)\+\\dots\+\\epsilon\_\{1\}^\{N\-1\}\\mathcal\{E\}\_\{N\-1\}\(t\)=0\\\\ \\mathcal\{E\}\_\{0\}\(t\)\+\\epsilon\_\{2\}\\mathcal\{E\}\_\{1\}\(t\)\+\\epsilon\_\{2\}^\{2\}\\mathcal\{E\}\_\{2\}\(t\)\+\\dots\+\\epsilon\_\{2\}^\{N\-1\}\\mathcal\{E\}\_\{N\-1\}\(t\)=0\\\\ \\quad\\vdots\\\\ \\mathcal\{E\}\_\{0\}\(t\)\+\\epsilon\_\{N\}\\mathcal\{E\}\_\{1\}\(t\)\+\\epsilon\_\{N\}^\{2\}\\mathcal\{E\}\_\{2\}\(t\)\+\\dots\+\\epsilon\_\{N\}^\{N\-1\}\\mathcal\{E\}\_\{N\-1\}\(t\)=0\.\\end\{cases\}\(11\.66\)For particular choice,
\[1ϵ1ϵ12…ϵ1N−11ϵ2ϵ22…ϵ2N−1⋱1ϵNϵN2…ϵNN−1\]\[ℰ0\(t\)ℰ1\(t\)ℰN−1\(t\)\]=\[000\],\\displaystyle\\begin\{bmatrix\}1&\\epsilon\_\{1\}&\\epsilon\_\{1\}^\{2\}&\\dots&\\epsilon\_\{1\}^\{N\-1\}\\\\ 1&\\epsilon\_\{2\}&\\epsilon\_\{2\}^\{2\}&\\dots&\\epsilon\_\{2\}^\{N\-1\}\\\\ \\vdots&\\vdots&\\vdots&\\ddots&\\vdots\\\\ 1&\\epsilon\_\{N\}&\\epsilon\_\{N\}^\{2\}&\\dots&\\epsilon\_\{N\}^\{N\-1\}\\end\{bmatrix\}\\begin\{bmatrix\}\\mathcal\{E\}\_\{0\}\(t\)\\\\ \\mathcal\{E\}\_\{1\}\(t\)\\\\ \\vdots\\\\ \\mathcal\{E\}\_\{N\-1\}\(t\)\\end\{bmatrix\}=\\begin\{bmatrix\}0\\\\ 0\\\\ \\vdots\\\\ 0\\end\{bmatrix\},\(11\.67\)which is Vandermonde and invertible, and writingVℰ\(t\)=0V\\mathcal\{E\}\(t\)=0, this has determinant
det\(V\)=∏1≤i<j≤N\(ϵj−ϵi\)\.\\displaystyle\\det\(V\)=\\prod\_\{1\\leq i<j\\leq N\}\(\\epsilon\_\{j\}\-\\epsilon\_\{i\}\)\.\(11\.68\)Use shorthand𝒟0:=ddt\\mathcal\{D\}\_\{0\}:=\\frac\{d\}\{dt\}denote the time derivative linear evolution operator infinite\-width, and we have
𝒟0v1\(t\)=−⟨𝒯def\(t\)⟩Ψv0\(t\),\\displaystyle\\mathcal\{D\}\_\{0\}v\_\{1\}\(t\)=\-\\langle\\mathcal\{T\}\_\{\\text\{def\}\}\(t\)\\rangle\_\{\\Psi\}v\_\{0\}\(t\),\(11\.69\)or equivalently,
vk\(t\)=−∫0t⟨𝒯def\(τ\)⟩Ψvk−1\(τ\)dτ\.\\displaystyle v\_\{k\}\(t\)=\-\\int\_\{0\}^\{t\}\\langle\\mathcal\{T\}\_\{\\text\{def\}\}\(\\tau\)\\rangle\_\{\\Psi\}v\_\{k\-1\}\(\\tau\)d\\tau\.\(11\.70\)What this primarily shows is that parameter evolution retains a "memory" of previous states, and that subsequent states are largely determined by their previous states, geometrically speaking\. In general, a Rayleigh quotient takes the form
Rayleigh\(M,x\)=x†Mxx†x,\\displaystyle\\text\{Rayleigh\}\(M,x\)=\\frac\{x^\{\\dagger\}Mx\}\{x^\{\\dagger\}x\},\(11\.71\)which is a way to stretch the space in the direction ofxx\. Moreover,λmin≤x†Mxx†x≤λmax\\lambda\_\{\\text\{min\}\}\\leq\\frac\{x^\{\\dagger\}Mx\}\{x^\{\\dagger\}x\}\\leq\\lambda\_\{\\text\{max\}\}for eigenvalues ofMM\.
### 11\.3Gradient flux searches
In this section, we attempt to characterize the total amount of flexibility across gradients in anϵ\\epsilon\-ball around initialization\. The motivation for this is we will examine the flux across varying radii via
∫0ϵ\(∮∂Br\(θ0\)dcℒ∧ωK−1\)𝑑r\.\\displaystyle\\int\_\{0\}^\{\\epsilon\}\\left\(\\oint\_\{\\partial B\_\{r\}\(\\theta\_\{0\}\)\}d^\{c\}\\mathcal\{L\}\\wedge\\omega^\{K\-1\}\\right\)dr\.\(11\.72\)and attempt to quantify this flux value\. Generally, a larger value is more desirable\. This implies there is greater variability for trajectory paths, and greater ability to escape theϵ\\epsilon\-ball in a quickly\-descending descent path\. In particular, we will try to characterize descent behavior pathologically using descent rules and the gradients\. For example, we examine a ball around initialization via flux around the boundary with Stokes’ theorem
∮∂Bϵ\(θ0\)dcℒ∧ωK−1=∫Bϵ\(θ0\)i∂∂¯ℒ∧ωK−1∝∫Bϵ\(θ0\)\(Δ∂¯ℒ\)ωKK\!\.\\displaystyle\\oint\_\{\\partial B\_\{\\epsilon\}\(\\theta\_\{0\}\)\}d^\{c\}\\mathcal\{L\}\\wedge\\omega^\{K\-1\}=\\int\_\{B\_\{\\epsilon\}\(\\theta\_\{0\}\)\}i\\partial\\overline\{\\partial\}\\mathcal\{L\}\\wedge\\omega^\{K\-1\}\\propto\\int\_\{B\_\{\\epsilon\}\(\\theta\_\{0\}\)\}\(\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}\)\\frac\{\\omega^\{K\}\}\{K\!\}\.\(11\.73\)The left\-hand side corresponds to the gradients on the loss sincedc=−i2\(∂−∂¯\)d^\{c\}=\-\\frac\{i\}\{2\}\(\\partial\-\\overline\{\\partial\}\), and we can link it to the Laplacian proportionally\. Under a well\-curvature\-conditioned landscape, the Laplacian is positive and behaves nicely\. This translates to robust gradient flux\. If the gradient of the loss behaves nicely, so do the descent directions\. In a profaned\-curvature landscape, the eigenvalue blowup is characterized with anisotropy and Laplacian contributions from distorted directions are affected, and overall descent paths are affected too\. Ill\-conditioned gradients on the loss affect learning negatively\. We can noted\(dcℒ∧ωK−1\)=ddcℒ∧ωK−1=i∂∂¯ℒ∧ωK−1d\(d^\{c\}\\mathcal\{L\}\\wedge\\omega^\{K\-1\}\)=dd^\{c\}\\mathcal\{L\}\\wedge\\omega^\{K\-1\}=i\\partial\\overline\{\\partial\}\\mathcal\{L\}\\wedge\\omega^\{K\-1\}and the Hessian of the loss is an exact 2\-form\. In de Rham cohomology, all exact forms are trivial; however, this is not to say the result of manipulating[11\.73](https://arxiv.org/html/2608.19584#S11.E73)is not affected by Ricci curvature, which it is\. We prove this in Appendix[11\.3](https://arxiv.org/html/2608.19584#S11.SS3), therefore this result is affected whether or not the manifold is Calabi\-Yau\.
Proof of Lemma 11\.We find the flux across the ball around initialization and accumulate it across radii by by integrating the\(2K−1\)\(2K\-1\)\-form over the boundary∂Br\(θ0\)\\partial B\_\{r\}\(\\theta\_\{0\}\)and pulling it across the radiusrr
𝒲\(ϵ\)=∫0ϵ\(∮∂Br\(θ0\)dcℒ∧ωK−1\)𝑑r\.\\displaystyle\\mathcal\{W\}\(\\epsilon\)=\\int\_\{0\}^\{\\epsilon\}\\left\(\\oint\_\{\\partial B\_\{r\}\(\\theta\_\{0\}\)\}d^\{c\}\\mathcal\{L\}\\wedge\\omega^\{K\-1\}\\right\)dr\.\(11\.74\)We apply Stokes’ theorem\. As we saw earlier in[11\.73](https://arxiv.org/html/2608.19584#S11.E73), but including constants, and using the relation between the Kähler form and the Laplacian,
∮∂Br\(θ0\)dcℒ∧ωK−1=∫Br\(θ0\)i∂∂¯ℒ∧ωK−1=1K∫Br\(θ0\)\(Δ∂¯ℒ\)ωK\.\\displaystyle\\oint\_\{\\partial B\_\{r\}\(\\theta\_\{0\}\)\}d^\{c\}\\mathcal\{L\}\\wedge\\omega^\{K\-1\}=\\int\_\{B\_\{r\}\(\\theta\_\{0\}\)\}i\\partial\\overline\{\\partial\}\\mathcal\{L\}\\wedge\\omega^\{K\-1\}=\\frac\{1\}\{K\}\\int\_\{B\_\{r\}\(\\theta\_\{0\}\)\}\(\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}\)\\omega^\{K\}\.\(11\.75\)The local expansion of the Kähler form in a local coordinate chartξ\\xiup to fourth order is[35](https://arxiv.org/html/2608.19584#bib.bib30)[60](https://arxiv.org/html/2608.19584#bib.bib31)
Φ\(ξ,ξ¯\)=∑i=1dξiξ¯i−14Rij¯kl¯ξiξ¯jξkξ¯l\+𝒪\(‖ξ‖5\)\.\\displaystyle\\Phi\(\\xi,\\overline\{\\xi\}\)=\\sum\_\{i=1\}^\{d\}\\xi^\{i\}\\overline\{\\xi\}^\{i\}\-\\frac\{1\}\{4\}R\_\{i\\overline\{j\}k\\overline\{l\}\}\\xi^\{i\}\\overline\{\\xi\}^\{j\}\\xi^\{k\}\\overline\{\\xi\}^\{l\}\+\\mathcal\{O\}\(\\\|\\xi\\\|^\{5\}\)\.\(11\.76\)This is a local property of a Kähler manifold under a fixed gauge\. Differentiating since the metric is the Wirtinger Hessian of the potential,
hij¯\(ξ\)=δij¯−Rij¯kl¯ξkξ¯l\+𝒪\(‖ξ‖3\)\.\\displaystyle h\_\{i\\overline\{j\}\}\(\\xi\)=\\delta\_\{i\\overline\{j\}\}\-R\_\{i\\overline\{j\}k\\overline\{l\}\}\\xi^\{k\}\\overline\{\\xi\}^\{l\}\+\\mathcal\{O\}\(\\\|\\xi\\\|^\{3\}\)\.\(11\.77\)Letω0=i2δij¯dθi∧dθ¯j\\omega\_\{0\}=\\frac\{i\}\{2\}\\delta\_\{i\\overline\{j\}\}d\\theta^\{i\}\\wedge d\\overline\{\\theta\}^\{j\}be the standard flat Kähler form\. The volume expansion can be written
ωKK\!=\(1−Rickl¯ξkξ¯l\+𝒪\(‖ξ‖3\)\)ω0KK\!\.\\displaystyle\\frac\{\\omega^\{K\}\}\{K\!\}=\\left\(1\-\\text\{Ric\}\_\{k\\overline\{l\}\}\\xi^\{k\}\\overline\{\\xi\}^\{l\}\+\\mathcal\{O\}\(\\\|\\xi\\\|^\{3\}\)\\right\)\\frac\{\\omega\_\{0\}^\{K\}\}\{K\!\}\.\(11\.78\)Let us take the inversehh,
hij¯\(ξ\)=δij¯\+Rkl¯ij¯ξkξ¯l\+𝒪\(‖ξ‖3\)\.\\displaystyle h^\{i\\overline\{j\}\}\(\\xi\)=\\delta^\{i\\overline\{j\}\}\+R^\{i\\overline\{j\}\}\_\{\\phantom\{i\\overline\{j\}\}k\\overline\{l\}\}\\xi^\{k\}\\overline\{\\xi\}^\{l\}\+\\mathcal\{O\}\(\\\|\\xi\\\|^\{3\}\)\.\(11\.79\)Therefore, contracting the inverse metric with the mixed Hessian
Δ∂¯ℒ=\(δij¯\+Rkl¯ij¯ξkξ¯l\+𝒪\(‖ξ‖3\)\)∂i∂¯jℒ\.\\displaystyle\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}=\\left\(\\delta^\{i\\overline\{j\}\}\+R^\{i\\overline\{j\}\}\_\{\\phantom\{i\\overline\{j\}\}k\\overline\{l\}\}\\xi^\{k\}\\overline\{\\xi\}^\{l\}\+\\mathcal\{O\}\(\\\|\\xi\\\|^\{3\}\)\\right\)\\partial\_\{i\}\\overline\{\\partial\}\_\{j\}\\mathcal\{L\}\.\(11\.80\)Multiplying by the volume form as in[11\.78](https://arxiv.org/html/2608.19584#S11.E78),
\(Δ∂¯ℒ\)ωKK\!\\displaystyle\(\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}\)\\frac\{\\omega^\{K\}\}\{K\!\}=\\bBigg@3\[Δ0ℒ\+Rkl¯ij¯ξkξ¯l∂i∂¯jℒ\+𝒪\(‖ξ‖3\)\\bBigg@3\]\\bBigg@3\[1−Ricmn¯ξmξ¯n\+𝒪\(‖ξ‖3\)\\bBigg@3\]ω0KK\!\\displaystyle=\\bBigg@\{3\}\[\\Delta\_\{0\}\\mathcal\{L\}\+R^\{i\\overline\{j\}\}\_\{\\phantom\{i\\overline\{j\}\}k\\overline\{l\}\}\\xi^\{k\}\\overline\{\\xi\}^\{l\}\\partial\_\{i\}\\overline\{\\partial\}\_\{j\}\\mathcal\{L\}\+\\mathcal\{O\}\(\\\|\\xi\\\|^\{3\}\)\\bBigg@\{3\}\]\\bBigg@\{3\}\[1\-\\text\{Ric\}\_\{m\\overline\{n\}\}\\xi^\{m\}\\overline\{\\xi\}^\{n\}\+\\mathcal\{O\}\(\\\|\\xi\\\|^\{3\}\)\\bBigg@\{3\}\]\\frac\{\\omega\_\{0\}^\{K\}\}\{K\!\}\(11\.81\)=\[δij¯∂i∂¯jℒ\+\(Rkl¯ij¯∂i∂¯jℒ−Rickl¯\(Δ0ℒ\)\)ξkξ¯l\+𝒪\(‖ξ‖3\)\]ω0KK\!\.\\displaystyle=\\left\[\\delta^\{i\\overline\{j\}\}\\partial\_\{i\}\\overline\{\\partial\}\_\{j\}\\mathcal\{L\}\+\\left\(R^\{i\\overline\{j\}\}\_\{\\phantom\{i\\overline\{j\}\}k\\overline\{l\}\}\\partial\_\{i\}\\overline\{\\partial\}\_\{j\}\\mathcal\{L\}\-\\text\{Ric\}\_\{k\\overline\{l\}\}\(\\Delta\_\{0\}\\mathcal\{L\}\)\\right\)\\xi^\{k\}\\overline\{\\xi\}^\{l\}\+\\mathcal\{O\}\(\\\|\\xi\\\|^\{3\}\)\\right\]\\frac\{\\omega\_\{0\}^\{K\}\}\{K\!\}\.\(11\.82\)Let us examine the integrand term
\(Rkl¯ij¯∂i∂¯jℒ−Rickl¯\(Δ0ℒ\)\)ξkξ¯lω0KK\!\\displaystyle\\left\(R^\{i\\overline\{j\}\}\_\{\\phantom\{i\\overline\{j\}\}k\\overline\{l\}\}\\partial\_\{i\}\\overline\{\\partial\}\_\{j\}\\mathcal\{L\}\-\\text\{Ric\}\_\{k\\overline\{l\}\}\(\\Delta\_\{0\}\\mathcal\{L\}\)\\right\)\\xi^\{k\}\\overline\{\\xi\}^\{l\}\\frac\{\\omega\_\{0\}^\{K\}\}\{K\!\}\(11\.83\)Under a Taylor expansion,
∂i∂¯jℒ\(ξ\)=∂i∂¯jℒ\(0\)\+Re\(cij¯mξm\)\+𝒪\(‖ξ‖2\)\.\\displaystyle\\partial\_\{i\}\\overline\{\\partial\}\_\{j\}\\mathcal\{L\}\(\\xi\)=\\partial\_\{i\}\\overline\{\\partial\}\_\{j\}\\mathcal\{L\}\(0\)\+\\text\{Re\}\\left\(c\_\{i\\overline\{j\}m\}\\xi^\{m\}\\right\)\+\\mathcal\{O\}\(\\\|\\xi\\\|^\{2\}\)\.\(11\.84\)When we contract with theξkξ¯l\\xi^\{k\}\\overline\{\\xi\}^\{l\}, we get a constant termRkl¯ij¯\(0\)⋅∂i∂¯jℒ\(0\)⋅ξkξ¯lR^\{i\\overline\{j\}\}\_\{\\phantom\{i\\overline\{j\}\}k\\overline\{l\}\}\(0\)\\cdot\\partial\_\{i\}\\overline\{\\partial\}\_\{j\}\\mathcal\{L\}\(0\)\\cdot\\xi^\{k\}\\overline\{\\xi\}^\{l\}, an odd power term involvingξkξ¯lξm\\xi^\{k\}\\overline\{\\xi\}^\{l\}\\xi^\{m\}, and a higher order term\. Because we are integrating a symmetric ball, the odd powers vanish due to symmetry\. Therefore, let us examine the curvature constant tensor
Ckl¯=Rkl¯ij¯\(0\)∂i∂¯jℒ\(0\)−Rickl¯\(0\)Δ0ℒ\(0\),\\displaystyle C\_\{k\\overline\{l\}\}=R^\{i\\overline\{j\}\}\_\{\\phantom\{i\\overline\{j\}\}k\\overline\{l\}\}\(0\)\\partial\_\{i\}\\overline\{\\partial\}\_\{j\}\\mathcal\{L\}\(0\)\-\\text\{Ric\}\_\{k\\overline\{l\}\}\(0\)\\Delta\_\{0\}\\mathcal\{L\}\(0\),\(11\.85\)and we must evaluate this with the contraction, the form, and the integral added
∫BrCkl¯ξkξ¯lω0KK\!=Ckl¯∫Brξkξ¯lω0KK\!\.\\displaystyle\\int\_\{B\_\{r\}\}C\_\{k\\overline\{l\}\}\\xi^\{k\}\\overline\{\\xi\}^\{l\}\\frac\{\\omega\_\{0\}^\{K\}\}\{K\!\}=C\_\{k\\overline\{l\}\}\\int\_\{B\_\{r\}\}\\xi^\{k\}\\overline\{\\xi\}^\{l\}\\frac\{\\omega\_\{0\}^\{K\}\}\{K\!\}\.\(11\.86\)Let us examine two cases of the above\. The first is that we assumek≠lk\\neq l\. In polar coordinatesθk=ρkeiϕk\\theta^\{k\}=\\rho\_\{k\}e^\{i\\phi\_\{k\}\}, the integral portion includes∫02πeiϕkdϕk∫02πe−iϕldϕl\\int\_\{0\}^\{2\\pi\}e^\{i\\phi\_\{k\}\}d\\phi\_\{k\}\\int\_\{0\}^\{2\\pi\}e^\{\-i\\phi\_\{l\}\}d\\phi\_\{l\}\. This integral vanishes to exactly 0\. Let us examine the diagonal casek=lk=l\. The integral of the squared magnitude of one coordinate is the same as any other coordinate by symmetry\. From this, we can note
∑m=1K∫Br‖zm‖2dV0=∫Br‖z‖2dV0\\displaystyle\\sum\_\{m=1\}^\{K\}\\int\_\{B\_\{r\}\}\\\|z^\{m\}\\\|^\{2\}dV\_\{0\}=\\int\_\{B\_\{r\}\}\\\|z\\\|^\{2\}dV\_\{0\}\(11\.87\)K∫Br‖zk‖2dV0=∫Br‖z‖2dV0⟹∫Br‖zk‖2dV0=1K∫Br‖z‖2dV0\.\\displaystyle K\\int\_\{B\_\{r\}\}\\\|z^\{k\}\\\|^\{2\}dV\_\{0\}=\\int\_\{B\_\{r\}\}\\\|z\\\|^\{2\}dV\_\{0\}\\implies\\int\_\{B\_\{r\}\}\\\|z^\{k\}\\\|^\{2\}dV\_\{0\}=\\frac\{1\}\{K\}\\int\_\{B\_\{r\}\}\\\|z\\\|^\{2\}dV\_\{0\}\.\(11\.88\)Combining the two cases, we can develop
∫Brξkξ¯lω0KK\!=δkl¯1K∫Br‖ξ‖2ω0KK\!=δkl¯πKr2K\+2K\!\(K\+1\)\.\\displaystyle\\int\_\{B\_\{r\}\}\\xi^\{k\}\\overline\{\\xi\}^\{l\}\\frac\{\\omega\_\{0\}^\{K\}\}\{K\!\}=\\delta^\{k\\overline\{l\}\}\\frac\{1\}\{K\}\\int\_\{B\_\{r\}\}\\\|\\xi\\\|^\{2\}\\frac\{\\omega\_\{0\}^\{K\}\}\{K\!\}=\\delta^\{k\\overline\{l\}\}\\frac\{\\pi^\{K\}r^\{2K\+2\}\}\{K\!\(K\+1\)\}\.\(11\.89\)The last equality is the evauluation of the integral, i\.e\. using polar coordinates\. Therefore, returning to[11\.86](https://arxiv.org/html/2608.19584#S11.E86), we can contract the curvature term with a Kronecker delta
Ckl¯δkl¯=Rkl¯ij¯\(0\)δkl¯∂i∂¯jℒ\(0\)−Rickl¯\(0\)δkl¯Δ0ℒ\(0\)\.\\displaystyle C\_\{k\\overline\{l\}\}\\delta^\{k\\overline\{l\}\}=R^\{i\\overline\{j\}\}\_\{\\phantom\{i\\overline\{j\}\}k\\overline\{l\}\}\(0\)\\delta^\{k\\overline\{l\}\}\\partial\_\{i\}\\overline\{\\partial\}\_\{j\}\\mathcal\{L\}\(0\)\-\\text\{Ric\}\_\{k\\overline\{l\}\}\(0\)\\delta^\{k\\overline\{l\}\}\\Delta\_\{0\}\\mathcal\{L\}\(0\)\.\(11\.90\)By definitionRkl¯ij¯δkl¯=Ricij¯R^\{i\\overline\{j\}\}\_\{\\phantom\{i\\overline\{j\}\}k\\overline\{l\}\}\\delta^\{k\\overline\{l\}\}=\\text\{Ric\}^\{i\\overline\{j\}\}andRickl¯δkl¯=R\\text\{Ric\}\_\{k\\overline\{l\}\}\\delta^\{k\\overline\{l\}\}=R\. Putting everything together, namely returning to the integral of[11\.75](https://arxiv.org/html/2608.19584#S11.E75), the first term being the Kronecker contraction and the interior using[11\.90](https://arxiv.org/html/2608.19584#S11.E90)
∫Br\(θ0\)\(Δ∂¯ℒ\)ω0KK\!=Δ0ℒ\(0\)πKr2KK\!\+\[Ricij¯\(0\)∂i∂¯jℒ\(0\)−RΔ0ℒ\(0\)\]πKr2K\+2K\!\(K\+1\)\+𝒪\(r2K\+4\)\.\\displaystyle\\int\_\{B\_\{r\}\(\\theta\_\{0\}\)\}\(\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}\)\\frac\{\\omega\_\{0\}^\{K\}\}\{K\!\}=\\Delta\_\{0\}\\mathcal\{L\}\(0\)\\frac\{\\pi^\{K\}r^\{2K\}\}\{K\!\}\+\\left\[\\text\{Ric\}^\{i\\overline\{j\}\}\(0\)\\partial\_\{i\}\\overline\{\\partial\}\_\{j\}\\mathcal\{L\}\(0\)\-R\\Delta\_\{0\}\\mathcal\{L\}\(0\)\\right\]\\frac\{\\pi^\{K\}r^\{2K\+2\}\}\{K\!\(K\+1\)\}\+\\mathcal\{O\}\(r^\{2K\+4\}\)\.\(11\.91\)The first term includes the integral of the volume form with identity coefficient integrand\. Multiplying through the factorial term,
∫Br\(θ0\)\(Δ∂¯ℒ\)ωK=Δ0ℒ\(0\)πKr2K\+\[Ricij¯\(0\)∂i∂¯jℒ\(0\)−RΔ0ℒ\(0\)\]πKr2K\+2K\+1\+𝒪\(r2K\+4\)\.\\displaystyle\\int\_\{B\_\{r\}\(\\theta\_\{0\}\)\}\(\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}\)\\omega^\{K\}=\\Delta\_\{0\}\\mathcal\{L\}\(0\)\\pi^\{K\}r^\{2K\}\+\\left\[\\text\{Ric\}^\{i\\overline\{j\}\}\(0\)\\partial\_\{i\}\\overline\{\\partial\}\_\{j\}\\mathcal\{L\}\(0\)\-R\\Delta\_\{0\}\\mathcal\{L\}\(0\)\\right\]\\frac\{\\pi^\{K\}r^\{2K\+2\}\}\{K\+1\}\+\\mathcal\{O\}\(r^\{2K\+4\}\)\.\(11\.92\)Returning to our Stokes’ theorem identity in[11\.75](https://arxiv.org/html/2608.19584#S11.E75),
∮∂Br\(θ0\)dcℒ∧ωK−1=1K∫Br\(θ0\)\(Δ∂¯ℒ\)ωK\\displaystyle\\oint\_\{\\partial B\_\{r\}\(\\theta\_\{0\}\)\}d^\{c\}\\mathcal\{L\}\\wedge\\omega^\{K\-1\}=\\frac\{1\}\{K\}\\int\_\{B\_\{r\}\(\\theta\_\{0\}\)\}\(\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}\)\\omega^\{K\}\(11\.93\)=πKKΔ0ℒ\(0\)r2K\+πKK\(K\+1\)\[Ricij¯\(0\)∂i∂¯jℒ\(0\)−RΔ0ℒ\(0\)\]r2K\+2\+𝒪\(r2K\+4\)\.\\displaystyle=\\frac\{\\pi^\{K\}\}\{K\}\\Delta\_\{0\}\\mathcal\{L\}\(0\)r^\{2K\}\+\\frac\{\\pi^\{K\}\}\{K\(K\+1\)\}\\left\[\\text\{Ric\}^\{i\\overline\{j\}\}\(0\)\\partial\_\{i\}\\overline\{\\partial\}\_\{j\}\\mathcal\{L\}\(0\)\-R\\Delta\_\{0\}\\mathcal\{L\}\(0\)\\right\]r^\{2K\+2\}\+\\mathcal\{O\}\(r^\{2K\+4\}\)\.\(11\.94\)Lastly, integrating over the ball,
𝒲\(ϵ\)=∫0ϵ\(∮∂Br\(θ0\)dcℒ∧ωK−1\)𝑑r\\displaystyle\\mathcal\{W\}\(\\epsilon\)=\\int\_\{0\}^\{\\epsilon\}\\left\(\\oint\_\{\\partial B\_\{r\}\(\\theta\_\{0\}\)\}d^\{c\}\\mathcal\{L\}\\wedge\\omega^\{K\-1\}\\right\)dr\(11\.95\)=∫0ϵ\(πKKΔ0ℒ\(0\)r2K\+πKK\(K\+1\)\[Ricij¯\(0\)∂i∂¯jℒ\(0\)−RΔ0ℒ\(0\)\]r2K\+2\+𝒪\(r2K\+4\)\)𝑑r\.\\displaystyle=\\int\_\{0\}^\{\\epsilon\}\\left\(\\frac\{\\pi^\{K\}\}\{K\}\\Delta\_\{0\}\\mathcal\{L\}\(0\)r^\{2K\}\+\\frac\{\\pi^\{K\}\}\{K\(K\+1\)\}\\left\[\\text\{Ric\}^\{i\\overline\{j\}\}\(0\)\\partial\_\{i\}\\overline\{\\partial\}\_\{j\}\\mathcal\{L\}\(0\)\-R\\Delta\_\{0\}\\mathcal\{L\}\(0\)\\right\]r^\{2K\+2\}\+\\mathcal\{O\}\(r^\{2K\+4\}\)\\right\)dr\.\(11\.96\)Finishing the integration, we have the final result
𝒲\(ϵ\)=πKK\(2K\+1\)Δ0ℒ\(0\)ϵ2K\+1\+πKK\(K\+1\)\(2K\+3\)\[Ricij¯\(0\)∂i∂¯jℒ\(0\)−RΔ0ℒ\(0\)\]ϵ2K\+3\+𝒪\(ϵ2K\+5\)\.\\displaystyle\\mathcal\{W\}\(\\epsilon\)=\\frac\{\\pi^\{K\}\}\{K\(2K\+1\)\}\\Delta\_\{0\}\\mathcal\{L\}\(0\)\\epsilon^\{2K\+1\}\+\\frac\{\\pi^\{K\}\}\{K\(K\+1\)\(2K\+3\)\}\\left\[\\text\{Ric\}^\{i\\overline\{j\}\}\(0\)\\partial\_\{i\}\\overline\{\\partial\}\_\{j\}\\mathcal\{L\}\(0\)\-R\\Delta\_\{0\}\\mathcal\{L\}\(0\)\\right\]\\epsilon^\{2K\+3\}\+\\mathcal\{O\}\(\\epsilon^\{2K\+5\}\)\.\(11\.97\)Shorthand, this becomes
𝒲\(ϵ\)=𝒪\(Δ0ℒϵ2K\+1\+\[⟨Ric,ℋℒ⟩−RΔ0ℒ\]ϵ2K\+3\+𝒪\(ϵ2K\+5\)\)\.\\displaystyle\\mathcal\{W\}\(\\epsilon\)=\\mathcal\{O\}\\left\(\\Delta\_\{0\}\\mathcal\{L\}\\epsilon^\{2K\+1\}\+\\big\[\\langle\\text\{Ric\},\\mathcal\{H\}\_\{\\mathcal\{L\}\}\\rangle\-R\\Delta\_\{0\}\\mathcal\{L\}\\big\]\\epsilon^\{2K\+3\}\+\\mathcal\{O\}\(\\epsilon^\{2K\+5\}\)\\right\)\.\(11\.98\)Let us tailor to our context\. Under a quadratic costℒ\(z\)=∑α\(fα\(z\)−yα\)2\\mathcal\{L\}\(z\)=\\sum\_\{\\alpha\}\(f\_\{\\alpha\}\(z\)\-y\_\{\\alpha\}\)^\{2\}, and since‖i∂∂¯f‖2=𝒪\(1m\)\\\|i\\partial\\overline\{\\partial\}f\\\|\_\{2\}=\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\), we can note via chain rule
∂i∂¯jℒ\(0\)=2∑α∂ifα\(0\)∂¯jfα\(0\)\+𝒪\(1m\)\.\\displaystyle\\partial\_\{i\}\\overline\{\\partial\}\_\{j\}\\mathcal\{L\}\(0\)=2\\sum\_\{\\alpha\}\\partial\_\{i\}f\_\{\\alpha\}\(0\)\\overline\{\\partial\}\_\{j\}f\_\{\\alpha\}\(0\)\+\\mathcal\{O\}\(\\frac\{1\}\{\\sqrt\{m\}\}\)\.\(11\.99\)Contracting with the flat metric, since we have definedΔ0\\Delta\_\{0\}this way,
Δ0ℒ\(0\)=2∑α‖∂fα\(0\)‖2\+𝒪\(Km\)\.\\displaystyle\\Delta\_\{0\}\\mathcal\{L\}\(0\)=2\\sum\_\{\\alpha\}\\\|\\partial f\_\{\\alpha\}\(0\)\\\|^\{2\}\+\\mathcal\{O\}\(\\frac\{K\}\{\\sqrt\{m\}\}\)\.\(11\.100\)Substituting back into our integrated flux of[11\.97](https://arxiv.org/html/2608.19584#S11.E97), we recover the quantity\. In the Calabi\-Yau case, the second term vanishes, leaving us with
𝒲CY\(ϵ\)=2πKK\(2K\+1\)\(∑α‖∂fα\(0\)‖2\)ϵ2K\+1\+𝒪\(ϵ2K\+1\(2K\+1\)m\)\+𝒪\(ϵ2K\+5\)\.\\displaystyle\\mathcal\{W\}\_\{CY\}\(\\epsilon\)=\\frac\{2\\pi^\{K\}\}\{K\(2K\+1\)\}\\left\(\\sum\_\{\\alpha\}\\\|\\partial f\_\{\\alpha\}\(0\)\\\|^\{2\}\\right\)\\epsilon^\{2K\+1\}\+\\mathcal\{O\}\\left\(\\frac\{\\epsilon^\{2K\+1\}\}\{\(2K\+1\)\\sqrt\{m\}\}\\right\)\+\\mathcal\{O\}\\left\(\\epsilon^\{2K\+5\}\\right\)\.\(11\.101\)This middle term can be reduced in the negatively curved scenario to
−c‖∂fα\(0\)‖2−\(−cK\)‖∂fα\(0\)‖2=c\(K−1\)‖∂fα\(0\)‖2\\displaystyle\-c\\\|\\partial f\_\{\\alpha\}\(0\)\\\|^\{2\}\-\(\-cK\)\\\|\\partial f\_\{\\alpha\}\(0\)\\\|^\{2\}=c\(K\-1\)\\\|\\partial f\_\{\\alpha\}\(0\)\\\|^\{2\}\(11\.102\)since the trace picks up a dimension, and this total term is positive\. We can deduce𝒲\(ϵ\)\\mathcal\{W\}\(\\epsilon\)is larger in the non\-Calabi\-Yau case, which means that the accumulated flux is higher\. This means that gradients are stronger in that region, and there is greater escape of initialization\.
□\\square
## 12Additional results with negative curvature
### 12\.1Asymptotic variance
Proof of Lemma 12\.Let us examine the steady state of the Fokker\-Planck equation on the manifold in the limitt→∞t\\rightarrow\\infty
ρ∞\(θ\)=1𝒵e−2ηℒ\(θ\),𝒵=∫Ue−2ηℒ\(θ\)ωKK\!,\\displaystyle\\rho\_\{\\infty\}\(\\theta\)=\\frac\{1\}\{\\mathcal\{Z\}\}e^\{\-\\frac\{2\}\{\\eta\}\\mathcal\{L\}\(\\theta\)\},\\quad\\mathcal\{Z\}=\\int\_\{U\}e^\{\-\\frac\{2\}\{\\eta\}\\mathcal\{L\}\(\\theta\)\}\\frac\{\\omega^\{K\}\}\{K\!\},\(12\.1\)which follows from the Fokker\-Planck equation when the time derivative vanishes, hence the distribution stabilizes, in the limit0=∇h⋅\(ρ∞\(θ\)∇hℒ\(θ\)\+η2∇hρ∞\(θ\)\)0=\\nabla\_\{h\}\\cdot\\left\(\\rho\_\{\\infty\}\(\\theta\)\\nabla\_\{h\}\\mathcal\{L\}\(\\theta\)\+\\frac\{\\eta\}\{2\}\\nabla\_\{h\}\\rho\_\{\\infty\}\(\\theta\)\\right\)\. This can be reformulated to yield an exponential using the log derivative definition and∇h\(logρ∞\(θ\)\)=−2η∇hℒ\(θ\)\\nabla\_\{h\}\(\\log\\rho\_\{\\infty\}\(\\theta\)\)=\-\\frac\{2\}\{\\eta\}\\nabla\_\{h\}\\mathcal\{L\}\(\\theta\)\. Define the Witten\-deformed Dolbeault operator acting on\(0,1\)\(0,1\)\-forms as
∂¯η=∂¯\+1η∂¯ℒ∧\.\\displaystyle\\overline\{\\partial\}\_\{\\eta\}=\\overline\{\\partial\}\+\\frac\{1\}\{\\eta\}\\overline\{\\partial\}\\mathcal\{L\}\\wedge\.\(12\.2\)Here,η\\etais the learning rate\. The above extension of the Dolbeault operator takes into account the loss and it is built in, whereas∂¯\\overline\{\\partial\}and its associated LaplacianΔ∂¯=∂¯∂¯†\+∂¯†∂¯\\Delta\_\{\\overline\{\\partial\}\}=\\overline\{\\partial\}\\overline\{\\partial\}^\{\\dagger\}\+\\overline\{\\partial\}^\{\\dagger\}\\overline\{\\partial\}are with respect to the manifold’s geometry but unassociated with the loss function\. We will need to work with the Laplacian, but the kernel of the Laplacian, i\.e\. a scalar harmonic function, are functions that are holomorphic\. On a compact Kähler manifold, Liouville’s theorem dictates that the only globally holomorphic functions are constants\. Therefore, to incorporate the effects of learning, we introduce the loss into the operator\. Acting on a formα\\alpha,
∂¯ηα=∂¯α\+1η∂¯ℒ∧α,\\displaystyle\\overline\{\\partial\}\_\{\\eta\}\\alpha=\\overline\{\\partial\}\\alpha\+\\frac\{1\}\{\\eta\}\\overline\{\\partial\}\\mathcal\{L\}\\wedge\\alpha,\(12\.3\)so a gradient term is built into the operator\. It is the "opposite of a directional derivative" since by∂¯ℒ∧∂¯ℒ=0\\overline\{\\partial\}\\mathcal\{L\}\\wedge\\overline\{\\partial\}\\mathcal\{L\}=0\. The division byη\\etais meant to match the definition ofρ∞\\rho\_\{\\infty\}in[12\.1](https://arxiv.org/html/2608.19584#S12.E1)\. We can note the form variety of[12\.3](https://arxiv.org/html/2608.19584#S12.E3)has an equivalent formulation
∂¯ηα=e−ℒ/η∂¯\(eℒ/ηα\)\.\\displaystyle\\overline\{\\partial\}\_\{\\eta\}\\alpha=e^\{\-\\mathcal\{L\}/\\eta\}\\overline\{\\partial\}\\left\(e^\{\\mathcal\{L\}/\\eta\}\\alpha\\right\)\.\(12\.4\)This follows since∂¯\(fα\)=\(∂¯f\)∧α\+f∂¯α\\overline\{\\partial\}\(f\\alpha\)=\(\\overline\{\\partial\}f\)\\wedge\\alpha\+f\\overline\{\\partial\}\\alphaand choose particularf=eℒ/ηf=e^\{\\mathcal\{L\}/\\eta\}\. Via chain rule,∂¯\(eℒ/η\)=eℒ/η∂¯\(ℒη\)=1ηeℒ/η∂¯ℒ\\overline\{\\partial\}\\left\(e^\{\\mathcal\{L\}/\\eta\}\\right\)=e^\{\\mathcal\{L\}/\\eta\}\\overline\{\\partial\}\\left\(\\frac\{\\mathcal\{L\}\}\{\\eta\}\\right\)=\\frac\{1\}\{\\eta\}e^\{\\mathcal\{L\}/\\eta\}\\overline\{\\partial\}\\mathcal\{L\}\. Therefore, we can see
∂¯\(eℒ/ηα\)=\(1ηeℒ/η∂¯ℒ\)∧α\+eℒ/η∂¯α\.\\displaystyle\\overline\{\\partial\}\\left\(e^\{\\mathcal\{L\}/\\eta\}\\alpha\\right\)=\\left\(\\frac\{1\}\{\\eta\}e^\{\\mathcal\{L\}/\\eta\}\\overline\{\\partial\}\\mathcal\{L\}\\right\)\\wedge\\alpha\+e^\{\\mathcal\{L\}/\\eta\}\\overline\{\\partial\}\\alpha\.\(12\.5\)and simplifying and canceling the exponential gives us[12\.3](https://arxiv.org/html/2608.19584#S12.E3)\. Now, the corresponding deformed Dolbeault Laplacian isΔη=∂¯η∂¯η†\+∂¯η†∂¯η\\Delta\_\{\\eta\}=\\overline\{\\partial\}\_\{\\eta\}\\overline\{\\partial\}\_\{\\eta\}^\{\\dagger\}\+\\overline\{\\partial\}\_\{\\eta\}^\{\\dagger\}\\overline\{\\partial\}\_\{\\eta\}\. Using the complex Bochner\-Weitzenböck identity \(with an operator formulation\), this deformed Laplacian expands into
Δη=Δ∂¯\+1η2‖∂¯ℒ‖h2\+1η∇ω1,1ℒ\+Ric\.\\displaystyle\\Delta\_\{\\eta\}=\\Delta\_\{\\overline\{\\partial\}\}\+\\frac\{1\}\{\\eta^\{2\}\}\\\|\\overline\{\\partial\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}\+\\frac\{1\}\{\\eta\}\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\+\\text\{Ric\}\.\(12\.6\)Now, we can notice the kernel of the Witten Laplacian corresponds to a steady\-state of the Fokker\-Planck equation\.
Claim\.Among allΨ∈kerΔη\\Psi\\in\\text\{ker\}\\Delta\_\{\\eta\}satisfying\|Ψ\|2=ρ∞\(θ\):=1𝒵e−2ℒ/η\|\\Psi\|^\{2\}=\\rho\_\{\\infty\}\(\\theta\):=\\frac\{1\}\{\\mathcal\{Z\}\}e^\{\-2\\mathcal\{L\}/\\eta\}pointwise, the solution is unique up to a constant unit\-modulus phase, and it is given byΨ0=1𝒵e−ℒ/η\\Psi\_\{0\}=\\frac\{1\}\{\\sqrt\{\\mathcal\{Z\}\}\}e^\{\-\\mathcal\{L\}/\\eta\}\.
Proof of claim\.By definitionΔη=∂¯η∂¯η†\+∂¯η†∂¯η\\Delta\_\{\\eta\}=\\overline\{\\partial\}\_\{\\eta\}\\overline\{\\partial\}\_\{\\eta\}^\{\\dagger\}\+\\overline\{\\partial\}\_\{\\eta\}^\{\\dagger\}\\overline\{\\partial\}\_\{\\eta\}\. LetΨ\\Psibe a smooth\(0,0\)\(0,0\)\-form\. Because there are no differential forms of negative degree, the adjoint applied to any\(0,0\)\(0,0\)\-form vanishes, meaning∂¯η†Ψ=0\\overline\{\\partial\}\_\{\\eta\}^\{\\dagger\}\\Psi=0\. Consequently,
ΔηΨ=∂¯η†∂¯ηΨ\.\\displaystyle\\Delta\_\{\\eta\}\\Psi=\\overline\{\\partial\}\_\{\\eta\}^\{\\dagger\}\\overline\{\\partial\}\_\{\\eta\}\\Psi\.\(12\.7\)To find the kernel, we must examineΔηΨ0=0\\Delta\_\{\\eta\}\\Psi\_\{0\}=0\. Let us examine the inner product
⟨Ψ0,ΔηΨ0⟩h=⟨Ψ0,∂¯η†∂¯ηΨ0⟩h\.\\displaystyle\\langle\\Psi\_\{0\},\\Delta\_\{\\eta\}\\Psi\_\{0\}\\rangle\_\{h\}=\\langle\\Psi\_\{0\},\\overline\{\\partial\}\_\{\\eta\}^\{\\dagger\}\\overline\{\\partial\}\_\{\\eta\}\\Psi\_\{0\}\\rangle\_\{h\}\.\(12\.8\)By definition of the adjoint,
⟨Ψ0,ΔηΨ0⟩h=⟨∂¯ηΨ0,∂¯ηΨ0⟩h=‖∂¯ηΨ0‖h2\.\\displaystyle\\langle\\Psi\_\{0\},\\Delta\_\{\\eta\}\\Psi\_\{0\}\\rangle\_\{h\}=\\langle\\overline\{\\partial\}\_\{\\eta\}\\Psi\_\{0\},\\overline\{\\partial\}\_\{\\eta\}\\Psi\_\{0\}\\rangle\_\{h\}=\\\|\\overline\{\\partial\}\_\{\\eta\}\\Psi\_\{0\}\\\|\_\{h\}^\{2\}\.\(12\.9\)We can note‖∂¯ηΨ0‖h2=0\\\|\\overline\{\\partial\}\_\{\\eta\}\\Psi\_\{0\}\\\|\_\{h\}^\{2\}=0if and only if∂¯ηΨ0=0\\overline\{\\partial\}\_\{\\eta\}\\Psi\_\{0\}=0\. By what we saw in[12\.4](https://arxiv.org/html/2608.19584#S12.E4), we can rewrite
e−ℒ/η∂¯\(eℒ/ηΨ0\)=0,\\displaystyle e^\{\-\\mathcal\{L\}/\\eta\}\\overline\{\\partial\}\\left\(e^\{\\mathcal\{L\}/\\eta\}\\Psi\_\{0\}\\right\)=0,\(12\.10\)and by positivity of the exponential, we must require∂¯\(eℒ/ηΨ0\)=0\\overline\{\\partial\}\\left\(e^\{\\mathcal\{L\}/\\eta\}\\Psi\_\{0\}\\right\)=0\. This implies that the scalar functionF\(θ\)=eℒ/ηΨ0F\(\\theta\)=e^\{\\mathcal\{L\}/\\eta\}\\Psi\_\{0\}is holomorphic\.
Now, we impose the defining condition of an exact pointwise match toρ∞\\rho\_\{\\infty\},
\|Ψ\(θ\)\|2=ρ∞\(θ\)=1𝒵e−2ℒ\(θ\)/ηfor allθ∈U,\\displaystyle\|\\Psi\(\\theta\)\|^\{2\}=\\rho\_\{\\infty\}\(\\theta\)=\\frac\{1\}\{\\mathcal\{Z\}\}e^\{\-2\\mathcal\{L\}\(\\theta\)/\\eta\}\\quad\\text\{for all \}\\theta\\in U,\(12\.11\)we obtain for the magnitude ofF\(θ\)F\(\\theta\)
\|F\(θ\)\|=\|eℒ/ηΨ\(θ\)\|=eℒ/η\|Ψ\(θ\)\|=eℒ/η⋅1𝒵e−ℒ/η=1𝒵,\\displaystyle\|F\(\\theta\)\|=\\left\|e^\{\\mathcal\{L\}/\\eta\}\\Psi\(\\theta\)\\right\|=e^\{\\mathcal\{L\}/\\eta\}\|\\Psi\(\\theta\)\|=e^\{\\mathcal\{L\}/\\eta\}\\cdot\\frac\{1\}\{\\sqrt\{\\mathcal\{Z\}\}\}e^\{\-\\mathcal\{L\}/\\eta\}=\\frac\{1\}\{\\sqrt\{\\mathcal\{Z\}\}\},\(12\.12\)which is a constant independent ofθ\\theta\. Thus,FFis holomorphic onUUwith a constant modulus\. By the open mapping theorem \(since a non\-constant holomorphic function on a connected domain maps open sets to open sets, its image cannot lie entirely on a circle\|w\|=1/𝒵\|w\|=1/\\sqrt\{\\mathcal\{Z\}\}\),FFmust be a constant function,F\(θ\)≡CF\(\\theta\)\\equiv Cwith\|C\|=1/𝒵\|C\|=1/\\sqrt\{\\mathcal\{Z\}\}\. This proves the claim\.
□\\square
Hence,
V∞\\displaystyle V\_\{\\infty\}=𝔼\[‖θ−θ∗‖h2\]\\displaystyle=\\mathbb\{E\}\\left\[\\\|\\theta\-\\theta^\{\*\}\\\|\_\{h\}^\{2\}\\right\]\(12\.13\)=∫U‖θ−θ∗‖h2ρ∞\(θ\)ωKK\!\\displaystyle=\\int\_\{U\}\\\|\\theta\-\\theta^\{\*\}\\\|\_\{h\}^\{2\}\\rho\_\{\\infty\}\(\\theta\)\\frac\{\\omega^\{K\}\}\{K\!\}\(12\.14\)=∫U‖θ−θ∗‖h2\|Ψ0\|2ωKK\!\.\\displaystyle=\\int\_\{U\}\\\|\\theta\-\\theta^\{\*\}\\\|\_\{h\}^\{2\}\|\\Psi\_\{0\}\|^\{2\}\\frac\{\\omega^\{K\}\}\{K\!\}\.\(12\.15\)Now we attempt to evaluate the integral\. Via the normal coordinates trick of[11\.3](https://arxiv.org/html/2608.19584#S11.SS3), we arrive at again
ωKK\!=\(1−Rickl¯ξkξ¯l\+𝒪\(‖ξ‖3\)\)ω0KK\!\\displaystyle\\frac\{\\omega^\{K\}\}\{K\!\}=\\left\(1\-\\text\{Ric\}\_\{k\\overline\{l\}\}\\xi^\{k\}\\overline\{\\xi\}^\{l\}\+\\mathcal\{O\}\(\\\|\\xi\\\|^\{3\}\)\\right\)\\frac\{\\omega\_\{0\}^\{K\}\}\{K\!\}\(12\.16\)forξ=θ−θ∗\\xi=\\theta\-\\theta^\{\*\}\. Via Taylor expansion, and sinceθ∗\\theta^\{\*\}is critical,
⟹ℒ\(ξ\)=ℒ\(0\)\+∇ij¯1,1ℒ\(0\)ξiξ¯j\+Re\(∇ij2,0ℒ\(0\)ξiξj\)\+𝒪\(‖ξ‖3\)\\displaystyle\\implies\\mathcal\{L\}\(\\xi\)=\\mathcal\{L\}\(0\)\+\\nabla\_\{i\\overline\{j\}\}^\{1,1\}\\mathcal\{L\}\(0\)\\xi^\{i\}\\overline\{\\xi\}^\{j\}\+\\text\{Re\}\\left\(\\nabla\_\{ij\}^\{2,0\}\\mathcal\{L\}\(0\)\\xi^\{i\}\\xi^\{j\}\\right\)\+\\mathcal\{O\}\(\\\|\\xi\\\|^\{3\}\)\(12\.17\)⟹×−2ηandexpe−2ηℒ\(ξ\)=e−2ηℒ\(0\)exp\(−2η∇ij¯1,1ℒ\(0\)ξiξ¯j−2ηRe\(∇ij2,0ℒ\(0\)ξiξj\)\)\+𝒪\(‖ξ‖3\)\.\\displaystyle\\stackrel\{\{\\scriptstyle\\times\-\\frac\{2\}\{\\eta\}\\ \\text\{and\}\\ \\exp\}\}\{\{\\implies\}\}e^\{\-\\frac\{2\}\{\\eta\}\\mathcal\{L\}\(\\xi\)\}=e^\{\-\\frac\{2\}\{\\eta\}\\mathcal\{L\}\(0\)\}\\exp\\left\(\-\\frac\{2\}\{\\eta\}\\nabla\_\{i\\overline\{j\}\}^\{1,1\}\\mathcal\{L\}\(0\)\\xi^\{i\}\\overline\{\\xi\}^\{j\}\-\\frac\{2\}\{\\eta\}\\text\{Re\}\\left\(\\nabla\_\{ij\}^\{2,0\}\\mathcal\{L\}\(0\)\\xi^\{i\}\\xi^\{j\}\\right\)\\right\)\+\\mathcal\{O\}\(\\\|\\xi\\\|^\{3\}\)\.\(12\.18\)Combining with the volume forms of[12\.16](https://arxiv.org/html/2608.19584#S12.E16),
e−2ηℒ\(ξ\)ωKK\!=e−2ηℒ\(0\)\\displaystyle e^\{\-\\frac\{2\}\{\\eta\}\\mathcal\{L\}\(\\xi\)\}\\frac\{\\omega^\{K\}\}\{K\!\}=e^\{\-\\frac\{2\}\{\\eta\}\\mathcal\{L\}\(0\)\}\(12\.19\)×exp\(−2η∇ij¯1,1ℒ\(0\)ξiξ¯j−2ηRe\(∇ij2,0ℒ\(0\)ξiξj\)\)\(1−Ricij¯\(0\)ξiξ¯j\)⏟=exp\(−\[2η∇ij¯1,1ℒ\(0\)\+Ricij¯\(0\)\]ξiξ¯j−2ηRe\(∇ij2,0ℒ\(0\)ξiξj\)\+𝒪\(‖ξ‖4\)\)ω0KK\!\+𝒪\(‖ξ‖3\)\.\\displaystyle\\quad\\quad\\times\\underbrace\{\\exp\\left\(\-\\frac\{2\}\{\\eta\}\\nabla\_\{i\\overline\{j\}\}^\{1,1\}\\mathcal\{L\}\(0\)\\xi^\{i\}\\overline\{\\xi\}^\{j\}\-\\frac\{2\}\{\\eta\}\\text\{Re\}\\left\(\\nabla\_\{ij\}^\{2,0\}\\mathcal\{L\}\(0\)\\xi^\{i\}\\xi^\{j\}\\right\)\\right\)\\left\(1\-\\text\{Ric\}\_\{i\\overline\{j\}\}\(0\)\\xi^\{i\}\\overline\{\\xi\}^\{j\}\\right\)\}\_\{=\\exp\\left\(\-\\left\[\\frac\{2\}\{\\eta\}\\nabla\_\{i\\overline\{j\}\}^\{1,1\}\\mathcal\{L\}\(0\)\+\\text\{Ric\}\_\{i\\overline\{j\}\}\(0\)\\right\]\\xi^\{i\}\\overline\{\\xi\}^\{j\}\-\\frac\{2\}\{\\eta\}\\text\{Re\}\\left\(\\nabla\_\{ij\}^\{2,0\}\\mathcal\{L\}\(0\)\\xi^\{i\}\\xi^\{j\}\\right\)\+\\mathcal\{O\}\(\\\|\\xi\\\|^\{4\}\)\\right\)\}\\frac\{\\omega\_\{0\}^\{K\}\}\{K\!\}\+\\mathcal\{O\}\(\\\|\\xi\\\|^\{3\}\)\.\(12\.20\)We have used1−x=e−x\+𝒪\(x2\)1\-x=e^\{\-x\}\+\\mathcal\{O\}\(x^\{2\}\)\. Define the vectorξ~=\[ξξ¯\]\\widetilde\{\\xi\}=\\begin\{bmatrix\}\\xi\\\\ \\overline\{\\xi\}\\end\{bmatrix\}\. Using the identityRe\(z\)=12\(z\+z¯\)\\text\{Re\}\(z\)=\\frac\{1\}\{2\}\(z\+\\overline\{z\}\), we can rewrite the exponential argument as a canonical quadratic form−12ξ~†𝒫ξ~\-\\frac\{1\}\{2\}\\widetilde\{\\xi\}^\{\\dagger\}\\mathcal\{P\}\\widetilde\{\\xi\}, where the augmented block precision matrix𝒫\\mathcal\{P\}is given by
𝒫=2η\[∇ω1,1ℒ\(θ∗\)\+η2Ric\(θ∗\)∇2,0¯ℒ\(θ∗\)∇ω2,0ℒ\(θ∗\)∇1,1¯ℒ\(θ∗\)\+η2Ric¯\(θ∗\)\]\.\\displaystyle\\mathcal\{P\}=\\frac\{2\}\{\\eta\}\\begin\{bmatrix\}\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\(\\theta^\{\*\}\)\+\\frac\{\\eta\}\{2\}\\text\{Ric\}\(\\theta^\{\*\}\)&\\overline\{\\nabla^\{2,0\}\}\\mathcal\{L\}\(\\theta^\{\*\}\)\\\\ \\nabla^\{2,0\}\_\{\\omega\}\\mathcal\{L\}\(\\theta^\{\*\}\)&\\overline\{\\nabla^\{1,1\}\}\\mathcal\{L\}\(\\theta^\{\*\}\)\+\\frac\{\\eta\}\{2\}\\overline\{\\text\{Ric\}\}\(\\theta^\{\*\}\)\\end\{bmatrix\}\.\(12\.21\)Returning to[12\.13](https://arxiv.org/html/2608.19584#S12.E13), we can note the asymptotic equivalence overUU
V∞=∫U‖ξ‖h2exp\(−12ξ~†𝒫ξ~\)ω0KK\!∫Uexp\(−12ξ~†𝒫ξ~\)ω0KK\!\+𝒪\(η2\)\.\\displaystyle V\_\{\\infty\}=\\frac\{\\int\_\{U\}\\\|\\xi\\\|\_\{h\}^\{2\}\\exp\\left\(\-\\frac\{1\}\{2\}\\widetilde\{\\xi\}^\{\\dagger\}\\mathcal\{P\}\\widetilde\{\\xi\}\\right\)\\frac\{\\omega\_\{0\}^\{K\}\}\{K\!\}\}\{\\int\_\{U\}\\exp\\left\(\-\\frac\{1\}\{2\}\\widetilde\{\\xi\}^\{\\dagger\}\\mathcal\{P\}\\widetilde\{\\xi\}\\right\)\\frac\{\\omega\_\{0\}^\{K\}\}\{K\!\}\}\+\\mathcal\{O\}\(\\eta^\{2\}\)\.\(12\.22\)The numerator is equivalent to evaluatingΣeff=𝔼\[ξξ†\]\\Sigma\_\{\\text\{eff\}\}=\\mathbb\{E\}\[\\xi\\xi^\{\\dagger\}\], and the denominator is a normalization constant\. To make the distribution to be represented via an expected value, we can absorb this denominator into the expected value implicitly\. We have the relation𝔼\[ξ~ξ~†\]=𝒫−1=Σaug\\mathbb\{E\}\[\\widetilde\{\\xi\}\\widetilde\{\\xi\}^\{\\dagger\}\]=\\mathcal\{P\}^\{\-1\}=\\Sigma\_\{\\text\{aug\}\}, since
Σaug=𝔼\[\[ξξ¯\]\[ξ†ξT\]\]=\[𝔼\[ξξ†\]𝔼\[ξξT\]𝔼\[ξ¯ξ†\]𝔼\[ξ¯ξT\]\]\.\\displaystyle\\Sigma\_\{\\text\{aug\}\}=\\mathbb\{E\}\\left\[\\begin\{bmatrix\}\\xi\\\\ \\overline\{\\xi\}\\end\{bmatrix\}\\begin\{bmatrix\}\\xi^\{\\dagger\}&\\xi^\{T\}\\end\{bmatrix\}\\right\]=\\begin\{bmatrix\}\\mathbb\{E\}\[\\xi\\xi^\{\\dagger\}\]&\\mathbb\{E\}\[\\xi\\xi^\{T\}\]\\\\ \\mathbb\{E\}\[\\overline\{\\xi\}\\xi^\{\\dagger\}\]&\\mathbb\{E\}\[\\overline\{\\xi\}\\xi^\{T\}\]\\end\{bmatrix\}\.\(12\.23\)By definition ofV∞V\_\{\\infty\}, we are only interested in the top\-left block\. This means the numerator evaluation corresponds to the top\-leftK×KK\\times Kblock of the augmented covariance matrixΣaug=𝒫−1\\Sigma\_\{\\text\{aug\}\}=\\mathcal\{P\}^\{\-1\}\. By inverting𝒫\\mathcal\{P\}, we find this top\-left block to be\(A−BD−1C\)−1\(A\-BD^\{\-1\}C\)^\{\-1\}, which is
Σeff=η2\[\(∇ω1,1ℒ\(θ∗\)\+η2Ric\(θ∗\)\)−∇ω2,0¯ℒ\(θ∗\)\(∇ω1,1¯ℒ\(θ∗\)\+η2Ric¯\(θ∗\)\)−1∇ω2,0ℒ\(θ∗\)\]−1\.\\displaystyle\\Sigma\_\{\\text\{eff\}\}=\\frac\{\\eta\}\{2\}\\left\[\\left\(\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\(\\theta^\{\*\}\)\+\\frac\{\\eta\}\{2\}\\text\{Ric\}\(\\theta^\{\*\}\)\\right\)\-\\overline\{\\nabla\_\{\\omega\}^\{2,0\}\}\\mathcal\{L\}\(\\theta^\{\*\}\)\\left\(\\overline\{\\nabla^\{1,1\}\_\{\\omega\}\}\\mathcal\{L\}\(\\theta^\{\*\}\)\+\\frac\{\\eta\}\{2\}\\overline\{\\text\{Ric\}\}\(\\theta^\{\*\}\)\\right\)^\{\-1\}\\nabla\_\{\\omega\}^\{2,0\}\\mathcal\{L\}\(\\theta^\{\*\}\)\\right\]^\{\-1\}\.\(12\.24\)The expected value of‖ξ‖h2\\\|\\xi\\\|\_\{h\}^\{2\}under the \(local\) metric is the trace of this covariance matrix,𝔼\[‖ξ‖h2\]=Trh\(Σeff\)\\mathbb\{E\}\[\\\|\\xi\\\|\_\{h\}^\{2\}\]=\\text\{Tr\}\_\{h\}\(\\Sigma\_\{\\text\{eff\}\}\)\. Therefore, we conclude
V∞=η2Trh\(\[∇ω1,1ℒ\+η2Ric−∇ω2,0¯ℒ\(∇ω1,1¯ℒ\+η2Ric¯\)−1∇ω2,0ℒ\]−1\)\|θ∗\+𝒪\(η2\)\.\\displaystyle V\_\{\\infty\}=\\frac\{\\eta\}\{2\}\\text\{Tr\}\_\{h\}\\left\(\\left\[\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\+\\frac\{\\eta\}\{2\}\\text\{Ric\}\-\\overline\{\\nabla\_\{\\omega\}^\{2,0\}\}\\mathcal\{L\}\\left\(\\overline\{\\nabla^\{1,1\}\_\{\\omega\}\}\\mathcal\{L\}\+\\frac\{\\eta\}\{2\}\\overline\{\\text\{Ric\}\}\\right\)^\{\-1\}\\nabla\_\{\\omega\}^\{2,0\}\\mathcal\{L\}\\right\]^\{\-1\}\\right\)\\Bigg\|\_\{\\theta^\{\*\}\}\+\\mathcal\{O\}\(\\eta^\{2\}\)\.\(12\.25\)
□\\square
Remark\.In the Calabi\-Yau case, we get
V∞,CY=η2Trh\(\[∇ω1,1ℒ−∇ω2,0¯ℒ\(∇ω1,1¯ℒ\)−1∇ω2,0ℒ\]−1\)\|θ∗\+𝒪\(η2\)\.\\displaystyle V\_\{\\infty,\\text\{CY\}\}=\\frac\{\\eta\}\{2\}\\text\{Tr\}\_\{h\}\\left\(\\left\[\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\-\\overline\{\\nabla\_\{\\omega\}^\{2,0\}\}\\mathcal\{L\}\\left\(\\overline\{\\nabla^\{1,1\}\_\{\\omega\}\}\\mathcal\{L\}\\right\)^\{\-1\}\\nabla\_\{\\omega\}^\{2,0\}\\mathcal\{L\}\\right\]^\{\-1\}\\right\)\\Bigg\|\_\{\\theta^\{\*\}\}\+\\mathcal\{O\}\(\\eta^\{2\}\)\.\(12\.26\)Moreover, we also get simplifications in the proofωKK\!=ω0KK\!\+𝒪\(‖ξ‖3\)\\frac\{\\omega^\{K\}\}\{K\!\}=\\frac\{\\omega\_\{0\}^\{K\}\}\{K\!\}\+\\mathcal\{O\}\(\\\|\\xi\\\|^\{3\}\)and
𝒫=2η\[∇ω1,1ℒ\(θ∗\)∇2,0¯ℒ\(θ∗\)∇ω2,0ℒ\(θ∗\)∇1,1¯ℒ\(θ∗\)\]\.\\displaystyle\\mathcal\{P\}=\\frac\{2\}\{\\eta\}\\begin\{bmatrix\}\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\(\\theta^\{\*\}\)&\\overline\{\\nabla^\{2,0\}\}\\mathcal\{L\}\(\\theta^\{\*\}\)\\\\ \\nabla^\{2,0\}\_\{\\omega\}\\mathcal\{L\}\(\\theta^\{\*\}\)&\\overline\{\\nabla^\{1,1\}\}\\mathcal\{L\}\(\\theta^\{\*\}\)\\end\{bmatrix\}\.\(12\.27\)Generally, high variance, meaning highV∞V\_\{\\infty\}is undesirable, since we more closely desire convergence to the local minima\. If we define,
\{𝒜general=∇ω1,1ℒ\+η2Ric−∇ω2,0¯ℒ\(∇ω1,1¯ℒ\+η2Ric¯\)−1∇ω2,0ℒ𝒜CY=∇ω1,1ℒ−∇ω2,0¯ℒ\(∇ω1,1¯ℒ\)−1∇ω2,0ℒ\\displaystyle\\begin\{cases\}&\\mathcal\{A\}\_\{\\text\{general\}\}=\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\+\\frac\{\\eta\}\{2\}\\text\{Ric\}\-\\overline\{\\nabla\_\{\\omega\}^\{2,0\}\}\\mathcal\{L\}\\left\(\\overline\{\\nabla^\{1,1\}\_\{\\omega\}\}\\mathcal\{L\}\+\\frac\{\\eta\}\{2\}\\overline\{\\text\{Ric\}\}\\right\)^\{\-1\}\\nabla\_\{\\omega\}^\{2,0\}\\mathcal\{L\}\\\\ &\\mathcal\{A\}\_\{\\text\{CY\}\}=\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\-\\overline\{\\nabla\_\{\\omega\}^\{2,0\}\}\\mathcal\{L\}\\left\(\\overline\{\\nabla^\{1,1\}\_\{\\omega\}\}\\mathcal\{L\}\\right\)^\{\-1\}\\nabla\_\{\\omega\}^\{2,0\}\\mathcal\{L\}\\end\{cases\}\(12\.28\)and we can see the relation
𝒜general≻𝒜CYwhenRic≻0\.\\displaystyle\\mathcal\{A\}\_\{\\text\{general\}\}\\succ\\mathcal\{A\}\_\{\\text\{CY\}\}\\quad\\text\{when \}\\text\{Ric\}\\succ 0\.\(12\.29\)We can note via matrix inversion𝒜−1\\mathcal\{A\}^\{\-1\}decreases trace as𝒜\\mathcal\{A\}increases trace\. A positive Ricci curvature acts as a "restoring force" that helps the descent path more closely fall into the optimum\. From this, we can note a positive Ricci curvature is most desirable\. In the Ricci\-flat scenario,ρ∞\\rho\_\{\\infty\}is determined solely by the Hessians, and there is no contribution to the trace\. Even more so, we can note a negatively curved space contributes most adversely to this trace value, and results in the highest variance\. Thus, a Calabi\-Yau manifold sits at the interface of the more and less desirable cases\.
### 12\.2Minimal eigenvalue collapse under negative Ricci curvature
Proof of Lemma 13\.LetU⊆MU\\subseteq Mbe an open set andK⊆UK\\subseteq Ua compact subset\. DenoteΔη=∂¯η∂¯η†\+∂¯η†∂¯η\\Delta\_\{\\eta\}=\\overline\{\\partial\}\_\{\\eta\}\\overline\{\\partial\}\_\{\\eta\}^\{\\dagger\}\+\\overline\{\\partial\}\_\{\\eta\}^\{\\dagger\}\\overline\{\\partial\}\_\{\\eta\}as in Appendix[12\.1](https://arxiv.org/html/2608.19584#S12.SS1)\. We bound the minimum eigenvalueλmin\\lambda\_\{\\text\{min\}\}\. The minimum eigenvalue follows a Rayleigh quotient over all valid forms by the min\-max theorem, or Rayleigh principle\. Let us define the test form
α=χ∂¯ℒe−ℒ/η\.\\displaystyle\\alpha=\\chi\\overline\{\\partial\}\\mathcal\{L\}e^\{\-\\mathcal\{L\}/\\eta\}\.\(12\.30\)whereχ∈C∞\\chi\\in C^\{\\infty\}is a smooth bump function similar to a compactly\-supported mollifier, andχ\|K=1,χ=0\\chi\|\_\{K\}=1,\\chi=0outside ofUU, and0<χ\(x\)<10<\\chi\(x\)<1in the annular regionU∖KU\\setminus K\. Sinceα\\alphais not necessarily optimal, we get the inequality
λ1≤⟨α,Δηα⟩h‖α‖h2\.\\displaystyle\\lambda\_\{1\}\\leq\\frac\{\\langle\\alpha,\\Delta\_\{\\eta\}\\alpha\\rangle\_\{h\}\}\{\\\|\\alpha\\\|\_\{h\}^\{2\}\}\.\(12\.31\)
Claim\.A modified Bochner identity gives
⟨α,Δηα⟩h=∫U\(‖∇1,0α‖h2\+1η2‖∂¯ℒ‖h2‖α‖h2\+Ric\(α,α¯\)\+1η⟨α,\(2∇ω1,1ℒ−Δ∂¯ℒ\)α⟩h\)ωKK\!\.\\displaystyle\\langle\\alpha,\\Delta\_\{\\eta\}\\alpha\\rangle\_\{h\}=\\int\_\{U\}\\left\(\\\|\\nabla^\{1,0\}\\alpha\\\|\_\{h\}^\{2\}\+\\frac\{1\}\{\\eta^\{2\}\}\\\|\\overline\{\\partial\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}\\\|\\alpha\\\|\_\{h\}^\{2\}\+\\text\{Ric\}\(\\alpha,\\overline\{\\alpha\}\)\+\\frac\{1\}\{\\eta\}\\langle\\alpha,\(2\\nabla\_\{\\omega\}^\{1,1\}\\mathcal\{L\}\-\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}\)\\alpha\\rangle\_\{h\}\\right\)\\frac\{\\omega^\{K\}\}\{K\!\}\.\(12\.32\)
We will prove the claim after our main result\. Using our test formα=χ∂¯ℒe−ℒ/η\\alpha=\\chi\\overline\{\\partial\}\\mathcal\{L\}e^\{\-\\mathcal\{L\}/\\eta\}, and since Ricci curvature is a bilinear form,Ric\(∇ℒ,∇¯ℒ\)≤−κ‖∇ℒ‖h2\\text\{Ric\}\(\\nabla\\mathcal\{L\},\\overline\{\\nabla\}\\mathcal\{L\}\)\\leq\-\\kappa\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{2\}\. Here, we have assumed negative and bounded Ricci curvature\. Therefore, the numerator of the Rayleigh quotient[12\.31](https://arxiv.org/html/2608.19584#S12.E31)is bounded by
⟨α,Δηα⟩h\\displaystyle\\langle\\alpha,\\Delta\_\{\\eta\}\\alpha\\rangle\_\{h\}≤∫Uχ2e−2ℒ/η\(∥∇1,0\(∂¯ℒ\)−1η∂ℒ⊗∂¯ℒ∥h2\\displaystyle\\leq\\int\_\{U\}\\chi^\{2\}e^\{\-2\\mathcal\{L\}/\\eta\}\\Bigg\(\\\|\\nabla^\{1,0\}\(\\overline\{\\partial\}\\mathcal\{L\}\)\-\\frac\{1\}\{\\eta\}\\partial\\mathcal\{L\}\\otimes\\overline\{\\partial\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}\(12\.33\)OPEN\+1η2‖∇ℒ‖h4\+1η⟨α,\(2∇ω1,1ℒ−Δ∂¯ℒ\)α⟩h−κ‖∇ℒ‖h2\)ωKK\!\+ℰ\(∇χ\)\.\\displaystyle\\quad\+\\frac\{1\}\{\\eta^\{2\}\}\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{4\}\+\\frac\{1\}\{\\eta\}\\langle\\alpha,\(2\\nabla\_\{\\omega\}^\{1,1\}\\mathcal\{L\}\-\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}\)\\alpha\\rangle\_\{h\}\-\\kappa\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{2\}\\Bigg\)\\frac\{\\omega^\{K\}\}\{K\!\}\+\\mathcal\{E\}\(\\nabla\\chi\)\.\(12\.34\)Hereℰ\(∇χ\)\\mathcal\{E\}\(\\nabla\\chi\)is an annular gradient term\. The denominator of the Rayleigh quotient[12\.31](https://arxiv.org/html/2608.19584#S12.E31)is
∥α∥h2=∫Uχ2e−2ℒ/η∥∇ℒ∥h2ωKK\!\.\\displaystyle\\\|\\alpha\\\|\_\{h\}^\{2\}=\\int\_\{U\}\\chi^\{2\}e^\{\-2\\mathcal\{L\}/\\eta\}\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{2\}\\frac\{\\omega^\{K\}\}\{K\!\}\.\(12\.35\)The integral of the last term cancels
∫Uχ2e−2ℒ/η\(−κ∥∇ℒ∥h2\)ωKK\!∫Uχ2e−2ℒ/η∥∇ℒ∥h2ωKK\!=−κ∫Uχ2e−2ℒ/η∥∇ℒ∥h2ωKK\!∫Uχ2e−2ℒ/η∥∇ℒ∥h2ωKK\!=−κ\.\\displaystyle\\frac\{\\int\_\{U\}\\chi^\{2\}e^\{\-2\\mathcal\{L\}/\\eta\}\(\-\\kappa\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{2\}\)\\frac\{\\omega^\{K\}\}\{K\!\}\}\{\\int\_\{U\}\\chi^\{2\}e^\{\-2\\mathcal\{L\}/\\eta\}\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{2\}\\frac\{\\omega^\{K\}\}\{K\!\}\}=\-\\kappa\\cancel\{\\frac\{\\int\_\{U\}\\chi^\{2\}e^\{\-2\\mathcal\{L\}/\\eta\}\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{2\}\\frac\{\\omega^\{K\}\}\{K\!\}\}\{\\int\_\{U\}\\chi^\{2\}e^\{\-2\\mathcal\{L\}/\\eta\}\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{2\}\\frac\{\\omega^\{K\}\}\{K\!\}\}\}=\-\\kappa\.\(12\.36\)We bound the middle term\. The quadratic form of the \(1,1\)\-Hessian acting on the gradient is bounded by its maximum eigenvalue
⟨∂¯ℒ,\(∇ω1,1ℒ\)∂¯ℒ⟩h≤‖∇ω1,1ℒ‖2‖∂¯ℒ‖h2\.\\displaystyle\\langle\\overline\{\\partial\}\\mathcal\{L\},\(\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\)\\overline\{\\partial\}\\mathcal\{L\}\\rangle\_\{h\}\\leq\\\|\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\\\|\_\{2\}\\\|\\overline\{\\partial\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}\.\(12\.37\)Letβ1,1=supθ∈U‖∇ω1,1ℒ‖2\\beta\_\{1,1\}=\\sup\_\{\\theta\\in U\}\\\|\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\\\|\_\{2\}\. Therefore, we examine the middle term and sinceΔ∂¯ℒ=Tr\(∇ω1,1ℒ\)\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}=\\text\{Tr\}\(\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\)
1η∫Uχ2e−2ℒ/η⟨∂¯ℒ,\(2∇1,1ωℒ−Δ∂¯ℒ\)∂¯ℒ⟩hωKK\!∫Uχ2e−2ℒ/η∥∇ℒ∥h2ωKK\!≤\(2\+K\)β1,1η∫Uχ2e−2ℒ/η∥∇ℒ∥h2ωKK\!∫Uχ2e−2ℒ/η∥∇ℒ∥h2ωKK\!=\(2\+K\)β1,1η\.\\displaystyle\\frac\{1\}\{\\eta\}\\frac\{\\int\_\{U\}\\chi^\{2\}e^\{\-2\\mathcal\{L\}/\\eta\}\\langle\\overline\{\\partial\}\\mathcal\{L\},\(2\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\-\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}\)\\overline\{\\partial\}\\mathcal\{L\}\\rangle\_\{h\}\\frac\{\\omega^\{K\}\}\{K\!\}\}\{\\int\_\{U\}\\chi^\{2\}e^\{\-2\\mathcal\{L\}/\\eta\}\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{2\}\\frac\{\\omega^\{K\}\}\{K\!\}\}\\leq\\frac\{\(2\+K\)\\beta\_\{1,1\}\}\{\\eta\}\\cancel\{\\frac\{\\int\_\{U\}\\chi^\{2\}e^\{\-2\\mathcal\{L\}/\\eta\}\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{2\}\\frac\{\\omega^\{K\}\}\{K\!\}\}\{\\int\_\{U\}\\chi^\{2\}e^\{\-2\\mathcal\{L\}/\\eta\}\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{2\}\\frac\{\\omega^\{K\}\}\{K\!\}\}\}=\\frac\{\(2\+K\)\\beta\_\{1,1\}\}\{\\eta\}\.\(12\.38\)Now we examine the‖∇1,0\(∂¯ℒ\)−1η∂ℒ⊗∂¯ℒ‖h2\+1η2‖∇ℒ‖h4\\\|\\nabla^\{1,0\}\(\\overline\{\\partial\}\\mathcal\{L\}\)\-\\frac\{1\}\{\\eta\}\\partial\\mathcal\{L\}\\otimes\\overline\{\\partial\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}\+\\frac\{1\}\{\\eta^\{2\}\}\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{4\}term,
∫Uχ2e−2ℒ/η\(∥∇1,0\(∂¯ℒ\)−1η∂ℒ⊗∂¯ℒ∥h2\+1η2∥∇ℒ∥h4\)ωKK\!∫Uχ2e−2ℒ/η∥∇ℒ∥h2ωKK\!,\\displaystyle\\frac\{\\int\_\{U\}\\chi^\{2\}e^\{\-2\\mathcal\{L\}/\\eta\}\(\\\|\\nabla^\{1,0\}\(\\overline\{\\partial\}\\mathcal\{L\}\)\-\\frac\{1\}\{\\eta\}\\partial\\mathcal\{L\}\\otimes\\overline\{\\partial\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}\+\\frac\{1\}\{\\eta^\{2\}\}\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{4\}\)\\frac\{\\omega^\{K\}\}\{K\!\}\}\{\\int\_\{U\}\\chi^\{2\}e^\{\-2\\mathcal\{L\}/\\eta\}\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{2\}\\frac\{\\omega^\{K\}\}\{K\!\}\},\(12\.39\)which is nontrivial to bound\. We will use Laplace’s method
∫h\(x\)eMg\(x\)𝑑x≈\(2πM\)n/21\|det∇2g\(x0\)\|h\(x0\)eMg\(x0\)\.\\displaystyle\\int h\(x\)e^\{Mg\(x\)\}dx\\approx\\left\(\\frac\{2\\pi\}\{M\}\\right\)^\{n/2\}\\frac\{1\}\{\\sqrt\{\|\\det\\nabla^\{2\}g\(x\_\{0\}\)\|\}\}h\(x\_\{0\}\)e^\{Mg\(x\_\{0\}\)\}\.\(12\.40\)The above is not manifold\-valued\. Instead, we can shrink the domain and locally use a flat approximation via the above\. Laplace’s method is an asymptotic localizer that ignored the topology and local curvature of the space, reducing the problem to flat Euclidean or Hermitian geometry on the tangent space[42](https://arxiv.org/html/2608.19584#bib.bib79)\. By the triangle inequality, and noting that‖∂ℒ⊗∂¯ℒ‖h=‖∂ℒ‖h‖∂¯ℒ‖h=‖∇ℒ‖h2\\\|\\partial\\mathcal\{L\}\\otimes\\overline\{\\partial\}\\mathcal\{L\}\\\|\_\{h\}=\\\|\\partial\\mathcal\{L\}\\\|\_\{h\}\\\|\\overline\{\\partial\}\\mathcal\{L\}\\\|\_\{h\}=\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{2\}, we bound
‖∇1,0\(∂¯ℒ\)−1η∂ℒ⊗∂¯ℒ‖h≤‖∇1,0\(∂¯ℒ\)‖h\+1η‖∇ℒ‖h2\.\\displaystyle\\left\\\|\\nabla^\{1,0\}\(\\overline\{\\partial\}\\mathcal\{L\}\)\-\\frac\{1\}\{\\eta\}\\partial\\mathcal\{L\}\\otimes\\overline\{\\partial\}\\mathcal\{L\}\\right\\\|\_\{h\}\\leq\\\|\\nabla^\{1,0\}\(\\overline\{\\partial\}\\mathcal\{L\}\)\\\|\_\{h\}\+\\frac\{1\}\{\\eta\}\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{2\}\.\(12\.41\)Attempting to bound via evaluating suprema will result in a diverging bound around a critical bound\. Instead, we examine the‖∇1,0\(∂¯ℒ\)−1η∂ℒ⊗∂¯ℒ‖h2\+1η2‖∇ℒ‖h4\\\|\\nabla^\{1,0\}\(\\overline\{\\partial\}\\mathcal\{L\}\)\-\\frac\{1\}\{\\eta\}\\partial\\mathcal\{L\}\\otimes\\overline\{\\partial\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}\+\\frac\{1\}\{\\eta^\{2\}\}\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{4\}term via Laplace asymptotics, using[12\.40](https://arxiv.org/html/2608.19584#S12.E40)\. Assumeℒ\\mathcal\{L\}has a unique minimumθ0∈U\\theta\_\{0\}\\in U, withθ0∈K\\theta\_\{0\}\\in Kso thatχ\(θ0\)=1\\chi\(\\theta\_\{0\}\)=1, and letH:=Hessℒ\(θ0\)≻0H:=\\mathrm\{Hess\}\\mathcal\{L\}\(\\theta\_\{0\}\)\\succ 0denote the real Hessian of dimension2K×2K2K\\times 2K\. With the integrals in[12\.39](https://arxiv.org/html/2608.19584#S12.E39), we will utilize Taylor expansions\. Denoteggthe2K×2K2K\\times 2Krealification ofhh\. In normal coordinatesξ~=θ−θ0∈ℂK\\widetilde\{\\xi\}=\\theta\-\\theta\_\{0\}\\in\\mathbb\{C\}^\{K\},ξ=\(ξ~ξ~¯\)∈ℂ2K\\xi=\\begin\{pmatrix\}\\widetilde\{\\xi\}\\\\ \\overline\{\\widetilde\{\\xi\}\}\\end\{pmatrix\}\\in\\mathbb\{C\}^\{2K\},
ℒ\(θ0\+ξ\)\\displaystyle\\mathcal\{L\}\(\\theta\_\{0\}\+\\xi\)=ℒ\(θ0\)\+12ξTHξ\+𝒪\(‖ξ‖3\),\\displaystyle=\\mathcal\{L\}\(\\theta\_\{0\}\)\+\\tfrac\{1\}\{2\}\\xi^\{T\}H\\xi\+\\mathcal\{O\}\(\\\|\\xi\\\|^\{3\}\),\(12\.42\)=ℒ\(θ0\)\+12\(ξ~ξ~¯\)†\(∂2ℒ∂θ¯∂θ∂2ℒ∂θ¯2∂2ℒ∂θ2∂2ℒ∂θ∂θ¯\)\(ξ~ξ~¯\)\+𝒪\(‖ξ‖3\)\\displaystyle=\\mathcal\{L\}\(\\theta\_\{0\}\)\+\\frac\{1\}\{2\}\\begin\{pmatrix\}\\widetilde\{\\xi\}\\\\ \\overline\{\\widetilde\{\\xi\}\}\\end\{pmatrix\}^\{\\dagger\}\\begin\{pmatrix\}\\frac\{\\partial^\{2\}\\mathcal\{L\}\}\{\\partial\\overline\{\\theta\}\\partial\\theta\}&\\frac\{\\partial^\{2\}\\mathcal\{L\}\}\{\\partial\\overline\{\\theta\}^\{2\}\}\\\\ \\frac\{\\partial^\{2\}\\mathcal\{L\}\}\{\\partial\\theta^\{2\}\}&\\frac\{\\partial^\{2\}\\mathcal\{L\}\}\{\\partial\\theta\\partial\\overline\{\\theta\}\}\\end\{pmatrix\}\\begin\{pmatrix\}\\widetilde\{\\xi\}\\\\ \\overline\{\\widetilde\{\\xi\}\}\\end\{pmatrix\}\+\\mathcal\{O\}\(\\\|\\xi\\\|^\{3\}\)\(12\.43\)e−2ℒ\(θ0\+ξ\)/η\\displaystyle e^\{\-2\\mathcal\{L\}\(\\theta\_\{0\}\+\\xi\)/\\eta\}=e−2ℒ\(θ0\)/ηe−ξTHξ/η\(1\+𝒪\(∥ξ∥\)\)\.\\displaystyle=e^\{\-2\\mathcal\{L\}\(\\theta\_\{0\}\)/\\eta\}e^\{\-\\xi^\{T\}H\\xi/\\eta\}\\big\(1\+\\mathcal\{O\}\(\\\|\\xi\\\|\)\\big\)\.\(12\.44\)
Let us examine the denominator of[12\.39](https://arxiv.org/html/2608.19584#S12.E39)\. Since∇ℒ\(θ0\+ξ\)=Hξ\+𝒪\(‖ξ‖2\)\\nabla\\mathcal\{L\}\(\\theta\_\{0\}\+\\xi\)=H\\xi\+\\mathcal\{O\}\(\\\|\\xi\\\|^\{2\}\),
‖∇ℒ\(θ0\+ξ\)‖h2\\displaystyle\\\|\\nabla\\mathcal\{L\}\(\\theta\_\{0\}\+\\xi\)\\\|\_\{h\}^\{2\}=ξTQξ\+𝒪\(‖ξ‖3\),Q:=H†g−1H\.\\displaystyle=\\xi^\{T\}Q\\xi\+\\mathcal\{O\}\(\\\|\\xi\\\|^\{3\}\),\\quad Q:=H^\{\\dagger\}g^\{\-1\}H\.\(12\.45\)We can use[12\.45](https://arxiv.org/html/2608.19584#S12.E45)and via Laplace’s method we can translate the manifold\-valued integral to something in terms of Lebesgue measure∫U∥∇ℒ∥h2e−2ℒ/ηωKK\!≈e−2ℒ\(θ0\)/η∫ℝ2K\(ξTQξ\)e−ξT\(H/η\)ξdξ\\int\_\{U\}\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{2\}e^\{\-2\\mathcal\{L\}/\\eta\}\\frac\{\\omega^\{K\}\}\{K\!\}\\approx e^\{\-2\\mathcal\{L\}\(\\theta\_\{0\}\)/\\eta\}\\int\_\{\\mathbb\{R\}^\{2K\}\}\(\\xi^\{T\}Q\\xi\)e^\{\-\\xi^\{T\}\(H/\\eta\)\\xi\}d\\xi\. Applying the Gaussian second\-moment identity∫ξTQξe−ξTAξ𝑑ξ=πK2detATr\(A−1Q\)\\int\\xi^\{T\}Q\\xi e^\{\-\\xi^\{T\}A\\xi\}d\\xi=\\frac\{\\pi^\{K\}\}\{2\\sqrt\{\\det A\}\}\\text\{Tr\}\(A^\{\-1\}Q\)withA=H/ηA=H/\\eta, and usingH−1Q=g−1HH^\{\-1\}Q=g^\{\-1\}H,
∫Uχ2e−2ℒ/η∥∇ℒ∥h2ωKK\!\\displaystyle\\int\_\{U\}\\chi^\{2\}e^\{\-2\\mathcal\{L\}/\\eta\}\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{2\}\\frac\{\\omega^\{K\}\}\{K\!\}=e−2ℒ\(θ0\)/ηηK\+1πK2detHTr\(g−1H\)\(1\+𝒪\(η\)\)\.\\displaystyle=e^\{\-2\\mathcal\{L\}\(\\theta\_\{0\}\)/\\eta\}\\eta^\{K\+1\}\\frac\{\\pi^\{K\}\}\{2\\sqrt\{\\det H\}\}\\text\{Tr\}\(g^\{\-1\}H\)\\big\(1\+\\mathcal\{O\}\(\\sqrt\{\\eta\}\)\\big\)\.\(12\.46\)
We are using a real\-Gaussian second moment on complexified variables treated as independent\. We can note integrating againstωKK\!\\frac\{\\omega^\{K\}\}\{K\!\}is the equivalent of Lebesgue measure along the complex manifold weighted by the metric, so the identity holds fiberwise\. Instead, Laplace’s method allows us to turn the manifold\-valued quantity to one on a flat space, then we use the identity\. Now we examine the numerator of[12\.39](https://arxiv.org/html/2608.19584#S12.E39)\. WriteH0:=∇1,0\(∂¯ℒ\)\(θ0\)H\_\{0\}:=\\nabla^\{1,0\}\(\\overline\{\\partial\}\\mathcal\{L\}\)\(\\theta\_\{0\}\)\. Since∂ℒ⊗∂¯ℒ=𝒪\(‖ξ‖2\)\\partial\\mathcal\{L\}\\otimes\\overline\{\\partial\}\\mathcal\{L\}=\\mathcal\{O\}\(\\\|\\xi\\\|^\{2\}\)and‖∇ℒ‖h4=𝒪\(‖ξ‖4\)\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{4\}=\\mathcal\{O\}\(\\\|\\xi\\\|^\{4\}\),
‖∇1,0\(∂¯ℒ\)−1η∂ℒ⊗∂¯ℒ‖h2\+1η2‖∇ℒ‖h4\\displaystyle\\left\\\|\\nabla^\{1,0\}\(\\overline\{\\partial\}\\mathcal\{L\}\)\-\\frac\{1\}\{\\eta\}\\partial\\mathcal\{L\}\\otimes\\overline\{\\partial\}\\mathcal\{L\}\\right\\\|\_\{h\}^\{2\}\+\\frac\{1\}\{\\eta^\{2\}\}\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{4\}\(12\.47\)=‖H0‖h2−2ηRe⟨H0,∂ℒ⊗∂¯ℒ⟩h\+1η2‖∂ℒ⊗∂¯ℒ‖h2\+1η2‖∇ℒ‖h4\+𝒪\(‖ξ‖\)\.\\displaystyle=\\\|H\_\{0\}\\\|\_\{h\}^\{2\}\-\\frac\{2\}\{\\eta\}\\text\{Re\}\\langle H\_\{0\},\\partial\\mathcal\{L\}\\otimes\\overline\{\\partial\}\\mathcal\{L\}\\rangle\_\{h\}\+\\frac\{1\}\{\\eta^\{2\}\}\\\|\\partial\\mathcal\{L\}\\otimes\\overline\{\\partial\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}\+\\frac\{1\}\{\\eta^\{2\}\}\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{4\}\+\\mathcal\{O\}\(\\\|\\xi\\\|\)\.\(12\.48\)Under the Laplace localization scalingξ=ηy\\xi=\\sqrt\{\\eta\}y, we note that∂ℒ⊗∂¯ℒ=𝒪\(η‖y‖2\)\\partial\\mathcal\{L\}\\otimes\\overline\{\\partial\}\\mathcal\{L\}=\\mathcal\{O\}\(\\eta\\\|y\\\|^\{2\}\)and‖∇ℒ‖h4=𝒪\(η2‖y‖4\)\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{4\}=\\mathcal\{O\}\(\\eta^\{2\}\\\|y\\\|^\{4\}\)\. The1/η1/\\etaand1/η21/\\eta^\{2\}coefficients balance this, meaning all terms in the expansion contribute to the leading\-order Gaussian integral\. LetC0\>0C\_\{0\}\>0be the constant denote the sum of these𝒪\(1\)\\mathcal\{O\}\(1\)Gaussian moments
∫Uχ2e−2ℒ/η\(∥∇1,0\(∂¯ℒ\)−1η∂ℒ⊗∂¯ℒ∥h2\+1η2∥∇ℒ∥h4\)ωKK\!=e−2ℒ\(θ0\)/ηηKπKdetHC0\(1\+𝒪\(η\)\)\.\\displaystyle\\int\_\{U\}\\chi^\{2\}e^\{\-2\\mathcal\{L\}/\\eta\}\\left\(\\Big\\\|\\nabla^\{1,0\}\(\\overline\{\\partial\}\\mathcal\{L\}\)\-\\frac\{1\}\{\\eta\}\\partial\\mathcal\{L\}\\otimes\\overline\{\\partial\}\\mathcal\{L\}\\Big\\\|\_\{h\}^\{2\}\+\\frac\{1\}\{\\eta^\{2\}\}\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{4\}\\right\)\\frac\{\\omega^\{K\}\}\{K\!\}=e^\{\-2\\mathcal\{L\}\(\\theta\_\{0\}\)/\\eta\}\\eta^\{K\}\\frac\{\\pi^\{K\}\}\{\\sqrt\{\\det H\}\}C\_\{0\}\\big\(1\+\\mathcal\{O\}\(\\sqrt\{\\eta\}\)\\big\)\.\(12\.49\)
The factorsηK\(detH\)−1/2πKe−2ℒ\(θ0\)/η\\eta^\{K\}\(\\det H\)^\{\-1/2\}\\pi^\{K\}e^\{\-2\\mathcal\{L\}\(\\theta\_\{0\}\)/\\eta\}cancel, leaving
∫Uχ2e−2ℒ/η\(∥∇1,0\(∂¯ℒ\)−1η∂ℒ⊗∂¯ℒ∥h2\+1η2∥∇ℒ∥h4\)ωKK\!∫Uχ2e−2ℒ/η∥∇ℒ∥h2ωKK\!\\displaystyle\\frac\{\\int\_\{U\}\\chi^\{2\}e^\{\-2\\mathcal\{L\}/\\eta\}\\left\(\\\|\\nabla^\{1,0\}\(\\overline\{\\partial\}\\mathcal\{L\}\)\-\\frac\{1\}\{\\eta\}\\partial\\mathcal\{L\}\\otimes\\overline\{\\partial\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}\+\\frac\{1\}\{\\eta^\{2\}\}\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{4\}\\right\)\\frac\{\\omega^\{K\}\}\{K\!\}\}\{\\int\_\{U\}\\chi^\{2\}e^\{\-2\\mathcal\{L\}/\\eta\}\\\|\\nabla\\mathcal\{L\}\\\|\_\{h\}^\{2\}\\frac\{\\omega^\{K\}\}\{K\!\}\}\(12\.50\)=ηK\(detH\)−1/2πKe−2ℒ\(θ0\)/ηηK\(detH\)−1/2πKe−2ℒ\(θ0\)/η⋅C0ηTr\(g−1⋅Hessℒ\(θ0\)\)⏟=:Mη\(1\+𝒪\(η\)\),η→0\+\.\\displaystyle=\\cancel\{\\frac\{\\eta^\{K\}\(\\det H\)^\{\-1/2\}\\pi^\{K\}e^\{\-2\\mathcal\{L\}\(\\theta\_\{0\}\)/\\eta\}\}\{\\eta^\{K\}\(\\det H\)^\{\-1/2\}\\pi^\{K\}e^\{\-2\\mathcal\{L\}\(\\theta\_\{0\}\)/\\eta\}\}\}\\cdot\\underbrace\{\\frac\{C\_\{0\}\}\{\\eta\\text\{Tr\}\\big\(g^\{\-1\}\\cdot\\mathrm\{Hess\}\\mathcal\{L\}\(\\theta\_\{0\}\)\\big\)\}\}\_\{=:M\_\{\\eta\}\}\\big\(1\+\\mathcal\{O\}\(\\sqrt\{\\eta\}\)\\big\),\\eta\\to 0^\{\+\}\.\(12\.51\)We will letC0C\_\{0\}absorb all constants\. SinceHessℒ\(θ0\)≻0\\mathrm\{Hess\}\\mathcal\{L\}\(\\theta\_\{0\}\)\\succ 0andh≻0h\\succ 0, the traceTr\(g−1⋅Hessℒ\(θ0\)\)\\text\{Tr\}\(g^\{\-1\}\\cdot\\mathrm\{Hess\}\\mathcal\{L\}\(\\theta\_\{0\}\)\)is positive, soMηM\_\{\\eta\}is finite forη\\etabounded below positively and well\-defined\. Combining \([12\.50](https://arxiv.org/html/2608.19584#S12.E50)\) with \([12\.36](https://arxiv.org/html/2608.19584#S12.E36)\) and \([12\.38](https://arxiv.org/html/2608.19584#S12.E38)\), the eigenvalue bound becomes
λ1≤\(2\+K\)β1,1η−κ\+Mη\+ℰ\(∇χ\)expansive term\.\\displaystyle\\lambda\_\{1\}\\leq\\frac\{\(2\+K\)\\beta\_\{1,1\}\}\{\\eta\}\-\\kappa\+M\_\{\\eta\}\+\\frac\{\\mathcal\{E\}\(\\nabla\\chi\)\}\{\\text\{expansive term\}\}\.\(12\.52\)Finally, we examine the error termℰ\\mathcal\{E\}\. We get under consideration of the larger space
limR→∞ℰ\(∇χR\)‖αR‖h2=0\.\\displaystyle\\lim\_\{R\\to\\infty\}\\frac\{\\mathcal\{E\}\(\\nabla\\chi\_\{R\}\)\}\{\\\|\\alpha\_\{R\}\\\|\_\{h\}^\{2\}\}=0\.\(12\.53\)We remark the above is slightly non\-rigorous, although sufficient for our purposes\. This completes the proof\.
□\\square
Proof of claim\.We have defined∂¯η=∂¯\+1η∂¯ℒ∧\\overline\{\\partial\}\_\{\\eta\}=\\overline\{\\partial\}\+\\frac\{1\}\{\\eta\}\\overline\{\\partial\}\\mathcal\{L\}\\wedgeand notice this has adjoint
∂¯η†=∂¯†\+1ηi∇¯ℒ\.\\displaystyle\\overline\{\\partial\}\_\{\\eta\}^\{\\dagger\}=\\overline\{\\partial\}^\{\\dagger\}\+\\frac\{1\}\{\\eta\}i\_\{\\overline\{\\nabla\}\\mathcal\{L\}\}\.\(12\.54\)Under an integral inner product⟨α,Δηα⟩h=∫U\(α,Δηα\)hωKK\!\\langle\\alpha,\\Delta\_\{\\eta\}\\alpha\\rangle\_\{h\}=\\int\_\{U\}\(\\alpha,\\Delta\_\{\\eta\}\\alpha\)\_\{h\}\\frac\{\\omega^\{K\}\}\{K\!\}\. Therefore, under integration by parts with decay \(recall we have definedχ\\chiwith decay\)
⟨α,Δηα⟩h=‖∂¯ηα‖h2\+‖∂¯η†α‖h2\.\\displaystyle\\langle\\alpha,\\Delta\_\{\\eta\}\\alpha\\rangle\_\{h\}=\\\|\\overline\{\\partial\}\_\{\\eta\}\\alpha\\\|\_\{h\}^\{2\}\+\\\|\\overline\{\\partial\}\_\{\\eta\}^\{\\dagger\}\\alpha\\\|\_\{h\}^\{2\}\.\(12\.55\)Expanding the inner products and using the definitions,
‖∂¯ηα‖h2=‖∂¯α\+1η∂¯ℒ∧α‖h2=‖∂¯α‖h2\+1η2‖∂¯ℒ∧α‖h2\+2ηRe⟨∂¯α,∂¯ℒ∧α⟩h\\displaystyle\\\|\\overline\{\\partial\}\_\{\\eta\}\\alpha\\\|\_\{h\}^\{2\}=\\\|\\overline\{\\partial\}\\alpha\+\\frac\{1\}\{\\eta\}\\overline\{\\partial\}\\mathcal\{L\}\\wedge\\alpha\\\|\_\{h\}^\{2\}=\\\|\\overline\{\\partial\}\\alpha\\\|\_\{h\}^\{2\}\+\\frac\{1\}\{\\eta^\{2\}\}\\\|\\overline\{\\partial\}\\mathcal\{L\}\\wedge\\alpha\\\|\_\{h\}^\{2\}\+\\frac\{2\}\{\\eta\}\\text\{Re\}\\langle\\overline\{\\partial\}\\alpha,\\overline\{\\partial\}\\mathcal\{L\}\\wedge\\alpha\\rangle\_\{h\}\(12\.56\)‖∂¯η†α‖h2=‖∂¯†α\+1ηi∇¯ℒα‖h2=‖∂¯†α‖h2\+1η2‖i∇¯ℒα‖h2\+2ηRe⟨∂¯†α,i∇¯ℒα⟩h\.\\displaystyle\\\|\\overline\{\\partial\}\_\{\\eta\}^\{\\dagger\}\\alpha\\\|\_\{h\}^\{2\}=\\\|\\overline\{\\partial\}^\{\\dagger\}\\alpha\+\\frac\{1\}\{\\eta\}i\_\{\\overline\{\\nabla\}\\mathcal\{L\}\}\\alpha\\\|\_\{h\}^\{2\}=\\\|\\overline\{\\partial\}^\{\\dagger\}\\alpha\\\|\_\{h\}^\{2\}\+\\frac\{1\}\{\\eta^\{2\}\}\\\|i\_\{\\overline\{\\nabla\}\\mathcal\{L\}\}\\alpha\\\|\_\{h\}^\{2\}\+\\frac\{2\}\{\\eta\}\\text\{Re\}\\langle\\overline\{\\partial\}^\{\\dagger\}\\alpha,i\_\{\\overline\{\\nabla\}\\mathcal\{L\}\}\\alpha\\rangle\_\{h\}\.\(12\.57\)On a Kähler manifold, the Weitzenböck/Bochner\-Kodaira\-Morrey\-Kohn identity[47](https://arxiv.org/html/2608.19584#bib.bib61)is
⟨α,Δ∂¯α⟩h=∫U\(‖∇1,0α‖h2\+Ric\(α,α¯\)\)ωKK\!\.\\displaystyle\\langle\\alpha,\\Delta\_\{\\overline\{\\partial\}\}\\alpha\\rangle\_\{h\}=\\int\_\{U\}\\left\(\\\|\\nabla^\{1,0\}\\alpha\\\|\_\{h\}^\{2\}\+\\text\{Ric\}\(\\alpha,\\overline\{\\alpha\}\)\\right\)\\frac\{\\omega^\{K\}\}\{K\!\}\.\(12\.58\)We can note the generalized relation of the interior product/insertion operator on a product ofppanti\-holomorphic basis generators[8](https://arxiv.org/html/2608.19584#bib.bib71)
i∇ℒ¯k\(dθ¯a1∧⋯∧dθ¯ap\)=k\!∑S⊆\{1,…,p\}\|S\|=k⋀j=1p\{\(∇ℒ¯\)aj,j∈Sdθ¯aj,j∉S\.\\displaystyle i\_\{\\overline\{\\nabla\\mathcal\{L\}\}\}^\{k\}\(d\\overline\{\\theta\}^\{a\_\{1\}\}\\wedge\\cdots\\wedge d\\overline\{\\theta\}^\{a\_\{p\}\}\)=k\!\\sum\_\{\\begin\{subarray\}\{c\}S\\subseteq\\\{1,\\dots,p\\\}\\\\ \|S\|=k\\end\{subarray\}\}\\bigwedge\_\{j=1\}^\{p\}\\begin\{cases\}\(\\overline\{\\nabla\\mathcal\{L\}\}\)^\{a\_\{j\}\},&j\\in S\\\\ d\\overline\{\\theta\}^\{a\_\{j\}\},&j\\notin S\.\\end\{cases\}\(12\.59\)For our\(0,1\)\(0,1\)\-test form, we havep=1p=1, and any iterationk≥2k\\geq 2annihilates the form entirely\. Applyingk=1k=1leads us to the1/η21/\\eta^\{2\}terms
1η2\(‖∂¯ℒ∧α‖h2\+‖i∇¯ℒα‖h2\)=1η2‖∂¯ℒ‖h2‖α‖h2,\\displaystyle\\frac\{1\}\{\\eta^\{2\}\}\\left\(\\\|\\overline\{\\partial\}\\mathcal\{L\}\\wedge\\alpha\\\|\_\{h\}^\{2\}\+\\\|i\_\{\\overline\{\\nabla\}\\mathcal\{L\}\}\\alpha\\\|\_\{h\}^\{2\}\\right\)=\\frac\{1\}\{\\eta^\{2\}\}\\\|\\overline\{\\partial\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}\\\|\\alpha\\\|\_\{h\}^\{2\},\(12\.60\)which is consistent with the fundamental relation for a Clifford algebra, differential forms,β∧iβ♯\+iβ♯β∧=‖β‖ω2\\beta\\wedge i\_\{\\beta^\{\\sharp\}\}\+i\_\{\\beta^\{\\sharp\}\}\\beta\\wedge=\\\|\\beta\\\|\_\{\\omega\}^\{2\}[64](https://arxiv.org/html/2608.19584#bib.bib59)[28](https://arxiv.org/html/2608.19584#bib.bib60)\. Examining the real term,
2ηRe\(⟨∂¯α,∂¯ℒ∧α⟩h\+⟨∂¯†α,i∇¯ℒα⟩h\)\.\\displaystyle\\frac\{2\}\{\\eta\}\\text\{Re\}\\left\(\\langle\\overline\{\\partial\}\\alpha,\\overline\{\\partial\}\\mathcal\{L\}\\wedge\\alpha\\rangle\_\{h\}\+\\langle\\overline\{\\partial\}^\{\\dagger\}\\alpha,i\_\{\\overline\{\\nabla\}\\mathcal\{L\}\}\\alpha\\rangle\_\{h\}\\right\)\.\(12\.61\)Via integration by parts,
=1η⟨α,\(∂¯†\(∂¯ℒ∧⋅\)\+∂¯ℒ∧∂¯†\+∂¯i∇¯ℒ\+i∇¯ℒ∂¯\)α⟩h\.\\displaystyle=\\frac\{1\}\{\\eta\}\\langle\\alpha,\\left\(\\overline\{\\partial\}^\{\\dagger\}\(\\overline\{\\partial\}\\mathcal\{L\}\\wedge\\cdot\)\+\\overline\{\\partial\}\\mathcal\{L\}\\wedge\\overline\{\\partial\}^\{\\dagger\}\+\\overline\{\\partial\}i\_\{\\overline\{\\nabla\}\\mathcal\{L\}\}\+i\_\{\\overline\{\\nabla\}\\mathcal\{L\}\}\\overline\{\\partial\}\\right\)\\alpha\\rangle\_\{h\}\.\(12\.62\)We can note the inside operator is the sum of two anticommutators
T=\{∂¯†,∂¯ℒ∧\}\+\{∂¯,i∇¯ℒ\}\.\\displaystyle T=\\\{\\overline\{\\partial\}^\{\\dagger\},\\overline\{\\partial\}\\mathcal\{L\}\\wedge\\\}\+\\\{\\overline\{\\partial\},i\_\{\\overline\{\\nabla\}\\mathcal\{L\}\}\\\}\.\(12\.63\)In local coordinates, the second term in[12\.63](https://arxiv.org/html/2608.19584#S12.E63)evaluates to
\{∂¯,i∇¯ℒ\}α=\(∇1,1ℒ\)α\+∇∇¯ℒα\.\\displaystyle\\\{\\overline\{\\partial\},i\_\{\\overline\{\\nabla\}\\mathcal\{L\}\}\\\}\\alpha=\(\\nabla^\{1,1\}\\mathcal\{L\}\)\\alpha\+\\nabla\_\{\\overline\{\\nabla\}\\mathcal\{L\}\}\\alpha\.\(12\.64\)The first anticommutator,\{∂¯†,∂¯ℒ∧\}\\\{\\overline\{\\partial\}^\{\\dagger\},\\overline\{\\partial\}\\mathcal\{L\}\\wedge\\\}, is the adjoint of the second, and from[12\.63](https://arxiv.org/html/2608.19584#S12.E63), we get
\{∂¯†,∂¯ℒ∧\}α=\(∇1,1ℒ\)α−\(Δ∂¯ℒ\)α−∇∇ℒα\.\\displaystyle\\\{\\overline\{\\partial\}^\{\\dagger\},\\overline\{\\partial\}\\mathcal\{L\}\\wedge\\\}\\alpha=\(\\nabla^\{1,1\}\\mathcal\{L\}\)\\alpha\-\(\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}\)\\alpha\-\\nabla\_\{\\nabla\\mathcal\{L\}\}\\alpha\.\(12\.65\)On a Kähler manifold acting on a\(0,1\)\(0,1\)\-form, the first\-order covariant differential terms components of these two anticommutators are skew\-adjoint relative to each other\. When summed, these covariant terms annihilate, leaving only
1η⟨α,Tα⟩h=1η⟨α,\(2∇ω1,1ℒ−Δ∂¯ℒ\)α⟩h\.\\displaystyle\\frac\{1\}\{\\eta\}\\langle\\alpha,T\\alpha\\rangle\_\{h\}=\\frac\{1\}\{\\eta\}\\langle\\alpha,\(2\\nabla\_\{\\omega\}^\{1,1\}\\mathcal\{L\}\-\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}\)\\alpha\\rangle\_\{h\}\.\(12\.66\)Gathering all terms, we see
⟨α,Δηα⟩h\\displaystyle\\langle\\alpha,\\Delta\_\{\\eta\}\\alpha\\rangle\_\{h\}=∫U\(‖∇1,0α‖h2\+1η2‖∂¯ℒ‖h2‖α‖h2\+Ric\(α,α¯\)\+1η⟨α,\(2∇ω1,1ℒ−Δ∂¯ℒ\)α⟩h\)ωKK\!\.\\displaystyle=\\int\_\{U\}\\left\(\\\|\\nabla^\{1,0\}\\alpha\\\|\_\{h\}^\{2\}\+\\frac\{1\}\{\\eta^\{2\}\}\\\|\\overline\{\\partial\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}\\\|\\alpha\\\|\_\{h\}^\{2\}\+\\text\{Ric\}\(\\alpha,\\overline\{\\alpha\}\)\+\\frac\{1\}\{\\eta\}\\langle\\alpha,\(2\\nabla\_\{\\omega\}^\{1,1\}\\mathcal\{L\}\-\\Delta\_\{\\overline\{\\partial\}\}\\mathcal\{L\}\)\\alpha\\rangle\_\{h\}\\right\)\\frac\{\\omega^\{K\}\}\{K\!\}\.\(12\.67\)This proves the claim\.
□\\square
### 12\.3Regions of some but not almost everywhere bad curvature and its linear bounds inKK
Proof of Lemma 14\.Let the parameter obey the natural gradient descent vector fieldV=−∇h1,0ℒV=\-\\nabla^\{1,0\}\_\{h\}\\mathcal\{L\}\. Define the\(1,1\)\(1,1\)\-form curvatureΘV=i⟨Θh\(T1,0M\)V,V⟩h\\Theta\_\{V\}=i\\langle\\Theta\_\{h\}\(T^\{1,0\}M\)V,V\\rangle\_\{h\}\. Writing the Riemann curvature asRikj¯l¯R\_\{ik\\overline\{j\}\\overline\{l\}\}, we have in local holomorphic coordinates
ΘV=iRikj¯l¯ViV¯jdθk∧dθ¯l\.\\displaystyle\\Theta\_\{V\}=iR\_\{ik\\overline\{j\}\\overline\{l\}\}V^\{i\}\\overline\{V\}^\{j\}d\\theta^\{k\}\\wedge d\\overline\{\\theta\}^\{l\}\.\(12\.68\)Applying the dual Lefschetz operatorΛω\\Lambda\_\{\\omega\}on this\(1,1\)\(1,1\)\-form yields the Ricci curvature, i\.e\.ΛωΘV=hkl¯Rikj¯l¯ViV¯j=Ric\(V,V¯\)\\Lambda\_\{\\omega\}\\Theta\_\{V\}=h^\{k\\overline\{l\}\}R\_\{ik\\overline\{j\}\\overline\{l\}\}V^\{i\}\\overline\{V\}^\{j\}=\\text\{Ric\}\(V,\\overline\{V\}\)\.
By the complex Lefschetz decomposition theorem on Kähler manifolds, any\(1,1\)\(1,1\)\-form decomposes into a trace component proportional to the Kähler form and a primitive component[71](https://arxiv.org/html/2608.19584#bib.bib66)[30](https://arxiv.org/html/2608.19584#bib.bib67)\. DecomposingΘV\\Theta\_\{V\}, we have
ΘV=1K\(ΛωΘV\)ω\+ΘV,prim=Ric\(V,V¯\)Kω\+ΘV,prim,\\displaystyle\\Theta\_\{V\}=\\frac\{1\}\{K\}\(\\Lambda\_\{\\omega\}\\Theta\_\{V\}\)\\omega\+\\Theta\_\{V,\\text\{prim\}\}=\\frac\{\\text\{Ric\}\(V,\\overline\{V\}\)\}\{K\}\\omega\+\\Theta\_\{V,\\text\{prim\}\},\(12\.69\)whereΘV,prim\\Theta\_\{V,\\text\{prim\}\}is primitive, i\.e\.,ΛωΘV,prim=0\\Lambda\_\{\\omega\}\\Theta\_\{V,\\text\{prim\}\}=0[71](https://arxiv.org/html/2608.19584#bib.bib66)\. Byω−q\\omega\-q\-semi\-positivity hypothesis, which is\{\(iΘh\(T1,0M\)∧ωq−1∧Ω\)V,V\}h≥0\\Bigg\\\{\(i\\Theta\_\{h\}\(T^\{1,0\}M\)\\wedge\\omega^\{q\-1\}\\wedge\\Omega\)V,V\\Bigg\\\}\_\{h\}\\geq 0, we get
\(Ric\(V,V¯\)Kω\+ΘV,prim\)∧ωq−1∧Ω≥0\.\\displaystyle\\left\(\\frac\{\\text\{Ric\}\(V,\\overline\{V\}\)\}\{K\}\\omega\+\\Theta\_\{V,\\text\{prim\}\}\\right\)\\wedge\\omega^\{q\-1\}\\wedge\\Omega\\geq 0\.\(12\.70\)The inner product\{⋅,⋅\}h\\\{\\cdot,\\cdot\\\}\_\{h\}is absorbed into the definition ofΘV\\Theta\_\{V\}since
\{\(iΘh\(T1,0M\)∧ωq−1∧Ω\)V,V\}h=i\{Θh\(T1,0M\)V,V\}h∧ωq−1∧Ω,\\displaystyle\\Bigg\\\{\(i\\Theta\_\{h\}\(T^\{1,0\}M\)\\wedge\\omega^\{q\-1\}\\wedge\\Omega\)V,V\\Bigg\\\}\_\{h\}=i\\Bigg\\\{\\Theta\_\{h\}\(T^\{1,0\}M\)V,V\\Bigg\\\}\_\{h\}\\wedge\\omega^\{q\-1\}\\wedge\\Omega,\(12\.71\)soΘV=i\{Θh\(T1,0M\)V,V\}h\\Theta\_\{V\}=i\\Bigg\\\{\\Theta\_\{h\}\(T^\{1,0\}M\)V,V\\Bigg\\\}\_\{h\}\. Now, recall the Hodge star operator⋆\\starmaps a\(K,K\)\(K,K\)\-form to a\(0,0\)\(0,0\)\-form, a scalar function\. Applying the Hodge star and distributing in[12\.70](https://arxiv.org/html/2608.19584#S12.E70),
Ric\(V,V¯\)K⋆\(ωq∧Ω\)\+⋆\(ΘV,prim∧ωq−1∧Ω\)≥0\.\\displaystyle\\frac\{\\text\{Ric\}\(V,\\overline\{V\}\)\}\{K\}\\star\(\\omega^\{q\}\\wedge\\Omega\)\+\\star\(\\Theta\_\{V,\\text\{prim\}\}\\wedge\\omega^\{q\-1\}\\wedge\\Omega\)\\geq 0\.\(12\.72\)By applying the Hodge star, we have successfully pulled the Ricci curvature out of the wedge product and what remains is a scalar function\. DenoteMq,Ω=⋆\(ωq∧Ω\)M\_\{q,\\Omega\}=\\star\(\\omega^\{q\}\\wedge\\Omega\), which is positive since bothω\\omegaandΩ\\Omegaare positive forms, and let𝒫q,Ω\(V\)=⋆\(ΘV,prim∧ωq−1∧Ω\)\\mathcal\{P\}\_\{q,\\Omega\}\(V\)=\\star\(\\Theta\_\{V,\\text\{prim\}\}\\wedge\\omega^\{q\-1\}\\wedge\\Omega\)denote the primitive curvature term for short\. The Ricci curvature term follows by rearranging[12\.72](https://arxiv.org/html/2608.19584#S12.E72)
Ric\(V,V¯\)≥−KMq,Ω𝒫q,Ω\(V\)\.\\displaystyle\\text\{Ric\}\(V,\\overline\{V\}\)\\geq\-\\frac\{K\}\{M\_\{q,\\Omega\}\}\\mathcal\{P\}\_\{q,\\Omega\}\(V\)\.\(12\.73\)We have kept track of signs to make sure the inequality is valid\. Let us turn to what we defined as the divergence in the lemma statement\. We setΘ=divh\(V\)\\Theta=\\text\{div\}\_\{h\}\(V\)\. From the complex Bochner\-Weitzenböck identity \(this identity is reminiscent of the one we saw in[12\.2](https://arxiv.org/html/2608.19584#S12.SS2), although taking a different form\),
Θ˙=−12Δh𝒱˙−‖∇ω1,1ℒ‖h2−‖∇ω2,0ℒ‖h2−Ric\(V,V¯\)\.\\displaystyle\\dot\{\\Theta\}=\-\\frac\{1\}\{2\}\\Delta\_\{h\}\\dot\{\\mathcal\{V\}\}\-\\\|\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}\-\\\|\\nabla^\{2,0\}\_\{\\omega\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}\-\\text\{Ric\}\(V,\\overline\{V\}\)\.\(12\.74\)Recall𝒱˙=−‖V‖h2\\dot\{\\mathcal\{V\}\}=\-\\\|V\\\|\_\{h\}^\{2\}\. Substituting in[12\.73](https://arxiv.org/html/2608.19584#S12.E73),
Θ˙\+12Δh𝒱˙≤−‖∇ω1,1ℒ‖h2−‖∇ω2,0ℒ‖h2\+KMq,Ω𝒫q,Ω\(V\)\.\\displaystyle\\dot\{\\Theta\}\+\\frac\{1\}\{2\}\\Delta\_\{h\}\\dot\{\\mathcal\{V\}\}\\leq\-\\\|\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}\-\\\|\\nabla^\{2,0\}\_\{\\omega\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}\+\\frac\{K\}\{M\_\{q,\\Omega\}\}\\mathcal\{P\}\_\{q,\\Omega\}\(V\)\.\(12\.75\)In particular, the left\-hand side is bounded by two subtracted squared norms and a primitive curvature term that scales linearly inKK\. The left\-hand side represents an expansion term, which is the divergence, with diffusion, which is the Laplacian\. SinceVVis a velocity, the material derivativeΘ˙\\dot\{\\Theta\}is an acceleration\. Hence, we lose a guarantee of convergence\. We can note the upper bound diverges−‖∇ω1,1ℒ‖h2−‖∇ω2,0ℒ‖h2=−𝒪\(1m\)\-\\\|\\nabla^\{1,1\}\_\{\\omega\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}\-\\\|\\nabla^\{2,0\}\_\{\\omega\}\\mathcal\{L\}\\\|\_\{h\}^\{2\}=\-\\mathcal\{O\}\(\\frac\{1\}\{m\}\)and
lim sup∑kmkmk\+1\+mL=K→∞\[−𝒪\(1m\)\+K⋅Ω\(1\)\]=\+∞\.\\displaystyle\\limsup\_\{\\sum\_\{k\}m\_\{k\}m\_\{k\+1\}\+m\_\{L\}=K\\to\\infty\}\\left\[\-\\mathcal\{O\}\(\\frac\{1\}\{m\}\)\+K\\cdot\\Omega\(1\)\\right\]=\+\\infty\.\(12\.76\)The1m\\frac\{1\}\{m\}follows from the results from[8\.2](https://arxiv.org/html/2608.19584#S8.SS2)and[8\.5](https://arxiv.org/html/2608.19584#S8.SS5)and since the norms are squared\. We have noted the primitive curvature term scales inΩ\(1\)\\Omega\(1\)and not𝒪\(1\)\\mathcal\{O\}\(1\), so it is at minimum a constant\. If it were𝒪\(1\)\\mathcal\{O\}\(1\), the bound may not hold, since it could be ideally scale more slowly, for example since1K≤𝒪\(1\)\\frac\{1\}\{K\}\\leq\\mathcal\{O\}\(1\)\. The lower bound follows almost immediately by noting−Ric\(V,V¯\)≥−κmax‖V‖h2\-\\text\{Ric\}\(V,\\overline\{V\}\)\\geq\-\\kappa\_\{\\max\}\\\|V\\\|\_\{h\}^\{2\}, and since the two Hessian terms in[12\.74](https://arxiv.org/html/2608.19584#S12.E74)are𝒪\(1m\)\\mathcal\{O\}\(\\frac\{1\}\{m\}\), we get
Θ˙\+12Δh𝒱˙≥−𝒪\(1m\)\+κmax𝒱˙,\\displaystyle\\dot\{\\Theta\}\+\\frac\{1\}\{2\}\\Delta\_\{h\}\\dot\{\\mathcal\{V\}\}\\geq\-\\mathcal\{O\}\\left\(\\frac\{1\}\{m\}\\right\)\+\\kappa\_\{\\max\}\\dot\{\\mathcal\{V\}\},\(12\.77\)as before noting1m\\frac\{1\}\{m\}and not1m\\frac\{1\}\{\\sqrt\{m\}\}due to the square\.
□\\squareSimilar Articles
State-Space NTK Collapse Near Bifurcations
This paper develops a local theory of gradient descent near bifurcations in dynamical models, showing that the state-space neural tangent kernel collapses to a rank-one operator that dominates learning dynamics, making optimization effectively low-dimensional and predictable from normal forms.
Spectral Asymptotics of Neural Network Loss Landscapes: An Exact Decomposition of the Curvature Exponent
This paper presents an exact decomposition of the curvature exponent α in neural network loss landscapes, explaining why it varies across layer types. It introduces the spectral alignment decomposition and derives a spectral transfer identity linking curvature, gradient rank decay, and Hessian exponents, validated across architectures and datasets.
Convergence Guarantees of Gradient Descent for Neural Networks via Generalized Lipschitz Smoothness
This paper establishes convergence guarantees for gradient descent on general feedforward neural networks of arbitrary width/depth, using a novel generalized Lipschitz smoothness condition that holds for common activations and mean-squared error, without special initialization or dataset requirements.
Energy Manifold Natural Gradient Descent: Riemannian Optimization for Neural PDE Solvers
Introduces Energy Manifold Natural Gradient Descent (EMNGD), a manifold optimization framework for neural PDE solvers that aligns parameter updates with function-space energy curvature while respecting parameter constraints. Theoretical guarantees and empirical results show improved accuracy and convergence.
The Field Knows: Cross-Dimensional Geometry from Navigation to Black Holes
This paper introduces a continuous metric field framework trained by a single causal contrastive loss that unifies geometric structure discovery from robot navigation to black hole emergence, demonstrating zero-shot generalization across dimensions.