The Loss Floor of Denoising Score Matching: Fisher Geometry from Schr\"odinger Bridges

arXiv cs.LG Papers

Summary

This paper isolates the irreducible excess in denoising score matching training loss, showing it equals the trace of the Fisher-Rao metric along diffusion trajectories, connecting variational principles, latent space geometry, and information theory.

arXiv:2608.23916v1 Announce Type: new Abstract: Denoising score matching trains diffusion models by regressing onto a conditional score, although generation ultimately requires the marginal score. The two objectives share the same population minimizer, but the conditional target remains random at fixed noisy state and introduces an irreducible excess in the training loss. We isolate this excess and show that, for a general corruption kernel under mild regularity assumptions, it is exactly the trace of the Fisher--Rao metric of the conditional endpoint family, integrated along the diffusion trajectory. This gives an exact conditional-variance decomposition of the denoising objective and identifies the information geometry observed in diffusion latent spaces as an intrinsic component of the training loss. We derive the result from a Schr"odinger bridge variational principle, in which the ideal objective arises as excess path-space relative entropy. For corruption diffusions, the Fisher term is proportional to the rate at which the noisy state loses mutual information about the clean data, separating the loss floor into an information flow determined by the data and a weight determined by the corruption schedule and objective. In the Gaussian case, this yields a closed form for the floor, recovers reparametrization invariance of the continuous-time objective, and relates its high-SNR divergence to the information dimension of the data. Finally, we show that raw losses obtained with different noise ranges or weightings need not rank models consistently because they contain different additive floors, and contrast the second-order geometry seen by training with the third-order conditional statistics entering numerical sampling error.
Original Article
View Cached Full Text

Cached at: 08/26/26, 09:29 AM

# The Loss Floor of Denoising Score Matching:Fisher Geometry from Schrödinger Bridges
Source: [https://arxiv.org/html/2608.23916](https://arxiv.org/html/2608.23916)
Avinash RajuAffiliation:Great Wall Motors Co\. Ltd\.Affiliation:Beijing, ChinaEmail:[avinashraju777@gmail\.com](mailto:)Kai ZhangAffiliation:China Patent Information CenterAffiliation:Beijing, ChinaEmail:[rachelzhang261@gmail\.com](mailto:)

###### Abstract

Denoising score matching trains diffusion models by regressing onto a conditional score, although the generative dynamics ultimately require the marginal score\. The two objectives share the same population minimizer, but the conditional target remains random at fixed noisy state and therefore introduces an irreducible excess in the training loss\. We isolate this excess and show that, for a general corruption kernel under mild regularity assumptions, it is exactly the trace of the Fisher–Rao metric of the conditional endpoint family, integrated along the diffusion trajectory\. The result provides an exact conditional\-variance decomposition of the denoising objective and identifies the information geometry recently observed in diffusion latent spaces as an intrinsic component of the training loss rather than an additional structure imposed on the model\. We derive the result from the Schrödinger bridge variational principle, within which the ideal objective arises as excess path\-space relative entropy\. For corruption diffusions, the Fisher term is proportional to the rate at which the noisy state loses mutual information about the clean data, separating the loss floor into an information flow determined by the data and a weight determined by the corruption schedule and training objective\. In the Gaussian case this yields a closed form for the floor, recovers the reparametrization invariance of the continuous\-time objective, and relates its high\-SNR divergence to the information dimension of the data\. Finally, we examine consequences for practice\. Raw training losses obtained with different noise ranges or weightings contain different additive floors and need not rank models consistently; subtracting the floor restores the correct ordering in our example\. We also contrast the second\-order geometry seen by the training objective with the third\-order conditional statistics that enter numerical sampling error\. Together, these results connect the variational origin of diffusion dynamics, the geometry of latent space, and the information\-theoretic structure of denoising score matching\.

## 1Introduction

Denoising diffusion models\[[1](https://arxiv.org/html/2608.23916#bib.bib1),[2](https://arxiv.org/html/2608.23916#bib.bib2),[3](https://arxiv.org/html/2608.23916#bib.bib3),[4](https://arxiv.org/html/2608.23916#bib.bib4),[5](https://arxiv.org/html/2608.23916#bib.bib5)\]are now a standard framework for generative modeling, with applications spanning image synthesis\[[6](https://arxiv.org/html/2608.23916#bib.bib6),[7](https://arxiv.org/html/2608.23916#bib.bib7),[8](https://arxiv.org/html/2608.23916#bib.bib8)\], video and audio generation\[[9](https://arxiv.org/html/2608.23916#bib.bib9),[10](https://arxiv.org/html/2608.23916#bib.bib10),[11](https://arxiv.org/html/2608.23916#bib.bib11),[12](https://arxiv.org/html/2608.23916#bib.bib12)\], molecular design\[[13](https://arxiv.org/html/2608.23916#bib.bib13),[14](https://arxiv.org/html/2608.23916#bib.bib14)\], and discrete domains such as language\[[15](https://arxiv.org/html/2608.23916#bib.bib15),[16](https://arxiv.org/html/2608.23916#bib.bib16),[17](https://arxiv.org/html/2608.23916#bib.bib17)\]\. Their formulations are by now well understood from several equivalent viewpoints: time\-reversed stochastic differential equations\[[18](https://arxiv.org/html/2608.23916#bib.bib18),[19](https://arxiv.org/html/2608.23916#bib.bib19),[4](https://arxiv.org/html/2608.23916#bib.bib4)\], variational objectives\[[20](https://arxiv.org/html/2608.23916#bib.bib20),[21](https://arxiv.org/html/2608.23916#bib.bib21),[22](https://arxiv.org/html/2608.23916#bib.bib22)\], flow matching and stochastic interpolants\[[23](https://arxiv.org/html/2608.23916#bib.bib23),[24](https://arxiv.org/html/2608.23916#bib.bib24),[25](https://arxiv.org/html/2608.23916#bib.bib25)\], and posterior denoising through Tweedie\-type identities\[[26](https://arxiv.org/html/2608.23916#bib.bib26),[27](https://arxiv.org/html/2608.23916#bib.bib27),[28](https://arxiv.org/html/2608.23916#bib.bib28)\]\. Common to all of these viewpoints is that training proceeds by denoising score matching \(DSM\)\[[29](https://arxiv.org/html/2608.23916#bib.bib29),[30](https://arxiv.org/html/2608.23916#bib.bib30),[31](https://arxiv.org/html/2608.23916#bib.bib31)\]\.

The generative dynamics, however, require the*marginal*score𝐬⁡\(𝐱,t\)=∇𝐱​log​Pt​\(𝐱\)\\mathbf\{s\}\(\\mathbf\{x\},t\)=\\nabla\_\{\\mathbf\{x\}\}\\log P\_\{t\}\(\\mathbf\{x\}\)of the noisy data distribution, which is intractable\. DSM circumvents this by regressing onto the*conditional*score𝐬cond​\(𝐱,t,𝐲\)=∇𝐱​log​q​\(𝐱,t∣𝐲\)\\mathbf\{s\}\_\{\\rm cond\}\(\\mathbf\{x\},t;\\mathbf\{y\}\)=\\nabla\_\{\\mathbf\{x\}\}\\log q\(\\mathbf\{x\},t\\mid\\mathbf\{y\}\)of the corruption kernel for a clean sample𝐲\\mathbf\{y\}\[[2](https://arxiv.org/html/2608.23916#bib.bib2),[30](https://arxiv.org/html/2608.23916#bib.bib30),[31](https://arxiv.org/html/2608.23916#bib.bib31)\]\. The conditional score is an unbiased estimator of the marginal score, so the two objectives share the same population minimizer; but at fixed𝐱\\mathbf\{x\}the conditional target remains a random variable, and this sample\-level stochasticity has a cost that the standard derivation leaves unquantified\.

That cost is the subject of this work\. Because sharing a minimizer does not imply sharing a loss value, replacing the marginal target by a random conditional target inflates the objective by a term that no model can reduce, and the question we answer is what this irreducible excess is and what it depends on\.

Our main result is that the excess admits an exact geometric characterization\. The fluctuation of the conditional score about its posterior mean is the score of the conditional endpoint distribution

p⁡\(𝐲∣𝐱,t\)=Pdata​\(𝐲\)​q​\(𝐱,t∣𝐲\)Pt​\(𝐱\)p\(\\mathbf\{y\}\\mid\\mathbf\{x\},t\)=\\frac\{P\_\{\\rm data\}\(\\mathbf\{y\}\)\\,q\(\\mathbf\{x\},t\\mid\\mathbf\{y\}\)\}\{P\_\{t\}\(\\mathbf\{x\}\)\}with respect to the latent coordinate𝐱\\mathbf\{x\}, so that its second moment is by definition the Fisher information of this family, the central object of information geometry\[[32](https://arxiv.org/html/2608.23916#bib.bib32),[33](https://arxiv.org/html/2608.23916#bib.bib33),[34](https://arxiv.org/html/2608.23916#bib.bib34)\]\. This yields an exact orthogonal decomposition of the denoising objective,

Δ​𝒮DSM=Δ​𝒮ideal\+12​∫0Tγ⁡\(t\)​𝔼𝐱∼Pt​\[tr⁡g⁡\(𝐱,t\)\]​𝑑t,\\Delta\\mathcal\{S\}\_\{\\rm DSM\}=\\Delta\\mathcal\{S\}\_\{\\rm ideal\}\+\\frac\{1\}\{2\}\\int\_\{0\}^\{T\}\\gamma\(t\)\\,\\mathbb\{E\}\_\{\\mathbf\{x\}\\sim P\_\{t\}\}\\\!\\left\[\\operatorname\{tr\}g\(\\mathbf\{x\},t\)\\right\]dt,\(1\.1\)whereΔ​𝒮ideal\\Delta\\mathcal\{S\}\_\{\\rm ideal\}is the ideal marginal\-score objective andg⁡\(𝐱,t\)g\(\\mathbf\{x\},t\)is the Fisher–Rao metric tensor of the conditional endpoint family\. The second term depends only on the corruption and the data, never on the model: it is an irreducible loss floor intrinsic to conditional denoising score matching\.

The Schrödinger bridge formulation\[[35](https://arxiv.org/html/2608.23916#bib.bib35),[36](https://arxiv.org/html/2608.23916#bib.bib36),[37](https://arxiv.org/html/2608.23916#bib.bib37),[38](https://arxiv.org/html/2608.23916#bib.bib38),[39](https://arxiv.org/html/2608.23916#bib.bib39)\]provides a variational foundation for the reverse diffusion dynamics and supplies the common origin of both terms\. A Schrödinger bridge is the relative\-entropy projection of a reference path measure onto prescribed endpoint constraints, and its optimal dynamics are given by a Doobhh\-transform\[[40](https://arxiv.org/html/2608.23916#bib.bib40)\], in which the reference kernel is tilted by a positive harmonic function that enforces the terminal marginal; bridges and their stochastic\-control interpretation have by now been used extensively in generative modeling\[[41](https://arxiv.org/html/2608.23916#bib.bib41),[42](https://arxiv.org/html/2608.23916#bib.bib42),[43](https://arxiv.org/html/2608.23916#bib.bib43),[44](https://arxiv.org/html/2608.23916#bib.bib44),[45](https://arxiv.org/html/2608.23916#bib.bib45),[46](https://arxiv.org/html/2608.23916#bib.bib46)\]\. Here the same variational structure plays a different role: the*first*variation of the bridge functional produces the score\-based generative drift and hence the ideal objectiveΔ​𝒮ideal\\Delta\\mathcal\{S\}\_\{\\rm ideal\}, while the*second*variation produces the quadratic form whose pullback to latent space is the Fisher metric appearing in the floor\. For affine Gaussian corruption,q⁡\(𝐱,t∣𝐲\)=𝒩⁡\(𝐱,αt​𝐲,σt2​I\)q\(\\mathbf\{x\},t\\mid\\mathbf\{y\}\)=\\mathcal\{N\}\(\\mathbf\{x\};\\alpha\_\{t\}\\mathbf\{y\},\\sigma\_\{t\}^\{2\}I\), this metric reduces to

gi​j​\(𝐱,t\)=1σt2​δi​j\+∂i∂jlog⁡Pt​\(𝐱\),g\_\{ij\}\(\\mathbf\{x\},t\)=\\frac\{1\}\{\\sigma\_\{t\}^\{2\}\}\\delta\_\{ij\}\+\\partial\_\{i\}\\partial\_\{j\}\\log P\_\{t\}\(\\mathbf\{x\}\),\(1\.2\)which is precisely the Hessian/Fisher geometry recently identified in diffusion latent spaces\[[47](https://arxiv.org/html/2608.23916#bib.bib47),[48](https://arxiv.org/html/2608.23916#bib.bib48)\]\. Equation \([1\.1](https://arxiv.org/html/2608.23916#S1.E1)\) thus provides a variational origin for that geometry: it is not an additional structure imposed on the model, but the variance of the conditional regression target already present in the training loss\.

The floor also admits an information\-theoretic evaluation\. For a general corruption diffusiond​𝐱t=f⁡\(𝐱t,t\)​d​t\+D⁡\(t\)​d​𝐖td\\mathbf\{x\}\_\{t\}=f\(\\mathbf\{x\}\_\{t\},t\)\\,dt\+\\sqrt\{D\(t\)\}\\,d\\mathbf\{W\}\_\{t\},

𝔼Pt​\[tr⁡g⁡\(𝐱,t\)\]=−2D⁡\(t\)​dd​t​I​\(𝐲,𝐱t\),\\mathbb\{E\}\_\{P\_\{t\}\}\\\!\\left\[\\operatorname\{tr\}g\(\\mathbf\{x\},t\)\\right\]=\-\\frac\{2\}\{D\(t\)\}\\frac\{d\}\{dt\}I\(\\mathbf\{y\};\\mathbf\{x\}\_\{t\}\),\(1\.3\)so the floor separates into a data\-dependent information flow and a schedule\-dependent weight; when the loss weighting coincides with the diffusion coefficient, the integral telescopes to a difference of mutual informations, and in the Gaussian case to a difference of differential entropies through the I–MMSE relation\[[49](https://arxiv.org/html/2608.23916#bib.bib49)\]\.

These identities have direct practical consequences\. Because the additive floor varies with the SNR range, schedule, and loss weighting, raw training losses are not comparable across training configurations; we exhibit a ranking inversion that floor subtraction repairs\. The factorization further separates what schedule design can and cannot affect, complementing recent work on information\-based scheduling and noise allocation\[[20](https://arxiv.org/html/2608.23916#bib.bib20),[50](https://arxiv.org/html/2608.23916#bib.bib50),[51](https://arxiv.org/html/2608.23916#bib.bib51),[52](https://arxiv.org/html/2608.23916#bib.bib52),[53](https://arxiv.org/html/2608.23916#bib.bib53),[54](https://arxiv.org/html/2608.23916#bib.bib54),[55](https://arxiv.org/html/2608.23916#bib.bib55)\]\. Finally, the training objective probes only the second cumulant of the conditional endpoint law, whereas numerical integration of the generative dynamics also probes the third; we use this contrast as a diagnostic for where third\-order geometric effects become significant\.

The paper is organized around the three identities displayed above: the loss floor is the integrated trace of a Fisher–Rao metric \(Eq\. \([1\.1](https://arxiv.org/html/2608.23916#S1.E1)\)\); that trace is the rate of mutual\-information loss along the corruption \(Eq\. \([1\.3](https://arxiv.org/html/2608.23916#S1.E3)\)\); and for Gaussian corruption the metric and the floor admit the closed forms that underlie the geometry of Eq\. \([1\.2](https://arxiv.org/html/2608.23916#S1.E2)\)\. Section[2](https://arxiv.org/html/2608.23916#S2)derives the ideal objective from the Schrödinger bridge variational principle and reduces it to DSM\. Section[3](https://arxiv.org/html/2608.23916#S3)constructs the Fisher geometry of the bridge and proves the decomposition theorem\. Section[4](https://arxiv.org/html/2608.23916#S4)evaluates the floor in information\-theoretic terms\. Section[5](https://arxiv.org/html/2608.23916#S5)develops the consequences for training and sampling, and Section[6](https://arxiv.org/html/2608.23916#S6)discusses scope, extensions, and limitations\. Rather than introducing a new diffusion objective or scheduling algorithm, this work exposes an exact decomposition already implicit in the standard objective and develops the geometric and information\-theoretic structure it entails\.

## 2From Path\-Space Entropy to Denoising Score Matching

This section derives the training objective of diffusion models from the Schrödinger bridge variational principle\. Our presentation proceeds in four steps:\(i\)the bridge problem selects an optimal path measure by relative\-entropy minimization and its optimality conditions take the form of a Doobhh\-transform\(ii\)the excess path entropy relative to the optimum reduces to a quadratic score objective,\(iii\)and a conditional decomposition of the score field reduces that objective to denoising score matching\.The derivation below is included primarily to fix the variational origin of the score\-matching objective\. Its role is modest but useful as it provides a direct route from path\-space entropy to the marginal\-score objective against which denoising score matching will later be compared\. The problem of reverse diffusion was studied from an action principle perspective in\[[56](https://arxiv.org/html/2608.23916#bib.bib56)\]whose notations we borrow\.

### 2\.1The path\-space variational principle

Our goal is to transform an initial distributionP0=PpriorP\_\{0\}=P\_\{\\rm prior\}into a target distributionPT=PdataP\_\{T\}=P\_\{\\rm data\}through a stochastic process\. Let𝒢\\mathcal\{G\}denote a reference path measure, assumed Markovian and governed by killed forward and backward Kolmogorov equations in its final and initial arguments, respectively\. The Schrödinger bridge problem\[[35](https://arxiv.org/html/2608.23916#bib.bib35),[36](https://arxiv.org/html/2608.23916#bib.bib36),[39](https://arxiv.org/html/2608.23916#bib.bib39)\]arises from a large\-deviation question: among the empirical path measures generated by many independent realizations of𝒢\\mathcal\{G\}, which one dominates when one conditions on the endpoint marginals beingP0P\_\{0\}andPTP\_\{T\}?

By Sanov’s theorem, which assigns to the empirical measure of independent samples a large\-deviation rate given by relative entropy\[[57](https://arxiv.org/html/2608.23916#bib.bib57),[58](https://arxiv.org/html/2608.23916#bib.bib58)\], the empirical path measureℋ\\mathcal\{H\}asymptotically carries the large\-deviation weight

ℙ\(ℋn≃ℋ\)≍exp\[−nDKLpath\(ℋ∥𝒢\)\],\\mathbb\{P\}\(\\mathcal\{H\}\_\{n\}\\simeq\\mathcal\{H\}\)\\asymp\\exp\\left\[\-nD\_\{\\rm KL\}^\{\\rm path\}\(\\mathcal\{H\}\\,\\\|\\,\\mathcal\{G\}\)\\right\],so that conditioning on the endpoint constraints restricts the admissible measures to𝒞=\{ℋ:H0=P0,HT=PT\}\\mathcal\{C\}=\\\{\\mathcal\{H\}:H\_\{0\}=P\_\{0\},\\;H\_\{T\}=P\_\{T\}\\\}and the dominant contribution is the constrained minimum of the rate function,

ℋ∗=argminℋ:H0=P0,HT=PTDpathKL\(ℋ∥𝒢\),DpathK​L\(ℋ∥𝒢\)=∫d𝐱0:Nℋ\(𝐱0:N\)lnℋ\(𝐱0:N\)𝒢\(𝐱0:N\)\.\\begin\{split\}\\mathcal\{H\}^\{\*\}&=\\argmin\_\{\\mathcal\{H\}:\\,H\_\{0\}=P\_\{0\},\\;H\_\{T\}=P\_\{T\}\}D^\{\\rm path\}\_\{KL\}\(\\mathcal\{H\}\\\|\\mathcal\{G\}\),\\\\ D^\{\\rm path\}\_\{KL\}\(\\mathcal\{H\}\\\|\\mathcal\{G\}\)&=\\int d\\mathbf\{x\}\_\{0:N\}\\,\\mathcal\{H\}\(\\mathbf\{x\}\_\{0:N\}\)\\ln\\frac\{\\mathcal\{H\}\(\\mathbf\{x\}\_\{0:N\}\)\}\{\\mathcal\{G\}\(\\mathbf\{x\}\_\{0:N\}\)\}\.\\end\{split\}\(2\.1\)Here the interval\[0,T\]\[0,T\]is discretized intoN≫1N\\gg 1segments of lengthΔ​t=T/N\\Delta t=T/N, and the reference measure factorizes as

𝒢\(𝐱0:N\)=P0\(𝐱0\)∏k=0N−1G\(𝐱k\+1∣𝐱k\)\.\\mathcal\{G\}\(\\mathbf\{x\}\_\{0:N\}\)=P\_\{0\}\(\\mathbf\{x\}\_\{0\}\)\\prod\_\{k=0\}^\{N\-1\}G\(\\mathbf\{x\}\_\{k\+1\}\\mid\\mathbf\{x\}\_\{k\}\)\.\(2\.2\)Note that𝒢\\mathcal\{G\}is required neither to be a probability measure nor to mapP0P\_\{0\}toPTP\_\{T\}, since the endpoint constraints are imposed variationally rather than by construction of the reference\.

Solving Eq\. \([2\.1](https://arxiv.org/html/2608.23916#S2.E1)\) by Lagrange multipliers \(Appendix[A](https://arxiv.org/html/2608.23916#A1)\) yields the factorized optimum

ℋ∗\(𝐱0:N\)=f\(𝐱0\)𝒢\(𝐱0:N\)g\(𝐱N\),\\mathcal\{H\}^\{\*\}\(\\mathbf\{x\}\_\{0:N\}\)=f\(\\mathbf\{x\}\_\{0\}\)\\,\\mathcal\{G\}\(\\mathbf\{x\}\_\{0:N\}\)\\,g\(\\mathbf\{x\}\_\{N\}\),\(2\.3\)whereffandggdepend only on the initial and final endpoints\. This is precisely the structure of a Doobhh\-transform\[[40](https://arxiv.org/html/2608.23916#bib.bib40)\], obtained here as a direct consequence of the variational principle: the optimal path measure is the reference diffusion biased at its endpoints by factors that enforce the marginal constraints\. Sinceℋ∗\\mathcal\{H\}^\{\*\}is itself Markovian, so that

ℋ\(𝐱0:N\)=P0\(𝐱0\)∏k=0N−1H\(𝐱k\+1∣𝐱k\),\\mathcal\{H\}\(\\mathbf\{x\}\_\{0:N\}\)=P\_\{0\}\(\\mathbf\{x\}\_\{0\}\)\\prod\_\{k=0\}^\{N\-1\}H\(\\mathbf\{x\}\_\{k\+1\}\\mid\\mathbf\{x\}\_\{k\}\),\(2\.4\)the path\-space divergence in Eq\. \([2\.1](https://arxiv.org/html/2608.23916#S2.E1)\) decomposes into a sum of per\-step divergences between transition kernels, a fact we exploit below\.

### 2\.2Optimal bridge dynamics

The bridge dynamics at intermediate times are most transparently expressed through two auxiliary fields: the backward fieldψ⁡\(𝐱,t\)\\psi\(\\mathbf\{x\},t\)propagates the terminal constraint backward in time, while the forward fieldψ¯​\(𝐱,t\)\\bar\{\\psi\}\(\\mathbf\{x\},t\)propagates the initial condition forward\. A direct consequence of the factorized form Eq\. \([2\.3](https://arxiv.org/html/2608.23916#S2.E3)\) is that the intermediate marginal is their product,

P⁡\(𝐱,t\)=ψ⁡\(𝐱,t\)​ψ¯​\(𝐱,t\),P\(\\mathbf\{x\},t\)=\\psi\(\\mathbf\{x\},t\)\\,\\bar\{\\psi\}\(\\mathbf\{x\},t\),\(2\.5\)so the density at each point of spacetime is the confluence of information flowing from the past and the future\. In the continuum limit the fields obey a dual pair of linear equations \(Appendix[A](https://arxiv.org/html/2608.23916#A1)\), namely the adjoint \(backward\) Kolmogorov equation

∂tψ\+𝐯⋅∇ψ\+γ⁡\(t\)2​∇2ψ−VG​ψ=0,\\partial\_\{t\}\\psi\+\\mathbf\{v\}\\cdot\\nabla\\psi\+\\frac\{\\gamma\(t\)\}\{2\}\\nabla^\{2\}\\psi\-V\_\{G\}\\psi=0,\(2\.6\)and the forward Fokker–Planck equation

∂tψ¯\+∇⋅\(𝐯​ψ¯\)−γ⁡\(t\)2​∇2ψ¯\+VG​ψ¯=0,\\partial\_\{t\}\\bar\{\\psi\}\+\\nabla\\cdot\(\\mathbf\{v\}\\bar\{\\psi\}\)\-\\frac\{\\gamma\(t\)\}\{2\}\\nabla^\{2\}\\bar\{\\psi\}\+V\_\{G\}\\bar\{\\psi\}=0,\(2\.7\)where𝐯\\mathbf\{v\}is the drift of the reference process,γ⁡\(t\)\\gamma\(t\)its diffusion coefficient, andVG​\(𝐱,t\)V\_\{G\}\(\\mathbf\{x\},t\)a possible killing potential\.

The optimal transition kernel is the reference kernel tilted by the ratio of backward fields,

H∗​\(𝐱k\+1∣𝐱k\)=ψk\+1​\(𝐱k\+1\)ψk​\(𝐱k\)​G​\(𝐱k\+1∣𝐱k\),H^\{\*\}\(\\mathbf\{x\}\_\{k\+1\}\\mid\\mathbf\{x\}\_\{k\}\)=\\frac\{\\psi\_\{k\+1\}\(\\mathbf\{x\}\_\{k\+1\}\)\}\{\\psi\_\{k\}\(\\mathbf\{x\}\_\{k\}\)\}\\,G\(\\mathbf\{x\}\_\{k\+1\}\\mid\\mathbf\{x\}\_\{k\}\),\(2\.8\)so that the bridge kernel is a local Doob tilt of the reference kernel\. Its short\-time expansion \(Appendix[A](https://arxiv.org/html/2608.23916#A1)\), which requires retaining second\-order terms inΔ​𝐱\\Delta\\mathbf\{x\}becauseΔ​𝐱∼Δ​t\\Delta\\mathbf\{x\}\\sim\\sqrt\{\\Delta t\}under the reference measure, preserves the diffusion coefficient and shifts the drift:

𝐯∗\(𝐱,t\)=𝐯\(𝐱,t\)\+γ\(t\)∇lnψ\(𝐱,t\)\.\\mathbf\{v\}\_\{\*\}\(\\mathbf\{x\},t\)=\\mathbf\{v\}\(\\mathbf\{x\},t\)\+\\gamma\(t\)\\,\\nabla\\ln\\psi\(\\mathbf\{x\},t\)\.\(2\.9\)The correctionγ∇lnψ\\gamma\\nabla\\ln\\psiis the control field by which the bridge steers the reference dynamics toward the terminal constraint\.

A particularly tractable sector of this construction underlies score\-based generative modeling\. If the reference process satisfiesVG=−∇⋅𝐯V\_\{G\}=\-\\nabla\\cdot\\mathbf\{v\}, the forward equation admits the trivial solutionψ¯=const\.\\bar\{\\psi\}=\{\\rm const\.\}, which we set to unity using the scaling symmetry of Eq\. \([2\.3](https://arxiv.org/html/2608.23916#S2.E3)\); the entire burden of the terminal constraint is then carried by the backward fieldψ\\psi, andPt=ψtP\_\{t\}=\\psi\_\{t\}by Eq\. \([2\.5](https://arxiv.org/html/2608.23916#S2.E5)\), so that the optimal drift Eq\. \([2\.9](https://arxiv.org/html/2608.23916#S2.E9)\) reduces to the score\-matching drift

𝐯∗\(𝐱,t\)=𝐯\(𝐱,t\)\+γ\(t\)∇lnPt\(𝐱\)\.\\mathbf\{v\}\_\{\*\}\(\\mathbf\{x\},t\)=\\mathbf\{v\}\(\\mathbf\{x\},t\)\+\\gamma\(t\)\\,\\nabla\\ln P\_\{t\}\(\\mathbf\{x\}\)\.\(2\.10\)The learned score is therefore not an arbitrary regression target but the optimal control field required to satisfy the terminal data constraint\.111Herettparameterizes the generative direction, so the sign convention of Eq\. \([2\.9](https://arxiv.org/html/2608.23916#S2.E9)\) differs from conventional reverse\-time SDE notation; the two are related by the corresponding change of orientation\.

### 2\.3Excess path entropy and the ideal score objective

We now turn to learning the score𝐬=∇ln⁡Pt\\mathbf\{s\}=\\nabla\\ln P\_\{t\}from a finite dataset𝒟=\{𝐲i\}i=1N∼Pdata\\mathcal\{D\}=\\\{\\mathbf\{y\}\_\{i\}\\\}\_\{i=1\}^\{N\}\\sim P\_\{\\rm data\}\. The optimumℋ∗\\mathcal\{H\}^\{\*\}is theII\-projection of𝒢\\mathcal\{G\}onto the constraint set\[[38](https://arxiv.org/html/2608.23916#bib.bib38),[58](https://arxiv.org/html/2608.23916#bib.bib58)\], for which relative entropy obeys the Pythagorean identity

DK​Lpath\(ℋ∥𝒢\)=DK​Lpath\(ℋ∗∥𝒢\)\+DK​Lpath\(ℋ∥ℋ∗\)D^\{\\text\{path\}\}\_\{KL\}\(\\mathcal\{H\}\\\|\\mathcal\{G\}\)=D^\{\\text\{path\}\}\_\{KL\}\(\\mathcal\{H\}^\{\*\}\\\|\\mathcal\{G\}\)\+D^\{\\text\{path\}\}\_\{KL\}\(\\mathcal\{H\}\\\|\\mathcal\{H\}^\{\*\}\)\(2\.11\)for any feasibleℋ\\mathcal\{H\}sharing the endpoint marginals ofℋ∗\\mathcal\{H\}^\{\*\}\. Since the first term on the right is independent ofℋ\\mathcal\{H\}, learning the bridge is equivalent to minimizing the excess action

Δ𝒮≡DK​Lpath\(ℋ∥𝒢\)−DK​Lpath\(ℋ∗∥𝒢\)=DK​Lpath\(ℋ∥ℋ∗\)≥0\.\\Delta\\mathcal\{S\}\\;\\equiv\\;D\_\{KL\}^\{\\text\{path\}\}\(\\mathcal\{H\}\\\|\\mathcal\{G\}\)\-D\_\{KL\}^\{\\text\{path\}\}\(\\mathcal\{H\}^\{\*\}\\\|\\mathcal\{G\}\)\\;=\\;D\_\{KL\}^\{\\text\{path\}\}\(\\mathcal\{H\}\\\|\\mathcal\{H\}^\{\*\}\)\\;\\geq\\;0\.\(2\.12\)For Markovian path measures with a common initial distribution, this path\-space divergence decomposes into per\-step kernel divergences, each of which is quadratic in the drift difference to leading order inΔ​t\\Delta t\(Appendix[A\.5](https://arxiv.org/html/2608.23916#A1.SS5)\), and summing over time slices gives

Δ​𝒮=12​∫0Td​tγ⁡\(t\)​𝔼𝐱∼PtH​\[‖𝐯H​\(𝐱,t\)−𝐯∗​\(𝐱,t\)‖2\],\\Delta\\mathcal\{S\}=\\frac\{1\}\{2\}\\int\_\{0\}^\{T\}\\frac\{dt\}\{\\gamma\(t\)\}\\,\\mathbb\{E\}\_\{\\mathbf\{x\}\\sim P\_\{t\}^\{H\}\}\\left\[\\\|\\mathbf\{v\}\_\{H\}\(\\mathbf\{x\},t\)\-\\mathbf\{v\}\_\{\*\}\(\\mathbf\{x\},t\)\\\|^\{2\}\\right\],\(2\.13\)where the expectation is taken underPtHP\_\{t\}^\{H\}, the marginal of the*model’s own*dynamics\. This on\-policy weighting is computationally prohibitive, since it requires simulating the generative process during training\. However, at the level of unrestricted drift fields the integrand is point\-wise minimized by𝐯H=𝐯∗\\mathbf\{v\}\_\{H\}=\\mathbf\{v\}\_\{\*\}, so the functional minimizer is independent of the positive weighting measure, and we therefore replacePtHP\_\{t\}^\{H\}by the reference noising marginalPtP\_\{t\}, which is available without simulation, obtaining the simulation\-free objective\[[45](https://arxiv.org/html/2608.23916#bib.bib45)\]

Δ​𝒮SM=12​∫0Td​tγ⁡\(t\)​𝔼𝐱∼Pt​\[‖𝐯H​\(𝐱,t\)−𝐯∗​\(𝐱,t\)‖2\]\.\\Delta\\mathcal\{S\}\_\{\\text\{SM\}\}=\\frac\{1\}\{2\}\\int\_\{0\}^\{T\}\\frac\{dt\}\{\\gamma\(t\)\}\\,\\mathbb\{E\}\_\{\\mathbf\{x\}\\sim P\_\{t\}\}\\left\[\\\|\\mathbf\{v\}\_\{H\}\(\\mathbf\{x\},t\)\-\\mathbf\{v\}\_\{\*\}\(\\mathbf\{x\},t\)\\\|^\{2\}\\right\]\.\(2\.14\)For a restricted parametric model this replacement does affect the projection, a point we return to in Sec\.[6](https://arxiv.org/html/2608.23916#S6)\.

To expose the quantity being learned, we parameterize the candidate drift as a control deformation of the reference,

𝐯H​\(𝐱,t\)=𝐯⁡\(𝐱,t\)\+γ⁡\(t\)​𝐬θ​\(𝐱,t\),\\mathbf\{v\}\_\{H\}\(\\mathbf\{x\},t\)=\\mathbf\{v\}\(\\mathbf\{x\},t\)\+\\gamma\(t\)\\,\\mathbf\{s\}\_\{\\theta\}\(\\mathbf\{x\},t\),\(2\.15\)mirroring the structure of the optimal drift Eq\. \([2\.9](https://arxiv.org/html/2608.23916#S2.E9)\)\. The reference drift then cancels from the difference𝐯H−𝐯∗\\mathbf\{v\}\_\{H\}\-\\mathbf\{v\}\_\{\*\}—the step by which the reference dynamics disappears from the objective—leaving the ideal score\-matching loss

Δ​𝒮ideal=12​∫0Td​t​γ​\(t\)​𝔼𝐱∼Pt​\[‖𝐬θ​\(𝐱,t\)−∇𝐱​log​Pt​\(𝐱\)‖2\],\\Delta\\mathcal\{S\}\_\{\\rm ideal\}\\;=\\;\\frac\{1\}\{2\}\\int\_\{0\}^\{T\}dt\\,\\gamma\(t\)\\,\\mathbb\{E\}\_\{\\mathbf\{x\}\\sim P\_\{t\}\}\\left\[\\left\\\|\\mathbf\{s\}\_\{\\theta\}\(\\mathbf\{x\},t\)\-\\nabla\_\{\\mathbf\{x\}\}\\log P\_\{t\}\(\\mathbf\{x\}\)\\right\\\|^\{2\}\\right\],\(2\.16\)where we have usedPt=ψtP\_\{t\}=\\psi\_\{t\}in the score\-matching sector\. Equation \([2\.16](https://arxiv.org/html/2608.23916#S2.E16)\) is*ideal*in the sense that its regression target, the marginal score, is not available in closed form\.

### 2\.4Denoising score matching as conditional regression

The ideal objective depends on the marginal score, whose evaluation requires the global fieldψ\\psi\. Becauseψ\\psiobeys a*linear*equation, the principle of superposition allows us to define a conditional fieldψ\(𝐲\)​\(𝐱,t\)\\psi^\{\(\\mathbf\{y\}\)\}\(\\mathbf\{x\},t\)corresponding to a bridge pinned to the Dirac terminal constraintg⁡\(𝐱T\)=δd​\(𝐱T−𝐲\)g\(\\mathbf\{x\}\_\{T\}\)=\\delta^\{d\}\(\\mathbf\{x\}\_\{T\}\-\\mathbf\{y\}\)\. The global field is then the continuous superposition of these conditional fields over the data distribution,

ψ⁡\(𝐱,t\)=∫d​𝐲​Pdata​\(𝐲\)​ψ\(𝐲\)​\(𝐱,t\)=𝔼𝐲∼Pdata​\[ψ\(𝐲\)​\(𝐱,t\)\]=P⁡\(𝐱,t\)\.\\psi\(\\mathbf\{x\},t\)=\\int d\\mathbf\{y\}\\,P\_\{\\rm data\}\(\\mathbf\{y\}\)\\,\\psi^\{\(\\mathbf\{y\}\)\}\(\\mathbf\{x\},t\)=\\mathbb\{E\}\_\{\\mathbf\{y\}\\sim P\_\{\\rm data\}\}\\\!\\left\[\\psi^\{\(\\mathbf\{y\}\)\}\(\\mathbf\{x\},t\)\\right\]=P\(\\mathbf\{x\},t\)\.\(2\.17\)Within the score\-matching sector, each conditional field is the corruption kernel itself,

ψ\(𝐲\)​\(𝐱,t\)=q⁡\(𝐱,t∣𝐲\)=ℋ∗​\(𝐱,𝐲\)Pdata​\(𝐲\)\.\\psi^\{\(\\mathbf\{y\}\)\}\(\\mathbf\{x\},t\)=q\(\\mathbf\{x\},t\\mid\\mathbf\{y\}\)=\\frac\{\\mathcal\{H\}^\{\*\}\(\\mathbf\{x\},\\mathbf\{y\}\)\}\{P\_\{\\rm data\}\(\\mathbf\{y\}\)\}\.\(2\.18\)Replacing the data expectation by its Monte Carlo estimate and applying Jensen’s inequality to the convex squared norm then defines the tractable upper bound

Δ​𝒮DSM\\displaystyle\\Delta\\mathcal\{S\}\_\{\\rm DSM\}:=12​∫0Td​t​γ​\(t\)​𝔼𝐲∼Pdata​𝔼𝐱∼q⁡\(𝐱,t∣𝐲\)​\[‖𝐬θ​\(𝐱,t\)−∇𝐱​ln​q​\(𝐱,t∣𝐲\)‖2\]\\displaystyle:=\\frac\{1\}\{2\}\\int\_\{0\}^\{T\}dt\\,\\gamma\(t\)\\,\\mathbb\{E\}\_\{\\mathbf\{y\}\\sim P\_\{\\rm data\}\}\\mathbb\{E\}\_\{\\mathbf\{x\}\\sim q\(\\mathbf\{x\},t\\mid\\mathbf\{y\}\)\}\\left\[\\left\\\|\\mathbf\{s\}\_\{\\theta\}\(\\mathbf\{x\},t\)\-\\nabla\_\{\\mathbf\{x\}\}\\ln q\(\\mathbf\{x\},t\\mid\\mathbf\{y\}\)\\right\\\|^\{2\}\\right\]≥12​∫0Td​t​γ​\(t\)​𝔼𝐱∼Pt​\[‖𝐬θ​\(𝐱,t\)−∇𝐱​ln​\(𝔼𝐲∼Pdata​\[ψ\(𝐲\)​\(𝐱,t\)\]\)‖2\]=Δ​𝒮ideal\.\\displaystyle\\geq\\frac\{1\}\{2\}\\int\_\{0\}^\{T\}dt\\,\\gamma\(t\)\\,\\mathbb\{E\}\_\{\\mathbf\{x\}\\sim P\_\{t\}\}\\left\[\\left\\\|\\mathbf\{s\}\_\{\\theta\}\(\\mathbf\{x\},t\)\-\\nabla\_\{\\mathbf\{x\}\}\\ln\\left\(\\mathbb\{E\}\_\{\\mathbf\{y\}\\sim P\_\{\\rm data\}\}\[\\psi^\{\(\\mathbf\{y\}\)\}\(\\mathbf\{x\},t\)\]\\right\)\\right\\\|^\{2\}\\right\]=\\Delta\\mathcal\{S\}\_\{\\rm ideal\}\.\(2\.19\)Every term inΔ​𝒮DSM\\Delta\\mathcal\{S\}\_\{\\rm DSM\}is computable: training samples a data point𝐲\\mathbf\{y\}, corrupts it through the reference kernel, and regresses𝐬θ​\(𝐱,t\)\\mathbf\{s\}\_\{\\theta\}\(\\mathbf\{x\},t\)onto the conditional score∇𝐱​ln​q​\(𝐱,t∣𝐲\)\\nabla\_\{\\mathbf\{x\}\}\\ln q\(\\mathbf\{x\},t\\mid\\mathbf\{y\}\), which for affine Gaussian corruption is linear in the noise and reduces Eq\. \([2\.19](https://arxiv.org/html/2608.23916#S2.E19)\) to the standard denoising score\-matching loss of\[[30](https://arxiv.org/html/2608.23916#bib.bib30),[31](https://arxiv.org/html/2608.23916#bib.bib31),[2](https://arxiv.org/html/2608.23916#bib.bib2)\]\.

The Jensen gapΔ​𝒮DSM−Δ​𝒮ideal\\Delta\\mathcal\{S\}\_\{\\rm DSM\}\-\\Delta\\mathcal\{S\}\_\{\\rm ideal\}is the price of this substitution\. Since the conditional score is an unbiased estimator of the marginal score, the gap does not shift the minimizer, but it does shift the value of the objective by an additive, model\-independent term; the remainder of the paper studies that term—what it is, what geometry it carries, and what it implies for practice\.

## 3The Fisher Geometry of the Denoising Loss

Section[2](https://arxiv.org/html/2608.23916#S2)obtained the generative dynamics from the*first*variation of a path\-space relative entropy\. We now show that the*second*variation of the same functional determines the geometry of the states the diffusion visits, and that this geometry is exactly what inflates the denoising objective above its ideal value\. The construction proceeds in three steps:\(1\)the second variation defines a canonical quadratic form on the bridge family, measurable with respect to the endpoints alone \(Sec\.[3\.1](https://arxiv.org/html/2608.23916#S3.SS1)\),\(2\)the requirements of a regular, computable corruption select a finite\-dimensional reduction of that form \(Secs\.[3\.2](https://arxiv.org/html/2608.23916#S3.SS2)–[3\.4](https://arxiv.org/html/2608.23916#S3.SS4)\), and\(3\)the Jensen gap of Sec\.[2\.4](https://arxiv.org/html/2608.23916#S2.SS4)is identified with its integrated trace \(Sec\.[3\.5](https://arxiv.org/html/2608.23916#S3.SS5)\)\.Derivations are collected in Appendix[B](https://arxiv.org/html/2608.23916#A2)\.

### 3\.1Endpoint geometry of the bridge

Every solution of the bridge problem is an endpoint tilt of the reference:

lnℋ\(u,v\)\(𝐱0:N\)=u\(𝐱0\)\+v\(𝐱N\)\+ln𝒢\(𝐱0:N\)−Λ\[u,v\],\\ln\\mathcal\{H\}^\{\(u,v\)\}\(\\mathbf\{x\}\_\{0:N\}\)=u\(\\mathbf\{x\}\_\{0\}\)\+v\(\\mathbf\{x\}\_\{N\}\)\+\\ln\\mathcal\{G\}\(\\mathbf\{x\}\_\{0:N\}\)\-\\Lambda\[u,v\],\(3\.1\)withΛ\[u,v\]=ln∫d𝐱0:N𝒢eu⁡\(𝐱0\)\+v⁡\(𝐱N\)\\Lambda\[u,v\]=\\ln\\int d\\mathbf\{x\}\_\{0:N\}\\,\\mathcal\{G\}\\,e^\{u\(\\mathbf\{x\}\_\{0\}\)\+v\(\\mathbf\{x\}\_\{N\}\)\}finite on a domain𝒟\\mathcal\{D\}, andρ\(u,v\)\\rho^\{\(u,v\)\}the joint endpoint law underℋ\(u,v\)\\mathcal\{H\}^\{\(u,v\)\}\.

###### Lemma 3\.1\(Ambient Fisher form\)\.

For\(u,v\)\(u,v\)interior to𝒟\\mathcal\{D\},

δ2​Λ​\[\(δ​u,δ​v\)\]=Varρ\(u,v\)​\[δ​u​\(𝐱0\)\+δ​v​\(𝐱N\)\]\.\\delta^\{2\}\\Lambda\\big\[\(\\delta u,\\delta v\)\\big\]\\;=\\;\\mathrm\{Var\}\_\{\\rho^\{\(u,v\)\}\}\\big\[\\delta u\(\\mathbf\{x\}\_\{0\}\)\+\\delta v\(\\mathbf\{x\}\_\{N\}\)\\big\]\.\(3\.2\)

The proof observes thatΛ\\Lambdais the cumulant\-generating functional of the endpoint evaluations under𝒢\\mathcal\{G\}\(Appendix[B](https://arxiv.org/html/2608.23916#A2)\)\. For the Schrödinger bridge problem𝒢\\mathcal\{G\}need not be a probability measure, owing to the killing potential of Sec\.[2](https://arxiv.org/html/2608.23916#S2), but for the exponential\-family argument it suffices that𝒢\\mathcal\{G\}beσ\\sigma\-finite with the relevant exponential moments, which places the bridge within the framework of infinite\-dimensional exponential families\[[59](https://arxiv.org/html/2608.23916#bib.bib59),[60](https://arxiv.org/html/2608.23916#bib.bib60)\]\.

Expanding Eq\. \([3\.2](https://arxiv.org/html/2608.23916#S3.E2)\) yields

δ2​Λ=Varρ​\[δ​u​\(𝐱0\)\]\+Varρ​\[δ​v​\(𝐱N\)\]\+2​Covρ​\[δ​u​\(𝐱0\),δ​v​\(𝐱N\)\],\\delta^\{2\}\\Lambda=\\mathrm\{Var\}\_\{\\rho\}\[\\delta u\(\\mathbf\{x\}\_\{0\}\)\]\+\\mathrm\{Var\}\_\{\\rho\}\[\\delta v\(\\mathbf\{x\}\_\{N\}\)\]\+2\\,\\mathrm\{Cov\}\_\{\\rho\}\[\\delta u\(\\mathbf\{x\}\_\{0\}\),\\delta v\(\\mathbf\{x\}\_\{N\}\)\],\(3\.3\)so the two endpoint sectors are*not*orthogonal in the ambient geometry\. In short, the cross term carries the endpoint dependence induced by the reference dynamics and the bridge constraints\.

A deeper implication of Eq\. \([3\.2](https://arxiv.org/html/2608.23916#S3.E2)\) concerns the tangent space rather than the metric itself\. For an infinitesimal deformation of the bridge family, the first\-order variation of the log\-likelihood, i\.e\. the tangent vector to the statistical manifold atℋ\(u,v\)\\mathcal\{H\}^\{\(u,v\)\}, is

δ​ℓ=δ​u​\(𝐱0\)\+δ​v​\(𝐱N\)−𝔼⁡\[δ​u​\(𝐱0\)\+δ​v​\(𝐱N\)\],\\delta\\ell=\\delta u\(\\mathbf\{x\}\_\{0\}\)\+\\delta v\(\\mathbf\{x\}\_\{N\}\)\-\\mathbb\{E\}\[\\delta u\(\\mathbf\{x\}\_\{0\}\)\+\\delta v\(\\mathbf\{x\}\_\{N\}\)\],so*every infinitesimal likelihood ratio within the bridge family is measurable with respect to the endpointσ\\sigma\-algebra*\. The endpoint pair thus constitutes a sufficient statistic for the local statistical experiment defined by the tilt family, and Eq\. \([3\.2](https://arxiv.org/html/2608.23916#S3.E2)\) is its second moment\. This places the construction within the classical framework in which Fisher information is monotone under Markov kernels and preserved exactly under sufficient reductions\[[61](https://arxiv.org/html/2608.23916#bib.bib61),[62](https://arxiv.org/html/2608.23916#bib.bib62)\]\.

### 3\.2A regular conditional endpoint family

A natural first candidate for a latent geometry is to condition the path measure on an interior stateXt=𝐱X\_\{t\}=\\mathbf\{x\}and to regard𝐱↦ℋ\(⋅∣Xt=𝐱\)\\mathbf\{x\}\\mapsto\\mathcal\{H\}\(\\cdot\\mid X\_\{t\}=\\mathbf\{x\}\)as a statistical family indexed by the spatial coordinate\. This construction, however, is not regular\. For distinct𝐱≠𝐱′\\mathbf\{x\}\\neq\\mathbf\{x\}^\{\\prime\}, the conditioned measures are supported on the disjoint path sets\{Xt=𝐱\}\\\{X\_\{t\}=\\mathbf\{x\}\\\}and\{Xt=𝐱′\}\\\{X\_\{t\}=\\mathbf\{x\}^\{\\prime\}\\\}and are therefore mutually singular, so that no finite local Kullback–Leibler expansion exists from which a Fisher metric could be extracted\. This obstruction is not a pathology of the bridge formulation but the generic consequence of exact point conditioning for continuous\-path measures\[[63](https://arxiv.org/html/2608.23916#bib.bib63), see, e\.g\., Section 2\.2\], and it rules out exact conditioned path measures as a regular Fisher family, motivating instead a reduction that retains the endpoint variables\.

The sufficiency of the endpoint tangent space suggests such a reduction: retain the endpoint experiment and condition*it*on the latent state,

ιt:𝐱⟼p⁡\(𝐱0,𝐱N∣Xt=𝐱\),\\iota\_\{t\}:\\;\\mathbf\{x\}\\;\\longmapsto\\;p\(\\mathbf\{x\}\_\{0\},\\mathbf\{x\}\_\{N\}\\mid X\_\{t\}=\\mathbf\{x\}\),\(3\.4\)defining a chain of reductions𝒫path→𝒫end→𝒫end\|lat\\mathcal\{P\}\_\{\\rm path\}\\to\\mathcal\{P\}\_\{\\rm end\}\\to\\mathcal\{P\}\_\{\\rm end\\mid lat\}, of which the first step is justified by the endpoint measurability of the bridge tangent space and the second is necessitated by the requirement of a regular parameterization\. Equation \([3\.4](https://arxiv.org/html/2608.23916#S3.E4)\) discards precisely the path information on which the ambient Fisher form does not depend, and while we do not claim that this reduction is unique, it is a minimal reduction consistent with the tangent structure, and it is the one the denoising objective itself uses\.

### 3\.3Gaussian corruption from structural requirements

Equation \([3\.4](https://arxiv.org/html/2608.23916#S3.E4)\) is still completely general, and the diffusion construction used in practice is obtained by imposing what a trainable model requires:

1. \(i\)the latent state lives in the data spaceℝd\\mathbb\{R\}^\{d\}rather than a learned code space;
2. \(ii\)the states form an ordered Markov hierarchy from the data to a fixed, data\-independent prior left invariant by the transition;
3. \(iii\)the corruption is specified before training \(hence state\-independent\), isotropic, and continuous in the continuum limit\.

The ordering already constrains the information flow\. Along the corruption chain, which runs from the data endpoint𝐱N\\mathbf\{x\}\_\{N\}toward the prior endpoint𝐱0\\mathbf\{x\}\_\{0\}, the data\-processing inequality givesI⁡\(𝐱N,𝐱k\)≤I⁡\(𝐱N,𝐱k\+1\)I\(\\mathbf\{x\}\_\{N\};\\mathbf\{x\}\_\{k\}\)\\leq I\(\\mathbf\{x\}\_\{N\};\\mathbf\{x\}\_\{k\+1\}\), and invariance of the prior under the transition givesDK​L\(Pk∥Pprior\)≤DK​L\(Pk\+1∥Pprior\)D\_\{KL\}\(P\_\{k\}\\\|P\_\{\\rm prior\}\)\\leq D\_\{KL\}\(P\_\{k\+1\}\\\|P\_\{\\rm prior\}\)whenever the step fromk\+1k\{\+\}1tokkmoves along the chain, so that information about the data can only be discarded, never created, as corruption proceeds\[[64](https://arxiv.org/html/2608.23916#bib.bib64)\]\.

A state\-independent, isotropic, continuous Markov corruption is an additive diffusiond​𝐱t=𝐛⁡\(t\)​d​t\+γ⁡\(t\)​d​𝐖td\\mathbf\{x\}\_\{t\}=\\mathbf\{b\}\(t\)dt\+\\sqrt\{\\gamma\(t\)\}\\,d\\mathbf\{W\}\_\{t\}, whose finite\-time kernel is Gaussian,

q⁡\(𝐱,t∣𝐱N\)=𝒩⁡\(𝐱,αt​𝐱N,σt2​𝐈\),q\(\\mathbf\{x\},t\\mid\\mathbf\{x\}\_\{N\}\)=\\mathcal\{N\}\\big\(\\mathbf\{x\};\\ \\alpha\_\{t\}\\mathbf\{x\}\_\{N\},\\ \\sigma\_\{t\}^\{2\}\\mathbf\{I\}\\big\),\(3\.5\)after absorbing the deterministic drift intoαt\\alpha\_\{t\}\. We present Eq\. \([3\.5](https://arxiv.org/html/2608.23916#S3.E5)\) not as a uniqueness theorem for all corruption processes but as the natural continuous realization of requirements \([\(i\)](https://arxiv.org/html/2608.23916#S3.I2.i1)\)–\([\(iii\)](https://arxiv.org/html/2608.23916#S3.I2.i3)\) and the standard family used in practice\[[2](https://arxiv.org/html/2608.23916#bib.bib2),[4](https://arxiv.org/html/2608.23916#bib.bib4),[20](https://arxiv.org/html/2608.23916#bib.bib20)\]; relaxing state independence, isotropy, or continuity leads to broader model classes discussed in Sec\.[6](https://arxiv.org/html/2608.23916#S6)\.

Read as a function of the clean endpoint, Eq\. \([3\.5](https://arxiv.org/html/2608.23916#S3.E5)\) exhibits a finite sufficient statistic:

ln⁡q⁡\(𝐱,t∣𝐱N\)=A⁡\(𝐱,t\)⋅𝐬⁡\(𝐱N\)\+B⁡\(𝐱,t\)\+C⁡\(𝐱N\),𝐬⁡\(𝐱N\)=\(𝐱N,−12​‖𝐱N‖2\),\\ln q\(\\mathbf\{x\},t\\mid\\mathbf\{x\}\_\{N\}\)=A\(\\mathbf\{x\},t\)\\cdot\\mathbf\{s\}\(\\mathbf\{x\}\_\{N\}\)\+B\(\\mathbf\{x\},t\)\+C\(\\mathbf\{x\}\_\{N\}\),\\qquad\\mathbf\{s\}\(\\mathbf\{x\}\_\{N\}\)=\\Big\(\\mathbf\{x\}\_\{N\},\\,\-\\tfrac\{1\}\{2\}\\left\\\|\\mathbf\{x\}\_\{N\}\\right\\\|^\{2\}\\Big\),\(3\.6\)withd\+1d\+1components:ddfirst moments coupling to the spatial latent coordinates and one second moment coupling to the noise coordinate\. The extra dimension of latent spacetime is thus the second\-moment statistic of the corruption\. Isotropy is what keeps this statistic one\-dimensional: an anisotropic Gaussian requires the full quadratic−12𝐱N⊗𝐱N\-\\tfrac\{1\}\{2\}\\mathbf\{x\}\_\{N\}\\otimes\\mathbf\{x\}\_\{N\}, givingm=d\+d⁡\(d\+1\)/2m=d\+d\(d\{\+\}1\)/2natural coordinates ford\+1d\+1latent coordinates, so the latent family would acquire codimension and become a*curved*exponential family\[[65](https://arxiv.org/html/2608.23916#bib.bib65),[66](https://arxiv.org/html/2608.23916#bib.bib66)\]\. Absorbing the𝐱\\mathbf\{x\}\-independent factors into a carrierν=eC​g\\nu=e^\{C\}gyields

p⁡\(𝐱N∣𝐱,t\)=exp⁡\[A⁡\(𝐱,t\)⋅𝐬⁡\(𝐱N\)−Λ⁡\(A⁡\(𝐱,t\)\)\]​ν​\(𝐱N\),p\(\\mathbf\{x\}\_\{N\}\\mid\\mathbf\{x\},t\)=\\exp\\big\[A\(\\mathbf\{x\},t\)\\cdot\\mathbf\{s\}\(\\mathbf\{x\}\_\{N\}\)\-\\Lambda\(A\(\\mathbf\{x\},t\)\)\\big\]\\,\\nu\(\\mathbf\{x\}\_\{N\}\),\(3\.7\)in which the learned model enters only throughν\\nu, while the map to natural coordinates is fixed by the corruption\. Explicitly, withSNRt=αt2/σt2\\mathrm\{SNR\}\_\{t\}=\\alpha\_\{t\}^\{2\}/\\sigma\_\{t\}^\{2\}and𝐱~=𝐱/αt\\tilde\{\\mathbf\{x\}\}=\\mathbf\{x\}/\\alpha\_\{t\},

A⁡\(𝐱,t\)=\(αtσt2​𝐱,αt2σt2\)=SNRt⋅\(𝐱~,1\)\.A\(\\mathbf\{x\},t\)\\;=\\;\\Big\(\\frac\{\\alpha\_\{t\}\}\{\\sigma\_\{t\}^\{2\}\}\\mathbf\{x\},\\ \\frac\{\\alpha\_\{t\}^\{2\}\}\{\\sigma\_\{t\}^\{2\}\}\\Big\)\\;=\\;\\mathrm\{SNR\}\_\{t\}\\cdot\\big\(\\tilde\{\\mathbf\{x\}\},\\,1\\big\)\.\(3\.8\)Latent spacetime is therefore a cone in natural\-parameter space: the direction ofAAis the rescaled spatial coordinate and its magnitude is the signal\-to\-noise ratio\. For a nondegenerate schedule, the Jacobian ofAAhas rankd\+1d\+1, so the latent coordinates locally parameterize an open subset of natural\-parameter space and inherit its dual affine structure, while a monotone reparametrization of time slides along the ray Eq\. \([3\.8](https://arxiv.org/html/2608.23916#S3.E8)\) without changing it—the geometric content of the schedule reparametrization invariance of diffusion objectives\[[20](https://arxiv.org/html/2608.23916#bib.bib20)\], to which we return in Sec\.[4](https://arxiv.org/html/2608.23916#S4)\.

### 3\.4The pullback metric

The immersionιlat:\(𝐱,t\)↦A⁡\(𝐱,t\)\\iota\_\{\\rm lat\}:\(\\mathbf\{x\},t\)\\mapsto A\(\\mathbf\{x\},t\)carries the ambient form to latent spacetime, and for coordinateszμ,zν∈\{x1,…,xd,t\}z^\{\\mu\},z^\{\\nu\}\\in\\\{x^\{1\},\\dots,x^\{d\},t\\\}the induced metric factorizes as

gμ​ν=∂μAa​∂νAb⏟fixed by the schedule​Covp⁡\(𝐱N∣𝐱,t\)​\[𝐬a,𝐬b\]⏟data and model\.g\_\{\\mu\\nu\}\\;=\\;\\underbrace\{\\partial\_\{\\mu\}A^\{a\}\\,\\partial\_\{\\nu\}A^\{b\}\}\_\{\\text\{fixed by the schedule\}\}\\;\\underbrace\{\\mathrm\{Cov\}\_\{p\(\\mathbf\{x\}\_\{N\}\\mid\\mathbf\{x\},t\)\}\\big\[\\mathbf\{s\}\_\{a\},\\mathbf\{s\}\_\{b\}\\big\]\}\_\{\\text\{data and model\}\}\.\(3\.9\)Once the corruption and its sufficient statistic are fixed, the metric follows from the second variation, and the two factors separate cleanly due to state independence of the corruption\.

###### Proposition 3\.2\(Endpoint split\)\.

The two endpoints are conditionally independent given the latent state,p⁡\(𝐱0,𝐱N∣Xt=𝐱\)=p⁡\(𝐱0∣Xt=𝐱\)​p​\(𝐱N∣Xt=𝐱\)p\(\\mathbf\{x\}\_\{0\},\\mathbf\{x\}\_\{N\}\\mid X\_\{t\}=\\mathbf\{x\}\)=p\(\\mathbf\{x\}\_\{0\}\\mid X\_\{t\}=\\mathbf\{x\}\)\\,p\(\\mathbf\{x\}\_\{N\}\\mid X\_\{t\}=\\mathbf\{x\}\), and consequently

gμ​ν=gμ​νpast\+gμ​νfuture\.g\_\{\\mu\\nu\}=g^\{\\rm past\}\_\{\\mu\\nu\}\+g^\{\\rm future\}\_\{\\mu\\nu\}\.\(3\.10\)

The endpoint sectors are not orthogonal in the ambient metric Eq\. \([3\.3](https://arxiv.org/html/2608.23916#S3.E3)\); the cross term vanishes only after conditioning on the intermediate state, by the Markov property\. The past sectorgpastg^\{\\rm past\}is the Fisher information of the seed posteriorp⁡\(𝐱0∣𝐱,t\)p\(\\mathbf\{x\}\_\{0\}\\mid\\mathbf\{x\},t\), measuring how much the corrupted state reveals about its initialization, and it is generically nonzero yet invisible to denoising score matching, whose regression target involves the clean endpoint alone\.

###### Proposition 3\.3\(The future sector is the loss metric\)\.

For the corruption Eq\. \([3\.5](https://arxiv.org/html/2608.23916#S3.E5)\),

gi​jfuture​\(𝐱,t\)=Covp⁡\(𝐱N∣𝐱,t\)\[∂ilnq,∂jlnq\]=αt2σt4Cov\[xNi,xNj∣Xt=𝐱\]=1σt2​δi​j\+∂i∂jln⁡Pt​\(𝐱\)\.\\begin\{split\}g^\{\\rm future\}\_\{ij\}\(\\mathbf\{x\},t\)&=\\mathrm\{Cov\}\_\{p\(\\mathbf\{x\}\_\{N\}\\mid\\mathbf\{x\},t\)\}\\big\[\\partial\_\{i\}\\ln q,\\ \\partial\_\{j\}\\ln q\\big\]=\\frac\{\\alpha\_\{t\}^\{2\}\}\{\\sigma\_\{t\}^\{4\}\}\\mathrm\{Cov\}\\big\[x\_\{N\}^\{i\},x\_\{N\}^\{j\}\\mid X\_\{t\}=\\mathbf\{x\}\\big\]\\\\ &=\\frac\{1\}\{\\sigma\_\{t\}^\{2\}\}\\delta\_\{ij\}\+\\partial\_\{i\}\\partial\_\{j\}\\ln P\_\{t\}\(\\mathbf\{x\}\)\.\\end\{split\}\(3\.11\)

Equation \([3\.11](https://arxiv.org/html/2608.23916#S3.E11)\) is the tensor whose trace appears in the decomposition theorem below\. This metric has been studied descriptively in recent work\[[47](https://arxiv.org/html/2608.23916#bib.bib47),[48](https://arxiv.org/html/2608.23916#bib.bib48)\]and the present derivation shows that it follows from the second variation of the bridge objective\. It is worth emphasizing two aspect of this construction\. First, the full latent metric is Eq\. \([3\.10](https://arxiv.org/html/2608.23916#S3.E10)\), while the training objective sees only its future sector, so the geometry is not an ad hoc construction reverse\-engineered from the loss but contains strictly more structure than the loss probes\. Second, the cone structure makes the information hierarchy quantitative\. WritingK\(θ\)=DK​L\(pθ∥ν\)K\(\\theta\)=D\_\{KL\}\(p\_\{\\theta\}\\\|\\nu\)for the information the latent carries about the clean endpoint, the exponential\-family identity∇θK=g⁡\(θ\)​θ\\nabla\_\{\\theta\}K=g\(\\theta\)\\thetagives, along the rayθ=r​n\\theta=r\\,nof Eq\. \([3\.8](https://arxiv.org/html/2608.23916#S3.E8)\),

d​Kd​r=r​gr​r,gr​r=n⊤​Covp​\[𝐬\]​n\.\\frac\{dK\}\{dr\}\\;=\\;r\\,g\_\{rr\},\\qquad g\_\{rr\}=n^\{\\top\}\\mathrm\{Cov\}\_\{p\}\[\\mathbf\{s\}\]\\,n\.\(3\.12\)Since the radial coordinate is the signal\-to\-noise ratio, the information the latent retains about the data decays along the corruption at a rate set by the radial component of the Fisher metric, which is the pointwise counterpart of the integrated identity of Sec\.[4](https://arxiv.org/html/2608.23916#S4)and a first indication that the loss floor and an accumulated information are the same quantity rather than merely equal numbers\.

The construction of this section assembles classical components\. The identification of the Fisher metric with the local Kullback–Leibler form and its monotonicity under Markov kernels are due to Chentsov\[[61](https://arxiv.org/html/2608.23916#bib.bib61),[62](https://arxiv.org/html/2608.23916#bib.bib62)\], the endpoint\-tilt family is an instance of the infinite\-dimensional exponential families of Pistone and Sempi\[[59](https://arxiv.org/html/2608.23916#bib.bib59),[60](https://arxiv.org/html/2608.23916#bib.bib60)\], the induced geometry of a finite\-dimensional submanifold is the theory of curved exponential families of Efron and Amari\[[65](https://arxiv.org/html/2608.23916#bib.bib65),[66](https://arxiv.org/html/2608.23916#bib.bib66)\], and the reparametrization invariance in the SNR coordinate is standard\[[20](https://arxiv.org/html/2608.23916#bib.bib20)\]\. What is new here is the assembly: the ambient form is fixed by the bridge variational principle, and the latent metric is a pullback of that fixed form rather than a separate postulate\.

### 3\.5The irreducible loss floor

It remains to identify the geometry of Secs\.[3\.1](https://arxiv.org/html/2608.23916#S3.SS1)–[3\.4](https://arxiv.org/html/2608.23916#S3.SS4)with the Jensen gap of Sec\.[2\.4](https://arxiv.org/html/2608.23916#S2.SS4)\. Denote byq⁡\(𝐲∣𝐱,t\)=Pdata​\(𝐲\)​q​\(𝐱,t∣𝐲\)/P⁡\(𝐱,t\)q\(\\mathbf\{y\}\\mid\\mathbf\{x\},t\)=P\_\{\\rm data\}\(\\mathbf\{y\}\)\\,q\(\\mathbf\{x\},t\\mid\\mathbf\{y\}\)/P\(\\mathbf\{x\},t\)the posterior over clean data given the noisy observation, and write∇ln⁡q​\(𝐲∣𝐱,t\)\\nabla\\ln q\(\\mathbf\{y\}\\mid\\mathbf\{x\},t\)for its score in𝐱\\mathbf\{x\}\. Expanding the squared norm in Eq\. \([2\.19](https://arxiv.org/html/2608.23916#S2.E19)\) around the ideal target splits the gap into two contributions:

Δ​𝒮DSM−Δ​𝒮ideal=12∫0Tdtγ\(t\)𝔼𝐱∼Pt𝔼𝐲∼q⁡\(𝐲∣𝐱,t\)\{‖∇lnq\(𝐲∣𝐱,t\)‖2−2∇lnq\(𝐲∣𝐱,t\)⋅\(𝐬θ−∇lnP\(𝐱,t\)\)\}\.\\displaystyle\\begin\{split\}\\Delta\\mathcal\{S\}\_\{\\rm DSM\}\-\\Delta\\mathcal\{S\}\_\{\\rm ideal\}&=\\frac\{1\}\{2\}\\int\_\{0\}^\{T\}dt\\,\\gamma\(t\)\\,\\mathbb\{E\}\_\{\\mathbf\{x\}\\sim P\_\{t\}\}\\mathbb\{E\}\_\{\\mathbf\{y\}\\sim q\(\\mathbf\{y\}\\mid\\mathbf\{x\},t\)\}\\Big\\\{\\\\ &\\qquad\\left\\\|\\nabla\\ln q\(\\mathbf\{y\}\\mid\\mathbf\{x\},t\)\\right\\\|^\{2\}\-2\\,\\nabla\\ln q\(\\mathbf\{y\}\\mid\\mathbf\{x\},t\)\\cdot\\big\(\\mathbf\{s\}\_\{\\theta\}\-\\nabla\\ln P\(\\mathbf\{x\},t\)\\big\)\\Big\\\}\.\\end\{split\}\(3\.13\)The cross term vanishes identically for*every*θ\\theta, not merely at convergence: by the score identity

𝔼𝐲∼q\(⋅∣𝐱,t\)\[∇𝐱lnq\(𝐲∣𝐱,t\)\]=∇𝐱∫d𝐲q\(𝐲∣𝐱,t\)=0,\\mathbb\{E\}\_\{\\mathbf\{y\}\\sim q\(\\cdot\\mid\\mathbf\{x\},t\)\}\\left\[\\nabla\_\{\\mathbf\{x\}\}\\ln q\(\\mathbf\{y\}\\mid\\mathbf\{x\},t\)\\right\]=\\nabla\_\{\\mathbf\{x\}\}\\int d\\mathbf\{y\}\\,q\(\\mathbf\{y\}\\mid\\mathbf\{x\},t\)=0,\(3\.14\)its posterior expectation is zero point\-wise in𝐱\\mathbf\{x\}, so only the first, model\-independent term survives and the gap carries no dependence on the model:

Δ​𝒮DSM−Δ​𝒮ideal\\displaystyle\\Delta\\mathcal\{S\}\_\{\\rm DSM\}\-\\Delta\\mathcal\{S\}\_\{\\rm ideal\}=12​∫0Td​t​γ​\(t\)​𝔼𝐱∼Pt​𝔼𝐲∼q⁡\(𝐲∣𝐱,t\)​‖∇𝐱​ln​q​\(𝐲∣𝐱,t\)‖2\\displaystyle=\\frac\{1\}\{2\}\\int\_\{0\}^\{T\}dt\\,\\gamma\(t\)\\;\\mathbb\{E\}\_\{\\mathbf\{x\}\\sim P\_\{t\}\}\\mathbb\{E\}\_\{\\mathbf\{y\}\\sim q\(\\mathbf\{y\}\\mid\\mathbf\{x\},t\)\}\\left\\\|\\nabla\_\{\\mathbf\{x\}\}\\ln q\(\\mathbf\{y\}\\mid\\mathbf\{x\},t\)\\right\\\|^\{2\}=12​∫0Td​t​γ​\(t\)​𝔼𝐱∼Pt​\[tr​Cov𝐲\|𝐱​\[∇𝐱​ln​q​\(𝐱,t∣𝐲\)\]\],\\displaystyle=\\frac\{1\}\{2\}\\int\_\{0\}^\{T\}dt\\,\\gamma\(t\)\\;\\mathbb\{E\}\_\{\\mathbf\{x\}\\sim P\_\{t\}\}\\left\[\\mathrm\{tr\}\\,\\mathrm\{Cov\}\_\{\\mathbf\{y\}\\mid\\mathbf\{x\}\}\\\!\\left\[\\nabla\_\{\\mathbf\{x\}\}\\ln q\(\\mathbf\{x\},t\\mid\\mathbf\{y\}\)\\right\]\\right\],\(3\.15\)where the second equality uses Bayes’ rule,∇𝐱​ln​q​\(𝐲∣𝐱,t\)=∇𝐱​ln​q​\(𝐱,t∣𝐲\)−∇𝐱​ln​P​\(𝐱,t\)\\nabla\_\{\\mathbf\{x\}\}\\ln q\(\\mathbf\{y\}\\mid\\mathbf\{x\},t\)=\\nabla\_\{\\mathbf\{x\}\}\\ln q\(\\mathbf\{x\},t\\mid\\mathbf\{y\}\)\-\\nabla\_\{\\mathbf\{x\}\}\\ln P\(\\mathbf\{x\},t\), by which the posterior score is the centered conditional score\.

###### Theorem 3\.4\(The loss floor is a Fisher information\)\.

Letq⁡\(𝐱,t∣𝐲\)q\(\\mathbf\{x\},t\\mid\\mathbf\{y\}\)be a corruption kernel satisfying the following regularity conditions:q⁡\(𝐱,t∣𝐲\)\>0q\(\\mathbf\{x\},t\\mid\\mathbf\{y\}\)\>0on the support of interest;𝐱↦ln⁡q⁡\(𝐱,t∣𝐲\)\\mathbf\{x\}\\mapsto\\ln q\(\\mathbf\{x\},t\\mid\\mathbf\{y\}\)is differentiable; and the gradient∇𝐱q​\(𝐱,t∣𝐲\)\\nabla\_\{\\mathbf\{x\}\}q\(\\mathbf\{x\},t\\mid\\mathbf\{y\}\)is uniformly dominated by an integrable function of𝐲\\mathbf\{y\}in a sufficiently small neighbourhood of each𝐱\\mathbf\{x\}, so that differentiation under the integral sign is permitted\. LetPtP\_\{t\}be the induced marginal and let

p⁡\(𝐲∣𝐱,t\)=Pdata​\(𝐲\)​q​\(𝐱,t∣𝐲\)Pt​\(𝐱\)p\(\\mathbf\{y\}\\mid\\mathbf\{x\},t\)\\;=\\;\\frac\{P\_\{\\rm data\}\(\\mathbf\{y\}\)\\,q\(\\mathbf\{x\},t\\mid\\mathbf\{y\}\)\}\{P\_\{t\}\(\\mathbf\{x\}\)\}\(3\.16\)be the*conditional endpoint family*: the family of laws over clean endpoints𝐲\\mathbf\{y\}indexed by the latent coordinate𝐱\\mathbf\{x\}, which plays the role of the parameter\. Let

gi​j​\(𝐱,t\):=𝔼𝐲\|𝐱​\[∂iln⁡p⁡\(𝐲∣𝐱,t\)​∂jln⁡p⁡\(𝐲∣𝐱,t\)\]g\_\{ij\}\(\\mathbf\{x\},t\)\\;:=\\;\\mathbb\{E\}\_\{\\mathbf\{y\}\\mid\\mathbf\{x\}\}\\big\[\\partial\_\{i\}\\ln p\(\\mathbf\{y\}\\mid\\mathbf\{x\},t\)\\,\\partial\_\{j\}\\ln p\(\\mathbf\{y\}\\mid\\mathbf\{x\},t\)\\big\]\(3\.17\)be its Fisher–Rao metric\. Then

Δ​𝒮DSM−Δ​𝒮ideal=12​∫0Td​t​γ​\(t\)​𝔼𝐱∼Pt​\[tr⁡g⁡\(𝐱,t\)\]\.\\Delta\\mathcal\{S\}\_\{\\rm DSM\}\-\\Delta\\mathcal\{S\}\_\{\\rm ideal\}\\;=\\;\\frac\{1\}\{2\}\\int\_\{0\}^\{T\}dt\\,\\gamma\(t\)\\;\\mathbb\{E\}\_\{\\mathbf\{x\}\\sim P\_\{t\}\}\\big\[\\operatorname\{tr\}g\(\\mathbf\{x\},t\)\\big\]\.\(3\.18\)

###### Proof\.

Sincep⁡\(𝐲∣𝐱,t\)=Pdata​\(𝐲\)​q​\(𝐱,t∣𝐲\)/Pt​\(𝐱\)p\(\\mathbf\{y\}\\mid\\mathbf\{x\},t\)=P\_\{\\rm data\}\(\\mathbf\{y\}\)q\(\\mathbf\{x\},t\\mid\\mathbf\{y\}\)/P\_\{t\}\(\\mathbf\{x\}\)andPdata​\(𝐲\)P\_\{\\rm data\}\(\\mathbf\{y\}\)does not depend on𝐱\\mathbf\{x\}, differentiating the logarithm with respect to𝐱\\mathbf\{x\}gives

∇𝐱​ln​p​\(𝐲∣𝐱,t\)=∇𝐱​ln​q​\(𝐱,t∣𝐲\)−∇𝐱​ln​Pt​\(𝐱\)\.\\nabla\_\{\\mathbf\{x\}\}\\ln p\(\\mathbf\{y\}\\mid\\mathbf\{x\},t\)\\;=\\;\\nabla\_\{\\mathbf\{x\}\}\\ln q\(\\mathbf\{x\},t\\mid\\mathbf\{y\}\)\\;\-\\;\\nabla\_\{\\mathbf\{x\}\}\\ln P\_\{t\}\(\\mathbf\{x\}\)\.\(3\.19\)Taking𝔼𝐲\|𝐱\\mathbb\{E\}\_\{\\mathbf\{y\}\\mid\\mathbf\{x\}\}of Eq\. \([3\.19](https://arxiv.org/html/2608.23916#S3.E19)\) and using𝔼𝐲\|𝐱\[∇𝐱lnp\]=∇𝐱∫d𝐲p\(𝐲∣𝐱,t\)=0\\mathbb\{E\}\_\{\\mathbf\{y\}\\mid\\mathbf\{x\}\}\[\\nabla\_\{\\mathbf\{x\}\}\\ln p\]=\\nabla\_\{\\mathbf\{x\}\}\\\!\\int d\\mathbf\{y\}\\,p\(\\mathbf\{y\}\\mid\\mathbf\{x\},t\)=0recovers the unbiasedness identity of denoising score matching: the conditional score is an unbiased estimator of the marginal score\. The posterior score is therefore centered under the posterior, and its second moment equals its covariance,

𝔼𝐲\|𝐱​\[∇𝐱​ln​p​\(𝐲∣𝐱,t\)​∇𝐱​ln⁡p​\(𝐲∣𝐱,t\)⊤\]=g⁡\(𝐱,t\),\\mathbb\{E\}\_\{\\mathbf\{y\}\\mid\\mathbf\{x\}\}\\left\[\\nabla\_\{\\mathbf\{x\}\}\\ln p\(\\mathbf\{y\}\\mid\\mathbf\{x\},t\)\\,\\nabla\_\{\\mathbf\{x\}\}\\ln p\(\\mathbf\{y\}\\mid\\mathbf\{x\},t\)^\{\\top\}\\right\]=g\(\\mathbf\{x\},t\),which is simultaneously the Fisher–Rao metric \([3\.17](https://arxiv.org/html/2608.23916#S3.E17)\) and the integrand of Eq\. \([3\.15](https://arxiv.org/html/2608.23916#S3.E15)\)\. Taking the trace and inserting into Eq\. \([3\.15](https://arxiv.org/html/2608.23916#S3.E15)\) gives Eq\. \([3\.18](https://arxiv.org/html/2608.23916#S3.E18)\)\. ∎

Beyond the stated regularity conditions, the theorem assumes nothing about the corruption kernel, such as Gaussianity, affine structure, or exponential\-family form, and its content is a conditional\-variance identity\. The ideal objectiveΔ​𝒮ideal\\Delta\\mathcal\{S\}\_\{\\rm ideal\}compares the model against∇ln⁡Pt\\nabla\\ln P\_\{t\}, which by Eq\. \([3\.19](https://arxiv.org/html/2608.23916#S3.E19)\) is the posterior*mean*of the conditional score\. The implementable objectiveΔ​𝒮DSM\\Delta\\mathcal\{S\}\_\{\\rm DSM\}cannot evaluate that mean and substitutes a single draw𝐲∼p\(⋅∣𝐱,t\)\\mathbf\{y\}\\sim p\(\\cdot\\mid\\mathbf\{x\},t\), which is unbiased but noisy\. The point\-wise decomposition now reads

𝔼​‖𝐬θ−𝐬cond‖2⏟denoising objective=‖𝐬θ−∇ln⁡Pt‖2⏟estimation error\+Var⁡\(𝐬cond∣𝐱\)⏟irreducible conditional variance,\\underbrace\{\\mathbb\{E\}\\big\\\|\\mathbf\{s\}\_\{\\theta\}\-\\mathbf\{s\}\_\{\\rm cond\}\\big\\\|^\{2\}\}\_\{\\text\{denoising objective\}\}\\;=\\;\\underbrace\{\\big\\\|\\mathbf\{s\}\_\{\\theta\}\-\\nabla\\\!\\ln P\_\{t\}\\big\\\|^\{2\}\}\_\{\\text\{estimation error\}\}\\;\+\\;\\underbrace\{\\mathrm\{Var\}\\big\(\\mathbf\{s\}\_\{\\rm cond\}\\mid\\mathbf\{x\}\\big\)\}\_\{\\text\{irreducible conditional variance\}\},\(3\.20\)in which the second term on the right\-hand side is the Fisher information of the latent immersionιt:𝐱↦p\(⋅∣𝐱,t\)\\iota\_\{t\}:\\mathbf\{x\}\\mapsto p\(\\cdot\\mid\\mathbf\{x\},t\)of Eq\. \([3\.4](https://arxiv.org/html/2608.23916#S3.E4)\)\. This is the same computation as that of a bias–variance decomposition, though the first term is a squared estimation error rather than a bias, since𝐬θ\\mathbf\{s\}\_\{\\theta\}is deterministic given𝐱\\mathbf\{x\}\. Since the conditional variance of a regression target does not involve the regressor, the floor cannot depend onθ\\theta\. In the language of Sec\.[3\.2](https://arxiv.org/html/2608.23916#S3.SS2):*the latent Fisher metric is the covariance of the random target introduced by denoising score matching*, and the loss floor is its integrated trace along the generative flow\.

For the affine Gaussian kernel, the conditional score is\(αt​𝐲−𝐱\)/σt2\(\\alpha\_\{t\}\\mathbf\{y\}\-\\mathbf\{x\}\)/\\sigma\_\{t\}^\{2\}, and Eq\. \([3\.19](https://arxiv.org/html/2608.23916#S3.E19)\) yields∇𝐱​ln​p​\(𝐲∣𝐱,t\)=\(αt/σt2\)​\(𝐲−𝔼⁡\[𝐲∣𝐱\]\)\\nabla\_\{\\mathbf\{x\}\}\\ln p\(\\mathbf\{y\}\\mid\\mathbf\{x\},t\)=\(\\alpha\_\{t\}/\\sigma\_\{t\}^\{2\}\)\(\\mathbf\{y\}\-\\mathbb\{E\}\[\\mathbf\{y\}\\mid\\mathbf\{x\}\]\), so the metric reduces to the closed form Eq\. \([3\.11](https://arxiv.org/html/2608.23916#S3.E11)\)\.

## 4The Loss Floor as Information Flow

Theorem[3\.4](https://arxiv.org/html/2608.23916#S3.Thmtheorem4)identified the irreducible part of the denoising objective with a Fisher–Rao metric of the endpoint posterior, and for affine kernels, the explicit expression is given by Eq\. \([3\.11](https://arxiv.org/html/2608.23916#S3.E11)\)\. The remaining question is what the metric measures when accumulated along the diffusion trajectory\. Remarkably, the time integral has a simple information\-theoretic form: it is the amount of information about the clean endpoint that is lost under corruption\.

This interpretation also clarifies several properties of the loss floor that otherwise appear unrelated\. The familiar schedule invariance of the diffusion objective becomes a consequence of reparametrization invariance\. Moreover, the divergence near the data end is controlled by the information dimension of the data, and the decomposition admits a natural thermodynamic interpretation\. Proofs are collected in Appendix[C](https://arxiv.org/html/2608.23916#A3), with numerical details in Appendix[D](https://arxiv.org/html/2608.23916#A4)\.

### 4\.1The mutual\-information identity

The Fisher–Rao representation of Theorem[3\.4](https://arxiv.org/html/2608.23916#S3.Thmtheorem4)is local in the corruption parameter\. To understand the accumulated floor, we first ask what its instantaneous integrand measures\. Although the explicit formulas below are often encountered in the special case of affine Gaussian corruption, the first step requires only the corruption kernel itself\. Let

J⁡\(p\):=𝔼p​‖∇ln⁡p‖2J\(p\):=\\mathbb\{E\}\_\{p\}\\\|\\nabla\\ln p\\\|^\{2\}denote the Fisher information of a density with respect to translations\.

The answer to the above question is particularly simple\. Averaged over the noisy marginal, the trace of the posterior Fisher metric is exactly the Fisher information lost when the endpoint\-conditioned corruption kernels are mixed over the data distribution\.

###### Lemma 4\.1\(Fisher\-information form of the floor\)\.

For any corruption kernel,

𝔼Pt\[trg\(⋅,t\)\]=𝔼𝐲\[J\(qt\(⋅∣𝐲\)\)\]−J\(Pt\)\.\\mathbb\{E\}\_\{P\_\{t\}\}\\big\[\\operatorname\{tr\}g\(\\cdot,t\)\\big\]=\\mathbb\{E\}\_\{\\mathbf\{y\}\}\\big\[J\(q\_\{t\}\(\\cdot\\mid\\mathbf\{y\}\)\)\\big\]\-J\(P\_\{t\}\)\.\(4\.1\)

Thus the instantaneous floor is the gap between the Fisher information available when the clean endpoint𝐲\\mathbf\{y\}is known and that remaining after𝐲\\mathbf\{y\}has been marginalized out\. Its non\-negativity is the corresponding monotonicity of Fisher information under mixing\[[64](https://arxiv.org/html/2608.23916#bib.bib64)\]\.

To turn this local identity into a statement about the full loss floor, we now let the corruption be generated continuously by a diffusion\. Suppose that

d​𝐱t=f⁡\(𝐱t,t\)​d​t\+D⁡\(t\)​d​𝐖td\\mathbf\{x\}\_\{t\}=f\(\\mathbf\{x\}\_\{t\},t\)\\,dt\+\\sqrt\{D\(t\)\}\\,d\\mathbf\{W\}\_\{t\}starts from the data distribution, whereffandDDare independent of𝐲\\mathbf\{y\}, andttincreases in the corruption direction\. We writeD⁡\(t\)D\(t\)rather thanγ⁡\(t\)\\gamma\(t\)to distinguish the corruption process considered here from the bridge dynamics of Sec\.[2](https://arxiv.org/html/2608.23916#S2)\. Along such a diffusion, the same Fisher\-information gap determines the rate at which the noisy variable forgets the clean endpoint:

###### Lemma 4\.2\(de Bruijn identity along a corruption diffusion\)\.

Under the assumptions above,

dd​tI\(𝐲;𝐱t\)=−D⁡\(t\)2\(𝔼𝐲J\(qt\(⋅∣𝐲\)\)−J\(Pt\)\)=−D⁡\(t\)2𝔼Pt\[trg\(⋅,t\)\]≤0\.\\frac\{d\}\{dt\}I\(\\mathbf\{y\};\\mathbf\{x\}\_\{t\}\)=\-\\frac\{D\(t\)\}\{2\}\\Big\(\\mathbb\{E\}\_\{\\mathbf\{y\}\}J\(q\_\{t\}\(\\cdot\\mid\\mathbf\{y\}\)\)\-J\(P\_\{t\}\)\\Big\)=\-\\frac\{D\(t\)\}\{2\}\\,\\mathbb\{E\}\_\{P\_\{t\}\}\\big\[\\operatorname\{tr\}g\(\\cdot,t\)\\big\]\\leq 0\.\(4\.2\)

Eq \([4\.2](https://arxiv.org/html/2608.23916#S4.E2)\) gives the information\-theoretic meaning of the metric trace\. It is, up to the local diffusion scale, the instantaneous rate at which information about the clean endpoint is erased by the corruption process\. Integrating this identity converts the local Fisher geometry of the denoising loss into a global information\-flow law\. The de Bruijn identity relates entropy production under Gaussian diffusion to Fisher information\[[64](https://arxiv.org/html/2608.23916#bib.bib64),[49](https://arxiv.org/html/2608.23916#bib.bib49)\], and we have shown that the same structure appears as the difference between the conditional and marginal Fisher informations and, through Theorem[3\.4](https://arxiv.org/html/2608.23916#S3.Thmtheorem4), as the trace of the endpoint\-posterior metric\.

###### Theorem 4\.3\(General factorization of the floor\)\.

Under the hypotheses of Lemma[4\.2](https://arxiv.org/html/2608.23916#S4.Thmtheorem2), for an arbitrary loss weightingw⁡\(t\)w\(t\),

Δ𝒮DSM−Δ𝒮ideal=12∫w\(t\)𝔼Pt\[trg\]dt=−∫w⁡\(t\)D⁡\(t\)⏟scheduled​I​\(𝐲,𝐱t\)d​t⏟datadt\.\\Delta\\mathcal\{S\}\_\{\\rm DSM\}\-\\Delta\\mathcal\{S\}\_\{\\rm ideal\}\\;=\\;\\frac\{1\}\{2\}\\int w\(t\)\\,\\mathbb\{E\}\_\{P\_\{t\}\}\\big\[\\operatorname\{tr\}g\\big\]\\,dt\\;=\\;\-\\int\\underbrace\{\\frac\{w\(t\)\}\{D\(t\)\}\}\_\{\\text\{schedule\}\}\\;\\underbrace\{\\frac\{dI\(\\mathbf\{y\};\\mathbf\{x\}\_\{t\}\)\}\{dt\}\}\_\{\\text\{data\}\}\\,dt\.\(4\.3\)In particular, if the loss weighting equals the diffusion coefficient,w=Dw=D, the floor is simply:Δ​𝒮DSM−Δ​𝒮ideal=I⁡\(𝐲,𝐱t0\)−I⁡\(𝐲,𝐱t1\)\\Delta\\mathcal\{S\}\_\{\\rm DSM\}\-\\Delta\\mathcal\{S\}\_\{\\rm ideal\}=I\(\\mathbf\{y\};\\mathbf\{x\}\_\{t\_\{0\}\}\)\-I\(\\mathbf\{y\};\\mathbf\{x\}\_\{t\_\{1\}\}\)\.

In making this statement, notice that once again, we made no assumptions about Gaussian, affine interpolant or signal\-to\-noise ratio\. The data enters only through the information flow−dI/dt\-dI/dt, and the schedule only through the ratiow/Dw/D\. We have verified Eq\. \([4\.2](https://arxiv.org/html/2608.23916#S4.E2)\) numerically for a nonlinear corruption drift by solving the Fokker–Planck equation directly \(Appendix[D](https://arxiv.org/html/2608.23916#A4)\)\.

### 4\.2Schedule and weighting

The word*schedule*is used in the diffusion literature for several different objects, and since Theorem[4\.3](https://arxiv.org/html/2608.23916#S4.Thmtheorem3)assigns these objects distinct roles, we separate them before evaluating the floor\. The primary one is the*corruption schedule*\(αt,σt\)\(\\alpha\_\{t\},\\sigma\_\{t\}\), which specifies how the data are gradually destroyed by noise; equivalently, one may specify the drift–diffusion pair\(𝐛,γ\)\(\\mathbf\{b\},\\gamma\)of the forward process or the SNR profileλt\\lambda\_\{t\}\. This is the only one of the four that shapes the integrand itself, because it determines the curve that the corruption traces in the information manifold\. A second object, the*loss weighting*w⁡\(t\)w\(t\), multiplies the contribution of each time to the objective\. A third, the*training\-time density*, is the distribution from which noise levels are drawn during optimization\. Neither of these two enters the integrand and together they determine only how the integral is weighted and how it is sampled\. The fourth, the*sampling\-time discretization*, is the step placement of a numerical solver integrating the generative dynamics and plays no role in the training objective at all, and we defer it to Sec\.[5\.3](https://arxiv.org/html/2608.23916#S5.SS3)\. With these distinctions in place, we return to the bridge construction of Sec\.[2](https://arxiv.org/html/2608.23916#S2), which singles out a distinguished combination of schedule and weighting:

###### Lemma 4\.4\(Bridge weighting is the SNR measure\)\.

Writeλt:=αt2/σt2\\lambda\_\{t\}:=\\alpha\_\{t\}^\{2\}/\\sigma\_\{t\}^\{2\}for the signal\-to\-noise ratio\. Among the marginal\-preserving diffusions that realize a given affine interpolant, the bridge construction of Sec\.[2](https://arxiv.org/html/2608.23916#S2)singles out the memoryless one, that is, the unique state\-independent coefficient for which the rescaled process𝐱t/αt\\mathbf\{x\}\_\{t\}/\\alpha\_\{t\}has independent increments, namely

γ⁡\(t\)=2​σt2​dd​t​ln⁡αtσt,\\gamma\(t\)\\;=\\;2\\,\\sigma\_\{t\}^\{2\}\\,\\frac\{d\}\{dt\}\\ln\\frac\{\\alpha\_\{t\}\}\{\\sigma\_\{t\}\},\(4\.4\)and for this coefficient, for any affine interpolant,

γ⁡\(t\)​αt2σt4=d​λtd​t\.\\gamma\(t\)\\,\\frac\{\\alpha\_\{t\}^\{2\}\}\{\\sigma\_\{t\}^\{4\}\}\\;=\\;\\frac\{d\\lambda\_\{t\}\}\{dt\}\.\(4\.5\)

The statement can be verified by a direct computation for each interpolant family and is given in Appendix[D](https://arxiv.org/html/2608.23916#A4)\. The lemma is also of independent interest beyond its use below, since it implies*the memoryless schedule is precisely the one whose loss weighting is the SNR measure*\. In other words, the bridge member of the marginal\-preserving family is not one convenient choice among many but the canonical object singled out by the objective itself\.

With the four notions separated, we can now factor the floor into a data contribution and a schedule contribution\. Suppose the training loss uses a weightw⁡\(t\)w\(t\)in place of the bridge weightγ\\gamma\. Changing variables fromtttoλ\\lambdawith Lemma[4\.4](https://arxiv.org/html/2608.23916#S4.Thmtheorem4)then gives

Δ​𝒮DSM−Δ​𝒮ideal=12​∫ω⁡\(λ\)​MMSE​\(λ\)​𝑑λ=12​∫ω⁡\(λ\)​S​\(λ\)​d​log⁡λ,ω:=wγ,\\Delta\\mathcal\{S\}\_\{\\rm DSM\}\-\\Delta\\mathcal\{S\}\_\{\\rm ideal\}\\;=\\;\\frac\{1\}\{2\}\\int\\omega\(\\lambda\)\\,\\mathrm\{MMSE\}\(\\lambda\)\\,d\\lambda\\;=\\;\\frac\{1\}\{2\}\\int\\omega\(\\lambda\)\\,S\(\\lambda\)\\,d\\log\\lambda,\\qquad\\omega:=\\frac\{w\}\{\\gamma\},\(4\.6\)whereMMSE⁡\(λ\):=𝔼​‖𝐲−𝔼⁡\[𝐲∣𝐱λ\]‖2\\mathrm\{MMSE\}\(\\lambda\):=\\mathbb\{E\}\\,\\\|\\mathbf\{y\}\-\\mathbb\{E\}\[\\mathbf\{y\}\\mid\\mathbf\{x\}\_\{\\lambda\}\]\\\|^\{2\}with𝐱λ=λ​𝐲\+𝐳\\mathbf\{x\}\_\{\\lambda\}=\\sqrt\{\\lambda\}\\,\\mathbf\{y\}\+\\mathbf\{z\},𝐳∼𝒩⁡\(0,𝐈\)\\mathbf\{z\}\\sim\\mathcal\{N\}\(0,\\mathbf\{I\}\), and we have defined the*information spectrum*of the data:

S⁡\(λ\):=λ​MMSE​\(λ\)=2​d​I​\(𝐲,𝐱λ\)d​log⁡λ\.S\(\\lambda\)\\;:=\\;\\lambda\\,\\mathrm\{MMSE\}\(\\lambda\)\\;=\\;2\\,\\frac\{dI\(\\mathbf\{y\};\\mathbf\{x\}\_\{\\lambda\}\)\}\{d\\log\\lambda\}\.\(4\.7\)It is easy to see that the data dependency enters only throughSS, a scalar function of log\-SNR determined by the distribution alone, while the schedule dependence enters only throughω\\omegaand the endpoints, with the bridge weighting corresponding toω≡1\\omega\\equiv 1\.

In this case, the spectrum has a transparent structure, shown in Fig\.[1](https://arxiv.org/html/2608.23916#S4.F1)\. Each structural scale of the data contributes a bump, so that for a mixture with modes separated bymmand within\-mode scaless, mode identity is resolved nearλ∼m−2\\lambda\\sim m^\{\-2\}and the within\-mode directions nearλ∼s−2\\lambda\\sim s^\{\-2\}, while at largeλ\\lambdathe spectrum saturates at the information dimension,S→d⁡\(𝐲\)S\\to d\(\\mathbf\{y\}\), which is Eq\. \([4\.11](https://arxiv.org/html/2608.23916#S4.E11)\) below\.

This structure suggests a criterion for allocating resolution\. If one asks that each unit of schedule carry equal information, a natural choice, since the excessΔ​𝒮ideal\\Delta\\mathcal\{S\}\_\{\\rm ideal\}is what the model must learn whileSSsets the scale of the noise it must learn through, then the allocation satisfies

d​td​log⁡λ∝S⁡\(λ\)\.\\frac\{dt\}\{d\\log\\lambda\}\\;\\propto\\;S\(\\lambda\)\.\(4\.8\)For featureless data the spectrumSSis flat, and the criterion reduces to uniform allocation inlog⁡λ\\log\\lambda\. This is qualitatively close to the commonly used cosine\-like schedules over finite, clipped SNR ranges, though not identical to them, and related derivations of cosine\-like schedules from the schedule\-dependent part of the geometry appear in\[[52](https://arxiv.org/html/2608.23916#bib.bib52),[51](https://arxiv.org/html/2608.23916#bib.bib51)\]\. Real data, however, are not featureless\. The spectrum carries a bump at each structural scale of the data, so a uniform\-in\-log⁡λ\\log\\lambdaallocation under\-samples precisely the noise levels at which the data are most informative\. By the discussion of Sec\.[4\.4](https://arxiv.org/html/2608.23916#S4.SS4), these are the neighbourhoods of the symmetry\-breaking transitions\. The criterion also has a direct information\-theoretic meaning\. Since−dI/dt=dH\(𝐲∣𝐱t\)/dt\-dI/dt=dH\(\\mathbf\{y\}\\mid\\mathbf\{x\}\_\{t\}\)/dt, Eq\. \([4\.8](https://arxiv.org/html/2608.23916#S4.E8)\) coincides with the criterion of entropic time schedulers\[[50](https://arxiv.org/html/2608.23916#bib.bib50)\], which we have thus recovered from the loss decomposition rather than from an entropy argument\. Equation \([4\.3](https://arxiv.org/html/2608.23916#S4.E3)\) shows that this criterion is not merely plausible but forced by the structure of the objective, because the floor factorizes into a data flow and a schedule ratio, the total is pinned by the endpoints, and equal\-information allocation is precisely the choice that makes the schedule factor uniform in the data’s own coordinate\. Finally, training\-time noise allocation has itself been studied\[[55](https://arxiv.org/html/2608.23916#bib.bib55)\]\. A narrower variant that we flag as an open question, and do not pursue here, is to useS⁡\(λ\)S\(\\lambda\)as an*online*importance density for Monte Carlo sampling of training noise levels, estimated from the denoiser’s posterior variance as training proceeds\.

Figure 1:Data/schedule separation of the loss floor\.Left:the information spectrumS⁡\(λ\)=λ​MMSE​\(λ\)S\(\\lambda\)=\\lambda\\,\\mathrm\{MMSE\}\(\\lambda\)for two mixtures\. Dotted lines mark the mode\-separation scaleλ=m−2\\lambda=m^\{\-2\}, dashed lines the within\-mode scaleλ=s−2\\lambda=s^\{\-2\}; the plateau at largeλ\\lambdais the intrinsic dimensionk=1k=1\.Right:the schedule density implied by Eq\. \([4\.8](https://arxiv.org/html/2608.23916#S4.E8)\), compared with a uniform\-in\-log\\logSNR allocation\. Uniform allocation is correct only for featureless data\.
### 4\.3Gaussian corruption and reparametrization invariance

The affine case is the specialization in which the information flow has a closed form\. For affine kernels, Eq\. \([3\.11](https://arxiv.org/html/2608.23916#S3.E11)\) makestr⁡g\\operatorname\{tr\}gproportional toMMSE⁡\(λ\)\\mathrm\{MMSE\}\(\\lambda\)\.

###### Theorem 4\.5\(Loss floor==information gained\)\.

With the bridge schedule of Lemma[4\.4](https://arxiv.org/html/2608.23916#S4.Thmtheorem4),

Δ​𝒮DSM−Δ​𝒮ideal=12​∫MMSE⁡\(λ\)​𝑑λ=I⁡\(𝐲,𝐱λ1\)−I⁡\(𝐲,𝐱λ0\)=h⁡\(𝐱λ1\)−h⁡\(𝐱λ0\),\\Delta\\mathcal\{S\}\_\{\\rm DSM\}\-\\Delta\\mathcal\{S\}\_\{\\rm ideal\}\\;=\\;\\frac\{1\}\{2\}\\int\\mathrm\{MMSE\}\(\\lambda\)\\,d\\lambda\\;=\\;I\(\\mathbf\{y\};\\mathbf\{x\}\_\{\\lambda\_\{1\}\}\)\-I\(\\mathbf\{y\};\\mathbf\{x\}\_\{\\lambda\_\{0\}\}\)\\;=\\;h\(\\mathbf\{x\}\_\{\\lambda\_\{1\}\}\)\-h\(\\mathbf\{x\}\_\{\\lambda\_\{0\}\}\),\(4\.9\)whereIIis mutual information andhhdifferential entropy\.

The proof combines Theorem[3\.4](https://arxiv.org/html/2608.23916#S3.Thmtheorem4)with Lemma[4\.4](https://arxiv.org/html/2608.23916#S4.Thmtheorem4)and the I–MMSE relation of Guo, Shamai and Verdú\[[49](https://arxiv.org/html/2608.23916#bib.bib49)\]\(Appendix[C](https://arxiv.org/html/2608.23916#A3)\)\. Figure[2](https://arxiv.org/html/2608.23916#S4.F2)verifies Eq\. \([4\.9](https://arxiv.org/html/2608.23916#S4.E9)\) numerically on a Gaussian mixture \(Appendix[D](https://arxiv.org/html/2608.23916#A4)\)\.

Figure 2:The loss floor is an entropy\.Left:the trace of the latent Fisher–Rao metric, equivalently the MMSE of endpoint estimation, along the flow\.Right:the accumulated floor12​∫MMSE​𝑑λ\\frac\{1\}\{2\}\\int\\mathrm\{MMSE\}\\,d\\lambda\(solid\) against the differential\-entropy changeΔ​h​\(𝐱λ\)\\Delta h\(\\mathbf\{x\}\_\{\\lambda\}\)\(dashed\), verifying Theorem[4\.5](https://arxiv.org/html/2608.23916#S4.Thmtheorem5)\. Two\-mode Gaussian mixture, entropies by quadrature\.Because the right\-hand side of Eq\. \([4\.9](https://arxiv.org/html/2608.23916#S4.E9)\) is a difference of endpoint values, the loss floor depends on the noise schedule only through the endpoints of the SNR range, not on the path between them\. This is the invariance theorem of variational diffusion models\[[20](https://arxiv.org/html/2608.23916#bib.bib20)\], obtained here as a integral rather than by direct computation\. The geometric interpretation of this is that, a schedule is a*parametrization*of a fixed curve in the information manifold, and the loss floor is a parametrization\-invariant line integral along that curve, pinned by its endpoints\.

Invariance is thus not an algebraic accident of the ELBO but the statement that a geometric quantity does not depend on how one traverses the curve, and this also delimits what schedule design can achieve\. A schedule cannot change the floor, it can only redistribute the estimation error along the path\. In that sense, the recent Fisher\-geometric schedule derivations\[[52](https://arxiv.org/html/2608.23916#bib.bib52),[51](https://arxiv.org/html/2608.23916#bib.bib51)\]and the present account are complementary\.

### 4\.4High\-SNR asymptotics and information dimension

For continuous data, the mutual informationI⁡\(𝐲,𝐱λ\)I\(\\mathbf\{y\};\\mathbf\{x\}\_\{\\lambda\}\)diverges asλ→∞\\lambda\\to\\infty, because perfect observation of a continuous variable carries unbounded information\. Theorem[4\.5](https://arxiv.org/html/2608.23916#S4.Thmtheorem5)then implies that the floor is infinite in the continuum limit and that it is finite in practice only because implementations truncate the SNR range\. The growth of the metric trace in Eq\. \([3\.11](https://arxiv.org/html/2608.23916#S3.E11)\) asσt→0\\sigma\_\{t\}\\to 0is therefore not a defect of the objective\. It is the divergence of a mutual information, and the familiartmint\_\{\\min\}cut\-off is more accurately understood as an information cut\-off\. The rate at which the floor diverges is itself informative\. We first give the Gaussian computation, which is elementary, and then state the general asymptotic, which rests on a deeper result from information theory\.

For*Gaussian*𝐲\\mathbf\{y\}with covariance eigenvaluessi2s\_\{i\}^\{2\}, the posterior is Gaussian as well, with variancesi2/\(1\+λ​si2\)s\_\{i\}^\{2\}/\(1\+\\lambda s\_\{i\}^\{2\}\)in each eigendirection\. Summing these variances gives

MMSE⁡\(λ\)=∑isi21\+λ​si2→λ→∞kλ,Δ​𝒮DSM−Δ​𝒮ideal\|λ≤Λ∼k2​log⁡Λ,\\mathrm\{MMSE\}\(\\lambda\)\\;=\\;\\sum\_\{i\}\\frac\{s\_\{i\}^\{2\}\}\{1\+\\lambda s\_\{i\}^\{2\}\}\\;\\xrightarrow\[\\lambda\\to\\infty\]\{\}\\;\\frac\{k\}\{\\lambda\},\\qquad\\Delta\\mathcal\{S\}\_\{\\rm DSM\}\-\\Delta\\mathcal\{S\}\_\{\\rm ideal\}\\Big\|\_\{\\lambda\\leq\\Lambda\}\\sim\\frac\{k\}\{2\}\\log\\Lambda,\(4\.10\)wherekkis the number of non\-zero eigenvalues\. We stress that this computation uses Gaussianity of the prior in an essential way and does not follow from the support and covariance of the data alone\.

The logarithmic rate itself is nevertheless general\. What changes in the general case is the coefficient, which is an information\-theoretic quantity rather than a linear\-algebraic one\. The high\-SNR behaviour ofI⁡\(𝐲,𝐱λ\)I\(\\mathbf\{y\};\\mathbf\{x\}\_\{\\lambda\}\)in a Gaussian channel is governed by the Rényi information dimension of the input\[[67](https://arxiv.org/html/2608.23916#bib.bib67)\], and correspondinglyλ​MMSE​\(λ\)\\lambda\\,\\mathrm\{MMSE\}\(\\lambda\)tends to the MMSE dimension of𝐲\\mathbf\{y\}\[[68](https://arxiv.org/html/2608.23916#bib.bib68),[69](https://arxiv.org/html/2608.23916#bib.bib69)\]\. Writingd⁡\(𝐲\)d\(\\mathbf\{y\}\)for this dimension, the floor behaves as

Δ​𝒮DSM−Δ​𝒮ideal\|λ≤Λ∼d⁡\(𝐲\)2​log⁡Λ\.\\Delta\\mathcal\{S\}\_\{\\rm DSM\}\-\\Delta\\mathcal\{S\}\_\{\\rm ideal\}\\Big\|\_\{\\lambda\\leq\\Lambda\}\\;\\sim\\;\\frac\{d\(\\mathbf\{y\}\)\}\{2\}\\,\\log\\Lambda\.\(4\.11\)This single formula covers several regimes\. For a distribution that is regular on akk\-dimensional manifold,d⁡\(𝐲\)=kd\(\\mathbf\{y\}\)=kand Eq\. \([4\.11](https://arxiv.org/html/2608.23916#S4.E11)\) reduces to Eq\. \([4\.10](https://arxiv.org/html/2608.23916#S4.E10)\)\. For a mixture of discrete and continuous components,d⁡\(𝐲\)d\(\\mathbf\{y\}\)equals the weight of the continuous part and need not be an integer\. For singular distributions, the information dimension may fail to exist altogether\. The divergence rate of the loss floor is therefore a property of the information dimension of the data, and it can be estimated from the denoiser’s posterior variance alone, without access to a likelihood or to the Jacobian of the network\. Figure[4](https://arxiv.org/html/2608.23916#S4.F4)confirms Eq\. \([4\.10](https://arxiv.org/html/2608.23916#S4.E10)\) numerically, and the fitted slopes matchk/2k/2to three decimals\. We note that estimating intrinsic dimension from diffusion models is an established topic\[[70](https://arxiv.org/html/2608.23916#bib.bib70),[71](https://arxiv.org/html/2608.23916#bib.bib71),[72](https://arxiv.org/html/2608.23916#bib.bib72)\]\. Our point here is not to propose a competitive estimator but to explain why the loss floor knows the dimension at all\.

Figure 3:The floor diverges logarithmically at the data end, with slopek/2k/2for Gaussian data of rankkkinℝ10\\mathbb\{R\}^\{10\}\.Figure 4:The metric trace is monotone \(left\), while its*rate*localizes the symmetry\-breaking transition \(right\), with dotted lines marking the predicted bimodality threshold\.
These asymptotics sit close to the symmetry\-breaking transitions studied in\[[73](https://arxiv.org/html/2608.23916#bib.bib73),[74](https://arxiv.org/html/2608.23916#bib.bib74),[75](https://arxiv.org/html/2608.23916#bib.bib75)\], and it is tempting to expect the floor itself to diverge at those transitions\. It does not\. To see why, writeM\(t\):=𝔼PttrCov\[𝐲∣𝐱,t\]M\(t\):=\\mathbb\{E\}\_\{P\_\{t\}\}\\operatorname\{tr\}\\mathrm\{Cov\}\[\\mathbf\{y\}\\mid\\mathbf\{x\},t\]for the intrinsic part of the floor\. This quantity decreases monotonically from its value at maximal ambiguity to the within\-mode variance, so it remains finite throughout the transition\. What localizes the transition is instead its rate\|d​M/d​t\|\|dM/dt\|, which sharpens as the modes separate, as shown in Fig\.[4](https://arxiv.org/html/2608.23916#S4.F4)\. The geometric signature of speciation is thus the rate of change of the latent metric rather than the metric itself, a statement close in content to the entropy\-rate criterion of\[[76](https://arxiv.org/html/2608.23916#bib.bib76)\]\.

### 4\.5A thermodynamic interpretation

We close this section with a thermodynamic interpretation of the decomposition\. This interpretation is an analogy rather than an identity, and we will be careful about where it holds and where it breaks\.

The starting point is the notion of entropy production in stochastic thermodynamics\. There, entropy production is a path\-space relative entropy between a process and its time reversal,στ=limd​t→0DK​L\(P∥P†\)/dt\\sigma\_\{\\tau\}=\\lim\_\{dt\\to 0\}D\_\{KL\}\(P\\\|P^\{\\dagger\}\)/dt, whereP†P^\{\\dagger\}denotes the time\-reversed path measure\[[77](https://arxiv.org/html/2608.23916#bib.bib77),[78](https://arxiv.org/html/2608.23916#bib.bib78)\]\. A related construction equips the space of path measures with a Fisher metric by interpolating between two velocity fields, and the induced line element isd​t2​μ​T​∫‖ν′−ν‖2​P\\frac\{dt\}\{2\\mu T\}\\int\\\|\\nu^\{\\prime\}\-\\nu\\\|^\{2\}P\[[78](https://arxiv.org/html/2608.23916#bib.bib78)\]\.

Our excess action Eq\. \([2\.13](https://arxiv.org/html/2608.23916#S2.E13)\) has exactly this form, with the diffusion coefficientγ\\gammaplaying the role of2​μ​T2\\mu T\. It can therefore be read as the dissipation incurred by running the model’s velocity field in place of the optimal one\. Theorem[4\.5](https://arxiv.org/html/2608.23916#S4.Thmtheorem5)complements this picture, since it identifies the floor withΔ​h\\Delta h, an entropy change of the state along the reference corruption channel\.

However, we need to be careful before combining these two observations\. The quantityΔ​h\\Delta his a change of differential entropy of the state\. Entropy production in stochastic thermodynamics is instead a path\-space quantity with separate system and environment contributions\. The two coincide only under the Gaussian\-channel reading in the SNR parametrization\. With this caveat in place, the decomposition resembles a dissipation\-plus\-entropy splitting,

Δ​𝒮DSM⏟denoising loss=Δ​𝒮ideal⏟dissipation\+Δ​h⏟entropy production\.\\underbrace\{\\Delta\\mathcal\{S\}\_\{\\rm DSM\}\}\_\{\\text\{denoising loss\}\}\\;=\\;\\underbrace\{\\Delta\\mathcal\{S\}\_\{\\rm ideal\}\}\_\{\\text\{dissipation\}\}\\;\+\\;\\underbrace\{\\Delta h\}\_\{\\text\{entropy production\}\}\.\(4\.12\)
That the score\-matching objective is itself an entropy\-production functional has been established independently and in more detail by\[[79](https://arxiv.org/html/2608.23916#bib.bib79)\], whose time\-asymmetry entropy production is proportional to the score\-matching loss\. Our contribution is complementary, since we identify the*irreducible*term with a Fisher–Rao metric and withΔ​h\\Delta h\. We also caution that Eq\. \([4\.12](https://arxiv.org/html/2608.23916#S4.E12)\) is*not*the excess/housekeeping decomposition of\[[78](https://arxiv.org/html/2608.23916#bib.bib78)\]\. That decomposition splits entropy production by a geometric criterion, separating the Wasserstein speed of the density from the remainder, whereas ours splits the loss by dependence on the model\. The two coincide in neither definition nor value\.

## 5Consequences for Training and Sampling

The decomposition of Sec\.[4](https://arxiv.org/html/2608.23916#S4)affects the two halves of the pipeline differently, because training and sampling probe different parts of the geometry\. The training objective involves only the second cumulant of the denoising posterior, while the discretization error of the sampler also involves the third\.

### 5\.1Why raw training losses are not comparable

For the bridge weighting, the decomposition readsΔ​𝒮DSM=Δ​𝒮ideal\+Δ​h\\Delta\\mathcal\{S\}\_\{\\rm DSM\}=\\Delta\\mathcal\{S\}\_\{\\rm ideal\}\+\\Delta h, and the floor depends on the schedule through its endpoints and on the data through its entropy\. Raw training losses are therefore not directly comparable across runs that use different SNR ranges, schedules, or loss weightings\. Under fixed training conditions the floor is common to all runs and the conventional loss remains a valid comparator, but cross\-configuration comparisons can be genuinely misleading\. Figure[5](https://arxiv.org/html/2608.23916#S5.F5)\(left, centre\) gives an example on the analytic mixture of Appendix[D](https://arxiv.org/html/2608.23916#A4)\. With two SNR ranges and two models of fixed known quality, the better model evaluated on the wider range reports a larger raw loss \(1\.54551\.5455\) than the worse model on the narrower range \(1\.12671\.1267\), purely because the wider range integrates more of the floor\. Subtracting the floors \(1\.04141\.0414and1\.51431\.5143\) gives excesses that order the models correctly on both ranges\. We therefore recommend reporting the floor\-subtracted excess alongside the conventional loss whenever models are compared across schedules or SNR ranges\. The correction is cheap, since Theorem[3\.4](https://arxiv.org/html/2608.23916#S3.Thmtheorem4)expresses the floor as an expectation of the posterior covariance, which the denoising loop already has the ingredients to estimate\.

Figure 5:Consequences of the decomposition\.Left:raw training loss for two models of known quality on two SNR ranges; the ranking inverts\.Centre:after subtracting the floor, the excess orders the models correctly on both ranges\.Right:step placement for a first\-order solver on the analytic mixture, mean endpoint error against the exactly integrated probability\-flow ODE\.
### 5\.2What schedule design can and cannot change

The decomposition also delimits schedule design\. Because the total floor is pinned by the endpoints of the SNR range, a schedule can only redistribute where the estimation error is incurred, and the equal\-information allocation discussed after Eq\. \([4\.8](https://arxiv.org/html/2608.23916#S4.E8)\) recovers the criterion already used by entropic time schedulers\[[50](https://arxiv.org/html/2608.23916#bib.bib50)\]\. The one variant that appears unclaimed is the training\-time analogue, namely samplingttfrom a density proportional toS⁡\(λ\)=tr⁡Cov⁡\[𝐲∣𝐱t\]S\(\\lambda\)=\\operatorname\{tr\}\\mathrm\{Cov\}\[\\mathbf\{y\}\\mid\\mathbf\{x\}\_\{t\}\]estimated online from the denoiser, which we record as an open question rather than a proposal\.

### 5\.3Second\- versus third\-order conditional statistics

This subsection is exploratory and reports diagnostic observations on an analytic example rather than sampler design principles\. Higher cumulants of the denoising posterior are absent from the training objective, because denoising score matching draws each noise level independently and never discretizes a trajectory\. They appear as soon as one integrates the sampler\. Writing the probability\-flow trajectory in the mean\-parameter coordinateη=∇Λ=𝔼⁡\[𝐲∣𝐱t\]\\eta=\\nabla\\Lambda=\\mathbb\{E\}\[\\mathbf\{y\}\\mid\\mathbf\{x\}\_\{t\}\]with tangentu=d​θ/d​λu=d\\theta/d\\lambda,

d​ηd​λ=g​u,d2​ηd​λ2=T⁡\[u,u\]\+g​d​ud​λ,T:=∇3Λ,\\frac\{d\\eta\}\{d\\lambda\}\\;=\\;g\\,u,\\qquad\\frac\{d^\{2\}\\eta\}\{d\\lambda^\{2\}\}\\;=\\;T\[u,u\]\\;\+\\;g\\,\\frac\{du\}\{d\\lambda\},\\qquad T:=\\nabla^\{3\}\\Lambda,\(5\.1\)whereTTis the third cumulant of the conditional endpoint law\. The leading truncation term of a first\-order step is therefored2​η/d​λ2d^\{2\}\\eta/d\\lambda^\{2\}, which depends on the third cumulant throughT⁡\[u,u\]T\[u,u\]in addition to the second\-order termg​d​u/d​λg\\,du/d\\lambda\. Measured along the exactly integrated trajectory on the analytic mixture \(Appendix[D](https://arxiv.org/html/2608.23916#A4)\), the contribution ofTTis strongly localized in time and peaks where the modes separate, with a median share of0\.470\.47of‖d2​η/d​λ2‖\\\|d^\{2\}\\eta/d\\lambda^\{2\}\\\|across trajectories\. Accordingly, Figure[5](https://arxiv.org/html/2608.23916#S5.F5)\(right\) shows that equalizing the true leading error‖d2​η/d​λ2‖​Δ​λ2\\\|d^\{2\}\\eta/d\\lambda^\{2\}\\\|\\Delta\\lambda^\{2\}outperforms equalizing Fisher–Rao arclength at every step count tested\. We report this as a diagnostic rather than a sampler proposal, because the comparison uses an explicit Euler solver while standard baselines such as DDIM\[[3](https://arxiv.org/html/2608.23916#bib.bib3)\]integrate the linear part exactly, and the decisive experiment with an exponential integrator at matched order remains to be carried out\.

## 6Discussion

The identity behind every result in this paper is elementary\. Denoising score matching regresses onto a random target whose posterior mean is the marginal score, and the fluctuation of that target is the score of the conditional endpoint family\. Its second moment is a Fisher information by definition, which is why the training loss contains a term no model can reduce and why that term is the metric of the latent manifold\. Once the corruption is a diffusion, the same quantity measures the rate at which the noisy state forgets the data, so the floor accumulates a mutual information\. The geometry, the information theory, and the objective are three descriptions of one number\.

### 6\.1Scope and relation to prior work

The primary results are the conditional\-variance identity of Theorem[3\.4](https://arxiv.org/html/2608.23916#S3.Thmtheorem4)and its factorization into an information flow and a schedule weight in Theorem[4\.3](https://arxiv.org/html/2608.23916#S4.Thmtheorem3)\. The ingredients of Sec\.[3](https://arxiv.org/html/2608.23916#S3)are classical, namely the second variation of a log\-partition functional, the theory of sufficient reductions, and the induced geometry of curved exponential families\[[61](https://arxiv.org/html/2608.23916#bib.bib61),[65](https://arxiv.org/html/2608.23916#bib.bib65),[66](https://arxiv.org/html/2608.23916#bib.bib66)\], and we claim novelty only for their assembly into a derivation of the latent metric from the bridge variational principle\. Several nearby lines of work address optimization problems that are easily conflated with ours, including sampler discretization\[[53](https://arxiv.org/html/2608.23916#bib.bib53),[54](https://arxiv.org/html/2608.23916#bib.bib54),[50](https://arxiv.org/html/2608.23916#bib.bib50)\], the training\-time noise density\[[55](https://arxiv.org/html/2608.23916#bib.bib55)\], loss weighting\[[20](https://arxiv.org/html/2608.23916#bib.bib20)\], and time reparametrization\[[52](https://arxiv.org/html/2608.23916#bib.bib52),[51](https://arxiv.org/html/2608.23916#bib.bib51)\]\. Our contribution to this thread is not a new member of any of these families but the factorization Eq\. \([4\.6](https://arxiv.org/html/2608.23916#S4.E6)\), which delimits what any of them can affect\. Two further boundaries are worth marking\. The reading of the score\-matching objective as an entropy\-production functional is developed independently and in more detail by\[[79](https://arxiv.org/html/2608.23916#bib.bib79)\], while our contribution is the identification of the irreducible term with a Fisher–Rao metric and with an entropy change\. Estimating intrinsic dimension from a trained diffusion model is established\[[70](https://arxiv.org/html/2608.23916#bib.bib70),[71](https://arxiv.org/html/2608.23916#bib.bib71),[72](https://arxiv.org/html/2608.23916#bib.bib72)\], and Eq\. \([4\.11](https://arxiv.org/html/2608.23916#S4.E11)\) explains why the floor carries that information rather than proposing a new estimator\.

### 6\.2The discrete case of masked diffusion

The decomposition also has a definite consequence beyond the continuous setting, which we believe is among the most useful observations of this work\. The masked\-diffusion objective is a cross\-entropy, and a cross\-entropy is the Bregman divergence generated by the negative entropy\. The same algebra that produced the conditional\-variance split Eq\. \([3\.20](https://arxiv.org/html/2608.23916#S3.E20)\) applies to any such loss\. The expected divergence to a random target separates into a model\-dependent term evaluated at the posterior mean of the target and an irreducible remainder known as the Bregman information of the target\[[80](https://arxiv.org/html/2608.23916#bib.bib80),[81](https://arxiv.org/html/2608.23916#bib.bib81),[82](https://arxiv.org/html/2608.23916#bib.bib82)\]\. For the cross\-entropy this remainder is a conditional entropy\. Writingmtm\_\{t\}for the masking probability, decreasing from11to00, andUUfor the set of unmasked tokens, a still\-masked token is revealed over an infinitesimal step with probability−m˙t/mt\-\\dot\{m\}\_\{t\}/m\_\{t\}at a costH⁡\(yi∣yU\)H\(y\_\{i\}\\mid y\_\{U\}\)\. With a perfect model the residual loss is therefore

floor=∫01d​mm​𝔼U​\[∑i∉UH⁡\(yi∣yU\)\],\\mathrm\{floor\}\\;=\\;\\int\_\{0\}^\{1\}\\frac\{dm\}\{m\}\\;\\mathbb\{E\}\_\{U\}\\Big\[\\sum\_\{i\\notin U\}H\\big\(y\_\{i\}\\mid y\_\{U\}\\big\)\\Big\],\(6\.1\)where each token is unmasked independently with probability1−m1\-m\. Two properties follow immediately\. Substitutingu=mtu=m\_\{t\}in the time integral removes the schedule, so the floor depends only on the endpoint masking rates and not on the shape ofmtm\_\{t\}\. Evaluating Eq\. \([6\.1](https://arxiv.org/html/2608.23916#S6.E1)\) then gives the entropyH⁡\(y1,…,yn\)H\(y\_\{1\},\\ldots,y\_\{n\}\)of the data\. We have verified both numerically for correlated tokens \(Appendix[D](https://arxiv.org/html/2608.23916#A4)\)\.

This is the discrete counterpart of the continuous statement\. There the floor is an accumulated mutual information fixed by the endpoints of the SNR range, and here it is the data entropy fixed by the endpoint masking rates\. In both cases the schedule redistributes where the cost is incurred without changing the total\. This also accounts for the known schedule invariance of the masked\-diffusion evidence bound\[[83](https://arxiv.org/html/2608.23916#bib.bib83),[84](https://arxiv.org/html/2608.23916#bib.bib84)\], because the objective is a Bregman divergence whose irreducible part telescopes along the filtration\. We leave open the harder question of a discrete analogue of the cone Eq\. \([3\.8](https://arxiv.org/html/2608.23916#S3.E8)\) and of the latent metric itself, which would require theα\\alpha\-geometry of the simplex rather than the argument given here\.

### 6\.3Limitations

Several limitations should be kept in view\. The closed\-form evaluation of the floor asΔ​I\\Delta Iassumes affine Gaussian corruption, and Theorem[4\.3](https://arxiv.org/html/2608.23916#S4.Thmtheorem3)covers general corruption diffusions only at the price of leaving the information flow implicit\. The high\-SNR rate is governed by the information dimension, which need not exist for singular data distributions\. The replacement of the on\-policy weighting by the simulation\-free one in Sec\.[2\.3](https://arxiv.org/html/2608.23916#S2.SS3)leaves the functional minimizer unchanged, but the projection onto a restricted parametric model does depend on the weighting measure, so our statements concern the value of the objective rather than the parametric projection\. Our experiments are analytic or low\-dimensional by design and establish identities rather than performance\. Finally, the floor is only as estimable as the posterior covariance, which is hardest to estimate at the low\-noise end where the floor diverges\.

### 6\.4Concluding remarks

We have shown that the denoising score\-matching objective decomposes exactly into a model\-dependent estimation error and a model\-independent floor, and that the floor is the trace of the Fisher–Rao metric of the conditional endpoint family integrated along the flow\. For corruption diffusions the floor factorizes into a data\-dependent information flow and a schedule\-dependent weight, and for affine Gaussian corruption it evaluates to the mutual information accumulated between data and noisy state\. The practical consequence we would emphasize is the least glamorous one\. A reported diffusion loss mixes model fit with a data\- and schedule\-dependent constant, and the two should be separated before the number is used to compare training runs\.

## Ethics Statement

In the preparation of this manuscript we used large language models for language improvement, editing, and literature search and summarization\. These tools were not used to produce research ideas, derivations, or results, and every suggested change was reviewed and verified by the authors before inclusion\. We take full responsibility for the content of this paper, and any mistakes that remain in the text are entirely ours\.

## Acknowledgements

This work was carried out as an independent research project and received no specific grant from any funding agency in the public, commercial, or not\-for\-profit sectors\. It was conducted outside the authors’ official duties, and the views expressed are those of the authors alone, not of their employers\.

## References

- \[1\]Jascha Sohl\-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli\.Deep unsupervised learning using nonequilibrium thermodynamics\.In Francis Bach and David Blei, editors,Proceedings of the 32nd International Conference on Machine Learning, volume 37 ofProceedings of Machine Learning Research, pages 2256–2265, Lille, France, 07–09 Jul 2015\. PMLR\.
- \[2\]Jonathan Ho, Ajay Jain, and Pieter Abbeel\.Denoising diffusion probabilistic models\.CoRR, abs/2006\.11239, 2020\.
- \[3\]Jiaming Song, Chenlin Meng, and Stefano Ermon\.Denoising diffusion implicit models\.ArXiv, abs/2010\.02502, 2020\.
- \[4\]Yang Song, Jascha Sohl\-Dickstein, Diederik P\. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole\.Score\-based generative modeling through stochastic differential equations\.In9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3\-7, 2021\. OpenReview\.net, 2021\.
- \[5\]Alexander Quinn Nichol and Prafulla Dhariwal\.Improved denoising diffusion probabilistic models\.In Marina Meila and Tong Zhang, editors,Proceedings of the 38th International Conference on Machine Learning, volume 139 ofProceedings of Machine Learning Research, pages 8162–8171\. PMLR, 18–24 Jul 2021\.
- \[6\]Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer\.High\-resolution image synthesis with latent diffusion models\.InCVPR, pages 10684–10695, 2022\.
- \[7\]Prafulla Dhariwal and Alexander Nichol\.Diffusion models beat GANs on image synthesis\.InAdvances in Neural Information Processing Systems, volume 34, pages 8780–8794, 2021\.
- \[8\]Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine\.Elucidating the design space of diffusion\-based generative models\.ArXiv, abs/2206\.00364, 2022\.
- \[9\]Jonathan Ho, Tim Salimans, Alexey A\. Gritsenko, William Chan, Mohammad Norouzi, and David J\. Fleet\.Video diffusion models\.ArXiv, abs/2204\.03458, 2022\.
- \[10\]A\. Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis\.Align your latents: High\-resolution video synthesis with latent diffusion models\.2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\), pages 22563–22575, 2023\.
- \[11\]Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro\.Diffwave: A versatile diffusion model for audio synthesis\.ArXiv, abs/2009\.09761, 2020\.
- \[12\]Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, and Mikhail Kudinov\.Grad\-tts: A diffusion probabilistic model for text\-to\-speech\.InInternational Conference on Machine Learning, 2021\.
- \[13\]Joseph L\. Watson, David Juergens, Nathaniel R\. Bennett, Brian L\. Trippe, Jason Yim, et al\.De novo design of protein structure and function with RFdiffusion\.Nature, 620:1089–1100, 2023\.
- \[14\]Claudio Zeni, Robert Pinsler, Daniel Zügner, Andrew Fowler, Matthew Horton, et al\.A generative model for inorganic materials design\.Nature, 639:624–632, 2025\.
- \[15\]Jacob Austin, Daniel D\. Johnson, Jonathan Ho, Daniel Tarlow, and Rianne van den Berg\.Structured denoising diffusion models in discrete state\-spaces\.InAdvances in Neural Information Processing Systems, volume 34, pages 17981–17993, 2021\.
- \[16\]Aaron Lou, Chenlin Meng, and Stefano Ermon\.Discrete diffusion modeling by estimating the ratios of the data distribution\.InICML, 2024\.
- \[17\]Shen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji\-Rong Wen, and Chongxuan Li\.Large language diffusion models, 2025\.
- \[18\]Brian D\.O\. Anderson\.Reverse\-time diffusion equation models\.Stochastic Processes and their Applications, 12\(3\):313–326, 1982\.
- \[19\]U\. G\. Haussmann and E\. Pardoux\.Time Reversal of Diffusions\.The Annals of Probability, 14\(4\):1188 – 1205, 1986\.
- \[20\]Diederik Kingma, Tim Salimans, Ben Poole, and Jonathan Ho\.Variational diffusion models\.Advances in neural information processing systems, 34:21696–21707, 2021\.
- \[21\]Chin\-Wei Huang, Jae Hyun Lim, and Aaron C Courville\.A variational perspective on diffusion\-based generative models and score matching\.In M\. Ranzato, A\. Beygelzimer, Y\. Dauphin, P\.S\. Liang, and J\. Wortman Vaughan, editors,Advances in Neural Information Processing Systems, volume 34, pages 22863–22876\. Curran Associates, Inc\., 2021\.
- \[22\]Yang Song, Conor Durkan, Iain Murray, and Stefano Ermon\.Maximum likelihood training of score\-based diffusion models\.In M\. Ranzato, A\. Beygelzimer, Y\. Dauphin, P\.S\. Liang, and J\. Wortman Vaughan, editors,Advances in Neural Information Processing Systems, volume 34, pages 1415–1428\. Curran Associates, Inc\., 2021\.
- \[23\]Yaron Lipman, Ricky T\. Q\. Chen, Heli Ben\-Hamu, Maximilian Nickel, and Matt Le\.Flow matching for generative modeling\.ArXiv, abs/2210\.02747, 2022\.
- \[24\]Xingchao Liu, Chengyue Gong, and Qiang Liu\.Flow straight and fast: Learning to generate and transfer data with rectified flow\.ArXiv, abs/2209\.03003, 2022\.
- \[25\]Michael S\. Albergo, Nicholas Matthew Boffi, and Eric Vanden\-Eijnden\.Stochastic interpolants: A unifying framework for flows and diffusions\.J\. Mach\. Learn\. Res\., 26:209:1–209:80, 2023\.
- \[26\]Herbert E\. Robbins\.An empirical bayes approach to statistics\.In Samuel Kotz and Norman L\. Johnson, editors,Breakthroughs in Statistics: Foundations and Basic Theory, pages 388–394\. Springer New York, New York, NY, 1992\.
- \[27\]Bradley Efron\.Tweedie’s formula and selection bias\.Journal of the American Statistical Association, 106:1602 – 1614, 2011\.
- \[28\]Hila Manor and Tomer Michaeli\.On the posterior distribution in denoising: Application to uncertainty quantification\.ArXiv, abs/2309\.13598, 2023\.
- \[29\]Aapo Hyvärinen\.Estimation of non\-normalized statistical models by score matching\.J\. Mach\. Learn\. Res\., 6:695–709, dec 2005\.
- \[30\]Pascal Vincent\.A connection between score matching and denoising autoencoders\.Neural Computation, 23\(7\):1661–1674, 2011\.
- \[31\]Yang Song and Stefano Ermon\.Generative modeling by estimating gradients of the data distribution\.In H\. Wallach, H\. Larochelle, A\. Beygelzimer, F\. d'Alché\-Buc, E\. Fox, and R\. Garnett, editors,Advances in Neural Information Processing Systems, volume 32\. Curran Associates, Inc\., 2019\.
- \[32\]Ole Barndorff\-Nielsen\.Information and exponential families in statistical theory\.John Wiley & Sons Ltd, 1978\.
- \[33\]S\. Amari and H\. Nagaoka\.Methods of Information Geometry\.Fields Institute Communications\. American Mathematical Society, 2000\.
- \[34\]Frank Nielsen\.An elementary introduction to information geometry\.Entropy, 22\(10\):1100, 2020\.
- \[35\]E\. Schrödinger\.Über die Umkehrung der Naturgesetze\.Sitzungsberichte der Preussischen Akademie der Wissenschaften\. Physikalisch\-mathematische Klasse\. Verlag der Akademie der Wissenschaften in Kommission bei Walter De Gruyter u\. Company, 1931\.
- \[36\]E\. Schrödinger\.Sur la théorie relativiste de l’électron et l’interprétation de la mécanique quantique\.Annales de l’institut Henri Poincaré, 2\(4\):269–310, 1932\.
- \[37\]Hans Föllmer\.Random fields and diffusion processes\.InÉcole d’Été de Probabilités de Saint\-Flour XV–XVII, 1985–87, volume 1362 ofLecture Notes in Mathematics, pages 101–203\. Springer, 1988\.
- \[38\]I\. Csiszár\.I\-divergence geometry of probability distributions and minimization problems\.The Annals of Probability, 13\(1\):146–158, 1975\.
- \[39\]Christian Léonard\.A survey of the schrödinger problem and some of its connections with optimal transport\.Discrete and Continuous Dynamical Systems, 34\(4\):1533–1574, 2014\.
- \[40\]Joseph L\. Doob\.Conditional brownian motion and the boundary limits of harmonic functions\.Bulletin de la Société Mathématique de France, 85:431–458, 1957\.
- \[41\]Valentin De Bortoli, James Thornton, Jeremy Heng, and Arnaud Doucet\.Diffusion schrödinger bridge with applications to score\-based generative modeling\.In M\. Ranzato, A\. Beygelzimer, Y\. Dauphin, P\.S\. Liang, and J\. Wortman Vaughan, editors,Advances in Neural Information Processing Systems, volume 34, pages 17695–17709\. Curran Associates, Inc\., 2021\.
- \[42\]Gefei Wang, Yuling Jiao, Qian Xu, Yang Wang, and Can Yang\.Deep generative learning via schrödinger bridge\.In Marina Meila and Tong Zhang, editors,Proceedings of the 38th International Conference on Machine Learning, volume 139 ofProceedings of Machine Learning Research, pages 10794–10804\. PMLR, 18–24 Jul 2021\.
- \[43\]Francisco Vargas, Pierre Thodoroff, Austen Lamacraft, and Neil Lawrence\.Solving schrödinger bridges via maximum likelihood\.Entropy, 23\(9\), 2021\.
- \[44\]Tianrong Chen, Guan\-Horng Liu, and Evangelos A\. Theodorou\.Likelihood training of Schrödinger bridge using forward\-backward SDEs theory\.InInternational Conference on Learning Representations, 2022\.
- \[45\]Alexander Tong, Esmeralda S\. Whitammer, Kilian Fatras, Lazar Atanackovic, Yanlei Zhang, Guillaume Huguet, Guy Wolf, and Yoshua Bengio\.Simulation\-free schrödinger bridges via score and flow matching\.InInternational Conference on Artificial Intelligence and Statistics, 2023\.
- \[46\]Yuyang Shi, Valentin De Bortoli, Andrew Campbell, and Arnaud Doucet\.Diffusion Schrödinger bridge matching\.InAdvances in Neural Information Processing Systems, volume 36, 2023\.
- \[47\]Rafał Karczewski, Markus Heinonen, Alison Pouplin, Søren Hauberg, and Vikas Garg\.The spacetime of diffusion models: An information geometry perspective, 2026\.
- \[48\]Alexander Lobashev, Dmitry Guskov, Maria Larchenko, and Mikhail Tamm\.Hessian geometry of latent space in generative models, 2025\.
- \[49\]Dongning Guo, Shlomo Shamai, and Sergio Verdú\.Mutual information and minimum mean\-square error in Gaussian channels\.IEEE Transactions on Information Theory, 51\(4\):1261–1282, 2005\.
- \[50\]Ognjen Stancevic, Julius Handke, and Luca Ambrogioni\.Entropic time schedulers for generative diffusion models\.InAdvances in Neural Information Processing Systems, 2025\.
- \[51\]Seo Taek Kong, Weina Wang, and R\. Srikant\.Noise schedule design for diffusion models: An optimal control perspective, 2026\.
- \[52\]The cosine schedule is Fisher\-Rao\-optimal for masked discrete diffusion models, 2025\.
- \[53\]Amirmojtaba Sabour, Sanja Fidler, and Karsten Kreis\.Align your steps: Optimizing sampling schedules in diffusion models\.InProceedings of the 41st International Conference on Machine Learning, 2024\.
- \[54\]Christopher Williams, Andrew Campbell, Arnaud Doucet, and Saifuddin Syed\.Score\-optimal diffusion schedules\.InAdvances in Neural Information Processing Systems, 2024\.
- \[55\]Luca Ambrogioni et al\.From atoms to entropy: Optimal noise allocation for diffusion training in the convex regime, 2026\.
- \[56\]Akhil Premkumar\.Generative diffusion from an action principle\.ArXiv, abs/2310\.04490, 2023\.
- \[57\]Amir Dembo and Ofer Zeitouni\.Large Deviations Techniques and Applications\.Springer, 2 edition, 1998\.
- \[58\]Imre Csiszár\.Information geometry and alternating minimization procedures\.Statistics and Decisions, Dedewicz, 1:205–237, 1984\.
- \[59\]Giovanni Pistone and Carlo Sempi\.An infinite\-dimensional geometric structure on the space of all the probability measures equivalent to a given one\.The Annals of Statistics, 23\(5\):1543–1561, 1995\.
- \[60\]Alberto Cena and Giovanni Pistone\.Exponential statistical manifold\.Annals of the Institute of Statistical Mathematics, 59\(1\):27–56, 2007\.
- \[61\]N\. N\. Chentsov\.Statistical Decision Rules and Optimal Inference, volume 53 ofTranslations of Mathematical Monographs\.American Mathematical Society, 1982\.
- \[62\]Nihat Ay, Jürgen Jost, Hông Vân Lê, and Lorenz Schwachhöfer\.Information Geometry, volume 64 ofErgebnisse der Mathematik und ihrer Grenzgebiete\.Springer, 2017\.
- \[63\]Olav Kallenberg\.Processes, Distributions, and Independence, pages 45–61\.Springer New York, New York, NY, 2002\.
- \[64\]Thomas M\. Cover and Joy A\. Thomas\.Elements of Information Theory\.Wiley, 2 edition, 2006\.
- \[65\]Bradley Efron\.Defining the curvature of a statistical problem \(with applications to second order efficiency\)\.The Annals of Statistics, 3\(6\):1189–1242, 1975\.
- \[66\]Shun\-ichi Amari\.Differential geometry of curved exponential families—curvatures and information loss\.The Annals of Statistics, 10\(2\):357–385, 1982\.
- \[67\]Alfréd Rényi\.On the dimension and entropy of probability distributions\.Acta Mathematica Academiae Scientiarum Hungarica, 10\(1–2\):193–215, 1959\.
- \[68\]Yihong Wu and Sergio Verdú\.Rényi information dimension: Fundamental limits of almost lossless analog compression\.IEEE Transactions on Information Theory, 56\(8\):3721–3748, 2010\.
- \[69\]Yihong Wu and Sergio Verdú\.MMSE dimension\.IEEE Transactions on Information Theory, 57\(8\):4857–4879, 2011\.
- \[70\]Phil Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum, and Tom Goldstein\.The intrinsic dimension of images and its impact on learning\.InInternational Conference on Learning Representations, 2021\.
- \[71\]Jan Pawel Stanczuk, Georgios Batzolis, Teo Tsolaki, and Carola\-Bibiane Schönlieb\.Diffusion models encode the intrinsic dimension of data manifolds\.InProceedings of the 41st International Conference on Machine Learning, 2024\.
- \[72\]Hamidreza Kamkari, Brendan Leigh Ross, Rasa Hosseinzadeh, Jesse C\. Cresswell, and Gabriel Loaiza\-Ganem\.A geometric view of data complexity: Efficient local intrinsic dimension estimation with diffusion models\.InAdvances in Neural Information Processing Systems, 2024\.
- \[73\]Gabriel Barroso Raya and Luca Ambrogioni\.Spontaneous symmetry breaking in generative diffusion models\.Journal of Statistical Mechanics: Theory and Experiment, 2024, 2023\.
- \[74\]Luca Ambrogioni\.The statistical thermodynamics of generative diffusion models: Phase transitions, symmetry breaking, and critical instability\.Entropy, 27, 2023\.
- \[75\]Giulio Biroli, Tony Bonnaire, Valentin de Bortoli, and Marc Mézard\.Dynamical regimes of diffusion models\.Nature Communications, 15, 2024\.
- \[76\]Anonymous\.The entropic signature of class speciation in diffusion models\.InProceedings of the 43rd International Conference on Machine Learning, 2026\.
- \[77\]Udo Seifert\.Stochastic thermodynamics, fluctuation theorems and molecular machines\.Reports on Progress in Physics, 75\(12\):126001, 2012\.
- \[78\]Sosuke Ito\.Geometric thermodynamics for the Fokker–Planck equation: Stochastic thermodynamic links between information geometry and optimal transport\.Information Geometry, 7:S441–S483, 2024\.
- \[79\]Xuehao Ding, H\. T\. Quan, and Yuhai Tu\.Stochastic thermodynamics of score matching in diffusion models, 2026\.
- \[80\]Leonard J\. Savage\.Elicitation of personal probabilities and expectations\.Journal of the American Statistical Association, 66\(336\):783–801, 1971\.
- \[81\]Arindam Banerjee, Srujana Merugu, Inderjit S\. Dhillon, and Joydeep Ghosh\.Clustering with bregman divergences\.Journal of Machine Learning Research, 6:1705–1749, 2005\.
- \[82\]Arindam Banerjee, Xin Guo, and Hui Wang\.On the optimality of conditional expectation as a Bregman predictor\.IEEE Transactions on Information Theory, 51\(7\):2664–2669, 2005\.
- \[83\]Subham Sekhar Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan, Edgar Marroquin, Justin T\. Chiu, Alexander Rush, and Volodymyr Kuleshov\.Simple and effective masked diffusion language models\.InAdvances in Neural Information Processing Systems, 2024\.
- \[84\]Jiaxin Shi, Kehang Han, Zhe Wang, Arnaud Doucet, and Michalis K\. Titsias\.Simplified and generalized masked diffusion for discrete data\.InAdvances in Neural Information Processing Systems, 2024\.
- \[85\]Hannes Risken\.The Fokker–Planck Equation: Methods of Solution and Applications\.Springer, 2 edition, 1996\.
- \[86\]Lawrence D\. Brown\.Fundamentals of statistical exponential families : with applications in statistical decision theory\.Lecture notes\-monograph series Fundamentals of statistical exponential families\. Institute of Mathematical Statistics, 1986\.

## Appendix

## Appendix AVariational Derivation of the Endpoint\-Tilted Path Measure

This appendix collects the derivations underlying the variational formulation of the Schrödinger bridge in Sec\.[2](https://arxiv.org/html/2608.23916#S2)\. We first solve the constrained minimization that yields the factorized optimal path measure, and then derive the recursion relations for the auxiliary fields together with their continuum limit\. The remaining subsections expand the tilted kernel at short times, construct the conditional fields that connect the bridge to the score\-matching objective, and compute the short\-time KL divergence between diffusion kernels that underlies the excess action\.

### A\.1Solution of the bridge problem

Consider the time interval\[0,T\]\[0,T\]discretized intoNNequal segments of lengthΔ​t=T/N\\Delta t=T/N, and let𝐱k\\mathbf\{x\}\_\{k\}denote the state at timetk=k​Δ​tt\_\{k\}=k\\Delta t\. The joint path measure under the reference process is

𝒢\(𝐱0:N\)=P0\(𝐱0\)∏k=0N−1G\(𝐱k\+1∣𝐱k\),\\mathcal\{G\}\(\\mathbf\{x\}\_\{0:N\}\)=P\_\{0\}\(\\mathbf\{x\}\_\{0\}\)\\prod\_\{k=0\}^\{N\-1\}G\(\\mathbf\{x\}\_\{k\+1\}\\mid\\mathbf\{x\}\_\{k\}\),\(A\.1\)and under the candidate process,

ℋ\(𝐱0:N\)=P0\(𝐱0\)H\(𝐱1:N∣𝐱0\)=P0\(𝐱0\)∏k=0N−1H\(𝐱k\+1∣𝐱k\)\.\\mathcal\{H\}\(\\mathbf\{x\}\_\{0:N\}\)=P\_\{0\}\(\\mathbf\{x\}\_\{0\}\)H\(\\mathbf\{x\}\_\{1:N\}\\mid\\mathbf\{x\}\_\{0\}\)=P\_\{0\}\(\\mathbf\{x\}\_\{0\}\)\\prod\_\{k=0\}^\{N\-1\}H\(\\mathbf\{x\}\_\{k\+1\}\\mid\\mathbf\{x\}\_\{k\}\)\.\(A\.2\)The Schrödinger bridge is formulated directly as the minimization of relative entropy on path space,

ℋ∗=argminℋ:ℋ0=P0,ℋN=PTDpathKL\(ℋ∥𝒢\),DpathK​L\(ℋ∥𝒢\)=∫d𝐱0:Nℋ\(𝐱0:N\)lnℋ\(𝐱0:N\)𝒢\(𝐱0:N\),\\begin\{split\}\\mathcal\{H\}^\{\*\}&=\\argmin\_\{\\mathcal\{H\}:\\;\\mathcal\{H\}\_\{0\}=P\_\{0\},\\ \\mathcal\{H\}\_\{N\}=P\_\{T\}\}\\ D^\{\\rm path\}\_\{KL\}\(\\mathcal\{H\}\\\|\\mathcal\{G\}\),\\\\ D^\{\\rm path\}\_\{KL\}\(\\mathcal\{H\}\\\|\\mathcal\{G\}\)&=\\int d\\mathbf\{x\}\_\{0:N\}\\,\\mathcal\{H\}\(\\mathbf\{x\}\_\{0:N\}\)\\ln\\frac\{\\mathcal\{H\}\(\\mathbf\{x\}\_\{0:N\}\)\}\{\\mathcal\{G\}\(\\mathbf\{x\}\_\{0:N\}\)\},\\end\{split\}\(A\.3\)with the endpoint marginals imposed as constraints\. No equivalence between a marginal\-propagator divergence and the path\-space divergence is required, and we do not assert one\. Equation \([A\.3](https://arxiv.org/html/2608.23916#A1.E3)\) is the definition of the problem in the sense of\[[37](https://arxiv.org/html/2608.23916#bib.bib37),[39](https://arxiv.org/html/2608.23916#bib.bib39)\], and the derivation below is a variation of that functional alone\.

To enforce the marginal constraints at the initial and final times, we introduce Lagrange multiplier functionsλ0​\(𝐱0\)\\lambda\_\{0\}\(\\mathbf\{x\}\_\{0\}\)andλN​\(𝐱N\)\\lambda\_\{N\}\(\\mathbf\{x\}\_\{N\}\)and form the augmented functional

ℱ\[ℋ;λ0,λN\]=DpathK​L\(ℋ∥𝒢\)\+∫d𝐲λ0\(𝐲\)\[∫d𝐱1:Nℋ\(𝐲,𝐱1:N\)−P0\(𝐲\)\]\+∫d𝐲λN\(𝐲\)\[∫d𝐱0:N−1ℋ\(𝐱0:N−1,𝐲\)−PT\(𝐲\)\]\.\\displaystyle\\begin\{split\}\\mathcal\{F\}\[\\mathcal\{H\};\\lambda\_\{0\},\\lambda\_\{N\}\]=D^\{\\text\{path\}\}\_\{KL\}\(\\mathcal\{H\}\\\|\\mathcal\{G\}\)\+\\int d\\mathbf\{y\}\\,\\lambda\_\{0\}\(\\mathbf\{y\}\)\\left\[\\int d\\mathbf\{x\}\_\{1:N\}\\,\\mathcal\{H\}\(\\mathbf\{y\},\\mathbf\{x\}\_\{1:N\}\)\-P\_\{0\}\(\\mathbf\{y\}\)\\right\]\\\\ \+\\int d\\mathbf\{y\}\\,\\lambda\_\{N\}\(\\mathbf\{y\}\)\\left\[\\int d\\mathbf\{x\}\_\{0:N\-1\}\\,\\mathcal\{H\}\(\\mathbf\{x\}\_\{0:N\-1\},\\mathbf\{y\}\)\-P\_\{T\}\(\\mathbf\{y\}\)\\right\]\.\\end\{split\}\(A\.4\)Varyingℱ\\mathcal\{F\}with respect toℋ\\mathcal\{H\}, treated as a functional derivative in the space of probability measures, and setting the variation to zero gives

ln⁡ℋ∗𝒢\+1\+λ0​\(𝐱0\)\+λN​\(𝐱N\)=0,\\ln\\frac\{\\mathcal\{H\}^\{\*\}\}\{\\mathcal\{G\}\}\+1\+\\lambda\_\{0\}\(\\mathbf\{x\}\_\{0\}\)\+\\lambda\_\{N\}\(\\mathbf\{x\}\_\{N\}\)=0,\(A\.5\)which rearranges toℋ∗\(𝐱0:N\)=f\(𝐱0\)𝒢\(𝐱0:N\)g\(𝐱N\)\\mathcal\{H\}^\{\*\}\(\\mathbf\{x\}\_\{0:N\}\)=f\(\\mathbf\{x\}\_\{0\}\)\\,\\mathcal\{G\}\(\\mathbf\{x\}\_\{0:N\}\)\\,g\(\\mathbf\{x\}\_\{N\}\), where the endpoint functions are identified as

f⁡\(𝐱0\)\\displaystyle f\(\\mathbf\{x\}\_\{0\}\)=e−12−c−λ0​\(𝐱0\),\\displaystyle=e^\{\-\\frac\{1\}\{2\}\-c\-\\lambda\_\{0\}\(\\mathbf\{x\}\_\{0\}\)\},\(A\.6a\)g⁡\(𝐱N\)\\displaystyle g\(\\mathbf\{x\}\_\{N\}\)=e−12\+c−λN​\(𝐱N\)\.\\displaystyle=e^\{\-\\frac\{1\}\{2\}\+c\-\\lambda\_\{N\}\(\\mathbf\{x\}\_\{N\}\)\}\.\(A\.6b\)The constantccreflects a global symmetry, since the transformationf→eα​ff\\rightarrow e^\{\\alpha\}f,g→e−α​gg\\rightarrow e^\{\-\\alpha\}gleaves the productf​gfginvariant\.

To fix the multipliers explicitly, we enforce the endpoint marginals\. The initial marginal ofℋ∗\\mathcal\{H\}^\{\*\}is

P0\(𝐱0\)=∫d𝐱1:Nℋ∗\(𝐱0:N\)=f\(𝐱0\)P0\(𝐱0\)∫d𝐱1:N𝒢\(𝐱1:N∣𝐱0\)g\(𝐱N\),P\_\{0\}\(\\mathbf\{x\}\_\{0\}\)=\\int d\\mathbf\{x\}\_\{1:N\}\\,\\mathcal\{H\}^\{\*\}\(\\mathbf\{x\}\_\{0:N\}\)=f\(\\mathbf\{x\}\_\{0\}\)P\_\{0\}\(\\mathbf\{x\}\_\{0\}\)\\int d\\mathbf\{x\}\_\{1:N\}\\,\\mathcal\{G\}\(\\mathbf\{x\}\_\{1:N\}\\mid\\mathbf\{x\}\_\{0\}\)\\,g\(\\mathbf\{x\}\_\{N\}\),\(A\.7\)which implies

f\(𝐱0\)=1ψ0​\(𝐱0\),ψ0\(𝐱0\):=∫d𝐱1:N𝒢\(𝐱1:N∣𝐱0\)g\(𝐱N\)\.f\(\\mathbf\{x\}\_\{0\}\)=\\frac\{1\}\{\\psi\_\{0\}\(\\mathbf\{x\}\_\{0\}\)\},\\qquad\\psi\_\{0\}\(\\mathbf\{x\}\_\{0\}\):=\\int d\\mathbf\{x\}\_\{1:N\}\\,\\mathcal\{G\}\(\\mathbf\{x\}\_\{1:N\}\\mid\\mathbf\{x\}\_\{0\}\)\\,g\(\\mathbf\{x\}\_\{N\}\)\.\(A\.8\)Similarly, enforcing the final marginal gives

PN\(𝐱N\)=g\(𝐱N\)∫d𝐱0:N−1ℋ∗\(𝐱0:N−1,𝐱N\)=g\(𝐱N\)ψ¯N\(𝐱N\),P\_\{N\}\(\\mathbf\{x\}\_\{N\}\)=g\(\\mathbf\{x\}\_\{N\}\)\\int d\\mathbf\{x\}\_\{0:N\-1\}\\,\\mathcal\{H\}^\{\*\}\(\\mathbf\{x\}\_\{0:N\-1\},\\mathbf\{x\}\_\{N\}\)=g\(\\mathbf\{x\}\_\{N\}\)\\,\\bar\{\\psi\}\_\{N\}\(\\mathbf\{x\}\_\{N\}\),\(A\.9\)where the integral over the preceding slices reproduces the forward fieldψ¯N\\bar\{\\psi\}\_\{N\}defined below\. These relations close the system, confirm the consistency of the forward and backward factorization, and recover the product identityPk​\(𝐱k\)=ψk​\(𝐱k\)​ψ¯k​\(𝐱k\)P\_\{k\}\(\\mathbf\{x\}\_\{k\}\)=\\psi\_\{k\}\(\\mathbf\{x\}\_\{k\}\)\\,\\bar\{\\psi\}\_\{k\}\(\\mathbf\{x\}\_\{k\}\)at the endpoints\.

### A\.2Recursion relations and the continuum limit

We now derive the recursion relations for the auxiliary fieldsψk\\psi\_\{k\}andψ¯k\\bar\{\\psi\}\_\{k\}that characterize the bridge at intermediate times\. The backward field is the propagation of the terminal tilt,

ψk\(𝐱k\)=∫d𝐱k\+1:Ng\(𝐱N\)∏s=kN−1G\(𝐱s\+1∣𝐱s\),\\psi\_\{k\}\(\\mathbf\{x\}\_\{k\}\)=\\int d\\mathbf\{x\}\_\{k\+1:N\}\\,g\(\\mathbf\{x\}\_\{N\}\)\\prod\_\{s=k\}^\{N\-1\}G\(\\mathbf\{x\}\_\{s\+1\}\\mid\\mathbf\{x\}\_\{s\}\),\(A\.10\)and by the Chapman–Kolmogorov equation it satisfies the backward recursion

ψk−1​\(𝐱k−1\)=∫d​𝐱k​ψk​\(𝐱k\)​G​\(𝐱k∣𝐱k−1\),\\psi\_\{k\-1\}\(\\mathbf\{x\}\_\{k\-1\}\)=\\int d\\mathbf\{x\}\_\{k\}\\,\\psi\_\{k\}\(\\mathbf\{x\}\_\{k\}\)\\,G\(\\mathbf\{x\}\_\{k\}\\mid\\mathbf\{x\}\_\{k\-1\}\),\(A\.11\)with terminal conditionψN​\(𝐱N\)=g⁡\(𝐱N\)\\psi\_\{N\}\(\\mathbf\{x\}\_\{N\}\)=g\(\\mathbf\{x\}\_\{N\}\)\. Similarly, the forward field is

ψ¯k\(𝐱k\)=∫d𝐱0:k−1f\(𝐱0\)P0\(𝐱0\)∏s=0k−1G\(𝐱s\+1∣𝐱s\),\\bar\{\\psi\}\_\{k\}\(\\mathbf\{x\}\_\{k\}\)=\\int d\\mathbf\{x\}\_\{0:k\-1\}\\,f\(\\mathbf\{x\}\_\{0\}\)P\_\{0\}\(\\mathbf\{x\}\_\{0\}\)\\prod\_\{s=0\}^\{k\-1\}G\(\\mathbf\{x\}\_\{s\+1\}\\mid\\mathbf\{x\}\_\{s\}\),\(A\.12\)and satisfies the forward recursion

ψ¯k\+1​\(𝐱k\+1\)=∫d​𝐱k​G​\(𝐱k\+1∣𝐱k\)​ψ¯k​\(𝐱k\),\\bar\{\\psi\}\_\{k\+1\}\(\\mathbf\{x\}\_\{k\+1\}\)=\\int d\\mathbf\{x\}\_\{k\}\\,G\(\\mathbf\{x\}\_\{k\+1\}\\mid\\mathbf\{x\}\_\{k\}\)\\,\\bar\{\\psi\}\_\{k\}\(\\mathbf\{x\}\_\{k\}\),\(A\.13\)with initial conditionψ¯0​\(𝐱0\)=f⁡\(𝐱0\)​P0​\(𝐱0\)\\bar\{\\psi\}\_\{0\}\(\\mathbf\{x\}\_\{0\}\)=f\(\\mathbf\{x\}\_\{0\}\)P\_\{0\}\(\\mathbf\{x\}\_\{0\}\)\. The product identityPk​\(𝐱k\)=ψk​\(𝐱k\)​ψ¯k​\(𝐱k\)P\_\{k\}\(\\mathbf\{x\}\_\{k\}\)=\\psi\_\{k\}\(\\mathbf\{x\}\_\{k\}\)\\,\\bar\{\\psi\}\_\{k\}\(\\mathbf\{x\}\_\{k\}\)follows directly from the factorized joint measure by integrating out all other time slices\.

To obtain the continuum limit of these recursions, consider the backward recursion Eq\. \([A\.11](https://arxiv.org/html/2608.23916#A1.E11)\)\. AsΔ​t→0\\Delta t\\rightarrow 0, the transition kernelG⁡\(𝐱k∣𝐱k−1\)G\(\\mathbf\{x\}\_\{k\}\\mid\\mathbf\{x\}\_\{k\-1\}\)is the short\-time propagator of a diffusion with drift𝐯\\mathbf\{v\}and diffusion coefficientγ⁡\(t\)\\gamma\(t\), and the standard Kramers–Moyal expansion\[[85](https://arxiv.org/html/2608.23916#bib.bib85)\]yields the adjoint \(backward\) Kolmogorov equation

∂tψ\+𝐯⋅∇ψ\+γ⁡\(t\)2​∇2ψ−VG​ψ=0,\\partial\_\{t\}\\psi\+\\mathbf\{v\}\\cdot\\nabla\\psi\+\\frac\{\\gamma\(t\)\}\{2\}\\nabla^\{2\}\\psi\-V\_\{G\}\\psi=0,\(A\.14\)with terminal conditionψ⁡\(𝐱,T\)=g⁡\(𝐱\)\\psi\(\\mathbf\{x\},T\)=g\(\\mathbf\{x\}\)\. The killing termVGV\_\{G\}appears if the reference process includes a potential allowing for killing, so that probability is not conserved\. Similarly, the forward recursion Eq\. \([A\.13](https://arxiv.org/html/2608.23916#A1.E13)\) yields the forward Fokker–Planck equation

∂tψ¯\+∇⋅\(𝐯​ψ¯\)−γ⁡\(t\)2​∇2ψ¯\+VG​ψ¯=0,\\partial\_\{t\}\\bar\{\\psi\}\+\\nabla\\cdot\(\\mathbf\{v\}\\bar\{\\psi\}\)\-\\frac\{\\gamma\(t\)\}\{2\}\\nabla^\{2\}\\bar\{\\psi\}\+V\_\{G\}\\bar\{\\psi\}=0,\(A\.15\)with initial conditionψ¯​\(𝐱,0\)=f⁡\(𝐱\)​P0​\(𝐱\)\\bar\{\\psi\}\(\\mathbf\{x\},0\)=f\(\\mathbf\{x\}\)P\_\{0\}\(\\mathbf\{x\}\)\. These two equations are dual under time reversal and form the continuous\-time backbone of the Schrödinger bridge\.

### A\.3Short\-time expansion of the tilted kernel

The optimal transition kernel is

H∗​\(𝐱k\+1∣𝐱k\)=ψk\+1​\(𝐱k\+1\)ψk​\(𝐱k\)​G​\(𝐱k\+1∣𝐱k\)\.H^\{\*\}\(\\mathbf\{x\}\_\{k\+1\}\\mid\\mathbf\{x\}\_\{k\}\)=\\frac\{\\psi\_\{k\+1\}\(\\mathbf\{x\}\_\{k\+1\}\)\}\{\\psi\_\{k\}\(\\mathbf\{x\}\_\{k\}\)\}\\,G\(\\mathbf\{x\}\_\{k\+1\}\\mid\\mathbf\{x\}\_\{k\}\)\.\(A\.16\)To extract its short\-time form, begin with the short\-time reference kernel

G⁡\(𝐱k\+1∣𝐱k\)=1\(2​π​γk​Δ​t\)d/2​exp⁡\[−‖Δ​𝐱−𝐯⁡\(𝐱k\)​Δ​t‖22​γk​Δ​t−VG​\(𝐱k\)​Δ​t\],G\(\\mathbf\{x\}\_\{k\+1\}\\mid\\mathbf\{x\}\_\{k\}\)=\\frac\{1\}\{\(2\\pi\\gamma\_\{k\}\\Delta t\)^\{d/2\}\}\\exp\\left\[\-\\frac\{\\\|\\Delta\\mathbf\{x\}\-\\mathbf\{v\}\(\\mathbf\{x\}\_\{k\}\)\\Delta t\\\|^\{2\}\}\{2\\gamma\_\{k\}\\Delta t\}\-V\_\{G\}\(\\mathbf\{x\}\_\{k\}\)\\Delta t\\right\],\(A\.17\)whereΔ​𝐱=𝐱k\+1−𝐱k\\Delta\\mathbf\{x\}=\\mathbf\{x\}\_\{k\+1\}\-\\mathbf\{x\}\_\{k\}andγk=γ⁡\(tk\)\\gamma\_\{k\}=\\gamma\(t\_\{k\}\)\. Under the reference measure the increment scales asΔ​𝐱∼Δ​t\\Delta\\mathbf\{x\}\\sim\\sqrt\{\\Delta t\}, so the log\-ratio of theψ\\psifields must be expanded to first order inΔ​t\\Delta tand second order inΔ​𝐱\\Delta\\mathbf\{x\},

ln⁡ψk\+1​\(𝐱k\+1\)ψk​\(𝐱k\)=Δ​𝐱⋅∇ln⁡ψk​\(𝐱k\)\+12​Δ​𝐱a​Δ​𝐱b​∂a∂bln⁡ψk​\(𝐱k\)\+Δt∂tlnψk\(𝐱k\)\+𝒪\(Δt3/2\)\.\\begin\{split\}\\ln\\frac\{\\psi\_\{k\+1\}\(\\mathbf\{x\}\_\{k\+1\}\)\}\{\\psi\_\{k\}\(\\mathbf\{x\}\_\{k\}\)\}=\{\}&\\Delta\\mathbf\{x\}\\cdot\\nabla\\ln\\psi\_\{k\}\(\\mathbf\{x\}\_\{k\}\)\+\\frac\{1\}\{2\}\\Delta\\mathbf\{x\}^\{a\}\\Delta\\mathbf\{x\}^\{b\}\\,\\partial\_\{a\}\\partial\_\{b\}\\ln\\psi\_\{k\}\(\\mathbf\{x\}\_\{k\}\)\\\\ &\+\\Delta t\\,\\partial\_\{t\}\\ln\\psi\_\{k\}\(\\mathbf\{x\}\_\{k\}\)\+\\mathcal\{O\}\(\\Delta t^\{3/2\}\)\.\\end\{split\}\(A\.18\)Substituting this expansion into Eq\. \([A\.16](https://arxiv.org/html/2608.23916#A1.E16)\) and combining with the Gaussian exponent of Eq\. \([A\.17](https://arxiv.org/html/2608.23916#A1.E17)\), the linear term inΔ​𝐱\\Delta\\mathbf\{x\}completes the square against the reference exponent and produces the effective drift𝐯∗=𝐯\+γk∇lnψk\\mathbf\{v\}\_\{\*\}=\\mathbf\{v\}\+\\gamma\_\{k\}\\nabla\\ln\\psi\_\{k\}\. The remaining terms of orderΔ​t\\Delta tcombine to give the residuals required by normalization\. The tilted kernel is therefore, to leading order, a Gaussian diffusion with the same noise coefficientγk\\gamma\_\{k\}and the effective drift

𝐯∗\(𝐱,t\)=𝐯\(𝐱,t\)\+γ\(t\)∇lnψ\(𝐱,t\)\\mathbf\{v\}\_\{\*\}\(\\mathbf\{x\},t\)=\\mathbf\{v\}\(\\mathbf\{x\},t\)\+\\gamma\(t\)\\,\\nabla\\ln\\psi\(\\mathbf\{x\},t\)\(A\.19\)in the continuum limit\. Equation \([A\.19](https://arxiv.org/html/2608.23916#A1.E19)\) is the central dynamical result of the Schrödinger bridge, since the optimal process is obtained by adding the gradient of the log\-backward field to the reference drift\.

### A\.4Conditional field construction and the posterior mixture

In the score\-matching sector of Sec\.[2](https://arxiv.org/html/2608.23916#S2), we setVG=−∇⋅𝐯V\_\{G\}=\-\\nabla\\cdot\\mathbf\{v\}so that the forward field becomes constant,ψ¯=1\\bar\{\\psi\}=1\. The marginal is thereforePt=ψtP\_\{t\}=\\psi\_\{t\}and the effective drift is𝐯∗=𝐯\+γ∇lnPt\\mathbf\{v\}\_\{\*\}=\\mathbf\{v\}\+\\gamma\\nabla\\ln P\_\{t\}\. It remains to express∇ln⁡Pt\\nabla\\ln P\_\{t\}in terms of the training data\.

Because the equation governingψ\\psiis linear, its solution can be obtained by superposition\. For a single data point𝐲\\mathbf\{y\}, define the conditional field

ψ\(𝐲\)\(𝐱,t\)=G\(𝐲,T∣𝐱,t\),\\psi^\{\(\\mathbf\{y\}\)\}\(\\mathbf\{x\},t\)=G\(\\mathbf\{y\},T\\mid\\mathbf\{x\},t\),\(A\.20\)the transition kernel from\(𝐱,t\)\(\\mathbf\{x\},t\)to the terminal point𝐲\\mathbf\{y\}at timeTT\. This field satisfies the same backward Kolmogorov equation asψ\\psi, with terminal conditionψ\(𝐲\)​\(𝐱,T\)=δd​\(𝐱−𝐲\)\\psi^\{\(\\mathbf\{y\}\)\}\(\\mathbf\{x\},T\)=\\delta^\{d\}\(\\mathbf\{x\}\-\\mathbf\{y\}\), and by linearity the full field is the average over the data distribution,

ψ⁡\(𝐱,t\)=∫d​𝐲​Pdata​\(𝐲\)​ψ\(𝐲\)​\(𝐱,t\)\.\\psi\(\\mathbf\{x\},t\)=\\int d\\mathbf\{y\}\\,P\_\{\\rm data\}\(\\mathbf\{y\}\)\\,\\psi^\{\(\\mathbf\{y\}\)\}\(\\mathbf\{x\},t\)\.\(A\.21\)
For the affine Gaussian diffusions used in practice, namely variance\-preserving or variance\-exploding schedules, the reference kernel is Gaussian, and reversing time viaτ=T−t\\tau=T\-tto pass from the generative orientation to the noising one gives

ψ\(𝐲\)​\(𝐱,t\)=qτ​\(𝐱∣𝐲\)=𝒩⁡\(𝐱,ατ​𝐲,στ2​𝐈\),\\psi^\{\(\\mathbf\{y\}\)\}\(\\mathbf\{x\},t\)=q\_\{\\tau\}\(\\mathbf\{x\}\\mid\\mathbf\{y\}\)=\\mathcal\{N\}\\\!\\left\(\\mathbf\{x\};\\alpha\_\{\\tau\}\\mathbf\{y\},\\;\\sigma\_\{\\tau\}^\{2\}\\mathbf\{I\}\\right\),\(A\.22\)whereατ\\alpha\_\{\\tau\}andστ\\sigma\_\{\\tau\}are the signal and noise schedules of the noising process\. Suppressing the orientation change in the notation from here on, the marginal is consequently

Pt​\(𝐱\)=∫d​𝐲​Pdata​\(𝐲\)​qt​\(𝐱∣𝐲\),P\_\{t\}\(\\mathbf\{x\}\)=\\int d\\mathbf\{y\}\\,P\_\{\\rm data\}\(\\mathbf\{y\}\)\\,q\_\{t\}\(\\mathbf\{x\}\\mid\\mathbf\{y\}\),\(A\.23\)which is precisely the forward noising marginal\.

The conditional bridge kernel for a fixed data point𝐲\\mathbf\{y\}follows from the tilted kernel formula,

H\(𝐲\)​\(𝐱k\+1∣𝐱k\)=ψk\+1\(𝐲\)​\(𝐱k\+1\)ψk\(𝐲\)​\(𝐱k\)​G​\(𝐱k\+1∣𝐱k\),H^\{\(\\mathbf\{y\}\)\}\(\\mathbf\{x\}\_\{k\+1\}\\mid\\mathbf\{x\}\_\{k\}\)=\\frac\{\\psi^\{\(\\mathbf\{y\}\)\}\_\{k\+1\}\(\\mathbf\{x\}\_\{k\+1\}\)\}\{\\psi^\{\(\\mathbf\{y\}\)\}\_\{k\}\(\\mathbf\{x\}\_\{k\}\)\}\\,G\(\\mathbf\{x\}\_\{k\+1\}\\mid\\mathbf\{x\}\_\{k\}\),\(A\.24\)and by the Markov property of the noising chain this simplifies to the posterior of the noising chain conditioned on the clean sample,

H\(𝐲\)​\(𝐱k\+1∣𝐱k\)=q⁡\(𝐱k\+1∣𝐱k,𝐲\)=q⁡\(𝐱k∣𝐱k\+1\)​q​\(𝐱k\+1∣𝐲\)q⁡\(𝐱k∣𝐲\)\.H^\{\(\\mathbf\{y\}\)\}\(\\mathbf\{x\}\_\{k\+1\}\\mid\\mathbf\{x\}\_\{k\}\)=q\(\\mathbf\{x\}\_\{k\+1\}\\mid\\mathbf\{x\}\_\{k\},\\mathbf\{y\}\)=\\frac\{q\(\\mathbf\{x\}\_\{k\}\\mid\\mathbf\{x\}\_\{k\+1\}\)\\,q\(\\mathbf\{x\}\_\{k\+1\}\\mid\\mathbf\{y\}\)\}\{q\(\\mathbf\{x\}\_\{k\}\\mid\\mathbf\{y\}\)\}\.\(A\.25\)Averaging over the data distribution with the posterior weight

p⁡\(𝐲∣𝐱k\)=Pdata​\(𝐲\)​q​\(𝐱k∣𝐲\)Pk​\(𝐱k\)p\(\\mathbf\{y\}\\mid\\mathbf\{x\}\_\{k\}\)=\\frac\{P\_\{\\rm data\}\(\\mathbf\{y\}\)\\,q\(\\mathbf\{x\}\_\{k\}\\mid\\mathbf\{y\}\)\}\{P\_\{k\}\(\\mathbf\{x\}\_\{k\}\)\}\(A\.26\)yields the full optimal reverse kernel,

H∗​\(𝐱k\+1∣𝐱k\)=∫d​𝐲​p​\(𝐲∣𝐱k\)​q​\(𝐱k\+1∣𝐱k,𝐲\)\.H^\{\*\}\(\\mathbf\{x\}\_\{k\+1\}\\mid\\mathbf\{x\}\_\{k\}\)=\\int d\\mathbf\{y\}\\,p\(\\mathbf\{y\}\\mid\\mathbf\{x\}\_\{k\}\)\\,q\(\\mathbf\{x\}\_\{k\+1\}\\mid\\mathbf\{x\}\_\{k\},\\mathbf\{y\}\)\.\(A\.27\)This is the posterior\-weighted mixture representation of the optimal bridge\. The generative process averages over the posterior of the clean data point given the current noisy observation and then steps according to the conditional noising posterior\.

Finally, the same construction expresses the marginal score as a posterior expectation,

∇ln⁡Pt​\(𝐱\)=∫d𝐲Pdata\(𝐲\)qt\(𝐱∣𝐲\)∇lnqt\(𝐱∣𝐲\)Pt​\(𝐱\)=𝔼p⁡\(𝐲∣𝐱\)​\[∇ln⁡qt​\(𝐱∣𝐲\)\]\.\\nabla\\ln P\_\{t\}\(\\mathbf\{x\}\)=\\frac\{\\int d\\mathbf\{y\}\\,P\_\{\\rm data\}\(\\mathbf\{y\}\)\\,q\_\{t\}\(\\mathbf\{x\}\\mid\\mathbf\{y\}\)\\,\\nabla\\ln q\_\{t\}\(\\mathbf\{x\}\\mid\\mathbf\{y\}\)\}\{P\_\{t\}\(\\mathbf\{x\}\)\}=\\mathbb\{E\}\_\{p\(\\mathbf\{y\}\\mid\\mathbf\{x\}\)\}\\left\[\\nabla\\ln q\_\{t\}\(\\mathbf\{x\}\\mid\\mathbf\{y\}\)\\right\]\.\(A\.28\)This is the identity underlying the simulation\-free score\-matching objective, in which the unknown marginal score is replaced by a tractable posterior expectation estimated in practice by sampling a clean data point and adding noise\.

### A\.5Short\-time KL divergence between diffusion kernels

We derive the per\-step KL divergence between two diffusion processes that share the same diffusion coefficient but differ in drift, the computation underlying the excess action of Sec\.[2\.3](https://arxiv.org/html/2608.23916#S2.SS3)\. Consider two transition kernels over a small intervalΔ​t\\Delta twith drifts𝐯1​\(𝐱k\)\\mathbf\{v\}\_\{1\}\(\\mathbf\{x\}\_\{k\}\),𝐯2​\(𝐱k\)\\mathbf\{v\}\_\{2\}\(\\mathbf\{x\}\_\{k\}\)and common diffusion coefficientγk\\gamma\_\{k\},

Hi\(𝐱k\+1∣𝐱k\)=1\(2​π​γk​Δ​t\)d/2exp\[−‖Δ​𝐱−𝐯i​\(𝐱k\)​Δ​t‖22​γk​Δ​t\],i=1,2,H\_\{i\}\(\\mathbf\{x\}\_\{k\+1\}\\mid\\mathbf\{x\}\_\{k\}\)=\\frac\{1\}\{\(2\\pi\\gamma\_\{k\}\\Delta t\)^\{d/2\}\}\\exp\\left\[\-\\frac\{\\\|\\Delta\\mathbf\{x\}\-\\mathbf\{v\}\_\{i\}\(\\mathbf\{x\}\_\{k\}\)\\Delta t\\\|^\{2\}\}\{2\\gamma\_\{k\}\\Delta t\}\\right\],\\quad i=1,2,\(A\.29\)whereΔ​𝐱=𝐱k\+1−𝐱k\\Delta\\mathbf\{x\}=\\mathbf\{x\}\_\{k\+1\}\-\\mathbf\{x\}\_\{k\}\. The drifts are evaluated at the initial point𝐱k\\mathbf\{x\}\_\{k\}, and higher\-order terms in the drift do not affect the leading\-order divergence\. The per\-step KL divergence is

DK​L\(H\(1\)\(⋅∣𝐱k\)∥H\(2\)\(⋅∣𝐱k\)\)=∫d𝐱kd𝐱k\+1P\(1\)k\(𝐱k\)H\(1\)\(𝐱k\+1∣𝐱k\)×ln⁡H\(1\)​\(𝐱k\+1∣𝐱k\)H\(2\)​\(𝐱k\+1∣𝐱k\),\\begin\{split\}D\_\{KL\}\\bigl\(H^\{\(1\)\}\(\\cdot\\mid\\mathbf\{x\}\_\{k\}\)\\,\\\|\\,H^\{\(2\)\}\(\\cdot\\mid\\mathbf\{x\}\_\{k\}\)\\bigr\)=\\int d\\mathbf\{x\}\_\{k\}\\,d\\mathbf\{x\}\_\{k\+1\}\\,P^\{\(1\)\}\_\{k\}\(\\mathbf\{x\}\_\{k\}\)\\,H^\{\(1\)\}\(\\mathbf\{x\}\_\{k\+1\}\\mid\\mathbf\{x\}\_\{k\}\)&\\\\ \\times\\ln\\frac\{H^\{\(1\)\}\(\\mathbf\{x\}\_\{k\+1\}\\mid\\mathbf\{x\}\_\{k\}\)\}\{H^\{\(2\)\}\(\\mathbf\{x\}\_\{k\+1\}\\mid\\mathbf\{x\}\_\{k\}\)\}&,\\end\{split\}\(A\.30\)wherePk\(1\)P^\{\(1\)\}\_\{k\}is the marginal under the first process, or any measure with the same support, since the leading\-order result is independent of the weighting\. The log\-ratio simplifies because the normalization factors cancel,

ln⁡H1H2=1γk​Δ​𝐱⋅\(𝐯1−𝐯2\)\+Δ​t2​γk​\(‖𝐯2‖2−‖𝐯1‖2\)\.\\ln\\frac\{H\_\{1\}\}\{H\_\{2\}\}=\\frac\{1\}\{\\gamma\_\{k\}\}\\Delta\\mathbf\{x\}\\cdot\(\\mathbf\{v\}\_\{1\}\-\\mathbf\{v\}\_\{2\}\)\+\\frac\{\\Delta t\}\{2\\gamma\_\{k\}\}\\left\(\\\|\\mathbf\{v\}\_\{2\}\\\|^\{2\}\-\\\|\\mathbf\{v\}\_\{1\}\\\|^\{2\}\\right\)\.\(A\.31\)UnderH1H\_\{1\}the increment is distributed as

Δ​𝐱=𝐯1​\(𝐱k\)​Δ​t\+γk​Δ​t​𝝃,𝝃∼𝒩⁡\(0,𝐈\),\\Delta\\mathbf\{x\}=\\mathbf\{v\}\_\{1\}\(\\mathbf\{x\}\_\{k\}\)\\Delta t\+\\sqrt\{\\gamma\_\{k\}\\Delta t\}\\,\\bm\{\\xi\},\\qquad\\bm\{\\xi\}\\sim\\mathcal\{N\}\(0,\\mathbf\{I\}\),\(A\.32\)and substituting into the log\-ratio and averaging over𝝃\\bm\{\\xi\}gives

𝔼H1​\[ln⁡H1H2\]\\displaystyle\\mathbb\{E\}\_\{H\_\{1\}\}\\left\[\\ln\\frac\{H\_\{1\}\}\{H\_\{2\}\}\\right\]=Δ​tγk​𝐯1⋅\(𝐯1−𝐯2\)\+Δ​t2​γk​\(‖𝐯2‖2−‖𝐯1‖2\)\\displaystyle=\\frac\{\\Delta t\}\{\\gamma\_\{k\}\}\\,\\mathbf\{v\}\_\{1\}\\cdot\(\\mathbf\{v\}\_\{1\}\-\\mathbf\{v\}\_\{2\}\)\+\\frac\{\\Delta t\}\{2\\gamma\_\{k\}\}\\left\(\\\|\\mathbf\{v\}\_\{2\}\\\|^\{2\}\-\\\|\\mathbf\{v\}\_\{1\}\\\|^\{2\}\\right\)=Δ​t2​γk​‖𝐯1−𝐯2‖2,\\displaystyle=\\frac\{\\Delta t\}\{2\\gamma\_\{k\}\}\\\|\\mathbf\{v\}\_\{1\}\-\\mathbf\{v\}\_\{2\}\\\|^\{2\},\(A\.33\)where the term proportional to𝔼⁡\[𝝃\]\\mathbb\{E\}\[\\bm\{\\xi\}\]vanishes\. This is exact for the expectation of the log\-ratio underH1H\_\{1\}, because the log\-ratio is linear inΔ​𝐱\\Delta\\mathbf\{x\}plus a deterministic constant\. No higher\-order corrections inΔ​t\\Delta tarise beyond the time\-dependence of the drifts over the interval, which contributes𝒪⁡\(Δ​t3/2\)\\mathcal\{O\}\(\\Delta t^\{3/2\}\)or higher\. Therefore

DK​L\(H1\(⋅∣𝐱k\)∥H2\(⋅∣𝐱k\)\)=Δ​t2​γk∥𝐯1\(𝐱k\)−𝐯2\(𝐱k\)∥2\+𝒪\(Δt3/2\)\.D\_\{KL\}\\bigl\(H\_\{1\}\(\\cdot\\mid\\mathbf\{x\}\_\{k\}\)\\,\\\|\\,H\_\{2\}\(\\cdot\\mid\\mathbf\{x\}\_\{k\}\)\\bigr\)=\\frac\{\\Delta t\}\{2\\gamma\_\{k\}\}\\\|\\mathbf\{v\}\_\{1\}\(\\mathbf\{x\}\_\{k\}\)\-\\mathbf\{v\}\_\{2\}\(\\mathbf\{x\}\_\{k\}\)\\\|^\{2\}\+\\mathcal\{O\}\(\\Delta t^\{3/2\}\)\.\(A\.34\)Integrating over the initial state𝐱k\\mathbf\{x\}\_\{k\}with weightPk\(1\)​\(𝐱k\)P^\{\(1\)\}\_\{k\}\(\\mathbf\{x\}\_\{k\}\)and taking the continuum limit yields the integral form used in the main text,

Δ​𝒮=12​∫0Td​tγ⁡\(t\)​𝔼𝐱∼Pt​\[‖𝐯H​\(𝐱,t\)−𝐯∗​\(𝐱,t\)‖2\]\+𝒪⁡\(Δ​t\),\\Delta\\mathcal\{S\}=\\frac\{1\}\{2\}\\int\_\{0\}^\{T\}\\frac\{dt\}\{\\gamma\(t\)\}\\,\\mathbb\{E\}\_\{\\mathbf\{x\}\\sim P\_\{t\}\}\\left\[\\\|\\mathbf\{v\}\_\{H\}\(\\mathbf\{x\},t\)\-\\mathbf\{v\}\_\{\*\}\(\\mathbf\{x\},t\)\\\|^\{2\}\\right\]\+\\mathcal\{O\}\(\\Delta t\),\(A\.35\)where the error term vanishes as the discretization is refined\. This establishes the connection between the path\-space KL divergence and the squared drift difference underlying the score\-matching objective\.

## Appendix BDetails of the Latent\-Geometry Construction

This appendix supplies the derivations for Sec\.[3](https://arxiv.org/html/2608.23916#S3)\. We compute the ambient second variation and establish the endpoint measurability of the bridge tangent space\. We then show that exact path conditioning is singular, work out the exponential\-family form of the Gaussian corruption, and carry out the pullback with its endpoint split\. The final subsections give the Gaussian evaluation and the radial information identity\.

### B\.1The ambient second variation

Let𝒢\\mathcal\{G\}be aσ\\sigma\-finite reference path measure and

Λ\[u,v\]=ln∫d𝐱0:N𝒢\(𝐱0:N\)eu⁡\(𝐱0\)\+v⁡\(𝐱N\),𝒟=\{\(u,v\):Λ\[u,v\]<∞\}\.\\Lambda\[u,v\]=\\ln\\int d\\mathbf\{x\}\_\{0:N\}\\,\\mathcal\{G\}\(\\mathbf\{x\}\_\{0:N\}\)\\,e^\{u\(\\mathbf\{x\}\_\{0\}\)\+v\(\\mathbf\{x\}\_\{N\}\)\},\\qquad\\mathcal\{D\}=\\\{\(u,v\):\\Lambda\[u,v\]<\\infty\\\}\.\(B\.1\)Writeρ\(u,v\)\\rho^\{\(u,v\)\}for the joint law of\(𝐱0,𝐱N\)\(\\mathbf\{x\}\_\{0\},\\mathbf\{x\}\_\{N\}\)underℋ\(u,v\)\\mathcal\{H\}^\{\(u,v\)\}of Eq\. \([3\.1](https://arxiv.org/html/2608.23916#S3.E1)\), and letS=δ​u​\(𝐱0\)\+δ​v​\(𝐱N\)S=\\delta u\(\\mathbf\{x\}\_\{0\}\)\+\\delta v\(\\mathbf\{x\}\_\{N\}\)for an admissible direction\(δ​u,δ​v\)\(\\delta u,\\delta v\)\.

Considerλ↦Λ⁡\[u\+λ​δ​u,v\+λ​δ​v\]\\lambda\\mapsto\\Lambda\[u\+\\lambda\\,\\delta u,\\ v\+\\lambda\\,\\delta v\]\. By construction

Λ⁡\[u\+λ​δ​u,v\+λ​δ​v\]=Λ⁡\[u,v\]\+ln⁡𝔼ρ\(u,v\)​\[eλ​S\],\\Lambda\[u\+\\lambda\\delta u,v\+\\lambda\\delta v\]=\\Lambda\[u,v\]\+\\ln\\mathbb\{E\}\_\{\\rho^\{\(u,v\)\}\}\\big\[e^\{\\lambda S\}\\big\],\(B\.2\)since tilting byλ​S\\lambda Sreweights the endpoint law by exactlyeλ​Se^\{\\lambda S\}\. The second term is the cumulant\-generating function ofSSunderρ\(u,v\)\\rho^\{\(u,v\)\}, and differentiating twice atλ=0\\lambda=0gives

dd​λ\|0=𝔼ρ​\[S\],d2d​λ2\|0=Varρ​\[S\],\\frac\{d\}\{d\\lambda\}\\Big\|\_\{0\}=\\mathbb\{E\}\_\{\\rho\}\[S\],\\qquad\\frac\{d^\{2\}\}\{d\\lambda^\{2\}\}\\Big\|\_\{0\}=\\mathrm\{Var\}\_\{\\rho\}\[S\],\(B\.3\)which is Lemma[3\.1](https://arxiv.org/html/2608.23916#S3.Thmtheorem1)\. Differentiation under the integral is justified for\(u,v\)\(u,v\)interior to𝒟\\mathcal\{D\}by the standard exponential\-integrability argument for exponential families\[[86](https://arxiv.org/html/2608.23916#bib.bib86),[32](https://arxiv.org/html/2608.23916#bib.bib32)\], because on the interior𝔼𝒢​\[e\(u\+λ​δ​u\)​\(𝐱0\)\+\(v\+λ​δ​v\)​\(𝐱N\)\]\\mathbb\{E\}\_\{\\mathcal\{G\}\}\[e^\{\(u\+\\lambda\\delta u\)\(\\mathbf\{x\}\_\{0\}\)\+\(v\+\\lambda\\delta v\)\(\\mathbf\{x\}\_\{N\}\)\}\]is finite on a neighbourhood ofλ=0\\lambda=0and analytic there\.

Nothing beyondσ\\sigma\-finiteness and exponential integrability was used\. In particular,𝒢\\mathcal\{G\}need not be a probability measure, need not be Markov, and no finite\-dimensional parameterization is involved\.

### B\.2Endpoint measurability of the tangent space

The log\-likelihood ratio between two nearby members of the tilt family is, directly from Eq\. \([3\.1](https://arxiv.org/html/2608.23916#S3.E1)\),

lnℋ\(u\+δ​u,v\+δ​v\)ℋ\(u,v\)\(𝐱0:N\)=δu\(𝐱0\)\+δv\(𝐱N\)−δΛ,\\ln\\frac\{\\mathcal\{H\}^\{\(u\+\\delta u,\\,v\+\\delta v\)\}\}\{\\mathcal\{H\}^\{\(u,v\)\}\}\(\\mathbf\{x\}\_\{0:N\}\)=\\delta u\(\\mathbf\{x\}\_\{0\}\)\+\\delta v\(\\mathbf\{x\}\_\{N\}\)\-\\delta\\Lambda,\(B\.4\)so the centred score is

δ​ℓ=δ​u​\(𝐱0\)\+δ​v​\(𝐱N\)−𝔼ρ​\[δ​u​\(𝐱0\)\+δ​v​\(𝐱N\)\],\\delta\\ell=\\delta u\(\\mathbf\{x\}\_\{0\}\)\+\\delta v\(\\mathbf\{x\}\_\{N\}\)\-\\mathbb\{E\}\_\{\\rho\}\\big\[\\delta u\(\\mathbf\{x\}\_\{0\}\)\+\\delta v\(\\mathbf\{x\}\_\{N\}\)\\big\],\(B\.5\)a function of the endpoints alone\. Every tangent vector of the bridge family is therefore represented by an endpoint\-measurable score\. The whole local statistical experiment, namely the collection of likelihood ratios distinguishing infinitesimally separated bridges, is measurable with respect toσ⁡\(𝐱0,𝐱N\)\\sigma\(\\mathbf\{x\}\_\{0\},\\mathbf\{x\}\_\{N\}\)\. Lemma[3\.1](https://arxiv.org/html/2608.23916#S3.Thmtheorem1)is the second moment of Eq\. \([B\.5](https://arxiv.org/html/2608.23916#A2.E5)\)\.

This is the precise sense in which passing from the path to the endpoint pair loses nothing relevant\. By the monotonicity of Fisher information under Markov kernels, any reduction of the path variable can only decrease the Fisher form, with equality for a sufficient reduction\[[61](https://arxiv.org/html/2608.23916#bib.bib61),[62](https://arxiv.org/html/2608.23916#bib.bib62)\]\. Equation \([B\.5](https://arxiv.org/html/2608.23916#A2.E5)\) exhibits the endpoint pair as sufficient for the tangent experiment, so the reduction is lossless at this order\.

### B\.3Singularity of exact path conditioning

Let𝒢\\mathcal\{G\}be a nondegenerate diffusion reference and considerℋ\(⋅∣Xt=𝐱\)\\mathcal\{H\}\(\\cdot\\mid X\_\{t\}=\\mathbf\{x\}\)\. The event\{Xt=𝐱\}\\\{X\_\{t\}=\\mathbf\{x\}\\\}has probability zero, and the conditional measures for distinct𝐱\\mathbf\{x\}are carried by the disjoint path sets\{ω:ω⁡\(t\)=𝐱\}\\\{\\omega:\\omega\(t\)=\\mathbf\{x\}\\\}\. Hence for𝐱≠𝐱′\\mathbf\{x\}\\neq\\mathbf\{x\}^\{\\prime\}the two measures are mutually singular and

DK​L\(ℋ\(⋅∣Xt=𝐱\)∥ℋ\(⋅∣Xt=𝐱′\)\)=\+∞\.D\_\{KL\}\\big\(\\mathcal\{H\}\(\\cdot\\mid X\_\{t\}=\\mathbf\{x\}\)\\,\\big\\\|\\,\\mathcal\{H\}\(\\cdot\\mid X\_\{t\}=\\mathbf\{x\}^\{\\prime\}\)\\big\)=\+\\infty\.\(B\.6\)A Fisher metric requiresDK​L\(Pθ\+d​θ∥Pθ\)=12gi​jdθidθj\+O\(\|dθ\|3\)D\_\{KL\}\(P\_\{\\theta\+d\\theta\}\\\|P\_\{\\theta\}\)=\\tfrac\{1\}\{2\}g\_\{ij\}d\\theta^\{i\}d\\theta^\{j\}\+O\(\|d\\theta\|^\{3\}\), but no such expansion exists here at any order\. The obstruction is not specific to Schrödinger bridges but is generic for exact point conditioning of continuous\-path measures, and it is the reason the reduction Eq\. \([3\.4](https://arxiv.org/html/2608.23916#S3.E4)\) conditions the endpoint experiment rather than the path\.

### B\.4Exponential\-family form of the Gaussian corruption

Withq⁡\(𝐱,t∣𝐱N\)=𝒩⁡\(𝐱,αt​𝐱N,σt2​𝐈\)q\(\\mathbf\{x\},t\\mid\\mathbf\{x\}\_\{N\}\)=\\mathcal\{N\}\(\\mathbf\{x\};\\alpha\_\{t\}\\mathbf\{x\}\_\{N\},\\sigma\_\{t\}^\{2\}\\mathbf\{I\}\),

ln⁡q=−‖𝐱−αt​𝐱N‖22​σt2\+const=αtσt2​𝐱⋅𝐱N−αt22​σt2​‖𝐱N‖2−‖𝐱‖22​σt2\+const,\\ln q=\-\\frac\{\\left\\\|\\mathbf\{x\}\-\\alpha\_\{t\}\\mathbf\{x\}\_\{N\}\\right\\\|^\{2\}\}\{2\\sigma\_\{t\}^\{2\}\}\+\\text\{const\}=\\frac\{\\alpha\_\{t\}\}\{\\sigma\_\{t\}^\{2\}\}\\,\\mathbf\{x\}\\cdot\\mathbf\{x\}\_\{N\}\\;\-\\;\\frac\{\\alpha\_\{t\}^\{2\}\}\{2\\sigma\_\{t\}^\{2\}\}\\left\\\|\\mathbf\{x\}\_\{N\}\\right\\\|^\{2\}\\;\-\\;\\frac\{\\left\\\|\\mathbf\{x\}\\right\\\|^\{2\}\}\{2\\sigma\_\{t\}^\{2\}\}\+\\text\{const\},\(B\.7\)which is Eq\. \([3\.6](https://arxiv.org/html/2608.23916#S3.E6)\) with

𝐬⁡\(𝐱N\)=\(𝐱N,−12​‖𝐱N‖2\),A⁡\(𝐱,t\)=\(αtσt2​𝐱,αt2σt2\),\\mathbf\{s\}\(\\mathbf\{x\}\_\{N\}\)=\\Big\(\\mathbf\{x\}\_\{N\},\-\\tfrac\{1\}\{2\}\\left\\\|\\mathbf\{x\}\_\{N\}\\right\\\|^\{2\}\\Big\),\\qquad A\(\\mathbf\{x\},t\)=\\Big\(\\frac\{\\alpha\_\{t\}\}\{\\sigma\_\{t\}^\{2\}\}\\mathbf\{x\},\\ \\frac\{\\alpha\_\{t\}^\{2\}\}\{\\sigma\_\{t\}^\{2\}\}\\Big\),\(B\.8\)andBB,CCcollecting the terms depending on𝐱\\mathbf\{x\}alone and on𝐱N\\mathbf\{x\}\_\{N\}alone\. The coefficient of the second\-moment statistic is the signal\-to\-noise ratioSNRt=αt2/σt2\\mathrm\{SNR\}\_\{t\}=\\alpha\_\{t\}^\{2\}/\\sigma\_\{t\}^\{2\}, so with𝐱~=𝐱/αt\\tilde\{\\mathbf\{x\}\}=\\mathbf\{x\}/\\alpha\_\{t\},

A⁡\(𝐱,t\)=SNRt⋅\(𝐱~,1\),A\(\\mathbf\{x\},t\)=\\mathrm\{SNR\}\_\{t\}\\cdot\(\\tilde\{\\mathbf\{x\}\},1\),\(B\.9\)which is Eq\. \([3\.8](https://arxiv.org/html/2608.23916#S3.E8)\)\. At fixed𝐱~\\tilde\{\\mathbf\{x\}\}the time derivative is∂tA=\(∂tSNRt\)​\(𝐱~,1\)\|A\\partial\_\{t\}A=\(\\partial\_\{t\}\\mathrm\{SNR\}\_\{t\}\)\\,\(\\tilde\{\\mathbf\{x\}\},1\)\\parallel A, so the temporal direction of latent spacetime is radial in natural\-parameter space and the spatial directions are angular\.

For an*anisotropic*Gaussian corruption with covarianceΣt\\Sigma\_\{t\}, the term−12​𝐱N⊤​Σt−1​𝐱N\-\\tfrac\{1\}\{2\}\\mathbf\{x\}\_\{N\}^\{\\top\}\\Sigma\_\{t\}^\{\-1\}\\mathbf\{x\}\_\{N\}does not reduce to a multiple of‖𝐱N‖2\\left\\\|\\mathbf\{x\}\_\{N\}\\right\\\|^\{2\}, and the sufficient statistic must carry the full symmetric tensor−12𝐱N⊗𝐱N\-\\tfrac\{1\}\{2\}\\mathbf\{x\}\_\{N\}\\otimes\\mathbf\{x\}\_\{N\}\. The natural\-parameter space then hasm=d\+d⁡\(d\+1\)/2m=d\+d\(d\+1\)/2dimensions while the latent manifold still hasd\+1d\+1\. The image ofιlat\\iota\_\{\\rm lat\}therefore acquires codimension, the immersion is no longer onto an open set, and the latent family becomes a curved exponential family with nonvanishing embedding curvatureΠ⟂​\[∂μ∂νA\]\\Pi^\{\\perp\}\[\\partial\_\{\\mu\}\\partial\_\{\\nu\}A\]\. Isotropy is exactly the conditionm=d\+1m=d\+1\.

### B\.5Pullback and endpoint split

The conditional endpoint law Eq\. \([3\.7](https://arxiv.org/html/2608.23916#S3.E7)\) is an exponential family with carrierν\\nuand natural parameterA⁡\(𝐱,t\)A\(\\mathbf\{x\},t\), so Lemma[3\.1](https://arxiv.org/html/2608.23916#S3.Thmtheorem1)applies to it verbatim and its Fisher form in natural coordinates isCovp​\[𝐬a,𝐬b\]\\mathrm\{Cov\}\_\{p\}\[\\mathbf\{s\}\_\{a\},\\mathbf\{s\}\_\{b\}\]\. Composing withιlat\\iota\_\{\\rm lat\}and using the chain rule gives Eq\. \([3\.9](https://arxiv.org/html/2608.23916#S3.E9)\),

gμ​ν=∂μAa​∂νAb​Covp​\[𝐬a,𝐬b\]\.g\_\{\\mu\\nu\}=\\partial\_\{\\mu\}A^\{a\}\\,\\partial\_\{\\nu\}A^\{b\}\\,\\mathrm\{Cov\}\_\{p\}\[\\mathbf\{s\}\_\{a\},\\mathbf\{s\}\_\{b\}\]\.\(B\.10\)
For the split, the bridge is Markov, so conditioning onXt=𝐱X\_\{t\}=\\mathbf\{x\}makes past and future independent,

p⁡\(𝐱0,𝐱N∣Xt=𝐱\)=p⁡\(𝐱0∣Xt=𝐱\)​p​\(𝐱N∣Xt=𝐱\)\.p\(\\mathbf\{x\}\_\{0\},\\mathbf\{x\}\_\{N\}\\mid X\_\{t\}=\\mathbf\{x\}\)=p\(\\mathbf\{x\}\_\{0\}\\mid X\_\{t\}=\\mathbf\{x\}\)\\,p\(\\mathbf\{x\}\_\{N\}\\mid X\_\{t\}=\\mathbf\{x\}\)\.\(B\.11\)The score of the joint conditional with respect to𝐱\\mathbf\{x\}is then the sum of the two individual scores, and because the two factors are independent under the conditional measure the cross\-covariance vanishes, giving Eq\. \([3\.10](https://arxiv.org/html/2608.23916#S3.E10)\) with

gi​jfuture\\displaystyle g^\{\\rm future\}\_\{ij\}=Covp⁡\(𝐱N∣𝐱,t\)​\[∂iln⁡q⁡\(𝐱,t∣𝐱N\),∂jln⁡q⁡\(𝐱,t∣𝐱N\)\],\\displaystyle=\\mathrm\{Cov\}\_\{p\(\\mathbf\{x\}\_\{N\}\\mid\\mathbf\{x\},t\)\}\\big\[\\partial\_\{i\}\\ln q\(\\mathbf\{x\},t\\mid\\mathbf\{x\}\_\{N\}\),\\ \\partial\_\{j\}\\ln q\(\\mathbf\{x\},t\\mid\\mathbf\{x\}\_\{N\}\)\\big\],\(B\.12\)gi​jpast\\displaystyle g^\{\\rm past\}\_\{ij\}=Covp⁡\(𝐱0∣𝐱,t\)​\[∂iln⁡G⁡\(𝐱,t∣𝐱0\),∂jln⁡G⁡\(𝐱,t∣𝐱0\)\]\.\\displaystyle=\\mathrm\{Cov\}\_\{p\(\\mathbf\{x\}\_\{0\}\\mid\\mathbf\{x\},t\)\}\\big\[\\partial\_\{i\}\\ln G\(\\mathbf\{x\},t\\mid\\mathbf\{x\}\_\{0\}\),\\ \\partial\_\{j\}\\ln G\(\\mathbf\{x\},t\\mid\\mathbf\{x\}\_\{0\}\)\\big\]\.\(B\.13\)No such cancellation occurs in Eq\. \([3\.3](https://arxiv.org/html/2608.23916#S3.E3)\), where the endpoints are not conditioned and their covariance is generically nonzero\.

### B\.6Gaussian evaluation

From Eq\. \([3\.5](https://arxiv.org/html/2608.23916#S3.E5)\),

∂iln⁡q⁡\(𝐱,t∣𝐱N\)=−xi−αt​xN,iσt2,\\partial\_\{i\}\\ln q\(\\mathbf\{x\},t\\mid\\mathbf\{x\}\_\{N\}\)=\-\\frac\{x\_\{i\}\-\\alpha\_\{t\}x\_\{N,i\}\}\{\\sigma\_\{t\}^\{2\}\},\(B\.14\)whose only𝐱N\\mathbf\{x\}\_\{N\}\-dependence is throughαt​xN,i/σt2\\alpha\_\{t\}x\_\{N,i\}/\\sigma\_\{t\}^\{2\}\. Hence

gi​jfuture\(𝐱,t\)=αt2σt4Cov\[xNi,xNj∣Xt=𝐱\]\.g^\{\\rm future\}\_\{ij\}\(\\mathbf\{x\},t\)=\\frac\{\\alpha\_\{t\}^\{2\}\}\{\\sigma\_\{t\}^\{4\}\}\\,\\mathrm\{Cov\}\\big\[x\_\{N\}^\{i\},x\_\{N\}^\{j\}\\mid X\_\{t\}=\\mathbf\{x\}\\big\]\.\(B\.15\)To convert this into Hessian form, write the marginal as an integral over the data distribution,Pt​\(𝐱\)=∫d​𝐱N​Pdata​\(𝐱N\)​q​\(𝐱,t∣𝐱N\)P\_\{t\}\(\\mathbf\{x\}\)=\\int d\\mathbf\{x\}\_\{N\}\\,P\_\{\\rm data\}\(\\mathbf\{x\}\_\{N\}\)\\,q\(\\mathbf\{x\},t\\mid\\mathbf\{x\}\_\{N\}\)\. Differentiating once,

∂iln⁡Pt​\(𝐱\)=𝔼⁡\[∂iln⁡q∣𝐱\]=αt​𝔼​\[xNi∣𝐱\]−xiσt2,\\partial\_\{i\}\\ln P\_\{t\}\(\\mathbf\{x\}\)=\\mathbb\{E\}\\big\[\\partial\_\{i\}\\ln q\\mid\\mathbf\{x\}\\big\]=\\frac\{\\alpha\_\{t\}\\mathbb\{E\}\[x\_\{N\}^\{i\}\\mid\\mathbf\{x\}\]\-x\_\{i\}\}\{\\sigma\_\{t\}^\{2\}\},\(B\.16\)which is the Tweedie relation\[[26](https://arxiv.org/html/2608.23916#bib.bib26),[27](https://arxiv.org/html/2608.23916#bib.bib27)\]\. Differentiating a second time and using∂j𝔼\[xNi∣𝐱\]=\(αt/σt2\)Cov\[xNi,xNj∣𝐱\]\\partial\_\{j\}\\mathbb\{E\}\[x\_\{N\}^\{i\}\\mid\\mathbf\{x\}\]=\(\\alpha\_\{t\}/\\sigma\_\{t\}^\{2\}\)\\mathrm\{Cov\}\[x\_\{N\}^\{i\},x\_\{N\}^\{j\}\\mid\\mathbf\{x\}\],

∂i∂jlnPt\(𝐱\)=−δi​jσt2\+αt2σt4Cov\[xNi,xNj∣𝐱\]\.\\partial\_\{i\}\\partial\_\{j\}\\ln P\_\{t\}\(\\mathbf\{x\}\)=\-\\frac\{\\delta\_\{ij\}\}\{\\sigma\_\{t\}^\{2\}\}\+\\frac\{\\alpha\_\{t\}^\{2\}\}\{\\sigma\_\{t\}^\{4\}\}\\mathrm\{Cov\}\\big\[x\_\{N\}^\{i\},x\_\{N\}^\{j\}\\mid\\mathbf\{x\}\\big\]\.\(B\.17\)Combining with Eq\. \([B\.15](https://arxiv.org/html/2608.23916#A2.E15)\) gives Eq\. \([3\.11](https://arxiv.org/html/2608.23916#S3.E11)\),

gi​jfuture​\(𝐱,t\)=1σt2​δi​j\+∂i∂jln⁡Pt​\(𝐱\)\.g^\{\\rm future\}\_\{ij\}\(\\mathbf\{x\},t\)=\\frac\{1\}\{\\sigma\_\{t\}^\{2\}\}\\delta\_\{ij\}\+\\partial\_\{i\}\\partial\_\{j\}\\ln P\_\{t\}\(\\mathbf\{x\}\)\.\(B\.18\)Comparison with Theorem[3\.4](https://arxiv.org/html/2608.23916#S3.Thmtheorem4)is immediate\. The theorem’s metric is𝔼⁡\[∂iln⁡p⁡\(𝐱N∣𝐱,t\)​∂jln⁡p⁡\(𝐱N∣𝐱,t\)\]\\mathbb\{E\}\[\\partial\_\{i\}\\ln p\(\\mathbf\{x\}\_\{N\}\\mid\\mathbf\{x\},t\)\\,\\partial\_\{j\}\\ln p\(\\mathbf\{x\}\_\{N\}\\mid\\mathbf\{x\},t\)\], and by Eq\. \([3\.19](https://arxiv.org/html/2608.23916#S3.E19)\) the posterior score is∂iln⁡q−∂iln⁡Pt\\partial\_\{i\}\\ln q\-\\partial\_\{i\}\\ln P\_\{t\}\. Its second moment under the posterior is the covariance of∂iln⁡q\\partial\_\{i\}\\ln q, which is exactlygfutureg^\{\\rm future\}\.

### B\.7The radial information identity

Letpθ​\(𝐱N\)∝ν⁡\(𝐱N\)​eθ⋅𝐬p\_\{\\theta\}\(\\mathbf\{x\}\_\{N\}\)\\propto\\nu\(\\mathbf\{x\}\_\{N\}\)e^\{\\theta\\cdot\\mathbf\{s\}\}andK\(θ\)=DK​L\(pθ∥ν\)K\(\\theta\)=D\_\{KL\}\(p\_\{\\theta\}\\\|\\nu\)\. UsingΛ\(θ\)=ln∫νeθ⋅𝐬\\Lambda\(\\theta\)=\\ln\\int\\nu\\,e^\{\\theta\\cdot\\mathbf\{s\}\}andη=∇Λ=𝔼pθ​\[𝐬\]\\eta=\\nabla\\Lambda=\\mathbb\{E\}\_\{p\_\{\\theta\}\}\[\\mathbf\{s\}\],

K⁡\(θ\)=𝔼pθ​\[θ⋅𝐬−Λ⁡\(θ\)\]=θ⋅η⁡\(θ\)−Λ⁡\(θ\),K\(\\theta\)=\\mathbb\{E\}\_\{p\_\{\\theta\}\}\\big\[\\theta\\cdot\\mathbf\{s\}\-\\Lambda\(\\theta\)\\big\]=\\theta\\cdot\\eta\(\\theta\)\-\\Lambda\(\\theta\),\(B\.19\)so that

∇θK=η\+θ⊤​∇2Λ−∇Λ=g⁡\(θ\)​θ,g=∇2Λ=Covpθ​\[𝐬\]\.\\nabla\_\{\\theta\}K=\\eta\+\\theta^\{\\top\}\\nabla^\{2\}\\Lambda\-\\nabla\\Lambda=g\(\\theta\)\\,\\theta,\\qquad g=\\nabla^\{2\}\\Lambda=\\mathrm\{Cov\}\_\{p\_\{\\theta\}\}\[\\mathbf\{s\}\]\.\(B\.20\)Restricting to the rayθ=r​n\\theta=r\\,nwithnnfixed and normalized,

d​Kd​r=n⋅∇θK=r​n⊤​g​n=r​gr​r,\\frac\{dK\}\{dr\}=n\\cdot\\nabla\_\{\\theta\}K=r\\,n^\{\\top\}g\\,n=r\\,g\_\{rr\},\(B\.21\)which is Eq\. \([3\.12](https://arxiv.org/html/2608.23916#S3.E12)\)\. Since by Eq\. \([3\.8](https://arxiv.org/html/2608.23916#S3.E8)\) the radial coordinate isSNRt\\mathrm\{SNR\}\_\{t\}and the angular coordinate is the rescaled latent, the corruption moves along the ray andKKdecreases monotonically asSNRt\\mathrm\{SNR\}\_\{t\}decreases, at a rate set by the radial component of the Fisher metric\. Atr=0r=0the conditional endpoint law coincides with the carrier and the latent retains nothing about the data\.

## Appendix CProofs of the Information\-Flow Identities

###### Proof of Lemma[4\.1](https://arxiv.org/html/2608.23916#S4.Thmtheorem1)\.

By Eq\. \([3\.19](https://arxiv.org/html/2608.23916#S3.E19)\) and the centering identity, for fixed𝐱\\mathbf\{x\}

𝔼𝐲\|𝐱​‖∇𝐱​ln​p​\(𝐲∣𝐱,t\)‖2\\displaystyle\\mathbb\{E\}\_\{\\mathbf\{y\}\\mid\\mathbf\{x\}\}\\big\\\|\\nabla\_\{\\\!\\mathbf\{x\}\}\\ln p\(\\mathbf\{y\}\\mid\\mathbf\{x\},t\)\\big\\\|^\{2\}=𝔼𝐲\|𝐱∥∇𝐱lnq\(𝐱,t∣𝐲\)∥2−2∇lnPt⋅𝔼𝐲\|𝐱\[∇𝐱lnq\]\\displaystyle=\\mathbb\{E\}\_\{\\mathbf\{y\}\\mid\\mathbf\{x\}\}\\big\\\|\\nabla\_\{\\\!\\mathbf\{x\}\}\\ln q\(\\mathbf\{x\},t\\mid\\mathbf\{y\}\)\\big\\\|^\{2\}\-2\\,\\nabla\\\!\\ln P\_\{t\}\\cdot\\mathbb\{E\}\_\{\\mathbf\{y\}\\mid\\mathbf\{x\}\}\\big\[\\nabla\_\{\\\!\\mathbf\{x\}\}\\ln q\\big\]\+‖∇ln⁡Pt‖2\\displaystyle\\quad\+\\big\\\|\\nabla\\\!\\ln P\_\{t\}\\big\\\|^\{2\}=𝔼𝐲\|𝐱​‖∇𝐱​ln​q​\(𝐱,t∣𝐲\)‖2−‖∇ln⁡Pt​\(𝐱\)‖2\.\\displaystyle=\\mathbb\{E\}\_\{\\mathbf\{y\}\\mid\\mathbf\{x\}\}\\big\\\|\\nabla\_\{\\\!\\mathbf\{x\}\}\\ln q\(\\mathbf\{x\},t\\mid\\mathbf\{y\}\)\\big\\\|^\{2\}\-\\big\\\|\\nabla\\\!\\ln P\_\{t\}\(\\mathbf\{x\}\)\\big\\\|^\{2\}\.Averaging over𝐱∼Pt\\mathbf\{x\}\\sim P\_\{t\}and using the tower property,𝔼Pt𝔼𝐲\|𝐱∥∇lnq∥2=𝔼𝐲𝔼𝐱\|𝐲∥∇lnq\(⋅∣𝐲\)∥2=𝔼𝐲J\(qt\(⋅∣𝐲\)\)\\mathbb\{E\}\_\{P\_\{t\}\}\\mathbb\{E\}\_\{\\mathbf\{y\}\\mid\\mathbf\{x\}\}\\\|\\nabla\\ln q\\\|^\{2\}=\\mathbb\{E\}\_\{\\mathbf\{y\}\}\\mathbb\{E\}\_\{\\mathbf\{x\}\\mid\\mathbf\{y\}\}\\\|\\nabla\\ln q\(\\cdot\\mid\\mathbf\{y\}\)\\\|^\{2\}=\\mathbb\{E\}\_\{\\mathbf\{y\}\}J\(q\_\{t\}\(\\cdot\\mid\\mathbf\{y\}\)\)\. ∎

###### Proof of Lemma[4\.2](https://arxiv.org/html/2608.23916#S4.Thmtheorem2)\.

Any solution of the Fokker–Planck equation∂tp=−∇⋅\(fp\)\+D2∇2p\\partial\_\{t\}p=\-\\nabla\\\!\\cdot\\\!\(fp\)\+\\tfrac\{D\}\{2\}\\nabla^\{2\}psatisfiesdd​t​h​\(pt\)=𝔼pt​\[∇⋅f\]\+D2​J​\(pt\)\\tfrac\{d\}\{dt\}h\(p\_\{t\}\)=\\mathbb\{E\}\_\{p\_\{t\}\}\[\\nabla\\\!\\cdot\\\!f\]\+\\tfrac\{D\}\{2\}J\(p\_\{t\}\), by two integrations by parts\. Apply this toPtP\_\{t\}and toqt\(⋅∣𝐲\)q\_\{t\}\(\\cdot\\mid\\mathbf\{y\}\)for each𝐲\\mathbf\{y\}, and average the latter over𝐲\\mathbf\{y\}\. The drift terms then agree by the tower property and cancel inI⁡\(𝐲,𝐱t\)=h⁡\(𝐱t\)−h⁡\(𝐱t∣𝐲\)I\(\\mathbf\{y\};\\mathbf\{x\}\_\{t\}\)=h\(\\mathbf\{x\}\_\{t\}\)\-h\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{y\}\)\. Eq\. \([4\.1](https://arxiv.org/html/2608.23916#S4.E1)\) converts the result into the metric trace\. ∎

###### Proof of Theorem[4\.5](https://arxiv.org/html/2608.23916#S4.Thmtheorem5)\.

Theorem[3\.4](https://arxiv.org/html/2608.23916#S3.Thmtheorem4)with the Gaussian evaluation Eq\. \([3\.11](https://arxiv.org/html/2608.23916#S3.E11)\) gives the integrandγ⁡\(t\)​\(αt2/σt4\)​tr⁡Cov⁡\[𝐲∣𝐱\]\\gamma\(t\)\(\\alpha\_\{t\}^\{2\}/\\sigma\_\{t\}^\{4\}\)\\operatorname\{tr\}\\mathrm\{Cov\}\[\\mathbf\{y\}\\mid\\mathbf\{x\}\]\. By Lemma[4\.4](https://arxiv.org/html/2608.23916#S4.Thmtheorem4)the prefactor isd​λ/d​td\\lambda/dt, so thett\-integral becomes12​∫MMSE​𝑑λ\\frac\{1\}\{2\}\\int\\mathrm\{MMSE\}\\,d\\lambda\. The second equality is the I–MMSE relation of Guo, Shamai and Verdú\[[49](https://arxiv.org/html/2608.23916#bib.bib49)\],d​I​\(𝐲,𝐱λ\)/d​λ=12​MMSE​\(λ\)dI\(\\mathbf\{y\};\\mathbf\{x\}\_\{\\lambda\}\)/d\\lambda=\\tfrac\{1\}\{2\}\\mathrm\{MMSE\}\(\\lambda\)\. The third holds becauseh⁡\(𝐱λ∣𝐲\)=h⁡\(𝐳\)h\(\\mathbf\{x\}\_\{\\lambda\}\\mid\\mathbf\{y\}\)=h\(\\mathbf\{z\}\)is independent ofλ\\lambda, soI=h⁡\(𝐱λ\)−h⁡\(𝐳\)I=h\(\\mathbf\{x\}\_\{\\lambda\}\)\-h\(\\mathbf\{z\}\)and the constant cancels in the difference\. ∎

## Appendix DNumerical and Experimental Details

This appendix collects the numerical verifications quoted in the main text\. All experiments are analytic or low\-dimensional by design, so that every quantity can be computed exactly or by quadrature\. They establish identities rather than performance\.

### D\.1de Bruijn identity for a nonlinear drift

For the nonlinear corruption driftf⁡\(x\)=−a​x3f\(x\)=\-ax^\{3\}, we solved the Fokker–Planck equation directly and evaluated both sides of Eq\. \([4\.2](https://arxiv.org/html/2608.23916#S4.E2)\)\. The two sides agree with a median relative error of1\.7×10−31\.7\\times 10^\{\-3\}point\-wise in time, and the integrated form agrees to0\.4%0\.4\\%\.

### D\.2The schedule identity

For the linear interpolantαt=t\\alpha\_\{t\}=t,σt=1−t\\sigma\_\{t\}=1\-t, the bridge coefficient of Eq\. \([4\.4](https://arxiv.org/html/2608.23916#S4.E4)\) isγ⁡\(t\)=2​σt2​dd​t​ln⁡\(αt/σt\)=2​\(1−t\)/t\\gamma\(t\)=2\\sigma\_\{t\}^\{2\}\\,\\frac\{d\}\{dt\}\\ln\(\\alpha\_\{t\}/\\sigma\_\{t\}\)=2\(1\-t\)/t, so the left\-hand side of Eq\. \([4\.5](https://arxiv.org/html/2608.23916#S4.E5)\) is2​1−tt​t2\(1−t\)4=2​t\(1−t\)32\\frac\{1\-t\}\{t\}\\frac\{t^\{2\}\}\{\(1\-t\)^\{4\}\}=\\frac\{2t\}\{\(1\-t\)^\{3\}\}, which equalsdd​t​t2\(1−t\)2=d​λtd​t\\frac\{d\}\{dt\}\\frac\{t^\{2\}\}\{\(1\-t\)^\{2\}\}=\\frac\{d\\lambda\_\{t\}\}\{dt\}\. The trigonometric and polynomial interpolants follow by the same computation, and we have verified Eq\. \([4\.5](https://arxiv.org/html/2608.23916#S4.E5)\) numerically to a relative error of3×10−63\\times 10^\{\-6\}for four distinct interpolants\.

### D\.3The floor–entropy identity

Figure[2](https://arxiv.org/html/2608.23916#S4.F2)evaluates Theorem[4\.5](https://arxiv.org/html/2608.23916#S4.Thmtheorem5)on a two\-mode Gaussian mixture for which the differential entropyh⁡\(𝐱λ\)h\(\\mathbf\{x\}\_\{\\lambda\}\)is computable by quadrature\. The accumulated floor12​∫MMSE​𝑑λ\\frac\{1\}\{2\}\\int\\mathrm\{MMSE\}\\,d\\lambdaand the entropy changeΔ​h\\Delta hagree to a relative error of2×10−52\\times 10^\{\-5\}\.

### D\.4High\-SNR slopes

Figure[4](https://arxiv.org/html/2608.23916#S4.F4)fits the divergence rate of the floor for Gaussian data of rankkkembedded inℝ10\\mathbb\{R\}^\{10\}\. The fitted slopes matchk/2k/2to three decimals, confirming Eq\. \([4\.10](https://arxiv.org/html/2608.23916#S4.E10)\)\.

### D\.5Ranking inversion

The example of Sec\.[5\.1](https://arxiv.org/html/2608.23916#S5.SS1)uses the analytic two\-mode mixture\. The two SNR ranges areAA\(narrow\) andBB\(wide\)\. The two models are exact scores scaled byc=0\.97c=0\.97\(better\) andc=0\.90c=0\.90\(worse\)\. Raw losses are1\.04901\.0490and1\.12671\.1267onAAand1\.54551\.5455and1\.86011\.8601onBB\. The floors, computed from Theorem[3\.4](https://arxiv.org/html/2608.23916#S3.Thmtheorem4)as expectations of the posterior covariance, are1\.04141\.0414and1\.51431\.5143\. The floor\-subtracted excesses are0\.0077<0\.08540\.0077<0\.0854onAAand0\.0311<0\.34580\.0311<0\.3458onBB\.

### D\.6Third\-cumulant localization

The measurements of Sec\.[5\.3](https://arxiv.org/html/2608.23916#S5.SS3)evaluate both terms ofd2​η/d​λ2d^\{2\}\\eta/d\\lambda^\{2\}in Eq\. \([5\.1](https://arxiv.org/html/2608.23916#S5.E1)\) along exactly integrated probability\-flow trajectories on the same mixture\. The share ofT⁡\[u,u\]T\[u,u\]in‖d2​η/d​λ2‖\\\|d^\{2\}\\eta/d\\lambda^\{2\}\\\|is0\.050\.05att=0\.10t=0\.10,0\.850\.85att=0\.30t=0\.30,1\.191\.19att=0\.45t=0\.45\(values above one indicating partial cancellation between the two terms\), and0\.110\.11att=0\.65t=0\.65, with a median of0\.470\.47across trajectories\. The step\-placement comparison of Fig\.[5](https://arxiv.org/html/2608.23916#S5.F5)uses an explicit first\-order Euler solver, with endpoint error measured against the exactly integrated probability\-flow ODE\.

### D\.7The discrete floor

Forn=3n=3tokens over an alphabet of size33with total correlation0\.460\.46nats, Eq\. \([6\.1](https://arxiv.org/html/2608.23916#S6.E1)\) reproduces the data entropyH\(y1:n\)=2\.58757H\(y\_\{1:n\}\)=2\.58757to a relative error of8×10−88\\times 10^\{\-8\}\. Integrating inttagainst the weight−m˙t/mt\-\\dot\{m\}\_\{t\}/m\_\{t\}returns the same value for the masking schedulesm=1−tm=1\-t,cos⁡\(π​t/2\)\\cos\(\\pi t/2\),\(1−t\)3\(1\-t\)^\{3\}, and1−t\\sqrt\{1\-t\}, confirming schedule independence\.

Similar Articles

Fisher Widths: Local Learning Geometry and Anisotropic Recovery

arXiv cs.LG

This paper introduces Fisher width and inverse-Fisher width on statistical manifolds, studying their roles in local learning bounds and anisotropic recovery. It proves a complementary relation between the two widths and obtains recovery estimates based on Fisher geometry.