Diffusion models recover accurate mixture weights despite score function insensitivity

arXiv cs.LG Papers

Summary

This paper resolves the paradox that diffusion models can accurately recover mixture weights even when the score function is insensitive to them, introducing the Diffusion Score Sensitivity Index (DSSI) and showing that intermediate noise levels provide informative signals for weight recovery.

arXiv:2607.15485v1 Announce Type: new Abstract: Score-based generative models exhibit a puzzling behavior: they often appear to cover all modes of a target multimodal distribution and yet may fail to learn the correct relative mode amplitudes, which can be interpreted as mixture weights. We resolve this apparent paradox by relating the diffusion score matching (DSM) loss to the error in estimating mixture weights from generated samples. We show that, even when the target score is insensitive to mixture weights, generated samples can recover the weights accurately if the scores at intermediate noise levels are informative about the weights. Accordingly, we define the diffusion score sensitivity index (DSSI) as the variation in the DSM loss relative to changes in a parameter. We then show that the DSSI governs the accuracy with which the parameter of the target distribution can be estimated from generated samples. For Gaussian mixtures in arbitrary dimensions, we prove that the mixture weight estimation errors are on the same order as the DSM loss under mild conditions. Empirically, we show the emergence of sensitivity during the noising process of benchmark data distributions under typical noise schedules, and that these sensitivity values predict how well a well-trained model recovers mixture weights. Furthermore, we show that the choice of noise schedule can reduce diffusion sensitivity, leading to mode amplification. Although we focus on mixture weights, the proposed sensitivity framework governs the recovery of any qualitative parameter of the target distribution.
Original Article
View Cached Full Text

Cached at: 07/20/26, 09:28 AM

# Diffusion models recover accurate mixture weights despite score function insensitivity
Source: [https://arxiv.org/html/2607.15485](https://arxiv.org/html/2607.15485)
Andrew DennehyCorresponding Author:[nishac@uchicago\.edu](https://arxiv.org/html/2607.15485v1/mailto:[email protected])Committee on Computational and Applied Mathematics, University of ChicagoDepartment of Statistics, University of ChicagoData Science Institute, University of Chicago Rebecca WillettCommittee on Computational and Applied Mathematics, University of ChicagoDepartment of Statistics, University of ChicagoData Science Institute, University of ChicagoDepartment of Computer Science, University of ChicagoNSF\-Simons National Institute for Theory and Mathematics in BiologyNSF\-Simons National AI Institute for the SkyNisha ChandramoorthyCommittee on Computational and Applied Mathematics, University of ChicagoDepartment of Statistics, University of ChicagoNSF\-Simons National Institute for Theory and Mathematics in BiologyNSF\-Simons National AI Institute for the Sky

###### Abstract

Score\-based generative models exhibit a puzzling behavior: they often appear to cover all modes of a target multimodal distribution and yet may fail to learn the correct relative mode amplitudes, which can be interpreted as mixture weights\. We resolve this apparent paradox by relating the diffusion score matching \(DSM\) loss to the error in estimating mixture weights from generated samples\. We show that, even when the target score is insensitive to mixture weights, generated samples can recover the weights accurately if the scores at intermediate noise levels are informative about the weights\. Accordingly, we define thediffusion score sensitivity index\(DSSI\) as the variation in the DSM loss relative to changes in a parameter\. We then show that the DSSI governs the accuracy with which the parameter of the target distribution can be estimated from generated samples\. For Gaussian mixtures in arbitrary dimensions, we prove that the mixture weight estimation errors are on the same order as the DSM loss under mild conditions\. Empirically, we show the emergence of sensitivity during the noising process of benchmark data distributions under typical noise schedules, and that these sensitivity values predict how well a well\-trained model recovers mixture weights\. Furthermore, we show that the choice of noise schedule can reduce diffusion sensitivity, leading to mode amplification\. Although we focus on mixture weights, the proposed sensitivity framework governs the recovery of any qualitative parameter of the target distribution\.

## 1Introduction

Generative models are powerful tools that are increasingly used in scientific research\[[29](https://arxiv.org/html/2607.15485#bib.bib4),[43](https://arxiv.org/html/2607.15485#bib.bib18),[19](https://arxiv.org/html/2607.15485#bib.bib17)\], medicine\[[35](https://arxiv.org/html/2607.15485#bib.bib16),[15](https://arxiv.org/html/2607.15485#bib.bib15),[56](https://arxiv.org/html/2607.15485#bib.bib14)\], the justice system\[[37](https://arxiv.org/html/2607.15485#bib.bib12),[42](https://arxiv.org/html/2607.15485#bib.bib13)\], and more\. In these and other settings, users rely on generative models to accurately reflect the qualitative properties of the true target distribution\. Past work notes that learning the score of a target distribution may be insufficient for accurate sampling\[[27](https://arxiv.org/html/2607.15485#bib.bib27),[11](https://arxiv.org/html/2607.15485#bib.bib26),[34](https://arxiv.org/html/2607.15485#bib.bib28)\], particularly for canonical multimodal distributions such as Gaussian mixtures\[[3](https://arxiv.org/html/2607.15485#bib.bib78),[60](https://arxiv.org/html/2607.15485#bib.bib46)\]\. Generative models that incorrectly reflect the amplitudes or variances of modes\[[63](https://arxiv.org/html/2607.15485#bib.bib58),[47](https://arxiv.org/html/2607.15485#bib.bib8),[18](https://arxiv.org/html/2607.15485#bib.bib10)\]may hinder uncertainty quantification and the recovery of qualitative parameters\[[8](https://arxiv.org/html/2607.15485#bib.bib51),[1](https://arxiv.org/html/2607.15485#bib.bib52),[38](https://arxiv.org/html/2607.15485#bib.bib57)\]\. Our work proposes an explicit evaluation of generative models through the lens of parameter recovery\. Such an evaluation augments sample\-quality metrics, such as the Frechet inception distance \(FID\)\[[22](https://arxiv.org/html/2607.15485#bib.bib31)\], which may not explicitly test for target\-specific structure, such as multimodality\[[48](https://arxiv.org/html/2607.15485#bib.bib7)\]\.

In this work, we present an information\-theoretic mechanism that determines when such key parameters may be represented accurately by diffusion models\.In the Gaussian mixture setting, even when the score matching loss is insensitive to mixture weights, learning a diffusion model anneals the score matching loss across steps of the diffusion process\[[52](https://arxiv.org/html/2607.15485#bib.bib61)\]; this annealing leads to sensitivity to mixture weights, resulting in accurate estimates that were unanticipated by past analyses that focused only on the \(unannealed\) score matching loss\[[27](https://arxiv.org/html/2607.15485#bib.bib27)\]\.

This idea is illustrated in[Figure1](https://arxiv.org/html/2607.15485#S1.F1)\. The top row plots the score functions for two Gaussian mixture distributions that are identical except for the mixture weightγ\\gammaand shows that they are nearly identical except in a region with very low probability mass\. We are thus unlikely to have sufficient training samples for the loss to reliably indicate the value ofγ\\gammathat best represents the training data\. However, for larger values ofttin the diffusion process \(t=0\.25t=0\.25ort=0\.50t=0\.50in the figure\), the corresponding distributions are less well separated, and the score functions are clearly distinguishable over a higher\-probability region\. Since training diffusion models uses the score across multiplet∈\[0,1\]t\\in\[0,1\], learned models hone in on the correct value ofγ\\gamma\.

![Refer to caption](https://arxiv.org/html/2607.15485v1/x1.png)Figure 1:Score sensitivity to mixture weights at different diffusion times\. Att=0t=0, the scores forγ=0\.8\\gamma=0\.8andγ=0\.5\\gamma=0\.5are nearly identical except on a low\-probability region\. Att\>0t\>0, the mixture components overlap more, and the region where the scores differ has a substantially higher probability\.To quantify this phenomenon, we measure the sensitivity of the diffusion score matching lossℓDSM\\ell\_\{\\mathrm\{DSM\}\}\([8](https://arxiv.org/html/2607.15485#S3.E8)\) to changes in parameters of interest\. Specifically, let\{p\(γ\)\}γ∈𝚪\\\{p^\{\(\\gamma\)\}\\\}\_\{\\gamma\\in\\bm\{\\Gamma\}\}be a set of parameterized distributions\. Letp\(γ∗\)p^\{\(\\gamma^\{\*\}\)\}be a target distribution with target parameterγ∗\\gamma^\{\*\}\. Thediffusion score sensitivity index \(DSSI\)of the target parameterγ∗\\gamma^\{\*\}is defined as,

L​\(γ∗\)=infγ∈𝚪ℓDSM​\(p\(γ\),p\(γ∗\)\)\(γ−γ∗\)2L\(\\gamma^\{\*\}\)=\\inf\_\{\\gamma\\in\\bm\{\\Gamma\}\}~\\frac\{\\ell\_\{\\mathrm\{DSM\}\}\(p^\{\(\\gamma\)\},p^\{\(\\gamma^\{\*\}\)\}\)\}\{\(\\gamma\-\\gamma^\{\*\}\)^\{2\}\}\(1\)
We note that bothℓDSM\\ell\_\{\\mathrm\{DSM\}\}andL​\(γ∗\)L\(\\gamma^\{\*\}\)depend on our choice of noise schedule\. If the target data distribution has mixture weightγ∗\\gamma^\{\*\}, and we learn a distribution with mixture parameterγ\\gamma, then the squared error of mixture weight estimationγ\\gammais bounded by the diffusion score matching loss divided by the DSSI\. Thus, when the DSSI is small, our learned distribution may be highly misrepresentative ofγ∗\\gamma^\{\*\}even when the diffusion score matching loss is low\. In contrast, whenL​\(γ∗\)L\(\\gamma^\{\*\}\)is large, we know that a small diffusion score matching loss indicates low error in the parameterγ\\gamma\. Our main contributions are as follows:

- •For general bimodal mixtures, we prove a strict minimax separation in mixture\-weight recovery between estimators that use only the target\-score data and estimators with access to score data at an intermediate, mode\-overlapping noise level\. \(Theorem[3\.5](https://arxiv.org/html/2607.15485#S3.Thmtheorem5)\)
- •For Gaussian mixtures, we prove a lower bound on the DSSI that holds in any dimension; thus, parameter estimation error is controlled by the DSM loss \(Theorem[4\.1](https://arxiv.org/html/2607.15485#S4.Thmtheorem1)\)\.
- •We further empirically evaluate the DSSI on a real\-world dataset\. These experiments illustrate how design decisions relating to diffusion model training and inference\-time sampling schedules can affect the DSSI and the fidelity of generated samples to the training data distribution\.

## 2Preliminaries: score\-based diffusion models

Score\-based diffusion models sample from a target distribution onℝd\\mathbb\{R\}^\{d\}by tracing out a one\-parameter family of distributions\{pt\}t∈\[0,1\]\\\{p\_\{t\}\\\}\_\{t\\in\[0,1\]\}indexed by a diffusion timett, withp0p\_\{0\}being the target data distribution andp1≈𝒩​\(𝟎,𝐈\)p\_\{1\}\\approx\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)the standard normal distribution\. A*forward process*progressively transforms samples fromp0p\_\{0\}into samples fromp1p\_\{1\}along this family\. A*reverse process*inverts the densities along the forward process, starting from𝒩​\(0,𝐈\)\\mathcal\{N\}\(0,\\mathbf\{I\}\), using learned approximations of the scores of the intermediate marginalsptp\_\{t\}\. Thus, score\-based diffusions have 3 components: \(a\) the forward diffusion process and its corresponding reverse process, \(b\) a training objective for learning the marginal score functions∇log⁡pt\\nabla\\log p\_\{t\}of the intermediate distributions fort∈\[0,1\]t\\in\[0,1\], and \(c\) a sampling procedure used to approximately generate samples from the target data distribution at inference time\. In this section, we provide a brief description of each component\.

##### Noising process\.

Given a data distributionp0p\_\{0\}onℝd\\mathbb\{R\}^\{d\}, the noising process we consider is the variance\-preserving forward SDE \(VP\-SDE\) proposed inSonget al\.\[[53](https://arxiv.org/html/2607.15485#bib.bib24)\]\. Give a noise schedule\{βt\}\\\{\\beta\_\{t\}\\\}, this takes the form

d​𝐱t=−12​βt​𝐱t​d​t\+βt​d​𝐰t,𝐱0∼p0,t∈\[0,1\],\\mathrm\{d\}\{\\mathbf\{x\}\}\_\{t\}\\;=\\;\-\\tfrac\{1\}\{2\}\\beta\_\{t\}\\,\{\\mathbf\{x\}\}\_\{t\}\\,\\mathrm\{d\}t\\;\+\\;\\sqrt\{\\beta\_\{t\}\}\\,\\mathrm\{d\}\\mathbf\{w\}\_\{t\},\\qquad\{\\mathbf\{x\}\}\_\{0\}\\sim p\_\{0\},\\quad t\\in\[0,1\],\(2\)where𝐰t\\mathbf\{w\}\_\{t\}is a standarddd\-dimensional Wiener process\. We denote byptp\_\{t\}the marginal distribution of𝐱t\{\\mathbf\{x\}\}\_\{t\}\. The forward conditional law of𝐱t\{\\mathbf\{x\}\}\_\{t\}given𝐱0\{\\mathbf\{x\}\}\_\{0\}under \([2](https://arxiv.org/html/2607.15485#S2.E2)\) is Gaussian,

pt​\(𝐱t∣𝐱0\)=𝒩​\(𝐱t;αt​𝐱0,\(1−αt\)​𝐈\),p\_\{t\}\(\{\\mathbf\{x\}\}\_\{t\}\\mid\{\\mathbf\{x\}\}\_\{0\}\)\\;=\\;\\mathcal\{N\}\\\!\\left\(\{\\mathbf\{x\}\}\_\{t\};\\;\\sqrt\{\\alpha\_\{t\}\}\\,\{\\mathbf\{x\}\}\_\{0\},\\;\(1\-\\alpha\_\{t\}\)\\,\\mathbf\{I\}\\right\),\(3\)with conditional meanαt​𝐱0\\sqrt\{\\alpha\_\{t\}\}\{\\mathbf\{x\}\}\_\{0\}and conditional covariance\(1−αt\)​𝐈\(1\-\\alpha\_\{t\}\)\\mathbf\{I\}, whereαt=e−∫0tβs​ds\\alpha\_\{t\}=e^\{\-\\int\_\{0\}^\{t\}\\beta\_\{s\}\\mathrm\{d\}s\}\.111As a slight abuse of terminology, the sequences\{αt\}\\\{\\alpha\_\{t\}\\\}and\{βt\}\\\{\\beta\_\{t\}\\\}will interchangably be referred to as the “noise schedule”\.

##### Denoising process\.

Anderson \[[4](https://arxiv.org/html/2607.15485#bib.bib25)\]established that the time\-reversal of \([2](https://arxiv.org/html/2607.15485#S2.E2)\) is itself an SDE:

d​𝐱t=\[−12​βt​𝐱t−βt​∇𝐱log⁡pt​\(𝐱t\)\]​d​t\+βt​d​𝐰¯t,𝐱1∼p1,\\mathrm\{d\}\{\\mathbf\{x\}\}\_\{t\}\\;=\\;\\left\[\-\\tfrac\{1\}\{2\}\{\\beta\_\{t\}\}\\,\{\\mathbf\{x\}\}\_\{t\}\\;\-\\;\{\\beta\_\{t\}\}\\,\\nabla\_\{\{\\mathbf\{x\}\}\}\\log p\_\{t\}\(\{\\mathbf\{x\}\}\_\{t\}\)\\right\]\\mathrm\{d\}t\\;\+\\;\\sqrt\{\{\\beta\_\{t\}\}\}\\,\\mathrm\{d\}\\bar\{\\mathbf\{w\}\}\_\{t\},\\qquad\{\\mathbf\{x\}\}\_\{1\}\\sim p\_\{1\},\(4\)where𝐰¯t\\bar\{\\mathbf\{w\}\}\_\{t\}is the reverse\-time Wiener process\. Starting from𝐱1∼p1\{\\mathbf\{x\}\}\_\{1\}\\sim p\_\{1\}and integrating the reverse SDE with the exact score function from timet=1t=1tot=0t=0produces samples𝐱0∼p0\{\\mathbf\{x\}\}\_\{0\}\\sim p\_\{0\}\. In practice, this reverse SDE is simulated starting from𝐱1∼𝒩​\(0,𝐈\)≈p1\{\\mathbf\{x\}\}\_\{1\}\\sim\\mathcal\{N\}\(0,\\mathbf\{I\}\)\\approx p\_\{1\}using the learned approximation𝐬𝜽:ℝd×\[0,1\]→ℝd\\mathbf\{s\}\_\{\{\\bm\{\\theta\}\}\}:\\mathbb\{R\}^\{d\}\\times\[0,1\]\\rightarrow\\mathbb\{R\}^\{d\}with parameters𝜽∈𝚯\{\\bm\{\\theta\}\}\\in\\bm\{\\Theta\}trained such that𝐬𝜽​\(𝐱,t\)≈∇log⁡pt​\(𝐱\)\\mathbf\{s\}\_\{\{\\bm\{\\theta\}\}\}\(\{\\mathbf\{x\}\},t\)\\approx\\nabla\\log p\_\{t\}\(\{\\mathbf\{x\}\}\)simultaneously for all noise levels to produce samples approximately distributed as the targetp0p\_\{0\}\.

For a given trained score model\{𝐬𝜽​\(⋅,t\)\}t∈\[0,1\]\\\{\\mathbf\{s\}\_\{\{\\bm\{\\theta\}\}\}\(\\cdot,t\)\\\}\_\{t\\in\[0,1\]\}\(or the true score function\{∇log⁡pt\}t∈\[0,1\]\\\{\\nabla\\log p\_\{t\}\\\}\_\{t\\in\[0,1\]\}\), generation of samples is an inference task: one may simulate the reverse SDE in \([4](https://arxiv.org/html/2607.15485#S2.E4)\) with any choice of consistent numerical integration schemes\[[25](https://arxiv.org/html/2607.15485#bib.bib5)\]\. The overall process has two distinct sources of error: the*approximation error*incurred in learning the score and the*sampling error*incurred by the choice of sampling algorithm \(e\.g\., discretization error, finite timesteps, etc\.\)\.

## 3Minimax bounds for parameter estimation from diffusion models

In this section, we show that when the score along the noising process exhibits sensitivity toγ\\gamma, the information\-theoretic barrier that disallows its recovery from the target score alone is removed\. We prove that when the score is sensitive at some noise level, we obtain a lower parameter estimation error when compared to only using the target score\. While stated explicitly for mixture weight as the parameter, the results in this section apply to any parameter\. We consider only bimodal mixtures for simplicity, but make no distributional assumptions about the components\.

##### Mixture Family\.

Letp\(0\)p^\{\(0\)\}andp\(1\)p^\{\(1\)\}be two fixed probability distributions onℝd\\mathbb\{R\}^\{d\}\. A probability densityp\(γ\),γ∈𝚪⊆\(0,1\),p^\{\(\\gamma\)\},\\gamma\\in\\bm\{\\Gamma\}\\subseteq\(0,1\),is a member of a family of two\-component mixtures if it is of the formp\(γ\)≔\(1−γ\)​p\(0\)\+γ​p\(1\)p^\{\(\\gamma\)\}\\coloneqq\(1\-\\gamma\)\\\>p^\{\(0\)\}\+\\gamma\\\>p^\{\(1\)\}\. The mixture parameterγ\\gammacontrols the relative weights of the two component distributions,p\(0\)p^\{\(0\)\}andp\(1\)p^\{\(1\)\}\. We now consider the evolution of a mixture,p\(γ\),p^\{\(\\gamma\)\},under a noising process or forward process in a score\-based diffusion model\. Specifically, whenp0≡p\(γ\),p\_\{0\}\\equiv p^\{\(\\gamma\)\},we denote bypt\(γ\)p\_\{t\}^\{\(\\gamma\)\}the marginals of𝐱t\\mathbf\{x\}\_\{t\}following the process \([2](https://arxiv.org/html/2607.15485#S2.E2)\)\. Let𝐬\(γ\)\\mathbf\{s\}^\{\(\\gamma\)\}denote the score of the densityp\(γ\)p^\{\(\\gamma\)\}, and similar for marginal densities of𝐱t\\mathbf\{x\}\_\{t\}:

𝐬\(γ\):=∇log⁡p\(γ\)and𝐬t\(γ\):=∇log⁡pt\(γ\)\.\\mathbf\{s\}^\{\(\\gamma\)\}:=\\nabla\\log p^\{\(\\gamma\)\}\\qquad\\text\{and\}\\qquad\\mathbf\{s\}\_\{t\}^\{\(\\gamma\)\}:=\\nabla\\log p\_\{t\}^\{\(\\gamma\)\}\.\(5\)The relative weights of the mixture are preserved under the forward process: the noisy marginalpt\(γ\)p\_\{t\}^\{\(\\gamma\)\}is itself aγ\\gamma\-mixture of the noisy component marginalspt\(0\)p\_\{t\}^\{\(0\)\}andpt\(1\)p\_\{t\}^\{\(1\)\}\(see[LemmaD\.1](https://arxiv.org/html/2607.15485#A4.Thmtheorem1)\)\. As a consequence the score of the noisy mixture densitypt\(γ\)p\_\{t\}^\{\(\\gamma\)\}can be decomposed as follows,

𝐬t\(γ\)​\(𝐱\)=wt\(γ,0\)​\(𝐱\)​𝐬t\(0\)​\(𝐱\)\+wt\(γ,1\)​\(𝐱\)​𝐬t\(1\)​\(𝐱\)\\displaystyle\\mathbf\{s\}\_\{t\}^\{\(\\gamma\)\}\(\{\\mathbf\{x\}\}\)=w\_\{t\}^\{\(\\gamma,0\)\}\(\{\\mathbf\{x\}\}\)\\,\\mathbf\{s\}\_\{t\}^\{\(0\)\}\(\{\\mathbf\{x\}\}\)\\;\+\\;w\_\{t\}^\{\(\\gamma,1\)\}\(\{\\mathbf\{x\}\}\)\\,\\mathbf\{s\}\_\{t\}^\{\(1\)\}\(\{\\mathbf\{x\}\}\)\(6\)with component weightswt\(γ,0\)​\(𝐱\)=\(1−γ\)​pt\(0\)​\(𝐱\)/pt\(γ\)​\(𝐱\)w\_\{t\}^\{\(\\gamma,0\)\}\(\{\\mathbf\{x\}\}\)=\(1\-\\gamma\)\\,p\_\{t\}^\{\(0\)\}\(\{\\mathbf\{x\}\}\)/p\_\{t\}^\{\(\\gamma\)\}\(\{\\mathbf\{x\}\}\)andwt\(γ,1\)​\(𝐱\)=γ​pt\(1\)​\(𝐱\)/pt\(γ\)​\(𝐱\)w\_\{t\}^\{\(\\gamma,1\)\}\(\{\\mathbf\{x\}\}\)=\\gamma\\,p\_\{t\}^\{\(1\)\}\(\{\\mathbf\{x\}\}\)/p\_\{t\}^\{\(\\gamma\)\}\(\{\\mathbf\{x\}\}\)\. This decomposition enables a simplified theoretical analysis, as we see in the next section\.

##### Classical score matching loss/relative Fisher information\.

The classical score matching \(SM\) loss measures the relative Fisher information between distributions\[[24](https://arxiv.org/html/2607.15485#bib.bib65)\]\. For a pair of mixtures,

ℓSM​\(p\(γ′\),p\(γ\)\):=𝔼𝐱∼p\(γ\)​‖𝐬\(γ′\)​\(𝐱\)−𝐬\(γ\)​\(𝐱\)‖2\.\\displaystyle\\ell\_\{\\mathrm\{SM\}\}\(p^\{\(\\gamma^\{\\prime\}\)\},p^\{\(\\gamma\)\}\):=\\mathbb\{E\}\_\{\\mathbf\{x\}\\sim p^\{\(\\gamma\)\}\}\\\|\\mathbf\{s\}^\{\(\\gamma^\{\\prime\}\)\}\(\\mathbf\{x\}\)\-\\mathbf\{s\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\)\\\|^\{2\}\.\(7\)The relative Fisher information between a pair of mixture distributions with identical components but different mixture weights is small when the components are well\-separated\. In[Figure1](https://arxiv.org/html/2607.15485#S1.F1), the scores of Gaussian mixtures may appear nearly identical outside a low\-probability region, an effect strengthened with increasing separation between mixture components as shown in[Figure2\(d\)](https://arxiv.org/html/2607.15485#S3.F2.sf4)\. Any sampling algorithm that relies solely on the score, such as Langevin sampling, exhibits near\-identical behavior across distinct mixture distributions; consequently, the generated samples may not accurately reflect the target mixture weights\.

##### Diffusion score matching loss for mixtures\.

The diffusion score matching \(DSM\) loss anneals the score matching loss across the noising process \([2](https://arxiv.org/html/2607.15485#S2.E2)\)\[[52](https://arxiv.org/html/2607.15485#bib.bib61)\]\. For a pair of mixtures,

ℓDSM​\(p\(γ′\),p\(γ\)\)\\displaystyle\\ell\_\{\\mathrm\{DSM\}\}\(p^\{\(\\gamma^\{\\prime\}\)\},p^\{\(\\gamma\)\}\):=𝔼t∼Unif​\(0,1\)𝐱∼pt\(γ\)​\[λt​‖𝐬t\(γ′\)​\(𝐱\)−𝐬t\(γ\)​\(𝐱\)‖2\]=𝔼t∼Unif​\(0,1\)​\[λt​ℓSM​\(pt\(γ′\),pt\(γ\)\)\]\.\\displaystyle:=\\mathbb\{E\}\_\{\\begin\{subarray\}\{c\}t\\sim\\mathrm\{Unif\}\(0,1\)\\\\ \\mathbf\{x\}\\sim p\_\{t\}^\{\(\\gamma\)\}\\end\{subarray\}\}\[\\lambda\_\{t\}\\\|\\mathbf\{s\}\_\{t\}^\{\(\\gamma^\{\\prime\}\)\}\(\\mathbf\{x\}\)\-\\mathbf\{s\}\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\)\\\|^\{2\}\]=\\mathbb\{E\}\_\{t\\sim\\mathrm\{Unif\}\(0,1\)\}\[\\lambda\_\{t\}\\\>\\ell\_\{\\mathrm\{SM\}\}\(p\_\{t\}^\{\(\\gamma^\{\\prime\}\)\},p\_\{t\}^\{\(\\gamma\)\}\)\]\.\(8\)Here,λt\\lambda\_\{t\}is a weighting factor set toλt=1−αt\\lambda\_\{t\}=1\-\\alpha\_\{t\}followingSonget al\.\[[53](https://arxiv.org/html/2607.15485#bib.bib24)\]\. The DSM loss measures the relative Fisher information over the entire sequence of noised densities generated by the process \([2](https://arxiv.org/html/2607.15485#S2.E2)\)\. Recall that the marginalsptp\_\{t\}’s and scores𝐬t\\mathbf\{s\}\_\{t\}’s depend on the noise schedule\{αt\}t\\\{\\alpha\_\{t\}\\\}\_\{t\}, as shown in \([3](https://arxiv.org/html/2607.15485#S2.E3)\) and \([5](https://arxiv.org/html/2607.15485#S3.E5)\), and this way, the diffusion score matching lossℓDSM\\ell\_\{\\mathrm\{DSM\}\}is dependent on the choice of noise schedule\. In practice, we may train a neural network𝐬𝜽\\mathbf\{s\}\_\{\{\\bm\{\\theta\}\}\}to approximate the score function\. With some abuse of notation, we denotep\(𝜽\)p^\{\(\{\\bm\{\\theta\}\}\)\}the distribution222The superscript here indicates the trainable parameters𝜽\{\\bm\{\\theta\}\}of the learned network\.of samples generated using the score functions\{𝐬𝜽​\(⋅,t\)\}t∈\(0,1\)\\\{\\mathbf\{s\}\_\{\{\\bm\{\\theta\}\}\}\(\\cdot,t\)\\\}\_\{t\\in\(0,1\)\}for a fixed noise schedule as per \([4](https://arxiv.org/html/2607.15485#S2.E4)\)\. Since the true score function𝐬t\(γ\)\\mathbf\{s\}\_\{t\}^\{\(\\gamma\)\}and indeed the target distributionp\(γ\)p^\{\(\\gamma\)\}are unknown, calculating the diffusion score matching loss is intractable\. Instead, we minimize the objective

ℓNN\(p\(𝜽\),p\(γ\)\):=𝔼t∼Unif​\(0,1\)𝐱0∼p\(γ\),𝐱t∼pt\(γ\)\(⋅\|𝐱0\)\[λt∥𝐬𝜽\(𝐱t,t\)−∇logpt\(γ\)\(𝐱t\|𝐱0\)∥2\]\.\\ell\_\{\\mathrm\{NN\}\}\(p^\{\(\{\\bm\{\\theta\}\}\)\},p^\{\(\\gamma\)\}\):=\\mathbb\{E\}\_\{\\begin\{subarray\}\{c\}t\\sim\\mathrm\{Unif\}\(0,1\)\\\\ \\mathbf\{x\}\_\{0\}\\sim p^\{\(\\gamma\)\},\\mathbf\{x\}\_\{t\}\\sim p\_\{t\}^\{\(\\gamma\)\}\(\\cdot\|\\mathbf\{x\}\_\{0\}\)\\end\{subarray\}\}\[\\lambda\_\{t\}\\\|\\mathbf\{\\mathbf\{s\}\}\_\{\{\\bm\{\\theta\}\}\}\(\\mathbf\{x\}\_\{t\},t\)\-\\nabla\\log p\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\|\\mathbf\{x\}\_\{0\}\)\\\|^\{2\}\]\.\(9\)
As perLABEL:\{lem:nn\-dsm\-ordering\}, there exists a constantC≥0C\\geq 0, depending only onp\(γ\)p^\{\(\\gamma\)\}, such thatℓNN​\(p\(𝜽\),p\(γ\)\)=ℓDSM​\(p\(𝜽\),pt\(γ\)\)\+C\\ell\_\{\\mathrm\{NN\}\}\(p^\{\(\{\\bm\{\\theta\}\}\)\},p^\{\(\\gamma\)\}\)=\\ell\_\{\\mathrm\{DSM\}\}\(p^\{\(\{\\bm\{\\theta\}\}\)\},p\_\{t\}^\{\(\\gamma\)\}\)\+C\. Thus,ℓDSM​\(p\(𝜽\),p\(γ\)\)≤ℓNN​\(p\(𝜽\),p\(γ\)\)\\ell\_\{\\mathrm\{DSM\}\}\(p^\{\(\{\\bm\{\\theta\}\}\)\},p^\{\(\\gamma\)\}\)\\leq\\ell\_\{\\mathrm\{NN\}\}\(p^\{\(\{\\bm\{\\theta\}\}\)\},p^\{\(\\gamma\)\}\)\.

##### Minimax theory for mixture weight recovery\.

From classical minimax theory \(see e\.g\.,\[[17](https://arxiv.org/html/2607.15485#bib.bib2)\]for a textbook exposition\), we know that the error in any estimator of a parameter is lower\-bounded by the informativeness of the data distribution about the parameter\. Now, the data distribution loses information about the parameter in the noising process\. How, then, would annealing benefit arameter recovery, if the original mixture is well\-separated and hence insensitive to parameter changes? The answer lies in the fact that diffusion models learn to approximate the*scores*at various noise levels\. By minimizingℓDSM,\\ell\_\{\\mathrm\{DSM\}\},diffusion models obtain information about the parameter through the scores even if such information is not present in the samples\. Our main result in this section shows the existence of a noise scale at which the score\-data law is sensitive, provably enabling parameter recovery rates that are unattainable from target\-score data alone\.

###### Definition 3\.1\(Score oracle and score\-restricted estimators\)\.

We consider a Gaussian White Noise \(GWN\) oracle from nonparametric regression\[[57](https://arxiv.org/html/2607.15485#bib.bib1)\]of the form:𝐬^t\(γ\)=𝐬t\(γ\)\+σt​W,\\hat\{\\mathbf\{s\}\}\_\{t\}^\{\(\\gamma\)\}=\\mathbf\{s\}\_\{t\}^\{\(\\gamma\)\}\\\>\+\\sigma\_\{t\}\\\>W,whereσt\>0\\sigma\_\{t\}\>0is a scalar andWWis an isonormal field indexed byL2​\(pt\(γ\)\)L^\{2\}\(p\_\{t\}^\{\(\\gamma\)\}\)\. The statistical experiment is to estimateγ\\gammafrom the observation,𝐬^t\(γ\)\.\\hat\{\\mathbf\{s\}\}^\{\(\\gamma\)\}\_\{t\}\.Letqt\(γ\)q\_\{t\}^\{\(\\gamma\)\}denote the distribution of the observation, which is a Gaussian measure with mean𝐬t\(γ\)\.\\mathbf\{s\}\_\{t\}^\{\(\\gamma\)\}\.We define the set,Γscore,t,\\Gamma\_\{\\mathrm\{score\},t\},of all measurable functions of the GWN observations to\(0,1\)\(0,1\)to be the class of “score\-restricted estimators”\. Further, we use𝔼γ∗,t\\mathbb\{E\}\_\{\\gamma^\{\*\},t\}as a shorthand for expectations with respect to the score oracle model,qt\(γ\)\.q\_\{t\}^\{\(\\gamma\)\}\.See also Remark[4](https://arxiv.org/html/2607.15485#Thmremark4)for more details on the estimator classes\.

We callΓscore,t\\Gamma\_\{\\mathrm\{score\},t\}*score\-restricted estimators*since they access the score vector field \(corrupt by Gaussian noise\) and not the target data samples\. In other words, our class of estimatorsΓscore,t\\Gamma\_\{\\mathrm\{score\},t\}models parameter recovery achievable from having oracle access to a corrupt score vector field at timet,t,reflecting the denoising/generation phase in diffusion models after scores have been trained\. Since the observation model,qt\(γ\),q\_\{t\}^\{\(\\gamma\)\},is Gaussian, we can compute, fixing the indexing Hilbert space forWWto beL2​\(pt\(γ\)\),L^\{2\}\(p\_\{t\}^\{\(\\gamma\)\}\),thatKL\(qt\(γ\)\|\|qt\(γ′\)\)=ℓSM\(pt\(γ′\),pt\(γ\)\)/\(2σt2\)\.\\mathrm\{KL\}\(q\_\{t\}^\{\(\\gamma\)\}\|\|q\_\{t\}^\{\(\\gamma^\{\\prime\}\)\}\)=\\ell\_\{\\mathrm\{SM\}\}\(p\_\{t\}^\{\(\\gamma^\{\\prime\}\)\},p\_\{t\}^\{\(\\gamma\)\}\)/\(2\\sigma\_\{t\}^\{2\}\)\.See Remark[5](https://arxiv.org/html/2607.15485#Thmremark5)for this computation\.

###### Assumption 3\.2\(Stable score\-matching at timett\)\.

We assume that at some timet\>0,t\>0,the score observation law,qt\(\(⋅\)\),q\_\{t\}^\{\(\(\\cdot\)\)\},satisfies\(γ−γ′\)2≤At2KL\(qt\(γ\)\|\|qt\(γ′\)\)\(\\gamma\-\\gamma^\{\\prime\}\)^\{2\}\\leq A\_\{t\}^\{2\}\\\>\\mathrm\{KL\}\(q\_\{t\}^\{\(\\gamma\)\}\|\|q\_\{t\}^\{\(\\gamma^\{\\prime\}\)\}\)for someAt\>0A\_\{t\}\>0and allγ,γ′∈\(0,1\)\.\\gamma,\\gamma^\{\\prime\}\\in\(0,1\)\.

SinceKL\(qt\(γ\)\|\|qt\(γ′\)\)=ℓSM\(pt\(γ′\),pt\(γ\)\)/\(2σt2\),\\mathrm\{KL\}\(q\_\{t\}^\{\(\\gamma\)\}\|\|q\_\{t\}^\{\(\\gamma^\{\\prime\}\)\}\)=\\ell\_\{\\mathrm\{SM\}\}\(p\_\{t\}^\{\(\\gamma^\{\\prime\}\)\},p\_\{t\}^\{\(\\gamma\)\}\)/\(2\\sigma\_\{t\}^\{2\}\),the above assumption implies that when the score matching loss is bounded above by someC,C,then, the squared difference in the corresponding parameters,\(γ−γ′\)2≤At2​C/\(2​σt2\)\.\(\\gamma\-\\gamma^\{\\prime\}\)^\{2\}\\leq\\\>A\_\{t\}^\{2\}\\\>C/\(2\\sigma\_\{t\}^\{2\}\)\.Intuitively, this means that, for anO​\(σt2\)O\(\\sigma\_\{t\}^\{2\}\)constantAt,A\_\{t\},the score observations predict distances on the space of weight parameters\. Our next definition allows us to make this setting rigorous by comparing score observations at timettto a different behavior of the scores at time0:0:the target score observations,𝐬^\(γ\),\\hat\{\\mathbf\{s\}\}^\{\(\\gamma\)\},are insensitive in law to changes inγ\\gammaand therefore, do not predict distances on parameter space\. In the next section, note that we obtain explicit estimates onAtA\_\{t\}for Gaussian mixtures\.

###### Definition 3\.3\(\(ϵ,δ\)\(\\epsilon,\\delta\)\-diffusion sensitivity\)\.

We say that a mixture family\{p\(γ\)\}\\\{p^\{\(\\gamma\)\}\\\}exhibits anϵ\\epsilon\-diffusion sensitivity if for everyγ∗∈\(0,1\),\\gamma^\{\*\}\\in\(0,1\),there existϵ\\epsilon\-packing setsS​\(γ∗\)⊂\(0,1\)S\(\\gamma^\{\*\}\)\\subset\(0,1\)containingN​\(γ∗,ϵ\)≥16N\(\\gamma^\{\*\},\\epsilon\)\\geq 16elements withKL\(q\(γ\)\|\|q\(γ′\)\)=ℓSM\(p\(γ′\),p\(γ\)\)/\(2σ02\)≤δ\(ϵ,N\)≤ϵ2/4,\\mathrm\{KL\}\(q^\{\(\\gamma\)\}\|\|q^\{\(\\gamma^\{\\prime\}\)\}\)=\\ell\_\{\\mathrm\{SM\}\}\(p^\{\(\\gamma^\{\\prime\}\)\},p^\{\(\\gamma\)\}\)/\(2\\sigma\_\{0\}^\{2\}\)\\leq\\delta\(\\epsilon,N\)\\leq\\epsilon^\{2\}/4,for allγ,γ′∈S​\(γ∗\)\\gamma,\\gamma^\{\\prime\}\\in S\(\\gamma^\{\*\}\)satisfying\|γ−γ′\|≥4​2​At​ϵ\|\\gamma\-\\gamma^\{\\prime\}\|\\geq 4\\sqrt\{2\}\\\>A\_\{t\}\\epsilon\. Note thatAtA\_\{t\}here is as defined in Assumption[3\.2](https://arxiv.org/html/2607.15485#S3.Thmtheorem2)\.

###### Definition 3\.4\(ϵ\\epsilon\-diffusion sensitivity\)\.

We say that a mixture family\{p\(γ\)\}\\\{p^\{\(\\gamma\)\}\\\}exhibitsϵ\\epsilon\-diffusion sensitivity if for everyγ∗∈\(0,1\),\\gamma^\{\*\}\\in\(0,1\),there existϵ\\epsilon\-packing setsS​\(γ∗\)⊂\(0,1\)S\(\\gamma^\{\*\}\)\\subset\(0,1\)containingN​\(γ∗,ϵ\)≥16N\(\\gamma^\{\*\},\\epsilon\)\\geq 16elements withKL\(q\(γ\)\|\|q\(γ′\)\)=ℓSM\(p\(γ′\),p\(γ\)\)/\(2σ02\)≤ϵ2/4,\\mathrm\{KL\}\(q^\{\(\\gamma\)\}\|\|q^\{\(\\gamma^\{\\prime\}\)\}\)=\\ell\_\{\\mathrm\{SM\}\}\(p^\{\(\\gamma^\{\\prime\}\)\},p^\{\(\\gamma\)\}\)/\(2\\sigma\_\{0\}^\{2\}\)\\leq\\epsilon^\{2\}/4,for allγ,γ′∈S​\(γ∗\)\\gamma,\\gamma^\{\\prime\}\\in S\(\\gamma^\{\*\}\)satisfying\|γ−γ′\|≥4​2​At​ϵ\|\\gamma\-\\gamma^\{\\prime\}\|\\geq 4\\sqrt\{2\}\\\>A\_\{t\}\\epsilon\. Note thatAtA\_\{t\}here is as defined in Assumption[3\.2](https://arxiv.org/html/2607.15485#S3.Thmtheorem2)\.

###### Theorem 3\.5\.

Fixt\>0t\>0satifying Assumption[3\.2](https://arxiv.org/html/2607.15485#S3.Thmtheorem2)\. Letϵt\>0\\epsilon\_\{t\}\>0be a covering radius that solvesϵt:=σt​V​\(ϵt\),\\epsilon\_\{t\}:=\\sigma\_\{t\}\\\>\\sqrt\{V\(\\epsilon\_\{t\}\)\},whereV​\(ϵt\)V\(\\epsilon\_\{t\}\)is theϵt\\epsilon\_\{t\}\-covering entropy333Refer[DefinitionC\.3](https://arxiv.org/html/2607.15485#A3.Thmtheorem3)for an explicit statementwith respect to square root KL divergence of the class of Gaussian measures\{qt\(γ\)\}\.\\\{q\_\{t\}^\{\(\\gamma\)\}\\\}\.Suppose the bimodal mixture family satisfies[3\.2](https://arxiv.org/html/2607.15485#S3.Thmtheorem2)and exhibits anϵ\\epsilon\-diffusion sensitivity withϵ=ϵt\.\\epsilon=\\epsilon\_\{t\}\.Then, we have the following minimax upper bound on the score\-restricted estimators at timett,

minγ^∈Γscore,t⁡maxγ∗⁡𝔼γ∗,t​\|γ∗−γ^\|2≤2​At2​ϵt2\.\\displaystyle\\min\_\{\\hat\{\\gamma\}\\in\\Gamma\_\{\\mathrm\{score\},t\}\}\\max\_\{\\gamma^\{\*\}\}~\\mathbb\{E\}\_\{\\gamma^\{\*\},t\}\|\\gamma^\{\*\}\-\\hat\{\\gamma\}\|^\{2\}\\leq 2\\\>A\_\{t\}^\{2\}\\\>\\epsilon\_\{t\}^\{2\}\\\>\.\(10\)Simultaneously, if the packing set cardinalityN​\(γ∗,ϵt\)≥eσt2​V​\(ϵt\),N\(\\gamma^\{\*\},\\epsilon\_\{t\}\)\\geq e^\{\\sigma\_\{t\}^\{2\}\\\>V\(\\epsilon\_\{t\}\)\},for allγ∗,\\gamma^\{\*\},it holds that the minimax lower bound achievable for target score\-restricted estimators is the following:

minγ^∈Γscore,0⁡maxγ∗⁡𝔼γ∗,0​\|γ∗−γ^\|2≥4​At2​ϵt2\.\\displaystyle\\min\_\{\\hat\{\\gamma\}\\in\\Gamma\_\{\\mathrm\{score\},0\}\}\\max\_\{\\gamma^\{\*\}\}~\\mathbb\{E\}\_\{\\gamma^\{\*\},0\}\|\\gamma^\{\*\}\-\\hat\{\\gamma\}\|^\{2\}\\geq 4\\\>A\_\{t\}^\{2\}\\\>\\epsilon\_\{t\}^\{2\}\.\(11\)

##### Proof\.

The upper bound in \([10](https://arxiv.org/html/2607.15485#S3.E10)\) is a direct application of Theorem 2 from\[[61](https://arxiv.org/html/2607.15485#bib.bib3)\]\. Then, using a generalized Fano’s method \(Lemma 3 of\[[62](https://arxiv.org/html/2607.15485#bib.bib30)\]\) and an appropriate choice ofδ\\deltaas a function ofϵ\\epsilonin[Definition3\.4](https://arxiv.org/html/2607.15485#S3.Thmtheorem4)that limits the packing entropy of the cover of\(0,1\)\(0,1\)with cardinalityeV​\(ϵt\),e^\{V\(\\epsilon\_\{t\}\)\},we obtain the desired lower bound in \([11](https://arxiv.org/html/2607.15485#S3.E11)\)\. We defer a more detailed proof to Appendix[C](https://arxiv.org/html/2607.15485#A3)\. ∎

Theorem[3\.5](https://arxiv.org/html/2607.15485#S3.Thmtheorem5)elucidates the conditions under which a diffusion model learns the mixture weights accurately when target score estimation followed by Langevin sampling does not\. The above analysis applies to any bimodal mixture density and to any scalar parameter in place of the mixture weight\. The main mechanism that enables parameter recovery is that the score field at some timettchanges more rapidly with parameter changes\. In contrast, the target score at time 0 may not vary, creating an information\-theoretic barrier that limits the accuracy of any estimator\. That is, no estimator can distinguish two close values ofγ\\gammausing a finite number of target score samples\.

We remark thatYang and Barron \[[61](https://arxiv.org/html/2607.15485#bib.bib3)\]constructs a Bayes density estimator that belongs to the mixture class and achieves the upper bound \([10](https://arxiv.org/html/2607.15485#S3.E10)\)\. Thus, we can conclude from Theorem[3\.5](https://arxiv.org/html/2607.15485#S3.Thmtheorem5)that using score observations at timettleads to provably better parameter recovery\.

Resolving data\-processing inequality\.Saying a family of distributions exhibits an\(ϵ,δ\)\(\\epsilon,\\delta\)\-sensitivity essentially means that more information about the parameter is created in the noising process, which, at first glance, appears to violate the data processing inequality\. The law of noised samples,pt\(γ\),p\_\{t\}^\{\(\\gamma\)\},cannot be more informative aboutγ\\gammaat any timettthanp\(γ\)p^\{\(\\gamma\)\}\(the target mixture\), due to the data processing inequality\. More precisely, sincept\(γ\)p\_\{t\}^\{\(\\gamma\)\}results from aγ\\gamma\-independent convolution operation onp0\(γ\),p\_\{0\}^\{\(\\gamma\)\},the data processing inequality gives:KL\(pt\(γ′\)\|\|pt\(γ\)\)≤KL\(p\(γ′\)\|\|p\(γ\)\)\.\\mathrm\{KL\}\(p\_\{t\}^\{\(\\gamma^\{\\prime\}\)\}\|\|p\_\{t\}^\{\(\\gamma\)\}\)\\leq\\mathrm\{KL\}\(p^\{\(\\gamma^\{\\prime\}\)\}\|\|p^\{\(\\gamma\)\}\)\.But, our key idea here is that diffusion models use the score observation fields,𝐬^t\(γ\),\\hat\{\\mathbf\{s\}\}^\{\(\\gamma\)\}\_\{t\},to recoverγ\\gamma, even when𝐬^\(γ\)\\hat\{\\mathbf\{s\}\}^\{\(\\gamma\)\}can not separateγ\\gammavalues sufficiently\. Specifically, under the assumptions of Theorem[3\.5](https://arxiv.org/html/2607.15485#S3.Thmtheorem5)on the mixture family, the score observations can be more informative aboutγ\\gammaat timettthan at time 0, i\.e\.,KL\(qt\(γ′\)\|\|qt\(γ\)\)≥KL\(q\(γ′\)\|\|q\(γ\)\),\\mathrm\{KL\}\(q\_\{t\}^\{\(\\gamma^\{\\prime\}\)\}\|\|q\_\{t\}^\{\(\\gamma\)\}\)\\geq\\mathrm\{KL\}\(q^\{\(\\gamma^\{\\prime\}\)\}\|\|q^\{\(\\gamma\)\}\),without violating data processing\. In other words,[Definition3\.4](https://arxiv.org/html/2607.15485#S3.Thmtheorem4)defines a structural property of mixture distributions: the family,\{qt\(γ\)\}\\\{q\_\{t\}^\{\(\\gamma\)\}\\\}at intermediatettcan be more separated inγ\\gammathan att=0t=0when mixture components do not overlap\. In Lemma[C\.1](https://arxiv.org/html/2607.15485#A3.Thmtheorem1), we prove why this counterintuitive inequality does not violate data processing and why no estimator inΓ0\\Gamma\_\{0\}can reconstruct the score observation model thatΓt\\Gamma\_\{t\}has access to\.

### 3\.1Effect of modified sensitivity to mixture weight

While Theorem[3\.5](https://arxiv.org/html/2607.15485#S3.Thmtheorem5)provides a mechanism for parameter recovery, we now discuss experimental evidence for the same mechanism but using the practically computable diffusion score sensitivity index \([1](https://arxiv.org/html/2607.15485#S1.E1)\), rather than the information\-theoretic quantities used in[Definition3\.4](https://arxiv.org/html/2607.15485#S3.Thmtheorem4)\. Our numerical results show that the score sensitivity index quantifies how accurately a diffusion model captures the mixture weight of a Gaussian mixture model\. We generate samples using the true analytical score function𝐬t\(γ∗\)​\(𝐱\)\\mathbf\{s\}\_\{t\}^\{\(\\gamma^\{\*\}\)\}\(\\mathbf\{x\}\)and a sequence of modified noise schedules with a coarse time discretization \(T=100T=100\), designed to artificially reduce the diffusion score sensitivity indexL​\(γ∗\)L\(\\gamma^\{\*\}\)\(for details, see[SectionE\.1](https://arxiv.org/html/2607.15485#A5.SS1)\)\. Recall thatL​\(γ∗\)L\(\\gamma^\{\*\}\)andℓDSM\\ell\_\{\\mathrm\{DSM\}\}depend on the choice of noise schedule\{αt\}t\\\{\\alpha\_\{t\}\\\}\_\{t\}\.

We modify both linear and squared\-cosine noise schedules and numerically estimateL​\(γ∗\)L\(\\gamma^\{\*\}\)for these schedules\.[Figure2\(c\)](https://arxiv.org/html/2607.15485#S3.F2.sf3)shows the estimated mixture weightγ^\\hat\{\\gamma\}\(calculated using Expectation\-Maximization\) plotted againstL​\(γ∗\)L\(\\gamma^\{\*\}\)\. We see that as the sensitivity indexL​\(γ∗\)L\(\\gamma^\{\*\}\)decreases, both the bias and variance inγ^\\hat\{\\gamma\}increase\. This demonstrates that noise schedules with lowerL​\(γ∗\)L\(\\gamma^\{\*\}\)yield mixture\-weight estimates that are less robust to discretization errors, even with a perfect score estimator\.

On the other hand, the means and covariances of the modes are uniformly well\-approximated across the different noise schedules\. This shows that errors in estimating the mixture weight can occur without noticeable artifacts in the geometry of individual modes\. Modern accelerated diffusion sampling schemes \(e\.g\.,\[[50](https://arxiv.org/html/2607.15485#bib.bib11)\]\) leverage coarse time discretizations to reduce the computational cost of sampling while maintaining comparable sample fidelity\. Our results suggest that such sampling approaches may be vulnerable to misrepresenting parameters of interest in the data distribution while appearing to generate accurate samples\.

![Refer to caption](https://arxiv.org/html/2607.15485v1/x2.png)\(a\)
![Refer to caption](https://arxiv.org/html/2607.15485v1/x3.png)\(b\)
![Refer to caption](https://arxiv.org/html/2607.15485v1/x4.png)\(c\)
![Refer to caption](https://arxiv.org/html/2607.15485v1/x5.png)\(d\)

Figure 2:Score sensitivity for Gaussian mixture models in dimensiond=10d=10\. \(a\)ℓSM​\(pt\(γ\),pt\(γ∗\)\)\\ell\_\{\\mathrm\{SM\}\}\(p\_\{t\}^\{\(\\gamma\)\},p\_\{t\}^\{\(\\gamma^\{\*\}\)\}\)as a function ofttandγ\\gamma\. \(b\)ℓDSM​\(p\(γ\),p\(γ∗\)\)\\ell\_\{\\mathrm\{DSM\}\}\(p^\{\(\\gamma\)\},p^\{\(\\gamma^\{\*\}\)\}\)vs\.γ−γ∗\\gamma\-\\gamma^\{\*\}for multipleγ∗\\gamma^\{\*\}, showing non\-negligable sensitivity throughout\. \(c\) Mixture weight estimatesγ^\\hat\{\\gamma\}from generated data forγ∗=0\.6\\gamma^\{\*\}=0\.6under modified linear and cosine noise schedules \(see[SectionE\.1](https://arxiv.org/html/2607.15485#A5.SS1)\)\. The bias and variance ofγ^\\hat\{\\gamma\}increase asL​\(γ∗\)L\(\\gamma^\{\*\}\)decreases\. \(d\)ℓSM\\ell\_\{\\mathrm\{SM\}\}andℓDSM\\ell\_\{\\mathrm\{DSM\}\}as a function of separation forγ∗=0\.8\\gamma^\{\*\}=0\.8andγ=0\.9\\gamma=0\.9\.ℓDSM\\ell\_\{\\mathrm\{DSM\}\}remains nonnegligible whileℓSM\\ell\_\{\\mathrm\{SM\}\}rapidly decays\.

## 4Diffusion sensitivity to GMM mixture weights

We show that for a bimodal isotropic Gaussian mixture in any dimension, the diffusion score sensitivity index \(DSSI\)L​\(γ∗\)L\(\\gamma^\{\*\}\)is non\-negligible\. As a result, a diffusion model with a well\-trained score generates samples with the correct weights even for extreme values ofγ\\gamma\.

###### Theorem 4\.1\(Diffusion sensitivity in bimodal Gaussian mixtures\)\.

Consider the familyp\(γ\)​\(𝐱\)=\(1−γ\)​𝒩​\(𝐱;μ0,𝐈\)\+γ​𝒩​\(𝐱;μ1,𝐈\)p^\{\(\\gamma\)\}\(\\mathbf\{x\}\)=\(1\-\\gamma\)\\;\\mathcal\{N\}\(\\mathbf\{x\};\\mu\_\{0\},\\mathbf\{I\}\)\+\\gamma\\;\\mathcal\{N\}\(\\mathbf\{x\};\\mu\_\{1\},\\mathbf\{I\}\)and a linear noise scheduleβt=β1​t\\beta\_\{t\}=\\beta\_\{1\}tforβ1\>0\\beta\_\{1\}\>0\. Letαt=exp⁡\(∫0tβs​ds\)\\alpha\_\{t\}=\\exp\\left\(\\int\_\{0\}^\{t\}\\beta\_\{s\}\\mathrm\{d\}s\\right\)\. Then for anya∈\(0,1/2\)a\\in\(0,1/2\)

minγ^∈\[a,1−a\]⁡ℓDSM​\(p\(γ^\),p\(γ∗\)\)\(γ^−γ∗\)2\>\(1−Φ​\(∥μ1−μ0∥​α1\)\)​\(k2\+16−k\)7​β1​γ∗​\(1−γ∗\)​exp⁡\(−\(k2\+k\)/4\)\\displaystyle\\min\_\{\\hat\{\\gamma\}\\in\[a,1\-a\]\}\\frac\{\\ell\_\{\\mathrm\{DSM\}\}\(p^\{\(\\hat\{\\gamma\}\)\},p^\{\(\\gamma^\{\*\}\)\}\)\}\{\(\\hat\{\\gamma\}\-\\gamma^\{\*\}\)^\{2\}\}\>\\frac\{\(1\-\\Phi\(\\lVert\\mu\_\{1\}\-\\mu\_\{0\}\\rVert\\sqrt\{\\alpha\_\{1\}\}\)\)\\left\(\\sqrt\{k^\{2\}\+16\}\-k\\right\)\}\{7\\beta\_\{1\}\\gamma^\{\*\}\(1\-\\gamma^\{\*\}\)\}\\exp\\left\(\-\(k^\{2\}\+k\)/4\\right\)\(12\)withk:=\|log⁡1−aa\|\+\|log⁡1−γ∗γ∗\|k:=\\left\|\\log\\frac\{1\-a\}\{a\}\\right\|\+\\left\|\\log\\frac\{1\-\\gamma^\{\*\}\}\{\\gamma^\{\*\}\}\\right\|andΦ​\(⋅\)\\Phi\(\\cdot\)the CDF of𝒩​\(0,1\)\\mathcal\{N\}\(0,1\)\.

*Proof Sketch:*The proof proceeds in three steps\. We first decomposeℓDSM​\(p\(γ^\),p\(γ∗\)\)\\ell\_\{\\mathrm\{DSM\}\}\(p^\{\(\\hat\{\\gamma\}\)\},p^\{\(\\gamma^\{\*\}\)\}\)by mixture component \([38](https://arxiv.org/html/2607.15485#A4.E38)\) to reduce the problem to lower\-bounding the expectations over paths fromp\(0\)p^\{\(0\)\}andp\(1\)p^\{\(1\)\}independently\. Second, using the Probability Flow ODE \(which is equivalent in probability to the forward process\)\[[53](https://arxiv.org/html/2607.15485#bib.bib24)\], we find an analytical form for∥𝐬t\(γ^\)−𝐬t\(γ∗\)∥2\\lVert\\mathbf\{s\}\_\{t\}^\{\(\\hat\{\\gamma\}\)\}\-\\mathbf\{s\}\_\{t\}^\{\(\\gamma^\{\*\}\)\}\\rVert^\{2\}along paths𝐳t\\mathbf\{z\}\_\{t\}starting fromp\(0\)p^\{\(0\)\}\([PropositionD\.6](https://arxiv.org/html/2607.15485#A4.Thmtheorem6)\)\. Finally, we identify the interval ofttvalues where this expression is bounded below \([PropositionD\.9](https://arxiv.org/html/2607.15485#A4.Thmtheorem9)\), integrate over this region oftt\([PropositionD\.10](https://arxiv.org/html/2607.15485#A4.Thmtheorem10)\), and finally integrate over initial conditions𝐲0\\mathbf\{y\}\_\{0\}\. The bound is dimension\-independent because for a given initial condition𝐲0\\mathbf\{y\}\_\{0\}starting from e\.g\. modep\(0\)p^\{\(0\)\}, the bound only depends on⟨𝐲0−μ0,μ1−μ0⟩/∥μ1−μ0∥∼𝒩​\(0,1\)\\left\\langle\\mathbf\{y\}\_\{0\}\-\\mu\_\{0\},\\mu\_\{1\}\-\\mu\_\{0\}\\right\\rangle/\\lVert\\mu\_\{1\}\-\\mu\_\{0\}\\rVert\\sim\\mathcal\{N\}\(0,1\)in any dimension\. For full details, see[SectionD\.2](https://arxiv.org/html/2607.15485#A4.SS2)\. ∎

##### Effect of parameters on DSSI\.

The DSSIL​\(γ∗\)L\(\\gamma^\{\*\}\)depends on the mode separation∥μ1−μ0∥\\lVert\\mu\_\{1\}\-\\mu\_\{0\}\\rVert, true mixture weightγ∗\\gamma^\{\*\}, and choice of noise schedule\{αt\}t\\\{\\alpha\_\{t\}\\\}\_\{t\}\. First, consistent with the fact thatℓDSM​\(p\(γ\),p\(γ∗\)\)\\ell\_\{\\mathrm\{DSM\}\}\(p^\{\(\\gamma\)\},p^\{\(\\gamma^\{\*\}\)\}\)is an increasing function of the separation \([Figure2\(d\)](https://arxiv.org/html/2607.15485#S3.F2.sf4)\),[Figure3\(a\)](https://arxiv.org/html/2607.15485#S4.F3.sf1)showsL​\(γ∗\)L\(\\gamma^\{\*\}\)is an increasing function of the mode separation\. Second, although[Theorem4\.1](https://arxiv.org/html/2607.15485#S4.Thmtheorem1)is only stated for linear noise schedules, different continuous noise schedules are equivalent up to time\-reparameterization, guaranteeingℓSM​\(pt\(γ^\),pt\(γ∗\)\)\\ell\_\{\\mathrm\{SM\}\}\(p\_\{t\}^\{\(\\hat\{\\gamma\}\)\},p\_\{t\}^\{\(\\gamma^\{\*\}\)\}\)is nonnegligible for some range ofttunder any noise schedule \([Figure3\(b\)](https://arxiv.org/html/2607.15485#S4.F3.sf2)\)\. Note however thatL​\(γ∗\)L\(\\gamma^\{\*\}\)is defined relative toℓDSM\\ell\_\{\\mathrm\{DSM\}\}, which averagesℓSM​\(pt\(γ^\),pt\(γ∗\)\)\\ell\_\{\\mathrm\{SM\}\}\(p\_\{t\}^\{\(\\hat\{\\gamma\}\)\},p\_\{t\}^\{\(\\gamma^\{\*\}\)\}\)overt∈\[0,1\]t\\in\[0,1\], meaning a noise schedule which reduces this range ofttto a small interval\[a,a\+ϵ\]\[a,a\+\\epsilon\]scales downL​\(γ∗\)L\(\\gamma^\{\*\}\)roughly proportional toϵ\\epsilon, regardless of the peak value ofℓSM​\(pt\(γ^\),pt\(γ∗\)\)\\ell\_\{\\mathrm\{SM\}\}\(p\_\{t\}^\{\(\\hat\{\\gamma\}\)\},p\_\{t\}^\{\(\\gamma^\{\*\}\)\}\)\. Third,[Figure3\(c\)](https://arxiv.org/html/2607.15485#S4.F3.sf3)shows that for mode separation∥μ1−μ0∥=9\\lVert\\mu\_\{1\}\-\\mu\_\{0\}\\rVert=9,L​\(γ∗\)L\(\\gamma^\{\*\}\)is bounded below by 0\.37 across allγ∗\\gamma^\{\*\}, implying that the squared prediction error inγ\\gammais at most≈2\.7\\approx 2\.7times larger than the squared diffusion score error\. Notably, this bound is dimension\-independent, and hence this holds for bimodal isotropic Gaussian mixtures in any dimension\. In[SectionE\.4](https://arxiv.org/html/2607.15485#A5.SS4), we showcase additional experiments on parameters of Gaussian mixtures other than the mixture weight\.

![Refer to caption](https://arxiv.org/html/2607.15485v1/x6.png)\(a\)
![Refer to caption](https://arxiv.org/html/2607.15485v1/x7.png)\(b\)
![Refer to caption](https://arxiv.org/html/2607.15485v1/x8.png)\(c\)

Figure 3:Effect of Gaussian mixture model parameters on score sensitivity in dimensiond=10d=10\. \(a\)L​\(γ∗\)L\(\\gamma^\{\*\}\)vs\. mode separation for multipleγ∗≤0\.5\\gamma^\{\*\}\\leq 0\.5\(by symmetryL​\(γ∗\)=L​\(1−γ∗\)L\(\\gamma^\{\*\}\)=L\(1\-\\gamma^\{\*\}\)\)\.L​\(γ∗\)L\(\\gamma^\{\*\}\)is an increasing function of separation\. \(b\)ℓSM​\(pt\(γ\),pt\(γ∗\)\)\\ell\_\{\\mathrm\{SM\}\}\(p\_\{t\}^\{\(\\gamma\)\},p\_\{t\}^\{\(\\gamma^\{\*\}\)\}\)as a function ofttandγ\\gammafor a cosine noise schedule\. Compared to a linear noise schedule, the noise schedule only affects the range ofttin the forward process at which the score is sensitive toγ\\gamma\. \(c\)ℓDSM​\(p\(γ\),p\(γ∗\)\)/\(γ−γ∗\)2\\ell\_\{\\mathrm\{DSM\}\}\(p^\{\(\\gamma\)\},p^\{\(\\gamma^\{\*\}\)\}\)/\(\\gamma\-\\gamma^\{\*\}\)^\{2\}vs\.γ\\gamma;L​\(γ∗\)≥0\.27L\(\\gamma^\{\*\}\)\\geq 0\.27for allγ∗\\gamma^\{\*\}, with the minimum atγ∗=0\.5\\gamma^\{\*\}=0\.5\.
### 4\.1Related Work

##### Multimodal capabilities of diffusion models\.

The ability of diffusion models to sample multimodal targets varies substantially across common choices of neural architecture\[[18](https://arxiv.org/html/2607.15485#bib.bib10)\], training recipes\[[25](https://arxiv.org/html/2607.15485#bib.bib5)\], and sampling algorithms\[[47](https://arxiv.org/html/2607.15485#bib.bib8),[63](https://arxiv.org/html/2607.15485#bib.bib58)\]\. Existing evaluation metrics\[[48](https://arxiv.org/html/2607.15485#bib.bib7),[28](https://arxiv.org/html/2607.15485#bib.bib69),[14](https://arxiv.org/html/2607.15485#bib.bib70),[39](https://arxiv.org/html/2607.15485#bib.bib71),[2](https://arxiv.org/html/2607.15485#bib.bib72)\]are typically structure\-agnostic and may fail to reliably detect mode coverage\[[45](https://arxiv.org/html/2607.15485#bib.bib73),[55](https://arxiv.org/html/2607.15485#bib.bib74)\]\. In contrast, the diffusion score sensitivity index \([1](https://arxiv.org/html/2607.15485#S1.E1)\) provides a rigorous framework to reason about multimodal capability\.

##### Learning GMMs with Diffusion Score Matching\.

Diffusion score matching \(DSM\) isin principleexpressive enough to learn Gaussian mixture: gradient descent recovers component means in the balanced isotropic regime444The mixture modelq≔∑i=1k1k​𝒩​\(𝝁\(i\),𝐈\)q\\coloneqq\\sum\_\{i=1\}^\{k\}\\frac\{1\}\{k\}\\,\\mathcal\{N\}\(\\bm\{\\mu\}^\{\(i\)\},\\,\\mathbf\{I\}\)where each component has equal mixture weight and identity covariance\[[49](https://arxiv.org/html/2607.15485#bib.bib34)\], learning algorithms can be tailored for density estimation of GMMs with polynomial sample complexity\[[16](https://arxiv.org/html/2607.15485#bib.bib36),[13](https://arxiv.org/html/2607.15485#bib.bib35)\]\. Here,Gatmiryet al\.\[[16](https://arxiv.org/html/2607.15485#bib.bib36)\]develop noise\-sensitivity bounds similar in spirit to ours\. Unlike the aforementioned articles, our work focuses on the ability of diffusion models to sample from target GMMs rather than on density estimation\. When the learned score isL2L^\{2\}\-accurate, the resulting diffusion sampler is provably close to the target in total variation\[[12](https://arxiv.org/html/2607.15485#bib.bib37),[32](https://arxiv.org/html/2607.15485#bib.bib38)\]\. However, this may be insufficient to recover structural parameters that shape the geometry of distributions in low\-probability regions\.Liet al\.\[[34](https://arxiv.org/html/2607.15485#bib.bib28)\]suggest diffusion models may fail to sample multimodal targets due to the lack of expressivity of the Gaussian transition steps in the reverse diffusion process\. Our work complementsLiet al\.\[[34](https://arxiv.org/html/2607.15485#bib.bib28)\]by identifying a different mechanism for multimodal sampling, the diffusion score sensitivity index \([1](https://arxiv.org/html/2607.15485#S1.E1)\), and by providing additional insights into the role of noise schedules in alleviating mode amplification\.

##### Benefits of Annealing\.

Classical \(unannealed\) score matching\[[24](https://arxiv.org/html/2607.15485#bib.bib65)\]is unsuitable for generative modeling of multimodal targets: the asymptotic efficiency of the score matching estimator relative to the maximum\-likelihood estimator is governed by the log\-Sobolev \(LS\) constant\[[27](https://arxiv.org/html/2607.15485#bib.bib27)\], which grows exponentially in mode separation for Gaussian mixtures\[[11](https://arxiv.org/html/2607.15485#bib.bib26)\]\.Song and Ermon \[[52](https://arxiv.org/html/2607.15485#bib.bib61)\]proposed annealing over noise scale to handle multimodality, inspired byNeal \[[40](https://arxiv.org/html/2607.15485#bib.bib50)\]\.Biroliet al\.\[[6](https://arxiv.org/html/2607.15485#bib.bib41)\]andRaya and Ambrogioni \[[46](https://arxiv.org/html/2607.15485#bib.bib42)\]observe a phase transition at intermediate noise scales linked to the emergence of multimodal structure in practice\. Amidst the empirical evidence,Qin and Risteski \[[44](https://arxiv.org/html/2607.15485#bib.bib20)\]establishes the first formal statistical benefit of annealing in the context of diffusion\.Qin and Risteski \[[44](https://arxiv.org/html/2607.15485#bib.bib20)\]adapts techniques from the acceleration of Markov Chain mixing to derive a better score matching loss, with the resulting outcome closely resembling the diffusion score matching loss\. WhileQin and Risteski \[[44](https://arxiv.org/html/2607.15485#bib.bib20)\]establishes sample complexity bounds, our work addresses a complementary question of parameter recovery from generated samples\.

## 5Sensitivity in real datasets

We empirically show that mixture weight information is captured in the score functions of intermediate forward process distributions, even outside the Gaussian mixture setting\. We additionally explore how altering the noise schedule affects the DSSI and empirically demonstrate that a small DSSI value quantitatively degrades mixture\-weight recovery on non\-Gaussian data\. This is particularly salient to methods like\[[51](https://arxiv.org/html/2607.15485#bib.bib43),[36](https://arxiv.org/html/2607.15485#bib.bib44),[25](https://arxiv.org/html/2607.15485#bib.bib5)\]where noise schedules are selected to increase the speed of inference\-time sampling; we show that such methods have the potential to amplify some modes over others, augmenting the recent observation byZhanget al\.\[[63](https://arxiv.org/html/2607.15485#bib.bib58)\]\.

To study this, we train a simple autoencoder on the ones and eights from MNIST \(\[[31](https://arxiv.org/html/2607.15485#bib.bib6)\]\)\. Our target distributionp\(γ∗\)p^\{\(\\gamma^\{\*\}\)\}is then the mixture distribution of latent representations of eights and ones \(p\(1\)p^\{\(1\)\}is the distribution of ones\)\. We setγ∗=0\.4\\gamma^\{\*\}=0\.4and use deep ResNets\[[21](https://arxiv.org/html/2607.15485#bib.bib56)\]with a linear noise schedule as our score matching models\.

![Refer to caption](https://arxiv.org/html/2607.15485v1/x9.png)
![Refer to caption](https://arxiv.org/html/2607.15485v1/x10.png)
![Refer to caption](https://arxiv.org/html/2607.15485v1/x11.png)
![Refer to caption](https://arxiv.org/html/2607.15485v1/x12.png)\(a\)
![Refer to caption](https://arxiv.org/html/2607.15485v1/x13.png)\(b\)
![Refer to caption](https://arxiv.org/html/2607.15485v1/x14.png)\(c\)

Figure 4:Influence of noise schedule on mixture weight recovery in MNIST\. \(a\) Linear noise schedule; \(b\-c\) two synthetic noise schedules designed to reduceL​\(γ∗\)L\(\\gamma^\{\*\}\)\(see[SectionE\.1](https://arxiv.org/html/2607.15485#A5.SS1)\)\. Top row:ℓSM​\(pt\(γ\),pt\(γ∗\)\)\\ell\_\{\\mathrm\{SM\}\}\(p\_\{t\}^\{\(\\gamma\)\},p\_\{t\}^\{\(\\gamma^\{\*\}\)\}\)as a function ofttandγ\\gammaestimated using trained diffusion models, withγ∗=0\.4\\gamma^\{\*\}=0\.4\. Bottom row: generated samples using each schedule, with estimatedγ^\\hat\{\\gamma\}below\. LargerL​\(γ∗\)L\(\\gamma^\{\*\}\)yields more accurate mixture weight recovery\.##### Sensitivity at intermediatett\.

In[Figure4](https://arxiv.org/html/2607.15485#S5.F4), the top row displays the score sensitivity over time andγ\\gammausing trained diffusion models \(each with a different noise schedule\) to approximate the true score function \(see[SectionE\.3](https://arxiv.org/html/2607.15485#A5.SS3)for more details\)\. The first column uses the same linear noise schedule as in[Figure2\(a\)](https://arxiv.org/html/2607.15485#S3.F2.sf1)\. Note that relative to the Gaussian mixture example,ℓSM​\(pt\(γ\),pt\(γ∗\)\)\\ell\_\{\\mathrm\{SM\}\}\(p\_\{t\}^\{\(\\gamma\)\},p\_\{t\}^\{\(\\gamma^\{\*\}\)\}\)is nonnegligable over a wider range ofttvalues\. We numerically estimate that the sensitivity index is≈5×\\approx 5\\timeslarger than in the Gaussian mixture model example, so our theory guarantees a five\-fold smaller error in the predicted mixture weight for the same level of error underℓDSM\\ell\_\{\\mathrm\{DSM\}\}as in the Gaussian setting\.

##### Sensitivity affects parameter recovery\.

In[Figure4\(a\)](https://arxiv.org/html/2607.15485#S5.F4.sf1), the bottom row shows generated samples from a diffusion model withT=100T=100\. We estimate the value ofγ∗\\gamma^\{\*\}using generated samples and a trained classifier, shown below the generated samples\. The generated samples accurately reflect the true value of the mixture weight, with only a1%1\\%relative error in estimatingγ∗\\gamma^\{\*\}\.

To demonstrate the effect of the score sensitivity indexL​\(γ∗\)L\(\\gamma^\{\*\}\)on the error in estimatingγ∗\\gamma^\{\*\}, we perform an ablation study on the trained diffusion model by modifying the noise schedule to reduceL​\(γ∗\)L\(\\gamma^\{\*\}\)as in[Section3\.1](https://arxiv.org/html/2607.15485#S3.SS1)\(for details see[SectionE\.1](https://arxiv.org/html/2607.15485#A5.SS1)\)\.[Figure4\(b\)](https://arxiv.org/html/2607.15485#S5.F4.sf2)and[Figure4\(c\)](https://arxiv.org/html/2607.15485#S5.F4.sf3)show the score sensitivity and generated samples for the modified noise schedules, along with the corresponding estimatesγ^\\hat\{\\gamma\}from the samples\. When the score sensitivity index reduces by a factor of1010\([Figure4\(c\)](https://arxiv.org/html/2607.15485#S5.F4.sf3)\), we see that the estimate ofγ∗\\gamma^\{\*\}has a relative error of31%31\\%\. Surprisingly, the samples from the modified noise schedules have comparable fidelity to those from the linear noise schedule \([Figure4\(a\)](https://arxiv.org/html/2607.15485#S5.F4.sf1)\)\. This is in line with the theoretical and empirical observations that say that diffusion models under errors and accelerated sampling schemes can still accurately learn the geometry of the support of the target distribution\[[51](https://arxiv.org/html/2607.15485#bib.bib43),[41](https://arxiv.org/html/2607.15485#bib.bib29),[54](https://arxiv.org/html/2607.15485#bib.bib68),[10](https://arxiv.org/html/2607.15485#bib.bib64)\]\. More details on our MNIST experiments can be found in[SectionE\.3](https://arxiv.org/html/2607.15485#A5.SS3)\.

## 6Conclusion

We investigate the ability of diffusion models to recover the relative mode amplitudes of a multimodal target, despite the insensitivity of the score of the target distribution to mixture weights under high separation\. Our analysis clarifies the impact of annealing: the score of the noisy target distribution across noise scales provides the necessary information for accurate sampling of multimodal distributions\. Our experiments indicate that inference\-time sampling algorithms can be tuned for diversity and fidelity, suggesting a mechanism for tailoring sampling algorithms to data distributions arising in scientific applications\.

##### Limitations\.

We prove a quantitative relationship between the sensitivity index and the diffusion score error only in the Gaussian mixture setting\. Theorem[3\.5](https://arxiv.org/html/2607.15485#S3.Thmtheorem5)provides an information theoretic mechanism that cannot be computationally verified since it involves minimax bounds\. Further, evaluation of the diffusion score\-sensitivity index requires knowledge of the true parameter of interestγ∗\\gamma^\{\*\}and the score function of the target distribution𝐬t\(γ∗\)\\mathbf\{s\}\_\{t\}^\{\(\\gamma^\{\*\}\)\}\. Thus, the application to benchmark datasets requires surrogate generative models in place of the unknown target score\.

## References

- \[1\]H\. Addison, E\. Kendon, S\. Ravuri, L\. Aitchison, and P\. A\. G\. Watson\(2026\)Machine learning emulation of precipitation from km\-scale UK regional climate simulations using a diffusion model\.Journal of Advances in Modeling Earth Systems\.Note:arXiv:2407\.14158External Links:[Document](https://dx.doi.org/10.1029/2025MS005140)Cited by:[§1](https://arxiv.org/html/2607.15485#S1.p1.1)\.
- \[2\]A\. M\. Alaa, B\. van Breugel, E\. S\. Saveliev, and M\. van der Schaar\(2022\)How faithful is your synthetic data? sample\-level metrics for evaluating and auditing generative models\.InProceedings of the 39th International Conference on Machine Learning,Proceedings of Machine Learning Research,pp\. 290–306\.Note:arXiv:2102\.08921Cited by:[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px1.p1.1)\.
- \[3\]D\. Alspach and H\. Sorenson\(1972\)Nonlinear bayesian estimation using gaussian sum approximations\.IEEE Transactions on Automatic Control17\(4\),pp\. 439–448\.External Links:[Document](https://dx.doi.org/10.1109/TAC.1972.1100034)Cited by:[§1](https://arxiv.org/html/2607.15485#S1.p1.1)\.
- \[4\]B\. D\.O\. Anderson\(1982\)Reverse\-time diffusion equation models\.Stochastic Processes and their Applications12\(3\),pp\. 313–326\.External Links:ISSN 0304\-4149,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/0304-4149%2882%2990051-5),[Link](https://www.sciencedirect.com/science/article/pii/0304414982900515)Cited by:[§2](https://arxiv.org/html/2607.15485#S2.SS0.SSS0.Px2.p1.11)\.
- \[5\]L\. Biewald\(2020\)Experiment tracking with weights and biases\.Note:Software available from wandb\.comExternal Links:[Link](https://www.wandb.com/)Cited by:[§E\.2](https://arxiv.org/html/2607.15485#A5.SS2.p2.1)\.
- \[6\]G\. Biroli, T\. Bonnaire, V\. De Bortoli, and M\. Mézard\(2024\)Dynamical regimes of diffusion models\.Nature Communications15,pp\. 9957\.External Links:[Document](https://dx.doi.org/10.1038/s41467-024-54281-3)Cited by:[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px3.p1.1)\.
- \[7\]C\. M\. Bishop\(2006\)Pattern recognition and machine learning \(information science and statistics\)\.Springer\-Verlag,Berlin, Heidelberg\.External Links:ISBN 0387310738Cited by:[§D\.1\.1](https://arxiv.org/html/2607.15485#A4.SS1.SSS1.p1.4)\.
- \[8\]F\. Bouchet, J\. Rolland, and J\. Wouters\(2019\)Rare event sampling methods\.Chaos: An Interdisciplinary Journal of Nonlinear Science29\(8\),pp\. 080402\.External Links:[Document](https://dx.doi.org/10.1063/1.5120509)Cited by:[§1](https://arxiv.org/html/2607.15485#S1.p1.1)\.
- \[9\]J\. Bradbury, R\. Frostig, P\. Hawkins, M\. J\. Johnson, Y\. Katariya, C\. Leary, D\. Maclaurin, G\. Necula, A\. Paszke, J\. VanderPlas, S\. Wanderman\-Milne, and Q\. Zhang\(2018\)JAX: composable transformations of Python\+NumPy programs\.External Links:[Link](http://github.com/jax-ml/jax)Cited by:[§E\.2](https://arxiv.org/html/2607.15485#A5.SS2.p2.1)\.
- \[10\]N\. Chandramoorthy and A\. A\. de Clercq\(2026\)When and how can inexact generative models still sample from the data manifold?\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=QBjuYL4gAX)Cited by:[§5](https://arxiv.org/html/2607.15485#S5.SS0.SSS0.Px2.p2.7)\.
- \[11\]H\. Chen, S\. Chewi, and J\. Niles\-Weed\(2021\)Dimension\-free log\-sobolev inequalities for mixture distributions\.Vol\.281\.External Links:ISSN 0022\-1236,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.jfa.2021.109236),[Link](https://www.sciencedirect.com/science/article/pii/S0022123621003189)Cited by:[§1](https://arxiv.org/html/2607.15485#S1.p1.1),[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px3.p1.1)\.
- \[12\]S\. Chen, S\. Chewi, J\. Li, Y\. Li, A\. Salim, and A\. Zhang\(2023\)Sampling is as easy as learning the score: theory for diffusion models with minimal data assumptions\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=zyLVMgsZ0U_)Cited by:[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px2.p1.1)\.
- \[13\]S\. Chen, V\. Kontonis, and K\. Shah\(2025\-30 Jun–04 Jul\)Learning general gaussian mixtures with efficient score matching\.InProceedings of Thirty Eighth Conference on Learning Theory,N\. Haghtalab and A\. Moitra \(Eds\.\),Proceedings of Machine Learning Research, Vol\.291,pp\. 1029–1090\.External Links:[Link](https://proceedings.mlr.press/v291/chen25e.html)Cited by:[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px2.p1.1)\.
- \[14\]J\. Djolonga, M\. Lucic, M\. Cuturi, O\. Bachem, O\. Bousquet, and S\. Gelly\(2020\)Precision\-recall curves using information divergence frontiers\.InProceedings of the 23rd International Conference on Artificial Intelligence and Statistics,Proceedings of Machine Learning Research, Vol\.108,pp\. 2550–2559\.Cited by:[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px1.p1.1)\.
- \[15\]L\. J\. Fahrner, E\. Chen, E\. Topol, and P\. Rajpurkar\(2025\)The generative era of medical ai\.Cell188\(14\),pp\. 3648–3660\.Cited by:[§1](https://arxiv.org/html/2607.15485#S1.p1.1)\.
- \[16\]K\. Gatmiry, J\. Kelner, and H\. Lee\(2025\-30 Jun–04 Jul\)Learning mixtures of gaussians using diffusion models\.InProceedings of Thirty Eighth Conference on Learning Theory,N\. Haghtalab and A\. Moitra \(Eds\.\),Proceedings of Machine Learning Research, Vol\.291,pp\. 2403–2456\.External Links:[Link](https://proceedings.mlr.press/v291/gatmiry25b.html)Cited by:[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px2.p1.1)\.
- \[17\]S\. A\. Geer\(2000\)Empirical processes in m\-estimation\.Vol\.6,Cambridge university press\.Cited by:[§3](https://arxiv.org/html/2607.15485#S3.SS0.SSS0.Px4.p1.1)\.
- \[18\]S\. Hakemi, N\. Akhtar, G\. M\. Hassan, and A\. Mian\(2025\)Deeper diffusion models amplify bias\.arXiv preprint arXiv:2505\.17560\.Cited by:[§1](https://arxiv.org/html/2607.15485#S1.p1.1),[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px1.p1.1)\.
- \[19\]M\. Han, J\. Devany, M\. Fruchart, M\. L\. Gardel, and V\. Vitelli\(2025\)Learning noisy tissue dynamics across time scales\.External Links:2510\.19090,[Link](https://arxiv.org/abs/2510.19090)Cited by:[§1](https://arxiv.org/html/2607.15485#S1.p1.1)\.
- \[20\]T\. Hang, S\. Gu, X\. Geng, and B\. Guo\(2024\)Improved noise schedule for diffusion training\.External Links:2407\.03297,[Link](https://arxiv.org/abs/2407.03297)Cited by:[§E\.1](https://arxiv.org/html/2607.15485#A5.SS1.p1.3)\.
- \[21\]K\. He, X\. Zhang, S\. Ren, and J\. Sun\(2015\)Deep residual learning for image recognition\.2016 IEEE Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 770–778\.External Links:[Link](https://api.semanticscholar.org/CorpusID:206594692)Cited by:[§E\.2](https://arxiv.org/html/2607.15485#A5.SS2.p1.6),[§5](https://arxiv.org/html/2607.15485#S5.p2.3)\.
- \[22\]M\. Heusel, H\. Ramsauer, T\. Unterthiner, B\. Nessler, and S\. Hochreiter\(2017\)GANs trained by a two time\-scale update rule converge to a local nash equilibrium\.InAdvances in Neural Information Processing Systems,I\. Guyon, U\. V\. Luxburg, S\. Bengio, H\. Wallach, R\. Fergus, S\. Vishwanathan, and R\. Garnett \(Eds\.\),Vol\.30,pp\.\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2017/file/8a1d694707eb0fefe65871369074926d-Paper.pdf)Cited by:[§1](https://arxiv.org/html/2607.15485#S1.p1.1)\.
- \[23\]J\. D\. Hunter\(2007\)Matplotlib: a 2d graphics environment\.Computing in Science & Engineering9\(3\),pp\. 90–95\.External Links:[Document](https://dx.doi.org/10.1109/MCSE.2007.55)Cited by:[§E\.2](https://arxiv.org/html/2607.15485#A5.SS2.p2.1)\.
- \[24\]A\. Hyvärinen\(2005\-12\)Estimation of non\-normalized statistical models by score matching\.J\. Mach\. Learn\. Res\.6,pp\. 695–709\.External Links:ISSN 1532\-4435Cited by:[§3](https://arxiv.org/html/2607.15485#S3.SS0.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px3.p1.1)\.
- \[25\]T\. Karras, M\. Aittala, T\. Aila, and S\. Laine\(2022\)Elucidating the design space of diffusion\-based generative models\.Advances in neural information processing systems35,pp\. 26565–26577\.Cited by:[§2](https://arxiv.org/html/2607.15485#S2.SS0.SSS0.Px2.p2.2),[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2607.15485#S5.p1.1)\.
- \[26\]P\. Kidger and C\. Garcia\(2021\)Equinox: neural networks in JAX via callable PyTrees and filtered transformations\.Differentiable Programming workshop at Neural Information Processing Systems 2021\.Cited by:[§E\.2](https://arxiv.org/html/2607.15485#A5.SS2.p2.1),[§E\.3](https://arxiv.org/html/2607.15485#A5.SS3.p1.5)\.
- \[27\]F\. Koehler, A\. Heckett, and A\. Risteski\(2023\)Statistical efficiency of score matching: the view from isoperimetry\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=TD7AnQjNzR6)Cited by:[Appendix C](https://arxiv.org/html/2607.15485#A3.p7.9),[§1](https://arxiv.org/html/2607.15485#S1.p1.1),[§1](https://arxiv.org/html/2607.15485#S1.p2.1),[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px3.p1.1)\.
- \[28\]T\. Kynkäänniemi, T\. Karras, S\. Laine, J\. Lehtinen, and T\. Aila\(2019\)Improved precision and recall metric for assessing generative models\.InAdvances in Neural Information Processing Systems,Vol\.32\.Cited by:[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px1.p1.1)\.
- \[29\]F\. Lanusse, R\. Mandelbaum, S\. Ravanbakhsh, C\. Li, P\. Freeman, and B\. Póczos\(2021\)Deep generative models for galaxy image simulations\.Monthly Notices of the Royal Astronomical Society504\(4\),pp\. 5543–5555\.Cited by:[§1](https://arxiv.org/html/2607.15485#S1.p1.1)\.
- \[30\]K\. Łatuszyński, M\. T\. Moores, and T\. Stumpf\-Fétizon\(2025\)MCMC for multi\-modal distributions\.arXiv preprint arXiv:2501\.05908\.Cited by:[Appendix B](https://arxiv.org/html/2607.15485#A2.p1.1)\.
- \[31\]Y\. LeCun, C\. Cortes, and C\. Burges\(2010\)MNIST handwritten digit database\.ATT Labs \[Online\]\. Available: http://yann\.lecun\.com/exdb/mnist2\.Cited by:[§5](https://arxiv.org/html/2607.15485#S5.p2.3)\.
- \[32\]H\. Lee, J\. Lu, and Y\. Tan\(2023\)Convergence of score\-based generative modeling for general data distributions\.InProceedings of the 34th International Conference on Algorithmic Learning Theory,Proceedings of Machine Learning Research, Vol\.201,pp\. 946–985\.Note:arXiv:2209\.12381Cited by:[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px2.p1.1)\.
- \[33\]H\. Lee, A\. Risteski, and R\. Ge\(2018\)Beyond log\-concavity: provable guarantees for sampling multi\-modal distributions using simulated tempering langevin monte carlo\.InAdvances in Neural Information Processing Systems,S\. Bengio, H\. Wallach, H\. Larochelle, K\. Grauman, N\. Cesa\-Bianchi, and R\. Garnett \(Eds\.\),Vol\.31,pp\.\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2018/file/c6ede20e6f597abf4b3f6bb30cee16c7-Paper.pdf)Cited by:[Appendix B](https://arxiv.org/html/2607.15485#A2.p1.1)\.
- \[34\]Y\. Li, B\. van Breugel, and M\. van der Schaar\(2024\)Soft mixture denoising: beyond the expressive bottleneck of diffusion models\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=aaBnFAyW9O)Cited by:[§1](https://arxiv.org/html/2607.15485#S1.p1.1),[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px2.p1.1)\.
- \[35\]M\. Loni, F\. Poursalim, M\. Asadi, and A\. Gharehbaghi\(2025\)A review on generative ai models for synthetic medical text, time series, and longitudinal data\.npj Digital Medicine8\(1\),pp\. 281\.Cited by:[§1](https://arxiv.org/html/2607.15485#S1.p1.1)\.
- \[36\]C\. Lu, Y\. Zhou, F\. Bao, J\. Chen, C\. Li, and J\. Zhu\(2022\)DPM\-solver: a fast ODE solver for diffusion probabilistic model sampling in around 10 steps\.InAdvances in Neural Information Processing Systems,Vol\.35\.Note:arXiv:2206\.00927Cited by:[§5](https://arxiv.org/html/2607.15485#S5.p1.1)\.
- \[37\]J\. Ludwig and S\. Mullainathan\(2024\)Machine learning as a tool for hypothesis generation\.The Quarterly Journal of Economics139\(2\),pp\. 751–827\.Cited by:[§1](https://arxiv.org/html/2607.15485#S1.p1.1)\.
- \[38\]N\. Mudur, C\. Cuesta\-Lazaro, and D\. P\. Finkbeiner\(2023\)Cosmological field emulation and parameter inference with diffusion models\.InNeurIPS Workshop on Machine Learning and the Physical Sciences,Note:arXiv:2312\.07534Cited by:[§1](https://arxiv.org/html/2607.15485#S1.p1.1)\.
- \[39\]M\. F\. Naeem, S\. J\. Oh, Y\. Uh, Y\. Choi, and J\. Yoo\(2020\)Reliable fidelity and diversity metrics for generative models\.InProceedings of the 37th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.119,pp\. 7176–7185\.Cited by:[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px1.p1.1)\.
- \[40\]R\. M\. Neal\(2001\-04\)Annealed importance sampling\.Statistics and Computing11\(2\),pp\. 125–139\.External Links:ISSN 1573\-1375,[Link](https://doi.org/10.1023/A:1008923215028),[Document](https://dx.doi.org/10.1023/A%3A1008923215028)Cited by:[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px3.p1.1)\.
- \[41\]A\. Q\. Nichol and P\. Dhariwal\(2021\-18–24 Jul\)Improved denoising diffusion probabilistic models\.InProceedings of the 38th International Conference on Machine Learning,M\. Meila and T\. Zhang \(Eds\.\),Proceedings of Machine Learning Research, Vol\.139,pp\. 8162–8171\.External Links:[Link](https://proceedings.mlr.press/v139/nichol21a.html)Cited by:[§5](https://arxiv.org/html/2607.15485#S5.SS0.SSS0.Px2.p2.7)\.
- \[42\]E\. A\. Posner and S\. Saran\(2026\)Judge ai: a case\-study of large language models as judges\.Journal of Law & Empirical Analysis0\(0\),pp\. 2755323X261433614\.External Links:[Document](https://dx.doi.org/10.1177/2755323X261433614),[Link](https://doi.org/10.1177/2755323X261433614),https://doi\.org/10\.1177/2755323X261433614Cited by:[§1](https://arxiv.org/html/2607.15485#S1.p1.1)\.
- \[43\]I\. Price, A\. Sanchez\-Gonzalez, F\. Alet, T\. R\. Andersson, A\. El\-Kadi, D\. Masters, T\. Ewalds, J\. Stott, S\. Mohamed, P\. Battaglia, R\. Lam, and M\. Willson\(2024\-12\)Probabilistic weather forecasting with machine learning\.Nature\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1038/s41586-024-08252-9)Cited by:[§1](https://arxiv.org/html/2607.15485#S1.p1.1)\.
- \[44\]Y\. Qin and A\. Risteski\(2024\)Fit like you sample: sample\-efficient generalized score matching from fast mixing diffusions\.InThe Thirty Seventh Annual Conference on Learning Theory,pp\. 4413–4457\.Cited by:[Appendix C](https://arxiv.org/html/2607.15485#A3.p7.9),[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px3.p1.1)\.
- \[45\]O\. Räisä, B\. van Breugel, and M\. van der Schaar\(2025\)Position: all current generative fidelity and diversity metrics are flawed\.InForty\-second International Conference on Machine Learning Position Paper Track,External Links:[Link](https://openreview.net/forum?id=DMRrbb36r5)Cited by:[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px1.p1.1)\.
- \[46\]G\. Raya and L\. Ambrogioni\(2023\)Spontaneous symmetry breaking in generative diffusion models\.InAdvances in Neural Information Processing Systems,Vol\.36\.External Links:[Link](https://arxiv.org/abs/2305.19693)Cited by:[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px3.p1.1)\.
- \[47\]N\. Roos, E\. Iakovleva, A\. Gjergji, V\. P\. Pastore, and E\. Tartaglione\(2026\)How i met your bias: investigating bias amplification in diffusion models\.InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision,pp\. 5374–5383\.Cited by:[§1](https://arxiv.org/html/2607.15485#S1.p1.1),[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px1.p1.1)\.
- \[48\]M\. S\. Sajjadi, O\. Bachem, M\. Lucic, O\. Bousquet, and S\. Gelly\(2018\)Assessing generative models via precision and recall\.Advances in neural information processing systems31\.Cited by:[§1](https://arxiv.org/html/2607.15485#S1.p1.1),[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px1.p1.1)\.
- \[49\]K\. Shah, S\. Chen, and A\. Klivans\(2023\)Learning mixtures of gaussians using the DDPM objective\.InThirty\-seventh Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=aig7sgdRfI)Cited by:[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px2.p1.1)\.
- \[50\]J\. Song, C\. Meng, and S\. Ermon\(2021\)Denoising diffusion implicit models\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=St1giarCHLP)Cited by:[§3\.1](https://arxiv.org/html/2607.15485#S3.SS1.p3.1)\.
- \[51\]J\. Song, C\. Meng, and S\. Ermon\(2021\)Denoising diffusion implicit models\.InInternational Conference on Learning Representations,Note:arXiv:2010\.02502Cited by:[§5](https://arxiv.org/html/2607.15485#S5.SS0.SSS0.Px2.p2.7),[§5](https://arxiv.org/html/2607.15485#S5.p1.1)\.
- \[52\]Y\. Song and S\. Ermon\(2019\)Generative modeling by estimating gradients of the data distribution\.InAdvances in Neural Information Processing Systems,Vol\.32,pp\. 11895–11907\.Cited by:[§1](https://arxiv.org/html/2607.15485#S1.p2.1),[§3](https://arxiv.org/html/2607.15485#S3.SS0.SSS0.Px3.p1.12),[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px3.p1.1)\.
- \[53\]Y\. Song, J\. Sohl\-Dickstein, D\. P\. Kingma, A\. Kumar, S\. Ermon, and B\. Poole\(2021\)Score\-based generative modeling through stochastic differential equations\.External Links:[Link](https://openreview.net/forum?id=PxTIG12RRHS)Cited by:[§D\.2](https://arxiv.org/html/2607.15485#A4.SS2.p1.4),[§E\.2](https://arxiv.org/html/2607.15485#A5.SS2.p1.6),[§E\.2](https://arxiv.org/html/2607.15485#A5.SS2.p2.1),[§2](https://arxiv.org/html/2607.15485#S2.SS0.SSS0.Px1.p1.3),[§3](https://arxiv.org/html/2607.15485#S3.SS0.SSS0.Px3.p1.11),[§4](https://arxiv.org/html/2607.15485#S4.p2.12)\.
- \[54\]J\. Stanczuk, G\. Batzolis, T\. Deveney, and C\. Schönlieb\(2022\)Your diffusion model secretly knows the dimension of the data manifold\.arXiv preprint arXiv:2212\.12611\.Cited by:[§5](https://arxiv.org/html/2607.15485#S5.SS0.SSS0.Px2.p2.7)\.
- \[55\]G\. Stein, J\. C\. Cresswell, R\. Hosseinzadeh, Y\. Sui, B\. L\. Ross, V\. Villecroze, Z\. Liu, A\. L\. Caterini, E\. Taylor, and G\. Loaiza\-Ganem\(2023\)Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models\.InThirty\-seventh Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=08zf7kTOoh)Cited by:[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px1.p1.1)\.
- \[56\]Z\. L\. Teo, A\. J\. Thirunavukarasu, K\. Elangovan, H\. Cheng, P\. Moova, B\. Soetikno, C\. Nielsen, A\. Pollreisz, D\. S\. J\. Ting, R\. J\. Morris,et al\.\(2025\)Generative artificial intelligence in medicine\.Nature medicine31\(10\),pp\. 3270–3282\.Cited by:[§1](https://arxiv.org/html/2607.15485#S1.p1.1)\.
- \[57\]A\. B\. Tsybakov\(2008\)Nonparametric estimators\.InIntroduction to Nonparametric Estimation,pp\. 1–76\.Cited by:[Definition 3\.1](https://arxiv.org/html/2607.15485#S3.Thmtheorem1.p1.12.12)\.
- \[58\]P\. Vincent\(2011\)A connection between score matching and denoising autoencoders\.Neural Computation23\(7\),pp\. 1661–1674\.External Links:[Document](https://dx.doi.org/10.1162/NECO%5Fa%5F00142)Cited by:[§D\.1\.3](https://arxiv.org/html/2607.15485#A4.SS1.SSS3.2.p1.7)\.
- \[59\]P\. Virtanen, R\. Gommers, T\. E\. Oliphant, M\. Haberland, T\. Reddy, D\. Cournapeau, E\. Burovski, P\. Peterson, W\. Weckesser, J\. Bright, S\. J\. van der Walt, M\. Brett, J\. Wilson, K\. J\. Millman, N\. Mayorov, A\. R\. J\. Nelson, E\. Jones, R\. Kern, E\. Larson, C\. J\. Carey, İ\. Polat, Y\. Feng, E\. W\. Moore, J\. VanderPlas, D\. Laxalde, J\. Perktold, R\. Cimrman, I\. Henriksen, E\. A\. Quintero, C\. R\. Harris, A\. M\. Archibald, A\. H\. Ribeiro, F\. Pedregosa, P\. van Mulbregt, and SciPy 1\.0 Contributors\(2020\)SciPy 1\.0: Fundamental Algorithms for Scientific Computing in Python\.Nature Methods17,pp\. 261–272\.External Links:[Document](https://dx.doi.org/10.1038/s41592-019-0686-2)Cited by:[§E\.2](https://arxiv.org/html/2607.15485#A5.SS2.p2.1)\.
- \[60\]S\. J\. Yakowitz and J\. D\. Spragins\(1968\)On the identifiability of finite mixtures\.The Annals of Mathematical Statistics39\(1\),pp\. 209–214\.External Links:[Document](https://dx.doi.org/10.1214/aoms/1177698520)Cited by:[§1](https://arxiv.org/html/2607.15485#S1.p1.1)\.
- \[61\]Y\. Yang and A\. Barron\(1999\)Information\-theoretic determination of minimax rates of convergence\.The Annals of Statistics27\(5\),pp\. 1564–1599\.Cited by:[Appendix C](https://arxiv.org/html/2607.15485#A3.p7.9),[§3](https://arxiv.org/html/2607.15485#S3.SS0.SSS0.Px5.p1.4),[§3](https://arxiv.org/html/2607.15485#S3.SS0.SSS0.Px5.p3.1)\.
- \[62\]B\. Yu\(1997\)Assouad, fano, and le cam\.InFestschrift for Lucien Le Cam: Research Papers in Probability and Statistics,pp\. 423–435\.External Links:ISBN 978\-1\-4612\-1880\-7,[Document](https://dx.doi.org/10.1007/978-1-4612-1880-7%5F29),[Link](https://doi.org/10.1007/978-1-4612-1880-7_29)Cited by:[Appendix C](https://arxiv.org/html/2607.15485#A3.p2.5),[Appendix C](https://arxiv.org/html/2607.15485#A3.p4.1),[Appendix C](https://arxiv.org/html/2607.15485#A3.p5.2),[§3](https://arxiv.org/html/2607.15485#S3.SS0.SSS0.Px5.p1.4)\.
- \[63\]Y\. Zhang, Z\. Liao, J\. Wu, and D\. Zou\(2025\)On the collapse errors induced by the deterministic sampler for diffusion models\.External Links:2508\.16154,[Link](https://arxiv.org/abs/2508.16154)Cited by:[§1](https://arxiv.org/html/2607.15485#S1.p1.1),[§4\.1](https://arxiv.org/html/2607.15485#S4.SS1.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2607.15485#S5.p1.1)\.
- \[64\]J\. Zhuang, T\. Tang, Y\. Ding, S\. C\. Tatikonda, N\. Dvornek, X\. Papademetris, and J\. Duncan\(2020\)AdaBelief optimizer: adapting stepsizes by the belief in observed gradients\.Advances in Neural Information Processing Systems33\.Cited by:[§E\.2](https://arxiv.org/html/2607.15485#A5.SS2.p2.1)\.

## Appendix ANotation

## Appendix BComparison of sampling algorithms for multimodal distributions

Sampling from multimodal distributions using classical Langevin sampling is known to be challenging\[[33](https://arxiv.org/html/2607.15485#bib.bib67),[30](https://arxiv.org/html/2607.15485#bib.bib77)\]\. For a two\-component Gaussian mixture,[Figure2\(d\)](https://arxiv.org/html/2607.15485#S3.F2.sf4)shows that as mode separation increases, the sensitivity of the classical score matching loss to the mixture weight decays rapidly to 0 while the diffusion score matching loss does not\. As a consequence, sampling algorithms based solely on the target score are unsuitable for sampling from multimodal distributions\.[Figure5](https://arxiv.org/html/2607.15485#A2.F5)illustrates this effect, comparing classical Langevin sampling to diffusion sampling\.

![Refer to caption](https://arxiv.org/html/2607.15485v1/x15.png)\(a\)
![Refer to caption](https://arxiv.org/html/2607.15485v1/x16.png)\(b\)
![Refer to caption](https://arxiv.org/html/2607.15485v1/x17.png)\(c\)
![Refer to caption](https://arxiv.org/html/2607.15485v1/x18.png)\(d\)

Figure 5:Comparison between different sampling approaches on a high\-separation Gaussian mixture model\. \(a\) shows a KDE of the true data distribution\. \(b\)\-\(c\) show, in order, KDEs of samples using \(b\) Langevin sampling \(run to timeT=105T=10^\{5\}\), \(c\) diffusion sampling using the analytic score function, and \(d\) diffusion sampling using a trained neural network\. We can see that, unlike Langevin sampling, diffusion sampling accurately reflects the relative masses of the two modes\.
## Appendix CMinimax upper and lower bounds

Here we provide a more detailed proof of Theorem[3\.5](https://arxiv.org/html/2607.15485#S3.Thmtheorem5)and provide additional remarks on the Theorem and the underlying assumptions\. We begin with a remark on score\-restricted estimators \(Definition[3\.1](https://arxiv.org/html/2607.15485#S3.Thmtheorem1)\) and explain the motivation for their definition\.

In minimax theory, and in particular, Fano’s method, the accuracy of an estimator for a parameter learned from a finite number of data samples is inherently limited by how well the data samples can distinguish between two parameter values\. This distinguishability is often measured in terms of the KL divergence between the data distributions at two nearby parameter values\. At first glance, the data samples given to a diffusion model appear to be the target samples, and as noted previously, the data distributions with well\-separated modes do not distinguish between nearby values of weight parameters\. Further, noteKL\(pt\(γ\)\|\|pt\(γ′\)\)≤KL\(p\(γ\)\|\|p\(γ′\)\)\.\\mathrm\{KL\}\(p\_\{t\}^\{\(\\gamma\)\}\|\|p\_\{t\}^\{\(\\gamma^\{\\prime\}\)\}\)\\leq\\mathrm\{KL\}\(p^\{\(\\gamma\)\}\|\|p^\{\(\\gamma^\{\\prime\}\)\}\)\.As a result, it might seem that the generated samples from the denoising process do not accurately recover the parameter \(particularly, given the role that the KL divergence plays in Fano’s inequality\[[62](https://arxiv.org/html/2607.15485#bib.bib30)\]\), but our empirical results disagree with this expectation\. In our empirical results, we see that the noising/denoising process affects parameter recovery compared with score approximation followed by Langevin dynamics\. A key idea that helps resolve this mismatch is to recognize that the data a diffusion model uses to accurately estimate parameters is the noisy score \(approximated by a neural network\), not the noisy sample data\. In other words, the generated samples, from the reverse dynamics, are a function of the noisy scores \(interpolated by neural networks\)\. These neural network interpolations at some timettcan have a distribution that distinguishes between parameter values more than at time0\. In that sense, the timettscores contain sufficient information to generate samples corresponding to a nearby parameter value, while the time0score \(target score\) may not\.

Our score\-restricted estimators define the notion of recoveringγ∗\\gamma^\{\*\}from “score data” rather than sample data\. Such a notion models parameter estimation from the generated samples in score\-based generative models where the score field is first approximated and then used for sampling\. To distinguish Langevin sampling from diffusion models, we define two separate classes of estimators: one class has access to only a target score oracle and the other class has access to a score oracle at an intermediate time along the noising process\. Hence, the estimator classesΓscore,t\\Gamma\_\{\\mathrm\{score\},t\}andΓscore,0\\Gamma\_\{\\mathrm\{score\},0\}differ in the data distributions given to them, which are respectively,\(qt\(γ\)\)\(q\_\{t\}^\{\(\\gamma\)\}\)and\(q0\(γ\)\)\.\(q\_\{0\}^\{\(\\gamma\)\}\)\.In Lemma[C\.1](https://arxiv.org/html/2607.15485#A3.Thmtheorem1), we show that the oracle modelqt\(γ\)q\_\{t\}^\{\(\\gamma\)\}cannot be reconstructed in aγ\\gamma\-independent manner from the oracle mannerq0\(γ\),q\_\{0\}^\{\(\\gamma\)\},and this induces a separation between the two estimator classes\.

###### Lemma C\.1\.

There is noγ\\gamma\-independent post processing of score observations at time 0,𝐬^0\(γ\),\\hat\{\\mathbf\{s\}\}^\{\(\\gamma\)\}\_\{0\},that can reproduce𝐬^t\(γ\)\.\\hat\{\\mathbf\{s\}\}^\{\(\\gamma\)\}\_\{t\}\.

###### Proof\.

We can prove why by contradiction\. If there exists aγ\\gamma\-independent processing ofq\(γ\)q^\{\(\\gamma\)\}that produces the distributionqt\(γ\),q\_\{t\}^\{\(\\gamma\)\},the data processing inequality givesKL\(qt\(γ\)\|\|qt\(γ′\)\)≤KL\(q\(γ\)\|\|q\(γ′\)\)\.\\mathrm\{KL\}\(q\_\{t\}^\{\(\\gamma\)\}\|\|q\_\{t\}^\{\(\\gamma^\{\\prime\}\)\}\)\\leq\\mathrm\{KL\}\(q^\{\(\\gamma\)\}\|\|q^\{\(\\gamma^\{\\prime\}\)\}\)\.But, we have assumed that at timet,t,our mixture family is such thatKL\(q\(γ\)\|\|q\(γ′\)\)≤ϵt2/4≤\(1/128\)\(γ−γ′\)2/At2≤\(1/128\)KL\(qt\(γ\)\|\|qt\(γ′\)\)\.\\mathrm\{KL\}\(q^\{\(\\gamma\)\}\|\|q^\{\(\\gamma^\{\\prime\}\)\}\)\\leq\\epsilon\_\{t\}^\{2\}/4\\leq\(1/128\)\(\\gamma\-\\gamma^\{\\prime\}\)^\{2\}/A\_\{t\}^\{2\}\\leq\(1/128\)\\mathrm\{KL\}\(q\_\{t\}^\{\(\\gamma\)\}\|\|q\_\{t\}^\{\(\\gamma^\{\\prime\}\)\}\)\.Thus, the oracle observations and the statistical experiment that can be performed with the modelqt\(γ\)q\_\{t\}^\{\(\\gamma\)\}cannot be simulated from having access to the oracle,q\(γ\)\.q^\{\(\\gamma\)\}\.∎

Before we give the proof of Theorem[3\.5](https://arxiv.org/html/2607.15485#S3.Thmtheorem5), we recapitulate a corollary of Generalized Fano’s method from Lemma 3 of\[[62](https://arxiv.org/html/2607.15485#bib.bib30)\]and define covering entropy, which is used in Theorem[3\.5](https://arxiv.org/html/2607.15485#S3.Thmtheorem5)\.

###### Lemma C\.2\.

\[Lack of identifiability of mixture weights\] LetG:=\{γ1,⋯,γN\}∈\(0,1\)NG:=\\\{\\gamma\_\{1\},\\cdots,\\gamma\_\{N\}\\\}\\in\(0,1\)^\{N\}be a set ofN≥16N\\geq 16values of mixture weights such that\|γi−γj\|≥ε,i≠j,1≤i,j≤N\.\|\\gamma\_\{i\}\-\\gamma\_\{j\}\|\\geq\\varepsilon,\\\>\\\>i\\neq j,\\;1\\leq i,j\\leq N\.Then, for the class of score\-restricted estimators \([Definition3\.1](https://arxiv.org/html/2607.15485#S3.Thmtheorem1)\),Γscore,0,\\Gamma\_\{\\mathrm\{score\},0\},with oracle access to the model,q\(γ∗\)q^\{\(\\gamma^\{\*\}\)\}, ifKL\(q\(γi\)\|\|q\(γj\)\)≤logN/4,\\mathrm\{KL\}\(q^\{\(\\gamma\_\{i\}\)\}\|\|q^\{\(\\gamma\_\{j\}\)\}\)\\leq\\log N/4,we have

infγ^∈Γscore,0maxγ∗∈G⁡𝔼q\(γ∗\)​\[\|γ^−γ∗\|2\]≥ε2/8\.\\displaystyle\\inf\_\{\\hat\{\\gamma\}\\in\\Gamma\_\{\\mathrm\{score\},0\}\}\\max\_\{\\gamma^\{\*\}\\in G\}\\mathbb\{E\}\_\{q^\{\(\\gamma^\{\*\}\)\}\}\[\|\\hat\{\\gamma\}\-\\gamma^\{\*\}\|^\{2\}\]\\geq\\varepsilon^\{2\}/8\.\(13\)

We note that the proof follows from\[[62](https://arxiv.org/html/2607.15485#bib.bib30), Lemma 3\], which is an application of Fano’s inequality\. To get the error in the squared norm,𝔼\[\|γ^−γ∗\|2,\\mathbb\{E\}\[\|\\hat\{\\gamma\}\-\\gamma^\{\*\}\|^\{2\},as we have in \([13](https://arxiv.org/html/2607.15485#A3.E13)\), we need to additionally use Markov’s inequality:𝔼\[\|γ^−γ\|2≥\(ε/2\)2ℙ\(\|γ^−γ∗\|\>ε/2\)≥\(ε/2\)2\(1/2\)\.\\mathbb\{E\}\[\|\\hat\{\\gamma\}\-\\gamma\|^\{2\}\\geq\(\\varepsilon/2\)^\{2\}\\mathbb\{P\}\(\|\\hat\{\\gamma\}\-\\gamma^\{\*\}\|\>\\varepsilon/2\)\\geq\(\\varepsilon/2\)^\{2\}\\\>\(1/2\)\.The second inequality directly follows from the proof of Lemma 3 of\[[62](https://arxiv.org/html/2607.15485#bib.bib30)\]\. The main difference here is that the data used by the estimator corresponds to the score \([Definition3\.1](https://arxiv.org/html/2607.15485#S3.Thmtheorem1)\), chosen to exemplify the diffusion model setting\.

###### Definition C\.3\.

\[ϵ\\epsilon\-covering entropy\] LetOtO\_\{t\}be a finite open cover of\[0,1\],\[0,1\],with centersγ1,⋯,γNt,\\gamma\_\{1\},\\cdots,\\gamma\_\{N\_\{t\}\},such that for anyγ∈\(0,1\),\\gamma\\in\(0,1\),minjKL\(qt\(γ\)\|\|qt\(γj\)\)<ϵ2\.\\min\_\{j\}\\mathrm\{KL\}\(q\_\{t\}^\{\(\\gamma\)\}\|\|q\_\{t\}^\{\(\\gamma\_\{j\}\)\}\)<\\epsilon^\{2\}\.Then,Vϵ,t:=log⁡\|Ot\|V\_\{\\epsilon,t\}:=\\log\|O\_\{t\}\|is called theϵ\\epsilon\-covering entropy\.

The main novelty in the proof of Theorem[3\.5](https://arxiv.org/html/2607.15485#S3.Thmtheorem5)is the setup of score oracle\-based estimators\. Note that the above standard definitions of packing and covering entropies extend verbatim to our infinite\-dimensional statistical models\.

Secondly, while annealing has been known to be statistically beneficial\[[27](https://arxiv.org/html/2607.15485#bib.bib27),[44](https://arxiv.org/html/2607.15485#bib.bib20)\], our result proves better minimax bounds for parameter recovery, rather than density estimation\. To complete the upper bound in \([10](https://arxiv.org/html/2607.15485#S3.E10)\), we apply directly Theorem 2\[[61](https://arxiv.org/html/2607.15485#bib.bib3)\]\. To obtain the lower bound, we use Lemma[C\.2](https://arxiv.org/html/2607.15485#S5.EGx7)with a possibly different packing set of sizeN​\(γ∗\)N\(\\gamma^\{\*\}\)around eachγ∗\.\\gamma^\{\*\}\.Using the definition ofϵ\\epsilon\-sensitivity and assuming thatlog⁡N​\(γ∗\)≥max⁡\{log⁡16,σt2​V​\(ϵt\)\}\\log N\(\\gamma^\{\*\}\)\\geq\\max\\\{\\log 16,\\sigma\_\{t\}^\{2\}\\\>V\(\\epsilon\_\{t\}\)\\\}at allγ∗,\\gamma^\{\*\},we obtain that for anyγ,γ′∈S​\(γ∗\),\\gamma,\\gamma^\{\\prime\}\\in S\(\\gamma^\{\*\}\),KL\(q\(γ\)\|\|q\(γ′\)\)≤δ=ϵt2/\(4σt2\)≤\(logN\(γ∗\)/V\(ϵt\)\)ϵt2/4=logN\(γ∗\)/4,\\mathrm\{KL\}\(q^\{\(\\gamma\)\}\|\|q^\{\(\\gamma^\{\\prime\}\)\}\)\\leq\\delta=\\epsilon\_\{t\}^\{2\}/\(4\\sigma\_\{t\}^\{2\}\)\\leq\(\\log N\(\\gamma^\{\*\}\)/V\(\\epsilon\_\{t\}\)\)\\epsilon\_\{t\}^\{2\}/4=\\log N\(\\gamma^\{\*\}\)/4,sinceϵt2​σt2=V​\(ϵt\)\.\\epsilon\_\{t\}^\{2\}\\\>\\sigma\_\{t\}^\{2\}=V\(\\epsilon\_\{t\}\)\.As a result, the conditions of Lemma[C\.2](https://arxiv.org/html/2607.15485#S5.EGx7)apply at everyγ∗\.\\gamma^\{\*\}\.

## Appendix DDiffusion process on Gaussian mixture models

### D\.1Mixture Family Preliminaries

Letp\(0\)p^\{\(0\)\}andp\(1\)p^\{\(1\)\}be two fixed probability distributions onℝd\\mathbb\{R\}^\{d\}\. We recall the family of mixed probability densities

𝒫​\(p\(0\),p\(1\)\)≔\{p\(γ\)≔\(1−γ\)​p\(0\)​\(𝐱\)\+γ​p\(1\)​\(𝐱\)\|γ∈\[0,1\]\}\.\\mathcal\{P\}\(p^\{\(0\)\},p^\{\(1\)\}\)\\;\\coloneqq\\;\\big\\\{\\,p^\{\(\\gamma\)\}\\coloneqq\(1\-\\gamma\)\\,p^\{\(0\)\}\(\{\\mathbf\{x\}\}\)\+\\gamma\\,p^\{\(1\)\}\(\{\\mathbf\{x\}\}\)\\;\\big\|\\;\\gamma\\in\[0,1\]\\,\\big\\\}\.The set of noisy marginals at noise levelttproduced by the forward process \([2](https://arxiv.org/html/2607.15485#S2.E2)\) applied to samples𝐱0∼p\(γ\)\{\\mathbf\{x\}\}\_\{0\}\\sim p^\{\(\\gamma\)\}is denoted\{pt\(γ\)\}t∈\[0,1\]\\\{p\_\{t\}^\{\(\\gamma\)\}\\\}\_\{t\\in\[0,1\]\}\. The forward process preserves the mixture structure: the noisy marginalpt\(γ\)p\_\{t\}^\{\(\\gamma\)\}is itself a mixture of the noisy component marginalspt\(0\)p\_\{t\}^\{\(0\)\}andpt\(1\)p\_\{t\}^\{\(1\)\}, with the same mixture weightγ\\gamma\. The family of noisy mixtures and the noisy family of mixtures therefore coincide\.

###### Lemma D\.1\(Mixture closure under forward VP\-SDE\)\.

For everyγ∈\(0,1\)\\gamma\\in\(0,1\)and everyt∈\[0,1\]t\\in\[0,1\],

pt\(γ\)​\(𝐱\)=\(1−γ\)​pt\(0\)​\(𝐱\)\+γ​pt\(1\)​\(𝐱\)p\_\{t\}^\{\(\\gamma\)\}\(\{\\mathbf\{x\}\}\)=\(1\-\\gamma\)p\_\{t\}^\{\(0\)\}\(\{\\mathbf\{x\}\}\)\+\\gamma p\_\{t\}^\{\(1\)\}\(\{\\mathbf\{x\}\}\)\(14\)

###### Proof\.

Letkt​\(𝐲∣𝐱\):=𝒩​\(𝐲;αt​𝐱,\(1−αt\)​𝐈\)k\_\{t\}\(\\mathbf\{y\}\\mid\{\\mathbf\{x\}\}\):=\\mathcal\{N\}\(\\mathbf\{y\};\\,\\sqrt\{\\alpha\_\{t\}\}\\,\{\\mathbf\{x\}\},\\,\(1\-\\alpha\_\{t\}\)\\,\\mathbf\{I\}\)denote the forward conditional density \([3](https://arxiv.org/html/2607.15485#S2.E3)\)\. Sincektk\_\{t\}does not depend on the initial distribution, the marginal at noise levelttis obtained by a linear convolution against the initial density:

pt\(γ\)​\(𝐱\)=∫ℝdkt​\(𝐱∣𝐱0\)​p\(γ\)​\(𝐱0\)​d𝐱0\.p\_\{t\}^\{\(\\gamma\)\}\(\{\\mathbf\{x\}\}\)\\;=\\;\\int\_\{\\mathbb\{R\}^\{d\}\}k\_\{t\}\(\{\\mathbf\{x\}\}\\mid\{\\mathbf\{x\}\}\_\{0\}\)\\,p^\{\(\\gamma\)\}\(\{\\mathbf\{x\}\}\_\{0\}\)\\,\\mathrm\{d\}\{\\mathbf\{x\}\}\_\{0\}\.Substituting the definition of the mixture and applying the linearity of integration,

pt\(γ\)​\(𝐱\)\\displaystyle p\_\{t\}^\{\(\\gamma\)\}\(\{\\mathbf\{x\}\}\)\\;=\(1−γ\)​∫ℝdkt​\(𝐱∣𝐱0\)​p\(0\)​\(𝐱0\)​d𝐱0\+γ​∫ℝdkt​\(𝐱∣𝐱0\)​p\(1\)​\(𝐱0\)​d𝐱0\\displaystyle=\\;\(1\-\\gamma\)\\int\_\{\\mathbb\{R\}^\{d\}\}k\_\{t\}\(\{\\mathbf\{x\}\}\\mid\{\\mathbf\{x\}\}\_\{0\}\)\\,p^\{\(0\)\}\(\{\\mathbf\{x\}\}\_\{0\}\)\\,\\mathrm\{d\}\{\\mathbf\{x\}\}\_\{0\}\\;\+\\;\\gamma\\int\_\{\\mathbb\{R\}^\{d\}\}k\_\{t\}\(\{\\mathbf\{x\}\}\\mid\{\\mathbf\{x\}\}\_\{0\}\)\\,p^\{\(1\)\}\(\{\\mathbf\{x\}\}\_\{0\}\)\\,\\mathrm\{d\}\{\\mathbf\{x\}\}\_\{0\}=\(1−γ\)​pt\(0\)​\(𝐱\)\+γ​pt\(1\)​\(𝐱\)\.\\displaystyle=\\;\(1\-\\gamma\)\\,p\_\{t\}^\{\(0\)\}\(\{\\mathbf\{x\}\}\)\+\\gamma\\,p\_\{t\}^\{\(1\)\}\(\{\\mathbf\{x\}\}\)\.∎

A useful consequence of[LemmaD\.1](https://arxiv.org/html/2607.15485#A4.Thmtheorem1)is that samples frompt\(γ\)p\_\{t\}^\{\(\\gamma\)\}can be drawn by first sampling a component labelc∼Ber​\(γ\)c\\sim\\mathrm\{Ber\}\(\\gamma\)and then drawing𝐲t∼pt\(c\)\\mathbf\{y\}\_\{t\}\\sim p\_\{t\}^\{\(c\)\}, a component\-wise decomposition that simplifies subsequent analysis\.

#### D\.1\.1The score of the noisy mixture

Differentiatinglog⁡pt\(γ\)\\log p\_\{t\}^\{\(\\gamma\)\}and applying[LemmaD\.1](https://arxiv.org/html/2607.15485#A4.Thmtheorem1)gives the score of the noisy mixture as a pointwise convex combination of the component scores,

∇𝐱log⁡pt\(γ\)​\(𝐱\)=w0t​\(𝐱\)​∇𝐱log⁡pt\(0\)​\(𝐱\)\+w1t​\(𝐱\)​∇𝐱log⁡pt\(1\)​\(𝐱\),\\nabla\_\{\\mathbf\{x\}\}\\log p\_\{t\}^\{\(\\gamma\)\}\(\{\\mathbf\{x\}\}\)\\;=\\;w\_\{0\}^\{t\}\(\{\\mathbf\{x\}\}\)\\,\\nabla\_\{\\mathbf\{x\}\}\\log p\_\{t\}^\{\(0\)\}\(\{\\mathbf\{x\}\}\)\\;\+\\;w\_\{1\}^\{t\}\(\{\\mathbf\{x\}\}\)\\,\\nabla\_\{\\mathbf\{x\}\}\\log p\_\{t\}^\{\(1\)\}\(\{\\mathbf\{x\}\}\),\(15\)where the weights are the EM responsibilities — the posterior probabilities under the generative model that𝐱\{\\mathbf\{x\}\}was produced by component0or component11, respectively\[[7](https://arxiv.org/html/2607.15485#bib.bib60), §9\.2\],

w0t​\(𝐱\)≔\(1−γ\)​pt\(0\)​\(𝐱\)pt\(γ\)​\(𝐱\),w1t​\(𝐱\)≔γ​pt\(1\)​\(𝐱\)pt\(γ\)​\(𝐱\),w\_\{0\}^\{t\}\(\{\\mathbf\{x\}\}\)\\;\\coloneqq\\;\\frac\{\(1\-\\gamma\)\\,p\_\{t\}^\{\(0\)\}\(\{\\mathbf\{x\}\}\)\}\{p\_\{t\}^\{\(\\gamma\)\}\(\{\\mathbf\{x\}\}\)\},\\qquad w\_\{1\}^\{t\}\(\{\\mathbf\{x\}\}\)\\;\\coloneqq\\;\\frac\{\\gamma\\,p\_\{t\}^\{\(1\)\}\(\{\\mathbf\{x\}\}\)\}\{p\_\{t\}^\{\(\\gamma\)\}\(\{\\mathbf\{x\}\}\)\},\(16\)and satisfyw0t​\(𝐱\)\+w1t​\(𝐱\)=1w\_\{0\}^\{t\}\(\{\\mathbf\{x\}\}\)\+w\_\{1\}^\{t\}\(\{\\mathbf\{x\}\}\)=1\.

#### D\.1\.2Specialization: Gaussian components

When the components are Gaussian,p\(c\)=𝒩​\(𝝁c∗,𝚺c∗\)p^\{\(c\)\}=\\mathcal\{N\}\(\\bm\{\\mu\}\_\{c\}^\{\*\},\\bm\{\\Sigma\}\_\{c\}^\{\*\}\)forc∈\{0,1\}c\\in\\\{0,1\\\}, the noisy marginals remain Gaussian,pt\(c\)=𝒩​\(𝝁c,t∗,𝚺c,t∗\)p\_\{t\}^\{\(c\)\}=\\mathcal\{N\}\(\\bm\{\\mu\}\_\{c,t\}^\{\*\},\\bm\{\\Sigma\}\_\{c,t\}^\{\*\}\), and the score \([15](https://arxiv.org/html/2607.15485#A4.E15)\) takes the explicit form

∇𝐱log⁡pt\(γ\)​\(𝐱\)=∑c∈\{0,1\}wct​\(𝐱\)​\(𝚺c,t∗\)−1​\(𝝁c,t∗−𝐱\),\\nabla\_\{\\mathbf\{x\}\}\\log p\_\{t\}^\{\(\\gamma\)\}\(\{\\mathbf\{x\}\}\)\\;=\\;\\sum\_\{c\\in\\\{0,1\\\}\}w\_\{c\}^\{t\}\(\{\\mathbf\{x\}\}\)\\,\(\\bm\{\\Sigma\}\_\{c,t\}^\{\*\}\)^\{\-1\}\(\\bm\{\\mu\}\_\{c,t\}^\{\*\}\-\{\\mathbf\{x\}\}\),\(17\)a responsibility\-weighted sum of precision\-scaled displacements\(𝚺c,t∗\)−1​\(𝝁c,t∗−𝐱\)\(\\bm\{\\Sigma\}\_\{c,t\}^\{\*\}\)^\{\-1\}\(\\bm\{\\mu\}\_\{c,t\}^\{\*\}\-\{\\mathbf\{x\}\}\)\. The responsibility weights \([16](https://arxiv.org/html/2607.15485#A4.E16)\) admit the explicit closed form

wct​\(𝐱\)=πc​𝒩​\(𝐱;𝝁c,t∗,𝚺c,t∗\)∑c′∈\{0,1\}πc′​𝒩​\(𝐱;𝝁c′,t∗,𝚺c′,t∗\),w\_\{c\}^\{t\}\(\{\\mathbf\{x\}\}\)\\;=\\;\\frac\{\\pi\_\{c\}\\,\\mathcal\{N\}\\\!\\left\(\{\\mathbf\{x\}\};\\,\\bm\{\\mu\}\_\{c,t\}^\{\*\},\\,\\bm\{\\Sigma\}\_\{c,t\}^\{\*\}\\right\)\}\{\\sum\_\{c^\{\\prime\}\\in\\\{0,1\\\}\}\\pi\_\{c^\{\\prime\}\}\\,\\mathcal\{N\}\\\!\\left\(\{\\mathbf\{x\}\};\\,\\bm\{\\mu\}\_\{c^\{\\prime\},t\}^\{\*\},\\,\\bm\{\\Sigma\}\_\{c^\{\\prime\},t\}^\{\*\}\\right\)\},\(18\)with mixture weightsπ0=1−γ\\pi\_\{0\}=1\-\\gammaandπ1=γ\\pi\_\{1\}=\\gamma\. The numerator and denominator together carry information aboutγ\\gamma, the Mahalanobis distance from𝐱\{\\mathbf\{x\}\}to each component mean, and the Gaussian normalizing constantsdet\(𝚺c,t∗\)−1/2\\det\(\\bm\{\\Sigma\}\_\{c,t\}^\{\*\}\)^\{\-1/2\}\.

#### D\.1\.3DDPM\-style forward process

###### Lemma D\.2\.

Forp\(0\):=𝒩​\(𝐱;μ0,𝐈\)p^\{\(0\)\}:=\\mathcal\{N\}\(\\mathbf\{x\};\\mu\_\{0\},\\mathbf\{I\}\)andp\(1\):=𝒩​\(𝐱;μ1,𝐈\)p^\{\(1\)\}:=\\mathcal\{N\}\(\\mathbf\{x\};\\mu\_\{1\},\\mathbf\{I\}\)denote two individual component distributions\. For a variance\-preserving forward process with noise schedule\(α¯t\)t∈\[0,T\]\(\\bar\{\\alpha\}\_\{t\}\)\_\{t\\in\[0,T\]\}, the parameters of the mixture modes at timettare

𝝁c,t∗=α¯t​𝝁c∗,𝚺c,t∗=α¯t​𝚺c∗\+\(1−α¯t\)​𝐈,\\bm\{\\mu\}\_\{c,t\}^\{\*\}\\;=\\;\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\bm\{\\mu\}\_\{c\}^\{\*\},\\qquad\\bm\{\\Sigma\}\_\{c,t\}^\{\*\}\\;=\\;\\bar\{\\alpha\}\_\{t\}\\,\\bm\{\\Sigma\}\_\{c\}^\{\*\}\+\(1\-\\bar\{\\alpha\}\_\{t\}\)\\,\\mathbf\{I\},\(19\)Hence, the noisy marginal distributions are,

pt\(0\)​\(𝐱\)=𝒩​\(𝐱;αt​μ0,𝐈\),pt\(1\)​\(𝐱\)=𝒩​\(𝐱;αt​μ1,𝐈\)\\displaystyle p\_\{t\}^\{\(0\)\}\(\\mathbf\{x\}\)=\\mathcal\{N\}\(\\mathbf\{x\};\\sqrt\{\\alpha\_\{t\}\}\\mu\_\{0\},\\mathbf\{I\}\),\\qquad p\_\{t\}^\{\(1\)\}\(\\mathbf\{x\}\)=\\mathcal\{N\}\(\\mathbf\{x\};\\sqrt\{\\alpha\_\{t\}\}\\mu\_\{1\},\\mathbf\{I\}\)

###### Proof\.

Let𝐱0∼𝒩​\(μ,𝚺\)\\mathbf\{x\}\_\{0\}\\sim\\mathcal\{N\}\(\\mu,\\mathbf\{\\Sigma\}\)\. We have that

𝐱t\|𝐱0\\displaystyle\\mathbf\{x\}\_\{t\}\|\\mathbf\{x\}\_\{0\}∼𝒩​\(αt​𝐱0,\(1−αt\)​𝐈\)\\displaystyle\\sim\\mathcal\{N\}\(\\sqrt\{\\alpha\_\{t\}\}\\mathbf\{x\}\_\{0\},\(1\-\\alpha\_\{t\}\)\\mathbf\{I\}\)⟹𝐱t\\displaystyle\\Longrightarrow\\mathbf\{x\}\_\{t\}=αt​𝐱0\+1−αt​𝐳\\displaystyle=\\sqrt\{\\alpha\_\{t\}\}\\mathbf\{x\}\_\{0\}\+\\sqrt\{1\-\\alpha\_\{t\}\}\\mathbf\{z\}where𝐳∼𝒩​\(0,𝐈\)\\mathbf\{z\}\\sim\\mathcal\{N\}\(0,\\mathbf\{I\}\)is independent of𝐱0\\mathbf\{x\}\_\{0\}\. Therefore,

𝐱t\\displaystyle\\mathbf\{x\}\_\{t\}∼𝒩​\(αt​μ,αt​𝚺\+\(1−αt\)​𝐈\)\\displaystyle\\sim\\mathcal\{N\}\\left\(\\sqrt\{\\alpha\_\{t\}\}\\mathbf\{\\mu\},\\alpha\_\{t\}\\mathbf\{\\Sigma\}\+\(1\-\\alpha\_\{t\}\)\\mathbf\{I\}\\right\)∎

As per \([19](https://arxiv.org/html/2607.15485#A4.E19)\)αt\\alpha\_\{t\}decreases from11to0, the component means contract toward the origin and the covariances inflate toward𝐈\\mathbf\{I\}\. Substituting \([19](https://arxiv.org/html/2607.15485#A4.E19)\) into \([18](https://arxiv.org/html/2607.15485#A4.E18)\) reveals two limits of interest\. At smalltt\(αt→1\\alpha\_\{t\}\\to 1\), the perturbed components remain well separated, and the responsibilities concentrate sharply on the closest mode, recovering the M\-step regime of the EM algorithm\. At largett\(αt→0\\alpha\_\{t\}\\to 0\), the perturbed components all approach𝒩​\(𝟎,𝐈\)\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)and the responsibilities flatten toward the prior weights\(1−γ,γ\)\(1\-\\gamma,\\gamma\), becoming nearly independent of𝐱\{\\mathbf\{x\}\}\. Sensitivity of the score to the mixture parameterγ\\gammais therefore concentrated at the noise scales for which the responsibility weights are neither saturated nor uninformative — a window whose location and width are determined by the geometry of the components\.

We now demonstrate an equivalence between the neural network loss \([9](https://arxiv.org/html/2607.15485#S3.E9)\) and the diffusion score matching loss \([8](https://arxiv.org/html/2607.15485#S3.E8)\)\.

###### Theorem D\.3\.

Letp\(γ\)p^\{\(\\gamma\)\}be a data distribution onℝd\\mathbb\{R\}^\{d\}with𝔼p\(γ\)​‖𝐱0‖2<∞\\mathbb\{E\}\_\{p^\{\(\\gamma\)\}\}\\\|\\mathbf\{x\}\_\{0\}\\\|^\{2\}<\\infty, and let\{pt\(γ\)\}t∈\[0,1\]\\\{p\_\{t\}^\{\(\\gamma\)\}\\\}\_\{t\\in\[0,1\]\}be its marginals under the VP\-SDE \([2](https://arxiv.org/html/2607.15485#S2.E2)\)–\([3](https://arxiv.org/html/2607.15485#S2.E3)\)\. Let𝐬𝛉:ℝd×\[0,1\]→ℝd\\mathbf\{s\}\_\{\\bm\{\\theta\}\}:\\mathbb\{R\}^\{d\}\\times\[0,1\]\\to\\mathbb\{R\}^\{d\}be a score model with𝔼pt\(γ\)​‖𝐬𝛉​\(⋅,t\)‖2<∞\\mathbb\{E\}\_\{p\_\{t\}^\{\(\\gamma\)\}\}\\\|\\mathbf\{s\}\_\{\\bm\{\\theta\}\}\(\\cdot,t\)\\\|^\{2\}<\\inftyfor almost everyt∈\[0,1\]t\\in\[0,1\]\. Then

ℓNN​\(p\(𝜽\),p\(γ\)\)=ℓDSM​\(p\(𝜽\),p\(γ\)\)\+C​\(p\(γ\)\),\\ell\_\{\\mathrm\{NN\}\}\(p^\{\(\{\\bm\{\\theta\}\}\)\},p^\{\(\\gamma\)\}\)\\;=\\;\\ell\_\{\\mathrm\{DSM\}\}\(p^\{\(\{\\bm\{\\theta\}\}\)\},p^\{\(\\gamma\)\}\)\\;\+\\;C\(p^\{\(\\gamma\)\}\),\(20\)where the constantC​\(p\(γ\)\)C\(p^\{\(\\gamma\)\}\)does not depend on𝛉\{\\bm\{\\theta\}\}and admits the closed form

C​\(p\(γ\)\)\\displaystyle C\(p^\{\(\\gamma\)\}\)=∫01λt⋅αt\(1−αt\)2​MMSEt​\(p\(γ\)\)​𝑑t,\\displaystyle\\;=\\;\\int\_\{0\}^\{1\}\\lambda\_\{t\}\\cdot\\frac\{\\alpha\_\{t\}\}\{\(1\-\\alpha\_\{t\}\)^\{2\}\}\\,\\mathrm\{MMSE\}\_\{t\}\(p^\{\(\\gamma\)\}\)\\,dt,\(21\)MMSEt​\(p\(γ\)\)\\displaystyle\\mathrm\{MMSE\}\_\{t\}\(p^\{\(\\gamma\)\}\):=𝔼𝐱0∼p\(γ\)𝐱t∼pt\(γ\)\(⋅∣𝐱0\)​\[‖𝐱0−𝐱^0​\(𝐱t\)‖2\],\\displaystyle\\;:=\\;\\mathbb\{E\}\_\{\\begin\{subarray\}\{c\}\\mathbf\{x\}\_\{0\}\\,\\sim\\,p^\{\(\\gamma\)\}\\\\ \\mathbf\{x\}\_\{t\}\\,\\sim\\,p\_\{t\}^\{\(\\gamma\)\}\(\\cdot\\mid\\mathbf\{x\}\_\{0\}\)\\end\{subarray\}\}\\\!\\left\[\\,\\big\\\|\\mathbf\{x\}\_\{0\}\-\\widehat\{\\mathbf\{x\}\}\_\{0\}\(\\mathbf\{x\}\_\{t\}\)\\big\\\|^\{2\}\\,\\right\],\(22\)𝐱^0​\(𝐱t\)\\displaystyle\\widehat\{\\mathbf\{x\}\}\_\{0\}\(\\mathbf\{x\}\_\{t\}\):=𝔼𝐱0′∼p\(γ\)\(⋅∣𝐱t\)​\[𝐱0′\]\.\\displaystyle\\;:=\\;\\mathbb\{E\}\_\{\\mathbf\{x\}\_\{0\}^\{\\prime\}\\,\\sim\\,p^\{\(\\gamma\)\}\(\\cdot\\mid\\mathbf\{x\}\_\{t\}\)\}\[\\mathbf\{x\}\_\{0\}^\{\\prime\}\]\.\(23\)Moreover,C​\(p\(γ\)\)≥0C\(p^\{\(\\gamma\)\}\)\\geq 0, with equality if and only ifp\(γ\)p^\{\(\\gamma\)\}is a Dirac mass\. In particular,

ℓNN​\(p\(𝜽\),p\(γ\)\)≥ℓDSM​\(p\(𝜽\),p\(γ\)\)\\ell\_\{\\mathrm\{NN\}\}\(p^\{\(\{\\bm\{\\theta\}\}\)\},p^\{\(\\gamma\)\}\)\\;\\geq\\;\\ell\_\{\\mathrm\{DSM\}\}\(p^\{\(\{\\bm\{\\theta\}\}\)\},p^\{\(\\gamma\)\}\)\(24\)for every𝛉∈𝚯\{\\bm\{\\theta\}\}\\in\\bm\{\\Theta\}\.

###### Proof\.

Our arguments follow the reasoning in\[[58](https://arxiv.org/html/2607.15485#bib.bib62)\]with an additional logic for guaranteeing the required relation between the neural network loss and the diffusion score matching loss\. Throughout, all expectations are taken underp\(γ\)​\(𝐱0\)p^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{0\}\),pt\(γ\)​\(𝐱t∣𝐱0\)p\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{0\}\)for the joint, and underpt\(γ\)​\(𝐱t\)p\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\)for the marginal\. We suppress the dependence onttin𝐬𝜽​\(𝐱t,t\)\\mathbf\{s\}\_\{\\bm\{\\theta\}\}\(\\mathbf\{x\}\_\{t\},t\)when convenient\. Fixt∈\[0,1\]t\\in\[0,1\]\. Expanding the squared norm in the inner expectation ofℓDSM\\ell\_\{\\mathrm\{DSM\}\},

𝔼𝐱t​‖𝐬𝜽​\(𝐱t,t\)−𝐬t\(γ\)​\(𝐱t\)‖2=𝔼𝐱t​‖𝐬𝜽​\(𝐱t,t\)‖2−2​bt​\(𝜽\)\+𝔼𝐱t​‖𝐬t\(γ\)​\(𝐱t\)‖2,\\mathbb\{E\}\_\{\\mathbf\{x\}\_\{t\}\}\\\!\\big\\\|\\mathbf\{s\}\_\{\\bm\{\\theta\}\}\(\\mathbf\{x\}\_\{t\},t\)\-\\mathbf\{s\}\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\)\\big\\\|^\{2\}=\\mathbb\{E\}\_\{\\mathbf\{x\}\_\{t\}\}\\\|\\mathbf\{s\}\_\{\\bm\{\\theta\}\}\(\\mathbf\{x\}\_\{t\},t\)\\\|^\{2\}\\,\-\\,2\\,b\_\{t\}\(\{\\bm\{\\theta\}\}\)\\,\+\\,\\mathbb\{E\}\_\{\\mathbf\{x\}\_\{t\}\}\\\!\\big\\\|\\mathbf\{s\}\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\)\\big\\\|^\{2\},\(25\)withbt​\(𝜽\):=𝔼𝐱t​⟨𝐬𝜽​\(𝐱t,t\),𝐬t\(γ\)​\(𝐱t\)⟩b\_\{t\}\(\{\\bm\{\\theta\}\}\):=\\mathbb\{E\}\_\{\\mathbf\{x\}\_\{t\}\}\\\!\\big\\langle\\mathbf\{s\}\_\{\\bm\{\\theta\}\}\(\\mathbf\{x\}\_\{t\},t\),\\,\\mathbf\{s\}\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\)\\big\\rangle\. Similarly,

𝔼𝐱0,𝐱t∥𝐬𝜽\(𝐱t,t\)−∇𝐱tlogpt\(γ\)\(𝐱t∣𝐱0\)∥2\\displaystyle\\mathbb\{E\}\_\{\\mathbf\{x\}\_\{0\},\\mathbf\{x\}\_\{t\}\}\\\!\\big\\\|\\mathbf\{s\}\_\{\\bm\{\\theta\}\}\(\\mathbf\{x\}\_\{t\},t\)\-\\nabla\_\{\\mathbf\{x\}\_\{t\}\}\\log p\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{0\}\)\\big\\\|^\{2\}\(26\)=𝔼𝐱0,𝐱t∥𝐬𝜽\(𝐱t,t\)∥2−2at\(𝜽\)\+𝔼𝐱0,𝐱t∥∇𝐱tlogpt\(γ\)\(𝐱t∣𝐱0\)∥2,\\displaystyle\\qquad=\\mathbb\{E\}\_\{\\mathbf\{x\}\_\{0\},\\mathbf\{x\}\_\{t\}\}\\\|\\mathbf\{s\}\_\{\\bm\{\\theta\}\}\(\\mathbf\{x\}\_\{t\},t\)\\\|^\{2\}\\,\-\\,2\\,a\_\{t\}\(\{\\bm\{\\theta\}\}\)\\,\+\\,\\mathbb\{E\}\_\{\\mathbf\{x\}\_\{0\},\\mathbf\{x\}\_\{t\}\}\\\!\\big\\\|\\nabla\_\{\\mathbf\{x\}\_\{t\}\}\\log p\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{0\}\)\\big\\\|^\{2\},withat​\(𝜽\):=𝔼𝐱0,𝐱t​⟨𝐬𝜽​\(𝐱t,t\),∇𝐱tlog⁡pt\(γ\)​\(𝐱t∣𝐱0\)⟩a\_\{t\}\(\{\\bm\{\\theta\}\}\):=\\mathbb\{E\}\_\{\\mathbf\{x\}\_\{0\},\\mathbf\{x\}\_\{t\}\}\\\!\\big\\langle\\mathbf\{s\}\_\{\\bm\{\\theta\}\}\(\\mathbf\{x\}\_\{t\},t\),\\,\\nabla\_\{\\mathbf\{x\}\_\{t\}\}\\log p\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{0\}\)\\big\\rangle\. The first term in each expansion coincides because𝐬𝜽​\(𝐱t,t\)\\mathbf\{s\}\_\{\\bm\{\\theta\}\}\(\\mathbf\{x\}\_\{t\},t\)depends only on𝐱t\\mathbf\{x\}\_\{t\}\. We now claim that for allθ\\theta,bt​\(𝜽\)=at​\(𝜽\)b\_\{t\}\(\{\\bm\{\\theta\}\}\)=a\_\{t\}\(\{\\bm\{\\theta\}\}\)\. Using𝐬t\(γ\)=∇pt\(γ\)/pt\(γ\)\\mathbf\{s\}\_\{t\}^\{\(\\gamma\)\}=\\nabla p\_\{t\}^\{\(\\gamma\)\}/p\_\{t\}^\{\(\\gamma\)\}and the marginalizationpt\(γ\)​\(𝐱t\)=∫pt\(γ\)​\(𝐱t∣𝐱0\)​p\(γ\)​\(𝐱0\)​𝑑𝐱0p\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\)=\\int p\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{0\}\)\\,p^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{0\}\)\\,d\\mathbf\{x\}\_\{0\},

bt​\(𝜽\)\\displaystyle b\_\{t\}\(\{\\bm\{\\theta\}\}\)=∫pt\(γ\)​\(𝐱t\)​⟨𝐬𝜽​\(𝐱t,t\),𝐬t\(γ\)​\(𝐱t\)⟩​𝑑𝐱t\\displaystyle=\\int p\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\)\\,\\big\\langle\\mathbf\{s\}\_\{\\bm\{\\theta\}\}\(\\mathbf\{x\}\_\{t\},t\),\\,\\mathbf\{s\}\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\)\\big\\rangle\\,d\\mathbf\{x\}\_\{t\}=∫⟨𝐬𝜽​\(𝐱t,t\),∇𝐱tpt\(γ\)​\(𝐱t\)⟩​𝑑𝐱t\\displaystyle=\\int\\big\\langle\\mathbf\{s\}\_\{\\bm\{\\theta\}\}\(\\mathbf\{x\}\_\{t\},t\),\\,\\nabla\_\{\\mathbf\{x\}\_\{t\}\}p\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\)\\big\\rangle\\,d\\mathbf\{x\}\_\{t\}=∫⟨𝐬𝜽​\(𝐱t,t\),∇𝐱t​∫pt\(γ\)​\(𝐱t∣𝐱0\)​p\(γ\)​\(𝐱0\)​𝑑𝐱0⟩​𝑑𝐱t\\displaystyle=\\int\\Big\\langle\\mathbf\{s\}\_\{\\bm\{\\theta\}\}\(\\mathbf\{x\}\_\{t\},t\),\\,\\nabla\_\{\\mathbf\{x\}\_\{t\}\}\\\!\\\!\\int p\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{0\}\)\\,p^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{0\}\)\\,d\\mathbf\{x\}\_\{0\}\\Big\\rangle\\,d\\mathbf\{x\}\_\{t\}=∬⟨𝐬𝜽​\(𝐱t,t\),∇𝐱tpt\(γ\)​\(𝐱t∣𝐱0\)⟩​p\(γ\)​\(𝐱0\)​𝑑𝐱0​𝑑𝐱t\\displaystyle=\\iint\\big\\langle\\mathbf\{s\}\_\{\\bm\{\\theta\}\}\(\\mathbf\{x\}\_\{t\},t\),\\,\\nabla\_\{\\mathbf\{x\}\_\{t\}\}p\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{0\}\)\\big\\rangle\\,p^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{0\}\)\\,d\\mathbf\{x\}\_\{0\}\\,d\\mathbf\{x\}\_\{t\}=∬pt\(γ\)​\(𝐱t∣𝐱0\)​p\(γ\)​\(𝐱0\)​⟨𝐬𝜽​\(𝐱t,t\),∇𝐱tlog⁡pt\(γ\)​\(𝐱t∣𝐱0\)⟩​𝑑𝐱0​𝑑𝐱t\\displaystyle=\\iint p\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{0\}\)\\,p^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{0\}\)\\,\\big\\langle\\mathbf\{s\}\_\{\\bm\{\\theta\}\}\(\\mathbf\{x\}\_\{t\},t\),\\,\\nabla\_\{\\mathbf\{x\}\_\{t\}\}\\log p\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{0\}\)\\big\\rangle\\,d\\mathbf\{x\}\_\{0\}\\,d\\mathbf\{x\}\_\{t\}=at​\(𝜽\),\\displaystyle=a\_\{t\}\(\{\\bm\{\\theta\}\}\),where Leibniz’s rule applies by smoothness and Gaussian decay of the forward kernel, and Fubini’s theorem applies due to the regularity assumption on𝐬𝜽\\mathbf\{s\}\_\{\{\\bm\{\\theta\}\}\}\. Combining the above,

𝔼𝐱0,𝐱t∥𝐬𝜽\(𝐱t,t\)−∇𝐱tlogpt\(γ\)\(𝐱t∣𝐱0\)∥2−𝔼𝐱t∥𝐬𝜽\(𝐱t,t\)−𝐬t\(γ\)\(𝐱t\)∥2=gap\(t\),\\mathbb\{E\}\_\{\\mathbf\{x\}\_\{0\},\\mathbf\{x\}\_\{t\}\}\\\!\\big\\\|\\mathbf\{s\}\_\{\\bm\{\\theta\}\}\(\\mathbf\{x\}\_\{t\},t\)\-\\nabla\_\{\\mathbf\{x\}\_\{t\}\}\\log p\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{0\}\)\\big\\\|^\{2\}\-\\mathbb\{E\}\_\{\\mathbf\{x\}\_\{t\}\}\\\!\\big\\\|\\mathbf\{s\}\_\{\\bm\{\\theta\}\}\(\\mathbf\{x\}\_\{t\},t\)\-\\mathbf\{s\}\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\)\\big\\\|^\{2\}=\\mathrm\{gap\}\(t\),\(27\)where

gap\(t\):=𝔼𝐱0,𝐱t∥∇𝐱tlogpt\(γ\)\(𝐱t∣𝐱0\)∥2−𝔼𝐱t∥𝐬t\(γ\)\(𝐱t\)∥2,\\mathrm\{gap\}\(t\):=\\mathbb\{E\}\_\{\\mathbf\{x\}\_\{0\},\\mathbf\{x\}\_\{t\}\}\\\!\\big\\\|\\nabla\_\{\\mathbf\{x\}\_\{t\}\}\\log p\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{0\}\)\\big\\\|^\{2\}\-\\mathbb\{E\}\_\{\\mathbf\{x\}\_\{t\}\}\\\!\\big\\\|\\mathbf\{s\}\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\)\\big\\\|^\{2\},\(28\)which is independent of𝜽\{\\bm\{\\theta\}\}\. The Gaussian forward kernel \([3](https://arxiv.org/html/2607.15485#S2.E3)\) gives

∇𝐱tlog⁡pt\(γ\)​\(𝐱t∣𝐱0\)=−𝐱t−αt​𝐱01−αt=−𝐳1−αt,𝐳∼𝒩​\(𝟎,𝐈\),\\nabla\_\{\\mathbf\{x\}\_\{t\}\}\\log p\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{0\}\)=\-\\frac\{\\mathbf\{x\}\_\{t\}\-\\sqrt\{\\alpha\_\{t\}\}\\mathbf\{x\}\_\{0\}\}\{1\-\\alpha\_\{t\}\}=\-\\frac\{\\mathbf\{z\}\}\{\\sqrt\{1\-\\alpha\_\{t\}\}\},\\quad\\mathbf\{z\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\),\(29\)hence the first term of \([28](https://arxiv.org/html/2607.15485#A4.E28)\) equals

𝔼∥∇𝐱tlogpt\(γ\)\(𝐱t∣𝐱0\)∥2=d1−αt\.\\mathbb\{E\}~\\\!\\big\\\|\\nabla\_\{\\mathbf\{x\}\_\{t\}\}\\log p\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{0\}\)\\big\\\|^\{2\}=\\frac\{d\}\{1\-\\alpha\_\{t\}\}\.\(30\)For the marginal score, take expectations of \([29](https://arxiv.org/html/2607.15485#A4.E29)\) conditional on𝐱t\\mathbf\{x\}\_\{t\}we obtain Tweedie’s formula:

𝐬t\(γ\)​\(𝐱t\)=𝔼​\[∇𝐱tlog⁡pt\(γ\)​\(𝐱t∣𝐱0\)\|𝐱t\]=αt​𝔼​\[𝐱0∣𝐱t\]−𝐱t1−αt\.\\mathbf\{s\}\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\)=\\mathbb\{E\}~\\\!\\big\[\\nabla\_\{\\mathbf\{x\}\_\{t\}\}\\log p\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{0\}\)\\,\\big\|\\,\\mathbf\{x\}\_\{t\}\\big\]=\\frac\{\\sqrt\{\\alpha\_\{t\}\}\\,\\mathbb\{E\}\[\\mathbf\{x\}\_\{0\}\\mid\\mathbf\{x\}\_\{t\}\]\-\\mathbf\{x\}\_\{t\}\}\{1\-\\alpha\_\{t\}\}\.\(31\)Setting𝜼^​\(𝐱t\):=αt​𝔼​\[𝐱0∣𝐱t\]=𝔼​\[αt​𝐱0∣𝐱t\]\\hat\{\\bm\{\\eta\}\}\(\\mathbf\{x\}\_\{t\}\):=\\sqrt\{\\alpha\_\{t\}\}\\,\\mathbb\{E\}\[\\mathbf\{x\}\_\{0\}\\mid\\mathbf\{x\}\_\{t\}\]=\\mathbb\{E\}\[\\sqrt\{\\alpha\_\{t\}\}\\mathbf\{x\}\_\{0\}\\mid\\mathbf\{x\}\_\{t\}\], we see that,

𝔼​‖𝐬t\(γ\)​\(𝐱t\)‖2=𝔼​‖𝐱t−𝜼^​\(𝐱t\)‖2\(1−αt\)2\.\\mathbb\{E\}~\\\!\\big\\\|\\mathbf\{s\}\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\)\\big\\\|^\{2\}=\\frac\{\\mathbb\{E\}\\\|\\mathbf\{x\}\_\{t\}\-\\hat\{\\bm\{\\eta\}\}\(\\mathbf\{x\}\_\{t\}\)\\\|^\{2\}\}\{\(1\-\\alpha\_\{t\}\)^\{2\}\}\.\(32\)Decomposing𝐱t−𝜼^=\(𝐱t−αt​𝐱0\)\+\(αt​𝐱0−𝜼^\)=1−αt​𝐳\+αt​\(𝐱0−𝔼​\[𝐱0∣𝐱t\]\)\\mathbf\{x\}\_\{t\}\-\\hat\{\\bm\{\\eta\}\}=\(\\mathbf\{x\}\_\{t\}\-\\sqrt\{\\alpha\_\{t\}\}\\mathbf\{x\}\_\{0\}\)\+\(\\sqrt\{\\alpha\_\{t\}\}\\mathbf\{x\}\_\{0\}\-\\hat\{\\bm\{\\eta\}\}\)=\\sqrt\{1\-\\alpha\_\{t\}\}\\,\\mathbf\{z\}\+\\sqrt\{\\alpha\_\{t\}\}\(\\mathbf\{x\}\_\{0\}\-\\mathbb\{E\}\[\\mathbf\{x\}\_\{0\}\\mid\\mathbf\{x\}\_\{t\}\]\)and expanding the expectation we get,

𝔼​‖𝐱t−𝜼^‖2=\(1−αt\)​d\+2​αt​\(1−αt\)​𝔼​⟨𝐳,𝐱0−𝔼​\[𝐱0∣𝐱t\]⟩\+αt​MMSEt​\(p\(γ\)\)\.\\mathbb\{E\}\\\|\\mathbf\{x\}\_\{t\}\-\\hat\{\\bm\{\\eta\}\}\\\|^\{2\}=\(1\-\\alpha\_\{t\}\)\\,d\+2\\sqrt\{\\alpha\_\{t\}\(1\-\\alpha\_\{t\}\)\}\\,\\mathbb\{E\}~\\\!\\big\\langle\\mathbf\{z\},\\,\\mathbf\{x\}\_\{0\}\-\\mathbb\{E\}\[\\mathbf\{x\}\_\{0\}\\mid\\mathbf\{x\}\_\{t\}\]\\big\\rangle\+\\alpha\_\{t\}\\,\\mathrm\{MMSE\}\_\{t\}\(p^\{\(\\gamma\)\}\)\.\(33\)For the cross term, use𝐳=\(𝐱t−αt​𝐱0\)/1−αt\\mathbf\{z\}=\(\\mathbf\{x\}\_\{t\}\-\\sqrt\{\\alpha\_\{t\}\}\\mathbf\{x\}\_\{0\}\)/\\sqrt\{1\-\\alpha\_\{t\}\}and orthogonality of conditional expectation,𝔼​\[\(𝐱0−𝔼​\[𝐱0∣𝐱t\]\)​𝐱^0​\(𝐱t\)\]=0\\mathbb\{E\}~\\\!\\big\[\(\\mathbf\{x\}\_\{0\}\-\\mathbb\{E\}\[\\mathbf\{x\}\_\{0\}\\mid\\mathbf\{x\}\_\{t\}\]\)\\,\\widehat\{\\mathbf\{x\}\}\_\{0\}\(\\mathbf\{x\}\_\{t\}\)\\big\]=0for anygg:

𝔼​⟨𝐳,𝐱0−𝔼​\[𝐱0∣𝐱t\]⟩\\displaystyle\\mathbb\{E\}~\\\!\\big\\langle\\mathbf\{z\},\\,\\mathbf\{x\}\_\{0\}\-\\mathbb\{E\}\[\\mathbf\{x\}\_\{0\}\\mid\\mathbf\{x\}\_\{t\}\]\\big\\rangle=11−αt​𝔼​⟨𝐱t−αt​𝐱0,𝐱0−𝔼​\[𝐱0∣𝐱t\]⟩\\displaystyle=\\tfrac\{1\}\{\\sqrt\{1\-\\alpha\_\{t\}\}\}\\,\\mathbb\{E\}~\\\!\\big\\langle\\mathbf\{x\}\_\{t\}\-\\sqrt\{\\alpha\_\{t\}\}\\mathbf\{x\}\_\{0\},\\,\\mathbf\{x\}\_\{0\}\-\\mathbb\{E\}\[\\mathbf\{x\}\_\{0\}\\mid\\mathbf\{x\}\_\{t\}\]\\big\\rangle=11−αt​𝐱^0​\(0−αt​MMSEt​\(p\(γ\)\)\)\\displaystyle=\\tfrac\{1\}\{\\sqrt\{1\-\\alpha\_\{t\}\}\}\\,\\widehat\{\\mathbf\{x\}\}\_\{0\}\(0\-\\sqrt\{\\alpha\_\{t\}\}\\,\\mathrm\{MMSE\}\_\{t\}\(p^\{\(\\gamma\)\}\)\\big\)=−αt​MMSEt​\(p\(γ\)\)1−αt\.\\displaystyle=\-\\frac\{\\sqrt\{\\alpha\_\{t\}\}\\,\\mathrm\{MMSE\}\_\{t\}\(p^\{\(\\gamma\)\}\)\}\{\\sqrt\{1\-\\alpha\_\{t\}\}\}\.Substituting into \([33](https://arxiv.org/html/2607.15485#A4.E33)\),

𝔼​‖𝐱t−𝜼^‖2=\(1−αt\)​d−αt​MMSEt​\(p\(γ\)\),\\mathbb\{E\}\\\|\\mathbf\{x\}\_\{t\}\-\\hat\{\\bm\{\\eta\}\}\\\|^\{2\}=\(1\-\\alpha\_\{t\}\)\\,d\-\\alpha\_\{t\}\\,\\mathrm\{MMSE\}\_\{t\}\(p^\{\(\\gamma\)\}\),\(34\)and therefore via \([32](https://arxiv.org/html/2607.15485#A4.E32)\),

𝔼​‖𝐬t\(γ\)​\(𝐱t\)‖2=d1−αt−αt​MMSEt​\(p\(γ\)\)\(1−αt\)2\.\\mathbb\{E\}~\\\!\\big\\\|\\mathbf\{s\}\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{x\}\_\{t\}\)\\big\\\|^\{2\}=\\frac\{d\}\{1\-\\alpha\_\{t\}\}\-\\frac\{\\alpha\_\{t\}\\,\\mathrm\{MMSE\}\_\{t\}\(p^\{\(\\gamma\)\}\)\}\{\(1\-\\alpha\_\{t\}\)^\{2\}\}\.\(35\)Combining \([30](https://arxiv.org/html/2607.15485#A4.E30)\) and \([35](https://arxiv.org/html/2607.15485#A4.E35)\) in \([28](https://arxiv.org/html/2607.15485#A4.E28)\),

gap​\(t\)=αt​MMSEt​\(p\(γ\)\)\(1−αt\)2≥0\.\\mathrm\{gap\}\(t\)=\\frac\{\\alpha\_\{t\}\\,\\mathrm\{MMSE\}\_\{t\}\(p^\{\(\\gamma\)\}\)\}\{\(1\-\\alpha\_\{t\}\)^\{2\}\}\\;\\geq\\;0\.\(36\)Hence, multiplying \([27](https://arxiv.org/html/2607.15485#A4.E27)\) byλt\\lambda\_\{t\}and integrating overt∼Unif​\(0,1\)t\\sim\\mathrm\{Unif\}\(0,1\), the difference in the neural network loss and the diffusion score matching loss can be written as,

ℓNN\(p\(𝜽\),p\(γ\)\)−ℓDSM\(p\(𝜽\),p\(γ\)\)=∫01λtgap\(t\)dt=∫01λtαt​MMSEt​\(p\(γ\)\)\(1−αt\)2dt=:C\(p\(γ\)\),\\ell\_\{\\mathrm\{NN\}\}\(p^\{\(\{\\bm\{\\theta\}\}\)\},p^\{\(\\gamma\)\}\)\-\\ell\_\{\\mathrm\{DSM\}\}\(p^\{\(\{\\bm\{\\theta\}\}\)\},p^\{\(\\gamma\)\}\)=\\int\_\{0\}^\{1\}\\lambda\_\{t\}\\,\\mathrm\{gap\}\(t\)\\,dt=\\int\_\{0\}^\{1\}\\lambda\_\{t\}\\,\\frac\{\\alpha\_\{t\}\\,\\mathrm\{MMSE\}\_\{t\}\(p^\{\(\\gamma\)\}\)\}\{\(1\-\\alpha\_\{t\}\)^\{2\}\}\\,dt=:C\(p^\{\(\\gamma\)\}\),\(37\)proving \([20](https://arxiv.org/html/2607.15485#A4.E20)\)–\([21](https://arxiv.org/html/2607.15485#A4.E21)\)\. For the equality case, note thatC​\(p\(γ\)\)=0C\(p^\{\(\\gamma\)\}\)=0iffMMSEt​\(p\(γ\)\)=0\\mathrm\{MMSE\}\_\{t\}\(p^\{\(\\gamma\)\}\)=0for almost everyt∈\(0,1\)t\\in\(0,1\)\. The latter requires𝐱0=𝔼​\[𝐱0∣𝐱t\]\\mathbf\{x\}\_\{0\}=\\mathbb\{E\}\[\\mathbf\{x\}\_\{0\}\\mid\\mathbf\{x\}\_\{t\}\]almost surely, i\.e\.,𝐱0\\mathbf\{x\}\_\{0\}isσ​\(𝐱t\)\\sigma\(\\mathbf\{x\}\_\{t\}\)\-measurable\. Since𝐱t=αt​𝐱0\+1−αt​𝐳\\mathbf\{x\}\_\{t\}=\\sqrt\{\\alpha\_\{t\}\}\\mathbf\{x\}\_\{0\}\+\\sqrt\{1\-\\alpha\_\{t\}\}\\mathbf\{z\}with𝐳\\mathbf\{z\}independent of𝐱0\\mathbf\{x\}\_\{0\}and1−αt\>01\-\\alpha\_\{t\}\>0, this forces𝐱0\\mathbf\{x\}\_\{0\}to be deterministic\. ∎

### D\.2Proof of[Theorem4\.1](https://arxiv.org/html/2607.15485#S4.Thmtheorem1)

Throughout the proof, letΔ:=∥μ1−μ0∥\\Delta:=\\lVert\\mu\_\{1\}\-\\mu\_\{0\}\\rVert\. From[LemmaD\.1](https://arxiv.org/html/2607.15485#A4.Thmtheorem1)and[LemmaD\.2](https://arxiv.org/html/2607.15485#A4.Thmtheorem2), we can generate samples frompt\(γ\)p\_\{t\}^\{\(\\gamma\)\}by samplingc∼Ber​\(γ\)c\\sim\\mathrm\{Ber\}\(\\gamma\)and sampling𝐲t∼pt\(c\)\\mathbf\{y\}\_\{t\}\\sim p\_\{t\}^\{\(c\)\}\. Furthermore, the next lemma will show that we can generate these samples using the Probability Flow ODE \(\[[53](https://arxiv.org/html/2607.15485#bib.bib24)\]\) instead of the stochastic forward process\.

First, note that

ℓDSM​\(p\(γ^\),p\(γ∗\)\)\\displaystyle\\ell\_\{\\mathrm\{DSM\}\}\(p^\{\(\\hat\{\\gamma\}\)\},p^\{\(\\gamma^\{\*\}\)\}\)=𝔼t,𝐱t∼pt\(γ∗\)​\[∥𝐬t\(γ∗\)​\(𝐱t\)−𝐬t\(γ^\)∥​\(𝐱t\)\]\\displaystyle=\\mathbb\{E\}\_\{t,\\mathbf\{x\}\_\{t\}\\sim p\_\{t\}^\{\(\\gamma^\{\*\}\)\}\}\\left\[\\lVert\\mathbf\{s\}\_\{t\}^\{\(\\gamma^\{\*\}\)\}\(\\mathbf\{x\}\_\{t\}\)\-\\mathbf\{s\}\_\{t\}^\{\(\\hat\{\\gamma\}\)\}\\rVert\(\\mathbf\{x\}\_\{t\}\)\\right\]\(38\)=\(1−γ∗\)​𝔼t,𝐱t∼pt\(γ∗\)\|c=0​\[∥𝐬t\(γ∗\)​\(𝐱t\)−𝐬t\(γ^\)∥​\(𝐱t\)\]\\displaystyle=\(1\-\\gamma^\{\*\}\)\\mathbb\{E\}\_\{t,\\mathbf\{x\}\_\{t\}\\sim p\_\{t\}^\{\(\\gamma^\{\*\}\)\}\|c=0\}\\left\[\\lVert\\mathbf\{s\}\_\{t\}^\{\(\\gamma^\{\*\}\)\}\(\\mathbf\{x\}\_\{t\}\)\-\\mathbf\{s\}\_\{t\}^\{\(\\hat\{\\gamma\}\)\}\\rVert\(\\mathbf\{x\}\_\{t\}\)\\right\]\(39\)\+γ∗​𝔼t,𝐱t∼pt\(γ∗\)\|c=1​\[∥𝐬t\(γ∗\)​\(𝐱t\)−𝐬t\(γ^\)∥​\(𝐱t\)\]\\displaystyle\\quad\\quad\+\\gamma^\{\*\}\\mathbb\{E\}\_\{t,\\mathbf\{x\}\_\{t\}\\sim p\_\{t\}^\{\(\\gamma^\{\*\}\)\}\|c=1\}\\left\[\\lVert\\mathbf\{s\}\_\{t\}^\{\(\\gamma^\{\*\}\)\}\(\\mathbf\{x\}\_\{t\}\)\-\\mathbf\{s\}\_\{t\}^\{\(\\hat\{\\gamma\}\)\}\\rVert\(\\mathbf\{x\}\_\{t\}\)\\right\]\(40\)=\(1−γ∗\)​𝔼t,𝐱t∼pt\(0\)​\[∥𝐬t\(γ∗\)​\(𝐱t\)−𝐬t\(γ^\)∥​\(𝐱t\)\]\\displaystyle=\(1\-\\gamma^\{\*\}\)\\mathbb\{E\}\_\{t,\\mathbf\{x\}\_\{t\}\\sim p\_\{t\}^\{\(0\)\}\}\\left\[\\lVert\\mathbf\{s\}\_\{t\}^\{\(\\gamma^\{\*\}\)\}\(\\mathbf\{x\}\_\{t\}\)\-\\mathbf\{s\}\_\{t\}^\{\(\\hat\{\\gamma\}\)\}\\rVert\(\\mathbf\{x\}\_\{t\}\)\\right\]\(41\)\+γ∗​𝔼t,𝐱t∼pt\(1\)​\[∥𝐬t\(γ∗\)​\(𝐱t\)−𝐬t\(γ^\)∥​\(𝐱t\)\]\\displaystyle\\quad\\quad\+\\gamma^\{\*\}\\mathbb\{E\}\_\{t,\\mathbf\{x\}\_\{t\}\\sim p\_\{t\}^\{\(1\)\}\}\\left\[\\lVert\\mathbf\{s\}\_\{t\}^\{\(\\gamma^\{\*\}\)\}\(\\mathbf\{x\}\_\{t\}\)\-\\mathbf\{s\}\_\{t\}^\{\(\\hat\{\\gamma\}\)\}\\rVert\(\\mathbf\{x\}\_\{t\}\)\\right\]\(42\)
In particular, we can lower\-boundℓDSM\\ell\_\{\\mathrm\{DSM\}\}by independently lower\-bounding each of these expectations\.

###### Lemma D\.4\.

Consider any distributionπ\\piand letπt\\pi\_\{t\}be the noisy marginal at timettof the forward process\. Then, if𝐲t\\mathbf\{y\}\_\{t\}satisfies the Probability Flow ODE

d​𝐲t=−12​βt​\(𝐲t\+∇log⁡p𝐲t​\(𝐲t\)\)​d​t\\mathrm\{d\}\\mathbf\{y\}\_\{t\}=\-\\frac\{1\}\{2\}\\beta\_\{t\}\\left\(\\mathbf\{y\}\_\{t\}\+\\nabla\\log p\_\{\\mathbf\{y\}\_\{t\}\}\(\\mathbf\{y\}\_\{t\}\)\\right\)\\mathrm\{d\}t\(43\)and𝐲0∼π\\mathbf\{y\}\_\{0\}\\sim\\pi, then𝐲t∼πt\\mathbf\{y\}\_\{t\}\\sim\\pi\_\{t\}\.

###### Proof\.

Let𝐱t\\mathbf\{x\}\_\{t\}satisfy the Variance Preserving SDE and𝐲t\\mathbf\{y\}\_\{t\}be as in the lemma statement\. Then, we have the Fokker\-Planck equations

∂p𝐱t∂t\\displaystyle\\frac\{\\partial p\_\{\\mathbf\{x\}\_\{t\}\}\}\{\\partial t\}=−∇⋅\(−12​βt​𝐱​p𝐱t\)\+12​βt​Δ​p𝐱t\\displaystyle=\-\\nabla\\cdot\\left\(\-\\frac\{1\}\{2\}\\beta\_\{t\}\\mathbf\{x\}p\_\{\\mathbf\{x\}\_\{t\}\}\\right\)\+\\frac\{1\}\{2\}\\beta\_\{t\}\\Delta p\_\{\\mathbf\{x\}\_\{t\}\}=12​βt​∇⋅\(𝐱​p𝐱t\)\+12​βt​Δ​p𝐱t\\displaystyle=\\frac\{1\}\{2\}\\beta\_\{t\}\\nabla\\cdot\\left\(\\mathbf\{x\}p\_\{\\mathbf\{x\}\_\{t\}\}\\right\)\+\\frac\{1\}\{2\}\\beta\_\{t\}\\Delta p\_\{\\mathbf\{x\}\_\{t\}\}∂p𝐲t∂t\\displaystyle\\frac\{\\partial p\_\{\\mathbf\{y\}\_\{t\}\}\}\{\\partial t\}=12​βt​∇⋅\(𝐲​p𝐲t\)\+12​βt​∇⋅\(p𝐲t​∇log⁡p𝐲t\)\\displaystyle=\\frac\{1\}\{2\}\\beta\_\{t\}\\nabla\\cdot\\left\(\\mathbf\{y\}p\_\{\\mathbf\{y\}\_\{t\}\}\\right\)\+\\frac\{1\}\{2\}\\beta\_\{t\}\\nabla\\cdot\\left\(p\_\{\\mathbf\{y\}\_\{t\}\}\\nabla\\log p\_\{\\mathbf\{y\}\_\{t\}\}\\right\)=12​βt​∇⋅\(𝐲​p𝐲t\)\+12​βt​Δ​p𝐲t\\displaystyle=\\frac\{1\}\{2\}\\beta\_\{t\}\\nabla\\cdot\\left\(\\mathbf\{y\}p\_\{\\mathbf\{y\}\_\{t\}\}\\right\)\+\\frac\{1\}\{2\}\\beta\_\{t\}\\Delta p\_\{\\mathbf\{y\}\_\{t\}\}∎

The Probability Flow ODE \(\([43](https://arxiv.org/html/2607.15485#A4.E43)\)\) substantially simplifies the analysis because of the following fact:

###### Lemma D\.5\.

For𝐲0\\mathbf\{y\}\_\{0\}, and𝐲t\\mathbf\{y\}\_\{t\}following the Probability Flow ODE associated withccand starting from𝐲0∼p\(c\)\\mathbf\{y\}\_\{0\}\\sim p^\{\(c\)\}, we have that

𝐲t−αt​μc=𝐲0−μc\\mathbf\{y\}\_\{t\}\-\\sqrt\{\\alpha\_\{t\}\}\\mu\_\{c\}=\\mathbf\{y\}\_\{0\}\-\\mu\_\{c\}

###### Proof\.

Note thatdd​t​αt=12​βt​αt\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\sqrt\{\\alpha\_\{t\}\}=\\frac\{1\}\{2\}\\beta\_\{t\}\\sqrt\{\\alpha\_\{t\}\}\. Therefore,

dd​t​\(𝐲t−αt​μc\)\\displaystyle\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\(\\mathbf\{y\}\_\{t\}\-\\sqrt\{\\alpha\_\{t\}\}\\mu\_\{c\}\)=−12​βt​\(𝐲t\+∇log⁡pt\(c\)​\(𝐲t\)\)−12​βt​αt​μc\\displaystyle=\-\\frac\{1\}\{2\}\\beta\_\{t\}\\left\(\\mathbf\{y\}\_\{t\}\+\\nabla\\log p\_\{t\}^\{\(c\)\}\(\\mathbf\{y\}\_\{t\}\)\\right\)\-\\frac\{1\}\{2\}\\beta\_\{t\}\\sqrt\{\\alpha\_\{t\}\}\\mu\_\{c\}=−12​βt​\(𝐲t\+αt​μc−𝐲t−αt​μc\)\\displaystyle=\-\\frac\{1\}\{2\}\\beta\_\{t\}\\left\(\\mathbf\{y\}\_\{t\}\+\\sqrt\{\\alpha\_\{t\}\}\\mu\_\{c\}\-\\mathbf\{y\}\_\{t\}\-\\sqrt\{\\alpha\_\{t\}\}\\mu\_\{c\}\\right\)=0\\displaystyle=0∎

We now focus on the case wherec=0c=0\(the case ofc=1c=1will follow via symmetry\)\. Let𝐳t=𝐲t\|c=0\\mathbf\{z\}\_\{t\}=\\mathbf\{y\}\_\{t\}\|c=0\. Our first goal is to lower bound∥𝐬t\(γ∗\)​\(𝐳t\)−𝐬t\(γ^\)​\(𝐳t\)∥2\\lVert\\mathbf\{s\}\_\{t\}^\{\(\\gamma^\{\*\}\)\}\(\\mathbf\{z\}\_\{t\}\)\-\\mathbf\{s\}\_\{t\}^\{\(\\hat\{\\gamma\}\)\}\(\\mathbf\{z\}\_\{t\}\)\\rVert^\{2\}pointwise inttand𝐳t\\mathbf\{z\}\_\{t\}\. We first show an analytic expression for this score difference, then identify a range oftt’s along each path where this expression is sufficiently large\. Finally, by integrating over initial conditions𝐳0\\mathbf\{z\}\_\{0\}, we get a lower bound for𝔼t,𝐳t∼pt\(0\)​\[∥𝐬t\(γ∗\)​\(𝐳t\)−𝐬t\(γ^\)​\(𝐳t\)∥2\]\\mathbb\{E\}\_\{t,\\mathbf\{z\}\_\{t\}\\sim p\_\{t\}^\{\(0\)\}\}\\left\[\\lVert\\mathbf\{s\}\_\{t\}^\{\(\\gamma^\{\*\}\)\}\(\\mathbf\{z\}\_\{t\}\)\-\\mathbf\{s\}\_\{t\}^\{\(\\hat\{\\gamma\}\)\}\(\\mathbf\{z\}\_\{t\}\)\\rVert^\{2\}\\right\]\.

###### Proposition D\.6\.

Letγ^,γ∗∈\[0,1\]\\hat\{\\gamma\},\\gamma^\{\*\}\\in\[0,1\]\. Then,

∥𝐬t\(γ^\)​\(𝐳t\)−𝐬t\(γ∗\)​\(𝐳t\)∥2=\(γ^−γ∗\)2​∥μ0−μ1∥2​αt​e2​f0​\(𝐳0,t\)\(1−γ^\+γ^​ef0​\(𝐳0,t\)\)2​\(1−γ∗\+γ∗​ef0​\(𝐳0,t\)\)2\\lVert\\mathbf\{s\}\_\{t\}^\{\(\\hat\{\\gamma\}\)\}\(\\mathbf\{z\}\_\{t\}\)\-\\mathbf\{s\}\_\{t\}^\{\(\\gamma^\{\*\}\)\}\(\\mathbf\{z\}\_\{t\}\)\\rVert^\{2\}=\(\{\\hat\{\\gamma\}\}\-\\gamma^\{\*\}\)^\{2\}\\frac\{\\lVert\\mu\_\{0\}\-\\mu\_\{1\}\\rVert^\{2\}\\alpha\_\{t\}e^\{2f\_\{0\}\(\\mathbf\{z\}\_\{0\},t\)\}\}\{\(1\-\{\\hat\{\\gamma\}\}\+\{\\hat\{\\gamma\}\}e^\{f\_\{0\}\(\\mathbf\{z\}\_\{0\},t\)\}\)^\{2\}\(1\-\\gamma^\{\*\}\+\\gamma^\{\*\}e^\{f\_\{0\}\(\\mathbf\{z\}\_\{0\},t\)\}\)^\{2\}\}where

f0​\(𝐳0,t\):=−12​∥μ1−μ0∥2​αt\+⟨𝐳0−μ0,μ1−μ0⟩​αtf\_\{0\}\(\\mathbf\{z\}\_\{0\},t\):=\-\\frac\{1\}\{2\}\\lVert\\mu\_\{1\}\-\\mu\_\{0\}\\rVert^\{2\}\\alpha\_\{t\}\+\\left\\langle\\mathbf\{z\}\_\{0\}\-\\mu\_\{0\},\\mu\_\{1\}\-\\mu\_\{0\}\\right\\rangle\\sqrt\{\\alpha\_\{t\}\}

###### Proof\.

We

pt\(0\)​\(𝐱\)=1Zt​exp⁡\(−12​∥𝐱−αt​μ0∥2\)p\_\{t\}^\{\(0\)\}\(\\mathbf\{x\}\)=\\frac\{1\}\{Z\_\{t\}\}\\exp\\left\(\-\\frac\{1\}\{2\}\\lVert\\mathbf\{x\}\-\\sqrt\{\\alpha\_\{t\}\}\\mu\_\{0\}\\rVert^\{2\}\\right\)pt\(1\)​\(𝐱\)=1Zt​exp⁡\(−12​∥𝐱−αt​μ1∥2\)p\_\{t\}^\{\(1\)\}\(\\mathbf\{x\}\)=\\frac\{1\}\{Z\_\{t\}\}\\exp\\left\(\-\\frac\{1\}\{2\}\\lVert\\mathbf\{x\}\-\\sqrt\{\\alpha\_\{t\}\}\\mu\_\{1\}\\rVert^\{2\}\\right\)
Expanding the log likelihoods, we have

log⁡p1​\(𝐳t\)p0​\(𝐳t\)\\displaystyle\\log\\frac\{p\_\{1\}\(\\mathbf\{z\}\_\{t\}\)\}\{p\_\{0\}\(\\mathbf\{z\}\_\{t\}\)\}=−12​∥μ0−μ1∥2​αt\+⟨𝐳0−μ0,μ1−μ0⟩​αt\\displaystyle=\-\\frac\{1\}\{2\}\\lVert\\mu\_\{0\}\-\\mu\_\{1\}\\rVert^\{2\}\\alpha\_\{t\}\+\\left\\langle\\mathbf\{z\}\_\{0\}\-\\mu\_\{0\},\\mu\_\{1\}\-\\mu\_\{0\}\\right\\rangle\\sqrt\{\\alpha\_\{t\}\}
Letf​\(𝐳0,t\):=log⁡p1​\(𝐳t\)p0​\(𝐳t\)f\(\\mathbf\{z\}\_\{0\},t\):=\\log\\frac\{p\_\{1\}\(\\mathbf\{z\}\_\{t\}\)\}\{p\_\{0\}\(\\mathbf\{z\}\_\{t\}\)\}\. Then,

𝐬t\(γ\)​\(𝐳t\)\\displaystyle\\mathbf\{s\}\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{z\}\_\{t\}\)=\(1−γ\)​pt\(0\)​\(𝐳t\)​𝐬t\(0\)​\(𝐳t\)\+γ​pt\(1\)​\(𝐳t\)​𝐬t\(1\)​\(𝐳t\)pt\(γ\)​\(𝐳t\)\\displaystyle=\\frac\{\(1\-\\gamma\)\\;p\_\{t\}^\{\(0\)\}\(\\mathbf\{z\}\_\{t\}\)\\mathbf\{s\}\_\{t\}^\{\(0\)\}\(\\mathbf\{z\}\_\{t\}\)\+\\gamma\\;p\_\{t\}^\{\(1\)\}\(\\mathbf\{z\}\_\{t\}\)\\mathbf\{s\}\_\{t\}^\{\(1\)\}\(\\mathbf\{z\}\_\{t\}\)\}\{p\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{z\}\_\{t\}\)\}=𝐬t\(0\)​\(𝐳t\)\+γ​pt\(1\)​\(𝐳t\)​\(𝐬t\(1\)​\(𝐳t\)−𝐬t\(0\)​\(𝐳t\)\)pt\(γ\)​\(𝐳t\)\\displaystyle=\\mathbf\{s\}\_\{t\}^\{\(0\)\}\(\\mathbf\{z\}\_\{t\}\)\+\\frac\{\\gamma\\;p\_\{t\}^\{\(1\)\}\(\\mathbf\{z\}\_\{t\}\)\\left\(\\mathbf\{s\}\_\{t\}^\{\(1\)\}\(\\mathbf\{z\}\_\{t\}\)\-\\mathbf\{s\}\_\{t\}^\{\(0\)\}\(\\mathbf\{z\}\_\{t\}\)\\right\)\}\{p\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{z\}\_\{t\}\)\}=𝐬t\(0\)​\(𝐳t\)\+γ​pt\(1\)​\(𝐳t\)pt\(γ\)​\(𝐳t\)​αt​\(μ1−μ0\)\\displaystyle=\\mathbf\{s\}\_\{t\}^\{\(0\)\}\(\\mathbf\{z\}\_\{t\}\)\+\\frac\{\\gamma\\;p\_\{t\}^\{\(1\)\}\(\\mathbf\{z\}\_\{t\}\)\}\{p\_\{t\}^\{\(\\gamma\)\}\(\\mathbf\{z\}\_\{t\}\)\}\\sqrt\{\\alpha\_\{t\}\}\\left\(\\mu\_\{1\}\-\\mu\_\{0\}\\right\)=𝐬t\(0\)​\(𝐳t\)\+γ​ef​\(𝐳0,t\)1−γ\+γ​ef​\(𝐳0,t\)​αt​\(μ1−μ0\)\\displaystyle=\\mathbf\{s\}\_\{t\}^\{\(0\)\}\(\\mathbf\{z\}\_\{t\}\)\+\\frac\{\\gamma\\;e^\{f\(\\mathbf\{z\}\_\{0\},t\)\}\}\{1\-\\gamma\+\\gamma\\;e^\{f\(\\mathbf\{z\}\_\{0\},t\)\}\}\\sqrt\{\\alpha\_\{t\}\}\\left\(\\mu\_\{1\}\-\\mu\_\{0\}\\right\)where we used∇log⁡pt\(c\)​\(𝐱\)=αt​μc−𝐱\\nabla\\log p\_\{t\}^\{\(c\)\}\(\\mathbf\{x\}\)=\\sqrt\{\\alpha\_\{t\}\}\\mu\_\{c\}\-\\mathbf\{x\}\. Therefore,

𝐬t\(γ^\)​\(𝐳t\)−𝐬t\(γ∗\)​\(𝐳t\)=\(γ^−γ∗\)​ef​\(𝐳0,t\)\(1−γ^\+γ^​ef​\(𝐳0,t\)\)​\(1−γ∗\+γ∗​ef​\(𝐳0,t\)\)​αt​\(μ1−μ0\)\\displaystyle\\mathbf\{s\}\_\{t\}^\{\(\\hat\{\\gamma\}\)\}\(\\mathbf\{z\}\_\{t\}\)\-\\mathbf\{s\}\_\{t\}^\{\(\\gamma^\{\*\}\)\}\(\\mathbf\{z\}\_\{t\}\)=\\frac\{\(\{\\hat\{\\gamma\}\}\-\\gamma^\{\*\}\)\\;e^\{f\(\\mathbf\{z\}\_\{0\},t\)\}\}\{\\left\(1\-\{\\hat\{\\gamma\}\}\+\{\\hat\{\\gamma\}\}\\;e^\{f\(\\mathbf\{z\}\_\{0\},t\)\}\\right\)\\left\(1\-\\gamma^\{\*\}\+\\gamma^\{\*\}\\;e^\{f\(\\mathbf\{z\}\_\{0\},t\)\}\\right\)\}\\sqrt\{\\alpha\_\{t\}\}\\left\(\\mu\_\{1\}\-\\mu\_\{0\}\\right\)
This gives

∥𝐬t\(γ^\)​\(𝐳t\)−𝐬t\(γ∗\)​\(𝐳t\)∥2=\(γ^−γ∗\)2​∥μ1−μ0∥2​αt​e2​f​\(𝐳0,t\)\(1−γ^\+γ^​ef​\(𝐳0,t\)\)2​\(1−γ∗\+γ∗​ef​\(𝐳0,t\)\)2\\lVert\\mathbf\{s\}\_\{t\}^\{\(\\hat\{\\gamma\}\)\}\(\\mathbf\{z\}\_\{t\}\)\-\\mathbf\{s\}\_\{t\}^\{\(\\gamma^\{\*\}\)\}\(\\mathbf\{z\}\_\{t\}\)\\rVert^\{2\}=\(\{\\hat\{\\gamma\}\}\-\\gamma^\{\*\}\)^\{2\}\\frac\{\\lVert\\mu\_\{1\}\-\\mu\_\{0\}\\rVert^\{2\}\\alpha\_\{t\}\\;e^\{2f\(\\mathbf\{z\}\_\{0\},t\)\}\}\{\\left\(1\-\{\\hat\{\\gamma\}\}\+\{\\hat\{\\gamma\}\}\\;e^\{f\(\\mathbf\{z\}\_\{0\},t\)\}\\right\)^\{2\}\\left\(1\-\\gamma^\{\*\}\+\\gamma^\{\*\}\\;e^\{f\(\\mathbf\{z\}\_\{0\},t\)\}\\right\)^\{2\}\}∎

###### Lemma D\.7\.

For allx∈ℝx\\in\\mathbb\{R\}andγ∈\[0,1\]\\gamma\\in\[0,1\], we have

ex\(1−γ\+γ​ex\)2≥14​γ​\(1−γ\)​e−14​\(x−log⁡1−γγ\)2\\frac\{e^\{x\}\}\{\(1\-\\gamma\+\\gamma e^\{x\}\)^\{2\}\}\\geq\\frac\{1\}\{4\\gamma\(1\-\\gamma\)\}e^\{\-\\frac\{1\}\{4\}\\left\(x\-\\log\\frac\{1\-\\gamma\}\{\\gamma\}\\right\)^\{2\}\}with equality only whenx=log⁡1−γγx=\\log\\frac\{1\-\\gamma\}\{\\gamma\}\. In particular,

e2​f0​\(𝐳0,t\)\(1−γ^\+γ^​ef0​\(𝐳0,t\)\)2​\(1−γ∗\+γ∗​ef0​\(𝐳0,t\)\)2≥e−14​\(f0​\(𝐳0,t\)−log⁡1−γ^γ^\)2−14​\(f0​\(𝐳0,t\)−log⁡1−γ∗γ∗\)216​γ^​\(1−γ^\)​γ∗​\(1−γ∗\)\\displaystyle\\frac\{e^\{2f\_\{0\}\(\\mathbf\{z\}\_\{0\},t\)\}\}\{\\left\(1\-\{\\hat\{\\gamma\}\}\+\{\\hat\{\\gamma\}\}e^\{f\_\{0\}\(\\mathbf\{z\}\_\{0\},t\)\}\\right\)^\{2\}\\left\(1\-\\gamma^\{\*\}\+\\gamma^\{\*\}e^\{f\_\{0\}\(\\mathbf\{z\}\_\{0\},t\)\}\\right\)^\{2\}\}\\geq\\frac\{e^\{\-\\frac\{1\}\{4\}\\left\(f\_\{0\}\(\\mathbf\{z\}\_\{0\},t\)\-\\log\\frac\{1\-\{\\hat\{\\gamma\}\}\}\{\{\\hat\{\\gamma\}\}\}\\right\)^\{2\}\-\\frac\{1\}\{4\}\\left\(f\_\{0\}\(\\mathbf\{z\}\_\{0\},t\)\-\\log\\frac\{1\-\\gamma^\{\*\}\}\{\\gamma^\{\*\}\}\\right\)^\{2\}\}\}\{16\{\\hat\{\\gamma\}\}\(1\-\{\\hat\{\\gamma\}\}\)\\gamma^\{\*\}\(1\-\\gamma^\{\*\}\)\}

###### Proof\.

Letg​\(x\)=ex\(1−γ\+γ​ex\)2g\(x\)=\\frac\{e^\{x\}\}\{\(1\-\\gamma\+\\gamma e^\{x\}\)^\{2\}\}andh​\(x\)=14​γ​\(1−γ\)​e−14​\(x−log⁡1−γγ\)2h\(x\)=\\frac\{1\}\{4\\gamma\(1\-\\gamma\)\}e^\{\-\\frac\{1\}\{4\}\\left\(x\-\\log\\frac\{1\-\\gamma\}\{\\gamma\}\\right\)^\{2\}\}\. Then,

dd​x​log⁡g​\(x\)\\displaystyle\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}x\}\\log g\(x\)=1−γ−γ​ex1−γ\+γ​ex\\displaystyle=\\frac\{1\-\\gamma\-\\gamma e^\{x\}\}\{1\-\\gamma\+\\gamma e^\{x\}\}
For allγ∈\[0,1\]\\gamma\\in\[0,1\]1−γ\+γ​ex\>01\-\\gamma\+\\gamma e^\{x\}\>0\. Therefore,x∗:=log⁡1−γγx^\{\*\}:=\\log\\frac\{1\-\\gamma\}\{\\gamma\}is the only critical point ofg​\(x\)g\(x\), andg​\(x∗\)=14​γ​\(1−γ\)g\(x^\{\*\}\)=\\frac\{1\}\{4\\gamma\\left\(1\-\\gamma\\right\)\}\. Similarlyh​\(x∗\)=14​γ​\(1−γ\)h\(x^\{\*\}\)=\\frac\{1\}\{4\\gamma\\left\(1\-\\gamma\\right\)\}anddd​x​log⁡h​\(x∗\)=0\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}x\}\\log h\(x^\{\*\}\)=0\. We have

dd​x​\[log⁡h​\(x\)−log⁡g​\(x\)\]\\displaystyle\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}x\}\[\\log h\(x\)\-\\log g\(x\)\]=−12​\(x−x∗\)−1−ex−x∗1\+ex−x∗\\displaystyle=\-\\frac\{1\}\{2\}\(x\-x^\{\*\}\)\-\\frac\{1\-e^\{x\-x^\{\*\}\}\}\{1\+e^\{x\-x^\{\*\}\}\}d2d​x2​\[log⁡h​\(x\)−log⁡g​\(x\)\]\\displaystyle\\frac\{\\mathrm\{d\}^\{2\}\}\{\\mathrm\{d\}x^\{2\}\}\[\\log h\(x\)\-\\log g\(x\)\]=−\(1−ex−x∗\)22​\(1\+ex−x∗\)2\\displaystyle=\-\\frac\{\(1\-e^\{x\-x^\{\*\}\}\)^\{2\}\}\{2\(1\+e^\{x\-x^\{\*\}\}\)^\{2\}\}
So,d2d​x2​\[log⁡h​\(x\)−log⁡g​\(x\)\]≤0\\frac\{\\mathrm\{d\}^\{2\}\}\{\\mathrm\{d\}x^\{2\}\}\[\\log h\(x\)\-\\log g\(x\)\]\\leq 0with equality only whenx=x∗x=x^\{\*\}\. Therefore,log⁡h−log⁡g\\log h\-\\log gis strictly concave away fromx∗x^\{\*\}and maximized atx∗x^\{\*\}, whereh​\(x∗\)=g​\(x∗\)h\(x^\{\*\}\)=g\(x^\{\*\}\)\. Therefore,h<gh<g\. ∎

###### Lemma D\.8\.

Let𝐳t\\mathbf\{z\}\_\{t\}and𝐳0\\mathbf\{z\}\_\{0\}be as above\. LetM​\(𝐳0\):=⟨𝐳0−μ0,μ1−μ0⟩M\(\\mathbf\{z\}\_\{0\}\):=\\left\\langle\\mathbf\{z\}\_\{0\}\-\\mu\_\{0\},\\mu\_\{1\}\-\\mu\_\{0\}\\right\\rangle, and define the log\-odds ratios

r^\\displaystyle\\hat\{r\}:=log⁡1−γ^γ^\\displaystyle:=\\log\\frac\{1\-\{\\hat\{\\gamma\}\}\}\{\{\\hat\{\\gamma\}\}\}r∗\\displaystyle r^\{\*\}:=log⁡1−γ∗γ∗\\displaystyle:=\\log\\frac\{1\-\\gamma^\{\*\}\}\{\\gamma^\{\*\}\}
Letϵ\>0\\epsilon\>0be such that

ϵ\\displaystyle\\epsilon≤exp⁡\(−18​\[\(r^−r∗\)2\+\(r^\+r∗−M​\(𝐳0\)2Δ2\)\+2\]\)\\displaystyle\\leq\\exp\\left\(\-\\frac\{1\}\{8\}\\left\[\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\+\\left\(\\hat\{r\}\+r^\{\*\}\-\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{\\Delta^\{2\}\}\\right\)\_\{\+\}^\{2\}\\right\]\\right\)where\(⋅\)\+:=max⁡\{⋅,0\}\(\\cdot\)\_\{\+\}:=\\max\\\{\\cdot,0\\\}\.

Let

a\\displaystyle a=max⁡\{M​\(𝐳0\)Δ2\+1Δ​\(M​\(𝐳0\)2Δ2−\(r^\+r∗\)−8​log⁡1ϵ−\(r^−r∗\)2\)\+,α1\}\\displaystyle=\\max\\left\\\{\\frac\{M\(\\mathbf\{z\}\_\{0\}\)\}\{\\Delta^\{2\}\}\+\\frac\{1\}\{\\Delta\}\\sqrt\{\\left\(\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{\\Delta^\{2\}\}\-\\left\(\\hat\{r\}\+r^\{\*\}\\right\)\-\\sqrt\{8\\log\\frac\{1\}\{\\epsilon\}\-\\left\(\\hat\{r\}\-r^\{\*\}\\right\)^\{2\}\}\\right\)\_\{\+\}\},\\sqrt\{\\alpha\_\{1\}\}\\right\\\}b\\displaystyle b=min⁡\{M​\(𝐳0\)Δ2\+1Δ​M​\(𝐳0\)2Δ2−\(r^\+r∗\)\+8​log⁡1ϵ−\(r^−r∗\)2,1\}\\displaystyle=\\min\\left\\\{\\frac\{M\(\\mathbf\{z\}\_\{0\}\)\}\{\\Delta^\{2\}\}\+\\frac\{1\}\{\\Delta\}\\sqrt\{\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{\\Delta^\{2\}\}\-\\left\(\\hat\{r\}\+r^\{\*\}\\right\)\+\\sqrt\{8\\log\\frac\{1\}\{\\epsilon\}\-\\left\(\\hat\{r\}\-r^\{\*\}\\right\)^\{2\}\}\},1\\right\\\}
Then for allttsuch thatαt∈\[a,b\]\\sqrt\{\\alpha\_\{t\}\}\\in\[a,b\],

e−14​\(f0​\(𝐳0,t\)−r^\)s​2−14​\(f0​\(𝐳0,t\)−r∗\)2≥ϵe^\{\-\\frac\{1\}\{4\}\\left\(f\_\{0\}\(\\mathbf\{z\}\_\{0\},t\)\-\\hat\{r\}\\right\)^\{s\}2\-\\frac\{1\}\{4\}\\left\(f\_\{0\}\(\\mathbf\{z\}\_\{0\},t\)\-r^\{\*\}\\right\)^\{2\}\}\\geq\\epsilon

###### Proof\.

First, the inequality is equivalent to

\(f0​\(𝐳0,t\)−r∗\)2\+\(f0​\(z0,t\)−r∗\)2≤4​log⁡1ϵ\\displaystyle\\left\(f\_\{0\}\(\\mathbf\{z\}\_\{0\},t\)\-r^\{\*\}\\right\)^\{2\}\+\\left\(f\_\{0\}\(z\_\{0\},t\)\-r^\{\*\}\\right\)^\{2\}\\leq 4\\log\\frac\{1\}\{\\epsilon\}
Plugging in the definition offf, the inequality holds whenever

\[Δ22​\(αt−M​\(𝐳0\)Δ2\)2−M​\(𝐳0\)22​Δ2\+r^\+r∗2\]2≤2​log⁡1ϵ−\(r^−r∗\)24\\left\[\\frac\{\\Delta^\{2\}\}\{2\}\\left\(\\sqrt\{\\alpha\_\{t\}\}\-\\frac\{M\(\\mathbf\{z\}\_\{0\}\)\}\{\\Delta^\{2\}\}\\right\)^\{2\}\-\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{2\\Delta^\{2\}\}\+\\frac\{\\hat\{r\}\+r^\{\*\}\}\{2\}\\right\]^\{2\}\\leq 2\\log\\frac\{1\}\{\\epsilon\}\-\\frac\{\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\}\{4\}
We wish to identify the values ofαt\\sqrt\{\\alpha\_\{t\}\}which make this inequality hold\. This function is continuous, so its sign can only change at the roots of the corresponding equality constraint:

\[Δ22​\(αt−M​\(𝐳0\)Δ2\)2−M​\(𝐳0\)22​Δ2\+r^\+r∗2\]2=2​log⁡1ϵ−\(r^−r∗\)24\\left\[\\frac\{\\Delta^\{2\}\}\{2\}\\left\(\\sqrt\{\\alpha\_\{t\}\}\-\\frac\{M\(\\mathbf\{z\}\_\{0\}\)\}\{\\Delta^\{2\}\}\\right\)^\{2\}\-\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{2\\Delta^\{2\}\}\+\\frac\{\\hat\{r\}\+r^\{\*\}\}\{2\}\\right\]^\{2\}=2\\log\\frac\{1\}\{\\epsilon\}\-\\frac\{\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\}\{4\}
Solving this quartic gives the roots

αt=M​\(𝐳0\)Δ2±M​\(𝐳0\)2Δ4−r^\+r∗Δ2±1Δ2​8​log⁡1ϵ−\(r^−r∗\)2\\sqrt\{\\alpha\_\{t\}\}=\\frac\{M\(\\mathbf\{z\}\_\{0\}\)\}\{\\Delta^\{2\}\}\\pm\\sqrt\{\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{\\Delta^\{4\}\}\-\\frac\{\\hat\{r\}\+r^\{\*\}\}\{\\Delta^\{2\}\}\\pm\\frac\{1\}\{\\Delta^\{2\}\}\\sqrt\{8\\log\\frac\{1\}\{\\epsilon\}\-\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\}\}
The condition onϵ\\epsilonis such that the square roots are all well\-defined\. Letb′b^\{\\prime\}be the largest of these roots \(both±\\pm’s are set to\+\+\)\. Asαt→∞\\sqrt\{\\alpha\_\{t\}\}\\rightarrow\\infty, we would havef0​\(𝐳0,t\)→−∞f\_\{0\}\(\\mathbf\{z\}\_\{0\},t\)\\rightarrow\-\\infty\(because−Δ2/2\-\\Delta^\{2\}/2is strictly negative\), and hence

e−14​\(f0​\(𝐳0,t\)−r^\)2−14​\(f0​\(𝐳0,t\)−r∗\)2→0e^\{\-\\frac\{1\}\{4\}\\left\(f\_\{0\}\(\\mathbf\{z\}\_\{0\},t\)\-\\hat\{r\}\\right\)^\{2\}\-\\frac\{1\}\{4\}\\left\(f\_\{0\}\(\\mathbf\{z\}\_\{0\},t\)\-r^\{\*\}\\right\)^\{2\}\}\\rightarrow 0
Therefore, forαt\>b′\\sqrt\{\\alpha\_\{t\}\}\>b^\{\\prime\}the inequality does not hold\. Note that the four roots above are distinct \(with two being complex forϵ\\epsilonsufficiently small\), meaning that each root has multiplicity 1, and hence the polynomial changes signs atb′b^\{\\prime\}\. Therefore, forαt\\sqrt\{\\alpha\_\{t\}\}betweenb′b^\{\\prime\}and the next largest real root, the inequality holds\. The next largest real root is

a′=\{M​\(𝐳0\)Δ2\+M​\(𝐳0\)2Δ4−r^\+r∗Δ2−1Δ2​8​log⁡1ϵ−\(r^−r∗\)2ifM​\(𝐳0\)2Δ2−\(r^\+r∗\)−8​log⁡1ϵ−\(r^−r∗\)2≥0M​\(𝐳0\)Δ2−M​\(𝐳0\)2Δ4−r^\+r∗Δ2\+1Δ2​8​log⁡1ϵ−\(r^−r∗\)2otherwisea^\{\\prime\}=\\begin\{cases\}\\frac\{M\(\\mathbf\{z\}\_\{0\}\)\}\{\\Delta^\{2\}\}\+\\sqrt\{\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{\\Delta^\{4\}\}\-\\frac\{\\hat\{r\}\+r^\{\*\}\}\{\\Delta^\{2\}\}\-\\frac\{1\}\{\\Delta^\{2\}\}\\sqrt\{8\\log\\frac\{1\}\{\\epsilon\}\-\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\}\}&\\text\{if\}\\quad\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{\\Delta^\{2\}\}\-\(\\hat\{r\}\+r^\{\*\}\)\-\\sqrt\{8\\log\\frac\{1\}\{\\epsilon\}\-\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\}\\geq 0\\\\ \\frac\{M\(\\mathbf\{z\}\_\{0\}\)\}\{\\Delta^\{2\}\}\-\\sqrt\{\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{\\Delta^\{4\}\}\-\\frac\{\\hat\{r\}\+r^\{\*\}\}\{\\Delta^\{2\}\}\+\\frac\{1\}\{\\Delta^\{2\}\}\\sqrt\{8\\log\\frac\{1\}\{\\epsilon\}\-\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\}\}&\\text\{otherwise\}\\end\{cases\}
In the second case we havea′≤M​\(𝐳0\)Δ2a^\{\\prime\}\\leq\\frac\{M\(\\mathbf\{z\}\_\{0\}\)\}\{\\Delta^\{2\}\}, so

a′≤M​\(𝐳0\)Δ2\+\(M​\(𝐳0\)2Δ4−r^\+r∗Δ2−1Δ2​8​log⁡1ϵ−\(r^−r∗\)2\)\+=:a′′a^\{\\prime\}\\leq\\frac\{M\(\\mathbf\{z\}\_\{0\}\)\}\{\\Delta^\{2\}\}\+\\sqrt\{\\left\(\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{\\Delta^\{4\}\}\-\\frac\{\\hat\{r\}\+r^\{\*\}\}\{\\Delta^\{2\}\}\-\\frac\{1\}\{\\Delta^\{2\}\}\\sqrt\{8\\log\\frac\{1\}\{\\epsilon\}\-\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\}\\right\)\_\{\+\}\}=:a^\{\\prime\\prime\}
Therefore, the inequality holds on\[a′,b′\]\[a^\{\\prime\},b^\{\\prime\}\]and\[a′′,b′\]⊆\[a′,b′\]\[a^\{\\prime\\prime\},b^\{\\prime\}\]\\subseteq\[a^\{\\prime\},b^\{\\prime\}\], so the inequality holds on\[a′′,b′\]\[a^\{\\prime\\prime\},b^\{\\prime\}\]\. Note that fort∈\[0,1\]t\\in\[0,1\],αt∈\[α1,1\]\\alpha\_\{t\}\\in\[\\alpha\_\{1\},1\]whereα1\>0\\alpha\_\{1\}\>0is small, so we letb:=min⁡\{b′,1\}b:=\\min\\\{b^\{\\prime\},1\\\}anda:=max⁡\{a,α1\}a:=\\max\\\{a,\\sqrt\{\\alpha\_\{1\}\}\\\}\. ∎

###### Proposition D\.9\.

For allϵ\\epsilonsatisfying the condition in[LemmaD\.8](https://arxiv.org/html/2607.15485#A4.Thmtheorem8), for alltts\.t\.αt∈\[a,b\]\\sqrt\{\\alpha\_\{t\}\}\\in\[a,b\]\(withaaandbbas defined in the previous lemma\) we have

∥𝐬t\(γ^\)​\(𝐳t\)−𝐬t\(γ∗\)​\(𝐳t\)∥2≥\(γ^−γ∗\)2​∥μ1−μ0∥2​αt16​γ^​\(1−γ^\)​γ∗​\(1−γ∗\)​ϵ\\displaystyle\\lVert\\mathbf\{s\}\_\{t\}^\{\(\\hat\{\\gamma\}\)\}\(\\mathbf\{z\}\_\{t\}\)\-\\mathbf\{s\}\_\{t\}^\{\(\\gamma^\{\*\}\)\}\(\\mathbf\{z\}\_\{t\}\)\\rVert^\{2\}\\geq\\frac\{\(\\hat\{\\gamma\}\-\\gamma^\{\*\}\)^\{2\}\\lVert\\mu\_\{1\}\-\\mu\_\{0\}\\rVert^\{2\}\\alpha\_\{t\}\}\{16\\hat\{\\gamma\}\(1\-\\hat\{\\gamma\}\)\\gamma^\{\*\}\(1\-\\gamma^\{\*\}\)\}\\epsilon

###### Proof\.

Combining the above two lemmas,

∥𝐬t\(γ^\)​\(𝐳t\)−𝐬t\(γ∗\)​\(𝐳t\)∥2\\displaystyle\\lVert\\mathbf\{s\}\_\{t\}^\{\(\\hat\{\\gamma\}\)\}\(\\mathbf\{z\}\_\{t\}\)\-\\mathbf\{s\}\_\{t\}^\{\(\\gamma^\{\*\}\)\}\(\\mathbf\{z\}\_\{t\}\)\\rVert^\{2\}≥\(γ^−γ∗\)2​∥μ1−μ0∥2​αt16​γ^​\(1−γ^\)​γ∗​\(1−γ∗\)​e−14​\(f0​\(𝐳0,t\)−r^\)2−14​\(f0​\(𝐳0,t\)−r∗\)2\\displaystyle\\geq\\frac\{\(\\hat\{\\gamma\}\-\\gamma^\{\*\}\)^\{2\}\\lVert\\mu\_\{1\}\-\\mu\_\{0\}\\rVert^\{2\}\\alpha\_\{t\}\}\{16\\hat\{\\gamma\}\(1\-\\hat\{\\gamma\}\)\\gamma^\{\*\}\(1\-\\gamma^\{\*\}\)\}e^\{\-\\frac\{1\}\{4\}\\left\(f\_\{0\}\(\\mathbf\{z\}\_\{0\},t\)\-\\hat\{r\}\\right\)^\{2\}\-\\frac\{1\}\{4\}\\left\(f\_\{0\}\(\\mathbf\{z\}\_\{0\},t\)\-r^\{\*\}\\right\)^\{2\}\}≥\(γ^−γ∗\)2​∥μ1−μ0∥2​αt16​γ^​\(1−γ^\)​γ∗​\(1−γ∗\)​ϵ\\displaystyle\\geq\\frac\{\(\\hat\{\\gamma\}\-\\gamma^\{\*\}\)^\{2\}\\lVert\\mu\_\{1\}\-\\mu\_\{0\}\\rVert^\{2\}\\alpha\_\{t\}\}\{16\\hat\{\\gamma\}\(1\-\\hat\{\\gamma\}\)\\gamma^\{\*\}\(1\-\\gamma^\{\*\}\)\}\\epsilon∎

We now turn our attention to lower bounding

∫01∥𝐬t\(γ^\)​\(𝐳t\)−𝐬t\(γ∗\)​\(𝐳t\)∥2​dt\\int\_\{0\}^\{1\}\\lVert\\mathbf\{s\}\_\{t\}^\{\(\\hat\{\\gamma\}\)\}\(\\mathbf\{z\}\_\{t\}\)\-\\mathbf\{s\}\_\{t\}^\{\(\\gamma^\{\*\}\)\}\(\\mathbf\{z\}\_\{t\}\)\\rVert^\{2\}\\mathrm\{d\}talong a given path starting from𝐳𝟎\\mathbf\{z\_\{0\}\}\. From the lemmas above, we have

∫01∥∇log⁡pt\(γ^\)​\(𝐳t\)−∇log⁡pt\(γ∗\)​\(𝐳t\)∥2​dt\>\(γ^−γ∗\)2​∥μ1−μ0∥216​γ^​\(1−γ^\)​γ∗​\(1−γ∗\)​ϵ​∫tb2ta2αt​dt\\int\_\{0\}^\{1\}\\lVert\\nabla\\log p\_\{t\}^\{\(\\hat\{\\gamma\}\)\}\(\\mathbf\{z\}\_\{t\}\)\-\\nabla\\log p\_\{t\}^\{\(\\gamma^\{\*\}\)\}\(\\mathbf\{z\}\_\{t\}\)\\rVert^\{2\}\\mathrm\{d\}t\>\\frac\{\(\{\\hat\{\\gamma\}\}\-\\gamma^\{\*\}\)^\{2\}\\lVert\\mu\_\{1\}\-\\mu\_\{0\}\\rVert^\{2\}\}\{16\{\\hat\{\\gamma\}\}\(1\-\{\\hat\{\\gamma\}\}\)\\gamma^\{\*\}\(1\-\\gamma^\{\*\}\)\}\\epsilon\\int\_\{t\_\{b^\{2\}\}\}^\{t\_\{a^\{2\}\}\}\\alpha\_\{t\}\\mathrm\{d\}twhereta2t\_\{a^\{2\}\}andtb2t\_\{b^\{2\}\}are such thatαta2=a2\\alpha\_\{t\_\{a^\{2\}\}\}=\{a^\{2\}\}andαtb2=b2\\alpha\_\{t\_\{b^\{2\}\}\}=\{b^\{2\}\}\. Noteαt\\alpha\_\{t\}is a decreasing function oftt, hence why the bounds of the integral are reversed\. If we letα−1​\(τ\)\\alpha^\{\-1\}\(\\tau\)be the inverse function ofαt\\alpha\_\{t\}, such thatαα−1​\(τ\)=τ\\alpha\_\{\\alpha^\{\-1\}\(\\tau\)\}=\\tau, then

∫tb2ta2αt​dt=2​∫abτβα−1​\(τ2\)​dτ\\int\_\{t\_\{b^\{2\}\}\}^\{t\_\{a^\{2\}\}\}\\alpha\_\{t\}\\mathrm\{d\}t=2\\int\_\{a\}^\{b\}\\frac\{\\tau\}\{\\beta\_\{\\alpha^\{\-1\}\(\\tau^\{2\}\)\}\}\\mathrm\{d\}\\tau
In particular, if we consider the linear noise schedule

βt:=β1​t\\beta\_\{t\}:=\\beta\_\{1\}tthen

∫tb2ta2αt​dt=1β1​∫abτ−ln⁡τ​dτ\\int\_\{t\_\{b^\{2\}\}\}^\{t\_\{a^\{2\}\}\}\\alpha\_\{t\}\\mathrm\{d\}t=\\frac\{1\}\{\\sqrt\{\\beta\_\{1\}\}\}\\int\_\{a\}^\{b\}\\frac\{\\tau\}\{\\sqrt\{\-\\ln\\tau\}\}\\mathrm\{d\}\\tau
###### Proposition D\.10\.

Letaa,bb,Δ\\Delta,z0z\_\{0\},r^\\hat\{r\},r∗r^\{\*\}, andMM, be as above, and assumeϵ\\epsilonsatisfies the bound from[LemmaD\.8](https://arxiv.org/html/2607.15485#A4.Thmtheorem8)\. Define

R0:=8​ln⁡1ϵ−\(r^−r∗\)2−\(r^\+r∗\)R\_\{0\}:=\\sqrt\{8\\ln\\frac\{1\}\{\\epsilon\}\-\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\}\-\(\\hat\{r\}\+r^\{\*\}\)
Then, whenM​\(𝐳0\)Δ2≥α1\\frac\{M\(\\mathbf\{z\}\_\{0\}\)\}\{\\Delta^\{2\}\}\\geq\\sqrt\{\\alpha\_\{1\}\}and

ϵ≥exp⁡\(−18​\[\(Δ2−2​M​\(𝐳0\)\+\(r^\+r∗\)\)2\+\(r^−r∗\)2\]\)\\epsilon\\geq\\exp\\left\(\-\\frac\{1\}\{8\}\\left\[\(\\Delta^\{2\}\-2M\(\\mathbf\{z\}\_\{0\}\)\+\(\\hat\{r\}\+r^\{\*\}\)\)^\{2\}\+\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\\right\]\\right\)then

∫abτ−ln⁡τ​dτ\\displaystyle\\int\_\{a\}^\{b\}\\frac\{\\tau\}\{\\sqrt\{\-\\ln\\tau\}\}\\mathrm\{d\}\\tau≥M​\(𝐳0\)Δ​M​\(𝐳0\)2Δ2\+R0\+M​\(𝐳0\)22​Δ2\+12​R0Δ2​ln⁡\(Δ\)−ln⁡\(M​\(𝐳0\)Δ\+12​M​\(𝐳0\)2Δ2\+R0\)\\displaystyle\\geq\\frac\{\\frac\{M\(\\mathbf\{z\}\_\{0\}\)\}\{\\Delta\}\\sqrt\{\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{\\Delta^\{2\}\}\+R\_\{0\}\}\+\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{2\\Delta^\{2\}\}\+\\frac\{1\}\{2\}R\_\{0\}\}\{\\Delta^\{2\}\\sqrt\{\\ln\(\\Delta\)\-\\ln\\left\(\\frac\{M\(\\mathbf\{z\}\_\{0\}\)\}\{\\Delta\}\+\\frac\{1\}\{2\}\\sqrt\{\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{\\Delta^\{2\}\}\+R\_\{0\}\}\\right\)\}\}

###### Proof\.

Let

g​\(t\):=τ−ln⁡τg\(t\):=\\frac\{\\tau\}\{\\sqrt\{\-\\ln\\tau\}\}
Note that

g′′​\(t\)\\displaystyle g^\{\\prime\\prime\}\(t\)=12​t​\(−ln⁡t\)3/2\+34​t​\(−ln⁡t\)5/2\>0\\displaystyle=\\frac\{1\}\{2t\(\-\\ln t\)^\{3/2\}\}\+\\frac\{3\}\{4t\(\-\\ln t\)^\{5/2\}\}\>0
So,g​\(t\)g\(t\)is convex, meaning

∫abg​\(τ\)​dτ≥\(b−a\)​g​\(a\+b2\)\\int\_\{a\}^\{b\}g\(\\tau\)\\mathrm\{d\}\\tau\\geq\(b\-a\)g\\left\(\\frac\{a\+b\}\{2\}\\right\)
The condition thatM​\(𝐳0\)Δ2≥α1\\frac\{M\(\\mathbf\{z\}\_\{0\}\)\}\{\\Delta^\{2\}\}\\geq\\sqrt\{\\alpha\_\{1\}\}guaranteesa≥α1a\\geq\\sqrt\{\\alpha\_\{1\}\}\. Similarly, the condition onϵ\\epsilonguaranteesb≤1b\\leq 1, so

b−a\\displaystyle b\-a=1Δ​M​\(𝐳0\)2Δ2−\(r^\+r∗\)\+8​log⁡1ϵ−\(r^−r∗\)2\\displaystyle=\\frac\{1\}\{\\Delta\}\\sqrt\{\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{\\Delta^\{2\}\}\-\\left\(\\hat\{r\}\+r^\{\*\}\\right\)\+\\sqrt\{8\\log\\frac\{1\}\{\\epsilon\}\-\\left\(\\hat\{r\}\-r^\{\*\}\\right\)^\{2\}\}\}=1Δ​M​\(𝐳0\)2Δ2\+R0\\displaystyle=\\frac\{1\}\{\\Delta\}\\sqrt\{\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{\\Delta^\{2\}\}\+R\_\{0\}\}
We have

a\+b2\\displaystyle\\frac\{a\+b\}\{2\}=M​\(𝐳0\)Δ2\+12​Δ​M​\(𝐳0\)2Δ2\+R0\\displaystyle=\\frac\{M\(\\mathbf\{z\}\_\{0\}\)\}\{\\Delta^\{2\}\}\+\\frac\{1\}\{2\\Delta\}\\sqrt\{\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{\\Delta^\{2\}\}\+R\_\{0\}\}and so

⟹g​\(a\+b2\)\\displaystyle\\Longrightarrow g\\left\(\\frac\{a\+b\}\{2\}\\right\)=M​\(𝐳0\)Δ\+12​M​\(𝐳0\)2Δ2\+R0Δ​ln⁡\(Δ\)−ln⁡\(M​\(𝐳0\)Δ\+12​M​\(𝐳0\)2Δ2\+R0\)\\displaystyle=\\frac\{\\frac\{M\(\\mathbf\{z\}\_\{0\}\)\}\{\\Delta\}\+\\frac\{1\}\{2\}\\sqrt\{\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{\\Delta^\{2\}\}\+R\_\{0\}\}\}\{\\Delta\\sqrt\{\\ln\(\\Delta\)\-\\ln\\left\(\\frac\{M\(\\mathbf\{z\}\_\{0\}\)\}\{\\Delta\}\+\\frac\{1\}\{2\}\\sqrt\{\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{\\Delta^\{2\}\}\+R\_\{0\}\}\\right\)\}\}
Combining these,

∫abg​\(τ\)​dτ\\displaystyle\\int\_\{a\}^\{b\}g\(\\tau\)\\mathrm\{d\}\\tau≥M​\(𝐳0\)Δ​M​\(𝐳0\)2Δ2\+R0\+M​\(𝐳0\)22​Δ2\+12​R0Δ2​ln⁡\(Δ\)−ln⁡\(M​\(𝐳0\)Δ\+12​M​\(𝐳0\)2Δ2\+R0\)\\displaystyle\\geq\\frac\{\\frac\{M\(\\mathbf\{z\}\_\{0\}\)\}\{\\Delta\}\\sqrt\{\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{\\Delta^\{2\}\}\+R\_\{0\}\}\+\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{2\\Delta^\{2\}\}\+\\frac\{1\}\{2\}R\_\{0\}\}\{\\Delta^\{2\}\\sqrt\{\\ln\(\\Delta\)\-\\ln\\left\(\\frac\{M\(\\mathbf\{z\}\_\{0\}\)\}\{\\Delta\}\+\\frac\{1\}\{2\}\\sqrt\{\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{\\Delta^\{2\}\}\+R\_\{0\}\}\\right\)\}\}∎

Combining this lower bound with what we showed before, we have

∫01∥∇log⁡pt\(γ^\)​\(𝐳t\)−∇log⁡pt\(γ∗\)​\(𝐳t\)∥2​dt\\displaystyle\\int\_\{0\}^\{1\}\\lVert\\nabla\\log p\_\{t\}^\{\(\\hat\{\\gamma\}\)\}\(\\mathbf\{z\}\_\{t\}\)\-\\nabla\\log p\_\{t\}^\{\(\\gamma^\{\*\}\)\}\(\\mathbf\{z\}\_\{t\}\)\\rVert^\{2\}\\mathrm\{d\}t≥\(γ^−γ∗\)2​ϵ16​β1​γ^​\(1−γ^\)​γ∗​\(1−γ∗\)​M​\(𝐳0\)Δ​M​\(𝐳0\)2Δ2\+R0\+M​\(𝐳0\)22​Δ2\+12​R0ln⁡\(Δ\)−ln⁡\(M​\(𝐳0\)Δ\+12​M​\(𝐳0\)2Δ2\+R0\)\\displaystyle\\quad\\quad\\geq\\frac\{\(\\hat\{\\gamma\}\-\\gamma^\{\*\}\)^\{2\}\\epsilon\}\{16\\sqrt\{\\beta\_\{1\}\}\\;\\hat\{\\gamma\}\(1\-\\hat\{\\gamma\}\)\\gamma^\{\*\}\(1\-\\gamma^\{\*\}\)\}\\frac\{\\frac\{M\(\\mathbf\{z\}\_\{0\}\)\}\{\\Delta\}\\sqrt\{\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{\\Delta^\{2\}\}\+R\_\{0\}\}\+\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{2\\Delta^\{2\}\}\+\\frac\{1\}\{2\}R\_\{0\}\}\{\\sqrt\{\\ln\(\\Delta\)\-\\ln\\left\(\\frac\{M\(\\mathbf\{z\}\_\{0\}\)\}\{\\Delta\}\+\\frac\{1\}\{2\}\\sqrt\{\\frac\{M\(\\mathbf\{z\}\_\{0\}\)^\{2\}\}\{\\Delta^\{2\}\}\+R\_\{0\}\}\\right\)\}\}
Recall that𝐳0∼𝒩​\(μ0,𝐈\)\\mathbf\{z\}\_\{0\}\\sim\\mathcal\{N\}\(\\mu\_\{0\},\\mathbf\{I\}\), so

M​\(𝐳0\)\\displaystyle M\(\\mathbf\{z\}\_\{0\}\)=⟨𝐳0−μ0,μ1−μ0⟩∼𝒩​\(0,∥μ1−μ0∥2\)\\displaystyle=\\left\\langle\\mathbf\{z\}\_\{0\}\-\\mu\_\{0\},\\mu\_\{1\}\-\\mu\_\{0\}\\right\\rangle\\sim\\mathcal\{N\}\(0,\\lVert\\mu\_\{1\}\-\\mu\_\{0\}\\rVert^\{2\}\)⟹M​\(𝐳0\)Δ\\displaystyle\\Longrightarrow\\frac\{M\(\\mathbf\{z\}\_\{0\}\)\}\{\\Delta\}∼𝒩​\(0,1\)\\displaystyle\\sim\\mathcal\{N\}\(0,1\)
###### Proposition D\.11\.

LetΦ\\Phibe the CDF of𝒩​\(0,1\)\\mathcal\{N\}\(0,1\)\. Letτ≥Δ​α1\\tau\\geq\\Delta\\sqrt\{\\alpha\_\{1\}\}, and letϵ\>0\\epsilon\>0satisfyϵ∈\(ϵmin0,ϵmax0\)\\epsilon\\in\(\\epsilon\_\{\\mathrm\{min\}\}^\{0\},\\epsilon\_\{\\mathrm\{max\}\}^\{0\}\)with

ϵmax0\\displaystyle\\epsilon\_\{\\mathrm\{max\}\}^\{0\}=exp⁡\(−18​\[\(r^−r∗\)2\+\(r^\+r∗−τ2\)\+2\]\)\\displaystyle=\\exp\\left\(\-\\frac\{1\}\{8\}\\left\[\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\+\\left\(\\hat\{r\}\+r^\{\*\}\-\\tau^\{2\}\\right\)\_\{\+\}^\{2\}\\right\]\\right\)ϵmin0\\displaystyle\\epsilon\_\{\\mathrm\{min\}\}^\{0\}=exp⁡\(−18​\[\(Δ​\(Δ−2​τ\)\+\(r^\+r∗\)\)2\+\(r^−r∗\)2\]\)\\displaystyle=\\exp\\left\(\-\\frac\{1\}\{8\}\\left\[\(\\Delta\(\\Delta\-2\\tau\)\+\(\\hat\{r\}\+r^\{\*\}\)\)^\{2\}\+\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\\right\]\\right\)
Then, with probability1−Φ​\(τ\)1\-\\Phi\(\\tau\)over samples of𝐳0\\mathbf\{z\}\_\{0\},

∫01∥∇log⁡pt\(γ^\)​\(𝐳t\)−∇log⁡pt\(γ∗\)​\(𝐳t\)∥2​dt\\displaystyle\\int\_\{0\}^\{1\}\\lVert\\nabla\\log p\_\{t\}^\{\(\\hat\{\\gamma\}\)\}\(\\mathbf\{z\}\_\{t\}\)\-\\nabla\\log p\_\{t\}^\{\(\\gamma^\{\*\}\)\}\(\\mathbf\{z\}\_\{t\}\)\\rVert^\{2\}\\mathrm\{d\}t≥\(γ^−γ∗\)2​ϵ16​β1​γ^​\(1−γ^\)​γ∗​\(1−γ∗\)​τ​τ2\+R0\+12​τ2\+12​R0ln⁡\(Δ\)−ln⁡\(τ\+12​τ2\+R0\)\\displaystyle\\quad\\quad\\geq\\frac\{\(\\hat\{\\gamma\}\-\\gamma^\{\*\}\)^\{2\}\\epsilon\}\{16\\sqrt\{\\beta\_\{1\}\}\\;\\hat\{\\gamma\}\(1\-\\hat\{\\gamma\}\)\\gamma^\{\*\}\(1\-\\gamma^\{\*\}\)\}\\frac\{\\tau\\sqrt\{\\tau^\{2\}\+R\_\{0\}\}\+\\frac\{1\}\{2\}\\tau^\{2\}\+\\frac\{1\}\{2\}R\_\{0\}\}\{\\sqrt\{\\ln\(\\Delta\)\-\\ln\\left\(\\tau\+\\frac\{1\}\{2\}\\sqrt\{\\tau^\{2\}\+R\_\{0\}\}\\right\)\}\}

###### Proof\.

If we letK:=M​\(𝐳0\)ΔK:=\\frac\{M\(\\mathbf\{z\}\_\{0\}\)\}\{\\Delta\}our bound becomes

∫01∥∇log⁡pt\(γ^\)​\(𝐳t\)−∇log⁡pt\(γ∗\)​\(𝐳t\)∥2​dt\\displaystyle\\int\_\{0\}^\{1\}\\lVert\\nabla\\log p\_\{t\}^\{\(\\hat\{\\gamma\}\)\}\(\\mathbf\{z\}\_\{t\}\)\-\\nabla\\log p\_\{t\}^\{\(\\gamma^\{\*\}\)\}\(\\mathbf\{z\}\_\{t\}\)\\rVert^\{2\}\\mathrm\{d\}t≥\(γ^−γ∗\)2​ϵ16​γ^​\(1−γ^\)​γ∗​\(1−γ∗\)​K​K2\+R0\+12​K2\+12​R0ln⁡\(Δ\)−ln⁡\(K\+12​K2\+R0\)\\displaystyle\\quad\\quad\\geq\\frac\{\(\\hat\{\\gamma\}\-\\gamma^\{\*\}\)^\{2\}\\epsilon\}\{16\\hat\{\\gamma\}\(1\-\\hat\{\\gamma\}\)\\gamma^\{\*\}\(1\-\\gamma^\{\*\}\)\}\\frac\{K\\sqrt\{K^\{2\}\+R\_\{0\}\}\+\\frac\{1\}\{2\}K^\{2\}\+\\frac\{1\}\{2\}R\_\{0\}\}\{\\sqrt\{\\ln\(\\Delta\)\-\\ln\\left\(K\+\\frac\{1\}\{2\}\\sqrt\{K^\{2\}\+R\_\{0\}\}\\right\)\}\}
Assume thatK≥τK\\geq\\tau, which occurs with probability1−Φ​\(τ\)1\-\\Phi\(\\tau\)\. Our bound is an increasing function inKKand is thus minimized atK=τK=\\tau\. Similarly, our upper bound from before onϵ\\epsilonis a decreasing function ofKKand our lower bound from before is an increasing function ofϵ\\epsilon, so our bounds onϵ\\epsilonare tightest when we plug inτ\\tau, givingϵmax0\\epsilon\_\{\\mathrm\{max\}\}^\{0\}andϵmin0\\epsilon\_\{\\mathrm\{min\}\}^\{0\}from the statement of the proposition\. ∎

We now state the corresponding proposition for the case wherec=1c=1\. Let𝐰t:=𝐲t\|c=1\\mathbf\{w\}\_\{t\}:=\\mathbf\{y\}\_\{t\}\|c=1\.

###### Proposition D\.12\.

Letτ≥Δ​α1\\tau\\geq\\Delta\\sqrt\{\\alpha\_\{1\}\}, and letϵ\>0\\epsilon\>0satisfyϵ∈\(ϵmin1,ϵmax1\)\\epsilon\\in\(\\epsilon\_\{\\mathrm\{min\}\}^\{1\},\\epsilon\_\{\\mathrm\{max\}\}^\{1\}\)with

ϵmax1\\displaystyle\\epsilon\_\{\\mathrm\{max\}\}^\{1\}=exp⁡\(−18​\[\(r^−r∗\)2\+\(−\(r^\+r∗\)−τ2\)\+2\]\)\\displaystyle=\\exp\\left\(\-\\frac\{1\}\{8\}\\left\[\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\+\\left\(\-\(\\hat\{r\}\+r^\{\*\}\)\-\\tau^\{2\}\\right\)\_\{\+\}^\{2\}\\right\]\\right\)ϵmin1\\displaystyle\\epsilon\_\{\\mathrm\{min\}\}^\{1\}=exp⁡\(−18​\[\(Δ​\(Δ−2​τ\)\+\(r^\+r∗\)\)2−\(r^−r∗\)2\]\)\\displaystyle=\\exp\\left\(\-\\frac\{1\}\{8\}\\left\[\(\\Delta\(\\Delta\-2\\tau\)\+\(\\hat\{r\}\+r^\{\*\}\)\)^\{2\}\-\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\\right\]\\right\)
Then, with probability1−Φ​\(τ\)1\-\\Phi\(\\tau\)over samples of𝐰0\\mathbf\{w\}\_\{0\},

∫01∥∇log⁡pt\(γ^\)​\(𝐰t\)−∇log⁡pt\(γ∗\)​\(𝐰t\)∥2​dt\\displaystyle\\int\_\{0\}^\{1\}\\lVert\\nabla\\log p\_\{t\}^\{\(\\hat\{\\gamma\}\)\}\(\\mathbf\{w\}\_\{t\}\)\-\\nabla\\log p\_\{t\}^\{\(\\gamma^\{\*\}\)\}\(\\mathbf\{w\}\_\{t\}\)\\rVert^\{2\}\\mathrm\{d\}t≥\(γ^−γ∗\)2​ϵ16​β1​γ^​\(1−γ^\)​γ∗​\(1−γ∗\)​τ​τ2\+R1\+12​τ2\+12​R1ln⁡\(Δ\)−ln⁡\(τ\+12​τ2\+R1\)\\displaystyle\\quad\\quad\\geq\\frac\{\(\\hat\{\\gamma\}\-\\gamma^\{\*\}\)^\{2\}\\epsilon\}\{16\\sqrt\{\\beta\_\{1\}\}\\;\\hat\{\\gamma\}\(1\-\\hat\{\\gamma\}\)\\gamma^\{\*\}\(1\-\\gamma^\{\*\}\)\}\\frac\{\\tau\\sqrt\{\\tau^\{2\}\+R\_\{1\}\}\+\\frac\{1\}\{2\}\\tau^\{2\}\+\\frac\{1\}\{2\}R\_\{1\}\}\{\\sqrt\{\\ln\(\\Delta\)\-\\ln\\left\(\\tau\+\\frac\{1\}\{2\}\\sqrt\{\\tau^\{2\}\+R\_\{1\}\}\\right\)\}\}where

R1:=8​ln⁡1ϵ−\(r^−r∗\)2\+\(r^\+r∗\)R\_\{1\}:=\\sqrt\{8\\ln\\frac\{1\}\{\\epsilon\}\-\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\}\+\(\\hat\{r\}\+r^\{\*\}\)

###### Proof\.

This follows from repeating all of our analysis above for thec=0c=0case withγ∗\\gamma^\{\*\}andγ^\\hat\{\\gamma\}replaced with1−γ∗1\-\\gamma^\{\*\}and1−γ^1\-\\hat\{\\gamma\}, respectively\. ∎

Combining[PropositionD\.11](https://arxiv.org/html/2607.15485#A4.Thmtheorem11)and[PropositionD\.12](https://arxiv.org/html/2607.15485#A4.Thmtheorem12), we have by our decomposition ofℓDSM\\ell\_\{\\mathrm\{DSM\}\}from the start of this section that

ℓDSM​\(p\(γ^\),p\(γ∗\)\)≥C​\(γ^−γ∗\)2\\ell\_\{\\mathrm\{DSM\}\}\(p^\{\(\\hat\{\\gamma\}\)\},p^\{\(\\gamma^\{\*\}\)\}\)\\geq C\(\\hat\{\\gamma\}\-\\gamma^\{\*\}\)^\{2\}where

C=ϵ​\(1−Φ​\(τ\)\)16​β1​γ^​\(1−γ^\)​γ∗​\(1−γ∗\)​\[\(1−γ∗\)​τ​τ2\+R0\+12​τ2\+12​R0ln⁡\(Δ\)−ln⁡\(τ\+12​τ2\+R0\)\+γ∗​τ​τ2\+R1\+12​τ2\+12​R1ln⁡\(Δ\)−ln⁡\(τ\+12​τ2\+R1\)\]C=\\frac\{\\epsilon\(1\-\\Phi\(\\tau\)\)\}\{16\\sqrt\{\\beta\_\{1\}\}\\;\\hat\{\\gamma\}\(1\-\\hat\{\\gamma\}\)\\gamma^\{\*\}\(1\-\\gamma^\{\*\}\)\}\\left\[\(1\-\\gamma^\{\*\}\)\\frac\{\\tau\\sqrt\{\\tau^\{2\}\+R\_\{0\}\}\+\\frac\{1\}\{2\}\\tau^\{2\}\+\\frac\{1\}\{2\}R\_\{0\}\}\{\\sqrt\{\\ln\(\\Delta\)\-\\ln\\left\(\\tau\+\\frac\{1\}\{2\}\\sqrt\{\\tau^\{2\}\+R\_\{0\}\}\\right\)\}\}\+\\gamma^\{\*\}\\frac\{\\tau\\sqrt\{\\tau^\{2\}\+R\_\{1\}\}\+\\frac\{1\}\{2\}\\tau^\{2\}\+\\frac\{1\}\{2\}R\_\{1\}\}\{\\sqrt\{\\ln\(\\Delta\)\-\\ln\\left\(\\tau\+\\frac\{1\}\{2\}\\sqrt\{\\tau^\{2\}\+R\_\{1\}\}\\right\)\}\}\\right\]for allτ≥Δ​α1\\tau\\geq\\Delta\\sqrt\{\\alpha\_\{1\}\}andϵ∈\[ϵmin,ϵmax\]\\epsilon\\in\[\\epsilon\_\{\\mathrm\{min\}\},\\epsilon\_\{\\mathrm\{\\max\}\}\], where

ϵmax\\displaystyle\\epsilon\_\{\\mathrm\{max\}\}=min⁡\{ϵmax0,ϵmax1\}=exp⁡\(−18​\[\(r^−r∗\)2\+max⁡\{\(r^\+r∗−τ2\)\+2,\(−\(r^\+r∗\)−τ2\)\+2\}\]\)\\displaystyle=\\min\\\{\\epsilon\_\{\\mathrm\{max\}\}^\{0\},\\epsilon\_\{\\mathrm\{max\}\}^\{1\}\\\}=\\exp\\left\(\-\\frac\{1\}\{8\}\\left\[\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\+\\max\\\{\\left\(\\hat\{r\}\+r^\{\*\}\-\\tau^\{2\}\\right\)\_\{\+\}^\{2\},\\left\(\-\(\\hat\{r\}\+r^\{\*\}\)\-\\tau^\{2\}\\right\)\_\{\+\}^\{2\}\\\}\\right\]\\right\)ϵmin\\displaystyle\\epsilon\_\{\\mathrm\{min\}\}=max⁡\{ϵmin0,ϵmin1\}=exp⁡\(−18​\[min⁡\{\(Δ​\(Δ−2​τ\)\+\(r^\+r∗\)\)2,\(Δ​\(Δ−2​τ\)−\(r^\+r∗\)\)2\}\+\(r^−r∗\)2\]\)\\displaystyle=\\max\\\{\\epsilon\_\{\\mathrm\{min\}\}^\{0\},\\epsilon\_\{\\mathrm\{min\}\}^\{1\}\\\}=\\exp\\left\(\-\\frac\{1\}\{8\}\\left\[\\min\\\{\(\\Delta\(\\Delta\-2\\tau\)\+\(\\hat\{r\}\+r^\{\*\}\)\)^\{2\},\(\\Delta\(\\Delta\-2\\tau\)\-\(\\hat\{r\}\+r^\{\*\}\)\)^\{2\}\\\}\+\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\\right\]\\right\)
To simplify this expression, note that if

g​\(x\):=x​\(τ\+12​x\)ln⁡\(Δ\)−ln⁡\(τ\+12​x\)g\(x\):=\\frac\{x\(\\tau\+\\frac\{1\}\{2\}x\)\}\{\\sqrt\{\\ln\(\\Delta\)\-\\ln\(\\tau\+\\frac\{1\}\{2\}x\)\}\}then direct calculation shows that ifh​\(x\):=ln⁡\(Δ\)−ln⁡\(τ\+12​x\)h\(x\):=\\ln\(\\Delta\)\-\\ln\(\\tau\+\\frac\{1\}\{2\}x\),

g′​\(x\)=x4\+\(τ\+x\)​h​\(x\)h​\(x\)3/2\>0g^\{\\prime\}\(x\)=\\frac\{\\frac\{x\}\{4\}\+\(\\tau\+x\)h\(x\)\}\{h\(x\)^\{3/2\}\}\>0
Therefore,ggis increasing, meaning

C\\displaystyle C=ϵ​\(1−Φ​\(τ\)\)16​β1​γ^​\(1−γ^\)​γ∗​\(1−γ∗\)​\[\(1−γ∗\)​g​\(τ2\+R0\)\+γ∗​g​\(τ2\+R1\)\]\\displaystyle=\\frac\{\\epsilon\(1\-\\Phi\(\\tau\)\)\}\{16\\sqrt\{\\beta\_\{1\}\}\\;\\hat\{\\gamma\}\(1\-\\hat\{\\gamma\}\)\\gamma^\{\*\}\(1\-\\gamma^\{\*\}\)\}\\left\[\(1\-\\gamma^\{\*\}\)g\\left\(\\sqrt\{\\tau^\{2\}\+R\_\{0\}\}\\right\)\+\\gamma^\{\*\}g\\left\(\\sqrt\{\\tau^\{2\}\+R\_\{1\}\}\\right\)\\right\]≥ϵ​\(1−Φ​\(τ\)\)16​β1​γ^​\(1−γ^\)​γ∗​\(1−γ∗\)​g​\(min⁡\{τ2\+R0,τ2\+R1\}\)\\displaystyle\\geq\\frac\{\\epsilon\(1\-\\Phi\(\\tau\)\)\}\{16\\sqrt\{\\beta\_\{1\}\}\\;\\hat\{\\gamma\}\(1\-\\hat\{\\gamma\}\)\\gamma^\{\*\}\(1\-\\gamma^\{\*\}\)\}g\\left\(\\min\\\{\\sqrt\{\\tau^\{2\}\+R\_\{0\}\},\\sqrt\{\\tau^\{2\}\+R\_\{1\}\}\\\}\\right\)=ϵ​\(1−Φ​\(τ\)\)16​β1​γ^​\(1−γ^\)​γ∗​\(1−γ∗\)​g​\(τ2\+8​ln⁡1ϵ−\(r^−r∗\)2−\|r^−r∗\|\)\\displaystyle=\\frac\{\\epsilon\(1\-\\Phi\(\\tau\)\)\}\{16\\sqrt\{\\beta\_\{1\}\}\\;\\hat\{\\gamma\}\(1\-\\hat\{\\gamma\}\)\\gamma^\{\*\}\(1\-\\gamma^\{\*\}\)\}g\\left\(\\sqrt\{\\tau^\{2\}\+\\sqrt\{8\\ln\\frac\{1\}\{\\epsilon\}\-\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\}\-\\lvert\\hat\{r\}\-r^\{\*\}\\rvert\}\\right\)
LetRmin:=8​ln⁡1ϵ−\(r^−r∗\)2−\|r^−r∗\|R\_\{\\mathrm\{min\}\}:=\\sqrt\{8\\ln\\frac\{1\}\{\\epsilon\}\-\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\}\-\\lvert\\hat\{r\}\-r^\{\*\}\\rvert\. To obtain the bound from the body of the paper, we now specify values forτ\\tauandϵ\\epsilonwhich simplify the expression\. First, we will setτ\\tauto be its minimum value:τ=Δ​α1\\tau=\\Delta\\sqrt\{\\alpha\_\{1\}\}\. We choose this because numerically, it seems to be the value ofτ\\tauthat maximizes the expression\. This gives us

C≥ϵ​\(1−Φ​\(Δ​α1\)\)16​β1​γ^​\(1−γ^\)​γ∗​\(1−γ∗\)​Δ2​α1\+Rmin​\(Δ​α1\+12​Δ2​α1\+Rmin\)log⁡\(Δ\)−log⁡\(α1\)−log⁡\(Δ\+12​Δ2\+R0α1\)C\\geq\\frac\{\\epsilon\(1\-\\Phi\(\\Delta\\sqrt\{\\alpha\_\{1\}\}\)\)\}\{16\\sqrt\{\\beta\_\{1\}\}\\;\\hat\{\\gamma\}\(1\-\\hat\{\\gamma\}\)\\gamma^\{\*\}\(1\-\\gamma^\{\*\}\)\}\\frac\{\\sqrt\{\\Delta^\{2\}\\alpha\_\{1\}\+R\_\{\\mathrm\{min\}\}\}\\left\(\\Delta\\sqrt\{\\alpha\_\{1\}\}\+\\frac\{1\}\{2\}\\sqrt\{\\Delta^\{2\}\\alpha\_\{1\}\+R\_\{\\mathrm\{min\}\}\}\\right\)\}\{\\sqrt\{\\log\(\\Delta\)\-\\log\\left\(\\sqrt\{\\alpha\_\{1\}\}\\right\)\-\\log\\left\(\\Delta\+\\frac\{1\}\{2\}\\sqrt\{\\Delta^\{2\}\+\\frac\{R\_\{0\}\}\{\\alpha\_\{1\}\}\}\\right\)\}\}
Noting thatlog⁡\(Δ\)<log⁡\(Δ\+12​Δ2\+R0α1\)\\log\(\\Delta\)<\\log\\left\(\\Delta\+\\frac\{1\}\{2\}\\sqrt\{\\Delta^\{2\}\+\\frac\{R\_\{0\}\}\{\\alpha\_\{1\}\}\}\\right\), the denominator of the second term is less than−log⁡\(α1\)=14​β1\\sqrt\{\-\\log\(\\sqrt\{\\alpha\_\{1\}\}\)\}=\\frac\{1\}\{4\}\\beta\_\{1\}\(recallαt=exp⁡\(−∫0tβs​ds\)=exp⁡\(−12​β1​t2\)\\alpha\_\{t\}=\\exp\\left\(\-\\int\_\{0\}^\{t\}\\beta\_\{s\}\\mathrm\{d\}s\\right\)=\\exp\\left\(\-\\frac\{1\}\{2\}\\beta\_\{1\}t^\{2\}\\right\)\)\. Additionally, note that the numerator of the second term is an increasing function ofΔ\\Delta, meaning we can lower\-bound it by settingΔ=0\\Delta=0\. Combining these gives us

C≥1−Φ​\(Δ​α1\)8​β1​γ^​\(1−γ^\)​γ∗​\(1−γ∗\)​ϵ​RminC\\geq\\frac\{1\-\\Phi\(\\Delta\\sqrt\{\\alpha\_\{1\}\}\)\}\{8\\beta\_\{1\}\\;\\hat\{\\gamma\}\(1\-\\hat\{\\gamma\}\)\\gamma^\{\*\}\(1\-\\gamma^\{\*\}\)\}\\epsilon\\,R\_\{\\mathrm\{min\}\}
Finally, direct calculation shows thatϵ​Rmin\\epsilon\\,R\_\{\\mathrm\{min\}\}is maximized whenϵ=exp⁡\(−18​\(\(r^−r∗\)2\+δ2\)\)\\epsilon=\\exp\\left\(\-\\frac\{1\}\{8\}\\left\(\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\+\\delta^\{2\}\\right\)\\right\)whereδ=\|r^−r∗\|\+\(r^−r∗\)2\+162\\delta=\\frac\{\|\\hat\{r\}\-r^\{\*\}\|\+\\sqrt\{\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\+16\}\}\{2\}, giving

ϵ​Rmin=\(r^−r∗\)2\+16−\|r^−r∗\|2​exp⁡\(−18​\(\(r^−r∗\)2\+δ2\)\)\\epsilon\\,R\_\{\\mathrm\{min\}\}=\\frac\{\\sqrt\{\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\+16\}\-\|\\hat\{r\}\-r^\{\*\}\|\}\{2\}\\exp\\left\(\-\\frac\{1\}\{8\}\\left\(\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\+\\delta^\{2\}\\right\)\\right\)
Noteδ≤\|r^−r∗\|\+2\\delta\\leq\\lvert\\hat\{r\}\-r^\{\*\}\\rvert\+2\(becausea2\+b2≤\|a\|\+\|b\|\\sqrt\{a^\{2\}\+b^\{2\}\}\\leq\|a\|\+\|b\|\), andδ2≤\|r^−r∗\|2\+2​\|r^−r∗\|\+4\\delta^\{2\}\\leq\\lvert\\hat\{r\}\-r^\{\*\}\\rvert^\{2\}\+2\\lvert\\hat\{r\}\-r^\{\*\}\\rvert\+4\. Applying this, and the fact thatγ^​\(1−γ^\)≤14\\hat\{\\gamma\}\(1\-\\hat\{\\gamma\}\)\\leq\\frac\{1\}\{4\}

C\\displaystyle C≥\(1−Φ​\(Δ​α1\)\)​\(\(r^−r∗\)2\+16−\|r^−r∗\|\)4​β1​γ∗​\(1−γ∗\)​exp⁡\(−14​\(\(r^−r∗\)2\+\|r^−r∗\|\)−12\)\\displaystyle\\geq\\frac\{\\left\(1\-\\Phi\(\\Delta\\sqrt\{\\alpha\_\{1\}\}\)\\right\)\\left\(\\sqrt\{\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\+16\}\-\|\\hat\{r\}\-r^\{\*\}\|\\right\)\}\{4\\beta\_\{1\}\\;\\gamma^\{\*\}\(1\-\\gamma^\{\*\}\)\}\\exp\\left\(\-\\frac\{1\}\{4\}\\left\(\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\+\\lvert\\hat\{r\}\-r^\{\*\}\\rvert\\right\)\-\\frac\{1\}\{2\}\\right\)=\(1−Φ​\(Δ​α1\)\)​\(\(r^−r∗\)2\+16−\|r^−r∗\|\)4​e​β1​γ∗​\(1−γ∗\)​exp⁡\(−14​\(\(r^−r∗\)2\+\|r^−r∗\|\)\)\\displaystyle=\\frac\{\\left\(1\-\\Phi\(\\Delta\\sqrt\{\\alpha\_\{1\}\}\)\\right\)\\left\(\\sqrt\{\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\+16\}\-\|\\hat\{r\}\-r^\{\*\}\|\\right\)\}\{4\\sqrt\{e\}\\beta\_\{1\}\\;\\gamma^\{\*\}\(1\-\\gamma^\{\*\}\)\}\\exp\\left\(\-\\frac\{1\}\{4\}\\left\(\(\\hat\{r\}\-r^\{\*\}\)^\{2\}\+\\lvert\\hat\{r\}\-r^\{\*\}\\rvert\\right\)\\right\)
Finally, forγ^∈\[a,1−a\]\\hat\{\\gamma\}\\in\[a,1\-a\],\|r^−r∗\|≤log⁡1−aa\+\|log⁡1−γ∗γ∗\|\\lvert\\hat\{r\}\-r^\{\*\}\\rvert\\leq\\log\\frac\{1\-a\}\{a\}\+\\lvert\\log\\frac\{1\-\\gamma^\{\*\}\}\{\\gamma^\{\*\}\}\\rvert\. Our expression is a decreasing function of\|r^−r∗\|\\lvert\\hat\{r\}\-r^\{\*\}\\rvert, and thus for allγ^∈\[a,1−a\]\\hat\{\\gamma\}\\in\[a,1\-a\], if we letk:=log⁡1−aa\+\|log⁡1−γ∗γ∗\|k:=\\log\\frac\{1\-a\}\{a\}\+\\lvert\\log\\frac\{1\-\\gamma^\{\*\}\}\{\\gamma^\{\*\}\}\\rvert

C≥\(1−Φ​\(Δ​α1\)\)​\(k2\+16−k\)7​β1​γ∗​\(1−γ∗\)​exp⁡\(−\(k2\+k\)/4\)C\\geq\\frac\{\\left\(1\-\\Phi\(\\Delta\\sqrt\{\\alpha\_\{1\}\}\)\\right\)\\left\(\\sqrt\{k^\{2\}\+16\}\-k\\right\)\}\{7\\beta\_\{1\}\\;\\gamma^\{\*\}\(1\-\\gamma^\{\*\}\)\}\\exp\\left\(\-\\left\(k^\{2\}\+k\\right\)/4\\right\)where we additionally used4​e<74\\sqrt\{e\}<7\. Bringing this all together,

minγ^∈\[a,1−a\]⁡ℓDSM​\(p\(γ^\),p\(γ∗\)\)\(γ^−γ∗\)2\>\(1−Φ​\(∥μ1−μ0∥​α1\)\)​\(k2\+16−k\)7​β1​γ∗​\(1−γ∗\)​exp⁡\(−\(k2\+k\)/4\)\\displaystyle\\min\_\{\\hat\{\\gamma\}\\in\[a,1\-a\]\}\\frac\{\\ell\_\{\\mathrm\{DSM\}\}\(p^\{\(\\hat\{\\gamma\}\)\},p^\{\(\\gamma^\{\*\}\)\}\)\}\{\(\\hat\{\\gamma\}\-\\gamma^\{\*\}\)^\{2\}\}\>\\frac\{\(1\-\\Phi\(\\lVert\\mu\_\{1\}\-\\mu\_\{0\}\\rVert\\sqrt\{\\alpha\_\{1\}\}\)\)\\left\(\\sqrt\{k^\{2\}\+16\}\-k\\right\)\}\{7\\beta\_\{1\}\\gamma^\{\*\}\(1\-\\gamma^\{\*\}\)\}\\exp\\left\(\-\(k^\{2\}\+k\)/4\\right\)\(44\)∎

## Appendix EAdditional numerical experiments

### E\.1Noise schedules and time reparameterization

In this work, we reference two noise schedules used in practice\. The linear noise schedule is defined byβt:=β1​t\\beta\_\{t\}:=\\beta\_\{1\}tfor a constantβ1\>0\\beta\_\{1\}\>0, such thatαt=exp⁡\(−β12​t2\)\\alpha\_\{t\}=\\exp\\left\(\-\\frac\{\\beta\_\{1\}\}\{2\}t^\{2\}\\right\)\. The squared\-cosine noise schedule\[[20](https://arxiv.org/html/2607.15485#bib.bib85)\]

α¯t=cos\(t\+s1\+s⋅π2\)2\\bar\{\\alpha\}\_\{t\}=\\cos\\left\(\\frac\{t\+s\}\{1\+s\}\\cdot\\frac\{\\pi\}\{2\}\\right\)^\{2\}andαt=α¯tα¯0\\alpha\_\{t\}=\\frac\{\\bar\{\\alpha\}\_\{t\}\}\{\\bar\{\\alpha\}\_\{0\}\}fors=0\.008s=0\.008\.

We want to analyze the effect of artificially reducing the sensitivity index on the estimation ofγ∗\\gamma^\{\*\}\. To do this, we can reparameterize the time in the forward process ast=ta,b,v​\(τ\)t=t\_\{a,b,v\}\(\\tau\), where

ta,b,v​\(τ\)=\{τ0≤τ<aa\+v​\(τ−a\)a≤τ<a\+b−avb\+1−b1−\(a\+b−av\)​\(t−a−b−av\)a\+b−av≤t≤1t\_\{a,b,v\}\(\\tau\)=\\begin\{cases\}\\tau&0\\leq\\tau<a\\\\ a\+v\(\\tau\-a\)&a\\leq\\tau<a\+\\frac\{b\-a\}\{v\}\\\\ b\+\\frac\{1\-b\}\{1\-\\left\(a\+\\frac\{b\-a\}\{v\}\\right\)\}\\left\(t\-a\-\\frac\{b\-a\}\{v\}\\right\)&a\+\\frac\{b\-a\}\{v\}\\leq t\\leq 1\\end\{cases\}
Under this parameterization, if we let𝐲τ:=𝐱t​\(τ\)\\mathbf\{y\}\_\{\\tau\}:=\\mathbf\{x\}\_\{t\(\\tau\)\}where𝐱t\\mathbf\{x\}\_\{t\}follows the Variance Preserving SDE with noise schedule\{αt\}\\\{\\alpha\_\{t\}\\\}, then𝐲t\\mathbf\{y\}\_\{t\}satisfies

d​𝐲τ\\displaystyle\\mathrm\{d\}\\mathbf\{y\}\_\{\\tau\}=−βt​\(τ\)2​𝐲τ​d​td​τ​d​τ\+βt​\(τ\)⋅d​td​τ​d​Wt\\displaystyle=\-\\frac\{\\beta\_\{t\(\\tau\)\}\}\{2\}\\mathbf\{y\}\_\{\\tau\}\\frac\{\\mathrm\{d\}t\}\{\\mathrm\{d\}\\tau\}\\mathrm\{d\}\\tau\+\\sqrt\{\\beta\_\{t\(\\tau\)\}\}\\cdot\\sqrt\{\\frac\{\\mathrm\{d\}t\}\{\\mathrm\{d\}\\tau\}\}\\mathrm\{d\}W\_\{t\}\(45\)=−β~τ2​𝐲τ​d​τ\+β~τ​d​Wt\\displaystyle=\-\\frac\{\\tilde\{\\beta\}\_\{\\tau\}\}\{2\}\\mathbf\{y\}\_\{\\tau\}\\mathrm\{d\}\\tau\+\\sqrt\{\\tilde\{\\beta\}\_\{\\tau\}\}\\mathrm\{d\}W\_\{t\}\(46\)whereβ~τ:=βt​\(τ\)⋅d​td​τ\\tilde\{\\beta\}\_\{\\tau\}:=\\beta\_\{t\(\\tau\)\}\\cdot\\frac\{\\mathrm\{d\}t\}\{\\mathrm\{d\}\\tau\}\. In particular, reparameterizing time is equivalent to choosing a different noise schedule\{α~τ\}\\\{\\tilde\{\\alpha\}\_\{\\tau\}\\\}\. Forv≥1v\\geq 1,ta,b,vt\_\{a,b,v\}maps the interval\[a,b\]\[a,b\]to\[a,a\+b−av\]\\left\[a,a\+\\frac\{b\-a\}\{v\}\\right\], scaling the length of the interval by1/v1/v\. To reduce the value ofL​\(γ∗\)L\(\\gamma^\{\*\}\), we let\[a,b\]\[a,b\]contain the region of time whereℓSM​\(pt\(γ\),pt\(γ∗\)\)\(γ−γ∗\)2\>ϵ\\frac\{\\ell\_\{\\mathrm\{SM\}\}\(p\_\{t\}^\{\(\\gamma\)\},p\_\{t\}^\{\(\\gamma^\{\*\}\)\}\)\}\{\(\\gamma\-\\gamma^\{\*\}\)^\{2\}\}\>\\epsilonfor fixedϵ\>0\\epsilon\>0and increase the value ofvv\.

### E\.2Neural network experiment details

For the score models𝐬𝜽\\mathbf\{s\}\_\{\\bm\{\\theta\}\}in our experiments, we use deep ResNets\[[21](https://arxiv.org/html/2607.15485#bib.bib56)\]with width equal to the ambient dimension and depth88\. Unless otherwise specified, we use a linear noise schedule withβ1=20\\beta\_\{1\}=20and a time discretization ofT=4000T=4000steps on the interval\[10−5,1\]\[10^\{\-5\},1\]for training and\[10−3,1\]\[10^\{\-3\},1\]for sampling, following\[[53](https://arxiv.org/html/2607.15485#bib.bib24)\]\.

We train our models using the loss described in this paper \(and\[[53](https://arxiv.org/html/2607.15485#bib.bib24)\]\) using AdaBelief\[[64](https://arxiv.org/html/2607.15485#bib.bib79)\]with a linear warmup and cosine decay learning rate\. We use JAX\[[9](https://arxiv.org/html/2607.15485#bib.bib81)\]and Equinox\[[26](https://arxiv.org/html/2607.15485#bib.bib80)\]for our neural network code, SciPy\[[59](https://arxiv.org/html/2607.15485#bib.bib84)\]for additional machine learning algorithms, Matplotlib\[[23](https://arxiv.org/html/2607.15485#bib.bib82)\]for figure generation, and Weights and Biases\[[5](https://arxiv.org/html/2607.15485#bib.bib83)\]for experiment tracking\.555We will make our code available upon publication\.

### E\.3MNIST ablation study details

The architecture of our autoencoder is simple: two convolutional layers followed by a dense layer for the encoder, and the reverse for the decoder\. For our MNIST experiments, we use a latent dimension of1212\. We train the autoencoder on all examples of11s and88s in MNIST\. We additionally train an MNIST classifier on only11s and88s using the architecture in the Equinox documentation\[[26](https://arxiv.org/html/2607.15485#bib.bib80)\]\.

Unlike in the case of Gaussian mixture models, we do not know the analytical score function for the latents of our autoencoder, and thus estimation ofℓDSM​\(pt\(γ\),pt\(γ∗\)\)\\ell\_\{\\mathrm\{DSM\}\}\(p\_\{t\}^\{\(\\gamma\)\},p\_\{t\}^\{\(\\gamma^\{\*\}\)\}\)andL​\(γ∗\)L\(\\gamma^\{\*\}\)analytically is intractable\. Instead, forγ∗∈\{0\.2,0\.22,…,0\.8\}\\gamma^\{\*\}\\in\\\{0\.2,0\.22,\.\.\.,0\.8\\\}, we train diffusion models𝐬𝜽γ\\mathbf\{s\}\_\{\{\\bm\{\\theta\}\}\_\{\\gamma\}\}on data fromp\(γ\)p^\{\(\\gamma\)\}as approximations for the analytical scores using a linear noise schedule\. We keep the number of training examples constant across all experiments \(5800, the smaller of the counts of 1s and 8s in MNIST\) and train the models identically\. For eacht∈\[0,1\]t\\in\[0,1\], we then have the estimate

ℓDSM​\(pt\(γ\),pt\(γ∗\)\)≈𝔼𝐱t∼pt\(γ∗\)​\[λt​∥𝐬𝜽γ∗​\(𝐱t,t\)−𝐬𝜽γ′​\(𝐱t,t\)∥2\]\\ell\_\{\\mathrm\{DSM\}\}\(p\_\{t\}^\{\(\\gamma\)\},p\_\{t\}^\{\(\\gamma^\{\*\}\)\}\)\\approx\\mathbb\{E\}\_\{\\mathbf\{x\}\_\{t\}\\sim p\_\{t\}^\{\(\\gamma^\{\*\}\)\}\}\\left\[\\lambda\_\{t\}\\lVert\\mathbf\{s\}\_\{\{\\bm\{\\theta\}\}\_\{\\gamma^\{\*\}\}\}\(\\mathbf\{x\}\_\{t\},t\)\-\\mathbf\{s\}\_\{\{\\bm\{\\theta\}\}\_\{\\gamma^\{\\prime\}\}\}\(\\mathbf\{x\}\_\{t\},t\)\\rVert^\{2\}\\right\]
We generate all samples used in[Figure4](https://arxiv.org/html/2607.15485#S5.F4)using𝐬𝜽0\.6\\mathbf\{s\}\_\{\{\\bm\{\\theta\}\}\_\{0\.6\}\}, varying the noise schedule used for sampling\.

### E\.4DSSI for other parameters

Here, we apply our DSSI framework to parameters other than the mixture weight in a Gaussian model\. In what follows, we will letΣ0\\Sigma\_\{0\}andΣ1\\Sigma\_\{1\}be fixed random positive definite matrices, such that

Σi=A⊤​A\+10−8​I,Ai​j∼𝒩​\(0,1\)\\Sigma\_\{i\}=A^\{\\top\}A\+10^\{\-8\}I,\\;\\;A\_\{ij\}\\sim\\mathcal\{N\}\(0,1\)fori∈\{0,1\}i\\in\\\{0,1\\\}and use a linear noise schedule\. We focus here on one\-dimensional families of distributions for consistency with the body of the paper, although we note that our framework can be extended to multi\-parameter recovery directly\. For this section, we examine GMMs in dimensiond=10d=10\.

For our analysis of the sensitivity ofℓDSM\\ell\_\{\\mathrm\{DSM\}\}to the modes of the mixture components, we consider the one\-dimensional family of distributions

p\(δ\):=0\.6​𝒩​\(𝟎,𝚺0\)\+0\.4​𝒩​\(12​δd​𝟏,𝚺1\)p^\{\(\\delta\)\}:=0\.6\\;\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{\\Sigma\}\_\{0\}\)\+0\.4\\;\\mathcal\{N\}\\left\(\\frac\{12\\;\\delta\}\{\\sqrt\{d\}\}\\mathbf\{1\},\\mathbf\{\\Sigma\}\_\{1\}\\right\)in dimensiond=10d=10where we rangeδ\\deltabetween0and11\.[Figure6\(a\)](https://arxiv.org/html/2607.15485#A5.F6.sf1)showsℓSM​\(pt\(δ\),pt\(δ∗\)\)\\ell\_\{\\mathrm\{SM\}\}\(p\_\{t\}^\{\(\\delta\)\},p\_\{t\}^\{\(\\delta^\{\*\}\)\}\)as a function ofδ\\deltaandtt, and[Figure6\(b\)](https://arxiv.org/html/2607.15485#A5.F6.sf2)showsℓDSM​\(pt\(δ\),pt\(δ∗\)\)\(δ−δ∗\)2\\frac\{\\ell\_\{\\mathrm\{DSM\}\}\(p\_\{t\}^\{\(\\delta\)\},p\_\{t\}^\{\(\\delta^\{\*\}\)\}\)\}\{\(\\delta\-\\delta^\{\*\}\)^\{2\}\}as a function ofδ\\deltafor multiple values ofδ∗\\delta^\{\*\}\. We see that, as with the mixture weight,ℓSM​\(pt\(δ\),pt\(δ∗\)\)\\ell\_\{\\mathrm\{SM\}\}\(p\_\{t\}^\{\(\\delta\)\},p\_\{t\}^\{\(\\delta^\{\*\}\)\}\)is nonnegligible only for intermediate times in the forward process\. However, note that the minimum value ofL​\(δ∗\)L\(\\delta^\{\*\}\)we see here is greater than22\(compared to0\.370\.37in[Figure2\(b\)](https://arxiv.org/html/2607.15485#S3.F2.sf2)\), meaning that our theory implies a5\.4×5\.4\\timesstronger control on the error in estimationμ\\muusing a diffusion model trained to the same level of error underℓDSM\\ell\_\{\\mathrm\{DSM\}\}\.

![Refer to caption](https://arxiv.org/html/2607.15485v1/x19.png)\(a\)
![Refer to caption](https://arxiv.org/html/2607.15485v1/x20.png)\(b\)

Figure 6:Sensitivity ofℓDSM\\ell\_\{\\mathrm\{DSM\}\}to means of mixture components\. \(a\)ℓSM​\(pt\(δ\),pt\(δ∗\)\)\\ell\_\{\\mathrm\{SM\}\}\(p\_\{t\}^\{\(\\delta\)\},p\_\{t\}^\{\(\\delta^\{\*\}\)\}\)as a function ofttandδ\\delta\. \(b\)ℓDSM​\(p\(δ\),p\(δ∗\)\)\(δ−δ∗\)2\\frac\{\\ell\_\{\\mathrm\{DSM\}\}\(p^\{\(\\delta\)\},p^\{\(\\delta^\{\*\}\)\}\)\}\{\(\\delta\-\\delta^\{\*\}\)^\{2\}\}as a function ofδ\\delta\.We now apply our DSSI framework to covariance matrix recovery\. Consider the one\-dimensional family of distributions

p\(κ\):=0\.6​𝒩​\(𝟎,𝐈\)\+0\.4​𝒩​\(9\.0d​𝟏,𝚺κ\)p^\{\(\\kappa\)\}:=0\.6\\;\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)\+0\.4\\;\\mathcal\{N\}\\left\(\\frac\{9\.0\}\{\\sqrt\{d\}\}\\mathbf\{1\},\\mathbf\{\\Sigma\}\_\{\\kappa\}\\right\)in dimensiond=10d=10where we rangeδ\\deltabetween0and11\. Here, we define

𝚺κ:=𝚺01/2​\(𝚺0−1/2​𝚺1​𝚺0−1/2\)κ​𝚺01/2\\mathbf\{\\Sigma\}\_\{\\kappa\}:=\\mathbf\{\\Sigma\}\_\{0\}^\{1/2\}\\left\(\\mathbf\{\\Sigma\}\_\{0\}^\{\-1/2\}\\mathbf\{\\Sigma\}\_\{1\}\\mathbf\{\\Sigma\}\_\{0\}^\{\-1/2\}\\right\)^\{\\kappa\}\\mathbf\{\\Sigma\}\_\{0\}^\{1/2\}to be the geodesic interpolating between𝚺0\\mathbf\{\\Sigma\}\_\{0\}and𝚺1\\mathbf\{\\Sigma\}\_\{1\}\.[Figure7\(a\)](https://arxiv.org/html/2607.15485#A5.F7.sf1)showsℓSM​\(pt\(κ\),pt\(κ∗\)\)\\ell\_\{\\mathrm\{SM\}\}\(p\_\{t\}^\{\(\\kappa\)\},p\_\{t\}^\{\(\\kappa^\{\*\}\)\}\)as a function ofκ\\kappaandtt, and[Figure7\(b\)](https://arxiv.org/html/2607.15485#A5.F7.sf2)showsℓDSM​\(p\(κ\),p\(κ∗\)\)\(κ−κ∗\)2\\frac\{\\ell\_\{\\mathrm\{DSM\}\}\(p^\{\(\\kappa\)\},p^\{\(\\kappa^\{\*\}\)\}\)\}\{\(\\kappa\-\\kappa^\{\*\}\)^\{2\}\}as a function ofκ\\kappa\.

Again,ℓSM​\(pt\(δ\),pt\(δ∗\)\)\\ell\_\{\\mathrm\{SM\}\}\(p\_\{t\}^\{\(\\delta\)\},p\_\{t\}^\{\(\\delta^\{\*\}\)\}\)is nonnegligable for intermediate times in the forward process\. We additionally see thatL​\(κ∗\)L\(\\kappa^\{\*\}\)is greater than0\.60\.6uniformly acrossκ\\kappa\.

![Refer to caption](https://arxiv.org/html/2607.15485v1/x21.png)\(a\)
![Refer to caption](https://arxiv.org/html/2607.15485v1/x22.png)\(b\)

Figure 7:Sensitivity ofℓDSM\\ell\_\{\\mathrm\{DSM\}\}to covariance matrix of mixture components\. \(a\)ℓSM​\(pt\(κ\),pt\(κ∗\)\)\\ell\_\{\\mathrm\{SM\}\}\(p\_\{t\}^\{\(\\kappa\)\},p\_\{t\}^\{\(\\kappa^\{\*\}\)\}\)as a function ofttandκ\\kappa\. \(b\)ℓDSM​\(p\(κ\),p\(κ∗\)\)\(κ−κ∗\)2\\frac\{\\ell\_\{\\mathrm\{DSM\}\}\(p^\{\(\\kappa\)\},p^\{\(\\kappa^\{\*\}\)\}\)\}\{\(\\kappa\-\\kappa^\{\*\}\)^\{2\}\}as a function ofκ\\kappa\.

Similar Articles

Elucidating the SNR-t Bias of Diffusion Probabilistic Models

Hugging Face Daily Papers

This paper identifies a Signal-to-Noise Ratio timestep (SNR-t) bias in diffusion probabilistic models during inference, where SNR-timestep alignment from training is disrupted at inference time. The authors propose a differential correction method that decomposes samples into frequency components and corrects each separately, improving generation quality across models like IDDPM, ADM, DDIM, EDM, and FLUX with minimal computational overhead.