Score-based Outlier Generation via Controlling the Radon-Nikodym Derivative
Summary
The paper introduces a method for controlled outlier generation using diffusion models by manipulating likelihood through Radon-Nikodym derivatives, enabling the creation of low-likelihood samples without retraining.
View Cached Full Text
Cached at: 09/14/26, 08:34 AM
# Score-based Outlier Generationvia Controlling the Radon-Nikodym Derivative
Source: [https://arxiv.org/html/2609.12113](https://arxiv.org/html/2609.12113)
Tristan MilneKry Yik\-Chau LuiStephanie HazlewoodJun Liu††thanks:This research was supported by Mitacs and the Royal Bank of Canada under the Mitacs Accelerate program award, Application Reference IT49478\.††thanks:Amartya Mukherjee and Jun Liu are with the Department of Applied Mathematics, University of Waterloo, Waterloo, Ontario, Canada N2L 3G1 \(email:\(a29mukhe,j\.liu\)@uwaterloo\.ca\)\.††thanks:Tristan Milne, Stephanie Hazlewood, and Kry Yik\-Chau Lui are with the Royal Bank of Canada, 1 Place Ville Marie, Montreal, Quebec, H3C 3A9 \(email:\(tristan\.milne,stephanie\.hazlewood\)@rbc\.com, yikchau\.y\.lui@borealisai\.com\)\.
###### Abstract
Outliers are important for stress\-testing algorithms and understanding system behaviour under rare conditions\. Despite being commonly described as low\-likelihood events, existing generative approaches rarely control likelihood explicitly\. In this work, we introduce a measure\-theoretic notion of outliers based on the distribution of log\-likelihood values, which is guaranteed to assign higher probability mass to low\-likelihood events with a specifiable magnitude\. Building on this formulation, we derive how likelihood reweighting modifies the diffusion score and use this relation to motivate a controlled modification of the reverse\-time dynamics\. In particular, likelihood reweighting implies a scaling of the score function with a control term derived from the Radon\-Nikodym derivative of the likelihood distributions\. Correspondingly, the updated score function can be obtained with no retraining of the diffusion model\. We exploit the Ornstein–Uhlenbeck semigroup underlying diffusion models to motivate an exponentially interpolated controller which approximates the true control\. Experiments demonstrate controlled generation of low\-likelihood samples while remaining consistent with the data geometry\.
Outliers play an important role in evaluating the reliability of algorithms and decision\-making systems\. In many applications, it is important to understand how systems behave under rare or atypical conditions that deviate from the patterns commonly observed in training or historical data\. For example, in time series such as financial data\[[13](https://arxiv.org/html/2609.12113#bib.bib12)\], outliers model how capital markets respond to anomalous shifts in the financial landscape\. In tabular data such as electronic health records\[[18](https://arxiv.org/html/2609.12113#bib.bib11)\], outliers are crucial for studying patients who deviate from typical disease profiles and for developing treatment strategies tailored to these atypical cases\. Generating such rare scenarios is therefore essential for stress\-testing algorithms, improving model robustness, and understanding system behaviour under distributional shifts\.
Existing approaches to outlier generation suffer from two important limitations\. First, although outliers are commonly motivated as low\-likelihood samples under a reference distribution\[[27](https://arxiv.org/html/2609.12113#bib.bib16),[11](https://arxiv.org/html/2609.12113#bib.bib15)\], most methods rely on heuristic notions such as reconstruction error, latent\-space distance, or classifier uncertainty rather than explicitly controlling likelihood itself\. Hence, they provide little formal interpretability over the statistical rarity of the generated samples\. Second, many existing approaches require specialized architectures or training objectives designed specifically for outlier synthesis\[[14](https://arxiv.org/html/2609.12113#bib.bib28)\]\. Outlier generation is therefore treated as a separate learning problem rather than a controllable sampling problem\.
Diffusion models \(DMs\) have emerged as a powerful framework in generative modelling, achieving remarkable success in domains such as image synthesis\[[17](https://arxiv.org/html/2609.12113#bib.bib1),[22](https://arxiv.org/html/2609.12113#bib.bib6),[23](https://arxiv.org/html/2609.12113#bib.bib7),[21](https://arxiv.org/html/2609.12113#bib.bib5)\]and video generation\[[16](https://arxiv.org/html/2609.12113#bib.bib8),[2](https://arxiv.org/html/2609.12113#bib.bib9)\]\. These models are typically trained to reverse a stochastic process that gradually transforms samples from a data distribution into noise via a learned score function\. These properties also make DMs particularly attractive for outlier generation\. DMs admit tractable likelihood estimation through the probability\-flow ordinary differential equation \(PF\-ODE\) and evolve according to continuous\-time dynamics that can be controlled\. In fact, as we will show, a DM trained on regular data can be steered toward low\-likelihood regions at inference time, without retraining or fine\-tuning\.
In this work, we propose a distributional control\-theoretic framework for outlier generation in DMs\. Unlike previous approaches that identify individual anomalous samples, we control the likelihood distribution\. The log\-likelihood map induces a one\-dimensional pushforward measure, and generating outliers corresponds to steering this measure towards lower values\. Our key observation is that likelihood reweighting implies a scaling of the score function with a controller determined by a Radon–Nikodym \(RN\) derivative on likelihood spaces\. This exact pointwise identity motivates a controlled PF\-ODE\. To obtain a practical controller, we exploit the exponential convergence of the Ornstein–Uhlenbeck \(OU\) semigroup that underlies the forward diffusion process\. We use an exponentially decaying approximation to the initial likelihood\-reweighting control\.
Numerical experiments on Gaussian mixture examples and the CIFAR\-10 dataset demonstrate that the proposed method successfully steers diffusion sampling toward low\-likelihood regions while remaining consistent with the underlying data geometry\. The results suggest that DMs can be interpreted as controllable systems and opens the door to distribution steering and generative modelling under functional constraints\.
## IBackground
### I\-AScore\-Based Diffusion Models
The diffusion process\[[26](https://arxiv.org/html/2609.12113#bib.bib3)\]can be parameterized in continuous time by the following Ornstein\-Uhlenbeck \(OU\) SDE
d𝐱t=−𝐱tdt\+2d𝐰t,𝐱0∼pdatad\\mathbf\{x\}\_\{t\}=\-\\mathbf\{x\}\_\{t\}dt\+\\sqrt\{2\}d\\mathbf\{w\}\_\{t\},\\quad\\mathbf\{x\}\_\{0\}\\sim p\_\{data\}\(1\)fort∈\[0,T\]t\\in\[0,T\], where𝐱0\\mathbf\{x\}\_\{0\}is sampled from a data distribution and𝐰t\\mathbf\{w\}\_\{t\}is Brownian motion\. Asttgrows,𝐱t\\mathbf\{x\}\_\{t\}diffuses from the clean data𝐱0\\mathbf\{x\}\_\{0\}into Gaussian noise\. We then generate realistic data by sampling Gaussian noise𝐱T\\mathbf\{x\}\_\{T\}and solving the reverse\-time SDE
d𝐱t=\[−𝐱t−2∇logpt\(𝐱t\)\]dt\+2d𝐰¯t,d\\mathbf\{x\}\_\{t\}=\[\-\\mathbf\{x\}\_\{t\}\-2\\nabla\\log p\_\{t\}\(\\mathbf\{x\}\_\{t\}\)\]dt\+\\sqrt\{2\}d\\overline\{\\mathbf\{w\}\}\_\{t\},\(2\)where𝐰¯t\\overline\{\\mathbf\{w\}\}\_\{t\}denotes a reverse Brownian motion\. Alternatively, we can solve the probability flow ODE \(PF\-ODE\[[26](https://arxiv.org/html/2609.12113#bib.bib3)\]\) in reverse\-time
𝐱˙t=−𝐱t−∇logpt\(𝐱t\),𝐱T∼𝒩\(0,I\)\.\\dot\{\\mathbf\{x\}\}\_\{t\}=\-\\mathbf\{x\}\_\{t\}\-\\nabla\\log p\_\{t\}\(\\mathbf\{x\}\_\{t\}\),\\quad\\mathbf\{x\}\_\{T\}\\sim\\mathcal\{N\}\(0,I\)\.\(3\)It is common practice in the DM literature to train a neural network𝐬θ\(𝐱t,t\)\\mathbf\{s\}\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t\)to approximate∇logpt\(𝐱t\)\\nabla\\log p\_\{t\}\(\\mathbf\{x\}\_\{t\}\), from which realistic training data can then be generated\.
### I\-BFokker\-Planck\-Kolmogorov Equation
The Fokker\-Planck\-Kolmogorov \(FPK\)\[[3](https://arxiv.org/html/2609.12113#bib.bib4)\]equation that governs the evolution of the underlying distribution from the SDEs in Equation[1](https://arxiv.org/html/2609.12113#S1.E1)\(forward time\) and Equation[2](https://arxiv.org/html/2609.12113#S1.E2)\(reverse time\) is given by the partial differential equation \(PDE\):
dpt\(𝐱t\)dt=∇⋅\[𝐱tpt\(𝐱t\)\]\+Δpt\(𝐱t\),\\displaystyle\\frac\{dp\_\{t\}\(\\mathbf\{x\}\_\{t\}\)\}\{dt\}=\\nabla\\cdot\[\\mathbf\{x\}\_\{t\}p\_\{t\}\(\\mathbf\{x\}\_\{t\}\)\]\+\\Delta p\_\{t\}\(\\mathbf\{x\}\_\{t\}\),\(4\)where∇⋅\\nabla\\cdotis the divergence andΔ\\Deltais the Laplacian operator\. Recently, density steering via controlling the FPK equation has been of interest to the control community\[[12](https://arxiv.org/html/2609.12113#bib.bib25),[6](https://arxiv.org/html/2609.12113#bib.bib26),[25](https://arxiv.org/html/2609.12113#bib.bib27)\]\. It is often convenient to express this evolution in terms of the log\-density to obtain log\-likelihoods\. The following result gives the corresponding PDE satisfied by the log\-density\.
###### Proposition 1\(Proposition 3\.1 of\[[19](https://arxiv.org/html/2609.12113#bib.bib2)\]\)
Assume the ground truth densitypt\(𝐱\)p\_\{t\}\(\\mathbf\{x\}\)is sufficiently smooth onℝn×\[0,T\]\\mathbb\{R\}^\{n\}\\times\[0,T\]with its log\-density denoted aslt\(𝐱\):=logpt\(𝐱\)l\_\{t\}\(\\mathbf\{x\}\):=\\log p\_\{t\}\(\\mathbf\{x\}\)\. Then for all\(𝐱,t\)\(\\mathbf\{x\},t\), its log\-density satisfies the PDE
∂tlt\(𝐱\)=𝐱⋅∇lt\(𝐱\)\+n\+Δlt\(𝐱\)\+‖∇lt\(𝐱\)‖2\.\\partial\_\{t\}l\_\{t\}\(\\mathbf\{x\}\)=\\mathbf\{x\}\\cdot\\nabla l\_\{t\}\(\\mathbf\{x\}\)\+n\+\\Delta l\_\{t\}\(\\mathbf\{x\}\)\+\\\|\\nabla l\_\{t\}\(\\mathbf\{x\}\)\\\|^\{2\}\.\(5\)
### I\-COrnstein\-Uhlenbeck Operator
The forward diffusion process underlying score\-based models corresponds to OU dynamics, whose generator plays an important role in our analysis of likelihood evolution\.
###### Definition 1\(Ornstein\-Uhlenbeck \(OU\) Generator\)
The OU generatorℒ\\mathcal\{L\}acts on smooth functionsf∈C2\(ℝn\)f\\in C^\{2\}\(\\mathbb\{R\}^\{n\}\)by
ℒf\(𝐱\)=Δf\(𝐱\)−𝐱⋅∇f\(𝐱\)\.\\mathcal\{L\}f\(\\mathbf\{x\}\)=\\Delta f\(\\mathbf\{x\}\)\-\\mathbf\{x\}\\cdot\\nabla f\(\\mathbf\{x\}\)\.
###### Theorem 1\(Theorem 3\.8 of\[[4](https://arxiv.org/html/2609.12113#bib.bib20)\]\)
Letγ:=𝒩\(0,I\)\\gamma:=\\mathcal\{N\}\(0,I\)be the standard Gaussian measure\. In the Hilbert spaceL2\(γ\),L^\{2\}\(\\gamma\),ℒ\\mathcal\{L\}is self\-adjoint and has discrete spectrum\{0,−1,−2,…\}\\\{0,\-1,\-2,\.\.\.\\\}of non\-positive integers\. The eigenfunctions are the Hermite polynomials\[[15](https://arxiv.org/html/2609.12113#bib.bib19)\]\(Hα\)α∈ℕn\(H\_\{\\alpha\}\)\_\{\\alpha\\in\\mathbb\{N\}^\{n\}\}, orthogonal inL2\(γ\),L^\{2\}\(\\gamma\),satisfyingℒHα=−\|α\|Hα\\mathcal\{L\}H\_\{\\alpha\}=\-\|\\alpha\|H\_\{\\alpha\}\.
## IIControlled Likelihood Generation
Outliers are commonly described as samples with low likelihood under a reference data distribution\. In likelihood\-based generative modelling, this intuition translates into identifying regions where the log\-likelihoodl\(𝐱\):=logp\(𝐱\)l\(\\mathbf\{x\}\):=\\log p\(\\mathbf\{x\}\)is small relative to typical data\. However, DMs learn*probability measures*, not individual samples\. Thus, if we wish to generate outliers in a principled way, we must define outliers at the level of distributions rather than individual points\.
In this section, we characterize outliers through the distribution of log\-likelihood values induced by a probability measure\. Since outliers correspond to samples with unusually low likelihood, our goal is to steer the likelihood distribution toward lower values\. This motivates the use of stochastic ordering to formalize the notion that the target likelihood distribution places more probability mass on lower\-likelihood regions\. Our main theoretical result shows that likelihood reweighting induces a suitable control through a simple RN structure, resulting in a multiplicative modification of the score function\. Motivated by this relation, we use the modified score in a controlled PF\-ODE and empirically evaluate its ability to generate low\-likelihood samples\.
### II\-AProblem Formulation and Definition of Outlier
Given a data distributionμ\\muonℝn\\mathbb\{R\}^\{n\}with densitypp, the log\-likelihood functionl\(𝐱\)=logp\(𝐱\)l\(\\mathbf\{x\}\)=\\log p\(\\mathbf\{x\}\)induces a scalar random variablel\(𝐱\)l\(\\mathbf\{x\}\)when𝐱∼μ\\mathbf\{x\}\\sim\\mu\. The pushforward measureL:=l\#μL:=l\_\{\\\#\}\\mutherefore describes the distribution of likelihood values under the model\. Shifting this distribution toward lower values corresponds to generating samples that are globally less likely\. By modifying this one\-dimensional likelihood distribution while preserving the conditional structure of the data given its likelihood level, we obtain a mechanism for generating structured outliers that remain consistent with the underlying data geometry\. We now formalize these notions\.
###### Definition 2\(Log\-likelihood pushforward measure\)
Letμ\\mube a probability measure onℝn\\mathbb\{R\}^\{n\}with densitypp\. Define the log\-likelihood functionl\(𝐱\):=logp\(𝐱\)\.l\(\\mathbf\{x\}\):=\\log p\(\\mathbf\{x\}\)\.The pushforward measure ofμ\\muunderllis the probability measureL:=l\#μL:=l\_\{\\\#\}\\muonℝ\\mathbb\{R\}\.
###### Definition 3\(Likelihood\-reweighted measure\)
Letμ\\mube a probability measure onℝn\\mathbb\{R\}^\{n\}with densityppand log\-likelihoodl\(𝐱\)=logp\(𝐱\)l\(\\mathbf\{x\}\)=\\log p\(\\mathbf\{x\}\)\. LetL=l\#μL=l\_\{\\\#\}\\mu\. Letη\\etabe a probability measure onℝ\\mathbb\{R\}such thatη≪L\\eta\\ll L\. A probability measureν\\nuonℝn\\mathbb\{R\}^\{n\}is called a likelihood\-reweighted measure ofμ\\muwith targetη\\etaifν\\nuadmits the disintegration
ν\(A\)=∫ℝμ\(A∣l\(𝐱\)=u\)η\(𝑑u\),A⊂ℝnBorel,\\nu\(A\)=\\int\_\{\\mathbb\{R\}\}\\mu\(A\\mid l\(\\mathbf\{x\}\)=u\)\\,\\eta\(du\),\\quad A\\subset\\mathbb\{R\}^\{n\}\\text\{ Borel\},whereμ\(⋅∣l\(𝐱\)=u\)\\mu\(\\cdot\\mid l\(\\mathbf\{x\}\)=u\)denotes a regular conditional probability ofμ\\mugivenl\(𝐱\)=ul\(\\mathbf\{x\}\)=u, which exists for Borel probability measures on Polish spaces\[[20](https://arxiv.org/html/2609.12113#bib.bib13)\]\.
As a consequence, ifν\\nuis a likelihood\-reweighted measure ofμ\\muwith targetη\\eta, thenl\#ν=ηl\_\{\\\#\}\\nu=\\eta\.
###### Definition 4\(First\-order stochastic dominance \(FOSD\)\[[1](https://arxiv.org/html/2609.12113#bib.bib14)\]\)
Letμ\\muandν\\nube probability measures onℝ\\mathbb\{R\}\. We say thatν\\nufirst\-order stochastically dominatesμ\\muand write
if and only ifμ\(\[x,∞\)\)≤ν\(\[x,∞\)\)\\mu\(\[x,\\infty\)\)\\leq\\nu\(\[x,\\infty\)\)for allx∈ℝx\\in\\mathbb\{R\}\.
Equivalently, if we letFμF\_\{\\mu\}andFνF\_\{\\nu\}be cumulative distribution functions \(CDFs\) ofμ\\muandν\\nurespectively, then
μ≤stν⇔Fμ\(u\)≥Fν\(u\)for allu∈ℝ\.\\mu\\leq\_\{st\}\\nu\\iff F\_\{\\mu\}\(u\)\\geq F\_\{\\nu\}\(u\)\\quad\\text\{for all \}u\\in\\mathbb\{R\}\.
Stochastic dominance constraints are studied in stochastic optimization and decision theory as a way of enforcing preference relations between random outcomes\[[8](https://arxiv.org/html/2609.12113#bib.bib23),[9](https://arxiv.org/html/2609.12113#bib.bib22),[10](https://arxiv.org/html/2609.12113#bib.bib24)\]\.
###### Definition 5\(ρ\\rho\-outlier measure\)
Letμ\\mube a probability measure onℝn\\mathbb\{R\}^\{n\}with log\-likelihood pushforwardL=l\#μL=l\_\{\\\#\}\\mu\. A likelihood\-reweighted measureν\\nuwith targetη\\etais called a*ρ\\rho\-outlier measure*if
\(1\)η≤stL,and \(2\)W1\(L,η\)≥ρ,\\text\{\(1\) \}\\eta\\leq\_\{st\}L,\\quad\\text\{ and \\hskip 10\.22217pt\(2\) \}W\_\{1\}\(L,\\eta\)\\geq\\rho,where≤st\\leq\_\{st\}represents FOSD \(see Definition[4](https://arxiv.org/html/2609.12113#Thmdefinition4)\) andW1\(⋅,⋅\)W\_\{1\}\(\\cdot,\\cdot\)is the Wasserstein\-1 distance \(see Equation \([6](https://arxiv.org/html/2609.12113#S2.E6)\) below\)\.
To generate samples from the likelihood\-reweighted measureν\\nu, we seek to characterize its score∇logq\\nabla\\log qin terms of the score∇logp\\nabla\\log pof the reference distribution, whereppandqqare the densities ofμ\\muandν\\nurespectively\. This allows us to use an existing diffusion score model while modifying its sampling dynamics through a likelihood\-dependent correction\.
### II\-BRadon\-Nikodym Structure of Likelihood Reweighting
Based on our formulation of likelihood\-reweighted distributions \(Definition[3](https://arxiv.org/html/2609.12113#Thmdefinition3)\), we derive some further properties\.
###### Theorem 2
Letμ\\mube a probability measure onℝn\\mathbb\{R\}^\{n\}with densityppand log\-likelihoodl\(𝐱\)=logp\(𝐱\)l\(\\mathbf\{x\}\)=\\log p\(\\mathbf\{x\}\)\. LetL=l\#μL=l\_\{\\\#\}\\mube the pushforward measure ofμ\\muunderll\. Letη\\etabe a probability measure onℝ\\mathbb\{R\}such thatη≪L\\eta\\ll L, and define
z\(u\):=dηdL\(u\)\.z\(u\):=\\frac\{d\\eta\}\{dL\}\(u\)\.
Letν\\nube the likelihood\-reweighted measure defined by
ν\(A\):=∫ℝμ\(A∣l\(𝐱\)=u\)η\(𝑑u\),A⊂ℝnBorel\.\\nu\(A\):=\\int\_\{\\mathbb\{R\}\}\\mu\(A\\mid l\(\\mathbf\{x\}\)=u\)\\,\\eta\(du\),\\qquad A\\subset\\mathbb\{R\}^\{n\}\\text\{ Borel\}\.
Thenν≪μ\\nu\\ll\\muand
dνdμ\(𝐱\)=z\(l\(𝐱\)\)=dηdL\(l\(𝐱\)\)\.\\frac\{d\\nu\}\{d\\mu\}\(\\mathbf\{x\}\)=z\(l\(\\mathbf\{x\}\)\)=\\frac\{d\\eta\}\{dL\}\(l\(\\mathbf\{x\}\)\)\.\(7\)
###### Proof:
Letf:ℝn→ℝf:\\mathbb\{R\}^\{n\}\\to\\mathbb\{R\}be bounded and measurable\. By definition ofν\\nu,
∫f\(𝐱\)ν\(𝑑𝐱\)=∫ℝ\[∫f\(𝐱\)p\(𝑑𝐱∣l\(𝐱\)=u\)\]η\(𝑑u\)\.\\int f\(\\mathbf\{x\}\)\\,\\nu\(d\\mathbf\{x\}\)=\\int\_\{\\mathbb\{R\}\}\\left\[\\int f\(\\mathbf\{x\}\)\\,p\(d\\mathbf\{x\}\\mid l\(\\mathbf\{x\}\)=u\)\\right\]\\eta\(du\)\.Sinceη≪L\\eta\\ll Lwith densityz\(u\)z\(u\), we can writeη\(du\)=z\(u\)L\(du\)\\eta\(du\)=z\(u\)\\,L\(du\)and obtain
∫f\(𝐱\)ν\(𝑑𝐱\)=∫ℝz\(u\)∫f\(𝐱\)μ\(𝑑𝐱∣l\(𝐱\)=u\)L\(𝑑u\)\.\\int f\(\\mathbf\{x\}\)\\,\\nu\(d\\mathbf\{x\}\)=\\int\_\{\\mathbb\{R\}\}z\(u\)\\,\\int f\(\\mathbf\{x\}\)\\,\\mu\(d\\mathbf\{x\}\\mid l\(\\mathbf\{x\}\)=u\)\\,L\(du\)\.By the disintegration theorem\[[7](https://arxiv.org/html/2609.12113#bib.bib18)\],
∫ℝ∫f\(𝐱\)μ\(𝑑𝐱∣l\(𝐱\)=u\)L\(𝑑u\)=∫f\(𝐱\)μ\(𝑑𝐱\)\.\\int\_\{\\mathbb\{R\}\}\\int f\(\\mathbf\{x\}\)\\,\\mu\(d\\mathbf\{x\}\\mid l\(\\mathbf\{x\}\)=u\)L\(du\)=\\int f\(\\mathbf\{x\}\)\\,\\mu\(d\\mathbf\{x\}\)\.Applying the same identity to the measurable functionx↦z\(l\(𝐱\)\)f\(𝐱\)x\\mapsto z\(l\(\\mathbf\{x\}\)\)f\(\\mathbf\{x\}\)gives
∫f\(𝐱\)z\(l\(𝐱\)\)μ\(𝑑𝐱\)=∫ℝz\(u\)∫f\(𝐱\)μ\(𝑑𝐱∣l\(𝐱\)=u\)L\(𝑑u\)\.\\int f\(\\mathbf\{x\}\)\\,z\(l\(\\mathbf\{x\}\)\)\\,\\mu\(d\\mathbf\{x\}\)=\\int\_\{\\mathbb\{R\}\}z\(u\)\\int f\(\\mathbf\{x\}\)\\,\\mu\(d\\mathbf\{x\}\\mid l\(\\mathbf\{x\}\)=u\)L\(du\)\.Comparing the two expressions, we conclude
∫f\(𝐱\)ν\(𝑑𝐱\)=∫f\(𝐱\)z\(l\(𝐱\)\)μ\(𝑑𝐱\)\.\\int f\(\\mathbf\{x\}\)\\,\\nu\(d\\mathbf\{x\}\)=\\int f\(\\mathbf\{x\}\)\\,z\(l\(\\mathbf\{x\}\)\)\\,\\mu\(d\\mathbf\{x\}\)\.Since this holds for all bounded measurableff, it follows that
dνdμ\(𝐱\)=z\(l\(𝐱\)\)\.\\frac\{d\\nu\}\{d\\mu\}\(\\mathbf\{x\}\)=z\(l\(\\mathbf\{x\}\)\)\.∎
###### Corollary 1
Letμt\\mu\_\{t\}be a family of probability measures onℝn\\mathbb\{R\}^\{n\}with densitiesptp\_\{t\}and log\-densitieslt=logptl\_\{t\}=\\log p\_\{t\}\. LetLt=\(lt\)\#μtL\_\{t\}=\(l\_\{t\}\)\_\{\\\#\}\\mu\_\{t\}\. Fix a family of target measures\(ηt\)t∈\[0,T\]\(\\eta\_\{t\}\)\_\{t\\in\[0,T\]\}onℝ\\mathbb\{R\}such thatηt≪Lt\\eta\_\{t\}\\ll L\_\{t\}for eachtt, and define
zt\(u\):=dηtdLt\(u\)\.z\_\{t\}\(u\):=\\frac\{d\\eta\_\{t\}\}\{dL\_\{t\}\}\(u\)\.Defineνt\\nu\_\{t\}to be the likelihood\-reweighted measure ofμt\\mu\_\{t\}with targetηt\\eta\_\{t\}in the sense of Definition[3](https://arxiv.org/html/2609.12113#Thmdefinition3), and letqtq\_\{t\}denote its density\. Thenνt≪μt\\nu\_\{t\}\\ll\\mu\_\{t\}and
dνtdμt\(𝐱t\)=zt\(lt\(𝐱t\)\)=dηtdLt\(lt\(𝐱t\)\)\.\\frac\{d\\nu\_\{t\}\}\{d\\mu\_\{t\}\}\(\\mathbf\{x\}\_\{t\}\)=z\_\{t\}\(l\_\{t\}\(\\mathbf\{x\}\_\{t\}\)\)=\\frac\{d\\eta\_\{t\}\}\{dL\_\{t\}\}\(l\_\{t\}\(\\mathbf\{x\}\_\{t\}\)\)\.\(8\)
The proof follows the same argument as Theorem[2](https://arxiv.org/html/2609.12113#Thmtheorem2)\.
###### Proposition 2
Let\(μt,pt,lt,zt,Lt,ηt,νt\)t∈\[0,T\]\(\\mu\_\{t\},p\_\{t\},l\_\{t\},z\_\{t\},L\_\{t\},\\eta\_\{t\},\\nu\_\{t\}\)\_\{t\\in\[0,T\]\}be defined as in Corollary[1](https://arxiv.org/html/2609.12113#Thmcorollary1)\. Assume further thatptp\_\{t\}is aC1C^\{1\}density onℝn\\mathbb\{R\}^\{n\}and thatztz\_\{t\}isC1C^\{1\}, and letqt\(𝐱\)q\_\{t\}\(\\mathbf\{x\}\)be the density ofνt\\nu\_\{t\}with respect to the Lebesgue measure\. Then
∇logqt\(𝐱\)=\[1\+ct\(lt\(𝐱\)\)\]∇logpt\(𝐱\),\\nabla\\log q\_\{t\}\(\\mathbf\{x\}\)=\[1\+c\_\{t\}\(l\_\{t\}\(\\mathbf\{x\}\)\)\]\\nabla\\log p\_\{t\}\(\\mathbf\{x\}\),\(9\)where
ct\(u\):=∂ulogzt\(u\)\.c\_\{t\}\(u\):=\\partial\_\{u\}\\log z\_\{t\}\(u\)\.\(10\)
The proof is a direct calculation using the chain rule\. The proposition shows that∇𝐱\[logzt\(lt\(𝐱\)\)\]\\nabla\_\{\\mathbf\{x\}\}\[\\log z\_\{t\}\(l\_\{t\}\(\\mathbf\{x\}\)\)\]lies in the span of∇logpt\\nabla\\log p\_\{t\}everywhere\. We are now ready to summarize the main theoretical results\.
###### Corollary 2
Let\(ηt\)t∈\[0,T\]\(\\eta\_\{t\}\)\_\{t\\in\[0,T\]\}be a family of target likelihood distributions satisfyingηt≪Lt\\eta\_\{t\}\\ll L\_\{t\}for alltt\. Then the corresponding likelihood\-reweighted densitiesqtq\_\{t\}satisfy the score relation
∇logqt\(𝐱\)=\[1\+∂u\(logdηtdLt\)\(logpt\(𝐱\)\)\]∇logpt\(𝐱\)\.\\nabla\\log q\_\{t\}\(\\mathbf\{x\}\)=\\left\[1\+\\partial\_\{u\}\\left\(\\log\\frac\{d\\eta\_\{t\}\}\{dL\_\{t\}\}\\right\)\(\\log p\_\{t\}\(\\mathbf\{x\}\)\)\\right\]\\nabla\\log p\_\{t\}\(\\mathbf\{x\}\)\.\(11\)Consequently, likelihood reweighting induces a multiplicative modification of the diffusion score function\. Furthermore, we choose a target likelihood distributionη0\\eta\_\{0\}that satisfies
η0≤stL0,W1\(L0,η0\)\>ρ,\\eta\_\{0\}\\leq\_\{st\}L\_\{0\},\\quad W\_\{1\}\(L\_\{0\},\\eta\_\{0\}\)\>\\rho,matching our specification ofρ\\rho\-outliers\.
Corollary[2](https://arxiv.org/html/2609.12113#Thmcorollary2)shows that likelihood control does not require learning a new score function\. Instead, the score of a likelihood\-reweighted density can be expressed using the original diffusion score and a scalar likelihood\-dependent correction, without learning an independent score function\. This provides a distribution\-steering perspective on outlier generation through likelihood reweighting\.
## IIIImplementation
Motivated by the score relation in Corollary[2](https://arxiv.org/html/2609.12113#Thmcorollary2), we introduce the controlled reverse\-time ODE ansatz
𝐱˙t=−𝐱t−\(1\+ct\(lt\(𝐱t\)\)\)∇logpt\(𝐱t\)⏟=:∇logqt\(𝐱t\),\\dot\{\\mathbf\{x\}\}\_\{t\}=\-\\mathbf\{x\}\_\{t\}\-\\underbrace\{\(1\+c\_\{t\}\(l\_\{t\}\(\\mathbf\{x\}\_\{t\}\)\)\)\\nabla\\log p\_\{t\}\(\\mathbf\{x\}\_\{t\}\)\}\_\{=:\\nabla\\log q\_\{t\}\(\\mathbf\{x\}\_\{t\}\)\},\(12\)wherect\(u\)=ddulog\(dηtdLt\(u\)\)c\_\{t\}\(u\)=\\frac\{d\}\{du\}\\log\\left\(\\frac\{d\\eta\_\{t\}\}\{dL\_\{t\}\}\(u\)\\right\)is a coefficient that needs to be determined\. To approximatectc\_\{t\}, it is therefore necessary to understand how bothηt\\eta\_\{t\}andLtL\_\{t\}evolve over time under the diffusion dynamics\. In particular, the forward OU process governing the diffusion model induces a contraction of density perturbations toward the Gaussian equilibrium, which we analyze next\.
###### Theorem 3
Let𝐱t\\mathbf\{x\}\_\{t\}solve the OU SDE \([1](https://arxiv.org/html/2609.12113#S1.E1)\), where𝐱0∼p\\mathbf\{x\}\_\{0\}\\sim p\. Letμt\\mu\_\{t\}denote the time\-marginal law of𝐱t\\mathbf\{x\}\_\{t\}, and letptp\_\{t\}denote its density\. The invariant measure is the standard Gaussian measureγ=𝒩\(0,I\)\\gamma=\\mathcal\{N\}\(0,I\)with densitypγp\_\{\\gamma\}\. Define the log\-densities
lt\(𝐱\):=logpt\(𝐱\),l∗\(𝐱\):=logpγ\(𝐱\)\.l\_\{t\}\(\\mathbf\{x\}\):=\\log p\_\{t\}\(\\mathbf\{x\}\),\\quad l^\{\*\}\(\\mathbf\{x\}\):=\\log p\_\{\\gamma\}\(\\mathbf\{x\}\)\.Assumeμt≪γ\\mu\_\{t\}\\ll\\gammaand denoteht\(𝐱\)h\_\{t\}\(\\mathbf\{x\}\)as the RN derivativeht\(𝐱\):=dμtdγ\(𝐱\)h\_\{t\}\(\\mathbf\{x\}\):=\\frac\{d\\mu\_\{t\}\}\{d\\gamma\}\(\\mathbf\{x\}\)\. Assumeh0∈H1\(γ\)h\_\{0\}\\in H^\{1\}\(\\gamma\)\. Then
‖ht−1‖L2\(γ\)2≤e−2t‖h0−1‖L2\(γ\)2\.\\\|h\_\{t\}\-1\\\|\_\{L^\{2\}\(\\gamma\)\}^\{2\}\\leq e^\{\-2t\}\\\|h\_\{0\}\-1\\\|^\{2\}\_\{L^\{2\}\(\\gamma\)\}\.Moreover,
‖∇ht‖L2\(γ\)≤e−t‖∇h0‖L2\(γ\)\.\\\|\\nabla h\_\{t\}\\\|\_\{L^\{2\}\(\\gamma\)\}\\leq e^\{\-t\}\\\|\\nabla h\_\{0\}\\\|\_\{L^\{2\}\(\\gamma\)\}\.Finally, suppose there existsm\>0m\>0such thatht\(𝐱\)≥mh\_\{t\}\(\\mathbf\{x\}\)\\geq mforμt\\mu\_\{t\}\-almost all𝐱\\mathbf\{x\}andtt\. Then
‖ℓt−ℓ∗‖H1\(γ\)=‖loght‖H1\(γ\)≤e−tm‖h0−1‖H1\(γ\)\.\\\|\\ell\_\{t\}\-\\ell^\{\*\}\\\|\_\{H^\{1\}\(\\gamma\)\}=\\\|\\log h\_\{t\}\\\|\_\{H^\{1\}\(\\gamma\)\}\\leq\\frac\{e^\{\-t\}\}\{m\}\\\|h\_\{0\}\-1\\\|\_\{H^\{1\}\(\\gamma\)\}\.
###### Proof:
The densityptp\_\{t\}satisfies the FPK equation
∂tpt=∇⋅\(𝐱pt\)\+Δpt\.\\partial\_\{t\}p\_\{t\}=\\nabla\\cdot\(\\mathbf\{x\}p\_\{t\}\)\+\\Delta p\_\{t\}\.
Define the density ratioht\(𝐱\):=pt\(𝐱\)γ\(𝐱\)\.h\_\{t\}\(\\mathbf\{x\}\):=\\frac\{p\_\{t\}\(\\mathbf\{x\}\)\}\{\\gamma\(\\mathbf\{x\}\)\}\.Using the identities
∇pγ=−𝐱pγ,Δpγ=\(‖𝐱‖2−n\)pγ,\\nabla p\_\{\\gamma\}=\-\\mathbf\{x\}p\_\{\\gamma\},\\quad\\Delta p\_\{\\gamma\}=\(\\\|\\mathbf\{x\}\\\|^\{2\}\-n\)p\_\{\\gamma\},one verifies thathth\_\{t\}satisfies
∂tht=ℒht,ℒ:=Δ−𝐱⋅∇,\\partial\_\{t\}h\_\{t\}=\\mathcal\{L\}h\_\{t\},\\quad\\mathcal\{L\}:=\\Delta\-\\mathbf\{x\}\\cdot\\nabla,whereℒ\\mathcal\{L\}is the OU generator introduced in Definition[1](https://arxiv.org/html/2609.12113#Thmdefinition1)\. Letgt:=ht−1\.g\_\{t\}:=h\_\{t\}\-1\.Since∫ht𝑑γ=1\\int h\_\{t\}\\,d\\gamma=1, we have∫gt𝑑γ=0\\int g\_\{t\}\\,d\\gamma=0\. The evolution equation becomes
∂tgt=ℒgt,gt\|t=0=g0\.\\partial\_\{t\}g\_\{t\}=\\mathcal\{L\}g\_\{t\},\\quad g\_\{t\}\|\_\{t=0\}=g\_\{0\}\.\(13\)By Theorem[1](https://arxiv.org/html/2609.12113#Thmtheorem1), the operatorℒ\\mathcal\{L\}is self\-adjoint onL2\(γ\)L^\{2\}\(\\gamma\)with eigenfunctions given by Hermite polynomials\[[15](https://arxiv.org/html/2609.12113#bib.bib19)\],ℒHα=−\|α\|Hα\.\\mathcal\{L\}H\_\{\\alpha\}=\-\|\\alpha\|H\_\{\\alpha\}\.Expandingg0g\_\{0\}in the Hermite basis gives
g0=∑α≠0gαHα,gα=⟨g0,Hα⟩L2\(γ\)\.g\_\{0\}=\\sum\_\{\\alpha\\neq 0\}g\_\{\\alpha\}H\_\{\\alpha\},\\quad g\_\{\\alpha\}=\\langle g\_\{0\},H\_\{\\alpha\}\\rangle\_\{L^\{2\}\(\\gamma\)\}\.Let\(Pt\)t≥0\(P\_\{t\}\)\_\{t\\geq 0\}be the OU semigroup onL2\(γ\)L^\{2\}\(\\gamma\)generated byℒ\\mathcal\{L\}\. Sincegtg\_\{t\}satisfies Equation \([13](https://arxiv.org/html/2609.12113#S3.E13)\), we identifygtg\_\{t\}with the unique semigroup solutiongt=Ptg0\.g\_\{t\}=P\_\{t\}g\_\{0\}\.SincePtP\_\{t\}acts diagonally on the Hermite basis, we can expand this solution
gt=∑α≠0e−\|α\|tgαHα\.g\_\{t\}=\\sum\_\{\\alpha\\neq 0\}e^\{\-\|\\alpha\|t\}g\_\{\\alpha\}H\_\{\\alpha\}\.TakingL2\(γ\)L^\{2\}\(\\gamma\)norms yields
‖gt‖L2\(γ\)2=∑k=1∞e−2kt∑\|α\|=kgα2≤e−2t‖g0‖L2\(γ\)2,\\\|g\_\{t\}\\\|\_\{L^\{2\}\(\\gamma\)\}^\{2\}=\\sum\_\{k=1\}^\{\\infty\}e^\{\-2kt\}\\sum\_\{\|\\alpha\|=k\}g\_\{\\alpha\}^\{2\}\\leq e^\{\-2t\}\\\|g\_\{0\}\\\|^\{2\}\_\{L^\{2\}\(\\gamma\)\},where the inequality follows immediately sincee−kt≤e−te^\{\-kt\}\\leq e^\{\-t\}fork≥1k\\geq 1\. By Lemma 1 of\[[5](https://arxiv.org/html/2609.12113#bib.bib17)\], we can bound the gradient:
‖∇gt‖L2\(γ\)2≤e−2t‖∇g0‖L2\(γ\)2\.\\\|\\nabla g\_\{t\}\\\|^\{2\}\_\{L^\{2\}\(\\gamma\)\}\\leq e^\{\-2t\}\\\|\\nabla g\_\{0\}\\\|^\{2\}\_\{L^\{2\}\(\\gamma\)\}\.Becausegt=ht−1g\_\{t\}=h\_\{t\}\-1, this is equivalent to
‖∇ht‖L2\(γ\)≤e−t‖∇h0‖L2\(γ\)\.\\\|\\nabla h\_\{t\}\\\|\_\{L^\{2\}\(\\gamma\)\}\\leq e^\{\-t\}\\\|\\nabla h\_\{0\}\\\|\_\{L^\{2\}\(\\gamma\)\}\.
Now assumeht≥m\>0h\_\{t\}\\geq m\>0\. Sincelog\\logis1/m1/m\-Lipschitz on\[m,∞\)\[m,\\infty\),
\|loght\(𝐱\)\|≤1m\|ht\(𝐱\)−1\|\.\|\\log h\_\{t\}\(\\mathbf\{x\}\)\|\\leq\\frac\{1\}\{m\}\|h\_\{t\}\(\\mathbf\{x\}\)\-1\|\.Also, by expanding∇loght\(𝐱\)=∇ht\(𝐱\)ht\(𝐱\),\\nabla\\log h\_\{t\}\(\\mathbf\{x\}\)=\\frac\{\\nabla h\_\{t\}\(\\mathbf\{x\}\)\}\{h\_\{t\}\(\\mathbf\{x\}\)\},we have
‖∇loght\(𝐱\)‖≤1m‖∇ht\(𝐱\)‖\.\\\|\\nabla\\log h\_\{t\}\(\\mathbf\{x\}\)\\\|\\leq\\frac\{1\}\{m\}\\\|\\nabla h\_\{t\}\(\\mathbf\{x\}\)\\\|\.Combining the two inequalities,
‖loght‖H1\(γ\)≤1m‖ht−1‖H1\(γ\)\.\\\|\\log h\_\{t\}\\\|\_\{H^\{1\}\(\\gamma\)\}\\leq\\frac\{1\}\{m\}\\\|h\_\{t\}\-1\\\|\_\{H^\{1\}\(\\gamma\)\}\.Finally, using theL2L^\{2\}and gradient estimates above,
‖ht−1‖H1\(γ\)≤e−t‖h0−1‖H1\(γ\)\.\\\|h\_\{t\}\-1\\\|\_\{H^\{1\}\(\\gamma\)\}\\leq e^\{\-t\}\\\|h\_\{0\}\-1\\\|\_\{H^\{1\}\(\\gamma\)\}\.Since
loght=ℓt−ℓ∗,\\log h\_\{t\}=\\ell\_\{t\}\-\\ell^\{\*\},the conclusion follows\. ∎
Theorem[3](https://arxiv.org/html/2609.12113#Thmtheorem3)shows thatLt=\(lt\)\#μtL\_\{t\}=\(l\_\{t\}\)\_\{\\\#\}\\mu\_\{t\}converges exponentially towards the Gaussian likelihood equilibrium\(l∗\)\#γ\(l^\{\*\}\)\_\{\\\#\}\\gamma\. Motivated by this, we choose the target family\{ηt\}\\\{\\eta\_\{t\}\\\}that likewise vanishes toward the same equilibrium\. In particular, the slowest nonconstant OU decay ratee−te^\{\-t\}motivates the exponentially decaying controller introduced below\.
### III\-AApproximating a control function
In this section, we claim that, provided the log\-likelihood pushforward measureL0L\_\{0\}and a targetη0\\eta\_\{0\}, the exponentially interpolated control function
c~t\(u\)=e−tc0\(u\)=e−t∂u\(logdη0dL0\(u\)\),\\tilde\{c\}\_\{t\}\(u\)=e^\{\-t\}c\_\{0\}\(u\)=e^\{\-t\}\\partial\_\{u\}\\left\(\\log\\frac\{d\\eta\_\{0\}\}\{dL\_\{0\}\}\(u\)\\right\),\(14\)provides a natural approximation choice for the controlled reverse dynamics\. The true time‑dependent controlct\(u\)c\_\{t\}\(u\)is computationally intractable\. Motivated by the slowest decaying nonconstant term,e−te^\{\-t\}, of the OU semigroup, we choose the exponentially interpolated controller\. We prove that this approximation satisfies a reasonable error bound over both small and largett, thus providing a natural admissible approximation that respects the boundary conditions and admits explicit error bounds near both endpoints\.
###### Theorem 4
Let\(𝐱t\)t∈\[0,T\]\(\\mathbf\{x\}\_\{t\}\)\_\{t\\in\[0,T\]\}be an OU diffusion \([1](https://arxiv.org/html/2609.12113#S1.E1)\), and letLtL\_\{t\}andηt\\eta\_\{t\}be two families of probability measures onℝ\\mathds\{R\}which are absolutely continuous with respect to Lebesgue measure, converge exponentially to the Gaussian likelihood equilibrium, and satisfy:
η0≪L0,η0≤stL0,W1\(η0,L0\)≥ρ\.\\eta\_\{0\}\\ll L\_\{0\},\\quad\\eta\_\{0\}\\leq\_\{st\}L\_\{0\},\\quad W\_\{1\}\(\\eta\_\{0\},L\_\{0\}\)\\geq\\rho\.Define the likelihood ratio
zt\(u\):=dηtdLt\(u\),ct\(u\):=∂ulogzt\(u\)\.z\_\{t\}\(u\):=\\frac\{d\\eta\_\{t\}\}\{dL\_\{t\}\}\(u\),\\quad c\_\{t\}\(u\):=\\partial\_\{u\}\\log z\_\{t\}\(u\)\.Assume:
1. \(A1\)There existsL\>0L\>0such that for alls,t∈\[0,T\]s,t\\in\[0,T\], ‖ct−cs‖L2\(Lt\)≤L\|t−s\|\.\\\|c\_\{t\}\-c\_\{s\}\\\|\_\{L^\{2\}\(L\_\{t\}\)\}\\leq L\|t\-s\|\.
2. \(A2\)There existsK≥1K\\geq 1such that for allt∈\[0,T\]t\\in\[0,T\]and all measurableff, K−1‖f‖L2\(L0\)≤‖f‖L2\(Lt\)≤K‖f‖L2\(L0\)\.K^\{\-1\}\\\|f\\\|\_\{L^\{2\}\(L\_\{0\}\)\}\\leq\\\|f\\\|\_\{L^\{2\}\(L\_\{t\}\)\}\\leq K\\\|f\\\|\_\{L^\{2\}\(L\_\{0\}\)\}\.
Define the interpolantc~t\(u\):=e−tc0\(u\)\.\\tilde\{c\}\_\{t\}\(u\):=e^\{\-t\}c\_\{0\}\(u\)\.Thenc~0=c0\\tilde\{c\}\_\{0\}=c\_\{0\}, and for everyt∈\[0,T\]t\\in\[0,T\],
‖c~t−ct‖L2\(Lt\)≤K\(\(1−e−t\)‖c0‖L2\(L0\)\+Lt\)\.\\\|\\tilde\{c\}\_\{t\}\-c\_\{t\}\\\|\_\{L^\{2\}\(L\_\{t\}\)\}\\leq K\\Big\(\(1\-e^\{\-t\}\)\\\|c\_\{0\}\\\|\_\{L^\{2\}\(L\_\{0\}\)\}\+Lt\\Big\)\.In particular,‖c~t−ct‖L2\(Lt\)=O\(t\),\\\|\\tilde\{c\}\_\{t\}\-c\_\{t\}\\\|\_\{L^\{2\}\(L\_\{t\}\)\}=O\(t\),
###### Proof:
Fixt∈\[0,T\]t\\in\[0,T\]\. Add and subtractc0c\_\{0\}:
c~t−ct=\(e−tc0−c0\)\+\(c0−ct\)\.\\tilde\{c\}\_\{t\}\-c\_\{t\}=\(e^\{\-t\}c\_\{0\}\-c\_\{0\}\)\+\(c\_\{0\}\-c\_\{t\}\)\.TakeL2\(Lt\)L^\{2\}\(L\_\{t\}\)norms and apply the triangle inequality:
‖c~t−ct‖L2\(Lt\)≤‖\(1−e−t\)c0‖L2\(Lt\)\+‖ct−c0‖L2\(Lt\)\.\\\|\\tilde\{c\}\_\{t\}\-c\_\{t\}\\\|\_\{L^\{2\}\(L\_\{t\}\)\}\\leq\\\|\(1\-e^\{\-t\}\)c\_\{0\}\\\|\_\{L^\{2\}\(L\_\{t\}\)\}\+\\\|c\_\{t\}\-c\_\{0\}\\\|\_\{L^\{2\}\(L\_\{t\}\)\}\.By \(A2\),‖\(1−e−t\)c0‖L2\(Lt\)≤K\(1−e−t\)‖c0‖L2\(L0\)\.\\displaystyle\\quad\\\|\(1\-e^\{\-t\}\)c\_\{0\}\\\|\_\{L^\{2\}\(L\_\{t\}\)\}\\leq K\(1\-e^\{\-t\}\)\\\|c\_\{0\}\\\|\_\{L^\{2\}\(L\_\{0\}\)\}\.By \(A1\),‖ct−c0‖L2\(Lt\)≤Lt\.\\displaystyle\\quad\\\|c\_\{t\}\-c\_\{0\}\\\|\_\{L^\{2\}\(L\_\{t\}\)\}\\leq Lt\.Combine the two bounds to obtain
‖c~t−ct‖L2\(Lt\)≤K\(\(1−e−t\)‖c0‖L2\(L0\)\+Lt\)\.\\\|\\tilde\{c\}\_\{t\}\-c\_\{t\}\\\|\_\{L^\{2\}\(L\_\{t\}\)\}\\leq K\\Big\(\(1\-e^\{\-t\}\)\\\|c\_\{0\}\\\|\_\{L^\{2\}\(L\_\{0\}\)\}\+Lt\\Big\)\.Finally, since1−e−t≤t1\-e^\{\-t\}\\leq t, the right\-hand side isO\(t\)O\(t\)\. ∎
###### Theorem 5
Let\(𝐱t\)t∈\[0,T\],Lt,ηt,γ,zt\(u\),ct\(u\)\(\\mathbf\{x\}\_\{t\}\)\_\{t\\in\[0,T\]\},L\_\{t\},\\eta\_\{t\},\\gamma,z\_\{t\}\(u\),c\_\{t\}\(u\)be defined as in Theorem[4](https://arxiv.org/html/2609.12113#Thmtheorem4), and letlt,μt,νtl\_\{t\},\\mu\_\{t\},\\nu\_\{t\}be defined as in Corollary[1](https://arxiv.org/html/2609.12113#Thmcorollary1)\. Moreover, assumec0∈L∞c\_\{0\}\\in L^\{\\infty\}and suppose the assumptions from Theorem[3](https://arxiv.org/html/2609.12113#Thmtheorem3)are satisfied forμt\\mu\_\{t\}andνt\\nu\_\{t\}\. Then the exponentially\-interpolated controllerc~t\(u\):=e−tc0\(u\)\\tilde\{c\}\_\{t\}\(u\):=e^\{\-t\}c\_\{0\}\(u\)satisfies the boundary conditionc~0=c0\\tilde\{c\}\_\{0\}=c\_\{0\}, and moreover
∥\(c~t−ct\)\(lt\(⋅\)\)∇logpt\(⋅\)∥L2\(γ\)≤Ce−t,t∈\[0,T\],\\\|\(\\tilde\{c\}\_\{t\}\-c\_\{t\}\)\(l\_\{t\}\(\\cdot\)\)\\,\\nabla\\log p\_\{t\}\(\\cdot\)\\\|\_\{L^\{2\}\(\\gamma\)\}\\leq Ce^\{\-t\},\\quad t\\in\[0,T\],for a constantC\>0C\>0\. In particular,
\(c~t−ct\)\(lt\(x\)\)∇logpt\(x\)\(\\tilde\{c\}\_\{t\}\-c\_\{t\}\)\(l\_\{t\}\(x\)\)\\,\\nabla\\log p\_\{t\}\(x\)converges exponentially fast to zero inL2\(γ\)L^\{2\}\(\\gamma\)\.
###### Proof:
logzt\\log z\_\{t\}can be expanded as
logzt\(lt\(𝐱\)\)=logdνtdμt\(𝐱\)=logdνtdγ\(𝐱\)−logdμtdγ\(𝐱\)\.\\log z\_\{t\}\(l\_\{t\}\(\\mathbf\{x\}\)\)=\\log\\frac\{d\\nu\_\{t\}\}\{d\\mu\_\{t\}\}\(\\mathbf\{x\}\)=\\log\\frac\{d\\nu\_\{t\}\}\{d\\gamma\}\(\\mathbf\{x\}\)\-\\log\\frac\{d\\mu\_\{t\}\}\{d\\gamma\}\(\\mathbf\{x\}\)\.From Theorem[3](https://arxiv.org/html/2609.12113#Thmtheorem3)and the triangle inequality, we have
‖logzt\(lt\(⋅\)\)‖H1\(γ\)\\displaystyle\\\|\\log z\_\{t\}\(l\_\{t\}\(\\cdot\)\)\\\|\_\{H^\{1\}\(\\gamma\)\}≤\(‖logdνtdγ‖H1\(γ\)\+‖logdμtdγ‖H1\(γ\)\)\\displaystyle\\leq\\left\(\\\|\\log\\frac\{d\\nu\_\{t\}\}\{d\\gamma\}\\\|\_\{H^\{1\}\(\\gamma\)\}\+\\\|\\log\\frac\{d\\mu\_\{t\}\}\{d\\gamma\}\\\|\_\{H^\{1\}\(\\gamma\)\}\\right\)≤e−tm\(‖dν0dγ−1‖H1\(γ\)\+‖dμ0dγ−1‖H1\(γ\)\)\.\\displaystyle\\leq\\frac\{e^\{\-t\}\}\{m\}\\left\(\\\|\\frac\{d\\nu\_\{0\}\}\{d\\gamma\}\-1\\\|\_\{H^\{1\}\(\\gamma\)\}\+\\\|\\frac\{d\\mu\_\{0\}\}\{d\\gamma\}\-1\\\|\_\{H^\{1\}\(\\gamma\)\}\\right\)\.By definition of theH1\(γ\)H^\{1\}\(\\gamma\)norm,
‖∇logzt‖L2\(γ\)\\displaystyle\\\|\\nabla\\log z\_\{t\}\\\|\_\{L^\{2\}\(\\gamma\)\}≤e−tm\(‖dν0dγ−1‖H1\(γ\)\+‖dμ0dγ−1‖H1\(γ\)\)\.\\displaystyle\\leq\\frac\{e^\{\-t\}\}\{m\}\\left\(\\\|\\frac\{d\\nu\_\{0\}\}\{d\\gamma\}\-1\\\|\_\{H^\{1\}\(\\gamma\)\}\+\\\|\\frac\{d\\mu\_\{0\}\}\{d\\gamma\}\-1\\\|\_\{H^\{1\}\(\\gamma\)\}\\right\)\.Since∇logzt\(lt\(𝐱t\)\)\\nabla\\log z\_\{t\}\(l\_\{t\}\(\\mathbf\{x\}\_\{t\}\)\)can be expanded as∂ulogzt∇lt\(𝐱t\)\\partial\_\{u\}\\log z\_\{t\}\\nabla l\_\{t\}\(\\mathbf\{x\}\_\{t\}\),
∥∇logzt\(lt\(⋅\)\)∥L2\(γ\)=∥∂ulogzt∇lt\(𝐱t\)∥L2\(γ\),\\\|\\nabla\\log z\_\{t\}\(l\_\{t\}\(\\cdot\)\)\\\|\_\{L^\{2\}\(\\gamma\)\}=\\\|\\partial\_\{u\}\\log z\_\{t\}\\nabla l\_\{t\}\(\\mathbf\{x\}\_\{t\}\)\\\|\_\{L^\{2\}\(\\gamma\)\},which decays exponentially\. Finally, asct=∂ulogztc\_\{t\}=\\partial\_\{u\}\\log z\_\{t\}andc~t=e−tc0\\tilde\{c\}\_\{t\}=e^\{\-t\}c\_\{0\}, we can write
c~t−ct=e−t∂ulogz0−∂ulogzt\.\\tilde\{c\}\_\{t\}\-c\_\{t\}=e^\{\-t\}\\partial\_\{u\}\\log z\_\{0\}\-\\partial\_\{u\}\\log z\_\{t\}\.Using the triangle inequality,
∥\(c~t−ct\)\(lt\(⋅\)\)∇lt∥L2\(γ\)≤\\displaystyle\\\|\(\\tilde\{c\}\_\{t\}\-c\_\{t\}\)\(l\_\{t\}\(\\cdot\)\)\\nabla l\_\{t\}\\\|\_\{L^\{2\}\(\\gamma\)\}\\leqe−t‖c0‖L∞‖∇lt‖L2\(γ\)\\displaystyle\\penalty\\ e^\{\-t\}\\\|c\_\{0\}\\\|\_\{L^\{\\infty\}\}\\\|\\nabla l\_\{t\}\\\|\_\{L^\{2\}\(\\gamma\)\}\+‖∇logzt\(lt\)‖L2\(γ\)\.\\displaystyle\+\\\|\\nabla\\log z\_\{t\}\(l\_\{t\}\)\\\|\_\{L^\{2\}\(\\gamma\)\}\.Since Theorem[3](https://arxiv.org/html/2609.12113#Thmtheorem3)yields
‖∇lt−∇logγ‖L2\(γ\)=‖∇lt\+x‖L2\(γ\)≤C1e−t,\\\|\\nabla l\_\{t\}\-\\nabla\\log\\gamma\\\|\_\{L^\{2\}\(\\gamma\)\}=\\\|\\nabla l\_\{t\}\+x\\\|\_\{L^\{2\}\(\\gamma\)\}\\leq C\_\{1\}e^\{\-t\},we obtain that‖∇lt‖L2\(γ\)\\\|\\nabla l\_\{t\}\\\|\_\{L^\{2\}\(\\gamma\)\}is uniformly bounded overtt\. Thus,
∥\(c~t\\displaystyle\\\|\(\\tilde\{c\}\_\{t\}−ct\)\(lt\(⋅\)\)∇logpt\(⋅\)∥L2\(γ\)≤Ce−t,\\displaystyle\-c\_\{t\}\)\(l\_\{t\}\(\\cdot\)\)\\,\\nabla\\log p\_\{t\}\(\\cdot\)\\\|\_\{L^\{2\}\(\\gamma\)\}\\leq Ce^\{\-t\},whereC=\\displaystyle\\text\{where \}C=‖c0‖L∞\(γ\)supt∈\[0,T\]‖∇lt‖L2\(γ\)\\displaystyle\\penalty\\ \\\|c\_\{0\}\\\|\_\{L^\{\\infty\}\(\\gamma\)\}\\sup\_\{t\\in\[0,T\]\}\\\|\\nabla l\_\{t\}\\\|\_\{L^\{2\}\(\\gamma\)\}\+1m\(‖dν0dγ−1‖H1\(γ\)\+‖dμ0dγ−1‖H1\(γ\)\)\\displaystyle\+\\frac\{1\}\{m\}\(\\\|\\frac\{d\\nu\_\{0\}\}\{d\\gamma\}\-1\\\|\_\{H^\{1\}\(\\gamma\)\}\+\\\|\\frac\{d\\mu\_\{0\}\}\{d\\gamma\}\-1\\\|\_\{H^\{1\}\(\\gamma\)\}\)and therefore the right\-hand side isO\(e−t\)O\(e^\{\-t\}\)\.∎
###### Corollary 3
Under the assumptions of Theorems[4](https://arxiv.org/html/2609.12113#Thmtheorem4)and[5](https://arxiv.org/html/2609.12113#Thmtheorem5), the error between the true control termct\(u\)∇ltc\_\{t\}\(u\)\\nabla l\_\{t\}and the estimatec~t\(u\)∇lt=e−tc0\(u\)∇lt\\tilde\{c\}\_\{t\}\(u\)\\nabla l\_\{t\}=e^\{\-t\}c\_\{0\}\(u\)\\nabla l\_\{t\}are bounded for small and largett:
∥\(c~t−ct\)\(lt\(⋅\)\)∇lt∥L2\(γ\)≤O\(min\(t,e−t\)\)\.\\\|\(\\tilde\{c\}\_\{t\}\-c\_\{t\}\)\(l\_\{t\}\(\\cdot\)\)\\nabla l\_\{t\}\\\|\_\{L^\{2\}\(\\gamma\)\}\\leq O\(\\min\(t,e^\{\-t\}\)\)\.\(15\)
### III\-BApproximating the target Radon\-Nikodym derivative
We will rely on chi\-square approximations forL0L\_\{0\}in this paper\. In our numerical experiments, these approximations will be verified and demonstrated\. We introduce the notation of a linear mapLa,CL\_\{a,C\}defined asLa,C\(x\)=−ax\+CL\_\{a,C\}\(x\)=\-ax\+C\. Under this approximation, the control coefficients admit closed\-form expressions, making the proposed controller straightforward to compute in practice\.
###### Assumption 1
LetL0L\_\{0\}be a likelihood\-pushforward measure onℝ\\mathbb\{R\}with finite mean\. Assume there existsa\>0a\>0andCCsuch that\(La,C\)\#\(χn2\)≤stL0\(L\_\{a,C\}\)\_\{\\\#\}\(\\chi\_\{n\}^\{2\}\)\\leq\_\{st\}L\_\{0\}\.
###### Fact 1
Leta\>0a\>0and consider the measurePa:=\(La,C\)\#\(χn2\)P\_\{a\}:=\(L\_\{a,C\}\)\_\{\\\#\}\(\\chi\_\{n\}^\{2\}\)onℝ\\mathbb\{R\}\. Then for anyρ\>0\\rho\>0, we can define a measurePb:=\(Lb,C\)\#\(χn2\)P\_\{b\}:=\(L\_\{b,C\}\)\_\{\\\#\}\(\\chi\_\{n\}^\{2\}\), whereb=a\+ρnb=a\+\\frac\{\\rho\}\{n\}\. This yields
Pb≤stPa,W1\(Pa,Pb\)=ρ\.\\displaystyle P\_\{b\}\\leq\_\{st\}P\_\{a\},\\quad W\_\{1\}\(P\_\{a\},P\_\{b\}\)=\\rho\.
###### Fact 2
LetPa=\(La,C\)\#\(χn2\)P\_\{a\}=\(L\_\{a,C\}\)\_\{\\\#\}\(\\chi\_\{n\}^\{2\}\)andPb=\(Lb,C\)\#\(χn2\)P\_\{b\}=\(L\_\{b,C\}\)\_\{\\\#\}\(\\chi\_\{n\}^\{2\}\), witha,b\>0a,b\>0andC∈ℝC\\in\\mathbb\{R\}\. Their RN derivative is
dPbdPa\(u\)=\(ab\)n/2exp\[\(C−u\)\(12a−12b\)\],u<C\.\\frac\{dP\_\{b\}\}\{dP\_\{a\}\}\(u\)=\\left\(\\frac\{a\}\{b\}\\right\)^\{n/2\}\\exp\\left\[\(C\-u\)\\left\(\\frac\{1\}\{2a\}\-\\frac\{1\}\{2b\}\\right\)\\right\],\\quad u<C\.Correspondingly, we obtain the constant control
c0:=ddulogdPbdPa\(u\)=−12a\+12b\.c\_\{0\}:=\\frac\{d\}\{du\}\\log\\frac\{dP\_\{b\}\}\{dP\_\{a\}\}\(u\)=\-\\frac\{1\}\{2a\}\+\\frac\{1\}\{2b\}\.
## IVNumerical Experiments
### IV\-AMixture of Two Gaussians
We first evaluate our method on a synthetic testbed consisting of a mixture of two Gaussians
𝐱∼12𝒩\(𝐦,I\)\+12𝒩\(−𝐦,I\)\.\\mathbf\{x\}\\sim\\tfrac\{1\}\{2\}\\mathcal\{N\}\(\\mathbf\{m\},I\)\+\\tfrac\{1\}\{2\}\\mathcal\{N\}\(\-\\mathbf\{m\},I\)\.\(16\)For this model the score function admits the closed\-form expression\[[24](https://arxiv.org/html/2609.12113#bib.bib10)\]
∇logpt\(𝐱\)=tanh\(𝐦t⋅𝐱\)𝐦t−𝐱,𝐦t=e−t𝐦\.\\nabla\\log p\_\{t\}\(\\mathbf\{x\}\)=\\tanh\(\\mathbf\{m\}\_\{t\}\\cdot\\mathbf\{x\}\)\\mathbf\{m\}\_\{t\}\-\\mathbf\{x\},\\quad\\mathbf\{m\}\_\{t\}=e^\{\-t\}\\mathbf\{m\}\.\(17\)
We choose𝐦=m^1n\\mathbf\{m\}=\\hat\{m\}1\_\{n\}, where1n1\_\{n\}denotes thenn\-dimensional vector of ones\. This closed\-form score allows us to solve the PF\-ODE \([3](https://arxiv.org/html/2609.12113#S1.E3)\) and the log\-FPK equation \([5](https://arxiv.org/html/2609.12113#S1.E5)\) directly without training a neural network\. Our goal is to guide the reverse diffusion process to generateρ\\rho\-outliers\.
#### Assessment of Assumption[1](https://arxiv.org/html/2609.12113#Thmassumption1)\.
We empirically assess the approximation ofL0L\_\{0\}by a translated−12χn2\-\\tfrac\{1\}\{2\}\\chi\_\{n\}^\{2\}distribution\. FOSD is assessed by sorting samples from two likelihood distributions and verifying that the respective inequality holds element\-wise\. To estimateL0L\_\{0\}, we sample 10,000 points by solving the reverse PF\-ODE \([3](https://arxiv.org/html/2609.12113#S1.E3)\), then use the log\-FPK equation \([5](https://arxiv.org/html/2609.12113#S1.E5)\) to compute their log\-likelihood\. Table[I](https://arxiv.org/html/2609.12113#S4.T1)reports their empiricalW1W\_\{1\}distance across several dimensions and values ofm^\\hat\{m\}\. The results indicate that the chi\-square model provides a good approximation of the log\-likelihood distribution\. An example empirical CDF is shown in Figure[2](https://arxiv.org/html/2609.12113#S4.F2)\.
Table I:Empirical assessment of the chi\-square approximation witha=1/2a=1/2\. Values report the empiricalW1W\_\{1\}distance betweenL0L\_\{0\}and the fitted chi\-square distribution\.
#### Outlier generation\.
For the remaining experiments we setm^=2\\hat\{m\}=2and diffusion horizonT=80T=80\. We vary the dimensionn∈\{1,4,16,64,256\}n\\in\\\{1,4,16,64,256\\\}and targetρ∈\{0\.2,0\.3,0\.4,0\.5\}\\rho\\in\\\{0\.2,0\.3,0\.4,0\.5\\\}\. In each experiment, we verify that FOSD is empirically satisfied, and report theW1W\_\{1\}distance between the generated likelihood distributionη0\\eta\_\{0\}and the referenceL0L\_\{0\}\.
Our results are posted in Table[II](https://arxiv.org/html/2609.12113#S4.T2)\. The generated samples match the targetρ\\rhovalues while maintaining stochastic dominance\. Figure[2](https://arxiv.org/html/2609.12113#S4.F2)compares histograms of samples generated by the uncontrolled and controlled PF\-ODE in the one\-dimensional case\. The controlled dynamics produce samples concentrated in lower\-likelihood regions such as the tails and the low\-density region between the mixture modes\.
Table II:Outlier generation results for the Gaussian mixture model\. Values report the empiricalW1W\_\{1\}distances between the generated and reference likelihood distributions\. The measured distances closely match the prescribed target valuesρ\\rho, demonstrating accurate control over outlier magnitude while satisfying stochastic dominance in every experiment\.
Figure 1:Histogram comparison of PF\-ODE samples in the 1D Gaussian mixture example withρ=0\.5\\rho=0\.5\.
Figure 2:Empirical CDFs of the likelihood distributionL0L\_\{0\}and the target distributionη0\\eta\_\{0\}in the 4D example withρ=0\.5\\rho=0\.5\.
### IV\-BImage Data: CIFAR\-10
We next evaluate the method on the CIFAR\-10 dataset using the pretrained diffusion modelgoogle/ddpm\-cifar10\-32\[[17](https://arxiv.org/html/2609.12113#bib.bib1)\]\. The likelihood distributionL0L\_\{0\}is approximated using samples generated by the reverse PF\-ODE together with the log\-FPK equation \([5](https://arxiv.org/html/2609.12113#S1.E5)\)\.
Empirically,L0L\_\{0\}stochastically dominates a translated−12χn2\-\\tfrac\{1\}\{2\}\\chi\_\{n\}^\{2\}distribution, supporting Assumption[1](https://arxiv.org/html/2609.12113#Thmassumption1)\. Their empiricalW1W\_\{1\}distance is9\.1379\.137\. We vary the targetρ\\rhoacross\{50,100,150,200,250\}\\\{50,100,150,200,250\\\}and generateρ\\rho\-outliers by modifying the reverse diffusion dynamics\. Table[III](https://arxiv.org/html/2609.12113#S4.T3)reports the empiricalW1W\_\{1\}distances between the generated likelihood distribution and the reference distribution\. Although the realizedW1W\_\{1\}shifts do not exactly match the prescribedρ\\rhovalues, they increase monotonically withρ\\rho\. Thus, even when the score is represented by a neural network, the control parameter provides a consistent mechanism for adjusting the degree of likelihood shift\.
Figure[3](https://arxiv.org/html/2609.12113#S4.F3)confirms that increasingρ\\rhoprogressively shifts the generated likelihood distribution toward lower values\. At moderate control strengths, the samples in Figures[4\(a\)](https://arxiv.org/html/2609.12113#S4.F4.sf1)and[4\(b\)](https://arxiv.org/html/2609.12113#S4.F4.sf2)remain visually consistent with CIFAR\-10, suggesting that the controller can access lower\-likelihood regions without immediately destroying the learned data structure\.
Figure 3:Empirical CDFs of likelihood distributions for different values ofρ\\rho\.Table III:Outlier generation on CIFAR\-10\. Increasing the target parameterρ\\rhoproduces progressively largerW1W\_\{1\}shifts in the likelihood distribution, demonstrating that the proposed controller remains effective on image data\.\(a\)Generated samples withρ=50\\rho=50\.
\(b\)Generated samples withρ=100\\rho=100\.
Figure 4:Generatedρ\\rho\-outliers in the CIFAR\-10 image dataset\. Although the likelihood values are lower, as demonstrated in Figure[3](https://arxiv.org/html/2609.12113#S4.F3), the generated samples remain visually consistent with the underlying data distribution\.Overall, these experiments demonstrate that the proposed framework can steer diffusion sampling toward prescribed likelihood statistics\. In both synthetic and image datasets, the controlled PF\-ODE successfully shifts the likelihood distribution toward lower\-probability regions while preserving the structure of the data distribution\.
## VConclusion
In this work, we introduced a distributional perspective on outliers motivated by the probabilistic structure of diffusion models\. Instead of defining outliers as individual low\-likelihood samples, we formalized them through the distribution of log\-likelihood values and proposed a notion ofρ\\rho\-outlier based on the discrepancy between likelihood distributions\. Using this formulation, we developed a method for generating outliers by modifying the reverse\-time diffusion dynamics\. The key insight is that likelihood reweighting induces a simple modification of the score function via the RN derivative, which motivates the controlled PF\-ODE used for sampling\. Numerical experiments support the theoretical predictions on synthetic and image datasets\. Future work will explore broader applications of this framework to distribution steering\.
## References
- \[1\]V\. S\. Bawa\(1975\)Optimal rules for ordering uncertain prospects\.Journal of Financial economics2\(1\),pp\. 95–121\.Cited by:[Definition 4](https://arxiv.org/html/2609.12113#Thmdefinition4.3)\.
- \[2\]A\. Blattmann, T\. Dockhorn, S\. Kulal, D\. Mendelevitch, M\. Kilian, D\. Lorenz, Y\. Levi, Z\. English, V\. Voleti, A\. Letts,et al\.\(2023\)Stable video diffusion: scaling latent video diffusion models to large datasets\.arXiv preprint arXiv:2311\.15127\.Cited by:[Score\-based Outlier Generation via Controlling the Radon\-Nikodym Derivative](https://arxiv.org/html/2609.12113#p3.1)\.
- \[3\]V\. I\. Bogachev, N\. V\. Krylov, M\. Röckner, and S\. V\. Shaposhnikov\(2015\)Fokker–planck–kolmogorov equations\.Vol\.207,Mathematical Surveys and Monographs\.Cited by:[§I\-B](https://arxiv.org/html/2609.12113#S1.SS2.p1.1)\.
- \[4\]V\. I\. Bogachev\(2018\)Ornstein–uhlenbeck operators and semigroups\.Russian Mathematical Surveys73\(2\),pp\. 191–260\.Cited by:[Theorem 1](https://arxiv.org/html/2609.12113#Thmtheorem1.3)\.
- \[5\]Y\. Chen\(2025\)Log\-sobolev inequalities and markov semigroups\.ETH Zurich\.Note:Lecture 2: Ornstein–Uhlenbeck Semigroup and Gaussian Log\-Sobolev Inequality, ETH Zurich, Week 3–4Cited by:[§III](https://arxiv.org/html/2609.12113#S3.p3.7.1)\.
- \[6\]R\. Chertovskih, N\. Pogodaev, M\. Staritsyn, and A\. P\. Aguiar\(2024\)Optimal control of diffusion processes: infinite\-order variational analysis and numerical solution\.IEEE Control Systems Letters8,pp\. 1469–1474\.Cited by:[§I\-B](https://arxiv.org/html/2609.12113#S1.SS2.p1.2)\.
- \[7\]C\. Dellacherie and P\. Meyer\(1978\)Probabilities and potential, vol\. 29 of north\-holland mathematics studies\.North\-Holland Publishing Co\., Amsterdam\.Cited by:[§II\-B](https://arxiv.org/html/2609.12113#S2.SS2.p2.3.1)\.
- \[8\]D\. Dentcheva and A\. Ruszczynski\(2003\)Optimization with stochastic dominance constraints\.SIAM Journal on Optimization14\(2\),pp\. 548–566\.Cited by:[§II\-A](https://arxiv.org/html/2609.12113#S2.SS1.p4.1)\.
- \[9\]D\. Dentcheva and A\. Ruszczyński\(2004\)Optimality and duality theory for stochastic optimization problems with nonlinear dominance constraints\.Mathematical Programming99\(2\),pp\. 329–350\.Cited by:[§II\-A](https://arxiv.org/html/2609.12113#S2.SS1.p4.1)\.
- \[10\]D\. Dentcheva, M\. Ye, and Y\. Yi\(2022\)Risk\-averse sequential decision problems with time\-consistent stochastic dominance constraints\.In2022 IEEE 61st Conference on Decision and Control \(CDC\),pp\. 3605–3610\.Cited by:[§II\-A](https://arxiv.org/html/2609.12113#S2.SS1.p4.1)\.
- \[11\]X\. Du, Y\. Sun, J\. Zhu, and Y\. Li\(2023\)Dream the impossible: outlier imagination with diffusion models\.Advances in Neural Information Processing Systems36,pp\. 60878–60901\.Cited by:[Score\-based Outlier Generation via Controlling the Radon\-Nikodym Derivative](https://arxiv.org/html/2609.12113#p2.1)\.
- \[12\]A\. Fleig and R\. Guglielmi\(2017\)Optimal control of the fokker–planck equation with space\-dependent controls\.Journal of Optimization Theory and Applications174\(2\),pp\. 408–427\.Cited by:[§I\-B](https://arxiv.org/html/2609.12113#S1.SS2.p1.2)\.
- \[13\]P\. H\. Franses and D\. Van Dijk\(2000\)Non\-linear time series models in empirical finance\.Cambridge university press\.Cited by:[Score\-based Outlier Generation via Controlling the Radon\-Nikodym Derivative](https://arxiv.org/html/2609.12113#p1.1)\.
- \[14\]J\. Gu, X\. Zhang, and G\. Wang\(2025\)Beyond the norm: a survey of synthetic data generation for rare events\.arXiv preprint arXiv:2506\.06380\.Cited by:[Score\-based Outlier Generation via Controlling the Radon\-Nikodym Derivative](https://arxiv.org/html/2609.12113#p2.1)\.
- \[15\]M\. Hermite\(1864\)Sur un nouveau développement en série des fonctions\.Imprimerie de Gauthier\-Villars\.Cited by:[§III](https://arxiv.org/html/2609.12113#S3.p3.4.1),[Theorem 1](https://arxiv.org/html/2609.12113#Thmtheorem1.p1.1.1)\.
- \[16\]J\. Ho, W\. Chan, C\. Saharia, J\. Whang, R\. Gao, A\. Gritsenko, D\. P\. Kingma, B\. Poole, M\. Norouzi, D\. J\. Fleet,et al\.\(2022\)Imagen video: high definition video generation with diffusion models\.arXiv preprint arXiv:2210\.02303\.Cited by:[Score\-based Outlier Generation via Controlling the Radon\-Nikodym Derivative](https://arxiv.org/html/2609.12113#p3.1)\.
- \[17\]J\. Ho, A\. Jain, and P\. Abbeel\(2020\)Denoising diffusion probabilistic models\.InAdvances in Neural Information Processing Systems,External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2020/file/4c5bcfec8584af0d967f1ab10179ca4b-Paper.pdf)Cited by:[§IV\-B](https://arxiv.org/html/2609.12113#S4.SS2.p1.1),[Score\-based Outlier Generation via Controlling the Radon\-Nikodym Derivative](https://arxiv.org/html/2609.12113#p3.1)\.
- \[18\]B\. Hu, A\. Saragadam, A\. Layton, and H\. Chen\(2024\)Synthetic data from diffusion models improves drug discovery prediction\.In2024 IEEE International Conference on Bioinformatics and Biomedicine \(BIBM\),pp\. 6278–6285\.Cited by:[Score\-based Outlier Generation via Controlling the Radon\-Nikodym Derivative](https://arxiv.org/html/2609.12113#p1.1)\.
- \[19\]C\. Lai, Y\. Takida, N\. Murata, T\. Uesaka, Y\. Mitsufuji, and S\. Ermon\(2023\)FP\-Diffusion: improving score\-based diffusion models by enforcing the underlying score Fokker\-Planck equation\.InProceedings of the 40th International Conference on Machine LearningInternational Conference on Machine Learning,Cited by:[Proposition 1](https://arxiv.org/html/2609.12113#Thmproposition1.3)\.
- \[20\]D\. Leao Jr, M\. Fragoso, and P\. Ruffino\(2004\)Regular conditional probability, disintegration of probability and radon spaces\.Proyecciones \(Antofagasta\)23\(1\),pp\. 15–29\.Cited by:[Definition 3](https://arxiv.org/html/2609.12113#Thmdefinition3.p1.2.1)\.
- \[21\]W\. Peebles and S\. Xie\(2023\)Scalable diffusion models with transformers\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 4195–4205\.Cited by:[Score\-based Outlier Generation via Controlling the Radon\-Nikodym Derivative](https://arxiv.org/html/2609.12113#p3.1)\.
- \[22\]R\. Rombach, A\. Blattmann, D\. Lorenz, P\. Esser, and B\. Ommer\(2022\)High\-resolution image synthesis with latent diffusion models\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 10684–10695\.Cited by:[Score\-based Outlier Generation via Controlling the Radon\-Nikodym Derivative](https://arxiv.org/html/2609.12113#p3.1)\.
- \[23\]A\. Sauer, F\. Boesel, T\. Dockhorn, A\. Blattmann, P\. Esser, and R\. Rombach\(2024\)Fast high\-resolution image synthesis with latent adversarial diffusion distillation\.InSIGGRAPH Asia 2024 Conference Papers,pp\. 1–11\.Cited by:[Score\-based Outlier Generation via Controlling the Radon\-Nikodym Derivative](https://arxiv.org/html/2609.12113#p3.1)\.
- \[24\]K\. Shah, S\. Chen, and A\. Klivans\(2023\)Learning mixtures of gaussians using the ddpm objective\.Advances in Neural Information Processing Systems36,pp\. 19636–19649\.Cited by:[§IV\-A](https://arxiv.org/html/2609.12113#S4.SS1.p1.2)\.
- \[25\]C\. Sinigaglia, A\. Manzoni, and F\. Braghin\(2022\)Density control of large\-scale particles swarm through pde\-constrained optimization\.IEEE Transactions on Robotics38\(6\),pp\. 3530–3549\.Cited by:[§I\-B](https://arxiv.org/html/2609.12113#S1.SS2.p1.2)\.
- \[26\]Y\. Song, J\. Sohl\-Dickstein, D\. P\. Kingma, A\. Kumar, S\. Ermon, and B\. Poole\(2021\)Score\-based generative modeling through stochastic differential equations\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=PxTIG12RRHS)Cited by:[§I\-A](https://arxiv.org/html/2609.12113#S1.SS1.p1.1),[§I\-A](https://arxiv.org/html/2609.12113#S1.SS1.p1.3)\.
- \[27\]L\. Tao, X\. Du, X\. Zhu, and Y\. Li\(2023\)Non\-parametric outlier synthesis\.arXiv preprint arXiv:2303\.02966\.Cited by:[Score\-based Outlier Generation via Controlling the Radon\-Nikodym Derivative](https://arxiv.org/html/2609.12113#p2.1)\.
- \[28\]C\. Villani\(2021\)Topics in optimal transportation\.Vol\.58,American Mathematical Soc\.\.Cited by:[Remark 1](https://arxiv.org/html/2609.12113#Thmremark1.p2.1.1)\.Similar Articles
Generator-Guided Inverse Sampling for L\'evy-Driven Generative Models
This paper studies inverse sampling for Lévy-driven generative models, proposing a structured reverse sampler that decomposes dynamics into diffusion, small jump, and large jump components, with neural networks amortizing jump rates. The method is applied to OFDM-SISO channel estimation under mixed Gaussian and impulsive noise.
Language Generation as Optimal Control: Closed-Loop Diffusion in Latent Control Space
This paper reformulates language generation as a stochastic optimal control problem, addressing limitations of autoregressive and diffusion models, and proposes a closed-loop diffusion method in latent control space using Flow Matching, achieving high-fidelity generation and efficient parallel sampling.
Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning
This paper introduces DOSER, a framework using diffusion models for out-of-distribution detection and selective regularization in offline reinforcement learning. It aims to improve performance on static datasets by distinguishing between beneficial and detrimental OOD actions.
Catastrophic Compositional Generation: Why Vanilla Diffusion Models Fail to Extrapolate
This paper argues that vanilla conditional diffusion models fundamentally fail at compositional generation when the target distribution is out-of-distribution, due to score estimation error, and that inference-time corrections cannot fully compensate.
Specificity-Aware Diffusion Steering via Variance-Reduced Sequential Monte Carlo
The paper introduces a specificity-aware diffusion steering method that designs an explicit target distribution from an overlap-based objective and samples it with a variance-reduced Sequential Monte Carlo sampler, improving negative-guidance effectiveness and sampling stability across image generation and peptide-MHC binder tasks.