Score-based Outlier Generation via Controlling the Radon-Nikodym Derivative

arXiv cs.LG Papers

Summary

The paper introduces a method for controlled outlier generation using diffusion models by manipulating likelihood through Radon-Nikodym derivatives, enabling the creation of low-likelihood samples without retraining.

arXiv:2609.12113v1 Announce Type: new Abstract: Outliers are important for stress-testing algorithms and understanding system behaviour under rare conditions. Despite being commonly described as low-likelihood events, existing generative approaches rarely control likelihood explicitly. In this work, we introduce a measure-theoretic notion of outliers based on the distribution of log-likelihood values, which is guaranteed to assign higher probability mass to low-likelihood events with a specifiable magnitude. Building on this formulation, we derive how likelihood reweighting modifies the diffusion score and use this relation to motivate a controlled modification of the reverse-time dynamics. In particular, likelihood reweighting implies a scaling of the score function with a control term derived from the Radon-Nikodym derivative of the likelihood distributions. Correspondingly, the updated score function can be obtained with no retraining of the diffusion model. We exploit the Ornstein-Uhlenbeck semigroup underlying diffusion models to motivate an exponentially interpolated controller which approximates the true control. Experiments demonstrate controlled generation of low-likelihood samples while remaining consistent with the data geometry.
Original Article
View Cached Full Text

Cached at: 09/14/26, 08:34 AM

# Score-based Outlier Generationvia Controlling the Radon-Nikodym Derivative
Source: [https://arxiv.org/html/2609.12113](https://arxiv.org/html/2609.12113)
Tristan MilneKry Yik\-Chau LuiStephanie HazlewoodJun Liu††thanks:This research was supported by Mitacs and the Royal Bank of Canada under the Mitacs Accelerate program award, Application Reference IT49478\.††thanks:Amartya Mukherjee and Jun Liu are with the Department of Applied Mathematics, University of Waterloo, Waterloo, Ontario, Canada N2L 3G1 \(email:\(a29mukhe,j\.liu\)@uwaterloo\.ca\)\.††thanks:Tristan Milne, Stephanie Hazlewood, and Kry Yik\-Chau Lui are with the Royal Bank of Canada, 1 Place Ville Marie, Montreal, Quebec, H3C 3A9 \(email:\(tristan\.milne,stephanie\.hazlewood\)@rbc\.com, yikchau\.y\.lui@borealisai\.com\)\.

###### Abstract

Outliers are important for stress\-testing algorithms and understanding system behaviour under rare conditions\. Despite being commonly described as low\-likelihood events, existing generative approaches rarely control likelihood explicitly\. In this work, we introduce a measure\-theoretic notion of outliers based on the distribution of log\-likelihood values, which is guaranteed to assign higher probability mass to low\-likelihood events with a specifiable magnitude\. Building on this formulation, we derive how likelihood reweighting modifies the diffusion score and use this relation to motivate a controlled modification of the reverse\-time dynamics\. In particular, likelihood reweighting implies a scaling of the score function with a control term derived from the Radon\-Nikodym derivative of the likelihood distributions\. Correspondingly, the updated score function can be obtained with no retraining of the diffusion model\. We exploit the Ornstein–Uhlenbeck semigroup underlying diffusion models to motivate an exponentially interpolated controller which approximates the true control\. Experiments demonstrate controlled generation of low\-likelihood samples while remaining consistent with the data geometry\.

Outliers play an important role in evaluating the reliability of algorithms and decision\-making systems\. In many applications, it is important to understand how systems behave under rare or atypical conditions that deviate from the patterns commonly observed in training or historical data\. For example, in time series such as financial data\[[13](https://arxiv.org/html/2609.12113#bib.bib12)\], outliers model how capital markets respond to anomalous shifts in the financial landscape\. In tabular data such as electronic health records\[[18](https://arxiv.org/html/2609.12113#bib.bib11)\], outliers are crucial for studying patients who deviate from typical disease profiles and for developing treatment strategies tailored to these atypical cases\. Generating such rare scenarios is therefore essential for stress\-testing algorithms, improving model robustness, and understanding system behaviour under distributional shifts\.

Existing approaches to outlier generation suffer from two important limitations\. First, although outliers are commonly motivated as low\-likelihood samples under a reference distribution\[[27](https://arxiv.org/html/2609.12113#bib.bib16),[11](https://arxiv.org/html/2609.12113#bib.bib15)\], most methods rely on heuristic notions such as reconstruction error, latent\-space distance, or classifier uncertainty rather than explicitly controlling likelihood itself\. Hence, they provide little formal interpretability over the statistical rarity of the generated samples\. Second, many existing approaches require specialized architectures or training objectives designed specifically for outlier synthesis\[[14](https://arxiv.org/html/2609.12113#bib.bib28)\]\. Outlier generation is therefore treated as a separate learning problem rather than a controllable sampling problem\.

Diffusion models \(DMs\) have emerged as a powerful framework in generative modelling, achieving remarkable success in domains such as image synthesis\[[17](https://arxiv.org/html/2609.12113#bib.bib1),[22](https://arxiv.org/html/2609.12113#bib.bib6),[23](https://arxiv.org/html/2609.12113#bib.bib7),[21](https://arxiv.org/html/2609.12113#bib.bib5)\]and video generation\[[16](https://arxiv.org/html/2609.12113#bib.bib8),[2](https://arxiv.org/html/2609.12113#bib.bib9)\]\. These models are typically trained to reverse a stochastic process that gradually transforms samples from a data distribution into noise via a learned score function\. These properties also make DMs particularly attractive for outlier generation\. DMs admit tractable likelihood estimation through the probability\-flow ordinary differential equation \(PF\-ODE\) and evolve according to continuous\-time dynamics that can be controlled\. In fact, as we will show, a DM trained on regular data can be steered toward low\-likelihood regions at inference time, without retraining or fine\-tuning\.

In this work, we propose a distributional control\-theoretic framework for outlier generation in DMs\. Unlike previous approaches that identify individual anomalous samples, we control the likelihood distribution\. The log\-likelihood map induces a one\-dimensional pushforward measure, and generating outliers corresponds to steering this measure towards lower values\. Our key observation is that likelihood reweighting implies a scaling of the score function with a controller determined by a Radon–Nikodym \(RN\) derivative on likelihood spaces\. This exact pointwise identity motivates a controlled PF\-ODE\. To obtain a practical controller, we exploit the exponential convergence of the Ornstein–Uhlenbeck \(OU\) semigroup that underlies the forward diffusion process\. We use an exponentially decaying approximation to the initial likelihood\-reweighting control\.

Numerical experiments on Gaussian mixture examples and the CIFAR\-10 dataset demonstrate that the proposed method successfully steers diffusion sampling toward low\-likelihood regions while remaining consistent with the underlying data geometry\. The results suggest that DMs can be interpreted as controllable systems and opens the door to distribution steering and generative modelling under functional constraints\.

## IBackground

### I\-AScore\-Based Diffusion Models

The diffusion process\[[26](https://arxiv.org/html/2609.12113#bib.bib3)\]can be parameterized in continuous time by the following Ornstein\-Uhlenbeck \(OU\) SDE

d​𝐱t=−𝐱t​d​t\+2​d​𝐰t,𝐱0∼pd​a​t​ad\\mathbf\{x\}\_\{t\}=\-\\mathbf\{x\}\_\{t\}dt\+\\sqrt\{2\}d\\mathbf\{w\}\_\{t\},\\quad\\mathbf\{x\}\_\{0\}\\sim p\_\{data\}\(1\)fort∈\[0,T\]t\\in\[0,T\], where𝐱0\\mathbf\{x\}\_\{0\}is sampled from a data distribution and𝐰t\\mathbf\{w\}\_\{t\}is Brownian motion\. Asttgrows,𝐱t\\mathbf\{x\}\_\{t\}diffuses from the clean data𝐱0\\mathbf\{x\}\_\{0\}into Gaussian noise\. We then generate realistic data by sampling Gaussian noise𝐱T\\mathbf\{x\}\_\{T\}and solving the reverse\-time SDE

d𝐱t=\[−𝐱t−2∇logpt\(𝐱t\)\]dt\+2d𝐰¯t,d\\mathbf\{x\}\_\{t\}=\[\-\\mathbf\{x\}\_\{t\}\-2\\nabla\\log p\_\{t\}\(\\mathbf\{x\}\_\{t\}\)\]dt\+\\sqrt\{2\}d\\overline\{\\mathbf\{w\}\}\_\{t\},\(2\)where𝐰¯t\\overline\{\\mathbf\{w\}\}\_\{t\}denotes a reverse Brownian motion\. Alternatively, we can solve the probability flow ODE \(PF\-ODE\[[26](https://arxiv.org/html/2609.12113#bib.bib3)\]\) in reverse\-time

𝐱˙t=−𝐱t−∇log⁡pt​\(𝐱t\),𝐱T∼𝒩⁡\(0,I\)\.\\dot\{\\mathbf\{x\}\}\_\{t\}=\-\\mathbf\{x\}\_\{t\}\-\\nabla\\log p\_\{t\}\(\\mathbf\{x\}\_\{t\}\),\\quad\\mathbf\{x\}\_\{T\}\\sim\\mathcal\{N\}\(0,I\)\.\(3\)It is common practice in the DM literature to train a neural network𝐬θ​\(𝐱t,t\)\\mathbf\{s\}\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t\)to approximate∇log⁡pt​\(𝐱t\)\\nabla\\log p\_\{t\}\(\\mathbf\{x\}\_\{t\}\), from which realistic training data can then be generated\.

### I\-BFokker\-Planck\-Kolmogorov Equation

The Fokker\-Planck\-Kolmogorov \(FPK\)\[[3](https://arxiv.org/html/2609.12113#bib.bib4)\]equation that governs the evolution of the underlying distribution from the SDEs in Equation[1](https://arxiv.org/html/2609.12113#S1.E1)\(forward time\) and Equation[2](https://arxiv.org/html/2609.12113#S1.E2)\(reverse time\) is given by the partial differential equation \(PDE\):

d​pt​\(𝐱t\)d​t=∇⋅\[𝐱t​pt​\(𝐱t\)\]\+Δ​pt​\(𝐱t\),\\displaystyle\\frac\{dp\_\{t\}\(\\mathbf\{x\}\_\{t\}\)\}\{dt\}=\\nabla\\cdot\[\\mathbf\{x\}\_\{t\}p\_\{t\}\(\\mathbf\{x\}\_\{t\}\)\]\+\\Delta p\_\{t\}\(\\mathbf\{x\}\_\{t\}\),\(4\)where∇⁣⋅\\nabla\\cdotis the divergence andΔ\\Deltais the Laplacian operator\. Recently, density steering via controlling the FPK equation has been of interest to the control community\[[12](https://arxiv.org/html/2609.12113#bib.bib25),[6](https://arxiv.org/html/2609.12113#bib.bib26),[25](https://arxiv.org/html/2609.12113#bib.bib27)\]\. It is often convenient to express this evolution in terms of the log\-density to obtain log\-likelihoods\. The following result gives the corresponding PDE satisfied by the log\-density\.

###### Proposition 1\(Proposition 3\.1 of\[[19](https://arxiv.org/html/2609.12113#bib.bib2)\]\)

Assume the ground truth densitypt​\(𝐱\)p\_\{t\}\(\\mathbf\{x\}\)is sufficiently smooth onℝn×\[0,T\]\\mathbb\{R\}^\{n\}\\times\[0,T\]with its log\-density denoted aslt​\(𝐱\):=log⁡pt​\(𝐱\)l\_\{t\}\(\\mathbf\{x\}\):=\\log p\_\{t\}\(\\mathbf\{x\}\)\. Then for all\(𝐱,t\)\(\\mathbf\{x\},t\), its log\-density satisfies the PDE

∂tlt​\(𝐱\)=𝐱⋅∇lt​\(𝐱\)\+n\+Δ​lt​\(𝐱\)\+‖∇lt​\(𝐱\)‖2\.\\partial\_\{t\}l\_\{t\}\(\\mathbf\{x\}\)=\\mathbf\{x\}\\cdot\\nabla l\_\{t\}\(\\mathbf\{x\}\)\+n\+\\Delta l\_\{t\}\(\\mathbf\{x\}\)\+\\\|\\nabla l\_\{t\}\(\\mathbf\{x\}\)\\\|^\{2\}\.\(5\)

### I\-COrnstein\-Uhlenbeck Operator

The forward diffusion process underlying score\-based models corresponds to OU dynamics, whose generator plays an important role in our analysis of likelihood evolution\.

###### Definition 1\(Ornstein\-Uhlenbeck \(OU\) Generator\)

The OU generatorℒ\\mathcal\{L\}acts on smooth functionsf∈C2​\(ℝn\)f\\in C^\{2\}\(\\mathbb\{R\}^\{n\}\)by

ℒ​f​\(𝐱\)=Δ​f​\(𝐱\)−𝐱⋅∇f​\(𝐱\)\.\\mathcal\{L\}f\(\\mathbf\{x\}\)=\\Delta f\(\\mathbf\{x\}\)\-\\mathbf\{x\}\\cdot\\nabla f\(\\mathbf\{x\}\)\.

###### Theorem 1\(Theorem 3\.8 of\[[4](https://arxiv.org/html/2609.12113#bib.bib20)\]\)

Letγ:=𝒩⁡\(0,I\)\\gamma:=\\mathcal\{N\}\(0,I\)be the standard Gaussian measure\. In the Hilbert spaceL2​\(γ\),L^\{2\}\(\\gamma\),ℒ\\mathcal\{L\}is self\-adjoint and has discrete spectrum\{0,−1,−2,…\}\\\{0,\-1,\-2,\.\.\.\\\}of non\-positive integers\. The eigenfunctions are the Hermite polynomials\[[15](https://arxiv.org/html/2609.12113#bib.bib19)\]\(Hα\)α∈ℕn\(H\_\{\\alpha\}\)\_\{\\alpha\\in\\mathbb\{N\}^\{n\}\}, orthogonal inL2​\(γ\),L^\{2\}\(\\gamma\),satisfyingℒ​Hα=−\|α\|​Hα\\mathcal\{L\}H\_\{\\alpha\}=\-\|\\alpha\|H\_\{\\alpha\}\.

## IIControlled Likelihood Generation

Outliers are commonly described as samples with low likelihood under a reference data distribution\. In likelihood\-based generative modelling, this intuition translates into identifying regions where the log\-likelihoodl⁡\(𝐱\):=log⁡p⁡\(𝐱\)l\(\\mathbf\{x\}\):=\\log p\(\\mathbf\{x\}\)is small relative to typical data\. However, DMs learn*probability measures*, not individual samples\. Thus, if we wish to generate outliers in a principled way, we must define outliers at the level of distributions rather than individual points\.

In this section, we characterize outliers through the distribution of log\-likelihood values induced by a probability measure\. Since outliers correspond to samples with unusually low likelihood, our goal is to steer the likelihood distribution toward lower values\. This motivates the use of stochastic ordering to formalize the notion that the target likelihood distribution places more probability mass on lower\-likelihood regions\. Our main theoretical result shows that likelihood reweighting induces a suitable control through a simple RN structure, resulting in a multiplicative modification of the score function\. Motivated by this relation, we use the modified score in a controlled PF\-ODE and empirically evaluate its ability to generate low\-likelihood samples\.

### II\-AProblem Formulation and Definition of Outlier

Given a data distributionμ\\muonℝn\\mathbb\{R\}^\{n\}with densitypp, the log\-likelihood functionl⁡\(𝐱\)=log⁡p⁡\(𝐱\)l\(\\mathbf\{x\}\)=\\log p\(\\mathbf\{x\}\)induces a scalar random variablel⁡\(𝐱\)l\(\\mathbf\{x\}\)when𝐱∼μ\\mathbf\{x\}\\sim\\mu\. The pushforward measureL:=l\#​μL:=l\_\{\\\#\}\\mutherefore describes the distribution of likelihood values under the model\. Shifting this distribution toward lower values corresponds to generating samples that are globally less likely\. By modifying this one\-dimensional likelihood distribution while preserving the conditional structure of the data given its likelihood level, we obtain a mechanism for generating structured outliers that remain consistent with the underlying data geometry\. We now formalize these notions\.

###### Definition 2\(Log\-likelihood pushforward measure\)

Letμ\\mube a probability measure onℝn\\mathbb\{R\}^\{n\}with densitypp\. Define the log\-likelihood functionl⁡\(𝐱\):=log⁡p⁡\(𝐱\)\.l\(\\mathbf\{x\}\):=\\log p\(\\mathbf\{x\}\)\.The pushforward measure ofμ\\muunderllis the probability measureL:=l\#​μL:=l\_\{\\\#\}\\muonℝ\\mathbb\{R\}\.

###### Definition 3\(Likelihood\-reweighted measure\)

Letμ\\mube a probability measure onℝn\\mathbb\{R\}^\{n\}with densityppand log\-likelihoodl⁡\(𝐱\)=log⁡p⁡\(𝐱\)l\(\\mathbf\{x\}\)=\\log p\(\\mathbf\{x\}\)\. LetL=l\#​μL=l\_\{\\\#\}\\mu\. Letη\\etabe a probability measure onℝ\\mathbb\{R\}such thatη≪L\\eta\\ll L\. A probability measureν\\nuonℝn\\mathbb\{R\}^\{n\}is called a likelihood\-reweighted measure ofμ\\muwith targetη\\etaifν\\nuadmits the disintegration

ν⁡\(A\)=∫ℝμ⁡\(A∣l⁡\(𝐱\)=u\)​η​\(𝑑u\),A⊂ℝn​Borel,\\nu\(A\)=\\int\_\{\\mathbb\{R\}\}\\mu\(A\\mid l\(\\mathbf\{x\}\)=u\)\\,\\eta\(du\),\\quad A\\subset\\mathbb\{R\}^\{n\}\\text\{ Borel\},whereμ\(⋅∣l\(𝐱\)=u\)\\mu\(\\cdot\\mid l\(\\mathbf\{x\}\)=u\)denotes a regular conditional probability ofμ\\mugivenl⁡\(𝐱\)=ul\(\\mathbf\{x\}\)=u, which exists for Borel probability measures on Polish spaces\[[20](https://arxiv.org/html/2609.12113#bib.bib13)\]\.

As a consequence, ifν\\nuis a likelihood\-reweighted measure ofμ\\muwith targetη\\eta, thenl\#​ν=ηl\_\{\\\#\}\\nu=\\eta\.

###### Definition 4\(First\-order stochastic dominance \(FOSD\)\[[1](https://arxiv.org/html/2609.12113#bib.bib14)\]\)

Letμ\\muandν\\nube probability measures onℝ\\mathbb\{R\}\. We say thatν\\nufirst\-order stochastically dominatesμ\\muand write

if and only ifμ⁡\(\[x,∞\)\)≤ν⁡\(\[x,∞\)\)\\mu\(\[x,\\infty\)\)\\leq\\nu\(\[x,\\infty\)\)for allx∈ℝx\\in\\mathbb\{R\}\.

Equivalently, if we letFμF\_\{\\mu\}andFνF\_\{\\nu\}be cumulative distribution functions \(CDFs\) ofμ\\muandν\\nurespectively, then

μ≤s​tν⇔Fμ\(u\)≥Fν\(u\)for allu∈ℝ\.\\mu\\leq\_\{st\}\\nu\\iff F\_\{\\mu\}\(u\)\\geq F\_\{\\nu\}\(u\)\\quad\\text\{for all \}u\\in\\mathbb\{R\}\.
Stochastic dominance constraints are studied in stochastic optimization and decision theory as a way of enforcing preference relations between random outcomes\[[8](https://arxiv.org/html/2609.12113#bib.bib23),[9](https://arxiv.org/html/2609.12113#bib.bib22),[10](https://arxiv.org/html/2609.12113#bib.bib24)\]\.

###### Definition 5\(ρ\\rho\-outlier measure\)

Letμ\\mube a probability measure onℝn\\mathbb\{R\}^\{n\}with log\-likelihood pushforwardL=l\#​μL=l\_\{\\\#\}\\mu\. A likelihood\-reweighted measureν\\nuwith targetη\\etais called a*ρ\\rho\-outlier measure*if

\(1\)η≤s​tL,and \(2\)W1\(L,η\)≥ρ,\\text\{\(1\) \}\\eta\\leq\_\{st\}L,\\quad\\text\{ and \\hskip 10\.22217pt\(2\) \}W\_\{1\}\(L,\\eta\)\\geq\\rho,where≤s​t\\leq\_\{st\}represents FOSD \(see Definition[4](https://arxiv.org/html/2609.12113#Thmdefinition4)\) andW1​\(⋅,⋅\)W\_\{1\}\(\\cdot,\\cdot\)is the Wasserstein\-1 distance \(see Equation \([6](https://arxiv.org/html/2609.12113#S2.E6)\) below\)\.

To generate samples from the likelihood\-reweighted measureν\\nu, we seek to characterize its score∇log⁡q\\nabla\\log qin terms of the score∇log⁡p\\nabla\\log pof the reference distribution, whereppandqqare the densities ofμ\\muandν\\nurespectively\. This allows us to use an existing diffusion score model while modifying its sampling dynamics through a likelihood\-dependent correction\.

### II\-BRadon\-Nikodym Structure of Likelihood Reweighting

Based on our formulation of likelihood\-reweighted distributions \(Definition[3](https://arxiv.org/html/2609.12113#Thmdefinition3)\), we derive some further properties\.

###### Theorem 2

Letμ\\mube a probability measure onℝn\\mathbb\{R\}^\{n\}with densityppand log\-likelihoodl⁡\(𝐱\)=log⁡p⁡\(𝐱\)l\(\\mathbf\{x\}\)=\\log p\(\\mathbf\{x\}\)\. LetL=l\#​μL=l\_\{\\\#\}\\mube the pushforward measure ofμ\\muunderll\. Letη\\etabe a probability measure onℝ\\mathbb\{R\}such thatη≪L\\eta\\ll L, and define

z​\(u\):=d​ηd​L​\(u\)\.z\(u\):=\\frac\{d\\eta\}\{dL\}\(u\)\.
Letν\\nube the likelihood\-reweighted measure defined by

ν⁡\(A\):=∫ℝμ⁡\(A∣l⁡\(𝐱\)=u\)​η​\(𝑑u\),A⊂ℝn​Borel\.\\nu\(A\):=\\int\_\{\\mathbb\{R\}\}\\mu\(A\\mid l\(\\mathbf\{x\}\)=u\)\\,\\eta\(du\),\\qquad A\\subset\\mathbb\{R\}^\{n\}\\text\{ Borel\}\.
Thenν≪μ\\nu\\ll\\muand

d​νd​μ​\(𝐱\)=z⁡\(l⁡\(𝐱\)\)=d​ηd​L​\(l⁡\(𝐱\)\)\.\\frac\{d\\nu\}\{d\\mu\}\(\\mathbf\{x\}\)=z\(l\(\\mathbf\{x\}\)\)=\\frac\{d\\eta\}\{dL\}\(l\(\\mathbf\{x\}\)\)\.\(7\)

###### Proof:

Letf:ℝn→ℝf:\\mathbb\{R\}^\{n\}\\to\\mathbb\{R\}be bounded and measurable\. By definition ofν\\nu,

∫f⁡\(𝐱\)​ν​\(𝑑𝐱\)=∫ℝ\[∫f⁡\(𝐱\)​p​\(𝑑𝐱∣l⁡\(𝐱\)=u\)\]​η​\(𝑑u\)\.\\int f\(\\mathbf\{x\}\)\\,\\nu\(d\\mathbf\{x\}\)=\\int\_\{\\mathbb\{R\}\}\\left\[\\int f\(\\mathbf\{x\}\)\\,p\(d\\mathbf\{x\}\\mid l\(\\mathbf\{x\}\)=u\)\\right\]\\eta\(du\)\.Sinceη≪L\\eta\\ll Lwith densityz⁡\(u\)z\(u\), we can writeη⁡\(d​u\)=z⁡\(u\)​L​\(d​u\)\\eta\(du\)=z\(u\)\\,L\(du\)and obtain

∫f⁡\(𝐱\)​ν​\(𝑑𝐱\)=∫ℝz⁡\(u\)​∫f⁡\(𝐱\)​μ​\(𝑑𝐱∣l⁡\(𝐱\)=u\)​L​\(𝑑u\)\.\\int f\(\\mathbf\{x\}\)\\,\\nu\(d\\mathbf\{x\}\)=\\int\_\{\\mathbb\{R\}\}z\(u\)\\,\\int f\(\\mathbf\{x\}\)\\,\\mu\(d\\mathbf\{x\}\\mid l\(\\mathbf\{x\}\)=u\)\\,L\(du\)\.By the disintegration theorem\[[7](https://arxiv.org/html/2609.12113#bib.bib18)\],

∫ℝ∫f⁡\(𝐱\)​μ​\(𝑑𝐱∣l⁡\(𝐱\)=u\)​L​\(𝑑u\)=∫f⁡\(𝐱\)​μ​\(𝑑𝐱\)\.\\int\_\{\\mathbb\{R\}\}\\int f\(\\mathbf\{x\}\)\\,\\mu\(d\\mathbf\{x\}\\mid l\(\\mathbf\{x\}\)=u\)L\(du\)=\\int f\(\\mathbf\{x\}\)\\,\\mu\(d\\mathbf\{x\}\)\.Applying the same identity to the measurable functionx↦z⁡\(l⁡\(𝐱\)\)​f​\(𝐱\)x\\mapsto z\(l\(\\mathbf\{x\}\)\)f\(\\mathbf\{x\}\)gives

∫f⁡\(𝐱\)​z​\(l⁡\(𝐱\)\)​μ​\(𝑑𝐱\)=∫ℝz⁡\(u\)​∫f⁡\(𝐱\)​μ​\(𝑑𝐱∣l⁡\(𝐱\)=u\)​L​\(𝑑u\)\.\\int f\(\\mathbf\{x\}\)\\,z\(l\(\\mathbf\{x\}\)\)\\,\\mu\(d\\mathbf\{x\}\)=\\int\_\{\\mathbb\{R\}\}z\(u\)\\int f\(\\mathbf\{x\}\)\\,\\mu\(d\\mathbf\{x\}\\mid l\(\\mathbf\{x\}\)=u\)L\(du\)\.Comparing the two expressions, we conclude

∫f⁡\(𝐱\)​ν​\(𝑑𝐱\)=∫f⁡\(𝐱\)​z​\(l⁡\(𝐱\)\)​μ​\(𝑑𝐱\)\.\\int f\(\\mathbf\{x\}\)\\,\\nu\(d\\mathbf\{x\}\)=\\int f\(\\mathbf\{x\}\)\\,z\(l\(\\mathbf\{x\}\)\)\\,\\mu\(d\\mathbf\{x\}\)\.Since this holds for all bounded measurableff, it follows that

d​νd​μ​\(𝐱\)=z​\(l​\(𝐱\)\)\.\\frac\{d\\nu\}\{d\\mu\}\(\\mathbf\{x\}\)=z\(l\(\\mathbf\{x\}\)\)\.∎

###### Corollary 1

Letμt\\mu\_\{t\}be a family of probability measures onℝn\\mathbb\{R\}^\{n\}with densitiesptp\_\{t\}and log\-densitieslt=log⁡ptl\_\{t\}=\\log p\_\{t\}\. LetLt=\(lt\)\#​μtL\_\{t\}=\(l\_\{t\}\)\_\{\\\#\}\\mu\_\{t\}\. Fix a family of target measures\(ηt\)t∈\[0,T\]\(\\eta\_\{t\}\)\_\{t\\in\[0,T\]\}onℝ\\mathbb\{R\}such thatηt≪Lt\\eta\_\{t\}\\ll L\_\{t\}for eachtt, and define

zt​\(u\):=d​ηtd​Lt​\(u\)\.z\_\{t\}\(u\):=\\frac\{d\\eta\_\{t\}\}\{dL\_\{t\}\}\(u\)\.Defineνt\\nu\_\{t\}to be the likelihood\-reweighted measure ofμt\\mu\_\{t\}with targetηt\\eta\_\{t\}in the sense of Definition[3](https://arxiv.org/html/2609.12113#Thmdefinition3), and letqtq\_\{t\}denote its density\. Thenνt≪μt\\nu\_\{t\}\\ll\\mu\_\{t\}and

d​νtd​μt​\(𝐱t\)=zt​\(lt​\(𝐱t\)\)=d​ηtd​Lt​\(lt​\(𝐱t\)\)\.\\frac\{d\\nu\_\{t\}\}\{d\\mu\_\{t\}\}\(\\mathbf\{x\}\_\{t\}\)=z\_\{t\}\(l\_\{t\}\(\\mathbf\{x\}\_\{t\}\)\)=\\frac\{d\\eta\_\{t\}\}\{dL\_\{t\}\}\(l\_\{t\}\(\\mathbf\{x\}\_\{t\}\)\)\.\(8\)

The proof follows the same argument as Theorem[2](https://arxiv.org/html/2609.12113#Thmtheorem2)\.

###### Proposition 2

Let\(μt,pt,lt,zt,Lt,ηt,νt\)t∈\[0,T\]\(\\mu\_\{t\},p\_\{t\},l\_\{t\},z\_\{t\},L\_\{t\},\\eta\_\{t\},\\nu\_\{t\}\)\_\{t\\in\[0,T\]\}be defined as in Corollary[1](https://arxiv.org/html/2609.12113#Thmcorollary1)\. Assume further thatptp\_\{t\}is aC1C^\{1\}density onℝn\\mathbb\{R\}^\{n\}and thatztz\_\{t\}isC1C^\{1\}, and letqt​\(𝐱\)q\_\{t\}\(\\mathbf\{x\}\)be the density ofνt\\nu\_\{t\}with respect to the Lebesgue measure\. Then

∇logqt\(𝐱\)=\[1\+ct\(lt\(𝐱\)\)\]∇logpt\(𝐱\),\\nabla\\log q\_\{t\}\(\\mathbf\{x\}\)=\[1\+c\_\{t\}\(l\_\{t\}\(\\mathbf\{x\}\)\)\]\\nabla\\log p\_\{t\}\(\\mathbf\{x\}\),\(9\)where

ct​\(u\):=∂ulog⁡zt​\(u\)\.c\_\{t\}\(u\):=\\partial\_\{u\}\\log z\_\{t\}\(u\)\.\(10\)

The proof is a direct calculation using the chain rule\. The proposition shows that∇𝐱\[log⁡zt​\(lt​\(𝐱\)\)\]\\nabla\_\{\\mathbf\{x\}\}\[\\log z\_\{t\}\(l\_\{t\}\(\\mathbf\{x\}\)\)\]lies in the span of∇log⁡pt\\nabla\\log p\_\{t\}everywhere\. We are now ready to summarize the main theoretical results\.

###### Corollary 2

Let\(ηt\)t∈\[0,T\]\(\\eta\_\{t\}\)\_\{t\\in\[0,T\]\}be a family of target likelihood distributions satisfyingηt≪Lt\\eta\_\{t\}\\ll L\_\{t\}for alltt\. Then the corresponding likelihood\-reweighted densitiesqtq\_\{t\}satisfy the score relation

∇logqt\(𝐱\)=\[1\+∂u\(logd​ηtd​Lt\)\(logpt\(𝐱\)\)\]∇logpt\(𝐱\)\.\\nabla\\log q\_\{t\}\(\\mathbf\{x\}\)=\\left\[1\+\\partial\_\{u\}\\left\(\\log\\frac\{d\\eta\_\{t\}\}\{dL\_\{t\}\}\\right\)\(\\log p\_\{t\}\(\\mathbf\{x\}\)\)\\right\]\\nabla\\log p\_\{t\}\(\\mathbf\{x\}\)\.\(11\)Consequently, likelihood reweighting induces a multiplicative modification of the diffusion score function\. Furthermore, we choose a target likelihood distributionη0\\eta\_\{0\}that satisfies

η0≤s​tL0,W1\(L0,η0\)\>ρ,\\eta\_\{0\}\\leq\_\{st\}L\_\{0\},\\quad W\_\{1\}\(L\_\{0\},\\eta\_\{0\}\)\>\\rho,matching our specification ofρ\\rho\-outliers\.

Corollary[2](https://arxiv.org/html/2609.12113#Thmcorollary2)shows that likelihood control does not require learning a new score function\. Instead, the score of a likelihood\-reweighted density can be expressed using the original diffusion score and a scalar likelihood\-dependent correction, without learning an independent score function\. This provides a distribution\-steering perspective on outlier generation through likelihood reweighting\.

## IIIImplementation

Motivated by the score relation in Corollary[2](https://arxiv.org/html/2609.12113#Thmcorollary2), we introduce the controlled reverse\-time ODE ansatz

𝐱˙t=−𝐱t−\(1\+ct\(lt\(𝐱t\)\)\)∇logpt\(𝐱t\)⏟=:∇log⁡qt​\(𝐱t\),\\dot\{\\mathbf\{x\}\}\_\{t\}=\-\\mathbf\{x\}\_\{t\}\-\\underbrace\{\(1\+c\_\{t\}\(l\_\{t\}\(\\mathbf\{x\}\_\{t\}\)\)\)\\nabla\\log p\_\{t\}\(\\mathbf\{x\}\_\{t\}\)\}\_\{=:\\nabla\\log q\_\{t\}\(\\mathbf\{x\}\_\{t\}\)\},\(12\)wherect​\(u\)=dd​u​log⁡\(d​ηtd​Lt​\(u\)\)c\_\{t\}\(u\)=\\frac\{d\}\{du\}\\log\\left\(\\frac\{d\\eta\_\{t\}\}\{dL\_\{t\}\}\(u\)\\right\)is a coefficient that needs to be determined\. To approximatectc\_\{t\}, it is therefore necessary to understand how bothηt\\eta\_\{t\}andLtL\_\{t\}evolve over time under the diffusion dynamics\. In particular, the forward OU process governing the diffusion model induces a contraction of density perturbations toward the Gaussian equilibrium, which we analyze next\.

###### Theorem 3

Let𝐱t\\mathbf\{x\}\_\{t\}solve the OU SDE \([1](https://arxiv.org/html/2609.12113#S1.E1)\), where𝐱0∼p\\mathbf\{x\}\_\{0\}\\sim p\. Letμt\\mu\_\{t\}denote the time\-marginal law of𝐱t\\mathbf\{x\}\_\{t\}, and letptp\_\{t\}denote its density\. The invariant measure is the standard Gaussian measureγ=𝒩⁡\(0,I\)\\gamma=\\mathcal\{N\}\(0,I\)with densitypγp\_\{\\gamma\}\. Define the log\-densities

lt​\(𝐱\):=log⁡pt​\(𝐱\),l∗​\(𝐱\):=log⁡pγ​\(𝐱\)\.l\_\{t\}\(\\mathbf\{x\}\):=\\log p\_\{t\}\(\\mathbf\{x\}\),\\quad l^\{\*\}\(\\mathbf\{x\}\):=\\log p\_\{\\gamma\}\(\\mathbf\{x\}\)\.Assumeμt≪γ\\mu\_\{t\}\\ll\\gammaand denoteht​\(𝐱\)h\_\{t\}\(\\mathbf\{x\}\)as the RN derivativeht​\(𝐱\):=d​μtd​γ​\(𝐱\)h\_\{t\}\(\\mathbf\{x\}\):=\\frac\{d\\mu\_\{t\}\}\{d\\gamma\}\(\\mathbf\{x\}\)\. Assumeh0∈H1​\(γ\)h\_\{0\}\\in H^\{1\}\(\\gamma\)\. Then

‖ht−1‖L2​\(γ\)2≤e−2​t​‖h0−1‖L2​\(γ\)2\.\\\|h\_\{t\}\-1\\\|\_\{L^\{2\}\(\\gamma\)\}^\{2\}\\leq e^\{\-2t\}\\\|h\_\{0\}\-1\\\|^\{2\}\_\{L^\{2\}\(\\gamma\)\}\.Moreover,

‖∇ht‖L2​\(γ\)≤e−t​‖∇h0‖L2​\(γ\)\.\\\|\\nabla h\_\{t\}\\\|\_\{L^\{2\}\(\\gamma\)\}\\leq e^\{\-t\}\\\|\\nabla h\_\{0\}\\\|\_\{L^\{2\}\(\\gamma\)\}\.Finally, suppose there existsm\>0m\>0such thatht​\(𝐱\)≥mh\_\{t\}\(\\mathbf\{x\}\)\\geq mforμt\\mu\_\{t\}\-almost all𝐱\\mathbf\{x\}andtt\. Then

‖ℓt−ℓ∗‖H1​\(γ\)=‖log⁡ht‖H1​\(γ\)≤e−tm​‖h0−1‖H1​\(γ\)\.\\\|\\ell\_\{t\}\-\\ell^\{\*\}\\\|\_\{H^\{1\}\(\\gamma\)\}=\\\|\\log h\_\{t\}\\\|\_\{H^\{1\}\(\\gamma\)\}\\leq\\frac\{e^\{\-t\}\}\{m\}\\\|h\_\{0\}\-1\\\|\_\{H^\{1\}\(\\gamma\)\}\.

###### Proof:

The densityptp\_\{t\}satisfies the FPK equation

∂tpt=∇⋅\(𝐱​pt\)\+Δ​pt\.\\partial\_\{t\}p\_\{t\}=\\nabla\\cdot\(\\mathbf\{x\}p\_\{t\}\)\+\\Delta p\_\{t\}\.
Define the density ratioht​\(𝐱\):=pt​\(𝐱\)γ⁡\(𝐱\)\.h\_\{t\}\(\\mathbf\{x\}\):=\\frac\{p\_\{t\}\(\\mathbf\{x\}\)\}\{\\gamma\(\\mathbf\{x\}\)\}\.Using the identities

∇pγ=−𝐱​pγ,Δ​pγ=\(‖𝐱‖2−n\)​pγ,\\nabla p\_\{\\gamma\}=\-\\mathbf\{x\}p\_\{\\gamma\},\\quad\\Delta p\_\{\\gamma\}=\(\\\|\\mathbf\{x\}\\\|^\{2\}\-n\)p\_\{\\gamma\},one verifies thathth\_\{t\}satisfies

∂tht=ℒht,ℒ:=Δ−𝐱⋅∇,\\partial\_\{t\}h\_\{t\}=\\mathcal\{L\}h\_\{t\},\\quad\\mathcal\{L\}:=\\Delta\-\\mathbf\{x\}\\cdot\\nabla,whereℒ\\mathcal\{L\}is the OU generator introduced in Definition[1](https://arxiv.org/html/2609.12113#Thmdefinition1)\. Letgt:=ht−1\.g\_\{t\}:=h\_\{t\}\-1\.Since∫ht​𝑑γ=1\\int h\_\{t\}\\,d\\gamma=1, we have∫gt​𝑑γ=0\\int g\_\{t\}\\,d\\gamma=0\. The evolution equation becomes

∂tgt=ℒ​gt,gt\|t=0=g0\.\\partial\_\{t\}g\_\{t\}=\\mathcal\{L\}g\_\{t\},\\quad g\_\{t\}\|\_\{t=0\}=g\_\{0\}\.\(13\)By Theorem[1](https://arxiv.org/html/2609.12113#Thmtheorem1), the operatorℒ\\mathcal\{L\}is self\-adjoint onL2​\(γ\)L^\{2\}\(\\gamma\)with eigenfunctions given by Hermite polynomials\[[15](https://arxiv.org/html/2609.12113#bib.bib19)\],ℒ​Hα=−\|α\|​Hα\.\\mathcal\{L\}H\_\{\\alpha\}=\-\|\\alpha\|H\_\{\\alpha\}\.Expandingg0g\_\{0\}in the Hermite basis gives

g0=∑α≠0gα​Hα,gα=⟨g0,Hα⟩L2​\(γ\)\.g\_\{0\}=\\sum\_\{\\alpha\\neq 0\}g\_\{\\alpha\}H\_\{\\alpha\},\\quad g\_\{\\alpha\}=\\langle g\_\{0\},H\_\{\\alpha\}\\rangle\_\{L^\{2\}\(\\gamma\)\}\.Let\(Pt\)t≥0\(P\_\{t\}\)\_\{t\\geq 0\}be the OU semigroup onL2​\(γ\)L^\{2\}\(\\gamma\)generated byℒ\\mathcal\{L\}\. Sincegtg\_\{t\}satisfies Equation \([13](https://arxiv.org/html/2609.12113#S3.E13)\), we identifygtg\_\{t\}with the unique semigroup solutiongt=Pt​g0\.g\_\{t\}=P\_\{t\}g\_\{0\}\.SincePtP\_\{t\}acts diagonally on the Hermite basis, we can expand this solution

gt=∑α≠0e−\|α\|​t​gα​Hα\.g\_\{t\}=\\sum\_\{\\alpha\\neq 0\}e^\{\-\|\\alpha\|t\}g\_\{\\alpha\}H\_\{\\alpha\}\.TakingL2​\(γ\)L^\{2\}\(\\gamma\)norms yields

‖gt‖L2​\(γ\)2=∑k=1∞e−2​k​t​∑\|α\|=kgα2≤e−2​t​‖g0‖L2​\(γ\)2,\\\|g\_\{t\}\\\|\_\{L^\{2\}\(\\gamma\)\}^\{2\}=\\sum\_\{k=1\}^\{\\infty\}e^\{\-2kt\}\\sum\_\{\|\\alpha\|=k\}g\_\{\\alpha\}^\{2\}\\leq e^\{\-2t\}\\\|g\_\{0\}\\\|^\{2\}\_\{L^\{2\}\(\\gamma\)\},where the inequality follows immediately sincee−k​t≤e−te^\{\-kt\}\\leq e^\{\-t\}fork≥1k\\geq 1\. By Lemma 1 of\[[5](https://arxiv.org/html/2609.12113#bib.bib17)\], we can bound the gradient:

‖∇gt‖L2​\(γ\)2≤e−2​t​‖∇g0‖L2​\(γ\)2\.\\\|\\nabla g\_\{t\}\\\|^\{2\}\_\{L^\{2\}\(\\gamma\)\}\\leq e^\{\-2t\}\\\|\\nabla g\_\{0\}\\\|^\{2\}\_\{L^\{2\}\(\\gamma\)\}\.Becausegt=ht−1g\_\{t\}=h\_\{t\}\-1, this is equivalent to

‖∇ht‖L2​\(γ\)≤e−t​‖∇h0‖L2​\(γ\)\.\\\|\\nabla h\_\{t\}\\\|\_\{L^\{2\}\(\\gamma\)\}\\leq e^\{\-t\}\\\|\\nabla h\_\{0\}\\\|\_\{L^\{2\}\(\\gamma\)\}\.
Now assumeht≥m\>0h\_\{t\}\\geq m\>0\. Sincelog\\logis1/m1/m\-Lipschitz on\[m,∞\)\[m,\\infty\),

\|log⁡ht​\(𝐱\)\|≤1m​\|ht​\(𝐱\)−1\|\.\|\\log h\_\{t\}\(\\mathbf\{x\}\)\|\\leq\\frac\{1\}\{m\}\|h\_\{t\}\(\\mathbf\{x\}\)\-1\|\.Also, by expanding∇log⁡ht​\(𝐱\)=∇ht​\(𝐱\)ht​\(𝐱\),\\nabla\\log h\_\{t\}\(\\mathbf\{x\}\)=\\frac\{\\nabla h\_\{t\}\(\\mathbf\{x\}\)\}\{h\_\{t\}\(\\mathbf\{x\}\)\},we have

‖∇log⁡ht​\(𝐱\)‖≤1m​‖∇ht​\(𝐱\)‖\.\\\|\\nabla\\log h\_\{t\}\(\\mathbf\{x\}\)\\\|\\leq\\frac\{1\}\{m\}\\\|\\nabla h\_\{t\}\(\\mathbf\{x\}\)\\\|\.Combining the two inequalities,

‖log⁡ht‖H1​\(γ\)≤1m​‖ht−1‖H1​\(γ\)\.\\\|\\log h\_\{t\}\\\|\_\{H^\{1\}\(\\gamma\)\}\\leq\\frac\{1\}\{m\}\\\|h\_\{t\}\-1\\\|\_\{H^\{1\}\(\\gamma\)\}\.Finally, using theL2L^\{2\}and gradient estimates above,

‖ht−1‖H1​\(γ\)≤e−t​‖h0−1‖H1​\(γ\)\.\\\|h\_\{t\}\-1\\\|\_\{H^\{1\}\(\\gamma\)\}\\leq e^\{\-t\}\\\|h\_\{0\}\-1\\\|\_\{H^\{1\}\(\\gamma\)\}\.Since

log⁡ht=ℓt−ℓ∗,\\log h\_\{t\}=\\ell\_\{t\}\-\\ell^\{\*\},the conclusion follows\. ∎

Theorem[3](https://arxiv.org/html/2609.12113#Thmtheorem3)shows thatLt=\(lt\)\#​μtL\_\{t\}=\(l\_\{t\}\)\_\{\\\#\}\\mu\_\{t\}converges exponentially towards the Gaussian likelihood equilibrium\(l∗\)\#​γ\(l^\{\*\}\)\_\{\\\#\}\\gamma\. Motivated by this, we choose the target family\{ηt\}\\\{\\eta\_\{t\}\\\}that likewise vanishes toward the same equilibrium\. In particular, the slowest nonconstant OU decay ratee−te^\{\-t\}motivates the exponentially decaying controller introduced below\.

### III\-AApproximating a control function

In this section, we claim that, provided the log\-likelihood pushforward measureL0L\_\{0\}and a targetη0\\eta\_\{0\}, the exponentially interpolated control function

c~t​\(u\)=e−t​c0​\(u\)=e−t​∂u\(log⁡d​η0d​L0​\(u\)\),\\tilde\{c\}\_\{t\}\(u\)=e^\{\-t\}c\_\{0\}\(u\)=e^\{\-t\}\\partial\_\{u\}\\left\(\\log\\frac\{d\\eta\_\{0\}\}\{dL\_\{0\}\}\(u\)\\right\),\(14\)provides a natural approximation choice for the controlled reverse dynamics\. The true time‑dependent controlct​\(u\)c\_\{t\}\(u\)is computationally intractable\. Motivated by the slowest decaying nonconstant term,e−te^\{\-t\}, of the OU semigroup, we choose the exponentially interpolated controller\. We prove that this approximation satisfies a reasonable error bound over both small and largett, thus providing a natural admissible approximation that respects the boundary conditions and admits explicit error bounds near both endpoints\.

###### Theorem 4

Let\(𝐱t\)t∈\[0,T\]\(\\mathbf\{x\}\_\{t\}\)\_\{t\\in\[0,T\]\}be an OU diffusion \([1](https://arxiv.org/html/2609.12113#S1.E1)\), and letLtL\_\{t\}andηt\\eta\_\{t\}be two families of probability measures onℝ\\mathds\{R\}which are absolutely continuous with respect to Lebesgue measure, converge exponentially to the Gaussian likelihood equilibrium, and satisfy:

η0≪L0,η0≤s​tL0,W1\(η0,L0\)≥ρ\.\\eta\_\{0\}\\ll L\_\{0\},\\quad\\eta\_\{0\}\\leq\_\{st\}L\_\{0\},\\quad W\_\{1\}\(\\eta\_\{0\},L\_\{0\}\)\\geq\\rho\.Define the likelihood ratio

zt​\(u\):=d​ηtd​Lt​\(u\),ct​\(u\):=∂ulog⁡zt​\(u\)\.z\_\{t\}\(u\):=\\frac\{d\\eta\_\{t\}\}\{dL\_\{t\}\}\(u\),\\quad c\_\{t\}\(u\):=\\partial\_\{u\}\\log z\_\{t\}\(u\)\.Assume:

1. \(A1\)There existsL\>0L\>0such that for alls,t∈\[0,T\]s,t\\in\[0,T\], ‖ct−cs‖L2​\(Lt\)≤L​\|t−s\|\.\\\|c\_\{t\}\-c\_\{s\}\\\|\_\{L^\{2\}\(L\_\{t\}\)\}\\leq L\|t\-s\|\.
2. \(A2\)There existsK≥1K\\geq 1such that for allt∈\[0,T\]t\\in\[0,T\]and all measurableff, K−1​‖f‖L2​\(L0\)≤‖f‖L2​\(Lt\)≤K​‖f‖L2​\(L0\)\.K^\{\-1\}\\\|f\\\|\_\{L^\{2\}\(L\_\{0\}\)\}\\leq\\\|f\\\|\_\{L^\{2\}\(L\_\{t\}\)\}\\leq K\\\|f\\\|\_\{L^\{2\}\(L\_\{0\}\)\}\.

Define the interpolantc~t​\(u\):=e−t​c0​\(u\)\.\\tilde\{c\}\_\{t\}\(u\):=e^\{\-t\}c\_\{0\}\(u\)\.Thenc~0=c0\\tilde\{c\}\_\{0\}=c\_\{0\}, and for everyt∈\[0,T\]t\\in\[0,T\],

‖c~t−ct‖L2​\(Lt\)≤K⁡\(\(1−e−t\)​‖c0‖L2​\(L0\)\+L​t\)\.\\\|\\tilde\{c\}\_\{t\}\-c\_\{t\}\\\|\_\{L^\{2\}\(L\_\{t\}\)\}\\leq K\\Big\(\(1\-e^\{\-t\}\)\\\|c\_\{0\}\\\|\_\{L^\{2\}\(L\_\{0\}\)\}\+Lt\\Big\)\.In particular,‖c~t−ct‖L2​\(Lt\)=O⁡\(t\),\\\|\\tilde\{c\}\_\{t\}\-c\_\{t\}\\\|\_\{L^\{2\}\(L\_\{t\}\)\}=O\(t\),

###### Proof:

Fixt∈\[0,T\]t\\in\[0,T\]\. Add and subtractc0c\_\{0\}:

c~t−ct=\(e−t​c0−c0\)\+\(c0−ct\)\.\\tilde\{c\}\_\{t\}\-c\_\{t\}=\(e^\{\-t\}c\_\{0\}\-c\_\{0\}\)\+\(c\_\{0\}\-c\_\{t\}\)\.TakeL2​\(Lt\)L^\{2\}\(L\_\{t\}\)norms and apply the triangle inequality:

‖c~t−ct‖L2​\(Lt\)≤‖\(1−e−t\)​c0‖L2​\(Lt\)\+‖ct−c0‖L2​\(Lt\)\.\\\|\\tilde\{c\}\_\{t\}\-c\_\{t\}\\\|\_\{L^\{2\}\(L\_\{t\}\)\}\\leq\\\|\(1\-e^\{\-t\}\)c\_\{0\}\\\|\_\{L^\{2\}\(L\_\{t\}\)\}\+\\\|c\_\{t\}\-c\_\{0\}\\\|\_\{L^\{2\}\(L\_\{t\}\)\}\.By \(A2\),‖\(1−e−t\)​c0‖L2​\(Lt\)≤K⁡\(1−e−t\)​‖c0‖L2​\(L0\)\.\\displaystyle\\quad\\\|\(1\-e^\{\-t\}\)c\_\{0\}\\\|\_\{L^\{2\}\(L\_\{t\}\)\}\\leq K\(1\-e^\{\-t\}\)\\\|c\_\{0\}\\\|\_\{L^\{2\}\(L\_\{0\}\)\}\.By \(A1\),‖ct−c0‖L2​\(Lt\)≤L​t\.\\displaystyle\\quad\\\|c\_\{t\}\-c\_\{0\}\\\|\_\{L^\{2\}\(L\_\{t\}\)\}\\leq Lt\.Combine the two bounds to obtain

‖c~t−ct‖L2​\(Lt\)≤K⁡\(\(1−e−t\)​‖c0‖L2​\(L0\)\+L​t\)\.\\\|\\tilde\{c\}\_\{t\}\-c\_\{t\}\\\|\_\{L^\{2\}\(L\_\{t\}\)\}\\leq K\\Big\(\(1\-e^\{\-t\}\)\\\|c\_\{0\}\\\|\_\{L^\{2\}\(L\_\{0\}\)\}\+Lt\\Big\)\.Finally, since1−e−t≤t1\-e^\{\-t\}\\leq t, the right\-hand side isO⁡\(t\)O\(t\)\. ∎

###### Theorem 5

Let\(𝐱t\)t∈\[0,T\],Lt,ηt,γ,zt​\(u\),ct​\(u\)\(\\mathbf\{x\}\_\{t\}\)\_\{t\\in\[0,T\]\},L\_\{t\},\\eta\_\{t\},\\gamma,z\_\{t\}\(u\),c\_\{t\}\(u\)be defined as in Theorem[4](https://arxiv.org/html/2609.12113#Thmtheorem4), and letlt,μt,νtl\_\{t\},\\mu\_\{t\},\\nu\_\{t\}be defined as in Corollary[1](https://arxiv.org/html/2609.12113#Thmcorollary1)\. Moreover, assumec0∈L∞c\_\{0\}\\in L^\{\\infty\}and suppose the assumptions from Theorem[3](https://arxiv.org/html/2609.12113#Thmtheorem3)are satisfied forμt\\mu\_\{t\}andνt\\nu\_\{t\}\. Then the exponentially\-interpolated controllerc~t​\(u\):=e−t​c0​\(u\)\\tilde\{c\}\_\{t\}\(u\):=e^\{\-t\}c\_\{0\}\(u\)satisfies the boundary conditionc~0=c0\\tilde\{c\}\_\{0\}=c\_\{0\}, and moreover

∥\(c~t−ct\)\(lt\(⋅\)\)∇logpt\(⋅\)∥L2​\(γ\)≤Ce−t,t∈\[0,T\],\\\|\(\\tilde\{c\}\_\{t\}\-c\_\{t\}\)\(l\_\{t\}\(\\cdot\)\)\\,\\nabla\\log p\_\{t\}\(\\cdot\)\\\|\_\{L^\{2\}\(\\gamma\)\}\\leq Ce^\{\-t\},\\quad t\\in\[0,T\],for a constantC\>0C\>0\. In particular,

\(c~t−ct\)\(lt\(x\)\)∇logpt\(x\)\(\\tilde\{c\}\_\{t\}\-c\_\{t\}\)\(l\_\{t\}\(x\)\)\\,\\nabla\\log p\_\{t\}\(x\)converges exponentially fast to zero inL2​\(γ\)L^\{2\}\(\\gamma\)\.

###### Proof:

log⁡zt\\log z\_\{t\}can be expanded as

log⁡zt​\(lt​\(𝐱\)\)=log⁡d​νtd​μt​\(𝐱\)=log⁡d​νtd​γ​\(𝐱\)−log⁡d​μtd​γ​\(𝐱\)\.\\log z\_\{t\}\(l\_\{t\}\(\\mathbf\{x\}\)\)=\\log\\frac\{d\\nu\_\{t\}\}\{d\\mu\_\{t\}\}\(\\mathbf\{x\}\)=\\log\\frac\{d\\nu\_\{t\}\}\{d\\gamma\}\(\\mathbf\{x\}\)\-\\log\\frac\{d\\mu\_\{t\}\}\{d\\gamma\}\(\\mathbf\{x\}\)\.From Theorem[3](https://arxiv.org/html/2609.12113#Thmtheorem3)and the triangle inequality, we have

‖log⁡zt​\(lt​\(⋅\)\)‖H1​\(γ\)\\displaystyle\\\|\\log z\_\{t\}\(l\_\{t\}\(\\cdot\)\)\\\|\_\{H^\{1\}\(\\gamma\)\}≤\(‖log⁡d​νtd​γ‖H1​\(γ\)\+‖log⁡d​μtd​γ‖H1​\(γ\)\)\\displaystyle\\leq\\left\(\\\|\\log\\frac\{d\\nu\_\{t\}\}\{d\\gamma\}\\\|\_\{H^\{1\}\(\\gamma\)\}\+\\\|\\log\\frac\{d\\mu\_\{t\}\}\{d\\gamma\}\\\|\_\{H^\{1\}\(\\gamma\)\}\\right\)≤e−tm​\(‖d​ν0d​γ−1‖H1​\(γ\)\+‖d​μ0d​γ−1‖H1​\(γ\)\)\.\\displaystyle\\leq\\frac\{e^\{\-t\}\}\{m\}\\left\(\\\|\\frac\{d\\nu\_\{0\}\}\{d\\gamma\}\-1\\\|\_\{H^\{1\}\(\\gamma\)\}\+\\\|\\frac\{d\\mu\_\{0\}\}\{d\\gamma\}\-1\\\|\_\{H^\{1\}\(\\gamma\)\}\\right\)\.By definition of theH1​\(γ\)H^\{1\}\(\\gamma\)norm,

‖∇log⁡zt‖L2​\(γ\)\\displaystyle\\\|\\nabla\\log z\_\{t\}\\\|\_\{L^\{2\}\(\\gamma\)\}≤e−tm​\(‖d​ν0d​γ−1‖H1​\(γ\)\+‖d​μ0d​γ−1‖H1​\(γ\)\)\.\\displaystyle\\leq\\frac\{e^\{\-t\}\}\{m\}\\left\(\\\|\\frac\{d\\nu\_\{0\}\}\{d\\gamma\}\-1\\\|\_\{H^\{1\}\(\\gamma\)\}\+\\\|\\frac\{d\\mu\_\{0\}\}\{d\\gamma\}\-1\\\|\_\{H^\{1\}\(\\gamma\)\}\\right\)\.Since∇log⁡zt​\(lt​\(𝐱t\)\)\\nabla\\log z\_\{t\}\(l\_\{t\}\(\\mathbf\{x\}\_\{t\}\)\)can be expanded as∂ulogzt∇lt\(𝐱t\)\\partial\_\{u\}\\log z\_\{t\}\\nabla l\_\{t\}\(\\mathbf\{x\}\_\{t\}\),

∥∇logzt\(lt\(⋅\)\)∥L2​\(γ\)=∥∂ulogzt∇lt\(𝐱t\)∥L2​\(γ\),\\\|\\nabla\\log z\_\{t\}\(l\_\{t\}\(\\cdot\)\)\\\|\_\{L^\{2\}\(\\gamma\)\}=\\\|\\partial\_\{u\}\\log z\_\{t\}\\nabla l\_\{t\}\(\\mathbf\{x\}\_\{t\}\)\\\|\_\{L^\{2\}\(\\gamma\)\},which decays exponentially\. Finally, asct=∂ulog⁡ztc\_\{t\}=\\partial\_\{u\}\\log z\_\{t\}andc~t=e−t​c0\\tilde\{c\}\_\{t\}=e^\{\-t\}c\_\{0\}, we can write

c~t−ct=e−t​∂ulog⁡z0−∂ulog⁡zt\.\\tilde\{c\}\_\{t\}\-c\_\{t\}=e^\{\-t\}\\partial\_\{u\}\\log z\_\{0\}\-\\partial\_\{u\}\\log z\_\{t\}\.Using the triangle inequality,

∥\(c~t−ct\)\(lt\(⋅\)\)∇lt∥L2​\(γ\)≤\\displaystyle\\\|\(\\tilde\{c\}\_\{t\}\-c\_\{t\}\)\(l\_\{t\}\(\\cdot\)\)\\nabla l\_\{t\}\\\|\_\{L^\{2\}\(\\gamma\)\}\\leqe−t​‖c0‖L∞​‖∇lt‖L2​\(γ\)\\displaystyle\\penalty\\ e^\{\-t\}\\\|c\_\{0\}\\\|\_\{L^\{\\infty\}\}\\\|\\nabla l\_\{t\}\\\|\_\{L^\{2\}\(\\gamma\)\}\+‖∇log⁡zt​\(lt\)‖L2​\(γ\)\.\\displaystyle\+\\\|\\nabla\\log z\_\{t\}\(l\_\{t\}\)\\\|\_\{L^\{2\}\(\\gamma\)\}\.Since Theorem[3](https://arxiv.org/html/2609.12113#Thmtheorem3)yields

‖∇lt−∇log⁡γ‖L2​\(γ\)=‖∇lt\+x‖L2​\(γ\)≤C1​e−t,\\\|\\nabla l\_\{t\}\-\\nabla\\log\\gamma\\\|\_\{L^\{2\}\(\\gamma\)\}=\\\|\\nabla l\_\{t\}\+x\\\|\_\{L^\{2\}\(\\gamma\)\}\\leq C\_\{1\}e^\{\-t\},we obtain that‖∇lt‖L2​\(γ\)\\\|\\nabla l\_\{t\}\\\|\_\{L^\{2\}\(\\gamma\)\}is uniformly bounded overtt\. Thus,

∥\(c~t\\displaystyle\\\|\(\\tilde\{c\}\_\{t\}−ct\)\(lt\(⋅\)\)∇logpt\(⋅\)∥L2​\(γ\)≤Ce−t,\\displaystyle\-c\_\{t\}\)\(l\_\{t\}\(\\cdot\)\)\\,\\nabla\\log p\_\{t\}\(\\cdot\)\\\|\_\{L^\{2\}\(\\gamma\)\}\\leq Ce^\{\-t\},where​C=\\displaystyle\\text\{where \}C=‖c0‖L∞​\(γ\)​supt∈\[0,T\]‖∇lt‖L2​\(γ\)\\displaystyle\\penalty\\ \\\|c\_\{0\}\\\|\_\{L^\{\\infty\}\(\\gamma\)\}\\sup\_\{t\\in\[0,T\]\}\\\|\\nabla l\_\{t\}\\\|\_\{L^\{2\}\(\\gamma\)\}\+1m​\(‖d​ν0d​γ−1‖H1​\(γ\)\+‖d​μ0d​γ−1‖H1​\(γ\)\)\\displaystyle\+\\frac\{1\}\{m\}\(\\\|\\frac\{d\\nu\_\{0\}\}\{d\\gamma\}\-1\\\|\_\{H^\{1\}\(\\gamma\)\}\+\\\|\\frac\{d\\mu\_\{0\}\}\{d\\gamma\}\-1\\\|\_\{H^\{1\}\(\\gamma\)\}\)and therefore the right\-hand side isO⁡\(e−t\)O\(e^\{\-t\}\)\.∎

###### Corollary 3

Under the assumptions of Theorems[4](https://arxiv.org/html/2609.12113#Thmtheorem4)and[5](https://arxiv.org/html/2609.12113#Thmtheorem5), the error between the true control termct\(u\)∇ltc\_\{t\}\(u\)\\nabla l\_\{t\}and the estimatec~t\(u\)∇lt=e−tc0\(u\)∇lt\\tilde\{c\}\_\{t\}\(u\)\\nabla l\_\{t\}=e^\{\-t\}c\_\{0\}\(u\)\\nabla l\_\{t\}are bounded for small and largett:

∥\(c~t−ct\)\(lt\(⋅\)\)∇lt∥L2​\(γ\)≤O\(min\(t,e−t\)\)\.\\\|\(\\tilde\{c\}\_\{t\}\-c\_\{t\}\)\(l\_\{t\}\(\\cdot\)\)\\nabla l\_\{t\}\\\|\_\{L^\{2\}\(\\gamma\)\}\\leq O\(\\min\(t,e^\{\-t\}\)\)\.\(15\)

### III\-BApproximating the target Radon\-Nikodym derivative

We will rely on chi\-square approximations forL0L\_\{0\}in this paper\. In our numerical experiments, these approximations will be verified and demonstrated\. We introduce the notation of a linear mapLa,CL\_\{a,C\}defined asLa,C​\(x\)=−a​x\+CL\_\{a,C\}\(x\)=\-ax\+C\. Under this approximation, the control coefficients admit closed\-form expressions, making the proposed controller straightforward to compute in practice\.

###### Assumption 1

LetL0L\_\{0\}be a likelihood\-pushforward measure onℝ\\mathbb\{R\}with finite mean\. Assume there existsa\>0a\>0andCCsuch that\(La,C\)\#\(χn2\)≤s​tL0\(L\_\{a,C\}\)\_\{\\\#\}\(\\chi\_\{n\}^\{2\}\)\\leq\_\{st\}L\_\{0\}\.

###### Fact 1

Leta\>0a\>0and consider the measurePa:=\(La,C\)\#​\(χn2\)P\_\{a\}:=\(L\_\{a,C\}\)\_\{\\\#\}\(\\chi\_\{n\}^\{2\}\)onℝ\\mathbb\{R\}\. Then for anyρ\>0\\rho\>0, we can define a measurePb:=\(Lb,C\)\#​\(χn2\)P\_\{b\}:=\(L\_\{b,C\}\)\_\{\\\#\}\(\\chi\_\{n\}^\{2\}\), whereb=a\+ρnb=a\+\\frac\{\\rho\}\{n\}\. This yields

Pb≤s​tPa,W1\(Pa,Pb\)=ρ\.\\displaystyle P\_\{b\}\\leq\_\{st\}P\_\{a\},\\quad W\_\{1\}\(P\_\{a\},P\_\{b\}\)=\\rho\.

###### Fact 2

LetPa=\(La,C\)\#​\(χn2\)P\_\{a\}=\(L\_\{a,C\}\)\_\{\\\#\}\(\\chi\_\{n\}^\{2\}\)andPb=\(Lb,C\)\#​\(χn2\)P\_\{b\}=\(L\_\{b,C\}\)\_\{\\\#\}\(\\chi\_\{n\}^\{2\}\), witha,b\>0a,b\>0andC∈ℝC\\in\\mathbb\{R\}\. Their RN derivative is

d​Pbd​Pa​\(u\)=\(ab\)n/2​exp⁡\[\(C−u\)​\(12​a−12​b\)\],u<C\.\\frac\{dP\_\{b\}\}\{dP\_\{a\}\}\(u\)=\\left\(\\frac\{a\}\{b\}\\right\)^\{n/2\}\\exp\\left\[\(C\-u\)\\left\(\\frac\{1\}\{2a\}\-\\frac\{1\}\{2b\}\\right\)\\right\],\\quad u<C\.Correspondingly, we obtain the constant control

c0:=dd​u​log⁡d​Pbd​Pa​\(u\)=−12​a\+12​b\.c\_\{0\}:=\\frac\{d\}\{du\}\\log\\frac\{dP\_\{b\}\}\{dP\_\{a\}\}\(u\)=\-\\frac\{1\}\{2a\}\+\\frac\{1\}\{2b\}\.

## IVNumerical Experiments

### IV\-AMixture of Two Gaussians

We first evaluate our method on a synthetic testbed consisting of a mixture of two Gaussians

𝐱∼12​𝒩​\(𝐦,I\)\+12​𝒩​\(−𝐦,I\)\.\\mathbf\{x\}\\sim\\tfrac\{1\}\{2\}\\mathcal\{N\}\(\\mathbf\{m\},I\)\+\\tfrac\{1\}\{2\}\\mathcal\{N\}\(\-\\mathbf\{m\},I\)\.\(16\)For this model the score function admits the closed\-form expression\[[24](https://arxiv.org/html/2609.12113#bib.bib10)\]

∇log⁡pt​\(𝐱\)=tanh⁡\(𝐦t⋅𝐱\)​𝐦t−𝐱,𝐦t=e−t​𝐦\.\\nabla\\log p\_\{t\}\(\\mathbf\{x\}\)=\\tanh\(\\mathbf\{m\}\_\{t\}\\cdot\\mathbf\{x\}\)\\mathbf\{m\}\_\{t\}\-\\mathbf\{x\},\\quad\\mathbf\{m\}\_\{t\}=e^\{\-t\}\\mathbf\{m\}\.\(17\)
We choose𝐦=m^​1n\\mathbf\{m\}=\\hat\{m\}1\_\{n\}, where1n1\_\{n\}denotes thenn\-dimensional vector of ones\. This closed\-form score allows us to solve the PF\-ODE \([3](https://arxiv.org/html/2609.12113#S1.E3)\) and the log\-FPK equation \([5](https://arxiv.org/html/2609.12113#S1.E5)\) directly without training a neural network\. Our goal is to guide the reverse diffusion process to generateρ\\rho\-outliers\.

#### Assessment of Assumption[1](https://arxiv.org/html/2609.12113#Thmassumption1)\.

We empirically assess the approximation ofL0L\_\{0\}by a translated−12​χn2\-\\tfrac\{1\}\{2\}\\chi\_\{n\}^\{2\}distribution\. FOSD is assessed by sorting samples from two likelihood distributions and verifying that the respective inequality holds element\-wise\. To estimateL0L\_\{0\}, we sample 10,000 points by solving the reverse PF\-ODE \([3](https://arxiv.org/html/2609.12113#S1.E3)\), then use the log\-FPK equation \([5](https://arxiv.org/html/2609.12113#S1.E5)\) to compute their log\-likelihood\. Table[I](https://arxiv.org/html/2609.12113#S4.T1)reports their empiricalW1W\_\{1\}distance across several dimensions and values ofm^\\hat\{m\}\. The results indicate that the chi\-square model provides a good approximation of the log\-likelihood distribution\. An example empirical CDF is shown in Figure[2](https://arxiv.org/html/2609.12113#S4.F2)\.

Table I:Empirical assessment of the chi\-square approximation witha=1/2a=1/2\. Values report the empiricalW1W\_\{1\}distance betweenL0L\_\{0\}and the fitted chi\-square distribution\.
#### Outlier generation\.

For the remaining experiments we setm^=2\\hat\{m\}=2and diffusion horizonT=80T=80\. We vary the dimensionn∈\{1,4,16,64,256\}n\\in\\\{1,4,16,64,256\\\}and targetρ∈\{0\.2,0\.3,0\.4,0\.5\}\\rho\\in\\\{0\.2,0\.3,0\.4,0\.5\\\}\. In each experiment, we verify that FOSD is empirically satisfied, and report theW1W\_\{1\}distance between the generated likelihood distributionη0\\eta\_\{0\}and the referenceL0L\_\{0\}\.

Our results are posted in Table[II](https://arxiv.org/html/2609.12113#S4.T2)\. The generated samples match the targetρ\\rhovalues while maintaining stochastic dominance\. Figure[2](https://arxiv.org/html/2609.12113#S4.F2)compares histograms of samples generated by the uncontrolled and controlled PF\-ODE in the one\-dimensional case\. The controlled dynamics produce samples concentrated in lower\-likelihood regions such as the tails and the low\-density region between the mixture modes\.

Table II:Outlier generation results for the Gaussian mixture model\. Values report the empiricalW1W\_\{1\}distances between the generated and reference likelihood distributions\. The measured distances closely match the prescribed target valuesρ\\rho, demonstrating accurate control over outlier magnitude while satisfying stochastic dominance in every experiment\.![Refer to caption](https://arxiv.org/html/2609.12113v1/figs_png/density_reverse_controller_sde_sample.png)

Figure 1:Histogram comparison of PF\-ODE samples in the 1D Gaussian mixture example withρ=0\.5\\rho=0\.5\.![Refer to caption](https://arxiv.org/html/2609.12113v1/figs_png/logp_density_controlled_reverse_sde_sample.png)

Figure 2:Empirical CDFs of the likelihood distributionL0L\_\{0\}and the target distributionη0\\eta\_\{0\}in the 4D example withρ=0\.5\\rho=0\.5\.

### IV\-BImage Data: CIFAR\-10

We next evaluate the method on the CIFAR\-10 dataset using the pretrained diffusion modelgoogle/ddpm\-cifar10\-32\[[17](https://arxiv.org/html/2609.12113#bib.bib1)\]\. The likelihood distributionL0L\_\{0\}is approximated using samples generated by the reverse PF\-ODE together with the log\-FPK equation \([5](https://arxiv.org/html/2609.12113#S1.E5)\)\.

Empirically,L0L\_\{0\}stochastically dominates a translated−12​χn2\-\\tfrac\{1\}\{2\}\\chi\_\{n\}^\{2\}distribution, supporting Assumption[1](https://arxiv.org/html/2609.12113#Thmassumption1)\. Their empiricalW1W\_\{1\}distance is9\.1379\.137\. We vary the targetρ\\rhoacross\{50,100,150,200,250\}\\\{50,100,150,200,250\\\}and generateρ\\rho\-outliers by modifying the reverse diffusion dynamics\. Table[III](https://arxiv.org/html/2609.12113#S4.T3)reports the empiricalW1W\_\{1\}distances between the generated likelihood distribution and the reference distribution\. Although the realizedW1W\_\{1\}shifts do not exactly match the prescribedρ\\rhovalues, they increase monotonically withρ\\rho\. Thus, even when the score is represented by a neural network, the control parameter provides a consistent mechanism for adjusting the degree of likelihood shift\.

Figure[3](https://arxiv.org/html/2609.12113#S4.F3)confirms that increasingρ\\rhoprogressively shifts the generated likelihood distribution toward lower values\. At moderate control strengths, the samples in Figures[4\(a\)](https://arxiv.org/html/2609.12113#S4.F4.sf1)and[4\(b\)](https://arxiv.org/html/2609.12113#S4.F4.sf2)remain visually consistent with CIFAR\-10, suggesting that the controller can access lower\-likelihood regions without immediately destroying the learned data structure\.

![Refer to caption](https://arxiv.org/html/2609.12113v1/figs_png/cifar10_cdf.png)Figure 3:Empirical CDFs of likelihood distributions for different values ofρ\\rho\.Table III:Outlier generation on CIFAR\-10\. Increasing the target parameterρ\\rhoproduces progressively largerW1W\_\{1\}shifts in the likelihood distribution, demonstrating that the proposed controller remains effective on image data\.![Refer to caption](https://arxiv.org/html/2609.12113v1/figs_png/sample_grid_rho_50.png)\(a\)Generated samples withρ=50\\rho=50\.
![Refer to caption](https://arxiv.org/html/2609.12113v1/figs_png/sample_grid_rho_100.png)\(b\)Generated samples withρ=100\\rho=100\.

Figure 4:Generatedρ\\rho\-outliers in the CIFAR\-10 image dataset\. Although the likelihood values are lower, as demonstrated in Figure[3](https://arxiv.org/html/2609.12113#S4.F3), the generated samples remain visually consistent with the underlying data distribution\.Overall, these experiments demonstrate that the proposed framework can steer diffusion sampling toward prescribed likelihood statistics\. In both synthetic and image datasets, the controlled PF\-ODE successfully shifts the likelihood distribution toward lower\-probability regions while preserving the structure of the data distribution\.

## VConclusion

In this work, we introduced a distributional perspective on outliers motivated by the probabilistic structure of diffusion models\. Instead of defining outliers as individual low\-likelihood samples, we formalized them through the distribution of log\-likelihood values and proposed a notion ofρ\\rho\-outlier based on the discrepancy between likelihood distributions\. Using this formulation, we developed a method for generating outliers by modifying the reverse\-time diffusion dynamics\. The key insight is that likelihood reweighting induces a simple modification of the score function via the RN derivative, which motivates the controlled PF\-ODE used for sampling\. Numerical experiments support the theoretical predictions on synthetic and image datasets\. Future work will explore broader applications of this framework to distribution steering\.

## References

- \[1\]V\. S\. Bawa\(1975\)Optimal rules for ordering uncertain prospects\.Journal of Financial economics2\(1\),pp\. 95–121\.Cited by:[Definition 4](https://arxiv.org/html/2609.12113#Thmdefinition4.3)\.
- \[2\]A\. Blattmann, T\. Dockhorn, S\. Kulal, D\. Mendelevitch, M\. Kilian, D\. Lorenz, Y\. Levi, Z\. English, V\. Voleti, A\. Letts,et al\.\(2023\)Stable video diffusion: scaling latent video diffusion models to large datasets\.arXiv preprint arXiv:2311\.15127\.Cited by:[Score\-based Outlier Generation via Controlling the Radon\-Nikodym Derivative](https://arxiv.org/html/2609.12113#p3.1)\.
- \[3\]V\. I\. Bogachev, N\. V\. Krylov, M\. Röckner, and S\. V\. Shaposhnikov\(2015\)Fokker–planck–kolmogorov equations\.Vol\.207,Mathematical Surveys and Monographs\.Cited by:[§I\-B](https://arxiv.org/html/2609.12113#S1.SS2.p1.1)\.
- \[4\]V\. I\. Bogachev\(2018\)Ornstein–uhlenbeck operators and semigroups\.Russian Mathematical Surveys73\(2\),pp\. 191–260\.Cited by:[Theorem 1](https://arxiv.org/html/2609.12113#Thmtheorem1.3)\.
- \[5\]Y\. Chen\(2025\)Log\-sobolev inequalities and markov semigroups\.ETH Zurich\.Note:Lecture 2: Ornstein–Uhlenbeck Semigroup and Gaussian Log\-Sobolev Inequality, ETH Zurich, Week 3–4Cited by:[§III](https://arxiv.org/html/2609.12113#S3.p3.7.1)\.
- \[6\]R\. Chertovskih, N\. Pogodaev, M\. Staritsyn, and A\. P\. Aguiar\(2024\)Optimal control of diffusion processes: infinite\-order variational analysis and numerical solution\.IEEE Control Systems Letters8,pp\. 1469–1474\.Cited by:[§I\-B](https://arxiv.org/html/2609.12113#S1.SS2.p1.2)\.
- \[7\]C\. Dellacherie and P\. Meyer\(1978\)Probabilities and potential, vol\. 29 of north\-holland mathematics studies\.North\-Holland Publishing Co\., Amsterdam\.Cited by:[§II\-B](https://arxiv.org/html/2609.12113#S2.SS2.p2.3.1)\.
- \[8\]D\. Dentcheva and A\. Ruszczynski\(2003\)Optimization with stochastic dominance constraints\.SIAM Journal on Optimization14\(2\),pp\. 548–566\.Cited by:[§II\-A](https://arxiv.org/html/2609.12113#S2.SS1.p4.1)\.
- \[9\]D\. Dentcheva and A\. Ruszczyński\(2004\)Optimality and duality theory for stochastic optimization problems with nonlinear dominance constraints\.Mathematical Programming99\(2\),pp\. 329–350\.Cited by:[§II\-A](https://arxiv.org/html/2609.12113#S2.SS1.p4.1)\.
- \[10\]D\. Dentcheva, M\. Ye, and Y\. Yi\(2022\)Risk\-averse sequential decision problems with time\-consistent stochastic dominance constraints\.In2022 IEEE 61st Conference on Decision and Control \(CDC\),pp\. 3605–3610\.Cited by:[§II\-A](https://arxiv.org/html/2609.12113#S2.SS1.p4.1)\.
- \[11\]X\. Du, Y\. Sun, J\. Zhu, and Y\. Li\(2023\)Dream the impossible: outlier imagination with diffusion models\.Advances in Neural Information Processing Systems36,pp\. 60878–60901\.Cited by:[Score\-based Outlier Generation via Controlling the Radon\-Nikodym Derivative](https://arxiv.org/html/2609.12113#p2.1)\.
- \[12\]A\. Fleig and R\. Guglielmi\(2017\)Optimal control of the fokker–planck equation with space\-dependent controls\.Journal of Optimization Theory and Applications174\(2\),pp\. 408–427\.Cited by:[§I\-B](https://arxiv.org/html/2609.12113#S1.SS2.p1.2)\.
- \[13\]P\. H\. Franses and D\. Van Dijk\(2000\)Non\-linear time series models in empirical finance\.Cambridge university press\.Cited by:[Score\-based Outlier Generation via Controlling the Radon\-Nikodym Derivative](https://arxiv.org/html/2609.12113#p1.1)\.
- \[14\]J\. Gu, X\. Zhang, and G\. Wang\(2025\)Beyond the norm: a survey of synthetic data generation for rare events\.arXiv preprint arXiv:2506\.06380\.Cited by:[Score\-based Outlier Generation via Controlling the Radon\-Nikodym Derivative](https://arxiv.org/html/2609.12113#p2.1)\.
- \[15\]M\. Hermite\(1864\)Sur un nouveau développement en série des fonctions\.Imprimerie de Gauthier\-Villars\.Cited by:[§III](https://arxiv.org/html/2609.12113#S3.p3.4.1),[Theorem 1](https://arxiv.org/html/2609.12113#Thmtheorem1.p1.1.1)\.
- \[16\]J\. Ho, W\. Chan, C\. Saharia, J\. Whang, R\. Gao, A\. Gritsenko, D\. P\. Kingma, B\. Poole, M\. Norouzi, D\. J\. Fleet,et al\.\(2022\)Imagen video: high definition video generation with diffusion models\.arXiv preprint arXiv:2210\.02303\.Cited by:[Score\-based Outlier Generation via Controlling the Radon\-Nikodym Derivative](https://arxiv.org/html/2609.12113#p3.1)\.
- \[17\]J\. Ho, A\. Jain, and P\. Abbeel\(2020\)Denoising diffusion probabilistic models\.InAdvances in Neural Information Processing Systems,External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2020/file/4c5bcfec8584af0d967f1ab10179ca4b-Paper.pdf)Cited by:[§IV\-B](https://arxiv.org/html/2609.12113#S4.SS2.p1.1),[Score\-based Outlier Generation via Controlling the Radon\-Nikodym Derivative](https://arxiv.org/html/2609.12113#p3.1)\.
- \[18\]B\. Hu, A\. Saragadam, A\. Layton, and H\. Chen\(2024\)Synthetic data from diffusion models improves drug discovery prediction\.In2024 IEEE International Conference on Bioinformatics and Biomedicine \(BIBM\),pp\. 6278–6285\.Cited by:[Score\-based Outlier Generation via Controlling the Radon\-Nikodym Derivative](https://arxiv.org/html/2609.12113#p1.1)\.
- \[19\]C\. Lai, Y\. Takida, N\. Murata, T\. Uesaka, Y\. Mitsufuji, and S\. Ermon\(2023\)FP\-Diffusion: improving score\-based diffusion models by enforcing the underlying score Fokker\-Planck equation\.InProceedings of the 40th International Conference on Machine LearningInternational Conference on Machine Learning,Cited by:[Proposition 1](https://arxiv.org/html/2609.12113#Thmproposition1.3)\.
- \[20\]D\. Leao Jr, M\. Fragoso, and P\. Ruffino\(2004\)Regular conditional probability, disintegration of probability and radon spaces\.Proyecciones \(Antofagasta\)23\(1\),pp\. 15–29\.Cited by:[Definition 3](https://arxiv.org/html/2609.12113#Thmdefinition3.p1.2.1)\.
- \[21\]W\. Peebles and S\. Xie\(2023\)Scalable diffusion models with transformers\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 4195–4205\.Cited by:[Score\-based Outlier Generation via Controlling the Radon\-Nikodym Derivative](https://arxiv.org/html/2609.12113#p3.1)\.
- \[22\]R\. Rombach, A\. Blattmann, D\. Lorenz, P\. Esser, and B\. Ommer\(2022\)High\-resolution image synthesis with latent diffusion models\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 10684–10695\.Cited by:[Score\-based Outlier Generation via Controlling the Radon\-Nikodym Derivative](https://arxiv.org/html/2609.12113#p3.1)\.
- \[23\]A\. Sauer, F\. Boesel, T\. Dockhorn, A\. Blattmann, P\. Esser, and R\. Rombach\(2024\)Fast high\-resolution image synthesis with latent adversarial diffusion distillation\.InSIGGRAPH Asia 2024 Conference Papers,pp\. 1–11\.Cited by:[Score\-based Outlier Generation via Controlling the Radon\-Nikodym Derivative](https://arxiv.org/html/2609.12113#p3.1)\.
- \[24\]K\. Shah, S\. Chen, and A\. Klivans\(2023\)Learning mixtures of gaussians using the ddpm objective\.Advances in Neural Information Processing Systems36,pp\. 19636–19649\.Cited by:[§IV\-A](https://arxiv.org/html/2609.12113#S4.SS1.p1.2)\.
- \[25\]C\. Sinigaglia, A\. Manzoni, and F\. Braghin\(2022\)Density control of large\-scale particles swarm through pde\-constrained optimization\.IEEE Transactions on Robotics38\(6\),pp\. 3530–3549\.Cited by:[§I\-B](https://arxiv.org/html/2609.12113#S1.SS2.p1.2)\.
- \[26\]Y\. Song, J\. Sohl\-Dickstein, D\. P\. Kingma, A\. Kumar, S\. Ermon, and B\. Poole\(2021\)Score\-based generative modeling through stochastic differential equations\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=PxTIG12RRHS)Cited by:[§I\-A](https://arxiv.org/html/2609.12113#S1.SS1.p1.1),[§I\-A](https://arxiv.org/html/2609.12113#S1.SS1.p1.3)\.
- \[27\]L\. Tao, X\. Du, X\. Zhu, and Y\. Li\(2023\)Non\-parametric outlier synthesis\.arXiv preprint arXiv:2303\.02966\.Cited by:[Score\-based Outlier Generation via Controlling the Radon\-Nikodym Derivative](https://arxiv.org/html/2609.12113#p2.1)\.
- \[28\]C\. Villani\(2021\)Topics in optimal transportation\.Vol\.58,American Mathematical Soc\.\.Cited by:[Remark 1](https://arxiv.org/html/2609.12113#Thmremark1.p2.1.1)\.

Similar Articles

Generator-Guided Inverse Sampling for L\'evy-Driven Generative Models

arXiv cs.LG

This paper studies inverse sampling for Lévy-driven generative models, proposing a structured reverse sampler that decomposes dynamics into diffusion, small jump, and large jump components, with neural networks amortizing jump rates. The method is applied to OFDM-SISO channel estimation under mixed Gaussian and impulsive noise.

Specificity-Aware Diffusion Steering via Variance-Reduced Sequential Monte Carlo

arXiv cs.LG

The paper introduces a specificity-aware diffusion steering method that designs an explicit target distribution from an overlap-based objective and samples it with a variance-reduced Sequential Monte Carlo sampler, improving negative-guidance effectiveness and sampling stability across image generation and peptide-MHC binder tasks.