Unifying Generative Models with Path Integrals

arXiv cs.LG Papers

Summary

This paper proposes a unified path-integral framework for generative modeling, showing that flow-based, diffusion-based, variational, and adversarial models arise from different evaluation principles of a single master action. It derives a one-loop correction that reduces tree-level error dramatically and introduces a response-weighted score-matching objective.

arXiv:2608.12438v1 Announce Type: new Abstract: We formulate generative modeling as a path integral in which flow-based, diffusion-based, variational, and adversarial models arise as different evaluation principles for a single master action. Its Martin-Siggia-Rose-Janssen-de~Dominicis (MSRJD) form separates free from interacting probability flows and opens them to diagrammatic perturbation theory. The expansion yields a one-loop correction to deterministic samplers at no stochastic-sampling cost, which we validate on solvable and nonlinear drifts, where it reduces a 53 % tree-level error to 1.6 %. Imperfect learned scores enter as insertions and yield a response-weighted score-matching objective, and symmetry-equivariant drift design becomes an operator expansion with EFT power counting.
Original Article
View Cached Full Text

Cached at: 08/14/26, 09:29 AM

# 1Introduction
Source: [https://arxiv.org/html/2608.12438](https://arxiv.org/html/2608.12438)
TIF\-UNIMI\-2026\-11

Unifying Generative Models with Path Integrals

Ramon Winterhalder

TIFLab, Dipartimento di Fisica, Università degli Studi di Milano,

and INFN, Sezione di Milano, Via Celoria 16, I\-20133 Milano, Italy

Abstract

We formulate generative modeling as a path integral in which flow\-based, diffusion\-based, variational, and adversarial models arise as different evaluation principles for a single master action\. Its Martin\-Siggia\-Rose\-Janssen\-de Dominicis \(MSRJD\) form separates free from interacting probability flows and opens them to diagrammatic perturbation theory\. The expansion yields a one\-loop correction to deterministic samplers at no stochastic\-sampling cost, which we validate on solvable and nonlinear drifts, where it reduces a53%53\\,\\%tree\-level error to1\.6%1\.6\\,\\%\. Imperfect learned scores enter as insertions and yield a response\-weighted score\-matching objective, and symmetry\-equivariant drift design becomes an operator expansion with EFT power counting\.

Contents

## 1Introduction

Generative models sample complex, high\-dimensional probability distributions fast and accurate enough to be relevant for physics, where probabilistic modeling carries simulation, inference, and uncertainty quantification\. In high\-energy physics \(HEP\) they are being developed as accelerators for Monte Carlo event generation, detector simulation, and unfolding, while maintaining physical fidelity\[[7](https://arxiv.org/html/2608.12438#bib.bib71),[60](https://arxiv.org/html/2608.12438#bib.bib72),[2](https://arxiv.org/html/2608.12438#bib.bib73)\]\. See the HEP\-ML Living Review\[[20](https://arxiv.org/html/2608.12438#bib.bib74),[29](https://arxiv.org/html/2608.12438#bib.bib75),[45](https://arxiv.org/html/2608.12438#bib.bib19)\]and HEP\-ML Living Guide\[[46](https://arxiv.org/html/2608.12438#bib.bib2)\]for a broader overview\.

Generative models are typically introduced within their own mathematical language and training paradigm\. For instance, normalizing flows construct explicit invertible maps with tractable Jacobians, and diffusion models define generation as a time\-reversed stochastic process\. From a physics perspective, this fragmentation is surprising\. Physics describes probabilistic dynamics through stochastic processes and path integrals, which provide a unified language for dynamics, fluctuations, and correlations and are routine tools in statistical mechanics and quantum field theory\. In particular, the Onsager–Machlup formulation\[[58](https://arxiv.org/html/2608.12438#bib.bib59),[51](https://arxiv.org/html/2608.12438#bib.bib60)\]and the Martin\-Siggia\-Rose\-Janssen\-de Dominicis \(MSRJD\) functional integral\[[52](https://arxiv.org/html/2608.12438#bib.bib68),[35](https://arxiv.org/html/2608.12438#bib.bib67),[17](https://arxiv.org/html/2608.12438#bib.bib66)\]provide path\-integral descriptions of stochastic dynamics that are tailor\-made for the kind of probability transport generative models perform\.

Prior work has already explored connections between subsets of these frameworks\. The score\-based SDE framework unifies score matching and diffusion through stochastic differential equations running forward and backward in time\[[69](https://arxiv.org/html/2608.12438#bib.bib55)\], flow matching and stochastic\-interpolant constructions provide deterministic training objectives for continuous\-time transport\[[49](https://arxiv.org/html/2608.12438#bib.bib54),[1](https://arxiv.org/html/2608.12438#bib.bib53)\], and Refs\.\[[12](https://arxiv.org/html/2608.12438#bib.bib52),[9](https://arxiv.org/html/2608.12438#bib.bib51)\]have developed the connections between diffusion models, optimal control, and Schrödinger bridges\. Path\-integral and action\-based viewpoints have appeared in Refs\.\[[57](https://arxiv.org/html/2608.12438#bib.bib65),[72](https://arxiv.org/html/2608.12438#bib.bib58),[33](https://arxiv.org/html/2608.12438#bib.bib44),[62](https://arxiv.org/html/2608.12438#bib.bib46),[31](https://arxiv.org/html/2608.12438#bib.bib45)\]\. Closest to our work, Ref\.\[[31](https://arxiv.org/html/2608.12438#bib.bib45)\]formulates score\-based diffusion models as path integrals and evaluates likelihood corrections in a semiclassical Wentzel–Kramers–Brillouin \(WKB\) expansion of an interpolating parameter connecting the probability\-flow ODE to the reverse SDE\. No existing formalism does both\. The works above unify subsets of these models, and Ref\.\[[31](https://arxiv.org/html/2608.12438#bib.bib45)\]extracts corrections from a path integral, but not within a structure that also classifies the models it corrects\. Our construction differs in both scope and output\. It embeds the full family of models above in one master action, its loop expansion corrects sample observables perturbatively without direct stochastic sampling, and it develops the operator side of the EFT analogy\.

Starting from latent\-variable models and their extension to chains of latent variables, we show that generative modeling can be formulated as a path integral over latent trajectories governed by an Onsager–Machlup action\. Within this framework, normalizing flows, diffusion models, variational autoencoders, and generative adversarial networks emerge as distinct limits or approximations of a single master action, providing a classification scheme\. The classification is the starting point, but the reformulation also produces explicit results\. Once generative transport is a field theory, the error of a deterministic sampler, the effect of an imperfect score, and the design of a symmetry\-respecting architecture become perturbative computations, with the same diagrammatic bookkeeping as corrections to a scattering process in quantum field theory\. The contributions of this work are:

1. \(i\)We give a self\-contained derivation that recovers normalizing flows, continuous normalizing flows, diffusion models, conditional flow matching, Schrödinger bridges, variational autoencoders, and generative adversarial networks as limits, special cases, or evaluation principles of a single Onsager–Machlup master action;
2. \(ii\)We recast this action in MSRJD form and identify the free/interacting split, its propagators, and a diagrammatic expansion of the transition kernel;
3. \(iii\)We use this structure to organize the discrepancy between a deterministic drift\-based sampler and the full stochastic reverse process as a loop expansion in the diffusion strength\. At one\-loop order the expansion yields a covariance correction and a tadpole mean shift, obtained by integrating auxiliary equations alongside the deterministic trajectory\. We compute these terms in closed form and validate them on an exactly solvable model and on nonlinear drifts;
4. \(iv\)We treat an imperfect learned score as a diagrammatic insertion and derive the corresponding response\-weighted score\-matching objective;
5. \(v\)We show that EFT power counting, combined with a symmetry requirement, organizes the construction of equivariant drift architectures into a finite operator enumeration with a predicted coupling hierarchy\.

Throughout, the emphasis is theoretical, as we train no generative network\. The numerical studies validate the loop formulas on synthetic drifts, where the exact answer is available\.

The paper is organized as follows\. In[Section2](https://arxiv.org/html/2608.12438#S2), we formulate generative models as continuous probabilistic paths and derive the corresponding path\-integral representation\.[Section2\.3](https://arxiv.org/html/2608.12438#S2.SS3)demonstrates how standard classes of generative models are recovered as limits of the master action\. In[Section3](https://arxiv.org/html/2608.12438#S3), we develop an analogy with statistical field theory, introducing free and interacting probability flows, a scattering interpretation, and the relation to the MSRJD formalism\. In[Section4](https://arxiv.org/html/2608.12438#S4), we apply this structure to derive and validate a loop expansion for the error of deterministic samplers\. In[Section5](https://arxiv.org/html/2608.12438#S5), we use effective\-field\-theory power counting to construct symmetry\-constrained probability flows\. We conclude with an outlook in[Section6](https://arxiv.org/html/2608.12438#S6)\.

## 2Generative models as continuous probabilistic paths

Latent\-variable models and their extension to chains of latent variables lead to a path\-integral representation\. Taking the continuum limit yields stochastic or deterministic dynamics in latent space, whose probability flow can be described by an effective action\. This construction provides the common foundation on which different classes of generative models are built\.

### 2\.1From latent\-variable models to path integrals

Latent\-variable models are among the most widely used probabilistic formulations of generative modeling\[[10](https://arxiv.org/html/2608.12438#bib.bib61),[11](https://arxiv.org/html/2608.12438#bib.bib62)\]\. In their simplest form, they express the data distributionpdata​\(x\)p\_\{\\rm data\}\(x\)as a marginal over latent degrees of freedomz∈Rdz\\in\\mdmathbb\{R\}^\{d\},

pdata​\(x\)=∫d​z​p​\(x\|z\)​pprior​\(z\),\\displaystyle p\_\{\\rm data\}\(x\)=\\int\\mathrm\{d\}z\\;p\(x\|z\)\\,p\_\{\\rm prior\}\(z\)\\;,\(2\.1\)wherepprior​\(z\)p\_\{\\rm prior\}\(z\)denotes a prior distribution andp⁡\(x\|z\)p\(x\|z\)a conditional likelihood\. The central challenge is that the marginalpdata​\(x\)p\_\{\\rm data\}\(x\)is generally intractable for expressive choices ofp⁡\(x\|z\)p\(x\|z\), since the integral over the latent space cannot be evaluated in closed form\. This also renders the true posterior,

p⁡\(z\|x\)=p⁡\(x\|z\)​pprior​\(z\)pdata​\(x\),\\displaystyle p\(z\|x\)=\\frac\{p\(x\|z\)\\,p\_\{\\rm prior\}\(z\)\}\{p\_\{\\rm data\}\(x\)\}\\;,\(2\.2\)intractable, as its normalization requires the very quantitypdata​\(x\)p\_\{\\rm data\}\(x\)one seeks to compute\. Variational autoencoders \(VAEs\)\[[40](https://arxiv.org/html/2608.12438#bib.bib63),[41](https://arxiv.org/html/2608.12438#bib.bib64)\]address this by approximating the true posterior with a parametric inference networkq�​\(z\|x\)q\_\{\\phi\}\(z\|x\), trained to be as close as possible top⁡\(z\|x\)p\(z\|x\)\. Normalizing flows avoid this difficulty entirely by restricting to the deterministic special casep⁡\(x\|z\)=�​\(x−�⁡\(z\)\)p\(x\|z\)=\\delta\(x\-\\Phi\(z\)\)with an invertible map�\\Phi, for which we can evaluate the marginal exactly\. We return to this in[Section2\.3](https://arxiv.org/html/2608.12438#S2.SS3)\.

A natural extension of[Eq\.2\.1](https://arxiv.org/html/2608.12438#S2.E1)is to introduce a sequence of latent variables\{zk\}k=0N\\\{z\_\{k\}\\\}\_\{k=0\}^\{N\}interpolating between a simple priorpprior​\(zN\)p\_\{\\rm prior\}\(z\_\{N\}\)and the data distributionpdata​\(z0\)≡p⁡\(x\)p\_\{\\rm data\}\(z\_\{0\}\)\\equiv p\(x\),

pdata​\(z0\)=∫\[∏k=1Nd​zk​p​\(zk−1\|zk\)\]​pprior​\(zN\),\\displaystyle p\_\{\\rm data\}\(z\_\{0\}\)=\\int\\left\[\\prod\_\{k=1\}^\{N\}\\mathrm\{d\}z\_\{k\}\\;p\(z\_\{k\-1\}\|z\_\{k\}\)\\right\]p\_\{\\rm prior\}\(z\_\{N\}\)\\;,\(2\.3\)which can be interpreted as a discrete path integral over latent trajectories\. Here and in the following we identify the data variable with the terminal element of the chain,z0≡xz\_\{0\}\\equiv x, writingz0z\_\{0\}whenever its role as the endpoint of a latent trajectory matters and keepingxxwhen we emphasize contact with the data\. Each transition kernelp⁡\(zk−1\|zk\)p\(z\_\{k\-1\}\|z\_\{k\}\)defines one step of a Markov chain in the generative direction, and marginalization over intermediate states amounts to summing over all paths connecting the prior to the observation\. Introducing the discrete action

SN\[z0:N\]=−∑k=1Nlogp\(zk−1\|zk\),\\displaystyle S\_\{N\}\[z\_\{0:N\}\]=\-\\sum\_\{k=1\}^\{N\}\\log p\(z\_\{k\-1\}\|z\_\{k\}\)\\;,\(2\.4\)we can write[Eq\.2\.3](https://arxiv.org/html/2608.12438#S2.E3)in the exponential form

pdata\(z0\)=∫\[∏k=1Ndzk\]pprior\(zN\)exp\[−SN\[z0:N\]\],\\displaystyle p\_\{\\rm data\}\(z\_\{0\}\)=\\int\\Biggl\[\\prod\_\{k=1\}^\{N\}\\mathrm\{d\}z\_\{k\}\\Biggr\]\\,p\_\{\\rm prior\}\(z\_\{N\}\)\\,\\exp\\\!\\big\[\-S\_\{N\}\[z\_\{0:N\}\]\\big\]\\;,\(2\.5\)which illustrates the analogy to discrete path integrals from statistical mechanics\[[42](https://arxiv.org/html/2608.12438#bib.bib39)\]\. At this stage, the actionSNS\_\{N\}fully specifies the generative model through its transition probabilities\.

### 2\.2Continuum limit and the Onsager–Machlup action

We now take the continuum limit of the discrete latent chain in[Eq\.2\.5](https://arxiv.org/html/2608.12438#S2.E5)\. Recall that the kernelp⁡\(zk−1\|zk\)p\(z\_\{k\-1\}\|z\_\{k\}\)in the discrete action[2\.4](https://arxiv.org/html/2608.12438#S2.E4)runs in the*generative*direction,i\.e\.from prior \(k=Nk=N\) to data \(k=0k=0\)\. So far this kernel is arbitrary, and[Eq\.2\.4](https://arxiv.org/html/2608.12438#S2.E4)is a rewriting valid for any Markov chain\. A continuum limit exists only once the one\-step kernel is specified and carries a single power of�​t\\Delta tsuch that the sum turns into a Riemann sum\. Diffusion processes provide such a kernel, which is why we define a forward SDE, invoke its time reversal as a known result, and discretize that reverse process to obtainp⁡\(zk−1\|zk\)p\(z\_\{k\-1\}\|z\_\{k\}\)in closed form\. The continuum limit of the action then follows\.

We set�​t=1/N\\Delta t=1/Nand place the chain on the uniform gridtk=k​�​tt\_\{k\}=k\\,\\Delta tfork=0,1,…,Nk=0,1,\\ldots,N, so thatt0=0t\_\{0\}=0corresponds to data andtN=1t\_\{N\}=1corresponds to the prior\. From here on we write a grid state asz⁡\(tk\)z\(t\_\{k\}\)and reserve a subscript onzzfor a time argument, so thatz0≡z⁡\(0\)z\_\{0\}\\equiv z\(0\)andz1≡z⁡\(1\)z\_\{1\}\\equiv z\(1\)denote the data and prior endpoints\. Hence, we adopt the following convention throughout:

- •the*forward*process runs fromt=0t=0\(data\) tot=1t=1\(prior\),
- •and the*reverse*\(generative\) process runs fromt=1t=1\(prior\) back tot=0t=0\(data\)\.

We start by modeling the forward process as a diffusion\. Over a short interval�​t\\Delta t, the state evolves from the earlier statez⁡\(tk−1\)z\(t\_\{k\-1\}\)to the later statez⁡\(tk\)z\(t\_\{k\}\)as

z⁡\(tk\)=z⁡\(tk−1\)\+f⁡\(z⁡\(tk−1\),tk−1\)​�​t\+g⁡\(tk−1\)​�​t​�k,�k∼𝒩⁡\(0,I\),\\displaystyle z\(t\_\{k\}\)=z\(t\_\{k\-1\}\)\+f\(z\(t\_\{k\-1\}\),t\_\{k\-1\}\)\\,\\Delta t\+g\(t\_\{k\-1\}\)\\,\\sqrt\{\\Delta t\}\\,\\xi\_\{k\}\\;,\\qquad\\xi\_\{k\}\\sim\\mathcal\{N\}\(0,\\mdmathbb\{I\}\)\\;,\(2\.6\)wheref⁡\(z,t\)f\(z,t\)is the forward drift andg⁡\(t\)≥0g\(t\)\\geq 0a scalar diffusion coefficient\.[Equation2\.6](https://arxiv.org/html/2608.12438#S2.E6)is the Euler–Maruyama step of a diffusion process where the state moves deterministically byf​�​tf\\,\\Delta tand receives a Gaussian kick of sizeg​�​tg\\,\\sqrt\{\\Delta t\}, the�​t\\sqrt\{\\Delta t\}scaling being characteristic of Brownian noise\. A common choice is the variance\-preserving processf⁡\(z,t\)=−12​�​\(t\)​zf\(z,t\)=\-\\tfrac\{1\}\{2\}\\beta\(t\)\\,zwithg⁡\(t\)=�​\(t\)g\(t\)=\\sqrt\{\\beta\(t\)\}, which drives any data distribution toward a standard\-normal prior, while the variance\-exploding process takesf=0f=0with a growing diffusiong⁡\(t\)g\(t\)\[[69](https://arxiv.org/html/2608.12438#bib.bib55)\]\. Taking�​t→0\\Delta t\\to 0defines the forward stochastic differential equation \(SDE\)

d​z=f⁡\(z,t\)​d​t\+g⁡\(t\)​d​Wt,\\displaystyle\\mathrm\{d\}z=f\(z,t\)\\,\\mathrm\{d\}t\+g\(t\)\\,\\mathrm\{d\}W\_\{t\}\\;,\(2\.7\)withWtW\_\{t\}a standard Wiener process,i\.e\.the continuous\-time limit of a random walk\[[23](https://arxiv.org/html/2608.12438#bib.bib57)\]\. The induced densityp⁡\(z,t\)p\(z,t\), with boundary conditions

p⁡\(z,t\)→\{p⁡\(z,0\)≡pdata​\(z\)t→0p⁡\(z,1\)≡pprior​\(z\)t→1,\\displaystyle p\(z,t\)\\to\\begin\{cases\}p\(z,0\)\\equiv p\_\{\\rm data\}\(z\)\\qquad&t\\to 0\\\\ p\(z,1\)\\equiv p\_\{\\rm prior\}\(z\)\\qquad&t\\to 1\\;,\\end\{cases\}\(2\.8\)evolves according to the Fokker–Planck equation\[[64](https://arxiv.org/html/2608.12438#bib.bib56),[23](https://arxiv.org/html/2608.12438#bib.bib57)\]

∂tp\(z,t\)=−∇z⋅\(f\(z,t\)p\(z,t\)\)\+12g2\(t\)∇z2p\(z,t\)\.\\displaystyle\\partial\_\{t\}p\(z,t\)=\-\\nabla\_\{z\}\\cdot\\big\(f\(z,t\)\\,p\(z,t\)\\big\)\+\\tfrac\{1\}\{2\}g^\{2\}\(t\)\\,\\nabla\_\{z\}^\{2\}p\(z,t\)\\;\.\(2\.9\)The generative model runs this process backward\. The time reversal of the diffusion process in[Eq\.2\.7](https://arxiv.org/html/2608.12438#S2.E7)is itself a diffusion, satisfying the SDE\[[3](https://arxiv.org/html/2608.12438#bib.bib69),[28](https://arxiv.org/html/2608.12438#bib.bib70)\]

d​z=frev​\(z,t\)​d​t\+g⁡\(t\)​d​W¯t,frev​\(z,t\)≡f⁡\(z,t\)−g2​\(t\)​∇z​log⁡p⁡\(z,t\),\\displaystyle\\mathrm\{d\}z=f\_\{\\rm rev\}\(z,t\)\\,\\mathrm\{d\}t\+g\(t\)\\,\\mathrm\{d\}\\bar\{W\}\_\{t\}\\;,\\qquad f\_\{\\rm rev\}\(z,t\)\\equiv f\(z,t\)\-g^\{2\}\(t\)\\,\\nabla\_\{z\}\\log p\(z,t\)\\;,\(2\.10\)whereW¯t\\bar\{W\}\_\{t\}is a Wiener process run in the reverse direction and the reverse driftfrevf\_\{\\rm rev\}differs from the forward drift by the score∇z​log​p​\(z,t\)\\nabla\_\{z\}\\log p\(z,t\)of the density\. The score term corrects the forward drift for the reversal of time, pointing toward regions of higher probability mass and counteracting the spreading induced by diffusion\.

To obtain a path\-integral representation, we discretize the reverse process\. On the gridtk=k​�​tt\_\{k\}=k\\,\\Delta t, the standard Euler–Maruyama discretization of the reverse SDE\[[53](https://arxiv.org/html/2608.12438#bib.bib40),[43](https://arxiv.org/html/2608.12438#bib.bib41),[30](https://arxiv.org/html/2608.12438#bib.bib42)\]expresses the earlier statez⁡\(tk−1\)z\(t\_\{k\-1\}\)in terms of the later statez⁡\(tk\)z\(t\_\{k\}\)as

z⁡\(tk−1\)=z⁡\(tk\)−frev​\(z⁡\(tk\),tk\)​�​t\+g⁡\(tk\)​�​t​�k,�k∼𝒩⁡\(0,I\),\\displaystyle z\(t\_\{k\-1\}\)=z\(t\_\{k\}\)\-f\_\{\\rm rev\}\(z\(t\_\{k\}\),t\_\{k\}\)\\,\\Delta t\+g\(t\_\{k\}\)\\,\\sqrt\{\\Delta t\}\\,\\xi\_\{k\}\\;,\\qquad\\xi\_\{k\}\\sim\\mathcal\{N\}\(0,\\mdmathbb\{I\}\)\\;,\(2\.11\)the reverse\-time counterpart of[Eq\.2\.6](https://arxiv.org/html/2608.12438#S2.E6), with the drift contributing−frev​�​t\-f\_\{\\rm rev\}\\,\\Delta tas the step moves backward intt\. Equivalently, since the Euler–Maruyama step is linear in the Gaussian noise�k\\xi\_\{k\}, the transition is Gaussian in terms of the forward increment�​zk≡z⁡\(tk\)−z⁡\(tk−1\)\\Delta z\_\{k\}\\equiv z\(t\_\{k\}\)\-z\(t\_\{k\-1\}\),

p\(z\(tk−1\),tk−1∣z\(tk\),tk\)\\displaystyle p\(z\(t\_\{k\-1\}\),\\,t\_\{k\-1\}\\mid z\(t\_\{k\}\),\\,t\_\{k\}\)=𝒩k​exp⁡\[−‖�​zk−frev​\(z⁡\(tk\),tk\)​�​t‖22​g2​\(tk\)​�​t\],\\displaystyle=\\mathcal\{N\}\_\{k\}\\,\\exp\\\!\\left\[\-\\frac\{\\\|\\Delta z\_\{k\}\-f\_\{\\rm rev\}\(z\(t\_\{k\}\),t\_\{k\}\)\\,\\Delta t\\\|^\{2\}\}\{2g^\{2\}\(t\_\{k\}\)\\,\\Delta t\}\\right\]\\;,with𝒩k\\displaystyle\\text\{with\}\\qquad\\mathcal\{N\}\_\{k\}=\(2�g2\(tk\)�t\)−d/2\.\\displaystyle=\\big\(2\\pi g^\{2\}\(t\_\{k\}\)\\,\\Delta t\\big\)^\{\-d/2\}\\;\.\(2\.12\)At fixedz⁡\(tk\)z\(t\_\{k\}\), the densities in the earlier statez⁡\(tk−1\)z\(t\_\{k\-1\}\)and in the increment�​zk\\Delta z\_\{k\}differ only by a constant shift with unit Jacobian, so the Gaussian in the increment is the conditional density itself\. This is exactly the transition kernelp⁡\(zk−1\|zk\)p\(z\_\{k\-1\}\|z\_\{k\}\)entering the discrete chain in[Eq\.2\.3](https://arxiv.org/html/2608.12438#S2.E3), now given an explicit Gaussian form by the reverse SDE\. Taking the logarithm,

−logp\(z\(tk−1\),tk−1∣z\(tk\),tk\)=‖�​zk−frev​\(z⁡\(tk\),tk\)​�​t‖22​g2​\(tk\)​�​t−log𝒩k\.\\displaystyle\-\\log p\(z\(t\_\{k\-1\}\),\\,t\_\{k\-1\}\\mid z\(t\_\{k\}\),\\,t\_\{k\}\)=\\frac\{\\\|\\Delta z\_\{k\}\-f\_\{\\rm rev\}\(z\(t\_\{k\}\),t\_\{k\}\)\\,\\Delta t\\\|^\{2\}\}\{2g^\{2\}\(t\_\{k\}\)\\,\\Delta t\}\-\\log\\mathcal\{N\}\_\{k\}\\;\.\(2\.13\)The second term iszz\-independent but we track it explicitly, as it requires care in the continuum limit\. For the first term, we introduce the forward velocity at stepkkas

z˙​\(tk\)≡�​zk�​t=z⁡\(tk\)−z⁡\(tk−1\)�​t,\\displaystyle\\dot\{z\}\(t\_\{k\}\)\\equiv\\frac\{\\Delta z\_\{k\}\}\{\\Delta t\}=\\frac\{z\(t\_\{k\}\)\-z\(t\_\{k\-1\}\)\}\{\\Delta t\}\\;,\(2\.14\)so that

‖�​zk−frev​\(z⁡\(tk\),tk\)​�​t‖22​g2​\(tk\)​�​t=�​t2​g2​\(tk\)​‖z˙​\(tk\)−frev​\(z⁡\(tk\),tk\)‖2\.\\displaystyle\\frac\{\\\|\\Delta z\_\{k\}\-f\_\{\\rm rev\}\(z\(t\_\{k\}\),t\_\{k\}\)\\,\\Delta t\\\|^\{2\}\}\{2g^\{2\}\(t\_\{k\}\)\\,\\Delta t\}=\\frac\{\\Delta t\}\{2g^\{2\}\(t\_\{k\}\)\}\\\|\\dot\{z\}\(t\_\{k\}\)\-f\_\{\\rm rev\}\(z\(t\_\{k\}\),t\_\{k\}\)\\\|^\{2\}\\;\.\(2\.15\)Inserting into the discrete action gives azz\-dependent sum together with the collected Gaussian normalizations,

SN\[z0:N\]\\displaystyle S\_\{N\}\[z\_\{0:N\}\]=∑k=1N�​t2​g2​\(tk\)​‖z˙​\(tk\)−frev​\(z⁡\(tk\),tk\)‖2\+𝒞N,\\displaystyle=\\sum\_\{k=1\}^\{N\}\\frac\{\\Delta t\}\{2g^\{2\}\(t\_\{k\}\)\}\\,\\bigl\\\|\\dot\{z\}\(t\_\{k\}\)\-f\_\{\\rm rev\}\(z\(t\_\{k\}\),t\_\{k\}\)\\bigr\\\|^\{2\}\+\\mathcal\{C\}\_\{N\}\\;,with𝒞N\\displaystyle\\text\{with\}\\qquad\\mathcal\{C\}\_\{N\}=−∑k=1Nlog𝒩k\.\\displaystyle=\-\\sum\_\{k=1\}^\{N\}\\log\\mathcal\{N\}\_\{k\}\\;\.\(2\.16\)Thezz\-dependent sum carries a single power of�​t\\Delta tper term and is thus a Riemann sum\. SendingN→∞N\\to\\infty,i\.e\.�​t→0\\Delta t\\to 0, it converges to the Onsager–Machlup action,

SOM​\[z\]=∫01d​t​‖z˙​\(t\)−frev​\(z⁡\(t\),t\)‖22​g2​\(t\)≡∫01d​t​ℒOM​\(z⁡\(t\),z˙​\(t\)\),\\displaystyle S\_\{\\text\{OM\}\}\[z\]=\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\frac\{\\\|\\dot\{z\}\(t\)\-f\_\{\\rm rev\}\(z\(t\),t\)\\\|^\{2\}\}\{2g^\{2\}\(t\)\}\\equiv\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\mathcal\{L\}\_\{\\text\{OM\}\}\\big\(z\(t\),\\dot\{z\}\(t\)\\big\)\\;,\(2\.17\)where the last line defines the corresponding Onsager–Machlup LagrangianℒOM\\mathcal\{L\}\_\{\\text\{OM\}\}\. Thezz\-independent term𝒞N\\mathcal\{C\}\_\{N\}, by contrast, does*not*converge and it diverges logarithmically as�​t→0\\Delta t\\to 0\. This is the familiar situation for path integrals where the action and the measure do not have separate continuum limits, only their combination does\. We therefore define the continuum path measure𝒟​z\\mathcal\{D\}zto carry the divergent normalization, integrating over the interior of the time grid only,

𝒟​z≡limN→∞e−𝒞N​∏k=1N−1d​z​\(tk\),\\displaystyle\\mathcal\{D\}z\\equiv\\lim\_\{N\\to\\infty\}\\;e^\{\-\\mathcal\{C\}\_\{N\}\}\\prod\_\{k=1\}^\{N\-1\}\\mathrm\{d\}z\(t\_\{k\}\)\\;,\(2\.18\)so that the product𝒟​z​e−SOM​\[z\]\\mathcal\{D\}z\\,e^\{\-S\_\{\\text\{OM\}\}\[z\]\}is finite even though neither factor is\. With this definition the continuum limit is taken for the path integral as a whole, not for the action in isolation\. The assignment of the finite parts between action and measure is itself a convention rather than uniquely fixed,i\.e\.only the combination𝒟​z​e−SOM​\[z\]\\mathcal\{D\}z\\,e^\{\-S\_\{\\text\{OM\}\}\[z\]\}carries meaning\. The same applies to the action itself, which is fixed only up to the discretization convention, cf\. the Itô–Stratonovich discussion below\. The endpointsz⁡\(t0\)z\(t\_\{0\}\)andz⁡\(tN\)z\(t\_\{N\}\)are deliberately excluded from the measure, as their values are held fixed and not integrated over\. The data endpointz0=z⁡\(t0\)z\_\{0\}=z\(t\_\{0\}\)is identified withxxas in[Section2\.1](https://arxiv.org/html/2608.12438#S2.SS1), andz1=z⁡\(tN\)z\_\{1\}=z\(t\_\{N\}\)is the prior endpoint\. This defines the transition kernelKKof the generative process as the continuum limit of the composed chain,

K\(z0,0∣z1,1\)≡limN→∞∫\[∏k=1N−1dz\(tk\)\]∏k=1Np\(z\(tk−1\)∣z\(tk\)\)=∫𝒟zexp\[−SOM\[z\]\],\\displaystyle K\(z\_\{0\},0\\mid z\_\{1\},1\)\\equiv\\lim\_\{N\\to\\infty\}\\int\\Biggl\[\\prod\_\{k=1\}^\{N\-1\}\\mathrm\{d\}z\(t\_\{k\}\)\\Biggr\]\\prod\_\{k=1\}^\{N\}p\\big\(z\(t\_\{k\-1\}\)\\mid z\(t\_\{k\}\)\\big\)=\\int\\mathcal\{D\}z\\;\\exp\\\!\\left\[\-S\_\{\\text\{OM\}\}\[z\]\\right\]\\;,\(2\.19\)whereK\(z0,0∣z1,1\)K\(z\_\{0\},0\\mid z\_\{1\},1\)is the conditional probability of arriving at the data pointz0z\_\{0\}att=0t=0given the prior pointz1z\_\{1\}att=1t=1\. This is a pinned path integral summing over all latent trajectories connecting the two endpoints, weighted by the Onsager–Machlup action\. The boundary values of the path integral are fixed by the arguments of the kernel,z⁡\(0\)=z0z\(0\)=z\_\{0\}andz⁡\(1\)=z1z\(1\)=z\_\{1\}, and by[Eq\.2\.18](https://arxiv.org/html/2608.12438#S2.E18)they are not part of the measure\. We therefore suppress boundary conditions on the integral sign throughout\. The same construction on any subinterval of the grid, with the path integral running over the interior of\[ti,tf\]\[t\_\{i\},t\_\{f\}\]and the valuesz⁡\(ti\)=ziz\(t\_\{i\}\)=z\_\{i\}andz⁡\(tf\)=zfz\(t\_\{f\}\)=z\_\{f\}fixed, yields the two\-time kernel

K\(zi,ti∣zf,tf\)=∫𝒟zexp\[−∫titfdtℒOM\(z\(t\),z˙\(t\)\)\]withti<tf,\\displaystyle K\(z\_\{i\},t\_\{i\}\\mid z\_\{f\},t\_\{f\}\)=\\int\\mathcal\{D\}z\\;\\exp\\\!\\left\[\-\\int\_\{t\_\{i\}\}^\{t\_\{f\}\}\\mathrm\{d\}t\\;\\mathcal\{L\}\_\{\\text\{OM\}\}\\big\(z\(t\),\\dot\{z\}\(t\)\\big\)\\right\]\\qquad\\text\{with\}\\quad t\_\{i\}<t\_\{f\}\\;,\(2\.20\)the transition density of the reverse process from the statezfz\_\{f\}at timetft\_\{f\}to the stateziz\_\{i\}at timetit\_\{i\}, with[Eq\.2\.19](https://arxiv.org/html/2608.12438#S2.E19)the special case connecting prior and data\. Discretizing\[ti,tf\]\[t\_\{i\},t\_\{f\}\]on a grid ofNNsteps with an interior seam at indexss, theNNper\-step normalizations combine intoe−𝒞Ne^\{\-\\mathcal\{C\}\_\{N\}\}while theN−1N\-1integrations run over the interior slices only\. Writing�​t​ℒ\(k\)\\Delta t\\,\\mathcal\{L\}^\{\(k\)\}for thekk\-th summand of the discrete action in[Eq\.2\.16](https://arxiv.org/html/2608.12438#S2.E16), the discrete counterpart ofℒOM\\mathcal\{L\}\_\{\\text\{OM\}\}, the product factorizes at the seam\. Splitting the interior integrals there and taking the continuum limit in the two factors separately yields the exact composition law

K\(zi,ti∣zf,tf\)\\displaystyle K\(z\_\{i\},t\_\{i\}\\mid z\_\{f\},t\_\{f\}\)≡∫𝒟zexp\[−∫titfdtℒOM\(z\(t\),z˙\(t\)\)\]\\displaystyle\\equiv\\int\\mathcal\{D\}z\\;\\exp\\\!\\left\[\-\\int\_\{t\_\{i\}\}^\{t\_\{f\}\}\\mathrm\{d\}t\\;\\mathcal\{L\}\_\{\\text\{OM\}\}\\big\(z\(t\),\\dot\{z\}\(t\)\\big\)\\right\]=limN→∞∫∏k=1N−1d​z​\(tk\)​∏k=1N𝒩k​e−�​t​ℒ\(k\)\\displaystyle=\\lim\_\{N\\to\\infty\}\\int\\prod\_\{k=1\}^\{N\-1\}\\mathrm\{d\}z\(t\_\{k\}\)\\;\\prod\_\{k=1\}^\{N\}\\mathcal\{N\}\_\{k\}\\,e^\{\-\\Delta t\\,\\mathcal\{L\}^\{\(k\)\}\}=limN→∞∫∏k=1N−1d​z​\(tk\)​\(∏k=1s𝒩k​e−�​t​ℒ\(k\)\)​\(∏k=s\+1N𝒩k​e−�​t​ℒ\(k\)\)\\displaystyle=\\lim\_\{N\\to\\infty\}\\int\\prod\_\{k=1\}^\{N\-1\}\\mathrm\{d\}z\(t\_\{k\}\)\\;\\Bigg\(\\prod\_\{k=1\}^\{s\}\\mathcal\{N\}\_\{k\}\\,e^\{\-\\Delta t\\,\\mathcal\{L\}^\{\(k\)\}\}\\Bigg\)\\Bigg\(\\prod\_\{k=s\+1\}^\{N\}\\mathcal\{N\}\_\{k\}\\,e^\{\-\\Delta t\\,\\mathcal\{L\}^\{\(k\)\}\}\\Bigg\)=limN→∞∫dzs∫∏k=1s−1d​z​\(tk\)​∏k=1s𝒩k​e−�​t​ℒ\(k\)⏟⟶K\(zi,ti∣zs,ts\)∫∏k=s\+1N−1d​z​\(tk\)​∏k=s\+1N𝒩k​e−�​t​ℒ\(k\)⏟⟶K\(zs,ts∣zf,tf\)\\displaystyle=\\lim\_\{N\\to\\infty\}\\int\\mathrm\{d\}z\_\{s\}\\;\\underbrace\{\\int\\\!\\prod\_\{k=1\}^\{s\-1\}\\\!\\mathrm\{d\}z\(t\_\{k\}\)\\;\\prod\_\{k=1\}^\{s\}\\mathcal\{N\}\_\{k\}\\,e^\{\-\\Delta t\\,\\mathcal\{L\}^\{\(k\)\}\}\}\_\{\\displaystyle\\longrightarrow\\;K\(z\_\{i\},t\_\{i\}\\mid z\_\{s\},t\_\{s\}\)\}\\;\\underbrace\{\\int\\\!\\prod\_\{k=s\+1\}^\{N\-1\}\\\!\\mathrm\{d\}z\(t\_\{k\}\)\\;\\prod\_\{k=s\+1\}^\{N\}\\mathcal\{N\}\_\{k\}\\,e^\{\-\\Delta t\\,\\mathcal\{L\}^\{\(k\)\}\}\}\_\{\\displaystyle\\longrightarrow\\;K\(z\_\{s\},t\_\{s\}\\mid z\_\{f\},t\_\{f\}\)\}=∫dzsK\(zi,ti∣zs,ts\)K\(zs,ts∣zf,tf\)forti<ts<tf,\\displaystyle=\\int\\mathrm\{d\}z\_\{s\}\\;K\(z\_\{i\},t\_\{i\}\\mid z\_\{s\},t\_\{s\}\)\\,K\(z\_\{s\},t\_\{s\}\\mid z\_\{f\},t\_\{f\}\)\\qquad\\text\{for\}\\qquad t\_\{i\}<t\_\{s\}<t\_\{f\}\\;,\(2\.21\)giving the Chapman–Kolmogorov equation\[[24](https://arxiv.org/html/2608.12438#bib.bib22),[21](https://arxiv.org/html/2608.12438#bib.bib23)\], the continuum remnant of the Markov property of the discrete chain\. As the transition density of the reverse diffusion, the kernel satisfies the associated Fokker–Planck equation in its first arguments\[[64](https://arxiv.org/html/2608.12438#bib.bib56),[23](https://arxiv.org/html/2608.12438#bib.bib57)\],

−∂tK\(z,t∣z′,t′\)\\displaystyle\-\\partial\_\{t\}K\(z,t\\mid z^\{\\prime\},t^\{\\prime\}\)=∇z⋅\(frev\(z,t\)K\(z,t∣z′,t′\)\)\+12g2\(t\)∇z2K\(z,t∣z′,t′\),\\displaystyle=\\nabla\_\{z\}\\cdot\\Big\(f\_\{\\mathrm\{rev\}\}\(z,t\)\\,K\(z,t\\mid z^\{\\prime\},t^\{\\prime\}\)\\Big\)\+\\tfrac\{1\}\{2\}\\,g^\{2\}\(t\)\\,\\nabla\_\{z\}^\{2\}K\(z,t\\mid z^\{\\prime\},t^\{\\prime\}\)\\;,\(2\.22\)K\(z,t′∣z′,t′\)\\displaystyle K\(z,t^\{\\prime\}\\mid z^\{\\prime\},t^\{\\prime\}\)=�​\(z−z′\)\.\\displaystyle=\\delta\(z\-z^\{\\prime\}\)\\;\.\(2\.23\)The first equation is written with respect to the sampling direction and is consistent with the forward equation[2\.9](https://arxiv.org/html/2608.12438#S2.E9)\. The delta\-function boundary condition states that, before any evolution has taken place, a process initialized atz′z^\{\\prime\}is localized atz′z^\{\\prime\}\. It is also the vanishing\-time limit of the one\-step kernel in[Eq\.2\.12](https://arxiv.org/html/2608.12438#S2.E12)\. Integrating[Eq\.2\.22](https://arxiv.org/html/2608.12438#S2.E22)overzz, the right\-hand side is a total divergence and vanishes for sufficiently rapid decay at infinity, such that

∂t∫dzK\(z,t∣z′,t′\)=0\.\\displaystyle\\partial\_\{t\}\\int\\mathrm\{d\}z\\;K\(z,t\\mid z^\{\\prime\},t^\{\\prime\}\)=0\\;\.\(2\.24\)The normalization is therefore independent oftt\. The boundary condition in[Eq\.2\.23](https://arxiv.org/html/2608.12438#S2.E23)gives

∫dzK\(z,t′∣z′,t′\)=∫dz�\(z−z′\)=1,\\displaystyle\\int\\mathrm\{d\}z\\;K\(z,t^\{\\prime\}\\mid z^\{\\prime\},t^\{\\prime\}\)=\\int\\mathrm\{d\}z\\;\\delta\(z\-z^\{\\prime\}\)=1\\;,\(2\.25\)and consequently

∫dzK\(z,t∣z′,t′\)=1fort≤t′\.\\displaystyle\\int\\mathrm\{d\}z\\;K\(z,t\\mid z^\{\\prime\},t^\{\\prime\}\)=1\\qquad\\text\{for\}\\qquad t\\leq t^\{\\prime\}\\;\.\(2\.26\)Thus, the reverse process initialized atz′z^\{\\prime\}at timet′t^\{\\prime\}reaches some state at the earlier timettwith unit probability\. This is the continuum counterpart of the normalization of every Gaussian factor in[Eq\.2\.19](https://arxiv.org/html/2608.12438#S2.E19)\. Reinstating the prior weight and integrating over the prior endpoint gives the data density,

pdata\(z0\)=∫dz1pprior\(z1\)K\(z0,0∣z1,1\),\\displaystyle p\_\{\\rm data\}\(z\_\{0\}\)=\\int\\mathrm\{d\}z\_\{1\}\\;p\_\{\\rm prior\}\(z\_\{1\}\)\\,K\(z\_\{0\},0\\mid z\_\{1\},1\)\\;,\(2\.27\)which is the latent\-variable representation we started from in[Eq\.2\.1](https://arxiv.org/html/2608.12438#S2.E1), now with the conditional likelihood realized as a pinned path integral over the intervening trajectory,i\.e\.

K\(z0,0∣z1,1\)≡p\(z0\|z1\)\.\\displaystyle K\(z\_\{0\},0\\mid z\_\{1\},1\)\\equiv p\(z\_\{0\}\|z\_\{1\}\)\\;\.\(2\.28\)Writing the kernel out, the data density takes the master path\-integral form

pdata\(z0\)=∫dz1pprior\(z1\)∫𝒟zexp\[−∫01dt‖z˙​\(t\)−frev​\(z⁡\(t\),t\)‖22​g2​\(t\)\],\\displaystyle p\_\{\\rm data\}\(z\_\{0\}\)=\\int\\mathrm\{d\}z\_\{1\}\\;p\_\{\\rm prior\}\(z\_\{1\}\)\\int\\mathcal\{D\}z\\;\\exp\\\!\\left\[\-\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\frac\{\\\|\\dot\{z\}\(t\)\-f\_\{\\rm rev\}\(z\(t\),t\)\\\|^\{2\}\}\{2g^\{2\}\(t\)\}\\right\]\\;,\(2\.29\)the prior\-weighted sum over all latent trajectories ending at the data point, used repeatedly below\. Different choices of forward driftffand diffusion coefficientg⁡\(t\)g\(t\)give rise to differentfrevf\_\{\\rm rev\}, and hence different generative models, as we show in[Section2\.3](https://arxiv.org/html/2608.12438#S2.SS3)\.

#### Itô versus Stratonovich discretization

We derived the action in[Eq\.2\.17](https://arxiv.org/html/2608.12438#S2.E17)using the Itô convention, which evaluates the drift at the later timez⁡\(tk\)z\(t\_\{k\}\)\. The Stratonovich convention instead uses the midpoint12​\[z⁡\(tk−1\)\+z⁡\(tk\)\]\\tfrac\{1\}\{2\}\[z\(t\_\{k\-1\}\)\+z\(t\_\{k\}\)\]\. Expanding the midpoint drift aboutz⁡\(tk\)z\(t\_\{k\}\)and using⟨�​zk⟩=frev​�​t\\langle\\Delta z\_\{k\}\\rangle=f\_\{\\mathrm\{rev\}\}\\Delta tgenerates an additional contribution of order𝒪⁡\(�​t\)\\mathcal\{O\}\(\\Delta t\)proportional to the divergence of the drift, giving\[[78](https://arxiv.org/html/2608.12438#bib.bib50),[6](https://arxiv.org/html/2608.12438#bib.bib49)\]

SNStrat​\[z\]=SNItô​\[z\]\+12​∫01d​t​∇z⋅frev​\(z,t\)\+𝒪⁡\(�​t1/2\)\.\\displaystyle S\_\{N\}^\{\\text\{Strat\}\}\[z\]=S\_\{N\}^\{\\text\{It\\^\{o\}\}\}\[z\]\+\\frac\{1\}\{2\}\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\nabla\_\{z\}\\cdot f\_\{\\rm rev\}\(z,t\)\+\\mathcal\{O\}\(\\Delta t^\{1/2\}\)\\;\.\(2\.30\)The additional term is a local functional ofzz\. It can shift the finite\-noise stationarity equations, but it is of orderg0g^\{0\}against the Onsager–Machlup term of orderg−2g^\{\-2\}and therefore leaves the leading zero\-noise trajectory unchanged\. In the Itô discretization used throughout, the causal functional Jacobian is field independent, and we absorb it into the normalization of the path measure, while a Stratonovich or midpoint discretization would instead retain the divergence term as a local contribution to the action\.

### 2\.3Known generative models as limits of the master action

The path\-integral formulation developed above provides a unified description of generative modeling as probability transport in latent space, governed by an effective action\. Different generative architectures correspond to particular limits, parameterizations, or approximations of this master construction\. Normalizing flows, diffusion models, conditional flow matching, Schrödinger bridges, and variational autoencoders then follow as different ways of evaluating or approximating the same path integral, and adversarial models appear in the case where the saddle point admits no likelihood to evaluate\.

#### 2\.3\.1Normalizing flows as deterministic saddle points

Normalizing flows arise as the deterministic saddle\-point limit of the path\-integral formulation, obtained by sendingg⁡\(t\)→0g\(t\)\\to 0\. In this limit the path integral is dominated by the single path on which the action is stationary, its saddle point, in the same way that classical mechanics emerges from the quantum path integral as˜​h→0\\mathord\{\\mathchar 126h\}\\to 0\. The normalizing flow density formula follows directly from the latent\-variable representation

pdata​\(x\)=∫d​z​p​\(x\|z\)​pprior​\(z\)=∫d​z​�​\(x−�⁡\(z\)\)​pprior​\(z\),\\displaystyle p\_\{\\rm data\}\(x\)=\\int\\mathrm\{d\}z\\;p\(x\|z\)\\,p\_\{\\rm prior\}\(z\)=\\int\\mathrm\{d\}z\\;\\delta\(x\-\\Phi\(z\)\)\\,p\_\{\\rm prior\}\(z\)\\;,\(2\.31\)where we have used that the generative kernel is deterministic\. For an invertible differentiable map�\\Phifrom latent to data space,x=�⁡\(z\)x=\\Phi\(z\), the delta function�​\(x−�​\(z\)\)\\delta\(x\-\\Phi\(z\)\)has a unique root atz⋆=�−1​\(x\)z\_\{\\star\}=\\Phi^\{\-1\}\(x\)\. Applying the change\-of\-variables rule for delta functions,

�​\(x−�⁡\(z\)\)=�​\(z−�−1​\(x\)\)\|det∂�⁡\(z\)∂z\|z=�−1​\(x\),\\displaystyle\\delta\(x\-\\Phi\(z\)\)=\\frac\{\\delta\(z\-\\Phi^\{\-1\}\(x\)\)\}\{\\left\|\\det\\dfrac\{\\partial\\Phi\(z\)\}\{\\partial z\}\\right\|\_\{z=\\Phi^\{\-1\}\(x\)\}\}\\;,\(2\.32\)we obtain

pdata​\(x\)=∫d​z​�​\(z−�−1​\(x\)\)​pprior​\(z\)\|det∂�⁡\(z\)∂z\|z=�−1​\(x\)=pprior​\(�−1​\(x\)\)\|det∂�⁡\(z\)∂z\|z=�−1​\(x\)\.\\displaystyle p\_\{\\rm data\}\(x\)=\\int\\mathrm\{d\}z\\;\\delta\(z\-\\Phi^\{\-1\}\(x\)\)\\frac\{p\_\{\\rm prior\}\(z\)\}\{\\left\|\\det\\dfrac\{\\partial\\Phi\(z\)\}\{\\partial z\}\\right\|\_\{z=\\Phi^\{\-1\}\(x\)\}\}=\\frac\{p\_\{\\rm prior\}\\\!\\left\(\\Phi^\{\-1\}\(x\)\\right\)\}\{\\left\|\\det\\dfrac\{\\partial\\Phi\(z\)\}\{\\partial z\}\\right\|\_\{z=\\Phi^\{\-1\}\(x\)\}\}\\;\.\(2\.33\)Using the inverse function theorem

∂�−1​\(x\)∂x=\(∂�⁡\(z\)∂z\)−1\|z=�−1​\(x\),\\displaystyle\\dfrac\{\\partial\\Phi^\{\-1\}\(x\)\}\{\\partial x\}=\\left\.\\left\(\\dfrac\{\\partial\\Phi\(z\)\}\{\\partial z\}\\right\)^\{\-1\}\\right\|\_\{z=\\Phi^\{\-1\}\(x\)\}\\;,\(2\.34\)the density takes the equivalent form

pdata​\(x\)=pprior​\(�−1​\(x\)\)​\|det∂�−1​\(x\)∂x\|,\\displaystyle p\_\{\\rm data\}\(x\)=p\_\{\\rm prior\}\\\!\\left\(\\Phi^\{\-1\}\(x\)\\right\)\\,\\left\|\\det\\frac\{\\partial\\Phi^\{\-1\}\(x\)\}\{\\partial x\}\\right\|\\;,\(2\.35\)which is exactly the standard change\-of\-variables formula for normalizing flows\.

#### Functional delta distribution

The path\-integral derivation recovers this result as theg→0g\\to 0saddle\-point limit of the Onsager–Machlup action, and further generalizes it to the continuous\-time setting of continuous normalizing flows \(CNFs\)\. We introduce the equation\-of\-motion residual

ℰ⁡\[z\]​\(t\)≡z˙​\(t\)−frev​\(z⁡\(t\),t\),\\displaystyle\\mathcal\{E\}\[z\]\(t\)\\equiv\\dot\{z\}\(t\)\-f\_\{\\rm rev\}\(z\(t\),t\)\\;,\(2\.36\)so that the Onsager–Machlup weight of[Eq\.2\.17](https://arxiv.org/html/2608.12438#S2.E17)reads

e−SOM​\[z\]=exp\[−∫01dt‖ℰ​\[z\]​\(t\)‖22​g2​\(t\)\]\.\\displaystyle e^\{\-S\_\{\\text\{OM\}\}\[z\]\}=\\exp\\\!\\left\[\-\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\frac\{\\\|\\mathcal\{E\}\[z\]\(t\)\\\|^\{2\}\}\{2g^\{2\}\(t\)\}\\right\]\\;\.\(2\.37\)Asg⁡\(t\)→0g\(t\)\\to 0, this weight is exponentially suppressed wheneverℰ​\[z\]​\(t\)≠0\\mathcal\{E\}\[z\]\(t\)\\neq 0for anytt\. For a fixed pathzzwith‖ℰ⁡\[z\]‖2=�\>0\\\|\\mathcal\{E\}\[z\]\\\|^\{2\}=\\epsilon\>0, it vanishes asexp\(−�/2g2\)→0\\exp\(\-\\epsilon/2g^\{2\}\)\\to 0\. The path integral therefore concentrates on the unique deterministic path satisfying the saddle condition

z˙⋆​\(t\)=frev​\(z⋆​\(t\),t\)withz⋆​\(1\)=z1,\\displaystyle\\dot\{z\}\_\{\\star\}\(t\)=f\_\{\\rm rev\}\(z\_\{\\star\}\(t\),t\)\\qquad\\text\{with\}\\qquad z\_\{\\star\}\(1\)=z\_\{1\}\\;,\(2\.38\)the generative trajectory launched from the prior samplez1z\_\{1\}att=1t=1and arriving at the data pointz⋆​\(0\)z\_\{\\star\}\(0\)att=0t=0\. Asg→0g\\to 0the score contribution to the reverse drift vanishes,frev→ff\_\{\\rm rev\}\\to f, so the saddle path is also governed by the ODEz˙⋆=f⁡\(z⋆,t\)\\dot\{z\}\_\{\\star\}=f\(z\_\{\\star\},t\)and depends deterministically on its endpoints\. Integrating it defines the generative flow map and its inverse,

�:z1⟼z⋆​\(0\)withz⋆​\(1\)=z1,�−1:z0⟼z⋆​\(1\)withz⋆​\(0\)=z0,\\displaystyle\\Phi:\\;z\_\{1\}\\longmapsto z\_\{\\star\}\(0\)\\quad\\text\{with\}\\quad z\_\{\\star\}\(1\)=z\_\{1\}\\;,\\qquad\\Phi^\{\-1\}:\\;z\_\{0\}\\longmapsto z\_\{\\star\}\(1\)\\quad\\text\{with\}\\quad z\_\{\\star\}\(0\)=z\_\{0\}\\;,\(2\.39\)The notation is deliberate\. Integrating the drift from the prior time to the data time realizes the invertible map�\\Phiof[Eq\.2\.33](https://arxiv.org/html/2608.12438#S2.E33), so the discrete and continuous constructions share one symbol because they construct one object\. The required uniqueness and invertibility are guaranteed under the standard Picard–Lindelöf assumptions, namely continuity inttand local Lipschitz continuity inzz, together with conditions ensuring existence on the full intervalt∈\[0,1\]t\\in\[0,1\]\[[16](https://arxiv.org/html/2608.12438#bib.bib36),[5](https://arxiv.org/html/2608.12438#bib.bib37),[27](https://arxiv.org/html/2608.12438#bib.bib38)\]\.

Since the normalization required for the Gaussian\-to\-delta limit is already carried by the continuum measure𝒟​z\\mathcal\{D\}z, the deterministic limit turns the weighted measure into a functional delta, in the sense of distributions,

𝒟zexp\[−∫01dt‖ℰ​\[z\]​\(t\)‖22​g2​\(t\)\]→g→0𝒟z�\[ℰ\[z\]\],\\displaystyle\\mathcal\{D\}z\\;\\exp\\\!\\left\[\-\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\frac\{\\\|\\mathcal\{E\}\[z\]\(t\)\\\|^\{2\}\}\{2g^\{2\}\(t\)\}\\right\]\\;\\xrightarrow\{\\;g\\to 0\\;\}\\;\\mathcal\{D\}z\\;\\delta\[\\mathcal\{E\}\[z\]\]\\;,\(2\.40\)enforcing the deterministic flow equation\. To evaluate the pinned path integral we return to the discrete chain from[Eq\.2\.19](https://arxiv.org/html/2608.12438#S2.E19), where the limit acts factor by factor\. Each Gaussian one\-step kernel in[Eq\.2\.12](https://arxiv.org/html/2608.12438#S2.E12)degenerates to a delta enforcing one Euler step of the flow,

p⁡\(z⁡\(tk−1\)∣z⁡\(tk\)\)→g→0�​\(z⁡\(tk−1\)−z⁡\(tk\)\+f⁡\(z⁡\(tk\),tk\)​�​t\)\.\\displaystyle p\\big\(z\(t\_\{k\-1\}\)\\mid z\(t\_\{k\}\)\\big\)\\;\\xrightarrow\{\\;g\\to 0\\;\}\\;\\delta\\big\(z\(t\_\{k\-1\}\)\-z\(t\_\{k\}\)\+f\(z\(t\_\{k\}\),t\_\{k\}\)\\,\\Delta t\\big\)\\;\.\(2\.41\)Each interior integration removes one delta with unit Jacobian, as the constraint is linear in the earlier statez⁡\(tk−1\)z\(t\_\{k\-1\}\), fixing the interior states successively along the classical trajectory from the prior end downward,z⁡\(tk\)=z⋆​\(tk\)z\(t\_\{k\}\)=z\_\{\\star\}\(t\_\{k\}\)\. After all interior integrations a single constraint remains, pinning the data endpoint to the arrival point of the flow,

K\(z0,0∣z1,1\)→g→0�\(z0−�\(z1\)\)\.\\displaystyle K\(z\_\{0\},0\\mid z\_\{1\},1\)\\;\\xrightarrow\{\\;g\\to 0\\;\}\\;\\delta\\big\(z\_\{0\}\-\\Phi\(z\_\{1\}\)\\big\)\\;\.\(2\.42\)In the deterministic limit the kernel degenerates to deterministic transport of the prior point along the flow\. Inserting[Eq\.2\.42](https://arxiv.org/html/2608.12438#S2.E42)into[Eq\.2\.27](https://arxiv.org/html/2608.12438#S2.E27)and resolving the ordinary delta with the same change\-of\-variables rule as in[Eq\.2\.33](https://arxiv.org/html/2608.12438#S2.E33), we obtain

pdata​\(z0\)\\displaystyle p\_\{\\rm data\}\(z\_\{0\}\)=∫d​z1​pprior​\(z1\)​�​\(z0−�⁡\(z1\)\)\\displaystyle=\\int\\mathrm\{d\}z\_\{1\}\\;p\_\{\\rm prior\}\(z\_\{1\}\)\\,\\delta\\big\(z\_\{0\}\-\\Phi\(z\_\{1\}\)\\big\)=\([2\.32](https://arxiv.org/html/2608.12438#S2.E32)\)​∫d​z1​pprior​\(z1\)​�​\(z1−�−1​\(z0\)\)\|det∂�⁡\(z1\)∂z1\|z1=�−1​\(z0\)\\displaystyle\\overset\{\\eqref\{eq:change\_of\_variable\_delta\}\}\{=\}\\int\\mathrm\{d\}z\_\{1\}\\;p\_\{\\rm prior\}\(z\_\{1\}\)\\,\\frac\{\\delta\\big\(z\_\{1\}\-\\Phi^\{\-1\}\(z\_\{0\}\)\\big\)\}\{\\left\|\\det\\dfrac\{\\partial\\Phi\(z\_\{1\}\)\}\{\\partial z\_\{1\}\}\\right\|\_\{z\_\{1\}=\\Phi^\{\-1\}\(z\_\{0\}\)\}\}=\([2\.34](https://arxiv.org/html/2608.12438#S2.E34)\)​pprior​\(�−1​\(z0\)\)​\|det∂�−1​\(z0\)∂z0\|\.\\displaystyle\\overset\{\\eqref\{eq:ift\}\}\{=\}p\_\{\\rm prior\}\\big\(\\Phi^\{\-1\}\(z\_\{0\}\)\\big\)\\,\\left\|\\det\\frac\{\\partial\\Phi^\{\-1\}\(z\_\{0\}\)\}\{\\partial z\_\{0\}\}\\right\|\\;\.\(2\.43\)This is the normalizing\-flow density formula, now with the continuous flow map�\\Phiplaying the role of the discrete invertible transformff\. Considering the trajectory as a function of its data endpointz⋆​\(0\)=z0z\_\{\\star\}\(0\)=z\_\{0\}and differentiating the flow equation[2\.38](https://arxiv.org/html/2608.12438#S2.E38)with respect toz0z\_\{0\}, we obtain the variational equation

dd​t​M​\(t\)=∂zf⁡\(z⋆​\(t\),t\)​M​\(t\)withM⁡\(t\)=∂z⋆​\(t\)∂z0andM⁡\(0\)=I\.\\displaystyle\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}M\(t\)=\\partial\_\{z\}f\(z\_\{\\star\}\(t\),t\)\\,M\(t\)\\qquad\\text\{with\}\\qquad M\(t\)=\\frac\{\\partial z\_\{\\star\}\(t\)\}\{\\partial z\_\{0\}\}\\qquad\\text\{and\}\\qquad M\(0\)=\\mdmathbb\{I\}\\;\.\(2\.44\)The Jacobian entering[Eq\.2\.43](https://arxiv.org/html/2608.12438#S2.E43)isM⁡\(1\)=∂z⋆​\(1\)/∂z0=∂�−1​\(z0\)/∂z0M\(1\)=\\partial z\_\{\\star\}\(1\)/\\partial z\_\{0\}=\\partial\\Phi^\{\-1\}\(z\_\{0\}\)/\\partial z\_\{0\}\. Using Jacobi’s formula,

dd​t​log⁡\|detM⁡\(t\)\|=tr⁡\(d​M​\(t\)d​t​M​\(t\)−1\),\\displaystyle\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\log\|\\det M\(t\)\|=\\mathrm\{tr\}\\left\(\\frac\{\\mathrm\{d\}M\(t\)\}\{\\mathrm\{d\}t\}\\,M\(t\)^\{\-1\}\\right\)\\;,\(2\.45\)and inserting[Eq\.2\.44](https://arxiv.org/html/2608.12438#S2.E44), we thus obtain

dd​t​log⁡\|detM⁡\(t\)\|=tr⁡\(∂zf⁡\(z⋆​\(t\),t\)\)=∇z⋅f⁡\(z⋆​\(t\),t\)\.\\displaystyle\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\log\\bigl\|\\det M\(t\)\\bigr\|=\\mathrm\{tr\}\\left\(\\partial\_\{z\}f\(z\_\{\\star\}\(t\),t\)\\right\)=\\nabla\_\{z\}\\cdot f\(z\_\{\\star\}\(t\),t\)\\;\.\(2\.46\)Integrating fromt=0t=0tot=1t=1withlog⁡\|detM⁡\(0\)\|=0\\log\|\\det M\(0\)\|=0gives

log⁡\|detM⁡\(1\)\|=∫01d​t​∇z⋅f⁡\(z⋆​\(t\),t\),\\displaystyle\\log\\big\|\\det M\(1\)\\big\|=\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\nabla\_\{z\}\\cdot f\(z\_\{\\star\}\(t\),t\)\\;,\(2\.47\)and substituting into[Eq\.2\.43](https://arxiv.org/html/2608.12438#S2.E43)yields the continuous normalizing\-flow log\-likelihood,

log⁡pdata​\(z0\)=log⁡pprior​\(�−1​\(z0\)\)\+∫01d​t​∇z⋅f⁡\(z⋆​\(t\),t\),\\displaystyle\\log p\_\{\\rm data\}\(z\_\{0\}\)=\\log p\_\{\\rm prior\}\\big\(\\Phi^\{\-1\}\(z\_\{0\}\)\\big\)\+\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\nabla\_\{z\}\\cdot f\(z\_\{\\star\}\(t\),t\)\\;,\(2\.48\)here derived directly from the saddle\-point evaluation of the Onsager–Machlup path integral\.

#### 2\.3\.2Diffusion models as stochastic probability flows

Diffusion\-based generative models correspond to the regime in which stochasticity plays an essential role and the master path integral cannot be reduced to a single dominant trajectory\. Starting again from the master representation of the data density,

pdata\(z0\)=∫dz1pprior\(z1\)∫𝒟zexp\[−∫01dt‖z˙​\(t\)−frev​\(z⁡\(t\),t\)‖22​g2​\(t\)\],\\displaystyle p\_\{\\rm data\}\(z\_\{0\}\)=\\int\\mathrm\{d\}z\_\{1\}\\;p\_\{\\rm prior\}\(z\_\{1\}\)\\int\\mathcal\{D\}z\\;\\exp\\\!\\left\[\-\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\frac\{\\\|\\dot\{z\}\(t\)\-f\_\{\\mathrm\{rev\}\}\(z\(t\),t\)\\\|^\{2\}\}\{2g^\{2\}\(t\)\}\\right\]\\;,\(2\.49\)we observe that forg⁡\(t\)≠0g\(t\)\\neq 0the functional weight does not collapse onto a single trajectory\. Generative sampling now corresponds to drawing entire latent pathsz⁡\(t\)z\(t\)from the path measure defined by the Onsager–Machlup action\. While[Eq\.2\.49](https://arxiv.org/html/2608.12438#S2.E49)provides a closed\-form expression forpdata​\(z0\)p\_\{\\rm data\}\(z\_\{0\}\), it is intractable in practice\. Evaluating it requires integrating over all trajectories*and*knowing the reverse driftfrev​\(z,t\)f\_\{\\mathrm\{rev\}\}\(z,t\), which itself depends on the unknown score∇z​log​p​\(z,t\)\\nabla\_\{z\}\\log p\(z,t\)\. Diffusion models are therefore not trained on the data density directly\. They rely on alternative evaluation principles that bypass the explicit computation of the path integral\.

#### Score matching

A common strategy is to specify a forward noising process with prescribed driftf⁡\(z,t\)f\(z,t\)and diffusiong⁡\(t\)g\(t\),

d​z=f⁡\(z,t\)​d​t\+g⁡\(t\)​d​Wt,\\displaystyle\\mathrm\{d\}z=f\(z,t\)\\,\\mathrm\{d\}t\+g\(t\)\\,\\mathrm\{d\}W\_\{t\}\\;,\(2\.50\)initialized atz0∼pdata​\(z0\)z\_\{0\}\\sim p\_\{\\rm data\}\(z\_\{0\}\)\. This forward process defines conditional densitiesp⁡\(z\|z0,t\)p\(z\|z\_\{0\},t\)and marginalsp⁡\(z,t\)p\(z,t\)for alltt, and is chosen such that the final distributionpprior​\(z1\)p\_\{\\rm prior\}\(z\_\{1\}\)is a simple prior, typically Gaussian\. Onceffandggare fixed, these conditionals are fully determined\. For affine forward drifts they are Gaussian in closed form,

p⁡\(z\|z0,t\)=𝒩⁡\(z,�​\(t\)​z0,�2​\(t\)​I\),\\displaystyle p\(z\|z\_\{0\},t\)=\\mathcal\{N\}\\big\(z;\\,\\alpha\(t\)\\,z\_\{0\},\\;\\sigma^\{2\}\(t\)\\,\\mdmathbb\{I\}\\big\)\\;,\(2\.51\)with�\\alphaand�\\sigmaset byffandgg, so that we can draw a noised state in a single step without simulating the SDE\. This one\-shot samplability makes the denoising objective below practical\. The exact generative dynamics are given again by the reverse\-time stochastic differential equation,

d​z=\(f⁡\(z,t\)−g2​\(t\)​∇z​log⁡p⁡\(z,t\)\)​d​t\+g⁡\(t\)​d​W¯t,\\displaystyle\\mathrm\{d\}z=\\Big\(f\(z,t\)\-g^\{2\}\(t\)\\,\\nabla\_\{z\}\\log p\(z,t\)\\Big\)\\,\\mathrm\{d\}t\+g\(t\)\\,\\mathrm\{d\}\\bar\{W\}\_\{t\}\\;,\(2\.52\)restated here explicitly for the learned setting\. The central difficulty is therefore the estimation of the score∇z​log​p​\(z,t\)\\nabla\_\{z\}\\log p\(z,t\)\.

Lets�​\(z,t\)s\_\{\\theta\}\(z,t\)be a parameterized approximation of the score, for instance a neural network with tunable parameters�\\theta\. Replacing∇z​log​p​\(z,t\)\\nabla\_\{z\}\\log p\(z,t\)bys�​\(z,t\)s\_\{\\theta\}\(z,t\)defines an approximate reverse drift, and matching the resulting transition kernel to the true one yields a natural training objective\. Both kernels carry the Gaussian form of[Eq\.2\.12](https://arxiv.org/html/2608.12438#S2.E12)with the*same*covarianceg2​\(tk\)​�​t​Ig^\{2\}\(t\_\{k\}\)\\,\\Delta t\\,\\mdmathbb\{I\}, so their Kullback–Leibler divergence per step is the squared mean difference in units of that covariance,

�​�k\\displaystyle\\Delta\\mu\_\{k\}≡g2​\(tk\)​�​t​\[s�​\(z,tk\)−∇z​log​p​\(z,tk\)\],\\displaystyle\\equiv g^\{2\}\(t\_\{k\}\)\\,\\Delta t\\,\\big\[s\_\{\\theta\}\(z,t\_\{k\}\)\-\\nabla\_\{z\}\\log p\(z,t\_\{k\}\)\\big\]\\;,KLk\\displaystyle\\mathrm\{KL\}\_\{k\}=‖�​�k‖22​g2​\(tk\)​�​t=12​g2​\(tk\)​�​t​‖s�​\(z,tk\)−∇z​log​p​\(z,tk\)‖2\.\\displaystyle=\\frac\{\\\|\\Delta\\mu\_\{k\}\\\|^\{2\}\}\{2\\,g^\{2\}\(t\_\{k\}\)\\,\\Delta t\}=\\tfrac\{1\}\{2\}\\,g^\{2\}\(t\_\{k\}\)\\,\\Delta t\\,\\big\\\|s\_\{\\theta\}\(z,t\_\{k\}\)\-\\nabla\_\{z\}\\log p\(z,t\_\{k\}\)\\big\\\|^\{2\}\\;\.\(2\.53\)Summing the steps into a Riemann sum and averaging over states yields, up tott\-dependent constants, the objective

ℒscore​\(�\)∝Et∼�​\(t\)​Ez∼p⁡\(⋅,t\)​\[g2​\(t\)​‖s�​\(z,t\)−∇z​log​p​\(z,t\)‖2\],\\displaystyle\\mathcal\{L\}\_\{\\rm score\}\(\\theta\)\\propto\\mdmathbb E\_\{t\\sim\\pi\(t\)\}\\,\\mdmathbb E\_\{z\\sim p\(\\cdot,t\)\}\\Big\[g^\{2\}\(t\)\\,\\big\\\|s\_\{\\theta\}\(z,t\)\-\\nabla\_\{z\}\\log p\(z,t\)\\big\\\|^\{2\}\\Big\]\\;,\(2\.54\)where�​\(t\)\\pi\(t\)is a uniform sampling distribution over timest∈\[0,1\]t\\in\[0,1\], and the short\-time computation in[Eq\.2\.53](https://arxiv.org/html/2608.12438#S2.E53)fixes the weightg2​\(t\)g^\{2\}\(t\)\. Since the marginal score∇z​log​p​\(z,t\)\\nabla\_\{z\}\\log p\(z,t\)is unknown, this loss is rewritten using the forward conditional distribution\. For diffusion processes, the marginal density at timettreads

p⁡\(z,t\)=∫d​z0​p​\(z\|z0,t\)​pdata​\(z0\)\.\\displaystyle p\(z,t\)=\\int\\mathrm\{d\}z\_\{0\}\\;p\(z\|z\_\{0\},t\)\\,p\_\{\\rm data\}\(z\_\{0\}\)\\;\.\(2\.55\)Differentiating with respect tozzand using Bayes’ rule gives the pointwise identity

∇zlogp\(z,t\)=Ez0∼p\(⋅\|z,t\)\[∇zlogp\(z\|z0,t\)\]\.\\displaystyle\\nabla\_\{z\}\\log p\(z,t\)=\\mdmathbb E\_\{z\_\{0\}\\sim p\(\\cdot\|z,t\)\}\\big\[\\nabla\_\{z\}\\log p\(z\|z\_\{0\},t\)\\big\]\\;\.\(2\.56\)Equivalently, once averaged overz∼p⁡\(⋅,t\)z\\sim p\(\\cdot,t\), this posterior expectation becomes a joint expectation over the forward noising process,

Ez∼p⁡\(⋅,t\)Ez0∼p\(⋅\|z,t\)\[F\(z0,z,t\)\]=Ez0∼pdataEz∼p\(⋅\|z0,t\)\[F\(z0,z,t\)\],\\displaystyle\\mdmathbb E\_\{z\\sim p\(\\cdot,t\)\}\\mdmathbb E\_\{z\_\{0\}\\sim p\(\\cdot\|z,t\)\}\\Big\[F\(z\_\{0\},z,t\)\\Big\]=\\mdmathbb E\_\{z\_\{0\}\\sim p\_\{\\rm data\}\}\\mdmathbb E\_\{z\\sim p\(\\cdot\|z\_\{0\},t\)\}\\Big\[F\(z\_\{0\},z,t\)\\Big\]\\;,\(2\.57\)for any suitable test functionFF\. This follows from the two equivalent factorizations of the same joint distribution,

p⁡\(z,t\)​p​\(z0\|z,t\)=p⁡\(z\|z0,t\)​pdata​\(z0\)\.\\displaystyle p\(z,t\)\\,p\(z\_\{0\}\|z,t\)=p\(z\|z\_\{0\},t\)\\,p\_\{\\rm data\}\(z\_\{0\}\)\\;\.\(2\.58\)Thus, we never sample the posteriorp⁡\(z0\|z,t\)p\(z\_\{0\}\|z,t\)explicitly\. Drawing firstz0∼pdataz\_\{0\}\\sim p\_\{\\rm data\}and thenz∼p\(⋅\|z0,t\)z\\sim p\(\\cdot\|z\_\{0\},t\)evaluates the same joint expectation\. Using this relation, we replace the intractable objective from[Eq\.2\.54](https://arxiv.org/html/2608.12438#S2.E54)by the denoising score\-matching \(DSM\) loss

ℒDSM\(�\)=Et∼�​\(t\)Ez0∼pdataEz∼p\(⋅\|z0,t\)\[g2\(t\)∥s�\(z,t\)−∇zlogp\(z\|z0,t\)∥2\]\.\\displaystyle\\mathcal\{L\}\_\{\\rm DSM\}\(\\theta\)=\\mdmathbb E\_\{t\\sim\\pi\(t\)\}\\,\\mdmathbb E\_\{z\_\{0\}\\sim p\_\{\\rm data\}\}\\,\\mdmathbb E\_\{z\\sim p\(\\cdot\|z\_\{0\},t\)\}\\Big\[g^\{2\}\(t\)\\,\\big\\\|s\_\{\\theta\}\(z,t\)\-\\nabla\_\{z\}\\log p\(z\|z\_\{0\},t\)\\big\\\|^\{2\}\\Big\]\\;\.\(2\.59\)For this choice of weighting,[Eq\.2\.59](https://arxiv.org/html/2608.12438#S2.E59)differs from[Eq\.2\.54](https://arxiv.org/html/2608.12438#S2.E54)only by a�\\theta\-independent constant, and therefore has the same optimum\[[73](https://arxiv.org/html/2608.12438#bib.bib48),[13](https://arxiv.org/html/2608.12438#bib.bib1)\]\. More generally, practical implementations often replaceg2​\(t\)g^\{2\}\(t\)by a positive weighting function�​\(t\)\\lambda\(t\),

ℒDSM\(�\)\(�\)=Et∼�​\(t\)Ez0∼pdataEz∼p\(⋅\|z0,t\)\[�\(t\)∥s�\(z,t\)−∇zlogp\(z\|z0,t\)∥2\],\\displaystyle\\mathcal\{L\}\_\{\\rm DSM\}^\{\(\\lambda\)\}\(\\theta\)=\\mdmathbb E\_\{t\\sim\\pi\(t\)\}\\,\\mdmathbb E\_\{z\_\{0\}\\sim p\_\{\\rm data\}\}\\,\\mdmathbb E\_\{z\\sim p\(\\cdot\|z\_\{0\},t\)\}\\Big\[\\lambda\(t\)\\,\\big\\\|s\_\{\\theta\}\(z,t\)\-\\nabla\_\{z\}\\log p\(z\|z\_\{0\},t\)\\big\\\|^\{2\}\\Big\]\\;,\(2\.60\)which changes the relative emphasis placed on different noise levels but not the pointwise target score\. In the idealized infinite\-capacity limit, any strictly positive�​\(t\)\\lambda\(t\)therefore has the same score\-matching optimum, while for finite networks and finite data the choice of�​\(t\)\\lambda\(t\)affects the practical training dynamics and accuracy across time\. The choice�=g2\\lambda=g^\{2\}is the*likelihood weighting*, for which the objective bounds the negative log\-likelihood\[[68](https://arxiv.org/html/2608.12438#bib.bib14),[33](https://arxiv.org/html/2608.12438#bib.bib44)\], while more general weightings admit a variational reading as a weighted integral of evidence lower bounds\[[39](https://arxiv.org/html/2608.12438#bib.bib15)\]\.

After training, generative sampling proceeds by integrating the learned reverse\-time SDE obtained from[Eq\.2\.52](https://arxiv.org/html/2608.12438#S2.E52)with the score replaced bys�s\_\{\\theta\}\. Exact likelihood evaluation via the stochastic path integral remains intractable in general\. With an exact score, integrating the probability\-flow ODE together with a log\-density evolution evaluates the model likelihood\. With a learned score, the result is the likelihood of the learned deterministic flow\. In the language of[Section2\.2](https://arxiv.org/html/2608.12438#S2.SS2), score matching is thus the estimation of the drift entering the master kernel\. The trained score definesfrev�=f−g2​s�f^\{\\theta\}\_\{\\mathrm\{rev\}\}=f\-g^\{2\}s\_\{\\theta\}, and inserting it into[Eq\.2\.19](https://arxiv.org/html/2608.12438#S2.E19)specifies the path integral of the learned model, which stochastic sampling evaluates trajectory by trajectory\. Discrete\-time diffusion models such as denoising diffusion probabilistic models \(DDPMs\) correspond to specific time discretizations of this construction, combined with particular parameterizations of the score function and the reverse\-time dynamics\.

#### Conditional flow matching

Conditional flow matching \(CFM\) provides a deterministic alternative to stochastic diffusion models that nevertheless avoids direct likelihood evaluation\. In the present framework, CFM replaces stochastic reverse\-time sampling by a*probability\-flow ordinary differential equation*that preserves the same time\-marginal distributions as an underlying diffusion process\. Given a diffusion process with driftf⁡\(z,t\)f\(z,t\)and diffusiong⁡\(t\)g\(t\), the probability\-flow ODE follows from rewriting the Fokker–Planck equation[2\.9](https://arxiv.org/html/2608.12438#S2.E9)as a continuity equation\. Using∇zp=p​∇z​log⁡p\\nabla\_\{z\}p=p\\,\\nabla\_\{z\}\\log p, we recast the diffusion term as a transport term,

12​g2​\(t\)​∇z2p=∇z⋅\(12​g2​\(t\)​p​∇z​log⁡p\),\\displaystyle\\tfrac\{1\}\{2\}g^\{2\}\(t\)\\,\\nabla\_\{z\}^\{2\}p=\\nabla\_\{z\}\\cdot\\big\(\\tfrac\{1\}\{2\}g^\{2\}\(t\)\\,p\\,\\nabla\_\{z\}\\log p\\big\)\\;,\(2\.61\)so that[Eq\.2\.9](https://arxiv.org/html/2608.12438#S2.E9)becomes a pure continuity equation,

∂tp\(z,t\)=−∇z⋅\(v\(z,t\)p\(z,t\)\),\\displaystyle\\partial\_\{t\}p\(z,t\)=\-\\nabla\_\{z\}\\\!\\cdot\\big\(v\(z,t\)\\,p\(z,t\)\\big\)\\;,\(2\.62\)with the effective velocity

z˙=v⁡\(z,t\)=f⁡\(z,t\)−12​g2​\(t\)​∇z​log⁡p⁡\(z,t\)\.\\displaystyle\\dot\{z\}=v\(z,t\)=f\(z,t\)\-\\tfrac\{1\}\{2\}g^\{2\}\(t\)\\,\\nabla\_\{z\}\\log p\(z,t\)\\;\.\(2\.63\)Because the marginals obey[Eq\.2\.62](https://arxiv.org/html/2608.12438#S2.E62), the deterministic flow along[Eq\.2\.63](https://arxiv.org/html/2608.12438#S2.E63)reproduces exactly the same marginalsp⁡\(z,t\)p\(z,t\)as the stochastic process, with no noise in the sampling dynamics\. The factor12\\tfrac\{1\}\{2\}originates from rewriting the diffusion term12​g2​∇z2p\\tfrac\{1\}\{2\}g^\{2\}\\nabla\_\{z\}^\{2\}pof[Eq\.2\.9](https://arxiv.org/html/2608.12438#S2.E9)as the transport term in[Eq\.2\.61](https://arxiv.org/html/2608.12438#S2.E61)\. It is therefore*not*the score coefficient of the reverse\-time SDE drift, which carries the fullg2​∇z​log⁡pg^\{2\}\\nabla\_\{z\}\\log p\. The two descriptions generate identical time\-marginalsp⁡\(z,t\)p\(z,t\)but different trajectory ensembles\. The ODE reproduces every single\-time marginal of the stochastic process while suppressing its path\-space fluctuations\. As in diffusion models, the score∇z​log​p​\(z,t\)\\nabla\_\{z\}\\log p\(z,t\)is unknown and must be approximated\.

Rather than estimating the score directly, conditional flow matching constructs tractable training targets by introducing conditional probability flows\. Let

z=​\(t,z0,z1\)withz0∼pdataandz1∼pprior,\\displaystyle z=\\psi\(t;z\_\{0\},z\_\{1\}\)\\qquad\\text\{with\}\\qquad z\_\{0\}\\sim p\_\{\\rm data\}\\qquad\\text\{and\}\\qquad z\_\{1\}\\sim p\_\{\\rm prior\}\\;,\(2\.64\)denote a prescribed interpolation between data samplesz0z\_\{0\}and latent samplesz1z\_\{1\}\. This defines a*conditional*,i\.e\.endpoint\-dependent, target velocity

v⋆\(z,t\|z0,z1\)=∂t\(t;z0,z1\),\\displaystyle v^\{\\star\}\(z,t\|z\_\{0\},z\_\{1\}\)=\\partial\_\{t\}\\psi\(t;z\_\{0\},z\_\{1\}\)\\;,\(2\.65\)which is known analytically by construction\. We then introduce a neural velocity fieldv�​\(z,t\)v\_\{\\theta\}\(z,t\)that only depends on the current state and time, not on the particular endpoints\(z0,z1\)\(z\_\{0\},z\_\{1\}\)used to construct the interpolation\. We train it to approximate the*conditional expectation*of this target given\(z,t\)\(z,t\),

v�\(z,t\)≈E\[v⋆\(z,t\|z0,z1\)∣z,t\]=v\(z,t\),\\displaystyle v\_\{\\theta\}\(z,t\)\\approx\\mdmathbb E\\\!\\left\[v^\{\\star\}\(z,t\|z\_\{0\},z\_\{1\}\)\\mid z,t\\right\]=v\(z,t\)\\;,\(2\.66\)leading to the conditional flow matching objective\[[49](https://arxiv.org/html/2608.12438#bib.bib54)\]

ℒCFM\(�\)=Et,z0,z1\[∥v�\(z,t\)−v⋆\(z,t\|z0,z1\)∥2\]\.\\displaystyle\\mathcal\{L\}\_\{\\rm CFM\}\(\\theta\)=\\mdmathbb E\_\{t,z\_\{0\},z\_\{1\}\}\\Big\[\\big\\\|v\_\{\\theta\}\(z,t\)\-v^\{\\star\}\(z,t\|z\_\{0\},z\_\{1\}\)\\big\\\|^\{2\}\\Big\]\\;\.\(2\.67\)Minimizing this conditional\-target loss is equivalent to minimizing the \(intractable\) marginal\-target loss, even though the marginal velocity itself is never evaluated, by the same conditional\-versus\-marginal replacement that carried the score\-matching objective in[Eq\.2\.54](https://arxiv.org/html/2608.12438#S2.E54)to its denoising form in[Eq\.2\.59](https://arxiv.org/html/2608.12438#S2.E59)\[[49](https://arxiv.org/html/2608.12438#bib.bib54),[50](https://arxiv.org/html/2608.12438#bib.bib35)\]\.

For the Gaussian conditional paths used here, CFM is a deterministic, time\-local evaluation principle for the stochastic path integral, sampling along the probability\-flow ODE in[Eq\.2\.63](https://arxiv.org/html/2608.12438#S2.E63)while regressing on conditional targets that respect the same marginals\. For more general interpolants the conditional paths need not derive from a prescribed diffusion, and CFM independently defines transport through the continuity equation\. Even for deterministic CNFs one may prefer CFM over likelihood\-based training, since the latter requires evaluating a continuous\-time divergence integral that can be computationally expensive in high dimensions\.

#### Schrödinger bridges and stochastic control

Schrödinger bridges \(SBs\) and stochastic\-control approaches provide an alternative evaluation principle\[[12](https://arxiv.org/html/2608.12438#bib.bib52),[67](https://arxiv.org/html/2608.12438#bib.bib43)\]\. Rather than fixing the forward diffusion and learning the reverse drift via score matching, SB methods determine an*optimal controlled diffusion*whose path measure interpolates between prescribed endpoint marginals\. In the language of the master action this is still nothing but a choice of reverse drift inserted into the same path integral\. It arises globally, as the Doobhh\-transform of the reference dynamics\[[19](https://arxiv.org/html/2608.12438#bib.bib20),[48](https://arxiv.org/html/2608.12438#bib.bib47),[15](https://arxiv.org/html/2608.12438#bib.bib21)\]\. The transform is a reweighting of the reference path measure by a space–time harmonic functionh⁡\(z,t\)h\(z,t\), which shifts the reference drift byg2​∇z​log⁡hg^\{2\}\\,\\nabla\_\{z\}\\log h, a correction of the same form as the score term in[Eq\.2\.10](https://arxiv.org/html/2608.12438#S2.E10), withhhfixed by the two endpoint marginals rather than locally by a prescribed forward process\. Concretely, let𝒫0​\[z\]\\mathcal\{P\}\_\{0\}\[z\]be a reference path measure induced by a known diffusion,e\.g\.Brownian motion or an Ornstein–Uhlenbeck process\. Among all path measures𝒫⁡\[z\]\\mathcal\{P\}\[z\]with the prescribed endpoint marginals, the Schrödinger bridge solves

𝒫⋆=argmin𝒫\{KL\(𝒫∥𝒫0\)\|𝒫\(z\(0\)\)=pdata,𝒫\(z\(1\)\)=pprior\}\.\\displaystyle\\mathcal\{P\}^\{\\star\}=\\arg\\min\_\{\\mathcal\{P\}\}\\Big\\\{\\mathrm\{KL\}\\\!\\big\(\\mathcal\{P\}\\,\\\|\\,\\mathcal\{P\}\_\{0\}\\big\)\\;\\big\|\\;\\mathcal\{P\}\(z\(0\)\)=p\_\{\\rm data\},\\;\\mathcal\{P\}\(z\(1\)\)=p\_\{\\rm prior\}\\Big\\\}\\;\.\(2\.68\)This is an entropy\-regularized optimal\-transport problem on path space\[[48](https://arxiv.org/html/2608.12438#bib.bib47)\], in which the SB is the path measure closest to the reference while transportingpdatap\_\{\\rm data\}toppriorp\_\{\\rm prior\}\.

In the path\-integral language,[Eq\.2\.68](https://arxiv.org/html/2608.12438#S2.E68)is a variational principle on the Onsager–Machlup path measure\. Writing the reference measure as𝒫0∝e−S0​\[z\]\\mathcal\{P\}\_\{0\}\\propto e^\{\-S\_\{0\}\[z\]\}, withS0S\_\{0\}its Onsager–Machlup action, the KL divergence is

KL\(𝒫∥𝒫0\)=E𝒫\[S0\[z\]\]−H\[𝒫\]\+const,\\displaystyle\\mathrm\{KL\}\(\\mathcal\{P\}\\,\\\|\\,\\mathcal\{P\}\_\{0\}\)=\\mdmathbb\{E\}\_\{\\mathcal\{P\}\}\\big\[\\,S\_\{0\}\[z\]\\,\\big\]\-\\mdmathbb\{H\}\[\\mathcal\{P\}\]\+\\mathrm\{const\}\\;,\(2\.69\)whereH⁡\[𝒫\]\\mdmathbb\{H\}\[\\mathcal\{P\}\]is the path entropy\. The SB is thus the entropy\-regularized minimum of the Onsager–Machlup action over path measures with the prescribed boundary marginals\. In contrast to score\-based diffusion, which fixes the forward process and matches the reverse drift*locally*, the SB determines the entire probability flow*globally*, subject to both endpoint constraints\.

*Iterative proportional fitting*\(IPF\) solves the constrained problem in[Eq\.2\.68](https://arxiv.org/html/2608.12438#S2.E68)\[[12](https://arxiv.org/html/2608.12438#bib.bib52)\]\. Starting from𝒫\(0\)=𝒫0\\mathcal\{P\}^\{\(0\)\}=\\mathcal\{P\}\_\{0\}, one alternates two half\-bridge KL projections,

𝒫\(2​n\+1\)\\displaystyle\\mathcal\{P\}^\{\(2n\+1\)\}=argmin𝒫KL\(𝒫∥𝒫\(2​n\)\)\\displaystyle=\\arg\\min\_\{\\mathcal\{P\}\}\\;\\mathrm\{KL\}\\\!\\big\(\\mathcal\{P\}\\,\\\|\\,\\mathcal\{P\}^\{\(2n\)\}\\big\)with𝒫⁡\(z⁡\(1\)\)\\displaystyle\\qquad\\text\{with\}\\qquad\\mathcal\{P\}\(z\(1\)\)=pprior,\\displaystyle=p\_\{\\rm prior\}\\;,𝒫\(2​n\+2\)\\displaystyle\\mathcal\{P\}^\{\(2n\+2\)\}=argmin𝒫KL\(𝒫∥𝒫\(2​n\+1\)\)\\displaystyle=\\arg\\min\_\{\\mathcal\{P\}\}\\;\\mathrm\{KL\}\\\!\\big\(\\mathcal\{P\}\\,\\\|\\,\\mathcal\{P\}^\{\(2n\+1\)\}\\big\)with𝒫⁡\(z⁡\(0\)\)\\displaystyle\\qquad\\text\{with\}\\qquad\\mathcal\{P\}\(z\(0\)\)=pdata,\\displaystyle=p\_\{\\rm data\}\\;,\(2\.70\)each step enforcing one endpoint marginal while staying as close as possible to the current iteration\. The iteration converges to𝒫⋆\\mathcal\{P\}^\{\\star\}\[[12](https://arxiv.org/html/2608.12438#bib.bib52)\]\. In practice each half\-bridge is itself a diffusion\. Its optimal drift is a score, and a score\-matching regression of the form of[Eq\.2\.59](https://arxiv.org/html/2608.12438#S2.E59)solves it, so that one IPF step reduces to training a diffusion model against the trajectories of the previous iteration\[[12](https://arxiv.org/html/2608.12438#bib.bib52)\]\. Subsequent formulations recast the same alternation as iterative Markovian fitting\[[67](https://arxiv.org/html/2608.12438#bib.bib43)\], which projects alternately onto the Markov and reciprocal classes of path measures\.

#### 2\.3\.3Variational autoencoders as variational approximations

Variational autoencoders \(VAEs\) operate directly at the level of the latent\-variable model, without introducing the temporal chain in[Eq\.2\.3](https://arxiv.org/html/2608.12438#S2.E3)\. In the language of this paper they correspond not to a particular dynamics but to the absence of one,i\.e\.a single latent variablezzwith priorpprior​\(z\)p\_\{\\rm prior\}\(z\)and decoderp�​\(x\|z\)p\_\{\\theta\}\(x\|z\),

p�​\(x\)=∫d​z​p�​\(x\|z\)​pprior​\(z\),\\displaystyle p\_\{\\theta\}\(x\)=\\int\\mathrm\{d\}z\\;p\_\{\\theta\}\(x\|z\)\\,p\_\{\\rm prior\}\(z\)\\;,\(2\.71\)and no path over intermediate states\. The marginal likelihood is intractable for expressive decoders, since the integral overzzcannot be computed in closed form\. Thus, the true posterior

p�​\(z\|x\)=p�​\(x\|z\)​pprior​\(z\)p�​\(x\),\\displaystyle p\_\{\\theta\}\(z\|x\)=\\frac\{p\_\{\\theta\}\(x\|z\)\\,p\_\{\\rm prior\}\(z\)\}\{p\_\{\\theta\}\(x\)\}\\;,\(2\.72\)is likewise intractable\. Introducing an approximate encoderq�​\(z\|x\)q\_\{\\phi\}\(z\|x\)and writing

log⁡p�​\(x\)\\displaystyle\\log p\_\{\\theta\}\(x\)=log∫dzq�\(z\|x\)p�​\(x\|z\)​pprior​\(z\)q�​\(z\|x\),\\displaystyle=\\log\\int\\mathrm\{d\}z\\;q\_\{\\phi\}\(z\|x\)\\,\\frac\{p\_\{\\theta\}\(x\|z\)\\,p\_\{\\rm prior\}\(z\)\}\{q\_\{\\phi\}\(z\|x\)\}\\;,\(2\.73\)Jensen’s inequality applied to the concave logarithm yields

log⁡p�​\(x\)\\displaystyle\\log p\_\{\\theta\}\(x\)≥Eq�\(⋅\|x\)\[logp�\(x\|z\)\]−KL\(q�\(z\|x\)∥pprior\(z\)\)≡ℒELBO\(x;�,�\)\.\\displaystyle\\geq\\mdmathbb\{E\}\_\{q\_\{\\phi\}\(\\cdot\|x\)\}\\\!\\left\[\\log p\_\{\\theta\}\(x\|z\)\\right\]\-\\mathrm\{KL\}\\big\(q\_\{\\phi\}\(z\|x\)\\,\\\|\\,p\_\{\\rm prior\}\(z\)\\big\)\\equiv\\mathcal\{L\}\_\{\\rm ELBO\}\(x;\\theta,\\phi\)\\;\.\(2\.74\)The gap between the ELBO and the true log\-likelihood is exactly the KL divergence between the approximate and true posteriors,

logp�\(x\)−ℒELBO=KL\(q�\(z\|x\)∥p�\(z\|x\)\)≥0,\\displaystyle\\log p\_\{\\theta\}\(x\)\-\\mathcal\{L\}\_\{\\rm ELBO\}=\\mathrm\{KL\}\\big\(q\_\{\\phi\}\(z\|x\)\\,\\\|\\,p\_\{\\theta\}\(z\|x\)\\big\)\\geq 0\\;,\(2\.75\)which follows directly from the definition of KL divergence\. Maximizing the ELBO simultaneously tightens this bound and minimizes the KL betweenq�q\_\{\\phi\}and the true posterior\. The bound is tight whenq�=p�\(⋅\|x\)q\_\{\\phi\}=p\_\{\\theta\}\(\\cdot\|x\)\. Both conditionals are distributions even though deterministic networks implement them\. Each network returns the parameters of a Gaussian, the decoder the mean of an observation model and the encoder a mean and a variance,

p�​\(x\|z\)=𝒩⁡\(x,��​\(z\),�2​I\)andq�​\(z\|x\)=𝒩⁡\(z,��​\(x\),��2​\(x\)​I\),\\displaystyle p\_\{\\theta\}\(x\|z\)=\\mathcal\{N\}\\big\(x;\\,\\mu\_\{\\theta\}\(z\),\\,\\sigma^\{2\}\\,\\mdmathbb\{I\}\\big\)\\qquad\\text\{and\}\\qquad q\_\{\\phi\}\(z\|x\)=\\mathcal\{N\}\\big\(z;\\,\\mu\_\{\\phi\}\(x\),\\,\\sigma\_\{\\phi\}^\{2\}\(x\)\\,\\mdmathbb\{I\}\\big\)\\;,\(2\.76\)so the reconstruction term of[Eq\.2\.74](https://arxiv.org/html/2608.12438#S2.E74)equals the familiar squared\-error loss with�2\\sigma^\{2\}setting its weight against the KL term\. Neither spread is decorative, and the two fail at opposite ends of[Eq\.2\.74](https://arxiv.org/html/2608.12438#S2.E74)\. A deterministic decoder means�→0\\sigma\\to 0, wherep�​\(x\|z\)→�​\(x−��​\(z\)\)p\_\{\\theta\}\(x\|z\)\\to\\delta\(x\-\\mu\_\{\\theta\}\(z\)\)and the reconstruction term diverges, while a deterministic encoder givesKL\(�∥pprior\)=∞\\mathrm\{KL\}\(\\delta\\,\\\|\\,p\_\{\\rm prior\}\)=\\inftyand drives the bound to−∞\-\\infty\. The decoder noise makes the model a density, while the encoder noise letsq�q\_\{\\phi\}approximate a posterior that has support\.

One may nonetheless view the single\-latent model as the degenerate endpoint of the path construction, in which the trajectory consists of a single transition of unit duration\. This analogy is formal rather than dynamical, since a VAE does not arise from a continuum limit, and we include it only for completeness\. The single transition kernelp⁡\(z0\|z1\)p\(z\_\{0\}\|z\_\{1\}\)plays the role of the short\-time kernel in[Eq\.2\.12](https://arxiv.org/html/2608.12438#S2.E12)integrated over a single step of unit duration\. The discrete action is

S�​\[x,z\]=−log⁡p�​\(x\|z\),\\displaystyle S\_\{\\theta\}\[x,z\]=\-\\log p\_\{\\theta\}\(x\|z\)\\;,\(2\.77\)so the path integral reduces to a single ordinary integral over the latent variable,

p�​\(x\)=∫d​z​pprior​\(z\)​e−S�​\[x,z\],\\displaystyle p\_\{\\theta\}\(x\)=\\int\\mathrm\{d\}z\\;p\_\{\\rm prior\}\(z\)\\,e^\{\-S\_\{\\theta\}\[x,z\]\}\\;,\(2\.78\)which is exactly the latent\-variable model above\. The ELBO is then a variational lower bound obtained by importance\-weighting this integral withq�q\_\{\\phi\},

log⁡p�​\(x\)\\displaystyle\\log p\_\{\\theta\}\(x\)=logEq�\(⋅\|x\)\[pprior​\(z\)​e−S�​\[x,z\]q�​\(z\|x\)\]≥Eq�\(⋅\|x\)\[−S�\[x,z\]−logq�​\(z\|x\)pprior​\(z\)\]\.\\displaystyle=\\log\\mdmathbb\{E\}\_\{q\_\{\\phi\}\(\\cdot\|x\)\}\\\!\\left\[\\frac\{p\_\{\\rm prior\}\(z\)\\,e^\{\-S\_\{\\theta\}\[x,z\]\}\}\{q\_\{\\phi\}\(z\|x\)\}\\right\]\\geq\\mdmathbb\{E\}\_\{q\_\{\\phi\}\(\\cdot\|x\)\}\\\!\\left\[\-S\_\{\\theta\}\[x,z\]\-\\log\\frac\{q\_\{\\phi\}\(z\|x\)\}\{p\_\{\\rm prior\}\(z\)\}\\right\]\\;\.\(2\.79\)In the path\-integral language, maximizing the ELBO is equivalent to finding the importance distributionq�q\_\{\\phi\}that best approximates the posterior path measurep�​\(z\|x\)p\_\{\\theta\}\(z\|x\)within the parametric family\{q�\}\\\{q\_\{\\phi\}\\\}\.

Within the present unifying perspective, VAEs are thus characterized by two approximations: \(i\) the use of a single latent variable rather than a latent trajectory, and \(ii\) a variational approximation, replacing the intractable posterior by the tractable familyq�​\(z\|x\)q\_\{\\phi\}\(z\|x\)\. The tightness of the bound is controlled by the expressiveness ofq�q\_\{\\phi\}\. Whenq�=p�\(⋅\|x\)q\_\{\\phi\}=p\_\{\\theta\}\(\\cdot\|x\), the gap in[Eq\.2\.75](https://arxiv.org/html/2608.12438#S2.E75)vanishes and the ELBO recovers the true log\-likelihood\.

#### Adversarial models as likelihood\-free saddle points

Setting�=0\\sigma=0in[Eq\.2\.76](https://arxiv.org/html/2608.12438#S2.E76)makes the decoder deterministic,p�​\(x\|z\)→�​\(x−��​\(z\)\)p\_\{\\theta\}\(x\|z\)\\to\\delta\(x\-\\mu\_\{\\theta\}\(z\)\), in the same way thatg→0g\\to 0collapses the path integral onto its saddle in[Section2\.3\.1](https://arxiv.org/html/2608.12438#S2.SS3.SSS1)\. The marginal in[Eq\.2\.71](https://arxiv.org/html/2608.12438#S2.E71)then reads

p�​\(x\)=∫d​z​�​\(x−��​\(z\)\)​pprior​\(z\)\.\\displaystyle p\_\{\\theta\}\(x\)=\\int\\mathrm\{d\}z\\;\\delta\\big\(x\-\\mu\_\{\\theta\}\(z\)\\big\)\\,p\_\{\\rm prior\}\(z\)\\;\.\(2\.80\)Whetherp�p\_\{\\theta\}still has a density depends on��\\mu\_\{\\theta\}alone, and five cases arise\.

1. \(i\)At�\>0\\sigma\>0the decoder is a Gaussian of finite width, so each��​\(z\)\\mu\_\{\\theta\}\(z\)contributes top�​\(x\)p\_\{\\theta\}\(x\)at everyxxand the marginal stays positive everywhere\. The decoder width alone guarantees the density, whatever the latent dimension, and we remain in the variational setting above\.
2. \(ii\)At�=0\\sigma=0with an invertible��\\mu\_\{\\theta\}between spaces of equal dimension,[Eq\.2\.32](https://arxiv.org/html/2608.12438#S2.E32)rewrites the delta and[Eq\.2\.80](https://arxiv.org/html/2608.12438#S2.E80)returns the change of variables, recovering the normalizing flow of[Section2\.3\.1](https://arxiv.org/html/2608.12438#S2.SS3.SSS1)\.
3. \(iii\)At�=0\\sigma=0with equal dimensions but no inverse,[Eq\.2\.32](https://arxiv.org/html/2608.12438#S2.E32)generalizes to a sum over the solutions of��​\(z\)=x\\mu\_\{\\theta\}\(z\)=x, �​\(x−��​\(z\)\)=∑y∈��−1​\(x\)�​\(z−y\)\|det∂��​\(z\)∂z\|z=y,\\displaystyle\\delta\\big\(x\-\\mu\_\{\\theta\}\(z\)\\big\)=\\sum\_\{y\\,\\in\\,\\mu\_\{\\theta\}^\{\-1\}\(x\)\}\\frac\{\\delta\(z\-y\)\}\{\\left\|\\det\\dfrac\{\\partial\\mu\_\{\\theta\}\(z\)\}\{\\partial z\}\\right\|\_\{z=y\}\}\\;,\(2\.81\)where��−1​\(x\)\\mu\_\{\\theta\}^\{\-1\}\(x\)is now a discrete set rather than a single point\. Inserting this into[Eq\.2\.80](https://arxiv.org/html/2608.12438#S2.E80)gives p�​\(x\)=∑y∈��−1​\(x\)pprior​\(y\)\|det∂��​\(z\)∂z\|z=y,\\displaystyle p\_\{\\theta\}\(x\)=\\sum\_\{y\\,\\in\\,\\mu\_\{\\theta\}^\{\-1\}\(x\)\}\\frac\{p\_\{\\rm prior\}\(y\)\}\{\\left\|\\det\\dfrac\{\\partial\\mu\_\{\\theta\}\(z\)\}\{\\partial z\}\\right\|\_\{z=y\}\}\\;,\(2\.82\)the multibranch form of[Eq\.2\.33](https://arxiv.org/html/2608.12438#S2.E33)\. A one\-dimensional example makes this concrete\. The map��​\(z\)=z2\\mu\_\{\\theta\}\(z\)=z^\{2\}reaches onlyx≥0x\\geq 0, sop�p\_\{\\theta\}vanishes forx<0x<0, while everyx\>0x\>0has the two solutionsy=±xy=\\pm\\sqrt\{x\}with derivative of modulus2​x2\\sqrt\{x\}, and[Eq\.2\.82](https://arxiv.org/html/2608.12438#S2.E82)gives p�​\(x\)=pprior​\(x\)\+pprior​\(−x\)2​xwithx\>0\.\\displaystyle p\_\{\\theta\}\(x\)=\\frac\{p\_\{\\rm prior\}\(\\sqrt\{x\}\)\+p\_\{\\rm prior\}\(\-\\sqrt\{x\}\)\}\{2\\sqrt\{x\}\}\\qquad\\text\{with\}\\qquad x\>0\\;\.\(2\.83\)The density is exact and normalized although the map has no inverse, and it diverges atx→0x\\to 0, the image of the critical pointz=0z=0where the derivative vanishes\. Enumerating the solutions of��​\(z\)=x\\mu\_\{\\theta\}\(z\)=xfor a generic network is a global root\-finding problem with no guarantee of completeness, which rendersp�​\(x\)p\_\{\\theta\}\(x\)intractable\.
4. \(iv\)At�=0\\sigma=0with a latent space larger than the data space, the solutions of��​\(z\)=x\\mu\_\{\\theta\}\(z\)=xform a surface of dimensiondimz−dimx\\dim z\-\\dim xrather than a discrete set\. The sum in[Eq\.2\.82](https://arxiv.org/html/2608.12438#S2.E82)becomes an integral over that surface with its induced measure and Jacobian factor, which again exists but is generally intractable\.
5. \(v\)At�=0\\sigma=0with a latent space smaller than the data space, or with a Jacobian that loses rank on an open set,��\\mu\_\{\\theta\}maps into a subset of the data space of zero volume\. The distribution of generated samples remains well defined as a measure, but it assigns all its weight to that subset, so no densityp�​\(x\)p\_\{\\theta\}\(x\)exists and[Eq\.2\.74](https://arxiv.org/html/2608.12438#S2.E74)has no log\-likelihood left to bound\. Samples can still be drawn, butp�p\_\{\\theta\}cannot be evaluated at any point\.

Invertibility is therefore sufficient for a tractable likelihood but not necessary, and a bottleneck is not by itself the obstruction, as VAEs use one freely\. Only at�=0\\sigma=0does the geometry of��\\mu\_\{\\theta\}decide the question, and in cases[\(iii\)](https://arxiv.org/html/2608.12438#S2.I2.i3),[\(iv\)](https://arxiv.org/html/2608.12438#S2.I2.i4)and[\(v\)](https://arxiv.org/html/2608.12438#S2.I2.i5)it leaves us without a likelihood we can evaluate\. Training then has to proceed from samples alone\.

Generative adversarial networks do this\[[26](https://arxiv.org/html/2608.12438#bib.bib11)\]\. A classifierDDand the map��\\mu\_\{\\theta\}share a single objective,

ℒGAN​\(D,�\)=Ex∼pdata​\[log⁡D⁡\(x\)\]\+Ex∼p�​\[log⁡\(1−D⁡\(x\)\)\],\\displaystyle\\mathcal\{L\}\_\{\\rm GAN\}\(D,\\theta\)=\\mdmathbb\{E\}\_\{x\\sim p\_\{\\rm data\}\}\\big\[\\log D\(x\)\\big\]\+\\mdmathbb\{E\}\_\{x\\sim p\_\{\\theta\}\}\\big\[\\log\\big\(1\-D\(x\)\\big\)\\big\]\\;,\(2\.84\)which the classifier maximizes and the generator minimizes\. When both densities exist, maximizing pointwise inD⁡\(x\)D\(x\)at fixed�\\thetagives

D⋆​\(x\)=pdata​\(x\)pdata​\(x\)\+p�​\(x\)andlog⁡D⋆​\(x\)1−D⋆​\(x\)=log⁡pdata​\(x\)p�​\(x\)\.\\displaystyle D^\{\\star\}\(x\)=\\frac\{p\_\{\\rm data\}\(x\)\}\{p\_\{\\rm data\}\(x\)\+p\_\{\\theta\}\(x\)\}\\qquad\\text\{and\}\\qquad\\log\\frac\{D^\{\\star\}\(x\)\}\{1\-D^\{\\star\}\(x\)\}=\\log\\frac\{p\_\{\\rm data\}\(x\)\}\{p\_\{\\theta\}\(x\)\}\\;\.\(2\.85\)The optimal classifier returns the likelihood ratio that the Neyman–Pearson lemma identifies as the most powerful test statistic between two simple hypotheses\[[56](https://arxiv.org/html/2608.12438#bib.bib12)\]\. InsertingD⋆D^\{\\star\}back into[Eq\.2\.84](https://arxiv.org/html/2608.12438#S2.E84)leaves the generator minimizing

ℒGAN\(D⋆,�\)=2JS\(pdata∥p�\)−log4,\\displaystyle\\mathcal\{L\}\_\{\\rm GAN\}\(D^\{\\star\},\\theta\)=2\\,\\mathrm\{JS\}\\big\(p\_\{\\rm data\}\\,\\\|\\,p\_\{\\theta\}\\big\)\-\\log 4\\;,\(2\.86\)whereJS\\mathrm\{JS\}is the Jensen–Shannon divergence, which vanishes whenp�=pdatap\_\{\\theta\}=p\_\{\\rm data\}\. The generator improvesp�p\_\{\\theta\}without ever evaluating it, since the density enters only through the ratio that the classifier estimates from samples\.

All three put the likelihood out of reach, but only case[\(v\)](https://arxiv.org/html/2608.12438#S2.I2.i5)breaks[Eq\.2\.85](https://arxiv.org/html/2608.12438#S2.E85)\. Whenp�p\_\{\\theta\}is a density but we merely cannot evaluate it, the ratio is well defined and the classifier estimates it from samples exactly as intended\. Whenp�p\_\{\\theta\}is singular, its support and that ofpdatap\_\{\\rm data\}generically fail to overlap, which a bottleneck generator enforces by construction and natural data is believed to do on its own, and no optimal discriminator as in[Eq\.2\.85](https://arxiv.org/html/2608.12438#S2.E85)exists\. A classifier that separates the two supports perfectly then maximizes[Eq\.2\.84](https://arxiv.org/html/2608.12438#S2.E84), the Jensen–Shannon divergence equalslog⁡2\\log 2for every such pair, and[Eq\.2\.86](https://arxiv.org/html/2608.12438#S2.E86)stays constant in�\\thetawith vanishing gradient\[[4](https://arxiv.org/html/2608.12438#bib.bib13)\]\. Smearing both distributions with noise before comparing them restores the overlap and with it a usable gradient, so the width removed in[Eq\.2\.80](https://arxiv.org/html/2608.12438#S2.E80)returns in the training criterion rather than in the model\. The smearing need not be applied to the samples\. Expanding the discriminator to leading order in the noise variance turns the convolution into a penalty on its gradient norm\[[65](https://arxiv.org/html/2608.12438#bib.bib8)\], so a gradient\-penalized discriminator minimizes the smoothed divergence while the objective stays untouched, and such penalties restore local convergence even when both distributions lie on lower\-dimensional manifolds\[[54](https://arxiv.org/html/2608.12438#bib.bib7)\]\. Constraining the discriminator through spectral normalization\[[55](https://arxiv.org/html/2608.12438#bib.bib9)\]limits how sharply it can separate the two supports and acts in the same direction\.

ModelForwardReverseLossLikelihoodPath\-integral pictureNF/CNFDet\.Det\.LikelihoodExactSaddle point \(g→0g\\to 0\)CFMCond\. pathDet\.Cond\. flow matchingODE\-basedProbability\-flow ODEDDPM/SMStoch\.Stoch\.Denoising/score matchingIntractableFull path measureSBStoch\.Stoch\.Path\-space KL \(IPF\)IntractableGlob\. path\-space var\.VAEStoch\.Stoch\.ELBOLower boundSingle\-latent var\.GAN–Det\.AdversarialNone/IntractableSaddle point \(g→0g\\to 0\)Table 1:Comparison of generative models within the common framework\. “Forward” and “Reverse” denote the dynamics used to propagate probability mass from data to latent space and back, while “Det\.” and “Stoch\.” indicate deterministic and stochastic sampling dynamics, and “Cond\. path” a prescribed conditional interpolation\. “ODE\-based” likelihood means the density is available only by integrating the probability\-flow ODE, not in closed form\. The final column gives the role each model plays within the path\-integral formulation of[Section2](https://arxiv.org/html/2608.12438#S2)\.
#### 2\.3\.4Summary: evaluation principles for the master path integral

All generative models discussed in this section are organized by the same path integral in[Eq\.2\.29](https://arxiv.org/html/2608.12438#S2.E29)\. Their differences arise not from distinct probabilistic foundations but from how that path integral is evaluated or approximated in practice, while variational autoencoders and adversarial models arise as degenerate single\-step and likelihood\-free realizations of the underlying latent\-variable formulation\.[Table1](https://arxiv.org/html/2608.12438#S2.T1)summarizes the resulting taxonomy\. Viewed through this unified lens, the apparent diversity of modern generative models reflects different evaluation strategies within a common latent\-variable construction, and for the continuous\-time transport models these strategies act on the same stochastic path integral\.

## 3Probability flows as interacting field theories

In[Section2](https://arxiv.org/html/2608.12438#S2), we wrote a wide class of generative models as a continuous probabilistic path integral, with the data density obtained by summing over latent trajectories connecting the prior att=1t=1to the observation att=0t=0\. Reverse\-time dynamics with driftfrev​\(z,t\)f\_\{\\mathrm\{rev\}\}\(z,t\)and diffusiong⁡\(t\)g\(t\)lead to the Onsager–Machlup action,

SOM​\[z\]=∫01d​t​‖z˙−frev​\(z,t\)‖22​g2​\(t\)\.\\displaystyle S\_\{\\text\{OM\}\}\[z\]=\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\frac\{\\\|\\dot\{z\}\-f\_\{\\mathrm\{rev\}\}\(z,t\)\\\|^\{2\}\}\{2g^\{2\}\(t\)\}\\;\.\(3\.1\)Linear \(or linearized\) probability flows define Gaussian path measures and play the role of free theories, while nonlinearities in the drift generate interactions\. This perspective leads to a diagrammatic expansion and a scattering interpretation of generative probability transport, which we develop below\.

### 3\.1Free–interacting split and MSRJD representation

The central object in the path\-integral formulation is the actionSOM​\[z\]S\_\{\\text\{OM\}\}\[z\], which determines the path measure and hence the generative model\. From a field\-theoretic point of view, the distinction between*free*and*interacting*probability flows is governed by whether the action is quadratic in the path variables\. To make this explicit, we split the reverse drift into a linear \(affine\) part and a nonlinear remainder,

frev​\(z,t\)=A⁡\(t\)​z\+b⁡\(t\)\+fint​\(z,t\),fint​\(⋅,t\)​nonlinear in​z\.\\displaystyle f\_\{\\mathrm\{rev\}\}\(z,t\)=A\(t\)\\,z\+b\(t\)\+f\_\{\\mathrm\{int\}\}\(z,t\)\\;,\\qquad f\_\{\\mathrm\{int\}\}\(\\,\\cdot\\,,t\)\\;\\text\{nonlinear in \}z\\;\.\(3\.2\)To build intuition, take the usual affine forward process transporting the data to a Gaussian prior\. If the data is also an isotropic Gaussian, every forward marginalp⁡\(z,t\)p\(z,t\)stays Gaussian, the exact score∇z​log​p\\nabla\_\{z\}\\log pis linear inzz, andfrevf\_\{\\mathrm\{rev\}\}is affine, sofint=0f\_\{\\mathrm\{int\}\}=0and the theory is free\. Non\-Gaussian data, such as multimodal targets, make the intermediate scores nonlinear and generatefintf\_\{\\mathrm\{int\}\}, which therefore measures how far the transport departs from connecting two Gaussians\. Inserting[Eq\.3\.2](https://arxiv.org/html/2608.12438#S3.E2)into the Onsager–Machlup action yields

SOM​\[z\]=∫01d​t​‖z˙−A⁡\(t\)​z−b⁡\(t\)‖22​g2​\(t\)\+Sint​\[z\],\\displaystyle S\_\{\\text\{OM\}\}\[z\]=\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\frac\{\\\|\\dot\{z\}\-A\(t\)\\,z\-b\(t\)\\\|^\{2\}\}\{2g^\{2\}\(t\)\}\\;\+\\;S\_\{\\mathrm\{int\}\}\[z\]\\;,\(3\.3\)where the interaction functional is defined by

Sint​\[z\]≡∫01d​t​12​g2​\(t\)​\(‖fint​\(z,t\)‖2−2​\(z˙−A⁡\(t\)​z−b⁡\(t\)\)⋅fint​\(z,t\)\)\.\\displaystyle S\_\{\\mathrm\{int\}\}\[z\]\\equiv\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\frac\{1\}\{2g^\{2\}\(t\)\}\\Big\(\\\|f\_\{\\mathrm\{int\}\}\(z,t\)\\\|^\{2\}\-2\\,\(\\dot\{z\}\-A\(t\)\\,z\-b\(t\)\)\\cdot f\_\{\\mathrm\{int\}\}\(z,t\)\\Big\)\\;\.\(3\.4\)Iffint=0f\_\{\\mathrm\{int\}\}=0, the action is quadratic inzzand the path integral is Gaussian, so all correlation functions are determined by the corresponding propagators\. Conversely, any nonlinear dependence of the drift generates non\-quadratic terms inSintS\_\{\\mathrm\{int\}\}, which act as interaction vertices in a perturbative expansion\.

While[Eq\.3\.4](https://arxiv.org/html/2608.12438#S3.E4)defines interactions directly at the level of the Onsager–Machlup action, it is not always the most convenient representation for perturbation theory\. In particular, the Onsager–Machlup form is second order in time derivatives and entangles drift and noise in a way that obscures the causal structure of the underlying dynamics\. For diagrammatic calculations it is therefore advantageous to switch to a first\-order functional integral representation, in which interactions enter linearly and the response structure is explicit\.

The MSRJD\[[52](https://arxiv.org/html/2608.12438#bib.bib68),[35](https://arxiv.org/html/2608.12438#bib.bib67),[17](https://arxiv.org/html/2608.12438#bib.bib66)\]construction provides exactly this representation, which we derive by considering again the reverse\-time dynamics

d​z=frev​\(z,t\)​d​t\+g⁡\(t\)​d​W¯t,\\displaystyle\\mathrm\{d\}z=f\_\{\\mathrm\{rev\}\}\(z,t\)\\,\\mathrm\{d\}t\+g\(t\)\\,\\mathrm\{d\}\\bar\{W\}\_\{t\}\\;,\(3\.5\)and rewrite the associated path weight as a first\-order functional integral overzzand an auxiliary field\. Our central object is the path weight from[Eq\.2\.37](https://arxiv.org/html/2608.12438#S2.E37),

e−SOM​\[z\]=exp\[−∫01dt‖z˙−frev​\(z,t\)‖22​g2​\(t\)\],\\displaystyle e^\{\-S\_\{\\text\{OM\}\}\[z\]\}=\\exp\\\!\\left\[\-\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\frac\{\\\|\\dot\{z\}\-f\_\{\\mathrm\{rev\}\}\(z,t\)\\\|^\{2\}\}\{2g^\{2\}\(t\)\}\\right\]\\;,\(3\.6\)which we bring to MSRJD form\. First, we represent this Gaussian as a noise average of a functional delta enforcing the equation of motion\. Introducing the reverse white\-noise field�​\(t\)\\xi\(t\), the continuum limit of the discrete noise�k\\xi\_\{k\}in[Eq\.2\.11](https://arxiv.org/html/2608.12438#S2.E11), the dynamics take the Langevin form

z˙​\(t\)=frev​\(z⁡\(t\),t\)\+g⁡\(t\)​�​\(t\)with⟨�​\(t\)​�​\(t′\)⊤⟩=�​\(t−t′\)​I,\\displaystyle\\dot\{z\}\(t\)=f\_\{\\mathrm\{rev\}\}\(z\(t\),t\)\+g\(t\)\\,\\xi\(t\)\\qquad\\text\{with\}\\qquad\\big\\langle\\xi\(t\)\\,\\xi\(t^\{\\prime\}\)^\{\\\!\\top\}\\big\\rangle=\\delta\(t\-t^\{\\prime\}\)\\,\\mdmathbb\{I\}\\;,\(3\.7\)and introducing the delta enforcing[Eq\.3\.7](https://arxiv.org/html/2608.12438#S3.E7)gives

e−SOM​\[z\]=∫𝒟�exp\[−12∫01dt∥�\(t\)∥2\]�\[z˙−frev\(z,t\)−g\(t\)�\(t\)\],\\displaystyle e^\{\-S\_\{\\text\{OM\}\}\[z\]\}=\\int\\mathcal\{D\}\\xi\\;\\exp\\\!\\left\[\-\\frac\{1\}\{2\}\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\\|\\xi\(t\)\\\|^\{2\}\\right\]\\delta\\\!\\left\[\\dot\{z\}\-f\_\{\\mathrm\{rev\}\}\(z,t\)\-g\(t\)\\,\\xi\(t\)\\right\]\\;,\(3\.8\)since the�\\xi\-integral against the delta setsg​�=z˙−frevg\\xi=\\dot\{z\}\-f\_\{\\mathrm\{rev\}\}and returns[Eq\.3\.6](https://arxiv.org/html/2608.12438#S3.E6)\. Second, we resolve the delta by its functional Fourier representation\[[78](https://arxiv.org/html/2608.12438#bib.bib50),[71](https://arxiv.org/html/2608.12438#bib.bib76),[36](https://arxiv.org/html/2608.12438#bib.bib25)\], which introduces the real*response field*z^​\(t\)\\hat\{z\}\(t\),

�\[z˙−frev−g�\]∝∫𝒟z^exp\[−i∫01dtz^\(t\)⋅\(z˙\(t\)−frev\(z\(t\),t\)−g\(t\)�\(t\)\)\]\.\\displaystyle\\delta\\\!\\left\[\\dot\{z\}\-f\_\{\\mathrm\{rev\}\}\-g\\xi\\right\]\\propto\\int\\mathcal\{D\}\\hat\{z\}\\;\\exp\\\!\\left\[\-\\mathrm\{i\}\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\hat\{z\}\(t\)\\cdot\\Big\(\\dot\{z\}\(t\)\-f\_\{\\mathrm\{rev\}\}\(z\(t\),t\)\-g\(t\)\\,\\xi\(t\)\\Big\)\\right\]\\;\.\(3\.9\)Third, the noise field�​\(t\)\\xi\(t\)now appears linearly and is integrated out by completing the square, a Hubbard–Stratonovich transformation\[[70](https://arxiv.org/html/2608.12438#bib.bib26),[34](https://arxiv.org/html/2608.12438#bib.bib27)\]that trades the linear noise coupling for the quadratic response term,

∫𝒟�exp\[−∫01dt\(12∥�∥2\+ig\(t\)z^⋅�\)\]=exp\[−12∫01dtg2\(t\)∥z^∥2\]\.\\displaystyle\\int\\mathcal\{D\}\\xi\\;\\exp\\\!\\left\[\-\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\,\\Big\(\\frac\{1\}\{2\}\\,\\\|\\xi\\\|^\{2\}\+\\mathrm\{i\}\\,g\(t\)\\,\\hat\{z\}\\cdot\\xi\\Big\)\\right\]=\\exp\\\!\\left\[\-\\frac\{1\}\{2\}\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;g^\{2\}\(t\)\\,\\\|\\hat\{z\}\\\|^\{2\}\\right\]\\;\.\(3\.10\)Collecting the three steps expresses the weight as a functional integral over the response field,

e−SOM​\[z\]=∫𝒟​z^​e−SMSRJD​\[z,z^\],\\displaystyle e^\{\-S\_\{\\text\{OM\}\}\[z\]\}=\\int\\mathcal\{D\}\\hat\{z\}\\;e^\{\-S\_\{\\mathrm\{MSRJD\}\}\[z,\\hat\{z\}\]\}\\;,\(3\.11\)with the measure𝒟​z^\\mathcal\{D\}\\hat\{z\}absorbing thezz\-independent normalization of the Fourier representation, as𝒟​z\\mathcal\{D\}zabsorbs that of the Onsager–Machlup weight in[Section2\.2](https://arxiv.org/html/2608.12438#S2.SS2), and the MSRJD action

SMSRJD​\[z,z^\]=∫01d​t​\[i​z^⋅\(z˙−frev​\(z,t\)\)\+12​g2​\(t\)​‖z^‖2\],\\displaystyle S\_\{\\mathrm\{MSRJD\}\}\[z,\\hat\{z\}\]=\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\left\[\\mathrm\{i\}\\,\\hat\{z\}\\cdot\\big\(\\dot\{z\}\-f\_\{\\mathrm\{rev\}\}\(z,t\)\\big\)\+\\frac\{1\}\{2\}\\,g^\{2\}\(t\)\\,\\\|\\hat\{z\}\\\|^\{2\}\\right\]\\;,\(3\.12\)where we use the Itô convention adopted in[Section2\.2](https://arxiv.org/html/2608.12438#S2.SS2)\. With causal time ordering, the functional determinant from the delta constraint is triangular with a field\-independent diagonal, and we absorb it into the normalization of𝒟​z^\\mathcal\{D\}\\hat\{z\}\. Other discretizations generate an additional local divergence term and would modify the interaction vertices accordingly\.

The MSRJD representation provides a first\-order functional integral in which nonlinearities in the drift enter linearly through the couplingi​z^⋅frev​\(z,t\)\\mathrm\{i\}\\,\\hat\{z\}\\cdot f\_\{\\mathrm\{rev\}\}\(z,t\)\. For realz^\\hat\{z\}the noise term damps the integrand, so the representation is convergent as it stands\. Equivalently, the factor ofi\\mathrm\{i\}can be absorbed into a response field integrated along the imaginary axis, which renders the action real\[[78](https://arxiv.org/html/2608.12438#bib.bib50),[71](https://arxiv.org/html/2608.12438#bib.bib76)\]\. Using the decomposition from[Eq\.3\.2](https://arxiv.org/html/2608.12438#S3.E2), the action becomes

SMSRJD​\[z,z^\]=S0​\[z,z^\]\+Sint​\[z,z^\],\\displaystyle S\_\{\\mathrm\{MSRJD\}\}\[z,\\hat\{z\}\]=S\_\{0\}\[z,\\hat\{z\}\]\+S\_\{\\mathrm\{int\}\}\[z,\\hat\{z\}\]\\;,\(3\.13\)with the free part

S0​\[z,z^\]=∫01d​t​\[i​z^⋅\(z˙−A⁡\(t\)​z−b⁡\(t\)\)\+12​g2​\(t\)​‖z^‖2\],\\displaystyle S\_\{0\}\[z,\\hat\{z\}\]=\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\left\[\\mathrm\{i\}\\,\\hat\{z\}\\cdot\\big\(\\dot\{z\}\-A\(t\)\\,z\-b\(t\)\\big\)\+\\frac\{1\}\{2\}\\,g^\{2\}\(t\)\\,\\\|\\hat\{z\}\\\|^\{2\}\\right\]\\;,\(3\.14\)and the interaction

Sint\[z,z^\]=−i∫01dtz^\(t\)⋅fint\(z\(t\),t\)\.\\displaystyle S\_\{\\mathrm\{int\}\}\[z,\\hat\{z\}\]=\-\\,\\mathrm\{i\}\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\hat\{z\}\(t\)\\cdot f\_\{\\mathrm\{int\}\}\(z\(t\),t\)\\;\.\(3\.15\)

### 3\.2Scattering of probability flows

The MSRJD representation introduced above allows us to reinterpret generative probability transport as a scattering process in latent space\. In this picture, probability mass is propagated from a simple prior distribution att=1t=1to the data manifold att=0t=0by an interacting dynamical evolution, in close analogy with time\-dependent scattering in quantum field theory\.

The evolution between the two ends is carried by the transition kernelK\(z0,0∣z1,1\)K\(z\_\{0\},0\\mid z\_\{1\},1\)constructed in[Section2\.2](https://arxiv.org/html/2608.12438#S2.SS2)\. Through the marginal equation[2\.27](https://arxiv.org/html/2608.12438#S2.E27), the priorpprior​\(z1\)p\_\{\\rm prior\}\(z\_\{1\}\)plays the role of an*in\-state*, the data densitypdata​\(z0\)p\_\{\\rm data\}\(z\_\{0\}\)defines an*out\-state*, and the kernel encodes the dynamical evolution between them, in analogy with a transition amplitude\.

In the absence of interactions,i\.e\.a linear reverse drift and hence a quadratic action, we can evaluate the kernelKKexactly\. The resulting Gaussian kernel describes a free probability flow, in which probability mass propagates without mode coupling or distortion beyond what the linear drift and the diffusion induce\. This situation is directly analogous to free\-particle propagation, where the transition amplitude is fully determined by the free propagator\.

Nonlinear probability flows correspond to interacting theories, for which the kernelKKcannot be evaluated in closed form\. Rewriting the Onsager–Machlup weight in[Eq\.2\.19](https://arxiv.org/html/2608.12438#S2.E19)with the MSRJD identity in[Eq\.3\.11](https://arxiv.org/html/2608.12438#S3.E11), the kernel becomes a two\-field functional integral,

K\(z0,0∣z1,1\)=∫𝒟z𝒟z^e−SMSRJD​\[z,z^\],\\displaystyle K\(z\_\{0\},0\\mid z\_\{1\},1\)=\\int\\mathcal\{D\}z\\,\\mathcal\{D\}\\hat\{z\}\\;e^\{\-S\_\{\\mathrm\{MSRJD\}\}\[z,\\hat\{z\}\]\}\\;,\(3\.16\)with the latent\-path boundary values fixed by the arguments of the kernel as in[Section2\.2](https://arxiv.org/html/2608.12438#S2.SS2), and the response field unconstrained at both ends\. For perturbation theory it is convenient to release the data endpoint and work in the ensemble of the sampler,i\.e\.trajectories launched from the prior pointz1z\_\{1\}with a free final state\. We therefore define the normalized free expectation value

⟨O⁡\[z,z^\]⟩0≡∫d​z0​∫𝒟​z​𝒟​z^​O​\[z,z^\]​e−S0​\[z,z^\]with⟨1⟩0=1,\\displaystyle\\big\\langle O\[z,\\hat\{z\}\]\\big\\rangle\_\{0\}\\equiv\\int\\mathrm\{d\}z\_\{0\}\\int\\mathcal\{D\}z\\,\\mathcal\{D\}\\hat\{z\}\\;O\[z,\\hat\{z\}\]\\,e^\{\-S\_\{0\}\[z,\\hat\{z\}\]\}\\qquad\\text\{with\}\\qquad\\langle 1\\rangle\_\{0\}=1\\;,\(3\.17\)where the inner path integral is pinned atz⁡\(0\)=z0z\(0\)=z\_\{0\}andz⁡\(1\)=z1z\(1\)=z\_\{1\}as in[Eq\.2\.19](https://arxiv.org/html/2608.12438#S2.E19)and the data endpoint is subsequently integrated over\. AtSint=0S\_\{\\mathrm\{int\}\}=0the inner integral is the transition density of the affine process, normalized as in[Eq\.2\.26](https://arxiv.org/html/2608.12438#S2.E26)\. The average describes realizations of the free reverse process started atz1z\_\{1\}\. Pinning and releasing the data endpoint are related exactly, since inserting a delta function of the endpoint collapses thez0z\_\{0\}\-integral back onto the pinned path integral,

⟨�​\(z⁡\(0\)−z0\)​O​\[z,z^\]⟩0=∫𝒟​z​𝒟​z^​O​\[z,z^\]​e−S0​\[z,z^\],\\displaystyle\\big\\langle\\delta\\big\(z\(0\)\-z\_\{0\}\\big\)\\,O\[z,\\hat\{z\}\]\\big\\rangle\_\{0\}=\\int\\mathcal\{D\}z\\,\\mathcal\{D\}\\hat\{z\}\\;O\[z,\\hat\{z\}\]\\,e^\{\-S\_\{0\}\[z,\\hat\{z\}\]\}\\;,\(3\.18\)with the boundary values on the right\-hand side again fixed atz⁡\(0\)=z0z\(0\)=z\_\{0\}andz⁡\(1\)=z1z\(1\)=z\_\{1\}\. SplittingSMSRJD=S0\+SintS\_\{\\mathrm\{MSRJD\}\}=S\_\{0\}\+S\_\{\\mathrm\{int\}\}into its free and interacting parts, cf\.[Eq\.3\.13](https://arxiv.org/html/2608.12438#S3.E13), and expanding the interaction exponential inside[Eq\.3\.16](https://arxiv.org/html/2608.12438#S3.E16), each order of the transition kernel becomes a one\-sided free average with an endpoint insertion,

K\(z0,0∣z1,1\)\\displaystyle K\(z\_\{0\},0\\mid z\_\{1\},1\)=⟨�​\(z⁡\(0\)−z0\)​e−Sint​\[z,z^\]⟩0\\displaystyle=\\big\\langle\\delta\\big\(z\(0\)\-z\_\{0\}\\big\)\\,e^\{\-S\_\{\\mathrm\{int\}\}\[z,\\hat\{z\}\]\}\\big\\rangle\_\{0\}=∑n=0∞\(−1\)nn\!⟨�\(z\(0\)−z0\)\(Sint\[z,z^\]\)n⟩0≡∑n=0∞K\(n\)\(z0,0∣z1,1\)\.\\displaystyle=\\sum\_\{n=0\}^\{\\infty\}\\frac\{\(\-1\)^\{n\}\}\{n\!\}\\big\\langle\\delta\\big\(z\(0\)\-z\_\{0\}\\big\)\\,\\big\(S\_\{\\mathrm\{int\}\}\[z,\\hat\{z\}\]\\big\)^\{n\}\\big\\rangle\_\{0\}\\equiv\\sum\_\{n=0\}^\{\\infty\}K^\{\(n\)\}\(z\_\{0\},0\\mid z\_\{1\},1\)\\;\.\(3\.19\)The zeroth orderK\(0\)=⟨�​\(z⁡\(0\)−z0\)⟩0K^\{\(0\)\}=\\langle\\delta\(z\(0\)\-z\_\{0\}\)\\rangle\_\{0\}is the free kernel, evaluated in closed form in[Section3\.3](https://arxiv.org/html/2608.12438#S3.SS3), and the first corrections read

K\(1\)\(z0,0∣z1,1\)\\displaystyle K^\{\(1\)\}\(z\_\{0\},0\\mid z\_\{1\},1\)=−⟨�​\(z⁡\(0\)−z0\)​Sint​\[z,z^\]⟩0,\\displaystyle=\-\\,\\big\\langle\\delta\\big\(z\(0\)\-z\_\{0\}\\big\)\\,S\_\{\\mathrm\{int\}\}\[z,\\hat\{z\}\]\\big\\rangle\_\{0\}\\;,K\(2\)\(z0,0∣z1,1\)\\displaystyle K^\{\(2\)\}\(z\_\{0\},0\\mid z\_\{1\},1\)=12​⟨�​\(z⁡\(0\)−z0\)​\(Sint​\[z,z^\]\)2⟩0\.\\displaystyle=\\frac\{1\}\{2\}\\,\\big\\langle\\delta\\big\(z\(0\)\-z\_\{0\}\\big\)\\,\\big\(S\_\{\\mathrm\{int\}\}\[z,\\hat\{z\}\]\\big\)^\{2\}\\big\\rangle\_\{0\}\\;\.\(3\.20\)Using the explicit form of[Eq\.3\.15](https://arxiv.org/html/2608.12438#S3.E15), these become

K\(1\)\(z0,0∣z1,1\)\\displaystyle K^\{\(1\)\}\(z\_\{0\},0\\mid z\_\{1\},1\)=i​∫01d​t​⟨�​\(z⁡\(0\)−z0\)​z^​\(t\)⋅fint​\(z,t\)⟩0,\\displaystyle=\\mathrm\{i\}\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\Big\\langle\\delta\\big\(z\(0\)\-z\_\{0\}\\big\)\\,\\hat\{z\}\(t\)\\cdot f\_\{\\mathrm\{int\}\}\(z,t\)\\Big\\rangle\_\{0\}\\;,K\(2\)\(z0,0∣z1,1\)\\displaystyle K^\{\(2\)\}\(z\_\{0\},0\\mid z\_\{1\},1\)=−12∫01dt1dt2⟨�\(z\(0\)−z0\)z^\(t1\)⋅fint\(z,t1\)z^\(t2\)⋅fint\(z,t2\)⟩0\.\\displaystyle=\-\\frac\{1\}\{2\}\\int\_\{0\}^\{1\}\\mathrm\{d\}t\_\{1\}\\,\\mathrm\{d\}t\_\{2\}\\;\\Big\\langle\\delta\\big\(z\(0\)\-z\_\{0\}\\big\)\\,\\hat\{z\}\(t\_\{1\}\)\\cdot f\_\{\\mathrm\{int\}\}\(z,t\_\{1\}\)\\;\\hat\{z\}\(t\_\{2\}\)\\cdot f\_\{\\mathrm\{int\}\}\(z,t\_\{2\}\)\\Big\\rangle\_\{0\}\\;\.\(3\.21\)Integrating[Eq\.3\.19](https://arxiv.org/html/2608.12438#S3.E19)over the data endpoint removes the insertion,

∫d​z0​K\(n\)=\(−1\)nn\!​⟨Sintn⟩0,\\displaystyle\\int\\mathrm\{d\}z\_\{0\}\\,K^\{\(n\)\}=\\frac\{\(\-1\)^\{n\}\}\{n\!\}\\langle S\_\{\\mathrm\{int\}\}^\{n\}\\rangle\_\{0\}\\;,\(3\.22\)and every such average vanishes forn≥1n\\geq 1\. Each response field must contract with a latent field at one of the vertices, so every Wick pattern contains a closed cycle of contractions⟨z⁡\(t\)​z^​\(t′\)⟩0\\langle z\(t\)\\,\\hat\{z\}\(t^\{\\prime\}\)\\rangle\_\{0\}\. These contractions are causal along the sampling direction and vanish at equal times in the Itô convention, so no such cycle survives\. The normalization∫d​z0​K=1\\int\\mathrm\{d\}z\_\{0\}\\,K=1of[Eq\.2\.26](https://arxiv.org/html/2608.12438#S2.E26)therefore holds order by order in perturbation theory\. This structure shows how nonlinearities in the drift deform free propagation\. The zeroth\-order term is Gaussian transport governed by the linearized dynamics, while higher orders correspond to successive*insertions*of the nonlinear driftfintf\_\{\\mathrm\{int\}\}at intermediate times\. This insertion structure is precise at the level of probability transport\. Splitting the Fokker–Planck generator of[Eq\.2\.22](https://arxiv.org/html/2608.12438#S2.E22)according to[Eq\.3\.2](https://arxiv.org/html/2608.12438#S3.E2), the Duhamel formula for a perturbed evolution operator\[[63](https://arxiv.org/html/2608.12438#bib.bib6)\]gives

K\(1\)\(z0,0∣z1,1\)=∫01dt∫dzK\(0\)\(z0,0∣z,t\)∇z⋅\(fint\(z,t\)K\(0\)\(z,t∣z1,1\)\),\\displaystyle K^\{\(1\)\}\(z\_\{0\},0\\mid z\_\{1\},1\)=\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\int\\mathrm\{d\}z\\;K^\{\(0\)\}\(z\_\{0\},0\\mid z,t\)\\;\\nabla\_\{z\}\\\!\\cdot\\\!\\Big\(f\_\{\\mathrm\{int\}\}\(z,t\)\\,K^\{\(0\)\}\(z,t\\mid z\_\{1\},1\)\\Big\)\\;,\(3\.23\)free propagation from the prior to the intermediate timett, a single interaction, and free propagation onward, with higher orders iterating this pattern in time\-ordered fashion\.[Figure1](https://arxiv.org/html/2608.12438#S3.F1)states the same expansion diagrammatically, withK\(n\)K^\{\(n\)\}carryingnninsertions offintf\_\{\\mathrm\{int\}\}whatever that drift happens to be\. The interaction vertex is therefore the interacting part of the Fokker–Planck generator, inserted between free kernels\. The MSRJD vertexi​z^⋅fint\\mathrm\{i\}\\,\\hat\{z\}\\cdot f\_\{\\mathrm\{int\}\}of[Eq\.3\.15](https://arxiv.org/html/2608.12438#S3.E15)is its path\-integral representation, where its diagrammatic counterpart is the vertex field content derived in[Section3\.3](https://arxiv.org/html/2608.12438#S3.SS3), with exactly one response leg per vertex\.

z1z\_\{1\}z0z\_\{0\}==K\(0\)K^\{\(0\)\}\+\+fint​\(t\)f\_\{\\mathrm\{int\}\}\(t\)K\(1\)K^\{\(1\)\}\+\+K\(2\)K^\{\(2\)\}\+…\+\\;\\dotsFigure 1:Diagrammatic form of the kernel expansion in[Eq\.3\.19](https://arxiv.org/html/2608.12438#S3.E19)\. The hatched blob is the full kernelK\(z0,0∣z1,1\)K\(z\_\{0\},0\\mid z\_\{1\},1\)and the plain arrowed line is free propagationK\(0\)K^\{\(0\)\}under the linearized drift\. Each dot is one insertion offintf\_\{\\mathrm\{int\}\}integrated over its time argument as in[Eq\.3\.23](https://arxiv.org/html/2608.12438#S3.E23), soK\(n\)K^\{\(n\)\}carriesnninsertions\. No form offintf\_\{\\mathrm\{int\}\}is assumed\. Resolving a dot into vertices with a definite number of legs requires expandingfintf\_\{\\mathrm\{int\}\}in powers ofzz, which is done in[Section3\.3](https://arxiv.org/html/2608.12438#S3.SS3)\. Arrows run along the generative direction, from the prior att=1t=1to the data att=0t=0\.The free MSRJD measure contracts response and latent fields according to the Gaussian propagators defined byS0S\_\{0\}, generating a systematic series of corrections that can be organized and resummed\. Like perturbative expansions in interacting field theories, this series is in general*asymptotic*rather than convergent, since the number of contractions contributing at ordernngrows factorially\. It is reliable as a low\-order approximation in the weak\-coupling, small\-g2g^\{2\}regime\[[78](https://arxiv.org/html/2608.12438#bib.bib50)\]\. In this sense, learning a generative model corresponds to tuning the interaction functionalfint​\(z,t\)f\_\{\\mathrm\{int\}\}\(z,t\)such that the resulting \(interacting\) transition kernel in[Eq\.3\.19](https://arxiv.org/html/2608.12438#S3.E19), when convoluted with the prior as in[Eq\.2\.27](https://arxiv.org/html/2608.12438#S2.E27), reproduces the observed data distribution\.

This scattering viewpoint clarifies the role of determinism and stochasticity in generative modeling\. In the limitg⁡\(t\)→0g\(t\)\\to 0, the quadratic noise term12​g2​\(t\)​‖z^‖2\\tfrac\{1\}\{2\}g^\{2\}\(t\)\\\|\\hat\{z\}\\\|^\{2\}in the free MSRJD action vanishes and the response fieldz^\\hat\{z\}becomes a Lagrange multiplier enforcing the deterministic flow,i\.e\.

∫𝒟z^exp\[−i∫01dtz^\(t\)⋅\(z˙\(t\)−frev\(z\(t\),t\)\)\]∝�\[z˙−frev\(z,t\)\]\.\\displaystyle\\int\\mathcal\{D\}\\hat\{z\}\\;\\exp\\\!\\left\[\-\\mathrm\{i\}\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\hat\{z\}\(t\)\\cdot\\big\(\\dot\{z\}\(t\)\-f\_\{\\mathrm\{rev\}\}\(z\(t\),t\)\\big\)\\right\]\\propto\\delta\\\!\\big\[\\dot\{z\}\-f\_\{\\mathrm\{rev\}\}\(z,t\)\\big\]\\;\.\(3\.24\)Hence, the path integral collapses onto deterministic trajectories constrained by

z˙​\(t\)=frev​\(z⁡\(t\),t\),\\displaystyle\\dot\{z\}\(t\)=f\_\{\\mathrm\{rev\}\}\(z\(t\),t\)\\;,\(3\.25\)and the transition kernel reduces to its tree\-level \(saddle\-point\) approximation, recovering normalizing flows as a leading\-order description\. With the response field no longer fluctuating,SintS\_\{\\mathrm\{int\}\}produces no loop contributions and enters only through the driftfrevf\_\{\\mathrm\{rev\}\}in[Eq\.3\.25](https://arxiv.org/html/2608.12438#S3.E25)\. For finite diffusion strength, the noise term induces fluctuations of the response field and generates higher\-order corrections in the expansion of[Eq\.3\.19](https://arxiv.org/html/2608.12438#S3.E19)\.

### 3\.3Diagrammatic expansion and interpretation

The perturbative expansion of the transition kernel derived in[Section3\.2](https://arxiv.org/html/2608.12438#S3.SS2)organizes into propagators and interaction vertices, in direct parallel with interacting field theories\. Solid lines with a centered arrow denote contracted response propagators, dashed lines denote noise contractions, and plain solid lines denote uncontractedzzlegs, which attach to external states or to the classical field and turn into an arrowed response line or a dashed noise line only once contracted\. Thez^\\hat\{z\}leg of a vertex or insertion passes its arrow to the response line it emits\. Each insertion of[Fig\.1](https://arxiv.org/html/2608.12438#S3.F1)resolves into vertices once we expand the drift in powers of the state\. Expandingfintf\_\{\\mathrm\{int\}\}in powers of the state, with couplings written as tensors�\(m\)\\lambda^\{\(m\)\}of rankm\+1m\+1, the interaction functional in[Eq\.3\.15](https://arxiv.org/html/2608.12438#S3.E15)becomes

Sint\[z,z^\]=i∑m≥2∫01dt�ij1⋯jm\(m\)\(t\)z^i\(t\)zj1\(t\)⋯zjm\(t\),\\displaystyle S\_\{\\mathrm\{int\}\}\[z,\\hat\{z\}\]=\\mathrm\{i\}\\sum\_\{m\\geq 2\}\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\lambda^\{\(m\)\}\_\{i\\,j\_\{1\}\\cdots j\_\{m\}\}\(t\)\\;\\hat\{z\}\_\{i\}\(t\)\\,z\_\{j\_\{1\}\}\(t\)\\cdots z\_\{j\_\{m\}\}\(t\)\\;,\(3\.26\)with repeated indices summed over theddlatent components\. The term of ordermmgives a vertex with onez^\\hat\{z\}leg andmmlegs ofzz\. The expansion is used at finite order rather than resummed, so what matters is not its convergence but the size of the couplings left out\. A drift quadratic inzzgenerates the cubic vertexz^⋅z2\\hat\{z\}\\cdot z^\{2\}, a cubic drift the quartic vertexz^⋅z3\\hat\{z\}\\cdot z^\{3\}, and so on\.[Figure2](https://arxiv.org/html/2608.12438#S3.F2)collects the resulting elements, drawn for representative vertex orders\. Which vertices a model carries is fixed by its drift, while the propagators and the counting rules are the same for all of them\.

Two features of[Eq\.3\.26](https://arxiv.org/html/2608.12438#S3.E26)are fixed by the construction rather than by the particular drift\. First, every interaction vertex carries at least one response legz^\\hat\{z\}\. Integratingz^\\hat\{z\}out of[Eq\.3\.12](https://arxiv.org/html/2608.12438#S3.E12)enforces the equation of motion through�​\[z˙−frev\]\\delta\[\\dot\{z\}\-f\_\{\\mathrm\{rev\}\}\], whosezz\-integral is the normalization in[Eq\.2\.26](https://arxiv.org/html/2608.12438#S2.E26)for*any*drift, which a term independent ofz^\\hat\{z\}would spoil\. With the Itô Jacobian absorbed into the normalization, the drift nonlinearities generate only single\-response verticesi​z^⋅fint\\mathrm\{i\}\\,\\hat\{z\}\\cdot f\_\{\\mathrm\{int\}\}, cf\. the kernel\-level statement in[Eq\.3\.23](https://arxiv.org/html/2608.12438#S3.E23)\. Second, no vertex carries more than two response legs\. The response fieldz^\\hat\{z\}enters quadratically only through the noise term12​g2​‖z^‖2\\tfrac\{1\}\{2\}g^\{2\}\\\|\\hat\{z\}\\\|^\{2\}, so for the state\-independent diffusion considered here the sole two\-response object is the free noise propagator\. Multiplicative noiseg⁡\(z\)g\(z\)would promote this to azz\-dependent two\-response vertex, and non\-Gaussian noise to higher powers ofz^\\hat\{z\}\.

G⁡\(t,t′\)G\(t,t^\{\\prime\}\)C⁡\(t,t′\)∝g2C\(t,t^\{\\prime\}\)\\propto g^\{2\}z^\\hat\{z\}zzzzzz⊗\\otimesz^\\hat\{z\}�​f\\delta\\\!fFigure 2:Diagrammatic elements of the MSRJD expansion, drawn for representative vertex orders\. \(a\) the response propagatorG⁡\(t,t′\)G\(t,t^\{\\prime\}\)of[Eq\.3\.31](https://arxiv.org/html/2608.12438#S3.E31)\. \(b\) the noise\-induced correlationC⁡\(t,t′\)∝g2C\(t,t^\{\\prime\}\)\\propto g^\{2\}of[Eq\.3\.33](https://arxiv.org/html/2608.12438#S3.E33)\. \(c\) a vertex from[Eq\.3\.26](https://arxiv.org/html/2608.12438#S3.E26), here the quarticz^⋅z3\\hat\{z\}\\cdot z^\{3\}of a cubic drift\. \(d\) the one\-point insertioni​z^⋅�​f\\mathrm\{i\}\\hat\{z\}\\cdot\\delta\\\!fof an imperfect learned score\. \(e\) a tree\-level contribution of a single vertex\. \(f\) an inter\-vertex loop, drawn with the cubic verticesz^⋅z2\\hat\{z\}\\cdot z^\{2\}of a quadratic drift\. \(g\) a tadpole\.The expansion in[Eq\.3\.19](https://arxiv.org/html/2608.12438#S3.E19)is now a sum over diagrams built from response propagators, noise correlations, and the vertices of[Eq\.3\.26](https://arxiv.org/html/2608.12438#S3.E26)\. Tree diagrams involve only response propagators and describe deterministic transport along classical trajectories, while loop diagrams contain at least onezz–zzcontraction and carry powers ofg2g^\{2\}\. In the deterministic limitg⁡\(t\)→0g\(t\)\\to 0the correlations vanish and the expansion truncates at tree level, while at finite diffusion strength the loops contribute the fluctuation corrections that diffusion models sample directly\. Three counting rules fix which diagrams survive at a given order\. Every vertex carries exactly onez^\\hat\{z\}leg, response fields never contract with each other, and each closedzz–zzcontraction costs one power ofg2g^\{2\}\. Together these fix the order of a diagram before any integral is evaluated, while the response propagators fix its causal support\. The two examples below apply this to the affine drift, where the expansion terminates, and to a cubic drift, where the first correction appears\.

The propagators are the Gaussian averages in[Eq\.3\.17](https://arxiv.org/html/2608.12438#S3.E17)over trajectories started from the prior draw with the data endpoint free, and the endpoint delta of[Eq\.3\.18](https://arxiv.org/html/2608.12438#S3.E18)enters those averages as one more insertion\. All propagators of the theory are moments and derivatives of the transition kernel\. The first moment of[Eq\.2\.20](https://arxiv.org/html/2608.12438#S2.E20)is the conditional mean, and its derivative with respect to the initial state defines the*state\-transition Jacobian*,

�\(t\|z′,t′\)≡∫dzzK\(z,t∣z′,t′\)andU\(t,t′\)≡∂�​\(t\|z′,t′\)∂z′,\\displaystyle\\mu\(t\\,\|\\,z^\{\\prime\},t^\{\\prime\}\)\\equiv\\int\\mathrm\{d\}z\\;z\\,K\(z,t\\mid z^\{\\prime\},t^\{\\prime\}\)\\qquad\\text\{and\}\\qquad U\(t,t^\{\\prime\}\)\\equiv\\frac\{\\partial\\mu\(t\\,\|\\,z^\{\\prime\},t^\{\\prime\}\)\}\{\\partial z^\{\\prime\}\}\\;,\(3\.27\)the average transport of a state fromt′t^\{\\prime\}tottand its sensitivity to the starting point\. In the deterministic limitUUis the Jacobian of the flow map�\\Phi, for any drift\. For an affine drift,UUis independent ofz′z^\{\\prime\}and reduces to the state\-transition matrix of the linear dynamics,

∂tU=A⁡\(t\)​UwithU⁡\(t′,t′\)=IandU⁡\(t,t′\)=M⁡\(t\)​M​\(t′\)−1,\\displaystyle\\partial\_\{t\}U=A\(t\)\\,U\\qquad\\text\{with\}\\qquad U\(t^\{\\prime\},t^\{\\prime\}\)=\\mdmathbb\{I\}\\qquad\\text\{and\}\\qquad U\(t,t^\{\\prime\}\)=M\(t\)\\,M\(t^\{\\prime\}\)^\{\-1\}\\;,\(3\.28\)whereM⁡\(t\)=U⁡\(t,0\)M\(t\)=U\(t,0\)solves the same equation from the origin, cf\.[Eq\.2\.44](https://arxiv.org/html/2608.12438#S2.E44)\. Perturbing the drift instead of the state defines the*response*\. Since the generative update subtracts the drift, cf\.[Eq\.2\.11](https://arxiv.org/html/2608.12438#S2.E11), a shiftb⁡\(t′\)→b⁡\(t′\)\+�​b​\(t′\)b\(t^\{\\prime\}\)\\to b\(t^\{\\prime\}\)\+\\delta b\(t^\{\\prime\}\)acts as a state kick−�​b​\(t′\)​d​t′\-\\delta b\(t^\{\\prime\}\)\\,\\mathrm\{d\}t^\{\\prime\}that is subsequently transported byUU,

G⁡\(t,t′\)≡�​⟨z⁡\(t\)⟩0�​b​\(t′\)=−�​\(t′−t\)​U​\(t,t′\),\\displaystyle G\(t,t^\{\\prime\}\)\\equiv\\frac\{\\delta\\langle z\(t\)\\rangle\_\{0\}\}\{\\delta b\(t^\{\\prime\}\)\}=\-\\,\\theta\(t^\{\\prime\}\-t\)\\,U\(t,t^\{\\prime\}\)\\;,\(3\.29\)causal along the sampling direction\. For a nonlinear drift the transported Jacobian becomes state dependent\. At equal times, the second cumulant ofK\(⋅,t∣z1,1\)K\(\\,\\cdot\\,,t\\mid z\_\{1\},1\), and across different times, the connected moment of the joint law, which by[Eq\.2\.21](https://arxiv.org/html/2608.12438#S2.E21)is a product of kernels,

⟨z\(t\)z\(t′\)⊤⟩0=∫dydy′yy′⁣⊤K\(y,t∣y′,t′\)K\(y′,t′∣z1,1\)fort<t′\.\\displaystyle\\big\\langle z\(t\)\\,z\(t^\{\\prime\}\)^\{\\\!\\top\}\\big\\rangle\_\{0\}=\\int\\mathrm\{d\}y\\,\\mathrm\{d\}y^\{\\prime\}\\;y\\,y^\{\\prime\\top\}\\,K\(y,t\\mid y^\{\\prime\},t^\{\\prime\}\)\\,K\(y^\{\\prime\},t^\{\\prime\}\\mid z\_\{1\},1\)\\qquad\\text\{for\}\\qquad t<t^\{\\prime\}\\;\.\(3\.30\)Integrals of the kernel thus give statistics, derivatives give responses, and products give multi\-time structure\. The response and the connected correlation can also be expressed as expectation values of[Eq\.3\.17](https://arxiv.org/html/2608.12438#S3.E17):

1. 1\.The first is the*response propagator*, G⁡\(t,t′\)≡⟨z⁡\(t\)​i​z^​\(t′\)⊤⟩0\.\\displaystyle G\(t,t^\{\\prime\}\)\\equiv\\big\\langle z\(t\)\\,\\mathrm\{i\}\\,\\hat\{z\}\(t^\{\\prime\}\)^\{\\\!\\top\}\\big\\rangle\_\{0\}\\;\.\(3\.31\)Shifting the drift byb⁡\(t\)→b⁡\(t\)\+�​b​\(t\)b\(t\)\\to b\(t\)\+\\delta b\(t\)changes the mean path by �​⟨z⁡\(t\)⟩=∫01d​t′​G​\(t,t′\)​�​b​\(t′\)\.\\displaystyle\\delta\\langle z\(t\)\\rangle=\\int\_\{0\}^\{1\}\\mathrm\{d\}t^\{\\prime\}\\;G\(t,t^\{\\prime\}\)\\,\\delta b\(t^\{\\prime\}\)\\;\.\(3\.32\)Causality refers to the direction in which the generative dynamics is integrated, fromt=1t=1towardt=0t=0\. A drift perturbation at timet′t^\{\\prime\}influences the path only at later stages of sampling,i\.e\.at smallertt, so thatG⁡\(t,t′\)=0G\(t,t^\{\\prime\}\)=0fort\>t′t\>t^\{\\prime\}and response propagators are inherently ordered along the sampling clocks=1−ts=1\-t\.
2. 2\.The second is the*noise\-induced correlation*, C⁡\(t,t′\)≡⟨z⁡\(t\)​z​\(t′\)⊤⟩0−⟨z⁡\(t\)⟩0​⟨z⁡\(t′\)⟩0⊤,\\displaystyle C\(t,t^\{\\prime\}\)\\equiv\\big\\langle z\(t\)\\,z\(t^\{\\prime\}\)^\{\\\!\\top\}\\big\\rangle\_\{0\}\-\\big\\langle z\(t\)\\big\\rangle\_\{0\}\\big\\langle z\(t^\{\\prime\}\)\\big\\rangle\_\{0\}^\{\\\!\\top\}\\;,\(3\.33\)the connected part of the two\-point function\. The subtraction of the means is essential, as the mean⟨z⁡\(t\)⟩0\\langle z\(t\)\\rangle\_\{0\}is the free classical trajectory started atz1z\_\{1\}which does not vanish\. The connected part isolates the noise\-induced covariance, which arises entirely from the quadratic noise term12​g2​\(t\)​‖z^‖2\\tfrac\{1\}\{2\}g^\{2\}\(t\)\\\|\\hat\{z\}\\\|^\{2\}in the free action, which vanishes in the deterministic limitg⁡\(t\)→0g\(t\)\\to 0\. The response propagatorG⁡\(t,t′\)G\(t,t^\{\\prime\}\)needs no such subtraction because⟨z^⟩0=0\\langle\\hat\{z\}\\rangle\_\{0\}=0, and no independent⟨z^​z^⟩0\\langle\\hat\{z\}\\,\\hat\{z\}\\rangle\_\{0\}propagator exists\.

#### Example: the affine drift

The first example switches the interaction off,

fint​\(z,t\)=0andfrev​\(z,t\)=A⁡\(t\)​z\+b⁡\(t\),\\displaystyle f\_\{\\mathrm\{int\}\}\(z,t\)=0\\qquad\\text\{and\}\\qquad f\_\{\\mathrm\{rev\}\}\(z,t\)=A\(t\)\\,z\+b\(t\)\\;,\(3\.34\)so every�\(m\)\\lambda^\{\(m\)\}of[Eq\.3\.26](https://arxiv.org/html/2608.12438#S3.E26)vanishes and no vertex exists\. The free action is quadratic, so both propagators follow in closed form from the Gaussian averages in[Eq\.3\.17](https://arxiv.org/html/2608.12438#S3.E17)\. The response propagator inverts the operator∂t−A\\partial\_\{t\}\-Athat couplesz^\\hat\{z\}tozzthroughi​z^⋅\(z˙−A​z\)\\mathrm\{i\}\\,\\hat\{z\}\\cdot\(\\dot\{z\}\-Az\),

\(∂t−A\(t\)\)G\(t,t′\)=�\(t−t′\)I\.\\displaystyle\(\\partial\_\{t\}\-A\(t\)\)\\,G\(t,t^\{\\prime\}\)=\\delta\(t\-t^\{\\prime\}\)\\,\\mdmathbb\{I\}\\;\.\(3\.35\)With the causal boundary condition along the sampling direction, the solution is precisely[Eq\.3\.29](https://arxiv.org/html/2608.12438#S3.E29)\. The noise term of[Eq\.3\.14](https://arxiv.org/html/2608.12438#S3.E14)enters the correlation dressed by two response propagators,

C⁡\(t,t′\)=∫01d​u​G​\(t,u\)​g2​\(u\)​G​\(t′,u\)⊤=∫max⁡\(t,t′\)1d​u​g2​\(u\)​U​\(t,u\)​U​\(t′,u\)⊤,\\displaystyle C\(t,t^\{\\prime\}\)=\\int\_\{0\}^\{1\}\\mathrm\{d\}u\\;G\(t,u\)\\,g^\{2\}\(u\)\\,G\(t^\{\\prime\},u\)^\{\\\!\\top\}=\\int\_\{\\max\(t,t^\{\\prime\}\)\}^\{1\}\\mathrm\{d\}u\\;g^\{2\}\(u\)\\,U\(t,u\)\\,U\(t^\{\\prime\},u\)^\{\\\!\\top\}\\;,\(3\.36\)using[Eq\.3\.29](https://arxiv.org/html/2608.12438#S3.E29)in the last step\. The free kernel then follows from the endpoint insertion in three explicit steps\. First, we write the delta in[Eq\.3\.18](https://arxiv.org/html/2608.12438#S3.E18)in its Fourier representation,

�\(z\(0\)−z0\)=∫d​k\(2​�\)dei​k⋅\(z⁡\(0\)−z0\)=∫d​k\(2​�\)de−ik⋅z0ei​k⋅z⁡\(0\),\\displaystyle\\delta\\big\(z\(0\)\-z\_\{0\}\\big\)=\\int\\frac\{\\mathrm\{d\}k\}\{\(2\\pi\)^\{d\}\}\\;e^\{\\mathrm\{i\}\\,k\\cdot\(z\(0\)\-z\_\{0\}\)\}=\\int\\frac\{\\mathrm\{d\}k\}\{\(2\\pi\)^\{d\}\}\\;e^\{\-\\mathrm\{i\}\\,k\\cdot z\_\{0\}\}\\,e^\{\\mathrm\{i\}\\,k\\cdot z\(0\)\}\\;,\(3\.37\)with integration variablekk\. Inside the free average,z⁡\(0\)z\(0\)is the fluctuating endpoint of the path, whereasz0z\_\{0\}andkkare external parameters\. By linearity of the average, all path\-independent factors pull out,

K\(0\)\(z0,0∣z1,1\)=⟨�\(z\(0\)−z0\)⟩0=∫d​k\(2​�\)de−ik⋅z0⟨ei​k⋅z⁡\(0\)⟩0,\\displaystyle K^\{\(0\)\}\(z\_\{0\},0\\mid z\_\{1\},1\)=\\big\\langle\\delta\\big\(z\(0\)\-z\_\{0\}\\big\)\\big\\rangle\_\{0\}=\\int\\frac\{\\mathrm\{d\}k\}\{\(2\\pi\)^\{d\}\}\\;e^\{\-\\mathrm\{i\}\\,k\\cdot z\_\{0\}\}\\,\\big\\langle e^\{\\mathrm\{i\}\\,k\\cdot z\(0\)\}\\big\\rangle\_\{0\}\\;,\(3\.38\)and the remaining average is the characteristic function of the endpoint state\. Second, we evaluate the characteristic function\. In the free ensemble the mean path is the classical trajectory\. Writingz⁡\(t\)=z⋆​\(t\)\+�​\(t\)z\(t\)=z\_\{\\star\}\(t\)\+\\eta\(t\)withz˙⋆=A⁡\(t\)​z⋆\+b⁡\(t\)\\dot\{z\}\_\{\\star\}=A\(t\)\\,z\_\{\\star\}\+b\(t\)andz⋆​\(1\)=z1z\_\{\\star\}\(1\)=z\_\{1\}, the affine terms cancel and the free action becomes homogeneous and quadratic in the fluctuations,

S0​\[z⋆\+�,z^\]=∫01d​t​\[i​z^⋅\(�˙−A⁡\(t\)​�\)\+12​g2​\(t\)​‖z^‖2\],\\displaystyle S\_\{0\}\[z\_\{\\star\}\+\\eta,\\hat\{z\}\]=\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\left\[\\mathrm\{i\}\\,\\hat\{z\}\\cdot\\big\(\\dot\{\\eta\}\-A\(t\)\\,\\eta\\big\)\+\\frac\{1\}\{2\}\\,g^\{2\}\(t\)\\,\\\|\\hat\{z\}\\\|^\{2\}\\right\]\\;,\(3\.39\)with�​\(1\)=0\\eta\(1\)=0inherited from the pinned prior end and�​\(0\)\\eta\(0\)free\. Three properties of this ensemble carry the evaluation\. The action has no linear term, so the fluctuation has zero mean,⟨�​\(t\)⟩0=0\\langle\\eta\(t\)\\rangle\_\{0\}=0, and the endpoint average is the classical arrival point,⟨z⁡\(0\)⟩0=z⋆​\(0\)\\langle z\(0\)\\rangle\_\{0\}=z\_\{\\star\}\(0\)\. The action is invariant under\(�,z^\)→\(−�,−z^\)\(\\eta,\\hat\{z\}\)\\to\(\-\\eta,\-\\hat\{z\}\), so all odd moments of�\\etavanish\. And since the action is quadratic, the even moments obey Wick’s theorem, decomposing into sums over pairwise contractions with the equal\-time covariance given by the connected correlator in[Eq\.3\.33](https://arxiv.org/html/2608.12438#S3.E33),i\.e\.⟨�​\(0\)​�​\(0\)⊤⟩0=C⁡\(0,0\)\\langle\\eta\(0\)\\,\\eta\(0\)^\{\\\!\\top\}\\rangle\_\{0\}=C\(0,0\)\. The endpoint phase then splits into a deterministic factor and a fluctuation average,

⟨ei​k⋅z⁡\(0\)⟩0=ei​k⋅z⋆​\(0\)​⟨ei​k⋅�​\(0\)⟩0\.\\displaystyle\\big\\langle e^\{\\mathrm\{i\}\\,k\\cdot z\(0\)\}\\big\\rangle\_\{0\}=e^\{\\mathrm\{i\}\\,k\\cdot z\_\{\\star\}\(0\)\}\\,\\big\\langle e^\{\\mathrm\{i\}\\,k\\cdot\\eta\(0\)\}\\big\\rangle\_\{0\}\\;\.\(3\.40\)Expanding the remaining exponential and averaging term by term, the odd moments drop out, while each even moment of the scalark⋅�​\(0\)k\\cdot\\eta\(0\)Wick\-decomposes into\(2​m−1\)\!\!\(2m\-1\)\!\!pairings, each contributing one factor ofk⊤​C​\(0,0\)​kk^\{\\\!\\top\}C\(0,0\)\\,k,

⟨ei​k⋅�​\(0\)⟩0=∑m=0∞\(i\)2​m\(2​m\)\!​⟨\(k⋅�​\(0\)\)2​m⟩0=∑m=0∞\(−1\)m​\(2​m−1\)\!\!\(2​m\)\!​\(k⊤​C​\(0,0\)​k\)m\.\\displaystyle\\big\\langle e^\{\\mathrm\{i\}\\,k\\cdot\\eta\(0\)\}\\big\\rangle\_\{0\}=\\sum\_\{m=0\}^\{\\infty\}\\frac\{\(\\mathrm\{i\}\)^\{2m\}\}\{\(2m\)\!\}\\,\\big\\langle\\big\(k\\cdot\\eta\(0\)\\big\)^\{2m\}\\big\\rangle\_\{0\}=\\sum\_\{m=0\}^\{\\infty\}\\frac\{\(\-1\)^\{m\}\\,\(2m\-1\)\!\!\}\{\(2m\)\!\}\\,\\big\(k^\{\\\!\\top\}C\(0,0\)\\,k\\big\)^\{m\}\\;\.\(3\.41\)Using\(2​m\)\!=2m​m\!​\(2​m−1\)\!\!\(2m\)\!=2^\{m\}\\,m\!\\,\(2m\-1\)\!\!, the series resums into an exponential,

⟨ei​k⋅�​\(0\)⟩0=∑m=0∞1m\!​\(−12​k⊤​C​\(0,0\)​k\)m=exp⁡\[−12​k⊤​C​\(0,0\)​k\],\\displaystyle\\big\\langle e^\{\\mathrm\{i\}\\,k\\cdot\\eta\(0\)\}\\big\\rangle\_\{0\}=\\sum\_\{m=0\}^\{\\infty\}\\frac\{1\}\{m\!\}\\left\(\-\\frac\{1\}\{2\}\\,k^\{\\\!\\top\}C\(0,0\)\\,k\\right\)^\{\\\!m\}=\\exp\\\!\\Big\[\-\\tfrac\{1\}\{2\}\\,k^\{\\\!\\top\}C\(0,0\)\\,k\\Big\]\\;,\(3\.42\)and, combining[Eqs\.3\.40](https://arxiv.org/html/2608.12438#S3.E40)and[3\.42](https://arxiv.org/html/2608.12438#S3.E42), the characteristic function is exactly quadratic inkk,

⟨ei​k⋅z⁡\(0\)⟩0=exp⁡\[i​k⋅z⋆​\(0\)−12​k⊤​C​\(0,0\)​k\]\.\\displaystyle\\big\\langle e^\{\\mathrm\{i\}\\,k\\cdot z\(0\)\}\\big\\rangle\_\{0\}=\\exp\\\!\\Big\[\\mathrm\{i\}\\,k\\cdot z\_\{\\star\}\(0\)\-\\tfrac\{1\}\{2\}\\,k^\{\\\!\\top\}C\(0,0\)\\,k\\Big\]\\;\.\(3\.43\)Third, inserting[Eq\.3\.43](https://arxiv.org/html/2608.12438#S3.E43)into[Eq\.3\.38](https://arxiv.org/html/2608.12438#S3.E38)leaves a Gaussiankk\-integral, which converges sinceC⁡\(0,0\)C\(0,0\)is positive definite for nonvanishing diffusion, and is evaluated by completing the square,

K\(0\)\(z0,0∣z1,1\)\\displaystyle K^\{\(0\)\}\(z\_\{0\},0\\mid z\_\{1\},1\)=∫d​k\(2​�\)dexp\[−ik⋅\(z0−z⋆\(0\)\)−12k⊤C\(0,0\)k\]\\displaystyle=\\int\\frac\{\\mathrm\{d\}k\}\{\(2\\pi\)^\{d\}\}\\;\\exp\\\!\\Big\[\-\\mathrm\{i\}\\,k\\cdot\\big\(z\_\{0\}\-z\_\{\\star\}\(0\)\\big\)\-\\tfrac\{1\}\{2\}\\,k^\{\\\!\\top\}C\(0,0\)\\,k\\Big\]=exp⁡\[−12​\(z0−z⋆​\(0\)\)⊤​C​\(0,0\)−1​\(z0−z⋆​\(0\)\)\]\(2​�\)d​detC⁡\(0,0\)\\displaystyle=\\frac\{\\exp\\\!\\Big\[\-\\tfrac\{1\}\{2\}\\,\\big\(z\_\{0\}\-z\_\{\\star\}\(0\)\\big\)^\{\\\!\\top\}C\(0,0\)^\{\-1\}\\big\(z\_\{0\}\-z\_\{\\star\}\(0\)\\big\)\\Big\]\}\{\\sqrt\{\(2\\pi\)^\{d\}\\det C\(0,0\)\}\}=𝒩⁡\(z0,z⋆​\(0\),C⁡\(0,0\)\),\\displaystyle=\\mathcal\{N\}\\big\(z\_\{0\};\\,z\_\{\\star\}\(0\),\\,C\(0,0\)\\big\)\\;,\(3\.44\)the normal density inz0z\_\{0\}, centered on the endpoint of the classical trajectoryz⋆​\(0\)z\_\{\\star\}\(0\), and broadened by the accumulated noiseC⁡\(0,0\)C\(0,0\)\. In the deterministic limit the covariance vanishes andK\(0\)→�​\(z0−z⋆​\(0\)\)K^\{\(0\)\}\\to\\delta\\big\(z\_\{0\}\-z\_\{\\star\}\(0\)\\big\), recovering the normalizing\-flow transport of[Section2\.3](https://arxiv.org/html/2608.12438#S2.SS3)\. The result is exact\. With every coupling�\(m\)\\lambda^\{\(m\)\}equal to zero there is no vertex to insert, onlyK\(0\)K^\{\(0\)\}contributes to[Fig\.1](https://arxiv.org/html/2608.12438#S3.F1), and the diagrammatic series terminates after its first term\.

#### Example: a cubic drift

The second example keeps the leading odd nonlinearity,

fint​\(z,t\)=�​\(t\)​z3,\\displaystyle f\_\{\\mathrm\{int\}\}\(z,t\)=\\lambda\(t\)\\,z^\{3\}\\;,\(3\.45\)so that�\(3\)≡�​\(t\)\\lambda^\{\(3\)\}\\equiv\\lambda\(t\)is the only nonvanishing coupling in[Eq\.3\.26](https://arxiv.org/html/2608.12438#S3.E26), whose indices we suppress for readability\. To first order in�\\lambda, the response propagator is corrected by one insertion of the interaction,

G⁡\(t,t′\)\\displaystyle G\(t,t^\{\\prime\}\)→G⁡\(t,t′\)\+G1​\(t,t′\)\\displaystyle\\to G\(t,t^\{\\prime\}\)\+G\_\{1\}\(t,t^\{\\prime\}\)withG1​\(t,t′\)\\displaystyle\\text\{with\}\\qquad G\_\{1\}\(t,t^\{\\prime\}\)=−⟨z⁡\(t\)​i​z^​\(t′\)​Sint​\[z,z^\]⟩0=∫01d​s​�​\(s\)​⟨z⁡\(t\)​i​z^​\(t′\)​i​z^​\(s\)​z​\(s\)3⟩0,\\displaystyle=\-\\big\\langle z\(t\)\\,\\mathrm\{i\}\\hat\{z\}\(t^\{\\prime\}\)\\,S\_\{\\mathrm\{int\}\}\[z,\\hat\{z\}\]\\big\\rangle\_\{0\}=\\int\_\{0\}^\{1\}\\mathrm\{d\}s\\;\\lambda\(s\)\\,\\big\\langle z\(t\)\\,\\mathrm\{i\}\\hat\{z\}\(t^\{\\prime\}\)\\;\\mathrm\{i\}\\hat\{z\}\(s\)\\,z\(s\)^\{3\}\\big\\rangle\_\{0\}\\;,\(3\.46\)where no disconnected subtraction is needed since⟨Sint⟩0=0\\langle S\_\{\\mathrm\{int\}\}\\rangle\_\{0\}=0, as established below[Eq\.3\.21](https://arxiv.org/html/2608.12438#S3.E21)\. To evaluate the average, we split the latent field into the free mean and its fluctuation,z=z⋆\+�z=z\_\{\\star\}\+\\etawith⟨�⟩0=0\\langle\\eta\\rangle\_\{0\}=0as in[Eq\.3\.39](https://arxiv.org/html/2608.12438#S3.E39), and expand the vertex,

z​\(s\)3=z⋆​\(s\)3\+3​z⋆​\(s\)2​�​\(s\)\+3​z⋆​\(s\)​�​\(s\)2\+�​\(s\)3\.\\displaystyle z\(s\)^\{3\}=z\_\{\\star\}\(s\)^\{3\}\+3\\,z\_\{\\star\}\(s\)^\{2\}\\,\\eta\(s\)\+3\\,z\_\{\\star\}\(s\)\\,\\eta\(s\)^\{2\}\+\\eta\(s\)^\{3\}\\;\.\(3\.47\)The external backgroundz⋆​\(t\)z\_\{\\star\}\(t\)drops immediately, as in every pairing of the remaining fields one response field is left uncontracted and⟨z^⟩0=0\\langle\\hat\{z\}\\rangle\_\{0\}=0\. Wick’s theorem then reduces each term to pairings of\{�​\(t\),i​z^​\(t′\),i​z^​\(s\),�​\(s\)k\}\\\{\\eta\(t\),\\,\\mathrm\{i\}\\hat\{z\}\(t^\{\\prime\}\),\\,\\mathrm\{i\}\\hat\{z\}\(s\),\\,\\eta\(s\)^\{k\}\\\}built fromGGandCC, and most candidates vanish identically\. Terms with an odd number of fluctuation fields drop by the\(�,z^\)→\(−�,−z^\)\(\\eta,\\hat\{z\}\)\\to\(\-\\eta,\-\\hat\{z\}\)symmetry of the free action, response fields never pair with each other,⟨z^​z^⟩0=0\\langle\\hat\{z\}\\,\\hat\{z\}\\rangle\_\{0\}=0, and any pairing ofz^​\(s\)\\hat\{z\}\(s\)with�​\(s\)\\eta\(s\)is an equal\-time response,G⁡\(s,s\)=0G\(s,s\)=0in the Itô convention\. This forcesz^​\(s\)\\hat\{z\}\(s\)to pair with the external�​\(t\)\\eta\(t\), producingG⁡\(t,s\)G\(t,s\), and leaves exactly two surviving patterns\. From the term3​z⋆2​�​\(s\)3z\_\{\\star\}^\{2\}\\,\\eta\(s\), the single fluctuation pairs withz^​\(t′\)\\hat\{z\}\(t^\{\\prime\}\), while from�​\(s\)3\\eta\(s\)^\{3\}one of the three legs pairs withz^​\(t′\)\\hat\{z\}\(t^\{\\prime\}\)and the remaining two close into the equal\-time loopC⁡\(s,s\)C\(s,s\)\. Collecting both,

G1​\(t,t′\)=3​∫01d​s​�​\(s\)​\[z⋆​\(s\)2\+C⁡\(s,s\)\]​G​\(t,s\)​G​\(s,t′\),\\displaystyle G\_\{1\}\(t,t^\{\\prime\}\)=3\\int\_\{0\}^\{1\}\\mathrm\{d\}s\\;\\lambda\(s\)\\,\\Big\[\\,z\_\{\\star\}\(s\)^\{2\}\+C\(s,s\)\\,\\Big\]\\,G\(t,s\)\\,G\(s,t^\{\\prime\}\)\\;,\(3\.48\)with the causal supportt<s<t′t<s<t^\{\\prime\}enforced by the two response propagators\. The two contributions realize the two elementary topologies of[Fig\.2](https://arxiv.org/html/2608.12438#S3.F2)and are drawn in[Fig\.3](https://arxiv.org/html/2608.12438#S3.F3)\.

G1​\(t,t′\)G\_\{1\}\(t,t^\{\\prime\}\)==z⋆z\_\{\\star\}z⋆z\_\{\\star\}G⁡\(t,s\)G\(t,s\)G⁡\(s,t′\)G\(s,t^\{\\prime\}\)\+\+C⁡\(s,s\)C\(s,s\)G⁡\(t,s\)G\(t,s\)G⁡\(s,t′\)G\(s,t^\{\\prime\}\)Figure 3:The two contributions to[Eq\.3\.48](https://arxiv.org/html/2608.12438#S3.E48), assembled from the elements of[Fig\.2](https://arxiv.org/html/2608.12438#S3.F2)\. Both place one vertex at an intermediate timessbetween two response propagators\. On the left the two remainingzzlegs attach to the classical field and give the tree\-level term∝z⋆​\(s\)2\\propto z\_\{\\star\}\(s\)^\{2\}, topology \(e\)\. On the right the same two legs contract with each other into the equal\-time correlationC⁡\(s,s\)∝g2C\(s,s\)\\propto g^\{2\}and give the tadpole, topology \(g\)\. The choice between these two attachments is what separates tree level from one loop\.The background term∝z⋆2\\propto z\_\{\\star\}^\{2\}is of the tree\-level type \(e\), where two vertex legs attach to the classical field, no noise contraction is involved, and the term survives the deterministic limitg→0g\\to 0\. It is a mass insertion, the first order of replacing the affine linearizationA⁡\(t\)A\(t\)by the linearization about the classical trajectory,

Acl​\(t\)=A⁡\(t\)\+3​�​\(t\)​z⋆​\(t\)2,\\displaystyle A\_\{\\mathrm\{cl\}\}\(t\)=A\(t\)\+3\\lambda\(t\)\\,z\_\{\\star\}\(t\)^\{2\}\\;,\(3\.49\)and it is resummed automatically once we expand about the classical path of the full drift, as done in[Section4](https://arxiv.org/html/2608.12438#S4)\. The tadpole term∝C⁡\(s,s\)∝g2\\propto C\(s,s\)\\propto g^\{2\}is the loop element \(g\), where two legs of the same vertex contract with each other, the combinatorial factor33counts the pairings, and the contribution vanishes in the deterministic limit\.

The two examples illustrate the expansion\. The affine drift carries no vertex and the series stops atK\(0\)K^\{\(0\)\}, while a single cubic coupling already produces both elementary topologies at first order, one surviving the deterministic limit and one carrying the leading power ofg2g^\{2\}\. Both follow from the leg count of the vertex and the two ways its remaining legs attach\. Different classes of generative models correspond to distinct truncations of the same expansion\. Normalizing flows retain only tree diagrams, diffusion models evaluate the full series including loops, and conditional flow matching absorbs the fluctuation corrections to the single\-time marginals into an effective drift\.

## 4Loop corrections to deterministic samplers

Deterministic samplers are fast, which is why they are used\. For an exact score the probability\-flow ODE even reproduces the time\-marginals of the diffusion process exactly\. This is a statement about marginals only\. The tree\-level truncation of the stochastic path integral is the ODE generated by the reverse drift itself, and it misses the trajectory\-level and joint statistics of the true reverse process already at exact score\. Once the score is learned, the marginals deviate as well\.

We show that the MSRJD representation organizes the discrepancy as a systematic loop expansion in the diffusion strengthg2g^\{2\}, derive the leading correction in closed form, requiring two auxiliary equations integrated alongside the deterministic trajectory at no stochastic\-sampling cost, and validate it on an exactly solvable model and on nonlinear drifts, including a2424\-dimensional equivariant example\. We further show that an imperfect learned score enters the expansion as a calculable insertion, from which a response\-weighted training objective follows as a corollary\.

The gap between deterministic and stochastic sampling has recently been approached from a path\-integral perspective in Ref\.\[[31](https://arxiv.org/html/2608.12438#bib.bib45)\], where an interpolating parameter between the probability\-flow ODE and the reverse SDE plays the role of Planck’s constant and the negative log\-likelihood is evaluated in a WKB \(semiclassical\) expansion in this parameter\. Our expansion differs in scope and output\. It corrects smooth observables of the generated samples rather than the likelihood, its leading correction reduces to two auxiliary ordinary differential equations, and its hybrid reformulation remains well conditioned in the focusing and defocusing regimes near the data manifold\.

We work with the MSRJD action from[Eq\.3\.12](https://arxiv.org/html/2608.12438#S3.E12)and recall the deterministic limit established in[Section3\.2](https://arxiv.org/html/2608.12438#S3.SS2), whereg⁡\(t\)→0g\(t\)\\to 0and integrating out the response fieldz^\\hat\{z\}enforces the constraintz˙=frev​\(z,t\)\\dot\{z\}=f\_\{\\mathrm\{rev\}\}\(z,t\), and the transition kernel collapses onto the classical trajectory defined by

z˙cl​\(t\)=frev​\(zcl​\(t\),t\)withzcl​\(1\)=z1\.\\displaystyle\\dot\{z\}\_\{\\mathrm\{cl\}\}\(t\)=f\_\{\\mathrm\{rev\}\}\\big\(z\_\{\\mathrm\{cl\}\}\(t\),t\\big\)\\qquad\\text\{with\}\\qquad z\_\{\\mathrm\{cl\}\}\(1\)=z\_\{1\}\\;\.\(4\.1\)This drift ODE is the deterministic sampler singled out by the path integral itself, the continuous normalizing flow of[Section2\.3\.1](https://arxiv.org/html/2608.12438#S2.SS3.SSS1)run with the reverse drift\. Initialized at a prior drawz1z\_\{1\}and integrated tot=0t=0, it returnszcl​\(0\)z\_\{\\mathrm\{cl\}\}\(0\)\. It is not the probability\-flow ODE, which carries the score with coefficient12​g2\\tfrac\{1\}\{2\}g^\{2\}rather thang2g^\{2\}, cf\. the discussion below[Eq\.2\.61](https://arxiv.org/html/2608.12438#S2.E61)\. The offset12​g2​∇z​log⁡p\\tfrac\{1\}\{2\}g^\{2\}\\,\\nabla\_\{z\}\\log pis itself of orderg2g^\{2\}and enters as a calculable insertion of the form derived in[Section4\.2](https://arxiv.org/html/2608.12438#S4.SS2), so corrections relative to either deterministic sampler follow from the same formulas, with the probability\-flow case combining the fluctuation terms below with this tree\-level insertion shift\. If the score enteringfrevf\_\{\\mathrm\{rev\}\}is*exact*, the probability\-flow ODE and the reverse\-time SDE share the same time\-marginalsp⁡\(z,t\)p\(z,t\)by construction, and differ only in their trajectory\-level and joint statistics\. In every practical setting, however, the score is replaced by a learned approximations�s\_\{\\theta\}, the marginals no longer coincide, and the loop expansion becomes an expansion of the*marginal*error as well\. The correction derived below is the leading term in both cases, and the score\-mismatch contribution to the marginals is made explicit in[Section4\.2](https://arxiv.org/html/2608.12438#S4.SS2)\.

### 4\.1The fluctuation action

We expand the latent field about the classical trajectory,

z⁡\(t\)=zcl​\(t\)\+�​\(t\)with�​\(1\)=0,\\displaystyle z\(t\)=z\_\{\\mathrm\{cl\}\}\(t\)\+\\eta\(t\)\\qquad\\text\{with\}\\qquad\\eta\(1\)=0\\;,\(4\.2\)where the boundary condition reflects that the sampler is initialized at a fixed prior draw\. This is the interacting counterpart of the free\-theory expansion aboutz⋆z\_\{\\star\}in[Eq\.3\.39](https://arxiv.org/html/2608.12438#S3.E39), now around the classical trajectory of the full drift\. Substituting[Eq\.4\.2](https://arxiv.org/html/2608.12438#S4.E2)into the MSRJD action and expandingfrevf\_\{\\mathrm\{rev\}\}to first order in�\\eta,

frev​\(zcl\+�,t\)=frev​\(zcl,t\)\+Acl​\(t\)​�​\(t\)\+𝒪⁡\(�2\),Acl​\(t\)≡∇zfrev​\(zcl,t\),\\displaystyle f\_\{\\mathrm\{rev\}\}\(z\_\{\\mathrm\{cl\}\}\+\\eta,t\)=f\_\{\\mathrm\{rev\}\}\(z\_\{\\mathrm\{cl\}\},t\)\+A\_\{\\mathrm\{cl\}\}\(t\)\\,\\eta\(t\)\+\\mathcal\{O\}\(\\eta^\{2\}\)\\;,\\qquad A\_\{\\mathrm\{cl\}\}\(t\)\\equiv\\nabla\_\{z\}f\_\{\\mathrm\{rev\}\}\\big\(z\_\{\\mathrm\{cl\}\},t\\big\)\\;,\(4\.3\)we sort the action by the total power of the fluctuation fields\(�,z^\)\(\\eta,\\hat\{z\}\)\. The field\-independent term is a constant, and the term of first order in the fluctuations vanishes becausezclz\_\{\\mathrm\{cl\}\}solves the classical equation of motion in[Eq\.4\.1](https://arxiv.org/html/2608.12438#S4.E1), the saddle\-point condition\. The leading nontrivial term is therefore the*quadratic*\(Gaussian\) fluctuation action, which is the free fluctuation action from[Eq\.3\.39](https://arxiv.org/html/2608.12438#S3.E39)with the affine kernel replaced by the linearization about the classical path,A→AclA\\to A\_\{\\mathrm\{cl\}\},

S2​\[�,z^\]=∫01d​t​\[i​z^⋅\(�˙−Acl​\(t\)​�\)\+12​g2​\(t\)​‖z^‖2\]\.\\displaystyle S\_\{2\}\[\\eta,\\hat\{z\}\]=\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\left\[\\mathrm\{i\}\\,\\hat\{z\}\\cdot\\big\(\\dot\{\\eta\}\-A\_\{\\mathrm\{cl\}\}\(t\)\\,\\eta\\big\)\+\\tfrac\{1\}\{2\}\\,g^\{2\}\(t\)\\,\\\|\\hat\{z\}\\\|^\{2\}\\right\]\\;\.\(4\.4\)The interaction functionalSint​\[�,z^\]S\_\{\\mathrm\{int\}\}\[\\eta,\\hat\{z\}\]collects all terms of cubic and higher order in�\\eta, so by the counting of[Section3\.3](https://arxiv.org/html/2608.12438#S3.SS3)each additional vertex costs a further factor ofg2g^\{2\}and[Eq\.4\.4](https://arxiv.org/html/2608.12438#S4.E4)is the one\-loop order of the expansion\. For an affine driftSintS\_\{\\mathrm\{int\}\}vanishes and this Gaussian order is exact, while for a nonlinear drift one insertion of the leading cubic term still contributes at the same order through the tadpole contraction of[Fig\.2](https://arxiv.org/html/2608.12438#S3.F2)\(g\), which we include below\.

BecauseS2S\_\{2\}is Gaussian, the fluctuation�\\etahas zero mean and is fully characterized by its equal\-time covariance,

C⁡\(t\)≡⟨�​\(t\)​�​\(t\)⊤⟩S2,\\displaystyle C\(t\)\\equiv\\big\\langle\\eta\(t\)\\,\\eta\(t\)^\{\\\!\\top\}\\big\\rangle\_\{S\_\{2\}\}\\;,\(4\.5\)where the average runs over fluctuations pinned at the prior end,�​\(1\)=0\\eta\(1\)=0, with the endpoint�​\(0\)\\eta\(0\)integrated out, as in[Eq\.3\.17](https://arxiv.org/html/2608.12438#S3.E17)\. The equal\-time covariance is therefore the free correlator withA→AclA\\to A\_\{\\mathrm\{cl\}\},

C⁡\(t\)=∫t1d​u​g2​\(u\)​Ucl​\(t,u\)​Ucl​\(t,u\)⊤,\\displaystyle C\(t\)=\\int\_\{t\}^\{1\}\\mathrm\{d\}u\\;g^\{2\}\(u\)\\,U\_\{\\mathrm\{cl\}\}\(t,u\)\\,U\_\{\\mathrm\{cl\}\}\(t,u\)^\{\\\!\\top\}\\;,\(4\.6\)whereUcl​\(t,u\)U\_\{\\mathrm\{cl\}\}\(t,u\)solves[Eq\.3\.28](https://arxiv.org/html/2608.12438#S3.E28)with the same replacement\. Differentiating[Eq\.4\.6](https://arxiv.org/html/2608.12438#S4.E6)with respect tottby the Leibniz rule,

d​C​\(t\)d​t\\displaystyle\\frac\{\\mathrm\{d\}C\(t\)\}\{\\mathrm\{d\}t\}=−g2​\(t\)​I\+∫t1d​u​g2​\(u\)​\[∂tUcl​Ucl⊤\+Ucl​∂tUcl⊤\]\\displaystyle=\-\\,g^\{2\}\(t\)\\,\\mdmathbb\{I\}\+\\int\_\{t\}^\{1\}\\mathrm\{d\}u\\;g^\{2\}\(u\)\\Big\[\\partial\_\{t\}U\_\{\\mathrm\{cl\}\}\\,U\_\{\\mathrm\{cl\}\}^\{\\\!\\top\}\+U\_\{\\mathrm\{cl\}\}\\,\\partial\_\{t\}U\_\{\\mathrm\{cl\}\}^\{\\\!\\top\}\\Big\]=−g2​\(t\)​I\+Acl​\(t\)​C​\(t\)\+C⁡\(t\)​Acl​\(t\)⊤,\\displaystyle=\-\\,g^\{2\}\(t\)\\,\\mdmathbb\{I\}\+A\_\{\\mathrm\{cl\}\}\(t\)\\,C\(t\)\+C\(t\)\\,A\_\{\\mathrm\{cl\}\}\(t\)^\{\\\!\\top\}\\;,\(4\.7\)whereUcl​\(t,t\)=IU\_\{\\mathrm\{cl\}\}\(t,t\)=\\mdmathbb\{I\}fixes the boundary term and∂tUcl=Acl​Ucl\\partial\_\{t\}U\_\{\\mathrm\{cl\}\}=A\_\{\\mathrm\{cl\}\}U\_\{\\mathrm\{cl\}\}the integrand\. This is a Lyapunov equation\[[23](https://arxiv.org/html/2608.12438#bib.bib57)\]\. Written in components, with summation over repeated latent\-space indices implied here and in the following, and with the derivative taken along the sampling direction of decreasingtt, it reads

−d​Ci​jd​t=−\(Acl,i​k​\(t\)​Ck​j\+Ci​k​Acl,j​k​\(t\)\)\+g2​\(t\)​�i​j,Ci​j​\(1\)=0\.\\displaystyle\-\\frac\{\\mathrm\{d\}C\_\{ij\}\}\{\\mathrm\{d\}t\}=\-\\big\(A\_\{\\mathrm\{cl\},ik\}\(t\)\\,C\_\{kj\}\+C\_\{ik\}\\,A\_\{\\mathrm\{cl\},jk\}\(t\)\\big\)\+g^\{2\}\(t\)\\,\\delta\_\{ij\}\\;,\\qquad C\_\{ij\}\(1\)=0\\;\.\(4\.8\)The two terms of[Eq\.4\.8](https://arxiv.org/html/2608.12438#S4.E8)carry signs for different reasons\. The diffusion term enters with a definite positive sign, because noise variance accumulates along the sampling direction irrespective of the orientation of the time axis\. The drift term, by contrast, changes sign relative to the forward Lyapunov equation because the derivative is taken with respect to the sampling clocks=1−ts=1\-t\. With the boundary conditionC⁡\(1\)=0C\(1\)=0, meaning no fluctuations at the initial prior draw,[Eq\.4\.8](https://arxiv.org/html/2608.12438#S4.E8)integrates to a positive semidefiniteC⁡\(t\)C\(t\)for allt<1t<1, as a covariance must be\.

LetOObe an observable of the generated sample,i\.e\.a function of the endpoint statez⁡\(0\)z\(0\)rather than of the density\. Its expectation over the reverse process is the corresponding moment of the endpoint kernel,

⟨O⟩=∫dz0O\(z0\)K\(z0,0∣z1,1\),\\displaystyle\\langle O\\rangle=\\int\\mathrm\{d\}z\_\{0\}\\;O\(z\_\{0\}\)\\,K\(z\_\{0\},0\\mid z\_\{1\},1\)\\;,\(4\.9\)which we evaluate perturbatively\. Two effects contribute at orderg2g^\{2\}\. First, the Gaussian fluctuations described by[Eq\.4\.4](https://arxiv.org/html/2608.12438#S4.E4)broaden the endpoint distribution aroundzcl​\(0\)z\_\{\\mathrm\{cl\}\}\(0\)with covarianceC⁡\(0\)C\(0\)\. Second, for a nonlinear drift the*mean*of the endpoint is shifted at the same order\. Performing the drift expansion to second order,

frev,i​\(zcl\+�,t\)=frev,i​\(zcl,t\)\+Acl,i​j​�j\+12​∂j∂kfrev,i​\(zcl,t\)​�j​�k\+𝒪⁡\(�3\),\\displaystyle f\_\{\\mathrm\{rev\},i\}\(z\_\{\\mathrm\{cl\}\}\+\\eta,t\)=f\_\{\\mathrm\{rev\},i\}\(z\_\{\\mathrm\{cl\}\},t\)\+A\_\{\\mathrm\{cl\},ij\}\\,\\eta\_\{j\}\+\\tfrac\{1\}\{2\}\\,\\partial\_\{j\}\\partial\_\{k\}f\_\{\\mathrm\{rev\},i\}\(z\_\{\\mathrm\{cl\}\},t\)\\,\\eta\_\{j\}\\eta\_\{k\}\+\\mathcal\{O\}\(\\eta^\{3\}\)\\;,\(4\.10\)the quadratic term is a cubic vertex inSintS\_\{\\mathrm\{int\}\}with onez^\\hat\{z\}and two�\\etalegs\. Read back as a stochastic equation,S2S\_\{2\}together with this vertex is the reverse Langevin dynamics of[Eq\.3\.7](https://arxiv.org/html/2608.12438#S3.E7), linearized about the classical trajectory and carried to second order, so that the fluctuation obeys

�˙i=Acl,i​j​\(t\)​�j\+12​∂j∂kfrev,i​\(zcl,t\)​�j​�k\+g⁡\(t\)​�i,\\displaystyle\\dot\{\\eta\}\_\{i\}=A\_\{\\mathrm\{cl\},ij\}\(t\)\\,\\eta\_\{j\}\+\\tfrac\{1\}\{2\}\\,\\partial\_\{j\}\\partial\_\{k\}f\_\{\\mathrm\{rev\},i\}\(z\_\{\\mathrm\{cl\}\},t\)\\,\\eta\_\{j\}\\eta\_\{k\}\+g\(t\)\\,\\xi\_\{i\}\\;,\(4\.11\)driven by the same white noise�\\xias the reverse process\. Averaging over the Gaussian noise, with⟨�i⟩=0\\langle\\xi\_\{i\}\\rangle=0and⟨�j​�k⟩=Cj​k\\langle\\eta\_\{j\}\\eta\_\{k\}\\rangle=C\_\{jk\}at leading order \(the tadpole contraction of[Fig\.2](https://arxiv.org/html/2608.12438#S3.F2)\(g\)\), gives a closed equation for the shift�​m​\(t\)≡⟨�​\(t\)⟩\\delta m\(t\)\\equiv\\langle\\eta\(t\)\\rangle,

−d​�​mid​t=−Acl,i​j​\(t\)​�​mj−12​∂j∂kfrev,i​\(zcl​\(t\),t\)​Cj​k​\(t\),�​mi​\(1\)=0,\\displaystyle\-\\frac\{\\mathrm\{d\}\\delta m\_\{i\}\}\{\\mathrm\{d\}t\}=\-A\_\{\\mathrm\{cl\},ij\}\(t\)\\,\\delta m\_\{j\}\-\\tfrac\{1\}\{2\}\\,\\partial\_\{j\}\\partial\_\{k\}f\_\{\\mathrm\{rev\},i\}\\big\(z\_\{\\mathrm\{cl\}\}\(t\),t\\big\)\\,C\_\{jk\}\(t\)\\;,\\qquad\\delta m\_\{i\}\(1\)=0\\;,\(4\.12\)written again with respect to the sampling direction\. The mean shift is sourced by the covariance through the curvature of the drift, and it vanishes identically for affine drifts\. ExpandingO⁡\(zcl​\(0\)\+�​\(0\)\)O\(z\_\{\\mathrm\{cl\}\}\(0\)\+\\eta\(0\)\)to second order and averaging, with⟨�​\(0\)⟩=�​m​\(0\)\\langle\\eta\(0\)\\rangle=\\delta m\(0\)and⟨�​\(0\)​�​\(0\)⊤⟩=C⁡\(0\)\\langle\\eta\(0\)\\,\\eta\(0\)^\{\\\!\\top\}\\rangle=C\(0\)at leading order, gives

⟨O⁡\[z⁡\(0\)\]⟩=O⁡\(zcl​\(0\)\)\+∂iO⁡\(zcl​\(0\)\)​�​mi​\(0\)\+12​∂i∂jO⁡\(zcl​\(0\)\)​Ci​j​\(0\)\+𝒪⁡\(g4\),\\displaystyle\\big\\langle O\[z\(0\)\]\\big\\rangle=O\\big\(z\_\{\\mathrm\{cl\}\}\(0\)\\big\)\+\\partial\_\{i\}O\\big\(z\_\{\\mathrm\{cl\}\}\(0\)\\big\)\\,\\delta m\_\{i\}\(0\)\+\\tfrac\{1\}\{2\}\\,\\partial\_\{i\}\\partial\_\{j\}O\\big\(z\_\{\\mathrm\{cl\}\}\(0\)\\big\)\\,C\_\{ij\}\(0\)\+\\mathcal\{O\}\(g^\{4\}\)\\;,\(4\.13\)where the first term is the tree\-level \(deterministic\-sampler\) value and the remaining two terms constitute the one\-loop correction\. Both are linear in the covariance and hence, by[Eq\.4\.8](https://arxiv.org/html/2608.12438#S4.E8), linear ing2g^\{2\}at leading order\.[Equation4\.13](https://arxiv.org/html/2608.12438#S4.E13)is the central formula of this section, the one\-loop sampler correction\. It expresses the leading discrepancy between a deterministic sampler and the true stochastic process through two auxiliary linear equations,[Eqs\.4\.8](https://arxiv.org/html/2608.12438#S4.E8)and[4\.12](https://arxiv.org/html/2608.12438#S4.E12), integrated alongside the classical trajectory\. Propagating the full covariance means carryingd2d^\{2\}components and𝒪⁡\(d\)\\mathcal\{O\}\(d\)Jacobian–vector products per step, so largeddcalls for low\-rank or diagonal structure, observable\-specific adjoints, or stochastic trace estimation\. Equivalently, it states that the endpoint kernel remains Gaussian at this order, with first and second cumulantszcl​\(0\)\+�​m​\(0\)z\_\{\\mathrm\{cl\}\}\(0\)\+\\delta m\(0\)andC⁡\(0\)C\(0\)\.

The numerical tests below use the second momentO⁡\[z\]=‖z‖2O\[z\]=\\\|z\\\|^\{2\}, the lowest\-order observable sensitive to both the mean shift�​m​\(0\)\\delta m\(0\)and the covarianceC⁡\(0\)C\(0\), since a linear observable has vanishing second derivative and reduces[Eq\.4\.13](https://arxiv.org/html/2608.12438#S4.E13)to the mean shift alone\. For an affine drift that shift vanishes as well, so the free\-theory check in[Section4\.3](https://arxiv.org/html/2608.12438#S4.SS3)rests entirely onC⁡\(0\)C\(0\)\. With∂iO=2​zi\\partial\_\{i\}O=2z\_\{i\}and∂i∂jO=2​�i​j\\partial\_\{i\}\\partial\_\{j\}O=2\\,\\delta\_\{ij\}, the correction becomes

⟨‖z⁡\(0\)‖2⟩=‖zcl​\(0\)‖2\+2​zcl,i​\(0\)​�​mi​\(0\)\+Ci​i​\(0\)\+𝒪⁡\(g4\)\.\\displaystyle\\big\\langle\\\|z\(0\)\\\|^\{2\}\\big\\rangle=\\\|z\_\{\\mathrm\{cl\}\}\(0\)\\\|^\{2\}\+2\\,z\_\{\\mathrm\{cl\},i\}\(0\)\\,\\delta m\_\{i\}\(0\)\+C\_\{ii\}\(0\)\+\\mathcal\{O\}\(g^\{4\}\)\\;\.\(4\.14\)

### 4\.2Imperfect scores as insertions

In practice the exact reverse drift is unavailable, since the score enteringfrevf\_\{\\mathrm\{rev\}\}is replaced by a learned approximations�​\(z,t\)s\_\{\\theta\}\(z,t\), and the sampler integrates the driftfrev�=frev\+�​ff\_\{\\mathrm\{rev\}\}^\{\\theta\}=f\_\{\\mathrm\{rev\}\}\+\\delta\\\!fwith

�​f​\(z,t\)=−g2​\(t\)​\(s�​\(z,t\)−∇z​log​p​\(z,t\)\)\.\\displaystyle\\delta\\\!f\(z,t\)=\-\\,g^\{2\}\(t\)\\,\\big\(s\_\{\\theta\}\(z,t\)\-\\nabla\_\{z\}\\log p\(z,t\)\\big\)\\;\.\(4\.15\)In the MSRJD action the mismatch enters linearly, as the additional term

�S\[z,z^\]=−i∫01dtz^\(t\)⋅�f\(z\(t\),t\),\\displaystyle\\Delta S\[z,\\hat\{z\}\]=\-\\,\\mathrm\{i\}\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;\\hat\{z\}\(t\)\\cdot\\delta\\\!f\(z\(t\),t\)\\;,\(4\.16\)a one\-point*insertion*, diagrammatically the crossed dot of[Fig\.2](https://arxiv.org/html/2608.12438#S3.F2)\(d\) attached to a response line\. Its leading effect is a tree\-level shift of the endpoint mean, given by a single contraction of this insertion with the response propagator in[Eq\.3\.31](https://arxiv.org/html/2608.12438#S3.E31),

�​m​\(0\)\\displaystyle\\Delta m\(0\)=∫01d​t​G​\(0,t\)​�​f​\(zcl​\(t\),t\)\+𝒪⁡\(�​f2,g2​�​f\),\\displaystyle=\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;G\(0,t\)\\,\\delta\\\!f\\big\(z\_\{\\mathrm\{cl\}\}\(t\),t\\big\)\+\\mathcal\{O\}\\big\(\\delta\\\!f^\{2\},\\,g^\{2\}\\,\\delta\\\!f\\big\)\\;,withG⁡\(0,t\)\\displaystyle\\text\{with\}\\qquad G\(0,t\)=−M⁡\(0\)​M​\(t\)−1,\\displaystyle=\-\\,M\(0\)\\,M\(t\)^\{\-1\}\\;,\(4\.17\)with the response propagator evaluated along the classical trajectory in terms of the variational solutionM⁡\(t\)M\(t\)of[Eq\.2\.44](https://arxiv.org/html/2608.12438#S2.E44)\. This contribution shifts the learned sampler away from the exact\-score marginals, and it is the leading term of the marginal error\. Because[Eq\.4\.17](https://arxiv.org/html/2608.12438#S4.E17)is linear in�​f\\delta\\\!f, it can also be read in reverse, propagating score\-validation errors into uncertainties on generated observables\. The mismatch is also measurable\. A classifier trained to separate data from model samples at noise levelttreturns the log density ratio of[Eq\.2\.85](https://arxiv.org/html/2608.12438#S2.E85), whose gradient is−�f/g2\-\\delta\\\!f/g^\{2\}, and discriminator guidance\[[38](https://arxiv.org/html/2608.12438#bib.bib10)\]trains one at every noise level to estimate this correction and add it back to the learned score during sampling\. The insertion computed here and the correction measured there are the same object, reached analytically in one case and empirically in the other\.

#### A response\-weighted training objective

[Equation4\.17](https://arxiv.org/html/2608.12438#S4.E17)states which score errors matter, since an error�​s​\(t\)≡s�−∇z​log​p\\delta s\(t\)\\equiv s\_\{\\theta\}\-\\nabla\_\{z\}\\log pcommitted at timettis transported to the endpoint with amplitudeG⁡\(0,t\)​g2​\(t\)G\(0,t\)\\,g^\{2\}\(t\)\. If score errors at different times are approximately uncorrelated, the induced endpoint mean\-squared error is

‖�​m​\(0\)‖2≃∫01d​t​g4​\(t\)​‖G⁡\(0,t\)​�​s​\(t\)‖2,\\displaystyle\\big\\\|\\Delta m\(0\)\\big\\\|^\{2\}\\simeq\\int\_\{0\}^\{1\}\\mathrm\{d\}t\\;g^\{4\}\(t\)\\,\\big\\\|G\(0,t\)\\,\\delta s\(t\)\\big\\\|^\{2\}\\;,\(4\.18\)so that minimizing the*endpoint*error at leading order, after averaging over error directions, corresponds to the denoising objective in[Eq\.2\.60](https://arxiv.org/html/2608.12438#S2.E60)with the weighting

�resp​\(t\)∝g4​\(t\)​‖G⁡\(0,t\)‖2,\\displaystyle\\lambda\_\{\\mathrm\{resp\}\}\(t\)\\;\\propto\\;g^\{4\}\(t\)\\,\\big\\\|G\(0,t\)\\big\\\|^\{2\}\\;,\(4\.19\)the conventional likelihood weightingg2​\(t\)g^\{2\}\(t\)\[[68](https://arxiv.org/html/2608.12438#bib.bib14)\]dressed by the squared response propagator, with∥⋅∥\\\|\\cdot\\\|the Frobenius norm\. For the affine reference theoryG⁡\(0,t\)G\(0,t\)is state independent and available in closed form,‖G⁡\(0,t\)‖=‖U⁡\(0,t\)‖\\\|G\(0,t\)\\\|=\\\|U\(0,t\)\\\|\. For nonlinear drifts it depends on the trajectory, so a usable weighting requires an average over the data flow that we do not evaluate here\. The weighting is also observable aware, and if a specific observableOOis targeted,‖G⁡\(0,t\)‖2\\\|G\(0,t\)\\\|^\{2\}is replaced by‖∇O​\(zcl​\(0\)\)⋅G⁡\(0,t\)‖2\\\|\\nabla O\(z\_\{\\mathrm\{cl\}\}\(0\)\)\\cdot G\(0,t\)\\\|^\{2\}\. Loss weightings for diffusion models are usually chosen empirically and are known to matter in practice\[[37](https://arxiv.org/html/2608.12438#bib.bib34)\]\.[Equation4\.19](https://arxiv.org/html/2608.12438#S4.E19)provides a first\-principles candidate, derived from the propagator structure of the theory\.

The same response propagator appears as the adjoint state in reward fine\-tuning of diffusion and flow models\[[59](https://arxiv.org/html/2608.12438#bib.bib16),[18](https://arxiv.org/html/2608.12438#bib.bib17)\], where it transports the gradient of an external objective back through the sampler\. Here it reweights the score\-matching loss itself, so that training effort follows the sensitivity of the generated sample rather than an imposed reward\. The observable\-aware form parallels goal\-oriented error estimation, where local residuals carry the weight of an adjoint solution that controls the error in a target functional\[[8](https://arxiv.org/html/2608.12438#bib.bib18)\]\.

### 4\.3Numerical validation

We validate the one\-loop sampler correction on three models\. A free theory, where[Eq\.4\.13](https://arxiv.org/html/2608.12438#S4.E13)must reproduce the quadratic observable exactly, a low\-dimensional nonlinear drift, where it must fail by exactly the next order ing2g^\{2\}, and a high\-dimensional equivariant drift, which runs the matrix\-valued machinery in many dimensions\. The code reproducing all three studies is publicly available, see the note on code availability at the end of the paper\.

#### The Ornstein–Uhlenbeck free theory

When the reverse drift is affine,frev​\(z,t\)=A⁡\(t\)​z\+b⁡\(t\)f\_\{\\mathrm\{rev\}\}\(z,t\)=A\(t\)\\,z\+b\(t\), the construction of[Section3\.1](https://arxiv.org/html/2608.12438#S3.SS1)identifies the theory as free\. The action is quadratic,SintS\_\{\\mathrm\{int\}\}vanishes identically, and the mean shift in[Eq\.4\.12](https://arxiv.org/html/2608.12438#S4.E12)vanishes as well since∂j∂kfrev,i=0\\partial\_\{j\}\\partial\_\{k\}f\_\{\\mathrm\{rev\},i\}=0\. The endpoint kernel is then exactly Gaussian, so[Eq\.4\.13](https://arxiv.org/html/2608.12438#S4.E13), which truncates the observable at second order, is exact for observables at most quadratic inzz\. This covers the second moment used throughout, while higher observables pick up Gaussian moments such as⟨�4⟩=3​C2\\langle\\eta^\{4\}\\rangle=3C^\{2\}at orderg4g^\{4\}\. This provides a stringent and fully analytic check of the formalism\.

For constantAAandbbthe classical trajectory in[Eq\.4\.1](https://arxiv.org/html/2608.12438#S4.E1)is a linear relaxation\. Integrating the linear ODEz˙cl=A​zcl\+b\\dot\{z\}\_\{\\mathrm\{cl\}\}=A\\,z\_\{\\mathrm\{cl\}\}\+bfromzcl​\(1\)=z1z\_\{\\mathrm\{cl\}\}\(1\)=z\_\{1\}gives

zcl​\(t\)=eA⁡\(t−1\)​z1\+\(eA⁡\(t−1\)−I\)​A−1​b,\\displaystyle z\_\{\\mathrm\{cl\}\}\(t\)=e^\{A\(t\-1\)\}\\,z\_\{1\}\+\\big\(e^\{A\(t\-1\)\}\-\\mdmathbb\{I\}\\big\)\\,A^\{\-1\}b\\;,\(4\.20\)and evaluating at the data endpointt=0t=0gives

zcl​\(0\)=e−A​z1\+\(I−e−A\)​z⋆withz⋆=−A−1​b,\\displaystyle z\_\{\\mathrm\{cl\}\}\(0\)=e^\{\-A\}\\,z\_\{1\}\+\\big\(\\mdmathbb\{I\}\-e^\{\-A\}\\big\)\\,z\_\{\\star\}\\qquad\\text\{with\}\\qquad z\_\{\\star\}=\-A^\{\-1\}b\\;,\(4\.21\)which relaxes from the prior draw toward the fixed pointz⋆z\_\{\\star\}of the drift\. For constantAAthe state\-transition matrix isUcl​\(0,u\)=e−A​uU\_\{\\mathrm\{cl\}\}\(0,u\)=e^\{\-Au\}, so the covariance integral in[Eq\.4\.6](https://arxiv.org/html/2608.12438#S4.E6)at the endpointt=0t=0evaluates to

C⁡\(0\)=g2​∫01d​u​e−A​u​e−A⊤​u\.\\displaystyle C\(0\)=g^\{2\}\\int\_\{0\}^\{1\}\\mathrm\{d\}u\\;e^\{\-Au\}\\,e^\{\-A^\{\\\!\\top\}u\}\\;\.\(4\.22\)A noise kick injected at sampling timeuubefore the endpoint is transported tot=0t=0by the state\-transition matrixe−A​ue^\{\-Au\}and contributes its propagated variance\.[Equation4\.22](https://arxiv.org/html/2608.12438#S4.E22)is manifestly linear ing2g^\{2\}and positive definite\. The reverse process is in this case an Ornstein–Uhlenbeck process whose endpoint is exactly Gaussian with the mean in[Eq\.4\.21](https://arxiv.org/html/2608.12438#S4.E21)and the covariance in[Eq\.4\.22](https://arxiv.org/html/2608.12438#S4.E22), so[Eq\.4\.14](https://arxiv.org/html/2608.12438#S4.E14)must reproduce the true second moment identically\. We verify this numerically for the reverse driftfrev​\(z\)=A​z\+bf\_\{\\mathrm\{rev\}\}\(z\)=Az\+bin two dimensions, with

A=\(−1\.30\.40\.2−0\.9\),b=\(0\.5−0\.3\),\\displaystyle A=\\begin\{pmatrix\}\-1\.3&\\phantom\{\-\}0\.4\\\\ \\phantom\{\-\}0\.2&\-0\.9\\end\{pmatrix\}\\;,\\qquad b=\\begin\{pmatrix\}\\phantom\{\-\}0\.5\\\\ \-0\.3\\end\{pmatrix\}\\;,\(4\.23\)and a fixed prior drawz1=\(1\.0,0\.5\)z\_\{1\}=\(1\.0,\\,0\.5\)\.

Figure 4:Transport of the second moment⟨‖z‖2⟩\\langle\\\|z\\\|^\{2\}\\ranglealong the sampling clocks=1−ts=1\-tfor the three validation models, at fixed diffusion strengthg2=0\.4g^\{2\}=0\.4\. Upper panels compare the deterministic sampler at tree level with its one\-loop correction against the reference, lower panels give the relative deviation from that reference on a logarithmic scale\. The shaded band gives the precision of the reference itself, statistical for the Euler–Maruyama ensembles of the free and equivariant models and set by grid refinement for the Fokker–Planck solution of the cubic drift\.Although both the trajectory and the covariance are available in the closed forms just derived, we deliberately run the free theory through the same numerical procedure as the interacting examples, as a check of the solver itself\. The one\-loop prediction integrates the classical trajectory in[Eq\.4\.1](https://arxiv.org/html/2608.12438#S4.E1)together with the Lyapunov equation[4\.8](https://arxiv.org/html/2608.12438#S4.E8)in the sampling clock, using a fourth\-order Runge–Kutta scheme with5×1045\\times 10^\{4\}uniform steps\. Since the drift is affine, the mean shift vanishes and the prediction is‖zcl​\(0\)‖2\+Tr​C​\(0\)\\\|z\_\{\\mathrm\{cl\}\}\(0\)\\\|^\{2\}\+\\mathrm\{Tr\}\\,C\(0\)with tree\-level value‖zcl​\(0\)‖2=5\.684\\\|z\_\{\\mathrm\{cl\}\}\(0\)\\\|^\{2\}=5\.684\. The reference value is a direct Euler–Maruyama simulation of the reverse SDE, cf\.[Eq\.2\.11](https://arxiv.org/html/2608.12438#S2.E11), with2×1052\\times 10^\{5\}independent trajectories of2×1032\\times 10^\{3\}steps each, all initialized at the same prior drawz1z\_\{1\}\.[Table2](https://arxiv.org/html/2608.12438#S4.T2)reports, for each diffusion strength, the loop correctionTr​C​\(0\)\\mathrm\{Tr\}\\,C\(0\), the resulting one\-loop prediction, the Monte Carlo reference with its sampling error, and the relative deviation between the two, which is consistent with zero within the quoted errors\.

g2g^\{2\}Tr​C​\(0\)\\mathrm\{Tr\}\\,C\(0\)one\-loopSDErel\. error0\.20\.21\.6601\.6607\.3447\.3447\.345​\(10\)7\.345\(10\)0\.020\.02%0\.40\.43\.3203\.3209\.0049\.0049\.003​\(15\)9\.003\(15\)0\.010\.01%0\.60\.64\.9804\.98010\.66410\.66410\.659​\(19\)10\.659\(19\)0\.040\.04%0\.80\.86\.6416\.64112\.32412\.32412\.315​\(24\)12\.315\(24\)0\.070\.07%1\.01\.08\.3018\.30113\.98413\.98413\.971​\(28\)13\.971\(28\)0\.100\.10%1\.41\.411\.62111\.62117\.30517\.30517\.281​\(37\)17\.281\(37\)0\.140\.14%Table 2:One\-loop prediction for the second moment⟨‖z⁡\(0\)‖2⟩\\langle\\\|z\(0\)\\\|^\{2\}\\rangleversus direct reverse\-SDE simulation for the two\-dimensional Ornstein–Uhlenbeck free theory, with tree\-level value‖zcl​\(0\)‖2=5\.684\\\|z\_\{\\mathrm\{cl\}\}\(0\)\\\|^\{2\}=5\.684and2×1052\\times 10^\{5\}Euler–Maruyama trajectories per row\. Monte Carlo errors are given in parentheses\.Two features of[Table2](https://arxiv.org/html/2608.12438#S4.T2)confirm the analysis\. First, the loop correctionTr​C​\(0\)\\mathrm\{Tr\}\\,C\(0\)is exactly linear ing2g^\{2\}, since a linear fit returns a vanishing intercept to machine precision, consistent with[Eq\.4\.22](https://arxiv.org/html/2608.12438#S4.E22)\. Second, the one\-loop prediction matches the stochastic simulation to a relative deviation below0\.15%0\.15\\,\\%everywhere, within the Monte Carlo sampling noise\. This is the expected outcome for a free theory, where[Eq\.4\.13](https://arxiv.org/html/2608.12438#S4.E13)terminates at one loop, and it validates both the sign structure of the Lyapunov equation[4\.8](https://arxiv.org/html/2608.12438#S4.E8)and the observable formula[4\.13](https://arxiv.org/html/2608.12438#S4.E13)\. The left panels of[Fig\.4](https://arxiv.org/html/2608.12438#S4.F4)show the same statement over the full trajectory atg2=0\.4g^\{2\}=0\.4, where the corrected curve tracks the reference within its own sampling precision while the deterministic sampler falls away\.

#### An interacting drift

For a nonlinear drift,[Eq\.4\.13](https://arxiv.org/html/2608.12438#S4.E13)is the leading term of a nontrivial series, and both one\-loop structures, the covariance and the tadpole mean shift, contribute\. We test this on the one\-dimensional cubic reverse drift

frev​\(z\)=a​z\+�​z3,a=1\.0,�=0\.3,z1=1\.5,\\displaystyle f\_\{\\mathrm\{rev\}\}\(z\)=a\\,z\+\\lambda\\,z^\{3\}\\;,\\qquad a=1\.0\\;,\\quad\\lambda=0\.3\\;,\\quad z\_\{1\}=1\.5\\;,\(4\.24\)for which the sampling dynamicsd​z/d​s=−frev​\(z\)\\mathrm\{d\}z/\\mathrm\{d\}s=\-f\_\{\\mathrm\{rev\}\}\(z\)is globally contracting toward the origin, so that neighbouring trajectories never cross and the fluctuation problem remains well conditioned over a wide range of diffusion strengths, cf\.[Section4\.4](https://arxiv.org/html/2608.12438#S4.SS4)\.

The one\-loop prediction integrates the classical trajectory, the Lyapunov equation[4\.8](https://arxiv.org/html/2608.12438#S4.E8), and the mean\-shift equation[4\.12](https://arxiv.org/html/2608.12438#S4.E12)side by side, three scalar equations in one dimension, again with a fourth\-order Runge–Kutta scheme on5×1045\\times 10^\{4\}steps\. We obtain the exact reference independently of any sampling by solving the Fokker–Planck equation of the reverse process in the sampling clocks=1−ts=1\-t,

∂sp⁡\(z,s\)=∂z\(frev​\(z\)​p​\(z,s\)\)\+12​g2​∂z2p⁡\(z,s\),\\displaystyle\\partial\_\{s\}p\(z,s\)=\\partial\_\{z\}\\big\(f\_\{\\mathrm\{rev\}\}\(z\)\\,p\(z,s\)\\big\)\+\\tfrac\{1\}\{2\}\\,g^\{2\}\\,\\partial\_\{z\}^\{2\}\\,p\(z,s\)\\;,\(4\.25\)with an explicit finite\-difference scheme on a grid of3×1033\\times 10^\{3\}points coveringz∈\[−3,4\.5\]z\\in\[\-3,\\,4\.5\], where doubling the grid resolution changes the quoted moments by less than4×10−64\\times 10^\{\-6\}\. For consistency with the other examples, the table also lists a direct Euler–Maruyama simulation with10610^\{6\}trajectories, which confirms the Fokker–Planck reference within its sampling errors\. Resolving the𝒪⁡\(g4\)\\mathcal\{O\}\(g^\{4\}\)residuals themselves, down to6×10−56\\times 10^\{\-5\}, would require some10910^\{9\}trajectories with correspondingly refined time steps, which is why the deterministic reference anchors the relative errors and the scaling fit\.[Table3](https://arxiv.org/html/2608.12438#S4.T3)lists, for each diffusion strength, the two components of the one\-loop correction of[Eq\.4\.14](https://arxiv.org/html/2608.12438#S4.E14), the covariance broadeningTr​C​\(0\)\\mathrm\{Tr\}\\,C\(0\)and the tadpole mean shift2​zcl​\(0\)​�​m​\(0\)2\\,z\_\{\\mathrm\{cl\}\}\(0\)\\,\\delta m\(0\), together with the resulting one\-loop prediction, the exact Fokker–Planck value, and the relative deviation from it, with and without the tadpole term\.

g2g^\{2\}Tr​C​\(0\)\\mathrm\{Tr\}\\,C\(0\)2​zcl​�​m​\(0\)2\\,z\_\{\\mathrm\{cl\}\}\\,\\delta m\(0\)one\-loopexactSDErel\. errorw/o�​m\\delta m0\.10\.10\.03490\.0349−0\.0070\-0\.00700\.22020\.22020\.22020\.22020\.2200​\(2\)0\.2200\(2\)0\.030\.03%3\.23\.2%0\.20\.20\.06980\.0698−0\.0139\-0\.01390\.24820\.24820\.24800\.24800\.2478​\(2\)0\.2478\(2\)0\.100\.10%5\.75\.7%0\.40\.40\.13970\.1397−0\.0278\-0\.02780\.30410\.30410\.30290\.30290\.3026​\(3\)0\.3026\(3\)0\.420\.42%9\.69\.6%0\.60\.60\.20950\.2095−0\.0417\-0\.04170\.36010\.36010\.35680\.35680\.3565​\(4\)0\.3565\(4\)0\.920\.92%12\.612\.6%0\.80\.80\.27930\.2793−0\.0556\-0\.05560\.41600\.41600\.40950\.40950\.4091​\(5\)0\.4091\(5\)1\.581\.58%15\.215\.2%Table 3:One\-loop sampler correction for the second moment⟨z​\(0\)2⟩\\langle z\(0\)^\{2\}\\rangleversus the exact Fokker–Planck reference for the interacting model, with tree\-level valuezcl​\(0\)2=0\.1923z\_\{\\mathrm\{cl\}\}\(0\)^\{2\}=0\.1923\. The reference is deterministic and stable under grid refinement at the4×10−64\\times 10^\{\-6\}level, and the Euler–Maruyama column confirms it with10610^\{6\}trajectories, where parentheses denote the Monte Carlo uncertainty in the last digits\. The second and third columns are the two components of the one\-loop correction, covariance broadening and tadpole mean shift, while the last column is the relative error when the tadpole term is dropped\.A correctly truncated one\-loop expansion has to leave a residual of orderg4g^\{4\}, and this is what[Table3](https://arxiv.org/html/2608.12438#S4.T3)shows\. The residual, the relative deviation multiplied by the exact value, gives a log–log slope againstg2g^\{2\}of2\.32\.3over the full range and2\.12\.1over the two smallest values, approaching22from above asg2g^\{2\}decreases, with the excess attributable to𝒪⁡\(g6\)\\mathcal\{O\}\(g^\{6\}\)contributions\. Dropping the tadpole degrades the slope to1\.01\.0\. The mean shift is therefore not a refinement but part of the one\-loop order for nonlinear drifts, and the slope measures the order of the leading term left out\. At the largest diffusion strength shown, the one\-loop result reduces a53%53\\,\\%tree\-level error to1\.6%1\.6\\,\\%\. The center panels of[Fig\.4](https://arxiv.org/html/2608.12438#S4.F4)follow the improvement over the full trajectory atg2=0\.4g^\{2\}=0\.4, showing where along the sampling clock the tree\-level gap opens\.

#### A high\-dimensional equivariant drift

The two tests above establish exactness and scaling in minimal settings, while the formalism itself is matrix\-valued and applies in any dimension\. As a demonstration in many dimensions we take a drift couplingN=8N=8points inm=3m=3dimensions, an interacting theory inN​m=24Nm=24latent dimensions,

frev,i​\(z\)\\displaystyle f\_\{\\mathrm\{rev\},i\}\(z\)=a​zi\+a¯​z¯\+�1​‖zi‖2​zi\+�2​S​zi\+�3​\(z¯⋅zi\)​z¯\\displaystyle=a\\,z\_\{i\}\+\\bar\{a\}\\,\\bar\{z\}\+\\lambda\_\{1\}\\,\\\|z\_\{i\}\\\|^\{2\}z\_\{i\}\+\\lambda\_\{2\}\\,S\\,z\_\{i\}\+\\lambda\_\{3\}\\,\(\\bar\{z\}\\\!\\cdot\\\!z\_\{i\}\)\\,\\bar\{z\}withz¯\\displaystyle\\text\{with\}\\qquad\\bar\{z\}=1N∑jzjandS=1N∑j∥zj∥2,\\displaystyle=\\frac\{1\}\{N\}\\sum\_\{j\}z\_\{j\}\\qquad\\text\{and\}\\qquad S=\\frac\{1\}\{N\}\\sum\_\{j\}\\\|z\_\{j\}\\\|^\{2\}\\;,\(4\.26\)which is equivariant under permutations of theNNpoints and under simultaneous rotations of all of them\. We seta=1\.0a=1\.0,a¯=0\.5\\bar\{a\}=0\.5,�1=0\.2\\lambda\_\{1\}=0\.2,�2=0\.1\\lambda\_\{2\}=0\.1and�3=0\.1\\lambda\_\{3\}=0\.1, with a fixed prior drawz1∼𝒩⁡\(0,I\)z\_\{1\}\\sim\\mathcal\{N\}\(0,\\mdmathbb I\)and the sampling dynamics again globally contracting\. The one\-loop prediction integrates the2424\-dimensional classical trajectory together with the full matrix\-valued Lyapunov and mean\-shift equations[4\.8](https://arxiv.org/html/2608.12438#S4.E8)and[4\.12](https://arxiv.org/html/2608.12438#S4.E12), with the Jacobian and the Hessian contraction of the drift evaluated exactly\. The reference is an Euler–Maruyama simulation with1\.3×1051\.3\\times 10^\{5\}trajectories and10310^\{3\}time steps, where doubling the step count shifts the reference by less than0\.010\.01, negligible against the deviations quoted below\. At this dimension no grid\-based reference is available, and the Monte Carlo precision suffices for the comparisons below, while the high\-precision scaling fit remains with the one\-dimensional model above\.

[Table4](https://arxiv.org/html/2608.12438#S4.T4)shows the result\. The couplings place this model in a strongly fluctuating regime, since already atg2=0\.1g^\{2\}=0\.1the loop correction is comparable to the tree\-level value, and atg2=0\.8g^\{2\}=0\.8it exceeds it fivefold, so an uncorrected deterministic sampler misreads the second moment by4141–83%83\\,\\%\. The one\-loop prediction reduces this to1\.1%1\.1\\,\\%atg2=0\.1g^\{2\}=0\.1and degrades gracefully to13%13\\,\\%at the strongest coupling, with the residual growing by a factor of about3\.53\.5per doubling ofg2g^\{2\}, consistent with the expected𝒪⁡\(g4\)\\mathcal\{O\}\(g^\{4\}\)behavior in a regime where higher orders are no longer negligible, while dropping the tadpole term degrades the deviation by up to a further factor of four\. The equivariant structure of the drift and the loop machinery of this section thus combine without modification, and the deviation at the strongest coupling marks the onset of the breakdown discussed next\. The right panels of[Fig\.4](https://arxiv.org/html/2608.12438#S4.F4)follow the same model over the full trajectory atg2=0\.4g^\{2\}=0\.4, where the one\-loop residual stays above the precision of the reference throughout and grows late along the sampling clock, as the𝒪⁡\(g4\)\\mathcal\{O\}\(g^\{4\}\)terms accumulate\.

g2g^\{2\}Tr​C​\(0\)\\mathrm\{Tr\}\\,C\(0\)2​zcl⋅�​m​\(0\)2\\,z\_\{\\mathrm\{cl\}\}\\\!\\cdot\\\!\\delta m\(0\)one\-loopSDErel\. errorw/o�​m\\delta m0\.10\.10\.91860\.9186−0\.0692\-0\.06922\.05522\.05522\.0336​\(13\)2\.0336\(13\)1\.061\.06%4\.54\.5%0\.20\.21\.83721\.8372−0\.1384\-0\.13842\.90462\.90462\.8307​\(20\)2\.8307\(20\)2\.612\.61%7\.57\.5%0\.40\.43\.67453\.6745−0\.2767\-0\.27674\.60354\.60354\.3353​\(31\)4\.3353\(31\)6\.196\.19%12\.612\.6%0\.80\.87\.34907\.3490−0\.5534\-0\.55348\.00138\.00137\.0519​\(50\)7\.0519\(50\)13\.4613\.46%21\.321\.3%Table 4:One\-loop sampler correction for the second moment⟨‖z⁡\(0\)‖2⟩\\langle\\\|z\(0\)\\\|^\{2\}\\rangleof the2424\-dimensional equivariant drift, with tree\-level value‖zcl​\(0\)‖2=1\.206\\\|z\_\{\\mathrm\{cl\}\}\(0\)\\\|^\{2\}=1\.206, versus Euler–Maruyama simulation with1\.3×1051\.3\\times 10^\{5\}trajectories\. The parentheses denote the Monte Carlo uncertainty in the last digits\. The second and third columns are the two components of the one\-loop correction, and the last column is the relative error when the tadpole term is dropped\.

### 4\.4Caustics and the range of validity

The interacting example above was globally contracting\. The first caveat on the range of validity is the size of the loop parameter\. Higher\-loop contributions are suppressed by additional powers ofg2g^\{2\}weighted by the curvature of the drift, and the one\-loop truncation is quantitatively reliable wheng2​‖∇2frev‖g^\{2\}\\,\\\|\\nabla^\{2\}f\_\{\\mathrm\{rev\}\}\\\|is small along the classical trajectory\. This is the regime of low\-to\-moderate diffusion, which includes near\-deterministic samplers\.

The second is more subtle and is specific to the generative setting\. The linearized operatorAcl​\(t\)A\_\{\\mathrm\{cl\}\}\(t\)in[Eq\.4\.3](https://arxiv.org/html/2608.12438#S4.E3)inherits, through the score term in[Eq\.2\.10](https://arxiv.org/html/2608.12438#S2.E10), the curvature of the log\-density, which grows without bound as the marginals sharpen toward the data manifold\. Two degenerate regimes result, distinguished by the sign of this curvature along a given direction\. Along directions that contract onto a sharply peaked mode, the sampling flow*focuses*,i\.e\.neighbouring trajectories are driven together and the map from prior to data degenerates in the limit\. At any finite time the flow map of a smooth drift remains invertible, and the degeneracy is only approached in the limit\. This focusing is the probability\-flow analogue of a*caustic*in optics, where a caustic is the envelope on which a family of light rays is driven together and the ray density diverges, familiar as the bright lines on the bottom of a swimming pool\. Here the rays are the classical trajectories of the probability flow\. The corresponding eigenvalues ofC⁡\(t\)C\(t\)are squeezed toward the small equilibrium value set by the balance of contraction and noise, the precisionC​\(t\)−1C\(t\)^\{\-1\}becomes large, and the propagation of[Eq\.4\.8](https://arxiv.org/html/2608.12438#S4.E8)becomes numerically stiff\. Along directions in which neighbouring trajectories instead separate toward different modes, near the boundaries between basins of attraction, the flow*defocuses*, and the corresponding eigenvalues ofC⁡\(t\)C\(t\)grow exponentially\. In each regime one of the two matricesCCorC−1C^\{\-1\}degenerates while the other remains well conditioned\. The free Ornstein–Uhlenbeck theory of[Section4\.3](https://arxiv.org/html/2608.12438#S4.SS3)has a constant, curvature\-free linearization, so neither degeneracy arises, which is why the one\-loop result is there exact and well behaved\.

#### Regular reformulation via a Riccati equation

Both degeneracies are artifacts of committing to a single variable, not of the underlying physics\. They are removed by propagating, alongside the covariance, its inverse, the precision matrix

P⁡\(t\)≡C​\(t\)−1\.\\displaystyle P\(t\)\\equiv C\(t\)^\{\-1\}\\;\.\(4\.27\)Writings=1−ts=1\-tfor the sampling clock, the Lyapunov equation[4\.8](https://arxiv.org/html/2608.12438#S4.E8)reads

d​Ci​jd​s=−\(Acl,i​k​Ck​j\+Ci​k​Acl,j​k\)\+g2​�i​j,\\displaystyle\\frac\{\\mathrm\{d\}C\_\{ij\}\}\{\\mathrm\{d\}s\}=\-\\big\(A\_\{\\mathrm\{cl\},ik\}\\,C\_\{kj\}\+C\_\{ik\}\\,A\_\{\\mathrm\{cl\},jk\}\\big\)\+g^\{2\}\\,\\delta\_\{ij\}\\;,\(4\.28\)and differentiatingP=C−1P=C^\{\-1\}through the identityd​P/d​s=−P⁡\(d​C/d​s\)​P\\mathrm\{d\}P/\\mathrm\{d\}s=\-P\\,\(\\mathrm\{d\}C/\\mathrm\{d\}s\)\\,Pgives the matrix Riccati equation

d​Pi​jd​s=Pi​k​Acl,k​j​\(s\)\+Acl,k​i​\(s\)​Pk​j−g2​\(s\)​Pi​k​Pk​j,\\displaystyle\\frac\{\\mathrm\{d\}P\_\{ij\}\}\{\\mathrm\{d\}s\}=P\_\{ik\}\\,A\_\{\\mathrm\{cl\},kj\}\(s\)\+A\_\{\\mathrm\{cl\},ki\}\(s\)\\,P\_\{kj\}\-g^\{2\}\(s\)\\,P\_\{ik\}\\,P\_\{kj\}\\;,\(4\.29\)quadratic in the unknown matrix through its last term\. The two equations are complementary\. A rapidly growing eigenvalue ofCC, encountered along defocusing directions, corresponds to an eigenvalue ofPPapproaching zero\. The right\-hand side of[Eq\.4\.29](https://arxiv.org/html/2608.12438#S4.E29)is polynomial inPPwith bounded coefficients, so the precision formulation stays regular in this limit even where propagation inCCbecomes poorly conditioned\. Conversely, in the focusing regime it isPPthat becomes large, while the Lyapunov equation forCCremains regular as the corresponding eigenvalues ofCCbecome small\. In particular, the boundary conditionC⁡\(1\)=0C\(1\)=0itself corresponds toP→∞P\\to\\infty, so the propagation must in any case start from the Lyapunov form\. The stable procedure is therefore hybrid\. We monitor the eigenvalue spectrum ofC⁡\(t\)C\(t\)and propagate whichever ofCCorPPis currently well conditioned, switching back when the spectrum recovers\. Because the two descriptions are exact inverses, the switch introduces no approximation\. By Jacobi’s formula,

dd​s​log​detC=Pi​j​d​Cj​id​s,\\displaystyle\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}s\}\\log\\det C=P\_\{ij\}\\,\\frac\{\\mathrm\{d\}C\_\{ji\}\}\{\\mathrm\{d\}s\}\\;,\(4\.30\)and inserting the Lyapunov equation[4\.28](https://arxiv.org/html/2608.12438#S4.E28)gives

dd​s​log​detC⁡\(s\)=−2​Acl,k​k​\(s\)\+g2​\(s\)​Pk​k​\(s\),\\displaystyle\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}s\}\\log\\det C\(s\)=\-2\\,A\_\{\\mathrm\{cl\},kk\}\(s\)\+g^\{2\}\(s\)\\,P\_\{kk\}\(s\)\\;,\(4\.31\)which is finite in the defocusing regime, whereP→0P\\to 0even as eigenvalues ofCCdiverge\. In the focusing regime its decrease reflects the genuine sharpening of the endpoint distribution rather than a breakdown of the method\. The hybrid propagation of[Eqs\.4\.8](https://arxiv.org/html/2608.12438#S4.E8)and[4\.29](https://arxiv.org/html/2608.12438#S4.E29), together with[Eq\.4\.31](https://arxiv.org/html/2608.12438#S4.E31), thus remains well conditioned whenever one regime dominates the spectrum, as in the examples above\. When focusing and defocusing coexist along different directions, neither matrix is globally well conditioned, and a factored square\-root propagation ofCCextends the scheme to the mixed case\.

At orderg2g^\{2\}the correction has three pieces, the Gaussian broadening, the tadpole mean shift, and the score insertion, and each reduces to an ordinary differential equation integrated alongside the classical trajectory\. The correction needs only Jacobian and Hessian products of the drift, at𝒪⁡\(d\)\\mathcal\{O\}\(d\)products per step for the full covariance\. The low\-rank and diagonal approximations of[Section4\.1](https://arxiv.org/html/2608.12438#S4.SS1)carry it to largedd\.

## 5Symmetry\-constrained flows from EFT power counting

[Section4](https://arxiv.org/html/2608.12438#S4)organized fluctuations around a Gaussian reference\. The drift itself was left arbitrary\. In an effective field theory the interaction terms are not arbitrary\. They are the most general local operators consistent with the symmetries of the problem, ordered by a power\-counting scheme that ranks their importance\. We apply the same logic to the reverse driftfrev​\(z,t\)f\_\{\\mathrm\{rev\}\}\(z,t\), which turns the choice of architecture from an empirical matter into an operator enumeration\.

### 5\.1The drift as an effective expansion

The interaction functionalSint​\[z,z^\]S\_\{\\mathrm\{int\}\}\[z,\\hat\{z\}\]of[Eq\.3\.15](https://arxiv.org/html/2608.12438#S3.E15)is linear in the drift by construction, sincefintf\_\{\\mathrm\{int\}\}appears contracted with a single response field\. Any structure imposed onfintf\_\{\\mathrm\{int\}\}therefore passes directly into the vertices of the diagrammatic expansion\. We fix this structure by the same two ingredients that fix an effective Lagrangian, a symmetry and a power\-counting scheme\.

LetGGbe a group acting on the latent space through a representation�​\(g\)\\rho\(g\),g∈Gg\\in G\. We require the generative model to beGG\-equivariant, so that if the data distribution is invariant,

pdata​\(�​\(g\)​z\)=pdata​\(z\)∀g∈G,\\displaystyle p\_\{\\mathrm\{data\}\}\(\\rho\(g\)\\,z\)=p\_\{\\mathrm\{data\}\}\(z\)\\qquad\\forall\\,g\\in G\\;,\(5\.1\)the generated distribution should be invariant as well\. At the level of the dynamics this is guaranteed if the reverse drift isGG\-equivariant,

frev​\(�​\(g\)​z,t\)=�​\(g\)​frev​\(z,t\)∀g∈G,\\displaystyle f\_\{\\mathrm\{rev\}\}\(\\rho\(g\)\\,z,t\)=\\rho\(g\)\\,f\_\{\\mathrm\{rev\}\}\(z,t\)\\qquad\\forall\\,g\\in G\\;,\(5\.2\)and the prior and diffusion areGG\-invariant\.[Equation5\.2](https://arxiv.org/html/2608.12438#S5.E2)restrictsfrevf\_\{\\mathrm\{rev\}\}to the space of equivariant vector fields on the latent space, and the free/interacting split in[Eq\.3\.2](https://arxiv.org/html/2608.12438#S3.E2)respects this, since the affine part and the nonlinear remainder are separately constrained\.

A symmetry alone leaves an infinite\-dimensional space of admissible drifts\. The second ingredient is a power\-counting scheme that orders this space\. We assign to each candidate operator a*degree*under the rescaling

z→�​zandt→t,\\displaystyle z\\;\\to\\;\\lambda\\,z\\qquad\\text\{and\}\\qquad t\\;\\to\\;t\\;,\(5\.3\)which is the natural dilation of latent space, and we count powers of the diffusion strengthgg\. An operator𝒪⁡\(z,t\)\\mathcal\{O\}\(z,t\)contributing to the drift that scales as𝒪⁡\(�​z,t\)=�n​𝒪​\(z,t\)\\mathcal\{O\}\(\\lambda z,t\)=\\lambda^\{n\}\\mathcal\{O\}\(z,t\)is said to have degreenn\. The affine drift contains the degree\-one operatorzzand, when the symmetry admits it, a constant, while the leading interactions are the lowest\-degree equivariant operators not already present in the free theory\.

This ordering by degree is equivalently a dimensional analysis\. Requiring the MSRJD action in[Eq\.3\.12](https://arxiv.org/html/2608.12438#S3.E12)to be invariant under the dilation[5\.3](https://arxiv.org/html/2608.12438#S5.E3), withttheld fixed, assigns the weights

\[z\]=1,\[z^\]=−1,and\[g\]=1,\\displaystyle\[z\]=1\\;,\\qquad\[\\hat\{z\}\]=\-1\\;,\\qquad\\text\{and\}\\qquad\[g\]=1\\;,\(5\.4\)so that the response propagator and covariance satisfy

\[G\]=0and\[C\]=2,\\displaystyle\[G\]=0\\qquad\\text\{and\}\\qquad\[C\]=2\\;,\(5\.5\)as required\. These weights are fixed by the free theory rather than chosen, which is what makes the counting more than a convention\. A degree\-nndrift operator then enters with a coupling of weight1−n1\-n, so the degree\-one coefficients are dimensionless while the leading degree\-three couplings carry weight−2\-2\. Weights alone do not order anything\. They become a hierarchy once measured against a scale, so we introduce a latent scale�⁡\(t\)\\Lambda\(t\)and write

frev​\(z,t\)=∑n,aca\(n\)​\(t\)​�1−n​\(t\)​𝒪a\(n\)​\(z\)\\displaystyle f\_\{\\mathrm\{rev\}\}\(z,t\)=\\sum\_\{n,a\}c^\{\(n\)\}\_\{a\}\(t\)\\,\\Lambda^\{1\-n\}\(t\)\\,\\mathcal\{O\}^\{\(n\)\}\_\{a\}\(z\)\(5\.6\)wherennruns over operator degrees,aalabels the independentGG\-equivariant operators at each degree, and theca\(n\)c^\{\(n\)\}\_\{a\}are dimensionless\. Two lengths control the expansion,

ℓ⁡\(t\)≡⟨‖z⁡\(t\)‖2⟩and�⁡\(t\)≡‖∇frev‖‖∇2frev‖\|zcl​\(t\),\\displaystyle\\ell\(t\)\\equiv\\sqrt\{\\big\\langle\\\|z\(t\)\\\|^\{2\}\\big\\rangle\}\\qquad\\text\{and\}\\qquad\\Lambda\(t\)\\equiv\\frac\{\\big\\\|\\nabla f\_\{\\mathrm\{rev\}\}\\big\\\|\}\{\\big\\\|\\nabla^\{2\}f\_\{\\mathrm\{rev\}\}\\big\\\|\}\\bigg\|\_\{z\_\{\\mathrm\{cl\}\}\(t\)\}\\;,\(5\.7\)whereℓ\\ellis the typical size of a latent configuration and�\\Lambdais the distance over which the drift departs from linearity\. The second is a curvature scale in the ordinary sense, the length on which the second derivative offrevf\_\{\\mathrm\{rev\}\}becomes comparable to the first, and it is the same object that controls the loop parameter in[Section4\.4](https://arxiv.org/html/2608.12438#S4.SS4)\. Degree\-nnoperators are suppressed by\(ℓ/�\)n−1\(\\ell/\\Lambda\)^\{n\-1\}when the dimensionless coefficients are of order unity, so the expansion is controlled wheneverℓ⁡\(t\)≪�⁡\(t\)\\ell\(t\)\\ll\\Lambda\(t\), meaning the configuration stays within the region where the drift is close to linear\. Without this separation the degree remains a well\-defined ordering but carries no error estimate\.

Whether operator degree and loop order are independent axes depends on the regime\. The latent size of[Eq\.5\.7](https://arxiv.org/html/2608.12438#S5.E7)splits asℓ2=‖zcl‖2\+Tr⁡C\\ell^\{2\}=\\\|z\_\{\\mathrm\{cl\}\}\\\|^\{2\}\+\\operatorname\{Tr\}CwithC∝g2C\\propto g^\{2\}, so when the classical trajectory dominates,ℓ/�\\ell/\\Lambdacarries no dependence onggand the two expansions factorize\. A higher\-degree operator then enters at tree level suppressed byℓ/�\\ell/\\Lambdaalone, while each loop costsg2g^\{2\}regardless of the degree of its vertices\. When fluctuations dominate instead,ℓ∝g\\ell\\propto gand each factor of\(ℓ/�\)2\(\\ell/\\Lambda\)^\{2\}is itself a factor ofg2/�2g^\{2\}/\\Lambda^\{2\}\. Degree and loop counting then collapse onto a single parameter, as in the Weinberg counting of chiral effective field theory where every loop costs\(p/��\)2\(p/\\Lambda\_\{\\chi\}\)^\{2\}\[[74](https://arxiv.org/html/2608.12438#bib.bib4),[75](https://arxiv.org/html/2608.12438#bib.bib5)\]\. The near\-deterministic samplers of[Section4](https://arxiv.org/html/2608.12438#S4)sit in the first regime, where the double expansion is genuine\. For models built from permutation\-invariant pooling a third axis appears, the inverse point number1/N1/N, which enters through the intensive normalization of the pooled operators and the multiplicity of the relative modes\. We do not develop this large\-NNcounting here, but it is the natural parameter organizing the expansion at large point number, and it is what renders the free\-theory loop blocks of[Section5\.2](https://arxiv.org/html/2608.12438#S5.SS2)independent ofNN\.

The resulting structure is that of a*worldline*effective field theory, in which the latent state traces out a trajectory – a worldline – in latent space, parametrized by the auxiliary timett, and the drift operators are local operators living on this worldline, organized by symmetry and power counting\. The construction closely parallels the worldline EFTs used for extended objects in gravity, where a compact body is reduced to a point particle dressed with a tower of symmetry\-allowed operators on its worldline, with coefficients fixed by matching\[[25](https://arxiv.org/html/2608.12438#bib.bib77),[61](https://arxiv.org/html/2608.12438#bib.bib78)\]\. Here the latent configuration plays the role of the point particle, and training performs the matching\.

Given a symmetry groupGGand a target accuracy, one enumerates theGG\-equivariant operators up to a chosen maximum degree, retains them with free coefficients, and discards the rest\. The retained coefficients are the learnable parameters of the model, and forℓ≪�\\ell\\ll\\Lambdathe truncation error is set by the degree of the first neglected operator\.

### 5\.2Permutation\-equivariant drift

Consider latent configurations consisting ofNNpoints inRm\\mdmathbb\{R\}^\{m\},

z=\(z1,…,zN\)withzi∈Rm,\\displaystyle z=\(z\_\{1\},\\dots,z\_\{N\}\)\\qquad\\text\{with\}\\qquad z\_\{i\}\\in\\mdmathbb\{R\}^\{m\}\\;,\(5\.8\)and take the symmetry group to beG=SN×O⁡\(m\)G=S\_\{N\}\\times O\(m\)\. Throughout this section, subscripts onzzlabel the points of the configuration, not times, as the trajectory endpointsz⁡\(0\)=z0z\(0\)=z\_\{0\},z⁡\(1\)=z1z\(1\)=z\_\{1\}do not appear here\. The symmetric groupSNS\_\{N\}permutes the points, whileO⁡\(m\)O\(m\)rotates and reflects each point simultaneously,

�​\(�\)​z=\(z�−1​\(1\),…,z�−1​\(N\)\)andzi→R​ziwithR∈O⁡\(m\)\.\\displaystyle\\rho\(\\sigma\)\\,z=\\big\(z\_\{\\sigma^\{\-1\}\(1\)\},\\dots,z\_\{\\sigma^\{\-1\}\(N\)\}\\big\)\\qquad\\text\{and\}\\qquad z\_\{i\}\\;\\to\\;R\\,z\_\{i\}\\qquad\\text\{with\}\\qquad R\\in O\(m\)\\;\.\(5\.9\)This is the relevant symmetry for generative models of sets of isotropic data, from point clouds and collections of particles to the constituents of a jet in the high\-energy\-physics setting that motivates this work, where the physics is invariant under reordering of the constituents and, for suitable coordinates, under rotations\. Equivariance requires the drift on pointiito be a symmetric function of the other points that transforms as a vector underO⁡\(m\)O\(m\)\. The rotation factor has an immediate structural consequence\. Since−I∈O⁡\(m\)\-\\mdmathbb\{I\}\\in O\(m\), the drift must be odd,

frev​\(−z,t\)=−frev​\(z,t\),\\displaystyle f\_\{\\mathrm\{rev\}\}\(\-z,t\)=\-f\_\{\\mathrm\{rev\}\}\(z,t\)\\;,\(5\.10\)so equivariant operators exist only at odd polynomial degree and a constant term is excluded\. Consistently, the prior must beGG\-invariant, as is the standard normal\. The absence of even\-degree operators is thus not an accident of the enumeration below but a prediction of the symmetry\. Equivariant architectures for such data are well established, from permutation\-invariant function representations like deep sets\[[77](https://arxiv.org/html/2608.12438#bib.bib28)\], to equivariant flows and diffusion models\[[44](https://arxiv.org/html/2608.12438#bib.bib29),[66](https://arxiv.org/html/2608.12438#bib.bib30),[32](https://arxiv.org/html/2608.12438#bib.bib31)\]and their applications to particle clouds in collider physics\[[14](https://arxiv.org/html/2608.12438#bib.bib32),[47](https://arxiv.org/html/2608.12438#bib.bib33)\]\. What the field\-theoretic construction adds is not equivariance itself but its organization\.

The equivariant operators are built from the per\-point vectorziz\_\{i\}together with permutation\-invariant “pooling” combinations of the configuration\. Completeness at each degree follows from the first fundamental theorem forO⁡\(m\)O\(m\)\[[76](https://arxiv.org/html/2608.12438#bib.bib3)\], by which everyO⁡\(m\)O\(m\)\-invariant of a set of vectors is a function of their inner products, so that every equivariant vector field is such an invariant multiplyingziz\_\{i\}or a pooled vector\. Up to degree three, the relevant building blocks are

z¯≡1N​∑jzj,S≡1N​∑j‖zj‖2,T≡1N​∑jzj​zj⊤,w≡1N​∑j‖zj‖2​zj,\\displaystyle\\bar\{z\}\\equiv\\frac\{1\}\{N\}\\sum\_\{j\}z\_\{j\}\\;,\\qquad S\\equiv\\frac\{1\}\{N\}\\sum\_\{j\}\\\|z\_\{j\}\\\|^\{2\}\\;,\\qquad T\\equiv\\frac\{1\}\{N\}\\sum\_\{j\}z\_\{j\}z\_\{j\}^\{\\\!\\top\}\\;,\\qquad w\\equiv\\frac\{1\}\{N\}\\sum\_\{j\}\\\|z\_\{j\}\\\|^\{2\}\\,z\_\{j\}\\;,\(5\.11\)of degrees one, two, two, and three respectively\. Organized by degree, the equivariant vector fields acting on pointziz\_\{i\}are

degree​1:\\displaystyle\\text\{degree \}1:\\quadzi,z¯,\\displaystyle z\_\{i\}\\;,\\qquad\\bar\{z\}\\;,degree​3:\\displaystyle\\text\{degree \}3:\\quad‖zi‖2​zi,S​zi,\(z¯⋅zi\)​z¯,…\\displaystyle\\\|z\_\{i\}\\\|^\{2\}z\_\{i\}\\;,\\qquad S\\,z\_\{i\}\\;,\\qquad\(\\bar\{z\}\\\!\\cdot\\\!z\_\{i\}\)\\,\\bar\{z\}\\;,\\qquad\\dots\(5\.12\)where the complete degree\-three basis contains eleven operators, obtained from allO⁡\(m\)O\(m\)\-contractions of three vectors drawn from\{zi,z¯\}\\\{z\_\{i\},\\bar\{z\}\\\}and the pooled moments[5\.11](https://arxiv.org/html/2608.12438#S5.E11), the six structures‖zi‖2​zi\\\|z\_\{i\}\\\|^\{2\}z\_\{i\},\(z¯⋅zi\)​zi\(\\bar\{z\}\\\!\\cdot\\\!z\_\{i\}\)\\,z\_\{i\},‖z¯‖2​zi\\\|\\bar\{z\}\\\|^\{2\}z\_\{i\},‖zi‖2​z¯\\\|z\_\{i\}\\\|^\{2\}\\,\\bar\{z\},\(z¯⋅zi\)​z¯\(\\bar\{z\}\\\!\\cdot\\\!z\_\{i\}\)\\,\\bar\{z\},‖z¯‖2​z¯\\\|\\bar\{z\}\\\|^\{2\}\\,\\bar\{z\}, the four moment\-dressed structuresS​ziS\\,z\_\{i\},S​z¯S\\,\\bar\{z\},T​ziT\\,z\_\{i\},T​z¯T\\,\\bar\{z\}, and the pooled vectorww\. For smallNNormmsome of these structures become linearly dependent,e\.g\.TTcoincides withSSform=1m=1\. The degree\-one operators reproduce, after summation with free coefficients, the most general linear equivariant drift, a self term∝zi\\propto z\_\{i\}and a mean\-field term∝z¯\\propto\\bar\{z\}\. This is the free theory of[Section3\.1](https://arxiv.org/html/2608.12438#S3.SS1), now identified as the degree\-one truncation of the equivariant expansion\. The degree\-three operators are the leading interactions, and each is cubic inzzand hence generates a four\-point vertex of the MSRJD expansion,i\.e\.one response leg and three latent legs, see[Fig\.2](https://arxiv.org/html/2608.12438#S3.F2)\(c\)\.

Retaining the complete basis up to degree three with time\-dependent coefficients defines the leading equivariant reverse drift,

frev,i​\(z,t\)=a​\(t\)​zi\+a¯​\(t\)​z¯⏟free, degree​1\+�1​\(t\)​‖zi‖2​zi\+�2​\(t\)​S​zi\+�3​\(t\)​\(z¯⋅zi\)​z¯\+…⏟leading interactions, degree​3,\\displaystyle f\_\{\\mathrm\{rev\},i\}\(z,t\)=\\underbrace\{a\(t\)\\,z\_\{i\}\+\\bar\{a\}\(t\)\\,\\bar\{z\}\}\_\{\\text\{free, degree \}1\}\+\\underbrace\{\\lambda\_\{1\}\(t\)\\,\\\|z\_\{i\}\\\|^\{2\}z\_\{i\}\+\\lambda\_\{2\}\(t\)\\,S\\,z\_\{i\}\+\\lambda\_\{3\}\(t\)\\,\(\\bar\{z\}\\\!\\cdot\\\!z\_\{i\}\)\\,\\bar\{z\}\+\\ldots\}\_\{\\text\{leading interactions, degree \}3\}\\;,\(5\.13\)where all scalar coefficients are functions ofttonly\. Truncated at degree three and with constant coefficients, this is the drift used for the2424\-dimensional test in[Section4\.3](https://arxiv.org/html/2608.12438#S4.SS3)\. The power counting makes the role of each term transparent\. The free coefficientsa,a¯a,\\bar\{a\}define the Gaussian reference flow and are fixed, as in[Section4](https://arxiv.org/html/2608.12438#S4), by the linearized dynamics\. The couplings�k\\lambda\_\{k\}multiply the leading interaction vertices, distinguished by how the four legs are routed through the point index\. They are the independent components that the symmetry leaves in the rank\-four tensor�\(3\)\\lambda^\{\(3\)\}of[Eq\.3\.26](https://arxiv.org/html/2608.12438#S3.E26), reduced from𝒪⁡\(\(N​m\)4\)\\mathcal\{O\}\\big\(\(Nm\)^\{4\}\\big\)free entries to eleven numbers\. Operators of degree five and higher are the first neglected terms, and their omission is controlled whenℓ≪�\\ell\\ll\\Lambdaand the dimensionless coefficients stay of order unity\.

#### Consequences of the power counting

The construction changes what a network has to learn, as we can see for score matching and conditional flow matching\. Score matching and conditional flow matching regress different objects, but the forward process fixesf⁡\(z,t\)f\(z,t\)andg⁡\(t\)g\(t\), so both targets follow fromfrevf\_\{\\mathrm\{rev\}\}alone,

s⁡\(z,t\)\\displaystyle s\(z,t\)≡∇z​log​p​\(z,t\)=1g2​\(t\)​\[f⁡\(z,t\)−frev​\(z,t\)\],\\displaystyle\\equiv\\nabla\_\{z\}\\log p\(z,t\)=\\frac\{1\}\{g^\{2\}\(t\)\}\\Big\[\\,f\(z,t\)\-f\_\{\\mathrm\{rev\}\}\(z,t\)\\,\\Big\]\\;,v⁡\(z,t\)\\displaystyle v\(z,t\)=12​\[f⁡\(z,t\)\+frev​\(z,t\)\],\\displaystyle=\\tfrac\{1\}\{2\}\\Big\[\\,f\(z,t\)\+f\_\{\\mathrm\{rev\}\}\(z,t\)\\,\\Big\]\\;,\(5\.14\)the score of[Eq\.2\.10](https://arxiv.org/html/2608.12438#S2.E10)and the probability\-flow velocity of[Eq\.2\.63](https://arxiv.org/html/2608.12438#S2.E63)\. Replacingfrevf\_\{\\mathrm\{rev\}\}by the truncated expansion in[Eq\.5\.13](https://arxiv.org/html/2608.12438#S5.E13), writtenf�EFTf^\{\\mathrm\{EFT\}\}\_\{\\theta\}with coefficient functionsa⁡\(t\)a\(t\),a¯​\(t\)\\bar\{a\}\(t\), and�k​\(t\)\\lambda\_\{k\}\(t\)parametrized by a network with parameters�\\theta, gives

s�​\(z,t\)\\displaystyle s\_\{\\theta\}\(z,t\)=1g2​\(t\)​\[f⁡\(z,t\)−f�EFT​\(z,t\)\],\\displaystyle=\\frac\{1\}\{g^\{2\}\(t\)\}\\Big\[\\,f\(z,t\)\-f^\{\\mathrm\{EFT\}\}\_\{\\theta\}\(z,t\)\\,\\Big\]\\;,v�​\(z,t\)\\displaystyle v\_\{\\theta\}\(z,t\)=12​\[f⁡\(z,t\)\+f�EFT​\(z,t\)\]\.\\displaystyle=\\tfrac\{1\}\{2\}\\Big\[\\,f\(z,t\)\+f^\{\\mathrm\{EFT\}\}\_\{\\theta\}\(z,t\)\\,\\Big\]\\;\.\(5\.15\)Standard score matching fits a generics�​\(z,t\)s\_\{\\theta\}\(z,t\)and enforces the symmetry through the architecture\. Here the functional form is fixed, so the network maps a single time argument to thirteen coefficients rather than an\(N×m\)\(N\\times m\)\-dimensional state to an\(N×m\)\(N\\times m\)\-dimensional output, and the losses in[Eqs\.2\.60](https://arxiv.org/html/2608.12438#S2.E60)and[2\.67](https://arxiv.org/html/2608.12438#S2.E67)apply unchanged\. Equivariance holds identically becausef�EFTf^\{\\mathrm\{EFT\}\}\_\{\\theta\}is assembled from equivariant operators, once the prior and the conditional interpolation share the symmetry of the data\. The score ansatz divides an𝒪⁡\(g2\)\\mathcal\{O\}\(g^\{2\}\)difference byg2g^\{2\}, so errors inf�EFTf^\{\\mathrm\{EFT\}\}\_\{\\theta\}are amplified at smallgg, while the velocity ansatz is free of this\.

First, the expansion is*equivariant order by order*\. Because the vertices in[Eq\.5\.13](https://arxiv.org/html/2608.12438#S5.E13)and the propagators of the equivariant free theory are separately equivariant, every diagram built from them is equivariant, and the symmetry is preserved order by order ing2g^\{2\}as an exact property of the truncation rather than an approximate property of a trained model\. The closure holds for the symmetry and not for the degree, since higher orders can generate symmetry\-allowed structures of higher degree, which the expansion then orders in the same way\.

Second, the scheme*predicts a hierarchy*\. The relative weight of the degree\-three interactions against the linear flow measures\(ℓ/�\)2\(\\ell/\\Lambda\)^\{2\}directly,

�​\(t\)≡maxk⁡\|�k​\(t\)\|​ℓ2​\(t\)\|a⁡\(t\)\|,\\displaystyle\\epsilon\(t\)\\equiv\\frac\{\\max\_\{k\}\|\\lambda\_\{k\}\(t\)\|\\,\\ell^\{2\}\(t\)\}\{\|a\(t\)\|\}\\;,\(5\.16\)and, under the same order\-unity assumption for the dimensionless coefficients, degree\-\(2​k\+1\)\(2k\{\+\}1\)operators contribute at relative order�k\\epsilon^\{k\}\. The parameter is invariant under rescalings of the latent space, since the degree\-three couplings carry two inverse powers of�\\Lambda, and it is a measurable property of the fitted model rather than a schedule choice\. This makes the hierarchy falsifiable\. Fitted degree\-five couplings should be suppressed relative to fitted degree\-three couplings by a further power of�\\epsilon, and a measured�\\epsilonof order unity would show that the truncation is not controlled for that data\.

Third, the symmetry*reduces the cost of the loop expansion*of[Section4](https://arxiv.org/html/2608.12438#S4)\. For the free part of[Eq\.5\.13](https://arxiv.org/html/2608.12438#S5.E13), theSNS\_\{N\}isotypic decomposition\[[22](https://arxiv.org/html/2608.12438#bib.bib24)\]– the decomposition of the linearized dynamics into its irreducible symmetry sectors – splits the\(N×m\)\(N\\times m\)\-dimensional linearized dynamics into the mean modez¯\\bar\{z\}, with drift eigenvaluea\+a¯a\+\\bar\{a\}, andN−1N\{\-\}1identical relative modes with eigenvalueaa\. The Lyapunov equation[4\.8](https://arxiv.org/html/2608.12438#S4.E8)then factorizes into twom×mm\\times mblocks, independent ofNN\. Around a generic, non\-symmetric classical trajectory the interacting linearization respects this block structure only approximately, but it remains the natural preconditioner for the one\-loop computation at largeNN\.

The combination of the equivariance condition in[Eq\.5\.2](https://arxiv.org/html/2608.12438#S5.E2)and the power counting in[Eq\.5\.3](https://arxiv.org/html/2608.12438#S5.E3)selects a finite, ordered set of operators from the infinite\-dimensional space of admissible drifts, replacing architecture search with operator enumeration\. The same construction applies to any symmetry for which the equivariant operators can be enumerated – Lorentz symmetry for the four\-momenta of high\-energy physics, orSNS\_\{N\}alone when no rotational structure is present, in which case additional even\-degree operators such as componentwise products appear – with the group changing only the operator list[5\.12](https://arxiv.org/html/2608.12438#S5.E12), not the logic of the expansion\. The measured�\\epsilonof a fitted model decides whether the truncation is controlled for a given dataset, which makes[Eq\.5\.13](https://arxiv.org/html/2608.12438#S5.E13)a hypothesis with a built\-in test rather than an architecture choice\.

## 6Outlook

Path integrals organize generative models into a single master action, and that organization turns sampler errors, score imperfections, and architecture design into perturbative computations\. The resulting picture interprets generative modeling as a worldline effective field theory on probability space, in the sense of[Section5](https://arxiv.org/html/2608.12438#S5)\. Latent variables evolve along an auxiliary time direction, an Onsager–Machlup action weighs their trajectories, and the common generative model classes occupy the familiar regimes of such a theory\. Normalizing flows are its classical limit, in which the path integral collapses onto a single trajectory and exact likelihoods follow from a change of variables\. Finite diffusion turns on fluctuations around this classical path, and affine drifts single out the free theory\. Nonlinear diffusion models are then interacting worldline theories\. Their exact likelihoods are generally intractable, common likelihood approximations truncate the underlying functional integral, and score\-based training implicitly reconstructs the full interacting drift\.

Conditional flow matching fits the same picture\. It is not a distinct probabilistic model but an evaluation principle that replaces stochastic sampling by a deterministic flow, retains the finite\-noise marginals, and absorbs the unsampled fluctuations into an effective transport field\. Which parts of the interacting expansion it retains and which it averages over, compared with score matching and Schrödinger bridges, ties empirical performance to the diagrammatic structure\. The scattering viewpoint of[Section3\.2](https://arxiv.org/html/2608.12438#S3.SS2)treats conditioning, partial observations, and constraints as operator insertions, which places conditional generation in the same formalism\.

The loop expansion of[Section4](https://arxiv.org/html/2608.12438#S4)carries over to trained diffusion models without modification, since the hybrid Lyapunov–Riccati formulation of[Section4\.4](https://arxiv.org/html/2608.12438#S4.SS4)needs only Jacobian and Hessian products of the drift\. A loop\-corrected deterministic sampler, the drift ODE supplemented by the𝒪⁡\(g2\)\\mathcal\{O\}\(g^\{2\}\)correction, then delivers stochastic\-process observables from two auxiliary linear equations solved along the same trajectory, at no stochastic\-sampling cost\. Whether that correction survives contact with a learned score, where the score error of[Section4\.2](https://arxiv.org/html/2608.12438#S4.SS2)competes with the fluctuation term, is the first question to settle\. The insertion analysis of[Section4\.2](https://arxiv.org/html/2608.12438#S4.SS2)adds two further tests\. The response weighting derived there can compete against empirically tuned loss weightings, and the loop integrand, which identifies where along the trajectory the deterministic sampler deviates most, can allocate a limited budget of stochastic steps where they buy the most accuracy\.

The power\-counting scheme of[Section5](https://arxiv.org/html/2608.12438#S5)suggests a complementary program on the architecture side\. Building generative models from the leading equivariant drift of[Section5\.2](https://arxiv.org/html/2608.12438#S5.SS2), measuring the fitted interaction couplings, and testing the predicted hierarchy between operator degrees confront the expansion with practice\. For data whose complexity exceeds the reach of a polynomial truncation, the EFT drift provides a symmetry\-exact analytic backbone, with a small network learning only the higher\-degree residual whose size the power counting bounds\. Because the operator basis is exhaustive at each degree, it also contains drift terms that standard architectures do not realize, and the sign structure of the coefficients controls whether the flow contracts toward the data manifold\.

The formulation connects deterministic and stochastic generative models within one action and identifies the analytically solvable limits\. Sampler accuracy is now a calculation, carried out and validated in[Section4](https://arxiv.org/html/2608.12438#S4)\. Score quality and architecture design are so far only organized, but turning that organization into numbers is the work this framework makes possible\.

### Code availability

The code reproducing the numerical validations of[Section4\.3](https://arxiv.org/html/2608.12438#S4.SS3)– the Ornstein–Uhlenbeck free theory, the interacting cubic drift, and the2424\-dimensional equivariant drift, including the Fokker–Planck reference solver and the Euler–Maruyama simulations, is publicly available at[https://github\.com/ramonpeter/generative\-path\-integrals](https://github.com/ramonpeter/generative-path-integrals)\.

### Acknowledgements

I thank Rafael Aoude for detailed feedback on the manuscript and for pointing me to the worldline effective field theory literature, and Claudius Krause for his questions about adversarial models that turned into a subsection of its own\. I am grateful to Sofia Palacios Schweitzer for a detailed round of comments that improved the manuscript throughout, and to Nina Elmer, who was involved in the early days of this project, and Henning Bahl for reading it and for their comments along the way\.

## References

- \[1\]\(2023\)Building normalizing flows with stochastic interpolants\.External Links:2209\.15571Cited by:[§1](https://arxiv.org/html/2608.12438#S1.p3.1)\.
- \[2\]O\. Amramet al\.\(2025\)CaloChallenge 2022: a community challenge for fast calorimeter simulation\.Rept\. Prog\. Phys\.88\(11\),pp\. 116201\.External Links:2410\.21611,[Document](https://dx.doi.org/10.1088/1361-6633/ae1304)Cited by:[§1](https://arxiv.org/html/2608.12438#S1.p1.1)\.
- \[3\]B\. D\.O\. Anderson\(1982\)Reverse\-time diffusion equation models\.Stochastic Processes and their Applications12\(3\),pp\. 313–326\.External Links:ISSN 0304\-4149,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/0304-4149%2882%2990051-5)Cited by:[§2\.2](https://arxiv.org/html/2608.12438#S2.SS2.p2.6)\.
- \[4\]M\. Arjovsky and L\. Bottou\(2017\)Towards principled methods for training generative adversarial networks\.InInternational Conference on Learning Representations,External Links:1701\.04862Cited by:[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx5.p3.1)\.
- \[5\]V\. I\. Arnold\(1992\)Ordinary differential equations\.Springer\.Cited by:[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx1.p1.5)\.
- \[6\]C\. Aron, G\. Biroli, and L\. F\. Cugliandolo\(2010\)Symmetries of generating functionals of Langevin processes with colored multiplicative noise\.J\. Stat\. Mech\.1011,pp\. P11018\.External Links:1007\.5059,[Document](https://dx.doi.org/10.1088/1742-5468/2010/11/P11018)Cited by:[§2\.2](https://arxiv.org/html/2608.12438#S2.SS2.SSSx1.p1.1)\.
- \[7\]S\. Badgeret al\.\(2023\)Machine learning and LHC event generation\.SciPost Phys\.14\(4\),pp\. 079\.External Links:2203\.07460,[Document](https://dx.doi.org/10.21468/SciPostPhys.14.4.079)Cited by:[§1](https://arxiv.org/html/2608.12438#S1.p1.1)\.
- \[8\]R\. Becker and R\. Rannacher\(2001\)An optimal control approach to a posteriori error estimation in finite element methods\.Acta Numerica10,pp\. 1–102\.Cited by:[§4\.2](https://arxiv.org/html/2608.12438#S4.SS2.SSSx1.p2.1)\.
- \[9\]J\. Berner, L\. Richter, and K\. Ullrich\(2024\)An optimal control perspective on diffusion\-based generative modeling\.External Links:2211\.01364Cited by:[§1](https://arxiv.org/html/2608.12438#S1.p3.1)\.
- \[10\]C\. M\. Bishop\(2006\)Pattern recognition and machine learning\.Springer\.External Links:ISBN 978\-0\-387\-31073\-2Cited by:[§2\.1](https://arxiv.org/html/2608.12438#S2.SS1.p1.1)\.
- \[11\]D\. M\. Blei, A\. Kucukelbir, and J\. D\. McAuliffe\(2017\)Variational inference: a review for statisticians\.Journal of the American Statistical Association112\(518\),pp\. 859–877\.External Links:[Document](https://dx.doi.org/10.1080/01621459.2017.1285773)Cited by:[§2\.1](https://arxiv.org/html/2608.12438#S2.SS1.p1.1)\.
- \[12\]V\. D\. Bortoli, J\. Thornton, J\. Heng, and A\. Doucet\(2023\)Diffusion schrödinger bridge with applications to score\-based generative modeling\.External Links:2106\.01357Cited by:[§1](https://arxiv.org/html/2608.12438#S1.p3.1),[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx4.p1.1),[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx4.p3.1),[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx4.p3.2)\.
- \[13\]J\. Brehmer, K\. Cranmer, G\. Louppe, and J\. Pavez\(2018\)A Guide to Constraining Effective Field Theories with Machine Learning\.Phys\. Rev\. D98\(5\),pp\. 052004\.External Links:1805\.00020,[Document](https://dx.doi.org/10.1103/PhysRevD.98.052004)Cited by:[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx2.p2.8)\.
- \[14\]E\. Buhmann, G\. Kasieczka, and J\. Thaler\(2023\)EPiC\-GAN: Equivariant point cloud generation for particle jets\.SciPost Phys\.15\(4\),pp\. 130\.External Links:2301\.08128,[Document](https://dx.doi.org/10.21468/SciPostPhys.15.4.130)Cited by:[§5\.2](https://arxiv.org/html/2608.12438#S5.SS2.p1.4)\.
- \[15\]Y\. Chen, T\. T\. Georgiou, and M\. Pavon\(2021\)Stochastic control liaisons: Richard Sinkhorn meets Gaspard Monge on a Schrödinger bridge\.SIAM Review63\(2\),pp\. 249–313\.Cited by:[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx4.p1.1)\.
- \[16\]E\. A\. Coddington and N\. Levinson\(1955\)Theory of ordinary differential equations\.McGraw\-Hill\.Cited by:[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx1.p1.5)\.
- \[17\]C\. De Dominicis\(1976\)Techniques de renormalisation de la théorie des champs et dynamique des phénomènes critiques\.Journal de Physique Colloques37\(C1\),pp\. C1–247–C1–253\.External Links:[Document](https://dx.doi.org/10.1051/jphyscol%3A1976138)Cited by:[§1](https://arxiv.org/html/2608.12438#S1.p2.1),[§3\.1](https://arxiv.org/html/2608.12438#S3.SS1.p3.1)\.
- \[18\]C\. Domingo\-Enrich, M\. Drozdzal, B\. Karrer, and R\. T\. Q\. Chen\(2025\)Adjoint matching: fine\-tuning flow and diffusion generative models with memoryless stochastic optimal control\.InInternational Conference on Learning Representations,External Links:2409\.08861Cited by:[§4\.2](https://arxiv.org/html/2608.12438#S4.SS2.SSSx1.p2.1)\.
- \[19\]J\. L\. Doob\(1984\)Classical potential theory and its probabilistic counterpart\.Springer,New York\.Cited by:[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx4.p1.1)\.
- \[20\]M\. Feickert and B\. Nachman\(2021\)A Living Review of Machine Learning for Particle Physics\.External Links:2102\.02770Cited by:[§1](https://arxiv.org/html/2608.12438#S1.p1.1)\.
- \[21\]R\. P\. Feynman and A\. R\. Hibbs\(1965\)Quantum mechanics and path integrals\.McGraw\-Hill,New York\.Cited by:[§2\.2](https://arxiv.org/html/2608.12438#S2.SS2.p3.12)\.
- \[22\]W\. Fulton and J\. Harris\(1991\)Representation theory: a first course\.Graduate Texts in Mathematics, Vol\.129,Springer\.Cited by:[§5\.2](https://arxiv.org/html/2608.12438#S5.SS2.SSSx1.p4.1)\.
- \[23\]C\. W\. Gardiner\(1985\)Handbook of stochastic methods for physics, chemistry and the natural sciences\.2 edition,Springer Series in Synergetics, Vol\.13,Springer,Berlin, Heidelberg\.External Links:[Document](https://dx.doi.org/10.1007/978-3-662-02452-2)Cited by:[§2\.2](https://arxiv.org/html/2608.12438#S2.SS2.p2.4),[§2\.2](https://arxiv.org/html/2608.12438#S2.SS2.p2.5),[§2\.2](https://arxiv.org/html/2608.12438#S2.SS2.p3.12),[§4\.1](https://arxiv.org/html/2608.12438#S4.SS1.p2.4)\.
- \[24\]C\. Gardiner\(2009\)Stochastic methods: a handbook for the natural and social sciences\.4 edition,Springer Series in Synergetics,Springer,Berlin\.Cited by:[§2\.2](https://arxiv.org/html/2608.12438#S2.SS2.p3.12)\.
- \[25\]W\. D\. Goldberger and I\. Z\. Rothstein\(2006\)An Effective field theory of gravity for extended objects\.Phys\. Rev\. D73,pp\. 104029\.External Links:hep\-th/0409156,[Document](https://dx.doi.org/10.1103/PhysRevD.73.104029)Cited by:[§5\.1](https://arxiv.org/html/2608.12438#S5.SS1.p6.1)\.
- \[26\]I\. J\. Goodfellow, J\. Pouget\-Abadie, M\. Mirza, B\. Xu, D\. Warde\-Farley, S\. Ozair, A\. Courville, and Y\. Bengio\(2014\)Generative adversarial nets\.InAdvances in Neural Information Processing Systems,Vol\.27\.External Links:1406\.2661Cited by:[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx5.p2.1)\.
- \[27\]J\. K\. Hale\(2009\)Ordinary differential equations\.Dover Publications\.Cited by:[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx1.p1.5)\.
- \[28\]U\. G\. Haussmann and É\. Pardoux\(1986\)TIME reversal of diffusions\.Annals of Probability14,pp\. 1188–1205\.Cited by:[§2\.2](https://arxiv.org/html/2608.12438#S2.SS2.p2.6)\.
- \[29\]HEP ML CommunityA Living Review of Machine Learning for Particle Physics\.External Links:[Link](https://iml-wg.github.io/HEPML-LivingReview/)Cited by:[§1](https://arxiv.org/html/2608.12438#S1.p1.1)\.
- \[30\]D\. J\. Higham\(2001\)An algorithmic introduction to numerical simulation of stochastic differential equations\.SIAM Review43\(3\),pp\. 525–546\.External Links:[Document](https://dx.doi.org/10.1137/S0036144500378302)Cited by:[§2\.2](https://arxiv.org/html/2608.12438#S2.SS2.p3.1)\.
- \[31\]Y\. Hirono, A\. Tanaka, and K\. Fukushima\(2024\)Understanding diffusion models by feynman’s path integral\.External Links:2403\.11262Cited by:[§1](https://arxiv.org/html/2608.12438#S1.p3.1),[§4](https://arxiv.org/html/2608.12438#S4.p3.1)\.
- \[32\]E\. Hoogeboom, V\. G\. Satorras, C\. Vignac, and M\. Welling\(2022\)Equivariant diffusion for molecule generation in 3d\.External Links:2203\.17003Cited by:[§5\.2](https://arxiv.org/html/2608.12438#S5.SS2.p1.4)\.
- \[33\]C\. Huang, J\. H\. Lim, and A\. Courville\(2021\)A variational perspective on diffusion\-based generative models and score matching\.External Links:2106\.02808Cited by:[§1](https://arxiv.org/html/2608.12438#S1.p3.1),[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx2.p2.9)\.
- \[34\]J\. Hubbard\(1959\)Calculation of partition functions\.Phys\. Rev\. Lett\.3,pp\. 77–80\.External Links:[Document](https://dx.doi.org/10.1103/PhysRevLett.3.77)Cited by:[§3\.1](https://arxiv.org/html/2608.12438#S3.SS1.p3.6)\.
- \[35\]H\. Janssen\(1976\)On a lagrangean for classical field dynamics and renormalization group calculations of dynamical critical properties\.Zeitschrift für Physik B Condensed Matter23\(4\),pp\. 377–380\.External Links:[Document](https://dx.doi.org/10.1007/BF01316547),ISBN 1431\-584XCited by:[§1](https://arxiv.org/html/2608.12438#S1.p2.1),[§3\.1](https://arxiv.org/html/2608.12438#S3.SS1.p3.1)\.
- \[36\]A\. Kamenev\(2011\)Field theory of non\-equilibrium systems\.Cambridge University Press,Cambridge\.External Links:ISBN 9780521760829,[Document](https://dx.doi.org/10.1017/CBO9781139003667)Cited by:[§3\.1](https://arxiv.org/html/2608.12438#S3.SS1.p3.5)\.
- \[37\]T\. Karras, M\. Aittala, T\. Aila, and S\. Laine\(2022\)Elucidating the design space of diffusion\-based generative models\.InAdvances in Neural Information Processing Systems,Vol\.35\.External Links:2206\.00364Cited by:[§4\.2](https://arxiv.org/html/2608.12438#S4.SS2.SSSx1.p1.3)\.
- \[38\]D\. Kim, Y\. Kim, S\. J\. Kwon, W\. Kang, and I\. Moon\(2023\)Refining generative process with discriminator guidance in score\-based diffusion models\.InProceedings of the 40th International Conference on Machine Learning,Vol\.202,pp\. 16567–16598\.External Links:2211\.17091Cited by:[§4\.2](https://arxiv.org/html/2608.12438#S4.SS2.p1.4)\.
- \[39\]D\. P\. Kingma and R\. Gao\(2023\)Understanding diffusion objectives as the ELBO with simple data augmentation\.InAdvances in Neural Information Processing Systems,Vol\.36\.External Links:2303\.00848Cited by:[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx2.p2.9)\.
- \[40\]D\. P\. Kingma and M\. Welling\(2014\)Auto\-encoding variational Bayes\.InInternational Conference on Learning Representations \(ICLR\),External Links:1312\.6114Cited by:[§2\.1](https://arxiv.org/html/2608.12438#S2.SS1.p1.3)\.
- \[41\]D\. P\. Kingma and M\. Welling\(2019\)An introduction to variational autoencoders\.Foundations and Trends in Machine Learning12\(4\),pp\. 307–392\.External Links:[Document](https://dx.doi.org/10.1561/2200000056)Cited by:[§2\.1](https://arxiv.org/html/2608.12438#S2.SS1.p1.3)\.
- \[42\]H\. Kleinert\(2009\)Path integrals in quantum mechanics, statistics, polymer physics, and financial markets\.5 edition,World Scientific\.Cited by:[§2\.1](https://arxiv.org/html/2608.12438#S2.SS1.p2.4)\.
- \[43\]P\. E\. Kloeden and E\. Platen\(1992\)Numerical solution of stochastic differential equations\.Applications of Mathematics, Vol\.23,Springer,Berlin\.External Links:[Document](https://dx.doi.org/10.1007/978-3-662-12616-5)Cited by:[§2\.2](https://arxiv.org/html/2608.12438#S2.SS2.p3.1)\.
- \[44\]J\. Köhler, L\. Klein, and F\. Noé\(2020\)Equivariant flows: exact likelihood generative learning for symmetric densities\.External Links:2006\.02425Cited by:[§5\.2](https://arxiv.org/html/2608.12438#S5.SS2.p1.4)\.
- \[45\]C\. Krause, R\. Winterhalder, M\. Feickert, B\. Nachman, J\. Raine, and HEP\-ML Community\(2026\)A living review of machine learning for particle physics\.Zenodo\.External Links:[Document](https://dx.doi.org/10.5281/zenodo.21626667),[Link](https://doi.org/10.5281/zenodo.21626667)Cited by:[§1](https://arxiv.org/html/2608.12438#S1.p1.1)\.
- \[46\]C\. Krause, R\. Winterhalder, M\. Feickert, and B\. Nachman\(2026\)The Living Guide of Machine Learning for Particle Physics\.External Links:2608\.09531Cited by:[§1](https://arxiv.org/html/2608.12438#S1.p1.1)\.
- \[47\]M\. Leigh, D\. Sengupta, G\. Quétant, J\. A\. Raine, K\. Zoch, and T\. Golling\(2024\)PC\-JeDi: Diffusion for particle cloud generation in high energy physics\.SciPost Phys\.16\(1\),pp\. 018\.External Links:2303\.05376,[Document](https://dx.doi.org/10.21468/SciPostPhys.16.1.018)Cited by:[§5\.2](https://arxiv.org/html/2608.12438#S5.SS2.p1.4)\.
- \[48\]C\. Léonard\(2013\)A survey of the schrödinger problem and some of its connections with optimal transport\.External Links:1308\.0215Cited by:[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx4.p1.1),[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx4.p1.2)\.
- \[49\]Y\. Lipman, R\. T\. Q\. Chen, H\. Ben\-Hamu, M\. Nickel, and M\. Le\(2023\)Flow matching for generative modeling\.External Links:2210\.02747Cited by:[§1](https://arxiv.org/html/2608.12438#S1.p3.1),[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx3.p2.4),[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx3.p2.5)\.
- \[50\]Y\. Lipman, M\. Havasi, P\. Holderrieth, N\. Shaul, M\. Le, B\. Karrer, R\. T\. Q\. Chen, D\. Lopez\-Paz, H\. Ben\-Hamu, and I\. Gat\(2024\)Flow matching guide and code\.External Links:2412\.06264Cited by:[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx3.p2.5)\.
- \[51\]S\. Machlup and L\. Onsager\(1953\)Fluctuations and irreversible processes\. II\. systems with kinetic energy\.Physical Review91\(6\),pp\. 1512–1515\.External Links:[Document](https://dx.doi.org/10.1103/PhysRev.91.1512)Cited by:[§1](https://arxiv.org/html/2608.12438#S1.p2.1)\.
- \[52\]P\. C\. Martin, E\. D\. Siggia, and H\. A\. Rose\(1973\)Statistical dynamics of classical systems\.Phys\. Rev\. A8,pp\. 423–437\.External Links:[Document](https://dx.doi.org/10.1103/PhysRevA.8.423),[Link](https://link.aps.org/doi/10.1103/PhysRevA.8.423)Cited by:[§1](https://arxiv.org/html/2608.12438#S1.p2.1),[§3\.1](https://arxiv.org/html/2608.12438#S3.SS1.p3.1)\.
- \[53\]G\. Maruyama\(1955\)Continuous markov processes and stochastic equations\.Rendiconti del Circolo Matematico di Palermo4,pp\. 48–90\.External Links:[Document](https://dx.doi.org/10.1007/BF02846028)Cited by:[§2\.2](https://arxiv.org/html/2608.12438#S2.SS2.p3.1)\.
- \[54\]L\. Mescheder, A\. Geiger, and S\. Nowozin\(2018\)Which training methods for GANs do actually converge?\.InProceedings of the 35th International Conference on Machine Learning,Vol\.80,pp\. 3481–3490\.External Links:1801\.04406Cited by:[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx5.p3.1)\.
- \[55\]T\. Miyato, T\. Kataoka, M\. Koyama, and Y\. Yoshida\(2018\)Spectral normalization for generative adversarial networks\.InInternational Conference on Learning Representations,External Links:1802\.05957Cited by:[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx5.p3.1)\.
- \[56\]J\. Neyman and E\. S\. Pearson\(1933\)On the problem of the most efficient tests of statistical hypotheses\.Philos\. Trans\. R\. Soc\. Lond\. A231,pp\. 289–337\.External Links:[Document](https://dx.doi.org/10.1098/rsta.1933.0009)Cited by:[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx5.p2.3)\.
- \[57\]D\. Nielsen, P\. Jaini, E\. Hoogeboom, O\. Winther, and M\. Welling\(2020\)SurVAE flows: surjections to bridge the gap between VAEs and flows\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Vol\.33,pp\. 12685–12696\.External Links:2007\.02731Cited by:[§1](https://arxiv.org/html/2608.12438#S1.p3.1)\.
- \[58\]L\. Onsager and S\. Machlup\(1953\)Fluctuations and irreversible processes\.Physical Review91\(6\),pp\. 1505–1512\.External Links:[Document](https://dx.doi.org/10.1103/PhysRev.91.1505)Cited by:[§1](https://arxiv.org/html/2608.12438#S1.p2.1)\.
- \[59\]J\. Pan, J\. H\. Liew, V\. Y\. F\. Tan, J\. Feng, and H\. Yan\(2024\)AdjointDPM: adjoint sensitivity method for gradient backpropagation of diffusion probabilistic models\.InInternational Conference on Learning Representations,External Links:2307\.10711Cited by:[§4\.2](https://arxiv.org/html/2608.12438#S4.SS2.SSSx1.p2.1)\.
- \[60\]T\. Plehn, A\. Butter, B\. Dillon, T\. Heimel, C\. Krause, and R\. Winterhalder\(2022\)Modern Machine Learning for LHC Physicists\.External Links:2211\.01421Cited by:[§1](https://arxiv.org/html/2608.12438#S1.p1.1)\.
- \[61\]R\. A\. Porto\(2016\)The effective field theorist’s approach to gravitational dynamics\.Phys\. Rept\.633,pp\. 1–104\.External Links:1601\.04914,[Document](https://dx.doi.org/10.1016/j.physrep.2016.04.003)Cited by:[§5\.1](https://arxiv.org/html/2608.12438#S5.SS1.p6.1)\.
- \[62\]A\. Premkumar\(2023\)Generative diffusion from an action principle\.External Links:2310\.04490Cited by:[§1](https://arxiv.org/html/2608.12438#S1.p3.1)\.
- \[63\]M\. Reed and B\. Simon\(1980\)Methods of modern mathematical physics i: functional analysis\.Revised and enlarged edition,Academic Press\.Cited by:[§3\.2](https://arxiv.org/html/2608.12438#S3.SS2.p4.8)\.
- \[64\]H\. Risken\(1989\)The Fokker–Planck equation: methods of solution and applications\.2 edition,Springer Series in Synergetics, Vol\.18,Springer,Berlin, Heidelberg\.External Links:[Document](https://dx.doi.org/10.1007/978-3-642-61544-3)Cited by:[§2\.2](https://arxiv.org/html/2608.12438#S2.SS2.p2.5),[§2\.2](https://arxiv.org/html/2608.12438#S2.SS2.p3.12)\.
- \[65\]K\. Roth, A\. Lucchi, S\. Nowozin, and T\. Hofmann\(2017\)Stabilizing training of generative adversarial networks through regularization\.InAdvances in Neural Information Processing Systems,Vol\.30,pp\. 2018–2028\.External Links:1705\.09367Cited by:[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx5.p3.1)\.
- \[66\]V\. G\. Satorras, E\. Hoogeboom, and M\. Welling\(2021\)E\(n\) equivariant graph neural networks\.External Links:2102\.09844Cited by:[§5\.2](https://arxiv.org/html/2608.12438#S5.SS2.p1.4)\.
- \[67\]Y\. Shi, V\. De Bortoli, A\. Campbell, and A\. Doucet\(2023\)Diffusion Schrödinger bridge matching\.Advances in Neural Information Processing Systems36\.External Links:2303\.16852Cited by:[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx4.p1.1),[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx4.p3.2)\.
- \[68\]Y\. Song, C\. Durkan, I\. Murray, and S\. Ermon\(2021\)Maximum likelihood training of score\-based diffusion models\.InAdvances in Neural Information Processing Systems,Vol\.34,pp\. 1415–1428\.External Links:2101\.09258Cited by:[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx2.p2.9),[§4\.2](https://arxiv.org/html/2608.12438#S4.SS2.SSSx1.p1.3)\.
- \[69\]Y\. Song, J\. Sohl\-Dickstein, D\. P\. Kingma, A\. Kumar, S\. Ermon, and B\. Poole\(2021\)Score\-based generative modeling through stochastic differential equations\.External Links:2011\.13456Cited by:[§1](https://arxiv.org/html/2608.12438#S1.p3.1),[§2\.2](https://arxiv.org/html/2608.12438#S2.SS2.p2.3)\.
- \[70\]R\. L\. Stratonovich\(1957\)On a Method of Calculating Quantum Distribution Functions\.Soviet Physics Doklady2,pp\. 416\.Cited by:[§3\.1](https://arxiv.org/html/2608.12438#S3.SS1.p3.6)\.
- \[71\]U\. C\. Täuber\(2014\)Critical Dynamics: A Field Theory Approach to Equilibrium and Non\-Equilibrium Scaling Behavior\.Cambridge University Press,Cambridge\.External Links:[Document](https://dx.doi.org/10.1017/CBO9781139046213)Cited by:[§3\.1](https://arxiv.org/html/2608.12438#S3.SS1.p3.5),[§3\.1](https://arxiv.org/html/2608.12438#S3.SS1.p4.1)\.
- \[72\]R\. Verheyen\(2022\)Event Generation and Density Estimation with Surjective Normalizing Flows\.SciPost Phys\.13\(3\),pp\. 047\.External Links:2205\.01697,[Document](https://dx.doi.org/10.21468/SciPostPhys.13.3.047)Cited by:[§1](https://arxiv.org/html/2608.12438#S1.p3.1)\.
- \[73\]P\. Vincent\(2011\)A connection between score matching and denoising autoencoders\.Neural Computation23\(7\),pp\. 1661–1674\.External Links:[Document](https://dx.doi.org/10.1162/NECO%5Fa%5F00142)Cited by:[§2\.3](https://arxiv.org/html/2608.12438#S2.SS3.SSSx2.p2.8)\.
- \[74\]S\. Weinberg\(1979\)Phenomenological Lagrangians\.Physica A96\(1\-2\),pp\. 327–340\.External Links:[Document](https://dx.doi.org/10.1016/0378-4371%2879%2990223-1)Cited by:[§5\.1](https://arxiv.org/html/2608.12438#S5.SS1.p5.1)\.
- \[75\]S\. Weinberg\(1990\)Nuclear forces from chiral Lagrangians\.Phys\. Lett\. B251,pp\. 288–292\.External Links:[Document](https://dx.doi.org/10.1016/0370-2693%2890%2990938-3)Cited by:[§5\.1](https://arxiv.org/html/2608.12438#S5.SS1.p5.1)\.
- \[76\]H\. Weyl\(1946\)The classical groups: their invariants and representations\.2nd edition,Princeton University Press\.Cited by:[§5\.2](https://arxiv.org/html/2608.12438#S5.SS2.p2.1)\.
- \[77\]M\. Zaheer, S\. Kottur, S\. Ravanbakhsh, B\. Póczos, R\. Salakhutdinov, and A\. J\. Smola\(2017\)Deep sets\.30\.External Links:1703\.06114Cited by:[§5\.2](https://arxiv.org/html/2608.12438#S5.SS2.p1.4)\.
- \[78\]J\. Zinn\-Justin\(2002\)Quantum field theory and critical phenomena\.Int\. Ser\. Monogr\. Phys\.113,pp\. 1–1054\.Cited by:[§2\.2](https://arxiv.org/html/2608.12438#S2.SS2.SSSx1.p1.1),[§3\.1](https://arxiv.org/html/2608.12438#S3.SS1.p3.5),[§3\.1](https://arxiv.org/html/2608.12438#S3.SS1.p4.1),[§3\.2](https://arxiv.org/html/2608.12438#S3.SS2.p5.1)\.

Similar Articles

All in One: Generative Modeling as Mean-Field Game Design

arXiv cs.LG

This paper unifies twelve continuous-time generative models under mean-field game theory via a cost tuple, introduces MFGLab (a PyTorch library that auto-shares training loops and solvers), and proposes DI-Flow with differentiable entropy for better mode coverage.

Energy Generative Modeling: A Lyapunov-based Energy Matching Perspective

arXiv cs.LG

This paper proposes a unified framework for energy-based generative models by casting density transport as a nonlinear control problem with KL divergence as a Lyapunov function. It derives finite-step stopping criteria and demonstrates how nonlinear control theory tools can be applied to static scalar energy models.

Perron--Frobenius Operator Matching for Generative Modeling

arXiv cs.LG

Introduces Perron–Frobenius Operator Matching (PFOM), a generative framework that unifies flow, diffusion, and jump models via integral PF operator matching, proving KL divergence yields a practical loss equivalent to Koopman path matching, and develops Nesterov-accelerated training and sampling for improved efficiency.