Flow Map Learning via Nongradient Vector Flow

arXiv cs.LG Papers

Summary

This paper introduces SGFlow, a method for learning flow maps for diffusion models that avoids invertibility constraints and backpropagation through model iterations, achieving competitive FID scores on CIFAR with a proven stationary-point guarantee.

arXiv:2607.26398v1 Announce Type: new Abstract: Diffusion and flow-based models benefit from simple regression losses, but inference incurs significant overhead because sampling requires integration. Consistency models address this by directly learning the flow maps along the ODE trajectory, opening a design space between one-step and many-step approaches. However, existing methods face computational challenges such as requiring model inverses or backpropagation through iterated model calls, and do not always prove that the desired ODE flow map is a solution to the loss. We introduce SGFlow, an approach for learning flow maps that bypasses explicit invertibility constraints and expensive differentiation through model iteration. SGFlow trains a model to compute both the ODE solutions and the implied velocity from scratch by following non-conservative dynamics with a stationary point at the desired flow map. On the CIFAR image benchmark, no single method attains the best FID at every step count: SGFlow attains the best FID at 10 sampling steps and remains competitive with flow matching, Meanflow, and Lagrangian map matching at other step counts, while being the only one with a proven stationary-point guarantee for its stopgrad-based dynamics.
Original Article
View Cached Full Text

Cached at: 07/30/26, 09:58 AM

# Flow Map Learning via Nongradient Vector Flow
Source: [https://arxiv.org/html/2607.26398](https://arxiv.org/html/2607.26398)
Mark Goldstein1, Anshuk Uppal2, Raghav Singhal3, Aahlad Puli3, & Rajesh Ranganath3 1Center for Computational Mathematics, Flatiron Institute\. 2Department of Applied Mathematics and Computer Science, Technical University of Denmark\. 3Courant Institute, New York University\.

###### Abstract

Diffusion and flow\-based models benefit from simple regression losses, but inference incurs significant overhead because sampling requires integration\. Consistency models address this by directly learning the flow maps along the ODE trajectory, opening a design space between one\-step and many\-step approaches\. However, existing methods face computational challenges such as requiring model inverses or backpropagation through iterated model calls, and do not always prove that the desired ODE flow map is a solution to the loss\. We introduce SGFlow, an approach for learning flow maps that bypasses explicit invertibility constraints and expensive differentiation through model iteration\. SGFlow trains a model to compute both the ODE solutions and the implied velocity from scratch by following non\-conservative dynamics with a stationary point at the desired flow map\. On the CIFAR image benchmark, no single method attains the best FID at every step count: SGFlow attains the best FID at 10 sampling steps and remains competitive with flow matching, Meanflow, and Lagrangian map matching at other step counts, while being the only one with a proven stationary\-point guarantee for its stopgrad\-based dynamics\.

## 1Introduction

Diffusion and flow models\(Sohl\-Dicksteinet al\.,[2015](https://arxiv.org/html/2607.26398#bib.bib22); Hoet al\.,[2020](https://arxiv.org/html/2607.26398#bib.bib29); Songet al\.,[2020](https://arxiv.org/html/2607.26398#bib.bib23); Kingmaet al\.,[2021](https://arxiv.org/html/2607.26398#bib.bib10); Albergo and Vanden\-Eijnden,[2022](https://arxiv.org/html/2607.26398#bib.bib20); Singhalet al\.,[2023](https://arxiv.org/html/2607.26398#bib.bib30); Pandey and Mandt,[2023](https://arxiv.org/html/2607.26398#bib.bib32); Bartoshet al\.,[2024](https://arxiv.org/html/2607.26398#bib.bib33); Singhalet al\.,[2024](https://arxiv.org/html/2607.26398#bib.bib31); Albergoet al\.,[2023](https://arxiv.org/html/2607.26398#bib.bib21); Lipmanet al\.,[2022](https://arxiv.org/html/2607.26398#bib.bib46); Liuet al\.,[2022](https://arxiv.org/html/2607.26398#bib.bib47)\)have improved generation in domains such as proteins\(Abramsonet al\.,[2024](https://arxiv.org/html/2607.26398#bib.bib19)\)and images\(Peebles and Xie,[2023](https://arxiv.org/html/2607.26398#bib.bib27); Esseret al\.,[2024](https://arxiv.org/html/2607.26398#bib.bib13)\)\. Sampling from these models typically requires numerically integrating an ordinary or stochastic differential equation\. Numerical integration requires multiple forward passes of a neural network, leading to increased sampling latency and cost\.

To ameliorate this generation cost by changing the training, recent approaches for consistency modeling and map matching\(Songet al\.,[2023](https://arxiv.org/html/2607.26398#bib.bib9); Song and Dhariwal,[2023](https://arxiv.org/html/2607.26398#bib.bib8); Kimet al\.,[2023](https://arxiv.org/html/2607.26398#bib.bib5); Lu and Song,[2024](https://arxiv.org/html/2607.26398#bib.bib11); Boffiet al\.,[2024](https://arxiv.org/html/2607.26398#bib.bib44);[2025](https://arxiv.org/html/2607.26398#bib.bib38)\)aim to learn direct mappings from noise to intermediate or final data points along trajectories defined by probability flow ODEs, thereby avoiding costly integration\. However, the methods have their respective complexities\. For example, flow map matching requires model invertibility, while consistency models need either to map in one step or introduce extra steps that leave the target ODE trajectory\.

We introduce SGFlow \(for StopGrad Flow\), an approach that builds on flows and map matching methods and

- •Has true flow map as a unique stationary point
- •Does not restrict the class of neural networks used \(e\.g\., to invertible functions\)
- •Does not require auxiliary losses involving invertibility or adversarial optimization
- •Does not require optimizing through nested calls to the model
- •Allows for generation along the ODE trajectory with any number of steps

Existing methods for learning flow maps fall into a few categories in terms of their challenges; all challenges relate to the idea that a flow map is characterized by certain derivative properties and that losses minimize squared error to make these properties hold\. Flow map matching and related methods rely on a fundamental relationship between invertible mappings and ordinary differential equations \(ODEs\)\. This relationship typically requires explicitly computing both the forward map \(the model being trained\) and its inverse during training, complicating training, or requires expensive backpropagation through nested model calls\.Boffiet al\.\([2025](https://arxiv.org/html/2607.26398#bib.bib38)\)propose stopgrad placements for the flow map matching losses ofBoffiet al\.\([2024](https://arxiv.org/html/2607.26398#bib.bib44)\)that bypass this expensive nested differentiation, but do not prove that the stopgrads preserve the stationary point at the true flow map\. Meanflow\(Genget al\.,[2025](https://arxiv.org/html/2607.26398#bib.bib36)\)does not explicitly enforce the model inverse identities, and avoids backpropagation through forward\-mode derivatives altogether, achieving good image generation performance with low step counts; Meanflow is also not shown to have a stationary point at the true flow map\.

SGFlow avoids the complexity of tracking a model and its inverse by exploiting an alternate identity involving only Jacobian\-vector products \(JVPs\) without inverse functions\. This identity allows us to formulate the objective purely in terms of the forward map, without needing explicit access to its inverse\. Since solutions to ODEs naturally produce invertible mappings, the SGFlow objective implicitly encourages invertibility without explicitly enforcing it\. Thus, at optimality, SGFlow yields a continuously differentiable function that precisely integrates the velocity field, directly generating the desired data distribution\. We summarize the trade\-offs among recent methods in[Table˜1](https://arxiv.org/html/2607.26398#S1.T1)and in[Section˜6](https://arxiv.org/html/2607.26398#S6)\.

Experimentally, for a basic training setup using the same common architecture, we ask how flow matching, Meanflow, SGFlow, and Lagrangian map matching compare in moderate dimensions \(CIFAR\-10\) on unconditional metrics \(FID\) when decreasing the number of sampling steps\.

Table 1:Comparison to prior works\. We categorize flow map learning \(orconsistency modeling\) techniques and our proposed SGFlow method according to: \(1\) ability to adjust sampling steps post\-training, \(2\) whether they follow the PF\-ODE\(Songet al\.,[2020](https://arxiv.org/html/2607.26398#bib.bib23)\), \(3\) whether they allow simulation\-free training, \(4\) whether their objectives use regression, \(5\) whether training is free of model inversion, \(6\) whether the true flow map is proven to be optimal or stationary, and \(7\) whether training avoids differentiation through nested model calls\. See[Section˜6](https://arxiv.org/html/2607.26398#S6)for details\.
## 2Background

Stochastic interpolants\(Lipmanet al\.,[2022](https://arxiv.org/html/2607.26398#bib.bib46); Albergoet al\.,[2023](https://arxiv.org/html/2607.26398#bib.bib21)\), and more generally most diffusion and flow methods, hereafter justflows, pose generative modeling as transport of a simple base density to a target density\. Interpolants tackle the problem as follows\. Fort∈\[0,1\]t\\in\[0,1\]:

1. 1\.Choose \(αt\\alpha\_\{t\},σt\\sigma\_\{t\}\) whereα0=σ1=1\\alpha\_\{0\}=\\sigma\_\{1\}=1andα1=σ0=0\\alpha\_\{1\}=\\sigma\_\{0\}=0\. Commonly,αt=1−t\\alpha\_\{t\}=1\-tandσt=t\\sigma\_\{t\}=t\.
2. 2\.DefineXt=αt​X0\+σt​X1X\_\{t\}=\\alpha\_\{t\}X\_\{0\}\+\\sigma\_\{t\}X\_\{1\}for base densityX0∼q0X\_\{0\}\\sim q\_\{0\}and dataX1∼q1X\_\{1\}\\sim q\_\{1\}\(or vice versa\)\.
3. 3\.Learn to produce new samples along the trajectory of densities\.

For a functionff, letf˙t:=dd​t​ft\\dot\{f\}\_\{t\}:=\\frac\{d\}\{dt\}f\_\{t\}\. ThusX˙t:=α˙t​X0\+σ˙t​X1\\dot\{X\}\_\{t\}:=\\dot\{\\alpha\}\_\{t\}X\_\{0\}\+\\dot\{\\sigma\}\_\{t\}X\_\{1\}\. It follows thatXtX\_\{t\}has densityqtq\_\{t\}satisfying:

∂tqt​\(x\)=−∇x⋅\(qt​\(x\)​v​\(t,x\)\),v​\(t,x\):=𝔼​\[X˙t\|Xt=x\],\\displaystyle\\partial\_\{t\}q\_\{t\}\(x\)=\-\\nabla\_\{x\}\\cdot\(q\_\{t\}\(x\)v\(t,x\)\),\\quad\\quad v\(t,x\):=\\mathbb\{E\}\[\\dot\{X\}\_\{t\}~\|~X\_\{t\}=x\],\(1\)wherevvis called the velocity\. The PDE in[Equation˜1](https://arxiv.org/html/2607.26398#S2.E1)is derived in the above works\. To accomplish step three, one starts by making the observation that a density satisfies[Equation˜1](https://arxiv.org/html/2607.26398#S2.E1)if and only if it is the density of the solution to theprobability flow ODEd​X=v​d​tdX=vdtintegrated forward fromX0∼q0X\_\{0\}\\sim q\_\{0\}or in reverse fromX1∼q1X\_\{1\}\\sim q\_\{1\}\(Albergo and Vanden\-Eijnden,[2024](https://arxiv.org/html/2607.26398#bib.bib39)\)\. One then proceeds by first approximatingvvusing the following \(simulation\-free\) loss:

ℒv​\(vθ\)=𝔼​\[‖vθ​\(t,Xt\)−\(α˙t​X0\+σ˙t​X1\)‖2\]Xt=αt​X0\+σt​X1,\\displaystyle\\mathcal\{L\}\_\{v\}\(v\_\{\\theta\}\)=\\mathbb\{E\}\\Big\[\\\|v\_\{\\theta\}\(t,X\_\{t\}\)\-\(\\dot\{\\alpha\}\_\{t\}X\_\{0\}\+\\dot\{\\sigma\}\_\{t\}X\_\{1\}\)\\\|^\{2\}\\Big\]\_\{X\_\{t\}=\\alpha\_\{t\}X\_\{0\}\+\\sigma\_\{t\}X\_\{1\}\},\(2\)which has minimizervθ=vv\_\{\\theta\}=vand then solvingd​x=vθ​d​tdx=v\_\{\\theta\}dt\.

#### Background on Consistency Methods\.

Sampling from flows requires integration, where each step evaluates a neural networkvθv\_\{\\theta\}modeling a score, velocity, or similar\. Knowing integrals ofvvdirectly could, in principle, speed up sampling\. The goal of consistency and map matching methods is to learn to map along the trajectory implied by the optimalvv\. We review an example here, with others in[Section˜6](https://arxiv.org/html/2607.26398#S6)\.Songet al\.\([2023](https://arxiv.org/html/2607.26398#bib.bib9)\); Song and Dhariwal \([2023](https://arxiv.org/html/2607.26398#bib.bib8)\)seek to learn a mappingg^\\hat\{g\}that takes interpolant samplesXt∼qtX\_\{t\}\\sim q\_\{t\}toX^0\\widehat\{X\}\_\{0\}, thet=0t=0solution tod​x=v​d​tdx=v\\,dtstarting atXtX\_\{t\}\(note thatX^0\\widehat\{X\}\_\{0\}usually differs from the endpointX0X\_\{0\}used to drawXtX\_\{t\}\)\. The loss measures the distance between modeled outputs at two nearby points\. LetSG​\[g^\]\\text\{SG\}\[\\hat\{g\}\]indicate stopgrad\. Then:

Consistency​\(g^\):=𝔼q​\(Xt\)​\[dist​\(g^​\(t,Xt\),SG​\[g^\]​\(t−Δ​t,X^t−Δ​t\)\)\]\.\\displaystyle\\text\{Consistency\}\(\\hat\{g\}\):=\\mathbb\{E\}\_\{q\(X\_\{t\}\)\}\[\\text\{dist\}\(\\hat\{g\}\(t,X\_\{t\}\),\\text\{SG\}\[\\hat\{g\}\]\(t\-\\Delta t,\\widehat\{X\}\_\{t\-\\Delta t\}\)\)\]\.\(3\)The targetX^t−Δ​t\\widehat\{X\}\_\{t\-\\Delta t\}should come from integrating the true velocity a small stepΔ​t\\Delta tfromXtX\_\{t\}, but sincevvis unknown, it is typically approximated using a pretrainedvθv\_\{\\theta\}or avθv\_\{\\theta\}derived jointly withg^\\hat\{g\}— increasing training cost and introducing approximation error\. Allowing multistep sampling requires a re\-noising step that takes the trajectory off the probability\-flow ODE, so the resulting updates no longer correspond to integrating the PF\-ODE\.Kimet al\.\([2023](https://arxiv.org/html/2607.26398#bib.bib5)\)observe that this multistep approach “exhibits degrading sample quality with increasing NFE, lacking a clear trade\-off between computational budget \(NFE\) and sample fidelity”\. Subsequent works introduce various training\- and inference\-time modifications to bridge the gap between one\-step and many\-step sampling\(Songet al\.,[2023](https://arxiv.org/html/2607.26398#bib.bib9); Lu and Song,[2024](https://arxiv.org/html/2607.26398#bib.bib11); Kimet al\.,[2023](https://arxiv.org/html/2607.26398#bib.bib5); Boffiet al\.,[2024](https://arxiv.org/html/2607.26398#bib.bib44); Sabouret al\.,[2025](https://arxiv.org/html/2607.26398#bib.bib37); Genget al\.,[2025](https://arxiv.org/html/2607.26398#bib.bib36); Zhouet al\.,[2025](https://arxiv.org/html/2607.26398#bib.bib4)\); see[Section˜6](https://arxiv.org/html/2607.26398#S6)\.

## 3Method

We present SGFlow, a method for learning to solve the probability flow ODE without adversarial training, without model inverse during training, without representing explicit derivative matrices, and without costly simulations from pretrained models\. SGFlow trains a model to compute both the ODE solutions and the implied velocity from scratch by following non\-conservative dynamics\.

Consider a two\-time mapffthat fort≤ut\\leq ubringsXtX\_\{t\}up toXuX\_\{u\}by solving the probability flow ODEd​x=v​d​tdx=vdt\. Such anffthat integratesvvcan be defined as follows:

f​\(t,u,x\)\\displaystyle f\(t,u,x\)=x\+∫tuv​\(s,Xs\)​𝑑s=x\+∫tuv​\(s,f​\(t,s,x\)\)​𝑑s\\displaystyle=x\+\\int\_\{t\}^\{u\}v\(s,X\_\{s\}\)ds=x\+\\int\_\{t\}^\{u\}v\(s,f\(t,s,x\)\)ds\(4\)Differentiating the recursive form on the RHS w\.r\.t\.ttusing the total \(material\) derivative yields:

∂tf\+\(∂xf\)​v​\(t,x\)=0,f​\(u,u,x\)=x\\displaystyle\\partial\_\{t\}f\+\(\\partial\_\{x\}f\)v\(t,x\)=0,\\quad f\(u,u,x\)=x\(5\)This is uniquely solved at the true flow mapff\. We can square the left\-hand side for a parameterizedfθf\_\{\\theta\}and take an expectation\.XtX\_\{t\}is sampled by drawing dataX1X\_\{1\}, noiseX0X\_\{0\}, and computingXt=αt​X0\+σt​X1X\_\{t\}=\\alpha\_\{t\}X\_\{0\}\+\\sigma\_\{t\}X\_\{1\}:

L\\displaystyle L:=𝔼Xt\[∥∂tfθ\+\(∂xfθ\)v∥2\]=𝔼Xt\[∥∂tfθ\+\(∂xfθ\)𝔼\[X˙t\|Xt\]∥2\]\.\\displaystyle:=\\mathbb\{E\}\_\{X\_\{t\}\}\[\\\|\\partial\_\{t\}f\_\{\\theta\}\+\(\\partial\_\{x\}f\_\{\\theta\}\)v\\\|^\{2\}\]=\\mathbb\{E\}\_\{X\_\{t\}\}\[\\\|\\partial\_\{t\}f\_\{\\theta\}\+\(\\partial\_\{x\}f\_\{\\theta\}\)\\mathbb\{E\}\[\\dot\{X\}\_\{t\}\|X\_\{t\}\]\\\|^\{2\}\]\.\(6\)

The true mapffis the unique minimizer of this loss\. Usingv​\(t,x\)=𝔼​\[X˙t\|Xt=x\]v\(t,x\)=\\mathbb\{E\}\[\\dot\{X\}\_\{t\}\|X\_\{t\}=x\], we can expand,

L\\displaystyle L=𝔼Xt\[∥∂tfθ\+\(∂xfθ\)X˙t∥2−∥\(∂xfθ\)\(X˙t−𝔼\[X˙t\|Xt\]\)∥2\]\\displaystyle=\\mathbb\{E\}\_\{X\_\{t\}\}\[\\\|\\partial\_\{t\}f\_\{\\theta\}\+\(\\partial\_\{x\}f\_\{\\theta\}\)\\dot\{X\}\_\{t\}\\\|^\{2\}\-\\\|\(\\partial\_\{x\}f\_\{\\theta\}\)\(\\dot\{X\}\_\{t\}\-\\mathbb\{E\}\[\\dot\{X\}\_\{t\}\|X\_\{t\}\]\)\\\|^\{2\}\]\(7\)For an underlying modelf~θ\\tilde\{f\}\_\{\\theta\}, we can then use the parameterization

fθ​\(t,u,x\):=x\+\(u−t\)​f~θ​\(t,u,x\)\\displaystyle f\_\{\\theta\}\(t,u,x\):=x\+\(u\-t\)\\tilde\{f\}\_\{\\theta\}\(t,u,x\)\(8\)This parameterization automatically satisfies the boundary conditionfθ​\(u,u,x\)=xf\_\{\\theta\}\(u,u,x\)=x\. The parameterization yields two additional properties:

- •time derivative:∂tfθ​\(t,t,x\)=−f~θ​\(t,t,x\)\\partial\_\{t\}f\_\{\\theta\}\(t,t,x\)=\-\\tilde\{f\}\_\{\\theta\}\(t,t,x\)
- •Jacobian:∂xfθ​\(t,t,x\)=I\\partial\_\{x\}f\_\{\\theta\}\(t,t,x\)=I

Using these properties and evaluating att=ut=u, we see that the minimization of[eq\.˜7](https://arxiv.org/html/2607.26398#S3.E7)reduces to flow matching wheref~θ​\(t,t,x\)\\tilde\{f\}\_\{\\theta\}\(t,t,x\)is trained to match the velocity:

L\|t=u=𝔼Xt​\[‖f~θ​\(t,t,Xt\)−X˙t‖2\],\\displaystyle L\\big\|\_\{t=u\}=\\mathbb\{E\}\_\{X\_\{t\}\}\[\\\|\\tilde\{f\}\_\{\\theta\}\(t,t,X\_\{t\}\)\-\\dot\{X\}\_\{t\}\\\|^\{2\}\],\(9\)which reveals that for the trueff, we have that

−∂tf​\(t,t,⋅\)=f~​\(t,t,⋅\)=v​\(t,x\)=𝔼​\[X˙t\|Xt=x\]\\displaystyle\-\\partial\_\{t\}f\(t,t,\\cdot\)=\\tilde\{f\}\(t,t,\\cdot\)=v\(t,x\)=\\mathbb\{E\}\[\\dot\{X\}\_\{t\}\|X\_\{t\}=x\]\(10\)This motivates replacing the unknownvvin[eq\.˜7](https://arxiv.org/html/2607.26398#S3.E7)withstopgrad​\[f~θ​\(t,t,⋅\)\]\\text\{stopgrad\}\[\\tilde\{f\}\_\{\\theta\}\(t,t,\\cdot\)\]\. The stopgrad is used under the principle that since the originalvvdid not provide gradient updates forff, neither should a term that approximates it\. LetSG​\(\)\\text\{SG\}\(\)denote thestopgrad​\(\)\\text\{stopgrad\}\(\)operator\. TheSGFlowmethod follows parameter updates toθ\\thetaby differentiating

Lsg:=𝔼Xt​\[‖\(∂tfθ\)\+\(∂xfθ\)​X˙t‖2−‖\(∂xfθ\)​\(X˙t−SG​\[f~θ\]\)‖2\]\.\\displaystyle L\_\{\\text\{sg\}\}:=\\mathbb\{E\}\_\{X\_\{t\}\}\[\\\|\(\\partial\_\{t\}f\_\{\\theta\}\)\+\(\\partial\_\{x\}f\_\{\\theta\}\)\\dot\{X\}\_\{t\}\\\|^\{2\}\-\\\|\(\\partial\_\{x\}f\_\{\\theta\}\)\(\\dot\{X\}\_\{t\}\-\\text\{SG\}\[\\tilde\{f\}\_\{\\theta\}\]\)\\\|^\{2\}\]\.\(11\)

Here, the modelfθf\_\{\\theta\}uses the parameterization from[eq\.˜8](https://arxiv.org/html/2607.26398#S3.E8)\. The derivatives∂tfθ\\partial\_\{t\}f\_\{\\theta\}and∂xfθ\\partial\_\{x\}f\_\{\\theta\}are evaluated at\(t,u,Xt\)\(t,u,X\_\{t\}\)andf~θ\\tilde\{f\}\_\{\\theta\}is evaluated at\(t,t,Xt\)\(t,t,X\_\{t\}\)\. The expectation is taken overXtX\_\{t\}sampled by drawing dataX1X\_\{1\}, noiseX0X\_\{0\}, and computingXt=αt​X0\+σt​X1X\_\{t\}=\\alpha\_\{t\}X\_\{0\}\+\\sigma\_\{t\}X\_\{1\}andX˙t=α˙t​X0\+σ˙t​X1\\dot\{X\}\_\{t\}=\\dot\{\\alpha\}\_\{t\}X\_\{0\}\+\\dot\{\\sigma\}\_\{t\}X\_\{1\}\.

In practice, the PDE in[eq\.˜5](https://arxiv.org/html/2607.26398#S3.E5)must hold for all pairst≤ut\\leq u\. Letq​\(t,u\)q\(t,u\)be a joint distribution with support overt≤ut\\leq uand with positive probability ont=ut=u\. We take expectations over time and define

ℒ=𝔼q​\(t,u\)​\[L\],ℒsg=𝔼q​\(t,u\)​\[Lsg\]\\displaystyle\\mathcal\{L\}=\\mathbb\{E\}\_\{q\(t,u\)\}\[L\],\\quad\\mathcal\{L\}\_\{\\text\{sg\}\}=\\mathbb\{E\}\_\{q\(t,u\)\}\[L\_\{\\text\{sg\}\}\]We now connectℒ\\mathcal\{L\}andℒsg\\mathcal\{L\}\_\{\\text\{sg\}\}formally\.[Theorem˜1](https://arxiv.org/html/2607.26398#Thmtheorem1)shows that parameter updates ofℒ\\mathcal\{L\}andℒsg\\mathcal\{L\}\_\{\\text\{sg\}\}are0at the same solutions\.

###### Theorem 1\. Letq​\(t,u\)q\(t,u\)be a joint distribution over time pairs with support overt≤ut\\leq uand with positive probability ont=ut=u\. Let the familyℱ~\\mathcal\{\\tilde\{F\}\}include functionsf~\\tilde\{f\}that are continuously differentiable in all arguments\. LetXt=αt​X0\+σt​X1X\_\{t\}=\\alpha\_\{t\}X\_\{0\}\+\\sigma\_\{t\}X\_\{1\}andX˙t=α˙t​X0\+σ˙t​X1\\dot\{X\}\_\{t\}=\\dot\{\\alpha\}\_\{t\}X\_\{0\}\+\\dot\{\\sigma\}\_\{t\}X\_\{1\}\. Definef​\(t,u,x\):=x\+\(u−t\)​f~​\(t,u,x\)f\(t,u,x\):=x\+\(u\-t\)\\tilde\{f\}\(t,u,x\)\. Let expectations be computed overq​\(X0\)​q​\(X1\)q\(X\_\{0\}\)q\(X\_\{1\}\)\. Lets​gsgstand for stop\-gradient\. Defineℒ=𝔼q​\(t,u\)​\[L\]\\mathcal\{L\}=\\mathbb\{E\}\_\{q\(t,u\)\}\[L\]andℒsg=𝔼q​\(t,u\)​\[Lsg\]\\mathcal\{L\}\_\{\\text\{sg\}\}=\\mathbb\{E\}\_\{q\(t,u\)\}\[L\_\{\\text\{sg\}\}\]\. Thenf~∗\\tilde\{f\}^\{\*\}is a stationary point ofℒsg\\mathcal\{L\}\_\{\\text\{sg\}\}with respect toℱ~\\mathcal\{\\tilde\{F\}\}if and only iff~∗\\tilde\{f\}^\{\*\}is a stationary point ofℒ\\mathcal\{L\}with respect toℱ~\\mathcal\{\\tilde\{F\}\}\.

This is shown in[Section˜A\.5](https://arxiv.org/html/2607.26398#A1.SS5)\.

#### Intuition\.

The purpose of the theorem is to establish thatℒsg\\mathcal\{L\}\_\{\\text\{sg\}\}has the same set of solutions asℒ\\mathcal\{L\}despite not having access tovv; here,solutionis defined as a stationary point\. The intuition is that, despite the stopgrad, whent=ut=u,ℒsg\\mathcal\{L\}\_\{\\text\{sg\}\}tries to match the velocity\. We show thatℒsg\\mathcal\{L\}\_\{\\text\{sg\}\}is not at a stationary point when this velocity estimate is inaccurate, so the optimization continues moving and does not become stuck at functions that integrate \(“distill"\) an incorrect velocity\. As this match improves so does the match between the parameter updates fromℒsg\\mathcal\{L\}\_\{\\text\{sg\}\}andℒ\\mathcal\{L\}att≠ut\\neq u\. The main reason this works is thatf~​\(t,t,⋅\)\\tilde\{f\}\(t,t,\\cdot\)appears in other terms outside of the stopgrad, and those terms tell it where to go\. This is crucial and not all stopgrad optimizations benefit from this property\.

#### Computation\.

Both terms inℒsg\\mathcal\{L\}\_\{\\text\{sg\}\}can be computed as expected squared norms of Jacobian\-vector products \(JVPs\), which use forward\-mode autodifferentiation to avoid explicitly materializing Jacobians, saving memory\. Let the derivatives\(∂tf,∂uf,∂xf\)\(\\partial\_\{t\}f,\\partial\_\{u\}f,\\partial\_\{x\}f\)be evaluated at\(t,u,x\)\(t,u,x\)\. Using PyTorch notation,

JVP​\[f,\(t,u,x\),\(a,b,c\)\]:=\(∂tf\)⋅a\+\(∂uf\)⋅b\+\(∂xf\)⋅c\\displaystyle\\text\{JVP\}\[f,\(t,u,x\),\(a,b,c\)\]:=\(\\partial\_\{t\}f\)\\cdot a\+\(\\partial\_\{u\}f\)\\cdot b\+\(\\partial\_\{x\}f\)\\cdot cFor the first loss term,a=1a=1,b=0b=0, andc=X˙tc=\\dot\{X\}\_\{t\}\. For the second term:

a=0,b=0,c=X˙t\+SG​\[∂tf​\(t,t,Xt\)\]=X˙t−SG​\[f~​\(t,t,Xt\)\]\.\\displaystyle a=0,\\quad b=0,\\quad c=\\dot\{X\}\_\{t\}\+\\text\{SG\}\[\\partial\_\{t\}f\(t,t,X\_\{t\}\)\]=\\dot\{X\}\_\{t\}\-\\text\{SG\}\[\\tilde\{f\}\(t,t,X\_\{t\}\)\]\.Though we have two distinct JVPs, we can split the batch and randomly assign either set of\(a,b,c\)\(a,b,c\)values to each batch element\.

#### Nongradient Flow Most analyses in machine learning implicitly assume that optimization performs gradient descent on a scalar loss, so that stationary points coincide with minima\. This assumption does not hold here: following the update rules ofℒsg\\mathcal\{L\}\_\{\\text\{sg\}\}does not correspond to following the gradients of any single scalar objective, as we demonstrate in[section˜B\.1](https://arxiv.org/html/2607.26398#A2.SS1)\. The stopgrad structure breaks the symmetry required for the updates to be expressible as the gradient of a scalar function; such broken symmetry renders the dynamicsnon\-conservative\(Balduzziet al\.,[2018](https://arxiv.org/html/2607.26398#bib.bib40)\), meaning they do not necessarily minimize any quantity\. Thus, in the limit of small step size, the optimization dynamics correspond to a nongradient vector flow, where it is crucial to analyze stationary points rather than minima\.

## 4When Can Stopgrads Fix Things?

The stopgrad in SGFlow is not just a computational convenience, but is essential for correctness\. To see why, consider what happens without it\. Meanflow\(Genget al\.,[2025](https://arxiv.org/html/2607.26398#bib.bib36)\)and SGFlow both replace the unknown conditional velocityv​\(t,x\)=𝔼\[X˙t\|Xt=x\]v\(t,x\)=\\mathop\{\\mathbb\{E\}\}\[\\dot\{X\}\_\{t\}~\|~X\_\{t\}=x\]with the sample\-levelX˙t\\dot\{X\}\_\{t\}\. In standard flow matching this substitution is harmless: the loss is linear inX˙t\\dot\{X\}\_\{t\}, so pulling the conditional expectation out of the square only adds a constant, the posterior variance,

V​\(t,x\):=𝕍​ar​\(X˙t\|Xt=x\),\\displaystyle V\(t,x\):=\\mathbb\{V\}\\textrm\{ar\}\(\\dot\{X\}\_\{t\}~\|~X\_\{t\}=x\),\(12\)and does not move the optimum\. SGFlow’s second term compensates for this\. In Meanflow without stopgrads, the term∂xf~​v\\partial\_\{x\}\\tilde\{f\}\\,vmakes the integrand*nonlinear*inX˙t\\dot\{X\}\_\{t\}\. Thev→X˙tv\\to\\dot\{X\}\_\{t\}replacement, without a compensating term, then introduces cross\-terms involvingV\>0V\>0that shift the gradient away from the truth\. By contrast, with the stopgrad,X˙t\\dot\{X\}\_\{t\}shows up linearly in the Meanflow parameter updates, meaning the expected update is equal to the correct one withvv\. This was also used to show correctness inSabouret al\.\([2025](https://arxiv.org/html/2607.26398#bib.bib37)\)\.

#### Gaussian example\.

We make this concrete in 1D\. LetX0∼𝒩​\(0,1\)X\_\{0\}\\sim\\mathcal\{N\}\(0,1\)andX1∼𝒩​\(μ,λ2\)X\_\{1\}\\sim\\mathcal\{N\}\(\\mu,\\lambda^\{2\}\)withλ\>0\\lambda\>0\. Let us use the linear interpolantXt=\(1−t\)​X0\+t​X1X\_\{t\}=\(1\-t\)X\_\{0\}\+tX\_\{1\}withX˙t=X1−X0\\dot\{X\}\_\{t\}=X\_\{1\}\-X\_\{0\}\. In this setting everything is in closed form \(Appendix[B\.2](https://arxiv.org/html/2607.26398#A2.SS2)\): the marginal variance isσt2=\(1−t\)2\+t2​λ2\\sigma\_\{t\}^\{2\}=\(1\-t\)^\{2\}\+t^\{2\}\\lambda^\{2\}, the flow map Jacobian is∂xf∗=m​\(t,u\)=σu/σt\\partial\_\{x\}f^\{\*\}=m\(t,u\)=\\sigma\_\{u\}/\\sigma\_\{t\}, and the posterior variance isV​\(t\)=λ2/σt2\>0V\(t\)=\\lambda^\{2\}/\\sigma\_\{t\}^\{2\}\>0\.

Consider the Meanflow functional without stopgrads,

ℒnosg​\[f~\]=𝔼\[‖f~−X˙t−\(u−t\)​\(∂tf~\+∂xf~​X˙t\)‖2\]\.\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{nosg\}\}\[\\tilde\{f\}\]=\\mathop\{\\mathbb\{E\}\}\\Big\[\\big\\\|\\tilde\{f\}\-\\dot\{X\}\_\{t\}\-\(u\-t\)\\big\(\\partial\_\{t\}\\tilde\{f\}\+\\partial\_\{x\}\\tilde\{f\}\\,\\dot\{X\}\_\{t\}\\big\)\\big\\\|^\{2\}\\Big\]\.\(13\)Evaluating the first variation at the true flow mapf~∗\\tilde\{f\}^\{\*\}in the directionh​\(t,u,x\)=xh\(t,u,x\)=xgives \(see Appendix[B\.2](https://arxiv.org/html/2607.26398#A2.SS2)for the calculation\):

δ​ℒnosg​\[f~∗;h=x\]=2​λ2​𝔼t,u\[\(u−t\)​m​\(t,u\)σt2\]\>0\.\\displaystyle\\delta\\mathcal\{L\}\_\{\\mathrm\{nosg\}\}\[\\tilde\{f\}^\{\*\};\\,h=x\]=2\\lambda^\{2\}\\,\\mathop\{\\mathbb\{E\}\}\_\{t,u\}\\\!\\left\[\\frac\{\(u\-t\)\\,m\(t,u\)\}\{\\sigma\_\{t\}^\{2\}\}\\right\]\>0\.\(14\)Every factor in the integrand is strictly positive, so the gradient at the truth is nonzero—the true flow map is*not*a stationary point of the no\-stopgrad functional\. With the stopgrad, Theorem[1](https://arxiv.org/html/2607.26398#Thmtheorem1)guarantees the true flow map is the unique stationary point\.

## 5Experiments

### 5\.1Gaussian illustration: Meanflow with and without stopgrads

As shown in §[4](https://arxiv.org/html/2607.26398#S4), removing stopgrads from Meanflow destroys stationarity at the true flow map\. We illustrate this concretely in the 1D Gaussian setting\. We parameterize the target distribution asX1∼𝒩​\(μ,λ2\)X\_\{1\}\\sim\\mathcal\{N\}\(\\mu,\\lambda^\{2\}\)and study the parameter update field of the Meanflow functional over the parameter space\(μ,λ\)\(\\mu,\\lambda\)\. The model family is well\-specified:

f~μ^,λ^=exact flow map for​𝒩​\(0,1\)→𝒩​\(μ^,λ^2\)\\displaystyle\\tilde\{f\}\_\{\\hat\{\\mu\},\\hat\{\\lambda\}\}=\\textrm\{exact flow map for \}\\mathcal\{N\}\(0,1\)\\to\\mathcal\{N\}\(\\hat\{\\mu\},\\hat\{\\lambda\}^\{2\}\)\(15\)The truth\(μ∗,λ∗\)\(\\mu^\{\*\},\\lambda^\{\*\}\)lies in the model family\. All expectations admit closed\-form expressions in the Gaussian case with linear interpolant\.

We plot the update fields in[Figure˜1](https://arxiv.org/html/2607.26398#S5.F1)\. With the stopgrad, the true flow map is the unique stationary point, so the negative update arrows converge to the truth\. Without the stopgrad, the true parameters are not stationary: the update field, which is a gradient field, is nonzero at the truth, and the field drifts toward a spurious point\. Due to the form of the Gaussian family, it happens thatμ^spurious=μ∗\\hat\{\\mu\}^\{\\text\{spurious\}\}=\\mu^\{\*\}exactly, with the bias entirely alongλ^\\hat\{\\lambda\}\.

![Refer to caption](https://arxiv.org/html/2607.26398v1/images/meanflow.png)Figure 1:Update vector field of the Meanflow functionalover the well\-specified Gaussian family, with data\(μ∗,λ∗\)=\(0\.5,1\.0\)\(\\mu^\{\*\},\\lambda^\{\*\}\)=\(0\.5,1\.0\)\(star,⋆\\star\)\. We use∇ℒ\\nabla\\mathcal\{L\}to denote the update field, though it is not a gradient field in the stopgrad case\. Arrows are unit\-normalized; color encodes‖∇ℒ‖\\\|\\nabla\\mathcal\{L\}\\\|on log scale\.Left: with stopgrad\. The truth is the unique stationary point; arrows converge to⋆\\starfrom all directions, and the update magnitude vanishes to machine precision at the truth\.Right: without stopgrad\. The truth is not stationary; the field flows past⋆\\starand converges to a spurious attractor at\(μ^,λ^\)≈\(0\.50,0\.41\)\(\\hat\{\\mu\},\\hat\{\\lambda\}\)\\approx\(0\.50,0\.41\)\(diamond,◆\\blacklozenge\)\. Due to the form of the Gaussian family, it happens thatμ^spurious=μ∗\\hat\{\\mu\}^\{\\text\{spurious\}\}=\\mu^\{\*\}exactly, with the bias entirely along theλ^\\hat\{\\lambda\}component\.
### 5\.2Image Modeling on CIFAR\-10

#### Architecture\.

We modify the time embedding of the existing diffusion U\-Net fromDhariwal and Nichol \([2021](https://arxiv.org/html/2607.26398#bib.bib26)\)to handle two times\. We embed both timesttanduuwith the usual Fourier embeddings, concatenate them, and pass the result through a small feedforward network that outputs a hidden representation for use in the usual U\-Net\. We use 128 U\-Net channels and channel multipliers set to \(1,2,2,2\) with attention bools set to \(False, False, True, False\)\.

#### Training Settings\.

We use dropout0\.10\.1\. We choose the unconditional image generative modeling task \(we do not condition on the class label\)\. We useαt=1−t\\alpha\_\{t\}=1\-tandσt=t\\sigma\_\{t\}=t, with noise atX0X\_\{0\}and data atX1X\_\{1\}\. We train for 200,000 steps at learning rate2×10−42\\times 10^\{\-4\}\. The FID is monitored every 25,000 training steps, and the best is reported\.

#### Losses\.

We benchmark based on Flow Matching\(Lipmanet al\.,[2022](https://arxiv.org/html/2607.26398#bib.bib46)\)as the reference method for sampling in theO​\(100\)O\(100\)step regime\. For few\-step sampling, we benchmark with Meanflow\(Genget al\.,[2025](https://arxiv.org/html/2607.26398#bib.bib36)\), which stopgrads all model derivatives in the loss to avoid backpropagation through differentiation\. We also benchmark against Lagrangian flow map matching\(Boffiet al\.,[2024](https://arxiv.org/html/2607.26398#bib.bib44)\)using the stopgrad placement suggested inBoffiet al\.\([2025](https://arxiv.org/html/2607.26398#bib.bib38)\)\. We do not compare against the Eulerian self\-distillation \(ESD\) loss, whichBoffiet al\.\([2025](https://arxiv.org/html/2607.26398#bib.bib38)\)report as unstable for image experiments, nor against the Progressive self\-distillation \(PSD\) loss, which enforces the semigroup identity rather than a transport PDE\. We focus here on the PDE\-based functionals\.

Table 2:FID scores versus sampling steps on CIFAR\-10, computed from 50,000 EMA samples\. The FID is monitored every 25,000 training steps up to 200,000 training steps, and the best is reported\. Lower is better;boldindicates best per column\. The “SG theory” column indicates whether the method has a proven stationary\-point guarantee for its stopgrad\-based dynamics \(i\.e\., whether the true flow map has been proven to be a stationary point of the optimization\)\.
#### Results\.

We report the Fréchet Inception Distance \(FID\)\(Heuselet al\.,[2017](https://arxiv.org/html/2607.26398#bib.bib35)\)in[Table˜2](https://arxiv.org/html/2607.26398#S5.T2)\. No single method dominates across all step counts: Meanflow is strongest at 1 step, the Lagrangian loss is strongest at 50 and 100 steps, and SGFlow attains the best FID at 10 steps\. SGFlow is the only method in the comparison whose stationary points under stopgrad\-induced dynamics are characterized, while remaining competitive with, and at 10 steps surpassing, methods that do not provide such guarantees\. This suggests that the stationary\-point analysis we develop does not come at a cost in empirical performance\.

## 6Related work

Sampling from continuous\-time generative models such as diffusion and flow models requires numerical integration\. Each integration step requires a forward pass of a neural network, leading to computational costs and slow sampling\. Current approaches to address this cost can be broadly categorized into two types: \(1\) distilling a pretrained diffusion or flow model into a few\-step solver\(Salimans and Ho,[2022](https://arxiv.org/html/2607.26398#bib.bib7); Kimet al\.,[2023](https://arxiv.org/html/2607.26398#bib.bib5); Liuet al\.,[2023](https://arxiv.org/html/2607.26398#bib.bib16)\), and \(2\) learning a few\-step solver\(Zhouet al\.,[2025](https://arxiv.org/html/2607.26398#bib.bib4)\)\. Some approaches in this area allow for distillation as well as training from scratch\(Songet al\.,[2023](https://arxiv.org/html/2607.26398#bib.bib9); Boffiet al\.,[2024](https://arxiv.org/html/2607.26398#bib.bib44); Boffi and Vanden\-Eijnden,[2023](https://arxiv.org/html/2607.26398#bib.bib14)\)\.

Consistency Models \(CMs\)\(Songet al\.,[2023](https://arxiv.org/html/2607.26398#bib.bib9); Song and Dhariwal,[2023](https://arxiv.org/html/2607.26398#bib.bib8); Lu and Song,[2024](https://arxiv.org/html/2607.26398#bib.bib11)\)learn a one\-step map from noise to data, either by distilling a pretrained model or by learning from scratch\. Distillation requires sampling trajectories from the teacher model\. To allow for more steps after either training approach, CMs iteratively re\-noise the one\-step solution back to successively smaller time under the interpolant and then denoise, but this can take the solver off the probability flow\.

Consistency trajectory models \(CTMs\)\(Kimet al\.,[2023](https://arxiv.org/html/2607.26398#bib.bib5)\)extend CMs to learn two\-time maps using a combination of consistency and adversarial objectives, which requires training an additional discriminator model\(Goodfellowet al\.,[2014](https://arxiv.org/html/2607.26398#bib.bib6)\)\. CTM and SGFlow both target the same mathematical object, the probability flow ODE flow map \(i\.e\., the integral of the ODE\), but they learn this map through different means\. CTM learns the map by distilling a teacher solver, and the losses for teacher and student involve several nested model evaluations \(with data atx0x\_\{0\}, for0≤s≤u≤t≤10\\leq s\\leq u\\leq t\\leq 1, the teacher integrates fromtttouu, then jumps fromuutoss, then fromssto0; and the student jumps fromtttossand then to0\)\. The objective depends on a chosen feature\-space distance and, in practice, includes DSM and GAN terms that further influence the optimum\. Consequently, the CTM loss is sensitive to the quality of the ODE discretization used by the teacher \(in practice CTM finds the need to use a 2nd order solver during training\) and necessitates the presence of the GAN\.

Inductive Moment Matching \(IMM\)\(Zhouet al\.,[2025](https://arxiv.org/html/2607.26398#bib.bib4)\)learns a few\-step model via an implicit generative model trained with MMD\(Smolaet al\.,[2006](https://arxiv.org/html/2607.26398#bib.bib15); Grettonet al\.,[2012](https://arxiv.org/html/2607.26398#bib.bib3)\), where the MMD is estimated with bias within subsets of data\. In practice, the authors must use time\-weighting schedules and specific curriculum/inductive procedure to stabilize optimization\. While IMM produces high\-quality image samples, it solves the problem of marginally sampling the data distribution rather than sampling along a probability flow, where the latter is the task studied in this work\.

Meanflow\(Genget al\.,[2025](https://arxiv.org/html/2607.26398#bib.bib36)\)derives a JVP\-based objective for flow maps from the same PDE in[eq\.˜5](https://arxiv.org/html/2607.26398#S3.E5):

ℒt,umeanflow:=𝔼​\[‖f~θ​\(t,u,Xt\)−X˙t−\(u−t\)​SG​\(∂xf~θ⋅X˙t\+∂tf~θ\)‖2\]\.\\displaystyle\\mathcal\{L\}\_\{t,u\}^\{\\text\{meanflow\}\}:=\\mathbb\{E\}\[\\\|\\tilde\{f\}\_\{\\theta\}\(t,u,X\_\{t\}\)\-\\dot\{X\}\_\{t\}\-\(u\-t\)\\text\{SG\}\(\\partial\_\{x\}\\tilde\{f\}\_\{\\theta\}\\cdot\\dot\{X\}\_\{t\}\+\\partial\_\{t\}\\tilde\{f\}\_\{\\theta\}\)\\\|^\{2\}\]\.\(16\)Here Meanflow is written in forward time withttderivatives rather than reverse time withuuderivatives \([section˜B\.3](https://arxiv.org/html/2607.26398#A2.SS3)\)\. Applying the stopgradSGto all model derivatives improves efficiency, but there are no differentiated loss terms that encourage the model derivatives∂tf~θ\\partial\_\{t\}\\tilde\{f\}\_\{\\theta\}and∂xf~θ\\partial\_\{x\}\\tilde\{f\}\_\{\\theta\}to move toward the true flow map derivatives\. This contrasts with SGFlow whereSG​\[f~θ​\(t,t,x\)\]\\text\{SG\}\[\\tilde\{f\}\_\{\\theta\}\(t,t,x\)\]is used in place of𝔼​\[X˙t\|Xt\]\\mathbb\{E\}\[\\dot\{X\}\_\{t\}\|X\_\{t\}\], but where another term in the loss trains these two quantities to match\. Finally, between equations \(10, 11\) inGenget al\.\([2025](https://arxiv.org/html/2607.26398#bib.bib36)\),vvis replaced withX˙t\\dot\{X\}\_\{t\}where it appears quadratically, thereby pulling an expectation through a square and missing a resulting trace covariance term\(Boffiet al\.,[2025](https://arxiv.org/html/2607.26398#bib.bib38)\); interestingly, this causes Meanflow to have wrong stationary points when the stopgradis notused\. Meanflowwithstopgrad does have correct stationary points\(Sabouret al\.,[2025](https://arxiv.org/html/2607.26398#bib.bib37)\)\.

Flow Map Matching\(Boffiet al\.,[2024](https://arxiv.org/html/2607.26398#bib.bib44)\)learns a two\-time flow map, enabling mapping along the probability flow in either direction without adversarial training\. The Lagrangian loss requires only time derivatives, but relies on an additional invertibility loss, which encourages invertibility via swapping the time arguments:

fθ​\(t,u,fθ​\(u,t,x\)\)≈x\.\\displaystyle f\_\{\\theta\}\(t,u,f\_\{\\theta\}\(u,t,x\)\)\\approx x\.Here the model is trained with the second time argument either larger or smaller than the first\. This is straightforward to compute, but gradient steps require evaluating the model and its inverse at each training step\.Boffiet al\.\([2025](https://arxiv.org/html/2607.26398#bib.bib38)\)introduce stopgrad placements for the LSD, ESD, and PSD loss functionals fromBoffiet al\.\([2024](https://arxiv.org/html/2607.26398#bib.bib44)\)that substantially improve image generation performance\. Given this empirical success, understanding why these stopgrads help is important: they likely have good theoretical properties, but this remains to be shown\.

Distillation methods\. A complementary line of work approaches flow map learning by*distilling*the outputs of fully pretrained flow matching velocity models into few\-step solvers\. Specifically, the unknownvvin the flow map identities is taken to be a pretrained network\. By contrast, we emphasize training from scratch, avoiding dependence on a teacher model and ensuring that all components of the flow map are learned end\-to\-end\. That said, distillation can be attractive in practice when a pretrained model is already trusted \(whenvθv\_\{\\theta\}corresponds to the endpoint distributions and the chosenαt\\alpha\_\{t\}andσt\\sigma\_\{t\}, or when the objective is weaker—for example, to marginally sample from the approximated data distribution without explicitly solving the probability flow ODE\)\.

## 7Discussion and Limitations

#### Identifying functions through PDEs

Consider a PDE solved by a sought\-after mappingff, featuring a combination of terms such as the time and space derivatives∂tf,∂xf\\partial\_\{t\}f,\\partial\_\{x\}f\. The PDE being solved means thatffsets the residual to0\. Such anffcan be found by minimizing a squared error loss built from the residual\. Which terms should be parameterized by the model and which should be approximated as part of a ground\-truth loss target? If the equations can be rewritten in several ways, which yield easier or more challenging objectives? Answering this is applicable to improving training objectives for generative models as well as solving more general PDE\-related tasks with neural networks and optimization\.

#### Invertibility

The loss targets an invertible function at optimum\. To simplify training, we explicitly give up knowing the inverse, meaning that we only learn maps in one direction\. Luckily, this is the usual scenario for generative modeling\. For likelihoods, one can still substitute−∂tfθ\-\\partial\_\{t\}f\_\{\\theta\}forvθv\_\{\\theta\}in the probability flow ODE\(Songet al\.,[2021](https://arxiv.org/html/2607.26398#bib.bib45); Boffi and Vanden\-Eijnden,[2023](https://arxiv.org/html/2607.26398#bib.bib14)\)\. Thus this method can be seen from the perspective of training a normalizing flow\(Tabak and Vanden\-Eijnden,[2010](https://arxiv.org/html/2607.26398#bib.bib34); Tabak and Turner,[2013](https://arxiv.org/html/2607.26398#bib.bib42); Rezende and Mohamed,[2015](https://arxiv.org/html/2607.26398#bib.bib43); Papamakarioset al\.,[2021](https://arxiv.org/html/2607.26398#bib.bib18)\)without requiring the invertible architecture or inverse\-dependent loss\.

#### Architectures\.

SGFlow, Flow Map Matching, Simplified Consistency Models, and Meanflow all specify models whosetime\-derivativesequal the target of diffusion model training, but directly adapt architectures meant for diffusion models themselves\. Example architectures used in these works are the UNet fromDhariwal and Nichol \([2021](https://arxiv.org/html/2607.26398#bib.bib26)\), the diffusion transformer fromPeebles and Xie \([2023](https://arxiv.org/html/2607.26398#bib.bib27)\); Maet al\.\([2024](https://arxiv.org/html/2607.26398#bib.bib24)\), and the EDM architecture fromKarraset al\.\([2022](https://arxiv.org/html/2607.26398#bib.bib17);[2024](https://arxiv.org/html/2607.26398#bib.bib28)\)\. These architectures may thus be suboptimal for the problem at hand, precisely because the target of interest is defined as a function often computed in many diffusion model forward passes \(an integral\)\. In this work, compute limitations did not allow for the thorough exploration of architectures, but the authors believe that rethinking architectures is a convincing direction to improve the quality and training\-efficiency of learned flow maps\.

## References

- J\. Abramson, J\. Adler, J\. Dunger, R\. Evans, T\. Green, A\. Pritzel, O\. Ronneberger, L\. Willmore, A\. J\. Ballard, J\. Bambrick,et al\.\(2024\)Accurate structure prediction of biomolecular interactions with alphafold 3\.Nature630\(8016\),pp\. 493–500\.Cited by:[§1](https://arxiv.org/html/2607.26398#S1.p1.1)\.
- M\. S\. Albergo, N\. M\. Boffi, and E\. Vanden\-Eijnden \(2023\)Stochastic interpolants: a unifying framework for flows and diffusions\.arXiv preprint arXiv:2303\.08797\.Cited by:[§1](https://arxiv.org/html/2607.26398#S1.p1.1),[§2](https://arxiv.org/html/2607.26398#S2.p1.1)\.
- M\. S\. Albergo and E\. Vanden\-Eijnden \(2022\)Building normalizing flows with stochastic interpolants\.arXiv preprint arXiv:2209\.15571\.Cited by:[§1](https://arxiv.org/html/2607.26398#S1.p1.1)\.
- M\. S\. Albergo and E\. Vanden\-Eijnden \(2024\)Learning to sample better\.Journal of Statistical Mechanics: Theory and Experiment2024\(10\),pp\. 104014\.Cited by:[§2](https://arxiv.org/html/2607.26398#S2.p1.11)\.
- D\. Balduzzi, S\. Racaniere, J\. Martens, J\. Foerster, K\. Tuyls, and T\. Graepel \(2018\)The mechanics of n\-player differentiable games\.InInternational Conference on Machine Learning,pp\. 354–363\.Cited by:[§3](https://arxiv.org/html/2607.26398#S3.SS0.SSS0.Px3.p1.1)\.
- G\. Bartosh, D\. P\. Vetrov, and C\. Andersson Naesseth \(2024\)Neural flow diffusion models: learnable forward process for improved diffusion modelling\.Advances in Neural Information Processing Systems37,pp\. 73952–73985\.Cited by:[§1](https://arxiv.org/html/2607.26398#S1.p1.1)\.
- N\. M\. Boffi, M\. S\. Albergo, and E\. Vanden\-Eijnden \(2024\)Flow map matching\.arXiv preprint arXiv:2406\.07507\.Cited by:[Table 1](https://arxiv.org/html/2607.26398#S1.T1.pic1.1.1.1.1.1.1.5.4.1),[§1](https://arxiv.org/html/2607.26398#S1.p2.1),[§1](https://arxiv.org/html/2607.26398#S1.p4.1),[§2](https://arxiv.org/html/2607.26398#S2.SS0.SSS0.Px1.p1.20),[§5\.2](https://arxiv.org/html/2607.26398#S5.SS2.SSS0.Px3.p1.1),[§6](https://arxiv.org/html/2607.26398#S6.p1.1),[§6](https://arxiv.org/html/2607.26398#S6.p6.1),[§6](https://arxiv.org/html/2607.26398#S6.p6.2)\.
- N\. M\. Boffi, M\. S\. Albergo, and E\. Vanden\-Eijnden \(2025\)How to build a consistency model: learning flow maps via self\-distillation\.arXiv preprint arXiv:2505\.18825\.Cited by:[Table 1](https://arxiv.org/html/2607.26398#S1.T1.pic1.1.1.1.1.1.1.6.5.1),[Table 1](https://arxiv.org/html/2607.26398#S1.T1.pic1.1.1.1.1.1.1.7.6.1),[Table 1](https://arxiv.org/html/2607.26398#S1.T1.pic1.1.1.1.1.1.1.8.7.1),[§1](https://arxiv.org/html/2607.26398#S1.p2.1),[§1](https://arxiv.org/html/2607.26398#S1.p4.1),[§5\.2](https://arxiv.org/html/2607.26398#S5.SS2.SSS0.Px3.p1.1),[§6](https://arxiv.org/html/2607.26398#S6.p5.9),[§6](https://arxiv.org/html/2607.26398#S6.p6.2)\.
- N\. M\. Boffi and E\. Vanden\-Eijnden \(2023\)Probability flow solution of the fokker–planck equation\.Machine Learning: Science and Technology4\(3\),pp\. 035012\.Cited by:[§6](https://arxiv.org/html/2607.26398#S6.p1.1),[§7](https://arxiv.org/html/2607.26398#S7.SS0.SSS0.Px2.p1.2)\.
- P\. Dhariwal and A\. Nichol \(2021\)Diffusion models beat gans on image synthesis\.Advances in neural information processing systems34,pp\. 8780–8794\.Cited by:[§5\.2](https://arxiv.org/html/2607.26398#S5.SS2.SSS0.Px1.p1.2),[§7](https://arxiv.org/html/2607.26398#S7.SS0.SSS0.Px3.p1.1)\.
- P\. Esser, S\. Kulal, A\. Blattmann, R\. Entezari, J\. Müller, H\. Saini, Y\. Levi, D\. Lorenz, A\. Sauer, F\. Boesel,et al\.\(2024\)Scaling rectified flow transformers for high\-resolution image synthesis\.arXiv preprint arXiv:2403\.03206\.Cited by:[§1](https://arxiv.org/html/2607.26398#S1.p1.1)\.
- Z\. Geng, M\. Deng, X\. Bai, J\. Z\. Kolter, and K\. He \(2025\)Mean flows for one\-step generative modeling\.arXiv preprint arXiv:2505\.13447\.Cited by:[§B\.3](https://arxiv.org/html/2607.26398#A2.SS3.p1.7),[Table 1](https://arxiv.org/html/2607.26398#S1.T1.pic1.1.1.1.1.1.1.9.8.1),[§1](https://arxiv.org/html/2607.26398#S1.p4.1),[§2](https://arxiv.org/html/2607.26398#S2.SS0.SSS0.Px1.p1.20),[§4](https://arxiv.org/html/2607.26398#S4.p1.3),[§5\.2](https://arxiv.org/html/2607.26398#S5.SS2.SSS0.Px3.p1.1),[§6](https://arxiv.org/html/2607.26398#S6.p5.10),[§6](https://arxiv.org/html/2607.26398#S6.p5.9)\.
- I\. J\. Goodfellow, J\. Pouget\-Abadie, M\. Mirza, B\. Xu, D\. Warde\-Farley, S\. Ozair, A\. Courville, and Y\. Bengio \(2014\)Generative adversarial nets\.Advances in neural information processing systems27\.Cited by:[§6](https://arxiv.org/html/2607.26398#S6.p3.11)\.
- A\. Gretton, K\. M\. Borgwardt, M\. J\. Rasch, B\. Schölkopf, and A\. Smola \(2012\)A kernel two\-sample test\.The Journal of Machine Learning Research13\(1\),pp\. 723–773\.Cited by:[§6](https://arxiv.org/html/2607.26398#S6.p4.1)\.
- M\. Heusel, H\. Ramsauer, T\. Unterthiner, B\. Nessler, and S\. Hochreiter \(2017\)Gans trained by a two time\-scale update rule converge to a local nash equilibrium\.Advances in neural information processing systems30\.Cited by:[§5\.2](https://arxiv.org/html/2607.26398#S5.SS2.SSS0.Px4.p1.1)\.
- J\. Ho, A\. Jain, and P\. Abbeel \(2020\)Denoising diffusion probabilistic models\.Advances in neural information processing systems33,pp\. 6840–6851\.Cited by:[§1](https://arxiv.org/html/2607.26398#S1.p1.1)\.
- T\. Karras, M\. Aittala, T\. Aila, and S\. Laine \(2022\)Elucidating the design space of diffusion\-based generative models\.Advances in neural information processing systems35,pp\. 26565–26577\.Cited by:[§7](https://arxiv.org/html/2607.26398#S7.SS0.SSS0.Px3.p1.1)\.
- T\. Karras, M\. Aittala, J\. Lehtinen, J\. Hellsten, T\. Aila, and S\. Laine \(2024\)Analyzing and improving the training dynamics of diffusion models\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 24174–24184\.Cited by:[§7](https://arxiv.org/html/2607.26398#S7.SS0.SSS0.Px3.p1.1)\.
- D\. Kim, C\. Lai, W\. Liao, N\. Murata, Y\. Takida, T\. Uesaka, Y\. He, Y\. Mitsufuji, and S\. Ermon \(2023\)Consistency trajectory models: learning probability flow ode trajectory of diffusion\.arXiv preprint arXiv:2310\.02279\.Cited by:[Table 1](https://arxiv.org/html/2607.26398#S1.T1.pic1.1.1.1.1.1.1.4.3.1),[§1](https://arxiv.org/html/2607.26398#S1.p2.1),[§2](https://arxiv.org/html/2607.26398#S2.SS0.SSS0.Px1.p1.20),[§6](https://arxiv.org/html/2607.26398#S6.p1.1),[§6](https://arxiv.org/html/2607.26398#S6.p3.11)\.
- D\. P\. Kingma, T\. Salimans, B\. Poole, and J\. Ho \(2021\)Variational diffusion models\.arXiv preprint arXiv:2107\.00630\.Cited by:[§1](https://arxiv.org/html/2607.26398#S1.p1.1)\.
- Y\. Lipman, R\. T\. Chen, H\. Ben\-Hamu, M\. Nickel, and M\. Le \(2022\)Flow matching for generative modeling\.arXiv preprint arXiv:2210\.02747\.Cited by:[§1](https://arxiv.org/html/2607.26398#S1.p1.1),[§2](https://arxiv.org/html/2607.26398#S2.p1.1),[§5\.2](https://arxiv.org/html/2607.26398#S5.SS2.SSS0.Px3.p1.1)\.
- X\. Liu, C\. Gong, and Q\. Liu \(2022\)Flow straight and fast: learning to generate and transfer data with rectified flow\.arXiv preprint arXiv:2209\.03003\.Cited by:[§1](https://arxiv.org/html/2607.26398#S1.p1.1)\.
- X\. Liu, X\. Zhang, J\. Ma, J\. Peng,et al\.\(2023\)Instaflow: one step is enough for high\-quality diffusion\-based text\-to\-image generation\.InThe Twelfth International Conference on Learning Representations,Cited by:[§6](https://arxiv.org/html/2607.26398#S6.p1.1)\.
- C\. Lu and Y\. Song \(2024\)Simplifying, stabilizing and scaling continuous\-time consistency models\.arXiv preprint arXiv:2410\.11081\.Cited by:[§1](https://arxiv.org/html/2607.26398#S1.p2.1),[§2](https://arxiv.org/html/2607.26398#S2.SS0.SSS0.Px1.p1.20),[§6](https://arxiv.org/html/2607.26398#S6.p2.1)\.
- N\. Ma, M\. Goldstein, M\. S\. Albergo, N\. M\. Boffi, E\. Vanden\-Eijnden, and S\. Xie \(2024\)Sit: exploring flow and diffusion\-based generative models with scalable interpolant transformers\.InEuropean Conference on Computer Vision,pp\. 23–40\.Cited by:[§7](https://arxiv.org/html/2607.26398#S7.SS0.SSS0.Px3.p1.1)\.
- K\. Pandey and S\. Mandt \(2023\)A complete recipe for diffusion generative models\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 4261–4272\.Cited by:[§1](https://arxiv.org/html/2607.26398#S1.p1.1)\.
- G\. Papamakarios, E\. Nalisnick, D\. J\. Rezende, S\. Mohamed, and B\. Lakshminarayanan \(2021\)Normalizing flows for probabilistic modeling and inference\.Journal of Machine Learning Research22\(57\),pp\. 1–64\.Cited by:[§7](https://arxiv.org/html/2607.26398#S7.SS0.SSS0.Px2.p1.2)\.
- W\. Peebles and S\. Xie \(2023\)Scalable diffusion models with transformers\.InProceedings of the IEEE/CVF international conference on computer vision,pp\. 4195–4205\.Cited by:[§1](https://arxiv.org/html/2607.26398#S1.p1.1),[§7](https://arxiv.org/html/2607.26398#S7.SS0.SSS0.Px3.p1.1)\.
- D\. Rezende and S\. Mohamed \(2015\)Variational inference with normalizing flows\.InInternational conference on machine learning,pp\. 1530–1538\.Cited by:[§7](https://arxiv.org/html/2607.26398#S7.SS0.SSS0.Px2.p1.2)\.
- A\. Sabour, S\. Fidler, and K\. Kreis \(2025\)Align your flow: scaling continuous\-time flow map distillation\.arXiv preprint arXiv:2506\.14603\.Cited by:[§2](https://arxiv.org/html/2607.26398#S2.SS0.SSS0.Px1.p1.20),[§4](https://arxiv.org/html/2607.26398#S4.p1.9),[§6](https://arxiv.org/html/2607.26398#S6.p5.9)\.
- T\. Salimans and J\. Ho \(2022\)Progressive distillation for fast sampling of diffusion models\.arXiv preprint arXiv:2202\.00512\.Cited by:[§6](https://arxiv.org/html/2607.26398#S6.p1.1)\.
- R\. Singhal, M\. Goldstein, and R\. Ranganath \(2023\)Where to diffuse, how to diffuse, and how to get back: automated learning for multivariate diffusions\.arXiv preprint arXiv:2302\.07261\.Cited by:[§1](https://arxiv.org/html/2607.26398#S1.p1.1)\.
- R\. Singhal, M\. Goldstein, and R\. Ranganath \(2024\)What’s the score? automated denoising score matching for nonlinear diffusions\.InInternational Conference on Machine Learning,pp\. 45734–45758\.Cited by:[§1](https://arxiv.org/html/2607.26398#S1.p1.1)\.
- A\. J\. Smola, A\. Gretton, and K\. Borgwardt \(2006\)Maximum mean discrepancy\.In13th international conference, ICONIP,Vol\.6\.Cited by:[§6](https://arxiv.org/html/2607.26398#S6.p4.1)\.
- J\. Sohl\-Dickstein, E\. Weiss, N\. Maheswaranathan, and S\. Ganguli \(2015\)Deep unsupervised learning using nonequilibrium thermodynamics\.InInternational conference on machine learning,pp\. 2256–2265\.Cited by:[§1](https://arxiv.org/html/2607.26398#S1.p1.1)\.
- Y\. Song, P\. Dhariwal, M\. Chen, and I\. Sutskever \(2023\)Consistency models\.arXiv preprint arXiv:2303\.01469\.Cited by:[Table 1](https://arxiv.org/html/2607.26398#S1.T1.pic1.1.1.1.1.1.1.2.1.1),[Table 1](https://arxiv.org/html/2607.26398#S1.T1.pic1.1.1.1.1.1.1.3.2.1),[§1](https://arxiv.org/html/2607.26398#S1.p2.1),[§2](https://arxiv.org/html/2607.26398#S2.SS0.SSS0.Px1.p1.13),[§2](https://arxiv.org/html/2607.26398#S2.SS0.SSS0.Px1.p1.20),[§6](https://arxiv.org/html/2607.26398#S6.p1.1),[§6](https://arxiv.org/html/2607.26398#S6.p2.1)\.
- Y\. Song and P\. Dhariwal \(2023\)Improved techniques for training consistency models\.arXiv preprint arXiv:2310\.14189\.Cited by:[§1](https://arxiv.org/html/2607.26398#S1.p2.1),[§2](https://arxiv.org/html/2607.26398#S2.SS0.SSS0.Px1.p1.13),[§6](https://arxiv.org/html/2607.26398#S6.p2.1)\.
- Y\. Song, C\. Durkan, I\. Murray, and S\. Ermon \(2021\)Maximum likelihood training of score\-based diffusion models\.Advances in neural information processing systems34,pp\. 1415–1428\.Cited by:[§7](https://arxiv.org/html/2607.26398#S7.SS0.SSS0.Px2.p1.2)\.
- Y\. Song, J\. Sohl\-Dickstein, D\. P\. Kingma, A\. Kumar, S\. Ermon, and B\. Poole \(2020\)Score\-based generative modeling through stochastic differential equations\.arXiv preprint arXiv:2011\.13456\.Cited by:[Table 1](https://arxiv.org/html/2607.26398#S1.T1),[§1](https://arxiv.org/html/2607.26398#S1.p1.1)\.
- E\. G\. Tabak and C\. V\. Turner \(2013\)A family of nonparametric density estimation algorithms\.Communications on Pure and Applied Mathematics66\(2\),pp\. 145–164\.Cited by:[§7](https://arxiv.org/html/2607.26398#S7.SS0.SSS0.Px2.p1.2)\.
- E\. G\. Tabak and E\. Vanden\-Eijnden \(2010\)Density estimation by dual ascent of the log\-likelihood\.Communications in Mathematical Sciences8\(1\),pp\. 217–233\.Cited by:[§7](https://arxiv.org/html/2607.26398#S7.SS0.SSS0.Px2.p1.2)\.
- L\. Zhou, S\. Ermon, and J\. Song \(2025\)Inductive moment matching\.arXiv preprint arXiv:2503\.07565\.Cited by:[§2](https://arxiv.org/html/2607.26398#S2.SS0.SSS0.Px1.p1.20),[§6](https://arxiv.org/html/2607.26398#S6.p1.1),[§6](https://arxiv.org/html/2607.26398#S6.p4.1)\.

## Acknowledgments

This work was partly supported by the NIH/NHLBI Award R01HL148248, NSF Award 1922658 NRT\-HDR: FUTURE Foundations, Translation, and Responsibility for Data Science, NSF CAREER Award 2145542, NSF Award 2404476, ONR N00014\-23\-1\-2634, Optum, and Apple\. This work was also supported by IITP with a grant funded by the MSIT of the Republic of Korea in connection with the Global AI Frontier Lab International Collaborative Research\. Mark Goldstein would like to thank the Simons Foundation and Flatiron Institute for funding and compute resources that supported this project\. Anshuk Uppal would like to thank the Centre for Basic Machine Learning Research in Life Sciences, and Technical University Denmark\. The authors thank Eric Vanden\-Eijnden and Amirmojtaba Sabour for fruitful discussion on related work and for providing feedback on the theoretical results\.

## Appendix AProofs for Stationary Points

### A\.1First Variation Definitions

We consider scalar\-valued loss functionsℒ:ℱ→ℝ\\mathcal\{L\}:\\mathcal\{F\}\\to\\mathbb\{R\}that map a functionf∈ℱf\\in\\mathcal\{F\}to a real value\.

Define the tangent space𝒯f​\(ℱ\)\\mathcal\{T\}\_\{f\}\(\\mathcal\{F\}\)atff\. This space contains functionsh∈𝒯f​\(ℱ\)h\\in\\mathcal\{T\}\_\{f\}\(\\mathcal\{F\}\)such that there exists a curve indexed by scalarϵ\\epsilonsuch that for eachϵ\\epsilon,fϵ∈ℱf\_\{\\epsilon\}\\in\\mathcal\{F\}, and we have thatf0=ff\_\{0\}=fand\(dd​ϵ​fϵ\)\|ϵ=0=h\(\\frac\{d\}\{d\\epsilon\}f\_\{\\epsilon\}\)\|\_\{\\epsilon=0\}=h\.

The first variationδ​ℒ\\delta\\mathcal\{L\}of such a functionalℒ\\mathcal\{L\}evaluated atf∈ℱf\\in\\mathcal\{F\}in directionh∈𝒯f​\(ℱ\)h\\in\\mathcal\{T\}\_\{f\}\(\\mathcal\{F\}\)is defined as:

δ​ℒ​\[f;h\]:=\(dd​ϵ​ℒ​\[f\+ϵ​h\]\)ϵ=0\\displaystyle\\delta\\mathcal\{L\}\[f;h\]:=\\Big\(\\frac\{d\}\{d\\epsilon\}\\mathcal\{L\}\[f\+\\epsilon h\]\\Big\)\_\{\\epsilon=0\}\(17\)We then have thatf∗f^\{\*\}is a stationary point w\.r\.t\.ℱ\\mathcal\{F\}ifδ​ℒ​\[f∗;h\]=0\\delta\\mathcal\{L\}\[f^\{\*\};h\]=0for allh∈𝒯f∗​\(ℱ\)h\\in\\mathcal\{T\}\_\{f^\{\*\}\}\(\\mathcal\{F\}\)\.

### A\.2Stopgrad for Functionals

We define the stopgrad symbolsgfor a functional as follows\. Let𝒪\\mathcal\{O\}be a functional that maps two functionsf,gf,gto a real value\. Letℒ​\[f\]\\mathcal\{L\}\[f\]be a functional that is written in terms of𝒪\\mathcal\{O\}with symbolsgasℒ​\[f\]:=𝒪​\[f,sg​\[f\]\]\\mathcal\{L\}\[f\]:=\\mathcal\{O\}\[f,\\text\{sg\}\[f\]\], then we evaluate the following two quantities as follows

ℒ​\[f\]\\displaystyle\\mathcal\{L\}\[f\]=𝒪​\[f,f\]\\displaystyle=\\mathcal\{O\}\[f,f\]\(18\)δ​ℒ​\[f;h\]\\displaystyle\\delta\\mathcal\{L\}\[f;h\]=δ​𝒪​\[f,f;h,0\]\\displaystyle=\\delta\\mathcal\{O\}\[f,f;h,0\]\(19\)That is, the functional evaluates as usual but in a first variation, we do not perturb terms insg\. This corresponds to the stopgrad or detach\(\) operation used in machine learning code with autodifferentiation\.

### A\.3First Variation of Original Loss Functional

Our functionalℒ​\[f~\]\\mathcal\{L\}\[\\tilde\{f\}\]acts on functionsf~\\tilde\{f\}\. According to the definitions in[Section˜A\.1](https://arxiv.org/html/2607.26398#A1.SS1), we need to computeδ​ℒ​\[f~;h\]=\(dd​ϵ​ℒ​\[f~\+ϵ​h\]\)ϵ=0\\delta\\mathcal\{L\}\[\\tilde\{f\};h\]=\(\\frac\{d\}\{d\\epsilon\}\\mathcal\{L\}\[\\tilde\{f\}\+\\epsilon h\]\)\_\{\\epsilon=0\}\. The functional is:

T1​\[f~\]:\\displaystyle T\_\{1\}\[\\tilde\{f\}\]:=‖\(∂tf\)\(t,u,xt\)\+\(∂xf\)\(t,u,xt\)​x˙t‖f=x\+\(u−t\)​f~2\\displaystyle=\\\|\(\\partial\_\{t\}f\)\_\{\(t,u,x\_\{t\}\)\}\+\(\\partial\_\{x\}f\)\_\{\(t,u,x\_\{t\}\)\}\\dot\{x\}\_\{t\}\\\|^\{2\}\_\{f=x\+\(u\-t\)\\tilde\{f\}\}T2​\[f~\]:\\displaystyle T\_\{2\}\[\\tilde\{f\}\]:=∥\(∂xf\)\(t,u,xt\)\(x˙t−𝔼\[x˙t\|xt\]\)∥f=x\+\(u−t\)​f~2\\displaystyle=\\\|\(\\partial\_\{x\}f\)\_\{\(t,u,x\_\{t\}\)\}\(\\dot\{x\}\_\{t\}\-\\mathbb\{E\}\[\\dot\{x\}\_\{t\}\|x\_\{t\}\]\)\\\|^\{2\}\_\{f=x\+\(u\-t\)\\tilde\{f\}\}A​\[f~\]\\displaystyle A\[\\tilde\{f\}\]=T1​\[f~\]−T2​\[f~\]\\displaystyle=T\_\{1\}\[\\tilde\{f\}\]\-T\_\{2\}\[\\tilde\{f\}\]ℒ​\[f~\]\\displaystyle\\mathcal\{L\}\[\\tilde\{f\}\]=𝔼q​\(t,u\),q​\(x0\),q​\(x1\)​\[A​\[f~\]\]\\displaystyle=\\mathbb\{E\}\_\{q\(t,u\),q\(x\_\{0\}\),q\(x\_\{1\}\)\}\\Big\[A\[\\tilde\{f\}\]\\Big\]Lets define the pathfϵf\_\{\\epsilon\}by replacingf~\\tilde\{f\}withf~ϵ:=f~\+ϵ​h\\tilde\{f\}\_\{\\epsilon\}:=\\tilde\{f\}\+\\epsilon h\. Then:

fϵ\\displaystyle f\_\{\\epsilon\}:=x\+\(u−t\)​f~ϵ=x\+\(u−t\)​\(f~\+ϵ​h\)=x\+\(u−t\)​f~\+ϵ​\(u−t\)​h\\displaystyle:=x\+\(u\-t\)\\tilde\{f\}\_\{\\epsilon\}=x\+\(u\-t\)\(\\tilde\{f\}\+\\epsilon h\)=x\+\(u\-t\)\\tilde\{f\}\+\\epsilon\(u\-t\)h\(20\)
Then

dd​ϵ​ℒ​\[f~\+ϵ​h\]=𝔼​\[dd​ϵ​A​\[f~ϵ\]\]=𝔼​\[dd​ϵ​T1​\[f~\+ϵ​h\]−dd​ϵ​T2​\[f~\+ϵ​h\]\]\\displaystyle\\frac\{d\}\{d\\epsilon\}\\mathcal\{L\}\[\\tilde\{f\}\+\\epsilon h\]=\\mathbb\{E\}\\Big\[\\frac\{d\}\{d\\epsilon\}A\[\\tilde\{f\}\_\{\\epsilon\}\]\\Big\]=\\mathbb\{E\}\\Big\[\\frac\{d\}\{d\\epsilon\}T\_\{1\}\[\\tilde\{f\}\+\\epsilon h\]\-\\frac\{d\}\{d\\epsilon\}T\_\{2\}\[\\tilde\{f\}\+\\epsilon h\]\\Big\]\(21\)We first compute this derivative and then evaluate it atϵ=0\\epsilon=0\.

So

∂tfϵ\\displaystyle\\partial\_\{t\}f\_\{\\epsilon\}=∂t\[x\+\(u−t\)​f~\+ϵ​\(u−t\)​h\]\\displaystyle=\\partial\_\{t\}\\Big\[x\+\(u\-t\)\\tilde\{f\}\+\\epsilon\(u\-t\)h\\Big\]\(22\)=∂tx\+∂t\[\(u−t\)​f~\]\+ϵ​∂t\[\(u−t\)​h\]\\displaystyle=\\partial\_\{t\}x\+\\partial\_\{t\}\\Big\[\(u\-t\)\\tilde\{f\}\\Big\]\+\\epsilon\\partial\_\{t\}\\Big\[\(u\-t\)h\\Big\]\(23\)=\(u−t\)​∂tf~−f~\+ϵ​\(\(u−t\)​∂th−h\)\\displaystyle=\(u\-t\)\\partial\_\{t\}\\tilde\{f\}\-\\tilde\{f\}\+\\epsilon\\Big\(\(u\-t\)\\partial\_\{t\}h\-h\\Big\)\(24\)=\(u−t\)​\(∂tf~\+ϵ​∂th\)−\(f~\+ϵ​h\)\\displaystyle=\(u\-t\)\(\\partial\_\{t\}\\tilde\{f\}\+\\epsilon\\partial\_\{t\}h\)\-\(\\tilde\{f\}\+\\epsilon h\)\(25\)and

∂xfϵ\\displaystyle\\partial\_\{x\}f\_\{\\epsilon\}=∂x\[x\+\(u−t\)​f~\+ϵ​\(u−t\)​h\]=I\+\(u−t\)​∂x\(f~\+ϵ​h\)\\displaystyle=\\partial\_\{x\}\\Big\[x\+\(u\-t\)\\tilde\{f\}\+\\epsilon\(u\-t\)h\\Big\]=I\+\(u\-t\)\\partial\_\{x\}\(\\tilde\{f\}\+\\epsilon h\)\(26\)and

dd​ϵ​∂tfϵ\\displaystyle\\frac\{d\}\{d\\epsilon\}\\partial\_\{t\}f\_\{\\epsilon\}=dd​ϵ​\[\(u−t\)​\(∂tf~\+ϵ​∂th\)−\(f~\+ϵ​h\)\]\\displaystyle=\\frac\{d\}\{d\\epsilon\}\\Big\[\(u\-t\)\(\\partial\_\{t\}\\tilde\{f\}\+\\epsilon\\partial\_\{t\}h\)\-\(\\tilde\{f\}\+\\epsilon h\)\\Big\]\(27\)=dd​ϵ​\[\(u−t\)​∂tf~\+ϵ​\(u−t\)​∂th−f~−ϵ​h\]\\displaystyle=\\frac\{d\}\{d\\epsilon\}\\Big\[\(u\-t\)\\partial\_\{t\}\\tilde\{f\}\+\\epsilon\(u\-t\)\\partial\_\{t\}h\-\\tilde\{f\}\-\\epsilon h\\Big\]\(28\)=dd​ϵ​\[ϵ​\(u−t\)​∂th−ϵ​h\]\\displaystyle=\\frac\{d\}\{d\\epsilon\}\\Big\[\\epsilon\(u\-t\)\\partial\_\{t\}h\-\\epsilon h\\Big\]\(29\)=\(u−t\)​∂th−h\\displaystyle=\(u\-t\)\\partial\_\{t\}h\-h\(30\)and

dd​ϵ​∂xfϵ\\displaystyle\\frac\{d\}\{d\\epsilon\}\\partial\_\{x\}f\_\{\\epsilon\}=dd​ϵ​\[I\+\(u−t\)​∂x\(f~\+ϵ​h\)\]\\displaystyle=\\frac\{d\}\{d\\epsilon\}\\Big\[I\+\(u\-t\)\\partial\_\{x\}\(\\tilde\{f\}\+\\epsilon h\)\\Big\]\(31\)=dd​ϵ​I\+dd​ϵ​\(u−t\)​∂xf~\+dd​ϵ​\(u−t\)​∂xϵ​h\\displaystyle=\\frac\{d\}\{d\\epsilon\}I\+\\frac\{d\}\{d\\epsilon\}\(u\-t\)\\partial\_\{x\}\\tilde\{f\}\+\\frac\{d\}\{d\\epsilon\}\(u\-t\)\\partial\_\{x\}\\epsilon h\(32\)=\(u−t\)​∂xh\\displaystyle=\(u\-t\)\\partial\_\{x\}h\(33\)
For the first term,

T1​\[f~\+ϵ​h\]=‖∂tfϵ\+\(∂xfϵ\)​x˙t‖2\\displaystyle T\_\{1\}\[\\tilde\{f\}\+\\epsilon h\]=\\\|\\partial\_\{t\}f\_\{\\epsilon\}\+\(\\partial\_\{x\}f\_\{\\epsilon\}\)\\dot\{x\}\_\{t\}\\\|^\{2\}Differentiating

dd​ϵ​T1​\[f~\+ϵ​h\]\\displaystyle\\frac\{d\}\{d\\epsilon\}T\_\{1\}\[\\tilde\{f\}\+\\epsilon h\]=2​\(∂tfϵ\+\(∂xfϵ\)​x˙t\)⊤​dd​ϵ​\(∂tfϵ\+\(∂xfϵ\)​x˙t\)\\displaystyle=2\\Big\(\\partial\_\{t\}f\_\{\\epsilon\}\+\(\\partial\_\{x\}f\_\{\\epsilon\}\)\\dot\{x\}\_\{t\}\\Big\)^\{\\top\}\\frac\{d\}\{d\\epsilon\}\\Big\(\\partial\_\{t\}f\_\{\\epsilon\}\+\(\\partial\_\{x\}f\_\{\\epsilon\}\)\\dot\{x\}\_\{t\}\\Big\)\(34\)=2​\(∂tfϵ\+\(∂xfϵ\)​x˙t\)⊤​\(\(u−t\)​∂th−h⏟\+\(u−t\)​∂xh⏟​x˙t\)\\displaystyle=2\\Big\(\\partial\_\{t\}f\_\{\\epsilon\}\+\(\\partial\_\{x\}f\_\{\\epsilon\}\)\\dot\{x\}\_\{t\}\\Big\)^\{\\top\}\\Big\(\\underbrace\{\(u\-t\)\\partial\_\{t\}h\-h\}\+\\underbrace\{\(u\-t\)\\partial\_\{x\}h\}\\dot\{x\}\_\{t\}\\Big\)\(35\)=2​\(\(u−t\)​\(∂tf~\+ϵ​∂th\)−\(f~\+ϵ​h\)⏟\+\(I\+\(u−t\)​∂x\(f~\+ϵ​h\)⏟\)​x˙t\)⊤\\displaystyle=2\\Big\(\\underbrace\{\(u\-t\)\(\\partial\_\{t\}\\tilde\{f\}\+\\epsilon\\partial\_\{t\}h\)\-\(\\tilde\{f\}\+\\epsilon h\)\}\+\(\\underbrace\{I\+\(u\-t\)\\partial\_\{x\}\(\\tilde\{f\}\+\\epsilon h\)\}\)\\dot\{x\}\_\{t\}\\Big\)^\{\\top\}\(36\)\(\(u−t\)​∂th−h⏟\+\(u−t\)​∂xh⏟​x˙t\)\\displaystyle\\quad\\Big\(\\underbrace\{\(u\-t\)\\partial\_\{t\}h\-h\}\+\\underbrace\{\(u\-t\)\\partial\_\{x\}h\}\\dot\{x\}\_\{t\}\\Big\)\(37\)So

dd​ϵ​T1​\[f~\+ϵ​h\]\|ϵ=0\\displaystyle\\frac\{d\}\{d\\epsilon\}T\_\{1\}\[\\tilde\{f\}\+\\epsilon h\]\\Big\|\_\{\\epsilon=0\}=2​\(\(u−t\)​\(∂tf~\+∂xf~​x˙t\)−f~\+x˙t\)⊤​\(\(u−t\)​\(∂th\+∂xh​x˙t\)−h\)\\displaystyle=2\\Big\(\(u\-t\)\(\\partial\_\{t\}\\tilde\{f\}\+\\partial\_\{x\}\\tilde\{f\}\\dot\{x\}\_\{t\}\)\-\\tilde\{f\}\+\\dot\{x\}\_\{t\}\\Big\)^\{\\top\}\\Big\(\(u\-t\)\(\\partial\_\{t\}h\+\\partial\_\{x\}h\\dot\{x\}\_\{t\}\)\-h\\Big\)\(38\)For the second term,

T2​\[f~\+ϵ​h\]\\displaystyle T\_\{2\}\[\\tilde\{f\}\+\\epsilon h\]=‖∂xfϵ​\(x˙t−v\)‖2\\displaystyle=\\\|\\partial\_\{x\}f\_\{\\epsilon\}\(\\dot\{x\}\_\{t\}\-v\)\\\|^\{2\}\(39\)and

dd​ϵ​T2​\[f~\+ϵ​h\]\\displaystyle\\frac\{d\}\{d\\epsilon\}T\_\{2\}\[\\tilde\{f\}\+\\epsilon h\]=2​\[∂xfϵ​\(x˙t−v\)\]⊤​dd​ϵ​\[∂xfϵ​\(x˙t−v\)\]\\displaystyle=2\\Big\[\\partial\_\{x\}f\_\{\\epsilon\}\(\\dot\{x\}\_\{t\}\-v\)\\Big\]^\{\\top\}\\frac\{d\}\{d\\epsilon\}\\Big\[\\partial\_\{x\}f\_\{\\epsilon\}\(\\dot\{x\}\_\{t\}\-v\)\\Big\]\(40\)=2​\[∂xfϵ​\(x˙t−v\)\]⊤​dd​ϵ​\(∂xfϵ\)​\(x˙t−v\)\\displaystyle=2\\Big\[\\partial\_\{x\}f\_\{\\epsilon\}\(\\dot\{x\}\_\{t\}\-v\)\\Big\]^\{\\top\}\\frac\{d\}\{d\\epsilon\}\(\\partial\_\{x\}f\_\{\\epsilon\}\)\(\\dot\{x\}\_\{t\}\-v\)\(41\)=2​\[\(I\+\(u−t\)​∂x\(f~\+ϵ​h\)\)​\(x˙t−v\)\]⊤​\(u−t\)​∂xh​\(x˙t−v\)\\displaystyle=2\\Big\[\\Big\(I\+\(u\-t\)\\partial\_\{x\}\(\\tilde\{f\}\+\\epsilon h\)\\Big\)\(\\dot\{x\}\_\{t\}\-v\)\\Big\]^\{\\top\}\(u\-t\)\\partial\_\{x\}h\(\\dot\{x\}\_\{t\}\-v\)\(42\)So

dd​ϵ​T2​\[f~\+ϵ​h\]\|ϵ=0\\displaystyle\\frac\{d\}\{d\\epsilon\}T\_\{2\}\[\\tilde\{f\}\+\\epsilon h\]\\Big\|\_\{\\epsilon=0\}=2​\[\(I\+\(u−t\)​∂xf~\)​\(x˙t−v\)\]⊤​\(u−t\)​∂xh​\(x˙t−v\)\\displaystyle=2\\Big\[\\Big\(I\+\(u\-t\)\\partial\_\{x\}\\tilde\{f\}\\Big\)\(\\dot\{x\}\_\{t\}\-v\)\\Big\]^\{\\top\}\(u\-t\)\\partial\_\{x\}h\(\\dot\{x\}\_\{t\}\-v\)\(43\)So combining

dd​ϵ​A\|ϵ=0\\displaystyle\\frac\{d\}\{d\\epsilon\}A\\Big\|\_\{\\epsilon=0\}=2​\(\(u−t\)​\(∂tf~\+∂xf~​x˙t\)−f~\+x˙t\)⊤​\(\(u−t\)​\(∂th\+∂xh​x˙t\)−h\)\\displaystyle=2\\Big\(\(u\-t\)\(\\partial\_\{t\}\\tilde\{f\}\+\\partial\_\{x\}\\tilde\{f\}\\dot\{x\}\_\{t\}\)\-\\tilde\{f\}\+\\dot\{x\}\_\{t\}\\Big\)^\{\\top\}\\Big\(\(u\-t\)\(\\partial\_\{t\}h\+\\partial\_\{x\}h\\dot\{x\}\_\{t\}\)\-h\\Big\)\(44\)−2​\[\(I\+\(u−t\)​∂xf~\)​\(x˙t−v\)\]⊤​\(u−t\)​∂xh​\(x˙t−v\)\\displaystyle\-2\\Big\[\\Big\(I\+\(u\-t\)\\partial\_\{x\}\\tilde\{f\}\\Big\)\(\\dot\{x\}\_\{t\}\-v\)\\Big\]^\{\\top\}\(u\-t\)\\partial\_\{x\}h\(\\dot\{x\}\_\{t\}\-v\)\(45\)Att=ut=u, this simplifies

1​\[t=u\]​dd​ϵ​A​\[f~ϵ\]\|ϵ=0=2​\(x˙t−f~\)⊤​\(−h\)\\displaystyle 1\[t=u\]\\frac\{d\}\{d\\epsilon\}A\[\\tilde\{f\}\_\{\\epsilon\}\]\\Big\|\_\{\\epsilon=0\}=2\(\\dot\{x\}\_\{t\}\-\\tilde\{f\}\)^\{\\top\}\(\-h\)\(46\)which is the first variation for regression that makesf~\\tilde\{f\}equal toE​\[x˙t\|xt\]E\[\\dot\{x\}\_\{t\}\|x\_\{t\}\]\. Summarizing,

δ​ℒ​\[f~;h\]\\displaystyle\\delta\\mathcal\{L\}\[\\tilde\{f\};h\]=𝔼\[2\(\(u−t\)\(∂tf~\+∂xf~x˙t\)−f~\+x˙t\)⊤\(\(u−t\)\(∂th\+∂xhx˙t\)−h\)\\displaystyle=\\mathbb\{E\}\\Big\[2\\Big\(\(u\-t\)\(\\partial\_\{t\}\\tilde\{f\}\+\\partial\_\{x\}\\tilde\{f\}\\dot\{x\}\_\{t\}\)\-\\tilde\{f\}\+\\dot\{x\}\_\{t\}\\Big\)^\{\\top\}\\Big\(\(u\-t\)\(\\partial\_\{t\}h\+\\partial\_\{x\}h\\dot\{x\}\_\{t\}\)\-h\\Big\)\(47\)−2\[\(I\+\(u−t\)∂xf~\)\(x˙t−v\)\]⊤\(u−t\)∂xh\(x˙t−v\)\]\\displaystyle\-2\\big\[\\big\(I\+\(u\-t\)\\partial\_\{x\}\\tilde\{f\}\\big\)\(\\dot\{x\}\_\{t\}\-v\)\\big\]^\{\\top\}\(u\-t\)\\partial\_\{x\}h\(\\dot\{x\}\_\{t\}\-v\)\\Big\]\(48\)

### A\.4Lemma: velocity matches at a stationary point of original functional

###### Lemma 1\.

Letf~∗\\tilde\{f\}^\{\*\}be a stationary point ofℒ\\mathcal\{L\}\. Assume thatf~∗\\tilde\{f\}^\{\*\}is bounded\. Assume thatf~∗,v∈C1\\tilde\{f\}^\{\*\},v\\in C^\{1\}in arguments\(t,u,x\)\(t,u,x\)and that all expectations of terms featured in the integrand ofℒ\\mathcal\{L\}\(i\.e\.,v,f~,∂tf~,∂uf~,∂xf~,…v,\\tilde\{f\},\\partial\_\{t\}\\tilde\{f\},\\partial\_\{u\}\\tilde\{f\},\\partial\_\{x\}\\tilde\{f\},\\ldots\) are finite\. Then we have thatf~∗​\(t,t,⋅\)=v​\(t,⋅\)\\tilde\{f\}^\{\*\}\(t,t,\\cdot\)=v\(t,\\cdot\)where the velocityv​\(t,x\)=𝔼​\[x˙t\|xt\]v\(t,x\)=\\mathbb\{E\}\[\\dot\{x\}\_\{t\}\|x\_\{t\}\]\.

###### Proof\.

We proceed by contradiction\. By the premise, we are at a stationary pointf~∗\\tilde\{f\}^\{\*\}\. Letf∗:=x\+\(u−t\)​f~∗f^\{\*\}:=x\+\(u\-t\)\\tilde\{f\}^\{\*\}\. By the definition of stationary point in[section˜A\.1](https://arxiv.org/html/2607.26398#A1.SS1), we have thatδ​ℒ​\[f~∗;h\]=0\\delta\\mathcal\{L\}\[\\tilde\{f\}^\{\*\};h\]=0for all admissiblehh\. Suppose for the sake of contradiction that at this stationary point, the velocity does not match, meaning

−∂tf∗​\(t,t,⋅\)=f~∗​\(t,t,⋅\)​≠⏟suppose for contradiction​v​\(t,⋅\)\\displaystyle\-\\partial\_\{t\}f^\{\*\}\(t,t,\\cdot\)=\\tilde\{f\}^\{\*\}\(t,t,\\cdot\)\\underbrace\{\\neq\}\_\{\\text\{suppose for contradiction\}\}v\(t,\\cdot\)\(49\)
The proof proceeds by picking a direction for which the first variation is nonzero, providing a contradiction to being at a stationary point\. The contradiction \(the direction for which the first variation is nonzero\) is constructed to arise from assuming that the velocity does not match, meaning that by contradiction the velocity does not match\. Specific care is taken to ensure that this direction is admissible, in this case meaning it is a continuous function\.

We name a sequence of functionsgηg\_\{\\eta\}such that there existsη∗\\eta^\{\*\}such thatgη∗g\_\{\\eta^\{\*\}\}is continuous but yields the nonzero variation when chosen as a direction\. To establish this existence under continuity, the dominated convergence theorem is used\.

Let us defineg​\(t,u,x\)=1​\[t=u\]​\(f~∗​\(t,u,x\)−v​\(t,x\)\)g\(t,u,x\)=1\[t=u\]\\Big\(\\tilde\{f\}^\{\*\}\(t,u,x\)\-v\(t,x\)\\Big\)and evaluate it att=ut=uso thatg​\(t,t,x\)=f~∗​\(t,t,x\)−v​\(t,x\)g\(t,t,x\)=\\tilde\{f\}^\{\*\}\(t,t,x\)\-v\(t,x\)\. We then define the soft indicatorIη​\(t,u\)I\_\{\\eta\}\(t,u\)that goes to1​\[t=u\]1\[t=u\]asη→0\\eta\\to 0and define it as:

Iη​\(t,u\)\\displaystyle I\_\{\\eta\}\(t,u\)=1​\[η\>0\]​2​\(1−11\+exp⁡\(−1η2​\(t−u\)2\)\)\+1​\[η=0\]​1​\[t=u\]\\displaystyle=1\[\\eta\>0\]2\\Big\(1\-\\frac\{1\}\{1\+\\exp\(\-\\frac\{1\}\{\\eta^\{2\}\}\(t\-u\)^\{2\}\)\}\\Big\)\+1\[\\eta=0\]1\[t=u\]\(50\)Using the soft indicator, we define a sequence of functionsgηg\_\{\\eta\}so that asη→0\\eta\\to 0we will have pointwise convergence ofgη​\(t,t,x\)→g​\(t,t,x\)g\_\{\\eta\}\(t,t,x\)\\to g\(t,t,x\)which also means thatgη​\(t,u,x\)g\_\{\\eta\}\(t,u,x\)fort≠ut\\neq ugoes to0\. We pick

gη​\(t,u,x\)=Iη​\(t,u\)​\(f~∗​\(t,t,x\)−v​\(t,x\)\)\\displaystyle g\_\{\\eta\}\(t,u,x\)=I\_\{\\eta\}\(t,u\)\\Big\(\\tilde\{f\}^\{\*\}\(t,t,x\)\-v\(t,x\)\\Big\)\(51\)Pointwise convergence of g in eta\.We first establish pointwise convergence ofgηg\_\{\\eta\}toggfor all arguments\(t,u,x\)\(t,u,x\)asη→0\\eta\\to 0from the right\.

∀η^≥0,limη→\(η^\)\+gη​\(t,u,x\)=gη^​\(t,u,x\)\\displaystyle\\forall\\hat\{\\eta\}\\geq 0,\\quad\\lim\_\{\\eta\\to\(\\hat\{\\eta\}\)^\{\+\}\}g\_\{\\eta\}\(t,u,x\)=g\_\{\\hat\{\\eta\}\}\(t,u,x\)\(52\)To show this, for anyη^\>0\\hat\{\\eta\}\>0, use continuity ofgη^​\(t,u,x\)g\_\{\\hat\{\\eta\}\}\(t,u,x\)inη^\\hat\{\\eta\}\(product of function withoutη\\etatimes the soft indicator which is continuous\)\. Then to establish forη^=0\\hat\{\\eta\}=0, we consider two casest=ut=uandt≠ut\\neq u\. For equality:

limη→0\+gη​\(t,t,x\)\\displaystyle\\lim\_\{\\eta\\to 0^\{\+\}\}g\_\{\\eta\}\(t,t,x\)\(53\)=limη→0\+Iη​\(t,t\)​\(f~∗​\(t,t,x\)−v​\(t,x\)\)\\displaystyle=\\lim\_\{\\eta\\to 0^\{\+\}\}I\_\{\\eta\}\(t,t\)\\Big\(\\tilde\{f\}^\{\*\}\(t,t,x\)\-v\(t,x\)\\Big\)\(54\)=limη→0\+\[1​\[η\>0\]​2​\(1−11\+exp⁡\(0\)\)\+1​\[η=0\]\]​\(f~∗​\(t,t,x\)−v​\(t,x\)\)\\displaystyle=\\lim\_\{\\eta\\to 0^\{\+\}\}\\Big\[1\[\\eta\>0\]2\\Big\(1\-\\frac\{1\}\{1\+\\exp\(0\)\}\\Big\)\+1\[\\eta=0\]\\Big\]\\Big\(\\tilde\{f\}^\{\*\}\(t,t,x\)\-v\(t,x\)\\Big\)\(55\)=limη→0\+\[1​\[η\>0\]​1\+1​\[η=0\]\]​\(f~∗​\(t,t,x\)−v​\(t,x\)\)\\displaystyle=\\lim\_\{\\eta\\to 0^\{\+\}\}\\Big\[1\[\\eta\>0\]1\+1\[\\eta=0\]\\Big\]\\Big\(\\tilde\{f\}^\{\*\}\(t,t,x\)\-v\(t,x\)\\Big\)\(56\)=limη→0\+1​\[η≥0\]​\(f~∗​\(t,t,x\)−v​\(t,x\)\)\\displaystyle=\\lim\_\{\\eta\\to 0^\{\+\}\}1\[\\eta\\geq 0\]\\Big\(\\tilde\{f\}^\{\*\}\(t,t,x\)\-v\(t,x\)\\Big\)\(57\)=f~∗−v\\displaystyle=\\tilde\{f\}^\{\*\}\-v\(58\)which equalsgη=0​\(t,t,x\)g\_\{\\eta=0\}\(t,t,x\), establishing continuity\. Now fort≠ut\\neq u\.∀δ\>0\\forall\\delta\>0we must name anη​\(δ\)\\eta\(\\delta\)such that\|gη​\(δ\)−g0\|<δ\|g\_\{\\eta\(\\delta\)\}\-g\_\{0\}\|<\\delta, i\.e\.,\|gη​\(δ\)−0\|<δ\|g\_\{\\eta\(\\delta\)\}\-0\|<\\delta\. Assume\|f~∗−v\|<k\|\\tilde\{f\}^\{\*\}\-v\|<kuniformly in all input values t,x\.

limη→0\+gη​\(t,u,x\)\\displaystyle\\lim\_\{\\eta\\to 0^\{\+\}\}g\_\{\\eta\}\(t,u,x\)=limη→0\+\[1​\[η\>0\]​2​\(1−11\+exp⁡\(−1η2​\(t−u\)2\)\)\+1​\[η=0\]​1​\[t=u\]\]​\(f~∗​\(t,t,x\)−v​\(t,x\)\)\\displaystyle=\\lim\_\{\\eta\\to 0^\{\+\}\}\\Big\[1\[\\eta\>0\]2\\Big\(1\-\\frac\{1\}\{1\+\\exp\(\-\\frac\{1\}\{\\eta^\{2\}\}\(t\-u\)^\{2\}\)\}\\Big\)\+1\[\\eta=0\]1\[t=u\]\\Big\]\\Big\(\\tilde\{f\}^\{\*\}\(t,t,x\)\-v\(t,x\)\\Big\)=limη→0\+\[1​\[η\>0\]​2​\(1−11\+exp⁡\(−1η2​\(t−u\)2\)\)\]​\(f~∗​\(t,t,x\)−v​\(t,x\)\)\\displaystyle=\\lim\_\{\\eta\\to 0^\{\+\}\}\\Big\[1\[\\eta\>0\]2\\Big\(1\-\\frac\{1\}\{1\+\\exp\(\-\\frac\{1\}\{\\eta^\{2\}\}\(t\-u\)^\{2\}\)\}\\Big\)\\Big\]\\Big\(\\tilde\{f\}^\{\*\}\(t,t,x\)\-v\(t,x\)\\Big\)Since we are finding aδ\\deltaandη​\(δ\)\\eta\(\\delta\)so that\|gη​\(δ\)−0\|<δ\|g\_\{\\eta\(\\delta\)\}\-0\|<\\deltawhich means\|gη​\(δ\)\|<δ\|g\_\{\\eta\(\\delta\)\}\|<\\delta, this just means we can setδ\\deltato an upper bound on the term we are limiting: let the indicator take on11as when it is 0 we are done\.

δ\\displaystyle\\delta=\|2​\(1−11\+exp⁡\(…\)\)\|​k\\displaystyle=\\Big\|2\(1\-\\frac\{1\}\{1\+\\exp\(\.\.\.\)\}\)\\Big\|k\(59\)⇔δk=2​\(1−11\+exp⁡\(−1η2​\(t−u\)2\)\)\\displaystyle\\iff\\frac\{\\delta\}\{k\}=2\(1\-\\frac\{1\}\{1\+\\exp\(\-\\frac\{1\}\{\\eta^\{2\}\}\(t\-u\)^\{2\}\)\}\)\(60\)⇔δ2​k=1−11\+exp⁡\(−1η2​\(t−u\)2\)\\displaystyle\\iff\\frac\{\\delta\}\{2k\}=1\-\\frac\{1\}\{1\+\\exp\(\-\\frac\{1\}\{\\eta^\{2\}\}\(t\-u\)^\{2\}\)\}\(61\)⇔1−δ2​k=11\+exp⁡\(−1η2​\(t−u\)2\)\\displaystyle\\iff 1\-\\frac\{\\delta\}\{2k\}=\\frac\{1\}\{1\+\\exp\(\-\\frac\{1\}\{\\eta^\{2\}\}\(t\-u\)^\{2\}\)\}\(62\)⇔2​k2​k−δ2​k=11\+exp⁡\(−1η2​\(t−u\)2\)\\displaystyle\\iff\\frac\{2k\}\{2k\}\-\\frac\{\\delta\}\{2k\}=\\frac\{1\}\{1\+\\exp\(\-\\frac\{1\}\{\\eta^\{2\}\}\(t\-u\)^\{2\}\)\}\(63\)⇔2​k−δ2​k=11\+exp⁡\(−1η2​\(t−u\)2\)\\displaystyle\\iff\\frac\{2k\-\\delta\}\{2k\}=\\frac\{1\}\{1\+\\exp\(\-\\frac\{1\}\{\\eta^\{2\}\}\(t\-u\)^\{2\}\)\}\(64\)⇔2​k2​k−δ=1\+exp⁡\(−1η2​\(t−u\)2\)\\displaystyle\\iff\\frac\{2k\}\{2k\-\\delta\}=1\+\\exp\(\-\\frac\{1\}\{\\eta^\{2\}\}\(t\-u\)^\{2\}\)\(65\)⇔2​k2​k−δ−1=exp⁡\(−1η2​\(t−u\)2\)\\displaystyle\\iff\\frac\{2k\}\{2k\-\\delta\}\-1=\\exp\(\-\\frac\{1\}\{\\eta^\{2\}\}\(t\-u\)^\{2\}\)\(66\)⇔2​k2​k−δ−2​k−δ2​k−δ=exp⁡\(−1η2​\(t−u\)2\)\\displaystyle\\iff\\frac\{2k\}\{2k\-\\delta\}\-\\frac\{2k\-\\delta\}\{2k\-\\delta\}=\\exp\(\-\\frac\{1\}\{\\eta^\{2\}\}\(t\-u\)^\{2\}\)\(67\)⇔δ2​k−δ=exp⁡\(−1η2​\(t−u\)2\)\\displaystyle\\iff\\frac\{\\delta\}\{2k\-\\delta\}=\\exp\(\-\\frac\{1\}\{\\eta^\{2\}\}\(t\-u\)^\{2\}\)\(68\)⇔log⁡δ2​k−δ=−1η2​\(t−u\)2\\displaystyle\\iff\\log\\frac\{\\delta\}\{2k\-\\delta\}=\-\\frac\{1\}\{\\eta^\{2\}\}\(t\-u\)^\{2\}\(69\)⇔−log⁡δ2​k−δ=1η2​\(t−u\)2\\displaystyle\\iff\-\\log\\frac\{\\delta\}\{2k\-\\delta\}=\\frac\{1\}\{\\eta^\{2\}\}\(t\-u\)^\{2\}\(70\)⇔−log⁡δ2​k−δ\(t−u\)2=1η2\\displaystyle\\iff\-\\frac\{\\log\\frac\{\\delta\}\{2k\-\\delta\}\}\{\(t\-u\)^\{2\}\}=\\frac\{1\}\{\\eta^\{2\}\}\(71\)⇔−\(t−u\)2log⁡δ2​k−δ=η2\\displaystyle\\iff\-\\frac\{\(t\-u\)^\{2\}\}\{\\log\\frac\{\\delta\}\{2k\-\\delta\}\}=\\eta^\{2\}\(72\)Now note that the soft indicator is strictly<1<1and thatgηg\_\{\\eta\}for fixed\(t,u,x\)\(t,u,x\)is between−k\-kand0or0andkkdepending on the sign off~∗−v\\tilde\{f\}^\{\*\}\-v, but never both\. So its magnitude is at mostkk\. So

\|gη​\(δ\)−g0\|=\|gη​\(δ\)−0\|<δ<k\\displaystyle\|g\_\{\\eta\(\\delta\)\}\-g\_\{0\}\|=\|g\_\{\\eta\(\\delta\)\}\-0\|<\\delta<k\(73\)This can help us ascertain that the above square root to solve forη\\etawill be well defined:

δ<k\\displaystyle\\delta<k⟹2​k−δ\>k\\displaystyle\\implies 2k\-\\delta\>k\(74\)⟹12​k−δ<1k\\displaystyle\\implies\\frac\{1\}\{2k\-\\delta\}<\\frac\{1\}\{k\}\(75\)⟹δ2​k−δ<δk\\displaystyle\\implies\\frac\{\\delta\}\{2k\-\\delta\}<\\frac\{\\delta\}\{k\}\(76\)⟹δ2​k−δ<1\\displaystyle\\implies\\frac\{\\delta\}\{2k\-\\delta\}<1\(77\)⟹log⁡δ2​k−δ<log⁡1\\displaystyle\\implies\\log\\frac\{\\delta\}\{2k\-\\delta\}<\\log 1\(78\)⟹log⁡δ2​k−δ<0\\displaystyle\\implies\\log\\frac\{\\delta\}\{2k\-\\delta\}<0\(79\)meaning

η​\(δ\)=\(t−u\)2\|log⁡δ2​k−δ\|\\displaystyle\\eta\(\\delta\)=\\sqrt\{\\frac\{\(t\-u\)^\{2\}\}\{\|\\log\\frac\{\\delta\}\{2k\-\\delta\}\|\}\}\(80\)thus establishing convergence ofgη→gg\_\{\\eta\}\\to gasη→0\+\\eta\\to 0^\{\+\}for each\(t,u,x\)\(t,u,x\)i\.e\. pointwise convergence\.

Now recall the first variation ofℒ\\mathcal\{L\}\([section˜A\.3](https://arxiv.org/html/2607.26398#A1.SS3)\) and consider it as a function ofη\\eta:

s​\(η\):=δ​ℒ​\[f~∗;gη\]\\displaystyle s\(\\eta\):=\\delta\\mathcal\{L\}\[\\tilde\{f\}^\{\*\};g\_\{\\eta\}\]=𝔼\[2\(\(u−t\)\(∂tf~∗\)−f~∗\+\(I\+\(u−t\)\(∂xf~∗\)\)x˙t\)⊤\\displaystyle=\\mathbb\{E\}\\Bigg\[2\\Big\(\(u\-t\)\(\\partial\_\{t\}\\tilde\{f\}^\{\*\}\)\-\\tilde\{f\}^\{\*\}\+\(I\+\(u\-t\)\(\\partial\_\{x\}\\tilde\{f\}^\{\*\}\)\)\\dot\{x\}\_\{t\}\\Big\)^\{\\top\}\(81\)\(\(u−t\)​∂tgη−gη\+\(u−t\)​\(∂xgη\)​x˙t\)\\displaystyle\\quad\\Big\(\(u\-t\)\\partial\_\{t\}g\_\{\\eta\}\-g\_\{\\eta\}\+\(u\-t\)\(\\partial\_\{x\}g\_\{\\eta\}\)\\dot\{x\}\_\{t\}\\Big\)\(82\)−2\[\(I\+\(u−t\)∂xf~∗\)\(x˙t−v\)\]⊤\(u−t\)∂xgη\(x˙t−v\)\]\\displaystyle\-2\\Big\[\\Big\(I\+\(u\-t\)\\partial\_\{x\}\\tilde\{f\}^\{\*\}\\Big\)\(\\dot\{x\}\_\{t\}\-v\)\\Big\]^\{\\top\}\(u\-t\)\\partial\_\{x\}g\_\{\\eta\}\(\\dot\{x\}\_\{t\}\-v\)\\Bigg\]\(83\)Pointwise convergence of integrand in eta\.Collect the variablesω=\(t,u,x0,x1\)\\omega=\(t,u,x\_\{0\},x\_\{1\}\)and recall thatxtx\_\{t\}andx˙t\\dot\{x\}\_\{t\}are functions of\(x0,x1\)\(x\_\{0\},x\_\{1\}\)\. Defineϕη​\(ω\)\\phi\_\{\\eta\}\(\\omega\)as shorthand for the expectand so thats​\(η\)=𝔼​\[ϕη​\(ω\)\]s\(\\eta\)=\\mathbb\{E\}\[\\phi\_\{\\eta\}\(\\omega\)\]\. Under the boundedness assumptions and noting thatϕη\\phi\_\{\\eta\}only polynomially combinesgηg\_\{\\eta\}with\(∂tf~∗,∂uf~∗,∂xf~∗,v,∂xv,…\)\(\\partial\_\{t\}\\tilde\{f\}^\{\*\},\\partial\_\{u\}\\tilde\{f\}^\{\*\},\\partial\_\{x\}\\tilde\{f\}^\{\*\},v,\\partial\_\{x\}v,\\ldots\), similar reasoning used to showgη→gg\_\{\\eta\}\\to gcan also be used to establish thatϕη→ϕ\\phi\_\{\\eta\}\\to\\phipointwise\.

Establish upper envelope\.In addition to pointwise convergence ofϕη→ϕ\\phi\_\{\\eta\}\\to\\phiasη→0\\eta\\to 0from the right, we need an upper envelopeG​\(ω\)G\(\\omega\)\. Beyond the assumptions, the only thing needed to show that an upper envelope exists is to control the term\|\(u−t\)​∂tIη​\(t,u\)\|\|\(u\-t\)\\partial\_\{t\}I\_\{\\eta\}\(t,u\)\|\. The derivative of the soft indicator appears since∂tgη=∂t\(Iη​g\)=\(∂tIη\)​g\+Iη​∂tg\\partial\_\{t\}g\_\{\\eta\}=\\partial\_\{t\}\(I\_\{\\eta\}g\)=\(\\partial\_\{t\}I\_\{\\eta\}\)g\+I\_\{\\eta\}\\partial\_\{t\}g\. Atη=0\\eta=0we haveIη​\(t,u\)=1​\[t=u\]I\_\{\\eta\}\(t,u\)=1\[t=u\], so\(u−t\)​∂tIη​\(t,u\)\(u\-t\)\\,\\partial\_\{t\}I\_\{\\eta\}\(t,u\)vanishes identically: it is0fort≠ut\\neq ubecauseI0I\_\{0\}is constant, and fort=ut=ubecause of the\(u−t\)\(u\-t\)prefactor\. Forη\>0\\eta\>0, define:

z:=\(t−u\)2η2,σ​\(z\):=11\+exp⁡\(−z\),σ′​\(z\):=exp⁡\(−z\)\[1\+exp⁡\(−z\)\]2\\displaystyle z:=\\frac\{\(t\-u\)^\{2\}\}\{\\eta^\{2\}\},\\quad\\sigma\(z\):=\\frac\{1\}\{1\+\\exp\(\-z\)\},\\quad\\sigma^\{\\prime\}\(z\):=\\frac\{\\exp\(\-z\)\}\{\[1\+\\exp\(\-z\)\]^\{2\}\}\(84\)Note that forη\>0\\eta\>0,

∂tIη​\(t,u\)=2​2​\(t−u\)η2​σ′​\(\(t−u\)2η2\)\\displaystyle\\partial\_\{t\}I\_\{\\eta\}\(t,u\)=2\\frac\{2\(t\-u\)\}\{\\eta^\{2\}\}\\sigma^\{\\prime\}\(\\frac\{\(t\-u\)^\{2\}\}\{\\eta^\{2\}\}\)\(85\)and so

\|\(u−t\)​∂tIη​\(t,u\)\|=2​\(u−t\)​2​\(u−t\)η2​σ′​\(\(t−u\)2η2\)=4​z​σ′​\(z\)≤supz≥04​z​σ′​\(z\)≤C0<∞\\displaystyle\|\(u\-t\)\\partial\_\{t\}I\_\{\\eta\}\(t,u\)\|=2\(u\-t\)\\frac\{2\(u\-t\)\}\{\\eta^\{2\}\}\\sigma^\{\\prime\}\(\\frac\{\(t\-u\)^\{2\}\}\{\\eta^\{2\}\}\)=4z\\sigma^\{\\prime\}\(z\)\\leq\\sup\_\{z\\geq 0\}4z\\sigma^\{\\prime\}\(z\)\\leq C\_\{0\}<\\infty\(86\)Becauser​\(z\):=4​z​σ′​\(z\)r\(z\):=4z\\sigma^\{\\prime\}\(z\)is continuous and satisfiesr​\(0\)=0r\(0\)=0andr​\(z\)→0r\(z\)\\to 0asz→∞z\\to\\infty, it attains a finite maximum\.This bound is independent of𝜼\\boldsymbol\{\\eta\}\. Thus every term containing\(u−t\)​∂tgη\(u\-t\)\\partial\_\{t\}g\_\{\\eta\}is uniformly bounded inη\\etaby a product of a constant \(from the bound andIη∈\[0,1\)I\_\{\\eta\}\\in\[0,1\)\)\. The other quantities inϕ\\phiare bounded by assumption\. Thus such a bounding envelopeG​\(ω\)G\(\\omega\)exists\.

Using dominated convergence\.First,

limη→0\+s​\(η\)=limη→0\+𝔼​\[ϕη\]=limη→0\+p​\(t=u\)​𝔼​\[ϕη\|t=u\]\+p​\(t<u\)​𝔼​\[ϕη\|t<u\]\\displaystyle\\lim\_\{\\eta\\to 0^\{\+\}\}s\(\\eta\)=\\lim\_\{\\eta\\to 0^\{\+\}\}\\mathbb\{E\}\[\\phi\_\{\\eta\}\]=\\lim\_\{\\eta\\to 0^\{\+\}\}p\(t=u\)\\mathbb\{E\}\[\\phi\_\{\\eta\}\\,\|\\,t=u\]\+p\(t<u\)\\mathbb\{E\}\[\\phi\_\{\\eta\}\\,\|\\,t<u\]\(87\)By the pointwise convergence ofϕη→ϕ\\phi\_\{\\eta\}\\to\\phiasη→0\+\\eta\\to 0^\{\+\}and by the envelope, we can compute the limit of the first and second terms separately\. Expanding the first term \(withp​\(t=u\)p\(t=u\)\):

limη→0\+p​\(t=u\)​𝔼​\[ϕη\|t=u\]\\displaystyle\\lim\_\{\\eta\\to 0^\{\+\}\}p\(t=u\)\\mathbb\{E\}\[\\phi\_\{\\eta\}\\,\|\\,t=u\]=limη→0\+p\(t=u\)𝔼\[2\(\(u−t\)\(∂tf~∗\)−f~∗\+\(I\+\(u−t\)\(∂xf~∗\)\)x˙t\)⊤\\displaystyle=\\lim\_\{\\eta\\to 0^\{\+\}\}p\(t=u\)\\mathbb\{E\}\\Bigg\[2\\Big\(\(u\-t\)\(\\partial\_\{t\}\\tilde\{f\}^\{\*\}\)\-\\tilde\{f\}^\{\*\}\+\(I\+\(u\-t\)\(\\partial\_\{x\}\\tilde\{f\}^\{\*\}\)\)\\dot\{x\}\_\{t\}\\Big\)^\{\\top\}\(\(u−t\)​∂tgη−gη\+\(u−t\)​\(∂xgη\)​x˙t\)\\displaystyle\\quad\\Big\(\(u\-t\)\\partial\_\{t\}g\_\{\\eta\}\-g\_\{\\eta\}\+\(u\-t\)\(\\partial\_\{x\}g\_\{\\eta\}\)\\dot\{x\}\_\{t\}\\Big\)−2\[\(I\+\(u−t\)∂xf~∗\)\(x˙t−v\)\]⊤\(u−t\)∂xgη\(x˙t−v\)\|t=u\]\\displaystyle\-2\\Big\[\\Big\(I\+\(u\-t\)\\partial\_\{x\}\\tilde\{f\}^\{\*\}\\Big\)\(\\dot\{x\}\_\{t\}\-v\)\\Big\]^\{\\top\}\(u\-t\)\\partial\_\{x\}g\_\{\\eta\}\(\\dot\{x\}\_\{t\}\-v\)\\,\|\\,t=u\\Bigg\]=limη→0\+p​\(t=u\)​𝔼​\[2​\(−f~∗\+x˙t\)⊤​\(−gη\)\|t=u\]\\displaystyle=\\lim\_\{\\eta\\to 0^\{\+\}\}p\(t=u\)\\mathbb\{E\}\\Bigg\[2\\Big\(\-\\tilde\{f\}^\{\*\}\+\\dot\{x\}\_\{t\}\\Big\)^\{\\top\}\\Big\(\-g\_\{\\eta\}\\Big\)\\,\|\\,t=u\\Bigg\]=limη→0\+p​\(t=u\)​𝔼​\[2​\(−f~∗\+𝔼\[x˙t\|xt\]\)⊤​\(−gη\)\|t=u\]\\displaystyle=\\lim\_\{\\eta\\to 0^\{\+\}\}p\(t=u\)\\mathbb\{E\}\\Bigg\[2\\Big\(\-\\tilde\{f\}^\{\*\}\+\\mathop\{\\mathbb\{E\}\}\[\\dot\{x\}\_\{t\}\\,\|\\,x\_\{t\}\]\\Big\)^\{\\top\}\\Big\(\-g\_\{\\eta\}\\Big\)\\,\|\\,t=u\\Bigg\]=limη→0\+p​\(t=u\)​𝔼​\[2​\(−f~∗\+v\)⊤​\(−gη\)\|t=u\]\\displaystyle=\\lim\_\{\\eta\\to 0^\{\+\}\}p\(t=u\)\\mathbb\{E\}\\Bigg\[2\\Big\(\-\\tilde\{f\}^\{\*\}\+v\\Big\)^\{\\top\}\\Big\(\-g\_\{\\eta\}\\Big\)\\,\|\\,t=u\\Bigg\]=p​\(t=u\)​𝔼​\[limη→0\+2​\(−f~∗\+v\)⊤​\(−gη\)\|t=u\]\\displaystyle=p\(t=u\)\\mathbb\{E\}\\Bigg\[\\lim\_\{\\eta\\to 0^\{\+\}\}2\\Big\(\-\\tilde\{f\}^\{\*\}\+v\\Big\)^\{\\top\}\\Big\(\-g\_\{\\eta\}\\Big\)\\,\|\\,t=u\\Bigg\]=p​\(t=u\)​𝔼​\[limη→0\+2​‖−f~∗\+v‖22\|t=u\]\\displaystyle=p\(t=u\)\\mathbb\{E\}\\Bigg\[\\lim\_\{\\eta\\to 0^\{\+\}\}2\|\|\-\\tilde\{f\}^\{\*\}\+v\|\|\_\{2\}^\{2\}\\,\|\\,t=u\\Bigg\]This term is greater than zero by the assumption that the velocity does not match at the stationary point and the assumption of positive probabilityp​\(t=u\)\>0p\(t=u\)\>0\.

Expanding the second term \(withp​\(t<u\)p\(t<u\)\):

limη→0\+p​\(t<u\)​𝔼​\[ϕη\|t<u\]\\displaystyle\\lim\_\{\\eta\\to 0^\{\+\}\}p\(t<u\)\\mathbb\{E\}\[\\phi\_\{\\eta\}\\,\|\\,t<u\]=limη→0\+p\(t<u\)𝔼\[2\(\(u−t\)\(∂tf~∗\)−f~∗\+\(I\+\(u−t\)\(∂xf~∗\)\)x˙t\)⊤\\displaystyle=\\lim\_\{\\eta\\to 0^\{\+\}\}p\(t<u\)\\mathbb\{E\}\\Bigg\[2\\Big\(\(u\-t\)\(\\partial\_\{t\}\\tilde\{f\}^\{\*\}\)\-\\tilde\{f\}^\{\*\}\+\(I\+\(u\-t\)\(\\partial\_\{x\}\\tilde\{f\}^\{\*\}\)\)\\dot\{x\}\_\{t\}\\Big\)^\{\\top\}\(\(u−t\)​∂tgη−gη\+\(u−t\)​\(∂xgη\)​x˙t\)\\displaystyle\\quad\\Big\(\(u\-t\)\\partial\_\{t\}g\_\{\\eta\}\-g\_\{\\eta\}\+\(u\-t\)\(\\partial\_\{x\}g\_\{\\eta\}\)\\dot\{x\}\_\{t\}\\Big\)−2\[\(I\+\(u−t\)∂xf~∗\)\(x˙t−v\)\]⊤\(u−t\)∂xgη\(x˙t−v\)\|t<u\]\\displaystyle\-2\\Big\[\\Big\(I\+\(u\-t\)\\partial\_\{x\}\\tilde\{f\}^\{\*\}\\Big\)\(\\dot\{x\}\_\{t\}\-v\)\\Big\]^\{\\top\}\(u\-t\)\\partial\_\{x\}g\_\{\\eta\}\(\\dot\{x\}\_\{t\}\-v\)\\,\|\\,t<u\\Bigg\]=p\(t<u\)𝔼\[limη→0\+2\(\(u−t\)\(∂tf~∗\)−f~∗\+\(I\+\(u−t\)\(∂xf~∗\)\)x˙t\)⊤\\displaystyle=p\(t<u\)\\mathbb\{E\}\\Bigg\[\\lim\_\{\\eta\\to 0^\{\+\}\}2\\Big\(\(u\-t\)\(\\partial\_\{t\}\\tilde\{f\}^\{\*\}\)\-\\tilde\{f\}^\{\*\}\+\(I\+\(u\-t\)\(\\partial\_\{x\}\\tilde\{f\}^\{\*\}\)\)\\dot\{x\}\_\{t\}\\Big\)^\{\\top\}\(\(u−t\)​∂tgη−gη\+\(u−t\)​\(∂xgη\)​x˙t\)\\displaystyle\\quad\\Big\(\(u\-t\)\\partial\_\{t\}g\_\{\\eta\}\-g\_\{\\eta\}\+\(u\-t\)\(\\partial\_\{x\}g\_\{\\eta\}\)\\dot\{x\}\_\{t\}\\Big\)−2\[\(I\+\(u−t\)∂xf~∗\)\(x˙t−v\)\]⊤\(u−t\)∂xgη\(x˙t−v\)\|t<u\]\\displaystyle\-2\\Big\[\\Big\(I\+\(u\-t\)\\partial\_\{x\}\\tilde\{f\}^\{\*\}\\Big\)\(\\dot\{x\}\_\{t\}\-v\)\\Big\]^\{\\top\}\(u\-t\)\\partial\_\{x\}g\_\{\\eta\}\(\\dot\{x\}\_\{t\}\-v\)\\,\|\\,t<u\\Bigg\]There’s noη\\etain the first term in each of the two dot products that make up the expectand, so we can focus on the second term in the dot products, wheret<ut<uFor the second term in the first dot product:

limη→0\+\(\(u−t\)​∂tgη−gη\+\(u−t\)​\(∂xgη\)​x˙t\)\\displaystyle\\lim\_\{\\eta\\to 0^\{\+\}\}\\Big\(\(u\-t\)\\partial\_\{t\}g\_\{\\eta\}\-g\_\{\\eta\}\+\(u\-t\)\(\\partial\_\{x\}g\_\{\\eta\}\)\\dot\{x\}\_\{t\}\\Big\)=limη→0\+\(\(u−t\)​∂tgη−Iη​\(t,u\)​\(f~∗​\(t,t,x\)−v​\(t,x\)\)\+\(u−t\)​\(Iη​\(t,u\)​∂x\[\(f~∗​\(t,t,x\)−v​\(t,x\)\)\]\)​x˙t\)\\displaystyle=\\lim\_\{\\eta\\to 0^\{\+\}\}\\Big\(\(u\-t\)\\partial\_\{t\}g\_\{\\eta\}\-I\_\{\\eta\}\(t,u\)\\Big\(\\tilde\{f\}^\{\*\}\(t,t,x\)\-v\(t,x\)\\Big\)\+\(u\-t\)\(I\_\{\\eta\}\(t,u\)\\partial\_\{x\}\[\\Big\(\\tilde\{f\}^\{\*\}\(t,t,x\)\-v\(t,x\)\\Big\)\]\)\\dot\{x\}\_\{t\}\\Big\)=limη→0\+\(u−t\)​∂tgη\\displaystyle=\\lim\_\{\\eta\\to 0^\{\+\}\}\(u\-t\)\\partial\_\{t\}g\_\{\\eta\}=limη→0\+\(u−t\)​∂t\[Iη​\(t,u\)​\(f~∗​\(t,t,x\)−v​\(t,x\)\)\]\\displaystyle=\\lim\_\{\\eta\\to 0^\{\+\}\}\(u\-t\)\\partial\_\{t\}\[I\_\{\\eta\}\(t,u\)\\Big\(\\tilde\{f\}^\{\*\}\(t,t,x\)\-v\(t,x\)\\Big\)\]=limη→0\+\(u−t\)​\(∂tIη\)​g\+Iη​∂tg\\displaystyle=\\lim\_\{\\eta\\to 0^\{\+\}\}\(u\-t\)\(\\partial\_\{t\}I\_\{\\eta\}\)g\+I\_\{\\eta\}\\partial\_\{t\}g=limη→0\+\(u−t\)​\(∂tIη\)​Iη​\(t,u\)​\(f~∗​\(t,t,x\)−v​\(t,x\)\)\\displaystyle=\\lim\_\{\\eta\\to 0^\{\+\}\}\(u\-t\)\(\\partial\_\{t\}I\_\{\\eta\}\)I\_\{\\eta\}\(t,u\)\\Big\(\\tilde\{f\}^\{\*\}\(t,t,x\)\-v\(t,x\)\\Big\)=\(f~∗​\(t,t,x\)−v​\(t,x\)\)​\(u−t\)​limη→0\+\(∂tIη\)​Iη​\(t,u\)=0\\displaystyle=\\Big\(\\tilde\{f\}^\{\*\}\(t,t,x\)\-v\(t,x\)\\Big\)\(u\-t\)\\lim\_\{\\eta\\to 0^\{\+\}\}\(\\partial\_\{t\}I\_\{\\eta\}\)I\_\{\\eta\}\(t,u\)=0The last equality holds because the function and its time derivative both go to zero\.

This means the first product in the expectation goes to zero\. By a similar argument the second term goes to zero as well\.

Putting it all together

L:=limη→0\+s​\(η\)=limη→0\+𝔼​\[ϕη\]=limη→0\+p​\(t=u\)​𝔼​\[ϕη\|t=u\]\+p​\(t<u\)​𝔼​\[ϕη\|t<u\]\>0\\displaystyle L:=\\lim\_\{\\eta\\to 0^\{\+\}\}s\(\\eta\)=\\lim\_\{\\eta\\to 0^\{\+\}\}\\mathbb\{E\}\[\\phi\_\{\\eta\}\]=\\lim\_\{\\eta\\to 0^\{\+\}\}p\(t=u\)\\mathbb\{E\}\[\\phi\_\{\\eta\}\\,\|\\,t=u\]\+p\(t<u\)\\mathbb\{E\}\[\\phi\_\{\\eta\}\\,\|\\,t<u\]\>0Resultingly,

∃η0\>0​s\.t\.​∀η∗​s\.t\.​0<η∗<η0⟹\|s​\(η∗\)−L\|<ϵ\\displaystyle\\exists\\eta\_\{0\}\>0\\text\{ s\.t\. \}\\forall\\eta^\{\*\}\\text\{ s\.t\. \}0<\\eta^\{\*\}<\\eta\_\{0\}\\implies\|s\(\\eta^\{\*\}\)\-L\|<\\epsilon\(88\)If we pickϵ=0\.5​L\\epsilon=0\.5Lthen

\|s​\(η∗\)−L\|<\.5​L⟹s​\(η∗\)\>L−\.5​L=\.5​L\>0\\displaystyle\|s\(\\eta^\{\*\}\)\-L\|<\.5L\\implies s\(\\eta^\{\*\}\)\>L\-\.5L=\.5L\>0\(89\)sos​\(η∗\)\>0s\(\\eta^\{\*\}\)\>0\. But this contradicts being at a stationary point\. It cannot be thatf~∗​\(t,t,⋅\)≠v​\(t,⋅\)\\tilde\{f\}^\{\*\}\(t,t,\\cdot\)\\neq v\(t,\\cdot\)\. Therefore the velocity must match\.

∎

### A\.5Theorem 1

We present a proof about the functional stationary points of the SGFlow functional\. We use the definitions of first variation and stationary point from[section˜A\.1](https://arxiv.org/html/2607.26398#A1.SS1)and the definition of functional stopgrad from[section˜A\.2](https://arxiv.org/html/2607.26398#A1.SS2)\.

###### Theorem\.

Letq​\(t,u\)q\(t,u\)be a joint distribution over time pairs with support overt≤ut\\leq uand with positive probability ont=ut=u\. Let the familyℱ~\\mathcal\{\\tilde\{F\}\}include functionsf~\\tilde\{f\}that are continuously differentiable in all arguments\. Letxt=αt​x0\+σt​x1x\_\{t\}=\\alpha\_\{t\}x\_\{0\}\+\\sigma\_\{t\}x\_\{1\}andx˙t=α˙t​x0\+σ˙t​x1\\dot\{x\}\_\{t\}=\\dot\{\\alpha\}\_\{t\}x\_\{0\}\+\\dot\{\\sigma\}\_\{t\}x\_\{1\}\. Definef​\(t,u,x\):=x\+\(u−t\)​f~​\(t,u,x\)f\(t,u,x\):=x\+\(u\-t\)\\tilde\{f\}\(t,u,x\)\. Let expectations be computed overq​\(X0\)​q​\(X1\)q\(X\_\{0\}\)q\(X\_\{1\}\)\. Lets​gsgstand for stop\-gradient\. Define:

ℒ​\[f~\]\\displaystyle\\mathcal\{L\}\[\\tilde\{f\}\]:=𝔼\[∥\(∂tf\)\|\(t,u,xt\)\+\(∂xf\)\|\(t,u,xt\)x˙t∥2−∥\(∂xf\)\(t,u,xt\)\(x˙t−𝔼\[x˙t\|xt\]\)∥2\]\\displaystyle:=\\mathbb\{E\}\\Bigg\[\\\|\(\\partial\_\{t\}f\)\|\_\{\(t,u,x\_\{t\}\)\}\+\(\\partial\_\{x\}f\)\|\_\{\(t,u,x\_\{t\}\)\}\\dot\{x\}\_\{t\}\\\|^\{2\}\-\\\|\(\\partial\_\{x\}f\)\_\{\(t,u,x\_\{t\}\)\}\(\\dot\{x\}\_\{t\}\-\\mathbb\{E\}\[\\dot\{x\}\_\{t\}\|x\_\{t\}\]\)\\\|^\{2\}\\Bigg\]ℒsg​\[f~\]\\displaystyle\\mathcal\{L\}^\{\\text\{sg\}\}\[\\tilde\{f\}\]:=𝔼​\[‖\(∂tf\)\(t,u,xt\)\+\(∂xf\)\(t,u,xt\)​x˙t‖2−‖\(∂xf\)\(t,u,xt\)​\(x˙t\+sg​\[\(∂tf\)\]\(t,t,xt\)\)‖2\]\\displaystyle:=\\mathbb\{E\}\\Bigg\[\\\|\(\\partial\_\{t\}f\)\_\{\(t,u,x\_\{t\}\)\}\+\(\\partial\_\{x\}f\)\_\{\(t,u,x\_\{t\}\)\}\\dot\{x\}\_\{t\}\\\|^\{2\}\-\\\|\(\\partial\_\{x\}f\)\_\{\(t,u,x\_\{t\}\)\}\(\\dot\{x\}\_\{t\}\+\\text\{sg\}\[\(\\partial\_\{t\}f\)\]\_\{\(t,t,x\_\{t\}\)\}\)\\\|^\{2\}\\Bigg\]Thenf~∗\\tilde\{f\}^\{\*\}is a stationary point ofℒsg\\mathcal\{L\}^\{\\text\{sg\}\}with respect toℱ~\\mathcal\{\\tilde\{F\}\}if and only iff~∗\\tilde\{f\}^\{\*\}is a stationary point ofℒ\\mathcal\{L\}with respect toℱ~\\mathcal\{\\tilde\{F\}\}\.

###### Proof\.

Case 1: Iff~∗\\tilde\{f\}^\{\*\}is a stationary point ofℒ\\mathcal\{L\}, thenf~∗\\tilde\{f\}^\{\*\}is a stationary point ofℒsg\\mathcal\{L\}^\{\\text\{sg\}\}\.

- •Sincef~∗\\tilde\{f\}^\{\*\}is a stationary point,δ​ℒ​\[f~∗;⋅\]=0\\delta\\mathcal\{L\}\[\\tilde\{f\}^\{\*\};\\cdot\]=0
- •by[section˜A\.4](https://arxiv.org/html/2607.26398#A1.SS4), we have that∂tf∗​\(t,t,x\)=−f~∗​\(t,t,x\)=−𝔼​\[x˙t\|xt\]\\partial\_\{t\}f^\{\*\}\(t,t,x\)=\-\\tilde\{f\}^\{\*\}\(t,t,x\)=\-\\mathbb\{E\}\[\\dot\{x\}\_\{t\}\|x\_\{t\}\]wheref∗=x\+\(u−t\)​f~∗f^\{\*\}=x\+\(u\-t\)\\tilde\{f\}^\{\*\}
- •ℒsg=ℒ\\mathcal\{L\}^\{\\text\{sg\}\}=\\mathcal\{L\}
- •Sinceℒsg=ℒ\\mathcal\{L\}^\{\\text\{sg\}\}=\\mathcal\{L\}, thenδ​ℒsg​\[f~∗;⋅\]=δ​ℒ​\[f~∗;⋅\]=0\\delta\\mathcal\{L\}^\{\\text\{sg\}\}\[\\tilde\{f\}^\{\*\};\\cdot\]=\\delta\\mathcal\{L\}\[\\tilde\{f\}^\{\*\};\\cdot\]=0

Case 2: Iff~∗\\tilde\{f\}^\{\*\}is not a stationary point ofℒ\\mathcal\{L\}, thenf~∗\\tilde\{f\}^\{\*\}is not a stationary point ofℒsg\\mathcal\{L\}^\{\\text\{sg\}\}\.

Sincef~∗\\tilde\{f\}^\{\*\}is not a stationary point ofℒ\\mathcal\{L\}, then∃h\\exists hthat is admissible \(continuous\) such thatδ​ℒ​\[f~∗;h\]≠0\\delta\\mathcal\{L\}\[\\tilde\{f\}^\{\*\};h\]\\neq 0\. Then,

δ​ℒ​\[f~∗;h\]⏟LHS\\displaystyle\\underbrace\{\\delta\\mathcal\{L\}\[\\tilde\{f\}^\{\*\};h\]\}\_\{\\text\{LHS\}\}=𝔼​\[1​\[t=u\]​…\]⏟RHS\-L\+𝔼​\[1​\[t≠u\]​…\]⏟RHS\-R\\displaystyle=\\underbrace\{\\mathbb\{E\}\\Big\[1\[t=u\]\\ldots\\Big\]\}\_\{\\text\{RHS\-L\}\}\+\\underbrace\{\\mathbb\{E\}\\Big\[1\[t\\neq u\]\\ldots\\Big\]\}\_\{\\text\{RHS\-R\}\}\(90\)If the LHS is nonzero, then one of RHS\-L or RHS\-R is nonzero\. Consider both cases\.

Case 2a: The RHS\-R is nonzero and RHS\-L is zero\.RHS\-L being zero means that the velocity matches, which means that RHS\-R has the same first variation betweenℒ\\mathcal\{L\}andℒsg\\mathcal\{L\}^\{\\text\{sg\}\}\. So they must coincide regarding stationary points\.

Case 2b: RHS\-L is nonzero and RHS\-R is either zero or nonzero\.

Case 2b\-i\.δ​ℒsg​\[f~∗;h\]≠0\\delta\\mathcal\{L\}^\{\\text\{sg\}\}\[\\tilde\{f\}^\{\*\};h\]\\neq 0directly holds\. This is all we are trying to ensure anyway, so we are done in this case\.

Case 2b\-ii\.Define the soft indicator,IηI\_\{\\eta\}:

Iη​\(t,u\)\\displaystyle I\_\{\\eta\}\(t,u\)=1​\[η\>0\]​2​\(1−11\+exp⁡\(−1η2​\(t−u\)2\)\)\+1​\[η=0\]​1​\[t=u\]\.\\displaystyle=1\[\\eta\>0\]2\\Big\(1\-\\frac\{1\}\{1\+\\exp\(\-\\frac\{1\}\{\\eta^\{2\}\}\(t\-u\)^\{2\}\)\}\\Big\)\+1\[\\eta=0\]1\[t=u\]\.\(91\)Define the direction,h^η\\hat\{h\}\_\{\\eta\}:

h^η​\(t,u,x\)=Iη​\(t,u\)​h​\(t,u,x\)\.\\displaystyle\\hat\{h\}\_\{\\eta\}\(t,u,x\)=I\_\{\\eta\}\(t,u\)h\(t,u,x\)\.\(92\)h^η\\hat\{h\}\_\{\\eta\}is continuously differentiable for anyη\>0\\eta\>0\. This is true because it is a product of two functions that are each continuously differentiable \(hhis assumed continuously differentiable\)\. Recall the mappings​\(η\)s\(\\eta\)from[Section˜A\.4](https://arxiv.org/html/2607.26398#A1.SS4), that mapsη\\etatoδ​ℒ​\[f~∗;h^η\]\\delta\\mathcal\{L\}\[\\tilde\{f\}^\{\*\};\\hat\{h\}\_\{\\eta\}\]\. Under the conditions of dominated convergence established in[Section˜A\.4](https://arxiv.org/html/2607.26398#A1.SS4), we know that∃η∗\>0\\exists\\eta^\{\*\}\>0such thath^η∗\\hat\{h\}\_\{\\eta^\{\*\}\}is a continuously differentiable function for whichδ​ℒ​\[f~∗;h^η∗\]≠0\\delta\\mathcal\{L\}\[\\tilde\{f\}^\{\*\};\\hat\{h\}\_\{\\eta^\{\*\}\}\]\\neq 0\. This must mean that the velocity is not matched\. But we know that if the velocity does not match, we are not at a stationary point ofℒsg\\mathcal\{L\}^\{\\text\{sg\}\}either, sinceℒsg\\mathcal\{L\}^\{\\text\{sg\}\}andℒ\\mathcal\{L\}coincide on penalizing velocity matching ont=ut=u\. ∎

## Appendix BOther useful derivations and results

### B\.1Gradient updates are not optimization of one scalar objective via gradients

We show here that there exists a data distribution and a model such that the SGFlow updates are not the gradient of any single scalar objective\. We illustrate this by considering an simple 1D setting\. The key point is that, even in this restricted case, the update field induced by the stopgrad operator has non\-zero curl and therefore cannot be written as the gradient of any scalar objectiveJ​\(θ\)J\(\\theta\)\. The form to be differentiated is:

Lsg​\(θ\)=𝔼Xt​\[‖∂tfθ​\(t,u,Xt\)\+\(∂xfθ​\(t,u,Xt\)\)​X˙t‖2\]−𝔼Xt​\[‖\(∂xfθ​\(t,u,Xt\)\)​\(X˙t−stopgrad⁡\[f~θ​\(t,t,Xt\)\]\)‖2\]\.\\displaystyle\\begin\{split\}L\_\{\\text\{sg\}\}\(\\theta\)&=\\mathbb\{E\}\_\{X\_\{t\}\}\\\!\\left\[\\big\\\|\\partial\_\{t\}f\_\{\\theta\}\(t,u,X\_\{t\}\)\+\(\\partial\_\{x\}f\_\{\\theta\}\(t,u,X\_\{t\}\)\)\\,\\dot\{X\}\_\{t\}\\big\\\|^\{2\}\\right\]\\\\ &\\quad\-\\mathbb\{E\}\_\{X\_\{t\}\}\\\!\\left\[\\big\\\|\(\\partial\_\{x\}f\_\{\\theta\}\(t,u,X\_\{t\}\)\)\\big\(\\dot\{X\}\_\{t\}\-\\operatorname\{stopgrad\}\[\\tilde\{f\}\_\{\\theta\}\(t,t,X\_\{t\}\)\]\\big\)\\big\\\|^\{2\}\\right\]\.\\end\{split\}\(93\)forfθ​\(t,u,x\)=x\+\(u−t\)​f~θ​\(t,u,x\)f\_\{\\theta\}\(t,u,x\)=x\+\(u\-t\)\\tilde\{f\}\_\{\\theta\}\(t,u,x\)\. Let us work in 1D and fix values ofXt=xX\_\{t\}=xandX˙t=d\\dot\{X\}\_\{t\}=d, a constant\. To accomplish this, we can chooseX1X\_\{1\}freely and setX0=X1−dX\_\{0\}=X\_\{1\}\-d, so that the interpolation satisfies bothXt=xX\_\{t\}=xandX˙t=d\\dot\{X\}\_\{t\}=d\. In this case the expectations in \([93](https://arxiv.org/html/2607.26398#A2.E93)\) collapse to evaluation at this point \(equivalently, think of us approximating with 1 Monte Carlo sample\)\. Now, consider the parametersθ=\(θ1,θ2\)⊤\\theta=\(\\theta\_\{1\},\\theta\_\{2\}\)^\{\\top\}and a single time pair\(t,u\)\(t,u\)such that at the positionXt=xX\_\{t\}=x,

∂tfθ​\(t,u,x\)=0,∂xfθ​\(t,u,x\)=θ1,f~θ​\(t,t,x\)=θ2\.\\partial\_\{t\}f\_\{\\theta\}\(t,u,x\)=0,\\qquad\\partial\_\{x\}f\_\{\\theta\}\(t,u,x\)=\\theta\_\{1\},\\qquad\\tilde\{f\}\_\{\\theta\}\(t,t,x\)=\\theta\_\{2\}\.\(94\)\(For a sufficiently expressive model, such local values can be realized; we only need existence of such a configuration\.\)\. Plugging these into \([93](https://arxiv.org/html/2607.26398#A2.E93)\) and dropping the expectation \(single point\), the SGFlow functional is:

Lsg=\(θ1​d\)2−\(θ1​\(d−stopgrad​\[θ2\]\)\)2\.L\_\{\\text\{sg\}\}=\(\\theta\_\{1\}d\)^\{2\}\-\\big\(\\theta\_\{1\}\\big\(d\-\\text\{stopgrad\}\[\\theta\_\{2\}\]\\big\)\\big\)^\{2\}\.\(95\)
Let∇~\\tilde\{\\nabla\}denote differentiation with the stopgrad applied toθ2\\theta\_\{2\}in the second term\. Differentiating \([95](https://arxiv.org/html/2607.26398#A2.E95)\) with respect toθ1\\theta\_\{1\}yields

∇~θ1​Lsg\\displaystyle\\tilde\{\\nabla\}\_\{\\theta\_\{1\}\}L\_\{\\text\{sg\}\}=2​θ1​d2−2​θ1​\(d−θ2\)2\\displaystyle=2\\theta\_\{1\}d^\{2\}\-2\\theta\_\{1\}\\big\(d\-\\theta\_\{2\}\\big\)^\{2\}\(96\)=2​θ1​\(d2−\(d−θ2\)2\)\\displaystyle=2\\theta\_\{1\}\\Big\(d^\{2\}\-\(d\-\\theta\_\{2\}\)^\{2\}\\Big\)\(97\)=2​θ1​\(d2−\(d2−2​d​θ2\+θ22\)\)\\displaystyle=2\\theta\_\{1\}\\Big\(d^\{2\}\-\(d^\{2\}\-2d\\theta\_\{2\}\+\\theta\_\{2\}^\{2\}\)\\Big\)\(98\)=2​θ1​\(2​d​θ2−θ22\)\\displaystyle=2\\theta\_\{1\}\\big\(2d\\theta\_\{2\}\-\\theta\_\{2\}^\{2\}\\big\)\(99\)=2​θ1​θ2​\(2​d−θ2\)\.\\displaystyle=2\\theta\_\{1\}\\theta\_\{2\}\(2d\-\\theta\_\{2\}\)\.\(100\)Here the stopgrad onθ2\\theta\_\{2\}only affects the second term and does not change the first term\. Forθ2\\theta\_\{2\}, all occurrences appear inside a stopgrad and the first term does not depend onθ2\\theta\_\{2\}, hence

∇~θ2​Lsg=0\.\\tilde\{\\nabla\}\_\{\\theta\_\{2\}\}L\_\{\\text\{sg\}\}=0\.\(101\)
Thus the update field induced by SGFlow in this simple example is

g​\(θ1,θ2\):=\(g1​\(θ1,θ2\),g2​\(θ1,θ2\)\)=\(2​θ1​θ2​\(2​d−θ2\),0\)\.g\(\\theta\_\{1\},\\theta\_\{2\}\):=\\big\(g\_\{1\}\(\\theta\_\{1\},\\theta\_\{2\}\),g\_\{2\}\(\\theta\_\{1\},\\theta\_\{2\}\)\\big\)=\\big\(2\\theta\_\{1\}\\theta\_\{2\}\(2d\-\\theta\_\{2\}\),\\ 0\\big\)\.\(102\)
If this update were the gradient of some scalar objectiveJ​\(θ1,θ2\)J\(\\theta\_\{1\},\\theta\_\{2\}\), then the mixed partial derivatives would commute:

g1=∂θ1J,g2=∂θ2J⇒∂θ2g1=∂θ1g2\.g\_\{1\}=\\partial\_\{\\theta\_\{1\}\}J,\\qquad g\_\{2\}=\\partial\_\{\\theta\_\{2\}\}J\\quad\\Rightarrow\\quad\\partial\_\{\\theta\_\{2\}\}g\_\{1\}=\\partial\_\{\\theta\_\{1\}\}g\_\{2\}\.\(103\)However,

∂g1∂θ2\\displaystyle\\frac\{\\partial g\_\{1\}\}\{\\partial\\theta\_\{2\}\}=∂∂θ2​\(2​θ1​θ2​\(2​d−θ2\)\)=2​θ1​\(2​d−θ2−θ2\)=4​θ1​\(d−θ2\),\\displaystyle=\\frac\{\\partial\}\{\\partial\\theta\_\{2\}\}\\Big\(2\\theta\_\{1\}\\theta\_\{2\}\(2d\-\\theta\_\{2\}\)\\Big\)=2\\theta\_\{1\}\\big\(2d\-\\theta\_\{2\}\-\\theta\_\{2\}\\big\)=4\\theta\_\{1\}\(d\-\\theta\_\{2\}\),\(104\)∂g2∂θ1\\displaystyle\\frac\{\\partial g\_\{2\}\}\{\\partial\\theta\_\{1\}\}=0,\\displaystyle=0,\(105\)and hence the curl of the update field is

∂g1∂θ2−∂g2∂θ1=4​θ1​\(d−θ2\),\\frac\{\\partial g\_\{1\}\}\{\\partial\\theta\_\{2\}\}\-\\frac\{\\partial g\_\{2\}\}\{\\partial\\theta\_\{1\}\}=4\\theta\_\{1\}\(d\-\\theta\_\{2\}\),\(106\)which is non\-zero for genericθ\\theta\(for example, wheneverθ1≠0\\theta\_\{1\}\\neq 0andθ2≠d\\theta\_\{2\}\\neq d\)\. Thereforeggis a smooth vector field with non\-zero curl and*cannot*be written as the gradient of any scalar objectiveJ​\(θ1,θ2\)J\(\\theta\_\{1\},\\theta\_\{2\}\)in general\. Put differently, a constraint on the relationship between the parameters of the model and the value of the chosen datapoint is necessary to ensure zero curl\.

This 1D example is a specific instantiation of SGFlow \([93](https://arxiv.org/html/2607.26398#A2.E93)\) with a simple model and a single training point\. It shows that, once we introduce the stopgrad onf~θ​\(t,t,x\)\\tilde\{f\}\_\{\\theta\}\(t,t,x\), the resulting optimization dynamics are in general*non\-conservative*\. The stopgrad structure breaks the symmetry required for the updates to be the gradient of a single scalar function\. In this sense, SGFlow is formally a \(two\-player\)*game*rather than standard gradient descent on one potential function—or, if one prefers, a nongradient vector flow\.

### B\.2Gaussian counterexample: Meanflow without stopgrads

We compute the first variation of the no\-stopgrad Meanflow functionalℒnosg\\mathcal\{L\}\_\{\\mathrm\{nosg\}\}at the true flow map for the 1D Gaussian caseX0∼𝒩​\(0,1\)X\_\{0\}\\sim\\mathcal\{N\}\(0,1\),X1∼𝒩​\(μ,λ2\)X\_\{1\}\\sim\\mathcal\{N\}\(\\mu,\\lambda^\{2\}\),Xt=\(1−t\)​X0\+t​X1X\_\{t\}=\(1\-t\)X\_\{0\}\+tX\_\{1\},X˙t=X1−X0\\dot\{X\}\_\{t\}=X\_\{1\}\-X\_\{0\}\.

#### Closed\-form Gaussian quantities\.

- •Marginal variance:σt2=\(1−t\)2\+t2​λ2\\sigma\_\{t\}^\{2\}=\(1\-t\)^\{2\}\+t^\{2\}\\lambda^\{2\}\.
- •Velocity:v​\(t,x\)=α​\(t\)​\(x−t​μ\)\+μv\(t,x\)=\\alpha\(t\)\(x\-t\\mu\)\+\\muwhereα​\(t\)=t​λ2−\(1−t\)σt2\\alpha\(t\)=\\frac\{t\\lambda^\{2\}\-\(1\-t\)\}\{\\sigma\_\{t\}^\{2\}\}\.
- •Flow map:f∗​\(t,u,x\)=m​\(t,u\)​\(x−t​μ\)\+u​μf^\{\*\}\(t,u,x\)=m\(t,u\)\(x\-t\\mu\)\+u\\muwherem​\(t,u\)=σu/σtm\(t,u\)=\\sigma\_\{u\}/\\sigma\_\{t\}\.
- •Jacobian:∂xf∗​\(t,u,x\)=m​\(t,u\)\\partial\_\{x\}f^\{\*\}\(t,u,x\)=m\(t,u\)\.

#### Posterior variance\.

We computeV​\(t\):=𝕍​ar​\(X˙t\|Xt\)V\(t\):=\\mathbb\{V\}\\textrm\{ar\}\(\\dot\{X\}\_\{t\}~\|~X\_\{t\}\)\. SinceX˙t=X1−X0\\dot\{X\}\_\{t\}=X\_\{1\}\-X\_\{0\}andXt=\(1−t\)​X0\+t​X1X\_\{t\}=\(1\-t\)X\_\{0\}\+tX\_\{1\}withX0⟂X1X\_\{0\}\\perp X\_\{1\}:

𝕍​ar​\(X˙t\)\\displaystyle\\mathbb\{V\}\\textrm\{ar\}\(\\dot\{X\}\_\{t\}\)=1\+λ2,\\displaystyle=1\+\\lambda^\{2\},\(107\)ℂ​ov​\(X˙t,Xt\)\\displaystyle\\mathbb\{C\}\\textrm\{ov\}\(\\dot\{X\}\_\{t\},X\_\{t\}\)=t​λ2−\(1−t\),\\displaystyle=t\\lambda^\{2\}\-\(1\-t\),\(108\)𝕍​ar​\(Xt\)\\displaystyle\\mathbb\{V\}\\textrm\{ar\}\(X\_\{t\}\)=σt2\.\\displaystyle=\\sigma\_\{t\}^\{2\}\.\(109\)By the Gaussian conditional variance formula:

V​\(t\)=𝕍​ar​\(X˙t\)−ℂ​ov​\(X˙t,Xt\)2𝕍​ar​\(Xt\)=\(1\+λ2\)−\(t​λ2−\(1−t\)\)2σt2=λ2σt2\.\\displaystyle V\(t\)=\\mathbb\{V\}\\textrm\{ar\}\(\\dot\{X\}\_\{t\}\)\-\\frac\{\\mathbb\{C\}\\textrm\{ov\}\(\\dot\{X\}\_\{t\},X\_\{t\}\)^\{2\}\}\{\\mathbb\{V\}\\textrm\{ar\}\(X\_\{t\}\)\}=\(1\+\\lambda^\{2\}\)\-\\frac\{\(t\\lambda^\{2\}\-\(1\-t\)\)^\{2\}\}\{\\sigma\_\{t\}^\{2\}\}=\\frac\{\\lambda^\{2\}\}\{\\sigma\_\{t\}^\{2\}\}\.\(110\)\(The last equality follows from expanding:\(1\+λ2\)​σt2−\(t​λ2−\(1−t\)\)2=λ2\(1\+\\lambda^\{2\}\)\\sigma\_\{t\}^\{2\}\-\(t\\lambda^\{2\}\-\(1\-t\)\)^\{2\}=\\lambda^\{2\}\.\) In particular,V​\(t\)\>0V\(t\)\>0for allt∈\[0,1\]t\\in\[0,1\]whenλ\>0\\lambda\>0\.

#### First variation ofℒnosg\\mathcal\{L\}\_\{\\mathrm\{nosg\}\}at the true flow map\.

Letf~ϵ=f~\+ϵ​h\\tilde\{f\}^\{\\epsilon\}=\\tilde\{f\}\+\\epsilon h\. The quantity inside the norm ofℒnosg\\mathcal\{L\}\_\{\\mathrm\{nosg\}\}isR​\(ϵ\)=f~ϵ−X˙t−\(u−t\)​\(∂tf~ϵ\+∂xf~ϵ​X˙t\)R\(\\epsilon\)=\\tilde\{f\}^\{\\epsilon\}\-\\dot\{X\}\_\{t\}\-\(u\-t\)\(\\partial\_\{t\}\\tilde\{f\}^\{\\epsilon\}\+\\partial\_\{x\}\\tilde\{f\}^\{\\epsilon\}\\,\\dot\{X\}\_\{t\}\), with derivativeR′​\(ϵ\)=h−\(u−t\)​\(∂th\+∂xh​X˙t\)R^\{\\prime\}\(\\epsilon\)=h\-\(u\-t\)\(\\partial\_\{t\}h\+\\partial\_\{x\}h\\,\\dot\{X\}\_\{t\}\)\. The first variation isδ​ℒ=2​𝔼\[R​\(0\)⋅R′​\(0\)\]\\delta\\mathcal\{L\}=2\\,\\mathop\{\\mathbb\{E\}\}\[R\(0\)\\cdot R^\{\\prime\}\(0\)\]\.

WriteX˙t=v\+ξ\\dot\{X\}\_\{t\}=v\+\\xiwhereξ=X˙t−v\\xi=\\dot\{X\}\_\{t\}\-vsatisfies𝔼\[ξ\|Xt\]=0\\mathop\{\\mathbb\{E\}\}\[\\xi~\|~X\_\{t\}\]=0\. At the true flow map, the transport equation kills thevv\-part ofR​\(0\)R\(0\), leavingR​\(0\)=−∂xf∗​ξR\(0\)=\-\\partial\_\{x\}f^\{\*\}\\,\\xiwhere∂xf∗=1\+\(u−t\)​∂xf~∗\\partial\_\{x\}f^\{\*\}=1\+\(u\-t\)\\partial\_\{x\}\\tilde\{f\}^\{\*\}\. Similarly,R′​\(0\)=ϕ−\(u−t\)​∂xh​ξR^\{\\prime\}\(0\)=\\phi\-\(u\-t\)\\partial\_\{x\}h\\,\\xiwhereϕ=h−\(u−t\)​\(∂th\+∂xh​v\)\\phi=h\-\(u\-t\)\(\\partial\_\{t\}h\+\\partial\_\{x\}h\\,v\)\.

Computing𝔼\[R​\(0\)⋅R′​\(0\)\]\\mathop\{\\mathbb\{E\}\}\[R\(0\)\\cdot R^\{\\prime\}\(0\)\]: the cross\-term𝔼\[∂xf∗​ξ​ϕ\]=0\\mathop\{\\mathbb\{E\}\}\[\\partial\_\{x\}f^\{\*\}\\,\\xi\\,\\phi\]=0by the tower property \(condition on\(t,u,Xt\)\(t,u,X\_\{t\}\);𝔼\[ξ\|Xt\]=0\\mathop\{\\mathbb\{E\}\}\[\\xi~\|~X\_\{t\}\]=0\)\. The remaining term gives

δ​ℒnosg​\[f~∗;h\]=2​𝔼\[\(u−t\)​∂xf∗​V​\(t,Xt\)​∂xh\]\.\\displaystyle\\delta\\mathcal\{L\}\_\{\\mathrm\{nosg\}\}\[\\tilde\{f\}^\{\*\};h\]=2\\,\\mathop\{\\mathbb\{E\}\}\\big\[\(u\-t\)\\,\\partial\_\{x\}f^\{\*\}\\,V\(t,X\_\{t\}\)\\,\\partial\_\{x\}h\\big\]\.\(111\)

#### Evaluating in the Gaussian case\.

Substituting∂xf∗=m​\(t,u\)\\partial\_\{x\}f^\{\*\}=m\(t,u\)andV​\(t\)=λ2/σt2V\(t\)=\\lambda^\{2\}/\\sigma\_\{t\}^\{2\}, and choosingh​\(t,u,x\)=xh\(t,u,x\)=x\(so∂xh=1\\partial\_\{x\}h=1\):

δ​ℒnosg​\[f~∗;h=x\]=2​λ2​𝔼t,u\[\(u−t\)​m​\(t,u\)σt2\]\.\\displaystyle\\delta\\mathcal\{L\}\_\{\\mathrm\{nosg\}\}\[\\tilde\{f\}^\{\*\};\\,h=x\]=2\\lambda^\{2\}\\,\\mathop\{\\mathbb\{E\}\}\_\{t,u\}\\\!\\left\[\\frac\{\(u\-t\)\\,m\(t,u\)\}\{\\sigma\_\{t\}^\{2\}\}\\right\]\.\(112\)Sinceλ\>0\\lambda\>0,\(u−t\)\>0\(u\-t\)\>0fort<ut<u,m​\(t,u\)=σu/σt\>0m\(t,u\)=\\sigma\_\{u\}/\\sigma\_\{t\}\>0, andσt\>0\\sigma\_\{t\}\>0, every factor in the integrand is strictly positive\. Thereforeδ​ℒnosg​\[f~∗;h=x\]\>0\\delta\\mathcal\{L\}\_\{\\mathrm\{nosg\}\}\[\\tilde\{f\}^\{\*\};\\,h=x\]\>0: the first variation is nonzero, sof~∗\\tilde\{f\}^\{\*\}is*not*a stationary point ofℒnosg\\mathcal\{L\}\_\{\\mathrm\{nosg\}\}\.

### B\.3Differentiating w\.r\.t\.ttversusuu

In this work,f​\(t,u,x\)f\(t,u,x\)maps forward in time fromXt=xX\_\{t\}=xtoXuX\_\{u\}along the flowd​Xs=v​\(s,Xs\)dX\_\{s\}=v\(s,X\_\{s\}\)\. In Meanflow\(Genget al\.,[2025](https://arxiv.org/html/2607.26398#bib.bib36)\), the map is written in reverse time\. The transport PDE can be obtained by differentiating the forward\-time flow map with respect tott, or the reverse\-time flow map with respect touu\. In both cases, lett≤ut\\leq u:

ffwd​\(t,u,x\)=x\+∫tuv​\(s,Xs\)​𝑑s=x\+∫tuv​\(s,ffwd​\(t,s,x\)\)​𝑑s\\displaystyle f\_\{\\text\{fwd\}\}\(t,u,x\)=x\+\\int\_\{t\}^\{u\}v\(s,X\_\{s\}\)ds=x\+\\int\_\{t\}^\{u\}v\(s,f\_\{\\text\{fwd\}\}\(t,s,x\)\)ds\(113\)On the other hand, if moving in reverse time:

frev​\(t,u,x\)=x\+∫utv​\(s,Xs\)​𝑑s=x−∫tuv​\(s,frev​\(s,u,x\)\)​𝑑s\\displaystyle f\_\{\\text\{rev\}\}\(t,u,x\)=x\+\\int\_\{u\}^\{t\}v\(s,X\_\{s\}\)ds=x\-\\int\_\{t\}^\{u\}v\(s,f\_\{\\text\{rev\}\}\(s,u,x\)\)ds\(114\)Differentiating via the Leibniz rule yields:

∂tffwd​\(t,u,x\)\+\(∂xffwd\)​\(t,u,x\)​v​\(t,x\)\\displaystyle\\partial\_\{t\}f\_\{\\text\{fwd\}\}\(t,u,x\)\+\(\\partial\_\{x\}f\_\{\\text\{fwd\}\}\)\(t,u,x\)v\(t,x\)=0\\displaystyle=0\(115\)∂ufrev​\(t,u,x\)\+\(∂xfrev\)​\(t,u,x\)​v​\(u,x\)\\displaystyle\\partial\_\{u\}f\_\{\\text\{rev\}\}\(t,u,x\)\+\(\\partial\_\{x\}f\_\{\\text\{rev\}\}\)\(t,u,x\)v\(u,x\)=0\\displaystyle=0\(116\)

Similar Articles

Modeling Unknown Nonlocal PDE Systems via Flow Map Learning

arXiv cs.LG

This paper presents a flow-map learning framework for modeling unknown nonlocal PDEs directly from solution data, avoiding explicit nonlocal operator evaluation. The method learns finite-time evolution operators in modal or nodal space and demonstrates accurate long-time prediction for fractional diffusion and wave equations.

Reinforcement Learning via Value Gradient Flow

Hugging Face Daily Papers

Value Gradient Flow (VGF) presents a scalable approach to behavior-regularized reinforcement learning by formulating it as an optimal transport problem solved through discrete gradient flow, achieving state-of-the-art results on offline RL and LLM RL benchmarks. The method eliminates explicit policy parameterization while enabling adaptive test-time scaling by controlling transport budget.