Potential Matching Optimal Transport: Continuous Normalizing Flows for Exact $p$-Wasserstein Dynamics
Summary
This paper introduces PMOT, a potential-flow framework for general p-cost optimal transport using continuous normalizing flows, with theoretical zero-loss exactness and promising results on synthetic and high-dimensional benchmarks.
View Cached Full Text
Cached at: 08/07/26, 07:51 AM
# Continuous Normalizing Flows for Exact 𝑝-Wasserstein Dynamics
Source: [https://arxiv.org/html/2608.05666](https://arxiv.org/html/2608.05666)
## Potential Matching Optimal Transport: Continuous Normalizing Flows for Exactpp\-Wasserstein Dynamics
Lishuo Zhang1Ruizhi Huang1Yang Yu1Lei Li1,2 1School of Mathematical Sciences, Shanghai Jiao Tong University 2Institute of Natural Sciences, MOE\-LSC, Shanghai Jiao Tong University Shanghai 200240, China ShawnLi9@sjtu\.edu\.cnhoperealizer@163\.com yuyang2357@sjtu\.edu\.cnleili2010@sjtu\.edu\.cn
###### Abstract
We introduce Potential Matching Optimal Transport \(PMOT\), a potential\-flow framework for generalpp\-cost optimal transport withcp\(x,y\)=‖x−y‖pc\_\{p\}\(x,y\)=\\\|x\-y\\\|^\{p\}\. PMOT parameterizes the CNF velocity field with a scalar potential in the generalized Benamou–Brenier form for the chosen exponentpp\. It trains the potential gradient with a self\-induced matching loss along straight bridges determined by the model’s own endpoints, while allowing flexible terminal distribution matching\. Our main result establishes zero\-loss exactness: under the stated regularity, exact terminal matching, and uniqueness assumptions, any zero\-loss solution satisfies the generalized Benamou–Brenier optimality system and recovers the correspondingpp\-optimal transport map and dynamics\. On synthetic benchmarks, PMOT learnspp\-specific maps that agree with the correspondingpp\-matched OT references\. It also remains competitive as a likelihood\-based density model on high\-dimensional tabular data, and an MMD\-based color transformation experiment demonstrates flexible sample\-based terminal matching\.
## 1Introduction
Continuous Normalizing Flows \(CNFs\) learn invertible transformations between data distributions and tractable reference priors through time\-dependent ODE dynamics\(Chenet al\.,[2018](https://arxiv.org/html/2608.05666#bib.bib7); Grathwohlet al\.,[2018](https://arxiv.org/html/2608.05666#bib.bib8)\)\. They support exact likelihood evaluation via the continuous change\-of\-variables formula while simultaneously defining a transport trajectory between distributions\. This makes CNFs a natural interface between likelihood\-based generative modeling and optimal transport \(OT\)\.
A particularly influential line of work connects CNFs with OT by regularizing the learned dynamics toward low\-cost transport paths\. OT\-Flow\(Onkenet al\.,[2021](https://arxiv.org/html/2608.05666#bib.bib9)\), for example, introduces a potential\-flow parameterization together with a kinetic\-energy regularizer, yielding a likelihood\-based model whose dynamics are closely related to the Benamou–Brenier formulation of quadratic\-cost optimal transport\(Benamou and Brenier,[2000](https://arxiv.org/html/2608.05666#bib.bib1)\)\. However, the resulting geometry is naturally tied to the quadratic cost, or equivalently theW2W\_\{2\}geometry\. Many OT problems are instead defined by more general costs
cp\(x,y\)=‖x−y‖p,c\_\{p\}\(x,y\)=\\\|x\-y\\\|^\{p\},where the exponentppcontrols the induced transport geometry and, especially in asymmetric transport problems, can lead to different optimal couplings and trajectories\(Villani and others,[2009](https://arxiv.org/html/2608.05666#bib.bib2); Peyré and Cuturi,[2019](https://arxiv.org/html/2608.05666#bib.bib3)\)\. This motivates likelihood\-based potential flows whose dynamics are aligned not only with the quadratic case, but also with generalpp\-cost OT\.
In this work, we introduce*Potential Matching Optimal Transport*\(PMOT\), a potential\-flow framework for generalpp\-cost optimal transport with terminal distribution matching\. PMOT preserves the main structural advantages of OT\-Flow\-style CNFs when instantiated with a KL/NLL terminal loss: the dynamics remain continuous, and likelihoods are evaluated through the CNF change\-of\-variables formula\. PMOT parameterizes the velocity field in the generalized Benamou–Brenier form with conjugate exponentq=p/\(p−1\)q=p/\(p\-1\), thereby making the dynamics depend on the targetpp\-cost geometry\. Rather than using a quadratic kinetic\-energy penalty, PMOT introduces a self\-induced potential matching residual\. For each data samplexx, the current CNF produces a terminal endpointFθ\(x\)F\_\{\\theta\}\(x\), and the residual matches the potential gradient along the induced straight bridge to
−‖Fθ\(x\)−x‖p−2\(Fθ\(x\)−x\)\.\-\\\|F\_\{\\theta\}\(x\)\-x\\\|^\{p\-2\}\\bigl\(F\_\{\\theta\}\(x\)\-x\\bigr\)\.This condition is equivalent to matching the induced velocity to the constant\-speed bridge velocity\. It requires no precomputed OT coupling and instead imposes consistency on the model’s own endpoint map\.
The central theoretical property of PMOT is a zero\-loss exactness result\. Under the stated regularity, exact terminal matching, and uniqueness assumptions, every zero\-loss solution satisfies the generalized Benamou–Brenier optimality system for the correspondingpp\-cost problem\. Consequently, in this idealized limit, the induced flow recovers the exactpp\-Wasserstein dynamics and attains the optimal transport cost\. In finite\-sample neural training, this exactness is subject to empirical sampling, model capacity, numerical integration, and optimization error; our empirical claims are therefore stated in terms of alignment withpp\-matched OT references rather than unconditional exact recovery\. Accordingly, we evaluate geometric fidelity using post\-hoc OT references computed with the same exponent as the trained model\.
We evaluate PMOT on synthetic transport, high\-dimensional tabular density estimation, and image color transformation\. The experiments assesspp\-matched geometric alignment, generation under mild objective weighting, likelihood\-based density modeling, and MMD\-based sample transport\.
Our contributions are summarized as follows:
- •We introduce PMOT, a potential\-flow CNF framework for generalpp\-cost optimal transport, trained through self\-induced potential matching without external OT couplings or inner OT optimization\.
- •We prove that, under the stated regularity, exact terminal matching, and uniqueness assumptions, every zero\-loss PMOT solution recovers the correspondingpp\-optimal transport map and dynamics\.
- •We evaluate each model against OT references computed with the matching cost exponent and demonstratepp\-specific geometric alignment, competitive likelihood modeling, and flexible terminal matching through KL/NLL and MMD objectives\.
## 2Preliminaries
### 2\.1Generalpp\-Cost OT and Displacement Interpolation
Letμ0,μ1\\mu\_\{0\},\\mu\_\{1\}be probability measures onℝd\\mathbb\{R\}^\{d\}with finitepp\-th moments\. Forp≥1p\\geq 1, thepp\-cost optimal transport problem is
Wpp\(μ0,μ1\)=infπ∈Π\(μ0,μ1\)∫‖x−y‖pdπ\(x,y\)\.W\_\{p\}^\{p\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\inf\_\{\\pi\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)\}\\int\\\|x\-y\\\|^\{p\}\\,\\mathrm\{d\}\\pi\(x,y\)\.When an optimal Monge mapTpT\_\{p\}exists, it satisfiesTp\#μ0=μ1T\_\{p\\\#\}\\mu\_\{0\}=\\mu\_\{1\}and attains the same cost\. Its displacement interpolation is
Xt\(x\)=\(1−t\)x\+tTp\(x\),ρt=\(Xt\)\#μ0,X\_\{t\}\(x\)=\(1\-t\)x\+tT\_\{p\}\(x\),\\qquad\\rho\_\{t\}=\(X\_\{t\}\)\_\{\\\#\}\\mu\_\{0\},with velocity
up\(t,Xt\(x\)\)=Tp\(x\)−x\.u\_\{p\}\(t,X\_\{t\}\(x\)\)=T\_\{p\}\(x\)\-x\.This path attains the dynamicpp\-cost:
∫01∫ℝd‖up\(t,z\)‖pdρt\(z\)dt=∫ℝd‖Tp\(x\)−x‖pdμ0\(x\)=Wpp\(μ0,μ1\)\.\\int\_\{0\}^\{1\}\\int\_\{\\mathbb\{R\}^\{d\}\}\\\|u\_\{p\}\(t,z\)\\\|^\{p\}\\,\\mathrm\{d\}\\rho\_\{t\}\(z\)\\,\\mathrm\{d\}t=\\int\_\{\\mathbb\{R\}^\{d\}\}\\\|T\_\{p\}\(x\)\-x\\\|^\{p\}\\,\\mathrm\{d\}\\mu\_\{0\}\(x\)=W\_\{p\}^\{p\}\(\\mu\_\{0\},\\mu\_\{1\}\)\.
### 2\.2Benamou–Brenier Optimality and Potential Velocities
The samepp\-cost transport problem has a dynamic Benamou–Brenier formulation\(Santambrogio,[2015](https://arxiv.org/html/2608.05666#bib.bib5)\)\. The dynamic problem is
infρ,v1p∫01∫ℝd‖v\(t,x\)‖p𝑑ρt\(x\)𝑑t\\inf\_\{\\rho,v\}\\frac\{1\}\{p\}\\int\_\{0\}^\{1\}\\int\_\{\\mathbb\{R\}^\{d\}\}\\\|v\(t,x\)\\\|^\{p\}\\,d\\rho\_\{t\}\(x\)\\,dtsubject to the endpoint constraintsρt=0=μ0\\rho\_\{t=0\}=\\mu\_\{0\},ρt=1=μ1\\rho\_\{t=1\}=\\mu\_\{1\}and the continuity equation
∂tρt\+∇x⋅\(ρtv\)=0\.\\partial\_\{t\}\\rho\_\{t\}\+\\nabla\_\{x\}\\cdot\(\\rho\_\{t\}v\)=0\.In a smooth formal derivation, the scalar fieldΦ\(t,x\)\\Phi\(t,x\)appears as the Lagrange multiplier for this continuity\-equation constraint\. The corresponding Euler–Lagrange condition with respect to the velocity gives on the support ofρt\\rho\_\{t\}that
‖v‖p−2v=−∇xΦ\.\\\|v\\\|^\{p\-2\}v=\-\\nabla\_\{x\}\\Phi\.Sincep\>1p\>1, the mapz↦‖z‖p−2zz\\mapsto\\\|z\\\|^\{p\-2\}zis invertible, and therefore the optimal velocity has the potential form
v=−‖∇xΦ‖q−2∇xΦ=−∇ξHq\(∇xΦ\),Hq\(ξ\)=1q‖ξ‖q\.v=\-\\\|\\nabla\_\{x\}\\Phi\\\|^\{q\-2\}\\nabla\_\{x\}\\Phi=\-\\nabla\_\{\\xi\}H\_\{q\}\(\\nabla\_\{x\}\\Phi\),\\qquad H\_\{q\}\(\\xi\)=\\frac\{1\}\{q\}\\\|\\xi\\\|^\{q\}\.Thus, the potential parameterization used by PMOT is not an arbitrary assumption: it mirrors the velocity–multiplier relation in the smooth Benamou–Brenier optimality system\. The Euler–Lagrange condition with respect toρ\\rhothen yields∂tΦ\+∇xΦ⋅v\+1p∥v∥p=0\\partial\_\{t\}\\Phi\+\\nabla\_\{x\}\\Phi\\cdot v\+\\frac\{1\}\{p\}\\lVert v\\rVert^\{p\}=0and thus on the support ofρt\\rho\_\{t\}that
∂tΦ−1q‖∇Φ‖q=0\.\\partial\_\{t\}\\Phi\-\\frac\{1\}\{q\}\\\|\\nabla\\Phi\\\|^\{q\}=0\.As we shall see, this corresponds to the fact that the velocity of a particle stays constant along its trajectory, a key feature of the OT\. During neural training,Φθ\\Phi\_\{\\theta\}is a model parameter; it coincides with the BB multiplier only in the ideal zero\-loss optimal regime, up to adding a function of time, which does not change∇xΦθ\\nabla\_\{x\}\\Phi\_\{\\theta\}or the induced velocity\.
### 2\.3Continuous Normalizing Flows
A Continuous Normalizing Flow defines an invertible map through an ordinary differential equation
dztdt=vθ\(t,zt\),z0=x\.\\frac\{dz\_\{t\}\}\{dt\}=v\_\{\\theta\}\(t,z\_\{t\}\),\\qquad z\_\{0\}=x\.LetFθF\_\{\\theta\}denote the time\-one flow map,Fθ\(x\)=z1F\_\{\\theta\}\(x\)=z\_\{1\}\. If the terminal density is chosen as a tractable prior, such as𝒩\(0,I\)\\mathcal\{N\}\(0,I\), the data likelihood is computed by the instantaneous change\-of\-variables formula:
ddtlogρt\(zt\)=−∇⋅vθ\(t,zt\)\.\\frac\{d\}\{dt\}\\log\\rho\_\{t\}\(z\_\{t\}\)=\-\\nabla\\cdot v\_\{\\theta\}\(t,z\_\{t\}\)\.Therefore,
logρ0\(x\)=logρ1\(z1\)\+∫01∇⋅vθ\(t,zt\)𝑑t,\\log\\rho\_\{0\}\(x\)=\\log\\rho\_\{1\}\(z\_\{1\}\)\+\\int\_\{0\}^\{1\}\\nabla\\cdot v\_\{\\theta\}\(t,z\_\{t\}\)\\,dt,up to the sign convention determined by the integration direction\. CNFs therefore provide likelihood evaluation while simultaneously defining a continuous transport trajectory\.
## 3Potential Matching Optimal Transport
We introduce*Potential Matching Optimal Transport*\(PMOT\), a variational framework for generalpp\-cost optimal transport\. PMOT parameterizes the CNF velocity field through a scalar potential and optimizes two terms: potential matching along model\-generated straight bridges and terminal distribution matching\. Under the assumptions stated below, every sufficiently regular zero\-loss solution recovers the uniquepp\-optimal Monge map and its corresponding Benamou–Brenier velocity field\.
### 3\.1Potential\-Induced Flow
Letp\>1p\>1andq=p/\(p−1\)q=p/\(p\-1\)\. Following the Benamou–Brenier optimality relation described above, we take a scalar potentialΦ:\[0,1\]×ℝd→ℝ\\Phi:\[0,1\]\\times\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}as the primary variable and define the induced velocity field
vΦ\(t,x\)=−∥∇xΦ\(t,x\)∥q−2∇xΦ\(t,x\)\.v\_\{\\Phi\}\(t,x\)=\-\\lVert\\nabla\_\{x\}\\Phi\(t,x\)\\rVert^\{q\-2\}\\nabla\_\{x\}\\Phi\(t,x\)\.\(1\)Forp=2p=2, equation[1](https://arxiv.org/html/2608.05666#S3.E1)reduces to the familiar potential\-flow relationvΦ\(t,x\)=−∇xΦ\(t,x\)v\_\{\\Phi\}\(t,x\)=\-\\nabla\_\{x\}\\Phi\(t,x\)\.
For eachx∈ℝdx\\in\\mathbb\{R\}^\{d\}, consider the initial\-value problem
∂tz\(t\)=vΦ\(t,z\(t\)\),for almost everyt∈\(0,1\),z\(0\)=x\.\\partial\_\{t\}z\(t\)=v\_\{\\Phi\}\\bigl\(t,z\(t\)\\bigr\),\\text\{for almost every \}t\\in\(0,1\),\\qquad z\(0\)=x\.\(2\)Throughout, we restrict attention to potentialsΦ\\Phisuch that, for everyx∈ℝdx\\in\\mathbb\{R\}^\{d\}, problem equation[2](https://arxiv.org/html/2608.05666#S3.E2)admits a unique solution inAC\(\[0,1\];ℝd\)\\mathrm\{AC\}\(\[0,1\];\\mathbb\{R\}^\{d\}\)\. We denote this solution byzΦ\(⋅,x\)z\_\{\\Phi\}\(\\cdot,x\)\. And we further require that the associated solution mapzΦ:\[0,1\]×ℝd→ℝdz\_\{\\Phi\}:\[0,1\]\\times\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}^\{d\}is jointly Borel measurable\.
The terminal flow mapFΦ:ℝd→ℝdF\_\{\\Phi\}:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}^\{d\}is defined by
FΦ\(x\)=zΦ\(1,x\)\.F\_\{\\Phi\}\(x\)=z\_\{\\Phi\}\(1,x\)\.The distribution induced at timettis
ρtΦ=\(zΦ\(t,⋅\)\)\#μ0\.\\rho\_\{t\}^\{\\Phi\}=\\bigl\(z\_\{\\Phi\}\(t,\\cdot\)\\bigr\)\_\{\\\#\}\\mu\_\{0\}\.In particular,ρ0Φ=μ0,ρ1Φ=\(FΦ\)\#μ0\\rho\_\{0\}^\{\\Phi\}=\\mu\_\{0\},\\rho\_\{1\}^\{\\Phi\}=\(F\_\{\\Phi\}\)\_\{\\\#\}\\mu\_\{0\}\.
### 3\.2The PMOT Variational Objective
The potential matching loss is defined by
ℒPM\(Φ\):=𝔼x∼μ0\[∫01‖∇xΦ\(t,\(1−t\)x\+tFΦ\(x\)\)\+‖FΦ\(x\)−x‖p−2\(FΦ\(x\)−x\)‖2dt\]\.\\mathcal\{L\}\_\{\\mathrm\{PM\}\}\(\\Phi\):=\\mathbb\{E\}\_\{x\\sim\\mu\_\{0\}\}\\left\[\\int\_\{0\}^\{1\}\\left\\\|\\nabla\_\{x\}\\Phi\\\!\\left\(t,\(1\-t\)x\+tF\_\{\\Phi\}\(x\)\\right\)\+\\left\\\|F\_\{\\Phi\}\(x\)\-x\\right\\\|^\{p\-2\}\\left\(F\_\{\\Phi\}\(x\)\-x\\right\)\\right\\\|^\{2\}\\,\\mathrm\{d\}t\\right\]\.
Let𝒟\\mathcal\{D\}be a nonnegative discrepancy between probability measures satisfying𝒟\(P,Q\)=0⇔P=Q\\mathcal\{D\}\(P,Q\)=0\\Leftrightarrow P=Q\. We define the terminal distribution loss by
ℒterm\(Φ\):=𝒟\(ρ1Φ,μ1\)\.\\mathcal\{L\}\_\{\\mathrm\{term\}\}\(\\Phi\):=\\mathcal\{D\}\(\\rho\_\{1\}^\{\\Phi\},\\mu\_\{1\}\)\.A principal example is the forward KL divergence, which givesℒterm\(Φ\)=KL\(ρ1Φ∥μ1\)\\mathcal\{L\}\_\{\\mathrm\{term\}\}\(\\Phi\)=\\mathrm\{KL\}\(\\rho\_\{1\}^\{\\Phi\}\\\|\\mu\_\{1\}\)\.
We define the PMOT objective to be
ℒPMOT\(Φ\):=λPMℒPM\(Φ\)\+λtermℒterm\(Φ\),\\mathcal\{L\}\_\{\\mathrm\{PMOT\}\}\(\\Phi\):=\\lambda\_\{\\mathrm\{PM\}\}\\mathcal\{L\}\_\{\\mathrm\{PM\}\}\(\\Phi\)\+\\lambda\_\{\\mathrm\{term\}\}\\mathcal\{L\}\_\{\\mathrm\{term\}\}\(\\Phi\),whereλPM,λterm\>0\\lambda\_\{\\mathrm\{PM\}\},\\lambda\_\{\\mathrm\{term\}\}\>0are fixed weights\. The domaindom\(ℒPMOT\)\\operatorname\{dom\}\(\\mathcal\{L\}\_\{\\mathrm\{PMOT\}\}\)consists of all potentialsΦ\\Phisatisfying the preceding flow well\-posedness and joint Borel\-measurability requirements for which bothℒPM\(Φ\)\\mathcal\{L\}\_\{\\mathrm\{PM\}\}\(\\Phi\)andℒterm\(Φ\)\\mathcal\{L\}\_\{\\mathrm\{term\}\}\(\\Phi\)are well defined\.
### 3\.3Zero Loss Implies Optimal Transport
Our main theorem is:
###### Theorem 3\.2\(Exact recovery and velocity\-field uniqueness at zero loss\)\.
Letp∈\(1,∞\)p\\in\(1,\\infty\), letμ0,μ1∈𝒫p\(ℝd\)\\mu\_\{0\},\\mu\_\{1\}\\in\\mathcal\{P\}\_\{p\}\(\\mathbb\{R\}^\{d\}\), and writedom\(ℒPMOT\)\\operatorname\{dom\}\(\\mathcal\{L\}\_\{\\mathrm\{PMOT\}\}\)for the domain of the PMOT objective\. Assume thatμ0\\mu\_\{0\}is absolutely continuous with respect to Lebesgue measure\. Assume that thepp\-Benamou–Brenier problem admits a minimizer\(ρ⋆,v⋆\)=\(ρΦ⋆,vΦ⋆\)\(\\rho^\{\\star\},v^\{\\star\}\)=\(\\rho^\{\\Phi^\{\\star\}\},v\_\{\\Phi^\{\\star\}\}\)for someΦ⋆∈dom\(ℒPMOT\)\\Phi^\{\\star\}\\in\\operatorname\{dom\}\(\\mathcal\{L\}\_\{\\mathrm\{PMOT\}\}\)\. Then
minΦ∈dom\(ℒPMOT\)ℒPMOT\(Φ\)=ℒPMOT\(Φ⋆\)=0\.\\min\_\{\\Phi\\in\\operatorname\{dom\}\(\\mathcal\{L\}\_\{\\mathrm\{PMOT\}\}\)\}\\mathcal\{L\}\_\{\\mathrm\{PMOT\}\}\(\\Phi\)=\\mathcal\{L\}\_\{\\mathrm\{PMOT\}\}\(\\Phi^\{\\star\}\)=0\.Furthermore, letΦ\\Phibe a global minimizer\. IfΦ∈C2\(\(0,1\)×ℝd\)\\Phi\\in C^\{2\}\(\(0,1\)\\times\\mathbb\{R\}^\{d\}\), and ifsupp\(ρtΦ\)=ℝd\\operatorname\{supp\}\(\\rho\_\{t\}^\{\\Phi\}\)=\\mathbb\{R\}^\{d\}for everyt∈\(0,1\)t\\in\(0,1\), thenvΦ=v⋆v\_\{\\Phi\}=v^\{\\star\},dtρtΦ\(dx\)\\mathrm\{d\}t\\,\\rho\_\{t\}^\{\\Phi\}\(\\mathrm\{d\}x\)\-almost everywhere, and the corresponding terminal mapFΦF\_\{\\Phi\}is the unique optimal Monge map fromμ0\\mu\_\{0\}toμ1\\mu\_\{1\}for the cost∥x−y∥p\\lVert x\-y\\rVert^\{p\}\.
Under the stated assumptions, this theorem provides the theoretical justification for PMOT: any sufficiently regular potentialΦ\\Phiattaining zero PMOT loss recovers the uniquepp\-optimal Monge map and the corresponding Benamou–Brenier velocity field\. Thus, in the ideal zero\-loss setting, the PMOT objective exactly identifies the dynamicpp\-optimal transport solution\. The proof is given in Appendix[A](https://arxiv.org/html/2608.05666#A1)\.
### 3\.4Neural Parameterization and Training
For computation, we restrict the potential to a neural family\{Φθ\}\\\{\\Phi\_\{\\theta\}\\\}and use the scalar\-valued ResNet potential architecture of OT\-Flow\(Onkenet al\.,[2021](https://arxiv.org/html/2608.05666#bib.bib9)\)\. Letgθ\(t,x\)=∇xΦθ\(t,x\)g\_\{\\theta\}\(t,x\)=\\nabla\_\{x\}\\Phi\_\{\\theta\}\(t,x\)\. The implementation uses the stabilized velocity
vθ,ϵ\(t,x\)=−\(‖gθ\(t,x\)‖2\+ϵ\)\(q−2\)/2gθ\(t,x\),ϵ=10−12,v\_\{\\theta,\\epsilon\}\(t,x\)=\-\\bigl\(\\\|g\_\{\\theta\}\(t,x\)\\\|^\{2\}\+\\epsilon\\bigr\)^\{\(q\-2\)/2\}g\_\{\\theta\}\(t,x\),\\qquad\\epsilon=10^\{\-12\},\(3\)which approaches the population velocity in equation[1](https://arxiv.org/html/2608.05666#S3.E1)asϵ→0\\epsilon\\to 0\. We writevθv\_\{\\theta\}forvθ,ϵv\_\{\\theta,\\epsilon\}below\. The resulting ODE defines the flowzθ\(t,x\)z\_\{\\theta\}\(t,x\)and terminal mapFθ\(x\)=zθ\(1,x\)F\_\{\\theta\}\(x\)=z\_\{\\theta\}\(1,x\)\.
Given a minibatch\{x0\(i\)\}i=1B\\\{x\_\{0\}^\{\(i\)\}\\\}\_\{i=1\}^\{B\}fromμ0\\mu\_\{0\}, a forward ODE solve produces the model\-induced endpointsx1\(i\)=Fθ\(x0\(i\)\)x\_\{1\}^\{\(i\)\}=F\_\{\\theta\}\(x\_\{0\}^\{\(i\)\}\)\. For independently sampledti∼𝒰\[0,1\]t\_\{i\}\\sim\\mathcal\{U\}\[0,1\], define
x¯ti\(i\)=\(1−ti\)x0\(i\)\+tix1\(i\),u\(i\)=x1\(i\)−x0\(i\)\.\\bar\{x\}\_\{t\_\{i\}\}^\{\(i\)\}=\(1\-t\_\{i\}\)x\_\{0\}^\{\(i\)\}\+t\_\{i\}x\_\{1\}^\{\(i\)\},\\qquad u^\{\(i\)\}=x\_\{1\}^\{\(i\)\}\-x\_\{0\}^\{\(i\)\}\.Let
mp\(u\)=‖u‖p−2u,mp\(0\)=0,m\_\{p\}\(u\)=\\\|u\\\|^\{p\-2\}u,\\qquad m\_\{p\}\(0\)=0,denote the momentum dual to thepp\-power kinetic energy\. The stochastic self\-induced potential\-matching loss is
ℒ^PM\(θ\)=1B∑i=1B‖∇xΦθ\(ti,x¯ti\(i\)\)\+mp\(u\(i\)\)‖2\.\\widehat\{\\mathcal\{L\}\}\_\{\\mathrm\{PM\}\}\(\\theta\)=\\frac\{1\}\{B\}\\sum\_\{i=1\}^\{B\}\\left\\\|\\nabla\_\{x\}\\Phi\_\{\\theta\}\\bigl\(t\_\{i\},\\bar\{x\}\_\{t\_\{i\}\}^\{\(i\)\}\\bigr\)\+m\_\{p\}\\bigl\(u^\{\(i\)\}\\bigr\)\\right\\\|^\{2\}\.\(4\)
The model\-induced endpoints remain attached to the computational graph, so gradients propagate throughFθF\_\{\\theta\}in the potential\-matching loss\.
The population terminal objective remains
ℒterm\(θ\)=𝒟\(\(Fθ\)\#μ0,μ1\)\.\\mathcal\{L\}\_\{\\mathrm\{term\}\}\(\\theta\)=\\mathcal\{D\}\\bigl\(\(F\_\{\\theta\}\)\_\{\\\#\}\\mu\_\{0\},\\mu\_\{1\}\\bigr\)\.\(5\)For likelihood\-based training, we use the forward\-KL discrepancy and optimize its equivalent CNF negative log\-likelihood\. Writing
ℓθ\(x\)=∫01∇⋅vθ\(t,zθ\(t,x\)\)dt,\\ell\_\{\\theta\}\(x\)=\\int\_\{0\}^\{1\}\\nabla\\cdot v\_\{\\theta\}\\bigl\(t,z\_\{\\theta\}\(t,x\)\\bigr\)\\,\\mathrm\{d\}t,the empirical NLL is
ℒ^NLL\(θ\)=−1B∑i=1B\[logφ1\(Fθ\(x0\(i\)\)\)\+ℓθ\(x0\(i\)\)\],\\widehat\{\\mathcal\{L\}\}\_\{\\mathrm\{NLL\}\}\(\\theta\)=\-\\frac\{1\}\{B\}\\sum\_\{i=1\}^\{B\}\\left\[\\log\\varphi\_\{1\}\\bigl\(F\_\{\\theta\}\(x\_\{0\}^\{\(i\)\}\)\\bigr\)\+\\ell\_\{\\theta\}\(x\_\{0\}^\{\(i\)\}\)\\right\],whereφ1\\varphi\_\{1\}is the terminal reference density\. The NLL differs from the forward\-KL objective by aθ\\theta\-independent source\-entropy constant\. Therefore, the zero\-loss statement in Theorem[3\.2](https://arxiv.org/html/2608.05666#S3.Thmtheorem2)refers to the population KL discrepancy, not to the numerical NLL value reported in experiments\.
For sample\-based terminal matching, we instead use a minibatch estimate of squared MMD between\{Fθ\(x0\(i\)\)\}\\\{F\_\{\\theta\}\(x\_\{0\}^\{\(i\)\}\)\\\}and independent samples fromμ1\\mu\_\{1\}\(Grettonet al\.,[2012](https://arxiv.org/html/2608.05666#bib.bib11)\)\.
In either case, the implemented objective is
𝒥^PMOT\(θ\)=λPMℒ^PM\(θ\)\+λtermℒ^term\(θ\)\.\\widehat\{\\mathcal\{J\}\}\_\{\\mathrm\{PMOT\}\}\(\\theta\)=\\lambda\_\{\\mathrm\{PM\}\}\\widehat\{\\mathcal\{L\}\}\_\{\\mathrm\{PM\}\}\(\\theta\)\+\\lambda\_\{\\mathrm\{term\}\}\\widehat\{\\mathcal\{L\}\}\_\{\\mathrm\{term\}\}\(\\theta\)\.\(6\)No external OT pairs or HJB residual are used during training\. We solve the dynamics using a differentiable fixed\-step RK4 scheme and backpropagate through the ODE solve\. Likelihood evaluation integrates data forward to the reference, whereas generation integrates reference samples backward to the data space\.
## 4Experiments
Our experiments are designed to evaluate three aspects of PMOT\. First, we test whether PMOT endpoint maps agree with post\-hoc OT references computed under the intendedpp\-cost\. Second, we examine whether the self\-induced matching objective yields stable generation under mild likelihood weighting\. Third, we evaluate whether PMOT remains competitive as a likelihood\-based CNF on high\-dimensional tabular density estimation\. Across all experiments, Sinkhorn references are used only for evaluation and visualization, and never to construct training pairs\.
### 4\.1Toy Transport Geometry
Two\-dimensional synthetic distributions provide a controlled setting in which post\-hoc OT references can be estimated\. For each exponentpp, we compute an entropically regularized Sinkhorn coupling withcp\(x,y\)=‖x−y‖pc\_\{p\}\(x,y\)=\\\|x\-y\\\|^\{p\}\(Cuturi,[2013](https://arxiv.org/html/2608.05666#bib.bib12)\)\. Its transport cost and barycentric projection provide the reference cost and endpoint map, respectively; the resulting coupling and its derived quantities are used only post hoc for evaluation and visualization, never for training\. We report the model\-inducedpp\-cost, its ratio to the reference cost, and endpoint MSE to thepp\-matched barycentric reference\. Sinkhorn references are never used to construct training pairs\. Additional unconditional generation results for the toy distributions are provided in Appendix[B\.1](https://arxiv.org/html/2608.05666#A2.SS1)\.
Table 1:OT\-geometry evaluation on 8\-Gaussians and Pinwheel\. PMOT endpoints are compared with post\-hoc Sinkhorn barycentric references under the matchingpp\-cost\. Sinkhorn is used only for evaluation, withε=0\.02\\varepsilon=0\.02and 1000 iterations on 2000 samples\. Ratios within a few percent of one indicate close agreement with the reference\.Table[1](https://arxiv.org/html/2608.05666#S4.T1)gives the main quantitative geometry evaluation\. On 8\-Gaussians, all cost ratios are within2\.4%2\.4\\%of one\. On Pinwheel, all ratios are within3\.4%3\.4\\%of one\. Thus, when each model is evaluated against the matching exponent, its endpoint map has both transport cost and barycentric discrepancy close to the corresponding post\-hoc OT reference\.
To test whether the learned maps depend on the training exponent, we compare each checkpoint trained withptrain∈\{1\.5,2,3\}p\_\{\\mathrm\{train\}\}\\in\\\{1\.5,2,3\\\}against barycentric reference maps computed with everypref∈\{1\.5,2,3\}p\_\{\\mathrm\{ref\}\}\\in\\\{1\.5,2,3\\\}\. Each entry in Figure[1](https://arxiv.org/html/2608.05666#S4.F1)reports the corresponding endpoint MSE\.
Figure 1:pp\-swap evaluation on Pinwheel\. Rows correspond to PMOT checkpoints trained with exponentptrainp\_\{\\mathrm\{train\}\}, and columns correspond to post\-hoc Sinkhorn barycentric reference maps computed with exponentprefp\_\{\\mathrm\{ref\}\}\. Each entry reports endpoint MSE for evaluation seed 0; lower is better\. Bold entries mark the minimum in each row\. Across five independently resampled evaluation batches with fixed checkpoints, the mean off\-diagonal endpoint MSE is1\.41±0\.051\.41\\pm 0\.05times the mean diagonal endpoint MSE\.Figure[1](https://arxiv.org/html/2608.05666#S4.F1)shows a diagonal preference: each PMOT checkpoint is closest to the reference computed with its training exponent\. The separation is strongest betweenp=1\.5p=1\.5andp=3p=3, while neighboring exponents remain closer on this smooth toy distribution\. This supportspp\-specific behavior without requiring visually distinct trajectories\.
Moons provides a complementary qualitative example in a less symmetric setting\. Because entropic regularization can noticeably affect barycentric references on this dataset, we use this experiment only to visualize the learned coupling and continuous trajectories\.
Figure 2:Learned transport geometry on two moons forp=2p=2\. \(a\) Coupling induced by the PMOT map\. \(b\) Post\-hoc Sinkhorn barycentric reference\. \(c\) Continuous PMOT trajectories\. The resulting coupling and its derived quantities are used only post hoc for evaluation and visualization, never for training\.Figure[2](https://arxiv.org/html/2608.05666#S4.F2)compares the PMOT\-induced map with a post\-hoc Sinkhorn reference and the continuous PMOT trajectories\.
### 4\.2Likelihood\-Weight Sensitivity
We next compare PMOT with OT\-Flow on 8\-Gaussians under equal training budgets\. PMOT uses only self\-induced potential matching and likelihood, with weights\(λPM,λNLL\)=\(1,1\)\(\\lambda\_\{\\mathrm\{PM\}\},\\lambda\_\{\\mathrm\{NLL\}\}\)=\(1,1\)\. OT\-Flow uses its standard kinetic, likelihood, and HJB/R objective; we fix the kinetic and HJB/R coefficients to one and vary only the likelihood coefficient\. This comparison tests whether good generation quality requires a strongly likelihood\-dominated objective\.
Table 2:Generation\-quality comparison on 8\-Gaussians\. PMOT uses self\-induced potential matching and likelihood weights\(λPM,λNLL\)=\(1,1\)\(\\lambda\_\{\\mathrm\{PM\}\},\\lambda\_\{\\mathrm\{NLL\}\}\)=\(1,1\)and has no HJB/R regularizer\. For OT\-Flow,\(λL,λC,λR\)=\(1,wC,1\)\(\\lambda\_\{L\},\\lambda\_\{C\},\\lambda\_\{R\}\)=\(1,w\_\{C\},1\)and onlywCw\_\{C\}is varied\. All models are trained for 10k iterations\. Lower is better\.Figure 3:Generation quality on 8\-Gaussians under a limited training budget of 1000 iterations\. PMOT uses the mild balanced weights\(λPM,λNLL\)=\(1,1\)\(\\lambda\_\{\\mathrm\{PM\}\},\\lambda\_\{\\mathrm\{NLL\}\}\)=\(1,1\), while OT\-Flow uses\(λL,λC,λR\)=\(1,1,1\)\(\\lambda\_\{L\},\\lambda\_\{C\},\\lambda\_\{R\}\)=\(1,1,1\)\. Under the same training budget and mild weighting, PMOT recovers the eight modes more clearly, while OT\-Flow exhibits stronger inter\-mode artifacts\.Figure[3](https://arxiv.org/html/2608.05666#S4.F3)shows the same effect visually under a shorter 1000\-iteration budget: PMOT already recovers the eight modes clearly with balanced weights, whereas OT\-Flow with\(1,1,1\)\(1,1,1\)exhibits stronger inter\-mode artifacts\. Table[2](https://arxiv.org/html/2608.05666#S4.T2)quantifies the trend after 10k iterations\. OT\-Flow improves substantially as the likelihood coefficient is increased, whereas PMOT obtains the best histogram and MMD scores with the balanced\(1,1\)\(1,1\)objective\. This supports the view that the self\-induced matching term provides useful transport structure without adding an HJB/R regularizer\.
### 4\.3High\-Dimensional Density Estimation
Table 3:High\-dimensional tabular density estimation\. For PMOT, weights are reported as\(λPM,λNLL\)\(\\lambda\_\{\\mathrm\{PM\}\},\\lambda\_\{\\mathrm\{NLL\}\}\); for OT\-Flow, weights are reported as\(λL,λC,λR\)\(\\lambda\_\{L\},\\lambda\_\{C\},\\lambda\_\{R\}\)\. PMOT uses self\-induced potential matching and likelihood without an HJB/R regularizer\. MMD compares inverse\-generated samples with held\-out data\. Lower is better\.Table[3](https://arxiv.org/html/2608.05666#S4.T3)compares PMOT and OT\-Flow on MiniBooNE, POWER, and HEPMASS\(Papamakarioset al\.,[2017](https://arxiv.org/html/2608.05666#bib.bib13); Grathwohlet al\.,[2018](https://arxiv.org/html/2608.05666#bib.bib8); Onkenet al\.,[2021](https://arxiv.org/html/2608.05666#bib.bib9)\)\. PMOT uses the same mild\(1,5\)\(1,5\)weighting across all three datasets and does not include an HJB/R regularizer\. OT\-Flow uses dataset\-specific likelihood and HJB/R coefficients\. In these runs, PMOT obtains lower NLL on all three datasets and comparable generation MMD, indicating that the simpler objective remains competitive in high\-dimensional tabular density estimation\.
#### Image color transformation\.
To illustrate sample\-based terminal matching, we apply MMD\-based PMOT to image color transformation\. Figure[4](https://arxiv.org/html/2608.05666#S4.F4)shows the transformation along the learned flow; experimental details and additional visualizations are provided in Appendix[B\.4](https://arxiv.org/html/2608.05666#A2.SS4)\.
Figure 4:MMD\-based image color transformation along the learned PMOT flow\. The leftmost panel is the source image and the rightmost panel is the target style image\.
## 5Related Work
#### CNFs and optimal\-transport dynamics\.
Continuous Normalizing Flows \(CNFs\) define invertible transformations through neural ordinary differential equations and evaluate likelihoods using the instantaneous change\-of\-variables formula\(Chenet al\.,[2018](https://arxiv.org/html/2608.05666#bib.bib7); Grathwohlet al\.,[2018](https://arxiv.org/html/2608.05666#bib.bib8)\)\. This makes CNFs natural candidates for learning continuous transport maps, but likelihood maximization alone does not identify a unique path between the data distribution and the reference prior\. OT\-Flow addresses this ambiguity by combining a potential\-flow CNF with kinetic\-energy and HJB regularization, yielding dynamics closely related to the Benamou–Brenier formulation of quadratic\-cost optimal transport\(Benamou and Brenier,[2000](https://arxiv.org/html/2608.05666#bib.bib1); Onkenet al\.,[2021](https://arxiv.org/html/2608.05666#bib.bib9); Jing and Li,[2025](https://arxiv.org/html/2608.05666#bib.bib20)\)\. Recent work has also used machine learning to approximate geodesic dynamics under related transport geometries, such as the spherical Wasserstein–Fisher–Rao metric\(Jinget al\.,[2024](https://arxiv.org/html/2608.05666#bib.bib19)\)\. PMOT likewise learns continuous transport dynamics, but focuses on likelihood\-based potential flows for generalpp\-cost Wasserstein geometry\(Villani and others,[2009](https://arxiv.org/html/2608.05666#bib.bib2); Peyré and Cuturi,[2019](https://arxiv.org/html/2608.05666#bib.bib3)\)\. It parameterizes the CNF velocity in the generalized Benamou–Brenier form for the chosen exponentppand replaces the quadratic kinetic\-energy regularizer with a self\-induced potential\-matching objective \.
#### Flow matching and neural OT with general costs\.
Flow matching trains continuous\-time generative models by regressing velocity fields along prescribed probability paths, avoiding the need to backpropagate through ODE solves during training\(Lipmanet al\.,[2023](https://arxiv.org/html/2608.05666#bib.bib10)\)\. OT\-coupled flow\-matching methods use minibatch OT couplings to define training pairs or conditional paths\(Tonget al\.,[2023](https://arxiv.org/html/2608.05666#bib.bib16)\)\. PMOT instead uses a potential matching residual whose endpoint pair is induced by the current CNF map, so no external OT coupling is used during training\. Neural OT methods also learn maps or couplings under general cost functions, including costs beyond the quadratic case\(Korotinet al\.,[2023](https://arxiv.org/html/2608.05666#bib.bib14); Asadulaevet al\.,[2024](https://arxiv.org/html/2608.05666#bib.bib15)\)\. PMOT is complementary to this literature: rather than learning only a static map or using minibatch OT pairs as supervision, it learns continuous potential\-flow dynamics while preserving CNF likelihood evaluation and allowing flexible terminal matching such as KL/NLL or MMD\.
## 6Conclusion
We presented Potential Matching Optimal Transport \(PMOT\), a potential\-flow CNF framework for generalpp\-cost optimal transport\. PMOT parameterizes the CNF velocity field with a scalar potential in the generalized Benamou–Brenier form and trains the potential gradient using a self\-induced matching loss on straight bridges determined by the model endpoints\. It requires no external transport pairs or inner OT optimization and supports both likelihood\-based training through a KL/NLL terminal objective and sample\-based terminal matching through MMD\.
Our main result establishes exactness at zero population PMOT loss: under the stated realizability, regularity, and full\-support assumptions, every global minimizer recovers the uniquepp\-optimal Monge map and its Benamou–Brenier dynamics\. On synthetic benchmarks, the learned maps agree withpp\-matched OT references and are closest to the references computed using their respective training exponents\. PMOT also remains competitive as a likelihood\-based density model on high\-dimensional tabular data, while an additional MMD\-based color transformation experiment demonstrates sample\-based terminal matching\. As with deterministic CNFs more broadly, PMOT represents Monge\-type maps, and finite\-sample training, numerical integration, and optimization introduce practical approximation errors\.
## References
- A\. Asadulaev, A\. Korotin, V\. Egiazarian, P\. Mokrov, and E\. Burnaev \(2024\)Neural optimal transport with general cost functionals\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=gIiz7tBtYZ)Cited by:[§5](https://arxiv.org/html/2608.05666#S5.SS0.SSS0.Px2.p1.1)\.
- J\. Benamou and Y\. Brenier \(2000\)A computational fluid mechanics solution to the monge\-kantorovich mass transfer problem\.Numerische Mathematik84\(3\),pp\. 375–393\.External Links:[Document](https://dx.doi.org/10.1007/s002110050002)Cited by:[§1](https://arxiv.org/html/2608.05666#S1.p2.1),[§5](https://arxiv.org/html/2608.05666#S5.SS0.SSS0.Px1.p1.2)\.
- R\. T\. Chen, Y\. Rubanova, J\. Bettencourt, and D\. K\. Duvenaud \(2018\)Neural ordinary differential equations\.Advances in neural information processing systems31\.Cited by:[§1](https://arxiv.org/html/2608.05666#S1.p1.1),[§5](https://arxiv.org/html/2608.05666#S5.SS0.SSS0.Px1.p1.2)\.
- M\. Cuturi \(2013\)Sinkhorn distances: lightspeed computation of optimal transport\.InAdvances in Neural Information Processing Systems,Cited by:[§4\.1](https://arxiv.org/html/2608.05666#S4.SS1.p1.4)\.
- J\. Deng, W\. Dong, R\. Socher, L\. Li, K\. Li, and L\. Fei\-Fei \(2009\)ImageNet: a large\-scale hierarchical image database\.In2009 IEEE Conference on Computer Vision and Pattern Recognition,Vol\.,pp\. 248–255\.External Links:[Document](https://dx.doi.org/10.1109/CVPR.2009.5206848)Cited by:[§B\.4](https://arxiv.org/html/2608.05666#A2.SS4.p2.6)\.
- W\. Grathwohl, R\. T\. Chen, J\. Bettencourt, I\. Sutskever, and D\. Duvenaud \(2018\)Ffjord: free\-form continuous dynamics for scalable reversible generative models\.arXiv preprint arXiv:1810\.01367\.Cited by:[§1](https://arxiv.org/html/2608.05666#S1.p1.1),[§4\.3](https://arxiv.org/html/2608.05666#S4.SS3.p1.1),[§5](https://arxiv.org/html/2608.05666#S5.SS0.SSS0.Px1.p1.2)\.
- A\. Gretton, K\. M\. Borgwardt, M\. J\. Rasch, B\. Schölkopf, and A\. Smola \(2012\)A kernel two\-sample test\.Journal of Machine Learning Research13\(25\),pp\. 723–773\.External Links:[Link](https://www.jmlr.org/papers/v13/gretton12a.html)Cited by:[§3\.4](https://arxiv.org/html/2608.05666#S3.SS4.p5.2)\.
- Y\. Jing, J\. Chen, L\. Li, and J\. Lu \(2024\)A machine learning framework for geodesics under spherical wasserstein–fisher–rao metric and its application for weighted sample generation\.Journal of Scientific Computing98\(1\),pp\. 5\.Cited by:[§5](https://arxiv.org/html/2608.05666#S5.SS0.SSS0.Px1.p1.2)\.
- Y\. Jing and L\. Li \(2025\)Convergence analysis of ot\-flow for sample generation\.Numerical Mathematics: Theory, Methods and Applications18\(2\),pp\. 325–352\.Cited by:[§5](https://arxiv.org/html/2608.05666#S5.SS0.SSS0.Px1.p1.2)\.
- A\. Korotin, L\. Li, A\. Genevay, J\. M\. Solomon, A\. Filippov, and E\. Burnaev \(2023\)Neural optimal transport\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=d8CBRlWNkqH)Cited by:[§5](https://arxiv.org/html/2608.05666#S5.SS0.SSS0.Px2.p1.1)\.
- Y\. Lipman, R\. T\. Q\. Chen, H\. Ben\-Hamu, M\. Nickel, and M\. Le \(2023\)Flow matching for generative modeling\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=PqvMRDCJT9t)Cited by:[§5](https://arxiv.org/html/2608.05666#S5.SS0.SSS0.Px2.p1.1)\.
- D\. Onken, S\. Wu Fung, X\. Li, and L\. Ruthotto \(2021\)OT\-Flow: fast and accurate continuous normalizing flows via optimal transport\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.35,pp\. 9223–9232\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v35i10.17113),[Link](https://ojs.aaai.org/index.php/AAAI/article/view/17113)Cited by:[§1](https://arxiv.org/html/2608.05666#S1.p2.1),[§3\.4](https://arxiv.org/html/2608.05666#S3.SS4.p1.2),[§4\.3](https://arxiv.org/html/2608.05666#S4.SS3.p1.1),[§5](https://arxiv.org/html/2608.05666#S5.SS0.SSS0.Px1.p1.2)\.
- G\. Papamakarios, T\. Pavlakou, and I\. Murray \(2017\)Masked autoregressive flow for density estimation\.InAdvances in Neural Information Processing Systems,Cited by:[§4\.3](https://arxiv.org/html/2608.05666#S4.SS3.p1.1)\.
- G\. Peyré and M\. Cuturi \(2019\)Computational optimal transport\.Foundations and Trends in Machine Learning11\(5–6\),pp\. 355–607\.External Links:[Document](https://dx.doi.org/10.1561/2200000073)Cited by:[§1](https://arxiv.org/html/2608.05666#S1.p2.3),[§5](https://arxiv.org/html/2608.05666#S5.SS0.SSS0.Px1.p1.2)\.
- O\. Russakovsky, J\. Deng, H\. Su, J\. Krause, S\. Satheesh, S\. Ma, Z\. Huang, A\. Karpathy, A\. Khosla, M\. Bernstein, A\. C\. Berg, and L\. Fei\-Fei \(2015\)ImageNet Large Scale Visual Recognition Challenge\.International Journal of Computer Vision \(IJCV\)115\(3\),pp\. 211–252\.External Links:[Document](https://dx.doi.org/10.1007/s11263-015-0816-y)Cited by:[§B\.4](https://arxiv.org/html/2608.05666#A2.SS4.p2.6)\.
- F\. Santambrogio \(2015\)Optimal transport for applied mathematicians: calculus of variations, pdes, and modeling\.Progress in Nonlinear Differential Equations and Their Applications, Vol\.87,Birkhäuser Cham\.External Links:[Document](https://dx.doi.org/10.1007/978-3-319-20828-2)Cited by:[§A\.1](https://arxiv.org/html/2608.05666#A1.SS1.11.p11.21),[§2\.2](https://arxiv.org/html/2608.05666#S2.SS2.p1.1)\.
- A\. Tong, K\. Fatras, N\. Malkin, G\. Huguet, Y\. Zhang, J\. Rector\-Brooks, G\. Wolf, and Y\. Bengio \(2023\)Improving and generalizing flow\-based generative models with minibatch optimal transport\.arXiv preprint arXiv:2302\.00482\.Cited by:[§5](https://arxiv.org/html/2608.05666#S5.SS0.SSS0.Px2.p1.1)\.
- C\. Villaniet al\.\(2009\)Optimal transport: old and new\.Vol\.338,Springer\.Cited by:[§1](https://arxiv.org/html/2608.05666#S1.p2.3),[§5](https://arxiv.org/html/2608.05666#S5.SS0.SSS0.Px1.p1.2)\.
## Appendix ADetails for Theorem[3\.2](https://arxiv.org/html/2608.05666#S3.Thmtheorem2)
### A\.1Proof of Theorem[3\.2](https://arxiv.org/html/2608.05666#S3.Thmtheorem2)
###### Proof\.
\(1\) By assumption,\(ρ⋆,v⋆\)=\(ρΦ⋆,vΦ⋆\)\(\\rho^\{\\star\},v^\{\\star\}\)=\(\\rho^\{\\Phi^\{\\star\}\},v\_\{\\Phi^\{\\star\}\}\)is a solution of thepp\-Benamou–Brenier problem for someΦ⋆∈dom\(ℒPMOT\)\\Phi^\{\\star\}\\in\\operatorname\{dom\}\(\\mathcal\{L\}\_\{\\mathrm\{PMOT\}\}\)\.
The endpoint condition of the Benamou–Brenier problem and the definition of the induced flow give
\(FΦ⋆\)♯μ0=ρ1Φ⋆=ρ1⋆=μ1\.\(F\_\{\\Phi^\{\\star\}\}\)\_\{\\sharp\}\\mu\_\{0\}=\\rho\_\{1\}^\{\\Phi^\{\\star\}\}=\\rho\_\{1\}^\{\\star\}=\\mu\_\{1\}\.Since\(ρΦ⋆,vΦ⋆\)\(\\rho^\{\\Phi^\{\\star\}\},v\_\{\\Phi^\{\\star\}\}\)minimizes the Benamou–Brenier action, the pushforward formula and the characteristic equation yield
1pWpp\(μ0,μ1\)\\displaystyle\\frac\{1\}\{p\}W\_\{p\}^\{p\}\(\\mu\_\{0\},\\mu\_\{1\}\)=1p∫01∫ℝd∥vΦ⋆\(t,y\)∥pρtΦ⋆\(dy\)dt\\displaystyle=\\frac\{1\}\{p\}\\int\_\{0\}^\{1\}\\int\_\{\\mathbb\{R\}^\{d\}\}\\lVert v\_\{\\Phi^\{\\star\}\}\(t,y\)\\rVert^\{p\}\\,\\rho\_\{t\}^\{\\Phi^\{\\star\}\}\(\\mathrm\{d\}y\)\\,\\mathrm\{d\}t=1p∫ℝd∫01‖∂tzΦ⋆\(t,x\)‖pdtμ0\(dx\)\\displaystyle=\\frac\{1\}\{p\}\\int\_\{\\mathbb\{R\}^\{d\}\}\\int\_\{0\}^\{1\}\\left\\lVert\\partial\_\{t\}z\_\{\\Phi^\{\\star\}\}\(t,x\)\\right\\rVert^\{p\}\\,\\mathrm\{d\}t\\,\\mu\_\{0\}\(\\mathrm\{d\}x\)≥1p∫ℝd‖FΦ⋆\(x\)−x‖pμ0\(dx\)\\displaystyle\\geq\\frac\{1\}\{p\}\\int\_\{\\mathbb\{R\}^\{d\}\}\\left\\lVert F\_\{\\Phi^\{\\star\}\}\(x\)\-x\\right\\rVert^\{p\}\\,\\mu\_\{0\}\(\\mathrm\{d\}x\)≥1pWpp\(μ0,μ1\)\.\\displaystyle\\geq\\frac\{1\}\{p\}W\_\{p\}^\{p\}\(\\mu\_\{0\},\\mu\_\{1\}\)\.The first inequality is Jensen’s inequality applied toμ0\\mu\_\{0\}\-almost every trajectory of finitepp\-energy, using
FΦ⋆\(x\)−x=∫01∂tzΦ⋆\(t,x\)dt,F\_\{\\Phi^\{\\star\}\}\(x\)\-x=\\int\_\{0\}^\{1\}\\partial\_\{t\}z\_\{\\Phi^\{\\star\}\}\(t,x\)\\,\\mathrm\{d\}t,and the second follows because\(Id,FΦ⋆\)♯μ0∈Π\(μ0,μ1\)\(\\operatorname\{Id\},F\_\{\\Phi^\{\\star\}\}\)\_\{\\sharp\}\\mu\_\{0\}\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)\. Since the leftmost and rightmost quantities in the preceding chain coincide, equality holds throughout\. In particular,\(Id,FΦ⋆\)♯μ0\(\\operatorname\{Id\},F\_\{\\Phi^\{\\star\}\}\)\_\{\\sharp\}\\mu\_\{0\}is an optimal coupling\. Sinceμ0\\mu\_\{0\}is absolutely continuous andp\>1p\>1, the optimal coupling is induced by the unique optimal Monge mapTpT\_\{p\}, and therefore
FΦ⋆=Tpμ0\-almost everywhere\.F\_\{\\Phi^\{\\star\}\}=T\_\{p\}\\qquad\\mu\_\{0\}\\text\{\-almost everywhere\}\.Moreover, the nonnegative gap in Jensen’s inequality has zero integral with respect toμ0\\mu\_\{0\}\. Thus equality in Jensen’s inequality holds forμ0\\mu\_\{0\}\-almost everyxx\. Sincew↦∥w∥pw\\mapsto\\lVert w\\rVert^\{p\}is strictly convex forp\>1p\>1, this implies
∂tzΦ⋆\(t,x\)=FΦ⋆\(x\)−xfor almost everyt∈\(0,1\)\\partial\_\{t\}z\_\{\\Phi^\{\\star\}\}\(t,x\)=F\_\{\\Phi^\{\\star\}\}\(x\)\-x\\qquad\\text\{for almost every \}t\\in\(0,1\)forμ0\\mu\_\{0\}\-almost everyxx\. SincezΦ⋆\(⋅,x\)z\_\{\\Phi^\{\\star\}\}\(\\cdot,x\)is absolutely continuous andzΦ⋆\(0,x\)=xz\_\{\\Phi^\{\\star\}\}\(0,x\)=x, integration in time gives
zΦ⋆\(t,x\)=\(1−t\)x\+tFΦ⋆\(x\)=\(1−t\)x\+tTp\(x\)z\_\{\\Phi^\{\\star\}\}\(t,x\)=\(1\-t\)x\+tF\_\{\\Phi^\{\\star\}\}\(x\)=\(1\-t\)x\+tT\_\{p\}\(x\)for everyt∈\[0,1\]t\\in\[0,1\]and forμ0\\mu\_\{0\}\-almost everyxx\. Combining this identity with the characteristic equation gives
vΦ⋆\(t,\(1−t\)x\+tFΦ⋆\(x\)\)=FΦ⋆\(x\)−xdt⊗μ0\-almost everywhere\.v\_\{\\Phi^\{\\star\}\}\\bigl\(t,\(1\-t\)x\+tF\_\{\\Phi^\{\\star\}\}\(x\)\\bigr\)=F\_\{\\Phi^\{\\star\}\}\(x\)\-x\\qquad\\mathrm\{d\}t\\otimes\\mu\_\{0\}\\text\{\-almost everywhere\}\.Forr\>1r\>1, letAr\(ξ\):=∥ξ∥r−2ξA\_\{r\}\(\\xi\):=\\lVert\\xi\\rVert^\{r\-2\}\\xi, withAr\(0\):=0A\_\{r\}\(0\):=0\. The duality mapsApA\_\{p\}andAqA\_\{q\}are odd and inverse to each other\. SincevΦ⋆=−Aq\(∇xΦ⋆\)v\_\{\\Phi^\{\\star\}\}=\-A\_\{q\}\(\\nabla\_\{x\}\\Phi^\{\\star\}\), the preceding velocity identity is equivalent to
∇xΦ⋆\(t,\(1−t\)x\+tFΦ⋆\(x\)\)=−∥FΦ⋆\(x\)−x∥p−2\(FΦ⋆\(x\)−x\)dt⊗μ0\-almost everywhere\.\\nabla\_\{x\}\\Phi^\{\\star\}\\bigl\(t,\(1\-t\)x\+tF\_\{\\Phi^\{\\star\}\}\(x\)\\bigr\)=\-\\lVert F\_\{\\Phi^\{\\star\}\}\(x\)\-x\\rVert^\{p\-2\}\\bigl\(F\_\{\\Phi^\{\\star\}\}\(x\)\-x\\bigr\)\\qquad\\mathrm\{d\}t\\otimes\\mu\_\{0\}\\text\{\-almost everywhere\}\.Thus the potential\-matching residual vanishes andℒPM\(Φ⋆\)=0\\mathcal\{L\}\_\{\\mathrm\{PM\}\}\(\\Phi^\{\\star\}\)=0\. The endpoint identity above also givesℒterm\(Φ⋆\)=0\\mathcal\{L\}\_\{\\mathrm\{term\}\}\(\\Phi^\{\\star\}\)=0\. Since the two terms in the PMOT objective are nonnegative and have positive weights,
minΦ∈dom\(ℒPMOT\)ℒPMOT\(Φ\)=ℒPMOT\(Φ⋆\)=0\.\\min\_\{\\Phi\\in\\operatorname\{dom\}\(\\mathcal\{L\}\_\{\\mathrm\{PMOT\}\}\)\}\\mathcal\{L\}\_\{\\mathrm\{PMOT\}\}\(\\Phi\)=\\mathcal\{L\}\_\{\\mathrm\{PMOT\}\}\(\\Phi^\{\\star\}\)=0\.
\(2\) Now letΦ\\Phibe a global minimizer satisfying the regularity and full\-support assumptions of the theorem\. By part \(1\), the minimum value of the PMOT objective is zero\. Hence
ℒPM\(Φ\)=ℒterm\(Φ\)=0,\\mathcal\{L\}\_\{\\mathrm\{PM\}\}\(\\Phi\)=\\mathcal\{L\}\_\{\\mathrm\{term\}\}\(\\Phi\)=0,and in particular,\(FΦ\)\#μ0=μ1\(F\_\{\\Phi\}\)\_\{\\\#\}\\mu\_\{0\}=\\mu\_\{1\}\.
Define
St\(x\):=\(1−t\)x\+tFΦ\(x\)S\_\{t\}\(x\):=\(1\-t\)x\+tF\_\{\\Phi\}\(x\)and the potential\-matching residual
R\(t,x\):=∇xΦ\(t,St\(x\)\)\+∥FΦ\(x\)−x∥p−2\(FΦ\(x\)−x\),\(t,x\)∈\(0,1\)×ℝd\.R\(t,x\):=\\nabla\_\{x\}\\Phi\\bigl\(t,S\_\{t\}\(x\)\\bigr\)\+\\lVert F\_\{\\Phi\}\(x\)\-x\\rVert^\{p\-2\}\\bigl\(F\_\{\\Phi\}\(x\)\-x\\bigr\),\\qquad\(t,x\)\\in\(0,1\)\\times\\mathbb\{R\}^\{d\}\.Since the solution mapzΦz\_\{\\Phi\}is jointly Borel measurable, its terminal\-time sliceFΦ=zΦ\(1,⋅\)F\_\{\\Phi\}=z\_\{\\Phi\}\(1,\\cdot\)is Borel measurable\. Consequently, the map\(t,x\)↦St\(x\)\(t,x\)\\mapsto S\_\{t\}\(x\)is jointly Borel measurable\. Moreover,∇xΦ\\nabla\_\{x\}\\Phiis continuous on\(0,1\)×ℝd\(0,1\)\\times\\mathbb\{R\}^\{d\}, and the mapξ↦∥ξ∥p−2ξ\\xi\\mapsto\\lVert\\xi\\rVert^\{p\-2\}\\xi, with value0atξ=0\\xi=0, is continuous forp\>1p\>1\. It follows thatRRis jointly Borel measurable\.
The function\(t,x\)↦∥R\(t,x\)∥2\(t,x\)\\mapsto\\lVert R\(t,x\)\\rVert^\{2\}is therefore nonnegative and jointly Borel measurable\. By Tonelli’s theorem, the function
h\(x\):=∫01∥R\(t,x\)∥2dth\(x\):=\\int\_\{0\}^\{1\}\\lVert R\(t,x\)\\rVert^\{2\}\\,\\mathrm\{d\}tis measurable\. SinceℒPM\(Φ\)=0\\mathcal\{L\}\_\{\\mathrm\{PM\}\}\(\\Phi\)=0,
0=∫ℝdh\(x\)μ0\(dx\)\.0=\\int\_\{\\mathbb\{R\}^\{d\}\}h\(x\)\\,\\mu\_\{0\}\(\\mathrm\{d\}x\)\.Becauseh≥0h\\geq 0, it follows thath=0h=0forμ0\\mu\_\{0\}\-almost everyxx\. Hence the measurable set
E:=\{x∈ℝd:h\(x\)=0\}E:=\\\{x\\in\\mathbb\{R\}^\{d\}:h\(x\)=0\\\}satisfiesμ0\(E\)=1\\mu\_\{0\}\(E\)=1\. For everyx∈Ex\\in E, we have
∫01∥R\(t,x\)∥2dt=0,\\int\_\{0\}^\{1\}\\lVert R\(t,x\)\\rVert^\{2\}\\,\\mathrm\{d\}t=0,and therefore
R\(t,x\)=0for almost everyt∈\(0,1\)\.R\(t,x\)=0\\qquad\\text\{for almost every \}t\\in\(0,1\)\.For each fixedx∈Ex\\in E, the mapt↦R\(t,x\)t\\mapsto R\(t,x\)is continuous on\(0,1\)\(0,1\)\. Since it vanishes for almost everyt∈\(0,1\)t\\in\(0,1\), it must vanish for everyt∈\(0,1\)t\\in\(0,1\)\.
Thus, for everyx∈Ex\\in Eandt∈\(0,1\)t\\in\(0,1\),
∇xΦ\(t,St\(x\)\)=−Ap\(FΦ\(x\)−x\)\.\\nabla\_\{x\}\\Phi\\bigl\(t,S\_\{t\}\(x\)\\bigr\)=\-A\_\{p\}\\bigl\(F\_\{\\Phi\}\(x\)\-x\\bigr\)\.UsingvΦ=−Aq\(∇xΦ\)v\_\{\\Phi\}=\-A\_\{q\}\(\\nabla\_\{x\}\\Phi\), together with the oddness and inverse relation of the duality maps defined in part \(1\), gives
vΦ\(t,St\(x\)\)=−Aq\(−Ap\(FΦ\(x\)−x\)\)=FΦ\(x\)−x\.v\_\{\\Phi\}\\bigl\(t,S\_\{t\}\(x\)\\bigr\)=\-A\_\{q\}\\bigl\(\-A\_\{p\}\(F\_\{\\Phi\}\(x\)\-x\)\\bigr\)=F\_\{\\Phi\}\(x\)\-x\.
Since
∂tSt\(x\)=FΦ\(x\)−x,\\partial\_\{t\}S\_\{t\}\(x\)=F\_\{\\Phi\}\(x\)\-x,it follows that, for everyx∈Ex\\in E,
∂tSt\(x\)=vΦ\(t,St\(x\)\),t∈\(0,1\),S0\(x\)=x\.\\partial\_\{t\}S\_\{t\}\(x\)=v\_\{\\Phi\}\\bigl\(t,S\_\{t\}\(x\)\\bigr\),\\qquad t\\in\(0,1\),\\qquad S\_\{0\}\(x\)=x\.SinceS⋅\(x\)S\_\{\\cdot\}\(x\)is affine in time, it is absolutely continuous on\[0,1\]\[0,1\]and satisfies the characteristic equation on\(0,1\)\(0,1\)\. HenceS⋅\(x\)S\_\{\\cdot\}\(x\)solves problem equation[2](https://arxiv.org/html/2608.05666#S3.E2), and uniqueness gives
zΦ\(t,x\)=St\(x\)for everyt∈\[0,1\]and everyx∈E\.z\_\{\\Phi\}\(t,x\)=S\_\{t\}\(x\)\\qquad\\text\{for every \}t\\in\[0,1\]\\text\{ and every \}x\\in E\.In particular,
vΦ\(t,zΦ\(t,x\)\)=FΦ\(x\)−xfor everyt∈\(0,1\)and everyx∈E\.v\_\{\\Phi\}\\bigl\(t,z\_\{\\Phi\}\(t,x\)\\bigr\)=F\_\{\\Phi\}\(x\)\-x\\qquad\\text\{for every \}t\\in\(0,1\)\\text\{ and every \}x\\in E\.Thus, forμ0\\mu\_\{0\}\-almost every initial point, the induced trajectory is a constant\-speed straight line and the velocity is constant along it\.
It remains to show that the pair\(ρtΦ,vΦ\)\(\\rho\_\{t\}^\{\\Phi\},v\_\{\\Phi\}\)induced by these straight\-line trajectories is a minimizer of thepp\-Benamou–Brenier problem\. We first prove thatFΦF\_\{\\Phi\}is the optimal map for the staticpp\-cost transport problem and then verify that the induced straight\-line flow attains the minimum Benamou–Brenier action\. Set
Q:=\(0,1\)×ℝd,q=pp−1,Hq\(ξ\)=1q∥ξ∥q,Aq\(ξ\)=∇Hq\(ξ\)=∥ξ∥q−2ξ,Q:=\(0,1\)\\times\\mathbb\{R\}^\{d\},\\qquad q=\\frac\{p\}\{p\-1\},\\qquad H\_\{q\}\(\\xi\)=\\frac\{1\}\{q\}\\lVert\\xi\\rVert^\{q\},\\qquad A\_\{q\}\(\\xi\)=\\nabla H\_\{q\}\(\\xi\)=\\lVert\\xi\\rVert^\{q\-2\}\\xi,and writeg=∇xΦg=\\nabla\_\{x\}\\Phi, so thatvΦ=−Aq\(g\)v\_\{\\Phi\}=\-A\_\{q\}\(g\)\. Define
G\(t,x\):=∂tΦ\(t,x\)−Hq\(g\(t,x\)\),\(t,x\)∈Q\.G\(t,x\):=\\partial\_\{t\}\\Phi\(t,x\)\-H\_\{q\}\(g\(t,x\)\),\\qquad\(t,x\)\\in Q\.We claim that∇xG=0\\nabla\_\{x\}G=0onQQ\. SinceΦ∈C2\(Q\)\\Phi\\in C^\{2\}\(Q\), the fieldg=∇xΦg=\\nabla\_\{x\}\\PhiisC1C^\{1\}onQQ\. Moreover, for everyq\>1q\>1, the mapAq\(ξ\)=∥ξ∥q−2ξA\_\{q\}\(\\xi\)=\\lVert\\xi\\rVert^\{q\-2\}\\xiisC1C^\{1\}onℝd∖\{0\}\\mathbb\{R\}^\{d\}\\setminus\\\{0\\\}, with
DAq\(ξ\)=∥ξ∥q−2I\+\(q−2\)∥ξ∥q−4ξ⊗ξ\.DA\_\{q\}\(\\xi\)=\\lVert\\xi\\rVert^\{q\-2\}I\+\(q\-2\)\\lVert\\xi\\rVert^\{q\-4\}\\xi\\otimes\\xi\.Sinceggis continuous, the set
U:=\{\(t,x\)∈Q:g\(t,x\)≠0\}U:=\\\{\(t,x\)\\in Q:g\(t,x\)\\neq 0\\\}is open\. Therefore,vΦ=−Aq∘gv\_\{\\Phi\}=\-A\_\{q\}\\circ gisC1C^\{1\}onUUby the chain rule\. For everyx∈Ex\\in E, the velocity is constant along the corresponding characteristic:
vΦ\(t,zΦ\(t,x\)\)=FΦ\(x\)−x,t∈\(0,1\)\.v\_\{\\Phi\}\\bigl\(t,z\_\{\\Phi\}\(t,x\)\\bigr\)=F\_\{\\Phi\}\(x\)\-x,\\qquad t\\in\(0,1\)\.By the straight\-line identity,zΦ\(⋅,x\)=S⋅\(x\)z\_\{\\Phi\}\(\\cdot,x\)=S\_\{\\cdot\}\(x\)is affine and thereforeC1C^\{1\}on\(0,1\)\(0,1\)\. SinceUUis open andzΦ\(⋅,x\)z\_\{\\Phi\}\(\\cdot,x\)is continuous, wheneverx∈Ex\\in Eand\(t,zΦ\(t,x\)\)∈U\(t,z\_\{\\Phi\}\(t,x\)\)\\in U, the trajectory remains inUUon a neighborhood oftt\. TheC1C^\{1\}chain rule therefore applies on this neighborhood, and differentiating the preceding identity with respect tottand using∂tzΦ\(t,x\)=vΦ\(t,zΦ\(t,x\)\)\\partial\_\{t\}z\_\{\\Phi\}\(t,x\)=v\_\{\\Phi\}\(t,z\_\{\\Phi\}\(t,x\)\)gives
0\\displaystyle 0=ddtvΦ\(t,zΦ\(t,x\)\)\\displaystyle=\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}v\_\{\\Phi\}\\bigl\(t,z\_\{\\Phi\}\(t,x\)\\bigr\)=∂tvΦ\(t,zΦ\(t,x\)\)\+DxvΦ\(t,zΦ\(t,x\)\)∂tzΦ\(t,x\)\\displaystyle=\\partial\_\{t\}v\_\{\\Phi\}\\bigl\(t,z\_\{\\Phi\}\(t,x\)\\bigr\)\+D\_\{x\}v\_\{\\Phi\}\\bigl\(t,z\_\{\\Phi\}\(t,x\)\\bigr\)\\,\\partial\_\{t\}z\_\{\\Phi\}\(t,x\)=\[∂tvΦ\+DxvΦvΦ\]\(t,zΦ\(t,x\)\)\.\\displaystyle=\\bigl\[\\partial\_\{t\}v\_\{\\Phi\}\+D\_\{x\}v\_\{\\Phi\}\\,v\_\{\\Phi\}\\bigr\]\\bigl\(t,z\_\{\\Phi\}\(t,x\)\\bigr\)\.Define the space–time measure as the pushforward
ν:=\(\(t,x\)↦\(t,zΦ\(t,x\)\)\)\#\(dt⊗μ0\)\.\\nu:=\\bigl\(\(t,x\)\\mapsto\(t,z\_\{\\Phi\}\(t,x\)\)\\bigr\)\_\{\\\#\}\\bigl\(\\mathrm\{d\}t\\otimes\\mu\_\{0\}\\bigr\)\.The joint Borel measurability ofzΦz\_\{\\Phi\}ensures that this measure is well defined\. SinceρtΦ=\(zΦ\(t,⋅\)\)\#μ0\\rho\_\{t\}^\{\\Phi\}=\(z\_\{\\Phi\}\(t,\\cdot\)\)\_\{\\\#\}\\mu\_\{0\}, it can equivalently be written as
ν\(dt,dy\)=dtρtΦ\(dy\)\.\\nu\(\\mathrm\{d\}t,\\mathrm\{d\}y\)=\\mathrm\{d\}t\\,\\rho\_\{t\}^\{\\Phi\}\(\\mathrm\{d\}y\)\.Sinceμ0\(E\)=1\\mu\_\{0\}\(E\)=1, the preceding identity implies
∂tvΦ\+DxvΦvΦ=0ν\-almost everywhere onU\.\\partial\_\{t\}v\_\{\\Phi\}\+D\_\{x\}v\_\{\\Phi\}\\,v\_\{\\Phi\}=0\\qquad\\nu\\text\{\-almost everywhere on \}U\.The assumptionsupp\(ρtΦ\)=ℝd\\operatorname\{supp\}\(\\rho\_\{t\}^\{\\Phi\}\)=\\mathbb\{R\}^\{d\}for everyt∈\(0,1\)t\\in\(0,1\)implies thatν\\nuhas full support onQQ\. Indeed, every nonempty open subset ofQQcontains a productI×BI\\times B, whereI⊂\(0,1\)I\\subset\(0,1\)andB⊂ℝdB\\subset\\mathbb\{R\}^\{d\}are nonempty and open, and
ν\(I×B\)=∫IρtΦ\(B\)dt\>0\.\\nu\(I\\times B\)=\\int\_\{I\}\\rho\_\{t\}^\{\\Phi\}\(B\)\\,\\mathrm\{d\}t\>0\.Since the map
\(t,x\)⟼∂tvΦ\(t,x\)\+DxvΦ\(t,x\)vΦ\(t,x\)\(t,x\)\\longmapsto\\partial\_\{t\}v\_\{\\Phi\}\(t,x\)\+D\_\{x\}v\_\{\\Phi\}\(t,x\)v\_\{\\Phi\}\(t,x\)is continuous onUU, it follows that
\(∂t\+vΦ⋅∇x\)vΦ=∂tvΦ\+DxvΦvΦ=0onU\.\(\\partial\_\{t\}\+v\_\{\\Phi\}\\cdot\\nabla\_\{x\}\)v\_\{\\Phi\}=\\partial\_\{t\}v\_\{\\Phi\}\+D\_\{x\}v\_\{\\Phi\}\\,v\_\{\\Phi\}=0\\qquad\\text\{on \}U\.To relate the material derivative ofvΦv\_\{\\Phi\}to∇xG\\nabla\_\{x\}G, note thatDxg=Dx2ΦD\_\{x\}g=D\_\{x\}^\{2\}\\Phi,DAq=D2HqDA\_\{q\}=D^\{2\}H\_\{q\}, and\(vΦ⋅∇x\)vΦ=DxvΦvΦ\(v\_\{\\Phi\}\\cdot\\nabla\_\{x\}\)v\_\{\\Phi\}=D\_\{x\}v\_\{\\Phi\}\\,v\_\{\\Phi\}\. Hence, the chain rule yields
\(∂t\+vΦ⋅∇x\)vΦ\\displaystyle\(\\partial\_\{t\}\+v\_\{\\Phi\}\\cdot\\nabla\_\{x\}\)v\_\{\\Phi\}=∂tvΦ\+DxvΦvΦ\\displaystyle=\\partial\_\{t\}v\_\{\\Phi\}\+D\_\{x\}v\_\{\\Phi\}\\,v\_\{\\Phi\}=−DAq\(g\)∂tg−DAq\(g\)DxgvΦ\\displaystyle=\-DA\_\{q\}\(g\)\\,\\partial\_\{t\}g\-DA\_\{q\}\(g\)D\_\{x\}g\\,v\_\{\\Phi\}=−D2Hq\(g\)\(∂tg\+Dx2ΦvΦ\)\.\\displaystyle=\-D^\{2\}H\_\{q\}\(g\)\\bigl\(\\partial\_\{t\}g\+D\_\{x\}^\{2\}\\Phi\\,v\_\{\\Phi\}\\bigr\)\.On the other hand, using the symmetry ofDx2ΦD\_\{x\}^\{2\}\\PhiandvΦ=−Aq\(g\)v\_\{\\Phi\}=\-A\_\{q\}\(g\),
∇xG\\displaystyle\\nabla\_\{x\}G=∂tg−∇x\(Hq\(g\)\)\\displaystyle=\\partial\_\{t\}g\-\\nabla\_\{x\}\\bigl\(H\_\{q\}\(g\)\\bigr\)=∂tg−\(Dxg\)⊤Aq\(g\)\\displaystyle=\\partial\_\{t\}g\-\(D\_\{x\}g\)^\{\\top\}A\_\{q\}\(g\)=∂tg−Dx2ΦAq\(g\)\\displaystyle=\\partial\_\{t\}g\-D\_\{x\}^\{2\}\\Phi\\,A\_\{q\}\(g\)=∂tg\+Dx2ΦvΦ\.\\displaystyle=\\partial\_\{t\}g\+D\_\{x\}^\{2\}\\Phi\\,v\_\{\\Phi\}\.Combining the preceding two identities yields
0=\(∂t\+vΦ⋅∇x\)vΦ=−D2Hq\(g\)∇xGonU\.0=\(\\partial\_\{t\}\+v\_\{\\Phi\}\\cdot\\nabla\_\{x\}\)v\_\{\\Phi\}=\-D^\{2\}H\_\{q\}\(g\)\\nabla\_\{x\}G\\qquad\\text\{on \}U\.Forg≠0g\\neq 0,
D2Hq\(g\)=∥g∥q−2I\+\(q−2\)∥g∥q−4g⊗gD^\{2\}H\_\{q\}\(g\)=\\lVert g\\rVert^\{q\-2\}I\+\(q\-2\)\\lVert g\\rVert^\{q\-4\}g\\otimes gis positive definite\. Indeed, for anyη∈ℝd\\eta\\in\\mathbb\{R\}^\{d\}, define its components parallel and orthogonal toggby
η∥:=η⋅g∥g∥2g,η⟂:=η−η∥\.\\eta\_\{\\parallel\}:=\\frac\{\\eta\\cdot g\}\{\\lVert g\\rVert^\{2\}\}g,\\qquad\\eta\_\{\\perp\}:=\\eta\-\\eta\_\{\\parallel\}\.Thenη=η⟂\+η∥\\eta=\\eta\_\{\\perp\}\+\\eta\_\{\\parallel\},η⟂⋅g=0\\eta\_\{\\perp\}\\cdot g=0, and\(g⋅η\)2=∥g∥2∥η∥∥2\(g\\cdot\\eta\)^\{2\}=\\lVert g\\rVert^\{2\}\\lVert\\eta\_\{\\parallel\}\\rVert^\{2\}\. Hence
η⊤D2Hq\(g\)η\\displaystyle\\eta^\{\\top\}D^\{2\}H\_\{q\}\(g\)\\eta=∥g∥q−2∥η∥2\+\(q−2\)∥g∥q−4\(g⋅η\)2\\displaystyle=\\lVert g\\rVert^\{q\-2\}\\lVert\\eta\\rVert^\{2\}\+\(q\-2\)\\lVert g\\rVert^\{q\-4\}\(g\\cdot\\eta\)^\{2\}=∥g∥q−2\(∥η⟂∥2\+∥η∥∥2\+\(q−2\)∥η∥∥2\)\\displaystyle=\\lVert g\\rVert^\{q\-2\}\\bigl\(\\lVert\\eta\_\{\\perp\}\\rVert^\{2\}\+\\lVert\\eta\_\{\\parallel\}\\rVert^\{2\}\+\(q\-2\)\\lVert\\eta\_\{\\parallel\}\\rVert^\{2\}\\bigr\)=∥g∥q−2\(∥η⟂∥2\+\(q−1\)∥η∥∥2\)\>0\\displaystyle=\\lVert g\\rVert^\{q\-2\}\\bigl\(\\lVert\\eta\_\{\\perp\}\\rVert^\{2\}\+\(q\-1\)\\lVert\\eta\_\{\\parallel\}\\rVert^\{2\}\\bigr\)\>0for everyη≠0\\eta\\neq 0, sinceg≠0g\\neq 0andq\>1q\>1\. Equivalently, the tangential eigenvalue is∥g∥q−2\\lVert g\\rVert^\{q\-2\}and the radial eigenvalue is\(q−1\)∥g∥q−2\(q\-1\)\\lVert g\\rVert^\{q\-2\}\. Therefore∇xG=0\\nabla\_\{x\}G=0onUU\.
It remains to consider the zero set
Z:=\{\(t,x\)∈Q:g\(t,x\)=0\}\.Z:=\\\{\(t,x\)\\in Q:g\(t,x\)=0\\\}\.At every interior point ofZZrelative toQQ, the fieldggvanishes in a neighborhood, so∂tg=0\\partial\_\{t\}g=0there and hence∇xG=0\\nabla\_\{x\}G=0\. Every other point ofZZis a limit point ofUU, and the continuity onQQof
∇xG=∂tg−Dx2ΦAq\(g\)\\nabla\_\{x\}G=\\partial\_\{t\}g\-D\_\{x\}^\{2\}\\Phi\\,A\_\{q\}\(g\)again yields∇xG=0\\nabla\_\{x\}G=0\. Thus∇xG=0\\nabla\_\{x\}G=0onZZas well\. Together with the conclusion onUU, this yields
∇xG=0onQ\.\\nabla\_\{x\}G=0\\qquad\\text\{on \}Q\.
Sinceℝd\\mathbb\{R\}^\{d\}is connected, there is a continuous functiona:\(0,1\)→ℝa:\(0,1\)\\to\\mathbb\{R\}such that
G\(t,x\)=a\(t\)\(t,x\)∈Q\.G\(t,x\)=a\(t\)\\qquad\(t,x\)\\in Q\.Define
α\(t\):=∫1/2ta\(s\)ds,t∈\(0,1\)\.\\alpha\(t\):=\\int\_\{1/2\}^\{t\}a\(s\)\\,\\mathrm\{d\}s,\\qquad t\\in\(0,1\)\.Thenα′\(t\)=a\(t\)\\alpha^\{\\prime\}\(t\)=a\(t\)on\(0,1\)\(0,1\)\. The gauge\-adjusted potential
Φ~\(t,x\):=Φ\(t,x\)−α\(t\)\\widetilde\{\\Phi\}\(t,x\):=\\Phi\(t,x\)\-\\alpha\(t\)induces the same velocity on\(0,1\)×ℝd\(0,1\)\\times\\mathbb\{R\}^\{d\}, and satisfies the Hamilton–Jacobi equation
∂tΦ~=Hq\(∇xΦ~\)onQ\.\\partial\_\{t\}\\widetilde\{\\Phi\}=H\_\{q\}\(\\nabla\_\{x\}\\widetilde\{\\Phi\}\)\\qquad\\text\{on \}Q\.Letcp\(x,y\)=∥x−y∥p/pc\_\{p\}\(x,y\)=\\lVert x\-y\\rVert^\{p\}/p\. Fenchel’s inequality, applied to the convex conjugate pairLp\(w\)=∥w∥p/pL\_\{p\}\(w\)=\\lVert w\\rVert^\{p\}/pandHq\(ξ\)=∥ξ∥q/qH\_\{q\}\(\\xi\)=\\lVert\\xi\\rVert^\{q\}/q, reads
Lp\(w\)\+Hq\(ξ\)\+w⋅ξ≥0,L\_\{p\}\(w\)\+H\_\{q\}\(\\xi\)\+w\\cdot\\xi\\geq 0,with equality if and only ifw=−Aq\(ξ\)w=\-A\_\{q\}\(\\xi\)\. Fix0<ε<1/20<\\varepsilon<1/2\. For everyx∈Ex\\in E,
FΦ\(x\)−x=−Aq\(∇xΦ~\(t,St\(x\)\)\),t∈\(0,1\),F\_\{\\Phi\}\(x\)\-x=\-A\_\{q\}\\bigl\(\\nabla\_\{x\}\\widetilde\{\\Phi\}\(t,S\_\{t\}\(x\)\)\\bigr\),\\qquad t\\in\(0,1\),so equality holds in Fenchel’s inequality along the straight\-line trajectorySt\(x\)S\_\{t\}\(x\)\. Since∂tSt\(x\)=FΦ\(x\)−x\\partial\_\{t\}S\_\{t\}\(x\)=F\_\{\\Phi\}\(x\)\-x, the chain rule and the Hamilton–Jacobi equation give
ddtΦ~\(t,St\(x\)\)\\displaystyle\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\widetilde\{\\Phi\}\\bigl\(t,S\_\{t\}\(x\)\\bigr\)=∂tΦ~\(t,St\(x\)\)\+∇xΦ~\(t,St\(x\)\)⋅∂tSt\(x\)\\displaystyle=\\partial\_\{t\}\\widetilde\{\\Phi\}\\bigl\(t,S\_\{t\}\(x\)\\bigr\)\+\\nabla\_\{x\}\\widetilde\{\\Phi\}\\bigl\(t,S\_\{t\}\(x\)\\bigr\)\\cdot\\partial\_\{t\}S\_\{t\}\(x\)=Hq\(∇xΦ~\(t,St\(x\)\)\)\+∇xΦ~\(t,St\(x\)\)⋅\(FΦ\(x\)−x\)\\displaystyle=H\_\{q\}\\bigl\(\\nabla\_\{x\}\\widetilde\{\\Phi\}\(t,S\_\{t\}\(x\)\)\\bigr\)\+\\nabla\_\{x\}\\widetilde\{\\Phi\}\\bigl\(t,S\_\{t\}\(x\)\\bigr\)\\cdot\\bigl\(F\_\{\\Phi\}\(x\)\-x\\bigr\)=−Lp\(FΦ\(x\)−x\)\\displaystyle=\-L\_\{p\}\\bigl\(F\_\{\\Phi\}\(x\)\-x\\bigr\)=−cp\(x,FΦ\(x\)\)\.\\displaystyle=\-c\_\{p\}\\bigl\(x,F\_\{\\Phi\}\(x\)\\bigr\)\.Therefore, the fundamental theorem of calculus on\[ε,1−ε\]\[\\varepsilon,1\-\\varepsilon\]yields
Φ~\(1−ε,S1−ε\(x\)\)−Φ~\(ε,Sε\(x\)\)\\displaystyle\\widetilde\{\\Phi\}\\bigl\(1\-\\varepsilon,S\_\{1\-\\varepsilon\}\(x\)\\bigr\)\-\\widetilde\{\\Phi\}\\bigl\(\\varepsilon,S\_\{\\varepsilon\}\(x\)\\bigr\)=∫ε1−εddtΦ~\(t,St\(x\)\)dt=−\(1−2ε\)cp\(x,FΦ\(x\)\)\.\\displaystyle\\qquad=\\int\_\{\\varepsilon\}^\{1\-\\varepsilon\}\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\widetilde\{\\Phi\}\\bigl\(t,S\_\{t\}\(x\)\\bigr\)\\,\\mathrm\{d\}t=\-\(1\-2\\varepsilon\)c\_\{p\}\\bigl\(x,F\_\{\\Phi\}\(x\)\\bigr\)\.Equivalently,
Φ~\(ε,Sε\(x\)\)−Φ~\(1−ε,S1−ε\(x\)\)=\(1−2ε\)cp\(x,FΦ\(x\)\)=∥S1−ε\(x\)−Sε\(x\)∥pp\(1−2ε\)p−1\.\\widetilde\{\\Phi\}\\bigl\(\\varepsilon,S\_\{\\varepsilon\}\(x\)\\bigr\)\-\\widetilde\{\\Phi\}\\bigl\(1\-\\varepsilon,S\_\{1\-\\varepsilon\}\(x\)\\bigr\)=\(1\-2\\varepsilon\)c\_\{p\}\\bigl\(x,F\_\{\\Phi\}\(x\)\\bigr\)=\\frac\{\\lVert S\_\{1\-\\varepsilon\}\(x\)\-S\_\{\\varepsilon\}\(x\)\\rVert^\{p\}\}\{p\(1\-2\\varepsilon\)^\{p\-1\}\}\.
Take finitely many pointsx1,…,xN∈Ex\_\{1\},\\ldots,x\_\{N\}\\in E, setyi:=FΦ\(xi\)y\_\{i\}:=F\_\{\\Phi\}\(x\_\{i\}\), and letσ\\sigmabe any permutation of\{1,…,N\}\\\{1,\\ldots,N\\\}\. For eachii, define the constant\-speed line
γi\(t\):=Sε\(xi\)\+t−ε1−2ε\(S1−ε\(xσ\(i\)\)−Sε\(xi\)\),t∈\[ε,1−ε\]\.\\gamma\_\{i\}\(t\):=S\_\{\\varepsilon\}\(x\_\{i\}\)\+\\frac\{t\-\\varepsilon\}\{1\-2\\varepsilon\}\\left\(S\_\{1\-\\varepsilon\}\(x\_\{\\sigma\(i\)\}\)\-S\_\{\\varepsilon\}\(x\_\{i\}\)\\right\),\\qquad t\\in\[\\varepsilon,1\-\\varepsilon\]\.Its velocity is
∂tγi\(t\)=S1−ε\(xσ\(i\)\)−Sε\(xi\)1−2ε\.\\partial\_\{t\}\\gamma\_\{i\}\(t\)=\\frac\{S\_\{1\-\\varepsilon\}\(x\_\{\\sigma\(i\)\}\)\-S\_\{\\varepsilon\}\(x\_\{i\}\)\}\{1\-2\\varepsilon\}\.By the chain rule, the Hamilton–Jacobi equation, and Fenchel’s inequality,
ddtΦ~\(t,γi\(t\)\)\\displaystyle\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\widetilde\{\\Phi\}\\bigl\(t,\\gamma\_\{i\}\(t\)\\bigr\)=Hq\(∇xΦ~\(t,γi\(t\)\)\)\+∇xΦ~\(t,γi\(t\)\)⋅∂tγi\(t\)≥−Lp\(∂tγi\(t\)\)\.\\displaystyle=H\_\{q\}\\bigl\(\\nabla\_\{x\}\\widetilde\{\\Phi\}\(t,\\gamma\_\{i\}\(t\)\)\\bigr\)\+\\nabla\_\{x\}\\widetilde\{\\Phi\}\(t,\\gamma\_\{i\}\(t\)\)\\cdot\\partial\_\{t\}\\gamma\_\{i\}\(t\)\\geq\-L\_\{p\}\\bigl\(\\partial\_\{t\}\\gamma\_\{i\}\(t\)\\bigr\)\.Integrating over\[ε,1−ε\]\[\\varepsilon,1\-\\varepsilon\]and using the fact that∂tγi\\partial\_\{t\}\\gamma\_\{i\}is constant gives
Φ~\(ε,Sε\(xi\)\)−Φ~\(1−ε,S1−ε\(xσ\(i\)\)\)≤\(1−2ε\)Lp\(∂tγi\)=∥S1−ε\(xσ\(i\)\)−Sε\(xi\)∥pp\(1−2ε\)p−1\.\\displaystyle\\widetilde\{\\Phi\}\\bigl\(\\varepsilon,S\_\{\\varepsilon\}\(x\_\{i\}\)\\bigr\)\-\\widetilde\{\\Phi\}\\bigl\(1\-\\varepsilon,S\_\{1\-\\varepsilon\}\(x\_\{\\sigma\(i\)\}\)\\bigr\)\\leq\(1\-2\\varepsilon\)L\_\{p\}\\bigl\(\\partial\_\{t\}\\gamma\_\{i\}\\bigr\)=\\frac\{\\lVert S\_\{1\-\\varepsilon\}\(x\_\{\\sigma\(i\)\}\)\-S\_\{\\varepsilon\}\(x\_\{i\}\)\\rVert^\{p\}\}\{p\(1\-2\\varepsilon\)^\{p\-1\}\}\.The preceding equality, invariance of the sum of the terminal potential values under permutation, and these cross\-pair inequalities give
∑i=1N∥S1−ε\(xi\)−Sε\(xi\)∥pp\(1−2ε\)p−1\\displaystyle\\sum\_\{i=1\}^\{N\}\\frac\{\\lVert S\_\{1\-\\varepsilon\}\(x\_\{i\}\)\-S\_\{\\varepsilon\}\(x\_\{i\}\)\\rVert^\{p\}\}\{p\(1\-2\\varepsilon\)^\{p\-1\}\}=∑i=1N\[Φ~\(ε,Sε\(xi\)\)−Φ~\(1−ε,S1−ε\(xi\)\)\]\\displaystyle=\\sum\_\{i=1\}^\{N\}\\left\[\\widetilde\{\\Phi\}\\bigl\(\\varepsilon,S\_\{\\varepsilon\}\(x\_\{i\}\)\\bigr\)\-\\widetilde\{\\Phi\}\\bigl\(1\-\\varepsilon,S\_\{1\-\\varepsilon\}\(x\_\{i\}\)\\bigr\)\\right\]=∑i=1N\[Φ~\(ε,Sε\(xi\)\)−Φ~\(1−ε,S1−ε\(xσ\(i\)\)\)\]\\displaystyle=\\sum\_\{i=1\}^\{N\}\\left\[\\widetilde\{\\Phi\}\\bigl\(\\varepsilon,S\_\{\\varepsilon\}\(x\_\{i\}\)\\bigr\)\-\\widetilde\{\\Phi\}\\bigl\(1\-\\varepsilon,S\_\{1\-\\varepsilon\}\(x\_\{\\sigma\(i\)\}\)\\bigr\)\\right\]≤∑i=1N∥S1−ε\(xσ\(i\)\)−Sε\(xi\)∥pp\(1−2ε\)p−1\.\\displaystyle\\leq\\sum\_\{i=1\}^\{N\}\\frac\{\\lVert S\_\{1\-\\varepsilon\}\(x\_\{\\sigma\(i\)\}\)\-S\_\{\\varepsilon\}\(x\_\{i\}\)\\rVert^\{p\}\}\{p\(1\-2\\varepsilon\)^\{p\-1\}\}\.Asε→0\+\\varepsilon\\to 0^\{\+\},
Sε\(xi\)⟶xi,S1−ε\(xi\)⟶yi\.S\_\{\\varepsilon\}\(x\_\{i\}\)\\longrightarrow x\_\{i\},\\qquad S\_\{1\-\\varepsilon\}\(x\_\{i\}\)\\longrightarrow y\_\{i\}\.Taking the limit in the preceding finite\-sum inequality yields
∑i=1Ncp\(xi,yi\)≤∑i=1Ncp\(xi,yσ\(i\)\)\.\\sum\_\{i=1\}^\{N\}c\_\{p\}\(x\_\{i\},y\_\{i\}\)\\leq\\sum\_\{i=1\}^\{N\}c\_\{p\}\(x\_\{i\},y\_\{\\sigma\(i\)\}\)\.Therefore,
Γ:=\{\(x,FΦ\(x\)\):x∈E\}\\Gamma:=\\left\\\{\(x,F\_\{\\Phi\}\(x\)\):x\\in E\\right\\\}iscpc\_\{p\}\-cyclically monotone\. Sinceℒterm\(Φ\)=0\\mathcal\{L\}\_\{\\mathrm\{term\}\}\(\\Phi\)=0, the measure
πΦ:=\(Id,FΦ\)\#μ0\\pi\_\{\\Phi\}:=\(\\operatorname\{Id\},F\_\{\\Phi\}\)\_\{\\\#\}\\mu\_\{0\}belongs toΠ\(μ0,μ1\)\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)and is concentrated onΓ\\Gamma\. Moreover,
∫ℝd×ℝdcp\(x,y\)πΦ\(dx,dy\)\\displaystyle\\int\_\{\\mathbb\{R\}^\{d\}\\times\\mathbb\{R\}^\{d\}\}c\_\{p\}\(x,y\)\\,\\pi\_\{\\Phi\}\(\\mathrm\{d\}x,\\mathrm\{d\}y\)=1p∫ℝd∥FΦ\(x\)−x∥pμ0\(dx\)\\displaystyle=\\frac\{1\}\{p\}\\int\_\{\\mathbb\{R\}^\{d\}\}\\lVert F\_\{\\Phi\}\(x\)\-x\\rVert^\{p\}\\,\\mu\_\{0\}\(\\mathrm\{d\}x\)≤2p−1p\(∫ℝd∥x∥pμ0\(dx\)\+∫ℝd∥y∥pμ1\(dy\)\)<∞\.\\displaystyle\\leq\\frac\{2^\{p\-1\}\}\{p\}\\left\(\\int\_\{\\mathbb\{R\}^\{d\}\}\\lVert x\\rVert^\{p\}\\,\\mu\_\{0\}\(\\mathrm\{d\}x\)\+\\int\_\{\\mathbb\{R\}^\{d\}\}\\lVert y\\rVert^\{p\}\\,\\mu\_\{1\}\(\\mathrm\{d\}y\)\\right\)<\\infty\.Sincecpc\_\{p\}is continuous and finite\-valued onℝd×ℝd\\mathbb\{R\}^\{d\}\\times\\mathbb\{R\}^\{d\}, andπΦ\\pi\_\{\\Phi\}has finitecpc\_\{p\}\-cost, the standard sufficiency theorem forcpc\_\{p\}\-cyclically monotone transport plans implies thatπΦ\\pi\_\{\\Phi\}is an optimal coupling\(Santambrogio,[2015](https://arxiv.org/html/2608.05666#bib.bib5)\)\. The factor1/p1/pdoes not affect the optimizer, soFΦF\_\{\\Phi\}is optimal for the cost∥x−y∥p\\lVert x\-y\\rVert^\{p\}\.
Sinceμ0\\mu\_\{0\}is absolutely continuous andp\>1p\>1, the optimal transport map for this cost is uniqueμ0\\mu\_\{0\}\-almost everywhere\. Consequently,
FΦ=Tpμ0\-almost everywhere\.F\_\{\\Phi\}=T\_\{p\}\\qquad\\mu\_\{0\}\\text\{\-almost everywhere\}\.We now verify the dynamic optimality explicitly\. SinceρtΦ=\(zΦ\(t,⋅\)\)\#μ0\\rho\_\{t\}^\{\\Phi\}=\(z\_\{\\Phi\}\(t,\\cdot\)\)\_\{\\\#\}\\mu\_\{0\}, the pushforward formula, the straight\-line identity, andFΦ=TpF\_\{\\Phi\}=T\_\{p\}μ0\\mu\_\{0\}\-almost everywhere give
1p∫01∫ℝd∥vΦ\(t,y\)∥pρtΦ\(dy\)dt\\displaystyle\\frac\{1\}\{p\}\\int\_\{0\}^\{1\}\\int\_\{\\mathbb\{R\}^\{d\}\}\\lVert v\_\{\\Phi\}\(t,y\)\\rVert^\{p\}\\,\\rho\_\{t\}^\{\\Phi\}\(\\mathrm\{d\}y\)\\,\\mathrm\{d\}t=1p∫01∫ℝd‖vΦ\(t,zΦ\(t,x\)\)‖pμ0\(dx\)dt\\displaystyle=\\frac\{1\}\{p\}\\int\_\{0\}^\{1\}\\int\_\{\\mathbb\{R\}^\{d\}\}\\left\\lVert v\_\{\\Phi\}\\bigl\(t,z\_\{\\Phi\}\(t,x\)\\bigr\)\\right\\rVert^\{p\}\\,\\mu\_\{0\}\(\\mathrm\{d\}x\)\\,\\mathrm\{d\}t=1p∫01∫ℝd∥Tp\(x\)−x∥pμ0\(dx\)dt\\displaystyle=\\frac\{1\}\{p\}\\int\_\{0\}^\{1\}\\int\_\{\\mathbb\{R\}^\{d\}\}\\lVert T\_\{p\}\(x\)\-x\\rVert^\{p\}\\,\\mu\_\{0\}\(\\mathrm\{d\}x\)\\,\\mathrm\{d\}t=1p∫ℝd∥Tp\(x\)−x∥pμ0\(dx\)\\displaystyle=\\frac\{1\}\{p\}\\int\_\{\\mathbb\{R\}^\{d\}\}\\lVert T\_\{p\}\(x\)\-x\\rVert^\{p\}\\,\\mu\_\{0\}\(\\mathrm\{d\}x\)=1pWpp\(μ0,μ1\)\.\\displaystyle=\\frac\{1\}\{p\}W\_\{p\}^\{p\}\(\\mu\_\{0\},\\mu\_\{1\}\)\.The generalized Benamou–Brenier formula states thatWpp\(μ0,μ1\)/pW\_\{p\}^\{p\}\(\\mu\_\{0\},\\mu\_\{1\}\)/pis the minimum of the dynamic action\. Since each absolutely continuous trajectoryzΦ\(⋅,x\)z\_\{\\Phi\}\(\\cdot,x\)satisfies the characteristic equation almost everywhere in time andρtΦ=\(zΦ\(t,⋅\)\)\#μ0\\rho\_\{t\}^\{\\Phi\}=\(z\_\{\\Phi\}\(t,\\cdot\)\)\_\{\\\#\}\\mu\_\{0\}, the standard characteristic\-flow argument implies that\(ρtΦ,vΦ\)\(\\rho\_\{t\}^\{\\Phi\},v\_\{\\Phi\}\)satisfies the continuity equation in the distributional sense\. Together with the endpoint conditionsρ0Φ=μ0\\rho\_\{0\}^\{\\Phi\}=\\mu\_\{0\}andρ1Φ=μ1\\rho\_\{1\}^\{\\Phi\}=\\mu\_\{1\}, the preceding action identity therefore shows that\(ρtΦ,vΦ\)\(\\rho\_\{t\}^\{\\Phi\},v\_\{\\Phi\}\)is a minimizer of thepp\-Benamou–Brenier problem\.
Finally, forμ0\\mu\_\{0\}\-almost everyxx,
zΦ\(t,x\)=\(1−t\)x\+tTp\(x\)for everyt∈\[0,1\],z\_\{\\Phi\}\(t,x\)=\(1\-t\)x\+tT\_\{p\}\(x\)\\qquad\\text\{for every \}t\\in\[0,1\],and
vΦ\(t,zΦ\(t,x\)\)=Tp\(x\)−xfor everyt∈\(0,1\)\.v\_\{\\Phi\}\\bigl\(t,z\_\{\\Phi\}\(t,x\)\\bigr\)=T\_\{p\}\(x\)\-x\\qquad\\text\{for every \}t\\in\(0,1\)\.
The Benamou–Brenier solution\(ρ⋆,v⋆\)\(\\rho^\{\\star\},v^\{\\star\}\)from part \(1\) is induced by the same optimal mapTpT\_\{p\}, and both velocity fields equalTp\(x\)−xT\_\{p\}\(x\)\-xalong the common trajectory\(1−t\)x\+tTp\(x\)\(1\-t\)x\+tT\_\{p\}\(x\),dt⊗μ0\\mathrm\{d\}t\\otimes\\mu\_\{0\}\-almost everywhere\. Since
ρtΦ=\(\(1−t\)Id\+tTp\)♯μ0=ρt⋆,\\rho\_\{t\}^\{\\Phi\}=\\bigl\(\(1\-t\)\\operatorname\{Id\}\+tT\_\{p\}\\bigr\)\_\{\\sharp\}\\mu\_\{0\}=\\rho\_\{t\}^\{\\star\},the pushforward formula gives
∫01∫ℝd‖vΦ\(t,y\)−v⋆\(t,y\)‖pρtΦ\(dy\)dt\\displaystyle\\int\_\{0\}^\{1\}\\int\_\{\\mathbb\{R\}^\{d\}\}\\left\\lVert v\_\{\\Phi\}\(t,y\)\-v^\{\\star\}\(t,y\)\\right\\rVert^\{p\}\\,\\rho\_\{t\}^\{\\Phi\}\(\\mathrm\{d\}y\)\\,\\mathrm\{d\}t=\\displaystyle=∫01∫ℝd‖vΦ\(t,\(1−t\)x\+tTp\(x\)\)−v⋆\(t,\(1−t\)x\+tTp\(x\)\)‖pμ0\(dx\)dt\\displaystyle\\int\_\{0\}^\{1\}\\int\_\{\\mathbb\{R\}^\{d\}\}\\left\\lVert v\_\{\\Phi\}\\bigl\(t,\(1\-t\)x\+tT\_\{p\}\(x\)\\bigr\)\-v^\{\\star\}\\bigl\(t,\(1\-t\)x\+tT\_\{p\}\(x\)\\bigr\)\\right\\rVert^\{p\}\\,\\mu\_\{0\}\(\\mathrm\{d\}x\)\\,\\mathrm\{d\}t=\\displaystyle=0\.\\displaystyle 0\.Since the integrand on the left\-hand side is nonnegative, it follows that
vΦ=v⋆dtρtΦ\(dy\)\-almost everywhere\.v\_\{\\Phi\}=v^\{\\star\}\\qquad\\mathrm\{d\}t\\,\\rho\_\{t\}^\{\\Phi\}\(\\mathrm\{d\}y\)\\text\{\-almost everywhere\}\.
This proves both exact recovery and velocity\-field uniqueness\. ∎
### A\.2The Role and Non\-Redundancy of Interior Full Support
Theorem[3\.2](https://arxiv.org/html/2608.05666#S3.Thmtheorem2)shows that a zero\-loss potentialΦ\\Phirecovers the optimal transport map provided thatΦ\\PhiisC2C^\{2\}on\(0,1\)×ℝd\(0,1\)\\times\\mathbb\{R\}^\{d\}and the induced distributionρtΦ\\rho\_\{t\}^\{\\Phi\}has full support for everyt∈\(0,1\)t\\in\(0,1\)\. The interiorC2C^\{2\}regularity ensures that the differential calculations and chain\-rule arguments in the proof are valid\. A natural question is whether the full\-support condition can be removed or replaced by a weaker assumption\. The following counterexamples address this question\. Counterexample 1 gives a zero\-loss potential for whichρtΦ\\rho\_\{t\}^\{\\Phi\}does not have full support and the induced terminal map is not optimal; in this example, the target measureμ1\\mu\_\{1\}also fails to have full support\. Counterexample 2 further shows that requiringμ1\\mu\_\{1\}itself to have full support is still insufficient to guarantee optimality\.
Both counterexamples usep=2p=2\. In this casevΦ=−∇xΦv\_\{\\Phi\}=\-\\nabla\_\{x\}\\Phi, and hence, withSt\(x\)=\(1−t\)x\+tFΦ\(x\)S\_\{t\}\(x\)=\(1\-t\)x\+tF\_\{\\Phi\}\(x\),
∇xΦ\(t,St\(x\)\)\+FΦ\(x\)−x=−\[vΦ\(t,St\(x\)\)−\(FΦ\(x\)−x\)\]\.\\nabla\_\{x\}\\Phi\\bigl\(t,S\_\{t\}\(x\)\\bigr\)\+F\_\{\\Phi\}\(x\)\-x=\-\\left\[v\_\{\\Phi\}\\bigl\(t,S\_\{t\}\(x\)\\bigr\)\-\\bigl\(F\_\{\\Phi\}\(x\)\-x\\bigr\)\\right\]\.Thus the potential\-matching residual is the negative of the velocity residual used in the former flow\-matching formulation\. Their squared norms are identical, so the two zero\-loss conditions coincide in these examples\.
#### Counterexample 1: failure after removing interior full support\.
Letd=2d=2,p=2p=2, andc2\(x,y\)=∥x−y∥2/2c\_\{2\}\(x,y\)=\\lVert x\-y\\rVert^\{2\}/2\. FixM\>8M\>\\sqrt\{8\}and define
a−\\displaystyle a\_\{\-\}:=\(−2,−M/2\),\\displaystyle=\(\-2,\-M/2\),u−\\displaystyle u\_\{\-\}:=\(1,M\),\\displaystyle=\(1,M\),b−\\displaystyle b\_\{\-\}:=a−\+u−=\(−1,M/2\),\\displaystyle=a\_\{\-\}\+u\_\{\-\}=\(\-1,M/2\),a\+\\displaystyle a\_\{\+\}:=\(2,M/2\),\\displaystyle=\(2,M/2\),u\+\\displaystyle u\_\{\+\}:=\(−1,−M\),\\displaystyle=\(\-1,\-M\),b\+\\displaystyle b\_\{\+\}:=a\+\+u\+=\(1,−M/2\)\.\\displaystyle=a\_\{\+\}\+u\_\{\+\}=\(1,\-M/2\)\.Choose0<r<1/40<r<1/4and set
A−:=B\(a−,r/2\),A\+:=B\(a\+,r/2\)\.A\_\{\-\}:=B\(a\_\{\-\},r/2\),\\qquad A\_\{\+\}:=B\(a\_\{\+\},r/2\)\.Define
μ0:=12Unif\(A−\)\+12Unif\(A\+\)\.\\mu\_\{0\}:=\\frac\{1\}\{2\}\\operatorname\{Unif\}\(A\_\{\-\}\)\+\\frac\{1\}\{2\}\\operatorname\{Unif\}\(A\_\{\+\}\)\.DefineFFμ0\\mu\_\{0\}\-almost everywhere by
F\(x\):=\{x\+u−,x∈A−,x\+u\+,x∈A\+\.F\(x\):=\\begin\{cases\}x\+u\_\{\-\},&x\\in A\_\{\-\},\\\\ x\+u\_\{\+\},&x\\in A\_\{\+\}\.\\end\{cases\}Define
μ1:=F♯μ0\.\\mu\_\{1\}:=F\_\{\\sharp\}\\mu\_\{0\}\.Then
supp\(μ1\)=B\(b−,r/2\)¯∪B\(b\+,r/2\)¯≠ℝ2\.\\operatorname\{supp\}\(\\mu\_\{1\}\)=\\overline\{B\(b\_\{\-\},r/2\)\}\\cup\\overline\{B\(b\_\{\+\},r/2\)\}\\neq\\mathbb\{R\}^\{2\}\.
We next construct a smooth potential whose flow agrees withFFon the support ofμ0\\mu\_\{0\}\. Define the moving centers
c−\(t\):=a−\+tu−,c\+\(t\):=a\+\+tu\+,t∈\[0,1\]\.c\_\{\-\}\(t\):=a\_\{\-\}\+tu\_\{\-\},\\qquad c\_\{\+\}\(t\):=a\_\{\+\}\+tu\_\{\+\},\\qquad t\\in\[0,1\]\.Their first coordinates satisfy
\(c\+\(t\)\)1−\(c−\(t\)\)1=4−2t≥2\.\\bigl\(c\_\{\+\}\(t\)\\bigr\)\_\{1\}\-\\bigl\(c\_\{\-\}\(t\)\\bigr\)\_\{1\}=4\-2t\\geq 2\.Letχ∈Cc∞\(ℝ2\)\\chi\\in C\_\{c\}^\{\\infty\}\(\\mathbb\{R\}^\{2\}\)satisfy
χ\(ξ\)=1for∥ξ∥≤r,supp\(χ\)⊂B\(0,2r\)\.\\chi\(\\xi\)=1\\quad\\text\{for \}\\lVert\\xi\\rVert\\leq r,\\qquad\\operatorname\{supp\}\(\\chi\)\\subset B\(0,2r\)\.Becauser<1/4r<1/4, the two moving cutoff supports remain disjoint for allt∈\[0,1\]t\\in\[0,1\]\. Define
Φ\(t,y\):=\\displaystyle\\Phi\(t,y\)=\{\}−χ\(y−c−\(t\)\)u−⋅\(y−c−\(t\)\)\\displaystyle\-\\chi\\bigl\(y\-c\_\{\-\}\(t\)\\bigr\)u\_\{\-\}\\cdot\\bigl\(y\-c\_\{\-\}\(t\)\\bigr\)−χ\(y−c\+\(t\)\)u\+⋅\(y−c\+\(t\)\),\\displaystyle\-\\chi\\bigl\(y\-c\_\{\+\}\(t\)\\bigr\)u\_\{\+\}\\cdot\\bigl\(y\-c\_\{\+\}\(t\)\\bigr\),and letvΦ=−∇yΦv\_\{\\Phi\}=\-\\nabla\_\{y\}\\Phi\. OnB\(c−\(t\),r\)B\(c\_\{\-\}\(t\),r\), the first cutoff is identically one and the second vanishes, so
vΦ\(t,y\)=u−\.v\_\{\\Phi\}\(t,y\)=u\_\{\-\}\.Similarly,vΦ\(t,y\)=u\+v\_\{\\Phi\}\(t,y\)=u\_\{\+\}onB\(c\+\(t\),r\)B\(c\_\{\+\}\(t\),r\)\. The potentialΦ\\Phiis smooth on\[0,1\]×ℝ2\[0,1\]\\times\\mathbb\{R\}^\{2\}, andvΦv\_\{\\Phi\}is smooth with uniformly compact spatial support\. Hence it generates a unique global smooth flow, and every time slicezΦ\(t,⋅\)z\_\{\\Phi\}\(t,\\cdot\)is a diffeomorphism ofℝ2\\mathbb\{R\}^\{2\}\.
Forx∈A−x\\in A\_\{\-\}, the curvez\(t\)=x\+tu−z\(t\)=x\+tu\_\{\-\}satisfies
∥z\(t\)−c−\(t\)∥=∥x−a−∥<r/2<r,\\lVert z\(t\)\-c\_\{\-\}\(t\)\\rVert=\\lVert x\-a\_\{\-\}\\rVert<r/2<r,and therefore
∂tz\(t\)=u−=vΦ\(t,z\(t\)\)\.\\partial\_\{t\}z\(t\)=u\_\{\-\}=v\_\{\\Phi\}\\bigl\(t,z\(t\)\\bigr\)\.Uniqueness of the flow giveszΦ\(t,x\)=x\+tu−z\_\{\\Phi\}\(t,x\)=x\+tu\_\{\-\}for everyt∈\[0,1\]t\\in\[0,1\]\. The same argument giveszΦ\(t,x\)=x\+tu\+z\_\{\\Phi\}\(t,x\)=x\+tu\_\{\+\}forx∈A\+x\\in A\_\{\+\}\. Consequently, forμ0\\mu\_\{0\}\-almost everyxx,
zΦ\(t,x\)=\(1−t\)x\+tF\(x\),FΦ\(x\)=zΦ\(1,x\)=F\(x\)\.z\_\{\\Phi\}\(t,x\)=\(1\-t\)x\+tF\(x\),\\qquad F\_\{\\Phi\}\(x\)=z\_\{\\Phi\}\(1,x\)=F\(x\)\.It follows that
vΦ\(t,\(1−t\)x\+tFΦ\(x\)\)=FΦ\(x\)−xv\_\{\\Phi\}\\bigl\(t,\(1\-t\)x\+tF\_\{\\Phi\}\(x\)\\bigr\)=F\_\{\\Phi\}\(x\)\-xfor everyt∈\[0,1\]t\\in\[0,1\]and forμ0\\mu\_\{0\}\-almost everyxx\. ThusℒPM\(Φ\)=0\\mathcal\{L\}\_\{\\mathrm\{PM\}\}\(\\Phi\)=0\. Since\(FΦ\)♯μ0=μ1\(F\_\{\\Phi\}\)\_\{\\sharp\}\\mu\_\{0\}=\\mu\_\{1\}, the terminal loss also vanishes, and hence
ℒPMOT\(Φ\)=0\.\\mathcal\{L\}\_\{\\mathrm\{PMOT\}\}\(\\Phi\)=0\.Because the PMOT objective is nonnegative,Φ\\Phiis a global minimizer\. Nevertheless, its intermediate distributions occupy only two moving tubes:
supp\(ρtΦ\)=B\(c−\(t\),r/2\)¯∪B\(c\+\(t\),r/2\)¯≠ℝ2,t∈\[0,1\]\.\\operatorname\{supp\}\(\\rho\_\{t\}^\{\\Phi\}\)=\\overline\{B\(c\_\{\-\}\(t\),r/2\)\}\\cup\\overline\{B\(c\_\{\+\}\(t\),r/2\)\}\\neq\\mathbb\{R\}^\{2\},\\qquad t\\in\[0,1\]\.
The induced coupling is not optimal\. Since the density ofμ0\\mu\_\{0\}is strictly positive neara−a\_\{\-\}anda\+a\_\{\+\}andFFis a translation on each ball, both\(a−,b−\)\(a\_\{\-\},b\_\{\-\}\)and\(a\+,b\+\)\(a\_\{\+\},b\_\{\+\}\)belong to the support of
π:=\(Id,F\)♯μ0\.\\pi:=\(\\operatorname\{Id\},F\)\_\{\\sharp\}\\mu\_\{0\}\.The cost of the assigned pairing is
c2\(a−,b−\)\+c2\(a\+,b\+\)=12∥u−∥2\+12∥u\+∥2=1\+M2\.c\_\{2\}\(a\_\{\-\},b\_\{\-\}\)\+c\_\{2\}\(a\_\{\+\},b\_\{\+\}\)=\\frac\{1\}\{2\}\\lVert u\_\{\-\}\\rVert^\{2\}\+\\frac\{1\}\{2\}\\lVert u\_\{\+\}\\rVert^\{2\}=1\+M^\{2\}\.After exchanging the endpoints,a−−b\+=\(−3,0\)a\_\{\-\}\-b\_\{\+\}=\(\-3,0\)anda\+−b−=\(3,0\)a\_\{\+\}\-b\_\{\-\}=\(3,0\), so
c2\(a−,b\+\)\+c2\(a\+,b−\)=9\.c\_\{2\}\(a\_\{\-\},b\_\{\+\}\)\+c\_\{2\}\(a\_\{\+\},b\_\{\-\}\)=9\.SinceM\>8M\>\\sqrt\{8\}, we have1\+M2\>91\+M^\{2\}\>9\. Thus the support ofπ\\piviolates the two\-pointc2c\_\{2\}\-cyclical\-monotonicity inequality, andπ\\piis not an optimal coupling betweenμ0\\mu\_\{0\}andμ1\\mu\_\{1\}\.
Thus, without interior full support, a smooth potential may attain zero PMOT loss while inducing a nonoptimal terminal map\. The obstruction is the persistent vacuum outside the two moving tubes: the potential\-matching loss constrains the potential gradient, and equivalently the velocity in the quadratic case, only along mass\-carrying trajectories, so the material\-derivative identity cannot be extended to the entire interior space–time domain\.
#### Counterexample 2: terminal full support does not replace interior full support\.
In Counterexample 1, the target measureμ1\\mu\_\{1\}does not have full support\. The following example further shows that even the conditionsupp\(μ1\)=ℝd\\operatorname\{supp\}\(\\mu\_\{1\}\)=\\mathbb\{R\}^\{d\}is insufficient to guarantee that a zero\-loss terminal map is optimal when the intermediate distributions fail to have full support\.
Letd=2d=2,p=2p=2, and fixM\>8M\>\\sqrt\{8\}\. Define
K−:=\{x∈ℝ2:x1<−1\},K\+:=\{x∈ℝ2:x1\>1\},K\_\{\-\}:=\\\{x\\in\\mathbb\{R\}^\{2\}:x\_\{1\}<\-1\\\},\\qquad K\_\{\+\}:=\\\{x\\in\\mathbb\{R\}^\{2\}:x\_\{1\}\>1\\\},and set
a−:=\(−2,−M/2\),a\+:=\(2,M/2\)\.a\_\{\-\}:=\(\-2,\-M/2\),\\qquad a\_\{\+\}:=\(2,M/2\)\.For suitable normalizing constantsZ−,Z\+\>0Z\_\{\-\},Z\_\{\+\}\>0, let
f−\(x\):=1Z−e−∥x−a−∥2𝟏K−\(x\),f\+\(x\):=1Z\+e−∥x−a\+∥2𝟏K\+\(x\),f\_\{\-\}\(x\):=\\frac\{1\}\{Z\_\{\-\}\}e^\{\-\\lVert x\-a\_\{\-\}\\rVert^\{2\}\}\\mathbf\{1\}\_\{K\_\{\-\}\}\(x\),\\qquad f\_\{\+\}\(x\):=\\frac\{1\}\{Z\_\{\+\}\}e^\{\-\\lVert x\-a\_\{\+\}\\rVert^\{2\}\}\\mathbf\{1\}\_\{K\_\{\+\}\}\(x\),and define
μ0\(dx\):=12f−\(x\)dx\+12f\+\(x\)dx\.\\mu\_\{0\}\(\\mathrm\{d\}x\):=\\frac\{1\}\{2\}f\_\{\-\}\(x\)\\,\\mathrm\{d\}x\+\\frac\{1\}\{2\}f\_\{\+\}\(x\)\\,\\mathrm\{d\}x\.Thenμ0≪ℒ2\\mu\_\{0\}\\ll\\mathcal\{L\}^\{2\}andμ0∈𝒫2\(ℝ2\)\\mu\_\{0\}\\in\\mathcal\{P\}\_\{2\}\(\\mathbb\{R\}^\{2\}\)\.
Let
u−:=\(1,M\),u\+:=\(−1,−M\),u\_\{\-\}:=\(1,M\),\\qquad u\_\{\+\}:=\(\-1,\-M\),and define
F\(x\):=\{x\+u−,x∈K−,x\+u\+,x∈K\+,μ1:=F♯μ0\.F\(x\):=\\begin\{cases\}x\+u\_\{\-\},&x\\in K\_\{\-\},\\\\ x\+u\_\{\+\},&x\\in K\_\{\+\},\\end\{cases\}\\qquad\\mu\_\{1\}:=F\_\{\\sharp\}\\mu\_\{0\}\.SinceF\(K−\)=\{y1<0\}F\(K\_\{\-\}\)=\\\{y\_\{1\}<0\\\}andF\(K\+\)=\{y1\>0\}F\(K\_\{\+\}\)=\\\{y\_\{1\}\>0\\\}, we have
supp\(μ1\)=ℝ2\.\\operatorname\{supp\}\(\\mu\_\{1\}\)=\\mathbb\{R\}^\{2\}\.ForSt\(x\):=\(1−t\)x\+tF\(x\)S\_\{t\}\(x\):=\(1\-t\)x\+tF\(x\)andρt:=\(St\)♯μ0\\rho\_\{t\}:=\(S\_\{t\}\)\_\{\\sharp\}\\mu\_\{0\},
supp\(ρt\)=\{y1≤−\(1−t\)\}∪\{y1≥1−t\}≠ℝ2,t∈\(0,1\)\.\\operatorname\{supp\}\(\\rho\_\{t\}\)=\\\{y\_\{1\}\\leq\-\(1\-t\)\\\}\\cup\\\{y\_\{1\}\\geq 1\-t\\\}\\neq\\mathbb\{R\}^\{2\},\\qquad t\\in\(0,1\)\.
Choose a smooth nondecreasing functionϑ∈C∞\(ℝ;\[0,1\]\)\\vartheta\\in C^\{\\infty\}\(\\mathbb\{R\};\[0,1\]\)such that
ϑ\(r\)=0forr≤−1,ϑ\(r\)=1forr≥1,\\vartheta\(r\)=0\\quad\\text\{for \}r\\leq\-1,\\qquad\\vartheta\(r\)=1\\quad\\text\{for \}r\\geq 1,and define, for\(t,y\)∈\(0,1\)×ℝ2\(t,y\)\\in\(0,1\)\\times\\mathbb\{R\}^\{2\},
Φ\(t,y\):=\(2ϑ\(y11−t\)−1\)\(y1\+My2\)\.\\Phi\(t,y\):=\\left\(2\\vartheta\\left\(\\frac\{y\_\{1\}\}\{1\-t\}\\right\)\-1\\right\)\(y\_\{1\}\+My\_\{2\}\)\.The values ofΦ\\Phiatt=0t=0andt=1t=1may be chosen arbitrarily\. ThenΦ∈C∞\(\(0,1\)×ℝ2\)\\Phi\\in C^\{\\infty\}\(\(0,1\)\\times\\mathbb\{R\}^\{2\}\)\. Ify1<−\(1−t\)y\_\{1\}<\-\(1\-t\), then
Φ\(t,y\)=−\(y1\+My2\),vΦ\(t,y\)=−∇yΦ\(t,y\)=\(1,M\)=u−\.\\Phi\(t,y\)=\-\(y\_\{1\}\+My\_\{2\}\),\\qquad v\_\{\\Phi\}\(t,y\)=\-\\nabla\_\{y\}\\Phi\(t,y\)=\(1,M\)=u\_\{\-\}\.Similarly, ify1\>1−ty\_\{1\}\>1\-t, then
Φ\(t,y\)=y1\+My2,vΦ\(t,y\)=\(−1,−M\)=u\+\.\\Phi\(t,y\)=y\_\{1\}\+My\_\{2\},\\qquad v\_\{\\Phi\}\(t,y\)=\(\-1,\-M\)=u\_\{\+\}\.For everyx∈K−∪K\+x\\in K\_\{\-\}\\cup K\_\{\+\}andt∈\(0,1\)t\\in\(0,1\), the straight\-line trajectorySt\(x\)S\_\{t\}\(x\)remains in the corresponding occupied component\. Hence
vΦ\(t,St\(x\)\)=F\(x\)−x=∂tSt\(x\)\.v\_\{\\Phi\}\\bigl\(t,S\_\{t\}\(x\)\\bigr\)=F\(x\)\-x=\\partial\_\{t\}S\_\{t\}\(x\)\.ThusFΦ=FF\_\{\\Phi\}=Fμ0\\mu\_\{0\}\-almost everywhere, and
ℒPM\(Φ\)=0,ℒterm\(Φ\)=0,ℒPMOT\(Φ\)=0\.\\mathcal\{L\}\_\{\\mathrm\{PM\}\}\(\\Phi\)=0,\\qquad\\mathcal\{L\}\_\{\\mathrm\{term\}\}\(\\Phi\)=0,\\qquad\\mathcal\{L\}\_\{\\mathrm\{PMOT\}\}\(\\Phi\)=0\.In particular,Φ∈dom\(ℒPMOT\)\\Phi\\in\\operatorname\{dom\}\(\\mathcal\{L\}\_\{\\mathrm\{PMOT\}\}\)and is a global minimizer\.
Nevertheless,FFis not optimal\. Let
x−:=a−,x\+:=a\+,y−:=F\(x−\)=\(−1,M/2\),y\+:=F\(x\+\)=\(1,−M/2\)\.x\_\{\-\}:=a\_\{\-\},\\qquad x\_\{\+\}:=a\_\{\+\},\\qquad y\_\{\-\}:=F\(x\_\{\-\}\)=\(\-1,M/2\),\\qquad y\_\{\+\}:=F\(x\_\{\+\}\)=\(1,\-M/2\)\.Forc2\(x,y\)=∥x−y∥2/2c\_\{2\}\(x,y\)=\\lVert x\-y\\rVert^\{2\}/2,
c2\(x−,y−\)\+c2\(x\+,y\+\)\\displaystyle c\_\{2\}\(x\_\{\-\},y\_\{\-\}\)\+c\_\{2\}\(x\_\{\+\},y\_\{\+\}\)=1\+M2,\\displaystyle=1\+M^\{2\},c2\(x−,y\+\)\+c2\(x\+,y−\)\\displaystyle c\_\{2\}\(x\_\{\-\},y\_\{\+\}\)\+c\_\{2\}\(x\_\{\+\},y\_\{\-\}\)=9\.\\displaystyle=9\.SinceM\>8M\>\\sqrt\{8\}, the first quantity is larger than the second\. Moreover, the density ofμ0\\mu\_\{0\}is positive nearx−x\_\{\-\}andx\+x\_\{\+\}, so\(x−,y−\)\(x\_\{\-\},y\_\{\-\}\)and\(x\+,y\+\)\(x\_\{\+\},y\_\{\+\}\)belong to the support of\(Id,F\)♯μ0\(\\operatorname\{Id\},F\)\_\{\\sharp\}\\mu\_\{0\}\. Hence this support is notc2c\_\{2\}\-cyclically monotone, and\(Id,F\)♯μ0\(\\operatorname\{Id\},F\)\_\{\\sharp\}\\mu\_\{0\}is not an optimal coupling\.
ThusℒPMOT\(Φ\)=0\\mathcal\{L\}\_\{\\mathrm\{PMOT\}\}\(\\Phi\)=0andsupp\(μ1\)=ℝ2\\operatorname\{supp\}\(\\mu\_\{1\}\)=\\mathbb\{R\}^\{2\}, butFΦF\_\{\\Phi\}is not an optimal Monge map\. Terminal full support therefore does not replace the interior full\-support condition in Theorem[3\.2](https://arxiv.org/html/2608.05666#S3.Thmtheorem2)\.
## Appendix BAdditional Experimental Results
### B\.1Additional Toy Generation Results
Figure[5](https://arxiv.org/html/2608.05666#A2.F5)shows unconditional generation results on the two\-dimensional toy benchmarks\. Samples are generated by transporting standard Gaussian samples through the learned inverse flow\. These results verify endpoint generation quality, while the main text focuses on transport geometry andpp\-specific alignment\.
Figure 5:Unconditional generation on two\-dimensional toy benchmarks\. The top row shows samples from the target distributions, and the bottom row shows samples generated by PMOT from a standard Gaussian prior\. All panels use the same two\-dimensional histogram visualization protocol\.
### B\.2Repeatedpp\-Swap Evaluation on Pinwheel
In the main text, Figure[1](https://arxiv.org/html/2608.05666#S4.F1)shows a representativepp\-swap evaluation on Pinwheel\. To verify that the diagonal preference is not due to a particular evaluation sample, we repeat the evaluation over five independently sampled source and reference batches\. For each repetition, we compute the mean endpoint MSE over the diagonal entries, corresponding to matched training and reference exponents, and over the off\-diagonal entries, corresponding to mismatched exponents\.
Table 4:Repeatedpp\-swap evaluation on Pinwheel\. Diagonal entries compare each PMOT checkpoint with the Sinkhorn barycentric reference computed using the matching exponent\. Off\-diagonal entries use mismatched reference exponents\. Values are averaged over the corresponding heatmap entries in each evaluation repetition and reported as mean±\\pmstandard deviation\.Across the five evaluation repetitions, the mean off\-diagonal endpoint MSE is1\.410±0\.0481\.410\\pm 0\.048times the mean diagonal MSE, corresponding to an average increase of41\.0%41\.0\\%\. This indicates that the fixed PMOT checkpoints are consistently closer to post\-hoc OT references computed with their matching cost exponents\. These repetitions measure evaluation\-sample variability rather than variability across independent training runs\.
### B\.3Likelihood\-Weight Sensitivity on POWER
Table[5](https://arxiv.org/html/2608.05666#A2.T5)reports an additional likelihood\-weight sensitivity study on POWER\. PMOT uses the same fixed weights\(λPM,λNLL\)=\(1,5\)\(\\lambda\_\{\\mathrm\{PM\}\},\\lambda\_\{\\mathrm\{NLL\}\}\)=\(1,5\)as in the main high\-dimensional experiments and does not include an HJB/R regularizer\. For OT\-Flow, we keep the kinetic and HJB/R coefficients fixed and vary only the likelihood coefficientλC\\lambda\_\{C\}\.
Table 5:Likelihood\-weight sensitivity on POWER\. PMOT weights are reported as\(λPM,λNLL\)\(\\lambda\_\{\\mathrm\{PM\}\},\\lambda\_\{\\mathrm\{NLL\}\}\), while OT\-Flow weights are reported as\(λL,λC,λR\)\(\\lambda\_\{L\},\\lambda\_\{C\},\\lambda\_\{R\}\)\. OT\-Flow keeps its kinetic and HJB/R coefficients fixed while varyingλC\\lambda\_\{C\}\. Evaluation uses biased MMD with 4000 generated and 4000 held\-out samples\. Lower is better\.Increasing the OT\-Flow likelihood coefficient improves NLL on POWER, although generation MMD varies non\-monotonically\. PMOT achieves lower NLL and MMD with the fixed\(1,5\)\(1,5\)weighting and no HJB/R regularizer\. Together with the 8\-Gaussians sensitivity experiment, these results suggest that PMOT remains effective without a heavily likelihood\-dominated objective\.
### B\.4Image Color Transformation
We use image color transformation to illustrate PMOT with sample\-based terminal matching in RGB space\. In this experiment, the terminal discrepancy in equation[5](https://arxiv.org/html/2608.05666#S3.E5)is instantiated as MMD rather than the NLL objective used for toy and tabular density estimation\. The learned forward flow transports the empirical color distribution of a content image toward that of a style image\. The resulting map is applied independently to each pixel, so the spatial layout of the content image is retained\.
We select two256×256256\\times 256images from ImageNet\(Denget al\.,[2009](https://arxiv.org/html/2608.05666#bib.bib17); Russakovskyet al\.,[2015](https://arxiv.org/html/2608.05666#bib.bib18)\): an American chameleon \(synset n01682714\) as the content image and a sulphur\-crested cockatoo \(synset n01819313\) as the style image\. Their empirical RGB distributions are treated as the source and target distributions, respectively\. RGB values are rescaled to\(0,1\)3\(0,1\)^\{3\}and mapped toℝ3\\mathbb\{R\}^\{3\}using a componentwise logit transform\. We train PMOT withp=2p=2,\(λPM,λMMD\)=\(1,10\)\(\\lambda\_\{\\mathrm\{PM\}\},\\lambda\_\{\\mathrm\{MMD\}\}\)=\(1,10\), and a mixture of five Gaussian kernels with bandwidthsσ∈\{0\.25,0\.5,1\.0,2\.0,4\.0\}\\sigma\\in\\\{0\.25,0\.5,1\.0,2\.0,4\.0\\\}for 25,000 iterations\.
Table[6](https://arxiv.org/html/2608.05666#A2.T6)reports the final training losses\. Figure[4](https://arxiv.org/html/2608.05666#S4.F4)in the main text visualizes the transformed image along the learned flow\. The marginal RGB distributions in Figure[6](https://arxiv.org/html/2608.05666#A2.F6)provide an additional qualitative comparison between the source, transformed, and target color distributions\.
Table 6:Final PMOT training losses for image color transformation\.Figure 6:Marginal red, green, and blue distributions of the content image, the PMOT\-transformed image, and the target style image\. The transformed marginals move toward those of the target, providing qualitative evidence of color\-distribution matching\.Similar Articles
Multimarginal flow matching with optimal transport potentials
Proposes OTP-FM, a novel method for multimarginal flow matching that uses optimal transport potentials to softly steer flows through intermediate marginals, achieving state-of-the-art performance on single-cell RNA sequencing, oceanographic, and meteorological datasets.
Perron--Frobenius Operator Matching for Generative Modeling
Introduces Perron–Frobenius Operator Matching (PFOM), a generative framework that unifies flow, diffusion, and jump models via integral PF operator matching, proving KL divergence yields a practical loss equivalent to Koopman path matching, and develops Nesterov-accelerated training and sampling for improved efficiency.
Reward Transport: Property Control in Flow Matching via Noise-Space Alignment
This paper introduces Reward Transport, a method that uses optimal transport coupling during flow matching training to align a scalar noise coordinate with molecular rewards, enabling monotone control over molecular properties like logP and QED at inference without additional computation.
Capturing non-Markovian dynamics in non-equilibrium stochastic systems using flow matching
This paper develops a generative flow matching method to capture non-Markovian dynamics in non-equilibrium stochastic systems, demonstrating improved predictions for the Kramers first passage time problem compared to Markovian baselines.
FMOPF: Latent Flow Matching with Constraint-Aware Interaction Priors for AC Optimal Power Flow
FMOPF uses latent flow matching with constraint-aware interaction priors to generate diverse, feasible near-optimal solutions for AC optimal power flow, scaling to hundreds of buses while preserving feasibility.