Dynamic Generalized Gromov-Wasserstein Optimal Transport

arXiv cs.LG Papers

Summary

This paper introduces TP-DATE, a simulation-free framework for dynamic Gromov-Wasserstein optimal transport, which enhances spatial structure preservation and continuous dynamics reconstruction in spatial transcriptomics data.

arXiv:2609.20008v1 Announce Type: new Abstract: Gromov--Wasserstein optimal transport (GW-OT) extends classical optimal transport by introducing structure-aware transport cost. This is particularly relevant for spatial transcriptomics, where dynamical reconstruction should preserve tissue structure in addition to matching expression patterns. While static formulations have been widely used for such structure-aware alignment, a general dynamic formulation for reconstructing continuous trajectories is still missing. We introduce Travelling Pair Dynamical Alignment and Trajectory Estimation (TP-DATE), a theoretical and computational framework to generalize GW-OT dynamically in a simulation-free manner. We formulate a broad class of static and dynamic Quadratic-form OT (QOT) through path actions and prove the static dynamic equivalence. We further develop travelling-pair flow matching, which allows interacting conditional paths and marginalizes their interactions into a single vector field. On synthetic and real spatial transcriptomics data, TP-DATE better preserves spatial structure and improves continuous 3D dynamics reconstruction.
Original Article
View Cached Full Text

Cached at: 09/18/26, 09:15 AM

# Dynamic Generalized Gromov-Wasserstein Optimal Transport
Source: [https://arxiv.org/html/2609.20008](https://arxiv.org/html/2609.20008)
Junda Ying††thanks:Equal contributionZhiwei Zeng11footnotemark:1Affiliation:Center for Quantitative Biology, Peking UniversityPeijie Zhou††thanks:Corresponding authors: pjzhou@pku\.edu\.cn, zhangl@math\.pku\.edu\.cnAffiliation:Center for Quantitative Biology, Peking UniversityAffiliation:Center for Machine Learning Research, Peking UniversityAffiliation:National Engineering Laboratory for Big Data Analysis and Applications, BeijingAffiliation:AI for Science Institute, BeijingLei Zhang22footnotemark:2Affiliation:Beijing International Center for Mathematical Research, Peking UniversityAffiliation:Center for Quantitative Biology, Peking UniversityAffiliation:Center for Machine Learning Research, Peking UniversityAffiliation:School of Mathematical Sciences, Peking UniversityAffiliation:Institute for Artificial Intelligence, Peking University

###### Abstract

Gromov–Wasserstein optimal transport \(GW\-OT\) extends classical optimal transport by introducing structure\-aware transport cost\. This is particularly relevant for spatial transcriptomics, where dynamical reconstruction should preserve tissue structure in addition to matching expression patterns\. While static formulations have been widely used for such structure\-aware alignment, a general dynamic formulation for reconstructing continuous trajectories is still missing\. We introduceTravellingPairDynamicalAlignment andTrajectoryEstimation \(TP\-DATE\), a theoretical and computational framework to generalize GW\-OT dynamically in a simulation\-free manner\. We formulate a broad class of static and dynamic Quadratic\-form OT \(QOT\) through path actions and prove the static dynamic equivalence\. We further develop travelling\-pair flow matching, which allows interacting conditional paths and marginalizes their interactions into a single vector field\. On synthetic and real spatial transcriptomics data, TP\-DATE better preserves spatial structure and improves continuous 3D dynamics reconstruction\.

## 1Introduction

Optimal transport \(OT\)\([Kantorovich, 1958](https://arxiv.org/html/2609.20008#bib.bib20)\)provides a principled way to match probability distributions by minimizing transport cost, while its dynamic formulation\([Benamou and Brenier, 2000](https://arxiv.org/html/2609.20008#bib.bib21)\)lifts this endpoint matching into a continuous time evolution of measures\. This perspective has made OT a natural tool for reconstructing dynamics from population snapshots, including single\-cell systems where individual cells cannot be tracked longitudinally\. Static OT aligns snapshots by infering couplings across time points, whereas dynamic OT enables continuous interpolation and trajectory inference between observed snapshots\([Schiebinger et al\., 2019](https://arxiv.org/html/2609.20008#bib.bib31);[Tong et al\., 2020](https://arxiv.org/html/2609.20008#bib.bib34);[Tong et al\., 2024a](https://arxiv.org/html/2609.20008#bib.bib19);[Sha et al\., 2024](https://arxiv.org/html/2609.20008#bib.bib16);[Zhang et al\., 2025a](https://arxiv.org/html/2609.20008#bib.bib15);[Peng et al\., 2026b](https://arxiv.org/html/2609.20008#bib.bib14)\)\. Extensions such as Schrödinger bridges\([Schrödinger, 1932](https://arxiv.org/html/2609.20008#bib.bib25);[Léonard, 2014](https://arxiv.org/html/2609.20008#bib.bib26)\)and Wasserstein–Fisher–Rao\([Chizat et al\., 2018b](https://arxiv.org/html/2609.20008#bib.bib23);[Chizat et al\., 2018a](https://arxiv.org/html/2609.20008#bib.bib22);[Liero et al\., 2018](https://arxiv.org/html/2609.20008#bib.bib24)\)further broaden this framework to stochastic and unbalanced dynamics\.

However, pointwise transport cost alone may be insufficient when the data carry meaningful internal structure such as spatial structure\. Gromov–Wasserstein OT \(GW\-OT\) compares pairwise relations within distributions and are therefore well suited to structure\-aware matching tasks[Mémoli \(2011\)](https://arxiv.org/html/2609.20008#bib.bib35)\. Fused GW\-OT \(FGW\-OT\) further combines feature and structural information\([Vayer et al\., 2020](https://arxiv.org/html/2609.20008#bib.bib9);[Klein et al\., 2025](https://arxiv.org/html/2609.20008#bib.bib32)\)\. Such methods have become useful for structured data including graphs, heterogeneous domains, and spatial omics\. However, most GW\-like OTs still lack of dynamic formulation, therefore can only align snapshots instead of reconstructing a continuous time dynamics\.

Though a dynamic GW\-like formulation may not be meaningful when aligning snapshots from different modalities, it becomes important when the snapshots represent different states of the same structured system\. Spatial transcriptomics is a canonical example\. Snapshots collected at different times describe the same tissue evolving over time, and a meaningful interpolation should recover not only gene expression changes but also the continuous evolution of spatial organization\([Chen et al\., 2022](https://arxiv.org/html/2609.20008#bib.bib4);[Wei et al\., 2022](https://arxiv.org/html/2609.20008#bib.bib3)\)\. The same principle applies to serial tissue sections, where the interpolation axis is tissue depth rather than time and unseen intermediate sections correspond to physically meaningful states\([Zhang et al\., 2026a](https://arxiv.org/html/2609.20008#bib.bib5)\)\.

Recent work has begun to generalize GW\-OT from complementary directions\. On the static side,\([Wang and Zhang, 2025](https://arxiv.org/html/2609.20008#bib.bib6)\)defines a family of static Quadratic\-form optimal transport \(QOT\) including GW\-OT as a special case\. On the dynamic side, inner\-product GW\-OT \(IGW\-OT\)\([Zhang et al\., 2026b](https://arxiv.org/html/2609.20008#bib.bib7)\)develops and studies a dynamic formulation for a particular GW\-like OT based on gradient flow and Riemannian geometry\. However, a general dynamic QOT formulation and its theory is still missing, as is an efficient simulation\-free algorithm for solving them\.

To address these limitations, We introduceTravellingPairDynamicalAlignment andTrajectoryEstimation \(TP\-DATE\), a general framework for dynamic QOT together with a simulation\-free travelling\-pair flow matching solver\. Our contributions are summarized as follows\.

- •We developed a mathematical theory of dynamic QOT and included standard OT, GW\-OT, and IGW\-OT as special cases of our QOT framework\.
- •We developed a simulation\-free travelling pair flow matching framework for learning dynamics with interacting conditional paths and solving the dynamic QOT problems\.
- •We proposed a dynamic fusion of OT and GW\-OT, and demonstrate the ability of it in several biological tasks including spatiotemporal dynamics and 3D structure reconstruction\.

## 2Related works

Optimal transport and extensions\.Optimal transport\([Kantorovich, 1958](https://arxiv.org/html/2609.20008#bib.bib20)\)and its dynamic formulations\([Benamou and Brenier, 2000](https://arxiv.org/html/2609.20008#bib.bib21)\)have been widely used to reconstruct dynamics from snapshot data\. Stochastic counterparts\([Schrödinger, 1932](https://arxiv.org/html/2609.20008#bib.bib25);[Léonard, 2014](https://arxiv.org/html/2609.20008#bib.bib26)\)and unbalanced extensions\([Chizat et al\., 2018a](https://arxiv.org/html/2609.20008#bib.bib22);[Liero et al\., 2018](https://arxiv.org/html/2609.20008#bib.bib24)\)have further been applied to model more complex dynamics\. Another important line of extension is QOT, represented by formulations such as GW\-OT and FGW\-OT\([Mémoli, 2011](https://arxiv.org/html/2609.20008#bib.bib35);[Vayer et al\., 2020](https://arxiv.org/html/2609.20008#bib.bib9);[Klein et al\., 2025](https://arxiv.org/html/2609.20008#bib.bib32)\), which can model global structure preservation during transport\. Recently,\([Wang and Zhang, 2025](https://arxiv.org/html/2609.20008#bib.bib6)\)introduced a static QOT framework that unifies a broad class of such extensions, while\([Zhang et al\., 2026b](https://arxiv.org/html/2609.20008#bib.bib7)\)developed IGW\-OT as a dynamic formulation for a particular QOT problem\. We propose TP\-DATE, a unified dynamic formulation for a broad class of QOT problems which is compatible with these existing formulations\.

Flow matching based optimal transport solvers\.Flow matching\([Lipman et al\., 2023](https://arxiv.org/html/2609.20008#bib.bib18)\)is an efficient simulation\-free generative modeling framework that has been used to solve OT\([Tong et al\., 2024a](https://arxiv.org/html/2609.20008#bib.bib19)\), Schrödinger bridge\([Tong et al\., 2024b](https://arxiv.org/html/2609.20008#bib.bib28)\), WFR\([Peng et al\., 2026b](https://arxiv.org/html/2609.20008#bib.bib14)\), and a variety of OT\-based dynamics reconstruction problems\([Eyring et al\., 2024](https://arxiv.org/html/2609.20008#bib.bib11);[Rathod et al\., 2026](https://arxiv.org/html/2609.20008#bib.bib2);[Klein et al\., 2024](https://arxiv.org/html/2609.20008#bib.bib33);[Ying et al\., 2026](https://arxiv.org/html/2609.20008#bib.bib38)\)\. Inspired by the conditional path technique and the travelling Dirac in optimal transport\([Chizat et al\., 2018a](https://arxiv.org/html/2609.20008#bib.bib22)\), we formulate dynamic QOT from a path action viewpoint and develop a flow matching solver for this class of problems\.

OT\-based spatial transcriptomics dynamics reconstruction\.Several recent methods have been developed for spatiotemporal dynamics inference from spatial transcriptomics\. On the static alignment side,\([Klein et al\., 2025](https://arxiv.org/html/2609.20008#bib.bib32)\)formulates cross\-time alignment via GW\-OT and FGW\-OT, while\([Halmos et al\., 2025](https://arxiv.org/html/2609.20008#bib.bib1)\)encodes temporal and spatial structural information into static transport objectives\. On the dynamic side,\([Rathod et al\., 2026](https://arxiv.org/html/2609.20008#bib.bib2)\)incorporates structure\-aware static couplings into flow matching\.\([Peng et al\., 2026a](https://arxiv.org/html/2609.20008#bib.bib12)\)introduces structure\-preserving terms directly into the dynamic transport objective, and\([Zhang et al\., 2025b](https://arxiv.org/html/2609.20008#bib.bib13)\)explicitly models interactions within the learned dynamics, both by simulation\-based NeuralODE\([Chen et al\., 2018](https://arxiv.org/html/2609.20008#bib.bib17)\)\. This leaves a clear gap for TP\-DATE, a simulation\-free flow\-matching framework that directly models structure\-aware transport dynamics\.

## 3Preliminaries

Static optimal transport\.Let𝒫⁡\(𝒳\)\\mathcal\{P\}\(\\mathcal\{X\}\)be the set of all probability densities supported on𝒳\\mathcal\{X\}for some𝒳⊂ℝd\\mathcal\{X\}\\subset\\mathbb\{R\}^\{d\}\.μ0,μ1\\mu\_\{0\},\\mu\_\{1\}are probability densities in𝒫⁡\(𝒳\)\\mathcal\{P\}\(\\mathcal\{X\}\)\.Π⁡\(μ0,μ1\)\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)denotes the set of all couplings\.

Π\(μ0,μ1\)=\{γ∈𝒫\(𝒳2\)\|∫𝒳γ\(𝒙,𝒚\)d𝒚=μ0\(𝒙\),∫𝒳γ\(𝒙,𝒚\)d𝒙=μ1\(𝒚\)\}\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)=\\left\\\{\\gamma\\in\\mathcal\{P\}\(\\mathcal\{X\}^\{2\}\)\|\\int\_\{\\mathcal\{X\}\}\\gamma\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{y\}=\\mu\_\{0\}\(\\bm\{x\}\),\\ \\int\_\{\\mathcal\{X\}\}\\gamma\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{x\}=\\mu\_\{1\}\(\\bm\{y\}\)\\right\\\}The static optimal transport \(OT\), also known as the Kantorovich form\([Kantorovich, 1958](https://arxiv.org/html/2609.20008#bib.bib20)\), is defined as

OT​\(μ0,μ1\)=infγ∈Π⁡\(μ0,μ1\)∫𝒳2C⁡\(𝒙,𝒚\)​γ​\(𝒙,𝒚\)​𝒅𝒙​𝒅𝒚\\displaystyle\\text\{OT\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\inf\_\{\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)\}\\int\_\{\\mathcal\{X\}^\{2\}\}\\,C\(\\bm\{x\},\\bm\{y\}\)\\gamma\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\(1\)whereC⁡\(𝒙,𝒚\)C\(\\bm\{x\},\\bm\{y\}\)represents the cost of transporting unit mass from𝒙\\bm\{x\}to𝒚\\bm\{y\}, andγ⁡\(𝒙,𝒚\)∈Π⁡\(μ0,μ1\)\\gamma\(\\bm\{x\},\\bm\{y\}\)\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)is called the coupling\. From an optimization perspective, \([1](https://arxiv.org/html/2609.20008#S3.E1)\) can be viewed as a linear programming w\.r\.t the couplingγ\\gamma\. As a well known example, when choosingC⁡\(𝒙,𝒚\)=12​‖𝒙−𝒚‖2C\(\\bm\{x\},\\bm\{y\}\)=\\frac\{1\}\{2\}\\\|\\bm\{x\}\-\\bm\{y\}\\\|^\{2\}, which is the square of the Euclidean distance, the infimum is called the square of 2\-Wasserstein distance \(𝒲2\\mathcal\{W\}\_\{2\}\)\.

Static quadratic\-form optimal transport\.Recently, a static quadrtic\-from optimal transport \(QOT\) problem is defined and studied mathematically\([Wang and Zhang, 2025](https://arxiv.org/html/2609.20008#bib.bib6)\)\. Given two spaces𝒳,𝒴\\mathcal\{X\},\\mathcal\{Y\}, and a cost functionC:\(𝒳×𝒴\)2→ℝC:\(\\mathcal\{X\}\\times\\mathcal\{Y\}\)^\{2\}\\to\\mathbb\{R\}, the static QOT is defined as

QOT​\(μ0,μ1\)=infγ∈Π⁡\(μ0,μ1\)∫𝒳4C⁡\(𝒙,𝒚,𝒙′,𝒚′\)​γ​\(𝒙,𝒚\)​γ​\(𝒙′,𝒚′\)​𝑑𝒙​𝑑𝒚​d​𝒙′​d​𝒚′\\displaystyle\\text\{QOT\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\inf\_\{\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)\}\\int\_\{\\mathcal\{X\}^\{4\}\}\\,C\(\\bm\{x\},\\bm\{y\},\\bm\{x\}^\{\\prime\},\\bm\{y\}^\{\\prime\}\)\\gamma\(\\bm\{x\},\\bm\{y\}\)\\gamma\(\\bm\{x\}^\{\\prime\},\\bm\{y\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}\\bm\{x\}^\{\\prime\}\\mathrm\{d\}\\bm\{y\}^\{\\prime\}\(2\)If we chooseC⁡\(𝒙,𝒚,𝒙′,𝒚′\)=\|d𝒳​\(𝒙,𝒙′\)−d𝒴​\(𝒚,𝒚′\)\|2C\(\\bm\{x\},\\bm\{y\},\\bm\{x\}^\{\\prime\},\\bm\{y\}^\{\\prime\}\)=\|d\_\{\\mathcal\{X\}\}\(\\bm\{x\},\\bm\{x\}^\{\\prime\}\)\-d\_\{\\mathcal\{Y\}\}\(\\bm\{y\},\\bm\{y\}^\{\\prime\}\)\|^\{2\}where\(𝒳,d𝒳\),\(𝒴,d𝒴\)\(\\mathcal\{X\},d\_\{\\mathcal\{X\}\}\),\(\\mathcal\{Y\},d\_\{\\mathcal\{Y\}\}\)are two metric spaces, \([2](https://arxiv.org/html/2609.20008#S3.E2)\) recovers the Gromov\-Wasserstein OT \(GW\-OT\)\. This type of QOT has been widely used in single\-cell trajectory inference\([Mémoli, 2011](https://arxiv.org/html/2609.20008#bib.bib35);[Klein et al\., 2025](https://arxiv.org/html/2609.20008#bib.bib32)\)\. Intuitively, GW\-OT preserves the local structure after transport\. If we chooseC⁡\(𝒙,𝒚,𝒙′,𝒚′\)=f⁡\(𝒙,𝒚\)\+g⁡\(𝒙′,𝒚′\)C\(\\bm\{x\},\\bm\{y\},\\bm\{x\}^\{\\prime\},\\bm\{y\}^\{\\prime\}\)=f\(\\bm\{x\},\\bm\{y\}\)\+g\(\\bm\{x\}^\{\\prime\},\\bm\{y\}^\{\\prime\}\), \([2](https://arxiv.org/html/2609.20008#S3.E2)\) just reduces to the standard static OT with costf\+gf\+g\. Different to static OT, static QOT is a quadratic programming w\.r\.t the couplingγ\\gamma\.

Dynamic optimal transport\.For static OT,\([Benamou and Brenier, 2000](https://arxiv.org/html/2609.20008#bib.bib21)\)establised a dynamic form for𝒲2\\mathcal\{W\}\_\{2\}case, also known as the BB\-form\.

OT​\(μ0,μ1\)=infρ,𝒖∫01∫𝒳12​‖𝒖⁡\(𝒙,t\)‖22​ρt​\(𝒙\)​𝑑𝒙​𝑑t\\displaystyle\\text\{OT\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\inf\_\{\\rho,\\bm\{u\}\}\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}\}\\,\\frac\{1\}\{2\}\\\|\\bm\{u\}\(\\bm\{x\},t\)\\\|\_\{2\}^\{2\}\\rho\_\{t\}\(\\bm\{x\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}t\(3\)s\.t\.∂tρ\+∇𝒙⋅\(ρ​𝒖\)=0,ρ0=μ0,ρ1=μ1\\displaystyle\\text\{s\.t\.\}\\ \\ \\ \\ \\ \\ \\ \\partial\_\{t\}\\rho\+\\nabla\_\{\\bm\{x\}\}\\cdot\(\\rho\\bm\{u\}\)=0,\\ \\rho\_\{0\}=\\mu\_\{0\},\\ \\rho\_\{1\}=\\mu\_\{1\}They proved the equivalence between \([3](https://arxiv.org/html/2609.20008#S3.E3)\) and \([1](https://arxiv.org/html/2609.20008#S3.E1)\) whenC⁡\(𝒙,𝒚\)=12​‖𝒙−𝒚‖2C\(\\bm\{x\},\\bm\{y\}\)=\\frac\{1\}\{2\}\\\|\\bm\{x\}\-\\bm\{y\}\\\|^\{2\}\. Intuitively, it aims to find a continuous probability flow connectingμ0,μ1\\mu\_\{0\},\\mu\_\{1\}which also minimizes the total kinetic energy\. With nice fluid dynamics interpretation, it has also been widely used in single\-cell trajectory inference\([Tong et al\., 2020](https://arxiv.org/html/2609.20008#bib.bib34);[Tong et al\., 2024a](https://arxiv.org/html/2609.20008#bib.bib19);[Klein et al\., 2024](https://arxiv.org/html/2609.20008#bib.bib33)\)\.

Travelling Dirac\.To solve dynamic OT, one can first consider the dynamic OT between two Dirac measures

OT​\(δ𝒙0,δ𝒙1\)=inf𝒙t12​∫01‖𝒙t˙‖22​𝑑ts\.t\.𝒙0=𝒙0,𝒙1=𝒙1\\text\{OT\}\(\\delta\_\{\\bm\{x\}\_\{0\}\},\\delta\_\{\\bm\{x\}\_\{1\}\}\)=\\inf\_\{\\bm\{x\}\_\{t\}\}\\frac\{1\}\{2\}\\int\_\{0\}^\{1\}\\\|\\dot\{\\bm\{x\}\_\{t\}\}\\\|\_\{2\}^\{2\}\\mathrm\{d\}t\\quad\\text\{s\.t\.\}\\quad\\bm\{x\}\_\{0\}=\\bm\{x\}\_\{0\},\\quad\\bm\{x\}\_\{1\}=\\bm\{x\}\_\{1\}\(4\)which yields the displacement interpolation𝒙t=\(1−t\)​𝒙0\+t​𝒙1\\bm\{x\}\_\{t\}=\(1\-t\)\\bm\{x\}\_\{0\}\+t\\bm\{x\}\_\{1\}, also known as the travelling Dirac\. More complex travelling Dirac can also be derived for different type of OT, such as unbalanced OT\([Chizat et al\., 2018b](https://arxiv.org/html/2609.20008#bib.bib23);[Chizat et al\., 2018a](https://arxiv.org/html/2609.20008#bib.bib22)\)\. For a specific dynamic OT, previous works obtained its solution by integrating these travelling Diracs over the static couplingγ\\gamma\([Tong et al\., 2024a](https://arxiv.org/html/2609.20008#bib.bib19);[Peng et al\., 2026b](https://arxiv.org/html/2609.20008#bib.bib14)\)\. Therefore, the dynamic OT can decoupled to two parts: the travelling Dirac and the static coupling\. Integration can be realized by flow matching\.

Flow matching\.Flow matching is a simulation\-free generative framework for learning a continuous probability flow from data\([Lipman et al\., 2023](https://arxiv.org/html/2609.20008#bib.bib18)\)\. Givenμ0,μ1\\mu\_\{0\},\\mu\_\{1\}, it aims to learn a marginal velocity field𝒖t​\(𝒙\)\\bm\{u\}\_\{t\}\(\\bm\{x\}\), such that∂tρ\+∇𝒙⋅\(ρ​𝒖\)=0\\partial\_\{t\}\\rho\+\\nabla\_\{\\bm\{x\}\}\\cdot\(\\rho\\bm\{u\}\)=0andρ0=μ0,ρ1=μ1\\rho\_\{0\}=\\mu\_\{0\},\\rho\_\{1\}=\\mu\_\{1\}, which means𝒖t​\(𝒙\)\\bm\{u\}\_\{t\}\(\\bm\{x\}\)transportsμ0\\mu\_\{0\}toμ1\\mu\_\{1\}\. They parameterized a neural network𝒖𝜽\\bm\{u\}\_\{\\bm\{\\theta\}\}to approximate the true𝒖\\bm\{u\}\. The neural networks are trained to minimize the marginal regression loss\.

ℒFM​\(𝜽\)=∫01∫𝒳‖𝒖𝜽​\(𝒙,t\)−𝒖t​\(𝒙\)‖22​ρt​\(𝒙\)​𝑑𝒙​𝑑t\\displaystyle\\mathcal\{L\}\_\{\\text\{FM\}\}\(\\bm\{\\theta\}\)=\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}\}\\left\\\|\\bm\{\\bm\{u\}\_\{\\theta\}\}\(\\bm\{x\},t\)\-\\bm\{u\}\_\{t\}\(\\bm\{x\}\)\\right\\\|\_\{2\}^\{2\}\\rho\_\{t\}\(\\bm\{x\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}t\(5\)Although the trueρ,𝒖\\rho,\\bm\{u\}are intractable, they proved that minimizing the loss above is equivalent to minimize the conditional regression loss

\\displaystyleℒCFM​\(𝜽\)=𝔼t∼𝒰⁡\[0,1\],𝒛∼q⁡\(𝒛\),𝒙∼ρt​\(𝒙\|𝒛\)​‖𝒖𝜽​\(𝒙,t\)−𝒖t​\(𝒙\|𝒛\)‖22\\displaystyle\\mathcal\{L\}\_\{\\text\{CFM\}\}\(\\bm\{\\theta\}\)=\\mathbb\{E\}\_\{t\\sim\\mathcal\{U\}\[0,1\],\\bm\{z\}\\sim q\(\\bm\{z\}\),\\bm\{x\}\\sim\\rho\_\{t\}\(\\bm\{x\}\|\\bm\{z\}\)\}\\left\\\|\\bm\{\\bm\{u\}\_\{\\theta\}\}\(\\bm\{x\},t\)\-\\bm\{u\}\_\{t\}\(\\bm\{x\}\|\\bm\{z\}\)\\right\\\|\_\{2\}^\{2\}\(6\)where𝒛∼q⁡\(𝒛\)\\bm\{z\}\\sim q\(\\bm\{z\}\)is some conditional variable and∂tρt​\(𝒙\|𝒛\)\+∇𝒙⋅\(ρt​\(𝒙\|𝒛\)​𝒖t​\(𝒙\|𝒛\)\)=0\\partial\_\{t\}\\rho\_\{t\}\(\\bm\{x\}\|\\bm\{z\}\)\+\\nabla\_\{\\bm\{x\}\}\\cdot\(\\rho\_\{t\}\(\\bm\{x\}\|\\bm\{z\}\)\\bm\{u\}\_\{t\}\(\\bm\{x\}\|\\bm\{z\}\)\)=0is called the conditional continuity equation, which the conditional probability flowρt​\(𝒙\|𝒛\)\\rho\_\{t\}\(\\bm\{x\}\|\\bm\{z\}\)and the conditional velocity𝒖t​\(𝒙\|𝒛\)\\bm\{u\}\_\{t\}\(\\bm\{x\}\|\\bm\{z\}\)satisfy\. The marginal probability flow satisfiesρt​\(𝒙\)=∫ρt​\(𝒙\|𝒛\)​q​\(𝒛\)​𝑑𝒛\\rho\_\{t\}\(\\bm\{x\}\)=\\int\\rho\_\{t\}\(\\bm\{x\}\|\\bm\{z\}\)q\(\\bm\{z\}\)\\mathrm\{d\}\\bm\{z\}\. The marginalization theorem states that the marginal velocity𝒖t​\(𝒙\)=∫𝒖t​\(𝒙\|𝒛\)​ρt​\(𝒙\|𝒛\)​q​\(𝒛\)ρt​\(𝒙\)​𝑑𝒛\\bm\{u\}\_\{t\}\(\\bm\{x\}\)=\\int\\bm\{u\}\_\{t\}\(\\bm\{x\}\|\\bm\{z\}\)\\frac\{\\rho\_\{t\}\(\\bm\{x\}\|\\bm\{z\}\)q\(\\bm\{z\}\)\}\{\\rho\_\{t\}\(\\bm\{x\}\)\}\\mathrm\{d\}\\bm\{z\}transportsμ0\\mu\_\{0\}toμ1\\mu\_\{1\}\. Therefore, one can learn a admissible𝒖t​\(𝒙\)\\bm\{u\}\_\{t\}\(\\bm\{x\}\)by designing tractableq⁡\(𝒛\)q\(\\bm\{z\}\)and conditional pathρt​\(𝒙\|𝒛\)\\rho\_\{t\}\(\\bm\{x\}\|\\bm\{z\}\)with tractable conditional velocity𝒖t​\(𝒙\|𝒛\)\\bm\{u\}\_\{t\}\(\\bm\{x\}\|\\bm\{z\}\), and applying flow matching\.

By careful design, flow matching can be used for solving dynamic OT\. Following\([Tong et al\., 2024a](https://arxiv.org/html/2609.20008#bib.bib19)\), one can choose𝒛=\(𝒙0,𝒙1\)∼γ⁡\(𝒙0,𝒙1\)\\bm\{z\}=\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\}\)\\sim\\gamma\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\}\)drawn from the static OT coupling, set the travelling Dirac as conditional path, and derive the conditional velocity from it\. The resulting flow are proved to recover the dynamic OT flow\. Under this framework, travelling Diracs are integrated over the static OT coupling independently, which means particles actually move independently\. In this work, we develop a flow matching framework which allows conditional paths to interact\.

## 4Dynamic QOT

In this section, we consider one metric space𝒳⊂ℝd\\mathcal\{X\}\\subset\\mathbb\{R\}^\{d\}, i\.e\.𝒴=𝒳\\mathcal\{Y\}=\\mathcal\{X\}, and develop a dynamic formulation for QOT under this condition\. Since the flow have to live in some space, the meaning of a dynamic flow connecting two totally different spaces needs further study\. All the proofs are left to[A](https://arxiv.org/html/2609.20008#A1)\.

### 4\.1Static and Dynamic form

Path action\.To model conditional paths with interaction, we study how a pair is transported to a pair rather than travelling Dirac\. We call the corresponding pathtravelling pair\. Let𝒛=\(𝒙0,𝒙1\),𝒛′=\(𝒚0,𝒚1\)\\bm\{z\}=\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\}\),\\bm\{z\}^\{\\prime\}=\(\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\), we define the action of a pair path for some Lagrangianℒ\\mathcal\{L\}with sufficient regularity

𝒜⁡\(𝒛,𝒛′\)=inf𝒙t,𝒚t∫01ℒ⁡\(t,𝒙t,𝒚t,𝒙˙t,𝒚˙t\)​𝑑t\\mathcal\{A\}\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)=\\inf\_\{\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\}\}\\int\_\{0\}^\{1\}\\mathcal\{L\}\(t,\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\},\\dot\{\\bm\{x\}\}\_\{t\},\\dot\{\\bm\{y\}\}\_\{t\}\)\\mathrm\{d\}t\(7\)The minimizer yields a conditional velocity pair𝒙˙t=𝒖t1\(𝒙t,𝒚t\|𝒛,𝒛′\),𝒚˙t=𝒖t2\(𝒙t,𝒚t\|𝒛,𝒛′\)\\dot\{\\bm\{x\}\}\_\{t\}=\\bm\{u\}\_\{t\}^\{1\}\(\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\),\\dot\{\\bm\{y\}\}\_\{t\}=\\bm\{u\}\_\{t\}^\{2\}\(\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\. It can also be viewed as over the travelling Pair measure pathπt\(𝒙,𝒚\|𝒛,𝒛′\)\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)

𝒜\(𝒛,𝒛′\)=infπ,𝒖1,𝒖2∫01∫𝒳2ℒ\(t,𝒙,𝒚,u1t\(𝒙,𝒚\|𝒛,𝒛′\),u2t\(𝒙,𝒚\|𝒛,𝒛′\)\)πt\(𝒙,𝒚\|𝒛,𝒛′\)d𝒙d𝒚dt\\displaystyle\\mathcal\{A\}\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)=\\inf\_\{\\pi,\\bm\{u\}^\{1\},\\bm\{u\}^\{2\}\}\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}^\{2\}\}\\mathcal\{L\}\(t,\\bm\{x\},\\bm\{y\},u^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\),u^\{2\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\)\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}t\(8\)∂tπt\(𝒙,𝒚\|𝒛,𝒛′\)\+∇𝒙⋅\(πt\(𝒙,𝒚\|𝒛,𝒛′\)𝒖t1\(𝒙,𝒚\|𝒛,𝒛′\)\)\+∇𝒚⋅\(πt\(𝒙,𝒚\|𝒛,𝒛′\)𝒖t2\(𝒙,𝒚\|𝒛,𝒛′\)\)=0\\displaystyle\\partial\_\{t\}\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\+\\nabla\_\{\\bm\{x\}\}\\cdot\(\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\bm\{u\}\_\{t\}^\{1\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\)\+\\nabla\_\{\\bm\{y\}\}\\cdot\(\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\bm\{u\}\_\{t\}^\{2\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\)=0\\with boundary conditionπ0\(𝒙,𝒚\|𝒛,𝒛′\)=δ\(𝒙0,𝒚0\)\(𝒙,𝒚\),π1\(𝒙,𝒚\|𝒛,𝒛′\)=δ\(𝒙1,𝒚1\)\(𝒙,𝒚\)\\pi\_\{0\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)=\\delta\_\{\(\\bm\{x\}\_\{0\},\\bm\{y\}\_\{0\}\)\}\(\\bm\{x\},\\bm\{y\}\),\\pi\_\{1\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)=\\delta\_\{\(\\bm\{x\}\_\{1\},\\bm\{y\}\_\{1\}\)\}\(\\bm\{x\},\\bm\{y\}\)\.

Static form\.Based on the path action𝒜\\mathcal\{A\}, we can define the corresponding static QOT as

QOTS​\(μ0,μ1\)=infγ∈Π⁡\(μ0,μ1\)∫𝒳4𝒜⁡\(𝒛,𝒛′\)​γ​\(𝒛\)​γ​\(𝒛′\)​𝑑𝒛​d​𝒛′\\displaystyle\\text\{QOT\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\inf\_\{\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)\}\\int\_\{\\mathcal\{X\}^\{4\}\}\\mathcal\{A\}\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\gamma\(\\bm\{z\}\)\\gamma\(\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}\(9\)The optimal coupling can be solved by standard quadratic programming algorithms, such as Frank\-Wolfe\([Kerdoncuff et al\., 2021](https://arxiv.org/html/2609.20008#bib.bib8);[Flamary et al\., 2021](https://arxiv.org/html/2609.20008#bib.bib30)\)\. We left the details to[B\.8](https://arxiv.org/html/2609.20008#A2.SS8)\.

Dynamic form\.Inspired by the travelling Dirac technique in standard OT, we define the dynamic QOT on the travelling pair path level\. The corresponding dynamic probability flow isπt\(𝒙,𝒚\)=∫πt\(𝒙,𝒚\|𝒛,𝒛′\)γ\(𝒛\)γ\(𝒛′\)d𝒛d𝒛′\.\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)=\\int\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\gamma\(\\bm\{z\}\)\\gamma\(\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}\.

QOTD​\(μ0,μ1\)=\\displaystyle\\text\{QOT\}\_\{D\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\(10\)inf∫01∫∫𝒳2ℒ\(t,𝒙,𝒚,u1t\(𝒙,𝒚\|𝒛,𝒛′\),u2t\(𝒙,𝒚\|𝒛,𝒛′\)\)πt\(𝒙,𝒚\|𝒛,𝒛′\)γ\(𝒛\)γ\(𝒛′\)d𝒙d𝒚d𝒛d𝒛′dt\\displaystyle\\inf\\int\_\{0\}^\{1\}\\int\\int\_\{\\mathcal\{X\}^\{2\}\}\\mathcal\{L\}\(t,\\bm\{x\},\\bm\{y\},u^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\),u^\{2\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\)\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\gamma\(\\bm\{z\}\)\\gamma\(\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}\\mathrm\{d\}t
The infimum is taken over the couplingγ\\gammaand all travelling pairs\. A fluid dynamics form or say BB\-form, analog to OT\([Benamou and Brenier, 2000](https://arxiv.org/html/2609.20008#bib.bib21)\), is also defined as

QOTB​B​\(μ0,μ1\)=infπ,𝒖1,𝒖2∫01∫𝒳2ℒ⁡\(t,𝒙,𝒚,ut1​\(𝒙,𝒚\),ut2​\(𝒙,𝒚\)\)​πt​\(𝒙,𝒚\)​𝑑𝒙​𝑑𝒚​𝑑t\\displaystyle\\text\{QOT\}\_\{BB\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\inf\_\{\\pi,\\bm\{u\}^\{1\},\\bm\{u\}^\{2\}\}\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}^\{2\}\}\\mathcal\{L\}\(t,\\bm\{x\},\\bm\{y\},u^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\),u^\{2\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\)\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}t\(11\)∂tπt​\(𝒙,𝒚\)\+∇𝒙⋅\(πt​\(𝒙,𝒚\)​ut1​\(𝒙,𝒚\)\)\+∇𝒚⋅\(πt​\(𝒙,𝒚\)​ut2​\(𝒙,𝒚\)\)=0\\displaystyle\\partial\_\{t\}\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\+\\nabla\_\{\\bm\{x\}\}\\cdot\(\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)u^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\)\+\\nabla\_\{\\bm\{y\}\}\\cdot\(\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)u^\{2\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\)=0whereπt\\pi\_\{t\}is a continuous probability flow on the two\-particle space𝒳2\\mathcal\{X\}^\{2\}\. The marginal velocity pair𝒖t1,𝒖t2\\bm\{u\}^\{1\}\_\{t\},\\bm\{u\}^\{2\}\_\{t\}are particle velocities with interaction effect\. The formulation intuitively seeks for a two\-particle flow minimizing the total action\. Our first result is the equivalence between the static and dynamic form\. However, these forms are upper bound of the BB\-form but not equal to it, hence a surrogate\. We discuss the details and difficulties in[F](https://arxiv.org/html/2609.20008#A6)\. Fortunately, since both the static and dynamic formulations provide upper bounds on the BB\-form, solving dynamic QOT also implicitly minimizes the total energy in the sense of the BB\-form\.

###### Theorem 4\.1\.

Ifℒ\\mathcal\{L\}is convex w\.r\.t\(𝐱˙,𝐲˙\)\(\\dot\{\\bm\{x\}\},\\dot\{\\bm\{y\}\}\), thenQOTS​\(μ0,μ1\)=QOTD​\(μ0,μ1\)≥QOTB​B​\(μ0,μ1\)\\text\{QOT\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\text\{QOT\}\_\{D\}\(\\mu\_\{0\},\\mu\_\{1\}\)\\geq\\text\{QOT\}\_\{BB\}\(\\mu\_\{0\},\\mu\_\{1\}\)\.

### 4\.2Lagrangian with tractable conditional velocity pair

In the next section, we established a flow matching framework for solving dynamic QOT flows\. In order to do that, we need to access to the conditional velocity pair\. Therefore, we then consider what kind ofℒ\\mathcal\{L\}leads to tractable conditional velocity\. Note that TP\-DATE is not restricted to the choices introduced below\. As long as the conditional velocity pair can be obtained in some way, TP\-DATE remains applicable\. To take kinetic energy into account, we consider a broad class ofℒ\\mathcal\{L\}with the form

ℒ⁡\(t,𝒙,𝒚,𝒙˙,𝒚˙\)=12​‖𝒙˙‖2\+12​‖𝒚˙‖2\+λ​Φ​\(𝒙−𝒚,𝒙˙−𝒚˙\)\\mathcal\{L\}\(t,\\bm\{x\},\\bm\{y\},\\dot\{\\bm\{x\}\},\\dot\{\\bm\{y\}\}\)=\\frac\{1\}\{2\}\\\|\\dot\{\\bm\{x\}\}\\\|^\{2\}\+\\frac\{1\}\{2\}\\\|\\dot\{\\bm\{y\}\}\\\|^\{2\}\+\\lambda\\Phi\(\\bm\{x\}\-\\bm\{y\},\\dot\{\\bm\{x\}\}\-\\dot\{\\bm\{y\}\}\)\(12\)whereΦ\\Phirepresents a convex interaction term, making it different to standard OT\. To simplify the pair dynamics, we change the variables to𝒄t=𝒙t\+𝒚t2,𝒒t=𝒚t−𝒙t\\bm\{c\}\_\{t\}=\\frac\{\\bm\{x\}\_\{t\}\+\\bm\{y\}\_\{t\}\}\{2\},\\bm\{q\}\_\{t\}=\\bm\{y\}\_\{t\}\-\\bm\{x\}\_\{t\}, which are the barycenter and the ralative displacement of the pair\. The Lagrangian becomes

ℒ⁡\(t,𝒙,𝒚,𝒙˙,𝒚˙\)=‖𝒄˙t‖2\+14​‖𝒒˙t‖2\+λ​Φ​\(𝒒,𝒒˙\)\\mathcal\{L\}\(t,\\bm\{x\},\\bm\{y\},\\dot\{\\bm\{x\}\},\\dot\{\\bm\{y\}\}\)=\\\|\\dot\{\\bm\{c\}\}\_\{t\}\\\|^\{2\}\+\\frac\{1\}\{4\}\\\|\\dot\{\\bm\{q\}\}\_\{t\}\\\|^\{2\}\+\\lambda\\Phi\(\\bm\{q\},\\dot\{\\bm\{q\}\}\)\(13\)Since the convexity preserves under affine transformation, it is sufficient to chooseΦ\\Phito be convex w\.r\.t𝒒˙\\dot\{\\bm\{q\}\}\. One natural choice isΦ⁡\(𝒒\)\\Phi\(\\bm\{q\}\), which is independent to𝒒˙\\dot\{\\bm\{q\}\}\. IfΦ\\Phiis further radial, then the travelling pair has a low dimensional structure, hence learnable without the curse of dimensionality\.

###### Proposition 4\.2\.

The optimal𝐜t\\bm\{c\}\_\{t\}is𝐜t=\(1−t\)​𝐜0\+t​𝐜1\\bm\{c\}\_\{t\}=\(1\-t\)\\bm\{c\}\_\{0\}\+t\\bm\{c\}\_\{1\}\. IfΦ=Φ⁡\(‖𝐪‖\)\\Phi=\\Phi\(\\\|\\bm\{q\}\\\|\),𝐪0∦𝐪1\\bm\{q\}\_\{0\}\\nparallel\\bm\{q\}\_\{1\}, then𝐪t∈span​\{𝐪0,𝐪1\}\\bm\{q\}\_\{t\}\\in\\text\{span\}\\\{\\bm\{q\}\_\{0\},\\bm\{q\}\_\{1\}\\\}\.

GW\-OT as a limit case\.Another family of meaningfulΦ\\PhiisΦ=\|dd​t​ϕ​\(𝒒t\)\|2\\Phi=\|\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\phi\(\\bm\{q\}\_\{t\}\)\|^\{2\}\. It penalizes the variation of the relative displacement\. The most explicit choice isΦ=\|dd​t​‖𝒒t‖\|2\\Phi=\|\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\\|\\bm\{q\}\_\{t\}\\\|\|^\{2\}\. Intuitively, this seeks for a transport not only saving kinetic energy, but also trying to minimize the distortion\. ThisΦ\\Phileads to analytic conditional velocity\.

###### Proposition 4\.3\.

IfΦ=\|dd​t​ϕ​\(𝐪t\)\|2\\Phi=\|\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\phi\(\\bm\{q\}\_\{t\}\)\|^\{2\},Φ\\Phiis convex w\.r\.t𝐪˙t\\dot\{\\bm\{q\}\}\_\{t\}\. IfΦ=\|dd​t​‖𝐪t‖\|2\\Phi=\|\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\\|\\bm\{q\}\_\{t\}\\\|\|^\{2\},𝐪t\\bm\{q\}\_\{t\}has analytic solution\.

A notable property of this choice ofΦ=\|dd​t​‖𝒒t‖\|2\\Phi=\|\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\\|\\bm\{q\}\_\{t\}\\\|\|^\{2\}is the following theorem\.

###### Theorem 4\.4\.

Φ=\|dd​t​‖𝒒t‖\|2\\Phi=\|\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\\|\\bm\{q\}\_\{t\}\\\|\|^\{2\}\. Whenλ=0\\lambda=0andd≥2d\\geq 2, the corresponding static QOT reduces to standard OT, while whenλ→\+∞\\lambda\\to\+\\infty, it reduces to standard GW\-OT:λ−1​QOTS​\(μ0,μ1\)→GW\-OT​\(μ0,μ1\)\\lambda^\{\-1\}\\text\{QOT\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)\\to\\text\{GW\-OT\}\(\\mu\_\{0\},\\mu\_\{1\}\)\.

We therefore refer to this choice of Lagrangian as adynamic fusion of OT and GW\-OT\. Of note, this is different from the existing FGW\-OT formulation\([Vayer et al\., 2020](https://arxiv.org/html/2609.20008#bib.bib9);[Klein et al\., 2025](https://arxiv.org/html/2609.20008#bib.bib32)\)\. FGW\-OT is obtained by taking a weighted combination directly at the static level, whereas dynamic fusion introduces the weighting at the level of the dynamic formulation\. Consequently, its induced static formulation is not FGW\-OT\. Also, by choosing properΦ\\Phi, the recently considered static IGW\-OT\([Zhang et al\., 2026b](https://arxiv.org/html/2609.20008#bib.bib7)\)can also be realized as a special case of our dynamic QOT framework\. We discuss the relation between these formulations in[D\.1](https://arxiv.org/html/2609.20008#A4.SS1)[D\.2](https://arxiv.org/html/2609.20008#A4.SS2)\.

Modality separation\.In some settings, the coordinates may consist of multiple modalities, and we may wish to impose different interaction terms on different modalities\. We illustrate the formulation using the two modality case, the extension to multiple modalities follows naturally\. Let𝒙=\(𝒙1,𝒙2\),𝒚=\(𝒚1,𝒚2\)\\bm\{x\}=\(\\bm\{x\}^\{1\},\\bm\{x\}^\{2\}\),\\bm\{y\}=\(\\bm\{y\}^\{1\},\\bm\{y\}^\{2\}\), and the Lagrangian is separated to the two modalities

ℒ⁡\(t,𝒙,𝒚,𝒙˙,𝒚˙\)=ℒ1​\(t,𝒙1,𝒚1,𝒙˙1,𝒚˙1\)\+ℒ2​\(t,𝒙2,𝒚2,𝒙˙2,𝒚˙2\)\.\\mathcal\{L\}\(t,\\bm\{x\},\\bm\{y\},\\dot\{\\bm\{x\}\},\\dot\{\\bm\{y\}\}\)=\\mathcal\{L\}\_\{1\}\(t,\\bm\{x\}^\{1\},\\bm\{y\}^\{1\},\\dot\{\\bm\{x\}\}^\{1\},\\dot\{\\bm\{y\}\}^\{1\}\)\+\\mathcal\{L\}\_\{2\}\(t,\\bm\{x\}^\{2\},\\bm\{y\}^\{2\},\\dot\{\\bm\{x\}\}^\{2\},\\dot\{\\bm\{y\}\}^\{2\}\)\.\(14\)Then one can easily check that the conditional velocities of the two modalities can also be computed separately\. The conditional velocity pair of\(𝒙,𝒚\)\(\\bm\{x\},\\bm\{y\}\)is exactly the concatenation of the conditional velocity pairs induced byℒi\\mathcal\{L\}\_\{i\}\.

## 5Travelling pair flow matching

In this section, we develop a flow matching framework for solving the dynamic QOT problem\. As OT\-CFM\([Tong et al\., 2024a](https://arxiv.org/html/2609.20008#bib.bib19)\)averages travelling Diracs over OT coupling, our framework allows to average travelling pairs over QOT coupling\. All the proofs are left to[A](https://arxiv.org/html/2609.20008#A1)\.

### 5\.1Marginalization theory

Pair marginalization\.Given conditional probability pathsπt\(𝒙,𝒚\|𝒛,𝒛′\)\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)and conditional velocity pairs𝒖t1\(𝒙,𝒚\|𝒛,𝒛′\),𝒖t2\(𝒙,𝒚\|𝒛,𝒛′\)\\bm\{u\}\_\{t\}^\{1\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\),\\bm\{u\}\_\{t\}^\{2\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)satisfying the conditional continuity equation, let the conditional variables\(𝒛,𝒛′\)∼q⁡\(𝒛,𝒛′\)\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\sim q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\), and define the marginal probability flow asπt\(𝒙,𝒚\)=∫πt\(𝒙,𝒚\|𝒛,𝒛′\)q\(𝒛,𝒛′\)d𝒛d𝒛′\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)=\\int\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}, the marginalization theorem on the two\-particle space holds\.

###### Theorem 5\.1\.

Define the marginal velocity as𝐮ti\(𝐱,𝐲\)=∫𝐮ti\(𝐱,𝐲\|𝐳,𝐳′\)πt\(𝐱,𝐲\|𝐳,𝐳′\)q\(𝐳,𝐳′\)πt​\(𝐱,𝐲\)d𝐳d𝐳′\\bm\{u\}^\{i\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)=\\int\\bm\{u\}^\{i\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\frac\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\}\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\}\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}, the marginals satisfiy the marginal continuity equation

∂tπt\+∇𝒙⋅\(πt​𝒖t1\)\+∇𝒚⋅\(πt​𝒖t2\)=0\\partial\_\{t\}\\pi\_\{t\}\+\\nabla\_\{\\bm\{x\}\}\\cdot\(\\pi\_\{t\}\\bm\{u\}^\{1\}\_\{t\}\)\+\\nabla\_\{\\bm\{y\}\}\\cdot\(\\pi\_\{t\}\\bm\{u\}^\{2\}\_\{t\}\)=0\(15\)Ifq=γ⊗γq=\\gamma\\otimes\\gammafor someγ∈Π⁡\(μ0,μ1\)\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\), the marginal velocity transportsμ0⊗μ0\\mu\_\{0\}\\otimes\\mu\_\{0\}toμ1⊗μ1\\mu\_\{1\}\\otimes\\mu\_\{1\}\.

Particle marginalization\.The marginal velocity𝒖t1​\(𝒙,𝒚\)\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)can be interpreted as how particle𝒙\\bm\{x\}travels given the influence of𝒚\\bm\{y\}\. From this perspective, one can obtain a marginal velocity𝒗t​\(𝒙\)\\bm\{v\}\_\{t\}\(\\bm\{x\}\)only depends on𝒙\\bm\{x\}by averaging the influences of all possible𝒚\\bm\{y\}\.

𝒗t​\(𝒙\)=𝔼πt​\[𝒖t1​\(𝑿t,𝒀t\)\|𝑿t=𝒙\]=∫𝒳𝒖t1​\(𝒙,𝒚\)​πt​\(𝒚\|𝒙\)​𝑑𝒚\\bm\{v\}\_\{t\}\(\\bm\{x\}\)=\\mathbb\{E\}\_\{\\pi\_\{t\}\}\[\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{X\}\_\{t\},\\bm\{Y\}\_\{t\}\)\|\\bm\{X\}\_\{t\}=\\bm\{x\}\]=\\int\_\{\\mathcal\{X\}\}\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\pi\_\{t\}\(\\bm\{y\}\|\\bm\{x\}\)\\mathrm\{d\}\\bm\{y\}\(16\)This velocity only depends on the particle itself, but it already contains all the interaction information by taking conditional expectation\. We further define the𝒙\\bm\{x\}marginalρt​\(𝒙\)=∫𝒳πt​\(𝒙,𝒚\)​𝑑𝒚\\rho\_\{t\}\(\\bm\{x\}\)=\\int\_\{\\mathcal\{X\}\}\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{y\}\. For these single particle marginals, we also have a marginalization theorem\.

###### Theorem 5\.2\.

The single particle marginals satisfy the continuity equation

∂tρt\+∇𝒙⋅\(ρt​𝒗t\)=0\\partial\_\{t\}\\rho\_\{t\}\+\\nabla\_\{\\bm\{x\}\}\\cdot\(\\rho\_\{t\}\\bm\{v\}\_\{t\}\)=0\(17\)

Equivalently, it means𝒗t​\(𝒙\)\\bm\{v\}\_\{t\}\(\\bm\{x\}\)transportsμ0\\mu\_\{0\}toμ1\\mu\_\{1\}\. It can be viewed as the projection of the pair flow on the single particle space\. However, it is worth clarifying that the two\-particle dynamics generally cannot be recovered from the single\-particle marginal dynamics\. They are not equivalent\.

Mean\-field approximation\.We further decouple the velocity𝒖t1\\bm\{u\}^\{1\}\_\{t\}into two parts𝒖t1​\(𝒙,𝒚\)=𝒃t​\(𝒙\)\+𝒇t​\(𝒙,𝒚\)\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)=\\bm\{b\}\_\{t\}\(\\bm\{x\}\)\+\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\.𝒃t\\bm\{b\}\_\{t\}represents the self\-driven term, and𝒇t\\bm\{f\}\_\{t\}represents the interaction term\. Under this decomposition, we have

𝒗t​\(𝒙\)=𝒃t​\(𝒙\)\+∫𝒳𝒇t​\(𝒙,𝒚\)​πt​\(𝒚\|𝒙\)​𝑑𝒚\\bm\{v\}\_\{t\}\(\\bm\{x\}\)=\\bm\{b\}\_\{t\}\(\\bm\{x\}\)\+\\int\_\{\\mathcal\{X\}\}\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\pi\_\{t\}\(\\bm\{y\}\|\\bm\{x\}\)\\mathrm\{d\}\\bm\{y\}\(18\)The particle continuity equation \([17](https://arxiv.org/html/2609.20008#S5.E17)\) becomes

∂tρt​\(𝒙\)\+∇𝒙⋅\[ρt​\(𝒙\)​\(𝒃t​\(𝒙\)\+∫𝒳𝒇t​\(𝒙,𝒚\)​πt​\(𝒚\|𝒙\)​𝑑𝒚\)\]=0\\partial\_\{t\}\\rho\_\{t\}\(\\bm\{x\}\)\+\\nabla\_\{\\bm\{x\}\}\\cdot\[\\rho\_\{t\}\(\\bm\{x\}\)\\big\(\\bm\{b\}\_\{t\}\(\\bm\{x\}\)\+\\int\_\{\\mathcal\{X\}\}\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\pi\_\{t\}\(\\bm\{y\}\|\\bm\{x\}\)\\mathrm\{d\}\\bm\{y\}\\big\)\]=0\(19\)
Under independence approximationπt​\(𝒙,𝒚\)≈ρt​\(𝒙\)​ρt​\(𝒚\)\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\approx\\rho\_\{t\}\(\\bm\{x\}\)\\rho\_\{t\}\(\\bm\{y\}\), \([19](https://arxiv.org/html/2609.20008#S5.E19)\) reduces to the mean\-field flow\.

∂tρt​\(𝒙\)\+∇𝒙⋅\[ρt​\(𝒙\)​\(𝒃t​\(𝒙\)\+∫𝒳𝒇t​\(𝒙,𝒚\)​ρt​\(𝒚\)​𝑑𝒚\)\]=0\\partial\_\{t\}\\rho\_\{t\}\(\\bm\{x\}\)\+\\nabla\_\{\\bm\{x\}\}\\cdot\[\\rho\_\{t\}\(\\bm\{x\}\)\\big\(\\bm\{b\}\_\{t\}\(\\bm\{x\}\)\+\\int\_\{\\mathcal\{X\}\}\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\rho\_\{t\}\(\\bm\{y\}\)\\mathrm\{d\}\\bm\{y\}\\big\)\]=0\(20\)Though the independence may not hold accurately due to the particle interactions, the difference between the continuous measure path generated by the mean\-field approximation \([20](https://arxiv.org/html/2609.20008#S5.E20)\) and the true dynamics \([19](https://arxiv.org/html/2609.20008#S5.E19)\) can be controlled by the difference betweenπt\\pi\_\{t\}andρt⊗ρt\\rho\_\{t\}\\otimes\\rho\_\{t\}\.

###### Proposition 5\.3\.

Let the true probability flow beρt\\rho\_\{t\}and the mean\-field approximation beρ^t\\hat\{\\rho\}\_\{t\}\. Assume𝐛t,𝐟t\\bm\{b\}\_\{t\},\\bm\{f\}\_\{t\}are both Lipschitz, andM=max0<t≤1max𝐱𝒲1\(ρt,πt\(⋅\|𝐱\)\)M=\\underset\{0<t\\leq 1\}\{\\max\}\\underset\{\\bm\{x\}\}\{\\max\}\\mathcal\{W\}\_\{1\}\(\\rho\_\{t\},\\pi\_\{t\}\(\\cdot\|\\bm\{x\}\)\), then∀0<t≤1,∃Ct\>0\\forall 0<t\\leq 1,\\exists C\_\{t\}\>0such that𝒲1​\(ρt,ρ^t\)≤Ct​M\\mathcal\{W\}\_\{1\}\(\\rho\_\{t\},\\hat\{\\rho\}\_\{t\}\)\\leq C\_\{t\}M\.

### 5\.2Training loss design

Marginal velocity pair\.Analog to conditional flow matching\([Lipman et al\., 2023](https://arxiv.org/html/2609.20008#bib.bib18)\), when travelling pairs are tractable, we can parameterize two neural networks𝒖𝜽i​\(𝒙,𝒚,t\),i=1,2\\bm\{u\}^\{i\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},\\bm\{y\},t\),i=1,2and regress the true marginal velocity pair by minimizing the conditional flow matching loss\. This further leads us to the solution to dynamic QOT problem \([10](https://arxiv.org/html/2609.20008#S4.E10)\)\.

ℒpair\(𝜽\)=𝔼t∼𝒰\[0,1\],\(𝒛,𝒛′\)∼q\(𝒛,𝒛′\),\(𝒙,𝒚\)∼πt\(𝒙,𝒚\|𝒛,𝒛′\)∑i=12‖𝒖𝜽𝒊\(𝒙,𝒚,t\)−𝒖ti\(𝒙,𝒚\|𝒛,𝒛′\)‖22\\mathcal\{L\}\_\{\\text\{pair\}\}\(\\bm\{\\theta\}\)=\\mathbb\{E\}\_\{t\\sim\\mathcal\{U\}\[0,1\],\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\sim q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\),\(\\bm\{x\},\\bm\{y\}\)\\sim\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\}\\sum\_\{i=1\}^\{2\}\\left\\\|\\bm\{\\bm\{u\}^\{i\}\_\{\\theta\}\}\(\\bm\{x\},\\bm\{y\},t\)\-\\bm\{u\}^\{i\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\right\\\|\_\{2\}^\{2\}\(21\)
###### Theorem 5\.4\.

The minimizer is𝐮𝛉𝐢​\(𝐱,𝐲,t\)=𝐮ti​\(𝐱,𝐲\)\\bm\{\\bm\{u\}^\{i\}\_\{\\theta\}\}\(\\bm\{x\},\\bm\{y\},t\)=\\bm\{u\}^\{i\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\. Whenq=γ⊗γq=\\gamma\\otimes\\gammawhereγ\\gammais the optimal QOT coupling of \([9](https://arxiv.org/html/2609.20008#S4.E9)\), and the travelling pair used is the minimizer of \([7](https://arxiv.org/html/2609.20008#S4.E7)\), then the learned velocity pair generates the probability flow of the dynamic QOT problem \([10](https://arxiv.org/html/2609.20008#S4.E10)\)\.

Particle velocity\.A new design is that we can also obtain the particle velocity𝒗t​\(𝒙\)\\bm\{v\}\_\{t\}\(\\bm\{x\}\)\([16](https://arxiv.org/html/2609.20008#S5.E16)\) by flow matching\. We can parameterize a single neural network𝒗𝜽​\(𝒙,t\)\\bm\{v\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},t\)and try to minimize the intractable marginal regression lossℒM​\(𝜽\)=𝔼t∼𝒰⁡\[0,1\],𝒙∼ρt​\(𝒙\)​‖𝒗𝜽​\(𝒙,t\)−𝒗t​\(𝒙\)‖2\\mathcal\{L\}\_\{M\}\(\\bm\{\\theta\)\}=\\mathbb\{E\}\_\{t\\sim\\mathcal\{U\}\[0,1\],\\bm\{x\}\\sim\\rho\_\{t\}\(\\bm\{x\}\)\}\\\|\\bm\{v\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},t\)\-\\bm\{v\}\_\{t\}\(\\bm\{x\}\)\\\|^\{2\}\. The minimizer is𝒗t​\(𝒙\)\\bm\{v\}\_\{t\}\(\\bm\{x\}\)\. A tractable conditional version is

ℒC\(𝜽\)=𝔼t∼𝒰\[0,1\],\(𝒛,𝒛′\)∼q\(𝒛,𝒛′\),\(𝒙,𝒚\)∼πt\(𝒙,𝒚\|𝒛,𝒛′\)∥𝒗𝜽\(𝒙,t\)−𝒖t1\(𝒙,𝒚\|𝒛,𝒛′\)∥2\\mathcal\{L\}\_\{C\}\(\\bm\{\\theta\)\}=\\mathbb\{E\}\_\{t\\sim\\mathcal\{U\}\[0,1\],\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\sim q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\),\(\\bm\{x\},\\bm\{y\}\)\\sim\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\}\\\|\\bm\{v\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},t\)\-\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\\|^\{2\}\(22\)
###### Theorem 5\.5\.

∇𝜽ℒM​\(𝜽\)=∇𝜽ℒC​\(𝜽\)\\nabla\_\{\\bm\{\\theta\}\}\\mathcal\{L\}\_\{M\}\(\\bm\{\\theta\}\)=\\nabla\_\{\\bm\{\\theta\}\}\\mathcal\{L\}\_\{C\}\(\\bm\{\\theta\}\)\.

In practice, if the symmetry condition holds, i\.e\.q=γ⊗γq=\\gamma\\otimes\\gammaandℒ\\mathcal\{L\}is symmetric to\(𝒙,𝒙˙\),\(𝒚,𝒚˙\)\(\\bm\{x\},\\dot\{\\bm\{x\}\}\),\(\\bm\{y\},\\dot\{\\bm\{y\}\}\), then it is easy to show that𝒖t1\(𝒚,𝒙\|𝒛,𝒛′\)=𝒖t2\(𝒙,𝒚\|𝒛,𝒛′\)\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{y\},\\bm\{x\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)=\\bm\{u\}^\{2\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\. In this case, the conditional loss \([22](https://arxiv.org/html/2609.20008#S5.E22)\) can be replaced by a convex combination of itself and∥𝒗𝜽\(𝒚,t\)−𝒖t2\(𝒙,𝒚\|𝒛,𝒛′\)∥2\\\|\\bm\{v\}\_\{\\bm\{\\theta\}\}\(\\bm\{y\},t\)\-\\bm\{u\}^\{2\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\\|^\{2\}to make the training more efficient\. In most QOT settings, this condition obviously holds\.

Velocity decomposition\.If there are some gauges of velocity decomposition𝒖t1\(𝒙,𝒚\|𝒛,𝒛′\)=𝒃t\(𝒙\)\+𝒇t\(𝒙,𝒚\|𝒛,𝒛′\)\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)=\\bm\{b\}\_\{t\}\(\\bm\{x\}\)\+\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\), it is also possible to regress the interaction part𝒇\\bm\{f\}via flow matching\. Let’s parameterize a neural network𝒇ϕ​\(𝒙,𝒚,t\)\\bm\{f\}\_\{\\bm\{\\phi\}\}\(\\bm\{x\},\\bm\{y\},t\)and minimize the conditional loss below\.

ℒf​\(ϕ\)\\displaystyle\\mathcal\{L\}\_\{f\}\(\\bm\{\\phi\}\)=𝔼t∼𝒰\[0,1\],\(𝒛,𝒛′\)∼q\(𝒛,𝒛′\),\(𝒙,𝒚\)∼πt\(𝒙,𝒚\|𝒛,𝒛′\)‖𝒇ϕ\(𝒙,𝒚,t\)−𝒇t\(𝒙,𝒚\|𝒛,𝒛′\)‖22\\displaystyle=\\mathbb\{E\}\_\{t\\sim\\mathcal\{U\}\[0,1\],\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\sim q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\),\(\\bm\{x\},\\bm\{y\}\)\\sim\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\}\\left\\\|\\bm\{\\bm\{f\}\_\{\\phi\}\}\(\\bm\{x\},\\bm\{y\},t\)\-\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\right\\\|\_\{2\}^\{2\}\(23\)
###### Theorem 5\.6\.

The minimizer of \([23](https://arxiv.org/html/2609.20008#S5.E23)\) is𝐟t​\(𝐱,𝐲\)\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\.

An interesting point of the decomposed velocity is that one can use it to simulate the mean\-field interaction dynamics\. ConsiderNNparticlesX0i,i=1,2⋯,NX\_\{0\}^\{i\},i=1,2\\cdots,Nindependently sampled fromρ0=μ0\\rho\_\{0\}=\\mu\_\{0\}following the ODE systemdXti=\(𝒃t\(Xti\)\+1N−1∑j≠i𝒇t\(Xti,Xtj\)\)dt,i=1,2,⋯,N\\mathrm\{d\}X\_\{t\}^\{i\}=\\Big\(\\bm\{b\}\_\{t\}\(X\_\{t\}^\{i\}\)\+\\frac\{1\}\{N\-1\}\\sum\_\{j\\neq i\}\\bm\{f\}\_\{t\}\(X\_\{t\}^\{i\},X\_\{t\}^\{j\}\)\\Big\)\\mathrm\{d\}t,i=1,2,\\cdots,N\. WhenN→∞N\\to\\infty, this dynamics approximates the mean\-field dynamics \([20](https://arxiv.org/html/2609.20008#S5.E20)\)\. Note that the velocity decomposition is not naturally unique, hence some gauges must be given\. One possible decomposition is to first train a single\-particle velocity field as𝒃\\bm\{b\}using TP\-DATE, OT\-CFM, or other methods, or to specify𝒃\\bm\{b\}directly based on domain specific prior knowledge\. When only a conditional form of𝒃\\bm\{b\}is available, additional conditions are required to make the regression well defined; we discuss this case in[E](https://arxiv.org/html/2609.20008#A5)\. Based on domain specific decomposition gauges, users may further exploit such a decomposition to provide more interpretable meanings for the components of the learned velocity\.

## 6Experiments

As a direct application of QOT, we demonstrate on both synthetic and real datasets that TP\-DATE outperforms existing methods on spatiotemporal transcriptomics slice interpolation and continuous 3D reconstruction tasks\. In this section, we useΦ=\|dd​t​‖𝒒t‖\|2\\Phi=\|\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\\|\\bm\{q\}\_\{t\}\\\|\|^\{2\}for space andΦ=0\\Phi=0for expression\. Details of this Lagrangian choice can be found in[B\.3](https://arxiv.org/html/2609.20008#A2.SS3)\. We train TP\-DATE under single\-particle version \([22](https://arxiv.org/html/2609.20008#S5.E22)\)\. We use \(F\)GW\-OT coupling for flow matching training, i\.e\. \(F\)GW\-CFM, as baselines to further demonstrate the necessity of dynamic QOT against static QOT\. See[B\.7](https://arxiv.org/html/2609.20008#A2.SS7)for details\. Ablation and scaling study of TP\-DATE are left to[C\.1](https://arxiv.org/html/2609.20008#A3.SS1),[C\.2](https://arxiv.org/html/2609.20008#A3.SS2)\.

TP\-DATE preserves spatial structure during transport\.We first use two synthetic data to verify that TP\-DATE better preserves spatial structure during the transport\. Each synthetic dataset consists of three cell types, and the ground truth spatial dynamics are defined by a90∘90^\{\\circ\}counterclockwise rigid rotation, so that the relative spatial structure of the data remains unchanged throughout the entire process\. In the Rotation data, the ground truth expression dynamics keep gene expression unchanged, whereas in the Rotation \+ expression distractor setting \(Distractor\), a type dependent expression change is applied to each cell type\. See[B\.4](https://arxiv.org/html/2609.20008#A2.SS4)for further details\.

![Refer to caption](https://arxiv.org/html/2609.20008v1/figures/toy_rotation_otcfm_tpdate_1x6.png)Figure 1:Trajectories on RotationTable 1:Hold one out experiments on synthetic data\. The best performing results are in bold\. We report mean and standard deviation over 5 random seeds\.We conducted the standard hold one out experiment commonly used in the field to evaluate the dynamics reconstruction performance of different methods\. That is, we trained the models on the initial and final time points and compared the model predictions at the intermediate unseen time point with the ground truth\. As shown in Figure[1](https://arxiv.org/html/2609.20008#S6.F1), on Rotation data, OT\-CFM follows nearly straight paths to minimize the kinetic energy and shrinks the global structure, whereas TP\-DATE recovers the more faithful curved rotation while better preserving the spatial structure\.

Quantitatively, we use the Spatial MSE and pairwise distortion to assess how well the model recovers the intermediate spatial structure, and the fused Wasserstein distance to evaluate whether the model can recover the joint pattern of gene expression and spatial coordinates\. See[B\.6](https://arxiv.org/html/2609.20008#A2.SS6)for further details\. As shown in Table[1](https://arxiv.org/html/2609.20008#S6.T1), TP\-DATE gives the lowest spatial and fused errors, and low pair distortions on both datasets\. These results show that TP\-DATE can better preserve the spatial structure during transport\. For these rotation datasets, to maintain the task non\-trivial, we implemented stVCR\([Peng et al\., 2026a](https://arxiv.org/html/2609.20008#bib.bib12)\)without performing rigid body transformation\.

TP\-DATE improves spatiotemporal dynamics reconstruction\.We further evaluate TP\-DATE by hold one out experiments on two real spatiotemporal transcriptomics datasets, Mouse brain\([Chen et al\., 2022](https://arxiv.org/html/2609.20008#bib.bib4)\)and ARTISTA\([Wei et al\., 2022](https://arxiv.org/html/2609.20008#bib.bib3)\)\. We use the Wasserstein distances in spatial coordinate and expression space, respectively, to evaluate whether the model prediction can accurately recover the spatial structure and gene expression pattern of the intermediate slice\. To jointly account for spatial and expression information, we use the fused Wasserstein distance introduced above\. In addition, we introduce the spatially coupled expression MSE \(SC\-eMSE\) to evaluate the recovery of spatial gene expression patterns in the intermediate slice\. Specifically, SC\-eMSE first computes a spatial optimal transport coupling between the predicted and ground truth slice using squared distance in spatial coordinate space, and then evaluates the transport cost under this coupling using squared distance in expression space\. Intuitively, this metric measures how different the gene expression profiles are between spatially corresponding locations in the two slices\. See[B\.5](https://arxiv.org/html/2609.20008#A2.SS5)for dataset details, and[B\.6](https://arxiv.org/html/2609.20008#A2.SS6)for metric details\.

Table 2:Hold one out experiments on spatiotemporal transcriptomics data\. The best performing results are in bold\. We report mean and standard deviation over 5 random seeds\.On ARTISTA, although TP\-DATE is less effective than the current state\-of\-the\-art method stVCR\([Peng et al\., 2026a](https://arxiv.org/html/2609.20008#bib.bib12)\)in recovering the overall spatial shape of the tissue and performs comparably to the recently proposed ContextFlow\([Rathod et al\., 2026](https://arxiv.org/html/2609.20008#bib.bib2)\), it outperforms existing baselines on the other evaluation metrics and on the Mouse Brain dataset \(Table[2](https://arxiv.org/html/2609.20008#S6.T2)\)\. We emphasize that spatiotemporal interpolation requires not only recovering the overall spatial morphology of the tissue, but also reconstructing the spatial patterns of gene expression\. TP\-DATE performs better on the latter aspect\. Therefore, these results suggest that TP\-DATE improves spatiotemporal dynamics reconstruction\.

TP\-DATE allows continuous 3D reconstruction\.Recently,[Zhang et al\. \(2026a\)](https://arxiv.org/html/2609.20008#bib.bib5)released a spatial transcriptomics dataset containing slices collected from the same tumor tissue at the same time point but at different depths\. The slices are ordered along the depth axis rather than time\. Since high\-throughput volumetric spatial transcriptomics remains experimentally challenging and is not routinely available for many tissues and platforms, continuous reconstruction of 3D spatial structure from serial tissue sections remains an important computational task\. If we assume that the spatial structure and gene expression patterns of the same tissue vary smoothly across adjacent depths, the QOT framework can likewise be used to interpolate unseen depths, thereby reconstructing the three dimensional tissue structure from a small number of sections sampled at different depths\. TP\-DATE achieves performance comparable to the best baseline in terms of spatial Wasserstein distance, while outperforming all baselines on the other evaluation metrics \(Table[3](https://arxiv.org/html/2609.20008#S6.T3)\), demonstrating its superior performance on this new and important computational task\.

Table 3:Hold one out experiments on 3D reconstruction\. The best performing results are in bold\. We report mean and standard deviation over 5 random seeds\.
## 7Conclusion and discussion

We proposed TP\-DATE, a dynamic QOT theory with a travelling pair flow matching framework for solving a family of dynamic QOT problem\. Theoretically, we formulated dynamic QOT and included GW\-OT, IGW\-OT as special cases\. Algorithmically, we developed a flow matching method that allows conditional paths to interact and used it to solve the dynamic QOT problem\. Empirically, we quantitatively demonstrated the advantages of dynamic QOT on spatiotemporal dynamics reconstruction and continuous 3D reconstruction\. We also theoretically offered a possibility to decompose the learned velocity into interpretable parts when specific domain knowledge is given\.

QOT is a relatively new concept, yet some of its special cases, such as GW\-OT, have already demonstrated substantial potential for applications\([Mémoli, 2011](https://arxiv.org/html/2609.20008#bib.bib35);[Vayer et al\., 2020](https://arxiv.org/html/2609.20008#bib.bib9);[Klein et al\., 2025](https://arxiv.org/html/2609.20008#bib.bib32)\)\. Beyond the dynamic formulation studied in this work, an important direction is to investigate whether other extensions of OT can also be introduced to the QOT setting, such as stochastic or unbalanced QOT formulations\. Future work can also try to combine the velocity decomposition paradigm with specific field knowledge for more interpretable dynamics modeling\. Another promising direction is to consider more general interaction terms, or even to learn these interactions directly from data\.

## AI Use Disclosure

In this work, we used generative AI tools to assist with translation and language polishing, and to help verify the correctness and technical details of the mathematical arguments and algorithms\. We did not use generative AI tools to develop the theoretical framework, formulate the main mathematical results, construct the proofs, design or implement the algorithms, or write the scientific content of the paper\. All AI\-assisted content was independently checked by the authors, and all final scientific and editorial decisions were made by the authors\. We take full responsibility for the final content of this work\.

## References

- Ambrosioet al\.\(2008\)L\. Ambrosio, N\. Gigli, and G\. SavaréGradient flows: in metric spaces and in the space of probability measures\.2 edition,Lectures in Mathematics\. ETH Zürich,Birkhäuser Basel\.External Links:[Document](https://dx.doi.org/10.1007/978-3-7643-8722-8),ISBN 978\-3\-7643\-8721\-1Cited by:[§F\.2](https://arxiv.org/html/2609.20008#A6.SS2.p2.1)\.
- Benamou and Brenier \(2000\)J\. Benamou and Y\. BrenierA computational fluid mechanics solution to the monge\-kantorovich mass transfer problem\.Numerische Mathematik84,pp\. 375–393\.External Links:[Document](https://dx.doi.org/10.1007/s002110050002)Cited by:[§F\.1](https://arxiv.org/html/2609.20008#A6.SS1.p1.2),[§F\.1](https://arxiv.org/html/2609.20008#A6.SS1.p1.3),[§1](https://arxiv.org/html/2609.20008#S1.p1.1),[§2](https://arxiv.org/html/2609.20008#S2.p1.1),[§3](https://arxiv.org/html/2609.20008#S3.p3.2),[§4\.1](https://arxiv.org/html/2609.20008#S4.SS1.p4.2)\.
- Chenet al\.\(2022\)A\. Chen, S\. Liao, M\. Cheng,et al\.Spatiotemporal transcriptomic atlas of mouse organogenesis using dna nanoball\-patterned arrays\.Cell185\(10\),pp\. 1777–1792\.External Links:ISSN 0092\-8674,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.cell.2022.04.003),[Link](https://www.sciencedirect.com/science/article/pii/S0092867422003993)Cited by:[§B\.5\.1](https://arxiv.org/html/2609.20008#A2.SS5.SSS1.p1.1),[§1](https://arxiv.org/html/2609.20008#S1.p3.1),[§6](https://arxiv.org/html/2609.20008#S6.p5.1)\.
- Chenet al\.\(2018\)R\. T\. Q\. Chen, Y\. Rubanova, J\. Bettencourt, and D\. K\. DuvenaudNeural ordinary differential equations\.InAdvances in Neural Information Processing Systems,S\. Bengio, H\. Wallach, H\. Larochelle, K\. Grauman, N\. Cesa\-Bianchi, and R\. Garnett \(Eds\.\),Vol\.31,pp\.\.Cited by:[§2](https://arxiv.org/html/2609.20008#S2.p3.1)\.
- Chizatet al\.\(2018a\)L\. Chizat, G\. Peyré, B\. Schmitzer, and F\. VialardAn interpolating distance between optimal transport and fisher–rao metrics\.Foundations of Computational Mathematics18\(1\),pp\. 1–44\.Cited by:[§1](https://arxiv.org/html/2609.20008#S1.p1.1),[§2](https://arxiv.org/html/2609.20008#S2.p1.1),[§2](https://arxiv.org/html/2609.20008#S2.p2.1),[§3](https://arxiv.org/html/2609.20008#S3.p4.2)\.
- Chizatet al\.\(2018b\)L\. Chizat, G\. Peyré, B\. Schmitzer, and F\. VialardUnbalanced optimal transport: dynamic and kantorovich formulations\.Journal of Functional Analysis274\(11\),pp\. 3090–3123\.Cited by:[§1](https://arxiv.org/html/2609.20008#S1.p1.1),[§3](https://arxiv.org/html/2609.20008#S3.p4.2)\.
- De Bortoliet al\.\(2021\)V\. De Bortoli, J\. Thornton, J\. Heng, and A\. DoucetDiffusion schrödinger bridge with applications to score\-based generative modeling\.InAdvances in Neural Information Processing Systems,M\. Ranzato, A\. Beygelzimer, Y\. Dauphin, P\.S\. Liang, and J\. W\. Vaughan \(Eds\.\),Vol\.34,pp\. 17695–17709\.Cited by:[§D\.8](https://arxiv.org/html/2609.20008#A4.SS8.p3.1)\.
- Elfwinget al\.\(2018\)S\. Elfwing, E\. Uchibe, and K\. DoyaSigmoid\-weighted linear units for neural network function approximation in reinforcement learning\.Neural Networks107,pp\. 3–11\.Note:Special issue on deep reinforcement learningExternal Links:ISSN 0893\-6080Cited by:[§B\.2](https://arxiv.org/html/2609.20008#A2.SS2.p1.1)\.
- Eyringet al\.\(2024\)L\. Eyring, D\. Klein, T\. Uscidda, G\. Palla, N\. Kilbertus, Z\. Akata, and F\. J\. TheisUnbalancedness in neural monge maps improves unpaired domain translation\.InThe Twelfth International Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2609.20008#S2.p2.1)\.
- Flamaryet al\.\(2021\)R\. Flamary, N\. Courty, A\. Gramfort, M\. Z\. Alaya, A\. Boisbunon, S\. Chambon, L\. Chapel, A\. Corenflos, K\. Fatras, N\. Fournier,et al\.Pot: python optimal transport\.Journal of Machine Learning Research22\(78\),pp\. 1–8\.Cited by:[§B\.8\.1](https://arxiv.org/html/2609.20008#A2.SS8.SSS1.p1.3),[§4\.1](https://arxiv.org/html/2609.20008#S4.SS1.p2.3)\.
- Halmoset al\.\(2025\)P\. Halmos, X\. Liu, J\. Gold, F\. Chen, L\. Ding, and B\. J\. RaphaelDeST\-ot: alignment of spatiotemporal transcriptomics data\.Cell Systems16\(2\),pp\. 101160\.External Links:[Document](https://dx.doi.org/10.1016/j.cels.2024.12.001)Cited by:[§2](https://arxiv.org/html/2609.20008#S2.p3.1)\.
- Kantorovich \(1958\)L\. V\. KantorovichOn the translocation of masses\.Management Science5\(1\),pp\. 1–4\.External Links:ISSN 00251909, 15265501Cited by:[§1](https://arxiv.org/html/2609.20008#S1.p1.1),[§2](https://arxiv.org/html/2609.20008#S2.p1.1),[§3](https://arxiv.org/html/2609.20008#S3.p1.3)\.
- Kerdoncuffet al\.\(2021\)T\. Kerdoncuff, R\. Emonet, and M\. SebbanSampled gromov wasserstein\.Machine Learning110,pp\. 2151–2186\.External Links:[Document](https://dx.doi.org/10.1007/s10994-021-06035-1)Cited by:[§B\.8\.1](https://arxiv.org/html/2609.20008#A2.SS8.SSS1.p1.3),[§4\.1](https://arxiv.org/html/2609.20008#S4.SS1.p2.3)\.
- Kingma and Ba \(2017\)D\. P\. Kingma and J\. BaAdam: a method for stochastic optimization\.External Links:1412\.6980,[Link](https://arxiv.org/abs/1412.6980)Cited by:[§B\.2](https://arxiv.org/html/2609.20008#A2.SS2.p1.1)\.
- Kleinet al\.\(2025\)D\. Klein, G\. Palla, M\. Lange,et al\.Mapping cells through time and space with moscot\.Nature638,pp\. 1065–1075\.External Links:[Document](https://dx.doi.org/10.1038/s41586-024-08453-2)Cited by:[§D\.1\.2](https://arxiv.org/html/2609.20008#A4.SS1.SSS2.p1.4),[§D\.1\.3](https://arxiv.org/html/2609.20008#A4.SS1.SSS3.p1.1),[§1](https://arxiv.org/html/2609.20008#S1.p2.1),[§2](https://arxiv.org/html/2609.20008#S2.p1.1),[§2](https://arxiv.org/html/2609.20008#S2.p3.1),[§3](https://arxiv.org/html/2609.20008#S3.p2.3),[§4\.2](https://arxiv.org/html/2609.20008#S4.SS2.p4.1),[§7](https://arxiv.org/html/2609.20008#S7.p2.1)\.
- Kleinet al\.\(2024\)D\. Klein, T\. Uscidda, F\. J\. Theis, and M\. cuturiGENOT: entropic \(gromov\) wasserstein flow matching with applications to single\-cell genomics\.InThe Thirty\-eighth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=hjspWd7jvg)Cited by:[§D\.4](https://arxiv.org/html/2609.20008#A4.SS4.p1.1),[§2](https://arxiv.org/html/2609.20008#S2.p2.1),[§3](https://arxiv.org/html/2609.20008#S3.p3.3)\.
- Léonard \(2014\)C\. LéonardA survey of the schrödinger problem and some of its connections with optimal transport\.Discrete and Continuous Dynamical Systems34\(4\),pp\. 1533–1574\.External Links:ISSN 1078\-0947,[Document](https://dx.doi.org/10.3934/dcds.2014.34.1533)Cited by:[§1](https://arxiv.org/html/2609.20008#S1.p1.1),[§2](https://arxiv.org/html/2609.20008#S2.p1.1)\.
- Lieroet al\.\(2018\)M\. Liero, A\. Mielke, and G\. SavaréOptimal entropy\-transport problems and a new hellinger–kantorovich distance between positive measures\.Inventiones mathematicae211,pp\. 969–1117\.Cited by:[§1](https://arxiv.org/html/2609.20008#S1.p1.1),[§2](https://arxiv.org/html/2609.20008#S2.p1.1)\.
- Lipmanet al\.\(2023\)Y\. Lipman, R\. T\. Q\. Chen, H\. Ben\-Hamu, M\. Nickel, and M\. LeFlow matching for generative modeling\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=PqvMRDCJT9t)Cited by:[§A\.5](https://arxiv.org/html/2609.20008#A1.SS5.p2.1.1),[§A\.8](https://arxiv.org/html/2609.20008#A1.SS8.p2.2.1),[§2](https://arxiv.org/html/2609.20008#S2.p2.1),[§3](https://arxiv.org/html/2609.20008#S3.p5.3),[§5\.2](https://arxiv.org/html/2609.20008#S5.SS2.p1.1)\.
- Lisini \(2007\)S\. LisiniCharacterization of absolutely continuous curves in wasserstein spaces\.Calculus of Variations and Partial Differential Equations28,pp\. 85–120\.External Links:[Document](https://dx.doi.org/10.1007/s00526-006-0032-2)Cited by:[§F\.2](https://arxiv.org/html/2609.20008#A6.SS2.p2.1)\.
- Mémoli \(2011\)F\. MémoliGromov–wasserstein distances and the metric approach to object matching\.Foundations of Computational Mathematics11,pp\. 417–487\.External Links:[Document](https://dx.doi.org/10.1007/s10208-011-9093-5)Cited by:[§D\.1\.2](https://arxiv.org/html/2609.20008#A4.SS1.SSS2.p1.4),[§1](https://arxiv.org/html/2609.20008#S1.p2.1),[§2](https://arxiv.org/html/2609.20008#S2.p1.1),[§3](https://arxiv.org/html/2609.20008#S3.p2.3),[§7](https://arxiv.org/html/2609.20008#S7.p2.1)\.
- Penget al\.\(2026a\)Q\. Peng, P\. Zhou, and T\. LiStVCR: spatiotemporal dynamics of single cells\.Nature Methods23,pp\. 542–553\.External Links:[Document](https://dx.doi.org/10.1038/s41592-026-03010-3)Cited by:[§C\.2](https://arxiv.org/html/2609.20008#A3.SS2.p1.1),[§D\.7](https://arxiv.org/html/2609.20008#A4.SS7.p1.1),[§D\.8](https://arxiv.org/html/2609.20008#A4.SS8.p3.1),[§2](https://arxiv.org/html/2609.20008#S2.p3.1),[Table 1](https://arxiv.org/html/2609.20008#S6.T1.2.1.5.1),[§6](https://arxiv.org/html/2609.20008#S6.p4.1),[§6](https://arxiv.org/html/2609.20008#S6.p6.1)\.
- Penget al\.\(2026b\)Q\. Peng, Z\. Wang, J\. Ying, Y\. Sun, Q\. Nie, L\. Zhang, T\. Li, and P\. ZhouWFR\-FM: simulation\-free dynamic unbalanced optimal transport\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=1nqu7bK1mm)Cited by:[§D\.7](https://arxiv.org/html/2609.20008#A4.SS7.p1.2),[§D\.8](https://arxiv.org/html/2609.20008#A4.SS8.p3.1),[§1](https://arxiv.org/html/2609.20008#S1.p1.1),[§2](https://arxiv.org/html/2609.20008#S2.p2.1),[§3](https://arxiv.org/html/2609.20008#S3.p4.2)\.
- Rathodet al\.\(2026\)S\. S\. Rathod, F\. Ceccarelli, S\. B\. Holden, P\. Liò, X\. Zhang, and J\. TanevskiContextFlow: context\-aware flow matching for trajectory inference from spatial omics data\.External Links:2510\.02952,[Link](https://arxiv.org/abs/2510.02952)Cited by:[§D\.6](https://arxiv.org/html/2609.20008#A4.SS6.p1.1),[§2](https://arxiv.org/html/2609.20008#S2.p2.1),[§2](https://arxiv.org/html/2609.20008#S2.p3.1),[Table 1](https://arxiv.org/html/2609.20008#S6.T1.2.1.7.1),[§6](https://arxiv.org/html/2609.20008#S6.p6.1)\.
- Sakalyanet al\.\(2025\)K\. Sakalyan, A\. Palma, F\. Guerranti, F\. J\. Theis, and S\. GünnemannModeling microenvironment trajectories on spatial transcriptomics with nicheflow\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=5ofJyjgrth)Cited by:[§D\.5](https://arxiv.org/html/2609.20008#A4.SS5.p1.1)\.
- Schiebingeret al\.\(2019\)G\. Schiebinger, J\. Shu, M\. Tabaka, B\. Cleary, V\. Subramanian, A\. Solomon, J\. Gould, S\. Liu, S\. Lin, P\. Berube,et al\.Optimal\-transport analysis of single\-cell gene expression identifies developmental trajectories in reprogramming\.Cell176,pp\. 928–943\.External Links:[Document](https://dx.doi.org/10.1016/j.cell.2019.01.006)Cited by:[§1](https://arxiv.org/html/2609.20008#S1.p1.1)\.
- Schrödinger \(1932\)E\. SchrödingerSur la théorie relativiste de l’électron et l’interprétation de la mécanique quantique\.Annales de l’Institut Henri Poincaré2\(4\),pp\. 269–310\.Cited by:[§1](https://arxiv.org/html/2609.20008#S1.p1.1),[§2](https://arxiv.org/html/2609.20008#S2.p1.1)\.
- Shaet al\.\(2024\)Y\. Sha, Y\. Qiu, P\. Zhou, and Q\. NieReconstructing growth and dynamic trajectories from single\-cell transcriptomics data\.Nature Machine Intelligence6\(1\),pp\. 25–39\.Cited by:[§1](https://arxiv.org/html/2609.20008#S1.p1.1)\.
- Shiet al\.\(2023\)Y\. Shi, V\. De Bortoli, A\. Campbell, and A\. DoucetDiffusion schrödinger bridge matching\.InAdvances in Neural Information Processing Systems,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),Vol\.36,pp\. 62183–62223\.Cited by:[§D\.8](https://arxiv.org/html/2609.20008#A4.SS8.p3.1)\.
- Tonget al\.\(2024a\)A\. Tong, K\. Fatras, N\. Malkin, G\. Huguet, Y\. Zhang, J\. Rector\-Brooks, G\. Wolf, and Y\. BengioImproving and generalizing flow\-based generative models with minibatch optimal transport\.Transactions on Machine Learning Research,pp\. 1–34\.External Links:ISSN 2835\-8856Cited by:[§A\.5](https://arxiv.org/html/2609.20008#A1.SS5.p2.1.1),[§A\.8](https://arxiv.org/html/2609.20008#A1.SS8.p2.2.1),[§D\.3](https://arxiv.org/html/2609.20008#A4.SS3.p1.2),[§D\.8](https://arxiv.org/html/2609.20008#A4.SS8.p3.1),[§1](https://arxiv.org/html/2609.20008#S1.p1.1),[§2](https://arxiv.org/html/2609.20008#S2.p2.1),[§3](https://arxiv.org/html/2609.20008#S3.p3.3),[§3](https://arxiv.org/html/2609.20008#S3.p4.2),[§3](https://arxiv.org/html/2609.20008#S3.p6.1),[§5](https://arxiv.org/html/2609.20008#S5.p1.1),[Table 1](https://arxiv.org/html/2609.20008#S6.T1.2.1.2.2)\.
- Tonget al\.\(2020\)A\. Tong, J\. Huang, G\. Wolf, D\. Van Dijk, and S\. KrishnaswamyTrajectoryNet: a dynamic optimal transport network for modeling cellular dynamics\.InProceedings of the 37th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.119,pp\. 9526–9536\.Cited by:[§D\.8](https://arxiv.org/html/2609.20008#A4.SS8.p3.1),[§1](https://arxiv.org/html/2609.20008#S1.p1.1),[§3](https://arxiv.org/html/2609.20008#S3.p3.3)\.
- Tonget al\.\(2024b\)A\. Y\. Tong, N\. Malkin, K\. Fatras, L\. Atanackovic, Y\. Zhang, G\. Huguet, G\. Wolf, and Y\. BengioSimulation\-free Schrödinger bridges via score and flow matching\.InProceedings of The 27th International Conference on Artificial Intelligence and Statistics,Proceedings of Machine Learning Research, Vol\.238,pp\. 1279–1287\.Cited by:[§D\.8](https://arxiv.org/html/2609.20008#A4.SS8.p3.1),[§2](https://arxiv.org/html/2609.20008#S2.p2.1)\.
- Traaget al\.\(2019\)V\. A\. Traag, L\. Waltman, and N\. J\. van EckFrom louvain to leiden: guaranteeing well\-connected communities\.Scientific Reports9,pp\. 5233\.External Links:[Document](https://dx.doi.org/10.1038/s41598-019-41695-z)Cited by:[§B\.5\.1](https://arxiv.org/html/2609.20008#A2.SS5.SSS1.p1.1)\.
- Vayeret al\.\(2020\)T\. Vayer, L\. Chapel, R\. Flamary, R\. Tavenard, and N\. CourtyFused gromov\-wasserstein distance for structured objects\.Algorithms13\(9\)\.External Links:[Link](https://www.mdpi.com/1999-4893/13/9/212),ISSN 1999\-4893,[Document](https://dx.doi.org/10.3390/a13090212)Cited by:[§D\.1\.3](https://arxiv.org/html/2609.20008#A4.SS1.SSS3.p1.1),[§1](https://arxiv.org/html/2609.20008#S1.p2.1),[§2](https://arxiv.org/html/2609.20008#S2.p1.1),[§4\.2](https://arxiv.org/html/2609.20008#S4.SS2.p4.1),[§7](https://arxiv.org/html/2609.20008#S7.p2.1)\.
- Wang and Zhang \(2025\)R\. Wang and Z\. ZhangQuadratic\-form optimal transport\.Mathematical Programming\.External Links:[Document](https://dx.doi.org/10.1007/s10107-025-02282-5)Cited by:[§1](https://arxiv.org/html/2609.20008#S1.p4.1),[§2](https://arxiv.org/html/2609.20008#S2.p1.1),[§3](https://arxiv.org/html/2609.20008#S3.p2.2)\.
- Weiet al\.\(2022\)X\. Wei, S\. Fu, H\. Li,et al\.Single\-cell stereo\-seq reveals induced progenitor cells involved in axolotl brain regeneration\.Science377\(6610\),pp\. eabp9444\.External Links:[Document](https://dx.doi.org/10.1126/science.abp9444),[Link](https://www.science.org/doi/abs/10.1126/science.abp9444)Cited by:[§B\.5\.2](https://arxiv.org/html/2609.20008#A2.SS5.SSS2.p1.1),[§1](https://arxiv.org/html/2609.20008#S1.p3.1),[§6](https://arxiv.org/html/2609.20008#S6.p5.1)\.
- Yinget al\.\(2026\)J\. Ying, Y\. Wang, B\. Yang, P\. Zhou, and L\. ZhangBeyond continuity: simulation\-free reconstruction of discrete branching dynamics from single\-cell snapshots\.InForty\-third International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=IABjbJlinz)Cited by:[§2](https://arxiv.org/html/2609.20008#S2.p2.1)\.
- Zhanget al\.\(2026a\)H\. Zhang, Z\. Zhang, P\. Wang, T\. Xu, X\. Chen, Y\. Zhao, S\. Lin, W\. Cai, P\. Ren, C\. Luo, P\. Zhang, Y\. Wang, S\. Hou, Y\. Zhao, H\. Zeng, Z\. Liu, C\. Wang, Z\. Gao, Y\. Feng, D\. Pan, and Z\. ZengUncovering spatially resolved functional genomics with crispr screen sequencing\.Cell189,pp\. 4594–4618\.External Links:[Document](https://dx.doi.org/10.1016/j.cell.2026.04.049)Cited by:[§B\.5\.3](https://arxiv.org/html/2609.20008#A2.SS5.SSS3.p1.1),[§C\.2](https://arxiv.org/html/2609.20008#A3.SS2.p1.1),[§1](https://arxiv.org/html/2609.20008#S1.p3.1),[§6](https://arxiv.org/html/2609.20008#S6.p7.1)\.
- Zhanget al\.\(2026b\)Z\. Zhang, Z\. Goldfeld, K\. Greenewald,et al\.Gradient flows and riemannian structure in the gromov\-wasserstein geometry\.Foundations of Computational Mathematics26,pp\. 1911–2003\.External Links:[Document](https://dx.doi.org/10.1007/s10208-025-09722-w)Cited by:[§D\.2](https://arxiv.org/html/2609.20008#A4.SS2.p1.1),[§1](https://arxiv.org/html/2609.20008#S1.p4.1),[§2](https://arxiv.org/html/2609.20008#S2.p1.1),[§4\.2](https://arxiv.org/html/2609.20008#S4.SS2.p4.1)\.
- Zhanget al\.\(2025a\)Z\. Zhang, T\. Li, and P\. ZhouLearning stochastic dynamics from snapshots through regularized unbalanced optimal transport\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=gQlxd3Mtru)Cited by:[§1](https://arxiv.org/html/2609.20008#S1.p1.1)\.
- Zhanget al\.\(2025b\)Z\. Zhang, Z\. Wang, Y\. Sun, T\. Li, and P\. ZhouModeling cell dynamics and interactions with unbalanced mean field schrödinger bridge\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=Z6DJJIN8IJ)Cited by:[§C\.2](https://arxiv.org/html/2609.20008#A3.SS2.p1.1),[§D\.8](https://arxiv.org/html/2609.20008#A4.SS8.p1.1),[§2](https://arxiv.org/html/2609.20008#S2.p3.1),[Table 1](https://arxiv.org/html/2609.20008#S6.T1.2.1.6.1)\.

## Appendix AProofs

Proofs of theorems and propositions\.

### A\.1Proof of Theorem[4\.1](https://arxiv.org/html/2609.20008#S4.Thmtheorem1)

Theorem[4\.1](https://arxiv.org/html/2609.20008#S4.Thmtheorem1)\.Ifℒ\\mathcal\{L\}is convex w\.r\.t\(𝐱˙,𝐲˙\)\(\\dot\{\\bm\{x\}\},\\dot\{\\bm\{y\}\}\), thenQOTS​\(μ0,μ1\)=QOTD​\(μ0,μ1\)≥QOTB​B​\(μ0,μ1\)\\text\{QOT\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\text\{QOT\}\_\{D\}\(\\mu\_\{0\},\\mu\_\{1\}\)\\geq\\text\{QOT\}\_\{BB\}\(\\mu\_\{0\},\\mu\_\{1\}\)\.

###### Proof\.

For convenience, we restate the three forms below\. The static form is

QOTS​\(μ0,μ1\)=infγ∈Π⁡\(μ0,μ1\)∫𝒳4𝒜⁡\(𝒛,𝒛′\)​γ​\(𝒛\)​γ​\(𝒛′\)​d𝒛​d​𝒛′\.\\displaystyle\\text\{QOT\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\inf\_\{\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)\}\\int\_\{\\mathcal\{X\}^\{4\}\}\\mathcal\{A\}\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\gamma\(\\bm\{z\}\)\\gamma\(\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}\.\(24\)The dynamic form is

QOTD​\(μ0,μ1\)=\\displaystyle\\text\{QOT\}\_\{D\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\(25\)inf∫01∫∫𝒳2ℒ\(t,𝒙,𝒚,u1t\(𝒙,𝒚\|𝒛,𝒛′\),u2t\(𝒙,𝒚\|𝒛,𝒛′\)\)πt\(𝒙,𝒚\|𝒛,𝒛′\)γ\(𝒛\)γ\(𝒛′\)d𝒙d𝒚d𝒛d𝒛′dt\.\\displaystyle\\inf\\int\_\{0\}^\{1\}\\int\\int\_\{\\mathcal\{X\}^\{2\}\}\\mathcal\{L\}\(t,\\bm\{x\},\\bm\{y\},u^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\),u^\{2\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\)\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\gamma\(\\bm\{z\}\)\\gamma\(\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}\\mathrm\{d\}t\.The BB\-form is

QOTB​B​\(μ0,μ1\)=infπ,𝒖1,𝒖2∫01∫𝒳2ℒ⁡\(t,𝒙,𝒚,ut1​\(𝒙,𝒚\),ut2​\(𝒙,𝒚\)\)​πt​\(𝒙,𝒚\)​𝑑𝒙​𝑑𝒚​𝑑t\\displaystyle\\text\{QOT\}\_\{BB\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\inf\_\{\\pi,\\bm\{u\}^\{1\},\\bm\{u\}^\{2\}\}\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}^\{2\}\}\\mathcal\{L\}\(t,\\bm\{x\},\\bm\{y\},u^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\),u^\{2\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\)\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}t\(26\)∂tπt​\(𝒙,𝒚\)\+∇𝒙⋅\(πt​\(𝒙,𝒚\)​ut1​\(𝒙,𝒚\)\)\+∇𝒚⋅\(πt​\(𝒙,𝒚\)​ut2​\(𝒙,𝒚\)\)=0\.\\displaystyle\\partial\_\{t\}\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\+\\nabla\_\{\\bm\{x\}\}\\cdot\(\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)u^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\)\+\\nabla\_\{\\bm\{y\}\}\\cdot\(\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)u^\{2\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\)=0\.
For any\(𝒛,𝒛′\)\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\), the optimal path action𝒜⁡\(𝒛,𝒛′\)\\mathcal\{A\}\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)is attained by the travelling pair𝒙˙t=𝒖t1\(𝒙t,𝒚t\|𝒛,𝒛′\),𝒚˙t=𝒖t2\(𝒙t,𝒚t\|𝒛,𝒛′\)\\dot\{\\bm\{x\}\}\_\{t\}=\\bm\{u\}\_\{t\}^\{1\}\(\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\),\\dot\{\\bm\{y\}\}\_\{t\}=\\bm\{u\}\_\{t\}^\{2\}\(\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\.

𝒜⁡\(𝒛,𝒛′\)\\displaystyle\\mathcal\{A\}\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)=inf𝒙t,𝒚t∫01ℒ⁡\(t,𝒙t,𝒚t,𝒙˙t,𝒚˙t\)​𝑑t\\displaystyle=\\inf\_\{\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\}\}\\int\_\{0\}^\{1\}\\mathcal\{L\}\(t,\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\},\\dot\{\\bm\{x\}\}\_\{t\},\\dot\{\\bm\{y\}\}\_\{t\}\)\\mathrm\{d\}t\(27\)=∫01ℒ\(t,𝒙t,𝒚t,𝒖t1\(𝒙t,𝒚t\|𝒛,𝒛′\),𝒖t2\(𝒙t,𝒚t\|𝒛,𝒛′\)\)dt\\displaystyle=\\int\_\{0\}^\{1\}\\mathcal\{L\}\(t,\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\},\\bm\{u\}\_\{t\}^\{1\}\(\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\),\\bm\{u\}\_\{t\}^\{2\}\(\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\)\\mathrm\{d\}t=∫01∫𝒳2ℒ\(t,𝒙,𝒚,𝒖t1\(𝒙,𝒚\|𝒛,𝒛′\),𝒖t2\(𝒙,𝒚\|𝒛,𝒛′\)\)πt\(𝒙,𝒚\|𝒛,𝒛′\)d𝒙d𝒚dt\\displaystyle=\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}^\{2\}\}\\mathcal\{L\}\(t,\\bm\{x\},\\bm\{y\},\\bm\{u\}\_\{t\}^\{1\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\),\\bm\{u\}\_\{t\}^\{2\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\)\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}twhereπt\(𝒙,𝒚\|𝒛,𝒛′\)\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)is the conditional Dirac probability path generated by the travelling pair velocity, which is

∂tπt\(𝒙,𝒚\|𝒛,𝒛′\)\+∇𝒙⋅\(πt\(𝒙,𝒚\|𝒛,𝒛′\)𝒖t1\(𝒙t,𝒚t\|𝒛,𝒛′\)\)\+∇𝒚⋅\(πt\(𝒙,𝒚\|𝒛,𝒛′\)𝒖t2\(𝒙t,𝒚t\|𝒛,𝒛′\)\)=0\\displaystyle\\partial\_\{t\}\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\+\\nabla\_\{\\bm\{x\}\}\\cdot\(\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\bm\{u\}\_\{t\}^\{1\}\(\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\)\+\\nabla\_\{\\bm\{y\}\}\\cdot\(\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\bm\{u\}\_\{t\}^\{2\}\(\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\)=0\(28\)π0\(𝒙,𝒚\|𝒛,𝒛′\)=δ\(𝒙0,𝒚0\)\(𝒙,𝒚\),π1\(𝒙,𝒚\|𝒛,𝒛′\)=δ\(𝒙1,𝒚1\)\(𝒙,𝒚\)\\displaystyle\\pi\_\{0\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)=\\delta\_\{\(\\bm\{x\}\_\{0\},\\bm\{y\}\_\{0\}\)\}\(\\bm\{x\},\\bm\{y\}\),\\quad\\pi\_\{1\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)=\\delta\_\{\(\\bm\{x\}\_\{1\},\\bm\{y\}\_\{1\}\)\}\(\\bm\{x\},\\bm\{y\}\)Therefore,QOTD​\(μ0,μ1\)=QOTS​\(μ0,μ1\)\\text\{QOT\}\_\{D\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\text\{QOT\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)\.

For any admissible couplingγ\\gamma,

∫𝒳4𝒜⁡\(𝒛,𝒛′\)​γ​\(𝒛\)​γ​\(𝒛′\)​𝑑𝒛​d​𝒛′\\displaystyle\\int\_\{\\mathcal\{X\}^\{4\}\}\\mathcal\{A\}\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\gamma\(\\bm\{z\}\)\\gamma\(\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}\(29\)=∫ℒ\(t,𝒙,𝒚,𝒖t1\(𝒙,𝒚\|𝒛,𝒛′\),𝒖t2\(𝒙,𝒚\|𝒛,𝒛′\)\)πt\(𝒙,𝒚\|𝒛,𝒛′\)γ\(𝒛\)γ\(𝒛′\)d𝒙d𝒚dtd𝒛d𝒛′\\displaystyle=\\int\\mathcal\{L\}\(t,\\bm\{x\},\\bm\{y\},\\bm\{u\}\_\{t\}^\{1\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\),\\bm\{u\}\_\{t\}^\{2\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\)\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\gamma\(\\bm\{z\}\)\\gamma\(\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}t\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}=∫\(∫ℒ\(t,𝒙,𝒚,𝒖t1\(𝒙,𝒚\|𝒛,𝒛′\),𝒖t2\(𝒙,𝒚\|𝒛,𝒛′\)\)πt\(𝒙,𝒚\|𝒛,𝒛′\)γ\(𝒛\)γ\(𝒛′\)πt​\(𝒙,𝒚\)d𝒛d𝒛′\)πt\(𝒙,𝒚\)d𝒙d𝒚dt\\displaystyle=\\int\\Big\(\\int\\mathcal\{L\}\(t,\\bm\{x\},\\bm\{y\},\\bm\{u\}\_\{t\}^\{1\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\),\\bm\{u\}\_\{t\}^\{2\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\)\\frac\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\gamma\(\\bm\{z\}\)\\gamma\(\\bm\{z\}^\{\\prime\}\)\}\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\}\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}\\Big\)\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}t≥∫ℒ\(t,𝒙,𝒚,𝒖t1\(𝒙,𝒚\),𝒖t2\(𝒙,𝒚\)\)πt\(𝒙,𝒚\)d𝒙d𝒚dt\(Jensen’s inequality\)\\displaystyle\\geq\\int\\mathcal\{L\}\(t,\\bm\{x\},\\bm\{y\},\\bm\{u\}\_\{t\}^\{1\}\(\\bm\{x\},\\bm\{y\}\),\\bm\{u\}\_\{t\}^\{2\}\(\\bm\{x\},\\bm\{y\}\)\)\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}t\\qquad\(\\text\{Jensen's inequality\}\)≥QOTB​B\(μ0,μ1\)\(Take infimum\)\\displaystyle\\geq\\text\{QOT\}\_\{BB\}\(\\mu\_\{0\},\\mu\_\{1\}\)\\qquad\(\\text\{Take infimum\}\)By taking infimum w\.r\.tγ\\gamma, we haveQOTS​\(μ0,μ1\)≥QOTB​B​\(μ0,μ1\)\\text\{QOT\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)\\geq\\text\{QOT\}\_\{BB\}\(\\mu\_\{0\},\\mu\_\{1\}\)\. ∎

### A\.2Proof of Proposition[4\.2](https://arxiv.org/html/2609.20008#S4.Thmtheorem2)

Proposition[4\.2](https://arxiv.org/html/2609.20008#S4.Thmtheorem2)\.The optimal𝐜t\\bm\{c\}\_\{t\}is𝐜t=\(1−t\)​𝐜0\+t​𝐜1\\bm\{c\}\_\{t\}=\(1\-t\)\\bm\{c\}\_\{0\}\+t\\bm\{c\}\_\{1\}\. IfΦ=Φ⁡\(‖𝐪‖\)\\Phi=\\Phi\(\\\|\\bm\{q\}\\\|\),𝐪0∦𝐪1\\bm\{q\}\_\{0\}\\nparallel\\bm\{q\}\_\{1\}, then𝐪t∈span​\{𝐪0,𝐪1\}\\bm\{q\}\_\{t\}\\in\\text\{span\}\\\{\\bm\{q\}\_\{0\},\\bm\{q\}\_\{1\}\\\}\.

###### Proof\.

Recall the Lagrangian

ℒ⁡\(t,𝒙,𝒚,𝒙˙,𝒚˙\)=‖𝒄˙t‖2\+14​‖𝒒˙t‖2\+λ​Φ​\(𝒒,𝒒˙\)\.\\mathcal\{L\}\(t,\\bm\{x\},\\bm\{y\},\\dot\{\\bm\{x\}\},\\dot\{\\bm\{y\}\}\)=\\\|\\dot\{\\bm\{c\}\}\_\{t\}\\\|^\{2\}\+\\frac\{1\}\{4\}\\\|\\dot\{\\bm\{q\}\}\_\{t\}\\\|^\{2\}\+\\lambda\\Phi\(\\bm\{q\},\\dot\{\\bm\{q\}\}\)\.\(30\)The Euler\-Lagrange equation is

\{2​𝒄¨t=𝟎\(dd​t​∂ℒ∂𝒄˙=∂ℒ∂𝒄\)12​𝒒¨t\+λ​dd​t​∂Φ∂𝒒˙=λ​∂Φ∂𝒒\(dd​t​∂ℒ∂𝒒˙=∂ℒ∂𝒒\)\\left\\\{\\begin\{aligned\} &2\\ddot\{\\bm\{c\}\}\_\{t\}=\\bm\{0\}\\qquad&\(\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\dot\{\\bm\{c\}\}\}=\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\bm\{c\}\}\)\\\\ &\\frac\{1\}\{2\}\\ddot\{\\bm\{q\}\}\_\{t\}\+\\lambda\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\frac\{\\partial\\Phi\}\{\\partial\\dot\{\\bm\{q\}\}\}=\\lambda\\frac\{\\partial\\Phi\}\{\\partial\\bm\{q\}\}\\qquad&\(\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\dot\{\\bm\{q\}\}\}=\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\bm\{q\}\}\)\\end\{aligned\}\\right\.\(31\)The first equation yields𝒄t=\(1−t\)​𝒄0\+t​𝒄1\\bm\{c\}\_\{t\}=\(1\-t\)\\bm\{c\}\_\{0\}\+t\\bm\{c\}\_\{1\}\. WhenΦ\\Phiis a radial function of𝒒\\bm\{q\}, the second equation reduces to

12​𝒒¨t=λ​∂Φ∂𝒒=λ​Φ′​\(‖𝒒t‖\)​𝒒t‖𝒒t‖\\frac\{1\}\{2\}\\ddot\{\\bm\{q\}\}\_\{t\}=\\lambda\\frac\{\\partial\\Phi\}\{\\partial\\bm\{q\}\}=\\lambda\\Phi^\{\\prime\}\(\\\|\\bm\{q\}\_\{t\}\\\|\)\\frac\{\\bm\{q\}\_\{t\}\}\{\\\|\\bm\{q\}\_\{t\}\\\|\}\(32\)That means the acceleration𝒒¨t\\ddot\{\\bm\{q\}\}\_\{t\}is parallel to𝒒t\\bm\{q\}\_\{t\}, and hence the angular momentum is conserved\. Therefore,𝒒t\\bm\{q\}\_\{t\}is constrained on a plane\. When𝒒0∦𝒒1\\bm\{q\}\_\{0\}\\nparallel\\bm\{q\}\_\{1\}, we have𝒒t∈span​\{𝒒0,𝒒1\}\\bm\{q\}\_\{t\}\\in\\text\{span\}\\\{\\bm\{q\}\_\{0\},\\bm\{q\}\_\{1\}\\\}\. ∎

### A\.3Proof of Proposition[4\.3](https://arxiv.org/html/2609.20008#S4.Thmtheorem3)

Proposition[4\.3](https://arxiv.org/html/2609.20008#S4.Thmtheorem3)\.IfΦ=\|dd​t​ϕ​\(𝐪t\)\|2\\Phi=\|\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\phi\(\\bm\{q\}\_\{t\}\)\|^\{2\},Φ\\Phiis convex w\.r\.t𝐪˙t\\dot\{\\bm\{q\}\}\_\{t\}\. IfΦ=\|dd​t​‖𝐪t‖\|2\\Phi=\|\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\\|\\bm\{q\}\_\{t\}\\\|\|^\{2\},𝐪t\\bm\{q\}\_\{t\}has analytic solution\.

###### Proof\.

By direct calculation,

Φ=\|dd​t​ϕ​\(𝒒t\)\|2=\|∇𝒒ϕ​\(𝒒t\)T​𝒒˙t\|2=𝒒˙tT​\[∇𝒒ϕ​\(𝒒t\)​∇𝒒ϕ​\(𝒒t\)T\]​𝒒˙t\\displaystyle\\Phi=\|\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\phi\(\\bm\{q\}\_\{t\}\)\|^\{2\}=\|\\nabla\_\{\\bm\{q\}\}\\phi\(\\bm\{q\}\_\{t\}\)^\{\\mathrm\{T\}\}\\dot\{\\bm\{q\}\}\_\{t\}\|^\{2\}=\\dot\{\\bm\{q\}\}\_\{t\}^\{\\mathrm\{T\}\}\[\\nabla\_\{\\bm\{q\}\}\\phi\(\\bm\{q\}\_\{t\}\)\\nabla\_\{\\bm\{q\}\}\\phi\(\\bm\{q\}\_\{t\}\)^\{\\mathrm\{T\}\}\]\\dot\{\\bm\{q\}\}\_\{t\}\(33\)Therefore, it is convex w\.r\.t𝒒˙t\\dot\{\\bm\{q\}\}\_\{t\}\. WhenΦ=\|dd​t​‖𝒒t‖\|2\\Phi=\|\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\\|\\bm\{q\}\_\{t\}\\\|\|^\{2\}, the action of\(𝒒,𝒒˙\)\(\\bm\{q\},\\dot\{\\bm\{q\}\}\)is

∫0114​‖𝒒˙t‖2\+λ​\|dd​t​‖𝒒t‖\|2​𝑑t\.\\int\_\{0\}^\{1\}\\frac\{1\}\{4\}\\\|\\dot\{\\bm\{q\}\}\_\{t\}\\\|^\{2\}\+\\lambda\|\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\\|\\bm\{q\}\_\{t\}\\\|\|^\{2\}\\mathrm\{d\}t\.\(34\)The corresponding Euler\-Lagrange equation is

12​𝒒¨t\+λ​dd​t​∂Φ∂𝒒˙=λ​∂Φ∂𝒒\.\\frac\{1\}\{2\}\\ddot\{\\bm\{q\}\}\_\{t\}\+\\lambda\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\frac\{\\partial\\Phi\}\{\\partial\\dot\{\\bm\{q\}\}\}=\\lambda\\frac\{\\partial\\Phi\}\{\\partial\\bm\{q\}\}\.\(35\)By direct calculation, we have

∂Φ∂𝒒˙=2​𝒒tT​𝒒˙t‖𝒒t‖2​𝒒t,∂Φ∂𝒒=2​𝒒tT​𝒒˙t‖𝒒t‖2​𝒒˙t−2​\(𝒒tT​𝒒˙t\)2‖𝒒t‖4​𝒒t\.\\frac\{\\partial\\Phi\}\{\\partial\\dot\{\\bm\{q\}\}\}=\\frac\{2\\bm\{q\}\_\{t\}^\{\\mathrm\{T\}\}\\dot\{\\bm\{q\}\}\_\{t\}\}\{\\\|\\bm\{q\}\_\{t\}\\\|^\{2\}\}\\bm\{q\}\_\{t\},\\qquad\\frac\{\\partial\\Phi\}\{\\partial\\bm\{q\}\}=\\frac\{2\\bm\{q\}\_\{t\}^\{\\mathrm\{T\}\}\\dot\{\\bm\{q\}\}\_\{t\}\}\{\\\|\\bm\{q\}\_\{t\}\\\|^\{2\}\}\\dot\{\\bm\{q\}\}\_\{t\}\-\\frac\{2\(\\bm\{q\}\_\{t\}^\{\\mathrm\{T\}\}\\dot\{\\bm\{q\}\}\_\{t\}\)^\{2\}\}\{\\\|\\bm\{q\}\_\{t\}\\\|^\{4\}\}\\bm\{q\}\_\{t\}\.\\\(36\)Therefore,

dd​t​∂Φ∂𝒒˙=2​\(‖𝒒˙t‖2\+𝒒tT​𝒒¨t‖𝒒t‖2−2​\(𝒒tT​𝒒˙t\)2‖𝒒t‖4\)​𝒒t\+2​𝒒tT​𝒒˙t‖𝒒t‖2​𝒒˙t\.\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\frac\{\\partial\\Phi\}\{\\partial\\dot\{\\bm\{q\}\}\}=2\\Big\(\\frac\{\\\|\\dot\{\\bm\{q\}\}\_\{t\}\\\|^\{2\}\+\\bm\{q\}\_\{t\}^\{\\mathrm\{T\}\}\\ddot\{\\bm\{q\}\}\_\{t\}\}\{\\\|\\bm\{q\}\_\{t\}\\\|^\{2\}\}\-2\\frac\{\(\\bm\{q\}\_\{t\}^\{\\mathrm\{T\}\}\\dot\{\\bm\{q\}\}\_\{t\}\)^\{2\}\}\{\\\|\\bm\{q\}\_\{t\}\\\|^\{4\}\}\\Big\)\\bm\{q\}\_\{t\}\+\\frac\{2\\bm\{q\}\_\{t\}^\{\\mathrm\{T\}\}\\dot\{\\bm\{q\}\}\_\{t\}\}\{\\\|\\bm\{q\}\_\{t\}\\\|^\{2\}\}\\dot\{\\bm\{q\}\}\_\{t\}\.\(37\)The Euler\-Lagrange equation reduces to

12​𝒒¨t\+2​λ​\(‖𝒒˙t‖2\+𝒒tT​𝒒¨t‖𝒒t‖2−\(𝒒tT​𝒒˙t\)2‖𝒒t‖4\)​𝒒t=𝟎\.\\frac\{1\}\{2\}\\ddot\{\\bm\{q\}\}\_\{t\}\+2\\lambda\\Big\(\\frac\{\\\|\\dot\{\\bm\{q\}\}\_\{t\}\\\|^\{2\}\+\\bm\{q\}\_\{t\}^\{\\mathrm\{T\}\}\\ddot\{\\bm\{q\}\}\_\{t\}\}\{\\\|\\bm\{q\}\_\{t\}\\\|^\{2\}\}\-\\frac\{\(\\bm\{q\}\_\{t\}^\{\\mathrm\{T\}\}\\dot\{\\bm\{q\}\}\_\{t\}\)^\{2\}\}\{\\\|\\bm\{q\}\_\{t\}\\\|^\{4\}\}\\Big\)\\bm\{q\}\_\{t\}=\\bm\{0\}\.\(38\)This leads to the conservation of angular momentum, therefore𝒒t\\bm\{q\}\_\{t\}is again constrained on the plane spanned by𝒒0,𝒒1\\bm\{q\}\_\{0\},\\bm\{q\}\_\{1\}\. For convenience, we definert=‖𝒒t‖,α=arccos⁡⟨𝒒0,𝒒1⟩‖𝒒0‖​‖𝒒1‖r\_\{t\}=\\\|\\bm\{q\}\_\{t\}\\\|,\\alpha=\\arccos\\frac\{\\langle\\bm\{q\}\_\{0\},\\bm\{q\}\_\{1\}\\rangle\}\{\\\|\\bm\{q\}\_\{0\}\\\|\\\|\\bm\{q\}\_\{1\}\\\|\}, and the polar coordinates on that plane\.

𝒒0=r0​\(1,0\),𝒒1=r1​\(cos⁡α,sin⁡α\),𝒒t=rt​\(cos⁡θt,sin⁡θt\)\\bm\{q\}\_\{0\}=r\_\{0\}\(1,0\),\\qquad\\bm\{q\}\_\{1\}=r\_\{1\}\(\\cos\\alpha,\\sin\\alpha\),\\qquad\\bm\{q\}\_\{t\}=r\_\{t\}\(\\cos\\theta\_\{t\},\\sin\\theta\_\{t\}\)\(39\)Under this representation,𝒒˙t=r˙t​\(cos⁡θt,sin⁡θt\)\+rt​θ˙t​\(−sin⁡θt,cos⁡θt\)\\dot\{\\bm\{q\}\}\_\{t\}=\\dot\{r\}\_\{t\}\(\\cos\\theta\_\{t\},\\sin\\theta\_\{t\}\)\+r\_\{t\}\\dot\{\\theta\}\_\{t\}\(\-\\sin\\theta\_\{t\},\\cos\\theta\_\{t\}\),\|𝒒˙t\|2=r˙t2\+rt2​θ˙t2\|\\dot\{\\bm\{q\}\}\_\{t\}\|^\{2\}=\\dot\{r\}\_\{t\}^\{2\}\+r\_\{t\}^\{2\}\\dot\{\\theta\}\_\{t\}^\{2\}\. The action \([34](https://arxiv.org/html/2609.20008#A1.E34)\) becomes

∫0114​rt2​θ˙t2\+\(14\+λ\)​r˙t2​𝑑t\.\\int\_\{0\}^\{1\}\\frac\{1\}\{4\}r\_\{t\}^\{2\}\\dot\{\\theta\}\_\{t\}^\{2\}\+\(\\frac\{1\}\{4\}\+\\lambda\)\\dot\{r\}\_\{t\}^\{2\}\\mathrm\{d\}t\.\(40\)Writea=14\+λ,k=11\+4​λ∈\(0,1\]a=\\frac\{1\}\{4\}\+\\lambda,k=\\sqrt\{\\frac\{1\}\{1\+4\\lambda\}\}\\in\(0,1\], and define𝒘t=a​rt​\(cos⁡k​θt,sin⁡k​θt\)\\bm\{w\}\_\{t\}=\\sqrt\{a\}r\_\{t\}\(\\cos k\\theta\_\{t\},\\sin k\\theta\_\{t\}\), we have

‖𝒘˙t‖2=a⁡\(r˙t2\+k2​rt2​θ˙t2\)=14​rt2​θ˙t2\+\(14\+λ\)​r˙t2\.\\\|\\dot\{\\bm\{w\}\}\_\{t\}\\\|^\{2\}=a\(\\dot\{r\}\_\{t\}^\{2\}\+k^\{2\}r\_\{t\}^\{2\}\\dot\{\\theta\}\_\{t\}^\{2\}\)=\\frac\{1\}\{4\}r\_\{t\}^\{2\}\\dot\{\\theta\}\_\{t\}^\{2\}\+\(\\frac\{1\}\{4\}\+\\lambda\)\\dot\{r\}\_\{t\}^\{2\}\.\(41\)Therefore, the minimization problem reduces to

inf𝒘t∫01‖𝒘˙t‖2​𝑑ts\.t\.𝒘0=a​r0​\(1,0\),𝒘1=a​r1​\(cos⁡k​α,sin⁡k​α\)\.\\inf\_\{\\bm\{w\}\_\{t\}\}\\int\_\{0\}^\{1\}\\\|\\dot\{\\bm\{w\}\}\_\{t\}\\\|^\{2\}\\mathrm\{d\}t\\quad\\text\{s\.t\.\}\\quad\\bm\{w\}\_\{0\}=\\sqrt\{a\}r\_\{0\}\(1,0\),\\bm\{w\}\_\{1\}=\\sqrt\{a\}r\_\{1\}\(\\cos k\\alpha,\\sin k\\alpha\)\.\(42\)Sinceα∈\(0,π\)\\alpha\\in\(0,\\pi\)andk∈\(0,1\]k\\in\(0,1\], the optimal𝒘t\\bm\{w\}\_\{t\}is the displacement interpolation

𝒘t=\(1−t\)​𝒘0\+t​𝒘1=a​\(\(1−t\)​r0\+t​r1​cos⁡k​α,t​r1​sin⁡k​α\)\.\\bm\{w\}\_\{t\}=\(1\-t\)\\bm\{w\}\_\{0\}\+t\\bm\{w\}\_\{1\}=\\sqrt\{a\}\(\(1\-t\)r\_\{0\}\+tr\_\{1\}\\cos k\\alpha,tr\_\{1\}\\sin k\\alpha\)\.\(43\)By direct calculation, we have

‖𝒘˙t‖2=a⁡\(r02\+r12−2​r0​r1​cos⁡k​α\)\.\\\|\\dot\{\\bm\{w\}\}\_\{t\}\\\|^\{2\}=a\(r\_\{0\}^\{2\}\+r\_\{1\}^\{2\}\-2r\_\{0\}r\_\{1\}\\cos k\\alpha\)\.\(44\)Therefore, given the conditional variables𝒛=\(𝒙0,𝒙1\),𝒛′=\(𝒚0,𝒚1\)\\bm\{z\}=\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\}\),\\bm\{z\}^\{\\prime\}=\(\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\), we can obtain the analytic solution of the travelling pair path and the path action via \([43](https://arxiv.org/html/2609.20008#A1.E43),[44](https://arxiv.org/html/2609.20008#A1.E44)\)\. We first calculate𝒄0=𝒙0\+𝒚02,𝒄1=𝒙1\+𝒚12,𝒒0=𝒙0−𝒚0,𝒒1=𝒙1−𝒚1\\bm\{c\}\_\{0\}=\\frac\{\\bm\{x\}\_\{0\}\+\\bm\{y\}\_\{0\}\}\{2\},\\bm\{c\}\_\{1\}=\\frac\{\\bm\{x\}\_\{1\}\+\\bm\{y\}\_\{1\}\}\{2\},\\bm\{q\}\_\{0\}=\\bm\{x\}\_\{0\}\-\\bm\{y\}\_\{0\},\\bm\{q\}\_\{1\}=\\bm\{x\}\_\{1\}\-\\bm\{y\}\_\{1\}andr0=‖𝒒0‖,r1=‖𝒒1‖,α=arccos⁡⟨𝒒0,𝒒1⟩‖𝒒0‖​‖𝒒1‖r\_\{0\}=\\\|\\bm\{q\}\_\{0\}\\\|,r\_\{1\}=\\\|\\bm\{q\}\_\{1\}\\\|,\\alpha=\\arccos\\frac\{\\langle\\bm\{q\}\_\{0\},\\bm\{q\}\_\{1\}\\rangle\}\{\\\|\\bm\{q\}\_\{0\}\\\|\\\|\\bm\{q\}\_\{1\}\\\|\}\. Then we get the analytic path action

𝒜⁡\(𝒛,𝒛′\)=‖𝒄1−𝒄0‖2\+\(14\+λ\)​\(r02\+r12−2​r0​r1​cos⁡k​α\)\.\\mathcal\{A\}\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)=\\\|\\bm\{c\}\_\{1\}\-\\bm\{c\}\_\{0\}\\\|^\{2\}\+\(\\frac\{1\}\{4\}\+\\lambda\)\(r\_\{0\}^\{2\}\+r\_\{1\}^\{2\}\-2r\_\{0\}r\_\{1\}\\cos k\\alpha\)\.\(45\)Then, we calculate the analytic solution of the polar coordinates\.

\{rt=‖𝒘t‖/a=\(1−t\)2​r02\+t2​r12\+2​t​\(1−t\)​r0​r1​cos⁡k​αθt=Arg​\(𝒘t\)/k=1\+4​λ⋅Arg​\(\(1−t\)​r0\+t​r1​cos⁡k​α\+i⋅t​r1​sin⁡k​α\)\\left\\\{\\begin\{aligned\} r\_\{t\}&=\\\|\\bm\{w\}\_\{t\}\\\|/\\sqrt\{a\}=\\sqrt\{\(1\-t\)^\{2\}r\_\{0\}^\{2\}\+t^\{2\}r\_\{1\}^\{2\}\+2t\(1\-t\)r\_\{0\}r\_\{1\}\\cos k\\alpha\}\\\\ \\theta\_\{t\}&=\\text\{Arg\}\(\\bm\{w\}\_\{t\}\)/k=\\sqrt\{1\+4\\lambda\}\\cdot\\text\{Arg\}\\Big\(\(1\-t\)r\_\{0\}\+tr\_\{1\}\\cos k\\alpha\+i\\cdot tr\_\{1\}\\sin k\\alpha\\Big\)\\end\{aligned\}\\right\.\(46\)The analytic solution of C\-Q decomposition can be further calculated\.

\{𝒄t=\(1−t\)​𝒄0\+t​𝒄1𝒒t=rt​\(cos⁡θt,sin⁡θt\)\\left\\\{\\begin\{aligned\} \\bm\{c\}\_\{t\}&=\(1\-t\)\\bm\{c\}\_\{0\}\+t\\bm\{c\}\_\{1\}\\\\ \\bm\{q\}\_\{t\}&=r\_\{t\}\(\\cos\\theta\_\{t\},\\sin\\theta\_\{t\}\)\\end\{aligned\}\\right\.\(47\)The analytic solution of the travelling pair is

\{𝒙t=𝒄t\+12​𝒒t𝒚t=𝒄t−12​𝒒t,\{𝒙˙t=𝒄1−𝒄0\+12​𝒒˙t𝒚˙t=𝒄1−𝒄0−12​𝒒˙t\\left\\\{\\begin\{aligned\} \\bm\{x\}\_\{t\}&=\\bm\{c\}\_\{t\}\+\\frac\{1\}\{2\}\\bm\{q\}\_\{t\}\\\\ \\bm\{y\}\_\{t\}&=\\bm\{c\}\_\{t\}\-\\frac\{1\}\{2\}\\bm\{q\}\_\{t\}\\end\{aligned\}\\right\.,\\qquad\\left\\\{\\begin\{aligned\} \\dot\{\\bm\{x\}\}\_\{t\}&=\\bm\{c\}\_\{1\}\-\\bm\{c\}\_\{0\}\+\\frac\{1\}\{2\}\\dot\{\\bm\{q\}\}\_\{t\}\\\\ \\dot\{\\bm\{y\}\}\_\{t\}&=\\bm\{c\}\_\{1\}\-\\bm\{c\}\_\{0\}\-\\frac\{1\}\{2\}\\dot\{\\bm\{q\}\}\_\{t\}\\end\{aligned\}\\right\.\(48\)∎

### A\.4Proof of Theorem[4\.4](https://arxiv.org/html/2609.20008#S4.Thmtheorem4)

Theorem[4\.4](https://arxiv.org/html/2609.20008#S4.Thmtheorem4)Φ=\|dd​t​‖𝐪t‖\|2\\Phi=\|\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\\|\\bm\{q\}\_\{t\}\\\|\|^\{2\}\. Whenλ=0\\lambda=0andd≥2d\\geq 2, the corresponding static QOT reduces to standard OT, while whenλ→\+∞\\lambda\\to\+\\infty, it reduces to standard GW\-OT:λ−1​QOTS​\(μ0,μ1\)→GW\-OT​\(μ0,μ1\)\\lambda^\{\-1\}\\text\{QOT\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)\\to\\text\{GW\-OT\}\(\\mu\_\{0\},\\mu\_\{1\}\)\.

###### Proof\.

See[D\.1](https://arxiv.org/html/2609.20008#A4.SS1)for details\. ∎

### A\.5Proof of Theorem[5\.1](https://arxiv.org/html/2609.20008#S5.Thmtheorem1)

Theorem[5\.1](https://arxiv.org/html/2609.20008#S5.Thmtheorem1)\.Define the marginal velocity as𝐮ti\(𝐱,𝐲\)=∫𝐮ti\(𝐱,𝐲\|𝐳,𝐳′\)πt\(𝐱,𝐲\|𝐳,𝐳′\)q\(𝐳,𝐳′\)πt​\(𝐱,𝐲\)d𝐳d𝐳′\\bm\{u\}^\{i\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)=\\int\\bm\{u\}^\{i\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\frac\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\}\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\}\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}, the marginals satisfiy the marginal continuity equation\.

∂tπt\+∇𝒙⋅\(πt​𝒖t1\)\+∇𝒚⋅\(πt​𝒖t2\)=0\\partial\_\{t\}\\pi\_\{t\}\+\\nabla\_\{\\bm\{x\}\}\\cdot\(\\pi\_\{t\}\\bm\{u\}^\{1\}\_\{t\}\)\+\\nabla\_\{\\bm\{y\}\}\\cdot\(\\pi\_\{t\}\\bm\{u\}^\{2\}\_\{t\}\)=0\(49\)Ifq=γ⊗γq=\\gamma\\otimes\\gammafor someγ∈Π⁡\(μ0,μ1\)\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\), the marginal velocity transportsμ0⊗μ0\\mu\_\{0\}\\otimes\\mu\_\{0\}toμ1⊗μ1\\mu\_\{1\}\\otimes\\mu\_\{1\}\.

###### Proof\.

Consider the probability pathπt​\(𝒙,𝒚\)\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)on the two\-particle augmented space𝒳2\\mathcal\{X\}^\{2\}, the corresponding velocity is𝒖t​\(𝒙,𝒚\)=\(𝒖t1​\(𝒙,𝒚\)𝒖t2​\(𝒙,𝒚\)\)\\bm\{u\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)=\\begin\{pmatrix\}\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\\\ \\bm\{u\}^\{2\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\end\{pmatrix\}\. Applying the marginalization theorem of standard conditional flow matching\([Lipman et al\., 2023](https://arxiv.org/html/2609.20008#bib.bib18);[Tong et al\., 2024a](https://arxiv.org/html/2609.20008#bib.bib19)\), we have

∂tπt\+∇\(𝒙,𝒚\)⋅\(πt​𝒖t\)=0\\partial\_\{t\}\\pi\_\{t\}\+\\nabla\_\{\(\\bm\{x\},\\bm\{y\}\)\}\\cdot\(\\pi\_\{t\}\\bm\{u\}\_\{t\}\)=0\(50\)By definition,∇\(𝒙,𝒚\)⋅\(πt​𝒖t\)=∇𝒙⋅\(πt​𝒖t1\)\+∇𝒚⋅\(πt​𝒖t2\)\\nabla\_\{\(\\bm\{x\},\\bm\{y\}\)\}\\cdot\(\\pi\_\{t\}\\bm\{u\}\_\{t\}\)=\\nabla\_\{\\bm\{x\}\}\\cdot\(\\pi\_\{t\}\\bm\{u\}^\{1\}\_\{t\}\)\+\\nabla\_\{\\bm\{y\}\}\\cdot\(\\pi\_\{t\}\\bm\{u\}^\{2\}\_\{t\}\)\. Therefore, identity \([49](https://arxiv.org/html/2609.20008#A1.E49)\) holds\. If we further haveq=γ⊗γq=\\gamma\\otimes\\gammafor someγ∈Π⁡\(μ0,μ1\)\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\), then by definition, we have

π0​\(𝒙,𝒚\)\\displaystyle\\pi\_\{0\}\(\\bm\{x\},\\bm\{y\}\)=∫π0\(𝒙,𝒚\|𝒛,𝒛′\)q\(𝒛,𝒛′\)d𝒛d𝒛′\\displaystyle=\\int\\pi\_\{0\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}\(51\)=∫δ\(𝒙0,𝒚0\)​\(𝒙,𝒚\)​γ​\(𝒙0,𝒙1\)​γ​\(𝒚0,𝒚1\)​d​𝒙0​d​𝒙1​d​𝒚0​d​𝒚1\\displaystyle=\\int\\delta\_\{\(\\bm\{x\}\_\{0\},\\bm\{y\}\_\{0\}\)\}\(\\bm\{x\},\\bm\{y\}\)\\gamma\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\}\)\\gamma\(\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\)\\mathrm\{d\}\\bm\{x\}\_\{0\}\\mathrm\{d\}\\bm\{x\}\_\{1\}\\mathrm\{d\}\\bm\{y\}\_\{0\}\\mathrm\{d\}\\bm\{y\}\_\{1\}=∫δ𝒙0​\(𝒙\)​γ​\(𝒙0,𝒙1\)​d​𝒙0​d​𝒙1​∫δ𝒚0​\(𝒚\)​γ​\(𝒚0,𝒚1\)​d​𝒚0​d​𝒚1\\displaystyle=\\int\\delta\_\{\\bm\{x\}\_\{0\}\}\(\\bm\{x\}\)\\gamma\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\}\)\\mathrm\{d\}\\bm\{x\}\_\{0\}\\mathrm\{d\}\\bm\{x\}\_\{1\}\\int\\delta\_\{\\bm\{y\}\_\{0\}\}\(\\bm\{y\}\)\\gamma\(\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\)\\mathrm\{d\}\\bm\{y\}\_\{0\}\\mathrm\{d\}\\bm\{y\}\_\{1\}=μ0​\(𝒙\)​μ0​\(𝒚\)\\displaystyle=\\mu\_\{0\}\(\\bm\{x\}\)\\mu\_\{0\}\(\\bm\{y\}\)where𝒛=\(𝒙0,𝒙𝟏\),𝒛′=\(𝒚0,𝒚1\)\\bm\{z\}=\(\\bm\{x\}\_\{0\},\\bm\{x\_\{1\}\}\),\\bm\{z\}^\{\\prime\}=\(\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\)\. Similarly,π1​\(𝒙,𝒚\)=μ1​\(𝒙\)​μ1​\(𝒚\)\\pi\_\{1\}\(\\bm\{x\},\\bm\{y\}\)=\\mu\_\{1\}\(\\bm\{x\}\)\\mu\_\{1\}\(\\bm\{y\}\)\. Therefore, the marginal velocity transportsμ0⊗μ0\\mu\_\{0\}\\otimes\\mu\_\{0\}toμ1⊗μ1\\mu\_\{1\}\\otimes\\mu\_\{1\}\. ∎

### A\.6Proof of Theorem[5\.2](https://arxiv.org/html/2609.20008#S5.Thmtheorem2)

Theorem[5\.2](https://arxiv.org/html/2609.20008#S5.Thmtheorem2)\.The single particle marginals satisfy the continuity equation

∂tρt\+∇𝒙⋅\(ρt​𝒗t\)=0\\partial\_\{t\}\\rho\_\{t\}\+\\nabla\_\{\\bm\{x\}\}\\cdot\(\\rho\_\{t\}\\bm\{v\}\_\{t\}\)=0\(52\)
###### Proof\.

Recall the definitions

ρt​\(𝒙\)=∫𝒳πt​\(𝒙,𝒚\)​𝑑𝒚,𝒗t​\(𝒙\)=∫𝒳𝒖t1​\(𝒙,𝒚\)​πt​\(𝒚\|𝒙\)​𝑑𝒚=∫𝒳𝒖t1​\(𝒙,𝒚\)​πt​\(𝒙,𝒚\)ρt​\(𝒙\)​𝑑𝒚\.\\rho\_\{t\}\(\\bm\{x\}\)=\\int\_\{\\mathcal\{X\}\}\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{y\},\\quad\\bm\{v\}\_\{t\}\(\\bm\{x\}\)=\\int\_\{\\mathcal\{X\}\}\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\pi\_\{t\}\(\\bm\{y\}\|\\bm\{x\}\)\\mathrm\{d\}\\bm\{y\}=\\int\_\{\\mathcal\{X\}\}\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\frac\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\}\{\\rho\_\{t\}\(\\bm\{x\}\)\}\\mathrm\{d\}\\bm\{y\}\.\(53\)By direct calculation, we have

∂tρt​\(𝒙\)\\displaystyle\\partial\_\{t\}\\rho\_\{t\}\(\\bm\{x\}\)=∂t∫𝒳πt​\(𝒙,𝒚\)​𝒅𝒚=∫𝒳∂tπt​\(𝒙,𝒚\)​𝒅𝒚\\displaystyle=\\partial\_\{t\}\\int\_\{\\mathcal\{X\}\}\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{y\}=\\int\_\{\\mathcal\{X\}\}\\partial\_\{t\}\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{y\}\(54\)=−∫𝒳∇𝒙⋅\(πt\(𝒙,𝒚\)𝒖1t\(𝒙,𝒚\)\)\+∇𝒚⋅\(πt\(𝒙,𝒚\)𝒖2t\(𝒙,𝒚\)\)d𝒚\\displaystyle=\-\\int\_\{\\mathcal\{X\}\}\\nabla\_\{\\bm\{x\}\}\\cdot\(\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\)\+\\nabla\_\{\\bm\{y\}\}\\cdot\(\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\bm\{u\}^\{2\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\)\\mathrm\{d\}\\bm\{y\}=−∇𝒙⋅∫𝒳πt\(𝒙,𝒚\)𝒖1t\(𝒙,𝒚\)d𝒚−ρt\(𝒙\)∫𝒳∇𝒚⋅\(πt\(𝒚\|𝒙\)𝒖2t\(𝒙,𝒚\)\)d𝒚\\displaystyle=\-\\nabla\_\{\\bm\{x\}\}\\cdot\\int\_\{\\mathcal\{X\}\}\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{y\}\-\\rho\_\{t\}\(\\bm\{x\}\)\\int\_\{\\mathcal\{X\}\}\\nabla\_\{\\bm\{y\}\}\\cdot\\Big\(\\pi\_\{t\}\(\\bm\{y\}\|\\bm\{x\}\)\\bm\{u\}^\{2\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\Big\)\\mathrm\{d\}\\bm\{y\}=−∇𝒙⋅∫𝒳πt\(𝒙,𝒚\)𝒖1t\(𝒙,𝒚\)d𝒚\\displaystyle=\-\\nabla\_\{\\bm\{x\}\}\\cdot\\int\_\{\\mathcal\{X\}\}\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{y\}=−∇𝒙⋅\(ρt\(𝒙\)∫𝒳πt​\(𝒙,𝒚\)ρt​\(𝒙\)𝒖1t\(𝒙,𝒚\)d𝒚\)\\displaystyle=\-\\nabla\_\{\\bm\{x\}\}\\cdot\\Big\(\\rho\_\{t\}\(\\bm\{x\}\)\\int\_\{\\mathcal\{X\}\}\\frac\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\}\{\\rho\_\{t\}\(\\bm\{x\}\)\}\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{y\}\\Big\)=−∇𝒙⋅\(ρt\(𝒙\)𝒗t\(𝒙\)\)\\displaystyle=\-\\nabla\_\{\\bm\{x\}\}\\cdot\(\\rho\_\{t\}\(\\bm\{x\}\)\\bm\{v\}\_\{t\}\(\\bm\{x\}\)\)∎

Remark\.If we further have∫q⁡\(𝒛,𝒛′\)​d​𝒛′=γ⁡\(𝒛\)\\int q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{z\}^\{\\prime\}=\\gamma\(\\bm\{z\}\)for someγ∈Π⁡\(μ0,μ1\)\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\), then it is easy to checkρ0=μ0,ρ1=μ1\\rho\_\{0\}=\\mu\_\{0\},\\rho\_\{1\}=\\mu\_\{1\}\. In this case,𝒗t\\bm\{v\}\_\{t\}transportsμ0\\mu\_\{0\}toμ1\\mu\_\{1\}\. We point out that this condition is slightly weaker thanq=γ⊗γq=\\gamma\\otimes\\gammafor someγ∈Π⁡\(μ0,μ1\)\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\), which we used in theorem[5\.1](https://arxiv.org/html/2609.20008#S5.Thmtheorem1)\.

### A\.7Proof of Proposition[5\.3](https://arxiv.org/html/2609.20008#S5.Thmtheorem3)

Proposition[5\.3](https://arxiv.org/html/2609.20008#S5.Thmtheorem3)\.Let the true probability flow beρt\\rho\_\{t\}and the mean\-field approximation beρ^t\\hat\{\\rho\}\_\{t\}\. Assume𝐛t,𝐟t\\bm\{b\}\_\{t\},\\bm\{f\}\_\{t\}are both Lipschitz, andM=max0<t≤1max𝐱𝒲1\(ρt,πt\(⋅\|𝐱\)\)M=\\underset\{0<t\\leq 1\}\{\\max\}\\underset\{\\bm\{x\}\}\{\\max\}\\mathcal\{W\}\_\{1\}\(\\rho\_\{t\},\\pi\_\{t\}\(\\cdot\|\\bm\{x\}\)\), then∀0<t≤1,∃Ct\>0\\forall 0<t\\leq 1,\\exists C\_\{t\}\>0such that𝒲1​\(ρt,ρ^t\)≤Ct​M\\mathcal\{W\}\_\{1\}\(\\rho\_\{t\},\\hat\{\\rho\}\_\{t\}\)\\leq C\_\{t\}M\.

###### Proof\.

LetLb,LfL\_\{b\},L\_\{f\}be the Lipschitz constant of𝒃t,𝒇t\\bm\{b\}\_\{t\},\\bm\{f\}\_\{t\}, andL≥Lb,LfL\\geq L\_\{b\},L\_\{f\}\. The true dynamics is

∂tρt​\(𝒙\)\+∇𝒙⋅\[ρt​\(𝒙\)​\(𝒃t​\(𝒙\)\+∫𝒳𝒇t​\(𝒙,𝒚\)​πt​\(𝒚\|𝒙\)​𝑑𝒚\)\]=0\.\\partial\_\{t\}\\rho\_\{t\}\(\\bm\{x\}\)\+\\nabla\_\{\\bm\{x\}\}\\cdot\[\\rho\_\{t\}\(\\bm\{x\}\)\\big\(\\bm\{b\}\_\{t\}\(\\bm\{x\}\)\+\\int\_\{\\mathcal\{X\}\}\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\pi\_\{t\}\(\\bm\{y\}\|\\bm\{x\}\)\\mathrm\{d\}\\bm\{y\}\\big\)\]=0\.\(55\)The mean\-field approximation is

∂tρ^t​\(𝒙\)\+∇𝒙⋅\[ρ^t​\(𝒙\)​\(𝒃t​\(𝒙\)\+∫𝒳𝒇t​\(𝒙,𝒚\)​ρ^t​\(𝒚\)​𝑑𝒚\)\]=0\.\\partial\_\{t\}\\hat\{\\rho\}\_\{t\}\(\\bm\{x\}\)\+\\nabla\_\{\\bm\{x\}\}\\cdot\[\\hat\{\\rho\}\_\{t\}\(\\bm\{x\}\)\\big\(\\bm\{b\}\_\{t\}\(\\bm\{x\}\)\+\\int\_\{\\mathcal\{X\}\}\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\hat\{\\rho\}\_\{t\}\(\\bm\{y\}\)\\mathrm\{d\}\\bm\{y\}\\big\)\]=0\.\(56\)LetRt​\(𝒙\)=∫𝒳𝒇t​\(𝒙,𝒚\)​\[πt​\(𝒚\|𝒙\)−ρ^t​\(𝒚\)\]​𝑑𝒚R\_\{t\}\(\\bm\{x\}\)=\\int\_\{\\mathcal\{X\}\}\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\[\\pi\_\{t\}\(\\bm\{y\}\|\\bm\{x\}\)\-\\hat\{\\rho\}\_\{t\}\(\\bm\{y\}\)\]\\mathrm\{d\}\\bm\{y\}represent the difference between the two velocity fields\. By the Lipschitz condition, we have

‖Rt​\(𝒙\)‖\\displaystyle\\\|R\_\{t\}\(\\bm\{x\}\)\\\|=‖∫𝒳𝒇t​\(𝒙,𝒚\)​\[πt​\(𝒚\|𝒙\)−ρ^t​\(𝒚\)\]​d𝒚‖\\displaystyle=\\\|\\int\_\{\\mathcal\{X\}\}\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\[\\pi\_\{t\}\(\\bm\{y\}\|\\bm\{x\}\)\-\\hat\{\\rho\}\_\{t\}\(\\bm\{y\}\)\]\\mathrm\{d\}\\bm\{y\}\\\|\(57\)≤\|∫𝒳𝒇t​\(𝒙,𝒚\)​\[πt​\(𝒚\|𝒙\)−ρt​\(𝒚\)\]​d𝒚\|\+‖∫𝒳𝒇t​\(𝒙,𝒚\)​\[ρt​\(𝒚\)−ρ^t​\(𝒚\)\]​d𝒚‖\\displaystyle\\leq\\\|\\int\_\{\\mathcal\{X\}\}\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\[\\pi\_\{t\}\(\\bm\{y\}\|\\bm\{x\}\)\-\\rho\_\{t\}\(\\bm\{y\}\)\]\\mathrm\{d\}\\bm\{y\}\\\|\+\\\|\\int\_\{\\mathcal\{X\}\}\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\[\\rho\_\{t\}\(\\bm\{y\}\)\-\\hat\{\\rho\}\_\{t\}\(\\bm\{y\}\)\]\\mathrm\{d\}\\bm\{y\}\\\|≤L\(𝒲1\(πt\(⋅\|𝒙\),ρt\)\+𝒲1\(ρt,ρ^t\)\)\\displaystyle\\leq L\\Big\(\\mathcal\{W\}\_\{1\}\(\\pi\_\{t\}\(\\cdot\|\\bm\{x\}\),\\rho\_\{t\}\)\+\\mathcal\{W\}\_\{1\}\(\\rho\_\{t\},\\hat\{\\rho\}\_\{t\}\)\\Big\)≤L​M\+L​𝒲1​\(ρt,ρ^t\)\.\\displaystyle\\leq LM\+L\\mathcal\{W\}\_\{1\}\(\\rho\_\{t\},\\hat\{\\rho\}\_\{t\}\)\.Sinceρ^0=ρ0\\hat\{\\rho\}\_\{0\}=\\rho\_\{0\}, the difference betweenρt,ρ^t\\rho\_\{t\},\\hat\{\\rho\}\_\{t\}can be fully described by the difference between the velocity fields\. For any starting pointX0=Y0=𝒙X\_\{0\}=Y\_\{0\}=\\bm\{x\}, the true dynamics is

dd​t​Xt=𝒃t​\(Xt\)\+∫𝒳𝒇t​\(Xt,𝒚\)​πt​\(𝒚\|Xt\)​𝑑𝒚≜A⁡\(Xt\)\.\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}X\_\{t\}=\\bm\{b\}\_\{t\}\(X\_\{t\}\)\+\\int\_\{\\mathcal\{X\}\}\\bm\{f\}\_\{t\}\(X\_\{t\},\\bm\{y\}\)\\pi\_\{t\}\(\\bm\{y\}\|X\_\{t\}\)\\mathrm\{d\}\\bm\{y\}\\triangleq A\(X\_\{t\}\)\.\(58\)The mean\-field dynamics is

dd​t​Yt=𝒃t​\(Yt\)\+∫𝒳𝒇t​\(Yt,𝒚\)​ρ^t​\(𝒚\)​𝑑𝒚≜B⁡\(Yt\)\.\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}Y\_\{t\}=\\bm\{b\}\_\{t\}\(Y\_\{t\}\)\+\\int\_\{\\mathcal\{X\}\}\\bm\{f\}\_\{t\}\(Y\_\{t\},\\bm\{y\}\)\\hat\{\\rho\}\_\{t\}\(\\bm\{y\}\)\\mathrm\{d\}\\bm\{y\}\\triangleq B\(Y\_\{t\}\)\.\(59\)LetZt=‖Xt−Yt‖Z\_\{t\}=\\\|X\_\{t\}\-Y\_\{t\}\\\|, we have

dd​t​\(Xt−Yt\)\\displaystyle\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\(X\_\{t\}\-Y\_\{t\}\)=A⁡\(Xt\)−B⁡\(Yt\)=A⁡\(Xt\)−B⁡\(Xt\)\+B⁡\(Xt\)−B⁡\(Yt\)\\displaystyle=A\(X\_\{t\}\)\-B\(Y\_\{t\}\)=A\(X\_\{t\}\)\-B\(X\_\{t\}\)\+B\(X\_\{t\}\)\-B\(Y\_\{t\}\)\(60\)=Rt​\(Xt\)\+\[𝒃t​\(Xt\)−𝒃t​\(Yt\)\+∫𝒳\(𝒇t​\(Xt,𝒚\)−𝒇t​\(Yt,𝒚\)\)​ρ^t​\(𝒚\)​d𝒚\]\\displaystyle=R\_\{t\}\(X\_\{t\}\)\+\[\\bm\{b\}\_\{t\}\(X\_\{t\}\)\-\\bm\{b\}\_\{t\}\(Y\_\{t\}\)\+\\int\_\{\\mathcal\{X\}\}\\Big\(\\bm\{f\}\_\{t\}\(X\_\{t\},\\bm\{y\}\)\-\\bm\{f\}\_\{t\}\(Y\_\{t\},\\bm\{y\}\)\\Big\)\\hat\{\\rho\}\_\{t\}\(\\bm\{y\}\)\\mathrm\{d\}\\bm\{y\}\]Therefore,

dd​t​Zt≤‖dd​t​\(Xt−Yt\)‖≤L​M\+L​𝒲1​\(ρt,ρ^t\)\+2​L​Zt\.\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}Z\_\{t\}\\leq\\\|\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\(X\_\{t\}\-Y\_\{t\}\)\\\|\\leq LM\+L\\mathcal\{W\}\_\{1\}\(\\rho\_\{t\},\\hat\{\\rho\}\_\{t\}\)\+2LZ\_\{t\}\.\(61\)By couplingXt,YtX\_\{t\},Y\_\{t\}generated by the same starting point𝒙\\bm\{x\}together, we obtain an inequality of𝒲1​\(ρt,ρ^t\)\\mathcal\{W\}\_\{1\}\(\\rho\_\{t\},\\hat\{\\rho\}\_\{t\}\)\.

𝒲1​\(ρt,ρ^t\)\\displaystyle\\mathcal\{W\}\_\{1\}\(\\rho\_\{t\},\\hat\{\\rho\}\_\{t\}\)=infγ∫𝒳2‖𝒙−𝒚‖​γ​\(𝒙,𝒚\)​𝒅𝒙​𝒅𝒚\\displaystyle=\\inf\_\{\\gamma\}\\int\_\{\\mathcal\{X\}^\{2\}\}\\\|\\bm\{x\}\-\\bm\{y\}\\\|\\gamma\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\(62\)≤∫𝒳2Zt​\(𝒙\)​ρ0​\(𝒙\)​d𝒙\.\\displaystyle\\leq\\int\_\{\\mathcal\{X\}^\{2\}\}Z\_\{t\}\(\\bm\{x\}\)\\rho\_\{0\}\(\\bm\{x\}\)\\mathrm\{d\}\\bm\{x\}\.Therefore, it is sufficient to estimate the upper bound ofZtZ\_\{t\}\. By Gronwall’s inequality, we have

Zt\\displaystyle Z\_\{t\}≤e2​L​t​L​∫0te−2​L​s​\(M\+𝒲1​\(ρs,ρ^s\)\)​𝑑s\\displaystyle\\leq e^\{2Lt\}L\\int\_\{0\}^\{t\}e^\{\-2Ls\}\(M\+\\mathcal\{W\}\_\{1\}\(\\rho\_\{s\},\\hat\{\\rho\}\_\{s\}\)\)\\mathrm\{d\}s\(63\)=e2​L​t−12​M\+e2​L​t​L​∫0te−2​L​s​𝒲1​\(ρs,ρ^s\)​ds\.\\displaystyle=\\frac\{e^\{2Lt\}\-1\}\{2\}M\+e^\{2Lt\}L\\int\_\{0\}^\{t\}e^\{\-2Ls\}\\mathcal\{W\}\_\{1\}\(\\rho\_\{s\},\\hat\{\\rho\}\_\{s\}\)\\mathrm\{d\}s\.Therefore, we get another inequality of𝒲1​\(ρt,ρ^t\)\\mathcal\{W\}\_\{1\}\(\\rho\_\{t\},\\hat\{\\rho\}\_\{t\}\)\.

𝒲1​\(ρt,ρ^t\)\\displaystyle\\mathcal\{W\}\_\{1\}\(\\rho\_\{t\},\\hat\{\\rho\}\_\{t\}\)≤∫𝒳2\(e2​L​t−12​M\+e2​L​t​L​∫0te−2​L​s​𝒲1​\(ρs,ρ^s\)​𝒅s\)​ρ0​\(𝒙\)​𝒅𝒙\\displaystyle\\leq\\int\_\{\\mathcal\{X\}^\{2\}\}\\Big\(\\frac\{e^\{2Lt\}\-1\}\{2\}M\+e^\{2Lt\}L\\int\_\{0\}^\{t\}e^\{\-2Ls\}\\mathcal\{W\}\_\{1\}\(\\rho\_\{s\},\\hat\{\\rho\}\_\{s\}\)\\mathrm\{d\}s\\Big\)\\rho\_\{0\}\(\\bm\{x\}\)\\mathrm\{d\}\\bm\{x\}\(64\)=e2​L​t−12​M\+e2​L​t​L​∫0te−2​L​s​𝒲1​\(ρs,ρ^s\)​ds\.\\displaystyle=\\frac\{e^\{2Lt\}\-1\}\{2\}M\+e^\{2Lt\}L\\int\_\{0\}^\{t\}e^\{\-2Ls\}\\mathcal\{W\}\_\{1\}\(\\rho\_\{s\},\\hat\{\\rho\}\_\{s\}\)\\mathrm\{d\}s\.Or equivalently, we have

e−2​L​t​𝒲1​\(ρt,ρ^t\)≤1−e−2​L​t2​M\+L​∫0te−2​L​s​𝒲1​\(ρs,ρ^s\)​𝑑s\.e^\{\-2Lt\}\\mathcal\{W\}\_\{1\}\(\\rho\_\{t\},\\hat\{\\rho\}\_\{t\}\)\\leq\\frac\{1\-e^\{\-2Lt\}\}\{2\}M\+L\\int\_\{0\}^\{t\}e^\{\-2Ls\}\\mathcal\{W\}\_\{1\}\(\\rho\_\{s\},\\hat\{\\rho\}\_\{s\}\)\\mathrm\{d\}s\.\(65\)LetF⁡\(t\)=∫0te−2​L​s​𝒲1​\(ρs,ρ^s\)​𝑑sF\(t\)=\\int\_\{0\}^\{t\}e^\{\-2Ls\}\\mathcal\{W\}\_\{1\}\(\\rho\_\{s\},\\hat\{\\rho\}\_\{s\}\)\\mathrm\{d\}s, the inequality can be reformulated as

F′​\(t\)≤1−e−2​L​t2​M\+L​F​\(t\)\.F^\{\\prime\}\(t\)\\leq\\frac\{1\-e^\{\-2Lt\}\}\{2\}M\+LF\(t\)\.\(66\)By Gronwall’s inequality again, we have

F⁡\(t\)≤M6​L​\(2​eL​t−3\+e−2​L​t\)\.F\(t\)\\leq\\frac\{M\}\{6L\}\(2e^\{Lt\}\-3\+e^\{\-2Lt\}\)\.\(67\)Plug in \([66](https://arxiv.org/html/2609.20008#A1.E66)\), we have

𝒲1​\(ρt,ρ^t\)≤e3​L​t−13​M\.\\mathcal\{W\}\_\{1\}\(\\rho\_\{t\},\\hat\{\\rho\}\_\{t\}\)\\leq\\frac\{e^\{3Lt\}\-1\}\{3\}M\.\(68\)
∎

### A\.8Proof of Theorem[5\.4](https://arxiv.org/html/2609.20008#S5.Thmtheorem4)

Theorem[5\.4](https://arxiv.org/html/2609.20008#S5.Thmtheorem4)\.The minimizer is𝐮𝛉𝐢​\(𝐱,𝐲,t\)=𝐮ti​\(𝐱,𝐲\)\\bm\{\\bm\{u\}^\{i\}\_\{\\theta\}\}\(\\bm\{x\},\\bm\{y\},t\)=\\bm\{u\}^\{i\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\. Whenq=γ⊗γq=\\gamma\\otimes\\gammawhereγ\\gammais the optimal QOT coupling of \([9](https://arxiv.org/html/2609.20008#S4.E9)\), and the travelling pair used is the minimizer of \([7](https://arxiv.org/html/2609.20008#S4.E7)\), then the learned velocity pair generates the probability flow of the dynamic QOT problem \([10](https://arxiv.org/html/2609.20008#S4.E10)\)\.

###### Proof\.

Recall the conditional loss

ℒpair\(𝜽\)=𝔼t∼𝒰\[0,1\],\(𝒛,𝒛′\)∼q\(𝒛,𝒛′\),\(𝒙,𝒚\)∼πt\(𝒙,𝒚\|𝒛,𝒛′\)∑i=12‖𝒖𝜽𝒊\(𝒙,𝒚,t\)−𝒖ti\(𝒙,𝒚\|𝒛,𝒛′\)‖22\\mathcal\{L\}\_\{\\text\{pair\}\}\(\\bm\{\\theta\}\)=\\mathbb\{E\}\_\{t\\sim\\mathcal\{U\}\[0,1\],\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\sim q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\),\(\\bm\{x\},\\bm\{y\}\)\\sim\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\}\\sum\_\{i=1\}^\{2\}\\left\\\|\\bm\{\\bm\{u\}^\{i\}\_\{\\theta\}\}\(\\bm\{x\},\\bm\{y\},t\)\-\\bm\{u\}^\{i\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\right\\\|\_\{2\}^\{2\}\(69\)Similarly to what we have done in the proof of theorem[5\.1](https://arxiv.org/html/2609.20008#S5.Thmtheorem1), on the two\-particle augmented space, the standard equivalence between the marginal flow matching loss and the conditional flow matching loss\([Lipman et al\., 2023](https://arxiv.org/html/2609.20008#bib.bib18);[Tong et al\., 2024a](https://arxiv.org/html/2609.20008#bib.bib19)\)tells us that minimizing the conditional loss \([69](https://arxiv.org/html/2609.20008#A1.E69)\) is equivalent to minimizing the marginal loss

ℒpair\-marginal​\(𝜽\)=𝔼t∼𝒰⁡\[0,1\],\(𝒙,𝒚\)∼πt​\(𝒙,𝒚\)​∑i=12‖𝒖𝜽𝒊​\(𝒙,𝒚,t\)−𝒖ti​\(𝒙,𝒚\)‖22\\mathcal\{L\}\_\{\\text\{pair\-marginal\}\}\(\\bm\{\\theta\}\)=\\mathbb\{E\}\_\{t\\sim\\mathcal\{U\}\[0,1\],\(\\bm\{x\},\\bm\{y\}\)\\sim\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\}\\sum\_\{i=1\}^\{2\}\\left\\\|\\bm\{\\bm\{u\}^\{i\}\_\{\\theta\}\}\(\\bm\{x\},\\bm\{y\},t\)\-\\bm\{u\}^\{i\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\right\\\|\_\{2\}^\{2\}\(70\)Therefore, the minimizer is𝒖𝜽𝒊​\(𝒙,𝒚,t\)=𝒖ti​\(𝒙,𝒚\)\\bm\{\\bm\{u\}^\{i\}\_\{\\theta\}\}\(\\bm\{x\},\\bm\{y\},t\)=\\bm\{u\}^\{i\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\. By the pair marginalization theorem[5\.1](https://arxiv.org/html/2609.20008#S5.Thmtheorem1), we know that the minimizer generates the probability flow of the dynamic QOT\. ∎

Remark\.Although the learned velocity induces the same probability flow as the dynamic QOT construction, the mechanism by which it generates this flow is not exactly the same as in the definition of dynamic QOT\. In dynamic QOT, the coupling first determines where each particle is transported, and the travelling\-pair path then specifies how each pairwise transport is carried out\. The resulting process is therefore generally non\-Markovian, since its evolution depends on the prescribed endpoint\. In realistic dynamics, however, future states are typically not known in advance\. Instead, the evolution is determined by the current state, possibly together with its history\. A natural approximation is therefore to replace the original process by a Markov process\. The velocity pair learned by flow matching defines a pair of time\-dependent velocity fields on the two\-particle space, so particles evolve according to their current positions and time, making the resulting dynamics Markovian\. The two processes share the same two\-particle marginal distributionπt​\(𝒙,𝒚\)\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)at every timett\. In this sense, the learned velocity pair can be viewed as providing a Markovian projection of the original endpoint\-conditioned process\.

### A\.9Proof of Theorem[5\.5](https://arxiv.org/html/2609.20008#S5.Thmtheorem5)

Theorem[5\.5](https://arxiv.org/html/2609.20008#S5.Thmtheorem5)\.∇𝜽ℒM​\(𝜽\)=∇𝜽ℒC​\(𝜽\)\\nabla\_\{\\bm\{\\theta\}\}\\mathcal\{L\}\_\{M\}\(\\bm\{\\theta\}\)=\\nabla\_\{\\bm\{\\theta\}\}\\mathcal\{L\}\_\{C\}\(\\bm\{\\theta\}\)\.

###### Proof\.

We first recall the two losses\.

ℒM​\(𝜽\)=𝔼t∼𝒰⁡\[0,1\],𝒙∼ρt​\(𝒙\)​‖𝒗𝜽​\(𝒙,t\)−𝒗t​\(𝒙\)‖2\\displaystyle\\mathcal\{L\}\_\{M\}\(\\bm\{\\theta\)\}=\\mathbb\{E\}\_\{t\\sim\\mathcal\{U\}\[0,1\],\\bm\{x\}\\sim\\rho\_\{t\}\(\\bm\{x\}\)\}\\\|\\bm\{v\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},t\)\-\\bm\{v\}\_\{t\}\(\\bm\{x\}\)\\\|^\{2\}\(71\)ℒC\(𝜽\)=𝔼t∼𝒰\[0,1\],\(𝒛,𝒛′\)∼q\(𝒛,𝒛′\),\(𝒙,𝒚\)∼πt\(𝒙,𝒚\|𝒛,𝒛′\)∥𝒗𝜽\(𝒙,t\)−𝒖1t\(𝒙,𝒚\|𝒛,𝒛′\)∥2\\displaystyle\\mathcal\{L\}\_\{C\}\(\\bm\{\\theta\)\}=\\mathbb\{E\}\_\{t\\sim\\mathcal\{U\}\[0,1\],\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\sim q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\),\(\\bm\{x\},\\bm\{y\}\)\\sim\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\}\\\|\\bm\{v\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},t\)\-\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\\|^\{2\}By direct calculation, we have

ℒC​\(𝜽\)\\displaystyle\\mathcal\{L\}\_\{C\}\(\\bm\{\\theta\)\}=𝔼t∼𝒰\[0,1\],\(𝒛,𝒛′\)∼q\(𝒛,𝒛′\),\(𝒙,𝒚\)∼πt\(𝒙,𝒚\|𝒛,𝒛′\)∥𝒗𝜽\(𝒙,t\)−𝒖1t\(𝒙,𝒚\|𝒛,𝒛′\)∥2\\displaystyle=\\mathbb\{E\}\_\{t\\sim\\mathcal\{U\}\[0,1\],\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\sim q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\),\(\\bm\{x\},\\bm\{y\}\)\\sim\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\}\\\|\\bm\{v\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},t\)\-\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\\|^\{2\}\(72\)=∫01∫∫𝒳2∥𝒗𝜽\(𝒙,t\)−𝒖1t\(𝒙,𝒚\|𝒛,𝒛′\)∥2πt\(𝒙,𝒚\|𝒛,𝒛′\)q\(𝒛,𝒛′\)d𝒙d𝒚d𝒛d𝒛′dt\\displaystyle=\\int\_\{0\}^\{1\}\\int\\int\_\{\\mathcal\{X\}^\{2\}\}\\\|\\bm\{v\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},t\)\-\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\\|^\{2\}\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}\\mathrm\{d\}t=∫01∫∫𝒳2∥𝒗𝜽\(𝒙,t\)∥2πt\(𝒙,𝒚\|𝒛,𝒛′\)q\(𝒛,𝒛′\)d𝒙d𝒚d𝒛d𝒛′dt\\displaystyle=\\int\_\{0\}^\{1\}\\int\\int\_\{\\mathcal\{X\}^\{2\}\}\\\|\\bm\{v\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},t\)\\\|^\{2\}\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}\\mathrm\{d\}t\+∫01∫∫𝒳2∥𝒖1t\(𝒙,𝒚\|𝒛,𝒛′\)∥2πt\(𝒙,𝒚\|𝒛,𝒛′\)q\(𝒛,𝒛′\)d𝒙d𝒚d𝒛d𝒛′dt\\displaystyle\+\\int\_\{0\}^\{1\}\\int\\int\_\{\\mathcal\{X\}^\{2\}\}\\\|\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\\|^\{2\}\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}\\mathrm\{d\}t−2∫01∫∫𝒳2⟨𝒗𝜽\(𝒙,t\),𝒖1t\(𝒙,𝒚\|𝒛,𝒛′\)⟩πt\(𝒙,𝒚\|𝒛,𝒛′\)q\(𝒛,𝒛′\)d𝒙d𝒚d𝒛d𝒛′dt\\displaystyle\-2\\int\_\{0\}^\{1\}\\int\\int\_\{\\mathcal\{X\}^\{2\}\}\\langle\\bm\{v\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},t\),\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\rangle\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}\\mathrm\{d\}t=∫01∫𝒳2‖𝒗𝜽​\(𝒙,t\)‖2​πt​\(𝒙,𝒚\)​𝑑𝒙​𝑑𝒚​𝑑t\+C\\displaystyle=\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}^\{2\}\}\\\|\\bm\{v\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},t\)\\\|^\{2\}\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}t\+C−2∫01∫𝒳2⟨𝒗𝜽\(𝒙,t\),∫𝒖1t\(𝒙,𝒚\|𝒛,𝒛′\)πt\(𝒙,𝒚\|𝒛,𝒛′\)q\(𝒛,𝒛′\)πt​\(𝒙,𝒚\)d𝒛d𝒛′⟩πt\(𝒙,𝒚\)d𝒙d𝒚dt\\displaystyle\-2\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}^\{2\}\}\\langle\\bm\{v\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},t\),\\int\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\frac\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\}\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\}\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}\\rangle\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}t=∫01∫𝒳2‖𝒗𝜽​\(𝒙,t\)‖2​πt​\(𝒙,𝒚\)​𝑑𝒙​𝑑𝒚​𝑑t\+C\\displaystyle=\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}^\{2\}\}\\\|\\bm\{v\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},t\)\\\|^\{2\}\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}t\+C−2∫01∫𝒳2⟨𝒗𝜽\(𝒙,t\),𝒖1t\(𝒙,𝒚\)⟩πt\(𝒙,𝒚\)d𝒙d𝒚dt\\displaystyle\-2\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}^\{2\}\}\\langle\\bm\{v\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},t\),\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\rangle\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}t=∫01∫𝒳2‖𝒗𝜽​\(𝒙,t\)‖2​πt​\(𝒙,𝒚\)​𝑑𝒙​𝑑𝒚​𝑑t\+C\\displaystyle=\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}^\{2\}\}\\\|\\bm\{v\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},t\)\\\|^\{2\}\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}t\+C−2∫01∫𝒳⟨𝒗𝜽\(𝒙,t\),∫𝒳𝒖1t\(𝒙,𝒚\)πt​\(𝒙,𝒚\)ρt​\(𝒙\)d𝒚⟩ρt\(𝒙\)d𝒙dt\\displaystyle\-2\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}\}\\langle\\bm\{v\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},t\),\\int\_\{\\mathcal\{X\}\}\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\frac\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\}\{\\rho\_\{t\}\(\\bm\{x\}\)\}\\mathrm\{d\}\\bm\{y\}\\rangle\\rho\_\{t\}\(\\bm\{x\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}t=∫01∫𝒳2‖𝒗𝜽​\(𝒙,t\)‖2​πt​\(𝒙,𝒚\)​𝑑𝒙​𝑑𝒚​𝑑t\+C−2​∫01∫𝒳⟨𝒗𝜽​\(𝒙,t\),𝒗t​\(𝒙\)⟩​ρt​\(𝒙\)​𝑑𝒙​𝑑t\\displaystyle=\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}^\{2\}\}\\\|\\bm\{v\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},t\)\\\|^\{2\}\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}t\+C\-2\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}\}\\langle\\bm\{v\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},t\),\\bm\{v\}\_\{t\}\(\\bm\{x\}\)\\rangle\\rho\_\{t\}\(\\bm\{x\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}t=∫01∫𝒳‖𝒗𝜽​\(𝒙,t\)‖2​ρt​\(𝒙\)​𝑑𝒙​𝑑t−2​∫01∫𝒳⟨𝒗𝜽​\(𝒙,t\),𝒗t​\(𝒙\)⟩​ρt​\(𝒙\)​𝑑𝒙​𝑑t\\displaystyle=\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}\}\\\|\\bm\{v\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},t\)\\\|^\{2\}\\rho\_\{t\}\(\\bm\{x\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}t\-2\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}\}\\langle\\bm\{v\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},t\),\\bm\{v\}\_\{t\}\(\\bm\{x\}\)\\rangle\\rho\_\{t\}\(\\bm\{x\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}t\+∫01∫𝒳∥𝒗t\(𝒙\)∥2ρt\(𝒙\)d𝒙dt\+C\\displaystyle\+\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}\}\\\|\\bm\{v\}\_\{t\}\(\\bm\{x\}\)\\\|^\{2\}\\rho\_\{t\}\(\\bm\{x\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}t\+C=∫01∫𝒳‖𝒗𝜽​\(𝒙,t\)−𝒗t​\(𝒙\)‖2​ρt​\(𝒙\)​𝑑𝒙​𝑑t\+C\\displaystyle=\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}\}\\\|\\bm\{v\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},t\)\-\\bm\{v\}\_\{t\}\(\\bm\{x\}\)\\\|^\{2\}\\rho\_\{t\}\(\\bm\{x\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}t\+C=ℒM​\(𝜽\)\+C\\displaystyle=\\mathcal\{L\}\_\{M\}\(\\bm\{\\theta\}\)\+CwhereCCis some constant independent to𝜽\\bm\{\\theta\}\. Therefore,∇𝜽ℒM​\(𝜽\)=∇𝜽ℒC​\(𝜽\)\\nabla\_\{\\bm\{\\theta\}\}\\mathcal\{L\}\_\{M\}\(\\bm\{\\theta\}\)=\\nabla\_\{\\bm\{\\theta\}\}\\mathcal\{L\}\_\{C\}\(\\bm\{\\theta\}\)\. ∎

### A\.10Proof of Theorem[5\.6](https://arxiv.org/html/2609.20008#S5.Thmtheorem6)

Theorem[5\.6](https://arxiv.org/html/2609.20008#S5.Thmtheorem6)\.The minimizer of \([23](https://arxiv.org/html/2609.20008#S5.E23)\) is𝐟t​\(𝐱,𝐲\)\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\.

###### Proof\.

We first recall the loss\.

ℒf​\(ϕ\)\\displaystyle\\mathcal\{L\}\_\{f\}\(\\bm\{\\phi\}\)=𝔼t∼𝒰\[0,1\],\(𝒛,𝒛′\)∼q\(𝒛,𝒛′\),\(𝒙,𝒚\)∼πt\(𝒙,𝒚\|𝒛,𝒛′\)‖𝒇ϕ\(𝒙,𝒚,t\)−𝒇t\(𝒙,𝒚\|𝒛,𝒛′\)‖22\\displaystyle=\\mathbb\{E\}\_\{t\\sim\\mathcal\{U\}\[0,1\],\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\sim q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\),\(\\bm\{x\},\\bm\{y\}\)\\sim\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\}\\left\\\|\\bm\{\\bm\{f\}\_\{\\phi\}\}\(\\bm\{x\},\\bm\{y\},t\)\-\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\right\\\|\_\{2\}^\{2\}\(73\)Similarly to theorem[5\.4](https://arxiv.org/html/2609.20008#S5.Thmtheorem4), we have

ℒf​\(ϕ\)\\displaystyle\\mathcal\{L\}\_\{f\}\(\\bm\{\\phi\}\)=𝔼t∼𝒰⁡\[0,1\],\(𝒙,𝒚\)∼πt​\(𝒙,𝒚\)​‖𝒇ϕ​\(𝒙,𝒚,t\)−𝒇t​\(𝒙,𝒚\)‖22\+C\\displaystyle=\\mathbb\{E\}\_\{t\\sim\\mathcal\{U\}\[0,1\],\(\\bm\{x\},\\bm\{y\}\)\\sim\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\}\\left\\\|\\bm\{\\bm\{f\}\_\{\\phi\}\}\(\\bm\{x\},\\bm\{y\},t\)\-\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\right\\\|\_\{2\}^\{2\}\+C\(74\)whereCCis a constant independent toϕ\\bm\{\\phi\}\. Therefore, the minimizer is𝒇t​\(𝒙,𝒚\)\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\. ∎

## Appendix BImplementation details

### B\.1Computation resources

All experiments were performed on a personal computer with NVIDIA 5070 Ti GPU\.

### B\.2Network architecture

The velocity net𝒗𝜽\\bm\{v\}\_\{\\bm\{\\theta\}\}was parameterized as a MLP with 128 hidden channels, 5 hidden layers and SiLU activations[Elfwing et al\. \(2018\)](https://arxiv.org/html/2609.20008#bib.bib39)\. It receives the scalar timettconcatenated with expression and spatial state\(𝒙,𝒔\)\(\\bm\{x\},\\bm\{s\}\), and outputs the velocity on both expression and spatial space\. Across all experiments, the networks were optimized by Adam\([Kingma and Ba, 2017](https://arxiv.org/html/2609.20008#bib.bib29)\)with learning rate2×10−32\\times 10^\{\-3\}, weight decay10−510^\{\-5\}, and gradient norm clipping at 5\.

### B\.3Lagrangian choice for experiments

As introduced in[4\.2](https://arxiv.org/html/2609.20008#S4.SS2), we used a modality separated Lagrangian for numerical experiments\. On both synthetic data and real data, we only added interaction term on spatial coordinates, and used standard kinetic energy on expression spaces\. That is, let two cells represented by\(𝒙,𝒔\),\(𝒚,𝒉\)\(\\bm\{x\},\\bm\{s\}\),\(\\bm\{y\},\\bm\{h\}\)where𝒙,𝒚\\bm\{x\},\\bm\{y\}are expressions,𝒔,𝒉\\bm\{s\},\\bm\{h\}are spatial coordinates, we choose the Lagrangian as

ℒ=\\displaystyle\\mathcal\{L\}=ηgene​\(12​‖𝒙˙‖2\+12​‖𝒚˙‖2\)\+λgene​\|dd​t​‖𝒙−𝒚‖\|2\\displaystyle\\eta\_\{\\text\{gene\}\}\\left\(\\frac\{1\}\{2\}\\\|\\dot\{\\bm\{x\}\}\\\|^\{2\}\+\\frac\{1\}\{2\}\\\|\\dot\{\\bm\{y\}\}\\\|^\{2\}\\right\)\+\\lambda\_\{\\text\{gene\}\}\|\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\\|\\bm\{x\}\-\\bm\{y\}\\\|\|^\{2\}\(75\)\+\\displaystyle\+ηspatial​\(12​‖𝒔˙‖2\+12​‖𝒉˙‖2\)\+λspatial​\|dd​t​‖𝒔−𝒉‖\|2\\displaystyle\\eta\_\{\\text\{spatial\}\}\\left\(\\frac\{1\}\{2\}\\\|\\dot\{\\bm\{s\}\}\\\|^\{2\}\+\\frac\{1\}\{2\}\\\|\\dot\{\\bm\{h\}\}\\\|^\{2\}\\right\)\+\\lambda\_\{\\text\{spatial\}\}\|\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\\|\\bm\{s\}\-\\bm\{h\}\\\|\|^\{2\}
Since the corresponding QOT problem remains the same under global scaling, we setηgene=1\\eta\_\{\\text\{gene\}\}=1\. The interaction term is only added to spatial Lagrangian, henceλgene=0\\lambda\_\{\\text\{gene\}\}=0\. The Lagrangian reduces to

ℒ=\\displaystyle\\mathcal\{L\}=12​‖𝒙˙‖2\+12​‖𝒚˙‖2\+ηspatial​\(12​‖𝒔˙‖2\+12​‖𝒉˙‖2\)\+λspatial​\|dd​t​‖𝒔−𝒉‖\|2\\displaystyle\\frac\{1\}\{2\}\\\|\\dot\{\\bm\{x\}\}\\\|^\{2\}\+\\frac\{1\}\{2\}\\\|\\dot\{\\bm\{y\}\}\\\|^\{2\}\+\\eta\_\{\\text\{spatial\}\}\\left\(\\frac\{1\}\{2\}\\\|\\dot\{\\bm\{s\}\}\\\|^\{2\}\+\\frac\{1\}\{2\}\\\|\\dot\{\\bm\{h\}\}\\\|^\{2\}\\right\)\+\\lambda\_\{\\text\{spatial\}\}\|\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\\|\\bm\{s\}\-\\bm\{h\}\\\|\|^\{2\}\(76\)
There are only two hyperparametersηspatial,λspatial\\eta\_\{\\text\{spatial\}\},\\lambda\_\{\\text\{spatial\}\}\. We conducted an ablation study in[C\.1](https://arxiv.org/html/2609.20008#A3.SS1)\.

### B\.4Synthetic rotation data details

In[6](https://arxiv.org/html/2609.20008#S6), we evaluated TP\-DATE on two synthetic rotation datasets\. Here, we introduce the generation details\. Each synthetic rotation dataset contains 2,000 paired cells from three cell types, with two spatial coordinates and 50 expression features\. The target time point is a90∘90^\{\\circ\}rotation of the source time point\. Therefore, the true midpoint is the corresponding45∘45^\{\\circ\}rotation\.

The source spatial coordinates were sampled from three noisy curved arms\. For a cell of typecc, we sampled

r∼𝒰\(0\.18,1\),ϵθ∼𝒩\(0,0\.0552\),ϵr,ϵq∼𝒩\(0,0\.0352\),r\\sim\\mathcal\{U\}\(0\.18,1\),\\qquad\\epsilon\_\{\\theta\}\\sim\\mathcal\{N\}\(0,0\.055^\{2\}\),\\qquad\\epsilon\_\{r\},\\epsilon\_\{q\}\\sim\\mathcal\{N\}\(0,0\.035^\{2\}\),and set

θ=θc\+1\.05​r\+ϵθ,𝒔0=\(0\.35\+2\.25​r\+ϵr\)​\[cos⁡θsin⁡θ\]\+ϵq​\[−sin⁡θcos⁡θ\],\\theta=\\theta\_\{c\}\+1\.05r\+\\epsilon\_\{\\theta\},\\qquad\\bm\{s\}\_\{0\}=\(0\.35\+2\.25r\+\\epsilon\_\{r\}\)\\begin\{bmatrix\}\\cos\\theta\\\\ \\sin\\theta\\end\{bmatrix\}\+\\epsilon\_\{q\}\\begin\{bmatrix\}\-\\sin\\theta\\\\ \\cos\\theta\\end\{bmatrix\},where\(θ0,θ1,θ2\)=\(15∘,140∘,260∘\)\(\\theta\_\{0\},\\theta\_\{1\},\\theta\_\{2\}\)=\(15^\{\\circ\},140^\{\\circ\},260^\{\\circ\}\)\.

After concatenating the three cell types, we randomly permuted the cells, subtracted the global spatial centroid, and divided all coordinates byn−1​∑i‖𝒔0,i‖22\\sqrt\{n^\{\-1\}\\sum\_\{i\}\\\|\\bm\{s\}\_\{0,i\}\\\|\_\{2\}^\{2\}\}\. The spatial trajectory was a rigid counterclockwise rotation,

𝒔t,i=R⁡\(90∘​t\)​𝒔0,i,t∈\[0,1\],\\bm\{s\}\_\{t,i\}=R\(90^\{\\circ\}t\)\\bm\{s\}\_\{0,i\},\\qquad t\\in\[0,1\],whereR⁡\(⋅\)R\(\\cdot\)is the two\-dimensional rotation matrix\. Consequently, the observed source, held out midpoint, and target correspond to rotations of0∘0^\{\\circ\},45∘45^\{\\circ\}, and90∘90^\{\\circ\}, respectively, and all pairwise spatial distances are exactly preserved\.

Source expression was initialized independently sampled from𝒩⁡\(0,0\.182\)\\mathcal\{N\}\(0,0\.18^\{2\}\)\. To generate cell\-type markers, we added2\.22\.2to genes−121\\\!\-\\\!12,−2413\\\!\-\\\!24, and−3625\\\!\-\\\!36for cell types 1, 2, and 3, respectively\. Genes−5037\\\!\-\\\!50additionally received0\.750\.75times the following 14 dimensional spatial signals, evaluated using the normalized source coordinates\(xi,yi\)\(x\_\{i\},y\_\{i\}\), radiusρi=xi2\+yi2\\rho\_\{i\}=\\sqrt\{x\_\{i\}^\{2\}\+y\_\{i\}^\{2\}\}, and angleϕi=atan2​yixi\\phi\_\{i\}=\\text\{atan2\}\\frac\{y\_\{i\}\}\{x\_\{i\}\}:

\(xi,yi,ρi,sin⁡ϕi,cos⁡ϕi,xi​yi,xi2−yi2,sin⁡2​ϕi,cos⁡2​ϕi,OPENρi​sin⁡\(ϕi\+0\.7\),ρi​cos⁡\(ϕi−0\.4\),exp⁡\(−0\.7​ρi2\),tanh⁡\(1\.5​xi\),tanh⁡\(1\.5​yi\)\)\.\\begin\{split\}\(&x\_\{i\},y\_\{i\},\\rho\_\{i\},\\sin\\phi\_\{i\},\\cos\\phi\_\{i\},x\_\{i\}y\_\{i\},x\_\{i\}^\{2\}\-y\_\{i\}^\{2\},\\sin 2\\phi\_\{i\},\\cos 2\\phi\_\{i\},\\\\ &\\rho\_\{i\}\\sin\(\\phi\_\{i\}\+0\.7\),\\rho\_\{i\}\\cos\(\\phi\_\{i\}\-0\.4\),\\exp\(\-0\.7\\rho\_\{i\}^\{2\}\),\\tanh\(1\.5x\_\{i\}\),\\tanh\(1\.5y\_\{i\}\)\)\.\\end\{split\}Each expression feature was then centered and divided by its population standard deviation\. For Rotation, expression remained unchanged over time\. For Distractor, we constructed a drift matrix𝚫\\bm\{\\Delta\}by adding0\.450\.45to genes−4137\\\!\-\\\!41,−4642\\\!\-\\\!46, and−5047\\\!\-\\\!50for cell types 1, 2, and 3, respectively, followed by independent noise sampled from𝒩⁡\(0,0\.162\)\\mathcal\{N\}\(0,0\.16^\{2\}\)added for each cell, all 50 features\. Denoted the expression matrix at timettas𝒆t\\bm\{e\}\_\{t\}, we set

𝒆0\.5=𝒆0\+12​𝚫,𝒆1=𝒆0\+𝚫\.\\bm\{e\}\_\{0\.5\}=\\bm\{e\}\_\{0\}\+\\tfrac\{1\}\{2\}\\bm\{\\Delta\},\\qquad\\bm\{e\}\_\{1\}=\\bm\{e\}\_\{0\}\+\\bm\{\\Delta\}\.Finally, each feature was standardized using its mean and population standard deviation computed jointly over the concatenated source, midpoint, and target expression matrices\. Therefore, the Distractor dataset contains a cell type dependent temporal expression shift, while its spatial dynamics remain the same rigid rotation as in Rotation\.

For the TP\-DATE experiments in the main text, the hyperparameter was set to\(ηspatial,λspatial\)=\(1,16\)\.\(\\eta\_\{\\mathrm\{spatial\}\},\\lambda\_\{\\mathrm\{spatial\}\}\)=\(1,16\)\.

### B\.5Real data details

#### B\.5\.1Developing Mouse brain

The developing mouse brain data was adopted from\([Chen et al\., 2022](https://arxiv.org/html/2609.20008#bib.bib4)\)\. This dataset contains spatial transcriptomics sections from mouse embryonic development at eight different time points\. We selected brain cells according to the cell annotations provided with the dataset and performed Leiden clustering\([Traag et al\., 2019](https://arxiv.org/html/2609.20008#bib.bib10)\)on gene expression space for visualization purpose only\. We use the snapshots from day 14\.5, 15\.5, and 16\.5, denoted E14\.5, E15\.5, and E16\.5 with 17591, 17031, 17296 cells, respectively\. We reduced the dimension by first selecting 1500 highly variable genes, then PCA to 50D\. The spatial coordinates were linearly scaled to\[−1,1\]\[\-1,1\]\. The hold one out experiment was trained on E14\.5, E16\.5, and evaluated on E15\.5\. The hyperparameters are set to\(ηspatial,λspatial\)=\(100,0\.01\)\(\\eta\_\{\\mathrm\{spatial\}\},\\lambda\_\{\\mathrm\{spatial\}\}\)=\(100,0\.01\)\. An overview of the data is shown in figure[2](https://arxiv.org/html/2609.20008#A2.F2)\. Different colors stand for different Leiden clusters\.

![Refer to caption](https://arxiv.org/html/2609.20008v1/figures/Mouse_Brain.png)Figure 2:Mouse Brain overview
#### B\.5\.2ARTISTA

The ARTISTA data was adopted from\([Wei et al\., 2022](https://arxiv.org/html/2609.20008#bib.bib3)\)\. This dataset describes the brain regeneration process in axolotl following injury and contains spatial transcriptomics sections collected at seven different time points from Day 2 to Day 60 after injury\. In the main text, we use snapshots at Day 2, Day 5, and Day 10, and subsample 7,500 cells per snapshot\. The expression vectors were also dimension reduced to 50D by PCA\. The spatial coordinates were linearly scaled to\[−1,1\]\[\-1,1\]\. For hold one out experiment, we trained on Day 2 and Day 10, and evaluated on Day 5\. The hyperparameters are set to\(ηspatial,λspatial\)=\(100,0\.01\)\(\\eta\_\{\\mathrm\{spatial\}\},\\lambda\_\{\\mathrm\{spatial\}\}\)=\(100,0\.01\)\. The normalized evaluation time is therefore0\.3750\.375instead of0\.50\.5\. An overview of the data is shown in figure[3](https://arxiv.org/html/2609.20008#A2.F3)\. Different colors stand for Niche annotations provided with the dataset\.

![Refer to caption](https://arxiv.org/html/2609.20008v1/figures/ARTISTA.png)Figure 3:ARTISTA overview
#### B\.5\.3Tumor

The Tumor data was adopted from\([Zhang et al\., 2026a](https://arxiv.org/html/2609.20008#bib.bib5)\)\. This dataset contains five spatial transcriptomics sections of the same MC38 tumor collected at different depths \(subQ\-1 to subQ\-5\)\. We use subQ\-3, subQ\-4, and subQ\-5 with 10,000 spatially sampled cells per snapshot\. The expression vectors are also reduced to 50D via PCA\. The spatial coordinates were linearly scaled to\[−1,1\]\[\-1,1\]\. For hold one out experiments, models are trained on subQ\-3 and subQ\-5, and evaluated on subQ\-4\. The hyperparameters are set to\(ηspatial,λspatial\)=\(0\.01,1\)\(\\eta\_\{\\mathrm\{spatial\}\},\\lambda\_\{\\mathrm\{spatial\}\}\)=\(0\.01,1\)\. An overview of the data is shown in figure[4](https://arxiv.org/html/2609.20008#A2.F4)\. Different colors stand for Niche annotations provided with the dataset\.

![Refer to caption](https://arxiv.org/html/2609.20008v1/figures/Tumor.png)Figure 4:Tumor overview

### B\.6Evaluation metrics

In the hold one out experiments on real data, we train the model on the two endpoint time points, infer the distribution at the intermediate time point, and compare the prediction with the ground truth\. In the spatial transcriptomics setting, each cell is jointly characterized by its gene expression𝒙\\bm\{x\}and spatial coordinate𝒔\\bm\{s\}, so that each snapshot can be regarded as a joint distributionp⁡\(𝒙,𝒔\)p\(\\bm\{x\},\\bm\{s\}\)\. At the intermediate time point, we evaluate the discrepancy between the ground truth distributionppand the predicted distributionp^\\hat\{p\}\.

#### B\.6\.1Marginal Wassersteins

To measure the discrepancy between two joint distribution, one can first measure the discrepancy between the marginal distributions\. Letpx​\(𝒙\)=∫p⁡\(𝒙,𝒔\)​𝑑𝒔,ps​\(𝒔\)=∫p⁡\(𝒙,𝒔\)​𝑑𝒙p\_\{x\}\(\\bm\{x\}\)=\\int p\(\\bm\{x\},\\bm\{s\}\)\\mathrm\{d\}\\bm\{s\},p\_\{s\}\(\\bm\{s\}\)=\\int p\(\\bm\{x\},\\bm\{s\}\)\\mathrm\{d\}\\bm\{x\}be the marginals on expression space and spatial space, the spatial Wasserstein distance and expression Wasserstein distance are defined as

Spatial​W22​\(p,p^\)=infγ∈Π⁡\(ps,p^s\)∫‖𝒔−𝒔′‖2​γ​\(𝒔,𝒔′\)​𝑑𝒔​d​𝒔′\\displaystyle\\text\{Spatial \}W\_\{2\}^\{2\}\(p,\\hat\{p\}\)=\\inf\_\{\\gamma\\in\\Pi\(p\_\{s\},\\hat\{p\}\_\{s\}\)\}\\int\\\|\\bm\{s\}\-\\bm\{s\}^\{\\prime\}\\\|^\{2\}\\gamma\(\\bm\{s\},\\bm\{s\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{s\}\\mathrm\{d\}\\bm\{s\}^\{\\prime\}\(77\)Expression​W22​\(p,p^\)=infγ∈Π⁡\(px,p^x\)∫‖𝒙−𝒙′‖2​γ​\(𝒙,𝒙′\)​𝑑𝒙​d​𝒙′\\displaystyle\\text\{Expression \}W\_\{2\}^\{2\}\(p,\\hat\{p\}\)=\\inf\_\{\\gamma\\in\\Pi\(p\_\{x\},\\hat\{p\}\_\{x\}\)\}\\int\\\|\\bm\{x\}\-\\bm\{x\}^\{\\prime\}\\\|^\{2\}\\gamma\(\\bm\{x\},\\bm\{x\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{x\}^\{\\prime\}Intuitively, these marginal discrepancies measure how well can the model reconstruct the spatial or expression information solely\. To further measure how well can the model reconstruct the spatial and expression information jointly, we introduce fused Wasserstein distance and spatially coupled expression MSE \(SC\-eMSE\)\.

#### B\.6\.2Fused Wasserstein

FusedW22W\_\{2\}^\{2\}is defined as

Fused​W22​\(p,p^\)=infγ∈Π⁡\(p,p^\)∫‖𝒛−𝒛′‖2​γ​\(𝒛,𝒛′\)​𝑑𝒛​d​𝒛′\\text\{Fused \}W\_\{2\}^\{2\}\(p,\\hat\{p\}\)=\\inf\_\{\\gamma\\in\\Pi\(p,\\hat\{p\}\)\}\\int\\\|\\bm\{z\}\-\\bm\{z\}^\{\\prime\}\\\|^\{2\}\\gamma\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}\(78\)where𝒛=\(𝒙,α​𝒔\)\\bm\{z\}=\(\\bm\{x\},\\alpha\\bm\{s\}\)\. The expression vector𝒙\\bm\{x\}was concatenated with a scaled spatial vectorα​𝒔\\alpha\\bm\{s\}, and the standard Wasserstein distance is calculated based on the concatenated vector\.α\\alphawas introduced to balance the importance of𝒙\\bm\{x\}and𝒔\\bm\{s\}, since they may have different dimension and scale\. For all datasets, we chooseα\\alphawith a common rule

α=∑d=1DxStd⁡\(𝒙d\)∑d=1DsStd⁡\(𝒔d\)\.\\alpha=\\frac\{\\sum\_\{d=1\}^\{D\_\{x\}\}\\operatorname\{Std\}\(\\bm\{x\}^\{d\}\)\}\{\\sum\_\{d=1\}^\{D\_\{s\}\}\\operatorname\{Std\}\(\\bm\{s\}^\{d\}\)\}\.\(79\)whereDx,DsD\_\{x\},D\_\{s\}are the dimensions of expression space and spatial space,𝒙d,𝒔d\\bm\{x\}^\{d\},\\bm\{s\}^\{d\}are thedd\-th components of𝒙,𝒔\\bm\{x\},\\bm\{s\}\. Intuitively, this selection ofα\\alphabalances the total standard deviation of expression and space\. In hold one out experiments, the standard deviations are calculated only on the training time points\.

#### B\.6\.3SC\-eMSE

To test whether expression is reconstructed at the correct spatial location, we additionally report spatially coupled expression MSE \(SC\-eMSE\)\. Letγ\\gammabe the optimal coupling betweenρ,ρ^\\rho,\\hat\{\\rho\}, obtained using spatial coordinates𝒔\\bm\{s\}alone with the spatial cost‖𝒔−𝒔′‖2\\\|\\bm\{s\}\-\\bm\{s\}^\{\\prime\}\\\|^\{2\}, we define

SC​\-​eMSE⁡\(ρ,ρ^\)=1Dx​∫‖𝒙−𝒙′‖2​γ​\(𝒛,𝒛′\)​𝑑𝒛​d​𝒛′\\operatorname\{SC\\text\{\-\}eMSE\}\(\\rho,\\hat\{\\rho\}\)=\\frac\{1\}\{D\_\{x\}\}\\int\\\|\\bm\{x\}\-\\bm\{x\}^\{\\prime\}\\\|^\{2\}\\gamma\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}where𝒛=\(𝒙,𝒔\)\\bm\{z\}=\(\\bm\{x\},\\bm\{s\}\)\. The spatial couplingγ\\gammais fitted without expression information, while the transport cost‖𝒙−𝒙′‖2\\\|\\bm\{x\}\-\\bm\{x\}^\{\\prime\}\\\|^\{2\}is calculated without spatial information\. Intuitively, SC\-eMSE measures whether the predicted snapshot shares similar spatial expression patterns with the ground truth\.

#### B\.6\.4Toydata metrics

On rotation toys, the ground truth dynamics are known rigid rotations\. Therefore, we can directly calculate the spatial MSE to measure the spatial level accuracy, instead of calculating distribution\-level Wasserstein distance\. At the unseen time pointtt, let the ground truth spatial states be𝒔⋆\\bm\{s\}^\{\\star\}, the predicted spatial states be𝒔^\\hat\{\\bm\{s\}\}\. The spatial MSE is defined as

MSEs=1n​∑i=1n‖𝒔^i−𝒔i⋆‖2\.\\operatorname\{MSE\}\_\{s\}=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\\|\\hat\{\\bm\{s\}\}\_\{i\}\-\\bm\{s\}\_\{i\}^\{\\star\}\\\|^\{2\}\.\(80\)
We also calculated the pair distortion to measure how well can the predicted dynamics preserve the spatial structure\. We uniformly sampledM=6000M=6000pairs𝒫=\{\(ir,kr\)\}r=1M\\mathcal\{P\}=\\\{\(i\_\{r\},k\_\{r\}\)\\\}\_\{r=1\}^\{M\}whereit≠kri\_\{t\}\\neq k\_\{r\}\. The pair distortion is defined as

PD=1M​∑r=1M\|‖s^ir−s^kr‖2−‖sir⋆−skr⋆‖2\|‖sir⋆−skr⋆‖2\.\\operatorname\{PD\}=\\frac\{1\}\{M\}\\sum\_\{r=1\}^\{M\}\\frac\{\\left\|\\\|\\hat\{s\}\_\{i\_\{r\}\}\-\\hat\{s\}\_\{k\_\{r\}\}\\\|\_\{2\}\-\\\|s\_\{i\_\{r\}\}^\{\\star\}\-s\_\{k\_\{r\}\}^\{\\star\}\\\|\_\{2\}\\right\|\}\{\\\|s\_\{i\_\{r\}\}^\{\\star\}\-s\_\{k\_\{r\}\}^\{\\star\}\\\|\_\{2\}\}\.\(81\)Intuitively, it measures the average relative error of pairwise spatial distance\. A lower PD means the predicted dynamics preserves the spatial structure better\. In practice,‖sir⋆−skr⋆‖2\\\|s\_\{i\_\{r\}\}^\{\\star\}\-s\_\{k\_\{r\}\}^\{\\star\}\\\|\_\{2\}can be easily calculated at time00

‖sir⋆−skr⋆‖2=‖sir0−skr0‖2\\\|s\_\{i\_\{r\}\}^\{\\star\}\-s\_\{k\_\{r\}\}^\{\\star\}\\\|\_\{2\}=\\\|s\_\{i\_\{r\}\}^\{0\}\-s\_\{k\_\{r\}\}^\{0\}\\\|\_\{2\}since the true dynamics are rigid transformations\. Note that, the ground truth pair distortions are zero on toy datasets, so lower values are meaningful here\. But it is not the case of real datasets where nonzero biological deformation may be correct, and the held out cells have no source\-cell correspondence\. Therefore, we only report pair distortions on toy datasets\.

### B\.7GW\-CFM and FGW\-CFM baselines

In the main text, we compared TP\-DATE with GW\-CFM and FGW\-CFM\. These baselines are introduced as a naive dynamic extension of static GW\-OT and FGW\-OT\. We used GW\-OT coupling or FGW\-OT coupling, and the displacement conditional path𝒙t=\(1−t\)​𝒙0\+t​𝒙1\\bm\{x\}\_\{t\}=\(1\-t\)\\bm\{x\}\_\{0\}\+t\\bm\{x\}\_\{1\}for conditional flow matching training\. That is, for two probability densitiesμ0,μ1\\mu\_\{0\},\\mu\_\{1\}, we first calculate the GW\-OT coupling or FGW\-OT couplingγ⁡\(𝒙0,𝒙1\)\\gamma\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\}\)\. Next, we parameterize a velocity network𝒗𝜽​\(𝒙,t\)\\bm\{v\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},t\), and train it by minimizing the conditional flow matching loss

\\displaystyleℒCFM​\(𝜽\)=𝔼t∼𝒰⁡\[0,1\],\(𝒙0,𝒙1\)∼γ⁡\(𝒙0,𝒙1\),𝒙=\(1−t\)​𝒙0\+t​𝒙1​‖𝒗𝜽​\(𝒙,t\)−\(𝒙1−𝒙0\)‖22\.\\displaystyle\\mathcal\{L\}\_\{\\text\{CFM\}\}\(\\bm\{\\theta\}\)=\\mathbb\{E\}\_\{t\\sim\\mathcal\{U\}\[0,1\],\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\}\)\\sim\\gamma\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\}\),\\bm\{x\}=\(1\-t\)\\bm\{x\}\_\{0\}\+t\\bm\{x\}\_\{1\}\}\\left\\\|\\bm\{\\bm\{v\}\_\{\\theta\}\}\(\\bm\{x\},t\)\-\(\\bm\{x\}\_\{1\}\-\\bm\{x\}\_\{0\}\)\\right\\\|\_\{2\}^\{2\}\.\(82\)It can be viewed as replacing the OT coupling in OT\-CFM by GW/FGW\-OT coupling\. Intuitively, it is a kind of linear interpolation betweenμ0,μ1\\mu\_\{0\},\\mu\_\{1\}induced by the corresponding coupling, therefore a natural choice for dynamic extension\. For FGW\-OT, we fused OT and GW\-OT with weight7:37:3\.

### B\.8Static QOT solver

The static QOT problem is

QOTS​\(μ0,μ1\)=infγ∈Π⁡\(μ0,μ1\)∫𝒳4𝒜⁡\(𝒛,𝒛′\)​γ​\(𝒛\)​γ​\(𝒛′\)​d𝒛​d​𝒛′\.\\displaystyle\\text\{QOT\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\inf\_\{\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)\}\\int\_\{\\mathcal\{X\}^\{4\}\}\\mathcal\{A\}\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\gamma\(\\bm\{z\}\)\\gamma\(\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}\.\(83\)In practice,μ0,μ1\\mu\_\{0\},\\mu\_\{1\}are represented by point cloud data\{Xi\}i=1:M,\{Yj\}j=1:N\\\{X\_\{i\}\\\}\_\{i=1:M\},\\\{Y\_\{j\}\\\}\_\{j=1:N\}\. The conditional variables𝒛=\(Xi,Yj\),𝒛′=\(Xk,Yl\)\\bm\{z\}=\(X\_\{i\},Y\_\{j\}\),\\bm\{z\}^\{\\prime\}=\(X\_\{k\},Y\_\{l\}\)are two end point pairs\. The static cost in \([83](https://arxiv.org/html/2609.20008#A2.E83)\) can be written as

Q⁡\(Γ\)=∑Ai​j,k​l​Γi​j​Γk​lQ\(\\Gamma\)=\\sum A\_\{ij,kl\}\\Gamma\_\{ij\}\\Gamma\_\{kl\}\(84\)whereΓ∈\{Γ∈ℝ≥0M×N\|Γ𝟏=1M𝟏,𝟏TΓ=1N𝟏T\}\\Gamma\\in\\\{\\Gamma\\in\\mathbb\{R\}\_\{\\geq 0\}^\{M\\times N\}\|\\Gamma\\bm\{1\}=\\frac\{1\}\{M\}\\bm\{1\},\\bm\{1\}^\{\\mathrm\{T\}\}\\Gamma=\\frac\{1\}\{N\}\\bm\{1\}^\{\\mathrm\{T\}\}\\\}is the empirical coupling,Ai​j,k​l=A⁡\(Xi,Yj,Xk,Yl\)A\_\{ij,kl\}=A\(X\_\{i\},Y\_\{j\},X\_\{k\},Y\_\{l\}\)is the path action\. The static QOT problem is then formulated as

min⁡∑Γ⁡Ai​j,k​l​Γi​j​Γk​l\\min\_\{\\Gamma\}\\sum A\_\{ij,kl\}\\Gamma\_\{ij\}\\Gamma\_\{kl\}\(85\)which is a quadratic programming w\.r\.tΓ\\Gamma\.

#### B\.8\.1Frank\-Wolfe algorithm

The Frank\-Wolfe algorithm\([Kerdoncuff et al\., 2021](https://arxiv.org/html/2609.20008#bib.bib8)\)is an iterative first\-order optimization algorithm for convex constraint optimization\. Though QOT problem may not be convex, it is standard and widely used in solving Gromov\-Wasserstein type static QOT problem\([Flamary et al\., 2021](https://arxiv.org/html/2609.20008#bib.bib30)\)\. Consider a convex setDDand an optimization problem

min𝒙∈D⁡f⁡\(𝒙\)\\min\_\{\\bm\{x\}\\in D\}f\(\\bm\{x\}\)\(86\)The Frank\-Wolfe algorithm solves it by iteratively doing

Initilization𝒙0∈D,k=0\\displaystyle\\text\{Initilization\}\\quad\\bm\{x\}\_\{0\}\\in D,\\quad k=0\(87\)Step1\.𝒔k=arg⁡min𝒔∈D𝒔T∇f\(𝒙k\)\\displaystyle\\text\{Step1\.\}\\quad\\bm\{s\}\_\{k\}=\\mathop\{\\arg\\min\}\_\{\\bm\{s\}\\in D\}\\bm\{s\}^\{\\mathrm\{T\}\}\\nabla f\(\\bm\{x\}\_\{k\}\)Step2\.𝒙k\+1=𝒙k\+2k\+2​\(𝒔k−𝒙k\),k=k\+1\\displaystyle\\text\{Step2\.\}\\quad\\bm\{x\}\_\{k\+1\}=\\bm\{x\}\_\{k\}\+\\frac\{2\}\{k\+2\}\(\\bm\{s\}\_\{k\}\-\\bm\{x\}\_\{k\}\),\\quad k=k\+1In our case, with symmetric conditions such asℒ\\mathcal\{L\}is symmetric to\(𝒙,𝒙˙\),\(𝒚,𝒚˙\)\(\\bm\{x\},\\dot\{\\bm\{x\}\}\),\(\\bm\{y\},\\dot\{\\bm\{y\}\}\), it is easy to checkAi​j,k​l=Ak​l,i​jA\_\{ij,kl\}=A\_\{kl,ij\}\. Therefore,∇Q​\(Γ\)=2​∑Ai​j,k​l​Γk​l\\nabla Q\(\\Gamma\)=2\\sum A\_\{ij,kl\}\\Gamma\_\{kl\}\. The Frank\-Wolfe algorithm \([87](https://arxiv.org/html/2609.20008#A2.E87)\) can be realized as

InitilizationΓ0,n=0\\displaystyle\\text\{Initilization\}\\quad\\Gamma^\{0\},\\quad n=0\(88\)Step1\.Gi​jn=2​∑Ai​j,k​l​Γk​ln\\displaystyle\\text\{Step1\.\}\\quad G\_\{ij\}^\{n\}=2\\sum A\_\{ij,kl\}\\Gamma\_\{kl\}^\{n\}Step2\.Πn=arg⁡minΠ⁡⟨Π,G⟩\\displaystyle\\text\{Step2\.\}\\quad\\Pi^\{n\}=\\mathop\{\\arg\\min\}\_\{\\Pi\}\\langle\\Pi,G\\rangleStep3\.Γn\+1=Γn\+2n\+2​\(Πn−Γn\),n=n\+1\\displaystyle\\text\{Step3\.\}\\quad\\Gamma^\{n\+1\}=\\Gamma^\{n\}\+\\frac\{2\}\{n\+2\}\(\\Pi^\{n\}\-\\Gamma^\{n\}\),\\quad n=n\+1In step 1, we estimate the gradientGG\. In step 2, we solve a standard optimal transport subproblem\. In step 3, we updateΓ\\Gamma\. The OT subproblem is relaxed and solved by Sinkhorn\.

#### B\.8\.2Gradient estimation

The full computation of step 1 is𝒪⁡\(M2​N2\)\\mathcal\{O\}\(M^\{2\}N^\{2\}\)which is expansive\. To reduce the computational cost, we estimate the gradient via Monte Carlo\.

G^i​j=2R​∑h=1RAi​j,kh​lh,\(kh,lh\)​∼i\.i\.d\.​Γ\\hat\{G\}\_\{ij\}=\\frac\{2\}\{R\}\\sum\_\{h=1\}^\{R\}A\_\{ij,k\_\{h\}l\_\{h\}\},\\quad\(k\_\{h\},l\_\{h\}\)\\overset\{i\.i\.d\.\}\{\\sim\}\\Gamma\(89\)The computational complexity is𝒪⁡\(R​M​N\)\\mathcal\{O\}\(RMN\)whereRRis the number of Monte Carlo samples\.

#### B\.8\.3Sparse KNN

If𝒪⁡\(R​M​N\)\\mathcal\{O\}\(RMN\)is still too expansive, we can further construct a bidirectional KNN graph between\{Xi\}i=1:M,\{Yj\}j=1:N\\\{X\_\{i\}\\\}\_\{i=1:M\},\\\{Y\_\{j\}\\\}\_\{j=1:N\}and restrictΓ\\Gammaon the graph\. Intuitively,XiX\_\{i\}is more likely to transport to someYjY\_\{j\}near to itself rather than far away from itself\. Therefore, we construct two KNN graphs on gene expression space\. EachXiX\_\{i\}is linked to itsKKnearestYjY\_\{j\}, and eachYjY\_\{j\}is also linked to itsKKnearestXiX\_\{i\}\. We take the union of the edges, denotedE0E\_\{0\}\. That is,\(i,j\)∈E0\(i,j\)\\in E\_\{0\}meansXiX\_\{i\}is one of theKKnearest neighbors ofYjY\_\{j\}, orYjY\_\{j\}is one of theKKnearest neighbors ofXiX\_\{i\}\. It is easy to check\|E0\|≤K⁡\(M\+N\)\|E\_\{0\}\|\\leq K\(M\+N\)\. We use bidirectional KNN rather than one way KNN so that all points have neighbors\. We then want to restrict the couplingΓ\\GammaonE0E\_\{0\}, that is

Γ∈ℝ≥0M×N,Γ​𝟏=1M​𝟏,𝟏T​Γ=1N​𝟏T,Γi​j=0,\(i,j\)∉E0\.\\Gamma\\in\\mathbb\{R\}\_\{\\geq 0\}^\{M\\times N\},\\quad\\Gamma\\bm\{1\}=\\frac\{1\}\{M\}\\bm\{1\},\\quad\\quad\\bm\{1\}^\{\\mathrm\{T\}\}\\Gamma=\\frac\{1\}\{N\}\\bm\{1\}^\{\\mathrm\{T\}\},\\quad\\Gamma\_\{ij\}=0,\(i,j\)\\notin E\_\{0\}\.\(90\)Although the bidirectional construction guarantees that every row and column has at least one admissible edge, this alone does not necessarily guarantee the existence of a coupling with the prescribed marginals\. We therefore augmentE0E\_\{0\}by a small number of additional edges to guarantee feasibility\. Specifically, initialize residual masses

ri=1M,sj=1N,r\_\{i\}=\\frac\{1\}\{M\},\\qquad s\_\{j\}=\\frac\{1\}\{N\},and an auxiliary couplingΓ¯=0\\bar\{\\Gamma\}=0\. We first scan all edges\(i,j\)∈E0\(i,j\)\\in E\_\{0\}in dictionary order\. For each edge, we assign

δi​j=min⁡\{ri,sj\},Γ¯i​j←Γ¯i​j\+δi​j,\\delta\_\{ij\}=\\min\\\{r\_\{i\},s\_\{j\}\\\},\\qquad\\bar\{\\Gamma\}\_\{ij\}\\leftarrow\\bar\{\\Gamma\}\_\{ij\}\+\\delta\_\{ij\},and updateri←ri−δi​jr\_\{i\}\\leftarrow r\_\{i\}\-\\delta\_\{ij\},sj←sj−δi​js\_\{j\}\\leftarrow s\_\{j\}\-\\delta\_\{ij\}\. After all KNN edges have been scanned, if residual masses remain, we repeatedly select a pairi,ji,jwithri\>0r\_\{i\}\>0andsj\>0s\_\{j\}\>0, add the edge\(i,j\)\(i,j\), and assign

δ=min⁡\{ri,sj\}\.\\delta=\\min\\\{r\_\{i\},s\_\{j\}\\\}\.Each added edge exhausts at least one residual row or column\. Therefore, at mostM\+N−1M\+N\-1additional edges are required\. Denoting these repair edges byErepE\_\{\\mathrm\{rep\}\}, we finally set

E=E0∪Erep\.E=E\_\{0\}\\cup E\_\{\\mathrm\{rep\}\}\.By construction,Γ¯\\bar\{\\Gamma\}satisfies

Γ¯​𝟏=1M​𝟏,𝟏T​Γ¯=1N​𝟏T,supp⁡\(Γ¯\)⊆E,\\bar\{\\Gamma\}\\bm\{1\}=\\frac\{1\}\{M\}\\bm\{1\},\\qquad\\bm\{1\}^\{\\mathrm\{T\}\}\\bar\{\\Gamma\}=\\frac\{1\}\{N\}\\bm\{1\}^\{\\mathrm\{T\}\},\\qquad\\operatorname\{supp\}\(\\bar\{\\Gamma\}\)\\subseteq E,which explicitly guarantees that the sparse coupling constraint is feasible\. Moreover,

\|E\|≤K⁡\(M\+N\)\+M\+N−1,\|E\|\\leq K\(M\+N\)\+M\+N\-1,and the feasibility repair requires only𝒪⁡\(K⁡\(M\+N\)\)\\mathcal\{O\}\(K\(M\+N\)\)additional computation\. We then restrict the couplingΓ\\Gammaon the augmented sparse supportEE, that is

Γ∈ℝ≥0M×N,Γ​𝟏=1M​𝟏,𝟏T​Γ=1N​𝟏T,Γi​j=0,\(i,j\)∉E\.\\Gamma\\in\\mathbb\{R\}\_\{\\geq 0\}^\{M\\times N\},\\quad\\Gamma\\bm\{1\}=\\frac\{1\}\{M\}\\bm\{1\},\\quad\\bm\{1\}^\{\\mathrm\{T\}\}\\Gamma=\\frac\{1\}\{N\}\\bm\{1\}^\{\\mathrm\{T\}\},\\quad\\Gamma\_\{ij\}=0,\\ \(i,j\)\\notin E\.\(91\)Mathematically, solving the OT subproblem under constraint \([91](https://arxiv.org/html/2609.20008#A2.E91)\) is equivalent to solving a standard OT problem with a masked cost

\{G~i​j=Gi​j,\(i,j\)∈EG~i​j=\+∞,\(i,j\)∉E\\left\\\{\\begin\{aligned\} &\\tilde\{G\}\_\{ij\}=G\_\{ij\},\\quad\(i,j\)\\in E\\\\ &\\tilde\{G\}\_\{ij\}=\+\\infty,\\quad\(i,j\)\\notin E\\end\{aligned\}\\right\.\(92\)In Sinkhorn iteration, infinite cost leads to zero element in Gibbs kernel

\{Ki​j=exp\(−G~i​j/τ\)=exp\(−Gi​j/τ\),\(i,j\)∈EKi​j=exp\(−G~i​j/τ\)=0,\(i,j\)∉E\\left\\\{\\begin\{aligned\} &K\_\{ij\}=\\exp\(\-\\tilde\{G\}\_\{ij\}/\\tau\)=\\exp\(\-G\_\{ij\}/\\tau\),\\quad\(i,j\)\\in E\\\\ &K\_\{ij\}=\\exp\(\-\\tilde\{G\}\_\{ij\}/\\tau\)=0,\\quad\(i,j\)\\notin E\\\\ \\end\{aligned\}\\right\.\(93\)whereτ\\tauis the coefficient of entropy relaxation\. Further, a zero element in Gibbs kernel leads to zero element in updated coupling

Γi​jnew=ui​Ki​j​vj=0,\(i,j\)∉E\\Gamma^\{\\text\{new\}\}\_\{ij\}=u\_\{i\}K\_\{ij\}v\_\{j\}=0,\\quad\(i,j\)\\notin E\(94\)whereu,vu,vare dual variables\. The analysis above shows that we can only estimateGi​jG\_\{ij\}and updateΓi​j\\Gamma\_\{ij\}for\(i,j\)∈E\(i,j\)\\in E\. For\(i,j\)∉E\(i,j\)\\notin E,Ki​j,Γi​jK\_\{ij\},\\Gamma\_\{ij\}are automatically zero\. The complexity of gradient estimation is therefore further reduced to𝒪⁡\(R​K​\(M\+N\)\)\\mathcal\{O\}\(RK\(M\+N\)\)\. We conducted an ablation study forKKin[C\.1](https://arxiv.org/html/2609.20008#A3.SS1)\.

## Appendix CAdditional results

### C\.1Ablation studies

#### C\.1\.1Loss weights

To demonstrate the robustness of TP\-DATE, we conducted an ablation study for hyperparameters\(ηspatial,λspatial\)\(\\eta\_\{\\text\{spatial\}\},\\lambda\_\{\\text\{spatial\}\}\)[B\.3](https://arxiv.org/html/2609.20008#A2.SS3)on ARTISTA\. As shown in Table[4](https://arxiv.org/html/2609.20008#A3.T4), TP\-DATE remains robust under different hyperparameter settings\.

Table 4:TP\-DATE hyperparameter sensitivity on ARTISTA\.
#### C\.1\.2Sparse KNN

We also conducted an ablation study forKKin sparse KNN construction on the synthetic rotation data \(Distractor\) and the Tumor dataset\. On the Distractor dataset, we testedK=64,128,256,512,1024,K=64,128,256,512,1024,and20002000\. Since both time points in Distractor contain 2,000 cells,K=2000K=2000corresponds to using the full problem without the sparse KNN approximation\. As shown in Table[5](https://arxiv.org/html/2609.20008#A3.T5), the performance is exactly identical across different values ofKKon this relatively simple synthetic dataset \(exceptKK=2000\)\. Interestingly, the method performs slightly worse when the sparse KNN approximation is not used\. This is because, in this simple setting, the mass of the correct coupling is already highly concentrated on the sparse KNN edges, so restricting the optimization to these edges can actually make the coupling problem easier to solve and lead to a more accurate solution\. To rule out the possibility that the observed results were due to the dataset being insufficiently complex, we further testedK=64,128,256,512,1024,K=64,128,256,512,1024,on the Tumor dataset\. The corresponding model performance are approximately the same, demonstrating the robustness of sparse KNN technique \(Table[6](https://arxiv.org/html/2609.20008#A3.T6)\)\.

Table 5:KKablation study on Distractor dataset\. The dagger marks the value used in the main experiments\.Table 6:KKablation study on Tumor dataset\. The dagger marks the value used in the main experiments\.We also recorded the training time and memory usage under different values ofKKon two datasets\. As shown in Figs\.[5](https://arxiv.org/html/2609.20008#A3.F5)and[6](https://arxiv.org/html/2609.20008#A3.F6), on a log–log scale, the slopes of the linear regressions for both training time and GPU memory usage with respect toKKare close to11, indicating that, as expected, the computational cost of TP\-DATE scales approximately linearly withKK\. CPU memory usage does not exhibit the same linear dependence, mainly because CPU memory is substantially affected by data loading, preprocessing, and other fixed overheads, which dominate the algorithmic memory cost whenKKis relatively small\.

![Refer to caption](https://arxiv.org/html/2609.20008v1/figures/KNN_Distractor.png)Figure 5:Training time and memory usage for differentKKon Distractor\.![Refer to caption](https://arxiv.org/html/2609.20008v1/figures/KNN_Tumor.png)Figure 6:Training time and memory usage for differentKKon Tumor\.
#### C\.1\.3Conditional path

In the Mouse Brain and ARTISTA experiments, we used relatively small weights for the interaction term\. To isolate the contribution of the interaction\-induced conditional path from that of the static QOT coupling, we trained an additional baseline using exactly the same QOT coupling as TP\-DATE but replacing the travelling\-pair conditional path with standard displacement interpolation\. All other training settings and random seeds were kept unchanged\. As shown in Table[7](https://arxiv.org/html/2609.20008#A3.T7), TP\-DATE achieves lower mean errors across all evaluation metrics on both datasets, with particularly clear improvements in expression reconstruction, SC\-eMSE, and fusedW2W\_\{2\}\. These results indicate that the performance gain cannot be attributed solely to the static QOT coupling\. Even with a relatively small interaction weight, the interaction induced dynamic conditional path provides an additional and consistent benefit\. We do not additionally consider the zero\-interaction coupling as a separate ablation, since removing the interaction term reduces the induced static QOT to the corresponding weighted standard OT formulation, whose flow\-matching counterpart is already represented by OT\-CFM\.

Table 7:Conditional path ablation on spatial transcriptomics hold one out experiment\. We report mean and standard deviation over 5 random seeds\.

### C\.2Scalability

We trained TP\-DATE and simulation\-based baselines \(stVCR\([Peng et al\., 2026a](https://arxiv.org/html/2609.20008#bib.bib12)\), CytoBridge\([Zhang et al\., 2025b](https://arxiv.org/html/2609.20008#bib.bib13)\)\) on different cell numbers\. We downsampled the subQ\-3, subQ\-5 slices of the Tumor dataset\([Zhang et al\., 2026a](https://arxiv.org/html/2609.20008#bib.bib5)\)to 10,000, 30,000, 50,000, 70,000, and 90,000 cells, respectively, and recorded the training time and memory usage of the three methods on datasets of the corresponding sizes\. As shown in Figure[7](https://arxiv.org/html/2609.20008#A3.F7), CytoBridge ran out of memory on 90000, stVCR and TP\-DATE both require moderate memory\. As shown in Figure[8](https://arxiv.org/html/2609.20008#A3.F8), TP\-DATE exhibits computational cost that scales linearly with dataset size and is more efficient than the two simulation\-based baselines\. This highlights the computational efficiency of simulation\-free approaches such as flow matching\.

![Refer to caption](https://arxiv.org/html/2609.20008v1/figures/CPU.png)

![Refer to caption](https://arxiv.org/html/2609.20008v1/figures/GPU.png)

Figure 7:Training memory on different size data![Refer to caption](https://arxiv.org/html/2609.20008v1/figures/Time.png)Figure 8:Training time on different size data

## Appendix DRelations to other works

### D\.1GW\-OT and FGW\-OT

Consider the Lagrangian

ℒ⁡\(t,𝒙,𝒚,𝒙˙,𝒚˙\)=12​‖𝒙˙‖2\+12​‖𝒚˙‖2\+λ​\|dd​t​‖𝒙t−𝒚t‖\|2\.\\mathcal\{L\}\(t,\\bm\{x\},\\bm\{y\},\\dot\{\\bm\{x\}\},\\dot\{\\bm\{y\}\}\)=\\frac\{1\}\{2\}\\\|\\dot\{\\bm\{x\}\}\\\|^\{2\}\+\\frac\{1\}\{2\}\\\|\\dot\{\\bm\{y\}\}\\\|^\{2\}\+\\lambda\|\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\\|\\bm\{x\}\_\{t\}\-\\bm\{y\}\_\{t\}\\\|\|^\{2\}\.\(95\)
#### D\.1\.1Standard static OT

Whenλ=0\\lambda=0, the static cost is

𝒜⁡\(𝒙0,𝒙1,𝒚0,𝒚1\)=inf𝒙t,𝒚t∫0112​‖𝒙˙t‖2\+12​‖𝒚˙t‖2​𝑑t=12​\(‖𝒙1−𝒙0‖2\+‖𝒚1−𝒚0‖2\)\.\\mathcal\{A\}\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\},\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\)=\\inf\_\{\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\}\}\\int\_\{0\}^\{1\}\\frac\{1\}\{2\}\\\|\\dot\{\\bm\{x\}\}\_\{t\}\\\|^\{2\}\+\\frac\{1\}\{2\}\\\|\\dot\{\\bm\{y\}\}\_\{t\}\\\|^\{2\}\\mathrm\{d\}t=\\frac\{1\}\{2\}\\Big\(\\\|\\bm\{x\}\_\{1\}\-\\bm\{x\}\_\{0\}\\\|^\{2\}\+\\\|\\bm\{y\}\_\{1\}\-\\bm\{y\}\_\{0\}\\\|^\{2\}\\Big\)\.\(96\)The corresponding static form is

QOTS​\(μ0,μ1\)\\displaystyle\\text\{QOT\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)=infγ∈Π⁡\(μ0,μ1\)∫𝒳4𝒜⁡\(𝒙0,𝒙1,𝒚0,𝒚1\)​γ​\(𝒙0,𝒙1\)​γ​\(𝒚0,𝒚1\)​d​𝒙0​d​𝒙1​d​𝒚0​d​𝒚1\\displaystyle=\\inf\_\{\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)\}\\int\_\{\\mathcal\{X\}^\{4\}\}\\mathcal\{A\}\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\},\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\)\\gamma\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\}\)\\gamma\(\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\)\\mathrm\{d\}\\bm\{x\}\_\{0\}\\mathrm\{d\}\\bm\{x\}\_\{1\}\\mathrm\{d\}\\bm\{y\}\_\{0\}\\mathrm\{d\}\\bm\{y\}\_\{1\}\(97\)=infγ∈Π⁡\(μ0,μ1\)12​\(∫𝒳2‖𝒙1−𝒙0‖2​γ​\(𝒙0,𝒙1\)​d​𝒙0​d​𝒙1\+∫𝒳2‖𝒚1−𝒚0‖2​γ​\(𝒚0,𝒚1\)​d​𝒚0​d​𝒚1\)\\displaystyle=\\inf\_\{\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)\}\\frac\{1\}\{2\}\\Big\(\\int\_\{\\mathcal\{X\}^\{2\}\}\\\|\\bm\{x\}\_\{1\}\-\\bm\{x\}\_\{0\}\\\|^\{2\}\\gamma\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\}\)\\mathrm\{d\}\\bm\{x\}\_\{0\}\\mathrm\{d\}\\bm\{x\}\_\{1\}\+\\int\_\{\\mathcal\{X\}^\{2\}\}\\\|\\bm\{y\}\_\{1\}\-\\bm\{y\}\_\{0\}\\\|^\{2\}\\gamma\(\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\)\\mathrm\{d\}\\bm\{y\}\_\{0\}\\mathrm\{d\}\\bm\{y\}\_\{1\}\\Big\)=infγ∈Π⁡\(μ0,μ1\)∫𝒳2‖𝒙1−𝒙0‖2​γ​\(𝒙0,𝒙1\)​d​𝒙0​d​𝒙1\\displaystyle=\\inf\_\{\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)\}\\int\_\{\\mathcal\{X\}^\{2\}\}\\\|\\bm\{x\}\_\{1\}\-\\bm\{x\}\_\{0\}\\\|^\{2\}\\gamma\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\}\)\\mathrm\{d\}\\bm\{x\}\_\{0\}\\mathrm\{d\}\\bm\{x\}\_\{1\}=OT​\(μ0,μ1\)\.\\displaystyle=\\text\{OT\}\(\\mu\_\{0\},\\mu\_\{1\}\)\.It is equivalent to the standard static OT\.

#### D\.1\.2Standard GW\-OT

When there are no kinetic terms, which isℒ=\|dd​t​‖𝒙t−𝒚t‖\|2\\mathcal\{L\}=\|\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\\|\\bm\{x\}\_\{t\}\-\\bm\{y\}\_\{t\}\\\|\|^\{2\}, letrt=‖𝒙t−𝒚t‖r\_\{t\}=\\\|\\bm\{x\}\_\{t\}\-\\bm\{y\}\_\{t\}\\\|, the static cost is

𝒜⁡\(𝒙0,𝒙1,𝒚0,𝒚1\)=infrt∫01\|r˙t\|2​𝑑ts\.t\.r0=‖𝒙0−𝒚0‖,r1=‖𝒙1−𝒚1‖\.\\mathcal\{A\}\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\},\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\)=\\inf\_\{r\_\{t\}\}\\int\_\{0\}^\{1\}\|\\dot\{r\}\_\{t\}\|^\{2\}\\mathrm\{d\}t\\quad\\text\{s\.t\.\}\\quad r\_\{0\}=\\\|\\bm\{x\}\_\{0\}\-\\bm\{y\}\_\{0\}\\\|,\\quad r\_\{1\}=\\\|\\bm\{x\}\_\{1\}\-\\bm\{y\}\_\{1\}\\\|\.\(98\)Using the Euler\-Lagrange equation, it is easy to show that whend≥2d\\geq 2, the optimalrtr\_\{t\}is attained byrt=\(1−t\)​r0\+t​r1r\_\{t\}=\(1\-t\)r\_\{0\}\+tr\_\{1\}\. Therefore, the path action is

𝒜⁡\(𝒙0,𝒙1,𝒚0,𝒚1\)=\(r1−r0\)2=\(‖𝒙1−𝒚1‖−‖𝒙0−𝒚0‖\)2\.\\mathcal\{A\}\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\},\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\)=\(r\_\{1\}\-r\_\{0\}\)^\{2\}=\(\\\|\\bm\{x\}\_\{1\}\-\\bm\{y\}\_\{1\}\\\|\-\\\|\\bm\{x\}\_\{0\}\-\\bm\{y\}\_\{0\}\\\|\)^\{2\}\.\(99\)The corresponding static form is

QOTS​\(μ0,μ1\)\\displaystyle\\text\{QOT\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)=infγ∈Π⁡\(μ0,μ1\)∫𝒳4𝒜⁡\(𝒙0,𝒙1,𝒚0,𝒚1\)​γ​\(𝒙0,𝒙1\)​γ​\(𝒚0,𝒚1\)​d​𝒙0​d​𝒙1​d​𝒚0​d​𝒚1\\displaystyle=\\inf\_\{\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)\}\\int\_\{\\mathcal\{X\}^\{4\}\}\\mathcal\{A\}\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\},\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\)\\gamma\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\}\)\\gamma\(\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\)\\mathrm\{d\}\\bm\{x\}\_\{0\}\\mathrm\{d\}\\bm\{x\}\_\{1\}\\mathrm\{d\}\\bm\{y\}\_\{0\}\\mathrm\{d\}\\bm\{y\}\_\{1\}\(100\)=infγ∈Π⁡\(μ0,μ1\)∫𝒳4\(‖𝒙1−𝒚1‖−‖𝒙0−𝒚0‖\)2​γ​\(𝒙0,𝒙1\)​γ​\(𝒚0,𝒚1\)​d​𝒙0​d​𝒙1​d​𝒚0​d​𝒚1\\displaystyle=\\inf\_\{\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)\}\\int\_\{\\mathcal\{X\}^\{4\}\}\(\\\|\\bm\{x\}\_\{1\}\-\\bm\{y\}\_\{1\}\\\|\-\\\|\\bm\{x\}\_\{0\}\-\\bm\{y\}\_\{0\}\\\|\)^\{2\}\\gamma\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\}\)\\gamma\(\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\)\\mathrm\{d\}\\bm\{x\}\_\{0\}\\mathrm\{d\}\\bm\{x\}\_\{1\}\\mathrm\{d\}\\bm\{y\}\_\{0\}\\mathrm\{d\}\\bm\{y\}\_\{1\}=GWOT​\(μ0,μ1\)\\displaystyle=\\text\{GWOT\}\(\\mu\_\{0\},\\mu\_\{1\}\)which is equivalent to standard GW\-OT\([Mémoli, 2011](https://arxiv.org/html/2609.20008#bib.bib35);[Klein et al\., 2025](https://arxiv.org/html/2609.20008#bib.bib32)\)\. We point out that different to static OT which has a dynamic form, GW\-OT does not naturally have an unique dynamic form\. Because the static cost \([98](https://arxiv.org/html/2609.20008#A4.E98)\) is defined by only minimizing w\.r\.trtr\_\{t\}, which can not uniquely determine𝒒t\\bm\{q\}\_\{t\}\. Intuitively, the travelling pair path action only penalizes the variation of the length of𝒒t\\bm\{q\}\_\{t\}, but doesn’t penalize the rotation\.

#### D\.1\.3Dynamic analog to FGW\-OT

Previous studies have combined the standard static OT and GW\-OT together to obtain a fused GW\-OT \(FGW\-OT\)\([Vayer et al\., 2020](https://arxiv.org/html/2609.20008#bib.bib9);[Klein et al\., 2025](https://arxiv.org/html/2609.20008#bib.bib32)\), which simultaneously minimizes the kinetic energy and the preserves the local structure\. Its cost is defined as a weighted sum of static OT cost and static GW\-OT cost\.

CFGW​\(μ0,μ1\)=α​CGW​\(μ0,μ1\)\+\(1−α\)​COT​\(μ0,μ1\)C\_\{\\text\{FGW\}\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\alpha C\_\{\\text\{GW\}\}\(\\mu\_\{0\},\\mu\_\{1\}\)\+\(1\-\\alpha\)C\_\{\\text\{OT\}\}\(\\mu\_\{0\},\\mu\_\{1\}\)\(101\)As discussed above, the Lagrangian \([95](https://arxiv.org/html/2609.20008#A4.E95)\) is also a weighted sum of OT and GW\-OT up to a scalar, but in a dynamic sense\. Hence, intuitively, the corresponding dynamic form can be viewed as a dynamic analog to FGW\-OT\.

As discussed in[A\.3](https://arxiv.org/html/2609.20008#A1.SS3), the Lagrangian \([95](https://arxiv.org/html/2609.20008#A4.E95)\) can be written as‖𝒄˙t‖2\+14​rt2​θ˙t2\+\(14\+λ\)​r˙t2\\\|\\dot\{\\bm\{c\}\}\_\{t\}\\\|^\{2\}\+\\frac\{1\}\{4\}r\_\{t\}^\{2\}\\dot\{\\theta\}\_\{t\}^\{2\}\+\(\\frac\{1\}\{4\}\+\\lambda\)\\dot\{r\}\_\{t\}^\{2\}under the C\-Q decomposition and the polar coordinates on the plane spanned by𝒒0,𝒒1\\bm\{q\}\_\{0\},\\bm\{q\}\_\{1\}\. The first term is the kinetic energy of the barycenter, which is an analog to the standard kinetic term12​‖𝒙˙‖2\+12​‖𝒚˙‖2\\frac\{1\}\{2\}\\\|\\dot\{\\bm\{x\}\}\\\|^\{2\}\+\\frac\{1\}\{2\}\\\|\\dot\{\\bm\{y\}\}\\\|^\{2\}\. The third term exactly leads to the static path action of GW\-OT \([98](https://arxiv.org/html/2609.20008#A4.E98)\)\. The second term further penalizes the variation ofθt\\theta\_\{t\}, i\.e\. the rotation of𝒒t\\bm\{q\}\_\{t\}\. Therefore, the Lagrangian \([95](https://arxiv.org/html/2609.20008#A4.E95)\) actually leads to a more strict preservation of local structure than FGW\-OT\.

We also derived the analytic form of the path action in[A\.3](https://arxiv.org/html/2609.20008#A1.SS3)\.

𝒜⁡\(𝒙0,𝒙1,𝒚0,𝒚1\)=‖𝒄1−𝒄0‖2\+\(14\+λ\)​\(r02\+r12−2​r0​r1​cos⁡k​α\)\\mathcal\{A\}\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\},\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\)=\\\|\\bm\{c\}\_\{1\}\-\\bm\{c\}\_\{0\}\\\|^\{2\}\+\(\\frac\{1\}\{4\}\+\\lambda\)\(r\_\{0\}^\{2\}\+r\_\{1\}^\{2\}\-2r\_\{0\}r\_\{1\}\\cos k\\alpha\)\(102\)Whenλ=0\\lambda=0,

𝒜⁡\(𝒙0,𝒙1,𝒚0,𝒚1\)\\displaystyle\\mathcal\{A\}\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\},\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\)=‖𝒄1−𝒄0‖2\+14​\(r02\+r12−2​r0​r1​cos⁡α\)\\displaystyle=\\\|\\bm\{c\}\_\{1\}\-\\bm\{c\}\_\{0\}\\\|^\{2\}\+\\frac\{1\}\{4\}\(r\_\{0\}^\{2\}\+r\_\{1\}^\{2\}\-2r\_\{0\}r\_\{1\}\\cos\\alpha\)\(103\)=‖𝒄1−𝒄0‖2\+14​‖𝒒1−𝒒0‖2=12​‖𝒙1−𝒙0‖2\+12​‖𝒚1−𝒚0‖2\.\\displaystyle=\\\|\\bm\{c\}\_\{1\}\-\\bm\{c\}\_\{0\}\\\|^\{2\}\+\\frac\{1\}\{4\}\\\|\\bm\{q\}\_\{1\}\-\\bm\{q\}\_\{0\}\\\|^\{2\}=\\frac\{1\}\{2\}\\\|\\bm\{x\}\_\{1\}\-\\bm\{x\}\_\{0\}\\\|^\{2\}\+\\frac\{1\}\{2\}\\\|\\bm\{y\}\_\{1\}\-\\bm\{y\}\_\{0\}\\\|^\{2\}\.It reduces to OT\. Whenλ→∞\\lambda\\to\\infty,

λ−1​𝒜​\(𝒙0,𝒙1,𝒚0,𝒚1\)\\displaystyle\\lambda^\{\-1\}\\mathcal\{A\}\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\},\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\)=𝒪⁡\(λ−1\)\+\(r02\+r12−2​r0​r1​cos⁡11\+4​λ​α\)\\displaystyle=\\mathcal\{O\}\(\\lambda^\{\-1\}\)\+\(r\_\{0\}^\{2\}\+r\_\{1\}^\{2\}\-2r\_\{0\}r\_\{1\}\\cos\\sqrt\{\\frac\{1\}\{1\+4\\lambda\}\}\\alpha\)\(104\)=𝒪⁡\(λ−1\)\+\(r02\+r12−2​r0​r1​\(1\+𝒪⁡\(λ−1\)\)\)\\displaystyle=\\mathcal\{O\}\(\\lambda^\{\-1\}\)\+\(r\_\{0\}^\{2\}\+r\_\{1\}^\{2\}\-2r\_\{0\}r\_\{1\}\(1\+\\mathcal\{O\}\(\\lambda^\{\-1\}\)\)\)=𝒪⁡\(λ−1\)\+\(r02\+r12−2​r0​r1\)\\displaystyle=\\mathcal\{O\}\(\\lambda^\{\-1\}\)\+\(r\_\{0\}^\{2\}\+r\_\{1\}^\{2\}\-2r\_\{0\}r\_\{1\}\)=\(r1−r0\)2\+𝒪⁡\(λ−1\)→\(‖𝒙1−𝒚1‖−‖𝒙0−𝒚0‖\)2\\displaystyle=\(r\_\{1\}\-r\_\{0\}\)^\{2\}\+\\mathcal\{O\}\(\\lambda^\{\-1\}\)\\to\(\\\|\\bm\{x\}\_\{1\}\-\\bm\{y\}\_\{1\}\\\|\-\\\|\\bm\{x\}\_\{0\}\-\\bm\{y\}\_\{0\}\\\|\)^\{2\}It reduces to GW\-OT\. Therefore, for generalλ∈\(0,\+∞\)\\lambda\\in\(0,\+\\infty\), it can be viewed as a dynamic fusion of OT and GW\-OT rather than FGW\-OT, which is just a static fusion\.

#### D\.1\.4The 1D case

We needd≥2d\\geq 2\(ddis the dimension of the space\) to have \([98](https://arxiv.org/html/2609.20008#A4.E98)\) be equivalent to the static GW\-OT cost\. To be self\-contained, we briefly discuss the 1D case\(d=1\)\(d=1\)here\. Whend=1d=1, the Lagrangianℒ=\(dd​t​\|xt−yt\|\)2=\(dd​t​\(xt−yt\)\)2\\mathcal\{L\}=\(\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\|x\_\{t\}\-y\_\{t\}\|\)^\{2\}=\\Big\(\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\(x\_\{t\}\-y\_\{t\}\)\\Big\)^\{2\}\. Letq1=x1−y1,q0=x0−y0q\_\{1\}=x\_\{1\}\-y\_\{1\},q\_\{0\}=x\_\{0\}\-y\_\{0\}, the path action is

𝒜⁡\(x0,x1,y0,y1\)=infqt∫01\|q˙t\|2​𝑑ts\.t\.q0=x0−y0,q1=x1−y1\.\\mathcal\{A\}\(x\_\{0\},x\_\{1\},y\_\{0\},y\_\{1\}\)=\\inf\_\{q\_\{t\}\}\\int\_\{0\}^\{1\}\|\\dot\{q\}\_\{t\}\|^\{2\}\\mathrm\{d\}t\\quad\\text\{s\.t\.\}\\quad q\_\{0\}=x\_\{0\}\-y\_\{0\},\\quad q\_\{1\}=x\_\{1\}\-y\_\{1\}\.\(105\)The optimal path isqt=\(1−t\)​q0\+t​q1q\_\{t\}=\(1\-t\)q\_\{0\}\+tq\_\{1\}, instead ofrt=\(1−t\)​rt\+t​r1r\_\{t\}=\(1\-t\)r\_\{t\}\+tr\_\{1\}\. The path action is therefore

𝒜⁡\(x0,x1,y0,y1\)=\(q1−q0\)2=\(x1−y1−x0\+y0\)2\\mathcal\{A\}\(x\_\{0\},x\_\{1\},y\_\{0\},y\_\{1\}\)=\(q\_\{1\}\-q\_\{0\}\)^\{2\}=\(x\_\{1\}\-y\_\{1\}\-x\_\{0\}\+y\_\{0\}\)^\{2\}\(106\)instead of\(\|x1−y1\|−\|x0−y0\|\)2\(\|x\_\{1\}\-y\_\{1\}\|\-\|x\_\{0\}\-y\_\{0\}\|\)^\{2\}\. Thus, whend=1d=1, our path action does not reduce to the static cost of GW\-OT\. In fact, the two cost are equal if and only if\(x1−y1\)​\(x0−y0\)≥0\(x\_\{1\}\-y\_\{1\}\)\(x\_\{0\}\-y\_\{0\}\)\\geq 0\. Geometrically, in a high dimensional space \(d≥2d\\geq 2\), one can continuously deform the configuration and reverse the relative positions of two points while keeping the magnitude of their relative displacement equal to the linear interpolation between its endpoint values\. But it is impossible in one dimension\. Reversing the relative positions of two points necessarily forces their relative displacement to vanish at some intermediate time, so its magnitude cannot follow the linear interpolation between the initial and final values\.

### D\.2IGW\-OT

Developing dynamic formulations of GW\-like OT has long attracted considerable attention\. A recent successful attempt is the inner product GW\-OT \(IGW\-OT\), proposed by\([Zhang et al\., 2026b](https://arxiv.org/html/2609.20008#bib.bib7)\), which replaces the cost based on discrepancies between pairwise distances in GW\-OT with one based on discrepancies between inner products\.

IGW\-OTS2​\(μ0,μ1\)=infγ∈Π⁡\(μ0,μ1\)∫𝒳4\(⟨𝒙1,𝒚1⟩−⟨𝒙0,𝒚0⟩\)2​γ​\(𝒙0,𝒙1\)​γ​\(𝒚0,𝒚1\)​d​𝒙0​d​𝒙1​d​𝒚0​d​𝒚1\.\\text\{IGW\-OT\}^\{2\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\inf\_\{\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)\}\\int\_\{\\mathcal\{X\}^\{4\}\}\(\\langle\\bm\{x\}\_\{1\},\\bm\{y\}\_\{1\}\\rangle\-\\langle\\bm\{x\}\_\{0\},\\bm\{y\}\_\{0\}\\rangle\)^\{2\}\\gamma\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\}\)\\gamma\(\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\)\\mathrm\{d\}\\bm\{x\}\_\{0\}\\mathrm\{d\}\\bm\{x\}\_\{1\}\\mathrm\{d\}\\bm\{y\}\_\{0\}\\mathrm\{d\}\\bm\{y\}\_\{1\}\.\(107\)
They developed a dynamic formulation for the above static QOT problem, denoted as IGW\-OT\. Intuitively, IGW\-OT likewise aims to preserve spatial structure as much as possible throughout the transport process\. Different to TP\-DATE, they focused on the gradient flow and Riemannian structure of IGW\-OT\.

We point out that, with an appropriate choice of Lagrangian within the TP\-DATE framework, IGW\-OT can also be recovered as a special case of the dynamic QOT formulation introduced in our work\. Take

ℒ⁡\(t,𝒙t,𝒚t,𝒙˙t,𝒚˙t\)=\|dd​t​⟨𝒙t,𝒚t⟩\|2=\|⟨𝒙˙t,𝒚t⟩\+⟨𝒙t,𝒚˙t⟩\|2\\mathcal\{L\}\(t,\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\},\\dot\{\\bm\{x\}\}\_\{t\},\\dot\{\\bm\{y\}\}\_\{t\}\)=\|\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\langle\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\}\\rangle\|^\{2\}=\|\\langle\\dot\{\\bm\{x\}\}\_\{t\},\\bm\{y\}\_\{t\}\\rangle\+\\langle\\bm\{x\}\_\{t\},\\dot\{\\bm\{y\}\}\_\{t\}\\rangle\|^\{2\}\(108\)The path action is

𝒜⁡\(𝒙0,𝒙1,𝒚0,𝒚1\)=infrt∫01\|r˙t\|2​𝑑ts\.t\.r0=⟨𝒙0,𝒚0⟩,r1=⟨𝒙1,𝒚1⟩\.\\mathcal\{A\}\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\},\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\)=\\inf\_\{r\_\{t\}\}\\int\_\{0\}^\{1\}\|\\dot\{r\}\_\{t\}\|^\{2\}\\mathrm\{d\}t\\quad\\text\{s\.t\.\}\\quad r\_\{0\}=\\langle\\bm\{x\}\_\{0\},\\bm\{y\}\_\{0\}\\rangle,\\quad r\_\{1\}=\\langle\\bm\{x\}\_\{1\},\\bm\{y\}\_\{1\}\\rangle\.\(109\)Using the Euler\-Lagrange equation, it is easy to show that whend≥2d\\geq 2, the optimalrtr\_\{t\}is attained byrt=\(1−t\)​r0\+t​r1r\_\{t\}=\(1\-t\)r\_\{0\}\+tr\_\{1\}\. Therefore, the path action is

𝒜⁡\(𝒙0,𝒙1,𝒚0,𝒚1\)=\(r1−r0\)2=\(⟨𝒙1,𝒚1⟩−⟨𝒙0,𝒚0⟩\)2\.\\mathcal\{A\}\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\},\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\)=\(r\_\{1\}\-r\_\{0\}\)^\{2\}=\(\\langle\\bm\{x\}\_\{1\},\\bm\{y\}\_\{1\}\\rangle\-\\langle\\bm\{x\}\_\{0\},\\bm\{y\}\_\{0\}\\rangle\)^\{2\}\.\(110\)The corresponding static form is

QOTS​\(μ0,μ1\)\\displaystyle\\text\{QOT\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)=infγ∈Π⁡\(μ0,μ1\)∫𝒳4𝒜⁡\(𝒙0,𝒙1,𝒚0,𝒚1\)​γ​\(𝒙0,𝒙1\)​γ​\(𝒚0,𝒚1\)​d​𝒙0​d​𝒙1​d​𝒚0​d​𝒚1\\displaystyle=\\inf\_\{\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)\}\\int\_\{\\mathcal\{X\}^\{4\}\}\\mathcal\{A\}\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\},\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\)\\gamma\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\}\)\\gamma\(\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\)\\mathrm\{d\}\\bm\{x\}\_\{0\}\\mathrm\{d\}\\bm\{x\}\_\{1\}\\mathrm\{d\}\\bm\{y\}\_\{0\}\\mathrm\{d\}\\bm\{y\}\_\{1\}\(111\)=infγ∈Π⁡\(μ0,μ1\)∫𝒳4\(⟨𝒙1,𝒚1⟩−⟨𝒙0,𝒚0⟩\)2​γ​\(𝒙0,𝒙1\)​γ​\(𝒚0,𝒚1\)​d​𝒙0​d​𝒙1​d​𝒚0​d​𝒚1\\displaystyle=\\inf\_\{\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)\}\\int\_\{\\mathcal\{X\}^\{4\}\}\(\\langle\\bm\{x\}\_\{1\},\\bm\{y\}\_\{1\}\\rangle\-\\langle\\bm\{x\}\_\{0\},\\bm\{y\}\_\{0\}\\rangle\)^\{2\}\\gamma\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\}\)\\gamma\(\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\)\\mathrm\{d\}\\bm\{x\}\_\{0\}\\mathrm\{d\}\\bm\{x\}\_\{1\}\\mathrm\{d\}\\bm\{y\}\_\{0\}\\mathrm\{d\}\\bm\{y\}\_\{1\}=IGW\-OTS2​\(μ0,μ1\)\\displaystyle=\\text\{IGW\-OT\}^\{2\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)The convexity of Lagrangian is also satisfied\. To show that, let

𝒗t=\(𝒙˙t𝒚˙t\)𝒛t=\(𝒙t𝒚t\)\\bm\{v\}\_\{t\}=\\begin\{pmatrix\}\\dot\{\\bm\{x\}\}\_\{t\}\\\\ \\dot\{\\bm\{y\}\}\_\{t\}\\end\{pmatrix\}\\qquad\\bm\{z\}\_\{t\}=\\begin\{pmatrix\}\\bm\{x\}\_\{t\}\\\\ \\bm\{y\}\_\{t\}\\end\{pmatrix\}\(112\)we have

\|⟨𝒙˙t,𝒚t⟩\+⟨𝒙t,𝒚˙t⟩\|2=\|𝒛t⋅𝒗t\|2=𝒗tT​\(𝒛t​𝒛tT\)​𝒗t\.\|\\langle\\dot\{\\bm\{x\}\}\_\{t\},\\bm\{y\}\_\{t\}\\rangle\+\\langle\\bm\{x\}\_\{t\},\\dot\{\\bm\{y\}\}\_\{t\}\\rangle\|^\{2\}=\|\\bm\{z\}\_\{t\}\\cdot\\bm\{v\}\_\{t\}\|^\{2\}=\\bm\{v\}\_\{t\}^\{\\mathrm\{T\}\}\(\\bm\{z\}\_\{t\}\\bm\{z\}\_\{t\}^\{\\mathrm\{T\}\}\)\\bm\{v\}\_\{t\}\.\(113\)Therefore, the Lagrangian is convex w\.r\.t𝒗t\\bm\{v\}\_\{t\}, hence\(𝒙˙t,𝒚˙t\)\(\\dot\{\\bm\{x\}\}\_\{t\},\\dot\{\\bm\{y\}\}\_\{t\}\)\. Thus, the static IGW objective is recovered as a special case of our path action construction, while the induced dynamic flow need not coincide with their intrinsic IGW flow \(for mathematical details, please refer to their paper\)\. From this perspective, our TP\-DATE framework can be viewed as a more general dynamic QOT framework\.

#### D\.2\.1The 1D case

In the IGW\-OT case, we also needd≥2d\\geq 2to makert=\(1−t\)​r0\+t​r1r\_\{t\}=\(1\-t\)r\_\{0\}\+tr\_\{1\}attainable\. We give an explicit construction here\. Let𝒑t=𝒙t\+𝒚t2,𝒒t=𝒙t−𝒚t2\\bm\{p\}\_\{t\}=\\frac\{\\bm\{x\}\_\{t\}\+\\bm\{y\}\_\{t\}\}\{\\sqrt\{2\}\},\\bm\{q\}\_\{t\}=\\frac\{\\bm\{x\}\_\{t\}\-\\bm\{y\}\_\{t\}\}\{\\sqrt\{2\}\}, we have⟨𝒙t,𝒚t⟩=12​\(‖𝒑t‖2−‖𝒒t‖2\)\\langle\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\}\\rangle=\\frac\{1\}\{2\}\(\\\|\\bm\{p\}\_\{t\}\\\|^\{2\}\-\\\|\\bm\{q\}\_\{t\}\\\|^\{2\}\)\. Whend≥2d\\geq 2, it is easy to construct𝒑t,𝒒t\\bm\{p\}\_\{t\},\\bm\{q\}\_\{t\}such that‖𝒑t‖2=\(1−t\)​‖𝒑0‖2\+t​‖𝒑1‖2,‖𝒒t‖2=\(1−t\)​‖𝒒0‖2\+t​‖𝒒1‖2\\\|\\bm\{p\}\_\{t\}\\\|^\{2\}=\(1\-t\)\\\|\\bm\{p\}\_\{0\}\\\|^\{2\}\+t\\\|\\bm\{p\}\_\{1\}\\\|^\{2\},\\\|\\bm\{q\}\_\{t\}\\\|^\{2\}=\(1\-t\)\\\|\\bm\{q\}\_\{0\}\\\|^\{2\}\+t\\\|\\bm\{q\}\_\{1\}\\\|^\{2\}, hence⟨𝒙t,𝒚t⟩=\(1−t\)​⟨𝒙0,𝒚0⟩\+t⁡⟨𝒙1,𝒚1⟩\\langle\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\}\\rangle=\(1\-t\)\\langle\\bm\{x\}\_\{0\},\\bm\{y\}\_\{0\}\\rangle\+t\\langle\\bm\{x\}\_\{1\},\\bm\{y\}\_\{1\}\\rangle\.

Whend=1d=1, considerx0=1,y0=−1,x1=−1,y1=1x\_\{0\}=1,y\_\{0\}=\-1,x\_\{1\}=\-1,y\_\{1\}=1\. By direct calculation, we havex0​y0=x1​y1=−1x\_\{0\}y\_\{0\}=x\_\{1\}y\_\{1\}=\-1\. But, it is impossible to construct a travelling pair\(xt,yt\)\(x\_\{t\},y\_\{t\}\)such thatxt​yt≡−1x\_\{t\}y\_\{t\}\\equiv\-1, since there must exists somettsuch thatxt=0x\_\{t\}=0, hencext​yt=0≠−1x\_\{t\}y\_\{t\}=0\\neq\-1\. The reason is similar to the GW\-OT case\. A reflection can be continuously realized by a rigid body rotation in high dimensional spaced≥2d\\geq 2, but can never be realized by a continuous rigid transformation in 1D spaces\.

Whend=1d=1, considerℒ=\(dd​t​\(xt​yt\)\)2\\mathcal\{L\}=\\Big\(\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\(x\_\{t\}y\_\{t\}\)\\Big\)^\{2\}, by direct calculation, one can easily check that the corresponding static cost is

𝒜\(x0,x1,y0,y1\)=\{\(\|x0​y0\|\+\|x1​y1\|\)2,x0​x1<0,y0​y1<0\(x1​y1−x0​y0\)2,otherwise\\mathcal\{A\}\(x\_\{0\},x\_\{1\},y\_\{0\},y\_\{1\}\)=\\left\\\{\\begin\{aligned\} &\(\|x\_\{0\}y\_\{0\}\|\+\|x\_\{1\}y\_\{1\}\|\)^\{2\},\\quad x\_\{0\}x\_\{1\}<0,y\_\{0\}y\_\{1\}<0\\\\ &\(x\_\{1\}y\_\{1\}\-x\_\{0\}y\_\{0\}\)^\{2\},\\quad\\text\{otherwise\}\\end\{aligned\}\\right\.\(114\)The infimum is attained whenx0​y0​x1​y1≠0x\_\{0\}y\_\{0\}x\_\{1\}y\_\{1\}\\neq 0orx0​y0=x1​y1=0x\_\{0\}y\_\{0\}=x\_\{1\}y\_\{1\}=0, but may not be attainable whenx0​y0=0,x1​y1≠0x\_\{0\}y\_\{0\}=0,x\_\{1\}y\_\{1\}\\neq 0orx0​y0≠0,x1​y1=0x\_\{0\}y\_\{0\}\\neq 0,x\_\{1\}y\_\{1\}=0\.

### D\.3OT\-CFM

\([Tong et al\., 2024a](https://arxiv.org/html/2609.20008#bib.bib19)\)has proposed OT\-CFM, a flow matching framework for solving standard dynamic OT problem\. It uses the optimal OT coupling under squared Euclidean distance cost, and the displacement conditional path for flow matching training\. The learned flow can be proved to solve the standard dynamic OT problem with kinetic energy cost\.

\\displaystyleOT​\(μ0,μ1\)=infρ,𝒖∫01∫𝒳12​‖𝒖⁡\(𝒙,t\)‖22​ρt​\(𝒙\)​𝑑𝒙​𝑑t\\displaystyle\\text\{OT\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\inf\_\{\\rho,\\bm\{u\}\}\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}\}\\,\\frac\{1\}\{2\}\\\|\\bm\{u\}\(\\bm\{x\},t\)\\\|\_\{2\}^\{2\}\\rho\_\{t\}\(\\bm\{x\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}t\(115\)s\.t\.∂tρ\+∇𝒙⋅\(ρ​𝒖\)=0,ρ0=μ0,ρ1=μ1\\displaystyle\\text\{s\.t\.\}\\ \\ \\ \\ \\ \\ \\ \\partial\_\{t\}\\rho\+\\nabla\_\{\\bm\{x\}\}\\cdot\(\\rho\\bm\{u\}\)=0,\\ \\rho\_\{0\}=\\mu\_\{0\},\\ \\rho\_\{1\}=\\mu\_\{1\}
Inspired by this pioneering work, we proposed a new travelling pair flow matching framework which can solve the dynamic QOT problem\. When choosing the Lagrangian in our dynamic QOT framework as

ℒ⁡\(t,𝒙,𝒚,𝒙˙,𝒚˙\)=12​‖𝒙˙‖2\+12​‖𝒚˙‖2\\mathcal\{L\}\(t,\\bm\{x\},\\bm\{y\},\\dot\{\\bm\{x\}\},\\dot\{\\bm\{y\}\}\)=\\frac\{1\}\{2\}\\\|\\dot\{\\bm\{x\}\}\\\|^\{2\}\+\\frac\{1\}\{2\}\\\|\\dot\{\\bm\{y\}\}\\\|^\{2\}\(116\)the corresponding dynamic QOT reduces to \([115](https://arxiv.org/html/2609.20008#A4.E115)\)\. Therefore, TP\-DATE reduces to OT\-CFM under this setting\.

### D\.4GENOT

\([Klein et al\., 2024](https://arxiv.org/html/2609.20008#bib.bib33)\)also combines flow matching with GW\-OT, but with a fundamentally different purpose\. GENOT is a neural OT solver that uses flow matching to learn couplings associated with static entropic OT, GW\-OT, or FGW\-OT\. In contrast, TP\-DATE does not use flow matching to solve the QOT coupling\. It uses flow matching to find the dynamic QOT probability flow\. The flow in GENOT serves as a parameterization of a static transport plan, while the flow in TP\-DATE represents the reconstructed dynamics between snapshots\. The two approaches therefore address different problems\.

### D\.5NicheFlow

NicheFlow\([Sakalyan et al\., 2025](https://arxiv.org/html/2609.20008#bib.bib36)\)advances spatial trajectory inference by modeling local cellular microenvironments as structured point clouds, jointly capturing changes in spatial coordinates and gene expression states\. It combines OT\-based source–target matching with Variational Flow Matching \(VFM\) to generate future microenvironments conditioned on observed source niches\. However, its flow evolves from noise to the target state conditioned on the source, so the internal flow time represents a generative process rather than the continuous time biological dynamics between two tissue states\. In contrast, TP\-DATE directly models source\-to\-target continuous time dynamics and imposes structure preservation at the dynamical level through dynamic QOT\.

### D\.6ContextFlow

ContextFlow\([Rathod et al\., 2026](https://arxiv.org/html/2609.20008#bib.bib2)\)is a recently proposed context\-aware generative modeling approach for transcriptomic snapshot interpolation\. It incorporates spatial information and cell–cell communication signals into a static OT cost, thereby obtaining couplings that account for these contextual factors\. These couplings are then used in conditional flow matching to reconstruct the dynamics\. From this perspective, ContextFlow can be viewed as introducing a class of context\-aware static OT formulations\. In contrast, TP\-DATE incorporates spatial structure directly at the dynamical level rather than indirectly through a static coupling\. Viewed this way, TP\-DATE extends the context\-aware principle to dynamic settings through the framework of dynamic QOT\.

### D\.7stVCR

stVCR\([Peng et al\., 2026a](https://arxiv.org/html/2609.20008#bib.bib12)\)models spatiotemporal dynamics of single cells with Wasserstein\-Fisher\-Rao dynamics \(WFR\)\. WFR can be viewed as an unbalanced extension of standard OT\. It extends OT from transporting probability densities to transporting unbalanced measures\.

WFRδ2​\(μ0,μ1\)=infρ,𝒖∫01∫𝒳12​\[‖𝒖⁡\(𝒙,t\)‖22\+δ2​gt2​\(𝒙\)\]​ρt​\(𝒙\)​𝑑𝒙​𝑑t\\displaystyle\\text\{WFR\}\_\{\\delta\}^\{2\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\inf\_\{\\rho,\\bm\{u\}\}\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}\}\\,\\frac\{1\}\{2\}\[\\\|\\bm\{u\}\(\\bm\{x\},t\)\\\|\_\{2\}^\{2\}\+\\delta^\{2\}g\_\{t\}^\{2\}\(\\bm\{x\}\)\]\\rho\_\{t\}\(\\bm\{x\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}t\(117\)s\.t\.∂tρ\+∇𝒙⋅\(ρ​𝒖\)=g​ρ,ρ0=μ0,ρ1=μ1\\displaystyle\\text\{s\.t\.\}\\ \\ \\ \\ \\ \\ \\ \\ \\ \\ \\ \\ \\ \\partial\_\{t\}\\rho\+\\nabla\_\{\\bm\{x\}\}\\cdot\(\\rho\\bm\{u\}\)=g\\rho,\\ \\rho\_\{0\}=\\mu\_\{0\},\\ \\rho\_\{1\}=\\mu\_\{1\}It models the cell proliferation and apoptosis by a growth rate termgt​\(𝒙\)g\_\{t\}\(\\bm\{x\}\)\. When applied on scRNA\-seq data without spatial information, this WFR problem can be solved by unbalanced flow matching\([Peng et al\., 2026b](https://arxiv.org/html/2609.20008#bib.bib14)\)\. On spatial transcriptomics data, to preserve the spatial structure, an optional spatial structure preserving loss is added\. To account for possible slight misalignment between the coordinate systems of spatial transcriptomics slices, stVCR further introduces a rigid body transformation invariant OT formulation\. Before each OT training step, it first solves an optimal rigid body transformation to align the coordinate systems\. Both techniques make it unsolvable for flow matching based methods\. Therefore, stVCR used NeuralODE method to solve the corresponding OT problem\.

In contrast, TP\-DATE naturally preserves the spatial structure and addresses slight misalignment by adding interaction termΦ\\Phi, and is simulation\-free\. Although TP\-DATE is more efficient than stVCR in modeling spatial structure preservation, it does not currently account for cell proliferation and apoptosis as stVCR does\. This limitation points to an important direction for future work: extending the dynamic QOT framework to the unbalanced setting\.

### D\.8CytoBridge

CytoBridge\([Zhang et al\., 2025b](https://arxiv.org/html/2609.20008#bib.bib13)\)is a NeuralODE framework for learning interactions and mean\-field interactive dynamics from snapshots\. They focuses on the unbalanced mean\-field Schrödinger bridge problem \(UMFSB\)\.

UMFSB​\(μ0,μ1\)=infρ,𝒖∫01∫𝒳12​\[‖𝒖⁡\(𝒙,t\)‖22\+gt2​\(𝒙\)\]​ρt​\(𝒙\)​𝑑𝒙​𝑑t\\displaystyle\\text\{UMFSB\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\inf\_\{\\rho,\\bm\{u\}\}\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}\}\\,\\frac\{1\}\{2\}\[\\\|\\bm\{u\}\(\\bm\{x\},t\)\\\|\_\{2\}^\{2\}\+g\_\{t\}^\{2\}\(\\bm\{x\}\)\]\\rho\_\{t\}\(\\bm\{x\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}t\(118\)s\.t\.∂tρ⁡\(𝒙\)\+∇𝒙⋅\[ρ⁡\(𝒙\)​\(𝒖t​\(𝒙\)−∫k⁡\(𝒙,𝒚\)​∇𝒙Φ​\(𝒙−𝒚\)​ρt​\(𝒚\)​𝑑𝒚\)\]=g​ρ\\displaystyle\\text\{s\.t\.\}\\ \\ \\ \\partial\_\{t\}\\rho\(\\bm\{x\}\)\+\\nabla\_\{\\bm\{x\}\}\\cdot\[\\rho\(\\bm\{x\}\)\(\\bm\{u\}\_\{t\}\(\\bm\{x\}\)\-\\int k\(\\bm\{x\},\\bm\{y\}\)\\nabla\_\{\\bm\{x\}\}\\Phi\(\\bm\{x\}\-\\bm\{y\}\)\\rho\_\{t\}\(\\bm\{y\}\)\\mathrm\{d\}\\bm\{y\}\)\]=g\\rhoρ0=μ0,ρ1=μ1\\displaystyle\\rho\_\{0\}=\\mu\_\{0\},\\ \\rho\_\{1\}=\\mu\_\{1\}whereΦ\\Phiis some interaction potential for modeling cell\-cell interactions, andkkis some gate function to control the intensity of the interaction\. The key distinction between CytoBridge and TP\-DATE is where the interaction term is introduced\. CytoBridge incorporates the interaction term directly into the dynamical equation as part of the constraint, whereas TP\-DATE retains the standard continuity equation constraint on the two\-particle space and instead places the interaction term in the action functional\.

One advantage of CytoBridge is that, because the interaction term is part of the dynamics to be learned, the interaction itself can be inferred from data through a NeuralODE\. This is feasible for NeuralODE based methods\. In principle, essentially any interaction term can be parameterized by a neural network and learned by minimizing the loss\. For flow matching based methods, however, this is generally much more challenging, since the conditional paths determined by the interaction term must typically be computed in advance\. Whether techniques from inverse problems can be used to overcome this limitation and enable flow matching methods to learn such interactions from the data is therefore a very interesting direction for future research\.

Another interesting question is whether CytoBridge can be made simulation\-free\. From TrajectoryNet\([Tong et al\., 2020](https://arxiv.org/html/2609.20008#bib.bib34)\)to OT\-CFM\([Tong et al\., 2024a](https://arxiv.org/html/2609.20008#bib.bib19)\), Diffusion Schrödinger bridge\([De Bortoli et al\., 2021](https://arxiv.org/html/2609.20008#bib.bib27);[Shi et al\., 2023](https://arxiv.org/html/2609.20008#bib.bib37)\)toSF2​M\\text\{SF\}^\{2\}\\text\{M\}\([Tong et al\., 2024b](https://arxiv.org/html/2609.20008#bib.bib28)\), stVCR\([Peng et al\., 2026a](https://arxiv.org/html/2609.20008#bib.bib12)\)to WFR\-FM[Peng et al\. \(2026b\)](https://arxiv.org/html/2609.20008#bib.bib14), previous works have already built simulation\-free flow matching based algorithms for OT, Schrödinger bridge, and WFR\. There is a trend of converting simulation\-based or iteration\-based Neural OT solvers to simulation\-free solvers by flow matching\. However, the success of these algorithms relies on an important mathematical property of the underlying OT problems: the dynamical constraint is linear, or equivalently, local, with respect to the measure flowρt\\rho\_\{t\}\. This property is precisely what allows the marginalization theorem to hold in these settings\. Once a CytoBridge style interaction term is introduced directly into the dynamical constraint, this linearity or locality is lost\. One can readily verify that, for the CytoBridge constraint equation, the usual marginalization argument no longer applies\. This limitation is also reflected in the independence of conditional paths in standard flow matching\. Once the initial and terminal states are fixed, each conditional path is determined independently of the others, with no coupling between different paths\. As a result, standard conditional path constructions are not naturally equipped to represent interactions between particles\. From this perspective, extending CytoBridge to a fully simulation\-free flow matching formulation appears to be fundamentally challenging\.

TP\-DATE’s travelling\-pair flow matching may provide a possible starting point for addressing this challenge\. By considering the problem on the two\-particle space, we simultaneously enable interactions between conditional paths while preserving the linearity of the dynamical constraint\. Moreover, through velocity decomposition, the resulting single\-particle dynamics can be interpreted as an approximation to mean\-field dynamics, in which interactions at the population level emerge from underlying pairwise interactions\. Although TP\-DATE was not originally designed for this purpose, a promising direction for future work is to build on travelling\-pair flow matching and investigate whether formulating the problem on the two\-particle space can enable efficient flow matching based modeling of nonlocal dynamics\.

## Appendix EVelocity decomposition

In general cases, one may want to only have a velocity decomposition at conditional velocity level𝒖t1\(𝒙,𝒚\|𝒛,𝒛′\)=𝒃t\(𝒙\|𝒛,𝒛′\)\+𝒇t\(𝒙,𝒚\|𝒛,𝒛′\)\.\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)=\\bm\{b\}\_\{t\}\(\\bm\{x\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\+\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\.In this case, we want to regress the marginal velocities via flow matching\.

ℒb​\(𝜽\)\\displaystyle\\mathcal\{L\}\_\{b\}\(\\bm\{\\theta\}\)=𝔼t∼𝒰\[0,1\],\(𝒛,𝒛′\)∼q\(𝒛,𝒛′\),\(𝒙,𝒚\)∼πt\(𝒙,𝒚\|𝒛,𝒛′\)‖𝒃𝜽\(𝒙,t\)−𝒃t\(𝒙\|𝒛,𝒛′\)‖22\\displaystyle=\\mathbb\{E\}\_\{t\\sim\\mathcal\{U\}\[0,1\],\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\sim q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\),\(\\bm\{x\},\\bm\{y\}\)\\sim\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\}\\left\\\|\\bm\{\\bm\{b\}\_\{\\theta\}\}\(\\bm\{x\},t\)\-\\bm\{b\}\_\{t\}\(\\bm\{x\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\right\\\|\_\{2\}^\{2\}\(119\)ℒf​\(ϕ\)\\displaystyle\\mathcal\{L\}\_\{f\}\(\\bm\{\\phi\}\)=𝔼t∼𝒰\[0,1\],\(𝒛,𝒛′\)∼q\(𝒛,𝒛′\),\(𝒙,𝒚\)∼πt\(𝒙,𝒚\|𝒛,𝒛′\)‖𝒇ϕ\(𝒙,𝒚,t\)−𝒇t\(𝒙,𝒚\|𝒛,𝒛′\)‖22\\displaystyle=\\mathbb\{E\}\_\{t\\sim\\mathcal\{U\}\[0,1\],\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\sim q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\),\(\\bm\{x\},\\bm\{y\}\)\\sim\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\}\\left\\\|\\bm\{\\bm\{f\}\_\{\\phi\}\}\(\\bm\{x\},\\bm\{y\},t\)\-\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\right\\\|\_\{2\}^\{2\}Similarly to theorem[5\.6](https://arxiv.org/html/2609.20008#S5.Thmtheorem6), we still have

ℒb​\(𝜽\)\\displaystyle\\mathcal\{L\}\_\{b\}\(\\bm\{\\theta\}\)=𝔼t∼𝒰⁡\[0,1\],\(𝒙,𝒚\)∼πt​\(𝒙,𝒚\)​‖𝒃𝜽​\(𝒙,t\)−𝒃t​\(𝒙\)‖22\+C1\\displaystyle=\\mathbb\{E\}\_\{t\\sim\\mathcal\{U\}\[0,1\],\(\\bm\{x\},\\bm\{y\}\)\\sim\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\}\\left\\\|\\bm\{\\bm\{b\}\_\{\\theta\}\}\(\\bm\{x\},t\)\-\\bm\{b\}\_\{t\}\(\\bm\{x\}\)\\right\\\|\_\{2\}^\{2\}\+C\_\{1\}\(120\)=𝔼t∼𝒰⁡\[0,1\],𝒙∼ρt​\(𝒙\)​‖𝒃𝜽​\(𝒙,t\)−𝒃t​\(𝒙\)‖22\+C1\\displaystyle=\\mathbb\{E\}\_\{t\\sim\\mathcal\{U\}\[0,1\],\\bm\{x\}\\sim\\rho\_\{t\}\(\\bm\{x\}\)\}\\left\\\|\\bm\{\\bm\{b\}\_\{\\theta\}\}\(\\bm\{x\},t\)\-\\bm\{b\}\_\{t\}\(\\bm\{x\}\)\\right\\\|\_\{2\}^\{2\}\+C\_\{1\}ℒf​\(ϕ\)\\displaystyle\\mathcal\{L\}\_\{f\}\(\\bm\{\\phi\}\)=𝔼t∼𝒰⁡\[0,1\],\(𝒙,𝒚\)∼πt​\(𝒙,𝒚\)​‖𝒇ϕ​\(𝒙,𝒚,t\)−𝒇t​\(𝒙,𝒚\)‖22\+C2\\displaystyle=\\mathbb\{E\}\_\{t\\sim\\mathcal\{U\}\[0,1\],\(\\bm\{x\},\\bm\{y\}\)\\sim\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\}\\left\\\|\\bm\{\\bm\{f\}\_\{\\phi\}\}\(\\bm\{x\},\\bm\{y\},t\)\-\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\right\\\|\_\{2\}^\{2\}\+C\_\{2\}whereC1,C2C\_\{1\},C\_\{2\}are constants independent to𝜽,ϕ\\bm\{\\theta\},\\bm\{\\phi\}\. Therefore, the minimizer is𝒃t​\(𝒙\)\\bm\{b\}\_\{t\}\(\\bm\{x\}\)and𝒇t​\(𝒙,𝒚\)\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\. The problem here is that the conditional velocity decomposition may not lead to a marginal velocity decomposition, i\.e\.𝒖t1​\(𝒙,𝒚\)=𝒃t​\(𝒙\)\+𝒇t​\(𝒙,𝒚\)\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)=\\bm\{b\}\_\{t\}\(\\bm\{x\}\)\+\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)does not naturally hold\. By marginalization theorems, we have

\{𝒖1t\(𝒙,𝒚\)=∫𝒖1t\(𝒙,𝒚\|𝒛,𝒛′\)πt\(𝒙,𝒚\|𝒛,𝒛′\)q\(𝒛,𝒛′\)πt​\(𝒙,𝒚\)d𝒛d𝒛′𝒃t​\(𝒙\)=∫\(∫𝒃t​\(𝒙\|𝒛,𝒛′\)​πt\(𝒙,𝒚\|𝒛,𝒛′\)q\(𝒛,𝒛′\)πt​\(𝒙,𝒚\)​𝒅𝒛​d​𝒛′\)​πt​\(𝒙,𝒚\)ρt​\(𝒙\)​𝒅𝒚𝒇t\(𝒙,𝒚\)=∫𝒇t\(𝒙,𝒚\|𝒛,𝒛′\)πt\(𝒙,𝒚\|𝒛,𝒛′\)q\(𝒛,𝒛′\)πt​\(𝒙,𝒚\)d𝒛d𝒛′\\left\\\{\\begin\{aligned\} &\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)=\\int\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\frac\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\}\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\}\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}\\\\ &\\bm\{b\}\_\{t\}\(\\bm\{x\}\)=\\int\\Big\(\\int\\bm\{b\}\_\{t\}\(\\bm\{x\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\frac\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\}\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\}\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}\\Big\)\\frac\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\}\{\\rho\_\{t\}\(\\bm\{x\}\)\}\\mathrm\{d\}\\bm\{y\}\\\\ &\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)=\\int\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\frac\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\}\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\}\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}\\end\{aligned\}\\right\.\(121\)The conditional velocity decomposition yields

𝔼\[𝒃t\(𝑿\|𝒛,𝒛′\)\|𝑿=𝒙,𝒀=𝒚\]\\displaystyle\\mathbb\{E\}\[\\bm\{b\}\_\{t\}\(\\bm\{X\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\|\\bm\{X\}=\\bm\{x\},\\bm\{Y\}=\\bm\{y\}\]\(122\)=\\displaystyle=∫𝒃t​\(𝒙\|𝒛,𝒛′\)​πt\(𝒙,𝒚\|𝒛,𝒛′\)q\(𝒛,𝒛′\)πt​\(𝒙,𝒚\)​𝑑𝒛​d​𝒛′\\displaystyle\\int\\bm\{b\}\_\{t\}\(\\bm\{x\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\frac\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\}\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\}\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}=\\displaystyle=∫\[𝒖1t\(𝒙,𝒚\|𝒛,𝒛′\)−𝒇t\(𝒙,𝒚\|𝒛,𝒛′\)\]πt\(𝒙,𝒚\|𝒛,𝒛′\)q\(𝒛,𝒛′\)πt​\(𝒙,𝒚\)d𝒛d𝒛′\\displaystyle\\int\[\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\-\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\]\\frac\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\}\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\}\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}=\\displaystyle=𝒖t1​\(𝒙,𝒚\)−𝒇t​\(𝒙,𝒚\)\\displaystyle\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\-\\bm\{f\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)By definition,

𝒃t​\(𝒙\)=𝔼⁡\[𝒃t​\(𝑿\|𝒛,𝒛′\)\|𝑿=𝒙\]\\bm\{b\}\_\{t\}\(\\bm\{x\}\)=\\mathbb\{E\}\[\\bm\{b\}\_\{t\}\(\\bm\{X\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\|\\bm\{X\}=\\bm\{x\}\]\(123\)Therefore, the necessary and sufficient condition of marginal velocity decomposition is

𝔼\[𝒃t\(𝑿\|𝒛,𝒛′\)\|𝑿=𝒙,𝒀=𝒚\]=𝔼\[𝒃t\(𝑿\|𝒛,𝒛′\)\|𝑿=𝒙\]\\mathbb\{E\}\[\\bm\{b\}\_\{t\}\(\\bm\{X\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\|\\bm\{X\}=\\bm\{x\},\\bm\{Y\}=\\bm\{y\}\]=\\mathbb\{E\}\[\\bm\{b\}\_\{t\}\(\\bm\{X\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\|\\bm\{X\}=\\bm\{x\}\]\(124\)Intuitively, the condition requires that, once the current state of the particle is given, the state of the interacting particle provides no additional information about the conditional mean of the self\-driven velocity\. It is a conditional mean\-independence condition, which is weaker than conditional independence\. The simplest example is what we introduced in the main text𝒃t\(𝒙,𝒚\|𝒛,𝒛′\)=𝒃t\(𝒙\)\\bm\{b\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)=\\bm\{b\}\_\{t\}\(\\bm\{x\}\)\. It is easy to check that it satisfies the condition\. Other kind of decompositions have to be designed carefully to satisfy this condition, in order to be learned by flow matching\.

## Appendix FThe BB\-form

For convenience, we restate the three forms below\. The static form is

QOTS​\(μ0,μ1\)=infγ∈Π⁡\(μ0,μ1\)∫𝒳4𝒜⁡\(𝒛,𝒛′\)​γ​\(𝒛\)​γ​\(𝒛′\)​d𝒛​d​𝒛′\.\\displaystyle\\text\{QOT\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\inf\_\{\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)\}\\int\_\{\\mathcal\{X\}^\{4\}\}\\mathcal\{A\}\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\gamma\(\\bm\{z\}\)\\gamma\(\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}\.\(125\)The dynamic form is

QOTD​\(μ0,μ1\)=\\displaystyle\\text\{QOT\}\_\{D\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\(126\)inf∫01∫∫𝒳2ℒ\(t,𝒙,𝒚,u1t\(𝒙,𝒚\|𝒛,𝒛′\),u2t\(𝒙,𝒚\|𝒛,𝒛′\)\)πt\(𝒙,𝒚\|𝒛,𝒛′\)γ\(𝒛\)γ\(𝒛′\)d𝒙d𝒚d𝒛d𝒛′dt\.\\displaystyle\\inf\\int\_\{0\}^\{1\}\\int\\int\_\{\\mathcal\{X\}^\{2\}\}\\mathcal\{L\}\(t,\\bm\{x\},\\bm\{y\},u^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\),u^\{2\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\)\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\gamma\(\\bm\{z\}\)\\gamma\(\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}\\mathrm\{d\}t\.The BB\-form is

QOTB​B​\(μ0,μ1\)=infπ,𝒖1,𝒖2∫01∫𝒳2ℒ⁡\(t,𝒙,𝒚,ut1​\(𝒙,𝒚\),ut2​\(𝒙,𝒚\)\)​πt​\(𝒙,𝒚\)​𝑑𝒙​𝑑𝒚​𝑑t\\displaystyle\\text\{QOT\}\_\{BB\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\inf\_\{\\pi,\\bm\{u\}^\{1\},\\bm\{u\}^\{2\}\}\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}^\{2\}\}\\mathcal\{L\}\(t,\\bm\{x\},\\bm\{y\},u^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\),u^\{2\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\)\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}t\(127\)∂tπt​\(𝒙,𝒚\)\+∇𝒙⋅\(πt​\(𝒙,𝒚\)​ut1​\(𝒙,𝒚\)\)\+∇𝒚⋅\(πt​\(𝒙,𝒚\)​ut2​\(𝒙,𝒚\)\)=0\.\\displaystyle\\partial\_\{t\}\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\+\\nabla\_\{\\bm\{x\}\}\\cdot\(\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)u^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\)\+\\nabla\_\{\\bm\{y\}\}\\cdot\(\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)u^\{2\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\)=0\.
### F\.1Difficulty of the equivalence between static QOT and BB\-form

Inspired by the pioneering work of\([Benamou and Brenier, 2000](https://arxiv.org/html/2609.20008#bib.bib21)\), one may expect to have an equivalence between the static form and the BB\-form also for QOT\. We explain why it can not\. Let𝒛=\(𝒙,𝒚\)\\bm\{z\}=\(\\bm\{x\},\\bm\{y\}\), and𝒖t​\(𝒛\)=\(ut1​\(𝒙,𝒚\)ut2​\(𝒙,𝒚\)\)\\bm\{u\}\_\{t\}\(\\bm\{z\}\)=\\begin\{pmatrix\}u^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\\\ u^\{2\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\end\{pmatrix\}, the BB\-form can be reformulated as

QOTB​B​\(μ0,μ1\)=infπ,𝒖∫01∫𝒳2ℒ⁡\(t,𝒛,ut​\(𝒛\)\)​πt​\(𝒛\)​𝑑𝒛​𝑑t\\displaystyle\\text\{QOT\}\_\{BB\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\inf\_\{\\pi,\\bm\{u\}\}\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}^\{2\}\}\\mathcal\{L\}\(t,\\bm\{z\},u\_\{t\}\(\\bm\{z\}\)\)\\pi\_\{t\}\(\\bm\{z\}\)\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}t\(128\)s\.t\.∂tπt​\(𝒛\)\+∇𝒛⋅\(πt​\(𝒛\)​ut​\(𝒛\)\)=0\.\\displaystyle\\text\{s\.t\.\}\\ \\ \\ \\ \\ \\ \\ \\ \\ \\ \\ \\ \\partial\_\{t\}\\pi\_\{t\}\(\\bm\{z\}\)\+\\nabla\_\{\\bm\{z\}\}\\cdot\(\\pi\_\{t\}\(\\bm\{z\}\)u\_\{t\}\(\\bm\{z\}\)\)=0\.which is just the standard OT on the two\-particle space𝒳2\\mathcal\{X\}^\{2\}\. When choosingℒ⁡\(t,𝒛,ut​\(𝒛\)\)=12​‖𝒛˙‖2\\mathcal\{L\}\(t,\\bm\{z\},u\_\{t\}\(\\bm\{z\}\)\)=\\frac\{1\}\{2\}\\\|\\dot\{\\bm\{z\}\}\\\|^\{2\},\([Benamou and Brenier, 2000](https://arxiv.org/html/2609.20008#bib.bib21)\)proved that it is equivalent to the static OT on the two\-particle space

OT​\(μ0,μ1\)=infq∈Π⁡\(μ0⊗μ0,μ1⊗μ1\)∫𝒳4𝒜⁡\(𝒛,𝒛′\)​q​\(𝒛,𝒛′\)​d𝒛​d​𝒛′\.\\displaystyle\\text\{OT\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\inf\_\{q\\in\\Pi\(\\mu\_\{0\}\\otimes\\mu\_\{0\},\\mu\_\{1\}\\otimes\\mu\_\{1\}\)\}\\int\_\{\\mathcal\{X\}^\{4\}\}\\mathcal\{A\}\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}\.\(129\)This static form is slightly different to the static QOT\. In static QOT \([125](https://arxiv.org/html/2609.20008#A6.E125)\), we restrict the couplingqqto a smaller set

Π2​\(μ0,μ1\)≜\{q=γ⊗γ\|γ∈Π⁡\(μ0,μ1\)\}⊂Π⁡\(μ0⊗μ0,μ1⊗μ1\)\.\\Pi^\{2\}\(\\mu\_\{0\},\\mu\_\{1\}\)\\triangleq\\\{q=\\gamma\\otimes\\gamma\|\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)\\\}\\subset\\Pi\(\\mu\_\{0\}\\otimes\\mu\_\{0\},\\mu\_\{1\}\\otimes\\mu\_\{1\}\)\.\(130\)From this perspective, the static QOT on𝒳\\mathcal\{X\}can be viewed as kind of special OT on𝒳2\\mathcal\{X\}^\{2\}with more strict coupling constraint\. Therefore, it is naturally to have onlyQOTS​\(μ0,μ1\)≥QOTB​B​\(μ0,μ1\)\\text\{QOT\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)\\geq\\text\{QOT\}\_\{BB\}\(\\mu\_\{0\},\\mu\_\{1\}\), and it is essentially hard to improve it\.

### F\.2Difficulty of the equivalence between static QOT and the constrained BB\-form

One may naturally ask whether the BB\-form can also be equipped with additional constraints so that it becomes compatible with the extra coupling constraint imposed in static QOT\. The answer is also no\.

By superposition principle\([Ambrosio et al\., 2008](https://arxiv.org/html/2609.20008#bib.bib41);[Lisini, 2007](https://arxiv.org/html/2609.20008#bib.bib40)\), for any admissible probability pathπt​\(𝒙,𝒚\)\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)satisfying the continuity equation∂tπt​\(𝒙,𝒚\)\+∇𝒙⋅\(πt​\(𝒙,𝒚\)​𝒖t1​\(𝒙t,𝒚t\)\)\+∇𝒚⋅\(πt​\(𝒙,𝒚\)​𝒖t2​\(𝒙t,𝒚t\)\)=0\\partial\_\{t\}\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\+\\nabla\_\{\\bm\{x\}\}\\cdot\(\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\bm\{u\}\_\{t\}^\{1\}\(\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\}\)\)\+\\nabla\_\{\\bm\{y\}\}\\cdot\(\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\bm\{u\}\_\{t\}^\{2\}\(\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\}\)\)=0, the superposition principle states that there exists a probability measureη\\etaover the path spaceC⁡\(\[0,1\],𝒳2\)C\(\[0,1\],\\mathcal\{X\}^\{2\}\)such that for a path\(𝒙t,𝒚t\)\(\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\}\),𝒙t˙=𝒖t1​\(𝒙,𝒚\),𝒚t˙=𝒖t2​\(𝒙,𝒚\)\\dot\{\\bm\{x\}\_\{t\}\}=\\bm\{u\}^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\),\\dot\{\\bm\{y\}\_\{t\}\}=\\bm\{u\}^\{2\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)holdsη−a\.e\.\\eta\-a\.e\., and the projection ofη\\etaat timettisπt\\pi\_\{t\}\. That means, for pathΓ\\Gammaand evaluation mappinget:et​\(Γ\)=Γte\_\{t\}:e\_\{t\}\(\\Gamma\)=\\Gamma\_\{t\}, we have\(et\)\#​η=πt\(e\_\{t\}\)\_\{\\\#\}\\eta=\\pi\_\{t\}\. Intuitively, it says that a probability path generated by a continuity equation can be viewed as a superposition of several Dirac\-Dirac paths which in this paper we also called the conditional paths\.

To be compatible to theγ⊗γ\\gamma\\otimes\\gammastructure of static QOT, the condition

∃γ∈Π⁡\(μ0,μ1\),s\.tπ0,1≜\(e0,e1\)\#​η=γ⊗γ\\exists\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\),\\quad\\text\{s\.t\}\\quad\\pi\_\{0,1\}\\triangleq\(e\_\{0\},e\_\{1\}\)\_\{\\\#\}\\eta=\\gamma\\otimes\\gamma\(131\)naturally follows\. We call it the QOT compatibility condition\. With this condition, we define the QOT compatible BB\-form

QOTC​\(μ0,μ1\)=infπ,𝒖1,𝒖2∫01∫𝒳2ℒ⁡\(t,𝒙,𝒚,ut1​\(𝒙,𝒚\),ut2​\(𝒙,𝒚\)\)​πt​\(𝒙,𝒚\)​𝑑𝒙​𝑑𝒚​𝑑t\\displaystyle\\text\{QOT\}\_\{C\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\inf\_\{\\pi,\\bm\{u\}^\{1\},\\bm\{u\}^\{2\}\}\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{X\}^\{2\}\}\\mathcal\{L\}\(t,\\bm\{x\},\\bm\{y\},u^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\),u^\{2\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\)\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}t\(132\)s\.t\.∃γ∈Π⁡\(μ0,μ1\),π0,1=γ⊗γ\\displaystyle\\text\{s\.t\.\}\\quad\\exists\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\),\\quad\\pi\_\{0,1\}=\\gamma\\otimes\\gamma∂tπt​\(𝒙,𝒚\)\+∇𝒙⋅\(πt​\(𝒙,𝒚\)​ut1​\(𝒙,𝒚\)\)\+∇𝒚⋅\(πt​\(𝒙,𝒚\)​ut2​\(𝒙,𝒚\)\)=0\.\\displaystyle\\partial\_\{t\}\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\+\\nabla\_\{\\bm\{x\}\}\\cdot\(\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)u^\{1\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\)\+\\nabla\_\{\\bm\{y\}\}\\cdot\(\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)u^\{2\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\)=0\.With this condition, one can prove the inverse inequalityQOTS​\(μ0,μ1\)≤QOTC​\(μ0,μ1\)\\text\{QOT\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)\\leq\\text\{QOT\}\_\{C\}\(\\mu\_\{0\},\\mu\_\{1\}\)\.

###### Proof\.

For any conditional path\(𝒙\[0,1\],𝒚\[0,1\]\)\(\\bm\{x\}\_\{\[0,1\]\},\\bm\{y\}\_\{\[0,1\]\}\), let𝒛=\(𝒙0,𝒙1\),𝒛′=\(𝒚0,𝒚1\)\\bm\{z\}=\(\\bm\{x\}\_\{0\},\\bm\{x\}\_\{1\}\),\\bm\{z\}^\{\\prime\}=\(\\bm\{y\}\_\{0\},\\bm\{y\}\_\{1\}\), we have

∫01ℒ⁡\(t,𝒙t,𝒚t,𝒙˙t,𝒚˙t\)​𝑑t≥𝒜⁡\(𝒛,𝒛′\)\\int\_\{0\}^\{1\}\\mathcal\{L\}\(t,\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\},\\dot\{\\bm\{x\}\}\_\{t\},\\dot\{\\bm\{y\}\}\_\{t\}\)\\mathrm\{d\}t\\geq\\mathcal\{A\}\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\(133\)by definition\. Now consider the projection ofη\\etaon the two time points\{0,1\}\\\{0,1\\\}\. Since\(e0,e1\)\#​η=π0,1\(e\_\{0\},e\_\{1\}\)\_\{\\\#\}\\eta=\\pi\_\{0,1\}, there exists a couplingγ∈Π⁡\(μ0,μ1\)\\gamma\\in\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)such that\(e0,e1\)\#​η=γ⊗γ\(e\_\{0\},e\_\{1\}\)\_\{\\\#\}\\eta=\\gamma\\otimes\\gamma\. By \([133](https://arxiv.org/html/2609.20008#A6.E133)\), we have

∫ℒ⁡\(t,𝒙,𝒚,𝒖t1​\(𝒙,𝒚\),𝒖t2​\(𝒙,𝒚\)\)​πt​\(𝒙,𝒚\)​𝑑𝒙​𝑑𝒚​𝑑t\\displaystyle\\int\\mathcal\{L\}\(t,\\bm\{x\},\\bm\{y\},\\bm\{u\}\_\{t\}^\{1\}\(\\bm\{x\},\\bm\{y\}\),\\bm\{u\}\_\{t\}^\{2\}\(\\bm\{x\},\\bm\{y\}\)\)\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{y\}\\mathrm\{d\}t\(134\)=∫ℒ⁡\(t,𝒙t,𝒚t,𝒖t1​\(𝒙t,𝒚t\),𝒖t2​\(𝒙t,𝒚t\)\)​𝑑η​\(𝒙\[0,1\],𝒚\[0,1\]\)​𝑑t\\displaystyle=\\int\\mathcal\{L\}\(t,\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\},\\bm\{u\}\_\{t\}^\{1\}\(\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\}\),\\bm\{u\}\_\{t\}^\{2\}\(\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\}\)\)\\mathrm\{d\}\\eta\(\\bm\{x\}\_\{\[0,1\]\},\\bm\{y\}\_\{\[0,1\]\}\)\\mathrm\{d\}t=∫ℒ⁡\(t,𝒙t,𝒚t,𝒙˙t,𝒚˙t\)​𝑑η​\(𝒙\[0,1\],𝒚\[0,1\]\)​𝑑t\\displaystyle=\\int\\mathcal\{L\}\(t,\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\},\\dot\{\\bm\{x\}\}\_\{t\},\\dot\{\\bm\{y\}\}\_\{t\}\)\\mathrm\{d\}\\eta\(\\bm\{x\}\_\{\[0,1\]\},\\bm\{y\}\_\{\[0,1\]\}\)\\mathrm\{d\}t=∫\(∫01ℒ⁡\(t,𝒙t,𝒚t,𝒙˙t,𝒚˙t\)​dt\)​dη​\(𝒙\[0,1\],𝒚\[0,1\]\)\\displaystyle=\\int\\Big\(\\int\_\{0\}^\{1\}\\mathcal\{L\}\(t,\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\},\\dot\{\\bm\{x\}\}\_\{t\},\\dot\{\\bm\{y\}\}\_\{t\}\)\\mathrm\{d\}t\\Big\)\\mathrm\{d\}\\eta\(\\bm\{x\}\_\{\[0,1\]\},\\bm\{y\}\_\{\[0,1\]\}\)=∫\(∫01ℒ⁡\(t,𝒙t,𝒚t,𝒙˙t,𝒚˙t\)​dt\)​d​\(e0,e1\)\#​η​\(𝒙\[0,1\],𝒚\[0,1\]\)\\displaystyle=\\int\\Big\(\\int\_\{0\}^\{1\}\\mathcal\{L\}\(t,\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\},\\dot\{\\bm\{x\}\}\_\{t\},\\dot\{\\bm\{y\}\}\_\{t\}\)\\mathrm\{d\}t\\Big\)\\mathrm\{d\}\(e\_\{0\},e\_\{1\}\)\_\{\\\#\}\\eta\(\\bm\{x\}\_\{\[0,1\]\},\\bm\{y\}\_\{\[0,1\]\}\)=∫\(∫01ℒ⁡\(t,𝒙t,𝒚t,𝒙˙t,𝒚˙t\)​𝑑t\)​γ​\(𝒛\)​γ​\(𝒛′\)​𝑑𝒛​d​𝒛′\\displaystyle=\\int\\Big\(\\int\_\{0\}^\{1\}\\mathcal\{L\}\(t,\\bm\{x\}\_\{t\},\\bm\{y\}\_\{t\},\\dot\{\\bm\{x\}\}\_\{t\},\\dot\{\\bm\{y\}\}\_\{t\}\)\\mathrm\{d\}t\\Big\)\\gamma\(\\bm\{z\}\)\\gamma\(\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}≥∫𝒜⁡\(𝒛,𝒛′\)​γ​\(𝒛\)​γ​\(𝒛′\)​𝑑𝒛​d​𝒛′\\displaystyle\\geq\\int\\mathcal\{A\}\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\gamma\(\\bm\{z\}\)\\gamma\(\\bm\{z\}^\{\\prime\}\)\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}≥QOTS​\(μ0,μ1\)\\displaystyle\\geq\\text\{QOT\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)By taking infimum w\.r\.tπt​\(𝒙,𝒚\)\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\), we haveQOTC​\(μ0,μ1\)≥QOTS​\(μ0,μ1\)\\text\{QOT\}\_\{C\}\(\\mu\_\{0\},\\mu\_\{1\}\)\\geq\\text\{QOT\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)\. ∎

Although we can prove the inverse inequality, unfortunately, the original side cannot be guaranteed in general under this condition\. That is, we no longer haveQOTC​\(μ0,μ1\)≤QOTS​\(μ0,μ1\)\\text\{QOT\}\_\{C\}\(\\mu\_\{0\},\\mu\_\{1\}\)\\leq\\text\{QOT\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)\. The reason is that, we proved the original inequality by Jensen’s inequality in[A\.1](https://arxiv.org/html/2609.20008#A1.SS1)\. To use Jensen’s inequality, we marginalize all the travelling pair conditional velocities into a marginal velocity pair

𝒖ti\(𝒙,𝒚\)=∫𝒖ti\(𝒙,𝒚\|𝒛,𝒛′\)πt\(𝒙,𝒚\|𝒛,𝒛′\)q\(𝒛,𝒛′\)πt​\(𝒙,𝒚\)d𝒛d𝒛′\\bm\{u\}^\{i\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\)=\\int\\bm\{u\}^\{i\}\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\\frac\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\|\\bm\{z\},\\bm\{z\}^\{\\prime\}\)q\(\\bm\{z\},\\bm\{z\}^\{\\prime\}\)\}\{\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)\}\\mathrm\{d\}\\bm\{z\}\\mathrm\{d\}\\bm\{z\}^\{\\prime\}\(135\)whereq=γ⊗γq=\\gamma\\otimes\\gamma\. The point here is that althoughqqhas the tensor structure, the probability flowπt​\(𝒙,𝒚\)\\pi\_\{t\}\(\\bm\{x\},\\bm\{y\}\)generated by the marginal velocity pair does not naturally satisfy the QOT compatible condition\. Therefore, we can only proveQOTS​\(μ0,μ1\)≥QOTB​B​\(μ0,μ1\)\\text\{QOT\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)\\geq\\text\{QOT\}\_\{BB\}\(\\mu\_\{0\},\\mu\_\{1\}\), but can never proveQOTS​\(μ0,μ1\)≥QOTC​\(μ0,μ1\)\\text\{QOT\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)\\geq\\text\{QOT\}\_\{C\}\(\\mu\_\{0\},\\mu\_\{1\}\)\. The optimal result we can obtain is thereforeQOTC​\(μ0,μ1\)≥QOTS​\(μ0,μ1\)≥QOTB​B​\(μ0,μ1\)\\text\{QOT\}\_\{C\}\(\\mu\_\{0\},\\mu\_\{1\}\)\\geq\\text\{QOT\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)\\geq\\text\{QOT\}\_\{BB\}\(\\mu\_\{0\},\\mu\_\{1\}\)\.

### F\.3The intuitive relation between the three forms

In the main text, we define the static QOT, the dynamic QOT and the BB\-form\. The intuition of the BB\-form is to minimize the population level action induced by the Lagrangian during the transport, which is what, in practice, one may actually want to minimize\. SinceQOTD​\(μ0,μ1\)=QOTS​\(μ0,μ1\)≥QOTB​B​\(μ0,μ1\)\\text\{QOT\}\_\{D\}\(\\mu\_\{0\},\\mu\_\{1\}\)=\\text\{QOT\}\_\{S\}\(\\mu\_\{0\},\\mu\_\{1\}\)\\geq\\text\{QOT\}\_\{BB\}\(\\mu\_\{0\},\\mu\_\{1\}\), minimizing the static QOT action is equivalent to minimizing a tractable upper bound of the BB\-form\. Intuitively, it also implicitly tries to minimize the BB\-form action\. But, the static form contains no dynamic information\. To model the continuous time dynamics, we have to extend it to the dynamic form, which is equivalent to the static form\. In the main text, we introduce flow matching to solve the dynamic QOT flow by obtaining a Markovian projection\. By Jensen’s inequality, the BB\-form action induced by this marginal flow is less than the corresponding static action, but greater than the infimum BB\-form action by definition\. Therefore, the learned flow can be viewed as a finer upper bound approximation of the true BB\-form action\.

Similar Articles

Diffusion enabled Optimal Transport distances for graph matching

arXiv cs.LG

Introduces Diffusion Semi-Relaxed Fused Gromov-Wasserstein (DsrFGW), a novel method that integrates node features and structural connectivity for graph comparison via optimal transport and diffusion processes, demonstrating improved robustness to noise and missing edges on synthetic tasks.

Multimarginal flow matching with optimal transport potentials

arXiv cs.LG

Proposes OTP-FM, a novel method for multimarginal flow matching that uses optimal transport potentials to softly steer flows through intermediate marginals, achieving state-of-the-art performance on single-cell RNA sequencing, oceanographic, and meteorological datasets.

Geometry-Aware Tabular Diffusion

arXiv cs.LG

Introduces Geometry-Aware Tabular Diffusion (GATD), which augments tabular diffusion denoisers with explicit pairwise geometric features. Achieves state-of-the-art performance on ten benchmarks while using significantly fewer parameters.