TracingFlow: A Simulation-Free Trajectory Inference Framework Based on Second-Order Dynamics
Summary
TracingFlow is a simulation-free framework using second-order dynamics for trajectory inference, improving accuracy in capturing high-curvature transitions in single-cell omics data.
View Cached Full Text
Cached at: 08/24/26, 04:36 AM
# A Simulation-Free Trajectory Inference FrameworkBased on Second-Order Dynamics
Source: [https://arxiv.org/html/2608.21070](https://arxiv.org/html/2608.21070)
## TracingFlow: A Simulation\-Free Trajectory Inference Framework Based on Second\-Order Dynamics
Yuhao SunZekun Wu11footnotemark:1Affiliation:School of Mathematical Sciences, Peking UniversityZixun HuangAffiliation:School of Mathematical Sciences, Peking UniversityPeijie ZhouThanks:Corresponding author:pjzhou@pku\.edu\.cnAffiliation:Center for Machine Learning Research, Peking UniversityAffiliation:Center for Quantitative Biology, Peking UniversityAffiliation:National Engineering Laboratory for Big Data Analysis and Applications, BeijingAffiliation:AI for Science Institute, Beijing
###### Abstract
Inferring continuous system evolution from sparse temporal snapshots is a key challenge in generative modeling and single\-cell omics\. While Optimal Transport \(OT\) is popular, existing frameworks are largely restricted to first\-order dynamics, assuming memoryless velocity fields\. This limits expressiveness, as first\-order systems fail to account for regulatory momentum and time\-delayed responses inherent in processes like cell differentiation\. Here, we introduce TracingFlow, a simulation\-free Flow Matching framework generalizing to second\-order dynamics\. By using neural networks to regress the acceleration field, TracingFlow provides an exact, efficient solution to the Dynamical Optimal Acceleration Transport \(DOAT\) problem\. Unlike first\-order methods yielding over\-smoothed trajectories, our second\-order formulation captures high\-curvature transitions and nonlinear evolutions by learning the underlying force fields\. Evaluated on complex synthetic and large\-scale scRNA\-seq datasets, TracingFlow achieves superior accuracy in distributional reconstruction and trajectory faithfulness\. Moreover, by integrating lineage tracing priors, it recovers dynamical structures that are both mathematically optimal and biologically plausible\.
โ โ footnotetext:Preprint\. August 21, 2026\.## 1Introduction
Recovering underlying dynamics from discrete observations is a pivotal task in single\-cell omics, known as Trajectory Inference \(TI\)\[[9](https://arxiv.org/html/2608.21070#bib.bib9),[72](https://arxiv.org/html/2608.21070#bib.bib72)\]\. To identify the least costly evolution between distributions, Optimal Transport \(OT\) frameworks have emerged\[[7](https://arxiv.org/html/2608.21070#bib.bib7),[71](https://arxiv.org/html/2608.21070#bib.bib71)\]\. These generally categorize into Static OT methods\[[51](https://arxiv.org/html/2608.21070#bib.bib51),[28](https://arxiv.org/html/2608.21070#bib.bib28),[22](https://arxiv.org/html/2608.21070#bib.bib22)\], which learn direct mappings, and Dynamical OT methods\[[57](https://arxiv.org/html/2608.21070#bib.bib57),[24](https://arxiv.org/html/2608.21070#bib.bib24),[67](https://arxiv.org/html/2608.21070#bib.bib67),[44](https://arxiv.org/html/2608.21070#bib.bib44)\], which capture continuous evolution via flow maps\.
Early Dynamical OT frameworks relied on Neural ODEs\[[9](https://arxiv.org/html/2608.21070#bib.bib9)\], incurring high computational costs due to numerical integration\. To address this, Flow Matching\[[35](https://arxiv.org/html/2608.21070#bib.bib35),[36](https://arxiv.org/html/2608.21070#bib.bib36),[48](https://arxiv.org/html/2608.21070#bib.bib48)\]was proposed as a simulation\-free alternative, regressing velocity fields to push source distributions to targets\. By decomposing transport costs into single\-particle trajectories, Flow Matching efficiently solves Dynamical OT problems\[[58](https://arxiv.org/html/2608.21070#bib.bib58),[27](https://arxiv.org/html/2608.21070#bib.bib27),[50](https://arxiv.org/html/2608.21070#bib.bib50)\], including its various variants\[[59](https://arxiv.org/html/2608.21070#bib.bib59),[15](https://arxiv.org/html/2608.21070#bib.bib15),[8](https://arxiv.org/html/2608.21070#bib.bib8),[45](https://arxiv.org/html/2608.21070#bib.bib45)\]\.
However, most TI frameworks assume first\-order dynamics \(๐ห=๐โก\(๐,t\)\\bm\{\\dot\{x\}\}=\\bm\{v\}\(\\bm\{x\},t\)\), inherently restricting๐\\bm\{v\}to be single\-valued w\.r\.t๐\\bm\{x\}\. This limits expressiveness when modeling complex biological priors, such as lineage tracing\[[30](https://arxiv.org/html/2608.21070#bib.bib30),[38](https://arxiv.org/html/2608.21070#bib.bib38)\], potentially yielding dynamics that contradict with additional biological prior information\. Furthermore,\[[19](https://arxiv.org/html/2608.21070#bib.bib19),[34](https://arxiv.org/html/2608.21070#bib.bib34),[23](https://arxiv.org/html/2608.21070#bib.bib23)\]pointed out that if more modalities such as chromatin accessibility or proteomics are considered, the biological kinetics may follow higher\-order dynamics or could be modelled in augmented space\. While 3MSBM\[[56](https://arxiv.org/html/2608.21070#bib.bib56)\]recently introduced a second\-order Momentum Schrรถdinger Bridge for smoother trajectories, it relies on iterative retraining the acceleration field similar to Rectified Flow\[[36](https://arxiv.org/html/2608.21070#bib.bib36)\]to determine optimal couplings\. Furthermore, it employs a heuristic method to estimate the initial velocity, rather than incorporating this estimation as part of the optimal control problem\.
To overcome these limitations, we introduceTracingFlow, a simulation\-free framework utilizing second\-order dynamical systems\. We propose the Dynamical Optimal Acceleration Transport \(DOAT\) problem, which determines evolution by minimizing acceleration costs\. TracingFlow directly regresses the acceleration field and initial velocity, avoiding ODE simulation\. Our experiments show that TracingFlow achieves competitive and often improved reconstruction accuracy relative to existing simulation\-free frameworks\[[58](https://arxiv.org/html/2608.21070#bib.bib58),[59](https://arxiv.org/html/2608.21070#bib.bib59),[69](https://arxiv.org/html/2608.21070#bib.bib69),[42](https://arxiv.org/html/2608.21070#bib.bib42)\], while effectively incorporating biological priors\. Our contributions include:
- โขWe propose TracingFlow, a simulation\-free TI framework based on second\-order dynamics\. It solves the proposed DOAT problem by regressing acceleration fields, reducing computational costs compared to ODE\-based methods\.
- โขWe provide theoretical guarantees by decoupling the design of the optimal transport plan and single\-particle control\. We further introduce an iterative strategy to transform the real\-world Velocity Missed\-DOAT \(VM\-DOAT\) problem into a standard DOAT problem, enabling approximate solutions\.
- โขWe demonstrate TracingFlowโs effectiveness on multi\-time\-point real world datasets\. Results indicate superior distribution reconstruction accuracy and enhanced capability in preserving biological priors compared to state\-of\-the\-art methods\.
Figure 1:An illustration figure for TracingFlow
## 2Related Works
#### Solving Optimal Transport via Flow Matching\.
Flow Matching\[[35](https://arxiv.org/html/2608.21070#bib.bib35),[36](https://arxiv.org/html/2608.21070#bib.bib36),[48](https://arxiv.org/html/2608.21070#bib.bib48),[1](https://arxiv.org/html/2608.21070#bib.bib1)\]transports probability distributions via learned flow maps, offering scalable simulation\-free training\. As an optimal control problem, Optimal Transport \(OT\)\[[47](https://arxiv.org/html/2608.21070#bib.bib47)\]decomposes into Optimal Coupling and single\-particle geodesics, facilitating solutions via Flow Matching\[[58](https://arxiv.org/html/2608.21070#bib.bib58),[27](https://arxiv.org/html/2608.21070#bib.bib27),[50](https://arxiv.org/html/2608.21070#bib.bib50),[15](https://arxiv.org/html/2608.21070#bib.bib15)\]\. This approach extends to various variants and generalizations\[[59](https://arxiv.org/html/2608.21070#bib.bib59),[15](https://arxiv.org/html/2608.21070#bib.bib15),[8](https://arxiv.org/html/2608.21070#bib.bib8),[13](https://arxiv.org/html/2608.21070#bib.bib13),[61](https://arxiv.org/html/2608.21070#bib.bib61),[45](https://arxiv.org/html/2608.21070#bib.bib45),[26](https://arxiv.org/html/2608.21070#bib.bib26),[68](https://arxiv.org/html/2608.21070#bib.bib68),[4](https://arxiv.org/html/2608.21070#bib.bib4),[46](https://arxiv.org/html/2608.21070#bib.bib46)\]\. However, these typically employ first\-order dynamics\. While second\-order approaches have emerged\[[56](https://arxiv.org/html/2608.21070#bib.bib56)\], they require iterative training of the acceleration field to achieve optimal coupling\. Moreover, they rely on heuristic methods for initial velocity estimation, rather than integrating it as a component of the optimal control problem\. TracingFlow addresses this by decomposing the Dynamical Acceleration Optimal Transport \(DOAT\) problem into optimal coupling and single particle optimal trajectory, achieving a truly simulation\-free process for DOAT\.
#### Overcoming Single\-Valued Velocity Constraint in Flow Matching\.
In standard Flow Matching, the velocity field is modeled as a function solely dependent on๐\\bm\{x\}andtt\. This single\-valued dependency prevents the model from representing intersecting trajectories with distinct velocities, often resulting in curved inference paths and increased computational costs\. To decouple the velocity from the current position, the velocity network can incorporate additional inputs like class labels\[[73](https://arxiv.org/html/2608.21070#bib.bib73),[54](https://arxiv.org/html/2608.21070#bib.bib54)\], initial positions\[[14](https://arxiv.org/html/2608.21070#bib.bib14)\], or hidden states\[[20](https://arxiv.org/html/2608.21070#bib.bib20)\]\. Additionally,\[[69](https://arxiv.org/html/2608.21070#bib.bib69)\]employs hierarchical generation to mitigate this issue\. However, these methods do not extend to second\-order dynamics\. While\[[42](https://arxiv.org/html/2608.21070#bib.bib42)\]utilized a second\-order system, it is limited to constant\-acceleration paths\. In contrast, TracingFlow offers a flexible second\-order setting\. By designing minimal\-cost paths, it effectively resolves the limitation of the single\-valued velocity field\.
#### Single\-cell Trajectory Inference and Lineage Integration\.
Optimal Transport \(OT\) is a robust framework for inferring cellular dynamics from scRNA\-seq data\[[51](https://arxiv.org/html/2608.21070#bib.bib51),[28](https://arxiv.org/html/2608.21070#bib.bib28),[57](https://arxiv.org/html/2608.21070#bib.bib57),[24](https://arxiv.org/html/2608.21070#bib.bib24),[40](https://arxiv.org/html/2608.21070#bib.bib40),[41](https://arxiv.org/html/2608.21070#bib.bib41),[2](https://arxiv.org/html/2608.21070#bib.bib2),[74](https://arxiv.org/html/2608.21070#bib.bib74),[37](https://arxiv.org/html/2608.21070#bib.bib37),[66](https://arxiv.org/html/2608.21070#bib.bib66),[12](https://arxiv.org/html/2608.21070#bib.bib12),[60](https://arxiv.org/html/2608.21070#bib.bib60),[33](https://arxiv.org/html/2608.21070#bib.bib33),[53](https://arxiv.org/html/2608.21070#bib.bib53),[29](https://arxiv.org/html/2608.21070#bib.bib29),[6](https://arxiv.org/html/2608.21070#bib.bib6),[10](https://arxiv.org/html/2608.21070#bib.bib10),[25](https://arxiv.org/html/2608.21070#bib.bib25),[70](https://arxiv.org/html/2608.21070#bib.bib70),[55](https://arxiv.org/html/2608.21070#bib.bib55),[44](https://arxiv.org/html/2608.21070#bib.bib44),[65](https://arxiv.org/html/2608.21070#bib.bib65),[52](https://arxiv.org/html/2608.21070#bib.bib52),[11](https://arxiv.org/html/2608.21070#bib.bib11)\]\. However, transcriptomic similarity alone often fails to resolve complex trajectories\. To mitigate this, studies incorporate lineage tracing priors such as clonal barcode information\. Existing approaches integrate such info via regularization\[[17](https://arxiv.org/html/2608.21070#bib.bib17),[49](https://arxiv.org/html/2608.21070#bib.bib49)\], structural alignment\[[32](https://arxiv.org/html/2608.21070#bib.bib32)\], velocity mapping\[[62](https://arxiv.org/html/2608.21070#bib.bib62)\], or sparse transition modeling\[[63](https://arxiv.org/html/2608.21070#bib.bib63),[21](https://arxiv.org/html/2608.21070#bib.bib21),[18](https://arxiv.org/html/2608.21070#bib.bib18)\]\. While improving accuracy, these generally focus on discrete couplings or static matrices\. TracingFlow advances this by embedding priors into a continuous OT\-based second\-order flow matching framework, capturing complex dynamics consistent with lineage information\.
## 3Mathematical Background
In this section, we provide the mathematical formulation of the problem addressed by TracingFlow\.
#### Dynamical Optimal Acceleration Transport \(DOAT\) Problem
Let๐ณโโd\\mathcal\{X\}\\subset\\mathbb\{R\}^\{d\}and๐ฑโโd\\mathcal\{V\}\\subset\\mathbb\{R\}^\{d\}denote the*position*and*velocity*spaces, respectively\. We define the*augmented space*as๐ฎโ๐ณร๐ฑ\\mathcal\{S\}\\coloneqq\\mathcal\{X\}\\times\\mathcal\{V\}, so that the full system state is๐=\(๐,๐\)โ๐ฎ\\bm\{s\}=\(\\bm\{x\},\\bm\{v\}\)\\in\\mathcal\{S\}\. Consider the second\-order dynamic system:
๐ห=๐๐ห=๐โก\(๐,๐,t\),tโ\[t0,tK\]\.\\dot\{\\bm\{x\}\}=\\bm\{v\}\\quad\\dot\{\\bm\{v\}\}=\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\),\\quad t\\in\[t\_\{0\},t\_\{K\}\]\.\(1\)Modeling a large ensemble of such particles via a probability densityฯt\\rho\_\{t\}in๐ฎ\\mathcal\{S\}, we follow\[[5](https://arxiv.org/html/2608.21070#bib.bib5),[11](https://arxiv.org/html/2608.21070#bib.bib11)\]to formulate the following optimal control problem:
min๐,ฯโก๐ฅDOATโ\(๐,ฯ\)=โซt0tKโซ๐ฎ12โโ๐โก\(๐,๐,t\)โ2โฯtโ\(๐,๐\)โ๐๐โ๐๐โ๐t\\small\{\\min\_\{\\bm\{a\},\\rho\}\\mathcal\{J\}\_\{\\text\{DOAT\}\}\(\\bm\{a\},\\rho\)=\\int\_\{t\_\{0\}\}^\{t\_\{K\}\}\\int\_\{\\mathcal\{S\}\}\\frac\{1\}\{2\}\\\|\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\)\\\|^\{2\}\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\\,\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{v\}\\mathrm\{d\}t\}\(2\)s\.t\.โtฯt\+โ๐โ
\(๐โฯt\)\+โ๐โ
\(๐โฯt\)=0,\\displaystyle\\partial\_\{t\}\\rho\_\{t\}\+\\nabla\_\{\\bm\{x\}\}\\cdot\(\\bm\{v\}\\rho\_\{t\}\)\+\\nabla\_\{\\bm\{v\}\}\\cdot\(\\bm\{a\}\\rho\_\{t\}\)=0,\(3\)ฯti\(๐,๐\)=ฮผi\(๐,๐\),i=0,โฆ,K\.\\displaystyle\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\)=\\mu\_\{i\}\(\\bm\{x\},\\bm\{v\}\),\\quad i=0,\\dots,K\.\(4\)Here, constraints are imposed atK\+1K\+1timestampstit\_\{i\}with given densitiesฮผi\\mu\_\{i\}\.[Equation3](https://arxiv.org/html/2608.21070#S3.E3)represents the continuity equation in the augmented space\. Analogous to standard Dynamical Optimal Transport\[[47](https://arxiv.org/html/2608.21070#bib.bib47)\], we term this the*Dynamical Optimal Acceleration Transport \(DOAT\)*problem\. To facilitate the subsequent numerical solution via particle methods, we assume the existence of an optimal flow map\.
#### Velocity\-Missed Dynamical Optimal Acceleration Transport \(VM\-DOAT\) Problem
However, in practical applications such as single\-cell datasets, only particle positions are observable, rendering velocities as unknown quantities\. This presents a discrepancy with the DOAT formulation\.
Consequently, TracingFlow addresses a relaxation of the problem:
min๐,ฯโก๐ฅVM\-DOATโ\(๐,ฯ\)=โซt0tKโซ๐ฎ12โโ๐โก\(๐,๐,t\)โ2โฯtโ\(๐,๐\)โ๐๐โ๐๐โ๐t\\small\{\\min\_\{\\bm\{a\},\\rho\}\\mathcal\{J\}\_\{\\text\{VM\-DOAT\}\}\(\\bm\{a\},\\rho\)=\\int\_\{t\_\{0\}\}^\{t\_\{K\}\}\\int\_\{\\mathcal\{S\}\}\\dfrac\{1\}\{2\}\\\|\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\)\\\|^\{2\}\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\\,\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{v\}\\mathrm\{d\}t\}\(5\)s\.t\.โtฯtโ\(๐,๐\)\+โ๐โ
\(๐โฯt\)\+โ๐โ
\(๐โฯt\)=0,\\displaystyle\\quad\\partial\_\{t\}\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\+\\nabla\_\{\\bm\{x\}\}\\cdot\(\\bm\{v\}\\rho\_\{t\}\)\+\\nabla\_\{\\bm\{v\}\}\\cdot\(\\bm\{a\}\\rho\_\{t\}\)=0,\(6\)โซฯti\(๐,๐\)d๐=ฮผi\(pos\)\(๐\),i=0,1,โฆ,K\.\\displaystyle\\int\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\)\\,\\mathrm\{d\}\\bm\{v\}=\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\),\\ i=0,1,\\dots,K\.\(7\)Here, the constraints are imposed only on the marginal distribution of positions,โซฯtjโ\(๐,๐\)โ๐๐\\int\\rho\_\{t\_\{j\}\}\(\\bm\{x\},\\bm\{v\}\)\\,\\mathrm\{d\}\\bm\{v\}, at each time point, with no constraints on the velocity distribution\. We refer to this problem as*Velocity Missed\-Dynamical Optimal Acceleration Transport \(VM\-DOAT\)*\.
## 4Methodology of TracingFlow
In this section, we first introduce the key properties of the DOAT problem that are essential for our algorithm design\. Subsequently, drawing inspiration from Flow Matching approaches for standard Dynamical OT, we devise an algorithm to exactly solve the DOAT problem by combining Optimal Coupling with single\-particle optimal trajectories\. Finally, we introduce a global path\-based velocity completion scheme that effectively bridges the gap between VM\-DOAT and the standard DOAT framework\. By inferring missing velocities through the lens of global trajectory consistency, our approach enables the application of optimal transport theory to solve the VM\-DOAT problem in a principled manner\.
### 4\.1Key Properties of DOAT Problem
#### Optimal Single Particle Trajectory
The DOAT problem is an optimal control problem over distributions\. We begin by considering the corresponding single\-particle optimal control problem\. Specifically, the optimal trajectory is a*cubic interpolation*determined by the boundary states\.
###### Proposition 4\.1\.
Letฮณ:\[0,T\]โโd\\gamma:\[0,T\]\\to\\mathbb\{R\}^\{d\}be a twice continuously differentiable curve satisfying the boundary conditionsฮณโก\(0\)=๐ฑ0\\gamma\(0\)=\\bm\{x\}\_\{0\},ฮณโฒโ\(0\)=๐ฏ0\\gamma^\{\\prime\}\(0\)=\\bm\{v\}\_\{0\},ฮณโก\(T\)=๐ฑT\\gamma\(T\)=\\bm\{x\}\_\{T\}, andฮณโฒโ\(T\)=๐ฏT\\gamma^\{\\prime\}\(T\)=\\bm\{v\}\_\{T\}\. The variational problem
minโกโซ0Tฮณโก12โโฮณโฒโฒโ\(t\)โ22โ๐t\\min\_\{\\gamma\}\\int\_\{0\}^\{T\}\\frac\{1\}\{2\}\\\|\\gamma^\{\\prime\\prime\}\(t\)\\\|\_\{2\}^\{2\}\\,\\mathrm\{d\}t\(8\)admits a unique minimizer, which is a cubic polynomial of the form
ฮณโก\(t\)=c3โt3\+c2โt2\+c1โt\+c0,\\gamma\(t\)=c\_\{3\}t^\{3\}\+c\_\{2\}t^\{2\}\+c\_\{1\}t\+c\_\{0\},\(9\)where the coefficients are detailed in[SectionA\.1](https://arxiv.org/html/2608.21070#A1.SS1)\.This result is got by*Euler\-Lagrange*formula\. See a detailed proof in[SectionA\.1](https://arxiv.org/html/2608.21070#A1.SS1)\.
#### Optimal Coupling
By utilizing the single\-particle optimal trajectory, we can immediately obtain the minimum cost for transporting a particle from\(๐0,๐0\)\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\)to\(๐T,๐T\)\(\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\):
###### Corollary 4\.2\.
Substitute[Equation9](https://arxiv.org/html/2608.21070#S4.E9)to[Equation8](https://arxiv.org/html/2608.21070#S4.E8), we get the minimum cost :
๐\[0,T\]\\displaystyle\\mathcal\{C\}\_\{\[0,T\]\}=2T3โ\{\(โ๐0โ2\+โจ๐0,๐Tโฉ\+โ๐Tโ2\)โT2\+3โโจ๐0\+๐T,๐0โ๐TโฉโT\+3โโ๐0โ๐Tโ2\}\\displaystyle=\\frac\{2\}\{T^\{3\}\}\\bigg\\\{\(\\\|\\bm\{v\}\_\{0\}\\\|^\{2\}\+\\langle\\bm\{v\}\_\{0\},\\bm\{v\}\_\{T\}\\rangle\+\\\|\\bm\{v\}\_\{T\}\\\|^\{2\}\)T^\{2\}\+3\\langle\\bm\{v\}\_\{0\}\+\\bm\{v\}\_\{T\},\\bm\{x\}\_\{0\}\-\\bm\{x\}\_\{T\}\\rangle T\+3\\\|\\bm\{x\}\_\{0\}\-\\bm\{x\}\_\{T\}\\\|^\{2\}\\bigg\\\}\(10\)
The minimum cost allows us to define an optimal transport problem betweent=0t=0andt=Tt=T:
ฯ\[0,T\]โ=minโกโซฯ\[0,T\]โก๐\[0,T\]โ\[\(๐0,๐0\),\(๐T,๐T\)\]โdโฯ\[0,T\]\\pi^\{\*\}\_\{\[0,T\]\}=\\min\_\{\\pi\_\{\[0,T\]\}\}\\int\\mathcal\{C\}\_\{\[0,T\]\}\[\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\),\(\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\)\]\\mathrm\{d\}\\pi\_\{\[0,T\]\}\(11\)whereฯ\[0,T\]โ\[\(๐0,๐0\),\(๐T,๐T\)\]\\pi\_\{\[0,T\]\}\[\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\),\(\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\)\]is any coupling satisfying the constraintsโซฯ\[0,T\]โdโ๐0โdโ๐0=ฮผTโ\(๐,๐\)\\int\\pi\_\{\[0,T\]\}\\mathrm\{d\}\\bm\{x\}\_\{0\}\\mathrm\{d\}\\bm\{v\}\_\{0\}=\\mu\_\{T\}\(\\bm\{x\},\\bm\{v\}\)andโซฯ\[0,T\]โdโ๐Tโdโ๐T=ฮผ0โ\(๐,๐\)\\int\\pi\_\{\[0,T\]\}\\mathrm\{d\}\\bm\{x\}\_\{T\}\\mathrm\{d\}\\bm\{v\}\_\{T\}=\\mu\_\{0\}\(\\bm\{x\},\\bm\{v\}\),ฮผ0โ\(๐,๐\)\\mu\_\{0\}\(\\bm\{x\},\\bm\{v\}\)andฮผTโ\(๐,๐\)\\mu\_\{T\}\(\\bm\{x\},\\bm\{v\}\)are two distributions in augmented space\. We term this formulation theStatic Optimal Acceleration Transport \(SOAT\)problem\.
It is worth noting that while DOAT permits crossings in the projected position space, it prohibits trajectory crossing in the augmented space, which guarentees the neatness of trajectories\. The following remark provides the mathematical justification for this non\-crossing property in the augmented space, demonstrating that any such intersection implies a strictly higher transport cost\.
[Equation2](https://arxiv.org/html/2608.21070#S3.E2)represents a multi\-marginal optimal transport problem over the interval\[t0,tK\]\[t\_\{0\},t\_\{K\}\]\. To address this, we formulate sequential SOAT problems between adjacent time stepstit\_\{i\}andti\+1t\_\{i\+1\}by applying the substitutions๐0โ๐i\\bm\{x\}\_\{0\}\\rightarrow\\bm\{x\}\_\{i\},๐Tโ๐i\+1\\bm\{x\}\_\{T\}\\rightarrow\\bm\{x\}\_\{i\+1\},๐0โ๐i\\bm\{v\}\_\{0\}\\rightarrow\\bm\{v\}\_\{i\},๐Tโ๐i\+1\\bm\{v\}\_\{T\}\\rightarrow\\bm\{v\}\_\{i\+1\}, andTโti\+1โtiT\\rightarrow t\_\{i\+1\}\-t\_\{i\}to[Equation11](https://arxiv.org/html/2608.21070#S4.E11)\. For simplicity, let๐iโi\+1\\mathcal\{C\}\_\{i\\rightarrow i\+1\}denote the resulting pairwise cost andฯiโi\+1โ\\pi^\{\*\}\_\{i\\rightarrow i\+1\}the optimal plan\. Due to the Markovian nature of the dynamics, the full multi\-marginal plan decomposes into the sequence of these local plans\{ฯiโi\+1โ\}\\\{\{\\pi^\{\*\}\_\{i\\rightarrow i\+1\}\}\\\}fori=0,โฆ,Kโ1i=0,\\dots,K\-1\.
### 4\.2Conditional and Marginal Probability Path
#### Obtaining Marginal Probability Path via Mixing Conditional Probability Paths
Similar to the approach in standard Flow Matching\[[35](https://arxiv.org/html/2608.21070#bib.bib35),[58](https://arxiv.org/html/2608.21070#bib.bib58)\], directly constructing the Marginal Probability Path is difficult\. Therefore, we condition on๐\\bm\{z\}and define the Marginal Probability Path as a mixture of Conditional Probability Paths:
ฯtโ\(๐,๐\)=โซฯtโ\(๐,๐\|๐\)โqโ\(๐\)โ๐๐\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)=\\int\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\)q\(\\bm\{z\}\)\\mathrm\{d\}\\bm\{z\}\(12\)The condition๐\\bm\{z\}consists ofK\+1K\+1points selected respectively from the snapshot data atK\+1K\+1time pointst=0,1,โฏ,Kt=0,1,\\cdots,K\. Ifฯtโ\(๐,๐\|๐\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\)is generated by a conditional acceleration field๐โก\(๐,๐,t\|๐\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\|\\bm\{z\}\)starting from the initial conditionฯ0โ\(๐,๐\|๐\)\\rho\_\{0\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\), then the marginal acceleration field:
๐โก\(๐,๐,t\):=๐ผqโก\(๐\)โ\{๐โก\(๐,๐,t\|๐\)โฯtโ\(๐,๐\|๐\)ฯtโ\(๐,๐\)\}\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\):=\\mathbb\{E\}\_\{q\(\\bm\{z\}\)\}\\left\\\{\\dfrac\{\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\|\\bm\{z\}\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\)\}\{\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\}\\right\\\}\(13\)will generate the marginal probability pathฯtโ\(๐,๐\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\.
###### Theorem 4\.4\.
The marginal acceleration field in Eq\. \([13](https://arxiv.org/html/2608.21070#S4.E13)\) generates the marginal probability path\. See the proof in[SectionA\.3](https://arxiv.org/html/2608.21070#A1.SS3)\.
#### Equivalence between Regressing Conditional Acceleration Fields and Marginal Acceleration Fields
Assuming the marginal acceleration field๐โก\(๐,๐,t\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\)is known and we can sample from the marginal probability pathฯtโ\(๐,๐\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\), we can directly regress๐โก\(๐,๐,t\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\)using a neural network\. Let๐๐ฝโ\(โ
,โ
,โ
\):โdรโdรโ1โโd\\bm\{a\}\_\{\\bm\{\\theta\}\}\(\\cdot,\\cdot,\\cdot\):\\mathbb\{R\}^\{d\}\\times\\mathbb\{R\}^\{d\}\\times\\mathbb\{R\}^\{1\}\\rightarrow\\mathbb\{R\}^\{d\}be a time\-dependent acceleration field parameterized by a neural network with parametersฮธ\\theta\. The acceleration flow matching loss is
โAFM=๐ผ\(๐,๐\)โผฯtโ\(๐,๐\),tโผ๐ฐโก\[t0,tK\]โ\{โ๐๐ฝโ๐โก\(๐,๐,t\)โ2\}\\begin\{split\}\\mathcal\{L\}\_\{\\text\{AFM\}\}=\\mathbb\{E\}\_\{\(\\bm\{x\},\\bm\{v\}\)\\sim\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\),t\\sim\\mathcal\{U\}\[t\_\{0\},t\_\{K\}\]\}\\left\\\{\\\|\\bm\{a\}\_\{\\bm\{\\theta\}\}\-\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\)\\\|^\{2\}\\right\\\}\\end\{split\}However,๐โก\(๐,๐,t\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\)is generally intractable because it is defined via an expectation \(integral\), and its denominatorฯtโ\(๐,๐\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)also requires integration to compute\. In contrast, the conditional acceleration field๐โก\(๐,๐,t\|z\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\|z\)and the conditional probability pathฯtโ\(๐,๐\|z\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|z\)have simple forms\. Therefore, in actual training, we use the following Conditional Acceleration Flow Matching \(CAFM\) objective:
โCAFM=๐ผ\(๐,๐\)โผฯtโ\(๐,๐\|z\),tโผ๐ฐโก\[t0,tK\],zโผqโก\(z\)โ\[โ๐ฮธโ\(๐,๐,t\)โ๐โก\(๐,๐,t\|z\)โ2\]\\begin\{split\}\\mathcal\{L\}\_\{\\text\{CAFM\}\}=\\mathbb\{E\}\_\{\\begin\{subarray\}\{c\}\(\\bm\{x\},\\bm\{v\}\)\\sim\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|z\),\\\\ t\\sim\\mathcal\{U\}\[t\_\{0\},t\_\{K\}\],z\\sim q\(z\)\\end\{subarray\}\}\\Big\[\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{x\},\\bm\{v\},t\)\-\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\|z\)\\\|^\{2\}\\Big\]\\end\{split\}These two objectives are equivalent in training, as described by the following theorem:
###### Theorem 4\.5\.
Ifฯtโ\(๐ฑ,๐ฏ\)\>0\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\>0for any๐ฑโโd,๐ฏโโd,tโ\[t0,tK\]\\bm\{x\}\\in\\mathbb\{R\}^\{d\},\\bm\{v\}\\in\\mathbb\{R\}^\{d\},t\\in\[t\_\{0\},t\_\{K\}\], thenโAFM\\mathcal\{L\}\_\{\\text\{AFM\}\}andโCAFM\\mathcal\{L\}\_\{\\text\{CAFM\}\}differ only by a constant independent of๐\\bm\{\\theta\}\. In other words,โ๐โAFM=โ๐โCAFM\\nabla\_\{\\bm\{\\theta\}\}\\mathcal\{L\}\_\{\\text\{AFM\}\}=\\nabla\_\{\\bm\{\\theta\}\}\\mathcal\{L\}\_\{\\text\{CAFM\}\}\. The proof is left to[SectionA\.4](https://arxiv.org/html/2608.21070#A1.SS4)\.
#### Design of the Conditional Probability Path
Consider the joint distributionsฮผ0โ\(๐,๐\)\\mu\_\{0\}\(\\bm\{x\},\\bm\{v\}\),ฮผ1โ\(๐,๐\)\\mu\_\{1\}\(\\bm\{x\},\\bm\{v\}\), โฆ,ฮผKโ\(๐,๐\)\\mu\_\{K\}\(\\bm\{x\},\\bm\{v\}\)atK\+1K\+1time points\. The optimal transport plans calculated via solving SOAT problem areฯ0โ1โ,ฯ1โ2โ,โฏ,ฯKโ1โKโ\\pi^\{\*\}\_\{0\\rightarrow 1\},\\pi^\{\*\}\_\{1\\rightarrow 2\},\\cdots,\\pi^\{\*\}\_\{K\-1\\rightarrow K\}\. Taking the condition variable๐=\[\(๐0,๐0\),โฏ,\(๐K,๐K\)\]\\bm\{z\}=\[\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\),\\cdots,\(\\bm\{x\}\_\{K\},\\bm\{v\}\_\{K\}\)\], and similar to standard Flow Matching, we choose a path with time\-varying mean and invariant variance as the conditional probability path, i\.e\.:
ฯtโ\(๐,๐\|z\)=๐ฉโก\(๐\|๐xโ\(t\),ฯx2โI\)โ
๐ฉโก\(๐\|๐vโ\(t\),ฯv2โI\)\\begin\{split\}\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|z\)=\\mathcal\{N\}\(\\bm\{x\}\|\\bm\{m\}\_\{x\}\(t\),\\sigma\_\{x\}^\{2\}I\)\\cdot\\mathcal\{N\}\(\\bm\{v\}\|\\bm\{m\}\_\{v\}\(t\),\\sigma\_\{v\}^\{2\}I\)\\end\{split\}\(14\)Here,๐x,๐v\\bm\{m\}\_\{x\},\\bm\{m\}\_\{v\}are given by the Minimum Cost Trajectory described in[Section4\.1](https://arxiv.org/html/2608.21070#S4.SS1.SSS0.Px1), where fortโ\[ti,ti\+1\]t\\in\[t\_\{i\},t\_\{i\+1\}\]:
๐xโ\(t\)=c3โ\(tโti\)3\+c2โ\(tโti\)2\+c1โ\(tโti\)\+c0๐vโ\(t\)โ3โc3โ\(tโti\)2\+2โc2โ\(tโti\)\+c1\\displaystyle\\bm\{m\}\_\{x\}\(t\)=c\_\{3\}\(t\-t\_\{i\}\)^\{3\}\+c\_\{2\}\(t\-t\_\{i\}\)^\{2\}\+c\_\{1\}\(t\-t\_\{i\}\)\+c\_\{0\}\\quad\\bm\{m\}\_\{v\}\(t\)3c\_\{3\}\(t\-t\_\{i\}\)^\{2\}\+2c\_\{2\}\(t\-t\_\{i\}\)\+c\_\{1\}, the coefficients are identical to those given in[SectionA\.1](https://arxiv.org/html/2608.21070#A1.SS1), with๐0,๐0\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}replaced by๐i,๐i\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\},๐T,๐T\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}replaced by๐i\+1,๐i\+1\\bm\{x\}\_\{i\+1\},\\bm\{v\}\_\{i\+1\}, andTTreplaced byti\+1โtit\_\{i\+1\}\-t\_\{i\}\. Based on the optimal transport plans, we define the optimal propagator \(transition kernel\) from timetit\_\{i\}toti\+1t\_\{i\+1\}as the probability of a particle at\(๐i,๐i\)\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\)at timetit\_\{i\}being transported to\(๐i\+1,๐i\+1\)\(\\bm\{x\}\_\{i\+1\},\\bm\{v\}\_\{i\+1\}\)at timeti\+1t\_\{i\+1\}:
๐ฆiโi\+1โโ\[\(๐i,๐i\),\(๐i\+1,๐i\+1\)\]=ฯiโi\+1โโ\[\(๐i,๐i\),\(๐i\+1,๐i\+1\)\]ฮผiโ\(๐i,๐i\)\\mathcal\{K\}^\{\*\}\_\{i\\rightarrow i\+1\}\[\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\),\(\\bm\{x\}\_\{i\+1\},\\bm\{v\}\_\{i\+1\}\)\]=\\frac\{\\pi^\{\*\}\_\{i\\rightarrow i\+1\}\[\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\),\(\\bm\{x\}\_\{i\+1\},\\bm\{v\}\_\{i\+1\}\)\]\}\{\\mu\_\{i\}\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\)\}\(15\)
Then,
๐โผฮผ0โ\(๐0,๐0\)โ
โi=0Kโ1๐ฆiโi\+1โโ\[\(๐i,๐i\),\(๐i\+1,๐i\+1\)\]\\bm\{z\}\\sim\\mu\_\{0\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\)\\cdot\\prod\_\{i=0\}^\{K\-1\}\\mathcal\{K\}^\{\*\}\_\{i\\rightarrow i\+1\}\[\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\),\(\\bm\{x\}\_\{i\+1\},\\bm\{v\}\_\{i\+1\}\)\]
In the following, we demonstrate that the constructed marginal probability path recovers the true distribution at every given time point\.
###### Theorem 4\.6\.
Whenฯxโ0,ฯvโ0\\sigma\_\{x\}\\rightarrow 0,\\sigma\_\{v\}\\rightarrow 0, the marginal probability density constructed by[Equation12](https://arxiv.org/html/2608.21070#S4.E12)satisfiesฯtiโ\(๐ฑ,๐ฏ\)โฮผiโ\(๐ฑ,๐ฏ\)\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\)\\rightarrow\\mu\_\{i\}\(\\bm\{x\},\\bm\{v\}\)at theK\+1K\+1given time points\. Therefore, the Marginal Probability Path constructed by mixing Conditional Probability Paths can correctly reconstruct the marginal distributions\. See the proof in[SectionA\.5](https://arxiv.org/html/2608.21070#A1.SS5)\.
Furthermore, we demonstrate that the design yields an exact solution to the DOAT problem\.
###### Proposition 4\.7\.
Asฯxโ0\\sigma\_\{x\}\\rightarrow 0andฯvโ0\\sigma\_\{v\}\\rightarrow 0, the marginal probability path and the marginal acceleration field will converge to the solution to the DOAT problem\. See the proof in[SectionA\.6](https://arxiv.org/html/2608.21070#A1.SS6)\.
### 4\.3Transform VM\-DOAT to DOAT
#### Relations between DOAT and VM\-DOAT
The optimal transport cost of the DOAT problem is a functional of the sequence of joint distributionsฮผiโ\(๐,๐\)\\mu\_\{i\}\(\\bm\{x\},\\bm\{v\}\)fori=0,โฆ,Ki=0,\\dots,K\. Formally, we denote this dependency as๐ฅDOATโ\[\{ฮผiโ\(๐,๐\)\}\|i=0K\]\\mathcal\{J\}\_\{\\text\{DOAT\}\}\[\\\{\\mu\_\{i\}\(\\bm\{x\},\\bm\{v\}\)\\\}\|\_\{i=0\}^\{K\}\]\. In the VM\-DOAT setting, we only have access to the marginal spatial distributionsฮผi\(pos\)โ\(๐\)=โซฮผiโ\(๐,๐\)โ๐๐\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\)=\\int\\mu\_\{i\}\(\\bm\{x\},\\bm\{v\}\)\\mathrm\{d\}\\bm\{v\}\. Consequently, the cost of VM\-DOAT, denoted as๐ฅVM\-DOATโ\[\{ฮผi\(pos\)โ\(๐\)\}\|i=0K\]\\mathcal\{J\}\_\{\\text\{VM\-DOAT\}\}\[\\\{\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\)\\\}\|\_\{i=0\}^\{K\}\], relates to๐ฅDOAT\\mathcal\{J\}\_\{\\text\{DOAT\}\}through the following minimization:
๐ฅVM\-DOATโ\[\{ฮผi\(pos\)โ\(๐\)\}\]=minฯiโ\(๐\|๐\)โก๐ฅDOATโ\[\{ฮผi\(pos\)โ\(๐\)โ
ฯiโ\(๐\|๐\)\}\],\\mathcal\{J\}\_\{\\text\{VM\-DOAT\}\}\[\\\{\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\)\\\}\]=\\min\_\{\\omega\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\)\}\\mathcal\{J\}\_\{\\text\{DOAT\}\}\[\\\{\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\)\\cdot\\omega\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\)\\\}\],\(16\)whereฯiโ\(๐\|๐\)\\omega\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\)represents the conditional distribution of velocity, such that the joint distribution factorizes asฮผiโ\(๐,๐\)=ฮผi\(pos\)โ\(๐\)โ
ฯiโ\(๐\|๐\)\\mu\_\{i\}\(\\bm\{x\},\\bm\{v\}\)=\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\)\\cdot\\omega\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\)\. Thus, our objective is to determine the optimalฯiโ\(๐\|๐\)\\omega\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\)that minimizes the expression above, thereby transforming the VM\-DOAT problem into a solvable DOAT instance\.
This transformation rests on two key properties: 1\) The SOAT problem can be solved efficiently using the Sinkhorn method\. Once the Transport Plan is computed, we can sample a trajectory\[\(๐0,๐0\),โฏ,\(๐K,๐K\)\]\[\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\),\\cdots,\(\\bm\{x\}\_\{K\},\\bm\{v\}\_\{K\}\)\]; 2\) Given the positions๐0:K\\bm\{x\}\_\{0:K\}of a trajectory, consider the following multi\-marginal optimal control problem:
minโซt0tK12โฅ๐ธโฒโฒ\(t\)โฅ2dt,s\.t\.๐ธ\(ti\)=๐i,i=0,โฆ,K\.\\min\\int\_\{t\_\{0\}\}^\{t\_\{K\}\}\\dfrac\{1\}\{2\}\\\|\\bm\{\\gamma\}^\{\\prime\\prime\}\(t\)\\\|^\{2\}\\mathrm\{d\}t,\\quad\\text\{s\.t\. \}\\bm\{\\gamma\}\(t\_\{i\}\)=\\bm\{x\}\_\{i\},\\ i=0,\\dots,K\.\(17\)The velocities of the optimal trajectory, denoted asV=\[๐ธโฒโ\(t0\),โฏ,๐ธโฒโ\(tK\)\]V=\[\\bm\{\\gamma\}^\{\\prime\}\(t\_\{0\}\),\\cdots,\\bm\{\\gamma\}^\{\\prime\}\(t\_\{K\}\)\], can be obtained by solving a sparse linear system, as illustrated in the following proposition:
###### Proposition 4\.8\.
The velocitiesVVdefined above are the solution to the sparse linear systemVโA=BVA=B, whereAโโ\(K\+1\)ร\(K\+1\)A\\in\\mathbb\{R\}^\{\(K\+1\)\\times\(K\+1\)\}is a tridiagonal matrix andBโโdร\(K\+1\)B\\in\\mathbb\{R\}^\{d\\times\(K\+1\)\}\. A detailed derivation and the explicit expressions forAAandBBare provided in[SectionA\.7](https://arxiv.org/html/2608.21070#A1.SS7)\.
It is worth noting that solving this linear system requires๐ชโก\(K\)\\mathcal\{O\}\(K\)time, ensuring high computational efficiency\.
#### Iterative Scheme to Perform the Transformation
Based on this decomposition, we employ an iterative scheme to estimate the conditional velocity distributions acrosst0,โฆ,tKt\_\{0\},\\dots,t\_\{K\}\. We initializeฯi\(0\)โ\(๐\|๐\)=ฮดโก\(๐โ๐\)\\omega\_\{i\}^\{\(0\)\}\(\\bm\{v\}\|\\bm\{x\}\)=\\delta\(\\bm\{v\}\-\\bm\{0\}\)and subsequently alternate between the following two steps:
- โข*Solve SOAT Plans*: Compute theKKoptimal SOAT plans between the distributionsฮผi\(pos\)โ\(๐\)โ
ฯi\(m\)โ\(๐\|๐\)\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\)\\cdot\\omega^\{\(m\)\}\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\)for adjacent time points\. This yields the plansฯiโi\+1โ\(m\)โ\[\(๐i,๐i\),\(๐i\+1,๐i\+1\)\]\\pi^\{\*\(m\)\}\_\{i\\rightarrow i\+1\}\[\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\),\(\\bm\{x\}\_\{i\+1\},\\bm\{v\}\_\{i\+1\}\)\]fori=0,โฆ,Kโ1i=0,\\dots,K\-1, wheremmdenotes the iteration index\. This step effectively provides the multi\-marginal optimal transport plan given the currentฯi\(m\)โ\(๐\|๐\)\\omega^\{\(m\)\}\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\)\.
- โข*Update Velocity Distributions*: Update the conditional velocity distributions based on the computed SOAT plans: ฯi\(m\+1\)โ\(๐\|๐i\)=โซq\(m\)\(๐\)ฮด\(๐^i\[๐0โฏ๐K\]โ๐\)d๐โiโซq\(m\)โ\(๐\)โdโ๐โi,\\omega^\{\(m\+1\)\}\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\_\{i\}\)=\\frac\{\\int q^\{\(m\)\}\(\\bm\{z\}\)\\delta\(\\bm\{\\hat\{v\}\}\_\{i\}\[\\bm\{x\}\_\{0\}\\cdots\\bm\{x\}\_\{K\}\]\-\\bm\{v\}\)\\,\\mathrm\{d\}\\bm\{z\}\_\{\\setminus i\}\}\{\\int q^\{\(m\)\}\(\\bm\{z\}\)\\,\\mathrm\{d\}\\bm\{z\}\_\{\\setminus i\}\},\(18\)wheredโ๐โi\\mathrm\{d\}\\bm\{z\}\_\{\\setminus i\}denotes the volume elementdโ๐0โdโ๐0โโฆโdโ๐Kโdโ๐K\\mathrm\{d\}\\bm\{x\}\_\{0\}\\mathrm\{d\}\\bm\{v\}\_\{0\}\\dots\\mathrm\{d\}\\bm\{x\}\_\{K\}\\mathrm\{d\}\\bm\{v\}\_\{K\}excludingdโ๐i\\mathrm\{d\}\\bm\{x\}\_\{i\}\(i\.e\., integrating over all variables except๐i\\bm\{x\}\_\{i\}\)\. The term๐^iโ\[๐0,โฆ,๐K\]\\bm\{\\hat\{v\}\}\_\{i\}\[\\bm\{x\}\_\{0\},\\dots,\\bm\{x\}\_\{K\}\]denotes the optimal velocity at timetit\_\{i\}derived from the sequence๐0,โฆ,๐K\\bm\{x\}\_\{0\},\\dots,\\bm\{x\}\_\{K\}via[Section4\.3](https://arxiv.org/html/2608.21070#S4.SS3.SSS0.Px1)\. Intuitively,ฯi\(m\+1\)โ\(๐\|๐i\)\\omega^\{\(m\+1\)\}\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\_\{i\}\)is set to the distribution induced by the velocities of all optimal control trajectories passing through๐i\\bm\{x\}\_\{i\}at timetit\_\{i\}\. Since the total transport cost is the sum of costs along individual trajectories, aligning๐\\bm\{v\}with the optimal velocity of each trajectory minimizes the global objective\.
In the update step above, the trajectory variable๐=\[\(๐0,๐0\),โฆ,\(๐K,๐K\)\]\\bm\{z\}=\[\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\),\\dots,\(\\bm\{x\}\_\{K\},\\bm\{v\}\_\{K\}\)\]follows the joint distributionq\(m\)โ\(๐\)q^\{\(m\)\}\(\\bm\{z\}\)defined as:
q\(m\)โ\(๐\)=ฮผ0\(pos\)โ\(๐0\)โ
ฯ0\(m\)โ\(๐0\|๐0\)รโi=0Kโ1๐ฆiโi\+1โ\(m\)โ\[\(๐i,๐i\),\(๐i\+1,๐i\+1\)\],\\displaystyle q^\{\(m\)\}\(\\bm\{z\}\)=\\ \\mu\_\{0\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\_\{0\}\)\\cdot\\omega\_\{0\}^\{\(m\)\}\(\\bm\{v\}\_\{0\}\|\\bm\{x\}\_\{0\}\)\\times\\prod\_\{i=0\}^\{K\-1\}\\mathcal\{K\}^\{\*\(m\)\}\_\{i\\rightarrow i\+1\}\[\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\),\(\\bm\{x\}\_\{i\+1\},\\bm\{v\}\_\{i\+1\}\)\],where the propagator๐ฆiโi\+1โ\(m\)\\mathcal\{K\}\_\{i\\rightarrow i\+1\}^\{\*\(m\)\}is given by:
๐ฆiโi\+1โ\(m\)โ\[\(๐i,๐i\),\(๐i\+1,๐i\+1\)\]=ฯiโi\+1โ\(m\)โ\[\(๐i,๐i\),\(๐i\+1,๐i\+1\)\]ฮผi\(pos\)โ\(๐i\)โ
ฯi\(m\)โ\(๐i\|๐i\)\.\\displaystyle\\mathcal\{K\}^\{\*\(m\)\}\_\{i\\rightarrow i\+1\}\[\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\),\(\\bm\{x\}\_\{i\+1\},\\bm\{v\}\_\{i\+1\}\)\]=\\frac\{\\pi^\{\*\(m\)\}\_\{i\\rightarrow i\+1\}\[\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\),\(\\bm\{x\}\_\{i\+1\},\\bm\{v\}\_\{i\+1\}\)\]\}\{\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\_\{i\}\)\\cdot\\omega\_\{i\}^\{\(m\)\}\(\\bm\{v\}\_\{i\}\|\\bm\{x\}\_\{i\}\)\}\.
During the iteration process, the objectiveโDOATโ\[ฮผi\(pos\)โ\(๐\)โ
ฯi\(m\)โ\(๐\|๐\)\]\\mathcal\{L\}\_\{\\text\{DOAT\}\}\[\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\)\\cdot\\omega\_\{i\}^\{\(m\)\}\(\\bm\{v\}\|\\bm\{x\}\)\]decreases monotonically\. Since the cost is bounded below by00, the iterative algorithm is guaranteed to converge\.
## 5TracingFlow Algorithm for Trajectory Inference
In the practical implementation of TracingFlow, we first perform an approximation of the conditional velocity distribution to reduce the VM\-DOAT problem to a DOAT problem\. Subsequently, we employ flow matching to solve for the acceleration field of the DOAT problem\. Notably, for lineage tracing data, we incorporate biological priors into this framework\.
### 5\.1Approximation of Conditional Velocity Distribution
Determining the full conditional velocity distributionฯiโ\(๐\|๐\)\\omega\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\)to transform VM\-DOAT to DOAT \([Section4\.3](https://arxiv.org/html/2608.21070#S4.SS3)\) is computationally prohibitive, potentially requiring an auxiliary generative model\. We therefore adopt a deterministic approximation, assigning a single velocity to each data point\. Consequently,[Equation18](https://arxiv.org/html/2608.21070#S4.E18)is modified to:
ฯi\(m\+1\)โ\(๐\|๐i\)\\displaystyle\\omega^\{\(m\+1\)\}\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\_\{i\}\)=ฮดโก\(๐โ๐exp\(m\+1\)โ\(๐i\)\),๐exp\(m\+1\)โ\(๐i\)=โซq\(m\)\(๐\)๐^i\[๐0โฏ๐K\]d๐โiโซq\(m\)โ\(๐\)โdโ๐โi\.\\displaystyle=\\delta\(\\bm\{v\}\-\\bm\{v\}^\{\(m\+1\)\}\_\{\\text\{exp\}\}\(\\bm\{x\}\_\{i\}\)\),\\ \\ \\bm\{v\}\_\{\\text\{exp\}\}^\{\(m\+1\)\}\(\\bm\{x\}\_\{i\}\)=\\frac\{\\int q^\{\(m\)\}\(\\bm\{z\}\)\\bm\{\\hat\{v\}\}\_\{i\}\[\\bm\{x\}\_\{0\}\\cdots\\bm\{x\}\_\{K\}\]\\,\\mathrm\{d\}\\bm\{z\}\_\{\\setminus i\}\}\{\\int q^\{\(m\)\}\(\\bm\{z\}\)\\,\\mathrm\{d\}\\bm\{z\}\_\{\\setminus i\}\}\.\(19\)Intuitively, this assigns๐exp\(m\+1\)โ\(๐i\)\\bm\{v\}^\{\(m\+1\)\}\_\{\\text\{exp\}\}\(\\bm\{x\}\_\{i\}\)as the mean velocity of all optimal control trajectories passing through๐i\\bm\{x\}\_\{i\}at timetit\_\{i\}\.
### 5\.2MiniBatch\-OT
Solving SOAT repeatedly is computationally intensive for large\-scale datasets\. However, since the conditional velocity distributionฯiโ\(๐\|๐\)\\omega\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\)is approximated as a Dirac delta, the joint empirical distributionฮผiโ\(๐,๐\)\\mu\_\{i\}\(\\bm\{x\},\\bm\{v\}\)effectively becomes a superposition of Dirac deltas\. This allows us to employ a minibatch strategy: we partitionฮผi\\mu\_\{i\}andฮผi\+1\\mu\_\{i\+1\}intoBBcorresponding minibatches\{ฮผi\(n\)\}n=1B\\\{\\mu\_\{i\}^\{\(n\)\}\\\}\_\{n=1\}^\{B\}and\{ฮผi\+1\(n\)\}n=1B\\\{\\mu\_\{i\+1\}^\{\(n\)\}\\\}\_\{n=1\}^\{B\}\. By solving the local plansฯiโi\+1โ\(n\)\\pi\_\{i\\rightarrow i\+1\}^\{\*\(n\)\}independently and aggregating them asฯโiโi\+1=โn=1Bฯiโi\+1โ\(n\)\\pi^\{\*\}\_\{i\\rightarrow i\+1\}=\\oplus\_\{n=1\}^\{B\}\\pi\_\{i\\rightarrow i\+1\}^\{\*\(n\)\}, we significantly enhance computational efficiency\.
### 5\.3Incorporation of Biological Priors
In lineage tracing applications, let indiceslil\_\{i\}andli\+1l\_\{i\+1\}denote individual samples at time pointstit\_\{i\}andti\+1t\_\{i\+1\}, respectively\. Each data point\(๐i,li,๐i,li\)\(\\bm\{x\}\_\{i,l\_\{i\}\},\\bm\{v\}\_\{i,l\_\{i\}\}\)is associated with a barcodebi,liโโ\+b\_\{i,l\_\{i\}\}\\in\\mathbb\{N\}^\{\+\}\. We incorporate this prior by modifying the cost matrix๐iโi\+1\\mathcal\{C\}\_\{i\\rightarrow i\+1\}\. Specifically, when computing the cost between thell\-th sample attit\_\{i\}and thekk\-th sample atti\+1t\_\{i\+1\}, we compare their barcodes: ifbi,liโ bi\+1,li\+1b\_\{i,l\_\{i\}\}\\neq b\_\{i\+1,l\_\{i\+1\}\}, the standard cost is multiplied by a penalty factorp0โ\(p0\>1\)p\_\{0\}\(p\_\{0\}\>1\)\. This soft constraint discourages transitions between distinct lineages, effectively embedding biological knowledge into the dynamics learning process\.
### 5\.4Algorithm of TracingFlow
The TracingFlow algorithm comprises three main steps: First, we apply the iterative approximation from[Section5\.1](https://arxiv.org/html/2608.21070#S5.SS1)to determine initial velocities, effectively transforming the VM\-DOAT problem into a DOAT problem\. Second, we compute the Optimal Transport Plan for SOAT as defined in[Equation11](https://arxiv.org/html/2608.21070#S4.E11)\. Third, we sample conditional probability paths to train the acceleration network๐๐ฝโ\(๐,๐,t\)\\bm\{a\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},\\bm\{v\},t\)via regression\. Additionally, to provide initial conditions for inference, we parameterize the unique initial velocities derived in the first step using a neural network๐0,๐โ\(๐\)\\bm\{v\}\_\{0,\\bm\{\\xi\}\}\(\\bm\{x\}\), trained similarly via regression\. The algorithmโs pseudocode is provided in[AppendixD](https://arxiv.org/html/2608.21070#A4)\.
Theoretically, TracingFlow recovers the exact distribution if the acceleration and initial velocity are perfectly fitted\. In practice, where losses are non\-zero, we prove that the positional๐ฒ2\\mathcal\{W\}\_\{2\}distance between the generated and true distributions is bounded by the sum of the velocity and AFM losses, and the relevant theorem is detailed in[TheoremA\.1](https://arxiv.org/html/2608.21070#A1.Thmtheorem1)\.
## 6Experiment Results
To evaluate the effectiveness of our algorithm, TracingFlow, we conducted three categories of experiments: 1\) verifying whether the acceleration field of TracingFlow can transport the initial data distributionฮผ0\(pos\)โ\(๐\)=โซฯt0โ\(๐,๐\)โ๐๐\\mu\_\{0\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\)=\\int\\rho\_\{t\_\{0\}\}\(\\bm\{x\},\\bm\{v\}\)\\mathrm\{d\}\\bm\{v\}to the data distributionฮผi\(pos\)โ\(๐\)\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\)at other given timestit\_\{i\}; 2\) verifying whether TracingFlow can effectively perform distribution interpolation and extrapolation; and 3\) verifying whether TracingFlow can more effectively infer the dynamical laws in Lineage Tracing data given biological prior knowledge\.
#### Distribution Transport
We evaluated TracingFlow on a 2\-dimensional synthetic dataset and real single\-cell omics datasets of varying dimensions\. The evaluation metrics were the 1\-Wasserstein \(๐ฒ1\\mathcal\{W\}\_\{1\}\) and 2\-Wasserstein \(๐ฒ2\\mathcal\{W\}\_\{2\}\) distances, measuring the discrepancy between the fitted and ground truth distributions\. The second\-order dynamics model of TracingFlow provides stronger expressive power than first\-order algorithms, yielding lower average๐ฒ1\\mathcal\{W\}\_\{1\}and๐ฒ2\\mathcal\{W\}\_\{2\}values across all datasets \([Table1](https://arxiv.org/html/2608.21070#S6.T1)\)\. We visualized the evolutionary trajectories learned by OT\-CFM and TracingFlow on the synthetic dataset in[Figure2](https://arxiv.org/html/2608.21070#S6.F2)\. TracingFlowโs second\-order model allows it to learn trajectories that cross in the position space๐1,๐2\\bm\{x\}\_\{1\},\\bm\{x\}\_\{2\}, whereas OT\-CFM is unable to do so\. In[SectionC\.1](https://arxiv.org/html/2608.21070#A3.SS1), we present the distribution reconstruction accuracy at each time point, as well as experimental results on additional datasets \(such as the Mexico Gulf dataset\)\.
Table 1:Average๐ฒ1\\mathcal\{W\}\_\{1\}and๐ฒ2\\mathcal\{W\}\_\{2\}distances between the generated and ground truth distributions at various time points on the 2D Simulation , Cite 5D, and Cite 100D datasets for different algorithms\.Figure 2:On the Simulation 2D dataset: a\) non\-crossing paths learned by OT\-CFM, and b\) crossing paths learned by TracingFlow\.
#### Interpolation and Extrapolation
To verify TracingFlowโs interpolation and extrapolation capabilities, we conducted a Hold\-One\-Out experiment on the 5\-dimensional EB dataset \(time pointst=0,โฆ,4t=0,\\dots,4\)\. In each iteration, we withheld one time point from the latter four for training and calculated the๐ฒ1\\mathcal\{W\}\_\{1\}and๐ฒ2\\mathcal\{W\}\_\{2\}distances between the interpolated and true distributions\. As shown in[Table3](https://arxiv.org/html/2608.21070#S6.T3), TracingFlow achieved the highest accuracy, demonstrating that models based on second\-order dynamics are superior in capturing high\-curvature and non\-linear trajectories\.
#### Lineage Tracing Data
To verify whether TracingFlow better handles lineage tracing data given biological priors, we experimented on a 3D Simulation Lineage Dataset and the Hematopoiesis Dataset\. Following[Section5\.3](https://arxiv.org/html/2608.21070#S5.SS3), we introduced biological priors into TracingFlow by modifying the transport cost matrix, other baselines also incorporates priors via conditional velocity fields\. We used lineage\-weighted๐ฒ1\\mathcal\{W\}\_\{1\}and๐ฒ2\\mathcal\{W\}\_\{2\}distances to evaluate consistency with biological priors while learning dynamics\. Results in[Table3](https://arxiv.org/html/2608.21070#S6.T3)show that incorporating priors enables TracingFlow to outperform other algorithms\. In[Figure4](https://arxiv.org/html/2608.21070#S6.F4), we visualized dynamics on the simulation dataset to demonstrate how TracingFlow learns trajectories conforming to priors\. In[Figure4](https://arxiv.org/html/2608.21070#S6.F4), we plotted the PCA\-reduced positions of two barcodes from the Hematopoiesis Dataset att=1t=1; the positions inferred by TracingFlow are significantly closer to the ground truth than those by OT\-CFM\.
Table 2:Average๐ฒ1\\mathcal\{W\}\_\{1\}and๐ฒ2\\mathcal\{W\}\_\{2\}distances between the predicted and ground truth distributions at held\-out time points on the EB 5D dataset for different algorithms\.
Table 3:Average lineage\-weighted๐ฒ1\\mathcal\{W\}\_\{1\}and๐ฒ2\\mathcal\{W\}\_\{2\}distances between the generated and ground truth distributions at various time points on the 3D Simulation Lineage and Hematopoiesis datasets for different algorithms\.
Figure 3:On the Sim\-Lineage dataset: a\) paths learned by OT\-CFM without considering biological priors, and b\) paths learned by TracingFlow that preserve biological priors\.
Figure 4:Comparison of generated \(circles\) and ground truth \(โxโ\) positions att=1t=1for two barcodes on the hematopoiesis dataset\[[64](https://arxiv.org/html/2608.21070#bib.bib64)\]\. \(a\) OT\-CFM\. \(b\) TracingFlow\. Visualized via 2D PCA\.
## 7Conclusion and Limitation
In this work, we introduce TracingFlow, a Flow Matching framework for the Dynamical Acceleration Optimal Transport \(DOAT\) problem\. By pre\-solving SOAT to learn acceleration and initial velocity, it enables efficient, simulation\-free solutions for large\-scale DOAT problem\. Experiments on simulated and real world data demonstrate that second\-order dynamics enhance expressivity, yielding precise temporal distribution recovery\. Additionally, incorporating biological priors better preserves intercellular lineages\.
A limitation is the computationally intensive SOAT pre\-calculation, though this is mitigable via minibatch\-OT\. Furthermore, explicitly modeling the theoretical conditional velocity distribution per data point is prohibitive, so we approximate it using its expectation\. The objectiveโซ0T12โโ๐โ2โ๐t\\int\_\{0\}^\{T\}\\frac\{1\}\{2\}\\\|\\bm\{a\}\\\|^\{2\}\\mathrm\{d\}talso lacks a clear physical interpretation, as standard classical actions exclude second time derivatives \(see[SectionE\.2](https://arxiv.org/html/2608.21070#A5.SS2)\)\. Despite current sequencing measurements lacking velocity data, priors suggest biological systems may follow higher\-order dynamics due to the existence of complex regulations\. Future work will extend TracingFlow to further model these dynamics\.
## References
- Albergo and Vanden\-Eijnden \[2022\]Michael S Albergo and Eric Vanden\-Eijnden\.Building normalizing flows with stochastic interpolants\.*arXiv preprint arXiv:2209\.15571*, 2022\.
- Albergo et al\. \[2023\]Michael S Albergo, Nicholas M Boffi, and Eric Vanden\-Eijnden\.Stochastic interpolants: A unifying framework for flows and diffusions\.*arXiv preprint arXiv:2303\.08797*, 2023\.
- Arnolโd \[2013\]Vladimir Igorevich Arnolโd\.*Mathematical methods of classical mechanics*, volume 60\.Springer Science & Business Media, 2013\.
- Atanackovic et al\. \[2025\]Lazar Atanackovic, Xi Zhang, Brandon Amos, Mathieu Blanchette, Leo J Lee, Yoshua Bengio, Alexander Tong, and Kirill Neklyudov\.Meta flow matching: Integrating vector fields on the wasserstein manifold\.In*The Thirteenth International Conference on Learning Representations*, 2025\.
- Benamou et al\. \[2019\]Jean\-David Benamou, Thomas O Gallouรซt, and Franรงois\-Xavier Vialard\.Second\-order models for optimal transport and cubic splines on the wasserstein space\.*Foundations of Computational Mathematics*, 19\(5\):1113โ1143, 2019\.
- Bunne et al\. \[2023\]Charlotte Bunne, Ya\-Ping Hsieh, Marco Cuturi, and Andreas Krause\.The schrรถdinger bridge between gaussian measures has a closed form\.In*International Conference on Artificial Intelligence and Statistics*, pages 5802โ5833\. PMLR, 2023\.
- Bunne et al\. \[2024\]Charlotte Bunne, Geoffrey Schiebinger, Andreas Krause, Aviv Regev, and Marco Cuturi\.Optimal transport for single\-cell and spatial omics\.*Nature Reviews Methods Primers*, 4\(1\):58, 2024\.
- Cao et al\. \[2025\]Zihan Cao, Yu Zhong, and Liang\-Jian Deng\.Taming flow matching with unbalanced optimal transport into fast pansharpening\.*arXiv preprint arXiv:2503\.14975*, 2025\.
- Chen et al\. \[2018\]Ricky TQ Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud\.Neural ordinary differential equations\.*Advances in neural information processing systems*, 31, 2018\.
- Chen et al\. \[2022\]Tianrong Chen, Guan\-Horng Liu, and Evangelos Theodorou\.Likelihood training of schrรถdinger bridge using forward\-backward SDEs theory\.In*International Conference on Learning Representations*, 2022\.
- Chen et al\. \[2019\]Yongxin Chen, Giovanni Conforti, Tryphon T Georgiou, and Luigia Ripani\.Multi\-marginal schrรถdinger bridges\.In*International Conference on Geometric Science of Information*, pages 725โ732\. Springer, 2019\.
- Chizat et al\. \[2022\]Lรฉnaรฏc Chizat, Stephen Zhang, Matthieu Heitz, and Geoffrey Schiebinger\.Trajectory inference via mean\-field langevin in path space\.*Advances in Neural Information Processing Systems*, 35:16731โ16742, 2022\.
- Corso et al\. \[2025\]Gabriele Corso, Vignesh Ram Somnath, Noah Getz, Regina Barzilay, Tommi Jaakkola, and Andreas Krause\.Composing unbalanced flows for flexible docking and relaxation\.In*The Thirteenth International Conference on Learning Representations*, 2025\.
- De Bortoli et al\. \[2023\]Valentin De Bortoli, Guan\-Horng Liu, Tianrong Chen, Evangelos A Theodorou, and Weilie Nie\.Augmented bridge matching\.*arXiv preprint arXiv:2311\.06978*, 2023\.
- Eyring et al\. \[2024\]Luca Eyring, Dominik Klein, Thรฉo Uscidda, Giovanni Palla, Niki Kilbertus, Zeynep Akata, and Fabian J Theis\.Unbalancedness in neural Monge maps improves unpaired domain translation\.In*The Twelfth International Conference on Learning Representations*, 2024\.
- Flamary et al\. \[2021\]Rรฉmi Flamary, Nicolas Courty, Alexandre Gramfort, Mokhtar Z Alaya, Aurรฉlie Boisbunon, Stanislas Chambon, Laetitia Chapel, Adrien Corenflos, Kilian Fatras, Nemo Fournier, et al\.Pot: Python optimal transport\.*Journal of Machine Learning Research*, 22\(78\):1โ8, 2021\.
- Forrow and Schiebinger \[2021\]Aden Forrow and Geoffrey Schiebinger\.Lineageot is a unified framework for lineage tracing and trajectory inference\.*Nature Communications*, 12\(1\), 2021\.
- Gao et al\. \[2025\]Mingze Gao, Melania Barile, Shirom Chabra, Myriam Haltalli, Emily F\. Calderbank, Yiming Chao, Weizhong Zheng, Nicola K\. Wilson, Elisa Laurenti, Berthold Gรถttgens, and Yuanhua Huang\.CLADES: a hybrid neuralode\-gillespie approach for unveiling clonal cell fate and differentiation dynamics\.*Nature Communications*, 16\(1\), 2025\.
- Gorin et al\. \[2020\]Gennady Gorin, Valentine Svensson, and Lior Pachter\.Protein velocity and acceleration from single\-cell multiomics experiments\.*Genome biology*, 21\(1\):39, 2020\.
- Guo and Schwing \[2025\]Pengsheng Guo and Alexander G Schwing\.Variational rectified flow matching\.*arXiv preprint arXiv:2502\.09616*, 2025\.
- Guo et al\. \[2025\]Wenbo Guo, Zeyu Chen, Xinqi Li, Jingmin Huang, Qifan Hu, and Jin Gu\.sctrace\+: Enhancing cell fate inference by integrating the lineage\-tracing and multi\-faceted transcriptomic similarity information\.*Cell Systems*, 16\(9\):101398, 2025\.
- Halmos et al\. \[2025\]Peter Halmos, Xinhao Liu, Julian Gold, Feng Chen, Li Ding, and Benjamin J Raphael\.Dest\-ot: Alignment of spatiotemporal transcriptomics data\.*Cell Systems*, 16\(2\), 2025\.
- Hong et al\. \[2025\]Ari Hong, Sangseon Lee, and Kwangsoo Kim\.Multi\-omic relay velocity modeling uncovers dynamic chromatin\-transcription regulation across cell states\.*Nature Communications*, 2025\.
- Huguet et al\. \[2022\]Guillaume Huguet, Daniel Sumner Magruder, Alexander Tong, Oluwadamilola Fasina, Manik Kuchroo, Guy Wolf, and Smita Krishnaswamy\.Manifold interpolating optimal\-transport flows for trajectory inference\.*Advances in neural information processing systems*, 35:29705โ29718, 2022\.
- Jiang and Wan \[2024\]Qi Jiang and Lin Wan\.A physics\-informed neural SDE network for learning cellular dynamics from time\-series scRNA\-seq data\.*Bioinformatics*, 40:ii120โii127, 09 2024\.ISSN 1367\-4811\.
- Kapusniak et al\. \[2024\]Kacper Kapusniak, Peter Potaptchik, Teodora Reu, Leo Zhang, Alexander Tong, Michael Bronstein, Joey Bose, and Francesco Di Giovanni\.Metric flow matching for smooth interpolations on the data manifold\.*Advances in Neural Information Processing Systems*, 37:135011โ135042, 2024\.
- Klein et al\. \[2024\]Dominik Klein, Thรฉo Uscidda, Fabian Theis, and Marco Cuturi\.Genot: Entropic \(gromov\) wasserstein flow matching with applications to single\-cell genomics\.*Advances in Neural Information Processing Systems*, 37:103897โ103944, 2024\.
- Klein et al\. \[2025\]Dominik Klein, Giovanni Palla, Marius Lange, Michal Klein, Zoe Piran, Manuel Gander, Laetitia Meng\-Papaxanthos, Michael Sterr, Lama Saber, Changying Jing, et al\.Mapping cells through time and space with moscot\.*Nature*, pages 1โ11, 2025\.
- Koshizuka and Sato \[2023\]Takeshi Koshizuka and Issei Sato\.Neural lagrangian schrรถdinger bridge: Diffusion modeling for population dynamics\.In*The Eleventh International Conference on Learning Representations*, 2023\.
- Kretzschmar and Watt \[2012\]Kai Kretzschmar and Fiona M\. Watt\.Lineage tracing\.*Cell*, 148\(1โ2\):33โ45, 2012\.
- Lance et al\. \[2022\]Christopher Lance, Malte D Luecken, Daniel B Burkhardt, Robrecht Cannoodt, Pia Rautenstrauch, Anna Laddach, Aidyn Ubingazhibov, Zhi\-Jie Cao, Kaiwen Deng, Sumeer Khan, et al\.Multimodal single cell data integration challenge: results and lessons learned\.*BioRxiv*, pages 2022โ04, 2022\.
- Lange et al\. \[2024\]Marius Lange, Zoe Piran, Michal Klein, Bastiaan Spanjaard, Dominik Klein, Jan Philipp Junker, Fabian J\. Theis, and Mor Nitzan\.Mapping lineage\-traced cells across time points with moslin\.*Genome Biology*, 25\(1\), 2024\.
- Lavenant et al\. \[2024\]Hugo Lavenant, Stephen Zhang, Young\-Heon Kim, Geoffrey Schiebinger, et al\.Toward a mathematical theory of trajectory inference\.*The Annals of Applied Probability*, 34\(1A\):428โ500, 2024\.
- Li et al\. \[2023\]Chen Li, Maria C Virgilio, Kathleen L Collins, and Joshua D Welch\.Multi\-omic single\-cell velocity models epigenomeโtranscriptome interactions and improves cell fate prediction\.*Nature biotechnology*, 41\(3\):387โ398, 2023\.
- Lipman et al\. \[2023\]Yaron Lipman, Ricky T\. Q\. Chen, Heli Ben\-Hamu, Maximilian Nickel, and Matthew Le\.Flow matching for generative modeling\.In*The Eleventh International Conference on Learning Representations*, 2023\.
- Liu et al\. \[2022\]Xingchao Liu, Chengyue Gong, and Qiang Liu\.Flow straight and fast: Learning to generate and transfer data with rectified flow\.*arXiv preprint arXiv:2209\.03003*, 2022\.
- Maddu et al\. \[2024\]Suryanarayana Maddu, Victor Chardรจs, Michael Shelley, et al\.Inferring biological processes with intrinsic noise from cross\-sectional data\.*arXiv preprint arXiv:2410\.07501*, 2024\.
- Mao et al\. \[2025\]Shanjun Mao, Chenyang Zhang, Runjiu Chen, Shan Tang, Xiaodan Fan, and Jie Hu\.Cell lineage tracing: Methods, applications, and challenges\.*Quantitative Biology*, 13\(4\), 2025\.
- Moon et al\. \[2019\]Kevin R Moon, David Van Dijk, Zheng Wang, Scott Gigante, Daniel B Burkhardt, William S Chen, Kristina Yim, Antonia van den Elzen, Matthew J Hirn, Ronald R Coifman, et al\.Visualizing structure and transitions in high\-dimensional biological data\.*Nature biotechnology*, 37\(12\):1482โ1492, 2019\.
- Neklyudov et al\. \[2023\]Kirill Neklyudov, Rob Brekelmans, Daniel Severo, and Alireza Makhzani\.Action matching: Learning stochastic dynamics from samples\.In*International conference on machine learning*, pages 25858โ25889\. PMLR, 2023\.
- Neklyudov et al\. \[2024\]Kirill Neklyudov, Rob Brekelmans, Alexander Tong, Lazar Atanackovic, Qiang Liu, and Alireza Makhzani\.A computational framework for solving wasserstein lagrangian flows\.In*Forty\-first International Conference on Machine Learning*, 2024\.
- Park et al\. \[2024\]Dogyun Park, Sojin Lee, Sihyeon Kim, Taehoon Lee, Youngjoon Hong, and Hyunwoo J Kim\.Constant acceleration flow\.*Advances in Neural Information Processing Systems*, 37:90030โ90060, 2024\.
- Paszke et al\. \[2017\]Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer\.Automatic differentiation in pytorch\.2017\.
- Peng et al\. \[2024\]Qiangwei Peng, Peijie Zhou, and Tiejun Li\.stVCR: Spatiotemporal dynamics of single cells\.*bioRxiv*, pages 2024โ06, 2024\.
- Peng et al\. \[2026\]Qiangwei Peng, Zihan Wang, Junda Ying, Yuhao Sun, Qing Nie, Lei Zhang, Tiejun Li, and Peijie Zhou\.WFR\-FM: Simulation\-free dynamic unbalanced optimal transport\.*arXiv preprint arXiv:2601\.06810*, 2026\.
- Petroviฤ et al\. \[2025\]Katarina Petroviฤ, Lazar Atanackovic, Kacper Kapusniak, Michael M\. Bronstein, Joey Bose, and Alexander Tong\.Curly flow matching for learning non\-gradient field dynamics\.In*Learning Meaningful Representations of Life \(LMRL\) Workshop at ICLR 2025*, 2025\.
- Peyrรฉ and Cuturi \[2019\]Gabriel Peyrรฉ and Marco Cuturi\.*Computational optimal transport: With applications to data science*\.Now Foundations and Trends, 2019\.
- Pooladian et al\. \[2023\]Aram\-Alexandre Pooladian, Heli Ben\-Hamu, Carles Domingo\-Enrich, Brandon Amos, Yaron Lipman, and Ricky TQ Chen\.Multisample flow matching: Straightening flows with minibatch couplings\.*arXiv preprint arXiv:2304\.14772*, 2023\.
- Prasad et al\. \[2020\]Neha Prasad, Karren Yang, and Caroline Uhler\.Optimal transport using gans for lineage tracing\.*arXiv preprint arXiv:2007\.12098*, 2020\.
- Rohbeck et al\. \[2025\]Martin Rohbeck, Edward De Brouwer, Charlotte Bunne, Jan\-Christian Huetter, Anne Biton, Kelvin Y Chen, Aviv Regev, and Romain Lopez\.Modeling complex system dynamics with flow matching across time and conditions\.In*The Thirteenth International Conference on Learning Representations*, 2025\.
- Schiebinger et al\. \[2019\]Geoffrey Schiebinger, Jian Shu, Marcin Tabaka, Brian Cleary, Vidya Subramanian, Aryeh Solomon, Joshua Gould, Siyan Liu, Stacie Lin, Peter Berube, et al\.Optimal\-transport analysis of single\-cell gene expression identifies developmental trajectories in reprogramming\.*Cell*, 176\(4\):928โ943, 2019\.
- Sha et al\. \[2024\]Yutong Sha, Yuchi Qiu, Peijie Zhou, and Qing Nie\.Reconstructing growth and dynamic trajectories from single\-cell transcriptomics data\.*Nature Machine Intelligence*, 6\(1\):25โ39, 2024\.
- Shi et al\. \[2024\]Yuyang Shi, Valentin De Bortoli, Andrew Campbell, and Arnaud Doucet\.Diffusion schrรถdinger bridge matching\.*Advances in Neural Information Processing Systems*, 36, 2024\.
- Shrestha and Fu \[2025\]Sagar Shrestha and Xiao Fu\.Diversified flow matching with translation identifiability\.*arXiv preprint arXiv:2511\.05558*, 2025\.
- Sun et al\. \[2025\]Yuhao Sun, Zhenyi Zhang, Zihan Wang, Tiejun Li, and Peijie Zhou\.Variational regularized unbalanced optimal transport: Single network, least action\.*arXiv preprint arXiv:2505\.11823*, 2025\.
- Theodoropoulos et al\. \[2025\]Panagiotis Theodoropoulos, Augustinos D Saravanos, Evangelos A Theodorou, and Guan\-Horng Liu\.Momentum multi\-marginal schr\\\\backslash" odinger bridge matching\.*arXiv preprint arXiv:2506\.10168*, 2025\.
- Tong et al\. \[2020\]Alexander Tong, Jessie Huang, Guy Wolf, David Van Dijk, and Smita Krishnaswamy\.Trajectorynet: A dynamic optimal transport network for modeling cellular dynamics\.In*International conference on machine learning*, pages 9526โ9536\. PMLR, 2020\.
- Tong et al\. \[2024a\]Alexander Tong, Kilian FATRAS, Nikolay Malkin, Guillaume Huguet, Yanlei Zhang, Jarrid Rector\-Brooks, Guy Wolf, and Yoshua Bengio\.Improving and generalizing flow\-based generative models with minibatch optimal transport\.*Transactions on Machine Learning Research*, 2024a\.ISSN 2835\-8856\.Expert Certification\.
- Tong et al\. \[2024b\]Alexander Tong, Nikolay Malkin, Kilian Fatras, Lazar Atanackovic, Yanlei Zhang, Guillaume Huguet, Guy Wolf, and Yoshua Bengio\.Simulation\-free schrรถdinger bridges via score and flow matching\.In*International Conference on Artificial Intelligence and Statistics*, pages 1279โ1287\. PMLR, 2024b\.
- Ventre et al\. \[2023\]Elias Ventre, Aden Forrow, Nitya Gadhiwala, Parijat Chakraborty, Omer Angel, and Geoffrey Schiebinger\.Trajectory inference for a branching sde model of cell differentiation\.*arXiv preprint arXiv:2307\.07687*, 2023\.
- Wang et al\. \[2025\]Dongyi Wang, Yuanwei Jiang, Zhenyi Zhang, Xiang Gu, Peijie Zhou, and Jian Sun\.Joint velocity\-growth flow matching for single\-cell dynamics modeling\.*arXiv preprint arXiv:2505\.13413*, 2025\.
- Wang et al\. \[2023\]Kun Wang, Liangzhen Hou, Xin Wang, Xiangwei Zhai, Zhaolian Lu, Zhike Zi, Weiwei Zhai, Xionglei He, Christina Curtis, Da Zhou, and Zheng Hu\.Phylovelo enhances transcriptomic velocity field mapping using monotonically expressed genes\.*Nature Biotechnology*, 42\(5\):778โ789, 2023\.
- Wang et al\. \[2022\]Shou\-Wen Wang, Michael J\. Herriges, Kilian Hurley, Darrell N\. Kotton, and Allon M\. Klein\.Cospar identifies early cell fate biases from single\-cell transcriptomic and lineage information\.*Nature Biotechnology*, 40\(7\):1066โ1074, 2022\.
- Weinreb et al\. \[2020\]Caleb Weinreb, Alejo Rodriguez\-Fraticelli, Fernando D Camargo, and Allon M Klein\.Lineage tracing on transcriptional landscapes links state to fate during differentiation\.*Science*, 367\(6479\):eaaw3381, 2020\.
- Yang \[2025\]Maosheng Yang\.Topological schrรถdinger bridge matching\.In*The Thirteenth International Conference on Learning Representations*, 2025\.
- Yeo et al\. \[2021\]Grace Hui Ting Yeo, Sachit D Saksena, and David K Gifford\.Generative modeling of single\-cell time series with prescient enables prediction of cell trajectories with interventions\.*Nature communications*, 12\(1\):3222, 2021\.
- Zhang et al\. \[2024a\]Jiaqi Zhang, Erica Larschan, Jeremy Bigness, and Ritambhara Singh\.scNODE: generative model for temporal single cell transcriptomic data prediction\.*Bioinformatics*, 40\(Supplement\_2\):ii146โii154, 09 2024a\.ISSN 1367\-4811\.
- Zhang et al\. \[2024b\]Xi Nicole Zhang, Yuan Pu, Yuki Kawamura, Andrew Loza, Yoshua Bengio, Dennis Shung, and Alexander Tong\.Trajectory flow matching with applications to clinical time series modelling\.*Advances in Neural Information Processing Systems*, 37:107198โ107224, 2024b\.
- Zhang et al\. \[2025a\]Yichi Zhang, Yici Yan, Alex Schwing, and Zhizhen Zhao\.Towards hierarchical rectified flow\.*arXiv preprint arXiv:2502\.17436*, 2025a\.
- Zhang et al\. \[2025b\]Zhenyi Zhang, Tiejun Li, and Peijie Zhou\.Learning stochastic dynamics from snapshots through regularized unbalanced optimal transport\.In*The Thirteenth International Conference on Learning Representations*, 2025b\.
- Zhang et al\. \[2025c\]Zhenyi Zhang, Yuhao Sun, Qiangwei Peng, Tiejun Li, and Peijie Zhou\.Integrating dynamical systems modeling with spatiotemporal scRNA\-Seq data analysis\.*Entropy*, 27\(5\), 2025c\.ISSN 1099\-4300\.
- Zheng et al\. \[2017\]Grace XY Zheng, Jessica M Terry, Phillip Belgrader, Paul Ryvkin, Zachary W Bent, Ryan Wilson, Solongo B Ziraldo, Tobias D Wheeler, Geoff P McDermott, Junjie Zhu, et al\.Massively parallel digital transcriptional profiling of single cells\.*Nature communications*, 8\(1\):14049, 2017\.
- Zhu and Lin \[2024\]Qunxi Zhu and Wei Lin\.Switched flow matching: Eliminating singularities via switching odes\.*arXiv preprint arXiv:2405\.11605*, 2024\.
- Zhu et al\. \[2024\]Qunxi Zhu, Bolin Zhao, Jingdong Zhang, Peiyang Li, and Wei Lin\.Governing equation discovery of a complex system from snapshots\.*arXiv preprint arXiv:2410\.16694*, 2024\.
## Contents of Appendix
## Appendix AProofs of Main Theorems
### A\.1Proof of Proposition 4\.1
###### Proof\.
First, we show that the unique minimizer is a cubic polynomial\. Consider the energy functional defined on the interval\[0,T\]\[0,T\]:
๐ฅโก\[ฮณ\]=12โโซ0Tโฮณโฒโฒโ\(t\)โ22โ๐t\.\\mathcal\{J\}\[\\gamma\]=\\frac\{1\}\{2\}\\int\_\{0\}^\{T\}\\\|\\gamma^\{\\prime\\prime\}\(t\)\\\|\_\{2\}^\{2\}\\,\\mathrm\{d\}t\.\(20\)The integrandLโก\(ฮณ,ฮณโฒ,ฮณโฒโฒ\)=12โโฮณโฒโฒโ22L\(\\gamma,\\gamma^\{\\prime\},\\gamma^\{\\prime\\prime\}\)=\\frac\{1\}\{2\}\\\|\\gamma^\{\\prime\\prime\}\\\|\_\{2\}^\{2\}depends on derivatives up to the second order\. The necessary condition for optimality is given by the EulerโLagrange equation:
โLโฮณโddโtโ\(โLโฮณโฒ\)\+d2dโt2โ\(โLโฮณโฒโฒ\)=0\.\\frac\{\\partial L\}\{\\partial\\gamma\}\-\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\left\(\\frac\{\\partial L\}\{\\partial\\gamma^\{\\prime\}\}\\right\)\+\\frac\{\\mathrm\{d\}^\{2\}\}\{\\mathrm\{d\}t^\{2\}\}\\left\(\\frac\{\\partial L\}\{\\partial\\gamma^\{\\prime\\prime\}\}\\right\)=0\.\(21\)SinceโL/โฮณ=0\\partial L/\\partial\\gamma=0,โL/โฮณโฒ=0\\partial L/\\partial\\gamma^\{\\prime\}=0, andโL/โฮณโฒโฒ=ฮณโฒโฒ\\partial L/\\partial\\gamma^\{\\prime\\prime\}=\\gamma^\{\\prime\\prime\}, the equation simplifies to:
d2dโt2โฮณโฒโฒโ\(t\)=ฮณ\(4\)โ\(t\)=0\.\\frac\{\\mathrm\{d\}^\{2\}\}\{\\mathrm\{d\}t^\{2\}\}\\gamma^\{\\prime\\prime\}\(t\)=\\gamma^\{\(4\)\}\(t\)=0\.\(22\)This implies that the optimal trajectoryฮณโก\(t\)\\gamma\(t\)is a cubic polynomial:
ฮณโก\(t\)=c3โt3\+c2โt2\+c1โt\+c0\.\\gamma\(t\)=c\_\{3\}t^\{3\}\+c\_\{2\}t^\{2\}\+c\_\{1\}t\+c\_\{0\}\.\(23\)Imposing the boundary conditionsฮณโก\(0\)=๐0,ฮณโฒโ\(0\)=๐0\\gamma\(0\)=\\bm\{x\}\_\{0\},\\gamma^\{\\prime\}\(0\)=\\bm\{v\}\_\{0\}andฮณโก\(T\)=๐T,ฮณโฒโ\(T\)=๐T\\gamma\(T\)=\\bm\{x\}\_\{T\},\\gamma^\{\\prime\}\(T\)=\\bm\{v\}\_\{T\}leads to the linear system:
\{c0=๐0,c1=๐0,c3โT3\+c2โT2\+c1โT\+c0=๐T,3โc3โT2\+2โc2โT\+c1=๐T\.\\begin\{cases\}c\_\{0\}=\\bm\{x\}\_\{0\},\\\\ c\_\{1\}=\\bm\{v\}\_\{0\},\\\\ c\_\{3\}T^\{3\}\+c\_\{2\}T^\{2\}\+c\_\{1\}T\+c\_\{0\}=\\bm\{x\}\_\{T\},\\\\ 3c\_\{3\}T^\{2\}\+2c\_\{2\}T\+c\_\{1\}=\\bm\{v\}\_\{T\}\.\\end\{cases\}\(24\)Solving for the coefficients yields the unique solution:
\{c0=๐0,c1=๐0,c2=3โ\(๐Tโ๐0\)โ\(2โ๐0\+๐T\)โTT2,c3=\(๐0\+๐T\)โTโ2โ\(๐Tโ๐0\)T3\.\\begin\{cases\}c\_\{0\}=\\bm\{x\}\_\{0\},\\\\ c\_\{1\}=\\bm\{v\}\_\{0\},\\\\ c\_\{2\}=\\dfrac\{3\(\\bm\{x\}\_\{T\}\-\\bm\{x\}\_\{0\}\)\-\(2\\bm\{v\}\_\{0\}\+\\bm\{v\}\_\{T\}\)T\}\{T^\{2\}\},\\\\ c\_\{3\}=\\dfrac\{\(\\bm\{v\}\_\{0\}\+\\bm\{v\}\_\{T\}\)T\-2\(\\bm\{x\}\_\{T\}\-\\bm\{x\}\_\{0\}\)\}\{T^\{3\}\}\.\\end\{cases\}\(25\)โ
### A\.2Proof of Remark 4\.3
###### Proof\.
First we propose a lemma:
Lemma\.Let๐\[0,t\]โ\(๐0,๐0,๐t,๐t\)\\mathcal\{C\}\_\{\[0,t\]\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\},\\bm\{x\}\_\{t\},\\bm\{v\}\_\{t\}\)be the least acceleration cost through\[0,t\]\[0,t\],\(๐๐,๐0\),\(๐T,๐T\)โ๐ณร๐ฑ\(\\bm\{x\_\{0\}\},\\bm\{v\}\_\{0\}\),\(\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\)\\in\\mathcal\{X\}\\times\\mathcal\{V\}are two different points\. Then for arbitrary midpointฯโ\(0,1\)\\tau\\in\(0,1\)and state\(๐,๐\)โ๐ณร๐ฑ\(\\bm\{x\},\\bm\{v\}\)\\in\\mathcal\{X\}\\times\\mathcal\{V\}, the following inequality holds:
๐\[0,ฯ\]โ\(๐๐,๐๐,๐,๐\)\+๐\[ฯ,T\]โ\(๐,๐,๐T,๐T\)โฅ๐\[0,T\]โ\(๐0,๐0,๐T,๐T\)\\mathcal\{C\}\_\{\[0,\\tau\]\}\(\\bm\{x\_\{0\}\},\\bm\{v\_\{0\}\},\\bm\{x\},\\bm\{v\}\)\+\\mathcal\{C\}\_\{\[\\tau,T\]\}\(\\bm\{x\},\\bm\{v\},\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\)\\geq\\mathcal\{C\}\_\{\[0,T\]\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\},\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\)\(26\)The "==" holds if and only if\(๐,๐\)=ฮณโโ\(ฯ\)\(\\bm\{x,v\}\)=\\gamma^\{\*\}\(\\tau\), whereฮณโโ\(t\)โ\(tโ\[0,T\]\)\\gamma^\{\*\}\(t\)\(t\\in\[0,T\]\)is the cubic interpolation trajectory connecting\(๐๐,๐๐\)\(\\bm\{x\_\{0\}\},\\bm\{v\_\{0\}\}\)and\(๐T,๐๐ป\)\(\\bm\{x\}\_\{T\},\\bm\{v\_\{T\}\}\)\.
###### Proof\.
Fixฯโ\(0,T\)\\tau\\in\(0,T\)\. Define objective functionFฯโ\(๐,๐\)=๐\[0,ฯ\]โ\(๐๐,๐๐,๐,๐\)\+๐\[ฯ,T\]โ\(๐,๐,๐T,๐T\)F\_\{\\tau\}\(\\bm\{x\},\\bm\{v\}\)=\\mathcal\{C\}\_\{\[0,\\tau\]\}\(\\bm\{x\_\{0\}\},\\bm\{v\_\{0\}\},\\bm\{x\},\\bm\{v\}\)\+\\mathcal\{C\}\_\{\[\\tau,T\]\}\(\\bm\{x\},\\bm\{v\},\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\)\. By Corollary[4\.1](https://arxiv.org/html/2608.21070#S4.SS1.SSS0.Px2), we can rewrite
Fฯโ\(๐,๐\)=2ฯ3โ\{\(โ๐0โ2\+โจ๐0,๐โฉ\+โ๐โ2\)โฯ2\+3โโจ๐0\+๐,๐0โ๐โฉโฯ\+3โโ๐0โ๐โ2\}\+\\displaystyle F\_\{\\tau\}\(\\bm\{x\},\\bm\{v\}\)=\\frac\{2\}\{\\tau^\{3\}\}\\left\\\{\(\\\|\\bm\{v\}\_\{0\}\\\|^\{2\}\+\\langle\\bm\{v\}\_\{0\},\\bm\{v\}\\rangle\+\\\|\\bm\{v\}\\\|^\{2\}\)\\tau^\{2\}\+3\\langle\\bm\{v\}\_\{0\}\+\\bm\{v\},\\bm\{x\}\_\{0\}\-\\bm\{x\}\\rangle\\tau\+3\\\|\\bm\{x\}\_\{0\}\-\\bm\{x\}\\\|^\{2\}\\right\\\}\+\(27\)2\(Tโฯ\)3โ\{\(โ๐Tโ2\+โจ๐T,๐โฉ\+โ๐โ2\)โ\(Tโฯ\)2\+3โโจ๐T\+๐,๐โ๐TโฉโT\+3โโ๐Tโ๐โ2\}\\displaystyle\\frac\{2\}\{\(T\-\\tau\)^\{3\}\}\\left\\\{\(\\\|\\bm\{v\}\_\{T\}\\\|^\{2\}\+\\langle\\bm\{v\}\_\{T\},\\bm\{v\}\\rangle\+\\\|\\bm\{v\}\\\|^\{2\}\)\(T\-\\tau\)^\{2\}\+3\\langle\\bm\{v\}\_\{T\}\+\\bm\{v\},\\bm\{x\}\-\\bm\{x\}\_\{T\}\\rangle T\+3\\\|\\bm\{x\}\_\{T\}\-\\bm\{x\}\\\|^\{2\}\\right\\\}Which is a quadratic form of\(๐,๐\)\(\\bm\{x\},\\bm\{v\}\)up to a constant\.Fฯโ\(๐,๐\)โฅ0F\_\{\\tau\}\(\\bm\{x\},\\bm\{v\}\)\\geq 0always holds, namely, the quadratic form has a finite lower bound, so it is semi\-quadratic\. To minimizeFฯโ\(๐,๐\)F\_\{\\tau\}\(\\bm\{x\},\\bm\{v\}\), we can take differentiations:
โ๐Fฯโ\(๐,๐\)=3โ\(โ\(๐0\+๐\)โฯ\+2โ\(๐โ๐0\)ฯ2\+\(๐\+๐T\)โ\(Tโฯ\)\+2โ\(๐โ๐T\)\(Tโฯ\)2\)=0\.\\nabla\_\{\\bm\{x\}\}F\_\{\\tau\}\(\\bm\{x\},\\bm\{v\}\)=3\\left\(\\frac\{\-\(\\bm\{v\}\_\{0\}\+\\bm\{v\}\)\\tau\+2\(\\bm\{x\}\-\\bm\{x\}\_\{0\}\)\}\{\{\\tau\}^\{2\}\}\+\\frac\{\(\\bm\{v\}\+\\bm\{v\}\_\{T\}\)\(T\-\\tau\)\+2\(\\bm\{x\}\-\\bm\{x\}\_\{T\}\)\}\{\(T\-\{\\tau\}\)^\{2\}\}\\right\)=0\.\(28\)โ๐Fฯโ\(๐,๐\)=\(๐0\+2โ๐\)โฯ\+3โ\(๐0โ๐\)ฯ2\+\(๐T\+2โ๐\)โ\(Tโฯ\)\+3โ\(๐โ๐T\)\(Tโฯ\)2=0,\\nabla\_\{\\bm\{v\}\}F\_\{\\tau\}\(\\bm\{x\},\\bm\{v\}\)=\\frac\{\(\\bm\{v\}\_\{0\}\+2\\bm\{v\}\)\{\\tau\}\+3\(\\bm\{x\}\_\{0\}\-\\bm\{x\}\)\}\{\{\\tau\}^\{2\}\}\+\\frac\{\(\\bm\{v\}\_\{T\}\+2\\bm\{v\}\)\(T\-\{\\tau\}\)\+3\(\\bm\{x\}\-\\bm\{x\}\_\{T\}\)\}\{\(T\-\{\\tau\}\)^\{2\}\}=0,\(29\)
Denote the cubic interpolation from\(๐๐,๐๐\)\(\\bm\{x\_\{0\}\},\\bm\{v\_\{0\}\}\)to\(๐,๐\)\(\\bm\{x\},\\bm\{v\}\)on\[0,ฯ\]\[0,\\tau\]and\(๐,๐\)\(\\bm\{x\},\\bm\{v\}\)to\(๐T,๐T\)\(\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\)on\[ฯ,T\]\[\\tau,T\]byฮณ\[0,ฯ\]โ\(t\)\\gamma\_\{\[0,\\tau\]\}\(t\)andฮณ\[ฯ,T\]โ\(t\)\\gamma\_\{\[\\tau,T\]\}\(\{t\}\), respectively\.The coefficients are known according to[SectionA\.1](https://arxiv.org/html/2608.21070#A1.SS1)\. Then we have:
ฮณ\[0,ฯ\]โฒโฒ\(ฯโ\)=6โ
\(๐0\+๐\)โฯโ2โ\(๐โ๐0\)ฯ2\+2โ
3โ\(๐โ๐0\)โ\(2โ๐0\+๐\)โฯฯ2=2โ
3โ\(๐๐โ๐\)\+\(๐๐\+2โ๐\)โฯฯ2,\\gamma\_\{\[0,\\tau\]\}^\{\{\}^\{\\prime\\prime\}\}\(\\tau^\{\-\}\)=6\\cdot\\frac\{\(\\bm\{v\}\_\{0\}\+\\bm\{v\}\)\\tau\-2\(\\bm\{x\}\-\\bm\{x\}\_\{0\}\)\}\{\\tau^\{2\}\}\+2\\cdot\\frac\{3\{\(\\bm\{x\}\-\\bm\{x\}\_\{0\}\)\-\(2\\bm\{v\}\_\{0\}\+\\bm\{v\}\)\}\\tau\}\{\\tau^\{2\}\}=2\\cdot\\frac\{3\(\\bm\{x\_\{0\}\}\-\\bm\{x\}\)\+\(\\bm\{v\_\{0\}\}\+2\\bm\{v\}\)\\tau\}\{\\tau^\{2\}\},\(30\)ฮณ\[ฯ,T\]โฒโฒ\(ฯ\+\)=6โ
\(๐\+๐T\)โฯโ2โ\(๐Tโ๐\)ฯ2\+2โ
3โ\(๐Tโ๐\)โ\(2โ๐\+๐T\)โ\(Tโฯ\)\(Tโฯ\)2\\displaystyle\\gamma\_\{\[\\tau,T\]\}^\{\{\}^\{\\prime\\prime\}\}\(\\tau^\{\+\}\)=6\\cdot\\frac\{\(\\bm\{v\}\+\\bm\{v\}\_\{T\}\)\\tau\-2\(\\bm\{x\}\_\{T\}\-\\bm\{x\}\)\}\{\\tau^\{2\}\}\+2\\cdot\\frac\{3\(\\bm\{x\}\_\{T\}\-\\bm\{x\}\)\-\(2\\bm\{v\}\+\\bm\{v\}\_\{T\}\)\(T\-\\tau\)\}\{\(T\-\\tau\)^\{2\}\}\(31\)=2โ
3โ\(๐โ๐T\)\+\(๐\+2โ๐T\)โ\(Tโฯ\)\(Tโฯ\)2\\displaystyle=2\\cdot\\frac\{3\(\\bm\{x\}\-\\bm\{x\}\_\{T\}\)\+\(\\bm\{v\}\+2\\bm\{v\}\_\{T\}\)\(T\-\\tau\)\}\{\(T\-\\tau\)^\{2\}\}ฮณ\[0,ฯ\]โฒโฒโฒ\(ฯโ\)=6โ
\(๐0\+๐\)โฯโ2โ\(๐โ๐0\)ฯ2\\gamma\_\{\[0,\\tau\]\}^\{\{\}^\{\\prime\\prime\\prime\}\}\(\\tau^\{\-\}\)=6\\cdot\\frac\{\(\\bm\{v\}\_\{0\}\+\\bm\{v\}\)\\tau\-2\(\\bm\{x\}\-\\bm\{x\}\_\{0\}\)\}\{\\tau^\{2\}\}\(32\)
ฮณ\[ฯ,T\]โฒโฒโฒ\(ฯ\+\)=6โ
\(๐\+๐T\)โ\(Tโฯ\)โ2โ\(๐Tโ๐\)\(Tโฯ\)2\\displaystyle\\gamma\_\{\[\\tau,T\]\}^\{\{\}^\{\\prime\\prime\\prime\}\}\(\\tau^\{\+\}\)=6\\cdot\\frac\{\(\\bm\{v\}\+\\bm\{v\}\_\{T\}\)\(T\-\\tau\)\-2\(\\bm\{x\}\_\{T\}\-\\bm\{x\}\)\}\{\(T\-\\tau\)^\{2\}\}\(33\)
Then we can find that[28](https://arxiv.org/html/2608.21070#A1.E28)and[29](https://arxiv.org/html/2608.21070#A1.E29)appear to beฮณ\[0,ฯ\]โฒโฒโฒ\(ฯโ\)=ฮณ\[ฯ,T\]โฒโฒโฒ\(ฯ\+\),ฮณ\[0,ฯ\]โฒโฒ\(ฯโ\)=ฮณ\[ฯ,T\]โฒโฒ\(ฯ\+\)\.\\gamma\_\{\[0,\\tau\]\}^\{\{\}^\{\\prime\\prime\\prime\}\}\(\\tau^\{\-\}\)=\\gamma\_\{\[\\tau,T\]\}^\{\{\}^\{\\prime\\prime\\prime\}\}\(\\tau^\{\+\}\),\\gamma\_\{\[0,\\tau\]\}^\{\{\}^\{\\prime\\prime\}\}\(\\tau^\{\-\}\)=\\gamma\_\{\[\\tau,T\]\}^\{\{\}^\{\\prime\\prime\}\}\(\\tau^\{\+\}\)\.Because the two curves are cubic, andฮณ\[0,ฯ\]\(i\)\(ฯโ\)=ฮณ\[ฯ,T\]\(i\)\(ฯ\+\),i=0,1,2,3\\gamma^\{\(i\)\}\_\{\[0,\\tau\]\}\(\\tau^\{\-\}\)=\\gamma\_\{\[\\tau,T\]\}^\{\(i\)\}\(\\tau^\{\+\}\),i=0,1,2,3, so we know that in the optimal caseฮณ\[0,ฯ\]=ฮณโ\|\[0,ฯ\],ฮณ\[ฯ,T\]=ฮณโ\|\[ฯ,T\]\.\\gamma\_\{\[0,\\tau\]\}=\\gamma^\{\*\}\\big\|\_\{\[0,\\tau\]\},\\gamma\_\{\[\\tau,T\]\}=\\gamma^\{\*\}\\big\|\_\{\[\\tau,T\]\}\.\. That is, the unique solution of[28](https://arxiv.org/html/2608.21070#A1.E28)and[29](https://arxiv.org/html/2608.21070#A1.E29)is\(ฮณโ\(ฯ\),ฮณโโฒ\(ฯ\)\),\(\\gamma^\{\*\}\(\\tau\),\{\\gamma^\{\*\}\}^\{\{\}^\{\\prime\}\}\(\\tau\)\),which completes the proof\. โ
Back to our target, assumeฮณ\\gammaandฮณ~\\tilde\{\\gamma\}cross at\(๐ฯ,๐ฯ\)\(\\bm\{x\}\_\{\\tau\},\\bm\{v\}\_\{\\tau\}\), then we have
๐\[0,T\]โ\(๐0,๐0,๐T,๐T\)\+๐\[0,T\]โ\(๐~0,๐~0,๐~T,๐~T\)\\displaystyle\\mathcal\{C\}\_\{\[0,T\]\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\},\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\)\+\\mathcal\{C\}\_\{\[0,T\]\}\(\\tilde\{\\bm\{x\}\}\_\{0\},\\tilde\{\\bm\{v\}\}\_\{0\},\\tilde\{\\bm\{x\}\}\_\{T\},\\tilde\{\\bm\{v\}\}\_\{T\}\)\(34\)=\\displaystyle=\{๐\[0,ฯ\]โ\(๐0,๐0,๐ฯ,๐ฯ\)\+๐\[ฯ,T\]โ\(๐ฯ,๐ฯ,๐T,๐T\)\}\+\{๐\[0,ฯ\]โ\(๐~0,๐~0,๐ฯ,๐ฯ\)\+๐\[ฯ,T\]โ\(๐ฯ,๐ฯ,๐~T,๐~T\)\}\\displaystyle\\left\\\{\\mathcal\{C\}\_\{\[0,\\tau\]\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\},\\bm\{x\}\_\{\\tau\},\\bm\{v\}\_\{\\tau\}\)\+\\mathcal\{C\}\_\{\[\\tau,T\]\}\(\\bm\{x\}\_\{\\tau\},\\bm\{v\}\_\{\\tau\},\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\)\\right\\\}\+\\left\\\{\\mathcal\{C\}\_\{\[0,\\tau\]\}\(\\tilde\{\\bm\{x\}\}\_\{0\},\\tilde\{\\bm\{v\}\}\_\{0\},\\bm\{x\}\_\{\\tau\},\\bm\{v\}\_\{\\tau\}\)\+\\mathcal\{C\}\_\{\[\\tau,T\]\}\(\\bm\{x\}\_\{\\tau\},\\bm\{v\}\_\{\\tau\},\\tilde\{\\bm\{x\}\}\_\{T\},\\tilde\{\\bm\{v\}\}\_\{T\}\)\\right\\\}=\\displaystyle=\{๐\[0,ฯ\]โ\(๐0,๐0,๐ฯ,๐ฯ\)\+๐\[ฯ,T\]โ\(๐ฯ,๐ฯ,๐~T,๐~T\)\}\+\{๐\[0,ฯ\]โ\(๐~0,๐~0,๐ฯ,๐ฯ\)\+๐\[ฯ,T\]โ\(๐ฯ,๐ฯ,๐T,๐T\)\}\\displaystyle\\left\\\{\\mathcal\{C\}\_\{\[0,\\tau\]\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\},\\bm\{x\}\_\{\\tau\},\\bm\{v\}\_\{\\tau\}\)\+\\mathcal\{C\}\_\{\[\\tau,T\]\}\(\\bm\{x\}\_\{\\tau\},\\bm\{v\}\_\{\\tau\},\\tilde\{\\bm\{x\}\}\_\{T\},\\tilde\{\\bm\{v\}\}\_\{T\}\)\\right\\\}\+\\left\\\{\\mathcal\{C\}\_\{\[0,\\tau\]\}\(\\tilde\{\\bm\{x\}\}\_\{0\},\\tilde\{\\bm\{v\}\}\_\{0\},\\bm\{x\}\_\{\\tau\},\\bm\{v\}\_\{\\tau\}\)\+\\mathcal\{C\}\_\{\[\\tau,T\]\}\(\\bm\{x\}\_\{\\tau\},\\bm\{v\}\_\{\\tau\},\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\)\\right\\\}โฅ\\displaystyle\\geq๐\[0,T\]โ\(๐0,๐0,๐~T,๐~T\)\+๐\[0,T\]โ\(๐~0,๐~0,๐T,๐T\)\.\\displaystyle\\mathcal\{C\}\_\{\[0,T\]\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\},\\tilde\{\\bm\{x\}\}\_\{T\},\\tilde\{\\bm\{v\}\}\_\{T\}\)\+\\mathcal\{C\}\_\{\[0,T\]\}\(\\tilde\{\\bm\{x\}\}\_\{0\},\\tilde\{\\bm\{v\}\}\_\{0\},\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\)\.
The equality holds if and only ifฮณ\\gammaandฮณ~\\tilde\{\\gamma\}coincide\. Consequently, the strict inequality holds for distinct trajectories, which implies the non\-crossing property\. This completes the proof\. โ
### A\.3Proof of Theorem 4\.4
###### Proof\.
Consider the time derivative of the marginal distribution:
โtฯtโ\(๐,๐\)\\displaystyle\\partial\_\{t\}\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)=โt\{โซฯtโ\(๐,๐\|z\)โqโ\(z\)โ๐z\}\\displaystyle=\\partial\_\{t\}\\left\\\{\\int\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|z\)q\(z\)\\mathrm\{d\}z\\right\\\}=โซโtฯtโ\(๐,๐\|z\)โqโ\(z\)โ๐z\\displaystyle=\\int\\partial\_\{t\}\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|z\)q\(z\)\\mathrm\{d\}z=โโซโ๐โ
\(ฯt\(๐,๐\|z\)๐\)q\(z\)dzโโซโ๐โ
\(ฯt\(๐,๐\|z\)๐\(๐,๐,t\|z\)\)q\(z\)dz\\displaystyle=\-\\int\\nabla\_\{\\bm\{x\}\}\\cdot\(\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|z\)\\bm\{v\}\)q\(z\)\\mathrm\{d\}z\-\\int\\nabla\_\{\\bm\{v\}\}\\cdot\(\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|z\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\|z\)\)q\(z\)\\mathrm\{d\}z=โโ๐โ
\(โซฯt\(๐,๐\|z\)๐q\(z\)dz\)โโ๐โ
\(โซฯt\(๐,๐\|z\)๐\(๐,๐,t\|z\)q\(z\)dz\)\\displaystyle=\-\\nabla\_\{\\bm\{x\}\}\\cdot\\left\(\\int\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|z\)\\bm\{v\}q\(z\)\\mathrm\{d\}z\\right\)\-\\nabla\_\{\\bm\{v\}\}\\cdot\\left\(\\int\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|z\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\|z\)q\(z\)\\mathrm\{d\}z\\right\)=โโ๐โ
\(ฯt\(๐,๐\)๐\)โโ๐โ
\(๐t\(๐,๐\)ฯt\(๐,๐\)\)\\displaystyle=\-\\nabla\_\{\\bm\{x\}\}\\cdot\(\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\\bm\{v\}\)\-\\nabla\_\{\\bm\{v\}\}\\cdot\(\\bm\{a\}\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\)\(35\)Therefore,ฯtโ\(๐,๐\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)satisfies the continuity equation:
โtฯt\(๐,๐\)=โโ๐โ
\(ฯt\(๐,๐\)๐\)โโ๐โ
\(๐t\(๐,๐\)ฯt\(๐,๐\)\)\\partial\_\{t\}\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)=\-\\nabla\_\{\\bm\{x\}\}\\cdot\(\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\\bm\{v\}\)\-\\nabla\_\{\\bm\{v\}\}\\cdot\(\\bm\{a\}\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\)\(36\)This completes the proof\. โ
### A\.4Proof of Theorem 4\.5
###### Proof\.
Consider the gradient of the AFM Loss and CAFM Loss:
โ๐ฝโAFM\\displaystyle\\nabla\_\{\\bm\{\\theta\}\}\\mathcal\{L\}\_\{\\text\{AFM\}\}=โฮธ๐ผฯtโ\(๐,๐\)โ\(โ๐ฮธโ\(๐,๐,t\)โ2โ2โ๐ฮธTโ\(๐,๐,t\)โ๐โ\(๐,๐,t\)\)\\displaystyle=\\nabla\_\{\\theta\}\\mathbb\{E\}\_\{\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\}\(\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{x\},\\bm\{v\},t\)\\\|^\{2\}\-2\\bm\{a\}\_\{\\theta\}^\{T\}\(\\bm\{x\},\\bm\{v\},t\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\)\)\(37\)โ๐ฝโCAFM\\displaystyle\\nabla\_\{\\bm\{\\theta\}\}\\mathcal\{L\}\_\{\\text\{CAFM\}\}=โฮธ๐ผฯtโ\(๐,๐\|๐\)โqโ\(๐\)โ\(โ๐ฮธโ\(๐,๐,t\)โ2โ2โ๐ฮธTโ\(๐,๐,t\)โ๐โ\(๐,๐,t\|๐\)\)\\displaystyle=\\nabla\_\{\\theta\}\\mathbb\{E\}\_\{\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\)q\(\\bm\{z\}\)\}\(\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{x\},\\bm\{v\},t\)\\\|^\{2\}\-2\\bm\{a\}\_\{\\theta\}^\{T\}\(\\bm\{x\},\\bm\{v\},t\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\|\\bm\{z\}\)\)\(38\)For the first term ofโ๐ฝโAFM\\nabla\_\{\\bm\{\\theta\}\}\\mathcal\{L\}\_\{\\text\{AFM\}\}:
๐ผฯtโ\(๐,๐\)โโ๐ฮธโ\(๐,๐,t\)โ2\\displaystyle\\mathbb\{E\}\_\{\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\}\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{x\},\\bm\{v\},t\)\\\|^\{2\}=โซฯtโ\(๐,๐\)โโ๐ฮธโ\(๐,๐,t\)โ2โ๐๐โ๐๐\\displaystyle=\\int\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{x\},\\bm\{v\},t\)\\\|^\{2\}\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{v\}=โซฯtโ\(๐,๐\|z\)โqโ\(๐\)โโ๐ฮธโ\(๐,๐,t\)โ2โ๐๐โ๐๐โ๐z\\displaystyle=\\int\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|z\)q\(\\bm\{z\}\)\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{x\},\\bm\{v\},t\)\\\|^\{2\}\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{v\}\\mathrm\{d\}z=๐ผฯtโ\(๐,๐\|๐\)โqโ\(๐\)โโ๐ฮธโ\(๐,๐,t\)โ2\\displaystyle=\\mathbb\{E\}\_\{\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\)q\(\\bm\{z\}\)\}\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{x\},\\bm\{v\},t\)\\\|^\{2\}\(39\)For the second term:
๐ผฯtโ\(๐,๐\)โ\[๐ฮธTโ\(๐,๐,t\)โ๐โ\(๐,๐,t\)\]\\displaystyle\\mathbb\{E\}\_\{\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\}\[\\bm\{a\}\_\{\\theta\}^\{T\}\(\\bm\{x\},\\bm\{v\},t\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\)\]=โซฯtโ\(๐,๐\)โ๐ฮธTโ\(๐,๐,t\)โ๐โ\(๐,๐,t\)โ๐๐โ๐๐\\displaystyle=\\int\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\\bm\{a\}\_\{\\theta\}^\{T\}\(\\bm\{x\},\\bm\{v\},t\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{v\}=โซฯtโ\(๐,๐\)โ๐ฮธTโ\(๐,๐,t\)โ\(โซqโก\(๐\)โ๐โก\(๐,๐,t\|๐\)โฯtโ\(๐,๐\|๐\)ฯtโ\(๐,๐\)โ๐z\)โ๐๐โ๐๐\\displaystyle=\\int\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\\bm\{a\}\_\{\\theta\}^\{T\}\(\\bm\{x\},\\bm\{v\},t\)\\left\(\\int q\(\\bm\{z\}\)\\dfrac\{\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\|\\bm\{z\}\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\)\}\{\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\}\\mathrm\{d\}z\\right\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{v\}=โซโซโก๐ฮธTโ\(๐,๐,t\)โ๐โ\(๐,๐,t\|๐\)โฯtโ\(๐,๐\|๐\)โqโ\(z\)โ๐๐โ๐๐โ๐๐\\displaystyle=\\int\\int\\bm\{a\}\_\{\\theta\}^\{T\}\(\\bm\{x\},\\bm\{v\},t\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\|\\bm\{z\}\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\)q\(z\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{v\}\\mathrm\{d\}\\bm\{z\}=๐ผฯtโ\(๐,๐\|๐\)โqโ\(๐\)โ\(๐ฮธTโ\(๐,๐,t\)โ๐โ\(๐,๐,t\|๐\)\)\\displaystyle=\\mathbb\{E\}\_\{\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\)q\(\\bm\{z\}\)\}\(\\bm\{a\}\_\{\\theta\}^\{T\}\(\\bm\{x\},\\bm\{v\},t\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\|\\bm\{z\}\)\)\(40\)โ
### A\.5Proof of Theorem 4\.6
###### Proof\.
Consider a given timetit\_\{i\}\. Based on the expressions for๐๐,๐๐\\bm\{m\}\_\{\\bm\{x\}\},\\bm\{m\}\_\{\\bm\{v\}\}, at this moment we exactly have:
ฯtiโ\(๐,๐\|๐\)=๐ฉโก\(๐\|๐i,ฯx2โ๐ฐ\)โ
๐ฉโก\(๐\|๐i,ฯv2โ๐ฐ\)\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\)=\\mathcal\{N\}\(\\bm\{x\}\|\\bm\{x\}\_\{i\},\\sigma\_\{x\}^\{2\}\\bm\{I\}\)\\cdot\\mathcal\{N\}\(\\bm\{v\}\|\\bm\{v\}\_\{i\},\\sigma\_\{v\}^\{2\}\\bm\{I\}\)\(41\)
For simplicity, letdโ๐j=dโ๐jโdโ๐j\\mathrm\{d\}\\bm\{s\}\_\{j\}=\\mathrm\{d\}\\bm\{x\}\_\{j\}\\mathrm\{d\}\\bm\{v\}\_\{j\}Therefore, the Marginal Probability Path is:
ฯtiโ\(๐,๐\)\\displaystyle\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\)=โซฯtiโ\(๐,๐\|๐\)โqโ\(๐\)โ๐๐\\displaystyle=\\int\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\)q\(\\bm\{z\}\)\\mathrm\{d\}\\bm\{z\}=โซฯti\(๐,๐\|\(๐i,๐i\)\)ฮผ0\(๐0,๐0\)โ
โj=0Kโ1๐ฆjโj\+1โ\[\(๐j,๐j\),\(๐j\+1,๐j\+1\)\]d๐0โฏd๐K\\displaystyle=\\int\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\|\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\)\)\\mu\_\{0\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\)\\cdot\\prod\_\{j=0\}^\{K\-1\}\\mathcal\{K\}^\{\*\}\_\{j\\rightarrow j\+1\}\[\(\\bm\{x\}\_\{j\},\\bm\{v\}\_\{j\}\),\(\\bm\{x\}\_\{j\+1\},\\bm\{v\}\_\{j\+1\}\)\]\\mathrm\{d\}\\bm\{s\}\_\{0\}\\cdots\\mathrm\{d\}\\bm\{s\}\_\{K\}\(42\)Using the definition of the propagator, we know that:
โซ๐ฆjโj\+1โโ\[\(๐j,๐j\),\(๐j\+1,๐j\+1\)\]โdโ๐j\+1โdโ๐j\+1=1\\int\\mathcal\{K\}^\{\*\}\_\{j\\rightarrow j\+1\}\[\(\\bm\{x\}\_\{j\},\\bm\{v\}\_\{j\}\),\(\\bm\{x\}\_\{j\+1\},\\bm\{v\}\_\{j\+1\}\)\]\\mathrm\{d\}\\bm\{x\}\_\{j\+1\}\\mathrm\{d\}\\bm\{v\}\_\{j\+1\}=1\(43\)
Thus, we can integrate out the future terms \(j\>ij\>i\) first , and then iteratively collapse the past terms \(j<ij<i\):
ฯtiโ\(๐,๐\)\\displaystyle\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\)=โซฯti\(๐,๐\|\(๐i,๐i\)\)ฮผ0\(๐0,๐0\)โ
โj=0Kโ1๐ฆjโj\+1โ\[\(๐j,๐j\),\(๐j\+1,๐j\+1\)\]d๐0โฏd๐K\\displaystyle=\\int\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\|\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\)\)\\mu\_\{0\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\)\\cdot\\prod\_\{j=0\}^\{K\-1\}\\mathcal\{K\}^\{\*\}\_\{j\\rightarrow j\+1\}\[\(\\bm\{x\}\_\{j\},\\bm\{v\}\_\{j\}\),\(\\bm\{x\}\_\{j\+1\},\\bm\{v\}\_\{j\+1\}\)\]\\mathrm\{d\}\\bm\{s\}\_\{0\}\\cdots\\mathrm\{d\}\\bm\{s\}\_\{K\}=โซฯti\(๐,๐\|\(๐i,๐i\)\)ฮผ0\(๐0,๐0\)โ
โj=0iโ1๐ฆjโj\+1โ\[\(๐j,๐j\),\(๐j\+1,๐j\+1\)\]d๐0โฏd๐i\\displaystyle=\\int\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\|\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\)\)\\mu\_\{0\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\)\\cdot\\prod\_\{j=0\}^\{i\-1\}\\mathcal\{K\}^\{\*\}\_\{j\\rightarrow j\+1\}\[\(\\bm\{x\}\_\{j\},\\bm\{v\}\_\{j\}\),\(\\bm\{x\}\_\{j\+1\},\\bm\{v\}\_\{j\+1\}\)\]\\mathrm\{d\}\\bm\{s\}\_\{0\}\\cdots\\mathrm\{d\}\\bm\{s\}\_\{i\}=โซฯti\(๐,๐\|\(๐i,๐i\)\)ฯ0โ1โ\[\(๐0,๐0\),\(๐1,๐1\)\]โj=1iโ1๐ฆjโj\+1โ\[\(๐j,๐j\),\(๐j\+1,๐j\+1\)\]d๐0โฏd๐i\\displaystyle=\\int\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\|\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\)\)\\pi^\{\*\}\_\{0\\rightarrow 1\}\[\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\),\(\\bm\{x\}\_\{1\},\\bm\{v\}\_\{1\}\)\]\\prod\_\{j=1\}^\{i\-1\}\\mathcal\{K\}^\{\*\}\_\{j\\rightarrow j\+1\}\[\(\\bm\{x\}\_\{j\},\\bm\{v\}\_\{j\}\),\(\\bm\{x\}\_\{j\+1\},\\bm\{v\}\_\{j\+1\}\)\]\\mathrm\{d\}\\bm\{s\}\_\{0\}\\cdots\\mathrm\{d\}\\bm\{s\}\_\{i\}=โซฯti\(๐,๐\|\(๐i,๐i\)\)ฮผ1\(๐1,๐1\)โj=1iโ1๐ฆjโj\+1โ\[\(๐j,๐j\),\(๐j\+1,๐j\+1\)\]d๐1โฏd๐i\\displaystyle=\\int\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\|\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\)\)\\mu\_\{1\}\(\\bm\{x\}\_\{1\},\\bm\{v\}\_\{1\}\)\\prod\_\{j=1\}^\{i\-1\}\\mathcal\{K\}^\{\*\}\_\{j\\rightarrow j\+1\}\[\(\\bm\{x\}\_\{j\},\\bm\{v\}\_\{j\}\),\(\\bm\{x\}\_\{j\+1\},\\bm\{v\}\_\{j\+1\}\)\]\\mathrm\{d\}\\bm\{s\}\_\{1\}\\cdots\\mathrm\{d\}\\bm\{s\}\_\{i\}โฎ\\displaystyle\\quad\\vdots=โซฯtiโ\(๐,๐\|\(๐i,๐i\)\)โฮผiโ\(๐i,๐i\)โdโ๐i\\displaystyle=\\int\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\|\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\)\)\\mu\_\{i\}\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\)\\mathrm\{d\}\\bm\{s\}\_\{i\}=โซ\(๐ฉโก\(๐\|๐i,ฯx2โ๐ฐ\)โ
๐ฉโก\(๐\|๐i,ฯv2โ๐ฐ\)\)โฮผiโ\(๐i,๐i\)โdโ๐i\\displaystyle=\\int\(\\mathcal\{N\}\(\\bm\{x\}\|\\bm\{x\}\_\{i\},\\sigma\_\{x\}^\{2\}\\bm\{I\}\)\\cdot\\mathcal\{N\}\(\\bm\{v\}\|\\bm\{v\}\_\{i\},\\sigma\_\{v\}^\{2\}\\bm\{I\}\)\)\\mu\_\{i\}\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\)\\mathrm\{d\}\\bm\{s\}\_\{i\}=\(ฮผiโg\)โ\(๐,๐\)\\displaystyle=\(\\mu\_\{i\}\*g\)\(\\bm\{x\},\\bm\{v\}\)\(44\)wheregโก\(๐,๐\)=\(๐ฉโก\(๐\|๐,ฯx2โ๐ฐ\)โ
๐ฉโก\(๐\|๐,ฯv2โ๐ฐ\)\)g\(\\bm\{x\},\\bm\{v\}\)=\(\\mathcal\{N\}\(\\bm\{x\}\|\\bm\{0\},\\sigma\_\{x\}^\{2\}\\bm\{I\}\)\\cdot\\mathcal\{N\}\(\\bm\{v\}\|\\bm\{0\},\\sigma\_\{v\}^\{2\}\\bm\{I\}\)\)\. When\(ฯ๐,ฯ๐\)โ๐\(\\sigma\_\{\\bm\{x\}\},\\sigma\_\{\\bm\{v\}\}\)\\rightarrow\\bm\{0\}, we haveฯtiโ\(๐,๐\)โฮผiโ\(๐,๐\)\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\)\\rightarrow\\mu\_\{i\}\(\\bm\{x\},\\bm\{v\}\)in distribution\.
Ifฮผiโ\(๐,๐\)\\mu\_\{i\}\(\\bm\{x\},\\bm\{v\}\)is absolutely continuous with respect to the Lebesgue measure and admits a density inL1โ\(โ2โd\)L^\{1\}\(\\mathbb\{R\}^\{2d\}\), we can guarantee strong convergence in theL1L^\{1\}norm:limฯโ0โฯtiโ\(๐,๐\)โฮผiโ\(๐,๐\)โL1=0\\lim\_\{\\sigma\\to 0\}\\\|\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\)\-\\mu\_\{i\}\(\\bm\{x\},\\bm\{v\}\)\\\|\_\{L^\{1\}\}=0for alli=0,1,โฆ,Ki=0,1,\.\.\.,K\.
โ
### A\.6Proof of Proposition 4\.7
First, we prove that the Marginal Probability Path constructed by TracingFlow converges to the optimal solution of the DOAT problem whenฯ๐โ0,ฯ๐โ0\\sigma\_\{\\bm\{x\}\}\\rightarrow 0,\\sigma\_\{\\bm\{v\}\}\\rightarrow 0\.
We investigate directly in the augmented space\. Let๐=\[๐,๐\]Tโ๐ณร๐ฑ\\bm\{s\}=\[\\bm\{x\},\\bm\{v\}\]^\{T\}\\in\\mathcal\{X\}\\times\\mathcal\{V\}, and consider the cost of the DOAT problem\. The subsequent position of a particle starting from the initial point๐0\\bm\{s\}\_\{0\}can be represented by the flow mapฯtโ\(๐0\)\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\. Therefore, the cost can be written as:
โDOATโ\(๐,ฯ\)\\displaystyle\\mathcal\{L\}\_\{\\text\{DOAT\}\}\(\\bm\{a\},\\rho\)=โซt0tkโซ๐ณร๐ฑ12โโ๐โก\(๐,t\)โ2โฯtโ\(๐\)โ๐๐โ๐๐โ๐t\\displaystyle=\\int\_\{t\_\{0\}\}^\{t\_\{k\}\}\\int\_\{\\mathcal\{X\}\\times\\mathcal\{V\}\}\\dfrac\{1\}\{2\}\\\|\\bm\{a\}\(\\bm\{s\},t\)\\\|^\{2\}\\rho\_\{t\}\(\\bm\{s\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{v\}\\mathrm\{d\}t=โซt0tkโซ๐ณร๐ฑ12\|๐โก\(ฯtโ\(๐0\),t\)\|det2โก\(โฯโ๐0\)โ1โฯ0โ\(๐0\)โdet\(โฯโ๐0\)โdโ๐0โ๐t\\displaystyle=\\int\_\{t\_\{0\}\}^\{t\_\{k\}\}\\int\_\{\\mathcal\{X\}\\times\\mathcal\{V\}\}\\dfrac\{1\}\{2\}\\\|\\bm\{a\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\\|^\{2\}\\det\\left\(\\dfrac\{\\partial\\bm\{\\phi\}\}\{\\partial\\bm\{s\}\_\{0\}\}\\right\)^\{\-1\}\\rho\_\{0\}\(\\bm\{s\}\_\{0\}\)\\det\\left\(\\dfrac\{\\partial\\bm\{\\phi\}\}\{\\partial\\bm\{s\}\_\{0\}\}\\right\)\\mathrm\{d\}\\bm\{s\}\_\{0\}\\mathrm\{d\}t=โซ๐ณร๐ฑฯ0โ\(๐0\)โdโ๐0โโซt0tk12โโ๐โก\(ฯtโ\(๐0\),t\)โ2โ๐t\\displaystyle=\\int\_\{\\mathcal\{X\}\\times\\mathcal\{V\}\}\\rho\_\{0\}\(\\bm\{s\}\_\{0\}\)\\mathrm\{d\}\\bm\{s\}\_\{0\}\\int\_\{t\_\{0\}\}^\{t\_\{k\}\}\\dfrac\{1\}\{2\}\\\|\\bm\{a\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\\|^\{2\}\\mathrm\{d\}t=โซฮฉ\(โซt0tk12โโ๐ธโฒโฒโ\(t\)โ2โ๐t\)โ๐โโ\[๐ธ\]\\displaystyle=\\int\_\{\\Omega\}\\left\(\\int\_\{t\_\{0\}\}^\{t\_\{k\}\}\\dfrac\{1\}\{2\}\\\|\\bm\{\\gamma\}^\{\\prime\\prime\}\(t\)\\\|^\{2\}\\mathrm\{d\}t\\right\)\\mathrm\{d\}\\mathbb\{P\}\[\\bm\{\\gamma\}\]\(45\)
whereโ\\mathbb\{P\}is a probability measure in the path space, satisfying the constraint\(eti\)\#โโ=ฮผiโ\(๐,๐\)\(e\_\{t\_\{i\}\}\)\_\{\\\#\}\\mathbb\{P\}=\\mu\_\{i\}\(\\bm\{x\},\\bm\{v\}\)\. Here,ete\_\{t\}is the evaluation mapetโ\[๐ธ\]=๐ธโ\(t\)e\_\{t\}\[\\bm\{\\gamma\}\]=\\bm\{\\gamma\}\(t\), and\(eti\)\#โโ\(e\_\{t\_\{i\}\}\)\_\{\\\#\}\\mathbb\{P\}denotes the push\-forward of the probability measure by the evaluation map\. According to[Section4\.1](https://arxiv.org/html/2608.21070#S4.SS1.SSS0.Px1), the minimum cost path connecting๐i\\bm\{s\}\_\{i\}and๐i\+1\\bm\{s\}\_\{i\+1\}is a cubic spline, and the minimum cost is๐iโi\+1โ\[๐i,๐i\+1\]\\mathcal\{C\}\_\{i\\rightarrow i\+1\}\[\\bm\{s\}\_\{i\},\\bm\{s\}\_\{i\+1\}\]\. Therefore, for any path๐ธ\\bm\{\\gamma\}:
โซtiti\+112โโ๐ธโฒโฒโ\(t\)โ2โ๐tโฅ๐iโi\+1โ\[๐i,๐i\+1\]\\int\_\{t\_\{i\}\}^\{t\_\{i\+1\}\}\\dfrac\{1\}\{2\}\\\|\\bm\{\\gamma\}^\{\\prime\\prime\}\(t\)\\\|^\{2\}\\mathrm\{d\}t\\geq\\mathcal\{C\}\_\{i\\rightarrow i\+1\}\[\\bm\{s\}\_\{i\},\\bm\{s\}\_\{i\+1\}\]\(46\)
Thus, the DOAT problem cost has a lower bound:
โDOATโ\(ฯ,๐\)\\displaystyle\\mathcal\{L\}\_\{\\text\{DOAT\}\}\(\\rho,\\bm\{a\}\)=โซฮฉโi=0Kโ1\(โซtiti\+112โโ๐ธโฒโฒโ\(t\)โ2โ๐t\)โ๐โโ\[๐ธ\]\\displaystyle=\\int\_\{\\Omega\}\\sum\_\{i=0\}^\{K\-1\}\\left\(\\int\_\{t\_\{i\}\}^\{t\_\{i\+1\}\}\\dfrac\{1\}\{2\}\\\|\\bm\{\\gamma\}^\{\\prime\\prime\}\(t\)\\\|^\{2\}\\mathrm\{d\}t\\right\)\\mathrm\{d\}\\mathbb\{P\}\[\\bm\{\\gamma\}\]โฅโซฮฉโi=0Kโ1๐iโi\+1โ\[๐i,๐i\+1\]โ๐โโ\[๐ธ\]\\displaystyle\\geq\\int\_\{\\Omega\}\\sum\_\{i=0\}^\{K\-1\}\\mathcal\{C\}\_\{i\\rightarrow i\+1\}\[\\bm\{s\}\_\{i\},\\bm\{s\}\_\{i\+1\}\]\\mathrm\{d\}\\mathbb\{P\}\[\\bm\{\\gamma\}\]โฅโซโi=0Kโ1๐iโi\+1โ\[๐i,๐i\+1\]โdโฯโโ\(๐i,๐i\+1\)\\displaystyle\\geq\\int\\sum\_\{i=0\}^\{K\-1\}\\mathcal\{C\}\_\{i\\rightarrow i\+1\}\[\\bm\{s\}\_\{i\},\\bm\{s\}\_\{i\+1\}\]\\mathrm\{d\}\\pi^\{\*\}\(\\bm\{s\}\_\{i\},\\bm\{s\}\_\{i\+1\}\)\(47\)
whereฯโโ\(๐i,๐i\+1\)\\pi^\{\*\}\(\\bm\{s\}\_\{i\},\\bm\{s\}\_\{i\+1\}\)is the Optimal Transport Plan for the SOAT problem betweenฮผโก\(๐i\)\\mu\(\\bm\{s\}\_\{i\}\)andฮผโก\(๐i\+1\)\\mu\(\\bm\{s\}\_\{i\+1\}\)\.
Consider the path constructed by TracingFlow\. In the case whereฯ๐,ฯ๐โ0\\sigma\_\{\\bm\{x\}\},\\sigma\_\{\\bm\{v\}\}\\rightarrow 0, we haveฯtโ\(๐,๐\|๐\)=ฯtโ\(๐\|๐\)โฮดโก\(๐โ๐ธ๐โโ\(t\)\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\)=\\rho\_\{t\}\(\\bm\{s\}\|\\bm\{z\}\)\\rightarrow\\delta\(\\bm\{s\}\-\\bm\{\\gamma\}^\{\*\}\_\{\\bm\{z\}\}\(t\)\), where๐ธ๐โ\\bm\{\\gamma\}^\{\*\}\_\{\\bm\{z\}\}is the optimal single\-particle trajectory connecting๐=\[๐0,๐1,โฏ,๐K\]\\bm\{z\}=\[\\bm\{s\}\_\{0\},\\bm\{s\}\_\{1\},\\cdots,\\bm\{s\}\_\{K\}\]\. We calculate the transport cost of the constructed path:
โTF\\displaystyle\\mathcal\{L\}\_\{\\text\{TF\}\}=โซq\(๐\)d๐โi=0Kโ1โซtiti\+1โฅ๐ธโ๐โฒโฒ\(t\)โฅ2dt\\displaystyle=\\int q\(\\bm\{z\}\)\\mathrm\{d\}\\bm\{z\}\\sum\_\{i=0\}^\{K\-1\}\\int\_\{t\_\{i\}\}^\{t\_\{i\+1\}\}\\\|\\bm\{\\gamma\}\{\{\}^\{\\prime\\prime\}\}\_\{\\bm\{z\}\}^\{\*\}\(t\)\\\|^\{2\}\\mathrm\{d\}t=โซqโก\(๐\)โ๐๐โโi=0Kโ1๐iโi\+1โ\[๐i,๐i\+1\]\\displaystyle=\\int q\(\\bm\{z\}\)\\mathrm\{d\}\\bm\{z\}\\sum\_\{i=0\}^\{K\-1\}\\mathcal\{C\}\_\{i\\rightarrow i\+1\}\[\\bm\{s\}\_\{i\},\\bm\{s\}\_\{i\+1\}\]=โซฮผ0โ\(๐0\)โโi=0Kโ1๐ฆiโi\+1โโ\[๐i,๐i\+1\]โ\(โj=0Kโ1๐jโj\+1โ\[๐j,๐j\+1\]\)โ๐๐\\displaystyle=\\int\\mu\_\{0\}\(\\bm\{s\}\_\{0\}\)\\prod\_\{i=0\}^\{K\-1\}\\mathcal\{K\}^\{\*\}\_\{i\\rightarrow i\+1\}\[\\bm\{s\}\_\{i\},\\bm\{s\}\_\{i\+1\}\]\\left\(\\sum\_\{j=0\}^\{K\-1\}\\mathcal\{C\}\_\{j\\rightarrow j\+1\}\[\\bm\{s\}\_\{j\},\\bm\{s\}\_\{j\+1\}\]\\right\)\\mathrm\{d\}\\bm\{z\}\(48\)
Here, using the definition of the propagator:
โซฮผ0โ\(๐0\)โโi=0Kโ1๐ฆiโi\+1โโ\[๐i,๐i\+1\]โ\(๐jโj\+1โ\[๐j,๐j\+1\]\)โ๐
๐=โซฯjโj\+1โโ\[๐j,๐j\+1\]โ๐jโj\+1โ\[๐j,๐j\+1\]โdโ๐jโdโ๐j\+1\\begin\{split\}&\\int\\mu\_\{0\}\(\\bm\{s\}\_\{0\}\)\\prod\_\{i=0\}^\{K\-1\}\\mathcal\{K\}^\{\*\}\_\{i\\rightarrow i\+1\}\[\\bm\{s\}\_\{i\},\\bm\{s\}\_\{i\+1\}\]\\left\(\\mathcal\{C\}\_\{j\\rightarrow j\+1\}\[\\bm\{s\}\_\{j\},\\bm\{s\}\_\{j\+1\}\]\\right\)\\mathrm\{d\}\\bm\{z\}\\\\ &\\qquad=\\int\\pi^\{\*\}\_\{j\\rightarrow j\+1\}\[\\bm\{s\}\_\{j\},\\bm\{s\}\_\{j\+1\}\]\\mathcal\{C\}\_\{j\\rightarrow j\+1\}\[\\bm\{s\}\_\{j\},\\bm\{s\}\_\{j\+1\}\]\\mathrm\{d\}\\bm\{s\}\_\{j\}\\mathrm\{d\}\\bm\{s\}\_\{j\+1\}\\end\{split\}\(49\)
Therefore:
โTF=โi=0Kโ1โซฯiโi\+1โโ\[๐i,๐i\+1\]โ๐iโi\+1โ\[๐i,๐i\+1\]โdโ๐iโdโ๐i\+1\\mathcal\{L\}\_\{\\text\{TF\}\}=\\sum\_\{i=0\}^\{K\-1\}\\int\\pi^\{\*\}\_\{i\\rightarrow i\+1\}\[\\bm\{s\}\_\{i\},\\bm\{s\}\_\{i\+1\}\]\\mathcal\{C\}\_\{i\\rightarrow i\+1\}\[\\bm\{s\}\_\{i\},\\bm\{s\}\_\{i\+1\}\]\\mathrm\{d\}\\bm\{s\}\_\{i\}\\mathrm\{d\}\\bm\{s\}\_\{i\+1\}\(50\)This coincides exactly with the lower bound of the DOAT problem cost derived above\. Thus, the path constructed by TracingFlow is exactly the optimal solution to the DOAT problem in the case whereฯ๐,ฯ๐โ0\\sigma\_\{\\bm\{x\}\},\\sigma\_\{\\bm\{v\}\}\\rightarrow 0\.
Next, we prove that under the construction of TracingFlow, the acceleration field๐โก\(๐,t\)\\bm\{a\}\(\\bm\{s\},t\)is single\-valued with respect to๐\\bm\{s\}andtt\. By[Section4\.1](https://arxiv.org/html/2608.21070#S4.SS1.SSS0.Px2), under the Optimal Transport Plan, the trajectories of flow maps starting from different points must not intersect\. This implies that for a given timett, any point๐\\bm\{s\}in the augmented space is traversed by at most one flow map trajectory\. Therefore, the acceleration field๐โก\(๐,t\)\\bm\{a\}\(\\bm\{s\},t\)possesses single\-valuedness\. In other words, the solution constructed by TracingFlow can indeed be expressed by a single optimal acceleration field๐โก\(๐,t\)\\bm\{a\}\(\\bm\{s\},t\)\.
### A\.7Proof of Proposition 4\.8
###### Proof\.
Letฮณ:\[0,T\]โโd\\gamma:\[0,T\]\\to\\mathbb\{R\}^\{d\}be a path that isC1C^\{1\}on\[0,T\]\[0,T\]and piecewiseC2C^\{2\}on each interval\[tjโ1,tj\]\[t\_\{j\-1\},t\_\{j\}\]\.
The optimal solution is given by a piecewise cubic polynomial:
ฮณโก\(tjโ1\+ฯ\)=\(๐jโ1\+๐j\)โฮโtjโ2โ\(๐jโ๐jโ1\)\(ฮโtj\)3โฯ3\+3โ\(๐jโ๐jโ1\)โ\(2โ๐jโ1\+๐j\)โฮโtj\(ฮโtj\)2โฯ2\+๐jโ1โฯ\+๐jโ1,\\begin\{split\}\\gamma\(t\_\{j\-1\}\+\\tau\)&=\\frac\{\(\\bm\{v\}\_\{j\-1\}\+\\bm\{v\}\_\{j\}\)\\Delta t\_\{j\}\-2\(\\bm\{x\}\_\{j\}\-\\bm\{x\}\_\{j\-1\}\)\}\{\(\\Delta t\_\{j\}\)^\{3\}\}\\tau^\{3\}\\\\ &\\quad\+\\frac\{3\(\\bm\{x\}\_\{j\}\-\\bm\{x\}\_\{j\-1\}\)\-\(2\\bm\{v\}\_\{j\-1\}\+\\bm\{v\}\_\{j\}\)\\Delta t\_\{j\}\}\{\(\\Delta t\_\{j\}\)^\{2\}\}\\tau^\{2\}\+\\bm\{v\}\_\{j\-1\}\\tau\+\\bm\{x\}\_\{j\-1\},\\end\{split\}\(51\)forฯโ\[0,ฮโtj\]\\tau\\in\[0,\\Delta t\_\{j\}\], whereฮโtj=tjโtjโ1\\Delta t\_\{j\}=t\_\{j\}\-t\_\{j\-1\},
According to Corollary[4\.1](https://arxiv.org/html/2608.21070#S4.SS1.SSS0.Px2), it suffices to minimize
๐ฅโก\(๐0,๐1,โฆ,๐K\)=\\displaystyle\\mathcal\{J\}\(\\bm\{v\}\_\{0\},\\bm\{v\}\_\{1\},\\dots,\\bm\{v\}\_\{K\}\)=2โj=1K1\(ฮโtj\)3\{\(โฅ๐jโ1โฅ2\+โจ๐jโ1,๐jโฉ\+โฅ๐jโฅ2\)\(ฮtj\)2\\displaystyle 2\\sum\_\{j=1\}^\{K\}\\frac\{1\}\{\(\\Delta t\_\{j\}\)^\{3\}\}\\bigg\\\{\(\\\|\\bm\{v\}\_\{j\-1\}\\\|^\{2\}\+\\langle\\bm\{v\}\_\{j\-1\},\\bm\{v\}\_\{j\}\\rangle\+\\\|\\bm\{v\}\_\{j\}\\\|^\{2\}\)\(\\Delta t\_\{j\}\)^\{2\}\+3โจ๐jโ1\+๐j,xjโ1โxjโฉฮtj\+3โฅxjโ1โxjโฅ2\}\\displaystyle\\qquad\\qquad\\qquad\+3\\langle\\bm\{v\}\_\{j\-1\}\+\\bm\{v\}\_\{j\},x\_\{j\-1\}\-x\_\{j\}\\rangle\\Delta t\_\{j\}\+3\\\|x\_\{j\-1\}\-x\_\{j\}\\\|^\{2\}\\bigg\\\}\(52\)๐ฅโฅ0\\mathcal\{J\}\\geq 0always holds, so the the formula above is a semi\-positive quadratic form\.
Take differentiation, we have
โ๐j๐ฅ\\displaystyle\\nabla\_\{\\bm\{v\}\_\{j\}\}\\mathcal\{J\}=2\(๐jโ1\+2โ๐jฮโtj\+2โ๐j\+๐j\+1ฮโtj\+1\)\+6\(๐jโ1โ๐j\(ฮโtj\)2\+๐jโ๐j\+1\(ฮโtj\+1\)2\),forj=1,โฆ,Kโ1,\\displaystyle=2\\left\(\\frac\{\\bm\{v\}\_\{j\-1\}\+2\\bm\{v\}\_\{j\}\}\{\\Delta t\_\{j\}\}\+\\frac\{2\\bm\{v\}\_\{j\}\+\\bm\{v\}\_\{j\+1\}\}\{\\Delta t\_\{j\+1\}\}\\right\)\+6\\left\(\\frac\{\\bm\{x\}\_\{j\-1\}\-\\bm\{x\}\_\{j\}\}\{\(\\Delta t\_\{j\}\)^\{2\}\}\+\\frac\{\\bm\{x\}\_\{j\}\-\\bm\{x\}\_\{j\+1\}\}\{\(\\Delta t\_\{j\+1\}\)^\{2\}\}\\right\),\\quad\\text\{for \}j=1,\\dots,K\-1,\(53\)โ๐0๐ฅ\\displaystyle\\nabla\_\{\\bm\{v\}\_\{0\}\}\\mathcal\{J\}=2โ\(2โ๐0\+๐1\)ฮโt1\+6โ\(๐0โ๐1\)\(ฮโt1\)2,โ๐K๐ฅ=2โ\(๐Kโ1\+2โ๐K\)ฮโtK\+6โ\(๐Kโ1โ๐K\)\(ฮโtK\)2\.\\displaystyle=\\frac\{2\(2\\bm\{v\}\_\{0\}\+\\bm\{v\}\_\{1\}\)\}\{\\Delta t\_\{1\}\}\+\\frac\{6\(\\bm\{x\}\_\{0\}\-\\bm\{x\}\_\{1\}\)\}\{\(\\Delta t\_\{1\}\)^\{2\}\},\\nabla\_\{\\bm\{v\}\_\{K\}\}\\mathcal\{J\}=\\frac\{2\(\\bm\{v\}\_\{K\-1\}\+2\\bm\{v\}\_\{K\}\)\}\{\\Delta t\_\{K\}\}\+\\frac\{6\(\\bm\{x\}\_\{K\-1\}\-\\bm\{x\}\_\{K\}\)\}\{\(\\Delta t\_\{K\}\)^\{2\}\}\.\(54\)
Set the gradients to zero and rearranging the terms to separate the velocities๐\\bm\{v\}and positions๐\\bm\{x\}, we obtain:
Forj=0j=0:
2ฮโt1โ๐0\+1ฮโt1โ๐1=3โ\(๐1โ๐0\)\(ฮโt1\)2\.\\frac\{2\}\{\\Delta t\_\{1\}\}\\bm\{v\}\_\{0\}\+\\frac\{1\}\{\\Delta t\_\{1\}\}\\bm\{v\}\_\{1\}=\\frac\{3\(\\bm\{x\}\_\{1\}\-\\bm\{x\}\_\{0\}\)\}\{\(\\Delta t\_\{1\}\)^\{2\}\}\.\(55\)Forj=1,โฆ,Kโ1j=1,\\dots,K\-1:
1ฮโtjโ๐jโ1\+2โ\(1ฮโtj\+1ฮโtj\+1\)โ๐j\+1ฮโtj\+1โ๐j\+1=3โ\(๐jโ๐jโ1\(ฮโtj\)2\+๐j\+1โ๐j\(ฮโtj\+1\)2\)\.\\frac\{1\}\{\\Delta t\_\{j\}\}\\bm\{v\}\_\{j\-1\}\+2\\left\(\\frac\{1\}\{\\Delta t\_\{j\}\}\+\\frac\{1\}\{\\Delta t\_\{j\+1\}\}\\right\)\\bm\{v\}\_\{j\}\+\\frac\{1\}\{\\Delta t\_\{j\+1\}\}\\bm\{v\}\_\{j\+1\}=3\\left\(\\frac\{\\bm\{x\}\_\{j\}\-\\bm\{x\}\_\{j\-1\}\}\{\(\\Delta t\_\{j\}\)^\{2\}\}\+\\frac\{\\bm\{x\}\_\{j\+1\}\-\\bm\{x\}\_\{j\}\}\{\(\\Delta t\_\{j\+1\}\)^\{2\}\}\\right\)\.\(56\)Forj=Kj=K:
1ฮโtKโ๐Kโ1\+2ฮโtKโ๐K=3โ\(๐Kโ๐Kโ1\)\(ฮโtK\)2\.\\frac\{1\}\{\\Delta t\_\{K\}\}\\bm\{v\}\_\{K\-1\}\+\\frac\{2\}\{\\Delta t\_\{K\}\}\\bm\{v\}\_\{K\}=\\frac\{3\(\\bm\{x\}\_\{K\}\-\\bm\{x\}\_\{K\-1\}\)\}\{\(\\Delta t\_\{K\}\)^\{2\}\}\.\(57\)
DenoteV=\[๐0,๐1,โฆ,๐K\]โโdร\(K\+1\)V=\[\\bm\{v\}\_\{0\},\\bm\{v\}\_\{1\},\\dots,\\bm\{v\}\_\{K\}\]\\in\\mathbb\{R\}^\{d\\times\(K\+1\)\}and writing this system in matrix formVโA=BVA=B, where
A=\[2ฮโt11ฮโt10โฏ001ฮโt12โ\(1ฮโt1\+1ฮโt2\)1ฮโt2โฏ0001ฮโt22โ\(1ฮโt2\+1ฮโt3\)โฑโฑโฑ1ฮโtKโ1000โฏ1ฮโtKโ12โ\(1ฮโtKโ1\+1ฮโtK\)1ฮโtK00โฏ01ฮโtK2ฮโtK\]โโ\(K\+1\)ร\(K\+1\)A=\\begin\{bmatrix\}\\frac\{2\}\{\\Delta t\_\{1\}\}&\\frac\{1\}\{\\Delta t\_\{1\}\}&0&\\cdots&0&0\\\\ \\frac\{1\}\{\\Delta t\_\{1\}\}&2\\left\(\\frac\{1\}\{\\Delta t\_\{1\}\}\+\\frac\{1\}\{\\Delta t\_\{2\}\}\\right\)&\\frac\{1\}\{\\Delta t\_\{2\}\}&\\cdots&0&0\\\\ 0&\\frac\{1\}\{\\Delta t\_\{2\}\}&2\\left\(\\frac\{1\}\{\\Delta t\_\{2\}\}\+\\frac\{1\}\{\\Delta t\_\{3\}\}\\right\)&\\ddots&\\vdots&\\vdots\\\\ \\vdots&\\vdots&\\ddots&\\ddots&\\frac\{1\}\{\\Delta t\_\{K\-1\}\}&0\\\\ 0&0&\\cdots&\\frac\{1\}\{\\Delta t\_\{K\-1\}\}&2\\left\(\\frac\{1\}\{\\Delta t\_\{K\-1\}\}\+\\frac\{1\}\{\\Delta t\_\{K\}\}\\right\)&\\frac\{1\}\{\\Delta t\_\{K\}\}\\\\ 0&0&\\cdots&0&\\frac\{1\}\{\\Delta t\_\{K\}\}&\\frac\{2\}\{\\Delta t\_\{K\}\}\\end\{bmatrix\}\\in\\mathbb\{R\}^\{\(K\+1\)\\times\(K\+1\)\}\(58\)
B=\[3โ\(๐1โ๐0\)\(ฮโt1\)2,3โ\(๐1โ๐0\(ฮโt1\)2\+๐2โ๐1\(ฮโt2\)2\),โฆ,OPEN3โ\(๐Kโ1โ๐Kโ2\(ฮโtKโ1\)2\+๐Kโ๐Kโ1\(ฮโtK\)2\),3โ\(๐Kโ๐Kโ1\)\(ฮโtK\)2\]โโdร\(K\+1\)\\begin\{split\}B=\\Bigg\[&\\frac\{3\(\\bm\{x\}\_\{1\}\-\\bm\{x\}\_\{0\}\)\}\{\(\\Delta t\_\{1\}\)^\{2\}\},3\\left\(\\frac\{\\bm\{x\}\_\{1\}\-\\bm\{x\}\_\{0\}\}\{\(\\Delta t\_\{1\}\)^\{2\}\}\+\\frac\{\\bm\{x\}\_\{2\}\-\\bm\{x\}\_\{1\}\}\{\(\\Delta t\_\{2\}\)^\{2\}\}\\right\),\\dots,\\\\ &3\\left\(\\frac\{\\bm\{x\}\_\{K\-1\}\-\\bm\{x\}\_\{K\-2\}\}\{\(\\Delta t\_\{K\-1\}\)^\{2\}\}\+\\frac\{\\bm\{x\}\_\{K\}\-\\bm\{x\}\_\{K\-1\}\}\{\(\\Delta t\_\{K\}\)^\{2\}\}\\right\),\\frac\{3\(\\bm\{x\}\_\{K\}\-\\bm\{x\}\_\{K\-1\}\)\}\{\(\\Delta t\_\{K\}\)^\{2\}\}\\Bigg\]\\in\\mathbb\{R\}^\{d\\times\(K\+1\)\}\\end\{split\}\(59\)The solution satisfiesฮณโฒโ\(tj\)=๐j\\gamma^\{\\prime\}\(t\_\{j\}\)=\\bm\{v\}\_\{j\}for allj=0,1,โฆโK\.j=0,1,\.\.\.K\.
SinceAAis strictly diagonally dominant, the solutionVVis unique\. โ
### A\.8Proof of Theorem A\.1
###### Theorem A\.1\.
Let๐๐โ\(๐ฑ,๐ฏ,t\)\\bm\{a\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},\\bm\{v\},t\)beLฮธL\_\{\\theta\}\-Lipschitz continuous of\(๐ฑ,๐ฏ\)\(\\bm\{x\},\\bm\{v\}\)\. The Wasserstein\-2 distance between the generated distributionฯ^ti\(pos\)โ\(๐ฑ\)=โซฯ^tiโ\(๐ฑ,๐ฏ\)โ๐๐ฏ\\hat\{\\rho\}\_\{t\_\{i\}\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\)=\\int\\hat\{\\rho\}\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\)\\mathrm\{d\}\\bm\{v\}and the true distributionฮผi\(pos\)โ\(๐ฑ\)\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\)in the position space is bounded by
๐ฒ22โ\(ฯ^ti\(pos\),ฮผi\(pos\)\)โค2โeฮบโก\(tiโt0\)โ\(โv0\+\(tiโt0\)โโAFM\|\[t0,ti\]\)\\mathcal\{W\}\_\{2\}^\{2\}\(\\hat\{\\rho\}\_\{t\_\{i\}\}^\{\(\\text\{pos\}\)\},\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\)\\leq 2e^\{\\kappa\(t\_\{i\}\-t\_\{0\}\)\}\\left\(\\mathcal\{L\}\_\{v\_\{0\}\}\+\(t\_\{i\}\-t\_\{0\}\)\\mathcal\{L\}\_\{\\text\{AFM\}\}\|\_\{\[t\_\{0\},t\_\{i\}\]\}\\right\)\(60\)where the cumulative AFM loss and the initial velocity loss are defined as:
โAFM\|\[t0,ti\]=๐ผt,\(๐,๐\)โผฯtโฅ๐๐ฝ\(๐,๐,t\)โ๐\(๐,๐,t\)โฅ2,โv0=๐ผ\(๐0,๐0\)โผฯt0โฅ๐0,๐\(๐0\)โ๐0โฅ2\.\\displaystyle\\mathcal\{L\}\_\{\\text\{AFM\}\}\|\_\{\[t\_\{0\},t\_\{i\}\]\}=\\mathbb\{E\}\_\{t,\(\\bm\{x\},\\bm\{v\}\)\\sim\\rho\_\{t\}\}\\\|\\bm\{a\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},\\bm\{v\},t\)\-\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\)\\\|^\{2\},\\mathcal\{L\}\_\{v\_\{0\}\}=\\mathbb\{E\}\_\{\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\)\\sim\\rho\_\{t\_\{0\}\}\}\\\|\\bm\{v\}\_\{0,\\bm\{\\xi\}\}\(\\bm\{x\}\_\{0\}\)\-\\bm\{v\}\_\{0\}\\\|^\{2\}\.
andฮบ\\kappais a constant only depending onLฮธL\_\{\\theta\}\.
###### Proof\.
We address the problem directly within the augmented space๐ณร๐ฑ\\mathcal\{X\}\\times\\mathcal\{V\}\. Let the system state be denoted by๐=\[๐,๐\]T\\bm\{s\}=\[\\bm\{x\},\\bm\{v\}\]^\{T\}, and letฮผiโ\(๐\)\\mu\_\{i\}\(\\bm\{s\}\)represent the data distribution in this space\. We define a flow mapฯt:\[t0,tK\]ร๐ณร๐ฑโ๐ณร๐ฑ\\bm\{\\phi\}\_\{t\}:\[t\_\{0\},t\_\{K\}\]\\times\\mathcal\{X\}\\times\\mathcal\{V\}\\rightarrow\\mathcal\{X\}\\times\\mathcal\{V\}, which is governed by the exact marginal acceleration field๐โก\(๐,t\)\\bm\{a\}\(\\bm\{s\},t\)defined in[Equation13](https://arxiv.org/html/2608.21070#S4.E13)\.
Using\(ฯt\)\#\(\\bm\{\\phi\}\_\{t\}\)\_\{\\\#\}to denote the push\-forward operator, the time\-dependent density evolves asฯt=\(ฯt\)\#โฯt0\\rho\_\{t\}=\(\\bm\{\\phi\}\_\{t\}\)\_\{\\\#\}\\rho\_\{t\_\{0\}\}, whereฯt0=ฮผ0\\rho\_\{t\_\{0\}\}=\\mu\_\{0\}\. Similarly, letฯt๐ฝ\\bm\{\\phi\}^\{\\bm\{\\theta\}\}\_\{t\}be the flow map induced by the learned acceleration field๐๐ฝโ\(๐,t\)\\bm\{a\}\_\{\\bm\{\\theta\}\}\(\\bm\{s\},t\)\. The generated distributionฯ^t\\hat\{\\rho\}\_\{t\}then evolves according toฯ^t=\(ฯt๐ฝ\)\#โฯ^t0\\hat\{\\rho\}\_\{t\}=\(\\bm\{\\phi\}^\{\\bm\{\\theta\}\}\_\{t\}\)\_\{\\\#\}\\hat\{\\rho\}\_\{t\_\{0\}\}\. It is important to note thatฯ^t0โ ฯt0\\hat\{\\rho\}\_\{t\_\{0\}\}\\neq\\rho\_\{t\_\{0\}\}, as the conditional distribution of the initial velocity is approximated via a neural network\.
We introduce a bridging measureฮฝt=\(ฯtฮธ\)\#โฯt0\\nu\_\{t\}=\(\\bm\{\\phi\}^\{\\theta\}\_\{t\}\)\_\{\\\#\}\\rho\_\{t\_\{0\}\}, then the Wasserstein\-2 distance between joint distributions can be decomposed as
๐ฒ22โ\(ฯt,ฯ^t\)โค\(๐ฒ2โ\(ฯt,ฮฝt\)\+๐ฒ2โ\(ฮฝt,ฯ^t\)\)2โค2โ\(๐ฒ22โ\(ฯt,ฮฝt\)\+๐ฒ22โ\(ฮฝt,ฯ^t\)\)\\mathcal\{W\}\_\{2\}^\{2\}\(\\rho\_\{t\},\\hat\{\\rho\}\_\{t\}\)\\leq\\left\(\{\\mathcal\{W\}\_\{2\}\(\\rho\_\{t\},\\nu\_\{t\}\)\}\+\{\\mathcal\{W\}\_\{2\}\(\\nu\_\{t\},\\hat\{\\rho\}\_\{t\}\)\}\\right\)^\{2\}\\leq 2\\left\(\\mathcal\{W\}\_\{2\}^\{2\}\(\\rho\_\{t\},\\nu\_\{t\}\)\+\\mathcal\{W\}\_\{2\}^\{2\}\(\\nu\_\{t\},\\hat\{\\rho\}\_\{t\}\)\\right\)\\\\\(61\)
Estimation of Term 1: DefineQt=โซฮผ0โ\(๐0\)โโฯtฮธโ\(๐0\)โฯtโ\(๐0\)โ2โdโ๐0,Q\_\{t\}=\\int\\mu\_\{0\}\(\\bm\{s\}\_\{0\}\)\\\|\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\\\|^\{2\}\\mathrm\{d\}\\bm\{s\}\_\{0\},hereฮผ0=ฯt0\\mu\_\{0\}=\\rho\_\{t\_\{0\}\}according to definition\. Noting thatQt0=0Q\_\{t\_\{0\}\}=0since the starting points are identical\. DefineA=\[0I00\]A=\\begin\{bmatrix\}0&I\\\\ 0&0\\end\{bmatrix\}\.
Taking the time derivative, we have:
dโQtdโt\\displaystyle\\frac\{\\mathrm\{d\}Q\_\{t\}\}\{\\mathrm\{d\}t\}=2โโซฮผ0โ\(๐0\)โ\(ฯtฮธโ\(๐0\)โฯtโ\(๐0\)\)Tโ\(dโฯtฮธโ\(๐0\)dโtโdโฯtโ\(๐0\)dโt\)โdโ๐0\\displaystyle=2\\int\\mu\_\{0\}\(\\bm\{s\}\_\{0\}\)\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\)^\{T\}\\left\(\\frac\{\\mathrm\{d\}\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\}\{\\mathrm\{d\}t\}\-\\frac\{\\mathrm\{d\}\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\}\{\\mathrm\{d\}t\}\\right\)\\mathrm\{d\}\\bm\{s\}\_\{0\}=2โโซฮผ0โ\(๐0\)โ\(ฯtฮธโ\(๐0\)โฯtโ\(๐0\)\)Tโ\(Aโฯtฮธโ\(๐0\)โAโฯtโ\(๐0\)\+\[๐๐ฮธโ\(ฯtฮธโ\(๐0\),t\)\]โ\[๐๐โก\(ฯtโ\(๐0\),t\)\]\)โdโ๐0\\displaystyle=2\\int\\mu\_\{0\}\(\\bm\{s\}\_\{0\}\)\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\)^\{T\}\\left\(A\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-A\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\+\\begin\{bmatrix\}\\bm\{0\}\\\\ \\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\),t\)\\end\{bmatrix\}\-\\begin\{bmatrix\}\\bm\{0\}\\\\ \\bm\{a\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\end\{bmatrix\}\\right\)\\mathrm\{d\}\\bm\{s\}\_\{0\}=2โซฮผ0\(๐0\)\[\(ฯtฮธ\(๐0\)โฯt\(๐0\)\)TA\(ฯtฮธ\(๐0\)โฯt\(๐0\)\)\\displaystyle=2\\int\\mu\_\{0\}\(\\bm\{s\}\_\{0\}\)\\Bigg\[\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\)^\{T\}A\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\)\+\(ฯtฮธโ\(๐0\)โฯtโ\(๐0\)\)Tโ\[๐๐ฮธโ\(ฯtฮธโ\(๐0\),t\)โ๐ฮธโ\(ฯtโ\(๐0\),t\)\]\\displaystyle\\quad\+\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\)^\{T\}\\begin\{bmatrix\}\\bm\{0\}\\\\ \\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\),t\)\-\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\end\{bmatrix\}\+\(ฯtฮธ\(๐0\)โฯt\(๐0\)\)T\[๐๐ฮธโ\(ฯtโ\(๐0\),t\)โ๐โก\(ฯtโ\(๐0\),t\)\]\]d๐0\\displaystyle\\quad\+\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\)^\{T\}\\begin\{bmatrix\}\\bm\{0\}\\\\ \\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\-\\bm\{a\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\end\{bmatrix\}\\Bigg\]\\mathrm\{d\}\\bm\{s\}\_\{0\}\(62\)
We bound the three terms inside the integral separately\. Letฮโฯtโ\(๐0\)=ฯtฮธโ\(๐0\)โฯtโ\(๐0\)\\Delta\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)=\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\.
1. 1\.SinceโAโโค1\\\|A\\\|\\leq 1, we haveฮโฯtโ\(๐0\)TโAโฮโฯtโ\(๐0\)โคโฮโฯtโ\(๐0\)โ2\\Delta\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)^\{T\}A\\Delta\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\\leq\\\|\\Delta\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\\\|^\{2\}\.
2. 2\.Using the Lipschitz continuity of the network๐ฮธ\\bm\{a\}\_\{\\theta\}with constantLฮธL\_\{\\theta\}: ฮโฯtโ\(๐0\)Tโ\[๐๐ฮธโ\(ฯtฮธโ\(๐0\),t\)โ๐ฮธโ\(ฯtโ\(๐0\),t\)\]\\displaystyle\\Delta\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)^\{T\}\\begin\{bmatrix\}\\bm\{0\}\\\\ \\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\),t\)\-\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\end\{bmatrix\}โคโฮโฯtโ\(๐0\)โโ
โ๐ฮธโ\(ฯtฮธโ\(๐0\),t\)โ๐ฮธโ\(ฯtโ\(๐0\),t\)โ\\displaystyle\\leq\\\|\\Delta\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\\\|\\cdot\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\),t\)\-\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\\|โคLฮธโโฮโฯtโ\(๐0\)โ2\.\\displaystyle\\leq L\_\{\\theta\}\\\|\\Delta\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\\\|^\{2\}\.\(63\)
3. 3\.Using Cauchy\-Schwarz and the inequality2โxโyโคx2\+y22xy\\leq x^\{2\}\+y^\{2\}: ฮโฯtโ\(๐0\)Tโ\[๐๐ฮธโ\(ฯtโ\(๐0\),t\)โ๐โก\(ฯtโ\(๐0\),t\)\]\\displaystyle\\Delta\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)^\{T\}\\begin\{bmatrix\}\\bm\{0\}\\\\ \\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\-\\bm\{a\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\end\{bmatrix\}โคโฮโฯtโ\(๐0\)โโ
โ๐ฮธโ\(ฯtโ\(๐0\),t\)โ๐โก\(ฯtโ\(๐0\),t\)โ\\displaystyle\\leq\\\|\\Delta\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\\\|\\cdot\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\-\\bm\{a\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\\|โค12โโฮโฯtโ\(๐0\)โ2\+12โโ๐ฮธโ\(ฯtโ\(๐0\),t\)โ๐โก\(ฯtโ\(๐0\),t\)โ2\.\\displaystyle\\leq\\frac\{1\}\{2\}\\\|\\Delta\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\\\|^\{2\}\+\\frac\{1\}\{2\}\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\-\\bm\{a\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\\|^\{2\}\.\(64\)
Substituting these bounds back into the derivative:
dโQtdโt\\displaystyle\\frac\{\\mathrm\{d\}Q\_\{t\}\}\{\\mathrm\{d\}t\}โค2โโซฮผ0โ\(๐0\)โ\[\(1\+Lฮธ\+12\)โโฮโฯtโ\(๐0\)โ2\+12โโ๐ฮธโ\(ฯtโ\(๐0\),t\)โ๐โก\(ฯtโ\(๐0\),t\)โ2\]โdโ๐0\\displaystyle\\leq 2\\int\\mu\_\{0\}\(\\bm\{s\}\_\{0\}\)\\left\[\(1\+L\_\{\\theta\}\+\\frac\{1\}\{2\}\)\\\|\\Delta\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\\\|^\{2\}\+\\frac\{1\}\{2\}\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\-\\bm\{a\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\\|^\{2\}\\right\]\\mathrm\{d\}\\bm\{s\}\_\{0\}=\(2โLฮธ\+3\)โQt\+โซฮผ0โ\(๐0\)โโ๐ฮธโ\(ฯtโ\(๐0\),t\)โ๐โก\(ฯtโ\(๐0\),t\)โ2โdโ๐0\\displaystyle=\(2L\_\{\\theta\}\+3\)Q\_\{t\}\+\\int\\mu\_\{0\}\(\\bm\{s\}\_\{0\}\)\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\-\\bm\{a\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\\|^\{2\}\\mathrm\{d\}\\bm\{s\}\_\{0\}\(65\)
By Gronwallโs inequality, sinceQt0=0Q\_\{t\_\{0\}\}=0:
Qtโคe\(2โLฮธ\+3\)โ\(tโt0\)โโซt0t\(โซฮผ0โ\(๐0\)โโ๐ฮธโ\(ฯฯโ\(๐0\),ฯ\)โ๐โก\(ฯฯโ\(๐0\),ฯ\)โ2โdโ๐0\)โ๐ฯQ\_\{t\}\\leq e^\{\(2L\_\{\\theta\}\+3\)\(t\-t\_\{0\}\)\}\\int\_\{t\_\{0\}\}^\{t\}\\left\(\\int\\mu\_\{0\}\(\\bm\{s\}\_\{0\}\)\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{\\tau\}\(\\bm\{s\}\_\{0\}\),\\tau\)\-\\bm\{a\}\(\\bm\{\\phi\}\_\{\\tau\}\(\\bm\{s\}\_\{0\}\),\\tau\)\\\|^\{2\}\\mathrm\{d\}\\bm\{s\}\_\{0\}\\right\)\\mathrm\{d\}\\tau\(66\)By the definition of Wasserstein\-2 distance, we have:
t0โ๐ฒ22โ\(ฯt,ฮฝt\)\\displaystyle t\_\{0\}\\mathcal\{W\}\_\{2\}^\{2\}\(\\rho\_\{t\},\\nu\_\{t\}\)=minโกโซฯโกโ๐1โ๐2โ2โ๐ฯs\.t\.โโซฯโdโ๐1=ฯt,โซฯโdโ๐2=ฮฝt\\displaystyle=\\min\_\{\\pi\}\\int\\\|\\bm\{s\}\_\{1\}\-\\bm\{s\}\_\{2\}\\\|^\{2\}\\mathrm\{d\}\\pi\\quad\\text\{s\.t\.\}\\int\\pi\\mathrm\{d\}\\bm\{s\}\_\{1\}=\\rho\_\{t\},\\int\\pi\\mathrm\{d\}\\bm\{s\}\_\{2\}=\\nu\_\{t\}โคโซฮผ0โ\(๐0\)โโฯt๐ฝโ\(๐0\)โฯtโ\(๐0\)โ2โdโ๐0\\displaystyle\\leq\\int\\mu\_\{0\}\(\\bm\{s\}\_\{0\}\)\\\|\\bm\{\\phi\}^\{\\bm\{\\theta\}\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\\\|^\{2\}\\mathrm\{d\}\\bm\{s\}\_\{0\}=Qtโคe\(2โLฮธ\+3\)โ\(tโt0\)โโซt0t๐ผ๐โผฮผฯโโ๐ฮธโ\(๐,ฯ\)โ๐โก\(๐,ฯ\)โ2โ๐ฯ\\displaystyle=Q\_\{t\}\\leq e^\{\(2L\_\{\\theta\}\+3\)\(t\-t\_\{0\}\)\}\\int\_\{t\_\{0\}\}^\{t\}\\mathbb\{E\}\_\{\\bm\{s\}\\sim\\mu\_\{\\tau\}\}\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{s\},\\tau\)\-\\bm\{a\}\(\\bm\{s\},\\tau\)\\\|^\{2\}\\mathrm\{d\}\\tau\(67\)=e\(2โLฮธ\+3\)โ\(tโt0\)โ\(tโt0\)โโAFM\|\[t0,t\]\.\\displaystyle=e^\{\(2L\_\{\\theta\}\+3\)\(t\-t\_\{0\}\)\}\(t\-t\_\{0\}\)\\mathcal\{L\}\_\{\\text\{AFM\}\}\|\_\{\[t\_\{0\},t\]\}\.\(68\)
Estimation of Term 2:Takeฯ0โ\(๐0,๐0ฮพ\)\\pi\_\{0\}\(\\bm\{s\}\_\{0\},\\bm\{s\}\_\{0\}^\{\\xi\}\)as the optimal coupling that attains๐ฒ22โ\(ฮฝt0,ฯ^t0\)=โซโ๐0โ๐0ฮพโ2โdโฯ0โ\(๐๐,๐0ฮพ\)\\mathcal\{W\}\_\{2\}^\{2\}\(\\nu\_\{t\_\{0\}\},\\hat\{\\rho\}\_\{t\_\{0\}\}\)=\\int\\\|\\bm\{s\}\_\{0\}\-\\bm\{s\}\_\{0\}^\{\\xi\}\\\|^\{2\}\\mathrm\{d\}\\pi\_\{0\}\(\\bm\{s\_\{0\}\},\\bm\{s\}\_\{0\}^\{\\xi\}\)\. Note thatฮฝt0=ฮผ0\\nu\_\{t\_\{0\}\}=\\mu\_\{0\}\(the true data distribution\)\.
From a pair of starting points\(๐0,๐0ฮพ\)โผฯ0\(\\bm\{s\}\_\{0\},\\bm\{s\}\_\{0\}^\{\\xi\}\)\\sim\\pi\_\{0\}, define the square distance between trajectories under the same approximate flow asฮโก\(t\)=โฯtฮธโ\(๐0\)โฯtฮธโ\(๐0ฮพ\)โ2\\Delta\(t\)=\\\|\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}^\{\\xi\}\)\\\|^\{2\}\. Taking the time derivative, we have:
dโฮโ\(t\)dโt\\displaystyle\\frac\{\\mathrm\{d\}\\Delta\(t\)\}\{\\mathrm\{d\}t\}=2โ\(ฯtฮธโ\(๐0\)โฯtฮธโ\(๐0ฮพ\)\)Tโ\(dโฯtฮธโ\(๐0\)dโtโdโฯtฮธโ\(๐0ฮพ\)dโt\)\\displaystyle=2\{\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}^\{\\xi\}\)\)^\{T\}\}\\left\(\\frac\{\\mathrm\{d\}\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\}\{\\mathrm\{d\}t\}\-\\frac\{\\mathrm\{d\}\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}^\{\\xi\}\)\}\{\\mathrm\{d\}t\}\\right\)=2โ\(ฯtฮธโ\(๐0\)โฯtฮธโ\(๐0ฮพ\)\)Tโ\(Aโก\(ฯtฮธโ\(๐0\)โฯtฮธโ\(๐0ฮพ\)\)\+\[๐๐ฮธโ\(ฯtฮธโ\(๐0\),t\)โ๐ฮธโ\(ฯtฮธโ\(๐0ฮพ\),t\)\]\)\\displaystyle=2\{\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}^\{\\xi\}\)\)^\{T\}\}\\left\(A\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}^\{\\xi\}\)\)\+\\begin\{bmatrix\}\\bm\{0\}\\\\ \\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\),t\)\-\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}^\{\\xi\}\),t\)\\end\{bmatrix\}\\right\)โค2โ\(โAโโ
โฯtฮธโ\(๐0\)โฯtฮธโ\(๐0ฮพ\)โ2\+โฯtฮธโ\(๐0\)โฯtฮธโ\(๐0ฮพ\)โโ
โ๐ฮธโ\(ฯtฮธโ\(๐0\),t\)โ๐ฮธโ\(ฯtฮธโ\(๐0ฮพ\),t\)โ\)\\displaystyle\\leq 2\\left\(\\\|A\\\|\\cdot\\\|\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}^\{\\xi\}\)\\\|^\{2\}\+\\\|\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}^\{\\xi\}\)\\\|\\cdot\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\),t\)\-\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}^\{\\xi\}\),t\)\\\|\\right\)โค2โ\(1\+Lฮธ\)โฮโ\(t\)\.\\displaystyle\\leq 2\(1\+L\_\{\\theta\}\)\\Delta\(t\)\.\(69\)Again by Gronwallโs inequality, we have:
ฮโก\(t\)โคe2โ\(1\+Lฮธ\)โ\(tโt0\)โฮโ\(t0\)=e2โ\(1\+Lฮธ\)โ\(tโt0\)โโ๐0โ๐0ฮพโ2\.\\Delta\(t\)\\leq e^\{2\(1\+L\_\{\\theta\}\)\(t\-t\_\{0\}\)\}\\Delta\(t\_\{0\}\)=e^\{2\(1\+L\_\{\\theta\}\)\(t\-t\_\{0\}\)\}\\\|\\bm\{s\}\_\{0\}\-\\bm\{s\}\_\{0\}^\{\\xi\}\\\|^\{2\}\.\(70\)Sinceฮฝt=\(ฯtฮธ\)\#โฮฝt0\\nu\_\{t\}=\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\)\_\{\\\#\}\\nu\_\{t\_\{0\}\}andฯ^t0=\(ฯtฮธ\)\#โฯ^t0\\hat\{\\rho\}\_\{t\_\{0\}\}=\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\)\_\{\\\#\}\\hat\{\\rho\}\_\{t\_\{0\}\}, the push\-forward of the optimal couplingฯt=\(ฯtฮธ,ฯtฮธ\)\#โฯ0\\pi\_\{t\}=\(\\bm\{\\phi\}\_\{t\}^\{\\theta\},\\bm\{\\phi\}\_\{t\}^\{\\theta\}\)\_\{\\\#\}\\pi\_\{0\}is a valid coupling for\(ฮฝt,ฯ^t\)\(\\nu\_\{t\},\\hat\{\\rho\}\_\{t\}\)\. Integrate the squared distance over all starting pairs with respect toฯ0\\pi\_\{0\}, we obtain:
๐ฒ22โ\(ฮฝt,ฯ^t\)\\displaystyle\\mathcal\{W\}\_\{2\}^\{2\}\(\\nu\_\{t\},\\hat\{\\rho\}\_\{t\}\)โคโซโฯtฮธโ\(๐0\)โฯtฮธโ\(๐0ฮพ\)โ2โdโฯ0โ\(๐0,๐0ฮพ\)\\displaystyle\\leq\\int\\\|\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}^\{\\xi\}\)\\\|^\{2\}\\mathrm\{d\}\\pi\_\{0\}\(\\bm\{s\}\_\{0\},\\bm\{s\}\_\{0\}^\{\\xi\}\)โคโซe2โ\(1\+Lฮธ\)โ\(tโt0\)โโ๐0โ๐0ฮพโ2โdโฯ0โ\(๐0,๐0ฮพ\)\\displaystyle\\leq\\int e^\{2\(1\+L\_\{\\theta\}\)\(t\-t\_\{0\}\)\}\\\|\\bm\{s\}\_\{0\}\-\\bm\{s\}\_\{0\}^\{\\xi\}\\\|^\{2\}\\mathrm\{d\}\\pi\_\{0\}\(\\bm\{s\}\_\{0\},\\bm\{s\}\_\{0\}^\{\\xi\}\)=e2โ\(1\+Lฮธ\)โ\(tโt0\)โ๐ฒ22โ\(ฮฝt0,ฯ^t0\)\\displaystyle=e^\{2\(1\+L\_\{\\theta\}\)\(t\-t\_\{0\}\)\}\\mathcal\{W\}\_\{2\}^\{2\}\(\\nu\_\{t\_\{0\}\},\\hat\{\\rho\}\_\{t\_\{0\}\}\)\(71\)Letฯ~โ\(๐๐,๐0ฮพ\)\\tilde\{\\pi\}\(\\bm\{s\_\{0\}\},\\bm\{s\}\_\{0\}^\{\\xi\}\)be the coupling betweenฮฝ0\\nu\_\{0\}andฯ^0\\hat\{\\rho\}\_\{0\}induced by the identity mapping on the position space, i\.e\., pairing๐0=\(๐0,๐0\)\\bm\{s\}\_\{0\}=\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\)with๐0ฮพ=\(๐0,๐0ฮพ\)\\bm\{s\}\_\{0\}^\{\\xi\}=\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}^\{\\xi\}\)\. We can bound Wasserstein distance byL2L\_\{2\}loss:
๐ฒ22โ\(ฮฝt0,ฯ^t0\)\\displaystyle\\mathcal\{W\}\_\{2\}^\{2\}\(\\nu\_\{t\_\{0\}\},\\hat\{\\rho\}\_\{t\_\{0\}\}\)โคโซโ๐0โ๐0ฮพโ2โ๐ฯ~โ\(๐0,๐0ฮพ\)\\displaystyle\\leq\\int\\\|\\bm\{s\}\_\{0\}\-\\bm\{s\}\_\{0\}^\{\\xi\}\\\|^\{2\}\\mathrm\{d\}\\tilde\{\\pi\}\(\\bm\{s\}\_\{0\},\\bm\{s\}\_\{0\}^\{\\xi\}\)=โซ\(โ๐0โ๐0ฮพโ2\+โ๐0โ๐0ฮพโ2\)โ๐ฯ~โ\(๐0,๐0,๐0ฮพ,๐0ฮพ\)\\displaystyle=\\int\\left\(\\\|\\bm\{x\}\_\{0\}\-\\bm\{x\}\_\{0\}^\{\\xi\}\\\|^\{2\}\+\\\|\\bm\{v\}\_\{0\}\-\\bm\{v\}\_\{0\}^\{\\xi\}\\\|^\{2\}\\right\)\\mathrm\{d\}\\tilde\{\\pi\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\},\\bm\{x\}\_\{0\}^\{\\xi\},\\bm\{v\}\_\{0\}^\{\\xi\}\)=โซโ๐0โ๐0ฮพโ2โ๐ฯ~โ\(๐0,๐0,๐0ฮพ,๐0ฮพ\)=โv0\\displaystyle=\\int\\\|\\bm\{v\}\_\{0\}\-\\bm\{v\}^\{\\xi\}\_\{0\}\\\|^\{2\}\\mathrm\{d\}\\tilde\{\\pi\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\},\{\\bm\{x\}\}\_\{0\}^\{\\xi\},\{\\bm\{v\}\}\_\{0\}^\{\\xi\}\)=\\mathcal\{L\}\_\{v\_\{0\}\}\(72\)
Thus๐ฒ22โ\(ฮฝt,ฯ^t\)โคe2โ\(1\+Lฮธ\)โ\(tโt0\)โโv0\.\\mathcal\{W\}\_\{2\}^\{2\}\(\\nu\_\{t\},\\hat\{\\rho\}\_\{t\}\)\\leq e^\{2\(1\+L\_\{\\theta\}\)\(t\-t\_\{0\}\)\}\\mathcal\{L\}\_\{v\_\{0\}\}\.
Combining the bounds above, we have
๐ฒ22โ\(ฯ^ti,ฯti\)\\displaystyle\\mathcal\{W\}\_\{2\}^\{2\}\(\\hat\{\\rho\}\_\{t\_\{i\}\},\\rho\_\{t\_\{i\}\}\)โค2โ\[๐ฒ22โ\(ฯ^ti,ฮฝti\)\+๐ฒ22โ\(ฮฝti,ฯti\)\]\\displaystyle\\leq 2\\left\[\\mathcal\{W\}\_\{2\}^\{2\}\(\\hat\{\\rho\}\_\{t\_\{i\}\},\\nu\_\{t\_\{i\}\}\)\+\\mathcal\{W\}\_\{2\}^\{2\}\(\\nu\_\{t\_\{i\}\},\\rho\_\{t\_\{i\}\}\)\\right\]โค2โ\[e\(2โLฮธ\+3\)โ\(tiโt0\)โ
\(tiโt0\)โโAFM\|\[t0,ti\]\+e\(2โLฮธ\+2\)โ\(tiโt0\)โโv0\]\.\\displaystyle\\leq 2\\left\[e^\{\(2L\_\{\\theta\}\+3\)\(t\_\{i\}\-t\_\{0\}\)\}\\cdot\(t\_\{i\}\-t\_\{0\}\)\\mathcal\{L\}\_\{\\text\{AFM\}\}\|\_\{\[t\_\{0\},t\_\{i\}\]\}\+e^\{\(2L\_\{\\theta\}\+2\)\(t\_\{i\}\-t\_\{0\}\)\}\\mathcal\{L\}\_\{v\_\{0\}\}\\right\]\.\(73\)
Defineฮบ=\(2โLฮธ\+3\)\\kappa=\{\(2L\_\{\\theta\}\+3\)\},then the final bound is obtained by:
๐ฒ22โ\(ฯ^ti\(pos\),ฯti\(pos\)=ฮผi\(pos\)\)\\displaystyle\\mathcal\{W\}\_\{2\}^\{2\}\(\\hat\{\\rho\}^\{\\text\{\(pos\)\}\}\_\{t\_\{i\}\},\\rho^\{\(\\text\{pos\}\)\}\_\{t\_\{i\}\}=\\mu\_\{i\}^\{\\text\{\(pos\)\}\}\)=minโกโซฯ๐โกโ๐1โ๐2โ2โdโฯ๐โ\(๐1,๐2\)s\.t\.โโซฯ๐โdโ๐2=ฯti\(pos\),โซฯ๐โdโ๐1=ฯ^ti\(pos\)\\displaystyle=\\min\_\{\\pi\_\{\\bm\{x\}\}\}\\int\\\|\\bm\{x\}\_\{1\}\-\\bm\{x\}\_\{2\}\\\|^\{2\}\\mathrm\{d\}\\pi\_\{\\bm\{x\}\}\(\\bm\{x\}\_\{1\},\\bm\{x\}\_\{2\}\)\\quad\\text\{s\.t\.\}\\int\\pi\_\{\\bm\{x\}\}\\mathrm\{d\}\\bm\{x\}\_\{2\}=\\rho^\{\(\\text\{pos\}\)\}\_\{t\_\{i\}\},\\quad\\int\\pi\_\{\\bm\{x\}\}\\mathrm\{d\}\\bm\{x\}\_\{1\}=\\hat\{\\rho\}\_\{t\_\{i\}\}^\{\\text\{\(pos\)\}\}โคminโกโซฯโกโ๐1โ๐2โ2โ๐ฯโ\(๐1,๐2\)s\.t\.โโซฯโdโ๐2=ฯti,โซฯโdโ๐1=ฯ^ti\\displaystyle\\leq\\min\_\{\\pi\}\\int\\\|\\bm\{x\}\_\{1\}\-\\bm\{x\}\_\{2\}\\\|^\{2\}\\mathrm\{d\}\\pi\(\\bm\{s\}\_\{1\},\\bm\{s\}\_\{2\}\)\\quad\\text\{s\.t\.\}\\int\\pi\\mathrm\{d\}\\bm\{s\}\_\{2\}=\\rho\_\{t\_\{i\}\},\\quad\\int\\pi\\mathrm\{d\}\\bm\{s\}\_\{1\}=\\hat\{\\rho\}\_\{t\_\{i\}\}โคminโกโซฯโกโ๐1โ๐2โ2โ๐ฯโ\(๐1,๐2\)\\displaystyle\\leq\\min\_\{\\pi\}\\int\\\|\\bm\{s\}\_\{1\}\-\\bm\{s\}\_\{2\}\\\|^\{2\}\\mathrm\{d\}\\pi\(\\bm\{s\}\_\{1\},\\bm\{s\}\_\{2\}\)\(74\)=๐ฒ22โ\(ฯ^ti,ฯti\)โค2โeฮบโก\(tiโt0\)โ\(\(tiโt0\)โโAFM\|\[t0,ti\]\+โv0\)\.\\displaystyle=\\mathcal\{W\}\_\{2\}^\{2\}\(\\hat\{\\rho\}\_\{t\_\{i\}\},\\rho\_\{t\_\{i\}\}\)\\leq 2e^\{\\kappa\(t\_\{i\}\-t\_\{0\}\)\}\\left\(\(t\_\{i\}\-t\_\{0\}\)\\mathcal\{L\}\_\{\\text\{AFM\}\}\|\_\{\[t\_\{0\},t\_\{i\}\]\}\+\\mathcal\{L\}\_\{v\_\{0\}\}\\right\)\.\(75\)
This completes the proof\. Note that our derivation holds even for multi\-valued or stochastic cases, since the upper bound on the Wasserstein distance is established via a valid coupling, which remains well\-defined regardless of whether the mapping is deterministic\.
โ
## Appendix BDatasets and Evaluation Metric
### B\.1Experiment Setup
The experiments were conducted on a shared high\-performance computing \(HPC\) cluster\. All experimental runs utilized GPUs equipped with 80GB of memory\. Both the network๐ฮธโ\(๐,๐,t\)\\bm\{a\}\_\{\\theta\}\(\\bm\{x\},\\bm\{v\},t\), employed to approximate acceleration, and the network๐ฯโ\(๐,๐,t\)\\bm\{v\}\_\{\\phi\}\(\\bm\{x\},\\bm\{v\},t\), employed to approximate velocity, utilize residual architectures\. Specifically,๐ฮธโ\(๐,๐,t\)\\bm\{a\}\_\{\\theta\}\(\\bm\{x\},\\bm\{v\},t\)consists of 4 hidden layers, while๐ฯโ\(๐,๐,t\)\\bm\{v\}\_\{\\phi\}\(\\bm\{x\},\\bm\{v\},t\)comprises 2 hidden layers\. The Optimal Transport \(OT\) calculations were implemented using the Python Optimal Transport \(POT\) library\[[16](https://arxiv.org/html/2608.21070#bib.bib16)\], and the training process of neural networks were implemented using Pytorch\[[43](https://arxiv.org/html/2608.21070#bib.bib43)\]\.
In all experiments, TracingFlow employs a completely consistent set of hyperparameters\. When solving for OT using the Sinkhorn algorithm, the regularization coefficient is set toฯต=1ร10โ3\\epsilon=1\\times 10^\{\-3\}\. When incorporating biological priors, the penalty coefficient for barcode mismatches is set top0=25p\_\{0\}=25\. Since the number of cells at each time point in the datasets processed in this paper is on the order of10310^\{3\}, we do not utilize mini\-batch OT; instead, we compute the Transport Plan on the full dataset\.
### B\.2Evaluation Metrics
We employ the 1\-Wasserstein Distance \(๐ฒ1\\mathcal\{W\}\_\{1\}\) and the 2\-Wasserstein Distance \(๐ฒ2\\mathcal\{W\}\_\{2\}\) to quantify the discrepancy between the predicted distribution and the ground truth distribution\. These metrics are defined as follows:
๐ฒ1โ\(ฮผ,ฮฝ\)=minโกโซฯโฮ โก\(ฮผ,ฮฝ\)โกโ๐โ๐โโ๐ฯโ\(๐,๐\)\\mathcal\{W\}\_\{1\}\(\\mu,\\nu\)=\\min\_\{\\pi\\in\\Pi\(\\mu,\\nu\)\}\\int\\\|\\bm\{x\}\-\\bm\{y\}\\\|\\mathrm\{d\}\\pi\(\\bm\{x\},\\bm\{y\}\)\(76\)๐ฒ2โ\(ฮผ,ฮฝ\)=\(minโกโซฯโฮ โก\(ฮผ,ฮฝ\)โกโ๐โ๐โ22โ๐ฯโ\(๐,๐\)\)1/2\\mathcal\{W\}\_\{2\}\(\\mu,\\nu\)=\\left\(\\min\_\{\\pi\\in\\Pi\(\\mu,\\nu\)\}\\int\\\|\\bm\{x\}\-\\bm\{y\}\\\|\_\{2\}^\{2\}\\mathrm\{d\}\\pi\(\\bm\{x\},\\bm\{y\}\)\\right\)^\{1/2\}\(77\)For the evaluation of datasets containing lineage information, we utilize the lineage\-weighted๐ฒ1\\mathcal\{W\}\_\{1\}and๐ฒ2\\mathcal\{W\}\_\{2\}distances\. These are defined based on the decomposition of the distributions\. Consider the distribution generated by the model, denoted asฮผโก\(๐\)\\mu\(\\bm\{x\}\), and the reference distribution provided by the dataset, denoted asฮฝโก\(๐\)\\nu\(\\bm\{x\}\)\. Both consist ofLLlineages and can be decomposed as:
ฮผโก\(๐\)=โi=1Lwiโฮผ\(i\)โ\(๐\),ฮฝโก\(๐\)=โi=1Lviโฮฝ\(i\)โ\(๐\)\\mu\(\\bm\{x\}\)=\\sum\_\{i=1\}^\{L\}w\_\{i\}\\mu^\{\(i\)\}\(\\bm\{x\}\),\\quad\\nu\(\\bm\{x\}\)=\\sum\_\{i=1\}^\{L\}v\_\{i\}\\nu^\{\(i\)\}\(\\bm\{x\}\)\(78\)whereฮผ\(i\)โ\(๐\)\\mu^\{\(i\)\}\(\\bm\{x\}\)andฮฝ\(i\)โ\(๐\)\\nu^\{\(i\)\}\(\\bm\{x\}\)represent the distributions of theii\-th lineage withinฮผโก\(๐\)\\mu\(\\bm\{x\}\)andฮฝโก\(๐\)\\nu\(\\bm\{x\}\), respectively, satisfyingโซฮผ\(i\)โ\(๐\)โ๐๐=1\\int\\mu^\{\(i\)\}\(\\bm\{x\}\)\\mathrm\{d\}\\bm\{x\}=1andโซฮฝ\(i\)โ\(๐\)โ๐๐=1\\int\\nu^\{\(i\)\}\(\\bm\{x\}\)\\mathrm\{d\}\\bm\{x\}=1\. The termswiw\_\{i\}andviv\_\{i\}denote the weights of theii\-th lineage inฮผโก\(๐\)\\mu\(\\bm\{x\}\)andฮฝโก\(๐\)\\nu\(\\bm\{x\}\), respectively, subject to the constraintโi=1Lwi=โi=1Lvi=1\\sum\_\{i=1\}^\{L\}w\_\{i\}=\\sum\_\{i=1\}^\{L\}v\_\{i\}=1\. The lineage\-weighted๐ฒ1\\mathcal\{W\}\_\{1\}and๐ฒ2\\mathcal\{W\}\_\{2\}distances are defined as:
๐ฒ1\(L\)\(ฮผโฅฮฝ\)=โi=1Lvi๐ฒ1\(ฮผ\(i\),ฮฝ\(i\)\),๐ฒ2\(L\)\(ฮผโฅฮฝ\)=โi=1Lvi๐ฒ2\(ฮผ\(i\),ฮฝ\(i\)\)\\mathcal\{W\}\_\{1\}^\{\(L\)\}\(\\mu\\\|\\nu\)=\\sum\_\{i=1\}^\{L\}v\_\{i\}\\mathcal\{W\}\_\{1\}\(\\mu^\{\(i\)\},\\nu^\{\(i\)\}\),\\quad\\mathcal\{W\}\_\{2\}^\{\(L\)\}\(\\mu\\\|\\nu\)=\\sum\_\{i=1\}^\{L\}v\_\{i\}\\mathcal\{W\}\_\{2\}\(\\mu^\{\(i\)\},\\nu^\{\(i\)\}\)\(79\)In essence, we first calculate the๐ฒ1\\mathcal\{W\}\_\{1\}and๐ฒ2\\mathcal\{W\}\_\{2\}distances between the distributions of corresponding lineages in the two datasets\. Subsequently, we compute a weighted average of these distances using the proportions of each lineage given by the reference distributionฮฝโก\(๐\)\\nu\(\\bm\{x\}\)\. It is important to note that the lineage\-weighted๐ฒ1\\mathcal\{W\}\_\{1\}and๐ฒ2\\mathcal\{W\}\_\{2\}metrics serve solely as indicators of the modelโs ability to preserve lineage prior information while learning the dynamics\. They are not true distance metrics in the mathematical sense, as they are evidently asymmetric with respect toฮผ\\muandฮฝ\\nu\.
### B\.3Datasets
3D Simulation Lineage Data\.Consider the following dynamics with three lineage barcodes and a twist structure:
\{dโXt\(0\)=vโdโt\+ฯxโdโBtXdโYt\(0\)=ฯyโdโBtYdโZt\(0\)=ฯzโdโBtZ\{dโXt\(1\)=vโdโt\+ฯxโdโBtXdโYt\(1\)=โฮผโZt\(1\)โdโt\+ฯyโdโBtYdโZt\(1\)=ฮผโYt\(1\)โdโt\+ฯzโdโBtZ\{dโXt\(2\)=vโdโt\+ฯxโdโBtXdโYt\(2\)=ฮผโZt\(2\)โdโt\+ฯyโdโBtYdโZt\(2\)=โฮผโYt\(2\)โdโt\+ฯzโdโBtZ\\left\\\{\\begin\{aligned\} dX\_\{t\}^\{\(0\)\}&=v\\,dt\+\\sigma\_\{x\}dB\_\{t\}^\{X\}\\\\ dY\_\{t\}^\{\(0\)\}&=\\sigma\_\{y\}dB\_\{t\}^\{Y\}\\\\ dZ\_\{t\}^\{\(0\)\}&=\\sigma\_\{z\}dB\_\{t\}^\{Z\}\\end\{aligned\}\\right\.\\quad\\left\\\{\\begin\{aligned\} dX\_\{t\}^\{\(1\)\}&=v\\,dt\+\\sigma\_\{x\}dB\_\{t\}^\{X\}\\\\ dY\_\{t\}^\{\(1\)\}&=\-\\mu Z\_\{t\}^\{\(1\)\}\\,dt\+\\sigma\_\{y\}dB\_\{t\}^\{Y\}\\\\ dZ\_\{t\}^\{\(1\)\}&=\\mu Y\_\{t\}^\{\(1\)\}\\,dt\+\\sigma\_\{z\}dB\_\{t\}^\{Z\}\\end\{aligned\}\\right\.\\quad\\left\\\{\\begin\{aligned\} dX\_\{t\}^\{\(2\)\}&=v\\,dt\+\\sigma\_\{x\}dB\_\{t\}^\{X\}\\\\ dY\_\{t\}^\{\(2\)\}&=\\mu Z\_\{t\}^\{\(2\)\}\\,dt\+\\sigma\_\{y\}dB\_\{t\}^\{Y\}\\\\ dZ\_\{t\}^\{\(2\)\}&=\-\\mu Y\_\{t\}^\{\(2\)\}\\,dt\+\\sigma\_\{z\}dB\_\{t\}^\{Z\}\\end\{aligned\}\\right\.\(80\)We takev=3\.0,ฮผ=3\.14,ฯx=ฯy=ฯz=0\.1v=3\.0,\\mu=3\.14,\\sigma\_\{x\}=\\sigma\_\{y\}=\\sigma\_\{z\}=0\.1and time pointst=0,1,2t=0,1,2\. Initialize\(X0\(0\),Y0\(0\),Z0\(0\)\)โโผโ๐ฉโ\(\[0,0,0\],0\.5โI\)\(X\_\{0\}^\{\(0\)\},Y\_\{0\}^\{\(0\)\},Z\_\{0\}^\{\(0\)\}\)\\overset\{\}\{\\sim\}\\mathcal\{N\}\(\[0,0,0\],0\.5I\),\(X0\(1\),Y0\(1\),Z0\(1\)\)โโผโ๐ฉโ\(\[0,โ2,0\],0\.5โI\),\(X0\(2\),Y0\(2\),Z0\(2\)\)โผ๐ฉโก\(\[0,2,0\],0\.5โI\)\(X\_\{0\}^\{\(1\)\},Y\_\{0\}^\{\(1\)\},Z\_\{0\}^\{\(1\)\}\)\\overset\{\}\{\\sim\}\\mathcal\{N\}\(\[0,\-2,0\],0\.5I\),\(X\_\{0\}^\{\(2\)\},Y\_\{0\}^\{\(2\)\},Z\_\{0\}^\{\(2\)\}\)\{\\sim\}\\mathcal\{N\}\(\[0,2,0\],0\.5I\)and sample500500particles for each\.


Figure 5:Dynamics of simulation datasets\. \(a\)2D simulation data, \(b\)3D simulation data and \(c\)3D simulation data projected to XY plane\.Cite Data\.We utilized the Cite\-seq dataset introduced in\[[31](https://arxiv.org/html/2608.21070#bib.bib31)\], which comprises 31,240 cells collected across 4 distinct time points\. To evaluate the performance of TracingFlow on real\-world biological data, we applied Principal Component Analysis \(PCA\) to reduce the dimensionality of the gene expression profiles to 100 and 5 dimensions, respectively\.
Embryoid Bodies Data\.We employed the Embryoid Bodies \(EB\) dataset from\[[39](https://arxiv.org/html/2608.21070#bib.bib39)\], consisting of 16,819 cells sampled at 5 time points\. We utilized PCA to reduce the dimensionality of the gene expression data to 5 dimensions\.
Hematopoiesis Data\.We employed the Hematopoiesis dataset from\[[64](https://arxiv.org/html/2608.21070#bib.bib64)\], consisting of 130,887 cells sampled at 3 time points\. Part of the cells are recorded with barcodes\. We selected 22,329 cells with barcodes from the dataset and utilized PCA to reduce the dimensionality of gene expression data to 50 dimensions\. When processing this dataset, we filtered for data where barcodes were present at all three time points\. The barcodes for the remaining data were set to undefined, incurring no additional transport cost during the incorporation of biological priors \([Section5\.3](https://arxiv.org/html/2608.21070#S5.SS3)\)\.
Mexico Gulf DataTo assess the capability and robustness of Tracing Flow in capturing long\-term system dynamics, we utilize the Mexico Gulf dataset, following\[[56](https://arxiv.org/html/2608.21070#bib.bib56)\]\. This two\-dimensional dataset comprises 9 time points\.
## Appendix CExperiment Details
### C\.1Experiment Details across various datasets
Simulation 2D Data[Table4](https://arxiv.org/html/2608.21070#A3.T4)presents the performance of TracingFlow and other algorithms on the 2D Simulation dataset\. On average, TracingFlow reconstructs the data distribution at each time point with higher accuracy\.
Table 4:๐ฒ1\\mathcal\{W\}\_\{1\}and๐ฒ2\\mathcal\{W\}\_\{2\}distances between the generated and ground\-truth distributions at each time point on the 2D Simulation dataset\.Cite 5D & Cite 100D DataWe present the distribution reconstruction performance of each algorithm on the Cite 5D and Cite 100D datasets in[Table5](https://arxiv.org/html/2608.21070#A3.T5)and[Table6](https://arxiv.org/html/2608.21070#A3.T6), respectively\. On these datasets TracingFlow achieves superior reconstruction accuracy at all time steps except fort=3t=3\. The trajectories learned by TracingFlow on these two datasets are illustrated in[Figure6](https://arxiv.org/html/2608.21070#A3.F6)\.
Table 5:๐ฒ1\\mathcal\{W\}\_\{1\}and๐ฒ2\\mathcal\{W\}\_\{2\}distances between the generated and ground\-truth distributions at each time point on the Cite 5D datasetTable 6:๐ฒ1\\mathcal\{W\}\_\{1\}and๐ฒ2\\mathcal\{W\}\_\{2\}distances between the generated and ground\-truth distributions at each time point on the Cite 100D datasetFigure 6:Trajectories learned by TracingFlow on the \(a\) Cite 5D and \(b\) Cite 100D datasets, projected to 2D using PCA\.EB 5D Data[Table7](https://arxiv.org/html/2608.21070#A3.T7)displays the interpolation accuracy on the EB 5D dataset, evaluated by holding out one time point during training\. TracingFlow outperformed other methods in reconstruction accuracy for all time points exceptt=1t=1\.[Figure7](https://arxiv.org/html/2608.21070#A3.F7)illustrates the interpolated distributions generated by TracingFlow across the four time points\.
Table 7:๐ฒ1\\mathcal\{W\}\_\{1\}and๐ฒ2\\mathcal\{W\}\_\{2\}distances between predicted and ground\-truth distributions at held\-out time points on the EB 5D dataset\.Figure 7:Predicted distributions by TracingFlow at four held\-out time points, projected to 2D using PCA\.3D Simulation Lineage Dataset[Table8](https://arxiv.org/html/2608.21070#A3.T8)demonstrates the ability of various algorithms to preserve biological priors on the 3D Simulation Lineage dataset\. Since the trajectories learned by TracingFlow align with biological priors \(see[Figure4](https://arxiv.org/html/2608.21070#S6.F4)\), it achieves lower lineage\-weighted๐ฒโ1\\mathcal\{W\}\{1\}and๐ฒโ2\\mathcal\{W\}\{2\}scores on average\.
Table 8:lineage\-weighted๐ฒ1\\mathcal\{W\}\_\{1\}and๐ฒ2\\mathcal\{W\}\_\{2\}distances between the generated and ground\-truth distributions at each time point on the 3D Simulation Lineage datasetHematopoiesis Dataset[Table9](https://arxiv.org/html/2608.21070#A3.T9)presents the performance of different algorithms in retaining biological priors on real\-world datasets, measured by the lineage\-weighted๐ฒโ1,๐ฒโ2\\mathcal\{W\}\{1\},\\mathcal\{W\}\{2\}distance\. TracingFlow surpassed other algorithms at all time points\. We also evaluated the algorithms without introducing biological priors by Vanilla๐ฒ1\\mathcal\{W\}\_\{1\},๐ฒ2\\mathcal\{W\}\_\{2\}distance, and as shown in[Table10](https://arxiv.org/html/2608.21070#A3.T10), TracingFlow maintained superior performance\.[Figure8](https://arxiv.org/html/2608.21070#A3.F8)visualizes the data points generated by TracingFlow on the real\-world dataset under both conditions \(with and without biological priors\)\.
Table 9:lineage\-weighted๐ฒ1\\mathcal\{W\}\_\{1\}and๐ฒ2\\mathcal\{W\}\_\{2\}distances between the generated and ground\-truth distributions at each time point on the Hematopoiesis datasetTable 10:๐ฒโ1\\mathcal\{W\}\{1\}and๐ฒโ2\\mathcal\{W\}\{2\}distances between the generated and ground\-truth distributions at each time point on the Hematopoiesis datasetFigure 8:Original data \(a\), and data points generated by TracingFlow on the Hematopoiesis dataset without biological priors \(b\) and with biological priors \(c\), projected to 2D using UMAP\.Mexico Gulf Dataset[Table11](https://arxiv.org/html/2608.21070#A3.T11)presents the performance of various algorithms on the Mexico Gulf dataset\. Tracing Flow achieves the best distribution reconstruction accuracy at all time points, yielding smooth trajectories\.
Table 11:๐ฒ1\\mathcal\{W\}\_\{1\}and๐ฒ2\\mathcal\{W\}\_\{2\}distances between the generated and ground\-truth distributions at each time point on the Mexico Gulf dataset
### C\.2Training time and scalability of TracingFlow
To evaluate the scalability of TracingFlow, we report the training time required by each method across various datasets, as presented in Table[12](https://arxiv.org/html/2608.21070#A3.T12)\. Experimental results indicate that Flow Matching algorithms based on second\-order dynamics \(e\.g\., 3 MSBM and our TracingFlow\) generally require more training time compared to first\-order dynamics algorithms \(e\.g\., OT\-CFM and SF2M\)\. Notably, the training time of TracingFlow is slightly lower than that of 3 MSBM, demonstrating that our method does not introduce additional computational overhead\.
Table 12:Training time comparison \(in seconds\) of different methods across various datasets\.
### C\.3Sensitivity Analysis on Minibatch\-OT
In all experiments presented in this paper, the full dataset was utilized to compute the OT plan\. However, as the scale of the data increases, the computational cost of the OT plan grows rapidly, making the adoption of the minibatch computation method described in[Section5\.2](https://arxiv.org/html/2608.21070#S5.SS2)inevitable\. We conducted a sensitivity analysis on the Cite\-5D dataset to evaluate the impact of batch size in minibatch\-OT computation on the accuracy of distribution reconstruction\. The results are presented in Table[13](https://arxiv.org/html/2608.21070#A3.T13)\. Experimental results demonstrate that TracingFlow consistently and accurately reconstructs the distributions at each time point when the batch size is adjusted within a certain range \( batch sizeโผ103\\sim 10^\{3\}\)\.
Table 13:Sensitivity analysis of batch size for Minibatch\-OT on the Cite\-5D dataset\.
## Appendix DPseudocode for Algorithm
The pseudocode for the training process of Tracing Flow is provided in[Algorithm1](https://arxiv.org/html/2608.21070#alg1)\.
Algorithm 1Tracing\-Flow Matching Training0:Snapshot datasets
๐k=\{๐k\(i\)\}i=1Nk\\mathcal\{D\}\_\{k\}=\\\{\\bm\{x\}\_\{k\}^\{\(i\)\}\\\}\_\{i=1\}^\{N\_\{k\}\}at times
t0<t1<โฏ<tKt\_\{0\}<t\_\{1\}<\\dots<t\_\{K\}; Acceleration Network
๐๐ฝ\\bm\{a\}\_\{\\bm\{\\theta\}\}; Velocity Network
๐๐\\bm\{v\}\_\{\\bm\{\\xi\}\}; Hyperparameters: batch size
BtrajB\_\{\\text\{traj\}\}, learning rates
ฮท๐,ฮท๐\\eta\_\{\\bm\{a\}\},\\eta\_\{\\bm\{v\}\}\.
1:Initialize:Estimated velocity sets
๐ฑk=\{๐k\(i\)\}i=1Nkโ\{๐\}\\mathcal\{V\}\_\{k\}=\\\{\\bm\{v\}\_\{k\}^\{\(i\)\}\\\}\_\{i=1\}^\{N\_\{k\}\}\\leftarrow\\\{\\mathbf\{0\}\\\}for all
kk\.
2:whileVelocity Estimates Not Convergeddo
3:Stage 1: Velocity Assignment \(Iterative SOAT\)
4:Initialize velocity accumulators
๐ฑ^k\(i\)โโ
\\hat\{\\mathcal\{V\}\}\_\{k\}^\{\(i\)\}\\leftarrow\\emptysetfor all
k,ik,i\.
5:// Step 1\.1: Solve Optimal Transport
6:for
k=0k=0to
Kโ1K\-1do
7:Compute cost matrix
๐iโi\+1\\mathcal\{C\}\_\{i\\rightarrow i\+1\}with / w\.o\. biological prior\.
8:Compute optimal coupling
ฯkโk\+1โ\\pi^\{\*\}\_\{k\\to k\+1\}between
\(๐k,๐k\)\(\\bm\{x\}\_\{k\},\\bm\{v\}\_\{k\}\)and
\(๐k\+1,๐k\+1\)\(\\bm\{x\}\_\{k\+1\},\\bm\{v\}\_\{k\+1\}\)by solving the SOAT Problem\.
9:Compute propagator \(transition\) matrix
๐ฆkโk\+1โ\\mathcal\{K\}^\{\*\}\_\{k\\to k\+1\}by normalizing
ฯkโk\+1โ\\pi^\{\*\}\_\{k\\to k\+1\}:
10:
๐ฆkโk\+1โโ\(i,j\)=ฯkโk\+1โโ\(i,j\)โjโฒฯkโk\+1โโ\(i,jโฒ\)\\mathcal\{K\}^\{\*\}\_\{k\\to k\+1\}\(i,j\)=\\frac\{\\pi^\{\*\}\_\{k\\to k\+1\}\(i,j\)\}\{\\sum\_\{j^\{\\prime\}\}\\pi^\{\*\}\_\{k\\to k\+1\}\(i,j^\{\\prime\}\)\}
11:endfor
12:// Step 1\.2: Trajectory Sampling & Velocity Update
13:Sample
BtrajB\_\{\\text\{traj\}\}discrete trajectories
Tm=\{\(๐k\(sm,k\),๐k\(sm,k\)\)\}k=0KT\_\{m\}=\\\{\(\\bm\{x\}\_\{k\}^\{\(s\_\{m,k\}\)\},\\bm\{v\}\_\{k\}^\{\(s\_\{m,k\}\)\}\)\\\}\_\{k=0\}^\{K\}for
m=1โโฆโBtrajm=1\\dots B\_\{\\text\{traj\}\}\.
14:
โ
\\cdotInitial indices
sm,0s\_\{m,0\}are sampled uniformly\.
15:
โ
\\cdotSubsequent indices
sm,k\+1โผ๐ฆkโk\+1โโ\(sm,k,โ
\)s\_\{m,k\+1\}\\sim\\mathcal\{K\}^\{\*\}\_\{k\\to k\+1\}\(s\_\{m,k\},\\cdot\)\.
16:for
m=1m=1to
BtrajB\_\{\\text\{traj\}\}do
17:Compute optimal path velocities
\{๐^k\(sm,k\)\}k=0K\\\{\\hat\{\\bm\{v\}\}\_\{k\}^\{\(s\_\{m,k\}\)\}\\\}\_\{k=0\}^\{K\}for trajectory
TmT\_\{m\}using Proposition[4\.3](https://arxiv.org/html/2608.21070#S4.SS3.SSS0.Px1)\.
18:Accumulate:
๐ฑ^k\(sm,k\)โ๐ฑ^k\(sm,k\)โช\{๐^k\(sm,k\)\}\\hat\{\\mathcal\{V\}\}\_\{k\}^\{\(s\_\{m,k\}\)\}\\leftarrow\\hat\{\\mathcal\{V\}\}\_\{k\}^\{\(s\_\{m,k\}\)\}\\cup\\\{\\hat\{\\bm\{v\}\}\_\{k\}^\{\(s\_\{m,k\}\)\}\\\}for
k=0โโฆโKk=0\\dots K\.
19:endfor
20:// Step 1\.3: Average Accumulated Velocities
21:for
k=0k=0to
KK,
i=1i=1to
NkN\_\{k\}do
22:if
๐ฑ^k\(i\)โ โ
\\hat\{\\mathcal\{V\}\}\_\{k\}^\{\(i\)\}\\neq\\emptysetthen
23:
๐k\(i\)โ1\|๐ฑ^k\(i\)\|โโ๐โ๐ฑ^k\(i\)๐\\bm\{v\}\_\{k\}^\{\(i\)\}\\leftarrow\\frac\{1\}\{\|\\hat\{\\mathcal\{V\}\}\_\{k\}^\{\(i\)\}\|\}\\sum\_\{\\bm\{v\}\\in\\hat\{\\mathcal\{V\}\}\_\{k\}^\{\(i\)\}\}\\bm\{v\}
24:endif
25:endfor
26:endwhile
27:Stage 2: Final Transport Plan Computation
28:Compute optimal couplings
ฯโ\\pi^\{\*\}and propagators
๐ฆโ\\mathcal\{K\}^\{\*\}using the converged velocities
๐ฑk\\mathcal\{V\}\_\{k\}\.
29:Stage 3: Neural Network Training
30:whileNot Convergeddo
31:Sample random time
tโผ๐ฐโก\[t0,tK\]t\\sim\\mathcal\{U\}\[t\_\{0\},t\_\{K\}\]and coupling
๐โผqโก\(๐\)\\bm\{z\}\\sim q\(\\bm\{z\}\)\.
32:Compute Acceleration Matching Loss:
33:
โCAFM=๐ผ\(๐,๐\)โผฯt\(โ
\|z\),๐โผq\(๐\),tโผ๐ฐ\[t0,tK\]\[โฅ๐๐ฝ\(๐,๐,t\)โ๐\(๐,๐,t\|z\)โฅ2\]\\mathcal\{L\}\_\{\\text\{CAFM\}\}=\\mathbb\{E\}\_\{\(\\bm\{x\},\\bm\{v\}\)\\sim\\rho\_\{t\}\(\\cdot\|z\),\\bm\{z\}\\sim q\(\\bm\{z\}\),t\\sim\\mathcal\{U\}\[t\_\{0\},t\_\{K\}\]\}\\Big\[\\\|\\bm\{a\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},\\bm\{v\},t\)\-\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\|z\)\\\|^\{2\}\\Big\]
34:Compute Initial Velocity Loss:
35:
โ๐0=๐ผ\(๐,๐\)โผฯ0โ\[โ๐๐โ\(๐\)โ๐โ2\]\\mathcal\{L\}\_\{\\bm\{v\}\_\{0\}\}=\\mathbb\{E\}\_\{\(\\bm\{x\},\\bm\{v\}\)\\sim\\rho\_\{0\}\}\\Big\[\\\|\\bm\{v\}\_\{\\bm\{\\xi\}\}\(\\bm\{x\}\)\-\\bm\{v\}\\\|^\{2\}\\Big\]
36:Update parameters \(Gradient Descent\):
37:
๐ฝโ๐ฝโฮท๐โโ๐ฝโCAFM\\bm\{\\theta\}\\leftarrow\\bm\{\\theta\}\-\\eta\_\{\\bm\{a\}\}\\nabla\_\{\\bm\{\\theta\}\}\\mathcal\{L\}\_\{\\text\{CAFM\}\}
38:
๐โ๐โฮท๐โโ๐โ๐0\\bm\{\\xi\}\\leftarrow\\bm\{\\xi\}\-\\eta\_\{\\bm\{v\}\}\\nabla\_\{\\bm\{\\xi\}\}\\mathcal\{L\}\_\{\\bm\{v\}\_\{0\}\}
39:endwhile
## Appendix EDiscussion
### E\.1Relation to other Algorithms
#### Relation to 3 MSBM
3 MSBM considers the following stochastic optimal control problem:
minโก๐ผฯtโโซ\[โ๐tโ2\]โ๐t\\min\\mathbb\{E\}\_\{\\rho\_\{t\}\}\\int\[\\\|\\bm\{a\}\_\{t\}\\\|^\{2\}\]\\mathrm\{d\}t\(81\)dโ๐t=Aโ๐tโ๐t\+๐tโ๐t\+gโdโ๐พt,๐nโผqn=โซฯnโ\(๐,๐\)โdโ๐n\\mathrm\{d\}\\bm\{m\}\_\{t\}=A\\bm\{m\}\_\{t\}\\mathrm\{d\}t\+\\bm\{u\}\_\{t\}\\mathrm\{d\}t\+g\\mathrm\{d\}\\bm\{W\}\_\{t\},\\quad\\bm\{x\}\_\{n\}\\sim q\_\{n\}=\\int\\pi\_\{n\}\(\\bm\{x\},\\bm\{v\}\)\\mathrm\{d\}\\bm\{v\}\_\{n\}\(82\)where๐t=\[๐t,๐t\]T\\bm\{m\}\_\{t\}=\[\\bm\{x\}\_\{t\},\\bm\{v\}\_\{t\}\]^\{T\},A=\[0001\]A=\\begin\{bmatrix\}0&0\\\\ 0&1\\end\{bmatrix\}, andg=\[000ฯ\]g=\\begin\{bmatrix\}0&0\\\\ 0&\\sigma\\end\{bmatrix\}\. This problem can be viewed as a โstochasticโ version of the VM\-DOAT problem proposed in this paper\. 3 MSBM also provides a simulation\-free training method to regress the acceleration field\. However, 3 MSBM does not explicitly solve this stochastic optimal control problem using the form of optimal coupling plus optimal single\-particle trajectories\. Its optimal coupling requires iterative solutions similar to Rectified Flow\[[36](https://arxiv.org/html/2608.21070#bib.bib36)\]: a process of repeatedly selecting the coupling and trainingaฮธa\_\{\\theta\}\. Furthermore, it relies on a heuristic method to estimate the initial velocity \(via normal initialization and iteratively solving forward and backward SDEs\), rather than incorporating this estimation as part of the optimal control problem\. Experiments show that Tracing Flow achieves better distribution reconstruction accuracy than 3MSBM\.
#### Relation to MMFM
In MMFM, the single\-particle path is similarly specified by:
โซโฮณโฒโฒโ\(t\)โ2โ๐t,๐k=ฮณโก\(tk\)\\int\\\|\\gamma^\{\\prime\\prime\}\(t\)\\\|^\{2\}\\mathrm\{d\}t,\\quad\\bm\{x\}\_\{k\}=\\gamma\(t\_\{k\}\)\(83\)which corresponds to a natural spline\. However, MMFM does not solve an optimal control problem over distributions\. Furthermore, the velocity field in MMFM remains single\-valued with respect to๐\\bm\{x\}, which prevents it from learning trajectories that cross in the position space\.
#### Relation to OAT\-FM
OAT\-FM introduces a DOAT problem formulation consistent with this paper; however, the objective of OAT\-FM is not to genuinely solve the DOAT problem within the augmented space๐ณร๐ฑ\\mathcal\{X\}\\times\\mathcal\{V\}, but rather to fine\-tune a pre\-trained FM model\. It employs the loss:
โCAFM\\displaystyle\\mathcal\{L\}\_\{\\text\{CAFM\}\}=๐ผฯโโ\[\(๐0,๐0\),\(๐1,๐1\)\]\[ฮฑโฅ๐tโ๐0tโ๐0\+๐๐ฝ2โฅ2\+\(1โฮฑ\)โฅ๐ฮธ\(๐t,t\)โ๐0โฅ2\\displaystyle=\\mathbb\{E\}\_\{\\pi^\{\*\}\[\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\),\(\\bm\{x\}\_\{1\},\\bm\{v\}\_\{1\}\)\]\}\\big\[\\alpha\\\|\\dfrac\{\\bm\{x\}\_\{t\}\-\\bm\{x\}\_\{0\}\}\{t\}\-\\dfrac\{\\bm\{v\}\_\{0\}\+\\bm\{v\}\_\{\\bm\{\\theta\}\}\}\{2\}\\\|^\{2\}\+\(1\-\\alpha\)\\\|\\bm\{v\}\_\{\\theta\}\(\\bm\{x\}\_\{t\},t\)\-\\bm\{v\}\_\{0\}\\\|^\{2\}\+ฮฑโฅ๐1โ๐t1โtโ๐๐ฝ\+๐12โฅ2\+\(1โฮฑ\)โฅ๐1โ๐ฮธ\(๐t,t\)โฅ2\]\\displaystyle\+\\alpha\\\|\\dfrac\{\\bm\{x\}\_\{1\}\-\\bm\{x\}\_\{t\}\}\{1\-t\}\-\\dfrac\{\\bm\{v\}\_\{\\bm\{\\theta\}\}\+\\bm\{v\}\_\{1\}\}\{2\}\\\|^\{2\}\+\(1\-\\alpha\)\\\|\\bm\{v\}\_\{1\}\-\\bm\{v\}\_\{\\theta\}\(\\bm\{x\}\_\{t\},t\)\\\|^\{2\}\\big\]\(84\)to enforce the velocity field along the line segment connecting any two samples to be as parallel to that line as possible\. As a fine\-tuning approach, it effectively still addresses a distribution transport problem involving only two time points within the position space๐ณ\\mathcal\{X\}, rather than tackling a multi\-marginal optimal control problem\. Furthermore, OAT\-FM does not establish a connection between its loss function and the cost along each trajectory12โโซ01โฮณยจโ2โ๐t\\dfrac\{1\}\{2\}\\int\_\{0\}^\{1\}\\\|\{\\ddot\{\\gamma\}\}\\\|^\{2\}\\mathrm\{d\}t, nor does it explicitly learn the acceleration field๐โก\(๐,๐,t\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\)\.
### E\.2Is๐ฎ=๐ณร๐ฑ\\mathcal\{S\}=\\mathcal\{X\}\\times\\mathcal\{V\}a Phase Space?
In the Conclusion and Limitation section, we noted that if๐\\bm\{x\}is regarded as the generalized coordinate and๐\\bm\{v\}as its time derivative, the termโซ12โโ๐โ2โ๐t\\int\\frac\{1\}\{2\}\\\|\\bm\{a\}\\\|^\{2\}\\mathrm\{d\}tcannot be interpreted as a physical action, as the action in classical mechanics typically does not involve second\-order time derivatives\. Here, we offer an alternative physical interpretation of this control objective\.
###### Proposition E\.1\.
Consider add\-dimensional second\-order dynamical system governed by the equations:
๐ห\\displaystyle\\dot\{\\bm\{x\}\}=๐\\displaystyle=\\bm\{v\}\(85\)๐ห\\displaystyle\\dot\{\\bm\{v\}\}=๐\\displaystyle=\\bm\{a\}\(86\)where๐ฑ,๐ฏ,๐โโd\\bm\{x\},\\bm\{v\},\\bm\{a\}\\in\\mathbb\{R\}^\{d\}\. The optimal control objective is given by:
minโก12โโซ0Tโ๐โ2โ๐t\\min\\frac\{1\}\{2\}\\int\_\{0\}^\{T\}\\\|\\bm\{a\}\\\|^\{2\}\\mathrm\{d\}t\(87\)
The evolution of this system under the optimal control law is equivalent to a Hamiltonian system in4โd4d\-dimensional phase space with the following Hamiltonian:
H=๐1Tโ๐2\+12โโ๐2โ2H=\\bm\{p\}\_\{1\}^\{T\}\\bm\{q\}\_\{2\}\+\\frac\{1\}\{2\}\\\|\\bm\{p\}\_\{2\}\\\|^\{2\}\(88\)where๐ช1,๐ช2โโd\\bm\{q\}\_\{1\},\\bm\{q\}\_\{2\}\\in\\mathbb\{R\}^\{d\}are the canonical coordinates, and๐ฉ1,๐ฉ2โโd\\bm\{p\}\_\{1\},\\bm\{p\}\_\{2\}\\in\\mathbb\{R\}^\{d\}are the canonical momenta \(note that๐ช1\\bm\{q\}\_\{1\}does not appear in the Hamiltonian\)\. Here,๐ช1\\bm\{q\}\_\{1\}and๐ช2\\bm\{q\}\_\{2\}correspond to๐ฑ\\bm\{x\}and๐ฏ\\bm\{v\}in the optimal control problem, respectively, while๐ฉ2\\bm\{p\}\_\{2\}corresponds to๐\\bm\{a\}\.
###### Proof\.
We derive the equations of motion directly using Hamiltonโs canonical equations:
๐ห1=โHโ๐1=๐2,๐ห2=โHโ๐2=๐2,๐ห1=โโHโ๐1=๐,๐ห2=โโHโ๐2=โ๐1\\dot\{\\bm\{q\}\}\_\{1\}=\\frac\{\\partial H\}\{\\partial\\bm\{p\}\_\{1\}\}=\\bm\{q\}\_\{2\},\\quad\\dot\{\\bm\{q\}\}\_\{2\}=\\frac\{\\partial H\}\{\\partial\\bm\{p\}\_\{2\}\}=\\bm\{p\}\_\{2\},\\quad\\dot\{\\bm\{p\}\}\_\{1\}=\-\\frac\{\\partial H\}\{\\partial\\bm\{q\}\_\{1\}\}=\\bm\{0\},\\quad\\dot\{\\bm\{p\}\}\_\{2\}=\-\\frac\{\\partial H\}\{\\partial\\bm\{q\}\_\{2\}\}=\-\\bm\{p\}\_\{1\}\(89\)
This implies that๐1\\bm\{p\}\_\{1\}is constant\. Solving these equations sequentially yields:
๐2โ\(t\)\\displaystyle\\bm\{p\}\_\{2\}\(t\)=๐2โฒโ๐1โt\\displaystyle=\\bm\{c\}\_\{2\}^\{\\prime\}\-\\bm\{p\}\_\{1\}t\(90\)๐2โ\(t\)\\displaystyle\\bm\{q\}\_\{2\}\(t\)=๐1โฒ\+๐2โฒโtโ12โ๐1โt2\\displaystyle=\\bm\{c\}\_\{1\}^\{\\prime\}\+\\bm\{c\}\_\{2\}^\{\\prime\}t\-\\frac\{1\}\{2\}\\bm\{p\}\_\{1\}t^\{2\}\(91\)๐1โ\(t\)\\displaystyle\\bm\{q\}\_\{1\}\(t\)=๐0โฒ\+๐1โฒ\+12โ๐2โฒโt2โ16โ๐1โt3\\displaystyle=\\bm\{c\}\_\{0\}^\{\\prime\}\+\\bm\{c\}\_\{1\}^\{\\prime\}\+\\frac\{1\}\{2\}\\bm\{c\}\_\{2\}^\{\\prime\}t^\{2\}\-\\frac\{1\}\{6\}\\bm\{p\}\_\{1\}t^\{3\}\(92\)Comparing this with the results in[Equation9](https://arxiv.org/html/2608.21070#S4.E9), we observe that๐1โ\(t\)\\bm\{q\}\_\{1\}\(t\)and๐2โ\(t\)\\bm\{q\}\_\{2\}\(t\)satisfy the same equations as๐โก\(t\)\\bm\{x\}\(t\)and๐โก\(t\)\\bm\{v\}\(t\), with the undetermined constants determined by the boundary conditions\. This completes the proof\. โ
Consequently, we observe that although the augmented space๐ฎ=๐ณร๐ฑ\\mathcal\{S\}=\\mathcal\{X\}\\times\\mathcal\{V\}concatenates position๐\\bm\{x\}and velocity \(momentum\)๐\\bm\{v\}, itcannotbe interpreted as aphase space\. Instead, it should be viewed as theconfiguration spaceof the classical mechanical system described by the Hamiltonian above\. We can attempt to recover the Lagrangian from this Hamiltonian\. The canonical equations established that๐ห1=๐2\\dot\{\\bm\{q\}\}\_\{1\}=\\bm\{q\}\_\{2\}and๐ห2=๐2\\dot\{\\bm\{q\}\}\_\{2\}=\\bm\{p\}\_\{2\}\. Applying the Legendre transformation:
L\\displaystyle L=๐1Tโ๐ห1\+๐2Tโ๐ห2โH\\displaystyle=\\bm\{p\}\_\{1\}^\{T\}\\dot\{\\bm\{q\}\}\_\{1\}\+\\bm\{p\}\_\{2\}^\{T\}\\dot\{\\bm\{q\}\}\_\{2\}\-H=๐1Tโ\(๐ห1โ๐2\)\+12โโ๐ห2โ2\\displaystyle=\\bm\{p\}\_\{1\}^\{T\}\(\\dot\{\\bm\{q\}\}\_\{1\}\-\\bm\{q\}\_\{2\}\)\+\\frac\{1\}\{2\}\\\|\\dot\{\\bm\{q\}\}\_\{2\}\\\|^\{2\}\(93\)We find that the term๐1T\\bm\{p\}\_\{1\}^\{T\}cannot be eliminated; the Lagrangian cannot be expressed solely as a function of the canonical coordinates and their first derivatives\. This is a characteristic of constrained systems \(in Diracโs theory of constraints, a constrained system is typically defined as one where the Hessian matrix of the Lagrangian is singular\. To see this intuitively, if we regard๐1\\bm\{p\}\_\{1\}in the Lagrangian asddnew canonical coordinates, their corresponding canonical momenta are zero, implying the motion is confined to a submanifold of the phase space\)\.
In[Equation10](https://arxiv.org/html/2608.21070#S4.E10), we derived the minimum cost for the system to transition from\(๐1,๐1\)\(\\bm\{x\}\_\{1\},\\bm\{v\}\_\{1\}\)to\(๐2,๐2\)\(\\bm\{x\}\_\{2\},\\bm\{v\}\_\{2\}\)\. While this cost is a symmetric positive\-definite quadratic form with respect to๐\\bm\{x\}and๐\\bm\{v\}, it does not constitute a metric on the augmented space๐ฎ=๐ณร๐ฑ\\mathcal\{S\}=\\mathcal\{X\}\\times\\mathcal\{V\}\. \(If it were a metric, the augmented space would be flat, and the optimal control trajectoriesโgeodesicsโwould be linear functions oftt\)\. Maupertuisโ principle suggests a close relationship between the metric in the configuration space and the systemโs Lagrangian\[[3](https://arxiv.org/html/2608.21070#bib.bib3)\]: if the Lagrangian takes the form:
Lโก\(๐,๐ห\)=12โmiโjโ\(๐\)โqหiโqหjโVโก\(๐\)L\(\\bm\{q\},\\dot\{\\bm\{q\}\}\)=\\frac\{1\}\{2\}m\_\{ij\}\(\\bm\{q\}\)\\dot\{q\}^\{i\}\\dot\{q\}^\{j\}\-V\(\\bm\{q\}\)\(94\)then the configuration space possesses the metric:
giโjโ\(๐\)=2โ\(EโVโก\(๐\)\)โmiโjโ\(๐\)g\_\{ij\}\(\\bm\{q\}\)=2\(E\-V\(\\bm\{q\}\)\)m\_\{ij\}\(\\bm\{q\}\)\(95\)whereEEis the total energy of the particle\. For the optimal control problem discussed herein, we have seen that the Lagrangian cannot be written in such a form; therefore, it is fundamentally impossible to equip the configuration space๐ฎ=๐ณร๐ฑ\\mathcal\{S\}=\\mathcal\{X\}\\times\\mathcal\{V\}with such a metric\.
### E\.3Broader Impacts
This paper presents work whose goal is to advance the field of machine learning\. There are many potential societal consequences of our work, none of which we feel must be specifically highlighted here\. However, as our algorithm is applied to single\-cell data analysis, the fidelity of the generated trajectories is heavily dependent on the quality of the input data\. Consequently, when utilizing our method for biological or medical research purposes, it is critical to employ high\-quality datasets, incorporate established and accurate biological priors, and ensure that the algorithmic outputs are rigorously validated by domain experts\.Similar Articles
Capturing non-Markovian dynamics in non-equilibrium stochastic systems using flow matching
This paper develops a generative flow matching method to capture non-Markovian dynamics in non-equilibrium stochastic systems, demonstrating improved predictions for the Kramers first passage time problem compared to Markovian baselines.
Trajectory as the Teacher: Few-Step Discrete Flow Matching via Energy-Navigated Distillation
This paper introduces Trajectory-Shaped Discrete Flow Matching (TS-DFM), which replaces blind stochastic jumps with guided navigation to significantly improve text generation efficiency and reduce computational costs. The method achieves superior perplexity and speed compared to traditional multi-step baselines while maintaining unchanged inference costs.
Recursive Flow Matching
Introduces Recursive Flow Matching (RecFM), a generative framework for forecasting complex spatiotemporal dynamics that achieves high fidelity with fewer steps and improved accuracy and speed, including up to 20x speedup over diffusion-based emulators.
Two-Parameter Flows for Learning Population Dynamics of Physical Systems
Proposes two-parameter flows to learn the dynamics of high-dimensional probability densities from unlabeled samples, using conditional flow matching to extract physics-time velocity fields.
PrismFlow: Residual Dynamics for Flow Matching in Time-Series Generation
PrismFlow introduces a Flow Matching method with Koopman-inspired dynamical experts to handle multimodal and multiscale time-series data, achieving state-of-the-art performance with significant improvements in Context-FID and Discriminative Score.