TracingFlow: A Simulation-Free Trajectory Inference Framework Based on Second-Order Dynamics

arXiv cs.LG Papers

Summary

TracingFlow is a simulation-free framework using second-order dynamics for trajectory inference, improving accuracy in capturing high-curvature transitions in single-cell omics data.

arXiv:2608.21070v1 Announce Type: new Abstract: Inferring continuous system evolution from sparse temporal snapshots is a key challenge in generative modeling and single-cell omics. While Optimal Transport (OT) is popular, existing frameworks are largely restricted to first-order dynamics, assuming memoryless velocity fields. This limits expressiveness, as first-order systems fail to account for regulatory momentum and time-delayed responses inherent in processes like cell differentiation. Here, we introduce TracingFlow, a simulation-free Flow Matching framework generalizing to second-order dynamics. By using neural networks to regress the acceleration field, TracingFlow provides an exact, efficient solution to the Dynamical Optimal Acceleration Transport (DOAT) problem. Unlike first-order methods yielding over-smoothed trajectories, our second-order formulation captures high-curvature transitions and nonlinear evolutions by learning the underlying force fields. Evaluated on complex synthetic and large-scale scRNA-seq datasets, TracingFlow achieves superior accuracy in distributional reconstruction and trajectory faithfulness. Moreover, by integrating lineage tracing priors, it recovers dynamical structures that are both mathematically optimal and biologically plausible.
Original Article
View Cached Full Text

Cached at: 08/24/26, 04:36 AM

# A Simulation-Free Trajectory Inference FrameworkBased on Second-Order Dynamics
Source: [https://arxiv.org/html/2608.21070](https://arxiv.org/html/2608.21070)
## TracingFlow: A Simulation\-Free Trajectory Inference Framework Based on Second\-Order Dynamics

Yuhao SunZekun Wu11footnotemark:1Affiliation:School of Mathematical Sciences, Peking UniversityZixun HuangAffiliation:School of Mathematical Sciences, Peking UniversityPeijie ZhouThanks:Corresponding author:pjzhou@pku\.edu\.cnAffiliation:Center for Machine Learning Research, Peking UniversityAffiliation:Center for Quantitative Biology, Peking UniversityAffiliation:National Engineering Laboratory for Big Data Analysis and Applications, BeijingAffiliation:AI for Science Institute, Beijing

###### Abstract

Inferring continuous system evolution from sparse temporal snapshots is a key challenge in generative modeling and single\-cell omics\. While Optimal Transport \(OT\) is popular, existing frameworks are largely restricted to first\-order dynamics, assuming memoryless velocity fields\. This limits expressiveness, as first\-order systems fail to account for regulatory momentum and time\-delayed responses inherent in processes like cell differentiation\. Here, we introduce TracingFlow, a simulation\-free Flow Matching framework generalizing to second\-order dynamics\. By using neural networks to regress the acceleration field, TracingFlow provides an exact, efficient solution to the Dynamical Optimal Acceleration Transport \(DOAT\) problem\. Unlike first\-order methods yielding over\-smoothed trajectories, our second\-order formulation captures high\-curvature transitions and nonlinear evolutions by learning the underlying force fields\. Evaluated on complex synthetic and large\-scale scRNA\-seq datasets, TracingFlow achieves superior accuracy in distributional reconstruction and trajectory faithfulness\. Moreover, by integrating lineage tracing priors, it recovers dynamical structures that are both mathematically optimal and biologically plausible\.

โ€ โ€ footnotetext:Preprint\. August 21, 2026\.## 1Introduction

Recovering underlying dynamics from discrete observations is a pivotal task in single\-cell omics, known as Trajectory Inference \(TI\)\[[9](https://arxiv.org/html/2608.21070#bib.bib9),[72](https://arxiv.org/html/2608.21070#bib.bib72)\]\. To identify the least costly evolution between distributions, Optimal Transport \(OT\) frameworks have emerged\[[7](https://arxiv.org/html/2608.21070#bib.bib7),[71](https://arxiv.org/html/2608.21070#bib.bib71)\]\. These generally categorize into Static OT methods\[[51](https://arxiv.org/html/2608.21070#bib.bib51),[28](https://arxiv.org/html/2608.21070#bib.bib28),[22](https://arxiv.org/html/2608.21070#bib.bib22)\], which learn direct mappings, and Dynamical OT methods\[[57](https://arxiv.org/html/2608.21070#bib.bib57),[24](https://arxiv.org/html/2608.21070#bib.bib24),[67](https://arxiv.org/html/2608.21070#bib.bib67),[44](https://arxiv.org/html/2608.21070#bib.bib44)\], which capture continuous evolution via flow maps\.

Early Dynamical OT frameworks relied on Neural ODEs\[[9](https://arxiv.org/html/2608.21070#bib.bib9)\], incurring high computational costs due to numerical integration\. To address this, Flow Matching\[[35](https://arxiv.org/html/2608.21070#bib.bib35),[36](https://arxiv.org/html/2608.21070#bib.bib36),[48](https://arxiv.org/html/2608.21070#bib.bib48)\]was proposed as a simulation\-free alternative, regressing velocity fields to push source distributions to targets\. By decomposing transport costs into single\-particle trajectories, Flow Matching efficiently solves Dynamical OT problems\[[58](https://arxiv.org/html/2608.21070#bib.bib58),[27](https://arxiv.org/html/2608.21070#bib.bib27),[50](https://arxiv.org/html/2608.21070#bib.bib50)\], including its various variants\[[59](https://arxiv.org/html/2608.21070#bib.bib59),[15](https://arxiv.org/html/2608.21070#bib.bib15),[8](https://arxiv.org/html/2608.21070#bib.bib8),[45](https://arxiv.org/html/2608.21070#bib.bib45)\]\.

However, most TI frameworks assume first\-order dynamics \(๐’™ห™=๐’—โก\(๐’™,t\)\\bm\{\\dot\{x\}\}=\\bm\{v\}\(\\bm\{x\},t\)\), inherently restricting๐’—\\bm\{v\}to be single\-valued w\.r\.t๐’™\\bm\{x\}\. This limits expressiveness when modeling complex biological priors, such as lineage tracing\[[30](https://arxiv.org/html/2608.21070#bib.bib30),[38](https://arxiv.org/html/2608.21070#bib.bib38)\], potentially yielding dynamics that contradict with additional biological prior information\. Furthermore,\[[19](https://arxiv.org/html/2608.21070#bib.bib19),[34](https://arxiv.org/html/2608.21070#bib.bib34),[23](https://arxiv.org/html/2608.21070#bib.bib23)\]pointed out that if more modalities such as chromatin accessibility or proteomics are considered, the biological kinetics may follow higher\-order dynamics or could be modelled in augmented space\. While 3MSBM\[[56](https://arxiv.org/html/2608.21070#bib.bib56)\]recently introduced a second\-order Momentum Schrรถdinger Bridge for smoother trajectories, it relies on iterative retraining the acceleration field similar to Rectified Flow\[[36](https://arxiv.org/html/2608.21070#bib.bib36)\]to determine optimal couplings\. Furthermore, it employs a heuristic method to estimate the initial velocity, rather than incorporating this estimation as part of the optimal control problem\.

To overcome these limitations, we introduceTracingFlow, a simulation\-free framework utilizing second\-order dynamical systems\. We propose the Dynamical Optimal Acceleration Transport \(DOAT\) problem, which determines evolution by minimizing acceleration costs\. TracingFlow directly regresses the acceleration field and initial velocity, avoiding ODE simulation\. Our experiments show that TracingFlow achieves competitive and often improved reconstruction accuracy relative to existing simulation\-free frameworks\[[58](https://arxiv.org/html/2608.21070#bib.bib58),[59](https://arxiv.org/html/2608.21070#bib.bib59),[69](https://arxiv.org/html/2608.21070#bib.bib69),[42](https://arxiv.org/html/2608.21070#bib.bib42)\], while effectively incorporating biological priors\. Our contributions include:

- โ€ขWe propose TracingFlow, a simulation\-free TI framework based on second\-order dynamics\. It solves the proposed DOAT problem by regressing acceleration fields, reducing computational costs compared to ODE\-based methods\.
- โ€ขWe provide theoretical guarantees by decoupling the design of the optimal transport plan and single\-particle control\. We further introduce an iterative strategy to transform the real\-world Velocity Missed\-DOAT \(VM\-DOAT\) problem into a standard DOAT problem, enabling approximate solutions\.
- โ€ขWe demonstrate TracingFlowโ€™s effectiveness on multi\-time\-point real world datasets\. Results indicate superior distribution reconstruction accuracy and enhanced capability in preserving biological priors compared to state\-of\-the\-art methods\.

![Refer to caption](https://arxiv.org/html/2608.21070v1/illustration.png)Figure 1:An illustration figure for TracingFlow
## 2Related Works

#### Solving Optimal Transport via Flow Matching\.

Flow Matching\[[35](https://arxiv.org/html/2608.21070#bib.bib35),[36](https://arxiv.org/html/2608.21070#bib.bib36),[48](https://arxiv.org/html/2608.21070#bib.bib48),[1](https://arxiv.org/html/2608.21070#bib.bib1)\]transports probability distributions via learned flow maps, offering scalable simulation\-free training\. As an optimal control problem, Optimal Transport \(OT\)\[[47](https://arxiv.org/html/2608.21070#bib.bib47)\]decomposes into Optimal Coupling and single\-particle geodesics, facilitating solutions via Flow Matching\[[58](https://arxiv.org/html/2608.21070#bib.bib58),[27](https://arxiv.org/html/2608.21070#bib.bib27),[50](https://arxiv.org/html/2608.21070#bib.bib50),[15](https://arxiv.org/html/2608.21070#bib.bib15)\]\. This approach extends to various variants and generalizations\[[59](https://arxiv.org/html/2608.21070#bib.bib59),[15](https://arxiv.org/html/2608.21070#bib.bib15),[8](https://arxiv.org/html/2608.21070#bib.bib8),[13](https://arxiv.org/html/2608.21070#bib.bib13),[61](https://arxiv.org/html/2608.21070#bib.bib61),[45](https://arxiv.org/html/2608.21070#bib.bib45),[26](https://arxiv.org/html/2608.21070#bib.bib26),[68](https://arxiv.org/html/2608.21070#bib.bib68),[4](https://arxiv.org/html/2608.21070#bib.bib4),[46](https://arxiv.org/html/2608.21070#bib.bib46)\]\. However, these typically employ first\-order dynamics\. While second\-order approaches have emerged\[[56](https://arxiv.org/html/2608.21070#bib.bib56)\], they require iterative training of the acceleration field to achieve optimal coupling\. Moreover, they rely on heuristic methods for initial velocity estimation, rather than integrating it as a component of the optimal control problem\. TracingFlow addresses this by decomposing the Dynamical Acceleration Optimal Transport \(DOAT\) problem into optimal coupling and single particle optimal trajectory, achieving a truly simulation\-free process for DOAT\.

#### Overcoming Single\-Valued Velocity Constraint in Flow Matching\.

In standard Flow Matching, the velocity field is modeled as a function solely dependent on๐’™\\bm\{x\}andtt\. This single\-valued dependency prevents the model from representing intersecting trajectories with distinct velocities, often resulting in curved inference paths and increased computational costs\. To decouple the velocity from the current position, the velocity network can incorporate additional inputs like class labels\[[73](https://arxiv.org/html/2608.21070#bib.bib73),[54](https://arxiv.org/html/2608.21070#bib.bib54)\], initial positions\[[14](https://arxiv.org/html/2608.21070#bib.bib14)\], or hidden states\[[20](https://arxiv.org/html/2608.21070#bib.bib20)\]\. Additionally,\[[69](https://arxiv.org/html/2608.21070#bib.bib69)\]employs hierarchical generation to mitigate this issue\. However, these methods do not extend to second\-order dynamics\. While\[[42](https://arxiv.org/html/2608.21070#bib.bib42)\]utilized a second\-order system, it is limited to constant\-acceleration paths\. In contrast, TracingFlow offers a flexible second\-order setting\. By designing minimal\-cost paths, it effectively resolves the limitation of the single\-valued velocity field\.

#### Single\-cell Trajectory Inference and Lineage Integration\.

Optimal Transport \(OT\) is a robust framework for inferring cellular dynamics from scRNA\-seq data\[[51](https://arxiv.org/html/2608.21070#bib.bib51),[28](https://arxiv.org/html/2608.21070#bib.bib28),[57](https://arxiv.org/html/2608.21070#bib.bib57),[24](https://arxiv.org/html/2608.21070#bib.bib24),[40](https://arxiv.org/html/2608.21070#bib.bib40),[41](https://arxiv.org/html/2608.21070#bib.bib41),[2](https://arxiv.org/html/2608.21070#bib.bib2),[74](https://arxiv.org/html/2608.21070#bib.bib74),[37](https://arxiv.org/html/2608.21070#bib.bib37),[66](https://arxiv.org/html/2608.21070#bib.bib66),[12](https://arxiv.org/html/2608.21070#bib.bib12),[60](https://arxiv.org/html/2608.21070#bib.bib60),[33](https://arxiv.org/html/2608.21070#bib.bib33),[53](https://arxiv.org/html/2608.21070#bib.bib53),[29](https://arxiv.org/html/2608.21070#bib.bib29),[6](https://arxiv.org/html/2608.21070#bib.bib6),[10](https://arxiv.org/html/2608.21070#bib.bib10),[25](https://arxiv.org/html/2608.21070#bib.bib25),[70](https://arxiv.org/html/2608.21070#bib.bib70),[55](https://arxiv.org/html/2608.21070#bib.bib55),[44](https://arxiv.org/html/2608.21070#bib.bib44),[65](https://arxiv.org/html/2608.21070#bib.bib65),[52](https://arxiv.org/html/2608.21070#bib.bib52),[11](https://arxiv.org/html/2608.21070#bib.bib11)\]\. However, transcriptomic similarity alone often fails to resolve complex trajectories\. To mitigate this, studies incorporate lineage tracing priors such as clonal barcode information\. Existing approaches integrate such info via regularization\[[17](https://arxiv.org/html/2608.21070#bib.bib17),[49](https://arxiv.org/html/2608.21070#bib.bib49)\], structural alignment\[[32](https://arxiv.org/html/2608.21070#bib.bib32)\], velocity mapping\[[62](https://arxiv.org/html/2608.21070#bib.bib62)\], or sparse transition modeling\[[63](https://arxiv.org/html/2608.21070#bib.bib63),[21](https://arxiv.org/html/2608.21070#bib.bib21),[18](https://arxiv.org/html/2608.21070#bib.bib18)\]\. While improving accuracy, these generally focus on discrete couplings or static matrices\. TracingFlow advances this by embedding priors into a continuous OT\-based second\-order flow matching framework, capturing complex dynamics consistent with lineage information\.

## 3Mathematical Background

In this section, we provide the mathematical formulation of the problem addressed by TracingFlow\.

#### Dynamical Optimal Acceleration Transport \(DOAT\) Problem

Let๐’ณโŠ‚โ„d\\mathcal\{X\}\\subset\\mathbb\{R\}^\{d\}and๐’ฑโŠ‚โ„d\\mathcal\{V\}\\subset\\mathbb\{R\}^\{d\}denote the*position*and*velocity*spaces, respectively\. We define the*augmented space*as๐’ฎโ‰”๐’ณร—๐’ฑ\\mathcal\{S\}\\coloneqq\\mathcal\{X\}\\times\\mathcal\{V\}, so that the full system state is๐’”=\(๐’™,๐’—\)โˆˆ๐’ฎ\\bm\{s\}=\(\\bm\{x\},\\bm\{v\}\)\\in\\mathcal\{S\}\. Consider the second\-order dynamic system:

๐’™ห™=๐’—๐’—ห™=๐’‚โก\(๐’™,๐’—,t\),tโˆˆ\[t0,tK\]\.\\dot\{\\bm\{x\}\}=\\bm\{v\}\\quad\\dot\{\\bm\{v\}\}=\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\),\\quad t\\in\[t\_\{0\},t\_\{K\}\]\.\(1\)Modeling a large ensemble of such particles via a probability densityฯt\\rho\_\{t\}in๐’ฎ\\mathcal\{S\}, we follow\[[5](https://arxiv.org/html/2608.21070#bib.bib5),[11](https://arxiv.org/html/2608.21070#bib.bib11)\]to formulate the following optimal control problem:

min๐’‚,ฯโก๐’ฅDOATโ€‹\(๐’‚,ฯ\)=โˆซt0tKโˆซ๐’ฎ12โ€‹โ€–๐’‚โก\(๐’™,๐’—,t\)โ€–2โ€‹ฯtโ€‹\(๐’™,๐’—\)โ€‹๐‘‘๐’™โ€‹๐‘‘๐’—โ€‹๐‘‘t\\small\{\\min\_\{\\bm\{a\},\\rho\}\\mathcal\{J\}\_\{\\text\{DOAT\}\}\(\\bm\{a\},\\rho\)=\\int\_\{t\_\{0\}\}^\{t\_\{K\}\}\\int\_\{\\mathcal\{S\}\}\\frac\{1\}\{2\}\\\|\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\)\\\|^\{2\}\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\\,\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{v\}\\mathrm\{d\}t\}\(2\)s\.t\.โˆ‚tฯt\+โˆ‡๐’™โ‹…\(๐’—โ€‹ฯt\)\+โˆ‡๐’—โ‹…\(๐’‚โ€‹ฯt\)=0,\\displaystyle\\partial\_\{t\}\\rho\_\{t\}\+\\nabla\_\{\\bm\{x\}\}\\cdot\(\\bm\{v\}\\rho\_\{t\}\)\+\\nabla\_\{\\bm\{v\}\}\\cdot\(\\bm\{a\}\\rho\_\{t\}\)=0,\(3\)ฯti\(๐’™,๐’—\)=ฮผi\(๐’™,๐’—\),i=0,โ€ฆ,K\.\\displaystyle\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\)=\\mu\_\{i\}\(\\bm\{x\},\\bm\{v\}\),\\quad i=0,\\dots,K\.\(4\)Here, constraints are imposed atK\+1K\+1timestampstit\_\{i\}with given densitiesฮผi\\mu\_\{i\}\.[Equation3](https://arxiv.org/html/2608.21070#S3.E3)represents the continuity equation in the augmented space\. Analogous to standard Dynamical Optimal Transport\[[47](https://arxiv.org/html/2608.21070#bib.bib47)\], we term this the*Dynamical Optimal Acceleration Transport \(DOAT\)*problem\. To facilitate the subsequent numerical solution via particle methods, we assume the existence of an optimal flow map\.

#### Velocity\-Missed Dynamical Optimal Acceleration Transport \(VM\-DOAT\) Problem

However, in practical applications such as single\-cell datasets, only particle positions are observable, rendering velocities as unknown quantities\. This presents a discrepancy with the DOAT formulation\.

Consequently, TracingFlow addresses a relaxation of the problem:

min๐’‚,ฯโก๐’ฅVM\-DOATโ€‹\(๐’‚,ฯ\)=โˆซt0tKโˆซ๐’ฎ12โ€‹โ€–๐’‚โก\(๐’™,๐’—,t\)โ€–2โ€‹ฯtโ€‹\(๐’™,๐’—\)โ€‹๐‘‘๐’™โ€‹๐‘‘๐’—โ€‹๐‘‘t\\small\{\\min\_\{\\bm\{a\},\\rho\}\\mathcal\{J\}\_\{\\text\{VM\-DOAT\}\}\(\\bm\{a\},\\rho\)=\\int\_\{t\_\{0\}\}^\{t\_\{K\}\}\\int\_\{\\mathcal\{S\}\}\\dfrac\{1\}\{2\}\\\|\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\)\\\|^\{2\}\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\\,\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{v\}\\mathrm\{d\}t\}\(5\)s\.t\.โˆ‚tฯtโ€‹\(๐’™,๐’—\)\+โˆ‡๐’™โ‹…\(๐’—โ€‹ฯt\)\+โˆ‡๐’—โ‹…\(๐’‚โ€‹ฯt\)=0,\\displaystyle\\quad\\partial\_\{t\}\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\+\\nabla\_\{\\bm\{x\}\}\\cdot\(\\bm\{v\}\\rho\_\{t\}\)\+\\nabla\_\{\\bm\{v\}\}\\cdot\(\\bm\{a\}\\rho\_\{t\}\)=0,\(6\)โˆซฯti\(๐’™,๐’—\)d๐’—=ฮผi\(pos\)\(๐’™\),i=0,1,โ€ฆ,K\.\\displaystyle\\int\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\)\\,\\mathrm\{d\}\\bm\{v\}=\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\),\\ i=0,1,\\dots,K\.\(7\)Here, the constraints are imposed only on the marginal distribution of positions,โˆซฯtjโ€‹\(๐’™,๐’—\)โ€‹๐‘‘๐’—\\int\\rho\_\{t\_\{j\}\}\(\\bm\{x\},\\bm\{v\}\)\\,\\mathrm\{d\}\\bm\{v\}, at each time point, with no constraints on the velocity distribution\. We refer to this problem as*Velocity Missed\-Dynamical Optimal Acceleration Transport \(VM\-DOAT\)*\.

## 4Methodology of TracingFlow

In this section, we first introduce the key properties of the DOAT problem that are essential for our algorithm design\. Subsequently, drawing inspiration from Flow Matching approaches for standard Dynamical OT, we devise an algorithm to exactly solve the DOAT problem by combining Optimal Coupling with single\-particle optimal trajectories\. Finally, we introduce a global path\-based velocity completion scheme that effectively bridges the gap between VM\-DOAT and the standard DOAT framework\. By inferring missing velocities through the lens of global trajectory consistency, our approach enables the application of optimal transport theory to solve the VM\-DOAT problem in a principled manner\.

### 4\.1Key Properties of DOAT Problem

#### Optimal Single Particle Trajectory

The DOAT problem is an optimal control problem over distributions\. We begin by considering the corresponding single\-particle optimal control problem\. Specifically, the optimal trajectory is a*cubic interpolation*determined by the boundary states\.

###### Proposition 4\.1\.

Letฮณ:\[0,T\]โ†’โ„d\\gamma:\[0,T\]\\to\\mathbb\{R\}^\{d\}be a twice continuously differentiable curve satisfying the boundary conditionsฮณโก\(0\)=๐ฑ0\\gamma\(0\)=\\bm\{x\}\_\{0\},ฮณโ€ฒโ€‹\(0\)=๐ฏ0\\gamma^\{\\prime\}\(0\)=\\bm\{v\}\_\{0\},ฮณโก\(T\)=๐ฑT\\gamma\(T\)=\\bm\{x\}\_\{T\}, andฮณโ€ฒโ€‹\(T\)=๐ฏT\\gamma^\{\\prime\}\(T\)=\\bm\{v\}\_\{T\}\. The variational problem

minโกโˆซ0Tฮณโก12โ€‹โ€–ฮณโ€ฒโ€ฒโ€‹\(t\)โ€–22โ€‹๐‘‘t\\min\_\{\\gamma\}\\int\_\{0\}^\{T\}\\frac\{1\}\{2\}\\\|\\gamma^\{\\prime\\prime\}\(t\)\\\|\_\{2\}^\{2\}\\,\\mathrm\{d\}t\(8\)admits a unique minimizer, which is a cubic polynomial of the form

ฮณโก\(t\)=c3โ€‹t3\+c2โ€‹t2\+c1โ€‹t\+c0,\\gamma\(t\)=c\_\{3\}t^\{3\}\+c\_\{2\}t^\{2\}\+c\_\{1\}t\+c\_\{0\},\(9\)where the coefficients are detailed in[SectionA\.1](https://arxiv.org/html/2608.21070#A1.SS1)\.This result is got by*Euler\-Lagrange*formula\. See a detailed proof in[SectionA\.1](https://arxiv.org/html/2608.21070#A1.SS1)\.

#### Optimal Coupling

By utilizing the single\-particle optimal trajectory, we can immediately obtain the minimum cost for transporting a particle from\(๐’™0,๐’—0\)\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\)to\(๐’™T,๐’—T\)\(\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\):

###### Corollary 4\.2\.

Substitute[Equation9](https://arxiv.org/html/2608.21070#S4.E9)to[Equation8](https://arxiv.org/html/2608.21070#S4.E8), we get the minimum cost :

๐’ž\[0,T\]\\displaystyle\\mathcal\{C\}\_\{\[0,T\]\}=2T3โ€‹\{\(โ€–๐’—0โ€–2\+โŸจ๐’—0,๐’—TโŸฉ\+โ€–๐’—Tโ€–2\)โ€‹T2\+3โ€‹โŸจ๐’—0\+๐’—T,๐’™0โˆ’๐’™TโŸฉโ€‹T\+3โ€‹โ€–๐’™0โˆ’๐’™Tโ€–2\}\\displaystyle=\\frac\{2\}\{T^\{3\}\}\\bigg\\\{\(\\\|\\bm\{v\}\_\{0\}\\\|^\{2\}\+\\langle\\bm\{v\}\_\{0\},\\bm\{v\}\_\{T\}\\rangle\+\\\|\\bm\{v\}\_\{T\}\\\|^\{2\}\)T^\{2\}\+3\\langle\\bm\{v\}\_\{0\}\+\\bm\{v\}\_\{T\},\\bm\{x\}\_\{0\}\-\\bm\{x\}\_\{T\}\\rangle T\+3\\\|\\bm\{x\}\_\{0\}\-\\bm\{x\}\_\{T\}\\\|^\{2\}\\bigg\\\}\(10\)

The minimum cost allows us to define an optimal transport problem betweent=0t=0andt=Tt=T:

ฯ€\[0,T\]โˆ—=minโกโˆซฯ€\[0,T\]โก๐’ž\[0,T\]โ€‹\[\(๐’™0,๐’—0\),\(๐’™T,๐’—T\)\]โ€‹dโ€‹ฯ€\[0,T\]\\pi^\{\*\}\_\{\[0,T\]\}=\\min\_\{\\pi\_\{\[0,T\]\}\}\\int\\mathcal\{C\}\_\{\[0,T\]\}\[\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\),\(\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\)\]\\mathrm\{d\}\\pi\_\{\[0,T\]\}\(11\)whereฯ€\[0,T\]โ€‹\[\(๐’™0,๐’—0\),\(๐’™T,๐’—T\)\]\\pi\_\{\[0,T\]\}\[\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\),\(\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\)\]is any coupling satisfying the constraintsโˆซฯ€\[0,T\]โ€‹dโ€‹๐’™0โ€‹dโ€‹๐’—0=ฮผTโ€‹\(๐’™,๐’—\)\\int\\pi\_\{\[0,T\]\}\\mathrm\{d\}\\bm\{x\}\_\{0\}\\mathrm\{d\}\\bm\{v\}\_\{0\}=\\mu\_\{T\}\(\\bm\{x\},\\bm\{v\}\)andโˆซฯ€\[0,T\]โ€‹dโ€‹๐’™Tโ€‹dโ€‹๐’—T=ฮผ0โ€‹\(๐’™,๐’—\)\\int\\pi\_\{\[0,T\]\}\\mathrm\{d\}\\bm\{x\}\_\{T\}\\mathrm\{d\}\\bm\{v\}\_\{T\}=\\mu\_\{0\}\(\\bm\{x\},\\bm\{v\}\),ฮผ0โ€‹\(๐’™,๐’—\)\\mu\_\{0\}\(\\bm\{x\},\\bm\{v\}\)andฮผTโ€‹\(๐’™,๐’—\)\\mu\_\{T\}\(\\bm\{x\},\\bm\{v\}\)are two distributions in augmented space\. We term this formulation theStatic Optimal Acceleration Transport \(SOAT\)problem\.

It is worth noting that while DOAT permits crossings in the projected position space, it prohibits trajectory crossing in the augmented space, which guarentees the neatness of trajectories\. The following remark provides the mathematical justification for this non\-crossing property in the augmented space, demonstrating that any such intersection implies a strictly higher transport cost\.

[Equation2](https://arxiv.org/html/2608.21070#S3.E2)represents a multi\-marginal optimal transport problem over the interval\[t0,tK\]\[t\_\{0\},t\_\{K\}\]\. To address this, we formulate sequential SOAT problems between adjacent time stepstit\_\{i\}andti\+1t\_\{i\+1\}by applying the substitutions๐’™0โ†’๐’™i\\bm\{x\}\_\{0\}\\rightarrow\\bm\{x\}\_\{i\},๐’™Tโ†’๐’™i\+1\\bm\{x\}\_\{T\}\\rightarrow\\bm\{x\}\_\{i\+1\},๐’—0โ†’๐’—i\\bm\{v\}\_\{0\}\\rightarrow\\bm\{v\}\_\{i\},๐’—Tโ†’๐’—i\+1\\bm\{v\}\_\{T\}\\rightarrow\\bm\{v\}\_\{i\+1\}, andTโ†’ti\+1โˆ’tiT\\rightarrow t\_\{i\+1\}\-t\_\{i\}to[Equation11](https://arxiv.org/html/2608.21070#S4.E11)\. For simplicity, let๐’žiโ†’i\+1\\mathcal\{C\}\_\{i\\rightarrow i\+1\}denote the resulting pairwise cost andฯ€iโ†’i\+1โˆ—\\pi^\{\*\}\_\{i\\rightarrow i\+1\}the optimal plan\. Due to the Markovian nature of the dynamics, the full multi\-marginal plan decomposes into the sequence of these local plans\{ฯ€iโ†’i\+1โˆ—\}\\\{\{\\pi^\{\*\}\_\{i\\rightarrow i\+1\}\}\\\}fori=0,โ€ฆ,Kโˆ’1i=0,\\dots,K\-1\.

### 4\.2Conditional and Marginal Probability Path

#### Obtaining Marginal Probability Path via Mixing Conditional Probability Paths

Similar to the approach in standard Flow Matching\[[35](https://arxiv.org/html/2608.21070#bib.bib35),[58](https://arxiv.org/html/2608.21070#bib.bib58)\], directly constructing the Marginal Probability Path is difficult\. Therefore, we condition on๐’›\\bm\{z\}and define the Marginal Probability Path as a mixture of Conditional Probability Paths:

ฯtโ€‹\(๐’™,๐’—\)=โˆซฯtโ€‹\(๐’™,๐’—\|๐’›\)โ€‹qโ€‹\(๐’›\)โ€‹๐‘‘๐’›\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)=\\int\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\)q\(\\bm\{z\}\)\\mathrm\{d\}\\bm\{z\}\(12\)The condition๐’›\\bm\{z\}consists ofK\+1K\+1points selected respectively from the snapshot data atK\+1K\+1time pointst=0,1,โ‹ฏ,Kt=0,1,\\cdots,K\. Ifฯtโ€‹\(๐’™,๐’—\|๐’›\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\)is generated by a conditional acceleration field๐’‚โก\(๐’™,๐’—,t\|๐’›\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\|\\bm\{z\}\)starting from the initial conditionฯ0โ€‹\(๐’™,๐’—\|๐’›\)\\rho\_\{0\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\), then the marginal acceleration field:

๐’‚โก\(๐’™,๐’—,t\):=๐”ผqโก\(๐’›\)โ€‹\{๐’‚โก\(๐’™,๐’—,t\|๐’›\)โ€‹ฯtโ€‹\(๐’™,๐’—\|๐’›\)ฯtโ€‹\(๐’™,๐’—\)\}\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\):=\\mathbb\{E\}\_\{q\(\\bm\{z\}\)\}\\left\\\{\\dfrac\{\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\|\\bm\{z\}\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\)\}\{\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\}\\right\\\}\(13\)will generate the marginal probability pathฯtโ€‹\(๐’™,๐’—\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\.

###### Theorem 4\.4\.

The marginal acceleration field in Eq\. \([13](https://arxiv.org/html/2608.21070#S4.E13)\) generates the marginal probability path\. See the proof in[SectionA\.3](https://arxiv.org/html/2608.21070#A1.SS3)\.

#### Equivalence between Regressing Conditional Acceleration Fields and Marginal Acceleration Fields

Assuming the marginal acceleration field๐’‚โก\(๐’™,๐’—,t\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\)is known and we can sample from the marginal probability pathฯtโ€‹\(๐’™,๐’—\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\), we can directly regress๐’‚โก\(๐’™,๐’—,t\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\)using a neural network\. Let๐’‚๐œฝโ€‹\(โ‹…,โ‹…,โ‹…\):โ„dร—โ„dร—โ„1โ†’โ„d\\bm\{a\}\_\{\\bm\{\\theta\}\}\(\\cdot,\\cdot,\\cdot\):\\mathbb\{R\}^\{d\}\\times\\mathbb\{R\}^\{d\}\\times\\mathbb\{R\}^\{1\}\\rightarrow\\mathbb\{R\}^\{d\}be a time\-dependent acceleration field parameterized by a neural network with parametersฮธ\\theta\. The acceleration flow matching loss is

โ„’AFM=๐”ผ\(๐’™,๐’—\)โˆผฯtโ€‹\(๐’™,๐’—\),tโˆผ๐’ฐโก\[t0,tK\]โ€‹\{โ€–๐’‚๐œฝโˆ’๐’‚โก\(๐’™,๐’—,t\)โ€–2\}\\begin\{split\}\\mathcal\{L\}\_\{\\text\{AFM\}\}=\\mathbb\{E\}\_\{\(\\bm\{x\},\\bm\{v\}\)\\sim\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\),t\\sim\\mathcal\{U\}\[t\_\{0\},t\_\{K\}\]\}\\left\\\{\\\|\\bm\{a\}\_\{\\bm\{\\theta\}\}\-\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\)\\\|^\{2\}\\right\\\}\\end\{split\}However,๐’‚โก\(๐’™,๐’—,t\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\)is generally intractable because it is defined via an expectation \(integral\), and its denominatorฯtโ€‹\(๐’™,๐’—\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)also requires integration to compute\. In contrast, the conditional acceleration field๐’‚โก\(๐’™,๐’—,t\|z\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\|z\)and the conditional probability pathฯtโ€‹\(๐’™,๐’—\|z\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|z\)have simple forms\. Therefore, in actual training, we use the following Conditional Acceleration Flow Matching \(CAFM\) objective:

โ„’CAFM=๐”ผ\(๐’™,๐’—\)โˆผฯtโ€‹\(๐’™,๐’—\|z\),tโˆผ๐’ฐโก\[t0,tK\],zโˆผqโก\(z\)โ€‹\[โ€–๐’‚ฮธโ€‹\(๐’™,๐’—,t\)โˆ’๐’‚โก\(๐’™,๐’—,t\|z\)โ€–2\]\\begin\{split\}\\mathcal\{L\}\_\{\\text\{CAFM\}\}=\\mathbb\{E\}\_\{\\begin\{subarray\}\{c\}\(\\bm\{x\},\\bm\{v\}\)\\sim\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|z\),\\\\ t\\sim\\mathcal\{U\}\[t\_\{0\},t\_\{K\}\],z\\sim q\(z\)\\end\{subarray\}\}\\Big\[\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{x\},\\bm\{v\},t\)\-\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\|z\)\\\|^\{2\}\\Big\]\\end\{split\}These two objectives are equivalent in training, as described by the following theorem:

###### Theorem 4\.5\.

Ifฯtโ€‹\(๐ฑ,๐ฏ\)\>0\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\>0for any๐ฑโˆˆโ„d,๐ฏโˆˆโ„d,tโˆˆ\[t0,tK\]\\bm\{x\}\\in\\mathbb\{R\}^\{d\},\\bm\{v\}\\in\\mathbb\{R\}^\{d\},t\\in\[t\_\{0\},t\_\{K\}\], thenโ„’AFM\\mathcal\{L\}\_\{\\text\{AFM\}\}andโ„’CAFM\\mathcal\{L\}\_\{\\text\{CAFM\}\}differ only by a constant independent of๐›‰\\bm\{\\theta\}\. In other words,โˆ‡๐›‰โ„’AFM=โˆ‡๐›‰โ„’CAFM\\nabla\_\{\\bm\{\\theta\}\}\\mathcal\{L\}\_\{\\text\{AFM\}\}=\\nabla\_\{\\bm\{\\theta\}\}\\mathcal\{L\}\_\{\\text\{CAFM\}\}\. The proof is left to[SectionA\.4](https://arxiv.org/html/2608.21070#A1.SS4)\.

#### Design of the Conditional Probability Path

Consider the joint distributionsฮผ0โ€‹\(๐’™,๐’—\)\\mu\_\{0\}\(\\bm\{x\},\\bm\{v\}\),ฮผ1โ€‹\(๐’™,๐’—\)\\mu\_\{1\}\(\\bm\{x\},\\bm\{v\}\), โ€ฆ,ฮผKโ€‹\(๐’™,๐’—\)\\mu\_\{K\}\(\\bm\{x\},\\bm\{v\}\)atK\+1K\+1time points\. The optimal transport plans calculated via solving SOAT problem areฯ€0โ†’1โˆ—,ฯ€1โ†’2โˆ—,โ‹ฏ,ฯ€Kโˆ’1โ†’Kโˆ—\\pi^\{\*\}\_\{0\\rightarrow 1\},\\pi^\{\*\}\_\{1\\rightarrow 2\},\\cdots,\\pi^\{\*\}\_\{K\-1\\rightarrow K\}\. Taking the condition variable๐’›=\[\(๐’™0,๐’—0\),โ‹ฏ,\(๐’™K,๐’—K\)\]\\bm\{z\}=\[\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\),\\cdots,\(\\bm\{x\}\_\{K\},\\bm\{v\}\_\{K\}\)\], and similar to standard Flow Matching, we choose a path with time\-varying mean and invariant variance as the conditional probability path, i\.e\.:

ฯtโ€‹\(๐’™,๐’—\|z\)=๐’ฉโก\(๐’™\|๐’Žxโ€‹\(t\),ฯƒx2โ€‹I\)โ‹…๐’ฉโก\(๐’—\|๐’Žvโ€‹\(t\),ฯƒv2โ€‹I\)\\begin\{split\}\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|z\)=\\mathcal\{N\}\(\\bm\{x\}\|\\bm\{m\}\_\{x\}\(t\),\\sigma\_\{x\}^\{2\}I\)\\cdot\\mathcal\{N\}\(\\bm\{v\}\|\\bm\{m\}\_\{v\}\(t\),\\sigma\_\{v\}^\{2\}I\)\\end\{split\}\(14\)Here,๐’Žx,๐’Žv\\bm\{m\}\_\{x\},\\bm\{m\}\_\{v\}are given by the Minimum Cost Trajectory described in[Section4\.1](https://arxiv.org/html/2608.21070#S4.SS1.SSS0.Px1), where fortโˆˆ\[ti,ti\+1\]t\\in\[t\_\{i\},t\_\{i\+1\}\]:

๐’Žxโ€‹\(t\)=c3โ€‹\(tโˆ’ti\)3\+c2โ€‹\(tโˆ’ti\)2\+c1โ€‹\(tโˆ’ti\)\+c0๐’Žvโ€‹\(t\)โ€‹3โ€‹c3โ€‹\(tโˆ’ti\)2\+2โ€‹c2โ€‹\(tโˆ’ti\)\+c1\\displaystyle\\bm\{m\}\_\{x\}\(t\)=c\_\{3\}\(t\-t\_\{i\}\)^\{3\}\+c\_\{2\}\(t\-t\_\{i\}\)^\{2\}\+c\_\{1\}\(t\-t\_\{i\}\)\+c\_\{0\}\\quad\\bm\{m\}\_\{v\}\(t\)3c\_\{3\}\(t\-t\_\{i\}\)^\{2\}\+2c\_\{2\}\(t\-t\_\{i\}\)\+c\_\{1\}, the coefficients are identical to those given in[SectionA\.1](https://arxiv.org/html/2608.21070#A1.SS1), with๐’™0,๐’—0\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}replaced by๐’™i,๐’—i\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\},๐’™T,๐’—T\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}replaced by๐’™i\+1,๐’—i\+1\\bm\{x\}\_\{i\+1\},\\bm\{v\}\_\{i\+1\}, andTTreplaced byti\+1โˆ’tit\_\{i\+1\}\-t\_\{i\}\. Based on the optimal transport plans, we define the optimal propagator \(transition kernel\) from timetit\_\{i\}toti\+1t\_\{i\+1\}as the probability of a particle at\(๐’™i,๐’—i\)\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\)at timetit\_\{i\}being transported to\(๐’™i\+1,๐’—i\+1\)\(\\bm\{x\}\_\{i\+1\},\\bm\{v\}\_\{i\+1\}\)at timeti\+1t\_\{i\+1\}:

๐’ฆiโ†’i\+1โˆ—โ€‹\[\(๐’™i,๐’—i\),\(๐’™i\+1,๐’—i\+1\)\]=ฯ€iโ†’i\+1โˆ—โ€‹\[\(๐’™i,๐’—i\),\(๐’™i\+1,๐’—i\+1\)\]ฮผiโ€‹\(๐’™i,๐’—i\)\\mathcal\{K\}^\{\*\}\_\{i\\rightarrow i\+1\}\[\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\),\(\\bm\{x\}\_\{i\+1\},\\bm\{v\}\_\{i\+1\}\)\]=\\frac\{\\pi^\{\*\}\_\{i\\rightarrow i\+1\}\[\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\),\(\\bm\{x\}\_\{i\+1\},\\bm\{v\}\_\{i\+1\}\)\]\}\{\\mu\_\{i\}\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\)\}\(15\)
Then,

๐’›โˆผฮผ0โ€‹\(๐’™0,๐’—0\)โ‹…โˆi=0Kโˆ’1๐’ฆiโ†’i\+1โˆ—โ€‹\[\(๐’™i,๐’—i\),\(๐’™i\+1,๐’—i\+1\)\]\\bm\{z\}\\sim\\mu\_\{0\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\)\\cdot\\prod\_\{i=0\}^\{K\-1\}\\mathcal\{K\}^\{\*\}\_\{i\\rightarrow i\+1\}\[\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\),\(\\bm\{x\}\_\{i\+1\},\\bm\{v\}\_\{i\+1\}\)\]
In the following, we demonstrate that the constructed marginal probability path recovers the true distribution at every given time point\.

###### Theorem 4\.6\.

Whenฯƒxโ†’0,ฯƒvโ†’0\\sigma\_\{x\}\\rightarrow 0,\\sigma\_\{v\}\\rightarrow 0, the marginal probability density constructed by[Equation12](https://arxiv.org/html/2608.21070#S4.E12)satisfiesฯtiโ€‹\(๐ฑ,๐ฏ\)โ†’ฮผiโ€‹\(๐ฑ,๐ฏ\)\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\)\\rightarrow\\mu\_\{i\}\(\\bm\{x\},\\bm\{v\}\)at theK\+1K\+1given time points\. Therefore, the Marginal Probability Path constructed by mixing Conditional Probability Paths can correctly reconstruct the marginal distributions\. See the proof in[SectionA\.5](https://arxiv.org/html/2608.21070#A1.SS5)\.

Furthermore, we demonstrate that the design yields an exact solution to the DOAT problem\.

###### Proposition 4\.7\.

Asฯƒxโ†’0\\sigma\_\{x\}\\rightarrow 0andฯƒvโ†’0\\sigma\_\{v\}\\rightarrow 0, the marginal probability path and the marginal acceleration field will converge to the solution to the DOAT problem\. See the proof in[SectionA\.6](https://arxiv.org/html/2608.21070#A1.SS6)\.

### 4\.3Transform VM\-DOAT to DOAT

#### Relations between DOAT and VM\-DOAT

The optimal transport cost of the DOAT problem is a functional of the sequence of joint distributionsฮผiโ€‹\(๐’™,๐’—\)\\mu\_\{i\}\(\\bm\{x\},\\bm\{v\}\)fori=0,โ€ฆ,Ki=0,\\dots,K\. Formally, we denote this dependency as๐’ฅDOATโ€‹\[\{ฮผiโ€‹\(๐’™,๐’—\)\}\|i=0K\]\\mathcal\{J\}\_\{\\text\{DOAT\}\}\[\\\{\\mu\_\{i\}\(\\bm\{x\},\\bm\{v\}\)\\\}\|\_\{i=0\}^\{K\}\]\. In the VM\-DOAT setting, we only have access to the marginal spatial distributionsฮผi\(pos\)โ€‹\(๐’™\)=โˆซฮผiโ€‹\(๐’™,๐’—\)โ€‹๐‘‘๐’—\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\)=\\int\\mu\_\{i\}\(\\bm\{x\},\\bm\{v\}\)\\mathrm\{d\}\\bm\{v\}\. Consequently, the cost of VM\-DOAT, denoted as๐’ฅVM\-DOATโ€‹\[\{ฮผi\(pos\)โ€‹\(๐’™\)\}\|i=0K\]\\mathcal\{J\}\_\{\\text\{VM\-DOAT\}\}\[\\\{\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\)\\\}\|\_\{i=0\}^\{K\}\], relates to๐’ฅDOAT\\mathcal\{J\}\_\{\\text\{DOAT\}\}through the following minimization:

๐’ฅVM\-DOATโ€‹\[\{ฮผi\(pos\)โ€‹\(๐’™\)\}\]=minฯ‰iโ€‹\(๐’—\|๐’™\)โก๐’ฅDOATโ€‹\[\{ฮผi\(pos\)โ€‹\(๐’™\)โ‹…ฯ‰iโ€‹\(๐’—\|๐’™\)\}\],\\mathcal\{J\}\_\{\\text\{VM\-DOAT\}\}\[\\\{\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\)\\\}\]=\\min\_\{\\omega\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\)\}\\mathcal\{J\}\_\{\\text\{DOAT\}\}\[\\\{\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\)\\cdot\\omega\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\)\\\}\],\(16\)whereฯ‰iโ€‹\(๐’—\|๐’™\)\\omega\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\)represents the conditional distribution of velocity, such that the joint distribution factorizes asฮผiโ€‹\(๐’™,๐’—\)=ฮผi\(pos\)โ€‹\(๐’™\)โ‹…ฯ‰iโ€‹\(๐’—\|๐’™\)\\mu\_\{i\}\(\\bm\{x\},\\bm\{v\}\)=\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\)\\cdot\\omega\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\)\. Thus, our objective is to determine the optimalฯ‰iโ€‹\(๐’—\|๐’™\)\\omega\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\)that minimizes the expression above, thereby transforming the VM\-DOAT problem into a solvable DOAT instance\.

This transformation rests on two key properties: 1\) The SOAT problem can be solved efficiently using the Sinkhorn method\. Once the Transport Plan is computed, we can sample a trajectory\[\(๐’™0,๐’—0\),โ‹ฏ,\(๐’™K,๐’—K\)\]\[\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\),\\cdots,\(\\bm\{x\}\_\{K\},\\bm\{v\}\_\{K\}\)\]; 2\) Given the positions๐’™0:K\\bm\{x\}\_\{0:K\}of a trajectory, consider the following multi\-marginal optimal control problem:

minโˆซt0tK12โˆฅ๐œธโ€ฒโ€ฒ\(t\)โˆฅ2dt,s\.t\.๐œธ\(ti\)=๐’™i,i=0,โ€ฆ,K\.\\min\\int\_\{t\_\{0\}\}^\{t\_\{K\}\}\\dfrac\{1\}\{2\}\\\|\\bm\{\\gamma\}^\{\\prime\\prime\}\(t\)\\\|^\{2\}\\mathrm\{d\}t,\\quad\\text\{s\.t\. \}\\bm\{\\gamma\}\(t\_\{i\}\)=\\bm\{x\}\_\{i\},\\ i=0,\\dots,K\.\(17\)The velocities of the optimal trajectory, denoted asV=\[๐œธโ€ฒโ€‹\(t0\),โ‹ฏ,๐œธโ€ฒโ€‹\(tK\)\]V=\[\\bm\{\\gamma\}^\{\\prime\}\(t\_\{0\}\),\\cdots,\\bm\{\\gamma\}^\{\\prime\}\(t\_\{K\}\)\], can be obtained by solving a sparse linear system, as illustrated in the following proposition:

###### Proposition 4\.8\.

The velocitiesVVdefined above are the solution to the sparse linear systemVโ€‹A=BVA=B, whereAโˆˆโ„\(K\+1\)ร—\(K\+1\)A\\in\\mathbb\{R\}^\{\(K\+1\)\\times\(K\+1\)\}is a tridiagonal matrix andBโˆˆโ„dร—\(K\+1\)B\\in\\mathbb\{R\}^\{d\\times\(K\+1\)\}\. A detailed derivation and the explicit expressions forAAandBBare provided in[SectionA\.7](https://arxiv.org/html/2608.21070#A1.SS7)\.

It is worth noting that solving this linear system requires๐’ชโก\(K\)\\mathcal\{O\}\(K\)time, ensuring high computational efficiency\.

#### Iterative Scheme to Perform the Transformation

Based on this decomposition, we employ an iterative scheme to estimate the conditional velocity distributions acrosst0,โ€ฆ,tKt\_\{0\},\\dots,t\_\{K\}\. We initializeฯ‰i\(0\)โ€‹\(๐’—\|๐’™\)=ฮดโก\(๐’—โˆ’๐ŸŽ\)\\omega\_\{i\}^\{\(0\)\}\(\\bm\{v\}\|\\bm\{x\}\)=\\delta\(\\bm\{v\}\-\\bm\{0\}\)and subsequently alternate between the following two steps:

- โ€ข*Solve SOAT Plans*: Compute theKKoptimal SOAT plans between the distributionsฮผi\(pos\)โ€‹\(๐’™\)โ‹…ฯ‰i\(m\)โ€‹\(๐’—\|๐’™\)\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\)\\cdot\\omega^\{\(m\)\}\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\)for adjacent time points\. This yields the plansฯ€iโ†’i\+1โˆ—\(m\)โ€‹\[\(๐’™i,๐’—i\),\(๐’™i\+1,๐’—i\+1\)\]\\pi^\{\*\(m\)\}\_\{i\\rightarrow i\+1\}\[\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\),\(\\bm\{x\}\_\{i\+1\},\\bm\{v\}\_\{i\+1\}\)\]fori=0,โ€ฆ,Kโˆ’1i=0,\\dots,K\-1, wheremmdenotes the iteration index\. This step effectively provides the multi\-marginal optimal transport plan given the currentฯ‰i\(m\)โ€‹\(๐’—\|๐’™\)\\omega^\{\(m\)\}\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\)\.
- โ€ข*Update Velocity Distributions*: Update the conditional velocity distributions based on the computed SOAT plans: ฯ‰i\(m\+1\)โ€‹\(๐’—\|๐’™i\)=โˆซq\(m\)\(๐’›\)ฮด\(๐’—^i\[๐’™0โ‹ฏ๐’™K\]โˆ’๐’—\)d๐’›โˆ–iโˆซq\(m\)โ€‹\(๐’›\)โ€‹dโ€‹๐’›โˆ–i,\\omega^\{\(m\+1\)\}\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\_\{i\}\)=\\frac\{\\int q^\{\(m\)\}\(\\bm\{z\}\)\\delta\(\\bm\{\\hat\{v\}\}\_\{i\}\[\\bm\{x\}\_\{0\}\\cdots\\bm\{x\}\_\{K\}\]\-\\bm\{v\}\)\\,\\mathrm\{d\}\\bm\{z\}\_\{\\setminus i\}\}\{\\int q^\{\(m\)\}\(\\bm\{z\}\)\\,\\mathrm\{d\}\\bm\{z\}\_\{\\setminus i\}\},\(18\)wheredโ€‹๐’›โˆ–i\\mathrm\{d\}\\bm\{z\}\_\{\\setminus i\}denotes the volume elementdโ€‹๐’™0โ€‹dโ€‹๐’—0โ€‹โ€ฆโ€‹dโ€‹๐’™Kโ€‹dโ€‹๐’—K\\mathrm\{d\}\\bm\{x\}\_\{0\}\\mathrm\{d\}\\bm\{v\}\_\{0\}\\dots\\mathrm\{d\}\\bm\{x\}\_\{K\}\\mathrm\{d\}\\bm\{v\}\_\{K\}excludingdโ€‹๐’™i\\mathrm\{d\}\\bm\{x\}\_\{i\}\(i\.e\., integrating over all variables except๐’™i\\bm\{x\}\_\{i\}\)\. The term๐’—^iโ€‹\[๐’™0,โ€ฆ,๐’™K\]\\bm\{\\hat\{v\}\}\_\{i\}\[\\bm\{x\}\_\{0\},\\dots,\\bm\{x\}\_\{K\}\]denotes the optimal velocity at timetit\_\{i\}derived from the sequence๐’™0,โ€ฆ,๐’™K\\bm\{x\}\_\{0\},\\dots,\\bm\{x\}\_\{K\}via[Section4\.3](https://arxiv.org/html/2608.21070#S4.SS3.SSS0.Px1)\. Intuitively,ฯ‰i\(m\+1\)โ€‹\(๐’—\|๐’™i\)\\omega^\{\(m\+1\)\}\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\_\{i\}\)is set to the distribution induced by the velocities of all optimal control trajectories passing through๐’™i\\bm\{x\}\_\{i\}at timetit\_\{i\}\. Since the total transport cost is the sum of costs along individual trajectories, aligning๐’—\\bm\{v\}with the optimal velocity of each trajectory minimizes the global objective\.

In the update step above, the trajectory variable๐’›=\[\(๐’™0,๐’—0\),โ€ฆ,\(๐’™K,๐’—K\)\]\\bm\{z\}=\[\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\),\\dots,\(\\bm\{x\}\_\{K\},\\bm\{v\}\_\{K\}\)\]follows the joint distributionq\(m\)โ€‹\(๐’›\)q^\{\(m\)\}\(\\bm\{z\}\)defined as:

q\(m\)โ€‹\(๐’›\)=ฮผ0\(pos\)โ€‹\(๐’™0\)โ‹…ฯ‰0\(m\)โ€‹\(๐’—0\|๐’™0\)ร—โˆi=0Kโˆ’1๐’ฆiโ†’i\+1โˆ—\(m\)โ€‹\[\(๐’™i,๐’—i\),\(๐’™i\+1,๐’—i\+1\)\],\\displaystyle q^\{\(m\)\}\(\\bm\{z\}\)=\\ \\mu\_\{0\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\_\{0\}\)\\cdot\\omega\_\{0\}^\{\(m\)\}\(\\bm\{v\}\_\{0\}\|\\bm\{x\}\_\{0\}\)\\times\\prod\_\{i=0\}^\{K\-1\}\\mathcal\{K\}^\{\*\(m\)\}\_\{i\\rightarrow i\+1\}\[\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\),\(\\bm\{x\}\_\{i\+1\},\\bm\{v\}\_\{i\+1\}\)\],where the propagator๐’ฆiโ†’i\+1โˆ—\(m\)\\mathcal\{K\}\_\{i\\rightarrow i\+1\}^\{\*\(m\)\}is given by:

๐’ฆiโ†’i\+1โˆ—\(m\)โ€‹\[\(๐’™i,๐’—i\),\(๐’™i\+1,๐’—i\+1\)\]=ฯ€iโ†’i\+1โˆ—\(m\)โ€‹\[\(๐’™i,๐’—i\),\(๐’™i\+1,๐’—i\+1\)\]ฮผi\(pos\)โ€‹\(๐’™i\)โ‹…ฯ‰i\(m\)โ€‹\(๐’—i\|๐’™i\)\.\\displaystyle\\mathcal\{K\}^\{\*\(m\)\}\_\{i\\rightarrow i\+1\}\[\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\),\(\\bm\{x\}\_\{i\+1\},\\bm\{v\}\_\{i\+1\}\)\]=\\frac\{\\pi^\{\*\(m\)\}\_\{i\\rightarrow i\+1\}\[\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\),\(\\bm\{x\}\_\{i\+1\},\\bm\{v\}\_\{i\+1\}\)\]\}\{\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\_\{i\}\)\\cdot\\omega\_\{i\}^\{\(m\)\}\(\\bm\{v\}\_\{i\}\|\\bm\{x\}\_\{i\}\)\}\.
During the iteration process, the objectiveโ„’DOATโ€‹\[ฮผi\(pos\)โ€‹\(๐’™\)โ‹…ฯ‰i\(m\)โ€‹\(๐’—\|๐’™\)\]\\mathcal\{L\}\_\{\\text\{DOAT\}\}\[\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\)\\cdot\\omega\_\{i\}^\{\(m\)\}\(\\bm\{v\}\|\\bm\{x\}\)\]decreases monotonically\. Since the cost is bounded below by00, the iterative algorithm is guaranteed to converge\.

## 5TracingFlow Algorithm for Trajectory Inference

In the practical implementation of TracingFlow, we first perform an approximation of the conditional velocity distribution to reduce the VM\-DOAT problem to a DOAT problem\. Subsequently, we employ flow matching to solve for the acceleration field of the DOAT problem\. Notably, for lineage tracing data, we incorporate biological priors into this framework\.

### 5\.1Approximation of Conditional Velocity Distribution

Determining the full conditional velocity distributionฯ‰iโ€‹\(๐’—\|๐’™\)\\omega\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\)to transform VM\-DOAT to DOAT \([Section4\.3](https://arxiv.org/html/2608.21070#S4.SS3)\) is computationally prohibitive, potentially requiring an auxiliary generative model\. We therefore adopt a deterministic approximation, assigning a single velocity to each data point\. Consequently,[Equation18](https://arxiv.org/html/2608.21070#S4.E18)is modified to:

ฯ‰i\(m\+1\)โ€‹\(๐’—\|๐’™i\)\\displaystyle\\omega^\{\(m\+1\)\}\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\_\{i\}\)=ฮดโก\(๐’—โˆ’๐’—exp\(m\+1\)โ€‹\(๐’™i\)\),๐’—exp\(m\+1\)โ€‹\(๐’™i\)=โˆซq\(m\)\(๐’›\)๐’—^i\[๐’™0โ‹ฏ๐’™K\]d๐’›โˆ–iโˆซq\(m\)โ€‹\(๐’›\)โ€‹dโ€‹๐’›โˆ–i\.\\displaystyle=\\delta\(\\bm\{v\}\-\\bm\{v\}^\{\(m\+1\)\}\_\{\\text\{exp\}\}\(\\bm\{x\}\_\{i\}\)\),\\ \\ \\bm\{v\}\_\{\\text\{exp\}\}^\{\(m\+1\)\}\(\\bm\{x\}\_\{i\}\)=\\frac\{\\int q^\{\(m\)\}\(\\bm\{z\}\)\\bm\{\\hat\{v\}\}\_\{i\}\[\\bm\{x\}\_\{0\}\\cdots\\bm\{x\}\_\{K\}\]\\,\\mathrm\{d\}\\bm\{z\}\_\{\\setminus i\}\}\{\\int q^\{\(m\)\}\(\\bm\{z\}\)\\,\\mathrm\{d\}\\bm\{z\}\_\{\\setminus i\}\}\.\(19\)Intuitively, this assigns๐’—exp\(m\+1\)โ€‹\(๐’™i\)\\bm\{v\}^\{\(m\+1\)\}\_\{\\text\{exp\}\}\(\\bm\{x\}\_\{i\}\)as the mean velocity of all optimal control trajectories passing through๐’™i\\bm\{x\}\_\{i\}at timetit\_\{i\}\.

### 5\.2MiniBatch\-OT

Solving SOAT repeatedly is computationally intensive for large\-scale datasets\. However, since the conditional velocity distributionฯ‰iโ€‹\(๐’—\|๐’™\)\\omega\_\{i\}\(\\bm\{v\}\|\\bm\{x\}\)is approximated as a Dirac delta, the joint empirical distributionฮผiโ€‹\(๐’™,๐’—\)\\mu\_\{i\}\(\\bm\{x\},\\bm\{v\}\)effectively becomes a superposition of Dirac deltas\. This allows us to employ a minibatch strategy: we partitionฮผi\\mu\_\{i\}andฮผi\+1\\mu\_\{i\+1\}intoBBcorresponding minibatches\{ฮผi\(n\)\}n=1B\\\{\\mu\_\{i\}^\{\(n\)\}\\\}\_\{n=1\}^\{B\}and\{ฮผi\+1\(n\)\}n=1B\\\{\\mu\_\{i\+1\}^\{\(n\)\}\\\}\_\{n=1\}^\{B\}\. By solving the local plansฯ€iโ†’i\+1โˆ—\(n\)\\pi\_\{i\\rightarrow i\+1\}^\{\*\(n\)\}independently and aggregating them asฯ€โˆ—iโ†’i\+1=โŠ•n=1Bฯ€iโ†’i\+1โˆ—\(n\)\\pi^\{\*\}\_\{i\\rightarrow i\+1\}=\\oplus\_\{n=1\}^\{B\}\\pi\_\{i\\rightarrow i\+1\}^\{\*\(n\)\}, we significantly enhance computational efficiency\.

### 5\.3Incorporation of Biological Priors

In lineage tracing applications, let indiceslil\_\{i\}andli\+1l\_\{i\+1\}denote individual samples at time pointstit\_\{i\}andti\+1t\_\{i\+1\}, respectively\. Each data point\(๐’™i,li,๐’—i,li\)\(\\bm\{x\}\_\{i,l\_\{i\}\},\\bm\{v\}\_\{i,l\_\{i\}\}\)is associated with a barcodebi,liโˆˆโ„•\+b\_\{i,l\_\{i\}\}\\in\\mathbb\{N\}^\{\+\}\. We incorporate this prior by modifying the cost matrix๐’žiโ†’i\+1\\mathcal\{C\}\_\{i\\rightarrow i\+1\}\. Specifically, when computing the cost between thell\-th sample attit\_\{i\}and thekk\-th sample atti\+1t\_\{i\+1\}, we compare their barcodes: ifbi,liโ‰ bi\+1,li\+1b\_\{i,l\_\{i\}\}\\neq b\_\{i\+1,l\_\{i\+1\}\}, the standard cost is multiplied by a penalty factorp0โ€‹\(p0\>1\)p\_\{0\}\(p\_\{0\}\>1\)\. This soft constraint discourages transitions between distinct lineages, effectively embedding biological knowledge into the dynamics learning process\.

### 5\.4Algorithm of TracingFlow

The TracingFlow algorithm comprises three main steps: First, we apply the iterative approximation from[Section5\.1](https://arxiv.org/html/2608.21070#S5.SS1)to determine initial velocities, effectively transforming the VM\-DOAT problem into a DOAT problem\. Second, we compute the Optimal Transport Plan for SOAT as defined in[Equation11](https://arxiv.org/html/2608.21070#S4.E11)\. Third, we sample conditional probability paths to train the acceleration network๐’‚๐œฝโ€‹\(๐’™,๐’—,t\)\\bm\{a\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},\\bm\{v\},t\)via regression\. Additionally, to provide initial conditions for inference, we parameterize the unique initial velocities derived in the first step using a neural network๐’—0,๐ƒโ€‹\(๐’™\)\\bm\{v\}\_\{0,\\bm\{\\xi\}\}\(\\bm\{x\}\), trained similarly via regression\. The algorithmโ€™s pseudocode is provided in[AppendixD](https://arxiv.org/html/2608.21070#A4)\.

Theoretically, TracingFlow recovers the exact distribution if the acceleration and initial velocity are perfectly fitted\. In practice, where losses are non\-zero, we prove that the positional๐’ฒ2\\mathcal\{W\}\_\{2\}distance between the generated and true distributions is bounded by the sum of the velocity and AFM losses, and the relevant theorem is detailed in[TheoremA\.1](https://arxiv.org/html/2608.21070#A1.Thmtheorem1)\.

## 6Experiment Results

To evaluate the effectiveness of our algorithm, TracingFlow, we conducted three categories of experiments: 1\) verifying whether the acceleration field of TracingFlow can transport the initial data distributionฮผ0\(pos\)โ€‹\(๐’™\)=โˆซฯt0โ€‹\(๐’™,๐’—\)โ€‹๐‘‘๐’—\\mu\_\{0\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\)=\\int\\rho\_\{t\_\{0\}\}\(\\bm\{x\},\\bm\{v\}\)\\mathrm\{d\}\\bm\{v\}to the data distributionฮผi\(pos\)โ€‹\(๐’™\)\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\)at other given timestit\_\{i\}; 2\) verifying whether TracingFlow can effectively perform distribution interpolation and extrapolation; and 3\) verifying whether TracingFlow can more effectively infer the dynamical laws in Lineage Tracing data given biological prior knowledge\.

#### Distribution Transport

We evaluated TracingFlow on a 2\-dimensional synthetic dataset and real single\-cell omics datasets of varying dimensions\. The evaluation metrics were the 1\-Wasserstein \(๐’ฒ1\\mathcal\{W\}\_\{1\}\) and 2\-Wasserstein \(๐’ฒ2\\mathcal\{W\}\_\{2\}\) distances, measuring the discrepancy between the fitted and ground truth distributions\. The second\-order dynamics model of TracingFlow provides stronger expressive power than first\-order algorithms, yielding lower average๐’ฒ1\\mathcal\{W\}\_\{1\}and๐’ฒ2\\mathcal\{W\}\_\{2\}values across all datasets \([Table1](https://arxiv.org/html/2608.21070#S6.T1)\)\. We visualized the evolutionary trajectories learned by OT\-CFM and TracingFlow on the synthetic dataset in[Figure2](https://arxiv.org/html/2608.21070#S6.F2)\. TracingFlowโ€™s second\-order model allows it to learn trajectories that cross in the position space๐’™1,๐’™2\\bm\{x\}\_\{1\},\\bm\{x\}\_\{2\}, whereas OT\-CFM is unable to do so\. In[SectionC\.1](https://arxiv.org/html/2608.21070#A3.SS1), we present the distribution reconstruction accuracy at each time point, as well as experimental results on additional datasets \(such as the Mexico Gulf dataset\)\.

Table 1:Average๐’ฒ1\\mathcal\{W\}\_\{1\}and๐’ฒ2\\mathcal\{W\}\_\{2\}distances between the generated and ground truth distributions at various time points on the 2D Simulation , Cite 5D, and Cite 100D datasets for different algorithms\.![Refer to caption](https://arxiv.org/html/2608.21070v1/SimLineage.png)Figure 2:On the Simulation 2D dataset: a\) non\-crossing paths learned by OT\-CFM, and b\) crossing paths learned by TracingFlow\.
#### Interpolation and Extrapolation

To verify TracingFlowโ€™s interpolation and extrapolation capabilities, we conducted a Hold\-One\-Out experiment on the 5\-dimensional EB dataset \(time pointst=0,โ€ฆ,4t=0,\\dots,4\)\. In each iteration, we withheld one time point from the latter four for training and calculated the๐’ฒ1\\mathcal\{W\}\_\{1\}and๐’ฒ2\\mathcal\{W\}\_\{2\}distances between the interpolated and true distributions\. As shown in[Table3](https://arxiv.org/html/2608.21070#S6.T3), TracingFlow achieved the highest accuracy, demonstrating that models based on second\-order dynamics are superior in capturing high\-curvature and non\-linear trajectories\.

#### Lineage Tracing Data

To verify whether TracingFlow better handles lineage tracing data given biological priors, we experimented on a 3D Simulation Lineage Dataset and the Hematopoiesis Dataset\. Following[Section5\.3](https://arxiv.org/html/2608.21070#S5.SS3), we introduced biological priors into TracingFlow by modifying the transport cost matrix, other baselines also incorporates priors via conditional velocity fields\. We used lineage\-weighted๐’ฒ1\\mathcal\{W\}\_\{1\}and๐’ฒ2\\mathcal\{W\}\_\{2\}distances to evaluate consistency with biological priors while learning dynamics\. Results in[Table3](https://arxiv.org/html/2608.21070#S6.T3)show that incorporating priors enables TracingFlow to outperform other algorithms\. In[Figure4](https://arxiv.org/html/2608.21070#S6.F4), we visualized dynamics on the simulation dataset to demonstrate how TracingFlow learns trajectories conforming to priors\. In[Figure4](https://arxiv.org/html/2608.21070#S6.F4), we plotted the PCA\-reduced positions of two barcodes from the Hematopoiesis Dataset att=1t=1; the positions inferred by TracingFlow are significantly closer to the ground truth than those by OT\-CFM\.

Table 2:Average๐’ฒ1\\mathcal\{W\}\_\{1\}and๐’ฒ2\\mathcal\{W\}\_\{2\}distances between the predicted and ground truth distributions at held\-out time points on the EB 5D dataset for different algorithms\.
Table 3:Average lineage\-weighted๐’ฒ1\\mathcal\{W\}\_\{1\}and๐’ฒ2\\mathcal\{W\}\_\{2\}distances between the generated and ground truth distributions at various time points on the 3D Simulation Lineage and Hematopoiesis datasets for different algorithms\.

![Refer to caption](https://arxiv.org/html/2608.21070v1/SimLineage3D.png)Figure 3:On the Sim\-Lineage dataset: a\) paths learned by OT\-CFM without considering biological priors, and b\) paths learned by TracingFlow that preserve biological priors\.
![Refer to caption](https://arxiv.org/html/2608.21070v1/RealLineage.png)Figure 4:Comparison of generated \(circles\) and ground truth \(โ€™xโ€™\) positions att=1t=1for two barcodes on the hematopoiesis dataset\[[64](https://arxiv.org/html/2608.21070#bib.bib64)\]\. \(a\) OT\-CFM\. \(b\) TracingFlow\. Visualized via 2D PCA\.

## 7Conclusion and Limitation

In this work, we introduce TracingFlow, a Flow Matching framework for the Dynamical Acceleration Optimal Transport \(DOAT\) problem\. By pre\-solving SOAT to learn acceleration and initial velocity, it enables efficient, simulation\-free solutions for large\-scale DOAT problem\. Experiments on simulated and real world data demonstrate that second\-order dynamics enhance expressivity, yielding precise temporal distribution recovery\. Additionally, incorporating biological priors better preserves intercellular lineages\.

A limitation is the computationally intensive SOAT pre\-calculation, though this is mitigable via minibatch\-OT\. Furthermore, explicitly modeling the theoretical conditional velocity distribution per data point is prohibitive, so we approximate it using its expectation\. The objectiveโˆซ0T12โ€‹โ€–๐’‚โ€–2โ€‹๐‘‘t\\int\_\{0\}^\{T\}\\frac\{1\}\{2\}\\\|\\bm\{a\}\\\|^\{2\}\\mathrm\{d\}talso lacks a clear physical interpretation, as standard classical actions exclude second time derivatives \(see[SectionE\.2](https://arxiv.org/html/2608.21070#A5.SS2)\)\. Despite current sequencing measurements lacking velocity data, priors suggest biological systems may follow higher\-order dynamics due to the existence of complex regulations\. Future work will extend TracingFlow to further model these dynamics\.

## References

- Albergo and Vanden\-Eijnden \[2022\]Michael S Albergo and Eric Vanden\-Eijnden\.Building normalizing flows with stochastic interpolants\.*arXiv preprint arXiv:2209\.15571*, 2022\.
- Albergo et al\. \[2023\]Michael S Albergo, Nicholas M Boffi, and Eric Vanden\-Eijnden\.Stochastic interpolants: A unifying framework for flows and diffusions\.*arXiv preprint arXiv:2303\.08797*, 2023\.
- Arnolโ€™d \[2013\]Vladimir Igorevich Arnolโ€™d\.*Mathematical methods of classical mechanics*, volume 60\.Springer Science & Business Media, 2013\.
- Atanackovic et al\. \[2025\]Lazar Atanackovic, Xi Zhang, Brandon Amos, Mathieu Blanchette, Leo J Lee, Yoshua Bengio, Alexander Tong, and Kirill Neklyudov\.Meta flow matching: Integrating vector fields on the wasserstein manifold\.In*The Thirteenth International Conference on Learning Representations*, 2025\.
- Benamou et al\. \[2019\]Jean\-David Benamou, Thomas O Gallouรซt, and Franรงois\-Xavier Vialard\.Second\-order models for optimal transport and cubic splines on the wasserstein space\.*Foundations of Computational Mathematics*, 19\(5\):1113โ€“1143, 2019\.
- Bunne et al\. \[2023\]Charlotte Bunne, Ya\-Ping Hsieh, Marco Cuturi, and Andreas Krause\.The schrรถdinger bridge between gaussian measures has a closed form\.In*International Conference on Artificial Intelligence and Statistics*, pages 5802โ€“5833\. PMLR, 2023\.
- Bunne et al\. \[2024\]Charlotte Bunne, Geoffrey Schiebinger, Andreas Krause, Aviv Regev, and Marco Cuturi\.Optimal transport for single\-cell and spatial omics\.*Nature Reviews Methods Primers*, 4\(1\):58, 2024\.
- Cao et al\. \[2025\]Zihan Cao, Yu Zhong, and Liang\-Jian Deng\.Taming flow matching with unbalanced optimal transport into fast pansharpening\.*arXiv preprint arXiv:2503\.14975*, 2025\.
- Chen et al\. \[2018\]Ricky TQ Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud\.Neural ordinary differential equations\.*Advances in neural information processing systems*, 31, 2018\.
- Chen et al\. \[2022\]Tianrong Chen, Guan\-Horng Liu, and Evangelos Theodorou\.Likelihood training of schrรถdinger bridge using forward\-backward SDEs theory\.In*International Conference on Learning Representations*, 2022\.
- Chen et al\. \[2019\]Yongxin Chen, Giovanni Conforti, Tryphon T Georgiou, and Luigia Ripani\.Multi\-marginal schrรถdinger bridges\.In*International Conference on Geometric Science of Information*, pages 725โ€“732\. Springer, 2019\.
- Chizat et al\. \[2022\]Lรฉnaรฏc Chizat, Stephen Zhang, Matthieu Heitz, and Geoffrey Schiebinger\.Trajectory inference via mean\-field langevin in path space\.*Advances in Neural Information Processing Systems*, 35:16731โ€“16742, 2022\.
- Corso et al\. \[2025\]Gabriele Corso, Vignesh Ram Somnath, Noah Getz, Regina Barzilay, Tommi Jaakkola, and Andreas Krause\.Composing unbalanced flows for flexible docking and relaxation\.In*The Thirteenth International Conference on Learning Representations*, 2025\.
- De Bortoli et al\. \[2023\]Valentin De Bortoli, Guan\-Horng Liu, Tianrong Chen, Evangelos A Theodorou, and Weilie Nie\.Augmented bridge matching\.*arXiv preprint arXiv:2311\.06978*, 2023\.
- Eyring et al\. \[2024\]Luca Eyring, Dominik Klein, Thรฉo Uscidda, Giovanni Palla, Niki Kilbertus, Zeynep Akata, and Fabian J Theis\.Unbalancedness in neural Monge maps improves unpaired domain translation\.In*The Twelfth International Conference on Learning Representations*, 2024\.
- Flamary et al\. \[2021\]Rรฉmi Flamary, Nicolas Courty, Alexandre Gramfort, Mokhtar Z Alaya, Aurรฉlie Boisbunon, Stanislas Chambon, Laetitia Chapel, Adrien Corenflos, Kilian Fatras, Nemo Fournier, et al\.Pot: Python optimal transport\.*Journal of Machine Learning Research*, 22\(78\):1โ€“8, 2021\.
- Forrow and Schiebinger \[2021\]Aden Forrow and Geoffrey Schiebinger\.Lineageot is a unified framework for lineage tracing and trajectory inference\.*Nature Communications*, 12\(1\), 2021\.
- Gao et al\. \[2025\]Mingze Gao, Melania Barile, Shirom Chabra, Myriam Haltalli, Emily F\. Calderbank, Yiming Chao, Weizhong Zheng, Nicola K\. Wilson, Elisa Laurenti, Berthold Gรถttgens, and Yuanhua Huang\.CLADES: a hybrid neuralode\-gillespie approach for unveiling clonal cell fate and differentiation dynamics\.*Nature Communications*, 16\(1\), 2025\.
- Gorin et al\. \[2020\]Gennady Gorin, Valentine Svensson, and Lior Pachter\.Protein velocity and acceleration from single\-cell multiomics experiments\.*Genome biology*, 21\(1\):39, 2020\.
- Guo and Schwing \[2025\]Pengsheng Guo and Alexander G Schwing\.Variational rectified flow matching\.*arXiv preprint arXiv:2502\.09616*, 2025\.
- Guo et al\. \[2025\]Wenbo Guo, Zeyu Chen, Xinqi Li, Jingmin Huang, Qifan Hu, and Jin Gu\.sctrace\+: Enhancing cell fate inference by integrating the lineage\-tracing and multi\-faceted transcriptomic similarity information\.*Cell Systems*, 16\(9\):101398, 2025\.
- Halmos et al\. \[2025\]Peter Halmos, Xinhao Liu, Julian Gold, Feng Chen, Li Ding, and Benjamin J Raphael\.Dest\-ot: Alignment of spatiotemporal transcriptomics data\.*Cell Systems*, 16\(2\), 2025\.
- Hong et al\. \[2025\]Ari Hong, Sangseon Lee, and Kwangsoo Kim\.Multi\-omic relay velocity modeling uncovers dynamic chromatin\-transcription regulation across cell states\.*Nature Communications*, 2025\.
- Huguet et al\. \[2022\]Guillaume Huguet, Daniel Sumner Magruder, Alexander Tong, Oluwadamilola Fasina, Manik Kuchroo, Guy Wolf, and Smita Krishnaswamy\.Manifold interpolating optimal\-transport flows for trajectory inference\.*Advances in neural information processing systems*, 35:29705โ€“29718, 2022\.
- Jiang and Wan \[2024\]Qi Jiang and Lin Wan\.A physics\-informed neural SDE network for learning cellular dynamics from time\-series scRNA\-seq data\.*Bioinformatics*, 40:ii120โ€“ii127, 09 2024\.ISSN 1367\-4811\.
- Kapusniak et al\. \[2024\]Kacper Kapusniak, Peter Potaptchik, Teodora Reu, Leo Zhang, Alexander Tong, Michael Bronstein, Joey Bose, and Francesco Di Giovanni\.Metric flow matching for smooth interpolations on the data manifold\.*Advances in Neural Information Processing Systems*, 37:135011โ€“135042, 2024\.
- Klein et al\. \[2024\]Dominik Klein, Thรฉo Uscidda, Fabian Theis, and Marco Cuturi\.Genot: Entropic \(gromov\) wasserstein flow matching with applications to single\-cell genomics\.*Advances in Neural Information Processing Systems*, 37:103897โ€“103944, 2024\.
- Klein et al\. \[2025\]Dominik Klein, Giovanni Palla, Marius Lange, Michal Klein, Zoe Piran, Manuel Gander, Laetitia Meng\-Papaxanthos, Michael Sterr, Lama Saber, Changying Jing, et al\.Mapping cells through time and space with moscot\.*Nature*, pages 1โ€“11, 2025\.
- Koshizuka and Sato \[2023\]Takeshi Koshizuka and Issei Sato\.Neural lagrangian schrรถdinger bridge: Diffusion modeling for population dynamics\.In*The Eleventh International Conference on Learning Representations*, 2023\.
- Kretzschmar and Watt \[2012\]Kai Kretzschmar and Fiona M\. Watt\.Lineage tracing\.*Cell*, 148\(1โ€“2\):33โ€“45, 2012\.
- Lance et al\. \[2022\]Christopher Lance, Malte D Luecken, Daniel B Burkhardt, Robrecht Cannoodt, Pia Rautenstrauch, Anna Laddach, Aidyn Ubingazhibov, Zhi\-Jie Cao, Kaiwen Deng, Sumeer Khan, et al\.Multimodal single cell data integration challenge: results and lessons learned\.*BioRxiv*, pages 2022โ€“04, 2022\.
- Lange et al\. \[2024\]Marius Lange, Zoe Piran, Michal Klein, Bastiaan Spanjaard, Dominik Klein, Jan Philipp Junker, Fabian J\. Theis, and Mor Nitzan\.Mapping lineage\-traced cells across time points with moslin\.*Genome Biology*, 25\(1\), 2024\.
- Lavenant et al\. \[2024\]Hugo Lavenant, Stephen Zhang, Young\-Heon Kim, Geoffrey Schiebinger, et al\.Toward a mathematical theory of trajectory inference\.*The Annals of Applied Probability*, 34\(1A\):428โ€“500, 2024\.
- Li et al\. \[2023\]Chen Li, Maria C Virgilio, Kathleen L Collins, and Joshua D Welch\.Multi\-omic single\-cell velocity models epigenomeโ€“transcriptome interactions and improves cell fate prediction\.*Nature biotechnology*, 41\(3\):387โ€“398, 2023\.
- Lipman et al\. \[2023\]Yaron Lipman, Ricky T\. Q\. Chen, Heli Ben\-Hamu, Maximilian Nickel, and Matthew Le\.Flow matching for generative modeling\.In*The Eleventh International Conference on Learning Representations*, 2023\.
- Liu et al\. \[2022\]Xingchao Liu, Chengyue Gong, and Qiang Liu\.Flow straight and fast: Learning to generate and transfer data with rectified flow\.*arXiv preprint arXiv:2209\.03003*, 2022\.
- Maddu et al\. \[2024\]Suryanarayana Maddu, Victor Chardรจs, Michael Shelley, et al\.Inferring biological processes with intrinsic noise from cross\-sectional data\.*arXiv preprint arXiv:2410\.07501*, 2024\.
- Mao et al\. \[2025\]Shanjun Mao, Chenyang Zhang, Runjiu Chen, Shan Tang, Xiaodan Fan, and Jie Hu\.Cell lineage tracing: Methods, applications, and challenges\.*Quantitative Biology*, 13\(4\), 2025\.
- Moon et al\. \[2019\]Kevin R Moon, David Van Dijk, Zheng Wang, Scott Gigante, Daniel B Burkhardt, William S Chen, Kristina Yim, Antonia van den Elzen, Matthew J Hirn, Ronald R Coifman, et al\.Visualizing structure and transitions in high\-dimensional biological data\.*Nature biotechnology*, 37\(12\):1482โ€“1492, 2019\.
- Neklyudov et al\. \[2023\]Kirill Neklyudov, Rob Brekelmans, Daniel Severo, and Alireza Makhzani\.Action matching: Learning stochastic dynamics from samples\.In*International conference on machine learning*, pages 25858โ€“25889\. PMLR, 2023\.
- Neklyudov et al\. \[2024\]Kirill Neklyudov, Rob Brekelmans, Alexander Tong, Lazar Atanackovic, Qiang Liu, and Alireza Makhzani\.A computational framework for solving wasserstein lagrangian flows\.In*Forty\-first International Conference on Machine Learning*, 2024\.
- Park et al\. \[2024\]Dogyun Park, Sojin Lee, Sihyeon Kim, Taehoon Lee, Youngjoon Hong, and Hyunwoo J Kim\.Constant acceleration flow\.*Advances in Neural Information Processing Systems*, 37:90030โ€“90060, 2024\.
- Paszke et al\. \[2017\]Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer\.Automatic differentiation in pytorch\.2017\.
- Peng et al\. \[2024\]Qiangwei Peng, Peijie Zhou, and Tiejun Li\.stVCR: Spatiotemporal dynamics of single cells\.*bioRxiv*, pages 2024โ€“06, 2024\.
- Peng et al\. \[2026\]Qiangwei Peng, Zihan Wang, Junda Ying, Yuhao Sun, Qing Nie, Lei Zhang, Tiejun Li, and Peijie Zhou\.WFR\-FM: Simulation\-free dynamic unbalanced optimal transport\.*arXiv preprint arXiv:2601\.06810*, 2026\.
- Petroviฤ‡ et al\. \[2025\]Katarina Petroviฤ‡, Lazar Atanackovic, Kacper Kapusniak, Michael M\. Bronstein, Joey Bose, and Alexander Tong\.Curly flow matching for learning non\-gradient field dynamics\.In*Learning Meaningful Representations of Life \(LMRL\) Workshop at ICLR 2025*, 2025\.
- Peyrรฉ and Cuturi \[2019\]Gabriel Peyrรฉ and Marco Cuturi\.*Computational optimal transport: With applications to data science*\.Now Foundations and Trends, 2019\.
- Pooladian et al\. \[2023\]Aram\-Alexandre Pooladian, Heli Ben\-Hamu, Carles Domingo\-Enrich, Brandon Amos, Yaron Lipman, and Ricky TQ Chen\.Multisample flow matching: Straightening flows with minibatch couplings\.*arXiv preprint arXiv:2304\.14772*, 2023\.
- Prasad et al\. \[2020\]Neha Prasad, Karren Yang, and Caroline Uhler\.Optimal transport using gans for lineage tracing\.*arXiv preprint arXiv:2007\.12098*, 2020\.
- Rohbeck et al\. \[2025\]Martin Rohbeck, Edward De Brouwer, Charlotte Bunne, Jan\-Christian Huetter, Anne Biton, Kelvin Y Chen, Aviv Regev, and Romain Lopez\.Modeling complex system dynamics with flow matching across time and conditions\.In*The Thirteenth International Conference on Learning Representations*, 2025\.
- Schiebinger et al\. \[2019\]Geoffrey Schiebinger, Jian Shu, Marcin Tabaka, Brian Cleary, Vidya Subramanian, Aryeh Solomon, Joshua Gould, Siyan Liu, Stacie Lin, Peter Berube, et al\.Optimal\-transport analysis of single\-cell gene expression identifies developmental trajectories in reprogramming\.*Cell*, 176\(4\):928โ€“943, 2019\.
- Sha et al\. \[2024\]Yutong Sha, Yuchi Qiu, Peijie Zhou, and Qing Nie\.Reconstructing growth and dynamic trajectories from single\-cell transcriptomics data\.*Nature Machine Intelligence*, 6\(1\):25โ€“39, 2024\.
- Shi et al\. \[2024\]Yuyang Shi, Valentin De Bortoli, Andrew Campbell, and Arnaud Doucet\.Diffusion schrรถdinger bridge matching\.*Advances in Neural Information Processing Systems*, 36, 2024\.
- Shrestha and Fu \[2025\]Sagar Shrestha and Xiao Fu\.Diversified flow matching with translation identifiability\.*arXiv preprint arXiv:2511\.05558*, 2025\.
- Sun et al\. \[2025\]Yuhao Sun, Zhenyi Zhang, Zihan Wang, Tiejun Li, and Peijie Zhou\.Variational regularized unbalanced optimal transport: Single network, least action\.*arXiv preprint arXiv:2505\.11823*, 2025\.
- Theodoropoulos et al\. \[2025\]Panagiotis Theodoropoulos, Augustinos D Saravanos, Evangelos A Theodorou, and Guan\-Horng Liu\.Momentum multi\-marginal schr\\\\backslash" odinger bridge matching\.*arXiv preprint arXiv:2506\.10168*, 2025\.
- Tong et al\. \[2020\]Alexander Tong, Jessie Huang, Guy Wolf, David Van Dijk, and Smita Krishnaswamy\.Trajectorynet: A dynamic optimal transport network for modeling cellular dynamics\.In*International conference on machine learning*, pages 9526โ€“9536\. PMLR, 2020\.
- Tong et al\. \[2024a\]Alexander Tong, Kilian FATRAS, Nikolay Malkin, Guillaume Huguet, Yanlei Zhang, Jarrid Rector\-Brooks, Guy Wolf, and Yoshua Bengio\.Improving and generalizing flow\-based generative models with minibatch optimal transport\.*Transactions on Machine Learning Research*, 2024a\.ISSN 2835\-8856\.Expert Certification\.
- Tong et al\. \[2024b\]Alexander Tong, Nikolay Malkin, Kilian Fatras, Lazar Atanackovic, Yanlei Zhang, Guillaume Huguet, Guy Wolf, and Yoshua Bengio\.Simulation\-free schrรถdinger bridges via score and flow matching\.In*International Conference on Artificial Intelligence and Statistics*, pages 1279โ€“1287\. PMLR, 2024b\.
- Ventre et al\. \[2023\]Elias Ventre, Aden Forrow, Nitya Gadhiwala, Parijat Chakraborty, Omer Angel, and Geoffrey Schiebinger\.Trajectory inference for a branching sde model of cell differentiation\.*arXiv preprint arXiv:2307\.07687*, 2023\.
- Wang et al\. \[2025\]Dongyi Wang, Yuanwei Jiang, Zhenyi Zhang, Xiang Gu, Peijie Zhou, and Jian Sun\.Joint velocity\-growth flow matching for single\-cell dynamics modeling\.*arXiv preprint arXiv:2505\.13413*, 2025\.
- Wang et al\. \[2023\]Kun Wang, Liangzhen Hou, Xin Wang, Xiangwei Zhai, Zhaolian Lu, Zhike Zi, Weiwei Zhai, Xionglei He, Christina Curtis, Da Zhou, and Zheng Hu\.Phylovelo enhances transcriptomic velocity field mapping using monotonically expressed genes\.*Nature Biotechnology*, 42\(5\):778โ€“789, 2023\.
- Wang et al\. \[2022\]Shou\-Wen Wang, Michael J\. Herriges, Kilian Hurley, Darrell N\. Kotton, and Allon M\. Klein\.Cospar identifies early cell fate biases from single\-cell transcriptomic and lineage information\.*Nature Biotechnology*, 40\(7\):1066โ€“1074, 2022\.
- Weinreb et al\. \[2020\]Caleb Weinreb, Alejo Rodriguez\-Fraticelli, Fernando D Camargo, and Allon M Klein\.Lineage tracing on transcriptional landscapes links state to fate during differentiation\.*Science*, 367\(6479\):eaaw3381, 2020\.
- Yang \[2025\]Maosheng Yang\.Topological schrรถdinger bridge matching\.In*The Thirteenth International Conference on Learning Representations*, 2025\.
- Yeo et al\. \[2021\]Grace Hui Ting Yeo, Sachit D Saksena, and David K Gifford\.Generative modeling of single\-cell time series with prescient enables prediction of cell trajectories with interventions\.*Nature communications*, 12\(1\):3222, 2021\.
- Zhang et al\. \[2024a\]Jiaqi Zhang, Erica Larschan, Jeremy Bigness, and Ritambhara Singh\.scNODE: generative model for temporal single cell transcriptomic data prediction\.*Bioinformatics*, 40\(Supplement\_2\):ii146โ€“ii154, 09 2024a\.ISSN 1367\-4811\.
- Zhang et al\. \[2024b\]Xi Nicole Zhang, Yuan Pu, Yuki Kawamura, Andrew Loza, Yoshua Bengio, Dennis Shung, and Alexander Tong\.Trajectory flow matching with applications to clinical time series modelling\.*Advances in Neural Information Processing Systems*, 37:107198โ€“107224, 2024b\.
- Zhang et al\. \[2025a\]Yichi Zhang, Yici Yan, Alex Schwing, and Zhizhen Zhao\.Towards hierarchical rectified flow\.*arXiv preprint arXiv:2502\.17436*, 2025a\.
- Zhang et al\. \[2025b\]Zhenyi Zhang, Tiejun Li, and Peijie Zhou\.Learning stochastic dynamics from snapshots through regularized unbalanced optimal transport\.In*The Thirteenth International Conference on Learning Representations*, 2025b\.
- Zhang et al\. \[2025c\]Zhenyi Zhang, Yuhao Sun, Qiangwei Peng, Tiejun Li, and Peijie Zhou\.Integrating dynamical systems modeling with spatiotemporal scRNA\-Seq data analysis\.*Entropy*, 27\(5\), 2025c\.ISSN 1099\-4300\.
- Zheng et al\. \[2017\]Grace XY Zheng, Jessica M Terry, Phillip Belgrader, Paul Ryvkin, Zachary W Bent, Ryan Wilson, Solongo B Ziraldo, Tobias D Wheeler, Geoff P McDermott, Junjie Zhu, et al\.Massively parallel digital transcriptional profiling of single cells\.*Nature communications*, 8\(1\):14049, 2017\.
- Zhu and Lin \[2024\]Qunxi Zhu and Wei Lin\.Switched flow matching: Eliminating singularities via switching odes\.*arXiv preprint arXiv:2405\.11605*, 2024\.
- Zhu et al\. \[2024\]Qunxi Zhu, Bolin Zhao, Jingdong Zhang, Peiyang Li, and Wei Lin\.Governing equation discovery of a complex system from snapshots\.*arXiv preprint arXiv:2410\.16694*, 2024\.

## Contents of Appendix

## Appendix AProofs of Main Theorems

### A\.1Proof of Proposition 4\.1

###### Proof\.

First, we show that the unique minimizer is a cubic polynomial\. Consider the energy functional defined on the interval\[0,T\]\[0,T\]:

๐’ฅโก\[ฮณ\]=12โ€‹โˆซ0Tโ€–ฮณโ€ฒโ€ฒโ€‹\(t\)โ€–22โ€‹๐‘‘t\.\\mathcal\{J\}\[\\gamma\]=\\frac\{1\}\{2\}\\int\_\{0\}^\{T\}\\\|\\gamma^\{\\prime\\prime\}\(t\)\\\|\_\{2\}^\{2\}\\,\\mathrm\{d\}t\.\(20\)The integrandLโก\(ฮณ,ฮณโ€ฒ,ฮณโ€ฒโ€ฒ\)=12โ€‹โ€–ฮณโ€ฒโ€ฒโ€–22L\(\\gamma,\\gamma^\{\\prime\},\\gamma^\{\\prime\\prime\}\)=\\frac\{1\}\{2\}\\\|\\gamma^\{\\prime\\prime\}\\\|\_\{2\}^\{2\}depends on derivatives up to the second order\. The necessary condition for optimality is given by the Eulerโ€“Lagrange equation:

โˆ‚Lโˆ‚ฮณโˆ’ddโ€‹tโ€‹\(โˆ‚Lโˆ‚ฮณโ€ฒ\)\+d2dโ€‹t2โ€‹\(โˆ‚Lโˆ‚ฮณโ€ฒโ€ฒ\)=0\.\\frac\{\\partial L\}\{\\partial\\gamma\}\-\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\left\(\\frac\{\\partial L\}\{\\partial\\gamma^\{\\prime\}\}\\right\)\+\\frac\{\\mathrm\{d\}^\{2\}\}\{\\mathrm\{d\}t^\{2\}\}\\left\(\\frac\{\\partial L\}\{\\partial\\gamma^\{\\prime\\prime\}\}\\right\)=0\.\(21\)Sinceโˆ‚L/โˆ‚ฮณ=0\\partial L/\\partial\\gamma=0,โˆ‚L/โˆ‚ฮณโ€ฒ=0\\partial L/\\partial\\gamma^\{\\prime\}=0, andโˆ‚L/โˆ‚ฮณโ€ฒโ€ฒ=ฮณโ€ฒโ€ฒ\\partial L/\\partial\\gamma^\{\\prime\\prime\}=\\gamma^\{\\prime\\prime\}, the equation simplifies to:

d2dโ€‹t2โ€‹ฮณโ€ฒโ€ฒโ€‹\(t\)=ฮณ\(4\)โ€‹\(t\)=0\.\\frac\{\\mathrm\{d\}^\{2\}\}\{\\mathrm\{d\}t^\{2\}\}\\gamma^\{\\prime\\prime\}\(t\)=\\gamma^\{\(4\)\}\(t\)=0\.\(22\)This implies that the optimal trajectoryฮณโก\(t\)\\gamma\(t\)is a cubic polynomial:

ฮณโก\(t\)=c3โ€‹t3\+c2โ€‹t2\+c1โ€‹t\+c0\.\\gamma\(t\)=c\_\{3\}t^\{3\}\+c\_\{2\}t^\{2\}\+c\_\{1\}t\+c\_\{0\}\.\(23\)Imposing the boundary conditionsฮณโก\(0\)=๐’™0,ฮณโ€ฒโ€‹\(0\)=๐’—0\\gamma\(0\)=\\bm\{x\}\_\{0\},\\gamma^\{\\prime\}\(0\)=\\bm\{v\}\_\{0\}andฮณโก\(T\)=๐’™T,ฮณโ€ฒโ€‹\(T\)=๐’—T\\gamma\(T\)=\\bm\{x\}\_\{T\},\\gamma^\{\\prime\}\(T\)=\\bm\{v\}\_\{T\}leads to the linear system:

\{c0=๐’™0,c1=๐’—0,c3โ€‹T3\+c2โ€‹T2\+c1โ€‹T\+c0=๐’™T,3โ€‹c3โ€‹T2\+2โ€‹c2โ€‹T\+c1=๐’—T\.\\begin\{cases\}c\_\{0\}=\\bm\{x\}\_\{0\},\\\\ c\_\{1\}=\\bm\{v\}\_\{0\},\\\\ c\_\{3\}T^\{3\}\+c\_\{2\}T^\{2\}\+c\_\{1\}T\+c\_\{0\}=\\bm\{x\}\_\{T\},\\\\ 3c\_\{3\}T^\{2\}\+2c\_\{2\}T\+c\_\{1\}=\\bm\{v\}\_\{T\}\.\\end\{cases\}\(24\)Solving for the coefficients yields the unique solution:

\{c0=๐’™0,c1=๐’—0,c2=3โ€‹\(๐’™Tโˆ’๐’™0\)โˆ’\(2โ€‹๐’—0\+๐’—T\)โ€‹TT2,c3=\(๐’—0\+๐’—T\)โ€‹Tโˆ’2โ€‹\(๐’™Tโˆ’๐’™0\)T3\.\\begin\{cases\}c\_\{0\}=\\bm\{x\}\_\{0\},\\\\ c\_\{1\}=\\bm\{v\}\_\{0\},\\\\ c\_\{2\}=\\dfrac\{3\(\\bm\{x\}\_\{T\}\-\\bm\{x\}\_\{0\}\)\-\(2\\bm\{v\}\_\{0\}\+\\bm\{v\}\_\{T\}\)T\}\{T^\{2\}\},\\\\ c\_\{3\}=\\dfrac\{\(\\bm\{v\}\_\{0\}\+\\bm\{v\}\_\{T\}\)T\-2\(\\bm\{x\}\_\{T\}\-\\bm\{x\}\_\{0\}\)\}\{T^\{3\}\}\.\\end\{cases\}\(25\)โˆŽ

### A\.2Proof of Remark 4\.3

###### Proof\.

First we propose a lemma:

Lemma\.Let๐’ž\[0,t\]โ€‹\(๐’™0,๐’—0,๐’™t,๐’—t\)\\mathcal\{C\}\_\{\[0,t\]\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\},\\bm\{x\}\_\{t\},\\bm\{v\}\_\{t\}\)be the least acceleration cost through\[0,t\]\[0,t\],\(๐’™๐ŸŽ,๐’—0\),\(๐’™T,๐’—T\)โˆˆ๐’ณร—๐’ฑ\(\\bm\{x\_\{0\}\},\\bm\{v\}\_\{0\}\),\(\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\)\\in\\mathcal\{X\}\\times\\mathcal\{V\}are two different points\. Then for arbitrary midpointฯ„โˆˆ\(0,1\)\\tau\\in\(0,1\)and state\(๐’™,๐’—\)โˆˆ๐’ณร—๐’ฑ\(\\bm\{x\},\\bm\{v\}\)\\in\\mathcal\{X\}\\times\\mathcal\{V\}, the following inequality holds:

๐’ž\[0,ฯ„\]โ€‹\(๐’™๐ŸŽ,๐’—๐ŸŽ,๐’™,๐’—\)\+๐’ž\[ฯ„,T\]โ€‹\(๐’™,๐’—,๐’™T,๐’—T\)โ‰ฅ๐’ž\[0,T\]โ€‹\(๐’™0,๐’—0,๐’™T,๐’—T\)\\mathcal\{C\}\_\{\[0,\\tau\]\}\(\\bm\{x\_\{0\}\},\\bm\{v\_\{0\}\},\\bm\{x\},\\bm\{v\}\)\+\\mathcal\{C\}\_\{\[\\tau,T\]\}\(\\bm\{x\},\\bm\{v\},\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\)\\geq\\mathcal\{C\}\_\{\[0,T\]\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\},\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\)\(26\)The "==" holds if and only if\(๐’™,๐’—\)=ฮณโˆ—โ€‹\(ฯ„\)\(\\bm\{x,v\}\)=\\gamma^\{\*\}\(\\tau\), whereฮณโˆ—โ€‹\(t\)โ€‹\(tโˆˆ\[0,T\]\)\\gamma^\{\*\}\(t\)\(t\\in\[0,T\]\)is the cubic interpolation trajectory connecting\(๐’™๐ŸŽ,๐’—๐ŸŽ\)\(\\bm\{x\_\{0\}\},\\bm\{v\_\{0\}\}\)and\(๐’™T,๐’—๐‘ป\)\(\\bm\{x\}\_\{T\},\\bm\{v\_\{T\}\}\)\.

###### Proof\.

Fixฯ„โˆˆ\(0,T\)\\tau\\in\(0,T\)\. Define objective functionFฯ„โ€‹\(๐’™,๐’—\)=๐’ž\[0,ฯ„\]โ€‹\(๐’™๐ŸŽ,๐’—๐ŸŽ,๐’™,๐’—\)\+๐’ž\[ฯ„,T\]โ€‹\(๐’™,๐’—,๐’™T,๐’—T\)F\_\{\\tau\}\(\\bm\{x\},\\bm\{v\}\)=\\mathcal\{C\}\_\{\[0,\\tau\]\}\(\\bm\{x\_\{0\}\},\\bm\{v\_\{0\}\},\\bm\{x\},\\bm\{v\}\)\+\\mathcal\{C\}\_\{\[\\tau,T\]\}\(\\bm\{x\},\\bm\{v\},\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\)\. By Corollary[4\.1](https://arxiv.org/html/2608.21070#S4.SS1.SSS0.Px2), we can rewrite

Fฯ„โ€‹\(๐’™,๐’—\)=2ฯ„3โ€‹\{\(โ€–๐’—0โ€–2\+โŸจ๐’—0,๐’—โŸฉ\+โ€–๐’—โ€–2\)โ€‹ฯ„2\+3โ€‹โŸจ๐’—0\+๐’—,๐’™0โˆ’๐’™โŸฉโ€‹ฯ„\+3โ€‹โ€–๐’™0โˆ’๐’™โ€–2\}\+\\displaystyle F\_\{\\tau\}\(\\bm\{x\},\\bm\{v\}\)=\\frac\{2\}\{\\tau^\{3\}\}\\left\\\{\(\\\|\\bm\{v\}\_\{0\}\\\|^\{2\}\+\\langle\\bm\{v\}\_\{0\},\\bm\{v\}\\rangle\+\\\|\\bm\{v\}\\\|^\{2\}\)\\tau^\{2\}\+3\\langle\\bm\{v\}\_\{0\}\+\\bm\{v\},\\bm\{x\}\_\{0\}\-\\bm\{x\}\\rangle\\tau\+3\\\|\\bm\{x\}\_\{0\}\-\\bm\{x\}\\\|^\{2\}\\right\\\}\+\(27\)2\(Tโˆ’ฯ„\)3โ€‹\{\(โ€–๐’—Tโ€–2\+โŸจ๐’—T,๐’—โŸฉ\+โ€–๐’—โ€–2\)โ€‹\(Tโˆ’ฯ„\)2\+3โ€‹โŸจ๐’—T\+๐’—,๐’™โˆ’๐’™TโŸฉโ€‹T\+3โ€‹โ€–๐’™Tโˆ’๐’™โ€–2\}\\displaystyle\\frac\{2\}\{\(T\-\\tau\)^\{3\}\}\\left\\\{\(\\\|\\bm\{v\}\_\{T\}\\\|^\{2\}\+\\langle\\bm\{v\}\_\{T\},\\bm\{v\}\\rangle\+\\\|\\bm\{v\}\\\|^\{2\}\)\(T\-\\tau\)^\{2\}\+3\\langle\\bm\{v\}\_\{T\}\+\\bm\{v\},\\bm\{x\}\-\\bm\{x\}\_\{T\}\\rangle T\+3\\\|\\bm\{x\}\_\{T\}\-\\bm\{x\}\\\|^\{2\}\\right\\\}Which is a quadratic form of\(๐’™,๐’—\)\(\\bm\{x\},\\bm\{v\}\)up to a constant\.Fฯ„โ€‹\(๐’™,๐’—\)โ‰ฅ0F\_\{\\tau\}\(\\bm\{x\},\\bm\{v\}\)\\geq 0always holds, namely, the quadratic form has a finite lower bound, so it is semi\-quadratic\. To minimizeFฯ„โ€‹\(๐’™,๐’—\)F\_\{\\tau\}\(\\bm\{x\},\\bm\{v\}\), we can take differentiations:

โˆ‡๐’™Fฯ„โ€‹\(๐’™,๐’—\)=3โ€‹\(โˆ’\(๐’—0\+๐’—\)โ€‹ฯ„\+2โ€‹\(๐’™โˆ’๐’™0\)ฯ„2\+\(๐’—\+๐’—T\)โ€‹\(Tโˆ’ฯ„\)\+2โ€‹\(๐’™โˆ’๐’™T\)\(Tโˆ’ฯ„\)2\)=0\.\\nabla\_\{\\bm\{x\}\}F\_\{\\tau\}\(\\bm\{x\},\\bm\{v\}\)=3\\left\(\\frac\{\-\(\\bm\{v\}\_\{0\}\+\\bm\{v\}\)\\tau\+2\(\\bm\{x\}\-\\bm\{x\}\_\{0\}\)\}\{\{\\tau\}^\{2\}\}\+\\frac\{\(\\bm\{v\}\+\\bm\{v\}\_\{T\}\)\(T\-\\tau\)\+2\(\\bm\{x\}\-\\bm\{x\}\_\{T\}\)\}\{\(T\-\{\\tau\}\)^\{2\}\}\\right\)=0\.\(28\)โˆ‡๐’—Fฯ„โ€‹\(๐’™,๐’—\)=\(๐’—0\+2โ€‹๐’—\)โ€‹ฯ„\+3โ€‹\(๐’™0โˆ’๐’™\)ฯ„2\+\(๐’—T\+2โ€‹๐’—\)โ€‹\(Tโˆ’ฯ„\)\+3โ€‹\(๐’™โˆ’๐’™T\)\(Tโˆ’ฯ„\)2=0,\\nabla\_\{\\bm\{v\}\}F\_\{\\tau\}\(\\bm\{x\},\\bm\{v\}\)=\\frac\{\(\\bm\{v\}\_\{0\}\+2\\bm\{v\}\)\{\\tau\}\+3\(\\bm\{x\}\_\{0\}\-\\bm\{x\}\)\}\{\{\\tau\}^\{2\}\}\+\\frac\{\(\\bm\{v\}\_\{T\}\+2\\bm\{v\}\)\(T\-\{\\tau\}\)\+3\(\\bm\{x\}\-\\bm\{x\}\_\{T\}\)\}\{\(T\-\{\\tau\}\)^\{2\}\}=0,\(29\)
Denote the cubic interpolation from\(๐’™๐ŸŽ,๐’—๐ŸŽ\)\(\\bm\{x\_\{0\}\},\\bm\{v\_\{0\}\}\)to\(๐’™,๐’—\)\(\\bm\{x\},\\bm\{v\}\)on\[0,ฯ„\]\[0,\\tau\]and\(๐’™,๐’—\)\(\\bm\{x\},\\bm\{v\}\)to\(๐’™T,๐’—T\)\(\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\)on\[ฯ„,T\]\[\\tau,T\]byฮณ\[0,ฯ„\]โ€‹\(t\)\\gamma\_\{\[0,\\tau\]\}\(t\)andฮณ\[ฯ„,T\]โ€‹\(t\)\\gamma\_\{\[\\tau,T\]\}\(\{t\}\), respectively\.The coefficients are known according to[SectionA\.1](https://arxiv.org/html/2608.21070#A1.SS1)\. Then we have:

ฮณ\[0,ฯ„\]โ€ฒโ€ฒ\(ฯ„โˆ’\)=6โ‹…\(๐’—0\+๐’—\)โ€‹ฯ„โˆ’2โ€‹\(๐’™โˆ’๐’™0\)ฯ„2\+2โ‹…3โ€‹\(๐’™โˆ’๐’™0\)โˆ’\(2โ€‹๐’—0\+๐’—\)โ€‹ฯ„ฯ„2=2โ‹…3โ€‹\(๐’™๐ŸŽโˆ’๐’™\)\+\(๐’—๐ŸŽ\+2โ€‹๐’—\)โ€‹ฯ„ฯ„2,\\gamma\_\{\[0,\\tau\]\}^\{\{\}^\{\\prime\\prime\}\}\(\\tau^\{\-\}\)=6\\cdot\\frac\{\(\\bm\{v\}\_\{0\}\+\\bm\{v\}\)\\tau\-2\(\\bm\{x\}\-\\bm\{x\}\_\{0\}\)\}\{\\tau^\{2\}\}\+2\\cdot\\frac\{3\{\(\\bm\{x\}\-\\bm\{x\}\_\{0\}\)\-\(2\\bm\{v\}\_\{0\}\+\\bm\{v\}\)\}\\tau\}\{\\tau^\{2\}\}=2\\cdot\\frac\{3\(\\bm\{x\_\{0\}\}\-\\bm\{x\}\)\+\(\\bm\{v\_\{0\}\}\+2\\bm\{v\}\)\\tau\}\{\\tau^\{2\}\},\(30\)ฮณ\[ฯ„,T\]โ€ฒโ€ฒ\(ฯ„\+\)=6โ‹…\(๐’—\+๐’—T\)โ€‹ฯ„โˆ’2โ€‹\(๐’™Tโˆ’๐’™\)ฯ„2\+2โ‹…3โ€‹\(๐’™Tโˆ’๐’™\)โˆ’\(2โ€‹๐’—\+๐’—T\)โ€‹\(Tโˆ’ฯ„\)\(Tโˆ’ฯ„\)2\\displaystyle\\gamma\_\{\[\\tau,T\]\}^\{\{\}^\{\\prime\\prime\}\}\(\\tau^\{\+\}\)=6\\cdot\\frac\{\(\\bm\{v\}\+\\bm\{v\}\_\{T\}\)\\tau\-2\(\\bm\{x\}\_\{T\}\-\\bm\{x\}\)\}\{\\tau^\{2\}\}\+2\\cdot\\frac\{3\(\\bm\{x\}\_\{T\}\-\\bm\{x\}\)\-\(2\\bm\{v\}\+\\bm\{v\}\_\{T\}\)\(T\-\\tau\)\}\{\(T\-\\tau\)^\{2\}\}\(31\)=2โ‹…3โ€‹\(๐’™โˆ’๐’™T\)\+\(๐’—\+2โ€‹๐’—T\)โ€‹\(Tโˆ’ฯ„\)\(Tโˆ’ฯ„\)2\\displaystyle=2\\cdot\\frac\{3\(\\bm\{x\}\-\\bm\{x\}\_\{T\}\)\+\(\\bm\{v\}\+2\\bm\{v\}\_\{T\}\)\(T\-\\tau\)\}\{\(T\-\\tau\)^\{2\}\}ฮณ\[0,ฯ„\]โ€ฒโ€ฒโ€ฒ\(ฯ„โˆ’\)=6โ‹…\(๐’—0\+๐’—\)โ€‹ฯ„โˆ’2โ€‹\(๐’™โˆ’๐’™0\)ฯ„2\\gamma\_\{\[0,\\tau\]\}^\{\{\}^\{\\prime\\prime\\prime\}\}\(\\tau^\{\-\}\)=6\\cdot\\frac\{\(\\bm\{v\}\_\{0\}\+\\bm\{v\}\)\\tau\-2\(\\bm\{x\}\-\\bm\{x\}\_\{0\}\)\}\{\\tau^\{2\}\}\(32\)
ฮณ\[ฯ„,T\]โ€ฒโ€ฒโ€ฒ\(ฯ„\+\)=6โ‹…\(๐’—\+๐’—T\)โ€‹\(Tโˆ’ฯ„\)โˆ’2โ€‹\(๐’™Tโˆ’๐’™\)\(Tโˆ’ฯ„\)2\\displaystyle\\gamma\_\{\[\\tau,T\]\}^\{\{\}^\{\\prime\\prime\\prime\}\}\(\\tau^\{\+\}\)=6\\cdot\\frac\{\(\\bm\{v\}\+\\bm\{v\}\_\{T\}\)\(T\-\\tau\)\-2\(\\bm\{x\}\_\{T\}\-\\bm\{x\}\)\}\{\(T\-\\tau\)^\{2\}\}\(33\)
Then we can find that[28](https://arxiv.org/html/2608.21070#A1.E28)and[29](https://arxiv.org/html/2608.21070#A1.E29)appear to beฮณ\[0,ฯ„\]โ€ฒโ€ฒโ€ฒ\(ฯ„โˆ’\)=ฮณ\[ฯ„,T\]โ€ฒโ€ฒโ€ฒ\(ฯ„\+\),ฮณ\[0,ฯ„\]โ€ฒโ€ฒ\(ฯ„โˆ’\)=ฮณ\[ฯ„,T\]โ€ฒโ€ฒ\(ฯ„\+\)\.\\gamma\_\{\[0,\\tau\]\}^\{\{\}^\{\\prime\\prime\\prime\}\}\(\\tau^\{\-\}\)=\\gamma\_\{\[\\tau,T\]\}^\{\{\}^\{\\prime\\prime\\prime\}\}\(\\tau^\{\+\}\),\\gamma\_\{\[0,\\tau\]\}^\{\{\}^\{\\prime\\prime\}\}\(\\tau^\{\-\}\)=\\gamma\_\{\[\\tau,T\]\}^\{\{\}^\{\\prime\\prime\}\}\(\\tau^\{\+\}\)\.Because the two curves are cubic, andฮณ\[0,ฯ„\]\(i\)\(ฯ„โˆ’\)=ฮณ\[ฯ„,T\]\(i\)\(ฯ„\+\),i=0,1,2,3\\gamma^\{\(i\)\}\_\{\[0,\\tau\]\}\(\\tau^\{\-\}\)=\\gamma\_\{\[\\tau,T\]\}^\{\(i\)\}\(\\tau^\{\+\}\),i=0,1,2,3, so we know that in the optimal caseฮณ\[0,ฯ„\]=ฮณโˆ—\|\[0,ฯ„\],ฮณ\[ฯ„,T\]=ฮณโˆ—\|\[ฯ„,T\]\.\\gamma\_\{\[0,\\tau\]\}=\\gamma^\{\*\}\\big\|\_\{\[0,\\tau\]\},\\gamma\_\{\[\\tau,T\]\}=\\gamma^\{\*\}\\big\|\_\{\[\\tau,T\]\}\.\. That is, the unique solution of[28](https://arxiv.org/html/2608.21070#A1.E28)and[29](https://arxiv.org/html/2608.21070#A1.E29)is\(ฮณโˆ—\(ฯ„\),ฮณโˆ—โ€ฒ\(ฯ„\)\),\(\\gamma^\{\*\}\(\\tau\),\{\\gamma^\{\*\}\}^\{\{\}^\{\\prime\}\}\(\\tau\)\),which completes the proof\. โˆŽ

Back to our target, assumeฮณ\\gammaandฮณ~\\tilde\{\\gamma\}cross at\(๐’™ฯ„,๐’—ฯ„\)\(\\bm\{x\}\_\{\\tau\},\\bm\{v\}\_\{\\tau\}\), then we have

๐’ž\[0,T\]โ€‹\(๐’™0,๐’—0,๐’™T,๐’—T\)\+๐’ž\[0,T\]โ€‹\(๐’™~0,๐’—~0,๐’™~T,๐’—~T\)\\displaystyle\\mathcal\{C\}\_\{\[0,T\]\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\},\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\)\+\\mathcal\{C\}\_\{\[0,T\]\}\(\\tilde\{\\bm\{x\}\}\_\{0\},\\tilde\{\\bm\{v\}\}\_\{0\},\\tilde\{\\bm\{x\}\}\_\{T\},\\tilde\{\\bm\{v\}\}\_\{T\}\)\(34\)=\\displaystyle=\{๐’ž\[0,ฯ„\]โ€‹\(๐’™0,๐’—0,๐’™ฯ„,๐’—ฯ„\)\+๐’ž\[ฯ„,T\]โ€‹\(๐’™ฯ„,๐’—ฯ„,๐’™T,๐’—T\)\}\+\{๐’ž\[0,ฯ„\]โ€‹\(๐’™~0,๐’—~0,๐’™ฯ„,๐’—ฯ„\)\+๐’ž\[ฯ„,T\]โ€‹\(๐’™ฯ„,๐’—ฯ„,๐’™~T,๐’—~T\)\}\\displaystyle\\left\\\{\\mathcal\{C\}\_\{\[0,\\tau\]\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\},\\bm\{x\}\_\{\\tau\},\\bm\{v\}\_\{\\tau\}\)\+\\mathcal\{C\}\_\{\[\\tau,T\]\}\(\\bm\{x\}\_\{\\tau\},\\bm\{v\}\_\{\\tau\},\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\)\\right\\\}\+\\left\\\{\\mathcal\{C\}\_\{\[0,\\tau\]\}\(\\tilde\{\\bm\{x\}\}\_\{0\},\\tilde\{\\bm\{v\}\}\_\{0\},\\bm\{x\}\_\{\\tau\},\\bm\{v\}\_\{\\tau\}\)\+\\mathcal\{C\}\_\{\[\\tau,T\]\}\(\\bm\{x\}\_\{\\tau\},\\bm\{v\}\_\{\\tau\},\\tilde\{\\bm\{x\}\}\_\{T\},\\tilde\{\\bm\{v\}\}\_\{T\}\)\\right\\\}=\\displaystyle=\{๐’ž\[0,ฯ„\]โ€‹\(๐’™0,๐’—0,๐’™ฯ„,๐’—ฯ„\)\+๐’ž\[ฯ„,T\]โ€‹\(๐’™ฯ„,๐’—ฯ„,๐’™~T,๐’—~T\)\}\+\{๐’ž\[0,ฯ„\]โ€‹\(๐’™~0,๐’—~0,๐’™ฯ„,๐’—ฯ„\)\+๐’ž\[ฯ„,T\]โ€‹\(๐’™ฯ„,๐’—ฯ„,๐’™T,๐’—T\)\}\\displaystyle\\left\\\{\\mathcal\{C\}\_\{\[0,\\tau\]\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\},\\bm\{x\}\_\{\\tau\},\\bm\{v\}\_\{\\tau\}\)\+\\mathcal\{C\}\_\{\[\\tau,T\]\}\(\\bm\{x\}\_\{\\tau\},\\bm\{v\}\_\{\\tau\},\\tilde\{\\bm\{x\}\}\_\{T\},\\tilde\{\\bm\{v\}\}\_\{T\}\)\\right\\\}\+\\left\\\{\\mathcal\{C\}\_\{\[0,\\tau\]\}\(\\tilde\{\\bm\{x\}\}\_\{0\},\\tilde\{\\bm\{v\}\}\_\{0\},\\bm\{x\}\_\{\\tau\},\\bm\{v\}\_\{\\tau\}\)\+\\mathcal\{C\}\_\{\[\\tau,T\]\}\(\\bm\{x\}\_\{\\tau\},\\bm\{v\}\_\{\\tau\},\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\)\\right\\\}โ‰ฅ\\displaystyle\\geq๐’ž\[0,T\]โ€‹\(๐’™0,๐’—0,๐’™~T,๐’—~T\)\+๐’ž\[0,T\]โ€‹\(๐’™~0,๐’—~0,๐’™T,๐’—T\)\.\\displaystyle\\mathcal\{C\}\_\{\[0,T\]\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\},\\tilde\{\\bm\{x\}\}\_\{T\},\\tilde\{\\bm\{v\}\}\_\{T\}\)\+\\mathcal\{C\}\_\{\[0,T\]\}\(\\tilde\{\\bm\{x\}\}\_\{0\},\\tilde\{\\bm\{v\}\}\_\{0\},\\bm\{x\}\_\{T\},\\bm\{v\}\_\{T\}\)\.
The equality holds if and only ifฮณ\\gammaandฮณ~\\tilde\{\\gamma\}coincide\. Consequently, the strict inequality holds for distinct trajectories, which implies the non\-crossing property\. This completes the proof\. โˆŽ

### A\.3Proof of Theorem 4\.4

###### Proof\.

Consider the time derivative of the marginal distribution:

โˆ‚tฯtโ€‹\(๐’™,๐’—\)\\displaystyle\\partial\_\{t\}\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)=โˆ‚t\{โˆซฯtโ€‹\(๐’™,๐’—\|z\)โ€‹qโ€‹\(z\)โ€‹๐‘‘z\}\\displaystyle=\\partial\_\{t\}\\left\\\{\\int\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|z\)q\(z\)\\mathrm\{d\}z\\right\\\}=โˆซโˆ‚tฯtโ€‹\(๐’™,๐’—\|z\)โ€‹qโ€‹\(z\)โ€‹๐‘‘z\\displaystyle=\\int\\partial\_\{t\}\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|z\)q\(z\)\\mathrm\{d\}z=โˆ’โˆซโˆ‡๐’™โ‹…\(ฯt\(๐’™,๐’—\|z\)๐’—\)q\(z\)dzโˆ’โˆซโˆ‡๐’—โ‹…\(ฯt\(๐’™,๐’—\|z\)๐’‚\(๐’™,๐’—,t\|z\)\)q\(z\)dz\\displaystyle=\-\\int\\nabla\_\{\\bm\{x\}\}\\cdot\(\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|z\)\\bm\{v\}\)q\(z\)\\mathrm\{d\}z\-\\int\\nabla\_\{\\bm\{v\}\}\\cdot\(\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|z\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\|z\)\)q\(z\)\\mathrm\{d\}z=โˆ’โˆ‡๐’™โ‹…\(โˆซฯt\(๐’™,๐’—\|z\)๐’—q\(z\)dz\)โˆ’โˆ‡๐’—โ‹…\(โˆซฯt\(๐’™,๐’—\|z\)๐’‚\(๐’™,๐’—,t\|z\)q\(z\)dz\)\\displaystyle=\-\\nabla\_\{\\bm\{x\}\}\\cdot\\left\(\\int\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|z\)\\bm\{v\}q\(z\)\\mathrm\{d\}z\\right\)\-\\nabla\_\{\\bm\{v\}\}\\cdot\\left\(\\int\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|z\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\|z\)q\(z\)\\mathrm\{d\}z\\right\)=โˆ’โˆ‡๐’™โ‹…\(ฯt\(๐’™,๐’—\)๐’—\)โˆ’โˆ‡๐’—โ‹…\(๐’‚t\(๐’™,๐’—\)ฯt\(๐’™,๐’—\)\)\\displaystyle=\-\\nabla\_\{\\bm\{x\}\}\\cdot\(\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\\bm\{v\}\)\-\\nabla\_\{\\bm\{v\}\}\\cdot\(\\bm\{a\}\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\)\(35\)Therefore,ฯtโ€‹\(๐’™,๐’—\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)satisfies the continuity equation:

โˆ‚tฯt\(๐’™,๐’—\)=โˆ’โˆ‡๐’™โ‹…\(ฯt\(๐’™,๐’—\)๐’—\)โˆ’โˆ‡๐’—โ‹…\(๐’‚t\(๐’™,๐’—\)ฯt\(๐’™,๐’—\)\)\\partial\_\{t\}\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)=\-\\nabla\_\{\\bm\{x\}\}\\cdot\(\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\\bm\{v\}\)\-\\nabla\_\{\\bm\{v\}\}\\cdot\(\\bm\{a\}\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\)\(36\)This completes the proof\. โˆŽ

### A\.4Proof of Theorem 4\.5

###### Proof\.

Consider the gradient of the AFM Loss and CAFM Loss:

โˆ‡๐œฝโ„’AFM\\displaystyle\\nabla\_\{\\bm\{\\theta\}\}\\mathcal\{L\}\_\{\\text\{AFM\}\}=โˆ‡ฮธ๐”ผฯtโ€‹\(๐’™,๐’—\)โ€‹\(โ€–๐’‚ฮธโ€‹\(๐’™,๐’—,t\)โ€–2โˆ’2โ€‹๐’‚ฮธTโ€‹\(๐’™,๐’—,t\)โ€‹๐’‚โ€‹\(๐’™,๐’—,t\)\)\\displaystyle=\\nabla\_\{\\theta\}\\mathbb\{E\}\_\{\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\}\(\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{x\},\\bm\{v\},t\)\\\|^\{2\}\-2\\bm\{a\}\_\{\\theta\}^\{T\}\(\\bm\{x\},\\bm\{v\},t\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\)\)\(37\)โˆ‡๐œฝโ„’CAFM\\displaystyle\\nabla\_\{\\bm\{\\theta\}\}\\mathcal\{L\}\_\{\\text\{CAFM\}\}=โˆ‡ฮธ๐”ผฯtโ€‹\(๐’™,๐’—\|๐’›\)โ€‹qโ€‹\(๐’›\)โ€‹\(โ€–๐’‚ฮธโ€‹\(๐’™,๐’—,t\)โ€–2โˆ’2โ€‹๐’‚ฮธTโ€‹\(๐’™,๐’—,t\)โ€‹๐’‚โ€‹\(๐’™,๐’—,t\|๐’›\)\)\\displaystyle=\\nabla\_\{\\theta\}\\mathbb\{E\}\_\{\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\)q\(\\bm\{z\}\)\}\(\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{x\},\\bm\{v\},t\)\\\|^\{2\}\-2\\bm\{a\}\_\{\\theta\}^\{T\}\(\\bm\{x\},\\bm\{v\},t\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\|\\bm\{z\}\)\)\(38\)For the first term ofโˆ‡๐œฝโ„’AFM\\nabla\_\{\\bm\{\\theta\}\}\\mathcal\{L\}\_\{\\text\{AFM\}\}:

๐”ผฯtโ€‹\(๐’™,๐’—\)โ€‹โ€–๐’‚ฮธโ€‹\(๐’™,๐’—,t\)โ€–2\\displaystyle\\mathbb\{E\}\_\{\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\}\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{x\},\\bm\{v\},t\)\\\|^\{2\}=โˆซฯtโ€‹\(๐’™,๐’—\)โ€‹โ€–๐’‚ฮธโ€‹\(๐’™,๐’—,t\)โ€–2โ€‹๐‘‘๐’™โ€‹๐‘‘๐’—\\displaystyle=\\int\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{x\},\\bm\{v\},t\)\\\|^\{2\}\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{v\}=โˆซฯtโ€‹\(๐’™,๐’—\|z\)โ€‹qโ€‹\(๐’›\)โ€‹โ€–๐’‚ฮธโ€‹\(๐’™,๐’—,t\)โ€–2โ€‹๐‘‘๐’™โ€‹๐‘‘๐’—โ€‹๐‘‘z\\displaystyle=\\int\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|z\)q\(\\bm\{z\}\)\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{x\},\\bm\{v\},t\)\\\|^\{2\}\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{v\}\\mathrm\{d\}z=๐”ผฯtโ€‹\(๐’™,๐’—\|๐’›\)โ€‹qโ€‹\(๐’›\)โ€‹โ€–๐’‚ฮธโ€‹\(๐’™,๐’—,t\)โ€–2\\displaystyle=\\mathbb\{E\}\_\{\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\)q\(\\bm\{z\}\)\}\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{x\},\\bm\{v\},t\)\\\|^\{2\}\(39\)For the second term:

๐”ผฯtโ€‹\(๐’™,๐’—\)โ€‹\[๐’‚ฮธTโ€‹\(๐’™,๐’—,t\)โ€‹๐’‚โ€‹\(๐’™,๐’—,t\)\]\\displaystyle\\mathbb\{E\}\_\{\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\}\[\\bm\{a\}\_\{\\theta\}^\{T\}\(\\bm\{x\},\\bm\{v\},t\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\)\]=โˆซฯtโ€‹\(๐’™,๐’—\)โ€‹๐’‚ฮธTโ€‹\(๐’™,๐’—,t\)โ€‹๐’‚โ€‹\(๐’™,๐’—,t\)โ€‹๐‘‘๐’™โ€‹๐‘‘๐’—\\displaystyle=\\int\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\\bm\{a\}\_\{\\theta\}^\{T\}\(\\bm\{x\},\\bm\{v\},t\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{v\}=โˆซฯtโ€‹\(๐’™,๐’—\)โ€‹๐’‚ฮธTโ€‹\(๐’™,๐’—,t\)โ€‹\(โˆซqโก\(๐’›\)โ€‹๐’‚โก\(๐’™,๐’—,t\|๐’›\)โ€‹ฯtโ€‹\(๐’™,๐’—\|๐’›\)ฯtโ€‹\(๐’™,๐’—\)โ€‹๐‘‘z\)โ€‹๐‘‘๐’™โ€‹๐‘‘๐’—\\displaystyle=\\int\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\\bm\{a\}\_\{\\theta\}^\{T\}\(\\bm\{x\},\\bm\{v\},t\)\\left\(\\int q\(\\bm\{z\}\)\\dfrac\{\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\|\\bm\{z\}\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\)\}\{\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\)\}\\mathrm\{d\}z\\right\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{v\}=โˆซโˆซโก๐’‚ฮธTโ€‹\(๐’™,๐’—,t\)โ€‹๐’‚โ€‹\(๐’™,๐’—,t\|๐’›\)โ€‹ฯtโ€‹\(๐’™,๐’—\|๐’›\)โ€‹qโ€‹\(z\)โ€‹๐‘‘๐’™โ€‹๐‘‘๐’—โ€‹๐‘‘๐’›\\displaystyle=\\int\\int\\bm\{a\}\_\{\\theta\}^\{T\}\(\\bm\{x\},\\bm\{v\},t\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\|\\bm\{z\}\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\)q\(z\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{v\}\\mathrm\{d\}\\bm\{z\}=๐”ผฯtโ€‹\(๐’™,๐’—\|๐’›\)โ€‹qโ€‹\(๐’›\)โ€‹\(๐’‚ฮธTโ€‹\(๐’™,๐’—,t\)โ€‹๐’‚โ€‹\(๐’™,๐’—,t\|๐’›\)\)\\displaystyle=\\mathbb\{E\}\_\{\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\)q\(\\bm\{z\}\)\}\(\\bm\{a\}\_\{\\theta\}^\{T\}\(\\bm\{x\},\\bm\{v\},t\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\|\\bm\{z\}\)\)\(40\)โˆŽ

### A\.5Proof of Theorem 4\.6

###### Proof\.

Consider a given timetit\_\{i\}\. Based on the expressions for๐’Ž๐’™,๐’Ž๐’—\\bm\{m\}\_\{\\bm\{x\}\},\\bm\{m\}\_\{\\bm\{v\}\}, at this moment we exactly have:

ฯtiโ€‹\(๐’™,๐’—\|๐’›\)=๐’ฉโก\(๐’™\|๐’™i,ฯƒx2โ€‹๐‘ฐ\)โ‹…๐’ฉโก\(๐’—\|๐’—i,ฯƒv2โ€‹๐‘ฐ\)\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\)=\\mathcal\{N\}\(\\bm\{x\}\|\\bm\{x\}\_\{i\},\\sigma\_\{x\}^\{2\}\\bm\{I\}\)\\cdot\\mathcal\{N\}\(\\bm\{v\}\|\\bm\{v\}\_\{i\},\\sigma\_\{v\}^\{2\}\\bm\{I\}\)\(41\)
For simplicity, letdโ€‹๐’”j=dโ€‹๐’™jโ€‹dโ€‹๐’—j\\mathrm\{d\}\\bm\{s\}\_\{j\}=\\mathrm\{d\}\\bm\{x\}\_\{j\}\\mathrm\{d\}\\bm\{v\}\_\{j\}Therefore, the Marginal Probability Path is:

ฯtiโ€‹\(๐’™,๐’—\)\\displaystyle\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\)=โˆซฯtiโ€‹\(๐’™,๐’—\|๐’›\)โ€‹qโ€‹\(๐’›\)โ€‹๐‘‘๐’›\\displaystyle=\\int\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\)q\(\\bm\{z\}\)\\mathrm\{d\}\\bm\{z\}=โˆซฯti\(๐’™,๐’—\|\(๐’™i,๐’—i\)\)ฮผ0\(๐’™0,๐’—0\)โ‹…โˆj=0Kโˆ’1๐’ฆjโ†’j\+1โˆ—\[\(๐’™j,๐’—j\),\(๐’™j\+1,๐’—j\+1\)\]d๐’”0โ‹ฏd๐’”K\\displaystyle=\\int\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\|\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\)\)\\mu\_\{0\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\)\\cdot\\prod\_\{j=0\}^\{K\-1\}\\mathcal\{K\}^\{\*\}\_\{j\\rightarrow j\+1\}\[\(\\bm\{x\}\_\{j\},\\bm\{v\}\_\{j\}\),\(\\bm\{x\}\_\{j\+1\},\\bm\{v\}\_\{j\+1\}\)\]\\mathrm\{d\}\\bm\{s\}\_\{0\}\\cdots\\mathrm\{d\}\\bm\{s\}\_\{K\}\(42\)Using the definition of the propagator, we know that:

โˆซ๐’ฆjโ†’j\+1โˆ—โ€‹\[\(๐’™j,๐’—j\),\(๐’™j\+1,๐’—j\+1\)\]โ€‹dโ€‹๐’™j\+1โ€‹dโ€‹๐’—j\+1=1\\int\\mathcal\{K\}^\{\*\}\_\{j\\rightarrow j\+1\}\[\(\\bm\{x\}\_\{j\},\\bm\{v\}\_\{j\}\),\(\\bm\{x\}\_\{j\+1\},\\bm\{v\}\_\{j\+1\}\)\]\\mathrm\{d\}\\bm\{x\}\_\{j\+1\}\\mathrm\{d\}\\bm\{v\}\_\{j\+1\}=1\(43\)
Thus, we can integrate out the future terms \(j\>ij\>i\) first , and then iteratively collapse the past terms \(j<ij<i\):

ฯtiโ€‹\(๐’™,๐’—\)\\displaystyle\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\)=โˆซฯti\(๐’™,๐’—\|\(๐’™i,๐’—i\)\)ฮผ0\(๐’™0,๐’—0\)โ‹…โˆj=0Kโˆ’1๐’ฆjโ†’j\+1โˆ—\[\(๐’™j,๐’—j\),\(๐’™j\+1,๐’—j\+1\)\]d๐’”0โ‹ฏd๐’”K\\displaystyle=\\int\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\|\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\)\)\\mu\_\{0\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\)\\cdot\\prod\_\{j=0\}^\{K\-1\}\\mathcal\{K\}^\{\*\}\_\{j\\rightarrow j\+1\}\[\(\\bm\{x\}\_\{j\},\\bm\{v\}\_\{j\}\),\(\\bm\{x\}\_\{j\+1\},\\bm\{v\}\_\{j\+1\}\)\]\\mathrm\{d\}\\bm\{s\}\_\{0\}\\cdots\\mathrm\{d\}\\bm\{s\}\_\{K\}=โˆซฯti\(๐’™,๐’—\|\(๐’™i,๐’—i\)\)ฮผ0\(๐’™0,๐’—0\)โ‹…โˆj=0iโˆ’1๐’ฆjโ†’j\+1โˆ—\[\(๐’™j,๐’—j\),\(๐’™j\+1,๐’—j\+1\)\]d๐’”0โ‹ฏd๐’”i\\displaystyle=\\int\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\|\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\)\)\\mu\_\{0\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\)\\cdot\\prod\_\{j=0\}^\{i\-1\}\\mathcal\{K\}^\{\*\}\_\{j\\rightarrow j\+1\}\[\(\\bm\{x\}\_\{j\},\\bm\{v\}\_\{j\}\),\(\\bm\{x\}\_\{j\+1\},\\bm\{v\}\_\{j\+1\}\)\]\\mathrm\{d\}\\bm\{s\}\_\{0\}\\cdots\\mathrm\{d\}\\bm\{s\}\_\{i\}=โˆซฯti\(๐’™,๐’—\|\(๐’™i,๐’—i\)\)ฯ€0โ†’1โˆ—\[\(๐’™0,๐’—0\),\(๐’™1,๐’—1\)\]โˆj=1iโˆ’1๐’ฆjโ†’j\+1โˆ—\[\(๐’™j,๐’—j\),\(๐’™j\+1,๐’—j\+1\)\]d๐’”0โ‹ฏd๐’”i\\displaystyle=\\int\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\|\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\)\)\\pi^\{\*\}\_\{0\\rightarrow 1\}\[\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\),\(\\bm\{x\}\_\{1\},\\bm\{v\}\_\{1\}\)\]\\prod\_\{j=1\}^\{i\-1\}\\mathcal\{K\}^\{\*\}\_\{j\\rightarrow j\+1\}\[\(\\bm\{x\}\_\{j\},\\bm\{v\}\_\{j\}\),\(\\bm\{x\}\_\{j\+1\},\\bm\{v\}\_\{j\+1\}\)\]\\mathrm\{d\}\\bm\{s\}\_\{0\}\\cdots\\mathrm\{d\}\\bm\{s\}\_\{i\}=โˆซฯti\(๐’™,๐’—\|\(๐’™i,๐’—i\)\)ฮผ1\(๐’™1,๐’—1\)โˆj=1iโˆ’1๐’ฆjโ†’j\+1โˆ—\[\(๐’™j,๐’—j\),\(๐’™j\+1,๐’—j\+1\)\]d๐’”1โ‹ฏd๐’”i\\displaystyle=\\int\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\|\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\)\)\\mu\_\{1\}\(\\bm\{x\}\_\{1\},\\bm\{v\}\_\{1\}\)\\prod\_\{j=1\}^\{i\-1\}\\mathcal\{K\}^\{\*\}\_\{j\\rightarrow j\+1\}\[\(\\bm\{x\}\_\{j\},\\bm\{v\}\_\{j\}\),\(\\bm\{x\}\_\{j\+1\},\\bm\{v\}\_\{j\+1\}\)\]\\mathrm\{d\}\\bm\{s\}\_\{1\}\\cdots\\mathrm\{d\}\\bm\{s\}\_\{i\}โ‹ฎ\\displaystyle\\quad\\vdots=โˆซฯtiโ€‹\(๐’™,๐’—\|\(๐’™i,๐’—i\)\)โ€‹ฮผiโ€‹\(๐’™i,๐’—i\)โ€‹dโ€‹๐’”i\\displaystyle=\\int\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\|\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\)\)\\mu\_\{i\}\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\)\\mathrm\{d\}\\bm\{s\}\_\{i\}=โˆซ\(๐’ฉโก\(๐’™\|๐’™i,ฯƒx2โ€‹๐‘ฐ\)โ‹…๐’ฉโก\(๐’—\|๐’—i,ฯƒv2โ€‹๐‘ฐ\)\)โ€‹ฮผiโ€‹\(๐’™i,๐’—i\)โ€‹dโ€‹๐’”i\\displaystyle=\\int\(\\mathcal\{N\}\(\\bm\{x\}\|\\bm\{x\}\_\{i\},\\sigma\_\{x\}^\{2\}\\bm\{I\}\)\\cdot\\mathcal\{N\}\(\\bm\{v\}\|\\bm\{v\}\_\{i\},\\sigma\_\{v\}^\{2\}\\bm\{I\}\)\)\\mu\_\{i\}\(\\bm\{x\}\_\{i\},\\bm\{v\}\_\{i\}\)\\mathrm\{d\}\\bm\{s\}\_\{i\}=\(ฮผiโˆ—g\)โ€‹\(๐’™,๐’—\)\\displaystyle=\(\\mu\_\{i\}\*g\)\(\\bm\{x\},\\bm\{v\}\)\(44\)wheregโก\(๐’™,๐’—\)=\(๐’ฉโก\(๐’™\|๐ŸŽ,ฯƒx2โ€‹๐‘ฐ\)โ‹…๐’ฉโก\(๐’—\|๐ŸŽ,ฯƒv2โ€‹๐‘ฐ\)\)g\(\\bm\{x\},\\bm\{v\}\)=\(\\mathcal\{N\}\(\\bm\{x\}\|\\bm\{0\},\\sigma\_\{x\}^\{2\}\\bm\{I\}\)\\cdot\\mathcal\{N\}\(\\bm\{v\}\|\\bm\{0\},\\sigma\_\{v\}^\{2\}\\bm\{I\}\)\)\. When\(ฯƒ๐’™,ฯƒ๐’—\)โ†’๐ŸŽ\(\\sigma\_\{\\bm\{x\}\},\\sigma\_\{\\bm\{v\}\}\)\\rightarrow\\bm\{0\}, we haveฯtiโ€‹\(๐’™,๐’—\)โ†’ฮผiโ€‹\(๐’™,๐’—\)\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\)\\rightarrow\\mu\_\{i\}\(\\bm\{x\},\\bm\{v\}\)in distribution\.

Ifฮผiโ€‹\(๐’™,๐’—\)\\mu\_\{i\}\(\\bm\{x\},\\bm\{v\}\)is absolutely continuous with respect to the Lebesgue measure and admits a density inL1โ€‹\(โ„2โ€‹d\)L^\{1\}\(\\mathbb\{R\}^\{2d\}\), we can guarantee strong convergence in theL1L^\{1\}norm:limฯƒโ†’0โ€–ฯtiโ€‹\(๐’™,๐’—\)โˆ’ฮผiโ€‹\(๐’™,๐’—\)โ€–L1=0\\lim\_\{\\sigma\\to 0\}\\\|\\rho\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\)\-\\mu\_\{i\}\(\\bm\{x\},\\bm\{v\}\)\\\|\_\{L^\{1\}\}=0for alli=0,1,โ€ฆ,Ki=0,1,\.\.\.,K\.

โˆŽ

### A\.6Proof of Proposition 4\.7

First, we prove that the Marginal Probability Path constructed by TracingFlow converges to the optimal solution of the DOAT problem whenฯƒ๐’™โ†’0,ฯƒ๐’—โ†’0\\sigma\_\{\\bm\{x\}\}\\rightarrow 0,\\sigma\_\{\\bm\{v\}\}\\rightarrow 0\.

We investigate directly in the augmented space\. Let๐’”=\[๐’™,๐’—\]Tโˆˆ๐’ณร—๐’ฑ\\bm\{s\}=\[\\bm\{x\},\\bm\{v\}\]^\{T\}\\in\\mathcal\{X\}\\times\\mathcal\{V\}, and consider the cost of the DOAT problem\. The subsequent position of a particle starting from the initial point๐’”0\\bm\{s\}\_\{0\}can be represented by the flow mapฯ•tโ€‹\(๐’”0\)\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\. Therefore, the cost can be written as:

โ„’DOATโ€‹\(๐’‚,ฯ\)\\displaystyle\\mathcal\{L\}\_\{\\text\{DOAT\}\}\(\\bm\{a\},\\rho\)=โˆซt0tkโˆซ๐’ณร—๐’ฑ12โ€‹โ€–๐’‚โก\(๐’”,t\)โ€–2โ€‹ฯtโ€‹\(๐’”\)โ€‹๐‘‘๐’™โ€‹๐‘‘๐’—โ€‹๐‘‘t\\displaystyle=\\int\_\{t\_\{0\}\}^\{t\_\{k\}\}\\int\_\{\\mathcal\{X\}\\times\\mathcal\{V\}\}\\dfrac\{1\}\{2\}\\\|\\bm\{a\}\(\\bm\{s\},t\)\\\|^\{2\}\\rho\_\{t\}\(\\bm\{s\}\)\\mathrm\{d\}\\bm\{x\}\\mathrm\{d\}\\bm\{v\}\\mathrm\{d\}t=โˆซt0tkโˆซ๐’ณร—๐’ฑ12\|๐’‚โก\(ฯ•tโ€‹\(๐’”0\),t\)\|det2โก\(โˆ‚ฯ•โˆ‚๐’”0\)โˆ’1โ€‹ฯ0โ€‹\(๐’”0\)โ€‹det\(โˆ‚ฯ•โˆ‚๐’”0\)โ€‹dโ€‹๐’”0โ€‹๐‘‘t\\displaystyle=\\int\_\{t\_\{0\}\}^\{t\_\{k\}\}\\int\_\{\\mathcal\{X\}\\times\\mathcal\{V\}\}\\dfrac\{1\}\{2\}\\\|\\bm\{a\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\\|^\{2\}\\det\\left\(\\dfrac\{\\partial\\bm\{\\phi\}\}\{\\partial\\bm\{s\}\_\{0\}\}\\right\)^\{\-1\}\\rho\_\{0\}\(\\bm\{s\}\_\{0\}\)\\det\\left\(\\dfrac\{\\partial\\bm\{\\phi\}\}\{\\partial\\bm\{s\}\_\{0\}\}\\right\)\\mathrm\{d\}\\bm\{s\}\_\{0\}\\mathrm\{d\}t=โˆซ๐’ณร—๐’ฑฯ0โ€‹\(๐’”0\)โ€‹dโ€‹๐’”0โ€‹โˆซt0tk12โ€‹โ€–๐’‚โก\(ฯ•tโ€‹\(๐’”0\),t\)โ€–2โ€‹๐‘‘t\\displaystyle=\\int\_\{\\mathcal\{X\}\\times\\mathcal\{V\}\}\\rho\_\{0\}\(\\bm\{s\}\_\{0\}\)\\mathrm\{d\}\\bm\{s\}\_\{0\}\\int\_\{t\_\{0\}\}^\{t\_\{k\}\}\\dfrac\{1\}\{2\}\\\|\\bm\{a\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\\|^\{2\}\\mathrm\{d\}t=โˆซฮฉ\(โˆซt0tk12โ€‹โ€–๐œธโ€ฒโ€ฒโ€‹\(t\)โ€–2โ€‹๐‘‘t\)โ€‹๐‘‘โ„™โ€‹\[๐œธ\]\\displaystyle=\\int\_\{\\Omega\}\\left\(\\int\_\{t\_\{0\}\}^\{t\_\{k\}\}\\dfrac\{1\}\{2\}\\\|\\bm\{\\gamma\}^\{\\prime\\prime\}\(t\)\\\|^\{2\}\\mathrm\{d\}t\\right\)\\mathrm\{d\}\\mathbb\{P\}\[\\bm\{\\gamma\}\]\(45\)
whereโ„™\\mathbb\{P\}is a probability measure in the path space, satisfying the constraint\(eti\)\#โ€‹โ„™=ฮผiโ€‹\(๐’™,๐’—\)\(e\_\{t\_\{i\}\}\)\_\{\\\#\}\\mathbb\{P\}=\\mu\_\{i\}\(\\bm\{x\},\\bm\{v\}\)\. Here,ete\_\{t\}is the evaluation mapetโ€‹\[๐œธ\]=๐œธโ€‹\(t\)e\_\{t\}\[\\bm\{\\gamma\}\]=\\bm\{\\gamma\}\(t\), and\(eti\)\#โ€‹โ„™\(e\_\{t\_\{i\}\}\)\_\{\\\#\}\\mathbb\{P\}denotes the push\-forward of the probability measure by the evaluation map\. According to[Section4\.1](https://arxiv.org/html/2608.21070#S4.SS1.SSS0.Px1), the minimum cost path connecting๐’”i\\bm\{s\}\_\{i\}and๐’”i\+1\\bm\{s\}\_\{i\+1\}is a cubic spline, and the minimum cost is๐’žiโ†’i\+1โ€‹\[๐’”i,๐’”i\+1\]\\mathcal\{C\}\_\{i\\rightarrow i\+1\}\[\\bm\{s\}\_\{i\},\\bm\{s\}\_\{i\+1\}\]\. Therefore, for any path๐œธ\\bm\{\\gamma\}:

โˆซtiti\+112โ€‹โ€–๐œธโ€ฒโ€ฒโ€‹\(t\)โ€–2โ€‹๐‘‘tโ‰ฅ๐’žiโ†’i\+1โ€‹\[๐’”i,๐’”i\+1\]\\int\_\{t\_\{i\}\}^\{t\_\{i\+1\}\}\\dfrac\{1\}\{2\}\\\|\\bm\{\\gamma\}^\{\\prime\\prime\}\(t\)\\\|^\{2\}\\mathrm\{d\}t\\geq\\mathcal\{C\}\_\{i\\rightarrow i\+1\}\[\\bm\{s\}\_\{i\},\\bm\{s\}\_\{i\+1\}\]\(46\)
Thus, the DOAT problem cost has a lower bound:

โ„’DOATโ€‹\(ฯ,๐’‚\)\\displaystyle\\mathcal\{L\}\_\{\\text\{DOAT\}\}\(\\rho,\\bm\{a\}\)=โˆซฮฉโˆ‘i=0Kโˆ’1\(โˆซtiti\+112โ€‹โ€–๐œธโ€ฒโ€ฒโ€‹\(t\)โ€–2โ€‹๐‘‘t\)โ€‹๐‘‘โ„™โ€‹\[๐œธ\]\\displaystyle=\\int\_\{\\Omega\}\\sum\_\{i=0\}^\{K\-1\}\\left\(\\int\_\{t\_\{i\}\}^\{t\_\{i\+1\}\}\\dfrac\{1\}\{2\}\\\|\\bm\{\\gamma\}^\{\\prime\\prime\}\(t\)\\\|^\{2\}\\mathrm\{d\}t\\right\)\\mathrm\{d\}\\mathbb\{P\}\[\\bm\{\\gamma\}\]โ‰ฅโˆซฮฉโˆ‘i=0Kโˆ’1๐’žiโ†’i\+1โ€‹\[๐’”i,๐’”i\+1\]โ€‹๐‘‘โ„™โ€‹\[๐œธ\]\\displaystyle\\geq\\int\_\{\\Omega\}\\sum\_\{i=0\}^\{K\-1\}\\mathcal\{C\}\_\{i\\rightarrow i\+1\}\[\\bm\{s\}\_\{i\},\\bm\{s\}\_\{i\+1\}\]\\mathrm\{d\}\\mathbb\{P\}\[\\bm\{\\gamma\}\]โ‰ฅโˆซโˆ‘i=0Kโˆ’1๐’žiโ†’i\+1โ€‹\[๐’”i,๐’”i\+1\]โ€‹dโ€‹ฯ€โˆ—โ€‹\(๐’”i,๐’”i\+1\)\\displaystyle\\geq\\int\\sum\_\{i=0\}^\{K\-1\}\\mathcal\{C\}\_\{i\\rightarrow i\+1\}\[\\bm\{s\}\_\{i\},\\bm\{s\}\_\{i\+1\}\]\\mathrm\{d\}\\pi^\{\*\}\(\\bm\{s\}\_\{i\},\\bm\{s\}\_\{i\+1\}\)\(47\)
whereฯ€โˆ—โ€‹\(๐’”i,๐’”i\+1\)\\pi^\{\*\}\(\\bm\{s\}\_\{i\},\\bm\{s\}\_\{i\+1\}\)is the Optimal Transport Plan for the SOAT problem betweenฮผโก\(๐’”i\)\\mu\(\\bm\{s\}\_\{i\}\)andฮผโก\(๐’”i\+1\)\\mu\(\\bm\{s\}\_\{i\+1\}\)\.

Consider the path constructed by TracingFlow\. In the case whereฯƒ๐’™,ฯƒ๐’—โ†’0\\sigma\_\{\\bm\{x\}\},\\sigma\_\{\\bm\{v\}\}\\rightarrow 0, we haveฯtโ€‹\(๐’™,๐’—\|๐’›\)=ฯtโ€‹\(๐’”\|๐’›\)โ†’ฮดโก\(๐’”โˆ’๐œธ๐’›โˆ—โ€‹\(t\)\)\\rho\_\{t\}\(\\bm\{x\},\\bm\{v\}\|\\bm\{z\}\)=\\rho\_\{t\}\(\\bm\{s\}\|\\bm\{z\}\)\\rightarrow\\delta\(\\bm\{s\}\-\\bm\{\\gamma\}^\{\*\}\_\{\\bm\{z\}\}\(t\)\), where๐œธ๐’›โˆ—\\bm\{\\gamma\}^\{\*\}\_\{\\bm\{z\}\}is the optimal single\-particle trajectory connecting๐’›=\[๐’”0,๐’”1,โ‹ฏ,๐’”K\]\\bm\{z\}=\[\\bm\{s\}\_\{0\},\\bm\{s\}\_\{1\},\\cdots,\\bm\{s\}\_\{K\}\]\. We calculate the transport cost of the constructed path:

โ„’TF\\displaystyle\\mathcal\{L\}\_\{\\text\{TF\}\}=โˆซq\(๐’›\)d๐’›โˆ‘i=0Kโˆ’1โˆซtiti\+1โˆฅ๐œธโˆ—๐’›โ€ฒโ€ฒ\(t\)โˆฅ2dt\\displaystyle=\\int q\(\\bm\{z\}\)\\mathrm\{d\}\\bm\{z\}\\sum\_\{i=0\}^\{K\-1\}\\int\_\{t\_\{i\}\}^\{t\_\{i\+1\}\}\\\|\\bm\{\\gamma\}\{\{\}^\{\\prime\\prime\}\}\_\{\\bm\{z\}\}^\{\*\}\(t\)\\\|^\{2\}\\mathrm\{d\}t=โˆซqโก\(๐’›\)โ€‹๐‘‘๐’›โ€‹โˆ‘i=0Kโˆ’1๐’žiโ†’i\+1โ€‹\[๐’”i,๐’”i\+1\]\\displaystyle=\\int q\(\\bm\{z\}\)\\mathrm\{d\}\\bm\{z\}\\sum\_\{i=0\}^\{K\-1\}\\mathcal\{C\}\_\{i\\rightarrow i\+1\}\[\\bm\{s\}\_\{i\},\\bm\{s\}\_\{i\+1\}\]=โˆซฮผ0โ€‹\(๐’”0\)โ€‹โˆi=0Kโˆ’1๐’ฆiโ†’i\+1โˆ—โ€‹\[๐’”i,๐’”i\+1\]โ€‹\(โˆ‘j=0Kโˆ’1๐’žjโ†’j\+1โ€‹\[๐’”j,๐’”j\+1\]\)โ€‹๐‘‘๐’›\\displaystyle=\\int\\mu\_\{0\}\(\\bm\{s\}\_\{0\}\)\\prod\_\{i=0\}^\{K\-1\}\\mathcal\{K\}^\{\*\}\_\{i\\rightarrow i\+1\}\[\\bm\{s\}\_\{i\},\\bm\{s\}\_\{i\+1\}\]\\left\(\\sum\_\{j=0\}^\{K\-1\}\\mathcal\{C\}\_\{j\\rightarrow j\+1\}\[\\bm\{s\}\_\{j\},\\bm\{s\}\_\{j\+1\}\]\\right\)\\mathrm\{d\}\\bm\{z\}\(48\)
Here, using the definition of the propagator:

โˆซฮผ0โ€‹\(๐’”0\)โ€‹โˆi=0Kโˆ’1๐’ฆiโ†’i\+1โˆ—โ€‹\[๐’”i,๐’”i\+1\]โ€‹\(๐’žjโ†’j\+1โ€‹\[๐’”j,๐’”j\+1\]\)โ€‹๐’…๐’›=โˆซฯ€jโ†’j\+1โˆ—โ€‹\[๐’”j,๐’”j\+1\]โ€‹๐’žjโ†’j\+1โ€‹\[๐’”j,๐’”j\+1\]โ€‹dโ€‹๐’”jโ€‹dโ€‹๐’”j\+1\\begin\{split\}&\\int\\mu\_\{0\}\(\\bm\{s\}\_\{0\}\)\\prod\_\{i=0\}^\{K\-1\}\\mathcal\{K\}^\{\*\}\_\{i\\rightarrow i\+1\}\[\\bm\{s\}\_\{i\},\\bm\{s\}\_\{i\+1\}\]\\left\(\\mathcal\{C\}\_\{j\\rightarrow j\+1\}\[\\bm\{s\}\_\{j\},\\bm\{s\}\_\{j\+1\}\]\\right\)\\mathrm\{d\}\\bm\{z\}\\\\ &\\qquad=\\int\\pi^\{\*\}\_\{j\\rightarrow j\+1\}\[\\bm\{s\}\_\{j\},\\bm\{s\}\_\{j\+1\}\]\\mathcal\{C\}\_\{j\\rightarrow j\+1\}\[\\bm\{s\}\_\{j\},\\bm\{s\}\_\{j\+1\}\]\\mathrm\{d\}\\bm\{s\}\_\{j\}\\mathrm\{d\}\\bm\{s\}\_\{j\+1\}\\end\{split\}\(49\)
Therefore:

โ„’TF=โˆ‘i=0Kโˆ’1โˆซฯ€iโ†’i\+1โˆ—โ€‹\[๐’”i,๐’”i\+1\]โ€‹๐’žiโ†’i\+1โ€‹\[๐’”i,๐’”i\+1\]โ€‹dโ€‹๐’”iโ€‹dโ€‹๐’”i\+1\\mathcal\{L\}\_\{\\text\{TF\}\}=\\sum\_\{i=0\}^\{K\-1\}\\int\\pi^\{\*\}\_\{i\\rightarrow i\+1\}\[\\bm\{s\}\_\{i\},\\bm\{s\}\_\{i\+1\}\]\\mathcal\{C\}\_\{i\\rightarrow i\+1\}\[\\bm\{s\}\_\{i\},\\bm\{s\}\_\{i\+1\}\]\\mathrm\{d\}\\bm\{s\}\_\{i\}\\mathrm\{d\}\\bm\{s\}\_\{i\+1\}\(50\)This coincides exactly with the lower bound of the DOAT problem cost derived above\. Thus, the path constructed by TracingFlow is exactly the optimal solution to the DOAT problem in the case whereฯƒ๐’™,ฯƒ๐’—โ†’0\\sigma\_\{\\bm\{x\}\},\\sigma\_\{\\bm\{v\}\}\\rightarrow 0\.

Next, we prove that under the construction of TracingFlow, the acceleration field๐’‚โก\(๐’”,t\)\\bm\{a\}\(\\bm\{s\},t\)is single\-valued with respect to๐’”\\bm\{s\}andtt\. By[Section4\.1](https://arxiv.org/html/2608.21070#S4.SS1.SSS0.Px2), under the Optimal Transport Plan, the trajectories of flow maps starting from different points must not intersect\. This implies that for a given timett, any point๐’”\\bm\{s\}in the augmented space is traversed by at most one flow map trajectory\. Therefore, the acceleration field๐’‚โก\(๐’”,t\)\\bm\{a\}\(\\bm\{s\},t\)possesses single\-valuedness\. In other words, the solution constructed by TracingFlow can indeed be expressed by a single optimal acceleration field๐’‚โก\(๐’”,t\)\\bm\{a\}\(\\bm\{s\},t\)\.

### A\.7Proof of Proposition 4\.8

###### Proof\.

Letฮณ:\[0,T\]โ†’โ„d\\gamma:\[0,T\]\\to\\mathbb\{R\}^\{d\}be a path that isC1C^\{1\}on\[0,T\]\[0,T\]and piecewiseC2C^\{2\}on each interval\[tjโˆ’1,tj\]\[t\_\{j\-1\},t\_\{j\}\]\.

The optimal solution is given by a piecewise cubic polynomial:

ฮณโก\(tjโˆ’1\+ฯ„\)=\(๐’—jโˆ’1\+๐’—j\)โ€‹ฮ”โ€‹tjโˆ’2โ€‹\(๐’™jโˆ’๐’™jโˆ’1\)\(ฮ”โ€‹tj\)3โ€‹ฯ„3\+3โ€‹\(๐’™jโˆ’๐’™jโˆ’1\)โˆ’\(2โ€‹๐’—jโˆ’1\+๐’—j\)โ€‹ฮ”โ€‹tj\(ฮ”โ€‹tj\)2โ€‹ฯ„2\+๐’—jโˆ’1โ€‹ฯ„\+๐’™jโˆ’1,\\begin\{split\}\\gamma\(t\_\{j\-1\}\+\\tau\)&=\\frac\{\(\\bm\{v\}\_\{j\-1\}\+\\bm\{v\}\_\{j\}\)\\Delta t\_\{j\}\-2\(\\bm\{x\}\_\{j\}\-\\bm\{x\}\_\{j\-1\}\)\}\{\(\\Delta t\_\{j\}\)^\{3\}\}\\tau^\{3\}\\\\ &\\quad\+\\frac\{3\(\\bm\{x\}\_\{j\}\-\\bm\{x\}\_\{j\-1\}\)\-\(2\\bm\{v\}\_\{j\-1\}\+\\bm\{v\}\_\{j\}\)\\Delta t\_\{j\}\}\{\(\\Delta t\_\{j\}\)^\{2\}\}\\tau^\{2\}\+\\bm\{v\}\_\{j\-1\}\\tau\+\\bm\{x\}\_\{j\-1\},\\end\{split\}\(51\)forฯ„โˆˆ\[0,ฮ”โ€‹tj\]\\tau\\in\[0,\\Delta t\_\{j\}\], whereฮ”โ€‹tj=tjโˆ’tjโˆ’1\\Delta t\_\{j\}=t\_\{j\}\-t\_\{j\-1\},

According to Corollary[4\.1](https://arxiv.org/html/2608.21070#S4.SS1.SSS0.Px2), it suffices to minimize

๐’ฅโก\(๐’—0,๐’—1,โ€ฆ,๐’—K\)=\\displaystyle\\mathcal\{J\}\(\\bm\{v\}\_\{0\},\\bm\{v\}\_\{1\},\\dots,\\bm\{v\}\_\{K\}\)=2โˆ‘j=1K1\(ฮ”โ€‹tj\)3\{\(โˆฅ๐’—jโˆ’1โˆฅ2\+โŸจ๐’—jโˆ’1,๐’—jโŸฉ\+โˆฅ๐’—jโˆฅ2\)\(ฮ”tj\)2\\displaystyle 2\\sum\_\{j=1\}^\{K\}\\frac\{1\}\{\(\\Delta t\_\{j\}\)^\{3\}\}\\bigg\\\{\(\\\|\\bm\{v\}\_\{j\-1\}\\\|^\{2\}\+\\langle\\bm\{v\}\_\{j\-1\},\\bm\{v\}\_\{j\}\\rangle\+\\\|\\bm\{v\}\_\{j\}\\\|^\{2\}\)\(\\Delta t\_\{j\}\)^\{2\}\+3โŸจ๐’—jโˆ’1\+๐’—j,xjโˆ’1โˆ’xjโŸฉฮ”tj\+3โˆฅxjโˆ’1โˆ’xjโˆฅ2\}\\displaystyle\\qquad\\qquad\\qquad\+3\\langle\\bm\{v\}\_\{j\-1\}\+\\bm\{v\}\_\{j\},x\_\{j\-1\}\-x\_\{j\}\\rangle\\Delta t\_\{j\}\+3\\\|x\_\{j\-1\}\-x\_\{j\}\\\|^\{2\}\\bigg\\\}\(52\)๐’ฅโ‰ฅ0\\mathcal\{J\}\\geq 0always holds, so the the formula above is a semi\-positive quadratic form\.

Take differentiation, we have

โˆ‡๐’—j๐’ฅ\\displaystyle\\nabla\_\{\\bm\{v\}\_\{j\}\}\\mathcal\{J\}=2\(๐’—jโˆ’1\+2โ€‹๐’—jฮ”โ€‹tj\+2โ€‹๐’—j\+๐’—j\+1ฮ”โ€‹tj\+1\)\+6\(๐’™jโˆ’1โˆ’๐’™j\(ฮ”โ€‹tj\)2\+๐’™jโˆ’๐’™j\+1\(ฮ”โ€‹tj\+1\)2\),forj=1,โ€ฆ,Kโˆ’1,\\displaystyle=2\\left\(\\frac\{\\bm\{v\}\_\{j\-1\}\+2\\bm\{v\}\_\{j\}\}\{\\Delta t\_\{j\}\}\+\\frac\{2\\bm\{v\}\_\{j\}\+\\bm\{v\}\_\{j\+1\}\}\{\\Delta t\_\{j\+1\}\}\\right\)\+6\\left\(\\frac\{\\bm\{x\}\_\{j\-1\}\-\\bm\{x\}\_\{j\}\}\{\(\\Delta t\_\{j\}\)^\{2\}\}\+\\frac\{\\bm\{x\}\_\{j\}\-\\bm\{x\}\_\{j\+1\}\}\{\(\\Delta t\_\{j\+1\}\)^\{2\}\}\\right\),\\quad\\text\{for \}j=1,\\dots,K\-1,\(53\)โˆ‡๐’—0๐’ฅ\\displaystyle\\nabla\_\{\\bm\{v\}\_\{0\}\}\\mathcal\{J\}=2โ€‹\(2โ€‹๐’—0\+๐’—1\)ฮ”โ€‹t1\+6โ€‹\(๐’™0โˆ’๐’™1\)\(ฮ”โ€‹t1\)2,โˆ‡๐’—K๐’ฅ=2โ€‹\(๐’—Kโˆ’1\+2โ€‹๐’—K\)ฮ”โ€‹tK\+6โ€‹\(๐’™Kโˆ’1โˆ’๐’™K\)\(ฮ”โ€‹tK\)2\.\\displaystyle=\\frac\{2\(2\\bm\{v\}\_\{0\}\+\\bm\{v\}\_\{1\}\)\}\{\\Delta t\_\{1\}\}\+\\frac\{6\(\\bm\{x\}\_\{0\}\-\\bm\{x\}\_\{1\}\)\}\{\(\\Delta t\_\{1\}\)^\{2\}\},\\nabla\_\{\\bm\{v\}\_\{K\}\}\\mathcal\{J\}=\\frac\{2\(\\bm\{v\}\_\{K\-1\}\+2\\bm\{v\}\_\{K\}\)\}\{\\Delta t\_\{K\}\}\+\\frac\{6\(\\bm\{x\}\_\{K\-1\}\-\\bm\{x\}\_\{K\}\)\}\{\(\\Delta t\_\{K\}\)^\{2\}\}\.\(54\)
Set the gradients to zero and rearranging the terms to separate the velocities๐’—\\bm\{v\}and positions๐’™\\bm\{x\}, we obtain:

Forj=0j=0:

2ฮ”โ€‹t1โ€‹๐’—0\+1ฮ”โ€‹t1โ€‹๐’—1=3โ€‹\(๐’™1โˆ’๐’™0\)\(ฮ”โ€‹t1\)2\.\\frac\{2\}\{\\Delta t\_\{1\}\}\\bm\{v\}\_\{0\}\+\\frac\{1\}\{\\Delta t\_\{1\}\}\\bm\{v\}\_\{1\}=\\frac\{3\(\\bm\{x\}\_\{1\}\-\\bm\{x\}\_\{0\}\)\}\{\(\\Delta t\_\{1\}\)^\{2\}\}\.\(55\)Forj=1,โ€ฆ,Kโˆ’1j=1,\\dots,K\-1:

1ฮ”โ€‹tjโ€‹๐’—jโˆ’1\+2โ€‹\(1ฮ”โ€‹tj\+1ฮ”โ€‹tj\+1\)โ€‹๐’—j\+1ฮ”โ€‹tj\+1โ€‹๐’—j\+1=3โ€‹\(๐’™jโˆ’๐’™jโˆ’1\(ฮ”โ€‹tj\)2\+๐’™j\+1โˆ’๐’™j\(ฮ”โ€‹tj\+1\)2\)\.\\frac\{1\}\{\\Delta t\_\{j\}\}\\bm\{v\}\_\{j\-1\}\+2\\left\(\\frac\{1\}\{\\Delta t\_\{j\}\}\+\\frac\{1\}\{\\Delta t\_\{j\+1\}\}\\right\)\\bm\{v\}\_\{j\}\+\\frac\{1\}\{\\Delta t\_\{j\+1\}\}\\bm\{v\}\_\{j\+1\}=3\\left\(\\frac\{\\bm\{x\}\_\{j\}\-\\bm\{x\}\_\{j\-1\}\}\{\(\\Delta t\_\{j\}\)^\{2\}\}\+\\frac\{\\bm\{x\}\_\{j\+1\}\-\\bm\{x\}\_\{j\}\}\{\(\\Delta t\_\{j\+1\}\)^\{2\}\}\\right\)\.\(56\)Forj=Kj=K:

1ฮ”โ€‹tKโ€‹๐’—Kโˆ’1\+2ฮ”โ€‹tKโ€‹๐’—K=3โ€‹\(๐’™Kโˆ’๐’™Kโˆ’1\)\(ฮ”โ€‹tK\)2\.\\frac\{1\}\{\\Delta t\_\{K\}\}\\bm\{v\}\_\{K\-1\}\+\\frac\{2\}\{\\Delta t\_\{K\}\}\\bm\{v\}\_\{K\}=\\frac\{3\(\\bm\{x\}\_\{K\}\-\\bm\{x\}\_\{K\-1\}\)\}\{\(\\Delta t\_\{K\}\)^\{2\}\}\.\(57\)
DenoteV=\[๐’—0,๐’—1,โ€ฆ,๐’—K\]โˆˆโ„dร—\(K\+1\)V=\[\\bm\{v\}\_\{0\},\\bm\{v\}\_\{1\},\\dots,\\bm\{v\}\_\{K\}\]\\in\\mathbb\{R\}^\{d\\times\(K\+1\)\}and writing this system in matrix formVโ€‹A=BVA=B, where

A=\[2ฮ”โ€‹t11ฮ”โ€‹t10โ‹ฏ001ฮ”โ€‹t12โ€‹\(1ฮ”โ€‹t1\+1ฮ”โ€‹t2\)1ฮ”โ€‹t2โ‹ฏ0001ฮ”โ€‹t22โ€‹\(1ฮ”โ€‹t2\+1ฮ”โ€‹t3\)โ‹ฑโ‹ฑโ‹ฑ1ฮ”โ€‹tKโˆ’1000โ‹ฏ1ฮ”โ€‹tKโˆ’12โ€‹\(1ฮ”โ€‹tKโˆ’1\+1ฮ”โ€‹tK\)1ฮ”โ€‹tK00โ‹ฏ01ฮ”โ€‹tK2ฮ”โ€‹tK\]โˆˆโ„\(K\+1\)ร—\(K\+1\)A=\\begin\{bmatrix\}\\frac\{2\}\{\\Delta t\_\{1\}\}&\\frac\{1\}\{\\Delta t\_\{1\}\}&0&\\cdots&0&0\\\\ \\frac\{1\}\{\\Delta t\_\{1\}\}&2\\left\(\\frac\{1\}\{\\Delta t\_\{1\}\}\+\\frac\{1\}\{\\Delta t\_\{2\}\}\\right\)&\\frac\{1\}\{\\Delta t\_\{2\}\}&\\cdots&0&0\\\\ 0&\\frac\{1\}\{\\Delta t\_\{2\}\}&2\\left\(\\frac\{1\}\{\\Delta t\_\{2\}\}\+\\frac\{1\}\{\\Delta t\_\{3\}\}\\right\)&\\ddots&\\vdots&\\vdots\\\\ \\vdots&\\vdots&\\ddots&\\ddots&\\frac\{1\}\{\\Delta t\_\{K\-1\}\}&0\\\\ 0&0&\\cdots&\\frac\{1\}\{\\Delta t\_\{K\-1\}\}&2\\left\(\\frac\{1\}\{\\Delta t\_\{K\-1\}\}\+\\frac\{1\}\{\\Delta t\_\{K\}\}\\right\)&\\frac\{1\}\{\\Delta t\_\{K\}\}\\\\ 0&0&\\cdots&0&\\frac\{1\}\{\\Delta t\_\{K\}\}&\\frac\{2\}\{\\Delta t\_\{K\}\}\\end\{bmatrix\}\\in\\mathbb\{R\}^\{\(K\+1\)\\times\(K\+1\)\}\(58\)
B=\[3โ€‹\(๐’™1โˆ’๐’™0\)\(ฮ”โ€‹t1\)2,3โ€‹\(๐’™1โˆ’๐’™0\(ฮ”โ€‹t1\)2\+๐’™2โˆ’๐’™1\(ฮ”โ€‹t2\)2\),โ€ฆ,OPEN3โ€‹\(๐’™Kโˆ’1โˆ’๐’™Kโˆ’2\(ฮ”โ€‹tKโˆ’1\)2\+๐’™Kโˆ’๐’™Kโˆ’1\(ฮ”โ€‹tK\)2\),3โ€‹\(๐’™Kโˆ’๐’™Kโˆ’1\)\(ฮ”โ€‹tK\)2\]โˆˆโ„dร—\(K\+1\)\\begin\{split\}B=\\Bigg\[&\\frac\{3\(\\bm\{x\}\_\{1\}\-\\bm\{x\}\_\{0\}\)\}\{\(\\Delta t\_\{1\}\)^\{2\}\},3\\left\(\\frac\{\\bm\{x\}\_\{1\}\-\\bm\{x\}\_\{0\}\}\{\(\\Delta t\_\{1\}\)^\{2\}\}\+\\frac\{\\bm\{x\}\_\{2\}\-\\bm\{x\}\_\{1\}\}\{\(\\Delta t\_\{2\}\)^\{2\}\}\\right\),\\dots,\\\\ &3\\left\(\\frac\{\\bm\{x\}\_\{K\-1\}\-\\bm\{x\}\_\{K\-2\}\}\{\(\\Delta t\_\{K\-1\}\)^\{2\}\}\+\\frac\{\\bm\{x\}\_\{K\}\-\\bm\{x\}\_\{K\-1\}\}\{\(\\Delta t\_\{K\}\)^\{2\}\}\\right\),\\frac\{3\(\\bm\{x\}\_\{K\}\-\\bm\{x\}\_\{K\-1\}\)\}\{\(\\Delta t\_\{K\}\)^\{2\}\}\\Bigg\]\\in\\mathbb\{R\}^\{d\\times\(K\+1\)\}\\end\{split\}\(59\)The solution satisfiesฮณโ€ฒโ€‹\(tj\)=๐’—j\\gamma^\{\\prime\}\(t\_\{j\}\)=\\bm\{v\}\_\{j\}for allj=0,1,โ€ฆโ€‹K\.j=0,1,\.\.\.K\.

SinceAAis strictly diagonally dominant, the solutionVVis unique\. โˆŽ

### A\.8Proof of Theorem A\.1

###### Theorem A\.1\.

Let๐š๐›‰โ€‹\(๐ฑ,๐ฏ,t\)\\bm\{a\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},\\bm\{v\},t\)beLฮธL\_\{\\theta\}\-Lipschitz continuous of\(๐ฑ,๐ฏ\)\(\\bm\{x\},\\bm\{v\}\)\. The Wasserstein\-2 distance between the generated distributionฯ^ti\(pos\)โ€‹\(๐ฑ\)=โˆซฯ^tiโ€‹\(๐ฑ,๐ฏ\)โ€‹๐‘‘๐ฏ\\hat\{\\rho\}\_\{t\_\{i\}\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\)=\\int\\hat\{\\rho\}\_\{t\_\{i\}\}\(\\bm\{x\},\\bm\{v\}\)\\mathrm\{d\}\\bm\{v\}and the true distributionฮผi\(pos\)โ€‹\(๐ฑ\)\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\(\\bm\{x\}\)in the position space is bounded by

๐’ฒ22โ€‹\(ฯ^ti\(pos\),ฮผi\(pos\)\)โ‰ค2โ€‹eฮบโก\(tiโˆ’t0\)โ€‹\(โ„’v0\+\(tiโˆ’t0\)โ€‹โ„’AFM\|\[t0,ti\]\)\\mathcal\{W\}\_\{2\}^\{2\}\(\\hat\{\\rho\}\_\{t\_\{i\}\}^\{\(\\text\{pos\}\)\},\\mu\_\{i\}^\{\(\\text\{pos\}\)\}\)\\leq 2e^\{\\kappa\(t\_\{i\}\-t\_\{0\}\)\}\\left\(\\mathcal\{L\}\_\{v\_\{0\}\}\+\(t\_\{i\}\-t\_\{0\}\)\\mathcal\{L\}\_\{\\text\{AFM\}\}\|\_\{\[t\_\{0\},t\_\{i\}\]\}\\right\)\(60\)where the cumulative AFM loss and the initial velocity loss are defined as:

โ„’AFM\|\[t0,ti\]=๐”ผt,\(๐’™,๐’—\)โˆผฯtโˆฅ๐’‚๐œฝ\(๐’™,๐’—,t\)โˆ’๐’‚\(๐’™,๐’—,t\)โˆฅ2,โ„’v0=๐”ผ\(๐’™0,๐’—0\)โˆผฯt0โˆฅ๐’—0,๐ƒ\(๐’™0\)โˆ’๐’—0โˆฅ2\.\\displaystyle\\mathcal\{L\}\_\{\\text\{AFM\}\}\|\_\{\[t\_\{0\},t\_\{i\}\]\}=\\mathbb\{E\}\_\{t,\(\\bm\{x\},\\bm\{v\}\)\\sim\\rho\_\{t\}\}\\\|\\bm\{a\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},\\bm\{v\},t\)\-\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\)\\\|^\{2\},\\mathcal\{L\}\_\{v\_\{0\}\}=\\mathbb\{E\}\_\{\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\)\\sim\\rho\_\{t\_\{0\}\}\}\\\|\\bm\{v\}\_\{0,\\bm\{\\xi\}\}\(\\bm\{x\}\_\{0\}\)\-\\bm\{v\}\_\{0\}\\\|^\{2\}\.
andฮบ\\kappais a constant only depending onLฮธL\_\{\\theta\}\.

###### Proof\.

We address the problem directly within the augmented space๐’ณร—๐’ฑ\\mathcal\{X\}\\times\\mathcal\{V\}\. Let the system state be denoted by๐’”=\[๐’™,๐’—\]T\\bm\{s\}=\[\\bm\{x\},\\bm\{v\}\]^\{T\}, and letฮผiโ€‹\(๐’”\)\\mu\_\{i\}\(\\bm\{s\}\)represent the data distribution in this space\. We define a flow mapฯ•t:\[t0,tK\]ร—๐’ณร—๐’ฑโ†’๐’ณร—๐’ฑ\\bm\{\\phi\}\_\{t\}:\[t\_\{0\},t\_\{K\}\]\\times\\mathcal\{X\}\\times\\mathcal\{V\}\\rightarrow\\mathcal\{X\}\\times\\mathcal\{V\}, which is governed by the exact marginal acceleration field๐’‚โก\(๐’”,t\)\\bm\{a\}\(\\bm\{s\},t\)defined in[Equation13](https://arxiv.org/html/2608.21070#S4.E13)\.

Using\(ฯ•t\)\#\(\\bm\{\\phi\}\_\{t\}\)\_\{\\\#\}to denote the push\-forward operator, the time\-dependent density evolves asฯt=\(ฯ•t\)\#โ€‹ฯt0\\rho\_\{t\}=\(\\bm\{\\phi\}\_\{t\}\)\_\{\\\#\}\\rho\_\{t\_\{0\}\}, whereฯt0=ฮผ0\\rho\_\{t\_\{0\}\}=\\mu\_\{0\}\. Similarly, letฯ•t๐œฝ\\bm\{\\phi\}^\{\\bm\{\\theta\}\}\_\{t\}be the flow map induced by the learned acceleration field๐’‚๐œฝโ€‹\(๐’”,t\)\\bm\{a\}\_\{\\bm\{\\theta\}\}\(\\bm\{s\},t\)\. The generated distributionฯ^t\\hat\{\\rho\}\_\{t\}then evolves according toฯ^t=\(ฯ•t๐œฝ\)\#โ€‹ฯ^t0\\hat\{\\rho\}\_\{t\}=\(\\bm\{\\phi\}^\{\\bm\{\\theta\}\}\_\{t\}\)\_\{\\\#\}\\hat\{\\rho\}\_\{t\_\{0\}\}\. It is important to note thatฯ^t0โ‰ ฯt0\\hat\{\\rho\}\_\{t\_\{0\}\}\\neq\\rho\_\{t\_\{0\}\}, as the conditional distribution of the initial velocity is approximated via a neural network\.

We introduce a bridging measureฮฝt=\(ฯ•tฮธ\)\#โ€‹ฯt0\\nu\_\{t\}=\(\\bm\{\\phi\}^\{\\theta\}\_\{t\}\)\_\{\\\#\}\\rho\_\{t\_\{0\}\}, then the Wasserstein\-2 distance between joint distributions can be decomposed as

๐’ฒ22โ€‹\(ฯt,ฯ^t\)โ‰ค\(๐’ฒ2โ€‹\(ฯt,ฮฝt\)\+๐’ฒ2โ€‹\(ฮฝt,ฯ^t\)\)2โ‰ค2โ€‹\(๐’ฒ22โ€‹\(ฯt,ฮฝt\)\+๐’ฒ22โ€‹\(ฮฝt,ฯ^t\)\)\\mathcal\{W\}\_\{2\}^\{2\}\(\\rho\_\{t\},\\hat\{\\rho\}\_\{t\}\)\\leq\\left\(\{\\mathcal\{W\}\_\{2\}\(\\rho\_\{t\},\\nu\_\{t\}\)\}\+\{\\mathcal\{W\}\_\{2\}\(\\nu\_\{t\},\\hat\{\\rho\}\_\{t\}\)\}\\right\)^\{2\}\\leq 2\\left\(\\mathcal\{W\}\_\{2\}^\{2\}\(\\rho\_\{t\},\\nu\_\{t\}\)\+\\mathcal\{W\}\_\{2\}^\{2\}\(\\nu\_\{t\},\\hat\{\\rho\}\_\{t\}\)\\right\)\\\\\(61\)
Estimation of Term 1: DefineQt=โˆซฮผ0โ€‹\(๐’”0\)โ€‹โ€–ฯ•tฮธโ€‹\(๐’”0\)โˆ’ฯ•tโ€‹\(๐’”0\)โ€–2โ€‹dโ€‹๐’”0,Q\_\{t\}=\\int\\mu\_\{0\}\(\\bm\{s\}\_\{0\}\)\\\|\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\\\|^\{2\}\\mathrm\{d\}\\bm\{s\}\_\{0\},hereฮผ0=ฯt0\\mu\_\{0\}=\\rho\_\{t\_\{0\}\}according to definition\. Noting thatQt0=0Q\_\{t\_\{0\}\}=0since the starting points are identical\. DefineA=\[0I00\]A=\\begin\{bmatrix\}0&I\\\\ 0&0\\end\{bmatrix\}\.

Taking the time derivative, we have:

dโ€‹Qtdโ€‹t\\displaystyle\\frac\{\\mathrm\{d\}Q\_\{t\}\}\{\\mathrm\{d\}t\}=2โ€‹โˆซฮผ0โ€‹\(๐’”0\)โ€‹\(ฯ•tฮธโ€‹\(๐’”0\)โˆ’ฯ•tโ€‹\(๐’”0\)\)Tโ€‹\(dโ€‹ฯ•tฮธโ€‹\(๐’”0\)dโ€‹tโˆ’dโ€‹ฯ•tโ€‹\(๐’”0\)dโ€‹t\)โ€‹dโ€‹๐’”0\\displaystyle=2\\int\\mu\_\{0\}\(\\bm\{s\}\_\{0\}\)\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\)^\{T\}\\left\(\\frac\{\\mathrm\{d\}\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\}\{\\mathrm\{d\}t\}\-\\frac\{\\mathrm\{d\}\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\}\{\\mathrm\{d\}t\}\\right\)\\mathrm\{d\}\\bm\{s\}\_\{0\}=2โ€‹โˆซฮผ0โ€‹\(๐’”0\)โ€‹\(ฯ•tฮธโ€‹\(๐’”0\)โˆ’ฯ•tโ€‹\(๐’”0\)\)Tโ€‹\(Aโ€‹ฯ•tฮธโ€‹\(๐’”0\)โˆ’Aโ€‹ฯ•tโ€‹\(๐’”0\)\+\[๐ŸŽ๐’‚ฮธโ€‹\(ฯ•tฮธโ€‹\(๐’”0\),t\)\]โˆ’\[๐ŸŽ๐’‚โก\(ฯ•tโ€‹\(๐’”0\),t\)\]\)โ€‹dโ€‹๐’”0\\displaystyle=2\\int\\mu\_\{0\}\(\\bm\{s\}\_\{0\}\)\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\)^\{T\}\\left\(A\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-A\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\+\\begin\{bmatrix\}\\bm\{0\}\\\\ \\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\),t\)\\end\{bmatrix\}\-\\begin\{bmatrix\}\\bm\{0\}\\\\ \\bm\{a\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\end\{bmatrix\}\\right\)\\mathrm\{d\}\\bm\{s\}\_\{0\}=2โˆซฮผ0\(๐’”0\)\[\(ฯ•tฮธ\(๐’”0\)โˆ’ฯ•t\(๐’”0\)\)TA\(ฯ•tฮธ\(๐’”0\)โˆ’ฯ•t\(๐’”0\)\)\\displaystyle=2\\int\\mu\_\{0\}\(\\bm\{s\}\_\{0\}\)\\Bigg\[\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\)^\{T\}A\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\)\+\(ฯ•tฮธโ€‹\(๐’”0\)โˆ’ฯ•tโ€‹\(๐’”0\)\)Tโ€‹\[๐ŸŽ๐’‚ฮธโ€‹\(ฯ•tฮธโ€‹\(๐’”0\),t\)โˆ’๐’‚ฮธโ€‹\(ฯ•tโ€‹\(๐’”0\),t\)\]\\displaystyle\\quad\+\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\)^\{T\}\\begin\{bmatrix\}\\bm\{0\}\\\\ \\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\),t\)\-\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\end\{bmatrix\}\+\(ฯ•tฮธ\(๐’”0\)โˆ’ฯ•t\(๐’”0\)\)T\[๐ŸŽ๐’‚ฮธโ€‹\(ฯ•tโ€‹\(๐’”0\),t\)โˆ’๐’‚โก\(ฯ•tโ€‹\(๐’”0\),t\)\]\]d๐’”0\\displaystyle\\quad\+\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\)^\{T\}\\begin\{bmatrix\}\\bm\{0\}\\\\ \\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\-\\bm\{a\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\end\{bmatrix\}\\Bigg\]\\mathrm\{d\}\\bm\{s\}\_\{0\}\(62\)
We bound the three terms inside the integral separately\. Letฮ”โ€‹ฯ•tโ€‹\(๐’”0\)=ฯ•tฮธโ€‹\(๐’”0\)โˆ’ฯ•tโ€‹\(๐’”0\)\\Delta\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)=\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\.

1. 1\.Sinceโ€–Aโ€–โ‰ค1\\\|A\\\|\\leq 1, we haveฮ”โ€‹ฯ•tโ€‹\(๐’”0\)Tโ€‹Aโ€‹ฮ”โ€‹ฯ•tโ€‹\(๐’”0\)โ‰คโ€–ฮ”โ€‹ฯ•tโ€‹\(๐’”0\)โ€–2\\Delta\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)^\{T\}A\\Delta\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\\leq\\\|\\Delta\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\\\|^\{2\}\.
2. 2\.Using the Lipschitz continuity of the network๐’‚ฮธ\\bm\{a\}\_\{\\theta\}with constantLฮธL\_\{\\theta\}: ฮ”โ€‹ฯ•tโ€‹\(๐’”0\)Tโ€‹\[๐ŸŽ๐’‚ฮธโ€‹\(ฯ•tฮธโ€‹\(๐’”0\),t\)โˆ’๐’‚ฮธโ€‹\(ฯ•tโ€‹\(๐’”0\),t\)\]\\displaystyle\\Delta\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)^\{T\}\\begin\{bmatrix\}\\bm\{0\}\\\\ \\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\),t\)\-\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\end\{bmatrix\}โ‰คโ€–ฮ”โ€‹ฯ•tโ€‹\(๐’”0\)โ€–โ‹…โ€–๐’‚ฮธโ€‹\(ฯ•tฮธโ€‹\(๐’”0\),t\)โˆ’๐’‚ฮธโ€‹\(ฯ•tโ€‹\(๐’”0\),t\)โ€–\\displaystyle\\leq\\\|\\Delta\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\\\|\\cdot\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\),t\)\-\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\\|โ‰คLฮธโ€‹โ€–ฮ”โ€‹ฯ•tโ€‹\(๐’”0\)โ€–2\.\\displaystyle\\leq L\_\{\\theta\}\\\|\\Delta\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\\\|^\{2\}\.\(63\)
3. 3\.Using Cauchy\-Schwarz and the inequality2โ€‹xโ€‹yโ‰คx2\+y22xy\\leq x^\{2\}\+y^\{2\}: ฮ”โ€‹ฯ•tโ€‹\(๐’”0\)Tโ€‹\[๐ŸŽ๐’‚ฮธโ€‹\(ฯ•tโ€‹\(๐’”0\),t\)โˆ’๐’‚โก\(ฯ•tโ€‹\(๐’”0\),t\)\]\\displaystyle\\Delta\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)^\{T\}\\begin\{bmatrix\}\\bm\{0\}\\\\ \\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\-\\bm\{a\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\end\{bmatrix\}โ‰คโ€–ฮ”โ€‹ฯ•tโ€‹\(๐’”0\)โ€–โ‹…โ€–๐’‚ฮธโ€‹\(ฯ•tโ€‹\(๐’”0\),t\)โˆ’๐’‚โก\(ฯ•tโ€‹\(๐’”0\),t\)โ€–\\displaystyle\\leq\\\|\\Delta\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\\\|\\cdot\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\-\\bm\{a\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\\|โ‰ค12โ€‹โ€–ฮ”โ€‹ฯ•tโ€‹\(๐’”0\)โ€–2\+12โ€‹โ€–๐’‚ฮธโ€‹\(ฯ•tโ€‹\(๐’”0\),t\)โˆ’๐’‚โก\(ฯ•tโ€‹\(๐’”0\),t\)โ€–2\.\\displaystyle\\leq\\frac\{1\}\{2\}\\\|\\Delta\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\\\|^\{2\}\+\\frac\{1\}\{2\}\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\-\\bm\{a\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\\|^\{2\}\.\(64\)

Substituting these bounds back into the derivative:

dโ€‹Qtdโ€‹t\\displaystyle\\frac\{\\mathrm\{d\}Q\_\{t\}\}\{\\mathrm\{d\}t\}โ‰ค2โ€‹โˆซฮผ0โ€‹\(๐’”0\)โ€‹\[\(1\+Lฮธ\+12\)โ€‹โ€–ฮ”โ€‹ฯ•tโ€‹\(๐’”0\)โ€–2\+12โ€‹โ€–๐’‚ฮธโ€‹\(ฯ•tโ€‹\(๐’”0\),t\)โˆ’๐’‚โก\(ฯ•tโ€‹\(๐’”0\),t\)โ€–2\]โ€‹dโ€‹๐’”0\\displaystyle\\leq 2\\int\\mu\_\{0\}\(\\bm\{s\}\_\{0\}\)\\left\[\(1\+L\_\{\\theta\}\+\\frac\{1\}\{2\}\)\\\|\\Delta\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\\\|^\{2\}\+\\frac\{1\}\{2\}\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\-\\bm\{a\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\\|^\{2\}\\right\]\\mathrm\{d\}\\bm\{s\}\_\{0\}=\(2โ€‹Lฮธ\+3\)โ€‹Qt\+โˆซฮผ0โ€‹\(๐’”0\)โ€‹โ€–๐’‚ฮธโ€‹\(ฯ•tโ€‹\(๐’”0\),t\)โˆ’๐’‚โก\(ฯ•tโ€‹\(๐’”0\),t\)โ€–2โ€‹dโ€‹๐’”0\\displaystyle=\(2L\_\{\\theta\}\+3\)Q\_\{t\}\+\\int\\mu\_\{0\}\(\\bm\{s\}\_\{0\}\)\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\-\\bm\{a\}\(\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\),t\)\\\|^\{2\}\\mathrm\{d\}\\bm\{s\}\_\{0\}\(65\)
By Gronwallโ€™s inequality, sinceQt0=0Q\_\{t\_\{0\}\}=0:

Qtโ‰คe\(2โ€‹Lฮธ\+3\)โ€‹\(tโˆ’t0\)โ€‹โˆซt0t\(โˆซฮผ0โ€‹\(๐’”0\)โ€‹โ€–๐’‚ฮธโ€‹\(ฯ•ฯ„โ€‹\(๐’”0\),ฯ„\)โˆ’๐’‚โก\(ฯ•ฯ„โ€‹\(๐’”0\),ฯ„\)โ€–2โ€‹dโ€‹๐’”0\)โ€‹๐‘‘ฯ„Q\_\{t\}\\leq e^\{\(2L\_\{\\theta\}\+3\)\(t\-t\_\{0\}\)\}\\int\_\{t\_\{0\}\}^\{t\}\\left\(\\int\\mu\_\{0\}\(\\bm\{s\}\_\{0\}\)\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{\\tau\}\(\\bm\{s\}\_\{0\}\),\\tau\)\-\\bm\{a\}\(\\bm\{\\phi\}\_\{\\tau\}\(\\bm\{s\}\_\{0\}\),\\tau\)\\\|^\{2\}\\mathrm\{d\}\\bm\{s\}\_\{0\}\\right\)\\mathrm\{d\}\\tau\(66\)By the definition of Wasserstein\-2 distance, we have:

t0โ€‹๐’ฒ22โ€‹\(ฯt,ฮฝt\)\\displaystyle t\_\{0\}\\mathcal\{W\}\_\{2\}^\{2\}\(\\rho\_\{t\},\\nu\_\{t\}\)=minโกโˆซฯ€โกโ€–๐’”1โˆ’๐’”2โ€–2โ€‹๐‘‘ฯ€s\.t\.โ€‹โˆซฯ€โ€‹dโ€‹๐’”1=ฯt,โˆซฯ€โ€‹dโ€‹๐’”2=ฮฝt\\displaystyle=\\min\_\{\\pi\}\\int\\\|\\bm\{s\}\_\{1\}\-\\bm\{s\}\_\{2\}\\\|^\{2\}\\mathrm\{d\}\\pi\\quad\\text\{s\.t\.\}\\int\\pi\\mathrm\{d\}\\bm\{s\}\_\{1\}=\\rho\_\{t\},\\int\\pi\\mathrm\{d\}\\bm\{s\}\_\{2\}=\\nu\_\{t\}โ‰คโˆซฮผ0โ€‹\(๐’”0\)โ€‹โ€–ฯ•t๐œฝโ€‹\(๐’”0\)โˆ’ฯ•tโ€‹\(๐’”0\)โ€–2โ€‹dโ€‹๐’”0\\displaystyle\\leq\\int\\mu\_\{0\}\(\\bm\{s\}\_\{0\}\)\\\|\\bm\{\\phi\}^\{\\bm\{\\theta\}\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}\(\\bm\{s\}\_\{0\}\)\\\|^\{2\}\\mathrm\{d\}\\bm\{s\}\_\{0\}=Qtโ‰คe\(2โ€‹Lฮธ\+3\)โ€‹\(tโˆ’t0\)โ€‹โˆซt0t๐”ผ๐’”โˆผฮผฯ„โ€‹โ€–๐’‚ฮธโ€‹\(๐’”,ฯ„\)โˆ’๐’‚โก\(๐’”,ฯ„\)โ€–2โ€‹๐‘‘ฯ„\\displaystyle=Q\_\{t\}\\leq e^\{\(2L\_\{\\theta\}\+3\)\(t\-t\_\{0\}\)\}\\int\_\{t\_\{0\}\}^\{t\}\\mathbb\{E\}\_\{\\bm\{s\}\\sim\\mu\_\{\\tau\}\}\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{s\},\\tau\)\-\\bm\{a\}\(\\bm\{s\},\\tau\)\\\|^\{2\}\\mathrm\{d\}\\tau\(67\)=e\(2โ€‹Lฮธ\+3\)โ€‹\(tโˆ’t0\)โ€‹\(tโˆ’t0\)โ€‹โ„’AFM\|\[t0,t\]\.\\displaystyle=e^\{\(2L\_\{\\theta\}\+3\)\(t\-t\_\{0\}\)\}\(t\-t\_\{0\}\)\\mathcal\{L\}\_\{\\text\{AFM\}\}\|\_\{\[t\_\{0\},t\]\}\.\(68\)
Estimation of Term 2:Takeฯ€0โ€‹\(๐’”0,๐’”0ฮพ\)\\pi\_\{0\}\(\\bm\{s\}\_\{0\},\\bm\{s\}\_\{0\}^\{\\xi\}\)as the optimal coupling that attains๐’ฒ22โ€‹\(ฮฝt0,ฯ^t0\)=โˆซโ€–๐’”0โˆ’๐’”0ฮพโ€–2โ€‹dโ€‹ฯ€0โ€‹\(๐’”๐ŸŽ,๐’”0ฮพ\)\\mathcal\{W\}\_\{2\}^\{2\}\(\\nu\_\{t\_\{0\}\},\\hat\{\\rho\}\_\{t\_\{0\}\}\)=\\int\\\|\\bm\{s\}\_\{0\}\-\\bm\{s\}\_\{0\}^\{\\xi\}\\\|^\{2\}\\mathrm\{d\}\\pi\_\{0\}\(\\bm\{s\_\{0\}\},\\bm\{s\}\_\{0\}^\{\\xi\}\)\. Note thatฮฝt0=ฮผ0\\nu\_\{t\_\{0\}\}=\\mu\_\{0\}\(the true data distribution\)\.

From a pair of starting points\(๐’”0,๐’”0ฮพ\)โˆผฯ€0\(\\bm\{s\}\_\{0\},\\bm\{s\}\_\{0\}^\{\\xi\}\)\\sim\\pi\_\{0\}, define the square distance between trajectories under the same approximate flow asฮ”โก\(t\)=โ€–ฯ•tฮธโ€‹\(๐’”0\)โˆ’ฯ•tฮธโ€‹\(๐’”0ฮพ\)โ€–2\\Delta\(t\)=\\\|\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}^\{\\xi\}\)\\\|^\{2\}\. Taking the time derivative, we have:

dโ€‹ฮ”โ€‹\(t\)dโ€‹t\\displaystyle\\frac\{\\mathrm\{d\}\\Delta\(t\)\}\{\\mathrm\{d\}t\}=2โ€‹\(ฯ•tฮธโ€‹\(๐’”0\)โˆ’ฯ•tฮธโ€‹\(๐’”0ฮพ\)\)Tโ€‹\(dโ€‹ฯ•tฮธโ€‹\(๐’”0\)dโ€‹tโˆ’dโ€‹ฯ•tฮธโ€‹\(๐’”0ฮพ\)dโ€‹t\)\\displaystyle=2\{\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}^\{\\xi\}\)\)^\{T\}\}\\left\(\\frac\{\\mathrm\{d\}\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\}\{\\mathrm\{d\}t\}\-\\frac\{\\mathrm\{d\}\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}^\{\\xi\}\)\}\{\\mathrm\{d\}t\}\\right\)=2โ€‹\(ฯ•tฮธโ€‹\(๐’”0\)โˆ’ฯ•tฮธโ€‹\(๐’”0ฮพ\)\)Tโ€‹\(Aโก\(ฯ•tฮธโ€‹\(๐’”0\)โˆ’ฯ•tฮธโ€‹\(๐’”0ฮพ\)\)\+\[๐ŸŽ๐’‚ฮธโ€‹\(ฯ•tฮธโ€‹\(๐’”0\),t\)โˆ’๐’‚ฮธโ€‹\(ฯ•tฮธโ€‹\(๐’”0ฮพ\),t\)\]\)\\displaystyle=2\{\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}^\{\\xi\}\)\)^\{T\}\}\\left\(A\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}^\{\\xi\}\)\)\+\\begin\{bmatrix\}\\bm\{0\}\\\\ \\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\),t\)\-\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}^\{\\xi\}\),t\)\\end\{bmatrix\}\\right\)โ‰ค2โ€‹\(โ€–Aโ€–โ‹…โ€–ฯ•tฮธโ€‹\(๐’”0\)โˆ’ฯ•tฮธโ€‹\(๐’”0ฮพ\)โ€–2\+โ€–ฯ•tฮธโ€‹\(๐’”0\)โˆ’ฯ•tฮธโ€‹\(๐’”0ฮพ\)โ€–โ‹…โ€–๐’‚ฮธโ€‹\(ฯ•tฮธโ€‹\(๐’”0\),t\)โˆ’๐’‚ฮธโ€‹\(ฯ•tฮธโ€‹\(๐’”0ฮพ\),t\)โ€–\)\\displaystyle\\leq 2\\left\(\\\|A\\\|\\cdot\\\|\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}^\{\\xi\}\)\\\|^\{2\}\+\\\|\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}^\{\\xi\}\)\\\|\\cdot\\\|\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\),t\)\-\\bm\{a\}\_\{\\theta\}\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}^\{\\xi\}\),t\)\\\|\\right\)โ‰ค2โ€‹\(1\+Lฮธ\)โ€‹ฮ”โ€‹\(t\)\.\\displaystyle\\leq 2\(1\+L\_\{\\theta\}\)\\Delta\(t\)\.\(69\)Again by Gronwallโ€™s inequality, we have:

ฮ”โก\(t\)โ‰คe2โ€‹\(1\+Lฮธ\)โ€‹\(tโˆ’t0\)โ€‹ฮ”โ€‹\(t0\)=e2โ€‹\(1\+Lฮธ\)โ€‹\(tโˆ’t0\)โ€‹โ€–๐’”0โˆ’๐’”0ฮพโ€–2\.\\Delta\(t\)\\leq e^\{2\(1\+L\_\{\\theta\}\)\(t\-t\_\{0\}\)\}\\Delta\(t\_\{0\}\)=e^\{2\(1\+L\_\{\\theta\}\)\(t\-t\_\{0\}\)\}\\\|\\bm\{s\}\_\{0\}\-\\bm\{s\}\_\{0\}^\{\\xi\}\\\|^\{2\}\.\(70\)Sinceฮฝt=\(ฯ•tฮธ\)\#โ€‹ฮฝt0\\nu\_\{t\}=\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\)\_\{\\\#\}\\nu\_\{t\_\{0\}\}andฯ^t0=\(ฯ•tฮธ\)\#โ€‹ฯ^t0\\hat\{\\rho\}\_\{t\_\{0\}\}=\(\\bm\{\\phi\}\_\{t\}^\{\\theta\}\)\_\{\\\#\}\\hat\{\\rho\}\_\{t\_\{0\}\}, the push\-forward of the optimal couplingฯ€t=\(ฯ•tฮธ,ฯ•tฮธ\)\#โ€‹ฯ€0\\pi\_\{t\}=\(\\bm\{\\phi\}\_\{t\}^\{\\theta\},\\bm\{\\phi\}\_\{t\}^\{\\theta\}\)\_\{\\\#\}\\pi\_\{0\}is a valid coupling for\(ฮฝt,ฯ^t\)\(\\nu\_\{t\},\\hat\{\\rho\}\_\{t\}\)\. Integrate the squared distance over all starting pairs with respect toฯ€0\\pi\_\{0\}, we obtain:

๐’ฒ22โ€‹\(ฮฝt,ฯ^t\)\\displaystyle\\mathcal\{W\}\_\{2\}^\{2\}\(\\nu\_\{t\},\\hat\{\\rho\}\_\{t\}\)โ‰คโˆซโ€–ฯ•tฮธโ€‹\(๐’”0\)โˆ’ฯ•tฮธโ€‹\(๐’”0ฮพ\)โ€–2โ€‹dโ€‹ฯ€0โ€‹\(๐’”0,๐’”0ฮพ\)\\displaystyle\\leq\\int\\\|\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}\)\-\\bm\{\\phi\}\_\{t\}^\{\\theta\}\(\\bm\{s\}\_\{0\}^\{\\xi\}\)\\\|^\{2\}\\mathrm\{d\}\\pi\_\{0\}\(\\bm\{s\}\_\{0\},\\bm\{s\}\_\{0\}^\{\\xi\}\)โ‰คโˆซe2โ€‹\(1\+Lฮธ\)โ€‹\(tโˆ’t0\)โ€‹โ€–๐’”0โˆ’๐’”0ฮพโ€–2โ€‹dโ€‹ฯ€0โ€‹\(๐’”0,๐’”0ฮพ\)\\displaystyle\\leq\\int e^\{2\(1\+L\_\{\\theta\}\)\(t\-t\_\{0\}\)\}\\\|\\bm\{s\}\_\{0\}\-\\bm\{s\}\_\{0\}^\{\\xi\}\\\|^\{2\}\\mathrm\{d\}\\pi\_\{0\}\(\\bm\{s\}\_\{0\},\\bm\{s\}\_\{0\}^\{\\xi\}\)=e2โ€‹\(1\+Lฮธ\)โ€‹\(tโˆ’t0\)โ€‹๐’ฒ22โ€‹\(ฮฝt0,ฯ^t0\)\\displaystyle=e^\{2\(1\+L\_\{\\theta\}\)\(t\-t\_\{0\}\)\}\\mathcal\{W\}\_\{2\}^\{2\}\(\\nu\_\{t\_\{0\}\},\\hat\{\\rho\}\_\{t\_\{0\}\}\)\(71\)Letฯ€~โ€‹\(๐’”๐ŸŽ,๐’”0ฮพ\)\\tilde\{\\pi\}\(\\bm\{s\_\{0\}\},\\bm\{s\}\_\{0\}^\{\\xi\}\)be the coupling betweenฮฝ0\\nu\_\{0\}andฯ^0\\hat\{\\rho\}\_\{0\}induced by the identity mapping on the position space, i\.e\., pairing๐’”0=\(๐’™0,๐’—0\)\\bm\{s\}\_\{0\}=\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\)with๐’”0ฮพ=\(๐’™0,๐’—0ฮพ\)\\bm\{s\}\_\{0\}^\{\\xi\}=\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}^\{\\xi\}\)\. We can bound Wasserstein distance byL2L\_\{2\}loss:

๐’ฒ22โ€‹\(ฮฝt0,ฯ^t0\)\\displaystyle\\mathcal\{W\}\_\{2\}^\{2\}\(\\nu\_\{t\_\{0\}\},\\hat\{\\rho\}\_\{t\_\{0\}\}\)โ‰คโˆซโ€–๐’”0โˆ’๐’”0ฮพโ€–2โ€‹๐‘‘ฯ€~โ€‹\(๐’”0,๐’”0ฮพ\)\\displaystyle\\leq\\int\\\|\\bm\{s\}\_\{0\}\-\\bm\{s\}\_\{0\}^\{\\xi\}\\\|^\{2\}\\mathrm\{d\}\\tilde\{\\pi\}\(\\bm\{s\}\_\{0\},\\bm\{s\}\_\{0\}^\{\\xi\}\)=โˆซ\(โ€–๐’™0โˆ’๐’™0ฮพโ€–2\+โ€–๐’—0โˆ’๐’—0ฮพโ€–2\)โ€‹๐‘‘ฯ€~โ€‹\(๐’™0,๐’—0,๐’™0ฮพ,๐’—0ฮพ\)\\displaystyle=\\int\\left\(\\\|\\bm\{x\}\_\{0\}\-\\bm\{x\}\_\{0\}^\{\\xi\}\\\|^\{2\}\+\\\|\\bm\{v\}\_\{0\}\-\\bm\{v\}\_\{0\}^\{\\xi\}\\\|^\{2\}\\right\)\\mathrm\{d\}\\tilde\{\\pi\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\},\\bm\{x\}\_\{0\}^\{\\xi\},\\bm\{v\}\_\{0\}^\{\\xi\}\)=โˆซโ€–๐’—0โˆ’๐’—0ฮพโ€–2โ€‹๐‘‘ฯ€~โ€‹\(๐’™0,๐’—0,๐’™0ฮพ,๐’—0ฮพ\)=โ„’v0\\displaystyle=\\int\\\|\\bm\{v\}\_\{0\}\-\\bm\{v\}^\{\\xi\}\_\{0\}\\\|^\{2\}\\mathrm\{d\}\\tilde\{\\pi\}\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\},\{\\bm\{x\}\}\_\{0\}^\{\\xi\},\{\\bm\{v\}\}\_\{0\}^\{\\xi\}\)=\\mathcal\{L\}\_\{v\_\{0\}\}\(72\)
Thus๐’ฒ22โ€‹\(ฮฝt,ฯ^t\)โ‰คe2โ€‹\(1\+Lฮธ\)โ€‹\(tโˆ’t0\)โ€‹โ„’v0\.\\mathcal\{W\}\_\{2\}^\{2\}\(\\nu\_\{t\},\\hat\{\\rho\}\_\{t\}\)\\leq e^\{2\(1\+L\_\{\\theta\}\)\(t\-t\_\{0\}\)\}\\mathcal\{L\}\_\{v\_\{0\}\}\.

Combining the bounds above, we have

๐’ฒ22โ€‹\(ฯ^ti,ฯti\)\\displaystyle\\mathcal\{W\}\_\{2\}^\{2\}\(\\hat\{\\rho\}\_\{t\_\{i\}\},\\rho\_\{t\_\{i\}\}\)โ‰ค2โ€‹\[๐’ฒ22โ€‹\(ฯ^ti,ฮฝti\)\+๐’ฒ22โ€‹\(ฮฝti,ฯti\)\]\\displaystyle\\leq 2\\left\[\\mathcal\{W\}\_\{2\}^\{2\}\(\\hat\{\\rho\}\_\{t\_\{i\}\},\\nu\_\{t\_\{i\}\}\)\+\\mathcal\{W\}\_\{2\}^\{2\}\(\\nu\_\{t\_\{i\}\},\\rho\_\{t\_\{i\}\}\)\\right\]โ‰ค2โ€‹\[e\(2โ€‹Lฮธ\+3\)โ€‹\(tiโˆ’t0\)โ‹…\(tiโˆ’t0\)โ€‹โ„’AFM\|\[t0,ti\]\+e\(2โ€‹Lฮธ\+2\)โ€‹\(tiโˆ’t0\)โ€‹โ„’v0\]\.\\displaystyle\\leq 2\\left\[e^\{\(2L\_\{\\theta\}\+3\)\(t\_\{i\}\-t\_\{0\}\)\}\\cdot\(t\_\{i\}\-t\_\{0\}\)\\mathcal\{L\}\_\{\\text\{AFM\}\}\|\_\{\[t\_\{0\},t\_\{i\}\]\}\+e^\{\(2L\_\{\\theta\}\+2\)\(t\_\{i\}\-t\_\{0\}\)\}\\mathcal\{L\}\_\{v\_\{0\}\}\\right\]\.\(73\)
Defineฮบ=\(2โ€‹Lฮธ\+3\)\\kappa=\{\(2L\_\{\\theta\}\+3\)\},then the final bound is obtained by:

๐’ฒ22โ€‹\(ฯ^ti\(pos\),ฯti\(pos\)=ฮผi\(pos\)\)\\displaystyle\\mathcal\{W\}\_\{2\}^\{2\}\(\\hat\{\\rho\}^\{\\text\{\(pos\)\}\}\_\{t\_\{i\}\},\\rho^\{\(\\text\{pos\}\)\}\_\{t\_\{i\}\}=\\mu\_\{i\}^\{\\text\{\(pos\)\}\}\)=minโกโˆซฯ€๐’™โกโ€–๐’™1โˆ’๐’™2โ€–2โ€‹dโ€‹ฯ€๐’™โ€‹\(๐’™1,๐’™2\)s\.t\.โ€‹โˆซฯ€๐’™โ€‹dโ€‹๐’™2=ฯti\(pos\),โˆซฯ€๐’™โ€‹dโ€‹๐’™1=ฯ^ti\(pos\)\\displaystyle=\\min\_\{\\pi\_\{\\bm\{x\}\}\}\\int\\\|\\bm\{x\}\_\{1\}\-\\bm\{x\}\_\{2\}\\\|^\{2\}\\mathrm\{d\}\\pi\_\{\\bm\{x\}\}\(\\bm\{x\}\_\{1\},\\bm\{x\}\_\{2\}\)\\quad\\text\{s\.t\.\}\\int\\pi\_\{\\bm\{x\}\}\\mathrm\{d\}\\bm\{x\}\_\{2\}=\\rho^\{\(\\text\{pos\}\)\}\_\{t\_\{i\}\},\\quad\\int\\pi\_\{\\bm\{x\}\}\\mathrm\{d\}\\bm\{x\}\_\{1\}=\\hat\{\\rho\}\_\{t\_\{i\}\}^\{\\text\{\(pos\)\}\}โ‰คminโกโˆซฯ€โกโ€–๐’™1โˆ’๐’™2โ€–2โ€‹๐‘‘ฯ€โ€‹\(๐’”1,๐’”2\)s\.t\.โ€‹โˆซฯ€โ€‹dโ€‹๐’”2=ฯti,โˆซฯ€โ€‹dโ€‹๐’”1=ฯ^ti\\displaystyle\\leq\\min\_\{\\pi\}\\int\\\|\\bm\{x\}\_\{1\}\-\\bm\{x\}\_\{2\}\\\|^\{2\}\\mathrm\{d\}\\pi\(\\bm\{s\}\_\{1\},\\bm\{s\}\_\{2\}\)\\quad\\text\{s\.t\.\}\\int\\pi\\mathrm\{d\}\\bm\{s\}\_\{2\}=\\rho\_\{t\_\{i\}\},\\quad\\int\\pi\\mathrm\{d\}\\bm\{s\}\_\{1\}=\\hat\{\\rho\}\_\{t\_\{i\}\}โ‰คminโกโˆซฯ€โกโ€–๐’”1โˆ’๐’”2โ€–2โ€‹๐‘‘ฯ€โ€‹\(๐’”1,๐’”2\)\\displaystyle\\leq\\min\_\{\\pi\}\\int\\\|\\bm\{s\}\_\{1\}\-\\bm\{s\}\_\{2\}\\\|^\{2\}\\mathrm\{d\}\\pi\(\\bm\{s\}\_\{1\},\\bm\{s\}\_\{2\}\)\(74\)=๐’ฒ22โ€‹\(ฯ^ti,ฯti\)โ‰ค2โ€‹eฮบโก\(tiโˆ’t0\)โ€‹\(\(tiโˆ’t0\)โ€‹โ„’AFM\|\[t0,ti\]\+โ„’v0\)\.\\displaystyle=\\mathcal\{W\}\_\{2\}^\{2\}\(\\hat\{\\rho\}\_\{t\_\{i\}\},\\rho\_\{t\_\{i\}\}\)\\leq 2e^\{\\kappa\(t\_\{i\}\-t\_\{0\}\)\}\\left\(\(t\_\{i\}\-t\_\{0\}\)\\mathcal\{L\}\_\{\\text\{AFM\}\}\|\_\{\[t\_\{0\},t\_\{i\}\]\}\+\\mathcal\{L\}\_\{v\_\{0\}\}\\right\)\.\(75\)
This completes the proof\. Note that our derivation holds even for multi\-valued or stochastic cases, since the upper bound on the Wasserstein distance is established via a valid coupling, which remains well\-defined regardless of whether the mapping is deterministic\.

โˆŽ

## Appendix BDatasets and Evaluation Metric

### B\.1Experiment Setup

The experiments were conducted on a shared high\-performance computing \(HPC\) cluster\. All experimental runs utilized GPUs equipped with 80GB of memory\. Both the network๐’‚ฮธโ€‹\(๐’™,๐’—,t\)\\bm\{a\}\_\{\\theta\}\(\\bm\{x\},\\bm\{v\},t\), employed to approximate acceleration, and the network๐’—ฯ•โ€‹\(๐’™,๐’—,t\)\\bm\{v\}\_\{\\phi\}\(\\bm\{x\},\\bm\{v\},t\), employed to approximate velocity, utilize residual architectures\. Specifically,๐’‚ฮธโ€‹\(๐’™,๐’—,t\)\\bm\{a\}\_\{\\theta\}\(\\bm\{x\},\\bm\{v\},t\)consists of 4 hidden layers, while๐’—ฯ•โ€‹\(๐’™,๐’—,t\)\\bm\{v\}\_\{\\phi\}\(\\bm\{x\},\\bm\{v\},t\)comprises 2 hidden layers\. The Optimal Transport \(OT\) calculations were implemented using the Python Optimal Transport \(POT\) library\[[16](https://arxiv.org/html/2608.21070#bib.bib16)\], and the training process of neural networks were implemented using Pytorch\[[43](https://arxiv.org/html/2608.21070#bib.bib43)\]\.

In all experiments, TracingFlow employs a completely consistent set of hyperparameters\. When solving for OT using the Sinkhorn algorithm, the regularization coefficient is set toฯต=1ร—10โˆ’3\\epsilon=1\\times 10^\{\-3\}\. When incorporating biological priors, the penalty coefficient for barcode mismatches is set top0=25p\_\{0\}=25\. Since the number of cells at each time point in the datasets processed in this paper is on the order of10310^\{3\}, we do not utilize mini\-batch OT; instead, we compute the Transport Plan on the full dataset\.

### B\.2Evaluation Metrics

We employ the 1\-Wasserstein Distance \(๐’ฒ1\\mathcal\{W\}\_\{1\}\) and the 2\-Wasserstein Distance \(๐’ฒ2\\mathcal\{W\}\_\{2\}\) to quantify the discrepancy between the predicted distribution and the ground truth distribution\. These metrics are defined as follows:

๐’ฒ1โ€‹\(ฮผ,ฮฝ\)=minโกโˆซฯ€โˆˆฮ โก\(ฮผ,ฮฝ\)โกโ€–๐’™โˆ’๐’šโ€–โ€‹๐‘‘ฯ€โ€‹\(๐’™,๐’š\)\\mathcal\{W\}\_\{1\}\(\\mu,\\nu\)=\\min\_\{\\pi\\in\\Pi\(\\mu,\\nu\)\}\\int\\\|\\bm\{x\}\-\\bm\{y\}\\\|\\mathrm\{d\}\\pi\(\\bm\{x\},\\bm\{y\}\)\(76\)๐’ฒ2โ€‹\(ฮผ,ฮฝ\)=\(minโกโˆซฯ€โˆˆฮ โก\(ฮผ,ฮฝ\)โกโ€–๐’™โˆ’๐’šโ€–22โ€‹๐‘‘ฯ€โ€‹\(๐’™,๐’š\)\)1/2\\mathcal\{W\}\_\{2\}\(\\mu,\\nu\)=\\left\(\\min\_\{\\pi\\in\\Pi\(\\mu,\\nu\)\}\\int\\\|\\bm\{x\}\-\\bm\{y\}\\\|\_\{2\}^\{2\}\\mathrm\{d\}\\pi\(\\bm\{x\},\\bm\{y\}\)\\right\)^\{1/2\}\(77\)For the evaluation of datasets containing lineage information, we utilize the lineage\-weighted๐’ฒ1\\mathcal\{W\}\_\{1\}and๐’ฒ2\\mathcal\{W\}\_\{2\}distances\. These are defined based on the decomposition of the distributions\. Consider the distribution generated by the model, denoted asฮผโก\(๐’™\)\\mu\(\\bm\{x\}\), and the reference distribution provided by the dataset, denoted asฮฝโก\(๐’™\)\\nu\(\\bm\{x\}\)\. Both consist ofLLlineages and can be decomposed as:

ฮผโก\(๐’™\)=โˆ‘i=1Lwiโ€‹ฮผ\(i\)โ€‹\(๐’™\),ฮฝโก\(๐’™\)=โˆ‘i=1Lviโ€‹ฮฝ\(i\)โ€‹\(๐’™\)\\mu\(\\bm\{x\}\)=\\sum\_\{i=1\}^\{L\}w\_\{i\}\\mu^\{\(i\)\}\(\\bm\{x\}\),\\quad\\nu\(\\bm\{x\}\)=\\sum\_\{i=1\}^\{L\}v\_\{i\}\\nu^\{\(i\)\}\(\\bm\{x\}\)\(78\)whereฮผ\(i\)โ€‹\(๐’™\)\\mu^\{\(i\)\}\(\\bm\{x\}\)andฮฝ\(i\)โ€‹\(๐’™\)\\nu^\{\(i\)\}\(\\bm\{x\}\)represent the distributions of theii\-th lineage withinฮผโก\(๐’™\)\\mu\(\\bm\{x\}\)andฮฝโก\(๐’™\)\\nu\(\\bm\{x\}\), respectively, satisfyingโˆซฮผ\(i\)โ€‹\(๐’™\)โ€‹๐‘‘๐’™=1\\int\\mu^\{\(i\)\}\(\\bm\{x\}\)\\mathrm\{d\}\\bm\{x\}=1andโˆซฮฝ\(i\)โ€‹\(๐’™\)โ€‹๐‘‘๐’™=1\\int\\nu^\{\(i\)\}\(\\bm\{x\}\)\\mathrm\{d\}\\bm\{x\}=1\. The termswiw\_\{i\}andviv\_\{i\}denote the weights of theii\-th lineage inฮผโก\(๐’™\)\\mu\(\\bm\{x\}\)andฮฝโก\(๐’™\)\\nu\(\\bm\{x\}\), respectively, subject to the constraintโˆ‘i=1Lwi=โˆ‘i=1Lvi=1\\sum\_\{i=1\}^\{L\}w\_\{i\}=\\sum\_\{i=1\}^\{L\}v\_\{i\}=1\. The lineage\-weighted๐’ฒ1\\mathcal\{W\}\_\{1\}and๐’ฒ2\\mathcal\{W\}\_\{2\}distances are defined as:

๐’ฒ1\(L\)\(ฮผโˆฅฮฝ\)=โˆ‘i=1Lvi๐’ฒ1\(ฮผ\(i\),ฮฝ\(i\)\),๐’ฒ2\(L\)\(ฮผโˆฅฮฝ\)=โˆ‘i=1Lvi๐’ฒ2\(ฮผ\(i\),ฮฝ\(i\)\)\\mathcal\{W\}\_\{1\}^\{\(L\)\}\(\\mu\\\|\\nu\)=\\sum\_\{i=1\}^\{L\}v\_\{i\}\\mathcal\{W\}\_\{1\}\(\\mu^\{\(i\)\},\\nu^\{\(i\)\}\),\\quad\\mathcal\{W\}\_\{2\}^\{\(L\)\}\(\\mu\\\|\\nu\)=\\sum\_\{i=1\}^\{L\}v\_\{i\}\\mathcal\{W\}\_\{2\}\(\\mu^\{\(i\)\},\\nu^\{\(i\)\}\)\(79\)In essence, we first calculate the๐’ฒ1\\mathcal\{W\}\_\{1\}and๐’ฒ2\\mathcal\{W\}\_\{2\}distances between the distributions of corresponding lineages in the two datasets\. Subsequently, we compute a weighted average of these distances using the proportions of each lineage given by the reference distributionฮฝโก\(๐’™\)\\nu\(\\bm\{x\}\)\. It is important to note that the lineage\-weighted๐’ฒ1\\mathcal\{W\}\_\{1\}and๐’ฒ2\\mathcal\{W\}\_\{2\}metrics serve solely as indicators of the modelโ€™s ability to preserve lineage prior information while learning the dynamics\. They are not true distance metrics in the mathematical sense, as they are evidently asymmetric with respect toฮผ\\muandฮฝ\\nu\.

### B\.3Datasets

3D Simulation Lineage Data\.Consider the following dynamics with three lineage barcodes and a twist structure:

\{dโ€‹Xt\(0\)=vโ€‹dโ€‹t\+ฯƒxโ€‹dโ€‹BtXdโ€‹Yt\(0\)=ฯƒyโ€‹dโ€‹BtYdโ€‹Zt\(0\)=ฯƒzโ€‹dโ€‹BtZ\{dโ€‹Xt\(1\)=vโ€‹dโ€‹t\+ฯƒxโ€‹dโ€‹BtXdโ€‹Yt\(1\)=โˆ’ฮผโ€‹Zt\(1\)โ€‹dโ€‹t\+ฯƒyโ€‹dโ€‹BtYdโ€‹Zt\(1\)=ฮผโ€‹Yt\(1\)โ€‹dโ€‹t\+ฯƒzโ€‹dโ€‹BtZ\{dโ€‹Xt\(2\)=vโ€‹dโ€‹t\+ฯƒxโ€‹dโ€‹BtXdโ€‹Yt\(2\)=ฮผโ€‹Zt\(2\)โ€‹dโ€‹t\+ฯƒyโ€‹dโ€‹BtYdโ€‹Zt\(2\)=โˆ’ฮผโ€‹Yt\(2\)โ€‹dโ€‹t\+ฯƒzโ€‹dโ€‹BtZ\\left\\\{\\begin\{aligned\} dX\_\{t\}^\{\(0\)\}&=v\\,dt\+\\sigma\_\{x\}dB\_\{t\}^\{X\}\\\\ dY\_\{t\}^\{\(0\)\}&=\\sigma\_\{y\}dB\_\{t\}^\{Y\}\\\\ dZ\_\{t\}^\{\(0\)\}&=\\sigma\_\{z\}dB\_\{t\}^\{Z\}\\end\{aligned\}\\right\.\\quad\\left\\\{\\begin\{aligned\} dX\_\{t\}^\{\(1\)\}&=v\\,dt\+\\sigma\_\{x\}dB\_\{t\}^\{X\}\\\\ dY\_\{t\}^\{\(1\)\}&=\-\\mu Z\_\{t\}^\{\(1\)\}\\,dt\+\\sigma\_\{y\}dB\_\{t\}^\{Y\}\\\\ dZ\_\{t\}^\{\(1\)\}&=\\mu Y\_\{t\}^\{\(1\)\}\\,dt\+\\sigma\_\{z\}dB\_\{t\}^\{Z\}\\end\{aligned\}\\right\.\\quad\\left\\\{\\begin\{aligned\} dX\_\{t\}^\{\(2\)\}&=v\\,dt\+\\sigma\_\{x\}dB\_\{t\}^\{X\}\\\\ dY\_\{t\}^\{\(2\)\}&=\\mu Z\_\{t\}^\{\(2\)\}\\,dt\+\\sigma\_\{y\}dB\_\{t\}^\{Y\}\\\\ dZ\_\{t\}^\{\(2\)\}&=\-\\mu Y\_\{t\}^\{\(2\)\}\\,dt\+\\sigma\_\{z\}dB\_\{t\}^\{Z\}\\end\{aligned\}\\right\.\(80\)We takev=3\.0,ฮผ=3\.14,ฯƒx=ฯƒy=ฯƒz=0\.1v=3\.0,\\mu=3\.14,\\sigma\_\{x\}=\\sigma\_\{y\}=\\sigma\_\{z\}=0\.1and time pointst=0,1,2t=0,1,2\. Initialize\(X0\(0\),Y0\(0\),Z0\(0\)\)โ€‹โˆผโ€‹๐’ฉโ€‹\(\[0,0,0\],0\.5โ€‹I\)\(X\_\{0\}^\{\(0\)\},Y\_\{0\}^\{\(0\)\},Z\_\{0\}^\{\(0\)\}\)\\overset\{\}\{\\sim\}\\mathcal\{N\}\(\[0,0,0\],0\.5I\),\(X0\(1\),Y0\(1\),Z0\(1\)\)โ€‹โˆผโ€‹๐’ฉโ€‹\(\[0,โˆ’2,0\],0\.5โ€‹I\),\(X0\(2\),Y0\(2\),Z0\(2\)\)โˆผ๐’ฉโก\(\[0,2,0\],0\.5โ€‹I\)\(X\_\{0\}^\{\(1\)\},Y\_\{0\}^\{\(1\)\},Z\_\{0\}^\{\(1\)\}\)\\overset\{\}\{\\sim\}\\mathcal\{N\}\(\[0,\-2,0\],0\.5I\),\(X\_\{0\}^\{\(2\)\},Y\_\{0\}^\{\(2\)\},Z\_\{0\}^\{\(2\)\}\)\{\\sim\}\\mathcal\{N\}\(\[0,2,0\],0\.5I\)and sample500500particles for each\.

![Refer to caption](https://arxiv.org/html/2608.21070v1/figures/2D.png)
![Refer to caption](https://arxiv.org/html/2608.21070v1/figures/3D.png)

Figure 5:Dynamics of simulation datasets\. \(a\)2D simulation data, \(b\)3D simulation data and \(c\)3D simulation data projected to XY plane\.Cite Data\.We utilized the Cite\-seq dataset introduced in\[[31](https://arxiv.org/html/2608.21070#bib.bib31)\], which comprises 31,240 cells collected across 4 distinct time points\. To evaluate the performance of TracingFlow on real\-world biological data, we applied Principal Component Analysis \(PCA\) to reduce the dimensionality of the gene expression profiles to 100 and 5 dimensions, respectively\.

Embryoid Bodies Data\.We employed the Embryoid Bodies \(EB\) dataset from\[[39](https://arxiv.org/html/2608.21070#bib.bib39)\], consisting of 16,819 cells sampled at 5 time points\. We utilized PCA to reduce the dimensionality of the gene expression data to 5 dimensions\.

Hematopoiesis Data\.We employed the Hematopoiesis dataset from\[[64](https://arxiv.org/html/2608.21070#bib.bib64)\], consisting of 130,887 cells sampled at 3 time points\. Part of the cells are recorded with barcodes\. We selected 22,329 cells with barcodes from the dataset and utilized PCA to reduce the dimensionality of gene expression data to 50 dimensions\. When processing this dataset, we filtered for data where barcodes were present at all three time points\. The barcodes for the remaining data were set to undefined, incurring no additional transport cost during the incorporation of biological priors \([Section5\.3](https://arxiv.org/html/2608.21070#S5.SS3)\)\.

Mexico Gulf DataTo assess the capability and robustness of Tracing Flow in capturing long\-term system dynamics, we utilize the Mexico Gulf dataset, following\[[56](https://arxiv.org/html/2608.21070#bib.bib56)\]\. This two\-dimensional dataset comprises 9 time points\.

## Appendix CExperiment Details

### C\.1Experiment Details across various datasets

Simulation 2D Data[Table4](https://arxiv.org/html/2608.21070#A3.T4)presents the performance of TracingFlow and other algorithms on the 2D Simulation dataset\. On average, TracingFlow reconstructs the data distribution at each time point with higher accuracy\.

Table 4:๐’ฒ1\\mathcal\{W\}\_\{1\}and๐’ฒ2\\mathcal\{W\}\_\{2\}distances between the generated and ground\-truth distributions at each time point on the 2D Simulation dataset\.Cite 5D & Cite 100D DataWe present the distribution reconstruction performance of each algorithm on the Cite 5D and Cite 100D datasets in[Table5](https://arxiv.org/html/2608.21070#A3.T5)and[Table6](https://arxiv.org/html/2608.21070#A3.T6), respectively\. On these datasets TracingFlow achieves superior reconstruction accuracy at all time steps except fort=3t=3\. The trajectories learned by TracingFlow on these two datasets are illustrated in[Figure6](https://arxiv.org/html/2608.21070#A3.F6)\.

Table 5:๐’ฒ1\\mathcal\{W\}\_\{1\}and๐’ฒ2\\mathcal\{W\}\_\{2\}distances between the generated and ground\-truth distributions at each time point on the Cite 5D datasetTable 6:๐’ฒ1\\mathcal\{W\}\_\{1\}and๐’ฒ2\\mathcal\{W\}\_\{2\}distances between the generated and ground\-truth distributions at each time point on the Cite 100D dataset![Refer to caption](https://arxiv.org/html/2608.21070v1/Cite.png)Figure 6:Trajectories learned by TracingFlow on the \(a\) Cite 5D and \(b\) Cite 100D datasets, projected to 2D using PCA\.EB 5D Data[Table7](https://arxiv.org/html/2608.21070#A3.T7)displays the interpolation accuracy on the EB 5D dataset, evaluated by holding out one time point during training\. TracingFlow outperformed other methods in reconstruction accuracy for all time points exceptt=1t=1\.[Figure7](https://arxiv.org/html/2608.21070#A3.F7)illustrates the interpolated distributions generated by TracingFlow across the four time points\.

Table 7:๐’ฒ1\\mathcal\{W\}\_\{1\}and๐’ฒ2\\mathcal\{W\}\_\{2\}distances between predicted and ground\-truth distributions at held\-out time points on the EB 5D dataset\.![Refer to caption](https://arxiv.org/html/2608.21070v1/EB.png)Figure 7:Predicted distributions by TracingFlow at four held\-out time points, projected to 2D using PCA\.3D Simulation Lineage Dataset[Table8](https://arxiv.org/html/2608.21070#A3.T8)demonstrates the ability of various algorithms to preserve biological priors on the 3D Simulation Lineage dataset\. Since the trajectories learned by TracingFlow align with biological priors \(see[Figure4](https://arxiv.org/html/2608.21070#S6.F4)\), it achieves lower lineage\-weighted๐’ฒโ€‹1\\mathcal\{W\}\{1\}and๐’ฒโ€‹2\\mathcal\{W\}\{2\}scores on average\.

Table 8:lineage\-weighted๐’ฒ1\\mathcal\{W\}\_\{1\}and๐’ฒ2\\mathcal\{W\}\_\{2\}distances between the generated and ground\-truth distributions at each time point on the 3D Simulation Lineage datasetHematopoiesis Dataset[Table9](https://arxiv.org/html/2608.21070#A3.T9)presents the performance of different algorithms in retaining biological priors on real\-world datasets, measured by the lineage\-weighted๐’ฒโ€‹1,๐’ฒโ€‹2\\mathcal\{W\}\{1\},\\mathcal\{W\}\{2\}distance\. TracingFlow surpassed other algorithms at all time points\. We also evaluated the algorithms without introducing biological priors by Vanilla๐’ฒ1\\mathcal\{W\}\_\{1\},๐’ฒ2\\mathcal\{W\}\_\{2\}distance, and as shown in[Table10](https://arxiv.org/html/2608.21070#A3.T10), TracingFlow maintained superior performance\.[Figure8](https://arxiv.org/html/2608.21070#A3.F8)visualizes the data points generated by TracingFlow on the real\-world dataset under both conditions \(with and without biological priors\)\.

Table 9:lineage\-weighted๐’ฒ1\\mathcal\{W\}\_\{1\}and๐’ฒ2\\mathcal\{W\}\_\{2\}distances between the generated and ground\-truth distributions at each time point on the Hematopoiesis datasetTable 10:๐’ฒโ€‹1\\mathcal\{W\}\{1\}and๐’ฒโ€‹2\\mathcal\{W\}\{2\}distances between the generated and ground\-truth distributions at each time point on the Hematopoiesis dataset![Refer to caption](https://arxiv.org/html/2608.21070v1/RealLineage_A.png)Figure 8:Original data \(a\), and data points generated by TracingFlow on the Hematopoiesis dataset without biological priors \(b\) and with biological priors \(c\), projected to 2D using UMAP\.Mexico Gulf Dataset[Table11](https://arxiv.org/html/2608.21070#A3.T11)presents the performance of various algorithms on the Mexico Gulf dataset\. Tracing Flow achieves the best distribution reconstruction accuracy at all time points, yielding smooth trajectories\.

Table 11:๐’ฒ1\\mathcal\{W\}\_\{1\}and๐’ฒ2\\mathcal\{W\}\_\{2\}distances between the generated and ground\-truth distributions at each time point on the Mexico Gulf dataset

### C\.2Training time and scalability of TracingFlow

To evaluate the scalability of TracingFlow, we report the training time required by each method across various datasets, as presented in Table[12](https://arxiv.org/html/2608.21070#A3.T12)\. Experimental results indicate that Flow Matching algorithms based on second\-order dynamics \(e\.g\., 3 MSBM and our TracingFlow\) generally require more training time compared to first\-order dynamics algorithms \(e\.g\., OT\-CFM and SF2M\)\. Notably, the training time of TracingFlow is slightly lower than that of 3 MSBM, demonstrating that our method does not introduce additional computational overhead\.

Table 12:Training time comparison \(in seconds\) of different methods across various datasets\.
### C\.3Sensitivity Analysis on Minibatch\-OT

In all experiments presented in this paper, the full dataset was utilized to compute the OT plan\. However, as the scale of the data increases, the computational cost of the OT plan grows rapidly, making the adoption of the minibatch computation method described in[Section5\.2](https://arxiv.org/html/2608.21070#S5.SS2)inevitable\. We conducted a sensitivity analysis on the Cite\-5D dataset to evaluate the impact of batch size in minibatch\-OT computation on the accuracy of distribution reconstruction\. The results are presented in Table[13](https://arxiv.org/html/2608.21070#A3.T13)\. Experimental results demonstrate that TracingFlow consistently and accurately reconstructs the distributions at each time point when the batch size is adjusted within a certain range \( batch sizeโˆผ103\\sim 10^\{3\}\)\.

Table 13:Sensitivity analysis of batch size for Minibatch\-OT on the Cite\-5D dataset\.

## Appendix DPseudocode for Algorithm

The pseudocode for the training process of Tracing Flow is provided in[Algorithm1](https://arxiv.org/html/2608.21070#alg1)\.

Algorithm 1Tracing\-Flow Matching Training0:Snapshot datasets

๐’Ÿk=\{๐’™k\(i\)\}i=1Nk\\mathcal\{D\}\_\{k\}=\\\{\\bm\{x\}\_\{k\}^\{\(i\)\}\\\}\_\{i=1\}^\{N\_\{k\}\}at times

t0<t1<โ‹ฏ<tKt\_\{0\}<t\_\{1\}<\\dots<t\_\{K\}; Acceleration Network

๐’‚๐œฝ\\bm\{a\}\_\{\\bm\{\\theta\}\}; Velocity Network

๐’—๐ƒ\\bm\{v\}\_\{\\bm\{\\xi\}\}; Hyperparameters: batch size

BtrajB\_\{\\text\{traj\}\}, learning rates

ฮท๐’‚,ฮท๐’—\\eta\_\{\\bm\{a\}\},\\eta\_\{\\bm\{v\}\}\.

1:Initialize:Estimated velocity sets

๐’ฑk=\{๐’—k\(i\)\}i=1Nkโ†\{๐ŸŽ\}\\mathcal\{V\}\_\{k\}=\\\{\\bm\{v\}\_\{k\}^\{\(i\)\}\\\}\_\{i=1\}^\{N\_\{k\}\}\\leftarrow\\\{\\mathbf\{0\}\\\}for all

kk\.

2:whileVelocity Estimates Not Convergeddo

3:Stage 1: Velocity Assignment \(Iterative SOAT\)

4:Initialize velocity accumulators

๐’ฑ^k\(i\)โ†โˆ…\\hat\{\\mathcal\{V\}\}\_\{k\}^\{\(i\)\}\\leftarrow\\emptysetfor all

k,ik,i\.

5:// Step 1\.1: Solve Optimal Transport

6:for

k=0k=0to

Kโˆ’1K\-1do

7:Compute cost matrix

๐’žiโ†’i\+1\\mathcal\{C\}\_\{i\\rightarrow i\+1\}with / w\.o\. biological prior\.

8:Compute optimal coupling

ฯ€kโ†’k\+1โˆ—\\pi^\{\*\}\_\{k\\to k\+1\}between

\(๐’™k,๐’—k\)\(\\bm\{x\}\_\{k\},\\bm\{v\}\_\{k\}\)and

\(๐’™k\+1,๐’—k\+1\)\(\\bm\{x\}\_\{k\+1\},\\bm\{v\}\_\{k\+1\}\)by solving the SOAT Problem\.

9:Compute propagator \(transition\) matrix

๐’ฆkโ†’k\+1โˆ—\\mathcal\{K\}^\{\*\}\_\{k\\to k\+1\}by normalizing

ฯ€kโ†’k\+1โˆ—\\pi^\{\*\}\_\{k\\to k\+1\}:

10:

๐’ฆkโ†’k\+1โˆ—โ€‹\(i,j\)=ฯ€kโ†’k\+1โˆ—โ€‹\(i,j\)โˆ‘jโ€ฒฯ€kโ†’k\+1โˆ—โ€‹\(i,jโ€ฒ\)\\mathcal\{K\}^\{\*\}\_\{k\\to k\+1\}\(i,j\)=\\frac\{\\pi^\{\*\}\_\{k\\to k\+1\}\(i,j\)\}\{\\sum\_\{j^\{\\prime\}\}\\pi^\{\*\}\_\{k\\to k\+1\}\(i,j^\{\\prime\}\)\}
11:endfor

12:// Step 1\.2: Trajectory Sampling & Velocity Update

13:Sample

BtrajB\_\{\\text\{traj\}\}discrete trajectories

Tm=\{\(๐’™k\(sm,k\),๐’—k\(sm,k\)\)\}k=0KT\_\{m\}=\\\{\(\\bm\{x\}\_\{k\}^\{\(s\_\{m,k\}\)\},\\bm\{v\}\_\{k\}^\{\(s\_\{m,k\}\)\}\)\\\}\_\{k=0\}^\{K\}for

m=1โ€‹โ€ฆโ€‹Btrajm=1\\dots B\_\{\\text\{traj\}\}\.

14:

โ‹…\\cdotInitial indices

sm,0s\_\{m,0\}are sampled uniformly\.

15:

โ‹…\\cdotSubsequent indices

sm,k\+1โˆผ๐’ฆkโ†’k\+1โˆ—โ€‹\(sm,k,โ‹…\)s\_\{m,k\+1\}\\sim\\mathcal\{K\}^\{\*\}\_\{k\\to k\+1\}\(s\_\{m,k\},\\cdot\)\.

16:for

m=1m=1to

BtrajB\_\{\\text\{traj\}\}do

17:Compute optimal path velocities

\{๐’—^k\(sm,k\)\}k=0K\\\{\\hat\{\\bm\{v\}\}\_\{k\}^\{\(s\_\{m,k\}\)\}\\\}\_\{k=0\}^\{K\}for trajectory

TmT\_\{m\}using Proposition[4\.3](https://arxiv.org/html/2608.21070#S4.SS3.SSS0.Px1)\.

18:Accumulate:

๐’ฑ^k\(sm,k\)โ†๐’ฑ^k\(sm,k\)โˆช\{๐’—^k\(sm,k\)\}\\hat\{\\mathcal\{V\}\}\_\{k\}^\{\(s\_\{m,k\}\)\}\\leftarrow\\hat\{\\mathcal\{V\}\}\_\{k\}^\{\(s\_\{m,k\}\)\}\\cup\\\{\\hat\{\\bm\{v\}\}\_\{k\}^\{\(s\_\{m,k\}\)\}\\\}for

k=0โ€‹โ€ฆโ€‹Kk=0\\dots K\.

19:endfor

20:// Step 1\.3: Average Accumulated Velocities

21:for

k=0k=0to

KK,

i=1i=1to

NkN\_\{k\}do

22:if

๐’ฑ^k\(i\)โ‰ โˆ…\\hat\{\\mathcal\{V\}\}\_\{k\}^\{\(i\)\}\\neq\\emptysetthen

23:

๐’—k\(i\)โ†1\|๐’ฑ^k\(i\)\|โ€‹โˆ‘๐’—โˆˆ๐’ฑ^k\(i\)๐’—\\bm\{v\}\_\{k\}^\{\(i\)\}\\leftarrow\\frac\{1\}\{\|\\hat\{\\mathcal\{V\}\}\_\{k\}^\{\(i\)\}\|\}\\sum\_\{\\bm\{v\}\\in\\hat\{\\mathcal\{V\}\}\_\{k\}^\{\(i\)\}\}\\bm\{v\}
24:endif

25:endfor

26:endwhile

27:Stage 2: Final Transport Plan Computation

28:Compute optimal couplings

ฯ€โˆ—\\pi^\{\*\}and propagators

๐’ฆโˆ—\\mathcal\{K\}^\{\*\}using the converged velocities

๐’ฑk\\mathcal\{V\}\_\{k\}\.

29:Stage 3: Neural Network Training

30:whileNot Convergeddo

31:Sample random time

tโˆผ๐’ฐโก\[t0,tK\]t\\sim\\mathcal\{U\}\[t\_\{0\},t\_\{K\}\]and coupling

๐’›โˆผqโก\(๐’›\)\\bm\{z\}\\sim q\(\\bm\{z\}\)\.

32:Compute Acceleration Matching Loss:

33:

โ„’CAFM=๐”ผ\(๐’™,๐’—\)โˆผฯt\(โ‹…\|z\),๐’›โˆผq\(๐’›\),tโˆผ๐’ฐ\[t0,tK\]\[โˆฅ๐’‚๐œฝ\(๐’™,๐’—,t\)โˆ’๐’‚\(๐’™,๐’—,t\|z\)โˆฅ2\]\\mathcal\{L\}\_\{\\text\{CAFM\}\}=\\mathbb\{E\}\_\{\(\\bm\{x\},\\bm\{v\}\)\\sim\\rho\_\{t\}\(\\cdot\|z\),\\bm\{z\}\\sim q\(\\bm\{z\}\),t\\sim\\mathcal\{U\}\[t\_\{0\},t\_\{K\}\]\}\\Big\[\\\|\\bm\{a\}\_\{\\bm\{\\theta\}\}\(\\bm\{x\},\\bm\{v\},t\)\-\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\|z\)\\\|^\{2\}\\Big\]
34:Compute Initial Velocity Loss:

35:

โ„’๐’—0=๐”ผ\(๐’™,๐’—\)โˆผฯ0โ€‹\[โ€–๐’—๐ƒโ€‹\(๐’™\)โˆ’๐’—โ€–2\]\\mathcal\{L\}\_\{\\bm\{v\}\_\{0\}\}=\\mathbb\{E\}\_\{\(\\bm\{x\},\\bm\{v\}\)\\sim\\rho\_\{0\}\}\\Big\[\\\|\\bm\{v\}\_\{\\bm\{\\xi\}\}\(\\bm\{x\}\)\-\\bm\{v\}\\\|^\{2\}\\Big\]
36:Update parameters \(Gradient Descent\):

37:

๐œฝโ†๐œฝโˆ’ฮท๐’‚โ€‹โˆ‡๐œฝโ„’CAFM\\bm\{\\theta\}\\leftarrow\\bm\{\\theta\}\-\\eta\_\{\\bm\{a\}\}\\nabla\_\{\\bm\{\\theta\}\}\\mathcal\{L\}\_\{\\text\{CAFM\}\}
38:

๐ƒโ†๐ƒโˆ’ฮท๐’—โ€‹โˆ‡๐ƒโ„’๐’—0\\bm\{\\xi\}\\leftarrow\\bm\{\\xi\}\-\\eta\_\{\\bm\{v\}\}\\nabla\_\{\\bm\{\\xi\}\}\\mathcal\{L\}\_\{\\bm\{v\}\_\{0\}\}
39:endwhile

## Appendix EDiscussion

### E\.1Relation to other Algorithms

#### Relation to 3 MSBM

3 MSBM considers the following stochastic optimal control problem:

minโก๐”ผฯtโ€‹โˆซ\[โ€–๐’‚tโ€–2\]โ€‹๐‘‘t\\min\\mathbb\{E\}\_\{\\rho\_\{t\}\}\\int\[\\\|\\bm\{a\}\_\{t\}\\\|^\{2\}\]\\mathrm\{d\}t\(81\)dโ€‹๐’Žt=Aโ€‹๐’Žtโ€‹๐‘‘t\+๐’–tโ€‹๐‘‘t\+gโ€‹dโ€‹๐‘พt,๐’™nโˆผqn=โˆซฯ€nโ€‹\(๐’™,๐’—\)โ€‹dโ€‹๐’—n\\mathrm\{d\}\\bm\{m\}\_\{t\}=A\\bm\{m\}\_\{t\}\\mathrm\{d\}t\+\\bm\{u\}\_\{t\}\\mathrm\{d\}t\+g\\mathrm\{d\}\\bm\{W\}\_\{t\},\\quad\\bm\{x\}\_\{n\}\\sim q\_\{n\}=\\int\\pi\_\{n\}\(\\bm\{x\},\\bm\{v\}\)\\mathrm\{d\}\\bm\{v\}\_\{n\}\(82\)where๐’Žt=\[๐’™t,๐’—t\]T\\bm\{m\}\_\{t\}=\[\\bm\{x\}\_\{t\},\\bm\{v\}\_\{t\}\]^\{T\},A=\[0001\]A=\\begin\{bmatrix\}0&0\\\\ 0&1\\end\{bmatrix\}, andg=\[000ฯƒ\]g=\\begin\{bmatrix\}0&0\\\\ 0&\\sigma\\end\{bmatrix\}\. This problem can be viewed as a โ€œstochasticโ€ version of the VM\-DOAT problem proposed in this paper\. 3 MSBM also provides a simulation\-free training method to regress the acceleration field\. However, 3 MSBM does not explicitly solve this stochastic optimal control problem using the form of optimal coupling plus optimal single\-particle trajectories\. Its optimal coupling requires iterative solutions similar to Rectified Flow\[[36](https://arxiv.org/html/2608.21070#bib.bib36)\]: a process of repeatedly selecting the coupling and trainingaฮธa\_\{\\theta\}\. Furthermore, it relies on a heuristic method to estimate the initial velocity \(via normal initialization and iteratively solving forward and backward SDEs\), rather than incorporating this estimation as part of the optimal control problem\. Experiments show that Tracing Flow achieves better distribution reconstruction accuracy than 3MSBM\.

#### Relation to MMFM

In MMFM, the single\-particle path is similarly specified by:

โˆซโ€–ฮณโ€ฒโ€ฒโ€‹\(t\)โ€–2โ€‹๐‘‘t,๐’™k=ฮณโก\(tk\)\\int\\\|\\gamma^\{\\prime\\prime\}\(t\)\\\|^\{2\}\\mathrm\{d\}t,\\quad\\bm\{x\}\_\{k\}=\\gamma\(t\_\{k\}\)\(83\)which corresponds to a natural spline\. However, MMFM does not solve an optimal control problem over distributions\. Furthermore, the velocity field in MMFM remains single\-valued with respect to๐’™\\bm\{x\}, which prevents it from learning trajectories that cross in the position space\.

#### Relation to OAT\-FM

OAT\-FM introduces a DOAT problem formulation consistent with this paper; however, the objective of OAT\-FM is not to genuinely solve the DOAT problem within the augmented space๐’ณร—๐’ฑ\\mathcal\{X\}\\times\\mathcal\{V\}, but rather to fine\-tune a pre\-trained FM model\. It employs the loss:

โ„’CAFM\\displaystyle\\mathcal\{L\}\_\{\\text\{CAFM\}\}=๐”ผฯ€โˆ—โ€‹\[\(๐’™0,๐’—0\),\(๐’™1,๐’—1\)\]\[ฮฑโˆฅ๐’™tโˆ’๐’™0tโˆ’๐’—0\+๐’—๐œฝ2โˆฅ2\+\(1โˆ’ฮฑ\)โˆฅ๐’—ฮธ\(๐’™t,t\)โˆ’๐’—0โˆฅ2\\displaystyle=\\mathbb\{E\}\_\{\\pi^\{\*\}\[\(\\bm\{x\}\_\{0\},\\bm\{v\}\_\{0\}\),\(\\bm\{x\}\_\{1\},\\bm\{v\}\_\{1\}\)\]\}\\big\[\\alpha\\\|\\dfrac\{\\bm\{x\}\_\{t\}\-\\bm\{x\}\_\{0\}\}\{t\}\-\\dfrac\{\\bm\{v\}\_\{0\}\+\\bm\{v\}\_\{\\bm\{\\theta\}\}\}\{2\}\\\|^\{2\}\+\(1\-\\alpha\)\\\|\\bm\{v\}\_\{\\theta\}\(\\bm\{x\}\_\{t\},t\)\-\\bm\{v\}\_\{0\}\\\|^\{2\}\+ฮฑโˆฅ๐’™1โˆ’๐’™t1โˆ’tโˆ’๐’—๐œฝ\+๐’—12โˆฅ2\+\(1โˆ’ฮฑ\)โˆฅ๐’—1โˆ’๐’—ฮธ\(๐’™t,t\)โˆฅ2\]\\displaystyle\+\\alpha\\\|\\dfrac\{\\bm\{x\}\_\{1\}\-\\bm\{x\}\_\{t\}\}\{1\-t\}\-\\dfrac\{\\bm\{v\}\_\{\\bm\{\\theta\}\}\+\\bm\{v\}\_\{1\}\}\{2\}\\\|^\{2\}\+\(1\-\\alpha\)\\\|\\bm\{v\}\_\{1\}\-\\bm\{v\}\_\{\\theta\}\(\\bm\{x\}\_\{t\},t\)\\\|^\{2\}\\big\]\(84\)to enforce the velocity field along the line segment connecting any two samples to be as parallel to that line as possible\. As a fine\-tuning approach, it effectively still addresses a distribution transport problem involving only two time points within the position space๐’ณ\\mathcal\{X\}, rather than tackling a multi\-marginal optimal control problem\. Furthermore, OAT\-FM does not establish a connection between its loss function and the cost along each trajectory12โ€‹โˆซ01โ€–ฮณยจโ€–2โ€‹๐‘‘t\\dfrac\{1\}\{2\}\\int\_\{0\}^\{1\}\\\|\{\\ddot\{\\gamma\}\}\\\|^\{2\}\\mathrm\{d\}t, nor does it explicitly learn the acceleration field๐’‚โก\(๐’™,๐’—,t\)\\bm\{a\}\(\\bm\{x\},\\bm\{v\},t\)\.

### E\.2Is๐’ฎ=๐’ณร—๐’ฑ\\mathcal\{S\}=\\mathcal\{X\}\\times\\mathcal\{V\}a Phase Space?

In the Conclusion and Limitation section, we noted that if๐’™\\bm\{x\}is regarded as the generalized coordinate and๐’—\\bm\{v\}as its time derivative, the termโˆซ12โ€‹โ€–๐’‚โ€–2โ€‹๐‘‘t\\int\\frac\{1\}\{2\}\\\|\\bm\{a\}\\\|^\{2\}\\mathrm\{d\}tcannot be interpreted as a physical action, as the action in classical mechanics typically does not involve second\-order time derivatives\. Here, we offer an alternative physical interpretation of this control objective\.

###### Proposition E\.1\.

Consider add\-dimensional second\-order dynamical system governed by the equations:

๐’™ห™\\displaystyle\\dot\{\\bm\{x\}\}=๐’—\\displaystyle=\\bm\{v\}\(85\)๐’—ห™\\displaystyle\\dot\{\\bm\{v\}\}=๐’‚\\displaystyle=\\bm\{a\}\(86\)where๐ฑ,๐ฏ,๐šโˆˆโ„d\\bm\{x\},\\bm\{v\},\\bm\{a\}\\in\\mathbb\{R\}^\{d\}\. The optimal control objective is given by:

minโก12โ€‹โˆซ0Tโ€–๐’‚โ€–2โ€‹๐‘‘t\\min\\frac\{1\}\{2\}\\int\_\{0\}^\{T\}\\\|\\bm\{a\}\\\|^\{2\}\\mathrm\{d\}t\(87\)
The evolution of this system under the optimal control law is equivalent to a Hamiltonian system in4โ€‹d4d\-dimensional phase space with the following Hamiltonian:

H=๐’‘1Tโ€‹๐’’2\+12โ€‹โ€–๐’‘2โ€–2H=\\bm\{p\}\_\{1\}^\{T\}\\bm\{q\}\_\{2\}\+\\frac\{1\}\{2\}\\\|\\bm\{p\}\_\{2\}\\\|^\{2\}\(88\)where๐ช1,๐ช2โˆˆโ„d\\bm\{q\}\_\{1\},\\bm\{q\}\_\{2\}\\in\\mathbb\{R\}^\{d\}are the canonical coordinates, and๐ฉ1,๐ฉ2โˆˆโ„d\\bm\{p\}\_\{1\},\\bm\{p\}\_\{2\}\\in\\mathbb\{R\}^\{d\}are the canonical momenta \(note that๐ช1\\bm\{q\}\_\{1\}does not appear in the Hamiltonian\)\. Here,๐ช1\\bm\{q\}\_\{1\}and๐ช2\\bm\{q\}\_\{2\}correspond to๐ฑ\\bm\{x\}and๐ฏ\\bm\{v\}in the optimal control problem, respectively, while๐ฉ2\\bm\{p\}\_\{2\}corresponds to๐š\\bm\{a\}\.

###### Proof\.

We derive the equations of motion directly using Hamiltonโ€™s canonical equations:

๐’’ห™1=โˆ‚Hโˆ‚๐’‘1=๐’’2,๐’’ห™2=โˆ‚Hโˆ‚๐’‘2=๐’‘2,๐’‘ห™1=โˆ’โˆ‚Hโˆ‚๐’’1=๐ŸŽ,๐’‘ห™2=โˆ’โˆ‚Hโˆ‚๐’’2=โˆ’๐’‘1\\dot\{\\bm\{q\}\}\_\{1\}=\\frac\{\\partial H\}\{\\partial\\bm\{p\}\_\{1\}\}=\\bm\{q\}\_\{2\},\\quad\\dot\{\\bm\{q\}\}\_\{2\}=\\frac\{\\partial H\}\{\\partial\\bm\{p\}\_\{2\}\}=\\bm\{p\}\_\{2\},\\quad\\dot\{\\bm\{p\}\}\_\{1\}=\-\\frac\{\\partial H\}\{\\partial\\bm\{q\}\_\{1\}\}=\\bm\{0\},\\quad\\dot\{\\bm\{p\}\}\_\{2\}=\-\\frac\{\\partial H\}\{\\partial\\bm\{q\}\_\{2\}\}=\-\\bm\{p\}\_\{1\}\(89\)
This implies that๐’‘1\\bm\{p\}\_\{1\}is constant\. Solving these equations sequentially yields:

๐’‘2โ€‹\(t\)\\displaystyle\\bm\{p\}\_\{2\}\(t\)=๐’„2โ€ฒโˆ’๐’‘1โ€‹t\\displaystyle=\\bm\{c\}\_\{2\}^\{\\prime\}\-\\bm\{p\}\_\{1\}t\(90\)๐’’2โ€‹\(t\)\\displaystyle\\bm\{q\}\_\{2\}\(t\)=๐’„1โ€ฒ\+๐’„2โ€ฒโ€‹tโˆ’12โ€‹๐’‘1โ€‹t2\\displaystyle=\\bm\{c\}\_\{1\}^\{\\prime\}\+\\bm\{c\}\_\{2\}^\{\\prime\}t\-\\frac\{1\}\{2\}\\bm\{p\}\_\{1\}t^\{2\}\(91\)๐’’1โ€‹\(t\)\\displaystyle\\bm\{q\}\_\{1\}\(t\)=๐’„0โ€ฒ\+๐’„1โ€ฒ\+12โ€‹๐’„2โ€ฒโ€‹t2โˆ’16โ€‹๐’‘1โ€‹t3\\displaystyle=\\bm\{c\}\_\{0\}^\{\\prime\}\+\\bm\{c\}\_\{1\}^\{\\prime\}\+\\frac\{1\}\{2\}\\bm\{c\}\_\{2\}^\{\\prime\}t^\{2\}\-\\frac\{1\}\{6\}\\bm\{p\}\_\{1\}t^\{3\}\(92\)Comparing this with the results in[Equation9](https://arxiv.org/html/2608.21070#S4.E9), we observe that๐’’1โ€‹\(t\)\\bm\{q\}\_\{1\}\(t\)and๐’’2โ€‹\(t\)\\bm\{q\}\_\{2\}\(t\)satisfy the same equations as๐’™โก\(t\)\\bm\{x\}\(t\)and๐’—โก\(t\)\\bm\{v\}\(t\), with the undetermined constants determined by the boundary conditions\. This completes the proof\. โˆŽ

Consequently, we observe that although the augmented space๐’ฎ=๐’ณร—๐’ฑ\\mathcal\{S\}=\\mathcal\{X\}\\times\\mathcal\{V\}concatenates position๐’™\\bm\{x\}and velocity \(momentum\)๐’—\\bm\{v\}, itcannotbe interpreted as aphase space\. Instead, it should be viewed as theconfiguration spaceof the classical mechanical system described by the Hamiltonian above\. We can attempt to recover the Lagrangian from this Hamiltonian\. The canonical equations established that๐’’ห™1=๐’’2\\dot\{\\bm\{q\}\}\_\{1\}=\\bm\{q\}\_\{2\}and๐’’ห™2=๐’‘2\\dot\{\\bm\{q\}\}\_\{2\}=\\bm\{p\}\_\{2\}\. Applying the Legendre transformation:

L\\displaystyle L=๐’‘1Tโ€‹๐’’ห™1\+๐’‘2Tโ€‹๐’’ห™2โˆ’H\\displaystyle=\\bm\{p\}\_\{1\}^\{T\}\\dot\{\\bm\{q\}\}\_\{1\}\+\\bm\{p\}\_\{2\}^\{T\}\\dot\{\\bm\{q\}\}\_\{2\}\-H=๐’‘1Tโ€‹\(๐’’ห™1โˆ’๐’’2\)\+12โ€‹โ€–๐’’ห™2โ€–2\\displaystyle=\\bm\{p\}\_\{1\}^\{T\}\(\\dot\{\\bm\{q\}\}\_\{1\}\-\\bm\{q\}\_\{2\}\)\+\\frac\{1\}\{2\}\\\|\\dot\{\\bm\{q\}\}\_\{2\}\\\|^\{2\}\(93\)We find that the term๐’‘1T\\bm\{p\}\_\{1\}^\{T\}cannot be eliminated; the Lagrangian cannot be expressed solely as a function of the canonical coordinates and their first derivatives\. This is a characteristic of constrained systems \(in Diracโ€™s theory of constraints, a constrained system is typically defined as one where the Hessian matrix of the Lagrangian is singular\. To see this intuitively, if we regard๐’‘1\\bm\{p\}\_\{1\}in the Lagrangian asddnew canonical coordinates, their corresponding canonical momenta are zero, implying the motion is confined to a submanifold of the phase space\)\.

In[Equation10](https://arxiv.org/html/2608.21070#S4.E10), we derived the minimum cost for the system to transition from\(๐’™1,๐’—1\)\(\\bm\{x\}\_\{1\},\\bm\{v\}\_\{1\}\)to\(๐’™2,๐’—2\)\(\\bm\{x\}\_\{2\},\\bm\{v\}\_\{2\}\)\. While this cost is a symmetric positive\-definite quadratic form with respect to๐’™\\bm\{x\}and๐’—\\bm\{v\}, it does not constitute a metric on the augmented space๐’ฎ=๐’ณร—๐’ฑ\\mathcal\{S\}=\\mathcal\{X\}\\times\\mathcal\{V\}\. \(If it were a metric, the augmented space would be flat, and the optimal control trajectoriesโ€”geodesicsโ€”would be linear functions oftt\)\. Maupertuisโ€™ principle suggests a close relationship between the metric in the configuration space and the systemโ€™s Lagrangian\[[3](https://arxiv.org/html/2608.21070#bib.bib3)\]: if the Lagrangian takes the form:

Lโก\(๐’’,๐’’ห™\)=12โ€‹miโ€‹jโ€‹\(๐’’\)โ€‹qห™iโ€‹qห™jโˆ’Vโก\(๐’’\)L\(\\bm\{q\},\\dot\{\\bm\{q\}\}\)=\\frac\{1\}\{2\}m\_\{ij\}\(\\bm\{q\}\)\\dot\{q\}^\{i\}\\dot\{q\}^\{j\}\-V\(\\bm\{q\}\)\(94\)then the configuration space possesses the metric:

giโ€‹jโ€‹\(๐’’\)=2โ€‹\(Eโˆ’Vโก\(๐’’\)\)โ€‹miโ€‹jโ€‹\(๐’’\)g\_\{ij\}\(\\bm\{q\}\)=2\(E\-V\(\\bm\{q\}\)\)m\_\{ij\}\(\\bm\{q\}\)\(95\)whereEEis the total energy of the particle\. For the optimal control problem discussed herein, we have seen that the Lagrangian cannot be written in such a form; therefore, it is fundamentally impossible to equip the configuration space๐’ฎ=๐’ณร—๐’ฑ\\mathcal\{S\}=\\mathcal\{X\}\\times\\mathcal\{V\}with such a metric\.

### E\.3Broader Impacts

This paper presents work whose goal is to advance the field of machine learning\. There are many potential societal consequences of our work, none of which we feel must be specifically highlighted here\. However, as our algorithm is applied to single\-cell data analysis, the fidelity of the generated trajectories is heavily dependent on the quality of the input data\. Consequently, when utilizing our method for biological or medical research purposes, it is critical to employ high\-quality datasets, incorporate established and accurate biological priors, and ensure that the algorithmic outputs are rigorously validated by domain experts\.

Similar Articles

Trajectory as the Teacher: Few-Step Discrete Flow Matching via Energy-Navigated Distillation

Hugging Face Daily Papers

This paper introduces Trajectory-Shaped Discrete Flow Matching (TS-DFM), which replaces blind stochastic jumps with guided navigation to significantly improve text generation efficiency and reduce computational costs. The method achieves superior perplexity and speed compared to traditional multi-step baselines while maintaining unchanged inference costs.

Recursive Flow Matching

Hugging Face Daily Papers

Introduces Recursive Flow Matching (RecFM), a generative framework for forecasting complex spatiotemporal dynamics that achieves high fidelity with fewer steps and improved accuracy and speed, including up to 20x speedup over diffusion-based emulators.