当Riemann与Wasserstein流动:流形上概率分布的生成建模
摘要
本文介绍了Riemannian Wasserstein Entropic Flow Matching(RWEFM),这是一个用于建模黎曼流形上概率分布的生成框架,应用于单细胞生物学和蛋白质构象等科学领域。
arXiv:2609.25659v1 Announce Type: new
Abstract: Many scientific datasets, such as molecular conformational ensembles or single-cell tissue measurements, are naturally modeled as meta-distributions: distributions over probability measures on non-Euclidean domains. Existing generative methods largely assume Euclidean geometry and fail to capture this structure. We introduce Riemannian Wasserstein Entropic Flow Matching (RWEFM), a generative framework on the Wasserstein space $\mathcal{P}_2(\mathcal{M})$ of a Riemannian manifold $(\mathcal{M},g)$. RWEFM is trained by regressing a neural vector field onto Riemannian optimal transport velocities, using McCann displacement interpolations as conditional paths. We confirm theoretically that this construction leads to a valid flow matching approach on $\mathcal{P}_2(\mathcal{M})$ and introduce the Riemannian Entropic Map, a GPU-efficient approximation of the optimal transport map on manifolds. Our experiments show that by respecting the intrinsic geometry of the data, RWEFM can generate whole single-cell samples in hyperspherical latent spaces and protein conformational ensembles on the torus. As RWEFM requires only a geodesic distance and a projection operator, it is not restricted to manifolds with closed-form geometry, which we demonstrate by generating distributions on a general triangulated mesh.
查看缓存全文
缓存时间: 2026/09/23 09:33
# When Riemann flows with Wasserstein:Generative Modeling of Probability Distributions on Manifolds
Source: [https://arxiv.org/html/2609.25659](https://arxiv.org/html/2609.25659)
Edward De Brouwer∗,1Rishabh Anand2Rex Ying2Aïcha Bentaieb1Gabriele Scalia1Hector Corrada Bravo1
###### Abstract
Many scientific datasets, such as molecular conformational ensembles or single\-cell tissue measurements, are naturally modeled as meta\-distributions: distributions over probability measures on non\-Euclidean domains\. Existing generative methods largely assume Euclidean geometry and fail to capture this structure\. We introduce Riemannian Wasserstein Entropic Flow Matching \(RWEFM\), a generative framework on the Wasserstein space𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)of a Riemannian manifold\(ℳ,g\)\(\\mathcal\{M\},g\)\. RWEFM is trained by regressing a neural vector field onto Riemannian optimal transport velocities, using McCann displacement interpolations as conditional paths\. We confirm theoretically that this construction leads to a valid flow matching approach on𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)and introduce the Riemannian Entropic Map, a GPU\-efficient approximation of the optimal transport map on manifolds\. Our experiments show that by respecting the intrinsic geometry of the data, RWEFM can generate whole single\-cell samples in hyperspherical latent spaces and protein conformational ensembles on the torus\. As RWEFM requires only a geodesic distance and a projection operator, it is not restricted to manifolds with closed\-form geometry, which we demonstrate by generating distributions on a general triangulated mesh\. Code and tutorials are available at[RWEFM](https://github.com/DoronHav/WassersteinFlowMatching)\.
11footnotetext:Equal contribution\.11footnotetext:Genentech Inc\., South San Francisco, CA\.havivd@gene\.com, debroue1@gene\.com\.22footnotetext:Yale University, New Haven, CT\.## 1Introduction
Many modern scientific datasets are most naturally represented as collections of*distributions*of data points living on structured non\-Euclidean spaces\. Molecular conformational ensembles\([Axelrod and Gomez\-Bombarelli, 2022](https://arxiv.org/html/2609.25659#bib.bib19)\)describe each molecule as a distribution over rotations and translations in 3D space; climate records\([Abatzoglou et al\., 2018](https://arxiv.org/html/2609.25659#bib.bib20)\)can be viewed as distributions of weather variables over the sphere; and tissue\-level single\-cell RNA\-seq\([CZI Cell Science Program et al\., 2025](https://arxiv.org/html/2609.25659#bib.bib18)\)represents each sample as an empirical distribution of cell states living on a biologically meaningful manifold \(*e\.g\.*spherical or hyperbolic\)\. In these settings, the object of interest is a distribution over probability measures on a manifold\.
Generative modeling provides a principled way to summarize such data and to enable scientific tasks, including in\-silico design\([Gruver et al\., 2023](https://arxiv.org/html/2609.25659#bib.bib17)\), hypothesis testing\([Candes et al\., 2018](https://arxiv.org/html/2609.25659#bib.bib16)\), and causal discovery in the underlying physical process\([Zhu et al\., 2023](https://arxiv.org/html/2609.25659#bib.bib9)\)\. While recent generative models have achieved impressive results in Euclidean data \(images\([Lin et al\., 2024](https://arxiv.org/html/2609.25659#bib.bib13)\), videos\([Jin et al\., 2025](https://arxiv.org/html/2609.25659#bib.bib15)\), single\-cell\([Klein et al\., 2025](https://arxiv.org/html/2609.25659#bib.bib11)\)\), and have increasingly incorporated non\-Euclidean structures \(*e\.g\.*molecules\([Schneuing et al\., 2024](https://arxiv.org/html/2609.25659#bib.bib10)\)\), most of this progress concerns generating individual data points defined as finite\-dimensional vectors\. In contrast, generative modeling of*distributions*requires operating in an infinite\-dimensional space: the Wasserstein space of probability measures𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\. Recent work has explored generative models directly on Wasserstein space\([Haviv et al\., 2024](https://arxiv.org/html/2609.25659#bib.bib26);[Atanackovic et al\., 2024](https://arxiv.org/html/2609.25659#bib.bib34);[Piening et al\., 2026](https://arxiv.org/html/2609.25659#bib.bib8)\)but these methods assume the underlying domain to be Euclidean \(*i\.e\.*ℳ≡ℝd\\mathcal\{M\}\\equiv\\mathbb\{R\}^\{d\}\), thereby failing to faithfully model datasets whose support lies on curved geometries\.
Figure 1:Riemannian Wasserstein Entropic Flow Matching\.Top:Base Riemannian manifold\(ℳ,dg\)\(\\mathcal\{M\},d\_\{g\}\)\. Each blob is a probability distributionμ,ν∈𝒫2\(ℳ\)\\mu,\\nu\\in\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)and a single training example is one blob, not one point\. The dashed line is the Riemannian OT path betweenμ\\muandν\\nuwhich is estimated via point\-cloud samples from each blob using the entropic map\.Bottom:Wasserstein manifold\(𝒫2\(ℳ\),𝒲2\)\(\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\),\\mathcal\{W\}\_\{2\}\), where ablobis a distribution\-over\-distributions and each point is a distribution in the top figure\. The same dashed path lifts here as the McCann interpolantμt\\mu\_\{t\}, the geodesic along which RWEFM learns its flow\.In this work, we introduce Riemannian Wasserstein Entropic Flow Matching \(RWEFM\), a flow\-matching \(FM\) framework\([Lipman et al\., 2022](https://arxiv.org/html/2609.25659#bib.bib22)\)for generative modeling on𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\. RWEFM lifts conditional FM to the Wasserstein space using McCann displacement interpolations as conditional paths between paired measures\(μ,ν\)\(\\mu,\\nu\)\. We establish that this yields a valid FM objective on𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)and that the induced marginal dynamics satisfy a weak continuity equation and admit a path\-space representation via the superposition principle\([Pinzi and Savaré, 2025](https://arxiv.org/html/2609.25659#bib.bib41)\)\.
A key practical bottleneck is the computation of optimal transport maps between empirical distributions on Riemannian manifolds\. To this end, we introduce the Riemannian Entropic Map, a GPU\-efficient approximation of the Monge map that extends the entropic map\([Pooladian and Niles\-Weed, 2021](https://arxiv.org/html/2609.25659#bib.bib21)\)to general geometries\. In essence, our estimator is obtained as the exponential map of the barycentric projection of tangent vectors\. We further provide an error bound that quantifies the approximation induced by regularization and finite sampling\.
We evaluate RWEFM across a range of geometries and scientific applications\. This includes generation of digits, letters and Kanji characters on the sphere𝕊2\\mathbb\{S\}^\{2\}, hyperbolic diskℍ2\\mathbb\{H\}^\{2\}and torus𝕋2\\mathbb\{T\}^\{2\}, as well as whole single\-cell sample generation in hyperspherical latent spaces𝕊128−1\\mathbb\{S\}^\{128\-1\}and protein conformational ensembles on𝕋2\\mathbb\{T\}^\{2\}\. To emphasize that our framework is not limited to closed\-form geometries, we further generate distributions on a general triangulated mesh\. Our experiments show that respecting the intrinsic geometry of the data and using the Riemannian Entropic Map improves generation quality\.
##### Contributions
\(i\)We propose RWEFM, a framework for generative modeling on the space of probability distributions on Riemannian manifolds—including general geometries without closed\-form geodesics, such as triangulated meshes—and show that the resulting marginal dynamics are theoretically well\-posed\.\(ii\)We introduce the Riemannian Entropic Map, a fast and accurate approximation of Riemannian optimal transport maps suitable for large\-scale GPU training, together with theoretical guarantees\.\(iii\)We benchmark RWEFM against FM baselines on real\-world datasets such as generation of single\-cell samples and protein conformational ensembles\.
## 2Background and Related Work
### 2\.1Optimal Transport on Riemannian Manifolds
Optimal Transport \(OT\)\([Villani, 2008](https://arxiv.org/html/2609.25659#bib.bib28)\)provides a geometric framework for comparing probability distributions on Riemannian manifolds\. Let\(ℳ,g\)\(\\mathcal\{M\},g\)be a complete Riemannian manifold equipped with a metricggand its corresponding geodesic distancedg\(⋅,⋅\)d\_\{g\}\(\\cdot,\\cdot\), and letμ,ν∈𝒫2\(ℳ\)\\mu,\\nu\\in\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)be two probability measures with finite second moments\. The Monge formulation seeks a deterministic mapT:ℳ→ℳT:\\mathcal\{M\}\\to\\mathcal\{M\}pushingμ\\muontoν\\nu\(denotedT\#μ=νT\_\{\\\#\}\\mu=\\nu\) that minimizes the total transportation cost:
infT\{∫ℳc\(x,T\(x\)\)𝑑μ\(x\):T\#μ=ν\},\\inf\_\{T\}\\left\\\{\\int\_\{\\mathcal\{M\}\}c\(x,T\(x\)\)d\\mu\(x\):T\_\{\\\#\}\\mu=\\nu\\right\\\},\(1\)
wherec\(x,y\)=12dg2\(x,y\)c\(x,y\)=\\frac\{1\}\{2\}d\_\{g\}^\{2\}\(x,y\)is the cost of moving a unit of mass fromxxtoyy\. Given mild assumptions, such asμ\\mubeing absolutely continuous with respect to the volume measure, the Brenier–McCann theorem\([Ambrosio et al\., 2021](https://arxiv.org/html/2609.25659#bib.bib3)\)states that there exists a unique optimal transport mapT0T\_\{0\}, known as the Monge map, of the formT0\(x\)=expx\(−∇φ0\(x\)\)T\_\{0\}\(x\)=\\text\{exp\}\_\{x\}\(\-\\nabla\\varphi\_\{0\}\(x\)\)whereφ0:ℳ→ℝ\\varphi\_\{0\}:\\mathcal\{M\}\\to\\mathbb\{R\}is a c\-concave function\.
The Kantorovich formulation relaxes the Monge problem by allowing for probabilistic couplings:
W22\(μ,ν\)=infπ∈Π\(μ,ν\)∫ℳ×ℳdg2\(x,y\)𝑑π\(x,y\)\.W\_\{2\}^\{2\}\(\\mu,\\nu\)=\\inf\_\{\\pi\\in\\Pi\(\\mu,\\nu\)\}\\int\_\{\\mathcal\{M\}\\times\\mathcal\{M\}\}d\_\{g\}^\{2\}\(x,y\)d\\pi\(x,y\)\.whereΠ\(μ,ν\)\\Pi\(\\mu,\\nu\)is the set of joint distributions onℳ×ℳ\\mathcal\{M\}\\times\\mathcal\{M\}with marginalsμ\\muandν\\nu\. The optimal valueW2\(μ,ν\)W\_\{2\}\(\\mu,\\nu\)defines the 2\-Wasserstein distance between the measures\. Unlike the Monge formulation, the Kantorovich problem does not require continuity of measures to admit a solution\. When the Monge mapT0T\_\{0\}exists, it can be recovered from the Kantorovich problem via the dual formulation:12W22\(μ,ν\)=supφ∫ℳφ𝑑μ\+∫ℳφc𝑑ν,\\frac\{1\}\{2\}W\_\{2\}^\{2\}\(\\mu,\\nu\)=\\sup\_\{\\varphi\}\\int\_\{\\mathcal\{M\}\}\\varphi d\\mu\+\\int\_\{\\mathcal\{M\}\}\\varphi^\{c\}d\\nu,withφc\(y\)=infx∈ℳ\{12d\(x,y\)2−φ\(x\)\}\\varphi^\{c\}\(y\)=\\inf\_\{x\\in\\mathcal\{M\}\}\\\{\\frac\{1\}\{2\}d\(x,y\)^\{2\}\-\\varphi\(x\)\\\}, and the Monge map is thenT0\(x\)=expx\(−∇φ0\(x\)\)T\_\{0\}\(x\)=\\text\{exp\}\_\{x\}\(\-\\nabla\\varphi\_\{0\}\(x\)\)\.
#### 2\.1\.1Statistical Estimation of Optimal Transport
Closed\-form solutions for optimal transport maps are typically limited to very simple distributions so we generally rely on statistical estimation from finite samples\{xi\}i=1m∼μ\\\{x\_\{i\}\\\}\_\{i=1\}^\{m\}\\sim\\muand\{yj\}j=1n∼ν\\\{y\_\{j\}\\\}\_\{j=1\}^\{n\}\\sim\\nu\. However, the Kantorovich problem on discrete samples is a linear program with cubic complexity in the number of samples, hindering its application to large datasets\. Instead, practical approaches rely on entropic OT, which regularizes the \(discrete\) objective with an entropy termH\(π\)H\(\\pi\)\. The problem is strictly convex and can be efficiently solved using the Sinkhorn algorithm\([Cuturi, 2013](https://arxiv.org/html/2609.25659#bib.bib43)\)\.
minπ∈Π\(μ^,ν^\)∑i=1m∑j=1nc\(xi,yj\)πij\+ε∑i=1m∑j=1nπijlogπij\.\\min\_\{\\pi\\in\\Pi\(\\hat\{\\mu\},\\hat\{\\nu\}\)\}\\sum\_\{i=1\}^\{m\}\\sum\_\{j=1\}^\{n\}c\(x\_\{i\},y\_\{j\}\)\\pi\_\{ij\}\+\\varepsilon\\sum\_\{i=1\}^\{m\}\\sum\_\{j=1\}^\{n\}\\pi\_\{ij\}\\log\\pi\_\{ij\}\.This problem admits the dual formulation:
supf∈L1\(μ\)g∈L1\(ν\)∑if\(xi\)\+∑jg\(yj\)−ε∑i=1m∑j=1ne\(f\(xi\)\+g\(yj\)−12dg\(xi,yj\)2\)/ε\+ε\\sup\_\{\\begin\{subarray\}\{c\}f\\in L^\{1\}\(\\mu\)\\\\ g\\in L^\{1\}\(\\nu\)\\end\{subarray\}\}\\;\\sum\_\{i\}f\(x\_\{i\}\)\+\\sum\_\{j\}g\(y\_\{j\}\)\-\\varepsilon\\sum\_\{i=1\}^\{m\}\\sum\_\{j=1\}^\{n\}e^\{\\left\(f\(x\_\{i\}\)\+g\(y\_\{j\}\)\-\\frac\{1\}\{2\}d\_\{g\}\(x\_\{i\},y\_\{j\}\)^\{2\}\\right\)/\\varepsilon\}\+\\varepsilon\(2\)While entropic optimal transport is widely used for computing the optimal transport distances,[Pooladian and Niles\-Weed \(2021\)](https://arxiv.org/html/2609.25659#bib.bib21)showed that it can be used to compute a tractable estimator of the Monge mapT0T\_\{0\}\. Given the optimal entropic couplingπε\\pi\_\{\\varepsilon\}, theentropic mapis the barycentric projectionTε\(xi\)=𝔼πε\[Y\|X=xi\]T\_\{\\varepsilon\}\(x\_\{i\}\)=\\mathbb\{E\}\_\{\\pi\_\{\\varepsilon\}\}\[Y\|X=x\_\{i\}\]which, under suitable regularity assumptions and an appropriate joint choice of regularization and sample size, consistently estimates the Monge mapT0T\_\{0\}\. The optimal entropic potentials\(fε,gε\)\(f\_\{\\varepsilon\},g\_\{\\varepsilon\}\)are the solutions of Eq \([2](https://arxiv.org/html/2609.25659#S2.E2)\)\.
#### 2\.1\.2Wasserstein Geometry
The Wasserstein space over a Riemannian manifold\(ℳ,g\)\(\\mathcal\{M\},g\), denoted𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\), is the space of probability measures onℳ\\mathcal\{M\}with finite second moments, equipped with the 2\-Wasserstein distanceW2W\_\{2\}\. While not rigorously a Riemannian manifold due to infinite dimensionality, it can be endowed with a Riemannian\-like structure\. There, the tangent space at a measureμ∈𝒫2\(ℳ\)\\mu\\in\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)is identified with the closure of the set of gradients of smooth functions inL2\(μ,Tℳ\)L^\{2\}\(\\mu,T\\mathcal\{M\}\):
Tμ𝒫2\(ℳ\)=\{v=∇ϕ:ϕ∈Cc∞\(ℳ\)\}¯L2\(μ\),T\_\{\\mu\}\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)=\\overline\{\\\{v=\\nabla\\phi:\\phi\\in C\_\{c\}^\{\\infty\}\(\\mathcal\{M\}\)\\\}\}^\{L^\{2\}\(\\mu\)\},and endowed with the norm∥v∥L2\(μ\)2=∫ℳ∥v∥g2𝑑μ\(x\)\\lVert v\\rVert^\{2\}\_\{L^\{2\}\(\\mu\)\}=\\int\_\{\\mathcal\{M\}\}\\lVert v\\rVert^\{2\}\_\{g\}d\\mu\(x\)\. The exponential and logarithm maps in this space are:expμ\(v\):=expx\(v\(x\)\)\#μ,logμ\(ν\):=−∇φ\(x\)\\exp\_\{\\mu\}\(v\):=\\exp\_\{x\}\(v\(x\)\)\_\{\\\#\}\\mu,\\,\\log\_\{\\mu\}\(\\nu\):=\-\\nabla\\varphi\(x\)\\,, whereTμ→ν=expx\(−∇φ\(x\)\)T^\{\\mu\\to\\nu\}=\\exp\_\{x\}\(\-\\nabla\\varphi\(x\)\)is the optimal transport map fromμ\\mutoν\\nu\. Geodesics in𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)between measuresμ\\muandν\\nuare probability paths\(μt\)t∈\[0,1\]\(\\mu\_\{t\}\)\_\{t\\in\[0,1\]\}where:
μt=expμ\(tlogμ\(ν\)\)=expx\(−t∇φ\(x\)\)♯μ\\mu\_\{t\}=\\exp\_\{\\mu\}\(t\\log\_\{\\mu\}\(\\nu\)\)=\\exp\_\{x\}\(\-t\\nabla\\varphi\(x\)\)\_\{\\sharp\}\\mu
### 2\.2Flow Matching on Riemannian Manifolds and the Wasserstein Space
Flow matching generates samples from a target distributionν\\nufrom samples of a source distributionμ\\muby learning a time\-dependent vector field that generates a probability pathμt\\mu\_\{t\}such thatμ0=μ\\mu\_\{0\}=\\muandμ1=ν\\mu\_\{1\}=\\nu\. Directly learning such a vector field is generally impossible\. Instead, flow matching defines conditional probability pathsμt\(⋅\|z\)\\mu\_\{t\}\(\\cdot\|z\)for which computing a generating vector field is tractable\.
##### Ingredients of flow matching
This construction relies on three key components: \(C1\) a condition for when a vector field generates a probability path, \(C2\) the marginalization of the conditional vector field generates the marginal probability path, and \(C3\) the losses from regressing the learnable vector field to the conditional or the marginal vector field are equivalent\. The continuity equation on\(ℳ,g\)\(\\mathcal\{M\},g\),∂tμt\(x\)\+∇g⋅\(μt\(x\)vt\(x\)\)=0\\partial\_\{t\}\\mu\_\{t\}\(x\)\+\\nabla\_\{g\}\\cdot\(\\mu\_\{t\}\(x\)v\_\{t\}\(x\)\)=0, plays the role of the first component\. The second is satisfied with the marginal vector field defined as:
vt\(x\)=∫vt\(x\|z\)μt\(x\|z\)μt\(x\)𝑑π\(z\),v\_\{t\}\(x\)=\\int v\_\{t\}\(x\|z\)\\frac\{\\mu\_\{t\}\(x\|z\)\}\{\\mu\_\{t\}\(x\)\}d\\pi\(z\),Indeed, if\(μt\(⋅\|z\),vt\(⋅\|z\)\)\(\\mu\_\{t\}\(\\cdot\|z\),v\_\{t\}\(\\cdot\|z\)\)solves the continuity equation,\(μt,vt\)\(\\mu\_\{t\},v\_\{t\}\)is also a solution\. We usez=\(x0,x1\)∼πz=\(x\_\{0\},x\_\{1\}\)\\sim\\piandπ\\piis a coupling fromμ\\muandν\\nu, for which multiple options have been identified in the literature\([Pooladian et al\., 2023](https://arxiv.org/html/2609.25659#bib.bib35);[Lipman et al\., 2022](https://arxiv.org/html/2609.25659#bib.bib22);[Tong et al\., 2023](https://arxiv.org/html/2609.25659#bib.bib23)\)\. A typical conditional path is then a dirac centered on a curvext\(x0,x1\)x\_\{t\}\(x\_\{0\},x\_\{1\}\)that interpolates betweenx0x\_\{0\}andx1x\_\{1\}\. In this case, the generating vector field is given byvt\(x\|z\)=x˙t=ddtxt\(x0,x1\)v\_\{t\}\(x\|z\)=\\dot\{x\}\_\{t\}=\\frac\{d\}\{dt\}x\_\{t\}\(x\_\{0\},x\_\{1\}\)\. Finally, the equivalence of losses \(C3\) follows from the second point, as shown in\([Lipman et al\., 2022](https://arxiv.org/html/2609.25659#bib.bib22)\)\(see[Lipman et al\. \(2024\)](https://arxiv.org/html/2609.25659#bib.bib25)for a comprehensive tutorial\)\. This leads to the celebrated flow matching objective:
𝔼t∼𝒰\[0,1\],\(x0,x1\)∼π,xt∼μt\(⋅∣z\)\[∥vθ\(xt,t\)−x˙t∥g\(xt\)2\]\\displaystyle\\mathbb\{E\}\_\{t\\sim\\mathcal\{U\}\[0,1\],\(x\_\{0\},x\_\{1\}\)\\sim\\pi,x\_\{t\}\\sim\\mu\_\{t\}\(\\cdot\\mid z\)\}\\left\[\\\|v\_\{\\theta\}\(x\_\{t\},t\)\-\\dot\{x\}\_\{t\}\\\|\_\{g\(x\_\{t\}\)\}^\{2\}\\right\]
Recently[Haviv et al\. \(2024\)](https://arxiv.org/html/2609.25659#bib.bib26)extended this framework to the Wasserstein space over Euclidean domains, termed Wasserstein Flow Matching \(WFM\)\. They consider distributions on the Wasserstein space \(*i\.e\.*probability distribution on the space of probability distributions\),ℙ0,ℙ1∈𝒫2\(𝒫2\(ℝd\)\)\\mathbb\{P\}\_\{0\},\\mathbb\{P\}\_\{1\}\\in\\mathcal\{P\}\_\{2\}\(\\mathcal\{P\}\_\{2\}\(\\mathbb\{R\}^\{d\}\)\)\. In practice, samples fromℙ0\\mathbb\{P\}\_\{0\}andℙ1\\mathbb\{P\}\_\{1\}are represented as empirical distributions or point clouds\(X0,X1\)\(X\_\{0\},X\_\{1\}\)andXtX\_\{t\}is defined as the McCann interpolant:Xt=\(1−t\)X0\+tT^X0→X1\(X0\)X\_\{t\}=\(1\-t\)X\_\{0\}\+t\\hat\{T\}^\{X\_\{0\}\\to X\_\{1\}\}\(X\_\{0\}\), withT^X0→X1\\hat\{T\}^\{X\_\{0\}\\to X\_\{1\}\}the optimal transport map betweenX0X\_\{0\}andX1X\_\{1\}\. One then learns the target vector fielduθ\(Xt,t\)∈T𝒫2\(ℝd\)u\_\{\\theta\}\(X\_\{t\},t\)\\in T\\mathcal\{P\}\_\{2\}\(\\mathbb\{R\}^\{d\}\)by regressing it againstX˙t=T^X0→X1\(X0\)−X0\\dot\{X\}\_\{t\}=\\hat\{T\}^\{X\_\{0\}\\to X\_\{1\}\}\(X\_\{0\}\)\-X\_\{0\}\.
Figure 2:Generation of distributions on non\-Euclidean spacesExamples of empirical distributions generated by the proposed Riemannian Wasserstein Entropic Flow Matching model on non\-Euclidean geometries\. The top row displays MNIST digits \(3, 4, and 8\) generated on the sphere𝕊2\\mathbb\{S\}^\{2\}shown in 2D via Mollweide projections\. The bottom row shows EMNIST characters \(H, W, and Y\) generated on the hyperbolic planeℍ2\\mathbb\{H\}^\{2\}, visualized using the Poincaré disk model\.
### 2\.3Related Work
##### Optimal Transport on Riemannian Manifolds\.
Optimal Transport on Riemannian manifolds has been extensively studied in the mathematical literature, with foundational results on the existence and uniqueness of Monge maps, duality theory, and regularity properties\([McCann, 2001](https://arxiv.org/html/2609.25659#bib.bib27);[Villani, 2008](https://arxiv.org/html/2609.25659#bib.bib28)\)\. These theoretical insights have paved the way for many applications, particularly in generative modeling for non\-Euclidean data\([Bose et al\., 2023](https://arxiv.org/html/2609.25659#bib.bib29);[De Bortoli et al\., 2022](https://arxiv.org/html/2609.25659#bib.bib12);[Huguet et al\., 2023](https://arxiv.org/html/2609.25659#bib.bib14)\)\. Relatedly,[You \(2026\)](https://arxiv.org/html/2609.25659#bib.bib5)study intrinsic and tangential barycentric projections of transport plans on Riemannian manifolds\. Recent works have also explored Optimal Transport on Wasserstein spaces themselves, proving existence of continuity equations and geodesics despite the infinite dimensionality, enriching the theoretical framework and expanding its applicability\([Bonet et al\., 2025](https://arxiv.org/html/2609.25659#bib.bib24);[Emami and Pass, 2025](https://arxiv.org/html/2609.25659#bib.bib30);[Pinzi and Savaré, 2025](https://arxiv.org/html/2609.25659#bib.bib41)\)\.
Table 1:Comparison of FM approaches\.Only RWEFM respects both the Riemannian structure of the data space and the Wasserstein structure of the distribution space\.
##### Generative Modeling of Distributions
Our work relates to recent efforts in defining generative models over spaces of probability measures\. Fisher FM\([Davis et al\., 2024](https://arxiv.org/html/2609.25659#bib.bib31)\)and Categorical FM\([Cheng et al\., 2024](https://arxiv.org/html/2609.25659#bib.bib32)\)apply the Flow Matching framework to the simplexΔd\\Delta\_\{d\}equipped with the Fisher–Rao geometry, focusing on categorical data\. Similarly,[Stark et al\. \(2024\)](https://arxiv.org/html/2609.25659#bib.bib33)utilize the Dirichlet distribution for discrete data generation\. For continuous distributions, Meta FM\([Atanackovic et al\., 2024](https://arxiv.org/html/2609.25659#bib.bib34)\)and Wasserstein FM\([Haviv et al\., 2024](https://arxiv.org/html/2609.25659#bib.bib26);[Piening et al\., 2026](https://arxiv.org/html/2609.25659#bib.bib8)\)learn flows directly on the Wasserstein space𝒫2\(ℝd\)\\mathcal\{P\}\_\{2\}\(\\mathbb\{R\}^\{d\}\)\. Crucially, these approaches are restricted to Euclidean base spaces, while our framework extends to distributions over general Riemannian manifolds\.
## 3Riemannian Wasserstein Entropic FM
### 3\.1Flow Matching on𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)
RWEFM learns generative flows on𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)by lifting the conditional FM paradigm to𝒫2\(𝒫2\(ℳ\)\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\), respecting the intrinsic geometry ofℳ\\mathcal\{M\}and𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\. Following Section[2\.2](https://arxiv.org/html/2609.25659#S2.SS2), this requires three components \(C1–C3\)\. We start with \(C1\), the conditions for when a vector fieldVt∈T𝒫2\(ℳ\)V\_\{t\}\\in T\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)generates a probability pathℙt∈𝒫2\(𝒫2\(ℳ\)\)\\mathbb\{P\}\_\{t\}\\in\\mathcal\{P\}\_\{2\}\(\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\)\.
#### 3\.1\.1Weak continuity equation and superposition on𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\(C1\)
Table 2:RWEFM on MNIST, EMNIST, and KMNIST\.Benchmarking of RWEFM against FM variants and other point\-cloud generative models on hyperbolic \(ℍ2\\mathbb\{H\}^\{2\}, EMNIST\), spherical \(𝕊2\\mathbb\{S\}^\{2\}, MNIST\), and toroidal \(𝕋2\\mathbb\{T\}^\{2\}, KMNIST\) data\. We report the classwise 1\-NN deviation \(1\-NN\-D\) using Chamfer Distance \(CD\) and Earth Mover’s Distance \(EMD\) between generated and test point clouds\. The score ranges from00to0\.50\.5, where00corresponds to50%50\\%accuracy in each class; lower is better\. Values are averaged over 3 random seeds and 5 samplings per seed\. The best three results per column are highlighted \(1st,2nd,3rd\); ties share a rank\. Dashes indicate methods not evaluated on that manifold\. Full results with MMD metrics and per\-manifold mean±\\pmstd are reported in[tablesS3](https://arxiv.org/html/2609.25659#A4.T3),[S5](https://arxiv.org/html/2609.25659#A4.T5)and[S6](https://arxiv.org/html/2609.25659#A4.T6)\(Appendix[D](https://arxiv.org/html/2609.25659#A4)\)\. Per\-dataset training times for all methods are in[tableS1](https://arxiv.org/html/2609.25659#A4.T1)\.The following result shows that if\(ℙt,Vt\)\(\\mathbb\{P\}\_\{t\},V\_\{t\}\)solves the weak continuity equation on𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\), thenVtV\_\{t\}*generates*ℙt\\mathbb\{P\}\_\{t\}in the Lagrangian sense:
###### Theorem 1\.
Letℙt\\mathbb\{P\}\_\{t\}be an absolutely continuous curve on𝒫2\(𝒫2\(ℳ\)\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\), and a vector fieldVt∈L2\(ℙt,T𝒫2\(ℳ\)\)V\_\{t\}\\in L^\{2\}\(\\mathbb\{P\}\_\{t\},T\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\)with∫𝒫2\(ℳ\)∥Vt\(μ\)∥L2\(μ\)2dℙt\(μ\)<∞\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\lVert V\_\{t\}\(\\mu\)\\rVert^\{2\}\_\{L^\{2\}\(\\mu\)\}d\\mathbb\{P\}\_\{t\}\(\\mu\)<\\infty\.
Assume that\(ℙt,Vt\)\(\\mathbb\{P\}\_\{t\},V\_\{t\}\)solves the weak continuity equation on𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\):
∫0T∫𝒫2\(ℳ\)\(∂tφt\(μ\)\+⟨∇𝒲φt\(μ\),Vt\(μ\)⟩L2\(μ\)\)dℙt\(μ\)𝑑t=0\\int\_\{0\}^\{T\}\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\bigg\(\\partial\_\{t\}\\varphi\_\{t\}\(\\mu\)\+\\langle\\nabla\_\{\\mathcal\{W\}\}\\varphi\_\{t\}\(\\mu\),V\_\{t\}\(\\mu\)\\rangle\_\{L^\{2\}\(\\mu\)\}\\bigg\)\\,d\\mathbb\{P\}\_\{t\}\(\\mu\)dt=0\(3\)
for all smooth, compactly supported cylinder functionalsφt\\varphi\_\{t\}\. Assume moreover thatVtV\_\{t\}satisfies suitable regularity condition specified in Appendix[C](https://arxiv.org/html/2609.25659#A3)\. Then, forℙ0−a\.e\\mathbb\{P\}\_\{0\}\-a\.eμ∈𝒫2\(ℳ\)\\mu\\in\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\), there exists a unique absolutely continuous curveγμ\(t\):\[0,T\]→𝒫2\(ℳ\)\\gamma\_\{\\mu\}\(t\):\[0,T\]\\rightarrow\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)solving
γ˙μ\(t\)=Vt\(γμ\(t\)\)for a\.e\. t, withγμ\(0\)=μ\\dot\{\\gamma\}\_\{\\mu\}\(t\)=V\_\{t\}\(\\gamma\_\{\\mu\}\(t\)\)\\quad\\text\{for a\.e\. t, with \}\\gamma\_\{\\mu\}\(0\)=\\muand the mapΦt\(μ\):=γμ\(t\)\\Phi\_\{t\}\(\\mu\):=\\gamma\_\{\\mu\}\(t\)defines a flow such thatℙt=\(Φt\)\#ℙ0for allt∈\[0,T\]\\mathbb\{P\}\_\{t\}=\(\\Phi\_\{t\}\)\_\{\\\#\}\\mathbb\{P\}\_\{0\}\\;\\text\{for all \}t\\in\[0,T\]
The proof uses a superposition principle for random measures and is given in Appendix[C\.1](https://arxiv.org/html/2609.25659#A3.SS1)\.
#### 3\.1\.2Marginalization of the conditional vector field \(C2\)
LetΠ\(μ,ν\)\\Pi\(\\mu,\\nu\)be a probability measure on𝒫2\(ℳ\)⊗𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\\otimes\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\), with marginals∫μdΠ\(μ,ν\)=ℙ1\\int\_\{\\mu\}d\\Pi\(\\mu,\\nu\)=\\mathbb\{P\}\_\{1\}and∫νdΠ\(μ,ν\)=ℙ0\\int\_\{\\nu\}d\\Pi\(\\mu,\\nu\)=\\mathbb\{P\}\_\{0\}\. Given a conditional vector fieldVt\(⋅\|μ0,μ1\)V\_\{t\}\(\\cdot\|\\mu\_\{0\},\\mu\_\{1\}\), conditioned on a sample\(μ0,μ1\)∼Π\(\\mu\_\{0\},\\mu\_\{1\}\)\\sim\\Pi, we define the marginal vector fieldVtV\_\{t\}as:
Vt\(μt\):=𝔼ℙt\(μ,ν\|μt\)\[Vt\(μt\|μ0,μ1\)\]=∬𝒫2\(ℳ\)×𝒫2\(ℳ\)Vt\(μt\|μ,ν\)dℙt\(μ,ν\|μt\)V\_\{t\}\(\\mu\_\{t\}\):=\\mathbb\{E\}\_\{\\mathbb\{P\}\_\{t\}\(\\mu,\\nu\|\\mu\_\{t\}\)\}\[V\_\{t\}\(\\mu\_\{t\}\|\\mu\_\{0\},\\mu\_\{1\}\)\]=\\iint\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\\times\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}V\_\{t\}\(\\mu\_\{t\}\|\\mu,\\nu\)\\ d\\mathbb\{P\}\_\{t\}\(\\mu,\\nu\|\\mu\_\{t\}\)wheredℙt\(μ,ν\|μt\)=dℙt\(μt\|μ,ν\)dΠ\(μ,ν\)dℙt\(μt\)d\\mathbb\{P\}\_\{t\}\(\\mu,\\nu\|\\mu\_\{t\}\)=\\frac\{d\\mathbb\{P\}\_\{t\}\(\\mu\_\{t\}\|\\mu,\\nu\)d\\Pi\(\\mu,\\nu\)\}\{d\\mathbb\{P\}\_\{t\}\(\\mu\_\{t\}\)\}is the conditional probability measure\. The following result establishes the weak continuity equation for the marginal pair and, under the hypotheses of Theorem[1](https://arxiv.org/html/2609.25659#Thmtheorem1), its Lagrangian interpretation\.
###### Proposition 2\.
Assume\(ℙt\(⋅\|μ,ν\),Vt\(⋅\|μ,ν\)\)\(\\mathbb\{P\}\_\{t\}\(\\cdot\|\\mu,\\nu\),V\_\{t\}\(\\cdot\|\\mu,\\nu\)\)solves the weak continuity equation \([3](https://arxiv.org/html/2609.25659#S3.E3)\), forμ,ν\\mu,\\nuΠ−a\.s\.\\Pi\-a\.s\.\. Then\(ℙt,Vt\)\(\\mathbb\{P\}\_\{t\},V\_\{t\}\)also solves \([3](https://arxiv.org/html/2609.25659#S3.E3)\)\. Under the hypotheses of the superposition theorem \(Theorem[1](https://arxiv.org/html/2609.25659#Thmtheorem1)\) for the marginal pair,VtV\_\{t\}generates the marginal probability pathℙt\\mathbb\{P\}\_\{t\}\.
#### 3\.1\.3RWEFM Training Objective \(C3\)
The ideal, yet intractable, Flow Matching objective is
ℒFM=𝔼t,μt∼ℙt\[∥Vt\(μt\)−Vtθ\(μt\)∥L2\(μt\)2\]\.\\displaystyle\\mathcal\{L\}\_\{FM\}=\\mathbb\{E\}\_\{t,\\mu\_\{t\}\\sim\\mathbb\{P\}\_\{t\}\}\[\\lVert V\_\{t\}\(\\mu\_\{t\}\)\-V\_\{t\}^\{\\theta\}\(\\mu\_\{t\}\)\\rVert^\{2\}\_\{L^\{2\}\(\\mu\_\{t\}\)\}\]\.To bypass this, we optimize the RWEFM objective:
ℒRWEFM=𝔼t,μ,ν∼Π,μt∼ℙt\(⋅∣μ,ν\)\[∥Vt\(μt∣μ,ν\)−Vtθ\(μt\)∥L2\(μt\)2\]\\mathcal\{L\}\_\{RWEFM\}=\\mathbb\{E\}\_\{\\begin\{subarray\}\{c\}t,\\mu,\\nu\\sim\\Pi,\\\\ \\mu\_\{t\}\\sim\\mathbb\{P\}\_\{t\}\(\\cdot\\mid\\mu,\\nu\)\\end\{subarray\}\}\\bigg\[\\big\\lVert V\_\{t\}\(\\mu\_\{t\}\\mid\\mu,\\nu\)\-V\_\{t\}^\{\\theta\}\(\\mu\_\{t\}\)\\big\\rVert^\{2\}\_\{L^\{2\}\(\\mu\_\{t\}\)\}\\bigg\]\(4\)As shown in Appendix[C\.3](https://arxiv.org/html/2609.25659#A3.SS3),∇θℒFM=∇θℒRWEFM\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{FM\}=\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{RWEFM\}\. CombiningC1–C3, we have successfully lifted flow matching to𝒫2\(𝒫2\(ℳ\)\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\)\.
#### 3\.1\.4RWEFM interpolants
Different choices of conditional probability pathsℙt\(⋅\|μ0,μ1\)\\mathbb\{P\}\_\{t\}\(\\cdot\|\\mu\_\{0\},\\mu\_\{1\}\)are possible\. In this work, we choose the McCann geodesics as the canonical interpolants betweenμ\\muandν\\nu:
ℙt\(⋅\|μ,ν\)=δμtwithμt=\(ψt\)\#μandψt\(x\)=expx\(tlogx\(Tμ→ν\(x\)\)\)\.\\mathbb\{P\}\_\{t\}\(\\cdot\|\\mu,\\nu\)=\\delta\_\{\\mu\_\{t\}\}\\text\{ with \}\\mu\_\{t\}=\(\\psi\_\{t\}\)\_\{\\\#\}\\mu\\text\{ and \}\\psi\_\{t\}\(x\)=\\exp\_\{x\}\(t\\log\_\{x\}\(T^\{\\mu\\to\\nu\}\(x\)\)\)\.\(5\)For such conditional probability path, the conditional vector field that solves \([3](https://arxiv.org/html/2609.25659#S3.E3)\) is given by:
Vt\(⋅\|μ,ν\)=11−tlogxtTμ→ν\(ψt−1\(xt\)\)\.\\displaystyle V\_\{t\}\(\\cdot\|\\mu,\\nu\)=\\frac\{1\}\{1\-t\}\\log\_\{x\_\{t\}\}T^\{\\mu\\to\\nu\}\(\\psi\_\{t\}^\{\-1\}\(x\_\{t\}\)\)\.\(6\)We verify thatVt\(⋅\|μ,ν\)∈Tμt𝒫2\(ℳ\)V\_\{t\}\(\\cdot\|\\mu,\\nu\)\\in T\_\{\\mu\_\{t\}\}\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)withlogxtTμ→ν\(ψt−1\(xt\)\)=logxtTμ→ν\(xt\)=−∇φ\(xt\)\\log\_\{x\_\{t\}\}T^\{\\mu\\to\\nu\}\(\\psi\_\{t\}^\{\-1\}\(x\_\{t\}\)\)=\\log\_\{x\_\{t\}\}T^\{\\mu\\to\\nu\}\(x\_\{t\}\)=\-\\nabla\\varphi\(x\_\{t\}\)for someφ∈Cc∞\(ℳ\)\\varphi\\in C\_\{c\}^\{\\infty\}\(\\mathcal\{M\}\)\.
Figure 3:Riemannian Entropic Map\.Gaussian\-to\-checkerboard transport on𝕊2\\mathbb\{S\}^\{2\}using Euclidean \(left\) and Riemannian \(middle\) entropic maps\. The Riemannian approach respects the underlying geometry, resulting in a more accurate mapping \(see[fig\.S2](https://arxiv.org/html/2609.25659#A4.F2)\)\.
### 3\.2The Riemannian Entropic Map
A crucial step of computing the RWEFM objective is obtaining the Monge mapTμ→νT^\{\\mu\\to\\nu\}between two distributions\. To this end, we propose theRiemannian Entropic MapTε\(x\)T\_\{\\varepsilon\}\(x\), a scalable estimator forTμ→νT^\{\\mu\\to\\nu\}on Riemannian manifolds based on finite samples fromμ\\muandν\\nuand fast entropic OT solvers\.
Like the Euclidean entropic map\([Pooladian and Niles\-Weed, 2021](https://arxiv.org/html/2609.25659#bib.bib21)\), our estimator leverages the entropic optimal couplingπε\\pi\_\{\\varepsilon\}obtained from the Sinkhorn algorithm, computed using the Riemannian distance cost matrixCij=12dg2\(xi,yj\)C\_\{ij\}=\\frac\{1\}\{2\}d\_\{g\}^\{2\}\(x\_\{i\},y\_\{j\}\)between samples\{xi\}i=1N∼μ\\\{x\_\{i\}\\\}\_\{i=1\}^\{N\}\\sim\\muand\{yj\}j=1N∼ν\\\{y\_\{j\}\\\}\_\{j=1\}^\{N\}\\sim\\nu\.
###### Definition 3\(Riemannian Entropic Map\)\.
Letπε\\pi\_\{\\varepsilon\}be the optimal entropic coupling\. We define the Riemannian entropic mapTε:ℳ→ℳT\_\{\\varepsilon\}:\\mathcal\{M\}\\to\\mathcal\{M\}as:
Tε\(x\):=expx\(∫ℳlogx\(y\)dπε\(y\|x\)\)\.T\_\{\\varepsilon\}\(x\):=\\exp\_\{x\}\\left\(\\int\_\{\\mathcal\{M\}\}\\log\_\{x\}\(y\)d\\pi\_\{\\varepsilon\}\(y\|x\)\\right\)\.\(7\)
###### Theorem 4\.
Letfεf\_\{\\varepsilon\}be the optimal dual potential of the entropic optimal transport problem \(Eq \([2](https://arxiv.org/html/2609.25659#S2.E2)\)\)\. Under primal feasibility conditions, the Riemannian entropic map arises as the gradient of the dual potential:Tε\(x\)=expx\(−∇fε\(x\)\)T\_\{\\varepsilon\}\(x\)=\\exp\_\{x\}\\left\(\-\\nabla f\_\{\\varepsilon\}\(x\)\\right\)\. The proof is in Appendix[A](https://arxiv.org/html/2609.25659#A1); whenℳ=ℝd\\mathcal\{M\}=\\mathbb\{R\}^\{d\}, Eq \([7](https://arxiv.org/html/2609.25659#S3.E7)\) recovers the Euclidean mapTε=𝔼πε\[Y\|X=x\]T\_\{\\varepsilon\}=\\mathbb\{E\}\_\{\\pi\_\{\\varepsilon\}\}\[Y\|X=x\]\.
Computationally, this definition implies a straightforwardlift\-average\-retractprocedure for empirical measures\. To evaluate the map at a source pointxx: \(i\)liftby computing the tangent vectorsvj=logx\(yj\)v\_\{j\}=\\log\_\{x\}\(y\_\{j\}\)for all target pointsyjy\_\{j\}, mapping the geometry fromℳ\\mathcal\{M\}to the vector spaceTxℳT\_\{x\}\\mathcal\{M\}; \(ii\)averageto calculate the weighted meanv¯=∑jπε\(yj\|x\)vj\\bar\{v\}=\\sum\_\{j\}\\pi\_\{\\varepsilon\}\(y\_\{j\}\|x\)v\_\{j\}; \(iii\)retractby applying the exponential map to project back to the manifold:Tε\(x\)=expx\(v¯\)T\_\{\\varepsilon\}\(x\)=\\exp\_\{x\}\(\\bar\{v\}\)\. We highlight that this procedure only requires access to the Riemannian exponential and logarithm maps, making it applicable to a wide range of manifolds\. When these maps are not available, our estimator can still be efficiently computed with only access to the distance functiondg\(x,y\)d\_\{g\}\(x,y\), as shown in Appendix[A\.6](https://arxiv.org/html/2609.25659#A1.SS6); we exploit this to apply RWEFM on a triangulated mesh with no closed\-form geometry \(Section[4\.2\.3](https://arxiv.org/html/2609.25659#S4.SS2.SSS3)\)\. Furthermore, the estimator has the same complexity as the Euclidean entropic map, making it equally scalable to large datasets\.
To justify usingTεT\_\{\\varepsilon\}as a reliable proxy for the true Monge mapT0T\_\{0\}in our learning objective, we establish the following statistical bound\. This result ensures that the bias introduced by entropic regularization and finite sampling is controlled\.
###### Theorem 5\(Statistical performance\)\.
LetT^ε,\(n,n\)\\widehat\{T\}\_\{\\varepsilon,\(n,n\)\}be the canonical out\-of\-sample entropic map computed fromnniid observations from each ofμ\\muandν\\nu, with the two samples independent\. Under Assumptions[9](https://arxiv.org/html/2609.25659#Thmtheorem9),[10](https://arxiv.org/html/2609.25659#Thmtheorem10), and[11](https://arxiv.org/html/2609.25659#Thmtheorem11), there exist constantsC<∞C<\\inftyandε0\>0\\varepsilon\_\{0\}\>0such that, forn≥2n\\geq 2and0<ε≤ε00<\\varepsilon\\leq\\varepsilon\_\{0\},
𝔼∫Ωdg2\(T^ε,\(n,n\)\(x\),T0\(x\)\)dμ\(x\)≤C\[εlog\(eε\)\+ε−sdlog\(n\+1\)n\],\\displaystyle\\mathbb\{E\}\\int\_\{\\Omega\}d\_\{g\}^\{2\}\\\!\\left\(\\widehat\{T\}\_\{\\varepsilon,\(n,n\)\}\(x\),T\_\{0\}\(x\)\\right\)\\,d\\mu\(x\)\\leq C\\left\[\\varepsilon\\log\\\!\\left\(\\frac\{e\}\{\\varepsilon\}\\right\)\+\\varepsilon^\{\-s\_\{d\}\}\\frac\{\\log\(n\+1\)\}\{\\sqrt\{n\}\}\\right\],wheresd=max\{1,d/2\}s\_\{d\}=\\max\\\{1,d/2\\\}andd=dimℳd=\\dim\\mathcal\{M\}\. The constants depend on the geometry and regularity assumptions\. The proof is given in Appendix[A\.7](https://arxiv.org/html/2609.25659#A1.SS7)\.
### 3\.3Algorithm and Implementation
Algorithm 1RWEFM Training StepInput:
ℙ0,ℙ1\\mathbb\{P\}\_\{0\},\\mathbb\{P\}\_\{1\}
Sample
μ∼ℙ0,ν∼ℙ1\\mu\\sim\\mathbb\{P\}\_\{0\},\\nu\\sim\\mathbb\{P\}\_\{1\}Sample
\{xi\}i=1N∼μ,\{yj\}j=1M∼ν\\\{x\_\{i\}\\\}\_\{i=1\}^\{N\}\\sim\\mu,\\\{y\_\{j\}\\\}\_\{j=1\}^\{M\}\\sim\\nu
Compute
π\\pivia Sinkhorn
\(C,ε\)\(C,\\varepsilon\), Sample
t∼𝒰\[0,1\]t\\sim\\mathcal\{U\}\[0,1\]
foreach
xix\_\{i\}do
wij=πij/∑kπikw\_\{ij\}=\\pi\_\{ij\}/\\sum\_\{k\}\\pi\_\{ik\}
Tε\(xi\)=expxi\(∑jwijlogxi\(yj\)\)T\_\{\\varepsilon\}\(x\_\{i\}\)=\\exp\_\{x\_\{i\}\}\\\!\\left\(\\sum\_\{j\}w\_\{ij\}\\log\_\{x\_\{i\}\}\(y\_\{j\}\)\\right\)
zi=expxi\(tlogxi\(Tε\(xi\)\)\)z\_\{i\}=\\exp\_\{x\_\{i\}\}\(t\\log\_\{x\_\{i\}\}\(T\_\{\\varepsilon\}\(x\_\{i\}\)\)\)vi=logzi\(Tε\(xi\)\)/\(1−t\)v\_\{i\}=\\log\_\{z\_\{i\}\}\(T\_\{\\varepsilon\}\(x\_\{i\}\)\)/\(1\-t\)
v^i=vθ\(zi,t,\{zi\}i=1N\)\\hat\{v\}\_\{i\}=v\_\{\\theta\}\(z\_\{i\},t,\\\{z\_\{i\}\\\}\_\{i=1\}^\{N\}\)
ℒ=1N∑i‖v^i−vi‖g\(zi\)2\\mathcal\{L\}=\\frac\{1\}\{N\}\\sum\_\{i\}\\\|\\hat\{v\}\_\{i\}\-v\_\{i\}\\\|\_\{g\(z\_\{i\}\)\}^\{2\}
Update
θ←θ−η∇θℒ\\theta\\leftarrow\\theta\-\\eta\\nabla\_\{\\theta\}\\mathcal\{L\}
Algorithm 2RWEFM GenerationInput:
μ∼ℙ0\\mu\\sim\\mathbb\{P\}\_\{0\}, discretization step
dtdt
Sample
X0=\{xi\}i=1N∼μX\_\{0\}=\\\{x\_\{i\}\\\}\_\{i=1\}^\{N\}\\sim\\mu
for
t←0t\\leftarrow 0to
1−dt1\-dtstep
dtdtdo
Compute
vi=vθ\(xi,t,\{xi\}i=1N\)v\_\{i\}=v\_\{\\theta\}\(x\_\{i\},t,\\\{x\_\{i\}\\\}\_\{i=1\}^\{N\}\)
foreach
xix\_\{i\}do
xi←expxi\(vi⋅dt\)x\_\{i\}\\leftarrow\\exp\_\{x\_\{i\}\}\(v\_\{i\}\\cdot dt\)
Return:
X1=\{xi\}i=1NX\_\{1\}=\\\{x\_\{i\}\\\}\_\{i=1\}^\{N\}
Here we detail how RWEFM is implemented in practice\. The training algorithm is presented in Algorithm[1](https://arxiv.org/html/2609.25659#alg1)\. At each training step, we sample distributionsμ\\muandν\\nufrom the meta\-distributionsℙ0\\mathbb\{P\}\_\{0\}andℙ1\\mathbb\{P\}\_\{1\}, represented as point cloudsX0=\{xi\}i=1N,xi∼μX\_\{0\}=\\\{x\_\{i\}\\\}\_\{i=1\}^\{N\},x\_\{i\}\\sim\\muandX1=\{yj\}j=1M,yj∼νX\_\{1\}=\\\{y\_\{j\}\\\}\_\{j=1\}^\{M\},y\_\{j\}\\sim\\nu\. We compute the Riemannian entropic map \(Eq \([7](https://arxiv.org/html/2609.25659#S3.E7)\)\) to generate an interpolating point cloudXt=\{zi\}i=1NX\_\{t\}=\\\{z\_\{i\}\\\}\_\{i=1\}^\{N\}withzi=expxi\(tlogxi\(Tε\(xi\)\)\)z\_\{i\}=\\exp\_\{x\_\{i\}\}\(t\\log\_\{x\_\{i\}\}\(T\_\{\\varepsilon\}\(x\_\{i\}\)\)\)\(Eq \([5](https://arxiv.org/html/2609.25659#S3.E5)\)\) and target velocity vectorsVt=\{11−tlogzi\(Tε\(xi\)\)\}i=1NV\_\{t\}=\\\{\\frac\{1\}\{1\-t\}\\log\_\{z\_\{i\}\}\(T\_\{\\varepsilon\}\(x\_\{i\}\)\)\\\}\_\{i=1\}^\{N\}\(Eq \([6](https://arxiv.org/html/2609.25659#S3.E6)\)\)\. Since the data is set\-structured, we parameterize the neural networkvθ\(Xt,t\):ℳN→TXtℳNv\_\{\\theta\}\(X\_\{t\},t\):\\mathcal\{M\}^\{N\}\\to T\_\{X\_\{t\}\}\\mathcal\{M\}^\{N\}using self\-attention architectures, mapping point sets to tangent vectors\. The loss is defined as the mean squared geodesic norm between the network predictionv^i\\hat\{v\}\_\{i\}and the targetviv\_\{i\}\(Eq \([4](https://arxiv.org/html/2609.25659#S3.E4)\)\)\. For generation \(Algorithm[2](https://arxiv.org/html/2609.25659#alg2)\), we start fromX0∼μ∼ℙ0X\_\{0\}\\sim\\mu\\sim\\mathbb\{P\}\_\{0\}and integratevθv\_\{\\theta\}over time to obtain the final sampleX1X\_\{1\}\.
Appendix[D](https://arxiv.org/html/2609.25659#A4)details hyperparameters and architecture\. RWEFM typically trains in a few hours on a single GPU\. For large datasets like the single\-cell atlas, we trained with 1024 cells sampled from each distribution, still achieving high\-quality results in less than 24 hours of training\. The size of the point clouds and the entropic regularizationε\\varepsilonare the main hyper\-parameters impacting computational complexity, with per\-dataset training times for all methods reported in[tableS1](https://arxiv.org/html/2609.25659#A4.T1)\. Training time can be meaningfully reduced by increasingε\\varepsilonor decreasing the number of sampled particles, at a marginal cost to generation quality, as studied in detail in Section[D\.9](https://arxiv.org/html/2609.25659#A4.SS9)\.
## 4Results
### 4\.1The Riemannian Entropic Map is a faithful estimator of the Monge map onℳ\\mathcal\{M\}
We begin by qualitatively demonstrating the effectiveness of the Riemannian entropic map\.[Figure3](https://arxiv.org/html/2609.25659#S3.F3)illustrates the transport error on the sphere𝕊2\\mathbb\{S\}^\{2\}, comparing our Riemannian estimator against a Euclidean baseline\. By strictly adhering to the underlying geometry, the Riemannian entropic map yields a significantly more accurate coupling, whereas the Euclidean approach incurs high distortion as it ignores the manifold curvature\. We provide quantitative evaluation of its performance on constructed examples on the sphere𝕊2\\mathbb\{S\}^\{2\}and hyperbolic spaceℍ2\\mathbb\{H\}^\{2\}in Appendix[D\.8](https://arxiv.org/html/2609.25659#A4.SS8)\. The ground truth OT map is generated by taking the exponential map of the gradient of a distance function between a pointx∈ℳx\\in\\mathcal\{M\}and a fixed attractorx¯∈ℳ\\bar\{x\}\\in\\mathcal\{M\}\. Our results confirm that our estimator is as accurate as the unregularized OT map \(Eq \([1](https://arxiv.org/html/2609.25659#S2.E1)\)\), albeit much faster, and outperforms the Euclidean entropic map, which neglects the underlying geometry \([fig\.S2](https://arxiv.org/html/2609.25659#A4.F2)\)\.
Figure 4:Generative Modeling of Single\-Cell Lung Samples\.UMAP visualization of real and generated scRNA\-seq samples from healthy and cancer lung tissues\. The generated samples demonstrate high congruence with the real distributions\. Furthermore, operating in the foundation model’s latent space enables direct analysis of generated data, such as accurate cell typing, highlighting the utility of generating in the \(non\-Euclidean\) latent spaces of foundation models\.
### 4\.2Benchmarking RWEFM on synthetic and real\-world scientific datasets
#### 4\.2\.1Baselines and Evaluation Metrics
We benchmark RWEFM against several baselines\.FM\([Lipman et al\., 2022](https://arxiv.org/html/2609.25659#bib.bib22)\)andRFM\([Chen and Lipman, 2023](https://arxiv.org/html/2609.25659#bib.bib47)\)operate on individual data points\.WFM\([Haviv et al\., 2024](https://arxiv.org/html/2609.25659#bib.bib26);[Atanackovic et al\., 2024](https://arxiv.org/html/2609.25659#bib.bib34)\)learns flows on the Wasserstein space but is restricted to Euclidean domains\. We also introduceSetFM/SetRFM, which apply the FM objective to point clouds without OT couplings\. For point cloud data, we includePVD\([Zhou et al\., 2021](https://arxiv.org/html/2609.25659#bib.bib42)\)andPSF\([Wu et al\., 2023](https://arxiv.org/html/2609.25659#bib.bib2)\), trained in Euclidean space and projected onto the manifolds\. All set\-level methods share the same self\-attention architecture, FM/RFM use an MLP and we project generated samples from non\-Riemannian methods onto the manifold for fair comparison\. We evaluate model performance using classwise 1\-NN deviation \(1\-NN\-D\) and MMD, with Chamfer Distance \(CD\) and Earth Mover’s Distance \(EMD\) between point clouds \(Appendix[D\.3](https://arxiv.org/html/2609.25659#A4.SS3)\)\. For equally sized sets of real and generated point clouds, letareala\_\{\\mathrm\{real\}\}andagena\_\{\\mathrm\{gen\}\}be their respective classwise nearest\-neighbor classification accuracies\. We report1\-NN\-D:=12\(\|areal−12\|\+\|agen−12\|\)\\text\{1\-NN\-D\}:=\\frac\{1\}\{2\}\\left\(\\left\|a\_\{\\mathrm\{real\}\}\-\\frac\{1\}\{2\}\\right\|\+\\left\|a\_\{\\mathrm\{gen\}\}\-\\frac\{1\}\{2\}\\right\|\\right\)This score lies in\[0,0\.5\]\[0,0\.5\]and is minimized when both classwise accuracies equal0\.50\.5\. Taking the absolute deviations before averaging prevents opposite classwise deviations from canceling\.
#### 4\.2\.2MNIST, EMNIST, and KMNIST on the Sphere, Hyperbolic Space, and Torus
We evaluate RWEFM on MNIST\([LeCun et al\., 1998](https://arxiv.org/html/2609.25659#bib.bib37)\), EMNIST\([Cohen et al\., 2017](https://arxiv.org/html/2609.25659#bib.bib36)\), and KMNIST\([ROIS\-CODH, 2018](https://arxiv.org/html/2609.25659#bib.bib1)\)datasets mapped onto Riemannian manifolds: MNIST digits on the sphere𝕊2\\mathbb\{S\}^\{2\}, EMNIST letters on hyperbolic spaceℍ2\\mathbb\{H\}^\{2\}, and KMNIST characters on the torus𝕋2\\mathbb\{T\}^\{2\}\. RWEFM learns high\-quality flows and outperforms baselines in all settings \(Figure[2](https://arxiv.org/html/2609.25659#S2.F2), Tables[2](https://arxiv.org/html/2609.25659#S3.T2),[S3](https://arxiv.org/html/2609.25659#A4.T3)\)\. Notably, on the torus𝕋2\\mathbb\{T\}^\{2\}, WFM and RWEFM achieve comparable performance\. This is expected as𝕋2\\mathbb\{T\}^\{2\}is a flat manifold with zero curvature and the geometric advantage of RWEFM’s Riemannian OT over WFM’s Euclidean OT diminishes accordingly\. Unsurprisingly, FM/RFM perform worst as they are unable to capture the dependencies between individual points required to generate meaningful point clouds\.
Figure 5:Distributions generated on the Stanford bunny\.RWEFM\-generated empirical distributions on a general triangulated mesh with no closed\-form geometry, showing two samples per digit class \(0, 2, 9\)\. Generated points lie on the curved surface and recover the digit morphology, illustrating RWEFM on arbitrary geometries \([table3](https://arxiv.org/html/2609.25659#S4.T3)\)\.
#### 4\.2\.3Beyond closed\-form manifolds: distributions on a general triangulated mesh
The manifolds above admit closed\-form geodesics, exponential and logarithm maps\. Many scientific geometries—triangulated surfaces, learned metrics, implicit surfaces—do not\.
RWEFM extends naturally to this setting: both the McCann interpolant and the Riemannian entropic map \(Section[3\.2](https://arxiv.org/html/2609.25659#S3.SS2)\) can be evaluated from only a geodesic distancedgd\_\{g\}and a projection operator onto the manifold, without analyticexp\\exp/log\\logmaps \(Appendix[A\.6](https://arxiv.org/html/2609.25659#A1.SS6)\)\.
As a proof of concept, we generate distributions on the surface of the Stanford bunny, a triangulated mesh with no analytic geometry, using MNIST digits as the empirical distributions\. We endow the mesh with the spectral \(biharmonic\) premetric of[Chen and Lipman \(2023\)](https://arxiv.org/html/2609.25659#bib.bib47), computed from its smallest Laplace–Beltrami eigenpairs, and lay each digit onto a fixed tangent chart on the bunny’s flank \(Appendix[D\.5](https://arxiv.org/html/2609.25659#A4.SS5)\)\.[Figure5](https://arxiv.org/html/2609.25659#S4.F5)shows point clouds generated by RWEFM for three digit classes and the samples lie on the curved surface and recover the digit morphology\.
Table 3:RWEFM on a general triangulated mesh \(Stanford bunny\)\.Hereℳmesh\\mathcal\{M\}\_\{\\mathrm\{mesh\}\}denotes the Stanford bunny surface represented by a triangulated mesh\. Classwise 1\-NN deviation \(1\-NN\-D; lower is better\) under the mesh spectral metric for MNIST digits generated on the bunny \(best per column bold; 3 seeds×\\times5 samplings\)\. RWEFM and WFM use the sampled map \(Appendix[D\.5](https://arxiv.org/html/2609.25659#A4.SS5)\); MMD in[tableS2](https://arxiv.org/html/2609.25659#A4.T2)\.Quantitatively \([table3](https://arxiv.org/html/2609.25659#S4.T3)\), the mesh\-aware methods \(RWEFM, SetRFM\) sharply outperform their Euclidean counterparts \(WFM, SetFM\), which ignore the surface and generally attain or approach the maximum classwise 1\-NN deviation \(D1NN=0\.5D\_\{\\mathrm\{1NN\}\}=0\.5\), indicating large departures from 50% classwise accuracy\. Respecting the intrinsic geometry therefore remains decisive even when the manifold has no closed\-form description, demonstrating that RWEFM applies to arbitrary geometries\.
#### 4\.2\.4Single\-cell RNA\-seq samples on Riemannian spaces
In single\-cell genomics, early generative models focused on synthesizing individual cells\. A growing body of work now targets a fundamentally harder task: generating*whole samples*—each sample being an empirical distribution of thousands of cells representing a patient or tissue\. This sample\-level perspective is essential for capturing inter\-sample heterogeneity in disease modeling and perturbation studies\([Boyeau et al\., 2025](https://arxiv.org/html/2609.25659#bib.bib50);[Boiarsky et al\., 2025](https://arxiv.org/html/2609.25659#bib.bib4);[Haviv et al\., 2024](https://arxiv.org/html/2609.25659#bib.bib26);[Atanackovic et al\., 2024](https://arxiv.org/html/2609.25659#bib.bib34);[Klein et al\., 2025](https://arxiv.org/html/2609.25659#bib.bib11)\)\. Concretely, each training example is a point cloudμ=\{xi\}i=1N\\mu=\\\{x\_\{i\}\\\}\_\{i=1\}^\{N\}of cell embeddings for one patient sample, and the goal is to generate new point clouds that faithfully reproduce the distribution of real samples\.
Table 4:Generation of cells and torsion angles on manifolds\.\(Left\) whole single\-cell RNAseq samples on𝕊128−1\\mathbb\{S\}^\{128\-1\}\(SCimilarity embeddings\); \(right\) per\-protein torsion\-angle distributions on𝕋2\\mathbb\{T\}^\{2\}\(ESM\-conditioned\)\. Extended metrics in[tableS4](https://arxiv.org/html/2609.25659#A4.T4); times in[tableS1](https://arxiv.org/html/2609.25659#A4.T1)\.A further challenge is that cell embeddings from foundation models such as SCimilarity\([Heimberg et al\., 2025](https://arxiv.org/html/2609.25659#bib.bib39)\)naturally live on non\-Euclidean manifolds such as hyperspheres \(𝕊d−1\\mathbb\{S\}^\{d\-1\}\) or hyperbolic spaces\([Ding and Regev, 2021](https://arxiv.org/html/2609.25659#bib.bib38)\), so naive Euclidean generation distorts the learned geometry\. RWEFM directly addresses both challenges: it generates*distributions*\(whole samples\) on the*manifold*by learning a flow on the Wasserstein space𝒫2\(𝕊128−1\)\\mathcal\{P\}\_\{2\}\(\\mathbb\{S\}^\{128\-1\}\)\. We evaluate on generating whole single\-cell blood samples in SCimilarity’s hyperspherical embedding space and find that RWEFM substantially outperforms Euclidean and non\-distributional baselines \(Table[4](https://arxiv.org/html/2609.25659#S4.T4), Figures[4](https://arxiv.org/html/2609.25659#S4.F4),[S1](https://arxiv.org/html/2609.25659#A4.F1)\)\. The generated samples accurately recapitulate real cell\-type compositions and enable direct cell\-typing analysis, demonstrating the utility of respecting the geometry of foundation model latent spaces\.
#### 4\.2\.5Protein torsion angle distributions on the torus
Figure 6:Torsion angle generation on the torus𝕋2\\mathbb\{T\}^\{2\}\.True and generated distributions of protein backbone torsion angles \(ϕ\\phi,ψ\\psi\) for held\-out mdCATH proteins, conditioned on ESM embeddings of their sequences \(first2020amino acids shown in subplot titles\)\. RWEFM accurately captures the multimodal structure of conformational distributions\.Protein backbone conformations can be characterized by dihedral anglesϕ\\phiandψ\\psi\. The joint angular distribution is depicted using the Ramachandran plot, which reveals energetically favorable conformations and structural motifs such asα\\alpha\-helices andβ\\beta\-sheets\. Each protein is thus characterized by a*distribution*of torsion angles\. Since both angles are periodic, this distribution lives on the flat torus𝕋2\\mathbb\{T\}^\{2\}, a common setting for Riemannian generative models\([Chen and Lipman, 2023](https://arxiv.org/html/2609.25659#bib.bib47);[Davis et al\., 2025](https://arxiv.org/html/2609.25659#bib.bib48)\)\.
We apply RWEFM to generate per\-protein torsion angle distributions from mdCATH\([Mirarchi et al\., 2024](https://arxiv.org/html/2609.25659#bib.bib7)\)\. For each protein, we aggregate the\(ϕ,ψ\)\(\\phi,\\psi\)angles across all residues during MD simulation, producing an empirical distribution over𝕋2\\mathbb\{T\}^\{2\}\. Unlike prior work that generates individual angle pairs, we generate entire distributions conditioned on ESM embeddings\([Hayes et al\., 2025](https://arxiv.org/html/2609.25659#bib.bib49)\)of the protein sequence\. On unseen test\-set proteins, we find that RWEFM accurately captures the complex, multimodal structure of torsion angle distributions \(Figure[6](https://arxiv.org/html/2609.25659#S4.F6)\)\. Quantitatively, RWEFM outperforms Euclidean and non\-distributional baselines \(Table[4](https://arxiv.org/html/2609.25659#S4.T4)\), demonstrating the benefits of integrating Riemannian geometry with flow matching on𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)for modeling protein conformational landscapes\.
## 5Discussion
We have introduced Riemannian Wasserstein Entropic Flow Matching \(RWEFM\), a novel framework for generative modeling of distributions on Riemannian manifolds\. We showed that our framework is theoretically sound, computationally tractable, and empirically outperforms previous approaches that either fail to capture the distributional or geometric aspects of the data\. Our proposed Riemannian entropic map is an essential part of our approach, providing highly scalable and accurate estimation of the OT map\. Our empirical benchmark spans multiple scientific applications, including single\-cell genomics and protein conformation, highlighting the diversity and abundance of scientific problems that could benefit from explicitly incorporating the geometry and distributional nature of the data\.
## References
- Abatzoglouet al\.\(2018\)J\. T\. Abatzoglou, S\. Z\. Dobrowski, S\. A\. Parks, and K\. C\. HegewischTerraClimate, a high\-resolution global dataset of monthly climate and climatic water balance from 1958–2015\.Scientific data5\(1\),pp\. 170191\.Cited by:[§1](https://arxiv.org/html/2609.25659#S1.p1.1)\.
- Ambrosioet al\.\(2021\)L\. Ambrosio, E\. Brué, and D\. SemolaLectures on optimal transport\.Vol\.130,Springer\.Cited by:[§2\.1](https://arxiv.org/html/2609.25659#S2.SS1.p2.1)\.
- Atanackovicet al\.\(2024\)L\. Atanackovic, X\. Zhang, B\. Amos, M\. Blanchette, L\. J\. Lee, Y\. Bengio, A\. Tong, and K\. NeklyudovMeta flow matching: integrating vector fields on the wasserstein manifold\.arXiv preprint arXiv:2408\.14608\.Cited by:[§1](https://arxiv.org/html/2609.25659#S1.p2.1),[§2\.3](https://arxiv.org/html/2609.25659#S2.SS3.SSS0.Px2.p1.1),[§4\.2\.1](https://arxiv.org/html/2609.25659#S4.SS2.SSS1.p1.1),[§4\.2\.4](https://arxiv.org/html/2609.25659#S4.SS2.SSS4.p1.1)\.
- Axelrod and Gomez\-Bombarelli \(2022\)S\. Axelrod and R\. Gomez\-BombarelliGEOM, energy\-annotated molecular conformations for property prediction and molecular generation\.Scientific Data9\(1\),pp\. 185\.Cited by:[§1](https://arxiv.org/html/2609.25659#S1.p1.1)\.
- Boiarskyet al\.\(2025\)R\. Boiarsky, J\. Wenckstern, N\. J\. Haradhvala, G\. Getz, and D\. SontagA diffusion\-based autoencoder for learning patient\-level representations from single\-cell data\.bioRxiv\.Cited by:[§4\.2\.4](https://arxiv.org/html/2609.25659#S4.SS2.SSS4.p1.1)\.
- Bonetet al\.\(2025\)C\. Bonet, C\. Vauthier, and A\. KorbaFlowing datasets with wasserstein over wasserstein gradient flows\.arXiv preprint arXiv:2506\.07534\.Cited by:[§D\.1](https://arxiv.org/html/2609.25659#A4.SS1.SSS0.Px3.p1.1),[§2\.3](https://arxiv.org/html/2609.25659#S2.SS3.SSS0.Px1.p1.1)\.
- Boseet al\.\(2023\)A\. J\. Bose, T\. Akhound\-Sadegh, G\. Huguet, K\. Fatras, J\. Rector\-Brooks, C\. Liu, A\. C\. Nica, M\. Korablyov, M\. Bronstein, and A\. TongSe \(3\)\-stochastic flow matching for protein backbone generation\.arXiv preprint arXiv:2310\.02391\.Cited by:[§2\.3](https://arxiv.org/html/2609.25659#S2.SS3.SSS0.Px1.p1.1)\.
- Boyeauet al\.\(2025\)P\. Boyeau, J\. Hong, A\. Gayoso, M\. Kim, J\. L\. McFaline\-Figueroa, M\. I\. Jordan, E\. Azizi, C\. Ergen, and N\. YosefDeep generative modeling of sample\-level heterogeneity in single\-cell genomics\.Nature Methods22\(11\),pp\. 2264–2274\.Cited by:[§4\.2\.4](https://arxiv.org/html/2609.25659#S4.SS2.SSS4.p1.1)\.
- Candeset al\.\(2018\)E\. Candes, Y\. Fan, L\. Janson, and J\. LvPanning for gold:‘model\-x’knockoffs for high dimensional controlled variable selection\.Journal of the Royal Statistical Society Series B: Statistical Methodology80\(3\),pp\. 551–577\.Cited by:[§1](https://arxiv.org/html/2609.25659#S1.p2.1)\.
- Chen and Lipman \(2023\)R\. T\. Chen and Y\. LipmanFlow matching on general geometries\.arXiv preprint arXiv:2302\.03660\.Cited by:[§D\.5](https://arxiv.org/html/2609.25659#A4.SS5.p1.1),[§4\.2\.1](https://arxiv.org/html/2609.25659#S4.SS2.SSS1.p1.1),[§4\.2\.3](https://arxiv.org/html/2609.25659#S4.SS2.SSS3.p3.1),[§4\.2\.5](https://arxiv.org/html/2609.25659#S4.SS2.SSS5.p1.1)\.
- Chenget al\.\(2024\)C\. Cheng, J\. Li, J\. Peng, and G\. LiuCategorical flow matching on statistical manifolds\.arXiv preprint arXiv:2405\.16441\.Cited by:[§2\.3](https://arxiv.org/html/2609.25659#S2.SS3.SSS0.Px2.p1.1)\.
- Cohenet al\.\(2017\)G\. Cohen, S\. Afshar, J\. Tapson, and A\. Van SchaikEMNIST: extending mnist to handwritten letters\.In2017 international joint conference on neural networks \(IJCNN\),pp\. 2921–2926\.Cited by:[§D\.4](https://arxiv.org/html/2609.25659#A4.SS4.p1.1),[§4\.2\.2](https://arxiv.org/html/2609.25659#S4.SS2.SSS2.p1.1)\.
- Cuturiet al\.\(2022\)M\. Cuturi, L\. Meng\-Papaxanthos, Y\. Tian, C\. Bunne, G\. Davis, and O\. TeboulOptimal transport tools \(ott\): a jax toolbox for all things wasserstein\.arXiv preprint arXiv:2201\.12324\.Cited by:[§D\.1](https://arxiv.org/html/2609.25659#A4.SS1.SSS0.Px3.p3.1)\.
- Cuturi \(2013\)M\. CuturiSinkhorn distances: lightspeed computation of optimal transport\.Advances in neural information processing systems26\.Cited by:[§2\.1\.1](https://arxiv.org/html/2609.25659#S2.SS1.SSS1.p1.1)\.
- CZI Cell Science Programet al\.\(2025\)CZI Cell Science Program, S\. Abdulla, B\. Aevermann, P\. Assis, S\. Badajoz, S\. M\. Bell, E\. Bezzi, B\. Cakir, J\. Chaffer, S\. Chambers,et al\.CZ cellxgene discover: a single\-cell data platform for scalable exploration, analysis and modeling of aggregated data\.Nucleic acids research53\(D1\),pp\. D886–D900\.Cited by:[§1](https://arxiv.org/html/2609.25659#S1.p1.1)\.
- Daviset al\.\(2025\)O\. Davis, M\. S\. Albergo, N\. M\. Boffi, M\. M\. Bronstein, and A\. J\. BoseGeneralised flow maps for few\-step generative modelling on riemannian manifolds\.arXiv preprint arXiv:2510\.21608\.Cited by:[§4\.2\.5](https://arxiv.org/html/2609.25659#S4.SS2.SSS5.p1.1)\.
- Daviset al\.\(2024\)O\. Davis, S\. Kessler, M\. Petrache, I\. I\. Ceylan, M\. Bronstein, and A\. J\. BoseFisher flow matching for generative modeling over discrete data\.Advances in Neural Information Processing Systems37,pp\. 139054–139084\.Cited by:[§2\.3](https://arxiv.org/html/2609.25659#S2.SS3.SSS0.Px2.p1.1)\.
- De Bortoliet al\.\(2022\)V\. De Bortoli, E\. Mathieu, M\. Hutchinson, J\. Thornton, Y\. W\. Teh, and A\. DoucetRiemannian score\-based generative modelling\.Advances in neural information processing systems35,pp\. 2406–2422\.Cited by:[§2\.3](https://arxiv.org/html/2609.25659#S2.SS3.SSS0.Px1.p1.1)\.
- Ding and Regev \(2021\)J\. Ding and A\. RegevDeep generative model embedding of single\-cell rna\-seq profiles on hyperspheres and hyperbolic spaces\.Nature communications12\(1\),pp\. 2554\.Cited by:[§4\.2\.4](https://arxiv.org/html/2609.25659#S4.SS2.SSS4.p2.1)\.
- Emami and Pass \(2025\)P\. Emami and B\. PassOptimal transport with optimal transport cost: the monge–kantorovich problem on wasserstein spaces\.Calculus of Variations and Partial Differential Equations64\(2\),pp\. 43\.Cited by:[§D\.1](https://arxiv.org/html/2609.25659#A4.SS1.SSS0.Px3.p1.1),[§2\.3](https://arxiv.org/html/2609.25659#S2.SS3.SSS0.Px1.p1.1)\.
- Frostiget al\.\(2019\)R\. Frostig, M\. J\. Johnson, and C\. LearyCompiling machine learning programs via high\-level tracing\.InSysML conference 2018,Cited by:[§D\.1](https://arxiv.org/html/2609.25659#A4.SS1.SSS0.Px3.p3.1)\.
- Gruveret al\.\(2023\)N\. Gruver, S\. Stanton, N\. Frey, T\. G\. Rudner, I\. Hotzel, J\. Lafrance\-Vanasse, A\. Rajpal, K\. Cho, and A\. G\. WilsonProtein design with guided discrete diffusion\.Advances in neural information processing systems36,pp\. 12489–12517\.Cited by:[§1](https://arxiv.org/html/2609.25659#S1.p2.1)\.
- Havivet al\.\(2024\)D\. Haviv, A\. Pooladian, D\. Pe’er, and B\. AmosWasserstein flow matching: generative modeling over families of distributions\.arXiv preprint arXiv:2411\.00698\.Cited by:[§D\.1](https://arxiv.org/html/2609.25659#A4.SS1.SSS0.Px4.p2.1),[§1](https://arxiv.org/html/2609.25659#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.25659#S2.SS2.SSS0.Px1.p2.1),[§2\.3](https://arxiv.org/html/2609.25659#S2.SS3.SSS0.Px2.p1.1),[§4\.2\.1](https://arxiv.org/html/2609.25659#S4.SS2.SSS1.p1.1),[§4\.2\.4](https://arxiv.org/html/2609.25659#S4.SS2.SSS4.p1.1)\.
- Hayeset al\.\(2025\)T\. Hayes, R\. Rao, H\. Akin, N\. J\. Sofroniew, D\. Oktay, Z\. Lin, R\. Verkuil, V\. Q\. Tran, J\. Deaton, M\. Wiggert,et al\.Simulating 500 million years of evolution with a language model\.Science387\(6736\),pp\. 850–858\.Cited by:[§4\.2\.5](https://arxiv.org/html/2609.25659#S4.SS2.SSS5.p2.1)\.
- Heimberget al\.\(2025\)G\. Heimberg, T\. Kuo, D\. J\. DePianto, O\. Salem, T\. Heigl, N\. Diamant, G\. Scalia, T\. Biancalani, S\. J\. Turley, J\. R\. Rock,et al\.A cell atlas foundation model for scalable search of similar human cells\.Nature638\(8052\),pp\. 1085–1094\.Cited by:[Figure S1](https://arxiv.org/html/2609.25659#A4.F1),[Figure S1](https://arxiv.org/html/2609.25659#A4.F1.5.1),[§D\.6](https://arxiv.org/html/2609.25659#A4.SS6.p1.1),[§4\.2\.4](https://arxiv.org/html/2609.25659#S4.SS2.SSS4.p2.1)\.
- Huanget al\.\(2022\)J\. Huang, Y\. Jiao, Z\. Li, S\. Liu, Y\. Wang, and Y\. YangAn error analysis of generative adversarial networks for learning distributions\.Journal of machine learning research23\(116\),pp\. 1–43\.Cited by:[§A\.7\.2](https://arxiv.org/html/2609.25659#A1.SS7.SSS2.p8.1.1)\.
- Huguetet al\.\(2023\)G\. Huguet, A\. Tong, E\. De Brouwer, Y\. Zhang, G\. Wolf, I\. Adelstein, and S\. KrishnaswamyA heat diffusion perspective on geodesic preserving dimensionality reduction\.Advances in Neural Information Processing Systems36,pp\. 6986–7016\.Cited by:[§2\.3](https://arxiv.org/html/2609.25659#S2.SS3.SSS0.Px1.p1.1)\.
- Jinet al\.\(2025\)Y\. Jin, Z\. Sun, N\. Li, K\. Xu, H\. Jiang, N\. Zhuang, Q\. Huang, Y\. Song, Y\. Mu, and Z\. LinPyramidal flow matching for efficient video generative modeling\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 23378–23402\.Cited by:[§1](https://arxiv.org/html/2609.25659#S1.p2.1)\.
- Kingma and Ba \(2014\)D\. P\. Kingma and J\. BaAdam: a method for stochastic optimization\.arXiv preprint arXiv:1412\.6980\.Cited by:[§D\.1](https://arxiv.org/html/2609.25659#A4.SS1.SSS0.Px2.p1.1)\.
- Kleinet al\.\(2025\)D\. Klein, J\. S\. Fleck, D\. Bobrovskiy, L\. Zimmermann, S\. Becker, A\. Palma, L\. Dony, A\. Tejada\-Lapuerta, G\. Huguet, H\. Lin,et al\.CellFlow enables generative single\-cell phenotype modeling with flow matching\.bioRxiv\.Cited by:[§1](https://arxiv.org/html/2609.25659#S1.p2.1),[§4\.2\.4](https://arxiv.org/html/2609.25659#S4.SS2.SSS4.p1.1)\.
- LeCunet al\.\(1998\)Y\. LeCun, L\. Bottou, Y\. Bengio, and P\. HaffnerGradient\-based learning applied to document recognition\.Proceedings of the IEEE86\(11\),pp\. 2278–2324\.Cited by:[§D\.4](https://arxiv.org/html/2609.25659#A4.SS4.p1.1),[§4\.2\.2](https://arxiv.org/html/2609.25659#S4.SS2.SSS2.p1.1)\.
- Linet al\.\(2024\)S\. Lin, A\. Wang, and X\. YangSdxl\-lightning: progressive adversarial diffusion distillation\.arXiv preprint arXiv:2402\.13929\.Cited by:[§1](https://arxiv.org/html/2609.25659#S1.p2.1)\.
- Lipmanet al\.\(2022\)Y\. Lipman, R\. T\. Chen, H\. Ben\-Hamu, M\. Nickel, and M\. LeFlow matching for generative modeling\.arXiv preprint arXiv:2210\.02747\.Cited by:[§1](https://arxiv.org/html/2609.25659#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.25659#S2.SS2.SSS0.Px1.p1.2),[§4\.2\.1](https://arxiv.org/html/2609.25659#S4.SS2.SSS1.p1.1)\.
- Lipmanet al\.\(2024\)Y\. Lipman, M\. Havasi, P\. Holderrieth, N\. Shaul, M\. Le, B\. Karrer, R\. T\. Chen, D\. Lopez\-Paz, H\. Ben\-Hamu, and I\. GatFlow matching guide and code\.arXiv preprint arXiv:2412\.06264\.Cited by:[Appendix B](https://arxiv.org/html/2609.25659#A2.p1.1),[§C\.3](https://arxiv.org/html/2609.25659#A3.SS3.p4.1),[§2\.2](https://arxiv.org/html/2609.25659#S2.SS2.SSS0.Px1.p1.2)\.
- McCann \(2001\)R\. J\. McCannPolar factorization of maps on riemannian manifolds\.Geometric & Functional Analysis GAFA11\(3\),pp\. 589–608\.Cited by:[§2\.3](https://arxiv.org/html/2609.25659#S2.SS3.SSS0.Px1.p1.1)\.
- Mirarchiet al\.\(2024\)A\. Mirarchi, T\. Giorgino, and G\. De FabritiisMdCATH: a large\-scale md dataset for data\-driven computational biophysics\.Scientific Data11\(1\),pp\. 1299\.Cited by:[§D\.7](https://arxiv.org/html/2609.25659#A4.SS7.p1.1),[§4\.2\.5](https://arxiv.org/html/2609.25659#S4.SS2.SSS5.p2.1)\.
- Pieninget al\.\(2026\)M\. Piening, R\. Duong, and G\. SteidlGeneralized wasserstein flow matching: transport plans, everywhere, all at once\.arXiv preprint arXiv:2605\.08424\.Cited by:[§1](https://arxiv.org/html/2609.25659#S1.p2.1),[§2\.3](https://arxiv.org/html/2609.25659#S2.SS3.SSS0.Px2.p1.1)\.
- Pinzi and Savaré \(2025\)A\. Pinzi and G\. SavaréNested superposition principle for random measures and the geometry of the wasserstein on wasserstein space\.arXiv preprint arXiv:2510\.07523\.Cited by:[§C\.1\.2](https://arxiv.org/html/2609.25659#A3.SS1.SSS2.p1.1),[§C\.1\.2](https://arxiv.org/html/2609.25659#A3.SS1.SSS2.p3.2.1),[§1](https://arxiv.org/html/2609.25659#S1.p3.1),[§2\.3](https://arxiv.org/html/2609.25659#S2.SS3.SSS0.Px1.p1.1)\.
- Pinzi \(2026\)A\. PinziA study of the metric measure space of probability measures via a purely atomic superposition principle\.Nonlinear Analysis273,pp\. 114207\.Cited by:[§C\.1\.2](https://arxiv.org/html/2609.25659#A3.SS1.SSS2.p4.1.1)\.
- Pooladianet al\.\(2023\)A\. Pooladian, H\. Ben\-Hamu, C\. Domingo\-Enrich, B\. Amos, Y\. Lipman, and R\. T\. ChenMultisample flow matching: straightening flows with minibatch couplings\.arXiv preprint arXiv:2304\.14772\.Cited by:[§D\.1](https://arxiv.org/html/2609.25659#A4.SS1.SSS0.Px3.p1.1),[§2\.2](https://arxiv.org/html/2609.25659#S2.SS2.SSS0.Px1.p1.2)\.
- Pooladian and Niles\-Weed \(2021\)A\. Pooladian and J\. Niles\-WeedEntropic estimation of optimal transport maps\.arXiv preprint arXiv:2109\.12004\.Cited by:[§A\.3](https://arxiv.org/html/2609.25659#A1.SS3.SSS0.Px2.p3.1),[Appendix A](https://arxiv.org/html/2609.25659#A1.p2.1),[§1](https://arxiv.org/html/2609.25659#S1.p4.1),[§2\.1\.1](https://arxiv.org/html/2609.25659#S2.SS1.SSS1.p1.3),[§3\.2](https://arxiv.org/html/2609.25659#S3.SS2.p2.1)\.
- ROIS\-CODH \(2018\)ROIS\-CODHKMNIST dataset\.Note:[https://github\.com/rois\-codh/kmnist](https://github.com/rois-codh/kmnist)Cited by:[§D\.4](https://arxiv.org/html/2609.25659#A4.SS4.p3.1),[§4\.2\.2](https://arxiv.org/html/2609.25659#S4.SS2.SSS2.p1.1)\.
- Schneuinget al\.\(2024\)A\. Schneuing, C\. Harris, Y\. Du, K\. Didi, A\. Jamasb, I\. Igashov, W\. Du, C\. Gomes, T\. L\. Blundell, P\. Lio,et al\.Structure\-based drug design with equivariant diffusion models\.Nature Computational Science4\(12\),pp\. 899–909\.Cited by:[§1](https://arxiv.org/html/2609.25659#S1.p2.1)\.
- Starket al\.\(2024\)H\. Stark, B\. Jing, C\. Wang, G\. Corso, B\. Berger, R\. Barzilay, and T\. JaakkolaDirichlet flow matching with applications to DNA sequence design\.arXiv preprint arXiv:2402\.05841\.Cited by:[§2\.3](https://arxiv.org/html/2609.25659#S2.SS3.SSS0.Px2.p1.1)\.
- Tonget al\.\(2023\)A\. Tong, N\. Malkin, G\. Huguet, Y\. Zhang, J\. Rector\-Brooks, K\. Fatras, G\. Wolf, and Y\. BengioConditional flow matching: simulation\-free dynamic optimal transport\.arXiv preprint arXiv:2302\.00482\.Cited by:[§2\.2](https://arxiv.org/html/2609.25659#S2.SS2.SSS0.Px1.p1.2)\.
- Villani \(2008\)C\. VillaniOptimal transport: old and new\.Vol\.338,Springer\.Cited by:[§2\.1](https://arxiv.org/html/2609.25659#S2.SS1.p1.1),[§2\.3](https://arxiv.org/html/2609.25659#S2.SS3.SSS0.Px1.p1.1)\.
- Wuet al\.\(2023\)L\. Wu, D\. Wang, C\. Gong, X\. Liu, Y\. Xiong, R\. Ranjan, R\. Krishnamoorthi, V\. Chandra, and Q\. LiuFast point cloud generation with straight flows\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 9445–9454\.Cited by:[§4\.2\.1](https://arxiv.org/html/2609.25659#S4.SS2.SSS1.p1.1)\.
- You \(2026\)K\. YouBarycentric projections of optimal transport plans on riemannian manifolds\.arXiv preprint arXiv:2606\.07926\.Cited by:[§2\.3](https://arxiv.org/html/2609.25659#S2.SS3.SSS0.Px1.p1.1)\.
- Zhouet al\.\(2021\)L\. Zhou, Y\. Du, and J\. Wu3d shape generation and completion through point\-voxel diffusion\.InProceedings of the IEEE/CVF international conference on computer vision,pp\. 5826–5835\.Cited by:[§D\.3](https://arxiv.org/html/2609.25659#A4.SS3.p1.1),[§4\.2\.1](https://arxiv.org/html/2609.25659#S4.SS2.SSS1.p1.1)\.
- Zhuet al\.\(2023\)Z\. Zhu, F\. Locatello, and V\. CevherSample complexity bounds for score\-matching: causal discovery and generative modeling\.Advances in Neural Information Processing Systems36,pp\. 3325–3337\.Cited by:[§1](https://arxiv.org/html/2609.25659#S1.p2.1)\.
Appendix contents
## Appendix AEntropic Optimal Transport in Riemannian spaces
Recovering the optimal transport mapT0T\_\{0\}between a source distributionμ\\muand a target distributionν\\nuis a central task in applied optimal transport\. However, closed\-form solutions forT0T\_\{0\}are exceedingly rare, typically existing only for univariate distributions or Gaussians in Euclidean space\. Consequently, in the general setting of Riemannian manifolds, we must resort to statistical estimation from finite samples\.
To estimate the map efficiently, we leverage Entropic Optimal Transport \(EOT\)\. Unlike unregularized OT, which requires cubic\-time combinatorial solvers \(e\.g\., the Hungarian algorithm\), EOT can be solved using Sinkhorn’s algorithm\. On a Riemannian manifold, this simply requires computing the geodesic distance matrix between samples\. Once the entropic coupling is obtained, we extend the estimator proposed by[Pooladian and Niles\-Weed \[2021\]](https://arxiv.org/html/2609.25659#bib.bib21)to the Riemannian setting\.
### A\.1Problem Formulation
Let\(ℳ,g\)\(\\mathcal\{M\},g\)be a complete Riemannian manifold without boundary, and letd\(x,y\)d\(x,y\)denote the geodesic distance\. The classical Monge problem seeks a mapTTminimizing the transport cost:
infT∈𝒯\(μ,ν\)∫ℳ12d2\(x,T\(x\)\)𝑑μ\(x\)\.\\inf\_\{T\\in\\mathcal\{T\}\(\\mu,\\nu\)\}\\int\_\{\\mathcal\{M\}\}\\frac\{1\}\{2\}d^\{2\}\(x,T\(x\)\)d\\mu\(x\)\.
The Brenier\-McCann theorem shows that this problem is equivalent to
supϕ∈L1\(P\)∫ℳϕ\(x\)𝑑μ\(x\)\+∫ℳϕc\(x\)𝑑ν\(x\)\.\\sup\_\{\\phi\\in L^\{1\}\(P\)\}\\int\_\{\\mathcal\{M\}\}\\phi\(x\)d\\mu\(x\)\+\\int\_\{\\mathcal\{M\}\}\\phi^\{c\}\(x\)d\\nu\(x\)\.\(8\)
where the c\-transform is defined as
ϕc\(y\)=infx∈ℳ\{12d2\(x,y\)−ϕ\(x\)\}\.\\phi^\{c\}\(y\)=\\inf\_\{x\\in\\mathcal\{M\}\}\\\{\\frac\{1\}\{2\}d^\{2\}\(x,y\)\-\\phi\(x\)\\\}\.
The optimal transport map is then given byT0\(x\)=expx\(−∇ϕ0\(x\)\)T\_\{0\}\(x\)=\\exp\_\{x\}\(\-\\nabla\\phi\_\{0\}\(x\)\)withϕ0\\phi\_\{0\}being the maximizer of equation \([8](https://arxiv.org/html/2609.25659#A1.E8)\) and the optimal plan is concentrated whereϕ0\(x\)\+ϕ0c\(y\)=12d2\(x,y\)\\phi\_\{0\}\(x\)\+\\phi\_\{0\}^\{c\}\(y\)=\\frac\{1\}\{2\}d^\{2\}\(x,y\)\.
The relaxation to the Kantorovich problem defines the 2\-Wasserstein distance:
12W22\(μ,ν\):=min∫ℳ×ℳπ∈Π\(μ,ν\)12d2\(x,y\)𝑑π\(x,y\),\\frac\{1\}\{2\}W\_\{2\}^\{2\}\(\\mu,\\nu\):=\\min\_\{\\pi\\in\\Pi\(\\mu,\\nu\)\}\\int\_\{\\mathcal\{M\}\\times\\mathcal\{M\}\}\\frac\{1\}\{2\}d^\{2\}\(x,y\)d\\pi\(x,y\),whereΠ\(μ,ν\)\\Pi\(\\mu,\\nu\)is the set of couplings with marginalsμ\\muandν\\nu\. For a regularization parameterε\>0\\varepsilon\>0, the Entropic Optimal Transport objective is:
Sε\(μ,ν\):=infπ∈Π\(μ,ν\)∫ℳ×ℳ12d2\(x,y\)dπ\(x,y\)\+εDKL\(π∥μ⊗ν\)\.S\_\{\\varepsilon\}\(\\mu,\\nu\):=\\inf\_\{\\pi\\in\\Pi\(\\mu,\\nu\)\}\\int\_\{\\mathcal\{M\}\\times\\mathcal\{M\}\}\\frac\{1\}\{2\}d^\{2\}\(x,y\)d\\pi\(x,y\)\+\\varepsilon D\_\{KL\}\(\\pi\\\|\\mu\\otimes\\nu\)\.\(9\)The unique solutionπε\\pi\_\{\\varepsilon\}has the formdπε\(x,y\)=e\(f\(x\)\+g\(y\)−12d2\(x,y\)\)/εdμ\(x\)dν\(y\)d\\pi\_\{\\varepsilon\}\(x,y\)=e^\{\(f\(x\)\+g\(y\)\-\\frac\{1\}\{2\}d^\{2\}\(x,y\)\)/\\varepsilon\}d\\mu\(x\)d\\nu\(y\), where the potentialsf,g∈C\(ℳ\)f,g\\in C\(\\mathcal\{M\}\)solve the dual problem:
supf,g∫f𝑑μ\+∫g𝑑ν−ε∫ℳ×ℳef\(x\)\+g\(y\)−12d2\(x,y\)ε𝑑μ\(x\)𝑑ν\(y\)\+ε\.\\sup\_\{f,g\}\\int fd\\mu\+\\int gd\\nu\-\\varepsilon\\int\_\{\\mathcal\{M\}\\times\\mathcal\{M\}\}e^\{\\frac\{f\(x\)\+g\(y\)\-\\frac\{1\}\{2\}d^\{2\}\(x,y\)\}\{\\varepsilon\}\}d\\mu\(x\)d\\nu\(y\)\+\\varepsilon\.
##### Schrödinger bridge
The Schrödinger bridge problem is closely linked to the Entropic Optimal Transport problem\. Let\(Xt\)t∈\[0,1\]\(X\_\{t\}\)\_\{t\\in\[0,1\]\}denote the Brownian motion on\(ℳ,g\)\(\\mathcal\{M\},g\)with generatorε2Δg\\frac\{\\varepsilon\}\{2\}\\Delta\_\{g\}, and letRεR\_\{\\varepsilon\}be its law on path spaceΩ:=C\(\[0,1\],ℳ\)\\Omega:=C\(\[0,1\],\\mathcal\{M\}\)\. Denote byRε01R\_\{\\varepsilon\}^\{01\}the joint law of the endpoints\(X0,X1\)\(X\_\{0\},X\_\{1\}\)underRεR\_\{\\varepsilon\}; it admits a density with respect to the product of volume measures given by the heat kernel:
dRε01\(x,y\)=pε\(1,x,y\)dvol\(x\)dvol\(y\),dR\_\{\\varepsilon\}^\{01\}\(x,y\)=p\_\{\\varepsilon\}\(1,x,y\)\\,d\\mathrm\{vol\}\(x\)\\,d\\mathrm\{vol\}\(y\),wherepε\(t,x,y\)p\_\{\\varepsilon\}\(t,x,y\)solves∂tpε=ε2Δgpε\\partial\_\{t\}p\_\{\\varepsilon\}=\\frac\{\\varepsilon\}\{2\}\\Delta\_\{g\}p\_\{\\varepsilon\}\.
The Schrödinger bridge problem seeks a couplingπ\\pithat minimizes
Cε\(μ,ν\)=infπ∈Π\(μ,ν\)DKL\(π∥Rε01\)\.C\_\{\\varepsilon\}\(\\mu,\\nu\)=\\inf\_\{\\pi\\in\\Pi\(\\mu,\\nu\)\}D\_\{\\mathrm\{KL\}\}\(\\pi\\,\\\|\\,R\_\{\\varepsilon\}^\{01\}\)\.
### A\.2The Riemannian Entropic Map
In the Euclidean setting, the entropic map is defined as the barycentric projectionTε\(x\)=𝔼πε\[Y\|X=x\]T\_\{\\varepsilon\}\(x\)=\\mathbb\{E\}\_\{\\pi\_\{\\varepsilon\}\}\[Y\|X=x\]\. On a manifold, the direct expectation is not well\-defined\. Instead, we define the map via the Riemannian exponential map and the conditional expectation in the tangent space\.
###### Definition 6\(Riemannian Entropic Map\)\.
Let\(fε,gε\)\(f\_\{\\varepsilon\},g\_\{\\varepsilon\}\)be the optimal entropic potentials\. We define the Riemannian entropic mapTε:ℳ→ℳT\_\{\\varepsilon\}:\\mathcal\{M\}\\to\\mathcal\{M\}as:
Tε\(x\):=expx\(∫ℳlogx\(y\)dπεx\(y\)\),T\_\{\\varepsilon\}\(x\):=\\exp\_\{x\}\\left\(\\int\_\{\\mathcal\{M\}\}\\log\_\{x\}\(y\)d\\pi\_\{\\varepsilon\}^\{x\}\(y\)\\right\),\(10\)
wherelogx=expx−1\\log\_\{x\}=\\exp\_\{x\}^\{\-1\}is the Riemannian logarithm andπεx\\pi\_\{\\varepsilon\}^\{x\}is the conditional distribution ofYYgivenX=xX=xunder the optimal plan\. With some abuse of notation,logx\\log\_\{x\}andexpx\\exp\_\{x\}are the \(Riemannian\) logarithm and exponential maps at pointxx\. When no subscript is given,log\\logandexp\\expdenote the natural logarithm and exponential functions\.
This definition implies a straightforward computational procedure\. To evaluate the map at a source pointxx, we rely on the optimal entropic couplingπε\\pi\_\{\\varepsilon\}\. The procedure involves three steps:
1. 1\.Lift to Tangent Space:For the fixed source pointxx, we compute the Riemannian logarithmlogx\(y\)\\log\_\{x\}\(y\)for every target pointyyin the support ofν\\nu\. This maps the geometry from the manifoldℳ\\mathcal\{M\}onto the vector spaceTxℳT\_\{x\}\\mathcal\{M\}\.
2. 2\.Euclidean Averaging:We compute the weighted average of these tangent vectors, where the weights correspond to the conditional probability massπε\(y\|x\)\\pi\_\{\\varepsilon\}\(y\|x\)\. BecauseTxℳT\_\{x\}\\mathcal\{M\}is a linear space, this is a simple Euclidean average\.
3. 3\.Retract to Manifold:We apply the exponential mapexpx\\exp\_\{x\}to this average vector to project the result back onto the manifold, yielding the estimated transport location\.
We now establish the relationship between this map and the dual potential gradient\.
###### Theorem 7\.
Let\(fε,gε\)\(f\_\{\\varepsilon\},g\_\{\\varepsilon\}\)be optimal entropic potentials\. The primal feasibility ofπε\\pi\_\{\\varepsilon\}requires that its first marginal isμ\\mu\. In terms of the potentials, this constraint implies:
∫ℳefε\(x\)\+gε\(y\)−12d2\(x,y\)ε𝑑ν\(y\)=1,∀x∈supp\(μ\)\.\\int\_\{\\mathcal\{M\}\}e^\{\\frac\{f\_\{\\varepsilon\}\(x\)\+g\_\{\\varepsilon\}\(y\)\-\\frac\{1\}\{2\}d^\{2\}\(x,y\)\}\{\\varepsilon\}\}d\\nu\(y\)=1,\\quad\\forall x\\in\\text\{supp\}\(\\mu\)\.\(11\)Under this condition, the Riemannian entropic map satisfies:
Tε\(x\)=expx\(−∇fε\(x\)\)\.T\_\{\\varepsilon\}\(x\)=\\exp\_\{x\}\\left\(\-\\nabla f\_\{\\varepsilon\}\(x\)\\right\)\.
###### Proof\.
We assume the potentials satisfy \([11](https://arxiv.org/html/2609.25659#A1.E11)\)\. Taking the logarithm, we isolatefε\(x\)f\_\{\\varepsilon\}\(x\):
fε\(x\)=−εlog∫ℳexp\(gε\(y\)−12d2\(x,y\)ε\)dν\(y\)\.f\_\{\\varepsilon\}\(x\)=\-\\varepsilon\\log\\int\_\{\\mathcal\{M\}\}\\exp\\left\(\\frac\{g\_\{\\varepsilon\}\(y\)\-\\frac\{1\}\{2\}d^\{2\}\(x,y\)\}\{\\varepsilon\}\\right\)d\\nu\(y\)\.Leth\(x,y\)=exp\(gε\(y\)−12d2\(x,y\)ε\)h\(x,y\)=\\exp\\left\(\\frac\{g\_\{\\varepsilon\}\(y\)\-\\frac\{1\}\{2\}d^\{2\}\(x,y\)\}\{\\varepsilon\}\\right\)\. We compute the Riemannian gradient∇fε\(x\)\\nabla f\_\{\\varepsilon\}\(x\):
∇fε\(x\)=−ε∇x\(∫ℳh\(x,y\)𝑑ν\(y\)\)∫ℳh\(x,y\)𝑑ν\(y\)\.\\nabla f\_\{\\varepsilon\}\(x\)=\-\\varepsilon\\frac\{\\nabla\_\{x\}\\left\(\\int\_\{\\mathcal\{M\}\}h\(x,y\)d\\nu\(y\)\\right\)\}\{\\int\_\{\\mathcal\{M\}\}h\(x,y\)d\\nu\(y\)\}\.Using the identity∇x\(12d2\(x,y\)\)=−logx\(y\)\\nabla\_\{x\}\(\\frac\{1\}\{2\}d^\{2\}\(x,y\)\)=\-\\log\_\{x\}\(y\), the gradient of the integrand is:
∇xh\(x,y\)=1εh\(x,y\)logx\(y\)\.\\nabla\_\{x\}h\(x,y\)=\\frac\{1\}\{\\varepsilon\}h\(x,y\)\\log\_\{x\}\(y\)\.Substituting this back, and identifying the conditional densitydπεx\(y\)=h\(x,y\)dν\(y\)∫h\(x,z\)𝑑ν\(z\)d\\pi\_\{\\varepsilon\}^\{x\}\(y\)=\\frac\{h\(x,y\)d\\nu\(y\)\}\{\\int h\(x,z\)d\\nu\(z\)\}, we obtain:
∇fε\(x\)=−∫ℳlogx\(y\)dπεx\(y\)\.\\nabla f\_\{\\varepsilon\}\(x\)=\-\\int\_\{\\mathcal\{M\}\}\\log\_\{x\}\(y\)d\\pi\_\{\\varepsilon\}^\{x\}\(y\)\.The integral term is exactly the argument of the exponential map in our definition ofTεT\_\{\\varepsilon\}\. Thus,Tε\(x\)=expx\(−∇fε\(x\)\)T\_\{\\varepsilon\}\(x\)=\\exp\_\{x\}\(\-\\nabla f\_\{\\varepsilon\}\(x\)\)\.
### A\.3Properties of the Estimator
##### Connection to the Fréchet Mean\.
The notion of a “center of mass” on a Riemannian manifold is formalized by the Fréchet mean\. Given the conditional distributionπεx\\pi\_\{\\varepsilon\}^\{x\}of the targetYYgivenX=xX=x, the true barycentric projection is the pointm∗m^\{\*\}that minimizes the expected squared distance:
m∗=argminz∈ℳF\(z\),whereF\(z\):=12∫ℳd2\(z,y\)dπεx\(y\)\.m^\{\*\}=\\arg\\min\_\{z\\in\\mathcal\{M\}\}F\(z\),\\quad\\text\{where \}F\(z\):=\\frac\{1\}\{2\}\\int\_\{\\mathcal\{M\}\}d^\{2\}\(z,y\)d\\pi\_\{\\varepsilon\}^\{x\}\(y\)\.Findingm∗m^\{\*\}generally requires an iterative optimization procedure, as the gradient of this objective is∇F\(z\)=−∫ℳlogz\(y\)dπεx\(y\)\\nabla F\(z\)=\-\\int\_\{\\mathcal\{M\}\}\\log\_\{z\}\(y\)d\\pi\_\{\\varepsilon\}^\{x\}\(y\), leading to the implicit condition∫logm∗\(y\)dπεx\(y\)=0\\int\\log\_\{m^\{\*\}\}\(y\)d\\pi\_\{\\varepsilon\}^\{x\}\(y\)=0\.
Our estimatorTε\(x\)T\_\{\\varepsilon\}\(x\)is computed directly as the exponential of the weighted average of tangent vectors, as defined in \([10](https://arxiv.org/html/2609.25659#A1.E10)\)\. This definition possesses a geometric connection to the true center of mass as the vector we compute,v=∫ℳlogx\(y\)dπεx\(y\)v=\\int\_\{\\mathcal\{M\}\}\\log\_\{x\}\(y\)d\\pi\_\{\\varepsilon\}^\{x\}\(y\), is exactly the negative gradient of the Fréchet objective evaluated at the source pointxx:
v=−∇F\(x\)\.v=\-\\nabla F\(x\)\.Consequently, the operationTε\(x\)=expx\(v\)T\_\{\\varepsilon\}\(x\)=\\exp\_\{x\}\(v\)geometrically corresponds to takinga single gradient descent stepon the objectiveF\(z\)F\(z\), initialized atxxwith a step size of 1\. We stress that our estimator is the correct generalization of the entropic map to the Riemannian setting, and we found the connection to the Fréchet mean to be a useful geometric intuition\.
##### Recovery of the Euclidean Case\.
If we reduceℳ\\mathcal\{M\}to the Euclidean spaceℝd\\mathbb\{R\}^\{d\}equipped with the standard metric, the geometry simplifies:d\(x,y\)=‖x−y‖d\(x,y\)=\\\|x\-y\\\|, the exponential map becomes translationexpx\(v\)=x\+v\\exp\_\{x\}\(v\)=x\+v, and the logarithm becomes subtractionlogx\(y\)=y−x\\log\_\{x\}\(y\)=y\-x\. Substituting these into our definition:
Tε\(x\)=expx\(∫\(y−x\)dπεx\(y\)\)=x\+\(∫ydπεx\(y\)−x∫dπεx\(y\)\)\.T\_\{\\varepsilon\}\(x\)=\\exp\_\{x\}\\left\(\\int\(y\-x\)d\\pi\_\{\\varepsilon\}^\{x\}\(y\)\\right\)=x\+\\left\(\\int yd\\pi\_\{\\varepsilon\}^\{x\}\(y\)\-x\\int d\\pi\_\{\\varepsilon\}^\{x\}\(y\)\\right\)\.
Sinceπεx\\pi\_\{\\varepsilon\}^\{x\}is a probability distribution, it integrates to 1\. Thus:
Tε\(x\)=∫ydπεx\(y\)=𝔼πε\[Y\|X=x\]\.T\_\{\\varepsilon\}\(x\)=\\int yd\\pi\_\{\\varepsilon\}^\{x\}\(y\)=\\mathbb\{E\}\_\{\\pi\_\{\\varepsilon\}\}\[Y\|X=x\]\.
This recovers the standard barycentric projection estimator from[Pooladian and Niles\-Weed \[2021\]](https://arxiv.org/html/2609.25659#bib.bib21)\.
### A\.4McCann Interpolation via the Entropic Map
A fundamental concept in optimal transport geometry is thedisplacement interpolation, or McCann interpolation, which describes the geodesic path between probability measures in Wasserstein space\. Specifically, ifT0T\_\{0\}is the optimal transport map pushingμ\\mutoν\\nu, the interpolation at timet∈\[0,1\]t\\in\[0,1\]is the distributionμt=\(T0,t\)\#μ\\mu\_\{t\}=\(T\_\{0,t\}\)\_\{\\\#\}\\mu, whereT0,t\(x\)T\_\{0,t\}\(x\)moves the mass atxxa fractionttof the way along the geodesic towardT0\(x\)T\_\{0\}\(x\)\.
Our Riemannian entropic map offers a constructive way to approximate this interpolation\. Because the mapTε\(x\)T\_\{\\varepsilon\}\(x\)is constructed via the exponential map of a tangent vectorvx=∫logx\(y\)dπεx\(y\)v\_\{x\}=\\int\\log\_\{x\}\(y\)d\\pi\_\{\\varepsilon\}^\{x\}\(y\), the geodesic connectingxxtoTε\(x\)T\_\{\\varepsilon\}\(x\)is simply the curveγ\(t\)=expx\(t⋅vx\)\\gamma\(t\)=\\exp\_\{x\}\(t\\cdot v\_\{x\}\)\.
###### Definition 8\(Entropic Interpolant\)\.
For any timet∈\[0,1\]t\\in\[0,1\], we define the time\-ttentropic mapTε,t:ℳ→ℳT\_\{\\varepsilon,t\}:\\mathcal\{M\}\\to\\mathcal\{M\}as:
Tε,t\(x\):=expx\(t⋅∫ℳlogx\(y\)dπεx\(y\)\)\.T\_\{\\varepsilon,t\}\(x\):=\\exp\_\{x\}\\left\(t\\cdot\\int\_\{\\mathcal\{M\}\}\\log\_\{x\}\(y\)d\\pi\_\{\\varepsilon\}^\{x\}\(y\)\\right\)\.The estimated McCann interpolation at timettis the pushforward measureμ^t:=\(Tε,t\)\#μ\\hat\{\\mu\}\_\{t\}:=\(T\_\{\\varepsilon,t\}\)\_\{\\\#\}\\mu\.
This formulation is geometrically intuitive: we compute the aggregate direction of transportvxv\_\{x\}in the tangent space and simply scale it byttbefore projecting back to the manifold\. Att=0t=0, the argument to the exponential map is the zero vector, recovering the identity map \(Tε,0\(x\)=xT\_\{\\varepsilon,0\}\(x\)=x\)\. Att=1t=1, we recover the full Riemannian entropic map \(Tε,1\(x\)=Tε\(x\)T\_\{\\varepsilon,1\}\(x\)=T\_\{\\varepsilon\}\(x\)\)\. For intermediatett, this generates a distribution supported on the geodesic flow between the source and the estimated target\.
### A\.5Out\-of\-Sample Estimation
A critical feature of the proposed estimator is its ability to generalize to points outside the initial training set\. Letν^=∑j=1n𝐛jδYj\\hat\{\\nu\}=\\sum\_\{j=1\}^\{n\}\\mathbf\{b\}\_\{j\}\\delta\_\{Y\_\{j\}\}be the discrete empirical measure of the target distribution, and let𝐠∈ℝn\\mathbf\{g\}\\in\\mathbb\{R\}^\{n\}be the optimal dual potential vector obtained from Sinkhorn’s algorithm on training samples\.
We seek to evaluate the mapTε\(x′\)T\_\{\\varepsilon\}\(x^\{\\prime\}\)for a new source pointx′∈ℳx^\{\\prime\}\\in\\mathcal\{M\}that was not present during training\. This evaluation is not an ad\-hoc interpolation but a direct consequence of theextension principlein semi\-discrete optimal transport\.
##### The\(c,ε\)\(c,\\varepsilon\)\-Transform and Dual Extension\.
In classical Optimal Transport \(ε=0\\varepsilon=0\), the relationship between optimal potentials is governed by thecc\-transform \(orcc\-conjugate\), which generalizes the Legendre\-Fenchel transform\. The source potentialffis obtained from the target potentialggvia a “hard” infimum:
f\(x\)=gc\(x\):=infy∈ℳ\(12d2\(x,y\)−g\(y\)\)\.f\(x\)=g^\{c\}\(x\):=\\inf\_\{y\\in\\mathcal\{M\}\}\\left\(\\frac\{1\}\{2\}d^\{2\}\(x,y\)\-g\(y\)\\right\)\.
In Entropic Optimal Transport, this non\-smooth operator is replaced by a “soft” minimum known as the\(c,ε\)\(c,\\varepsilon\)\-transform:
fε\(x\)=g\(c,ε\)\(x\):=−εlog\(∫ℳexp\(g\(y\)−12d2\(x,y\)ε\)𝑑ν\(y\)\)\.f\_\{\\varepsilon\}\(x\)=g^\{\(c,\\varepsilon\)\}\(x\):=\-\\varepsilon\\log\\left\(\\int\_\{\\mathcal\{M\}\}\\exp\\left\(\\frac\{g\(y\)\-\\frac\{1\}\{2\}d^\{2\}\(x,y\)\}\{\\varepsilon\}\\right\)d\\nu\(y\)\\right\)\.
In our estimation setting, we solve the dual problem using the discrete empirical measureν^\\hat\{\\nu\}\. While the resulting optimal potential𝐠\\mathbf\{g\}is a vector inℝn\\mathbb\{R\}^\{n\}, the\(c,ε\)\(c,\\varepsilon\)\-transform interprets these values as coefficients of a Riemannian kernel expansion\. Substituting the discrete measureQ^\\hat\{Q\}into the integral definition yields a globally defined, smooth potentialfε:ℳ→ℝf\_\{\\varepsilon\}:\\mathcal\{M\}\\to\\mathbb\{R\}:
fε\(x′\)=−εlog\(∑j=1n𝐛jexp\(𝐠j−12d2\(x′,Yj\)ε\)\),f\_\{\\varepsilon\}\(x^\{\\prime\}\)=\-\\varepsilon\\log\\left\(\\sum\_\{j=1\}^\{n\}\\mathbf\{b\}\_\{j\}\\exp\\left\(\\frac\{\\mathbf\{g\}\_\{j\}\-\\frac\{1\}\{2\}d^\{2\}\(x^\{\\prime\},Y\_\{j\}\)\}\{\\varepsilon\}\\right\)\\right\),\(12\)where𝐛j\\mathbf\{b\}\_\{j\}are the target weights \(typically1/n1/n\)\.
##### Derivation of the Out\-of\-Sample Map\.
We apply our main theorem,Tε\(x′\)=expx′\(−∇fε\(x′\)\)T\_\{\\varepsilon\}\(x^\{\\prime\}\)=\\exp\_\{x^\{\\prime\}\}\(\-\\nabla f\_\{\\varepsilon\}\(x^\{\\prime\}\)\), to this extended potential\. Differentiating \([12](https://arxiv.org/html/2609.25659#A1.E12)\) with respect tox′x^\{\\prime\}involves the gradient of the log\-sum\-exp function\. By the chain rule:
∇fε\(x′\)=−ε∑j=1n∇x′\[𝐛jexp\(𝐠j−12d2\(x′,Yj\)ε\)\]∑k=1n𝐛kexp\(𝐠k−12d2\(x′,Yk\)ε\)=−ε∑j=1n𝐛jexp\(𝐠j−12d2\(x′,Yj\)ε\)∑𝐛kexp\(𝐠k−12d2\(x′,Yk\)ε\)⋅∇x′\(𝐠j−12d2\(x′,Yj\)ε\)\.\\begin\{split\}\\nabla f\_\{\\varepsilon\}\(x^\{\\prime\}\)&=\-\\varepsilon\\frac\{\\sum\_\{j=1\}^\{n\}\\nabla\_\{x^\{\\prime\}\}\\left\[\\mathbf\{b\}\_\{j\}\\exp\\left\(\\frac\{\\mathbf\{g\}\_\{j\}\-\\frac\{1\}\{2\}d^\{2\}\(x^\{\\prime\},Y\_\{j\}\)\}\{\\varepsilon\}\\right\)\\right\]\}\{\\sum\_\{k=1\}^\{n\}\\mathbf\{b\}\_\{k\}\\exp\\left\(\\frac\{\\mathbf\{g\}\_\{k\}\-\\frac\{1\}\{2\}d^\{2\}\(x^\{\\prime\},Y\_\{k\}\)\}\{\\varepsilon\}\\right\)\}\\\\ &=\-\\varepsilon\\sum\_\{j=1\}^\{n\}\\frac\{\\mathbf\{b\}\_\{j\}\\exp\(\\frac\{\\mathbf\{g\}\_\{j\}\-\\frac\{1\}\{2\}d^\{2\}\(x^\{\\prime\},Y\_\{j\}\)\}\{\\varepsilon\}\)\}\{\\sum\\mathbf\{b\}\_\{k\}\\exp\(\\frac\{\\mathbf\{g\}\_\{k\}\-\\frac\{1\}\{2\}d^\{2\}\(x^\{\\prime\},Y\_\{k\}\)\}\{\\varepsilon\}\)\}\\cdot\\nabla\_\{x^\{\\prime\}\}\\left\(\\frac\{\\mathbf\{g\}\_\{j\}\-\\frac\{1\}\{2\}d^\{2\}\(x^\{\\prime\},Y\_\{j\}\)\}\{\\varepsilon\}\\right\)\.\\end\{split\}Since𝐠j\\mathbf\{g\}\_\{j\}is constant with respect tox′x^\{\\prime\}, the inner gradient is simply−12ε∇x′d2\(x′,Yj\)=1εlogx′\(Yj\)\-\\frac\{1\}\{2\\varepsilon\}\\nabla\_\{x^\{\\prime\}\}d^\{2\}\(x^\{\\prime\},Y\_\{j\}\)=\\frac\{1\}\{\\varepsilon\}\\log\_\{x^\{\\prime\}\}\(Y\_\{j\}\)\. Substituting this back yields:
−∇fε\(x′\)=∑j=1nwj\(x′\)logx′\(Yj\),\-\\nabla f\_\{\\varepsilon\}\(x^\{\\prime\}\)=\\sum\_\{j=1\}^\{n\}w\_\{j\}\(x^\{\\prime\}\)\\log\_\{x^\{\\prime\}\}\(Y\_\{j\}\),where the attention weightswj\(x′\)w\_\{j\}\(x^\{\\prime\}\)are given by the softmax function:
wj\(x′\)=𝐛jexp\(𝐠j−12d2\(x′,Yj\)ε\)∑k=1n𝐛kexp\(𝐠k−12d2\(x′,Yk\)ε\)\.w\_\{j\}\(x^\{\\prime\}\)=\\frac\{\\mathbf\{b\}\_\{j\}\\exp\\left\(\\frac\{\\mathbf\{g\}\_\{j\}\-\\frac\{1\}\{2\}d^\{2\}\(x^\{\\prime\},Y\_\{j\}\)\}\{\\varepsilon\}\\right\)\}\{\\sum\_\{k=1\}^\{n\}\\mathbf\{b\}\_\{k\}\\exp\\left\(\\frac\{\\mathbf\{g\}\_\{k\}\-\\frac\{1\}\{2\}d^\{2\}\(x^\{\\prime\},Y\_\{k\}\)\}\{\\varepsilon\}\\right\)\}\.Finally, the transport map for the new pointx′x^\{\\prime\}is computed by projecting this weighted average back onto the manifold:
Tε\(x′\)=expx′\(∑j=1nwj\(x′\)logx′\(Yj\)\)\.T\_\{\\varepsilon\}\(x^\{\\prime\}\)=\\exp\_\{x^\{\\prime\}\}\\left\(\\sum\_\{j=1\}^\{n\}w\_\{j\}\(x^\{\\prime\}\)\\log\_\{x^\{\\prime\}\}\(Y\_\{j\}\)\\right\)\.
### A\.6Computing the Riemannian entropic map without analyticexp\\expandlog\\logmaps
When closed\-formlog\\logandexp\\expmaps are not available, the entropic estimator remains computationally viable as long as the geodesic distance functiondg\(⋅,⋅\)d\_\{g\}\(\\cdot,\\cdot\)onℳ\\mathcal\{M\}is accessible\. Specifically, we assume we have access to: \(i\) the geodesic distancedg\(x,y\)d\_\{g\}\(x,y\)and its extrinsic gradient∇xdg\(x,y\)∈ℝd\\nabla\_\{x\}d\_\{g\}\(x,y\)\\in\\mathbb\{R\}^\{d\}, \(ii\) a projection operatorΠℳ:ℝd→ℳ\\Pi\_\{\\mathcal\{M\}\}:\\mathbb\{R\}^\{d\}\\to\\mathcal\{M\}mapping ambient points to the closest point onℳ\\mathcal\{M\}, and \(iii\) the Jacobian of the projection operatorJΠ\(x\)∈ℝd×dJ\_\{\\Pi\}\(x\)\\in\\mathbb\{R\}^\{d\\times d\}, which at pointsx∈ℳx\\in\\mathcal\{M\}projects ambient vectors onto the tangent spaceTxℳT\_\{x\}\\mathcal\{M\}\.
##### Approximating the logarithmic map\.
TheN2N^\{2\}logarithmic map evaluations in the lift step can be obtained from the gradient of the squared distance\. DefineD\(x,y\)=12dg\(x,y\)2D\(x,y\)=\\frac\{1\}\{2\}d\_\{g\}\(x,y\)^\{2\}\. Its extrinsic gradient with respect toxxis∇xD\(x,y\)=dg\(x,y\)∇xdg\(x,y\)\\nabla\_\{x\}D\(x,y\)=d\_\{g\}\(x,y\)\\,\\nabla\_\{x\}d\_\{g\}\(x,y\)\. Projecting onto the tangent space atxxvia the Jacobian of the manifold projection gives:
logx\(y\)=−JΠ\(x\)\(dg\(x,y\)∇xdg\(x,y\)\)\.\\log\_\{x\}\(y\)=\-J\_\{\\Pi\}\(x\)\\big\(d\_\{g\}\(x,y\)\\,\\nabla\_\{x\}d\_\{g\}\(x,y\)\\big\)\.Substituting into the Riemannian entropic map, the full lift\-average\-retract computation becomes:
Tε\(xi\)=expxi\(−∑j=1Nπε\(yj\|xi\)JΠ\(xi\)\(dg\(xi,yj\)∇xidg\(xi,yj\)\)\)\.T\_\{\\varepsilon\}\(x\_\{i\}\)=\\exp\_\{x\_\{i\}\}\\\!\\left\(\-\\sum\_\{j=1\}^\{N\}\\pi\_\{\\varepsilon\}\(y\_\{j\}\|x\_\{i\}\)\\,J\_\{\\Pi\}\(x\_\{i\}\)\\big\(d\_\{g\}\(x\_\{i\},y\_\{j\}\)\\,\\nabla\_\{x\_\{i\}\}d\_\{g\}\(x\_\{i\},y\_\{j\}\)\\big\)\\right\)\.Since the Sinkhorn algorithm already requires the cost matrixCij=12dg2\(xi,yj\)C\_\{ij\}=\\frac\{1\}\{2\}d\_\{g\}^\{2\}\(x\_\{i\},y\_\{j\}\), the distance gradients can be obtained by a backward pass through the same computation at negligible additional cost and the tangent projection can be performed with a Jacobian\-vector product\.
##### Approximating the exponential map via projected Euler integration\.
TheNNexponential map evaluations in the retract step can be approximated by numerically integrating the geodesic using only the projection operator\. Givenx∈ℳx\\in\\mathcal\{M\}andv∈Txℳv\\in T\_\{x\}\\mathcal\{M\}, we seeky=expx\(v\)y=\\exp\_\{x\}\(v\)\. For a given number of stepsKK, we can approximate this geodesic by iteratively taking small steps in the ambient space and projecting the position and velocity back onto the manifold\.
1. 1\.Euler step:y~k\+1=yk\+1Kvk\\tilde\{y\}\_\{k\+1\}=y\_\{k\}\+\\frac\{1\}\{K\}v\_\{k\}
2. 2\.Position projection:yk\+1=Πℳ\(y~k\+1\)y\_\{k\+1\}=\\Pi\_\{\\mathcal\{M\}\}\(\\tilde\{y\}\_\{k\+1\}\)
3. 3\.Velocity projection:v~k\+1=JΠ\(yk\+1\)vk\\tilde\{v\}\_\{k\+1\}=J\_\{\\Pi\}\(y\_\{k\+1\}\)v\_\{k\}
4. 4\.Speed renormalization:vk\+1=∥v∥⋅v~k\+1/∥v~k\+1∥v\_\{k\+1\}=\\lVert v\\rVert\\cdot\\tilde\{v\}\_\{k\+1\}/\\lVert\\tilde\{v\}\_\{k\+1\}\\rVert
Where we start withy0=xy\_\{0\}=xandv0=vv\_\{0\}=v\. The last step is to ensure that the speed of the geodesic is preserved\. This procedure requires only the projection operator and its Jacobian\-vector product\. This formulation opens the door to applying RWEFM on geometries where analyticexp\\expandlog\\logmaps are unavailable, such as triangulated meshes, learned Riemannian metrics, or implicit surfaces; we demonstrate a proof of concept on a triangulated mesh \(the Stanford bunny\) in Section[4\.2\.3](https://arxiv.org/html/2609.25659#S4.SS2.SSS3)\.
### A\.7Statistical performance of the estimator
We now establish a finite\-sample guarantee for the Riemannian entropic map\.
Throughout this section, we write
c\(x,y\):=12dg2\(x,y\),c\(x,y\):=\\frac\{1\}\{2\}d\_\{g\}^\{2\}\(x,y\),and let\(φ0,φ0c\)\(\\varphi\_\{0\},\\varphi\_\{0\}^\{c\}\)be a pair of optimal Kantorovich potentials for the unregularized problem\. Forx∈Ωx\\in\\Omega, define
x⋆:=T0\(x\),v0\(x\):=logxT0\(x\),x^\{\\star\}:=T\_\{0\}\(x\),\\qquad v\_\{0\}\(x\):=\\log\_\{x\}T\_\{0\}\(x\),and the Kantorovich duality gap
Dx\(y\):=c\(x,y\)−φ0\(x\)−φ0c\(y\)\.D\_\{x\}\(y\):=c\(x,y\)\-\\varphi\_\{0\}\(x\)\-\\varphi\_\{0\}^\{c\}\(y\)\.\(13\)By optimality,Dx\(y\)≥0D\_\{x\}\(y\)\\geq 0andDx\(T0\(x\)\)=0D\_\{x\}\(T\_\{0\}\(x\)\)=0\.
#### A\.7\.1Regularity assumptions onMMandT0T\_\{0\}
###### Assumption 9\(Geometry and convexity\)\.
There exists a compact, strongly geodesically convex setΩ⊂M\\Omega\\subset Msuch thatsupp\(μ\),supp\(ν\)⊂Ω\\operatorname\{supp\}\(\\mu\),\\operatorname\{supp\}\(\\nu\)\\subset\\Omega\. For everyx∈Ωx\\in\\Omega, the logarithmic maplogx:Ω→TxM\\log\_\{x\}:\\Omega\\to T\_\{x\}Mis single\-valued and smooth\. Moreover, on the relevant tangent balls the exponential map obeys the uniform Lipschitz bound
dg\(expx\(u\),expx\(v\)\)≤Cexp‖u−v‖g,x∈Ω\.d\_\{g\}\\\!\\left\(\\exp\_\{x\}\(u\),\\exp\_\{x\}\(v\)\\right\)\\leq C\_\{\\exp\}\\,\\\|u\-v\\\|\_\{g\},\\qquad x\\in\\Omega\.
###### Assumption 10\(Density bounds\)\.
The measuresμ\\muandν\\nuadmit densitiesfμ,fνf\_\{\\mu\},f\_\{\\nu\}with respect to Riemannian volume onΩ\\Omega, and
0<mρ≤fμ\(x\),fν\(x\)≤Mρ<∞,x∈Ω\.0<m\_\{\\rho\}\\leq f\_\{\\mu\}\(x\),f\_\{\\nu\}\(x\)\\leq M\_\{\\rho\}<\\infty,\\qquad x\\in\\Omega\.
###### Assumption 11\(Regularity and quadratic detachment of the Monge map\)\.
The Monge problem for the costc\(x,y\)=12dg2\(x,y\)c\(x,y\)=\\frac\{1\}\{2\}d\_\{g\}^\{2\}\(x,y\)admits a unique optimal mapT0:Ω→ΩT\_\{0\}:\\Omega\\to\\Omega, which is a diffeomorphism\. There exist constants0<λ≤Λ<∞0<\\lambda\\leq\\Lambda<\\inftysuch that, uniformly forx,y∈Ωx,y\\in\\Omega,
λ2dg2\(y,T0\(x\)\)≤Dx\(y\)≤Λ2dg2\(y,T0\(x\)\)\.\\frac\{\\lambda\}\{2\}d\_\{g\}^\{2\}\(y,T\_\{0\}\(x\)\)\\leq D\_\{x\}\(y\)\\leq\\frac\{\\Lambda\}\{2\}d\_\{g\}^\{2\}\(y,T\_\{0\}\(x\)\)\.\(14\)
Assumption[9](https://arxiv.org/html/2609.25659#Thmtheorem9)implies thatccis smooth on the compact setΩ×Ω\\Omega\\times\\Omega, hence all derivatives ofccthat appear below are uniformly bounded\. It also implies the following two basic geometric facts\. First, there is a constantClogC\_\{\\log\}such that
‖logx\(y\)−logx\(z\)‖g≤Clogdg\(y,z\),x,y,z∈Ω\.\\\|\\log\_\{x\}\(y\)\-\\log\_\{x\}\(z\)\\\|\_\{g\}\\leq C\_\{\\log\}d\_\{g\}\(y,z\),\\qquad x,y,z\\in\\Omega\.\(15\)Second, sinceΩ\\Omegais a compact subset of add\-dimensional smooth manifold, it admits measurable partitions into at mostCδ−dC\\delta^\{\-d\}sets of geodesic diameter at mostδ\\delta, for every sufficiently smallδ\>0\\delta\>0\.
Let
μn=1n∑i=1nδXi,νn=1n∑j=1nδYj,\\mu\_\{n\}=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\delta\_\{X\_\{i\}\},\\qquad\\nu\_\{n\}=\\frac\{1\}\{n\}\\sum\_\{j=1\}^\{n\}\\delta\_\{Y\_\{j\}\},whereXi∼iidμX\_\{i\}\\stackrel\{\{\\scriptstyle\\mathrm\{iid\}\}\}\{\{\\sim\}\}\\muandYj∼iidνY\_\{j\}\\stackrel\{\{\\scriptstyle\\mathrm\{iid\}\}\}\{\{\\sim\}\}\\nu, with the two samples independent\. We denote byπε,n\\pi\_\{\\varepsilon,n\}the optimal entropic plan betweenμ\\muandνn\\nu\_\{n\}, and byTε,nT\_\{\\varepsilon,n\}its Riemannian barycentric map,
Tε,n\(x\):=expx\(bε,n\(x\)\),bε,n\(x\):=∫Ωlogx\(y\)dπε,nx\(y\)\.T\_\{\\varepsilon,n\}\(x\):=\\exp\_\{x\}\\\!\\left\(b\_\{\\varepsilon,n\}\(x\)\\right\),\\qquad b\_\{\\varepsilon,n\}\(x\):=\\int\_\{\\Omega\}\\log\_\{x\}\(y\)\\,d\\pi\_\{\\varepsilon,n\}^\{x\}\(y\)\.
For the two\-sample estimator, let\(f^,g^\)\(\\widehat\{f\},\\widehat\{g\}\)be optimal entropic potentials for\(μn,νn\)\(\\mu\_\{n\},\\nu\_\{n\}\)\. We use the canonical out\-of\-sample extension already described in Section A\.5: for everyx∈Ωx\\in\\Omega,
f^\(x\):=−εlog\[1n∑j=1nexp\(g^\(Yj\)−c\(x,Yj\)ε\)\],\\widehat\{f\}\(x\):=\-\\varepsilon\\log\\\!\\left\[\\frac\{1\}\{n\}\\sum\_\{j=1\}^\{n\}\\exp\\\!\\left\(\\frac\{\\widehat\{g\}\(Y\_\{j\}\)\-c\(x,Y\_\{j\}\)\}\{\\varepsilon\}\\right\)\\right\],\(16\)and
w^j\(x\)\\displaystyle\\widehat\{w\}\_\{j\}\(x\):=exp\(\(g^\(Yj\)−c\(x,Yj\)\)/ε\)∑k=1nexp\(\(g^\(Yk\)−c\(x,Yk\)\)/ε\),\\displaystyle:=\\frac\{\\exp\\\!\\left\(\(\\widehat\{g\}\(Y\_\{j\}\)\-c\(x,Y\_\{j\}\)\)/\\varepsilon\\right\)\}\{\\sum\_\{k=1\}^\{n\}\\exp\\\!\\left\(\(\\widehat\{g\}\(Y\_\{k\}\)\-c\(x,Y\_\{k\}\)\)/\\varepsilon\\right\)\},\(17\)b^ε,\(n,n\)\(x\)\\displaystyle\\widehat\{b\}\_\{\\varepsilon,\(n,n\)\}\(x\):=∑j=1nw^j\(x\)logx\(Yj\),T^ε,\(n,n\)\(x\):=expx\(b^ε,\(n,n\)\(x\)\)\.\\displaystyle:=\\sum\_\{j=1\}^\{n\}\\widehat\{w\}\_\{j\}\(x\)\\log\_\{x\}\(Y\_\{j\}\),\\qquad\\widehat\{T\}\_\{\\varepsilon,\(n,n\)\}\(x\):=\\exp\_\{x\}\\\!\\left\(\\widehat\{b\}\_\{\\varepsilon,\(n,n\)\}\(x\)\\right\)\.\(18\)At the sampled source points, this agrees with the barycentric projection of the empirical optimal entropic coupling\.
Define
sd:=max\{1,d2\}\.s\_\{d\}:=\\max\\left\\\{1,\\frac\{d\}\{2\}\\right\\\}\.\(19\)The use ofsds\_\{d\}only matters in dimension one; for everyd≥2d\\geq 2,sd=d/2s\_\{d\}=d/2\.
###### Theorem 12\(Statistical performance of the Riemannian entropic map\)\.
Under Assumptions[9](https://arxiv.org/html/2609.25659#Thmtheorem9)–[11](https://arxiv.org/html/2609.25659#Thmtheorem11), there exist constantsC<∞C<\\inftyandε0∈\(0,1\]\\varepsilon\_\{0\}\\in\(0,1\], depending only on the regularity constants and the geometry ofΩ\\Omega, such that for every0<ε≤ε00<\\varepsilon\\leq\\varepsilon\_\{0\}andn≥2n\\geq 2,
𝔼∫Ωdg2\(T^ε,\(n,n\)\(x\),T0\(x\)\)dμ\(x\)≤C\[εlog\(eε\)\+ε−sdlog\(n\+1\)n\]\.\\begin\{split\}\\mathbb\{E\}\\int\_\{\\Omega\}d\_\{g\}^\{2\}\\\!\\left\(\\widehat\{T\}\_\{\\varepsilon,\(n,n\)\}\(x\),T\_\{0\}\(x\)\\right\)\\,d\\mu\(x\)\\leq C\\left\[\\varepsilon\\log\\\!\\left\(\\frac\{e\}\{\\varepsilon\}\\right\)\+\\varepsilon^\{\-s\_\{d\}\}\\frac\{\\log\(n\+1\)\}\{\\sqrt\{n\}\}\\right\]\.\\end\{split\}\(20\)In particular, ford≥2d\\geq 2,
𝔼∫Ωdg2\(T^ε,\(n,n\)\(x\),T0\(x\)\)dμ\(x\)≲εlog\(eε\)\+ε−d/2log\(n\+1\)n\.\\mathbb\{E\}\\int\_\{\\Omega\}d\_\{g\}^\{2\}\\\!\\left\(\\widehat\{T\}\_\{\\varepsilon,\(n,n\)\}\(x\),T\_\{0\}\(x\)\\right\)\\,d\\mu\(x\)\\lesssim\\varepsilon\\log\\\!\\left\(\\frac\{e\}\{\\varepsilon\}\\right\)\+\\varepsilon^\{\-d/2\}\\frac\{\\log\(n\+1\)\}\{\\sqrt\{n\}\}\.\(21\)Consequently, choosingε≍n−1/\(d\+2\)\\varepsilon\\asymp n^\{\-1/\(d\+2\)\}ford≥2d\\geq 2gives, up to logarithmic factors,
𝔼∫Ωdg2\(T^ε,\(n,n\)\(x\),T0\(x\)\)dμ\(x\)=O~\(n−1/\(d\+2\)\)\.\\mathbb\{E\}\\int\_\{\\Omega\}d\_\{g\}^\{2\}\\\!\\left\(\\widehat\{T\}\_\{\\varepsilon,\(n,n\)\}\(x\),T\_\{0\}\(x\)\\right\)\\,d\\mu\(x\)=\\widetilde\{O\}\\\!\\left\(n^\{\-1/\(d\+2\)\}\\right\)\.
#### A\.7\.2Deterministic and empirical ingredients
We first record a deterministic observation converting the Kantorovich duality gap into map error\.
###### Lemma 13\(Duality gap controls barycentric map error\)\.
Letπ∈Π\(μ,η\)\\pi\\in\\Pi\(\\mu,\\eta\)be any coupling whose second marginalη\\etais supported inΩ\\Omega\. Define
bπ\(x\):=∫Ωlogx\(y\)dπx\(y\),Tπ\(x\):=expx\(bπ\(x\)\)\.b\_\{\\pi\}\(x\):=\\int\_\{\\Omega\}\\log\_\{x\}\(y\)\\,d\\pi^\{x\}\(y\),\\qquad T\_\{\\pi\}\(x\):=\\exp\_\{x\}\(b\_\{\\pi\}\(x\)\)\.Then
∫Ωdg2\(Tπ\(x\),T0\(x\)\)𝑑μ\(x\)≤C∫Ω×ΩDx\(y\)𝑑π\(x,y\)\.\\int\_\{\\Omega\}d\_\{g\}^\{2\}\(T\_\{\\pi\}\(x\),T\_\{0\}\(x\)\)\\,d\\mu\(x\)\\leq C\\int\_\{\\Omega\\times\\Omega\}D\_\{x\}\(y\)\\,d\\pi\(x,y\)\.\(22\)
###### Proof\.
By Assumption[9](https://arxiv.org/html/2609.25659#Thmtheorem9)and Jensen’s inequality,
dg2\(Tπ\(x\),T0\(x\)\)\\displaystyle d\_\{g\}^\{2\}\(T\_\{\\pi\}\(x\),T\_\{0\}\(x\)\)≤Cexp2‖bπ\(x\)−v0\(x\)‖g2\\displaystyle\\leq C\_\{\\exp\}^\{2\}\\\|b\_\{\\pi\}\(x\)\-v\_\{0\}\(x\)\\\|\_\{g\}^\{2\}≤Cexp2∫Ω‖logx\(y\)−logx\(T0\(x\)\)‖g2dπx\(y\)\.\\displaystyle\\leq C\_\{\\exp\}^\{2\}\\int\_\{\\Omega\}\\\|\\log\_\{x\}\(y\)\-\\log\_\{x\}\(T\_\{0\}\(x\)\)\\\|\_\{g\}^\{2\}\\,d\\pi^\{x\}\(y\)\.Using \([15](https://arxiv.org/html/2609.25659#A1.E15)\) and the lower bound in \([14](https://arxiv.org/html/2609.25659#A1.E14)\),
‖logx\(y\)−logx\(T0\(x\)\)‖g2≤Cdg2\(y,T0\(x\)\)≤CDx\(y\)\.\\\|\\log\_\{x\}\(y\)\-\\log\_\{x\}\(T\_\{0\}\(x\)\)\\\|\_\{g\}^\{2\}\\leq Cd\_\{g\}^\{2\}\(y,T\_\{0\}\(x\)\)\\leq CD\_\{x\}\(y\)\.Integrating first inyyand then inxxproves the claim\.
The next lemma bounds the regularization bias without any heat\-kernel representation\.
###### Lemma 14\(Entropic cost bias\)\.
There existsC<∞C<\\inftysuch that, for all sufficiently smallε∈\(0,1\]\\varepsilon\\in\(0,1\],
0≤Sε\(μ,ν\)−12W22\(μ,ν\)≤Cεlog\(eε\)\.0\\leq S\_\{\\varepsilon\}\(\\mu,\\nu\)\-\\frac\{1\}\{2\}W\_\{2\}^\{2\}\(\\mu,\\nu\)\\leq C\\varepsilon\\log\\\!\\left\(\\frac\{e\}\{\\varepsilon\}\\right\)\.\(23\)
###### Proof\.
The lower bound is immediate because the transport part of the entropic objective is at least the unregularized optimal cost and the relative entropy is nonnegative\.
For the upper bound, fix0<δ<10<\\delta<1\. By compactness ofΩ\\Omega, choose a measurable partition\{Bk\}k=1N\\\{B\_\{k\}\\\}\_\{k=1\}^\{N\}ofΩ\\Omegasuch that
diamg\(Bk\)≤δ,N≤Cδ−d\.\\operatorname\{diam\}\_\{g\}\(B\_\{k\}\)\\leq\\delta,\\qquad N\\leq C\\delta^\{\-d\}\.SetAk:=T0−1\(Bk\)A\_\{k\}:=T\_\{0\}^\{\-1\}\(B\_\{k\}\)andpk:=μ\(Ak\)=ν\(Bk\)p\_\{k\}:=\\mu\(A\_\{k\}\)=\\nu\(B\_\{k\}\)\. Ignoring cells withpk=0p\_\{k\}=0, define
πδ:=∑k=1N1pkμ\|Ak⊗ν\|Bk\.\\pi^\{\\delta\}:=\\sum\_\{k=1\}^\{N\}\\frac\{1\}\{p\_\{k\}\}\\,\\mu\|\_\{A\_\{k\}\}\\otimes\\nu\|\_\{B\_\{k\}\}\.\(24\)Thenπδ∈Π\(μ,ν\)\\pi^\{\\delta\}\\in\\Pi\(\\mu,\\nu\)and
DKL\(πδ∥μ⊗ν\)=∑k=1Npklog1pk≤logN≤C\+dlog1δ\.D\_\{\\mathrm\{KL\}\}\(\\pi^\{\\delta\}\\\|\\mu\\otimes\\nu\)=\\sum\_\{k=1\}^\{N\}p\_\{k\}\\log\\frac\{1\}\{p\_\{k\}\}\\leq\\log N\\leq C\+d\\log\\frac\{1\}\{\\delta\}\.\(25\)Moreover, if\(x,y\)∈Ak×Bk\(x,y\)\\in A\_\{k\}\\times B\_\{k\}, then bothT0\(x\)T\_\{0\}\(x\)andyybelong toBkB\_\{k\}, so the upper bound in \([14](https://arxiv.org/html/2609.25659#A1.E14)\) givesDx\(y\)≤Cδ2D\_\{x\}\(y\)\\leq C\\delta^\{2\}\. Sinceπδ\\pi^\{\\delta\}and the optimal Monge plan have the same marginals, Kantorovich duality yields
∫cdπδ−12W22\(μ,ν\)=∫Dx\(y\)dπδ\(x,y\)≤Cδ2\.\\int c\\,d\\pi^\{\\delta\}\-\\frac\{1\}\{2\}W\_\{2\}^\{2\}\(\\mu,\\nu\)=\\int D\_\{x\}\(y\)\\,d\\pi^\{\\delta\}\(x,y\)\\leq C\\delta^\{2\}\.\(26\)Usingπδ\\pi^\{\\delta\}as a competitor in the entropic problem therefore gives
Sε\(μ,ν\)−12W22\(μ,ν\)≤Cδ2\+Cε\+dεlog1δ\.S\_\{\\varepsilon\}\(\\mu,\\nu\)\-\\frac\{1\}\{2\}W\_\{2\}^\{2\}\(\\mu,\\nu\)\\leq C\\delta^\{2\}\+C\\varepsilon\+d\\varepsilon\\log\\frac\{1\}\{\\delta\}\.Takingδ=ε\\delta=\\sqrt\{\\varepsilon\}proves \([23](https://arxiv.org/html/2609.25659#A1.E23)\)\.
We next state the empirical\-process estimate used twice below\. We include the argument because it is also the point at which the intrinsic dimensionddenters the statistical rate\.
###### Lemma 15\(Empirical process bound for entropic transforms\)\.
Letρ\\rhobe any probability measure supported onΩ\\Omega, and letρn\\rho\_\{n\}be its empirical measure based onnniid observations\. Consider functions of the form
F\(x\)=−εlog∫Ωexp\(a\(y\)−c\(x,y\)ε\)dξ\(y\),F\(x\)=\-\\varepsilon\\log\\int\_\{\\Omega\}\\exp\\\!\\left\(\\frac\{a\(y\)\-c\(x,y\)\}\{\\varepsilon\}\\right\)d\\xi\(y\),\(27\)whereξ\\xiis an arbitrary probability measure onΩ\\Omegaanda:Ω→ℝa:\\Omega\\to\\mathbb\{R\}is bounded and measurable\. Since adding a constant toFFdoes not change\(ρn−ρ\)F\(\\rho\_\{n\}\-\\rho\)F, normalize all such functions by fixingF\(x∘\)=0F\(x\_\{\\circ\}\)=0at one reference pointx∘∈Ωx\_\{\\circ\}\\in\\Omega, and denote the resulting class byℱε\\mathcal\{F\}\_\{\\varepsilon\}\. Then
𝔼supF∈ℱε\|∫ΩFd\(ρn−ρ\)\|≤C\(1\+ε1−sd\)log\(n\+1\)n\.\\mathbb\{E\}\\sup\_\{F\\in\\mathcal\{F\}\_\{\\varepsilon\}\}\\left\|\\int\_\{\\Omega\}F\\,d\(\\rho\_\{n\}\-\\rho\)\\right\|\\leq C\\left\(1\+\\varepsilon^\{1\-s\_\{d\}\}\\right\)\\frac\{\\log\(n\+1\)\}\{\\sqrt\{n\}\}\.\(28\)The constant is uniform overρ\\rho,ξ\\xi, andaa\.
###### Proof\.
Becauseccis smooth on the compact setΩ×Ω\\Omega\\times\\Omega, all of its partial derivatives are uniformly bounded\. Differentiating \([27](https://arxiv.org/html/2609.25659#A1.E27)\) shows first that
∇F\(x\)=∫Ω∇xc\(x,y\)dωx\(y\),\\nabla F\(x\)=\\int\_\{\\Omega\}\\nabla\_\{x\}c\(x,y\)\\,d\\omega\_\{x\}\(y\),whereωx\\omega\_\{x\}is the Gibbs probability measure proportional toexp\(\(a\(y\)−c\(x,y\)\)/ε\)dξ\(y\)\\exp\(\(a\(y\)\-c\(x,y\)\)/\\varepsilon\)d\\xi\(y\)\. Repeated differentiation gives, for every integerk≥1k\\geq 1,
‖F‖Ck\(Ω\)≤Ck\(1\+ε1−k\)\.\\\|F\\\|\_\{C^\{k\}\(\\Omega\)\}\\leq C\_\{k\}\\left\(1\+\\varepsilon^\{1\-k\}\\right\)\.\(29\)Indeed, thekk\-th derivative is a finite sum of products of derivatives ofccand centered moments underωx\\omega\_\{x\}, with at mostk−1k\-1powers ofε−1\\varepsilon^\{\-1\}\. Interpolation between consecutive integer orders gives the corresponding estimate for noninteger Hölder exponents\. Thus, withsds\_\{d\}from \([19](https://arxiv.org/html/2609.25659#A1.E19)\),
supF∈ℱε‖F‖Csd\(Ω\)≤C\(1\+ε1−sd\)\.\\sup\_\{F\\in\\mathcal\{F\}\_\{\\varepsilon\}\}\\\|F\\\|\_\{C^\{s\_\{d\}\}\(\\Omega\)\}\\leq C\\left\(1\+\\varepsilon^\{1\-s\_\{d\}\}\\right\)\.\(30\)The normalizationF\(x∘\)=0F\(x\_\{\\circ\}\)=0, together with the uniform first\-derivative bound, also gives a uniform envelope‖F‖∞≤C\\\|F\\\|\_\{\\infty\}\\leq C\.
A finite smooth atlas reduces the metric entropy calculation to the standard one for Hölder balls on bounded subsets ofℝd\\mathbb\{R\}^\{d\}\[[Huang et al\., 2022](https://arxiv.org/html/2609.25659#bib.bib6)\]\. Consequently, withMε=C\(1\+ε1−sd\)M\_\{\\varepsilon\}=C\(1\+\\varepsilon^\{1\-s\_\{d\}\}\),
logN\(η,ℱε,∥⋅∥∞\)≤C\(Mεη\)d/sd\.\\log N\\\!\\left\(\\eta,\\mathcal\{F\}\_\{\\varepsilon\},\\\|\\cdot\\\|\_\{\\infty\}\\right\)\\leq C\\left\(\\frac\{M\_\{\\varepsilon\}\}\{\\eta\}\\right\)^\{d/s\_\{d\}\}\.\(31\)By definition ofsds\_\{d\}, the exponentd/sdd/s\_\{d\}is at most22\. Standard symmetrization followed by the truncated Dudley entropy integral therefore gives
𝔼supF∈ℱε\|\(ρn−ρ\)F\|≤CMεlog\(n\+1\)n\.\\mathbb\{E\}\\sup\_\{F\\in\\mathcal\{F\}\_\{\\varepsilon\}\}\|\(\\rho\_\{n\}\-\\rho\)F\|\\leq CM\_\{\\varepsilon\}\\frac\{\\log\(n\+1\)\}\{\\sqrt\{n\}\}\.This is \([28](https://arxiv.org/html/2609.25659#A1.E28)\)\.
As a direct consequence, the one\-sample entropic cost is stable under empirical replacement of one marginal\.
###### Corollary 16\(One\-sample entropic\-cost deviation\)\.
For0<ε≤10<\\varepsilon\\leq 1,
𝔼\|Sε\(μ,νn\)−Sε\(μ,ν\)\|≤Cε−sdlog\(n\+1\)n\.\\mathbb\{E\}\\left\|S\_\{\\varepsilon\}\(\\mu,\\nu\_\{n\}\)\-S\_\{\\varepsilon\}\(\\mu,\\nu\)\\right\|\\leq C\\varepsilon^\{\-s\_\{d\}\}\\frac\{\\log\(n\+1\)\}\{\\sqrt\{n\}\}\.\(32\)
###### Proof\.
Use the semi\-dual representation with the first marginalμ\\mufixed\. For a source\-side functionff, optimize the dual objective over the target potential to obtain
f\(c,ε\)\(y\):=−εlog∫Ωexp\(f\(x\)−c\(x,y\)ε\)dμ\(x\)\.f^\{\(c,\\varepsilon\)\}\(y\):=\-\\varepsilon\\log\\int\_\{\\Omega\}\\exp\\\!\\left\(\\frac\{f\(x\)\-c\(x,y\)\}\{\\varepsilon\}\\right\)d\\mu\(x\)\.Then
Sε\(μ,η\)=supf\{∫f𝑑μ\+∫f\(c,ε\)𝑑η\}\.S\_\{\\varepsilon\}\(\\mu,\\eta\)=\\sup\_\{f\}\\left\\\{\\int f\\,d\\mu\+\\int f^\{\(c,\\varepsilon\)\}\\,d\\eta\\right\\\}\.Hence
\|Sε\(μ,νn\)−Sε\(μ,ν\)\|≤supf\|∫f\(c,ε\)d\(νn−ν\)\|\.\\left\|S\_\{\\varepsilon\}\(\\mu,\\nu\_\{n\}\)\-S\_\{\\varepsilon\}\(\\mu,\\nu\)\\right\|\\leq\\sup\_\{f\}\\left\|\\int f^\{\(c,\\varepsilon\)\}\\,d\(\\nu\_\{n\}\-\\nu\)\\right\|\.The transformed functions belong, up to additive constants, to the class in Lemma[15](https://arxiv.org/html/2609.25659#Thmtheorem15)\. Therefore
𝔼\|Sε\(μ,νn\)−Sε\(μ,ν\)\|≤C\(1\+ε1−sd\)log\(n\+1\)n\.\\mathbb\{E\}\\left\|S\_\{\\varepsilon\}\(\\mu,\\nu\_\{n\}\)\-S\_\{\\varepsilon\}\(\\mu,\\nu\)\\right\|\\leq C\(1\+\\varepsilon^\{1\-s\_\{d\}\}\)\\frac\{\\log\(n\+1\)\}\{\\sqrt\{n\}\}\.Since0<ε≤10<\\varepsilon\\leq 1andsd≥1s\_\{d\}\\geq 1, the right\-hand side is bounded by the one in \([32](https://arxiv.org/html/2609.25659#A1.E32)\)\.
For completeness, we also record the modified duality inequality needed to compare the one\- and two\-sample maps\. Importantly, it is valid for any cost function and does not use Euclidean linear structure\.
###### Lemma 17\(Modified entropic duality\)\.
LetP,QP,Qbe probability measures onΩ\\Omega, letπε\\pi\_\{\\varepsilon\}be their optimal entropic plan, and letccbe any bounded measurable cost\. Then
Sε\(P,Q\)≥∫ηdπε−ε∬exp\(η\(x,y\)−c\(x,y\)ε\)𝑑P\(x\)𝑑Q\(y\)\+εS\_\{\\varepsilon\}\(P,Q\)\\geq\\int\\eta\\,d\\pi\_\{\\varepsilon\}\-\\varepsilon\\iint\\exp\\\!\\left\(\\frac\{\\eta\(x,y\)\-c\(x,y\)\}\{\\varepsilon\}\\right\)dP\(x\)dQ\(y\)\+\\varepsilon\(33\)for everyη∈L1\(πε\)\\eta\\in L^\{1\}\(\\pi\_\{\\varepsilon\}\)\.
###### Proof\.
Letγ=dπε/d\(P⊗Q\)\\gamma=d\\pi\_\{\\varepsilon\}/d\(P\\otimes Q\)\. The elementary inequality
aloga≥ab−eb\+a,a≥0,b∈ℝ,a\\log a\\geq ab\-e^\{b\}\+a,\\qquad a\\geq 0,\\ b\\in\\mathbb\{R\},applied with
b=η\(x,y\)−c\(x,y\)εb=\\frac\{\\eta\(x,y\)\-c\(x,y\)\}\{\\varepsilon\}and integrated againstP⊗QP\\otimes Qgives the claim after using
Sε\(P,Q\)=∫cdπε\+ε∫logγdπε\.S\_\{\\varepsilon\}\(P,Q\)=\\int c\\,d\\pi\_\{\\varepsilon\}\+\\varepsilon\\int\\log\\gamma\\,d\\pi\_\{\\varepsilon\}\.
###### Proposition 18\(Stability to empirical sampling of the source\)\.
LetTε,nT\_\{\\varepsilon,n\}be the entropic map fromμ\\mutoνn\\nu\_\{n\}, and letT^ε,\(n,n\)\\widehat\{T\}\_\{\\varepsilon,\(n,n\)\}be the out\-of\-sample extension in \([18](https://arxiv.org/html/2609.25659#A1.E18)\)\. Then, for0<ε≤10<\\varepsilon\\leq 1,
𝔼∫Ωdg2\(T^ε,\(n,n\)\(x\),Tε,n\(x\)\)𝑑μ\(x\)≤Cε−sdlog\(n\+1\)n\.\\mathbb\{E\}\\int\_\{\\Omega\}d\_\{g\}^\{2\}\\\!\\left\(\\widehat\{T\}\_\{\\varepsilon,\(n,n\)\}\(x\),T\_\{\\varepsilon,n\}\(x\)\\right\)d\\mu\(x\)\\leq C\\varepsilon^\{\-s\_\{d\}\}\\frac\{\\log\(n\+1\)\}\{\\sqrt\{n\}\}\.\(34\)
###### Proof\.
Let\(fε,n,gε,n\)\(f\_\{\\varepsilon,n\},g\_\{\\varepsilon,n\}\)be optimal entropic potentials for\(μ,νn\)\(\\mu,\\nu\_\{n\}\), normalized so that
∫Ωexp\(fε,n\(x\)\+gε,n\(y\)−c\(x,y\)ε\)dνn\(y\)=1\\int\_\{\\Omega\}\\exp\\\!\\left\(\\frac\{f\_\{\\varepsilon,n\}\(x\)\+g\_\{\\varepsilon,n\}\(y\)\-c\(x,y\)\}\{\\varepsilon\}\\right\)d\\nu\_\{n\}\(y\)=1for everyx∈Ωx\\in\\Omega\. Likewise, by the definition \([16](https://arxiv.org/html/2609.25659#A1.E16)\),
γ^\(x,y\):=exp\(f^\(x\)\+g^\(y\)−c\(x,y\)ε\)\\widehat\{\\gamma\}\(x,y\):=\\exp\\\!\\left\(\\frac\{\\widehat\{f\}\(x\)\+\\widehat\{g\}\(y\)\-c\(x,y\)\}\{\\varepsilon\}\\right\)satisfies
∫Ωγ^\(x,y\)dνn\(y\)=1for everyx∈Ω\.\\int\_\{\\Omega\}\\widehat\{\\gamma\}\(x,y\)\\,d\\nu\_\{n\}\(y\)=1\\qquad\\text\{for every \}x\\in\\Omega\.\(35\)On the support ofμn⊗νn\\mu\_\{n\}\\otimes\\nu\_\{n\},γ^\\widehat\{\\gamma\}is the density of the empirical optimal entropic plan\. In addition,
b^ε,\(n,n\)\(x\)=∫Ωlogx\(y\)γ^\(x,y\)dνn\(y\)\.\\widehat\{b\}\_\{\\varepsilon,\(n,n\)\}\(x\)=\\int\_\{\\Omega\}\\log\_\{x\}\(y\)\\widehat\{\\gamma\}\(x,y\)\\,d\\nu\_\{n\}\(y\)\.
Apply Lemma[17](https://arxiv.org/html/2609.25659#Thmtheorem17)to the optimal planπε,n\\pi\_\{\\varepsilon,n\}betweenμ\\muandνn\\nu\_\{n\}, with
η\(x,y\)=εχ\(x,y\)\+f^\(x\)\+g^\(y\)\.\\eta\(x,y\)=\\varepsilon\\chi\(x,y\)\+\\widehat\{f\}\(x\)\+\\widehat\{g\}\(y\)\.Using \([35](https://arxiv.org/html/2609.25659#A1.E35)\) gives, for every integrableχ\\chi,
∫χdπε,n−∬\(eχ\(x,y\)−1\)γ^\(x,y\)dμ\(x\)dνn\(y\)≤ε−1\[Sε\(μ,νn\)−∫f^𝑑μ−∫g^dνn\]\.\\begin\{split\}&\\int\\chi\\,d\\pi\_\{\\varepsilon,n\}\-\\iint\(e^\{\\chi\(x,y\)\}\-1\)\\widehat\{\\gamma\}\(x,y\)\\,d\\mu\(x\)d\\nu\_\{n\}\(y\)\\\\ &\\qquad\\leq\\varepsilon^\{\-1\}\\left\[S\_\{\\varepsilon\}\(\\mu,\\nu\_\{n\}\)\-\\int\\widehat\{f\}\\,d\\mu\-\\int\\widehat\{g\}\\,d\\nu\_\{n\}\\right\]\.\\end\{split\}\(36\)
We now control the right\-hand side\. Because\(f^,g^\)\(\\widehat\{f\},\\widehat\{g\}\)is optimal for\(μn,νn\)\(\\mu\_\{n\},\\nu\_\{n\}\), while\(fε,n,gε,n\)\(f\_\{\\varepsilon,n\},g\_\{\\varepsilon,n\}\)is an admissible dual pair for that same empirical problem,
∫f^dμn\+∫g^dνn≥∫fε,ndμn\+∫gε,ndνn\.\\int\\widehat\{f\}\\,d\\mu\_\{n\}\+\\int\\widehat\{g\}\\,d\\nu\_\{n\}\\geq\\int f\_\{\\varepsilon,n\}\\,d\\mu\_\{n\}\+\\int g\_\{\\varepsilon,n\}\\,d\\nu\_\{n\}\.Here no exponential correction remains because both source potentials are chosen as the entropic\(c,ε\)\(c,\\varepsilon\)\-transform of their corresponding target potential, so their exponential term integrates to one pointwise inxx\. Also,
Sε\(μ,νn\)=∫fε,n𝑑μ\+∫gε,ndνn\.S\_\{\\varepsilon\}\(\\mu,\\nu\_\{n\}\)=\\int f\_\{\\varepsilon,n\}\\,d\\mu\+\\int g\_\{\\varepsilon,n\}\\,d\\nu\_\{n\}\.Therefore
Sε\(μ,νn\)−∫f^𝑑μ−∫g^dνn≤∫\(fε,n−f^\)d\(μ−μn\)\.\\begin\{split\}&S\_\{\\varepsilon\}\(\\mu,\\nu\_\{n\}\)\-\\int\\widehat\{f\}\\,d\\mu\-\\int\\widehat\{g\}\\,d\\nu\_\{n\}\\\\ &\\qquad\\leq\\int\(f\_\{\\varepsilon,n\}\-\\widehat\{f\}\)\\,d\(\\mu\-\\mu\_\{n\}\)\.\\end\{split\}\(37\)Conditionally onνn\\nu\_\{n\}, the functionfε,nf\_\{\\varepsilon,n\}is independent ofμn\\mu\_\{n\}, and hence its contribution has expectation zero\. The random functionf^\\widehat\{f\}, after an irrelevant additive normalization, belongs pathwise to the classℱε\\mathcal\{F\}\_\{\\varepsilon\}from Lemma[15](https://arxiv.org/html/2609.25659#Thmtheorem15)\. Taking expectations in \([36](https://arxiv.org/html/2609.25659#A1.E36)\)–\([37](https://arxiv.org/html/2609.25659#A1.E37)\) therefore gives
𝔼supχ\{∫χdπε,n−∬\(eχ−1\)γ^dμdνn\}≤Cε−1\(1\+ε1−sd\)log\(n\+1\)n≤Cε−sdlog\(n\+1\)n\.\\begin\{split\}\\mathbb\{E\}\\sup\_\{\\chi\}\\Bigg\\\{&\\int\\chi\\,d\\pi\_\{\\varepsilon,n\}\-\\iint\(e^\{\\chi\}\-1\)\\widehat\{\\gamma\}\\,d\\mu d\\nu\_\{n\}\\Bigg\\\}\\\\ &\\leq C\\varepsilon^\{\-1\}\\left\(1\+\\varepsilon^\{1\-s\_\{d\}\}\\right\)\\frac\{\\log\(n\+1\)\}\{\\sqrt\{n\}\}\\\\ &\\leq C\\varepsilon^\{\-s\_\{d\}\}\\frac\{\\log\(n\+1\)\}\{\\sqrt\{n\}\}\.\\end\{split\}\(38\)
It remains to extract the map error from this functional inequality\. For a measurable tangent vector fieldh\(x\)∈TxMh\(x\)\\in T\_\{x\}M, set
χh\(x,y\):=⟨h\(x\),logx\(y\)−b^ε,\(n,n\)\(x\)⟩g−a‖h\(x\)‖g2,\\chi\_\{h\}\(x,y\):=\\left\\langle h\(x\),\\log\_\{x\}\(y\)\-\\widehat\{b\}\_\{\\varepsilon,\(n,n\)\}\(x\)\\right\\rangle\_\{g\}\-a\\\|h\(x\)\\\|\_\{g\}^\{2\},\(39\)wherea\>0a\>0is a sufficiently large geometric constant\. BecauseΩ\\Omegais compact andlogx\(y\)\\log\_\{x\}\(y\)is uniformly bounded onΩ×Ω\\Omega\\times\\Omega, Hoeffding’s lemma applied conditionally inxxshows thataacan be chosen so that
∫Ω\(eχh\(x,y\)−1\)γ^\(x,y\)dνn\(y\)≤0for everyx∈Ω\.\\int\_\{\\Omega\}\(e^\{\\chi\_\{h\}\(x,y\)\}\-1\)\\widehat\{\\gamma\}\(x,y\)\\,d\\nu\_\{n\}\(y\)\\leq 0\\qquad\\text\{for every \}x\\in\\Omega\.\(40\)On the other hand, disintegratingπε,n\\pi\_\{\\varepsilon,n\}gives
∫χhdπε,n=∫Ω\[⟨h\(x\),bε,n\(x\)−b^ε,\(n,n\)\(x\)⟩g−a‖h\(x\)‖g2\]𝑑μ\(x\)\.\\int\\chi\_\{h\}\\,d\\pi\_\{\\varepsilon,n\}=\\int\_\{\\Omega\}\\left\[\\left\\langle h\(x\),b\_\{\\varepsilon,n\}\(x\)\-\\widehat\{b\}\_\{\\varepsilon,\(n,n\)\}\(x\)\\right\\rangle\_\{g\}\-a\\\|h\(x\)\\\|\_\{g\}^\{2\}\\right\]d\\mu\(x\)\.Taking the pointwise supremum overhhyields
suph∫χhdπε,n=14a∫Ω‖bε,n\(x\)−b^ε,\(n,n\)\(x\)‖g2𝑑μ\(x\)\.\\sup\_\{h\}\\int\\chi\_\{h\}\\,d\\pi\_\{\\varepsilon,n\}=\\frac\{1\}\{4a\}\\int\_\{\\Omega\}\\\|b\_\{\\varepsilon,n\}\(x\)\-\\widehat\{b\}\_\{\\varepsilon,\(n,n\)\}\(x\)\\\|\_\{g\}^\{2\}d\\mu\(x\)\.\(41\)Combining \([38](https://arxiv.org/html/2609.25659#A1.E38)\)–\([41](https://arxiv.org/html/2609.25659#A1.E41)\) and finally applying the Lipschitz bound for the exponential map from Assumption[9](https://arxiv.org/html/2609.25659#Thmtheorem9)proves \([34](https://arxiv.org/html/2609.25659#A1.E34)\)\.
#### A\.7\.3Proof of Theorem[12](https://arxiv.org/html/2609.25659#Thmtheorem12)
We first bound the one\-sample error\. Applying Lemma[13](https://arxiv.org/html/2609.25659#Thmtheorem13)toπε,n\\pi\_\{\\varepsilon,n\}gives
∫dg2\(Tε,n\(x\),T0\(x\)\)𝑑μ\(x\)≤C∫Dx\(y\)dπε,n\(x,y\)\.\\int d\_\{g\}^\{2\}\(T\_\{\\varepsilon,n\}\(x\),T\_\{0\}\(x\)\)\\,d\\mu\(x\)\\leq C\\int D\_\{x\}\(y\)\\,d\\pi\_\{\\varepsilon,n\}\(x,y\)\.\(42\)Sinceπε,n\\pi\_\{\\varepsilon,n\}is optimal for the entropic problem betweenμ\\muandνn\\nu\_\{n\},
∫Dx\(y\)dπε,n\(x,y\)\\displaystyle\\int D\_\{x\}\(y\)\\,d\\pi\_\{\\varepsilon,n\}\(x,y\)≤Sε\(μ,νn\)−∫φ0𝑑μ−∫φ0cdνn\.\\displaystyle\\leq S\_\{\\varepsilon\}\(\\mu,\\nu\_\{n\}\)\-\\int\\varphi\_\{0\}\\,d\\mu\-\\int\\varphi\_\{0\}^\{c\}\\,d\\nu\_\{n\}\.\(43\)Indeed, the difference between the right\-hand side and the left\-hand side is exactlyεDKL\(πε,n∥μ⊗νn\)\\varepsilon D\_\{\\mathrm\{KL\}\}\(\\pi\_\{\\varepsilon,n\}\\\|\\mu\\otimes\\nu\_\{n\}\), which is nonnegative\. Taking expectations and using𝔼νn=ν\\mathbb\{E\}\\nu\_\{n\}=\\nutogether with Kantorovich duality yields
𝔼∫Dx\(y\)dπε,n\(x,y\)\\displaystyle\\mathbb\{E\}\\int D\_\{x\}\(y\)\\,d\\pi\_\{\\varepsilon,n\}\(x,y\)≤𝔼\[Sε\(μ,νn\)−Sε\(μ,ν\)\]\\displaystyle\\leq\\mathbb\{E\}\\left\[S\_\{\\varepsilon\}\(\\mu,\\nu\_\{n\}\)\-S\_\{\\varepsilon\}\(\\mu,\\nu\)\\right\]\(44\)\+Sε\(μ,ν\)−12W22\(μ,ν\)\\displaystyle\\quad\+S\_\{\\varepsilon\}\(\\mu,\\nu\)\-\\frac\{1\}\{2\}W\_\{2\}^\{2\}\(\\mu,\\nu\)\(45\)≤Cε−sdlog\(n\+1\)n\+Cεlog\(eε\),\\displaystyle\\leq C\\varepsilon^\{\-s\_\{d\}\}\\frac\{\\log\(n\+1\)\}\{\\sqrt\{n\}\}\+C\\varepsilon\\log\\\!\\left\(\\frac\{e\}\{\\varepsilon\}\\right\),\(46\)where the last line follows from Corollary[16](https://arxiv.org/html/2609.25659#Thmtheorem16)and Lemma[14](https://arxiv.org/html/2609.25659#Thmtheorem14)\. Combining \([42](https://arxiv.org/html/2609.25659#A1.E42)\) and \([46](https://arxiv.org/html/2609.25659#A1.E46)\),
𝔼∫dg2\(Tε,n\(x\),T0\(x\)\)𝑑μ\(x\)≤C\[εlog\(eε\)\+ε−sdlog\(n\+1\)n\]\.\\mathbb\{E\}\\int d\_\{g\}^\{2\}\(T\_\{\\varepsilon,n\}\(x\),T\_\{0\}\(x\)\)\\,d\\mu\(x\)\\leq C\\left\[\\varepsilon\\log\\\!\\left\(\\frac\{e\}\{\\varepsilon\}\\right\)\+\\varepsilon^\{\-s\_\{d\}\}\\frac\{\\log\(n\+1\)\}\{\\sqrt\{n\}\}\\right\]\.\(47\)
Finally, the squared triangle inequality and Proposition[18](https://arxiv.org/html/2609.25659#Thmtheorem18)give
𝔼∫dg2\(T^ε,\(n,n\)\(x\),T0\(x\)\)𝑑μ\(x\)\\displaystyle\\mathbb\{E\}\\int d\_\{g\}^\{2\}\(\\widehat\{T\}\_\{\\varepsilon,\(n,n\)\}\(x\),T\_\{0\}\(x\)\)\\,d\\mu\(x\)≤2𝔼∫dg2\(T^ε,\(n,n\)\(x\),Tε,n\(x\)\)𝑑μ\(x\)\+2𝔼∫dg2\(Tε,n\(x\),T0\(x\)\)𝑑μ\(x\)\\displaystyle\\qquad\\leq 2\\mathbb\{E\}\\int d\_\{g\}^\{2\}\(\\widehat\{T\}\_\{\\varepsilon,\(n,n\)\}\(x\),T\_\{\\varepsilon,n\}\(x\)\)\\,d\\mu\(x\)\+2\\mathbb\{E\}\\int d\_\{g\}^\{2\}\(T\_\{\\varepsilon,n\}\(x\),T\_\{0\}\(x\)\)\\,d\\mu\(x\)≤C\[εlog\(eε\)\+ε−sdlog\(n\+1\)n\],\\displaystyle\\qquad\\leq C\\left\[\\varepsilon\\log\\\!\\left\(\\frac\{e\}\{\\varepsilon\}\\right\)\+\\varepsilon^\{\-s\_\{d\}\}\\frac\{\\log\(n\+1\)\}\{\\sqrt\{n\}\}\\right\],which proves Theorem[12](https://arxiv.org/html/2609.25659#Thmtheorem12)\.
## Appendix BA short introduction to flow matching
In this section, we provide a very concise introduction to flow matching and highlight the theoretical steps needed to show the validity of the procedure on the space of probability distributions\. We refer the reader to[Lipman et al\. \[2024\]](https://arxiv.org/html/2609.25659#bib.bib25)for a complete introduction to flow matching\.
#### B\.0\.1Vector fields generating probability paths
At its core, flow matching aims to find a tractable way to generate samples from a target distributionν\\nufrom a source distributionμ\\mu\. One way to do so is by learning a vector fieldutu\_\{t\}that*generates*a probability path with the right boundary conditions\. By*generation*, we understand the Lagrangian interpretation\. That is,utu\_\{t\}generates a probability pathμt\\mu\_\{t\}iff, for allt∈\[0,1\]t\\in\[0,1\]
∂tϕt\(x\)=ut\(ϕt\(x\)\)Xt=ϕt\(X0\)∼μt\.\\begin\{split\}\\partial\_\{t\}\\phi\_\{t\}\(x\)&=u\_\{t\}\(\\phi\_\{t\}\(x\)\)\\\\ X\_\{t\}&=\\phi\_\{t\}\(X\_\{0\}\)\\sim\\mu\_\{t\}\.\\end\{split\}
With such autu\_\{t\}, one can then sample fromμ\\muand integrate over time to generate samples from distributionsμt\\mu\_\{t\}\. An important result to verify thatutu\_\{t\}generates a given probability path is the mass conservation formula:
utgeneratesμt⇔∂tμt\(x\)\+∇⋅\(μtut\)\(x\)=0u\_\{t\}\\text\{ generates \}\\mu\_\{t\}\\Leftrightarrow\\partial\_\{t\}\\mu\_\{t\}\(x\)\+\\nabla\\cdot\(\\mu\_\{t\}u\_\{t\}\)\(x\)=0
That is,\(ut,μt\)\(u\_\{t\},\\mu\_\{t\}\)solving the continuity equation is equivalent toutu\_\{t\}generatingμt\\mu\_\{t\}\. However, directly using this relation is impractical as we typically don’t have information about the target distributionν\\nu\.
#### B\.0\.2Conditional vector fields and probability paths
A more promising strategy is to construct probability paths using conditional probability paths:
μt\(x\)=∫μt\|z\(x\|z\)𝑑π\(z\)\\mu\_\{t\}\(x\)=\\int\\mu\_\{t\|z\}\(x\|z\)d\\pi\(z\)
Of course, one needs to design the conditional probability paths and marginal distribution for the conditioning variable such that the boundaries coincide with the constraints \(μ\\muandν\\nu\)\.
Ifμ\\muis known \(*e\.g\.*a standard normal\), one can choose
dπ\(z\)=dν\(x1\),pt\|x1=𝒩\(tx1,\(1−t\)2\)d\\pi\(z\)=d\\nu\(x\_\{1\}\),\\quad p\_\{t\|x\_\{1\}\}=\\mathcal\{N\}\(tx\_\{1\},\(1\-t\)^\{2\}\)
We verify that in that case:
∫𝒩\(x,0,1\)dν\(x1\)=𝒩\(0,1\)∫δx1dν\(x1\)=ν\\begin\{split\}\\int\\mathcal\{N\}\(x;0,1\)d\\nu\(x\_\{1\}\)&=\\mathcal\{N\}\(0,1\)\\\\ \\int\\delta\_\{x\_\{1\}\}d\\nu\(x\_\{1\}\)&=\\nu\\end\{split\}
Whenμ\\muis an arbitrary \(unknown\) distribution, one can instead choose
dπ\(z\)=dπ\(x0,x1\),∫dπ\(x0,x1\)dx1=dμ\(x0\),∫dπ\(x0,x1\)dx0=dν\(x1\),pt\|x0,x1=δ\(1−t\)x0\+tx1\(x\)\\begin\{split\}&d\\pi\(z\)=d\\pi\(x\_\{0\},x\_\{1\}\),\\quad\\int d\\pi\(x\_\{0\},x\_\{1\}\)dx\_\{1\}=d\\mu\(x\_\{0\}\),\\\\ &\\int d\\pi\(x\_\{0\},x\_\{1\}\)dx\_\{0\}=d\\nu\(x\_\{1\}\),\\quad p\_\{t\|x\_\{0\},x\_\{1\}\}=\\delta\_\{\(1\-t\)x\_\{0\}\+tx\_\{1\}\}\(x\)\\end\{split\}
In which case, we verify again that the boundary conditions are respected
∬δx0dπ\(x0,x1\)=∫δx0dμ\(x0\)=μ\(x\)∬δx1dπ\(x0,x1\)=∫δx1dν\(x1\)=ν\(x\)\\begin\{split\}\\iint\\delta\_\{x\_\{0\}\}d\\pi\(x\_\{0\},x\_\{1\}\)&=\\int\\delta\_\{x\_\{0\}\}d\\mu\(x\_\{0\}\)=\\mu\(x\)\\\\ \\iint\\delta\_\{x\_\{1\}\}d\\pi\(x\_\{0\},x\_\{1\}\)&=\\int\\delta\_\{x\_\{1\}\}d\\nu\(x\_\{1\}\)=\\nu\(x\)\\end\{split\}
Given the simple form of the conditional probability paths, one can easily derive corresponding vector fields that generate the conditional distributions,vt\(x∣z\)v\_\{t\}\(x\\mid z\)\.
The first crucial realization of flow matching is that the marginal conditional vector fieldvtv\_\{t\}also generates the marginal probability pathμt\\mu\_\{t\}
∂tμt\(x\)\+∇⋅\(μtvt\)\(x\)=0\\partial\_\{t\}\\mu\_\{t\}\(x\)\+\\nabla\\cdot\(\\mu\_\{t\}v\_\{t\}\)\(x\)=0
where
vt\(x\)=∫vt\(x∣z\)μt\|z\(x\|z\)π\(z\)μt\(x\)vt\(x\)=𝔼\[vt\(Xt\|Z\)\|Xt=x\]\\begin\{split\}v\_\{t\}\(x\)&=\\int v\_\{t\}\(x\\mid z\)\\frac\{\\mu\_\{t\|z\}\(x\|z\)\\pi\(z\)\}\{\\mu\_\{t\}\(x\)\}\\\\ v\_\{t\}\(x\)&=\\mathbb\{E\}\[v\_\{t\}\(X\_\{t\}\|Z\)\|X\_\{t\}=x\]\\end\{split\}
That is, from simple conditional vector fields that generate simple conditional probability paths, we can construct the marginal vector field that generates the probability path of interest\. Now, of course, the equation above is not tractable asP\(Z∣Xt\)P\(Z\\mid X\_\{t\}\)is generally not available\. The second realization of flow matching is that learningvtθ\(x\)v\_\{t\}^\{\\theta\}\(x\)by regressing it to conditional vector fields has the same minimizer than regressingvtθ\(x\)v\_\{t\}^\{\\theta\}\(x\)against its true value, as shown next\.
#### B\.0\.3Flow matching loss
Our intended flow matching loss is to regress a learnable vector field to the marginal vector field:
ℒFM=𝔼t,Xt\[∥vt\(Xt\)−vtθ\(Xt\)∥2\]\\mathcal\{L\}\_\{FM\}=\\mathbb\{E\}\_\{t,X\_\{t\}\}\[\\lVert v\_\{t\}\(X\_\{t\}\)\-v\_\{t\}^\{\\theta\}\(X\_\{t\}\)\\rVert^\{2\}\]
Instead we can use a loss that involves only tractable distributions:
ℒCFM=𝔼t,Z∼π\(Z\),Xt∼μt\(Xt∣Z\)\[∥vt\(Xt\|Z\)−vtθ\(Xt\)∥2\]\\mathcal\{L\}\_\{CFM\}=\\mathbb\{E\}\_\{t,Z\\sim\\pi\(Z\),X\_\{t\}\\sim\\mu\_\{t\}\(X\_\{t\}\\mid Z\)\}\[\\lVert v\_\{t\}\(X\_\{t\}\|Z\)\-v\_\{t\}^\{\\theta\}\(X\_\{t\}\)\\rVert^\{2\}\]
Crucially, one can show that the gradients of both losses are identical\.
∇θℒCFM=∇θℒFM\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{CFM\}=\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{FM\}
#### B\.0\.4Roadmap to lifting to empirical distributions
In this work, we want to leverage the same strategy for training our generative model but operating on the space of probability distributions defined on a Riemannian manifoldℳ\\mathcal\{M\}:𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\), which is an infinite dimensional space\. For this to be possible, we need to show the three following statements apply on that space:
- •\(C1\) The existence of an equivalent of the mass conservation formula on𝒫2\(𝒫2\(ℳ\)\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\)\(Section[C\.1](https://arxiv.org/html/2609.25659#A3.SS1)\)
- •\(C2\) Showing that the marginal vector fieldvtv\_\{t\}generatesμt\\mu\_\{t\}\(Section[C\.2](https://arxiv.org/html/2609.25659#A3.SS2)\)
- •\(C3\) The gradient of the conditional flow matching loss is identical to the gradient of the flow matching loss \(Section[C\.3](https://arxiv.org/html/2609.25659#A3.SS3)\)
## Appendix CLifting Flow Matching to𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)
#### Setup and Definitions
We work on the Wasserstein space𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\), the space of probability measures on a manifoldℳ\\mathcal\{M\}with finite second moments\. We endow𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)with the Borelσ−\\sigma\-algebra generated by theW2W\_\{2\}metric \(with the underlying geodesic distance\)\.
### C\.1Continuity equation and Lagrangian flows on𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)
##### Objectives of this section
In the following, we show that if\(ℙt,Vt\)\(\\mathbb\{P\}\_\{t\},V\_\{t\}\)solve the continuity equation on𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\(*i\.e\.*with test functionals on𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\), thenVtV\_\{t\}generatesℙt\\mathbb\{P\}\_\{t\}in the Lagrangian senseℙt=\(Φt\)\#ℙ0\\mathbb\{P\}\_\{t\}=\(\\Phi\_\{t\}\)\_\{\\\#\}\\mathbb\{P\}\_\{0\}\. That is, conceptually, one can generate the probability pathℙt\\mathbb\{P\}\_\{t\}by flowing individual elements from00tottaccording to the vector fieldVtV\_\{t\}, which underlies the generation procedure used in flow matching\.
#### C\.1\.1Definitions
##### Tangent space and path space of continuous curves
For eachμ∈𝒫2\(ℳ\)\\mu\\in\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\), we identify the tangent spaceTμ𝒫2\(ℳ\)T\_\{\\mu\}\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)with the closure of the gradient vector fields inL2\(μ,Tℳ\)L^\{2\}\(\\mu;T\\mathcal\{M\}\)and endow it with the norm∥v∥L2\(μ\)2=∫ℳ∥v\(x\)∥g2𝑑μ\(x\)\\lVert v\\rVert^\{2\}\_\{L^\{2\}\(\\mu\)\}=\\int\_\{\\mathcal\{M\}\}\\lVert v\(x\)\\rVert\_\{g\}^\{2\}d\\mu\(x\)\. We further writeL2\(ℙt,T𝒫2\(ℳ\)\)L^\{2\}\(\\mathbb\{P\}\_\{t\},T\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\)the space of functionsVt:𝒫2\(ℳ\)→T𝒫2\(ℳ\)V\_\{t\}:\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\\rightarrow T\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)such that∫𝒫2\(ℳ\)∥Vt\(μ\)∥L2\(μ\)2\(μ\)dℙt<∞\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\lVert V\_\{t\}\(\\mu\)\\rVert^\{2\}\_\{L^\{2\}\(\\mu\)\}\(\\mu\)d\\mathbb\{P\}\_\{t\}<\\infty\. We writeΓT:=C\(\[0,T\],𝒫2\(ℳ\)\)\\Gamma\_\{T\}:=C\(\[0,T\];\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\)the path space of continuous curvesγ\\gammafrom\[0,T\]\[0,T\]to𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\.
We define the elevation mapsete\_\{t\}as
et:\(γ\)∈ΓT→γ\(t\)∈𝒫2\(ℳ\)fort∈\[0,T\]e\_\{t\}:\(\\gamma\)\\in\\Gamma\_\{T\}\\rightarrow\\gamma\(t\)\\in\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\\quad\\text\{for \}t\\in\[0,T\]
##### Cylinder Functions
###### Definition 19\.
A functionalℱ:𝒫2\(ℳ\)→ℝ\\mathcal\{F\}:\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\\to\\mathbb\{R\}is called a cylinder function if there exists an integerk≥1k\\geq 1, a smooth function with compact supportF∈Cc∞\(ℝk\)F\\in C\_\{c\}^\{\\infty\}\(\\mathbb\{R\}^\{k\}\), and a set of smooth functions with compact support\{V1,…,Vk\}⊂Cc∞\(ℳ\)\\\{V\_\{1\},\\dots,V\_\{k\}\\\}\\subset C\_\{c\}^\{\\infty\}\(\\mathcal\{M\}\), such that for any measureμ∈𝒫2\(ℳ\)\\mu\\in\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\):
ℱ\(μ\)=F\(∫ℳV1𝑑μ,…,∫ℳVk𝑑μ\)\\mathcal\{F\}\(\\mu\)=F\\left\(\\int\_\{\\mathcal\{M\}\}V\_\{1\}d\\mu,\\dots,\\int\_\{\\mathcal\{M\}\}V\_\{k\}d\\mu\\right\)
This definition can be extended to time\-dependent functionalsφt\(μ\)=φ\(t,μ\)\\varphi\_\{t\}\(\\mu\)=\\varphi\(t,\\mu\), whereφ\(t,μ\)=F\(t,∫V1𝑑μ,…,∫Vk𝑑μ\)\\varphi\(t,\\mu\)=F\(t,\\int V\_\{1\}d\\mu,\\dots,\\int V\_\{k\}d\\mu\)for someF∈Cc∞\(I×ℝk\)F\\in C\_\{c\}^\{\\infty\}\(I\\times\\mathbb\{R\}^\{k\}\)\.
In this text, we will assume that time\-dependent cylinder functions have a compact support in time\(0,T\)\(0,T\), such thatφ0\(x\)=φT\(x\)=0\\varphi\_\{0\}\(x\)=\\varphi\_\{T\}\(x\)=0for allxx\.
Using the chain rule, the gradient of a cylinder functionalℱ\\mathcal\{F\}, the Wasserstein gradient, at a measureμ\\mucan be computed\.
###### Definition 20\.
The Wasserstein gradient of a cylinder functionℱ\\mathcal\{F\}, denoted∇𝒲ℱ\(μ\)\\nabla\_\{\\mathcal\{W\}\}\\mathcal\{F\}\(\\mu\), is the vector field onℳ\\mathcal\{M\}:
∇𝒲ℱ\(μ\)=∑i=1k∂F∂xi\(∫V1dμ,…,∫Vkdμ\)∇Vi\\nabla\_\{\\mathcal\{W\}\}\\mathcal\{F\}\(\\mu\)=\\sum\_\{i=1\}^\{k\}\\frac\{\\partial F\}\{\\partial x\_\{i\}\}\\left\(\\int V\_\{1\}d\\mu,\\dots,\\int V\_\{k\}d\\mu\\right\)\\nabla V\_\{i\}
Here,∂F∂xi\\frac\{\\partial F\}\{\\partial x\_\{i\}\}is the partial derivative ofFFwith respect to itsii\-th argument, and∇Vi\\nabla V\_\{i\}is the gradient of the functionViV\_\{i\}on the manifoldℳ\\mathcal\{M\}\.
#### C\.1\.2Continuity equation and superposition principle on𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)
We give precise sufficient hypotheses for the Lagrangian interpretation used in the main text\. The argument adapts the superposition principle for random measures of[Pinzi and Savaré \[2025, Theorem 1\.2\]](https://arxiv.org/html/2609.25659#bib.bib41)through a smooth embedding\.
Recall that\(ℙt,Vt\)\(\\mathbb\{P\}\_\{t\},V\_\{t\}\)satisfies the weak continuity equation on𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)if
∫0T∫𝒫2\(ℳ\)\[∂tφt\(μ\)\\displaystyle\\int\_\{0\}^\{T\}\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\Bigl\[\\partial\_\{t\}\\varphi\_\{t\}\(\\mu\)\(Weak CE\)\+⟨∇𝒲φt\(μ\),Vt\(μ\)⟩L2\(μ\)\]dℙt\(μ\)dt=0,\\displaystyle\+\\bigl\\langle\\nabla\_\{\\mathcal\{W\}\}\\varphi\_\{t\}\(\\mu\),V\_\{t\}\(\\mu\)\\bigr\\rangle\_\{L^\{2\}\(\\mu\)\}\\Bigr\]\\,d\\mathbb\{P\}\_\{t\}\(\\mu\)\\,dt=0,for every smooth cylinder functionalφt\\varphi\_\{t\}with compact support in time in\(0,T\)\(0,T\)\.
###### Theorem 21\(Weak continuity equation and superposition in𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\)\.
Let\(ℳ,g\)\(\\mathcal\{M\},g\)be a connected, complete, smooth Riemannian manifold without boundary, and let\(ℙt\)t∈\[0,T\]\(\\mathbb\{P\}\_\{t\}\)\_\{t\\in\[0,T\]\}be an absolutely continuous curve in𝒫2\(𝒫2\(ℳ\)\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\)\. Assume there is a compact setK⊂ℳK\\subset\\mathcal\{M\}such thatℙt\(𝒫\(K\)\)=1\\mathbb\{P\}\_\{t\}\(\\mathcal\{P\}\(K\)\)=1for everytt, where𝒫\(K\)\\mathcal\{P\}\(K\)denotes the probability measures supported inKK\. LetVt\(μ\)∈Tμ𝒫2\(ℳ\)V\_\{t\}\(\\mu\)\\in T\_\{\\mu\}\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)admit a jointly Borel representativeb\(t,x,μ\)=Vt\(μ\)\(x\)∈Txℳb\(t,x,\\mu\)=V\_\{t\}\(\\mu\)\(x\)\\in T\_\{x\}\\mathcal\{M\}, and assume
∫0T∫𝒫2\(ℳ\)‖Vt\(μ\)‖L2\(μ\)2dℙt\(μ\)𝑑t<∞\.\\int\_\{0\}^\{T\}\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\\|V\_\{t\}\(\\mu\)\\\|\_\{L^\{2\}\(\\mu\)\}^\{2\}\\,d\\mathbb\{P\}\_\{t\}\(\\mu\)\\,dt<\\infty\.Suppose\(ℙt,Vt\)\(\\mathbb\{P\}\_\{t\},V\_\{t\}\)satisfies the weak continuity equation \([Weak CE](https://arxiv.org/html/2609.25659#A3.Ex82)\)\. Then:
1. 1\.There exists a probability measureΓ\\GammaonΓT=C\(\[0,T\],𝒫2\(ℳ\)\)\\Gamma\_\{T\}=C\(\[0,T\];\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\), concentrated onW2W\_\{2\}\-absolutely continuous curves taking values in𝒫\(K\)\\mathcal\{P\}\(K\), such that \(et\)\#Γ=ℙt,t∈\[0,T\]\.\(e\_\{t\}\)\_\{\\\#\}\\Gamma=\\mathbb\{P\}\_\{t\},\\qquad t\\in\[0,T\]\.
2. 2\.ForΓ\\Gamma\-almost every trajectoryγ\\gamma, the measure\-valued continuity equation ∂tγ\(t\)\+divℳ\(γ\(t\)Vt\(γ\(t\)\)\)=0\\partial\_\{t\}\\gamma\(t\)\+\\operatorname\{div\}\_\{\\mathcal\{M\}\}\\bigl\(\\gamma\(t\)V\_\{t\}\(\\gamma\(t\)\)\\bigr\)=0\(Measure CE\)holds weakly onℳ\\mathcal\{M\}\. Consequently, for each smooth time\-dependent cylinder functionalφt\\varphi\_\{t\}, alongΓ\\Gamma\-almost every trajectory and for almost everytt, ddtφt\(γ\(t\)\)=∂tφt\(γ\(t\)\)\+⟨∇𝒲φt\(γ\(t\)\),Vt\(γ\(t\)\)⟩L2\(γ\(t\)\)\.\\frac\{d\}\{dt\}\\varphi\_\{t\}\(\\gamma\(t\)\)=\\partial\_\{t\}\\varphi\_\{t\}\(\\gamma\(t\)\)\+\\bigl\\langle\\nabla\_\{\\mathcal\{W\}\}\\varphi\_\{t\}\(\\gamma\(t\)\),V\_\{t\}\(\\gamma\(t\)\)\\bigr\\rangle\_\{L^\{2\}\(\\gamma\(t\)\)\}\.This is the weak interpretation ofγ˙\(t\)=Vt\(γ\(t\)\)\\dot\{\\gamma\}\(t\)=V\_\{t\}\(\\gamma\(t\)\)used here\.
3. 3\.Assume additionally that, forℙ0\\mathbb\{P\}\_\{0\}\-almost everyμ\\mu, equation \([Measure CE](https://arxiv.org/html/2609.25659#A3.Ex85)\) has at most oneW2W\_\{2\}\-absolutely continuous solution in𝒫\(K\)\\mathcal\{P\}\(K\)withγ\(0\)=μ\\gamma\(0\)=\\mu\. Then, forℙ0\\mathbb\{P\}\_\{0\}\-almost everyμ\\mu, there exists a uniqueW2W\_\{2\}\-absolutely continuous curveγμ:\[0,T\]→𝒫\(K\)⊂𝒫2\(ℳ\)\\gamma\_\{\\mu\}:\[0,T\]\\to\\mathcal\{P\}\(K\)\\subset\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)solving γ˙μ\(t\)\\displaystyle\\dot\{\\gamma\}\_\{\\mu\}\(t\)=Vt\(γμ\(t\)\),for a\.e\.t∈\[0,T\],\\displaystyle=V\_\{t\}\(\\gamma\_\{\\mu\}\(t\)\),\\qquad\\text\{for a\.e\. \}t\\in\[0,T\],γμ\(0\)\\displaystyle\\gamma\_\{\\mu\}\(0\)=μ,\\displaystyle=\\mu,where the evolution equation is understood in the weak sense of \([Measure CE](https://arxiv.org/html/2609.25659#A3.Ex85)\), equivalently through the cylinder chain rule above\. The mapΦt\(μ\):=γμ\(t\)\\Phi\_\{t\}\(\\mu\):=\\gamma\_\{\\mu\}\(t\)defines a measurable flow, up toℙ0\\mathbb\{P\}\_\{0\}\-null sets, such that ℙt=\(Φt\)\#ℙ0,t∈\[0,T\]\.\\mathbb\{P\}\_\{t\}=\(\\Phi\_\{t\}\)\_\{\\\#\}\\mathbb\{P\}\_\{0\},\\qquad t\\in\[0,T\]\.ThusVtV\_\{t\}generates the probability pathℙt\\mathbb\{P\}\_\{t\}by evolving each initial measure along its trajectoryγμ\\gamma\_\{\\mu\}\.
###### Proof\.
Choose a smooth embeddingȷ:ℳ→ℝD\\jmath:\\mathcal\{M\}\\to\\mathbb\{R\}^\{D\}and writeJ\(μ\)=ȷ\#μJ\(\\mu\)=\\jmath\_\{\\\#\}\\muforμ∈𝒫\(K\)\\mu\\in\\mathcal\{P\}\(K\)\. Setℙ~t=J\#ℙt\\widetilde\{\\mathbb\{P\}\}\_\{t\}=J\_\{\\\#\}\\mathbb\{P\}\_\{t\}and define
b~\(t,ȷ\(x\),J\(μ\)\)=dȷxb\(t,x,μ\),x∈K,μ∈𝒫\(K\),\\widetilde\{b\}\(t,\\jmath\(x\),J\(\\mu\)\)=d\\jmath\_\{x\}\\,b\(t,x,\\mu\),\\qquad x\\in K,\\quad\\mu\\in\\mathcal\{P\}\(K\),extendingb~\\widetilde\{b\}by zero elsewhere\. On the compact setKK, the embedding has bounded differential and its inverse has bounded metric distortion\. Pulling back cylinder tests, using a smooth cutoff equal to one nearKK, transfers the weak continuity equation to\(ℙ~t,b~\)\(\\widetilde\{\\mathbb\{P\}\}\_\{t\},\\widetilde\{b\}\)on𝒫\(ℝD\)\\mathcal\{P\}\(\\mathbb\{R\}^\{D\}\)\. The integrated squared\-velocity bound is preserved up to a constant and implies the integrability required by[Pinzi and Savaré \[2025, Theorem 1\.2\]](https://arxiv.org/html/2609.25659#bib.bib41)\.
That theorem gives a measureΓ~\\widetilde\{\\Gamma\}on measure\-valued trajectories with marginalsℙ~t\\widetilde\{\\mathbb\{P\}\}\_\{t\}, concentrated on solutions of the continuity equation driven byb~\\widetilde\{b\}\. These trajectories remain in𝒫\(ȷ\(K\)\)\\mathcal\{P\}\(\\jmath\(K\)\): the marginal identities imply this simultaneously at rational times, and continuity and closedness extend it to every time\. We can therefore pull them back throughJ−1J^\{\-1\}to obtainΓ\\Gammawith the desired marginals\. Extension of smooth tests from the embedded manifold and the chain rule give \([Measure CE](https://arxiv.org/html/2609.25659#A3.Ex85)\)\[[Pinzi, 2026](https://arxiv.org/html/2609.25659#bib.bib40)\]\.
By Fubini and the marginal identities, the assumed energy is finite along almost every represented trajectory\. The velocity\-energy estimate for the Euclidean continuity equation givesW2W\_\{2\}absolute continuity there; metric comparison onKKtransfers this to the intrinsicW2W\_\{2\}metric\. Applying the cylinder chain rule to \([Measure CE](https://arxiv.org/html/2609.25659#A3.Ex85)\) proves the second claim\.
##### Disintegration and the deterministic flow\.
For the third claim, consider the joint law of the initial measure and the entire trajectory:
Γ^:=\(e0,id\)\#Γ∈𝒫\(𝒫2\(ℳ\)×ΓT\)\.\\widehat\{\\Gamma\}:=\(e\_\{0\},\\operatorname\{id\}\)\_\{\\\#\}\\Gamma\\in\\mathcal\{P\}\(\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\\times\\Gamma\_\{T\}\)\.Its first marginal is\(e0\)\#Γ=ℙ0\(e\_\{0\}\)\_\{\\\#\}\\Gamma=\\mathbb\{P\}\_\{0\}, and its second marginal isΓ\\Gamma\. Since𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)andΓT\\Gamma\_\{T\}are Polish spaces, disintegration with respect to the first marginal gives a measurable family of conditional path measures\{Γμ\}μ∈𝒫2\(ℳ\)\\\{\\Gamma\_\{\\mu\}\\\}\_\{\\mu\\in\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}such that
Γ^=∫𝒫2\(ℳ\)δμ⊗Γμdℙ0\(μ\)\.\\widehat\{\\Gamma\}=\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\delta\_\{\\mu\}\\otimes\\Gamma\_\{\\mu\}\\,d\\mathbb\{P\}\_\{0\}\(\\mu\)\.HereΓμ\\Gamma\_\{\\mu\}is the conditional law of the trajectory given its initial valueμ\\mu\. In particular, forℙ0\\mathbb\{P\}\_\{0\}\-almost everyμ\\mu,
Γμ\(\{γ∈ΓT:γ\(0\)=μ\}\)=1\.\\Gamma\_\{\\mu\}\\bigl\(\\\{\\gamma\\in\\Gamma\_\{T\}:\\gamma\(0\)=\\mu\\\}\\bigr\)=1\.The concentration properties ofΓ\\Gammaalso hold underΓμ\\Gamma\_\{\\mu\}forℙ0\\mathbb\{P\}\_\{0\}\-almost everyμ\\mu: its trajectories areW2W\_\{2\}\-absolutely continuous, remain in𝒫\(K\)\\mathcal\{P\}\(K\), and solve \([Measure CE](https://arxiv.org/html/2609.25659#A3.Ex85)\)\.
Under the additional uniqueness assumption, there is at most one such trajectory starting fromμ\\mu\. SinceΓμ\\Gamma\_\{\\mu\}is a probability measure concentrated on these trajectories, there is exactly one, denotedγμ\\gamma\_\{\\mu\}, and
Γμ=δγμforℙ0\-almost everyμ\.\\Gamma\_\{\\mu\}=\\delta\_\{\\gamma\_\{\\mu\}\}\\qquad\\text\{for \}\\mathbb\{P\}\_\{0\}\\text\{\-almost every \}\\mu\.The measurability of the conditional measures therefore gives a measurable trajectory mapμ↦γμ\\mu\\mapsto\\gamma\_\{\\mu\}up to null sets\. DefineΦt\(μ\):=γμ\(t\)\\Phi\_\{t\}\(\\mu\):=\\gamma\_\{\\mu\}\(t\), so thatΦ0\(μ\)=μ\\Phi\_\{0\}\(\\mu\)=\\muforℙ0\\mathbb\{P\}\_\{0\}\-almost everyμ\\mu\.
To identify the law at timett, take any bounded Borel functionalφ:𝒫2\(ℳ\)→ℝ\\varphi:\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\\to\\mathbb\{R\}\. The marginal identity and disintegration give
∫𝒫2\(ℳ\)φ\(ν\)dℙt\(ν\)\\displaystyle\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\varphi\(\\nu\)\\,d\\mathbb\{P\}\_\{t\}\(\\nu\)=∫ΓTφ\(γ\(t\)\)𝑑Γ\(γ\)\\displaystyle=\\int\_\{\\Gamma\_\{T\}\}\\varphi\(\\gamma\(t\)\)\\,d\\Gamma\(\\gamma\)=∫𝒫2\(ℳ\)\(∫ΓTφ\(γ\(t\)\)dΓμ\(γ\)\)dℙ0\(μ\)\\displaystyle=\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\left\(\\int\_\{\\Gamma\_\{T\}\}\\varphi\(\\gamma\(t\)\)\\,d\\Gamma\_\{\\mu\}\(\\gamma\)\\right\)d\\mathbb\{P\}\_\{0\}\(\\mu\)=∫𝒫2\(ℳ\)φ\(γμ\(t\)\)dℙ0\(μ\)\\displaystyle=\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\varphi\(\\gamma\_\{\\mu\}\(t\)\)\\,d\\mathbb\{P\}\_\{0\}\(\\mu\)=∫𝒫2\(ℳ\)φ\(Φt\(μ\)\)dℙ0\(μ\)\.\\displaystyle=\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\varphi\(\\Phi\_\{t\}\(\\mu\)\)\\,d\\mathbb\{P\}\_\{0\}\(\\mu\)\.This is precisely the pushforward identity
ℙt=\(Φt\)\#ℙ0,t∈\[0,T\],\\mathbb\{P\}\_\{t\}=\(\\Phi\_\{t\}\)\_\{\\\#\}\\mathbb\{P\}\_\{0\},\\qquad t\\in\[0,T\],which establishes the deterministic Lagrangian representation\.
#### C\.1\.3Example: Flowing between two diracs centered atμ0\\mu\_\{0\}andμ1\\mu\_\{1\}, withℳ=ℝd\\mathcal\{M\}=\\mathbb\{R\}^\{d\}
We consider a probability path in𝒫2\(𝒫2\(ℳ\)\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\)between two dirac distributions, centered atμ0\\mu\_\{0\}andμ1\\mu\_\{1\}\. In that case, we haveℙt=δμt\\mathbb\{P\}\_\{t\}=\\delta\_\{\\mu\_\{t\}\}\. We first consider the case where the underlying manifold is the Euclidean space\. We treat the general case whereℳ\\mathcal\{M\}is an arbitrary Riemannian manifold in the next section\.
We take the pathγ\\gammato be the constant\-speed geodesics connectingμ0\\mu\_\{0\}andμ1\\mu\_\{1\}:
γ\(t\)=μt=\(ψt\)\#μ0\\gamma\(t\)=\\mu\_\{t\}=\(\\psi\_\{t\}\)\_\{\\\#\}\\mu\_\{0\}
ψt\(x\):=\(1−t\)x\+tT0\(x\),t∈\[0,1\],\\psi\_\{t\}\(x\):=\(1\-t\)x\+t\\,T\_\{0\}\(x\),\\qquad t\\in\[0,1\],
andT0T\_\{0\}the optimal transport map onℝd\\mathbb\{R\}^\{d\}betweenμ0\\mu\_\{0\}andμ1\\mu\_\{1\}\.
The minimal norm tangent vector is the vector fieldvt∈Tμt𝒫2\(ℳ\)v\_\{t\}\\in T\_\{\\mu\_\{t\}\}\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\):
vt\(z\)=∂tψt\(x\)\|x=St\(z\)=T\(St\(z\)\)−St\(z\)=\(T−Id\)\(ψt−1\(z\)\),v\_\{t\}\(z\)=\\left\.\\partial\_\{t\}\\psi\_\{t\}\(x\)\\right\|\_\{x=S\_\{t\}\(z\)\}=T\(S\_\{t\}\(z\)\)\-S\_\{t\}\(z\)=\(T\-Id\)\(\\psi\_\{t\}^\{\-1\}\(z\)\),
withSt\(x\)=ψt−1S\_\{t\}\(x\)=\\psi\_\{t\}^\{\-1\}\.
We now need to define a vectorVtV\_\{t\}on𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)that generates this path:
Vt\(μ\)=\{vtifμ=μt0otherwiseV\_\{t\}\(\\mu\)=\\begin\{cases\}v\_\{t\}\\quad\\text\{if \}\\mu=\\mu\_\{t\}\\\\ 0\\quad\\text\{otherwise\}\\end\{cases\}
We now show that\(δμt,Vt\)\(\\delta\_\{\\mu\_\{t\}\},V\_\{t\}\)solves the continuity equation on𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\. We first show that\(μt,vt\)\(\\mu\_\{t\},v\_\{t\}\)follows the continuity equation onℝd\\mathbb\{R\}^\{d\}, then lift it to show that it follows the continuity equation on𝒫2\(ℝd\)\\mathcal\{P\}\_\{2\}\(\\mathbb\{R\}^\{d\}\)\.
##### \(μt,vt\)\(\\mu\_\{t\},v\_\{t\}\)onℝd\\mathbb\{R\}^\{d\}
From the definition ofμt\\mu\_\{t\}and the change of variable formula, we have that
∫ℝdξ\(z\)dμt\(z\)=∫ℝdξ\(ψt\(x\)\)dμ0\(x\)\\int\_\{\\mathbb\{R\}^\{d\}\}\\xi\(z\)d\\mu\_\{t\}\(z\)=\\int\_\{\\mathbb\{R\}^\{d\}\}\\xi\(\\psi\_\{t\}\(x\)\)d\\mu\_\{0\}\(x\)
for any continuous test functionξ\\xi\. Differentiating with respect tott\(and noting thatψ˙t=T\(x\)−x\\dot\{\\psi\}\_\{t\}=T\(x\)\-x\), we obtain
ddt∫ℝdξ\(z\)dμt\(z\)=∫ℝd∇ξ\(ψt\(x\)\)⋅\(T\(x\)−x\)dμ0\(x\)=∫ℝd∇ξ\(z\)⋅\(T−Id\)\(ψt−1\(z\)\)dμt\(z\)=∫ℝd∇ξ\(z\)⋅vt\(z\)dμt\(z\)\\begin\{split\}\\frac\{d\}\{dt\}\\int\_\{\\mathbb\{R\}^\{d\}\}\\xi\(z\)d\\mu\_\{t\}\(z\)&=\\int\_\{\\mathbb\{R\}^\{d\}\}\\nabla\\xi\(\\psi\_\{t\}\(x\)\)\\cdot\(T\(x\)\-x\)d\\mu\_\{0\}\(x\)\\\\ &=\\int\_\{\\mathbb\{R\}^\{d\}\}\\nabla\\xi\(z\)\\cdot\(T\-Id\)\(\\psi\_\{t\}^\{\-1\}\(z\)\)d\\mu\_\{t\}\(z\)=\\int\_\{\\mathbb\{R\}^\{d\}\}\\nabla\\xi\(z\)\\cdot v\_\{t\}\(z\)d\\mu\_\{t\}\(z\)\\end\{split\}
which is the weak formulation of the continuity equation onℝd\\mathbb\{R\}^\{d\}\.
We can now adapt it to time cylindrical functions of the form
φt\(μ\)=F\(t,∫ℝdϕ1𝑑μ,…,∫ℝdϕk𝑑μ\)\.\\varphi\_\{t\}\(\\mu\)=F\\\!\\Bigl\(t,\\int\_\{\\mathbb\{R\}^\{d\}\}\\phi\_\{1\}\\,d\\mu,\\ldots,\\int\_\{\\mathbb\{R\}^\{d\}\}\\phi\_\{k\}\\,d\\mu\\Bigr\)\.Using the identity above, we have
ddt∫ℝdϕi\(z\)dμt\(z\)=∫ℝd∇ϕi\(z\)⋅vt\(z\)dμt\(z\)\\frac\{d\}\{dt\}\\int\_\{\\mathbb\{R\}^\{d\}\}\\phi\_\{i\}\(z\)d\\mu\_\{t\}\(z\)=\\int\_\{\\mathbb\{R\}^\{d\}\}\\nabla\\phi\_\{i\}\(z\)\\cdot v\_\{t\}\(z\)d\\mu\_\{t\}\(z\)
We thus have
ddtφt\(μt\)=∂tF\(t,G\(t\)\)\+∑i=1k∂xiF\(t,∫ℝdϕ1𝑑μ,…,∫ℝdϕk𝑑μ\)⋅∫ℝd∇ϕi\(z\)⋅vt\(z\)dμt\(z\)\\frac\{d\}\{dt\}\\varphi\_\{t\}\(\\mu\_\{t\}\)=\\partial\_\{t\}F\(t,G\(t\)\)\+\\sum\_\{i=1\}^\{k\}\\partial\_\{x\_\{i\}\}F\(t,\\int\_\{\\mathbb\{R\}^\{d\}\}\\phi\_\{1\}d\\mu,\.\.\.,\\int\_\{\\mathbb\{R\}^\{d\}\}\\phi\_\{k\}d\\mu\)\\cdot\\int\_\{\\mathbb\{R\}^\{d\}\}\\nabla\\phi\_\{i\}\(z\)\\cdot v\_\{t\}\(z\)d\\mu\_\{t\}\(z\)
Using the definition of the Wasserstein gradient of the cylinder function, we write
ddtφt\(μt\)=∂tφt\(μt\)\+∫ℝd∇Wφt\(μt\)⋅vtdμt=∂tφt\(μt\)\+⟨∇Wφt\(μt\),vt⟩L2\(μt\)\\begin\{split\}\\frac\{d\}\{dt\}\\varphi\_\{t\}\(\\mu\_\{t\}\)&=\\partial\_\{t\}\\varphi\_\{t\}\(\\mu\_\{t\}\)\+\\int\_\{\\mathbb\{R\}^\{d\}\}\\nabla\_\{W\}\\varphi\_\{t\}\(\\mu\_\{t\}\)\\cdot v\_\{t\}d\\mu\_\{t\}\\\\ &=\\partial\_\{t\}\\varphi\_\{t\}\(\\mu\_\{t\}\)\+\\langle\\nabla\_\{W\}\\varphi\_\{t\}\(\\mu\_\{t\}\),v\_\{t\}\\rangle\_\{L^\{2\}\(\\mu\_\{t\}\)\}\\end\{split\}
which is the our desired results onℝd\\mathbb\{R\}^\{d\}that we now need to lift to𝒫2\(ℝd\)\\mathcal\{P\}\_\{2\}\(\\mathbb\{R\}^\{d\}\)\.
##### \(ℙt,Vt\)\(\\mathbb\{P\}\_\{t\},V\_\{t\}\)on𝒫2\(ℝd\)\\mathcal\{P\}\_\{2\}\(\\mathbb\{R\}^\{d\}\)
We have
∫01∫𝒫2\(ℝd\)\(∂tφt\(μ\)\+⟨∇𝒲φt\(μ\),Vt\(μ\)⟩L2\(μ\)\)dℙt\(μ\)dt=∫01\(∂tφt\(μt\)\+⟨∇𝒲φt\(μt\),Vt\(μt\)⟩L2\(μt\)\)𝑑t\\begin\{split\}&\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathbb\{R\}^\{d\}\)\}\\left\(\\partial\_\{t\}\\varphi\_\{t\}\(\\mu\)\+\\langle\\nabla\_\{\\mathcal\{W\}\}\\varphi\_\{t\}\(\\mu\),V\_\{t\}\(\\mu\)\\rangle\_\{L^\{2\}\(\\mu\)\}\\right\)\\ d\\mathbb\{P\}\_\{t\}\(\\mu\)dt=\\\\ &\\int\_\{0\}^\{1\}\\left\(\\partial\_\{t\}\\varphi\_\{t\}\(\\mu\_\{t\}\)\+\\langle\\nabla\_\{\\mathcal\{W\}\}\\varphi\_\{t\}\(\\mu\_\{t\}\),V\_\{t\}\(\\mu\_\{t\}\)\\rangle\_\{L^\{2\}\(\\mu\_\{t\}\)\}\\right\)dt\\end\{split\}
which is equal to∫01ddtφt\(μt\)𝑑t=φ1\(μ1\)−φ0\(μ0\)\\int\_\{0\}^\{1\}\\frac\{d\}\{dt\}\\varphi\_\{t\}\(\\mu\_\{t\}\)dt=\\varphi\_\{1\}\(\\mu\_\{1\}\)\-\\varphi\_\{0\}\(\\mu\_\{0\}\)and vanishes with the assumption thatφt\\varphi\_\{t\}has a compact support in time\. Hence, we have
∫01∫𝒫2\(ℝd\)\(∂tφt\(μ\)\+⟨∇𝒲φt\(μ\),Vt\(μ\)⟩L2\(μ\)\)dℙt\(μ\)𝑑t=0,\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathbb\{R\}^\{d\}\)\}\\left\(\\partial\_\{t\}\\varphi\_\{t\}\(\\mu\)\+\\langle\\nabla\_\{\\mathcal\{W\}\}\\varphi\_\{t\}\(\\mu\),V\_\{t\}\(\\mu\)\\rangle\_\{L^\{2\}\(\\mu\)\}\\right\)\\ d\\mathbb\{P\}\_\{t\}\(\\mu\)dt=0,
which shows that\(ℙt,Vt\)\(\\mathbb\{P\}\_\{t\},V\_\{t\}\)solves the weak continuity equation on𝒫2\(ℝd\)\\mathcal\{P\}\_\{2\}\(\\mathbb\{R\}^\{d\}\)\.
#### C\.1\.4Example: Flowing between two diracs on arbitrary Riemannian manifolds
We now generalize the example above to general Riemannian manifolds\. Considering arbitrary manifolds\(ℳ,g\)\(\\mathcal\{M\},g\), and writing geodesics betweenxxandyy∈ℳ\\in\\mathcal\{M\}\.
αt\(x,y\):=expx\(tlogxy\),t∈\[0,1\],\\alpha\_\{t\}\(x,y\)\\;:=\\;\\exp\_\{x\}\\\!\\bigl\(t\\,\\log\_\{x\}y\\bigr\),\\qquad t\\in\[0,1\],
where,
- •expx:TxM→M\\exp\_\{x\}:T\_\{x\}M\\\!\\to\\\!Mis the exponential map,
- •logxy∈TxM\\log\_\{x\}y\\in T\_\{x\}Mis the inverse \(expx\(logxy\)=y\\exp\_\{x\}\(\\log\_\{x\}y\)=y\)\.
The McCann displacement interpolation onMMis defined as
ψt\(x\):=αt\(x,T\(x\)\)=expx\(tlogxT\(x\)\),μt:=\(ψt\)\#μ0,\\psi\_\{t\}\(x\):=\\alpha\_\{t\}\(x,T\(x\)\)=\\exp\_\{x\}\(t\\,\\log\_\{x\}T\(x\)\),\\qquad\\mu\_\{t\}\\;:=\\;\(\\psi\_\{t\}\)\_\{\\\#\}\\mu\_\{0\},
which is the Riemannian analogue of the Euclidean straight\-line interpolation and is still a constant\-speedW2W\_\{2\}\-geodesic on𝒫2\(M\)\\mathcal\{P\}\_\{2\}\(M\)\.
The velocity field then writes
vt\(⋅\|μ0,μ1\)=ddtψt\(x\)\.v\_\{t\}\(\\cdot\|\\mu\_\{0\},\\mu\_\{1\}\)=\\frac\{d\}\{dt\}\\psi\_\{t\}\(x\)\.
We defineℙt=δμt\\mathbb\{P\}\_\{t\}=\\delta\_\{\\mu\_\{t\}\}andVt\(μ\)=vtV\_\{t\}\(\\mu\)=v\_\{t\}ifμ=μt\\mu=\\mu\_\{t\}and00otherwise\.
We can now follow the same procedure as in the Euclidean case to show that\(Vt,ℙt\)\(V\_\{t\},\\mathbb\{P\}\_\{t\}\)solves the weak continuity equation on𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\. We first verify it onℳ\\mathcal\{M\}and then lift it to𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\.
##### \(μt,vt\)\(\\mu\_\{t\},v\_\{t\}\)onℳ\\mathcal\{M\}
From the definition ofμt\\mu\_\{t\}and the change of variable formula, we have that
∫ℳξ\(z\)dμt\(z\)=∫ℳξ\(ψt\(x\)\)dμ0\(x\)\\int\_\{\\mathcal\{M\}\}\\xi\(z\)d\\mu\_\{t\}\(z\)=\\int\_\{\\mathcal\{M\}\}\\xi\(\\psi\_\{t\}\(x\)\)d\\mu\_\{0\}\(x\)
for any continuous test functionξ\\xi\. Differentiating with respect tott\(and using thatvt\(ψt\(x\)\)=ψ˙t\(x\)v\_\{t\}\(\\psi\_\{t\}\(x\)\)=\\dot\{\\psi\}\_\{t\}\(x\)\), we obtain
ddt∫ℳξ\(z\)dμt\(z\)=∫ℳ⟨∇ξ\(ψt\(x\)\),ψ˙t⟩gdμ0\(x\)=∫ℳ⟨∇ξ\(ψt\(x\)\),vt\(ψt\(x\)\)⟩gdμ0\(x\)=∫ℳ⟨∇ξ\(z\),vt\(z\)⟩gdμt\(z\)\\begin\{split\}\\frac\{d\}\{dt\}\\int\_\{\\mathcal\{M\}\}\\xi\(z\)d\\mu\_\{t\}\(z\)&=\\int\_\{\\mathcal\{M\}\}\\langle\\nabla\\xi\(\\psi\_\{t\}\(x\)\),\\dot\{\\psi\}\_\{t\}\\rangle\_\{g\}d\\mu\_\{0\}\(x\)\\\\ &=\\int\_\{\\mathcal\{M\}\}\\langle\\nabla\\xi\(\\psi\_\{t\}\(x\)\),v\_\{t\}\(\\psi\_\{t\}\(x\)\)\\rangle\_\{g\}d\\mu\_\{0\}\(x\)\\\\ &=\\int\_\{\\mathcal\{M\}\}\\langle\\nabla\\xi\(z\),v\_\{t\}\(z\)\\rangle\_\{g\}d\\mu\_\{t\}\(z\)\\end\{split\}
This can again be extended to time\-dependent cylinder functionsφt\\varphi\_\{t\}to yield
ddtφt\(μt\)=∂tφt\(μt\)\+∫ℳ⟨∇Wφt\(μt\),vt⟩gdμt=∂tφt\(μt\)\+⟨∇Wφt\(μt\),vt⟩L2\(μt\)\\begin\{split\}\\frac\{d\}\{dt\}\\varphi\_\{t\}\(\\mu\_\{t\}\)&=\\partial\_\{t\}\\varphi\_\{t\}\(\\mu\_\{t\}\)\+\\int\_\{\\mathcal\{M\}\}\\langle\\nabla\_\{W\}\\varphi\_\{t\}\(\\mu\_\{t\}\),v\_\{t\}\\rangle\_\{g\}d\\mu\_\{t\}\\\\ &=\\partial\_\{t\}\\varphi\_\{t\}\(\\mu\_\{t\}\)\+\\langle\\nabla\_\{W\}\\varphi\_\{t\}\(\\mu\_\{t\}\),v\_\{t\}\\rangle\_\{L^\{2\}\(\\mu\_\{t\}\)\}\\end\{split\}
where⟨⋅,⋅⟩L2\(μt\)\\langle\\cdot,\\cdot\\rangle\_\{L^\{2\}\(\\mu\_\{t\}\)\}implicitly encodes the Riemannian metric in the inner product\.
##### \(ℙt,Vt\)\(\\mathbb\{P\}\_\{t\},V\_\{t\}\)on𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)
Sinceℙt\\mathbb\{P\}\_\{t\}is concentrated onμt\\mu\_\{t\}, we have:
∫01∫𝒫2\(ℳ\)\(∂tφt\(μ\)\+⟨∇𝒲φt\(μ\),Vt\(μ\)⟩L2\(μ\)\)dℙt\(μ\)dt=∫01\(∂tφt\(μt\)\+⟨∇𝒲φt\(μt\),Vt\(μt\)⟩L2\(μt\)\)𝑑t\\begin\{split\}&\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\left\(\\partial\_\{t\}\\varphi\_\{t\}\(\\mu\)\+\\langle\\nabla\_\{\\mathcal\{W\}\}\\varphi\_\{t\}\(\\mu\),V\_\{t\}\(\\mu\)\\rangle\_\{L^\{2\}\(\\mu\)\}\\right\)\\ d\\mathbb\{P\}\_\{t\}\(\\mu\)dt=\\\\ &\\int\_\{0\}^\{1\}\\left\(\\partial\_\{t\}\\varphi\_\{t\}\(\\mu\_\{t\}\)\+\\langle\\nabla\_\{\\mathcal\{W\}\}\\varphi\_\{t\}\(\\mu\_\{t\}\),V\_\{t\}\(\\mu\_\{t\}\)\\rangle\_\{L^\{2\}\(\\mu\_\{t\}\)\}\\right\)dt\\end\{split\}
which, from our derivation above, is equal to∫01ddtφt\(μt\)𝑑t=φ1\(μ1\)−φ0\(μ0\)\\int\_\{0\}^\{1\}\\frac\{d\}\{dt\}\\varphi\_\{t\}\(\\mu\_\{t\}\)dt=\\varphi\_\{1\}\(\\mu\_\{1\}\)\-\\varphi\_\{0\}\(\\mu\_\{0\}\)and vanishes with the assumption thatφt\\varphi\_\{t\}has a compact support in time\. Hence, we have
∫01∫𝒫2\(ℳ\)\(∂tφt\(μ\)\+⟨∇𝒲φt\(μ\),Vt\(μ\)⟩L2\(μ\)\)dℙt\(μ\)𝑑t=0,\\int\_\{0\}^\{1\}\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\left\(\\partial\_\{t\}\\varphi\_\{t\}\(\\mu\)\+\\langle\\nabla\_\{\\mathcal\{W\}\}\\varphi\_\{t\}\(\\mu\),V\_\{t\}\(\\mu\)\\rangle\_\{L^\{2\}\(\\mu\)\}\\right\)\\ d\\mathbb\{P\}\_\{t\}\(\\mu\)dt=0,
which shows that\(ℙt,Vt\)\(\\mathbb\{P\}\_\{t\},V\_\{t\}\)solves the weak continuity equation on𝒫2\(ℳ\)\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\.
### C\.2Marginal vector field generates marginal probability path
In this section, we show that marginalizing conditional velocity fields preserves the weak continuity equation\. Under the hypotheses of Theorem[21](https://arxiv.org/html/2609.25659#Thmtheorem21)for the marginal pair, including trajectory uniqueness, the marginal vector field also generates the marginal probability path\.
###### Assumption 22\(Conditional velocity and probability path solve continuity equation\)\.
We assume that for any given boundary measuresμ0,μ1\\mu\_\{0\},\\mu\_\{1\}, the conditional flow\(ℙt\(⋅\|μ0,μ1\)\)t\(\\mathbb\{P\}\_\{t\}\(\\cdot\|\\mu\_\{0\},\\mu\_\{1\}\)\)\_\{t\}and its velocity fieldVt\(μ\|μ0,μ1\)V\_\{t\}\(\\mu\|\\mu\_\{0\},\\mu\_\{1\}\)satisfy the weak continuity equation\. This means for any suitable test \(cylinder\) functionφt\(μ\)\\varphi\_\{t\}\(\\mu\), the following holds:
∫0T∫𝒫2\(ℳ\)\(∂tφt\(μ\)\+⟨∇𝒲φt\(μ\),Vt\(μ\|μ0,μ1\)⟩L2\(μ\)\)dℙt\(μ\|μ0,μ1\)𝑑t=0\.\\int\_\{0\}^\{T\}\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\left\(\\partial\_\{t\}\\varphi\_\{t\}\(\\mu\)\+\\langle\\nabla\_\{\\mathcal\{W\}\}\\varphi\_\{t\}\(\\mu\),V\_\{t\}\(\\mu\|\\mu\_\{0\},\\mu\_\{1\}\)\\rangle\_\{L^\{2\}\(\\mu\)\}\\right\)\\ d\\mathbb\{P\}\_\{t\}\(\\mu\|\\mu\_\{0\},\\mu\_\{1\}\)dt=0\.
###### Definition 23\(Marginal velocity field\)\.
Analogously to classical flow matching, we define the marginal velocity fieldVt\(μ\)V\_\{t\}\(\\mu\)as the conditional expectation of the conditional velocity fieldsVt\(μ\|μ0,μ1\)V\_\{t\}\(\\mu\|\\mu\_\{0\},\\mu\_\{1\}\):
Vt\(μ\):=𝔼ℙt\(μ0,μ1\|μ\)\[Vt\(μ\|μ0,μ1\)\]=∫𝒫2\(ℳ\)Vt\(μ\|μ0,μ1\)dℙt\(μ0,μ1\|μ\)V\_\{t\}\(\\mu\):=\\mathbb\{E\}\_\{\\mathbb\{P\}\_\{t\}\(\\mu\_\{0\},\\mu\_\{1\}\|\\mu\)\}\[V\_\{t\}\(\\mu\|\\mu\_\{0\},\\mu\_\{1\}\)\]=\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}V\_\{t\}\(\\mu\|\\mu\_\{0\},\\mu\_\{1\}\)\\ d\\mathbb\{P\}\_\{t\}\(\\mu\_\{0\},\\mu\_\{1\}\|\\mu\)
wheredℙt\(μ0,μ1\|μ\)=dℙt\(μ\|μ0,μ1\)dΠ\(μ0,μ1\)dℙt\(μ\)d\\mathbb\{P\}\_\{t\}\(\\mu\_\{0\},\\mu\_\{1\}\|\\mu\)=\\frac\{d\\mathbb\{P\}\_\{t\}\(\\mu\|\\mu\_\{0\},\\mu\_\{1\}\)d\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)\}\{d\\mathbb\{P\}\_\{t\}\(\\mu\)\}is the conditional probability measure\.
This definition implies the key relation:
Vt\(μ\)dℙt\(μ\)=∫𝒫2\(ℳ\)×𝒫2\(ℳ\)Vt\(μ\|μ0,μ1\)dℙt\(μ\|μ0,μ1\)𝑑Π\(μ0,μ1\)V\_\{t\}\(\\mu\)d\\mathbb\{P\}\_\{t\}\(\\mu\)=\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\\times\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}V\_\{t\}\(\\mu\|\\mu\_\{0\},\\mu\_\{1\}\)\\ d\\mathbb\{P\}\_\{t\}\(\\mu\|\\mu\_\{0\},\\mu\_\{1\}\)d\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)
###### Theorem 24\(Marginal vector field generates the marginal probability path\)\.
Assume the conditional pair\(ℙt\(⋅\|μ0,μ1\),Vt\(⋅\|μ0,μ1\)\)\(\\mathbb\{P\}\_\{t\}\(\\cdot\|\\mu\_\{0\},\\mu\_\{1\}\),V\_\{t\}\(\\cdot\|\\mu\_\{0\},\\mu\_\{1\}\)\)satisfies Assumption[22](https://arxiv.org/html/2609.25659#Thmtheorem22)\. Then the marginal vector fieldVtV\_\{t\}from Definition[23](https://arxiv.org/html/2609.25659#Thmtheorem23)and the marginal probability pathℙt\\mathbb\{P\}\_\{t\}solve \([Weak CE](https://arxiv.org/html/2609.25659#A3.Ex82)\)\. Under the hypotheses of the superposition theorem \(Theorem[21](https://arxiv.org/html/2609.25659#Thmtheorem21)\) for the marginal pair, including trajectory uniqueness,VtV\_\{t\}generates the marginal probability pathℙt\\mathbb\{P\}\_\{t\}\.
###### Proof of Theorem[24](https://arxiv.org/html/2609.25659#Thmtheorem24)\.
We want to prove that the marginal flowℙt\\mathbb\{P\}\_\{t\}also satisfies the weak continuity equation with an appropriately defined marginal velocity fieldVt\(μ\)V\_\{t\}\(\\mu\)\. That is, we aim to show that:
∫0T∫𝒫2\(ℳ\)\(∂tφt\(μ\)\+⟨∇𝒲φt\(μ\),Vt\(μ\)⟩L2\(μ\)\)dℙt\(μ\)𝑑t=0\.\\int\_\{0\}^\{T\}\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\left\(\\partial\_\{t\}\\varphi\_\{t\}\(\\mu\)\+\\langle\\nabla\_\{\\mathcal\{W\}\}\\varphi\_\{t\}\(\\mu\),V\_\{t\}\(\\mu\)\\rangle\_\{L^\{2\}\(\\mu\)\}\\right\)\\ d\\mathbb\{P\}\_\{t\}\(\\mu\)dt=0\.
Let’s evaluate the left\-hand side of the target equation by substituting the definition of the marginal flowdℙt\(μ\)d\\mathbb\{P\}\_\{t\}\(\\mu\)\. We can split the expression into two parts\.
##### Temporal derivative term
We begin with the term containing the partial derivative with respect to time,∂tφt\(μ\)\\partial\_\{t\}\\varphi\_\{t\}\(\\mu\)\. We have
∫0T∫𝒫2\(ℳ\)∂tφt\(μ\)dℙt\(μ\)𝑑t\\int\_\{0\}^\{T\}\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\partial\_\{t\}\\varphi\_\{t\}\(\\mu\)\\ d\\mathbb\{P\}\_\{t\}\(\\mu\)dt
1\. Substitute the definition of the marginal measuredℙt\(μ\)d\\mathbb\{P\}\_\{t\}\(\\mu\):
=∫0T∫𝒫2\(ℳ\)∂tφt\(μ\)\(∫𝒫2\(ℳ\)×𝒫2\(ℳ\)dℙt\(μ\|μ0,μ1\)𝑑Π\(μ0,μ1\)\)𝑑t=\\int\_\{0\}^\{T\}\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\partial\_\{t\}\\varphi\_\{t\}\(\\mu\)\\left\(\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\\times\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}d\\mathbb\{P\}\_\{t\}\(\\mu\|\\mu\_\{0\},\\mu\_\{1\}\)d\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)\\right\)dt
2\. By Fubini’s theorem, we can exchange the order of integration:
=∫𝒫2\(ℳ\)×𝒫2\(ℳ\)\(∫0T∫𝒫2\(ℳ\)∂tφt\(μ\)dℙt\(μ\|μ0,μ1\)𝑑t\)𝑑Π\(μ0,μ1\)=\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\\times\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\left\(\\int\_\{0\}^\{T\}\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\partial\_\{t\}\\varphi\_\{t\}\(\\mu\)\\ d\\mathbb\{P\}\_\{t\}\(\\mu\|\\mu\_\{0\},\\mu\_\{1\}\)dt\\right\)d\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)
3\. Using the assumed conditional continuity equation, we replace the inner integral:
=∫𝒫2\(ℳ\)×𝒫2\(ℳ\)\(−∫0T∫𝒫2\(ℳ\)⟨∇𝒲φt\(μ\),Vt\(μ\|μ0,μ1\)⟩L2\(μ\)⋅dℙt\(μ\|μ0,μ1\)dt\)dΠ\(μ0,μ1\)=\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\\times\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\Bigl\(\-\\int\_\{0\}^\{T\}\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\langle\\nabla\_\{\\mathcal\{W\}\}\\varphi\_\{t\}\(\\mu\),V\_\{t\}\(\\mu\|\\mu\_\{0\},\\mu\_\{1\}\)\\rangle\_\{L^\{2\}\(\\mu\)\}\\\\ \\cdot\\,d\\mathbb\{P\}\_\{t\}\(\\mu\|\\mu\_\{0\},\\mu\_\{1\}\)\\,dt\\Bigr\)d\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)
4\. Combining the integrals gives our final expression for the first part:
Part 1=−∫𝒫2\(ℳ\)×𝒫2\(ℳ\)∫0T∫𝒫2\(ℳ\)⟨∇𝒲φt\(μ\),Vt\(μ\|μ0,μ1\)⟩L2\(μ\)⋅dℙt\(μ\|μ0,μ1\)dtdΠ\(μ0,μ1\)\\text\{Part 1\}=\-\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\\times\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\int\_\{0\}^\{T\}\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\langle\\nabla\_\{\\mathcal\{W\}\}\\varphi\_\{t\}\(\\mu\),V\_\{t\}\(\\mu\|\\mu\_\{0\},\\mu\_\{1\}\)\\rangle\_\{L^\{2\}\(\\mu\)\}\\\\ \\cdot\\,d\\mathbb\{P\}\_\{t\}\(\\mu\|\\mu\_\{0\},\\mu\_\{1\}\)\\,dt\\,d\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)
##### Velocity term
Now we analyze the second term involving the marginal velocity fieldVt\(μ\)V\_\{t\}\(\\mu\):
Part 2=∫0T∫𝒫2\(ℳ\)⟨∇𝒲φt\(μ\),Vt\(μ\)⟩L2\(μ\)dℙt\(μ\)𝑑t\\text\{Part 2\}=\\int\_\{0\}^\{T\}\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\langle\\nabla\_\{\\mathcal\{W\}\}\\varphi\_\{t\}\(\\mu\),V\_\{t\}\(\\mu\)\\rangle\_\{L^\{2\}\(\\mu\)\}\\ d\\mathbb\{P\}\_\{t\}\(\\mu\)dt
1\. Substituting the definition of the marginal velocity field:
Part 2=∫0T∫𝒫2\(ℳ\)⟨∇𝒲φt\(μ\),∫𝒫2\(ℳ\)×𝒫2\(ℳ\)Vt\(μ\|μ0,μ1\)dℙt\(μ\|μ0,μ1\)⟩L2\(μ\)⋅dΠ\(μ0,μ1\)dt\\text\{Part 2\}=\\int\_\{0\}^\{T\}\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\left\\langle\\nabla\_\{\\mathcal\{W\}\}\\varphi\_\{t\}\(\\mu\),\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\\times\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}V\_\{t\}\(\\mu\|\\mu\_\{0\},\\mu\_\{1\}\)\\ d\\mathbb\{P\}\_\{t\}\(\\mu\|\\mu\_\{0\},\\mu\_\{1\}\)\\right\\rangle\_\{L^\{2\}\(\\mu\)\}\\\\ \\cdot\\,d\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)\\,dt
2\. Using the linearity of the inner product and the definition of conditional probability, we can combine the integrals:
=∫0T∫𝒫2\(ℳ\)∫𝒫2\(ℳ\)×𝒫2\(ℳ\)⟨∇𝒲φt\(μ\),Vt\(μ\|μ0,μ1\)⟩L2\(μ\)dℙt\(μ\|μ0,μ1\)𝑑Π\(μ0,μ1\)𝑑t=\\int\_\{0\}^\{T\}\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\int\_\{\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\\times\\mathcal\{P\}\_\{2\}\(\\mathcal\{M\}\)\}\\langle\\nabla\_\{\\mathcal\{W\}\}\\varphi\_\{t\}\(\\mu\),V\_\{t\}\(\\mu\|\\mu\_\{0\},\\mu\_\{1\}\)\\rangle\_\{L^\{2\}\(\\mu\)\}\\ d\\mathbb\{P\}\_\{t\}\(\\mu\|\\mu\_\{0\},\\mu\_\{1\}\)d\\Pi\(\\mu\_\{0\},\\mu\_\{1\}\)dt
Hence,Part 2=−Part 1\\text\{Part 2\}=\-\\text\{Part 1\}, establishing the weak continuity equation for the marginal pair\. Under the stated additional hypotheses, Theorem[21](https://arxiv.org/html/2609.25659#Thmtheorem21)then givesℙt=\(Φt\)\#ℙ0\\mathbb\{P\}\_\{t\}=\(\\Phi\_\{t\}\)\_\{\\\#\}\\mathbb\{P\}\_\{0\}, completing the proof\.
We have thus shown that if we define the marginal velocity field as the conditional expectation of the conditional velocity fields, the resulting marginal flow satisfies the weak continuity equation on the Wasserstein manifold\.
### C\.3Equivalence of conditional and non\-conditional flow matching losses
A natural loss for the flow matching objective is
ℒFM=𝔼t,μt∼ℙt\[∥Vt\(μt\)−Vtθ\(μt\)∥L2\(μt\)2\]\\mathcal\{L\}\_\{FM\}=\\mathbb\{E\}\_\{t,\\mu\_\{t\}\\sim\\mathbb\{P\}\_\{t\}\}\[\\lVert V\_\{t\}\(\\mu\_\{t\}\)\-V\_\{t\}^\{\\theta\}\(\\mu\_\{t\}\)\\rVert^\{2\}\_\{L^\{2\}\(\\mu\_\{t\}\)\}\]
where∥Vt\(μ\)∥L2\(μ\)2=∫ℳ∥Vt\(μ\)\(x\)∥g2𝑑μ\(x\)\\lVert V\_\{t\}\(\\mu\)\\rVert^\{2\}\_\{L^\{2\}\(\\mu\)\}=\\int\_\{\\mathcal\{M\}\}\\lVert V\_\{t\}\(\\mu\)\(x\)\\rVert\_\{g\}^\{2\}d\\mu\(x\)\.
The conditional version would then be
ℒCFM=𝔼t,μ0,μ1∼Π,μt∼ℙt\(⋅∣μ0,μ1\)\[∥Vt\(μt∣μ0,μ1\)−Vtθ\(μt\)∥L2\(μt\)2\]\\mathcal\{L\}\_\{CFM\}=\\mathbb\{E\}\_\{t,\\mu\_\{0\},\\mu\_\{1\}\\sim\\Pi,\\mu\_\{t\}\\sim\\mathbb\{P\}\_\{t\}\(\\cdot\\mid\\mu\_\{0\},\\mu\_\{1\}\)\}\[\\lVert V\_\{t\}\(\\mu\_\{t\}\\mid\\mu\_\{0\},\\mu\_\{1\}\)\-V\_\{t\}^\{\\theta\}\(\\mu\_\{t\}\)\\rVert^\{2\}\_\{L^\{2\}\(\\mu\_\{t\}\)\}\]
Following the original flow matching proof[Lipman et al\. \[2024\]](https://arxiv.org/html/2609.25659#bib.bib25), we obtain
∇θℒCFM=𝔼t,μ0,μ1∼Π,μt∼ℙt\(⋅∣μ0,μ1\)\[∇θ∥Vt\(μt∣μ0,μ1\)−Vtθ\(μt\)∥2L2\(μt\)\]=𝔼t,μ0,μ1∼Π,μt∼ℙt\(⋅∣μ0,μ1\)\[2⟨∇θVtθ\(μt\),Vtθ\(μt\)⟩−2⟨Vt\(μt∣μ0,μ1\),∇θVtθ\(μt\)⟩\]\\begin\{split\}&\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{CFM\}=\\mathbb\{E\}\_\{t,\\mu\_\{0\},\\mu\_\{1\}\\sim\\Pi,\\mu\_\{t\}\\sim\\mathbb\{P\}\_\{t\}\(\\cdot\\mid\\mu\_\{0\},\\mu\_\{1\}\)\}\[\\nabla\_\{\\theta\}\\lVert V\_\{t\}\(\\mu\_\{t\}\\mid\\mu\_\{0\},\\mu\_\{1\}\)\-V\_\{t\}^\{\\theta\}\(\\mu\_\{t\}\)\\rVert^\{2\}\_\{L^\{2\}\(\\mu\_\{t\}\)\}\]\\\\ &=\\mathbb\{E\}\_\{t,\\mu\_\{0\},\\mu\_\{1\}\\sim\\Pi,\\mu\_\{t\}\\sim\\mathbb\{P\}\_\{t\}\(\\cdot\\mid\\mu\_\{0\},\\mu\_\{1\}\)\}\[2\\langle\\nabla\_\{\\theta\}V\_\{t\}^\{\\theta\}\(\\mu\_\{t\}\),V\_\{t\}^\{\\theta\}\(\\mu\_\{t\}\)\\rangle\-2\\langle V\_\{t\}\(\\mu\_\{t\}\\mid\\mu\_\{0\},\\mu\_\{1\}\),\\nabla\_\{\\theta\}V\_\{t\}^\{\\theta\}\(\\mu\_\{t\}\)\\rangle\]\\\\ \\end\{split\}
Focusing on the right handside in the expectation, we have
𝔼t,μ0,μ1∼Π,μt∼ℙt\(⋅∣μ0,μ1\)\[⟨Vt\(μt∣μ0,μ1\),∇θVtθ\(μt\)⟩\]\\displaystyle\\mathbb\{E\}\_\{t,\\mu\_\{0\},\\mu\_\{1\}\\sim\\Pi,\\mu\_\{t\}\\sim\\mathbb\{P\}\_\{t\}\(\\cdot\\mid\\mu\_\{0\},\\mu\_\{1\}\)\}\[\\langle V\_\{t\}\(\\mu\_\{t\}\\mid\\mu\_\{0\},\\mu\_\{1\}\),\\nabla\_\{\\theta\}V\_\{t\}^\{\\theta\}\(\\mu\_\{t\}\)\\rangle\]=𝔼t,μt∼ℙt\[𝔼μ0,μ1∼ℙ\(⋅∣μt\)\[⟨Vt\(μt∣μ0,μ1\),∇θVtθ\(μt\)⟩\]\]\\displaystyle=\\mathbb\{E\}\_\{t,\\mu\_\{t\}\\sim\\mathbb\{P\}\_\{t\}\}\\bigl\[\\mathbb\{E\}\_\{\\mu\_\{0\},\\mu\_\{1\}\\sim\\mathbb\{P\}\(\\cdot\\mid\\mu\_\{t\}\)\}\[\\langle V\_\{t\}\(\\mu\_\{t\}\\mid\\mu\_\{0\},\\mu\_\{1\}\),\\nabla\_\{\\theta\}V\_\{t\}^\{\\theta\}\(\\mu\_\{t\}\)\\rangle\]\\bigr\]=𝔼t,μt∼ℙt\[⟨𝔼μ0,μ1∼ℙ\(⋅∣μt\)\[Vt\(μt∣μ0,μ1\)\],∇θVtθ\(μt\)⟩\]\\displaystyle=\\mathbb\{E\}\_\{t,\\mu\_\{t\}\\sim\\mathbb\{P\}\_\{t\}\}\[\\langle\\mathbb\{E\}\_\{\\mu\_\{0\},\\mu\_\{1\}\\sim\\mathbb\{P\}\(\\cdot\\mid\\mu\_\{t\}\)\}\[V\_\{t\}\(\\mu\_\{t\}\\mid\\mu\_\{0\},\\mu\_\{1\}\)\],\\nabla\_\{\\theta\}V\_\{t\}^\{\\theta\}\(\\mu\_\{t\}\)\\rangle\]=𝔼t,μt∼ℙt\[⟨Vt\(μt\),∇θVtθ\(μt\)⟩\]\\displaystyle=\\mathbb\{E\}\_\{t,\\mu\_\{t\}\\sim\\mathbb\{P\}\_\{t\}\}\[\\langle V\_\{t\}\(\\mu\_\{t\}\),\\nabla\_\{\\theta\}V\_\{t\}^\{\\theta\}\(\\mu\_\{t\}\)\\rangle\]
Hence,
∇θℒCFM=𝔼t,μt∼ℙt\[2⟨∇θVtθ\(μt\),Vtθ\(μt\)⟩−2⟨Vt\(μt\),∇θVtθ\(μt\)⟩\]\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{CFM\}=\\mathbb\{E\}\_\{t,\\mu\_\{t\}\\sim\\mathbb\{P\}\_\{t\}\}\[2\\langle\\nabla\_\{\\theta\}V\_\{t\}^\{\\theta\}\(\\mu\_\{t\}\),V\_\{t\}^\{\\theta\}\(\\mu\_\{t\}\)\\rangle\-2\\langle V\_\{t\}\(\\mu\_\{t\}\),\\nabla\_\{\\theta\}V\_\{t\}^\{\\theta\}\(\\mu\_\{t\}\)\\rangle\]\(48\)
To finish the proof, we derive the gradient ofℒFM\\mathcal\{L\}\_\{FM\}
∇θℒFM=𝔼t,μt∼ℙt\[∇θ∥Vt\(μt\)−Vtθ\(μt\)∥L2\(μt\)2\]=𝔼t,μt∼ℙt\[2⟨∇θVtθ\(μt\),Vtθ\(μt\)⟩−2⟨Vt\(μt\),∇θVtθ\(μt\)⟩\],\\begin\{split\}\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{FM\}&=\\mathbb\{E\}\_\{t,\\mu\_\{t\}\\sim\\mathbb\{P\}\_\{t\}\}\[\\nabla\_\{\\theta\}\\lVert V\_\{t\}\(\\mu\_\{t\}\)\-V\_\{t\}^\{\\theta\}\(\\mu\_\{t\}\)\\rVert^\{2\}\_\{L^\{2\}\(\\mu\_\{t\}\)\}\]\\\\ &=\\mathbb\{E\}\_\{t,\\mu\_\{t\}\\sim\\mathbb\{P\}\_\{t\}\}\[2\\langle\\nabla\_\{\\theta\}V\_\{t\}^\{\\theta\}\(\\mu\_\{t\}\),V\_\{t\}^\{\\theta\}\(\\mu\_\{t\}\)\\rangle\-2\\langle V\_\{t\}\(\\mu\_\{t\}\),\\nabla\_\{\\theta\}V\_\{t\}^\{\\theta\}\(\\mu\_\{t\}\)\\rangle\],\\end\{split\}
which coincides exactly with \([48](https://arxiv.org/html/2609.25659#A3.E48)\)\.
## Appendix DExperimental Details and additional results
Here we lay out the settings and hyperparameters used in our experiments\. We discuss how datasets were processed, errors and metrics were computed, and model architectures and training procedures\.
### D\.1Neural Network Architecture & Training
We parameterize the velocity field as a neural networkvθ\(Xt,t,c\):ℳN×\[0,1\]×𝒞→TXtℳNv\_\{\\theta\}\(X\_\{t\},t,c\):\\mathcal\{M\}^\{N\}\\times\[0,1\]\\times\\mathcal\{C\}\\to T\_\{X\_\{t\}\}\\mathcal\{M\}^\{N\}, which maps a point cloudXt=\{x1,…,xN\}X\_\{t\}=\\\{x\_\{1\},\\ldots,x\_\{N\}\\\}withxi∈ℳx\_\{i\}\\in\\mathcal\{M\}, timet∈\[0,1\]t\\in\[0,1\], and optional conditionc∈𝒞c\\in\\mathcal\{C\}to a set of tangent vectors atXtX\_\{t\}\. The network architecture is designed to respect the permutation equivariance of point clouds and optimal transport maps\.
The architecture consists of three components: \(1\) an embedding layer that maps each pointxi∈ℳx\_\{i\}\\in\\mathcal\{M\}to a latent representation, \(2\) six multi\-head self\-attention blocks that process the embedded point cloud, and \(3\) an unembedding layer that projects back to the tangent spaceTXtℳNT\_\{X\_\{t\}\}\\mathcal\{M\}^\{N\}\. Each self\-attention block processes the latent representationh∈ℝN×dh\\in\\mathbb\{R\}^\{N\\times d\}as follows\. First, we incorporate temporal and conditional information by adding learned embeddings to formh~=h\+Embedt\(ϕ\(t\)\)\+Embedc\(c\)\\tilde\{h\}=h\+\\text\{Embed\}\_\{t\}\(\\phi\(t\)\)\+\\text\{Embed\}\_\{c\}\(c\), whereϕ\(t\)\\phi\(t\)denotes Fourier features of timett\. The block then applies:
h′\\displaystyle h^\{\\prime\}←h\+MHA\(LayerNorm\(h~\)\)\\displaystyle\\leftarrow h\+\\text\{MHA\}\(\\text\{LayerNorm\}\(\\tilde\{h\}\)\)h\\displaystyle h←h′\+MLP\(LayerNorm\(h′\)\)\\displaystyle\\leftarrow h^\{\\prime\}\+\\text\{MLP\}\(\\text\{LayerNorm\}\(h^\{\\prime\}\)\)where the multi\-head attention \(MHA\) uses 4 heads, and the residual connections are to the original streamhh\. After the final attention block, the unembedding layer projects the latent representation back to the ambient dimension ofℳ\\mathcal\{M\}to produce the tangent velocity vectors\.
##### Classifier\-Free Guidance
For conditional generation, we employ classifier\-free guidance to improve sample quality and controllability\. The conditioning embeddingEmbedc\(c\)\\text\{Embed\}\_\{c\}\(c\)projects the conditionccto the unit sphere in the embedding dimension, which creates a meaningful semantic separation between null and real conditions\. Real conditions lie on the unit sphere while the null condition is represented by the zero vector at the origin\. During training, we randomly drop conditions with a probability ofpp, and for these null conditions, we set the embeddingEmbedc\(c\)\\text\{Embed\}\_\{c\}\(c\)to the zero vector \(not the inputccitself\)\. At inference time, we use the classifier\-free guidance formula:
vθCFG\(Xt,t,c\)=vθ\(Xt,t,∅\)\+w⋅\(vθ\(Xt,t,c\)−vθ\(Xt,t,∅\)\)v\_\{\\theta\}^\{\\text\{CFG\}\}\(X\_\{t\},t,c\)=v\_\{\\theta\}\(X\_\{t\},t,\\emptyset\)\+w\\cdot\\left\(v\_\{\\theta\}\(X\_\{t\},t,c\)\-v\_\{\\theta\}\(X\_\{t\},t,\\emptyset\)\\right\)wherew≥1w\\geq 1is the guidance weight, and∅\\emptysetdenotes the null condition\. By default we setp=0\.1p=0\.1andw=2w=2\.
##### Training Details
By default, we train all models for 500,000 steps using the Adam optimizer\[[Kingma and Ba, 2014](https://arxiv.org/html/2609.25659#bib.bib46)\]with a learning rate of3×10−43\\times 10^\{\-4\}\. We apply learning rate decay with a factor of 0\.99 every 5,000 steps\. During training, we sample point clouds ofmin\(N,1024\)\\min\(N,1024\)\(where N is the number of particles in the empirical measure\) points from each distribution with a batch size of 32\. The time variablettis sampled uniformly from\[0,1\]\[0,1\]for each training step, and generation is performed using integration via 1000 Euler steps\.
For computing the Riemannian entropic map, we set the entropic regularization parameterε=0\.002\\varepsilon=0\.002, where all distance matrices are scaled to have a maximum value of 1\. The number of Sinkhorn iterations is determined automatically before training by sampling 100 pairs of distributions and finding the minimum number of iterations required to achieve convergence for 95% of the pairs\.
##### Mini\-batch OT Coupling
For unconditional generation, we enhance training by matching samples within mini\-batches using optimal transport\[[Pooladian et al\., 2023](https://arxiv.org/html/2609.25659#bib.bib35)\]\. This approach effectively performs OT in the Wasserstein space itself\[[Emami and Pass, 2025](https://arxiv.org/html/2609.25659#bib.bib30),[Bonet et al\., 2025](https://arxiv.org/html/2609.25659#bib.bib24)\]\. Given source distributions\{μi\}i=1Bs∼ℙ0\\\{\\mu\_\{i\}\\\}\_\{i=1\}^\{B\_\{s\}\}\\sim\\mathbb\{P\}\_\{0\}\(noise\) and target distributions\{νj\}j=1Bt∼ℙ1\\\{\\nu\_\{j\}\\\}\_\{j=1\}^\{B\_\{t\}\}\\sim\\mathbb\{P\}\_\{1\}\(data\) within a mini\-batch, we construct a cost matrixC∈ℝBs×BtC\\in\\mathbb\{R\}^\{B\_\{s\}\\times B\_\{t\}\}and sample training pairs\(μi,νj\)∼π∗\(\\mu\_\{i\},\\nu\_\{j\}\)\\sim\\pi^\{\*\}from the resulting optimal couplingπ∗\\pi^\{\*\}\.
To ensure scalability, we compute the cost matrix using the geometric Chamfer distance rather than the Wasserstein distance:
Ci,j=CD\(μi,νj\)=1\|μi\|∑x∈μiminy∈νjdℳ\(x,y\)\+1\|νj\|∑y∈νjminx∈μidℳ\(x,y\)C\_\{i,j\}=\\text\{CD\}\(\\mu\_\{i\},\\nu\_\{j\}\)=\\frac\{1\}\{\|\\mu\_\{i\}\|\}\\sum\_\{x\\in\\mu\_\{i\}\}\\min\_\{y\\in\\nu\_\{j\}\}d\_\{\\mathcal\{M\}\}\(x,y\)\+\\frac\{1\}\{\|\\nu\_\{j\}\|\}\\sum\_\{y\\in\\nu\_\{j\}\}\\min\_\{x\\in\\mu\_\{i\}\}d\_\{\\mathcal\{M\}\}\(x,y\)wheredℳd\_\{\\mathcal\{M\}\}is the intrinsic geodesic distance on the manifoldℳ\\mathcal\{M\}\. The optimal transport planπ∗\\pi^\{\*\}is computed using the Sinkhorn algorithm with entropic regularizationε=0\.001\\varepsilon=0\.001, where distance matrices are scaled to maximum value 1 and the number of iterations is determined automatically as described above\.
All models were implemented in JAX\[[Frostig et al\., 2019](https://arxiv.org/html/2609.25659#bib.bib44)\]using optimal transport tools from\[[Cuturi et al\., 2022](https://arxiv.org/html/2609.25659#bib.bib45)\]and trained on a single NVIDIA B200 GPU, taking 3\-6 hours depending on the dataset and manifold geometry\.
##### Source Noise Generation
While rigorously defined distributions exist on non\-Euclidean domains \(e\.g\., von Mises or wrapped normal distributions\), flow matching only requires a source that is easy to sample from, without needing a closed\-form likelihood\. We note that naively sampling noise asprojℳ\(𝒩\(0,I\)\)\\text\{proj\}\_\{\\mathcal\{M\}\}\(\\mathcal\{N\}\(0,I\)\)\(whereprojℳ\\text\{proj\}\_\{\\mathcal\{M\}\}is the projection onto the manifold\) fails to produce meaningful variability\. Our source distribution must be a distribution overdistributionsonℳ\\mathcal\{M\}\. Indeed, for large enough sample sizesnn, all generated “noise” point clouds become nearly identical, resulting in a degenerate distribution\.
Instead, following[Haviv et al\. \[2024\]](https://arxiv.org/html/2609.25659#bib.bib26), we employ a two\-level hierarchical sampling scheme\. First, we randomly sample a meanμ\\muand covarianceΣ\\Sigma, then generate points fromπ\(𝒩\(μ,Σ\)\)\\pi\(\\mathcal\{N\}\(\\mu,\\Sigma\)\)\. The distribution forμ\\muis derived from the empirical mean and covariance of all points in the training data, while the distribution forΣ\\Sigmais based on the per\-sample covariances\. Specifically, let\{Li\}\\\{L\_\{i\}\\\}be the Cholesky factors of the per\-sample covariances\. We compute the element\-wise meanL¯\\bar\{L\}and standard deviationσL\\sigma\_\{L\}of these lower\-triangular matrices, sample a random lower\-triangular matrixL~∼𝒩\(L¯,diag\(σL2\)\)\\tilde\{L\}\\sim\\mathcal\{N\}\(\\bar\{L\},\\text\{diag\}\(\\sigma\_\{L\}^\{2\}\)\), and setΣ=L~L~⊤\\Sigma=\\tilde\{L\}\\tilde\{L\}^\{\\top\}\. This construction ensuresΣ\\Sigmais positive semidefinite\. Samples from the resulting mean and covariance are then projected onto the manifold geometry\. For high\-dimensional datasets \(e\.g\., single\-cell data on𝕊128−1\\mathbb\{S\}^\{128\-1\}\), we only sample diagonal covariances for computational efficiency\.
### D\.2Geometric Operations on Manifolds
We provide explicit formulas for the geometric operations used in our method across different manifolds\. For each geometry, we define: \(1\) the squared geodesic distancedℳ2\(p,q\)d^\{2\}\_\{\\mathcal\{M\}\}\(p,q\), \(2\) the geodesic interpolantγ\(t,p0,p1\)\\gamma\(t;p\_\{0\},p\_\{1\}\)connectingp0p\_\{0\}top1p\_\{1\}at timet∈\[0,1\]t\\in\[0,1\], \(3\) the tangent velocityv\(t,p0,p1\)=ddtγ\(t,p0,p1\)v\(t;p\_\{0\},p\_\{1\}\)=\\frac\{d\}\{dt\}\\gamma\(t;p\_\{0\},p\_\{1\}\), \(4\) the exponential mapexpp\(v,Δt\)\\exp\_\{p\}\(v,\\Delta t\)that moves from pointppalong tangent vectorvvby timeΔt\\Delta t, and \(5\) the tangent norm∥⋅∥Tpℳ2\\\|\\cdot\\\|\_\{T\_\{p\}\\mathcal\{M\}\}^\{2\}used for computing the training loss between predicted and target velocities\.
Important note on representation:We always represent manifolds using extrinsic coordinates embedded in Euclidean space \(e\.g\.,𝕊2\\mathbb\{S\}^\{2\}as 3D unit vectors inℝ3\\mathbb\{R\}^\{3\},ℍ2\\mathbb\{H\}^\{2\}as 3D vectors in Lorentz space\)\. This extrinsic representation significantly simplifies neural network modeling, as the network can operate on fixed\-dimensional Euclidean vectors while geometric constraints are enforced through projection operations\. These formulas correspond directly to our implementation and are included here for completeness and reproducibility\.
##### Euclidean Spaceℝd\\mathbb\{R\}^\{d\}
The standard flat geometry\. Pointsp∈ℝdp\\in\\mathbb\{R\}^\{d\}have trivial tangent spacesTpℝd≅ℝdT\_\{p\}\\mathbb\{R\}^\{d\}\\cong\\mathbb\{R\}^\{d\}\. The geodesics are straight lines\.
Distance and interpolation:
dℝd2\(p,q\)=‖p−q‖22,γ\(t,p0,p1\)=\(1−t\)p0\+tp1d^\{2\}\_\{\\mathbb\{R\}^\{d\}\}\(p,q\)=\\\|p\-q\\\|\_\{2\}^\{2\},\\qquad\\gamma\(t;p\_\{0\},p\_\{1\}\)=\(1\-t\)p\_\{0\}\+tp\_\{1\}
Velocity and exponential map:
v\(t,p0,p1\)=p1−p0,expp\(v,Δt\)=p\+v⋅Δtv\(t;p\_\{0\},p\_\{1\}\)=p\_\{1\}\-p\_\{0\},\\qquad\\exp\_\{p\}\(v,\\Delta t\)=p\+v\\cdot\\Delta t
Tangent space:‖v−w‖Tpℝd2=‖v−w‖2\\\|v\-w\\\|\_\{T\_\{p\}\\mathbb\{R\}^\{d\}\}^\{2\}=\\\|v\-w\\\|^\{2\}\. Projection:proj\(p\)=p\\text\{proj\}\(p\)=p\.
##### dd\-Torus𝕋d\\mathbb\{T\}^\{d\}
Thedd\-dimensional torus𝕋d=\(𝕊1\)d\\mathbb\{T\}^\{d\}=\(\\mathbb\{S\}^\{1\}\)^\{d\}\. Points are represented asddanglesp∈\[0,2π\)dp\\in\[0,2\\pi\)^\{d\}\. The geodesic distance accounts for periodic wraparound in each coordinate\.
Distance and logarithmic map:
d𝕋d2\(p,q\)\\displaystyle d^\{2\}\_\{\\mathbb\{T\}^\{d\}\}\(p,q\)=∑i=1dmin\(\|pi−qi\|,2π−\|pi−qi\|\)2,\\displaystyle=\\sum\_\{i=1\}^\{d\}\\min\(\\lvert p\_\{i\}\-q\_\{i\}\\rvert,2\\pi\-\\lvert p\_\{i\}\-q\_\{i\}\\rvert\)^\{2\},logp0\(p1\)\\displaystyle\\log\_\{p\_\{0\}\}\(p\_\{1\}\)=arctan2\(sin\(p1−p0\),cos\(p1−p0\)\)\\displaystyle=\\arctan 2\(\\sin\(p\_\{1\}\-p\_\{0\}\),\\cos\(p\_\{1\}\-p\_\{0\}\)\)
Interpolation and velocity:
γ\(t,p0,p1\)=\(p0\+t⋅logp0\(p1\)\)mod2π,v\(t,p0,p1\)=logp0\(p1\)\\gamma\(t;p\_\{0\},p\_\{1\}\)=\(p\_\{0\}\+t\\cdot\\log\_\{p\_\{0\}\}\(p\_\{1\}\)\)\\mod 2\\pi,\\qquad v\(t;p\_\{0\},p\_\{1\}\)=\\log\_\{p\_\{0\}\}\(p\_\{1\}\)
Exponential map:expp\(v,Δt\)=\(p\+v⋅Δt\)mod2π\\exp\_\{p\}\(v,\\Delta t\)=\(p\+v\\cdot\\Delta t\)\\mod 2\\pi\.
Tangent space:‖v−w‖Tp𝕋d2=‖v−w‖2\\\|v\-w\\\|\_\{T\_\{p\}\\mathbb\{T\}^\{d\}\}^\{2\}=\\\|v\-w\\\|^\{2\}\. Projection:proj\(p\)=pmod2π\\text\{proj\}\(p\)=p\\mod 2\\pi\.
##### dd\-Sphere𝕊d\\mathbb\{S\}^\{d\}
Thedd\-dimensional sphere embedded inℝd\+1\\mathbb\{R\}^\{d\+1\}\. Points are unit vectorsp∈ℝd\+1p\\in\\mathbb\{R\}^\{d\+1\}with‖p‖=1\\\|p\\\|=1\. The tangent space atppisTp𝕊d=\{v∈ℝd\+1:⟨v,p⟩=0\}T\_\{p\}\\mathbb\{S\}^\{d\}=\\\{v\\in\\mathbb\{R\}^\{d\+1\}:\\langle v,p\\rangle=0\\\}\. We use spherical linear interpolation \(SLERP\)\.
Distance:d𝕊d2\(p,q\)=arccos\(clip\(⟨p,q⟩,−1,1\)\)2d^\{2\}\_\{\\mathbb\{S\}^\{d\}\}\(p,q\)=\\arccos\(\\text\{clip\}\(\\langle p,q\\rangle,\-1,1\)\)^\{2\}, whereθ=arccos\(⟨p0,p1⟩\)\\theta=\\arccos\(\\langle p\_\{0\},p\_\{1\}\\rangle\)\.
Interpolation:
γ\(t,p0,p1\)=\{sin\(\(1−t\)θ\)sin\(θ\)p0\+sin\(tθ\)sin\(θ\)p1ifsin\(θ\)≥10−6\(1−t\)p0\+tp1otherwise\\gamma\(t;p\_\{0\},p\_\{1\}\)=\\begin\{cases\}\\frac\{\\sin\(\(1\-t\)\\theta\)\}\{\\sin\(\\theta\)\}p\_\{0\}\+\\frac\{\\sin\(t\\theta\)\}\{\\sin\(\\theta\)\}p\_\{1\}&\\text\{if \}\\sin\(\\theta\)\\geq 10^\{\-6\}\\\\ \(1\-t\)p\_\{0\}\+tp\_\{1\}&\\text\{otherwise\}\\end\{cases\}
Velocity:
v\(t,p0,p1\)=\{−θcos\(\(1−t\)θ\)sin\(θ\)p0\+θcos\(tθ\)sin\(θ\)p1ifsin\(θ\)≥10−6−p0\+p1otherwisev\(t;p\_\{0\},p\_\{1\}\)=\\begin\{cases\}\-\\frac\{\\theta\\cos\(\(1\-t\)\\theta\)\}\{\\sin\(\\theta\)\}p\_\{0\}\+\\frac\{\\theta\\cos\(t\\theta\)\}\{\\sin\(\\theta\)\}p\_\{1\}&\\text\{if \}\\sin\(\\theta\)\\geq 10^\{\-6\}\\\\ \-p\_\{0\}\+p\_\{1\}&\\text\{otherwise\}\\end\{cases\}
Exponential map:
expp\(v,Δt\)=\{cos\(‖v‖Δt\)p\+sin\(‖v‖Δt\)v‖v‖if‖v‖≥10−6normalize\(p\+vΔt\)otherwise\\exp\_\{p\}\(v,\\Delta t\)=\\begin\{cases\}\\cos\(\\\|v\\\|\\Delta t\)p\+\\sin\(\\\|v\\\|\\Delta t\)\\frac\{v\}\{\\\|v\\\|\}&\\text\{if \}\\\|v\\\|\\geq 10^\{\-6\}\\\\ \\text\{normalize\}\(p\+v\\Delta t\)&\\text\{otherwise\}\\end\{cases\}
Tangent space:‖v−w‖Tp𝕊d2=‖vtan−wtan‖2\\\|v\-w\\\|\_\{T\_\{p\}\\mathbb\{S\}^\{d\}\}^\{2\}=\\\|v\_\{\\tan\}\-w\_\{\\tan\}\\\|^\{2\}wherevtan=v−⟨v,p⟩pv\_\{\\tan\}=v\-\\langle v,p\\rangle p\. Projection:proj\(p\)=p/‖p‖\\text\{proj\}\(p\)=p/\\\|p\\\|\.
##### Hyperbolic Spaceℍd\\mathbb\{H\}^\{d\}\(Lorentz Model\)
Hyperbolic space in the Lorentz \(hyperboloid\) model\. Points lie on\{x∈ℝd\+1:⟨x,x⟩L=−1,x0\>0\}\\\{x\\in\\mathbb\{R\}^\{d\+1\}:\\langle x,x\\rangle\_\{L\}=\-1,x\_\{0\}\>0\\\}with Minkowski inner product⟨x,y⟩L=−x0y0\+∑i=1dxiyi\\langle x,y\\rangle\_\{L\}=\-x\_\{0\}y\_\{0\}\+\\sum\_\{i=1\}^\{d\}x\_\{i\}y\_\{i\}\. The tangent space atppisTpℍd=\{v:⟨v,p⟩L=0\}T\_\{p\}\\mathbb\{H\}^\{d\}=\\\{v:\\langle v,p\\rangle\_\{L\}=0\\\}\.
Distance:
dℍd2\(p,q\)=arccosh\(clip\(−⟨p,q⟩L,1\+10−7,∞\)\)2,ω=arccosh\(−⟨p0,p1⟩L\)\.d^\{2\}\_\{\\mathbb\{H\}^\{d\}\}\(p,q\)=\\mathrm\{arccosh\}\\\!\\bigl\(\\mathrm\{clip\}\(\-\\langle p,q\\rangle\_\{L\},\\,1\+10^\{\-7\},\\,\\infty\)\\bigr\)^\{2\},\\quad\\omega=\\mathrm\{arccosh\}\(\-\\langle p\_\{0\},p\_\{1\}\\rangle\_\{L\}\)\.
Interpolation:
γ\(t,p0,p1\)=\{sinh\(\(1−t\)ω\)sinh\(ω\)p0\+sinh\(tω\)sinh\(ω\)p1ifsinh\(ω\)≥10−6\(1−t\)p0\+tp1otherwise\\gamma\(t;p\_\{0\},p\_\{1\}\)=\\begin\{cases\}\\frac\{\\sinh\(\(1\-t\)\\omega\)\}\{\\sinh\(\\omega\)\}p\_\{0\}\+\\frac\{\\sinh\(t\\omega\)\}\{\\sinh\(\\omega\)\}p\_\{1\}&\\text\{if \}\\sinh\(\\omega\)\\geq 10^\{\-6\}\\\\ \(1\-t\)p\_\{0\}\+tp\_\{1\}&\\text\{otherwise\}\\end\{cases\}
Velocity:
v\(t,p0,p1\)=\{−ωcosh\(\(1−t\)ω\)sinh\(ω\)p0\+ωcosh\(tω\)sinh\(ω\)p1sinh\(ω\)≥10−6−p0\+p1otherwisev\(t;p\_\{0\},p\_\{1\}\)=\\begin\{cases\}\\tfrac\{\-\\omega\\cosh\(\(1\-t\)\\omega\)\}\{\\sinh\(\\omega\)\}p\_\{0\}\+\\tfrac\{\\omega\\cosh\(t\\omega\)\}\{\\sinh\(\\omega\)\}p\_\{1\}&\\sinh\(\\omega\)\\geq 10^\{\-6\}\\\\ \-p\_\{0\}\+p\_\{1\}&\\text\{otherwise\}\\end\{cases\}
Exponential map\(withc=‖v‖Lc=\\\|v\\\|\_\{L\},‖v‖L=max\(⟨v,v⟩L,0\)\\\|v\\\|\_\{L\}=\\sqrt\{\\max\(\\langle v,v\\rangle\_\{L\},0\)\}\):
expp\(v,Δt\)=\{cosh\(cΔt\)p\+sinh\(cΔt\)cvc≥10−6proj\(p\+vΔt\)otherwise\\exp\_\{p\}\(v,\\Delta t\)=\\begin\{cases\}\\cosh\(c\\Delta t\)\\,p\+\\tfrac\{\\sinh\(c\\Delta t\)\}\{c\}\\,v&c\\geq 10^\{\-6\}\\\\ \\mathrm\{proj\}\(p\+v\\Delta t\)&\\text\{otherwise\}\\end\{cases\}
Tangent space:‖v−w‖Tpℍd2=⟨vtan−wtan,vtan−wtan⟩L\\\|v\-w\\\|\_\{T\_\{p\}\\mathbb\{H\}^\{d\}\}^\{2\}=\\langle v\_\{\\tan\}\-w\_\{\\tan\},v\_\{\\tan\}\-w\_\{\\tan\}\\rangle\_\{L\}wherevtan=v\+⟨v,p⟩Lpv\_\{\\tan\}=v\+\\langle v,p\\rangle\_\{L\}p\. Projection:proj\(p\)=\[1\+∥p1:d∥2,p1,…,pd\]⊤\\text\{proj\}\(p\)=\[\\sqrt\{1\+\\\|p\_\{1:d\}\\\|^\{2\}\},p\_\{1\},\\ldots,p\_\{d\}\]^\{\\top\}\.
DatasetMethodTrainNNTestNN\|𝒫\|\|\\mathcal\{P\}\|Sink\. itersTimeMNIST 3FM6,1311,010237–2h 15mRFM–1h 39mSet\-FM–4h 17mSet\-RFM–3h 08mWFM17909h 12mRWEFM8506h 04mPSF–∼\\sim12hPVD–∼\\sim12hMNIST 4FM5,842982215–2h 11mRFM–1h 36mSet\-FM–4h 13mSet\-RFM–3h 11mWFM240010h 54mRWEFM9306h 00mPSF–∼\\sim12hPVD–∼\\sim12hMNIST 8FM5,851974276–2h 12mRFM–1h 41mSet\-FM–4h 02mSet\-RFM–3h 12mWFM171010h 29mRWEFM8106h 34mPSF–∼\\sim12hPVD–∼\\sim12hEMNIST hFM4,800800300–46mRFM–46mSet\-FM–2h 03mSet\-RFM–1h 58mWFM4903h 33mRWEFM6004h 06mPSF–∼\\sim12hPVD–∼\\sim12hEMNIST wFM4,800800306–46mRFM–47mSet\-FM–2h 01mSet\-RFM–2h 00mWFM5303h 52mRWEFM7204h 22mPSF–∼\\sim12hPVD–∼\\sim12hEMNIST yFM4,800800278–43mRFM–44mSet\-FM–2h 09mSet\-RFM–1h 50mWFM5403h 42mRWEFM6504h 00mPSF–∼\\sim12hPVD–∼\\sim12hKMNIST kiFM6,0001,000464–1h 30mRFM–1h 44mSet\-FM–2h 13mSet\-RFM–2h 15mWFM4204h 51mRWEFM4405h 11mKMNIST naFM6,0001,000418–1h 51mRFM–2h 03mSet\-FM–2h 43mSet\-RFM–2h 53mWFM3504h 23mRWEFM4805h 09mKMNIST maFM6,0001,000428–1h 47mRFM–1h 37mSet\-FM–2h 13mSet\-RFM–2h 41mWFM4204h 37mRWEFM4605h 05mMDcathSet\-FM3,432858127,000–4h 23mSet\-RFM–5h 11mWFM6007h 32mRWEFM10909h 15mscRNA\-seqSet\-FM1,2002405,000–7h 33mSet\-RFM–8h 44mWFM42016h 08mRWEFM31014h 47mTable S1:Wall\-clock training timesSet\-FM/Set\-RFM pay an extra self\-attention cost over point\-cloud pairs; WFM and RWEFM pay a per\-step Sinkhorn OT cost whose iteration count \(column 6\) is determined automatically per experiment\. Even on the same dataset, WFM and RWEFM require different iteration counts because they use different distance kernels \(ambient Euclidean vs\. Riemannian geodesic\) leading to different convergence rates\. Dashes \(–\): no OT coupling\.
##### Handling cut loci on the manifolds\.
The validity of our theorems relies on the assumption that thelog\\logmap is single valued and smooth\. That said, for the manifolds considered above, the cut locus of a fixed point is a measure\-zero set, such that ambiguity occurs with probability zero\. In practice, our implementation numerically disambiguates these situations by returning a single value for the logarithmic map, as the above velocity expressions show\.
### D\.3Benchmarking metrics for generation of distributions on Manifolds
Figure S1:Individual whole single\-cell samples generated by RWEFM\.We show additional individual samples generated by RWEFM in the latent space of SCimilarity\[[Heimberg et al\., 2025](https://arxiv.org/html/2609.25659#bib.bib39)\], a foundation model for single\-cell data that operates in𝕊128−1\\mathbb\{S\}^\{128\-1\}\. Each panel shows the UMAP visualization of a single\-cell sample, with cells colored by their cell type\. Generated samples \(left\) closely match the cellular composition and structure of true samples \(right\)\.For unconditional generation, we compare real and generated distributions using a classwise nearest\-neighbor accuracy deviation and Maximum Mean Discrepancy \(MMD\), building on point\-cloud evaluation measures such as those in[Zhou et al\. \[2021\]](https://arxiv.org/html/2609.25659#bib.bib42)\. Each observation in this evaluation is an entire point cloud\. We generate as many clouds as there are clouds in the held\-out test set, pool the two sets, and classify each cloud as real or generated using the label of its nearest neighbor, excluding the cloud itself\. Distances between clouds are computed using geometric CD or EMD\.
Letareala\_\{\\mathrm\{real\}\}be the fraction of real clouds correctly classified as real andagena\_\{\\mathrm\{gen\}\}the fraction of generated clouds correctly classified as generated\. We report the classwise 1\-NN deviation abbreviated 1\-NN\-D in the tables\. In particular, this is not the raw pooled classification accuracy\. With equally sized classes, the pooled accuracy is\(areal\+agen\)/2\(a\_\{\\mathrm\{real\}\}\+a\_\{\\mathrm\{gen\}\}\)/2\. Both\(areal,agen\)=\(1,0\)\(a\_\{\\mathrm\{real\}\},a\_\{\\mathrm\{gen\}\}\)=\(1,0\)and\(0\.5,0\.5\)\(0\.5,0\.5\)give pooled accuracy0\.50\.5, although the first case labels every cloud as real\. Our score distinguishes these cases, giving0\.50\.5and00, respectively\.
Lower values therefore indicate smaller classwise deviations from 50% accuracy\. The minimum value00means that both empirical classwise accuracies equal0\.50\.5; the maximum is0\.50\.5\. Finite\-sample fluctuations can yield a nonzero score even when the real and generated distributions coincide\.
We also compute the Maximum Mean Discrepancy \(MMD\) using both geometric CD and EMD kernels\. The MMD is defined as:
MMD2\(μ,ν\)=𝔼x,x′∼μ\[k\(x,x′\)\]\+𝔼y,y′∼ν\[k\(y,y′\)\]−2𝔼x∼μ,y∼ν\[k\(x,y\)\]\\text\{MMD\}^\{2\}\(\\mu,\\nu\)=\\mathbb\{E\}\_\{x,x^\{\\prime\}\\sim\\mu\}\[k\(x,x^\{\\prime\}\)\]\+\\mathbb\{E\}\_\{y,y^\{\\prime\}\\sim\\nu\}\[k\(y,y^\{\\prime\}\)\]\-2\\mathbb\{E\}\_\{x\\sim\\mu,y\\sim\\nu\}\[k\(x,y\)\]wherek\(⋅,⋅\)k\(\\cdot,\\cdot\)is a kernel function\. We usek\(x,y\)=exp\(−d\(x,y\)/σ\)k\(x,y\)=\\exp\(\-d\(x,y\)/\\sigma\)whereddis either the geometric CD or EMD distance, andσ=0\.1\\sigma=0\.1is a scaling factor\. We report both MMD\-CD and MMD\-EMD values\.
For conditional generation, where we have a known ground\-truth target distribution we are trying to match, we directly compare the generated distribution to the ground\-truth using geometric Wasserstein distancesW1W\_\{1\}andW2W\_\{2\}, as well as MMD\. All distances are computed using the intrinsic geometry of the underlying manifold\.
### D\.4MNIST & EMNIST on Sphere and Hyperbolic Space
In[fig\.2](https://arxiv.org/html/2609.25659#S2.F2)we show how RWEFM can learn to generate distributions on the sphere and hyperbolic space\. We use the MNIST digit dataset\[[LeCun et al\., 1998](https://arxiv.org/html/2609.25659#bib.bib37)\]for the sphere experiments and EMNIST letters\[[Cohen et al\., 2017](https://arxiv.org/html/2609.25659#bib.bib36)\]for the hyperbolic experiments\. Both datasets consist of28×2828\\times 28grayscale images of handwritten digits/letters\. We convert each image to a point cloud inℝ2\\mathbb\{R\}^\{2\}by thresholding each pixel and normalizing the coordinates to\[−1,1\]\[\-1,1\]\. This transforms each image from a point inℝ28×28\\mathbb\{R\}^\{28\\times 28\}to a distribution overℝ2\\mathbb\{R\}^\{2\}\.
For MNIST, we convert the point\-cloud overℝ2\\mathbb\{R\}^\{2\}to a distribution over𝕊2\\mathbb\{S\}^\{2\}\. We treat thexxandyycoordinates of each point as longitude and latitude on the sphere and produce the3D3Dcoordinates via the spherical to Cartesian conversion\. For EMNIST, we use the Lorentz model of hyperbolic spaceℍ2\\mathbb\{H\}^\{2\}\. We convert the point\-cloud overℝ2\\mathbb\{R\}^\{2\}to a distribution overℍ2\\mathbb\{H\}^\{2\}by simply keeping thexxandyycoordinates as is and adding azzcoordinate such that each point lies on the hyperboloid defined by−x2−y2\+z2=1\-x^\{2\}\-y^\{2\}\+z^\{2\}=1\.
For the torus experiments, we use KMNIST\[[ROIS\-CODH, 2018](https://arxiv.org/html/2609.25659#bib.bib1)\], a dataset of28×2828\\times 28grayscale images of 10 handwritten Kanji characters\. As with MNIST and EMNIST, we threshold each pixel and normalize to obtain a point cloud inℝ2\\mathbb\{R\}^\{2\}\. To embed on the 2\-torus𝕋2=𝕊1×𝕊1\\mathbb\{T\}^\{2\}=\\mathbb\{S\}^\{1\}\\times\\mathbb\{S\}^\{1\}, we rescale thexxandyycoordinates from\[−1,1\]\[\-1,1\]to\[0,2π\]\[0,2\\pi\], treating each as an angular coordinate on the torus\.
Figure S2:Entropic Map Benchmark\.a\.We evaluate our Riemannian entropic map estimator on constructed examples on the sphere𝕊2\\mathbb\{S\}^\{2\}and hyperbolic spaceℍ2\\mathbb\{H\}^\{2\}where the ground\-truth map is known\. This ground truth arises from shifting an original distribution via the Riemannian gradient of a convex function over each space\. We compare the estimation error against the Euclidean entropic map and the unregularized OT map\.b\.Computational efficiency of our Riemannian entropic map implementation compared the Euclidean map and solving unregularized OT\.
### D\.5Distributions on a General Mesh \(Stanford Bunny\)
To demonstrate RWEFM on a geometry without closed\-formexp\\exp/log\\logmaps, we generate MNIST digit distributions on the surface of the Stanford bunny\. We reduce the mesh resolution to∼\\sim7k faces by vertex clustering, then center it at its centroid and rescale by the largest absolute coordinate to map it isotropically into\[−1,1\]3\[\-1,1\]^\{3\}\. We equip the mesh with the spectral \(biharmonic\) premetric of[Chen and Lipman \[2023\]](https://arxiv.org/html/2609.25659#bib.bib47)\. Concretely, we build the discrete Laplace–Beltrami operator from the cotangent stiffness matrixLL—with edge weights12\(cotαij\+cotβij\)\\tfrac\{1\}\{2\}\(\\cot\\alpha\_\{ij\}\+\\cot\\beta\_\{ij\}\), whereαij,βij\\alpha\_\{ij\},\\beta\_\{ij\}are the two angles opposite edge\(i,j\)\(i,j\)—and the lumped \(barycentric\) mass matrixM=diag\(mi\)M=\\mathrm\{diag\}\(m\_\{i\}\), wheremi=13∑f∋iAfm\_\{i\}=\\tfrac\{1\}\{3\}\\sum\_\{f\\ni i\}A\_\{f\}sums one third of the areasAfA\_\{f\}of the faces incident to vertexii\. We solve the generalized eigenproblemLϕi=λiMϕiL\\phi\_\{i\}=\\lambda\_\{i\}M\\phi\_\{i\}for itsk=100k=100smallest eigenpairs\(λi,ϕi\)\(\\lambda\_\{i\},\\phi\_\{i\}\)and define the squared distanced2\(x,y\)=∑iλi−2\(ϕi\(x\)−ϕi\(y\)\)2d^\{2\}\(x,y\)=\\sum\_\{i\}\\lambda\_\{i\}^\{\-2\}\\,\(\\phi\_\{i\}\(x\)\-\\phi\_\{i\}\(y\)\)^\{2\}, evaluated at surface points by barycentric interpolation over the nearest triangle\. Because the spectral premetric is not geodesic \(∥∇d∥≠1\\lVert\\nabla d\\rVert\\neq 1\), we integrate the premetric conditional vector field of[Chen and Lipman \[2023\]](https://arxiv.org/html/2609.25659#bib.bib47)directly for the interpolant, and evaluate the Riemannian entropic map from the distance function and mesh projection only \(Appendix[A\.6](https://arxiv.org/html/2609.25659#A1.SS6)\)\. Each MNIST image is binarized with Otsu’s threshold; from the foreground pixels we sample a fixedN=150N=150points \(with replacement when fewer than150150are available\), normalize them to\[−1,1\]2\[\-1,1\]^\{2\}, and add small Gaussian jitter\. Each resulting planar cloud is then laid onto a fixed tangent chart on the bunny’s flank via the mesh exponential map\.
We train one model per digit class\{0,2,9\}\\\{0,2,9\\\}for100,000100\{,\}000steps, using the same network architecture and optimizer as the closed\-form experiments\.In this experiment, the methods denoted RWEFM and WFM use the sampled OT maprather than the barycentric entropic map used elsewhere in the paper: instead of transporting each source particle to a weighted average of target particles, we assign it to a single target particle drawn from the entropic plan \(the same sampled map used for the single\-cell experiments\)\. RWEFM and SetRFM operate on the mesh with the spectral metric, whereas WFM and SetFM operate in ambientℝ3\\mathbb\{R\}^\{3\}and their generated clouds are projected onto the surface for scoring; SetRFM and SetFM use random \(identity\) couplings without OT\. Extended MMD metrics are reported in[tableS2](https://arxiv.org/html/2609.25659#A4.T2)\.
Table S2:MMD on the Stanford bunny mesh\(lower is better\), with Chamfer \(CD\) and Earth Mover’s \(EMD\) ground metrics under the mesh spectral distance, for MNIST digits generated on the bunny\. Companion to[table3](https://arxiv.org/html/2609.25659#S4.T3); best per column in bold\. RWEFM and WFM use the sampled OT map\.
### D\.6de novo generation of Single\-Cell Samples on Spherical spaces
We demonstrate the ability of RWEFM to generate de\-novo single\-cell samples in the latent space of SCimilarity\[[Heimberg et al\., 2025](https://arxiv.org/html/2609.25659#bib.bib39)\], which is a foundation model for single\-cell data that operates in𝕊128−1\\mathbb\{S\}^\{128\-1\}\. SCimilarity was used to embed a single\-cell atlas from healthy blood samples of1,2001,200human donors, each consisting of5,0005,000cells profiled with20,00020,000genes across various patient conditions\. For benchmarking \(as in Tables[S4](https://arxiv.org/html/2609.25659#A4.T4)&[4](https://arxiv.org/html/2609.25659#S4.T4)\), we perform unconditional generation, randomly holding out240240donors and evaluating the quality of240240generated samples against the true held\-out samples\. We note that due to the high dimensionality of the space, we used asampledmap instead of the entropic\. Briefly, instead of assigning each particle in the source distribution to a weighted average of all particles in the target distribution, we assign each particle to a single particle in the target distribution based on the optimal transport plan\. We believe the curse of dimensionality makes the entropic map less effective in this setting, as the entropic map points to unrealistic barycenters of target samples, whereas the sampled map points to actual target samples\.
Furthermore, we demonstrate class\-conditional generation by training a flow to cell distribution conditioned on tissue status \(healthy vs diseased\) on an expanded dataset which included pathological samples\. In[fig\.4](https://arxiv.org/html/2609.25659#S4.F4), we show that generated samples match the profiles of true samples, and display a marked shift in cellular composition between healthy and diseased samples\. This in turn demonstrates the value of using a foundation model, as opposed to Euclidean generation on raw gene expression data or other lower\-dimensional Euclidean embeddings\. Since RWEFM operates directly in the spherical latent space of SCimilarity, it can leverage the functionality that the foundation model provides, such as accurate cell typing and batch effect robust embeddings\.
In Figure[S1](https://arxiv.org/html/2609.25659#A4.F1), we present examples of individual samples generated with RWEFM for different tissue and disease combinations\. For each generated sample, we match it with the closest sample in the real data by Riemannian Wasserstein distance\. We observe that RWEFM can generate realistic tissue samples, showing great potential for downstream biological applications\.
### D\.7Protein Torsion Angle Generation on Torus
To generate the underlying data, we utilized the MD\-CATH dataset\[[Mirarchi et al\., 2024](https://arxiv.org/html/2609.25659#bib.bib7)\], which contains Molecular Dynamics simulations for domain structures from the CATH database\. For each protein in the dataset, we extracted the backbone torsion angles \(ϕ,ψ\\phi,\\psi\) for every residue at every time step of the simulation\. We then aggregated these angle pairs across all residues and time points into a single collection for each protein\. Sinceϕ\\phiandψ\\psiare periodic, this process effectively converts each protein into a single empirical distribution over the flat torus𝕋2\\mathbb\{T\}^\{2\}\.
We performed conditional generation by embedding the specific amino acid sequence of each protein using the ESM\-2 protein language model to obtain a conditioning vector\. We randomly withheld 20% of the proteins as a test set\. For these test proteins, we generated their torsion angle distributions \(point clouds of sizeN=2048N=2048\) conditioned on their ESM embeddings \([fig\.6](https://arxiv.org/html/2609.25659#S4.F6)\)\. Finally, we compared the generated distributions to the ground\-truth distributions derived from the MD simulations using standard distributional distance metrics on the torus \([table4](https://arxiv.org/html/2609.25659#S4.T4)\)\.
Figure S3:Entropic map reconstruction errorFor the spherical grid example from[fig\.3](https://arxiv.org/html/2609.25659#S3.F3), we benchmark the map quality of entropic and euclidean map estimators as a function of sample size and regularization strength\. We report the \(Spherical\) EMD between the source samples pushed forward by the estimated map and the target samples\.
### D\.8Benchmarking of Riemannian Entropic Map
In this manuscript we introduced the Riemannian analogue of the entropic optimal transport map estimator\. In[fig\.S2](https://arxiv.org/html/2609.25659#A4.F2), we benchmark the accuracy and computational efficiency of this estimator against the Euclidean entropic map and the unregularized OT map on constructed examples on the sphere𝕊2\\mathbb\{S\}^\{2\}and hyperbolic spaceℍ2\\mathbb\{H\}^\{2\}\. The ground\-truth map is known in these examples, as it arises from shifting an original distribution via the Riemannian gradient of a convex function over each space\. In both cases, the source is a \(projected\) Gaussian distribution centered at the north pole, and the target is obtained by applying the ground\-truth map to the source samples\. We vary the number of samples and the entropic regularization strength, and report the estimation error of each method in terms of the \(spherical/hyperbolic\) tangent norm between the estimated and ground\-truth map velocities\.
Next, we ask how does the quality of the pushed forward distribution vary as a function of map estimator, sample size and regularization strength\. In[fig\.S3](https://arxiv.org/html/2609.25659#A4.F3), we report the \(spherical\) EMD between the source samples pushed forward by the estimated map and the target samples, for varying sample sizes and regularization strengths\. Here too the source sample is a \(projected\) Gaussian centered at the north pole, and the target is the gridded sphere as shown in[fig\.3](https://arxiv.org/html/2609.25659#S3.F3)\. In each experiment, we estimate the map from a limited number of samples, and use the out\-of\-sample extension of map to push forward the entire \(n=15,000n=15,000\) set of source samples\. We see that the Riemannian entropic map consistently outperforms the Euclidean entropic map across sample sizes and regularization strengths\.
### D\.9Training Time versus Generation Quality Tradeoff
A key practical consideration when applying RWEFM is the tradeoff between computational cost and generation quality\. To provide users with concrete guidance on this tradeoff, we conduct a comprehensive ablation study on the two primary hyperparameters that control both training efficiency and sample quality: the entropic regularization parameterε\\varepsilonand the number of particlesnnsampled from each distribution during training\.
Figure S4:Training time versus generation quality tradeoff\.We benchmark RWEFM on generating MNIST digit 3 on𝕊2\\mathbb\{S\}^\{2\}, varying the entropic regularization parameterε\\varepsilonand number of particlesnnper distribution \. The quantity plotted on both vertical axes is the classwise 1\-NN deviation \(1\-NN\-D\), computed using Chamfer Distance and Earth Mover’s Distance and shown against total training time in seconds\. Lower scores indicate smaller classwise deviations from 50% accuracy; the minimum is 0\. Asε\\varepsilondecreases andnnincreases, training time grows substantially due to increased Sinkhorn iterations required for convergence and larger network capacity needed to process more particles, but generation quality improves markedly as the estimated optimal transport map becomes more accurate\. This demonstrates the concrete tradeoff between computational cost and sample quality that practitioners can tune based on their application requirements and computational budget\.We benchmark RWEFM on the task of generating MNIST digit 3 on the sphere𝕊2\\mathbb\{S\}^\{2\}, systematically varyingε∈\{0\.0002,0\.002,0\.02,0\.2\}\\varepsilon\\in\\\{0\.0002,0\.002,0\.02,0\.2\\\}andn∈\{32,64,128,256\}n\\in\\\{32,64,128,256\\\}across 16 experimental configurations\. For each configuration, we train the model for 500,000 steps and measure both total wall\-clock training time and generation quality on a held\-out test set\. Generation quality is assessed using classwise 1\-NN deviation \(1\-NN\-D\) with both Chamfer Distance \(CD\) and Earth Mover’s Distance \(EMD\) as ground metrics, where lower values indicate smaller classwise deviations from 50% accuracy\.
The results in[fig\.S4](https://arxiv.org/html/2609.25659#A4.F4)reveal a clear and predictable tradeoff\. Asε\\varepsilondecreases, the entropic optimal transport map becomes less regularized and thus more accurate, requiring more Sinkhorn iterations to converge during training\. Similarly, asnnincreases, the model must process more particles per distribution, necessitating larger kernel sizes in the Sinkhorn algorithm and greater capacity in the transformer feedforward layers\. Both factors increase training time—models withε=0\.0002\\varepsilon=0\.0002andn=256n=256take over 4 hours to train, compared to under an hour forε=0\.2\\varepsilon=0\.2andn=32n=32\.
However, this computational investment yields substantial improvements in generation quality\. The most accurate models use smallε\\varepsilonand largenn, while the fastest models exhibit significantly worse quality\. Interestingly, the relationship is roughly monotonic: intermediate configurations provide intermediate performance on both axes, allowing practitioners to select hyperparameters that balance their computational budget against their quality requirements\. All experiments were conducted on a single NVIDIA B200 GPU, with the reported timings providing practitioners with realistic expectations for their own deployments\. This analysis demonstrates that RWEFM offers a tunable spectrum of performance characteristics\.
### D\.10Additional Metrics and Error Values
These results complement the main\-text benchmarks with additional ground metrics and variability estimates\. We first report the overall MMD comparison and the single\-cell results, then compare the sampled and entropic RWEFM variants across random seeds\. Lower values are better for all metrics; 1\-NN\-D is reported as the classwise deviation from 0\.5\.
##### Benchmark summaries
[TableS3](https://arxiv.org/html/2609.25659#A4.T3)complements the 1\-NN\-D results in[table2](https://arxiv.org/html/2609.25659#S3.T2)with MMD under both CD and EMD ground metrics\. The comparison includes point\-cloud baselines alongside the flow\-matching methods\.
Table S3:MMD on MNIST, EMNIST, and KMNIST\.Extended metrics for[table2](https://arxiv.org/html/2609.25659#S3.T2), using Chamfer Distance \(CD\) and Earth Mover’s Distance \(EMD\)\. Dashes indicate methods not evaluated on𝕋2\\mathbb\{T\}^\{2\}\.
##### Single\-cell sample generation
[TableS4](https://arxiv.org/html/2609.25659#A4.T4)expands the single\-cell benchmark in[table4](https://arxiv.org/html/2609.25659#S4.T4)to both ground metrics\. Reporting 1\-NN\-D alongside MMD summarizes classwise nearest\-neighbor label mixing and kernel\-based distributional similarity\.
Table S4:Extended metrics for scRNA\-seq generation\.Whole\-sample generation on𝕊128−1\\mathbb\{S\}^\{128\-1\}, with mean±\\pmstandard deviation for classwise 1\-NN deviation \(1\-NN\-D\) and MMD\.
##### Variability and transport\-map variants
[TablesS5](https://arxiv.org/html/2609.25659#A4.T5)and[S6](https://arxiv.org/html/2609.25659#A4.T6)report mean±\\pmstandard deviation over three seeds and five samplings per seed, with separate columns for sampled OT assignments and the barycentric entropic map\.
Table S5:1\-NN\-D across manifolds\.Mean±\\pmstandard deviation of the classwise deviation from 0\.5; lower is better\. Dashes indicate unevaluated configurations\.The MMD results below follow the same dataset and method ordering\.
Table S6:MMD across manifolds\.Mean±\\pmstandard deviation under CD and EMD ground metrics; lower is better\. Dashes indicate unevaluated configurations\.相似文章
扩散和流匹配背后的几何:Wasserstein空间中的梯度流和测地线
本文揭示了扩散模型和流匹配是同一Wasserstein几何的两面:扩散遵循自由能梯度流(初值问题),而流匹配遵循Wasserstein测地线(边值问题),它们通过JKO格式统一起来。
扩散、基于分数和流匹配生成模型的统一测度论视角
本预印本提出了一个统一的测度论框架,用于理解扩散、基于分数和流匹配生成模型。它通过连续性/福克-普朗克方程建立了这些方法之间的联系,并分析了它们的采样方案及其理论保证。
基于功能流匹配的量子分布生成建模
提出量子流匹配(Quantum Flow Matching, QFM),一种利用自旋Wigner函数和功能流匹配来学习并生成多量子比特量子分布的生成模型,能够准确捕捉纯度和纠缠熵等物理性质。
用于可扩展局部生成建模的重整化群流匹配
介绍了重整化群流匹配(RGFM),这是一种生成框架,利用重整化群流进行可扩展的局部生成建模,从而提高全局一致性和计算效率。
利用流匹配捕获非平衡随机系统中的非马尔可夫动力学
本文开发了一种生成式流匹配方法,用于捕获非平衡随机系统中的非马尔可夫动力学,并展示了与马尔可夫基线相比,在Kramers首次通过时间问题上的改进预测。