SE(3)-MeanFlow: Few-Step Protein Backbone Generation on Lie Groups

arXiv cs.LG Papers

Summary

SE(3)-MeanFlow introduces a few-step generative framework for protein backbone generation on Lie groups, extending MeanFlow to SE(3) with closed-form average-velocity training targets and a rectification-based post-training that matches or exceeds flow-matching baselines at reduced sampling steps.

arXiv:2607.27431v1 Announce Type: new Abstract: Generative modeling of protein backbones promises the de novo design of proteins with prescribed structural and functional properties. Existing diffusion and flow-matching models produce high-quality backbones on SE(3)^N, but inference requires numerically integrating an ODE over hundreds of network evaluations, each involving a Lie group exponential map - a bottleneck for high-throughput design campaigns. We introduce SE(3)-MeanFlow, a few-step generative framework that extends MeanFlow from Euclidean space to the Lie group geometry of protein frames. Working natively in the Lie algebra so(3) and in R^3, we derive closed-form average-velocity identities for rotations and translations, giving simulation-free training targets. We further introduce an SE(3) alpha-Flow objective that removes the Jacobian-vector product from the rotation branch and serves as a warm-up stage, after which training switches to a small-t stabilized MeanFlow loss that is used for the remainder of pretraining and for rectification-based post-training. In protein backbone generation, SE(3)-MeanFlow matches or exceeds flow-matching baselines that use several times more sampling steps, and its advantage widens in the few-step regime, where rectification lets it lead at every matched budget - at a modest cost in diversity.
Original Article
View Cached Full Text

Cached at: 07/31/26, 10:02 AM

# SE(3)-MeanFlow: Few-Step Protein Backbone Generation on Lie Groups
Source: [https://arxiv.org/html/2607.27431](https://arxiv.org/html/2607.27431)
Binghang Lu2Yikai Liu3Elaheh Akbari4Soheil Kolouri4Linxuan Wang5Ping He4Shuchan Wang6Ruqi Zhang711footnotemark:1Guang Lin1,2,5,711footnotemark:1 1Department of Mathematics, Purdue University 2School of Electrical and Computer Engineering, Purdue University 3Department of Biochemistry, Purdue University 4Department of Computer Science, Vanderbilt University 5Department of Statistics, Purdue University 6Department of Mathematics and Computer Science, Freie Universität Berlin 7Department of Computer Science, Purdue UniversityCorresponding authors:Guanglin@purdue\.edu, ruqiz@purdue\.edu\.

###### Abstract

Generative modeling of protein backbones promises the de novo design of proteins with prescribed structural and functional properties\. Existing diffusion and flow\-matching models produce high\-quality backbones onSE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}, but inference requires numerically integrating an ODE over hundreds of network evaluations, each involving a Lie group exponential map—a bottleneck for high\-throughput design campaigns\. We introduceSE\(3\)\-MeanFlow, a few\-step generative framework that extends MeanFlow from Euclidean space to the Lie group geometry of protein frames\. Working natively in the Lie algebra𝔰​𝔬​\(3\)\\mathfrak\{so\}\(3\)and inℝ3\\mathbb\{R\}^\{3\}, we derive closed\-form average\-velocity identities for rotations and translations, giving simulation\-free training targets\. We further introduce anSE​\(3\)\\mathrm\{SE\}\(3\)α\\alpha\-Flow objective that removes the Jacobian–vector product from the rotation branch and serves as a warm\-up stage, after which training switches to a small\-ttstabilized MeanFlow loss that is used for the remainder of pretraining and for rectification\-based post\-training\. In protein backbone generation, SE\(3\)\-MeanFlow matches or exceeds flow\-matching baselines that use several times more sampling steps, and its advantage widens in the few\-step regime, where rectification lets it lead at every matched budget—at a modest cost in diversity\.

## 1Introduction

Proteins are one of the basic building blocks of life\. Their complex geometric structure enables specific inter\-molecular interactions that allow for crucial biological functions—acting as catalysts in chemical reactions, transporters for molecules, and mediators of immune responses\. With the emergence of computational techniques, it has become possible to rationally design novel proteins with desired structures that program their functions, opening pathways to solutions for long\-standing global health challenges including influenza\(Strauchet al\.,[2017](https://arxiv.org/html/2607.27431#bib.bib1)\), COVID\-19\(Caoet al\.,[2020](https://arxiv.org/html/2607.27431#bib.bib4)\)and cancer immunotherapy\(Silvaet al\.,[2019](https://arxiv.org/html/2607.27431#bib.bib6)\)\.

A protein backbone can be modeled as a sequence ofNNrigid bodies, one per residue, each associated with a frame under orientation\-preserving rigid transformations —the special Euclidean groupSE​\(3\)\\mathrm\{SE\}\(3\)\(Jumperet al\.,[2021](https://arxiv.org/html/2607.27431#bib.bib7)\)\. The full backbone is thus described by the product groupSE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}, and the problem of de novo protein design reduces to sampling from a learned distribution over this space\. Recent work has made substantial progress on generative modeling overSE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}\. Diffusion\-based methods such as FrameDiff\(Yimet al\.,[2023b](https://arxiv.org/html/2607.27431#bib.bib8)\)and RFDiffusion\(Watsonet al\.,[2023](https://arxiv.org/html/2607.27431#bib.bib5)\)achieve strong designability, while flow matching approaches—FoldFlow\(Boseet al\.,[2024](https://arxiv.org/html/2607.27431#bib.bib9)\)and FrameFlow\(Yimet al\.,[2023a](https://arxiv.org/html/2607.27431#bib.bib37)\)—further improve training stability and flexibility by learning time\-dependent vector fields onSE​\(3\)\\mathrm\{SE\}\(3\)in a simulation\-free manner\. Despite these advances, all existingSE​\(3\)\\mathrm\{SE\}\(3\)generative models share a common bottleneck: inference still requires numerically integrating an ODE over many steps, typically 100–500 network function evaluations \(NFE\)\. OnSE​\(3\)\\mathrm\{SE\}\(3\), each step involves evaluating the network and applying the Lie group exponential map, making inference substantially more expensive than in Euclidean space\. This limits practical deployment in high\-throughput drug discovery pipelines, where millions of candidate structures must be generated per campaign\.

Reducing the number of inference steps has been studied extensively in Euclidean generative modeling\. Consistency models\(Songet al\.,[2023](https://arxiv.org/html/2607.27431#bib.bib11); Song and Dhariwal,[2024](https://arxiv.org/html/2607.27431#bib.bib12)\)and progressive distillation\(Salimans and Ho,[2022](https://arxiv.org/html/2607.27431#bib.bib13)\)compress inference into one or few steps, but require a pretrained teacher or carefully staged training curricula\. MeanFlow\(Genget al\.,[2026a](https://arxiv.org/html/2607.27431#bib.bib35)\)offers a more principled and self\-contained alternative\. Rather than modeling the instantaneous velocityv​\(t,xt\)v\(t,x\_\{t\}\)as in standard flow matching\(Lipmanet al\.,[2022](https://arxiv.org/html/2607.27431#bib.bib14); Tonget al\.,[2023](https://arxiv.org/html/2607.27431#bib.bib20)\), MeanFlow introduces the notion of*average velocity*u​\(s,t,xt\)u\(s,t,x\_\{t\}\)—the mean velocity of the flow trajectory over the interval\[s,t\]\[s,t\]\. A well\-defined identity relatinguuandvvis derived purely from the definition of average velocity via the product rule and the fundamental theorem of calculus\. This identity yields a tractable, simulation\-free training objective that uses only instantaneous velocity as supervision\. At inference, the entire flow path is approximated in a single network evaluationuθ​\(0,1,x1\)u\_\{\\theta\}\(0,1,x\_\{1\}\), enabling 1\-NFE generation without distillation or pretraining\. However, MeanFlow is formulated in Euclidean spaceℝd\\mathbb\{R\}^\{d\}and does not account for the non\-trivial geometry ofSE​\(3\)\\mathrm\{SE\}\(3\)\. Naively lifting MeanFlow to a curved manifold is ill\-posed: velocities at different points along the trajectory live in different tangent spaces, and their average requires parallel transport along the path\.

Riemannian MeanFlow \(denoted as RMF\-PT\)Zhonget al\.\([2026](https://arxiv.org/html/2607.27431#bib.bib15)\)addresses this by extending MeanFlow to general Riemannian manifolds, defining average velocity via parallel transport and deriving a corresponding Riemannian MeanFlow identity\. While principled, this general\-purpose framework does not exploit the specific algebraic structure ofSE​\(3\)\\mathrm\{SE\}\(3\)as a Lie group\. Another concurrent work, Riemannian MeanFlow \(RMF\)Wooet al\.\([2026](https://arxiv.org/html/2607.27431#bib.bib23)\), defines the average velocity in the Lie algebra; we compare with it theoretically in Section[E\.6](https://arxiv.org/html/2607.27431#A5.SS6)and empirically in Section[4](https://arxiv.org/html/2607.27431#S4)\.

In this work, we proposeSE\(3\)\-MeanFlow, a few\-step generative model for de novo protein backbone design grounded in the Lie group structure ofSE​\(3\)\\mathrm\{SE\}\(3\)\. Our central insight is that the decompositionSE​\(3\)≅SO​\(3\)×ℝ3\\mathrm\{SE\}\(3\)\\cong\\mathrm\{SO\}\(3\)\\times\\mathbb\{R\}^\{3\}allows the MeanFlow identity to be derived separately in the Lie algebra𝔰​𝔬​\(3\)\\mathfrak\{so\}\(3\)and inℝ3\\mathbb\{R\}^\{3\}, both admitting closed\-form, simulation\-free training targets without parallel transport\.

Our main contributions are:

- •Integration\-based SE\(3\)\-MeanFlow with theory\.We propose an integration\-based SE\(3\)\-MeanFlow formulation conditioned on the current state\(Rt,xt\)\(R\_\{t\},x\_\{t\}\), and derive closed\-form average\-velocity identities in the Lie algebra𝔰​𝔬​\(3\)\\mathfrak\{so\}\(3\)and inℝ3\\mathbb\{R\}^\{3\}, without parallel transport\. We provide theoretical validation showing that the resulting training objectives are well\-defined and that the corresponding losses are consistent with the intended MeanFlow targets onSE​\(3\)\\mathrm\{SE\}\(3\)\.
- •Stable training algorithms for proteins\.We introduce anSE​\(3\)\\mathrm\{SE\}\(3\)α\\alpha\-Flow objective together with practical numerical stabilization techniques that make training and few\-step generation reliable for protein backbones\.
- •Protein design results\.On the SCOPe benchmark, our approach matches or exceeds flow\-matching baselines that use several times more sampling steps, with the largest gains in the few\-step regime, in both pretraining and post\-training settings, at a modest cost in diversity\.

## 2Background and Notations

### 2\.1Lie\-group notation onSE​\(3\)\\mathrm\{SE\}\(3\)

In this work, theSE​\(3\)\\mathrm\{SE\}\(3\)space is a collection of rigid motions in

SE​\(3\)≅SO​\(3\)×ℝ3,\\mathrm\{SE\}\(3\)\\cong\\mathrm\{SO\}\(3\)\\times\\mathbb\{R\}^\{3\},whereSO​\(3\)\\mathrm\{SO\}\(3\)denotes the space of3×33\\times 3rotation matrices, which forms a Lie group:

SO​\(3\)=\{R∈ℝ3×3:R​R⊤=I,det\(R\)=1\}\.\\mathrm\{SO\}\(3\)=\\\{R\\in\\mathbb\{R\}^\{3\\times 3\}:RR^\{\\top\}=I,\\ \\det\(R\)=1\\\}\.We followBoseet al\.\([2024](https://arxiv.org/html/2607.27431#bib.bib9)\)and use the decoupledSO​\(3\)×ℝ3\\mathrm\{SO\}\(3\)\\times\\mathbb\{R\}^\{3\}geometry when defining metrics and losses\.

A \(continuous\-time\) flow onSO​\(3\)\\mathrm\{SO\}\(3\)can be defined by a curveRt∈SO​\(3\)R\_\{t\}\\in\\mathrm\{SO\}\(3\)satisfying the left\-trivialized ODE

R˙t=Rt​Ωt,Ωt∈𝔰​𝔬​\(3\),\\dot\{R\}\_\{t\}=R\_\{t\}\\,\\Omega\_\{t\},\\qquad\\Omega\_\{t\}\\in\\mathfrak\{so\}\(3\),\(1\)where𝔰​𝔬​\(3\)\\mathfrak\{so\}\(3\)denotes the corresponding Lie algebra

𝔰​𝔬​\(3\)=\{Ω∈ℝ3×3:Ω⊤=−Ω\}\.\\mathfrak\{so\}\(3\)=\\\{\\Omega\\in\\mathbb\{R\}^\{3\\times 3\}:\\Omega^\{\\top\}=\-\\Omega\\\}\.
𝔰​𝔬​\(3\)\\mathfrak\{so\}\(3\)is isomorphic toℝ3\\mathbb\{R\}^\{3\}space, and we use the hat/vee maps\(⋅\)∧,\(⋅\)∨\(\\cdot\)^\{\\wedge\},\(\\cdot\)^\{\\vee\}to identifyℝ3↔𝔰​𝔬​\(3\)\\mathbb\{R\}^\{3\}\\leftrightarrow\\mathfrak\{so\}\(3\)\(e\.g\.,ω∈ℝ3↦ω∧∈𝔰​𝔬​\(3\)\\omega\\in\\mathbb\{R\}^\{3\}\\mapsto\\omega^\{\\wedge\}\\in\\mathfrak\{so\}\(3\)\)\.

Given the angular velocity fieldΩt:\[0,1\]→𝔰​𝔬​\(3\)\\Omega\_\{t\}:\[0,1\]\\to\\mathfrak\{so\}\(3\)along the trajectory, the solution of \([1](https://arxiv.org/html/2607.27431#S2.E1)\) can be written as

Rt=R0​𝒯​exp⁡\(∫0tΩτ​𝑑τ\),R\_\{t\}=R\_\{0\}\\,\\mathcal\{T\}\\exp\\Big\(\\int\_\{0\}^\{t\}\\Omega\_\{\\tau\}\\,d\\tau\\Big\),where𝒯​exp\\mathcal\{T\}\\expcomposes the infinitesimal rotations along time\. Hereexp\\expdenotes the \(matrix\) exponential map from the Lie algebra to the Lie group\. In particular, ifΩτ≡Ω\\Omega\_\{\\tau\}\\equiv\\Omegais constant, then

Rt=R0​exp⁡\(t​Ω\)\.R\_\{t\}=R\_\{0\}\\exp\(t\\Omega\)\.
For a detailed discussion ofSO​\(3\)\\mathrm\{SO\}\(3\)/SE​\(3\)\\mathrm\{SE\}\(3\)flows \(including the closed\-form/geodesic expressions under common metrics\), see Appendix[B](https://arxiv.org/html/2607.27431#A2)111All appendices referenced in this paper are provided in the technical supplement\.\.

### 2\.2Protein Backbone Parametrization

Following the setting inBoseet al\.\([2024](https://arxiv.org/html/2607.27431#bib.bib9)\), we parametrize the protein backbone as a sequence ofNNrigid bodies\. Each residuei∈\[N\]i\\in\[N\]is associated with a frameTi=\(Ri,xi\)∈SE​\(3\)T\_\{i\}=\(R\_\{i\},x\_\{i\}\)\\in\\mathrm\{SE\}\(3\), whereRi∈SO​\(3\)R\_\{i\}\\in\\mathrm\{SO\}\(3\)represents the orientation andxi∈ℝ3x\_\{i\}\\in\\mathbb\{R\}^\{3\}denotes the position of theCαC\_\{\\alpha\}atom\.

The 3D coordinates of the backbone atoms\{N,Cα,C,O\}i\\\{N,C\_\{\\alpha\},C,O\\\}\_\{i\}are recovered by applying the rigid transformationTiT\_\{i\}to a set of idealized coordinates\{N∗,Cα∗,C∗,O∗\}\\\{N^\{\*\},C\_\{\\alpha\}^\{\*\},C^\{\*\},O^\{\*\}\\\}:

\[N,Cα,C,O\]i=Ti∘\[N∗,Cα∗,C∗,O∗\],\[N,C\_\{\\alpha\},C,O\]\_\{i\}=T\_\{i\}\\circ\[N^\{\*\},C\_\{\\alpha\}^\{\*\},C^\{\*\},O^\{\*\}\],\(2\)whereCα∗=\(0,0,0\)C\_\{\\alpha\}^\{\*\}=\(0,0,0\)is fixed at the origin\. For a self\-contained review ofSO​\(3\)\\mathrm\{SO\}\(3\)/SE​\(3\)\\mathrm\{SE\}\(3\)geometry and notation \(including the definitions of\(⋅\)∧\(\\cdot\)^\{\\wedge\}and\(⋅\)∨\(\\cdot\)^\{\\vee\}\), see Appendix[B](https://arxiv.org/html/2607.27431#A2)–[C](https://arxiv.org/html/2607.27431#A3)\.

### 2\.3MeanFlow in Euclidean space

One limitation of the above flow\-matching method is that, to guarantee accuracy, we generally require the step size1T\\frac\{1\}\{T\}to be sufficiently small, equivalently requiring large inference stepsTT\. To address this issue, in Euclidean space,Genget al\.\([2026a](https://arxiv.org/html/2607.27431#bib.bib35),[b](https://arxiv.org/html/2607.27431#bib.bib27)\)propose the MeanFlow method\. In particular, given a probability pathptp\_\{t\}generated by velocityvtv\_\{t\}, it defines the average velocity:

\(t−s\)​vavg​\(s,t,xt\)=∫stv​\(τ,xτ\)​𝑑τ\.\(t\-s\)v^\{\\text\{avg\}\}\(s,t,x\_\{t\}\)=\\int\_\{s\}^\{t\}v\(\\tau,x\_\{\\tau\}\)d\\tau\.
It yields the training loss:

𝔼​\[‖vθavg​\(s,t,xt\)\+\(t−s\)​sg​\(dd​t​v^avg​\(s,t,xt\)\)−v​\(t,xt\)‖2\]\.\\mathbb\{E\}\\left\[\\left\\\|v^\{\\text\{avg\}\}\_\{\\theta\}\(s,t,x\_\{t\}\)\+\(t\-s\)\\,\\text\{sg\}\\\!\\left\(\\frac\{d\}\{dt\}\\hat\{v\}^\{\\text\{avg\}\}\(s,t,x\_\{t\}\)\\right\)\-v\(t,x\_\{t\}\)\\right\\\|^\{2\}\\right\]\.\(3\)where the randomness is given byt,s∼Unif​\(\[0,1\]\),s≤tt,s\\sim\\text\{Unif\}\(\[0,1\]\),s\\leq tand\(x0,x1\)∼q\(x\_\{0\},x\_\{1\}\)\\sim qfor some joint distributionqq\. By default,qqis an independent coupling betweenpd​a​t​ap\_\{data\}andpp​r​i​o​rp\_\{prior\},

dd​t​v^avg​\(s,t,xt\)=∂∂t​vθavg​\(s,t,xt\)\+∇xvθavg​\(s,t,xt\)⋅vt\.\\frac\{d\}\{dt\}\\hat\{v\}^\{\\text\{avg\}\}\(s,t,x\_\{t\}\)=\\frac\{\\partial\}\{\\partial t\}v^\{\\text\{avg\}\}\_\{\\theta\}\(s,t,x\_\{t\}\)\+\\nabla\_\{x\}v^\{\\text\{avg\}\}\_\{\\theta\}\(s,t,x\_\{t\}\)\\cdot v\_\{t\}\.\(4\)In addition,vtv\_\{t\}can be replacedvθavg​\(t,t,xt\)v^\{\\text\{avg\}\}\_\{\\theta\}\(t,t,x\_\{t\}\)inGenget al\.\([2026b](https://arxiv.org/html/2607.27431#bib.bib27)\)\.

Herext=\(1−t\)​x0\+t​x1x\_\{t\}=\(1\-t\)x\_\{0\}\+tx\_\{1\},v​\(t,xt\)=x1−x0v\(t,x\_\{t\}\)=x\_\{1\}\-x\_\{0\}, and sg is the stop\-gradient operator\. Intuitively, the integration is estimated by the parameterized model during the training process; thus, in the inference step, it yields a T\-step generation, whereT∈ℕT\\in\\mathbb\{N\}:

x^s=xt−1T​vθavg​\(s,t,xt\)\\displaystyle\\hat\{x\}\_\{s\}=x\_\{t\}\-\\frac\{1\}\{T\}v^\{\\text\{avg\}\}\_\{\\theta\}\(s,t,x\_\{t\}\)t=1−i/T,s=t−1/T,i=0,1,…​\(T−1\)\.\\displaystyle t=1\-i/T,s=t\-1/T,i=0,1,\.\.\.\(T\-1\)\.\(5\)When we setT=1T=1, it yields a one\-step generation\.

![Refer to caption](https://arxiv.org/html/2607.27431v1/figures/overview_main.png)Figure 1:Few\-step inference with SE\(3\)\-MeanFlow\. The two\-time networkfθf\_\{\\theta\}\(see Appendix[N](https://arxiv.org/html/2607.27431#A14)\) predicts the average velocity over\[s,t\]\[s,t\]rather than the instantaneous velocity attt, so one evaluation advances the state by the full step—a Lie group exponential onSO​\(3\)\\mathrm\{SO\}\(3\), a Euclidean update onℝ3\\mathbb\{R\}^\{3\}—givingNFE=T\\mathrm\{NFE\}=T\. The objective that makes this accurate is Proposition[1](https://arxiv.org/html/2607.27431#Thmproposition1)and its JVP\-freeα\\alpha\-Flow surrogate \(Section[3\.3](https://arxiv.org/html/2607.27431#S3.SS3)\)\.

## 3Method: MeanFlow inSE​\(3\)\\mathrm\{SE\}\(3\)space

SinceSE​\(3\)=SO​\(3\)×ℝ3\\mathrm\{SE\}\(3\)=\\mathrm\{SO\}\(3\)\\times\\mathbb\{R\}^\{3\}, theℝ3\\mathbb\{R\}^\{3\}component reduces to the conventional Euclidean MeanFlow; hence, we focus on introducing the MeanFlow onSO​\(3\)\\mathrm\{SO\}\(3\)below\. Let a distribution path\{pt∈𝒫​\(SO​\(3\)\):t∈\[0,1\]\}\\\{p\_\{t\}\\in\\mathcal\{P\}\(\\mathrm\{SO\}\(3\)\):t\\in\[0,1\]\\\}be generated by a velocity fieldvt:\[0,1\]×SO​\(3\)→𝒯​SO​\(3\)v\_\{t\}:\[0,1\]\\times\\mathrm\{SO\}\(3\)\\to\\mathcal\{T\}\\mathrm\{SO\}\(3\)\. We define the*average velocity*over\[s,t\]\[s,t\]through the time\-ordered exponential

exp⁡\(\(t−s\)​Ωavg​\(s,t,Rt,xt\)\):=𝒯​exp⁡\(∫stΩ​\(τ,Rτ,xτ\)​𝑑τ\)\.\\displaystyle\\exp\\\!\\big\(\(t\-s\)\\,\\Omega^\{\\text\{avg\}\}\(s,t,R\_\{t\},x\_\{t\}\)\\big\):=\\mathcal\{T\}\\exp\\\!\\Big\(\\int\_\{s\}^\{t\}\\Omega\(\\tau,R\_\{\\tau\},x\_\{\\tau\}\)\\,d\\tau\\Big\)\.\(6\)We denote the instantaneous \(body\-frame\) angular velocity \(vector form\) by

ωt:=Ω​\(t,Rt,xt\)∨∈ℝ3\.\\displaystyle\\omega\_\{t\}:=\\Omega\(t,R\_\{t\},x\_\{t\}\)^\{\\vee\}\\in\\mathbb\{R\}^\{3\}\.\(7\)Analogous to Euclidean MeanFlow, differentiating \([6](https://arxiv.org/html/2607.27431#S3.E6)\) gives an identity linking the average velocity, its trajectory derivative, and the instantaneous velocity\.

###### Proposition 1\.

Differentiating \([6](https://arxiv.org/html/2607.27431#S3.E6)\) with respect tottyields

J​\(\(t−s\)​Aavg\)​\(Aavg​\(s,t,Rt,xt\)\+\(t−s\)​dd​t​Aavg\)=ωt,\\displaystyle J\\big\(\(t\-s\)A^\{\\text\{avg\}\}\\big\)\\Big\(A^\{\\text\{avg\}\}\(s,t,R\_\{t\},x\_\{t\}\)\+\(t\-s\)\\tfrac\{d\}\{dt\}A^\{\\text\{avg\}\}\\Big\)=\\omega\_\{t\},\(8\)where​Aavg=\(Ωavg\)∨,\\displaystyle\\text\{where \}A^\{\\text\{avg\}\}=\(\\Omega^\{\\text\{avg\}\}\)^\{\\vee\},dd​t​Aavg=∂∂t​Aavg\+⟨∇RAavg,R˙t⟩\+⟨∇xAavg,x˙t⟩,\\displaystyle\\hskip 17\.00024pt\\frac\{d\}\{dt\}A^\{\\text\{avg\}\}=\\frac\{\\partial\}\{\\partial t\}A^\{\\text\{avg\}\}\+\\langle\\nabla\_\{R\}A^\{\\text\{avg\}\},\\dot\{R\}\_\{t\}\\rangle\+\\langle\\nabla\_\{x\}A^\{\\text\{avg\}\},\\dot\{x\}\_\{t\}\\rangle,\(9\)and the right JacobianJ:ℝ3→ℝ3×3J:\\mathbb\{R\}^\{3\}\\to\\mathbb\{R\}^\{3\\times 3\}is

J​\(A\):=I−1−cos⁡‖A‖2‖A‖22​A∧\+‖A‖2−sin⁡‖A‖2‖A‖23​\(A∧\)2,\\displaystyle J\(A\):=I\-\\frac\{1\-\\cos\\\|A\\\|\_\{2\}\}\{\\\|A\\\|\_\{2\}^\{2\}\}A^\{\\wedge\}\+\\frac\{\\\|A\\\|\_\{2\}\-\\sin\\\|A\\\|\_\{2\}\}\{\\\|A\\\|\_\{2\}^\{3\}\}\(A^\{\\wedge\}\)^\{2\},\(10\)with‖A‖2\\\|A\\\|\_\{2\}theℓ2\\ell\_\{2\}norm ofA∈ℝ3A\\in\\mathbb\{R\}^\{3\}; as‖A‖→0\\\|A\\\|\\to 0,J​\(A\)→IJ\(A\)\\to I\(we use its Taylor expansion there\)\.

This proposition is split into Propositions[3](https://arxiv.org/html/2607.27431#Thmproposition3)and[4](https://arxiv.org/html/2607.27431#Thmproposition4)in Appendix[F\.1](https://arxiv.org/html/2607.27431#A6.SS1), with proofs\.

### 3\.1Model parametrization and training loss

The network exposes two interchangeable heads: an*endpoint*\(xx\-\)head predicting the clean state\(R^0θ,x^0θ\)\(\\hat\{R\}\_\{0\}^\{\\theta\},\\hat\{x\}\_\{0\}^\{\\theta\}\), and an*average\-velocity*\(uu\-\)headAθavg:\[0,1\]×\[0,1\]×SO​\(3\)×ℝ3→ℝ3A\_\{\\theta\}^\{\\text\{avg\}\}:\[0,1\]\\times\[0,1\]\\times\\mathrm\{SO\}\(3\)\\times\\mathbb\{R\}^\{3\}\\to\\mathbb\{R\}^\{3\}andvθavgv\_\{\\theta\}^\{\\text\{avg\}\}, withΩθavg=\(Aθavg\)∧\\Omega\_\{\\theta\}^\{\\text\{avg\}\}=\(A\_\{\\theta\}^\{\\text\{avg\}\}\)^\{\\wedge\}\. The two carry the same information, related by the invertible map

Aθavg=1tlog\(\(R^0θ\)⊤Rt\)∨⇔R^0θ=Rtexp\(−tAθavg∧\),\\displaystyle A\_\{\\theta\}^\{\\text\{avg\}\}=\\tfrac\{1\}\{t\}\\log\\\!\\big\(\(\\hat\{R\}\_\{0\}^\{\\theta\}\)^\{\\top\}R\_\{t\}\\big\)^\{\\vee\}\\\!\\iff\\\!\\hat\{R\}\_\{0\}^\{\\theta\}=R\_\{t\}\\exp\\\!\\big\(\-t\\,A\_\{\\theta\}^\{\\text\{avg\}\\wedge\}\\big\),vθavg=1t​\(xt−x^0θ\)⇔x^0θ=xt−t​vθavg,\\displaystyle v\_\{\\theta\}^\{\\text\{avg\}\}=\\tfrac\{1\}\{t\}\\big\(x\_\{t\}\-\\hat\{x\}\_\{0\}^\{\\theta\}\\big\)\\\!\\iff\\\!\\hat\{x\}\_\{0\}^\{\\theta\}=x\_\{t\}\-t\\,v\_\{\\theta\}^\{\\text\{avg\}\},which, on the constant\-velocity \(xx\-prediction\) geodesic, is consistent with the\(t−s\)\(t\-s\)\-normalized average of \([6](https://arxiv.org/html/2607.27431#S3.E6)\)\. The losses below regress the velocity head; the endpoint head is produced at inference\.

Based on \([8](https://arxiv.org/html/2607.27431#S3.E8)\), we derive two equivalent mean flow training losses:

ℒSO​\(3\)J:=𝔼s<tq​\(R0,R1\)​\[‖ωθ−ωt‖22\],\\displaystyle\\mathcal\{L\}^\{J\}\_\{\\text\{SO\}\(3\)\}:=\\mathbb\{E\}\_\{\\begin\{subarray\}\{c\}s<t\\\\ q\(R\_\{0\},R\_\{1\}\)\\end\{subarray\}\}\\Big\[\\big\\\|\\omega\_\{\\theta\}\-\\omega\_\{t\}\\big\\\|\_\{2\}^\{2\}\\Big\],\(11\)ωθ=sg​\(J​\(\(t−s\)​Aθavg\)\)​\(Aθavg\+\(t−s\)​sg​\(dd​t​Aθavg\)\),\\displaystyle\\quad\\omega\_\{\\theta\}=\\text\{sg\}\\big\(J\(\(t\-s\)A\_\{\\theta\}^\{\\text\{avg\}\}\)\\big\)\\Big\(A\_\{\\theta\}^\{\\text\{avg\}\}\+\(t\-s\)\\,\\text\{sg\}\(\\tfrac\{d\}\{dt\}A\_\{\\theta\}^\{\\text\{avg\}\}\)\\Big\),ℒSO​\(3\)J−1:=𝔼​\[‖Aθavg−sg​\(Aθt​g​t\)‖2\]\\displaystyle\\mathcal\{L\}^\{J^\{\-1\}\}\_\{\\mathrm\{SO\}\(3\)\}:=\\mathbb\{E\}\\left\[\\\|A\_\{\\theta\}^\{\\text\{avg\}\}\-\\text\{sg\}\(A\_\{\\theta\}^\{tgt\}\)\\\|^\{2\}\\right\]\(12\)Aθt​g​t=J−1​\(\(t−s\)​Aθavg\)​ωt−\(t−s\)​dd​t​Aθavg\\displaystyle\\quad A\_\{\\theta\}^\{tgt\}=J^\{\-1\}\(\(t\-s\)A\_\{\\theta\}^\{\\text\{avg\}\}\)\\omega\_\{t\}\-\(t\-s\)\\frac\{d\}\{dt\}A\_\{\\theta\}^\{\\text\{avg\}\}where all model arguments are\(s,t,Rt,xt\)\(s,t,R\_\{t\},x\_\{t\}\)anddd​t​Aθavg\\tfrac\{d\}\{dt\}A\_\{\\theta\}^\{\\text\{avg\}\}is given by \([9](https://arxiv.org/html/2607.27431#S3.E9)\)\. Loss \([12](https://arxiv.org/html/2607.27431#S3.E12)\) is well\-defined sinceJJis invertible on the principal branch‖\(t−s\)​Aθavg‖2≤π\\\|\(t\-s\)A\_\{\\theta\}^\{\\text\{avg\}\}\\\|\_\{2\}\\leq\\pi\. We refer to Appendix[F\.1](https://arxiv.org/html/2607.27431#A6.SS1)for details\.

##### Manifold derivatives via Euclidean JVPs\.

AlthoughRt∈SO​\(3\)R\_\{t\}\\in\\mathrm\{SO\}\(3\)is manifold\-valued,dd​t​Aθavg\\tfrac\{d\}\{dt\}A\_\{\\theta\}^\{\\text\{avg\}\}is a directional derivative along the interpolation trajectory\(Rt,xt\)\(R\_\{t\},x\_\{t\}\), which we evaluate with a standard Euclidean Jacobian–vector product \(e\.g\.torch\.jvp\) through the chosen matrix representation ofRtR\_\{t\}\. A formal equivalence is given in Appendix[F\.1](https://arxiv.org/html/2607.27431#A6.SS1), Proposition[6](https://arxiv.org/html/2607.27431#Thmproposition6)\.

###### Proposition 2\(Correctness of the MeanFlow objective, informal\)\.

IfℒSO​\(3\)\\mathcal\{L\}\_\{\\text\{SO\}\(3\)\}in \([11](https://arxiv.org/html/2607.27431#S3.E11)\) or \([12](https://arxiv.org/html/2607.27431#S3.E12)\) is zero, the model recovers the correct relative rotations along the path:exp⁡\(\(\(t−s\)​Aθavg​\(s,t,Rt,xt\)\)∧\)=Rs⊤​Rt\\exp\\\!\\big\(\(\(t\-s\)A\_\{\\theta\}^\{\\text\{avg\}\}\(s,t,R\_\{t\},x\_\{t\}\)\)^\{\\wedge\}\\big\)=R\_\{s\}^\{\\top\}R\_\{t\}for all0≤s<t≤10\\leq s<t\\leq 1\.

A formal statement and proof are in Proposition[5](https://arxiv.org/html/2607.27431#Thmproposition5)\(Appendix\)\.

##### Inference\.

We step backwards using the learned average flow,

Rs←Rt​exp⁡\(−\(t−s\)​Ωθavg​\(s,t,Rt,xt\)\),t=1,…,1T,R\_\{s\}\\leftarrow R\_\{t\}\\exp\\\!\\big\(\-\(t\-s\)\\,\\Omega\_\{\\theta\}^\{\\text\{avg\}\}\(s,t,R\_\{t\},x\_\{t\}\)\\big\),\\qquad t=1,\\ldots,\\tfrac\{1\}\{T\},and analogouslyxs←xt−\(t−s\)​vθavgx\_\{s\}\\leftarrow x\_\{t\}\-\(t\-s\)\\,v\_\{\\theta\}^\{\\text\{avg\}\}\.See Figure[1](https://arxiv.org/html/2607.27431#S2.F1)\. Reported results use a variant of this update, the exponential rotation schedule of Appendix[I](https://arxiv.org/html/2607.27431#A9)\.

### 3\.2Practical implementation:SE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}adaptation

Each protein sample isRN∈ℝN×3×3R^\{N\}\\in\\mathbb\{R\}^\{N\\times 3\\times 3\},xN∈ℝN×3x^\{N\}\\in\\mathbb\{R\}^\{N\\times 3\}, withNNthe number of residues\. The model isfθ​\(s,t,RtN,xtN\)=\[Aθavg,N;vθavg,N\]∈ℝN×3×ℝN×3f\_\{\\theta\}\(s,t,R^\{N\}\_\{t\},x^\{N\}\_\{t\}\)=\[A\_\{\\theta\}^\{\\text\{avg\},N\};\\,v\_\{\\theta\}^\{\\text\{avg\},N\}\]\\in\\mathbb\{R\}^\{N\\times 3\}\\times\\mathbb\{R\}^\{N\\times 3\}, where theii\-th blockAθ,iavg,N∈ℝ3A\_\{\\theta,i\}^\{\\text\{avg\},N\}\\in\\mathbb\{R\}^\{3\}parametrizes the𝔰​𝔬​\(3\)\\mathfrak\{so\}\(3\)component viaΩθ,iavg=\(Aθ,iavg,N\)∧\\Omega\_\{\\theta,i\}^\{\\text\{avg\}\}=\(A\_\{\\theta,i\}^\{\\text\{avg\},N\}\)^\{\\wedge\}\. The loss \([11](https://arxiv.org/html/2607.27431#S3.E11)\) \(or \([12](https://arxiv.org/html/2607.27431#S3.E12)\)\) is summed over residues\.

### 3\.3SE​\(3\)\\mathrm\{SE\}\(3\)α\\alpha\-Flow: rotation formulation and MeanFlow limit

The differential target \([8](https://arxiv.org/html/2607.27431#S3.E8)\) requires the trajectory derivativedd​t​Aθavg\\tfrac\{d\}\{dt\}A^\{\\text\{avg\}\}\_\{\\theta\}via a JVP, which could be fragile on the rotation branch\. We therefore also adopt anα\\alpha\-Flow formulationZhanget al\.\([2025](https://arxiv.org/html/2607.27431#bib.bib25)\)that replaces this derivative by a two\-segment construction using only forward evaluations ofAθavgA^\{\\text\{avg\}\}\_\{\\theta\}\.

Given0≤s<t≤10\\leq s<t\\leq 1and a ratioα∈\(0,1\]\\alpha\\in\(0,1\], insert an intermediate timem=α​s\+\(1−α\)​tm=\\alpha s\+\(1\-\\alpha\)twith stepδ:=t−m=α​\(t−s\)\\delta:=t\-m=\\alpha\(t\-s\), splitting\[s,t\]\[s,t\]into a*far*segment\[s,m\]\[s,m\]and a*near*segment\[m,t\]\[m,t\]\. WithD​\(a,b\):=Ra⊤​Rb∈SO​\(3\)D\(a,b\):=R\_\{a\}^\{\\top\}R\_\{b\}\\in\\mathrm\{SO\}\(3\)the accumulated relative rotation, the segments compose*multiplicatively*\(rather than additively as in Euclidean space\),

D​\(s,t\)=D​\(s,m\)​D​\(m,t\)\(Appendix[J](https://arxiv.org/html/2607.27431#A10), Prop\.[7](https://arxiv.org/html/2607.27431#Thmproposition7)\)\.\\displaystyle D\(s,t\)=D\(s,m\)\\,D\(m,t\)\\qquad\\text\{\(Appendix~\\ref\{sec:alpha\-flow\}, Prop\.~\\ref\{prop:af\-additivity\}\)\}\.We anchor the near factor to data and bootstrap the far factor from the model\. Stepping back fromRtR\_\{t\}along the data angular velocityωt:=Ω​\(t,Rt,xt\)∨\\omega\_\{t\}:=\\Omega\(t,R\_\{t\},x\_\{t\}\)^\{\\vee\}gives the intermediate stateRm=Rt​exp⁡\(−δ​ωt∧\)R\_\{m\}=R\_\{t\}\\exp\(\-\\delta\\,\\omega\_\{t\}^\{\\wedge\}\)and the near factorD​\(m,t\)=exp⁡\(δ​ωt∧\)D\(m,t\)=\\exp\(\\delta\\,\\omega\_\{t\}^\{\\wedge\}\); a stop\-gradient model query at\(s,m,Rm,xm\)\(s,m,R\_\{m\},x\_\{m\}\)returnsAm:=Aθavg​\(s,m,Rm,xm\)A\_\{m\}:=A^\{\\text\{avg\}\}\_\{\\theta\}\(s,m,R\_\{m\},x\_\{m\}\)and the far factorD​\(s,m\)=exp⁡\(\(m−s\)​Am∧\)D\(s,m\)=\\exp\(\(m\-s\)A\_\{m\}^\{\\wedge\}\)\. Composing and mapping back to the Lie algebra gives the target average generator over\[s,t\]\[s,t\]:

Atgtavg\(s,t\)=1t−slog\(exp⁡\(\(m−s\)​Am∧\)⏟D​\(s,m\)​\(model\)exp⁡\(δ​ωt∧\)⏟D​\(m,t\)​\(data\)\)∨\.\\displaystyle A^\{\\text\{avg\}\}\_\{\\text\{tgt\}\}\(s,t\)=\\frac\{1\}\{t\-s\}\\,\\log\\\!\\Big\(\\underbrace\{\\exp\\\!\\big\(\(m\-s\)A\_\{m\}^\{\\wedge\}\\big\)\}\_\{D\(s,m\)\\ \(\\text\{model\}\)\}\\;\\underbrace\{\\exp\\\!\\big\(\\delta\\,\\omega\_\{t\}^\{\\wedge\}\\big\)\}\_\{D\(m,t\)\\ \(\\text\{data\}\)\}\\Big\)^\{\\vee\}\.\(13\)The scalar1t−s\\tfrac\{1\}\{t\-s\}must stay outside thelog⁡\(exp⋅exp\)\\log\(\\exp\\cdot\\exp\): sinceAmA\_\{m\}andωt\\omega\_\{t\}do not commute, folding it in would corrupt the Baker–Campbell–Hausdorff cross term and no longer give the generator ofD​\(s,t\)D\(s,t\)\. We regress

ℒrotα=1α​∑i=1N‖Aθ,iavg​\(s,t,Rt,xt\)−sg​\(Atgt,iavg\)‖22,\\displaystyle\\mathcal\{L\}^\{\\alpha\}\_\{\\text\{rot\}\}=\\frac\{1\}\{\\alpha\}\\sum\_\{i=1\}^\{N\}\\big\\\|A^\{\\text\{avg\}\}\_\{\\theta,i\}\(s,t,R\_\{t\},x\_\{t\}\)\-\\mathrm\{sg\}\\big\(A^\{\\text\{avg\}\}\_\{\\text\{tgt\},i\}\\big\)\\big\\\|\_\{2\}^\{2\},\(14\)which uses only forward evaluations \(one log, two exps; no JVP\)\. The numerically stabilized form and the abelian translation branch — where composition reduces to a convex combination of velocities as in the Euclideanα\\alpha\-FlowZhanget al\.\([2025](https://arxiv.org/html/2607.27431#bib.bib25)\)— are given in Appendix[J\.3](https://arxiv.org/html/2607.27431#A10.SS3)\.

The ratioα\\alphainterpolates between flow matching and MeanFlow: atα=1\\alpha=1the far factor vanishes \(m=sm=s\) andAtgtavg=ωtA^\{\\text\{avg\}\}\_\{\\text\{tgt\}\}=\\omega\_\{t\}, while asα→0\\alpha\\to 0a first\-order BCH expansion recovers the differential MeanFlow objective \([8](https://arxiv.org/html/2607.27431#S3.E8)\) in gradient \(Appendix[J\.5](https://arxiv.org/html/2607.27431#A10.SS5)\)\.

## 4Experiments

We evaluate our method on unconditional protein backbone generation and compare it against recent diffusion\- and flow\-based baselines\. Our experiments are designed to assess both generation quality and sampling efficiency, with a particular focus on few\-step generation\. To this end, we report performance under different numbers of sampling steps, allowing us to examine how well each method maintains designability, diversity, and novelty as the computational budget is reduced\.

### 4\.1Datasets

Following prior worksYueet al\.\([2025](https://arxiv.org/html/2607.27431#bib.bib22)\); Wooet al\.\([2026](https://arxiv.org/html/2607.27431#bib.bib23)\), we conduct experiments on the SCOPe datasetChandoniaet al\.\([2022](https://arxiv.org/html/2607.27431#bib.bib26)\); Yimet al\.\([2023a](https://arxiv.org/html/2607.27431#bib.bib37)\), which consists of 3,673 preprocessed protein backbones with residue lengths between 60 and 128\. We refer to Figure[9](https://arxiv.org/html/2607.27431#A14.F9)in the Appendix for a visualization of the backbone length\-frequency distribution\.

![Refer to caption](https://arxiv.org/html/2607.27431v1/x1.png)

![Refer to caption](https://arxiv.org/html/2607.27431v1/x2.png)

Figure 2:Designable fraction versus length \(top:T=100T\{=\}100, bottom:T=20T\{=\}20\)\.![Refer to caption](https://arxiv.org/html/2607.27431v1/x3.png)

![Refer to caption](https://arxiv.org/html/2607.27431v1/x4.png)

Figure 3:scRMSD distributions \(top:T=100T\{=\}100, bottom:T=20T\{=\}20\)\.StepsMethodDesignabilityDiversityTM↓\\mathrm\{TM\}\\downarrowNoveltyTM↓\\mathrm\{TM\}\\downarrowFraction↑\\uparrowscRMSD↓\\downarrowscTM↑\\uparrow500 / 100FrameFlow \(500\)0\.8491\.439±1\.1371\.439\\pm 1\.1370\.879±0\.0840\.879\\pm 0\.0840\.3690\.654FrameFlow \(100\)0\.8031\.576±1\.3671\.576\\pm 1\.3670\.872±0\.0900\.872\\pm 0\.0900\.3600\.638QFlow \(500\)0\.9001\.271±1\.1131\.271\\pm 1\.1130\.897±0\.0780\.897\\pm 0\.0780\.3990\.720QFlow \(100\)0\.8851\.319±1\.0181\.319\\pm 1\.0180\.889±0\.0770\.889\\pm 0\.0770\.3930\.699RMF \(100\)0\.8321\.425±1\.1451\.425\\pm 1\.1450\.885±0\.0810\.885\\pm 0\.0810\.3430\.719SE3MF\(100\)0\.9361\.106±0\.870\\mathbf\{1\.106\}\\pm 0\.8700\.909±0\.067\\mathbf\{0\.909\}\\pm 0\.0670\.4150\.73650QFlow0\.8701\.432±1\.2181\.432\\pm 1\.2180\.882±0\.0800\.882\\pm 0\.0800\.3790\.684RMF0\.8251\.446±1\.1891\.446\\pm 1\.1890\.885±0\.0810\.885\\pm 0\.0810\.3440\.717SE3MF0\.9061\.246±1\.190\\mathbf\{1\.246\}\\pm 1\.1900\.902±0\.078\\mathbf\{0\.902\}\\pm 0\.0780\.4130\.72220QFlow0\.7781\.703±1\.6581\.703\\pm 1\.6580\.853±0\.0980\.853\\pm 0\.0980\.3770\.648RMF0\.8061\.554±1\.4381\.554\\pm 1\.4380\.877±0\.0910\.877\\pm 0\.0910\.3450\.722SE3MF0\.8671\.369±1\.144\\mathbf\{1\.369\}\\pm 1\.1440\.883±0\.087\\mathbf\{0\.883\}\\pm 0\.0870\.4080\.70310QFlow0\.5592\.691±2\.2632\.691\\pm 2\.2630\.773±0\.1440\.773\\pm 0\.1440\.3790\.615RMF0\.7781\.641±1\.432\\mathbf\{1\.641\}\\pm 1\.4320\.870±0\.088\\mathbf\{0\.870\}\\pm 0\.0880\.3460\.713SE3MF0\.7281\.940±1\.7411\.940\\pm 1\.7410\.832±0\.1180\.832\\pm 0\.1180\.4020\.659Table 1:Unconditional protein backbone generation on SCOPe, grouped by sampling budget\. Baselines: FrameFlowYimet al\.\([2023a](https://arxiv.org/html/2607.27431#bib.bib37)\), QFlowYueet al\.\([2025](https://arxiv.org/html/2607.27431#bib.bib22)\), Riemannian MeanFlow \(RMF\)Wooet al\.\([2026](https://arxiv.org/html/2607.27431#bib.bib23)\)\. Within each budget group, best inboldand second\-bestunderlined\.Table 2:Post\-train comparison on SCOPe\.Within each budget, best inbold\. Training budgets are in Table[4](https://arxiv.org/html/2607.27431#S4.T4)\.
### 4\.2Baselines

We compare recent flow\-based backbone generators: FrameFlowYimet al\.\([2023a](https://arxiv.org/html/2607.27431#bib.bib37)\), QFlow and ReQFlowYueet al\.\([2025](https://arxiv.org/html/2607.27431#bib.bib22)\), and Riemannian MeanFlow \(RMF\)Wooet al\.\([2026](https://arxiv.org/html/2607.27431#bib.bib23)\)\. Among these, ReQFlow and RMF are most closely related to our approach, as they also target efficient few\-step or accelerated generation\. As a representative of earlier \(pre\-2024\) diffusion\- and flow\-matching methods we include FrameFlow, which demonstrated strongest generation capacity on SCOPe; we discuss the earlier baselines \(e\.g\. FrameDiff, Genie\) in Appendix[E\.1](https://arxiv.org/html/2607.27431#A5.SS1)\.

### 4\.3Implementation Details

For experiments on SCOPe, we use publicly available checkpoints when possible to reproduce baseline results under a consistent evaluation protocol\. All experiments are conducted using 4 H\-100 GPUs\. Unless otherwise stated, we follow the evaluation settings used in prior workYueet al\.\([2025](https://arxiv.org/html/2607.27431#bib.bib22)\), including the same datasets, sampling protocols, and evaluation metrics, to ensure a fair comparison across methods\.

### 4\.4Training details

##### Model and trainer\.

We implement our method within the public QFlow/ReQFlow codebaseYueet al\.\([2025](https://arxiv.org/html/2607.27431#bib.bib22)\)and inherit its overall training pipeline, adapting it in the following respects to fit the MeanFlow model\. *\(i\) Representation and interpolation\.*QFlow/ReQFlow parametrize rotations by unit quaternions; we instead work directly with rotation matrices under the decoupledSE​\(3\)=SO​\(3\)×ℝ3\\mathrm\{SE\}\(3\)=\\mathrm\{SO\}\(3\)\\times\\mathbb\{R\}^\{3\}representation ofYimet al\.\([2023b](https://arxiv.org/html/2607.27431#bib.bib8)\)\. Accordingly, we build the data–noise interpolation and the mini\-batch coupling with our ownSE​\(3\)\\mathrm\{SE\}\(3\)geodesic interpolant and optimal\-transport \(OT\) coupling in rotation\-matrix form \(Appendix[D](https://arxiv.org/html/2607.27431#A4); global\-OT coupling in Eq\. \([95](https://arxiv.org/html/2607.27431#A13.E95)\)\), rather than the quaternion interpolation of QFlow\. *\(ii\) Model\.*We keep the ReQFlow IPA trunk but make it consume a rotation\-matrix state\(Rt,xt\)∈SE​\(3\)N\(R\_\{t\},x\_\{t\}\)\\in\\mathrm\{SE\}\(3\)^\{N\}and condition on*two*times\(s,t\)\(s,t\)—via a shared two\-time embedding and a per\-block AdaLN\-Zero gate—and predict endpoints\(R^0θ,x^0θ\)\(\\hat\{R\}\_\{0\}^\{\\theta\},\\hat\{x\}\_\{0\}^\{\\theta\}\)\(Appendix[N](https://arxiv.org/html/2607.27431#A14)\)\. *\(iii\) Objective\.*In place of the \(V\-\)QFlow flow\-matching loss, we train with ourSE​\(3\)\\mathrm\{SE\}\(3\)MeanFlow objective \(Section[3](https://arxiv.org/html/2607.27431#S3)\) and itsα\\alpha\-Flow variant \(Section[3\.3](https://arxiv.org/html/2607.27431#S3.SS3)\); derivations and the stable implementation are in Appendix[F\.1](https://arxiv.org/html/2607.27431#A6.SS1)and[M](https://arxiv.org/html/2607.27431#A13)\. The time sampler is the two\-time extension \(s≤ts\\leq t\) of the QFlow sampler, with the marginal schedule ofttleft unchanged, ands∣t∼𝒰​\[tmin,t\]s\\mid t\\sim\\mathcal\{U\}\[t\_\{\\min\},t\]withtmin=10−6t\_\{\\min\}=10^\{\-6\}\(Table[7](https://arxiv.org/html/2607.27431#A13.T7)\)\. All remaining pipeline components are inherited from the QFlow/ReQFlow codebase\.

##### Stable training\.

The differential MeanFlow target requires a time derivative obtained via a Jacobian–vector product \(JVP\), which can be numerically brittle through the backbone network\. We use two remedies: the JVP\-freeα\\alpha\-Flow objective \(Section[3\.3](https://arxiv.org/html/2607.27431#S3.SS3)\) as a warm\-up, and a small\-ttstabilized form of the JVP\-based loss \(Appendix[I](https://arxiv.org/html/2607.27431#A9), Algorithm[3](https://arxiv.org/html/2607.27431#alg3)\)\.

##### Pre\-training \(two stages\)\.

Stage 1 is theα\\alpha\-Flow warm\-up onSE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}\(Section[3\.3](https://arxiv.org/html/2607.27431#S3.SS3); Appendix[M\.1](https://arxiv.org/html/2607.27431#A13.SS1)\); Stage 2 switches to the endpoint\+\+MeanFlow objective \(Appendix[M\.2](https://arxiv.org/html/2607.27431#A13.SS2)\)\.

##### Post\-training\.

We further apply a rectification \(self\-reflow\) stage following ReQFlow’s rectified\-flow strategy, keeping the Stage\-2 MeanFlow objective \(Appendix[M\.3](https://arxiv.org/html/2607.27431#A13.SS3)\)\. We denote the resulting model RecSE3MF\.

Table 3:Evaluation scores at different sampling steps\. Methods follow Table[1](https://arxiv.org/html/2607.27431#S4.T1)and Table[2](https://arxiv.org/html/2607.27431#S4.T2)\.Table 4:Model size, training budget, and sampling cost\. Post\-training steps are given as pre\-training\+\+rectification\. Time/step is wall\-clock per sampling step on a single NVIDIA H100 80GB \(single process, batch1010, length128128, fp32\); each method runs one network forward per step, so total sampling time≈\\approxsteps×\\timestime/step\.

### 4\.5Evaluation metrics and settings

We evaluate generated protein backbones with four metrics, following prior work\(Yueet al\.,[2025](https://arxiv.org/html/2607.27431#bib.bib22)\): designability, diversity, novelty, and efficiency\. For each chain lengthNN\(from6060to128128\) we generate1010backbones\. Designability is the primary measure of sample quality and assesses whether a generated backbone admits an amino\-acid sequence that folds back into a consistent structure\. For each backbone we design88sequences with ProteinMPNN\(Dauparaset al\.,[2022](https://arxiv.org/html/2607.27431#bib.bib2)\), predict the folded structure of each with ESMFold\(Linet al\.,[2023](https://arxiv.org/html/2607.27431#bib.bib39)\), and take the minimum self\-consistency RMSD \(scRMSD\) over the88designs\. A backbone is deemed*designable*when this scRMSD is at most2​Å2\\,\\text\{\\AA \}\. We report the fraction of designable backbones \(per length, averaged over lengths\), denoted Fraction, the mean scRMSD, and the mean self\-consistency TM\-score \(scTM\), where higher scTM and Fraction and lower scRMSD are better\.

To assess distributional properties, we measure structural diversity and novelty over the designable subset\. Diversity is the pairwise TM\-score\(Zhang and Skolnick,[2004](https://arxiv.org/html/2607.27431#bib.bib41)\)among designable structures of the same length, averaged across lengths, where lower values indicate a less redundant set\. Novelty compares each designable sample against the Protein Data Bank \(PDB\) with Foldseek\(van Kempenet al\.,[2022](https://arxiv.org/html/2607.27431#bib.bib40)\)and records the maximum TM\-score to the retrieved structures; the average of these maxima summarizes similarity to known proteins, with lower values indicating greater novelty\. Finally, we quantify efficiency by the number of sampling steps used to generate the backbones\.

### 4\.6Results and Discussion

![Refer to caption](https://arxiv.org/html/2607.27431v1/x5.png)Figure 4:Qualitative samples on SCOPe generated at different sampling budgets\. See Appendix for additional diagnostics\.SE\(3\)\-MeanFlow \(SE3MF\) consistently improves few\-step protein backbone generation\. Across the 20–100\-step regime, it achieves the strongest designability among pretrained baselines, ranking first in designable fraction, scRMSD, and scTM \(Table[1](https://arxiv.org/html/2607.27431#S4.T1)\)\. The gains are preserved across protein lengths \(Figures[2](https://arxiv.org/html/2607.27431#S4.F2)and[3](https://arxiv.org/html/2607.27431#S4.F3)\), indicating that the improvement is not restricted to short or structurally simple backbones\. At 20 steps, SE\(3\)\-MeanFlow retains a designable fraction of0\.8670\.867, compared with0\.8060\.806for RMF and0\.7780\.778for QFlow, demonstrating a strong designability–efficiency trade\-off under substantial step reduction; Figure[4](https://arxiv.org/html/2607.27431#S4.F4)shows designable backbones generated at this budget\.

These results support the effectiveness of specializing MeanFlow to the Lie\-group structure of protein frames\. Empirically, the complete formulation maintains strong designability as the sampling budget is reduced while preserving stable local backbone geometry, withC​α\\mathrm\{C\}\\alpha\-validity remaining near0\.970\.97from 100 to 10 steps \(Table[3](https://arxiv.org/html/2607.27431#S4.T3)\)\. The secondary\-structure statistics are likewise stable across sampling budgets \(Figure[10](https://arxiv.org/html/2607.27431#A15.F10)\), suggesting that accelerated generation does not substantially alter the structural composition of the samples\.

Rectification further strengthens the aggressive few\-step regime\. RecSE3MF outperforms ReQFlow on the three designability metrics at nearly all sampling budgets and reaches a designable fraction of0\.8940\.894at 10 steps \(Table[2](https://arxiv.org/html/2607.27431#S4.T2)\)\. This suggests that rectification and MeanFlow play complementary roles: rectification simplifies the transport paths, while MeanFlow learns accurate finite\-interval motion along those paths\.

The improved designability comes at a cost in coverage\. Both SE\(3\)\-MeanFlow and RecSE3MF have higher diversity TM\-scores than the coverage\-oriented baselines \(lower is better\), and their novelty is weakest at large budgets \(0\.7360\.736at100100steps, against0\.7190\.719for RMF\); the gap narrows as the budget falls, with novelty second\-best in its group at2020and1010steps\. Diversity is nearly flat in the number of steps \(0\.4150\.415to0\.4020\.402\), so the concentration reflects the learned model rather than step reduction\. Overall, SE\(3\)\-MeanFlow advances the designability–efficiency frontier while preserving stable geometric and structural statistics under substantially reduced sampling budgets\.

##### Training budget and model size\.

Table[4](https://arxiv.org/html/2607.27431#S4.T4)compares model size, training budget and per\-step sampling cost\. Our network is the ReQFlow/QFlow trunk with a small \(0\.33%0\.33\\%\) AdaLN addition, hence comparable in size to QFlow and checkpoint\-compatible with it \(Appendix[N](https://arxiv.org/html/2607.27431#A14)\), and∼26×\\sim\\\!26\\timessmaller than RMF; its total budget is on the same order as the flow\-matching baselines from the same codebase, and far below RMF’s∼598\\sim\\\!598k steps \(their cap is10001000k\)\. We conjecture that part of this gap reflects RMF’s semigroup consistency objective rather than its model size alone: straightening the transport paths enough for few\-step sampling appears to demand a large budget, and this cost grows with the number of residues\. Appendix[O\.2](https://arxiv.org/html/2607.27431#A15.SS2)tests this by fixing the model, the initialization and the training budget, and swapping only the objective; under this control the semigroup target does not reach the few\-step regime while ours does\.

On the low\-dimensionalSO​\(3\)2\\mathrm\{SO\}\(3\)^\{2\}benchmark of Appendix[H](https://arxiv.org/html/2607.27431#A8), a second controlled ablation in which only the loss varies, every average\-velocity method reaches the few\-step regime under a much smaller budget\.

## 5Conclusion

We introducedSE\(3\)\-MeanFlow, a few\-step generative framework onSE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}\. Exploiting the decoupledSE​\(3\)=SO​\(3\)×ℝ3\\mathrm\{SE\}\(3\)=\\mathrm\{SO\}\(3\)\\times\\mathbb\{R\}^\{3\}structure, we derive simulation\-free average\-velocity objectives directly in the Lie algebra, avoiding parallel transport, and stabilize training with a JVP\-freeα\\alpha\-Flow warm\-up and a small\-ttMeanFlow scheme supporting both training from scratch and rectification\-based post\-training\. On SCOPe, SE\(3\)\-MeanFlow advances the designability–efficiency frontier while using fewer sampling steps and a smaller model\. Its main limitation is a modest reduction in diversity; future work will seek objectives that better balance coverage and designability, extend to conditional and motif\-scaffolding tasks, and push toward one\-step generation\.

## Acknowledgments

Y\.B\. was supported in part by the Purdue Institute for Physical AI \(IPAI\) Postdoctoral Fellows Program\. L\.G\. would like to thank the support of National Science Foundation \(DMS\-2533878, DMS\-2053746, DMS\-2134209, ECCS\-2328241, CBET\-2347401 and OAC\-2311848\), and U\.S\. Department of Energy \(DOE\) Office of Science Advanced Scientific Computing Research program, under the ”Uncertainty Quantification for Multifidelity Operator Learning \(MOLUcQ\)” project \(Project No\. 81739\), DE\-SC0023161, the SciDAC LEADS Institute, and DOE–Fusion Energy Science, under grant number: DE\-SC0024583\.

## References

- A\. J\. Bose, T\. Akhound\-Sadegh, G\. Huguet, K\. Fatras, J\. Rector\-Brooks, C\. Liu, A\. C\. Nica, M\. Korablyov, M\. Bronstein, and A\. Tong \(2024\)SE\(3\)\-stochastic flow matching for protein backbone generation\.InThe Twelfth International Conference on Learning Representations,Cited by:[§C\.3](https://arxiv.org/html/2607.27431#A3.SS3.p1.7),[§D\.2](https://arxiv.org/html/2607.27431#A4.SS2.p1.4),[§D\.3](https://arxiv.org/html/2607.27431#A4.SS3.p1.6),[§E\.2](https://arxiv.org/html/2607.27431#A5.SS2),[§E\.2](https://arxiv.org/html/2607.27431#A5.SS2.p1.5),[§G\.1](https://arxiv.org/html/2607.27431#A7.SS1.p1.11),[Table 5](https://arxiv.org/html/2607.27431#A8.T5.14.19.5.1),[§1](https://arxiv.org/html/2607.27431#S1.p2.7),[§2\.1](https://arxiv.org/html/2607.27431#S2.SS1.p1.4),[§2\.2](https://arxiv.org/html/2607.27431#S2.SS2.p1.6),[Remark 2](https://arxiv.org/html/2607.27431#Thmremark2.p1.6)\.
- L\. Cao, I\. Goreshnik, B\. Coventry, J\. B\. Case, L\. Miller, L\. Kozodoy, R\. E\. Chen, L\. Carter, A\. C\. Walls, Y\. Park,et al\.\(2020\)De novo design of picomolar sars\-cov\-2 miniprotein inhibitors\.Science370\(6515\),pp\. 426–431\.Cited by:[§1](https://arxiv.org/html/2607.27431#S1.p1.1)\.
- J\. Chandonia, L\. Guan, S\. Lin, C\. Yu, N\. K\. Fox, and S\. E\. Brenner \(2022\)SCOPe: improvements to the structural classification of proteins–extended database to facilitate variant interpretation and machine learning\.Nucleic acids research50\(D1\),pp\. D553–D559\.Cited by:[§4\.1](https://arxiv.org/html/2607.27431#S4.SS1.p1.1)\.
- J\. Dauparas, I\. Anishchenko, N\. Bennett, H\. Bai, R\. J\. Ragotte, L\. F\. Milles, B\. I\. Wicky, A\. Courbet, R\. J\. de Haas, N\. Bethel,et al\.\(2022\)Robust deep learning–based protein sequence design using proteinmpnn\.Science378\(6615\),pp\. 49–56\.Cited by:[§4\.5](https://arxiv.org/html/2607.27431#S4.SS5.p1.7)\.
- R\. Flamary, N\. Courty, A\. Gramfort, M\. Z\. Alaya, A\. Boisbunon, S\. Chambon, L\. Chapel, A\. Corenflos, K\. Fatras, N\. Fournier,et al\.\(2021\)Pot: python optimal transport\.Journal of Machine Learning Research22\(78\),pp\. 1–8\.Cited by:[§M\.1](https://arxiv.org/html/2607.27431#A13.SS1.SSS0.Px2.p1.12),[§D\.3](https://arxiv.org/html/2607.27431#A4.SS3.p1.6)\.
- Z\. Geng, M\. Deng, X\. Bai, Z\. Kolter, and K\. He \(2026a\)Mean flows for one\-step generative modeling\.Advances in Neural Information Processing Systems38,pp\. 75460–75482\.Cited by:[18](https://arxiv.org/html/2607.27431#A1.E18.m1.1.1.1.1.1.1.7),[18](https://arxiv.org/html/2607.27431#A1.E18.m1.1.1.1.1.1.1.mf.7),[Appendix A](https://arxiv.org/html/2607.27431#A1.p2.1),[§G\.2](https://arxiv.org/html/2607.27431#A7.SS2.p1.2),[§1](https://arxiv.org/html/2607.27431#S1.p3.8),[§2\.3](https://arxiv.org/html/2607.27431#S2.SS3.p1.4)\.
- Z\. Geng, Y\. Lu, Z\. Wu, E\. Shechtman, J\. Z\. Kolter, and K\. He \(2026b\)Improved mean flows: on the challenges of fastforward generative models\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 30467–30476\.Cited by:[18](https://arxiv.org/html/2607.27431#A1.E18.m1.2.2.2.2.1.1.7),[18](https://arxiv.org/html/2607.27431#A1.E18.m1.2.2.2.2.1.1.mf.7),[§G\.2](https://arxiv.org/html/2607.27431#A7.SS2.p1.2),[§2\.3](https://arxiv.org/html/2607.27431#S2.SS3.p1.4),[§2\.3](https://arxiv.org/html/2607.27431#S2.SS3.p2.8)\.
- A\. J\. Hanson \(2005\)Visualizing quaternions\.InACM SIGGRAPH 2005 Courses,pp\. 1–es\.Cited by:[§L\.1](https://arxiv.org/html/2607.27431#A12.SS1.p1.8)\.
- G\. Huguet, J\. Vuckovic, K\. Fatras, E\. Thibodeau\-Laufer, P\. Lemos, R\. Islam, C\. Liu, J\. Rector\-Brooks, T\. Akhound\-Sadegh, M\. Bronstein,et al\.\(2024\)Sequence\-augmented se \(3\)\-flow matching for conditional protein generation\.Advances in neural information processing systems37,pp\. 33007–33036\.Cited by:[§E\.3](https://arxiv.org/html/2607.27431#A5.SS3),[§E\.3](https://arxiv.org/html/2607.27431#A5.SS3.p1.2)\.
- J\. Jumper, R\. Evans, A\. Pritzel, T\. Green, M\. Figurnov, O\. Ronneberger, K\. Tunyasuvunakool, R\. Bates, A\. Žídek, A\. Potapenko,et al\.\(2021\)Highly accurate protein structure prediction with alphafold\.nature596\(7873\),pp\. 583–589\.Cited by:[Appendix N](https://arxiv.org/html/2607.27431#A14.p1.6),[§1](https://arxiv.org/html/2607.27431#S1.p2.7)\.
- Y\. Lin, M\. Lee, Z\. Zhang, and M\. AlQuraishi \(2024\)Out of many, one: designing and scaffolding proteins at the scale of the structural universe with genie 2\.arXiv preprint arXiv:2405\.15489\.Cited by:[§E\.1](https://arxiv.org/html/2607.27431#A5.SS1.p1.9)\.
- Z\. Lin, H\. Akin, R\. Rao, B\. Hie, Z\. Zhu, W\. Lu, N\. Smetanin, R\. Verkuil, O\. Kabeli, Y\. Shmueli,et al\.\(2023\)Evolutionary\-scale prediction of atomic\-level protein structure with a language model\.Science379\(6637\),pp\. 1123–1130\.Cited by:[§4\.5](https://arxiv.org/html/2607.27431#S4.SS5.p1.7)\.
- Y\. Lipman, R\. T\. Chen, H\. Ben\-Hamu, M\. Nickel, and M\. Le \(2022\)Flow matching for generative modeling\.arXiv preprint arXiv:2210\.02747\.Cited by:[§D\.1](https://arxiv.org/html/2607.27431#A4.SS1.p1.4),[§1](https://arxiv.org/html/2607.27431#S1.p3.8)\.
- L\. Pauling, R\. B\. Corey, and H\. R\. Branson \(1951\)The structure of proteins: two hydrogen\-bonded helical configurations of the polypeptide chain\.Proceedings of the National Academy of Sciences37\(4\),pp\. 205–211\.Cited by:[§O\.3](https://arxiv.org/html/2607.27431#A15.SS3.p1.2)\.
- W\. Peebles and S\. Xie \(2023\)Scalable diffusion models with transformers\.InProceedings of the IEEE/CVF international conference on computer vision,pp\. 4195–4205\.Cited by:[§N\.2](https://arxiv.org/html/2607.27431#A14.SS2.p1.11)\.
- A\. Pooladian, H\. Ben\-Hamu, C\. Domingo\-Enrich, B\. Amos, Y\. Lipman, and R\. T\. Chen \(2023\)Multisample flow matching: straightening flows with minibatch couplings\.arXiv preprint arXiv:2304\.14772\.Cited by:[§D\.3](https://arxiv.org/html/2607.27431#A4.SS3.p1.6)\.
- T\. Salimans and J\. Ho \(2022\)Progressive distillation for fast sampling of diffusion models\.arXiv preprint arXiv:2202\.00512\.Cited by:[§1](https://arxiv.org/html/2607.27431#S1.p3.8)\.
- D\. Silva, S\. Yu, U\. Y\. Ulge, J\. B\. Spangler, K\. M\. Jude, C\. Labão\-Almeida, L\. R\. Ali, A\. Quijano\-Rubio, M\. Ruterbusch, I\. Leung,et al\.\(2019\)De novo design of potent and selective mimics of il\-2 and il\-15\.Nature565\(7738\),pp\. 186–191\.Cited by:[§1](https://arxiv.org/html/2607.27431#S1.p1.1)\.
- Y\. Song, P\. Dhariwal, M\. Chen, and I\. Sutskever \(2023\)Consistency models\.Cited by:[§1](https://arxiv.org/html/2607.27431#S1.p3.8)\.
- Y\. Song and P\. Dhariwal \(2024\)Improved techniques for training consistency models\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 15078–15097\.Cited by:[§1](https://arxiv.org/html/2607.27431#S1.p3.8)\.
- E\. Strauch, S\. M\. Bernard, D\. La, A\. J\. Bohn, P\. S\. Lee, C\. E\. Anderson, T\. Nieusma, C\. A\. Holstein, N\. K\. Garcia, K\. A\. Hooper,et al\.\(2017\)Computational design of trimeric influenza\-neutralizing proteins targeting the hemagglutinin receptor binding site\.Nature biotechnology35\(7\),pp\. 667–671\.Cited by:[§1](https://arxiv.org/html/2607.27431#S1.p1.1)\.
- A\. Tong, K\. Fatras, N\. Malkin, G\. Huguet, Y\. Zhang, J\. Rector\-Brooks, G\. Wolf, and Y\. Bengio \(2023\)Improving and generalizing flow\-based generative models with minibatch optimal transport\.arXiv preprint arXiv:2302\.00482\.Cited by:[§D\.1](https://arxiv.org/html/2607.27431#A4.SS1.p1.4),[§D\.3](https://arxiv.org/html/2607.27431#A4.SS3.p1.6),[§1](https://arxiv.org/html/2607.27431#S1.p3.8)\.
- M\. van Kempen, S\. S\. Kim, C\. Tumescheit, M\. Mirdita, C\. L\. Gilchrist, J\. Söding, and M\. Steinegger \(2022\)Foldseek: fast and accurate protein structure search\.Biorxiv,pp\. 2022–02\.Cited by:[§4\.5](https://arxiv.org/html/2607.27431#S4.SS5.p2.1)\.
- C\. Villaniet al\.\(2009\)Optimal transport: old and new\.Vol\.338,Springer\.Cited by:[§D\.3](https://arxiv.org/html/2607.27431#A4.SS3.p1.6)\.
- J\. L\. Watson, D\. Juergens, N\. R\. Bennett, B\. L\. Trippe, J\. Yim, H\. E\. Eisenach, W\. Ahern, A\. J\. Borst, R\. J\. Ragotte, L\. F\. Milles,et al\.\(2023\)De novo design of protein structure and function with rfdiffusion\.Nature620\(7976\),pp\. 1089–1100\.Cited by:[§1](https://arxiv.org/html/2607.27431#S1.p2.7)\.
- D\. Woo, M\. Skreta, S\. Park, K\. Neklyudov, and S\. Ahn \(2026\)Riemannian meanflow\.arXiv preprint arXiv:2602\.07744\.Cited by:[§K\.5](https://arxiv.org/html/2607.27431#A11.SS5.p1.1),[Table 8](https://arxiv.org/html/2607.27431#A14.T8),[Appendix N](https://arxiv.org/html/2607.27431#A14.p1.6),[Table 10](https://arxiv.org/html/2607.27431#A15.T10),[§E\.6](https://arxiv.org/html/2607.27431#A5.SS6),[§E\.6](https://arxiv.org/html/2607.27431#A5.SS6.p1.1),[§E\.7](https://arxiv.org/html/2607.27431#A5.SS7.p1.1),[§H\.4](https://arxiv.org/html/2607.27431#A8.SS4.SSS0.Px5.p1.6),[Table 5](https://arxiv.org/html/2607.27431#A8.T5.14.14.1),[Table 5](https://arxiv.org/html/2607.27431#A8.T5.14.17.3.1),[§1](https://arxiv.org/html/2607.27431#S1.p4.1),[§4\.1](https://arxiv.org/html/2607.27431#S4.SS1.p1.1),[§4\.2](https://arxiv.org/html/2607.27431#S4.SS2.p1.1),[Table 1](https://arxiv.org/html/2607.27431#S4.T1)\.
- J\. Yim, A\. Campbell, A\. Y\. Foong, M\. Gastegger, J\. Jiménez\-Luna, S\. Lewis, V\. G\. Satorras, B\. S\. Veeling, R\. Barzilay, T\. Jaakkola,et al\.\(2023a\)Fast protein backbone generation with se \(3\) flow matching\.arXiv preprint arXiv:2310\.05297\.Cited by:[§M\.1](https://arxiv.org/html/2607.27431#A13.SS1.SSS0.Px1.p1.4),[§E\.1](https://arxiv.org/html/2607.27431#A5.SS1.p1.9),[§1](https://arxiv.org/html/2607.27431#S1.p2.7),[§4\.1](https://arxiv.org/html/2607.27431#S4.SS1.p1.1),[§4\.2](https://arxiv.org/html/2607.27431#S4.SS2.p1.1),[Table 1](https://arxiv.org/html/2607.27431#S4.T1)\.
- J\. Yim, B\. L\. Trippe, V\. De Bortoli, E\. Mathieu, A\. Doucet, R\. Barzilay, and T\. Jaakkola \(2023b\)SE \(3\) diffusion model with application to protein backbone generation\.arXiv preprint arXiv:2302\.02277\.Cited by:[§E\.1](https://arxiv.org/html/2607.27431#A5.SS1.p1.9),[§G\.1](https://arxiv.org/html/2607.27431#A7.SS1.p1.11),[§G\.1](https://arxiv.org/html/2607.27431#A7.SS1.p1.3),[§1](https://arxiv.org/html/2607.27431#S1.p2.7),[§4\.4](https://arxiv.org/html/2607.27431#S4.SS4.SSS0.Px1.p1.11)\.
- A\. Yue, Z\. Wang, and H\. Xu \(2025\)Reqflow: rectified quaternion flow for efficient and high\-quality protein backbone generation\.arXiv preprint arXiv:2502\.14637\.Cited by:[Appendix L](https://arxiv.org/html/2607.27431#A12.p1.1),[§M\.1](https://arxiv.org/html/2607.27431#A13.SS1.SSS0.Px1.p1.4),[§M\.3](https://arxiv.org/html/2607.27431#A13.SS3.p1.1),[Table 8](https://arxiv.org/html/2607.27431#A14.T8),[Appendix N](https://arxiv.org/html/2607.27431#A14.p1.6),[§O\.1](https://arxiv.org/html/2607.27431#A15.SS1.p1.7),[§O\.3](https://arxiv.org/html/2607.27431#A15.SS3.p1.2),[§E\.4](https://arxiv.org/html/2607.27431#A5.SS4),[§E\.4](https://arxiv.org/html/2607.27431#A5.SS4.p1.2),[Table 5](https://arxiv.org/html/2607.27431#A8.T5.14.20.6.1),[Table 5](https://arxiv.org/html/2607.27431#A8.T5.14.21.7.1),[§I\.2](https://arxiv.org/html/2607.27431#A9.SS2.p1.2),[§4\.1](https://arxiv.org/html/2607.27431#S4.SS1.p1.1),[§4\.2](https://arxiv.org/html/2607.27431#S4.SS2.p1.1),[§4\.3](https://arxiv.org/html/2607.27431#S4.SS3.p1.1),[§4\.4](https://arxiv.org/html/2607.27431#S4.SS4.SSS0.Px1.p1.11),[§4\.5](https://arxiv.org/html/2607.27431#S4.SS5.p1.7),[Table 1](https://arxiv.org/html/2607.27431#S4.T1),[Remark 2](https://arxiv.org/html/2607.27431#Thmremark2.p1.6)\.
- O\. Zaghen, F\. Eijkelboom, A\. Pouplin, C\. Liu, M\. Welling, J\. van de Meent, and E\. J\. Bekkers \(2025\)Riemannian variational flow matching for material and protein design\.arXiv preprint arXiv:2502\.12981\.Cited by:[§E\.5](https://arxiv.org/html/2607.27431#A5.SS5),[§E\.5](https://arxiv.org/html/2607.27431#A5.SS5.p1.2)\.
- H\. Zhang, A\. Siarohin, W\. Menapace, M\. Vasilkovsky, S\. Tulyakov, Q\. Qu, and I\. Skorokhodov \(2025\)Alphaflow: understanding and improving meanflow models\.arXiv preprint arXiv:2510\.20771\.Cited by:[§J\.3](https://arxiv.org/html/2607.27431#A10.SS3.p1.4),[Appendix J](https://arxiv.org/html/2607.27431#A10.p1.10),[§3\.3](https://arxiv.org/html/2607.27431#S3.SS3.p1.3),[§3\.3](https://arxiv.org/html/2607.27431#S3.SS3.p2.22),[Remark 9](https://arxiv.org/html/2607.27431#Thmremark9.p1.7)\.
- Y\. Zhang and J\. Skolnick \(2004\)Scoring function for automated assessment of protein structure template quality\.Proteins: Structure, Function, and Bioinformatics57\(4\),pp\. 702–710\.Cited by:[§4\.5](https://arxiv.org/html/2607.27431#S4.SS5.p2.1)\.
- Y\. Zhang and C\. Sagui \(2015\)Secondary structure assignment for conformationally irregular peptides: comparison between dssp, stride and kaksi\.Journal of Molecular Graphics and Modelling55,pp\. 72–84\.Cited by:[§O\.3](https://arxiv.org/html/2607.27431#A15.SS3.p1.2)\.
- Z\. Zhong, H\. Sun, Y\. Zhao, Y\. Gong, and Y\. Yin \(2026\)Riemannian meanflow for one\-step generation on manifolds\.arXiv preprint arXiv:2603\.10718\.Cited by:[Appendix O](https://arxiv.org/html/2607.27431#A15.SS0.SSS0.Px3.p1.20),[§E\.7](https://arxiv.org/html/2607.27431#A5.SS7.p1.1),[§E\.7](https://arxiv.org/html/2607.27431#A5.SS7.p3.3),[§H\.4](https://arxiv.org/html/2607.27431#A8.SS4.SSS0.Px3.p1.8),[§1](https://arxiv.org/html/2607.27431#S1.p4.1)\.

## Appendix Contents

## Appendix ABackground: MeanFlow in Euclidean Space

Consider a probability path\{pt:t∈\[0,1\]\}\\\{p\_\{t\}:t\\in\[0,1\]\\\}withp0=pd​a​t​ap\_\{0\}=p\_\{data\}andp1=pp​r​i​o​rp\_\{1\}=p\_\{prior\}, we assumeptp\_\{t\}is generated by the following ODE system:

\{X0∼p0initial distributiond​Xt=v​\(t,Xt\)​d​tODEpt=Law​\(Xt\)Marginal distribution in time\\displaystyle\\begin\{cases\}X\_\{0\}\\sim p\_\{0\}&\\text\{initial distribution\}\\\\ dX\_\{t\}=v\(t,X\_\{t\}\)dt&\\text\{ODE\}\\\\ p\_\{t\}=\\text\{Law\}\(X\_\{t\}\)&\\text\{Marginal distribution in time\}\\end\{cases\}\(15\)Equivalently, we write\(vt,xt\)\(v\_\{t\},x\_\{t\}\)satisfies the continuity equation:

dd​t​pt=−∇⋅\(vt​pt\)\\displaystyle\\frac\{d\}\{dt\}p\_\{t\}=\-\\nabla\\cdot\(v\_\{t\}p\_\{t\}\)\(16\)with the same initial distribution condition\.

Genget al\.\[[2026a](https://arxiv.org/html/2607.27431#bib.bib35)\]proposed the mean velocity fieldu​\(s,t,xt\)u\(s,t,x\_\{t\}\)defined as

u​\(s,t,xt\):=1t−s​∫stv​\(τ,xτ\)​𝑑τ∈ℝd\.u\(s,t,x\_\{t\}\):=\\frac\{1\}\{t\-s\}\\int\_\{s\}^\{t\}v\(\\tau,x\_\{\\tau\}\)d\\tau\\in\\mathbb\{R\}^\{d\}\.\(17\)
Multiplying\(t−s\)\(t\-s\)on both sides and differentiating both sides with respect tottyields the*MeanFlow identity*:

\{L\.H\.S\.=dd​t​\(t−s\)​u​\(s,t,xt\)=u​\(s,t,xt\)\+\(t−s\)​dd​t​u​\(s,t,xt\)R\.H\.S\.=v​\(t,xt\)=x1−x0\\displaystyle\\begin\{cases\}\\text\{L\.H\.S\.\}=\\frac\{d\}\{dt\}\(t\-s\)u\(s,t,x\_\{t\}\)=u\(s,t,x\_\{t\}\)\+\(t\-s\)\\frac\{d\}\{dt\}u\(s,t,x\_\{t\}\)\\\\ \\text\{R\.H\.S\.\}=v\(t,x\_\{t\}\)=x\_\{1\}\-x\_\{0\}\\\\ \\end\{cases\}
We compute R\.H\.S\. from the data and L\.H\.S\. from the model, the training loss is derived:

ℒ=\{𝔼t,\(x0,x1\)∼π0​\[‖uθ\+sg​\(\(t−s\)​\(dd​t​ut\)−v​\(t,xt\)\)‖2\]​Genget al\.\[[2026a](https://arxiv.org/html/2607.27431#bib.bib35)\]𝔼t,\(x0,x1\)∼π0​\[‖uθ\+\(t−s\)​sg​\(dd​t​ut\)−v​\(t,xt\)‖2\]​Genget al\.\[[2026b](https://arxiv.org/html/2607.27431#bib.bib27)\]\\displaystyle\\mathcal\{L\}=\\begin\{cases\}&\\mathbb\{E\}\_\{t,\(x\_\{0\},x\_\{1\}\)\\sim\\pi\_\{0\}\}\\left\[\\left\\\|u\_\{\\theta\}\+\\text\{sg\}\(\(t\-s\)\(\\frac\{d\}\{dt\}u\_\{t\}\)\-v\(t,x\_\{t\}\)\)\\right\\\|^\{2\}\\right\]\\text\{\\cite\[cite\]\{\\@@bibref\{Authors Phrase1YearPhrase2\}\{geng2026mean\}\{\\@@citephrase\{\[\}\}\{\\@@citephrase\{\]\}\}\}\}\\\\ &\\mathbb\{E\}\_\{t,\(x\_\{0\},x\_\{1\}\)\\sim\\pi\_\{0\}\}\\left\[\\left\\\|u\_\{\\theta\}\+\(t\-s\)\\text\{sg\}\(\\frac\{d\}\{dt\}u\_\{t\}\)\-v\(t,x\_\{t\}\)\\right\\\|^\{2\}\\right\]\\text\{\\cite\[cite\]\{\\@@bibref\{Authors Phrase1YearPhrase2\}\{geng2026improved\}\{\\@@citephrase\{\[\}\}\{\\@@citephrase\{\]\}\}\}\}\\end\{cases\}\(18\)where​v​\(t,xt\)=x1−x0\\displaystyle\\text\{where \}v\(t,x\_\{t\}\)=x\_\{1\}\-x\_\{0\}dd​t​uθ​\(r,t,xt\)=∂uθ∂xt​v​\(t,xt\)\+∂uθ∂t\\displaystyle\\frac\{d\}\{dt\}u\_\{\\theta\}\(r,t,x\_\{t\}\)=\\frac\{\\partial u\_\{\\theta\}\}\{\\partial x\_\{t\}\}v\(t,x\_\{t\}\)\+\\frac\{\\partial u\_\{\\theta\}\}\{\\partial t\}andsgdenotes the stop\-gradient operation, which is used to prevent backpropagation through the target\.

##### One\-step training objective\.

MeanFlow learns a parametric fielduθ​\(s,t,xt\)u\_\{\\theta\}\(s,t,x\_\{t\}\)via regression, using the standard linear interpolation between a data samplex0∼pdatax\_\{0\}\\sim p\_\{\\text\{data\}\}and a prior samplex1∼ppriorx\_\{1\}\\sim p\_\{\\text\{prior\}\}:

xt=\(1−t\)​x0\+t​x1,v=x1−x0\.x\_\{t\}=\(1\-t\)x\_\{0\}\+tx\_\{1\},\\qquad v=x\_\{1\}\-x\_\{0\}\.\(19\)Using the MeanFlow identity, the target mean flow can be written as

utgt​\(s,t,xt\)=v−\(t−s\)​dd​t​uθ​\(s,t,xt\),u\_\{\\text\{tgt\}\}\(s,t,x\_\{t\}\)=v\-\(t\-s\)\\frac\{d\}\{dt\}u\_\{\\theta\}\(s,t,x\_\{t\}\),\(20\)wheredd​t​uθ\\frac\{d\}\{dt\}u\_\{\\theta\}is the total derivative \(implemented as a Jacobian\-vector product\), andutgtu\_\{\\text\{tgt\}\}is treated as a stop\-gradient target\. The MeanFlow training loss is

ℒMF​\(θ\)=𝔼x0∼pdata,x1∼pprior,s<t​\[‖uθ​\(s,t,x\)−sg​\(utgt​\(s,t,xt\)\)‖22\]\.\\mathcal\{L\}\_\{\\text\{MF\}\}\(\\theta\)=\\mathbb\{E\}\_\{x\_\{0\}\\sim p\_\{\\text\{data\}\},\\ x\_\{1\}\\sim p\_\{\\text\{prior\}\},\\ s<t\}\\left\[\\left\\\|u\_\{\\theta\}\(s,t,x\)\-\\mathrm\{sg\}\\big\(u\_\{\\text\{tgt\}\}\(s,t,x\_\{t\}\)\\big\)\\right\\\|\_\{2\}^\{2\}\\right\]\.\(21\)

## Appendix BBackground: SO\(3\) Space

### B\.1Basic concepts in SO\(3\) space

The special orthogonal group in three dimensions, denoted asSO​\(3\)\\mathrm\{SO\}\(3\), is defined as

SO​\(3\):=\{R∈ℝ3×3:R​R⊤=R⊤​R=I3,det\(R\)=1\}\.\\mathrm\{SO\}\(3\):=\\\{R\\in\\mathbb\{R\}^\{3\\times 3\}:RR^\{\\top\}=R^\{\\top\}R=I\_\{3\},\\det\(R\)=1\\\}\.The constraintR​R⊤=I3RR^\{\\top\}=I\_\{3\}consists of smooth polynomial equations, which implies thatSO​\(3\)\\mathrm\{SO\}\(3\)is a smooth embedded submanifold ofℝ9\\mathbb\{R\}^\{9\}\.

Equipped with matrix multiplicationSO​\(3\)×2∋\(S1,S2\)↦S1​S2∈SO​\(3\)\\mathrm\{SO\}\(3\)^\{\\times 2\}\\ni\(S\_\{1\},S\_\{2\}\)\\mapsto S\_\{1\}S\_\{2\}\\in\\mathrm\{SO\}\(3\),SO​\(3\)\\mathrm\{SO\}\(3\)forms a Lie group\. The identity element isI3I\_\{3\}, and the inverse of any elementR∈SO​\(3\)R\\in\\mathrm\{SO\}\(3\)is given byR−1=R⊤R^\{\-1\}=R^\{\\top\}\.

##### Tangent space\.

LetR∈SO​\(3\)R\\in\\mathrm\{SO\}\(3\)be fixed\. Consider a smooth curveR​\(t\)∈SO​\(3\)R\(t\)\\in\\mathrm\{SO\}\(3\)such that

SinceR​\(t\)​R​\(t\)⊤=I3R\(t\)R\(t\)^\{\\top\}=I\_\{3\}holds for alltt, differentiating att=0t=0yields

R˙​\(0\)​R⊤\+R​R˙​\(0\)⊤=0\.\\dot\{R\}\(0\)R^\{\\top\}\+R\\dot\{R\}\(0\)^\{\\top\}=0\.This implies thatR⊤​R˙​\(0\)R^\{\\top\}\\dot\{R\}\(0\)is skew\-symmetric\. Therefore, the tangent space atRRis given by

TR​\(SO​\(3\)\)=\{R​Ω:Ω∈𝔰​𝔬​\(3\)\},T\_\{R\}\(\\mathrm\{SO\}\(3\)\)=\\\{R\\Omega:\\Omega\\in\\mathfrak\{so\}\(3\)\\\},where the Lie algebra𝔰​𝔬​\(3\)\\mathfrak\{so\}\(3\)is defined as

𝔰​𝔬​\(3\):=\{Ω∈ℝ3×3:Ω⊤=−Ω\}\.\\mathfrak\{so\}\(3\):=\\\{\\Omega\\in\\mathbb\{R\}^\{3\\times 3\}:\\Omega^\{\\top\}=\-\\Omega\\\}\.This representation provides a global linear parameterization of tangent vectors onSO​\(3\)\\mathrm\{SO\}\(3\)\. We also use the standard hat/vee identification betweenℝ3\\mathbb\{R\}^\{3\}and𝔰​𝔬​\(3\)\\mathfrak\{so\}\(3\): forω=\(ω1,ω2,ω3\)∈ℝ3\\omega=\(\\omega\_\{1\},\\omega\_\{2\},\\omega\_\{3\}\)\\in\\mathbb\{R\}^\{3\},

ω∧:=\[0−ω3ω2ω30−ω1−ω2ω10\]∈𝔰𝔬\(3\),\(⋅\)∨:𝔰𝔬\(3\)→ℝ3is its inverse\.\\omega^\{\\wedge\}:=\\begin\{bmatrix\}0&\-\\omega\_\{3\}&\\omega\_\{2\}\\\\ \\omega\_\{3\}&0&\-\\omega\_\{1\}\\\\ \-\\omega\_\{2\}&\\omega\_\{1\}&0\\end\{bmatrix\}\\in\\mathfrak\{so\}\(3\),\\qquad\(\\cdot\)^\{\\vee\}:\\mathfrak\{so\}\(3\)\\to\\mathbb\{R\}^\{3\}\\ \\text\{is its inverse\}\.

##### Lie algebra and Lie bracket\.

The vector space𝔰​𝔬​\(3\)\\mathfrak\{so\}\(3\)is the Lie algebra associated with the Lie groupSO​\(3\)\\mathrm\{SO\}\(3\)\. The Lie bracket on𝔰​𝔬​\(3\)\\mathfrak\{so\}\(3\)is defined as the matrix commutator

\[Ω1,Ω2\]=Ω1​Ω2−Ω2​Ω1,Ω1,Ω2∈𝔰​𝔬​\(3\)\.\[\\Omega\_\{1\},\\Omega\_\{2\}\]=\\Omega\_\{1\}\\Omega\_\{2\}\-\\Omega\_\{2\}\\Omega\_\{1\},\\quad\\Omega\_\{1\},\\Omega\_\{2\}\\in\\mathfrak\{so\}\(3\)\.It is straightforward to verify that this operation preserves skew\-symmetry, hence𝔰​𝔬​\(3\)\\mathfrak\{so\}\(3\)is closed under the Lie bracket\.

### B\.2Curves and flows in SO\(3\)

In Euclidean space, a curve connectingx0x\_\{0\}andx1x\_\{1\}can be defined by the ODE

\{X0=x0,X˙t=v​\(t,Xt\)\.\\displaystyle\\begin\{cases\}X\_\{0\}=x\_\{0\},\\\\ \\dot\{X\}\_\{t\}=v\(t,X\_\{t\}\)\.\\end\{cases\}\(22\)and thusX1=∫01v​\(t,xt\)​𝑑tX\_\{1\}=\\int\_\{0\}^\{1\}v\(t,x\_\{t\}\)dt\. We now extend this construction to the Lie groupSO​\(3\)\\mathrm\{SO\}\(3\)\.

##### ODE\-defined curves onSO​\(3\)\\mathrm\{SO\}\(3\)\.

LetRt:\[0,1\]→SO​\(3\)R\_\{t\}:\[0,1\]\\to\\mathrm\{SO\}\(3\)be a time\-dependent curve\. A natural intrinsic definition of its evolution is given by

\{R0=r0R˙t=Rt​Ωt,Ωt∈𝔰​𝔬​\(3\)\.\\begin\{cases\}R\_\{0\}=r\_\{0\}\\\\ \\dot\{R\}\_\{t\}=R\_\{t\}\\Omega\_\{t\},\\quad\\Omega\_\{t\}\\in\\mathfrak\{so\}\(3\)\.\\end\{cases\}This ODE guarantees thatRtR\_\{t\}remains inSO​\(3\)\\mathrm\{SO\}\(3\)for alltt, since

dd​t​\(Rt​Rt⊤\)=Rt​Ωt​Rt⊤\+Rt​Ωt⊤​Rt⊤=0\.\\frac\{d\}\{dt\}\(R\_\{t\}R\_\{t\}^\{\\top\}\)=R\_\{t\}\\Omega\_\{t\}R\_\{t\}^\{\\top\}\+R\_\{t\}\\Omega\_\{t\}^\{\\top\}R\_\{t\}^\{\\top\}=0\.Given an initial conditionR0∈SO​\(3\)R\_\{0\}\\in\\mathrm\{SO\}\(3\), the above equation defines a smooth curve on the manifold\.

##### Exponential map and geodesics\.

When the angular velocity is constant, namelyΩt=Ω\\Omega\_\{t\}=\\Omega, the ODE admits a closed\-form solution

Rt=R0​exp⁡\(t​Ω\),R\_\{t\}=R\_\{0\}\\exp\(t\\Omega\),whereexp\\expdenotes the matrix exponential\. Under the canonical bi\-invariant Riemannian metric onSO​\(3\)\\mathrm\{SO\}\(3\), curves of this form are geodesics\.

Given two pointsR0,R1∈SO​\(3\)R\_\{0\},R\_\{1\}\\in\\mathrm\{SO\}\(3\), the geodesic connecting them is given by

Rt=R0​exp⁡\(t​Log​\(R0⊤​R1\)\),t∈\[0,1\],R\_\{t\}=R\_\{0\}\\exp\\big\(t\\text\{Log\}\(R\_\{0\}^\{\\top\}R\_\{1\}\)\\big\),\\quad t\\in\[0,1\],whereLogdenotes the matrix logarithm mappingSO​\(3\)\\mathrm\{SO\}\(3\)to𝔰​𝔬​\(3\)\\mathfrak\{so\}\(3\)\. This curve minimizes the path length among all smooth curves onSO​\(3\)\\mathrm\{SO\}\(3\)connectingR0R\_\{0\}andR1R\_\{1\}\.

##### Inner product and geodesic distance inSO​\(3\)\\mathrm\{SO\}\(3\)\.

Under the canonical bi\-invariant Riemannian metric, we identifyTR​SO​\(3\)=\{R​Ω:Ω∈𝔰​𝔬​\(3\)\}T\_\{R\}\\mathrm\{SO\}\(3\)=\\\{R\\Omega:\\Omega\\in\\mathfrak\{so\}\(3\)\\\}and define the inner product

⟨R​Ω1,R​Ω2⟩R=12​tr​\(Ω1⊤​Ω2\)=12​tr​\(\(R⊤​R˙1\)⊤​\(R⊤​R˙2\)\),\\langle R\\Omega\_\{1\},R\\Omega\_\{2\}\\rangle\_\{R\}=\\tfrac\{1\}\{2\}\\,\\mathrm\{tr\}\(\\Omega\_\{1\}^\{\\top\}\\Omega\_\{2\}\)=\\tfrac\{1\}\{2\}\\,\\mathrm\{tr\}\\big\(\(R^\{\\top\}\\dot\{R\}\_\{1\}\)^\{\\top\}\(R^\{\\top\}\\dot\{R\}\_\{2\}\)\\big\),which is invariant under left and right multiplication\. The induced geodesic distance betweenR0,R1∈SO​\(3\)R\_\{0\},R\_\{1\}\\in\\mathrm\{SO\}\(3\)is

dist​\(R0,R1\)=12​‖Log​\(R0⊤​R1\)‖F,\\mathrm\{dist\}\(R\_\{0\},R\_\{1\}\)=\\frac\{1\}\{\\sqrt\{2\}\}\\,\\big\\\|\\text\{Log\}\(R\_\{0\}^\{\\top\}R\_\{1\}\)\\big\\\|\_\{F\},where∥⋅∥F\\\|\\cdot\\\|\_\{F\}denotes the Frobenius norm\.

##### Integration in SO\(3\)\.

In Euclidean space,x1=∫01v​\(t,xt\)​𝑑t\+x0x\_\{1\}=\\int\_\{0\}^\{1\}v\(t,x\_\{t\}\)dt\+x\_\{0\}\.

Due to the non\-commutative group structure, the endpoint inSO​\(3\)\\mathrm\{SO\}\(3\)cannot be written as a simple integral\. Instead, the solution att=1t=1admits the group\-valued representation

R​\(1\)=R​\(0\)​𝒯​exp⁡\(∫01Ωt​𝑑t\),R\(1\)=R\(0\)\\mathcal\{T\}\\exp\\Big\(\\int\_\{0\}^\{1\}\\Omega\_\{t\}dt\\Big\),where𝒯\\mathcal\{T\}denotes the time\-ordering operator:

𝒯​exp⁡\(∫01A​\(t\)​𝑑t\)=I3\+∑k=1∞∫0≤t1≤⋯≤tk≤1A​\(t1\)​⋯​A​\(tk\)​𝑑t1​⋯​𝑑tk,\\mathcal\{T\}\\exp\\Big\(\\int\_\{0\}^\{1\}A\(t\)dt\\Big\)=I\_\{3\}\+\\sum\_\{k=1\}^\{\\infty\}\\int\_\{0\\leq t\_\{1\}\\leq\\cdots\\leq t\_\{k\}\\leq 1\}A\(t\_\{1\}\)\\cdots A\(t\_\{k\}\)dt\_\{1\}\\cdots dt\_\{k\},where the time\-ordering operator enforces the chronological ordering of the matrix products\.

## Appendix CBackground: SE\(3\) Space

### C\.1Basic concepts in theSE​\(3\)\\mathrm\{SE\}\(3\)space

The special Euclidean group in three dimensions is defined as

SE​\(3\)=\{\(R,x\):R∈SO​\(3\),x∈ℝ3\},\\mathrm\{SE\}\(3\)=\\\{\(R,x\):R\\in\\mathrm\{SO\}\(3\),\\ x\\in\\mathbb\{R\}^\{3\}\\\},representing orientation\-preserving rigid motions inℝ3\\mathbb\{R\}^\{3\}\. Algebraically,SE​\(3\)\\mathrm\{SE\}\(3\)is the semidirect product

SE​\(3\)=SO​\(3\)⋉\(ℝ3,\+\),\\mathrm\{SE\}\(3\)=\\mathrm\{SO\}\(3\)\\ltimes\(\\mathbb\{R\}^\{3\},\+\),where rotations act on translations\. The group operation is given by

\(R1,x1\)​\(R2,x2\)=\(R1​R2,x1\+R1​x2\),\\displaystyle\(R\_\{1\},x\_\{1\}\)\(R\_\{2\},x\_\{2\}\)=\(R\_\{1\}R\_\{2\},\\ x\_\{1\}\+R\_\{1\}x\_\{2\}\),\(23\)and the inverse by

\(R,x\)−1=\(R⊤,−R⊤​x\)\.\(R,x\)^\{\-1\}=\(R^\{\\top\},\\ \-R^\{\\top\}x\)\.As a smooth manifold,SE​\(3\)\\mathrm\{SE\}\(3\)is six\-dimensional and diffeomorphic toSO​\(3\)×ℝ3\\mathrm\{SO\}\(3\)\\times\\mathbb\{R\}^\{3\}\.

##### Lie algebra and tangent space\.

The Lie algebra𝔰​𝔢​\(3\)\\mathfrak\{se\}\(3\)consists of pairs\(Ω,v\)\(\\Omega,v\)withΩ∈𝔰​𝔬​\(3\)\\Omega\\in\\mathfrak\{so\}\(3\)andv∈ℝ3v\\in\\mathbb\{R\}^\{3\}, i\.e\.,

𝔰​𝔢​\(3\)=\{\(Ω,v\):Ω⊤=−Ω,v∈ℝ3\}\.\\mathfrak\{se\}\(3\)=\\\{\(\\Omega,v\):\\Omega^\{\\top\}=\-\\Omega,\\ v\\in\\mathbb\{R\}^\{3\}\\\}\.The Lie bracket is given by

\[\(Ω1,v1\),\(Ω2,v2\)\]=\(\[Ω1,Ω2\],Ω1​v2−Ω2​v1\)\.\[\(\\Omega\_\{1\},v\_\{1\}\),\(\\Omega\_\{2\},v\_\{2\}\)\]=\(\[\\Omega\_\{1\},\\Omega\_\{2\}\],\\ \\Omega\_\{1\}v\_\{2\}\-\\Omega\_\{2\}v\_\{1\}\)\.For any\(R,x\)∈SE​\(3\)\(R,x\)\\in\\mathrm\{SE\}\(3\), the tangent space is obtained by left translation:

T\(R,x\)​SE​\(3\)=\{\(R​Ω,R​v\):Ω∈𝔰​𝔬​\(3\),v∈ℝ3\}\.T\_\{\(R,x\)\}\\mathrm\{SE\}\(3\)=\\\{\(R\\Omega,\\ Rv\):\\Omega\\in\\mathfrak\{so\}\(3\),\\ v\\in\\mathbb\{R\}^\{3\}\\\}\.

### C\.2Curves and flows in SE\(3\)

A time\-dependent rigid motion is represented by a curve\(Rt,st\)∈SE​\(3\)\(R\_\{t\},s\_\{t\}\)\\in\\mathrm\{SE\}\(3\)\. A natural way to define its intrinsic evolution is through the left\-trivialized velocity:

\{R˙t=Rt​Ωt,x˙t=Rt​vt,\(Ωt,vt\)∈𝔰​𝔢​\(3\),\\begin\{cases\}\\dot\{R\}\_\{t\}=R\_\{t\}\\Omega\_\{t\},\\\\ \\dot\{x\}\_\{t\}=R\_\{t\}v\_\{t\},\\end\{cases\}\\qquad\(\\Omega\_\{t\},v\_\{t\}\)\\in\\mathfrak\{se\}\(3\),whereΩt\\Omega\_\{t\}governs angular motion andvtv\_\{t\}determines translational motion in the body frame\. Note that the translational velocity is expressed in the body frame and is therefore rotated byRtR\_\{t\}: in the full semidirect\-product geometry the two components are coupled\.

### C\.3The decoupled product geometry used in this work

Following the setting in Section 3 ofBoseet al\.\[[2024](https://arxiv.org/html/2607.27431#bib.bib9)\], we instead work with the product manifold

SO​\(3\)×ℝ3\\mathrm\{SO\}\(3\)\\times\\mathbb\{R\}^\{3\}equipped with a product Riemannian metric, under which the composition simplifies to

\(R1,s1\)​\(R2,s2\)=\(R1​R2,s1\+s2\)\.\\displaystyle\(R\_\{1\},s\_\{1\}\)\(R\_\{2\},s\_\{2\}\)=\(R\_\{1\}R\_\{2\},\\ s\_\{1\}\+s\_\{2\}\)\.\(24\)Comparing with \([23](https://arxiv.org/html/2607.27431#A3.E23)\), this replacess1\+R1​s2s\_\{1\}\+R\_\{1\}s\_\{2\}bys1\+s2s\_\{1\}\+s\_\{2\}, i\.e\. it drops the action ofR1R\_\{1\}ons2s\_\{2\}, so that the rotational and translational coordinates evolve independently\. The left\-trivialized flow correspondingly reduces toR˙t=Rt​Ωt\\dot\{R\}\_\{t\}=R\_\{t\}\\Omega\_\{t\}ands˙t=vt\\dot\{s\}\_\{t\}=v\_\{t\}\.

## Appendix DBackground: SE\(3\) Flow Matching

Our method builds on flow matching, which we briefly review here—first on a general Riemannian manifold, and then in the decoupledSE​\(3\)\\mathrm\{SE\}\(3\)geometry used throughout the paper\.

### D\.1Riemannian flow matching

Flow matching\[Lipmanet al\.,[2022](https://arxiv.org/html/2607.27431#bib.bib14), Tonget al\.,[2023](https://arxiv.org/html/2607.27431#bib.bib20)\]learns a time\-dependent velocity field whose flow transports a prior densityp1p\_\{1\}to the data densityp0p\_\{0\}\. On a Riemannian manifoldℳ\\mathcal\{M\}, one fixes a conditional coupling\(z0,z1\)∼q\(z\_\{0\},z\_\{1\}\)\\sim qand connects the endpoints by the minimizing geodesic

zt=expz0⁡\(t​logz0⁡\(z1\)\),t∈\[0,1\],z\_\{t\}=\\exp\_\{z\_\{0\}\}\\\!\\big\(t\\,\\log\_\{z\_\{0\}\}\(z\_\{1\}\)\\big\),\\qquad t\\in\[0,1\],whose time derivative is the conditional \(target\) velocityz˙t∈𝒯zt​ℳ\\dot\{z\}\_\{t\}\\in\\mathcal\{T\}\_\{z\_\{t\}\}\\mathcal\{M\}\. Regressing a model fielduθ​\(t,zt\)u\_\{\\theta\}\(t,z\_\{t\}\)onto this target,

𝔼t,q​\(z0,z1\)​\[‖uθ​\(t,zt\)−z˙t‖g2\],\\mathbb\{E\}\_\{t,\\,q\(z\_\{0\},z\_\{1\}\)\}\\big\[\\,\\\|u\_\{\\theta\}\(t,z\_\{t\}\)\-\\dot\{z\}\_\{t\}\\\|\_\{g\}^\{2\}\\,\\big\],recovers at its minimizer the marginal velocity that generates the interpolating pathptp\_\{t\}\. Sampling then integrates the learned field, e\.g\. backward from a prior drawz1z\_\{1\}to a data samplez0z\_\{0\}\.

### D\.2SE\(3\) flow matching

FollowingBoseet al\.\[[2024](https://arxiv.org/html/2607.27431#bib.bib9)\], we adopt the decoupled product geometrySE​\(3\)≅SO​\(3\)×ℝ3\\mathrm\{SE\}\(3\)\\cong\\mathrm\{SO\}\(3\)\\times\\mathbb\{R\}^\{3\}, so that the geodesic and its velocity split into independent rotational and translational parts\. For a data frame\(R0,x0\)\(R\_\{0\},x\_\{0\}\)and a prior frame\(R1,x1\)\(R\_\{1\},x\_\{1\}\), the conditional path is theSO​\(3\)\\mathrm\{SO\}\(3\)geodesic paired with the Euclidean straight line,

Rt=R0​exp⁡\(t​log⁡\(R0⊤​R1\)\),xt=\(1−t\)​x0\+t​x1\.R\_\{t\}=R\_\{0\}\\exp\\\!\\big\(t\\,\\log\(R\_\{0\}^\{\\top\}R\_\{1\}\)\\big\),\\qquad x\_\{t\}=\(1\-t\)\\,x\_\{0\}\+t\\,x\_\{1\}\.Differentiating gives the closed\-form conditional velocities, which are constant along each path:

ωt:=log\(R0⊤R1\)∨∈ℝ3,vt:=x1−x0∈ℝ3,\\omega\_\{t\}:=\\log\(R\_\{0\}^\{\\top\}R\_\{1\}\)^\{\\vee\}\\in\\mathbb\{R\}^\{3\},\\qquad v\_\{t\}:=x\_\{1\}\-x\_\{0\}\\in\\mathbb\{R\}^\{3\},so thatR˙t=Rt​ωt∧\\dot\{R\}\_\{t\}=R\_\{t\}\\,\\omega\_\{t\}^\{\\wedge\}andx˙t=vt\\dot\{x\}\_\{t\}=v\_\{t\}\. The model predicts the body angular and translational velocitiesAθ​\(t,Rt,xt\)A\_\{\\theta\}\(t,R\_\{t\},x\_\{t\}\)andvθ​\(t,xt\)v\_\{\\theta\}\(t,x\_\{t\}\)—equivalently the instantaneous \(s=ts\{=\}t\) evaluation of the two\-time head,Aθ​\(t,Rt,xt\)=Aθavg​\(t,t,Rt,xt\)A\_\{\\theta\}\(t,R\_\{t\},x\_\{t\}\)=A^\{\\mathrm\{avg\}\}\_\{\\theta\}\(t,t,R\_\{t\},x\_\{t\}\)—and theSE​\(3\)\\mathrm\{SE\}\(3\)flow\-matching objective regresses them onto these targets,

ℒFM=𝔼t,q​\(z0,z1\)​∑i=1N\[‖Aθ,i​\(t,Rt,i,xt,i\)−ωt,i‖22\+‖vθ,i​\(t,xt,i\)−vt,i‖22\],\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{FM\}\}=\\mathbb\{E\}\_\{t,\\,q\(z\_\{0\},z\_\{1\}\)\}\\sum\_\{i=1\}^\{N\}\\Big\[\\big\\\|A\_\{\\theta,i\}\(t,R\_\{t,i\},x\_\{t,i\}\)\-\\omega\_\{t,i\}\\big\\\|\_\{2\}^\{2\}\+\\big\\\|v\_\{\\theta,i\}\(t,x\_\{t,i\}\)\-v\_\{t,i\}\\big\\\|\_\{2\}^\{2\}\\Big\],\(25\)summed over theNNresidues\. Equation \([25](https://arxiv.org/html/2607.27431#A4.E25)\) is the data\-anchored objective to which our MeanFlow andα\\alpha\-Flow targets reduce in the appropriate limit—for instance atα=1\\alpha=1, whereAtgtavg=ωtA^\{\\mathrm\{avg\}\}\_\{\\mathrm\{tgt\}\}=\\omega\_\{t\}—and it supplies the boundary condition that keeps the consistency objective from collapsing\.

### D\.3Mini\-batch optimal\-transport coupling

The flow\-matching objective \([25](https://arxiv.org/html/2607.27431#A4.E25)\) is defined for any couplingq​\(z0,z1\)q\(z\_\{0\},z\_\{1\}\)of the data and prior marginals\. The simplest choice is the independent couplingq=p0⊗p1q=p\_\{0\}\\otimes p\_\{1\}, under which trajectories from different pairs cross and the marginal velocity field is far from constant along each path\. FollowingTonget al\.\[[2023](https://arxiv.org/html/2607.27431#bib.bib20)\], Pooladianet al\.\[[2023](https://arxiv.org/html/2607.27431#bib.bib47)\]and FoldFlowBoseet al\.\[[2024](https://arxiv.org/html/2607.27431#bib.bib9)\], we instead re\-pair each batch by discrete optimal transport\. GivenMMdata frames\{z0k\}\\\{z\_\{0\}^\{k\}\\\}andMMprior frames\{z1k\}\\\{z\_\{1\}^\{k\}\\\}with uniform weights, the Optimal transport problemVillani and others \[[2009](https://arxiv.org/html/2607.27431#bib.bib46)\], Flamaryet al\.\[[2021](https://arxiv.org/html/2607.27431#bib.bib43)\]

minΠ∈𝒰M​∑k,lΠk​l​c​\(z0k,z1l\),𝒰M=\{Π≥0:Π​𝟏=Π⊤​𝟏=1M​𝟏\},\\displaystyle\\min\_\{\\Pi\\in\\mathcal\{U\}\_\{M\}\}\\ \\sum\_\{k,l\}\\Pi\_\{kl\}\\,c\(z\_\{0\}^\{k\},z\_\{1\}^\{l\}\),\\qquad\\mathcal\{U\}\_\{M\}=\\\{\\Pi\\geq 0:\\Pi\\mathbf\{1\}=\\Pi^\{\\top\}\\mathbf\{1\}=\\tfrac\{1\}\{M\}\\mathbf\{1\}\\\},attains its optimum at a vertex of𝒰M\\mathcal\{U\}\_\{M\}, i\.e\. \(Birkhoff\) at a permutationπ⋆∈𝒮M\\pi^\{\\star\}\\in\\mathcal\{S\}\_\{M\}, so the plan reduces to a linear assignment solvable exactly inO​\(M3\)O\(M^\{3\}\)by the Hungarian algorithm\. On the decoupled product geometry \([24](https://arxiv.org/html/2607.27431#A3.E24)\) we take the squaredSE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}geodesic cost,

c​\(\(R0,x0\),\(R1,x1\)\)=∑i=1N\[λR​dSO​\(3\)2​\(R0i,R1i\)\+λx​‖x0i−x1i‖22\],dSO​\(3\)​\(R,R′\)=12​‖log⁡\(R⊤​R′\)‖F,\\displaystyle c\\big\(\(R\_\{0\},x\_\{0\}\),\(R\_\{1\},x\_\{1\}\)\\big\)=\\sum\_\{i=1\}^\{N\}\\Big\[\\lambda\_\{R\}\\,d\_\{\\mathrm\{SO\}\(3\)\}^\{2\}\(R\_\{0\}^\{i\},R\_\{1\}^\{i\}\)\+\\lambda\_\{x\}\\,\\\|x\_\{0\}^\{i\}\-x\_\{1\}^\{i\}\\\|\_\{2\}^\{2\}\\Big\],\\quad d\_\{\\mathrm\{SO\}\(3\)\}\(R,R^\{\\prime\}\)=\\tfrac\{1\}\{\\sqrt\{2\}\}\\\|\\log\(R^\{\\top\}R^\{\\prime\}\)\\\|\_\{F\},with\(λR,λx\)=\(0\.5,0\.5\)\(\\lambda\_\{R\},\\lambda\_\{x\}\)=\(0\.5,0\.5\)\. Re\-pairing shortens the average transport distance and straightens the induced probability path; the resulting velocity field is closer to constant along each trajectory, which is exactly the regime in which few\-step average\-velocity sampling is accurate\. This is also why the controlled benchmark of Appendix[H](https://arxiv.org/html/2607.27431#A8)disables OT \(Section[H\.2](https://arxiv.org/html/2607.27431#A8.SS2)\): with OT the instantaneous\- and average\-velocity parameterisations nearly coincide, and the comparison would no longer be informative\.

## Appendix ERelated Work

### E\.1Diffusion and early flow\-matching baselines

The first wave ofSE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}backbone generators are diffusion models over residue frames\. FrameDiffYimet al\.\[[2023b](https://arxiv.org/html/2607.27431#bib.bib8)\]formulates SE\(3\)\-invariant score\-based diffusion on multiple frames and generates designable monomers up to500500residues without a pretrained structure\-prediction network \(17\.417\.4M parameters\)\. GenieLinet al\.\[[2024](https://arxiv.org/html/2607.27431#bib.bib38)\]instead diffuses orientedCα\\mathrm\{C\}\_\{\\alpha\}residue clouds with SE\(3\)\-equivariant triangle updates; it is parameter\-light \(4\.14\.1M\) and attains high diversity and novelty, but its designability is comparatively low and degrades sharply as the number of sampling steps is reduced\. FrameFlowYimet al\.\[[2023a](https://arxiv.org/html/2607.27431#bib.bib37)\]keeps the frame representation but replaces diffusion withSE​\(3\)\\mathrm\{SE\}\(3\)flow matching, reporting roughly2×2\\timeshigher designability than FrameDiff at about5×5\\timesfewer sampling steps, and a∼23×\\sim\\\!23\\timessampling speed\-up over Genie at markedly higher designability\. Because FrameFlow dominates both earlier models on the designability–efficiency axis that is our focus, we take it as the representative of this pre\-2024 generation in the main\-text comparison \(Table[1](https://arxiv.org/html/2607.27431#S4.T1)\) and do not separately tabulate FrameDiff or Genie\. The flow\-matching methods most closely related to ours—FoldFlow, ReQFlow, and Riemannian MeanFlow—are discussed next\.

### E\.2FoldFlowBoseet al\.\[[2024](https://arxiv.org/html/2607.27431#bib.bib9)\]

Boseet al\.\[[2024](https://arxiv.org/html/2607.27431#bib.bib9)\]introducesSE​\(3\)\\mathrm\{SE\}\(3\)stochastic flow matching for protein backbone generation\. Their training objective is a flow\-matching regression that fits a time\-dependent vector field\(vθR,vθx\)\(v^\{R\}\_\{\\theta\},v^\{x\}\_\{\\theta\}\)to the drift of a chosen probability path \(anSE​\(3\)\\mathrm\{SE\}\(3\)bridge\) between data\(R0,x0\)\(R\_\{0\},x\_\{0\}\)and noise\(R1,x1\)\(R\_\{1\},x\_\{1\}\):

ℒFM\(θ\)=𝔼t,\(R0,x0\),\(R1,x1\)\[∥vθR\(Rt,t\)−vR\(Rt,t∣R0,R1\)∥SO​\(3\)2\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{FM\}\}\(\\theta\)=\\mathbb\{E\}\_\{t,\(R\_\{0\},x\_\{0\}\),\(R\_\{1\},x\_\{1\}\)\}\\left\[\\\|v\_\{\\theta\}^\{R\}\(R\_\{t\},t\)\-v^\{R\}\(R\_\{t\},t\\mid R\_\{0\},R\_\{1\}\)\\\|^\{2\}\_\{\\mathrm\{SO\}\(3\)\}\\right\.\+∥vθx\(xt,t\)−vtx\(xt,t\|x0,x1\)∥2\]\\displaystyle\\hskip 40\.00006pt\\left\.\+\\\|v\_\{\\theta\}^\{x\}\(x\_\{t\},t\)\-v^\{x\}\_\{t\}\(x\_\{t\},t\|x\_\{0\},x\_\{1\}\)\\\|^\{2\}\\right\]where​vtR​\(Rt\|R0,R1\):=R˙t∣R0,R1:=Rt​Ωt,Ωt=log⁡\(R0⊤​R1\),Rt=R0​et​Ωt\\displaystyle\\text\{where \}v^\{R\}\_\{t\}\(R\_\{t\}\|R\_\{0\},R\_\{1\}\):=\\dot\{R\}\_\{t\}\\mid\_\{R\_\{0\},R\_\{1\}\}:=R\_\{t\}\\Omega\_\{t\},\\Omega\_\{t\}=\\log\(R\_\{0\}^\{\\top\}R\_\{1\}\),R\_\{t\}=R\_\{0\}e^\{t\\Omega\_\{t\}\}vtx​\(xt\|x0,x1\)=x1−x0,xt=\(1−t\)​x0\+t​x1\.\\displaystyle\\hskip 28\.00006ptv^\{x\}\_\{t\}\(x\_\{t\}\|x\_\{0\},x\_\{1\}\)=x\_\{1\}\-x\_\{0\},x\_\{t\}=\(1\-t\)x\_\{0\}\+tx\_\{1\}\.In this work, the authors introduce three variants:

- •In the base setting,qqis the independent coupling between\(p0=pd​a​t​a,p1=pp​r​i​o​r\)\(p\_\{0\}=p\_\{data\},p\_\{1\}=p\_\{prior\}\)\.
- •In addition, the authors introduce the \(mini\-batch\) optimal transport coupling, obtained by batch\-size samplep0B∼i\.i\.d\.​p0,p1B∼i\.i\.d\.​p1p\_\{0\}^\{B\}\\sim\\text\{i\.i\.d\. \}p\_\{0\},p\_\{1\}^\{B\}\\sim\\text\{i\.i\.d\. \}p\_\{1\}and the method is denoted as FoldFlow\-OT\.
- •Furthermore, they introduce perturbation in the rotation interpretationRtR\_\{t\}by the IGSO\-distribution: R~t∣R0,R1∼ℐ​𝒢S​O​\(3\)​\(Rt,γ2​\(t\)​t​\(1−t\)\)\.\\tilde\{R\}\_\{t\}\\mid\_\{R\_\{0\},R\_\{1\}\}\\sim\\mathcal\{IG\}\_\{SO\(3\)\}\(R\_\{t\},\\gamma^\{2\}\(t\)t\(1\-t\)\)\.whereγ2​\(t\)\>0\\gamma^\{2\}\(t\)\>0is a predefined function\.

Adaptation to protein backbone space:SE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}As discussed in the main text, a protein backbone withNNresidues can be represented as a product spaceSE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}, i\.e\.,x=\(g1,…,gN\)x=\(g\_\{1\},\\dots,g\_\{N\}\)withgi=\(Ri,xi\)∈SE​\(3\)g\_\{i\}=\(R\_\{i\},x\_\{i\}\)\\in\\mathrm\{SE\}\(3\)\. TheSE​\(3\)\\mathrm\{SE\}\(3\)flow\-matching loss then extends by summing \(or averaging\) the per\-residueSE​\(3\)\\mathrm\{SE\}\(3\)losses, which is equivalent to*concatenating*all residue\-wise tangent vectors into a single6​N6N\-dimensional stacked vector:

ℒFMSE​\(3\)N​\(θ\)=\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{FM\}\}^\{\\mathrm\{SE\}\(3\)^\{N\}\}\(\\theta\)=𝔼1N\[∑i=1N∥vθ,iR\(Rt,i,t\)−vt,iR\(Rt,i∣R0,i,R1,i\)∥SO​\(3\)2\\displaystyle\\mathbb\{E\}\\frac\{1\}\{N\}\\Big\[\\,\\sum\_\{i=1\}^\{N\}\\\|v^\{R\}\_\{\\theta,i\}\(R\_\{t,i\},t\)\-v^\{R\}\_\{t,i\}\(R\_\{t,i\}\\mid R\_\{0,i\},R\_\{1,i\}\)\\\|^\{2\}\_\{\\mathrm\{SO\}\(3\)\}\+∑i=1N∥vθ,ix\(xt,i,t\)−vt,ix\(xt,i∣x0,i,x1,i\)∥2\]\\displaystyle\\qquad\+\\sum\_\{i=1\}^\{N\}\\\|v^\{x\}\_\{\\theta,i\}\(x\_\{t,i\},t\)\-v^\{x\}\_\{t,i\}\(x\_\{t,i\}\\mid x\_\{0,i\},x\_\{1,i\}\)\\\|^\{2\}\\,\\Big\]
In addition, the auxiliary loss \(see section[G\.1](https://arxiv.org/html/2607.27431#A7.SS1)\) is included in the final training loss\.

### E\.3FoldFlow\-2Huguetet al\.\[[2024](https://arxiv.org/html/2607.27431#bib.bib21)\]

Huguetet al\.\[[2024](https://arxiv.org/html/2607.27431#bib.bib21)\]extends FoldFlow to a sequence\-conditioned generative model\. The underlying generative task remains SE\(3\)Nflow matching as in FoldFlow, and the loss structure is identical\. The key novelty is that the vector field is now conditioned on a \(possibly masked\) amino acid sequencea¯\\bar\{a\}:

ℒFMFF2\(θ\)=𝔼t,π​\(x0,x1\),a¯1N\[\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{FM\}\}^\{\\mathrm\{FF2\}\}\(\\theta\)=\\mathbb\{E\}\_\{t,\\,\\pi\(x\_\{0\},x\_\{1\}\),\\,\\bar\{a\}\}\\frac\{1\}\{N\}\\Big\[∑i=1N∥vθ,iR\(Rt,i,t\|a¯\)−vt,iR\(Rt,i∣R0,i,R1,i,a¯\)∥SO​\(3\)2\\displaystyle\\sum\_\{i=1\}^\{N\}\\\|v^\{R\}\_\{\\theta,i\}\(R\_\{t,i\},t\|\\bar\{a\}\)\-v^\{R\}\_\{t,i\}\(R\_\{t,i\}\\mid R\_\{0,i\},R\_\{1,i\},\\bar\{a\}\)\\\|^\{2\}\_\{\\mathrm\{SO\}\(3\)\}\+\\displaystyle\+∑i=1N∥vθ,ix\(xt,i,t\|a¯\)−vt,ix\(xt,i∣x0,i,x1,i,a¯\)∥2\],\\displaystyle\\sum\_\{i=1\}^\{N\}\\\|v^\{x\}\_\{\\theta,i\}\(x\_\{t,i\},t\|\\bar\{a\}\)\-v^\{x\}\_\{t,i\}\(x\_\{t,i\}\\mid x\_\{0,i\},x\_\{1,i\},\\bar\{a\}\)\\\|^\{2\}\\Big\],\(26\)
whereπ​\(x0,x1\)\\pi\(x\_\{0\},x\_\{1\}\)is the minibatch OT coupling \(as in FoldFlow\-OT\), anda¯=a⊙m\\bar\{a\}=a\\odot mis the masked sequence with maskm∼Bern​\(0\.5\)m\\sim\\mathrm\{Bern\}\(0\.5\)applied uniformly across all residues\. This stochastic masking enables a single model to handle both conditional and unconditional generation:

- •a¯=\[∅\]N\\bar\{a\}=\[\\varnothing\]^\{N\}\(fully masked, probability 0\.5\): unconditional backbone generation, equivalent to FoldFlow\-OT\.
- •a¯=a\\bar\{a\}=a\(unmasked, probability 0\.5\): sequence\-conditioned generation, i\.e\. protein folding\.
- •a¯=a⊙m\\bar\{a\}=a\\odot m\(partially masked\): structure in\-painting and motif scaffolding\.

Reinforced Fine\-Tuning \(ReFT\)\.FoldFlow\-2 further introduces a fine\-tuning objective to align generations towards an auxiliary rewardrauxr\_\{\\mathrm\{aux\}\}\. Given a preferential dataset𝒟pref\\mathcal\{D\}\_\{\\mathrm\{pref\}\}filtered byrauxr\_\{\\mathrm\{aux\}\}, the ReFT objective is:

maxpθ⁡ℒReFT​\(θ\)=𝔼\(x,a\)∼𝒟pref​\[raux​\(x\)​log⁡pθ​\(x∣a\)\]\.\\displaystyle\\max\_\{p\_\{\\theta\}\}\\;\\mathcal\{L\}\_\{\\mathrm\{ReFT\}\}\(\\theta\)=\\mathbb\{E\}\_\{\(x,a\)\\sim\\mathcal\{D\}\_\{\\mathrm\{pref\}\}\}\\left\[r\_\{\\mathrm\{aux\}\}\(x\)\\log p\_\{\\theta\}\(x\\mid a\)\\right\]\.\(27\)This is applied, for example, to improve secondary structure diversity by upweighting samples rich inβ\\beta\-sheets and coils\.

### E\.4ReQFlowYueet al\.\[[2025](https://arxiv.org/html/2607.27431#bib.bib22)\]

Yueet al\.\[[2025](https://arxiv.org/html/2607.27431#bib.bib22)\]propose a quaternion\-based flow model for protein backbone generation\. Similar to FrameFlow and FoldFlow, the protein backbone is represented as a collection of residue\-wise rigid frames inSE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}\. However, instead of representing rotations by matrices inSO​\(3\)\\mathrm\{SO\}\(3\), ReQFlow parameterizes each residue frame as

gi=\(xi,qi\),xi∈ℝ3,qi∈𝕊3,g\_\{i\}=\(x\_\{i\},q\_\{i\}\),\\qquad x\_\{i\}\\in\\mathbb\{R\}^\{3\},\\;\\;q\_\{i\}\\in\\mathbb\{S\}^\{3\},wherexix\_\{i\}is the local translation andqiq\_\{i\}is a unit quaternion representing the 3D rotation\.

In particular, given a rotation matrixR∈SO​\(3\)R\\in\\mathrm\{SO\}\(3\), letω=ϕ​u=\(log⁡\(R\)\)∨∈ℝ3\\omega=\\phi u=\(\\log\(R\)\)^\{\\vee\}\\in\\mathbb\{R\}^\{3\}be its axis\-angle vector, whereϕ=‖ω‖,u=ω‖ω‖\\phi=\\\|\\omega\\\|,u=\\frac\{\\omega\}\{\\\|\\omega\\\|\}the corresponding unit quaternion is given by

q=exp⁡\(12​ω\)=\[cos⁡ϕ2,sin⁡ϕ2​u⊤\]⊤∈𝕊3\.\\displaystyle q=\\exp\(\\frac\{1\}\{2\}\\omega\)=\[\\cos\\frac\{\\phi\}\{2\},\\sin\\frac\{\\phi\}\{2\}u^\{\\top\}\]^\{\\top\}\\in\\mathbb\{S\}^\{3\}\.Conversely, given a unit quaternionq=\(w,x,y,z\)⊤∈𝕊3q=\(w,x,y,z\)^\{\\top\}\\in\\mathbb\{S\}^\{3\}, the corresponding rotation matrix is

R=exp⁡\(\(2​log⁡\(q\)\)∧\)\.R=\\exp\(\(2\\log\(q\)\)^\{\\wedge\}\)\.
Moreover, given 3D rotationsR1,R2R\_\{1\},R\_\{2\}with corresponding quaternion representationsq1:=\(s1,u1\),q2:=\(s2,u2\)q\_\{1\}:=\(s\_\{1\},u\_\{1\}\),q\_\{2\}:=\(s\_\{2\},u\_\{2\}\), their group action \(matrix multiplication\) can be expressed by quaternion multiplication:

q1⊗q2=\[s1​s2−u1⊤​u2s1​u2\+s2​u1\+u1⊗u2\]\.q\_\{1\}\\otimes q\_\{2\}=\\begin\{bmatrix\}s\_\{1\}s\_\{2\}\-u\_\{1\}^\{\\top\}u\_\{2\}\\\\ s\_\{1\}u\_\{2\}\+s\_\{2\}u\_\{1\}\+u\_\{1\}\\otimes u\_\{2\}\\end\{bmatrix\}\.
This quaternion algebra leads to a more numerically stable and efficient treatment of rotations, especially when the rotation angle is very small or close toπ\\pi\.

##### Quaternion Flow Matching \(QFlow\)\.

In ReQFlow, the authors denote by𝒯0\\mathcal\{T\}\_\{0\}the prior distribution

𝒩​\(0,I3\)×ℐ​𝒢S​O​\(3\),\\mathcal\{N\}\(0,I\_\{3\}\)\\times\\mathcal\{IG\}\_\{SO\(3\)\},where, with a slight abuse of notation,ℐ​𝒢S​O​\(3\)\\mathcal\{IG\}\_\{SO\(3\)\}denotes the isotropic Gaussian distribution in rotation space under the quaternion representation\. The target distribution𝒯1\\mathcal\{T\}\_\{1\}corresponds to the real protein data distribution\.

ReQFlow parameterizes the model as

Tθ,1​\(xt,qt\)=\(xθ,1,qθ,1\),T\_\{\\theta,1\}\(x\_\{t\},q\_\{t\}\)=\(x\_\{\\theta,1\},q\_\{\\theta,1\}\),wherexθ,1x\_\{\\theta,1\}andqθ,1q\_\{\\theta,1\}denote the predicted translation and rotation at terminal timet=1t=1, conditioned on the current state\(xt,qt\)\(x\_\{t\},q\_\{t\}\)\.

At timet∈\[0,1\]t\\in\[0,1\], the translation path follows the standard linear interpolation

xt=\(1−t\)​x0\+t​x1,x\_\{t\}=\(1\-t\)x\_\{0\}\+tx\_\{1\},with translation velocity

vtx=x1−x0=x1−xt1−t,vθ,tx=xθ,1−xt1−t\.v\_\{t\}^\{x\}=x\_\{1\}\-x\_\{0\}=\\frac\{x\_\{1\}\-x\_\{t\}\}\{1\-t\},\\qquad v\_\{\\theta,t\}^\{x\}=\\frac\{x\_\{\\theta,1\}\-x\_\{t\}\}\{1\-t\}\.
For the rotational component, ReQFlow adopts quaternion geodesic interpolation:

qt=q0⊗exp⁡\(t​log⁡\(q0−1⊗q1\)\)\.q\_\{t\}=q\_\{0\}\\otimes\\exp\\\!\\bigl\(t\\log\(q\_\{0\}^\{\-1\}\\otimes q\_\{1\}\)\\bigr\)\.The corresponding angular velocity and its model\-based estimate are

vtq=2​log⁡\(q0−1⊗q1\)=2​log⁡\(qt−1⊗q1\)1−t,vθ,tq=2​log⁡\(qt−1⊗qθ,1\)1−t\.v\_\{t\}^\{q\}=2\\log\(q\_\{0\}^\{\-1\}\\otimes q\_\{1\}\)=\\frac\{2\\log\(q\_\{t\}^\{\-1\}\\otimes q\_\{1\}\)\}\{1\-t\},v\_\{\\theta,t\}^\{q\}=\\frac\{2\\log\(q\_\{t\}^\{\-1\}\\otimes q\_\{\\theta,1\}\)\}\{1\-t\}\.
The inference dynamics are then approximated by

\{xt\+Δ​t=xt\+Δ​t​vθ,tx,qt\+Δ​t=qt⊗exp⁡\(12​Δ​t​vθ,tq\)\.\\begin\{cases\}x\_\{t\+\\Delta t\}=x\_\{t\}\+\\Delta t\\,v\_\{\\theta,t\}^\{x\},\\\\ q\_\{t\+\\Delta t\}=q\_\{t\}\\otimes\\exp\\\!\\left\(\\frac\{1\}\{2\}\\Delta t\\,v\_\{\\theta,t\}^\{q\}\\right\)\.\\end\{cases\}Accordingly, the flow\-matching loss is defined as

ℒQFlow=𝔼t,\(x0,q0\),\(x1,q1\)​\[‖vtx−vθ,tx‖2\]\+𝔼t,\(x0,q0\),\(x1,q1\)​\[‖vtq−vθ,tq‖2\]\.\\mathcal\{L\}\_\{\\mathrm\{QFlow\}\}=\\mathbb\{E\}\_\{t,\(x\_\{0\},q\_\{0\}\),\(x\_\{1\},q\_\{1\}\)\}\\bigl\[\\\|v^\{x\}\_\{t\}\-v^\{x\}\_\{\\theta,t\}\\\|^\{2\}\\bigr\]\+\\mathbb\{E\}\_\{t,\(x\_\{0\},q\_\{0\}\),\(x\_\{1\},q\_\{1\}\)\}\\bigl\[\\\|v^\{q\}\_\{t\}\-v^\{q\}\_\{\\theta,t\}\\\|^\{2\}\\bigr\]\.In addition, the auxiliary loss \(see section[G\.1](https://arxiv.org/html/2607.27431#A7.SS1)\) is incorporated into the training process\.

### E\.5Riemannian Gaussian Variational Flow Matching \(RG\-VFM\)Zaghenet al\.\[[2025](https://arxiv.org/html/2607.27431#bib.bib3)\]

Zaghenet al\.\[[2025](https://arxiv.org/html/2607.27431#bib.bib3)\]approach manifold generation from the*variational*rather than the velocity\-matching side\. Building on Variational Flow Matching, they replace the instantaneous\-velocity regression used by CFM/RFM—and, in our setting, by FrameFlow, FoldFlow, and ReQFlow—with an*endpoint*objective: a network predicts a terminal meanμθ​\(xt\)∈ℳ\\mu\_\{\\theta\}\(x\_\{t\}\)\\in\\mathcal\{M\}, and the posteriorqθ​\(x1∣xt\)q\_\{\\theta\}\(x\_\{1\}\\mid x\_\{t\}\)is modeled as a Riemannian Gaussian

𝒩Riem​\(x1∣μθ​\(xt\),σ\)∝exp⁡\(−distg​\(x1,μθ​\(xt\)\)22​σ2\)\.\\mathcal\{N\}\_\{\\mathrm\{Riem\}\}\\big\(x\_\{1\}\\mid\\mu\_\{\\theta\}\(x\_\{t\}\),\\sigma\\big\)\\;\\propto\\;\\exp\\\!\\Big\(\-\\tfrac\{\\mathrm\{dist\}\_\{g\}\\big\(x\_\{1\},\\mu\_\{\\theta\}\(x\_\{t\}\)\\big\)^\{2\}\}\{2\\sigma^\{2\}\}\\Big\)\.On a homogeneous manifold with closed\-form geodesics, the normalizing constant is independent ofμ\\mu, and the objective collapses to a squared geodesic \(endpoint\) distance:

ℒRG​\-​VFM​\(θ\)=𝔼t,x1,xt​\[‖logx1⁡\(μθ​\(xt\)\)‖g2\]=𝔼t,x1,xt​\[distg​\(x1,μθ​\(xt\)\)2\],\\mathcal\{L\}\_\{\\mathrm\{RG\\text\{\-\}VFM\}\}\(\\theta\)=\\mathbb\{E\}\_\{t,x\_\{1\},x\_\{t\}\}\\big\[\\\|\\log\_\{x\_\{1\}\}\(\\mu\_\{\\theta\}\(x\_\{t\}\)\)\\\|\_\{g\}^\{2\}\\big\]=\\mathbb\{E\}\_\{t,x\_\{1\},x\_\{t\}\}\\big\[\\mathrm\{dist\}\_\{g\}\\big\(x\_\{1\},\\mu\_\{\\theta\}\(x\_\{t\}\)\\big\)^\{2\}\\big\],which recovers the Euclidean VFM/MSE loss‖μθ​\(xt\)−x1‖2\\\|\\mu\_\{\\theta\}\(x\_\{t\}\)\-x\_\{1\}\\\|^\{2\}whenℳ=ℝd\\mathcal\{M\}=\\mathbb\{R\}^\{d\}\. Unlike vanilla RFM, whose vector field lives inTxt​ℳT\_\{x\_\{t\}\}\\mathcal\{M\}and requiressupp​\(p0\)\\mathrm\{supp\}\(p\_\{0\}\)to lie onℳ\\mathcal\{M\}, the variational objective compares endpoints in the single tangent spaceTx1​ℳT\_\{x\_\{1\}\}\\mathcal\{M\}and only needs the local geometry aroundp1p\_\{1\}\.

Adaptation to the backbone setting \(“variational ReQFlow”\)\.Instantiating this objective in ReQFlow’s frame representationgi=\(xi,qi\)∈ℝ3×𝕊3g\_\{i\}=\(x\_\{i\},q\_\{i\}\)\\in\\mathbb\{R\}^\{3\}\\times\\mathbb\{S\}^\{3\}yields an endpoint\-matching counterpart of QFlow\. Rather than regressing the translation/quaternion velocities as inℒQFlow\\mathcal\{L\}\_\{\\mathrm\{QFlow\}\}, the model predicts the terminal frame\(xθ,1,qθ,1\)\(x\_\{\\theta,1\},q\_\{\\theta,1\}\)and minimizes the endpoint geodesic distance directly:

ℒv​\-​ReQFlow​\(θ\)=𝔼t​\[∑i=1N‖x1,i−xθ,1,i‖2\+λ​∑i=1NdistSO​\(3\)​\(q1,i,qθ,1,i\)2\],distSO​\(3\)​\(q1,qθ,1\)=2​‖log⁡\(qθ,1−1⊗q1\)‖,\\mathcal\{L\}\_\{\\mathrm\{v\\text\{\-\}ReQFlow\}\}\(\\theta\)=\\mathbb\{E\}\_\{t\}\\Big\[\\,\\sum\_\{i=1\}^\{N\}\\\|x\_\{1,i\}\-x\_\{\\theta,1,i\}\\\|^\{2\}\+\\lambda\\sum\_\{i=1\}^\{N\}\\mathrm\{dist\}\_\{\\mathrm\{SO\}\(3\)\}\\big\(q\_\{1,i\},q\_\{\\theta,1,i\}\\big\)^\{2\}\\,\\Big\],\\qquad\\mathrm\{dist\}\_\{\\mathrm\{SO\}\(3\)\}\(q\_\{1\},q\_\{\\theta,1\}\)=2\\big\\\|\\log\(q\_\{\\theta,1\}^\{\-1\}\\otimes q\_\{1\}\)\\big\\\|,i\.e\. the flow\-matching term is replaced by a squared endpoint distance on the product spaceSE​\(3\)N=\(SO​\(3\)×ℝ3\)N\\mathrm\{SE\}\(3\)^\{N\}=\(\\mathrm\{SO\}\(3\)\\times\\mathbb\{R\}^\{3\}\)^\{N\}\.

Relation to our work\.At the loss level, RG\-VFM can be read as a variant of ReQFlow in which the velocity\-matching FM objective is replaced by an endpoint geodesic distance: the network still predicts a terminal meanμθ​\(xt\)\\mu\_\{\\theta\}\(x\_\{t\}\)\(anx1x\_\{1\}\-prediction parameterization\), so sampling remains an ODE integration and inherits the same limited few\-step generation behavior as standard flow matching\. Our method is instead a MeanFlow\-type model that directly learns the average\-velocity flow mapΦs,t\\Phi\_\{s,t\}via the time\-ordered exponential \(Eq[30](https://arxiv.org/html/2607.27431#A6.E30)\), which is what enables few\-step generation\. The two are nonetheless complementary rather than opposed: in the second phase of our training \(Section[M](https://arxiv.org/html/2607.27431#A13)\), we combine this endpoint geodesic loss with the MeanFlow objective, using the former to stabilize the JVP \(total\-derivative\) term in the MeanFlow loss\. Finally, RG\-VFM is validated only on a synthetic spherical dataset and releases noSE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}protein checkpoint, so we do not include it as a standalone baseline \(see below\)\.

### E\.6Riemannian MeanFlow \(RMF\)Wooet al\.\[[2026](https://arxiv.org/html/2607.27431#bib.bib23)\]

Concurrent with our work,Wooet al\.\[[2026](https://arxiv.org/html/2607.27431#bib.bib23)\]generalize MeanFlow from Euclidean space to Riemannian manifolds, with applications to scientific generative modeling on manifolds such as the simplex andSE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}\. Instead of learning an instantaneous velocity field and numerically integrating it during sampling, RMF directly learns the flow map

Φs,t:ℳ→ℳ,\\Phi\_\{s,t\}:\\mathcal\{M\}\\to\\mathcal\{M\},which transports a pointxs∼psx\_\{s\}\\sim p\_\{s\}at timessto its corresponding pointxt∼ptx\_\{t\}\\sim p\_\{t\}at timettalong the same integral curve\.

The key geometric quantity in RMF is the*average velocity*, defined by

us,t​\(xs\)=\{1t−s​logxs⁡\(xt\),t≠s,vs​\(xs\),t=s,u\_\{s,t\}\(x\_\{s\}\)=\\begin\{cases\}\\dfrac\{1\}\{t\-s\}\\log\_\{x\_\{s\}\}\(x\_\{t\}\),&t\\neq s,\\\\\[6\.0pt\] v\_\{s\}\(x\_\{s\}\),&t=s,\\end\{cases\}\(28\)
wherelogxs⁡\(xt\)∈Txs​ℳ\\log\_\{x\_\{s\}\}\(x\_\{t\}\)\\in T\_\{x\_\{s\}\}\\mathcal\{M\}denotes the Riemannian logarithmic map\. Geometrically,us,t​\(xs\)u\_\{s,t\}\(x\_\{s\}\)is the constant tangent velocity that transportsxsx\_\{s\}toxtx\_\{t\}along a geodesic over time intervalt−st\-s\. By differentiating both sides with respect tott, we obtain:

us,t​\(xs\)=d​\(logxs\)xt​\[vt\]−\(t−s\)​∂tus,t​\(xs\)\.\\displaystyle u\_\{s,t\}\(x\_\{s\}\)=d\(\\log\_\{x\_\{s\}\}\)\_\{x\_\{t\}\}\[v\_\{t\}\]\-\(t\-s\)\\partial\_\{t\}u\_\{s,t\}\(x\_\{s\}\)\.\(29\)and it induces the training loss:

ℒRMF​\(θ\)=𝔼xs,s,t​\[‖us,tθ​\(xs\)−sg​\(u^t​g​t\)‖2\],\\mathcal\{L\}\_\{\\mathrm\{RMF\}\}\(\\theta\)=\\mathbb\{E\}\_\{x\_\{s\},s,t\}\[\\\|u\_\{s,t\}^\{\\theta\}\(x\_\{s\}\)\-\\text\{sg\}\(\\hat\{u\}\_\{tgt\}\)\\\|^\{2\}\],whereu^t​g​t=\(t−s\)​Ds​us,tθ​\(xs\)−∇vs1logxs⁡Φθ​\(xs\)\\hat\{u\}\_\{tgt\}=\(t\-s\)D\_\{s\}u\_\{s,t\}^\{\\theta\}\(x\_\{s\}\)\-\\nabla\_\{v\_\{s\}\}^\{1\}\\log\_\{x\_\{s\}\}\\Phi^\{\\theta\}\(x\_\{s\}\)andΦs,tθ​\(xs\)=exp⁡\(\(t−s\)​us,tθ​\(x\)\)\\Phi\_\{s,t\}^\{\\theta\}\(x\_\{s\}\)=\\exp\(\(t\-s\)u\_\{s,t\}^\{\\theta\}\(x\)\)\.

They further include a cycle\-consistency regularizer

ℒc​y​c​\(θ\)=𝔼xt,s,t​\[dg​\(ϕs,tθ​\(ϕt,sθ​\(xt\)\),xt\)2\]\.\\mathcal\{L\}\_\{cyc\}\(\\theta\)=\\mathbb\{E\}\_\{x\_\{t\},s,t\}\\big\[d\_\{g\}\(\\phi\_\{s,t\}^\{\\theta\}\(\\phi\_\{t,s\}^\{\\theta\}\(x\_\{t\}\)\),x\_\{t\}\)^\{2\}\\big\]\.
##### Training objective in the protein experiments\.

In theSE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}protein setting, RMF is trained with a flow\-matching loss together with a semigroup \(flow\-map composition\) consistency loss\. The semigroup term enforcesΦr,t=Φs,t∘Φr,s\\Phi\_\{r,t\}=\\Phi\_\{s,t\}\\circ\\Phi\_\{r,s\}, but measures the residual as an*endpoint geodesic distance*—theSE​\(3\)\\mathrm\{SE\}\(3\)distance between the one\-step mapΦr,t​\(xr\)\\Phi\_\{r,t\}\(x\_\{r\}\)and the composed two\-step mapΦs,t​\(Φr,s​\(xr\)\)\\Phi\_\{s,t\}\(\\Phi\_\{r,s\}\(x\_\{r\}\)\)—rather than regressing log\-displacements in the Lie algebra, as in our finite semigroup loss \(Appendix[K](https://arxiv.org/html/2607.27431#A11)\)\. The two agree at the minimizer, but differ in conditioning: the endpoint\-distance form couples the rotation and translation branches through a singleSE​\(3\)\\mathrm\{SE\}\(3\)metric, whereas our log\-displacement form keeps them separated and, on the rotation branch, makes the BCH/right\-Jacobian structure explicit\. We compare both average\-velocity formulations on a controlledSO​\(3\)2\\mathrm\{SO\}\(3\)^\{2\}benchmark in Appendix[H](https://arxiv.org/html/2607.27431#A8)\.

### E\.7Riemannian MeanFlow via Parallel Transport \(RMF\-PT\)

Also concurrent with our work,Zhonget al\.\[[2026](https://arxiv.org/html/2607.27431#bib.bib15)\]extend MeanFlow to general Riemannian manifolds by defining the average velocity as a parallel\-transport integral: the instantaneous velocities along the trajectory are transported to a common tangent space before being averaged \(Eq\. 5 therein\)\. To distinguish it from the Riemannian MeanFlow ofWooet al\.\[[2026](https://arxiv.org/html/2607.27431#bib.bib23)\]discussed above, we refer to this method asRMF\-PT\. Our approach differs in three respects\.

\(i\) Definition\.RMF\-PT defines the mean velocity through a path\-dependent parallel\-transport integral, which requires the full trajectory\{xτ\}τ∈\[s,t\]\\\{x\_\{\\tau\}\\\}\_\{\\tau\\in\[s,t\]\}\. We instead define it via the time\-ordered exponential \([30](https://arxiv.org/html/2607.27431#A6.E30)\), which depends only on the endpointsRsR\_\{s\}andRtR\_\{t\}and reduces to the logarithmic mapΩavg=1t−s​log⁡\(Rs⊤​Rt\)\\Omega^\{\\mathrm\{avg\}\}=\\tfrac\{1\}\{t\-s\}\\log\(R\_\{s\}^\{\\top\}R\_\{t\}\)\. The two definitions coincide when the velocity field is constant along the trajectory \(the geodesic case\); for general fields they differ by BCH correction terms arising from the non\-commutativity ofSO​\(3\)\\mathrm\{SO\}\(3\)\.

\(ii\) Identity and loss\.Differentiating the two definitions yields structurally different identities\. That of RMF\-PT involves the covariant derivative∇γ˙u\\nabla\_\{\\dot\{\\gamma\}\}uwith Christoffel corrections \(Appendix A ofZhonget al\.[2026](https://arxiv.org/html/2607.27431#bib.bib15)\), which their implementation drops via a log\-map approximation\. Ours involves the*exact*right JacobianJJofSO​\(3\)\\mathrm\{SO\}\(3\)\(Proposition[3](https://arxiv.org/html/2607.27431#Thmproposition3)\), which admits a closed form and requires no approximation\.

\(iii\) Application\.RMF\-PT targets general manifolds and is evaluated on synthetic data \(spheres, tori, andSO​\(3\)\\mathrm\{SO\}\(3\)rotations\)\. Our method is built forSE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}protein backbone generation: we exploit the decoupled product geometrySO​\(3\)×ℝ3\\mathrm\{SO\}\(3\)\\times\\mathbb\{R\}^\{3\}\(Appendix[C](https://arxiv.org/html/2607.27431#A3)\) to obtain separate, simulation\-free objectives for rotations and translations, and validate them on*de novo*backbone design\.

### E\.8Baseline selection

Since FoldFlow/FoldFlow2 and v\-QFlow/v\-ReQFlow do not have public checkpoints for the SCOPe dataset, these methods are not included in the experiments\. In addition, FoldFlow and QFlow are essentially the same in nature—both are flow\-matching models—and differ only in the mathematical representation of rotations \(rotation matrices vs\. quaternions\) and in the backbone architecture\. Since QFlow outperforms FoldFlow on the PDB dataset, we consider QFlow already sufficiently representative of this family\. v\-QFlow, in turn, can be viewed as a variant of QFlow whose performance largely matches QFlow’s—strong at a moderate number of sampling steps \(e\.g\.100/500100/500steps\) but degrading as the step count is reduced\. We therefore select QFlow alone as the representative method\.

## Appendix FMeanFlow onSE​\(3\)\\mathrm\{SE\}\(3\): Derivation, Theoretical Properties, andNN\-component Implementation

### F\.1Model and loss function

We first introduce some fundamental results in Riemannian/SO3 manifold:

With these preliminaries in place, we now define the average angular velocity onSO​\(3\)\\mathrm\{SO\}\(3\)and derive the corresponding MeanFlow identity\.

We define theaverage velocityas

exp⁡\(\(t−s\)​Ωavg​\(s,t,Rt,xt\)⏟Ωs→t\):=𝒯​exp⁡\(∫stΩ​\(τ,Rτ,xτ\)​𝑑τ\)\\displaystyle\\exp\(\\underbrace\{\(t\-s\)\\Omega^\{\\text\{avg\}\}\(s,t,R\_\{t\},x\_\{t\}\)\}\_\{\\Omega^\{s\\to t\}\}\):=\\mathcal\{T\}\\exp\\Big\(\\int\_\{s\}^\{t\}\\Omega\(\\tau,R\_\{\\tau\},x\_\{\\tau\}\)\\,d\\tau\\Big\)\(30\)
Note that since eachΩ​\(τ,Rτ,xτ\)∈𝔰​𝔬​\(3\)\\Omega\(\\tau,R\_\{\\tau\},x\_\{\\tau\}\)\\in\\mathfrak\{so\}\(3\), we haveΩavg​\(s,t,Rt,xt\)∈𝔰​𝔬​\(3\)\\Omega^\{\\text\{avg\}\}\(s,t,R\_\{t\},x\_\{t\}\)\\in\\mathfrak\{so\}\(3\)\.

In the extreme case, givens=0,t=1s=0,t=1, the above mean velocity recoversR0R\_\{0\}fromR1R\_\{1\}:

\{R1=R0​exp⁡\(Ωavg​\(0,1,R1,x1\)\)ForwardR0=R1​exp⁡\(−Ωavg​\(0,1,R1,x1\)\)Backward\\begin\{cases\}R\_\{1\}=R\_\{0\}\\exp\(\\Omega^\{\\text\{avg\}\}\(0,1,R\_\{1\},x\_\{1\}\)\)&\\text\{Forward\}\\\\ R\_\{0\}=R\_\{1\}\\exp\(\-\\Omega^\{\\text\{avg\}\}\(0,1,R\_\{1\},x\_\{1\}\)\)&\\text\{Backward\}\\end\{cases\}We apply this convention to the ground truth as well:

Ωavg​\(s,t,Rt,xt\)=A​\(s,t,Rt,xt\)∧\.\\Omega^\{\\text\{avg\}\}\(s,t,R\_\{t\},x\_\{t\}\)=A\(s,t,R\_\{t\},x\_\{t\}\)^\{\\wedge\}\.
With this definition, we can now derive a differential identity relating the average angular velocity to the instantaneous velocity, which will form the basis of our training objective\.

###### Proposition 3\(Derivative of the averaged angular velocity\)\.

Under the assumptions,ts,tare independent and using the above notations, setAs→t=\(t−s\)​Aavg​\(s,t,Rt,xt\)A^\{s\\to t\}=\(t\-s\)A^\{\\text\{avg\}\}\(s,t,R\_\{t\},x\_\{t\}\)\. Taking the derivative of both sides of \([30](https://arxiv.org/html/2607.27431#A6.E30)\) with respect tottyields

\{L\.H\.S\.=Rs⊤​Rt​\(J​\(As→t\)​dd​t​As→t\)∧R\.H\.S\.=Rs⊤​Rt​Ω​\(t,Rt,xt\)\.\\displaystyle\\begin\{cases\}\\text\{L\.H\.S\.\}&=R\_\{s\}^\{\\top\}R\_\{t\}\\left\(J\(A^\{s\\to t\}\)\\frac\{d\}\{dt\}A^\{s\\to t\}\\right\)^\{\\wedge\}\\\\ \\text\{R\.H\.S\.\}&=R\_\{s\}^\{\\top\}R\_\{t\}\\Omega\(t,R\_\{t\},x\_\{t\}\)\\end\{cases\}\.\(31\)SinceRs⊤​RtR\_\{s\}^\{\\top\}R\_\{t\}is invertible, equating the two sides and applying\(⋅\)∨\(\\cdot\)^\{\\vee\}gives

J​\(As→t\)​dd​t​As→t=ωt⟺dd​t​As→t=J−1​\(As→t\)​ωt,\\displaystyle J\(A^\{s\\to t\}\)\\,\\tfrac\{d\}\{dt\}A^\{s\\to t\}=\\omega\_\{t\}\\qquad\\Longleftrightarrow\\qquad\\tfrac\{d\}\{dt\}A^\{s\\to t\}=J^\{\-1\}\(A^\{s\\to t\}\)\\,\\omega\_\{t\},\(32\)the two forms underlying \([11](https://arxiv.org/html/2607.27431#S3.E11)\) and \([12](https://arxiv.org/html/2607.27431#S3.E12)\) respectively; the equivalence uses invertibility ofJJ\(Remark[6](https://arxiv.org/html/2607.27431#Thmremark6)\)\. Expandingdd​t​As→t=Aavg\+\(t−s\)​dd​t​Aavg\\tfrac\{d\}\{dt\}A^\{s\\to t\}=A^\{\\text\{avg\}\}\+\(t\-s\)\\tfrac\{d\}\{dt\}A^\{\\text\{avg\}\}by \([38](https://arxiv.org/html/2607.27431#A6.E38)\) recovers \([8](https://arxiv.org/html/2607.27431#S3.E8)\)\.

###### Proof\.

Differentiate both sides of \([30](https://arxiv.org/html/2607.27431#A6.E30)\) with respect tott\.

For the right\-hand side, the standard derivative rule for the time\-ordered exponential gives

dd​t​𝒯​exp⁡\(∫stΩ​\(τ,Rτ,xτ\)​𝑑τ\)\\displaystyle\\frac\{d\}\{dt\}\\,\\mathcal\{T\}\\exp\\Big\(\\int\_\{s\}^\{t\}\\Omega\(\\tau,R\_\{\\tau\},x\_\{\\tau\}\)\\,d\\tau\\Big\)=\(𝒯​exp⁡\(∫stΩ​\(τ,Rτ,xτ\)​𝑑τ\)\)​Ω​\(t,Rt,xt\)\\displaystyle=\\Big\(\\mathcal\{T\}\\exp\\big\(\\int\_\{s\}^\{t\}\\Omega\(\\tau,R\_\{\\tau\},x\_\{\\tau\}\)\\,d\\tau\\big\)\\Big\)\\,\\Omega\(t,R\_\{t\},x\_\{t\}\)=exp⁡\(Ωs→t\)​Ω​\(t,Rt,xt\),\\displaystyle=\\exp\(\\Omega^\{s\\to t\}\)\\,\\Omega\(t,R\_\{t\},x\_\{t\}\),\(33\)whereexp⁡\(Ωs→t\)=Rs⊤​Rt\\exp\(\\Omega^\{s\\to t\}\)=R\_\{s\}^\{\\top\}R\_\{t\}by definition\.

For the left\-hand side, the Wilcox formula for the Fréchet derivative of the matrix exponential gives

dd​t​exp⁡\(Ωs→t\)\\displaystyle\\frac\{d\}\{dt\}\\exp\(\\Omega^\{s\\to t\}\)=exp⁡\(Ωs→t\)​∫01exp⁡\(−α​Ωs→t\)​\(dd​t​Ωs→t\)​exp⁡\(α​Ωs→t\)​𝑑α\\displaystyle=\\exp\(\\Omega^\{s\\to t\}\)\\int\_\{0\}^\{1\}\\exp\\big\(\-\\alpha\\Omega^\{s\\to t\}\\big\)\\Big\(\\frac\{d\}\{dt\}\\Omega^\{s\\to t\}\\Big\)\\exp\\big\(\\alpha\\Omega^\{s\\to t\}\\big\)\\,d\\alpha\(34\)=exp⁡\(Ωs→t\)​∫01exp⁡\(−α​Ωs→t\)​\(dd​t​As→t\)∧​exp⁡\(α​Ωs→t\)​𝑑α\\displaystyle=\\exp\(\\Omega^\{s\\to t\}\)\\int\_\{0\}^\{1\}\\exp\\big\(\-\\alpha\\Omega^\{s\\to t\}\\big\)\\Big\(\\frac\{d\}\{dt\}A^\{s\\to t\}\\Big\)^\{\\wedge\}\\exp\\big\(\\alpha\\Omega^\{s\\to t\}\\big\)\\,d\\alpha=exp⁡\(Ωs→t\)​∫01\(exp⁡\(−α​\(As→t\)∧\)​dd​t​As→t\)∧​𝑑α\\displaystyle=\\exp\(\\Omega^\{s\\to t\}\)\\int\_\{0\}^\{1\}\\left\(\\exp\\big\(\-\\alpha\(A^\{s\\to t\}\)^\{\\wedge\}\\big\)\\,\\frac\{d\}\{dt\}A^\{s\\to t\}\\right\)^\{\\wedge\}\\,d\\alpha=exp⁡\(Ωs→t\)​\(∫01exp⁡\(−α​\(As→t\)∧\)​𝑑α​dd​t​As→t\)∧\.\\displaystyle=\\exp\(\\Omega^\{s\\to t\}\)\\left\(\\int\_\{0\}^\{1\}\\exp\\big\(\-\\alpha\(A^\{s\\to t\}\)^\{\\wedge\}\\big\)\\,d\\alpha\\;\\frac\{d\}\{dt\}A^\{s\\to t\}\\right\)^\{\\wedge\}\.\(35\)Here the second line usesΩs→t=\(As→t\)∧\\Omega^\{s\\to t\}=\(A^\{s\\to t\}\)^\{\\wedge\}and the linearity of\(⋅\)∧\(\\cdot\)^\{\\wedge\}\. The third line uses the conjugation identityQ​x∧​Q⊤=\(Q​x\)∧Qx^\{\\wedge\}Q^\{\\top\}=\(Qx\)^\{\\wedge\}forQ∈SO​\(3\)Q\\in\\mathrm\{SO\}\(3\)andx∈ℝ3x\\in\\mathbb\{R\}^\{3\}\. The last line follows because\(⋅\)∧\(\\cdot\)^\{\\wedge\}is linear and independent ofα\\alpha\. We complete the proof by the identity of right Jacobian defined in \([10](https://arxiv.org/html/2607.27431#S3.E10)\)

J​\(As→t\)=∫01exp⁡\(−α​\(As→t\)∧\)​𝑑α\.\\displaystyle J\(A^\{s\\to t\}\)=\\int\_\{0\}^\{1\}\\exp\(\-\\alpha\(A^\{s\\to t\}\)^\{\\wedge\}\)d\\alpha\.∎

![Refer to caption](https://arxiv.org/html/2607.27431v1/x6.png)Figure 5:Numerical verification of Proposition[3](https://arxiv.org/html/2607.27431#Thmproposition3)\. On an analytic, non\-geodesicSO​\(3\)\\mathrm\{SO\}\(3\)path we evaluate both sides of \([31](https://arxiv.org/html/2607.27431#A6.E31)\) and plot the median residual∥J​\(As→t\)​dd​t​As→t−ω​\(t\)∥\\lVert J\(A^\{s\\to t\}\)\\,\\tfrac\{d\}\{dt\}A^\{s\\to t\}\-\\omega\(t\)\\rVertagainst the segment gapt−st\-s\. With the Jacobian termJ​\(As→t\)J\(A^\{s\\to t\}\)\(green\) the residual stays at float64 machine precision \(∼10−15\\sim\\\!10^\{\-15\}, at worst10−1110^\{\-11\}\), uniformly in the gap; dropping it \(J:=IJ:=I, red\) breaks the identity by𝒪​\(t−s\)\\mathcal\{O\}\(t\-s\), confirming both the proposition and the necessity of the Jacobian term\.We verify Proposition[3](https://arxiv.org/html/2607.27431#Thmproposition3)numerically\. We sample a smooth, non\-geodesic curveR​\(τ\)=exp⁡\(\(a\+b​τ\+c​τ2\)∧\)R\(\\tau\)=\\exp\\big\(\(a\+b\\tau\+c\\tau^\{2\}\)^\{\\wedge\}\\big\)onSO​\(3\)\\mathrm\{SO\}\(3\)withb,cb,cnon\-parallel, so the averaged generatorAs→t=Log​\(Rs⊤​Rt\)A^\{s\\to t\}=\\mathrm\{Log\}\(R\_\{s\}^\{\\top\}R\_\{t\}\)genuinely differs from the instantaneous one andJ​\(As→t\)J\(A^\{s\\to t\}\)acts non\-trivially\. For a batch of times we compute the instantaneous body angular velocityω​\(t\):=\(Ω​\(t\)\)∨=\(Rt⊤​R˙t\)∨\\omega\(t\):=\\big\(\\Omega\(t\)\\big\)^\{\\vee\}=\(R\_\{t\}^\{\\top\}\\dot\{R\}\_\{t\}\)^\{\\vee\}and the derivativedd​t​As→t\\tfrac\{d\}\{dt\}A^\{s\\to t\}by forward\-mode automatic differentiation, and evaluateJ​\(As→t\)J\(A^\{s\\to t\}\)from its closed form, all using the sameSO​\(3\)\\mathrm\{SO\}\(3\)exp/log/Jacobian routines as the training loss in double precision\. Figure[5](https://arxiv.org/html/2607.27431#A6.F5)reports the median residual of \([31](https://arxiv.org/html/2607.27431#A6.E31)\) over the batch versust−st\-s\.

The identity in Proposition[3](https://arxiv.org/html/2607.27431#Thmproposition3)involves the total time derivativedd​t​As→t\\frac\{d\}\{dt\}A^\{s\\to t\}, which we now compute explicitly\.

###### Proposition 4\.

In the above notation, we have the total derivative

dd​t​A​\(s,t,Rt,xt\)\\displaystyle\\frac\{d\}\{dt\}A\(s,t,R\_\{t\},x\_\{t\}\)=∂tA​\(s,t,Rt,xt\)\+⟨∇RA​\(s,t,Rt,xt\),R˙t⟩\+⟨∇xA​\(s,t,Rt,xt\),x˙t⟩\.\\displaystyle=\\partial\_\{t\}A\(s,t,R\_\{t\},x\_\{t\}\)\+\\langle\\nabla\_\{R\}A\(s,t,R\_\{t\},x\_\{t\}\),\\dot\{R\}\_\{t\}\\rangle\+\\langle\\nabla\_\{x\}A\(s,t,R\_\{t\},x\_\{t\}\),\\dot\{x\}\_\{t\}\\rangle\.

###### Proof\.

By the above remark, forv∈TRt​SO​\(3\)v\\in T\_\{R\_\{t\}\}\\mathrm\{SO\}\(3\),

dR​A​\(s,t,Rt,xt\)​\[v\]=∇RA​\(s,t,Rt,xt\)​\[v\]\.d\_\{R\}A\(s,t,R\_\{t\},x\_\{t\}\)\[v\]=\\nabla\_\{R\}A\(s,t,R\_\{t\},x\_\{t\}\)\[v\]\.Similarly, in the Euclidean component,dx​A​\(s,t,Rt,xt\)​\[w\]=⟨∇xA​\(s,t,Rt,xt\),w⟩d\_\{x\}A\(s,t,R\_\{t\},x\_\{t\}\)\[w\]=\\langle\\nabla\_\{x\}A\(s,t,R\_\{t\},x\_\{t\}\),w\\rangleforw∈ℝ3w\\in\\mathbb\{R\}^\{3\}\. Applying the chain rule along the curvet↦\(Rt,xt\)t\\mapsto\(R\_\{t\},x\_\{t\}\)gives the claimed identity\. ∎

SubstitutingAθavg​\(s,t,Rt,xt\)A\_\{\\theta\}^\{\\text\{avg\}\}\(s,t,R\_\{t\},x\_\{t\}\)into the two identities of \([32](https://arxiv.org/html/2607.27431#A6.E32)\) and definingAθs→t=\(t−s\)​Aθavg​\(s,t,Rt,xt\)A\_\{\\theta\}^\{s\\to t\}=\(t\-s\)A\_\{\\theta\}^\{\\text\{avg\}\}\(s,t,R\_\{t\},x\_\{t\}\), we obtain the two training losses of Section[3\.1](https://arxiv.org/html/2607.27431#S3.SS1):

ℒSO​\(3\)J\\displaystyle\\mathcal\{L\}^\{J\}\_\{\\mathrm\{SO\}\(3\)\}:=𝔼s<t,\(\(R0,x0\),\(R1,x1\)\)∼q0,1​\[‖sg​\(J​\(Aθs→t\)\)​\(Aθavg\+\(t−s\)​sg​\(dd​t​Aθavg\)\)−ωt‖22\],\\displaystyle:=\\mathbb\{E\}\_\{s<t,\\,\(\(R\_\{0\},x\_\{0\}\),\(R\_\{1\},x\_\{1\}\)\)\\sim q\_\{0,1\}\}\\Big\[\\big\\\|\\,\\mathrm\{sg\}\\big\(J\(A\_\{\\theta\}^\{s\\to t\}\)\\big\)\\big\(A\_\{\\theta\}^\{\\text\{avg\}\}\+\(t\-s\)\\,\\mathrm\{sg\}\(\\tfrac\{d\}\{dt\}A\_\{\\theta\}^\{\\text\{avg\}\}\)\\big\)\-\\omega\_\{t\}\\,\\big\\\|\_\{2\}^\{2\}\\Big\],\(36\)ℒSO​\(3\)J−1\\displaystyle\\mathcal\{L\}^\{J^\{\-1\}\}\_\{\\mathrm\{SO\}\(3\)\}:=𝔼s<t,\(\(R0,x0\),\(R1,x1\)\)∼q0,1​\[‖Aθavg−sg​\(J−1​\(Aθs→t\)​ωt−\(t−s\)​dd​t​Aθavg\)‖22\],\\displaystyle:=\\mathbb\{E\}\_\{s<t,\\,\(\(R\_\{0\},x\_\{0\}\),\(R\_\{1\},x\_\{1\}\)\)\\sim q\_\{0,1\}\}\\Big\[\\big\\\|\\,A\_\{\\theta\}^\{\\text\{avg\}\}\-\\mathrm\{sg\}\\big\(J^\{\-1\}\(A\_\{\\theta\}^\{s\\to t\}\)\\,\\omega\_\{t\}\-\(t\-s\)\\tfrac\{d\}\{dt\}A\_\{\\theta\}^\{\\text\{avg\}\}\\big\)\\,\\big\\\|\_\{2\}^\{2\}\\Big\],\(37\)whereq0,1q\_\{0,1\}is a coupling betweenp0=pd​a​t​ap\_\{0\}=p\_\{data\}andp1=pp​r​i​o​r=pn​o​i​s​ep\_\{1\}=p\_\{prior\}=p\_\{noise\}\(e\.g\. product measure or optimal transport coupling\), and the stop\-gradient is placed so that the gradient flows throughAθavgA\_\{\\theta\}^\{\\text\{avg\}\}only\. The two residuals are related byJ​\(Aθs→t\)J\(A\_\{\\theta\}^\{s\\to t\}\)and hence share their zero set; we train with \([37](https://arxiv.org/html/2607.27431#A6.E37)\)\. In addition, the interpolationR​\(t\)R\(t\)\(and its velocityR˙​\(t\)\\dot\{R\}\(t\)\) is obtained from the geodesic interpolation:

Ω​\(t\)=Log​\(R0⊤​R1\),R​\(t\)=R​\(0\)​exp⁡\(t​Ω​\(t\)\),R˙t=Rt​Ω​\(t\)\.\\Omega\(t\)=\\text\{Log\}\(R\_\{0\}^\{\\top\}R\_\{1\}\),\\quad R\(t\)=R\(0\)\\exp\(t\\Omega\(t\)\),\\quad\\dot\{R\}\_\{t\}=R\_\{t\}\\Omega\(t\)\.
The following proposition confirms that minimizing this loss recovers the correct relative rotation between any two points along the path\.

###### Proposition 5\.

IfℒSO​\(3\)J=0\\mathcal\{L\}^\{J\}\_\{\\mathrm\{SO\}\(3\)\}=0orℒSO​\(3\)J−1=0\\mathcal\{L\}^\{J^\{\-1\}\}\_\{\\mathrm\{SO\}\(3\)\}=0, i\.e\.

\(J​\(Aθs→t\)​dd​t​Aθs→t\)∧=Ω​\(t,Rt,xt\)∀s<t∈\[0,1\],\\big\(J\(A\_\{\\theta\}^\{s\\to t\}\)\\tfrac\{d\}\{dt\}A\_\{\\theta\}^\{s\\to t\}\\big\)^\{\\wedge\}=\\Omega\(t,R\_\{t\},x\_\{t\}\)\\qquad\\forall s<t\\in\[0,1\],then

exp⁡\(\(Aθs→t\)∧\)=Rs⊤​Rt\.\\exp\\big\(\(A\_\{\\theta\}^\{s\\to t\}\)^\{\\wedge\}\\big\)=R\_\{s\}^\{\\top\}R\_\{t\}\.

###### Proof\.

For simplicity we takes=0,t=1s=0,t=1\. Define the candidate reconstruction

R~t:=R0​exp⁡\(\(Aθ0→t\)∧\)\.\\tilde\{R\}\_\{t\}:=R\_\{0\}\\exp\(\(A\_\{\\theta\}^\{0\\to t\}\)^\{\\wedge\}\)\.By the closed\-form Fréchet derivative of the exponential onSO​\(3\)\\mathrm\{SO\}\(3\),

dd​t​exp⁡\(\(Aθ0→t\)∧\)=exp⁡\(\(Aθ0→t\)∧\)​\(J​\(Aθ0→t\)​dd​t​Aθ0→t\)∧\.\\frac\{d\}\{dt\}\\exp\(\(A\_\{\\theta\}^\{0\\to t\}\)^\{\\wedge\}\)=\\exp\(\(A\_\{\\theta\}^\{0\\to t\}\)^\{\\wedge\}\)\\big\(J\(A\_\{\\theta\}^\{0\\to t\}\)\\tfrac\{d\}\{dt\}A\_\{\\theta\}^\{0\\to t\}\\big\)^\{\\wedge\}\.Using the conditionℒSO​\(3\)J=0\\mathcal\{L\}^\{J\}\_\{\\mathrm\{SO\}\(3\)\}=0\(orℒSO​\(3\)J−1=0\\mathcal\{L\}^\{J^\{\-1\}\}\_\{\\mathrm\{SO\}\(3\)\}=0\) we have

\(J​\(Aθ0→t\)​dd​t​Aθ0→t\)∧=Ω​\(t,Rt,xt\)\.\\big\(J\(A\_\{\\theta\}^\{0\\to t\}\)\\tfrac\{d\}\{dt\}A\_\{\\theta\}^\{0\\to t\}\\big\)^\{\\wedge\}=\\Omega\(t,R\_\{t\},x\_\{t\}\)\.Hence

R~˙t=R~t​Ω​\(t,Rt,xt\)\.\\dot\{\\tilde\{R\}\}\_\{t\}=\\tilde\{R\}\_\{t\}\\Omega\(t,R\_\{t\},x\_\{t\}\)\.On the other hand the geodesic interpolation satisfies

R˙t=Rt​Ω​\(t,Rt,xt\),R0​given\.\\dot\{R\}\_\{t\}=R\_\{t\}\\Omega\(t,R\_\{t\},x\_\{t\}\),\\qquad R\_\{0\}\\ \\text\{given\}\.ThusRtR\_\{t\}andR~t\\tilde\{R\}\_\{t\}solve the same ODE with the same initial condition\. By uniqueness of solutions onSO​\(3\)\\mathrm\{SO\}\(3\)we obtain

R~t=Rt,∀t∈\[0,1\]\.\\tilde\{R\}\_\{t\}=R\_\{t\},\\qquad\\forall t\\in\[0,1\]\.Takingt=1t=1yields

R0​exp⁡\(\(Aθ0→1\)∧\)=R1⟺exp⁡\(\(Aθ0→1\)∧\)=R0⊤​R1\.R\_\{0\}\\exp\(\(A\_\{\\theta\}^\{0\\to 1\}\)^\{\\wedge\}\)=R\_\{1\}\\quad\\Longleftrightarrow\\quad\\exp\(\(A\_\{\\theta\}^\{0\\to 1\}\)^\{\\wedge\}\)=R\_\{0\}^\{\\top\}R\_\{1\}\.For generals<ts<tthe same argument gives

exp⁡\(\(Aθs→t\)∧\)=Rs⊤​Rt,\\exp\(\(A\_\{\\theta\}^\{s\\to t\}\)^\{\\wedge\}\)=R\_\{s\}^\{\\top\}R\_\{t\},which completes the proof\. ∎

To implement the loss in \([36](https://arxiv.org/html/2607.27431#A6.E36)\), it remains to compute the total derivativedd​t​Aθs→t\\frac\{d\}\{dt\}A\_\{\\theta\}^\{s\\to t\}\.

###### Proposition 6\.

From the product rule and chain rule, we have

dd​t​\(t−s\)​Aθ​\(s,t,Rt,xt\)\\displaystyle\\frac\{d\}\{dt\}\(t\-s\)A\_\{\\theta\}\(s,t,R\_\{t\},x\_\{t\}\)=Aθ​\(s,t,Rt,xt\)\+\(t−s\)​\(∂tAθ​\(s,t,Rt,xt\)\+⟨∇RAθ​\(s,t,Rt,xt\),R˙t⟩\+⟨∇xAθ​\(s,t,Rt,xt\),x˙t⟩\)\\displaystyle=A\_\{\\theta\}\(s,t,R\_\{t\},x\_\{t\}\)\+\(t\-s\)\\Big\(\\partial\_\{t\}A\_\{\\theta\}\(s,t,R\_\{t\},x\_\{t\}\)\+\\langle\\nabla\_\{R\}A\_\{\\theta\}\(s,t,R\_\{t\},x\_\{t\}\),\\dot\{R\}\_\{t\}\\rangle\+\\langle\\nabla\_\{x\}A\_\{\\theta\}\(s,t,R\_\{t\},x\_\{t\}\),\\dot\{x\}\_\{t\}\\rangle\\Big\)\(38\)

###### Proof\.

Fixs,ts,tand\(Rt,xt\)\(R\_\{t\},x\_\{t\}\)\. For convenience, we denote

Aθ:=Aθ​\(s,t,Rt,xt\)\.A\_\{\\theta\}:=A\_\{\\theta\}\(s,t,R\_\{t\},x\_\{t\}\)\.
By definition,\(t−s\)​Aθ\(t\-s\)A\_\{\\theta\}is a product of the scalar functiont−st\-sand the vector\-valued functionAθ​\(s,t,Rt,xt\)A\_\{\\theta\}\(s,t,R\_\{t\},x\_\{t\}\)\. Hence, by the product rule,

dd​t​\(t−s\)​Aθ\\displaystyle\\frac\{d\}\{dt\}\(t\-s\)A\_\{\\theta\}=dd​t​\(t−s\)​Aθ\+\(t−s\)​dd​t​Aθ\\displaystyle=\\frac\{d\}\{dt\}\(t\-s\)\\,A\_\{\\theta\}\+\(t\-s\)\\frac\{d\}\{dt\}A\_\{\\theta\}=Aθ\+\(t−s\)​dd​t​Aθ\.\\displaystyle=A\_\{\\theta\}\+\(t\-s\)\\frac\{d\}\{dt\}A\_\{\\theta\}\.\(39\)
Next, we compute the total derivative ofAθ​\(s,t,Rt,xt\)A\_\{\\theta\}\(s,t,R\_\{t\},x\_\{t\}\)with respect tott\. SinceAθA\_\{\\theta\}depends onttexplicitly and implicitly through bothRtR\_\{t\}andxtx\_\{t\}, by the chain rule,

dd​t​Aθ​\(s,t,Rt,xt\)\\displaystyle\\frac\{d\}\{dt\}A\_\{\\theta\}\(s,t,R\_\{t\},x\_\{t\}\)=∂tAθ​\(s,t,Rt,xt\)\+dR​Aθ​\(s,t,Rt,xt\)​\[R˙t\]\+dx​Aθ​\(s,t,Rt,xt\)​\[x˙t\]\.\\displaystyle=\\partial\_\{t\}A\_\{\\theta\}\(s,t,R\_\{t\},x\_\{t\}\)\+d\_\{R\}A\_\{\\theta\}\(s,t,R\_\{t\},x\_\{t\}\)\[\\dot\{R\}\_\{t\}\]\+d\_\{x\}A\_\{\\theta\}\(s,t,R\_\{t\},x\_\{t\}\)\[\\dot\{x\}\_\{t\}\]\.\(40\)
Substituting \([40](https://arxiv.org/html/2607.27431#A6.E40)\) into \([39](https://arxiv.org/html/2607.27431#A6.E39)\) yields

dd​t​\(t−s\)​Aθ​\(s,t,Rt,xt\)\\displaystyle\\frac\{d\}\{dt\}\(t\-s\)A\_\{\\theta\}\(s,t,R\_\{t\},x\_\{t\}\)=Aθ\+\(t−s\)​\(∂tAθ\+⟨∇RAθ,R˙t⟩\+⟨∇xAθ,x˙t⟩\)\.\\displaystyle=A\_\{\\theta\}\+\(t\-s\)\\Big\(\\partial\_\{t\}A\_\{\\theta\}\+\\langle\\nabla\_\{R\}A\_\{\\theta\},\\dot\{R\}\_\{t\}\\rangle\+\\langle\\nabla\_\{x\}A\_\{\\theta\},\\dot\{x\}\_\{t\}\\rangle\\Big\)\.
This proves \([38](https://arxiv.org/html/2607.27431#A6.E38)\)\. ∎

### F\.2Practical implementation:SE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}adaptation

In practice, each sample of the protein data is in shapeRN∈ℝN×3×3,xN∈ℝN×3R^\{N\}\\in\\mathbb\{R\}^\{N\\times 3\\times 3\},x^\{N\}\\in\\mathbb\{R\}^\{N\\times 3\}andNNcan be treated as the dimension of the data\. The model is defined as

fθ:\(RtN,xtN,t,s\)=\[Aθa​v​g,N;vθa​v​g,N\]∈ℝN×3×ℝN×3\.f\_\{\\theta\}:\(R^\{N\}\_\{t\},x^\{N\}\_\{t\},t,s\)=\[A\_\{\\theta\}^\{avg,N\};v\_\{\\theta\}^\{avg,N\}\]\\in\\mathbb\{R\}^\{N\\times 3\}\\times\\mathbb\{R\}^\{N\\times 3\}\.HereAθ,ia​v​g,N∈ℝ3A\_\{\\theta,i\}^\{avg,N\}\\in\\mathbb\{R\}^\{3\}parametrizes theii\-th𝔰​𝔬​\(3\)\\mathfrak\{so\}\(3\)component viaΩθ,i​\(s,t,Rt,i\)=\(Aθ,ia​v​g,N​\(s,t,Rt,i\)\)∧\\Omega\_\{\\theta,i\}\(s,t,R\_\{t,i\}\)=\(A\_\{\\theta,i\}^\{avg,N\}\(s,t,R\_\{t,i\}\)\)^\{\\wedge\}\.

##### Independent𝔰​𝔬​\(3\)N\\mathfrak\{so\}\(3\)^\{N\}loss\.

Extending \([36](https://arxiv.org/html/2607.27431#A6.E36)\) to𝔰​𝔬​\(3\)N\\mathfrak\{so\}\(3\)^\{N\}by treating each component independently gives

ℒ𝔰​𝔬​\(3\)N\\displaystyle\\mathcal\{L\}\_\{\\mathfrak\{so\}\(3\)^\{N\}\}:=𝔼s<t,q​\(R0N,R1N\)​∑i=1N‖sg​\(J​\(\(t−s\)​Aθ,iavg\)\)​\(Aθ,iavg\+\(t−s\)​sg​\(dd​t​A^θ,iavg\)\)−ωt,i‖22,\\displaystyle:=\\mathbb\{E\}\_\{\\begin\{subarray\}\{c\}s<t,\\\\ q\(R^\{N\}\_\{0\},R^\{N\}\_\{1\}\)\\end\{subarray\}\}\\sum\_\{i=1\}^\{N\}\\left\\\|\\mathrm\{sg\}\\Big\(J\\big\(\(t\-s\)A\_\{\\theta,i\}^\{\\mathrm\{avg\}\}\\big\)\\Big\)\\Big\(A\_\{\\theta,i\}^\{\\mathrm\{avg\}\}\+\(t\-s\)\\,\\mathrm\{sg\}\\big\(\\tfrac\{d\}\{dt\}\\hat\{A\}\_\{\\theta,i\}^\{\\mathrm\{avg\}\}\\big\)\\Big\)\-\\omega\_\{t,i\}\\right\\\|\_\{2\}^\{2\},\(41\)whereωt,i:=Ωi​\(t,Rt,i,xt,i\)∨\\omega\_\{t,i\}:=\\Omega\_\{i\}\(t,R\_\{t,i\},x\_\{t,i\}\)^\{\\vee\}is the data instantaneous angular velocity \(cf\. Eq\. \([7](https://arxiv.org/html/2607.27431#S3.E7)\)\) andJ​\(⋅\)J\(\\cdot\)is theSO​\(3\)\\mathrm\{SO\}\(3\)right Jacobian applied per component\. TheJ−1J^\{\-1\}form \([37](https://arxiv.org/html/2607.27431#A6.E37)\) extends in the same way, with theii\-th summand replaced by‖Aθ,iavg−sg​\(J−1​\(\(t−s\)​Aθ,iavg\)​ωt,i−\(t−s\)​dd​t​A^θ,iavg\)‖22\\big\\\|A\_\{\\theta,i\}^\{\\mathrm\{avg\}\}\-\\mathrm\{sg\}\\big\(J^\{\-1\}\\big\(\(t\-s\)A\_\{\\theta,i\}^\{\\mathrm\{avg\}\}\\big\)\\omega\_\{t,i\}\-\(t\-s\)\\tfrac\{d\}\{dt\}\\hat\{A\}\_\{\\theta,i\}^\{\\mathrm\{avg\}\}\\big\)\\big\\\|\_\{2\}^\{2\}\.

##### Translationℝ3×N\\mathbb\{R\}^\{3\\times N\}loss\.

For the translation component, the loss in Euclidean space \([3](https://arxiv.org/html/2607.27431#S2.E3)\) extends directly toNNindependent particles\. With the model predictionvθ,iavg∈ℝ3v\_\{\\theta,i\}^\{\\text\{avg\}\}\\in\\mathbb\{R\}^\{3\}for residueii, the translation loss is

ℒℝ3×N\\displaystyle\\mathcal\{L\}\_\{\\mathbb\{R\}^\{3\\times N\}\}:=𝔼s<t,q​\(x0N,x1N\)​∑i=1N‖vθ,iavg\+\(t−s\)​sg​\(dd​t​v^θ,iavg\)−vi​\(t,xt,i\)‖22,\\displaystyle:=\\mathbb\{E\}\_\{\\begin\{subarray\}\{c\}s<t,\\\\ q\(x^\{N\}\_\{0\},x^\{N\}\_\{1\}\)\\end\{subarray\}\}\\sum\_\{i=1\}^\{N\}\\left\\\|v\_\{\\theta,i\}^\{\\text\{avg\}\}\+\(t\-s\)\\,\\text\{sg\}\\\!\\left\(\\frac\{d\}\{dt\}\\hat\{v\}\_\{\\theta,i\}^\{\\text\{avg\}\}\\right\)\-v\_\{i\}\(t,x\_\{t,i\}\)\\right\\\|\_\{2\}^\{2\},\(42\)wherevi​\(t,xt,i\)=x1,i−x0,iv\_\{i\}\(t,x\_\{t,i\}\)=x\_\{1,i\}\-x\_\{0,i\}is the constant ground\-truth translation velocity along the linear interpolation pathxt,i=\(1−t\)​x0,i\+t​x1,ix\_\{t,i\}=\(1\-t\)x\_\{0,i\}\+tx\_\{1,i\}\. The combined MeanFlow loss onSE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}is thereforeℒM​FSE​\(3\)N=ℒ𝔰​𝔬​\(3\)N\+ℒℝ3×N\\mathcal\{L\}^\{\\mathrm\{SE\}\(3\)^\{N\}\}\_\{MF\}=\\mathcal\{L\}\_\{\\mathfrak\{so\}\(3\)^\{N\}\}\+\\mathcal\{L\}\_\{\\mathbb\{R\}^\{3\\times N\}\}\.

## Appendix GCommon Settings: Auxiliary Loss, Prior and Inference\.

### G\.1Auxiliary Loss

We include auxiliary losses fromYimet al\.\[[2023b](https://arxiv.org/html/2607.27431#bib.bib8)\]to enforce geometric consistency at the atomic level\. Specifically, let𝒜0∈ℝN×4×3\\mathcal\{A\}\_\{0\}\\in\\mathbb\{R\}^\{N\\times 4\\times 3\}denote the ground\-truth backbone atom coordinates \(in Å\) of the four heavy atoms\(N,Cα,C,O\)\(\\text\{N\},\\text\{C\}\_\{\\alpha\},\\text\{C\},\\text\{O\}\)per residue, and let𝒜^0\\hat\{\\mathcal\{A\}\}\_\{0\}be the corresponding coordinates predicted by the model\. We define a direct regression loss on backbone atom positions and a pairwise distance loss in a local neighbourhood,

ℒbb=14​N​∑‖𝒜0−𝒜^0‖2,ℒ2​D=‖𝟏​\{D<6​Å\}​\(D−D^\)‖2∑𝟏D<6​Å−N,\\mathcal\{L\}\_\{\\mathrm\{bb\}\}=\\frac\{1\}\{4N\}\\sum\\\|\\mathcal\{A\}\_\{0\}\-\\hat\{\\mathcal\{A\}\}\_\{0\}\\\|^\{2\},\\qquad\\mathcal\{L\}\_\{\\mathrm\{2D\}\}=\\frac\{\\\|\\mathbf\{1\}\\\{D<6\\text\{\\AA \}\\\}\(D\-\\hat\{D\}\)\\\|^\{2\}\}\{\\sum\\mathbf\{1\}\_\{D<6\\text\{\\AA \}\}\-N\},\(43\)whereD∈ℝN×N×4×4D\\in\\mathbb\{R\}^\{N\\times N\\times 4\\times 4\}is the tensor of pairwise distances between heavy atoms, i\.e\.Di​j​a​b=‖𝒜i​a−𝒜j​b‖D\_\{ijab\}=\\\|\\mathcal\{A\}\_\{ia\}\-\\mathcal\{A\}\_\{jb\}\\\|, andD^\\hat\{D\}is defined analogously from𝒜^0\\hat\{\\mathcal\{A\}\}\_\{0\}\. The auxiliary loss is then

ℒaux=𝔼𝒬​\[ℒbb\+ℒ2​D\],\\mathcal\{L\}\_\{\\mathrm\{aux\}\}=\\mathbb\{E\}\_\{\\mathcal\{Q\}\}\\\!\\left\[\\mathcal\{L\}\_\{\\mathrm\{bb\}\}\+\\mathcal\{L\}\_\{\\mathrm\{2D\}\}\\right\],\(44\)where𝒬​\(t,x0,x1,x~t\):=𝒰​\(0,1\)⊗π¯​\(x0,x1\)⊗ρt​\(x~t∣x0,x1\)\\mathcal\{Q\}\(t,x\_\{0\},x\_\{1\},\\tilde\{x\}\_\{t\}\):=\\mathcal\{U\}\(0,1\)\\otimes\\bar\{\\pi\}\(x\_\{0\},x\_\{1\}\)\\otimes\\rho\_\{t\}\(\\tilde\{x\}\_\{t\}\\mid x\_\{0\},x\_\{1\}\)is the factorized joint distribution\. And following the settings inBoseet al\.\[[2024](https://arxiv.org/html/2607.27431#bib.bib9)\], Yimet al\.\[[2023b](https://arxiv.org/html/2607.27431#bib.bib8)\], we only applyℒaux\\mathcal\{L\}\_\{\\mathrm\{aux\}\}fort<0\.75t<0\.75, scaling it byλaux\\lambda\_\{\\mathrm\{aux\}\}\.

### G\.2Prior Distribution

Following the notation inGenget al\.\[[2026a](https://arxiv.org/html/2607.27431#bib.bib35),[b](https://arxiv.org/html/2607.27431#bib.bib27)\], we denote protein backbones by\(R0N,x0N\)\(R\_\{0\}^\{N\},x\_\{0\}^\{N\}\)and sample\(R1N,x1N\)∈SO​\(3\)N×ℝN×3\(R\_\{1\}^\{N\},x\_\{1\}^\{N\}\)\\in\\mathrm\{SO\}\(3\)^\{N\}\\times\\mathbb\{R\}^\{N\\times 3\}from a simple prior\.

##### Translation prior\.

In Euclidean space, the natural analogue of a “standard” prior is an isotropic Gaussian,

x1N∼𝒩​\(0,σx2​IN×3\),x\_\{1\}^\{N\}\\sim\\mathcal\{N\}\\bigl\(0,\\sigma\_\{x\}^\{2\}I\_\{N\\times 3\}\\bigr\),\(45\)where we identifyx1N∈ℝN×3x\_\{1\}^\{N\}\\in\\mathbb\{R\}^\{N\\times 3\}with a vector inℝ3​N\\mathbb\{R\}^\{3N\}and typically setσx=1\\sigma\_\{x\}=1\.

##### Rotation prior: Gaussian analogue onSO​\(3\)\\mathrm\{SO\}\(3\)\.

On the rotation group, the closest analogue of an isotropic Gaussian is an*isotropic \(heat\-kernel\) Gaussian*onSO​\(3\)\\mathrm\{SO\}\(3\), denotedIGSO​\(3\)​\(σR\)\\mathrm\{IGSO\(3\)\}\(\\sigma\_\{R\}\)\. It is the transition density of Brownian motion onSO​\(3\)\\mathrm\{SO\}\(3\)at timeσR2\\sigma\_\{R\}^\{2\}, and is isotropic in the sense that its density depends only on the geodesic rotation angle

θ\(R\):=∥Log\(R\)∥2∈\[0,π\],Log:SO\(3\)→𝔰𝔬\(3\)≅ℝ3,\\theta\(R\):=\\\|\\text\{Log\}\(R\)\\\|\_\{2\}\\in\[0,\\pi\],\\qquad\\text\{Log\}:\\mathrm\{SO\}\(3\)\\to\\mathfrak\{so\}\(3\)\\cong\\mathbb\{R\}^\{3\},\(46\)with respect to the Haar measured​R\\mathrm\{d\}R\.

LetI​G​S​O​\(3\)NIGSO\(3\)^\{N\}denote the product measure ofI​G​S​O​\(3\)IGSO\(3\)\. Concretely, we use the factorized prior

p​\(R1N,x1N\)=I​G​S​O​\(3\)​\(σR\)N×𝒩​\(0,IN×3\)\.p\(R\_\{1\}^\{N\},x\_\{1\}^\{N\}\)=IGSO\(3\)\(\\sigma\_\{R\}\)^\{N\}\\times\\mathcal\{N\}\(0,I\_\{N\\times 3\}\)\.\(47\)The typical choice ofσR\\sigma\_\{R\}is1\.51\.5\. In addition, asσR→∞\\sigma\_\{R\}\\to\\infty,IGSO​\(3\)​\(σR\)\\mathrm\{IGSO\(3\)\}\(\\sigma\_\{R\}\)approaches the uniform \(Haar\) distribution onSO​\(3\)\\mathrm\{SO\}\(3\), which is another typical choice of the prior\.

Algorithm 1MeanFlow training forSE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}\(idealized form\)\. This is the plain MeanFlow objective, stated to make the structure of the method explicit; it is*not*the recipe used for our reported results\. The trained objective is the Stage\-2 loss \([97](https://arxiv.org/html/2607.27431#A13.E97)\)—which adds the endpoint anchor and the auxiliary term—with the rotation and translation losses computed in the small\-ttstabilized form of Algorithm[3](https://arxiv.org/html/2607.27431#alg3), after theα\\alpha\-Flow warm\-up of Algorithm[5](https://arxiv.org/html/2607.27431#alg5)\. See Appendix[M](https://arxiv.org/html/2607.27431#A13)\.1:Data distribution

pdatap\_\{\\text\{data\}\}over

SE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}; MeanFlow network

fθf\_\{\\theta\}; OT\-coupling flag

2:foreach training iterationdo

3:Sample a mini\-batch

\{\(R0N,x0N\)\}b=1B∼pdata\\\{\(R\_\{0\}^\{N\},x\_\{0\}^\{N\}\)\\\}\_\{b=1\}^\{B\}\\sim p\_\{\\text\{data\}\}\.

4:

R0N∈ℝB×N×3×3R\_\{0\}^\{N\}\\in\\mathbb\{R\}^\{B\\times N\\times 3\\times 3\},

x0N∈ℝB×N×3x\_\{0\}^\{N\}\\in\\mathbb\{R\}^\{B\\times N\\times 3\}\.

5:Sample

\(R1N,x1N\)\(R\_\{1\}^\{N\},x\_\{1\}^\{N\}\)from the prior; see Section[G\.2](https://arxiv.org/html/2607.27431#A7.SS2)\.

6:ifOT\-couplingthen

7:Solve

O​T​\(\(R0,x0\),\(R1,x1\)\)OT\(\(R\_\{0\},x\_\{0\}\),\(R\_\{1\},x\_\{1\}\)\)and couple

\(R0,x0\)\(R\_\{0\},x\_\{0\}\)and

\(R1,x1\)\(R\_\{1\},x\_\{1\}\)using the optimal transport plan\.

8:endif

9:Sample times

0≤s<t≤10\\leq s<t\\leq 1\.

10:Forward\-interpolate to time

ttto obtain

\(RtN,ΩtN\)\(R\_\{t\}^\{N\},\\Omega\_\{t\}^\{N\}\)and

\(xtN,vtN\)\(x\_\{t\}^\{N\},v\_\{t\}^\{N\}\)\.

11:Evaluate the network:

\[Aθa​v​g,N,vθa​v​g,N\]←fθ​\(s,t,RtN,xtN\)\[A\_\{\\theta\}^\{avg,N\},v\_\{\\theta\}^\{avg,N\}\]\\leftarrow f\_\{\\theta\}\(s,t,R\_\{t\}^\{N\},x\_\{t\}^\{N\}\)\.

12:Compute

ℒ𝔰​𝔬​\(3\)N\\mathcal\{L\}\_\{\\mathfrak\{so\}\(3\)^\{N\}\}via Eq\. \([41](https://arxiv.org/html/2607.27431#A6.E41)\)\.

13:Compute

ℒℝ3×N\\mathcal\{L\}\_\{\\mathbb\{R\}^\{3\\times N\}\}via Eq\. \([42](https://arxiv.org/html/2607.27431#A6.E42)\)\.

14:Compute the auxiliary loss

ℒaux\\mathcal\{L\}\_\{\\text\{aux\}\}; see Section[G\.1](https://arxiv.org/html/2607.27431#A7.SS1)\.

15:

ℒ←ℒ𝔰​𝔬​\(3\)N\+ℒℝ3×N\+λaux​1​\[t<0\.75\]​ℒaux\\mathcal\{L\}\\leftarrow\\mathcal\{L\}\_\{\\mathfrak\{so\}\(3\)^\{N\}\}\+\\mathcal\{L\}\_\{\\mathbb\{R\}^\{3\\times N\}\}\+\\lambda\_\{\\text\{aux\}\}\\,\\mathbb\{1\}\[t<0\.75\]\\,\\mathcal\{L\}\_\{\\text\{aux\}\}\.

16:Update

θ\\thetausing a gradient step \(or any optimizer\)\.

17:endfor

### G\.3Inference process

Similar to the original MeanFlow method in Euclidean space, we can adapt multi\-step inference for theSE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}MeanFlow framework\.

In particular, letT∈ℕT\\in\\mathbb\{N\}denote the number of inference steps and letΔ​t=1T\\Delta t=\\frac\{1\}\{T\}be the step size\. We apply the following updates onSO​\(3\)\\mathrm\{SO\}\(3\)andℝ3\\mathbb\{R\}^\{3\}, respectively:

Rt−1=Rt​e−Δ​t​Aθ∧​\(s,t,Rt,xt\),\\displaystyle R\_\{t\-1\}=R\_\{t\}e^\{\-\\Delta tA\_\{\\theta\}^\{\\wedge\}\(s,t,R\_\{t\},x\_\{t\}\)\},xt−1=xt−Δ​t​vθavg​\(s,t,Rt,xt\)\.\\displaystyle x\_\{t\-1\}=x\_\{t\}\-\\Delta tv^\{\\text\{avg\}\}\_\{\\theta\}\(s,t,R\_\{t\},x\_\{t\}\)\.wherettdecreases from11to1/T1/Tands=t−Δ​ts=t\-\\Delta t\. We summarize the procedure in Algorithm[2](https://arxiv.org/html/2607.27431#alg2)\.

Algorithm 2SE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}MeanFlow inference1:Batch size

BB; pretrained network

fθf\_\{\\theta\}; number of steps

TT
2:

\(R^0N,x^0N\)\(\\hat\{R\}\_\{0\}^\{N\},\\hat\{x\}\_\{0\}^\{N\}\)
3:

Δ​t=1/T\\Delta t=1/T
4:Sample

\(R1N,x1N\)\(R\_\{1\}^\{N\},x\_\{1\}^\{N\}\)from the prior over

SE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}\.

5:

\(RtN,xtN\)←\(R1N,x1N\)\(R\_\{t\}^\{N\},x\_\{t\}^\{N\}\)\\leftarrow\(R\_\{1\}^\{N\},x\_\{1\}^\{N\}\)\.

6:for

t=1,1−Δ​t,…,Δ​tt=1,1\-\\Delta t,\\ldots,\\Delta tdo

7:

s←t−Δ​ts\\leftarrow t\-\\Delta t\.

8:

\[Aθa​v​g,N,vθa​v​g,N\]←fθ​\(s,t,RtN,xtN\)\[A\_\{\\theta\}^\{avg,N\},v\_\{\\theta\}^\{avg,N\}\]\\leftarrow f\_\{\\theta\}\(s,t,R\_\{t\}^\{N\},x\_\{t\}^\{N\}\)\.

9:

RtN←RtN​exp⁡\(−Δ​t​\(Aθa​v​g,N\)∧\)R\_\{t\}^\{N\}\\leftarrow R\_\{t\}^\{N\}\\exp\(\-\\Delta t\\,\(A\_\{\\theta\}^\{avg,N\}\)^\{\\wedge\}\)\.

10:

xtN←xtN−Δ​t​vθa​v​g,Nx\_\{t\}^\{N\}\\leftarrow x\_\{t\}^\{N\}\-\\Delta t\\,v\_\{\\theta\}^\{avg,N\}\.

11:endfor

12:

R^0N=RtN,x^0N=xtN\\hat\{R\}^\{N\}\_\{0\}=R\_\{t\}^\{N\},\\hat\{x\}^\{N\}\_\{0\}=x\_\{t\}^\{N\}\.

## Appendix HAblation Study: A ControlledSO​\(3\)\\mathrm\{SO\}\(3\)Benchmark for Few\-Step Generation

Designability on SCOPe confounds the generative objective with the protein\-specific IPA trunk, the auxiliary losses, the self\-conditioning recipe, and the folding oracle, so a few\-step win there is hard to attribute\. This appendix therefore tests the central claim of the paper — that it is the*average\-velocity parameterisation*that enables few\-step generation, rather than the network, coupling, or protein\-specific machinery — on a controlledSO​\(3\)2\\mathrm\{SO\}\(3\)^\{2\}benchmark, where we hold all confounders fixed and vary only the training loss\.

A full ablation on real proteins is prohibitively expensive: the design space includes many hyperparameter combinations \(e\.g\., OT variants in sampling such as per\-GPU OT, global OT, or no OT; loss normalizations and clipping; different loss mixtures; architectural toggles such as enabling/disabling AdaLN; and optimizer / learning\-rate choices\)\. Each setting requires substantial compute \(4×\\times80GB GPUs\), long GPU\-hour budgets, and careful checkpoint selection to reliably judge performance\. Under our limited training budget we therefore do*not*run exhaustive real\-dataset ablations; instead, on the syntheticSO​\(3\)2\\mathrm\{SO\}\(3\)^\{2\}benchmark we fix the model, learning rate,SE​\(3\)\\mathrm\{SE\}\(3\)interpolator, and training steps so that performance differences are attributable to the losses themselves\.

### H\.1Data

A sample is a pair of rotations\(R\(0\),R\(1\)\)∈SO​\(3\)2\(R^\{\(0\)\},R^\{\(1\)\}\)\\in\\mathrm\{SO\}\(3\)^\{2\}: the rotational part of a two\-residue backbone, with the translation branch switched off\. Throughout we visualise a rotation at its axis\-angle coordinate

w=Log​\(R\)∈ℝ3,‖w‖≤π,\\displaystyle w=\\mathrm\{Log\}\(R\)\\in\\mathbb\{R\}^\{3\},\\qquad\\\|w\\\|\\leq\\pi,\(48\)so a distribution onSO​\(3\)\\mathrm\{SO\}\(3\)becomes a point cloud inside theπ\\pi\-ball \(Figure[6](https://arxiv.org/html/2607.27431#A8.F6)\)\.

The two atoms carry deliberately different geometry\.

##### Atom 0 \(cube6\): a mode\-seeking task\.

A mixture of six isotropic components centred at the six cube\-face rotations \(rotations byπ/2\\pi/2about±ex,±ey,±ez\\pm e\_\{x\},\\pm e\_\{y\},\\pm e\_\{z\}, hence‖w‖=π/2\\\|w\\\|=\\pi/2\)\. The modes are sharp and well separated, so a sampler must*commit*to one basin; hedging between two modes is immediately visible as mass in the empty region between them\.

##### Atom 1 \(moons3d\): a curved\-manifold task\.

The classical two\-moons distribution, lifted into the Lie algebra so that its support is a pair of interlocking crescents\. Here the model must trace a thin, curved, one\-dimensional structure rather than collapse onto a few points\.

![Refer to caption](https://arxiv.org/html/2607.27431v1/x7.png)Figure 6:TheSO​\(3\)2\\mathrm\{SO\}\(3\)^\{2\}target, drawn in axis\-angle coordinatesw=Log​\(R\)w=\\mathrm\{Log\}\(R\),‖w‖≤π\\\|w\\\|\\leq\\pi\.Left, atom 0 \(cube6\): six sharp modes at the cube\-face rotations — a mode\-seeking task\.Right, atom 1 \(moons3d\): two interlocking crescents — a curved\-manifold task\. Both must be solved by the same network under the same objective\.![Refer to caption](https://arxiv.org/html/2607.27431v1/x8.png)Figure 7:Two\-step \(T=2T\{=\}2\) samples, in the same axis\-angle coordinates as Figure[6](https://arxiv.org/html/2607.27431#A8.F6), from the*identical*frozen prior noise for every method\. Top row: atom 0 \(cube6\); bottom row: atom 1 \(moons3d\)\. The average\-velocity methods \(ours and both RMF variants\) already resolve the six modes and the crescent geometry after two steps\. FoldFlow and QFlow smear across the ball, and V\-QFlow contracts onto a shell\.
##### Prior\.

The prior is the Haar \(uniform\) measure onSO​\(3\)\\mathrm\{SO\}\(3\), independently per atom\. A Haar draw sits a mean geodesic distance of126\.5∘126\.5^\{\\circ\}from the identity, so the transport distance is large and the prior carries no information about either target\. We draw10,00010\{,\}000training and2,0002\{,\}000held\-out samples; the noise is resampled every epoch, from a seeded stream \(below\)\.

### H\.2Experiment setting

All methods share the following\. Only the training loss differs\.

1. 1\.Network\.The same1\.071\.07M\-parameterSE​\(3\)\\mathrm\{SE\}\(3\)MLP for every method, mapping\(Rt,t,s\)↦ω∈ℝ2×3\(R\_\{t\},t,s\)\\mapsto\\omega\\in\\mathbb\{R\}^\{2\\times 3\}, a body\-frame angular velocity per atom\. Single\-time methods \(FoldFlow\) are feds=ts\{=\}t; endpoint methods \(QFlow, V\-QFlow\) use the quaternion head of the same trunk, at matched rotation\-branch capacity\.
2. 2\.Time sampler\.t∼𝒰​\(0,1\)t\\sim\\mathcal\{U\}\(0,1\)with*no*lower cut\-off \(tε=0t\_\{\\varepsilon\}=0; the regular\-ttlosses never divide bytt\), ands∼𝒰​\(0,t\)s\\sim\\mathcal\{U\}\(0,t\)for the two\-time methods, with a50%50\\%s=ts\{=\}tanchor branch\. The three\-time method \(RMF semigroup\) additionally draws an interiorm∈\(s,t\)m\\in\(s,t\)and gets nos=ts\{=\}tbranch, since a degenerate interval makes the semigroup identity vacuous\.
3. 3\.Budget\.2020k steps, batch500500, the same optimiser and learning rate, and EMA weights at evaluation\.
4. 4\.Data and noise stream\.*Bit\-identical*across methods\. This is enforced by three independent RNG streams — batch indices, coupling, and times — so that methods which consume different numbers of random draws \(OT draws a multinomial; a three\-time method draws three uniforms where a single\-time method draws one\) nevertheless see the same\(R0,R1\)\(R\_\{0\},R\_\{1\}\)pair at every step\. With a single shared stream the sequences desynchronise after the first step, silently, and the comparison is no longer controlled\.

##### No optimal\-transport coupling\.

No method uses mini\-batch OT\. This is deliberate and it is the crux of the protocol: OT is a flow\-*straightening*device\. Under an OT coupling the learned probability path is close to a straight geodesic, in which case the average velocity and the instantaneous velocity nearly coincide — so a plain flow\-matching model gets few\-step generation for free, and the experiment can no longer distinguish the two parameterisations\. Enabling OT would hand the instantaneous\-velocity baselines exactly the property under test\. We therefore run every method*without*OT, so that few\-step behaviour is attributable to the objective alone\.

### H\.3Evaluation Metric

We report the Wasserstein\-2 distance between2,0002\{,\}000generated and2,0002\{,\}000held\-out samples, under the squared geodesic cost summed over the two atoms:

c​\(\(R\(0\),R\(1\)\),\(S\(0\),S\(1\)\)\)=∑a∈\{0,1\}dSO​\(3\)2​\(R\(a\),S\(a\)\),dSO​\(3\)​\(R,S\)=‖Log​\(R⊤​S\)‖,\\displaystyle c\\big\(\(R^\{\(0\)\},R^\{\(1\)\}\),\(S^\{\(0\)\},S^\{\(1\)\}\)\\big\)=\\sum\_\{a\\in\\\{0,1\\\}\}d\_\{\\mathrm\{SO\}\(3\)\}^\{2\}\\big\(R^\{\(a\)\},S^\{\(a\)\}\\big\),\\qquad d\_\{\\mathrm\{SO\}\(3\)\}\(R,S\)=\\big\\\|\\mathrm\{Log\}\(R^\{\\top\}S\)\\big\\\|,\(49\)solved as an exact assignment and reported in degrees\.

##### The floor\.

At finite sample size,𝒲2\\mathcal\{W\}\_\{2\}between two*independent*draws of the target is not zero\. We measure this floor with the identical estimator and the identical sample size, and report it in Table[5](https://arxiv.org/html/2607.27431#A8.T5):21\.70∘21\.70^\{\\circ\}\. No model can go below it, and a model at the floor is indistinguishable from the target at this sample size\.

##### Sampling\.

Models are integrated on a uniform Euler gridt∈linspace​\(1,0,T\)t\\in\\mathrm\{linspace\}\(1,0,T\)forT∈\{1,2,5,10,20\}T\\in\\\{1,2,5,10,20\\\}, with no final\-step special case and notmint\_\{\\min\}floor\. Every method is initialised from the same frozen prior\-noise batch, so the sample sets — and hence the panels of Figure[7](https://arxiv.org/html/2607.27431#A8.F7)— are directly comparable\.

Table 5:Few\-step generation on theSO​\(3\)2\\mathrm\{SO\}\(3\)^\{2\}toy benchmark\.𝒲2\\mathcal\{W\}\_\{2\}\(degrees; lower is better\) between2,0002\{,\}000generated and2,0002\{,\}000held\-out samples, as the number of Euler stepsTTfalls from2020to11\. All methods share one network, one data/noise stream, one time sampler and one budget \(Section[H\.2](https://arxiv.org/html/2607.27431#A8.SS2)\);*none*uses an OT coupling\.Flooris the𝒲2\\mathcal\{W\}\_\{2\}between two independent draws of the target and is the lowest attainable value\. Best per column in bold\. The separation is by*parameterisation*, not by method family: average\-velocity methods surviveT=1T\{=\}1; instantaneous\-velocity and endpoint\-prediction methods collapse\. Removing theSO​\(3\)\\mathrm\{SO\}\(3\)Jacobian \(J−1:=IJ^\{\-1\}\\\!:=\\\!I\) leaves the average\-velocity model unchanged within noise forT≥2T\\geq 2and costs0\.4∘0\.4^\{\\circ\}atT=1T\{=\}1, where the sampler queries the largest interval\.

### H\.4Results

Table[5](https://arxiv.org/html/2607.27431#A8.T5)shows a clean separation by*parameterisation*, not by method family or by method quality\.

##### 20 steps: the fine\-grid regime\.

AtT=20T\{=\}20the two parameterisations are indistinguishable in quality: every average\-velocity method, together with FoldFlow, sits within a few degrees of the21\.70∘21\.70^\{\\circ\}floor, and the single best number is in fact FoldFlow’s \(24\.77∘24\.77^\{\\circ\}\)\. With a fine enough Euler grid the instantaneous\-velocity parameterisation is entirely adequate and our objective buys nothing\. The exceptions are the quaternion\-parameterised methods, QFlow and V\-QFlow, which plateau roughly10∘10^\{\\circ\}above the floor\. Since QFlow optimises a flow\-matching loss and V\-QFlow an endpoint loss, both should in principle track FoldFlow in this regime; we attribute the gap to the rotation\-matrix backbone used throughout this ablation, which forces repeated quaternion–matrix conversions inside the model and the loss\. Appendix[L\.4](https://arxiv.org/html/2607.27431#A12.SS4)reports the same effect at protein scale\.

##### Few steps: MeanFlow methods outperform flow matching\.

AtT=2T\{=\}2\(andT=5T\{=\}5\), the MeanFlow\-style average\-velocity methods — ours and both RMF variants — already approach the21\.70∘21\.70^\{\\circ\}floor, with our model being slightly better overall\. In contrast, the flow\-matching baselines FoldFlow, QFlow and V\-QFlow exhibit substantially higher errors at these low step counts\. Figure[7](https://arxiv.org/html/2607.27431#A8.F7)corroborates this qualitatively: after two steps our samples already show six distinct modes and a recognisable pair of crescents, whereas FoldFlow and QFlow do not\.

AtT=1T\{=\}1, SE\(3\)\-MeanFlow remains the best method, but all approaches are relatively far from the floor\. This is expected: in the extreme case where two rotations differ by an angle ofπ\\pi, the shortest geodesic \(and hence the induced velocity\) is not well\-defined due to non\-uniqueness \(analogous to the ambiguity of shortest paths between the north and south poles on a sphere\)\. This suggests thatSO​\(3\)\\mathrm\{SO\}\(3\)is intrinsically ill\-suited for one\-step generation\. However, few\-step generation methods can still be applied\.

##### Effect of the right Jacobian\.

ReplacingJ−1J^\{\-1\}by the identity — the log\-map approximation ofZhonget al\.\[[2026](https://arxiv.org/html/2607.27431#bib.bib15)\]— costs0\.41∘0\.41^\{\\circ\}atT=1T\{=\}1and is within0\.13∘0\.13^\{\\circ\}of our model elsewhere \(Table[5](https://arxiv.org/html/2607.27431#A8.T5)\)\. This is the expected pattern: sinceJ−1​\(A\)=I\+12​A∧\+⋯J^\{\-1\}\(A\)=I\+\\tfrac\{1\}\{2\}A^\{\\wedge\}\+\\cdotsacts onAs→t=\(t−s\)​AavgA^\{s\\to t\}=\(t\-s\)A^\{\\mathrm\{avg\}\}, the correction scales with the interval, and aTT\-step sampler queries onlyt−s=1/Tt\-s=1/T\. The same scaling holds on real training data \(Table[9](https://arxiv.org/html/2607.27431#A15.T9)\)\. The Jacobian is closed\-form and free, so we retain it; its benefit is concentrated in the aggressive few\-step regime\.

##### Theα\\alpha\-Flow ablation isolates the mechanism\.

Ourα\\alpha\-Flow objective \(Section[J](https://arxiv.org/html/2607.27431#A10)\), trained alone withα\\alphaannealed1→0\.11\\to 0\.1, performs competitively, but is consistently slightly worse than SE\(3\)\-MeanFlow at low step counts \(Table[5](https://arxiv.org/html/2607.27431#A8.T5)\)\. This is expected:α\\alpha\-Flow is a JVP\-free surrogate that relaxes the full average\-velocity consistency enforced by MeanFlow\. Appending a MeanFlow phase to the last40%40\\%of the same schedule — same network, same budget, same data — improves the one\-step error by about1∘1^\{\\circ\}, with smaller gains atT=2T\{=\}2, indicating that few\-step capability is primarily supplied by the MeanFlow objective\. Training with MeanFlow throughout gives the better one\- and two\-step error, while the warm\-up variant is marginally better forT≥5T\\geq 5; neither ordering is large relative to the within\-block spread\. The few\-step advantage reported in the main text is therefore not an artefact of a short fine\-tuning phase\.

##### Concurrent work, Riemannian Meanflow\.

Both RMF variantsWooet al\.\[[2026](https://arxiv.org/html/2607.27431#bib.bib23)\]sit in the same block as our method and perform comparably \(RMF semigroup\+\+FM reaches31\.52∘31\.52^\{\\circ\}atT=1T\{=\}1, statistically indistinguishable from our30\.74∘30\.74^\{\\circ\}at this sample size\)\. We regard this as supporting the paper’s thesis rather than undercutting it: the thesis is about the average\-velocity parameterisation, and RMF is an average\-velocity method\. TheSO​\(3\)\\mathrm\{SO\}\(3\)\-specific contribution of our work — the exact right\-Jacobian identity of Proposition[1](https://arxiv.org/html/2607.27431#Thmproposition1)and the resultingSE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}objective — is evaluated on protein backbones in the main text, where the two methods do separate\.

##### Caveat\.

The numbers in Table[5](https://arxiv.org/html/2607.27431#A8.T5)are from a single seed\. The∼50∘\\sim\\\!50^\{\\circ\}gap between the two blocks is far beyond any plausible seed variation, but small differences*within*a block \(e\.g\. our30\.74∘30\.74^\{\\circ\}versus RMF’s31\.52∘31\.52^\{\\circ\}atT=1T\{=\}1\) should not be read as a ranking\.

## Appendix IStable Training Framework, Part 1: MeanFlow Loss \(Vanilla and Small\-ttStabilized\)

##### Vanilla \(no small\-ttreparameterisation\)\.

The MeanFlow identity onSO​\(3\)\\mathrm\{SO\}\(3\)\(Proposition[1](https://arxiv.org/html/2607.27431#Thmproposition1), with full derivation in Appendix[F\.1](https://arxiv.org/html/2607.27431#A6.SS1)\) yields a consistency target involving the trajectory derivativedd​t​Aθavg\\tfrac\{d\}\{dt\}A^\{\\mathrm\{avg\}\}\_\{\\theta\}\. Our practical deep learning model is*diff\-frame*: it takes the current state\(Rt,xt\)\(R\_\{t\},x\_\{t\}\)and a time embedding \(constructed from\(t,s\)\(t,s\)\) as input, and directly predicts the endpoints\(R^0θ,x^0θ\)\(\\hat\{R\}\_\{0\}^\{\\theta\},\\hat\{x\}\_\{0\}^\{\\theta\}\)\. We then define the average velocities using the endpoint \(“xx\-prediction”\) parameterisation

Aθavg​\(Rt,xt,t,s\)\\displaystyle A\_\{\\theta\}^\{\\mathrm\{avg\}\}\(R\_\{t\},x\_\{t\},t,s\)=1tlog\(\(R^0θ\)⊤Rt\)∨,\\displaystyle=\\frac\{1\}\{t\}\\,\\log\\\!\\big\(\(\\hat\{R\}\_\{0\}^\{\\theta\}\)^\{\\top\}R\_\{t\}\\big\)^\{\\vee\},\(50\)vθavg​\(Rt,xt,t,s\)\\displaystyle v\_\{\\theta\}^\{\\mathrm\{avg\}\}\(R\_\{t\},x\_\{t\},t,s\)=1t​\(xt−x^0θ\),\\displaystyle=\\frac\{1\}\{t\}\\,\(x\_\{t\}\-\\hat\{x\}\_\{0\}^\{\\theta\}\),\(51\)which is theSE​\(3\)\\mathrm\{SE\}\(3\)analogue of pixel\-space MeanFlow\. This is the “normal” objective; if one samples times bounded away from0and computesdd​t​Aθavg\\tfrac\{d\}\{dt\}A\_\{\\theta\}^\{\\mathrm\{avg\}\}via a JVP, it can be used directly\.

##### Motivation: small\-ttnumerical stability\.

Both branches contain explicit1/t1/tfactors and the rotation target further requires differentiating throughlog⁡\(\(R^0θ\)⊤​Rt\)\\log\(\(\\hat\{R\}\_\{0\}^\{\\theta\}\)^\{\\top\}R\_\{t\}\), so naive autodiff becomes numerically fragile ast→0t\\to 0\. Below we describe our implementation that preserves the same objective but avoids explicit1/t1/tduring target construction\.

#### Positivetmint\_\{\\min\}and last\-step endpoint prediction in baselines

Following prior work we enforce a strictly positivetmin\>0t\_\{\\min\}\>0and only sample training times from\[tmin,1\]\[t\_\{\\min\},1\]\. In our method we extend the lower bound totmin=10−6t\_\{\\min\}=10^\{\-6\}\.

At inference time, we integrate on a fixed gridt∈linspace​\(1,tmin,T\)t\\in\\mathrm\{linspace\}\(1,t\_\{\\min\},T\)\. In the final iteration, we set\(t,s\)=\(tmin,0\)\(t,s\)=\(t\_\{\\min\},0\)and directly use the model’s endpoint prediction\(R^0θ,x^0θ\)\(\\hat\{R\}\_\{0\}^\{\\theta\},\\hat\{x\}\_\{0\}^\{\\theta\}\)as the generated sample rather than performing another Euler update\. For the small\-ttconsistency losses we use a denominator clamptεt\_\{\\varepsilon\}when dividing byt2t^\{2\}\(only applied at the final loss normalization step\)\.

### I\.1AuxiliaryBB\-variables for stable time derivatives

The MeanFlow consistency loss involves a time derivative term of the form\(t−s\)​dd​t​Aθavg\(t\-s\)\\,\\tfrac\{d\}\{dt\}A\_\{\\theta\}^\{\\mathrm\{avg\}\}\(and similarly forvθavgv\_\{\\theta\}^\{\\mathrm\{avg\}\}\)\. Directly differentiatingAθavg=1tlog\(\(R^0θ\)⊤Rt\)∨A\_\{\\theta\}^\{\\mathrm\{avg\}\}=\\tfrac\{1\}\{t\}\\log\(\(\\hat\{R\}\_\{0\}^\{\\theta\}\)^\{\\top\}R\_\{t\}\)^\{\\vee\}can lead to large gradients whenttis small\.

We therefore introduce auxiliary variables that absorb the problematic factortt,

BθA​\(Rt,xt,t,s\)\\displaystyle B^\{A\}\_\{\\theta\}\(R\_\{t\},x\_\{t\},t,s\):=t​Aθavg​\(Rt,xt,t,s\),\\displaystyle:=t\\,A\_\{\\theta\}^\{\\mathrm\{avg\}\}\(R\_\{t\},x\_\{t\},t,s\),\(52\)Bθv​\(Rt,xt,t,s\)\\displaystyle B^\{v\}\_\{\\theta\}\(R\_\{t\},x\_\{t\},t,s\):=t​vθavg​\(Rt,xt,t,s\),\\displaystyle:=t\\,v\_\{\\theta\}^\{\\mathrm\{avg\}\}\(R\_\{t\},x\_\{t\},t,s\),\(53\)and compute time\-derivative information viadd​t​BθA\\tfrac\{d\}\{dt\}B^\{A\}\_\{\\theta\}anddd​t​Bθv\\tfrac\{d\}\{dt\}B^\{v\}\_\{\\theta\}\. In implementation we avoid dividing byttthroughout the derivative\-target construction; all1/t1/tfactors are deferred, and we only apply thet−2t^\{\-2\}normalization after forming the squared residuals \(withttclamped below bytεt\_\{\\varepsilon\}\)\. UsingBθA=t​AθavgB^\{A\}\_\{\\theta\}=tA\_\{\\theta\}^\{\\mathrm\{avg\}\}andh=t−sh=t\-s, one can rewrite

Aθs→t\\displaystyle A\_\{\\theta\}^\{s\\to t\}=h​Aθavg=ht​BθA,\\displaystyle=h\\,A\_\{\\theta\}^\{\\mathrm\{avg\}\}=\\frac\{h\}\{t\}\\,B^\{A\}\_\{\\theta\},t​dd​t​Aθs→t\\displaystyle t\\,\\frac\{d\}\{dt\}A\_\{\\theta\}^\{s\\to t\}=BθA\+h​\(dd​t​BθA−Aθavg\)=BθA\+\(h​dd​t​BθA−ht​BθA\)\.\\displaystyle=B^\{A\}\_\{\\theta\}\+h\\Big\(\\frac\{d\}\{dt\}B^\{A\}\_\{\\theta\}\-A\_\{\\theta\}^\{\\mathrm\{avg\}\}\\Big\)=B^\{A\}\_\{\\theta\}\+\\Big\(h\\,\\frac\{d\}\{dt\}B^\{A\}\_\{\\theta\}\-\\frac\{h\}\{t\}B^\{A\}\_\{\\theta\}\\Big\)\.\(54\)whereht≤1\\frac\{h\}\{t\}\\leq 1\. In implementation, we absorb the factorhhinto the JVP tangents, so the computed JVP output corresponds toh​dd​t​BθAh\\,\\tfrac\{d\}\{dt\}B^\{A\}\_\{\\theta\}directly \(and analogously forBθvB^\{v\}\_\{\\theta\}\)\. Heredd​t​BθA\\tfrac\{d\}\{dt\}B^\{A\}\_\{\\theta\}is the*total*time derivative ofBθA​\(Rt,xt,t,s\)B^\{A\}\_\{\\theta\}\(R\_\{t\},x\_\{t\},t,s\)\(withssheld fixed\)\. By the chain rule,

h​dd​t​BθA\\displaystyle h\\frac\{d\}\{dt\}B^\{A\}\_\{\\theta\}=∂BθA∂t\+⟨∂BθA∂Rt,h​d​Rtd​t⟩\+⟨∂BθA∂xt,h​d​xtd​t⟩,\\displaystyle=\\frac\{\\partial B^\{A\}\_\{\\theta\}\}\{\\partial t\}\+\\Big\\langle\\frac\{\\partial B^\{A\}\_\{\\theta\}\}\{\\partial R\_\{t\}\},\\,h\\frac\{dR\_\{t\}\}\{dt\}\\Big\\rangle\+\\Big\\langle\\frac\{\\partial B^\{A\}\_\{\\theta\}\}\{\\partial x\_\{t\}\},\\,h\\frac\{dx\_\{t\}\}\{dt\}\\Big\\rangle,\(55\)where∂BθA∂Rt\\tfrac\{\\partial B^\{A\}\_\{\\theta\}\}\{\\partial R\_\{t\}\}and∂BθA∂xt\\tfrac\{\\partial B^\{A\}\_\{\\theta\}\}\{\\partial x\_\{t\}\}include the implicit dependence through the network prediction\(R^0θ,x^0θ\)=fθ​\(Rt,xt,t,s\)\(\\hat\{R\}\_\{0\}^\{\\theta\},\\hat\{x\}\_\{0\}^\{\\theta\}\)=f\_\{\\theta\}\(R\_\{t\},x\_\{t\},t,s\)\. In practice we computedd​t​BθA\\tfrac\{d\}\{dt\}B^\{A\}\_\{\\theta\}\(anddd​t​Bθv\\tfrac\{d\}\{dt\}B^\{v\}\_\{\\theta\}\) with a JVP given the instantaneous velocities\(d​Rtd​t,d​xtd​t\)\(\\tfrac\{dR\_\{t\}\}\{dt\},\\tfrac\{dx\_\{t\}\}\{dt\}\)\.

which avoids explicitly formingdd​t​Aθavg\\tfrac\{d\}\{dt\}A\_\{\\theta\}^\{\\mathrm\{avg\}\}\. We apply the same construction to translation withBθv=t​vθavgB^\{v\}\_\{\\theta\}=t\\,v\_\{\\theta\}^\{\\mathrm\{avg\}\}\.

##### Small\-ttconsistency losses \(implementation form\)\.

To match the implementation, we form the residuals in a scaled way and defer all divisions byttto the final normalization step\. Let\(t−s\)\(t\-s\)denote the step size and define

Aθs→t\\displaystyle A\_\{\\theta\}^\{s\\to t\}:=\(t−s\)​Aθ=t−st​BθA,vθs→t:=\(t−s\)​vθ=t−st​Bθv,\\displaystyle:=\(t\-s\)A\_\{\\theta\}=\\frac\{t\-s\}\{t\}B^\{A\}\_\{\\theta\},\\qquad v\_\{\\theta\}^\{s\\to t\}:=\(t\-s\)v\_\{\\theta\}=\\frac\{t\-s\}\{t\}B^\{v\}\_\{\\theta\},\(56\)whereBθA=t​AθavgB^\{A\}\_\{\\theta\}=tA^\{\\mathrm\{avg\}\}\_\{\\theta\}andBθv=t​vθavgB^\{v\}\_\{\\theta\}=tv^\{\\mathrm\{avg\}\}\_\{\\theta\}\. With the\(t−s\)\(t\-s\)\-absorbed JVP outputs\(t−s\)​dd​t​BθA\(t\-s\)\\,\\tfrac\{d\}\{dt\}B^\{A\}\_\{\\theta\}and\(t−s\)​dd​t​Bθv\(t\-s\)\\,\\tfrac\{d\}\{dt\}B^\{v\}\_\{\\theta\}, we construct the scaled derivative targets

t​dd​t​Aθs→t\\displaystyle t\\,\\frac\{d\}\{dt\}A\_\{\\theta\}^\{s\\to t\}:=BθA\+sg​\(\(t−s\)​dd​t​BθA−t−st​BθA\),\\displaystyle:=B^\{A\}\_\{\\theta\}\+\\mathrm\{sg\}\\\!\\left\(\(t\-s\)\\,\\frac\{d\}\{dt\}B^\{A\}\_\{\\theta\}\-\\frac\{t\-s\}\{t\}B^\{A\}\_\{\\theta\}\\right\),t​dd​t​xθs→t\\displaystyle t\\,\\frac\{d\}\{dt\}x\_\{\\theta\}^\{s\\to t\}:=Bθv\+sg​\(\(t−s\)​dd​t​Bθv−t−st​Bθv\)\.\\displaystyle:=B^\{v\}\_\{\\theta\}\+\\mathrm\{sg\}\\\!\\left\(\(t\-s\)\\,\\frac\{d\}\{dt\}B^\{v\}\_\{\\theta\}\-\\frac\{t\-s\}\{t\}B^\{v\}\_\{\\theta\}\\right\)\.\(57\)where the JVP\-derived directional derivatives\(t−s\)​dd​t​BθA\(t\-s\)\\,\\tfrac\{d\}\{dt\}B^\{A\}\_\{\\theta\}and\(t−s\)​dd​t​Bθv\(t\-s\)\\,\\tfrac\{d\}\{dt\}B^\{v\}\_\{\\theta\}are computed as a single Jacobian–vector product offθ​\(Rt,xt,t,s\)f\_\{\\theta\}\(R\_\{t\},x\_\{t\},t,s\):

\(BθA,Bθv,…\),\(\(t−s\)​dd​t​BθA,\(t−s\)​dd​t​Bθv,…\)\\displaystyle\(B^\{A\}\_\{\\theta\},B^\{v\}\_\{\\theta\},\\ldots\),\\ \\big\(\(t\-s\)\\,\\tfrac\{d\}\{dt\}B^\{A\}\_\{\\theta\},\\ \(t\-s\)\\,\\tfrac\{d\}\{dt\}B^\{v\}\_\{\\theta\},\\ldots\\big\):=JVP\(R˙t,x˙t,t˙,s˙\)​\[fθ​\(Rt,xt,t,s\)\]\.\\displaystyle\\qquad:=\\ \\mathrm\{JVP\}\_\{\(\\dot\{R\}\_\{t\},\\dot\{x\}\_\{t\},\\dot\{t\},\\dot\{s\}\)\}\\big\[f\_\{\\theta\}\(R\_\{t\},x\_\{t\},t,s\)\\big\]\.\(58\)with the \(step\-size absorbed\) input tangent

\{R˙t=Rt​\(\(t−s\)​Ωt\),x˙t=\(t−s\)​vt\(MF\)\\left\\\{\\begin\{aligned\} \\dot\{R\}\_\{t\}&=R\_\{t\}\\,\\big\(\(t\-s\)\\,\\Omega\_\{t\}\\big\),\\\\ \\dot\{x\}\_\{t\}&=\(t\-s\)\\,v\_\{t\}\\end\{aligned\}\\right\.\\qquad\(\\text\{MF\}\)\(59\)or, for IMF where the instantaneous velocities are taken from a single shared forward at\(t,t\)\(t,t\),

\{R˙t=Rt​\(t−st​\(BθA​\(t,t,Rt,xt\)\)∧\),x˙t=t−st​Bθv​\(t,t,Rt,xt\)\(IMF\)\.\\left\\\{\\begin\{aligned\} \\dot\{R\}\_\{t\}&=R\_\{t\}\\,\\big\(\\frac\{t\-s\}\{t\}\\,\(B\_\{\\theta\}^\{\\mathrm\{A\}\}\(t,t,R\_\{t\},x\_\{t\}\)\)^\{\\wedge\}\\big\),\\\\ \\dot\{x\}\_\{t\}&=\\frac\{t\-s\}\{t\}B^\{v\}\_\{\\theta\}\(t,t,R\_\{t\},x\_\{t\}\)\\end\{aligned\}\\right\.\\qquad\(\\text\{IMF\}\)\.\(60\)The time tangents aret˙=\(t−s\)\\dot\{t\}=\(t\-s\)in the small\-ttparameterization \(otherwiset˙=1\\dot\{t\}=1\), ands˙=0\\dot\{s\}=0\. The per\-sample small\-ttlosses \(summing over residues\) are computed fromBθA,BθvB^\{A\}\_\{\\theta\},B^\{v\}\_\{\\theta\}and the JVP\-derived directional derivatives\(t−s\)​dd​t​BθA\(t\-s\)\\,\\tfrac\{d\}\{dt\}B^\{A\}\_\{\\theta\}and\(t−s\)​dd​t​Bθv\(t\-s\)\\,\\tfrac\{d\}\{dt\}B^\{v\}\_\{\\theta\}as

ℒrot\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{rot\}\}:=1max\(t,tε\)2​∑i=1N\{‖sg​\(J​\(Aθ,is→t\)\)​\(t​dd​t​Aθ,is→t\)−t​ωt,i‖22,jmode=J,‖t​dd​t​Aθ,is→t−sg​\(J−1​\(Aθ,is→t\)\)​t​ωt,i‖22,jmode=J−1,\\displaystyle:=\\frac\{1\}\{\\max\(t,t\_\{\\varepsilon\}\)^\{2\}\}\\sum\_\{i=1\}^\{N\}\\begin\{cases\}\\left\\\|\\mathrm\{sg\}\\\!\\left\(J\\\!\\left\(A\_\{\\theta,i\}^\{s\\to t\}\\right\)\\right\)\\left\(t\\,\\tfrac\{d\}\{dt\}A\_\{\\theta,i\}^\{s\\to t\}\\right\)\-t\\,\\omega\_\{t,i\}\\right\\\|\_\{2\}^\{2\},&\\texttt\{jmode\}=J,\\\\\[6\.0pt\] \\left\\\|t\\,\\tfrac\{d\}\{dt\}A\_\{\\theta,i\}^\{s\\to t\}\-\\mathrm\{sg\}\\\!\\left\(J^\{\-1\}\\\!\\left\(A\_\{\\theta,i\}^\{s\\to t\}\\right\)\\right\)t\\,\\omega\_\{t,i\}\\right\\\|\_\{2\}^\{2\},&\\texttt\{jmode\}=J^\{\-1\},\\end\{cases\}\(61\)ℒtrans\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{trans\}\}:=1max\(t,tε\)2​∑i=1N‖t​vt,i−t​dd​t​vθ,is→t‖22\.\\displaystyle:=\\frac\{1\}\{\\max\(t,t\_\{\\varepsilon\}\)^\{2\}\}\\sum\_\{i=1\}^\{N\}\\left\\\|t\\,v\_\{t,i\}\-t\\,\\tfrac\{d\}\{dt\}v\_\{\\theta,i\}^\{s\\to t\}\\right\\\|\_\{2\}^\{2\}\.\(62\)This is algebraically equivalent to the unscaled MeanFlow objectives but avoids explicit1/t1/tfactors during target construction; the only division byttis the finalmax\(t,tε\)−2\\max\(t,t\_\{\\varepsilon\}\)^\{\-2\}normalization applied after squaring the residuals\. The twojmodebranches share the same JVP \([58](https://arxiv.org/html/2607.27431#A9.E58)\) and scaled targets \([57](https://arxiv.org/html/2607.27431#A9.E57)\), differing only in the final residual\.

Algorithm 3Stable X\-prediction style Mean\-flow training loss1:State

\(Rt,xt\)\(R\_\{t\},x\_\{t\}\); times

s<ts<t; instantaneous velocities

\(Ωt,vt\)\(\\Omega\_\{t\},v\_\{t\}\)\(MF\) or

\(BθA​\(t,t\),Bθv​\(t,t\)\)\(B^\{A\}\_\{\\theta\}\(t,t\),B^\{v\}\_\{\\theta\}\(t,t\)\)\(IMF\); network

fθf\_\{\\theta\}; clamp

tεt\_\{\\varepsilon\}; Jacobian placement

jmode∈\{J,J−1\}\\texttt\{jmode\}\\in\\\{J,J^\{\-1\}\\\}
2:

Δ​t←t−s\\Delta t\\leftarrow t\-s
3:Construct the \(step\-size absorbed\) input tangent:

4:\(MF\)

R˙t←Rt​\(Δ​t​Ωt\)\\dot\{R\}\_\{t\}\\leftarrow R\_\{t\}\(\\Delta t\\,\\Omega\_\{t\}\),

x˙t←Δ​t​vt\\dot\{x\}\_\{t\}\\leftarrow\\Delta t\\,v\_\{t\}
5:\(IMF\)

R˙t←Rt​\(Δ​tt​\(BθA​\(t,t\)\)∧\)\\dot\{R\}\_\{t\}\\leftarrow R\_\{t\}\(\\frac\{\\Delta t\}\{t\}\\,\(B^\{A\}\_\{\\theta\}\(t,t\)\)^\{\\wedge\}\),

x˙t←Δ​tt​Bθv​\(t,t\)\\dot\{x\}\_\{t\}\\leftarrow\\frac\{\\Delta t\}\{t\}B^\{v\}\_\{\\theta\}\(t,t\)See \(Eq\. \([54](https://arxiv.org/html/2607.27431#A9.E54)\)\)\.

6:Set

t˙←Δ​t\\dot\{t\}\\leftarrow\\Delta tin the small\-

ttparameterisation \(otherwise

t˙←1\\dot\{t\}\\leftarrow 1\), and

s˙←0\\dot\{s\}\\leftarrow 0\.

7:Compute a single JVP of

fθ​\(Rt,xt,t,s\)f\_\{\\theta\}\(R\_\{t\},x\_\{t\},t,s\)along

\(R˙t,x˙t,t˙,s˙\)\(\\dot\{R\}\_\{t\},\\dot\{x\}\_\{t\},\\dot\{t\},\\dot\{s\}\)to obtain

8:primals

\(BθA,Bθv\)\(B^\{A\}\_\{\\theta\},B^\{v\}\_\{\\theta\}\)and directional derivatives

\(Δ​t​dd​t​BθA,Δ​t​dd​t​Bθv\)\\big\(\\Delta t\\,\\tfrac\{d\}\{dt\}B^\{A\}\_\{\\theta\},\\Delta t\\,\\tfrac\{d\}\{dt\}B^\{v\}\_\{\\theta\}\\big\)\(Eq\. \([58](https://arxiv.org/html/2607.27431#A9.E58)\)\)

9:Form the scaled derivative targets via Eq\. \([57](https://arxiv.org/html/2607.27431#A9.E57)\)\.

10:

𝒥←sg​\(J​\(Aθs→t\)\)\\mathcal\{J\}\\leftarrow\\mathrm\{sg\}\\big\(J\(A^\{s\\to t\}\_\{\\theta\}\)\\big\)
11:if

jmode=J\\texttt\{jmode\}=Jthen

12:

resrot←𝒥​\(t​dd​t​Aθs→t\)−t​ωt\\mathrm\{res\}\_\{\\mathrm\{rot\}\}\\leftarrow\\mathcal\{J\}\\big\(t\\,\\tfrac\{d\}\{dt\}A^\{s\\to t\}\_\{\\theta\}\\big\)\-t\\,\\omega\_\{t\}
13:else

14:

resrot←t​dd​t​Aθs→t−𝒥−1​\(t​ωt\)\\mathrm\{res\}\_\{\\mathrm\{rot\}\}\\leftarrow t\\,\\tfrac\{d\}\{dt\}A^\{s\\to t\}\_\{\\theta\}\-\\mathcal\{J\}^\{\-1\}\\big\(t\\,\\omega\_\{t\}\\big\)⊳\\trianglerightRemark[6](https://arxiv.org/html/2607.27431#Thmremark6)

15:endif

16:

restrans←t​dd​t​vθs→t−t​vt\\mathrm\{res\}\_\{\\mathrm\{trans\}\}\\leftarrow t\\,\\tfrac\{d\}\{dt\}v^\{s\\to t\}\_\{\\theta\}\-t\\,v\_\{t\}⊳\\trianglerightJ≡IJ\\equiv Ion the abelian branch

17:Normalise at the end by

max\(t,tε\)−2\\max\(t,t\_\{\\varepsilon\}\)^\{\-2\}\(Eq\. \([62](https://arxiv.org/html/2607.27431#A9.E62)\)\)

Algorithm 4Stable inference withtmin\>0t\_\{\\min\}\>0and endpoint prediction \(linear and exponential rotation schedules\)1:Prior sample

\(R1N,x1N\)\(R\_\{1\}^\{N\},x\_\{1\}^\{N\}\); network

fθf\_\{\\theta\}; steps

TT;

tmin\>0t\_\{\\min\}\>0; rotation schedule

∈\{linear,exp\}\\in\\\{\\textsc\{linear\},\\textsc\{exp\}\\\}, exp rate

c\>0c\>0
2:

\(R^0N,x^0N\)\(\\hat\{R\}\_\{0\}^\{N\},\\hat\{x\}\_\{0\}^\{N\}\)
3:

\{ti\}i=0T−1←linspace​\(1,tmin,T\)\\\{t\_\{i\}\\\}\_\{i=0\}^\{T\-1\}\\leftarrow\\mathrm\{linspace\}\(1,t\_\{\\min\},T\)
4:

\(R,x\)←\(R1N,x1N\)\(R,x\)\\leftarrow\(R\_\{1\}^\{N\},x\_\{1\}^\{N\}\)
5:for

i=0,…,T−2i=0,\\ldots,T\-2do

6:

t←ti,s←ti\+1,Δ​t←t−st\\leftarrow t\_\{i\},\\ \\ s\\leftarrow t\_\{i\+1\},\\ \\ \\Delta t\\leftarrow t\-s
7:

\(R^0θ,x^0θ\)←fθ​\(s,t,R,x\)\(\\hat\{R\}\_\{0\}^\{\\theta\},\\hat\{x\}\_\{0\}^\{\\theta\}\)\\leftarrow f\_\{\\theta\}\(s,t,R,x\)
8:

uθ←log\(\(R^0θ\)⊤R\)∨u\_\{\\theta\}\\leftarrow\\log\\big\(\(\\hat\{R\}\_\{0\}^\{\\theta\}\)^\{\\top\}R\\big\)^\{\\vee\}⊳\\trianglerightlog\-map fromRRtoward endpointR^0θ\\hat\{R\}\_\{0\}^\{\\theta\}

9:ifschedule

=linear=\\textsc\{linear\}then

10:

Aθavg←1t​uθA^\{\\mathrm\{avg\}\}\_\{\\theta\}\\leftarrow\\tfrac\{1\}\{t\}\\,u\_\{\\theta\}⊳\\trianglerightrate1/t1/t\(remaining time\)

11:else⊳\\trianglerightexp

12:

Aθavg←c​uθA^\{\\mathrm\{avg\}\}\_\{\\theta\}\\leftarrow c\\,u\_\{\\theta\}⊳\\trianglerightconstant ratecc

13:endif

14:

vθavg←1t​\(x−x^0θ\)v^\{\\mathrm\{avg\}\}\_\{\\theta\}\\leftarrow\\tfrac\{1\}\{t\}\(x\-\\hat\{x\}\_\{0\}^\{\\theta\}\)⊳\\trianglerighttranslation: linear in both schedules

15:

R←R​exp⁡\(−Δ​t​\(Aθavg\)∧\)R\\leftarrow R\\exp\\\!\\big\(\-\\Delta t\\,\(A^\{\\mathrm\{avg\}\}\_\{\\theta\}\)^\{\\wedge\}\\big\)
16:

x←x−Δ​t​vθavgx\\leftarrow x\-\\Delta t\\,v^\{\\mathrm\{avg\}\}\_\{\\theta\}
17:endfor

18:

\(R^0θ,x^0θ\)←fθ​\(0,tm​i​n,R,x\)\(\\hat\{R\}\_\{0\}^\{\\theta\},\\hat\{x\}\_\{0\}^\{\\theta\}\)\\leftarrow f\_\{\\theta\}\(0,t\_\{min\},R,x\)
19:

\(R^0N,x^0N\)←\(R^0θ,x^0θ\)\(\\hat\{R\}\_\{0\}^\{N\},\\hat\{x\}\_\{0\}^\{N\}\)\\leftarrow\(\\hat\{R\}\_\{0\}^\{\\theta\},\\hat\{x\}\_\{0\}^\{\\theta\}\)

### I\.2Rescaled MeanFlow: training and inference pseudocode

The rescaled MeanFlow training objective and the corresponding stable inference procedure are summarised in Algorithm[3](https://arxiv.org/html/2607.27431#alg3)and Algorithm[4](https://arxiv.org/html/2607.27431#alg4), respectively\. In inference, in addition to the standard linear rotation schedule, we also use an exponential \(exp\) rotation schedule inspired by ReQFlowYueet al\.\[[2025](https://arxiv.org/html/2607.27431#bib.bib22)\]\. The main idea is to redefine the effective rotation step sizes so that steps are larger on the noise side \(t≈1t\\approx 1\) and become smaller and more fine\-grained near the data side \(t≈0t\\approx 0\), which empirically improves few\-step generation stability\.

## Appendix JStable Training Framework, Part 2: JVP\-Free Training via SE\(3\)α\\alpha\-Flow

The stabilisation in Section[I](https://arxiv.org/html/2607.27431#A9)mitigates, but does not remove, the core difficulty of the*differential*consistency target: it still requires the trajectory derivativedd​t​Aθavg\\tfrac\{d\}\{dt\}A^\{\\mathrm\{avg\}\}\_\{\\theta\}via a Jacobian–vector product \(JVP\) offθf\_\{\\theta\}\. SinceAθavg=1tlog\(\(R^0θ\)⊤Rt\)∨A^\{\\mathrm\{avg\}\}\_\{\\theta\}=\\tfrac\{1\}\{t\}\\log\(\(\\hat\{R\}\_\{0\}^\{\\theta\}\)^\{\\top\}R\_\{t\}\)^\{\\vee\}already carries a1/t1/tfactor, differentiating it injects a second1/t1/t, so a head errorε\\varepsilonon the rotation output propagates to the target at orderε/t2\\varepsilon/t^\{2\}\. This is benign for the flat translation branch but dominates the instability of the rotation branch, where it compounds withSO​\(3\)\\mathrm\{SO\}\(3\)curvature and fragile forward\-mode autodiff through the IPA trunk\. We adopt theα\\alpha\-Flow framework ofZhanget al\.\[[2025](https://arxiv.org/html/2607.27431#bib.bib25)\], which replaces the differential target by a two\-evaluation, JVP\-free consistency target, reducing the amplification to𝒪​\(ε/t\)\\mathcal\{O\}\(\\varepsilon/t\)\.

### J\.1Model parameterization: endpoint and average\-velocity heads

The networkfθf\_\{\\theta\}, evaluated at input\(s,t,Rt,xt\)\(s,t,R\_\{t\},x\_\{t\}\), may be read either as an*endpoint*\(xx\-\)predictor returning\(R^0θ,x^0θ\)\(\\hat\{R\}\_\{0\}^\{\\theta\},\\hat\{x\}\_\{0\}^\{\\theta\}\), or as an*average\-velocity*\(uu\-\)predictor returning\(Aθavg,vθavg\)\(A^\{\\mathrm\{avg\}\}\_\{\\theta\},v^\{\\mathrm\{avg\}\}\_\{\\theta\}\)— the mean body angular velocity and mean translational velocity of the predicted geodesic from the endpoint to\(Rt,xt\)\(R\_\{t\},x\_\{t\}\)\. The two views are equivalent and can be transferred to each other:

\(rotation\)Aθavg=1tlog\(\(R^0θ\)⊤Rt\)∨⟺R^0θ=Rtexp\(−tAθavg∧\),\\displaystyle A^\{\\mathrm\{avg\}\}\_\{\\theta\}=\\tfrac\{1\}\{t\}\\log\\\!\\big\(\(\\hat\{R\}\_\{0\}^\{\\theta\}\)^\{\\top\}R\_\{t\}\\big\)^\{\\vee\}\\;\\;\\Longleftrightarrow\\;\\;\\hat\{R\}\_\{0\}^\{\\theta\}=R\_\{t\}\\exp\\\!\\big\(\-t\\,A^\{\\mathrm\{avg\}\\wedge\}\_\{\\theta\}\\big\),\(63\)\(translation\)vθavg=1t​\(xt−x^0θ\)⟺x^0θ=xt−t​vθavg\.\\displaystyle v^\{\\mathrm\{avg\}\}\_\{\\theta\}=\\tfrac\{1\}\{t\}\\big\(x\_\{t\}\-\\hat\{x\}\_\{0\}^\{\\theta\}\\big\)\\;\\;\\Longleftrightarrow\\;\\;\\hat\{x\}\_\{0\}^\{\\theta\}=x\_\{t\}\-t\\,v^\{\\mathrm\{avg\}\}\_\{\\theta\}\.\(64\)Both maps depend on\(s,t,Rt,xt\)\(s,t,R\_\{t\},x\_\{t\}\); we suppress the arguments where clear, and writeAθavg​\(s,t,Rt,xt\)A^\{\\mathrm\{avg\}\}\_\{\\theta\}\(s,t,R\_\{t\},x\_\{t\}\),R^0θ​\(Rt,xt,s,t\)\\hat\{R\}\_\{0\}^\{\\theta\}\(R\_\{t\},x\_\{t\},s,t\)etc\. when needed\. We develop theα\\alpha\-Flow targets and losses in the average\-velocity variables\(Aθavg,vθavg\)\(A^\{\\mathrm\{avg\}\}\_\{\\theta\},v^\{\\mathrm\{avg\}\}\_\{\\theta\}\), where the construction is most transparent \(the*regular\-tt*form\), and switch to the endpoint variables only to expose and cancel the small\-denominator factors, giving the numerically stabilised*small\-tt*form\. Throughout, the data\-side instantaneous quantities are the body angular velocityωt=Ω​\(t,Rt\)∨\\omega\_\{t\}=\\Omega\(t,R\_\{t\}\)^\{\\vee\}and the translational velocityvt=v​\(t,xt\)=x1−x0v\_\{t\}=v\(t,x\_\{t\}\)=x\_\{1\}\-x\_\{0\}\.

### J\.2Time grid

Letα∈\[αmin,1\]\\alpha\\in\[\\alpha\_\{\\min\},1\]be the consistency\-step ratio, floored by a smallαmin\>0\\alpha\_\{\\min\}\>0\. With0≤s≤t≤10\\leq s\\leq t\\leq 1, define the intermediate time and step

m=α​s\+\(1−α\)​t,δ=t−m=α​\(t−s\),m−s=\(1−α\)​\(t−s\),s≤m≤t,\\displaystyle m=\\alpha s\+\(1\-\\alpha\)t,\\qquad\\delta=t\-m=\\alpha\(t\-s\),\\qquad m\-s=\(1\-\\alpha\)\(t\-s\),\\qquad s\\leq m\\leq t,so that\(m−s\)\+δ=t−s\(m\-s\)\+\\delta=t\-s\. The stepped\-back \(intermediate\) state, obtained by integrating the shift velocities backward over\[m,t\]\[m,t\], is

xm=xt−δ​v~t,Rm=Rt​exp⁡\(−δ​ω~t∧\),\\displaystyle x\_\{m\}=x\_\{t\}\-\\delta\\,\\tilde\{v\}\_\{t\},\\qquad R\_\{m\}=R\_\{t\}\\exp\(\-\\delta\\,\\tilde\{\\omega\}\_\{t\}^\{\\wedge\}\),with the shift velocitiesv~t,ω~t\\tilde\{v\}\_\{t\},\\tilde\{\\omega\}\_\{t\}fixed below\.

### J\.3Translation target

FollowingZhanget al\.\[[2025](https://arxiv.org/html/2607.27431#bib.bib25)\], the average velocity over\[s,t\]\[s,t\]decomposes across the split pointmminto a*near*segment\[m,t\]\[m,t\]and a*far*segment\[s,m\]\[s,m\]\. The near segment uses the shift velocity

v~t=\{vt,\(data velocity; flow\-matching / MeanFlow choice\),uθx​\(m,t,xt\)=vθavg​\(m,t,xt\),\(model prediction; Shortcut choice\),\\tilde\{v\}\_\{t\}=\\begin\{cases\}v\_\{t\},&\\text\{\(data velocity; flow\-matching / MeanFlow choice\)\},\\\\\[2\.0pt\] u^\{x\}\_\{\\theta\}\(m,t,x\_\{t\}\)=v^\{\\mathrm\{avg\}\}\_\{\\theta\}\(m,t,x\_\{t\}\),&\\text\{\(model prediction; Shortcut choice\)\},\\end\{cases\}and we adopt the data velocityv~t=vt\\tilde\{v\}\_\{t\}=v\_\{t\}\. The far segment uses the stop\-gradient model average velocity at the stepped\-back statexm=xt−δ​v~tx\_\{m\}=x\_\{t\}\-\\delta\\tilde\{v\}\_\{t\},

vθavg​\(s,m,xm\)=1m​\(xm−x^0θ​\(Rm,xm,s,m\)\),\\displaystyle v^\{\\mathrm\{avg\}\}\_\{\\theta\}\(s,m,x\_\{m\}\)=\\tfrac\{1\}\{m\}\\big\(x\_\{m\}\-\\hat\{x\}\_\{0\}^\{\\theta\}\(R\_\{m\},x\_\{m\},s,m\)\\big\),and theα\\alpha\-Flow target is the convex combination

vtgtavg=α​v~t\+\(1−α\)​vθavg​\(s,m,xm\)\.\\displaystyle v^\{\\mathrm\{avg\}\}\_\{\\mathrm\{tgt\}\}=\\alpha\\,\\tilde\{v\}\_\{t\}\+\(1\-\\alpha\)\\,v^\{\\mathrm\{avg\}\}\_\{\\theta\}\(s,m,x\_\{m\}\)\.\(65\)
##### Regular\-ttloss\.

Regressing the model average velocity directly,

ℒtransreg=1α​∑i=1N‖vθ,iavg​\(s,t,xt\)−sg​\(vtgt,iavg\)‖22\.\\displaystyle\\mathcal\{L\}^\{\\mathrm\{reg\}\}\_\{\\mathrm\{trans\}\}=\\frac\{1\}\{\\alpha\}\\sum\_\{i=1\}^\{N\}\\big\\\|v^\{\\mathrm\{avg\}\}\_\{\\theta,i\}\(s,t,x\_\{t\}\)\-\\mathrm\{sg\}\\big\(v^\{\\mathrm\{avg\}\}\_\{\\mathrm\{tgt\},i\}\\big\)\\big\\\|\_\{2\}^\{2\}\.\(66\)The1/α1/\\alphaprefactor matches theα→0\\alpha\\to 0scaling of the residual \(Prop\.[8](https://arxiv.org/html/2607.27431#Thmproposition8)\); no furthers,ts,t\-dependent weight is needed because the abelian target is a plain convex combination\.

##### Small\-tt\(stabilised\) form\.

The average\-velocity variables carry an explicit1/t1/t\(viavavg=1t​\(xt−x^0\)v^\{\\mathrm\{avg\}\}=\\tfrac\{1\}\{t\}\(x\_\{t\}\-\\hat\{x\}\_\{0\}\)\) and1/m1/m\(in the far term\), which amplify head error ast,m→0t,m\\to 0\. Deferring the1/t1/tinto the displacement variableBθx=t​vθavg=xt−x^0θB^\{x\}\_\{\\theta\}=t\\,v^\{\\mathrm\{avg\}\}\_\{\\theta\}=x\_\{t\}\-\\hat\{x\}\_\{0\}^\{\\theta\}removes the model\-side division, and the corresponding displacement target is

Btgtx=t​vtgtavg=α​t​v~t\+\(1−α\)​tm​\(xm−x^0θ​\(Rm,xm,s,m\)\),\\displaystyle B^\{x\}\_\{\\mathrm\{tgt\}\}=t\\,v^\{\\mathrm\{avg\}\}\_\{\\mathrm\{tgt\}\}=\\alpha t\\,\\tilde\{v\}\_\{t\}\+\\frac\{\(1\-\\alpha\)t\}\{m\}\\big\(x\_\{m\}\-\\hat\{x\}\_\{0\}^\{\\theta\}\(R\_\{m\},x\_\{m\},s,m\)\\big\),where0<\(1−α\)​tm≤10<\\tfrac\{\(1\-\\alpha\)t\}\{m\}\\leq 1is numerically bounded\. Regressing on displacements with a floored normaliser gives

ℒtrans=1αmax\(t,tε\)2​∑i=1N‖Bθ,ix−sg​\(Btgt,ix\)‖22\.\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{trans\}\}=\\frac\{1\}\{\\alpha\\,\\max\(t,t\_\{\\varepsilon\}\)^\{2\}\}\\sum\_\{i=1\}^\{N\}\\big\\\|B^\{x\}\_\{\\theta,i\}\-\\mathrm\{sg\}\(B^\{x\}\_\{\\mathrm\{tgt\},i\}\)\\big\\\|\_\{2\}^\{2\}\.\(67\)Since‖vθavg−vtgtavg‖22=t−2​‖Bθx−Btgtx‖22\\\|v^\{\\mathrm\{avg\}\}\_\{\\theta\}\-v^\{\\mathrm\{avg\}\}\_\{\\mathrm\{tgt\}\}\\\|\_\{2\}^\{2\}=t^\{\-2\}\\\|B^\{x\}\_\{\\theta\}\-B^\{x\}\_\{\\mathrm\{tgt\}\}\\\|\_\{2\}^\{2\}, \([67](https://arxiv.org/html/2607.27431#A10.E67)\) coincides with \([66](https://arxiv.org/html/2607.27431#A10.E66)\) fort≥tεt\\geq t\_\{\\varepsilon\}; the floor only regularises the vanishing\-ttlimit\.

### J\.4Rotation target

We mirror the translation branch onSO​\(3\)\\mathrm\{SO\}\(3\), replacing vector addition by group composition\. For0≤a≤b≤10\\leq a\\leq b\\leq 1, the accumulated relative rotation is

D​\(a,b\):=𝒯​exp⁡\(∫abΩ​\(τ\)​𝑑τ\)=exp⁡\(\(b−a\)​Aavg​\(a,b\)∧\)=Ra⊤​Rb∈SO​\(3\),\\displaystyle D\(a,b\):=\\mathcal\{T\}\\exp\\\!\\Big\(\\int\_\{a\}^\{b\}\\Omega\(\\tau\)\\,d\\tau\\Big\)=\\exp\\\!\\big\(\(b\-a\)\\,A^\{\\mathrm\{avg\}\}\(a,b\)^\{\\wedge\}\\big\)=R\_\{a\}^\{\\top\}R\_\{b\}\\in\\mathrm\{SO\}\(3\),\(68\)where the middle equality collapses the time\-ordered exponential to a single generator and is exact under the geodesic \(constant\-body\-velocity\) assumption, which holds for the data path and for thexx\-prediction geodesic toR^0θ\\hat\{R\}\_\{0\}^\{\\theta\}\.

###### Proposition 7\(Interval additivity\)\.

For anys≤m≤ts\\leq m\\leq t,D​\(s,t\)=D​\(s,m\)​D​\(m,t\)\\;D\(s,t\)=D\(s,m\)\\,D\(m,t\)\.

###### Proof\.

D​\(s,t\)=Rs⊤​Rt=\(Rs⊤​Rm\)​\(Rm⊤​Rt\)=D​\(s,m\)​D​\(m,t\)D\(s,t\)=R\_\{s\}^\{\\top\}R\_\{t\}=\(R\_\{s\}^\{\\top\}R\_\{m\}\)\(R\_\{m\}^\{\\top\}R\_\{t\}\)=D\(s,m\)D\(m,t\)\. The regrouping is exact \(independent of the geodesic assumption\) and is the non\-commutativeSO​\(3\)\\mathrm\{SO\}\(3\)analogue of Euclidean displacement additivity: addition becomes group multiplication, ordered far \(\[s,m\]\[s,m\]\) before near \(\[m,t\]\[m,t\]\)\. ∎

Mirroring \([65](https://arxiv.org/html/2607.27431#A10.E65)\), the near segment\[m,t\]\[m,t\]uses the shift angular velocity

ω~t=\{ωt=Ω​\(t,Rt\)∨,\(data angular velocity\),Aθavg​\(m,t,Rt,xt\),\(model prediction; Shortcut choice\),\\tilde\{\\omega\}\_\{t\}=\\begin\{cases\}\\omega\_\{t\}=\\Omega\(t,R\_\{t\}\)^\{\\vee\},&\\text\{\(data angular velocity\)\},\\\\\[2\.0pt\] A^\{\\mathrm\{avg\}\}\_\{\\theta\}\(m,t,R\_\{t\},x\_\{t\}\),&\\text\{\(model prediction; Shortcut choice\)\},\\end\{cases\}and we adopt the data angular velocityω~t=ωt\\tilde\{\\omega\}\_\{t\}=\\omega\_\{t\}\. The far segment\[s,m\]\[s,m\]uses the stop\-gradient model average angular velocity at the stepped\-back state\(Rm,xm\)\(R\_\{m\},x\_\{m\}\),Rm=Rt​exp⁡\(−δ​ω~t∧\)R\_\{m\}=R\_\{t\}\\exp\(\-\\delta\\,\\tilde\{\\omega\}\_\{t\}^\{\\wedge\}\):

Am:=Aθavg\(s,m,Rm,xm\)=1mlog\(\(R^0θ\(Rm,xm,s,m\)\)⊤Rm\)∨\.\\displaystyle A\_\{m\}:=A^\{\\mathrm\{avg\}\}\_\{\\theta\}\(s,m,R\_\{m\},x\_\{m\}\)=\\tfrac\{1\}\{m\}\\log\\\!\\big\(\(\\hat\{R\}\_\{0\}^\{\\theta\}\(R\_\{m\},x\_\{m\},s,m\)\)^\{\\top\}R\_\{m\}\\big\)^\{\\vee\}\.ThenD​\(s,m\)=exp⁡\(\(m−s\)​Am∧\)D\(s,m\)=\\exp\\\!\\big\(\(m\-s\)A\_\{m\}^\{\\wedge\}\\big\)andD​\(m,t\)=exp⁡\(δ​ω~t∧\)=Rm⊤​RtD\(m,t\)=\\exp\\\!\\big\(\\delta\\,\\tilde\{\\omega\}\_\{t\}^\{\\wedge\}\\big\)=R\_\{m\}^\{\\top\}R\_\{t\}, withxm=xt−δ​v~tx\_\{m\}=x\_\{t\}\-\\delta\\tilde\{v\}\_\{t\}as in Section[J\.3](https://arxiv.org/html/2607.27431#A10.SS3)\.

##### Regular\-tttarget and loss\.

By Proposition[7](https://arxiv.org/html/2607.27431#Thmproposition7)and \([68](https://arxiv.org/html/2607.27431#A10.E68)\),log⁡D​\(s,t\)∨=\(t−s\)​Atgtavg\\log D\(s,t\)^\{\\vee\}=\(t\-s\)\\,A^\{\\mathrm\{avg\}\}\_\{\\mathrm\{tgt\}\}, so the average\-velocity target is the group\-composition analogue of \([65](https://arxiv.org/html/2607.27431#A10.E65)\):

Atgtavg=1t−slog\(exp⁡\(\(m−s\)​Am∧\)⏟D​\(s,m\)exp⁡\(δ​ω~t∧\)⏟D​\(m,t\)\)∨\.\\displaystyle A^\{\\mathrm\{avg\}\}\_\{\\mathrm\{tgt\}\}=\\frac\{1\}\{t\-s\}\\,\\log\\\!\\Big\(\\underbrace\{\\exp\\\!\\big\(\(m\-s\)A\_\{m\}^\{\\wedge\}\\big\)\}\_\{D\(s,m\)\}\\;\\underbrace\{\\exp\\\!\\big\(\\delta\\,\\tilde\{\\omega\}\_\{t\}^\{\\wedge\}\\big\)\}\_\{D\(m,t\)\}\\Big\)^\{\\\!\\vee\}\.\(69\)The loss mirrors \([66](https://arxiv.org/html/2607.27431#A10.E66)\):

ℒrotreg=1α​∑i=1N‖Aθ,iavg​\(s,t\)−sg​\(Atgt,iavg\)‖22\.\\displaystyle\\mathcal\{L\}^\{\\mathrm\{reg\}\}\_\{\\mathrm\{rot\}\}=\\frac\{1\}\{\\alpha\}\\sum\_\{i=1\}^\{N\}\\big\\\|A^\{\\mathrm\{avg\}\}\_\{\\theta,i\}\(s,t\)\-\\mathrm\{sg\}\\big\(A^\{\\mathrm\{avg\}\}\_\{\\mathrm\{tgt\},i\}\\big\)\\big\\\|\_\{2\}^\{2\}\.\(70\)The scalar1t−s\\tfrac\{1\}\{t\-s\}must remain*outside*log⁡\(exp⋅exp\)\\log\(\\exp\\cdot\\exp\): becauseAmA\_\{m\}andω~t\\tilde\{\\omega\}\_\{t\}do not commute, folding it into the two exponentials would rescale each segment angle and alter the Baker–Campbell–Hausdorff \(BCH\) cross term, no longer yieldinglog⁡D​\(s,t\)\\log D\(s,t\)\. This is the non\-commutative counterpart of the translation branch, where the same normalization is instead absorbed into the linear convex weightsα,1−α\\alpha,\\,1\-\\alpha\.

##### Small\-tt\(stabilised\) form\.

Two small\-denominator factors appear: the overall1/t1/tshared with translation, and the rotation\-specific1/\(t−s\)1/\(t\-s\)together withlog⁡\(⋅\)\\log\(\\cdot\)nearII\. Deferring the1/t1/tinto the displacementBθA=tAθavg=log\(\(R^0θ\)⊤Rt\)∨B^\{A\}\_\{\\theta\}=t\\,A^\{\\mathrm\{avg\}\}\_\{\\theta\}=\\log\(\(\\hat\{R\}\_\{0\}^\{\\theta\}\)^\{\\top\}R\_\{t\}\)^\{\\vee\}gives the displacement targetBtgtA=t​AtgtavgB^\{A\}\_\{\\mathrm\{tgt\}\}=t\\,A^\{\\mathrm\{avg\}\}\_\{\\mathrm\{tgt\}\},

BtgtA=tt−slog\(exp⁡\(\(m−sm​BmA\)∧\)⏟D​\(s,m\)exp⁡\(δ​ω~t∧\)⏟D​\(m,t\)\)∨,BmA=log\(\(R^0θ\(Rm,xm,s,m\)\)⊤Rm\)∨,\\displaystyle\\boxed\{\\;B^\{A\}\_\{\\mathrm\{tgt\}\}=\\frac\{t\}\{t\-s\}\\,\\log\\\!\\Big\(\\underbrace\{\\exp\\\!\\big\(\(\\tfrac\{m\-s\}\{m\}B^\{A\}\_\{m\}\)^\{\\wedge\}\\big\)\}\_\{D\(s,m\)\}\\;\\underbrace\{\\exp\\\!\\big\(\\delta\\,\\tilde\{\\omega\}\_\{t\}^\{\\wedge\}\\big\)\}\_\{D\(m,t\)\}\\Big\)^\{\\\!\\vee\},\\qquad B^\{A\}\_\{m\}=\\log\\\!\\big\(\(\\hat\{R\}\_\{0\}^\{\\theta\}\(R\_\{m\},x\_\{m\},s,m\)\)^\{\\top\}R\_\{m\}\\big\)^\{\\vee\},\\;\}\(71\)using\(m−s\)​Am=m−sm​BmA\(m\-s\)A\_\{m\}=\\tfrac\{m\-s\}\{m\}B^\{A\}\_\{m\}\. Both segment angles scale asO​\(t−s\)O\(t\-s\)and are bounded:m−sm​BmA\\tfrac\{m\-s\}\{m\}B^\{A\}\_\{m\}hasm−sm≤1\\tfrac\{m\-s\}\{m\}\\leq 1and main log≤π\\leq\\pi, whileδ​ω~t=α​\(t−s\)​ωt\\delta\\,\\tilde\{\\omega\}\_\{t\}=\\alpha\(t\-s\)\\,\\omega\_\{t\}with‖ωt‖≤π\\\|\\omega\_\{t\}\\\|\\leq\\pi\. Hence‖log⁡D​\(s,t\)∨‖=O​\(t−s\)\\\|\\log D\(s,t\)^\{\\vee\}\\\|=O\(t\-s\), the1t−s\\tfrac\{1\}\{t\-s\}factor is cancelled at the same rate, and

‖BtgtA‖≤π​\(\(1−α\)​tm\+α​t\)≤2​π,\\displaystyle\\big\\\|B^\{A\}\_\{\\mathrm\{tgt\}\}\\big\\\|\\leq\\pi\\Big\(\\tfrac\{\(1\-\\alpha\)t\}\{m\}\+\\alpha t\\Big\)\\leq 2\\pi,mirroring the bounded weights\(1−α\)​tm,α​t≤1\\tfrac\{\(1\-\\alpha\)t\}\{m\},\\,\\alpha t\\leq 1of the translation target: the target is analytically free of small\-denominator blow\-up\. For numerical stability whent−st\-sis below machine tolerancetεt\_\{\\varepsilon\}\(wherelog⁡\(⋅\)\\log\(\\cdot\)nearIIloses precision\), we avoid forming1t−s\\tfrac\{1\}\{t\-s\}and use the first\-order BCH limit, whoseO​\(t−s\)O\(t\-s\)correction is then negligible:

BtgtA=\{tt−slog\(D\(s,m\)D\(m,t\)\)∨,t−s≥tε,t​\[\(1−α\)​Am\+α​ω~t\],t−s<tε\.\\displaystyle B^\{A\}\_\{\\mathrm\{tgt\}\}=\\begin\{cases\}\\dfrac\{t\}\{t\-s\}\\,\\log\\\!\\big\(D\(s,m\)\\,D\(m,t\)\\big\)^\{\\vee\},&t\-s\\geq t\_\{\\varepsilon\},\\\\\[8\.0pt\] t\\big\[\(1\-\\alpha\)A\_\{m\}\+\\alpha\\,\\tilde\{\\omega\}\_\{t\}\\big\],&t\-s<t\_\{\\varepsilon\}\.\\end\{cases\}The rotation loss is then identical in form to \([67](https://arxiv.org/html/2607.27431#A10.E67)\),

ℒrot=1αmax\(t,tε\)2​∑i=1N‖Bθ,iA−sg​\(Btgt,iA\)‖22,\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{rot\}\}=\\frac\{1\}\{\\alpha\\,\\max\(t,t\_\{\\varepsilon\}\)^\{2\}\}\\sum\_\{i=1\}^\{N\}\\big\\\|B^\{A\}\_\{\\theta,i\}\-\\mathrm\{sg\}\(B^\{A\}\_\{\\mathrm\{tgt\},i\}\)\\big\\\|\_\{2\}^\{2\},\(72\)and, via‖Aθavg−Atgtavg‖22=t−2​‖BθA−BtgtA‖22\\\|A^\{\\mathrm\{avg\}\}\_\{\\theta\}\-A^\{\\mathrm\{avg\}\}\_\{\\mathrm\{tgt\}\}\\\|\_\{2\}^\{2\}=t^\{\-2\}\\\|B^\{A\}\_\{\\theta\}\-B^\{A\}\_\{\\mathrm\{tgt\}\}\\\|\_\{2\}^\{2\}, coincides with \([70](https://arxiv.org/html/2607.27431#A10.E70)\) whenevert≥tεt\\geq t\_\{\\varepsilon\}andt−s≥tεt\-s\\geq t\_\{\\varepsilon\}\. The target requires one logarithm and two exponentials and no differentiation offθf\_\{\\theta\}\.

### J\.5Theα→0\\alpha\\to 0Limit Recovers MeanFlow

In this subsection we show that theα\\alpha\-Flow loss converges to the differential MeanFlow loss asα→0\\alpha\\to 0\. We treat the rotation branch; the translation branch is the abelian case \(J≡IJ\\equiv I, no BCH correction\) and reduces to Euclidean MeanFlow \([3](https://arxiv.org/html/2607.27431#S2.E3)\) verbatim\.

###### Proposition 8\(α→0\\alpha\\to 0recovers the MeanFlow loss\)\.

Letg​\(τ\):=\(τ−s\)​Aθavg​\(s,τ,Rτ,xτ\)g\(\\tau\):=\(\\tau\-s\)A^\{\\mathrm\{avg\}\}\_\{\\theta\}\(s,\\tau,R\_\{\\tau\},x\_\{\\tau\}\)denote the model log\-displacement along the interpolation path, so thatg​\(t\)=\(t−s\)​Aθavg​\(s,t\)g\(t\)=\(t\-s\)A^\{\\mathrm\{avg\}\}\_\{\\theta\}\(s,t\)and, by the product rule,g˙​\(t\)=Aθavg\+\(t−s\)​dd​t​Aθavg\\dot\{g\}\(t\)=A^\{\\mathrm\{avg\}\}\_\{\\theta\}\+\(t\-s\)\\tfrac\{d\}\{dt\}A^\{\\mathrm\{avg\}\}\_\{\\theta\}is the bracketed quantity of \([8](https://arxiv.org/html/2607.27431#S3.E8)\)\. IfggisC2C^\{2\}ontt, then the target \([69](https://arxiv.org/html/2607.27431#A10.E69)\) expands as

Atgtavg=Aθavg​\(s,t\)\+α​\(J​\(g​\(t\)\)−1​ωt−g˙​\(t\)\)\+O​\(α2\)\\displaystyle A^\{\\mathrm\{avg\}\}\_\{\\mathrm\{tgt\}\}=A^\{\\mathrm\{avg\}\}\_\{\\theta\}\(s,t\)\+\\alpha\\Big\(J\\big\(g\(t\)\\big\)^\{\-1\}\\omega\_\{t\}\-\\dot\{g\}\(t\)\\Big\)\+O\(\\alpha^\{2\}\)\(74\)and consequently, per residue,

ℒ~rot=‖g˙​\(t\)−J−1​ωt‖22\+O​\(α\),\\displaystyle\\widetilde\{\\mathcal\{L\}\}\_\{\\mathrm\{rot\}\}=\\big\\\|\\dot\{g\}\(t\)\-J^\{\-1\}\\omega\_\{t\}\\big\\\|\_\{2\}^\{2\}\+O\(\\alpha\),\(75\)i\.e\. exactly the residual of theJ−1J^\{\-1\}form \([12](https://arxiv.org/html/2607.27431#S3.E12)\) of the MeanFlow identity\.

###### Proof\.

Since \([73](https://arxiv.org/html/2607.27431#A10.E73)\) decouples across residues, we takeN=1N=1and drop the indexii\. Writem=t−δm=t\-\\deltawithδ=α​\(t−s\)\\delta=\\alpha\(t\-s\), andJ:=J​\(g​\(t\)\)J:=J\(g\(t\)\)\.

Since the data velocitiesωt\\omega\_\{t\}andvtv\_\{t\}are constant along the interpolation path \(Appendix[D](https://arxiv.org/html/2607.27431#A4)\), the stepped\-back state satisfiesRt​exp⁡\(−δ​ωt∧\)=RmR\_\{t\}\\exp\(\-\\delta\\,\\omega\_\{t\}^\{\\wedge\}\)=R\_\{m\}andxt−δ​vt=xmx\_\{t\}\-\\delta v\_\{t\}=x\_\{m\}*exactly*\. The two factors of \([69](https://arxiv.org/html/2607.27431#A10.E69)\) are thereforeD​\(s,m\)=exp⁡\(g​\(m\)∧\)D\(s,m\)=\\exp\(g\(m\)^\{\\wedge\}\)andD​\(m,t\)=exp⁡\(δ​ωt∧\)D\(m,t\)=\\exp\(\\delta\\,\\omega\_\{t\}^\{\\wedge\}\), withggthe same function appearing in the statement; this is what lets the finite difference below capture the*total*derivative \([9](https://arxiv.org/html/2607.27431#S3.E9)\) rather than∂t\\partial\_\{t\}alone\.

The first\-order BCH formula giveslog⁡\(eX​eY\)=X\+adX1−e−adX​Y\+O​\(‖Y‖2\)\\log\(e^\{X\}e^\{Y\}\)=X\+\\frac\{\\mathrm\{ad\}\_\{X\}\}\{1\-e^\{\-\\mathrm\{ad\}\_\{X\}\}\}Y\+O\(\\\|Y\\\|^\{2\}\)\. On𝔰​𝔬​\(3\)\\mathfrak\{so\}\(3\)one hasadϕ∧​ψ∧=\(ϕ∧​ψ\)∧\\mathrm\{ad\}\_\{\\phi^\{\\wedge\}\}\\psi^\{\\wedge\}=\(\\phi^\{\\wedge\}\\psi\)^\{\\wedge\}, so in vector coordinates\(1−e−adXadX\)∨=∫01e−u​ϕ∧​𝑑u=J​\(ϕ\)\\big\(\\tfrac\{1\-e^\{\-\\mathrm\{ad\}\_\{X\}\}\}\{\\mathrm\{ad\}\_\{X\}\}\\big\)^\{\\vee\}=\\int\_\{0\}^\{1\}e^\{\-u\\phi^\{\\wedge\}\}du=J\(\\phi\)by \([10](https://arxiv.org/html/2607.27431#S3.E10)\)\. This is invertible:J​\(ϕ\)J\(\\phi\)is a polynomial in the skew matrixϕ∧\\phi^\{\\wedge\}, so its eigenvalues are11and1−e∓i​θ±i​θ\\tfrac\{1\-e^\{\\mp i\\theta\}\}\{\\pm i\\theta\}withθ=‖ϕ‖≤π\\theta=\\\|\\phi\\\|\\leq\\pi, all nonzero\. TakingX=g​\(m\)∧X=g\(m\)^\{\\wedge\}andY=δ​ωt∧=O​\(α\)Y=\\delta\\,\\omega\_\{t\}^\{\\wedge\}=O\(\\alpha\),

log\(D\(s,m\)D\(m,t\)\)∨=g\(m\)\+δJ\(g\(m\)\)−1ωt\+O\(δ2\)\.\\displaystyle\\log\\\!\\big\(D\(s,m\)D\(m,t\)\\big\)^\{\\vee\}=g\(m\)\+\\delta\\,J\\big\(g\(m\)\\big\)^\{\-1\}\\omega\_\{t\}\+O\(\\delta^\{2\}\)\.Substitutingg​\(m\)=g​\(t\)−δ​g˙​\(t\)\+O​\(δ2\)g\(m\)=g\(t\)\-\\delta\\dot\{g\}\(t\)\+O\(\\delta^\{2\}\)andJ​\(g​\(m\)\)−1=J−1\+O​\(δ\)J\(g\(m\)\)^\{\-1\}=J^\{\-1\}\+O\(\\delta\), we have:

Atgtavg\\displaystyle A^\{\\mathrm\{avg\}\}\_\{\\mathrm\{tgt\}\}=1t−slog\(D\(s,m\)D\(m,t\)\)∨\\displaystyle=\\frac\{1\}\{t\-s\}\\log\\\!\\big\(D\(s,m\)\\,D\(m,t\)\\big\)^\{\\vee\}=1t−s​\(g​\(m\)\+δ​J​\(g​\(m\)\)−1​ωt\+O​\(δ2\)\)\\displaystyle=\\frac\{1\}\{t\-s\}\\Big\(g\(m\)\+\\delta\\,J\\big\(g\(m\)\\big\)^\{\-1\}\\omega\_\{t\}\+O\(\\delta^\{2\}\)\\Big\)=1t−s​\(g​\(t\)−δ​g˙​\(t\)\+δ​J−1​ωt\+O​\(δ2\)\)\\displaystyle=\\frac\{1\}\{t\-s\}\\Big\(g\(t\)\-\\delta\\,\\dot\{g\}\(t\)\+\\delta\\,J^\{\-1\}\\omega\_\{t\}\+O\(\\delta^\{2\}\)\\Big\)=g​\(t\)t−s\+α​\(t−s\)t−s​\(J−1​ωt−g˙​\(t\)\)\+O​\(α2​\(t−s\)2\)t−s\\displaystyle=\\frac\{g\(t\)\}\{t\-s\}\+\\frac\{\\alpha\(t\-s\)\}\{t\-s\}\\Big\(J^\{\-1\}\\omega\_\{t\}\-\\dot\{g\}\(t\)\\Big\)\+\\frac\{O\\big\(\\alpha^\{2\}\(t\-s\)^\{2\}\\big\)\}\{t\-s\}=Aθavg​\(s,t\)−α​\(g˙​\(t\)−J−1​ωt\)\+O​\(α2\),\\displaystyle=A^\{\\mathrm\{avg\}\}\_\{\\theta\}\(s,t\)\-\\alpha\\Big\(\\dot\{g\}\(t\)\-J^\{\-1\}\\omega\_\{t\}\\Big\)\+O\(\\alpha^\{2\}\),which is \([74](https://arxiv.org/html/2607.27431#A10.E74)\)\. Substituting into \([73](https://arxiv.org/html/2607.27431#A10.E73)\),

ℒ~rot\\displaystyle\\widetilde\{\\mathcal\{L\}\}\_\{\\mathrm\{rot\}\}=1α2​‖Aθavg​\(s,t\)−sg​\(Atgtavg\)‖22\\displaystyle=\\frac\{1\}\{\\alpha^\{2\}\}\\Big\\\|A^\{\\mathrm\{avg\}\}\_\{\\theta\}\(s,t\)\-\\mathrm\{sg\}\\big\(A^\{\\mathrm\{avg\}\}\_\{\\mathrm\{tgt\}\}\\big\)\\Big\\\|\_\{2\}^\{2\}=1α2​‖Aθavg​\(s,t\)−\(Aθavg​\(s,t\)−α​\(g˙​\(t\)−J−1​ωt\)\+O​\(α2\)\)‖22\\displaystyle=\\frac\{1\}\{\\alpha^\{2\}\}\\Big\\\|A^\{\\mathrm\{avg\}\}\_\{\\theta\}\(s,t\)\-\\Big\(A^\{\\mathrm\{avg\}\}\_\{\\theta\}\(s,t\)\-\\alpha\\,\\big\(\\dot\{g\}\(t\)\-J^\{\-1\}\\omega\_\{t\}\\big\)\+O\(\\alpha^\{2\}\)\\Big\)\\Big\\\|\_\{2\}^\{2\}=1α2​‖α​\(g˙​\(t\)−J−1​ωt\)\+O​\(α2\)‖22\\displaystyle=\\frac\{1\}\{\\alpha^\{2\}\}\\Big\\\|\\alpha\\,\\big\(\\dot\{g\}\(t\)\-J^\{\-1\}\\omega\_\{t\}\\big\)\+O\(\\alpha^\{2\}\)\\Big\\\|\_\{2\}^\{2\}=‖\(g˙​\(t\)−J−1​ωt\)‖22\+O​\(α\),\\displaystyle=\\big\\\|\\big\(\\dot\{g\}\(t\)\-J^\{\-1\}\\omega\_\{t\}\\big\)\\big\\\|\_\{2\}^\{2\}\+O\(\\alpha\),which is \([75](https://arxiv.org/html/2607.27431#A10.E75)\)\. ∎

![Refer to caption](https://arxiv.org/html/2607.27431v1/x9.png)Figure 8:Theα\\alpha\-Flow target interpolates between flow matching and MeanFlow\. On a smooth generator fieldg​\(τ\)g\(\\tau\)we form the rotation targetBtgtA​\(α\)=t​AtgtavgB^\{A\}\_\{\\mathrm\{tgt\}\}\(\\alpha\)=tA^\{\\text\{avg\}\}\_\{\\mathrm\{tgt\}\}of \([71](https://arxiv.org/html/2607.27431#A10.E71)\) over a range ofα\\alphaand plot its distance to the two endpoints: to the flow\-matching targett​ωtt\\,\\omega\_\{t\}\(red\), which vanishes atα=1\\alpha=1\(Remark[8](https://arxiv.org/html/2607.27431#Thmremark8)\), and to the MeanFlow targett​Aavgt\\,A^\{\\mathrm\{avg\}\}\(green\), which vanishes asα→0\\alpha\\to 0\(Proposition[8](https://arxiv.org/html/2607.27431#Thmproposition8)\)\. The plot is drawn in the displacement variableBA=t​AavgB^\{A\}=t\\,A^\{\\mathrm\{avg\}\}of Section[J\.4](https://arxiv.org/html/2607.27431#A10.SS4); the common factorttaffects both curves equally\.We verify Proposition[8](https://arxiv.org/html/2607.27431#Thmproposition8)numerically\. We take a smooth generator fieldg​\(τ\)g\(\\tau\)and a data velocityωt\\omega\_\{t\}drawn*independently*ofgg, so thatJ​\(g​\(t\)\)−1​ωt≠g˙​\(t\)J\(g\(t\)\)^\{\-1\}\\omega\_\{t\}\\neq\\dot\{g\}\(t\)in general and the first\-order term of \([74](https://arxiv.org/html/2607.27431#A10.E74)\) is non\-degenerate\. For a range ofα\\alphawe form theα\\alpha\-Flow rotation targetAtgtavg​\(α\)=1t−s​Log​\(eg​\(m\)∧​eδ​ωt∧\)∨A^\{\\mathrm\{avg\}\}\_\{\\mathrm\{tgt\}\}\(\\alpha\)=\\tfrac\{1\}\{t\-s\}\\mathrm\{Log\}\\\!\\big\(e^\{g\(m\)^\{\\wedge\}\}e^\{\\delta\\,\\omega\_\{t\}^\{\\wedge\}\}\\big\)^\{\\vee\}withm=α​s\+\(1−α\)​tm=\\alpha s\+\(1\-\\alpha\)tandδ=α​\(t−s\)\\delta=\\alpha\(t\-s\), and compare it to the zeroth\- and first\-order predictions of \([74](https://arxiv.org/html/2607.27431#A10.E74)\), computingg˙\\dot\{g\}by forward\-mode AD andJ​\(g​\(t\)\)J\(g\(t\)\)from its closed form\. The𝒪​\(α\)\\mathcal\{O\}\(\\alpha\)/𝒪​\(α2\)\\mathcal\{O\}\(\\alpha^\{2\}\)scalings \(Figure[8](https://arxiv.org/html/2607.27431#A10.F8)\) confirm the expansion and its coefficient; consistently,ℒ~rot\\widetilde\{\\mathcal\{L\}\}\_\{\\mathrm\{rot\}\}converges to∥g˙−J​\(g​\(t\)\)−1​ωt∥22\\lVert\\dot\{g\}\-J\(g\(t\)\)^\{\-1\}\\omega\_\{t\}\\rVert\_\{2\}^\{2\}\.

Algorithm 5JVP\-free SE\(3\)α\\alpha\-Flow training step1:

\(RtN,xtN\)\(R\_\{t\}^\{N\},x\_\{t\}^\{N\}\); times

s≤ts\\leq t; data

\(ΩtN,vtN\)\(\\Omega\_\{t\}^\{N\},v\_\{t\}^\{N\}\); head

fθf\_\{\\theta\};

α∈\[αmin,1\]\\alpha\\in\[\\alpha\_\{\\min\},1\]; clamp

tεt\_\{\\varepsilon\}
2:

m←α​s\+\(1−α\)​tm\\leftarrow\\alpha s\+\(1\-\\alpha\)t,

δ←t−m\\delta\\leftarrow t\-m,

ωt:=ωtN←\(ΩtN\)∨\\omega\_\{t\}:=\\omega\_\{t\}^\{N\}\\leftarrow\(\\Omega\_\{t\}^\{N\}\)^\{\\vee\}
3:

RmN←RtN​exp⁡\(−δ​ωt∧\)R\_\{m\}^\{N\}\\leftarrow R\_\{t\}^\{N\}\\exp\(\-\\delta\\omega\_\{t\}^\{\\wedge\}\),

xmN←xtN−δ​vtNx\_\{m\}^\{N\}\\leftarrow x\_\{t\}^\{N\}\-\\delta\\,v\_\{t\}^\{N\}
4:

\(R^0m,x^0m\)←sg​fθ​\(RmN,xmN,s,m\)\(\\hat\{R\}\_\{0\}^\{m\},\\hat\{x\}\_\{0\}^\{m\}\)\\leftarrow\\mathrm\{sg\}\\,f\_\{\\theta\}\(R\_\{m\}^\{N\},x\_\{m\}^\{N\},s,m\)
5:

mAm←log\(\(R^0m\)⊤RmN\)∨mA\_\{m\}\\leftarrow\\log\(\(\\hat\{R\}\_\{0\}^\{m\}\)^\{\\top\}R\_\{m\}^\{N\}\)^\{\\vee\},

m​um←xmN−x^0mmu\_\{m\}\\leftarrow x\_\{m\}^\{N\}\-\\hat\{x\}\_\{0\}^\{m\}⊳\\trianglerightdisplacements

6:

BtgtA←tt−slog\(exp\(\(m−smmAm\)∧\)exp\(δωt∧\)\)∨B^\{A\}\_\{\\mathrm\{tgt\}\}\\leftarrow\\frac\{t\}\{t\-s\}\\log\\big\(\\exp\(\(\\tfrac\{m\-s\}\{m\}\\,mA\_\{m\}\)^\{\\wedge\}\)\\exp\(\\delta\\,\\omega\_\{t\}^\{\\wedge\}\)\\big\)^\{\\vee\}
7:

Btgtx←α​t​vtN\+\(1−α\)​tm​m​umB^\{x\}\_\{\\mathrm\{tgt\}\}\\leftarrow\\alpha t\\,v\_\{t\}^\{N\}\+\\tfrac\{\(1\-\\alpha\)t\}\{m\}\\,mu\_\{m\}
8:

\(R^0θ,x^0θ\)←fθ​\(RtN,xtN,s,t\)\(\\hat\{R\}\_\{0\}^\{\\theta\},\\hat\{x\}\_\{0\}^\{\\theta\}\)\\leftarrow f\_\{\\theta\}\(R\_\{t\}^\{N\},x\_\{t\}^\{N\},s,t\)⊳\\trianglerightonly graph path

9:

BθA←log\(\(R^0θ\)⊤RtN\)∨B^\{A\}\_\{\\theta\}\\leftarrow\\log\(\(\\hat\{R\}\_\{0\}^\{\\theta\}\)^\{\\top\}R\_\{t\}^\{N\}\)^\{\\vee\},

Bθx←xtN−x^0θB^\{x\}\_\{\\theta\}\\leftarrow x\_\{t\}^\{N\}\-\\hat\{x\}\_\{0\}^\{\\theta\}
10:Form

ℒrot,ℒtrans\\mathcal\{L\}\_\{\\mathrm\{rot\}\},\\mathcal\{L\}\_\{\\mathrm\{trans\}\}\(Eqs\. \([72](https://arxiv.org/html/2607.27431#A10.E72)\),\([67](https://arxiv.org/html/2607.27431#A10.E67)\)\);

## Appendix KStable Training Framework, Part 3: Semigroup Loss of SE\(3\)\-MeanFlow

The MeanFlow identity \([8](https://arxiv.org/html/2607.27431#S3.E8)\) is the*differential*form of our average velocity: it is obtained by differentiating \([30](https://arxiv.org/html/2607.27431#A6.E30)\) intt, and its training target carries both a Jacobian–vector product \(fordd​t​Aθavg\\tfrac\{d\}\{dt\}A^\{\\mathrm\{avg\}\}\_\{\\theta\}\) and the right JacobianJJ\. In this section, we show that the*same*definition \([30](https://arxiv.org/html/2607.27431#A6.E30)\) also admits a*finite*, derivative\-free characterization: the average\-velocity flow map is a genuine two\-parameter semigroup onSE​\(3\)\\mathrm\{SE\}\(3\), and the consistency with this semigroup yields a JVP\-free,JJ\-free training objective\. Throughout, lets≤m≤t∈\[0,1\]s\\leq m\\leq t\\in\[0,1\]\. We write

ωt:=Ω​\(t,Rt,xt\)∨=Log​\(R0⊤​R1\)∨,vt:=v​\(t,xt\)=x1−x0\\omega\_\{t\}:=\\Omega\(t,R\_\{t\},x\_\{t\}\)^\{\\vee\}=\\mathrm\{Log\}\(R\_\{0\}^\{\\top\}R\_\{1\}\)^\{\\vee\},\\qquad v\_\{t\}:=v\(t,x\_\{t\}\)=x\_\{1\}\-x\_\{0\}for the data\-side instantaneous \(constant body\-frame\) velocities used in flow matching \([25](https://arxiv.org/html/2607.27431#A4.E25)\)\.

### K\.1The average\-velocity flow map

Recall from \([30](https://arxiv.org/html/2607.27431#A6.E30)\) that the average angular velocity over\[s,t\]\[s,t\]is defined by

exp\(\(t−s\)Ωavg\(s,t,Rt,xt\)\)=𝒯exp\(∫stΩ\(τ,Rτ,xτ\)dτ\)=:D\(s,t\)∈SO\(3\),\\displaystyle\\exp\\\!\\big\(\(t\-s\)\\,\\Omega^\{\\mathrm\{avg\}\}\(s,t,R\_\{t\},x\_\{t\}\)\\big\)=\\mathcal\{T\}\\exp\\\!\\Big\(\\int\_\{s\}^\{t\}\\Omega\(\\tau,R\_\{\\tau\},x\_\{\\tau\}\)\\,d\\tau\\Big\)=:D\(s,t\)\\in\\mathrm\{SO\}\(3\),\(76\)where the last equalityD​\(s,t\)=Rs⊤​RtD\(s,t\)=R\_\{s\}^\{\\top\}R\_\{t\}holds by integrating the left\-trivialized ODER˙τ=Rτ​Ω​\(τ,Rτ\)\\dot\{R\}\_\{\\tau\}=R\_\{\\tau\}\\Omega\(\\tau,R\_\{\\tau\}\); this is exactly \([68](https://arxiv.org/html/2607.27431#A10.E68)\)\. We writeAavg=\(Ωavg\)∨A^\{\\mathrm\{avg\}\}=\(\\Omega^\{\\mathrm\{avg\}\}\)^\{\\vee\}, and define the rotation log\-displacement by

As→t:=\(t−s\)​Aavg​\(s,t,Rt,xt\)∈ℝ3,thusexp⁡\(\(As→t\)∧\)=D​\(s,t\)\.A^\{s\\to t\}:=\(t\-s\)A^\{\\mathrm\{avg\}\}\(s,t,R\_\{t\},x\_\{t\}\)\\in\\mathbb\{R\}^\{3\},\\quad\\text\{thus\}\\quad\\exp\\\!\\big\(\(A^\{s\\to t\}\)^\{\\wedge\}\\big\)=D\(s,t\)\.For the translation branch onℝ3\\mathbb\{R\}^\{3\}, the decoupled product metric \([24](https://arxiv.org/html/2607.27431#A3.E24)\) gives the displacementp​\(s,t\):=xt−xs=\(t−s\)​vavg​\(s,t,xt\)p\(s,t\):=x\_\{t\}\-x\_\{s\}=\(t\-s\)\\,v^\{\\mathrm\{avg\}\}\(s,t,x\_\{t\}\)\. Based on the two notations of displacement above, we give the following definition\.

###### Definition 1\(Average\-velocity flow map\)\.

Fors≤ts\\leq t, we defineΨs→t:SE​\(3\)→SE​\(3\)\\Psi\_\{s\\to t\}:\\mathrm\{SE\}\(3\)\\to\\mathrm\{SE\}\(3\)by

Ψs→t​\(R,x\):=\(R​D​\(s,t\),x\+p​\(s,t\)\)=\(R​exp⁡\(\(As→t\)∧\),x\+\(t−s\)​vavg​\(s,t\)\)\.\\displaystyle\\Psi\_\{s\\to t\}\(R,x\):=\\Big\(\\,R\\,D\(s,t\),\\ \\ x\+p\(s,t\)\\,\\Big\)=\\Big\(\\,R\\,\\exp\\\!\\big\(\(A^\{s\\to t\}\)^\{\\wedge\}\\big\),\\ \\ x\+\(t\-s\)\\,v^\{\\mathrm\{avg\}\}\(s,t\)\\,\\Big\)\.By construction,Ψs→t​\(Rs,xs\)=\(Rt,xt\)\\Psi\_\{s\\to t\}\(R\_\{s\},x\_\{s\}\)=\(R\_\{t\},x\_\{t\}\)along the interpolation path\.

### K\.2The semigroup property

Having defined the average\-velocity flow map, we next show that it forms a semigroup under time composition\.

###### Proposition 9\(Semigroup property of the average\-velocity flow\)\.

For alls≤m≤ts\\leq m\\leq t, the family\{Ψs→t\}0≤s≤t≤1\\\{\\Psi\_\{s\\to t\}\\\}\_\{0\\leq s\\leq t\\leq 1\}defined in Definition[1](https://arxiv.org/html/2607.27431#Thmdefinition1)satisfies

Ψm→t∘Ψs→m=Ψs→t,Ψt→t=id\.\\displaystyle\\Psi\_\{m\\to t\}\\circ\\Psi\_\{s\\to m\}=\\Psi\_\{s\\to t\},\\qquad\\Psi\_\{t\\to t\}=\\mathrm\{id\}\.\(77\)Equivalently, in terms of the average velocities,

\(rotation\)exp⁡\(\(As→m\)∧\)​exp⁡\(\(Am→t\)∧\)=exp⁡\(\(As→t\)∧\),\\displaystyle\\textbf\{\(rotation\)\}\\quad\\exp\\\!\\big\(\(A^\{s\\to m\}\)^\{\\wedge\}\\big\)\\,\\exp\\\!\\big\(\(A^\{m\\to t\}\)^\{\\wedge\}\\big\)=\\exp\\\!\\big\(\(A^\{s\\to t\}\)^\{\\wedge\}\\big\),\(78\)\(translation\)\(m−s\)​vavg​\(s,m\)\+\(t−m\)​vavg​\(m,t\)=\(t−s\)​vavg​\(s,t\)\.\\displaystyle\\textbf\{\(translation\)\}\\quad\(m\-s\)\\,v^\{\\mathrm\{avg\}\}\(s,m\)\+\(t\-m\)\\,v^\{\\mathrm\{avg\}\}\(m,t\)=\(t\-s\)\\,v^\{\\mathrm\{avg\}\}\(s,t\)\.\(79\)

###### Proof\.

OnSO​\(3\)\\mathrm\{SO\}\(3\), the transition operators compose by insertingRm​Rm⊤=IR\_\{m\}R\_\{m\}^\{\\top\}=I:

D​\(s,t\)=Rs⊤​Rt=\(Rs⊤​Rm\)​\(Rm⊤​Rt\)=D​\(s,m\)​D​\(m,t\),D\(s,t\)=R\_\{s\}^\{\\top\}R\_\{t\}=\(R\_\{s\}^\{\\top\}R\_\{m\}\)\(R\_\{m\}^\{\\top\}R\_\{t\}\)=D\(s,m\)\\,D\(m,t\),which is Proposition[7](https://arxiv.org/html/2607.27431#Thmproposition7)\. \(Equivalently, this is the multiplicativity of the time\-ordered exponential𝒯​exp⁡\(∫st\)=𝒯​exp⁡\(∫sm\)​𝒯​exp⁡\(∫mt\)\\mathcal\{T\}\\exp\(\\int\_\{s\}^\{t\}\)=\\mathcal\{T\}\\exp\(\\int\_\{s\}^\{m\}\)\\,\\mathcal\{T\}\\exp\(\\int\_\{m\}^\{t\}\), so the statement holds for an arbitrary velocity field, not only for the geodesic path\.\) SubstitutingD​\(a,b\)=exp⁡\(\(Aa→b\)∧\)D\(a,b\)=\\exp\(\(A^\{a\\to b\}\)^\{\\wedge\}\)gives \([78](https://arxiv.org/html/2607.27431#A11.E78)\)\. On the flat translation factor, displacements add,xt−xs=\(xt−xm\)\+\(xm−xs\)x\_\{t\}\-x\_\{s\}=\(x\_\{t\}\-x\_\{m\}\)\+\(x\_\{m\}\-x\_\{s\}\), which is \([79](https://arxiv.org/html/2607.27431#A11.E79)\)\. Combining the two factors and using the product law \([24](https://arxiv.org/html/2607.27431#A3.E24)\) givesΨm→t​\(Ψs→m​\(R,x\)\)=\(R​D​\(s,m\)​D​\(m,t\),x\+p​\(s,m\)\+p​\(m,t\)\)=\(R​D​\(s,t\),x\+p​\(s,t\)\)=Ψs→t​\(R,x\)\\Psi\_\{m\\to t\}\(\\Psi\_\{s\\to m\}\(R,x\)\)=\(R\\,D\(s,m\)D\(m,t\),\\,x\+p\(s,m\)\+p\(m,t\)\)=\(R\\,D\(s,t\),\\,x\+p\(s,t\)\)=\\Psi\_\{s\\to t\}\(R,x\), i\.e\. \([77](https://arxiv.org/html/2607.27431#A11.E77)\)\. ∎

### K\.3The semigroup\-consistency loss

Proposition[9](https://arxiv.org/html/2607.27431#Thmproposition9)characterizes the correct average\-velocity field*without*any time derivative: the field is consistent iff its one\-step prediction on\[s,t\]\[s,t\]equals the two\-step composition through any intermediatem∈\[s,t\]m\\in\[s,t\]\. We turn this into a regression objective\. Letfθf\_\{\\theta\}outputAθavg​\(⋅\)A^\{\\mathrm\{avg\}\}\_\{\\theta\}\(\\cdot\)andvθavg​\(⋅\)v^\{\\mathrm\{avg\}\}\_\{\\theta\}\(\\cdot\)\(in the endpoint parameterization of Section[I](https://arxiv.org/html/2607.27431#A9),Aθavg\(s,t,Rt\)=1tlog\(\(R^0θ\)⊤Rt\)∨A^\{\\mathrm\{avg\}\}\_\{\\theta\}\(s,t,R\_\{t\}\)=\\tfrac\{1\}\{t\}\\log\(\(\\hat\{R\}\_\{0\}^\{\\theta\}\)^\{\\top\}R\_\{t\}\)^\{\\vee\}\)\. Given a triples<m<ts<m<t, form the intermediate state by stepping back the near segment\[m,t\]\[m,t\],

Rm=Rt​exp⁡\(−\(t−m\)​Aθavg​\(m,t,Rt,xt\)∧\),xm=xt−\(t−m\)​vθavg​\(m,t,xt\),\\displaystyle R\_\{m\}=R\_\{t\}\\exp\\\!\\big\(\-\(t\-m\)A^\{\\mathrm\{avg\}\}\_\{\\theta\}\(m,t,R\_\{t\},x\_\{t\}\)^\{\\wedge\}\\big\),\\qquad x\_\{m\}=x\_\{t\}\-\(t\-m\)\\,v^\{\\mathrm\{avg\}\}\_\{\\theta\}\(m,t,x\_\{t\}\),\(80\)and define the composed \(two\-step\) targets by \([78](https://arxiv.org/html/2607.27431#A11.E78)\)–\([79](https://arxiv.org/html/2607.27431#A11.E79)\):

Atgts→t\\displaystyle A^\{s\\to t\}\_\{\\mathrm\{tgt\}\}:=log\(exp⁡\(\(m−s\)​Aθavg​\(s,m,Rm,xm\)∧\)⏟D​\(s,m\)exp⁡\(\(t−m\)​Aθavg​\(m,t,Rt,xt\)∧\)⏟D​\(m,t\)\)∨,\\displaystyle:=\\log\\\!\\Big\(\\underbrace\{\\exp\\\!\\big\(\(m\-s\)A^\{\\mathrm\{avg\}\}\_\{\\theta\}\(s,m,R\_\{m\},x\_\{m\}\)^\{\\wedge\}\\big\)\}\_\{D\(s,m\)\}\\;\\underbrace\{\\exp\\\!\\big\(\(t\-m\)A^\{\\mathrm\{avg\}\}\_\{\\theta\}\(m,t,R\_\{t\},x\_\{t\}\)^\{\\wedge\}\\big\)\}\_\{D\(m,t\)\}\\Big\)^\{\\\!\\vee\},\(81\)ptgts→t\\displaystyle p^\{s\\to t\}\_\{\\mathrm\{tgt\}\}:=\(m−s\)​vθavg​\(s,m,xm\)\+\(t−m\)​vθavg​\(m,t,xt\)\.\\displaystyle:=\(m\-s\)\\,v^\{\\mathrm\{avg\}\}\_\{\\theta\}\(s,m,x\_\{m\}\)\+\(t\-m\)\\,v^\{\\mathrm\{avg\}\}\_\{\\theta\}\(m,t,x\_\{t\}\)\.\(82\)The model regresses its*direct*one\-step prediction on\[s,t\]\[s,t\]onto these stop\-gradient targets:

ℒrotsg\\displaystyle\\mathcal\{L\}^\{\\mathrm\{sg\}\}\_\{\\mathrm\{rot\}\}=𝔼s<m<t​‖\(t−s\)​Aθavg​\(s,t,Rt,xt\)−sg​\(Atgts→t\)‖22,\\displaystyle=\\mathbb\{E\}\_\{s<m<t\}\\,\\Big\\\|\(t\-s\)A^\{\\mathrm\{avg\}\}\_\{\\theta\}\(s,t,R\_\{t\},x\_\{t\}\)\-\\mathrm\{sg\}\\big\(A^\{s\\to t\}\_\{\\mathrm\{tgt\}\}\\big\)\\Big\\\|\_\{2\}^\{2\},\(83\)ℒtranssg\\displaystyle\\mathcal\{L\}^\{\\mathrm\{sg\}\}\_\{\\mathrm\{trans\}\}=𝔼s<m<t​‖\(t−s\)​vθavg​\(s,t,xt\)−sg​\(ptgts→t\)‖22\.\\displaystyle=\\mathbb\{E\}\_\{s<m<t\}\\,\\Big\\\|\(t\-s\)v^\{\\mathrm\{avg\}\}\_\{\\theta\}\(s,t,x\_\{t\}\)\-\\mathrm\{sg\}\\big\(p^\{s\\to t\}\_\{\\mathrm\{tgt\}\}\\big\)\\Big\\\|\_\{2\}^\{2\}\.\(84\)The semigroup constraint alone admits trivial \(collapsed\) minimizers; it must be anchored by thes→ts\\to tboundary, where \([76](https://arxiv.org/html/2607.27431#A11.E76)\) degenerates to the instantaneous velocity\. This boundary term is exactly flow matching:

ℒrotbd=𝔼t​‖Aθavg​\(t,t,Rt,xt\)−ωt‖22,ℒtransbd=𝔼t​‖vθavg​\(t,t,xt\)−vt‖22\.\\displaystyle\\mathcal\{L\}^\{\\mathrm\{bd\}\}\_\{\\mathrm\{rot\}\}=\\mathbb\{E\}\_\{t\}\\big\\\|A^\{\\mathrm\{avg\}\}\_\{\\theta\}\(t,t,R\_\{t\},x\_\{t\}\)\-\\omega\_\{t\}\\big\\\|\_\{2\}^\{2\},\\qquad\\mathcal\{L\}^\{\\mathrm\{bd\}\}\_\{\\mathrm\{trans\}\}=\\mathbb\{E\}\_\{t\}\\big\\\|v^\{\\mathrm\{avg\}\}\_\{\\theta\}\(t,t,x\_\{t\}\)\-v\_\{t\}\\big\\\|\_\{2\}^\{2\}\.\(85\)The total objective is

ℒsg​\-​MF=ℒrotbd\+ℒtransbd⏟flow matching \(anchor\)\+ℒrotsg\+ℒtranssg⏟semigroup consistency\.\\displaystyle\\mathcal\{L\}^\{\\mathrm\{sg\\text\{\-\}MF\}\}=\\underbrace\{\\mathcal\{L\}^\{\\mathrm\{bd\}\}\_\{\\mathrm\{rot\}\}\+\\mathcal\{L\}^\{\\mathrm\{bd\}\}\_\{\\mathrm\{trans\}\}\}\_\{\\text\{flow matching \(anchor\)\}\}\+\\underbrace\{\\mathcal\{L\}^\{\\mathrm\{sg\}\}\_\{\\mathrm\{rot\}\}\+\\mathcal\{L\}^\{\\mathrm\{sg\}\}\_\{\\mathrm\{trans\}\}\}\_\{\\text\{semigroup consistency\}\}\.\(86\)Every term uses only forward evaluations offθf\_\{\\theta\}together withexp/log\\exp/\\logonSO​\(3\)\\mathrm\{SO\}\(3\)and addition onℝ3\\mathbb\{R\}^\{3\}: there is no Jacobian–vector product and no explicit right JacobianJJ\.

### K\.4Consistency: zero loss implies exact reconstruction

The boundary term fixes the instantaneous limit and the semigroup term propagates it to all intervals; together they pin down the unique correct flow map\. The following is the finite \(JVP\-free\) analogue of Proposition[5](https://arxiv.org/html/2607.27431#Thmproposition5)\.

###### Proposition 10\(Uniqueness of the semigroup minimizer\)\.

SupposeAθavgA^\{\\mathrm\{avg\}\}\_\{\\theta\}is continuous and, along the interpolation path, satisfies the boundary conditionAθavg​\(t,t,Rt,xt\)=ωtA^\{\\mathrm\{avg\}\}\_\{\\theta\}\(t,t,R\_\{t\},x\_\{t\}\)=\\omega\_\{t\}for alltttogether with the rotation semigroup identity

exp⁡\(\(Aθs→m\)∧\)​exp⁡\(\(Aθm→t\)∧\)=exp⁡\(\(Aθs→t\)∧\),∀s≤m≤t,\\exp\\\!\\big\(\(A^\{s\\to m\}\_\{\\theta\}\)^\{\\wedge\}\\big\)\\,\\exp\\\!\\big\(\(A^\{m\\to t\}\_\{\\theta\}\)^\{\\wedge\}\\big\)=\\exp\\\!\\big\(\(A^\{s\\to t\}\_\{\\theta\}\)^\{\\wedge\}\\big\),\\qquad\\forall\\,s\\leq m\\leq t,whereAθa→b:=\(b−a\)​Aθavg​\(a,b,Rb,xb\)A^\{a\\to b\}\_\{\\theta\}:=\(b\-a\)A^\{\\mathrm\{avg\}\}\_\{\\theta\}\(a,b,R\_\{b\},x\_\{b\}\)\. Thenexp⁡\(\(Aθs→t\)∧\)=Rs⊤​Rt\\exp\(\(A^\{s\\to t\}\_\{\\theta\}\)^\{\\wedge\}\)=R\_\{s\}^\{\\top\}R\_\{t\}for alls≤ts\\leq t; in particular, withs=0,t=1s=0,t=1,R^0:=R1​exp⁡\(−\(Aθ0→1\)∧\)=R0\\hat\{R\}\_\{0\}:=R\_\{1\}\\exp\(\-\(A^\{0\\to 1\}\_\{\\theta\}\)^\{\\wedge\}\)=R\_\{0\}\. The analogous statement holds for translation withvθavg​\(t,t\)=vtv^\{\\mathrm\{avg\}\}\_\{\\theta\}\(t,t\)=v\_\{t\}\.

###### Proof\.

Fixttand setG​\(s\):=exp⁡\(\(Aθs→t\)∧\)∈SO​\(3\)G\(s\):=\\exp\(\(A^\{s\\to t\}\_\{\\theta\}\)^\{\\wedge\}\)\\in\\mathrm\{SO\}\(3\), soG​\(t\)=IG\(t\)=I\. Forh\>0h\>0the semigroup identity givesG​\(s\)=exp⁡\(\(Aθs→s\+h\)∧\)​G​\(s\+h\)G\(s\)=\\exp\(\(A^\{s\\to s\+h\}\_\{\\theta\}\)^\{\\wedge\}\)\\,G\(s\+h\), hence

G​\(s\+h\)=exp⁡\(−\(Aθs→s\+h\)∧\)​G​\(s\)\.G\(s\+h\)=\\exp\\\!\\big\(\-\(A^\{s\\to s\+h\}\_\{\\theta\}\)^\{\\wedge\}\\big\)\\,G\(s\)\.By the boundary condition and continuity,Aθs→s\+h=h​Aθavg​\(s,s\+h\)=h​ωs\+o​\(h\)A^\{s\\to s\+h\}\_\{\\theta\}=h\\,A^\{\\mathrm\{avg\}\}\_\{\\theta\}\(s,s\+h\)=h\\,\\omega\_\{s\}\+o\(h\), soexp⁡\(−\(Aθs→s\+h\)∧\)=I−h​ωs∧\+o​\(h\)\\exp\(\-\(A^\{s\\to s\+h\}\_\{\\theta\}\)^\{\\wedge\}\)=I\-h\\,\\omega\_\{s\}^\{\\wedge\}\+o\(h\)and therefore

dd​s​G​\(s\)=−ωs∧​G​\(s\),G​\(t\)=I\.\\frac\{d\}\{ds\}G\(s\)=\-\\,\\omega\_\{s\}^\{\\wedge\}\\,G\(s\),\\qquad G\(t\)=I\.On the other handD​\(s,t\)=Rs⊤​RtD\(s,t\)=R\_\{s\}^\{\\top\}R\_\{t\}obeys, usingR˙s=Rs​ωs∧\\dot\{R\}\_\{s\}=R\_\{s\}\\omega\_\{s\}^\{\\wedge\}and skew\-symmetry,

dd​s​D​\(s,t\)=\(∂sRs⊤\)​Rt=−ωs∧​Rs⊤​Rt=−ωs∧​D​\(s,t\),D​\(t,t\)=I\.\\frac\{d\}\{ds\}D\(s,t\)=\\big\(\\partial\_\{s\}R\_\{s\}^\{\\top\}\\big\)R\_\{t\}=\-\\,\\omega\_\{s\}^\{\\wedge\}R\_\{s\}^\{\\top\}R\_\{t\}=\-\\,\\omega\_\{s\}^\{\\wedge\}D\(s,t\),\\qquad D\(t,t\)=I\.GGandD​\(⋅,t\)D\(\\cdot,t\)solve the same linear ODE with the same terminal condition, so by uniquenessG​\(s\)=D​\(s,t\)=Rs⊤​RtG\(s\)=D\(s,t\)=R\_\{s\}^\{\\top\}R\_\{t\}for alls≤ts\\leq t\. Takings=0,t=1s=0,t=1givesexp⁡\(\(Aθ0→1\)∧\)=R0⊤​R1\\exp\(\(A^\{0\\to 1\}\_\{\\theta\}\)^\{\\wedge\}\)=R\_\{0\}^\{\\top\}R\_\{1\}, i\.e\.R^0=R0\\hat\{R\}\_\{0\}=R\_\{0\}\. For translation, the additive cocyclepθ​\(s,t\)p\_\{\\theta\}\(s,t\)withpθ​\(t,t\)p\_\{\\theta\}\(t,t\)\-derivativevθavg​\(t,t\)=vtv^\{\\mathrm\{avg\}\}\_\{\\theta\}\(t,t\)=v\_\{t\}integrates topθ​\(s,t\)=xt−xsp\_\{\\theta\}\(s,t\)=x\_\{t\}\-x\_\{s\}, givingx^0=x0\\hat\{x\}\_\{0\}=x\_\{0\}\. ∎

Proposition[10](https://arxiv.org/html/2607.27431#Thmproposition10)shows \([86](https://arxiv.org/html/2607.27431#A11.E86)\) is a valid training objective: its global minimizer is the exact average\-velocity field, and the two ingredients are both necessary—the boundary term supplies the instantaneous data velocity, and the semigroup term is the \(curvature\-exact\) propagation rule that extends it to every interval\.

### K\.5Relation to the differential identity and toα\\alpha\-Flow

Connection to Riemannian MeanFlow semigroup\.The semigroup objective \([86](https://arxiv.org/html/2607.27431#A11.E86)\) can be viewed as a variant of the semigroup formulation proposed in Riemannian MeanFlowWooet al\.\[[2026](https://arxiv.org/html/2607.27431#bib.bib23)\]: their construction is based on the endpoint \(geodesic\) loss, while ours is derived from and anchored by the flow\-matching boundary condition \([85](https://arxiv.org/html/2607.27431#A11.E85)\)\. We provide code for this semigroup objective in our repository; however, in our experiments we found that with a limited training budget \(e\.g\., 10–20k steps\) semigroup training yields weaker few\-step improvements than our JVP\-based MeanFlow loss, and therefore we do not use the semigroup loss in our final training recipe\. We include it here for theoretical completeness\.

## Appendix LQuaternion Formulation of SE\(3\)\-MeanFlow

The SE\(3\)\-MeanFlow objectives admit an equivalent formulation with unit quaternions on𝕊3\\mathbb\{S\}^\{3\}, following the rotation parametrisation of ReQFlowYueet al\.\[[2025](https://arxiv.org/html/2607.27431#bib.bib22)\]\. We derive it here for completeness; an implementation is provided in our repository alongside the rotation\-matrix formulation used throughout the paper\.

### L\.1Quaternion algebra and conventions

We follow the standard conventions; seeHanson \[[2005](https://arxiv.org/html/2607.27431#bib.bib42)\]for a fuller treatment\. A quaternion is a pair

q=\(q0,𝐪\)∈ℝ×ℝ3,‖q‖=1,\\displaystyle q=\(q\_\{0\},\\mathbf\{q\}\)\\in\\mathbb\{R\}\\times\\mathbb\{R\}^\{3\},\\qquad\\\|q\\\|=1,with multiplication and inverse

q⊗p=\(q0​p0−𝐪⊤​𝐩,q0​𝐩\+p0​𝐪\+𝐪×𝐩\),q−1=\(q0,−𝐪\)\.\\displaystyle q\\otimes p=\\big\(q\_\{0\}p\_\{0\}\-\\mathbf\{q\}^\{\\top\}\\mathbf\{p\},\\;q\_\{0\}\\mathbf\{p\}\+p\_\{0\}\\mathbf\{q\}\+\\mathbf\{q\}\\times\\mathbf\{p\}\\big\),\\qquad q^\{\-1\}=\(q\_\{0\},\-\\mathbf\{q\}\)\.The unit quaternions form the manifold𝕊3\\mathbb\{S\}^\{3\}, a double cover ofSO​\(3\)\\mathrm\{SO\}\(3\): the quaternionsqqand−q\-qrepresent the same rotation

R​\(q\)=\[1−2​\(y2\+z2\)2​\(x​y−w​z\)2​\(x​z\+w​y\)2​\(x​y\+w​z\)1−2​\(x2\+z2\)2​\(y​z−w​x\)2​\(x​z−w​y\)2​\(y​z\+w​x\)1−2​\(x2\+y2\)\],q=\(w,x,y,z\)\.\\displaystyle R\(q\)=\\begin\{bmatrix\}1\-2\(y^\{2\}\+z^\{2\}\)&2\(xy\-wz\)&2\(xz\+wy\)\\\\ 2\(xy\+wz\)&1\-2\(x^\{2\}\+z^\{2\}\)&2\(yz\-wx\)\\\\ 2\(xz\-wy\)&2\(yz\+wx\)&1\-2\(x^\{2\}\+y^\{2\}\)\\end\{bmatrix\},\\qquad q=\(w,x,y,z\)\.RecoveringqqfromRRis standard \(trace\-based\) up to this sign; we fix it by enforcingw≥0w\\geq 0\.

##### Exponential and logarithm\.

We defineExp:ℝ3→𝕊3\\mathrm\{Exp\}:\\mathbb\{R\}^\{3\}\\to\\mathbb\{S\}^\{3\}with the half\-angle absorbed,

Exp​\(ω\):=\(cos⁡‖ω‖2,sin⁡‖ω‖2​ω‖ω‖\),Exp​\(0\)=\(1,𝟎\),\\displaystyle\\mathrm\{Exp\}\(\\omega\):=\\Big\(\\cos\\tfrac\{\\\|\\omega\\\|\}\{2\},\\;\\sin\\tfrac\{\\\|\\omega\\\|\}\{2\}\\,\\tfrac\{\\omega\}\{\\\|\\omega\\\|\}\\Big\),\\qquad\\mathrm\{Exp\}\(0\)=\(1,\\mathbf\{0\}\),\(87\)and letLog\\mathrm\{Log\}be its inverse on the hemispherew≥0w\\geq 0, so‖Log​\(q\)‖≤π\\\|\\mathrm\{Log\}\(q\)\\\|\\leq\\pi\. With this convention the covering map is compatible with the matrix exponential,

R​\(Exp​\(ω\)\)=exp⁡\(ω∧\),\\displaystyle R\\big\(\\mathrm\{Exp\}\(\\omega\)\\big\)=\\exp\(\\omega^\{\\wedge\}\),\(88\)soω\\omegacarries the same axis\-angle meaning as in the rest of the paper\. \(The ReQFlow subsection above writes the same object asexp⁡\(12​ω\)\\exp\(\\tfrac\{1\}\{2\}\\omega\); the factor of two is bookkeeping in where the half\-angle is placed\.\) The tangent space isTq​𝕊3=\{v∈ℝ4:q⊤​v=0\}T\_\{q\}\\mathbb\{S\}^\{3\}=\\\{v\\in\\mathbb\{R\}^\{4\}:q^\{\\top\}v=0\\\}\.

##### Kinematics\.

Letω​\(t\)∈ℝ3\\omega\(t\)\\in\\mathbb\{R\}^\{3\}be the*body\-frame*angular velocity, i\.e\. the same quantity as in \([1](https://arxiv.org/html/2607.27431#S2.E1)\)\. The left\-trivialised kinematics are

q˙t=12​qt⊗\(0,ω​\(t\)\),\\displaystyle\\dot\{q\}\_\{t\}=\\tfrac\{1\}\{2\}\\,q\_\{t\}\\otimes\(0,\\omega\(t\)\),\(89\)whose solution is the time\-ordered exponential in𝕊3\\mathbb\{S\}^\{3\},

qt=qs⊗𝒯​Exp​\(∫stω​\(τ\)​𝑑τ\),andqt=qs⊗Exp​\(\(t−s\)​ω\)​if​ω​is constant\.\\displaystyle q\_\{t\}=q\_\{s\}\\otimes\\mathcal\{T\}\\mathrm\{Exp\}\\Big\(\\int\_\{s\}^\{t\}\\omega\(\\tau\)\\,d\\tau\\Big\),\\qquad\\text\{and\}\\qquad q\_\{t\}=q\_\{s\}\\otimes\\mathrm\{Exp\}\\big\(\(t\-s\)\\,\\omega\\big\)\\ \\text\{ if \}\\omega\\text\{ is constant\}\.The half\-angle in \([87](https://arxiv.org/html/2607.27431#A12.E87)\) absorbs the factor12\\tfrac\{1\}\{2\}of \([89](https://arxiv.org/html/2607.27431#A12.E89)\), so no stray factors of two appear below\.

##### Interpolation and data velocity\.

Mirroring Appendix[D](https://arxiv.org/html/2607.27431#A4), for a data rotationq0q\_\{0\}and a prior rotationq1q\_\{1\}\(sign\-aligned so thatq0⊤​q1≥0q\_\{0\}^\{\\top\}q\_\{1\}\\geq 0, which selects the shorter arc\) the conditional path and its constant body angular velocity are

qt=q0⊗Exp​\(t​ωt\),ωt:=Log​\(q0−1⊗q1\)∈ℝ3\.\\displaystyle q\_\{t\}=q\_\{0\}\\otimes\\mathrm\{Exp\}\\big\(t\\,\\omega\_\{t\}\\big\),\\qquad\\omega\_\{t\}:=\\mathrm\{Log\}\\big\(q\_\{0\}^\{\-1\}\\otimes q\_\{1\}\\big\)\\in\\mathbb\{R\}^\{3\}\.\(90\)

### L\.2Average velocity and the MeanFlow identity

Exactly as in \([30](https://arxiv.org/html/2607.27431#A6.E30)\), we define the average angular velocity over\[s,t\]\[s,t\]through the time\-ordered exponential\. Writing the state aszt=\(xt,qt\)∈ℝ3×𝕊3z\_\{t\}=\(x\_\{t\},q\_\{t\}\)\\in\\mathbb\{R\}^\{3\}\\times\\mathbb\{S\}^\{3\}to make the conditioning explicit,

Exp​\(\(t−s\)​ωavg​\(s,t,xt,qt\)\):=𝒯​Exp​\(∫stω​\(τ,qτ,xτ\)​𝑑τ\)=qs−1⊗qt\.\\displaystyle\\mathrm\{Exp\}\\big\(\(t\-s\)\\,\\omega^\{\\mathrm\{avg\}\}\(s,t,x\_\{t\},q\_\{t\}\)\\big\):=\\mathcal\{T\}\\mathrm\{Exp\}\\Big\(\\int\_\{s\}^\{t\}\\omega\(\\tau,q\_\{\\tau\},x\_\{\\tau\}\)\\,d\\tau\\Big\)=q\_\{s\}^\{\-1\}\\otimes q\_\{t\}\.\(91\)
###### Proposition 11\(Quaternion MeanFlow identity\)\.

Differentiating \([91](https://arxiv.org/html/2607.27431#A12.E91)\) with respect tottyields

J​\(\(t−s\)​ωavg\)​\(ωavg​\(s,t,xt,qt\)\+\(t−s\)​dd​t​ωavg​\(s,t,xt,qt\)\)=ωt,\\displaystyle J\\big\(\(t\-s\)\\,\\omega^\{\\mathrm\{avg\}\}\\big\)\\Big\(\\omega^\{\\mathrm\{avg\}\}\(s,t,x\_\{t\},q\_\{t\}\)\+\(t\-s\)\\tfrac\{d\}\{dt\}\\omega^\{\\mathrm\{avg\}\}\(s,t,x\_\{t\},q\_\{t\}\)\\Big\)=\\omega\_\{t\},\(92\)withJJthe*same*right Jacobian \([10](https://arxiv.org/html/2607.27431#S3.E10)\) as in the rotation\-matrix formulation, and

dd​t​ωavg=∂∂t​ωavg\+⟨∇qωavg,q˙t⟩\+⟨∇xωavg,x˙t⟩\.\\displaystyle\\frac\{d\}\{dt\}\\omega^\{\\mathrm\{avg\}\}=\\frac\{\\partial\}\{\\partial t\}\\omega^\{\\mathrm\{avg\}\}\+\\big\\langle\\nabla\_\{q\}\\omega^\{\\mathrm\{avg\}\},\\dot\{q\}\_\{t\}\\big\\rangle\+\\big\\langle\\nabla\_\{x\}\\omega^\{\\mathrm\{avg\}\},\\dot\{x\}\_\{t\}\\big\\rangle\.

###### Proof\.

The covering map𝕊3→SO​\(3\)\\mathbb\{S\}^\{3\}\\to\\mathrm\{SO\}\(3\)of \([88](https://arxiv.org/html/2607.27431#A12.E88)\) is a two\-to\-one Lie group homomorphism and a local diffeomorphism, and under the identification\(0,ω\)↔ω∧\(0,\\omega\)\\leftrightarrow\\omega^\{\\wedge\}it induces the identity on Lie algebras \(𝔰​𝔲​\(2\)≅𝔰​𝔬​\(3\)\\mathfrak\{su\}\(2\)\\cong\\mathfrak\{so\}\(3\)\)\. Applying it to \([91](https://arxiv.org/html/2607.27431#A12.E91)\) returns \([30](https://arxiv.org/html/2607.27431#A6.E30)\) verbatim, so every step of the proofs of Propositions[3](https://arxiv.org/html/2607.27431#Thmproposition3)and[4](https://arxiv.org/html/2607.27431#Thmproposition4)carries over unchanged\. ∎

### L\.3Parametrisation, loss and inference

As in Section[I](https://arxiv.org/html/2607.27431#A9), the network is an endpoint predictor returning\(q^0θ,x^0θ\)\(\\hat\{q\}\_\{0\}^\{\\theta\},\\hat\{x\}\_\{0\}^\{\\theta\}\), converted to the average\-velocity head by

ωθavg=1t​Log​\(\(q^0θ\)−1⊗qt\)⇔q^0θ=qt⊗Exp​\(−t​ωθavg\)\.\\displaystyle\\omega^\{\\mathrm\{avg\}\}\_\{\\theta\}=\\tfrac\{1\}\{t\}\\,\\mathrm\{Log\}\\big\(\(\\hat\{q\}\_\{0\}^\{\\theta\}\)^\{\-1\}\\otimes q\_\{t\}\\big\)\\iff\\hat\{q\}\_\{0\}^\{\\theta\}=q\_\{t\}\\otimes\\mathrm\{Exp\}\\big\(\-t\\,\\omega^\{\\mathrm\{avg\}\}\_\{\\theta\}\\big\)\.Regressing the left\-hand side of \([92](https://arxiv.org/html/2607.27431#A12.E92)\) onto the data velocityωt\\omega\_\{t\}of \([90](https://arxiv.org/html/2607.27431#A12.E90)\) gives the rotation loss

ℒrot𝕊3=𝔼s<t,q​\(z0,z1\)​\[‖J​\(\(t−s\)​ωθavg\)​\(ωθavg\+\(t−s\)​sg​\(dd​t​ωθavg\)\)−ωt‖22\],\\displaystyle\\mathcal\{L\}^\{\\mathbb\{S\}^\{3\}\}\_\{\\mathrm\{rot\}\}=\\mathbb\{E\}\_\{s<t,\\,q\(z\_\{0\},z\_\{1\}\)\}\\Big\[\\big\\\|\\,J\\big\(\(t\-s\)\\omega^\{\\mathrm\{avg\}\}\_\{\\theta\}\\big\)\\Big\(\\omega^\{\\mathrm\{avg\}\}\_\{\\theta\}\+\(t\-s\)\\,\\mathrm\{sg\}\\big\(\\tfrac\{d\}\{dt\}\\omega^\{\\mathrm\{avg\}\}\_\{\\theta\}\\big\)\\Big\)\-\\omega\_\{t\}\\,\\big\\\|\_\{2\}^\{2\}\\Big\],\(93\)the exact analogue of \([11](https://arxiv.org/html/2607.27431#S3.E11)\); the translation branch and theα\\alpha\-Flow variant of Section[J](https://arxiv.org/html/2607.27431#A10)are unchanged, since neither touches the rotation representation\. Inference steps backwards byqs←qt⊗Exp​\(−\(t−s\)​ωθavg\)q\_\{s\}\\leftarrow q\_\{t\}\\otimes\\mathrm\{Exp\}\\big\(\-\(t\-s\)\\,\\omega^\{\\mathrm\{avg\}\}\_\{\\theta\}\\big\)\.

### L\.4Practical note

Empirically the quaternion implementation consistently underperformed the rotation\-matrix one, and we use the latter for every result in this paper\. Two causes are consistent with what we observed\. First, the surrounding pipeline \(dataset, IPA trunk, auxiliary losses, evaluation\) operates on rotation matrices, so the quaternion path inserts repeated matrix↔\\,\\leftrightarrow\\,quaternion conversions whose error accumulates\. Second, and more specific to MeanFlow, the double coverq∼−qq\\sim\-qmeans the enforced sign conventionw≥0w\\geq 0can flip along a trajectory; the endpoint predictionq^0θ\\hat\{q\}\_\{0\}^\{\\theta\}then jumps discontinuously, and the forward\-mode derivativedd​t​ωθavg\\tfrac\{d\}\{dt\}\\omega^\{\\mathrm\{avg\}\}\_\{\\theta\}in \([93](https://arxiv.org/html/2607.27431#A12.E93)\) is corrupted at exactly those steps\. Rotation matrices have no such ambiguity\. Together with Remark[15](https://arxiv.org/html/2607.27431#Thmremark15)—the quaternion target is not analytically simpler—this left no reason to prefer the quaternion branch, and we report no quaternion\-based results\.

## Appendix MTraining Curriculum: Pre\-training and Post\-training

Our final model is obtained in three phases: a two\-stage*pre\-training*that builds a strong multi\-step backbone generator, followed by a*post\-training*\(rectification\) stage that sharpens few\-step generation\. All phases share the same network \(Section[N](https://arxiv.org/html/2607.27431#A14)\) and the same decoupledSE​\(3\)=SO​\(3\)×ℝ3\\mathrm\{SE\}\(3\)=\\mathrm\{SO\}\(3\)\\times\\mathbb\{R\}^\{3\}objective; they differ only in which consistency target is used and in how the data–noise pairs are coupled\.

### M\.1Stage 1: JVP\-freeα\\alpha\-Flow warm\-up

Training the endpoint\-parameterized MeanFlow target directly from initialization is fragile: as discussed in Sections[I](https://arxiv.org/html/2607.27431#A9)and[J](https://arxiv.org/html/2607.27431#A10), the differential rotation target amplifies head error as𝒪​\(ε/t2\)\\mathcal\{O\}\(\\varepsilon/t^\{2\}\)and requires a forward\-mode derivative \(JVP\) through the IPA trunk\. We therefore warm up with the JVP\-freeα\\alpha\-Flow objective of Section[J](https://arxiv.org/html/2607.27431#A10), annealing the consistency\-step ratioα\\alphafrom a capαmax=1\\alpha\_\{\\max\}=1toward a floorαmin=0\.1\\alpha\_\{\\min\}=0\.1along a logistic schedule:

α​\(k\)=\{αmax,k≤ks,αmin\+\(αmax−αmin\)​σγ​\(τk\),ks<k<ke,αmin,k≥ke,σγ​\(τ\)=11\+eγ​\(τ−12\),τk=k−kske−ks,\\alpha\(k\)=\\begin\{cases\}\\alpha\_\{\\max\},&k\\leq k\_\{s\},\\\\\[2\.0pt\] \\alpha\_\{\\min\}\+\(\\alpha\_\{\\max\}\-\\alpha\_\{\\min\}\)\\,\\sigma\_\{\\gamma\}\(\\tau\_\{k\}\),&k\_\{s\}<k<k\_\{e\},\\\\\[2\.0pt\] \\alpha\_\{\\min\},&k\\geq k\_\{e\},\\end\{cases\}\\qquad\\sigma\_\{\\gamma\}\(\\tau\)=\\frac\{1\}\{1\+e^\{\\,\\gamma\(\\tau\-\\tfrac\{1\}\{2\}\)\}\},\\quad\\tau\_\{k\}=\\frac\{k\-k\_\{s\}\}\{k\_\{e\}\-k\_\{s\}\},\(94\)wherekkis the optimizer step andτk∈\[0,1\]\\tau\_\{k\}\\in\[0,1\]is the normalized progress through the anneal window\. Hereksk\_\{s\}is the hold phase—the ratio is pinned at the cap for the firstksk\_\{s\}steps—kek\_\{e\}is the step by which the floor is reached, andγ\\gammasets the steepness of the logistic transition, centered at the window midpointks\+12​\(ke−ks\)k\_\{s\}\+\\tfrac\{1\}\{2\}\(k\_\{e\}\-k\_\{s\}\)\. We useks=2000k\_\{s\}\\\!=\\\!2000,ke=150​kk\_\{e\}\\\!=\\\!150\\mathrm\{k\}, andγ=8\\gamma\\\!=\\\!8\.

By Remark[8](https://arxiv.org/html/2607.27431#Thmremark8),α=1\\alpha=1reduces the objective exactly to flow matching \([25](https://arxiv.org/html/2607.27431#A4.E25)\), so during the hold phase training is anchored to the data velocity and cannot collapse; decreasingα\\alphathereafter progressively injects the few\-step average\-velocity consistency that underlies fast sampling\. The anneal was scheduled over150​k150\\mathrm\{k\}steps, but we stopped Stage 1 early at100​k100\\mathrm\{k\}—whereα≈0\.29\\alpha\\\!\\approx\\\!0\.29—because validation designability had already plateaued;α\\alphatherefore never reaches its floorαmin=0\.1\\alpha\_\{\\min\}=0\.1during pre\-training\. This checkpoint initializes Stage 2\. Full settings are in Table[7](https://arxiv.org/html/2607.27431#A13.T7)\.

##### Self\-conditioning\.

In all phases we use50%50\\%self\-conditioning, following FrameFlowYimet al\.\[[2023a](https://arxiv.org/html/2607.27431#bib.bib37)\]/ReQFlowYueet al\.\[[2025](https://arxiv.org/html/2607.27431#bib.bib22)\]\. With probability0\.50\.5a detached preliminary forward pass produces an endpoint predictionx^0θ\\hat\{x\}\_\{0\}^\{\\theta\}whose pairwiseCα\\mathrm\{C\}\_\{\\alpha\}distogram is appended as an extra edge feature to a second, gradient\-carrying forward pass; the remaining half of each batch passes zeros in that slot\. At inference the previous step’s endpoint prediction is used as the self\-conditioning input \(Algorithm[4](https://arxiv.org/html/2607.27431#alg4)\)\. This improves sample quality at no additional inference cost\.

##### Global\-OT coupling\.

Both pre\-training stages instantiate the mini\-batch OT coupling of Appendix[D\.3](https://arxiv.org/html/2607.27431#A4.SS3), but solve a single plan over the*entire*global batch rather than one plan per rank\. At each step atorch\.distributed\.all\_gatherassembles allM=∑r=1WBrM=\\sum\_\{r=1\}^\{W\}B\_\{r\}backbones across theWWranks, and the independent data–noise pairing is replaced by the assignment

π⋆=arg⁡minπ∈𝒮M​∑k=1Mc​\(z0k,z1π​\(k\)\),\\displaystyle\\pi^\{\\star\}=\\arg\\min\_\{\\pi\\in\\mathcal\{S\}\_\{M\}\}\\sum\_\{k=1\}^\{M\}c\\big\(z\_\{0\}^\{k\},\\,z\_\{1\}^\{\\pi\(k\)\}\\big\),\(95\)withccthe decoupledSE​\(3\)N\\mathrm\{SE\}\(3\)^\{N\}transport cost of Appendix[D\.3](https://arxiv.org/html/2607.27431#A4.SS3)and\(λR,λx\)=\(0\.5,0\.5\)\(\\lambda\_\{R\},\\lambda\_\{x\}\)=\(0\.5,0\.5\)\. The exact Hungarian assignment is solved on rank 0 with POTFlamaryet al\.\[[2021](https://arxiv.org/html/2607.27431#bib.bib43)\]andπ⋆\\pi^\{\\star\}broadcast back; theO​\(M3\)O\(M^\{3\}\)solve is negligible against a forward/backward pass at these batch sizes\. Solving \([95](https://arxiv.org/html/2607.27431#A13.E95)\) once over allMMbackbones uses the full pairing information, whereasWWper\-rank plans search a block\-diagonal subset of𝒮M\\mathcal\{S\}\_\{M\}and under\-use it by a factorWW; empirically the global coupling roughly halves the early\-plateau time on SCOPe\. Since mini\-batch OT is a biased estimate of the true coupling whose bias decreases with batch size, the global plan yields a straighter induced path thanWWindependent per\-rank plans at the same per\-GPU memory cost\.

Post\-training disables it \(its self\-reflow pairs are already matched; Section[M\.3](https://arxiv.org/html/2607.27431#A13.SS3)\)\.

### M\.2Stage 2: endpoint\+\+MeanFlow objective

Starting from the Stage\-1 checkpoint, we continue with the endpoint\-anchored small\-ttMeanFlow objective for a maximum budget of10​k10\\mathrm\{k\}steps, and select the released checkpoint at6\.1​k6\.1\\mathrm\{k\}steps by validation designability\. We retain the50%50\\%self\-conditioning and the global\-OT coupling of Stage 1; the only changes are the switch to the differential MeanFlow loss and the added endpoint anchor\. The loss combines three terms\. The endpoint distance is defiend as:

ℒend:=∑i=1N\[DSO​\(3\)2\(Rθ,0i,R0i\)\+∥xθ,0i−x0i∥22\],DSO​\(3\)\(R,Q\):=∥Log\(R⊤Q\)∨∥2=arccos\(tr⁡\(R⊤​Q\)−12\),\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{end\}\}:=\\sum\_\{i=1\}^\{N\}\\Big\[\\,D\_\{\\mathrm\{SO\}\(3\)\}^\{2\}\\\!\\big\(R^\{i\}\_\{\\theta,0\},\\,R^\{i\}\_\{0\}\\big\)\+\\big\\\|x^\{i\}\_\{\\theta,0\}\-x^\{i\}\_\{0\}\\big\\\|\_\{2\}^\{2\}\\,\\Big\],\\qquad D\_\{\\mathrm\{SO\}\(3\)\}\(R,Q\):=\\big\\\|\\operatorname\{Log\}\(R^\{\\top\}Q\)^\{\\vee\}\\big\\\|\_\{2\}=\\arccos\\\!\\Big\(\\tfrac\{\\operatorname\{tr\}\(R^\{\\top\}Q\)\-1\}\{2\}\\Big\),\(96\)
##### Loss normalization and stability\.

Our optional loss\-magnitude normalization \(our\_loss\_norm\) is*disabled*in all phases\. For stability we instead rely on three mechanisms: a global gradient\-norm clip at1\.01\.0, per\-residue rotation\- and translation\-loss clamps \(5050and55\), and the endpoint near\-data reweighting1/max\(t,tε\)21/\\max\(t,t\_\{\\varepsilon\}\)^\{2\}\(tε=0\.1t\_\{\\varepsilon\}=0\.1, Stage 2 onward\)\. The rotation and translation branches are weighted by\(wR,wx\)=\(1\.0,1\.0\)\(w\_\{R\},w\_\{x\}\)=\(1\.0,1\.0\)\.

Table 6:Per\-term loss magnitudes atλMF=0\.05\\lambda\_\{\\mathrm\{MF\}\}=0\.05\(median over2424SCOPe training batches;our\_loss\_normdisabled, right\-Jacobian inverse form\)\. The raw MeanFlow residual is∼33×\\sim\\\!33\\timeslarger than the endpoint term—owing to the near\-data reweighting1/max\(t,tε\)21/\\max\(t,t\_\{\\varepsilon\}\)^\{2\}—andλMF=0\.05\\lambda\_\{\\mathrm\{MF\}\}=0\.05brings the two to the same order \(≈1\.6:1\\approx 1\.6\{:\}1\)\.
##### Total objective\.

The Stage\-2 loss is the weighted sum

λend​ℒend\+λMF​\(ℒrotMF\+ℒtransMF\)\+λaux​1​\[t<0\.75\]​ℒaux,\\displaystyle\\lambda\_\{\\mathrm\{end\}\}\\,\\mathcal\{L\}\_\{\\mathrm\{end\}\}\+\\lambda\_\{\\mathrm\{MF\}\}\\,\\big\(\\mathcal\{L\}^\{\\text\{MF\}\}\_\{\\mathrm\{rot\}\}\+\\mathcal\{L\}^\{\\text\{MF\}\}\_\{\\mathrm\{trans\}\}\\big\)\+\\lambda\_\{\\mathrm\{aux\}\}\\,\\mathbf\{1\}\[t<0\.75\]\\,\\mathcal\{L\}\_\{\\mathrm\{aux\}\},\(97\)with\(λend,λMF,λaux\)=\(1\.0,0\.05,2\.0\)\(\\lambda\_\{\\mathrm\{end\}\},\\lambda\_\{\\mathrm\{MF\}\},\\lambda\_\{\\mathrm\{aux\}\}\)=\(1\.0,\\,0\.05,\\,2\.0\); the endpoint term anchors the model and stabilizes the MeanFlow loss, and all remaining settings are listed in Table[7](https://arxiv.org/html/2607.27431#A13.T7)\. HereλMF=0\.05\\lambda\_\{\\mathrm\{MF\}\}=0\.05is a scale normalizer rather than a tuned trade\-off\. Because the loss\-magnitude normalization is disabled in all phases \(see above\), the terms of \([97](https://arxiv.org/html/2607.27431#A13.E97)\) enter at their raw magnitudes, and these differ by more than an order of magnitude: the small\-ttMeanFlow residuals \([62](https://arxiv.org/html/2607.27431#A9.E62)\) carry the near\-data reweighting1/max\(t,tε\)21/\\max\(t,t\_\{\\varepsilon\}\)^\{2\}, which reachestε−2=100t\_\{\\varepsilon\}^\{\-2\}=100fort≤tε=0\.1t\\leq t\_\{\\varepsilon\}=0\.1, whereas the endpoint term is unweighted\. WithλMF=0\.05\\lambda\_\{\\mathrm\{MF\}\}=0\.05the weighted MeanFlow contribution is comparable in magnitude toλend​ℒend\\lambda\_\{\\mathrm\{end\}\}\\mathcal\{L\}\_\{\\mathrm\{end\}\}throughout training—a ratio of roughly1\.6:11\.6\{:\}1\(Table[6](https://arxiv.org/html/2607.27431#A13.T6)\)\. The small numerical value therefore reflects the scale of the raw residual, not a small role in the objective; equivalently, one may enable the loss normalization and setλMF=1\\lambda\_\{\\mathrm\{MF\}\}=1\.

Table 7:Training configuration for the two pre\-training stages\. Post\-training \(self\-reflow\) reuses the Stage\-2 column verbatim, changing only the data coupling: the OT coupling is replaced by the model’s own deterministic self\-reflow pairs, with OT disabled \(Section[M\.3](https://arxiv.org/html/2607.27431#A13.SS3)\)\.

### M\.3Post\-training: self\-reflow rectification

Few\-step designability is further improved by a rectification stage that replaces the random data–noise coupling with the model’s own deterministic transport, following the rectified\-flow strategy of ReQFlowYueet al\.\[[2025](https://arxiv.org/html/2607.27431#bib.bib22)\]\.

##### Self\-reflow dataset\.

Using the Stage\-2 model, for each residue lengthN∈\[60,128\]N\\in\[60,128\]we draw prior samplesz1∼ppriorz\_\{1\}\\sim p\_\{\\mathrm\{prior\}\}and integrate the model to obtain coupled pairs\(z1,z^0\)\(z\_\{1\},\\hat\{z\}\_\{0\}\)\(noise→\\togenerated backbone\)\. Each generated backbone is scored with the self\-consistency pipeline \(ProteinMPNN followed by ESMFold\), and we retain only*designable*samples \(scRMSD<2​Å\\mathrm\{scRMSD\}<2\\text\{\\AA \}\)\. We keep5050designable pairs per length—over\-sampling100100generations per length to meet the quota—yielding69×50=345069\\times 50=3450coupled pairs over the SCOPe length range\.

##### Rectification objective\.

We then continue training from the Stage\-2 checkpoint on these pairs using the*identical*Stage\-2 objective \([97](https://arxiv.org/html/2607.27431#A13.E97)\), but with the independent couplingq​\(z0,z1\)q\(z\_\{0\},z\_\{1\}\)replaced by the model\-induced deterministic coupling\(z1,z^0\)\(z\_\{1\},\\hat\{z\}\_\{0\}\)and with the OT re\-coupling disabled \(the pairs are already transport\-consistent\)\. Because the target transport is now \(approximately\) straight, the average\-velocity field it must match is closer to constant along each trajectory, which is precisely the regime in which few\-step MeanFlow sampling is exact\. Checkpoints are saved densely and selected on validation designability and secondary\-structure content\.

### M\.4Checkpoint selection

Rather than selecting the final model by a grid search over training hyper\-parameters, we retain the training checkpoint according to a fixed, pre\-specified rule on validation\-time generation quality\. During training we periodically sample unconditional backbones at two inference budgets,T=100T\{=\}100andT=20T\{=\}20integration steps, and monitor \(i\) the Cα\\alpha–Cα\\alphabond\-geometry validity \(ca​\_​ca\\mathrm\{ca\\\_ca\}\) and \(ii\) the secondary\-structure composition \(helix and strand fractions\)\. We keep the checkpoint that \(a\) attainsca​\_​ca\>0\.97\\mathrm\{ca\\\_ca\}\>0\.97at*both*T=100T\{=\}100andT=20T\{=\}20, so that backbones remain geometrically valid in both the high\- and low\-step regimes; \(b\) keeps the strand fraction near or above0\.20\.2, matching the reference distribution and guarding against strand collapse; and \(c\) among the checkpoints satisfying \(a\)–\(b\), maximizes the helix fraction\. This criterion was fixed before final evaluation and applied identically across all runs\.

## Appendix NArchitecture

Our network is the ReQFlowYueet al\.\[[2025](https://arxiv.org/html/2607.27431#bib.bib22)\]trunk left*unchanged*, augmented with only two additions required to turn a single\-time flow\-matching model into a two\-time MeanFlow model: \(i\) a shared two\-time embedding for the MeanFlow inputs\(t,s\)\(t,s\)\(§[N\.3](https://arxiv.org/html/2607.27431#A14.SS3)\), and \(ii\) a per\-block, zero\-initialised AdaLN\-Zero gate on each block’s rigid update \(§[N\.2](https://arxiv.org/html/2607.27431#A14.SS2)\)\. The trunk itself—the invariant point attention \(IPA\) stack ofJumperet al\.\[[2021](https://arxiv.org/html/2607.27431#bib.bib7)\]as instantiated by ReQFlow, comprising per block an IPA module, a two\-layer sequence transformer, node/edge transitions and a backbone frame update—is not modified\. The full model has∼16\.8\\sim\\\!16\.8M parameters, of which the AdaLN additions account for only0\.33%0\.33\\%\(∼55\\sim\\\!55k\); it is therefore∼26×\\sim\\\!26\\timessmaller than RMF \(Wooet al\.[2026](https://arxiv.org/html/2607.27431#bib.bib23),437437M\) while remaining checkpoint\-compatible with ReQFlow/QFlow \(Table[8](https://arxiv.org/html/2607.27431#A14.T8)\)\.

### N\.1Backbone trunk

The trunk follows the ReQFlow configuration: node and edge embedding sizes256256and128128, IPA hidden widthchidden=128c\_\{\\text\{hidden\}\}=128with88attention heads,88query/key points and1212value points, a44\-head /22\-layer sequence transformer, and66IPA blocks\. The input embedder is ReQFlow’s: a self\-conditioned pairwise distogram \(2222bins\), a relative\-position encoding \(relpos​\_​k=64\\mathrm\{relpos\}\\\_k=64\), and a diffuse\-mask embedding\. Because the trunk and embedder are inherited unchanged, our model loads directly from our pretrained checkpoint \(in stage 2\)\.

OursReQFlowRMFnode embed size256256768edge embed size128128384chiddenc\_\{\\text\{hidden\}\}12812848\# attention heads8816\# query/key points88—\# value points121216\# IPA blocks6616seq\. tfmr \(heads/layers\)4 / 24 / 212 / 3per\-block AdaLN gateyesnoyestotal parameters∼16\.8\\sim\\\!16\.8M∼16\.8\\sim\\\!16\.8M∼437\\sim\\\!437MTable 8:Trunk configuration\. “Ours” is the ReQFlowYueet al\.\[[2025](https://arxiv.org/html/2607.27431#bib.bib22)\]trunk with the only architectural addition being a per\-block AdaLN\-Zero gate \(\+0\.05\+0\.05M,0\.33%0\.33\\%of parameters\)\. Both are∼26×\\sim\\\!26\\timessmaller than RMFWooet al\.\[[2026](https://arxiv.org/html/2607.27431#bib.bib23)\]while recovering most of the structural fidelity\.
### N\.2Block\-wise AdaLN\-Zero gate

The sole structural change to the trunk is a per\-block gate on the rigid update\. Each IPA block emits a six\-vector backbone update \(three rotational, three translational\) that is composed onto the running frame\. We modulate this update with an AdaLN\-ZeroPeebles and Xie \[[2023](https://arxiv.org/html/2607.27431#bib.bib44)\]scale,

updateb←\(1\+γb​\(𝐜b\)\)⊙updateb,\\mathrm\{update\}\_\{b\}\\;\\leftarrow\\;\\big\(1\+\\gamma\_\{b\}\(\\mathbf\{c\}\_\{b\}\)\\big\)\\odot\\mathrm\{update\}\_\{b\},where the gateγb\\gamma\_\{b\}is produced by a per\-block linear head from the conditioning vector𝐜b=\[𝐡b∥𝐜\(t\)\]\\mathbf\{c\}\_\{b\}=\[\\,\\mathbf\{h\}\_\{b\}\\;\\\|\\;\\mathbf\{c\}^\{\(t\)\}\\,\], formed by concatenating \(i\) the block’s current*node features*𝐡b\\mathbf\{h\}\_\{b\}—which areSE​\(3\)\\mathrm\{SE\}\(3\)\-invariant and state\-aware \(they depend on the current\(Rt,xt\)\(R\_\{t\},x\_\{t\}\)through IPA\)—and \(ii\) a joint time/interval embedding𝐜\(t\)\\mathbf\{c\}^\{\(t\)\}built fromttandh=t−sh=t\-s\. Every gate head is zero\-initialised, so at initialisationγb≡0\\gamma\_\{b\}\\equiv 0, the update is scaled by11, and the model is*forward\-identical*to the unmodified ReQFlow trunk\. This lets the two\-time MeanFlow conditioning enter each block’s geometry update per residue and per step, without perturbing the pretrained trunk at the start of training\. In implementation the gate is attached via a forward hook, so the trunk’s forward code is untouched\.

### N\.3Shared two\-time embedding

A MeanFlow prediction is conditioned on the pair\(t,s\)\(t,s\)with0≤s≤t≤10\\leq s\\leq t\\leq 1; we leth=t−sh=t\-s\. Bothttandhhare passed through a*single*shared sinusoidal\-plus\-MLP embedderϕ​\(⋅\)\\phi\(\\cdot\)and concatenated in a fixed order,𝐜\(t\)=\[ϕ​\(t\)∥ϕ​\(h\)\]\\mathbf\{c\}^\{\(t\)\}=\[\\phi\(t\)\\,\\\|\\,\\phi\(h\)\]\. Concatenation preserves the\(t,h\)\(t,h\)ordering; the same embedding feeds both the node/edge input features and the per\-block AdaLN gate of §[N\.2](https://arxiv.org/html/2607.27431#A14.SS2)\. This is the only place the single\-time ReQFlow trunk is made time\-pair aware\.

![Refer to caption](https://arxiv.org/html/2607.27431v1/x10.png)Figure 9:Residue\-length distribution of the SCOPe backbone dataset used in our experiments \(3,673 backbones with lengths in\[60,128\]\[60,128\]\)\.

## Appendix OExtra Results/Settings in the Protein Experiment

![Refer to caption](https://arxiv.org/html/2607.27431v1/x11.png)

![Refer to caption](https://arxiv.org/html/2607.27431v1/x12.png)

Figure 10:Secondary\-structure statistics \(top:T=100T\{=\}100, bottom:T=20T\{=\}20\)\.We provide additional visualizations in this section\. Figure[10](https://arxiv.org/html/2607.27431#A15.F10)reports the joint distribution of per\-sample helix and strand content \(via mdtraj DSSP\) for the SCOPe\-generated backbones of each method, atT=100T\{=\}100\(top\) andT=20T\{=\}20\(bottom\)\. Each panel is a22\-D histogram over the helix\-fraction \(xx\) and strand\-fraction \(yy\) plane, sharing a common colour scale; the anti\-diagonal reflects the intrinsic helix–strand trade\-off within a single chain\. Across all methods the mass concentrates along this trade\-off with an additional helix\-rich mode, consistent with the SCOPe length regime\. Our models \(SE3MF and RecSE3MF\) recover a distribution comparable to the flow\-matching baselines rather than collapsing to a single motif\. Importantly, the distribution is largely preserved when the sampling budget is reduced fromT=100T\{=\}100toT=20T\{=\}20: the few\-step regime does not visibly distort the secondary\-structure statistics, indicating that the average\-velocity consistency learned during training transfers to aggressive step reduction without a mode shift\.

##### Visualization of SCOPe\.

Figure[9](https://arxiv.org/html/2607.27431#A14.F9)shows the residue\-length distribution of the SCOPe backbones used in our experiments\.

##### Random seed in evaluation\.

All evaluations use a fixed random seed for exact reproducibility: for each sampled backbone the RNG seed is set deterministically to12345\+105​T\+103​N\+i12345\+10^\{5\}\\,T\+10^\{3\}\\,N\+i, whereTTis the number of sampling steps,NNthe chain length, andiithe sample index within that length\.

##### How large is the Jacobian correction on the training distribution?

Figure[5](https://arxiv.org/html/2607.27431#A6.F5)shows that dropping the Jacobian breaks the identity by𝒪​\(t−s\)\\mathcal\{O\}\(t\-s\)on an analytic path\. To check that this regime is actually visited during training, we evaluateρ=‖\(J−1​\(Aθs→t\)−I\)​ωt‖2/‖ωt‖2\\rho=\\\|\(J^\{\-1\}\(A^\{s\\to t\}\_\{\\theta\}\)\-I\)\\,\\omega\_\{t\}\\\|\_\{2\}/\\\|\\omega\_\{t\}\\\|\_\{2\}—the relative change the Jacobian makes to the regression target of \([12](https://arxiv.org/html/2607.27431#S3.E12)\)—over1\.2×1051\.2\\times 10^\{5\}residue\-level samples from held\-out SCOPe backbones under the Stage\-2 time sampler \(SE\(3\)\-OT coupling, Table[9](https://arxiv.org/html/2607.27431#A15.T9)\)\. The correction grows monotonically with the interval: medianρ\\rhois0\.8%0\.8\\%fort−s<0\.1t\-s<0\.1,7\.2%7\.2\\%fort−s∈\[0\.25,0\.5\)t\-s\\in\[0\.25,0\.5\), and28%28\\%fort−s\>0\.75t\-s\>0\.75\(P90:2\.6%2\.6\\%,19%19\\%,58%58\\%\)\. Since the sampler drawsssacross the full range belowtt,22%22\\%of training targets fall in the regime where the correction exceeds10%10\\%\. The right Jacobian is therefore not a negligible term on the distribution the objective is trained over, even though the intervals queried at inference \(t−s=1/Tt\-s=1/T\) are short enough thatJ−1≈IJ^\{\-1\}\\approx Ithere; the log\-map approximationJ−1:=IJ^\{\-1\}:=Iused by RMF\-PTZhonget al\.\[[2026](https://arxiv.org/html/2607.27431#bib.bib15)\]discards a correction of this size during training\.

Table 9:The inverse right Jacobian’s relative change to the rotation target,ρ=‖\(J−1​\(Aθs→t\)−I\)​ωt‖2/‖ωt‖2\\rho=\\\|\(J^\{\-1\}\(A^\{s\\to t\}\_\{\\theta\}\)\-I\)\\,\\omega\_\{t\}\\\|\_\{2\}/\\\|\\omega\_\{t\}\\\|\_\{2\}, and the driving angleϕ=‖Aθs→t‖2\\phi=\\\|A^\{s\\to t\}\_\{\\theta\}\\\|\_\{2\}\(rad\), binned by the intervalt−st\-sover1\.2×1051\.2\\times 10^\{5\}residue\-level samples from held\-out SCOPe backbones under the Stage\-2 \(SE\(3\)\-OT\) sampler\.
### O\.1Training budget: steps versus epochs

We report the training budget in optimizer steps, not epochs: the QFlow/ReQFlow loader \(from FrameFlow\) oversamples each epoch to≈460\\approx\\\!460steps, while ours does one pass at≈35\\approx\\\!35steps—a∼13×\\sim\\\!13\\timesgap on the identical SCOPe set that makes epoch counts incomparable\. Importantly, both QFlow/ReQFlow and our method use the same hardware \(4×\\times80GB GPUs\) and the same batch cap \(seeNmax2N^\{2\}\_\{\\max\}in Table[7](https://arxiv.org/html/2607.27431#A13.T7)\), so each optimizer step processes the same amount of data; the only difference is that under our loader each sample is seen once per epoch, whereas QFlow/ReQFlow repeats samples multiple times due to oversampling \(which depends on the number of GPUs\)\. Steps are directly comparable; as a check, the QFlow checkpoint’s90,16090\{,\}160steps equal its reported195195epochs\[Yueet al\.,[2025](https://arxiv.org/html/2607.27431#bib.bib22)\]\.

Table 10:Controlled objective ablation on SCOPe\.All rows start from the same Stage\-1α\\alpha\-Flow checkpoint \(step100100k\) and share the IPA trunk, data, coupling, self\-conditioning and optimizer; only the Stage\-2 consistency target is swapped\.*Train Steps*counts Stage\-2 steps only\. The semigroup rows instantiate the objectives ofWooet al\.\[[2026](https://arxiv.org/html/2607.27431#bib.bib23)\]inside our pipeline; they are*not*a reproduction of that work, whose hyperparameters, model size and training budget we do not replicate\. Best per performance metric within each sampling\-step group in bold\.
### O\.2Ablation Study: A Controlled Comparison of Consistency Objectives

The RMF rows of Table[1](https://arxiv.org/html/2607.27431#S4.T1)come from the released checkpoint, which uses a different trunk \(437437M versus our16\.816\.8M\) and a much larger budget \(598598k versus our Stage\-2 budget\), so that comparison does not isolate the objective\. Table[10](https://arxiv.org/html/2607.27431#A15.T10)removes the confound: all three rows start from the*same*Stage\-1α\\alpha\-Flow checkpoint and share the trunk, data, coupling, self\-conditioning and optimizer, so the only variable is the Stage\-2 consistency target\.

Under this matched budget the semigroup target is a weaker self\-consistency signal than the differential MeanFlow target\. The FM\+\+semigroup variant tracks plain flow matching closely at every budget — its designable fraction0\.886/0\.880/0\.781/0\.5330\.886/0\.880/0\.781/0\.533is within noise of QFlow’s0\.885/0\.870/0\.778/0\.5590\.885/0\.870/0\.778/0\.559in Table[1](https://arxiv.org/html/2607.27431#S4.T1), including the same collapse atT=10T\{=\}10, which is the signature of the instantaneous\-velocity parameterisation\. The endpoint\+\+semigroup variant is stronger in the aggressive regime but still falls short of MeanFlow at every budget\. Swapping in our MeanFlow target, on the same trunk and from the same initialization, recovers the few\-step regime \(0\.8670\.867atT=20T\{=\}20,0\.7280\.728atT=10T\{=\}10\)\.

We read this as a matter of training budget rather than of correctness: the two objectives share a minimizer \(Remark[13](https://arxiv.org/html/2607.27431#Thmremark13)\), but the semigroup constraint propagates the boundary condition only through the interval triples it happens to sample, whereas the differential target imposes the same consistency pointwise at every\(s,t\)\(s,t\)\. The semigroup route therefore appears to need a substantially longer schedule to reach the few\-step regime — consistent with RMF’s∼598​k\\sim\\\!598\\mathrm\{k\}steps, and with our observation in Appendix[K](https://arxiv.org/html/2607.27431#A11)that semigroup training gave weaker few\-step gains than the JVP\-based loss at1010–2020k steps\.

Table 11:Evaluation results for V\-QFlow and V\-ReQFlow\. These methods use a different data/evaluation pipeline from the other baselines, so we report them separately\.
### O\.3Additional results for V\-QFlow and V\-ReQFlow

V\-QFlow and V\-ReQFlowYueet al\.\[[2025](https://arxiv.org/html/2607.27431#bib.bib22)\]do not provide checkpoints trained on SCOPe; therefore, the scores in Table[11](https://arxiv.org/html/2607.27431#A15.T11)are computed using the released checkpoints trained on PDB\. Because this evaluation uses a different training dataset and pipeline from the SCOPe\-based baselines, the results are not directly comparable; we thus report them in this separate section\. Notably, while the overall score is close to 0\.99, the helix accuracy is high whereas the strand secondary\-structure accuracy is around 0\.1 or lower\. This imbalance is concerning biologically becauseβ\\beta\-strands typically assemble intoβ\\beta\-sheets via inter\-strand hydrogen bonding and often depend on non\-local sequence interactions; correspondingly, strand prediction is generally harder than helix prediction, and very low strand accuracy may indicate that the model fails to capture such long\-range constraintsPaulinget al\.\[[1951](https://arxiv.org/html/2607.27431#bib.bib45)\], Zhang and Sagui \[[2015](https://arxiv.org/html/2607.27431#bib.bib28)\]\.

## Appendix PComputational Resources

All experiments are conducted on a single GPU node with the following configuration\. Each node is equipped with4×4\\timesNVIDIA H100 80 GB HBM3 GPUs and dual AMD EPYC 9654 96\-core CPUs\. Training uses all four GPUs with PyTorch Lightning distributed data\-parallel \(DDP\),1414CPU worker threads per GPU process \(5656cores in total\), and gradient accumulation of22; the effective batch is set by a length\-based sampler with capNmax2=5×105N^\{2\}\_\{\\max\}=5\\times 10^\{5\}residues2\(Table[7](https://arxiv.org/html/2607.27431#A13.T7)\)\.

Similar Articles

Follow the Mean: Reference-Guided Flow Matching

Hugging Face Daily Papers

This paper introduces a method for controllable generation in flow matching by adjusting the conditional endpoint mean using a reference set, offering both training-free and semi-parametric guidance for style and content control.

Variable-Length Generative Protein Design via Generalized Poisson Flow

arXiv cs.LG

Introduces Generalized Poisson Flow (GPFlow), a variable-length generative framework for protein design that learns an inhomogeneous generalized Poisson process, enabling flexible length exploration and improving designability across structure, sequence, and peptide co-design tasks.

Continuous Adversarial MeanFlow Transfer

arXiv cs.LG

This paper proposes MeanFlow-Transfer (MF-T) and Continuous Adversarial MeanFlow (CAMF) to unify the adaptation and acceleration of pretrained diffusion and flow models, enabling high-quality few-step generation on new domains with limited data.

MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators

Hugging Face Daily Papers

MeanFlowNFT introduces a forward-process reinforcement learning method for average-velocity generators, enabling efficient alignment with human preferences while preserving fast few-step sampling. Experiments show it outperforms prior RL-tuned few-step generators on most metrics and even surpasses multi-step RL-tuned diffusion models.