Second Order Drifting Models
Summary
This paper proposes Second-Order Drifting Models, which augment drifting generative models with artificial velocity variables to achieve accelerated second-order dynamics in Fourier space, mitigating spectral stiffness while preserving one-step inference. The method is evaluated on synthetic distribution matching, sequential data generation, and robotic control, showing improved convergence and competitive performance.
View Cached Full Text
Cached at: 08/11/26, 08:08 AM
# Second Order Drifting Models
Source: [https://arxiv.org/html/2608.07924](https://arxiv.org/html/2608.07924)
Drake Brown\*, Yuhao Huang\*, Shih\-Hsin Wang, Bao Wang Department of Mathematics Scientific Computing and Imaging \(SCI\) Institute University of Utah 111Drake Brown and Yuhao Huang are co\-first authors\.
###### Abstract
Drifting models are a recent class of one\-step generative models that evolve the model distribution during training using a predefined sample\-based drift field\. Although they avoid iterative inference, their kernel\-based drift fields induce frequency\-dependent training dynamics: In the linearized regime, each Fourier mode of the density residual decays at a rate determined by the kernel spectrum, leading to slow recovery of fine\-scale structure\. We propose Second\-Order Drifting Models, which lift drifting dynamics into phase space by augmenting generated samples with artificial velocity variables\. We show that the resulting density perturbations obey accelerated second\-order dynamics in Fourier space, connecting drifting models to the celebrated Nesterov acceleration from optimization theory\. This provides a principled mechanism for mitigating the spectral stiffness of first\-order drifting while preserving one\-step inference\. We derive a practical semi\-implicit training algorithm and evaluate it on synthetic distribution matching, sequential data generation, and robotic control\. Across these settings, the second\-order drifting model improves convergence behavior and achieves competitive or superior performance over first\-order drifting baselines\.
## 1Introduction
Flow\-based generative models—particularly those based on diffusion \(cf\.\[[43](https://arxiv.org/html/2608.07924#bib.bib43),[18](https://arxiv.org/html/2608.07924#bib.bib18),[46](https://arxiv.org/html/2608.07924#bib.bib46)\]\) and flow matching \(cf\.\[[28](https://arxiv.org/html/2608.07924#bib.bib28),[1](https://arxiv.org/html/2608.07924#bib.bib1),[29](https://arxiv.org/html/2608.07924#bib.bib29),[20](https://arxiv.org/html/2608.07924#bib.bib20)\]\) mechanisms—can be formulated as learning a transformationffsuch that the pushforward of a source distributionqqmatches the target data distributionpp, i\.e\.,f♯q≈pf\_\{\\sharp\}q\\approx p222The mapf:𝒳→𝒳f:\\mathcal\{X\}\\to\\mathcal\{X\}that transports a source distributionqqto the target data distributionpp\. Concretely, ifX∼qX\\sim qandY=f\(X\)Y=f\(X\), thenYYfollows the pushforward distributionf♯qf\_\{\\sharp\}q, defined by\(f♯q\)\(A\):=q\(f−1\(A\)\)for all measurable setsA⊆𝒳\.\(f\_\{\\sharp\}q\)\(A\):=q\\bigl\(f^\{\-1\}\(A\)\\bigr\)\\quad\\text\{for all measurable sets \}A\\subseteq\\mathcal\{X\}\.The learning objective is to constructffsuch thatf♯q≈pf\_\{\\sharp\}q\\approx p\.\. Despite their remarkable performance in high\-fidelity generation\[[39](https://arxiv.org/html/2608.07924#bib.bib39),[19](https://arxiv.org/html/2608.07924#bib.bib19),[52](https://arxiv.org/html/2608.07924#bib.bib52),[48](https://arxiv.org/html/2608.07924#bib.bib48),[22](https://arxiv.org/html/2608.07924#bib.bib22)\], these transformations are defined via differential equations and require iterative evaluations at inference time—often requires hundreds of neural function evaluations \(NFEs\) per sample\. This makes generation computationally expensive and limits real\-time applicability, motivating the development of high\-quality single\-step models, including consistency models\[[45](https://arxiv.org/html/2608.07924#bib.bib45),[44](https://arxiv.org/html/2608.07924#bib.bib44)\], flow maps\[[5](https://arxiv.org/html/2608.07924#bib.bib5),[6](https://arxiv.org/html/2608.07924#bib.bib6)\], and Meanflows\[[14](https://arxiv.org/html/2608.07924#bib.bib14),[21](https://arxiv.org/html/2608.07924#bib.bib21)\]\.
Recently, drifting models\[[10](https://arxiv.org/html/2608.07924#bib.bib10)\]have been proposed to address this inefficiency by shifting distribution matching entirely to training\. In this framework, the neural networkfθf\_\{\\theta\}acts as a single\-step generator, mapping a simple prior distributionqq\(e\.g\., Gaussian noise\) directly to samples in data space\. At any training timett, the model induces a pushforward distributionqt=fθt♯qq\_\{t\}=f\_\{\\theta\_\{t\}\\sharp\}q\. The goal of learning is to evolveqtq\_\{t\}so that it matches the target data distributionpp\. Crucially, the model does not explicitly learn a vector field parameterized by a neural network\. Instead, training is guided by a pre\-designed drifting fieldVp,qt\(x\)V\_\{p,q\_\{t\}\}\(x\), which depends on both the current generated distributionqtq\_\{t\}and the target distributionpp\. This field prescribes how a samplext∼qtx\_\{t\}\\sim q\_\{t\}should move so as to reduce the discrepancy betweenqtq\_\{t\}andpp\. Intuitively,Vp,qt\(x\)V\_\{p,q\_\{t\}\}\(x\)can be viewed as a velocity that attracts generated samples toward regions of high data density while repelling them from over\-represented regions ofqtq\_\{t\}\.
During training, the neural network parameters evolve over iterations,θt→θt\+Δt\\theta\_\{t\}\\to\\theta\_\{t\+\\Delta t\}, which induces a corresponding evolution of generated samples\. For a fixed noise inputϵ∼q\\epsilon\\sim q, the generated sample follows a trajectoryxt=fθt\(ϵ\)x\_\{t\}=f\_\{\\theta\_\{t\}\}\(\\epsilon\)\. The parameter update thus results in a displacement
xt\+Δt−xt≈Vp,qt\(xt\)Δt,x\_\{t\+\\Delta t\}\-x\_\{t\}\\approx V\_\{p,q\_\{t\}\}\(x\_\{t\}\)\\,\\Delta t,meaning that the change in samples is driven by the drifting field\. In this sense, the optimization of network parameters implicitly transports samples according toVp,qtV\_\{p,q\_\{t\}\}\. From a continuous\-time viewpoint\[[49](https://arxiv.org/html/2608.07924#bib.bib49)\], taking the limitΔt→0\\Delta t\\to 0leads to the ordinary differential equation \(ODE\)
dxtdt=Vp,qt\(xt\)\.\\frac\{dx\_\{t\}\}\{dt\}=V\_\{p,q\_\{t\}\}\(x\_\{t\}\)\.\(1\)This ODE characterizes the training dynamics at the level of samples: the generated distributionqtq\_\{t\}evolves by flowing along the vector fieldVp,qtV\_\{p,q\_\{t\}\}\. At equilibrium, whenqt=pq\_\{t\}=p, the drifting field vanishes and the dynamics reach a fixed point\. See Section[3](https://arxiv.org/html/2608.07924#S3)for a more detailed derivation\.
The drifting field is constructed via kernel smoothing \(cf\. Section[2](https://arxiv.org/html/2608.07924#S2)\), where interactions between samples are weighted by a similarity kernel \(e\.g\., Gaussian\), resulting in an attraction–repulsion mechanism that quantifies the discrepancy between the data distributionppand the generated distributionqq\. However, this kernel\-based smoothing inherently acts as a low\-pass filter in the Fourier domain, attenuating high\-frequency components of the distributions\[[49](https://arxiv.org/html/2608.07924#bib.bib49)\]\. As a consequence, the resulting dynamics exhibit a pronounced spectral bias: while low\-frequency modes converge rapidly, high\-frequency modes decay much more slowly—exponentially so under a Gaussian kernel\. A similar, albeit less severe, issue persists when using a Laplacian kernel\. This spectral bias creates a fundamental bottleneck, slowing convergence and limiting the model’s ability to accurately capture fine\-scale features of the data; see detailed analysis in Section[3](https://arxiv.org/html/2608.07924#S3)\.
### 1\.1Our Contributions
Motivated by the spectral bottleneck described above, we propose*second\-order drifting models*, which lift the dynamics into phase space by introducing an auxiliary velocity variablevv:
\{dxdt=v,dvdt=Vp,qt\(x\)−γv,\\begin\{cases\}\\frac\{dx\}\{dt\}=v,\\\\ \\frac\{dv\}\{dt\}=V\_\{p,q\_\{t\}\}\(x\)\-\\gamma v,\\end\{cases\}\(2\)whereγ\>0\\gamma\>0is a damping coefficient \(can bett\-dependent\)\. By incorporating inertia/velocity, the resulting dynamics can effectively mitigate the low\-pass bias induced by kernel smoothing\.
At the distribution level, the linearized dynamics follow second\-order ODEs in Fourier space, turning each mode into an accelerated system analogous to the Nesterov\[[32](https://arxiv.org/html/2608.07924#bib.bib32)\]dynamics\. As a result, convergence is no longer constrained by the kernel spectrumλ\(ω\)\\lambda\(\\omega\), but can be regulated via the damping schedule, effectively reducing spectral stiffness and accelerating high\-frequency recovery\. Guided by this insight, we develop a semi\-implicit second\-order training algorithm that preserves one\-step generation while improving stability, convergence, and fine\-scale fidelity\. Empirically, our method achieves consistent improvements across synthetic generation, sequential data generation, and robotic control tasks\. In summary, our contributions are threefold:
- •Acceleration perspective\.We view first\-order drifting as spectrally imbalanced and introduce momentum to counter its low\-pass bias\.
- •Second\-order dynamics\.We propose a phase\-space formulation whose Fourier dynamics recover the Nesterov acceleration, reducing dependence on the kernel spectrum\.
- •Algorithm and validation\.We develop a semi\-implicit method that preserves one\-step inference and improves convergence and high\-frequency fidelity across tasks\.
### 1\.2Additional Related Works
Acceleration and Second\-Order Optimization Dynamics\.Momentum\-based acceleration has a long history in optimization \(cf\.\[[38](https://arxiv.org/html/2608.07924#bib.bib38),[32](https://arxiv.org/html/2608.07924#bib.bib32)\]\)\. Continuous\-time second\-order ODE formulations of these methods have been extensively studied to characterize convergence rates and guide algorithm design \(cf\.\[[4](https://arxiv.org/html/2608.07924#bib.bib4),[47](https://arxiv.org/html/2608.07924#bib.bib47),[2](https://arxiv.org/html/2608.07924#bib.bib2),[3](https://arxiv.org/html/2608.07924#bib.bib3),[40](https://arxiv.org/html/2608.07924#bib.bib40),[53](https://arxiv.org/html/2608.07924#bib.bib53),[24](https://arxiv.org/html/2608.07924#bib.bib24)\]\)\. Such perspectives have also inspired acceleration techniques for neural ODE\-related models\[[35](https://arxiv.org/html/2608.07924#bib.bib35),[34](https://arxiv.org/html/2608.07924#bib.bib34),[54](https://arxiv.org/html/2608.07924#bib.bib54),[51](https://arxiv.org/html/2608.07924#bib.bib51),[33](https://arxiv.org/html/2608.07924#bib.bib33)\]and diffusion models\[[11](https://arxiv.org/html/2608.07924#bib.bib11)\]\. In this paper, we bring this optimization viewpoint to drifting models: By linearizing the dynamics, each Fourier mode of the density perturbation behaves like a scalar optimization problem with conditioning governed by the kernel spectrum\. This perspective enables the design of second\-order dynamics to accelerate distribution matching\.
Advances in Drifting Models\.Several recent works have extended the drifting framework in complementary directions\. We highlight a few representative advances here, noting that a comprehensive review is beyond the scope of this paper due to space limitations\. The authors of\[[49](https://arxiv.org/html/2608.07924#bib.bib49)\]relate drifting model training to score matching, providing a theoretical foundation for our work\. Sinkhorn drifting exploits connections to Sinkhorn gradient flows to improve performance in low\-temperature regimes\[[16](https://arxiv.org/html/2608.07924#bib.bib16)\]\. A unified perspective relates drifting to score\-based models by showing that the mean\-shift direction corresponds to score mismatch, providing a coherent view across temperature regimes\[[25](https://arxiv.org/html/2608.07924#bib.bib25)\]\. From a transport perspective, a long–short flow\-map formulation derives drifting from flow\-based models and achieves competitive performance with reduced batch sizes\[[27](https://arxiv.org/html/2608.07924#bib.bib27)\]\. Gradient flow drifting interprets the method as a Wasserstein gradient flow of a kernel density estimation \(KDE\)\-based divergence, enabling the design of alternative velocity fields that mitigate mode collapse\[[7](https://arxiv.org/html/2608.07924#bib.bib7)\]\. In addition, analytical correction methods address minibatch\-induced bias in the drifting field through closed\-form adjustments\[[55](https://arxiv.org/html/2608.07924#bib.bib55)\], while friction\-augmented variants introduce damping mechanisms to stabilize the dynamics and improve performance in domain translation tasks\[[23](https://arxiv.org/html/2608.07924#bib.bib23)\]\.
## 2Background: Drifting Models
In this section, we review the drifting model\[[10](https://arxiv.org/html/2608.07924#bib.bib10)\]\. Letppbe a target distribution onℝd\\mathbb\{R\}^\{d\},fθ:ℝk→ℝdf\_\{\\theta\}:\\mathbb\{R\}^\{k\}\\to\\mathbb\{R\}^\{d\}a generator, andq:=\(fθ\)♯𝒩\(0,I\)q:=\(f\_\{\\theta\}\)\_\{\\sharp\}\\mathcal\{N\}\(0,I\)the induced model distribution, where𝒩\(0,I\)\{\\mathcal\{N\}\}\(0,I\)is standard Gaussian\. The drifting model learns a kernel\-based*drift vector field*directly from samples\.
For a positive kernelk\(x,y\)k\(x,y\), a particular construction of the drift field\[[10](https://arxiv.org/html/2608.07924#bib.bib10)\]is given by
Vp,q\(x\)=Vp\(x\)−Vq\(x\),V\_\{p,q\}\(x\)=V\_\{p\}\(x\)\-V\_\{q\}\(x\),with
Vp\(x\)=𝔼y∼p\[k\(x,y\)\(y−x\)\]𝔼y∼p\[k\(x,y\)\],Vq\(x\)=𝔼y∼q\[k\(x,y\)\(y−x\)\]𝔼y∼q\[k\(x,y\)\]\.V\_\{p\}\(x\)=\\frac\{\\mathbb\{E\}\_\{y\\sim p\}\[k\(x,y\)\(y\-x\)\]\}\{\\mathbb\{E\}\_\{y\\sim p\}\[k\(x,y\)\]\},\\ V\_\{q\}\(x\)=\\frac\{\\mathbb\{E\}\_\{y\\sim q\}\[k\(x,y\)\(y\-x\)\]\}\{\\mathbb\{E\}\_\{y\\sim q\}\[k\(x,y\)\]\}\.The drifting fieldVp,qV\_\{p,q\}defines stop\-gradient targets for generated samples, andfθf\_\{\\theta\}\(with initialization beingθ0\\theta\_\{0\}\) is updated so that its pushforward distribution matches the drifted distribution following:
- •Sampleϵ∼𝒩\(0,I\)\\epsilon\\sim\\mathcal\{N\}\(0,I\)and computext=fθt\(ϵ\)x\_\{t\}=f\_\{\\theta\_\{t\}\}\(\\epsilon\)
- •Driftxt\+Δt=xt\+Vp,qt\(xt\)x\_\{t\+\\Delta t\}=x\_\{t\}\+V\_\{p,q\_\{t\}\}\(x\_\{t\}\)
- •Updateθt\\theta\_\{t\}by backpropagating the following loss function: ℒ\(θ\)=𝔼ϵ\[‖fθt\(ϵ\)−sg\[xt\+Δt\]‖2\],\{\\mathcal\{L\}\}\(\\theta\)=\\mathbb\{E\}\_\{\\epsilon\}\[\\\|f\_\{\\theta\_\{t\}\}\(\\epsilon\)\-\\texttt\{sg\}\[x\_\{t\+\\Delta t\}\]\\\|^\{2\}\],wheresgdenotes the stop\-gradient operation\.
After convergence,fθf\_\{\\theta\}directly maps the prior distribution𝒩\(0,I\)\{\\mathcal\{N\}\}\(0,I\)to the data distributionpp, enabling data generation with a single function evaluation \(1\-NFE\)\.
## 3A Spectral Viewpoint of Training Drifting Models
In this section, we revisit the spectral analysis of drifting models\[[49](https://arxiv.org/html/2608.07924#bib.bib49)\], demonstrating that first\-order drifting causes frequency\-dependent convergence and intrinsic spectral stiffness\.
Starting from the ODE equation[1](https://arxiv.org/html/2608.07924#S1.E1), and letqt\(x\)q\_\{t\}\(x\)denote the density of the data, then we have\[[49](https://arxiv.org/html/2608.07924#bib.bib49)\]
∂tqt\(x\)\+∇⋅\(qt\(x\)Vp,qt\(x\)\)=0\.\\partial\_\{t\}q\_\{t\}\(x\)\+\\nabla\\cdot\\big\(q\_\{t\}\(x\)V\_\{p,q\_\{t\}\}\(x\)\\big\)=0\.\(3\)
##### Linearization\.
We consider linearizing equation[3](https://arxiv.org/html/2608.07924#S3.E3)for a general positive, radial symmetric kernelk\(x,y\)=K\(‖x−y‖\)k\(x,y\)=K\(\\\|x\-y\\\|\)withKKbeing integrable and decaying in‖x−y‖\\\|x\-y\\\|; notice that both Gaussian and Laplacian kernels satisfy these conditions\. To linearize equation[3](https://arxiv.org/html/2608.07924#S3.E3)around the equilibriumpp, letρt:=qt−p\\rho\_\{t\}:=q\_\{t\}\-p\. Under mild assumptions \(see Appendix[A](https://arxiv.org/html/2608.07924#A1)\), we can show thatρt\\rho\_\{t\}satisfies the PDE:
∂tρt\(x\)=−∇⋅\(∫M\(x−y\)ρt\(y\)𝑑y\):=−∇⋅\(M∗ρt\)\(x\),\\partial\_\{t\}\\rho\_\{t\}\(x\)=\-\\nabla\\cdot\\Big\(\\int M\(x\-y\)\\rho\_\{t\}\(y\)dy\\Big\):=\-\\nabla\\cdot\(M\*\\rho\_\{t\}\)\(x\),\(4\)whereM\(x\)=cxK\(x\)M\(x\)=cxK\(x\)withccbeing a constant\.
Taking the Fourier transform of equation[4](https://arxiv.org/html/2608.07924#S3.E4)yields a decoupled system:
∂tρ^t\(ω\)=λ\(ω\)ρ^t\(ω\),λ\(ω\):=ω⋅∇ωK^\(ω\),\\partial\_\{t\}\\hat\{\\rho\}\_\{t\}\(\\omega\)=\\lambda\(\\omega\)\\hat\{\\rho\}\_\{t\}\(\\omega\),\\quad\\lambda\(\\omega\):=\\omega\\cdot\\nabla\_\{\\omega\}\\hat\{K\}\(\\omega\),\(5\)whereK^\(ω\)\\hat\{K\}\(\\omega\)is the Fourier transform ofKK\. Thus, each frequency evolves independently and the dynamics are fully characterized by the spectral rateλ\(ω\)\\lambda\(\\omega\)\.
###### Proposition 3\.1\(Spectral stiffness of first\-order drifting\)\.
Under the above assumptions, the Fourier modes of the residual satisfy
∂tρ^t\(ω\)=λ\(ω\)ρ^t\(ω\),λ\(ω\)≤0\.\\partial\_\{t\}\\hat\{\\rho\}\_\{t\}\(\\omega\)=\\lambda\(\\omega\)\\hat\{\\rho\}\_\{t\}\(\\omega\),\\quad\\lambda\(\\omega\)\\leq 0\.Consequently, the time required to contract modeω\\omegaby a factorε\\varepsilonis
Tε\(ω\)=1\|λ\(ω\)\|log1ε\.T\_\{\\varepsilon\}\(\\omega\)=\\frac\{1\}\{\|\\lambda\(\\omega\)\|\}\\log\\frac\{1\}\{\\varepsilon\}\.
This proposition reveals a key limitation: convergence is controlled by the inverse spectral rate\|λ\(ω\)\|−1\|\\lambda\(\\omega\)\|^\{\-1\}\. Modes with small\|λ\(ω\)\|\|\\lambda\(\\omega\)\|become bottlenecks, resulting in slow convergence\.
##### Examples\.
We now examine the spectral rateλ\(ω\)\\lambda\(\\omega\)for two standard kernels\.
- •Gaussian kernel\.ForKG\(z\)=exp\(−‖z‖2/\(2σ2\)\)K\_\{G\}\(z\)=\\exp\(\-\\\|z\\\|^\{2\}/\(2\\sigma^\{2\}\)\), we obtain λ\(ω\)=−σ2\(2πσ2\)d/2‖ω‖2exp\(−σ22‖ω‖2\)<0\.\\lambda\(\\omega\)=\-\\frac\{\\sigma^\{2\}\}\{\(2\\pi\\sigma^\{2\}\)^\{d/2\}\}\\\|\\omega\\\|^\{2\}\\exp\\Big\(\-\\frac\{\\sigma^\{2\}\}\{2\}\\\|\\omega\\\|^\{2\}\\Big\)<0\.In particular,\|λ\(ω\)\|\|\\lambda\(\\omega\)\|decays exponentially in‖ω‖\\\|\\omega\\\|, which implies Tε\(ω\)∝exp\(σ22‖ω‖2\)‖ω‖2log1ε\.T\_\{\\varepsilon\}\(\\omega\)\\propto\\frac\{\\exp\(\\frac\{\\sigma^\{2\}\}\{2\}\\\|\\omega\\\|^\{2\}\)\}\{\\\|\\omega\\\|^\{2\}\}\\log\\frac\{1\}\{\\varepsilon\}\.Thus, high\-frequency modes converge exponentially slowly\.
- •Laplacian kernel\.ForKL\(z\)=exp\(−‖z‖/σ\)K\_\{L\}\(z\)=\\exp\(\-\\\|z\\\|/\\sigma\), we obtain λ\(ω\)=−Cd\(d\+1\)σ2‖ω‖2\(1\+σ2‖ω‖2\)\(d\+3\)/2<0,\\lambda\(\\omega\)=\-\\frac\{C\_\{d\}\(d\+1\)\\sigma^\{2\}\\\|\\omega\\\|^\{2\}\}\{\(1\+\\sigma^\{2\}\\\|\\omega\\\|^\{2\}\)^\{\(d\+3\)/2\}\}<0,whereCdC\_\{d\}is a dimension\-dependent constant\. This yields Tε\(ω\)∝\(1\+σ2‖ω‖2\)\(d\+3\)/2‖ω‖2log1ε\.T\_\{\\varepsilon\}\(\\omega\)\\propto\\frac\{\(1\+\\sigma^\{2\}\\\|\\omega\\\|^\{2\}\)^\{\(d\+3\)/2\}\}\{\\\|\\omega\\\|^\{2\}\}\\log\\frac\{1\}\{\\varepsilon\}\.Here, the decay is polynomial, but still leads to severe spectral imbalance\.
##### Interpretation \(low\-pass dynamics\)\.
SinceK^\(ω\)\\hat\{K\}\(\\omega\)decays with‖ω‖\\\|\\omega\\\|, the spectral rateλ\(ω\)\\lambda\(\\omega\)vanishes at high frequencies\. As a result, the dynamics preferentially eliminate low\-frequency errors, while high\-frequency components persist for long times\. In effect, the kernel induces a low\-pass filter on the residual dynamics\. This intrinsic spectral stiffness is a structural consequence of kernel smoothing, and explains why first\-order drifting alone leads to slow convergence of fine\-scale structure\.
## 4Second\-Order Drifting Models
From the previous analysis, the convergence ofρ^t\(ω\)\\hat\{\\rho\}\_\{t\}\(\\omega\)is frequency\-dependent: for Gaussian kernels, it is exponentially slow at largeω\\omega, while for Laplacian kernels it is slow at both high and low frequencies\. For the Gaussian case,\[[49](https://arxiv.org/html/2608.07924#bib.bib49)\]mitigates this via an annealed bandwidthσ\(t\)\\sigma\(t\)\. However, this approach is kernel\-specific, raising the question:
*Can we accelerate drifting model training for positive, radial, and translation\-invariant kernels, ideally making iteration complexity frequency\-independent?*
To address this question, we reinterpret the linearized first\-order dynamics from an optimization perspective\. Sinceλ\(ω\)≤0\\lambda\(\\omega\)\\leq 0\(cf\. Proposition[B\.1](https://arxiv.org/html/2608.07924#A2.Thmtheorem1)\), the Fourier residual evolves as
ρ^t\(ω\)=ρ^0\(ω\)exp\(λ\(ω\)t\)\.\\hat\{\\rho\}\_\{t\}\(\\omega\)=\\hat\{\\rho\}\_\{0\}\(\\omega\)\\exp\(\\lambda\(\\omega\)t\)\.When\|λ\(ω\)\|\|\\lambda\(\\omega\)\|is small, the corresponding mode decays slowly\. Equivalently, for fixedω\\omega, equation[5](https://arxiv.org/html/2608.07924#S3.E5)can be viewed as the continuous\-time limit of gradient descent on the convex quadratic objective
minρ^−12\(λ\(ω\)\)ρ^2\(ω\)\.\\min\_\{\\hat\{\\rho\}\}\-\\frac\{1\}\{2\}\(\\lambda\(\\omega\)\)\\hat\{\\rho\}^\{2\}\(\\omega\)\.This viewpoint motivates importing acceleration from optimization\. Rather than evolving the density residual via a first\-order ODE, we introduce second\-order damped ODE\.
### 4\.1Accelerated Second\-Order Dynamics via Nesterov ODE
To overcome the spectral stiffness of first\-order drifting, we directly adopt a time\-dependent second\-order dynamics inspired by Nesterov acceleration\. In the Fourier domain, we propose the evolution
∂t2ρ^t\(ω\)\+αt∂tρ^t\(ω\)−λ\(ω\)ρ^t\(ω\)=0,\\partial\_\{t\}^\{2\}\\hat\{\\rho\}\_\{t\}\(\\omega\)\+\\alpha\_\{t\}\\,\\partial\_\{t\}\\hat\{\\rho\}\_\{t\}\(\\omega\)\-\\lambda\(\\omega\)\\hat\{\\rho\}\_\{t\}\(\\omega\)=0,\(6\)whereαt=αt\\alpha\_\{t\}=\\frac\{\\alpha\}\{t\}withα≥3\\alpha\\geq 3\.
Equation[6](https://arxiv.org/html/2608.07924#S4.E6)can be viewed as the continuous\-time limit of Nesterov’s accelerated gradient method\[[47](https://arxiv.org/html/2608.07924#bib.bib47)\]\. Compared to the first\-order ODE, which evolve each frequency independently as∂tρ^t\(ω\)=λ\(ω\)ρ^t\(ω\)\\partial\_\{t\}\\hat\{\\rho\}\_\{t\}\(\\omega\)=\\lambda\(\\omega\)\\hat\{\\rho\}\_\{t\}\(\\omega\), the second\-order ODE introduces inertia, allowing faster decay of modes with small\|λ\(ω\)\|\|\\lambda\(\\omega\)\|\. In particular, classical results suggest that such dynamics achieve accelerated convergence rates of order𝒪\(1/t2\)\\mathcal\{O\}\(1/t^\{2\}\)for convex objectives, significantly mitigating the dependence on the spectral rateλ\(ω\)\\lambda\(\\omega\)\.
##### Why constant damping is insufficient\.
In optimization theory, another natural alternative is the constant\-damping second\-order system
∂t2ρ^t\(ω\)\+γ∂tρ^t\(ω\)−λ\(ω\)ρ^t\(ω\)=0,\\partial\_\{t\}^\{2\}\\hat\{\\rho\}\_\{t\}\(\\omega\)\+\\gamma\\,\\partial\_\{t\}\\hat\{\\rho\}\_\{t\}\(\\omega\)\-\\lambda\(\\omega\)\\hat\{\\rho\}\_\{t\}\(\\omega\)=0,\(7\)which corresponds to the heavy\-ball method\[[38](https://arxiv.org/html/2608.07924#bib.bib38)\]\. Its dynamical behavior depends on the roots of the characteristic equation
r2\+γr−λ\(ω\)=0\.r^\{2\}\+\\gamma r\-\\lambda\(\\omega\)=0\.The dynamics exhibit three regimes depending on the relation betweenγ2\\gamma^\{2\}and−4λ\(ω\)\-4\\lambda\(\\omega\):
- •Underdampedγ2<−4λ\\gamma^\{2\}<\-4\\lambda:the two roots arer=−γ/2±i\|4λ\+γ2\|/2,r=\-\\gamma/2\\pm i\\sqrt\{\|4\\lambda\+\\gamma^\{2\}\|\}/2,and the solution is ρ^t\(ω\)=e−γ2t\(C1cos\(\|4λ\+γ2\|/2t\)\+C2sin\(\|4λ\+γ2\|/2t\)\),C1,C2are constants\.\\hat\{\\rho\}\_\{t\}\(\\omega\)=e^\{\-\\frac\{\\gamma\}\{2\}t\}\\big\(C\_\{1\}\\cos\(\\sqrt\{\|4\\lambda\+\\gamma^\{2\}\|\}/2t\)\+C\_\{2\}\\sin\(\\sqrt\{\|4\\lambda\+\\gamma^\{2\}\|\}/2t\)\\big\),\\ C\_\{1\},C\_\{2\}\\ \\text\{are constants\}\.In this regime, the residual oscillates while its envelope decays at ratee−γt/2e^\{\-\\gamma t/2\}\.
- •Critically\-dampedγ2=−4λ\\gamma^\{2\}=\-4\\lambda:the characteristic equation has one repeated real rootr=−γ/2r=\-\\gamma/2, and the solution is ρ^t\(ω\)=e−γ2t\(C1\+C2t\),C1,C2are constants\.\\hat\{\\rho\}\_\{t\}\(\\omega\)=e^\{\-\\frac\{\\gamma\}\{2\}t\}\(C\_\{1\}\+C\_\{2\}t\),\\ C\_\{1\},C\_\{2\}\\ \\text\{are constants\}\.This regime achieves the fastest non\-oscillatory decay for a given mode\.
- •Overdampedγ2\>−4λ\\gamma^\{2\}\>\-4\\lambda:the two roots arer1,2=−γ/2±γ2\+4λ/2,r\_\{1,2\}=\-\\gamma/2\\pm\\sqrt\{\\gamma^\{2\}\+4\\lambda\}/2,and the solution is ρ^t\(ω\)=C1er1t\+C2er2t,C1,C2are constants\.\\hat\{\\rho\}\_\{t\}\(\\omega\)=C\_\{1\}e^\{r\_\{1\}t\}\+C\_\{2\}e^\{r\_\{2\}t\},\\ C\_\{1\},C\_\{2\}\\ \\text\{are constants\}\.In this case, the decay can become slow when the larger root is close to zero\.
Whenγ2≤−4λ\(ω\)\\gamma^\{2\}\\leq\-4\\lambda\(\\omega\), the modeρ^t\(ω\)\\hat\{\\rho\}\_\{t\}\(\\omega\)decays at a rate governed byγ\\gammarather than\|λ\(ω\)\|\|\\lambda\(\\omega\)\|, mitigating frequency dependence in the underdamped or critically damped regimes\. However, no fixedγ\\gammacan critically damp all frequencies: modes withγ2\>−4λ\(ω\)\\gamma^\{2\}\>\-4\\lambda\(\\omega\)become overdamped and may converge more slowly\. Therefore, constant damping cannot fully eliminate spectral imbalance\. Designing more effective time\-dependent damping schedules is an important yet challenging problem\. Such designs could not only lead to new training algorithms for drifting models, but also provide new insights for acceleration in optimization\. This remains a challenging direction for future research\.
### 4\.2Designing Second\-Order Drifting Models
In this subsection, we design a drifting model whose frequencies follow the second\-order Nesterov ODE in the previous subsection\. We start from the following particle ODE governing the data evolution:
x˙\(i\)\\displaystyle\\dot\{x\}^\{\(i\)\}=v\(i\),\\displaystyle=v^\{\(i\)\},\(8\)v˙\(i\)\\displaystyle\\dot\{v\}^\{\(i\)\}=Vp,qt\(x\(i\)\)−αtv\(i\)\.\\displaystyle=V\_\{p,q\_\{t\}\}\(x^\{\(i\)\}\)\-\\frac\{\\alpha\}\{t\}v^\{\(i\)\}\.A stable semi\-implicit discretization gives
vn\+1\(i\)\\displaystyle v^\{\(i\)\}\_\{n\+1\}=\(1−αn\+1\)vn\(i\)\+η1⋅Vp,qn\(xn\(i\)\),\\displaystyle=\\Big\(1\-\\frac\{\\alpha\}\{n\+1\}\\Big\)v^\{\(i\)\}\_\{n\}\+\\eta\_\{1\}\\cdot V\_\{p,q^\{n\}\}\(x^\{\(i\)\}\_\{n\}\),\(9\)xn\+1\(i\)\\displaystyle x^\{\(i\)\}\_\{n\+1\}=xn\(i\)\+η2⋅vn\+1\(i\),\\displaystyle=x^\{\(i\)\}\_\{n\}\+\\eta\_\{2\}\\cdot v^\{\(i\)\}\_\{n\+1\},whereη1,η2\>0\\eta\_\{1\},\\eta\_\{2\}\>0are two constant\.
In training, the stop gradient is applied toVp,qV\_\{p,q\}exactly as in the original drifting model\. In particular, letxNx\_\{N\}be the result ofNNinertial drifting steps fromx0=fθt\(ϵ\)x\_\{0\}=f\_\{\\theta\_\{t\}\}\(\\epsilon\), the loss is given by
ℒ\(θ\)=𝔼ϵ\[‖fθt\(ϵ\)−sg\[xN\]‖2\],\{\\mathcal\{L\}\}\(\\theta\)=\\mathbb\{E\}\_\{\\epsilon\}\[\\\|f\_\{\\theta\_\{t\}\}\(\\epsilon\)\-\\texttt\{sg\}\[x\_\{N\}\]\\\|^\{2\}\],\(10\)
Next, we show that, under the second\-order drifting dynamics in equation[9](https://arxiv.org/html/2608.07924#S4.E9), each Fourier mode evolves according to the Nesterov ODE in equation[6](https://arxiv.org/html/2608.07924#S4.E6)\. From the particle dynamics equation[8](https://arxiv.org/html/2608.07924#S4.E8)and kinetic theory\[[50](https://arxiv.org/html/2608.07924#bib.bib50)\], we know the joint density of particle and velocityQt\(x,v\)Q\_\{t\}\(x,v\)satisfies the PDE:
∂tQt\+∇x⋅\(vQt\)\+∇v⋅\(\[Vp,qt\(x\)−αtv\]Qt\)=0\.\\partial\_\{t\}Q\_\{t\}\+\\nabla\_\{x\}\\cdot\(vQ\_\{t\}\)\+\\nabla\_\{v\}\\cdot\\Big\(\\Big\[V\_\{p,q\_\{t\}\}\(x\)\-\\frac\{\\alpha\}\{t\}v\\Big\]Q\_\{t\}\\Big\)=0\.\(11\)Integrating equation[11](https://arxiv.org/html/2608.07924#S4.E11)over allvv, we obtain the standard continuity equation \(cf\. Appendix[C](https://arxiv.org/html/2608.07924#A3)\):
∂tqt\+∇x⋅\(qtut\)=0,\\partial\_\{t\}q\_\{t\}\+\\nabla\_\{x\}\\cdot\(q\_\{t\}u\_\{t\}\)=0,\(12\)whereqt\(x\)=∫Qt\(x,v\)𝑑vq\_\{t\}\(x\)=\\int Q\_\{t\}\(x,v\)dvandut\(x\)=1qt\(x\)∫vQt\(x,v\)𝑑vu\_\{t\}\(x\)=\\frac\{1\}\{q\_\{t\}\(x\)\}\\int vQ\_\{t\}\(x,v\)dv\.
Furthermore, multiply equation[11](https://arxiv.org/html/2608.07924#S4.E11)byvvand integrate overvv, we arrive at the momentum equation \(cf\. Appendix[C](https://arxiv.org/html/2608.07924#A3)\):
∂t\(qtut\)\+αt\(qtut\)=qtVp,qt\(x\)\.\\partial\_\{t\}\(q\_\{t\}u\_\{t\}\)\+\\frac\{\\alpha\}\{t\}\(q\_\{t\}u\_\{t\}\)=q\_\{t\}V\_\{p,q\_\{t\}\}\(x\)\.\(13\)
Letqt=p\+ρtq\_\{t\}=p\+\\rho\_\{t\}\. Sinceppis the equilibrium, its velocityuuand driftVVare zero\. Thus, for smallρt\\rho\_\{t\}, the termsqtut≈putq\_\{t\}u\_\{t\}\\approx pu\_\{t\}andqtVp,qt≈pVp,qtq\_\{t\}V\_\{p,q\_\{t\}\}\\approx pV\_\{p,q\_\{t\}\}\. Moreover, we have already shown that the linearized drift field isVp,qt\(x\)≈−\(M∗ρt\)\(x\)V\_\{p,q\_\{t\}\}\(x\)\\approx\-\(M\*\\rho\_\{t\}\)\(x\)\. Substituting these into equation[12](https://arxiv.org/html/2608.07924#S4.E12)and equation[13](https://arxiv.org/html/2608.07924#S4.E13), we have
∂tρt\+∇x⋅\(put\)\\displaystyle\\partial\_\{t\}\\rho\_\{t\}\+\\nabla\_\{x\}\\cdot\(pu\_\{t\}\)=0\\displaystyle=0∂t\(put\)\+αt\(put\)\\displaystyle\\partial\_\{t\}\(pu\_\{t\}\)\+\\frac\{\\alpha\}\{t\}\(pu\_\{t\}\)=−p\(M∗ρt\)\.\\displaystyle=\-p\(M\*\\rho\_\{t\}\)\.Take the time derivative of the linearized equation[12](https://arxiv.org/html/2608.07924#S4.E12):
∂t2ρt\+∇x⋅∂t\(put\)=0\.\\partial\_\{t\}^\{2\}\\rho\_\{t\}\+\\nabla\_\{x\}\\cdot\\partial\_\{t\}\(pu\_\{t\}\)=0\.Substitute∂t\(put\)\\partial\_\{t\}\(pu\_\{t\}\)from equation[13](https://arxiv.org/html/2608.07924#S4.E13):
∂t2ρt−αt∇x⋅\(put\)⏟−∂tρt−∇x⋅\(p\(M∗ρt\)\)=0\.\\partial\_\{t\}^\{2\}\\rho\_\{t\}\-\\frac\{\\alpha\}\{t\}\\underbrace\{\\nabla\_\{x\}\\cdot\(pu\_\{t\}\)\}\_\{\-\\partial\_\{t\}\\rho\_\{t\}\}\-\\nabla\_\{x\}\\cdot\(p\(M\*\\rho\_\{t\}\)\)=0\.Using the substitution from the first derivative ofρt\\rho\_\{t\}:
∂t2ρt\+αt∂tρt=∇⋅\(p\(M∗ρt\)\)\.\\partial\_\{t\}^\{2\}\\rho\_\{t\}\+\\frac\{\\alpha\}\{t\}\\partial\_\{t\}\\rho\_\{t\}=\\nabla\\cdot\(p\(M\*\\rho\_\{t\}\)\)\.Under the assumption of a locally flat priorpp, we obtain the final form:
∂t2ρt\(x\)\+αt∂tρt\(x\)=∇⋅\(M∗ρt\)\(x\)\.\\partial\_\{t\}^\{2\}\\rho\_\{t\}\(x\)\+\\frac\{\\alpha\}\{t\}\\partial\_\{t\}\\rho\_\{t\}\(x\)=\\nabla\\cdot\(M\*\\rho\_\{t\}\)\(x\)\.Applying Fourier transform to the above equation, we have
∂t2ρ^t\(ω\)\+αt∂tρ^t\(ω\)=λ\(ω\)ρ^t\(ω\)\.\\partial\_\{t\}^\{2\}\\hat\{\\rho\}\_\{t\}\(\\omega\)\+\\frac\{\\alpha\}\{t\}\\partial\_\{t\}\\hat\{\\rho\}\_\{t\}\(\\omega\)=\\lambda\(\\omega\)\\hat\{\\rho\}\_\{t\}\(\\omega\)\.
## 5Numerical Experiments
In this section, we numerically verify the advantages of the proposed second\-order models over other celebrated baseline models\. All experiments were conducted on NVIDIA GeForce RTX 3090 GPUs\. We use Laplacian kernel\-based drift field for all the experiments\.
Table 1:Comparison of KL divergence and MMD between the drifting and second\-order drifting models on the 2D Swiss roll synthetic task\.### 5\.1Synthetic Tasks
The Swiss roll is a smooth but highly folded manifold where high\-frequency structure shows up through rapid changes in direction and curvature\. Even though the density varies smoothly, the folding creates fine\-grained geometric variation that requires tracking sharp changes in trajectory to follow the manifold correctly\. This makes it a useful test for identifying whether a model can preserve high\-frequency behavior\. Additional experiment details are found in Appendix[D\.1](https://arxiv.org/html/2608.07924#A4.SS1)\.
At each training iteration, the model generates20482048samples, from which we estimate the empirical density using histograms and compute the KL divergence to the data distribution\. As shown in Figure[2](https://arxiv.org/html/2608.07924#S5.F2)\(a\), our model exhibits better convergence in terms of KL divergence\. At test time, we apply the same histogram\-based density estimation procedure to the final models and report both KL divergence\[[41](https://arxiv.org/html/2608.07924#bib.bib41)\]and maximum mean discrepancy \(MMD\)\[[42](https://arxiv.org/html/2608.07924#bib.bib42)\]in Table[1](https://arxiv.org/html/2608.07924#S5.T1), where our model consistently outperforms the baseline\. Furthermore, we transform the estimated sample and data densities into the Fourier domain\. Figure[2](https://arxiv.org/html/2608.07924#S5.F2)\(b\) shows that our model achieves smaller Fourier coefficient errors, particularly at high\-frequency modes\.
\(a\)\(b\)\(c\)Figure 1:Swiss Roll experiments: \(a\) Ground truth \(b\) Drifting model samples\. \(v\) Second\-order drifting model samples\. Our model improves frequency fidelity by reducing samples outside theSwiss rollregion\.\(a\)\(b\)Figure 2:Swiss Roll experiments: \(a\) Running\-average KL divergence vs\. training iteration \(5\-run average\)\. \(b\) Fourier coefficient error vs\. frequency for the sampled distributions generated by the drifting model and our second\-order drifting model, averaged over five runs\. We estimate the empirical sample densities using histograms and compute the KL divergence and Fourier coefficient errors based on the resulting density estimates\.
### 5\.2MNIST Image Generation
Table 2:FID of different methods on MNIST dataset\.We evaluate the drifting model and our second\-order drifting model on MNIST image generation\. MNIST consists of28×2828\\times 28grayscale handwritten digit images and is a standard benchmark for evaluating generative models on low\-resolution image distributions\[[26](https://arxiv.org/html/2608.07924#bib.bib26)\]\.
In this experiment, both the baseline drifting model and the proposed second\-order drifting model are implemented with a Diffusion Transformer \(DiT\) backbone\[[36](https://arxiv.org/html/2608.07924#bib.bib36)\]\. At inference time, both models generate samples from Gaussian noise\. We evaluate generation quality using the Fréchet Inception Distance \(FID\)\[[17](https://arxiv.org/html/2608.07924#bib.bib17)\], where a lower FID indicates that the generated distribution is closer to the real MNIST distribution\. The detailed experimental setup is in Appendix[D\.2](https://arxiv.org/html/2608.07924#A4.SS2)\.
Figure 3:Samples generated by \(left\) drifting and \(right\) our second\-order drifting models on MNIST dataset\.Table[2](https://arxiv.org/html/2608.07924#S5.T2)shows that the second\-order drifting model consistently improves over the original drifting model in terms of FID\. Figure[3](https://arxiv.org/html/2608.07924#S5.F3)shows some randomly generated samples by the baseline and our proposed second\-order drifting models\. We further compare the drifting\-based models with small\-NFE flow\-based generative models, including rectified flow\[[29](https://arxiv.org/html/2608.07924#bib.bib29)\], consistency models\[[45](https://arxiv.org/html/2608.07924#bib.bib45)\], and variational rectified flow \(VRFM\)\[[15](https://arxiv.org/html/2608.07924#bib.bib15)\]\.
### 5\.3Sequential Data Generation: Dynamical Systems
We further evaluate our proposed second\-order drifting model on dynamical system trajectory generation tasks\. Sampling trajectories in dynamical systems is a fundamental problem for forecasting and understanding complex time\-dependent phenomena, such as chaotic dynamics, biological oscillations, and extreme events\[[37](https://arxiv.org/html/2608.07924#bib.bib37),[31](https://arxiv.org/html/2608.07924#bib.bib31)\]\. Recent works\[[12](https://arxiv.org/html/2608.07924#bib.bib12),[20](https://arxiv.org/html/2608.07924#bib.bib20)\]have studied diffusion, flow\-matching\-based and 1\-NFE models for event\-guided dynamical system trajectory sampling\.
In this experiment, we formulate trajectory generation for the Lorenz\[[30](https://arxiv.org/html/2608.07924#bib.bib30)\]and FitzHugh–Nagumo\[[13](https://arxiv.org/html/2608.07924#bib.bib13)\]systems as a time\-series generation problem by discretizing the continuous time variabletton a uniform grid, following the experimental setup in\[[12](https://arxiv.org/html/2608.07924#bib.bib12),[20](https://arxiv.org/html/2608.07924#bib.bib20)\]\. Each trajectory is represented as a sequence of states concatenated into𝒙data=\[𝒙\(τm\)\]m=1M∈ℝMd\{\\bm\{x\}\}\_\{\\rm data\}=\[\{\\bm\{x\}\}\(\\tau\_\{m\}\)\]\_\{m=1\}^\{M\}\\in\\mathbb\{R\}^\{Md\}, whereMMdenotes the total number of time steps,ddis the system dimension, and𝒙\(τm\)∈ℝd\{\\bm\{x\}\}\(\\tau\_\{m\}\)\\in\\mathbb\{R\}^\{d\}is the system state at discretized timeτm\\tau\_\{m\}\. Given this representation, our goal is to learn a generative model that can produce realistic dynamical trajectories𝒙data\{\\bm\{x\}\}\_\{\\rm data\}\. For event\-guided trajectory generation, where events are defined by a constraint setE=\{𝒙data∣C\(𝒙data\)\>0\}E=\\\{\{\\bm\{x\}\}\_\{\\rm data\}\\mid C\(\{\\bm\{x\}\}\_\{\\rm data\}\)\>0\\\}\. Figure[4](https://arxiv.org/html/2608.07924#S5.F4)illustrates the event structures for the Lorenz and FitzHugh–Nagumo systems\. We condition the model on event guidance by using the corresponding class label as an additional input without relying on Tweedie’s formula as in prior diffusion\-based approaches\[[12](https://arxiv.org/html/2608.07924#bib.bib12),[20](https://arxiv.org/html/2608.07924#bib.bib20)\]\.
\(a\)\(b\)Figure 4:Events Illustration: \(a\) Lorenz system trajectories; red paths remain in one attractor arm, corresponding to the conditionC\(𝒙\)≥0C\(\{\\bm\{x\}\}\)\\geq 0\. \(b\) FitzHugh–Nagumo trajectories; neuron spike events are shown in red, defined byC\(𝒙\)≥0C\(\{\\bm\{x\}\}\)\\geq 0\.We implement both the original drifting model and our second\-order drifting model for this task Following prior work\[[12](https://arxiv.org/html/2608.07924#bib.bib12),[20](https://arxiv.org/html/2608.07924#bib.bib20)\], we use a U\-Net backbone\[[12](https://arxiv.org/html/2608.07924#bib.bib12)\]to model the trajectory distribution\. Model performance is evaluated using the KL divergence between the histogram\-based density estimates of the event\-value distributionsC\(𝒙\)C\(\{\\bm\{x\}\}\)for the sampled and dataset trajectories\. The detailed experimental setup is in Appendix[D\.3](https://arxiv.org/html/2608.07924#A4.SS3)\.
Table[3](https://arxiv.org/html/2608.07924#S5.T3)shows that our second\-order drifting model consistently improves trajectory generation quality over the original drifting model\. In particular, the second\-order update leads to lower distributional discrepancy under KL divergence, demonstrating its advantage in modeling complex dynamical system trajectories\.
Table 3:KL divergence between the eventC\(𝒙\)C\(\{\\bm\{x\}\}\)distributions of the generated trajectories and the dataset trajectories, estimated from histogram\-based density approximations, with/without conditioning on the event\.
### 5\.4Robotic Control
Finally we consider the robotics tasks from\[[10](https://arxiv.org/html/2608.07924#bib.bib10),[9](https://arxiv.org/html/2608.07924#bib.bib9)\]\. Robotics manipulation is a sequential decision\-making problem characterized by long\-horizon dependencies, contact\-rich interactions, and multi\-modal action distributions\. In such settings, control policies must maintain temporal consistency while adapting rapidly to changing object dynamics and task constraints\.
We evaluate on a diverse suite of robotic manipulation benchmarks including Can, Tool Hang, Lift, and PushT, as well as multi\-stage tasks such as BlockPush and Kitchen following the evaluation protocol of Drifting Policy\[[10](https://arxiv.org/html/2608.07924#bib.bib10),[9](https://arxiv.org/html/2608.07924#bib.bib9)\]\. Following a similar experimental setup as that in\[[10](https://arxiv.org/html/2608.07924#bib.bib10)\], we replace the diffusion\-based action generator with our second\-order drifting model while retaining the same state\-action representation and training pipeline\. The detailed experimental setup is in Appendix[D\.4](https://arxiv.org/html/2608.07924#A4.SS4)\. Table[4](https://arxiv.org/html/2608.07924#S5.T4)shows that we achieve comparable or higher success rates across tasks while requiring substantially fewer training epochs\. We compare our results with previous state\-of\-the\-art models namely Diffusion Policy\[[9](https://arxiv.org/html/2608.07924#bib.bib9)\]and Drifting policy\[[10](https://arxiv.org/html/2608.07924#bib.bib10)\]\. Specifically, second\-order updates lead to significantly increased training efficiency on robotics tasks\.
We further study different initializations for the initial velocityv0\(i\)v^\{\(i\)\}\_\{0\}, including learned initialization, zero initialization, and historical warm\-starting\. We observe that zero initialization \(v0\(i\)=0v\_\{0\}^\{\(i\)\}=0\) provides the best trade\-off between performance and efficiency, while historical warm\-starting leads to instability due to non\-stationary optimization dynamics\. Learned initializations achieved the best success rates, but doubled the training time\. Further details are provided in Appendix[D\.5\.1](https://arxiv.org/html/2608.07924#A4.SS5.SSS1)\.
Diffusion PolicyDriftingDrifting 2nd \(ours\)TaskSettingSuccess \(↑\\uparrow\)Epochs \(↓\\downarrow\)Success \(↑\\uparrow\)Epochs \(↓\\downarrow\)Success \(↑\\uparrow\)Epochs \(↓\\downarrow\)CanVisual\.973050\.9950\.99±0\.016\\mathbf\{\.99\_\{\\pm 0\.016\}\}16State\.965000\.9850\.992±0\.01\\mathbf\{\.992\_\{\\pm 0\.01\}\}20Tool HangVisual\.773050\.6725\.79±\.051\\mathbf\{\.79\_\{\\pm\.051\}\}16State\.305000\.3850\.71±\.06\\mathbf\{\.71\_\{\\pm\.06\}\}30LiftVisual1\.030501\.0501\.0±0\.0\\mathbf\{1\.0\_\{\\pm 0\.0\}\}8State\.9850001\.0501\.0±0\.0\\mathbf\{1\.0\_\{\\pm 0\.0\}\}25PushTVisual\.843050\.86100\.87±\.018\\textbf\{\.87\}\_\{\\pm\.018\}100State\.915000\.8650\.872±\.023\.872\_\{\\pm\.023\}30BlockPushPhase 10\.3650000\.565000\.372±\.061\.372\_\{\\pm\.061\}3500Phase 20\.1150000\.165000\.168±\.041\\mathbf\{\.168\_\{\\pm\.041\}\}3500KitchenPhase 11\.0050001\.003001\.00±0\.0\\mathbf\{1\.00\_\{\\pm 0\.0\}\}100Phase 21\.0050001\.003001\.00±0\.0\\mathbf\{1\.00\_\{\\pm 0\.0\}\}100Phase 31\.0050000\.993001\.00±0\.0\\mathbf\{1\.00\_\{\\pm 0\.0\}\}100Phase 40\.9950000\.96300\.97±\.015\{\.97\_\{\\pm\.015\}\}100Table 4:Success rates \(↑\) and training epochs \(↓\) across tasks for diffusion policy and drifting variants\.
## 6Conclusion
We proposed second\-order drifting models, a momentum\-based extension of drifting that mitigates the spectral bias of kernel\-driven dynamics\. By lifting the evolution into phase space, we showed that the linearized dynamics follow accelerated second\-order systems, reducing the dependence of convergence rates on the kernel spectrum\. This provides a principled acceleration mechanism consistent with the interpretation of drifting as score\-based gradient flow\. We further introduced a semi\-implicit training scheme that preserves one\-step inference while improving convergence and stability\. Experiments across multiple domains demonstrate consistent gains over first\-order drifting\. Overall, second\-order drifting offers a simple and effective approach to accelerating one\-step generative modeling\. As a limitation, the current formulation is primarily designed to accelerate the training dynamics of drifting models, and its applicability to broader classes of generative frameworks remains an important direction for future work\.
## References
- \[1\]Michael S Albergo, Nicholas M Boffi, and Eric Vanden\-Eijnden\.Stochastic interpolants: A unifying framework for flows and diffusions\.arXiv preprint arXiv:2303\.08797, 2023\.
- \[2\]Hédy Attouch, Zaki Chbani, Juan Peypouquet, and Patrick Redont\.Fast convergence of inertial dynamics and algorithms with asymptotic vanishing viscosity\.Mathematical Programming, 168\(1–2\):123–175, 2018\.
- \[3\]Hédy Attouch, Juan Peypouquet, and Patrick Redont\.Fast convex optimization via inertial dynamics with Hessian driven damping\.Journal of Differential Equations, 261\(10\):5734–5783, 2016\.
- \[4\]Hedy Attouch and Benar Fux Svaiter\.A continuous dynamical newton\-like approach to solving monotone inclusions\.SIAM Journal on Control and Optimization, 49\(2\):574–598, 2011\.
- \[5\]Nicholas M Boffi, Michael S Albergo, and Eric Vanden\-Eijnden\.Flow map matching\.arXiv preprint arXiv:2406\.07507, 2024\.
- \[6\]Nicholas M Boffi, Michael S Albergo, and Eric Vanden\-Eijnden\.How to build a consistency model: Learning flow maps via self\-distillation\.arXiv preprint arXiv:2505\.18825, 2025\.
- \[7\]Jiarui Cao, Zixuan Wei, and Yuxin Liu\.Gradient flow drifting: Generative modeling via wasserstein gradient flows of kde\-approximated divergences\.arXiv preprint arXiv:2603\.10592, 2026\.
- \[8\]José A Carrillo, Young\-Pil Choi, and Jinwook Jung\.Quantifying the hydrodynamic limit of vlasov\-type equations with alignment and nonlocal forces\.Mathematical Models and Methods in Applied Sciences, 31\(02\):327–408, 2021\.
- \[9\]Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song\.Diffusion policy: Visuomotor policy learning via action diffusion\.The International Journal of Robotics Research, 2024\.
- \[10\]Mingyang Deng, He Li, Tianhong Li, Yilun Du, and Kaiming He\.Generative modeling via drifting\.arXiv preprint arXiv:2602\.04770, 2026\.
- \[11\]Tim Dockhorn, Arash Vahdat, and Karsten Kreis\.Score\-based generative modeling with critically\-damped Langevin diffusion\.InInternational Conference on Learning Representations \(ICLR\), 2022\.
- \[12\]Marc Anton Finzi, Anudhyan Boral, Andrew Gordon Wilson, Fei Sha, and Leonardo Zepeda\-Núñez\.User\-defined event sampling and uncertainty quantification in diffusion models for physical dynamical systems\.InInternational Conference on Machine Learning, pages 10136–10152\. PMLR, 2023\.
- \[13\]Richard FitzHugh\.Impulses and physiological states in theoretical models of nerve membrane\.Biophysical journal, 1\(6\):445–466, 1961\.
- \[14\]Zhengyang Geng, Mingyang Deng, Xingjian Bai, J Zico Kolter, and Kaiming He\.Mean flows for one\-step generative modeling\.arXiv preprint arXiv:2505\.13447, 2025\.
- \[15\]Pengsheng Guo and Alexander G Schwing\.Variational rectified flow matching\.arXiv preprint arXiv:2502\.09616, 2025\.
- \[16\]Ping He, Om Khangaonkar, Hamed Pirsiavash, Yikun Bai, and Soheil Kolouri\.Sinkhorn\-drifting generative models\.arXiv preprint arXiv:2603\.12366, 2026\.
- \[17\]Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter\.Gans trained by a two time\-scale update rule converge to a local nash equilibrium\.InAdvances in Neural Information Processing Systems, 2017\.
- \[18\]Jonathan Ho, Ajay Jain, and Pieter Abbeel\.Denoising diffusion probabilistic models\.Advances in neural information processing systems, 33:6840–6851, 2020\.
- \[19\]Emiel Hoogeboom, Vıctor Garcia Satorras, Clément Vignac, and Max Welling\.Equivariant diffusion for molecule generation in 3d\.InInternational conference on machine learning, pages 8867–8887\. PMLR, 2022\.
- \[20\]Yuhao Huang, Taos Transue, Shih\-Hsin Wang, William Feldman, Hong Zhang, and Bao Wang\.Improving flow matching by aligning flow divergence\.arXiv preprint arXiv:2602\.00869, 2026\.
- \[21\]Yuhao Huang, Shih\-Hsin Wang, Andrea L Bertozzi, and Bao Wang\.Rmflow: Refined mean flow by a noise\-injection step for multimodal generation\.arXiv preprint arXiv:2602\.00849, 2026\.
- \[22\]Fan Jia, Yuhao Huang, Shih\-Hsin Wang, Cristina Garcia\-Cardona, Andrea L Bertozzi, and Bao Wang\.Plug\-and\-play image restoration with flow matching: A continuous viewpoint\.arXiv preprint arXiv:2512\.04283, 2025\.
- \[23\]Arkadii Kazanskii, Tatiana Petrova, Konstantin Bagrianskii, Aleksandr Puzikov, and Radu State\.Attraction, repulsion, and friction: Introducing dmf, a friction\-augmented drifting model\.arXiv preprint arXiv:2604\.18194, 2026\.
- \[24\]Nikola B Kovachki and Andrew M Stuart\.Continuous time analysis of momentum methods\.Journal of Machine Learning Research, 22\(17\):1–40, 2021\.
- \[25\]Chieh\-Hsin Lai, Bac Nguyen, Naoki Murata, Yuhta Takida, Toshimitsu Uesaka, Yuki Mitsufuji, Stefano Ermon, and Molei Tao\.A unified view of drifting and score\-based models\.arXiv preprint arXiv:2603\.07514, 2026\.
- \[26\]Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner\.Gradient\-based learning applied to document recognition\.Proceedings of the IEEE, 86\(11\):2278–2324, 2002\.
- \[27\]Zhiqi Li and Bo Zhu\.A long\-short flow\-map perspective for drifting models\.arXiv preprint arXiv:2602\.20463, 2026\.
- \[28\]Yaron Lipman, Ricky T\. Q\. Chen, Heli Ben\-Hamu, Maximilian Nickel, and Matthew Le\.Flow matching for generative modeling\.InThe Eleventh International Conference on Learning Representations, 2023\.
- \[29\]Xingchao Liu, Chengyue Gong, and qiang liu\.Flow straight and fast: Learning to generate and transfer data with rectified flow\.InThe Eleventh International Conference on Learning Representations, 2023\.
- \[30\]Edward N Lorenz\.Deterministic nonperiodic flow\.Journal of the Atmospheric Sciences, 20\(2\):130–141, 1963\.
- \[31\]Amir Mosavi, Pinar Ozturk, and Kwok\-wing Chau\.Flood prediction using machine learning models: Literature review\.Water, 10\(11\):1536, 2018\.
- \[32\]Yurii Nesterov\.A method for unconstrained convex minimization problem with the rate of convergence o \(1/k2\)\.InDokl\. Akad\. Nauk\. SSSR, volume 269, page 543, 1983\.
- \[33\]Ho Huu Nghia Nguyen, Tan Nguyen, Huyen Vo, Stanley Osher, and Thieu Vo\.Improving neural ordinary differential equations with nesterov’s accelerated gradient method\.Advances in Neural Information Processing Systems, 35:7712–7726, 2022\.
- \[34\]Tan Nguyen, Richard Baraniuk, Andrea Bertozzi, Stanley Osher, and Bao Wang\.Momentumrnn: Integrating momentum into recurrent neural networks\.Advances in neural information processing systems, 33:1924–1936, 2020\.
- \[35\]Alexander Norcliffe, Cristian Bodnar, Ben Day, Nikola Simidjievski, and Pietro Liò\.On second order behaviour in augmented neural odes\.Advances in neural information processing systems, 33:5911–5921, 2020\.
- \[36\]William Peebles and Saining Xie\.Scalable diffusion models with transformers\.InProceedings of the IEEE/CVF international conference on computer vision, pages 4195–4205, 2023\.
- \[37\]Sarah E Perkins and Lisa V Alexander\.On the measurement of heat waves\.Journal of climate, 26\(13\):4500–4517, 2013\.
- \[38\]Boris T Polyak\.Some methods of speeding up the convergence of iteration methods\.Ussr computational mathematics and mathematical physics, 4\(5\):1–17, 1964\.
- \[39\]Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer\.High\-resolution image synthesis with latent diffusion models\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022\.
- \[40\]Bin Shi, Simon S Du, Weijie Su, and Michael I Jordan\.Acceleration via symplectic discretization of high\-resolution differential equations\.Advances in Neural Information Processing Systems, 32, 2019\.
- \[41\]Jonathon Shlens\.Notes on kullback\-leibler divergence and likelihood\.arXiv preprint arXiv:1404\.2000, 2014\.
- \[42\]Alexander J Smola, A Gretton, and K Borgwardt\.Maximum mean discrepancy\.In13th international conference, ICONIP, volume 6, 2006\.
- \[43\]Jascha Sohl\-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli\.Deep unsupervised learning using nonequilibrium thermodynamics\.InInternational conference on machine learning, pages 2256–2265\. pmlr, 2015\.
- \[44\]Yang Song and Prafulla Dhariwal\.Improved techniques for training consistency models\.InThe Twelfth International Conference on Learning Representations, 2024\.
- \[45\]Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever\.Consistency models\.arXiv preprint arXiv:2303\.01469, 2023\.
- \[46\]Yang Song, Jascha Sohl\-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole\.Score\-based generative modeling through stochastic differential equations\.InInternational Conference on Learning Representations, 2021\.
- \[47\]Weijie Su, Stephen Boyd, and Emmanuel J Candes\.A differential equation for modeling nesterov’s accelerated gradient method: Theory and insights\.Journal of Machine Learning Research, 17\(153\):1–43, 2016\.
- \[48\]Taos Transue, Bohan Chen, So Takao, and Bao Wang\.Flow matching for efficient and scalable data assimilation\.arXiv preprint arXiv:2508\.13313, 2025\.
- \[49\]Erkan Turan and Maks Ovsjanikov\.Generative drifting is secretly score matching: a spectral and variational perspective\.arXiv preprint arXiv:2603\.09936, 2026\.
- \[50\]Cédric Villani\.A review of mathematical topics in collisional kinetic theory\.Handbook of mathematical fluid dynamics, 1:71–74, 2002\.
- \[51\]Bao Wang, Hedi Xia, Tan Nguyen, and Stanley Osher\.How does momentum benefit deep neural networks architecture design? a few case studies\.Research in the Mathematical Sciences, 9\(3\):57, 2022\.
- \[52\]Joseph L Watson, David Juergens, Nathaniel R Bennett, Brian L Trippe, Jason Yim, Helen E Eisenach, Woody Ahern, Andrew J Borst, Robert J Ragotte, Lukas F Milles, Basile I\. M\. Wicky, Nikita Hanikel, Samuel J\. Pellock, Alexis Courbet, William Sheffler, Jue Wang, Preetham Venkatesh, Isaac Sappington, Susana Vázquez Torres, Anna Lauko, Valentin De Bortoli, Emile Mathieu, Sergey Ovchinnikov, Regina Barzilay, Tommi S\. Jaakkola, Frank DiMaio, Minkyung Baek, and David Baker\.De novo design of protein structure and function with RFdiffusion\.Nature, 620\(7976\):1089–1100, 2023\.
- \[53\]Andre Wibisono, Ashia C Wilson, and Michael I Jordan\.A variational perspective on accelerated methods in optimization\.proceedings of the National Academy of Sciences, 113\(47\):E7351–E7358, 2016\.
- \[54\]Hedi Xia, Vai Suliafu, Hangjie Ji, Tan Nguyen, Andrea Bertozzi, Stanley Osher, and Bao Wang\.Heavy ball neural ordinary differential equations\.Advances in Neural Information Processing Systems, 34:18646–18659, 2021\.
- \[55\]Jiaru Zhang, Zeyun Deng, Juanwu Lu, Ziran Wang, and Ruqi Zhang\.Analytical correction for subsampling bias in drifting models\.arXiv preprint arXiv:2604\.27239, 2026\.
## Appendix ALinearized PDE
To derive the linearized PDE forρt=qt−p\\rho\_\{t\}=q\_\{t\}\-p, we start from the continuous\-limit evolution of the drifting model and linearize it around the equilibrium distributionpp\. The evolution of the model densityqt\(x\)q\_\{t\}\(x\)is governed by:
∂tqt\(x\)\+∇⋅\(qt\(x\)Vp,qt\(x\)\)=0,\\partial\_\{t\}q\_\{t\}\(x\)\+\\nabla\\cdot\\big\(q\_\{t\}\(x\)V\_\{p,q\_\{t\}\}\(x\)\\big\)=0,\(14\)where the drift field is defined as:
Vp,qt=Vp\(x\)−Vqt\(x\),V\_\{p,q\_\{t\}\}=V\_\{p\}\(x\)\-V\_\{q\_\{t\}\}\(x\),withVp\(x\)=𝔼y∼p\[k\(x,y\)\(y−x\)\]𝔼y∼p\[k\(x,y\)\]V\_\{p\}\(x\)=\\frac\{\\mathbb\{E\}\_\{y\\sim p\}\[k\(x,y\)\(y\-x\)\]\}\{\\mathbb\{E\}\_\{y\\sim p\}\[k\(x,y\)\]\}andVqtV\_\{q\_\{t\}\}is defined similarly\.
Substitutingqt=p\+ρtq\_\{t\}=p\+\\rho\_\{t\}into equation[14](https://arxiv.org/html/2608.07924#A1.E14):
- •Time derivative:∂t\(p\+ρt\)=∂tρt\\partial\_\{t\}\(p\+\\rho\_\{t\}\)=\\partial\_\{t\}\\rho\_\{t\}sinceppis stationary\.
- •Drift at equilibrium: at equilibrium \(qt=pq\_\{t\}=p\), the drift field vanishes, i\.e\.,Vp,p\(x\)=Vp\(x\)−Vp\(x\)=0V\_\{p,p\}\(x\)=V\_\{p\}\(x\)\-V\_\{p\}\(x\)=0\.
- •Expansion ofVqtV\_\{q\_\{t\}\}: we expandVqt\(x\)V\_\{q\_\{t\}\}\(x\)to the first order inρt\\rho\_\{t\}\. LetDρ\(x\)=∫k\(x,y\)p\(y\)𝑑yD\_\{\\rho\}\(x\)=\\int k\(x,y\)p\(y\)dybe the normalization factor\. Then Vp\+ρt\(x\)≈Vp\(x\)\+1Dp\(x\)∫k\(x,y\)\[\(y−x\)−Vp\(x\)\]ρt\(y\)𝑑y\.V\_\{p\+\\rho\_\{t\}\}\(x\)\\approx V\_\{p\}\(x\)\+\\frac\{1\}\{D\_\{p\}\(x\)\}\\int k\(x,y\)\[\(y\-x\)\-V\_\{p\}\(x\)\]\\rho\_\{t\}\(y\)dy\.Thus, the relative drift fieldVp,qtV\_\{p,q\_\{t\}\}becomes: Vp,p\+ρt\(x\)=Vp\(x\)−Vp\+ρt\(x\)≈−1Dp\(x\)∫k\(x,y\)\[\(y−x\)−Vp\(x\)\]ρt\(y\)𝑑y\.V\_\{p,p\+\\rho\_\{t\}\}\(x\)=V\_\{p\}\(x\)\-V\_\{p\+\\rho\_\{t\}\}\(x\)\\approx\-\\frac\{1\}\{D\_\{p\}\(x\)\}\\int k\(x,y\)\[\(y\-x\)\-V\_\{p\}\(x\)\]\\rho\_\{t\}\(y\)dy\.
Substitute the above linearized drift back into the continuity equation and keep only terms linear inρt\\rho\_\{t\}:
∂tρt\(x\)\\displaystyle\\partial\_\{t\}\\rho\_\{t\}\(x\)\+∇⋅\(p\(x\)⋅Vp,p\+ρt\(x\)\)≈0,\\displaystyle\+\\nabla\\cdot\\big\(p\(x\)\\cdot V\_\{p,p\+\\rho\_\{t\}\}\(x\)\\big\)\\approx 0,\(15\)∂tρt\(x\)\\displaystyle\\partial\_\{t\}\\rho\_\{t\}\(x\)=∇⋅\(p\(x\)Dp\(x\)∫k\(x,y\)\[\(y−x\)−Vp\(x\)\]ρt\(y\)𝑑y\)\.\\displaystyle=\\nabla\\cdot\\Big\(\\frac\{p\(x\)\}\{D\_\{p\}\(x\)\}\\int k\(x,y\)\[\(y\-x\)\-V\_\{p\}\(x\)\]\\rho\_\{t\}\(y\)dy\\Big\)\.
Furthermore, we make the following assumptions:
- •Radial symmetry kernel:k\(x,y\)=K\(‖x−y‖\)k\(x,y\)=K\(\\\|x\-y\\\|\)\.
- •We assumep\(x\)p\(x\)is approximately constant relative to the scale of the kernel\. This implies: \(1\)Vp\(x\)≈0V\_\{p\}\(x\)\\approx 0—no local bias in the target distribution\. \(2\)p\(x\)Dp\(x\)≈c\\frac\{p\(x\)\}\{D\_\{p\}\(x\)\}\\approx c, whereccis a constant, typically1∫K\(‖z‖\)𝑑z\\frac\{1\}\{\\int K\(\\\|z\\\|\)dz\}\.
Letz=x−yz=x\-y, then equation[15](https://arxiv.org/html/2608.07924#A1.E15)simplifies to
∂tρt\(x\)=∇⋅\(c∫K\(‖x−y‖\)\(x−y\)ρt\(y\)𝑑y\)=∇⋅\(M∗ρt\)\(x\)\.\\partial\_\{t\}\\rho\_\{t\}\(x\)=\\nabla\\cdot\\Big\(c\\int K\(\\\|x\-y\\\|\)\(x\-y\)\\rho\_\{t\}\(y\)dy\\Big\)=\\nabla\\cdot\(M\*\\rho\_\{t\}\)\(x\)\.
## Appendix BSpectral Estimation
###### Proposition B\.1\.
Letk\(x,y\)=K\(‖x−y‖\)k\(x,y\)=K\(\\\|x\-y\\\|\)be a positive, radial symmetric kerne withKKbeing integrable and decaying in‖x−y‖\\\|x\-y\\\|, then
λ\(ω\)=ω⋅∇ωK^\(ω\)≤0,\\lambda\(\\omega\)=\\omega\\cdot\\nabla\_\{\\omega\}\\hat\{K\}\(\\omega\)\\leq 0,whereK^\(ω\)\\hat\{K\}\(\\omega\)is the Fourier transform ofKK\.
###### Proof\.
We will show thatλ\(ω\)≤0\\lambda\(\\omega\)\\leq 0for allω\\omega, andλ\(ω\)=0\\lambda\(\\omega\)=0if and only ifω=0\\omega=0\.
By definition, we have
K^\(ω\)=∫ℝdK\(z\)e−iω⋅z𝑑z,\\hat\{K\}\(\\omega\)=\\int\_\{\\mathbb\{R\}^\{d\}\}K\(z\)e^\{\-i\\omega\\cdot z\}dz,which is real and even sinceKKis real and even\. Moreover, taking absolute values for the above equation and by triangle inequality, we have
\|K^\(ω\)\|≤∫K\(z\)\|e−ω⋅z\|𝑑z=∫K\(z\)𝑑z=K^\(0\)\.\|\\hat\{K\}\(\\omega\)\|\\leq\\int K\(z\)\|e^\{\-\\omega\\cdot z\}\|dz=\\int K\(z\)dz=\\hat\{K\}\(0\)\.That is,ω=0\\omega=0is a global maximum\. Moreover,K^\(ω\)=ϕ\(\|ω\|\)\\hat\{K\}\(\\omega\)=\\phi\(\|\\omega\|\)for some scalar functionϕ\(r\)\\phi\(r\)withϕ′\(r\)≤0\\phi^\{\\prime\}\(r\)\\leq 0forr\>0r\>0\. Therefore, we have
ω⋅∇ωK^\(ω\)=ω⋅\(ϕ′\(\|ω\|\)ω\|ω\|\)=\|ω\|ϕ′\(\|ω\|\)≤0,\\omega\\cdot\\nabla\_\{\\omega\}\\hat\{K\}\(\\omega\)=\\omega\\cdot\\Big\(\\phi^\{\\prime\}\(\|\\omega\|\)\\frac\{\\omega\}\{\|\\omega\|\}\\Big\)=\|\\omega\|\\phi^\{\\prime\}\(\|\\omega\|\)\\leq 0,with equality only atω=0\\omega=0\.
∎
## Appendix CPDE Limit for the Second\-Order Drifting Models
### C\.1Mass Conservation
We define the marginal densityqt\(x\)=∫Qt\(x,v\)𝑑vq\_\{t\}\(x\)=\\int Q\_\{t\}\(x,v\)dvand the local mean velocityut\(x\)=1qt\(x\)∫vQt\(x,v\)𝑑vu\_\{t\}\(x\)=\\frac\{1\}\{q\_\{t\}\(x\)\}\\int vQ\_\{t\}\(x,v\)dv\. Integrating the phase space equation over allvv:
- •1\.∫∂tQtdv=∂tqt\(x\)\\int\\partial\_\{t\}Q\_\{t\}dv=\\partial\_\{t\}q\_\{t\}\(x\)
- •2\.∫∇x⋅\(vQt\)𝑑v=∇x⋅\(qt\(x\)ut\(x\)\)\\int\\nabla\_\{x\}\\cdot\(vQ\_\{t\}\)dv=\\nabla\_\{x\}\\cdot\(q\_\{t\}\(x\)u\_\{t\}\(x\)\)
- •3\. The∇v\\nabla\_\{v\}term vanishes by the divergence theorem \(assumingQt→0Q\_\{t\}\\to 0as\|v\|→∞\|v\|\\to\\infty\)
This gives the standard continuity equation:
∂tqt\+∇x⋅\(qtut\)=0\.\\partial\_\{t\}q\_\{t\}\+\\nabla\_\{x\}\\cdot\(q\_\{t\}u\_\{t\}\)=0\.\(16\)
### C\.2Momentum Balance
Multiply the kinetic equation byvvand integrate overvv:
- •∫v∂tQtdv=∂t\(qtut\)\\int v\\partial\_\{t\}Q\_\{t\}dv=\\partial\_\{t\}\(q\_\{t\}u\_\{t\}\)
- •∫v\[v⋅∇xQt\]𝑑v=∇x⋅∫\(v⊗v\)Qt𝑑v≈0\\int v\[v\\cdot\\nabla\_\{x\}Q\_\{t\}\]dv=\\nabla\_\{x\}\\cdot\\int\(v\\otimes v\)Q\_\{t\}dv\\approx 0\(assuming low temperature/pressure where the velocity is negligible relative to the mean flow\)\[[8](https://arxiv.org/html/2608.07924#bib.bib8)\]\.
- •∫v\[∇v⋅\(VQt\)\]𝑑v=−Vp,qt\(x\)qt\(x\)\\int v\[\\nabla\_\{v\}\\cdot\(VQ\_\{t\}\)\]dv=\-V\_\{p,q\_\{t\}\}\(x\)q\_\{t\}\(x\)\(integration by parts\)\.
- •∫v\[∇v⋅\(−αtvQt\)\]𝑑v=αtqtut\\int v\[\\nabla\_\{v\}\\cdot\(\-\\frac\{\\alpha\}\{t\}vQ\_\{t\}\)\]dv=\\frac\{\\alpha\}\{t\}q\_\{t\}u\_\{t\}\(integration by parts\)
This gives the momentum equation:
∂t\(qtut\)\+αt\(qtut\)=qtVp,qt\(x\)\.\\partial\_\{t\}\(q\_\{t\}u\_\{t\}\)\+\\frac\{\\alpha\}\{t\}\(q\_\{t\}u\_\{t\}\)=q\_\{t\}V\_\{p,q\_\{t\}\}\(x\)\.\(17\)
## Appendix DExperimental Setup
### D\.1Implementation Details for 2D Synthetic
#### D\.1\.1Hyperparameters of Drifting Models
Table[5](https://arxiv.org/html/2608.07924#A4.T5)shows the hyperparameters used for the second\-order drifting model on the 2D Swiss roll trajectory sampling task\. Our model utilizes a Nesterov ODE\-inspired second\-order formulation\.
Table 5:Hyperparameters of the second\-order drifting model on the 2D Swiss roll task\.
#### D\.1\.2Dataset and Training Setup
The 2D Swiss roll dataset consists of 2048 samples with additive Gaussian noise of standard deviation0\.030\.03\. The task is purely two\-dimensional and no normalization or scaling is applied to the data\. The Swiss roll distribution contains highly folded local structure and sharp changes in density along the manifold, making it useful for evaluating a model’s ability to capture fine\-scale geometric structure during trajectory sampling\.
Models are trained for 10,000 epochs using AdamW with a linearly annealed learning rate schedule\. We do not use exponential moving average \(EMA\) during training\.
#### D\.1\.3Model Architecture
We use a residual MLP backbone for all 2D synthetic experiments\. The network operates on a 32\-dimensional latent representation with hidden dimension 256 and consists of six residual blocks\. Each block applies LayerNorm, followed by a SiLU activation and a linear projection within a residual connection\. The final output is scaled to remain within the range\[−4,4\]\[\-4,4\]\.
ResidualMLPBlock\(dd\)LayerNorm\(dd\)Linear\(d→dd\\rightarrow d\)SiLULinear\(d→dd\\rightarrow d\)ResidualConnection
Residual MLP ArchitectureInput Projection\(2→322\\rightarrow 32\)Hidden Projection\(32→25632\\rightarrow 256\)ResidualMLPBlock\(256256\)×6\\times 6Output Projection\(256→2256\\rightarrow 2\)Output Scaling\(\[−4,4\]\[\-4,4\]\)
Figure 5:Residual MLP backbone architecture used for the 2D Swiss roll experiments\.
### D\.2Experiment Details of the MNIST Task
#### D\.2\.1Hyperparameters of Drifting Models
Table[6](https://arxiv.org/html/2608.07924#A4.T6)shows the details of our second\-order drifting model for dynamical system trajectory sampling tasks\. Our model utilizes a Nesterov ODE\-inspired second\-order formulation\.
Table 6:Hyperparameters of the Second\-order drifting model on MNIST tasks\.
#### D\.2\.2Model and Training Setup
Model Architecture: We use diffusion transformer as backbone to learn the models\. Figure[6](https://arxiv.org/html/2608.07924#A4.F6)shows the detailed model architecture\.
Training Setup\. We train both models with a batch size of256256\. All models are trained for200200with learning rate2×10−42\\times 10^\{\-4\}\. To evaluate performance, we sample1000010000images from each model and1000010000images from the real dataset, then compute the FID score\.
Resource Usage and Time\. We run all experiments on a single RTX 3090 GPU\. Each training run requires approximately 8 GiB of GPU memory and takes around 10 hours\.
DiTBlock\(dd\):RMSNorm\(dd\)MultiHeadAttention\(heads=44, QK\-Norm, RoPE\)adaLN\-Zero modulationResidualConnectionRMSNorm\(dd\)SwiGLU\(mlp\_ratio=44\)adaLN\-Zero modulationResidualConnection
DriftDiT\-Tiny \(9M\) Architecture:PatchEmbed\(img=3232, patch=44, in=11,d=256d=256\)RegisterTokens\(88\)LabelEmbed\(10→25610\\to 256\)AlphaEmbed\(Fourier→\\toMLP\)StyleEmbed\(tokens=3232, codebook=6464\)DiTBlock\(dd\)×6\\times 6FinalLayer\(dd, patch=44, out=11\)Unpatchify\(32×32×132\\times 32\\times 1\)
Figure 6:DiT backbone architecture and the corresponding residual block for MNIST tasks\.
### D\.3Implementation Details for Dynamical Systems
This section presents the implementation details for Lorenz and FitzHugh–Nagumo trajectory sampling, including the second\-order drifting hyperparameters, the shared U\-Net backbone and the training settings\.
#### D\.3\.1Hyperparameters of Drifting Models
Table[7](https://arxiv.org/html/2608.07924#A4.T7)shows the details of our second\-order drifting model for dynamical system trajectory sampling tasks\. Our model utilizes a Nesterov\-inspired second\-order formulation\.
Table 7:Hyperparameters of the Second\-order drifting model on Dynamical system tasks\.
#### D\.3\.2Model and Training Setup
Model Architecture\. We use the same U\-Net backbone as described in Section H of\[[12](https://arxiv.org/html/2608.07924#bib.bib12)\]to parameterize both the drifting model and our second\-order drifting model\. Figure[7](https://arxiv.org/html/2608.07924#A4.F7)shows the architecture\.
ResBlock\(cc\):GroupNorm\(groups=c/4c/4\)SwishConv\(channels=33, ksize=33\)GroupNorm\(groups=c/4c/4\)SwishConv\(channels=33, ksize=33\)SkipConnection
Convolutional U\-Net Architecture:ResBlock\(cc\)×4\\times 4Downsample\(22\)ResBlock\(2c2c\)×8\\times 8Downsample\(22\)ResBlock\(4c4c\)×8\\times 8SkipResBlock\(4c4c\)×8\\times 8Upsample\(22\)SkipResBlock\(2c2c\)×8\\times 8Upsample\(22\)SkipResBlock\(cc\)×4\\times 4Conv\(128128\)Conv\(dd\)
Figure 7:U\-Net backbone architecture and the corresponding residual block for Dynamical system tasks\.Training Setup\. For each task, we train both models with a batch size of256256using80008000trajectories generated bytorchdiffeq, following\[[12](https://arxiv.org/html/2608.07924#bib.bib12)\]\. All models are trained for30003000epochs using a linearly decayed learning rate initialized at2×10−42\\times 10^\{\-4\}\. To evaluate performance, we sample20002000trajectories from each model, estimate the distribution of the constraint valueC\(𝒙\)C\(\{\\bm\{x\}\}\)using histograms, and compute the KL divergence between the generated and dataset distributions\.
Resource Usage and Time\. We run all experiments on a single RTX 3090 GPU\. Each training run requires approximately 12 GiB of GPU memory and takes around 20 hours\.
### D\.4Implementation Details for Robotics
#### D\.4\.1Hyperparameters for Robotics experiments
We summarize the configuration of our second\-order drifting model across all benchmark tasks\. The model is based on a Nesterov\-inspired second\-order update with semi\-implicit discretization\. Full per\-task configurations are provided in[Table˜8](https://arxiv.org/html/2608.07924#A4.T8)\.
Table 8:Hyperparameters used in robotics experimentsTable 9:Per\-task base channel widthccand its scaled variants2c2cand4c4c, used at successive downsampling stages of the Conditional 1D U\-Net \(Figure[10](https://arxiv.org/html/2608.07924#A4.T10)\)\.ConditionalResBlock\(c,c′,gc,\\;c^\{\\prime\},\\;g\)Conv1d\(c→c′c\\rightarrow c^\{\\prime\}\),k=5k\{=\}5GroupNorm\(c′/8c^\{\\prime\}/8\) \+ MishFiLM\(g→c′g\\rightarrow c^\{\\prime\}\)Conv1d\(c′→c′c^\{\\prime\}\\rightarrow c^\{\\prime\}\),k=5k\{=\}5GroupNorm\(c′/8c^\{\\prime\}/8\) \+ MishSkipConnection\(c→c′c\\rightarrow c^\{\\prime\}\)FiLM conditioning signal:g=ϕ\(t\)∥gglobalg=\\phi\(t\)\\;\\\|\\;g\_\{\\text\{global\}\}Linear→\(scale,bias\)\\rightarrow\(scale,\\;bias\)
Conditional 1D U\-NetInput:\(B,T,d\)\(B,\\;T,\\;d\)ResBlock\(cc\)×2\\times 2Downsample\(×2\\times 2\)ResBlock\(2c2c\)×2\\times 2Downsample\(×2\\times 2\)ResBlock\(4c4c\)×2\\times 2Downsample\(×2\\times 2\)ResBlock\(8c8c\) \[bottleneck\]×2\\times 2Upsample\(×2\\times 2\)ResBlock\(4c4c\)×2\\times 2Upsample\(×2\\times 2\)ResBlock\(2c2c\)×2\\times 2Upsample\(×2\\times 2\)ResBlock\(cc\)×2\\times 2Conv1d\(c→dc\\rightarrow d\)Output:\(B,T,d\)\(B,\\;T,\\;d\)
Table 10:Conditional U\-Net backbone architecture and the corresponding residual block for the Robotics tasks
### D\.5Model and Training Setup
##### Model Architecture\.
We adopt the Conditional 1D U\-Net from Diffusion Policy\[[9](https://arxiv.org/html/2608.07924#bib.bib9)\], matching its architecture and capacity to enable a controlled comparison\. Tables in[10](https://arxiv.org/html/2608.07924#A4.T10)show the architecture\.
##### Training Setup
All models were trained with AdamW and a variety of learning rate schedules following\[[9](https://arxiv.org/html/2608.07924#bib.bib9)\]\. Performance is evaluated by average success rate per task\.
Resource Usage and Time\. We run all experiments on a single RTX 3090 GPU\. Each training run requires approximately 12 GiB of GPU memory less than 24 hours\. With the exception of the blockpush task which took 48 hours\.
#### D\.5\.1Ablation Study on Parameterizations
We evaluate three initialization strategies for the initial velocityv0\(i\)v^\{\(i\)\}\_\{0\}:
- •Learned Initialization:A secondary network predictsv0\(i\)v^\{\(i\)\}\_\{0\}conditioned on the state\. This improves peak performance but increases computational overhead\.
- •Zero Initialization:We setv0\(i\)=0v^\{\(i\)\}\_\{0\}=0at inference time\. This yields strong performance while reducing training complexity and is used as the default\.
- •Historical Warm\-starting:We initializev0\(i\)v^\{\(i\)\}\_\{0\}using the velocity from the previous optimization step\. This consistently leads to instability and degraded performance\.
We hypothesize that the failure of historical warm\-starting is due to non\-stationarity in the optimization landscape, where evolving model parameters render past velocities stale\.
We observe slower convergence on PushT and BlockPush compared to other tasks\. We attribute this to the contact\-rich and highly multimodal nature of these environments, where small action errors can lead to large deviations in long\-term object motion\.Similar Articles
Drifting Objectives for Refining Discrete Diffusion Language Models
This paper introduces TokenDrift, a drifting objective that refines discrete diffusion language models by lifting categorical predictions to a continuous semantic space for anti-symmetric drifting, significantly improving generation quality under a fixed number of denoising steps.
DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models
DRIFT is a framework that adapts pretrained vision-language models for continuous output decoding by combining coarse prediction with iterative flow matching refinement, improving performance on perception and planning tasks.
DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization
This paper proposes DRIFT, a framework that combines offline trajectories with importance-weighted supervised fine-tuning to efficiently achieve multi-turn interactive learning performance comparable to reinforcement learning.
Attention Drift: What Autoregressive Speculative Decoding Models Learn
This paper identifies 'attention drift' in autoregressive speculative decoding models, where drafters' attention shifts from the prompt to their own generated tokens. The authors propose architectural changes, such as post-norm and RMSNorm, which improve acceptance rates and robustness across various benchmarks.
Drift Q-Learning
Proposes DriftQL, which combines a drift-based behavioral regularizer with critic-driven policy improvement for offline RL, outperforming diffusion and flow methods on D4RL and OGBench while maintaining simplicity and efficiency.