ResilPhase: Plug-and-Play Phase Mapping and Noise-Resilient Macro-Trajectory Extrapolation for Diffusion Acceleration
Summary
ResilPhase is a training-free acceleration framework for diffusion models that reformulates accelerated inference as stable macro-trajectory extrapolation in ODE space, using derivative-free barycentric Lagrange extrapolation and bounded phase mapping to achieve state-of-the-art fidelity under high acceleration ratios.
View Cached Full Text
Cached at: 06/26/26, 05:17 AM
# ResilPhase: Plug-and-Play Phase Mapping and Noise-Resilient Macro-Trajectory Extrapolation for Diffusion Acceleration
Source: [https://arxiv.org/html/2606.26769](https://arxiv.org/html/2606.26769)
11institutetext:Zhejiang University, Hangzhou, China
11email:zyan2@zju\.edu\.cn###### Abstract
The adoption of powerful diffusion models is hindered by their significant inference latency\. Recent “cache\-then\-forecast” schemes alleviate this issue by accelerating DiTs using derivative\-based polynomials, but they suffer from severe quality degradation at high acceleration ratios\. Our analysis reveals its root cause: the discrete extrapolation performed on representations that are misaligned with the continuous diffusion trajectory and are numerically unstable\. Thus, accelerated DiTs suffer from accumulated spatial errors, noisy derivative amplification, and high\-order instability\. We therefore reformulate accelerated inference as stable macro\-trajectory extrapolation in ordinary differential equation \(ODE\) space\. Instead of predicting intermediate features, we align forecasting with the model’s Global Drift \(GD\), i\.e\., the end\-to\-end state evolution, thereby eliminating feature inconsistency and memory overhead\. However, even this smooth macro\-trajectory remains vulnerable to the derivative fallacy: its higher\-order temporal derivatives are intrinsically noisy\. Thus, we introduce a derivative\-free barycentric Lagrange extrapolator to effectively bypass derivative instability and approximation error\. We further propose a bounded Phase Mapping that regularizes the extrapolation domain, suppressing oscillatory error growth\. These elements collectively constitute ResilPhase, a noise\-resilient acceleration framework\. Experiments on FLUX\.1\-dev and HunyuanVideo demonstrate state\-of\-the\-art fidelity under aggressive acceleration ratios\. Code is publicly available at[https://github\.com/zqc214/ResilPhase](https://github.com/zqc214/ResilPhase)\.
## 1Introduction
Diffusion Transformers \(DiTs\)\[[9](https://arxiv.org/html/2606.26769#bib.bib9),[25](https://arxiv.org/html/2606.26769#bib.bib25)\]have become the gold standard for high\-fidelity visual generation\[[1](https://arxiv.org/html/2606.26769#bib.bib1),[27](https://arxiv.org/html/2606.26769#bib.bib27),[23](https://arxiv.org/html/2606.26769#bib.bib23),[45](https://arxiv.org/html/2606.26769#bib.bib45),[33](https://arxiv.org/html/2606.26769#bib.bib33)\], because of their scalable architecture\. However, their performance is achieved through an iterative denoising process, requiring tens to hundreds of sequential forward steps that can not be parallelized\. This has become a significant latency bottleneck that obstructs real\-time and large\-scale deployment of DiTs\. FLUX\.1\-dev takes 23\.69 seconds to generate an image on an A100 graphics card\.
To address this latency bottleneck, recent training\-free acceleration efforts have evolved from passive “cache\-then\-reuse” paradigms\[[24](https://arxiv.org/html/2606.26769#bib.bib24),[37](https://arxiv.org/html/2606.26769#bib.bib37),[13](https://arxiv.org/html/2606.26769#bib.bib13)\]to “cache\-then\-forecast” approaches\[[18](https://arxiv.org/html/2606.26769#bib.bib18),[19](https://arxiv.org/html/2606.26769#bib.bib19),[44](https://arxiv.org/html/2606.26769#bib.bib44)\]that employ polynomials to predict feature trajectories\. However, this prevailing paradigm suffers from severe quality degradation at high acceleration ratios\. We argue this failure stems from a single root cause: discrete extrapolation is performed on intermediate representations misaligned with the continuous diffusion trajectory and numerically unstable\. This root cause manifests in three intertwined limitations\.
First, from a spatial perspective, most methods rely on a computationally intensive layer\-wise paradigm\. Fitting polynomials to highly erratic micro\-features within Transformer blocks cascades errors across the network depth, incurring massive memory overhead\. While directly forecasting absolute outputs \(e\.g\., FreqCa\[[17](https://arxiv.org/html/2606.26769#bib.bib17)\]\) bypasses this, as highlighted byΔ\\Delta\-DiT\[[3](https://arxiv.org/html/2606.26769#bib.bib3)\], it discards highly correlated input priors, restricting prediction fidelity bounds\. Conversely, methods reusing residual displacements \(e\.g\.,Δ\\Delta\-DiT\[[3](https://arxiv.org/html/2606.26769#bib.bib3)\]\) remain trapped in localized micro\-updates, missing macroscopic continuous dynamics\. Thus, existing paradigms lack an end\-to\-end target that simultaneously preserves input priors and aligns strictly with the ODE vector field\.

ă
Figure 1:Suppressing numerical instability via Phase Mapping\.\(a\) Standard trajectory extrapolation on uniform timesteps suffers from Runge’s phenomenon, causing chaotic edge oscillations\. \(b\) Our method maps discrete steps to a bounded phase space, minimizing the error bound for a stable fit\.Second, from a temporal perspective, current forecasting methods \(e\.g\., TaylorSeer\[[18](https://arxiv.org/html/2606.26769#bib.bib18)\]and HiCache\[[5](https://arxiv.org/html/2606.26769#bib.bib5)\]\) rely exclusively on derivative\-based approximations\. They critically overlook the mathematical nature of the dynamic trajectory: while the macro\-trajectory itself is smooth, its higher\-order derivatives over time are intrinsically chaotic \(as shown in Fig\.[2](https://arxiv.org/html/2606.26769#S1.F2)\(b\)\)\. Estimating these derivatives via finite differences exponentially amplifies the noise\.
Third, from a numerical approximation standpoint, performing polynomial extrapolation over uniform, discrete time steps inherently triggers Runge’s phenomenon\. This numerical instability causes the extrapolation error bound to grow uncontrollably at the interval edges, inevitably leading to catastrophic quality degradation at aggressive acceleration ratios\.
To overcome these intertwined bottlenecks, we proposeResilPhase, a noise\-resilient acceleration framework explicitly addressing the spatial, temporal, and numerical limitations of existing paradigms\. First, to resolve spatial cascading errors inherent in layer\-wise forecasting, ResilPhase shifts the prediction objective to the network’s macroscopic dynamic evolution\. We propose ODE\-Aligned Macro\-Trajectory Targeting, formulating the prediction target as the Global Drift \(GD\): the end\-to\-end state displacement between the model’s final output and initial input\. Unlike layer\-wise approaches fitting highly oscillatory micro\-signals within Transformer blocks, forecasting the GD aligns the extrapolator with the continuous probability flow ODE\. This paradigm shift eliminates the memory overhead of caching block features and strictly severs error accumulation across network layers\.
Figure 2:Effectiveness of the derivative\-free prediction strategy\.Unlike \(a\) our smooth global macro\-trajectory, \(b\) the finite\-difference derivative trajectory used by prior methods is intrinsically chaotic\. Applying polynomial extrapolation to such unstable dynamics causes severe prediction deviations\. As shown in \(c\), evaluating pure mathematical predictors reveals that our derivative\-free Lagrange extrapolator maintains significantly lowerL1L\_\{1\}relative error across varying extrapolation intervals \(NN\) than derivative\-based solvers \(Taylor, Hermite\)\.While predicting the GD yields a smooth macro\-trajectory ideal for forecasting \(Fig\.[2](https://arxiv.org/html/2606.26769#S1.F2)\(a\)\), it paradoxically amplifies the weakness of derivative\-based methods\. Like layer\-wise micro\-features, the GD’s higher\-order temporal derivatives remain intrinsically chaotic\. Applying derivative\-based solvers like Taylor or Hermite to such noisy signals causes divergent predictions\. Furthermore, finite\-difference estimation of these derivatives across discrete steps introduces compounding approximation errors\. To avoid this, we propose a derivative\-free extrapolator inspired by Lagrange interpolation\. Avoiding the standardO\(N2\)O\(N^\{2\}\)formulation, we develop a barycentric prediction scheme caching and reusing weights\. This reduces complexity toO\(N\)O\(N\)while eliminating derivative noise\. As a purely mathematical predictor, our approach achieves significantly lower and more stable error across extrapolation intervals than prior solvers \(Fig\.[2](https://arxiv.org/html/2606.26769#S1.F2)\(c\)\)\.
Even with a robust derivative\-free formulation, polynomial extrapolation over uniform discrete time steps remains susceptible to Runge’s phenomenon\. As shown in Fig\.[1](https://arxiv.org/html/2606.26769#S1.F1)\(a\), this numerical instability causes severe edge oscillations; ResilPhase is the first work to identify this as a critical bottleneck in diffusion acceleration\. To address this, we introduce a plug\-and\-play Phase Mapping mechanism \(Fig\.[1](https://arxiv.org/html/2606.26769#S1.F1)\(b\)\)\. As the first theoretically grounded technique to remap the discrete temporal extrapolation domain, we bridge classical numerical analysis and diffusion acceleration by utilizing Chebyshev nodes to project linear time steps \(tt\) into a bounded phase domain \(ss\), achieving optimal stability for class\-conditional generation\. Furthermore, we propose a data\-driven Balanced Mapping tailored to complex text\-to\-image and video tasks\. By performing extrapolation within this bounded continuous space, our mappings effectively transform a divergent numerical issue into a highly stable prediction, strictly minimizing the mathematical upper bound of the extrapolation error\.
To summarize, our primary contributions are as follows:
- •We propose ODE\-Aligned Macro\-Trajectory Targeting\. By forecasting the model’s Global Drift \(GD\) rather than layer\-wise features, we strictly sever spatial cascading errors, align with the continuous diffusion ODE, and eliminate heavy memory overhead\.
- •We introduce a noise\-resilient, derivative\-free Barycentric Lagrange extrapolator\. ThisO\(N\)O\(N\)framework perfectly synergizes with the GD by completely bypassing the inherent noise and chaotic high\-order derivatives of finite\-difference approximations\.
- •We design a plug\-and\-play Phase Mapping mechanism\. By non\-linearly projecting discrete time steps into a bounded phase space, it regularizes the extrapolation domain to suppress oscillatory error growth, strictly minimizing the polynomial extrapolation error bound\.
- •Extensive experiments show ResilPhase achieves∼\\sim5×\\timesspeedups on FLUX\.1\-dev and HunyuanVideo while maintaining highly competitive fidelity\. Furthermore, Phase Mapping acts as a plug\-and\-play stabilizer, mitigating extrapolation errors in existing derivative\-based accelerators\.
## 2Related Work
### 2\.1Diffusion Transformers and Traditional Acceleration
Diffusion Transformers \(DiTs\) achieve high\-fidelity generation through a sequential noise\-reversal process, which fundamentally imposes a severe inference latency bottleneck\. To mitigate this, early acceleration efforts primarily focused on model compression, such as network pruning\[[4](https://arxiv.org/html/2606.26769#bib.bib4),[48](https://arxiv.org/html/2606.26769#bib.bib48),[39](https://arxiv.org/html/2606.26769#bib.bib39)\]and quantization\[[11](https://arxiv.org/html/2606.26769#bib.bib11),[14](https://arxiv.org/html/2606.26769#bib.bib14),[32](https://arxiv.org/html/2606.26769#bib.bib32),[6](https://arxiv.org/html/2606.26769#bib.bib6)\], or step reduction via efficient solvers\[[21](https://arxiv.org/html/2606.26769#bib.bib21),[22](https://arxiv.org/html/2606.26769#bib.bib22),[41](https://arxiv.org/html/2606.26769#bib.bib41),[42](https://arxiv.org/html/2606.26769#bib.bib42)\]and knowledge distillation\[[30](https://arxiv.org/html/2606.26769#bib.bib30),[35](https://arxiv.org/html/2606.26769#bib.bib35),[15](https://arxiv.org/html/2606.26769#bib.bib15),[47](https://arxiv.org/html/2606.26769#bib.bib47)\]\. However, these approaches typically demand computationally expensive retraining to recover generation quality and often rely on complex algorithmic designs\. This reliance fundamentally limits their plug\-and\-play generality, driving recent research toward more flexible, training\-free acceleration paradigms\.
### 2\.2Acceleration Based on Feature Caching
Feature caching, the main training\-free diffusion acceleration strategy, evolved from passively reusing temporal features \(DeepCache\[[24](https://arxiv.org/html/2606.26769#bib.bib24)\], FORA\[[31](https://arxiv.org/html/2606.26769#bib.bib31)\],Δ\\Delta\-DiT\[[3](https://arxiv.org/html/2606.26769#bib.bib3)\], TeaCache\[[16](https://arxiv.org/html/2606.26769#bib.bib16)\], ToCa\[[49](https://arxiv.org/html/2606.26769#bib.bib49)\]\) to actively predicting them \(TaylorSeer\[[18](https://arxiv.org/html/2606.26769#bib.bib18)\], FoCa\[[44](https://arxiv.org/html/2606.26769#bib.bib44)\], HiCache\[[5](https://arxiv.org/html/2606.26769#bib.bib5)\], SpeCa\[[19](https://arxiv.org/html/2606.26769#bib.bib19)\], ClusCa\[[46](https://arxiv.org/html/2606.26769#bib.bib46)\], FreqCa\[[17](https://arxiv.org/html/2606.26769#bib.bib17)\]\)\. While recent methods like DiCache\[[2](https://arxiv.org/html/2606.26769#bib.bib2)\]dynamically adjust caching intervals, these predictive approaches share critical limitations\. Spatially, layer\-wise forecasting incurs memory overhead and amplifies micro\-feature errors across network depth\. Temporally, these methods implicitly assume smooth high\-order derivatives, relying on finite\-difference approximations\. However, approaches using Taylor series\[[18](https://arxiv.org/html/2606.26769#bib.bib18),[19](https://arxiv.org/html/2606.26769#bib.bib19)\], Hermite polynomials\[[5](https://arxiv.org/html/2606.26769#bib.bib5),[17](https://arxiv.org/html/2606.26769#bib.bib17)\], or ODE\-based forecasting\[[44](https://arxiv.org/html/2606.26769#bib.bib44)\]encounter a mathematical bottleneck: while the macro\-trajectory is smooth, its temporal derivatives remain intrinsically chaotic\. Consequently, applying derivative\-based forecasting to these noisy signals makes predictions susceptible to numerical instabilities like Runge’s phenomenon, restricting maximum acceleration and robustness\.
Figure 3:The ResilPhase acceleration framework\. \(a\) We shift from layer\-wise features to the ODE\-aligned Global Drift \(GD\), the end\-to\-end state displacement\. \(b\) Phase Mapping non\-linearly projects discrete time steps into a bounded phase space via Chebyshev \(CM\) or Balanced \(BM\) to suppress numerical instability\. \(c\) The pipeline runs full computations everyNNsteps, while our derivative\-free Barycentric Lagrange extrapolator of ordermmestimates the GD for intermediate steps\.
## 3Methodology
### 3\.1Preliminaries and the Derivative\-Induced Instability
Diffusion Transformers \(DiT\)A DiT model functions as a learnable velocity field in a probability flow ODE\. It consists of a stack ofLLTransformer blocks\. We denote the transformation of thell\-th block asglg\_\{l\}, wherel∈\{1,…,L\}l\\in\\\{1,\\dots,L\\\}\. For the input latent state𝐱t\\mathbf\{x\}\_\{t\}at timestepttand conditioncc, the end\-to\-end mapping is a composition of these blocks:G\(𝐱t,c\)=gL\(gL−1\(…g1\(𝐱t,c\)…\)\)G\(\\mathbf\{x\}\_\{t\},c\)=g\_\{L\}\(g\_\{L\-1\}\(\\dots g\_\{1\}\(\\mathbf\{x\}\_\{t\},c\)\\dots\)\)\. Thus, the input\(𝐱t,c\)\(\\mathbf\{x\}\_\{t\},c\)is transformed into the final output velocity \(or noise prediction\) through this complete processGG\.
Predictive Caching for Acceleration\.To reduce inference latency, predictive caching uses a sparse schedule\. Given an acceleration rateNN, full forward passes occur only at anchor timesteps \(e\.g\.,t,t−N,…t,t\-N,\\dots\)\. For intermediate steps, computations are skipped and features are estimated via a lightweight predictor\.
Conventionally, this prediction is layer\-wise\. LetF\(xtl\)F\(x\_\{t\}^\{l\}\)denote the feature output of layerllat timesteptt\. Existing methods \(e\.g\., TaylorSeer\[[18](https://arxiv.org/html/2606.26769#bib.bib18)\], SpeCa\[[19](https://arxiv.org/html/2606.26769#bib.bib19)\]\) assume smooth feature trajectories and employ polynomial expansions \(like Taylor series\) for forecasting:
Fpred,m\(xt−kl\)=F\(xtl\)\+∑i=1mΔiF\(xtl\)i\!\(−k\)i,\\vskip\-5\.69046ptF\_\{\\text\{pred\},m\}\(x\_\{t\-k\}^\{l\}\)=F\(x\_\{t\}^\{l\}\)\+\\sum\_\{i=1\}^\{m\}\\frac\{\\Delta^\{i\}F\(x\_\{t\}^\{l\}\)\}\{i\!\}\(\-k\)^\{i\},\(1\)whereΔiF\\Delta^\{i\}Fis theii\-th order finite difference derivative approximation\.
The Derivative\-Induced Instability\.Applying Equation \([1](https://arxiv.org/html/2606.26769#S3.E1)\) to discrete diffusion dynamics is inherently unstable\. While feature values are smooth, their finite\-difference derivatives \(ΔiF\\Delta^\{i\}F\) are dominated by chaos \(Fig\.[2](https://arxiv.org/html/2606.26769#S1.F2)\)\. Because finite differences act as high\-pass filters that exponentially amplify noise, derivative\-based solvers \(including Hermite splines\[[5](https://arxiv.org/html/2606.26769#bib.bib5)\]\) inherently destabilize trajectory reconstruction\. This necessitates our derivative\-free paradigm\.
### 3\.2The ResilPhase Framework
To bypass derivative instability, we adopt a derivative\-free approach\. Form\+1m\+1historical data points\{\(t0,F0\),…,\(tm,Fm\)\}\\\{\(t\_\{0\},F\_\{0\}\),\\dots,\(t\_\{m\},F\_\{m\}\)\\\}, the unique degree\-mminterpolation polynomialP\(t\)P\(t\)is:
P\(t\)=∑j=0mFjLj\(t\)\.P\(t\)=\\sum\_\{j=0\}^\{m\}F\_\{j\}L\_\{j\}\(t\)\.\(2\)ResilPhase \(Fig\.[3](https://arxiv.org/html/2606.26769#S2.F3)\) makes two choices\. First, the targetFjF\_\{j\}is the Global Drift, replacing erratic layer\-wise features with a smooth macro\-trajectory\. Second, we stabilize the basisLjL\_\{j\}via Barycentric Interpolation and Phase Mapping\. This enables accurate extrapolation everyNNsteps without gradient noise\.
### 3\.3ODE\-Aligned Macro\-Trajectory Targeting
Previous caching\-based acceleration methods rely on layer\-wise prediction\. However, from a dynamic systems perspective, forcing polynomials to fit highly non\-linear, erratic micro\-features within individual Transformer blocks causes prediction errors to cascade across the network depth\.
The Spatial Cascading Error of Layer\-wise Forecasting\.Letxtl−1x\_\{t\}^\{l\-1\}be the latent input to thell\-th Transformer block at timesteptt\. In a DiT block, the exact forward computation updates the hidden state via residual connections across the self\-attention \(fattn,lf\_\{\\text\{attn\},l\}\) and MLP \(fmlp,lf\_\{\\text\{mlp\},l\}\) modules\. For mathematical simplicity, we group the sum of these residual updates at layerllinto a single functionflf\_\{l\}, meaning the exact transformation isxtl=xtl−1\+fl\(xtl−1\)x\_\{t\}^\{l\}=x\_\{t\}^\{l\-1\}\+f\_\{l\}\(x\_\{t\}^\{l\-1\}\)\. During a skipped step, layer\-wise methods must approximate these internal updates via a polynomial predictor𝒫l\\mathcal\{P\}\_\{l\}, yielding the estimated feature:
x^tl=x^tl−1\+𝒫l\(x^tl−1\)\.\\hat\{x\}\_\{t\}^\{l\}=\\hat\{x\}\_\{t\}^\{l\-1\}\+\\mathcal\{P\}\_\{l\}\(\\hat\{x\}\_\{t\}^\{l\-1\}\)\.\(3\)
Letel=‖𝒫l\(x^tl−1\)−fl\(x^tl−1\)‖e\_\{l\}=\\\|\\mathcal\{P\}\_\{l\}\(\\hat\{x\}\_\{t\}^\{l\-1\}\)\-f\_\{l\}\(\\hat\{x\}\_\{t\}^\{l\-1\}\)\\\|andEl=‖x^tl−xtl‖E\_\{l\}=\\\|\\hat\{x\}\_\{t\}^\{l\}\-x\_\{t\}^\{l\}\\\|denote the local prediction error and the total accumulated state error up to layerll, respectively\. Assuming the transformationflf\_\{l\}is Lipschitz continuous with constantLfL\_\{f\}, we bound the accumulated error via the triangle inequality:
El\\displaystyle E\_\{l\}=‖\(x^tl−1\+𝒫l\(x^tl−1\)\)−\(xtl−1\+fl\(xtl−1\)\)‖\\displaystyle=\|\|\(\\hat\{x\}\_\{t\}^\{l\-1\}\+\\mathcal\{P\}\_\{l\}\(\\hat\{x\}\_\{t\}^\{l\-1\}\)\)\-\(x\_\{t\}^\{l\-1\}\+f\_\{l\}\(x\_\{t\}^\{l\-1\}\)\)\|\|\(4\)≤‖x^tl−1−xtl−1‖\+‖𝒫l\(x^tl−1\)−fl\(x^tl−1\)‖\+‖fl\(x^tl−1\)−fl\(xtl−1\)‖\\displaystyle\\leq\|\|\\hat\{x\}\_\{t\}^\{l\-1\}\-x\_\{t\}^\{l\-1\}\|\|\+\|\|\\mathcal\{P\}\_\{l\}\(\\hat\{x\}\_\{t\}^\{l\-1\}\)\-f\_\{l\}\(\\hat\{x\}\_\{t\}^\{l\-1\}\)\|\|\+\|\|f\_\{l\}\(\\hat\{x\}\_\{t\}^\{l\-1\}\)\-f\_\{l\}\(x\_\{t\}^\{l\-1\}\)\|\|≤El−1\+el\+LfEl−1=\(1\+Lf\)El−1\+el\.\\displaystyle\\leq E\_\{l\-1\}\+e\_\{l\}\+L\_\{f\}E\_\{l\-1\}=\(1\+L\_\{f\}\)E\_\{l\-1\}\+e\_\{l\}\.
Solving this recurrence relation from the first layerl=1l=1to the final layerLL\(noting that the input is exact, thusE0=0E\_\{0\}=0\), the total spatial cascading error is bounded by:
EL≤∑l=1L\(1\+Lf\)L−lel\.E\_\{L\}\\leq\\sum\_\{l=1\}^\{L\}\(1\+L\_\{f\}\)^\{L\-l\}e\_\{l\}\.\(5\)
This rigorous mathematical bound reveals a fundamental flaw in existing paradigms: spatial prediction errors amplify exponentially across the network depthLL\. Fitting high\-frequency micro\-features at each intermediate layer merely exacerbates the local errorele\_\{l\}, driving the total divergenceELE\_\{L\}uncontrollably high at aggressive acceleration ratios\.
Formulating the Global Drift\.To fundamentally sever this spatial error accumulation, we must elevate the prediction objective from intermediate micro\-features to the macroscopic evolution of the ODE\. We define the Global Drift \(GD\), denoted asD\(xt\)D\(x\_\{t\}\), as the end\-to\-end state displacement between the model’s final outputG\(xt\)G\(x\_\{t\}\)and its initial inputxtx\_\{t\}:
D\(xt\)=G\(xt\)−xt\.D\(x\_\{t\}\)=G\(x\_\{t\}\)\-x\_\{t\}\.\(6\)
Instead of predictingLLdistinct layer updates, a single macro\-predictor𝒫macro\\mathcal\{P\}\_\{\\text\{macro\}\}forecasts this trajectory during skipped steps, yieldingD^\(xt\)\\hat\{D\}\(x\_\{t\}\)\. The final output is reconstructed via addition:G^\(xt\)=xt\+D^\(xt\)\\hat\{G\}\(x\_\{t\}\)=x\_\{t\}\+\\hat\{D\}\(x\_\{t\}\)\.
Mathematically, the approximation error of our macro\-trajectory framework is strictly determined by a single term:
Emacro=‖G^\(xt\)−G\(xt\)‖=‖D^\(xt\)−D\(xt\)‖=emacro\.E\_\{\\text\{macro\}\}=\|\|\\hat\{G\}\(x\_\{t\}\)\-G\(x\_\{t\}\)\|\|=\|\|\\hat\{D\}\(x\_\{t\}\)\-D\(x\_\{t\}\)\|\|=e\_\{\\text\{macro\}\}\.\(7\)
By treating the entire DiT stack as a unified probability flow step, we completely bypass the cascading recurrence relation\. The error bound is decoupled from the network depthLL, successfully reducing the spatial error accumulation fromO\(\(1\+Lf\)L\)O\(\(1\+L\_\{f\}\)^\{L\}\)toO\(1\)O\(1\)\.
While the GD formulation successfully eliminates spatial error amplification, forecasting this macro\-trajectory with derivative\-based solvers remains vulnerable to the chaotic noise highlighted in Section[3\.1](https://arxiv.org/html/2606.26769#S3.SS1)\. This necessitates a robust, derivative\-free extrapolation framework, which we introduce next\.
### 3\.4Derivative\-Free Barycentric Interpolation
To circumvent the catastrophic noise amplification inherent in derivative\-based solvers as highlighted in the Derivative Paradox, we strictly ground our framework in a derivative\-free approach\. We define our target feature valueFjF\_\{j\}as the fully computed Global DriftD\(xt\)D\(x\_\{t\}\)at historical timesteptjt\_\{j\}\. A straightforward yet effective formulation of the interpolation polynomialP\(t\)P\(t\)relies on the classical Lagrange basis polynomialsLjL\_\{j\}, where
Lj\(t\)=∏k=0,k≠jmt−tktj−tk\.L\_\{j\}\(t\)=\\prod\_\{k=0,k\\neq j\}^\{m\}\\frac\{t\-t\_\{k\}\}\{t\_\{j\}\-t\_\{k\}\}\.\(8\)
However, the standard Lagrange formula is computationally expensive with a computational complexity ofO\(m2\)O\(m^\{2\}\)\. To address this issue, we adopt a more stable and efficient variant: the Barycentric Lagrange Interpolation Formula:
P\(t\)=∑j=0mwjt−tjFj∑j=0mwjt−tj,P\(t\)=\\frac\{\\sum\_\{j=0\}^\{m\}\\frac\{w\_\{j\}\}\{t\-t\_\{j\}\}F\_\{j\}\}\{\\sum\_\{j=0\}^\{m\}\\frac\{w\_\{j\}\}\{t\-t\_\{j\}\}\},\(9\)wherewj=1∏k=0,k≠jm\(tj−tk\)\.\\text\{where\}\\quad w\_\{j\}=\\frac\{1\}\{\\prod\_\{k=0,k\\neq j\}^\{m\}\(t\_\{j\}\-t\_\{k\}\)\}\.\(10\)
In this formulation, the barycentric weightswjw\_\{j\}depend only on the relative positions of the interpolation nodes\{tj\}\\\{t\_\{j\}\\\}, not the target feature values\{Fj\}\\\{F\_\{j\}\\\}\. This pivotal property allows the weights to be precomputed and cached during the full computation steps, making them readily available for all subsequent predictions\. Thus, the complexity can be reduced toO\(m\)O\(m\)\.
### 3\.5Phase Mapping for Prediction Stability
While the barycentric extrapolator isolates gradient noise, applying polynomial extrapolation over uniform discrete timesteps triggers severe numerical instability\. This classic issue, Runge’s phenomenon, causes uncontrollable edge oscillations that degrade fidelity at high acceleration ratios\. To suppress this error growth, we introduce a plug\-and\-play Phase Mapping mechanism \(Fig\.[3](https://arxiv.org/html/2606.26769#S2.F3)\(b\)\)\.
The Source of Numerical Instability\.To understand this instability mathematically, we examine the error term for Lagrange interpolation\. For a functionF\(t\)F\(t\)that ism\+1m\+1times differentiable, there exists a real numberξ∈\[tmin,tmax\]\\xi\\in\[t\_\{\\min\},t\_\{\\max\}\]such that the errorE\(t\)=F\(t\)−P\(t\)E\(t\)=F\(t\)\-P\(t\)of its degree\-mmpolynomial interpolationP\(t\)P\(t\)is:
E\(t\)=F\(m\+1\)\(ξ\)\(m\+1\)\!∏j=0m\(t−tj\),E\(t\)=\\frac\{F^\{\(m\+1\)\}\(\\xi\)\}\{\(m\+1\)\!\}\\prod\_\{j=0\}^\{m\}\(t\-t\_\{j\}\),\(11\)wheretjt\_\{j\}s are time steps used for interpolation\. There are two independent components of this function:Dm\+1=F\(m\+1\)\(ξ\)\(m\+1\)\!D\_\{m\+1\}=\\frac\{F^\{\(m\+1\)\}\(\\xi\)\}\{\(m\+1\)\!\}, ande\(t\)=∏j=0m\(t−tj\)e\(t\)=\\prod\_\{j=0\}^\{m\}\(t\-t\_\{j\}\)\.
Because theDm\+1D\_\{m\+1\}term is an inherent property of the DiT model and cannot be modified, the only way to reduce the total errorE\(t\)E\(t\)is by minimizing the node\-dependent terme\(t\)e\(t\)\. We achieve this by remapping the uniformly spaced time stepstjt\_\{j\}to a new, non\-uniform distribution of nodessjs\_\{j\}\(i\.e\., a mapping oft→st\\to s\)\. Based on the properties of DiTs, we propose two such mapping techniques suitable for different tasks: Chebyshev Mapping and Balanced Mapping\.
#### 3\.5\.1Chebyshev Mapping
For a given set ofm\+1m\+1recent, fully computed timesteps\{t0,…,tm\}\\\{t\_\{0\},\\dots,t\_\{m\}\\\}, we generate the corresponding Chebyshev nodes\{sk\}\\\{s\_\{k\}\\\}as follows:
sk=cos\(\(2k\+1\)π2\(m\+1\)\),fork=0,…,m\.s\_\{k\}=\\cos\\left\(\\frac\{\(2k\+1\)\\pi\}\{2\(m\+1\)\}\\right\),\\quad\\text\{for \}k=0,\\dots,m\.\(12\)However, this formula only prescribes the discrete nodes\{sk\}\\\{s\_\{k\}\\\}for known timesteps and is not a continuous function applicable to an arbitraryttargett\_\{\\text\{target\}\}\. To extrapolate the phase coordinatestargets\_\{\\text\{target\}\}for a target stepttarget<tmt\_\{\\text\{target\}\}<t\_\{m\}, we employ stable linear extrapolation using the two most recent points:
starget=sm\+sm−sm−1tm−tm−1\(ttarget−tm\)\.s\_\{\\text\{target\}\}=s\_\{m\}\+\\frac\{s\_\{m\}\-s\_\{m\-1\}\}\{t\_\{m\}\-t\_\{m\-1\}\}\(t\_\{\\text\{target\}\}\-t\_\{m\}\)\.\(13\)
#### 3\.5\.2Balanced Mapping
While the fixed structure of Chebyshev nodes effectively bounds the error space, its predetermined nature lacks the flexibility to dynamically adjust to varying temporal distributions encountered during highly complex text\-conditioned tasks\. To provide a more robust and adaptable phase transformation, we propose a novel, data\-driven mechanism called Balanced Mapping\.
Unlike Chebyshev mapping, this adaptive strategy process begins by analyzing the current set ofm\+1m\+1fully computed timesteps\{tj\}j=0m\\\{t\_\{j\}\\\}\_\{j=0\}^\{m\}to compute their meanμt\\mu\_\{t\}and maximum absolute deviationdmax=maxj\|tj−μt\|d\_\{\\max\}=\\max\_\{j\}\|t\_\{j\}\-\\mu\_\{t\}\|\.
The complete non\-linear mapping is then given by:
s=tanh\(α⋅t−μtdmax\),s=\\tanh\\left\(\\alpha\\cdot\\frac\{t\-\\mu\_\{t\}\}\{d\_\{\\max\}\}\\right\),\(14\)
whereα\>0\\alpha\>0is a configurable hyperparameter\. This data\-driven transformation dynamically confines mapped coordinates via a hyperbolic tangent function, adapting to the spread of recent timesteps\. Althoughα\\alphaenables fine\-grained control, a defaultα=0\.55\\alpha=0\.55consistently ensures robust performance across diverse settings\. Ultimately, Balanced Mapping delivers a resilient, plug\-and\-play extrapolation domain robust to numerical outliers\.
#### 3\.5\.3Error Analysis
We quantitatively compare the node\-dependent error bound,maxt\|e\(t\)\|\\max\_\{t\}\|e\(t\)\|, with and without phase mapping\.
##### Without Phase Mapping\.
Without phase mapping, the maximum value of the interpolation error polynomial is:
\|eori\(ttarget\)\|\\displaystyle\|e\_\{\\text\{ori\}\}\(t\_\{\\text\{target\}\}\)\|=∏j=0m\|ttarget−tj\|\\displaystyle=\\prod\_\{j=0\}^\{m\}\|t\_\{\\text\{target\}\}\-t\_\{j\}\|≤\|eori\(tm−N\+1\)\|=∏j=0m\|\(m−j\+1\)N−1\|\.\\displaystyle\\leq\|e\_\{\\text\{ori\}\}\(t\_\{m\}\-N\+1\)\|=\\prod\_\{j=0\}^\{m\}\|\(m\-j\+1\)N\-1\|\.\(15\)
##### Chebyshev Mapping\.
With Chebyshev mapping, the corresponding error polynomial is given by:
e\(s\)=∏j=0m\(s−sj\)\.\\vskip\-5\.69046pte\(s\)=\\prod\_\{j=0\}^\{m\}\(s\-s\_\{j\}\)\.\(16\)The point of maximum error still corresponds to the physical timettarget=tm−N\+1t\_\{\\text\{target\}\}=t\_\{m\}\-N\+1, which is mapped to the coordinatestargets\_\{\\text\{target\}\}in the phase space\. Although the extrapolated target coordinatestargets\_\{\\text\{target\}\}slightly exceeds the interval\(−1,1\)\(\-1,1\), it maintains a mathematically bounded distance to the historical Chebyshev nodes\. By algebraically mapping the physical time differences into the phase space, we derive the rigorous upper bound for the error polynomial’s magnitude:
\|echeby\(starget\)\|≤∏j=0m\|2\(m−j\+1\)−2N\|\.\\vskip\-5\.69046pt\|e\_\{\\text\{cheby\}\}\(s\_\{\\text\{target\}\}\)\|\\leq\\prod\_\{j=0\}^\{m\}\\left\|2\(m\-j\+1\)\-\\frac\{2\}\{N\}\\right\|\.\(17\)
##### Balanced Mapping\.
Similar to Chebyshev Mapping, we have:
\|ebal\(starget\)\|≤∏j=0m\|Tbal\|,\\vskip\-5\.69046pt\|e\_\{\\text\{bal\}\}\(s\_\{\\text\{target\}\}\)\|\\leq\\prod\_\{j=0\}^\{m\}\|T\_\{bal\}\|,\(18\)whereTbal=sinh\(2α\(N−1\)mN\)cosh\(α\)⋅cosh\(α2−mN−2NmN\)\+2\(m−j\)T\_\{bal\}=\\frac\{\\sinh\\left\(\\frac\{2\\alpha\(N\-1\)\}\{mN\}\\right\)\}\{\\cosh\(\\alpha\)\\cdot\\cosh\\left\(\\alpha\\frac\{2\-mN\-2N\}\{mN\}\\right\)\}\+2\(m\-j\)\.
#### 3\.5\.4Comparative Analysis\.
We now compare the three error bounds\. Let us define a functionf\(N,j,m\)f\(N,j,m\)as the difference between the terms inside the products\. The difference between no mapping and Chebyshev Node Mapping is:
f\(N,j,m\)=\(m−j\)\(N−2\)\+N\+2N−3\.f\(N,j,m\)=\(m\-j\)\(N\-2\)\+N\+\\frac\{2\}\{N\}\-3\.\(19\)In our acceleration task, constraints areN≥2N\\geq 2andm−j\>0m\-j\>0\. Asf\(N,j,m\)f\(N,j,m\)increases monotonically withN≥2N\\geq 2, it follows thatf\(N,j,m\)\>0f\(N,j,m\)\>0\. This demonstrates that Chebyshev Mapping significantly reduces the polynomial’s error upper bound\. A similar analysis shows Balanced Mapping also lowers the error bound versus the no\-mapping baseline\.
When comparing the two mapping schemes, extensive empirical evaluations reveal a distinct, task\-dependent superiority\. Specifically, Chebyshev Mapping proves to be highly effective for class\-conditional image generation tasks\. In contrast, for highly complex, text\-conditioned generation tasks such as text\-to\-image and text\-to\-video, our novel Balanced Mapping consistently demonstrates a clear advantage in preserving semantic alignment and visual fidelity\.
### 3\.6Methodology Summary
In conclusion, ResilPhase integrates three core components to construct a noise\-resilient acceleration framework: \(1\) targeting the ODE\-aligned Global Drift to bypass spatial cascading errors; \(2\) a derivative\-free Barycentric Lagrange Interpolator to eliminate temporal gradient noise; and \(3\) a bounded Phase Mapping mechanism to suppress numerical extrapolation instability\.
The final predicted macro\-trajectory displacement is formulated as:
D^\(xpred\)=P\(spred\)=∑j=0mwjspred−sjDj∑j=0mwjspred−sj\.\\hat\{D\}\(x\_\{\\text\{pred\}\}\)=P\(s\_\{\\text\{pred\}\}\)=\\frac\{\\sum\_\{j=0\}^\{m\}\\frac\{w\_\{j\}\}\{s\_\{\\text\{pred\}\}\-s\_\{j\}\}D\_\{j\}\}\{\\sum\_\{j=0\}^\{m\}\\frac\{w\_\{j\}\}\{s\_\{\\text\{pred\}\}\-s\_\{j\}\}\}\.\(20\)
Ultimately, ResilPhase delivers remarkably stable, high\-fidelity extrapolations across diverse generative tasks, maintaining exceptional generation quality even under extreme acceleration regimes\.
## 4Experiments
### 4\.1Settings
Model Configurations\.We evaluate four mainstream models using a 50\-step sampling schedule for fair comparison\. For text\-to\-image, we use FLUX\.1\-dev\[[12](https://arxiv.org/html/2606.26769#bib.bib12)\]with the Rectified Flow\[[20](https://arxiv.org/html/2606.26769#bib.bib20)\]sampler and SDXL\-base\-1\.0\[[26](https://arxiv.org/html/2606.26769#bib.bib26)\]\. For text\-to\-video, we use HunyuanVideo\-Large\[[36](https://arxiv.org/html/2606.26769#bib.bib36)\]\. For class\-conditional generation, we employ DiT\-XL/2\[[25](https://arxiv.org/html/2606.26769#bib.bib25)\]with DDIM\[[34](https://arxiv.org/html/2606.26769#bib.bib34)\]\. Detailed settings and SDXL analyses are in the Supplementary Material\.
Evaluation\.We evaluate across three tasks\. For the text\-to\-image task, we use 200 DrawBench\[[29](https://arxiv.org/html/2606.26769#bib.bib29)\]prompts, assessing quality and alignment with ImageReward\[[40](https://arxiv.org/html/2606.26769#bib.bib40)\]and CLIP Score\[[7](https://arxiv.org/html/2606.26769#bib.bib7)\]\. For the text\-to\-video task, we use 946 prompts, evaluating on 16 core dimensions\. For both tasks, we also report PSNR, SSIM\[[38](https://arxiv.org/html/2606.26769#bib.bib38)\], and LPIPS\[[43](https://arxiv.org/html/2606.26769#bib.bib43)\]for fidelity against original results\. For the class\-conditional task, we generate images from 1,000 ImageNet\[[28](https://arxiv.org/html/2606.26769#bib.bib28)\]categories and evaluate using FID\-50k\[[8](https://arxiv.org/html/2606.26769#bib.bib8)\], sFID, and Inception Score \(IS\)\. Further details are in the supplementary material\.
Table 1:Quantitative comparison of text\-to\-image generation for FLUX\.1\-dev\. Methods are compared at similar latency acceleration ratios across three tiers of speedups to evaluate generation quality\.MethodAccelerationImageReward↑\\uparrowCLIP↑\\uparrowPSNR↑\\uparrowSSIM↑\\uparrowLPIPS↓\\downarrowFLUX\.1\-devLatency\(s\)↓\\downarrowSpeed↑\\uparrowDrawBenchScore\[dev\]: 50 steps23\.691\.00×\\times1\.080432\.711\-\-\-\[dev\]: 11 steps5\.214\.55×\\times0\.954132\.48528\.3970\.59390\.5001FORA\(𝒩=6\)\(\\mathcal\{N\}=6\)\[[31](https://arxiv.org/html/2606.26769#bib.bib31)\]5\.204\.56×\\times0\.846832\.17828\.2300\.57000\.5405TeaCache\(l=1\.4\)\(\{l\}=1\.4\)\[[16](https://arxiv.org/html/2606.26769#bib.bib16)\]4\.924\.82×4\.82\\times0\.785032\.58827\.9540\.38370\.8349ToCa\(𝒩=12,R=90%\)\(\\mathcal\{N\}=12,R=90\\%\)\[[49](https://arxiv.org/html/2606.26769#bib.bib49)\]9\.192\.58×\\times0\.728431\.47528\.4440\.53800\.5768ClusCa\(𝒩=12,O=2\)\(\\mathcal\{N\}=12,O=2\)\[[46](https://arxiv.org/html/2606.26769#bib.bib46)\]4\.884\.85×\\times0\.495830\.35328\.0790\.37240\.7340SpeCa\(τ0=12,β=0\.3\)\(\{\\tau\_\{0\}\}=12,\{\\beta\}=0\.3\)\[[19](https://arxiv.org/html/2606.26769#bib.bib19)\]4\.964\.78×\\times0\.979832\.57128\.3660\.55670\.5324PFDiff\(𝒦=3,H=3\)\(\\mathcal\{K\}=3,H=3\)\[[37](https://arxiv.org/html/2606.26769#bib.bib37)\]5\.824\.07×\\times1\.038632\.81628\.6710\.61620\.4670HiCache\(𝒩=11,O=2\)\(\\mathcal\{N\}=11,O=2\)\[[5](https://arxiv.org/html/2606.26769#bib.bib5)\]5\.084\.66×\\times0\.804031\.60428\.2680\.52610\.5955FreqCa\(𝒩=6,O=2\)\(\\mathcal\{N\}=6,O=2\)\[[17](https://arxiv.org/html/2606.26769#bib.bib17)\]4\.854\.88×\\times1\.013032\.11428\.1200\.40230\.6877TaylorSeer\(𝒩=11,O=2\)\(\\mathcal\{N\}=11,O=2\)\[[18](https://arxiv.org/html/2606.26769#bib.bib18)\]5\.104\.65×\\times0\.624131\.89527\.9400\.30140\.8012ResilPhase\(𝒩=6,O=1\)\(\\mathcal\{N\}=6,O=1\)4\.774\.97×\\times1\.025832\.84729\.5360\.66550\.3834\[dev\]: 13 steps6\.113\.88×\\times0\.982132\.58628\.4980\.61360\.4682FORA\(𝒩=5\)\(\\mathcal\{N\}=5\)\[[31](https://arxiv.org/html/2606.26769#bib.bib31)\]5\.694\.16×\\times0\.923532\.37028\.2880\.56940\.5156TeaCache\(l=1\.0\)\(\{l\}=1\.0\)\[[16](https://arxiv.org/html/2606.26769#bib.bib16)\]5\.824\.07×4\.07\\times0\.851832\.75027\.9600\.38320\.8253ToCa\(𝒩=10,R=90%\)\(\\mathcal\{N\}=10,R=90\\%\)\[[49](https://arxiv.org/html/2606.26769#bib.bib49)\]10\.002\.37×\\times0\.886732\.03828\.6010\.57780\.5176ClusCa\(𝒩=8,O=2\)\(\\mathcal\{N\}=8,O=2\)\[[46](https://arxiv.org/html/2606.26769#bib.bib46)\]5\.924\.00×\\times0\.991332\.41028\.4300\.55280\.5310SpeCa\(τ0=8,β=0\.3\)\(\{\\tau\_\{0\}\}=8,\{\\beta\}=0\.3\)\[[19](https://arxiv.org/html/2606.26769#bib.bib19)\]5\.844\.06×\\times1\.045932\.76428\.7150\.61760\.4460PFDiff\(𝒦=3,H=2\)\(\\mathcal\{K\}=3,H=2\)\[[37](https://arxiv.org/html/2606.26769#bib.bib37)\]5\.834\.06×\\times1\.044232\.73128\.9230\.64810\.4223HiCache\(𝒩=7,O=2\)\(\\mathcal\{N\}=7,O=2\)\[[5](https://arxiv.org/html/2606.26769#bib.bib5)\]5\.963\.97×\\times1\.007932\.62228\.7050\.61470\.4542FreqCa\(𝒩=5,O=2\)\(\\mathcal\{N\}=5,O=2\)\[[17](https://arxiv.org/html/2606.26769#bib.bib17)\]5\.764\.11×\\times1\.049632\.71628\.1320\.41010\.6835TaylorSeer\(𝒩=7,O=2\)\(\\mathcal\{N\}=7,O=2\)\[[18](https://arxiv.org/html/2606.26769#bib.bib18)\]5\.983\.96×\\times0\.940632\.65728\.0490\.39310\.7119ResilPhase\(𝒩=5,O=1\)\(\\mathcal\{N\}=5,O=1\)5\.684\.17×\\times1\.059132\.90129\.5560\.70240\.3342\[dev\]: 15 steps7\.003\.38×\\times1\.004432\.61828\.6120\.63450\.4375Δ\\Delta\-DiT \(𝒩=10\\mathcal\{N\}=10\)7\.603\.12×\\times0\.116430\.36928\.1570\.55870\.6408FORA\(𝒩=4\)\(\\mathcal\{N\}=4\)\[[31](https://arxiv.org/html/2606.26769#bib.bib31)\]7\.503\.16×\\times0\.979132\.68428\.3420\.60780\.4725TeaCache\(l=0\.7\)\(\{l\}=0\.7\)\[[16](https://arxiv.org/html/2606.26769#bib.bib16)\]7\.443\.18×3\.18\\times0\.953632\.79427\.9610\.38130\.8176ToCa\(𝒩=8,R=90%\)\(\\mathcal\{N\}=8,R=90\\%\)\[[49](https://arxiv.org/html/2606.26769#bib.bib49)\]10\.782\.20×\\times0\.953732\.49028\.7860\.59540\.4764ClusCa\(𝒩=6,O=2\)\(\\mathcal\{N\}=6,O=2\)\[[46](https://arxiv.org/html/2606.26769#bib.bib46)\]6\.773\.50×\\times1\.042932\.79528\.8130\.62050\.4384SpeCa\(τ0=4,β=0\.3\)\(\{\\tau\_\{0\}\}=4,\{\\beta\}=0\.3\)\[[19](https://arxiv.org/html/2606.26769#bib.bib19)\]6\.783\.49×\\times1\.059833\.10829\.1200\.66620\.3786PFDiff\(𝒦=2,H=1\)\(\\mathcal\{K\}=2,H=1\)\[[37](https://arxiv.org/html/2606.26769#bib.bib37)\]7\.663\.09×\\times1\.056632\.98929\.5570\.70300\.3517HiCache\(𝒩=5,O=2\)\(\\mathcal\{N\}=5,O=2\)\[[5](https://arxiv.org/html/2606.26769#bib.bib5)\]7\.343\.23×\\times1\.043832\.89829\.2400\.68300\.3568FreqCa\(𝒩=4,O=2\)\(\\mathcal\{N\}=4,O=2\)\[[17](https://arxiv.org/html/2606.26769#bib.bib17)\]6\.743\.51×\\times1\.045733\.04128\.1460\.40860\.6850TaylorSeer\(𝒩=5,O=2\)\(\\mathcal\{N\}=5,O=2\)\[[18](https://arxiv.org/html/2606.26769#bib.bib18)\]7\.413\.20×\\times1\.056632\.81129\.1320\.67010\.3757ResilPhase\(𝒩=4,O=1\)\(\\mathcal\{N\}=4,O=1\)6\.623\.58×\\times1\.064732\.87430\.0200\.72160\.2989
- •Note:ResilPhase results were obtained utilizing Balanced Mapping\.
### 4\.2Text\-to\-Image Generation
Figure 4:Qualitative comparison on FLUX\.1\-dev\. While baseline methods suffer from visual artifacts and semantic errors at high speedup ratios, ResilPhase preserves both high fidelity and prompt accuracy, showing its superiority\.As shown in Tab\.[1](https://arxiv.org/html/2606.26769#S4.T1), ResilPhase consistently outperforms competitors in both speed and generation quality, with its advantage widening at higher acceleration ratios\. Notably, at an aggressive 4\.97x speedup on FLUX\.1\-dev, ResilPhase maintains an exceptional ImageReward of 1\.0258, delivering a\>\>64% improvement over the leading forecasting baseline TaylorSeer \(0\.6241\), alongside a 52% lower LPIPS score\. This quantitative robustness translates directly to superior visual fidelity\. As visually confirmed in Fig\.[4](https://arxiv.org/html/2606.26769#S4.F4), while baseline methods suffer from severe distortions and artifacts at extreme speedups, ResilPhase successfully preserves detailed structures and color accuracy, establishing a new state\-of\-the\-art balance between inference acceleration and exceptional image quality\.
### 4\.3Text\-to\-Video Generation
As shown in Tab\.[2](https://arxiv.org/html/2606.26769#S4.T2), ResilPhase achieves the optimal speed\-quality balance across all acceleration levels\. At an aggressive∼\\sim5×\\timesspeedup, it outperforms the SOTA TeaCache with a higher VBench\[[10](https://arxiv.org/html/2606.26769#bib.bib10)\]score \(79\.78\) and a 10% lower LPIPS, an advantage that widens further at the∼\\sim4\.3×\\timestier\. These quantitative gains translate directly to superior visual fidelity \(Fig\.[5](https://arxiv.org/html/2606.26769#S4.F5)\)\. While competing methods suffer from severe artifacts like object blur, incorrect scales, and distorted motions at high speeds, ResilPhase consistently delivers smooth, temporally coherent videos with crisp, semantically accurate details\.
Figure 5:Qualitative comparison of text\-to\-video generation methods\. While competing methods suffer from object distortion, missing details, and motion inconsistency, ResilPhase maintains superior temporal coherence and visual quality\.Table 2:Quantitative comparison for HunyuanVideo text\-to\-video generation, grouped by two tiers of similar latency acceleration ratio\.MethodSpeed↑\\uparrowPSNR↑\\uparrowSSIM↑\\uparrowLPIPS↓\\downarrowVBench↑\\uparrow50\-steps1\.00×\\times\-\-\-80\.87FORA\(𝒩=6\)\(\\mathcal\{N\}=6\)4\.87×\\times15\.8710\.60730\.458478\.70TeaCache\(l=0\.4\)\(\{l\}=0\.4\)4\.62×\\times17\.9230\.65470\.376079\.77ToCa\(𝒩=10,R=90%\)\(\\mathcal\{N\}=10,R=90\\%\)4\.86×\\times17\.5810\.58570\.450176\.37ClusCa\(𝒩=9,O=1\)\(\\mathcal\{N\}=9,O=1\)4\.88×\\times14\.6900\.53530\.513976\.98SpeCa\(τ0=1\.5,β=0\.2\)\(\{\\tau\_\{0\}\}=1\.5,\{\\beta\}=0\.2\)4\.65×\\times16\.4610\.58830\.421979\.59TaylorSeer\(𝒩=7,O=1\)\(\\mathcal\{N\}=7,O=1\)4\.63×\\times15\.5200\.56410\.458179\.07ResilPhase\(𝒩=6,O=1\)\(\\mathcal\{N\}=6,O=1\)4\.98×\\times18\.9200\.67090\.334179\.78FORA\(𝒩=4\)\(\\mathcal\{N\}=4\)3\.57×\\times16\.5820\.62440\.401080\.10TeaCache\(l=0\.3\)\(\{l\}=0\.3\)4\.01×\\times19\.0500\.68460\.320780\.38ToCa\(𝒩=7,R=90%\)\(\\mathcal\{N\}=7,R=90\\%\)4\.09×\\times18\.2990\.62460\.368879\.00ClusCa\(𝒩=6,O=1\)\(\\mathcal\{N\}=6,O=1\)4\.05×\\times16\.1160\.58810\.423079\.74SpeCa\(τ0=1\.2,β=0\.1\)\(\{\\tau\_\{0\}\}=1\.2,\{\\beta\}=0\.1\)4\.20×\\times16\.4880\.58880\.419879\.77TaylorSeer\(𝒩=5,O=1\)\(\\mathcal\{N\}=5,O=1\)3\.70×\\times17\.1170\.63160\.369080\.27ResilPhase\(𝒩=5,O=1\)\(\\mathcal\{N\}=5,O=1\)4\.28×\\times19\.4810\.69920\.295480\.42
Table 3:Quantitative comparison for DiT\-XL/2 class\-to\-image generation on ImageNet\.MethodSpeed↑\\uparrowFID↓\\downarrowsFID↓\\downarrowIS↑\\uparrowDDIM\-50 steps1\.00×\\times2\.3674\.387236\.06DDIM\-12 steps4\.20×\\times8\.4658\.664180\.41FORA\(𝒩=8\\mathcal\{N\}=8\)3\.64×\\times16\.69423\.396128\.08ToCa\(𝒩=12,R=93%\\mathcal\{N\}=12,R=93\\%\)3\.53×\\times34\.11030\.34185\.51ClusCa\(𝒩=12,K=4,O=2\)\(\\mathcal\{N\}=12,K=4,O=2\)3\.84×\\times15\.2069\.405124\.03SpeCa\(τ0=1\.5,β=0\.5\)\(\{\\tau\_\{0\}\}=1\.5,\{\\beta\}=0\.5\)4\.36×\\times6\.8668\.250171\.68TaylorSeer\(𝒩=13,O=2\)\(\\mathcal\{N\}=13,O=2\)2\.72×\\times15\.41516\.005120\.76ResilPhase\(𝒩=5,O=3\)\(\\mathcal\{N\}=5,O=3\)4\.41×\\times2\.8325\.026219\.19DDIM\-20 steps2\.50×\\times3\.8585\.201216\.76FORA\(𝒩=4\\mathcal\{N\}=4\)2\.60×\\times4\.7707\.591210\.98ToCa\(𝒩=5,R=93%\\mathcal\{N\}=5,R=93\\%\)2\.67×\\times6\.3217\.027196\.06ClusCa\(𝒩=6,K=8,O=4\)\(\\mathcal\{N\}=6,K=8,O=4\)2\.61×\\times3\.2515\.207217\.60SpeCa\(τ0=0\.1,β=0\.5\)\(\{\\tau\_\{0\}\}=0\.1,\{\\beta\}=0\.5\)2\.65×\\times2\.6334\.848231\.32TaylorSeer\(𝒩=8,O=3\)\(\\mathcal\{N\}=8,O=3\)2\.14×\\times4\.8077\.088197\.22ResilPhase\(𝒩=3,O=4\)\(\\mathcal\{N\}=3,O=4\)2\.78×\\times2\.3474\.672233\.57
### 4\.4Class\-Conditional Image Generation
On DiT\-XL/2, ResilPhase uniquely improves quality while accelerating\. At a2\.78×2\.78\\timesspeedup, it achieves an FID of 2\.347, outperforming both the full 50\-step DDIM baseline \(2\.367\) and the SOTA competitor ClusCa \(3\.251\)\. Its robustness shines at a4\.41×4\.41\\timesspeedup, maintaining a low FID of 2\.832 while most caching baselines collapse \(FID\>15\>15\)\. Even against the strongest competitor in this tier \(SpeCa, FID 6\.866\), ResilPhase delivers a massive 58% FID reduction\. Consistently leading in sFID and IS, ResilPhase proves highly effective at preserving original quality under aggressive acceleration\.
Table 4:Ablation study of ResilPhase components on the FLUX\-1\.dev\.ConfigurationPredictive ObjectivePhase MappingSpeed↑\\uparrowImageReward↑\\uparrowCLIP↑\\uparrowPSNR↑\\uparrowSSIM↑\\uparrowLPIPS↓\\downarrowGlobal DriftLayer\-wise OutputChebyshevBalanceDrawBenchScore✔3\.08×\\times1\.054632\.93729\.4350\.69480\.3323✔3\.57×\\times1\.057332\.92929\.4470\.70030\.3284Lagrange✔✔3\.08×\\times1\.055932\.97529\.4360\.69470\.3324\(𝒩=4,O=1\)\(\\mathcal\{N\}=4,O=1\)✔✔3\.08×\\times1\.058032\.87429\.9290\.71390\.2995✔✔3\.58×\\times1\.059532\.91429\.4480\.70040\.3281✔✔3\.58×\\times1\.064732\.91830\.0200\.72160\.2989✔3\.64×\\times1\.040432\.68229\.0130\.66590\.3853✔4\.17×\\times1\.055732\.75829\.0930\.67860\.3808Lagrange✔✔3\.64×\\times1\.040932\.65529\.0140\.66590\.3853\(𝒩=5,O=1\)\(\\mathcal\{N\}=5,O=1\)✔✔3\.64×\\times1\.052832\.99829\.4720\.69550\.3343✔✔4\.17×\\times1\.051732\.79729\.0340\.67150\.3811✔✔4\.17×\\times1\.059132\.90129\.5560\.70240\.3342✔4\.06×\\times1\.009832\.80328\.7060\.61840\.4479✔4\.96×\\times1\.015032\.65828\.7340\.62920\.4388Lagrange✔✔4\.06×\\times1\.012632\.68828\.7040\.61830\.4478\(𝒩=6,O=1\)\(\\mathcal\{N\}=6,O=1\)✔✔4\.05×\\times1\.034732\.70729\.1920\.65260\.3831✔✔4\.97×\\times1\.018632\.70828\.7340\.62910\.4388✔✔4\.97×\\times1\.025832\.84729\.5360\.66550\.3834
- •Note:Lagrange refers to the baseline configuration utilizing only Barycentric Lagrange extrapolation, without any phase mapping\.
### 4\.5Ablation Study
Ablation Study of ResilPhase Components\.Tab\.[4](https://arxiv.org/html/2606.26769#S4.T4)shows forecasting Global Drift outperforms layer\-wise prediction by eliminating cascading errors\. Phase Mapping further boosts quality, with Balanced Mapping proving more suitable for text\-to\-image tasks than Chebyshev Mapping\. We also compared Global Drift against predicting the Final Output, an alternative avoiding intermediate caching\. Fig\.[6](https://arxiv.org/html/2606.26769#S4.F6)confirms Global Drift consistently maintains superiority over the Final Output across extrapolation intervals\.
Figure 6:Comparing the performance of Global Drift \(GD\) and Final Output \(FO\) predictions across extrapolation intervals \(NN\) with and without Balance Phase Mapping\.Generalizability of the Phase Mapping\.Integrating Phase Mapping into TaylorSeer and HiCache consistently boosts ImageReward \(Fig\.[7](https://arxiv.org/html/2606.26769#S4.F7)\) and other perceptual metrics\. Furthermore, our appendix provides theoretical proofs that this mechanism strictly reduces their extrapolation error bounds, which strongly confirms its plug\-and\-play generalizability for existing polynomial accelerators\.
Figure 7:Comparing the performance of TaylorSeer and HiCache with different phase mapping strategies on FLUX\.1\-dev across acceleration intervals\.
## 5Conclusion
We introduce ResilPhase, a noise\-resilient, training\-free acceleration framework for Diffusion Transformers\. It overcomes spatial, temporal, and numerical bottlenecks in existing paradigms by synergizing an ODE\-aligned Global Drift target to eliminate cascading errors, a derivative\-free Barycentric Lagrange extrapolator to bypass gradient noise, and a bounded Phase Mapping mechanism to suppress extrapolation instability\. Experiments confirm it achieves state\-of\-the\-art speed and quality at aggressive acceleration ratios, establishing a robust real\-time inference paradigm and plug\-and\-play stabilizer for existing accelerators\.
## Acknowledgements
This study is partially supported by the National Natural Science Foundation of China \(Grant No\.62504204\)\.
## References
- \[1\]Blattmann, A\., Dockhorn, T\., Kulal, S\., Mendelevitch, D\., Kilian, M\., Lorenz, D\., Levi, Y\., English, Z\., Voleti, V\., Letts, A\., et al\.: Stable video diffusion: Scaling latent video diffusion models to large datasets\. arXiv preprint arXiv:2311\.15127 \(2023\)
- \[2\]Bu, J\., Ling, P\., Zhou, Y\., Wang, Y\., Zang, Y\., Lin, D\., Wang, J\.: Dicache: Let diffusion model determine its own cache\. arXiv preprint arXiv:2508\.17356 \(2025\)
- \[3\]Chen, P\., Shen, M\., Ye, P\., Cao, J\., Tu, C\., Bouganis, C\.S\., Zhao, Y\., Chen, T\.:δ\\delta\-dit: A training\-free acceleration method tailored for diffusion transformers\. arXiv preprint arXiv:2406\.01125 \(2024\)
- \[4\]Fang, G\., Ma, X\., Wang, X\.: Structural pruning for diffusion models \(2023\),[https://arxiv\.org/abs/2305\.10924](https://arxiv.org/abs/2305.10924)
- \[5\]Feng, L\., Zheng, S\., Liu, J\., Lin, Y\., Zhou, Q\., Cai, P\., Wang, X\., Chen, J\., Zou, C\., Ma, Y\., et al\.: Hicache: Training\-free acceleration of diffusion models via hermite polynomial\-based feature caching\. arXiv preprint arXiv:2508\.16984 \(2025\)
- \[6\]Guo, Y\., Yan, Z\., Yu, X\., Kong, Q\., Xie, J\., Luo, K\., Zeng, D\., Wu, Y\., Jia, Z\., Shi, Y\.: Hardware design and the fairness of a neural network\. Nature Electronics7\(8\), 714–723 \(2024\)
- \[7\]Hessel, J\., Holtzman, A\., Forbes, M\., Le Bras, R\., Choi, Y\.: Clipscore: A reference\-free evaluation metric for image captioning\. In: Proceedings of the 2021 conference on empirical methods in natural language processing\. pp\. 7514–7528 \(2021\)
- \[8\]Heusel, M\., Ramsauer, H\., Unterthiner, T\., Nessler, B\., Hochreiter, S\.: Gans trained by a two time\-scale update rule converge to a local nash equilibrium\. Advances in neural information processing systems30\(2017\)
- \[9\]Ho, J\., Jain, A\., Abbeel, P\.: Denoising diffusion probabilistic models\. Advances in neural information processing systems33, 6840–6851 \(2020\)
- \[10\]Huang, Z\., He, Y\., Yu, J\., Zhang, F\., Si, C\., Jiang, Y\., Zhang, Y\., Wu, T\., Jin, Q\., Chanpaisit, N\., et al\.: Vbench: Comprehensive benchmark suite for video generative models\. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition\. pp\. 21807–21818 \(2024\)
- \[11\]Kim, S\., Lee, H\., Cho, W\., Park, M\., Ro, W\.W\.: Ditto: Accelerating diffusion model via temporal value similarity\. In: 2025 IEEE International Symposium on High Performance Computer Architecture \(HPCA\)\. pp\. 338–352\. IEEE \(2025\)
- \[12\]Labs, B\.F\., Batifol, S\., Blattmann, A\., Boesel, F\., Consul, S\., Diagne, C\., Dockhorn, T\., English, J\., English, Z\., Esser, P\., et al\.: Flux\. 1 kontext: Flow matching for in\-context image generation and editing in latent space\. arXiv preprint arXiv:2506\.15742 \(2025\)
- \[13\]Li, S\., Hu, T\., van de Weijer, J\., Khan, F\.S\., Liu, T\., Li, L\., Yang, S\., Wang, Y\., Cheng, M\.M\., Yang, J\.: Faster diffusion: Rethinking the role of the encoder for diffusion model inference\. Advances in Neural Information Processing Systems37, 85203–85240 \(2024\)
- \[14\]Li, X\., Liu, Y\., Lian, L\., Yang, H\., Dong, Z\., Kang, D\., Zhang, S\., Keutzer, K\.: Q\-diffusion: Quantizing diffusion models\. In: Proceedings of the IEEE/CVF International Conference on Computer Vision\. pp\. 17535–17545 \(2023\)
- \[15\]Li, Y\., Wang, H\., Jin, Q\., Hu, J\., Chemerys, P\., Fu, Y\., Wang, Y\., Tulyakov, S\., Ren, J\.: Snapfusion: Text\-to\-image diffusion model on mobile devices within two seconds\. Advances in Neural Information Processing Systems36, 20662–20678 \(2023\)
- \[16\]Liu, F\., Zhang, S\., Wang, X\., Wei, Y\., Qiu, H\., Zhao, Y\., Zhang, Y\., Ye, Q\., Wan, F\.: Timestep embedding tells: It’s time to cache for video diffusion model\. In: Proceedings of the Computer Vision and Pattern Recognition Conference\. pp\. 7353–7363 \(2025\)
- \[17\]Liu, J\., Cai, P\., Zhou, Q\., Lin, Y\., Kong, D\., Huang, B\., Pan, Y\., Xu, H\., Zou, C\., Tang, J\., et al\.: Freqca: Accelerating diffusion models via frequency\-aware caching\. arXiv preprint arXiv:2510\.08669 \(2025\)
- \[18\]Liu, J\., Zou, C\., Lyu, Y\., Chen, J\., Zhang, L\.: From reusing to forecasting: Accelerating diffusion models with taylorseers\. arXiv preprint arXiv:2503\.06923 \(2025\)
- \[19\]Liu, J\., Zou, C\., Lyu, Y\., Ren, F\., Wang, S\., Li, K\., Zhang, L\.: Speca: Accelerating diffusion transformers with speculative feature caching\. In: Proceedings of the 33rd ACM International Conference on Multimedia\. pp\. 10024–10033 \(2025\)
- \[20\]Liu, X\., Gong, C\., et al\.: Flow straight and fast: Learning to generate and transfer data with rectified flow\. In: The Eleventh International Conference on Learning Representations
- \[21\]Lu, C\., Zhou, Y\., Bao, F\., Chen, J\., Li, C\., Zhu, J\.: Dpm\-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps\. Advances in neural information processing systems35, 5775–5787 \(2022\)
- \[22\]Lu, C\., Zhou, Y\., Bao, F\., Chen, J\., Li, C\., Zhu, J\.: Dpm\-solver\+\+: Fast solver for guided sampling of diffusion probabilistic models\. Machine Intelligence Research pp\. 1–22 \(2025\)
- \[23\]Ma, X\., Wang, Y\., Jia, G\., Chen, X\., Liu, Z\., Li, Y\.F\., Chen, C\., Qiao, Y\.: Latte: Latent diffusion transformer for video generation\. arXiv e\-prints pp\. arXiv–2401 \(2024\)
- \[24\]Ma, X\., Fang, G\., Wang, X\.: Deepcache: Accelerating diffusion models for free\. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition\. pp\. 15762–15772 \(2024\)
- \[25\]Peebles, W\., Xie, S\.: Scalable diffusion models with transformers\. In: Proceedings of the IEEE/CVF international conference on computer vision\. pp\. 4195–4205 \(2023\)
- \[26\]Podell, D\., English, Z\., Lacey, K\., Blattmann, A\., Dockhorn, T\., Müller, J\., Penna, J\., Rombach, R\.: Sdxl: Improving latent diffusion models for high\-resolution image synthesis\. arXiv preprint arXiv:2307\.01952 \(2023\)
- \[27\]Rombach, R\., Blattmann, A\., Lorenz, D\., Esser, P\., Ommer, B\.: High\-resolution image synthesis with latent diffusion models\. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition\. pp\. 10684–10695 \(2022\)
- \[28\]Russakovsky, O\., Deng, J\., Su, H\., Krause, J\., Satheesh, S\., Ma, S\., Huang, Z\., Karpathy, A\., Khosla, A\., Bernstein, M\., et al\.: Imagenet large scale visual recognition challenge\. International journal of computer vision115\(3\), 211–252 \(2015\)
- \[29\]Saharia, C\., Chan, W\., Saxena, S\., Li, L\., Whang, J\., Denton, E\.L\., Ghasemipour, K\., Gontijo Lopes, R\., Karagol Ayan, B\., Salimans, T\., et al\.: Photorealistic text\-to\-image diffusion models with deep language understanding\. Advances in neural information processing systems35, 36479–36494 \(2022\)
- \[30\]Salimans, T\., Ho, J\.: Progressive distillation for fast sampling of diffusion models\. In: International Conference on Learning Representations
- \[31\]Selvaraju, P\., Ding, T\., Chen, T\., Zharkov, I\., Liang, L\.: Fora: Fast\-forward caching in diffusion transformer acceleration\. arXiv preprint arXiv:2407\.01425 \(2024\)
- \[32\]Shang, Y\., Yuan, Z\., Xie, B\., Wu, B\., Yan, Y\.: Post\-training quantization on diffusion models\. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition\. pp\. 1972–1981 \(2023\)
- \[33\]Song, J\., Meng, C\., Ermon, S\.: Denoising diffusion implicit models\. In: International Conference on Learning Representations
- \[34\]Song, J\., Meng, C\., Ermon, S\.: Denoising diffusion implicit models\. arXiv preprint arXiv:2010\.02502 \(2020\)
- \[35\]Song, Y\., Dhariwal, P\., Chen, M\., Sutskever, I\.: Consistency models \(2023\)
- \[36\]Sun, X\., Chen, Y\., Huang, Y\., Xie, R\., Zhu, J\., Zhang, K\., Li, S\., Yang, Z\., Han, J\., Shu, X\., et al\.: Hunyuan\-large: An open\-source moe model with 52 billion activated parameters by tencent\. arXiv preprint arXiv:2411\.02265 \(2024\)
- \[37\]Wang, G\., Cai, Y\., Li, L\., Peng, W\., Su, S\.: Pfdiff: Training\-free acceleration of diffusion models combining past and future scores\. arXiv preprint arXiv:2408\.08822 \(2024\)
- \[38\]Wang, Z\., Bovik, A\.C\., Sheikh, H\.R\., Simoncelli, E\.P\.: Image quality assessment: from error visibility to structural similarity\. IEEE transactions on image processing13\(4\), 600–612 \(2004\)
- \[39\]Wu, Y\., Yan, Z\., Yin, X\., He, L\., Zhuo, C\.: Anas: Software–hardware co\-design of approximate neural network accelerators via neural architecture search\. Integration104, 102469 \(2025\)
- \[40\]Xu, J\., Liu, X\., Wu, Y\., Tong, Y\., Li, Q\., Ding, M\., Tang, J\., Dong, Y\.: Imagereward: Learning and evaluating human preferences for text\-to\-image generation\. Advances in Neural Information Processing Systems36, 15903–15935 \(2023\)
- \[41\]Yin, J\., Chen, T\., Chen, Y\., Pei, G\., Shu, X\., Yao, Y\., Shen, F\.: Pca\-seg: Revisiting cost aggregation for open\-vocabulary semantic and part segmentation\. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\)\. pp\. 27633–27643 \(June 2026\)
- \[42\]Yin, J\., Jiang, X\., Chen, T\., Pei, G\., Yao, Y\., Shen, F\., Shen, H\.T\.: Depmatch: Boosting semi\-supervised semantic segmentation by exploring depth difference knowledge\. IEEE Transactions on Image Processing35, 3256–3270 \(2026\)
- \[43\]Zhang, R\., Isola, P\., Efros, A\.A\., Shechtman, E\., Wang, O\.: The unreasonable effectiveness of deep features as a perceptual metric\. In: Proceedings of the IEEE conference on computer vision and pattern recognition\. pp\. 586–595 \(2018\)
- \[44\]Zheng, S\., Feng, L\., Wang, X\., Zhou, Q\., Cai, P\., Zou, C\., Liu, J\., Lin, Y\., Chen, J\., Ma, Y\., et al\.: Forecast then calibrate: Feature caching as ode for efficient diffusion transformers\. arXiv preprint arXiv:2508\.16211 \(2025\)
- \[45\]Zheng, Z\., Peng, X\., Yang, T\., Shen, C\., Li, S\., Liu, H\., Zhou, Y\., Li, T\., You, Y\.: Open\-sora: Democratizing efficient video production for all\. arXiv preprint arXiv:2412\.20404 \(2024\)
- \[46\]Zheng, Z\., Wang, X\., Zou, C\., Wang, S\., Zhang, L\.: Compute only 16 tokens in one timestep: Accelerating diffusion transformers with cluster\-driven feature caching\. In: Proceedings of the 33rd ACM International Conference on Multimedia\. pp\. 10181–10189 \(2025\)
- \[47\]Zhou, X\., Zhang, Y\., Sun, S\., Wen, C\., Yan, Z\.: Deploying edge llms for wafer defect detection in chip manufacturing\. In: 2025 International Symposium of Electronics Design Automation \(ISEDA\)\. pp\. 7–11\. IEEE \(2025\)
- \[48\]Zhu, H\., Tang, D\., Liu, J\., Lu, M\., Zheng, J\., Peng, J\., Li, D\., Wang, Y\., Jiang, F\., Tian, L\., et al\.: Dip\-go: A diffusion pruner via few\-step gradient optimization\. Advances in Neural Information Processing Systems37, 92581–92604 \(2024\)
- \[49\]Zou, C\., Liu, X\., Liu, T\., Huang, S\., Zhang, L\.: Accelerating diffusion transformers with token\-wise feature caching\. arXiv preprint arXiv:2410\.05317 \(2024\)Similar Articles
Least-Action-Guided Diffusion for Physical Extrapolation
Introduces LAPG, a diffusion framework guided by the principle of least action to improve physical consistency during inference for out-of-distribution extrapolation tasks in physics.
DiRecT: Safe Diffusion-Based Planning via Receding-Horizon Denoising
DiRecT introduces a training-free algorithm for safe diffusion-based planning that enforces constraints only on final clean trajectories using receding-horizon denoising, improving safety and performance over existing methods.
Differencing the Diffusion Trajectory toward Uncertain Components for Time Series Forecasting
This paper proposes DiffDiff, a diffusion framework for probabilistic time series forecasting that embeds predictability asymmetry into the diffusion trajectory, outperforming six diffusion baselines on seven benchmarks across four prediction horizons.
Coarse-to-Fine Multi-Resolution Diffusion Models for Trajectory Generation in Urban Systems
The paper proposes MR-Traj, a multi-resolution diffusion framework for synthetic trajectory generation in urban systems, which captures complex spatial-temporal dependencies at multiple resolutions and improves performance in fine-grained mobility modeling for downstream tasks.
Read the Trace, Steer the Path: Trajectory-Aware Reinforcement Learning for Diffusion Language Models
This paper introduces CAPR (Cached-Amortized Path Refinement), a reinforcement learning algorithm for diffusion large language models that extracts tree-like supervision signals from the denoising trace without the compute cost of full tree rollouts. CAPR achieves state-of-the-art performance on reasoning benchmarks like GSM8K, Math500, Sudoku, and Countdown at roughly 0.75x the cost of flat rollouts.