Differencing the Diffusion Trajectory toward Uncertain Components for Time Series Forecasting
Summary
This paper proposes DiffDiff, a diffusion framework for probabilistic time series forecasting that embeds predictability asymmetry into the diffusion trajectory, outperforming six diffusion baselines on seven benchmarks across four prediction horizons.
View Cached Full Text
Cached at: 07/28/26, 06:26 AM
# Differencing the Diffusion Trajectory toward Uncertain Components for Time Series Forecasting
Source: [https://arxiv.org/html/2607.22599](https://arxiv.org/html/2607.22599)
11footnotetext:Corresponding author\.Chen Su♠, Yuanhe Tian, Yan Song♠∗ ♠University of Science and Technology of China Zhongguancun Academy ♠suchen4565@mail\.ustc\.edu\.cnyhtian94@gmail\.com♠clksong@gmail\.com
###### Abstract
Diffusion models have become a widely used framework for probabilistic time series forecasting, modeling the distribution of future values given an observed history\. In time series forecasting, however, the future continues the observed history, creating an asymmetry the standard diffusion process leaves unaddressed, with slowly\-varying content largely determined by the observed continuity while higher\-frequency dynamics carry most of the residual uncertainty\. Existing diffusion\-based forecasters decouple this asymmetry through an external rule before generation, leaving the corruption trajectory blind to which parts of the target the history can already anchor\. We proposeDiffDiff, a diffusion framework that embeds this predictability asymmetry into the diffusion trajectory itself, so that a single end\-to\-end diffusion process becomes aware of which parts of the target the history can already anchor\.DiffDiffmakes the forward operator step\-dependent so that the noisy intermediate state progressively shifts from the target itself toward its second\-order differenced structure, while a conditioning pathway supplies the denoiser with both value\-domain and differential history information balanced by a stage\-adaptive gate at each diffusion step\. The terminal distribution approaches a standard Gaussian, preserving compatibility with existing samplers\. On seven benchmarks across four prediction horizons,DiffDiffoutperforms six diffusion baselines, and our analysis confirms thatDiffDiffconcentrates the diffusion’s generative effort on the most uncertain components of the target while relieving it from rebuilding the history\-anchored content\.111The code is available at[https://github\.com/synlp/DiffDiff](https://github.com/synlp/DiffDiff)\.
## 1Introduction
Probabilistic time series forecasting aims to predict the conditional distribution of future values from a historical observation window, providing both accurate point estimates and calibrated uncertainty quantification\. Reliable distributional forecasts are essential in risk\-sensitive applications such as energy dispatch\[[39](https://arxiv.org/html/2607.22599#bib.bib56),[8](https://arxiv.org/html/2607.22599#bib.bib72)\], financial hedging\[[6](https://arxiv.org/html/2607.22599#bib.bib57),[11](https://arxiv.org/html/2607.22599#bib.bib73)\], and extreme weather preparedness\[[16](https://arxiv.org/html/2607.22599#bib.bib71),[21](https://arxiv.org/html/2607.22599#bib.bib74)\], where decisions depend not only on expected outcomes but also on the range of plausible futures\. Denoising diffusion probabilistic models \(DDPMs\)\[[49](https://arxiv.org/html/2607.22599#bib.bib10),[17](https://arxiv.org/html/2607.22599#bib.bib9),[51](https://arxiv.org/html/2607.22599#bib.bib39),[46](https://arxiv.org/html/2607.22599#bib.bib40),[41](https://arxiv.org/html/2607.22599#bib.bib41),[35](https://arxiv.org/html/2607.22599#bib.bib42),[14](https://arxiv.org/html/2607.22599#bib.bib69),[64](https://arxiv.org/html/2607.22599#bib.bib58)\]have recently emerged as a powerful generative framework for this task because they model complex predictive distributions without restrictive parametric assumptions\.
Within the forecasting target, however, predictive uncertainty is not uniformly distributed\. Components that remain strongly correlated with the observed history, such as smooth trends, tend to carry little remaining uncertainty, whereas the components the history under\-determines, such as rapid fluctuations beyond what can be extrapolated from the past, concentrate the residual uncertainty and are where a generative model’s capacity is best spent\[[26](https://arxiv.org/html/2607.22599#bib.bib23),[61](https://arxiv.org/html/2607.22599#bib.bib22),[52](https://arxiv.org/html/2607.22599#bib.bib70)\]\. Recent diffusion\-based forecasters have begun to exploit this asymmetry by decoupling the target into a deterministic component and an uncertain component before any diffusion modeling, applying diffusion only to the latter\[[26](https://arxiv.org/html/2607.22599#bib.bib23),[67](https://arxiv.org/html/2607.22599#bib.bib24)\]\. Such external decoupling validates that this heterogeneity matters for forecasting, but it also localizes the asymmetry outside the diffusion pipeline, leaving the diffusion process itself task\-agnostic to the structure it is meant to model\. Modifications to the diffusion process itself have also been explored within forecasting pipelines, such as introducing learnable, data\-aware components into the forward process\[[28](https://arxiv.org/html/2607.22599#bib.bib14),[58](https://arxiv.org/html/2607.22599#bib.bib68),[44](https://arxiv.org/html/2607.22599#bib.bib26)\], matching the diffusion endpoint variance to the data’s non\-stationary variation\[[65](https://arxiv.org/html/2607.22599#bib.bib25)\], or slowing the decay of low\-frequency content along the diffusion trajectory\[[60](https://arxiv.org/html/2607.22599#bib.bib8)\]\. Yet none of these designs reshapes what the trajectory preserves of the target so that intermediate states emphasize the parts the history cannot already supply\. The diffusion’s capacity is therefore spent recovering both the parts the history could already anchor and those it cannot, instead of being focused where forecasting uncertainty actually concentrates\.
We proposeDiffDiff\(Differencing Diffusion\), a forecasting diffusion framework that embeds this asymmetry directly into the diffusion trajectory so that the diffusion process itself becomes aware of which parts of the target the history can already anchor, rather than offloading the asymmetry to an external decomposition\. The forward operator is made step\-dependent so that, as diffusion progresses, components readily reproducible from history are progressively suppressed in the noisy state, while components the history under\-determines remain salient and absorb most of the trajectory’s representational budget\. To compensate for this suppression, the conditioning pathway combines a value\-domain encoding of the observed history with a differencing branch that captures its temporal evolution, and a stage\-adaptive gate then re\-supplies the suppressed content to the denoiser at the diffusion stages where the noisy target alone is no longer informative\. The diffusion process thus remains a single end\-to\-end estimator over the full target, while its modeling capacity is reallocated along the trajectory, with history\-anchored content sustained through conditioning and diffusion concentrating on components for which the history is least informative\. Although the intermediate states are reshaped, the terminal distribution approaches a standard Gaussian, preserving compatibility with existing noise schedules and accelerated samplers\.
## 2Related Work
### 2\.1Diffusion Models for Time Series Forecasting
Diffusion\-based forecasting evolved from autoregressive designs\[[43](https://arxiv.org/html/2607.22599#bib.bib1),[34](https://arxiv.org/html/2607.22599#bib.bib12)\]to non\-autoregressive successors\[[57](https://arxiv.org/html/2607.22599#bib.bib2),[2](https://arxiv.org/html/2607.22599#bib.bib51),[48](https://arxiv.org/html/2607.22599#bib.bib13),[24](https://arxiv.org/html/2607.22599#bib.bib52),[55](https://arxiv.org/html/2607.22599#bib.bib67),[53](https://arxiv.org/html/2607.22599#bib.bib61)\]that avoid error accumulation\. Subsequent work enriched the conditioning pathway through multi\-resolution\[[47](https://arxiv.org/html/2607.22599#bib.bib16),[12](https://arxiv.org/html/2607.22599#bib.bib17),[54](https://arxiv.org/html/2607.22599#bib.bib60)\], seasonal\-trend\[[66](https://arxiv.org/html/2607.22599#bib.bib18)\], variational autoencoder \(VAE\)\-hybrid\[[27](https://arxiv.org/html/2607.22599#bib.bib19)\], retrieval\-augmented\[[30](https://arxiv.org/html/2607.22599#bib.bib15),[5](https://arxiv.org/html/2607.22599#bib.bib53),[20](https://arxiv.org/html/2607.22599#bib.bib54)\]designs\. Broader multimodal representation studies also preserve heterogeneous evidence through contrastive alignment, contextual augmentation, or adaptive fusion before prediction\[[42](https://arxiv.org/html/2607.22599#bib.bib63),[1](https://arxiv.org/html/2607.22599#bib.bib62),[13](https://arxiv.org/html/2607.22599#bib.bib64),[56](https://arxiv.org/html/2607.22599#bib.bib65)\]\. Components strongly correlated with the history are recoverable by deterministic predictors\[[38](https://arxiv.org/html/2607.22599#bib.bib43),[32](https://arxiv.org/html/2607.22599#bib.bib44),[62](https://arxiv.org/html/2607.22599#bib.bib47),[40](https://arxiv.org/html/2607.22599#bib.bib48),[33](https://arxiv.org/html/2607.22599#bib.bib49),[70](https://arxiv.org/html/2607.22599#bib.bib46),[68](https://arxiv.org/html/2607.22599#bib.bib50),[31](https://arxiv.org/html/2607.22599#bib.bib45)\], while the residual carries most of the predictive entropy\. D3U\[[26](https://arxiv.org/html/2607.22599#bib.bib23)\]freezes a point predictor over trend and seasonality and CDPM\[[67](https://arxiv.org/html/2607.22599#bib.bib24)\]routes the trend through a polynomial module, in each case pre\-cleaning the residual by a fixed external rule that locks the deterministic\-uncertain boundary before diffusion runs\. A complementary strand modifies the diffusion process itself end\-to\-end\. TMDM\[[28](https://arxiv.org/html/2607.22599#bib.bib14)\]and CN\-Diff\[[44](https://arxiv.org/html/2607.22599#bib.bib26)\]reshape the terminal prior, with TMDM centering it on a transformer point forecast and CN\-Diff training a learned multilayer perceptron \(MLP\) forward transformation jointly with a conditional prior\. NsDiff\[[65](https://arxiv.org/html/2607.22599#bib.bib25)\]learns a history\-conditional per\-cell diagonal noise variance, but the reshaping acts only on the noise while leaving the target encoding structurally fixed\. MA\-TSD\[[60](https://arxiv.org/html/2607.22599#bib.bib8)\]adopts a non\-isotropic forward operator built from progressive moving\-average smoothing that preserves low\-frequency content across the chain\.DiffDifflikewise reshapes what the intermediate state encodes, but redirects the trajectory’s budget toward components the history under\-determines while the conditioning pathway adaptively re\-supplies the suppressed content\. The asymmetry between deterministic and uncertain content thus emerges along the trajectory itself rather than being preset by an external decomposition\.
### 2\.2Non\-Isotropic Diffusion Processes
Several works have explored diffusion models with non\-isotropic forward processes in domains beyond time series\. Blurring Diffusion\[[18](https://arxiv.org/html/2607.22599#bib.bib27)\]and Inverse Heat Dissipation\[[45](https://arxiv.org/html/2607.22599#bib.bib28)\]replace isotropic Gaussian corruption with discrete cosine transform \(DCT\)\-domain frequency decay and heat\-equation partial differential equation \(PDE\) dynamics, respectively, for image generation\[[19](https://arxiv.org/html/2607.22599#bib.bib59)\]\. Soft Diffusion\[[9](https://arxiv.org/html/2607.22599#bib.bib29)\]formulates a general training objective for arbitrary linear corruptions combined with Gaussian noise, while Cold Diffusion\[[4](https://arxiv.org/html/2607.22599#bib.bib30)\]removes the Gaussian noise entirely and inverts arbitrary deterministic degradations such as blurring and masking\. Whitened Score Diffusion\[[3](https://arxiv.org/html/2607.22599#bib.bib31)\]enables arbitrary Gaussian forward processes with non\-identity covariance via whitened score matching, targeting imaging inverse problems\.DiffDiffspecializes this non\-isotropic framework for forecasting, where the operator is shaped by the within\-target predictability heterogeneity inherent to forecasting targets rather than by image\-domain priors such as blur or frequency decay\.
## 3Method
Given an observed history window𝐱hist\\mathbf\{x\}\_\{\\text\{hist\}\}, probabilistic time series forecasting aims to model the conditional distributionp\(𝐱fut∣𝐱hist\)p\(\\mathbf\{x\}\_\{\\text\{fut\}\}\\mid\\mathbf\{x\}\_\{\\text\{hist\}\}\)of a future window𝐱fut\\mathbf\{x\}\_\{\\text\{fut\}\}\. To learn this distribution with diffusion modeling, we denote the clean forecasting target by𝐱0\\mathbf\{x\}\_\{0\}and define a forward processq\(𝐱t∣𝐱0\)q\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{0\}\)that progressively transforms𝐱0\\mathbf\{x\}\_\{0\}into noisy latent states𝐱t\\mathbf\{x\}\_\{t\}over diffusion stepst=1,…,Tt=1,\\ldots,T, with the terminal state𝐱T\\mathbf\{x\}\_\{T\}approaching a standard Gaussian prior𝒩\(𝟎,𝐈\)\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)\. A parameterized reverse processpθ\(𝐱t−1∣𝐱t,𝐱hist\)p\_\{\\theta\}\(\\mathbf\{x\}\_\{t\-1\}\\mid\\mathbf\{x\}\_\{t\},\\mathbf\{x\}\_\{\\text\{hist\}\}\), implemented by a denoising networkfθ\(𝐱t,t,𝐱hist\)f\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t,\\mathbf\{x\}\_\{\\text\{hist\}\}\), learns to invert this transformation, so that iteratively applying the reverse transitions from a terminal sample yields a draw frompθ\(𝐱0∣𝐱hist\)p\_\{\\theta\}\(\\mathbf\{x\}\_\{0\}\\mid\\mathbf\{x\}\_\{\\text\{hist\}\}\)\. As depicted in Figure[1](https://arxiv.org/html/2607.22599#S3.F1),DiffDiffcouples a step\-dependent forward process that gradually encodes differential structure into the noisy state with a stage\-adaptive conditional denoiser that fuses value\-domain and differential history information\.
Figure 1:Overall architecture ofDiffDiff\.Top: the differencing forward process gradually shifts the clean target𝐱0\\mathbf\{x\}\_\{0\}toward its differenced structure under the step\-dependent operator𝐀t\\mathbf\{A\}\_\{t\}, and the deterministic reverse step recovers𝐱s\\mathbf\{x\}\_\{s\}from the denoiser prediction𝐱^0\\hat\{\\mathbf\{x\}\}\_\{0\}\.Bottom: the stage\-adaptive conditional denoiser encodes both the value\-domain history𝐱hist\\mathbf\{x\}\_\{\\text\{hist\}\}and its first\-order differencesΔ𝐱hist\\Delta\\mathbf\{x\}\_\{\\text\{hist\}\}, fuses them through a sigmoid gate conditioned on the diffusion timestep, and outputs the value\-domain prediction along with auxiliary statistics for instance denormalization\.### 3\.1Forward Process with Progressive Differencing
The information that a denoiser must recover from a noisy intermediate state depends on what signal content that state preserves\. We design the forward process so that, as diffusion progresses, the noisy intermediate state retains less of the components readily reproducible from the observed history and more of the components the history under\-determines\. The diffusion’s modeling capacity is thereby aligned with the predictability heterogeneity within the forecasting target rather than spread uniformly across it\. Given a clean target sequence𝐱0\\mathbf\{x\}\_\{0\}, we define a generalized forward marginal
q\(𝐱t∣𝐱0\)=𝒩\(α¯t𝐀t𝐱0,\(1−α¯t\)𝐈\)q\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{0\}\)=\\mathcal\{N\}\\\!\\left\(\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathbf\{A\}\_\{t\}\\,\\mathbf\{x\}\_\{0\},\\;\(1\-\\bar\{\\alpha\}\_\{t\}\)\\,\\mathbf\{I\}\\right\)\(1\)whereα¯t=∏τ=1t\(1−βτ\)\\bar\{\\alpha\}\_\{t\}=\\prod\_\{\\tau=1\}^\{t\}\(1\-\\beta\_\{\\tau\}\)is the cumulative noise coefficient following a cosine schedule\[[37](https://arxiv.org/html/2607.22599#bib.bib6)\]withα¯0=1\\bar\{\\alpha\}\_\{0\}=1andα¯T≈0\\bar\{\\alpha\}\_\{T\}\\approx 0, and𝐀t\\mathbf\{A\}\_\{t\}is a step\-dependent transition matrix that determines what signal representation is preserved at diffusion steptt\. Equivalently, by the reparameterization trick,
𝐱t=α¯t𝐀t𝐱0\+1−α¯tϵ,ϵ∼𝒩\(𝟎,𝐈\)\\mathbf\{x\}\_\{t\}=\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathbf\{A\}\_\{t\}\\,\\mathbf\{x\}\_\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\,\\boldsymbol\{\\epsilon\},\\qquad\\boldsymbol\{\\epsilon\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)\(2\)where the marginals at different steps share the same latent noiseϵ\\boldsymbol\{\\epsilon\}but apply different signal transformations𝐀t\\mathbf\{A\}\_\{t\}, defining a non\-Markovian forward family analogous to the denoising diffusion implicit model \(DDIM\) construction\[[50](https://arxiv.org/html/2607.22599#bib.bib7)\]\. When𝐀t=𝐈\\mathbf\{A\}\_\{t\}=\\mathbf\{I\}for alltt, all components are treated uniformly and the forward trajectory remains isotropic, which recovers the standard diffusion formulation as a special case\. Appendix[A](https://arxiv.org/html/2607.22599#A1)reviews this standard framework and its connection to our generalized formulation\. Specifically, for a length\-SSsequence𝐳=\[z1,z2,…,zS\]⊤\\mathbf\{z\}=\[z\_\{1\},z\_\{2\},\\ldots,z\_\{S\}\]^\{\\top\}, we define the second\-order differencing operator𝐃2\\mathbf\{D\}\_\{2\}entrywise as
\[𝐃2𝐳\]s=\{0,s=1zs−zs−1,s=2zs−2zs−1\+zs−2,s≥3\[\\mathbf\{D\}\_\{2\}\\mathbf\{z\}\]\_\{s\}=\\begin\{cases\}0,&s=1\\\\ z\_\{s\}\-z\_\{s\-1\},&s=2\\\\ z\_\{s\}\-2z\_\{s\-1\}\+z\_\{s\-2\},&s\\geq 3\\end\{cases\}\(3\)The first position removes the absolute level, the second retains the first\-order difference, and the remaining positions encode second\-order differences\. The transition matrix is then instantiated as
𝐀t=\(1−λt\)𝐈\+λt𝐃2\\mathbf\{A\}\_\{t\}=\(1\-\\lambda\_\{t\}\)\\,\\mathbf\{I\}\+\\lambda\_\{t\}\\,\\mathbf\{D\}\_\{2\}\(4\)whereλt∈\[0,λmax\]\\lambda\_\{t\}\\in\[0,\\lambda\_\{\\max\}\]increases monotonically with the diffusion step,λ0=0\\lambda\_\{0\}=0\(so that𝐀0=𝐈\\mathbf\{A\}\_\{0\}=\\mathbf\{I\}\), andλmax<1\\lambda\_\{\\max\}<1\. Early steps therefore remain close to the target𝐱0\\mathbf\{x\}\_\{0\}itself, while later steps increasingly emphasize differential structure\. The effective signal energy at stepttscales withα¯t‖𝐀t𝐱0‖2\\bar\{\\alpha\}\_\{t\}\\\|\\mathbf\{A\}\_\{t\}\\mathbf\{x\}\_\{0\}\\\|^\{2\}while the noise covariance\(1−α¯t\)𝐈\(1\-\\bar\{\\alpha\}\_\{t\}\)\\mathbf\{I\}remains isotropic, so𝐃2\\mathbf\{D\}\_\{2\}’s low\-frequency suppression and high\-frequency preservation produce a non\-uniform per\-frequency signal\-to\-noise ratio \(SNR\) that shifts the components retained by the noisy intermediate state toward those less reproducible from the observed history\. Appendices[B\.1](https://arxiv.org/html/2607.22599#A2.SS1)and[B\.5](https://arxiv.org/html/2607.22599#A2.SS5)formalize this frequency\-selective profile, and Section[5\.3](https://arxiv.org/html/2607.22599#S5.SS3)provides empirical validation\. Because𝐃2\\mathbf\{D\}\_\{2\}couples each position to its two preceding neighbors, defining the clean target as the future window alone would prevent it from encoding cross\-boundary curvature at the start of the forecast\. We therefore instantiate𝐱0\\mathbf\{x\}\_\{0\}as an extended forecasting segment that includes a short overlap with the observed history
𝐱0=\[𝐱label,𝐱fut\]\\mathbf\{x\}\_\{0\}=\[\\mathbf\{x\}\_\{\\text\{label\}\},\\,\\mathbf\{x\}\_\{\\text\{fut\}\}\]\(5\)where𝐱label\\mathbf\{x\}\_\{\\text\{label\}\}is a short label window of lengthLl≥2L\_\{l\}\\geq 2taken from the end of the observed history, and the extended target has total lengthS=Ll\+HS=L\_\{l\}\+H\. Appendix[B\.7](https://arxiv.org/html/2607.22599#A2.SS7)provides the formal analysis\. Although𝐀t\\mathbf\{A\}\_\{t\}reshapes the intermediate trajectory, the diffusion endpoint is unchanged\. Asα¯t→0\\bar\{\\alpha\}\_\{t\}\\to 0whent→Tt\\to T, the signal term vanishes regardless of𝐀t\\mathbf\{A\}\_\{t\}, and the terminal distribution approaches a standard Gaussian𝒩\(𝟎,𝐈\)\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\), as verified in Appendix[B\.6](https://arxiv.org/html/2607.22599#A2.SS6)\.
### 3\.2Conditional Reverse Process
Since the transition matrix𝐀t\\mathbf\{A\}\_\{t\}varies with the diffusion step, the standard DDPM posteriorq\(𝐱t−1∣𝐱t,𝐱0\)q\(\\mathbf\{x\}\_\{t\-1\}\\mid\\mathbf\{x\}\_\{t\},\\mathbf\{x\}\_\{0\}\)no longer directly applies to our non\-isotropic forward process\. We therefore follow the non\-Markovian formulation of DDIM\[[50](https://arxiv.org/html/2607.22599#bib.bib7)\]and parameterize the reverse dynamics through a denoiser that predicts the clean target in the value domain\. Specifically, given a noisy state𝐱t\\mathbf\{x\}\_\{t\}at stepttand the observed history𝐱hist\\mathbf\{x\}\_\{\\text\{hist\}\}, the denoiserfθ\(𝐱t,t,𝐱hist\)f\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t,\\mathbf\{x\}\_\{\\text\{hist\}\}\)predicts a clean target𝐱^0\\hat\{\\mathbf\{x\}\}\_\{0\}\. This value\-domain prediction is then used to recover the corresponding noise estimate
ϵ^=𝐱t−α¯t𝐀t𝐱^01−α¯t\\hat\{\\boldsymbol\{\\epsilon\}\}=\\frac\{\\mathbf\{x\}\_\{t\}\-\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathbf\{A\}\_\{t\}\\,\\hat\{\\mathbf\{x\}\}\_\{0\}\}\{\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\}\(6\)For a reverse step from𝐱t\\mathbf\{x\}\_\{t\}to an earlier state𝐱s\\mathbf\{x\}\_\{s\}withs<ts<t, we use the update
𝐱s=α¯s𝐀s𝐱^0\+1−α¯s−σt2ϵ^\+σtϵ′\\mathbf\{x\}\_\{s\}=\\sqrt\{\\bar\{\\alpha\}\_\{s\}\}\\,\\mathbf\{A\}\_\{s\}\\,\\hat\{\\mathbf\{x\}\}\_\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{s\}\-\\sigma\_\{t\}^\{2\}\}\\,\\hat\{\\boldsymbol\{\\epsilon\}\}\+\\sigma\_\{t\}\\,\\boldsymbol\{\\epsilon\}^\{\\prime\}\(7\)whereϵ′∼𝒩\(𝟎,𝐈\)\\boldsymbol\{\\epsilon\}^\{\\prime\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)andσt\\sigma\_\{t\}controls the stochasticity of the reverse transition under0≤σt2≤1−α¯s0\\leq\\sigma\_\{t\}^\{2\}\\leq 1\-\\bar\{\\alpha\}\_\{s\}, which we set toσt=0\\sigma\_\{t\}=0in all experiments\. Although this update involves step\-dependent matrices𝐀t\\mathbf\{A\}\_\{t\}and𝐀s\\mathbf\{A\}\_\{s\}, it remains consistent with the forward marginals\. In particular, if the denoiser is perfect so that𝐱^0=𝐱0\\hat\{\\mathbf\{x\}\}\_\{0\}=\\mathbf\{x\}\_\{0\}, then \([7](https://arxiv.org/html/2607.22599#S3.E7)\) yields
𝐱s∼q\(𝐱s∣𝐱0\)=𝒩\(α¯s𝐀s𝐱0,\(1−α¯s\)𝐈\)\\mathbf\{x\}\_\{s\}\\sim q\(\\mathbf\{x\}\_\{s\}\\mid\\mathbf\{x\}\_\{0\}\)=\\mathcal\{N\}\\\!\\left\(\\sqrt\{\\bar\{\\alpha\}\_\{s\}\}\\,\\mathbf\{A\}\_\{s\}\\,\\mathbf\{x\}\_\{0\},\\,\(1\-\\bar\{\\alpha\}\_\{s\}\)\\,\\mathbf\{I\}\\right\)\(8\)for any choice of\{𝐀t\}\\\{\\mathbf\{A\}\_\{t\}\\\}, because the𝐀t\\mathbf\{A\}\_\{t\}term cancels exactly in the noise residual𝐱t−α¯t𝐀t𝐱0\\mathbf\{x\}\_\{t\}\-\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathbf\{A\}\_\{t\}\\,\\mathbf\{x\}\_\{0\}\. The formal derivation is provided in Appendix[B\.2](https://arxiv.org/html/2607.22599#A2.SS2)\. The reverse process operates across two representation domains\. The denoiser always predicts𝐱^0\\hat\{\\mathbf\{x\}\}\_\{0\}in the value domain, while the reverse state𝐱s\\mathbf\{x\}\_\{s\}remains in the step\-dependent representation induced by𝐀s\\mathbf\{A\}\_\{s\}, so the term𝐀s𝐱^0\\mathbf\{A\}\_\{s\}\\hat\{\\mathbf\{x\}\}\_\{0\}maps the predicted clean target into the correct signal domain at stepss\. Asssdecreases,𝐀s\\mathbf\{A\}\_\{s\}gradually approaches𝐈\\mathbf\{I\}and the reverse trajectory returns to the original value representation\.
### 3\.3Stage\-Adaptive Conditional Denoiser
Since the forward process progressively shifts the signal preserved by the noisy intermediate state from the target itself toward its differenced structure, the reverse model should receive conditioning information that supplies both the value\-domain content from the observed history and cues about the history’s temporal evolution\. We therefore design a stage\-adaptive conditional denoiserfθ\(𝐱t,t,𝐱hist\)f\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t,\\mathbf\{x\}\_\{\\text\{hist\}\}\)that builds its conditioning from two complementary views of𝐱hist\\mathbf\{x\}\_\{\\text\{hist\}\}\. The value path encodes the history directly, and the difference path encodes its first\-order temporal differences,
𝐡val=MLPval\(𝐱hist\),𝐡diff=MLPdiff\(Δ𝐱hist\)\\mathbf\{h\}\_\{\\text\{val\}\}=\\mathrm\{MLP\}\_\{\\text\{val\}\}\(\\mathbf\{x\}\_\{\\text\{hist\}\}\),\\qquad\\mathbf\{h\}\_\{\\text\{diff\}\}=\\mathrm\{MLP\}\_\{\\text\{diff\}\}\(\\Delta\\mathbf\{x\}\_\{\\text\{hist\}\}\)\(9\)whereΔ𝐱hist\\Delta\\mathbf\{x\}\_\{\\text\{hist\}\}denotes the sequence of consecutive first\-order differences of𝐱hist\\mathbf\{x\}\_\{\\text\{hist\}\}\. The value representation𝐡val\\mathbf\{h\}\_\{\\text\{val\}\}preserves level, scale, and trend information, while the difference representation𝐡diff\\mathbf\{h\}\_\{\\text\{diff\}\}captures local velocity and short\-term fluctuations\. Although the forward process uses the second\-order operator𝐃2\\mathbf\{D\}\_\{2\}, first\-order differences suffice for the conditioning path because it only needs to convey the local velocity of the history, while the second\-order structure is already encoded in the noisy target through𝐀t\\mathbf\{A\}\_\{t\}\. Since the signal geometry changes with the diffusion step, a fixed combination of these two paths is not appropriate\. We therefore introduce a timestep\-adaptive gate
𝐠=sigmoid\(MLPgate\(\[𝐭emb,𝐡¯val\]\)\),𝐜=𝐡val\+𝐠⊙𝐡diff\\mathbf\{g\}=\\mathrm\{sigmoid\}\\\!\\left\(\\mathrm\{MLP\}\_\{\\text\{gate\}\}\\\!\\left\(\[\\mathbf\{t\}\_\{\\text\{emb\}\},\\,\\bar\{\\mathbf\{h\}\}\_\{\\text\{val\}\}\]\\right\)\\right\),\\qquad\\mathbf\{c\}=\\mathbf\{h\}\_\{\\text\{val\}\}\+\\mathbf\{g\}\\odot\\mathbf\{h\}\_\{\\text\{diff\}\}\(10\)where𝐭emb\\mathbf\{t\}\_\{\\text\{emb\}\}is a learned diffusion\-step embedding,𝐡¯val\\bar\{\\mathbf\{h\}\}\_\{\\text\{val\}\}is the mean of𝐡val\\mathbf\{h\}\_\{\\text\{val\}\},MLPgate\\mathrm\{MLP\}\_\{\\text\{gate\}\}is a two\-layer MLP with SiLU activation, and⊙\\odotdenotes element\-wise multiplication\. The gate raises the contribution of𝐡diff\\mathbf\{h\}\_\{\\text\{diff\}\}at diffusion stages where the noisy state retains more differenced structure and suppresses it when the value\-domain content alone suffices\. The conditioning feature𝐜\\mathbf\{c\}enters the denoiser through two pathways\. It serves as global context for the denoising backbonehθh\_\{\\theta\}, a stack of residual MLP layers, and is separately projected into the target sequence space to form a value\-level baseline\. The full denoiser output is
fθ\(𝐱t,t,𝐱hist\)=hθ\(𝐱t,t,𝐜\)\+𝐖p𝐜f\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t,\\mathbf\{x\}\_\{\\text\{hist\}\}\)=h\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t,\\mathbf\{c\}\)\+\\mathbf\{W\}\_\{p\}\\,\\mathbf\{c\}\(11\)where𝐖p\\mathbf\{W\}\_\{p\}is a learned linear projection\. The second term supplies a value\-domain baseline directly from the conditioning, whilehθh\_\{\\theta\}recovers the remaining structure from the noisy input\. Alternative conditioning variants are discussed in Appendix[E](https://arxiv.org/html/2607.22599#A5)\.
### 3\.4Training and Sampling
Training objective\.Before diffusion, the extended target𝐱0=\[𝐱label,𝐱fut\]\\mathbf\{x\}\_\{0\}=\[\\mathbf\{x\}\_\{\\text\{label\}\},\\,\\mathbf\{x\}\_\{\\text\{fut\}\}\]is instance\-normalized along the time dimension to obtain𝐱~0\\tilde\{\\mathbf\{x\}\}\_\{0\}\[[22](https://arxiv.org/html/2607.22599#bib.bib32)\]\. This normalization removes sample\-specific level and scale variation, so the diffusion’s modeling capacity is not consumed by absorbing per\-sample magnitude shifts\. The denoiser is trained to reconstruct this normalized target with a position\-weighted objective,
ℒrecon=𝔼t,𝐱~0,ϵ\[∑i=1Swi\(fθ,i\(𝐱t,t,𝐱hist\)−x~0,i\)2\],\\mathcal\{L\}\_\{\\mathrm\{recon\}\}=\\mathbb\{E\}\_\{t,\\,\\tilde\{\\mathbf\{x\}\}\_\{0\},\\,\\boldsymbol\{\\epsilon\}\}\\left\[\\sum\_\{i=1\}^\{S\}w\_\{i\}\\left\(f\_\{\\theta,i\}\(\\mathbf\{x\}\_\{t\},t,\\mathbf\{x\}\_\{\\text\{hist\}\}\)\-\\tilde\{x\}\_\{0,i\}\\right\)^\{\\\!2\}\\right\],\(12\)wheret∼Uniform\{0,…,T−1\}t\\sim\\mathrm\{Uniform\}\\\{0,\\ldots,T\-1\\\},𝐱t\\mathbf\{x\}\_\{t\}is sampled from the forward marginal defined on𝐱~0\\tilde\{\\mathbf\{x\}\}\_\{0\}, andwi=wfutw\_\{i\}=w\_\{\\mathrm\{fut\}\}is a fixed hyperparameter \(wfut≥1w\_\{\\mathrm\{fut\}\}\\geq 1\) for future positions andwi=1w\_\{i\}=1for label positions\. This asymmetric weighting keeps the optimization focused on the forecasting segment while ensuring the label window remains well\-reconstructed to preserve cross\-boundary continuity\. Appendix[B\.3](https://arxiv.org/html/2607.22599#A2.SS3)details the full denoising formulation underlying this objective\. To recover the absolute scale removed by normalization, an auxiliary linear head estimates the instance statistics\(μ^,ν^\)\(\\hat\{\\mu\},\\hat\{\\nu\}\)from the conditioning feature𝐜\\mathbf\{c\}, whereμ\\muandν\\nuare the per\-instance mean and standard deviation of𝐱0\\mathbf\{x\}\_\{0\}along the time dimension\. The total loss is
ℒ=ℒrecon\+\(μ^−μ\)2\+\(ν^−ν\)2\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{recon\}\}\+\(\\hat\{\\mu\}\-\\mu\)^\{2\}\+\(\\hat\{\\nu\}\-\\nu\)^\{2\}\(13\)where\(μ,ν\)\(\\mu,\\nu\)are the true instance normalization statistics\. Appendix[B\.4](https://arxiv.org/html/2607.22599#A2.SS4)establishes that this joint objective also controls the unnormalized prediction error\. At inference, the predicted statistics map the denoised output from the normalized target space back to the original value scale\.
Sampling\.At inference, forecasting starts from𝐱T∼𝒩\(𝟎,𝐈\)\\mathbf\{x\}\_\{T\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)and applies the reverse update in \([7](https://arxiv.org/html/2607.22599#S3.E7)\) withσt=0\\sigma\_\{t\}=0\. Since the reverse step is defined for arbitrary step pairs\(t,s\)\(t,s\)and the terminal distribution approaches a standard Gaussian, the standard DDIM sub\-sampling strategy applies directly\. We therefore use a subset of reverse steps during inference to reduce sampling cost without retraining\. After de\-normalization using the predicted statistics\(μ^,ν^\)\(\\hat\{\\mu\},\\hat\{\\nu\}\), the firstLlL\_\{l\}label positions are discarded and the remainingHHpositions serve as the forecast sample\.
## 4Experiment Settings
### 4\.1Datasets
We evaluate on seven widely used time series datasets spanning diverse temporal characteristics, namelyElectricity,ETTm2\[[69](https://arxiv.org/html/2607.22599#bib.bib33)\],Exchange Rate,Traffic,Solar\-Energy\[[25](https://arxiv.org/html/2607.22599#bib.bib34)\],Weather\[[63](https://arxiv.org/html/2607.22599#bib.bib35)\], andWind\[[60](https://arxiv.org/html/2607.22599#bib.bib8)\]\. Each dataset is split chronologically into training, validation, and test sets following standard forecasting protocol, with a ratio of 6:2:2 for ETTm2 and 7:1:2 for the remaining six datasets\. Following the prior studies\[[63](https://arxiv.org/html/2607.22599#bib.bib35),[38](https://arxiv.org/html/2607.22599#bib.bib43)\], we use a lookback lengthL=96L=96for bothDiffDiffand all baselines, and evaluate across four prediction horizonsH∈\{96,192,336,720\}H\\in\\\{96,192,336,720\\\}\. Probabilistic forecasting quality is measured by the continuous ranked probability score \(CRPS\)\[[15](https://arxiv.org/html/2607.22599#bib.bib37)\]where lower values indicate better calibrated and sharper forecasts\.
### 4\.2Baselines and Comparing Approaches
We compareDiffDiffagainst three controlled ablations that independently isolate its forward process and its denoiser\. The*Standard*variant pairs an isotropic Gaussian forward with a plain history\-encoded denoiser, removing both design choices\. The*Diff\-forward*variant keeps the differencing forward process but reverts the denoiser to a standard conditioning pathway that encodes only the value\-domain content of the history, so that any gain over Standard isolates the forward process\. The*Std\+G*variant pairs an isotropic Gaussian forward with the stage\-adaptive conditional denoiser, so that any gain over Standard isolates the denoiser\. For comparison with existing diffusion\-based forecasters, we include isotropic\-forward designs CSDI\[[57](https://arxiv.org/html/2607.22599#bib.bib2)\]and TimeDiff\[[48](https://arxiv.org/html/2607.22599#bib.bib13)\], methods that modify the diffusion process from within, namely TMDM\[[28](https://arxiv.org/html/2607.22599#bib.bib14)\], NsDiff\[[65](https://arxiv.org/html/2607.22599#bib.bib25)\], and MA\-TSD\[[60](https://arxiv.org/html/2607.22599#bib.bib8)\], and D3U\[[26](https://arxiv.org/html/2607.22599#bib.bib23)\], which pre\-cleans the residual through an external decomposition before applying diffusion\.
### 4\.3Implementation Details
ForDiffDiff,λt\\lambda\_\{t\}increases linearly from0toλmax=0\.5\\lambda\_\{\\max\}=0\.5along the diffusion chain, with label window lengthLl=48L\_\{l\}=48and future\-position loss weightwfut=5w\_\{\\text\{fut\}\}=5\. The diffusion usesT=50T=50steps under a cosineα¯t\\bar\{\\alpha\}\_\{t\}schedule, with inference via 10\-step DDIM sub\-sampling and 100 stochastic samples\. The denoiser is a stack of three residual MLP layers with model dimension 128, feedforward dimension 512, and dropout 0\.1\. Optimization uses Adam\[[23](https://arxiv.org/html/2607.22599#bib.bib36)\]at learning rate2×10−42\\times 10^\{\-4\}, weight decay10−510^\{\-5\}, and batch size 64 for up to 100 epochs, with all reported results averaged over five random seeds\.
## 5Results and Analysis
### 5\.1Overall Results
Table 1:Ablation results on seven benchmarks across four prediction horizons\. For every column the best value across the four methods isboldand the second\-best isunderlined\. Standard uses an isotropic forward with a plain denoiser\. The other three rows replace the forward only \(Diff\-forward\), the denoiser only \(Std\+G\), or both \(DiffDiff\)\. The rightmost column reports per\-row wins\.Table 2:Forecasting results on seven benchmarks across four prediction horizons\. For every column the best value across methods isboldand the second\-best isunderlined\. The rightmost column reports per\-row wins\.In Table[1](https://arxiv.org/html/2607.22599#S5.T1), Standard wins no setting while the two differencing\-forward variants account for every win, identifying the forward process as the primary source of improvement\. Std\+G improves over Standard on most settings yet wins no cell in the four\-way comparison, confirming that the denoiser carries genuine modeling power yet cannot substitute for the reshaped corruption path\. AtH=720H\{=\}720, where long\-range reconstruction is hardest,DiffDiffwins every setting while Diff\-forward wins none, leaving the two designs complementary rather than redundant\.
In Table[2](https://arxiv.org/html/2607.22599#S5.T2),DiffDiffwins 53 of 84 settings ahead of six diffusion baselines\. CSDI, originally an imputation model, deteriorates sharply at long horizons since conditioning on observed points fails to extrapolate over long gaps\. TimeDiff combines a history\-conditioned UNet with an isotropic forward, showing that denoiser\-side conditioning alone cannot compensate for a task\-agnostic corruption path\. NsDiff concentrates all 8 of its wins atH=720H\{=\}720, where its learned non\-stationary variance schedule meets prediction windows long enough to include regime changes\. TMDM is the strongest baseline with 16 wins, concentrated on Solar atH≤336H\\leq 336and Exchange atH=720H\{=\}720, where a persistent regime mean leaves the differencing operator little additional structure to exploit\. D3U obtains no first\-place result, with degradation sharpest on regime\-shifting series at long horizons where its externally trained point predictor’s errors enter the residual the diffusion stage models and the deterministic\-stochastic boundary cannot adjust once frozen\. Internalizing the asymmetry along the trajectory removes this dependence\. MA\-TSD’s wins concentrate on Exchange and Wind, whereasDiffDiff’s largest margins appear on Electricity, ETTm2, Traffic, and Weather, where forecast windows place substantial energy in components the history under\-determines\.
Figure 2:Reshaping of the corruption trajectory by the differencing forward process\.Top: Per\-frequency power spectral density \(PSD\) ratio ofDiffDiffto the standard isotropic forward at three diffusion progresses onelectricity, with values above \(below\) unity marking enhanced \(attenuated\) bands\.Bottom: Retained energy fraction of the window mean and the first\- and second\-order increments under each forward path along the diffusion progress onsolar\.Figure 3:Component\-level corruption alignment and gate dynamics\.Left:Per\-component coefficient of determinationR2R^\{2\}\(horizontal\) versus corruption aggressiveness \(vertical\) onwindandexchange\-rate, with each component reducing a window to a scalar, namely the window mean \(level\), summed fast Fourier transform \(FFT\) magnitudes in three frequency bands \(low, mid, high\), and the mean first\-order absolute difference \(diff1\)\. Components are traced across five diffusion progresses, with open markers for the standard forward and filled markers forDiffDiff, jittered horizontally for visibility\.Right:Gate activation of the differential branch versus diffusion timestep across seven datasets, with shaded bands showing±1\\pm 1standard deviation\.
### 5\.2Spectral Reshaping by the Differencing Forward Process
As shown in Figure[2](https://arxiv.org/html/2607.22599#S5.F2), we strip the additive noise from both forwards and compare what each path retains at three diffusion progresses, in the frequency domain \(top\) and across three operator\-aligned components \(bottom\)\. The reshape unfolds progressively along the chain rather than as a single fixed transform, with the PSD ratio sitting near unity att/T≈0\.25t/T\{\\approx\}0\.25and carving a sharp split between attenuated low\-frequency and amplified high\-frequency content byt/T≈0\.75t/T\{\\approx\}0\.75\. In the operator domain,DiffDiffattenuates the level energy faster than the standard forward, preserves the first\-order increment longer, and preserves the second\-order increment most of all, identifying𝐃2\\mathbf\{D\}\_\{2\}as a structural filter aligned with the order of differencing rather than tuned to arbitrary frequency bands\. With isotropic noise and selectively attenuated signal, the SNR redistributes toward high\-frequency content and away from low\-frequency content, turning the corruption trajectory from a uniform attenuator into a frequency\-selective preserver\.
### 5\.3Predictability\-Driven Corruption Reallocation
To check whether the corruption budget tracks history\-availability, we trace each signal component along the diffusion chain in an\(R2,aggressiveness\)\(R^\{2\},\\text\{aggressiveness\}\)plane, withR2R^\{2\}measuring the linear\-regression fit from the matched component in the history window to that in the future window and aggressiveness equal to1−α¯t‖ϕ\(𝐀t𝐱0\)‖2/‖ϕ\(𝐱0\)‖21\-\\bar\{\\alpha\}\_\{t\}\\\|\\phi\(\\mathbf\{A\}\_\{t\}\\mathbf\{x\}\_\{0\}\)\\\|^\{2\}/\\\|\\phi\(\\mathbf\{x\}\_\{0\}\)\\\|^\{2\}at diffusion progresst/Tt/T, capturing both operator\-induced and noise\-schedule\-induced energy decay\. Wind concentrates all five components into a narrowR2R^\{2\}band while exchange rate spreads them broadly, and on both the predictability ranking lines up with the spectral one, making selectivity along frequency equivalent to selectivity along predictability\. Under the standard forward the five components climb in lockstep regardless ofR2R^\{2\}, whereasDiffDiffbreaks this into a fan\-shaped spread, with high\-R2R^\{2\}components rising above the standard reference and low\-R2R^\{2\}ones falling below\. Att/T=0\.5t/T\{=\}0\.5the component\-wise correlation betweenR2R^\{2\}and aggressiveness rises from zero under the standard forward to\+0\.76\+0\.76on wind and\+0\.58\+0\.58on exchange rate underDiffDiff, with more history\-predictable components attenuated more heavily and less predictable ones preserved\. The forward process therefore spends its attenuation budget on the components the history can already supply, freeing the diffusion’s representational capacity for the harder, history\-under\-determined ones\.
### 5\.4Stage\-Adaptive Gating of the Differential Branch
To examine the gate’s adaptation to timestep and dataset, we sweep the diffusion timestep on the trained model and record the mean gate value over batch samples and hidden dimensions at each step\. As shown in the right panel of Figure[3](https://arxiv.org/html/2607.22599#S5.F3), the gate is far from constant, and the seven datasets separate into three behavioral regimes\. Electricity and Traffic, both with strong periodic structure, hold the gate above0\.50\.5from the start and lift it further withtt, keeping the differential branch open throughout the chain\. ETTm2, Exchange Rate, and Weather hold the gate near its initialization in early steps and admit it sharply in the late low\-SNR portion, drawing on differential evidence only as the value signal degrades\. Solar and Wind suppress the gate throughout the chain, indicating that for these highly stochastic series the differential branch extracts mostly noise rather than usable structure\. The gate’s stage\- and dataset\-adaptive admission complements the forward operator’s predictability\-aligned suppression by re\-supplying history\-anchored content when the noisy target no longer carries it\.
### 5\.5Case Study
As shown in Figure[4](https://arxiv.org/html/2607.22599#S5.F4), we visualize the95%95\\%prediction interval ofDiffDiffalongside the three closest diffusion baselines on an illustrative electricity window \(top, periodic\) and an exchange\-rate window \(bottom, regime\-shifting\), each selected as a top\-CRPS\-improvement window ofDiffDiffover the strongest baseline on that dataset\. TMDM’s bands hedge widely over the latter half of each horizon, with the median tracking the cycle but the spread leaving considerable residual ambiguity\. NsDiff produces visibly jagged predictions whose median and band fluctuate between adjacent time steps, a symptom of its per\-cell variance design\. MA\-TSD narrows the bands most aggressively on the exchange\-rate window, yet the median misses the trend reversal there and overshoots the periodic peaks and troughs on electricity\.DiffDiff’s forecast simultaneously achieves a smooth median, a narrow band, and faithful tracking of both the periodic cycle and the trend reversal\.
Figure 4:Per\-window95%95\\%prediction interval ofDiffDiffand three diffusion baselines on illustrative windows from electricity \(top\) and exchange rate \(bottom\) atH=96H\{=\}96, each selected as a top\-CRPS\-improvement window ofDiffDiffover the strongest baseline on that dataset\. The gray line shows the history, the black line the ground truth, the colored line the median forecast, and the shaded band the95%95\\%prediction interval\.
## 6Conclusion
We show that probabilistic time series forecasting benefits from encoding predictability asymmetry directly into the diffusion trajectory\.DiffDiffprogressively transforms the forward state from the target itself toward its second\-order differenced structure, while a stage\-adaptive conditioning pathway combines value\-domain and differential history cues\. This design reallocates model capacity from reconstructing history\-anchored content to recovering history\-under\-determined dynamics\. Across seven benchmarks,DiffDiffoutperforms six diffusion baselines, and component\-wise analysis confirms that the operator and gate adapt to history\-predictability along the chain\.
## References
- \[1\]J\. Alayrac, J\. Donahue, P\. Luc, A\. Miech, I\. Barr, Y\. Hasson, K\. Lenc, A\. Mensch, K\. Millican, M\. Reynolds,et al\.\(2022\)Flamingo: a visual language model for few\-shot learning\.Advances in neural information processing systems35,pp\. 23716–23736\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[2\]\(2022\)Diffusion\-based time series imputation and forecasting with structured state space models\.arXiv preprint arXiv:2208\.09399\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[3\]J\. Alido, T\. Li, Y\. Sun, and L\. Tian\(2025\)Whitened score diffusion: a structured prior for imaging inverse problems\.arXiv preprint arXiv:2505\.10311\.Cited by:[§2\.2](https://arxiv.org/html/2607.22599#S2.SS2.p1.1)\.
- \[4\]A\. Bansal, E\. Borgnia, H\. Chu, J\. Li, H\. Kazemi, F\. Huang, M\. Goldblum, J\. Geiping, and T\. Goldstein\(2023\)Cold diffusion: inverting arbitrary image transforms without noise\.Vol\.36,pp\. 41259–41282\.Cited by:[§2\.2](https://arxiv.org/html/2607.22599#S2.SS2.p1.1)\.
- \[5\]A\. Blattmann, R\. Rombach, K\. Oktay, J\. Müller, and B\. Ommer\(2022\)Retrieval\-augmented diffusion models\.Vol\.35,pp\. 15309–15324\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[6\]J\. Cao, Z\. Li, and J\. Li\(2019\)Financial time series forecasting model based on ceemdan and lstm\.Physica A: Statistical mechanics and its applications519,pp\. 127–139\.Cited by:[§1](https://arxiv.org/html/2607.22599#S1.p1.1)\.
- \[7\]P\. Chen, Y\. Zhang, Y\. Cheng, Y\. Shu, Y\. Wang, Q\. Wen, B\. Yang, and C\. Guo\(2024\)Pathformer: multi\-scale transformers with adaptive pathways for time series forecasting\.arXiv preprint arXiv:2402\.05956\.Cited by:[Appendix I](https://arxiv.org/html/2607.22599#A9.p1.1)\.
- \[8\]J\. Chou and D\. Tran\(2018\)Forecasting energy consumption time series using machine learning techniques based on usage patterns of residential householders\.Energy165,pp\. 709–726\.Cited by:[§1](https://arxiv.org/html/2607.22599#S1.p1.1)\.
- \[9\]G\. Daras, M\. Delbracio, H\. Talebi, A\. G\. Dimakis, and P\. Milanfar\(2022\)Soft diffusion: score matching for general corruptions\.arXiv preprint arXiv:2209\.05442\.Cited by:[§2\.2](https://arxiv.org/html/2607.22599#S2.SS2.p1.1)\.
- \[10\]J\. Demšar\(2006\)Statistical comparisons of classifiers over multiple data sets\.Journal of Machine learning research7\(Jan\),pp\. 1–30\.Cited by:[Appendix H](https://arxiv.org/html/2607.22599#A8.p1.1)\.
- \[11\]A\. Dingli and K\. S\. Fournier\(2017\)Financial time series forecasting\-a machine learning approach\.Machine Learning and Applications: An International Journal4\(1/2\),pp\. 3\.Cited by:[§1](https://arxiv.org/html/2607.22599#S1.p1.1)\.
- \[12\]X\. Fan, Y\. Wu, C\. Xu, Y\. Huang, W\. Liu, and J\. Bian\(2024\)Mg\-tsd: multi\-granularity time series diffusion models with guided learning process\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[13\]R\. Gan, Y\. Tian, K\. Pan, Y\. Song, and Y\. Zhang\(2026\)Reinforced context augmentation for multimodal emotion analysis\.IEEE Transactions on Multimedia\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[14\]R\. Gan, X\. Wu, J\. Lu, Y\. Tian, D\. Zhang, Z\. Wu, R\. Sun, C\. Liu, J\. Zhang, P\. Zhang,et al\.\(2023\)IDesigner: a high\-resolution and complex\-prompt following text\-to\-image diffusion model for interior design\.arXiv preprint arXiv:2312\.04326\.Cited by:[§1](https://arxiv.org/html/2607.22599#S1.p1.1)\.
- \[15\]T\. Gneiting and A\. E\. Raftery\(2007\)Strictly proper scoring rules, prediction, and estimation\.Journal of the American statistical Association102\(477\),pp\. 359–378\.Cited by:[§4\.1](https://arxiv.org/html/2607.22599#S4.SS1.p1.2)\.
- \[16\]P\. Hewage, A\. Behera, M\. Trovati, E\. Pereira, M\. Ghahremani, F\. Palmieri, and Y\. Liu\(2020\)Temporal convolutional neural \(tcn\) network for an effective weather forecasting using time\-series data from the local weather station: p\. hewage et al\.\.Soft Computing24\(21\),pp\. 16453–16482\.Cited by:[§1](https://arxiv.org/html/2607.22599#S1.p1.1)\.
- \[17\]J\. Ho, A\. Jain, and P\. Abbeel\(2020\)Denoising diffusion probabilistic models\.Vol\.33,pp\. 6840–6851\.Cited by:[§A\.1](https://arxiv.org/html/2607.22599#A1.SS1.p1.1),[§A\.2](https://arxiv.org/html/2607.22599#A1.SS2.p1.3),[§A\.3](https://arxiv.org/html/2607.22599#A1.SS3.SSS0.Px1.p1.5),[Appendix A](https://arxiv.org/html/2607.22599#A1.p1.1),[§B\.3](https://arxiv.org/html/2607.22599#A2.SS3.SSS0.Px2.p2.6),[§1](https://arxiv.org/html/2607.22599#S1.p1.1)\.
- \[18\]E\. Hoogeboom and T\. Salimans\(2022\)Blurring diffusion models\.Cited by:[§2\.2](https://arxiv.org/html/2607.22599#S2.SS2.p1.1)\.
- \[19\]Y\. Jin, W\. Chen, Y\. Tian, Y\. Song, and C\. Yan\(2024\)Improving radiology report generation with multi\-grained abnormality prediction\.Neurocomputing600,pp\. 128122\.Cited by:[§2\.2](https://arxiv.org/html/2607.22599#S2.SS2.p1.1)\.
- \[20\]B\. Jing, S\. Zhang, Y\. Zhu, B\. Peng, K\. Guan, A\. Margenot, and H\. Tong\(2022\)Retrieval based time series forecasting\.arXiv preprint arXiv:2209\.13525\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[21\]Z\. Karevan and J\. A\. Suykens\(2020\)Transductive lstm for time\-series prediction: an application to weather forecasting\.Neural Networks125,pp\. 1–9\.Cited by:[§1](https://arxiv.org/html/2607.22599#S1.p1.1)\.
- \[22\]T\. Kim, J\. Kim, Y\. Tae, C\. Park, J\. Choi, and J\. Choo\(2021\)Reversible instance normalization for accurate time\-series forecasting against distribution shift\.InInternational conference on learning representations,Cited by:[§B\.3](https://arxiv.org/html/2607.22599#A2.SS3.SSS0.Px3.p1.1),[§3\.4](https://arxiv.org/html/2607.22599#S3.SS4.p1.2)\.
- \[23\]D\. P\. Kingma and J\. Ba\(2014\)Adam: a method for stochastic optimization\.Cited by:[§4\.3](https://arxiv.org/html/2607.22599#S4.SS3.p1.9)\.
- \[24\]M\. Kollovieh, A\. F\. Ansari, M\. Bohlke\-Schneider, J\. Zschiegner, H\. Wang, and Y\. B\. Wang\(2023\)Predict, refine, synthesize: self\-guiding diffusion models for probabilistic time series forecasting\.Vol\.36,pp\. 28341–28364\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[25\]G\. Lai, W\. Chang, Y\. Yang, and H\. Liu\(2018\)Modeling long\-and short\-term temporal patterns with deep neural networks\.InThe 41st international ACM SIGIR conference on research & development in information retrieval,pp\. 95–104\.Cited by:[§4\.1](https://arxiv.org/html/2607.22599#S4.SS1.p1.2)\.
- \[26\]Q\. Li, Z\. Zhang, L\. Yao, Z\. Li, T\. Zhong, and Y\. Zhang\(2025\)Diffusion\-based decoupled deterministic and uncertain framework for probabilistic multivariate time series forecasting\.InThe Thirteenth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2607.22599#S1.p2.1),[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1),[§4\.2](https://arxiv.org/html/2607.22599#S4.SS2.p1.1)\.
- \[27\]Y\. Li, X\. Lu, Y\. Wang, and D\. Dou\(2022\)Generative time series forecasting with diffusion, denoise, and disentanglement\.Vol\.35,pp\. 23009–23022\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[28\]Y\. Li, W\. Chen, X\. Hu, B\. Chen, B\. Sun, and M\. Zhou\(2024\)Transformer\-modulated diffusion models for probabilistic multivariate time series forecasting\.InThe Twelfth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2607.22599#S1.p2.1),[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1),[§4\.2](https://arxiv.org/html/2607.22599#S4.SS2.p1.1)\.
- \[29\]S\. Lin, H\. Chen, H\. Wu, C\. Qiu, and W\. Lin\(2025\)Temporal query network for efficient multivariate time series forecasting\.arXiv preprint arXiv:2505\.12917\.Cited by:[Appendix I](https://arxiv.org/html/2607.22599#A9.p1.1)\.
- \[30\]J\. Liu, L\. Yang, H\. Li, and S\. Hong\(2024\)Retrieval\-augmented diffusion models for time series forecasting\.Vol\.37,pp\. 2766–2786\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[31\]S\. Liu, H\. Yu, C\. Liao, J\. Li, W\. Lin, A\. X\. Liu, and S\. Dustdar\(2021\)Pyraformer: low\-complexity pyramidal attention for long\-range time series modeling and forecasting\.InInternational conference on learning representations,Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[32\]Y\. Liu, T\. Hu, H\. Zhang, H\. Wu, S\. Wang, L\. Ma, and M\. Long\(2023\)Itransformer: inverted transformers are effective for time series forecasting\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[33\]Y\. Liu, H\. Wu, J\. Wang, and M\. Long\(2022\)Non\-stationary transformers: exploring the stationarity in time series forecasting\.Vol\.35,pp\. 9881–9893\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[34\]Y\. Liu, S\. Wijewickrema, D\. Hu, C\. Bester, S\. O’Leary, and J\. Bailey\(2025\)Stochastic diffusion: a diffusion based model for stochastic time series forecasting\.InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\. 2,pp\. 1939–1950\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[35\]C\. Lu, Y\. Zhou, F\. Bao, J\. Chen, C\. Li, and J\. Zhu\(2022\)Dpm\-solver: a fast ode solver for diffusion probabilistic model sampling in around 10 steps\.Vol\.35,pp\. 5775–5787\.Cited by:[§1](https://arxiv.org/html/2607.22599#S1.p1.1)\.
- \[36\]R\. Nematirad, A\. Pahwa, and B\. Natarajan\(2025\)Times2d: multi\-period decomposition and derivative mapping for general time series forecasting\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.39,pp\. 19651–19658\.Cited by:[Appendix I](https://arxiv.org/html/2607.22599#A9.p1.1)\.
- \[37\]A\. Q\. Nichol and P\. Dhariwal\(2021\)Improved denoising diffusion probabilistic models\.InInternational conference on machine learning,pp\. 8162–8171\.Cited by:[§3\.1](https://arxiv.org/html/2607.22599#S3.SS1.p1.6)\.
- \[38\]Y\. Nie, N\. H\. Nguyen, P\. Sinthong, and J\. Kalagnanam\(2022\)A time series is worth 64 words: long\-term forecasting with transformers\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1),[§4\.1](https://arxiv.org/html/2607.22599#S4.SS1.p1.2)\.
- \[39\]J\. Nowotarski and R\. Weron\(2018\)Recent advances in electricity price forecasting: a review of probabilistic forecasting\.Renewable and Sustainable Energy Reviews81,pp\. 1548–1568\.Cited by:[§1](https://arxiv.org/html/2607.22599#S1.p1.1)\.
- \[40\]B\. N\. Oreshkin, D\. Carpov, N\. Chapados, and Y\. Bengio\(2019\)N\-beats: neural basis expansion analysis for interpretable time series forecasting\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[41\]W\. Peebles and S\. Xie\(2023\)Scalable diffusion models with transformers\.InProceedings of the IEEE/CVF international conference on computer vision,pp\. 4195–4205\.Cited by:[§1](https://arxiv.org/html/2607.22599#S1.p1.1)\.
- \[42\]A\. Radford, J\. W\. Kim, C\. Hallacy, A\. Ramesh, G\. Goh, S\. Agarwal, G\. Sastry, A\. Askell, P\. Mishkin, J\. Clark,et al\.\(2021\)Learning transferable visual models from natural language supervision\.InInternational conference on machine learning,pp\. 8748–8763\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[43\]K\. Rasul, C\. Seward, I\. Schuster, and R\. Vollgraf\(2021\)Autoregressive denoising diffusion models for multivariate probabilistic time series forecasting\.InInternational conference on machine learning,pp\. 8857–8868\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[44\]J\. Rishi, G\. Mothish, and D\. Subramani\(2025\)Conditional diffusion model with nonlinear data transformation for time series forecasting\.InForty\-second International Conference on Machine Learning,Cited by:[§1](https://arxiv.org/html/2607.22599#S1.p2.1),[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[45\]S\. Rissanen, M\. Heinonen, and A\. Solin\(2022\)Generative modelling with inverse heat dissipation\.Cited by:[§2\.2](https://arxiv.org/html/2607.22599#S2.SS2.p1.1)\.
- \[46\]R\. Rombach, A\. Blattmann, D\. Lorenz, P\. Esser, and B\. Ommer\(2022\)High\-resolution image synthesis with latent diffusion models\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 10684–10695\.Cited by:[§1](https://arxiv.org/html/2607.22599#S1.p1.1)\.
- \[47\]L\. Shen, W\. Chen, and J\. Kwok\(2024\)Multi\-resolution diffusion models for time series forecasting\.InThe Twelfth International Conference on Learning Representations,Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[48\]L\. Shen and J\. Kwok\(2023\)Non\-autoregressive conditional diffusion models for time series prediction\.InInternational Conference on Machine Learning,pp\. 31016–31029\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1),[§4\.2](https://arxiv.org/html/2607.22599#S4.SS2.p1.1)\.
- \[49\]J\. Sohl\-Dickstein, E\. Weiss, N\. Maheswaranathan, and S\. Ganguli\(2015\)Deep unsupervised learning using nonequilibrium thermodynamics\.InInternational conference on machine learning,pp\. 2256–2265\.Cited by:[§A\.1](https://arxiv.org/html/2607.22599#A1.SS1.p1.1),[§1](https://arxiv.org/html/2607.22599#S1.p1.1)\.
- \[50\]J\. Song, C\. Meng, and S\. Ermon\(2020\)Denoising diffusion implicit models\.Cited by:[§A\.4](https://arxiv.org/html/2607.22599#A1.SS4.p1.4),[Appendix A](https://arxiv.org/html/2607.22599#A1.p1.1),[§3\.1](https://arxiv.org/html/2607.22599#S3.SS1.p1.13),[§3\.2](https://arxiv.org/html/2607.22599#S3.SS2.p1.7)\.
- \[51\]Y\. Song, J\. Sohl\-Dickstein, D\. P\. Kingma, A\. Kumar, S\. Ermon, and B\. Poole\(2020\)Score\-based generative modeling through stochastic differential equations\.Cited by:[§1](https://arxiv.org/html/2607.22599#S1.p1.1)\.
- \[52\]C\. Su, Z\. Cai, Y\. Tian, Z\. Chang, Z\. Zheng, and Y\. Song\(2025\)Diffusion models for time series forecasting: a survey\.arXiv preprint arXiv:2507\.14507\.Cited by:[§1](https://arxiv.org/html/2607.22599#S1.p2.1)\.
- \[53\]C\. Su, Y\. Tian, Q\. Liu, J\. Zhang, and Y\. Song\(2025\)Fusing large language models with temporal transformers for time series forecasting\.arXiv preprint arXiv:2507\.10098\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[54\]C\. Su, Y\. Tian, Y\. Song, and Y\. Zhang\(2025\)Text reinforcement for multimodal time series forecasting\.arXiv preprint arXiv:2509\.00687\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[55\]C\. Su, Y\. Tian, and Y\. Song\(2025\)Multimodal conditioned diffusive time series forecasting\.arXiv preprint arXiv:2504\.19669\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[56\]C\. Su, Y\. Tian, and Y\. Song\(2026\)Learning shared sentiment prototypes for adaptive multimodal sentiment analysis\.arXiv preprint arXiv:2604\.05873\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[57\]Y\. Tashiro, J\. Song, Y\. Song, and S\. Ermon\(2021\)Csdi: conditional score\-based diffusion models for probabilistic time series imputation\.Vol\.34,pp\. 24804–24816\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1),[§4\.2](https://arxiv.org/html/2607.22599#S4.SS2.p1.1)\.
- \[58\]Y\. Tian, F\. Xia, and Y\. Song\(2024\)Diffusion networks with task\-specific noise control for radiology report generation\.InProceedings of the 32nd ACM International Conference on Multimedia,pp\. 1771–1780\.Cited by:[§1](https://arxiv.org/html/2607.22599#S1.p2.1)\.
- \[59\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, and I\. Polosukhin\(2017\)Attention is all you need\.Vol\.30\.Cited by:[1st item](https://arxiv.org/html/2607.22599#A5.I1.i1.p1.3)\.
- \[60\]C\. Wang, L\. Yang, Z\. Wang, L\. Sun, and Y\. Wang\(2025\)A non\-isotropic time series diffusion model with moving average transitions\.InForty\-second International Conference on Machine Learning,Cited by:[§1](https://arxiv.org/html/2607.22599#S1.p2.1),[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1),[§4\.1](https://arxiv.org/html/2607.22599#S4.SS1.p1.2),[§4\.2](https://arxiv.org/html/2607.22599#S4.SS2.p1.1)\.
- \[61\]X\. Wang, R\. Dai, K\. Liu, and X\. Chu\(2025\)Effective probabilistic time series forecasting with fourier adaptive noise\-separated diffusion\.arXiv preprint arXiv:2505\.11306\.Cited by:[§1](https://arxiv.org/html/2607.22599#S1.p2.1)\.
- \[62\]H\. Wu, T\. Hu, Y\. Liu, H\. Zhou, J\. Wang, and M\. Long\(2022\)Timesnet: temporal 2d\-variation modeling for general time series analysis\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[63\]H\. Wu, J\. Xu, J\. Wang, and M\. Long\(2021\)Autoformer: decomposition transformers with auto\-correlation for long\-term series forecasting\.Vol\.34,pp\. 22419–22430\.Cited by:[§4\.1](https://arxiv.org/html/2607.22599#S4.SS1.p1.2)\.
- \[64\]X\. Wu, D\. Zhang, R\. Gan, J\. Lu, Z\. Wu, R\. Sun, J\. Zhang, P\. Zhang, and Y\. Song\(2024\)Taiyi\-diffusion\-xl: advancing bilingual text\-to\-image generation with large vision\-language model support\.arXiv preprint arXiv:2401\.14688\.Cited by:[§1](https://arxiv.org/html/2607.22599#S1.p1.1)\.
- \[65\]W\. Ye, Z\. Xu, and N\. Gui\(2025\)Non\-stationary diffusion for probabilistic time series forecasting\.Cited by:[§1](https://arxiv.org/html/2607.22599#S1.p2.1),[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1),[§4\.2](https://arxiv.org/html/2607.22599#S4.SS2.p1.1)\.
- \[66\]X\. Yuan and Y\. Qiao\(2024\)Diffusion\-ts: interpretable diffusion for general time series generation\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[67\]J\. Zhang, M\. Cheng, X\. Tao, Z\. Liu, and D\. Wang\(2024\)Conditional denoising meets polynomial modeling: a flexible decoupled framework for time series forecasting\.arXiv preprint arXiv:2410\.13253\.Cited by:[§1](https://arxiv.org/html/2607.22599#S1.p2.1),[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[68\]Y\. Zhang and J\. Yan\(2023\)Crossformer: transformer utilizing cross\-dimension dependency for multivariate time series forecasting\.InThe eleventh international conference on learning representations,Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
- \[69\]H\. Zhou, S\. Zhang, J\. Peng, S\. Zhang, J\. Li, H\. Xiong, and W\. Zhang\(2021\)Informer: beyond efficient transformer for long sequence time\-series forecasting\.InProceedings of the AAAI conference on artificial intelligence,Vol\.35,pp\. 11106–11115\.Cited by:[§4\.1](https://arxiv.org/html/2607.22599#S4.SS1.p1.2)\.
- \[70\]T\. Zhou, Z\. Ma, Q\. Wen, X\. Wang, L\. Sun, and R\. Jin\(2022\)Fedformer: frequency enhanced decomposed transformer for long\-term series forecasting\.InInternational conference on machine learning,pp\. 27268–27286\.Cited by:[§2\.1](https://arxiv.org/html/2607.22599#S2.SS1.p1.1)\.
## Appendix ADiffusion Model Background
This appendix reviews the standard denoising diffusion framework that underpins theDiffDiffdesign\. We follow the formulations in prior studies\[[17](https://arxiv.org/html/2607.22599#bib.bib9),[50](https://arxiv.org/html/2607.22599#bib.bib7)\], and conclude by showing how the standard framework specializes to a particular case of the generalized forward process introduced in Section[3\.1](https://arxiv.org/html/2607.22599#S3.SS1)\.
### A\.1Forward Process and Reparameterization
A denoising diffusion probabilistic model\[[49](https://arxiv.org/html/2607.22599#bib.bib10),[17](https://arxiv.org/html/2607.22599#bib.bib9)\]defines a forward Markov chain that progressively adds Gaussian noise to a data sample𝐱0∼q\(𝐱0\)\\mathbf\{x\}\_\{0\}\\sim q\(\\mathbf\{x\}\_\{0\}\)\. The one\-step transition is
q\(𝐱t∣𝐱t−1\)=𝒩\(𝐱t;αt𝐱t−1,βt𝐈\),q\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{t\-1\}\)=\\mathcal\{N\}\\\!\\left\(\\mathbf\{x\}\_\{t\};\\;\\sqrt\{\\alpha\_\{t\}\}\\,\\mathbf\{x\}\_\{t\-1\},\\;\\beta\_\{t\}\\,\\mathbf\{I\}\\right\),\(14\)whereβt∈\(0,1\)\\beta\_\{t\}\\in\(0,1\)is the noise variance at stepttandαt=1−βt\\alpha\_\{t\}=1\-\\beta\_\{t\}\. By the reparameterization trick,
𝐱t=αt𝐱t−1\+βtϵt−1,ϵt−1∼𝒩\(𝟎,𝐈\)\.\\mathbf\{x\}\_\{t\}=\\sqrt\{\\alpha\_\{t\}\}\\,\\mathbf\{x\}\_\{t\-1\}\+\\sqrt\{\\beta\_\{t\}\}\\,\\boldsymbol\{\\epsilon\}\_\{t\-1\},\\qquad\\boldsymbol\{\\epsilon\}\_\{t\-1\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)\.\(15\)Substituting𝐱t−1=αt−1𝐱t−2\+βt−1ϵt−2\\mathbf\{x\}\_\{t\-1\}=\\sqrt\{\\alpha\_\{t\-1\}\}\\,\\mathbf\{x\}\_\{t\-2\}\+\\sqrt\{\\beta\_\{t\-1\}\}\\,\\boldsymbol\{\\epsilon\}\_\{t\-2\}and using the fact that the sum of two independent Gaussian noise terms𝒩\(𝟎,σ12𝐈\)\+𝒩\(𝟎,σ22𝐈\)=𝒩\(𝟎,\(σ12\+σ22\)𝐈\)\\mathcal\{N\}\(\\mathbf\{0\},\\sigma\_\{1\}^\{2\}\\mathbf\{I\}\)\+\\mathcal\{N\}\(\\mathbf\{0\},\\sigma\_\{2\}^\{2\}\\mathbf\{I\}\)=\\mathcal\{N\}\(\\mathbf\{0\},\(\\sigma\_\{1\}^\{2\}\+\\sigma\_\{2\}^\{2\}\)\\mathbf\{I\}\), we obtain
𝐱t=αtαt−1𝐱t−2\+1−αtαt−1ϵ,ϵ∼𝒩\(𝟎,𝐈\),\\mathbf\{x\}\_\{t\}=\\sqrt\{\\alpha\_\{t\}\\alpha\_\{t\-1\}\}\\,\\mathbf\{x\}\_\{t\-2\}\+\\sqrt\{1\-\\alpha\_\{t\}\\alpha\_\{t\-1\}\}\\,\\boldsymbol\{\\epsilon\},\\qquad\\boldsymbol\{\\epsilon\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\),\(16\)sinceαtβt−1\+βt=αt\(1−αt−1\)\+\(1−αt\)=1−αtαt−1\\alpha\_\{t\}\\beta\_\{t\-1\}\+\\beta\_\{t\}=\\alpha\_\{t\}\(1\-\\alpha\_\{t\-1\}\)\+\(1\-\\alpha\_\{t\}\)=1\-\\alpha\_\{t\}\\alpha\_\{t\-1\}\. Applying this recursion from stepttdown to step0and definingα¯t=∏s=1tαs\\bar\{\\alpha\}\_\{t\}=\\prod\_\{s=1\}^\{t\}\\alpha\_\{s\}gives the closed\-form forward marginal
q\(𝐱t∣𝐱0\)=𝒩\(𝐱t;α¯t𝐱0,\(1−α¯t\)𝐈\),q\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{0\}\)=\\mathcal\{N\}\\\!\\left\(\\mathbf\{x\}\_\{t\};\\;\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathbf\{x\}\_\{0\},\\;\(1\-\\bar\{\\alpha\}\_\{t\}\)\\,\\mathbf\{I\}\\right\),\(17\)equivalently𝐱t=α¯t𝐱0\+1−α¯tϵ\\mathbf\{x\}\_\{t\}=\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathbf\{x\}\_\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\,\\boldsymbol\{\\epsilon\}withϵ∼𝒩\(𝟎,𝐈\)\\boldsymbol\{\\epsilon\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)\. This reparameterization allows sampling any noisy state𝐱t\\mathbf\{x\}\_\{t\}directly from𝐱0\\mathbf\{x\}\_\{0\}without iterating through intermediate steps, and it forms the foundation for both training and the non\-Markovian extensions discussed below\.
### A\.2Evidence Lower Bound
The reverse process is parameterized as
pθ\(𝐱0:T\)=p\(𝐱T\)∏t=1Tpθ\(𝐱t−1∣𝐱t\),pθ\(𝐱t−1∣𝐱t\)=𝒩\(𝐱t−1;𝝁θ\(𝐱t,t\),β~t𝐈\),p\_\{\\theta\}\(\\mathbf\{x\}\_\{0:T\}\)=p\(\\mathbf\{x\}\_\{T\}\)\\prod\_\{t=1\}^\{T\}p\_\{\\theta\}\(\\mathbf\{x\}\_\{t\-1\}\\mid\\mathbf\{x\}\_\{t\}\),\\qquad p\_\{\\theta\}\(\\mathbf\{x\}\_\{t\-1\}\\mid\\mathbf\{x\}\_\{t\}\)=\\mathcal\{N\}\\\!\\left\(\\mathbf\{x\}\_\{t\-1\};\\;\\boldsymbol\{\\mu\}\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t\),\\;\\tilde\{\\beta\}\_\{t\}\\,\\mathbf\{I\}\\right\),\(18\)wherep\(𝐱T\)=𝒩\(𝟎,𝐈\)p\(\\mathbf\{x\}\_\{T\}\)=\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)\. Applying Jensen’s inequality to the log\-marginal likelihood yields the evidence lower bound \(ELBO\):
logpθ\(𝐱0\)≥𝔼q\(𝐱1:T∣𝐱0\)\[logpθ\(𝐱0:T\)−logq\(𝐱1:T∣𝐱0\)\]\.\\log p\_\{\\theta\}\(\\mathbf\{x\}\_\{0\}\)\\;\\geq\\;\\mathbb\{E\}\_\{q\(\\mathbf\{x\}\_\{1:T\}\\mid\\mathbf\{x\}\_\{0\}\)\}\\\!\\left\[\\log p\_\{\\theta\}\(\\mathbf\{x\}\_\{0:T\}\)\-\\log q\(\\mathbf\{x\}\_\{1:T\}\\mid\\mathbf\{x\}\_\{0\}\)\\right\]\.\(19\)Using the Markov factorizations ofqqandpθp\_\{\\theta\}, the ELBO decomposes into\[[17](https://arxiv.org/html/2607.22599#bib.bib9)\]
ELBO=\\displaystyle\\mathrm\{ELBO\}=𝔼q\[logpθ\(𝐱0∣𝐱1\)\]−DKL\(q\(𝐱T∣𝐱0\)∥p\(𝐱T\)\)\\displaystyle\{\\mathbb\{E\}\_\{q\}\\\!\\left\[\\log p\_\{\\theta\}\(\\mathbf\{x\}\_\{0\}\\mid\\mathbf\{x\}\_\{1\}\)\\right\]\}\\;\-\\;\{D\_\{\\mathrm\{KL\}\}\\\!\\left\(q\(\\mathbf\{x\}\_\{T\}\\mid\\mathbf\{x\}\_\{0\}\)\\,\\\|\\,p\(\\mathbf\{x\}\_\{T\}\)\\right\)\}\(20\)−∑t=2TDKL\(q\(𝐱t−1∣𝐱t,𝐱0\)∥pθ\(𝐱t−1∣𝐱t\)\)\.\\displaystyle\\;\-\\;\\sum\_\{t=2\}^\{T\}\{D\_\{\\mathrm\{KL\}\}\\\!\\left\(q\(\\mathbf\{x\}\_\{t\-1\}\\mid\\mathbf\{x\}\_\{t\},\\mathbf\{x\}\_\{0\}\)\\,\\\|\\,p\_\{\\theta\}\(\\mathbf\{x\}\_\{t\-1\}\\mid\\mathbf\{x\}\_\{t\}\)\\right\)\}\.The forward posteriorq\(𝐱t−1∣𝐱t,𝐱0\)q\(\\mathbf\{x\}\_\{t\-1\}\\mid\\mathbf\{x\}\_\{t\},\\mathbf\{x\}\_\{0\}\)is Gaussian with
q\(𝐱t−1∣𝐱t,𝐱0\)=𝒩\(𝐱t−1;𝝁~t\(𝐱t,𝐱0\),β~t𝐈\),q\(\\mathbf\{x\}\_\{t\-1\}\\mid\\mathbf\{x\}\_\{t\},\\mathbf\{x\}\_\{0\}\)=\\mathcal\{N\}\\\!\\left\(\\mathbf\{x\}\_\{t\-1\};\\;\\tilde\{\\boldsymbol\{\\mu\}\}\_\{t\}\(\\mathbf\{x\}\_\{t\},\\mathbf\{x\}\_\{0\}\),\\;\\tilde\{\\beta\}\_\{t\}\\,\\mathbf\{I\}\\right\),\(21\)where
𝝁~t=α¯t−1βt1−α¯t𝐱0\+αt\(1−α¯t−1\)1−α¯t𝐱t,β~t=\(1−α¯t−1\)\(1−α¯t\)βt\.\\tilde\{\\boldsymbol\{\\mu\}\}\_\{t\}=\\frac\{\\sqrt\{\\bar\{\\alpha\}\_\{t\-1\}\}\\,\\beta\_\{t\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\\,\\mathbf\{x\}\_\{0\}\+\\frac\{\\sqrt\{\\alpha\_\{t\}\}\\,\(1\-\\bar\{\\alpha\}\_\{t\-1\}\)\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\\,\\mathbf\{x\}\_\{t\},\\qquad\\tilde\{\\beta\}\_\{t\}=\\frac\{\(1\-\\bar\{\\alpha\}\_\{t\-1\}\)\}\{\(1\-\\bar\{\\alpha\}\_\{t\}\)\}\\,\\beta\_\{t\}\.\(22\)Sinceq\(𝐱t−1∣𝐱t,𝐱0\)q\(\\mathbf\{x\}\_\{t\-1\}\\mid\\mathbf\{x\}\_\{t\},\\mathbf\{x\}\_\{0\}\)andpθ\(𝐱t−1∣𝐱t\)p\_\{\\theta\}\(\\mathbf\{x\}\_\{t\-1\}\\mid\\mathbf\{x\}\_\{t\}\)share the same varianceβ~t\\tilde\{\\beta\}\_\{t\}, each KL term reduces to a squared difference between their means:
DKL\(q\(𝐱t−1∣𝐱t,𝐱0\)∥pθ\(𝐱t−1∣𝐱t\)\)=12β~t∥𝝁~t−𝝁θ\(𝐱t,t\)∥2\.D\_\{\\mathrm\{KL\}\}\\\!\\left\(q\(\\mathbf\{x\}\_\{t\-1\}\\mid\\mathbf\{x\}\_\{t\},\\mathbf\{x\}\_\{0\}\)\\,\\\|\\,p\_\{\\theta\}\(\\mathbf\{x\}\_\{t\-1\}\\mid\\mathbf\{x\}\_\{t\}\)\\right\)=\\frac\{1\}\{2\\tilde\{\\beta\}\_\{t\}\}\\left\\\|\\tilde\{\\boldsymbol\{\\mu\}\}\_\{t\}\-\\boldsymbol\{\\mu\}\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t\)\\right\\\|^\{2\}\.\(23\)
### A\.3Prediction Parameterizations
Two equivalent network parameterizations are commonly used\.
#### ϵ\\boldsymbol\{\\epsilon\}\-prediction\.
Substituting𝐱0=\(𝐱t−1−α¯tϵ\)/α¯t\\mathbf\{x\}\_\{0\}=\(\\mathbf\{x\}\_\{t\}\-\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\,\\boldsymbol\{\\epsilon\}\)/\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}into \([22](https://arxiv.org/html/2607.22599#A1.E22)\) and letting𝝁θ\\boldsymbol\{\\mu\}\_\{\\theta\}predictϵ\\boldsymbol\{\\epsilon\}via a networkϵθ\(𝐱t,t\)\\boldsymbol\{\\epsilon\}\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t\), each KL term becomes proportional to‖ϵ−ϵθ\(𝐱t,t\)‖2\\\|\\boldsymbol\{\\epsilon\}\-\\boldsymbol\{\\epsilon\}\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t\)\\\|^\{2\}\. Dropping the timestep\-dependent coefficient yields the simplified training objective\[[17](https://arxiv.org/html/2607.22599#bib.bib9)\]:
ℒsimple=𝔼t,𝐱0,ϵ\[‖ϵ−ϵθ\(𝐱t,t\)‖2\],t∼Uniform\{1,…,T\}\.\\mathcal\{L\}\_\{\\mathrm\{simple\}\}=\\mathbb\{E\}\_\{t,\\,\\mathbf\{x\}\_\{0\},\\,\\boldsymbol\{\\epsilon\}\}\\\!\\left\[\\left\\\|\\boldsymbol\{\\epsilon\}\-\\boldsymbol\{\\epsilon\}\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t\)\\right\\\|^\{2\}\\right\],\\qquad t\\sim\\mathrm\{Uniform\}\\\{1,\\ldots,T\\\}\.\(24\)
#### 𝐱0\\mathbf\{x\}\_\{0\}\-prediction\.
Alternatively, a networkfθ\(𝐱t,t\)f\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t\)can directly predict the clean target𝐱0\\mathbf\{x\}\_\{0\}\. The two parameterizations are linked by the bijection
ϵ=𝐱t−α¯t𝐱01−α¯t,𝐱0=𝐱t−1−α¯tϵα¯t\.\\boldsymbol\{\\epsilon\}=\\frac\{\\mathbf\{x\}\_\{t\}\-\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathbf\{x\}\_\{0\}\}\{\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\},\\qquad\\mathbf\{x\}\_\{0\}=\\frac\{\\mathbf\{x\}\_\{t\}\-\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\,\\boldsymbol\{\\epsilon\}\}\{\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\}\.\(25\)Given𝐱t\\mathbf\{x\}\_\{t\}andtt, predicting𝐱0\\mathbf\{x\}\_\{0\}is equivalent to predictingϵ\\boldsymbol\{\\epsilon\}, and the corresponding simplified objective takes the form
ℒ𝐱0=𝔼t,𝐱0,ϵ\[‖𝐱0−fθ\(𝐱t,t\)‖2\]\.\\mathcal\{L\}\_\{\{\\mathbf\{x\}\_\{0\}\}\}=\\mathbb\{E\}\_\{t,\\,\\mathbf\{x\}\_\{0\},\\,\\boldsymbol\{\\epsilon\}\}\\\!\\left\[\\left\\\|\\mathbf\{x\}\_\{0\}\-f\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t\)\\right\\\|^\{2\}\\right\]\.\(26\)
### A\.4DDIM: Non\-Markovian Reverse Process
Prior study\[[50](https://arxiv.org/html/2607.22599#bib.bib7)\]observe that the training objective in \([24](https://arxiv.org/html/2607.22599#A1.E24)\) depends only on the marginalsq\(𝐱t∣𝐱0\)q\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{0\}\)and not on the full jointq\(𝐱1:T∣𝐱0\)q\(\\mathbf\{x\}\_\{1:T\}\\mid\\mathbf\{x\}\_\{0\}\)\. This observation permits defining a family of non\-Markovian reverse transitions that share the same marginals\. Specifically, a reverse transition from stepttto stept−1t\{\-\}1is defined as
qσ\(𝐱t−1∣𝐱t,𝐱0\)=𝒩\(𝐱t−1;α¯t−1𝐱0\+1−α¯t−1−σt2⋅𝐱t−α¯t𝐱01−α¯t,σt2𝐈\)\.q\_\{\\sigma\}\(\\mathbf\{x\}\_\{t\-1\}\\mid\\mathbf\{x\}\_\{t\},\\mathbf\{x\}\_\{0\}\)=\\mathcal\{N\}\\\!\\left\(\\mathbf\{x\}\_\{t\-1\};\\;\\sqrt\{\\bar\{\\alpha\}\_\{t\-1\}\}\\,\\mathbf\{x\}\_\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\-1\}\-\\sigma\_\{t\}^\{2\}\}\\cdot\\frac\{\\mathbf\{x\}\_\{t\}\-\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathbf\{x\}\_\{0\}\}\{\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\},\\;\\sigma\_\{t\}^\{2\}\\,\\mathbf\{I\}\\right\)\.\(27\)To verify marginal consistency, substitute𝐱t=α¯t𝐱0\+1−α¯tϵ\\mathbf\{x\}\_\{t\}=\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathbf\{x\}\_\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\,\\boldsymbol\{\\epsilon\}withϵ∼𝒩\(𝟎,𝐈\)\\boldsymbol\{\\epsilon\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)into \([27](https://arxiv.org/html/2607.22599#A1.E27)\):
𝐱t−1=α¯t−1𝐱0\+1−α¯t−1−σt2ϵ\+σtϵ′,ϵ′∼𝒩\(𝟎,𝐈\)\.\\mathbf\{x\}\_\{t\-1\}=\\sqrt\{\\bar\{\\alpha\}\_\{t\-1\}\}\\,\\mathbf\{x\}\_\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\-1\}\-\\sigma\_\{t\}^\{2\}\}\\,\\boldsymbol\{\\epsilon\}\+\\sigma\_\{t\}\\,\\boldsymbol\{\\epsilon\}^\{\\prime\},\\qquad\\boldsymbol\{\\epsilon\}^\{\\prime\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)\.\(28\)Marginalizing overϵ\\boldsymbol\{\\epsilon\}andϵ′\\boldsymbol\{\\epsilon\}^\{\\prime\}gives meanα¯t−1𝐱0\\sqrt\{\\bar\{\\alpha\}\_\{t\-1\}\}\\,\\mathbf\{x\}\_\{0\}and variance\(1−α¯t−1−σt2\)\+σt2=1−α¯t−1\(1\-\\bar\{\\alpha\}\_\{t\-1\}\-\\sigma\_\{t\}^\{2\}\)\+\\sigma\_\{t\}^\{2\}=1\-\\bar\{\\alpha\}\_\{t\-1\}, confirming thatq\(𝐱t−1∣𝐱0\)=𝒩\(α¯t−1𝐱0,\(1−α¯t−1\)𝐈\)q\(\\mathbf\{x\}\_\{t\-1\}\\mid\\mathbf\{x\}\_\{0\}\)=\\mathcal\{N\}\(\\sqrt\{\\bar\{\\alpha\}\_\{t\-1\}\}\\,\\mathbf\{x\}\_\{0\},\\,\(1\-\\bar\{\\alpha\}\_\{t\-1\}\)\\mathbf\{I\}\)is preserved\.
Settingσt=0\\sigma\_\{t\}=0yields a deterministic reverse step:
𝐱t−1=α¯t−1𝐱0\+1−α¯t−11−α¯t\(𝐱t−α¯t𝐱0\)\.\\mathbf\{x\}\_\{t\-1\}=\\sqrt\{\\bar\{\\alpha\}\_\{t\-1\}\}\\,\\mathbf\{x\}\_\{0\}\+\\sqrt\{\\frac\{1\-\\bar\{\\alpha\}\_\{t\-1\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\}\\left\(\\mathbf\{x\}\_\{t\}\-\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathbf\{x\}\_\{0\}\\right\)\.\(29\)This extends to arbitrary step pairs\(t,s\)\(t,s\)withs<ts<t, enabling accelerated sampling via a subset of reverse steps without retraining\.
### A\.5Connection toDiffDiff
The standard forward marginal in \([17](https://arxiv.org/html/2607.22599#A1.E17)\) is a special case of the generalized marginal in \([1](https://arxiv.org/html/2607.22599#S3.E1)\) with𝐀t=𝐈\\mathbf\{A\}\_\{t\}=\\mathbf\{I\}for alltt\.DiffDiffintroduces a step\-dependent transition matrix𝐀t=\(1−λt\)𝐈\+λt𝐃2\\mathbf\{A\}\_\{t\}=\(1\-\\lambda\_\{t\}\)\\mathbf\{I\}\+\\lambda\_\{t\}\\mathbf\{D\}\_\{2\}, so the signal mean becomesα¯t𝐀t𝐱0\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathbf\{A\}\_\{t\}\\mathbf\{x\}\_\{0\}while the noise covariance remains isotropic\(1−α¯t\)𝐈\(1\-\\bar\{\\alpha\}\_\{t\}\)\\mathbf\{I\}\. The DDIM reverse framework extends naturally to this setting\. Given a denoiserfθf\_\{\\theta\}that predicts𝐱^0\\hat\{\\mathbf\{x\}\}\_\{0\}in the value domain, the noise estimate and reverse step become \([6](https://arxiv.org/html/2607.22599#S3.E6)\) and \([7](https://arxiv.org/html/2607.22599#S3.E7)\) in the main text\. The signal term uses𝐀s\\mathbf\{A\}\_\{s\}rather than𝐀t\\mathbf\{A\}\_\{t\}to ensure the reverse state lies in the correct representation at stepss\. Marginal consistency of this generalized reverse process is proved in Appendix[B\.2](https://arxiv.org/html/2607.22599#A2.SS2)\. A practical consequence of introducing𝐀t\\mathbf\{A\}\_\{t\}is thatϵ\\boldsymbol\{\\epsilon\}\-prediction is no longer interchangeable with𝐱0\\mathbf\{x\}\_\{0\}\-prediction\. Underϵ\\boldsymbol\{\\epsilon\}\-prediction, the noise\-matching loss‖ϵ−ϵ^‖2\\\|\\boldsymbol\{\\epsilon\}\-\\hat\{\\boldsymbol\{\\epsilon\}\}\\\|^\{2\}is equivalent to the weighted reconstruction errorα¯t1−α¯t‖𝐀t\(𝐱0−𝐱^0\)‖2\\frac\{\\bar\{\\alpha\}\_\{t\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\\\|\\mathbf\{A\}\_\{t\}\(\\mathbf\{x\}\_\{0\}\-\\hat\{\\mathbf\{x\}\}\_\{0\}\)\\\|^\{2\}\(see Appendix[B\.3](https://arxiv.org/html/2607.22599#A2.SS3)\), with the quadratic form𝐀t⊤𝐀t\\mathbf\{A\}\_\{t\}^\{\\top\}\\mathbf\{A\}\_\{t\}underpenalizing low\-frequency reconstruction error and overpenalizing high\-frequency error \(Corollary 1\)\. Since the forward operator𝐀t\\mathbf\{A\}\_\{t\}already encodes the predictability asymmetry by reshaping the noisy intermediate states, this loss\-side reweighting would compound the inductive bias rather than complement it, while standard evaluation metrics weight all frequencies uniformly in the value domain\.DiffDifftherefore confines the frequency selectivity to the forward process and adopts the𝐱0\\mathbf\{x\}\_\{0\}\-prediction parameterization, which directly minimizes the value\-domain error and stays aligned with the downstream evaluation\.
## Appendix BTheoretical Analysis
This appendix provides theoretical analysis of theDiffDiffframework\. We establish the spectral properties of the differencing operator \(Section[B\.1](https://arxiv.org/html/2607.22599#A2.SS1)\), prove marginal consistency of the generalized reverse process \(Section[B\.2](https://arxiv.org/html/2607.22599#A2.SS2)\), motivate the training objective from denoising consistency and task alignment \(Section[B\.3](https://arxiv.org/html/2607.22599#A2.SS3)\), show the hybrid loss is an upper bound on the unnormalized reconstruction error \(Section[B\.4](https://arxiv.org/html/2607.22599#A2.SS4)\), analyze the frequency\-dependent signal\-to\-noise ratio \(Section[B\.5](https://arxiv.org/html/2607.22599#A2.SS5)\), verify terminal distribution convergence \(Section[B\.6](https://arxiv.org/html/2607.22599#A2.SS6)\), and formalize the role of the label overlap in preserving cross\-boundary information \(Section[B\.7](https://arxiv.org/html/2607.22599#A2.SS7)\)\. All analyses are stated in terms of a single length\-SSsequence and apply per\-channel; for multivariate inputs the same operator and bounds apply independently to each channel\.
### B\.1Spectral Properties of the Differencing Operator
#### Proposition 1 \(Spectral characterization of𝐃2\\mathbf\{D\}\_\{2\}\)\.
Let𝐃2∈ℝS×S\\mathbf\{D\}\_\{2\}\\in\\mathbb\{R\}^\{S\\times S\}be the second\-order differencing operator defined in \([3](https://arxiv.org/html/2607.22599#S3.E3)\)\. For a discrete complex exponential𝐞ω=\[1,eiω,…,ei\(S−1\)ω\]⊤\\mathbf\{e\}\_\{\\omega\}=\[1,e^\{i\\omega\},\\ldots,e^\{i\(S\-1\)\\omega\}\]^\{\\top\}withω∈\[0,π\]\\omega\\in\[0,\\pi\], the per\-element gain at any internal positions≥3s\\geq 3satisfies
\|\[𝐃2𝐞ω\]s\|2=\|1−e−iω\|4=\(2−2cosω\)2\.\\left\|\[\\mathbf\{D\}\_\{2\}\\,\\mathbf\{e\}\_\{\\omega\}\]\_\{s\}\\right\|^\{2\}=\\left\|1\-e^\{\-i\\omega\}\\right\|^\{4\}=\(2\-2\\cos\\omega\)^\{2\}\.\(30\)
#### Proof\.
Letzs=ei\(s−1\)ωz\_\{s\}=e^\{i\(s\-1\)\\omega\}denote thess\-th entry of𝐞ω\\mathbf\{e\}\_\{\\omega\}under 1\-based indexing\. Fors≥3s\\geq 3:
\[𝐃2𝐞ω\]s\\displaystyle\[\\mathbf\{D\}\_\{2\}\\,\\mathbf\{e\}\_\{\\omega\}\]\_\{s\}=zs−2zs−1\+zs−2\\displaystyle=z\_\{s\}\-2z\_\{s\-1\}\+z\_\{s\-2\}=ei\(s−1\)ω−2ei\(s−2\)ω\+ei\(s−3\)ω\\displaystyle=e^\{i\(s\-1\)\\omega\}\-2e^\{i\(s\-2\)\\omega\}\+e^\{i\(s\-3\)\\omega\}=ei\(s−3\)ω\(e2iω−2eiω\+1\)\\displaystyle=e^\{i\(s\-3\)\\omega\}\\\!\\left\(e^\{2i\\omega\}\-2e^\{i\\omega\}\+1\\right\)=ei\(s−3\)ω\(eiω−1\)2\.\\displaystyle=e^\{i\(s\-3\)\\omega\}\(e^\{i\\omega\}\-1\)^\{2\}\.\(31\)Taking the squared modulus and using\|eiθ\|=1\|e^\{i\\theta\}\|=1:
\|\[𝐃2𝐞ω\]s\|2=\|eiω−1\|4=\[\(cosω−1\)2\+sin2ω\]2=\(2−2cosω\)2\.\\left\|\[\\mathbf\{D\}\_\{2\}\\,\\mathbf\{e\}\_\{\\omega\}\]\_\{s\}\\right\|^\{2\}=\|e^\{i\\omega\}\-1\|^\{4\}=\\bigl\[\(\\cos\\omega\-1\)^\{2\}\+\\sin^\{2\}\\omega\\bigr\]^\{2\}=\(2\-2\\cos\\omega\)^\{2\}\.Since\|eiω−1\|=\|1−e−iω\|\|e^\{i\\omega\}\-1\|=\|1\-e^\{\-i\\omega\}\|, the two forms in \([30](https://arxiv.org/html/2607.22599#A2.E30)\) are equal\.
#### Corollary 1 \(Frequency response of𝐀t\\mathbf\{A\}\_\{t\}\)\.
For𝐀t=\(1−λt\)𝐈\+λt𝐃2\\mathbf\{A\}\_\{t\}=\(1\-\\lambda\_\{t\}\)\\mathbf\{I\}\+\\lambda\_\{t\}\\mathbf\{D\}\_\{2\}withλt∈\[0,λmax\]\\lambda\_\{t\}\\in\[0,\\lambda\_\{\\max\}\], the per\-element frequency response at internal positionss≥3s\\geq 3is
H𝐀t\(ω\)=\(1−λt\)\+λt\(1−e−iω\)2\.H\_\{\\mathbf\{A\}\_\{t\}\}\(\\omega\)=\(1\-\\lambda\_\{t\}\)\+\\lambda\_\{t\}\(1\-e^\{\-i\\omega\}\)^\{2\}\.\(32\)In particular:
- •DC \(ω=0\\omega=0\):\|H𝐀t\(0\)\|2=\(1−λt\)2\|H\_\{\\mathbf\{A\}\_\{t\}\}\(0\)\|^\{2\}=\(1\-\\lambda\_\{t\}\)^\{2\}, so low\-frequency content is attenuated\.
- •Nyquist \(ω=π\\omega=\\pi\):\|H𝐀t\(π\)\|2=\(1\+3λt\)2\|H\_\{\\mathbf\{A\}\_\{t\}\}\(\\pi\)\|^\{2\}=\(1\+3\\lambda\_\{t\}\)^\{2\}, so high\-frequency content is amplified\.
#### Proof\.
By linearity, fors≥3s\\geq 3:
\[𝐀t𝐞ω\]s=\(1−λt\)ei\(s−1\)ω\+λtei\(s−3\)ω\(eiω−1\)2=ei\(s−1\)ω\[\(1−λt\)\+λte−2iω\(eiω−1\)2\]\.\[\\mathbf\{A\}\_\{t\}\\,\\mathbf\{e\}\_\{\\omega\}\]\_\{s\}=\(1\-\\lambda\_\{t\}\)\\,e^\{i\(s\-1\)\\omega\}\+\\lambda\_\{t\}\\,e^\{i\(s\-3\)\\omega\}\(e^\{i\\omega\}\-1\)^\{2\}=e^\{i\(s\-1\)\\omega\}\\\!\\left\[\(1\-\\lambda\_\{t\}\)\+\\lambda\_\{t\}\\,e^\{\-2i\\omega\}\(e^\{i\\omega\}\-1\)^\{2\}\\right\]\.Sincee−2iω\(eiω−1\)2=1−2e−iω\+e−2iω=\(1−e−iω\)2e^\{\-2i\\omega\}\(e^\{i\\omega\}\-1\)^\{2\}=1\-2e^\{\-i\\omega\}\+e^\{\-2i\\omega\}=\(1\-e^\{\-i\\omega\}\)^\{2\}, the bracketed term equalsH𝐀t\(ω\)H\_\{\\mathbf\{A\}\_\{t\}\}\(\\omega\)\.
Atω=0\\omega=0:\(1−e0\)2=0\(1\-e^\{0\}\)^\{2\}=0, givingH𝐀t\(0\)=1−λtH\_\{\\mathbf\{A\}\_\{t\}\}\(0\)=1\-\\lambda\_\{t\}\.
Atω=π\\omega=\\pi:\(1−e−iπ\)2=\(1\+1\)2=4\(1\-e^\{\-i\\pi\}\)^\{2\}=\(1\+1\)^\{2\}=4, givingH𝐀t\(π\)=1\+3λtH\_\{\\mathbf\{A\}\_\{t\}\}\(\\pi\)=1\+3\\lambda\_\{t\}\.
Note that𝐃2\\mathbf\{D\}\_\{2\}is not a circulant matrix, soH𝐀t\(ω\)H\_\{\\mathbf\{A\}\_\{t\}\}\(\\omega\)is not an eigenvalue of𝐀t\\mathbf\{A\}\_\{t\}but rather the per\-element gain for internal rows\. Rowss=1s=1\(zero output\) ands=2s=2\(first\-order difference\) have boundary\-specific behavior\. WhenS≫1S\\gg 1, the fraction of boundary rows is negligible\.
### B\.2Marginal Consistency of the Generalized Reverse Process
#### Theorem 1 \(Marginal consistency\)\.
Let the forward marginal beq\(𝐱t∣𝐱0\)=𝒩\(α¯t𝐀t𝐱0,\(1−α¯t\)𝐈\)q\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{0\}\)=\\mathcal\{N\}\\\!\\left\(\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathbf\{A\}\_\{t\}\\,\\mathbf\{x\}\_\{0\},\\;\(1\-\\bar\{\\alpha\}\_\{t\}\)\\,\\mathbf\{I\}\\right\)with step\-dependent transition matrices\{𝐀t\}t=0T\\\{\\mathbf\{A\}\_\{t\}\\\}\_\{t=0\}^\{T\}\. Consider the reverse update from stepttto stepss\(s<ts<t\):
𝐱s=α¯s𝐀s𝐱^0\+1−α¯s−σt2𝐱t−α¯t𝐀t𝐱^01−α¯t\+σtϵ′,\\mathbf\{x\}\_\{s\}=\\sqrt\{\\bar\{\\alpha\}\_\{s\}\}\\,\\mathbf\{A\}\_\{s\}\\,\\hat\{\\mathbf\{x\}\}\_\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{s\}\-\\sigma\_\{t\}^\{2\}\}\\;\\frac\{\\mathbf\{x\}\_\{t\}\-\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathbf\{A\}\_\{t\}\\,\\hat\{\\mathbf\{x\}\}\_\{0\}\}\{\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\}\+\\sigma\_\{t\}\\,\\boldsymbol\{\\epsilon\}^\{\\prime\},\(33\)whereϵ′∼𝒩\(𝟎,𝐈\)\\boldsymbol\{\\epsilon\}^\{\\prime\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)is independent of𝐱t\\mathbf\{x\}\_\{t\}andσt2≤1−α¯s\\sigma\_\{t\}^\{2\}\\leq 1\-\\bar\{\\alpha\}\_\{s\}\. If the denoiser is perfect, i\.e\.,𝐱^0=𝐱0\\hat\{\\mathbf\{x\}\}\_\{0\}=\\mathbf\{x\}\_\{0\}, then𝐱s\\mathbf\{x\}\_\{s\}has the correct forward marginal:
𝐱s∼q\(𝐱s∣𝐱0\)=𝒩\(α¯s𝐀s𝐱0,\(1−α¯s\)𝐈\)\.\\mathbf\{x\}\_\{s\}\\sim q\(\\mathbf\{x\}\_\{s\}\\mid\\mathbf\{x\}\_\{0\}\)=\\mathcal\{N\}\\\!\\left\(\\sqrt\{\\bar\{\\alpha\}\_\{s\}\}\\,\\mathbf\{A\}\_\{s\}\\,\\mathbf\{x\}\_\{0\},\\;\(1\-\\bar\{\\alpha\}\_\{s\}\)\\,\\mathbf\{I\}\\right\)\.\(34\)
#### Proof\.
Under𝐱^0=𝐱0\\hat\{\\mathbf\{x\}\}\_\{0\}=\\mathbf\{x\}\_\{0\}, write𝐱t\\mathbf\{x\}\_\{t\}using the forward marginal:
𝐱t=α¯t𝐀t𝐱0\+1−α¯tϵ,ϵ∼𝒩\(𝟎,𝐈\)\.\\mathbf\{x\}\_\{t\}=\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathbf\{A\}\_\{t\}\\,\\mathbf\{x\}\_\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\,\\boldsymbol\{\\epsilon\},\\qquad\\boldsymbol\{\\epsilon\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)\.\(35\)Substituting \([35](https://arxiv.org/html/2607.22599#A2.E35)\) and𝐱^0=𝐱0\\hat\{\\mathbf\{x\}\}\_\{0\}=\\mathbf\{x\}\_\{0\}into \([33](https://arxiv.org/html/2607.22599#A2.E33)\):
𝐱s\\displaystyle\\mathbf\{x\}\_\{s\}=α¯s𝐀s𝐱0\+1−α¯s−σt21−α¯tϵ1−α¯t\+σtϵ′\\displaystyle=\\sqrt\{\\bar\{\\alpha\}\_\{s\}\}\\,\\mathbf\{A\}\_\{s\}\\,\\mathbf\{x\}\_\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{s\}\-\\sigma\_\{t\}^\{2\}\}\\;\\frac\{\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\,\\boldsymbol\{\\epsilon\}\}\{\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\}\+\\sigma\_\{t\}\\,\\boldsymbol\{\\epsilon\}^\{\\prime\}=α¯s𝐀s𝐱0\+1−α¯s−σt2ϵ\+σtϵ′\.\\displaystyle=\\sqrt\{\\bar\{\\alpha\}\_\{s\}\}\\,\\mathbf\{A\}\_\{s\}\\,\\mathbf\{x\}\_\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{s\}\-\\sigma\_\{t\}^\{2\}\}\\;\\boldsymbol\{\\epsilon\}\+\\sigma\_\{t\}\\,\\boldsymbol\{\\epsilon\}^\{\\prime\}\.\(36\)The term𝐀t\\mathbf\{A\}\_\{t\}cancels entirely:𝐱t−α¯t𝐀t𝐱0=1−α¯tϵ\\mathbf\{x\}\_\{t\}\-\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathbf\{A\}\_\{t\}\\,\\mathbf\{x\}\_\{0\}=\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\,\\boldsymbol\{\\epsilon\}, which is independent of𝐀t\\mathbf\{A\}\_\{t\}\.
Sinceϵ\\boldsymbol\{\\epsilon\}andϵ′\\boldsymbol\{\\epsilon\}^\{\\prime\}are independent standard Gaussians,𝐱s\\mathbf\{x\}\_\{s\}conditioned on𝐱0\\mathbf\{x\}\_\{0\}is Gaussian with
𝔼\[𝐱s∣𝐱0\]\\displaystyle\\mathbb\{E\}\[\\mathbf\{x\}\_\{s\}\\mid\\mathbf\{x\}\_\{0\}\]=α¯s𝐀s𝐱0,\\displaystyle=\\sqrt\{\\bar\{\\alpha\}\_\{s\}\}\\,\\mathbf\{A\}\_\{s\}\\,\\mathbf\{x\}\_\{0\},Cov\[𝐱s∣𝐱0\]\\displaystyle\\mathrm\{Cov\}\[\\mathbf\{x\}\_\{s\}\\mid\\mathbf\{x\}\_\{0\}\]=\(1−α¯s−σt2\+σt2\)𝐈=\(1−α¯s\)𝐈\.\\displaystyle=\\left\(1\-\\bar\{\\alpha\}\_\{s\}\-\\sigma\_\{t\}^\{2\}\+\\sigma\_\{t\}^\{2\}\\right\)\\mathbf\{I\}=\(1\-\\bar\{\\alpha\}\_\{s\}\)\\,\\mathbf\{I\}\.This matchesq\(𝐱s∣𝐱0\)q\(\\mathbf\{x\}\_\{s\}\\mid\\mathbf\{x\}\_\{0\}\)exactly\. Since𝐀s\\mathbf\{A\}\_\{s\}and𝐀t\\mathbf\{A\}\_\{t\}are arbitrary \(they need not be related\), the result holds for any sequence of transition matrices\.
The proof shows that marginal consistency holds for*any*choice of step\-dependent transition matrices\{𝐀t\}\\\{\\mathbf\{A\}\_\{t\}\\\}, not only the specific𝐀t=\(1−λt\)𝐈\+λt𝐃2\\mathbf\{A\}\_\{t\}=\(1\-\\lambda\_\{t\}\)\\mathbf\{I\}\+\\lambda\_\{t\}\\mathbf\{D\}\_\{2\}used byDiffDiff\. The only requirement is that the reverse update \([33](https://arxiv.org/html/2607.22599#A2.E33)\) uses the correct𝐀t\\mathbf\{A\}\_\{t\}and𝐀s\\mathbf\{A\}\_\{s\}at their respective steps\. Setting𝐀t=𝐈\\mathbf\{A\}\_\{t\}=\\mathbf\{I\}for allttrecovers the standard DDIM marginal consistency result\.
### B\.3Training Objective Derivation
We derive the training objective used byDiffDiff, starting from the denoising formulation and explaining why𝐱0\\mathbf\{x\}\_\{0\}\-prediction is the natural parameterization for our generalized forward process\.
#### From denoising to reconstruction loss\.
The reverse process in \([7](https://arxiv.org/html/2607.22599#S3.E7)\) is fully determined by the denoiser output𝐱^0=fθ\(𝐱t,t,𝐱hist\)\\hat\{\\mathbf\{x\}\}\_\{0\}=f\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t,\\mathbf\{x\}\_\{\\text\{hist\}\}\)\. By Theorem 1, perfect denoising \(𝐱^0=𝐱0\\hat\{\\mathbf\{x\}\}\_\{0\}=\\mathbf\{x\}\_\{0\}\) guarantees exact marginal recovery at every reverse step\. The natural training objective therefore minimizes the denoising error:
ℒdenoise=𝔼t,𝐱0,ϵ\[‖𝐱0−fθ\(𝐱t,t,𝐱hist\)‖2\],t∼Uniform\{0,…,T−1\}\.\\mathcal\{L\}\_\{\\mathrm\{denoise\}\}=\\mathbb\{E\}\_\{t,\\,\\mathbf\{x\}\_\{0\},\\,\\boldsymbol\{\\epsilon\}\}\\\!\\left\[\\left\\\|\\mathbf\{x\}\_\{0\}\-f\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t,\\mathbf\{x\}\_\{\\text\{hist\}\}\)\\right\\\|^\{2\}\\right\],\\qquad t\\sim\\mathrm\{Uniform\}\\\{0,\\ldots,T\{\-\}1\\\}\.\(37\)
#### Why𝐱0\\mathbf\{x\}\_\{0\}\-prediction is task\-aligned\.
For the standard forward process \(𝐀t=𝐈\\mathbf\{A\}\_\{t\}=\\mathbf\{I\}\),ϵ\\boldsymbol\{\\epsilon\}\-prediction and𝐱0\\mathbf\{x\}\_\{0\}\-prediction are equivalent up to a scalar rescaling:‖ϵ−ϵ^‖2=α¯t1−α¯t‖𝐱0−𝐱^0‖2\\\|\\boldsymbol\{\\epsilon\}\-\\hat\{\\boldsymbol\{\\epsilon\}\}\\\|^\{2\}=\\frac\{\\bar\{\\alpha\}\_\{t\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\\\|\\mathbf\{x\}\_\{0\}\-\\hat\{\\mathbf\{x\}\}\_\{0\}\\\|^\{2\}\. For our generalized forward process, this equivalence breaks\. Given𝐱t=α¯t𝐀t𝐱0\+1−α¯tϵ\\mathbf\{x\}\_\{t\}=\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathbf\{A\}\_\{t\}\\,\\mathbf\{x\}\_\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\,\\boldsymbol\{\\epsilon\}andϵ^=\(𝐱t−α¯t𝐀t𝐱^0\)/1−α¯t\\hat\{\\boldsymbol\{\\epsilon\}\}=\(\\mathbf\{x\}\_\{t\}\-\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathbf\{A\}\_\{t\}\\,\\hat\{\\mathbf\{x\}\}\_\{0\}\)/\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}, the noise\-matching error becomes
‖ϵ−ϵ^‖2=α¯t1−α¯t‖𝐀t\(𝐱0−𝐱^0\)‖2=α¯t1−α¯t\(𝐱0−𝐱^0\)⊤𝐀t⊤𝐀t\(𝐱0−𝐱^0\)\.\\\|\\boldsymbol\{\\epsilon\}\-\\hat\{\\boldsymbol\{\\epsilon\}\}\\\|^\{2\}=\\frac\{\\bar\{\\alpha\}\_\{t\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\\,\\left\\\|\\mathbf\{A\}\_\{t\}\(\\mathbf\{x\}\_\{0\}\-\\hat\{\\mathbf\{x\}\}\_\{0\}\)\\right\\\|^\{2\}=\\frac\{\\bar\{\\alpha\}\_\{t\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\\,\(\\mathbf\{x\}\_\{0\}\-\\hat\{\\mathbf\{x\}\}\_\{0\}\)^\{\\top\}\\mathbf\{A\}\_\{t\}^\{\\top\}\\mathbf\{A\}\_\{t\}\\,\(\\mathbf\{x\}\_\{0\}\-\\hat\{\\mathbf\{x\}\}\_\{0\}\)\.\(38\)The quadratic form𝐀t⊤𝐀t\\mathbf\{A\}\_\{t\}^\{\\top\}\\mathbf\{A\}\_\{t\}introduces a position\-coupling that weights prediction errors non\-uniformly\. Under the interior\-symbol approximation \(Corollary 1, valid for positionss≥3s\\geq 3\),𝐀t\\mathbf\{A\}\_\{t\}amplifies high\-frequency components \(gain\(1\+3λt\)2\(1\+3\\lambda\_\{t\}\)^\{2\}at Nyquist\) and attenuates low\-frequency ones \(gain\(1−λt\)2\(1\-\\lambda\_\{t\}\)^\{2\}at DC\)\. The forward operator𝐀t\\mathbf\{A\}\_\{t\}already encodes the predictability asymmetry by reshaping the noisy intermediate states themselves, so a loss\-side frequency reweighting would compound this inductive bias rather than complement it, while standard evaluation metrics \(MSE, CRPS\) weight all frequencies uniformly in the value domain\. We therefore confine the frequency selectivity to the forward process and adopt the𝐱0\\mathbf\{x\}\_\{0\}\-prediction parameterization, which directly minimizes the value\-domain error‖𝐱0−fθ‖2\\\|\\mathbf\{x\}\_\{0\}\-f\_\{\\theta\}\\\|^\{2\}and stays aligned with the downstream evaluation\. The noise estimate needed for the reverse step is recovered from the value\-domain prediction via \([6](https://arxiv.org/html/2607.22599#S3.E6)\) without requiring matrix inversion\.
Under𝐱0\\mathbf\{x\}\_\{0\}\-prediction, the noise\-matching error at steptttakes the form
‖ϵ−ϵ^‖2=α¯t1−α¯t‖𝐀t\(𝐱0−fθ\(𝐱t,t,𝐱hist\)\)‖2,\\\|\\boldsymbol\{\\epsilon\}\-\\hat\{\\boldsymbol\{\\epsilon\}\}\\\|^\{2\}=\\frac\{\\bar\{\\alpha\}\_\{t\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\\,\\left\\\|\\mathbf\{A\}\_\{t\}\(\\mathbf\{x\}\_\{0\}\-f\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t,\\mathbf\{x\}\_\{\\text\{hist\}\}\)\)\\right\\\|^\{2\},\(39\)where the quadratic form‖𝐀t𝚫‖2=𝚫⊤𝐀t⊤𝐀t𝚫\\\|\\mathbf\{A\}\_\{t\}\\boldsymbol\{\\Delta\}\\\|^\{2\}=\\boldsymbol\{\\Delta\}^\{\\top\}\\mathbf\{A\}\_\{t\}^\{\\top\}\\mathbf\{A\}\_\{t\}\\,\\boldsymbol\{\\Delta\}couples different positions of the prediction error through the transition matrix\. The simplified objective in \([37](https://arxiv.org/html/2607.22599#A2.E37)\) drops both the timestep\-dependent scalarα¯t/\(1−α¯t\)\\bar\{\\alpha\}\_\{t\}/\(1\-\\bar\{\\alpha\}\_\{t\}\)and the matrix weighting𝐀t⊤𝐀t\\mathbf\{A\}\_\{t\}^\{\\top\}\\mathbf\{A\}\_\{t\}, treating all positions and timesteps uniformly\. This simplification is analogous to the unweightedϵ\\boldsymbol\{\\epsilon\}\-prediction loss\[[17](https://arxiv.org/html/2607.22599#bib.bib9)\]and is justified empirically by improved sample quality\.
#### Instance normalization and position weighting\.
To focus the denoiser on temporal dynamics rather than sample\-specific level and scale, the extended target𝐱0=\[𝐱label;𝐱fut\]\\mathbf\{x\}\_\{0\}=\[\\mathbf\{x\}\_\{\\text\{label\}\};\\,\\mathbf\{x\}\_\{\\text\{fut\}\}\]is instance\-normalized along the time dimension\[[22](https://arxiv.org/html/2607.22599#bib.bib32)\]:
𝐱~0=𝐱0−μ1ν,μ=mean\(𝐱0\),ν=Var\(𝐱0\)\+ϵnorm,\\tilde\{\\mathbf\{x\}\}\_\{0\}=\\frac\{\\mathbf\{x\}\_\{0\}\-\\mu\\,\\mathbf\{1\}\}\{\\nu\},\\qquad\\mu=\\mathrm\{mean\}\(\\mathbf\{x\}\_\{0\}\),\\quad\\nu=\\sqrt\{\\mathrm\{Var\}\(\\mathbf\{x\}\_\{0\}\)\+\\epsilon\_\{\\mathrm\{norm\}\}\},\(40\)whereμ,ν∈ℝ\\mu,\\nu\\in\\mathbb\{R\}are per\-instance scalars broadcast to allSSpositions,ϵnorm\>0\\epsilon\_\{\\mathrm\{norm\}\}\>0is a small constant ensuringν\>0\\nu\>0even for near\-constant windows \(we useϵnorm=10−6\\epsilon\_\{\\mathrm\{norm\}\}=10^\{\-6\}\), and the forward process operates on𝐱~0\\tilde\{\\mathbf\{x\}\}\_\{0\}\. The position\-weighted reconstruction loss in \([12](https://arxiv.org/html/2607.22599#S3.E12)\) assigns weightwi=wfut≥1w\_\{i\}=w\_\{\\mathrm\{fut\}\}\\geq 1to future positions andwi=1w\_\{i\}=1to label positions, directing optimization toward the forecasting segment\.
#### Auxiliary statistics prediction\.
Since the denoiser operates in the normalized space, an auxiliary head predicts\(μ^,ν^\)\(\\hat\{\\mu\},\\hat\{\\nu\}\)from the conditioning feature𝐜\\mathbf\{c\}\. These predictions are trained with MSE losses and used at inference to map the denoised output back to the original scale\. The total training objective is
ℒ=ℒrecon\+‖μ^−μ‖2\+‖ν^−ν‖2,\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{recon\}\}\+\\\|\\hat\{\\mu\}\-\\mu\\\|^\{2\}\+\\\|\\hat\{\\nu\}\-\\nu\\\|^\{2\},\(41\)reproducing \([13](https://arxiv.org/html/2607.22599#S3.E13)\) in the main text\.
### B\.4Hybrid Loss Upper Bound
#### Proposition 2 \(Hybrid loss as upper bound\)\.
Let𝐱~0=\(𝐱0−μ\)/ν\\tilde\{\\mathbf\{x\}\}\_\{0\}=\(\\mathbf\{x\}\_\{0\}\-\\mu\)/\\nube the instance\-normalized target whereμ,ν∈ℝ\\mu,\\nu\\in\\mathbb\{R\}are the per\-instance mean and standard deviation, and let\(μ^,ν^\)\(\\hat\{\\mu\},\\hat\{\\nu\}\)be the predicted statistics\. The unnormalized reconstruction error satisfies
‖ν^fθ\(𝐱t,t,𝐱hist\)\+μ^1−𝐱0‖2≤3\(ν^2ℛ\+\(ν^−ν\)2‖𝐱~0‖2\+S\(μ^−μ\)2\),\\left\\\|\\hat\{\\nu\}\\,f\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t,\\mathbf\{x\}\_\{\\text\{hist\}\}\)\+\\hat\{\\mu\}\\,\\mathbf\{1\}\-\\mathbf\{x\}\_\{0\}\\right\\\|^\{2\}\\;\\leq\\;3\\left\(\\hat\{\\nu\}^\{2\}\\,\\mathcal\{R\}\+\(\\hat\{\\nu\}\-\\nu\)^\{2\}\\,\\\|\\tilde\{\\mathbf\{x\}\}\_\{0\}\\\|^\{2\}\+S\\,\(\\hat\{\\mu\}\-\\mu\)^\{2\}\\right\),\(42\)whereℛ=‖fθ\(𝐱t,t,𝐱hist\)−𝐱~0‖2\\mathcal\{R\}=\\\|f\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t,\\mathbf\{x\}\_\{\\text\{hist\}\}\)\-\\tilde\{\\mathbf\{x\}\}\_\{0\}\\\|^\{2\}is the normalized reconstruction error\.
#### Proof\.
Write the unnormalized prediction as𝐱^0unnorm=ν^fθ\+μ^1\\hat\{\\mathbf\{x\}\}\_\{0\}^\{\\text\{unnorm\}\}=\\hat\{\\nu\}\\,f\_\{\\theta\}\+\\hat\{\\mu\}\\,\\mathbf\{1\}and the true target as𝐱0=ν𝐱~0\+μ1\\mathbf\{x\}\_\{0\}=\\nu\\,\\tilde\{\\mathbf\{x\}\}\_\{0\}\+\\mu\\,\\mathbf\{1\}\. Decompose the error:
𝐱^0unnorm−𝐱0\\displaystyle\\hat\{\\mathbf\{x\}\}\_\{0\}^\{\\text\{unnorm\}\}\-\\mathbf\{x\}\_\{0\}=ν^fθ\+μ^1−ν𝐱~0−μ1\\displaystyle=\\hat\{\\nu\}\\,f\_\{\\theta\}\+\\hat\{\\mu\}\\,\\mathbf\{1\}\-\\nu\\,\\tilde\{\\mathbf\{x\}\}\_\{0\}\-\\mu\\,\\mathbf\{1\}=ν^\(fθ−𝐱~0\)\+\(ν^−ν\)𝐱~0\+\(μ^−μ\)1\.\\displaystyle=\\hat\{\\nu\}\\,\(f\_\{\\theta\}\-\\tilde\{\\mathbf\{x\}\}\_\{0\}\)\+\(\\hat\{\\nu\}\-\\nu\)\\,\\tilde\{\\mathbf\{x\}\}\_\{0\}\+\(\\hat\{\\mu\}\-\\mu\)\\,\\mathbf\{1\}\.\(43\)Apply‖𝐚\+𝐛\+𝐜‖2≤3\(‖𝐚‖2\+‖𝐛‖2\+‖𝐜‖2\)\\\|\\mathbf\{a\}\+\\mathbf\{b\}\+\\mathbf\{c\}\\\|^\{2\}\\leq 3\(\\\|\\mathbf\{a\}\\\|^\{2\}\+\\\|\\mathbf\{b\}\\\|^\{2\}\+\\\|\\mathbf\{c\}\\\|^\{2\}\)\(from convexity of∥⋅∥2\\\|\\cdot\\\|^\{2\}and Jensen’s inequality\) to bound each term:
‖ν^\(fθ−𝐱~0\)‖2\\displaystyle\\\|\\hat\{\\nu\}\\,\(f\_\{\\theta\}\-\\tilde\{\\mathbf\{x\}\}\_\{0\}\)\\\|^\{2\}=ν^2‖fθ−𝐱~0‖2=ν^2ℛ,\\displaystyle=\\hat\{\\nu\}^\{2\}\\,\\\|f\_\{\\theta\}\-\\tilde\{\\mathbf\{x\}\}\_\{0\}\\\|^\{2\}=\\hat\{\\nu\}^\{2\}\\,\\mathcal\{R\},‖\(ν^−ν\)𝐱~0‖2\\displaystyle\\\|\(\\hat\{\\nu\}\-\\nu\)\\,\\tilde\{\\mathbf\{x\}\}\_\{0\}\\\|^\{2\}=\(ν^−ν\)2‖𝐱~0‖2,\\displaystyle=\(\\hat\{\\nu\}\-\\nu\)^\{2\}\\,\\\|\\tilde\{\\mathbf\{x\}\}\_\{0\}\\\|^\{2\},‖\(μ^−μ\)1‖2\\displaystyle\\\|\(\\hat\{\\mu\}\-\\mu\)\\,\\mathbf\{1\}\\\|^\{2\}=S\(μ^−μ\)2\.\\displaystyle=S\\,\(\\hat\{\\mu\}\-\\mu\)^\{2\}\.Combining gives \([42](https://arxiv.org/html/2607.22599#A2.E42)\)\.
#### Connection to the weighted training loss\.
The bound in \([42](https://arxiv.org/html/2607.22599#A2.E42)\) uses the unweighted reconstruction errorℛ=‖fθ−𝐱~0‖2\\mathcal\{R\}=\\\|f\_\{\\theta\}\-\\tilde\{\\mathbf\{x\}\}\_\{0\}\\\|^\{2\}\. Since the actual training loss employs position weightswi≥1w\_\{i\}\\geq 1\(withwi=wfut≥1w\_\{i\}=w\_\{\\mathrm\{fut\}\}\\geq 1for future positions andwi=1w\_\{i\}=1for label positions\), the per\-instance weighted lossℒreconinst=∑iwi\(fθ,i−x~0,i\)2\\mathcal\{L\}\_\{\\mathrm\{recon\}\}^\{\\mathrm\{inst\}\}=\\sum\_\{i\}w\_\{i\}\(f\_\{\\theta,i\}\-\\tilde\{x\}\_\{0,i\}\)^\{2\}satisfiesℛ≤ℒreconinst\\mathcal\{R\}\\leq\\mathcal\{L\}\_\{\\mathrm\{recon\}\}^\{\\mathrm\{inst\}\}, so minimizingℒrecon\\mathcal\{L\}\_\{\\mathrm\{recon\}\}also drivesℛ→0\\mathcal\{R\}\\to 0\. Combined with the auxiliary losses drivingμ^→μ\\hat\{\\mu\}\\to\\muandν^→ν\\hat\{\\nu\}\\to\\nu, the three\-term objective jointly controls the unnormalized prediction error\.
#### Boundedness ofν^\\hat\{\\nu\}\.
The factorν^2\\hat\{\\nu\}^\{2\}in the first term of \([42](https://arxiv.org/html/2607.22599#A2.E42)\) requires thatν^\\hat\{\\nu\}be bounded for the bound to be operationally meaningful\. In our implementation,ν^\\hat\{\\nu\}is produced by a linear projection without activation, so it is not constrained to be positive\. We rely on the auxiliary MSE loss‖ν^−ν‖2\\\|\\hat\{\\nu\}\-\\nu\\\|^\{2\}to driveν^\\hat\{\\nu\}toward the trueν\>0\\nu\>0\(which is guaranteed positive by theϵnorm\\epsilon\_\{\\mathrm\{norm\}\}stabilization in \([40](https://arxiv.org/html/2607.22599#A2.E40)\)\), ensuring the bound remains tight during training\.
### B\.5Frequency\-Dependent Signal\-to\-Noise Ratio
This section analyzes how the forward process design ofDiffDiffcreates a frequency\-dependent signal\-to\-noise ratio \(SNR\) profile that differs qualitatively from the isotropic corruption of standard diffusion\.
#### Per\-frequency SNR definition\.
Consider the forward marginalq\(𝐱t∣𝐱0\)=𝒩\(α¯t𝐀t𝐱0,\(1−α¯t\)𝐈\)q\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{0\}\)=\\mathcal\{N\}\(\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathbf\{A\}\_\{t\}\\,\\mathbf\{x\}\_\{0\},\\,\(1\-\\bar\{\\alpha\}\_\{t\}\)\\mathbf\{I\}\)\. The signal component at stepttisα¯t𝐀t𝐱0\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathbf\{A\}\_\{t\}\\,\\mathbf\{x\}\_\{0\}and the noise variance is\(1−α¯t\)\(1\-\\bar\{\\alpha\}\_\{t\}\)per dimension\. Since𝐃2\\mathbf\{D\}\_\{2\}is not a circulant matrix, discrete Fourier modes are not exact eigenvectors of𝐀t\\mathbf\{A\}\_\{t\}\. However, for the interior positions \(s≥3s\\geq 3\), which constitute a fraction\(S−2\)/S\(S\{\-\}2\)/Sof the sequence, the operator acts as a shift\-invariant filter with frequency responseH𝐀t\(ω\)H\_\{\\mathbf\{A\}\_\{t\}\}\(\\omega\)from Corollary 1\. ForS≫1S\\gg 1, we therefore analyze the*effective*per\-frequency signal power in the interior of𝐀t𝐱0\\mathbf\{A\}\_\{t\}\\mathbf\{x\}\_\{0\}as being modulated by\|H𝐀t\(ω\)\|2\|H\_\{\\mathbf\{A\}\_\{t\}\}\(\\omega\)\|^\{2\}relative to the original signal power\|X0\(ω\)\|2\|X\_\{0\}\(\\omega\)\|^\{2\}\. This motivates the effective per\-frequency SNR:
SNReff\(ω,t\)=α¯t\|H𝐀t\(ω\)\|2\|X0\(ω\)\|21−α¯t\.\\mathrm\{SNR\}\_\{\\mathrm\{eff\}\}\(\\omega,t\)=\\frac\{\\bar\{\\alpha\}\_\{t\}\\,\|H\_\{\\mathbf\{A\}\_\{t\}\}\(\\omega\)\|^\{2\}\\,\|X\_\{0\}\(\\omega\)\|^\{2\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\.\(44\)This definition captures the dominant behavior for long sequences; boundary corrections affect only the first two positions and become negligible asSSgrows\.
#### Standard diffusion \(𝐀t=𝐈\\mathbf\{A\}\_\{t\}=\\mathbf\{I\}\)\.
When𝐀t=𝐈\\mathbf\{A\}\_\{t\}=\\mathbf\{I\}for alltt,\|H𝐈\(ω\)\|2=1\|H\_\{\\mathbf\{I\}\}\(\\omega\)\|^\{2\}=1at all frequencies \(with no boundary issue since𝐈\\mathbf\{I\}is exactly Fourier\-diagonal\), so
SNReff,std\(ω,t\)=α¯t\|X0\(ω\)\|21−α¯t\.\\mathrm\{SNR\}\_\{\\mathrm\{eff,std\}\}\(\\omega,t\)=\\frac\{\\bar\{\\alpha\}\_\{t\}\\,\|X\_\{0\}\(\\omega\)\|^\{2\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\.\(45\)All frequency components share the same SNR decay profileα¯t/\(1−α¯t\)\\bar\{\\alpha\}\_\{t\}/\(1\-\\bar\{\\alpha\}\_\{t\}\), scaled only by the signal power\|X0\(ω\)\|2\|X\_\{0\}\(\\omega\)\|^\{2\}\. The isotropic forward process treats low\-frequency and high\-frequency information identically\.
#### DiffDiff\(𝐀t=\(1−λt\)𝐈\+λt𝐃2\\mathbf\{A\}\_\{t\}=\(1\{\-\}\\lambda\_\{t\}\)\\mathbf\{I\}\+\\lambda\_\{t\}\\mathbf\{D\}\_\{2\}\)\.
Using the frequency response from Corollary 1:
SNReff,DD\(ω,t\)=α¯t\|\(1−λt\)\+λt\(1−e−iω\)2\|2\|X0\(ω\)\|21−α¯t\.\\mathrm\{SNR\}\_\{\\mathrm\{eff,DD\}\}\(\\omega,t\)=\\frac\{\\bar\{\\alpha\}\_\{t\}\\,\\left\|\(1\-\\lambda\_\{t\}\)\+\\lambda\_\{t\}\(1\-e^\{\-i\\omega\}\)^\{2\}\\right\|^\{2\}\\,\|X\_\{0\}\(\\omega\)\|^\{2\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\.\(46\)
#### Proposition 3 \(Frequency\-selective SNR modulation, interior approximation\)\.
Under the interior symbol approximation \(valid for positionss≥3s\\geq 3; see Proposition 1\), the ratio ofDiffDiffto standard effective per\-frequency SNR is
SNReff,DD\(ω,t\)SNReff,std\(ω,t\)=\|H𝐀t\(ω\)\|2=\|\(1−λt\)\+λt\(1−e−iω\)2\|2\.\\frac\{\\mathrm\{SNR\}\_\{\\mathrm\{eff,DD\}\}\(\\omega,t\)\}\{\\mathrm\{SNR\}\_\{\\mathrm\{eff,std\}\}\(\\omega,t\)\}=\|H\_\{\\mathbf\{A\}\_\{t\}\}\(\\omega\)\|^\{2\}=\\left\|\(1\-\\lambda\_\{t\}\)\+\\lambda\_\{t\}\(1\-e^\{\-i\\omega\}\)^\{2\}\\right\|^\{2\}\.\(47\)This ratio satisfies:
1. 1\.At DC \(ω=0\\omega=0\):SNReff,DD/SNReff,std=\(1−λt\)2<1\\mathrm\{SNR\}\_\{\\mathrm\{eff,DD\}\}/\\mathrm\{SNR\}\_\{\\mathrm\{eff,std\}\}=\(1\-\\lambda\_\{t\}\)^\{2\}<1forλt∈\(0,1\)\\lambda\_\{t\}\\in\(0,1\)\.
2. 2\.At Nyquist \(ω=π\\omega=\\pi\):SNReff,DD/SNReff,std=\(1\+3λt\)2\>1\\mathrm\{SNR\}\_\{\\mathrm\{eff,DD\}\}/\\mathrm\{SNR\}\_\{\\mathrm\{eff,std\}\}=\(1\+3\\lambda\_\{t\}\)^\{2\}\>1forλt\>0\\lambda\_\{t\}\>0\.
3. 3\.Forλt∈\(0,1\)\\lambda\_\{t\}\\in\(0,1\), there exists a unique crossover frequencyω∗∈\(0,π\)\\omega^\{\*\}\\in\(0,\\pi\)where\|H𝐀t\(ω∗\)\|2=1\|H\_\{\\mathbf\{A\}\_\{t\}\}\(\\omega^\{\*\}\)\|^\{2\}=1\.
#### Proof\.
Items \(1\) and \(2\) follow directly from Corollary 1\. For item \(1\),\(1−λt\)2<1\(1\-\\lambda\_\{t\}\)^\{2\}<1holds wheneverλt∈\(0,1\)\\lambda\_\{t\}\\in\(0,1\)\(equivalently0<1−λt<10<1\-\\lambda\_\{t\}<1\)\. For item \(3\), we have\|H𝐀t\(0\)\|2=\(1−λt\)2<1\|H\_\{\\mathbf\{A\}\_\{t\}\}\(0\)\|^\{2\}=\(1\-\\lambda\_\{t\}\)^\{2\}<1and\|H𝐀t\(π\)\|2=\(1\+3λt\)2\>1\|H\_\{\\mathbf\{A\}\_\{t\}\}\(\\pi\)\|^\{2\}=\(1\+3\\lambda\_\{t\}\)^\{2\}\>1forλt∈\(0,1\)\\lambda\_\{t\}\\in\(0,1\)\. Since\|H𝐀t\(ω\)\|2\|H\_\{\\mathbf\{A\}\_\{t\}\}\(\\omega\)\|^\{2\}is a continuous function ofω\\omega\(as a composition of continuous functions\), the intermediate value theorem guarantees the existence ofω∗∈\(0,π\)\\omega^\{\*\}\\in\(0,\\pi\)with\|H𝐀t\(ω∗\)\|2=1\|H\_\{\\mathbf\{A\}\_\{t\}\}\(\\omega^\{\*\}\)\|^\{2\}=1\. To show uniqueness, substitutec=1−cosω∈\[0,2\]c=1\-\\cos\\omega\\in\[0,2\]and expand\|H𝐀t\(ω\)\|2=1\|H\_\{\\mathbf\{A\}\_\{t\}\}\(\\omega\)\|^\{2\}=1into the quadratic4c2−4\(1−λt\)c−\(2−λt\)=04c^\{2\}\-4\(1\{\-\}\\lambda\_\{t\}\)c\-\(2\{\-\}\\lambda\_\{t\}\)=0\. The two roots arec=\[\(1−λt\)±\(1−λt\)2\+\(2−λt\)\]/2c=\\bigl\[\(1\{\-\}\\lambda\_\{t\}\)\\pm\\sqrt\{\(1\{\-\}\\lambda\_\{t\}\)^\{2\}\+\(2\{\-\}\\lambda\_\{t\}\)\}\\bigr\]/2; the “−\-” root is negative \(since the discriminant exceeds1−λt1\{\-\}\\lambda\_\{t\}forλt∈\(0,1\)\\lambda\_\{t\}\\in\(0,1\)\), so exactly one root lies in\(0,2\)\(0,2\), giving a uniqueω∗=arccos\(1−c\)\\omega^\{\*\}=\\arccos\(1\-c\)\. In practice,λt≤λmax=0\.5<1\\lambda\_\{t\}\\leq\\lambda\_\{\\max\}=0\.5<1, so the condition is always satisfied\.
#### Forecasting interpretation\.
The SNR ratio in \([47](https://arxiv.org/html/2607.22599#A2.E47)\) reveals the inductive bias introduced by the differencing forward process:
- •*Low\-frequency suppression*\(ω\\omeganear0\)\. Level and trend components see their SNR reduced by a factor\(1−λt\)2\(1\-\\lambda\_\{t\}\)^\{2\}relative to standard diffusion\. Asλt\\lambda\_\{t\}increases with the diffusion step, these components are destroyed faster, reflecting the assumption that they can be anchored by the observed history\.
- •*High\-frequency preservation*\(ω\\omeganearπ\\pi\)\. Temporal dynamics, fluctuations, and local change patterns see their SNR amplified by\(1\+3λt\)2\(1\+3\\lambda\_\{t\}\)^\{2\}\. These components persist longer in the noisy intermediate states, giving the denoiser more signal about short\-term dynamics at later reverse steps\.
- •*Progressive transition*\. Sinceλt\\lambda\_\{t\}increases monotonically from0toλmax\\lambda\_\{\\max\}, the frequency selectivity grows with the diffusion step\. Early steps \(λt≈0\\lambda\_\{t\}\\approx 0\) remain nearly isotropic, while later steps emphasize the frequency separation\.
This frequency\-dependent SNR profile aligns with the forecasting\-specific division of labor: the history already anchors low\-frequency structure, so the generative process can focus its capacity on the harder\-to\-predict high\-frequency dynamics\.
The above analysis uses the interior symbolH𝐀t\(ω\)H\_\{\\mathbf\{A\}\_\{t\}\}\(\\omega\), which is exact only for positionss≥3s\\geq 3of the finite operator𝐀t\\mathbf\{A\}\_\{t\}\. For a length\-SSsequence, boundary effects ats∈\{1,2\}s\\in\\\{1,2\\\}introduce deviations from the idealized frequency response\. In practice, withS≥48S\\geq 48\(the shortest target sequence in our experiments\), boundary positions constitute at most≈4\.2%\\approx 4\.2\\%of the sequence\. Furthermore, the noise covariance remains exactly\(1−α¯t\)𝐈\(1\-\\bar\{\\alpha\}\_\{t\}\)\\mathbf\{I\}regardless of𝐀t\\mathbf\{A\}\_\{t\}\(which modulates only the signal mean, not the noise\), so the flat noise spectrum assumed in the SNR definition is exact\. The interior symbol therefore provides an accurate characterization of the dominant SNR behavior for long sequences\.
### B\.6Terminal Distribution Convergence
#### Proposition 4 \(Terminal convergence\)\.
For any𝐱0\\mathbf\{x\}\_\{0\}with‖𝐱0‖<∞\\\|\\mathbf\{x\}\_\{0\}\\\|<\\inftyand any sequence of matrices\{𝐀t\}\\\{\\mathbf\{A\}\_\{t\}\\\}with uniformly bounded operator norm‖𝐀t‖op≤C\\\|\\mathbf\{A\}\_\{t\}\\\|\_\{\\mathrm\{op\}\}\\leq Cfor alltt,
DKL\(q\(𝐱T∣𝐱0\)∥𝒩\(𝟎,𝐈\)\)→0asα¯T→0\.D\_\{\\mathrm\{KL\}\}\\\!\\left\(q\(\\mathbf\{x\}\_\{T\}\\mid\\mathbf\{x\}\_\{0\}\)\\;\\\|\\;\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)\\right\)\\;\\to\\;0\\qquad\\text\{as \}\\bar\{\\alpha\}\_\{T\}\\to 0\.\(48\)
#### Proof\.
Forq\(𝐱T∣𝐱0\)=𝒩\(𝝁T,𝚺T\)q\(\\mathbf\{x\}\_\{T\}\\mid\\mathbf\{x\}\_\{0\}\)=\\mathcal\{N\}\(\\boldsymbol\{\\mu\}\_\{T\},\\boldsymbol\{\\Sigma\}\_\{T\}\)with𝝁T=α¯T𝐀T𝐱0\\boldsymbol\{\\mu\}\_\{T\}=\\sqrt\{\\bar\{\\alpha\}\_\{T\}\}\\,\\mathbf\{A\}\_\{T\}\\mathbf\{x\}\_\{0\}and𝚺T=\(1−α¯T\)𝐈\\boldsymbol\{\\Sigma\}\_\{T\}=\(1\-\\bar\{\\alpha\}\_\{T\}\)\\mathbf\{I\}, the KL divergence to the standard Gaussian is
DKL=12\[tr\(𝚺T\)−S\+‖𝝁T‖2−lndet\(𝚺T\)\],D\_\{\\mathrm\{KL\}\}=\\frac\{1\}\{2\}\\\!\\left\[\\mathrm\{tr\}\(\\boldsymbol\{\\Sigma\}\_\{T\}\)\-S\+\\\|\\boldsymbol\{\\mu\}\_\{T\}\\\|^\{2\}\-\\ln\\det\(\\boldsymbol\{\\Sigma\}\_\{T\}\)\\right\],\(49\)whereSSis the dimensionality\. Substituting:
tr\(𝚺T\)\\displaystyle\\mathrm\{tr\}\(\\boldsymbol\{\\Sigma\}\_\{T\}\)=S\(1−α¯T\),\\displaystyle=S\(1\-\\bar\{\\alpha\}\_\{T\}\),lndet\(𝚺T\)\\displaystyle\\ln\\det\(\\boldsymbol\{\\Sigma\}\_\{T\}\)=Sln\(1−α¯T\),\\displaystyle=S\\ln\(1\-\\bar\{\\alpha\}\_\{T\}\),‖𝝁T‖2\\displaystyle\\\|\\boldsymbol\{\\mu\}\_\{T\}\\\|^\{2\}=α¯T‖𝐀T𝐱0‖2≤α¯TC2‖𝐱0‖2\.\\displaystyle=\\bar\{\\alpha\}\_\{T\}\\,\\\|\\mathbf\{A\}\_\{T\}\\mathbf\{x\}\_\{0\}\\\|^\{2\}\\leq\\bar\{\\alpha\}\_\{T\}\\,C^\{2\}\\\|\\mathbf\{x\}\_\{0\}\\\|^\{2\}\.Therefore
DKL=12\[−Sα¯T\+α¯T‖𝐀T𝐱0‖2−Sln\(1−α¯T\)\]\.D\_\{\\mathrm\{KL\}\}=\\frac\{1\}\{2\}\\\!\\left\[\-S\\bar\{\\alpha\}\_\{T\}\+\\bar\{\\alpha\}\_\{T\}\\\|\\mathbf\{A\}\_\{T\}\\mathbf\{x\}\_\{0\}\\\|^\{2\}\-S\\ln\(1\-\\bar\{\\alpha\}\_\{T\}\)\\right\]\.\(50\)Asα¯T→0\\bar\{\\alpha\}\_\{T\}\\to 0, usingln\(1−α¯T\)=−α¯T\+O\(α¯T2\)\\ln\(1\-\\bar\{\\alpha\}\_\{T\}\)=\-\\bar\{\\alpha\}\_\{T\}\+O\(\\bar\{\\alpha\}\_\{T\}^\{2\}\):
DKL=12\[−Sα¯T\+α¯T‖𝐀T𝐱0‖2\+Sα¯T\+O\(α¯T2\)\]=α¯T2‖𝐀T𝐱0‖2\+O\(α¯T2\)\.D\_\{\\mathrm\{KL\}\}=\\frac\{1\}\{2\}\\\!\\left\[\-S\\bar\{\\alpha\}\_\{T\}\+\\bar\{\\alpha\}\_\{T\}\\\|\\mathbf\{A\}\_\{T\}\\mathbf\{x\}\_\{0\}\\\|^\{2\}\+S\\bar\{\\alpha\}\_\{T\}\+O\(\\bar\{\\alpha\}\_\{T\}^\{2\}\)\\right\]=\\frac\{\\bar\{\\alpha\}\_\{T\}\}\{2\}\\\|\\mathbf\{A\}\_\{T\}\\mathbf\{x\}\_\{0\}\\\|^\{2\}\+O\(\\bar\{\\alpha\}\_\{T\}^\{2\}\)\.\(51\)Since‖𝐀T𝐱0‖2≤C2‖𝐱0‖2<∞\\\|\\mathbf\{A\}\_\{T\}\\mathbf\{x\}\_\{0\}\\\|^\{2\}\\leq C^\{2\}\\\|\\mathbf\{x\}\_\{0\}\\\|^\{2\}<\\infty, the KL divergence vanishes asα¯T→0\\bar\{\\alpha\}\_\{T\}\\to 0\. For theDiffDiffparameterization with𝐀t=\(1−λt\)𝐈\+λt𝐃2\\mathbf\{A\}\_\{t\}=\(1\-\\lambda\_\{t\}\)\\mathbf\{I\}\+\\lambda\_\{t\}\\mathbf\{D\}\_\{2\}, the operator norm satisfies‖𝐀t‖op≤\(1−λt\)\+λt‖𝐃2‖op\\\|\\mathbf\{A\}\_\{t\}\\\|\_\{\\mathrm\{op\}\}\\leq\(1\-\\lambda\_\{t\}\)\+\\lambda\_\{t\}\\\|\\mathbf\{D\}\_\{2\}\\\|\_\{\\mathrm\{op\}\}\. Since𝐃2\\mathbf\{D\}\_\{2\}is a fixed finite\-dimensional matrix,‖𝐃2‖op\\\|\\mathbf\{D\}\_\{2\}\\\|\_\{\\mathrm\{op\}\}is bounded, confirming the uniform bound assumption\.
### B\.7Extended Target and Cross\-Boundary Information
The extended target𝐱0=\[𝐱label;𝐱fut\]∈ℝS\\mathbf\{x\}\_\{0\}=\[\\mathbf\{x\}\_\{\\text\{label\}\};\\,\\mathbf\{x\}\_\{\\text\{fut\}\}\]\\in\\mathbb\{R\}^\{S\}withS=Ll\+HS=L\_\{l\}\+Hincludes a label window of lengthLl≥2L\_\{l\}\\geq 2from the end of the observed history\. This section formalizes why the label overlap is beneficial under the differencing forward process\.
#### Proposition 5 \(Cross\-boundary information in extended vs\. future\-only targets\)\.
Let𝐱fut=\[x1f,…,xHf\]⊤∈ℝH\\mathbf\{x\}\_\{\\text\{fut\}\}=\[x\_\{1\}^\{f\},\\ldots,x\_\{H\}^\{f\}\]^\{\\top\}\\in\\mathbb\{R\}^\{H\}be the future window and𝐱ext=\[x−Ll\+1,…,x0,x1f,…,xHf\]⊤∈ℝS\\mathbf\{x\}\_\{\\text\{ext\}\}=\[x\_\{\-L\_\{l\}\+1\},\\ldots,x\_\{0\},x\_\{1\}^\{f\},\\ldots,x\_\{H\}^\{f\}\]^\{\\top\}\\in\\mathbb\{R\}^\{S\}be the extended target wherex0x\_\{0\}is the last observed value\. Under𝐃2\\mathbf\{D\}\_\{2\}as defined in \([3](https://arxiv.org/html/2607.22599#S3.E3)\):
1. 1\.The future\-only target satisfies\[𝐃2\(H\)𝐱fut\]1=0\[\\mathbf\{D\}\_\{2\}^\{\(H\)\}\\mathbf\{x\}\_\{\\text\{fut\}\}\]\_\{1\}=0and\[𝐃2\(H\)𝐱fut\]2=x2f−x1f\[\\mathbf\{D\}\_\{2\}^\{\(H\)\}\\mathbf\{x\}\_\{\\text\{fut\}\}\]\_\{2\}=x\_\{2\}^\{f\}\-x\_\{1\}^\{f\}, encoding only differences within the future window\.
2. 2\.The extended target satisfies\[𝐃2\(S\)𝐱ext\]Ll\+1=x1f−2x0\+x−1\[\\mathbf\{D\}\_\{2\}^\{\(S\)\}\\mathbf\{x\}\_\{\\text\{ext\}\}\]\_\{L\_\{l\}\+1\}=x\_\{1\}^\{f\}\-2x\_\{0\}\+x\_\{\-1\}, encoding the second\-order difference across the history–future boundary\.
Consequently, the differencing forward process applied to the extended target preserves cross\-boundary curvature information at positions near the junction, which is lost when the target is restricted to the future window alone\.
#### Proof\.
Item \(1\) follows directly from the definition of𝐃2\(H\)\\mathbf\{D\}\_\{2\}^\{\(H\)\}operating onℝH\\mathbb\{R\}^\{H\}: position 1 outputs zero \(removes absolute level\), and position 2 outputsx2f−x1fx\_\{2\}^\{f\}\-x\_\{1\}^\{f\}\(first\-order difference within the future\)\. The first future valuex1fx\_\{1\}^\{f\}and its relationship to the last observed valuex0x\_\{0\}do not appear anywhere in𝐃2\(H\)𝐱fut\\mathbf\{D\}\_\{2\}^\{\(H\)\}\\mathbf\{x\}\_\{\\text\{fut\}\}\.
For item \(2\), in the extended target the position corresponding tox1fx\_\{1\}^\{f\}is at indexLl\+1L\_\{l\}\+1\(1\-based\)\. SinceLl\+1≥3L\_\{l\}\+1\\geq 3\(givenLl≥2L\_\{l\}\\geq 2\), the full second\-order difference applies:
\[𝐃2\(S\)𝐱ext\]Ll\+1=x1f−2x0\+x−1,\[\\mathbf\{D\}\_\{2\}^\{\(S\)\}\\mathbf\{x\}\_\{\\text\{ext\}\}\]\_\{L\_\{l\}\+1\}=x\_\{1\}^\{f\}\-2x\_\{0\}\+x\_\{\-1\},\(52\)which encodes the local curvature at the history–future boundary\. Similarly, positionLlL\_\{l\}encodesx0−2x−1\+x−2x\_\{0\}\-2x\_\{\-1\}\+x\_\{\-2\}, providing the curvature just before the boundary\.
The proposition shows that the label overlap serves a specific structural role beyond general “boundary anchoring”: it ensures that the differencing operator can compute cross\-boundary differences that link the future trajectory to the observed history\. Under the forward process𝐱t=α¯t𝐀t𝐱0\+1−α¯tϵ\\mathbf\{x\}\_\{t\}=\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathbf\{A\}\_\{t\}\\,\\mathbf\{x\}\_\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\,\\boldsymbol\{\\epsilon\}, these cross\-boundary differences are preserved in the signal component at intermediate diffusion steps \(with strength modulated byλt\\lambda\_\{t\}\)\. This provides the denoiser with explicit continuity cues that would be absent if only the future window were used as the diffusion target\.
## Appendix CDataset Statistics
Table 3:Dataset statistics\.
## Appendix DEvaluation Metrics
Letyi,ty\_\{i,t\}denote the ground\-truth value andy^i,t\\hat\{y\}\_\{i,t\}the point prediction for theii\-th forecast window at time steptt, withi=1,…,Ni=1,\\ldots,Nandt=1,…,Ht=1,\\ldots,H\. The point\-prediction metrics are mean squared error \(MSE\) and mean absolute error \(MAE\):
MSE=1NH∑i=1N∑t=1H\(yi,t−y^i,t\)2,MAE=1NH∑i=1N∑t=1H\|yi,t−y^i,t\|\.\\mathrm\{MSE\}=\\frac\{1\}\{NH\}\\sum\_\{i=1\}^\{N\}\\sum\_\{t=1\}^\{H\}\(y\_\{i,t\}\-\\hat\{y\}\_\{i,t\}\)^\{2\},\\qquad\\mathrm\{MAE\}=\\frac\{1\}\{NH\}\\sum\_\{i=1\}^\{N\}\\sum\_\{t=1\}^\{H\}\\lvert y\_\{i,t\}\-\\hat\{y\}\_\{i,t\}\\rvert\.\(53\)
Probabilistic forecast quality is measured by the continuous ranked probability score \(CRPS\), computed via a multi\-quantile approximation with\|Q\|=9\|Q\|=9uniformly spaced quantile levelsQ=\{0\.1,0\.2,…,0\.9\}Q=\\\{0\.1,0\.2,\\ldots,0\.9\\\}\. For each window we drawK=100K=100samples from the predictive distribution, and lety^i,t\(q\)\\hat\{y\}\_\{i,t\}^\{\(q\)\}denote the empiricalqq\-quantile at position\(i,t\)\(i,t\)\. Then
CRPS=2NH\|Q\|∑i=1N∑t=1H∑q∈Qρq\(yi,t−y^i,t\(q\)\),ρq\(u\)=max\{qu,\(q−1\)u\}\.\\mathrm\{CRPS\}=\\frac\{2\}\{NH\|Q\|\}\\sum\_\{i=1\}^\{N\}\\sum\_\{t=1\}^\{H\}\\sum\_\{q\\in Q\}\\rho\_\{q\}\\\!\\left\(y\_\{i,t\}\-\\hat\{y\}\_\{i,t\}^\{\(q\)\}\\right\),\\quad\\rho\_\{q\}\(u\)=\\max\\\{qu,\\,\(q\-1\)u\\\}\.\(54\)Lower values of all three metrics indicate better forecasts\.
## Appendix EConditioning Mechanism Variants
To verify the necessity of each design choice, we compareDiffDiffagainst four ablation variants, each replacing one component of this pathway\.
#### Variant descriptions\.
All variants share the same forward operator𝐀t=\(1−λt\)𝐈\+λt𝐃2\\mathbf\{A\}\_\{t\}=\(1\{\-\}\\lambda\_\{t\}\)\\mathbf\{I\}\+\\lambda\_\{t\}\\mathbf\{D\}\_\{2\}, the same value\-domain main encoding𝐡val=MLPval\(𝐱hist\)\\mathbf\{h\}\_\{\\text\{val\}\}=\\text\{MLP\}\_\{\\text\{val\}\}\(\\mathbf\{x\}\_\{\\text\{hist\}\}\), and the same denoiser backbone\. They differ only in how the differential signal is incorporated into the conditioning feature\.
- •CrossAttn\.The sigmoid gate is replaced by a single\-head cross\-attention layer\[[59](https://arxiv.org/html/2607.22599#bib.bib55)\]that selects between two views, the value\-domain encoding𝐡¯val\\bar\{\\mathbf\{h\}\}\_\{\\text\{val\}\}and the differential encoding𝐡¯diff\\bar\{\\mathbf\{h\}\}\_\{\\text\{diff\}\}\(both channel\-averaged\)\. The query is projected from\[𝐭emb;𝐡¯val\]\[\\mathbf\{t\}\_\{\\text\{emb\}\};\\,\\bar\{\\mathbf\{h\}\}\_\{\\text\{val\}\}\], and the keys and values are the two view tokens\. The softmax produces a convex combination of the two views, in contrast to the unconstrained per\-dimension sigmoid inDiffDiff\.
- •OpenGate\.The same architecture asDiffDiffexcept the gate bias is set to a large positive value \(σ≈1\\sigma\{\\approx\}1\), so the differential branch is always fully applied without timestep\-dependent modulation\. This isolates whether the gate’s selective\-suppression ability is necessary\.
- •DualOrder\.A second\-order differential branch encoding\[Δ2𝐱hist\]τ=𝐱hist,τ−2𝐱hist,τ−1\+𝐱hist,τ−2\[\\Delta^\{2\}\\mathbf\{x\}\_\{\\text\{hist\}\}\]\_\{\\tau\}=\\mathbf\{x\}\_\{\\text\{hist\},\\tau\}\-2\\mathbf\{x\}\_\{\\text\{hist\},\\tau\-1\}\+\\mathbf\{x\}\_\{\\text\{hist\},\\tau\-2\}\(withτ\\tauindexing positions within the history window\) is added alongside the first\-order branch, each with its own independent sigmoid gate:𝐡=𝐡val\+𝐠1⊙𝐡Δ1\+𝐠2⊙𝐡Δ2\\mathbf\{h\}=\\mathbf\{h\}\_\{\\text\{val\}\}\+\\mathbf\{g\}\_\{1\}\\odot\\mathbf\{h\}\_\{\\Delta^\{1\}\}\+\\mathbf\{g\}\_\{2\}\\odot\\mathbf\{h\}\_\{\\Delta^\{2\}\}\. This tests whether the conditioning pathway benefits from explicit second\-order differential information beyond the first\-order branch\.
- •Concat\.The sigmoid gate is removed and the two branches are fused by concatenation followed by a small MLP:𝐜=MLPfuse\(\[𝐡val;𝐡diff\]\)\\mathbf\{c\}=\\mathrm\{MLP\}\_\{\\mathrm\{fuse\}\}\(\[\\mathbf\{h\}\_\{\\text\{val\}\};\\,\\mathbf\{h\}\_\{\\text\{diff\}\}\]\)\. This tests whether the gain attributed to the stage\-adaptive gate could be explained by the additional model capacity introduced when both branches are present, rather than by the selective\-admission mechanism itself\.
Table[4](https://arxiv.org/html/2607.22599#A5.T4)compares all five variants on the seven main benchmarks across four prediction horizons \(five\-seed average\)\.DiffDiffobtains 54 of 84 wins, supporting the chosen design\.
Replacing the sigmoid gate with cross\-attention \(CrossAttn\) introduces a softmax\-induced zero\-sum trade\-off, where the convex combination must shift mass away from the value view to admit the differential signal\. CrossAttn wins concentrate on Exchange at short horizons and on Solar and Wind at long horizons, but the variant remains behindDiffDiffon the strongly periodic and stationary benchmarks where the per\-dimension sigmoid preserves both views simultaneously\.
Forcing the gate fully open \(OpenGate\) removes the model’s ability to suppress the differential branch when its signal is uninformative, exposing the gate’s selective\-suppression role as the source of cross\-regime robustness\.
Adding a second\-order branch \(DualOrder\) doubles the gating optimization load without consistent gain, with the variant retaining only one win atH=720H\{=\}720, consistent with extra capacity introducing optimization difficulty rather than additional usable structure\.
Replacing the gate with concatenation\-and\-MLP fusion \(Concat\) introduces strictly more parameters thanDiffDiffyet collapses to two wins, mirroring OpenGate’s failure mode and ruling out larger fusion capacity as the source ofDiffDiff’s gain\. The benefit thus stems specifically from the gate’s ability to suppress the differential branch when its signal is uninformative\.
A single first\-order differential branch fused with the value\-domain encoding through a learned sigmoid gate suffices to deliver the predictability\-aligned inductive bias while preserving robustness on data where differential information is uninformative\.
Table 4:Conditioning ablation\. Within each \(dataset, metric\) column the best value across variants is inboldand the second\-best isunderlined\. The rightmost column reports per\-row wins across the 21 \(dataset, metric\) settings\.
## Appendix FDifferencing Order Ablation
The default forward process uses𝐃2\\mathbf\{D\}\_\{2\}, whose first row removes the absolute level, second row applies first\-order differencing, and remaining rows apply second\-order differencing\. To validate this choice, we test two alternative constructions:
- •𝐃1\\mathbf\{D\}\_\{1\}: all rows use first\-order differencing\[−1,1\]\[\-1,\\,1\]\(captures velocity only\)\.
- •𝐃3\\mathbf\{D\}\_\{3\}: rows 1–2 same as𝐃2\\mathbf\{D\}\_\{2\}; rows 3\+ use third\-order differencing\[−1,3,−3,1\]\[\-1,\\,3,\\,\-3,\\,1\]\(adds jerk information\)\.
All three orders use theDiffDiffconditioning pathway with horizon\-adaptiveλ\\lambda\. Table[5](https://arxiv.org/html/2607.22599#A6.T5)reports results on the seven main benchmarks across four prediction horizons\.
Table 5:Differencing order ablation\.𝐃k\\mathbf\{D\}\_\{k\}denotes the order of the differencing operator in the forward process, with𝐃2\\mathbf\{D\}\_\{2\}as the default used byDiffDiff\. Within each column the best value across orders is inbold\. The rightmost column reports per\-row wins\.#### Analysis\.
𝐃2\\mathbf\{D\}\_\{2\}achieves the most wins overall, supporting its choice as the default differencing operator\.
𝐃1\\mathbf\{D\}\_\{1\}uses only first\-order differences and loses to𝐃2\\mathbf\{D\}\_\{2\}on most settings, since𝐃2\\mathbf\{D\}\_\{2\}’s additional curvature information provides a more discriminating inductive bias for the denoiser\. It remains competitive only on benchmarks where the historical signal is dominated by smooth or strongly\-trending dynamics\.
𝐃3\\mathbf\{D\}\_\{3\}shows striking improvements on benchmarks dominated by periodic structure, where the third\-order jerk information helps capture fast local change\. However, its more aggressive differencing amplifies extrapolation errors on non\-stationary series, producing severe long\-horizon degradation on benchmarks where the future trajectory deviates strongly from a smooth continuation of the history\. The strongly asymmetric behavior across data regimes makes𝐃3\\mathbf\{D\}\_\{3\}a less robust default than𝐃2\\mathbf\{D\}\_\{2\}\.
𝐃2\\mathbf\{D\}\_\{2\}therefore strikes a robust trade\-off: enough differential structure to inform the denoiser without amplifying extrapolation errors on non\-stationary regimes\.
## Appendix GSensitivity Analysis
We conduct sensitivity analyses on four key hyperparameters ofDiffDiff: the differencing strengthλmax\\lambda\_\{\\max\}, the label window lengthLlL\_\{l\}, the future loss weightwfutw\_\{\\mathrm\{fut\}\}, and the number of DDIM sampling stepsSS\.
Table 6:Sensitivity toλmax\\lambda\_\{\\max\}\.Table 7:Sensitivity to label window lengthLlL\_\{l\}\.Table 8:Sensitivity to future loss weightwfutw\_\{\\mathrm\{fut\}\}\.Table 9:Sensitivity to DDIM sampling stepsSS\.
## Appendix HStatistical Significance
We report a paired Wilcoxon signed\-rank test\[[10](https://arxiv.org/html/2607.22599#bib.bib38)\]on per\-window CRPS comparingDiffDiffagainst strong baselines, namely, MA\-TSD, NsDiff, and TMDM \(Table[10](https://arxiv.org/html/2607.22599#A8.T10)\)\.
Table 10:Paired Wilcoxon signed\-rank test on per\-window CRPS\.DiffDiffis strongly significant against MA\-TSD withp<10−13p<10^\{\-13\}on nearly all settings, confirming that the per\-window CRPS gains in the main table are not driven by seed variance\. Comparisons against NsDiff and TMDM reachp<0\.05p<0\.05on the majority of settings but fall short on a small set of long\-horizon Solar and Wind settings\.
Table 11:Exploratory comparison ofDiffDiffwith three recent deterministic forecasters, namely TQNet, Times2D, and PathFormer, on the same seven benchmarks across four prediction horizons\.Boldandunderlinemark the best and second\-best value within each column\.
## Appendix IComparison with Deterministic Forecasters
Recent point\-forecast architectures optimise anℓ2\\ell\_\{2\}training objective that aligns directly with MSE and MAE evaluation, and therefore carry a structural advantage on these point metrics\. They do not emit predictive distributions and cannot be scored on CRPS\. Table[11](https://arxiv.org/html/2607.22599#A8.T11)presents a comparison betweenDiffDiffand three recent state\-of\-the\-art models, namely TQNet\[[29](https://arxiv.org/html/2607.22599#bib.bib75)\], Times2D\[[36](https://arxiv.org/html/2607.22599#bib.bib76)\], and PathFormer\[[7](https://arxiv.org/html/2607.22599#bib.bib77)\]\.DiffDiffremains competitive on point metrics across this exploratory comparison\.
## Appendix JComputational Efficiency
We report peak GPU memory and per\-batch wall\-clock for both training and inference\. All measurements use ETTm2 withH=720H\{=\}720and batch size 32 on a single RTX 5090\. Training memory and time cover one optimizer step of each method\. Inference is unified across all methods to serial sampling: each method generates exactlynsample=100n\_\{\\text\{sample\}\}\{=\}100trajectories one trajectory at a time\.
Table 12:Computation efficiency comparison\. Bold face indicates the best result per row\.DiffDiffattains the best CRPS while keeping training memory and per\-step training time close to the lightest baseline \(MA\-TSD\), and also achieves the fastest serial inference among all six diffusion methods\. The predictability\-aligned forward operator and stage\-adaptive conditioning thus add negligible computational overhead in this measured setting\.
## Appendix KLimitations
The forward transition is parameterized by a fixed second\-order differencing matrix𝐃2\\mathbf\{D\}\_\{2\}, with only the interpolation strengthλt\\lambda\_\{t\}varying along the chain\. While the𝐃2\\mathbf\{D\}\_\{2\}choice is principled, namely a linear operator whose interior rows have a closed\-form spectral characterization that aligns the suppressed components with the slowly\-varying part of the target, it is not learned from data\. Datasets whose history\-anchored content is not concentrated at low frequencies, or whose under\-determined content is not concentrated at high frequencies, would benefit from a data\-aware𝐀t\\mathbf\{A\}\_\{t\}that places its attenuation budget along directions other than the differencing one\. We leave such a parameterized family of forward operators, jointly learned with the denoiser under a similar marginal\-consistency constraint, to future work\.Similar Articles
DynG-Diff: A State-Aware Dynamic Guidance Diffusion Framework for Probabilistic Time Series Forecasting
DynG-Diff proposes a state-aware dynamic guidance diffusion framework for probabilistic multivariate time series forecasting, improving robustness by adaptively handling variable heterogeneity.
Temporal Difference Learning for Diffusion Models
This paper introduces a temporal difference (TD) learning objective for diffusion models that enforces cross-time consistency along the denoising trajectory. It reformulates denoising as a reinforcement learning policy evaluation problem, showing significant improvements in sample quality (FID), especially for few-step samplers.
SCENARIODIFF: A Scenario-level Guidance Framework for Multimodal Time Series Forecasting--Extended Version
ScenarioDiff is a hierarchical contextual reasoning framework for multimodal time series forecasting that organizes textual context into three levels to guide a Multimodal Diffusion Transformer, showing effectiveness in event-driven domains.
When Denoising Hurts: Rethinking the Terminal Step of Diffusion Time Series Forecasters -- Extended Version
This paper challenges the assumption that iterative denoising always improves forecasts in diffusion-based time series forecasting, proposing a global stopping criterion and a Bernoulli sampler to enhance accuracy and speed.
Denoising the Future: Context-Aware Spectral Diffusion for Temporal Knowledge Graph Extrapolation
The paper proposes FreqDiff, a frequency-aware diffusion framework for temporal knowledge graph extrapolation that improves uncertainty modeling and achieves state-of-the-art performance on benchmarks.