泄漏积分器重建:控制递归差分时间序列预测中的误差积累

arXiv cs.AI 论文

摘要

介绍了泄漏积分器重建,以解决递归差分时间序列预测中的误差积累问题,展示了在不同架构和数据集上显著减少误差的效果。

arXiv:2609.23378v1 Announce Type: new Abstract: We introduce leaky-integrator reconstruction, a training-free method that cures the error accumulation of recursive differenced forecasting. Our first contribution is diagnostic: predicting one-step changes and integrating them by cumulative summation, the standard remedy for non-stationarity, is a discrete integrator with a pole on the unit circle, and we show this makes recursive rollout of a nonlinear model diverge, its 336-step error reaching several times that of a well-behaved forecaster (normalised MAE 1.6-3.8 versus about 0.8) across every neural architecture tested. Our second, central contribution is the fix: move the pole inside the unit circle with a leaky integrator H(z) = 1/(1 - gamma z^-1), gamma < 1, which provably bounds the accumulated error variance. Applied at reconstruction time with a single fixed gamma=0.9 (no retraining, a two-line change to any deployed one-step or foundation-model forecaster), it shrinks error at every horizon, the mean gain over seven diverging architectures and twenty datasets growing from ~3% at H=24 to 23% at H=96, 37% at H=192 and 51% (43-74% across those architectures) at H=336 (78% with an oracle pole). Crucially, it is provably inert where no pathology exists (stable or joint predictors already at the irreducible rate), making it a safe, general default.
查看原文
查看缓存全文

缓存时间: 2026/09/23 09:24

# Taming Error Accumulationin Recursive Differenced Time-Series Forecasting
Source: [https://arxiv.org/html/2609.23378](https://arxiv.org/html/2609.23378)
## Leaky\-Integrator Reconstruction: Taming Error Accumulation in Recursive Differenced Time\-Series Forecasting

###### Abstract

We introduce*leaky\-integrator reconstruction*, a training\-free method that cures the error accumulation of recursive differenced forecasting\. Our first contribution is diagnostic: predicting one\-step changes and integrating them by cumulative summation, the standard remedy for non\-stationarity, is a discrete integrator with a pole on the*unit circle*, and we show this makes recursive rollout of a nonlinear model*diverge*, its336336\-step error reaching several times that of a well\-behaved forecaster \(normalised MAE1\.61\.6–3\.83\.8versus about0\.80\.8\) across every neural architecture tested\. Our second, central contribution is the fix: move the pole*inside*the unit circle with a leaky integratorH⁡\(z\)=1/\(1−γ​z−1\)H\(z\)\{=\}1/\(1\{\-\}\\gamma z^\{\-1\}\),γ<1\\gamma\{<\}1, which provably bounds the accumulated error variance\. Applied at reconstruction time with a*single fixed*γ=0\.9\\gamma\{=\}0\.9\(no retraining, a two\-line change to any deployed one\-step or foundation\-model forecaster\), it shrinks error at every horizon, the mean gain over seven diverging architectures and twenty datasets*growing*from∼3%\\sim\\\!3\\%atH=24H\{=\}24to23%23\\%atH=96H\{=\}96,37%37\\%atH=192H\{=\}192and51%51\\%\(4343–74%74\\%across those architectures\) atH=336H\{=\}336\(78%78\\%with an oracle pole\)\. Crucially, it is provably inert where no pathology exists \(stable or joint predictors already at the irreducible rate\), making it a safe, general default\.

###### Index Terms:

Time\-series forecasting, differencing, error accumulation, recursive forecasting, leaky integrator

††address:New York University, Tandon School of Engineering, Brooklyn, NY 11201zy3110@nyu\.edu## 1Introduction

A pervasive way to forecast a non\-stationary series\{yt\}\\\{y\_\{t\}\\\}is to predict its first differenceΔ​yt=yt−yt−1\\Delta y\_\{t\}=y\_\{t\}\-y\_\{t\-1\}and reconstruct levels by cumulative summation\[[3](https://arxiv.org/html/2609.23378#bib.bib1),[6](https://arxiv.org/html/2609.23378#bib.bib12)\]\. This underlies ARIMA, autoregressive and probabilistic neural forecasters\[[16](https://arxiv.org/html/2609.23378#bib.bib10)\], recursively sampled foundation models, and is kin to the re\-centering of instance normalization\[[9](https://arxiv.org/html/2609.23378#bib.bib3)\]and linear detrending\[[19](https://arxiv.org/html/2609.23378#bib.bib4)\]\. When such a forecaster is rolled out*recursively*\(each predicted increment fed back to form the next input\), the reconstruction is a discrete integrator1/\(1−z−1\)1/\(1\-z^\{\-1\}\)with a pole atz=1z=1, and the per\-step errors are integrated without bound\. We show this is not benign: recursively rolled out*nonlinear*forecasters diverge, their336336\-step error growing several\-fold above a stable baseline\.

We propose a direct remedy from classical signal processing: move the pole inside the unit circle, i\.e\. reconstruct with a*leaky*integrator1/\(1−γ​z−1\)1/\(1\-\\gamma z^\{\-1\}\),γ<1\\gamma<1\. The fix is applied purely at reconstruction time \(no retraining, one scalar\), yet it rescues every diverging architecture\. Our contributions are: \(i\) a signal\-processing account of recursive differenced forecasting as a unit\-pole integrator, and evidence that nonlinear rollout diverges \(Sec\.[2](https://arxiv.org/html/2609.23378#S2),[3](https://arxiv.org/html/2609.23378#S3)\); \(ii\) the leaky\-integrator reconstruction, which provably bounds accumulation and, with a*single fixed*pole and no tuning, cutsH=336H\{=\}336error by4343–74%74\\%across the diverging architectures over twenty datasets, the gain growing with horizon \(Sec\.[5](https://arxiv.org/html/2609.23378#S5)\); and \(iii\) a clear delineation of*when*the remedy applies \(and, by design, does nothing\), separating recursive over\-accumulation from the irreducible growth of stable and joint predictors \(Sec\.[5](https://arxiv.org/html/2609.23378#S5)\)\.

## 2Method: Leaky\-Integrator Reconstruction

Let a one\-step model predict the incrementΔ​y^\\Delta\\hat\{y\}\. Rolling it out and reconstructing levels from the last observationyt−1y\_\{t\-1\}gives

y^t\+h=yt−1\+∑j=1hΔ​y^t\+j,\\hat\{y\}\_\{t\+h\}=y\_\{t\-1\}\+\\sum\_\{j=1\}^\{h\}\\Delta\\hat\{y\}\_\{t\+j\},\(1\)the discrete integratorH⁡\(z\)=1/\(1−z−1\)H\(z\)=1/\(1\-z^\{\-1\}\)with a pole atz=1z=1\. Writing the increment error asej=Δ​y^t\+j−Δ​yt\+je\_\{j\}=\\Delta\\hat\{y\}\_\{t\+j\}\-\\Delta y\_\{t\+j\}, the reconstructed level error is the running sumεh=∑j≤hej\\varepsilon\_\{h\}=\\sum\_\{j\\leq h\}e\_\{j\}; for zero\-mean, weakly correlatedeje\_\{j\}of varianceσ2\\sigma^\{2\},Var⁡\(εh\)≈h​σ2\\operatorname\{Var\}\(\\varepsilon\_\{h\}\)\\\!\\approx\\\!h\\sigma^\{2\}, soRMS⁡\(εh\)∝h\\mathrm\{RMS\}\(\\varepsilon\_\{h\}\)\\\!\\propto\\\!\\sqrt\{h\}for white errors; a learned model’s errors are*biased and correlated*, so the sum accumulates faster\. This excess is a property of the integration, not of recursive feedback \(Sec\.[3](https://arxiv.org/html/2609.23378#S3)\); the remedy is to damp the integrator\.

Our reconstruction\.We replace the pure integrator with a*leaky*one,Hγ​\(z\)=1/\(1−γ​z−1\)H\_\{\\gamma\}\(z\)=1/\(1\-\\gamma z^\{\-1\}\)with0<γ≤10<\\gamma\\leq 1:

y^t\+h=yt−1\+sh,sh=γ​sh−1\+Δ​y^t\+h\.\\hat\{y\}\_\{t\+h\}=y\_\{t\-1\}\+s\_\{h\},\\qquad s\_\{h\}=\\gamma\\,s\_\{h\-1\}\+\\Delta\\hat\{y\}\_\{t\+h\}\.\(2\)Older increments now decay geometrically asγh−j\\gamma^\{h\-j\}, moving the pole toz=γz=\\gammainside the unit circle\. Equivalently, \([2](https://arxiv.org/html/2609.23378#S2.E2)\) is a first\-order IIR low\-pass \(an exponentially\-weighted integrator\) on the predicted\-increment stream, withγ\\gammainterpolating between pure integration \(γ=1\\gamma\{=\}1\) and pure persistence \(γ=0\\gamma\{=\}0\)\. The accumulated error variance is then*bounded*,Var⁡\(εh\)→σ2/\(1−γ2\)\\operatorname\{Var\}\(\\varepsilon\_\{h\}\)\\\!\\to\\\!\\sigma^\{2\}/\(1\-\\gamma^\{2\}\): the pole trades anO⁡\(1−γ\)O\(1\{\-\}\\gamma\)steady\-state trend bias for arrested variance growth, so the optimalγ\\gammatracks the ratio of accumulated error variance to genuine trend: strong damping when the recursion diverges, none when it is stable\. Crucially \([2](https://arxiv.org/html/2609.23378#S2.E2)\) is post\-hoc: the increment model is untouched, so the fix costs a single scalar and no retraining\.

Figure 1:Synthetic validation of the error model\.*Left:*integrating an AR\(1\) increment\-error sequence of autocorrelationρ\\rhogives growth exponentα\\alpha; white noise yieldsα=0\.5\\alpha\{=\}0\.5, and mild anti\-correlation \(ρ≈−0\.2\\rho\\\!\\approx\\\!\-0\.2\) reproduces the observed0\.440\.44\.*Right:*the leaky integrator caps the accumulated RMS at1/1−γ21/\\sqrt\{1\-\\gamma^\{2\}\}, versus the unboundedh\\sqrt\{h\}of the pure integrator\.Synthetic validation\.Fig\.[1](https://arxiv.org/html/2609.23378#S2.F1)verifies both halves of the model directly\. Integrating a white increment\-error sequence givesα=0\.50\\alpha=0\.50to three digits; endowing the errors with AR\(1\) autocorrelationρ\\rhoshiftsα\\alphasmoothly, and the mild anti\-correlation of real increment errors \(ρ≈−0\.2\\rho\\\!\\approx\\\!\-0\.2\) recovers the empirical0\.440\.44\. Passing the same sequence through the leaky filter caps its accumulated RMS at1/1−γ21/\\sqrt\{1\-\\gamma^\{2\}\}\(5\.05\.0atγ=0\.98\\gamma\{=\}0\.98,2\.32\.3atγ=0\.9\\gamma\{=\}0\.9\), matching theory to within1%1\\%\.

## 3Why Traditional Recursion Fails

Figure 2:Per\-step error growthMAE⁡\(k\)/MAE⁡\(1\)\\mathrm\{MAE\}\(k\)/\\mathrm\{MAE\}\(1\)atH=336H\{=\}336, averaged over five of our architectures \(LSTM, GRU, DLinear, PatchTST, iTransformer\), sixteen datasets and three seeds\. Differenced \(change\) error tracks thek\\sqrt\{k\}integrator law and coincides with the random\-walk \(irreducible\) rate; level prediction grows far more slowly\.Table 1:Per\-step growth exponentα\\alphaand factorg⁡\(336\)g\(336\)of differenced \(change\) vs\. level \(direct\) forecasting, by model \(benchmark, joint prediction\)\. The change exponent is≈k\\approx\\\!\\sqrt\{k\}for every architecture\.The integrator makes error grow ask\\sqrt\{k\}with forecast stepkk\. On a benchmark of five forecasters over sixteen datasets and seven horizons, the change formulation grows with exponentα=0\.44\\alpha\{=\}0\.44\(Fig\.[2](https://arxiv.org/html/2609.23378#S3.F2)\), reaching14×14\\timesat336336steps versus5\.0×5\.0\\timesfor level prediction; the exponent is architecture\-independent \(α∈\[0\.41,0\.45\]\\alpha\\in\[0\.41,0\.45\], Table[1](https://arxiv.org/html/2609.23378#S3.T1)\) and coincides with the random\-walk baseline \(α=0\.41\\alpha\{=\}0\.41\), the*irreducible*rate for which persistence is optimal\. Thisk\\sqrt\{k\}is the floor for*white*errors; a learned model’s increment errors are biased and correlated, so integrating them sendsH=336H\{=\}336normalised MAE to1\.61\.6–3\.83\.8\(pure column, Table[3](https://arxiv.org/html/2609.23378#S5.T3)\), well above the≈0\.8\\approx\\\!0\.8floor\. Crucially this excess is intrinsic to the*integration*, not to feedback: a*teacher\-forced*rollout \(every increment predicted from the true past, no feedback\) accumulates*as much or more*\(meanH=336H\{=\}336nMAE over MLP/GRU/Transformer:2\.32\.3teacher\-forced vs\.1\.81\.8recursive\), and free\-running recursion even self\-damps as its inputs contract\. Damping the integrator removes this excess, leaving the irreduciblek\\sqrt\{k\}\.

## 4Experimental Setup

Data\.We use twenty univariate target series spanning ten domains and a wide range of sampling rates, from daily financial data to1010–1515min energy, climate and traffic sensors: cryptocurrencies \(BTC, ETH, SOL, BNB, XRP\), equity indices \(S&P 500, Nikkei\), commodities \(oil, gold\), rates/volatility \(10y yield, VIX\), FX \(exchange rate\), energy \(electricity\), climate \(weather, temperature\), traffic, and industrial ETT sensors \(ETTh1/2, ETTm1/2\)\. Each series is split70/10/2070/10/20*chronologically*\(train/val/test\) to preclude look\-ahead leakage, andzz\-standardised using only training\-window statistics so that normalised MAE is scale\-free and comparable across domains\. We group the series into three regimes:*drift*\(12 trending financial/commodity series, where differencing is needed\),*structured*\(4 ETT sensors with strong daily/weekly seasonality\), and*near\-stationary*\(4: electricity, weather, temperature, traffic\), which isolate where the integrated errors drift most\.

Architectures\.We evaluate nine one\-step increment predictors: linear autoregression, MLP, LSTM\[[8](https://arxiv.org/html/2609.23378#bib.bib2)\], GRU, DLinear\[[19](https://arxiv.org/html/2609.23378#bib.bib4)\], a dilated TCN, PatchTST\[[14](https://arxiv.org/html/2609.23378#bib.bib6)\], iTransformer\[[11](https://arxiv.org/html/2609.23378#bib.bib7)\]and a small Transformer\. Our LSTM, GRU and Transformer rollouts are genuine autoregressive decoders \(the native mode of ARIMA, DeepAR and sampled foundation models\), where exposure bias\[[2](https://arxiv.org/html/2609.23378#bib.bib11)\]is expected; DLinear, PatchTST and iTransformer are natively*joint*multi\-output models that we additionally roll out to test whether the mechanism is architecture\-general \(their joint use is inert, Sec\.[5](https://arxiv.org/html/2609.23378#S5)\)\. These are compact univariate reimplementations that keep each design’s core \(patching, inverted attention, linear decomposition\); the univariate iTransformer in particular loses its cross\-variate mixing, so its rollout figures probe the mechanism rather than rank the architecture\.

Training\.Every model is trained once \(single seed\) to predict the next normalised incrementΔ​zt\\Delta z\_\{t\}from the previousL=96L\{=\}96increments, minimising a one\-step MSE with Adam \(learning rate10−310^\{\-3\}, batch256256, a few epochs over the sliding windows of the training split, sub\-sampled for the longest series\)\. The networks are deliberately compact, so that the reconstruction pole, not model capacity, is the variable under study, with a6464\-unit MLP,3232–6464\-unit recurrent cells, a three\-block dilated TCN, a one\-layer width\-3232four\-head Transformer, and lightweight univariate PatchTST/iTransformer encoders\. The direct multi\-horizon \(direct\-MH\) variant of each shares this backbone but replaces the one\-step head with anHH\-output head trained jointly on allHHfuture increments\.

Rollout, reconstruction and metric\.At each of up to200200stride\-spaced forecast origins per test series \(series yielding fewer than1515valid origins are dropped\), we roll the one\-step model out autoregressively toH=336H\{=\}336, feeding each predicted increment back as the next input, and reconstruct levels from the last observation either with the traditional pure integrator \(γ=1\\gamma\{=\}1, a cumulative sum\) or with our leaky filter \(Eq\. \([2](https://arxiv.org/html/2609.23378#S2.E2)\), a first\-order IIR\)\. We report*normalised MAE*\(MAE/σtrain\\text\{MAE\}/\\sigma\_\{\\text\{train\}\}, scale\-free and averageable across series\) at horizons\{24,96,192,336\}\\\{24,96,192,336\\\}, averaged over origins then datasets\. For the oracle upper bound the pole is swept over a1313\-point grid in\[0,1\]\[0,1\]; unless stated we use the deployableγ=0\.9\\gamma\{=\}0\.9\.

## 5Results

Table 2:Per\-dataset normalised MAE atH=336H\{=\}336\(mean over the seven nonlinear architectures\)\. Our fixed\-γ=0\.9\\gamma\{=\}0\.9reconstruction lowers error over the traditional recursive integrator \(rec\.\) on*nineteen of twenty*datasets, dramatically where recursion diverges \(e\.g\. Traffic→1\.067\.07\\\!\\to\\\!1\.06, Weather→1\.062\.93\\\!\\to\\\!1\.06\), the sole exception being the near\-random\-walk SP500 \(→0\.960\.90\\\!\\to\\\!0\.96\)\.Boldmarks the better of rec\. and ours; a well\-trained direct multi\-horizon model \(d\-MH\) is shown for reference\.Table 3:Our leaky reconstruction vs\. the traditional pure\-recursive integrator, atH=336H\{=\}336\(normalised MAE,2020datasets, nine architectures\)\. “ours” is a*single fixed*poleγ=0\.9\\gamma\{=\}0\.9\(deployable, no tuning;γ=1\\gamma\{=\}1for stable linear\-AR\); “oracle” is the best per\-dataset pole \(an upper bound\); “Impr\.” is the fixed\-γ\\gammareduction over the pure integrator\. Nonlinear recursion diverges and the fixed pole rescues it; stable linear\-AR is correctly unchanged\.Figure 3:*\(a\)*Improvement of our leaky reconstruction \(fixedγ=0\.9\\gamma\{=\}0\.9\) over the traditional pure\-recursive integrator vs\. horizon, the advantage grows monotonically withHHfor every nonlinear architecture and is flat for stable linear\-AR\.*\(b\)*AtH=336H\{=\}336the pure integrator diverges \(nMAE1\.61\.6–3\.83\.8\); our method restores every diverging architecture to≈0\.87\\approx\\\!0\.87–0\.970\.97\.Improvement over traditional recursion \(nine architectures\)\.We train a one\-step increment predictor for each of nine architectures \(linear\-AR, MLP, LSTM, GRU, DLinear, TCN, PatchTST, iTransformer, Transformer\), roll each out toH=336H\{=\}336on twenty datasets, and reconstruct with the pure \(γ=1\\gamma\{=\}1\) versus our leaky pole\. Table[3](https://arxiv.org/html/2609.23378#S5.T3)and Fig\.[3](https://arxiv.org/html/2609.23378#S5.F3)give the result\. Every nonlinear model diverges under the pure integrator; a*single fixed*poleγ=0\.9\\gamma\{=\}0\.9pulls all of them back to≈0\.87\\approx\\\!0\.87–0\.970\.97nMAE, a4343–74%74\\%error reduction\(the Transformer, which diverges worst at3\.83\.8, gains most; an oracle per\-dataset pole reaches0\.810\.81–0\.860\.86, i\.e\. up to78%78\\%\); Table[2](https://arxiv.org/html/2609.23378#S5.T2)gives the full per\-dataset breakdown\. Linear\-AR is stable and, correctly, left atγ=1\\gamma\{=\}1\. Fig\.[4](https://arxiv.org/html/2609.23378#S5.F4)shows the mechanism on individual forecasts, the pure integrator drifts away while our reconstruction stays anchored, and Table[4](https://arxiv.org/html/2609.23378#S5.T4)breaks the effect down by regime: divergence is worst on*near\-stationary*series \(pure nMAE3\.533\.53, where a trendless target gives the integrated increment errors no real trend to anchor to\), which the fixed pole rescues to1\.101\.10\. Over three seeds the fixed\-γ=0\.9\\gamma\{=\}0\.9error is stable \(std≤0\.24\\leq 0\.24nMAE\) while the pure integrator is seed\-erratic \(Transformer std5\.45\.4\), so the fix also removes recursion’s run\-to\-run instability\.

Figure 4:A single336336\-step recursive Transformer forecast on two near\-stationary series\. The pure integrator \(γ=1\\gamma\{=\}1\) drifts away, down on Weather, up on Temp, as the integrated increment errors accumulate, while the leaky reconstruction \(γ=0\.9\\gamma\{=\}0\.9\) stays anchored to the true level\.Table 4:Recursive rollout by regime atH=336H\{=\}336\(nMAE, mean over the seven nonlinear architectures\)\. The pure integrator diverges most on near\-stationary series; our reconstruction rescues every regime, approaching a well\-trained direct multi\-horizon model \(direct\-MH, shown for reference\) at no training cost\.Table 5:Improvement of our reconstruction \(fixedγ=0\.9\\gamma\{=\}0\.9\) over the traditional recursive integrator \(% reduction in nMAE\)*by horizon*\. The advantage grows monotonically withHH; at short horizons the fixed pole can slightly over\-damp \(negative entries\), asγ∗→1\\gamma^\{\\ast\}\\\!\\to\\\!1whenH→0H\{\\to\}0\(Fig\.[5](https://arxiv.org/html/2609.23378#S5.F5)\)\. Linear\-AR is held atγ=1\\gamma\{=\}1\.The advantage grows with the horizon\.Fig\.[3](https://arxiv.org/html/2609.23378#S5.F3)\(a\) and Table[5](https://arxiv.org/html/2609.23378#S5.T5)show the improvement is negligible at short horizons \(where little has accumulated, and a fixed pole can even over\-damp\) and rises monotonically to4343–74%74\\%atH=336H\{=\}336for the recurrent and attention models, so the method is most valuable exactly at the long horizons where naive recursion is most damaged\.

Figure 5:Setting the pole\.*Left:*H=336H\{=\}336error vs\.γ\\gamma\(γ\\gammadecreasing rightward, log scale\), nonlinear models fall steeply and plateau belowγ≈0\.8\\gamma\\\!\\approx\\\!0\.8; stable linear\-AR degrades under any damping\.*Right:*the optimal poleγ∗\\gamma^\{\\ast\}falls with the horizon\.The pole needs no tuning\.Fig\.[5](https://arxiv.org/html/2609.23378#S5.F5)shows the error is a smooth function ofγ\\gammathat plateaus onceγ≲0\.9\\gamma\\\!\\lesssim\\\!0\.9, with the optimum shifting to stronger damping at longer horizons\. This is why we headline a*single fixed*γ=0\.9\\gamma\{=\}0\.9: it recovers almost all of the oracle gain \(\+43%\+43\\%vs\.\+46%\+46\\%for MLP;\+74%\+74\\%vs\.\+78%\+78\\%for the Transformer, Table[3](https://arxiv.org/html/2609.23378#S5.T3)\) with no per\-dataset search\. Selectingγ\\gammaper \(model,dataset\) on a held\-out validation split reaches test nMAE0\.760\.76, just below the fixedγ=0\.9\\gamma\{=\}0\.9\(0\.810\.81\) and near the oracle \(0\.730\.73\): the fixed default leaves little on the table\. The one rule: withhold damping from a*stable*recursion\. Applied to linear\-AR,γ=0\.9\\gamma\{=\}0\.9instead*costs*9\.3%9\.3\\%, as there is no divergence to arrest, a case the super\-k\\sqrt\{k\}diagnostic \(Sec\.[6](https://arxiv.org/html/2609.23378#S6)\) detects automatically\.

Relation to direct multi\-horizon prediction\.The usual alternative, a*direct*multi\-horizon \(direct\-MH\) model emitting allHHsteps jointly, is a strong baseline overall competitive with our reconstructed one\-step model \(Tables[2](https://arxiv.org/html/2609.23378#S5.T2),[4](https://arxiv.org/html/2609.23378#S5.T4)\), but it is a separate, wider model with its own training run; our fix instead reuses an existing predictor at no cost, so the two are complementary\.

Scope\.The remedy targets integration*over*\-accumulation, not the irreduciblek\\sqrt\{k\}: it is inert for stable linear recursion and for*joint*multi\-output forecasters \(which never integrate one\-step increments, gaining≈0%\\approx\\\!0\\%\), where accuracy is governed by the signal and the target choice\.

## 6Discussion

Practical guidance\.Deployment is a two\-line change \(reconstruct with a fixedγ=0\.9\\gamma\{=\}0\.9, no retraining\) gated by a one\-line diagnostic: damp only when the pure\-integrator error grows super\-k\\sqrt\{k\}on a validation split, else leave a stable or joint model atγ=1\\gamma\{=\}1\.

Cost and generality\.The remedy adds no parameters and no training: at inference it is anO⁡\(H\)O\(H\)first\-order filter over the predicted increments\. That a singleγ=0\.9\\gamma\{=\}0\.9works across every diverging architecture is no accident: the divergence is a property of the unit\-pole reconstruction, not of any one model, and the error\-versus\-γ\\gammacurve is flat over a wide plateau \(Fig\.[5](https://arxiv.org/html/2609.23378#S5.F5)\)\.

## 7Relation to Prior Work

Differencing and stationarisation\.Differencing to induce stationarity is the foundation of Box–Jenkins ARIMA modelling\[[3](https://arxiv.org/html/2609.23378#bib.bib1)\], and error\-correction / cointegration models\[[6](https://arxiv.org/html/2609.23378#bib.bib12)\]formalise how integrated series are recombined\. The same idea reappears in deep forecasting as per\-window re\-centring: reversible instance normalisation\[[9](https://arxiv.org/html/2609.23378#bib.bib3)\], non\-stationary attention\[[12](https://arxiv.org/html/2609.23378#bib.bib13)\]and DLinear’s linear detrending\[[19](https://arxiv.org/html/2609.23378#bib.bib4)\]all forecast a de\-trended or differenced target, exactly the setting we analyse; the leaky reconstruction we adopt is itself a first\-order exponential smoother of the increment stream\[[7](https://arxiv.org/html/2609.23378#bib.bib16)\]\.

Long\-horizon architectures and the multi\-horizon strategy\.Most modern long\-sequence forecasters \(Informer\[[20](https://arxiv.org/html/2609.23378#bib.bib5)\], Autoformer\[[18](https://arxiv.org/html/2609.23378#bib.bib8)\], FEDformer\[[21](https://arxiv.org/html/2609.23378#bib.bib19)\], PatchTST\[[14](https://arxiv.org/html/2609.23378#bib.bib6)\]and iTransformer\[[11](https://arxiv.org/html/2609.23378#bib.bib7)\]\) emit allHHsteps*jointly*\(direct multi\-horizon\), and N\-BEATS\[[15](https://arxiv.org/html/2609.23378#bib.bib9)\]and its hierarchical successor N\-HiTS\[[4](https://arxiv.org/html/2609.23378#bib.bib20)\]use direct basis\-expansion heads, precisely to*avoid*recursion\. Whether to forecast directly or iterate one step at a time is a long\-standing question in the forecasting literature\[[13](https://arxiv.org/html/2609.23378#bib.bib14)\]\. Recursive rollout nonetheless remains the native mode for classical ARIMA, autoregressive probabilistic models such as DeepAR\[[16](https://arxiv.org/html/2609.23378#bib.bib10)\], and autoregressively sampled foundation models such as TimesFM\[[5](https://arxiv.org/html/2609.23378#bib.bib21)\], the regime in which our fix is most needed\.

Exposure bias and our position\.In recurrent sequence models\[[17](https://arxiv.org/html/2609.23378#bib.bib17)\]recursive error growth is known as exposure bias and is usually attacked at*training*time \(scheduled sampling\[[2](https://arxiv.org/html/2609.23378#bib.bib11)\], professor forcing\[[10](https://arxiv.org/html/2609.23378#bib.bib18)\]and probabilistic rollout\[[16](https://arxiv.org/html/2609.23378#bib.bib10)\]\), while its bias–variance trade\-off has been studied for multistep forecasting\[[1](https://arxiv.org/html/2609.23378#bib.bib15)\]\. In contrast, we contribute \(i\) a signal\-processing account that identifies recursive differenced reconstruction as a*unit\-pole integrator*, whose error accumulates ask\\sqrt\{k\}and, for a biased learned model, diverges; and \(ii\) a*training\-free*, single\-scalar remedy by pole placement, with a closed\-form variance bound, that rescues every diverging architecture at no training cost\. To our knowledge the leaky\-integrator reconstruction and its bias–variance analysis are new to time\-series forecasting\.

## 8Conclusion

We identified the root cause of long\-horizon failure in differenced forecasting: its reconstruction is a discrete integrator with a pole on the unit circle, so a learned model’s biased increment errors are integrated without bound, diverging to several times the irreduciblek\\sqrt\{k\}rate \(up to3\.83\.8normalised MAE\)\. Our contribution turns this diagnosis into a cure: placing the pole*inside*the unit circle \(a leaky integrator with one scalarγ\\gamma\) provably bounds the accumulated variance, and applied post hoc with a fixedγ=0\.9\\gamma\{=\}0\.9and no retraining it cuts336336\-step error by4343–74%74\\%on the seven diverging architectures over twenty datasets, and is*principled*,*general*, and*safe*by design\.

## References

- \[1\]S\. Ben Taieb and A\. F\. Atiya\(2016\)A bias and variance analysis for multistep\-ahead time series forecasting\.IEEE Trans\. on Neural Networks and Learning Systems27\(1\),pp\. 62–76\.Cited by:[§7](https://arxiv.org/html/2609.23378#S7.p3.1)\.
- \[2\]S\. Bengio, O\. Vinyals, N\. Jaitly, and N\. Shazeer\(2015\)Scheduled sampling for sequence prediction with recurrent neural networks\.InAdvances in Neural Information Processing Systems \(NeurIPS\),pp\. 1171–1179\.Cited by:[§4](https://arxiv.org/html/2609.23378#S4.p2.1),[§7](https://arxiv.org/html/2609.23378#S7.p3.1)\.
- \[3\]G\. E\. P\. Box and G\. M\. Jenkins\(1970\)Time series analysis: forecasting and control\.Holden\-Day\.Cited by:[§1](https://arxiv.org/html/2609.23378#S1.p1.1),[§7](https://arxiv.org/html/2609.23378#S7.p1.1)\.
- \[4\]C\. Challu, K\. G\. Olivares, B\. N\. Oreshkin, F\. Garza, M\. Mergenthaler\-Canseco, and A\. Dubrawski\(2023\)N\-HiTS: neural hierarchical interpolation for time series forecasting\.InAAAI Conf\. on Artificial Intelligence,Cited by:[§7](https://arxiv.org/html/2609.23378#S7.p2.1)\.
- \[5\]A\. Das, W\. Kong, R\. Sen, and Y\. Zhou\(2024\)A decoder\-only foundation model for time\-series forecasting\.InInt\. Conf\. on Machine Learning \(ICML\),Cited by:[§7](https://arxiv.org/html/2609.23378#S7.p2.1)\.
- \[6\]R\. F\. Engle and C\. W\. J\. Granger\(1987\)Co\-integration and error correction: representation, estimation, and testing\.Econometrica55\(2\),pp\. 251–276\.Cited by:[§1](https://arxiv.org/html/2609.23378#S1.p1.1),[§7](https://arxiv.org/html/2609.23378#S7.p1.1)\.
- \[7\]E\. S\. Gardner Jr\.\(2006\)Exponential smoothing: the state of the art—part II\.Int\. J\. of Forecasting22\(4\),pp\. 637–666\.Cited by:[§7](https://arxiv.org/html/2609.23378#S7.p1.1)\.
- \[8\]S\. Hochreiter and J\. Schmidhuber\(1997\)Long short\-term memory\.Neural Computation9\(8\),pp\. 1735–1780\.Cited by:[§4](https://arxiv.org/html/2609.23378#S4.p2.1)\.
- \[9\]T\. Kim, J\. Kim, Y\. Tae, C\. Park, J\. Choi, and J\. Choo\(2022\)Reversible instance normalization for accurate time\-series forecasting against distribution shift\.InInt\. Conf\. on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2609.23378#S1.p1.1),[§7](https://arxiv.org/html/2609.23378#S7.p1.1)\.
- \[10\]A\. Lamb, A\. Goyal, Y\. Zhang, S\. Zhang, A\. Courville, and Y\. Bengio\(2016\)Professor forcing: a new algorithm for training recurrent networks\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§7](https://arxiv.org/html/2609.23378#S7.p3.1)\.
- \[11\]Y\. Liu, T\. Hu, H\. Zhang, H\. Wu, S\. Wang, L\. Ma, and M\. Long\(2024\)iTransformer: inverted transformers are effective for time series forecasting\.InInt\. Conf\. on Learning Representations \(ICLR\),Cited by:[§4](https://arxiv.org/html/2609.23378#S4.p2.1),[§7](https://arxiv.org/html/2609.23378#S7.p2.1)\.
- \[12\]Y\. Liu, H\. Wu, J\. Wang, and M\. Long\(2022\)Non\-stationary transformers: exploring the stationarity in time series forecasting\.InAdvances in Neural Information Processing Systems \(NeurIPS\),pp\. 9881–9893\.Cited by:[§7](https://arxiv.org/html/2609.23378#S7.p1.1)\.
- \[13\]M\. Marcellino, J\. H\. Stock, and M\. W\. Watson\(2006\)A comparison of direct and iterated multistep AR methods for forecasting macroeconomic time series\.Journal of Econometrics135\(1–2\),pp\. 499–526\.Cited by:[§7](https://arxiv.org/html/2609.23378#S7.p2.1)\.
- \[14\]Y\. Nie, N\. H\. Nguyen, P\. Sinthong, and J\. Kalagnanam\(2023\)A time series is worth 64 words: long\-term forecasting with transformers\.InInt\. Conf\. on Learning Representations \(ICLR\),Cited by:[§4](https://arxiv.org/html/2609.23378#S4.p2.1),[§7](https://arxiv.org/html/2609.23378#S7.p2.1)\.
- \[15\]B\. N\. Oreshkin, D\. Carpov, N\. Chapados, and Y\. Bengio\(2020\)N\-BEATS: neural basis expansion analysis for interpretable time series forecasting\.InInt\. Conf\. on Learning Representations \(ICLR\),Cited by:[§7](https://arxiv.org/html/2609.23378#S7.p2.1)\.
- \[16\]D\. Salinas, V\. Flunkert, J\. Gasthaus, and T\. Januschowski\(2020\)DeepAR: probabilistic forecasting with autoregressive recurrent networks\.Int\. J\. of Forecasting36\(3\),pp\. 1181–1191\.Cited by:[§1](https://arxiv.org/html/2609.23378#S1.p1.1),[§7](https://arxiv.org/html/2609.23378#S7.p2.1),[§7](https://arxiv.org/html/2609.23378#S7.p3.1)\.
- \[17\]I\. Sutskever, O\. Vinyals, and Q\. V\. Le\(2014\)Sequence to sequence learning with neural networks\.InAdvances in Neural Information Processing Systems \(NeurIPS\),pp\. 3104–3112\.Cited by:[§7](https://arxiv.org/html/2609.23378#S7.p3.1)\.
- \[18\]H\. Wu, J\. Xu, J\. Wang, and M\. Long\(2021\)Autoformer: decomposition transformers with auto\-correlation for long\-term series forecasting\.InAdvances in Neural Information Processing Systems \(NeurIPS\),pp\. 22419–22430\.Cited by:[§7](https://arxiv.org/html/2609.23378#S7.p2.1)\.
- \[19\]A\. Zeng, M\. Chen, L\. Zhang, and Q\. Xu\(2023\)Are transformers effective for time series forecasting?\.InAAAI Conf\. on Artificial Intelligence,pp\. 11121–11128\.Cited by:[§1](https://arxiv.org/html/2609.23378#S1.p1.1),[§4](https://arxiv.org/html/2609.23378#S4.p2.1),[§7](https://arxiv.org/html/2609.23378#S7.p1.1)\.
- \[20\]H\. Zhou, S\. Zhang, J\. Peng, S\. Zhang, J\. Li, H\. Xiong, and W\. Zhang\(2021\)Informer: beyond efficient transformer for long sequence time\-series forecasting\.InAAAI Conf\. on Artificial Intelligence,pp\. 11106–11115\.Cited by:[§7](https://arxiv.org/html/2609.23378#S7.p2.1)\.
- \[21\]T\. Zhou, Z\. Ma, Q\. Wen, X\. Wang, L\. Sun, and R\. Jin\(2022\)FEDformer: frequency enhanced decomposed transformer for long\-term series forecasting\.InInt\. Conf\. on Machine Learning \(ICML\),Cited by:[§7](https://arxiv.org/html/2609.23378#S7.p2.1)\.

相似文章