AdaST: Adaptive Coupling for Spatial-Temporal Forecasting

arXiv cs.AI Papers

Summary

AdaST introduces an adaptive spatial-temporal forecasting framework that decomposes inputs into components with distinct coupling regimes using heterogeneity-aware experts, then recomposes them with a correlation-informed module to outperform state-of-the-art baselines on traffic, climate, and energy forecasting.

arXiv:2609.36119v1 Announce Type: new Abstract: Spatial-temporal (ST) forecasting underpins many real-world systems such as traffic, climate, and energy networks. While existing methods implicitly assume strong spatiotemporal coupling, we observe that real-world ST data exhibits distinct coupling regimes, ranging from temporal-dominated and spatial-dominated to strongly coupled patterns. This mismatch causes current models to suffer from spurious dependencies and degraded performance when one correlation dominates. To overcome this limitation, we aim to dynamically modulate spatial and temporal modeling based on the data's inherent coupling structure. However, three key challenges exist: unknown coupling structure, heterogeneous coupling dynamics, and suboptimal spatial modeling. We propose AdaST, an adaptive ST forecasting framework that tackles these challenges through a decompose-recompose paradigm. AdaST factorizes inputs into components capturing different coupling patterns using heterogeneity-aware experts. Each component is processed by role-aligned modules, and a correlation-informed adaptive recomposer integrates them for final prediction. Extensive experiments confirm that AdaST significantly outperforms state-of-the-art baselines, validating the necessity of an adaptive approach.
Original Article
View Cached Full Text

Cached at: 09/30/26, 09:41 AM

# AdaST: Adaptive Coupling for Spatial-Temporal Forecasting
Source: [https://arxiv.org/html/2609.36119](https://arxiv.org/html/2609.36119)
Zhenyu LeiAffiliation:University of VirginiaAffiliation:Charlottesville, VA, USAEmail:[vjd5zr@virginia\.edu](mailto:)Chenghao Liu††thanks:This work was completed prior to joining Datadog\.Yushun DongAffiliation:Florida State UniversityAffiliation:Tallahassee, FL, USAEmail:[yd24f@fsu\.edu](mailto:)Qi R\. WangAffiliation:Northeastern UniversityAffiliation:Boston, MA, USAEmail:[q\.wang@northeastern\.edu](mailto:)Jundong LiAffiliation:University of VirginiaAffiliation:Charlottesville, VA, USAEmail:[jundong@virginia\.edu](mailto:)

###### Abstract

Spatial\-temporal \(ST\) forecasting underpins many real\-world systems such as traffic, climate, and energy networks\. While existing methods implicitly assume strong spatiotemporal coupling, we observe that real\-world ST data exhibits distinct coupling regimes, ranging from temporal\-dominated and spatial\-dominated to strongly coupled patterns\. This mismatch causes current models to suffer from spurious dependencies and degraded performance when one correlation dominates\. To overcome this limitation, we aim to dynamically modulate spatial and temporal modeling based on the data’s inherent coupling structure\. However, three key challenges exist: unknown coupling structure, heterogeneous coupling dynamics, and suboptimal spatial modeling\. We propose AdaST, an adaptive ST forecasting framework that tackles these challenges through a decompose\-recompose paradigm\. AdaST factorizes inputs into components capturing different coupling patterns using heterogeneity\-aware experts\. Each component is processed by role\-aligned modules, and a correlation\-informed adaptive recomposer integrates them for final prediction\. Extensive experiments confirm that AdaST significantly outperforms state\-of\-the\-art baselines, validating the necessity of an adaptive approach\.

## 1Introduction

Spatial\-temporal data, which incorporates both spatial and temporal information, plays a critical role in numerous real\-world applications[Wang et al\. \(2020\)](https://arxiv.org/html/2609.36119#bib.bib22);[Atluri et al\. \(2018\)](https://arxiv.org/html/2609.36119#bib.bib23);[Birant and Kut \(2007\)](https://arxiv.org/html/2609.36119#bib.bib24);[Li et al\. \(2025\)](https://arxiv.org/html/2609.36119#bib.bib46)\. Forecasting based on such data has proven invaluable across diverse domains by leveraging the inherent correlations among data points distributed across space and time[Li and Zhu \(2021\)](https://arxiv.org/html/2609.36119#bib.bib25);[Lei et al\. \(2025\)](https://arxiv.org/html/2609.36119#bib.bib47);[Guo et al\. \(2021\)](https://arxiv.org/html/2609.36119#bib.bib26)\. Spatial\-temporal data is characterized by two fundamental correlations: spatial correlation, describing interactions among different locations, and temporal correlation, capturing how historical observations influence future states[Vuran et al\. \(2004\)](https://arxiv.org/html/2609.36119#bib.bib27)\. These correlations are often intertwined, giving rise to complex spatial\-temporal dependencies that complicate forecasting[Shao et al\. \(2022b\)](https://arxiv.org/html/2609.36119#bib.bib12);[Yi et al\. \(2024\)](https://arxiv.org/html/2609.36119#bib.bib28)\. To capture such dependencies, existing research has proposed various spatiotemporal coupling frameworks\. Representative approaches include graph neural networks[Zhao et al\. \(2026\)](https://arxiv.org/html/2609.36119#bib.bib48)combined with temporal convolutional networks[Diao et al\. \(2019\)](https://arxiv.org/html/2609.36119#bib.bib29), attention\-based architectures[Zhang et al\. \(2019\)](https://arxiv.org/html/2609.36119#bib.bib30), and diffusion\-based frameworks[Yang et al\. \(2024\)](https://arxiv.org/html/2609.36119#bib.bib31)\. Despite architectural differences, these approaches share a common design philosophy: spatial and temporal correlations are jointly modeled at each layer, implicitly assuming strong and homogeneous coupling throughout the data\.

However, real\-world ST data exhibits substantially different coupling structures across domains and scenarios[He et al\. \(2017\)](https://arxiv.org/html/2609.36119#bib.bib32)\. Through systematic analysis, we categorize spatial\-temporal data into three regimes based on coupling strength\.\(1\)In weakly\-coupled temporal\-dominated scenarios, historical patterns at each location independently govern future values with minimal cross\-location influence\. For example, household energy consumption depends primarily on its own historical habits rather than neighboring behaviors[Pierce et al\. \(2010\)](https://arxiv.org/html/2609.36119#bib.bib38)\.\(2\)In weakly\-coupled spatial\-dominated scenarios, neighboring locations in short time windows exert primary influence while historical trends become less relevant\. For example, traffic incidents immediately trigger downstream congestion regardless of historical patterns[Qi et al\. \(2018\)](https://arxiv.org/html/2609.36119#bib.bib39)\.\(3\)In strongly\-coupled scenarios, both correlations are tightly intertwined and predictions require balanced integration of temporal evolution and spatial diffusion\. For example, urban air quality prediction requires balancing local historical emissions and wind\-driven diffusion from neighbors[Wang and Song \(2018\)](https://arxiv.org/html/2609.36119#bib.bib40)\. Uniformly applying coupled frameworks across these structures leads to systematic performance degradation\. For temporal\-dominated data, forced spatial coupling introduces spurious cross\-location dependencies\. For spatial\-dominated data, enforced temporal coupling incorporates misleading historical correlations\. Even for strongly\-coupled data, fixed architectures cannot adapt to heterogeneous coupling strengths across regions or time periods\. Our preliminary experiments in Section[2\.3](https://arxiv.org/html/2609.36119#S2.SS3)further validate that mismatches between coupling structures and model assumptions significantly impair forecasting accuracy\.

These observations motivate a fundamental question: Can we develop adaptive forecasting frameworks that dynamically modulate spatial and temporal modeling according to the data’s inherent coupling structure? Nevertheless, addressing this question is challenging for three reasons\.\(1\) Unknown Coupling Structure\.The intrinsic coupling structure of a dataset is typically unknown a priori, where practitioners must rely on trial\-and\-error or domain expertise to select suitable architectures\.\(2\) Heterogeneous Coupling Dynamics\.Coupling strength varies significantly across space and time within a single dataset\. Business streets exhibit strong spatial dependencies as congestion cascades upstream within minutes[Xiong et al\. \(2018\)](https://arxiv.org/html/2609.36119#bib.bib42), while highway interchanges follow predictable temporal patterns[Kerner \(2002\)](https://arxiv.org/html/2609.36119#bib.bib41)\. Similarly, rush hours demonstrate temporal regularity as traffic converges on major routes, whereas accidents trigger spatial propagation regardless of historical patterns\.\(3\) Suboptimal Spatial Modeling\.Effective coupling requires accurately modeling each correlation\. To capture the spatial topology, most existing approaches use either fixed pre\-defined graphs that cannot capture evolving relationships[Lan et al\. \(2022\)](https://arxiv.org/html/2609.36119#bib.bib33), or separately learned graphs that can be noisy when misestimated[Kang et al\. \(2019\)](https://arxiv.org/html/2609.36119#bib.bib34)\. Others employ attention over all location pairs, but but recomputing pairwise scores for every sample incurs substantial overhead, and spurious long\-range links harm robustness[Wang et al\. \(2022\)](https://arxiv.org/html/2609.36119#bib.bib35)\. These limitations demand more reliable spatial modeling\.

We proposeAdaST, a simple but effective framework that adaptively modulates spatial and temporal modeling according to data\-dependent coupling structures\. Rather than uniformly enforcing joint spatial\-temporal modeling, AdaST follows a*decompose–recompose*principle: disentangling different correlation patterns can help prevent confounding interactions in adaptive coupling, which might otherwise obscure each component’s contribution and induce negative transfer[Xia et al\. \(2023\)](https://arxiv.org/html/2609.36119#bib.bib36);[Deng et al\. \(2024\)](https://arxiv.org/html/2609.36119#bib.bib37)\. Specifically, AdaST employs a decomposer that factorizes the input into three components: \(i\) a temporal\-specific component capturing intra\-series dynamics, \(ii\) a spatial\-specific component capturing cross\-location dependencies, and \(iii\) a spatiotemporal\-coupling component capturing joint interactions\. Each component is then processed by a role\-aligned module – temporal\-only, spatial\-only, or joint spatiotemporal modeling, respectively\. Finally, an adaptive recomposer learns data\-dependent mixture weights to recombine these components for prediction, while a correlation\-informed modulation further reduces spurious correlations within each component\. This design allows the model to automatically allocate capacity: favoring temporal cues when history dominates, spatial cues when propagation dominates, and joint cues when both are salient\.

To accommodate heterogeneous coupling dynamics, AdaST employs heterogeneity\-aware experts that capture spatial and temporal variations through learned embeddings, guiding decomposition based on location\-specific and period\-specific characteristics\. For spatial modeling, a lightweight spatial mixer efficiently mixes features along the spatial dimension\. This design offers a balanced trade\-off between expressiveness and robustness, enabling global interaction while avoiding excessive computational overhead and unstable long\-range dependencies\. Across diverse benchmarks, AdaST consistently outperforms all baselines\. Further analysis demonstrates that AdaST provides excellent interpretability, making it a transparent and adaptable solution for various forecasting scenarios\.

## 2Preliminaries

### 2\.1Spatial\-Temporal Forecasting

Spatial\-temporal forecasting aims to predict future values across multiple locations based on historical observations that exhibit both spatial and temporal dependencies\. Formally, let𝒱\\mathcal\{V\}denote a set ofNNnodes representing spatial locations, and let𝐱i∈ℝT×D\\mathbf\{x\}\_\{i\}\\in\\mathbb\{R\}^\{T\\times D\}be the time series associated with each nodevi∈𝒱v\_\{i\}\\in\\mathcal\{V\}, whereTTis the number of historical time steps andDDis the feature dimensionality\. The complete spatial\-temporal data is represented as a tensor𝐗=\[𝐱1,…,𝐱N\]⊤∈ℝN×T×D\\mathbf\{X\}=\[\\mathbf\{x\}\_\{1\},\\ldots,\\mathbf\{x\}\_\{N\}\]^\{\\top\}\\in\\mathbb\{R\}^\{N\\times T\\times D\}\. The task is to learn a mapping that predicts future observations𝐘^∈ℝN×T′×D\\hat\{\\mathbf\{Y\}\}\\in\\mathbb\{R\}^\{N\\times T^\{\\prime\}\\times D\}over the nextT′T^\{\\prime\}time steps:

𝐘^=f⁡\(𝐗,𝐀∗\),\\hat\{\\mathbf\{Y\}\}=f\(\\mathbf\{X\};\\mathbf\{A\}^\{\*\}\),\(1\)where𝐀∗∈\{𝐀,∅\}\\mathbf\{A\}^\{\*\}\\in\\\{\\mathbf\{A\},\\emptyset\\\}denotes an optional adjacency matrix𝐀∈ℝN×N\\mathbf\{A\}\\in\\mathbb\{R\}^\{N\\times N\}encoding pre\-existing spatial connectivity when prior structural knowledge is available\.

### 2\.2Coupling Structure

We characterize how spatial and temporal dependencies interact by distinguishing coupling regimes\.

Weakly Coupled Temporal\-Dominated\.Future values at locationiiare governed primarily by its own history, with negligible cross\-location influence:

yit≈gtemp\(𝐱i1:t−1\),y\_\{i\}^\{t\}\\approx g\_\{\\text\{temp\}\}\\\!\\big\(\\mathbf\{x\}\_\{i\}^\{1:t\-1\}\\big\),\(2\)wheregtempg\_\{\\text\{temp\}\}models intra\-series dynamics\.

Weakly Coupled Spatial\-Dominated\.Future values are driven by short\-horizon propagation from other nodes, while long\-term temporal trends contribute minimally:

yit≈gspat​\(\{𝐱jt−1\}j∈𝒱\),y\_\{i\}^\{t\}\\approx g\_\{\\text\{spat\}\}\\\!\\big\(\\\{\\mathbf\{x\}\_\{j\}^\{t\-1\}\\\}\_\{j\\in\\mathcal\{V\}\}\\big\),\(3\)wheregspatg\_\{\\text\{spat\}\}aggregates other nodes’ most recent states\.

Strongly Coupled\.Spatial and temporal effects are tightly intertwined and must be modeled jointly:

yit≈gst\(𝐱i1:t−1,\{𝐱j1:t−1\}j∈𝒱\),y\_\{i\}^\{t\}\\approx g\_\{\\text\{st\}\}\\\!\\big\(\\mathbf\{x\}\_\{i\}^\{1:t\-1\},\\,\\\{\\mathbf\{x\}\_\{j\}^\{1:t\-1\}\\\}\_\{j\\in\\mathcal\{V\}\}\\big\),\(4\)wheregstg\_\{\\text\{st\}\}integrates histories across space and time to capture their mutual dependence\.

Different datasets exhibit different coupling regimes, but many existing methods apply a single, strongly\-coupled spatial\-temporal architecture across the board\. This*mismatch*between the model’s inductive bias and the data’s coupling pattern encourages the models to exploit incidental cross\-branch associations that are spurious and not generalizable, leading to suboptimal performance\.

Table 1:Average temporal and spatial correlation coefficient in three synthetic datasets\. Bold values indicate dominant correlations matching the intended coupling regime\.
### 2\.3Preliminary Experiments

To empirically validate the importance of architecture\-structure alignment, we conduct controlled experiments on synthetic datasets with known coupling structures\.

Synthetic Data Generation\.We construct three datasets corresponding to the three coupling structures:

- •Temporal\-Dominated \(TD\):For each time seriesii, future values are generated solely from its own history\. Each follows an order\-12 Auto\-Regressive process \(AR\(12\)\) with coefficients\{ϕi,k\}k=112\\\{\\phi\_\{i,k\}\\\}\_\{k=1\}^\{12\}and an exogenous multi\-frequency signalsi​\(t\)s\_\{i\}\(t\): yit=∑k=112ϕi,k​𝐱it−k\+0\.4​si​\(t\)\+ϵit,ϵit∼𝒩⁡\(0,0\.152\)\.y\_\{i\}^\{t\}=\\sum\_\{k=1\}^\{12\}\\phi\_\{i,k\}\\,\\mathbf\{x\}\_\{i\}^\{t\-k\}\+0\.4\\,s\_\{i\}\(t\)\+\\epsilon\_\{i\}^\{t\},\\quad\\epsilon\_\{i\}^\{t\}\\sim\\mathcal\{N\}\(0,0\.15^\{2\}\)\.
- •Spatial\-Dominated \(SD\):Next\-step values depend only on neighbors’ current states\. We build a directed sparse random graph𝐀\\mathbf\{A\}\(no self\-loops\) and ensure each node has at least one in\- and out\-edge\. LetP=Dout−1​𝐀P=D\_\{\\text\{out\}\}^\{\-1\}\\mathbf\{A\}whereDout=diag⁡\(𝐀𝟏\)D\_\{\\text\{out\}\}=\\mathrm\{diag\}\(\\mathbf\{A\}\\mathbf\{1\}\): yit=∑j=1NPj​i​𝐱jt−1\+ϵit,ϵit∼𝒩⁡\(0,0\.32\)\.y\_\{i\}^\{t\}=\\sum\_\{j=1\}^\{N\}P\_\{ji\}\\,\\mathbf\{x\}\_\{j\}^\{t\-1\}\+\\epsilon\_\{i\}^\{t\},\\quad\\epsilon\_\{i\}^\{t\}\\sim\\mathcal\{N\}\(0,0\.3^\{2\}\)\.
- •Strongly\-Coupled \(SC\):Future values depend on history and spatial neighbors, combining per\-node AR\(6\) with graph diffusion and a weak exogenous term withϵit∼𝒩⁡\(0,0\.152\)\\epsilon\_\{i\}^\{t\}\\sim\\mathcal\{N\}\(0,0\.15^\{2\}\): yit=∑k=16ϕi,k​𝐱it−k\+∑j=1NPj​i​𝐱jt−1\+0\.2​si​\(t\)\+ϵit\.\\displaystyle y\_\{i\}^\{t\}=\\sum\_\{k=1\}^\{6\}\\phi\_\{i,k\}\\,\\mathbf\{x\}\_\{i\}^\{t\-k\}\+\\sum^\{N\}\_\{j=1\}P\_\{ji\}\\,\\mathbf\{x\}\_\{j\}^\{t\-1\}\+0\.2\\,s\_\{i\}\(t\)\+\\epsilon\_\{i\}^\{t\}\.

We computed the correlation coefficients for the three synthetic datasets and observed that each synthetic dataset exhibits its intended coupling structure in Table[1](https://arxiv.org/html/2609.36119#S2.T1)\. Details are provided in Appendix[A](https://arxiv.org/html/2609.36119#A1)\.

![Refer to caption](https://arxiv.org/html/2609.36119v1/prelim.png)Figure 1:Performance of different architectures on synthetic datasets with known coupling structures\. Darker colors indicate lower normalized MAE\.Architecture Categorization\.We evaluate three architecture types with four backbone implementations:

- •Temporal\-Only \(T\)models intra\-series dependencies from historical values at each location independently, with no cross\-location information exchange\. y^it=ftemp\(𝐱i1:t−1\),\\hat\{y\}\_\{i\}^\{t\}=f\_\{\\text\{temp\}\}\\left\(\\mathbf\{x\}\_\{i\}^\{1:t\-1\}\\right\),\(5\)
- •Spatial\-Only \(S\)aggregates information from other nodes at each time step through spatial convolution or attention, with simply averaging along the temporal dimension to suppress temporal pattern extraction\. y^it=1T​∑t=1Tfspat​\(\{𝐱jt\}j∈𝒱,𝐱it\),\\hat\{y\}\_\{i\}^\{t\}=\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}f\_\{\\text\{spat\}\}\\left\(\\left\\\{\\mathbf\{x\}\_\{j\}^\{t\}\\right\\\}\_\{j\\in\\mathcal\{V\}\},\\mathbf\{x\}\_\{i\}^\{t\}\\right\),\(6\)
- •Spatial\-Temporal \(ST\)jointly models spatial and temporal correlations through coupled architectures\. y^it=ftemp​\(fspat​\(\{𝐱jt\}j∈𝒱,𝐱it\)\),\\hat\{y\}\_\{i\}^\{t\}=f\_\{\\text\{temp\}\}\\left\(f\_\{\\text\{spat\}\}\\left\(\\left\\\{\\mathbf\{x\}\_\{j\}^\{t\}\\right\\\}\_\{j\\in\\mathcal\{V\}\},\\mathbf\{x\}\_\{i\}^\{t\}\\right\)\\right\),\(7\)

For each architecture type, we implement four backbone variants: attention\-based, convolutional, MLP\-based, and GWNet[Wu et al\. \(2019\)](https://arxiv.org/html/2609.36119#bib.bib4), a representative spatial\-temporal forecasting framework\.

Experimental Results\.Figure[1](https://arxiv.org/html/2609.36119#S2.F1)presents the normalized Mean Absolute Error \(MAE\) across all architecture\-data combinations\. Each heatmap corresponds to one backbone type, where the x\-axis represents data types \(TD, SD, SC\) and the y\-axis represents architecture types \(T, S, ST\)\. Colors are row\-normalized for visualization clarity, with darker colors indicating better performance\. From the figure, we could observe that diagonal entries consistently exhibit the darkest colors across all backbones and off\-diagonal entries show significant performance degradation, demonstrating that architecture\-structure alignment yields optimal performance \(e\.g\., temporal\-only architectures excel on temporal\-dominated data\)\. In addition, more expressive backbones \(e\.g\., attention vs\. convolution\) exhibit larger performance gaps for mismatched architectures, suggesting that complex models are more susceptible to overfitting spurious correlations when the architecture does not align with the underlying coupling structure\. These findings motivate the need for adaptive frameworks that can dynamically modulate spatial and temporal modeling based on the inherent coupling structure of the data, rather than enforcing fixed architectural priors\.

## 3Methodology

In this section, we elaborate our methodology AdaST, which is a simple and effective framework for adaptively coupling spatial and temporal signals in forecasting\. AdaST is decomposed into three modules: \(1\) Heterogeneity\-Aware Decoupling, which utilizes several heterogeneity\-aware experts to provide different views for decoupling\. \(2\) Spatial\-Temporal Modeling, which models temporal correlation with attention and spatial correlation with spatial mixer framework\. \(3\) Correlation\-Regularized Recoupling, which adaptively recouples different components with gated neural network with an correlation\-aware regularization to remove spurious correlation\. The overview is in Figure[2](https://arxiv.org/html/2609.36119#S3.F2)\.

### 3\.1Heterogeneity\-Aware Decomposition

To disentangle confounding cross\-pattern interactions, we decompose the input spatial\-temporal data into three components: a temporal\-specific component𝐗\(t\)\\mathbf\{X\}^\{\(t\)\}, a spatial\-specific component𝐗\(s\)\\mathbf\{X\}^\{\(s\)\}, and a spatiotemporal\-coupling component𝐗\(s​t\)\\mathbf\{X\}^\{\(st\)\}\. Each component will then be processed by its corresponding role\-aligned modeling module\. However, coupling structures exhibit considerable heterogeneity across spatial locations and temporal periods\. A fixed, uniform decomposition mechanism would fail to capture such variations, as it imposes identical decomposition patterns regardless of local characteristics\. To address this limitation, we propose a heterogeneity\-aware expert architecture wherein multiple specialized experts provide distinct decomposition perspectives that adapt to the underlying heterogeneity in coupling structures\.

![Refer to caption](https://arxiv.org/html/2609.36119v1/overview.png)Figure 2:The overall framework of AdaST\. Different colors denote different components\.We introduce four expert types that provide heterogeneous decomposition perspectives: \(1\)Spatial Expertfor spatial heterogeneity, \(2\)Time\-of\-Dayand \(3\)Day\-of\-Week Expertsfor temporal heterogeneity at different granularities, and \(4\)Spatiotemporal Expertfor joint variations\. Each maintains learnable contextual embeddings𝐄n∈ℝN×Dn\\mathbf\{E\}\_\{n\}\\in\\mathbb\{R\}^\{N\\times D\_\{n\}\}for spatial,𝐄h∈ℝTday×Dh\\mathbf\{E\}\_\{h\}\\in\\mathbb\{R\}^\{T\_\{\\text\{day\}\}\\times D\_\{h\}\}and𝐄w∈ℝTweek×Dw\\mathbf\{E\}\_\{w\}\\in\\mathbb\{R\}^\{T\_\{\\text\{week\}\}\\times D\_\{w\}\}for temporal, and𝐄a∈ℝT×N×Da\\mathbf\{E\}\_\{a\}\\in\\mathbb\{R\}^\{T\\times N\\times D\_\{a\}\}for spatial\-temporal contexts, encoding time and location\-specific attributes\. Given input𝐗\\mathbf\{X\}, we first project it to an embedding space𝐇0=Linear​\(𝐗\)∈ℝT×N×D0\\mathbf\{H\}\_\{0\}=\\text\{Linear\}\(\\mathbf\{X\}\)\\in\\mathbb\{R\}^\{T\\times N\\times D\_\{0\}\}\. Each expertkkthen constructs contextualized representations by concatenating𝐇0\\mathbf\{H\}\_\{0\}with its embeddings\. For the spatial and spatiotemporal experts, we have:

𝐇n\\displaystyle\\mathbf\{H\}\_\{n\}=\[𝐇0∥𝐄~n\],𝐇a=\[𝐇0∥𝐄a\],\\displaystyle=\[\\mathbf\{H\}\_\{0\}\\\|\\tilde\{\\mathbf\{E\}\}\_\{n\}\],\\quad\\mathbf\{H\}\_\{a\}=\[\\mathbf\{H\}\_\{0\}\\\|\\mathbf\{E\}\_\{a\}\],\(8\)where𝐄~n\\tilde\{\\mathbf\{E\}\}\_\{n\}denotes𝐄n\\mathbf\{E\}\_\{n\}expanded along temporal dimensions\. For temporal experts, we first index embeddings by timestamp\. Let𝐭,𝐝∈ℝT×N\\mathbf\{t\},\\mathbf\{d\}\\in\\mathbb\{R\}^\{T\\times N\}denote time\-of\-day and day\-of\-week indices:

𝐇h\\displaystyle\\mathbf\{H\}\_\{h\}=\[𝐇0∥𝐄~h\[𝐭\]\],𝐇w=\[𝐇0∥𝐄~w\[𝐝\]\]\.\\displaystyle=\[\\mathbf\{H\}\_\{0\}\\\|\\tilde\{\\mathbf\{E\}\}\_\{h\}\[\\mathbf\{t\}\]\],\\quad\\mathbf\{H\}\_\{w\}=\[\\mathbf\{H\}\_\{0\}\\\|\\tilde\{\\mathbf\{E\}\}\_\{w\}\[\\mathbf\{d\}\]\]\.\(9\)where𝐄~h\\tilde\{\\mathbf\{E\}\}\_\{h\}and𝐄~w\\tilde\{\\mathbf\{E\}\}\_\{w\}are expanded along spatial dimensions\. Each expert then applies a projection headfk:ℝD0\+Dk→ℝ3/4​DHf\_\{k\}:\\mathbb\{R\}^\{D\_\{0\}\+D\_\{k\}\}\\rightarrow\\mathbb\{R\}^\{3/4D\_\{H\}\}to extract three components:

\[𝐙k\(s​t\);𝐙k\(t\);𝐙k\(s\)\]=fk​\(𝐇k\),\[\\mathbf\{Z\}\_\{k\}^\{\(st\)\};\\mathbf\{Z\}\_\{k\}^\{\(t\)\};\\mathbf\{Z\}\_\{k\}^\{\(s\)\}\]=f\_\{k\}\(\\mathbf\{H\}\_\{k\}\),\(10\)wherek∈\{n,h,w,a\}k\\in\\\{n,h,w,a\\\}and each𝐙k\(⋅\)∈ℝT×N×1/4​DH\\mathbf\{Z\}\_\{k\}^\{\(\\cdot\)\}\\in\\mathbb\{R\}^\{T\\times N\\times 1/4D\_\{H\}\}\. Finally, we integrate perspectives from all experts via concatenation for each component:

𝐇\(⋅\)=\[𝐙n\(⋅\)∥𝐙h\(⋅\)∥𝐙w\(⋅\)∥𝐙a\(⋅\)\]∈ℝT×N×DH,\\mathbf\{H\}^\{\(\\cdot\)\}=\[\\mathbf\{Z\}\_\{n\}^\{\(\\cdot\)\}\\\|\\mathbf\{Z\}\_\{h\}^\{\(\\cdot\)\}\\\|\\mathbf\{Z\}\_\{w\}^\{\(\\cdot\)\}\\\|\\mathbf\{Z\}\_\{a\}^\{\(\\cdot\)\}\]\\in\\mathbb\{R\}^\{T\\times N\\times D\_\{H\}\},\(11\)where\(⋅\)∈\{s​t,t,s\}\(\\cdot\)\\in\\\{st,t,s\\\}\. The resulting𝐇\(s​t\),𝐇\(t\),𝐇\(s\)\\mathbf\{H\}^\{\(st\)\},\\mathbf\{H\}^\{\(t\)\},\\mathbf\{H\}^\{\(s\)\}encode heterogeneity\-aware coupling patterns and serve as inputs to subsequent role\-aligned modeling modules\.

### 3\.2Spatial\-Temporal Modeling

After decomposition, each component is processed by role\-aligned modules specialized for its correlation pattern\. The temporal\-specific component𝐗\(t\)\\mathbf\{X\}^\{\(t\)\}is processed by temporal attention layers, the spatial\-specific component𝐗\(s\)\\mathbf\{X\}^\{\(s\)\}by spatial mixer layers, and the spatiotemporal\-coupling component𝐗\(s​t\)\\mathbf\{X\}^\{\(st\)\}by interleaved temporal and spatial operations\.

For temporal modeling, we employ multi\-head self\-attention to capture long\-range dependencies:

𝐇′=MHA​\(𝐇𝐖Q,𝐇𝐖K,𝐇𝐖V\),\\displaystyle\\mathbf\{H\}^\{\\prime\}=\\text\{MHA\}\(\\mathbf\{H\}\\mathbf\{W\}^\{Q\},\\mathbf\{H\}\\mathbf\{W\}^\{K\},\\mathbf\{H\}\\mathbf\{W\}^\{V\}\),\(12\)TempAttn​\(𝐇\)=FFN​\(LN​\(𝐇\+Dropout​\(𝐇′\)\)\),\\displaystyle\\text\{TempAttn\}\(\\mathbf\{H\}\)=\\text\{FFN\}\(\\text\{LN\}\(\\mathbf\{H\}\+\\text\{Dropout\}\(\\mathbf\{H\}^\{\\prime\}\)\)\),\(13\)where LN denotes layer normalization and FFN denotes a feed\-forward network\. For spatial modeling, we adopt a spatial mixer with learnable assignment matrix𝐀∈ℝN×N\\mathbf\{A\}\\in\\mathbb\{R\}^\{N\\times N\}:

SpatMix​\(𝐇\)=FFN​\(𝐇\+\(softmax​\(𝐀\)​𝐇⊤\)⊤\),\\text\{SpatMix\}\(\\mathbf\{H\}\)=\\text\{FFN\}\(\\mathbf\{H\}\+\(\\text\{softmax\}\(\\mathbf\{A\}\)\\mathbf\{H\}^\{\\top\}\)^\{\\top\}\),\(14\)where𝐇⊤∈ℝN×\(T×D\)\\mathbf\{H\}^\{\\top\}\\in\\mathbb\{R\}^\{N\\times\(T\\times D\)\}denotes the transpose along the spatial dimension\. This design captures all pairwise spatial interactions without a predefined graph and substantial computation overhead\.

We stackLLtemporal attention layers for𝐇\(t\)\\mathbf\{H\}^\{\(t\)\}andLLspatial mixer layers for𝐇\(s\)\\mathbf\{H\}^\{\(s\)\}\. For𝐇\(s​t\)\\mathbf\{H\}^\{\(st\)\}, we interleave temporal attention and spatial mixing in each of theLLlayers\. The final outputs are𝐇^\(t\),𝐇^\(s\),𝐇^\(s​t\)∈ℝT×N×DH\\hat\{\\mathbf\{H\}\}^\{\(t\)\},\\hat\{\\mathbf\{H\}\}^\{\(s\)\},\\hat\{\\mathbf\{H\}\}^\{\(st\)\}\\in\\mathbb\{R\}^\{T\\times N\\times D\_\{H\}\}\.

### 3\.3Correlation\-Informed Recomposition

After role\-aligned modeling, we obtain three processed components𝐇^\(t\),𝐇^\(s\),𝐇^\(s​t\)\\hat\{\\mathbf\{H\}\}^\{\(t\)\},\\hat\{\\mathbf\{H\}\}^\{\(s\)\},\\hat\{\\mathbf\{H\}\}^\{\(st\)\}that capture distinct correlation patterns\. To generate final predictions, we adaptively recombine these components through a gated aggregation mechanism that learns data\-dependent mixture weights based on component representations and their correlation strengths\.

We employ a lightweight gated neural network that computes adaptive weights for each component\. For each component\(⋅\)∈\{t,s,s​t\}\(\\cdot\)\\in\\\{t,s,st\\\}, the gate is computed as:

g\(⋅\)=σ⁡\(W\(⋅\)⋅𝐇^\(⋅\)\)∈ℝT×N,g^\{\(\\cdot\)\}=\\sigma\(W\_\{\(\\cdot\)\}\\cdot\\hat\{\\mathbf\{H\}\}^\{\(\\cdot\)\}\)\\in\\mathbb\{R\}^\{T\\times N\},\(15\)whereσ⁡\(⋅\)\\sigma\(\\cdot\)is the sigmoid function,W\(⋅\)∈ℝDH×1W\_\{\(\\cdot\)\}\\in\\mathbb\{R\}^\{D\_\{H\}\\times 1\}are learnable parameters\. This allows the model to automatically adjust component contributions based on input characteristics and learned patterns\.

To further enhance the gating mechanism, we incorporate correlation\-informed weights that reflect the strength of inherent correlations within each component\. The rationale is that when a component exhibits weak correlations, spurious noise dominates over meaningful signals, and excessive reliance on such components degrades predictions\. We use cosine similarity as our correlation measures\. For the temporal component, we measure along the temporal dimension:

𝒞t\(𝐇\(t\)\)=1T−1∑i=1T−1CosineSim\(𝐇^:,i,:,:\(t\),𝐇^:,i\+1,:,:\(t\)\),\\mathcal\{C\}\_\{t\}\(\\mathbf\{H\}^\{\(t\)\}\)=\\frac\{1\}\{T\-1\}\\sum\_\{i=1\}^\{T\-1\}\\text\{CosineSim\}\(\\hat\{\\mathbf\{H\}\}^\{\(t\)\}\_\{:,i,:,:\},\\hat\{\\mathbf\{H\}\}^\{\(t\)\}\_\{:,i\+1,:,:\}\),\(16\)and for the spatial component, along the spatial dimension:

𝒞s\(𝐇\(s\)\)=1N−1∑j=1N−1CosineSim\(𝐇^:,:,j,:\(s\),𝐇^:,:,j\+1,:\(s\)\)\.\\mathcal\{C\}\_\{s\}\(\\mathbf\{H\}^\{\(s\)\}\)=\\frac\{1\}\{N\-1\}\\sum\_\{j=1\}^\{N\-1\}\\text\{CosineSim\}\(\\hat\{\\mathbf\{H\}\}^\{\(s\)\}\_\{:,:,j,:\},\\hat\{\\mathbf\{H\}\}^\{\(s\)\}\_\{:,:,j\+1,:\}\)\.\(17\)For the spatiotemporal component, we average both:

𝒞s​t​\(𝐇\(s​t\)\)=12​\(𝒞s​\(𝐇\(s​t\)\)\+𝒞t​\(𝐇\(s​t\)\)\)\.\\mathcal\{C\}\_\{st\}\(\\mathbf\{H\}^\{\(st\)\}\)=\\frac\{1\}\{2\}\\big\(\\mathcal\{C\}\_\{s\}\(\\mathbf\{H\}^\{\(st\)\}\)\+\\mathcal\{C\}\_\{t\}\(\\mathbf\{H\}^\{\(st\)\}\)\\big\)\.\(18\)We then modulate the gate scores by correlation measures:

g~\(⋅\)=g\(⋅\)∗\[𝒞\(⋅\)​\(𝐇\(⋅\)\)\]α,for​\(⋅\)∈\{t,s,s​t\},\\tilde\{g\}^\{\(\\cdot\)\}=g^\{\(\\cdot\)\}\*\[\\mathcal\{C\}\_\{\(\\cdot\)\}\(\\mathbf\{H\}^\{\(\\cdot\)\}\)\]^\{\\alpha\},\\quad\\text\{for \}\(\\cdot\)\\in\\\{t,s,st\\\},\(19\)whereα\\alphacontrols the influence of correlation strength\. The final weights are obtained via softmax:

w\(t\),w\(s\),w\(s​t\)=softmax​\(g~\(t\),g~\(s\),g~\(s​t\)\),w^\{\(t\)\},w^\{\(s\)\},w^\{\(st\)\}=\\text\{softmax\}\(\\tilde\{g\}^\{\(t\)\},\\tilde\{g\}^\{\(s\)\},\\tilde\{g\}^\{\(st\)\}\),\(20\)and the recomposed representation is:

𝐇^=w\(t\)⊙𝐇^\(t\)\+w\(s\)⊙𝐇^\(s\)\+w\(s​t\)⊙𝐇^\(s​t\),\\hat\{\\mathbf\{H\}\}=w^\{\(t\)\}\\odot\\hat\{\\mathbf\{H\}\}^\{\(t\)\}\+w^\{\(s\)\}\\odot\\hat\{\\mathbf\{H\}\}^\{\(s\)\}\+w^\{\(st\)\}\\odot\\hat\{\\mathbf\{H\}\}^\{\(st\)\},\(21\)where⊙\\odotdenotes element\-wise multiplication with broadcasting\. Finally, we apply a temporal projection followed by an output layer to generate predictions𝐘^\\hat\{\\mathbf\{Y\}\}\.

## 4Experiments

We empirically evaluate AdaST and organize this section around five research questions:RQ1\.How does AdaST compare with state\-of\-the\-art baselines across diverse benchmarks?RQ2\.What is the contribution of each module to overall component?RQ3\.How do different datasets manifest distinct coupling structures?RQ4\.How does the coupling structure evolve over time and vary across locations?RQ5\.How well are coupling components disentangled?

Table 2:Main results on PurpleAir and PEMS benchmarks\. The best and second\-best scores are highlighted inboldandunderlined\. AdaST achieves the best performance across all datasets\.### 4\.1Experimental Settings

Datasets\.We evaluate AdaST on four spatial\-temporal forecasting benchmarks with diverse coupling characteristics\. PEMS04, PEMS07, and PEMS08 are traffic flow datasets from the California Transportation Performance Management System, sampled at 5\-minute intervals \(Tday=288T\_\{\\text\{day\}\}=288time steps per day\)[Song et al\. \(2020\)](https://arxiv.org/html/2609.36119#bib.bib17)\. PurpleAir contains PM2\.5 measurements from air quality sensors in the Boston area\. We preprocess the raw 2\-minute data by resampling to 6\-minute intervals to mitigate noise and missing values \(Tday=240T\_\{\\text\{day\}\}=240time steps per day\)\. All datasets use weekly periodicity withTweek=7T\_\{\\text\{week\}\}=7\. Following standard practice, we apply Z\-score normalization and split each dataset into 60% training, 20% validation, and 20% testing\.

Baselines\.We compare AdaST against 16 representative methods\. Temporal\-only baselines include HI[Cui et al\. \(2021\)](https://arxiv.org/html/2609.36119#bib.bib1), DeepAR[Salinas et al\. \(2020\)](https://arxiv.org/html/2609.36119#bib.bib2), and NBeats[Oreshkin et al\. \(2019\)](https://arxiv.org/html/2609.36119#bib.bib3)\. Spatial\-temporal baselines include GWNet[Wu et al\. \(2019\)](https://arxiv.org/html/2609.36119#bib.bib4), DCRNN[Li et al\. \(2017\)](https://arxiv.org/html/2609.36119#bib.bib5), AGCRN[Bai et al\. \(2020\)](https://arxiv.org/html/2609.36119#bib.bib6), STGCN[Yu et al\. \(2017\)](https://arxiv.org/html/2609.36119#bib.bib7), GTS[Shang et al\. \(2021\)](https://arxiv.org/html/2609.36119#bib.bib8), MTGNN[Wu et al\. \(2020\)](https://arxiv.org/html/2609.36119#bib.bib9), StemGNN[Cao et al\. \(2020\)](https://arxiv.org/html/2609.36119#bib.bib10), STNorm[Deng et al\. \(2021\)](https://arxiv.org/html/2609.36119#bib.bib11), D2STGNN[Shao et al\. \(2022b\)](https://arxiv.org/html/2609.36119#bib.bib12), STID[Shao et al\. \(2022a\)](https://arxiv.org/html/2609.36119#bib.bib13), HimNet[Dong et al\. \(2024\)](https://arxiv.org/html/2609.36119#bib.bib14), STAEformer[Liu et al\. \(2023\)](https://arxiv.org/html/2609.36119#bib.bib15), and STDN[Cao et al\. \(2025\)](https://arxiv.org/html/2609.36119#bib.bib16)\.

Implementation Details\.We set expert embedding dimensions toDn=Dh=Dw=Da=24D\_\{n\}=D\_\{h\}=D\_\{w\}=D\_\{a\}=24and hidden dimension toDH=256D\_\{H\}=256\. The model usesL=3L=3layers with 4\-head temporal attention\. Both input and prediction lengths are 12 time steps \(T=T′=12T=T^\{\\prime\}=12\)\. We train with the Adam optimizer \(initial learning rate 0\.001, exponential decay, batch size 16\) and set the correlation weight toα=0\.1\\alpha=0\.1\. All experiments are conducted on an NVIDIA A100 80GB GPU using official baseline implementations with recommended hyperparameters\. Our code is available at[https://github\.com/LzyFischer/AdaST](https://github.com/LzyFischer/AdaST)\.

Evaluation Metrics\.We use three standard metrics: Mean Absolute Error \(MAE\), Root Mean Square Error \(RMSE\), and Mean Absolute Percentage Error \(MAPE\)\. Following convention, we report average performance across all 12 prediction horizons\.

### 4\.2Main Results

To answerRQ1, we present the comprehensive comparison results in Table[2](https://arxiv.org/html/2609.36119#S4.T2)\. We make the following key observations: \(1\) AdaST achieves state\-of\-the\-art results across all four benchmarks, consistently outperforming the strongest baselines\. Compared to the best\-performing competitors, AdaST improves MAE by up to1\.7%,1\.7\\%,RMSE by0\.7%0\.7\\%, and MAPE by0\.8%0\.8\\%on average, demonstrating the effectiveness of adaptive coupling mechanisms for spatial\-temporal forecasting\. \(2\) Spatial\-temporal methods generally outperform temporal\-only baselines, confirming that explicitly modeling spatial dependencies is crucial for forecasting tasks where cross\-location interactions exist\. However, the performance gap varies significantly across datasets, suggesting heterogeneous coupling structures\. \(3\) AdaST’s improvement is most pronounced on PurpleAir, where it outperforms the best baseline by4\.3%4\.3\\%in MAE\. This larger gain can be attributed to PurpleAir’s temporal\-dominated coupling \(as shown in Section[4\.4](https://arxiv.org/html/2609.36119#S4.SS4)\), where traditional spatial\-temporal methods introduce spurious spatial correlations\. AdaST’s adaptive mechanism effectively suppresses unnecessary spatial modeling through lower spatial gate scores, avoiding spurious dependencies while retaining beneficial signals\.

Table 3:Ablation study on PurpleAir and PEMS07\. Removing each component results in performance drop\.
### 4\.3Ablation Study

To answerRQ2, we systematically ablate key modules to evaluate their individual contributions on PurpleAir and PEMS07\. Table[3](https://arxiv.org/html/2609.36119#S4.T3)presents results for the following variants: removing individual heterogeneity experts \(w/o𝐄𝐧\\mathbf\{E\_\{n\}\},w/o𝐄𝐡\\mathbf\{E\_\{h\}\},w/o𝐄𝐰\\mathbf\{E\_\{w\}\},w/o𝐄𝐚\\mathbf\{E\_\{a\}\}\), replacing spatial mixer with spatial attention \(SpaAtt\), removing learned gate scores by using uniform averaging \(w/ogg\), and removing correlation measures from recomposition \(w/o𝒞\\mathcal\{C\}\)\. \(1\) All four experts contribute to performance, though their relative importance varies across datasets\. On PurpleAir, removing the spatial expert𝐄𝐧\\mathbf\{E\_\{n\}\}causes the largest degradation, indicating that location\-specific coupling patterns are most critical for air quality data where region characteristics differ substantially\. Conversely, on PEMS07, removing the time\-of\-day expert𝐄𝐡\\mathbf\{E\_\{h\}\}results in the greatest performance drop, reflecting the importance of diurnal traffic patterns\. \(2\) Replacing spatial mixer with spatial attention \(SpaAtt\) degrades performance, particularly on temporal\-dominated PurpleAir\. This suggests that attention can introduce spurious long\-range correlations when spatial coupling is weak\. \(3\) Both the learned gate scores and correlation measures are essential for effective recomposition\. Removing gates eliminates dataset\-adaptive weighting, while removing correlation measures allows spurious correlations to mislead recomposition\. Together, these mechanisms enable AdaST to dynamically balance components based on their reliability and relevance\.

![Refer to caption](https://arxiv.org/html/2609.36119v1/gate_scatter.png)Figure 3:Average gate scores of different datasets, revealing data\-specific coupling structures\.
### 4\.4Coupling Structure Analysis

To addressRQ3, we investigate how coupling structures vary across common spatial\-temporal datasets by visualizing the averaged gate scoresw\(t\)w^\{\(t\)\},w\(s\)w^\{\(s\)\}, andw\(s​t\)w^\{\(st\)\}in Figure[3](https://arxiv.org/html/2609.36119#S4.F3)\. We have the following observations: \(1\) Different datasets exhibit distinct coupling structures, confirming that a one\-size\-fits\-all architectural approach is suboptimal\. \(2\) Traditional spatial\-temporal forecasting benchmarks PEMS03\-PEMS08 exhibit relatively spatial\-dominant patterns\. This reflects the physical reality of traffic networks where congestion propagates spatially through road connections\. \(3\) PurpleAir and BeijingAirQuality demonstrate higher temporal gate scores, which can be attributed to the localized nature of air quality measurements, where sensor readings are primarily governed by local emission sources and meteorological conditions rather than immediate spatial diffusion from neighboring sensors\. This observation aligns with our main results \(Table[2](https://arxiv.org/html/2609.36119#S4.T2)\), where the temporal\-only baseline NBeats achieves the second\-best performance\. \(4\) For the single\-location multivariate ETTh and ETTm datasets, we repurpose the spatial module to model feature dimensions as pseudo\-spatial nodes\. We observe significant cross\-feature coupling, evidenced by spatial gate scores comparable in magnitude to temporal ones\. This confirms that inter\-variate interactions provide critical predictive signals, aligning with findings on these datasets[Grigsby et al\. \(2021\)](https://arxiv.org/html/2609.36119#bib.bib21)\. \(5\) METR\-LA and PEMS\-BAY exhibit stronger temporal dominance compared to PEMS03\-PEMS08\. This difference may reflect variations in road network characteristics: highway systems \(PEMS\-BAY covers the Bay Area highway network\) often exhibit more predictable temporal patterns driven by commuting schedules, while urban road networks \(PEMS from various California districts\) demonstrate stronger spatial propagation due to higher intersection density and more complex traffic interactions\.

### 4\.5Case Study

To addressRQ4, we present a case study on PEMS07 visualizing the learned gate scores across different time periods and locations in Figure[4](https://arxiv.org/html/2609.36119#S4.F4)\. We observe four key patterns\. \(1\) Gate scores are relatively stable across the whole datasets, which demonstrate the whole datasets share similar coupling structure\. \(2\) All gate scores exhibit clear periodic patterns that align with the underlying data periodicity,

Figure 4:Gate scores across different time periods and locations, illustrating dynamic coupling\.demonstrating that coupling structures are temporally heterogeneous and recurrent\. \(3\) Different spatial locations exhibit distinct coupling structures, validating the use of spatial experts that can adaptively capture location\-specific dependencies\. \(4\) Temporal gate scores peak during stable, gradual changes, while spatial gate scores spike during abrupt trend shifts\. Notably, when temporal and spatial gate scores are more comparable in magnitude, the spatial\-temporal gate score increases, indicating strong coupling\.

Figure 5:T\-SNE visualization of learned representations for temporal \(T\), spatial \(S\), and spatial\-temporal\-coupling \(ST\) components, demonstrating clear separation and effective disentanglement\.
### 4\.6Representation Analysis

To addressRQ5, we visualize the learned representations of the three components𝐇^\(t\)\\hat\{\\mathbf\{H\}\}^\{\(t\)\},𝐇^\(s\)\\hat\{\\mathbf\{H\}\}^\{\(s\)\}, and𝐇^\(s​t\)\\hat\{\\mathbf\{H\}\}^\{\(st\)\}before adaptive recomposition\. We apply t\-SNE dimensionality reduction to project the embeddings into 2D space, as shown in Figure[5](https://arxiv.org/html/2609.36119#S4.F5)\. The results reveal clear separation between the three component clusters, which demonstrates that our heterogeneity\-aware decomposition and role\-aligned modeling successfully enforce specialization, thereby providing a reliable foundation for adaptive recomposition\.

## 5Related Works

Spatial\-temporal data encodes both temporal dynamics and spatial dependencies across multiple locations\.Temporal\-Only Approaches\.Early forecasting methods such as ARIMA[Contreras et al\. \(2003\)](https://arxiv.org/html/2609.36119#bib.bib18), NBeats[Oreshkin et al\. \(2019\)](https://arxiv.org/html/2609.36119#bib.bib3), and DeepAR[Salinas et al\. \(2020\)](https://arxiv.org/html/2609.36119#bib.bib2)focused on temporal modeling, treating each time series independently\. While effective for intra\-series dependencies, they fail to capture critical spatial correlations\.Coupled Spatial\-Temporal Models\.To capture spatial dependencies, recent works explicitly couple temporal and spatial modeling\. Graph\-based methods leverage GNNs to model spatial correlations: DCRNN[Li et al\. \(2017\)](https://arxiv.org/html/2609.36119#bib.bib5)combines diffusion convolution with GRUs, STGCN[Yu et al\. \(2017\)](https://arxiv.org/html/2609.36119#bib.bib7)stacks temporal and spatial convolutions, and GWNet[Wu et al\. \(2019\)](https://arxiv.org/html/2609.36119#bib.bib4)introduces adaptive adjacency matrices\. More recently, attention\-based architectures[Guo et al\. \(2019\)](https://arxiv.org/html/2609.36119#bib.bib19)have achieved state\-of\-the\-art performance by capturing flexible, data\-driven correlations\. GMAN[Zheng et al\. \(2020\)](https://arxiv.org/html/2609.36119#bib.bib20)employs spatial\-temporal attention for long\-range dependencies, while STAEformer[Liu et al\. \(2023\)](https://arxiv.org/html/2609.36119#bib.bib15)introduces hierarchical attention for multi\-scale modeling\. These methods demonstrate substantial improvements by jointly modeling spatial and temporal correlations\.Decoupling and Decomposition\.Existing methods overlook that coupling structures vary significantly across datasets, leading to spurious correlations when architectural assumptions mismatch data characteristics\. D2STGNN[Shao et al\. \(2022b\)](https://arxiv.org/html/2609.36119#bib.bib12)decomposes temporal from spatial\-temporal components to preserve location\-specific patterns\. However, it lacks a systematic decomposition of all three correlation types, and employs fixed recomposition without adaptive mechanisms\. In this work, we propose AdaST, which adaptively couple spatial and temporal signals in forecasting, enabling automatic adaptation to diverse coupling structures\.

## 6Conclusion

In this paper, we highlight a fundamental challenge in spatial\-temporal forecasting: coupling structures vary significantly across datasets, yet existing approaches apply uniform architectural priors that fail to accommodate this variability\. To address this mismatch, we propose AdaST, an adaptive framework that dynamically modulates spatial and temporal modeling through a decompose\-recompose paradigm\. By disentangling temporal\-specific, spatial\-specific, and spatiotemporally coupled patterns and adaptively recombining them in a data\-dependent manner, AdaST aligns model inductive biases with intrinsic coupling structures\. Extensive experiments demonstrate that AdaST consistently outperforms state\-of\-the\-art baselines across diverse benchmarks, while providing interpretable insights into dataset\-specific coupling characteristics\.

## Acknowledgments and Disclosure of Funding

The authors declare no competing interests\. This work was supported in part by the National Science Foundation \(NSF\) under Grants 2125326, 2144209, 2223769, 2228534, 2402438, 2411248, and 2601942; the Office of Naval Research \(ONR\) under Grant N000142412636; the Commonwealth Cyber Initiative \(CCI\) under Grant HV\-4Q26\-073\.

## References

- Atluriet al\.\(2018\)G\. Atluri, A\. Karpatne, and V\. KumarSpatio\-temporal data mining: a survey of problems and methods\.ACM Computing Surveys \(CSUR\)51\(4\),pp\. 1–41\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p1.1)\.
- Baiet al\.\(2020\)L\. Bai, L\. Yao, C\. Li, X\. Wang, and C\. WangAdaptive graph convolutional recurrent network for traffic forecasting\.Advances in neural information processing systems33,pp\. 17804–17815\.Cited by:[§4\.1](https://arxiv.org/html/2609.36119#S4.SS1.p2.1)\.
- Birant and Kut \(2007\)D\. Birant and A\. KutST\-dbscan: an algorithm for clustering spatial–temporal data\.Data & knowledge engineering60\(1\),pp\. 208–221\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p1.1)\.
- Caoet al\.\(2020\)D\. Cao, Y\. Wang, J\. Duan, C\. Zhang, X\. Zhu, C\. Huang, Y\. Tong, B\. Xu, J\. Bai, J\. Tong,et al\.Spectral temporal graph neural network for multivariate time\-series forecasting\.Advances in neural information processing systems33,pp\. 17766–17778\.Cited by:[§4\.1](https://arxiv.org/html/2609.36119#S4.SS1.p2.1)\.
- Caoet al\.\(2025\)L\. Cao, B\. Wang, G\. Jiang, Y\. Yu, and J\. DongSpatiotemporal\-aware trend\-seasonality decomposition network for traffic flow forecasting\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.39,pp\. 11463–11471\.Cited by:[§4\.1](https://arxiv.org/html/2609.36119#S4.SS1.p2.1)\.
- Contreraset al\.\(2003\)J\. Contreras, R\. Espinola, F\. J\. Nogales, and A\. J\. ConejoARIMA models to predict next\-day electricity prices\.IEEE transactions on power systems18\(3\),pp\. 1014–1020\.Cited by:[§5](https://arxiv.org/html/2609.36119#S5.p1.1)\.
- Cuiet al\.\(2021\)Y\. Cui, J\. Xie, and K\. ZhengHistorical inertia: a neglected but powerful baseline for long sequence time\-series forecasting\.InProceedings of the 30th ACM international conference on information & knowledge management,pp\. 2965–2969\.Cited by:[§4\.1](https://arxiv.org/html/2609.36119#S4.SS1.p2.1)\.
- Denget al\.\(2021\)J\. Deng, X\. Chen, R\. Jiang, X\. Song, and I\. W\. TsangSt\-norm: spatial and temporal normalization for multi\-variate time series forecasting\.InProceedings of the 27th ACM SIGKDD conference on knowledge discovery & data mining,pp\. 269–278\.Cited by:[§4\.1](https://arxiv.org/html/2609.36119#S4.SS1.p2.1)\.
- Denget al\.\(2024\)J\. Deng, X\. Chen, R\. Jiang, D\. Yin, Y\. Yang, X\. Song, and I\. W\. TsangDisentangling structured components: towards adaptive, interpretable and scalable time series forecasting\.IEEE Transactions on Knowledge and Data Engineering36\(8\),pp\. 3783–3800\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p4.1)\.
- Diaoet al\.\(2019\)Z\. Diao, X\. Wang, D\. Zhang, Y\. Liu, K\. Xie, and S\. HeDynamic spatial\-temporal graph convolutional neural networks for traffic forecasting\.InProceedings of the AAAI conference on artificial intelligence,Vol\.33,pp\. 890–897\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p1.1)\.
- Donget al\.\(2024\)Z\. Dong, R\. Jiang, H\. Gao, H\. Liu, J\. Deng, Q\. Wen, and X\. SongHeterogeneity\-informed meta\-parameter learning for spatiotemporal time series forecasting\.InProceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining,pp\. 631–641\.Cited by:[§4\.1](https://arxiv.org/html/2609.36119#S4.SS1.p2.1)\.
- Grigsbyet al\.\(2021\)J\. Grigsby, Z\. Wang, N\. Nguyen, and Y\. QiLong\-range transformers for dynamic spatiotemporal forecasting\.arXiv preprint arXiv:2109\.12218\.Cited by:[§4\.4](https://arxiv.org/html/2609.36119#S4.SS4.p1.1)\.
- Guoet al\.\(2019\)S\. Guo, Y\. Lin, N\. Feng, C\. Song, and H\. WanAttention based spatial\-temporal graph convolutional networks for traffic flow forecasting\.InProceedings of the AAAI conference on artificial intelligence,Vol\.33,pp\. 922–929\.Cited by:[§5](https://arxiv.org/html/2609.36119#S5.p1.1)\.
- Guoet al\.\(2021\)S\. Guo, Y\. Lin, H\. Wan, X\. Li, and G\. CongLearning dynamics and heterogeneity of spatial\-temporal graph data for traffic forecasting\.IEEE Transactions on Knowledge and Data Engineering34\(11\),pp\. 5415–5428\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p1.1)\.
- Heet al\.\(2017\)G\. He, G\. Jin, and Y\. YangSpace\-time correlations and dynamic coupling in turbulent flows\.Annual Review of Fluid Mechanics49\(1\),pp\. 51–70\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p2.1)\.
- Kanget al\.\(2019\)Z\. Kang, H\. Pan, S\. C\. Hoi, and Z\. XuRobust graph learning from noisy data\.IEEE transactions on cybernetics50\(5\),pp\. 1833–1843\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p3.1)\.
- Kerner \(2002\)B\. S\. KernerEmpirical macroscopic features of spatial\-temporal traffic patterns at highway bottlenecks\.Physical Review E65\(4\),pp\. 046138\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p3.1)\.
- Lanet al\.\(2022\)S\. Lan, Y\. Ma, W\. Huang, W\. Wang, H\. Yang, and P\. LiDstagnn: dynamic spatial\-temporal aware graph neural network for traffic flow forecasting\.InInternational conference on machine learning,pp\. 11906–11917\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p3.1)\.
- Leiet al\.\(2025\)Z\. Lei, Y\. Dong, J\. Li, and C\. ChenSt\-fit: inductive spatial\-temporal forecasting with limited training data\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.39,pp\. 12031–12039\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p1.1)\.
- Liet al\.\(2025\)L\. Li, E\. E\. Ozguven, Y\. Zhao, G\. Wang, Y\. Xie, and Y\. DongTyphoFormer: language\-augmented transformer for accurate typhoon track forecasting\.InProceedings of the 33rd ACM International Conference on Advances in Geographic Information Systems,pp\. 1174–1177\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p1.1)\.
- Li and Zhu \(2021\)M\. Li and Z\. ZhuSpatial\-temporal fusion graph neural networks for traffic flow forecasting\.InProceedings of the AAAI conference on artificial intelligence,Vol\.35,pp\. 4189–4196\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p1.1)\.
- Liet al\.\(2017\)Y\. Li, R\. Yu, C\. Shahabi, and Y\. LiuDiffusion convolutional recurrent neural network: data\-driven traffic forecasting\.arXiv preprint arXiv:1707\.01926\.Cited by:[§4\.1](https://arxiv.org/html/2609.36119#S4.SS1.p2.1),[§5](https://arxiv.org/html/2609.36119#S5.p1.1)\.
- Liuet al\.\(2023\)H\. Liu, Z\. Dong, R\. Jiang, J\. Deng, J\. Deng, Q\. Chen, and X\. SongStaeformer: spatio\-temporal adaptive embedding makes vanilla transformer sota for traffic forecasting\.arXiv preprint arXiv:2308\.10425\.Cited by:[§4\.1](https://arxiv.org/html/2609.36119#S4.SS1.p2.1),[§5](https://arxiv.org/html/2609.36119#S5.p1.1)\.
- Nieet al\.\(2022\)Y\. Nie, N\. H\. Nguyen, P\. Sinthong, and J\. KalagnanamA time series is worth 64 words: long\-term forecasting with transformers\.arXiv preprint arXiv:2211\.14730\.Cited by:[Appendix C](https://arxiv.org/html/2609.36119#A3.p1.1)\.
- Oreshkinet al\.\(2019\)B\. N\. Oreshkin, D\. Carpov, N\. Chapados, and Y\. BengioN\-beats: neural basis expansion analysis for interpretable time series forecasting\.arXiv preprint arXiv:1905\.10437\.Cited by:[§4\.1](https://arxiv.org/html/2609.36119#S4.SS1.p2.1),[§5](https://arxiv.org/html/2609.36119#S5.p1.1)\.
- Pierceet al\.\(2010\)J\. Pierce, D\. J\. Schiano, and E\. PaulosHome, habits, and energy: examining domestic interactions and energy consumption\.InProceedings of the SIGCHI conference on human factors in computing systems,pp\. 1985–1994\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p2.1)\.
- Qiet al\.\(2018\)L\. Qi, M\. Zhou, and W\. LuanA dynamic road incident information delivery strategy to reduce urban traffic congestion\.IEEE/CAA Journal of Automatica Sinica5\(5\),pp\. 934–945\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p2.1)\.
- Salinaset al\.\(2020\)D\. Salinas, V\. Flunkert, J\. Gasthaus, and T\. JanuschowskiDeepAR: probabilistic forecasting with autoregressive recurrent networks\.International journal of forecasting36\(3\),pp\. 1181–1191\.Cited by:[§4\.1](https://arxiv.org/html/2609.36119#S4.SS1.p2.1),[§5](https://arxiv.org/html/2609.36119#S5.p1.1)\.
- Shanget al\.\(2021\)C\. Shang, J\. Chen, and J\. BiDiscrete graph structure learning for forecasting multiple time series\.arXiv preprint arXiv:2101\.06861\.Cited by:[§4\.1](https://arxiv.org/html/2609.36119#S4.SS1.p2.1)\.
- Shaoet al\.\(2024\)Z\. Shao, F\. Wang, Y\. Xu, W\. Wei, C\. Yu, Z\. Zhang, D\. Yao, T\. Sun, G\. Jin, X\. Cao,et al\.Exploring progress in multivariate time series forecasting: comprehensive benchmarking and heterogeneity analysis\.IEEE Transactions on Knowledge and Data Engineering37\(1\),pp\. 291–305\.Cited by:[NeurIPS Paper Checklist](https://arxiv.org/html/2609.36119#Ax2.I1.ix21.p1.1)\.
- Shaoet al\.\(2022a\)Z\. Shao, Z\. Zhang, F\. Wang, W\. Wei, and Y\. XuSpatial\-temporal identity: a simple yet effective baseline for multivariate time series forecasting\.InProceedings of the 31st ACM international conference on information & knowledge management,pp\. 4454–4458\.Cited by:[§4\.1](https://arxiv.org/html/2609.36119#S4.SS1.p2.1)\.
- Shaoet al\.\(2022b\)Z\. Shao, Z\. Zhang, W\. Wei, F\. Wang, Y\. Xu, X\. Cao, and C\. S\. JensenDecoupled dynamic spatial\-temporal graph neural network for traffic forecasting\.arXiv preprint arXiv:2206\.09112\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p1.1),[§4\.1](https://arxiv.org/html/2609.36119#S4.SS1.p2.1),[§5](https://arxiv.org/html/2609.36119#S5.p1.1)\.
- Songet al\.\(2020\)C\. Song, Y\. Lin, S\. Guo, and H\. WanSpatial\-temporal synchronous graph convolutional networks: a new framework for spatial\-temporal network data forecasting\.InProceedings of the AAAI conference on artificial intelligence,Vol\.34,pp\. 914–921\.Cited by:[§4\.1](https://arxiv.org/html/2609.36119#S4.SS1.p1.1)\.
- Vuranet al\.\(2004\)M\. C\. Vuran, Ö\. B\. Akan, and I\. F\. AkyildizSpatio\-temporal correlation: theory and applications for wireless sensor networks\.Computer Networks45\(3\),pp\. 245–259\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p1.1)\.
- Wang and Song \(2018\)J\. Wang and G\. SongA deep spatial\-temporal ensemble model for air quality prediction\.Neurocomputing314,pp\. 198–206\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p2.1)\.
- Wanget al\.\(2022\)P\. Wang, X\. Wang, F\. Wang, M\. Lin, S\. Chang, H\. Li, and R\. JinKvt: k\-nn attention for boosting vision transformers\.InEuropean conference on computer vision,pp\. 285–302\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p3.1)\.
- Wanget al\.\(2020\)S\. Wang, J\. Cao, and S\. Y\. PhilipDeep learning for spatio\-temporal data mining: a survey\.IEEE transactions on knowledge and data engineering34\(8\),pp\. 3681–3700\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p1.1)\.
- Wuet al\.\(2020\)Z\. Wu, S\. Pan, G\. Long, J\. Jiang, X\. Chang, and C\. ZhangConnecting the dots: multivariate time series forecasting with graph neural networks\.InProceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining,pp\. 753–763\.Cited by:[§4\.1](https://arxiv.org/html/2609.36119#S4.SS1.p2.1)\.
- Wuet al\.\(2019\)Z\. Wu, S\. Pan, G\. Long, J\. Jiang, and C\. ZhangGraph wavenet for deep spatial\-temporal graph modeling\.arXiv preprint arXiv:1906\.00121\.Cited by:[§2\.3](https://arxiv.org/html/2609.36119#S2.SS3.p5.2),[§4\.1](https://arxiv.org/html/2609.36119#S4.SS1.p2.1),[§5](https://arxiv.org/html/2609.36119#S5.p1.1)\.
- Xiaet al\.\(2023\)Y\. Xia, Y\. Liang, H\. Wen, X\. Liu, K\. Wang, Z\. Zhou, and R\. ZimmermannDeciphering spatio\-temporal graph forecasting: a causal lens and treatment\.Advances in Neural Information Processing Systems36,pp\. 37068–37088\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p4.1)\.
- Xionget al\.\(2018\)H\. Xiong, A\. Vahedian, X\. Zhou, Y\. Li, and J\. LuoPredicting traffic congestion propagation patterns: a propagation graph approach\.InProceedings of the 11th ACM SIGSPATIAL international workshop on computational transportation science,pp\. 60–69\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p3.1)\.
- Yanget al\.\(2024\)Y\. Yang, M\. Jin, H\. Wen, C\. Zhang, Y\. Liang, L\. Ma, Y\. Wang, C\. Liu, B\. Yang, Z\. Xu,et al\.A survey on diffusion models for time series and spatio\-temporal data\.arXiv preprint arXiv:2404\.18886\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p1.1)\.
- Yiet al\.\(2024\)K\. Yi, Q\. Zhang, H\. He, K\. Shi, L\. Hu, N\. An, and Z\. NiuDeep coupling network for multivariate time series forecasting\.ACM Transactions on Information Systems42\(5\),pp\. 1–28\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p1.1)\.
- Yuet al\.\(2017\)B\. Yu, H\. Yin, and Z\. ZhuSpatio\-temporal graph convolutional networks: a deep learning framework for traffic forecasting\.arXiv preprint arXiv:1709\.04875\.Cited by:[§4\.1](https://arxiv.org/html/2609.36119#S4.SS1.p2.1),[§5](https://arxiv.org/html/2609.36119#S5.p1.1)\.
- Zenget al\.\(2023\)A\. Zeng, M\. Chen, L\. Zhang, and Q\. XuAre transformers effective for time series forecasting?\.InProceedings of the AAAI conference on artificial intelligence,Vol\.37,pp\. 11121–11128\.Cited by:[Appendix C](https://arxiv.org/html/2609.36119#A3.p1.1)\.
- Zhanget al\.\(2019\)C\. Zhang, J\. James, and Y\. LiuSpatial\-temporal graph attention networks: a deep learning approach for traffic forecasting\.Ieee Access7,pp\. 166246–166256\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p1.1)\.
- Zhaoet al\.\(2026\)K\. Zhao, B\. Shen, Y\. Dai, S\. Chakraborty, and Y\. DongGraphIP\-bench: how hard is it to steal a graph neural network, and can we stop it?\.arXiv preprint arXiv:2605\.12827\.Cited by:[§1](https://arxiv.org/html/2609.36119#S1.p1.1)\.
- Zhenget al\.\(2020\)C\. Zheng, X\. Fan, C\. Wang, and J\. QiGman: a graph multi\-attention network for traffic prediction\.InProceedings of the AAAI conference on artificial intelligence,Vol\.34,pp\. 1234–1241\.Cited by:[§5](https://arxiv.org/html/2609.36119#S5.p1.1)\.

## Appendix Overview

1. A\.Synthetic Dataset Details\.[A](https://arxiv.org/html/2609.36119#A1)
2. B\.Implementation and Efficiency Analysis\.[B](https://arxiv.org/html/2609.36119#A2)
3. C\.Additional Main Results\.[C](https://arxiv.org/html/2609.36119#A3)
4. D\.Additional Ablation Study\.[D](https://arxiv.org/html/2609.36119#A4)
5. E\.Extended Interpretability Analysis\.[E](https://arxiv.org/html/2609.36119#A5)
6. F\.Limitation and Broader Impact\.[F](https://arxiv.org/html/2609.36119#A6)

## Appendix ASynthetic Dataset Details

In this section, we provide comprehensive details of the synthetic datasets used in Section[2\.3](https://arxiv.org/html/2609.36119#S2.SS3)\.

Data Generation\.For each synthetic dataset, we generateN=10N=10spatially distributed time series, each consisting ofT=200T=200time steps\. The autoregressive coefficients\{ϕi,k\}k=112\\\{\\phi\_\{i,k\}\\\}\_\{k=1\}^\{12\}are randomly sampled from a uniform distribution𝒰⁡\(0,0\.5\)\\mathcal\{U\}\(0,0\.5\)for each seriesii, ensuring stable AR processes\. The multi\-frequency exogenous signalssi​\(t\)s\_\{i\}\(t\)are constructed by summing three sinusoidal components with randomly selected frequencies and amplitudes uniformly sampled from\[0,0\.5\]\[0,0\.5\]\. For the spatial\-dominated dataset, the adjacency matrix𝐀\\mathbf\{A\}is generated as a directed sparse random graph with an edge density of0\.430\.43, where we enforce that each node has at least one incoming and one outgoing edge to ensure connectivity\.

Qualitative Analysis\.Figure[6](https://arxiv.org/html/2609.36119#A1.F6)visualizes representative time series from the three synthetic datasets\. We observe distinct patterns that reflect their underlying coupling structures: \(1\) The temporal\-dominated \(TD\) dataset exhibits smooth trajectories where historical information provides reliable predictive signals with minimal noise interference from neighboring series\. \(2\) The spatial\-dominated \(SD\) dataset displays more volatile, noisy patterns that prevent long\-term historical information from being informative; instead, predictions rely primarily on short\-term spatial information from neighbors at the most recent time step\. \(3\) The strongly\-coupled \(SC\) dataset shows intermediate characteristics, combining both smooth temporal trends and spatial propagation effects\.

Quantitative Correlation Analysis\.To quantitatively validate the intended coupling structures, we analyze temporal and spatial correlations in each synthetic dataset\. Figure[7](https://arxiv.org/html/2609.36119#A1.F7)shows the autocorrelation function \(ACF\) for temporal correlation, while Figure[8](https://arxiv.org/html/2609.36119#A1.F8)visualizes spatial correlation measured using Pearson correlation coefficients between spatially connected node pairs\. The results confirm our design intent: temporal\-dominated data exhibits significantly higher temporal correlation \(0\.63±0\.070\.63\\pm 0\.07\) but negligible spatial correlation \(−0\.05±0\.06\-0\.05\\pm 0\.06\), while spatial\-dominated data shows the opposite pattern with stronger spatial correlation \(0\.33±0\.120\.33\\pm 0\.12\) and weaker temporal correlation \(0\.25±0\.010\.25\\pm 0\.01\)\. The strongly\-coupled dataset demonstrates moderate levels of both correlations \(0\.54±0\.050\.54\\pm 0\.05temporal,0\.21±0\.020\.21\\pm 0\.02spatial\), validating the successful construction of three distinct coupling regimes\.

Figure 6:Visualization of three synthetic datasets with distinct coupling structures: \(a\) Temporal\-Dominated \(TD\), \(b\) Spatial\-Dominated \(SD\), and \(c\) Strongly\-Coupled \(SC\)\.Figure 7:Temporal correlation \(ACF\) analysis: TD data exhibits strong autocorrelation while SD data shows weak temporal dependencies\.![Refer to caption](https://arxiv.org/html/2609.36119v1/SpatialCorr.png)Figure 8:Spatial correlation \(Pearson\) analysis: SD data exhibits the strongest spatial dependencies while TD data shows negligible spatial correlation\.
## Appendix BImplementation and Efficiency Analysis

Hyperparameter Sensitivity\.We set expert embedding dimensions toDn=Dh=Dw=Da=24D\_\{n\}=D\_\{h\}=D\_\{w\}=D\_\{a\}=24and hidden dimension toDH=256D\_\{H\}=256\. The model usesL=3L=3layers with 4\-head temporal attention\. The correlation modulation weightα=0\.1\\alpha=0\.1is selected via grid search over\{0\.01,0\.05,0\.1,0\.5,1\.0\}\\\{0\.01,0\.05,0\.1,0\.5,1\.0\\\}on the validation set\. Both input and prediction lengths are fixed at 12 time steps \(T=T′=12T=T^\{\\prime\}=12\)\.

Efficiency Analysis of Spatial Mixer\.To evaluate the computational efficiency of our spatial mixer, we compare training time per epoch against spatial attention across all four main benchmarks\. As shown in Table[4](https://arxiv.org/html/2609.36119#A2.T4), the spatial mixer consistently achieves substantial speedups over spatial attention, ranging from1\.5×1\.5\\timeson PEMS04 to2\.0×2\.0\\timeson PurpleAir\. This efficiency gain stems from the spatial mixer’s input\-independent mixing matrix: although both operators share the same𝒪⁡\(N2​T​D\)\\mathcal\{O\}\(N^\{2\}TD\)aggregation cost, the spatial mixer computessoftmax​\(𝐀\)\\text\{softmax\}\(\\mathbf\{A\}\)only once per forward pass, avoiding the per\-sample query\-key scoring, softmax normalization, and attention\-map storage required by spatial attention\.

Table 4:Average training time per epoch \(seconds\) comparing spatial mixer and spatial attention\.
## Appendix CAdditional Main Results

We conduct additional experiments on three diverse dataset categories to further validate AdaST’s generalizability: ExchangeRate \(financial\), ETTh1 \(energy\), and METR\-LA \(highway traffic\)\. We expand the baseline comparison to include two state\-of\-the\-art time\-series\-only methods, PatchTST\[[24](https://arxiv.org/html/2609.36119#bib.bib43)\]and DLinear\[[45](https://arxiv.org/html/2609.36119#bib.bib44)\]\. For spatial\-temporal baselines that require a predefined graph structure, we only compare against STID, HimNet, and STNorm, which operate without predefined graphs\. Since some datasets contain large amounts of missing data and zero values that render RMSE and MAPE unstable, we report MAE as the primary metric\. All experiments follow the same short\-term forecasting setting as the main paper, with input and prediction lengths both fixed at 12 time steps\.

Table[5](https://arxiv.org/html/2609.36119#A3.T5)presents the results, from which we draw the following observations: \(1\) AdaST achieves the best average rank across all datasets, demonstrating strong generalizability beyond the traffic and air quality domains evaluated in the main paper\. \(2\) Spatial\-temporal baselines generally match or outperform time\-series\-only baselines on datasets with non\-negligible spatial correlations \(e\.g\., METR\-LA, PEMS08\), confirming the value of explicit spatial modeling when cross\-location dependencies are present\. \(3\) Conversely, time\-series\-only baselines such as PatchTST and DLinear outperform spatial\-temporal counterparts on temporally\-dominated datasets such as ExchangeRate and ETTh1\. This is consistent with our central argument: forcing spatial coupling onto temporally\-dominated data introduces spurious cross\-location dependencies that degrade predictive performance\.

Table 5:Results on additional benchmarks\. Best and second\-best scores are inboldandunderlined\.MethodExchangeRateETTh1METR\-LAPurpleAirPEMS08Avg\. RankMAE / RankMAE / RankMAE / RankMAE / RankMAE / RankPatchTST0\.073/ 10\.4610\.461/ 44\.7624\.762/ 50\.505¯\\underline\{0\.505\}/ 222\.0722\.07/ 53\.4DLinear0\.0770\.077/ 30\.444¯\\underline\{0\.444\}/ 24\.8204\.820/ 60\.5250\.525/ 322\.5122\.51/ 64\.0STID0\.075¯\\underline\{0\.075\}/ 20\.4660\.466/ 53\.1463\.146/ 30\.5630\.563/ 414\.2114\.21/ 43\.6STNorm0\.1080\.108/ 50\.4920\.492/ 63\.1533\.153/ 40\.5740\.574/ 515\.4115\.41/ 34\.6HimNet0\.1240\.124/ 60\.4580\.458/ 33\.131¯\\underline\{3\.131\}/ 20\.5740\.574/ 513\.52¯\\underline\{13\.52\}/ 23\.6AdaST0\.0830\.083/ 40\.407/ 13\.128/ 10\.489/ 113\.45/ 11\.6

## Appendix DAdditional Ablation Study

We present additional ablation experiments on PEMS04 and PEMS08 to complement the analysis in Section[4\.3](https://arxiv.org/html/2609.36119#S4.SS3)\. Table[6](https://arxiv.org/html/2609.36119#A4.T6)leads to the following observations: \(1\) Consistent with our findings on PurpleAir and PEMS07, all four heterogeneity\-aware experts contribute meaningfully to overall performance, though their relative importance varies across datasets\. On both PEMS04 and PEMS08, removing the time\-of\-day expert𝐄𝐡\\mathbf\{E\_\{h\}\}causes the most substantial performance degradation, reflecting the critical role of diurnal traffic patterns in these datasets\. \(2\) Replacing the spatial mixer with spatial attention \(SpaAtt\) yields only marginal differences on PEMS04 and PEMS08\. This suggests that the performance gap between these two spatial modeling strategies is less pronounced on strongly spatial\-coupled traffic datasets, compared to temporal\-dominated data like PurpleAir where attention introduces spurious long\-range correlations\. \(3\) Both learned gate scores \(gg\) and correlation measures \(𝒞\\mathcal\{C\}\) are essential for effective recomposition, enabling AdaST to dynamically balance component contributions based on their reliability and relevance across diverse coupling structures\.

Table 6:Ablation study on PEMS04 and PEMS08\.
## Appendix EExtended Interpretability Analysis

This section provides deeper insights into AdaST’s learned representations and adaptive gating behavior, complementing the coupling structure analysis and case study in the main paper\.

### E\.1Correlation Measure Analysis

We investigate how the normalized correlation measures𝒞t​\(𝐇\(t\)\)\\mathcal\{C\}\_\{t\}\(\\mathbf\{H\}^\{\(t\)\}\),𝒞s​\(𝐇\(s\)\)\\mathcal\{C\}\_\{s\}\(\\mathbf\{H\}^\{\(s\)\}\), and𝒞s​t​\(𝐇\(s​t\)\)\\mathcal\{C\}\_\{st\}\(\\mathbf\{H\}^\{\(st\)\}\)vary across datasets, providing further insight into the relationship between learned representations and adaptive recomposition\. Figure[9](https://arxiv.org/html/2609.36119#A5.F9)visualizes the average normalized correlation measures across multiple datasets\. We observe a strong correspondence between correlation measures and gate scores \(Figure[3](https://arxiv.org/html/2609.36119#S4.F3)in the main paper\): datasets with high spatial correlation measures \(e\.g\., PEMS\) also exhibit high spatial gate scores, validating our design rationale that both metrics reflect the information content and reliability of each component\.

We further identify two systematic relationships among components:\(1\) Temporal\-Spatial Trade\-off:Temporal and spatial correlation measures exhibit a negative relationship — datasets with stronger temporal correlations tend to show weaker spatial correlations, and vice versa\.\(2\) Temporal\-Spatiotemporal Alignment:Temporal correlation measures show a positive relationship with spatiotemporal correlation measures\. This can be attributed to the fact that reduced spatial correlation leads to enhanced temporal and spatiotemporal correlations, with the spatiotemporal component capturing residual joint patterns that complement pure temporal dynamics\.

![Refer to caption](https://arxiv.org/html/2609.36119v1/corr_scatter.png)Figure 9:Normalized correlation measures across datasets, further validating the adaptive recomposition mechanism\. Strong correspondence with gate scores in Figure[3](https://arxiv.org/html/2609.36119#S4.F3)confirms that correlation measures reliably reflect each component’s information content\.
### E\.2Additional Case Studies

Temporal Dynamics of Gate Scores\.Figures[10](https://arxiv.org/html/2609.36119#A5.F10)and[11](https://arxiv.org/html/2609.36119#A5.F11)visualize prediction results alongside learned gate scores across fine\-grained time periods for two representative locations in PEMS07\. Beyond confirming that gate scores vary across time periods \(consistent with Figure[4](https://arxiv.org/html/2609.36119#S4.F4)in the main paper\), we observe an interesting phenomenon: more accurate predictions correlate with more dynamic, temporally\-varying gate scores\. This is particularly evident in Figure[11](https://arxiv.org/html/2609.36119#A5.F11), where location00achieves better prediction accuracy than location11, accompanied by more turbulent gate score trajectories\. We hypothesize that this reflects the model’s ability to capture fine\-grained coupling dynamics: accurate predictions require adaptive, context\-sensitive weighting that responds to local temporal variations in coupling structure, whereas smoother gate scores may indicate insufficient adaptation to changing conditions\.

Spatial Distribution of Gate Scores\.Figure[12](https://arxiv.org/html/2609.36119#A5.F12)visualizes the learned gate scores across160160spatial locations for a fixed time period\. We observe two complementary patterns that validate our heterogeneity\-aware design: \(1\) Most locations exhibit similar gate score distributions, indicating a stable dataset\-level coupling structure\. \(2\) Despite this overall similarity, certain locations display distinctly different gate score patterns, underscoring the necessity of incorporating spatial heterogeneity experts \(𝐄n\\mathbf\{E\}\_\{n\}\) to capture location\-specific coupling characteristics\.

Figure 10:Fine\-grained temporal visualization \(1\-day span\) of predictions and gate scores for 2 locations in PEMS07\. More accurate predictions correspond to more dynamic gate score trajectories\.Figure 11:Fine\-grained temporal visualization \(10\-hour span\) of predictions and gate scores for 2 locations in PEMS07\. Location 0 achieves better accuracy alongside more adaptive gate score dynamics than location 1\.![Refer to caption](https://arxiv.org/html/2609.36119v1/gate_maps.png)Figure 12:Heatmap of gate scores across 160 spatial locations, showing a globally stable coupling structure with local variations that motivate the use of spatial heterogeneity experts\.

## Appendix FLimitation and Broader Impact

### F\.1Limitation

While AdaST demonstrates strong and consistent performance across diverse benchmarks, two limitations remain worth noting\. First, the decompose\-recompose paradigm introduces additional parameters relative to single\-branch architectures, which may increase memory consumption on datasets with very large numbers of nodes\. Second, our current evaluation focuses on short\-term forecasting with fixed input and output lengths of 12 time steps; extending AdaST to long\-term forecasting settings remains an open direction for future work\.

### F\.2Broader Impact

The primary positive impact of this work is enabling more reliable and transparent forecasting in applications such as traffic, air quality, and climate, which can support better planning and resource allocation\. As with many forecasting models, potential risks include misuse for overly confident decision\-making, distribution shift leading to degraded performance in deployment, and the amplification of biases or measurement errors present in sensor data\. We recommend careful validation under domain\-specific conditions, monitoring for performance drift, and transparency about model uncertainty when used in real\-world decision pipelines\.

## NeurIPS Paper Checklist

1. 1\.Claims
2. Question: Do the main claims made in the abstract and introduction accurately reflect the paper’s contributions and scope?
3. Answer:\[Yes\]
4. Justification: The abstract and introduction clearly state the three key contributions of AdaST: \(1\) identifying coupling structure mismatch as a fundamental challenge in ST forecasting, \(2\) proposing the decompose\-recompose framework with heterogeneity\-aware experts and correlation\-informed recomposition, and \(3\) demonstrating state\-of\-the\-art performance across diverse benchmarks\. All claims are empirically supported by preliminary experiments in Section[2\.3](https://arxiv.org/html/2609.36119#S2.SS3)and comprehensive evaluations in Section 4\.
5. 2\.Limitations
6. Question: Does the paper discuss the limitations of the work performed by the authors?
7. Answer:\[Yes\]
8. Justification: Limitations are discussed in Appendix F\.1, noting that the decompose\-recompose paradigm introduces additional parameters compared to single\-branch architectures, and that the current evaluation is restricted to short\-term forecasting with fixed horizon length of 12 steps\.
9. 3\.Theory assumptions and proofs
10. Question: For each theoretical result, does the paper provide the full set of assumptions and a complete \(and correct\) proof?
11. Answer:\[N/A\]
12. Justification: This paper does not include formal theoretical results or proofs\. The coupling structure formulations in Section 2\.2 are descriptive definitions used to motivate the framework design rather than theorems requiring proof\.
13. 4\.Experimental result reproducibility
14. Question: Does the paper fully disclose all the information needed to reproduce the main experimental results of the paper to the extent that it affects the main claims and/or conclusions of the paper \(regardless of whether the code and data are provided or not\)?
15. Answer:\[Yes\]
16. Justification: Full implementation details are provided in Section 4\.1 and Appendix[B](https://arxiv.org/html/2609.36119#A2), including model architecture, all hyperparameters, optimizer settings, hardware specifications, data splits, preprocessing procedures, and evaluation metrics\. An anonymized code repository is available at[https://anonymous\.4open\.science/r/AdaST\-70BC](https://anonymous.4open.science/r/AdaST-70BC)\.
17. 5\.Open access to data and code
18. Question: Does the paper provide open access to the data and code, with sufficient instructions to faithfully reproduce the main experimental results, as described in supplemental material?
19. Answer:\[Yes\]
20. Justification: An anonymized code repository is available at[https://anonymous\.4open\.science/r/AdaST\-70BC](https://anonymous.4open.science/r/AdaST-70BC)\. All datasets used \(PEMS04/07/08, PurpleAir, METR\-LA, ETTh1, ExchangeRate\) are publicly available benchmarks, with preprocessing procedures described in Section 4\.1 and Appendix[B](https://arxiv.org/html/2609.36119#A2)\.
21. 6\.Experimental setting/details
22. Question: Does the paper specify all the training and test details \(e\.g\., data splits, hyperparameters, how they were chosen, type of optimizer\) necessary to understand the results?
23. Answer:\[Yes\]
24. Justification: Section 4\.1 provides dataset statistics, data splits \(60/20/20\), Z\-score normalization, optimizer settings \(Adam, lr=0\.001, exponential decay, batch size=16\), and all key hyperparameters \(DHD\_\{H\}=256,LL=3,α\\alpha=0\.1\)\. Appendix[B](https://arxiv.org/html/2609.36119#A2)further details hyperparameter selection via grid search\.
25. 7\.Experiment statistical significance
26. Question: Does the paper report error bars suitably and correctly defined or other appropriate information about the statistical significance of the experiments?
27. Answer:\[No\]
28. Justification: Error bars are not reported due to the high computational cost of running multiple seeds across 16 baselines on 4 datasets\. This is consistent with standard practice in the spatial\-temporal forecasting community\[[30](https://arxiv.org/html/2609.36119#bib.bib45)\], and results follow the same single\-run evaluation protocol as all compared baselines\.
29. 8\.Experiments compute resources
30. Question: For each experiment, does the paper provide sufficient information on the computer resources \(type of compute workers, memory, time of execution\) needed to reproduce the experiments?
31. Answer:\[Yes\]
32. Justification: Section 4\.1 states that all experiments are conducted on an NVIDIA A100 80GB GPU\. Per\-epoch training times for all four main benchmarks are reported in Table[4](https://arxiv.org/html/2609.36119#A2.T4)in Appendix[B](https://arxiv.org/html/2609.36119#A2)\.
33. 9\.Code of ethics
35. Answer:\[Yes\]
36. Justification: This work uses only publicly available sensor datasets for forecasting tasks and involves no human subjects, sensitive personal data, or applications with direct harm potential\. The research fully conforms with the NeurIPS Code of Ethics\.
37. 10\.Broader impacts
38. Question: Does the paper discuss both potential positive societal impacts and negative societal impacts of the work performed?
39. Answer:\[Yes\]
40. Justification: Broader impacts are discussed in Appendix F\.2\. Positive impacts include improved forecasting reliability for traffic management, air quality monitoring, and climate applications\. Potential risks include overconfident decision\-making, performance degradation under distribution shift, and amplification of sensor biases, along with recommended mitigation strategies\.
41. 11\.Safeguards
42. Question: Does the paper describe safeguards that have been put in place for responsible release of data or models that have a high risk for misuse \(e\.g\., pre\-trained language models, image generators, or scraped datasets\)?
43. Answer:\[N/A\]
44. Justification: This paper proposes a general spatial\-temporal forecasting framework trained on public sensor datasets\. It does not release pre\-trained generative models or scraped data that carry significant misuse risk\.
45. 12\.Licenses for existing assets
46. Question: Are the creators or original owners of assets \(e\.g\., code, data, models\), used in the paper, properly credited and are the license and terms of use explicitly mentioned and properly respected?
47. Answer:\[Yes\]
48. Justification: All datasets and baseline methods are properly cited in Section 4\.1 and the references\. Baselines are evaluated using their official implementations with recommended hyperparameters, as stated in Section 4\.1\.
49. 13\.New assets
50. Question: Are new assets introduced in the paper well documented and is the documentation provided alongside the assets?
51. Answer:\[Yes\]
52. Justification: The AdaST codebase is released at[https://anonymous\.4open\.science/r/AdaST\-70BC](https://anonymous.4open.science/r/AdaST-70BC)with documentation covering model architecture, training procedures, and scripts to reproduce all main experimental results reported in the paper\.
53. 14\.Crowdsourcing and research with human subjects
54. Question: For crowdsourcing experiments and research with human subjects, does the paper include the full text of instructions given to participants and screenshots, if applicable, as well as details about compensation \(if any\)?
55. Answer:\[N/A\]
56. Justification: This paper does not involve crowdsourcing or research with human subjects\. All experiments are conducted on publicly available sensor and traffic datasets\.
57. 15\.Institutional review board \(IRB\) approvals or equivalent for research with human subjects
58. Question: Does the paper describe potential risks incurred by study participants, whether such risks were disclosed to the subjects, and whether Institutional Review Board \(IRB\) approvals \(or an equivalent approval/review based on the requirements of your country or institution\) were obtained?
59. Answer:\[N/A\]
60. Justification: This paper does not involve human subjects and therefore requires no IRB approval or equivalent review\.
61. 16\.Declaration of LLM usage
62. Question: Does the paper describe the usage of LLMs if it is an important, original, or non\-standard component of the core methods in this research?
63. Answer:\[N/A\]
64. Justification: LLMs are not used as any part of the core methodology\. Any incidental use of LLMs was limited to writing assistance only and does not affect the scientific contributions, experimental design, or reported results\.

Similar Articles

STKAN: Kolmogorov-Arnold Networks for Spatio-Temporal Forecasting

arXiv cs.LG

This paper introduces STKAN, a spatio-temporal forecasting architecture that integrates Taylor-polynomial Kolmogorov-Arnold Network modules for spatial and temporal token mixing. Experiments on five traffic benchmarks show competitive performance, suggesting nonlinear function approximators can complement architectural design.