RiskTraf: Risk-Extrapolated Residual Learning for Multi-Variate Traffic Flow Prediction
Summary
This paper introduces RiskTraf, a residual learning plug-in for multi-variate traffic flow prediction, and presents a new benchmark PEMSB-3V to enhance forecasting accuracy by extrapolating risk from speed and occupancy data.
View Cached Full Text
Cached at: 08/24/26, 04:32 AM
# Risk-Extrapolated Residual Learning for Multi-Variate Traffic Flow Prediction Source: [https://arxiv.org/html/2608.20656](https://arxiv.org/html/2608.20656) ## RiskTraf: Risk\-Extrapolated Residual Learning for Multi\-Variate Traffic Flow PredictionConference:Proceedings of the 35th ACM International Conference on Information and Knowledge Management; November 7–11, 2026; Rome, Italy\.Proceedings of the 35th ACM International Conference on Information and Knowledge Management \(CIKM ’26\), November 7–11, 2026, Rome, ItalyISBN:979\-8\-4007\-2539\-5/2026/11DOI:[10\.1145/3799682\.3840707](https://doi.org/10.1145/3799682.3840707)CCS:Computing methodologies Spatial and temporal reasoningCCS:Computing methodologies Neural networksCCS:Information systems Data mining Guangyu Wangemail:[hulnegy@gmail\.com](mailto:[email protected])Affiliation:Dongbei University of Finance & Economics,Dalian,Liaoning,ChinaZhidan LiuNote:Corresponding author\.email:[zhidanliu@hkust\-gz\.edu\.cn](mailto:[email protected])Affiliation:Hong Kong University of Science and Technology \(Guangzhou\),Guangzhou,Guangdong,China 2026; © cc ###### Abstract\. Traffic sensors commonly record flow, speed, and occupancy, but standard traffic flow forecasting benchmarks and models rarely exploit all three raw measurements reliably\. Although speed and occupancy provide sensor\-native traffic\-state information beyond flow alone, existing releases often omit these variables, replace them with proxies, or contain logically inconsistent records\. Moreover, direct empirical risk minimization over three\-variable inputs may exploit regime\-dependent shortcuts, as the relationships among flow, speed, and occupancy vary substantially between free\-flow and congested states\. We introducePEMSB\-3V, a public benchmark suite that preserves raw flow, speed, and occupancy measurements from PeMS detectors for flow prediction\. We also proposeRiskTraf, a model\-agnostic risk\-extrapolated residual plug\-in\. For each trained spatio\-temporal backbone, RiskTraf freezes the selected checkpoint and learns a lightweight zero\-start residual head from historical speed and occupancy\. The residual head constructs ordered traffic\-risk environments and optimizes horizon\-wise flow corrections with a risk extrapolation objective, thereby mitigating regime\-specific shortcut correlations without modifying the backbone\. Extensive experiments demonstrate that RiskTraf consistently improves diverse forecasting backbones and outperforms debiasing and distribution\-shift adaptation methods\. Our code and benchmark are available at[https://github\.com/Guangyu4/RiskTraf](https://github.com/Guangyu4/RiskTraf)\. ###### Keywords: Traffic Flow Prediction, Multi\-Variate Time Series Forecasting, Risk Extrapolation, Traffic Benchmark ††cc\-license:by## 1\.Introduction Traffic flow prediction aims to forecast the number of vehicles passing through road segment stations over future time intervals based on historical observations\. In widely used benchmarks such as PeMS \(Freeway Performance Measurement System\)\([8](https://arxiv.org/html/2608.20656#bib.bib14)\), traffic flow is typically defined as the vehicle count recorded by loop detectors within fixed time windows, e\.g\., every 5 minutes\. As a fundamental indicator of road utilization, accurate flow prediction plays a critical role in intelligent transportation systems\([43](https://arxiv.org/html/2608.20656#bib.bib13)\), enabling traffic authorities to anticipate congestion, optimize signal control, and allocate transportation resources more effectively\([35](https://arxiv.org/html/2608.20656#bib.bib17)\)\. Recent studies have made substantial progress in traffic forecasting by designing increasingly expressive spatio\-temporal models, including graph neural networks\([31](https://arxiv.org/html/2608.20656#bib.bib18);[45](https://arxiv.org/html/2608.20656#bib.bib19)\)and attention\-based architectures\([32](https://arxiv.org/html/2608.20656#bib.bib20)\)\. However, as the field matures, performance gains from purely architectural innovations have become increasingly marginal\([40](https://arxiv.org/html/2608.20656#bib.bib21)\), while model complexity and computational costs continue to grow\. This trend has motivated researchers to exploit richer contextual information for more accurate prediction\. Existing efforts have incorporated external sources such as weather conditions\([15](https://arxiv.org/html/2608.20656#bib.bib26);[30](https://arxiv.org/html/2608.20656#bib.bib27)\)and textual information\([49](https://arxiv.org/html/2608.20656#bib.bib28)\)\. Nevertheless, incorporating such auxiliary sources often requires additional collection, storage, and maintenance, thereby increasing the data\-side overhead of prediction systems\([10](https://arxiv.org/html/2608.20656#bib.bib48)\)\. More importantly, they are difficult to align with traffic measurements: weather typically affects broad regions while sensors observe localized road segments\([50](https://arxiv.org/html/2608.20656#bib.bib50)\), and textual signals such as news are updated at much coarser temporal granularity than traffic sensors\([22](https://arxiv.org/html/2608.20656#bib.bib49)\)\. Such spatial and temporal mismatches may introduce additional noise during data fusion\. Compared with external auxiliary sources, a more natural way to enrich traffic representations is to leverage multiple variables reported by the same traffic sensor\. In addition to flow, loop detectors commonly recordspeedandoccupancy, which describe complementary aspects of traffic states\. Speed captures vehicle movement dynamics, while occupancy measures the fraction of time during which detectors are occupied by vehicles\. Since these variables are collected from the same sensors and at the same temporal resolution as flow, they provide sensor\-native auxiliary information without the spatial and temporal misalignment of external data\. From an information\-theoretic perspective\([3](https://arxiv.org/html/2608.20656#bib.bib22)\), jointly modeling heterogeneous sensor variables can provide additional state information beyond any single variable\. A few studies have therefore begun to exploit flow, speed, and occupancy for traffic forecasting\([24](https://arxiv.org/html/2608.20656#bib.bib23);[39](https://arxiv.org/html/2608.20656#bib.bib24);[36](https://arxiv.org/html/2608.20656#bib.bib25)\)\. However, incorporating speed and occupancy into deep forecasting models is not a straightforward extension of flow\-only prediction\. Their effective use requires satisfying two coupled requirements: obtaining reliable three\-variable measurements from the same set of sensors, and learning from these auxiliary variables without relying on regime\-specific shortcuts\. Challenge 1: Lack of reliable multi\-variate benchmarks\.Existing traffic sensor releases do not guarantee that flow, speed, and occupancy are simultaneously valid and usable\. Raw loop\-detector records often require strict screening based on validity thresholds and traffic\-flow consistency\([42](https://arxiv.org/html/2608.20656#bib.bib15)\), and field detectors may suffer from faults such as stuck\-off, stuck\-on, or hanging\-on behaviors\([7](https://arxiv.org/html/2608.20656#bib.bib16)\)\. Moreover, speed and occupancy are sensitive to detector configurations, lane definitions, loop lengths, and road types\. As a result, a sensor that provides usable flow measurements may still be unreliable for raw three\-variable modeling\. This makes it difficult to systematically evaluate methods that exploit flow, speed, and occupancy under a standardized benchmark setting\. Challenge 2: Regime\-dependent auxiliary correlations\.Although speed and occupancy contain useful traffic\-state information, their relationships with flow are not invariant across traffic regimes\. For example, under congested conditions, slower vehicles may occupy loop detectors for longer durations, increasing occupancy even when the observed flow is similar\. Consequently, naively treating speed and occupancy as ordinary covariates may cause models to exploit regime\-specific shortcuts rather than robust predictive patterns\. Models trained with empirical risk minimization \(ERM\) can overfit such spurious correlations, leading to degraded generalization across traffic states\. Simply discarding speed and occupancy avoids this risk but also wastes valuable sensor information\. Existing distribution\-shift adaptation\([11](https://arxiv.org/html/2608.20656#bib.bib46);[16](https://arxiv.org/html/2608.20656#bib.bib11)\)and debiasing methods\([53](https://arxiv.org/html/2608.20656#bib.bib47)\)partially address related issues, but they often rely on specialized shift or bias detectors, whose effectiveness is limited by the accuracy and availability of such detectors\. To address these challenges, we introduce both a reliable benchmark suite and a risk\-aware learning framework\. First, we construct thePEMSB\-3V Benchmarkconsisting of four datasets: PEMS03\-B, PEMS04\-B, PEMS07\-B, and PEMS08\-B\. The benchmark is district\-aligned with widely used PeMS datasets while preserving raw flow, speed, and occupancy measurements reported by PeMS detectors\. By conducting sensor\-level validation, PEMSB\-3V provides a standardized testbed for studying historical three\-variable inputs with flow\-only forecasting targets\. Second, we proposeRiskTraf, ariskextrapolation plug\-in fortraffic prediction\. Rather than using speed and occupancy as unrestricted covariates or replacing existing forecasting architectures, RiskTraf treats them as historical auxiliary signals for residual correction\. Specifically, given a trained three\-variable backbone, we freeze its parameters and attach a lightweight zero\-initialized residual head driven by historical speed and occupancy\. The residual head learns horizon\-wise flow corrections across ordered low\-to\-high traffic\-risk environments, encouraging robust improvements while reducing reliance on regime\-specific auxiliary correlations\. RiskTraf performs rollback when validation MAE does not improve\. The main contributions of this work are summarized as follows: - •We releasePEMSB\-3V, a district\-aligned benchmark suite that enables standardized evaluation of flow prediction with historical flow, speed, and occupancy measurements\. - •We formulate regime\-dependent auxiliary correlations as a key obstacle in multi\-variate traffic forecasting and proposeRiskTraf, a REx\-based residual plug\-in that improves trained backbones without replacing their architectures\. - •We conduct extensive experiments on multiple PEMSB\-3V datasets and diverse spatio\-temporal backbones, showing that RiskTraf achieves consistent improvements over strong baselines, debiasing methods, and distribution\-shift adaptation approaches\. The rest of this paper is organized as follows\. Section[2](https://arxiv.org/html/2608.20656#S2)introduces the PEMSB\-3V benchmark\. Section[3](https://arxiv.org/html/2608.20656#S3)presents the RiskTraf framework\. Section[4](https://arxiv.org/html/2608.20656#S4)reports experimental results and analyses\. Section[5](https://arxiv.org/html/2608.20656#S5)discusses related work, and Section[6](https://arxiv.org/html/2608.20656#S6)concludes the paper\. ## 2\.PEMSB\-3V Benchmark We constructPEMSB\-3V, a benchmark suite consisting of four district\-level datasets:PEMS03\-B,PEMS04\-B,PEMS07\-B, andPEMS08\-B\. The suite is built from the California Department of Transportation Performance Measurement System \(PeMS\)111[https://pems\.dot\.ca\.gov/](https://pems.dot.ca.gov/)and its official source page222[https://dot\.ca\.gov/programs/traffic\-operations/mpr/pems\-source](https://dot.ca.gov/programs/traffic-operations/mpr/pems-source)\. PeMS is described by Caltrans as a statewide freeway monitoring system with nearly 40,000 detectors\. Unlike existing releases that replace missing variables with proxies or temporal codes, PEMSB\-3V retains only detectors whose native PeMS records contain all three raw measurements:flow,speed, andoccupancy\. The overall construction pipeline is summarized in Algorithm[1](https://arxiv.org/html/2608.20656#alg1)\. Algorithm 1PEMSB\-3V Benchmark Construction Pipeline0:District ID, time range \[tstart,tend\]\[t\_\{\\text\{start\}\},t\_\{\\text\{end\}\}\], completeness threshold ρ\\rho 0:Multi\-variate traffic benchmark 𝒟\\mathcal\{D\}and adjacency matrix 𝐀\\mathbf\{A\} 1:Query detector metadata from PeMS 2:Keep only detectors with raw flow, speed, and occupancy channels 3:Filter detectors whose temporal completeness is below ρ\\rho 4:foreach day in \[tstart,tend\]\[t\_\{\\text\{start\}\},t\_\{\\text\{end\}\}\]do 5:Download 5\-minute flow, speed, and occupancy records 6:Remove malformed timestamps and interpolate only short gaps 7:endfor 8:Aggregate the cleaned records into 𝐗∈ℝT×N×3\\mathbf\{X\}\\in\\mathbb\{R\}^\{T\\times N\\times 3\} 9:Construct the road topology and compute pairwise distances 10:Build 𝐀\\mathbf\{A\}with a distance kernel and sparsification threshold 11:Split the benchmark into train/val/test by ratio 6:2:2 12:return 𝒟\\mathcal\{D\}, 𝐀\\mathbf\{A\} Figure 1\.High\-quality PEMS03\-B sensors after metadata\-based screening\. The map shows 1,013 merged mainline/HOV detectors with complete metadata\.Map of high\-quality PEMS03\-B freeway mainline and HOV detector locations after metadata\-based screening\.To ensure that the benchmark is built from sensors with physically interpretable three\-variable measurements, we apply a metadata\-based progressive screening procedure before constructing the final sensor set\. The screening follows the earliest\-failed\-check rule\. We first perform an ID coverage check\. A target sensor without a corresponding PeMS metadata record is marked asMissing Metadata, since its static information cannot be verified\. We then perform a spatial localization check\. If latitude or longitude is missing, invalid, or outside the legal geographic range, the detector is marked asMissing Geolocationand excluded from map\-based topology construction\. Next, we check static structural completeness using detector length and lane count\. Records with empty, non\-numeric, or non\-positive values are marked asInvalid Static Attributes, because flow, speed, and occupancy depend on the detector segment and lane definition\. Finally, we check road\-type consistency\. Detectors withType∈\{ML,HV\}\\mathrm\{Type\}\\in\\\{\\mathrm\{ML\},\\mathrm\{HV\}\\\}are retained as freeway mainline or HOV sensors, whereasType∈\{OR,FR,FF\}\\mathrm\{Type\}\\in\\\{\\mathrm\{OR\},\\mathrm\{FR\},\\mathrm\{FF\}\\\}are treated as ramp or connector sensors and excluded from the mainline subset\. Figure[1](https://arxiv.org/html/2608.20656#S2.F1)illustrates the resulting PEMS03\-B sensor footprint after metadata\-based screening\. Among 1,708 target sensor IDs, 41 are removed due to missing metadata and 2 due to invalid or missing geolocation\. Among the remaining 1,665 geolocated records, 1,039 pass all metadata checks as complete mainline/HOV sensors, while 626 are identified as ramp or connector sensors\. After merging duplicate records with the same location, road direction, and detector type, the final map contains 1,013 high\-quality points, including 762 mainline and 251 HOV points\. This procedure makes each discarded sensor traceable to an interpretable failure category, rather than removing sensors through opaque filtering\. The subset design follows the administrative organization of PeMS rather than an arbitrary partition\. Specifically, PEMS03\-B, PEMS04\-B, PEMS07\-B, and PEMS08\-B correspond to Caltrans districts 03, 04, 07, and 08, respectively333[https://cwwp2\.dot\.ca\.gov/documentation/district\-map\-county\-chart\.htm](https://cwwp2.dot.ca.gov/documentation/district-map-county-chart.htm)\. This preserves geographic provenance while maintaining comparability with widely used PeMS benchmarks\. Table 1\.Availability and usability of raw flow, speed, and occupancy measurements in representative traffic benchmarks\.✓: usable raw measurement;✗: not provided;○\\bigcirc: provided but logically inconsistent\.SourceBenchmarkNodesIntervalFlowSpeedOccupancyDCRNN\([31](https://arxiv.org/html/2608.20656#bib.bib18)\)METR\-LA2075 mins✗✓✗PEMS\-BAY3255 mins✗✓✗LSTNet\([29](https://arxiv.org/html/2608.20656#bib.bib2)\)Traffic8621 hour✓○\\bigcirc○\\bigcircSTGCN\([48](https://arxiv.org/html/2608.20656#bib.bib7)\)PEMSD7\(M\)2285 mins✓○\\bigcirc○\\bigcircPEMSD7\(L\)1,0265 mins✓○\\bigcirc○\\bigcircASTGCN\([21](https://arxiv.org/html/2608.20656#bib.bib6)\)PEMSD4\-I2285 mins✓○\\bigcirc○\\bigcircPEMSD8\-I1,9795 mins✓○\\bigcirc○\\bigcircSTSGCN\([41](https://arxiv.org/html/2608.20656#bib.bib5)\)PEMS033585 mins✓○\\bigcirc○\\bigcircPEMS043075 mins✓○\\bigcirc○\\bigcircPEMS078835 mins✓○\\bigcirc○\\bigcircPEMS081705 mins✓○\\bigcirc○\\bigcircLargeST\([33](https://arxiv.org/html/2608.20656#bib.bib3)\)CA8,6005 mins✓✗✗GLA3,8345 mins✓✗✗GBA2,3525 mins✓✗✗SD7165 mins✓✗✗XXLTraffic\([47](https://arxiv.org/html/2608.20656#bib.bib4)\)Full\_PEMS031,8095 mins✓○\\bigcirc○\\bigcircFull\_PEMS044,0895 mins✓✗✗Full\_PEMS055735 mins✓✗✗Full\_PEMS067055 mins✓✗✗Full\_PEMS074,8885 mins✓✗✗Full\_PEMS082,0595 mins✓✗✗Full\_PEMS101,3785 mins✓✗✗Full\_PEMS111,4405 mins✓✗✗Full\_PEMS122,5875 mins✓✗✗tfNSW2760 mins✓✗✗TraffiDent\([19](https://arxiv.org/html/2608.20656#bib.bib51)\)TraffiDent∗16,9725 mins✓✓✓PEMSB\-3VPEMS03\-B1,0135 mins✓✓✓PEMS04\-B2,4745 mins✓✓✓PEMS07\-B2,7885 mins✓✓✓PEMS08\-B1,5155 mins✓✓✓ ∗TraffiDent provides all three variables but does not apply sensor\-type filtering or provide sufficient metadata to distinguish detector types\. We further audit representative public traffic benchmarks to examine whether they support the raw flow\-speed\-occupancy forecasting setting considered in this paper\. As shown in Table[1](https://arxiv.org/html/2608.20656#S2.T1), most existing benchmarks either omit speed and occupancy, replace them with proxy variables, or contain logically inconsistent records\. For example, PEMS03 and PEMS07 use daily and weekly temporal encodings in place of raw speed and occupancy, while PEMS04 and PEMS08 include physically implausible speed records, such as positive speed when flow is zero and truncated low\-speed ranges\. These issues prevent existing benchmarks from serving as reliable testbeds for studying how raw speed and occupancy contribute to flow prediction\. ## 3\.Methodology In this section, we presentRiskTraf, a model\-agnostic risk\-extrapolation plug\-in for multi\-variate traffic forecasting\. As illustrated in Figure[2](https://arxiv.org/html/2608.20656#S3.F2), RiskTraf adopts a paired two\-stage design without modifying the forecasting backbone: - •Stage Itrains a given standard spatio\-temporal backbone with 12 historical steps of traffic flow, speed, and occupancy to predict future flow\. - •Stage IIfreezes the trained backbone and attaches a lightweight residual head driven by historical speed and occupancy\. The residual head is optimized with a risk extrapolation objective to learn robust flow corrections across traffic\-risk environments\. Figure 2\.Overview of RiskTraf\. Stage I trains a backbone to predict future flow from historical flow, speed, and occupancy\. Stage II freezes the backbone, constructs traffic\-risk environments from historical speed and occupancy, and learns a lightweight residual head with risk extrapolation to correct the baseline prediction\.### 3\.1\.Problem Formulation and Paired Protocol LetFF,SS, andOOdenote traffic flow, speed, and occupancy, respectively\. For a sample ending at timet−1t\-1, the historical input is \(1\)Xt=\{Ft−T:t−1,St−T:t−1,Ot−T:t−1\},X\_\{t\}=\\\{F\_\{t\-T:t\-1\},S\_\{t\-T:t\-1\},O\_\{t\-T:t\-1\}\\\},whereTTis the input length and each variable is observed over all sensor nodes\. The prediction target is future flow only: \(2\)Yt=Ft:t\+H−1,Y\_\{t\}=F\_\{t:t\+H\-1\},whereHHis the prediction horizon\. Thus, speed and occupancy are available only as historical auxiliary measurements\. RiskTraf is applied under a paired protocol for each dataset\-backbone pair\. LetBϕB\_\{\\phi\}be an arbitrary spatio\-temporal backbone\. The Stage\-I baseline prediction is given by \(3\)Y^base=Bϕ\(Xt\)\.\\hat\{Y\}\_\{\\text\{base\}\}=B\_\{\\phi\}\(X\_\{t\}\)\.After selecting the baseline checkpoint by validation MAE, RiskTraf starts from that checkpoint, freezes the backbone parametersϕ\\phi, and learns a residual correction headgψg\_\{\\psi\}from historical speed and occupancy: \(4\)Y^risk=Y^base\+sr⋅gψ\(St−T:t−1,Ot−T:t−1\),\\hat\{Y\}\_\{\\text\{risk\}\}=\\hat\{Y\}\_\{\\text\{base\}\}\+s\_\{r\}\\cdot g\_\{\\psi\}\(S\_\{t\-T:t\-1\},O\_\{t\-T:t\-1\}\),wheresrs\_\{r\}is a small residual scaling factor; during Stage II only the residual parametersψ\\psiare updated\. ### 3\.2\.Stage I: Baseline Backbone Training The first stage follows the standard empirical risk minimization \(ERM\) setting adopted by existing traffic forecasting models\. Given a historical three\-variable sequenceXX, the backboneBϕB\_\{\\phi\}predicts the future flow sequenceYtY\_\{t\}\. The optimization objective is \(5\)minϕ𝔼\(X,Y\)\[ℓ\(Bϕ\(X\),Y\)\],\\min\_\{\\phi\}\\;\\mathbb\{E\}\_\{\(X,Y\)\}\\left\[\\ell\(B\_\{\\phi\}\(X\),Y\)\\right\],whereℓ\\ellis the masked mean absolute error \(MAE\) computed on inverse\-scaled flow values\. The checkpoint with the best validation MAE is selected as the reference checkpoint for both the vanilla backbone and the subsequent RiskTraf stage\. ### 3\.3\.Stage II: Risk\-Aware Residual Plug\-in The second stage attaches RiskTraf to the trained backbone obtained in Stage I\. Although the Stage\-I backbone has access to all three variables, its ERM objective may exploit regime\-specific correlations between speed, occupancy, and flow\. RiskTraf therefore freezes the backbone predictionY^base\\hat\{Y\}\_\{\\mathrm\{base\}\}and learns only a constrained horizon\-wise residual correction, which keeps the plug\-in model\-agnostic and preserves the spatio\-temporal representations learned by the backbone\. RiskTraf consists of three components: a zero\-start residual head, risk environment construction from historical covariates, and a REx\-based objective for robust residual learning\. #### 3\.3\.1\.Zero\-Start Residual Head RiskTraf uses a lightweight residual head to correct the fixed backbone prediction\. Since traffic sensors are not exchangeable, the same speed–occupancy pattern may imply different flow corrections at different road segments due to location, lane configuration, detector type, and upstream–downstream context\. Motivated by recent studies on spatio\-temporal and heterogeneity\-aware embeddings\([32](https://arxiv.org/html/2608.20656#bib.bib20);[14](https://arxiv.org/html/2608.20656#bib.bib36)\), RiskTraf introduces a minimal node\-identity embedding to provide node\-specific context without modifying the frozen backbone\. Specifically, RiskTraf maintains a trainable node\-identity embedding tableE=\[e1,…,eN\]⊤∈ℝN×deE=\[e\_\{1\},\\ldots,e\_\{N\}\]^\{\\top\}\\in\\mathbb\{R\}^\{N\\times d\_\{e\}\}, whereene\_\{n\}denotes the embedding of sensor nodenn\. The table is initialized at the beginning of Stage II and optimized together withψ\\psi; it is not copied from the backbone and requires no external sensor metadata\. In addition, the head uses node\-level historical summariesQQ, including mean speed, mean occupancy, and occupancy\-minus\-speed over the input window\. The residual correction is computed as \(6\)ΔY^=gψ\(St−T:t−1,Ot−T:t−1,E,Q\),\\Delta\\hat\{Y\}=g\_\{\\psi\}\(S\_\{t\-T:t\-1\},O\_\{t\-T:t\-1\},E,Q\),whereEEandQQare broadcast to the batch and horizon dimensions as needed\. The last linear layer ofgψg\_\{\\psi\}is initialized to zero, so RiskTraf initially reproduces the baseline prediction and any nonzero correction must be learned during the debiasing stage\. #### 3\.3\.2\.Risk Environment Construction Risk environments are constructed only from historical speed and occupancy\. For each sample in a training batch, we summarize the covariates over the historical window and all nodes: \(7\)v¯\\displaystyle\\bar\{v\}=1TN∑t,nvt,n,o¯=1TN∑t,not,n,\\displaystyle=\\frac\{1\}\{TN\}\\sum\_\{t,n\}v\_\{t,n\},\\quad\\bar\{o\}=\\frac\{1\}\{TN\}\\sum\_\{t,n\}o\_\{t,n\},whereTTis the input length andNNis the number of sensor nodes\. Inspired by mean alignment\([18](https://arxiv.org/html/2608.20656#bib.bib10)\), we define a normalized traffic\-risk score as \(8\)z\(x\)=x−μxσx,ρ=z\(o¯\)−z\(v¯\),z\(x\)=\\frac\{x\-\\mu\_\{x\}\}\{\\sigma\_\{x\}\},\\quad\\rho=z\(\\bar\{o\}\)\-z\(\\bar\{v\}\),whereμx\\mu\_\{x\}andσx\\sigma\_\{x\}are estimated from the training data\. A largerρ\\rhoindicates higher occupancy and lower speed, corresponding to more congested or high\-risk traffic states\. This ordering is grounded in the fundamental diagram of traffic flow, where congestion is characterized jointly by low speed and high occupancy, soρ\\rhoranks samples along the physical free\-flow–congestion axis\. Samples are sorted byρ\\rhoand partitioned intoKKordered environments as follows: \(9\)ei=\{\(Xj,Yj\)∣ρj∈quantilei\},i=1,…,K\.e^\{i\}=\\\{\(X\_\{j\},Y\_\{j\}\)\\mid\\rho\_\{j\}\\in\\text\{quantile\}\_\{i\}\\\},\\quad i=1,\\ldots,K\.This produces low\-to\-high traffic\-risk environments without requiring incident annotations or manually defined congestion labels\. #### 3\.3\.3\.REx\-Based Residual Learning Given the ordered risk environments, RiskTraf trains the residual head to improve flow prediction while keeping residual errors stable across traffic regimes\. For each environmenteie^\{i\}, we compute \(10\)ℒei=1\|ei\|∑\(X,Y\)∈eiℓ\(Y^base\+sr⋅gψ\(S,O,E,Q\),Y\)\.\\mathcal\{L\}^\{e^\{i\}\}=\\frac\{1\}\{\|e^\{i\}\|\}\\sum\_\{\(X,Y\)\\in e^\{i\}\}\\ell\(\\hat\{Y\}\_\{\\text\{base\}\}\+s\_\{r\}\\cdot g\_\{\\psi\}\(S,O,E,Q\),Y\)\.Letℒ¯=1K∑iℒei\\bar\{\\mathcal\{L\}\}=\\frac\{1\}\{K\}\\sum\_\{i\}\\mathcal\{L\}^\{e^\{i\}\}\. The REx penalty measures the variance of environment losses: \(11\)ℛREx=1K∑i=1K\(ℒei−ℒ¯\)2\.\\mathcal\{R\}\_\{\\text\{REx\}\}=\\frac\{1\}\{K\}\\sum\_\{i=1\}^\{K\}\(\\mathcal\{L\}^\{e^\{i\}\}\-\\bar\{\\mathcal\{L\}\}\)^\{2\}\. The full Stage\-II objective is \(12\)ℒtotal=ℒ¯\+λτ\(ℛREx\+β⋅𝒫pair\+η⋅𝒫extra\),\\mathcal\{L\}\_\{\\text\{total\}\}=\\bar\{\\mathcal\{L\}\}\+\\lambda\_\{\\tau\}\\left\(\\mathcal\{R\}\_\{\\text\{REx\}\}\+\\beta\\cdot\\mathcal\{P\}\_\{\\text\{pair\}\}\+\\eta\\cdot\\mathcal\{P\}\_\{\\text\{extra\}\}\\right\),whereβ\\betaandη\\etacontrol the relative weights of the two order\-aware penalties, defined on the ordered environment losses as \(13\)𝒫pair\\displaystyle\\mathcal\{P\}\_\{\\text\{pair\}\}=1K−1∑i=1K−1\[max\(0,ℒei\+1−ℒei\)\]2,\\displaystyle=\\frac\{1\}\{K\-1\}\\sum\_\{i=1\}^\{K\-1\}\\Big\[\\max\\big\(0,\\,\\mathcal\{L\}^\{e^\{i\+1\}\}\-\\mathcal\{L\}^\{e^\{i\}\}\\big\)\\Big\]^\{2\},𝒫extra\\displaystyle\\mathcal\{P\}\_\{\\text\{extra\}\}=\[max\(0,ℒeK−sg\(ℒe1\)\)\]2,\\displaystyle=\\Big\[\\max\\big\(0,\\,\\mathcal\{L\}^\{e^\{K\}\}\-\\operatorname\{sg\}\(\\mathcal\{L\}^\{e^\{1\}\}\)\\big\)\\Big\]^\{2\},wheresg\(⋅\)\\operatorname\{sg\}\(\\cdot\)denotes the stop\-gradient operator\.𝒫pair\\mathcal\{P\}\_\{\\text\{pair\}\}penalizes any adjacent pair whose higher\-risk environment incurs the larger residual loss, so that errors do not grow with the risk level, and𝒫extra\\mathcal\{P\}\_\{\\text\{extra\}\}penalizes the highest\-risk loss when it exceeds the detached lowest\-risk reference, pushing down only the former\. The penalty weight is warmed up as \(14\)λτ=λ⋅min\(1,max\(0,τ−τwarmupτtotal−τwarmup\)\),\\lambda\_\{\\tau\}=\\lambda\\cdot\\min\\left\(1,\\max\\left\(0,\\frac\{\\tau\-\\tau\_\{\\text\{warmup\}\}\}\{\\tau\_\{\\text\{total\}\}\-\\tau\_\{\\text\{warmup\}\}\}\\right\)\\right\),whereτ\\taudenotes the current Stage\-II epoch\. This warmup allows the residual head to first learn useful corrections before enforcing stronger cross\-environment consistency\. #### 3\.3\.4\.Validation Safeguard Since RiskTraf is designed as a plug\-in correction, it should not degrade a strong backbone\. We therefore evaluate the Stage\-II checkpoint on the validation set and keep the corrected model only when it improves validation MAE over the frozen baseline\. Otherwise, RiskTraf rolls back to the Stage\-I prediction\. This keeps the plug\-in conservative: it improves the backbone when useful residual patterns exist and avoids harmful corrections otherwise\. ### 3\.4\.Inference At inference, the selected Stage\-I backbone first produces the baseline forecastY^base\\hat\{Y\}\_\{\\text\{base\}\}\. If the Stage\-II residual head passes the validation safeguard, RiskTraf computesΔY^\\Delta\\hat\{Y\}from historical speed and occupancy, node embeddings, and node\-level summaries\. The final prediction is \(15\)Y^risk=Y^base\+sr⋅ΔY^\.\\hat\{Y\}\_\{\\text\{risk\}\}=\\hat\{Y\}\_\{\\text\{base\}\}\+s\_\{r\}\\cdot\\Delta\\hat\{Y\}\.If the residual head is rolled back, the output remains the Stage\-I baseline prediction\. ## 4\.Experiments We conduct experiments on the PEMSB\-3V benchmark suite introduced in Section[2](https://arxiv.org/html/2608.20656#S2)\. The experiments are designed to evaluate whether the proposed benchmark provides physically meaningful three\-variable traffic data, and whether RiskTraf can improve traffic flow prediction as a model\-agnostic residual plug\-in\. We organize the evaluation around the following research questions: - •RQ1: Data validity\.Does PEMSB\-3V exhibit physically meaningful flow–speed–occupancy relationships? - •RQ2: Overall effectiveness\.Can RiskTraf consistently improve diverse spatio\-temporal forecasting backbones? - •RQ3: Method comparison\.How does RiskTraf compare with existing debiasing and distribution\-shift adaptation methods? - •RQ4: Component analysis\.How do auxiliary variables, risk penalties, and environment granularity affect RiskTraf? - •RQ5: Practicality and interpretability\.What overhead does RiskTraf introduce, and what corrections does it learn under ambiguous traffic states? ### 4\.1\.Experimental Setup ##### Datasets and forecasting setting\. We evaluate RiskTraf on four PEMSB\-3V subsets: PEMS03\-B, PEMS04\-B, PEMS07\-B, and PEMS08\-B\. All experiments follow the same three\-variable\-to\-flow forecasting setting\. The input contains 12 historical time steps of flow, speed, and occupancy, corresponding to one hour of observations, and the target contains the next 12 time steps of flow\. All metrics are computed on inverse\-scaled flow values\. We report MAE, RMSE, and MAPE, where lower values indicate better performance\. ##### Backbones\. To evaluate model\-agnostic applicability, we apply RiskTraf to representative spatio\-temporal forecasting backbones from different architectural generations, including STGCN\([48](https://arxiv.org/html/2608.20656#bib.bib7)\), DCRNN\([31](https://arxiv.org/html/2608.20656#bib.bib18)\), AGCRN\([4](https://arxiv.org/html/2608.20656#bib.bib30)\), Graph WaveNet\([46](https://arxiv.org/html/2608.20656#bib.bib8)\), GMAN\([51](https://arxiv.org/html/2608.20656#bib.bib31)\), GTS\([38](https://arxiv.org/html/2608.20656#bib.bib34)\), STEMGNN\([5](https://arxiv.org/html/2608.20656#bib.bib32)\), STNorm\([13](https://arxiv.org/html/2608.20656#bib.bib33)\), STWA\([17](https://arxiv.org/html/2608.20656#bib.bib35)\), MegaCRN\([26](https://arxiv.org/html/2608.20656#bib.bib9)\), HimNet\([14](https://arxiv.org/html/2608.20656#bib.bib36)\), and STDN\([6](https://arxiv.org/html/2608.20656#bib.bib29)\)\. These backbones cover recurrent, graph\-based, attention\-based, normalization\-based, and recent heterogeneity\-aware designs\. ##### Paired plug\-in protocol\. For each dataset–backbone pair, we first train a vanilla three\-variable backbone using standard empirical risk minimization\. RiskTraf then starts from the same validation\-selected checkpoint, freezes the backbone, and trains only the lightweight residual head driven by historical speed and occupancy\. This paired protocol ensures that any improvement comes from the RiskTraf plug\-in rather than from a different backbone initialization or training recipe\. Table 2\.Main plug\-in results on PEMSB\-3V, averaged over horizons 3, 6, and 12\. Lower is better; MAPE is in percentage form\. Green/red arrows mark improvement/degradation over the backbone row\. ##### Implementation details\. For RiskTraf, each training batch is sorted by the speed–occupancy risk score and partitioned into four ordered traffic\-risk environments unless otherwise specified\. The residual head is trained with masked MAE and the warmed\-up risk extrapolation penalty described in Section[3\.3](https://arxiv.org/html/2608.20656#S3.SS3)\. Baseline backbones and compared methods are implemented using official code when available or reproduced following the original papers, with hyper\-parameters selected by validation MAE under the same protocol\. We use Adam for optimization, validation MAE for model selection, and rollback to the Stage\-I checkpoint when the plug\-in does not improve validation MAE\. The batch size is 32 unless otherwise specified\. All experiments are conducted on NVIDIA L20 GPUs\. Figure 3\.Flow–speed relationships under different occupancy states on the four PEMSB\-3V subsets\. Samples are grouped into low\-, mid\-, and high\-occupancy states; each curve shows the average speed at different flow levels\.Line plots showing that higher occupancy states correspond to lower speeds under similar traffic flow levels on PEMS03\-B and PEMS04\-B\.Figure 4\.Per\-horizon MAE decomposition for GTS, GWNet, and STNorm on the four PEMSB\-3V subsets\. For each horizon, the full bar denotes the baseline MAE, the solid segment denotes the MAE after adding RiskTraf, and the hatched cap denotes the absolute error reduced by RiskTraf\.Stacked bar charts for PEMS03\-B, PEMS04\-B, PEMS07\-B, and PEMS08\-B showing baseline MAE, RiskTraf MAE, and reduced MAE across prediction horizons 1 to 12 for representative backbones\. ### 4\.2\.Dataset Characteristics \(RQ1\) We first examine whether PEMSB\-3V preserves meaningful relationships among flow, speed, and occupancy\. Figure[3](https://arxiv.org/html/2608.20656#S4.F3)plots flow–speed curves under low\-, mid\-, and high\-occupancy states\. Across the four displayed subsets, higher occupancy generally corresponds to lower speed, which is consistent with the loop\-detector mechanism that vehicles occupy detectors for longer durations under congested conditions\. For the same flow level, speed can vary substantially across occupancy states, indicating that flow alone is insufficient to identify the underlying traffic regime\. The figure also shows that occupancy is not a trivial proxy for flow\. High\-occupancy states span a wide range of flow levels and exhibit distinct speed–flow patterns, rather than simply corresponding to the highest flow values\. Moreover, the slope and shape of the flow–speed curves differ across occupancy groups and benchmark subsets, suggesting regime\-dependent auxiliary correlations\. These observations support the use of speed and occupancy as informative traffic\-state indicators, while motivating RiskTraf to exploit them through risk\-aware residual learning rather than direct unconstrained fusion\. ### 4\.3\.Overall Performance \(RQ2\) Table[2](https://arxiv.org/html/2608.20656#S4.T2)reports MAE, RMSE, and MAPE averaged over horizons 3, 6, and 12\. For each backbone, the vanilla row denotes the Stage\-I model trained with historical flow, speed, and occupancy, while the \+RiskTraf row denotes the paired residual plug\-in result\. The results lead to the following observations\. - •RiskTraf consistently improves heterogeneous backbones\.Across the displayed backbones and PEMSB\-3V subsets, adding RiskTraf reduces MAE and RMSE in nearly all cases, supporting its model\-agnostic plug\-in design\. - •Larger gains appear on backbones more sensitive to shifted auxiliary correlations\.DCRNN\([31](https://arxiv.org/html/2608.20656#bib.bib18)\)and GTS\([38](https://arxiv.org/html/2608.20656#bib.bib34)\)obtain particularly large improvements, suggesting that risk\-aware residual learning is especially useful when direct three\-variable ERM is unstable\. - •Strong recent backbones still benefit from the plug\-in\.Competitive models such as STWA\([17](https://arxiv.org/html/2608.20656#bib.bib35)\), MegaCRN\([26](https://arxiv.org/html/2608.20656#bib.bib9)\), HimNet\([14](https://arxiv.org/html/2608.20656#bib.bib36)\), and STDN\([6](https://arxiv.org/html/2608.20656#bib.bib29)\)also improve with RiskTraf, showing that the plug\-in complements rather than replaces advanced backbone architectures\. - •MAPE degradation reflects denominator sensitivity\.A few worse MAPE results do not contradict the MAE/RMSE gains, since MAPE is unstable for zero or near\-zero actual values\([23](https://arxiv.org/html/2608.20656#bib.bib52)\)and behaves like MAE weighted by inverse target magnitude\([12](https://arxiv.org/html/2608.20656#bib.bib53)\)\. Thus, better abnormal\-regime corrections can still be penalized by small\-flow denominators\. Figure[4](https://arxiv.org/html/2608.20656#S4.F4)further examines whether the improvements persist at individual prediction horizons\. Each bar decomposes the baseline MAE into the remaining error after applying RiskTraf and the absolute error reduced by the plug\-in, where the hatched cap directly represents the removed MAE\. Across all four subsets, RiskTraf reduces errors on most horizons, showing that the gains are not caused by averaging a few favorable steps\. The improvement is especially visible for GTS\([38](https://arxiv.org/html/2608.20656#bib.bib34)\), while GWNet\([46](https://arxiv.org/html/2608.20656#bib.bib8)\)and STNorm\([13](https://arxiv.org/html/2608.20656#bib.bib33)\)obtain smaller but still consistent reductions\. This suggests that RiskTraf provides larger corrections for backbones with larger horizon\-wise errors, while still refining more stable backbones\. The reduction pattern is also horizon\-dependent\. On PEMS04\-B, PEMS07\-B, and PEMS08\-B, the hatched caps often become more pronounced from short to medium and long horizons, although some subsets also show clear short\-horizon corrections when the baseline is unstable\. This pattern is consistent with the increasing uncertainty of long\-range forecasting: near\-term flow is constrained by recent observations, whereas later horizons require recognizing the underlying traffic regime from speed and occupancy\. Therefore, the per\-horizon decomposition supports that RiskTraf does not simply apply a uniform offset, but learns residual corrections that are more useful when regime\-dependent uncertainty accumulates\. Table 3\.Comparison with robust traffic forecasting methods on PEMSB\-3V\. Results are averaged over horizons 3, 6, and 12\. The first four methods are standalone robustness methods, while the last column reports RiskTraf plugged into the strongest backbone from Table[2](https://arxiv.org/html/2608.20656#S4.T2)\. Best results are in bold\. ### 4\.4\.Comparison with Robust Forecasting Methods \(RQ3\) Table 4\.Ablation study on auxiliary measurements used for traffic\-risk environment construction\. Results are averaged over horizons 3, 6, and 12; MAPE is reported in percentage form\. The best result within each backbone–dataset block is in bold\.To answer RQ3, we compare RiskTraf with representative robust traffic forecasting methods, including debiasing\-based methods CauSTG\([52](https://arxiv.org/html/2608.20656#bib.bib1)\)and ST\-SSDL\([18](https://arxiv.org/html/2608.20656#bib.bib10)\), and distribution\-shift adaptation methods Dish\-TS\([16](https://arxiv.org/html/2608.20656#bib.bib11)\)and STEVE\([25](https://arxiv.org/html/2608.20656#bib.bib12)\)\. Since these methods are standalone forecasting frameworks rather than backbone models, Table[3](https://arxiv.org/html/2608.20656#S4.T3)reports their all\-horizon MAE/RMSE separately from the paired plug\-in comparison in Table[2](https://arxiv.org/html/2608.20656#S4.T2)\. We include STDN\([6](https://arxiv.org/html/2608.20656#bib.bib29)\)\+RiskTraf as the representative RiskTraf configuration, as STDN is the strongest backbone in the main comparison\. RiskTraf obtains the best MAE and RMSE on all four PEMSB\-3V subsets\. Compared with the strongest standalone competitor on each subset, it reduces MAE by 8\.75%, 5\.12%, 13\.68%, and 6\.09%, and reduces RMSE by 7\.99%, 3\.51%, 7\.87%, and 1\.04% on PEMS03\-B, PEMS04\-B, PEMS07\-B, and PEMS08\-B, respectively\. The consistent improvements on both MAE and RMSE indicate that RiskTraf reduces not only average prediction bias but also larger forecasting deviations\. Moreover, existing robust methods exhibit dataset\-dependent rankings, suggesting that their robustness mechanisms may be sensitive to district\-level traffic dynamics\. In contrast, RiskTraf remains consistently strong across subsets while preserving the original backbone architecture\. This supports the effectiveness of risk\-aware residual correction for handling regime\-dependent auxiliary correlations without replacing the forecasting model\. ### 4\.5\.Component and Sensitivity Analysis \(RQ4\) ##### Auxiliary\-variable ablation\. Table[4](https://arxiv.org/html/2608.20656#S4.T4)ablates the auxiliary measurements used for traffic\-risk environment construction while keeping the backbone, residual head, and training protocol unchanged\. Thew/o speedvariant constructs environments using occupancy only, thew/o speed&occvariant removes both speed and occupancy from environment partitioning, and the full variant uses both measurements\. The results show that occupancy alone already provides useful traffic\-state information:w/o speedconsistently outperformsw/o speed&occon most backbone–dataset pairs\. This is expected because occupancy directly reflects detector occupation intensity and is closely related to congestion states\. However, the fullspeed\+occvariant achieves the best overall performance across all backbone blocks, showing that speed provides complementary information about vehicle movement dynamics\. The improvement is particularly clear for DCRNN, while stronger backbones such as GWNet and STWA obtain smaller but still consistent gains\. These results confirm that RiskTraf benefits from constructing traffic\-risk environments with both auxiliary measurements\. Figure 5\.Sensitivity analysis of RiskTraf on PEMS03\-B with respect to the risk penalty weightλ\\lambdaand the number of risk environmentsKK\. Stars mark the lowest MAE in each sweep\. ##### Sensitivity toλ\\lambdaandKK\. Figure[5](https://arxiv.org/html/2608.20656#S4.F5)studies the sensitivity of RiskTraf on PEMS03\-B by varying the risk penalty weightλ\\lambdaand the number of environmentsKK, while keepingβ\\betaandη\\etafixed\. The weightλ\\lambdacontrols the strength of the REx, pairwise, and extrapolation regularizers\. A smallλ\\lambdaweakens cross\-environment constraints, whereas an overly largeλ\\lambdamay sacrifice pointwise accuracy for environment consistency\. The best performance is achieved aroundλ=0\.05\\lambda=0\.05\. The environment numberKKcontrols the granularity of traffic\-risk partitioning\. The non\-monotonic curve shows that environment construction is not simply improved by using more partitions\. AlthoughK=3K=3obtains the lowest MAE in this PEMS03\-B sweep, we useK=4K=4in the main experiments as a fixed quartile split for consistent evaluation across all dataset–backbone pairs\. This avoids tuningKKseparately for each setting\. Figure 6\.Targeted STAEformer forecasting cases on PEMS04\-B and PEMS08\-B\. The black curve denotes observed flow, the red dashed curve denotes the baseline prediction, and the blue curve denotes the prediction after adding RiskTraf\. The vertical dashed line marks the prediction start, and the shaded region highlights the residual correction\. ### 4\.6\.Efficiency Study \(RQ5\) Table 5\.Plug\-in efficiency comparison on PEMS03\-B with batch size 32\. Each pair compares a backbone with the same backbone equipped with RiskTraf\. Training time denotes the observed wall\-clock time of the recorded run, where RiskTraf trains only the residual head with the backbone frozen\.Table[5](https://arxiv.org/html/2608.20656#S4.T5)evaluates RiskTraf from a plug\-in efficiency perspective\. Since RiskTraf is not designed as a standalone forecasting architecture, the relevant comparison is between each backbone and the same backbone equipped with RiskTraf\. Across GTS\([38](https://arxiv.org/html/2608.20656#bib.bib34)\), GWNet\([46](https://arxiv.org/html/2608.20656#bib.bib8)\), and STNorm\([13](https://arxiv.org/html/2608.20656#bib.bib33)\)on PEMS03\-B, RiskTraf introduces only about 0\.004M additional trainable parameters and marginal GFLOP increases\. The inference latency and inference memory remain in the same practical range, indicating that the residual head adds limited deployment overhead\. The small overhead is accompanied by clear accuracy gains\. Test MAE decreases from 28\.45 to 21\.45 for GTS, from 19\.65 to 17\.43 for GWNet, and from 17\.31 to 15\.78 for STNorm\. These improvements show that the auxiliary risk environments and residual correction provide a favorable accuracy–efficiency trade\-off\. Training wall\-clock time is reported as an auxiliary measurement, since it is affected by Stage\-II early stopping and by the fact that only the residual head is optimized while the backbone is frozen\. Figure 7\.RiskTraf correction over STAEformer on PEMS03\-B\. Samples are grouped by speed and occupancy bins; each cell reportsmean\(Y^RiskTraf−Y^Base\)\\mathrm\{mean\}\(\\hat\{Y\}\_\{\\mathrm\{RiskTraf\}\}\-\\hat\{Y\}\_\{\\mathrm\{Base\}\}\), where larger positive values indicate stronger upward residual corrections\. ### 4\.7\.Case Study \(RQ5\) ##### Trajectory\-level correction\. Figure[6](https://arxiv.org/html/2608.20656#S4.F6)visualizes representative STAEformer forecasting cases on PEMS04\-B and PEMS08\-B\. The baseline sometimes over\-reacts or under\-reacts after the prediction start, leading to trajectories that deviate from the observed flow\. RiskTraf adjusts the baseline prediction toward the observed trajectory in these cases, especially around turning points or short\-term trend changes after the prediction boundary\. This shows that the residual head does not simply apply a constant shift, but can produce horizon\-wise corrections that reshape the forecast trajectory\. ##### State\-dependent correction\. Figure[7](https://arxiv.org/html/2608.20656#S4.F7)further groups samples by speed and occupancy states and reports the average correctionmean\(Y^RiskTraf−Y^Base\)\\mathrm\{mean\}\(\\hat\{Y\}\_\{\\mathrm\{RiskTraf\}\}\-\\hat\{Y\}\_\{\\mathrm\{Base\}\}\)\. The correction is positive in all bins, indicating that RiskTraf generally raises the STAEformer baseline in the selected PEMS03\-B setting\. More importantly, the correction magnitude is strongly associated with occupancy: low\-occupancy states receive relatively small corrections, whereas high\-occupancy states receive much larger corrections\. In contrast, the variation along the speed axis is less monotonic\. This pattern suggests that occupancy provides a stronger congestion\-state signal for the residual head, while speed helps refine the correction within each occupancy regime\. Figure 8\.t\-SNE visualization of PEMS03\-B samples colored by RiskTraf risk environments\. The three colors denote ordered environments constructed from historical speed and occupancy, showing separated regimes with overlapping transition regions\. ##### Risk\-environment structure\. Figure[8](https://arxiv.org/html/2608.20656#S4.F8)visualizes the risk\-environment partition on PEMS03\-B using t\-SNE\. The three environments form a coarse low\-to\-high risk organization with visible separation, while still retaining overlap in transition regions\. The separated regions suggest that the speed–occupancy risk score captures distinct traffic regimes, whereas the overlapping regions correspond to samples whose raw traffic patterns are less clearly separable\. This structure is consistent with the motivation of REx\-based residual learning: RiskTraf does not require perfectly separated environments, but encourages the residual head to remain stable across related risk regimes while still adapting to high\-risk states\. Overall, these case studies provide qualitative evidence for the mechanism behind RiskTraf\. The plug\-in changes individual trajectories, applies larger corrections under high\-occupancy traffic states, and relies on risk environments that reflect meaningful structure in the raw traffic data\. ## 5\.Related Work ### 5\.1\.Traffic Flow Prediction Early deep approaches relied on recurrent architectures such as LSTM, which are limited in capturing complex spatial dependencies of road networks; later work turned to graph\-based designs: DCRNN\([31](https://arxiv.org/html/2608.20656#bib.bib18)\)combines diffusion convolutions with sequence\-to\-sequence learning, and STGCN\([48](https://arxiv.org/html/2608.20656#bib.bib7)\)adopts a fully convolutional graph architecture, both relying on predefined adjacency matrices from road topology\. Graph WaveNet\([46](https://arxiv.org/html/2608.20656#bib.bib8)\)instead learns adaptive graphs for latent spatial dependencies, and MegaCRN\([26](https://arxiv.org/html/2608.20656#bib.bib9)\)adds memory and meta\-learning for heterogeneous patterns\. STAEformer\([32](https://arxiv.org/html/2608.20656#bib.bib20)\)showed that Transformers with spatio\-temporal adaptive embeddings perform strongly, STDN\([6](https://arxiv.org/html/2608.20656#bib.bib29)\)and HimNet\([14](https://arxiv.org/html/2608.20656#bib.bib36)\)explored seasonality and heterogeneity, and ST\-SSDL\([18](https://arxiv.org/html/2608.20656#bib.bib10)\)aligns inputs with historical means to reduce stochastic deviations\. As architectural gains become marginal\([40](https://arxiv.org/html/2608.20656#bib.bib21)\), some works turn to richer sources such as event information \(EastNet\([44](https://arxiv.org/html/2608.20656#bib.bib37)\)\) or global\-scale multimodal data \(Terra\([9](https://arxiv.org/html/2608.20656#bib.bib38)\)\), which require additional collection and alignment\. In contrast, speed and occupancy are sensor\-native measurements collected with flow, yet no existing benchmark systematically studies these raw auxiliary measurements for robust flow prediction; PEMSB\-3V fills this gap\. ### 5\.2\.Invariant Learning and Causal Inference Invariant learning assumes that stable mechanisms remain predictive across environments while spurious correlations vary\. IRM\([2](https://arxiv.org/html/2608.20656#bib.bib39)\)seeks representations that admit a shared optimal classifier across environments, IRM Games\([1](https://arxiv.org/html/2608.20656#bib.bib40)\)recasts this as Nash equilibrium finding, REx\([28](https://arxiv.org/html/2608.20656#bib.bib41)\)minimizes the variance of risks across environments to improve extrapolation, and Group DRO\([37](https://arxiv.org/html/2608.20656#bib.bib42)\)minimizes worst\-group risk for subgroup robustness\. Their effectiveness, however, hinges on how environments are defined: DomainBed\([20](https://arxiv.org/html/2608.20656#bib.bib43)\)shows limited gains over empirical risk minimization when partitions do not align with the relevant spurious correlations\. This is central in traffic prediction, where environments are not directly given\. RiskTraf therefore builds ordered risk environments from historical speed and occupancy—physically interpretable indicators of traffic regimes—and applies REx\-style regularization to a residual correction head\. ### 5\.3\.Distribution Shift in Time Series Non\-stationarity between training and test periods is commonly handled by statistical normalization: RevIN\([27](https://arxiv.org/html/2608.20656#bib.bib44)\)removes instance\-level statistics before prediction and restores them afterward, Dish\-TS\([16](https://arxiv.org/html/2608.20656#bib.bib11)\)models shifts through learnable transformations, and Non\-stationary Transformers\([34](https://arxiv.org/html/2608.20656#bib.bib45)\)introduce de\-stationary attention against over\-stationarization, while the multi\-order wavelet derivative transform\([54](https://arxiv.org/html/2608.20656#bib.bib54)\)captures evolving non\-stationary dynamics in the wavelet domain\. These methods target shifts in the statistical or spectral properties of the series itself\. The challenge studied here is different: the relationships among flow, speed, and occupancy themselves vary across traffic regimes, so treating speed and occupancy as ordinary covariates may introduce shortcut learning\. RiskTraf instead uses them to define traffic\-risk environments and learns a constrained residual correction that remains stable across these environments\. ## 6\.Conclusion and Future Works In this paper, we study how to effectively use speed and occupancy for traffic flow prediction\. While these sensor\-native variables provide useful traffic\-state information, direct three\-variable training can exploit regime\-dependent correlations and hurt generalization\. To address this, we propose RiskTraf, a paired plug\-in that freezes a trained three\-variable backbone and learns a lightweight residual correction from historical speed and occupancy under a REx objective\. By constructing traffic\-risk environments, RiskTraf performs state\-dependent flow correction on the frozen backbone\. We also introduce PEMSB\-3V, a public multi\-variate traffic benchmark with raw flow, speed, and occupancy measurements\. Experiments show that RiskTraf consistently improves diverse spatio\-temporal backbones and outperforms robust forecasting methods\. Future work may explore batch\-independent REx objectives and extend RiskTraf to other spatio\-temporal prediction tasks with auxiliary measurements\. ###### Acknowledgements\. This work was supported in part by National Natural Science Foundations of China under Grant No\. 62572416 and the Guangdong Provincial Key Lab of Integrated Communication, Sensing and Computation for Ubiquitous Internet of Things under Grant No\. 2023B1212010007\. ## GenAI Usage Disclosure Generative AI tools were used only to assist with language polishing, wording refinement, and formatting during manuscript preparation\. The research ideas, method design, experiments, analyses, figures, tables, and conclusions were produced and verified by the authors, who take full responsibility for the content of this paper\. ## References - Ahujaet al\.\(2020\)K\. Ahuja, K\. Shanmugam, K\. R\. Varshney, and A\. DhurandharInvariant risk minimization games\.InProceedings of the 37th International Conference on Machine Learning,ICML’20\.Cited by:[§5\.2](https://arxiv.org/html/2608.20656#S5.SS2.p1.1)\. - Arjovskyet al\.\(2020\)M\. Arjovsky, L\. Bottou, I\. Gulrajani, and D\. Lopez\-PazInvariant risk minimization\.External Links:1907\.02893,[Link](https://arxiv.org/abs/1907.02893)Cited by:[§5\.2](https://arxiv.org/html/2608.20656#S5.SS2.p1.1)\. - Ash \(2012\)R\. B\. AshInformation theory\.Courier Corporation\.Cited by:[§1](https://arxiv.org/html/2608.20656#S1.p3.1)\. - Baiet al\.\(2020\)L\. Bai, L\. Yao, C\. Li, X\. Wang, and C\. WangAdaptive graph convolutional recurrent network for traffic forecasting\.Advances in Neural Information Processing Systems33,pp\. 17804–17815\.Cited by:[§4\.1](https://arxiv.org/html/2608.20656#S4.SS1.SSS0.Px2.p1.1)\. - Caoet al\.\(2020\)D\. Cao, Y\. Wang, J\. Duan, C\. Zhang, X\. Zhu, C\. Huang, Y\. Tong, B\. Xu, J\. Bai, J\. Tong,et al\.Spectral temporal graph neural network for multivariate time\-series forecasting\.Advances in Neural Information Processing Systems33,pp\. 17766–17778\.Cited by:[§4\.1](https://arxiv.org/html/2608.20656#S4.SS1.SSS0.Px2.p1.1)\. - Caoet al\.\(2025\)L\. Cao, B\. Wang, G\. Jiang, Y\. Yu, and J\. DongSpatiotemporal\-aware trend\-seasonality decomposition network for traffic flow forecasting\.Proceedings of the AAAI Conference on Artificial Intelligence39\(11\),pp\. 11463–11471\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/33247),[Document](https://dx.doi.org/10.1609/aaai.v39i11.33247)Cited by:[3rd item](https://arxiv.org/html/2608.20656#S4.I2.i3.p1.1),[§4\.1](https://arxiv.org/html/2608.20656#S4.SS1.SSS0.Px2.p1.1),[§4\.4](https://arxiv.org/html/2608.20656#S4.SS4.p1.1),[§5\.1](https://arxiv.org/html/2608.20656#S5.SS1.p1.1)\. - Chenet al\.\(2003\)C\. Chen, J\. Kwon, J\. Rice, A\. Skabardonis, and P\. VaraiyaDetecting errors and imputing missing data for single\-loop surveillance systems\.Transportation Research Record1855\(1\),pp\. 160–167\.External Links:[Document](https://dx.doi.org/10.3141/1855-20)Cited by:[§1](https://arxiv.org/html/2608.20656#S1.p5.1)\. - Chenet al\.\(2001\)C\. Chen, K\. Petty, A\. Skabardonis, P\. Varaiya, and Z\. JiaFreeway performance measurement system: mining loop detector data\.Transportation Research Record1748\(1\),pp\. 96–102\.External Links:[Document](https://dx.doi.org/10.3141/1748-12)Cited by:[§1](https://arxiv.org/html/2608.20656#S1.p1.1)\. - Chenet al\.\(2024\)W\. Chen, X\. Hao, Y\. Wu, and Y\. LiangTerra: a multimodal spatio\-temporal dataset spanning the earth\.InThe Thirty\-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track,External Links:[Link](https://openreview.net/forum?id=I0zpivK0A0)Cited by:[§5\.1](https://arxiv.org/html/2608.20656#S5.SS1.p1.1)\. - Chen and Liang \(2025\)W\. Chen and Y\. LiangExpand and compress: exploring tuning principles for continual spatio\-temporal graph forecasting\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=FRzCIlkM7I)Cited by:[§1](https://arxiv.org/html/2608.20656#S1.p2.1)\. - Chenet al\.\(2021\)X\. Chen, J\. Wang, and K\. XieTrafficStream: a streaming traffic flow forecasting framework based on graph neural networks and continual learning\.InProceedings of the Thirtieth International Joint Conference on Artificial Intelligence,pp\. 3620–3626\.Cited by:[§1](https://arxiv.org/html/2608.20656#S1.p6.1)\. - de Myttenaereet al\.\(2016\)A\. de Myttenaere, B\. Golden, B\. Le Grand, and F\. RossiMean absolute percentage error for regression models\.Neurocomputing192,pp\. 38–48\.Note:Advances in artificial neural networks, machine learning and computational intelligenceExternal Links:ISSN 0925\-2312,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.neucom.2015.12.114),[Link](https://www.sciencedirect.com/science/article/pii/S0925231216003325)Cited by:[4th item](https://arxiv.org/html/2608.20656#S4.I2.i4.p1.1)\. - Denget al\.\(2021\)J\. Deng, X\. Chen, R\. Jiang, X\. Song, and I\. W\. TsangST\-norm: spatial and temporal normalization for multi\-variate time series forecasting\.InProceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining,pp\. 269–278\.Cited by:[§4\.1](https://arxiv.org/html/2608.20656#S4.SS1.SSS0.Px2.p1.1),[§4\.3](https://arxiv.org/html/2608.20656#S4.SS3.p4.1),[§4\.6](https://arxiv.org/html/2608.20656#S4.SS6.p1.1)\. - Donget al\.\(2024\)Z\. Dong, R\. Jiang, H\. Gao, H\. Liu, J\. Deng, Q\. Wen, and X\. SongHeterogeneity\-informed meta\-parameter learning for spatiotemporal time series forecasting\.InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining,pp\. 631–641\.Cited by:[§3\.3\.1](https://arxiv.org/html/2608.20656#S3.SS3.SSS1.p1.1),[3rd item](https://arxiv.org/html/2608.20656#S4.I2.i3.p1.1),[§4\.1](https://arxiv.org/html/2608.20656#S4.SS1.SSS0.Px2.p1.1),[§5\.1](https://arxiv.org/html/2608.20656#S5.SS1.p1.1)\. - Dunne and Ghosh \(2013\)S\. Dunne and B\. GhoshWeather adaptive traffic prediction using neurowavelet models\.IEEE Transactions on Intelligent Transportation Systems14\(1\),pp\. 370–379\.Cited by:[§1](https://arxiv.org/html/2608.20656#S1.p2.1)\. - Fanet al\.\(2023\)W\. Fan, P\. Wang, D\. Wang, D\. Wang, Y\. Zhou, and Y\. FuDish\-ts: a general paradigm for alleviating distribution shift in time series forecasting\.InProceedings of the Thirty\-Seventh AAAI Conference on Artificial Intelligence,AAAI’23\.External Links:[Link](https://doi.org/10.1609/aaai.v37i6.25914),[Document](https://dx.doi.org/10.1609/aaai.v37i6.25914)Cited by:[§1](https://arxiv.org/html/2608.20656#S1.p6.1),[§4\.4](https://arxiv.org/html/2608.20656#S4.SS4.p1.1),[§5\.3](https://arxiv.org/html/2608.20656#S5.SS3.p1.1)\. - Fanget al\.\(2024\)Y\. Fang, Y\. Liang, B\. Hui, Z\. Shao, L\. Deng, X\. Liu, X\. Jiang, and K\. ZhengEfficient large\-scale traffic forecasting with transformers: a spatial data management perspective\.arXiv preprint arXiv:2412\.09972\.Cited by:[3rd item](https://arxiv.org/html/2608.20656#S4.I2.i3.p1.1),[§4\.1](https://arxiv.org/html/2608.20656#S4.SS1.SSS0.Px2.p1.1)\. - Gaoet al\.\(2025\)H\. Gao, Z\. Dong, J\. Yong, S\. Fukushima, K\. Taura, and R\. JiangHow different from the past? spatio\-temporal time series forecasting with self\-supervised deviation learning\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=TgGH1bY6kl)Cited by:[§3\.3\.2](https://arxiv.org/html/2608.20656#S3.SS3.SSS2.p1.2),[§4\.4](https://arxiv.org/html/2608.20656#S4.SS4.p1.1),[§5\.1](https://arxiv.org/html/2608.20656#S5.SS1.p1.1)\. - Gouet al\.\(2026\)X\. Gou, Z\. Li, T\. Lan, J\. Lin, Z\. Li, B\. Zhao, C\. Zhang, D\. Wang, and X\. ZhangTraffiDent: a dataset for understanding the interplay between traffic dynamics and incidents\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track,External Links:[Link](https://openreview.net/forum?id=sZvXXPqONQ)Cited by:[Table 1](https://arxiv.org/html/2608.20656#S2.T1.6.1.27.1)\. - Gulrajani and Lopez\-Paz \(2021\)I\. Gulrajani and D\. Lopez\-PazIn search of lost domain generalization\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=lQdXeXDoWtI)Cited by:[§5\.2](https://arxiv.org/html/2608.20656#S5.SS2.p1.1)\. - Guoet al\.\(2019\)S\. Guo, Y\. Lin, N\. Feng, C\. Song, and H\. WanAttention based spatial\-temporal graph convolutional networks for traffic flow forecasting\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.33,pp\. 922–929\.Cited by:[Table 1](https://arxiv.org/html/2608.20656#S2.T1.6.1.7.1.1)\. - Heet al\.\(2013\)J\. He, W\. Shen, P\. Divakaruni, L\. Wynter, and R\. LawrenceImproving traffic prediction with tweet semantics\.InProceedings of the Twenty\-Third International Joint Conference on Artificial Intelligence,IJCAI ’13,pp\. 1387–1393\.Cited by:[§1](https://arxiv.org/html/2608.20656#S1.p2.1)\. - Hyndman and Koehler \(2006\)R\. J\. Hyndman and A\. B\. KoehlerAnother look at measures of forecast accuracy\.International Journal of Forecasting22\(4\),pp\. 679–688\.External Links:ISSN 0169\-2070,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.ijforecast.2006.03.001),[Link](https://www.sciencedirect.com/science/article/pii/S0169207006000239)Cited by:[4th item](https://arxiv.org/html/2608.20656#S4.I2.i4.p1.1)\. - Jiet al\.\(2022\)J\. Ji, J\. Wang, Z\. Jiang, J\. Jiang, and H\. ZhangSTDEN: towards physics\-guided neural networks for traffic flow prediction\.InProceedings of the AAAI conference on artificial intelligence,Vol\.36,pp\. 4048–4056\.Cited by:[§1](https://arxiv.org/html/2608.20656#S1.p3.1)\. - Jiet al\.\(2025\)J\. Ji, W\. Zhang, J\. Wang, and C\. HuangSeeing the unseen: learning basis confounder representations for robust traffic prediction\.InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\.1,KDD ’25,pp\. 577–588\.External Links:[Link](https://doi.org/10.1145/3690624.3709201),[Document](https://dx.doi.org/10.1145/3690624.3709201)Cited by:[§4\.4](https://arxiv.org/html/2608.20656#S4.SS4.p1.1)\. - Jianget al\.\(2023\)R\. Jiang, Z\. Wang, J\. Yong, P\. Jeph, Q\. Chen, Y\. Kobayashi, X\. Song, S\. Fukushima, and T\. SuzumuraSpatio\-temporal meta\-graph learning for traffic forecasting\.Proceedings of the AAAI Conference on Artificial Intelligence37,pp\. 8078–8086\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/25976),[Document](https://dx.doi.org/10.1609/aaai.v37i7.25976)Cited by:[3rd item](https://arxiv.org/html/2608.20656#S4.I2.i3.p1.1),[§4\.1](https://arxiv.org/html/2608.20656#S4.SS1.SSS0.Px2.p1.1),[§5\.1](https://arxiv.org/html/2608.20656#S5.SS1.p1.1)\. - Kimet al\.\(2021\)T\. Kim, J\. Kim, Y\. Tae, C\. Park, J\. Choi, and J\. ChooReversible instance normalization for accurate time\-series forecasting against distribution shift\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=cGDAkQo1C0p)Cited by:[§5\.3](https://arxiv.org/html/2608.20656#S5.SS3.p1.1)\. - Kruegeret al\.\(2021\)D\. Krueger, E\. Caballero, J\. Jacobsen, A\. Zhang, J\. Binas, D\. Zhang, R\. L\. Priol, and A\. CourvilleOut\-of\-distribution generalization via risk extrapolation \(rex\)\.InProceedings of the 38th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.139,pp\. 5815–5826\.Cited by:[§5\.2](https://arxiv.org/html/2608.20656#S5.SS2.p1.1)\. - Laiet al\.\(2018\)G\. Lai, W\. Chang, Y\. Yang, and H\. LiuModeling long\- and short\-term temporal patterns with deep neural networks\.InProceedings of the 41st International ACM SIGIR Conference on Research & Development in Information Retrieval \(SIGIR ’18\),pp\. 105–114\.External Links:[Document](https://dx.doi.org/10.1145/3209978.3210009)Cited by:[Table 1](https://arxiv.org/html/2608.20656#S2.T1.6.1.4.1)\. - Leeet al\.\(2015\)J\. Lee, B\. Hong, K\. Lee, and Y\. JangA prediction model of traffic congestion using weather data\.In2015 IEEE International Conference on Data Science and Data Intensive Systems,pp\. 81–88\.Cited by:[§1](https://arxiv.org/html/2608.20656#S1.p2.1)\. - Liet al\.\(2018\)Y\. Li, R\. Yu, C\. Shahabi, and Y\. LiuDiffusion convolutional recurrent neural network: data\-driven traffic forecasting\.InInternational Conference on Learning Representations \(ICLR ’18\),Cited by:[§1](https://arxiv.org/html/2608.20656#S1.p2.1),[Table 1](https://arxiv.org/html/2608.20656#S2.T1.6.1.2.1.1),[2nd item](https://arxiv.org/html/2608.20656#S4.I2.i2.p1.1),[§4\.1](https://arxiv.org/html/2608.20656#S4.SS1.SSS0.Px2.p1.1),[§5\.1](https://arxiv.org/html/2608.20656#S5.SS1.p1.1)\. - Liuet al\.\(2023a\)H\. Liu, Z\. Dong, R\. Jiang, J\. Deng, J\. Deng, Q\. Chen, and X\. SongSpatio\-temporal adaptive embedding makes vanilla transformer sota for traffic forecasting\.InProceedings of the 32nd ACM International Conference on Information and Knowledge Management,pp\. 4125–4129\.Cited by:[§1](https://arxiv.org/html/2608.20656#S1.p2.1),[§3\.3\.1](https://arxiv.org/html/2608.20656#S3.SS3.SSS1.p1.1),[§5\.1](https://arxiv.org/html/2608.20656#S5.SS1.p1.1)\. - Liuet al\.\(2023b\)X\. Liu, Y\. Xia, Y\. Liang, J\. Hu, Y\. Wang, L\. Bai, C\. Huang, Z\. Liu, B\. Hooi, and R\. ZimmermannLargeST: a benchmark dataset for large\-scale traffic forecasting\.InAdvances in Neural Information Processing Systems,Cited by:[Table 1](https://arxiv.org/html/2608.20656#S2.T1.6.1.13.1.1)\. - Liuet al\.\(2022a\)Y\. Liu, H\. Wu, J\. Wang, and M\. LongNon\-stationary transformers: exploring the stationarity in time series forecasting\.InAdvances in Neural Information Processing Systems,Cited by:[§5\.3](https://arxiv.org/html/2608.20656#S5.SS3.p1.1)\. - Liuet al\.\(2022b\)Z\. Liu, J\. Li, and K\. WuContext\-aware taxi dispatching at city\-scale using deep reinforcement learning\.IEEE Transactions on Intelligent Transportation Systems23\(3\),pp\. 1996–2009\.External Links:[Document](https://dx.doi.org/10.1109/TITS.2020.3030252)Cited by:[§1](https://arxiv.org/html/2608.20656#S1.p1.1)\. - Luet al\.\(2023\)J\. Lu, C\. Li, X\. B\. Wu, and X\. S\. ZhouPhysics\-informed neural networks for integrated traffic state and queue profile estimation: a differentiable programming approach on layered computational graphs\.Transportation Research Part C: Emerging Technologies153,pp\. 104224\.Cited by:[§1](https://arxiv.org/html/2608.20656#S1.p3.1)\. - Sagawaet al\.\(2020\)S\. Sagawa, P\. W\. Koh, T\. B\. Hashimoto, and P\. LiangDistributionally robust neural networks\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=ryxGuJrFvS)Cited by:[§5\.2](https://arxiv.org/html/2608.20656#S5.SS2.p1.1)\. - Shanget al\.\(2021\)C\. Shang, J\. Chen, and J\. BiDiscrete graph structure learning for forecasting multiple time series\.arXiv preprint arXiv:2101\.06861\.Cited by:[2nd item](https://arxiv.org/html/2608.20656#S4.I2.i2.p1.1),[§4\.1](https://arxiv.org/html/2608.20656#S4.SS1.SSS0.Px2.p1.1),[§4\.3](https://arxiv.org/html/2608.20656#S4.SS3.p4.1),[§4\.6](https://arxiv.org/html/2608.20656#S4.SS6.p1.1)\. - Shaoet al\.\(2025\)F\. Shao, H\. Shao, X\. Wu, Q\. Cheng, and W\. H\. LamA physics\-informed machine learning framework for speed\-flow prediction: integrating an s\-shaped traffic stream model with deep learning models\.Transportation Research Part C: Emerging Technologies180,pp\. 105362\.Cited by:[§1](https://arxiv.org/html/2608.20656#S1.p3.1)\. - Shaoet al\.\(2024\)Z\. Shao, F\. Wang, Y\. Xu, W\. Wei, C\. Yu, Z\. Zhang, D\. Yao, T\. Sun, G\. Jin, X\. Cao,et al\.Exploring progress in multivariate time series forecasting: comprehensive benchmarking and heterogeneity analysis\.IEEE Transactions on Knowledge and Data Engineering37\(1\),pp\. 291–305\.Cited by:[§1](https://arxiv.org/html/2608.20656#S1.p2.1),[§5\.1](https://arxiv.org/html/2608.20656#S5.SS1.p1.1)\. - Songet al\.\(2020\)C\. Song, Y\. Lin, S\. Guo, and H\. WanSpatial\-temporal synchronous graph convolutional networks: a new framework for spatial\-temporal network data forecasting\.Proceedings of the AAAI Conference on Artificial Intelligence34\(01\),pp\. 914–921\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/5438),[Document](https://dx.doi.org/10.1609/aaai.v34i01.5438)Cited by:[Table 1](https://arxiv.org/html/2608.20656#S2.T1.6.1.9.1.1)\. - Turochy and Smith \(2000\)R\. E\. Turochy and B\. L\. SmithNew procedure for detector data screening in traffic management systems\.Transportation Research Record1727\(1\),pp\. 127–131\.External Links:[Document](https://dx.doi.org/10.3141/1727-16)Cited by:[§1](https://arxiv.org/html/2608.20656#S1.p5.1)\. - Wang and Tong \(2026\)G\. Wang and J\. TongFrequency as identity: a fourier hypernetwork for spatiotemporal forecasting\.Applied Soft Computing187,pp\. 114297\.External Links:[Document](https://dx.doi.org/10.1016/j.asoc.2025.114297)Cited by:[§1](https://arxiv.org/html/2608.20656#S1.p1.1)\. - Wanget al\.\(2022\)Z\. Wang, R\. Jiang, H\. Xue, F\. D\. Salim, X\. Song, and R\. ShibasakiEvent\-aware multimodal mobility nowcasting\.Proceedings of the AAAI Conference on Artificial Intelligence36\(4\),pp\. 4228–4236\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/20342),[Document](https://dx.doi.org/10.1609/aaai.v36i4.20342)Cited by:[§5\.1](https://arxiv.org/html/2608.20656#S5.SS1.p1.1)\. - Wuet al\.\(2020\)Z\. Wu, S\. Pan, G\. Long, J\. Jiang, X\. Chang, and C\. ZhangConnecting the dots: multivariate time series forecasting with graph neural networks\.InProceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining,Cited by:[§1](https://arxiv.org/html/2608.20656#S1.p2.1)\. - Wuet al\.\(2019\)Z\. Wu, S\. Pan, G\. Long, J\. Jiang, and C\. ZhangGraph wavenet for deep spatial\-temporal graph modeling\.InProceedings of the 28th International Joint Conference on Artificial Intelligence,IJCAI’19,pp\. 1907–1913\.Cited by:[§4\.1](https://arxiv.org/html/2608.20656#S4.SS1.SSS0.Px2.p1.1),[§4\.3](https://arxiv.org/html/2608.20656#S4.SS3.p4.1),[§4\.6](https://arxiv.org/html/2608.20656#S4.SS6.p1.1),[§5\.1](https://arxiv.org/html/2608.20656#S5.SS1.p1.1)\. - Yinet al\.\(2025\)D\. Yin, H\. Xue, A\. Prabowo, S\. Ao, and F\. SalimXXLTraffic: Expanding and Extremely Long Traffic forecasting beyond test adaptation\.arXiv\.External Links:2406\.12693,[Document](https://dx.doi.org/10.48550/arXiv.2406.12693)Cited by:[Table 1](https://arxiv.org/html/2608.20656#S2.T1.6.1.17.1.1)\. - Yuet al\.\(2018\)B\. Yu, H\. Yin, and Z\. ZhuSpatio\-temporal graph convolutional networks: a deep learning framework for traffic forecasting\.InProceedings of the 27th International Joint Conference on Artificial Intelligence \(IJCAI 2018\),pp\. 3634–3640\.External Links:[Link](https://www.ijcai.org/proceedings/2018/0505.pdf)Cited by:[Table 1](https://arxiv.org/html/2608.20656#S2.T1.6.1.5.1.1),[§4\.1](https://arxiv.org/html/2608.20656#S4.SS1.SSS0.Px2.p1.1),[§5\.1](https://arxiv.org/html/2608.20656#S5.SS1.p1.1)\. - Zhanget al\.\(2024\)C\. Zhang, Y\. Zhang, Q\. Shao, B\. Li, Y\. Lv, X\. Piao, and B\. YinChatTraffic: text\-to\-traffic generation via diffusion model\.IEEE Transactions on Intelligent Transportation Systems,pp\. 1–13\.External Links:[Document](https://dx.doi.org/10.1109/TITS.2024.3510402)Cited by:[§1](https://arxiv.org/html/2608.20656#S1.p2.1)\. - Zhanget al\.\(2017\)J\. Zhang, Y\. Zheng, and D\. QiDeep spatio\-temporal residual networks for citywide crowd flows prediction\.Proceedings of the AAAI Conference on Artificial Intelligence31\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/10735),[Document](https://dx.doi.org/10.1609/aaai.v31i1.10735)Cited by:[§1](https://arxiv.org/html/2608.20656#S1.p2.1)\. - Zhenget al\.\(2020\)C\. Zheng, X\. Fan, C\. Wang, and J\. QiGman: a graph multi\-attention network for traffic prediction\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.34,pp\. 1234–1241\.Cited by:[§4\.1](https://arxiv.org/html/2608.20656#S4.SS1.SSS0.Px2.p1.1)\. - Zhouet al\.\(2023a\)Z\. Zhou, Q\. Huang, K\. Yang, K\. Wang, X\. Wang, Y\. Zhang, Y\. Liang, and Y\. WangMaintaining the status quo: capturing invariant relations for ood spatiotemporal learning\.InProceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining,KDD ’23,New York, NY, USA,pp\. 3603–3614\.External Links:ISBN 9798400701030,[Link](https://doi.org/10.1145/3580305.3599421),[Document](https://dx.doi.org/10.1145/3580305.3599421)Cited by:[§4\.4](https://arxiv.org/html/2608.20656#S4.SS4.p1.1)\. - Zhouet al\.\(2023b\)Z\. Zhou, Q\. Huang, K\. Yang, K\. Wang, X\. Wang, Y\. Zhang, Y\. Liang, and Y\. WangMaintaining the status quo: capturing invariant relations for ood spatiotemporal learning\.InProceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining,KDD ’23,pp\. 3603–3614\.Cited by:[§1](https://arxiv.org/html/2608.20656#S1.p6.1)\. - Zhouet al\.\(2026\)Z\. Zhou, J\. Hu, Q\. Wen, J\. T\. Kwok, and Y\. LiangMulti\-order wavelet derivative transform for deep time series forecasting\.External Links:2505\.11781Cited by:[§5\.3](https://arxiv.org/html/2608.20656#S5.SS3.p1.1)\.
Similar Articles
MambaLSTM: A Spatio-Temporal Framework for Enhanced Traffic Accident Risk Prediction
The paper proposes MambaLSTM, a framework combining Mamba state-space models and LSTM for spatio-temporal traffic accident risk prediction, addressing noise in feature fusion and global spatial correlation.
TrajRS: Towards Certified Robustness in Pedestrian Trajectory Prediction
This paper introduces TrajRS, an extension of Randomized Smoothing that provides certified robust radii for pedestrian trajectory predictors, offering verifiable safety guarantees against adversarial perturbations.
Trajectory-Aware Clinical Risk Prediction via Severity-Grounded Knowledge Graphs and Retrieval-Augmented Generation
Proposes TRACER, a framework that integrates severity-grounded knowledge graphs and retrieval-augmented generation for trajectory-aware clinical risk prediction, achieving large gains in mortality and readmission prediction on MIMIC-III and MIMIC-IV datasets.
Learning A Unified Risk Map for Autonomous Driving in Partially Observable Environments
Proposes a unified risk map modeling framework for autonomous driving that integrates traffic flow and collision risks in partially observable environments, using spatiotemporal modeling and diffusion-based scenario generation. Outperforms state-of-the-art occlusion-aware baselines on the Waymo Open Motion Dataset.
PRB-RUPFormer: A Recursive Unified Probabilistic Transformer for Residual PRB Forecasting
Proposes PRB-RUPFormer, a recursive unified probabilistic Transformer for forecasting residual Physical Resource Blocks in cellular networks, achieving high accuracy and uncertainty quantification on commercial LTE data.