Forecast Collapse in Time-Series Foundation Models
Summary
The paper identifies forecast collapse in time-series foundation models for hourly equity return prediction and introduces CalibRank to balance calibration and ranking, significantly improving cross-sectional correlation.
View Cached Full Text
Cached at: 08/17/26, 07:44 AM
Paper page - Forecast Collapse in Time-Series Foundation Models
Source: https://huggingface.co/papers/2608.14106 Published on Aug 14
·
Submitted byhttps://huggingface.co/Shuwan
Shuon Aug 17
Abstract
Forecast collapse in hourly equity return prediction stems from low predictability and per-series objectives, and the proposed CalibRank objective balances calibration and ranking to restore cross-sectional structure.
When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly flat and show poor stock ranking, as measured bycross-sectional correlation. We call thisforecast collapse. Surprisingly, the phenomenon largely disappears when forecasting trading volume under the same setting. We investigateforecast collapseacrosstime-series foundation models(TSFMs), twelve deep-learning forecasting models, and 97 public benchmark configurations, and find that it is closely tied to target predictability. We identify two distinct reasons behind it: low predictability limits the amplitude of calibrated point forecasts, whileper-series objectivesleave cross-series structure unidentified. These findings reveal acalibration-ranking tradeoff: optimizing squared error leads to flat predictions, whereas directly optimizingcross-sectional correlationimproves ranking but can inflate forecast amplitude by more than an order of magnitude. To address this tradeoff, we introduceCalibRank, a simple objective that balances calibration and ranking. On Finance1K,CalibRanknearly triplescross-sectional correlationwhile keeping amplitude close to the target, and improves correlation on all tested models. Our results reveal a blind spot in conventional time-series evaluation: per-series metrics can hide failures in cross-series structure needed by downstream decisions.
View arXiv pageView PDFAdd to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.14106 in a model README.md to link it from this page.
Datasets citing this paper1
#### abel-lab/finance1k Viewer• Updatedabout 2 hours ago • 57k • 50 • 1
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.14106 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Do Time Series Foundation Model Benchmarks Hide Regime-Dependent Failures? Evidence from Traffic Speed Forecasting
This paper introduces regime-stratified evaluation for time series foundation models, revealing that aggregate metrics hide severe failures during traffic regime transitions, and proposes bimodal mixture augmentation to improve coverage while preserving overall accuracy.
TS-Fault: Benchmarking Time Series Forecasters Against Structural Faults
This paper introduces TS-Fault, a benchmark for evaluating time series forecasting models under structured fault scenarios like broken dependencies and regime changes, finding that clean-data accuracy often anti-correlates with robustness and that foundation models are especially fragile.
Dynamic Regime-Aware Conformal Calibration for Reliable Economic Forecast Intervals under Multiple Distribution Shifts
The paper introduces DRACP, a conformal prediction method that unifies multiple adaptation mechanisms for reliable economic forecast intervals under distribution shifts, achieving calibration with wider intervals.
Assessing the Operational Viability of Foundation Models for Time Series Forecasting
This paper presents an applied evaluation of foundation models for time series forecasting compared to supervised approaches across four operational domains, and proposes a Complexity Router to selectively assign series to the optimal model class for balancing accuracy and inference cost.
How Good Are Time-Series Foundation Models for Pedestrian Crowd Count Forecasting? A Cross-Dataset Comparative Study
The paper benchmarks time-series foundation models for pedestrian crowd count forecasting across datasets, finding that foundation models excel in data-rich, seasonal regimes while simpler models can be competitive in limited data scenarios.