Forecast Collapse in Time-Series Foundation Models

Hugging Face Daily Papers Papers

Summary

The paper identifies forecast collapse in time-series foundation models for hourly equity return prediction and introduces CalibRank to balance calibration and ranking, significantly improving cross-sectional correlation.

When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly flat and show poor stock ranking, as measured by cross-sectional correlation. We call this forecast collapse. Surprisingly, the phenomenon largely disappears when forecasting trading volume under the same setting. We investigate forecast collapse across time-series foundation models (TSFMs), twelve deep-learning forecasting models, and 97 public benchmark configurations, and find that it is closely tied to target predictability. We identify two distinct reasons behind it: low predictability limits the amplitude of calibrated point forecasts, while per-series objectives leave cross-series structure unidentified. These findings reveal a calibration-ranking tradeoff: optimizing squared error leads to flat predictions, whereas directly optimizing cross-sectional correlation improves ranking but can inflate forecast amplitude by more than an order of magnitude. To address this tradeoff, we introduce CalibRank, a simple objective that balances calibration and ranking. On Finance1K, CalibRank nearly triples cross-sectional correlation while keeping amplitude close to the target, and improves correlation on all tested models. Our results reveal a blind spot in conventional time-series evaluation: per-series metrics can hide failures in cross-series structure needed by downstream decisions.
Original Article
View Cached Full Text

Cached at: 08/17/26, 07:44 AM

Paper page - Forecast Collapse in Time-Series Foundation Models

Source: https://huggingface.co/papers/2608.14106 Published on Aug 14

·

Submitted byhttps://huggingface.co/Shuwan

Shuon Aug 17

Abstract

Forecast collapse in hourly equity return prediction stems from low predictability and per-series objectives, and the proposed CalibRank objective balances calibration and ranking to restore cross-sectional structure.

When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly flat and show poor stock ranking, as measured bycross-sectional correlation. We call thisforecast collapse. Surprisingly, the phenomenon largely disappears when forecasting trading volume under the same setting. We investigateforecast collapseacrosstime-series foundation models(TSFMs), twelve deep-learning forecasting models, and 97 public benchmark configurations, and find that it is closely tied to target predictability. We identify two distinct reasons behind it: low predictability limits the amplitude of calibrated point forecasts, whileper-series objectivesleave cross-series structure unidentified. These findings reveal acalibration-ranking tradeoff: optimizing squared error leads to flat predictions, whereas directly optimizingcross-sectional correlationimproves ranking but can inflate forecast amplitude by more than an order of magnitude. To address this tradeoff, we introduceCalibRank, a simple objective that balances calibration and ranking. On Finance1K,CalibRanknearly triplescross-sectional correlationwhile keeping amplitude close to the target, and improves correlation on all tested models. Our results reveal a blind spot in conventional time-series evaluation: per-series metrics can hide failures in cross-series structure needed by downstream decisions.

View arXiv pageView PDFAdd to collection

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2608.14106 in a model README.md to link it from this page.

Datasets citing this paper1

#### abel-lab/finance1k Viewer• Updatedabout 2 hours ago • 57k • 50 • 1

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2608.14106 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

TS-Fault: Benchmarking Time Series Forecasters Against Structural Faults

arXiv cs.LG

This paper introduces TS-Fault, a benchmark for evaluating time series forecasting models under structured fault scenarios like broken dependencies and regime changes, finding that clean-data accuracy often anti-correlates with robustness and that foundation models are especially fragile.