Forecasting Scientific Progress with Artificial Intelligence
Summary
This paper introduces CUSP, a benchmark for evaluating AI systems' ability to forecast scientific progress, finding that current models show systematic overconfidence and domain-dependent limitations, failing to reliably predict scientific advances.
View Cached Full Text
Cached at: 05/22/26, 06:21 PM
Paper page - Forecasting Scientific Progress with Artificial Intelligence
Source: https://huggingface.co/papers/2605.22681
Abstract
Current AI systems demonstrate limited capability in predicting scientific progress, showing inconsistent performance across domains and systematic overconfidence in forecasts.
Artificial intelligence(AI) is increasingly embedded in scientific discovery, yet whether it can anticipatescientific progressremains unclear. To study this question, we introduce a temporally grounded evaluation framework for forecastingscientific progressunder controlled knowledge constraints. We present CUSP (Cutoff-conditioned UnseenScientific Progress), a multi-disciplinary and event-level benchmark that evaluatesscientific forecastingin AI systems throughfeasibility assessment,mechanistic reasoning,generative solution design, andtemporal prediction. Across 4,760 scientific events, we observe systematic anddomain-dependent limitationsin current frontier models. While models can identify plausible research directions from competing candidates, they fail to reliably predict whether scientific advances will be realized and systematically misestimate when they will occur. Performance is highly heterogeneous across domains, with the timing of AI progress more predictable than advances in biology, chemistry, and physics. Performance is largely insensitive to whether events occur before or after the training cutoff, suggesting these limitations cannot be explained solely by knowledge exposure in training data. Under controlled information access, additionalpre-cutoff knowledgeimproves performance but does not close the gap to full-information settings, which becomes more pronounced for high-citation advances. Models also exhibit systematic overconfidence and strong response biases, indicating unreliableuncertainty estimation. Taken together, current AI systems fall short as predictive tools forscientific progress. Access to prior knowledge does not translate into reliable forecasting, and performance benefits more frompost-event informationthan from forward-looking prediction.
View arXiv pageView PDFProject pageGitHub13Add to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.22681 in a model README.md to link it from this page.
Datasets citing this paper1
#### SeanWu25/CUSP Viewer• Updatedabout 12 hours ago • 4.76k • 253 • 8
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.22681 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
@dair_ai: Can frontier models forecast scientific progress? Mostly no, but here is why. This work looks at 4,760 scientific event…
A study evaluates frontier models' ability to forecast scientific progress across 4,760 events, finding they can identify plausible directions but cannot reliably predict outcomes or timelines, with systematic overconfidence.
SciPaths: Forecasting Pathways to Scientific Discovery
Introduces SciPaths, a benchmark for forecasting the enabling contributions required to realize a target scientific discovery, and evaluates frontier and open-weight language models, finding significant room for improvement in reasoning backward from contributions to enabling building blocks.
Does anyone else feel like AI benchmarks are becoming less useful for predicting real-world performance?
The article discusses the growing disconnect between high AI benchmark scores and actual real-world performance, highlighting issues like consistency, latency, and context handling.
Can collective AI intelligence outperform collective human intelligence?
Explores whether ensembles of AI models could outperform human crowds in prediction markets, questioning if AI consensus will eventually surpass human forecasting accuracy.
@nickscamara_: New discoveries are gonna come from models that can reason over the latest science The rate of scientific progress beco…
Firecrawl released a state-of-the-art research index for AI/ML papers, claiming 18% better recall on arXivQA than competitors, designed for autonomous research agents.