@VraserX: The AI scientist I’d pay for would spend a lot of its time checking other people’s results. In the new RECLAIM preprint…
Summary
A tweet discusses a new RECLAIM preprint showing that AI agents can reproduce research results at varying rates depending on code availability, highlighting the potential value of AI in replicating existing studies.
View Cached Full Text
Cached at: 09/29/26, 07:53 PM
The AI scientist I’d pay for would spend a lot of its time checking other people’s results.
In the new RECLAIM preprint, the best agent reproduced 41% of results in the tier with code, data and weights available. Without code, the best rate was 15%.
A common failure was implementing the method without checking the paper’s numbers.
Would you fund an AI lab that mostly replicated existing research? I would.
Similar Articles
RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?
RECLAIM is a benchmark for AI agents to reproduce claims from machine learning papers, with difficulty tiers based on available code, data, and weights. It shows that agents struggle to reproduce results, with best success rates of 41% in the easiest tier.
AI Coding Agents Can Reproduce Social Science Findings
This paper introduces SocSci-Repro-Bench, a benchmark of 221 tasks to evaluate AI coding agents' ability to reproduce social science findings from original data and code. It finds that frontier agents like Claude Code and Codex can reproduce a large share of results, with Claude substantially outperforming Codex, and that results are not primarily driven by memorization.
@askalphaxiv: 70% of AI research isn’t reproducible. With ICML 2026 happening last week, over 6000+ research papers have dropped, but…
Alphaxiv and Hugging Face launch a community challenge to test the reproducibility of AI research papers from ICML 2026, offering $4500 in GPU credits and an autoresearch agent to help participants.
PaperBench: Evaluating AI’s Ability to Replicate AI Research
OpenAI introduces PaperBench, a benchmark evaluating AI agents' ability to replicate state-of-the-art AI research by replicating 20 ICML 2024 papers with 8,316 gradable tasks. The best-performing model (Claude 3.5 Sonnet) achieves only 21% replication score, below human PhD-level performance, highlighting current limitations in autonomous research capabilities.
@tunguz: Wow. This cold be way bigger than the replication crisis.
An AI tool named Astra was used to analyze academic replication packages and uncovered numerous coding errors, some of which overturn central results in high-ranking journals, while also revealing that many models were not run properly.