@VraserX: The AI scientist I’d pay for would spend a lot of its time checking other people’s results. In the new RECLAIM preprint…

X AI KOLs Following News

Summary

A tweet discusses a new RECLAIM preprint showing that AI agents can reproduce research results at varying rates depending on code availability, highlighting the potential value of AI in replicating existing studies.

The AI scientist I’d pay for would spend a lot of its time checking other people’s results. In the new RECLAIM preprint, the best agent reproduced 41% of results in the tier with code, data and weights available. Without code, the best rate was 15%. A common failure was implementing the method without checking the paper’s numbers. Would you fund an AI lab that mostly replicated existing research? I would.
Original Article
View Cached Full Text

Cached at: 09/29/26, 07:53 PM

The AI scientist I’d pay for would spend a lot of its time checking other people’s results.

In the new RECLAIM preprint, the best agent reproduced 41% of results in the tier with code, data and weights available. Without code, the best rate was 15%.

A common failure was implementing the method without checking the paper’s numbers.

Would you fund an AI lab that mostly replicated existing research? I would.

Similar Articles

AI Coding Agents Can Reproduce Social Science Findings

arXiv cs.CL

This paper introduces SocSci-Repro-Bench, a benchmark of 221 tasks to evaluate AI coding agents' ability to reproduce social science findings from original data and code. It finds that frontier agents like Claude Code and Codex can reproduce a large share of results, with Claude substantially outperforming Codex, and that results are not primarily driven by memorization.

PaperBench: Evaluating AI’s Ability to Replicate AI Research

OpenAI Blog

OpenAI introduces PaperBench, a benchmark evaluating AI agents' ability to replicate state-of-the-art AI research by replicating 20 ICML 2024 papers with 8,316 gradable tasks. The best-performing model (Claude 3.5 Sonnet) achieves only 21% replication score, below human PhD-level performance, highlighting current limitations in autonomous research capabilities.