The best AI “science critics” are also the most overconfident — a benchmark on calibration vs. skill
Summary
The article introduces the Refute benchmark, which tests LLMs on critiquing science paper summaries and measures their calibration. Results show that the best critic models are often the most overconfident when wrong.
Similar Articles
AI research tools are still too eager to turn public signals into certainty
The author critiques AI research tools for overconfidence in weak signals, praising Komo AI's rapid discovery and source-attached summaries but highlighting the need for better uncertainty and contradiction handling. They describe a workflow that splits discovery, verification, and structured checking across multiple AI tools.
the problem isnt that AI is wrong, its that it's wrong so confidently
Discusses the issue of AI models producing incorrect answers with high confidence, highlighting the problem of overconfidence in AI outputs.
AI is confidently wrong way more than people give it credit for, change my mind
A user shares concerns about AI models presenting thin or ambiguous data with the same confidence as well-supported findings, citing a case where a complaint appearing only twice in 200 comments was ranked as a top concern. The piece questions whether this is a fixable prompting issue or a fundamental limitation requiring manual verification.
SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones?
SoundnessBench is a benchmark of 1,099 machine-learning research proposals that evaluates LLMs' ability to assess methodological validity, finding a pervasive optimism bias in current models.
Confidence Calibration in Large Language Models
This paper analyzes the confidence calibration of 11 popular LLMs, finding that they are generally overconfident, especially on hard tasks, and underconfident on easy tasks. It introduces LifeEval, a test for evaluating calibration across difficulty levels.