science-critique

Tag

Cards List
#science-critique

The best AI “science critics” are also the most overconfident — a benchmark on calibration vs. skill

Reddit r/artificial · 2026-06-05

The article introduces the Refute benchmark, which tests LLMs on critiquing science paper summaries and measures their calibration. Results show that the best critic models are often the most overconfident when wrong.

0 favorites 0 likes
← Back to home

Submit Feedback