brier-score

Tag

Cards List
#brier-score

ConfidenceBench: Evaluating Confidence Calibration in Large Language Models

arXiv cs.AI · yesterday Cached

ConfidenceBench is a new benchmark that evaluates verbalized confidence estimates in large language models using Brier scores, revealing that accuracy and calibration diverge and that even highly accurate models can be severely miscalibrated.

0 favorites 0 likes
#brier-score

The best AI “science critics” are also the most overconfident — a benchmark on calibration vs. skill

Reddit r/artificial · 2026-06-05

The article introduces the Refute benchmark, which tests LLMs on critiquing science paper summaries and measures their calibration. Results show that the best critic models are often the most overconfident when wrong.

0 favorites 0 likes
#brier-score

Teaching AI models to say “I’m not sure”

MIT News — Artificial Intelligence · 2026-04-22 Cached

MIT CSAIL researchers introduce RLCR, a method using Brier scores in reinforcement learning to train AI models to output calibrated confidence estimates, significantly reducing overconfidence without sacrificing accuracy.

0 favorites 0 likes
← Back to home

Submit Feedback