Calibration as a First-Class Criterion in LLM Evaluation

Hugging Face Daily Papers Papers

Summary

This paper argues that calibration should be a first-class criterion in LLM evaluation to enhance trustworthiness and address miscalibration problems in deployment and research pipelines.

Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is a well-studied subfield within NLP. Methods to measure it already exist. The problem is adoption: outside this subfield, NLP research regularly introduces new models, datasets, and benchmarks without checking whether the model's confidence scores are meaningful. We argue that this adoption gap is a major obstacle to trustworthy LLM evaluation. Miscalibration causes problems in two distinct areas: at deployment, where overconfident mistakes cause real harm, and inside the research pipeline, where methods like LLM-as-a-judge, synthetic data generation, and active learning rely on calibrated confidence without verifying it. Standard calibration metrics only require two inputs per example: a confidence score and a correctness judgment. Most benchmarks in use today already provide both, meaning calibration can be reported immediately. For open-ended generation, however, defining these two inputs is still an open challenge. We argue that each NLP subfield should pair its main performance metric with a calibration score and call for treating calibration as an essential property of every model rather than a niche topic.
Original Article
View Cached Full Text

Cached at: 09/24/26, 11:39 AM

Paper page - Calibration as a First-Class Criterion in LLM Evaluation

Source: https://huggingface.co/papers/2609.26489

Abstract

Calibrationoflanguagemodels--thealignmentbetweenexpressedorimplicitconfidenceandempiricalcorrectness--isawell-studiedsubfieldwithinNLP.Methodstomeasureitalreadyexist.Theproblemisadoption:outsidethissubfield,NLPresearchregularlyintroducesnewmodels,datasets,andbenchmarkswithoutcheckingwhetherthemodel’sconfidencescoresaremeaningful.WearguethatthisadoptiongapisamajorobstacletotrustworthyLLMevaluation.Miscalibrationcausesproblemsintwodistinctareas:atdeployment,whereoverconfidentmistakescauserealharm,andinsidetheresearchpipeline,wheremethodslikeLLM-as-a-judge,syntheticdatageneration,andactivelearningrelyoncalibratedconfidencewithoutverifyingit.Standardcalibrationmetricsonlyrequiretwoinputsperexample:aconfidencescoreandacorrectnessjudgment.Mostbenchmarksinusetodayalreadyprovideboth,meaningcalibrationcanbereportedimmediately.Foropen-endedgeneration,however,definingthesetwoinputsisstillanopenchallenge.WearguethateachNLPsubfieldshouldpairitsmainperformancemetricwithacalibrationscoreandcallfortreatingcalibrationasanessentialpropertyofeverymodelratherthananichetopic.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2609\.26489

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.26489 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.26489 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.26489 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Confidence Calibration in Large Language Models

arXiv cs.AI

This paper analyzes the confidence calibration of 11 popular LLMs, finding that they are generally overconfident, especially on hard tasks, and underconfident on easy tasks. It introduces LifeEval, a test for evaluating calibration across difficulty levels.