Calibration as a First-Class Criterion in LLM Evaluation
Summary
This paper argues that calibration should be a first-class criterion in LLM evaluation to enhance trustworthiness and address miscalibration problems in deployment and research pipelines.
View Cached Full Text
Cached at: 09/24/26, 11:39 AM
Paper page - Calibration as a First-Class Criterion in LLM Evaluation
Source: https://huggingface.co/papers/2609.26489
Abstract
Calibrationoflanguagemodels--thealignmentbetweenexpressedorimplicitconfidenceandempiricalcorrectness--isawell-studiedsubfieldwithinNLP.Methodstomeasureitalreadyexist.Theproblemisadoption:outsidethissubfield,NLPresearchregularlyintroducesnewmodels,datasets,andbenchmarkswithoutcheckingwhetherthemodel’sconfidencescoresaremeaningful.WearguethatthisadoptiongapisamajorobstacletotrustworthyLLMevaluation.Miscalibrationcausesproblemsintwodistinctareas:atdeployment,whereoverconfidentmistakescauserealharm,andinsidetheresearchpipeline,wheremethodslikeLLM-as-a-judge,syntheticdatageneration,andactivelearningrelyoncalibratedconfidencewithoutverifyingit.Standardcalibrationmetricsonlyrequiretwoinputsperexample:aconfidencescoreandacorrectnessjudgment.Mostbenchmarksinusetodayalreadyprovideboth,meaningcalibrationcanbereportedimmediately.Foropen-endedgeneration,however,definingthesetwoinputsisstillanopenchallenge.WearguethateachNLPsubfieldshouldpairitsmainperformancemetricwithacalibrationscoreandcallfortreatingcalibrationasanessentialpropertyofeverymodelratherthananichetopic.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2609\.26489
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.26489 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.26489 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.26489 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
When Calibration Rankings Reverse: Accuracy-Controlled Evaluation for Fair Comparison of LLMs
This paper demonstrates that global calibration metrics like Expected Calibration Error are confounded by model accuracy, and proposes ACE, an accuracy-controlled evaluation framework for fair comparison of large language models.
Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Data
This paper introduces Self-Evaluation Elicitation (SEE), which uses calibration-coupled reinforcement learning and masked distillation to elicit latent judge calibration in base LLMs with minimal data, improving calibration across benchmarks while preserving answer quality.
Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration
This study investigates capability-dependent biases in LLM judges and introduces a calibrated weighted majority voting ensemble method to enhance automated evaluation reliability without requiring labeled data.
Confidence Calibration in Large Language Models
This paper analyzes the confidence calibration of 11 popular LLMs, finding that they are generally overconfident, especially on hard tasks, and underconfident on easy tasks. It introduces LifeEval, a test for evaluating calibration across difficulty levels.
Faithful uncertainty in LLM agents: calibration vs utility tradeoff in practice[D]
A practitioner discusses the calibration vs. utility tradeoff in LLM agents, sharing experience with a verifier-based pipeline that reduces hallucinated tool calls by ~60% but introduces latency costs and drops easy correct answers.