「very likely」意味着「不确定」?大模型与人类在语言不确定性量化上的差异
摘要
本文以心理学文献中的人类言语不确定性标记为基准对大模型进行评测,并提出 VOCAL——一种基于优化的算法,可直接从大模型的输出中学习最优的标记-不确定性映射,揭示了大模型与人类在语言表达置信度方面存在的系统性差异。
arXiv:2610.00083v1 Announce Type: new
Abstract: Humans express uncertainty verbally via markers (e.g., "possible," "likely"), yet most LLM uncertainty quantification (UQ) relies on costing likelihood- or consistency-based signals. From a cognitive perspective, accurate verbal uncertainty reflects metacognitive monitoring, representing knowledge boundaries ("knowing that you don't know") to support regulation and information seeking. In this paper, we investigate how LLMs diverge from humans in verbal uncertainty quantification and whether verbal markers can reliably quantify LLM uncertainty. We curate a corpus of human uncertainty markers from psychology and decision-science literature and benchmark LLMs against it. We observe that LLMs encode verbal uncertainty with numerical levels that differ substantially from those of humans. We then introduce METHODNAME, a novel optimization-based algorithm that learns an optimal uncertainty profile over uncertainty markers directly from LLM outputs. By fitting a marker-uncertainty mapping to best explain empirical correctness, METHODNAME discovers how much probability mass each verbal marker should convey, rather than estimating uncertainty via repeated sampling. METHODNAME enables a direct, marker-level comparison of confidence semantics between humans and LLMs, disentangling mismatch and revealing systematic confidence disparities in verbal expressions.
查看缓存全文
缓存时间: 2026/10/03 09:50
# “very likely” Means “uncertain”? How LLMs Diverge from Humans in Linguistic Uncertainty Quantification Source: [https://arxiv.org/html/2610.00083](https://arxiv.org/html/2610.00083) Zicheng LiuAffiliation:Department of Mathematics, The University of Hong KongZijie LiuAffiliation:Department of Computer Science, UNC\-Chapel HillKaidi XuAffiliation:Department of Data Science, City University of Hong KongTianlong ChenAffiliation:Department of Computer Science, UNC\-Chapel HillCorrespondence to:[tianlong@cs\.unc\.edu](mailto:[email protected]) ###### Abstract Humans express uncertainty verbally via markers \(e\.g\., “possible,” “likely”\), yet most LLM uncertainty quantification \(UQ\) relies on costing likelihood\- or consistency\-based signals\. From a cognitive perspective, accurate verbal uncertainty reflects metacognitive monitoring, representing knowledge boundaries \(“knowing that you don’t know”\) to support regulation and information seeking\. In this paper, we investigateHow LLMs diverge from humans in verbal uncertainty quantification? Can verbal markers reliably quantify LLM uncertainty?We curate a corpus of human uncertainty markers from psychology and decision\-science literature and benchmark LLMs against it\. We observe that LLMs encode verbal uncertainty with numerical levels that differ substantially from those of humans\. We then introduceVOCAL, a novel optimization\-based algorithm that learns an optimal uncertainty profile over uncertainty markers directly from LLM outputs\. By fitting a marker–uncertainty mapping to best explain empirical correctness,VOCALdiscovers how much probability mass each verbal marker should convey, rather than estimating uncertainty via repeated sampling\.VOCALenables a direct, marker\-level comparison of confidence semantics between humans and LLMs, disentangling mismatch and revealing systematic confidence disparities in verbal expressions\. ###### Keywords: Machine Learning, ICML ††affiliationnotice:Equal contribution## 1Introduction Despite large language models’\(LLMs\) recent remarkable success across diverse domains\([Xie et al\., 2023](https://arxiv.org/html/2610.00083#bib.bib40);[Colombo et al\., 2024](https://arxiv.org/html/2610.00083#bib.bib41);[Yang et al\., 2024](https://arxiv.org/html/2610.00083#bib.bib38);[Thapa et al\., 2025](https://arxiv.org/html/2610.00083#bib.bib39)\), a fundamental question remains: when should we trust LLMs’ responses? Complementary evaluations have examined strategic LLM reasoning through game\-theoretic tasks\([Duan et al\., 2024b](https://arxiv.org/html/2610.00083#bib.bib13)\)and privacy leakage through membership inference attacks on diffusion models\([Duan et al\., 2023](https://arxiv.org/html/2610.00083#bib.bib14)\), illustrating that trustworthy generative AI spans multiple dimensions\. This question highlights the need to make LLMs more trustworthy and responsible\. Hallucinations are not only mistakes but also risks that can reduce users’ trust and cause harm in sensitive applications\([Asgari et al\., 2025](https://arxiv.org/html/2610.00083#bib.bib42);[Das et al\., 2025](https://arxiv.org/html/2610.00083#bib.bib43)\), like giving unsafe treatment advice in biomedicine\. One promising approach to mitigating this phenomenon is uncertainty quantification \(UQ\)\([Malinin and Gales, 2020](https://arxiv.org/html/2610.00083#bib.bib15);[Kuhn et al\., 2023a](https://arxiv.org/html/2610.00083#bib.bib9);[Duan et al\., 2024a](https://arxiv.org/html/2610.00083#bib.bib12)\), which provides a probabilistic signal to estimate trustworthiness without labeled data\. This capability facilitates distinguishing valid predictions from hallucinations or extrapolation, essential for real\-world deployment\. Figure 1:Comparison of traditional uncertainty quantification \(UQ\) methods and our methodVOCAL\. Traditional UQ methods \(sampling\-based and logits\-based\) exhibit a gap with human uncertainty expressions\. InVOCAL, UM\-lookup tables derived from human data alone cannot fully capture model uncertainty, so they are optimized with the model’s confidence distribution to better align with its internal uncertainty expressions\.However, existing approaches for quantifying hallucination in LLMs still have some limitations, primarily dividing into two main groups: sampling\-based techniques\([Farquhar et al\., 2024](https://arxiv.org/html/2610.00083#bib.bib44);[Kossen et al\., 2024](https://arxiv.org/html/2610.00083#bib.bib45);[Li et al\., 2025a](https://arxiv.org/html/2610.00083#bib.bib46);[McCabe et al\., 2025](https://arxiv.org/html/2610.00083#bib.bib47)\)and logits\-based techniques\([Sriramanan et al\., 2024](https://arxiv.org/html/2610.00083#bib.bib51);[Nguyen et al\., 2025](https://arxiv.org/html/2610.00083#bib.bib48);[Ma et al\., 2025](https://arxiv.org/html/2610.00083#bib.bib49);[Yang et al\., 2025](https://arxiv.org/html/2610.00083#bib.bib50)\)\. While logits\-based methods such as Predictive Entropy \(PE\)\([Malinin and Gales, 2020](https://arxiv.org/html/2610.00083#bib.bib15)\)offer reliability, they suffer from lexical sensitivity and computational overhead\. Sampling\-based methods, such as semantic entropy \(SE\)\([Kuhn et al\., 2023b](https://arxiv.org/html/2610.00083#bib.bib52)\), ensemble variance, or consistency checks across multiple generations, can be more robust but are also slow and costly\. Therefore, a recent direction focuses on verbal uncertainty, where models output numerical uncertainty scores\([Tian et al\., 2023](https://arxiv.org/html/2610.00083#bib.bib53)\)\. Although improving calibration compared to raw probability outputs, precise quantification remains unnatural in contrast to qualitative terms such as “possible,” “likely,” or “almost certain” that better capture the nuance of human reasoning\. This contrast shows a gap in current methods and points to the need for approaches that allow models to express uncertainty in a way that is more natural, human\-like, and trustworthy for real\-world use \([Figure1](https://arxiv.org/html/2610.00083#S1.F1)\(left\)\)\. Motivated by this gap, we ask: How LLMs diverge from humans in verbal uncertainty quantification? Can verbal markers reliably quantify LLM uncertainty? To study this question, we first construct the verbal uncertainty marker lookup table \(UM\-Lookup\) that maps qualitative expressions of uncertainty to numerical representations\. The lookup table is built through a literature review grounded in psychology and decision science\([Lichtenstein and Newman, 1967](https://arxiv.org/html/2610.00083#bib.bib2);[Beyth\-Marom, 1982](https://arxiv.org/html/2610.00083#bib.bib3);[Wesson and Pulford, 2009](https://arxiv.org/html/2610.00083#bib.bib4)\), followed by a debiasing procedure to refine ambiguous cases\. We then aggregate judgments from more than 300 human annotators, resulting in a curated resource of 115 distinct verbal uncertainty markers with associated numeric interpretations\. Leveraging this resource, we evaluate the ability of LLMs to align their verbal expressions of uncertainty with human interpretations\. Our results show that LLMs demonstrate non\-trivial UQ performance when assessed against theUM\-Lookup\. For example, when evaluated with GPT\-4o\([Achiam et al\., 2023](https://arxiv.org/html/2610.00083#bib.bib28)\)model on SciQ dataset\([Welbl et al\., 2017](https://arxiv.org/html/2610.00083#bib.bib23)\), verbal uncertainty outperforms representative logits\-based and sampling\-based methods such as PE and SE, achieving an improvement of 4\.7% AUROC and 5\.6% AUROC, respectively\. However, across broader benchmarks, verbalized UQ remains weaker than strong UQ baselines, reflecting a gap between human\-derived lookup tables for uncertainty markers and LLM uncertainty signals\. This gap largely arises from divergent interpretations: LLMs often associate terms like “possible” with significantly lower uncertainty than humans\. Furthermore, unlike humans often combine multiple verbal uncertainty markers to convey more fine\-grained or complex levels of uncertainty, LLMs typically rely on single markers at each time\([Vogel et al\., 2022](https://arxiv.org/html/2610.00083#bib.bib5)\)\. These differences suggest a gap between human communication patterns and how LLMs currently express verbal uncertainty\. To address this gap, we proposeVOCAL, an approximation algorithm that provides an optimal mapping solution between uncertainty markers and numerical uncertainty levels by adapting to the uncertainty distribution of each model \([Figure1](https://arxiv.org/html/2610.00083#S1.F1)\(right\)\)\.VOCALis evaluated over comprehensive experiments on a wide range of models and datasets\. Our results demonstrate that theVOCALsignificantly outperforms single\-sample UQ methods, such as\([Aichberger et al\., 2025](https://arxiv.org/html/2610.00083#bib.bib8)\), and achieve comparable performance as multi\-sample UQs, without additional sampling or computational requests\. Our contributions are: - •We present a lookup table that maps human verbal uncertainty markers to numerical uncertainty scores, grounded in psychology and decision science\. This lookup table is a foundational resource that could not only benefit follow\-up verbal uncertainty quantification methods but also inspire the metacognition monitoring research in the future\. - •We propose a simple yet effective method,VOCAL, that optimizes the alignment between verbal markers and model uncertainty distributions\. - •We conduct comprehensive experiments across multiple models and datasets, providing in\-depth analysis and demonstrating the effectiveness of our method\. We demonstrate thatVOCALsignificantly outperforms single\-sample UQ methods and achieves comparable performances as multi\-sample UQ methods, with significantly reduced computational cost\. ## 2Related Work #### Sampling\-Based LLM Uncertainty Quantification\. The need to mitigate untrustworthy outputs from LLMs, such as hallucinations, has made UQ a critical area of research\. UQ for free\-form generative models is uniquely challenging because a correct answer can be expressed in countless semantically equivalent ways\([Lin et al\., 2023](https://arxiv.org/html/2610.00083#bib.bib10);[Kuhn et al\., 2023a](https://arxiv.org/html/2610.00083#bib.bib9)\)\. This renders early methods like predictive entropy \(PE\) insufficient, as they often misinterpret this benign lexical variance as genuine semantic uncertainty\([Kuhn et al\., 2023a](https://arxiv.org/html/2610.00083#bib.bib9)\)\. To address this, a significant body of work has shifted towards semantic\-aware UQ\. Semantic Entropy \(SE\)\([Kuhn et al\., 2023a](https://arxiv.org/html/2610.00083#bib.bib9)\)clusters equivalent outputs to estimate uncertainty, while Semantic Density \(SD\)\([Qiu and Miikkulainen, 2024](https://arxiv.org/html/2610.00083#bib.bib11)\)quantifies a response’s uncertainty by measuring its density within a semantic space\. In contrast, other methods probe the internal states or consistency of the LLM\. Deg\([Lin et al\., 2023](https://arxiv.org/html/2610.00083#bib.bib10)\)and its successor INSIDE\([Chen et al\., 2024](https://arxiv.org/html/2610.00083#bib.bib6)\)analyze consistency across multiple generations to quantify uncertainty from a black\-box perspective\. Furthermore, Shifting Attention to Relevance \(SAR\)\([Duan et al\., 2024a](https://arxiv.org/html/2610.00083#bib.bib12)\)addresses the generative imbalance by assigning more weight to semantically relevant parts of a generation\. In more complex scenarios, UProp\([Duan et al\., 2025](https://arxiv.org/html/2610.00083#bib.bib7)\)introduces a framework to decompose and quantify uncertainty propagation in multi\-step decision processes\. Alternatively, G\-NLL\([Aichberger et al\., 2025](https://arxiv.org/html/2610.00083#bib.bib8)\)offers a computationally efficient UQ method based on the negative log\-likelihood of a single greedy\-decoded output, challenging the necessity of multi\-sampling\. These diverse approaches highlight the evolution of LLM UQ from simple lexical metrics to more semantically robust, context\-aware, and computationally efficient solutions\. #### Linguistic Uncertainty Quantification\. Verbalized uncertainty, which employs natural language to articulate uncertainty, was pioneered by[Mielke et al\. \(2022\)](https://arxiv.org/html/2610.00083#bib.bib57);[Lin et al\. \(2022\)](https://arxiv.org/html/2610.00083#bib.bib18)\. Early black\-box evaluations demonstrated that inherent overconfidence can be mitigated via carefully designed prompts and aggregation methods\([Tian et al\., 2023](https://arxiv.org/html/2610.00083#bib.bib53);[Xiong et al\., 2023](https://arxiv.org/html/2610.00083#bib.bib19)\)\. To achieve anthropomimetic uncertainty, in which models emulate nuanced human expression to enhance trust\([Ulmer et al\., 2025](https://arxiv.org/html/2610.00083#bib.bib1)\), recent approaches actively optimize alignment via supervised fine\-tuning and RLHF\([Chaudhry et al\., 2024](https://arxiv.org/html/2610.00083#bib.bib59);[Liu et al\., 2024](https://arxiv.org/html/2610.00083#bib.bib67);[Leng et al\., 2025](https://arxiv.org/html/2610.00083#bib.bib62);[Li et al\., 2025b](https://arxiv.org/html/2610.00083#bib.bib60)\)\. Despite these optimization efforts, empirical studies reveal a persistent faithfulness gap\.[Ji et al\. \(2025\)](https://arxiv.org/html/2610.00083#bib.bib17)and[Yona et al\. \(2024\)](https://arxiv.org/html/2610.00083#bib.bib61)identify that verbal decisiveness often diverges from intrinsic probabilities or internal verbal uncertainty features \(VUF\), leading to confident hallucinations\. However,[Band et al\. \(2024\)](https://arxiv.org/html/2610.00083#bib.bib58);[Yoon et al\. \(2025\)](https://arxiv.org/html/2610.00083#bib.bib63)observe that utilizing chain\-of\-thought reasoning can improve calibration trajectories during long\-form generation\.[Zhou et al\. \(2023a\)](https://arxiv.org/html/2610.00083#bib.bib65)attribute sensitivity to pretraining mimicry, where markers often signal ignorance, while[Belem et al\. \(2024\)](https://arxiv.org/html/2610.00083#bib.bib64)find numerical interpretations biased by encoded priors\. Extending this to human\-AI interaction,[Kim et al\. \(2024a\)](https://arxiv.org/html/2610.00083#bib.bib66)demonstrates that the specific phrasing of these uncertainty expressions modulates user reliance and trust in decision\-making tasks\. Instead of expensive fine\-tuning, we treat verbal uncertainty as an inherent predictive signal\. We proposeVOCAL, which optimizes uncertainty mappings via a lightweight lookup table without parameter updates\. Unlike sampling\-based methods,VOCALextracts optimized uncertainty from a single response, offering a computationally efficient UQ alternative\. ## 3Preliminary: How LLMs Diverge from Humans in Verbal Uncertainty ### 3\.1Problem Statement: Uncertainty Quantification Uncertainty quantification \(UQ\) aims to measure the degree of doubt that a model exhibits with respect to its generations\. In the context of LLMs, UQ evaluates the doubt that an LLM parameterized by𝜽\{\\bm\{\\theta\}\}assigns to a generation𝒚∼p𝜽\(𝒚∣𝒙\)\{\\bm\{y\}\}\\sim p\_\{\{\\bm\{\\theta\}\}\}\(\{\\bm\{y\}\}\\mid\{\\bm\{x\}\}\), given an input𝒙\{\\bm\{x\}\}\. Formally, let𝒬\\mathcal\{Q\}denote a UQ method\. The corresponding uncertainty scoreqqassociated with𝒚\{\\bm\{y\}\}is defined asq=𝒬\(𝒚,𝒙,𝜽\)∈ℝq=\\mathcal\{Q\}\(\{\\bm\{y\}\},\{\\bm\{x\}\},\{\\bm\{\\theta\}\}\)\\in\\mathbb\{R\}\. The specific realization of𝒬\\mathcal\{Q\}varies across different UQ approaches, depending on the underlying assumptions and techniques employed\. In[AppendixA](https://arxiv.org/html/2610.00083#A1), we present the realizations of popular LLM UQ methods in detail\. Performance Evaluation\.The performance evaluation of UQ usually follows a “correctness prediction” manner, measuring the correlation between the calculated uncertainty score from a UQ method𝒬\\mathcal\{Q\}and the correctness of model generations, with metrics such as AUROC and hallucination detection accuracy\. A higher AUROC or detection accuracy means𝒬\\mathcal\{Q\}correctly predicts the correctness of model generations, indicating a good uncertainty estimator\. ### 3\.2Human Verbal Uncertainty and its Numerical Representation Humans usually express their uncertainty in verbal form, with uncertainty markers \(UMs\) such as “might”, or “probably”, which encode a speaker’s degree of uncertainty\. Formally, we denote by𝒬VU\\mathcal\{Q\}\_\{\\text\{VU\}\}a UQ that quantifies uncertainty from UMs\. Then, given a model generation𝒚\{\\bm\{y\}\}, its verbal uncertaintyqqis denoted byq𝒚=𝒬VU\(𝒱𝒚\)q\_\{\{\\bm\{y\}\}\}=\\mathcal\{Q\}\_\{\\text\{VU\}\}\(\\mathcal\{V\}\_\{\{\\bm\{y\}\}\}\), where𝒰𝒚=\{𝒖1,𝒖2,⋯\}\\mathcal\{U\}\_\{\{\\bm\{y\}\}\}=\\\{\{\\bm\{u\}\}\_\{1\},\{\\bm\{u\}\}\_\{2\},\\cdots\\\}are the extracted UMs from𝒚\{\\bm\{y\}\}\. However, there are two challenges blocking the quantitative evaluation: ①How to convert human UMs to numerical representations?, even though we obtained their numerical scores, ②how to aggregate numerical scores from multiple UMs? To address these challenges, we introduce the first large\-scale lookup table of human uncertainty,UM\-Lookuptable, that maps human UMs to numerical probabilities\. OurUM\-Lookupis grounded in foundational empirical studies from psychology and decision science, including the seminal works of[Lichtenstein and Newman \(1967\)](https://arxiv.org/html/2610.00083#bib.bib2),[Beyth\-Marom \(1982\)](https://arxiv.org/html/2610.00083#bib.bib3),[Wesson and Pulford \(2009\)](https://arxiv.org/html/2610.00083#bib.bib4), and the comprehensive meta\-analysis by[Vogel et al\. \(2022\)](https://arxiv.org/html/2610.00083#bib.bib5)\. Statistically, we collect 115 unique UMs, with each phrase’s value derived from an average of 336 human ratings\. This process yields a standardized confidence scale on a probabilistic\[0,1\]\[0,1\]range, containing expressions like “impossible” \(0\.0\), “tossup” \(0\.50\), and “definite” \(0\.99\)\. To remove the bias during the aggregation, we standardize the varied data formats from these sources, via direct probability estimates\([Lichtenstein and Newman, 1967](https://arxiv.org/html/2610.00083#bib.bib2)\), numerical ranges\([Beyth\-Marom, 1982](https://arxiv.org/html/2610.00083#bib.bib3)\), Likert scales\([Wesson and Pulford, 2009](https://arxiv.org/html/2610.00083#bib.bib4)\), and meta\-analytic weighted means\([Vogel et al\., 2022](https://arxiv.org/html/2610.00083#bib.bib5)\), resulting in a consistent structure of a phrase, its mean value, and its frequency \(N\)\. The detailed methodology for this normalization and aggregation, along with the complete human VUE lookup table, is provided in[AppendixB](https://arxiv.org/html/2610.00083#A2)\. With theUM\-Lookup, each UM could be effectively converted to a numerical representation\. In terms of the aggregation strategy of multiple UMs, empirical work shows that when people use multiple verbal probability terms in one statement, listeners \(and coders\) tend to average them into a single “middle” probability\([Budescu and Wallsten, 1995](https://arxiv.org/html/2610.00083#bib.bib20)\)\. Thus, we simply average all theUM\-Lookup\(UMs\)\\texttt\{UM\-Lookup\}\(\\text\{UMs\}\)as the final quantified uncertainty: q𝒚=𝒬VU\-H\(𝒱𝒚\)=1N∑i\(1−UM\-Lookup\(𝒖i\)\),q\_\{\{\\bm\{y\}\}\}=\\mathcal\{Q\}\_\{\\text\{VU\-H\}\}\(\\mathcal\{V\}\_\{\{\\bm\{y\}\}\}\)=\\frac\{1\}\{N\}\\sum\_\{i\}\{\(1\-\\texttt\{UM\-Lookup\}\(\{\\bm\{u\}\}\_\{i\}\)\)\},whereNNis the number of UMs from𝒚\{\\bm\{y\}\}and𝒖i\{\\bm\{u\}\}\_\{i\}is theii\-th UM in𝒱𝒚\\mathcal\{V\}\_\{\{\\bm\{y\}\}\}\. We use\(1−UM\-Lookup\(𝒖i\)CLOSE\(1\-\\texttt\{UM\-Lookup\}\(\{\\bm\{u\}\}\_\{i\}\)to convert from confidence to uncertainty\. In the rest of this paper, we denote by𝒬VU\-H\\mathcal\{Q\}\_\{\\text\{VU\-H\}\}the verbal UQ method equipped with human verbal uncertainty mappingUM\-Lookup\. Figure 2:The results of verbal uncertainty quantification𝒬VU\-H\\mathcal\{Q\}\_\{\\text\{VU\-H\}\}withUM\-Lookuptable collected from human\.𝒬VU\-H\\mathcal\{Q\}\_\{\\text\{VU\-H\}\}achieves non\-trivial UQ performance in many cases, indicating that LLMs share similar uncertainty expression as humans to a certain degree\. ### 3\.3Analytical Insights We evaluate GPT\-4o\([Achiam et al\., 2023](https://arxiv.org/html/2610.00083#bib.bib28)\)and DeepSeek\-V3\.1\([DeepSeek\-AI, 2024](https://arxiv.org/html/2610.00083#bib.bib32)\)over diverse datasets, such as GSM\-Hard\([Gao et al\., 2022](https://arxiv.org/html/2610.00083#bib.bib25)\), GSM8K\([Cobbe et al\., 2021](https://arxiv.org/html/2610.00083#bib.bib26)\), MedQA\([Jin et al\., 2020](https://arxiv.org/html/2610.00083#bib.bib24)\), PIQA\([Bisk et al\., 2020](https://arxiv.org/html/2610.00083#bib.bib27)\)\. We prompt LLMs to express verbal uncertainty and quantify uncertainty via𝒬VU\-H\\mathcal\{Q\}\_\{\\text\{VU\-H\}\}\. Specifically, we use two five\-shot strategies: a standard Chain\-of\-Thought\(CoT\) prompting\([Wei et al\., 2023](https://arxiv.org/html/2610.00083#bib.bib34)\)and CoT with verbal uncertainty prompting, where the latter incorporates the UM list \(see[AppendixC](https://arxiv.org/html/2610.00083#A3)for details\)\. Figure 3:The distributions of uncertainty markers expressed by LLMs\. We show that advanced LLMs, such as GPT\-4o, express uncertainty in a more diverse manner compared to small LLMs \(e\.g\., Llama\-3\.1\-8B\-Instruct and Llama\-3\.2\-3B\-Instruct\)\. This also reveals that small LLMs tend to be over confident\.In[SectionF\.1](https://arxiv.org/html/2610.00083#A6.SS1), we demonstrate that verbal uncertainty maintains general performance as the CoT\. As illustrated in[Figure7](https://arxiv.org/html/2610.00083#A6.F7), we evaluate model accuracy under both our verbal uncertainty prompting and a standard CoT baseline\. Across all evaluated models on the GSM8K dataset, from GPT\-4o to Llama\-3\.2\-3B\-Instruct, performance remains on par, with no statistically significant degradation in accuracy\. This result provides an important validation: the elicitation of verbal uncertainty does not impose a significant performance penalty, thereby preserving the models’ core problem\-solving efficacy\. 𝒬VU\-H\\mathcal\{Q\}\_\{\\text\{VU\-H\}\}Achieves Non\-Trivial UQ PerformanceIn[Figure2](https://arxiv.org/html/2610.00083#S3.F2), our primary finding is that quantifying uncertainty via a human\-sourced verbal lookup table,𝒬VU\-H\\mathcal\{Q\}\_\{\\text\{VU\-H\}\}, provides a meaningful signal for UQ\. This method achieves non\-trivial performance \(where AUROC is significantly greater than 0\.5\) in 7 out of the 8 evaluated model\-dataset configurations\. In several cases, its performance is highly competitive with or even surpasses popular UQ baselines\. For instance, with GPT\-4o on the SciQ dataset,𝒬VU\-H\\mathcal\{Q\}\_\{\\text\{VU\-H\}\}outperforms both Probability Entropy \(PE\) and Semantic Entropy \(SE\)\. Similarly, for DeepSeek\-V3\.1 on MedQA, our method’s performance is on par with both baselines\. However, we also identify clear limitations\. While often effective,𝒬VU\-H\\mathcal\{Q\}\_\{\\text\{VU\-H\}\}is frequently outperformed by PE and can fail notably, such as with GPT\-4o on GSM8K, where its AUROC falls below random chance\. We attribute these mixed results to a fundamental discrepancy: the uncertainty score assigned to a UM via our human\-sourceUM\-Lookuptable does not always match the LLM’s internal uncertainty state at the moment it generates that expression\. Advanced LLMs Express More Diverse Uncertainty Expressions\.With proper prompting, we find that advanced LLMs can express a diverse and frequent set of verbal uncertainty markers\. As shown in[Figure3](https://arxiv.org/html/2610.00083#S3.F3), large\-scale models such as GPT\-4o and DeepSeek\-V3\.1 achieve the highest diversity scores \(calculated by entropy\)\. Results are obtained on the GSM8K dataset\. Conversely, smaller models demonstrate a limited capacity for expressing nuanced uncertainty\. This tendency is consistent with the well\-documented challenge of overconfidence in LLMs\([Jiang et al\., 2021](https://arxiv.org/html/2610.00083#bib.bib33);[Xiong et al\., 2023](https://arxiv.org/html/2610.00083#bib.bib19);[Tian et al\., 2023](https://arxiv.org/html/2610.00083#bib.bib53)\)\. Such overconfidence is a critical issue, as it can lead to significant errors\([Zhou et al\., 2023b](https://arxiv.org/html/2610.00083#bib.bib35)\), reduce user trust\([Kim et al\., 2024b](https://arxiv.org/html/2610.00083#bib.bib36)\), and result in harmful downstream consequences\([Li, 2023](https://arxiv.org/html/2610.00083#bib.bib37)\)\. The complete distributions for all evaluated models are provided in[SectionF\.2](https://arxiv.org/html/2610.00083#A6.SS2)\. ## 4VOCAL: Optimizing the Numerical Scores of Verbal Uncertainty Markers for LLMs In[Section3\.3](https://arxiv.org/html/2610.00083#S3.SS3), we observe that although the human\-derived verbal uncertainty lookup table \(UM\-Lookup\) provides non\-trivial UQ performance, it often lags behind logit\- and sampling\-based baselines\. This naturally raises an important question: rather than relying solely on human estimates, can we instead learnUM\-Lookupthat are tailored to LLMs themselves? ### 4\.1Setup To achieve an LLM\-tailored probabilistic UM\-Lookup, we introduceVOCAL, a simple yet effective algorithm that optimizes the numerical levels of uncertainty markers for LLMs\.VOCALis a data\-driven method that learns appropriate uncertainty scores from model generations\. To obtain reliable estimations of these scores, we first collect diverse generations across multiple domains, such as mathematics \(GSM8K, GSM\-Hard\), science \(PIQA, SciQ\), and the medical domain \(MedQA\)\. We then apply a verbal uncertainty prompting strategy \(see Appendix[AppendixC](https://arxiv.org/html/2610.00083#A3)for detailed templates\) to elicit responses with explicit verbal uncertainty expressions and extract UMs together with the correctness of the corresponding generations\. ### 4\.2Optimized Uncertainty Levels for LLMs Formally, we denote by𝒰=\{𝒖1,𝒖2,…,𝒖N\}\\mathcal\{U\}=\\\{\{\\bm\{u\}\}\_\{1\},\{\\bm\{u\}\}\_\{2\},\\dots,\{\\bm\{u\}\}\_\{N\}\\\}the intended UM set extracted from LLM generations\. The optimization objective ofVOCALis to learn a suitable numerical score mappingcic\_\{i\}for each UMuiu\_\{i\}\. Formally, given a LLM generation𝒚\{\\bm\{y\}\}, the aggregated verbal uncertainty of𝒚\{\\bm\{y\}\}is then given byq𝒚=QVU\-L\(V𝒚\)=1N𝒚∑i=1N𝒚\(1−ci\),q\_\{\\bm\{y\}\}=Q\_\{\\text\{VU\-L\}\}\(V\_\{\\bm\{y\}\}\)=\\frac\{1\}\{N\_\{\\bm\{y\}\}\}\\sum\_\{i=1\}^\{N\_\{\\bm\{y\}\}\}\(1\-c\_\{i\}\),whereN𝒚N\_\{\\bm\{y\}\}is the number of UMs in𝒚\{\\bm\{y\}\}andQVU\-LQ\_\{\\text\{VU\-L\}\}denotes the LLM\-specific verbal uncertainty quantifier\.𝒖𝒚,i∈𝒰\{\\bm\{u\}\}\_\{\{\\bm\{y\}\},i\}\\in\\mathcal\{U\}is theii\-th UM in𝒚\{\\bm\{y\}\}\. The objective ofVOCALis to optimizecic\_\{i\}so thatq𝒚q\_\{\\bm\{y\}\}faithfully reflects the uncertainty of the LLM with respect to its generation𝒚\{\\bm\{y\}\}, in particular assigning higher uncertainty to incorrect generations and lower uncertainty to correct generations\. Then, the optimization objective ofVOCALcan be formalized in a Binary Cross\-Entropy \(BCE\) manner: ℒ\(𝒄\)=min𝒄𝔼\(𝒙,𝒚\)\[−zlog𝒄𝒚−\(1−z\)log\(1−𝒄𝒚\)\],\\mathcal\{L\}\(\{\\bm\{c\}\}\)=\\min\_\{\{\\bm\{c\}\}\}\\;\\mathbb\{E\}\_\{\(\{\\bm\{x\}\},\{\\bm\{y\}\}\)\}\\Big\[\-z\\log\{\\bm\{c\}\}\_\{\{\\bm\{y\}\}\}\-\(1\-z\)\\log\(1\-\{\\bm\{c\}\}\_\{\{\\bm\{y\}\}\}\)\\Big\],where𝐜\\mathbf\{c\}denotes the learnable uncertainty assignments for all markers,𝒄𝒚=1N𝒚∑i=1N𝒚ci\{\\bm\{c\}\}\_\{\{\\bm\{y\}\}\}=\\frac\{1\}\{N\_\{\\bm\{y\}\}\}\\sum\_\{i=1\}^\{N\_\{\\bm\{y\}\}\}c\_\{i\}is the aggregated uncertainty in generation𝒚\{\\bm\{y\}\}, andz=𝟙\[𝒚=𝒚∗\]∈\{0,1\}z=\\mathds\{1\}\[\{\{\\bm\{y\}\}=\{\\bm\{y\}\}^\{\*\}\}\]\\in\\\{0,1\\\}is the correctness indicator\. This formulation defines a convex optimization problem under the logistic loss, and ensures that the learned numerical scores yield effective verbal uncertainty\. ### 4\.3Semantic Smoothing via Graph Laplacian Regularization A key challenge in learning numerical scores for verbal uncertainty markers is data sparsity: some markers such as “likely” or “possible” appear frequently, while others like “faint chance” or “virtually certain” may occur rarely, making their learned values unstable\. Intuitively, semantically similar markers should share similar uncertainty levels, unless strong evidence from data suggests otherwise\. To achieve that, we adopt graph Laplacian regularization to enforce smoothness by encouraging semantically similar verbal uncertainty markers to share consistent scores\. This choice is consistent with established formulations in graph\-based learning, where the Laplacian energy is used to promote smoothness over similarity graphs, and with recent applications of semantic graph smoothing in NLP\([Fu et al\., 2022](https://arxiv.org/html/2610.00083#bib.bib54);[Maskey et al\., 2023](https://arxiv.org/html/2610.00083#bib.bib55);[Fettal et al\., 2024](https://arxiv.org/html/2610.00083#bib.bib56)\)\. Concretely, we construct a weighted similarity graphG=\(𝒰,E\)G=\(\\mathcal\{U\},E\), where each edge weight𝑾ij\{\\bm\{W\}\}\_\{ij\}captures the semantic similarity between markers𝒖i\{\\bm\{u\}\}\_\{i\}and𝒖j\{\\bm\{u\}\}\_\{j\}, i\.e\.,𝑾ij=s\(𝒖i,𝒖j\)\{\\bm\{W\}\}\_\{ij\}=s\(\{\\bm\{u\}\}\_\{i\},\{\\bm\{u\}\}\_\{j\}\)\. By default, we use 3\-gram Jaccard similarity as the semantic similarity measurements\(⋅,⋅\)s\(\\cdot,\\cdot\)\. Let𝑳=𝑫−𝑾\{\\bm\{L\}\}=\{\\bm\{D\}\}\-\{\\bm\{W\}\}be the corresponding graph Laplacian, with𝑫\{\\bm\{D\}\}as the degree matrix\. The semantic smoothing regularizer is then defined as ℒlap\(𝐜\)=γ𝐜⊤𝑳𝐜=γ∑i,j𝑾ij\(ci−cj\)2\.\\mathcal\{L\}\_\{\\text\{lap\}\}\(\\mathbf\{c\}\)=\\gamma\\,\\mathbf\{c\}^\{\\top\}\{\\bm\{L\}\}\\mathbf\{c\}=\\gamma\\sum\_\{i,j\}\{\\bm\{W\}\}\_\{ij\}\(c\_\{i\}\-c\_\{j\}\)^\{2\}\.where𝐜\\mathbf\{c\}denotes the vector of learnable numerical scores for all markers andγ\>0\\gamma\>0is a hyperparameter controlling the regularization strength\. This quadratic Dirichlet\-energy penalty is the standard form for promoting smoothness on graphs; in thep=2p\{=\}2case used here, the Laplacian regularizer is a convex quadratic\([Fu et al\., 2022](https://arxiv.org/html/2610.00083#bib.bib54)\), while related variants such as fractional\- andpp\-Laplacian formulations modulate the extent of smoothing and robustness\([Fu et al\., 2022](https://arxiv.org/html/2610.00083#bib.bib54);[Maskey et al\., 2023](https://arxiv.org/html/2610.00083#bib.bib55)\)\. By penalizing large discrepancies between semantically similar markers, this convex quadratic regularizer promotes smoother uncertainty assignments and leads to more robust verbal uncertainty, particularly for rare markers—empirically consistent with semantic graph smoothing on textual representations\([Fettal et al\., 2024](https://arxiv.org/html/2610.00083#bib.bib56)\)\. The overall optimization objective is defined as the joint minimization of the BCE loss and the semantic smoothing regularizer, i\.e\.,ℒ\(𝒄\)\+ℒlap\(𝐜\)\\mathcal\{L\}\(\{\\bm\{c\}\}\)\+\\mathcal\{L\}\_\{\\text\{lap\}\}\(\\mathbf\{c\}\)\. We utilize Adam to optimize our uncertainty scores\. In[Section5\.1](https://arxiv.org/html/2610.00083#S5.SS1), we provide detailed training protocols and hyperparameters\.VOCALconstructs theUM\-Lookupthrough a one\-time optimization and can be directly applied to test\-time generations for uncertainty quantification\. Unlike logits\- or sampling\-based UQ methods,VOCALdoes not require additional sampling or inference\-time computation\. In this way,VOCALprovides an efficient and effective approach for LLM uncertainty quantification\. ## 5Experiments ### 5\.1Experimental Setup Models\.Our evaluation is conducted on a set of state\-of\-the\-art LLMs, including GPT\-4o\([Achiam et al\., 2023](https://arxiv.org/html/2610.00083#bib.bib28)\), DeepSeek\-V3\.1\([DeepSeek\-AI, 2024](https://arxiv.org/html/2610.00083#bib.bib32)\), GPT\-3\.5\-Turbo\([Brown et al\., 2020](https://arxiv.org/html/2610.00083#bib.bib29)\), Qwen2\.5\-7B\-Instruct\([Qwen et al\., 2025](https://arxiv.org/html/2610.00083#bib.bib31)\), Qwen2\.5\-72B\-Instruct\([Qwen et al\., 2025](https://arxiv.org/html/2610.00083#bib.bib31)\), Llama\-3\.2\-3B\-Instruct and Meta\-Llama\-3\.1\-8B\-Instruct\([Grattafiori et al\., 2024](https://arxiv.org/html/2610.00083#bib.bib30)\)\. To collect LLM generations forVOCAL, we adopt a verbal uncertainty prompting strategy \(CoT with verbal uncertainty prompting\)\. For other UQ baselines, we adopt the naive CoT prompt strategy for all the LLMs\. Please refer to[AppendixC](https://arxiv.org/html/2610.00083#A3)for detailed prompt templates\. A full specification of our generative configurations is provided in[SectionD\.1](https://arxiv.org/html/2610.00083#A4.SS1)\. Figure 4:The evaluation results ofVOCALwhen comparing with human\-sourcedUM\-Lookup, i\.e\.,𝒬VU\-H\\mathcal\{Q\}\_\{\\text\{VU\-H\}\}\. It demonstrates thatVOCALproduces LLM\-tailoredUM\-Lookuptable\.Datasets and Training Data Curation\.We consider 6 popular question\-answering datasets: GSM\-Hard\([Gao et al\., 2022](https://arxiv.org/html/2610.00083#bib.bib25)\), GSM8K\([Cobbe et al\., 2021](https://arxiv.org/html/2610.00083#bib.bib26)\), MedQA\([Jin et al\., 2020](https://arxiv.org/html/2610.00083#bib.bib24)\), PIQA\([Bisk et al\., 2020](https://arxiv.org/html/2610.00083#bib.bib27)\), SciQ\([Welbl et al\., 2017](https://arxiv.org/html/2610.00083#bib.bib23)\), and Trivia QA\([Joshi et al\., 2017](https://arxiv.org/html/2610.00083#bib.bib22)\)\. For a complete description of the datasets, please refer to[SectionD\.2](https://arxiv.org/html/2610.00083#A4.SS2)\. We randomly select 300 questions from each dataset to curate the training set ofVOCALand randomly select 1,000 questions in the rest of each dataset\. We will introduce the sample efficiency in this section\. Hyperparameters\.By default, we set the graph Laplacian regularization strength toγ=5×10−3\\gamma=5\\times 10^\{\-3\}and use a learning rate of1×10−31\\times 10^\{\-3\}\. Training is conducted for up to 100 epochs with early stopping, where optimization terminates if the loss does not decrease within the most recent 10 epochs\. In[AppendixG](https://arxiv.org/html/2610.00083#A7), we provide the ablation study forγ\\gamma\. Figure 5:The evaluation results ofVOCALand multi\-sample based UQ methods\. It is worth noting that sampling\-based methods rely on semantic consistency calculations, which are expensive and introduce latency in real\-world deployment\. It is shown thatVOCALachieves comparable performance to sampling\-based UQ methods\.LLM UQ Baselines\.We consider popular logits\- and sampling\-based LLM UQ methods: Lexical Similarity \(LS\)\([Fomicheva et al\., 2020](https://arxiv.org/html/2610.00083#bib.bib21)\), Predictive Entropy \(PE\)\([Malinin and Gales, 2020](https://arxiv.org/html/2610.00083#bib.bib15)\), Semantic Entropy \(SE\)\([Kuhn et al\., 2023a](https://arxiv.org/html/2610.00083#bib.bib9)\), Deg\([Lin et al\., 2023](https://arxiv.org/html/2610.00083#bib.bib10)\), sentSAR\([Duan et al\., 2024a](https://arxiv.org/html/2610.00083#bib.bib12)\), G\-NLL\([Aichberger et al\., 2025](https://arxiv.org/html/2610.00083#bib.bib8)\), and Semantic Density \(SD\)\([Qiu and Miikkulainen, 2024](https://arxiv.org/html/2610.00083#bib.bib11)\)\. For sampling\-based UQ baselines, we generate 5 samples for each question with a temperature of 0\.8, and follow the default settings of these baselines\. Evaluation Metrics\.Consistent with prior work\([Kuhn et al\., 2023a](https://arxiv.org/html/2610.00083#bib.bib9)\), we evaluate uncertainty quantification by measuring its ability to predict the correctness of a model’s generated answers, using the Area Under the Receiver Operating Characteristic Curve \(AUROC\) as the evaluation metric\. VOCALis More Tailored for LLMs than Human\-SourcedUM\-Lookup\.As shown in[Figure4](https://arxiv.org/html/2610.00083#S5.F4),VOCALconsistently improves AUROC over the human\-sourcedUM\-Lookup\(𝒬VU\-H\\mathcal\{Q\}\_\{\\text\{VU\-H\}\}\) across models and datasets, indicating a more effective mapping from linguistic uncertainty cues to correctness\. The most representative gain appears on GSM8K, where the human lookup struggles: for GPT\-4o,VOCALboosts AUROC from 0\.41 to around 0\.6, demonstrating a substantial calibration improvement on math reasoning\. Similar, though smaller, gains are observed on other benchmarks \(e\.g\., GSM\-Hard and PIQA/TriviaQA\) for multiple models, suggesting that the benefit is not tied to a single dataset\. Overall, the key conclusion is that uncertainty expressions are model\-dependent, and learning an LLM\-tailored uncertainty lookup \(VOCAL\) is more reliable than relying on a fixed, human\-defined table, especially in failure modes where human heuristics do not align with the model’s actual confidence behavior\. Table 1:The comparison results betweenVOCALand single\-sample UQ baselines\. It is shown thatVOCALis significantly better than these methods\.DatasetModelG\-NLLPPLVOCALTrivia QAGPT\-4o0\.5380\.5750\.573Qwen2\.5\-72B\-Ins\.0\.6270\.6190\.645SciQGPT\-4o0\.6630\.6480\.700Qwen2\.5\-72B\-Ins\.0\.5680\.5550\.717GSM\-HardDeepSeek\-V3\.10\.5200\.5670\.715Qwen2\.5\-72B\-Ins0\.5070\.5800\.679 VOCALOutperforms 1\-Sample UQ Methods\.As demonstrated in[Table1](https://arxiv.org/html/2610.00083#S5.T1),VOCALsignificantly outperforms single\-sample UQ baselines such as G\-NLL and Perplexity \(PPL\)\. Our method achieves the highest AUROC score in 5 out of the 6 evaluated settings\. While PPL is marginally better on Trivia QA with GPT\-4o,VOCAL’s superiority is pronounced on more challenging reasoning datasets\. For instance, on GSM\-Hard with DeepSeek\-V3\.1,VOCALachieves an AUROC of 0\.715, a substantial improvement over both G\-NLL \(0\.520\) and PPL \(0\.567\)\. These results underscore the limitations of UQ methods that rely on a single greedy\-decoded output and highlight the robustness of our approach\. VOCALis Comparable to Multi\-Sampling UQ Methods\.[Figure5](https://arxiv.org/html/2610.00083#S5.F5)further shows thatVOCAL, despite using only a single response, can reach performance that is competitive with multi\-sample UQ baselines that require multiple generations and expensive consistency computations\. In particular,VOCALoften tracks the upper envelope of these sampling\-based methods and can even be the best performer in some settings \(e\.g\., on SciQ with Qwen2\.5\-72B,VOCALachieves the top AUROC of 0\.717\)\. At the same time, we observe occasional gaps on certain datasets, such as TriviaQA, where several multi\-sample metrics can surpassVOCAL\. Overall, these results support the main conclusion:VOCALoffers a strong accuracy–efficiency trade\-off, delivering near state\-of\-the\-art uncertainty discrimination in many cases while avoiding the latency and cost of multi\-sample semantic\-consistency pipelines, which makes it more practical for real\-world deployment\. #### Calibration performance\. We evaluate the calibration quality of different uncertainty estimation methods using Expected Calibration Error \(ECE\) on GSM\-Hard, GSM8K, and MedQA\. For the vanilla setting, we min\-max normalize the estimated uncertainty scores before computing ECE\. We also apply two standard post\-hoc calibration methods, temperature scaling \(TS\) and isotonic regression \(IR\), with all results obtained using 5\-fold cross\-validation\. As shown in[Table3](https://arxiv.org/html/2610.00083#S5.T3), VOCAL is competitive in the vanilla setting, achieving an average ECE of 0\.1314, which is close to the best baseline\. After calibration, VOCAL achieves stronger performance: VOCAL w/ TS obtains the lowest average ECE among all TS\-calibrated methods, while VOCAL w/ IR further reduces the average ECE to 0\.0281, achieving the best overall calibration performance\. These results indicate that VOCAL provides a reliable uncertainty signal and can be effectively improved by standard post\-hoc calibration methods, with IR yielding the strongest calibration performance\. Number of Training Samples\.In[Figure9](https://arxiv.org/html/2610.00083#A7.F9)\(left\), we find a strong positive correlation between the number of training samples and uncertainty quantification performance\. Our results show that increasing the training data from 100 to 500 samples leads to a significant AUROC score improvement from approximately 0\.52 to 0\.60, demonstrating the benefit of a larger training set\. Figure 6:Cross\-LLM transferability\. LLMs share a substantial common structure in verbal uncertainty expression\.Table 2:Mean probabilities of verbal uncertainty markers for GPT\-4o and humans, sorted by the GPT\-4o score\. Row colors indicate the relationship between probabilities: uncolored for aligned values \(within a 0\.05 tolerance\), and light gray where the GPT\-4o probability is higher or the human probability is higher\.PhraseGPT\-4o Prob\.Human Prob\.absolutely certain1\.0000\.920confident0\.8390\.900positive0\.8390\.900sure0\.8390\.830i think0\.7100\.630almost certain0\.6770\.920think0\.6450\.490can0\.3550\.570reasonable to assume0\.3550\.605very likely0\.3550\.853likely0\.0000\.655Do LLMs Share Similar Confidence Level?In[Figure6](https://arxiv.org/html/2610.00083#S5.F6), we conduct cross\-LLM transferability experiments by training and testingVOCALon different LLM responses over the SciQ dataset\. We follow the same training protocols \(e\.g\., hyperparameters\) as before\. It is shown that the learned verbal\-uncertainty indicators are largely transferable across models, with AUROC values remaining consistently in a competitive range \(0\.60–0\.72\)\. These results suggest that LLMs share a substantial common structure in verbal uncertainty expression, yet model\-specific expression differences still prevent a fully universal, one\-size\-fits\-all uncertainty mapping\. How LLMs diverge from humans in verbal uncertainty quantification?We compare our human\-sourcedUM\-Lookupwith a version optimized for GPT\-4o on the SciQ dataset to analyze the alignment between human and LLM uncertainty expressions \([Table2](https://arxiv.org/html/2610.00083#S5.T2)\)\. Our analysis reveals a significant divergence between the two, demonstrating that LLMs are not aligned with human verbal uncertainty\. For instance, GPT\-4o expresses 0\.677 confidence for the phrase “almost certain”, a term humans use with far more confidence \(0\.92\), while conversely, it assigns a low probability to “very likely” \(0\.355\), which humans rate with high confidence \(0\.853\)\. This fundamental misalignment shows that human\-derived tables are not directly transferable to LLMs, opening a new research direction into developing model\-specific quantification methods like VOCAL\. Table 3:Expected Calibration Error \(ECE\) of GPT\-4o on GSM\-Hard, GSM8K, and MedQA\. Lower ECE indicates better calibration\. For vanilla results, uncertainty scores are min\-max normalized before computing ECE\. Temperature scaling \(TS\) and isotonic regression \(IR\) are applied as post\-hoc calibration methods\.MethodVanillaw/ TSw/ IRGSM\-HardGSM8KMedQAAvgGSM\-HardGSM8KMedQAAvgGSM\-HardGSM8KMedQAAvgVOCAL0\.22120\.04620\.12690\.13140\.05450\.01970\.03390\.03600\.05380\.00510\.02550\.0281G\-NLL0\.07310\.14760\.15160\.12410\.03140\.06230\.09060\.06140\.03460\.02980\.05420\.0395PPL0\.06860\.25530\.13510\.15300\.09520\.13160\.06910\.09860\.05910\.02860\.03120\.0396SE0\.71000\.54900\.75680\.67190\.39640\.45890\.44680\.43400\.05910\.03900\.05120\.0498PE0\.20060\.05980\.13110\.13050\.04540\.02900\.04480\.03970\.06910\.03640\.06490\.0568Deg0\.22450\.04820\.11190\.12820\.12770\.02280\.04410\.06490\.05200\.02580\.04970\.0425SD0\.70910\.54950\.75760\.67210\.40060\.46150\.44930\.43710\.04490\.03220\.03100\.0360 ## 6Conclusion This work investigates how LLMs diverge from humans in expressing verbal uncertainty\. By constructing the first large\-scale lookup table of human uncertainty markers and introducingVOCAL, an optimization\-based alignment algorithm, we show that human\-derived mappings only partially capture model behavior, while LLM\-specific calibrations offer more reliable quantification\.VOCALachieves performance comparable to costly multi\-sample UQ methods with much lower computational overhead\. Our findings highlight the importance of grounding LLM uncertainty in verbal expressions, offering both practical benefits for trustworthy deployment and new directions for human–AI alignment research\. ## Limitations Verbal uncertainty, while intuitive, faces several limitations\. Its representation capacity is relatively weak, providing only coarse signals compared to probabilistic or semantic approaches\. The extraction and cleaning of uncertainty markers also introduce challenges, as model outputs may contain ambiguous or overlapping expressions\. Moreover, interpretations of verbal markers vary across domains and cultural contexts, limiting the generalizability of a singleUM\-Lookup\. These issues highlight promising directions for future work on more expressive, robust, and context\-aware verbal UQ methods\. ## Impact Statement This paper presents work whose goal is to advance the field of Machine Learning\. There are many potential societal consequences of our work, none of which we feel must be specifically highlighted here\. ## Acknowledgment This work was partially supported by the Amazon Research Award and Cisco Faculty Award\. ## References - Achiamet al\.\(2023\)J\. Achiam, S\. Adler, S\. Agarwal, L\. Ahmad, I\. Akkaya, F\. L\. Aleman, D\. Almeida, J\. Altenschmidt, S\. Altman, S\. Anadkat,et al\.Gpt\-4 technical report\.arXiv preprint arXiv:2303\.08774\.Cited by:[§1](https://arxiv.org/html/2610.00083#S1.p5.1),[§3\.3](https://arxiv.org/html/2610.00083#S3.SS3.p1.1),[§5\.1](https://arxiv.org/html/2610.00083#S5.SS1.p1.1)\. - Aichbergeret al\.\(2025\)L\. Aichberger, K\. Schweighofer, and S\. HochreiterRethinking uncertainty estimation in natural language generation\.InICLR Workshop: Quantify Uncertainty and Hallucination in Foundation Models: The Next Frontier in Reliable AI,Cited by:[§1](https://arxiv.org/html/2610.00083#S1.p6.1),[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px1.p1.1),[§5\.1](https://arxiv.org/html/2610.00083#S5.SS1.p4.1)\. - Asgariet al\.\(2025\)E\. Asgari, N\. Montaña\-Brown, M\. Dubois, S\. Khalil, J\. Balloch, J\. A\. Yeung, and D\. PimentaA framework to assess clinical safety and hallucination rates of llms for medical text summarisation\.NPJ Digital Medicine8\(1\),pp\. 274\.External Links:[Document](https://dx.doi.org/10.1038/s41746-025-01670-7),[Link](https://doi.org/10.1038/s41746-025-01670-7)Cited by:[§1](https://arxiv.org/html/2610.00083#S1.p1.1)\. - Bandet al\.\(2024\)N\. Band, X\. Li, T\. Ma, and T\. HashimotoLinguistic calibration of long\-form generations\.External Links:2404\.00474,[Link](https://arxiv.org/abs/2404.00474)Cited by:[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px2.p1.1)\. - Belemet al\.\(2024\)C\. G\. Belem, M\. Kelly, M\. Steyvers, S\. Singh, and P\. SmythPerceptions of linguistic uncertainty by language models and humans\.External Links:2407\.15814,[Link](https://arxiv.org/abs/2407.15814)Cited by:[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px2.p1.1)\. - Beyth\-Marom \(1982\)R\. Beyth\-MaromHow probable is probable? a numerical translation of verbal probability expressions\.Journal of forecasting1\(3\),pp\. 257–269\.Cited by:[Appendix B](https://arxiv.org/html/2610.00083#A2.p1.1),[§1](https://arxiv.org/html/2610.00083#S1.p5.1),[§3\.2](https://arxiv.org/html/2610.00083#S3.SS2.p2.1)\. - Bisket al\.\(2020\)Y\. Bisk, R\. Zellers, R\. L\. Bras, J\. Gao, and Y\. ChoiPIQA: reasoning about physical commonsense in natural language\.InThirty\-Fourth AAAI Conference on Artificial Intelligence,Cited by:[§D\.2](https://arxiv.org/html/2610.00083#A4.SS2.p1.1),[§3\.3](https://arxiv.org/html/2610.00083#S3.SS3.p1.1),[§5\.1](https://arxiv.org/html/2610.00083#S5.SS1.p2.1)\. - Brownet al\.\(2020\)T\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. D\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell,et al\.Language models are few\-shot learners\.Advances in neural information processing systems33,pp\. 1877–1901\.Cited by:[§5\.1](https://arxiv.org/html/2610.00083#S5.SS1.p1.1)\. - Budescu and Wallsten \(1995\)D\. V\. Budescu and T\. S\. WallstenProcessing linguistic probabilities: general principles and empirical evidence\.InPsychology of learning and motivation,Vol\.32,pp\. 275–318\.Cited by:[§3\.2](https://arxiv.org/html/2610.00083#S3.SS2.p3.1)\. - Chaudhryet al\.\(2024\)A\. Chaudhry, S\. Thiagarajan, and D\. GorurFinetuning language models to emit linguistic expressions of uncertainty\.External Links:2409\.12180,[Link](https://arxiv.org/abs/2409.12180)Cited by:[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px2.p1.1)\. - Chenet al\.\(2024\)C\. Chen, K\. Liu, Z\. Chen, Y\. Gu, Y\. Wu, M\. Tao, Z\. Fu, and J\. YeINSIDE: llms’ internal states retain the power of hallucination detection\.arXiv preprint arXiv:2402\.03744\.Cited by:[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px1.p1.1)\. - Cobbeet al\.\(2021\)K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano, C\. Hesse, and J\. SchulmanTraining verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168\.Cited by:[§D\.2](https://arxiv.org/html/2610.00083#A4.SS2.p1.1),[§3\.3](https://arxiv.org/html/2610.00083#S3.SS3.p1.1),[§5\.1](https://arxiv.org/html/2610.00083#S5.SS1.p2.1)\. - Colomboet al\.\(2024\)P\. Colombo, T\. P\. Pires, M\. Boudiaf, D\. Culver, R\. Melo, C\. Corro, A\. F\. T\. Martins, F\. Esposito, V\. L\. Raposo, S\. Morgado, and M\. DesaSaulLM\-7b: a pioneering large language model for law\.External Links:2403\.03883,[Link](https://arxiv.org/abs/2403.03883)Cited by:[§1](https://arxiv.org/html/2610.00083#S1.p1.1)\. - Daset al\.\(2025\)A\. B\. Das, S\. Ahmed, and S\. K\. SakibHallucinations and key information extraction in medical texts: a comprehensive assessment of open\-source large language models\.External Links:2504\.19061,[Link](https://arxiv.org/abs/2504.19061)Cited by:[§1](https://arxiv.org/html/2610.00083#S1.p1.1)\. - DeepSeek\-AI \(2024\)DeepSeek\-AIDeepSeek\-v3 technical report\.External Links:2412\.19437,[Link](https://arxiv.org/abs/2412.19437)Cited by:[§3\.3](https://arxiv.org/html/2610.00083#S3.SS3.p1.1),[§5\.1](https://arxiv.org/html/2610.00083#S5.SS1.p1.1)\. - Duanet al\.\(2024a\)J\. Duan, H\. Cheng, S\. Wang, A\. Zavalny, C\. Wang, R\. Xu, B\. Kailkhura, and K\. XuShifting attention to relevance: towards the predictive uncertainty quantification of free\-form large language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 5050–5063\.Cited by:[§1](https://arxiv.org/html/2610.00083#S1.p1.1),[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px1.p1.1),[§5\.1](https://arxiv.org/html/2610.00083#S5.SS1.p4.1)\. - Duanet al\.\(2025\)J\. Duan, J\. Diffenderfer, S\. Madireddy, T\. Chen, B\. Kailkhura, and K\. XuUProp: investigating the uncertainty propagation of llms in multi\-step agentic decision\-making\.arXiv preprint arXiv:2506\.17419\.Cited by:[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px1.p1.1)\. - Duanet al\.\(2023\)J\. Duan, F\. Kong, S\. Wang, X\. Shi, and K\. XuAre diffusion models vulnerable to membership inference attacks?\.InInternational Conference on Machine Learning,pp\. 8717–8730\.Cited by:[§1](https://arxiv.org/html/2610.00083#S1.p1.1)\. - Duanet al\.\(2024b\)J\. Duan, R\. Zhang, J\. Diffenderfer, B\. Kailkhura, L\. Sun, E\. Stengel\-Eskin, M\. Bansal, T\. Chen, and K\. XuGtbench: uncovering the strategic reasoning capabilities of llms via game\-theoretic evaluations\.Advances in Neural Information Processing Systems37,pp\. 28219–28253\.Cited by:[§1](https://arxiv.org/html/2610.00083#S1.p1.1)\. - Farquharet al\.\(2024\)S\. Farquhar, J\. Kossen, L\. Kuhn, and Y\. GalDetecting hallucinations in large language models using semantic entropy\.Nature630\(8017\),pp\. 625–630\.Cited by:[§1](https://arxiv.org/html/2610.00083#S1.p2.1)\. - Fettalet al\.\(2024\)C\. Fettal, L\. Labiod, and M\. NadifMore discriminative sentence embeddings via semantic graph smoothing\.arXiv preprint arXiv:2402\.12890\.Cited by:[§4\.3](https://arxiv.org/html/2610.00083#S4.SS3.p2.1),[§4\.3](https://arxiv.org/html/2610.00083#S4.SS3.p2.2)\. - Fomichevaet al\.\(2020\)M\. Fomicheva, S\. Sun, L\. Yankovskaya, F\. Blain, F\. Guzmán, M\. Fishel, N\. Aletras, V\. Chaudhary, and L\. SpeciaUnsupervised quality estimation for neural machine translation\.Transactions of the Association for Computational Linguistics8,pp\. 539–555\.Cited by:[§5\.1](https://arxiv.org/html/2610.00083#S5.SS1.p4.1)\. - Fuet al\.\(2022\)G\. Fu, P\. Zhao, and Y\. Bianpp\-Laplacian based graph neural networks\.InInternational conference on machine learning,pp\. 6878–6917\.Cited by:[§4\.3](https://arxiv.org/html/2610.00083#S4.SS3.p2.1),[§4\.3](https://arxiv.org/html/2610.00083#S4.SS3.p2.2)\. - Gaoet al\.\(2022\)L\. Gao, A\. Madaan, S\. Zhou, U\. Alon, P\. Liu, Y\. Yang, J\. Callan, and G\. NeubigPAL: program\-aided language models\.arXiv preprint arXiv:2211\.10435\.Cited by:[§D\.2](https://arxiv.org/html/2610.00083#A4.SS2.p1.1),[§3\.3](https://arxiv.org/html/2610.00083#S3.SS3.p1.1),[§5\.1](https://arxiv.org/html/2610.00083#S5.SS1.p2.1)\. - Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. D\. Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. MaThe llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[§5\.1](https://arxiv.org/html/2610.00083#S5.SS1.p1.1)\. - Jiet al\.\(2025\)Z\. Ji, L\. Yu, Y\. Koishekenov, Y\. Bang, A\. Hartshorn, A\. Schelten, C\. Zhang, P\. Fung, and N\. CanceddaCalibrating verbal uncertainty as a linear feature to reduce hallucinations\.arXiv preprint arXiv:2503\.14477\.Cited by:[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px2.p1.1)\. - Jianget al\.\(2021\)Z\. Jiang, J\. Araki, H\. Ding, and G\. NeubigHow can we know when language models know? on the calibration of language models for question answering\.Transactions of the Association for Computational Linguistics9,pp\. 962–977\.Cited by:[§3\.3](https://arxiv.org/html/2610.00083#S3.SS3.p5.1)\. - Jinet al\.\(2020\)D\. Jin, E\. Pan, N\. Oufattole, W\. Weng, H\. Fang, and P\. SzolovitsWhat disease does this patient have? a large\-scale open domain question answering dataset from medical exams\.arXiv preprint arXiv:2009\.13081\.Cited by:[§D\.2](https://arxiv.org/html/2610.00083#A4.SS2.p1.1),[§3\.3](https://arxiv.org/html/2610.00083#S3.SS3.p1.1),[§5\.1](https://arxiv.org/html/2610.00083#S5.SS1.p2.1)\. - Joshiet al\.\(2017\)M\. Joshi, E\. Choi, D\. S\. Weld, and L\. ZettlemoyerTriviaqa: a large scale distantly supervised challenge dataset for reading comprehension\.arXiv preprint arXiv:1705\.03551\.Cited by:[§D\.2](https://arxiv.org/html/2610.00083#A4.SS2.p1.1),[§5\.1](https://arxiv.org/html/2610.00083#S5.SS1.p2.1)\. - Kimet al\.\(2024a\)S\. S\. Y\. Kim, Q\. V\. Liao, M\. Vorvoreanu, S\. Ballard, and J\. W\. Vaughan"I’m not sure, but…": examining the impact of large language models’ uncertainty expression on user reliance and trust\.InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency,FAccT ’24,New York, NY, USA,pp\. 822–835\.External Links:ISBN 9798400704505,[Link](https://doi.org/10.1145/3630106.3658941),[Document](https://dx.doi.org/10.1145/3630106.3658941)Cited by:[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px2.p1.1)\. - Kimet al\.\(2024b\)S\. S\. Kim, Q\. V\. Liao, M\. Vorvoreanu, S\. Ballard, and J\. W\. Vaughan" I’m not sure, but…": examining the impact of large language models’ uncertainty expression on user reliance and trust\.InProceedings of the 2024 ACM conference on fairness, accountability, and transparency,pp\. 822–835\.Cited by:[§3\.3](https://arxiv.org/html/2610.00083#S3.SS3.p5.1)\. - Kossenet al\.\(2024\)J\. Kossen, J\. Han, M\. Razzak, L\. Schut, S\. Malik, and Y\. GalSemantic entropy probes: robust and cheap hallucination detection in llms\.arXiv preprint arXiv:2406\.15927\.Cited by:[§1](https://arxiv.org/html/2610.00083#S1.p2.1)\. - Kuhnet al\.\(2023a\)L\. Kuhn, Y\. Gal, and S\. FarquharSemantic uncertainty: linguistic invariances for uncertainty estimation in natural language generation\.arXiv preprint arXiv:2302\.09664\.Cited by:[§1](https://arxiv.org/html/2610.00083#S1.p1.1),[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px1.p1.1),[§5\.1](https://arxiv.org/html/2610.00083#S5.SS1.p4.1),[§5\.1](https://arxiv.org/html/2610.00083#S5.SS1.p5.1)\. - Kuhnet al\.\(2023b\)L\. Kuhn, Y\. Gal, and S\. FarquharSemantic uncertainty: linguistic invariances for uncertainty estimation in natural language generation\.External Links:2302\.09664,[Link](https://arxiv.org/abs/2302.09664)Cited by:[§1](https://arxiv.org/html/2610.00083#S1.p2.1)\. - Kuhnet al\.\(2023c\)L\. Kuhn, Y\. Gal, and S\. FarquharSemantic uncertainty: linguistic invariances for uncertainty estimation in natural language generation\.InThe Eleventh International Conference on Learning Representations,Cited by:[Appendix A](https://arxiv.org/html/2610.00083#A1.p1.2)\. - Lenget al\.\(2025\)J\. Leng, C\. Huang, B\. Zhu, and J\. HuangTaming overconfidence in llms: reward calibration in rlhf\.External Links:2410\.09724,[Link](https://arxiv.org/abs/2410.09724)Cited by:[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px2.p1.1)\. - Liet al\.\(2025a\)X\. Li, Z\. Yu, Z\. Zhang, Y\. Zhuang, S\. Shah, N\. Sadagopan, and A\. BeniwalSemantic volume: quantifying and detecting both external and internal uncertainty in llms\.arXiv preprint arXiv:2502\.21239\.Cited by:[§1](https://arxiv.org/html/2610.00083#S1.p2.1)\. - Liet al\.\(2025b\)Y\. Li, M\. Xiong, J\. Wu, and B\. HooiConfTuner: training large language models to express their confidence verbally\.External Links:2508\.18847,[Link](https://arxiv.org/abs/2508.18847)Cited by:[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px2.p1.1)\. - Li \(2023\)Z\. LiThe dark side of chatgpt: legal and ethical challenges from stochastic parrots and hallucination\.arXiv preprint arXiv:2304\.14347\.Cited by:[§3\.3](https://arxiv.org/html/2610.00083#S3.SS3.p5.1)\. - Lichtenstein and Newman \(1967\)S\. Lichtenstein and J\. R\. NewmanEmpirical scaling of common verbal phrases associated with numerical probabilities\.Psychonomic science9\(10\),pp\. 563–564\.Cited by:[Appendix B](https://arxiv.org/html/2610.00083#A2.p1.1),[§1](https://arxiv.org/html/2610.00083#S1.p5.1),[§3\.2](https://arxiv.org/html/2610.00083#S3.SS2.p2.1)\. - Linet al\.\(2022\)S\. Lin, J\. Hilton, and O\. EvansTeaching models to express their uncertainty in words\.arXiv preprint arXiv:2205\.14334\.Cited by:[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px2.p1.1)\. - Linet al\.\(2023\)Z\. Lin, S\. Trivedi, and J\. SunGenerating with confidence: uncertainty quantification for black\-box large language models\.arXiv preprint arXiv:2305\.19187\.Cited by:[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px1.p1.1),[§5\.1](https://arxiv.org/html/2610.00083#S5.SS1.p4.1)\. - Liuet al\.\(2024\)S\. Liu, Z\. Li, X\. Liu, R\. Zhan, D\. F\. Wong, L\. S\. Chao, and M\. ZhangCan LLMs learn uncertainty on their own? expressing uncertainty effectively in a self\-training manner\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 21635–21645\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.1205/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.1205)Cited by:[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px2.p1.1)\. - Maet al\.\(2025\)H\. Ma, J\. Chen, J\. T\. Zhou, G\. Wang, and C\. ZhangEstimating llm uncertainty with evidence\.External Links:2502\.00290,[Link](https://arxiv.org/abs/2502.00290)Cited by:[§1](https://arxiv.org/html/2610.00083#S1.p2.1)\. - Malinin and Gales \(2020\)A\. Malinin and M\. GalesUncertainty estimation in autoregressive structured prediction\.arXiv preprint arXiv:2002\.07650\.Cited by:[Appendix A](https://arxiv.org/html/2610.00083#A1.p1.1),[§1](https://arxiv.org/html/2610.00083#S1.p1.1),[§1](https://arxiv.org/html/2610.00083#S1.p2.1),[§5\.1](https://arxiv.org/html/2610.00083#S5.SS1.p4.1)\. - Maskeyet al\.\(2023\)S\. Maskey, R\. Paolino, A\. Bacho, and G\. KutyniokA fractional graph laplacian approach to oversmoothing\.Advances in Neural Information Processing Systems36,pp\. 13022–13063\.Cited by:[§4\.3](https://arxiv.org/html/2610.00083#S4.SS3.p2.1),[§4\.3](https://arxiv.org/html/2610.00083#S4.SS3.p2.2)\. - McCabeet al\.\(2025\)L\. H\. McCabe, R\. Melamed, T\. Hartvigsen, and H\. H\. HuangEstimating semantic alphabet size for llm uncertainty quantification\.External Links:2509\.14478,[Link](https://arxiv.org/abs/2509.14478)Cited by:[§1](https://arxiv.org/html/2610.00083#S1.p2.1)\. - Mielkeet al\.\(2022\)S\. J\. Mielke, A\. Szlam, E\. Dinan, and Y\. BoureauReducing conversational agents’ overconfidence through linguistic calibration\.External Links:2012\.14983,[Link](https://arxiv.org/abs/2012.14983)Cited by:[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px2.p1.1)\. - Nguyenet al\.\(2025\)D\. Nguyen, A\. Payani, and B\. MirzasoleimanBeyond semantic entropy: boosting llm uncertainty quantification with pairwise semantic similarity\.External Links:2506\.00245,[Link](https://arxiv.org/abs/2506.00245)Cited by:[§1](https://arxiv.org/html/2610.00083#S1.p2.1)\. - Qiu and Miikkulainen \(2024\)X\. Qiu and R\. MiikkulainenSemantic density: uncertainty quantification for large language models through confidence measurement in semantic space\.Advances in neural information processing systems37,pp\. 134507–134533\.Cited by:[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px1.p1.1),[§5\.1](https://arxiv.org/html/2610.00083#S5.SS1.p4.1)\. - Qwenet al\.\(2025\)Qwen, :, A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, H\. Wei, H\. Lin, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Lin, K\. Dang, K\. Lu, K\. Bao, K\. Yang, L\. Yu, M\. Li, M\. Xue, P\. Zhang, Q\. Zhu, R\. Men, R\. Lin, T\. Li, T\. Tang, T\. Xia, X\. Ren, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Cui, Z\. Zhang, and Z\. QiuQwen2\.5 technical report\.External Links:2412\.15115,[Link](https://arxiv.org/abs/2412.15115)Cited by:[§5\.1](https://arxiv.org/html/2610.00083#S5.SS1.p1.1)\. - Sriramananet al\.\(2024\)G\. Sriramanan, S\. Bharti, V\. S\. Sadasivan, S\. Saha, P\. Kattakinda, and S\. FeiziLlm\-check: investigating detection of hallucinations in large language models\.Advances in Neural Information Processing Systems37,pp\. 34188–34216\.Cited by:[§1](https://arxiv.org/html/2610.00083#S1.p2.1)\. - Thapaet al\.\(2025\)R\. Thapa, Q\. Wu, K\. Wu, H\. Zhang, A\. Zhang, E\. Wu, H\. Ye, S\. Bedi, N\. Aresh, J\. Boen, S\. Reddy, B\. Athiwaratkun, S\. L\. Song, and J\. ZouDisentangling reasoning and knowledge in medical large language models\.External Links:2505\.11462,[Link](https://arxiv.org/abs/2505.11462)Cited by:[§1](https://arxiv.org/html/2610.00083#S1.p1.1)\. - Tianet al\.\(2023\)K\. Tian, E\. Mitchell, A\. Zhou, A\. Sharma, R\. Rafailov, H\. Yao, C\. Finn, and C\. D\. ManningJust ask for calibration: strategies for eliciting calibrated confidence scores from language models fine\-tuned with human feedback\.External Links:2305\.14975,[Link](https://arxiv.org/abs/2305.14975)Cited by:[§1](https://arxiv.org/html/2610.00083#S1.p2.1),[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px2.p1.1),[§3\.3](https://arxiv.org/html/2610.00083#S3.SS3.p5.1)\. - Ulmeret al\.\(2025\)D\. Ulmer, A\. Lorson, I\. Titov, and C\. HardmeierAnthropomimetic uncertainty: what verbalized uncertainty in language models is missing\.arXiv preprint arXiv:2507\.10587\.Cited by:[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px2.p1.1)\. - Vogelet al\.\(2022\)H\. Vogel, S\. Appelbaum, H\. Haller, and T\. OstermannThe interpretation of verbal probabilities: a systematic literature review and meta\-analysis\.German Medical Data Sciences 2022–Future Medicine: More Precise, More Integrative, More Sustainable\!,pp\. 9–16\.Cited by:[Appendix B](https://arxiv.org/html/2610.00083#A2.p1.1),[§1](https://arxiv.org/html/2610.00083#S1.p6.1),[§3\.2](https://arxiv.org/html/2610.00083#S3.SS2.p2.1)\. - Weiet al\.\(2023\)J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, B\. Ichter, F\. Xia, E\. Chi, Q\. Le, and D\. ZhouChain\-of\-thought prompting elicits reasoning in large language models\.External Links:2201\.11903,[Link](https://arxiv.org/abs/2201.11903)Cited by:[§3\.3](https://arxiv.org/html/2610.00083#S3.SS3.p1.1)\. - Welblet al\.\(2017\)J\. Welbl, N\. F\. Liu, and M\. GardnerCrowdsourcing multiple choice science questions\.arXiv preprint arXiv:1707\.06209\.Cited by:[§D\.2](https://arxiv.org/html/2610.00083#A4.SS2.p1.1),[§1](https://arxiv.org/html/2610.00083#S1.p5.1),[§5\.1](https://arxiv.org/html/2610.00083#S5.SS1.p2.1)\. - Wesson and Pulford \(2009\)C\. J\. Wesson and B\. D\. PulfordVerbal expressions of confidence and doubt\.Psychological Reports105\(1\),pp\. 151–160\.Cited by:[Appendix B](https://arxiv.org/html/2610.00083#A2.p1.1),[§1](https://arxiv.org/html/2610.00083#S1.p5.1),[§3\.2](https://arxiv.org/html/2610.00083#S3.SS2.p2.1)\. - Xieet al\.\(2023\)Q\. Xie, W\. Han, X\. Zhang, Y\. Lai, M\. Peng, A\. Lopez\-Lira, and J\. HuangPIXIU: a large language model, instruction data and evaluation benchmark for finance\.External Links:2306\.05443,[Link](https://arxiv.org/abs/2306.05443)Cited by:[§1](https://arxiv.org/html/2610.00083#S1.p1.1)\. - Xionget al\.\(2023\)M\. Xiong, Z\. Hu, X\. Lu, Y\. Li, J\. Fu, J\. He, and B\. HooiCan llms express their uncertainty? an empirical evaluation of confidence elicitation in llms\.arXiv preprint arXiv:2306\.13063\.Cited by:[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px2.p1.1),[§3\.3](https://arxiv.org/html/2610.00083#S3.SS3.p5.1)\. - Yanget al\.\(2024\)A\. Yang, B\. Zhang, B\. Hui, B\. Gao, B\. Yu, C\. Li, D\. Liu, J\. Tu, J\. Zhou, J\. Lin, K\. Lu, M\. Xue, R\. Lin, T\. Liu, X\. Ren, and Z\. ZhangQwen2\.5\-math technical report: toward mathematical expert model via self\-improvement\.External Links:2409\.12122,[Link](https://arxiv.org/abs/2409.12122)Cited by:[§1](https://arxiv.org/html/2610.00083#S1.p1.1)\. - Yanget al\.\(2025\)Y\. Yang, H\. Yoo, and H\. LeeMAQA: evaluating uncertainty quantification in llms regarding data uncertainty\.External Links:2408\.06816,[Link](https://arxiv.org/abs/2408.06816)Cited by:[§1](https://arxiv.org/html/2610.00083#S1.p2.1)\. - Yonaet al\.\(2024\)G\. Yona, R\. Aharoni, and M\. GevaCan large language models faithfully express their intrinsic uncertainty in words?\.External Links:2405\.16908,[Link](https://arxiv.org/abs/2405.16908)Cited by:[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px2.p1.1)\. - Yoonet al\.\(2025\)D\. Yoon, S\. Kim, S\. Yang, S\. Kim, S\. Kim, Y\. Kim, E\. Choi, Y\. Kim, and M\. SeoReasoning models better express their confidence\.External Links:2505\.14489,[Link](https://arxiv.org/abs/2505.14489)Cited by:[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px2.p1.1)\. - Zhouet al\.\(2023a\)K\. Zhou, D\. Jurafsky, and T\. HashimotoNavigating the grey area: how expressions of uncertainty and overconfidence affect language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 5506–5524\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.335/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.335)Cited by:[§2](https://arxiv.org/html/2610.00083#S2.SS0.SSS0.Px2.p1.1)\. - Zhouet al\.\(2023b\)K\. Zhou, D\. Jurafsky, and T\. HashimotoNavigating the grey area: how expressions of uncertainty and overconfidence affect language models\.arXiv preprint arXiv:2302\.13439\.Cited by:[§3\.3](https://arxiv.org/html/2610.00083#S3.SS3.p5.1)\. ## Appendix AUncertainty Quantification in LLMs For instance, from the Bayesian perspective, UQ can be derived by measuring the total uncertainty in the predictive distributionp𝜽\(𝒚∣𝒙\)p\_\{\{\\bm\{\\theta\}\}\}\(\{\\bm\{y\}\}\\mid\{\\bm\{x\}\}\), where a common choice is the Predictive Entropy \(PE\)\([Malinin and Gales, 2020](https://arxiv.org/html/2610.00083#bib.bib15)\), defined as 𝒬PE\(𝒙\)=∫p𝜽\(𝒚\|𝒙\)log\(p𝜽\(𝒚\|𝒙\)\)d𝒚≈−1N∑iNlogp𝜽\(𝒚\(i\)\|𝒙\),𝒚\(i\)∼p𝜽\(𝒚\|𝒙\),\\mathcal\{Q\}\_\{\\mathrm\{PE\}\}\(\{\\bm\{x\}\}\)=\\int\{p\_\{\{\\bm\{\\theta\}\}\}\(\{\\bm\{y\}\}\|\{\\bm\{x\}\}\)\\log\(p\_\{\{\\bm\{\\theta\}\}\}\(\{\\bm\{y\}\}\|\{\\bm\{x\}\}\)\)\}\\,d\{\\bm\{y\}\}\\approx\-\\frac\{1\}\{N\}\\sum\_\{i\}^\{N\}\{\\log\{p\_\{\{\\bm\{\\theta\}\}\}\(\{\\bm\{y\}\}^\{\(i\)\}\|\{\\bm\{x\}\}\)\}\},\\,\\,\{\\bm\{y\}\}^\{\(i\)\}\\sim p\_\{\{\\bm\{\\theta\}\}\}\(\{\\bm\{y\}\}\|\{\\bm\{x\}\}\),whereNNis the number of samples andp𝜽\(𝒚\(i\)\|𝒙\)=∏iLip𝜽\(zi\|z<i,𝒙\)p\_\{\{\\bm\{\\theta\}\}\}\(\{\\bm\{y\}\}^\{\(i\)\}\|\{\\bm\{x\}\}\)=\\prod\_\{i\}^\{L\_\{i\}\}p\_\{\{\\bm\{\\theta\}\}\}\(z\_\{i\}\|z\_\{<i\},\{\\bm\{x\}\}\)is the generative probability of𝒚\(i\)\{\\bm\{y\}\}^\{\(i\)\}with lengthLiL\_\{i\}\.ziz\_\{i\}is theii\-th token of𝒚\(i\)\{\\bm\{y\}\}^\{\(i\)\}\. Moreover,[Kuhn et al\. \(2023c\)](https://arxiv.org/html/2610.00083#bib.bib16)proposes Semantic Entropy \(SE\), which aggregates probability mass over semantic clusters of outputs: 𝒬SE\(𝒙\)=−1C∑iClog\(p𝜽\(𝒄i\|𝒙\)\),p𝜽\(𝒄i\|𝒙\)=∑𝒚∈𝒄ip𝜽\(𝒚\|𝒙\),\\mathcal\{Q\}\_\{\\mathrm\{SE\}\}\(\{\\bm\{x\}\}\)=\-\\frac\{1\}\{C\}\\sum\_\{i\}^\{C\}\{\\log\(p\_\{\{\\bm\{\\theta\}\}\}\(\{\\bm\{c\}\}\_\{i\}\|\{\\bm\{x\}\}\)\)\},\\,\\,p\_\{\{\\bm\{\\theta\}\}\}\(\{\\bm\{c\}\}\_\{i\}\|\{\\bm\{x\}\}\)=\\sum\_\{\{\\bm\{y\}\}\\in\{\\bm\{c\}\}\_\{i\}\}\{p\_\{\{\\bm\{\\theta\}\}\}\(\{\\bm\{y\}\}\|\{\\bm\{x\}\}\)\},whereCCis the number of semantic clusters and𝒄i\{\\bm\{c\}\}\_\{i\}is theii\-th cluster consisting of generations𝒚i\{\\bm\{y\}\}\_\{i\}sharing the same semantics\. These two examples illustrate how different realizations of𝒬\\mathcal\{Q\}target distinct aspects of output uncertainty\. ## Appendix BHuman Verbal Uncertainty Expression The lookup table presented below consolidates numerical probabilities for verbal uncertainty expressions \(VUEs\) from several key empirical studies\. The aggregation process involved several steps to harmonize the data\. For sources providing mean probability values, such as[Lichtenstein and Newman \(1967\)](https://arxiv.org/html/2610.00083#bib.bib2), the values were used directly \(e\.g\., "likely" with mean=0\.72\)\. For studies reporting ranges, like[Beyth\-Marom \(1982\)](https://arxiv.org/html/2610.00083#bib.bib3), we calculated the midpoint of the interquartile range to represent the central tendency \(e\.g\., "likely" \[0\.55, 0\.85\]→\\rightarrow0\.70\)\. Data from[Wesson and Pulford \(2009\)](https://arxiv.org/html/2610.00083#bib.bib4), originally on a 1–7 point scale, was linearly rescaled to the probabilistic range\[0,1\]\[0,1\]\. Meta\-analytic estimates from[Vogel et al\. \(2022\)](https://arxiv.org/html/2610.00083#bib.bib5)were incorporated to refine values and ensure cross\-study consistency\. The final probability for each VUE in[Table4](https://arxiv.org/html/2610.00083#A2.T4)was derived by averaging these processed values, weighted by study prominence and term frequency where applicable\. This table serves as the human\-grounded benchmark for our analysis\. Table 4:Full Lookup Table for Verbalized Uncertainty Expressions \(VUE\) with their associated probabilities and frequencies\.Uncertainty ExpressionUncertainty ProbabilityFrequency \(N\)Definite0\.990447\.0Certain0\.962905\.0Virtually certain0\.950447\.0Almost certain0\.920782\.0Absolutely certain0\.92096\.0Very high chance0\.91527\.0I know for a fact that it’s…0\.91096\.0I know it’s…0\.90096\.0Positive0\.90096\.0Confident0\.90096\.0Highly probable0\.8981081\.0Nearly certain0\.89527\.0No doubt0\.87096\.0Very probable0\.870187\.0Very likely0\.8531079\.0Most likely0\.85027\.0Close to certain0\.83527\.0Sure0\.83096\.0High chance0\.81027\.0I have no doubt, I mean I’m sure it’s…0\.81096\.0Reasonably certain0\.800447\.0Usually0\.770187\.0Fairly confident0\.76096\.0Reasonable assurance0\.750447\.0Remember0\.75096\.0Predictable0\.740146\.0Good chance0\.724858\.0Quite likely0\.717970\.0Meaningful chance0\.71527\.0Rather likely0\.690188\.0Probable0\.6822311\.0Believe0\.67096\.0Pretty good chance0\.670188\.0Fairly likely0\.660188\.0Likely0\.6552227\.0Suspect0\.64096\.0I would say it’s…0\.64096\.0I could be mistaken but I’m sure it’s…0\.64096\.0I think it’s…0\.63096\.0Reasonable chance0\.61527\.0One should assume0\.61027\.0It seems to me0\.60527\.0Reasonable to assume0\.60527\.0Non\-negligible chance0\.60027\.0I’m not completely confident, but I think it’s…0\.60096\.0Quite probable0\.600447\.0It seems0\.59027\.0Somewhat likely0\.590187\.0Rather0\.580124\.0Better than even0\.580187\.0I can’t say for sure, but I think it’s…0\.57096\.0One can expect0\.57027\.0I’m not certain, but it could be…0\.56096\.0Slight odds in favor0\.550185\.0I think it’s…\. but I can’t be sure\.0\.55096\.0Slightly more than half the time0\.550188\.0I guess it’s…0\.53096\.0I could be wrong, but I think it’s…0\.53096\.0I’m not sure, but it may be…0\.53096\.0Possible \(again?\)0\.520447\.0It’s…\. I think\.0\.52096\.0Fair chance0\.510188\.0Tossup0\.500188\.0Reasonably possible0\.500447\.0It could be0\.49527\.0May0\.49527\.0Think0\.49096\.0There is a chance0\.48527\.0One must consider0\.48027\.0Perhaps0\.478474\.0Could be0\.47096\.0Fighting chance0\.470186\.0I think it’s…\. isn’t it?0\.47096\.0Possible0\.4642663\.0Not inevitable0\.45527\.0Maybe0\.450670\.0Slight odds against0\.450185\.0I’m guessing, but I would say it’s…0\.45096\.0Slightly less than half the time0\.450188\.0Not quite even0\.440180\.0Inconclusive0\.430153\.0Don’t know0\.43096\.0Chance0\.420447\.0Not sure0\.42096\.0Not certain0\.400447\.0Possibly0\.380447\.0Can’t rule out entirely0\.36527\.0Uncertain0\.3561402\.0Chances are not great0\.34527\.0Somewhat unlikely0\.310186\.0Somewhat doubtful0\.300447\.0Small chance0\.29027\.0Low chance0\.28027\.0Fairly unlikely0\.250187\.0Doubtful0\.250474\.0Quite unlikely0\.2451193\.0Rather unlikely0\.225374\.0Not likely0\.213474\.0Not very probable0\.200187\.0Unlikely0\.1981752\.0Not probable0\.180559\.0Poor chance0\.18027\.0Seldom0\.160188\.0Not much chance0\.160186\.0Improbable0\.1451081\.0Very low chance0\.14027\.0Barely possible0\.130180\.0Faintly possible0\.130184\.0Very unlikely0\.1161304\.0Not possible0\.100559\.0Almost impossible0\.080559\.0Rare0\.070187\.0Remote0\.070447\.0Highly improbable0\.052851\.0Impossible0\.000559\.0 ## Appendix CPrompt LLMs to Express Verbal Uncertainty This appendix details the two Chain\-of\-Thought \(CoT\) system prompts used in our experiments\. The baselineStandard CoT Promptrequests a standard two\-field JSON answer\. In contrast, theCoT with Verbal Uncertainty Promptextends this by requiring the model to incorporate UMs into its response and to report these expressions in an additional ‘vue’ field within a three\-field JSON output\. CoT PromptYou are a helpful and conversational AI assistant\. Respond to questions in a natural, human\-like tone\.Your response MUST be in valid JSON format with these two fields:``` { "answer": "[Your conversational answer]", "final_answer": "[Your most specific answer]" } ``` The "final\_answer" should contain the most specific information possible, like a name, date, or place\. The "answer" should be a natural explanation, as if you’re talking to a friend\. CoT with Verbal Uncertainty PromptYou are a knowledgeable and conversational AI assistant\. Answer questions naturally with a human\-like tone\.Your response should include:1\.A natural, conversational answer that incorporates verbalized uncertainty expressions \(VUE\) naturally within the text2\.A VUE section that lists all the uncertainty phrases you used in your answer3\.A final\_answer section with the most specific answer you can provideIMPORTANT:You MUST respond in valid JSON format with exactly these three fields:``` { "answer": "[Your natural answer with embedded VUE expressions]", "vue": ["phrase1", "phrase2", "phrase3"], "final_answer": "[Your most specific answer]" } ``` In your answer, naturally include uncertainty expressions including: \{VUE\_LIST: ‘definite’, ‘certain’, ‘virtually certain’, ‘almost certain’, …\}Then in the vue field, provide an array of the uncertainty phrases you used\. In the final\_answer field, provide the most specific answer you can give \(e\.g\., a name, place, date, etc\.\)\. Make your answer sound natural and conversational, as if explaining to a friend\. Ensure your response is valid JSON that can be parsed\. ## Appendix DExperimental Settings ### D\.1Details of LLMs Generation All models were queried using two distinct configurations\. To assess correctness, we employed greedy decoding\. To quantify uncertainty, we utilized multinomial sampling to draw 5 samples at a temperature of 0\.8\. All generated outputs were constrained by a maximum length of 512 tokens and atop\_pvalue of 1\.0\. ### D\.2Datasets GSM8K[Cobbe et al\. \(2021\)](https://arxiv.org/html/2610.00083#bib.bib26)is a benchmark dataset featuring over 8,000 high\-quality grade school math word problems\. It is specifically designed to measure multi\-step quantitative reasoning, with a key feature being that problems require several reasoning steps to solve\.GSM\-Hard[Gao et al\. \(2022\)](https://arxiv.org/html/2610.00083#bib.bib25)is a challenging subset of GSM8K, curated to include only problems that necessitate the most complex and lengthy reasoning chains\.MedQA[Jin et al\. \(2020\)](https://arxiv.org/html/2610.00083#bib.bib24)is a large\-scale multiple\-choice dataset with over 11,000 questions derived from U\.S\. medical licensing exams, created to evaluate a model’s capacity for deep medical knowledge\.PIQA[Bisk et al\. \(2020\)](https://arxiv.org/html/2610.00083#bib.bib27)is a commonsense reasoning benchmark containing over 18,000 examples in its training and validation sets\. It is structured as a two\-choice task that tests a model’s understanding of physical interactions\.SciQ[Welbl et al\. \(2017\)](https://arxiv.org/html/2610.00083#bib.bib23)consists of approximately 13,700 crowdsourced science exam questions\. Each question is multiple\-choice and paired with a supporting text, testing both knowledge and comprehension\.TriviaQA[Joshi et al\. \(2017\)](https://arxiv.org/html/2610.00083#bib.bib22)is a high\-quality reading comprehension dataset that contains over 650k question\-answer pairs\. Its distinct challenge lies in requiring models to find answers within large, unstructured evidence documents\. ## Appendix EFailure Case Analysis In this section, we present failure cases of our method, including one where the model generates a correct answer with a low confidence score and another where it outputs an incorrect answer with a high confidence score\. These analyses help clarify when verbal uncertainty fails and point toward possible directions for addressing these issues\. ### E\.1Under\-Confidence Despite Correctness In the first case, the sampling\-based baseline shows high confidence because the outputs across different generations are consistent and match the ground truth\. In contrast, our method includes hedging expressions such as “not completely confident” or “I think,” which drive the confidence score much lower, even though the final answer is correct\. Case 1: DeepSeek\-V3\.1 Model on Dataset PIQAQuestion:How can I sneeze when I feel like I want to sneeze but can’t? Choices: \(1\) Look directly in front of you this will cause a reaction in your body and you will sneeze\. \(2\) Look directly into sunlight, this will cause a reaction in your body and you will sneeze\. Ground Truth Answer:Look directly into sunlight, this will cause a reaction in your body and you will sneeze\. Baseline Predictions: Answer: “If you’re trying to trigger a sneeze, looking directly into sunlight often works because it stimulates the optic nerve and can cause a sneezing reflex\.” Final Answer: Look directly into sunlight, this will cause a reaction in your body and you will sneeze\. Answer: “If you feel like you need to sneeze but can’t, try looking into bright sunlight—it often triggers a sneeze reflex\.” Final Answer: Look directly into sunlight, this will cause a reaction in your body and you will sneeze\. Answer: “If you’re trying to trigger a sneeze, looking directly into sunlight can often help because it stimulates a reflex in some people\.” Final Answer: Look directly into sunlight, this will cause a reaction in your body and you will sneeze\. Answer: “If you’re trying to sneeze but can’t, looking directly into sunlight often triggers a reaction that makes you sneeze\.” Final Answer: Look directly into sunlight, this will cause a reaction in your body and you will sneeze\. Answer: “If you feel like you need to sneeze but can’t, try looking into bright sunlight\. This often triggers a photic sneeze reflex in many people\.” Final Answer: Look directly into sunlight, this will cause a reaction in your body and you will sneeze\. VUE Output: Answer: I’mnot completely confident, butI thinklooking directly into sunlight is a common trick to trigger a sneeze\. Final Answer: Look directly into sunlight, this will cause a reaction in your body and you will sneeze\. VUE Phrases: \[“not completely confident”, “think”\] ### E\.2Over\-Confidence Despite Incorrectness In the second case, the sampling\-based baseline shows low confidence because the generated answers across different samples are inconsistent: some responses point to ectopic gastrin secretion, others to autoimmune destruction, and only one aligns with the ground truth\. This disagreement shows the uncertainty when the model answers this question and aligns with the incorrect response\. In contrast, our method uses high confidence expressions such as “fairly confident” or “most likely,” but the final answer is incorrect\. Case 2: DeepSeek\-V3\.1 Model on Dataset MedQAQuestion:An otherwise healthy 50\-year\-old man comes to the physician because of a 6\-month history of increasingly frequent episodes of upper abdominal pain, nausea, vomiting, and diarrhea\. He has had a 3\.2\-kg \(7\-lb\) weight loss during this time\. Physical examination shows bilateral pitting pedal edema\. An endoscopy shows prominent rugae in the gastric fundus\. Biopsy shows parietal cell atrophy\. Which of the following is the most likely underlying cause? Choices: \(1\) Serotonin\-secreting gastric tumor \(2\) Proliferation of gastric mucus\-producing cells \(3\) Excessive somatostatin secretion \(4\) Ectopic secretion of gastrin Ground Truth Answer:Proliferation of gastric mucus\-producing cells Baseline Predictions: Answer: “The findings—prominent gastric rugae, parietal cell atrophy, weight loss, and edema—point toward protein loss, likely due to gastric hypertrophy from ectopic gastrin secretion, as seen in Zollinger\-Ellison syndrome\.” Final Answer: Ectopic secretion of gastrin Answer: “This sounds like a classic case of pernicious anemia with gastric parietal cell atrophy, which is often autoimmune\-mediated\. The prominent rugae, weight loss, and edema point toward protein loss and malabsorption, likely due to the lack of intrinsic factor from parietal cell damage\.” Final Answer: Autoimmune destruction of parietal cells Answer: “This sounds like a classic case of pernicious anemia due to autoimmune destruction of parietal cells, leading to atrophic gastritis and vitamin B12 deficiency\. The prominent rugae, weight loss, and pedal edema are consistent with protein\-losing enteropathy from the gastric atrophy\.” Final Answer: Proliferation of gastric mucus\-producing cells Answer: “The symptoms and findings—like prominent rugae, parietal cell atrophy, weight loss, and edema—point toward a condition where there’s excessive gastrin secretion, often due to a gastrinoma\. This leads to high acid production, which can cause peptic ulcers and malabsorption issues, explaining the diarrhea and weight loss\. The edema might be from protein loss or other complications\. While other choices could affect the stomach, ectopic gastrin secretion fits best here\.” Final Answer: Ectopic secretion of gastrin Answer: “The symptoms and findings—like prominent rugae, parietal cell atrophy, weight loss, and edema—suggest a condition where there’s excessive gastrin secretion, leading to gastric hyperplasia and malabsorption\. This is most characteristic of Zollinger\-Ellison syndrome, often caused by a gastrin\-secreting tumor\.” Final Answer: Ectopic secretion of gastrin VUE Output: Answer:I’m fairly confident this describes pernicious anemia with gastric atrophy, which is most likely caused by autoimmune destruction of parietal cells leading to intrinsic factor deficiency and vitamin B12 malabsorption\. Final Answer: Autoimmune destruction of parietal cells\. Phrases: \[“fairly confident”, “most likely”\] ### E\.3Discussion\. These cases show that verbal expressions of uncertainty do not always align with a model’s internal confidence\. In some cases, hedging expression lowers the confidence even the predictions are correct\. In other cases, the model conveys strong certainty while producing incorrect responses, which undermines trust and reliability\. To address these challenges, future work should aim to capture uncertainty signals within specific domains and develop prompting strategies that foster clearer, more faithful representations of uncertainty\. ## Appendix FVerbal Uncertainty Quantification with HumanUM\-LookupTable ### F\.1Verbal Uncertainty Prompting Maintains General Performance In[Figure7](https://arxiv.org/html/2610.00083#A6.F7), we show that our verbal uncertainty prompting strategy does not significantly hurt the general performance of LLMs, which demonstrate the utility ofVOCALin applications\. Figure 7:Verbal uncertainty prompting maintains general performance\. ### F\.2Advanced LLMs Express Diverse Uncertainty Markers The UM distributions of each LLMs over all the datasets are presented in[Figure8](https://arxiv.org/html/2610.00083#A6.F8)\. Figure 8:Verbal uncertainty marker distributions of LLMs\. ## Appendix GOptimizedUM\-LookupTable To complement our analysis, we provide optimized lookup tables that map verbal uncertainty markers to numerical uncertainty values\. Specifically, Table[5](https://arxiv.org/html/2610.00083#A7.T5)presents the optimizedUM\-Lookupfor GPT\-4o on the SciQ dataset\. In addition, we report results for GPT\-3\.5\-Turbo on MedQA \([Table6](https://arxiv.org/html/2610.00083#A7.T6)\) and on SciQ \([Table7](https://arxiv.org/html/2610.00083#A7.T7)\)\. Semantic smoothingγ\\gammaAblation StudyOur analysis also reveals the model’s sensitivity to the semantic smoothing hyperparameter,γ\\gamma\. The results indicate that performance is not monotonic with this value; the optimal AUROC is achieved atγ=0\.005\\gamma=0\.005, while lower or higher values lead to performance degradation, highlighting the importance of careful hyperparameter tuning\. Figure 9:Ablation study on train samples andγ\\gammameasured by AUROC\.Table 5:Verbal uncertainty markers and their mean probabilities for GPT\-4o on the SciQ dataset, sorted by probability\.PhraseProbabilityabsolutely certain1\.000i’m sure1\.000pretty sure1\.000quite certain1\.000confident0\.839positive0\.839sure0\.839i’m pretty sure0\.742i think0\.710almost certain0\.677i think it’s safe to say0\.677i’m confident0\.645think0\.645because0\.355can0\.355closely tied0\.355pretty clear0\.355quite similar0\.355reasonable to assume0\.355typically0\.355very likely0\.355likely0\.000might have0\.000Table 6:Verbal uncertainty markers and their mean probabilities for GPT\-3\.5\-Turbo on the MedQA dataset, sorted by probability\.PhraseProbabilitybest course of action1\.000choice1\.000given1\.000highly probable1\.000increased risk1\.000most likely1\.000pretty sure1\.000quite confident1\.000suggestive1\.000based on0\.999may be0\.999may be needed0\.999most appropriate0\.999not definite0\.999would expect0\.999indication0\.998most common0\.998would most strongly0\.998would suspect0\.998likely0\.997one would expect0\.997should be0\.997could be0\.995i think0\.908would say0\.905seems0\.739recommend0\.506consider0\.504seems like0\.494i would say0\.034would be0\.013highly likely0\.008sounds like0\.004not completely confident0\.003most concerning0\.002indicating0\.001likelihood0\.001suspect0\.001important0\.000understandable0\.000Table 7:Verbal uncertainty markers and their mean probabilities for GPT\-3\.5\-Turbo on the SciQ dataset, sorted by probability\.PhraseProbabilityi know for a fact1\.000pretty sure1\.000but i think0\.960inclined to say0\.960not certain0\.960quite certain0\.920almost certain0\.880definite0\.880definitely0\.880i know for a fact that it’s…0\.880like0\.880primarily0\.880not completely confident0\.760can0\.720i think0\.720over time0\.720quite sure0\.560typically0\.000 ## Appendix HThe Use of Large Language Models \(LLMs\) For improved clarity and readability, we used OpenAI GPT\-4o strictly as an editing aid\. Its function was limited to correcting grammar, refining style, and polishing language, much like conventional grammar\-checking tools or dictionaries\. The model was not involved in generating scientific content or ideas, and its use remains in line with common standards for manuscript preparation\.
相似文章
自我评价之言:大语言模型在机器翻译中的口头化置信度研究
本文研究了从大语言模型中提取机器翻译输出置信度的口头化方法,并将其与内部token概率进行了比较。研究发现,尽管两种方法在错误检测和校准方面表现相似,但内部置信度与口头化置信度之间几乎没有相关性。
一致但校准不佳:评估LLM在自然语言风险沟通中的局限性
本文评估了九个LLM在自然语言中准确传达概率预测的能力,发现模型表现一致但校准不佳,尤其是在不确定性任务上。
“不太可能”有多不可能?评估大型语言模型对言语概率的感知
本文对大型语言模型如何解释言语概率表达式进行了系统性的跨模型评估,发现它们能够忠实追踪人类基准,但存在偏差,尤其是在负面表达方面,这对人类与AI之间的不确定性交流有影响。
LLM置信度估计的不同方法基准测试
本文对LLM置信度估计的各种黑盒与白盒方法进行了基准测试,包括口头化置信度、语言不确定性、推理长度、P(Answer)、P(True)和自我一致性,比较它们在主动学习和安全分类等任务中的有效性。
观点:大型语言模型中的不确定性量化仅是无监督聚类
这篇观点论文认为,当前大型语言模型的不确定性量化方法本质上属于无监督聚类,测量的是内部一致性而非外部正确性,因此无法检测出自信的幻觉。作者主张进行范式转变,将不确定性建立在客观真理之上。