Is This Your Final Answer? Cross-Contextual Consistency as a Measure of LLM Credibility
Summary
This paper introduces Cross-Contextual Consistency (C3), a behavioral property for measuring LLM credibility by checking whether answers remain stable under topic-aligned, content-neutral perturbations. Across 26 models and six benchmarks, they find that higher consistency correlates with correctness, offering a complementary evaluation axis.
View Cached Full Text
Cached at: 08/12/26, 08:34 AM
# Is This Your Final Answer? Cross-Contextual Consistency as a Measure of LLM Credibility
Source: [https://arxiv.org/html/2608.10315](https://arxiv.org/html/2608.10315)
Siyang Wu Data Science Institute University of Chicago Chicago, USA siyangwu@uchicago\.edu&Yibo Jiang Department of Computer Science University of Chicago Chicago, USA yiboj@uchicago\.edu&Bryon Aragam Booth School of Business University of Chicago Chicago, USA bryon@chicagobooth\.edu
###### Abstract
Large language models \(LLMs\) are powerful black\-box systems, making it difficult to discern whether their answers reflect stable internal beliefs or superficial pattern matching\. We identifycross\-contextual consistencyas an underutilized behavioral property of LLMs: a credible answer should remain stable when the same task is placed under topic\-aligned, content\-neutral contextual variation\. Building on this intuition, we operationalize Cross\-Contextual Consistency \(C3\) by comparing model generations under original and perturbed prompts\. Across 26 models and six benchmarks spanning reasoning, factuality, and code generation, we find that answers with smaller cross\-contextual shifts are more likely to be correct or factual\. We demonstrate that C3 provides a complementary axis of evaluation and can serve as a benchmark usefulness diagnostic, identifying which portions of a benchmark remain informative even when aggregated scores are widely considered “saturate”\.
## 1Introduction
Large language models \(LLMs\) achieve strong performance across a wide range of tasks, yet the internal basis of their responses remains poorly understood\. In contrast to this strong performance, LLMs often fail to preserve logical consistency: For example, the*reversal curse*\(Berglundet al\.,[2024](https://arxiv.org/html/2608.10315#bib.bib85)\)illustrates the failure to correctly associate facts that logically entail one another, while*context hijacking*\(Jianget al\.,[2024b](https://arxiv.org/html/2608.10315#bib.bib86)\)shows that appending semantically neutral context can trick LLMs into changing their answers\. These examples show that despite their incredible performance on a variety of challenging tasks, LLMs can still fail at basic logical reasoning\. What’s more, these failures are often unpredictable and manifest in surprising and unusual behaviours\. This motivates the need for principled approaches to directly assess model credibility and its relationship to performance and accuracy on downstream tasks\.
One popular approach is to exploit certain internal signals of large language models for evaluation\. Token\-level probabilities, for example, are commonly used in multiple\-choice benchmarks to assess calibration\(Kapooret al\.,[2024](https://arxiv.org/html/2608.10315#bib.bib73); Renet al\.,[2023](https://arxiv.org/html/2608.10315#bib.bib74)\)\. As powerful as these approaches can be in principle, these methods cannot be applied on closed\-source models and, more critically, cannot handle free\-form questions and open\-ended responses\. Other widely used methods such as self\-report\(Linet al\.,[2022](https://arxiv.org/html/2608.10315#bib.bib57)\), self\-consistency\(Wanget al\.,[2023](https://arxiv.org/html/2608.10315#bib.bib8)\), and paraphrasing consistency\(Portillo Wightmanet al\.,[2023](https://arxiv.org/html/2608.10315#bib.bib71)\)are straightforward to implement, but each can suffer from mechanistic failures such as overconfidence\(Rathiet al\.,[2025](https://arxiv.org/html/2608.10315#bib.bib59)\), trapping on the same incorrect answer\(Chenet al\.,[2023](https://arxiv.org/html/2608.10315#bib.bib91)\), or changing of the prompt’s meaning\(Chataigneret al\.,[2025](https://arxiv.org/html/2608.10315#bib.bib61)\), leading to inflated confidence estimates\.
An alternative approach is to attempt to elicit the confidence of LLM responses*indirectly*, without probing model parameters or asking for confidence measurements directly\. This approach can be likened to placing a model under judicial scrutiny\. As in legal proceedings, credibility is established not by a single answer but throughcross\-examination: deposition, repeated questioning, and extensive fact checking\. Accordingly, a model’s response should be examined for consistency under alternative but semantically equivalent formulations of the same question\. If a model truly*affirms*its answer, it should be robust to cross\-examination under logically and semantically equivalent contexts\. However, if a model’s outputs vary substantially across different framings, the credibility of the model is undermined\.
In this paper, we study a simple behavioral principle: an answer is more credible when it remains stable across contexts that change the surrounding wording or premise but not the task\-relevant meaning\. In an LLM, we call this property Cross\-Contextual Consistency \(C3\), illustrated in Figure[1](https://arxiv.org/html/2608.10315#S1.F1)\. By sampling a large number of prompt variations that preserve the original semantic content and then measuring the resulting distributional differences in model responses, we translate this behavioral principle into a quantitative metric for evaluating the credibility of a model\. Rather than judging the correctness of individual answers, we evaluate the consistency of a model’s responses, which is a much simpler objective that relies less on external knowledge, does not require access to model parameters, and does not rely on self\-reported measures\. Surprisingly, despite not being explicitly designed to test correctness—indeed, factuality and correctness are never explicitly evaluated by this metric—C3 proves to be a useful indicator of model correctness\.
This metric, which is measured at the instance level, can be aggregated into a model\-level robustness profile across different benchmarks\. Evaluation across 16 models and six benchmarks shows that perturbation\-based probing not only provides a practical means of quantifying how much a model’s generations shift under controlled, content\-neutral prompt perturbations, but also reveals a consistent association between local stability and downstream reliability: Models \(and individual task instances\) that exhibit smaller distributional shifts under perturbation are more likely to produce correct answers, whereas larger shifts are frequently associated with errors\.*This is despite the fact that accuracy is never explicitly measured by C3\.*Building on this observation, the results show that C3 provides a complementary evaluation axis that remains informative even when conventional benchmarks become less discriminative due to saturation and contamination\.
Figure 1:Overview of the proposed workflow\. From left to right, we sample generations from the original prompt and from semantically neutral perturbed prompts, measure the resulting distributional shift, and compute C3\. The right panel shows C3 for 26 models plotted by release month, revealing an increasing trend on the MMLU High School Statistics benchmark: newer models become progressively more credible\.Our main contributions are:
1. 1\.An underutilized behavioral property of LLMs\.We identify C3 as an underutilized behavioral property of LLMs’ credibility of generations: when a model’s answer is well\-supported, it should remain stable under cross\-examination\.
2. 2\.An operationalization of cross\-contextual consistency\.We instantiate this idea through C3, a black\-box evaluation protocol that compares model output distributions under original and perturbed contexts while adapting across multiple\-choice, short\-answer, long\-form factuality, and code\-generation tasks\.
3. 3\.Empirical evidence for the evaluation utility of cross\-contextual consistency\.Across six benchmarks and 16 main models, with an additional 10\-model case study, we show that C3 aligns with correctness and factuality, remains robust across perturbation sources and comparison estimators, and provides a complementary diagnostic for identifying saturated, brittle, biased, and unlearned benchmark regions\.
## 2Motivation: A tale from logical entailment
When prompted with a question, what does it intuitively mean for an LLM to answer faithfully? One natural interpretation is that the model responds in accordance with the knowledge it possesses\. Fortunately, this intuition can be formalized using the language of propositional logic\.
Let’s write the internal parametric knowledge base of a model asΓ=\{γ1,…,γn\}\\Gamma=\\\{\\gamma\_\{1\},\\ldots,\\gamma\_\{n\}\\\}, consisting ofnnatomic binary formulae, from which certain statements produced by the model can be constructed via logical operators \(e\.g\., conjunction, disjunction\)\. For example, we can have knowledge base with\{γ1=“Paris is in France”,γ2=“France is in Europe”\}\\\{\\gamma\_\{1\}\\text\{=\`\`Paris is in France"\},\\gamma\_\{2\}\\text\{=\`\`France is in Europe"\}\\\}and statements like “Paris is in an European country” \(γ1∧γ2\\gamma\_\{1\}\\land\\gamma\_\{2\}\) or “France is a European country” \(γ2\\gamma\_\{2\}\)\.
An evaluation mapvvassigns True or False values to atomic formulae, i\.e\.v\(γi\)v\(\\gamma\_\{i\}\), which induces truth values for statements built from atomic formulae\. A valuationvv, under which a statementffevaluates to true, is said to satisfy the statement, or to be a model of the statement\. LetM\(f\)M\(f\)be the set of models/worlds offf\. Intuitively,M\(f\)M\(f\)is the subset of all worlds under which statementffis True\. Due to the stochastic nature of language models, each evaluation is associated with a probabilityP\(v\)P\(v\)reflecting its likelihood of being true or false\. In other words, rather than modeling an LLM as having a single coherent view, we model it as a mixture of views, each associated with a probability\.
Given two statementsf,gf,g, one can say thatffentailsgg\(i\.e\.,f⊧gf\\models g\) ifM\(f\)⊆M\(g\)M\(f\)\\subseteq M\(g\)\. That is, ifffentailsgg, then in every world whereffis True,ggmust be True\. For example, “Paris is in an European country” entails “France is a European country” with the knowledge base that contains “Paris is in France” and “France is in Europe” becauseγ1∧γ2⊧γ2\\gamma\_\{1\}\\land\\gamma\_\{2\}\\models\\gamma\_\{2\}\. Interestingly, the two statements do not imply one another in isolation\. The implication only holds once the internal knowledge base is taken into account\.
When we do not have perfect entailment, the entailment probability can be computed as:P\(f⊧g\)=1−∑v∈M\(f\)andv∉M\(g\)P\(v\)P\(f\\models g\)=1\-\\sum\_\{v\\in M\(f\)\\text\{ and \}v\\not\\in M\(g\)\}P\(v\)This is the complement of the probability mass assigned to worlds in whichffis true andggis false\. In the context of LLMs, we letffdenote a prompt andggdenote an response\. We interpretf⊧gf\\models gas meaning that the promptffentails the responsegg\.
When evaluating LLMs, we seek to determine whether a given prompt\-response pair is genuinely implied by the model’s internal world view, or, in stochastic terms, with what probability the model supports that implication\. Under this framework, answering this question amounts to computing the aforementioned probability\. However, two challenges remain: \(1\) How to sample different worldsvv? \(2\) How to estimate different probabilitiesP\(v\)P\(v\)?
Our approach is motivated by considering different contexts as samples from different worlds\. Each prompt corresponds to one world\. Although the LLM’s output is stochastic, in practice the most probable answer typically dominates\. This is because a single prompt may constrain induced behaviors\. Therefore, we must perturb the prompts to induce diverse evaluation maps, while ensuring that the perturbed prompts remain within the semantic scope ofM\(f\)M\(f\)\.
Still, directly estimating these underlying probabilities is ill\-defined\. Instead, we adopt a different approach\. If for every worldvvwe havef⊧gf\\models g, then the entailment probability is11\. Conversely, if there is substantial disagreement across worlds, the entailment probability is low\. We therefore measure entailment by estimating the consistency of the model’s answers across different contexts, which can be interpreted as a notion of credibility\. This is how we characterize \(stochastic\) entailment through cross\-contextual consistency\.
## 3Operationalizing Cross\-Contextual Consistency
Building on the motivation in[Section2](https://arxiv.org/html/2608.10315#S2), we operationalizeCross\-Contextual Consistency \(C3\)as a behavioral probe for LLM credibility\. The central intuition is simple: if a model’s answer is well\-supported, then its answer behavior should remain stable when the same task is placed in a different but answer\-neutral context\. Conversely, if the model’s answer changes substantially under contextual variation that does not alter the task\-relevant meaning, then the answer is more context\-fragile\. We therefore treat C3 as a cross\-contextual comparison signal rather than as a single fixed metric: the goal is to compare model behavior under the original prompt and under controlled perturbed contexts\.
Controlled contextual perturbations\.For each original task queryx∈Xx\\in X, we sample a set of contextual perturbationsℰ=\{ϵ1,…,ϵn\}\\mathcal\{E\}=\\\{\\epsilon\_\{1\},\\dots,\\epsilon\_\{n\}\\\}\. Each perturbationϵ∈ℰ\\epsilon\\in\\mathcal\{E\}is prefixed to the original query, producing a perturbed queryx′=\[ϵ;x\]x^\{\\prime\}=\[\\epsilon;x\]\. These perturbations are designed to change the surrounding context while preserving the task itself\. In particular, we require perturbations to satisfy three properties: topic alignment, content neutrality, and non\-trivial contextual variation\. Topic alignment means that the perturbation remains within the broad domain or capability being tested\. Content neutrality means that the perturbation does not reveal, support, contradict, or otherwise change the correct answer\. Non\-triviality means that the perturbation introduces meaningful contextual variation rather than merely restating the original question\. We verify these three requirements separately in the appendix\. Content neutrality is evaluated in Appendix[A\.3](https://arxiv.org/html/2608.10315#A1.SS3), where we examine whether the added context avoids introducing answer\-relevant information or systematically shifting model performance\. Topic alignment is examined in Appendix[A\.4](https://arxiv.org/html/2608.10315#A1.SS4), where we assess whether perturbations remain within the same broad domain or capability as the original query\. Non\-trivial contextual variation is verified in Appendix[A\.2](https://arxiv.org/html/2608.10315#A1.SS2), where we check that accepted perturbations introduce diverse contextual variation rather than near\-duplicate restatements\.
Some examples of perturbed prompts are:
- •SVAMP:*Many tourists visited the ancient castle during the weekend\.*Rachel learned that 317 visitors came to Buckingham Palace that day\. If there were 295 visitors the previous day, how many more visitors visited Buckingham Palace that day than on the previous day?
- •SimpleQA:*An artist might release an EP during any year of their career\.*What EP did Rosalía release in 2019?
Definition of C3\.We formalize model generation as a stochastic process\. Given an inputxx, an LLM induces a conditional distributionP\(Y∣x\)P\(Y\\mid x\)over possible outputs\. This view follows the standard autoregressive perspective, where tokens are sampled sequentially based on the accumulated context and hidden state dynamics\(Vaswaniet al\.,[2017](https://arxiv.org/html/2608.10315#bib.bib46); Holtzmanet al\.,[2020](https://arxiv.org/html/2608.10315#bib.bib47); Geshkovskiet al\.,[2025](https://arxiv.org/html/2608.10315#bib.bib48)\)\. For an original queryxx, we sample a set of outputs𝒴=\{y1,y2,…,yn\}∼P\(Y∣x\)\\mathcal\{Y\}=\\\{y\_\{1\},y\_\{2\},\\dots,y\_\{n\}\\\}\\sim P\(Y\\mid x\)\. For each contextual perturbationϵ∈ℰ\\epsilon\\in\\mathcal\{E\}, we construct a perturbed queryxϵ=\[ϵ;x\]x\_\{\\epsilon\}=\[\\epsilon;x\]\. The perturbation setℰ\\mathcal\{E\}induces a perturbed output distribution,Pℰ\(Y∣x\)=1\|ℰ\|∑ϵ∈ℰP\(Y∣xϵ\)P^\{\\mathcal\{E\}\}\(Y\\mid x\)=\\frac\{1\}\{\|\\mathcal\{E\}\|\}\\sum\_\{\\epsilon\\in\\mathcal\{E\}\}P\(Y\\mid x\_\{\\epsilon\}\)from which we sample perturbed outputs𝒴ℰ=\{y1′,y2′,…,ym′\}∼Pℰ\(Y∣x\)\\mathcal\{Y\}^\{\\mathcal\{E\}\}=\\\{y^\{\\prime\}\_\{1\},y^\{\\prime\}\_\{2\},\\dots,y^\{\\prime\}\_\{m\}\\\}\\sim P^\{\\mathcal\{E\}\}\(Y\\mid x\)\.
Cross\-Contextual Consistency \(C3\) measures how stable the model’s answer behavior remains between the original and perturbed conditions\. Formally, we define C3 as a normalized inverse distance between the original and perturbed output distributions:C3\(x;ℰ\)=1−D~\(P\(Y∣x\),Pℰ\(Y∣x\)\)\\mathrm\{C3\}\(x;\\mathcal\{E\}\)=1\-\\widetilde\{D\}\\left\(P\(Y\\mid x\),P^\{\\mathcal\{E\}\}\(Y\\mid x\)\\right\), whereD~\(⋅,⋅\)\\widetilde\{D\}\(\\cdot,\\cdot\)is a task\-adaptive distance or disagreement function normalized to\[0,1\]\[0,1\]\. A higher C3 score indicates that the model’s output distribution changes less under topic\-aligned, content\-neutral contextual variation, while a lower C3 score indicates greater context\-fragility\. This definition makes C3 a general cross\-contextual comparison framework rather than a metric tied to a single distance\.
Distance metric: MMD\.In our main implementation, we instantiateD~\\widetilde\{D\}using Maximum Mean Discrepancy \(MMD\)\(Grettonet al\.,[2012](https://arxiv.org/html/2608.10315#bib.bib49)\)\. MMD provides a flexible non\-parametric estimator for comparing empirical answer distributions, making it suitable for both fixed\-format and open\-ended generations\. Given sampled outputs from the original condition𝒴\\mathcal\{Y\}and the perturbed condition𝒴ℰ\\mathcal\{Y\}^\{\\mathcal\{E\}\}, we compute an empirical MMD distance using task\-adaptive feature maps and kernel choices\. We then normalize the resulting distance into a consistency score in\[0,1\]\[0,1\], where larger values indicate smaller cross\-contextual shift\. Details on the empirical MMD estimator, task\-adaptive feature maps, kernel choices, and normalization are provided in Appendix[H](https://arxiv.org/html/2608.10315#A8)\.
Importantly, MMD provides one instantiation of C3, not the definition of C3 itself\. C3 is defined by the comparison between model behavior under original and perturbed contexts\. In Appendix[G](https://arxiv.org/html/2608.10315#A7), we replace MMD with a simpler cross\-comparison distance and show that the same cross\-contextual signal is largely preserved, supporting the view that C3 is driven by the original\-versus\-perturbed comparison rather than by a particular choice of distance\.
## 4Experiments
In this section, we describe our experimental setup, including data, methods, LLM models, and benchmarks\.
Data\.We evaluate C3 on six widely used benchmarks spanning arithmetic reasoning, multiple\-choice reasoning, commonsense inference, short\-form factual QA, long\-form factuality, and code generation\. Specifically, we use SVAMP\(Patelet al\.,[2021](https://arxiv.org/html/2608.10315#bib.bib55)\), MMLU High School Statistics\(Hendryckset al\.,[2021b](https://arxiv.org/html/2608.10315#bib.bib53),[a](https://arxiv.org/html/2608.10315#bib.bib54)\), CommonsenseQA\(Talmoret al\.,[2019](https://arxiv.org/html/2608.10315#bib.bib56)\), SimpleQA Verified\(Haaset al\.,[2025](https://arxiv.org/html/2608.10315#bib.bib50)\), FActScore\(Minet al\.,[2023](https://arxiv.org/html/2608.10315#bib.bib51)\), and HumanEval\(Chenet al\.,[2021](https://arxiv.org/html/2608.10315#bib.bib52)\)\. Together, these benchmarks cover both fixed\-format and open\-ended generations, allowing us to evaluate C3 across diverse task formats and answer types\. Detailed benchmark descriptions and prompts are provided in Appendix[I](https://arxiv.org/html/2608.10315#A9)\.
Baselines\.Since C3 is intended to serve as a proxy for the credibility of model generations, we compare it with related notions of confidence, consistency, and factuality\. We include both vanilla black\-box confidence estimators and a non\-vanilla factuality checking tool\. The vanilla baselines includeself\-reported confidence\(Linet al\.,[2022](https://arxiv.org/html/2608.10315#bib.bib57)\), where the model outputs an explicit confidence score together with its answer;self\-consistency\(Wanget al\.,[2023](https://arxiv.org/html/2608.10315#bib.bib8)\), which estimates confidence from agreement among repeated stochastic generations under the same prompt; andparaphrasing consistency\(Portillo Wightmanet al\.,[2023](https://arxiv.org/html/2608.10315#bib.bib71)\), which measures whether answers remain stable under meaning\-preserving prompt paraphrases\. As a non\-vanilla baseline, we compare against FActScore\(Minet al\.,[2023](https://arxiv.org/html/2608.10315#bib.bib51)\), which evaluates factual support for long\-form generations using external evidence\. For sampling\-based baselines, agreement is computed with a generalized consistency score, using exact\-match indicators for fixed\-format tasks and semantic similarity for open\-ended generations\. Full implementation details are provided in Appendix[J](https://arxiv.org/html/2608.10315#A10)\.
Models\.We evaluate C3 on 16 widely used LLMs spanning multiple families and scales\. We also conduct case studies on 10 additional models; however, due to their characteristics, such as heavier reasoning processes or deprecated designs, they are often costly in time and computational resources\. We therefore restrict these case studies to the MMLU High School Statistics benchmark\. The full list of models is provided in Appendix[E\.1](https://arxiv.org/html/2608.10315#A5.SS1)\.
Evaluation metrics\.We evaluate C3 and all baseline scores after normalizing each score to the range\[0,1\]\[0,1\]\. To measure calibration, we report theExpected Calibration Error \(ECE\), where lower values indicate better calibration\. For ranking\-based evaluation, we reportAUROC, as well as the area under the precision–recall curve for detecting correct outputs \(AUPRC\-P\) and detecting incorrect outputs \(AUPRC\-N\), with the latter computed using1−si1\-s\_\{i\}\. Details and implementation of these metrics are provided in Appendix[E\.2](https://arxiv.org/html/2608.10315#A5.SS2)\.
CommonsenseQASVAMPMethodECE↓\\downarrowAUROC↑\\uparrowAUPRC\-P↑\\uparrowAUPRC\-N↑\\uparrowECE↓\\downarrowAUROC↑\\uparrowAUPRC\-P↑\\uparrowAUPRC\-N↑\\uparrowSelf\-consistency0\.2240\.6430\.7560\.4610\.1420\.8610\.8930\.740Paraphrasing0\.2710\.6340\.7640\.4840\.1570\.8410\.8730\.720Self\-report0\.1910\.5420\.7460\.3320\.2400\.5320\.7630\.277C3 \(Ours\)\\cellcolorgray\!200\.189\\cellcolorgray\!20 0\.722\\cellcolorgray\!200\.805\\cellcolorgray\!200\.541\\cellcolorgray\!200\.072\\cellcolorgray\!200\.917\\cellcolorgray\!200\.936\\cellcolorgray\!200\.818MMLU High School StatisticsFActScoreMethodECE↓\\downarrowAUROC↑\\uparrowAUPRC\-P↑\\uparrowAUPRC\-N↑\\uparrowECE↓\\downarrowAUROC↑\\uparrowAUPRC\-P↑\\uparrowAUPRC\-N↑\\uparrowSelf\-consistency0\.3820\.5770\.5110\.6050\.2250\.8580\.909\\cellcolorgray\!200\.731Paraphrasing0\.4060\.5470\.4890\.5910\.2490\.8280\.8870\.716Self\-report0\.3780\.496\\cellcolorgray\!200\.5940\.4110\.3300\.4960\.7060\.335C3 \(Ours\)\\cellcolorgray\!200\.338\\cellcolorgray\!200\.5970\.530\\cellcolorgray\!200\.621\\cellcolorgray\!20 0\.160\\cellcolorgray\!200\.873\\cellcolorgray\!200\.9160\.701HumanEvalSimpleQAMethodECE↓\\downarrowAUROC↑\\uparrowAUPRC\-P↑\\uparrowAUPRC\-N↑\\uparrowECE↓\\downarrowAUROC↑\\uparrowAUPRC\-P↑\\uparrowAUPRC\-N↑\\uparrowSelf\-consistency0\.1310\.8120\.8960\.4870\.3930\.7920\.3680\.924Paraphrasing0\.1650\.7780\.8620\.4530\.4110\.7730\.3490\.906Self\-report0\.2870\.5280\.8070\.2110\.7780\.4900\.1510\.846C3 \(Ours\)\\cellcolorgray\!200\.110\\cellcolorgray\!200\.847\\cellcolorgray\!200\.943\\cellcolorgray\!200\.520\\cellcolorgray\!200\.166\\cellcolorgray\!200\.823\\cellcolorgray\!200\.389\\cellcolorgray\!200\.944Table 1:Comparison of C3 with baselines across six benchmarks average over 16 models we tested\. Shaded cells denote the best performance for each metric among the four methods\. CommonsenseQA evaluates commonsense world knowledge; SVAMP evaluates mathematical reasoning; MMLU High School Statistics evaluates statistical knowledge; FactScore and SimpleQA evaluate factuality in long\-form and short\-form generation, respectively; and HumanEval evaluates code generation\. Lower is better for ECE, while higher is better for AUROC, AUPRC\-P, and AUPRC\-N\. The noise for perturbation in the table is sampled by GPT\-4\.1, and we also shown that C3 does not rely on advanced model and the results can be still reproducible by smaller models with very few costs as we shown in Appendix[F](https://arxiv.org/html/2608.10315#A6)\.C3 elicitation\.For both perturbed and unperturbed sampling, we collected 30 trials per instance across 16 standard models to assess the alignment of C3 against other baselines\. The number 30 is supported by an empirical study shown in Appendix[B](https://arxiv.org/html/2608.10315#A2)\. In this study we use GPT\-4\.1 for noise sampling, however, an ablation study in Appendix[F](https://arxiv.org/html/2608.10315#A6)shows that perturbations from much smaller models are nearly as effective as those from larger ones, but with much lower computational overhead\. The noise was prefixed to each prompt, separated by a single space\. To determine answer equivalence, we use an LLM through the prompt detailed in Appendix[D\.5](https://arxiv.org/html/2608.10315#A4.SS5)\. Given the high volume of evaluations required \(tens of millions given the scale of our experiments\), we performed offline inference using a Qwen3\-8B\(Team,[2025](https://arxiv.org/html/2608.10315#bib.bib76)\)model temperature set to 0, supported by the VLLM framework\(Kwonet al\.,[2023](https://arxiv.org/html/2608.10315#bib.bib75)\)on two NVIDIA H200 GPUs\.
## 5Results
We find that C3 provides substantially better signals than the baseline methods, and aligns especially well on challenging benchmarks such as SimpleQA, where other approaches often produce inflated assessments\.
### 5\.1C3 is calibrated with truthfulness
#### Correctness\.
The results in Table[1](https://arxiv.org/html/2608.10315#S4.T1)demonstrate that C3 consistently outperforms existing baselines in aligning model “credibility” with truthfulness\. In math tasks like SVAMP, C3 achieves an AUROC of 0\.917, a noticeable jump over Self\-Consistency, with AUROC of 0\.861\. This suggests that measuring the credibility of LLMs with perturbations is more effective for capturing the logical coherence of a reasoning chain than simple sampling strategies using identical prompts\.
Factuality\.Furthermore, the consistent gains on FActScore, CommonsenseQA, and SimpleQA suggest that C3 is effective for detecting factuality\-related errors\. In these settings, traditional consistency metrics can be poorly calibrated because models may remain highly consistent even when they are confidently relying on incorrect parametric memories or strong priors\. By measuring the distributional shift in generations induced by content\-neutral noise, C3 increases contrast between well\-known facts, small shift, and poorly known facts, large shift, yielding a more informative signal of the model’s knowledge state\.
Figure 2:Scaling trends of Cross\-Contextual Consistency \(C3\) calibration across benchmarks\. Colors distinguish different benchmarks, while line styles represent metric types: solid lines denote AUROC \(higher is better\) and dashed lines denote ECE \(lower is better\)\. The results show that as model scale increases, the C3 becomes significantly better calibrated to correctness, evidenced by rising AUROC and declining ECE\.Long form generation\.Across the six benchmarks tested, HumanEval and FActScore are considered as “long\-form generation” because they do not assume a pre\-defined output format \(such as multiple choice or keywords\)\. C3 achieves nearly the best scores across four metrics in these two forms of generation: 0\.110 ECE and 0\.847 AUROC on HumanEval, and 0\.160 ECE and 0\.873 AUROC on FActScore\. This is a scenario where many existing works fail\. Our method scales cleanly to long\-form generation, keeping nearly as simple as fixed\-format tasks, which even allows us to evaluate code generation\. While recent efforts likeSharma and David \([2025](https://arxiv.org/html/2608.10315#bib.bib58)\)assess uncertainty in code via symbolic execution, our method achieves superior calibration without requiring external execution environments\.
Why they fail?Self\-reported measures exhibit miscalibration: While they occasionally achieve seemingly reasonable ECE \(e\.g\., 0\.191 in CommonsenseQA\), their discriminative power, measured by AUROC and AUPRC, remains near\-random\. This confirms that LLMs struggle to introspectively verbalize uncertainty, often overconfidently yielding near\-100% score for wrong answers\(Mielkeet al\.,[2022](https://arxiv.org/html/2608.10315#bib.bib60); Rathiet al\.,[2025](https://arxiv.org/html/2608.10315#bib.bib59)\)\. We also find that paraphrasing suffers from fundamental mechanistic flaws; since the paraphrasing is often performed by an LLM assistant, it can introduce semantic drift\. For instance, high word overlap can mask cases where swapping arguments changes the underlying meaning\(Zhanget al\.,[2019](https://arxiv.org/html/2608.10315#bib.bib62)\), or the assistant may unintentionally alter the core intent of the prompt\(Chataigneret al\.,[2025](https://arxiv.org/html/2608.10315#bib.bib61)\)\. Methods like self\-consistency paraphrasing do not fail entirely aligning with truthfulness but are often trapped in the “confidently wrong” loop of LLM generation when the model sticks with an incorrect answer\. As reflected in SimpleQA, which is a hard benchmark for many LLMs, self\-consistency fails and is mis\-calibrated with actual correctness on ECE, and paraphrasing makes results even worse\. Measuring the distributional shift with C3, we successfully nudge the model to sample different answers when the knowledge is not grounded in LLMs’ knowledge base, providing a much better signal in ECE and AUROC \(0\.166 and 0\.823, respectively\) in SimpleQA\.
C3 becomes more informative as model capability increases\. As shown in Figure[2](https://arxiv.org/html/2608.10315#S5.F2), larger models within the Llama and Mistral families generally show stronger alignment between C3 and downstream correctness\. This suggests that cross\-contextual consistency is most useful once a model has enough capability for stable answer behavior to emerge; at that point, residual instability more clearly marks fragile answers rather than broad capability failure\. However, scale does not eliminate cross\-contextual fragility: even the strongest models remain imperfectly stable under answer\-neutral contextual variation\.
### 5\.2A closer inspection on facutality
CommonsenseQA, FActScore, and SimpleQA all test a model’s world knowledge and factuality\. C3 shows a close alignment with the quality of generation across all three\. While FActScore probes externally, scoring by decomposing long\-form generation into atomic facts and verifying them against external sources, C3 probes internally\. It uses variations of noise to perturb the LLM to see if it “insists” on an answer; the resulting distributional shift provides a significant signal regarding the quality of generation\. As shown in Figure[3](https://arxiv.org/html/2608.10315#S5.F3), when aggregating instance\-level C3 and correctness, we observe a 0\.624 Spearman rank correlation on FActScore and an 0\.831 AUROC\. We also include a comparison with SimpleQA, a benchmark where most models fail, making it a strong indicator for detecting overconfident metrics; C3 shows a 0\.642 Spearman correlation and an 0\.821 AUROC\. Even without probing external information, C3 successfully aligns with FActScore\. This suggests that the model’s internal state contains a latent representation of its own knowledge boundaries: when a model “knows” a fact, its output distribution is resilient to input noise, whereas hallucinated facts reside in low\-probability regions that collapse or shift significantly under even minor perturbations\.
Figure 3:A detailed comparison of C3 alignment with factuality on FActScore and SimpleQA, showing Spearman rank correlation \(left\) and AUROC \(right\)\. C3 exhibits moderate to strong rank correlation with generation factuality\. Minor AUROC differences from Table 1 are due to different aggregation strategies\.
### 5\.3C3 for benchmark diagnosis
By jointly inspecting instance\-level C3 and instance\-level benchmark performance, we can show cross\-model typical behavior for each benchmark instance after averaging across the same 16 models\. We find that C3 captures complementary information about model behavior beyond aggregate performance alone, as shown in Figure[4](https://arxiv.org/html/2608.10315#S5.F4)\.
Why do we need C3 as an additional axis?Relying exclusively on performance scores often mask the underlying mechanism of a model’s knowledge\. A key phenomenon in current LLM evaluation is that benchmarks may suffer from data contamination or overfitting, where models memorize specific prompt\-answer pairs without acquiring the underlying reasoning\. C3 provide additional information: “Brittle” \(Yellow\) region where models achieve high accuracy but fail to maintain consistency under perturbation\. While high performance typically suggests capability, the low C3 in this region supports the alternative hypothesis: Success on these instances is driven by surface\-level pattern matching rather than robust semantic understanding\. This divergence serves as a proxy for detecting potential benchmark leakage to training processes of current LLMs\.
C3 also helps differentiating mastered from systematic bias\. The C3 axis further clarifies the status of the benchmark by distinguishing between “solved” and “biased” generation\. Instances in the “Mastered” \(Green\) region represent tasks where models have converged on a stable solution\. A high density of instances in this region supports the hypothesis of benchmark saturation, indicating that these specific questions no longer possess the discriminative power to distinguish between the capabilities of different models\. Conversely, the “Biased” region highlights instances where models are not merely guessing, but are consistently trapped on incorrect answers\. This supports the hypothesis that these benchmark instances trigger strong, incorrect priors or common misconceptions shared across models that possibly arise from model training processes\.
Benchmark by Benchmark Comparison\.Figure[4](https://arxiv.org/html/2608.10315#S5.F4)visualizes instance\-level C3 against performance, revealing distinct patterns that characterize the status of each benchmark\. CommonsenseQA and HumanEval both exhibit the signature of saturation, where a significant proportion of instances in the “Mastered” region suggests these tasks are well solved by modern models\. However, in contrast to CommonsenseQA, HumanEval displays a heavy “tail” extending into the “Brittle” quadrant; this pattern implies that its high performance may be partially inflated by overfitting, where models succeed via surface pattern matching but fail under perturbation, lacking the true reasoning process of coding\(Riddellet al\.,[2024](https://arxiv.org/html/2608.10315#bib.bib72)\)\. SVAMP and FActScore show a balanced proportion of unlearned and mastered instances, while MMLU High School Stats shows a high number of mastered but also a significant cluster of biased instances; this suggests that while basics are understood, specific statistical concepts trigger consistent, systematic misconceptions\. Finally, SimpleQA represents the true “hard” benchmark, dominated by the “Unlearned” region\. The scarcity of mastered instances and the prevalence of stable poor performance indicate that this benchmark is not saturated, but rather captures specific knowledge gaps that remain out of reach for the current generation of models\.
Figure 4:For each benchmark instance, we compute the mean performance and mean C3 across the 16 models, then partition instances into four regions using the median performance and median C3 computed over instances across the six benchmarks\. Green \(“mastered”\) indicates high performance with contextually consistent answers; yellow \(“brittle”\) indicates high performance but sensitive to perturbations; red \(“biased”\) indicates low performance yet perturbation\-invariant \(consistently wrong\) answers; gray indicates low performance with inconsistency, suggesting knowledge that is not reliably learned\.
## 6Related work
A variety of methods to estimate model confidence have been proposed in the literature\.White\-box\.Methods rely on logits, hidden states, or parameters, which are infeasible for closed models; moreover, OpenAI reported degraded calibration after post\-training, underscoring instability in probability\-based confidence\(OpenAIet al\.,[2024b](https://arxiv.org/html/2608.10315#bib.bib37); Xieet al\.,[2024](https://arxiv.org/html/2608.10315#bib.bib16); Shenet al\.,[2024](https://arxiv.org/html/2608.10315#bib.bib29)\)\.Self\-verbalized \(black\-box\)\.Such confidence can be miscalibrated and drift with prompting and post\-training, and often shows over\-confidence, even though some elicitation schemes help in specific settings\(Linet al\.,[2022](https://arxiv.org/html/2608.10315#bib.bib57); Xionget al\.,[2024](https://arxiv.org/html/2608.10315#bib.bib1); Kumaret al\.,[2024](https://arxiv.org/html/2608.10315#bib.bib2); Zhanget al\.,[2024b](https://arxiv.org/html/2608.10315#bib.bib3); Heoet al\.,[2025](https://arxiv.org/html/2608.10315#bib.bib4); OpenAIet al\.,[2024b](https://arxiv.org/html/2608.10315#bib.bib37)\)\.Agreement\-based \(black\-box\)\.Operating without input perturbation, plurality voting can conflate repetition bias with genuine certainty and underuse informative minority signals\(Huanget al\.,[2024](https://arxiv.org/html/2608.10315#bib.bib6); Wanget al\.,[2023](https://arxiv.org/html/2608.10315#bib.bib8)\)\.Semantic paraphrase \(black\-box\)\.Paraphrase\-based uncertainty presumes meaning preservation, yet small wording changes often shift semantics and behavior; entropy over paraphrases can therefore reflect semantic drift rather than true confidence\(Zhanget al\.,[2019](https://arxiv.org/html/2608.10315#bib.bib62); Wahleet al\.,[2024](https://arxiv.org/html/2608.10315#bib.bib38); Melamedet al\.,[2024](https://arxiv.org/html/2608.10315#bib.bib39); Mizrahiet al\.,[2024](https://arxiv.org/html/2608.10315#bib.bib13)\)\. Meanwhile paraphrasing with LLMs leads to biased collection of model performance,\(Lunardiet al\.,[2025](https://arxiv.org/html/2608.10315#bib.bib44)\)shows that the absolute accuracy scores drop significantly when when paraphrased the quesiton\.Answer\-calibration for long\-form\.Treating confidence as probability of the correct answer” is ill\-posed for multi\-sentence, open\-ended generations; automatic checks correlate unevenly with human judgment and preferences\(Xuet al\.,[2023](https://arxiv.org/html/2608.10315#bib.bib40); Fabbriet al\.,[2021](https://arxiv.org/html/2608.10315#bib.bib41); Chenet al\.,[2024](https://arxiv.org/html/2608.10315#bib.bib42)\)\.Other families \(conformal, refusal\)\.These typically require labels, specialized scoring access, or finetuning—constraints that hinder post\-hoc use in closed APIs\(Quachet al\.,[2024](https://arxiv.org/html/2608.10315#bib.bib17); Mohri and Hashimoto,[2024](https://arxiv.org/html/2608.10315#bib.bib34); Zhanget al\.,[2024a](https://arxiv.org/html/2608.10315#bib.bib43)\)\.
#### Our approach
We measure credibility via the distributional shift in generation induced by semantically equivalent, meaning\-preserving prompt perturbations, using this perturbation gap as a reference\-free proxy for confidence\(Kuhnet al\.,[2023](https://arxiv.org/html/2608.10315#bib.bib10); Mizrahiet al\.,[2024](https://arxiv.org/html/2608.10315#bib.bib13)\)\. This lens differs from existing prompt\-sensitivity and robustness notions in several important ways\. Some prior work quantifies sensitivity in probability space by tracking changes in log\-likelihoods or likelihood ratios across prompt variants, which typically requires access to internal scoring signals and reflects probability drift rather than semantic drift of the generated content\(Chatterjeeet al\.,[2024](https://arxiv.org/html/2608.10315#bib.bib88)\)\. Others study robustness through format\-induced variability by measuring sensitivity to spurious formatting features in prompt design, diagnosing evaluation volatility but not yielding an instance\-level credibility signal for open\-ended generations\(Sclaret al\.,[2024](https://arxiv.org/html/2608.10315#bib.bib90)\)\. Related measures are also often defined for classification by analyzing instability of predicted label distributions under rephrasing, which does not transfer cleanly to free\-form text where outputs must be compared semantically rather than as discrete labels\(Erricaet al\.,[2025](https://arxiv.org/html/2608.10315#bib.bib89)\)\. Another line evaluates the reliability of knowledge probing methods by testing whether accept/reject decisions remain consistent under perturbations, emphasizing probe stability rather than the stability of the generated answer itself\(Zhaoet al\.,[2025](https://arxiv.org/html/2608.10315#bib.bib87)\)\. In contrast, C3 directly measures invariance to meaning\-preserving perturbations, is black\-box and format\-adaptive, and provides a comparable credibility diagnostic across heterogeneous tasks and open\-ended generation formats\.
## 7Conclusion
This work identifies cross\-contextual consistency as a behavioral signal for evaluating LLM credibility: when a model’s answer is well\-supported, it should remain stable under topic\-aligned, content\-neutral contextual variation\. We operationalize this idea through C3, a black\-box framework that compares generation distributions under original and perturbed contexts\. Across 26 models and six benchmarks, C3 consistently aligns with correctness and factuality across reasoning, factual recall, long\-form generation, and code generation\. Beyond instance\-level evaluation, C3 also provides a complementary diagnostic for benchmark analysis, helping distinguish stable understanding from brittle or systematically biased performance\. These results suggest that controlled contextual perturbation offers a practical way to reveal answer fragility that standard accuracy\-based evaluation can miss\.
## 8Limitations
This work studies C3 across a broad but necessarily finite set of models, benchmarks, and experimental settings\. Future work may extend the evaluation to additional tasks, languages, and application domains\. Further exploration of alternative implementation choices may also help better understand how C3 can be adapted across different evaluation scenarios\.
## 9Societal Impact
This work aims to improve the evaluation of LLM reliability by providing a black\-box signal for identifying context\-fragile generations\. Such tools may help practitioners better detect uncertain, brittle, or potentially hallucinated outputs before deploying LLMs in higher\-stakes settings\. At the same time, C3 should not be treated as a guarantee of truthfulness or safety; a model can be consistent and still wrong\. In accordance with the NeurIPS Code of Ethics, we note that this method is intended as a diagnostic aid rather than a replacement for human oversight, domain expertise, or task\-specific safety evaluation\.
## References
- The reversal curse: LLMs trained on “a is b” fail to learn “b is a”\.External Links:[Link](https://openreview.net/forum?id=GPKTIktA0k)Cited by:[§1](https://arxiv.org/html/2608.10315#S1.p1.1)\.
- T\. B\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell, S\. Agarwal, A\. Herbert\-Voss, G\. Krueger, T\. Henighan, R\. Child, A\. Ramesh, D\. M\. Ziegler, J\. Wu, C\. Winter, C\. Hesse, M\. Chen, E\. Sigler, M\. Litwin, S\. Gray, B\. Chess, J\. Clark, C\. Berner, S\. McCandlish, A\. Radford, I\. Sutskever, and D\. Amodei \(2020\)Language models are few\-shot learners\.External Links:2005\.14165,[Link](https://arxiv.org/abs/2005.14165)Cited by:[§E\.1](https://arxiv.org/html/2608.10315#A5.SS1.p2.1)\.
- C\. Chataigner, R\. Ma, P\. Ganesh, Y\. Chen, A\. Taïk, E\. Creager, and G\. Farnadi \(2025\)Say it another way: auditing llms with a user\-grounded automated paraphrasing framework\.External Links:2505\.03563,[Link](https://arxiv.org/abs/2505.03563)Cited by:[§1](https://arxiv.org/html/2608.10315#S1.p2.1),[§5\.1](https://arxiv.org/html/2608.10315#S5.SS1.SSS0.Px1.p4.1)\.
- A\. Chatterjee, H\. S\. V\. N\. S\. K\. Renduchintala, S\. Bhatia, and T\. Chakraborty \(2024\)POSIX: a prompt sensitivity index for large language models\.Miami, Florida, USA,pp\. 14550–14565\.External Links:[Link](https://aclanthology.org/2024.findings-emnlp.852/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.852)Cited by:[§6](https://arxiv.org/html/2608.10315#S6.SS0.SSS0.Px1.p1.1)\.
- G\. H\. Chen, S\. Chen, Z\. Liu, F\. Jiang, and B\. Wang \(2024\)Humans or LLMs as the judge? a study on judgement bias\.Miami, Florida, USA,pp\. 8301–8327\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.474/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.474)Cited by:[§6](https://arxiv.org/html/2608.10315#S6.p1.1)\.
- J\. Chen, J\. Yoon, S\. Ebrahimi, S\. Arik, T\. Pfister, and S\. Jha \(2023\)Adaptation with self\-evaluation to improve selective prediction in LLMs\.Singapore,pp\. 5190–5213\.External Links:[Link](https://aclanthology.org/2023.findings-emnlp.345/),[Document](https://dx.doi.org/10.18653/v1/2023.findings-emnlp.345)Cited by:[§1](https://arxiv.org/html/2608.10315#S1.p2.1)\.
- M\. Chen, J\. Tworek, H\. Jun, Q\. Yuan, H\. P\. de Oliveira Pinto, J\. Kaplan, H\. Edwards, Y\. Burda, N\. Joseph, G\. Brockman, A\. Ray, R\. Puri, G\. Krueger, M\. Petrov, H\. Khlaaf, G\. Sastry, P\. Mishkin, B\. Chan, S\. Gray, N\. Ryder, M\. Pavlov, A\. Power, L\. Kaiser, M\. Bavarian, C\. Winter, P\. Tillet, F\. P\. Such, D\. Cummings, M\. Plappert, F\. Chantzis, E\. Barnes, A\. Herbert\-Voss, W\. H\. Guss, A\. Nichol, A\. Paino, N\. Tezak, J\. Tang, I\. Babuschkin, S\. Balaji, S\. Jain, W\. Saunders, C\. Hesse, A\. N\. Carr, J\. Leike, J\. Achiam, V\. Misra, E\. Morikawa, A\. Radford, M\. Knight, M\. Brundage, M\. Murati, K\. Mayer, P\. Welinder, B\. McGrew, D\. Amodei, S\. McCandlish, I\. Sutskever, and W\. Zaremba \(2021\)Evaluating large language models trained on code\.CoRRabs/2107\.03374\.External Links:[Link](https://arxiv.org/abs/2107.03374),2107\.03374Cited by:[Appendix I](https://arxiv.org/html/2608.10315#A9.p2.1),[§4](https://arxiv.org/html/2608.10315#S4.p2.1)\.
- G\. Comanici, E\. Bieber, M\. Schaekermann, I\. Pasupat, N\. Sachdeva, I\. Dhillon, M\. Blistein,et al\.\(2025\)Gemini 2\.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities\.External Links:2507\.06261,[Link](https://arxiv.org/abs/2507.06261)Cited by:[§E\.1](https://arxiv.org/html/2608.10315#A5.SS1.p1.1)\.
- T\. Computer \(2023\)RedPajama: an open source recipe to reproduce llama training datasetExternal Links:[Link](https://github.com/togethercomputer/RedPajama-Data)Cited by:[Appendix F](https://arxiv.org/html/2608.10315#A6.p1.1),[Appendix F](https://arxiv.org/html/2608.10315#A6.p4.1)\.
- F\. Errica, D\. Sanvito, G\. Siracusano, and R\. Bifulco \(2025\)What did I do wrong? quantifying LLMs’ sensitivity and consistency to prompt engineering\.Albuquerque, New Mexico,pp\. 1543–1558\.External Links:[Link](https://aclanthology.org/2025.naacl-long.73/),[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.73),ISBN 979\-8\-89176\-189\-6Cited by:[§6](https://arxiv.org/html/2608.10315#S6.SS0.SSS0.Px1.p1.1)\.
- A\. R\. Fabbri, W\. Kryściński, B\. McCann, C\. Xiong, R\. Socher, and D\. Radev \(2021\)SummEval: re\-evaluating summarization evaluation\.Transactions of the Association for Computational Linguistics9,pp\. 391–409\.External Links:[Link](https://aclanthology.org/2021.tacl-1.24/),[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00373)Cited by:[§6](https://arxiv.org/html/2608.10315#S6.p1.1)\.
- B\. Geshkovski, C\. Letrouit, Y\. Polyanskiy, and P\. Rigollet \(2025\)A mathematical perspective on transformers\.External Links:2312\.10794,[Link](https://arxiv.org/abs/2312.10794)Cited by:[§3](https://arxiv.org/html/2608.10315#S3.p4.9)\.
- A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian,et al\.\(2024\)The llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[§E\.1](https://arxiv.org/html/2608.10315#A5.SS1.p1.1)\.
- A\. Gretton, K\. M\. Borgwardt, M\. J\. Rasch, B\. Schölkopf, and A\. Smola \(2012\)A kernel two\-sample test\.Journal of Machine Learning Research13\(25\),pp\. 723–773\.External Links:[Link](http://jmlr.org/papers/v13/gretton12a.html)Cited by:[Appendix H](https://arxiv.org/html/2608.10315#A8.p1.6),[§3](https://arxiv.org/html/2608.10315#S3.p6.4)\.
- L\. Haas, G\. Yona, G\. D’Antonio, S\. Goldshtein, and D\. Das \(2025\)SimpleQA verified: a reliable factuality benchmark to measure parametric knowledge\.External Links:2509\.07968,[Link](https://arxiv.org/abs/2509.07968)Cited by:[Appendix I](https://arxiv.org/html/2608.10315#A9.p2.1),[§4](https://arxiv.org/html/2608.10315#S4.p2.1)\.
- D\. Hendrycks, C\. Burns, S\. Basart, A\. Critch, J\. Li, D\. Song, and J\. Steinhardt \(2021a\)Aligning ai with shared human values\.Proceedings of the International Conference on Learning Representations \(ICLR\)\.Cited by:[Appendix I](https://arxiv.org/html/2608.10315#A9.p1.1),[§4](https://arxiv.org/html/2608.10315#S4.p2.1)\.
- D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. Steinhardt \(2021b\)Measuring massive multitask language understanding\.Proceedings of the International Conference on Learning Representations \(ICLR\)\.Cited by:[Appendix I](https://arxiv.org/html/2608.10315#A9.p1.1),[§4](https://arxiv.org/html/2608.10315#S4.p2.1)\.
- J\. Heo, M\. Xiong, C\. Heinze\-Deml, and J\. Narain \(2025\)Do LLMs estimate uncertainty well in instruction\-following?\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=IHp3vOVQO2)Cited by:[§6](https://arxiv.org/html/2608.10315#S6.p1.1)\.
- A\. Holtzman, J\. Buys, L\. Du, M\. Forbes, and Y\. Choi \(2020\)The curious case of neural text degeneration\.External Links:[Link](https://openreview.net/forum?id=rygGQyrFvH)Cited by:[§3](https://arxiv.org/html/2608.10315#S3.p4.9)\.
- S\. Huang, Z\. Ma, J\. Du, C\. Meng, W\. Wang, and Z\. Lin \(2024\)Mirror\-consistency: harnessing inconsistency in majority voting\.InFindings of the Association for Computational Linguistics: EMNLP 2024,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 2408–2420\.External Links:[Link](https://aclanthology.org/2024.findings-emnlp.135/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.135)Cited by:[§6](https://arxiv.org/html/2608.10315#S6.p1.1)\.
- A\. Q\. Jiang, A\. Sablayrolles, A\. Roux, A\. Mensch, B\. Savary, C\. Bamford, D\. S\. Chaplot, D\. de las Casas, E\. B\. Hanna, F\. Bressand, G\. Lengyel, G\. Bour, G\. Lample, L\. R\. Lavaud, L\. Saulnier, M\. Lachaux, P\. Stock, S\. Subramanian, S\. Yang, S\. Antoniak, T\. L\. Scao, T\. Gervet, T\. Lavril, T\. Wang, T\. Lacroix, and W\. E\. Sayed \(2024a\)Mixtral of experts\.External Links:2401\.04088,[Link](https://arxiv.org/abs/2401.04088)Cited by:[§E\.1](https://arxiv.org/html/2608.10315#A5.SS1.p1.1)\.
- Y\. Jiang, G\. Rajendran, P\. Ravikumar, and B\. Aragam \(2024b\)Do llms dream of elephants \(when told not to\)? latent concept association and associative memory in transformers\.Advances in Neural Information Processing Systems37,pp\. 67712–67757\.Cited by:[§1](https://arxiv.org/html/2608.10315#S1.p1.1)\.
- S\. Kapoor, N\. Gruver, M\. Roberts, K\. M\. Collins, A\. Pal, U\. Bhatt, A\. Weller, S\. Dooley, M\. Goldblum, and A\. G\. Wilson \(2024\)Large language models must be taught to know what they don’t know\.External Links:[Link](https://openreview.net/forum?id=QzvWyggrYB)Cited by:[§1](https://arxiv.org/html/2608.10315#S1.p2.1)\.
- L\. Kuhn, Y\. Gal, and S\. Farquhar \(2023\)Semantic uncertainty: linguistic invariances for uncertainty estimation in natural language generation\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=VD-AYtP0dve)Cited by:[§6](https://arxiv.org/html/2608.10315#S6.SS0.SSS0.Px1.p1.1)\.
- A\. Kumar, R\. Morabito, S\. Umbet, J\. Kabbara, and A\. Emami \(2024\)Confidence under the hood: an investigation into the confidence\-probability alignment in large language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 315–334\.External Links:[Link](https://aclanthology.org/2024.acl-long.20/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.20)Cited by:[§6](https://arxiv.org/html/2608.10315#S6.p1.1)\.
- W\. Kwon, Z\. Li, S\. Zhuang, Y\. Sheng, L\. Zheng, C\. H\. Yu, J\. E\. Gonzalez, H\. Zhang, and I\. Stoica \(2023\)Efficient memory management for large language model serving with pagedattention\.Cited by:[§4](https://arxiv.org/html/2608.10315#S4.p6.1)\.
- S\. Lin, J\. Hilton, and O\. Evans \(2022\)Teaching models to express their uncertainty in words\.Transactions on Machine Learning Research\.Note:External Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=8s8K2UZGTZ)Cited by:[Appendix J](https://arxiv.org/html/2608.10315#A10.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2608.10315#S1.p2.1),[§4](https://arxiv.org/html/2608.10315#S4.p3.1),[§6](https://arxiv.org/html/2608.10315#S6.p1.1)\.
- R\. Lunardi, V\. D\. Mea, S\. Mizzaro, and K\. Roitero \(2025\)On robustness and reliability of benchmark\-based evaluation of llms\.External Links:2509\.04013,[Link](https://arxiv.org/abs/2509.04013)Cited by:[§6](https://arxiv.org/html/2608.10315#S6.p1.1)\.
- R\. Melamed, L\. H\. McCabe, T\. Wakhare, Y\. Kim, H\. H\. Huang, and E\. Boix\-Adserà \(2024\)Prompts have evil twins\.Miami, Florida, USA,pp\. 46–74\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.4/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.4)Cited by:[§6](https://arxiv.org/html/2608.10315#S6.p1.1)\.
- S\. J\. Mielke, A\. Szlam, E\. Dinan, and Y\. Boureau \(2022\)Reducing conversational agents’ overconfidence through linguistic calibration\.Transactions of the Association for Computational Linguistics10,pp\. 857–872\.External Links:[Link](https://aclanthology.org/2022.tacl-1.50/),[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00494)Cited by:[§5\.1](https://arxiv.org/html/2608.10315#S5.SS1.SSS0.Px1.p4.1)\.
- S\. Min, K\. Krishna, X\. Lyu, M\. Lewis, W\. Yih, P\. Koh, M\. Iyyer, L\. Zettlemoyer, and H\. Hajishirzi \(2023\)FActScore: fine\-grained atomic evaluation of factual precision in long form text generation\.Singapore,pp\. 12076–12100\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.741/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.741)Cited by:[Appendix J](https://arxiv.org/html/2608.10315#A10.SS0.SSS0.Px2.p1.1),[Appendix I](https://arxiv.org/html/2608.10315#A9.p2.1),[§4](https://arxiv.org/html/2608.10315#S4.p2.1),[§4](https://arxiv.org/html/2608.10315#S4.p3.1)\.
- M\. Mizrahi, G\. Kaplan, D\. Malkin, R\. Dror, D\. Shahaf, and G\. Stanovsky \(2024\)State of what art? a call for multi\-prompt LLM evaluation\.Transactions of the Association for Computational Linguistics12,pp\. 933–949\.External Links:[Link](https://aclanthology.org/2024.tacl-1.52/),[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00681)Cited by:[§6](https://arxiv.org/html/2608.10315#S6.SS0.SSS0.Px1.p1.1),[§6](https://arxiv.org/html/2608.10315#S6.p1.1)\.
- C\. Mohri and T\. Hashimoto \(2024\)Language models with conformal factuality guarantees\.InProceedings of the 41st International Conference on Machine LearningICLR Workshop: Quantify Uncertainty and Hallucination in Foundation Models: The Next Frontier in Reliable AIThe Thirteenth International Conference on Learning RepresentationsProceedings of the 2024 Conference on Empirical Methods in Natural Language ProcessingProceedings of the 2024 Conference on Empirical Methods in Natural Language ProcessingProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)Proceedings of the 2024 Conference on Empirical Methods in Natural Language ProcessingProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\)Advances in Neural Information Processing SystemsInternational Conference on Learning RepresentationsProceedings of the 2023 Conference on Empirical Methods in Natural Language ProcessingProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language TechnologiesProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\)Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\)Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\)Proceedings of the Thirty\-Eighth AAAI Conference on Artificial Intelligence and Thirty\-Sixth Conference on Innovative Applications of Artificial Intelligence and Fourteenth Symposium on Educational Advances in Artificial IntelligenceAdvances in Neural Information Processing SystemsProceedings of the 38th International Conference on Neural Information Processing SystemsProceedings of the 38th International Conference on Neural Information Processing SystemsProceedings of the 58th Annual Meeting of the Association for Computational LinguisticsProceedings of the 3rd Workshop on Trustworthy Natural Language Processing \(TrustNLP 2023\)Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)The Thirty\-eighth Annual Conference on Neural Information Processing SystemsProceedings on ”I Can’t Believe It’s Not Better: Failure Modes in the Age of Foundation Models” at NeurIPS 2023 WorkshopsProceedings of the ACM SIGOPS 29th Symposium on Operating Systems PrinciplesThe Twelfth International Conference on Learning RepresentationsFindings of the Association for Computational Linguistics: EMNLP 2025Findings of the Association for Computational Linguistics: EMNLP 2024Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\)Findings of the Association for Computational Linguistics: EMNLP 2023,R\. Salakhutdinov, Z\. Kolter, K\. Heller, A\. Weller, N\. Oliver, J\. Scarlett, F\. Berkenkamp, Y\. Al\-Onaizan, M\. Bansal, Y\. Chen, Y\. Al\-Onaizan, M\. Bansal, Y\. Chen, A\. Rogers, J\. Boyd\-Graber, N\. Okazaki, Y\. Al\-Onaizan, M\. Bansal, Y\. Chen, K\. Duh, H\. Gomez, S\. Bethard, I\. Guyon, U\. V\. Luxburg, S\. Bengio, H\. Wallach, R\. Fergus, S\. Vishwanathan, R\. Garnett, H\. Bouamor, J\. Pino, K\. Bali, J\. Burstein, C\. Doran, T\. Solorio, K\. Duh, H\. Gomez, S\. Bethard, A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, C\. Zhang, D\. Jurafsky, J\. Chai, N\. Schluter, J\. Tetreault, A\. Ovalle, K\. Chang, N\. Mehrabi, Y\. Pruksachatkun, A\. Galystan, J\. Dhamala, A\. Verma, T\. Cao, A\. Kumar, R\. Gupta, L\. Ku, A\. Martins, V\. Srikumar, J\. Antorán, A\. Blaas, K\. Buchanan, F\. Feng, V\. Fortuin, S\. Ghalebikesabi, A\. Kriegler, I\. Mason, D\. Rohde, F\. J\. R\. Ruiz, T\. Uelwer, Y\. Xie, R\. Yang, C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, V\. Peng, Y\. Al\-Onaizan, M\. Bansal, Y\. Chen, L\. Chiruzzo, A\. Ritter, L\. Wang, H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Proceedings of Machine Learning ResearchAAAI’24/IAAI’24/EAAI’24NIPS ’24NIPS ’24Proceedings of Machine Learning Research, Vol\.2353037239,pp\. 36029–36047\.External Links:[Link](https://proceedings.mlr.press/v235/mohri24a.html)Cited by:[§6](https://arxiv.org/html/2608.10315#S6.p1.1)\.
- OpenAI, :, A\. Hurst, A\. Lerer, A\. P\. Goucher, A\. Perelman, A\. Ramesh, A\. Clark,et al\.\(2024a\)GPT\-4o system card\.External Links:2410\.21276,[Link](https://arxiv.org/abs/2410.21276)Cited by:[§E\.1](https://arxiv.org/html/2608.10315#A5.SS1.p2.1)\.
- OpenAI, J\. Achiam, S\. Adler, S\. Agarwal, L\. Ahmad, I\. Akkaya,et al\.\(2024b\)GPT\-4 technical report\.External Links:2303\.08774,[Link](https://arxiv.org/abs/2303.08774)Cited by:[§E\.1](https://arxiv.org/html/2608.10315#A5.SS1.p1.1),[§6](https://arxiv.org/html/2608.10315#S6.p1.1)\.
- A\. Patel, S\. Bhattamishra, and N\. Goyal \(2021\)Are NLP models really able to solve simple math word problems?\.Online,pp\. 2080–2094\.External Links:[Link](https://aclanthology.org/2021.naacl-main.168),[Document](https://dx.doi.org/10.18653/v1/2021.naacl-main.168)Cited by:[Appendix I](https://arxiv.org/html/2608.10315#A9.p1.1),[§4](https://arxiv.org/html/2608.10315#S4.p2.1)\.
- G\. Portillo Wightman, A\. Delucia, and M\. Dredze \(2023\)Strength in numbers: estimating confidence of large language models by prompt agreement\.Toronto, Canada,pp\. 326–362\.External Links:[Link](https://aclanthology.org/2023.trustnlp-1.28/),[Document](https://dx.doi.org/10.18653/v1/2023.trustnlp-1.28)Cited by:[Appendix J](https://arxiv.org/html/2608.10315#A10.SS0.SSS0.Px1.p3.1),[§1](https://arxiv.org/html/2608.10315#S1.p2.1),[§4](https://arxiv.org/html/2608.10315#S4.p3.1)\.
- V\. Quach, A\. Fisch, T\. Schuster, A\. Yala, J\. H\. Sohn, T\. S\. Jaakkola, and R\. Barzilay \(2024\)Conformal language modeling\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=pzUhfQ74c5)Cited by:[§6](https://arxiv.org/html/2608.10315#S6.p1.1)\.
- N\. Rathi, D\. Jurafsky, and K\. Zhou \(2025\)Humans overrely on overconfident language models, across languages\.External Links:2507\.06306,[Link](https://arxiv.org/abs/2507.06306)Cited by:[§1](https://arxiv.org/html/2608.10315#S1.p2.1),[§5\.1](https://arxiv.org/html/2608.10315#S5.SS1.SSS0.Px1.p4.1)\.
- J\. Ren, Y\. Zhao, T\. Vu, P\. J\. Liu, and B\. Lakshminarayanan \(2023\)Self\-evaluation improves selective generation in large language models\.pp\. 49–64\.External Links:[Link](https://proceedings.mlr.press/v239/ren23a.html)Cited by:[§1](https://arxiv.org/html/2608.10315#S1.p2.1)\.
- M\. Riddell, A\. Ni, and A\. Cohan \(2024\)Quantifying contamination in evaluating code generation capabilities of language models\.Bangkok, Thailand,pp\. 14116–14137\.External Links:[Link](https://aclanthology.org/2024.acl-long.761/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.761)Cited by:[§5\.3](https://arxiv.org/html/2608.10315#S5.SS3.p4.1)\.
- M\. Sclar, Y\. Choi, Y\. Tsvetkov, and A\. Suhr \(2024\)Quantifying language models’ sensitivity to spurious features in prompt design or: how i learned to start worrying about prompt formatting\.External Links:2310\.11324,[Link](https://arxiv.org/abs/2310.11324)Cited by:[§6](https://arxiv.org/html/2608.10315#S6.SS0.SSS0.Px1.p1.1)\.
- A\. Sharma and C\. David \(2025\)Assessing correctness in llm\-based code generation via uncertainty estimation\.External Links:2502\.11620,[Link](https://arxiv.org/abs/2502.11620)Cited by:[§5\.1](https://arxiv.org/html/2608.10315#S5.SS1.SSS0.Px1.p3.1)\.
- M\. Shen, S\. Das, K\. Greenewald, P\. Sattigeri, G\. W\. Wornell, and S\. Ghosh \(2024\)Thermometer: towards universal calibration for large language models\.InProceedings of the 41st International Conference on Machine Learning,R\. Salakhutdinov, Z\. Kolter, K\. Heller, A\. Weller, N\. Oliver, J\. Scarlett, and F\. Berkenkamp \(Eds\.\),Proceedings of Machine Learning Research, Vol\.235,pp\. 44687–44711\.External Links:[Link](https://proceedings.mlr.press/v235/shen24c.html)Cited by:[§6](https://arxiv.org/html/2608.10315#S6.p1.1)\.
- A\. Singh, A\. Fry, A\. Perelman, A\. Tart,et al\.\(2025\)OpenAI gpt\-5 system card\.External Links:2601\.03267,[Link](https://arxiv.org/abs/2601.03267)Cited by:[§E\.1](https://arxiv.org/html/2608.10315#A5.SS1.p2.1)\.
- K\. Song, X\. Tan, T\. Qin, J\. Lu, and T\. Liu \(2020\)MPNet: masked and permuted pre\-training for language understanding\.arXiv preprint arXiv:2004\.09297\.Cited by:[§A\.2](https://arxiv.org/html/2608.10315#A1.SS2.p2.2)\.
- A\. Talmor, J\. Herzig, N\. Lourie, and J\. Berant \(2019\)CommonsenseQA: a question answering challenge targeting commonsense knowledge\.Minneapolis, Minnesota,pp\. 4149–4158\.External Links:[Link](https://aclanthology.org/N19-1421),[Document](https://dx.doi.org/10.18653/v1/N19-1421),1811\.00937Cited by:[Appendix I](https://arxiv.org/html/2608.10315#A9.p1.1),[§4](https://arxiv.org/html/2608.10315#S4.p2.1)\.
- G\. Team, R\. Anil, S\. Borgeaud, J\. Alayrac, J\. Yu,et al\.\(2025a\)Gemini: a family of highly capable multimodal models\.External Links:2312\.11805,[Link](https://arxiv.org/abs/2312.11805)Cited by:[§E\.1](https://arxiv.org/html/2608.10315#A5.SS1.p2.1)\.
- G\. Team, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova,et al\.\(2025b\)Gemma 3 technical report\.External Links:2503\.19786,[Link](https://arxiv.org/abs/2503.19786)Cited by:[§E\.1](https://arxiv.org/html/2608.10315#A5.SS1.p1.1)\.
- Q\. Team \(2025\)Qwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§4](https://arxiv.org/html/2608.10315#S4.p6.1)\.
- A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, and I\. Polosukhin \(2017\)Attention is all you need\.pp\.\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf)Cited by:[§3](https://arxiv.org/html/2608.10315#S3.p4.9)\.
- J\. P\. Wahle, T\. Ruas, Y\. Xu, and B\. Gipp \(2024\)Paraphrase types elicit prompt engineering capabilities\.Miami, Florida, USA,pp\. 11004–11033\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.617/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.617)Cited by:[§6](https://arxiv.org/html/2608.10315#S6.p1.1)\.
- X\. Wang, J\. Wei, D\. Schuurmans, Q\. V\. Le, E\. H\. Chi, S\. Narang, A\. Chowdhery, and D\. Zhou \(2023\)Self\-consistency improves chain of thought reasoning in language models\.InThe Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1\-5, 2023,External Links:[Link](https://openreview.net/forum?id=1PL1NIMMrw)Cited by:[Appendix J](https://arxiv.org/html/2608.10315#A10.SS0.SSS0.Px1.p2.1),[§1](https://arxiv.org/html/2608.10315#S1.p2.1),[§4](https://arxiv.org/html/2608.10315#S4.p3.1),[§6](https://arxiv.org/html/2608.10315#S6.p1.1)\.
- J\. Xie, A\. S\. Chen, Y\. Lee, E\. Mitchell, and C\. Finn \(2024\)Calibrating language models with adaptive temperature scaling\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 18128–18138\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.1007/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.1007)Cited by:[§6](https://arxiv.org/html/2608.10315#S6.p1.1)\.
- M\. Xiong, Z\. Hu, X\. Lu, Y\. LI, J\. Fu, J\. He, and B\. Hooi \(2024\)Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=gjeQKFxFpZ)Cited by:[§6](https://arxiv.org/html/2608.10315#S6.p1.1)\.
- F\. Xu, Y\. Song, M\. Iyyer, and E\. Choi \(2023\)A critical evaluation of evaluations for long\-form question answering\.Toronto, Canada,pp\. 3225–3245\.External Links:[Link](https://aclanthology.org/2023.acl-long.181/),[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.181)Cited by:[§6](https://arxiv.org/html/2608.10315#S6.p1.1)\.
- H\. Zhang, S\. Diao, Y\. Lin, Y\. Fung, Q\. Lian, X\. Wang, Y\. Chen, H\. Ji, and T\. Zhang \(2024a\)R\-tuning: instructing large language models to say ‘I don’t know’\.Mexico City, Mexico,pp\. 7113–7139\.External Links:[Link](https://aclanthology.org/2024.naacl-long.394/),[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.394)Cited by:[§6](https://arxiv.org/html/2608.10315#S6.p1.1)\.
- M\. Zhang, M\. Huang, R\. Shi, L\. Guo, C\. Peng, P\. Yan, Y\. Zhou, and X\. Qiu \(2024b\)Calibrating the confidence of large language models by eliciting fidelity\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 2959–2979\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.173/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.173)Cited by:[§6](https://arxiv.org/html/2608.10315#S6.p1.1)\.
- Y\. Zhang, J\. Baldridge, and L\. He \(2019\)PAWS: paraphrase adversaries from word scrambling\.Minneapolis, Minnesota,pp\. 1298–1308\.External Links:[Link](https://aclanthology.org/N19-1131/),[Document](https://dx.doi.org/10.18653/v1/N19-1131)Cited by:[§5\.1](https://arxiv.org/html/2608.10315#S5.SS1.SSS0.Px1.p4.1),[§6](https://arxiv.org/html/2608.10315#S6.p1.1)\.
- R\. Zhao, A\. Köksal, A\. Modarressi, M\. A\. Hedderich, and H\. Schuetze \(2025\)Do we know what LLMs don’t know? a study of consistency in knowledge probing\.Suzhou, China,pp\. 23254–23280\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.1263/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.1263),ISBN 979\-8\-89176\-335\-7Cited by:[§6](https://arxiv.org/html/2608.10315#S6.SS0.SSS0.Px1.p1.1)\.
## Appendix ADiscussion on Perturbation Noise source
### A\.1LLMs as Sampler of Noise
We utilize LLMs as samplers to generate semantic perturbations\. The noise generation process adheres to three critical principles to ensure the resulting samples are both challenging and informative:
- •Topic Alignment:Samples are generated to remain semantically aligned with the original input query to ensure the underlying task remains constant\.
- •Content Neutrality:The samples are designed to be content\-neutral, as non\-neutral content could introduce systematic bias or shifts in the model’s performance\.
- •Non\-Trivial Contextual Variations:The generated noise constitutes non\-trivial contextual variations of the input, testing the model’s robustness while maintaining the core meaning\.
we employ the prompts detailed in Appendix[D\.1](https://arxiv.org/html/2608.10315#A4.SS1)\. By utilizing LLMs as samplers constrained by these principles, we collect a set of perturbationsℰ=\{ϵ1,…,ϵn\}\\mathcal\{E\}=\\\{\\epsilon\_\{1\},\\dots,\\epsilon\_\{n\}\\\}\. To ensure the diversity ofℰ\\mathcal\{E\}, we filter out candidate noise that is semantically redundant with existing entries \(details are provided in Appendix[A\.2](https://arxiv.org/html/2608.10315#A1.SS2)\)\.
### A\.2Diversity of Noise Sampling
We utilize GPT\-4\.1 to perform noise sampling, aiming to maximize the diversity of noise per question\. It is well known that repeated generations from LLMs using the same prompt often produce semantically similar or even identical outputs\. To ensure variety, we maintain a clean pool of samples by comparing each newly generated data pointϵ∗\\epsilon^\{\*\}against an existing setℰ=\{ϵ1,ϵ2,…,ϵn\}\\mathcal\{E\}=\\\{\\epsilon\_\{1\},\\epsilon\_\{2\},\\dots,\\epsilon\_\{n\}\\\}in embedding space\. A new sampleϵ∗\\epsilon^\{\*\}is retained only if its pairwise semantic similarity with the existing elements in𝒩\\mathcal\{N\}remains below a thresholdkk\.
We compute embeddings usingall\-mpnet\-base\-v2\[Songet al\.,[2020](https://arxiv.org/html/2608.10315#bib.bib45)\]and measure similarity via cosine similarity\. However, the choice of similarity thresholdkkis typically heuristic and task\-specific\. In this work, we propose to empirically determine an appropriatekkspecific to GPT\-4\.1’s generation behavior, enabling us to maintain a semantically diverse pool\.
To do this, we sampled 30 premises generated by GPT\-4\.1 and used the same model to produce paraphrased versions using the prompt provided in Appendix[D\.2](https://arxiv.org/html/2608.10315#A4.SS2)\. For each original premise, we generated 50 paraphrases and embedded them usingall\-mpnet\-base\-v2\. Within each group of 50 rewrites, we computed all pairwise cosine similarities, resulting in\(502\)\\binom\{50\}\{2\}similarity scores per group\. Aggregating these across all groups yields the overall distribution of cosine similarities, representing GPT\-4\.1’s implicit notion of semantic equivalence\. The histogram of these results is shown in Figure[5](https://arxiv.org/html/2608.10315#A1.F5)\.
Figure 5:The distribution of pairwise cosine similarities among GPT\-4\.1\-generated paraphrases\. The red line indicates the 5th percentile \(k=0\.8124k=0\.8124\), which serves as our empirical threshold for filtering semantic redundancy\.The threshold is determined at a significance level ofα=0\.05\\alpha=0\.05\. This ensures that by retaining only those samples with a similarity belowk=0\.8124k=0\.8124, we have a statistical confidence that at most 5% of the accepted samples are semantically redundant paraphrases\.
### A\.3Semantic Neutrality of Perturbations
Figure[6](https://arxiv.org/html/2608.10315#A1.F6)evaluates whether our sampled perturbations introduce systematic answer\-relevant bias\. We conduct this check on MMLU High School Statistics across 26 models\. For each model\-question pair, we sample 30 generations from the original prompt and 30 generations from the perturbed prompt\. We then compute:
Δ=\#Correctperturbed−\#Correctoriginal,\\Delta=\\\#\\text\{Correct\}\_\{\\text\{perturbed\}\}\-\\\#\\text\{Correct\}\_\{\\text\{original\}\},where positive values indicate that the perturbation improves performance and negative values indicate that it hurts performance\.
If the added context revealed information about the answer, contradicted the question, or otherwise biased the model toward or away from the correct option, we would expect the distribution ofΔ\\Deltato shift systematically above or below zero\. In contrast, Figure[6](https://arxiv.org/html/2608.10315#A1.F6)shows that the distribution is centered near zero across all 26 models, with no consistent positive or negative shift\. This suggests that the perturbations do not systematically help or mislead the models on this benchmark\.
This analysis does not prove that every individual perturbation is perfectly neutral, but it provides empirical evidence that the perturbation procedure does not introduce a systematic directional bias in model performance\. We therefore treat these perturbations as approximately content\-neutral for the purposes of measuring cross\-contextual consistency\.
Figure 6:The difference of performance of each model before and after the perturbation noises are added\. The red line denote the 0 meaning no difference in the performances\.
### A\.4Topic Alignment
We also evaluate whether the sampled perturbations remain topic\-aligned with the benchmark from which they are generated\. Topic alignment means that the perturbation should stay within the broad domain or capability being tested, without revealing, contradicting, or otherwise modifying the answer\. This requirement separates our perturbations from arbitrary distractor text: the added context should create meaningful contextual variation, but should still be relevant to the type of task being evaluated\.
To assess topic alignment, we conduct an automatic topic\-classification check\. For each benchmark, we used whole perturbations generated by our pipeline with GPT\-4\.1 and remove the corresponding original question and answer\. We then ask an independent Qwen3\-8B judge to classify each perturbation into one of the broad benchmark\-level domains: arithmetic reasoning, statistical reasoning, commonsense reasoning, short\-form factual recall, long\-form factual or biographical generation, and code generation\. The judge observes only the perturbation itself, not the original query, answer, or benchmark label\. A perturbation is counted as topic\-aligned if the judge assigns it to the same broad domain as the benchmark from which it was sampled\.
Across the perturbations, approximately 96% are classified into their intended benchmark domain\. This suggests that the perturbations are not arbitrary out\-of\-domain distractors but instead remain aligned with the capability being evaluated\. Together with the neutrality analysis in Appendix[A\.3](https://arxiv.org/html/2608.10315#A1.SS3), this supports the use of our sampled perturbations as topic\-aligned, answer\-neutral contextual variations for estimating cross\-contextual consistency\.
## Appendix BEmpirical Choice of number of trials
To choose an appropriate number of trials, we conducted a case study to determine how many samples per question are needed to reliably characterize model behavior before and after perturbation\. As shown in Figure[7](https://arxiv.org/html/2608.10315#A2.F7), the estimated behavior stabilizes after about 20 samples\. We therefore use 30 trials per question as a more conservative choice in this setting\.
Figure 7:The performance of models across number of trials of samples we collected when perturbation is presented\.
## Appendix CC3 Evaluation on Model\-Level
Table 2:MMLU High School Stats Model Performance MetricsMetricECEAUROCPR\-PPR\-Ngemini\-2\.5\-flashself\-consistency0\.3590\.5060\.5270\.495Paraphrasing0\.3940\.4560\.4520\.472C30\.3270\.5490\.5690\.501Self\-report0\.2710\.4920\.7410\.237gemini\-2\.5\-flash\-liteself\-consistency0\.7950\.5030\.1660\.835Paraphrasing0\.8080\.4500\.1500\.789C30\.8180\.5230\.1720\.848Self\-report0\.3740\.5190\.5790\.442gemma\-3\-12b\-itself\-consistency0\.3040\.6460\.6500\.568Paraphrasing0\.3100\.6090\.6770\.590C30\.2410\.6790\.6890\.592Self\-report0\.3950\.5110\.5890\.446gemma\-3\-27b\-itself\-consistency0\.3260\.6470\.6630\.543Paraphrasing0\.3290\.6080\.6600\.555C30\.2600\.7060\.7210\.613Self\-report0\.4230\.4580\.5490\.420gemma\-3\-4b\-itself\-consistency0\.6020\.6280\.3350\.798Paraphrasing0\.6660\.5960\.3480\.785C30\.5040\.5950\.3240\.776Self\-report0\.5300\.4960\.3070\.716gpt\-4\.1self\-consistency0\.2020\.7390\.8390\.503Paraphrasing0\.2260\.6980\.8270\.482C30\.1800\.7530\.8560\.528Self\-report0\.2730\.4990\.7400\.269gpt\-4\.1\-miniself\-consistency0\.4160\.5360\.4930\.561Paraphrasing0\.4180\.5060\.4940\.506C30\.4040\.5210\.5020\.546Self\-report0\.2330\.5920\.8790\.232gpt\-4\.1\-nanoself\-consistency0\.3740\.4870\.4940\.490Paraphrasing0\.3570\.5160\.5040\.455C30\.3580\.4930\.5100\.502Self\-report0\.3670\.4800\.6440\.332
MetricECEAUROCPR\-PPR\-Nllama\-3\.1\-405b\-instructself\-consistency0\.1220\.7840\.8490\.618Paraphrasing0\.1640\.7570\.8280\.633C30\.0770\.7930\.8580\.639Self\-report0\.3000\.5080\.7430\.274llama\-3\.1\-70b\-instructself\-consistency0\.2110\.5980\.6570\.500Paraphrasing0\.2090\.5360\.6170\.523C30\.2090\.5490\.6350\.456Self\-report0\.2760\.5200\.7470\.271llama\-3\.1\-8b\-instructself\-consistency0\.3660\.4900\.3080\.686Paraphrasing0\.3950\.4800\.2600\.625C30\.2850\.4940\.3080\.700Self\-report0\.4890\.4910\.3730\.632llama\-3\.2\-1b\-instructself\-consistency0\.4920\.5380\.2200\.832Paraphrasing0\.5400\.5170\.1800\.773C30\.3600\.5690\.2150\.853Self\-report0\.5760\.5170\.1630\.862llama\-3\.2\-3b\-instructself\-consistency0\.4700\.4770\.1920\.814Paraphrasing0\.4760\.4800\.1980\.759C30\.3580\.5440\.2280\.846Self\-report0\.5400\.4770\.2860\.696mistral\-large\-2411self\-consistency0\.3710\.4840\.5480\.433Paraphrasing0\.4320\.4240\.5070\.423C30\.3800\.5210\.5730\.463Self\-report0\.3610\.4030\.6850\.227mistral\-medium\-3self\-consistency0\.2860\.6290\.7310\.445Paraphrasing0\.3170\.6000\.6540\.493C30\.2200\.7510\.8230\.547Self\-report0\.2740\.4990\.7720\.242mistral\-small\-24bself\-consistency0\.4200\.5420\.5050\.557Paraphrasing0\.4640\.5210\.4720\.586C30\.4220\.5150\.4910\.526Self\-report0\.3700\.4780\.7110\.280
Table 3:CommonsenseQA Model Performance MetricsMetricECEAUROCPR\-PPR\-Ngemini\-2\.5\-flashself\-consistency0\.1460\.7370\.8890\.394Paraphrasing0\.1820\.7030\.8840\.401C30\.1500\.7410\.9030\.381Self\-report0\.1890\.5110\.7690\.261gemini\-2\.5\-flash\-liteself\-consistency0\.7850\.5520\.1570\.868Paraphrasing0\.8390\.5520\.1580\.880C30\.8080\.5840\.1680\.892Self\-report0\.1780\.5820\.7890\.373gemma\-3\-12b\-itself\-consistency0\.2280\.5710\.7550\.358Paraphrasing0\.2610\.5690\.7640\.396C30\.1660\.7050\.8320\.512Self\-report0\.2190\.4930\.7230\.309gemma\-3\-27b\-itself\-consistency0\.2060\.5970\.7800\.385Paraphrasing0\.2440\.5470\.7470\.399C30\.1540\.7310\.8540\.539Self\-report0\.2080\.5750\.7750\.321gemma\-3\-4b\-itself\-consistency0\.3340\.5580\.6540\.444Paraphrasing0\.3690\.5090\.6460\.440C30\.2250\.6880\.7450\.572Self\-report0\.3380\.4760\.5880\.409gpt\-4\.1self\-consistency0\.1460\.6500\.8520\.419Paraphrasing0\.1810\.6370\.8750\.440C30\.1390\.7210\.8860\.408Self\-report0\.1100\.5610\.8410\.226gpt\-4\.1\-miniself\-consistency0\.1780\.5830\.8150\.269Paraphrasing0\.2110\.5660\.8100\.295C30\.1560\.6910\.8630\.462Self\-report0\.1440\.5710\.8290\.251gpt\-4\.1\-nanoself\-consistency0\.2280\.6620\.7960\.401Paraphrasing0\.2620\.7040\.8520\.453C30\.1970\.6900\.8160\.475Self\-report0\.1580\.5510\.7280\.412
MetricECEAUROCPR\-PPR\-Nllama\-3\.1\-405b\-instructself\-consistency0\.1420\.7520\.8690\.531Paraphrasing0\.1990\.7620\.8930\.571C30\.1170\.7910\.8940\.576Self\-report0\.1600\.5780\.7660\.368llama\-3\.1\-70b\-instructself\-consistency0\.1320\.7800\.8930\.533Paraphrasing0\.1940\.7420\.8880\.534C30\.1150\.8130\.9240\.525Self\-report0\.1790\.5720\.7600\.358llama\-3\.1\-8b\-instructself\-consistency0\.1390\.7220\.8020\.584Paraphrasing0\.1900\.7190\.8100\.591C30\.1100\.7450\.8510\.573Self\-report0\.2260\.5380\.7140\.366llama\-3\.2\-1b\-instructself\-consistency0\.1300\.6750\.6990\.623Paraphrasing0\.2270\.6800\.7180\.627C30\.0930\.6840\.7180\.607Self\-report0\.2840\.5070\.6010\.450llama\-3\.2\-3b\-instructself\-consistency0\.2190\.7030\.7430\.599Paraphrasing0\.3020\.6860\.7250\.609C30\.1280\.7380\.7970\.637Self\-report0\.2740\.4720\.5820\.431mistral\-large\-2411self\-consistency0\.1890\.5770\.7990\.329Paraphrasing0\.2250\.6180\.8490\.380C30\.1440\.7600\.8850\.556Self\-report0\.1370\.5760\.8430\.241mistral\-medium\-3self\-consistency0\.1830\.5620\.8030\.270Paraphrasing0\.2190\.5240\.7800\.276C30\.1460\.7420\.8810\.484Self\-report0\.1160\.5420\.8100\.243mistral\-small\-24b\-instruct\-2501self\-consistency0\.2020\.6130\.7960\.368Paraphrasing0\.2360\.6310\.8200\.449C30\.1720\.7220\.8610\.460Self\-report0\.1340\.5650\.8110\.300
Table 4:FactScore Model Performance MetricsMetricECEAUROCPR\-PPR\-Ngemini\-2\.5\-flashself\-consistency0\.0750\.8930\.9880\.399Paraphrasing0\.1100\.8440\.9120\.376C30\.0500\.9240\.9920\.464Self\-report0\.3470\.3670\.8560\.095gemini\-2\.5\-flash\-liteself\-consistency0\.1430\.7180\.9450\.279Paraphrasing0\.1560\.6650\.9300\.233C30\.0690\.7410\.9570\.247Self\-report0\.2770\.5530\.8950\.162gemma\-3\-12b\-itself\-consistency0\.3430\.9150\.9200\.926Paraphrasing0\.3490\.8770\.9470\.948C30\.2650\.9310\.9360\.930Self\-report0\.4110\.4430\.5220\.416gemma\-3\-27b\-itself\-consistency0\.3010\.9230\.9300\.929Paraphrasing0\.3040\.8840\.9270\.941C30\.2220\.9730\.9810\.971Self\-report0\.3510\.4790\.5970\.391gemma\-3\-4b\-itself\-consistency0\.4110\.9350\.9360\.948Paraphrasing0\.4750\.9030\.9490\.935C30\.3420\.9370\.9030\.961Self\-report0\.3610\.5600\.4680\.699gpt\-4\.1self\-consistency0\.0801\.0001\.0001\.000Paraphrasing0\.1030\.9590\.9880\.979C30\.0590\.9840\.9990\.804Self\-report0\.2890\.3320\.8920\.068gpt\-4\.1\-miniself\-consistency0\.0820\.9770\.9970\.735Paraphrasing0\.0840\.9480\.9990\.680C30\.0640\.9890\.9990\.915Self\-report0\.3160\.4750\.8710\.128gpt\-4\.1\-nanoself\-consistency0\.1360\.7880\.9690\.633Paraphrasing0\.1190\.8170\.9790\.597C30\.0940\.8700\.9880\.342Self\-report0\.3560\.4920\.9360\.093
MetricECEAUROCPR\-PPR\-Nllama\-3\.1\-405b\-instructself\-consistency0\.0980\.9070\.9850\.642Paraphrasing0\.1400\.8800\.9640\.657C30\.0690\.9200\.9880\.540Self\-report0\.2970\.5600\.8850\.219llama\-3\.1\-70b\-instructself\-consistency0\.2160\.9060\.9690\.771Paraphrasing0\.2150\.8450\.9280\.794C30\.1410\.9130\.9710\.780Self\-report0\.3210\.4770\.7160\.302llama\-3\.1\-8b\-instructself\-consistency0\.2930\.7390\.8120\.678Paraphrasing0\.3230\.7290\.7640\.617C30\.2150\.7890\.8170\.737Self\-report0\.3930\.5150\.6110\.450llama\-3\.2\-1b\-instructself\-consistency0\.4110\.7030\.5280\.874Paraphrasing0\.4590\.6820\.4880\.815C30\.2780\.6960\.5800\.837Self\-report0\.2980\.5410\.2920\.782llama\-3\.2\-3b\-instructself\-consistency0\.2990\.6980\.7180\.767Paraphrasing0\.3040\.7020\.7240\.713C30\.2150\.6920\.7020\.719Self\-report0\.3760\.5050\.3940\.641mistral\-large\-2411self\-consistency0\.2060\.9290\.9850\.804Paraphrasing0\.2660\.8680\.9440\.793C30\.1130\.9490\.9910\.764Self\-report0\.3050\.5430\.8340\.273mistral\-medium\-3self\-consistency0\.2400\.9290\.9790\.737Paraphrasing0\.2710\.9000\.9020\.786C30\.1570\.8900\.9680\.599Self\-report0\.2680\.5120\.7740\.277mistral\-small\-24b\-instruct\-2501self\-consistency0\.2680\.7700\.8770\.568Paraphrasing0\.3130\.7480\.8450\.598C30\.2020\.7730\.8790\.613Self\-report0\.3160\.5850\.7580\.369
Table 5:HumanEval Model Performance MetricsMetricECEAUROCPR\-PPR\-Ngemini\-2\.5\-flashself\-consistency0\.1790\.8650\.9580\.193Paraphrasing0\.2130\.8310\.9250\.159C30\.1650\.9110\.9980\.237Self\-report0\.2700\.4960\.9850\.025gemini\-2\.5\-flash\-liteself\-consistency0\.2790\.9090\.9580\.249Paraphrasing0\.3300\.8590\.9070\.199C30\.2370\.9410\.9990\.282Self\-report0\.2180\.8100\.9960\.075gemma\-3\-12b\-itself\-consistency0\.0870\.7820\.9160\.493Paraphrasing0\.1190\.7500\.8830\.461C30\.0550\.8620\.9690\.605Self\-report0\.3110\.4330\.8100\.152gemma\-3\-27b\-itself\-consistency0\.1160\.7140\.9020\.302Paraphrasing0\.1590\.6710\.8600\.260C30\.0690\.7680\.9540\.360Self\-report0\.2070\.5090\.8720\.142gemma\-3\-4b\-itself\-consistency0\.0930\.7400\.8230\.635Paraphrasing0\.1340\.6990\.7820\.594C30\.0490\.7920\.8580\.691Self\-report0\.2940\.5350\.7210\.315gpt\-4\.1self\-consistency0\.1100\.8570\.9640\.236Paraphrasing0\.1400\.8260\.9340\.205C30\.0940\.8440\.9920\.239Self\-report0\.2720\.6180\.9660\.079gpt\-4\.1\-miniself\-consistency0\.1360\.8540\.9580\.179Paraphrasing0\.1520\.8390\.9420\.163C30\.1130\.8830\.9950\.190Self\-report0\.2540\.4830\.9570\.052gpt\-4\.1\-nanoself\-consistency0\.1000\.8120\.9510\.547Paraphrasing0\.1290\.7830\.9220\.518C30\.1030\.8470\.9760\.553Self\-report0\.2730\.5290\.9240\.100
MetricECEAUROCPR\-PPR\-Nllama\-3\.1\-405b\-instructself\-consistency0\.1050\.8190\.9270\.585Paraphrasing0\.1430\.7810\.8890\.547C30\.1040\.8060\.9440\.560Self\-report0\.2840\.4870\.7800\.205llama\-3\.1\-70b\-instructself\-consistency0\.1400\.8350\.9280\.700Paraphrasing0\.1610\.8140\.9070\.679C30\.1130\.8520\.9560\.656Self\-report0\.3630\.3670\.7340\.170llama\-3\.1\-8b\-instructself\-consistency0\.1160\.8380\.8760\.787Paraphrasing0\.1480\.8050\.8430\.754C30\.0930\.8750\.9180\.841Self\-report0\.3650\.5840\.6840\.467llama\-3\.2\-1b\-instructself\-consistency0\.2360\.7510\.4950\.850Paraphrasing0\.2560\.7300\.4740\.829C30\.1300\.8510\.7540\.917Self\-report0\.3370\.5170\.3090\.728llama\-3\.2\-3b\-instructself\-consistency0\.0940\.8680\.8680\.873Paraphrasing0\.1300\.8320\.8320\.837C30\.0890\.8710\.8740\.874Self\-report0\.3810\.4970\.5110\.491mistral\-large\-2411self\-consistency0\.1000\.7720\.9370\.452Paraphrasing0\.1380\.7330\.8980\.413C30\.0970\.7810\.9580\.467Self\-report0\.2450\.5830\.9060\.154mistral\-medium\-3self\-consistency0\.1050\.8130\.9530\.286Paraphrasing0\.1370\.7810\.9210\.253C30\.1210\.8620\.9810\.401Self\-report0\.2540\.6140\.9560\.091mistral\-small\-24b\-instruct\-2501self\-consistency0\.1040\.7690\.9190\.434Paraphrasing0\.1570\.7160\.8660\.381C30\.1230\.8020\.9580\.445Self\-report0\.2630\.3810\.8040\.128
Table 6:SimpleQA Model Performance MetricsMetricECEAUROCPR\-PPR\-Ngemini\-2\.5\-flashself\-consistency0\.2020\.8380\.6940\.909Paraphrasing0\.2190\.8210\.6770\.892C30\.1140\.8270\.6330\.924Self\-report0\.6990\.4620\.2460\.735gemini\-2\.5\-flash\-liteself\-consistency0\.3080\.8720\.5130\.954Paraphrasing0\.3590\.8220\.4620\.904C30\.1290\.8350\.3560\.974Self\-report0\.7880\.4530\.1170\.865gemma\-3\-12b\-itself\-consistency0\.4870\.7810\.1120\.970Paraphrasing0\.5020\.7650\.0960\.955C30\.1740\.8290\.1990\.988Self\-report0\.8980\.4480\.0580\.928gemma\-3\-27b\-itself\-consistency0\.5840\.6750\.1570\.906Paraphrasing0\.6190\.6390\.1210\.870C30\.2440\.7540\.2310\.961Self\-report0\.8490\.3640\.0690\.886gemma\-3\-4b\-itself\-consistency0\.5400\.7850\.0570\.983Paraphrasing0\.5700\.7540\.0260\.952C30\.1930\.8280\.0930\.994Self\-report0\.9520\.6690\.0400\.986gpt\-4\.1self\-consistency0\.2390\.7620\.6590\.807Paraphrasing0\.2490\.7520\.6480\.796C30\.2030\.7420\.6490\.797Self\-report0\.5590\.4560\.3550\.609gpt\-4\.1\-miniself\-consistency0\.3010\.8500\.4820\.946Paraphrasing0\.2830\.8680\.5000\.965C30\.1360\.8640\.5580\.968Self\-report0\.7440\.5110\.1810\.836gpt\-4\.1\-nanoself\-consistency0\.2330\.8970\.4150\.999Paraphrasing0\.2410\.8890\.4060\.991C30\.1330\.8480\.3260\.979Self\-report0\.8120\.4510\.0900\.912
MetricECEAUROCPR\-PPR\-Nllama\-3\.1\-405b\-instructself\-consistency0\.1720\.6350\.5080\.713Paraphrasing0\.1980\.6080\.4820\.686C30\.2310\.6740\.5170\.745Self\-report0\.5260\.5380\.4380\.623llama\-3\.1\-70b\-instructself\-consistency0\.1670\.8340\.5480\.957Paraphrasing0\.1580\.8430\.5560\.965C30\.0860\.8900\.6000\.972Self\-report0\.7560\.4480\.1410\.825llama\-3\.1\-8b\-instructself\-consistency0\.1860\.7500\.3070\.924Paraphrasing0\.2010\.7350\.2920\.909C30\.0520\.8060\.2860\.974Self\-report0\.8670\.5240\.0680\.944llama\-3\.2\-1b\-instructself\-consistency0\.8570\.7980\.0100\.975Paraphrasing0\.8480\.8070\.0180\.983C30\.2460\.9420\.1250\.999Self\-report0\.9000\.8730\.0530\.996llama\-3\.2\-3b\-instructself\-consistency0\.1950\.9560\.4330\.989Paraphrasing0\.2170\.9340\.4100\.967C30\.0490\.9470\.3000\.998Self\-report0\.8480\.2340\.0270\.952mistral\-large\-2411self\-consistency0\.6280\.7610\.4220\.911Paraphrasing0\.6550\.7340\.3940\.883C30\.2020\.8240\.5410\.936Self\-report0\.7200\.4640\.2140\.753mistral\-medium\-3self\-consistency0\.6980\.7200\.3100\.895Paraphrasing0\.7120\.7060\.2960\.881C30\.2330\.8100\.5160\.944Self\-report0\.7220\.4890\.2170\.798mistral\-small\-24b\-instruct\-2501self\-consistency0\.4900\.7540\.2570\.948Paraphrasing0\.5470\.6970\.2000\.891C30\.2320\.7540\.2990\.949Self\-report0\.8100\.4510\.1070\.886
Table 7:SVAMP Model Performance MetricsMetricECEAUROCPR\-PPR\-Ngemini\-2\.5\-flashself\-consistency0\.0470\.9340\.9860\.766Paraphrasing0\.0620\.9140\.9660\.746C30\.0230\.9780\.9980\.817Self\-report0\.1060\.5230\.8940\.150gemini\-2\.5\-flash\-liteself\-consistency0\.1210\.9420\.9760\.783Paraphrasing0\.1360\.9220\.9560\.763C30\.0340\.9520\.9860\.860Self\-report0\.2160\.5000\.7840\.216gemma\-3\-12b\-itself\-consistency0\.2080\.7990\.8570\.678Paraphrasing0\.2230\.7790\.8370\.658C30\.0750\.9130\.9460\.834Self\-report0\.3000\.5000\.7000\.300gemma\-3\-27b\-itself\-consistency0\.1780\.8330\.9090\.649Paraphrasing0\.1930\.8130\.8890\.629C30\.0640\.9300\.9700\.797Self\-report0\.2400\.5000\.7600\.240gemma\-3\-4b\-itself\-consistency0\.3760\.7770\.7220\.756Paraphrasing0\.3910\.7570\.7020\.736C30\.1170\.8460\.8350\.805Self\-report0\.4350\.4960\.5630\.435gpt\-4\.1self\-consistency0\.0620\.8630\.9710\.709Paraphrasing0\.0770\.8430\.9510\.689C30\.0260\.9270\.9850\.787Self\-report0\.0930\.5530\.9140\.190gpt\-4\.1\-miniself\-consistency0\.0780\.8720\.9640\.724Paraphrasing0\.0930\.8520\.9440\.704C30\.0420\.9290\.9830\.786Self\-report0\.1010\.6050\.9090\.150gpt\-4\.1\-nanoself\-consistency0\.1590\.8530\.9160\.723Paraphrasing0\.1740\.8330\.8960\.703C30\.0900\.8780\.9430\.688Self\-report0\.2530\.5570\.7460\.308
MetricECEAUROCPR\-PPR\-Nllama\-3\.1\-405b\-instructself\-consistency0\.0320\.9530\.9930\.711Paraphrasing0\.0470\.9330\.9730\.691C30\.0780\.9450\.9920\.691Self\-report0\.1280\.5470\.8810\.186llama\-3\.1\-70b\-instructself\-consistency0\.0620\.9270\.9800\.693Paraphrasing0\.0770\.9070\.9600\.673C30\.0670\.9430\.9880\.741Self\-report0\.1650\.5120\.8380\.175llama\-3\.1\-8b\-instructself\-consistency0\.0530\.9040\.9020\.901Paraphrasing0\.0680\.8840\.8820\.881C30\.1560\.9070\.9040\.917Self\-report0\.3620\.5380\.6390\.419llama\-3\.2\-1b\-instructself\-consistency0\.1410\.8560\.7510\.908Paraphrasing0\.1560\.8360\.7310\.888C30\.1000\.8810\.7970\.935Self\-report0\.4580\.6130\.5680\.578llama\-3\.2\-3b\-instructself\-consistency0\.1510\.8440\.7150\.905Paraphrasing0\.1660\.8240\.6950\.885C30\.0840\.8560\.7770\.907Self\-report0\.4180\.5090\.5730\.430mistral\-large\-2411self\-consistency0\.1610\.7690\.9010\.590Paraphrasing0\.1760\.7490\.8810\.570C30\.0700\.9470\.9840\.790Self\-report0\.1700\.4970\.8290\.170mistral\-medium\-3self\-consistency0\.2250\.7960\.8690\.633Paraphrasing0\.2400\.7760\.8490\.613C30\.0400\.9370\.9700\.869Self\-report0\.1660\.5560\.8460\.256mistral\-small\-24b\-instruct\-2501self\-consistency0\.2220\.8510\.8800\.703Paraphrasing0\.2370\.8310\.8600\.683C30\.0800\.9010\.9220\.865Self\-report0\.2300\.5000\.7700\.230
## Appendix DPrompts
### D\.1The Prompt for Sampling
`Prompt: LLMs as Noise Samplers`
`D\.2 The Prompt for Sampling Prompt: Paraphrasing D\.3 The Prompt for Self\-Confidence Prompt: Self\-Confidence D\.4 The Prompt for Benchmarks Prompt: SVAMP Prompt: MMLU High School Stats Prompt: SimpleQA Prompt: FActScore Prompt: HumanEval Prompt: CommonsenseQA D\.5 The Prompt for LLMs as Judges Prompt: Judge Same Answer Appendix E Experiment Setup E\.1 Models We evaluate C3 on 16 widely used LLMs spanning multiple families and scales: OpenAI’s GPT\-4\.1 series \(4\.1, 4\.1\-mini, 4\.1\-nano\) \[OpenAI et al\., 2024b\]; Google’s Gemma\-3 Instruct models \(4B, 12B, 27B\) \[Team et al\., 2025b\] and Gemini models \(Gemini\-2\.5\-Flash, Gemini\-2\.5\-Flash\-Lite\) \[Comanici et al\., 2025\]; Meta’s Llama\-3 Instruct models \(1B, 3B, 8B, 70B, 405B\) \[Grattafiori et al\., 2024\]; and Mistral models \(Small, Medium, Large\) \[Jiang et al\., 2024a\]\. We also conducted case studies on additional models; however, due to their characteristics: such as heavier reasoning processes or deprecated designs, they are often costly in time and computational resources\. We therefore restrict these case studies to the MMLU High School Statistics benchmark\. To study the temporal evolution of C3 in Figure 1, we additionally include earlier and newer frontier models, including GPT\-3\.5\-Turbo \[Brown et al\., 2020\], GPT\-4\-Turbo, GPT\-4o, and GPT\-4o\-mini \[OpenAI et al\., 2024a\]; GPT\-5 \(5, 5\-mini, 5\-nano\) \[Singh et al\., 2025\]; and additional Gemini releases \(Gemini\-2\.5\-Pro, Gemini\-2\.0\-Flash, and Gemini\-2\.0\-Flash\-Lite\) \[Team et al\., 2025a\]\. Across all experiments, we use temperature T=1T\{=\}1 to probe typical stochastic generation behavior under standard decoding\. E\.2 Evaluation Metrics We evaluate C3 and other baseline scores, all normalized to the range \[0,1\]\[0,1\]\. A key aspect of these methods is to disclose the reliability of model generations: Their alignment with correctness or truthfulness\. For each benchmark instance, we estimate an empirical model performance by aggregating outcomes over nn sampling trials \(repeated samples of stochastic decoding\) and taking the per\-instance average\. We then evaluate how well each score associated with instances aligns with this per\-instance average performance using a combination of calibration and ranking metrics\. To measure calibration, we report the Expected Calibration Error \(ECE\), computed by partitioning instances into BB bins according to their scores as ECE=∑b=1B\|Ib\|N\|perf\(Ib\)−score\(Ib\)\|\\mathrm\{ECE\}=\\sum\_\{b=1\}^\{B\}\\frac\{\|I\_\{b\}\|\}\{N\}\\left\|\\mathrm\{perf\}\(I\_\{b\}\)\-\\mathrm\{score\}\(I\_\{b\}\)\\right\|, where NN is the number of instances, perf\(Ib\)\\mathrm\{perf\}\(I\_\{b\}\) denotes the average performance of instances in bin bb, and score\(Ib\)\\mathrm\{score\}\(I\_\{b\}\) denotes the average predicted score in bin bb; lower ECE values indicate better calibration\. For ranking\-based evaluation, we define binary labels by thresholding the per\-instance average performance as yi=𝟏\[p¯i≥0\.5\]y\_\{i\}=\\mathbf\{1\}\[\\bar\{p\}\_\{i\}\\geq 0\.5\], and report AUROC by treating the score for instance ii as a ranking signal to separate instances with yi=1y\_\{i\}=1 from those with yi=0y\_\{i\}=0\. In addition, we report precision\-recall metrics that evaluate how well a score ranks instances by correctness, including the area under the precision–recall curve for detecting correct outputs \(AUPRC\-P\) and for detecting incorrect outputs \(AUPRC\-N\), with the latter using 1−si1\-s\_\{i\}\. Appendix F Ablation Study on Source of Noises SimpleQA Method ECE ↓\\downarrow AUROC ↑\\uparrow AUPRC\-P ↑\\uparrow AUPRC\-N ↑\\uparrow Self\-consistency 0\.393 0\.792 0\.368 0\.924 Paraphrasing 0\.411 0\.773 0\.349 0\.906 Self\-report 0\.778 0\.490 0\.151 0\.846 C3 \(GPT\-4\.1\) \\cellcolorgray\!200\.166 0\.823 \\cellcolorgray\!200\.389 \\cellcolorgray\!200\.944 C3 \(Qwen3\-8B\) 0\.235 \\cellcolorgray\!200\.833 0\.378 0\.931 C3 \(Web Source\) 0\.247 0\.801 0\.355 0\.894 Table 8: Calibration results on SimpleQA comparing C3 under different perturbation sources\. We compare perturbations generated by GPT\-4\.1, perturbations generated by the open model Qwen3\-8B, and randomly sampled web\-sourced noise\. Results show that C3 remains competitive across perturbation sources, suggesting that its calibration signal is not solely dependent on frontier\-model\-generated perturbations\. Although perturbations generated by GPT\-4\.1 are preferred in our main experiments because they consistently produce higher\-quality generations and better satisfy the intended properties of diversity and content neutrality, C3 does not fundamentally depend on GPT\-4\.1 as the perturbation source\. To test this, we compare GPT\-4\.1 perturbations with two alternative sources: \(1\) perturbations generated by the open\-weight Qwen3\-8B model, and \(2\) random web\-sourced noise of approximately the same length sampled from RedPajama \[Computer, 2023\], which contains internet\-scraped text from a broad range of domains\. Effectiveness of noise\. As shown in Table 8, C3 remains competitive when perturbations are generated from non\-frontier sources\. GPT\-4\.1 achieves the best ECE, AUPRC\-P, and AUPRC\-N, suggesting that higher\-quality perturbations can improve calibration, especially in terms of probability calibration and precision\-recall behavior\. However, Qwen3\-8B achieves the highest AUROC among the C3 variants and remains close to GPT\-4\.1 on AUPRC\-P and AUPRC\-N\. Random web\-sourced noise also preserves a meaningful calibration signal, outperforming or remaining competitive with the baseline methods on several metrics\. These results suggest that the effectiveness of C3 is not solely an artifact of using a strong frontier model to generate perturbations\. Instead, the core signal appears to come from measuring whether model outputs remain stable under semantically neutral contextual variation\. Computational efficiency\. It has to be admitted that using advanced model like GPT\-4\.1 can be costly and using smaller open\-weight models substantially reduces the cost of perturbation generation\. In our Qwen3\-8B setting, we generated perturbations for 200 SimpleQA questions on a single GPU NVIDIA A100\. We generated 30 samples per question, this corresponds to 6,000 accepted noise samples in 32\.36 minutes after we applied filtering for diversity\. This is the exact setting in the main experiment\. The generation pipeline achieved 3\.09 completed accepted samples per second\. These results suggest that perturbation generation with smaller and open sourced model is practically feasible and can substantially reduce dependence on expensive frontier\-model APIs\. Randomly sampled in\-the\-wild corpus noise provides an even cheaper alternative\. Unlike model\-generated perturbations, this source does not require inference or prompt\-specific generation: once a corpus such as RedPajama \[Computer, 2023\] is available, noise snippets of the desired length can be sampled almost instantly\. Because these snippets are sampled independently of the original question, they are unlikely to contain answer\-specific information, which gives them a degree of semantic neutrality\. However, this neutrality comes at the cost of weaker topic alignment: unlike GPT\-4\.1 or Qwen3\-8B perturbations, in\-the\-wild snippets are not explicitly generated to match the benchmark domain or capability being tested\. As shown in Table 8, corpus\-based noise still preserves a useful C3 signal, but it is generally less effective than model\-generated perturbations\. This pattern suggests that topic alignment improves perturbation quality and strengthens the resulting C3 signal, while not being strictly necessary for C3 to remain informative\. Web\-sourced perturbations can therefore be viewed as a practical low\-cost approximation when generation cost is a concern, rather than as a replacement for topic\-aligned model\-generated perturbations\. Appendix G Ablation Study on MMD Table 9: Calibration results on SimpleQA comparing the original MMD C3 score with a simpler cross\-comparison methods\. The cross\-comparison method replaces the MMD distance with the average pairwise similarity between generations from the original and perturbed prompts\. The noises are still sampled from GPT\-4\.1\. SimpleQA Method ECE ↓\\downarrow AUROC ↑\\uparrow AUPRC\-P ↑\\uparrow AUPRC\-N ↑\\uparrow Self\-consistency 0\.393 0\.792 0\.368 0\.924 Paraphrasing 0\.411 0\.773 0\.349 0\.906 Self\-report 0\.778 0\.490 0\.151 0\.846 C3 \(MMD\) 0\.166 \\cellcolorgray\!200\.823 0\.389 \\cellcolorgray\!200\.944 C3 \(Cross Comparison\) \\cellcolorgray\!200\.155 0\.813 \\cellcolorgray\!200\.402 0\.931 To test whether the effectiveness of C3 depends specifically on the MMD formulation, we replace the original MMD distributional distance with a simpler cross\-comparison score\. Given generations from the original prompt X=\{x1,…,xn\}X=\\\{x\_\{1\},\\ldots,x\_\{n\}\\\} and generations from the perturbed prompt Y=\{y1,…,ym\}Y=\\\{y\_\{1\},\\ldots,y\_\{m\}\\\}, the cross\-comparison variant directly measures the average pairwise agreement between the two sets: Scross\(X,Y\)=1nm∑i=1n∑j=1mk\(xi,yj\),S\_\{\\mathrm\{cross\}\}\(X,Y\)=\\frac\{1\}\{nm\}\\sum\_\{i=1\}^\{n\}\\sum\_\{j=1\}^\{m\}k\(x\_\{i\},y\_\{j\}\), where k\(xi,yj\)k\(x\_\{i\},y\_\{j\}\) is an indicator function for answer equivalence: k\(xi,yj\)=𝟏\[xi≡yj\]\.k\(x\_\{i\},y\_\{j\}\)=\\mathbf\{1\}\[x\_\{i\}\\equiv y\_\{j\}\]\. Here, xi≡yjx\_\{i\}\\equiv y\_\{j\} means that the two generations give the same answer\. For fixed\-format tasks such as SimpleQA, this can be implemented by exact answer matching or by an equivalence judge when surface forms differ but the answer is semantically the same\. As shown in Table 9, the cross\-comparison remains competitive with the original MMD C3 score\. It slightly improves ECE and AUPRC\-P on SimpleQA, while MMD achieves higher AUROC and AUPRC\-N\. More importantly, both C3 variants outperform self\-consistency, paraphrasing consistency, and self\-report on most calibration and ranking metrics\. This suggests that the C3 signal is not merely an artifact of the specific MMD distance\. Instead, the useful signal appears to come from the broader perturbation comparison: when a model’s generations remain equivalent across original and semantically perturbed contexts, its answers are more likely to be reliable; when the cross\-context generations diverge, the answer is more likely to be fragile or incorrect\. MMD remains our main choice because it provides a principled distributional distance that accounts for both within\-set and cross\-set similarities, but this ablation shows that a simpler indicator cross\-comparison variant can preserve much of the same credibility signal\. Appendix H Operationalization through MMD To quantify the distance between the generative distributions P\(Y\|x\)P\(Y\|x\) and P\(Y\|x′\)P\(Y\|x^\{\\prime\}\), we use Maximum Mean Discrepancy \(MMD\) \[Gretton et al\., 2012\], a non\-parametric kernel\-based statistic for comparing two empirical distributions\. Given finite sample sets 𝒴\\mathcal\{Y\} and 𝒴′\\mathcal\{Y\}^\{\\prime\} of size nn and mm, we compute the unbiased empirical estimate: MMD^2\(𝒴,𝒴′\)\\displaystyle\\widehat\{\\mathrm\{MMD\}\}^\{2\}\(\\mathcal\{Y\},\\mathcal\{Y\}^\{\\prime\}\) =1n\(n−1\)∑i≠jnk\(ϕ\(yi\),ϕ\(yj\)\)\+1m\(m−1\)∑i≠jmk\(ϕ\(yi′\),ϕ\(yj′\)\)\\displaystyle=\\frac\{1\}\{n\(n\-1\)\}\\sum\_\{i\\neq j\}^\{n\}k\(\\phi\(y\_\{i\}\),\\phi\(y\_\{j\}\)\)\+\\frac\{1\}\{m\(m\-1\)\}\\sum\_\{i\\neq j\}^\{m\}k\(\\phi\(y^\{\\prime\}\_\{i\}\),\\phi\(y^\{\\prime\}\_\{j\}\)\) \(1\) −2nm∑i=1n∑j=1mk\(ϕ\(yi\),ϕ\(yj′\)\)\.\\displaystyle\\qquad\-\\frac\{2\}\{nm\}\\sum\_\{i=1\}^\{n\}\\sum\_\{j=1\}^\{m\}k\(\\phi\(y\_\{i\}\),\\phi\(y^\{\\prime\}\_\{j\}\)\)\. Here, ϕ\\phi maps model generations into a task\-appropriate representation, and kk is a kernel function that compares represented outputs\. This estimator measures how much the empirical answer distribution under the original prompt differs from the answer distribution under the perturbed prompt\. In the main experiments, we normalize this distance into a C3 score in \[0,1\]\[0,1\], where larger values indicate smaller cross\-contextual shift and therefore greater answer stability\. Task\-adaptive feature maps and kernel selection The flexibility of choices of feature mapping function ϕ\(⋅\)\\phi\(\\cdot\) and kernel function k\(⋅,⋅\)k\(\\cdot,\\cdot\) provide the flexibility of assessing generation of various types\. The realization of ϕ\\phi and kk are adapted to the specific format of the model’s output yy of the underlying task\. For tasks with fixed output formats \(e\.g,\. keywords, numbers, or multiple\-choices\), we could choose ϕ\\phi to be a mapping to categories, resulting in an indicator kernel k\(y,y′\)=𝕀\(y=y′\)k\(y,y^\{\\prime\}\)=\\mathbb\{I\}\(y=y^\{\\prime\}\)\. This setting enables the C3 to function as an distance between empirical probability mass functions, measuring categorical inconsistencies\. Conversely, for open\-ended generation \(e\.g,\. coding, summarization, and essay writing\), ϕ\\phi could be a embedding process from an embedding model that maps the generation to high dimensional spaces h∈ℝDh\\in\\mathbb\{R\}^\{D\}\. In these high\-dimensional space, we employ a semantic kernel \(typically a linear dot\-product to measure the cos similarities\)\. Even more flexibly, LLMs as Judges frameworks can be used directly to compare the if the two answers are the same or not regardless of the output formats\. Normalization and range of C3 To ensure C3 is an interpretable proxy for credibility, we transform the raw MMD distance into a normalized range of \[0,1\]\[0,1\]\. In our framework, we assume a characteristic kernel kk that is bounded and normalized, satisfying k\(y,y\)=1k\(y,y\)=1 and 0≤k\(y,y′\)≤10\\leq k\(y,y^\{\\prime\}\)\\leq 1 \(e\.g\., an indicator kernel or cosine similarity\)\. Under these conditions, the squared MMD admits the theoretical upper bound MMD2\(𝒴,𝒴′\)≤2MMD^\{2\}\(\\mathcal\{Y\},\\mathcal\{Y\}^\{\\prime\}\)\\leq 2, which is attained when the two generative distributions are maximally separated \(i\.e\., their cross\-similarity approaches zero; equivalently, in fixed\-format settings, outputs from the two sets completely mismatch\)\. We define the final C3 by scaling this distance: C3\(x,x′\)=1−12MMD^2\(𝒴,𝒴′\)C3\(x,x^\{\\prime\}\)=1\-\\frac\{1\}\{2\}\\widehat\{MMD\}^\{2\}\(\\mathcal\{Y\},\\mathcal\{Y\}^\{\\prime\}\) \.A value of C3≈1C3\\approx 1 indicates high credibility, where the model’s output distribution remains invariant to semantic perturbations, suggesting consistent internal reasoning patterns and parametric memories\. Conversely, C3≈0C3\\approx 0 indicates low credibility, signaling that semantic variation in the input has caused a complete shift in the model’s generative behavior, which is a symptom of factual fragility or reasoning inconsistencies\. Appendix I Benchmark Details We provide additional details on the six benchmarks used in our evaluation\. SVAMP \[Patel et al\., 2021\] contains arithmetic word problems that require multi\-step numerical reasoning and typically produce single numeric answers\. MMLU High School Statistics \[Hendrycks et al\., 2021b, a\] is an exam\-style multiple\-choice benchmark that evaluates statistical concepts, conceptual understanding, and quantitative reasoning\. CommonsenseQA \[Talmor et al\., 2019\] evaluates commonsense and relational inference in a multiple\-choice format\. For factuality\-oriented evaluation, we use SimpleQA Verified \[Haas et al\., 2025\] and FActScore \[Min et al\., 2023\]\. SimpleQA Verified consists of short\-form fact\-retrieval questions with verified answers, while FActScore evaluates long\-form generations by decomposing outputs into atomic claims and checking whether each claim is supported by external evidence\. Finally, HumanEval \[Chen et al\., 2021\] evaluates code synthesis through unit tests\. Together, these benchmarks span reasoning, factual recall, long\-form factuality, and code generation, covering both constrained answer formats and open\-ended generations\. We include the prompts used for each benchmark below\. Appendix J Baseline Details Since C3 serves as a proxy for the credibility of model generations, we compare it with related notions of confidence, consistency, and factuality\. We include both vanilla black\-box confidence estimators and a non\-vanilla factuality checking tool\. Vanilla approaches\. We consider three vanilla black\-box baselines\. Self\-reported confidence \[Lin et al\., 2022\] asks the model to output an explicit numeric confidence score, with clearly defined upper and lower bounds, alongside its answer\. This baseline tests whether the model can verbalize its own uncertainty in a way that aligns with correctness or factual support\. The prompts used for self\-reported confidence are provided in Appendix D\.3\. Self\-consistency \[Wang et al\., 2023\] estimates confidence from repeated stochastic decoding under the same prompt\. We sample KK completions and measure agreement among the generated outputs, where higher agreement indicates greater confidence in the model’s answer\. Paraphrasing consistency \[Portillo Wightman et al\., 2023\] measures whether model outputs remain stable under meaning\-preserving prompt\-level paraphrases\. We generate KK paraphrased variants of the original question and compare the resulting answers using the same consistency framework as self\-consistency\. For sampling\-based methods, including self\-consistency and paraphrasing consistency, we compute agreement over a set of KK generated outputs 𝒴=\{y1,y2,…,yK\}\\mathcal\{Y\}=\\\{y\_\{1\},y\_\{2\},\\dots,y\_\{K\}\\\}: 𝒞\(𝒴\)=1K\(K−1\)∑i=1K∑j≠iKk\(yi,yj\),\\mathcal\{C\}\(\\mathcal\{Y\}\)=\\frac\{1\}\{K\(K\-1\)\}\\sum\_\{i=1\}^\{K\}\\sum\_\{j\\neq i\}^\{K\}k\(y\_\{i\},y\_\{j\}\), \(2\) where k\(yi,yj\)k\(y\_\{i\},y\_\{j\}\) is a similarity function\. For fixed\-format tasks, we use an indicator function 𝟏\[yi=yj\]\\mathbf\{1\}\[y\_\{i\}=y\_\{j\}\]\. For open\-ended generation, we use a semantic similarity metric\. Non\-vanilla approach\. We also compare C3 with FActScore \[Min et al\., 2023\], a factuality checking tool for long\-form generation\. FActScore decomposes each output into atomic facts, retrieves evidence from an external knowledge source, and checks whether the evidence supports each claim\. In the benchmark setting, Wikipedia is used as the external knowledge source\. The final FActScore is computed as the fraction of supported atomic facts\.`Similar Articles
When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals
This paper audits whether self-consistency and cross-model agreement are reliable indicators of correctness in LLMs, finding that agreement is a weak, regime-dependent proxy and that frontier models exhibit overconfidence.
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy
This paper investigates how LLMs' answers change under meaning-preserving paraphrases, finding that instance-level behavior is unstable (flip rates >23%) and that single-prompt accuracy masks substantial inconsistency, while a self-paraphrasing strategy can partially recover latent knowledge.
Enhancing Assessment of Self-Consistency in LLM Explanations using Perturbation Strength
This paper introduces an LLM-as-a-judge method to measure perturbation strength for assessing self-consistency in LLM explanations, showing that input perturbations generally affect LLMs more strongly than CoT perturbations.
Two Axes of LLM Abstention: Answer Correctness and Question Answerability
This paper investigates the two axes of LLM abstention: answer correctness and question answerability. It shows that a single confidence threshold conflates these two failure modes, and proposes a three-class selective acceptance framework with separate budgets. Experiments across five instruction-tuned models reveal that answerability is internally legible but poorly captured by output confidence or self-assessments.
When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings
The paper evaluates local open-weight LLM judges against human ratings, finding high self-consistency but limited agreement with human judgments, highlighting the need for dual assessment.