Diagnosing Correctness Probes under Self-Judgement Confounding
Summary
This paper investigates whether neural network probes that predict correctness of language model outputs actually capture objective correctness or the model's own self-judgement, using conflict cases where the two disagree. The authors find that transferable directions predominantly preserve self-judgement polarity, challenging the interpretation of correctness readouts.
View Cached Full Text
Cached at: 07/21/26, 06:43 AM
# Diagnosing Correctness Probes under Self-Judgement Confounding
Source: [https://arxiv.org/html/2607.16799](https://arxiv.org/html/2607.16799)
###### Abstract
Hidden\-state readouts can predict whether language\-model outputs are correct, but objective correctness \(OC\) usually agrees with the model’s own self\-judgement \(SJ\), leaving the decoded signal semantically ambiguous\. We construct conflict cases in which OC and SJ predict opposite readout orderings\. On high\-confidence disagreements, conventional correctness\-labelled contrasts often rank incorrect/self\-endorsed responses above correct/self\-rejected responses, following SJ rather than OC\. We estimate factorial SJ\- and OC\-associated directions and evaluate their polarity across mathematical reasoning and factual recall\. Across four instruction\-tuned models up to 14B parameters, the SJ\-associated direction transfers above chance in both cross\-domain directions for every model, whereas the OC\-associated direction has a below\-chance point estimate for the expected OC ordering in every corresponding condition\. This transfer asymmetry develops across middle\-to\-late layers, persists under answer\-likelihood, sequence\-length, and null\-direction controls, and extends to MMLU and binary TruthfulQA without target\-domain direction fitting\. Across the studied models and diagnostic subsets, the most reliably transferable component preserves SJ\-associated polarity\. Transferability alone therefore does not establish objective\-correctness semantics\.
## Introduction
Large language models \(LLMs\) can produce fluent answers that are factually wrong, motivating efforts to read factual reliability directly from their internal representations\(Linet al\.[2022b](https://arxiv.org/html/2607.16799#bib.bib1); Huanget al\.[2025](https://arxiv.org/html/2607.16799#bib.bib23)\)\. Early work showed that latent knowledge and statement truthfulness are often decodable from hidden activations\(Burnset al\.[2023](https://arxiv.org/html/2607.16799#bib.bib10); Azaria and Mitchell[2023](https://arxiv.org/html/2607.16799#bib.bib11); Liet al\.[2023](https://arxiv.org/html/2607.16799#bib.bib12)\), while subsequent analyses reported a simple linear geometry for factual truth across datasets\(Marks and Tegmark[2024](https://arxiv.org/html/2607.16799#bib.bib13)\)\. Some of these readouts transfer across logical transformations, question\-answering tasks, and external knowledge settings\(Baoet al\.[2025](https://arxiv.org/html/2607.16799#bib.bib38)\), suggesting that truth\-related information may be encoded in a reusable form\. Yet reusability alone does not determine which variable a readout represents\.111Code and processed data are available athttps://github\.com/Yilong\-Lu/Diagnosing˙Correctness\.
The central issue is therefore not only whether a truth\-related direction transfers, but which variable preserves its polarity under transfer\. Such directions vary with layer, task type, task complexity, and instructions\(Azizianet al\.[2025](https://arxiv.org/html/2607.16799#bib.bib34); Pouliset al\.[2026](https://arxiv.org/html/2607.16799#bib.bib39)\); recent evidence places broadly shared and domain\-specific directions on a continuum\(Yinget al\.[2026](https://arxiv.org/html/2607.16799#bib.bib37)\)\. Separability can degrade under distribution shift, and some factual errors share internal geometry with successful knowledge recall\(Halleret al\.[2025](https://arxiv.org/html/2607.16799#bib.bib29); Cheanget al\.[2026](https://arxiv.org/html/2607.16799#bib.bib32)\)\. These findings leave the semantic interpretation of a transferable correctness readout unresolved\.
Figure 1:Conflict\-based semantic validation of correctness probes\. \(a\) Objective correctness \(OC\) and self\-judgement \(SJ\) define four response types\. On B/C conflicts, OC predictsB\>CB\>C, whereas SJ predictsC\>BC\>B\. \(b\) Hidden states are extracted at the final answer token, before the separate judgement prompt used to define SJ\. \(c\) The conventional contrastWmix=μA−μDW\_\{\\mathrm\{mix\}\}=\\mu\_\{A\}\-\\mu\_\{D\}decomposes into the factorial SJ\- and OC\-associated directionsWmetaW\_\{\\mathrm\{meta\}\}andWtruthW\_\{\\mathrm\{truth\}\}\. \(d\) Source\-domain directions are frozen and evaluated on B/C conflicts in another domain, including zero\-target\-fitting transfer to MMLU and binary TruthfulQA\. Successful component transfer corresponds toAUC\(Wmeta→SJ\)\>0\.5\\mathrm\{AUC\}\(W\_\{\\mathrm\{meta\}\}\\\!\\rightarrow\\\!\\mathrm\{SJ\}\)\>0\.5andAUC\(Wtruth→OC\)\>0\.5\\mathrm\{AUC\}\(W\_\{\\mathrm\{truth\}\}\\\!\\rightarrow\\\!\\mathrm\{OC\}\)\>0\.5\.One candidate source of this ambiguity is the model’s own evaluative stance\. Recent work reports a shared representational dimension spanning subjective evaluation and assent to factual claims\(Luet al\.[2025](https://arxiv.org/html/2607.16799#bib.bib24)\)\. A broader literature studies response evaluation through behavioural confidence, failure prediction, and internal activation analyses\(Kadavathet al\.[2022](https://arxiv.org/html/2607.16799#bib.bib15); Wanget al\.[2025a](https://arxiv.org/html/2607.16799#bib.bib25); Liet al\.[2025](https://arxiv.org/html/2607.16799#bib.bib31); Kumaranet al\.[2026a](https://arxiv.org/html/2607.16799#bib.bib28)\); aggregating self\-evaluation\-related states can also improve uncertainty calibration\(Xiaoet al\.[2026](https://arxiv.org/html/2607.16799#bib.bib44)\)\. At the feature level, correctness and output uncertainty can depend on partly distinct representations\(Patelet al\.[2026](https://arxiv.org/html/2607.16799#bib.bib36)\)\. These findings motivate a sharper diagnostic question for truth probing: when a correctness\-labelled readout transfers, does it preserve the relation between an answer and external ground truth, or the model’s own evaluation of that answer?
To make this distinction operational, we separate two response\-level variables\. Objective correctness \(OC\) records whether an answer is externally scored as correct\. Self\-judgement \(SJ\) records whether the model subsequently judges that answer to be correct; this correctness\-directed, second\-order assessment serves as our operational measure of metacognitive judgement\(Steyvers and Peters[2026](https://arxiv.org/html/2607.16799#bib.bib33)\)\. Because OC and SJ usually agree, a conventional correct\-versus\-incorrect contrast can conflate their activation correlates\. Their competing interpretations become testable when the two variables disagree\.
The empirical design crosses OC with SJ to form four response types\. The critical comparison is between objectively correct answers that the model rejects \(B\) and objectively wrong answers that it endorses \(C\)\. The two candidate signals predict opposite orderings: an OC\-associated direction should rank B above C, whereas an SJ\-associated direction should rank C above B\. Even an OC\-only mass\-mean control fitted without SJ labels follows SJ in all eight cross\-domain comparisons\. Factorial OC\- and SJ\-associated directions then assess which polarity transfers\. Across four instruction\-tuned LLMs, the SJ\-associatedWmetaW\_\{\\mathrm\{meta\}\}direction has the expected polarity in every cross\-domain condition spanning mathematical reasoning, factual recall, MMLU, and binary TruthfulQA, whereas the OC\-associatedWtruthW\_\{\\mathrm\{truth\}\}direction does not\. Self\-judgement is therefore a substantial source of semantic ambiguity in transferable correctness readouts\.
## Related Work
#### Truth\-direction transfer and semantic validation\.
Latent knowledge and factual truthfulness are often decodable from LLM activations using unsupervised, linear, or nonlinear readouts\(Burnset al\.[2023](https://arxiv.org/html/2607.16799#bib.bib10); Azaria and Mitchell[2023](https://arxiv.org/html/2607.16799#bib.bib11); Liet al\.[2023](https://arxiv.org/html/2607.16799#bib.bib12); Marks and Tegmark[2024](https://arxiv.org/html/2607.16799#bib.bib13)\)\. Some truth directions generalize across logical transformations and QA settings\(Baoet al\.[2025](https://arxiv.org/html/2607.16799#bib.bib38)\); others depend strongly on layer, task family, instructions, or truth type\(Azizianet al\.[2025](https://arxiv.org/html/2607.16799#bib.bib34); Pouliset al\.[2026](https://arxiv.org/html/2607.16799#bib.bib39); Yinget al\.[2026](https://arxiv.org/html/2607.16799#bib.bib37); Schoutenet al\.[2025](https://arxiv.org/html/2607.16799#bib.bib47)\)\. Apparent truth signals can also reflect superficial task features or knowledge recall\(Halleret al\.[2025](https://arxiv.org/html/2607.16799#bib.bib29); Cheanget al\.[2026](https://arxiv.org/html/2607.16799#bib.bib32)\)\. Their generalization to realistic model\-generated responses remains challenging\(Servedioet al\.[2025](https://arxiv.org/html/2607.16799#bib.bib46)\)\. Query–probe disagreement can reflect calibration or heterogeneous errors\(Liuet al\.[2023](https://arxiv.org/html/2607.16799#bib.bib49)\), while inter\-model disagreement can reveal domain\-specific privileged correctness information\(Ashuachet al\.[2026](https://arxiv.org/html/2607.16799#bib.bib48)\)\. Predictive decodability therefore does not identify the variable a readout exploits\(Hewitt and Liang[2019](https://arxiv.org/html/2607.16799#bib.bib19); Belinkov[2022](https://arxiv.org/html/2607.16799#bib.bib20)\); instructions and interaction can further redirect truth\-related representations\(Longet al\.[2025](https://arxiv.org/html/2607.16799#bib.bib42); Wanget al\.[2026](https://arxiv.org/html/2607.16799#bib.bib43)\)\. OC–SJ conflicts make semantic polarity directly testable because the two variables prescribe opposite rankings\.
#### Self\-evaluation, uncertainty, and introspection\.
LLM self\-evaluation has been studied through P\(True\), verbal confidence, uncertainty elicitation, unknown detection, and failure prediction\(Kadavathet al\.[2022](https://arxiv.org/html/2607.16799#bib.bib15); Linet al\.[2022a](https://arxiv.org/html/2607.16799#bib.bib16); Tianet al\.[2023](https://arxiv.org/html/2607.16799#bib.bib17); Xionget al\.[2024](https://arxiv.org/html/2607.16799#bib.bib18); Yinet al\.[2023](https://arxiv.org/html/2607.16799#bib.bib26); Wanget al\.[2025a](https://arxiv.org/html/2607.16799#bib.bib25)\)\. Hidden\-state trajectories have also been used to predict response correctness without an explicit verbal report\(Wanget al\.[2025b](https://arxiv.org/html/2607.16799#bib.bib45)\)\. Correctness and output uncertainty can be functionally dissociated at the sparse\-feature level\(Patelet al\.[2026](https://arxiv.org/html/2607.16799#bib.bib36)\), while confidence\-related information around answer completion can exceed token log\-probability\(Kumaranet al\.[2026a](https://arxiv.org/html/2607.16799#bib.bib28),[b](https://arxiv.org/html/2607.16799#bib.bib27)\)\. Stronger claims of privileged self\-access face different tests: prompted reports can fail to recover a model’s own linguistic knowledge, and apparent introspective performance can be reproduced from surface cues\(Songet al\.[2025](https://arxiv.org/html/2607.16799#bib.bib40); Singhet al\.[2026](https://arxiv.org/html/2607.16799#bib.bib41)\)\. Existing studies have separately decoded correctness\-related and response\-evaluation signals; we ask which correlated variable preserves its semantic polarity when they conflict\. SJ serves here as a correctness\-directed behavioural label whose answer\-token correlate is evaluated on those conflict cases\.
## Method
#### Setup and labels\.
For an answered itemiiat layerℓ\\ell, letxi,ℓx\_\{i,\\ell\}be the hidden state at the answer\-specific final token\. OC indicates objective correctness, and SJ indicates the model’s binary judgement of whether its own answer is correct\. The judgement turn asks, “Do you believe the answer above is correct? Answer only with Yes or No\.” Next\-token Yes/No scores define the binary correctness probabilitypjudgep\_\{\\mathrm\{judge\}\}; this separate turn is excluded from activation extraction\. The factorial contrasts below quantify the marginal activation associations of OC and SJ across the observed response types\. Confidence is used only to exclude weak judgements: the main analyses retainpjudge≤0\.3p\_\{\\mathrm\{judge\}\}\\leq 0\.3orpjudge≥0\.7p\_\{\\mathrm\{judge\}\}\\geq 0\.7and define SJ by the retained side\. This definesA=\(1,1\)A=\(1,1\),B=\(1,0\)B=\(1,0\),C=\(0,1\)C=\(0,1\), andD=\(0,0\)D=\(0,0\)over \(OC,SJ\);B∪CB\\cup Cis the conflict set where correctness and self\-judgement disagree\.
#### Factorial directions\.
LetμA,μB,μC,μD\\mu\_\{A\},\\mu\_\{B\},\\mu\_\{C\},\\mu\_\{D\}be quadrant mean activations after source/train\-only centering\. We compare an aligned mixed contrast with two factorial main effects:
Wmix=μA−μD,Wmeta=\[\(μA−μB\)\+\(μC−μD\)\]/2,Wtruth=\[\(μA−μC\)\+\(μB−μD\)\]/2,\\begin\{array\}\[\]\{rcl\}W\_\{\\mathrm\{mix\}\}&=&\\mu\_\{A\}\-\\mu\_\{D\},\\\\ W\_\{\\mathrm\{meta\}\}&=&\[\(\\mu\_\{A\}\-\\mu\_\{B\}\)\+\(\\mu\_\{C\}\-\\mu\_\{D\}\)\]/2,\\\\ W\_\{\\mathrm\{truth\}\}&=&\[\(\\mu\_\{A\}\-\\mu\_\{C\}\)\+\(\\mu\_\{B\}\-\\mu\_\{D\}\)\]/2,\\end\{array\}whereWmetaW\_\{\\mathrm\{meta\}\}andWtruthW\_\{\\mathrm\{truth\}\}are operational labels for the SJ\-associated and OC\-associated factorial contrasts, respectively\. The interpretation follows from a minimal associative representation model with signed labelsoi=2OCi−1o\_\{i\}=2\\mathrm\{OC\}\_\{i\}\-1andsi=2SJi−1s\_\{i\}=2\\mathrm\{SJ\}\_\{i\}\-1:
xi=αtoivtOC\+βtsivtSJ\+ηtoisivtINT\+ϵi\.x\_\{i\}=\\alpha\_\{t\}o\_\{i\}v^\{\\mathrm\{OC\}\}\_\{t\}\+\\beta\_\{t\}s\_\{i\}v^\{\\mathrm\{SJ\}\}\_\{t\}\+\\eta\_\{t\}o\_\{i\}s\_\{i\}v^\{\\mathrm\{INT\}\}\_\{t\}\+\\epsilon\_\{i\}\.HerevtOCv^\{\\mathrm\{OC\}\}\_\{t\}andvtSJv^\{\\mathrm\{SJ\}\}\_\{t\}are task\-specific latent directions associated with correctness and self\-judgement,vtINTv^\{\\mathrm\{INT\}\}\_\{t\}captures their interaction, andαt,βt,ηt\\alpha\_\{t\},\\beta\_\{t\},\\eta\_\{t\}are task\-dependent effect magnitudes\. Under this model, substituting the four cell means givesWmeta=2βtvtSJW\_\{\\mathrm\{meta\}\}=2\\beta\_\{t\}v^\{\\mathrm\{SJ\}\}\_\{t\},Wtruth=2αtvtOCW\_\{\\mathrm\{truth\}\}=2\\alpha\_\{t\}v^\{\\mathrm\{OC\}\}\_\{t\}, andWmix=Wtruth\+WmetaW\_\{\\mathrm\{mix\}\}=W\_\{\\mathrm\{truth\}\}\+W\_\{\\mathrm\{meta\}\}; the interaction term cancels in the main effects\. On B/C conflicts, OC predictsB\>CB\>Cwhereas SJ predictsC\>BC\>B\. The corresponding mean\-score difference under the mixed direction is
Δmix=E\[⟨x,Wmix⟩\|C\]−E\[⟨x,Wmix⟩\|B\]=4βt2‖vtSJ‖2−4αt2‖vtOC‖2\.\\begin\{array\}\[\]\{rcl\}\\Delta\_\{\\mathrm\{mix\}\}&=&\\mathrm\{E\}\[\\langle x,W\_\{\\mathrm\{mix\}\}\\rangle\|C\]\-\\mathrm\{E\}\[\\langle x,W\_\{\\mathrm\{mix\}\}\\rangle\|B\]\\\\ &=&4\\beta\_\{t\}^\{2\}\\\|v^\{\\mathrm\{SJ\}\}\_\{t\}\\\|^\{2\}\-4\\alpha\_\{t\}^\{2\}\\\|v^\{\\mathrm\{OC\}\}\_\{t\}\\\|^\{2\}\.\\end\{array\}Thus, when these responses are scored byWmixW\_\{\\mathrm\{mix\}\}, an AUC above 0\.5 for ranking C over B indicates that the SJ\-associated component dominates the conflict\-set ordering\. Cross\-domain transfer then tests whether a source\-estimated SJ\-associated direction preserves SJ polarity in the target domain more reliably than the corresponding OC\-associated direction preserves OC polarity\.
#### Strict pairs\.
For each model–domain combination, we sample eight stochastic free\-response answers per question and retain questions with at least one correct and one incorrect usable response\. We randomly select one of each and ask the same model to judge whether its response is correct\. Under the strict thresholdτ=0\.7\\tau=0\.7, a pair is retained only when both self\-judgements satisfy the symmetric high\-confidence criterion\. These analyses therefore characterize a high\-confidence diagnostic subset of questions exhibiting response variability; they do not estimate the prevalence of B/C cases in unconstrained model outputs\. Across model–domain cells, the resulting sets contain 362–2,775 retained questions and 249–2,146 B/C conflict responses\. Full attrition statistics and quadrant counts are reported in the supplement\.
#### Activation site\.
Activations are extracted from the answer prompt together with the corresponding assistant answer: the final token of the numeric answer span for Math, the final token of the actor\-name answer for Movies, and the forced answer\-letter token for OOD\. At layerℓ\\ell,xi,ℓx\_\{i,\\ell\}denotes the hidden state at this token after transformer blockℓ\\ell, that is, the post\-block residual\-stream state\. The judgement prompt is used only to computepjudgep\_\{\\mathrm\{judge\}\}; abbreviated prompt cores and full templates are provided in the supplement\.
Figure 2:Traditional mixed direction on OC–SJ conflicts\. Curves show the AUC for ranking C \(wrong/self\-endorsed\) above B \(correct/self\-rejected\) across layers\. Values above 0\.5 indicate thatWmixW\_\{\\mathrm\{mix\}\}ranks wrong/self\-endorsed samples above correct/self\-rejected samples\. Directions are derived on Math \(teal\) or Movies \(amber\) datasets; solid and dashed lines denote cross\-domain and within\-domain evaluation, respectively\. Shaded bands denote bootstrap 95% CIs conditional on the fitted source direction; all conflict responses from a question are resampled together\.Figure 3:Cross\-domain component transfer across layers\. Directions fitted in one domain are evaluated on B/C conflicts in the other; rows show transfer directions and columns show models\. Blue curves showAUC\(Wmeta→SJ\)\\mathrm\{AUC\}\(W\_\{\\mathrm\{meta\}\}\\\!\\rightarrow\\\!\\mathrm\{SJ\}\), and brown\-orange curves showAUC\(Wtruth→OC\)\\mathrm\{AUC\}\(W\_\{\\mathrm\{truth\}\}\\\!\\rightarrow\\\!\\mathrm\{OC\}\)\. The dotted horizontal line marks chance, and the pale vertical band marks the fixed normalized\-depth window\[0\.40,0\.80\]\[0\.40,0\.80\]summarized in Table[1](https://arxiv.org/html/2607.16799#Sx4.T1)\. Shaded bands are 95% target\-question cluster\-bootstrap CIs conditional on the fitted source direction; in\-panelnndenotes the number of target B/C responses contributing to each layer\-wise AUC\.
## Experimental Protocol
#### Datasets and models\.
For mathematical reasoning, we curate a Math dataset from an initial pool of 16,700 questions collected from public sources, including GSM8K and MATH\-style items\(Cobbeet al\.[2021](https://arxiv.org/html/2607.16799#bib.bib4); Hendryckset al\.[2021b](https://arxiv.org/html/2607.16799#bib.bib7)\)\. A Qwen2\.5\-1\.5B\-Instruct pilot removes questions solved on all eight attempts and the longest approximately 10% of responses, yielding 5,549 questions\. This filtering preserves response variability while limiting extreme generation lengths\. Movies is a person\-name factual\-recall domain using 17,856 actor\-question records from the Movies QA dataset described byOrgadet al\.\([2025](https://arxiv.org/html/2607.16799#bib.bib30)\)\. OC is determined by domain\-specific numeric\-answer parsing for Math and normalized actor\-name matching for Movies\. OOD evaluations use 4\-choice MMLU\(Hendryckset al\.[2021a](https://arxiv.org/html/2607.16799#bib.bib3)\)and binary TruthfulQA\(Linet al\.[2022b](https://arxiv.org/html/2607.16799#bib.bib1)\)\. MMLU contributes four options per question and binary TruthfulQA two, with dataset answer keys defining OC\. Candidate letters are inserted as assistant responses and judged with the same SJ query; neither target dataset contributes to direction fitting\. We evaluate four instruction models with at most 14B parameters: Qwen2\.5\-7B\-Instruct, Llama\-3\.1\-8B\-Instruct, Qwen2\.5\-14B\-Instruct, and OLMo\-3\-7B\-Instruct\(Qwen[2025](https://arxiv.org/html/2607.16799#bib.bib21); Grattafiori and others[2024](https://arxiv.org/html/2607.16799#bib.bib22); Team Olmo[2026](https://arxiv.org/html/2607.16799#bib.bib35)\)\. Full details appear in the supplement\.
#### Implementation\.
For each model, we performed forward passes in bfloat16 and extracted hidden states at the final answer token\. All subsequent analyses were performed offline using the saved activations\. We set the temperature to 1\.0 for both main\-model pass@8 generation and the one\-token Yes/No judgement\. Maximum generation lengths are 384 tokens for Math and 256 for Movies\. The resulting judgement score defines SJ\.
#### Experimental comparisons\.
Evaluation proceeds from a conflict\-set diagnostic to within\-domain validation and cross\-domain transfer\. Experiment 1 tests whether the B/C ordering induced by the conventional mixed directionWmixW\_\{\\mathrm\{mix\}\}follows OC or SJ when the two variables disagree\. We compute the binary AUC with C \(wrong/self\-endorsed\) as the positive class and B \(correct/self\-rejected\) as negative class\. Values above 0\.5 indicate that C responses receive higher scores, opposite to the ordering predicted by an OC\-specific direction\. The OC\-only mass\-mean control forms a canonical unwhitened class\-mean contrast from all retained source rows,WOC−only=E\[x∣OC=1\]−E\[x∣OC=0\]W\_\{\\mathrm\{OC\-only\}\}=\\mathrm\{E\}\[x\\mid\\mathrm\{OC\}=1\]\-\\mathrm\{E\}\[x\\mid\\mathrm\{OC\}=0\], without using SJ labels\(Marks and Tegmark[2024](https://arxiv.org/html/2607.16799#bib.bib13); Zouet al\.[2025](https://arxiv.org/html/2607.16799#bib.bib14)\)\. Because each retained question contributes one correct and one incorrect response, this contrast is equivalent to the mean within\-question correct\-minus\-incorrect activation difference\.
Experiment 2A tests whether the factorial directions preserve their expected SJ and OC rankings on held\-out B/C conflicts within domain\. We use five question\-grouped folds, fitting centering and both directions on training questions and scoring their held\-out responses\. Experiment 2B assesses bidirectional cross\-domain transfer by fitting the directions on all retained rows in one source domain, freezing them, and evaluating them on B/C conflicts in the other domain:Wmeta→SJW\_\{\\mathrm\{meta\}\}\\\!\\rightarrow\\\!\\mathrm\{SJ\}withWtruth→OCW\_\{\\mathrm\{truth\}\}\\\!\\rightarrow\\\!\\mathrm\{OC\}\.
#### Primary estimands and inference\.
We quantify component transfer using two AUCs\. Because OC and SJ are complementary labels on B/C conflicts, we defineΔCB=AUC\(Wmeta→SJ\)−AUC\(Wtruth→OC\)\\Delta\_\{\\mathrm\{CB\}\}=\\mathrm\{AUC\}\(W\_\{\\mathrm\{meta\}\}\\rightarrow\\mathrm\{SJ\}\)\-\\mathrm\{AUC\}\(W\_\{\\mathrm\{truth\}\}\\rightarrow\\mathrm\{OC\}\)to summarize the relative strength of the two component transfers\. Each component AUC is evaluated separately against chance \(0\.50\.5\), testing whetherWmetaW\_\{\\mathrm\{meta\}\}ranks SJ andWtruthW\_\{\\mathrm\{truth\}\}ranks OC on the same target responses\. All\-layer curves show layer\-wise structure, with shaded bands from a target\-question bootstrap that keeps all conflict rows from a question together\. All reported intervals are two\-sided 95% confidence intervals \(CIs\)\. All\-layer and OOD intervals condition on the fitted source direction\. Primary Exp2B fixed\-window inference independently resamples source and target questions, refits both factorial directions in every replicate, and averages layer\-wise AUCs over an architecture\-normalized middle\-to\-late window, selecting layers withℓ/ℓmax∈\[0\.40,0\.80\]\\ell/\\ell\_\{\\max\}\\in\[0\.40,0\.80\]\. This fixed window excludes architecture endpoints and avoids selecting a model\-specific peak layer\. Fixed\-window bootstrap estimates provide inference, whereas the full profiles characterize depth\-dependent structure\.
#### Question\-level controls\.
The source\-side sensitivity analysis takes the within\-question correct\-minus\-incorrect activation difference, eliminating additive question intercepts before direction fitting\. At each layer, multivariate least squares jointly estimates OC, SJ, and interaction coefficient vectors from these paired differences; the resultingb^SJ\\widehat\{b\}\_\{\\mathrm\{SJ\}\}andb^OC\\widehat\{b\}\_\{\\mathrm\{OC\}\}replaceWmetaW\_\{\\mathrm\{meta\}\}andWtruthW\_\{\\mathrm\{truth\}\}in the unchanged fixed\-window B/C evaluation\. Independent source\- and target\-question bootstrap resampling refits the adjusted directions\.
The complementary target\-side analysis keeps the pooled source directions fixed and uses every retained target row\. Component scores are standardized within layer and target domain, averaged over the fixed window, and regressed on signed OC, SJ, and their interaction with target\-question fixed effects\. Its cross\-component contrast isΔFE=βSJ\(Wmeta\)−βOC\(Wtruth\)\\Delta\_\{\\mathrm\{FE\}\}=\\beta^\{\(W\_\{\\mathrm\{meta\}\}\)\}\_\{\\mathrm\{SJ\}\}\-\\beta^\{\(W\_\{\\mathrm\{truth\}\}\)\}\_\{\\mathrm\{OC\}\}\. Target\-question bootstrap replicates repeat score standardization and model fitting\. This all\-cell analysis restores the aligned cells while absorbing target\-question\-specific differences\. Within\-component coefficient differences additionally test whetherWmetaW\_\{\\mathrm\{meta\}\}loads more strongly on SJ than OC, andWtruthW\_\{\\mathrm\{truth\}\}on OC than SJ\.
#### Main\-domain controls\.
Source\-label\-shuffle nulls test whether transfer depends on the fitted source label geometry, while norm\-matched random directions test whether arbitrary activation directions yield the same pattern\. B/C matching on rendered\-sequence token count tests prompt\-plus\-response length composition, and symmetric confidence\-threshold sensitivity tests dependence on judgement certainty\. For Math and Movies, we also score the generated assistant response tokens\. At each layer, projection scores are residualized against mean response log\-probability and response\-token count before computing component AUCs; the fixed\-window statistic then averages those layer\-wise AUCs, matching the primary estimand\. Each bootstrap replicate refits the layer\-wise nuisance regressions\.
#### OOD controls\.
Forced\-choice analyses add answer\-letter fixed effects to separate the projection signal from option identity\. A target\-side elicitation control replaces Yes/No with reversed X/Y mappings to test sensitivity to the judgement\-to\-label assignment\. Their correctness\-oriented log odds are averaged before applying the same symmetric confidence threshold; the source directions remain fixed\.
Table 1:Exp2B fixed\-window component transfer and relative asymmetry\. Estimates are averaged over normalized layer depths\[0\.40,0\.80\]\[0\.40,0\.80\]using 1,000 independent bootstrap resamples of source and target questions\.
## Results
#### Exp1: correctness\-labelled contrasts reverse the OC ordering on conflicts\.
IfWmixW\_\{\\mathrm\{mix\}\}were a clean OC readout, it should prefer B samples \(correct/self\-rejected\) over C samples \(wrong/self\-endorsed\), despite the model’s opposite judgement\. Equivalently, the pairwise AUC that treats C as the positive class,Pr\[s\(C\)\>s\(B\)\]\\Pr\[s\(C\)\>s\(B\)\], should fall below 0\.5\. Figure[2](https://arxiv.org/html/2607.16799#Sx3.F2)shows the opposite pattern: after the early layers, the curves are usually above 0\.5, indicating thatWmixW\_\{\\mathrm\{mix\}\}often assigns higher scores to wrong/self\-endorsed answers than to correct/self\-rejected answers\.
The OC\-only mass\-mean control strengthens this diagnosis without using SJ labels to fit the source direction\. Its fixed\-window AUC for ranking C above B exceeds 0\.5 in all eight model\-by\-direction cross\-domain evaluations, with 95% CIs above chance and estimates ranging from 0\.567 to 0\.874\. Correctness\-labelled mean contrasts can therefore follow the model’s judgement on cases where judgement and correctness disagree, even when SJ is absent from probe fitting\.
#### Exp2A: the SJ\-associated contrast is stable within domain\.
Question\-grouped folds keep the paired responses to one question in the same partition\. In held\-out questions,WmetaW\_\{\\mathrm\{meta\}\}predicts SJ above chance in all eight model–domain conditions \(AUC 0\.649–0\.915\), with every 95% CI above 0\.5\. By contrast,WtruthW\_\{\\mathrm\{truth\}\}predicts OC above chance only for Llama\-3\.1\-8B Math; its other seven point estimates are below 0\.5 \(0\.219–0\.483\)\. Consequently,ΔCB\\Delta\_\{\\mathrm\{CB\}\}is positive in seven of eight conditions \(0\.173–0\.696\), with all seven 95% CIs above zero\. The sole negative contrast is Llama\-3\.1\-8B Math, where both components discriminate above chance andWtruthW\_\{\\mathrm\{truth\}\}is stronger \(−0\.042\-0\.042\[−0\.097\-0\.097, 0\.012\]\)\. Exp2A therefore supports the expected in\-domain SJ ordering but does not support a generally stable OC\-associated readout; complete layer\-wise results are reported in the supplement\.
Figure 4:OOD transfer without target\-domain direction fitting\. Rows show \(a\) Math→\\rightarrowMMLU, \(b\) Movies→\\rightarrowMMLU, \(c\) Math→\\rightarrowbinary TruthfulQA, and \(d\) Movies→\\rightarrowbinary TruthfulQA; columns denote models\. Blue curves showAUC\(Wmeta→SJ\)\\mathrm\{AUC\}\(W\_\{\\mathrm\{meta\}\}\\\!\\rightarrow\\\!\\mathrm\{SJ\}\), and orange curves showAUC\(Wtruth→OC\)\\mathrm\{AUC\}\(W\_\{\\mathrm\{truth\}\}\\\!\\rightarrow\\\!\\mathrm\{OC\}\)on OOD B/C conflicts\. The dashed line marks chance\. Shaded bands show 95% CIs from a target\-question cluster bootstrap, conditional on the fitted source direction\.
#### Exp2B: only the SJ\-associated direction preserves cross\-domain polarity\.
Directions fitted on Math are evaluated on Movies, and directions fitted on Movies are evaluated on Math\. The layer\-wise profiles in Figure[3](https://arxiv.org/html/2607.16799#Sx3.F3)reveal a common progression\. The component curves are near chance in early layers and then separate across middle\-to\-late layers: the SJ\-associated curve rises above chance, whereas the OC\-associated curve remains near chance or reverses polarity\. The separation is sustained through the fixed analysis window rather than driven by a single isolated layer\. Across layers and transfer conditions,ΔCB\>0\\Delta\_\{\\mathrm\{CB\}\}\>0in 255/280 \(91%\) cross\-domain model–transfer–layer rows\.
The fixed\-window estimates confirm this visual pattern\. All eightWmeta→SJW\_\{\\mathrm\{meta\}\}\\rightarrow\\mathrm\{SJ\}AUCs exceed 0\.5 with joint bootstrap 95% CIs above chance \(0\.572–0\.928\), whereas all eightWtruth→OCW\_\{\\mathrm\{truth\}\}\\rightarrow\\mathrm\{OC\}point estimates are below chance \(0\.167–0\.441\), with seven joint 95% CIs excluding 0\.5\. The relativeΔCB\\Delta\_\{\\mathrm\{CB\}\}is positive for every model and transfer direction, ranging from 0\.138 to 0\.761, with all eight joint source–target bootstrap 95% CIs above zero \(Table[1](https://arxiv.org/html/2607.16799#Sx4.T1)\)\.
#### Question\- and response\-level controls\.
Source\-side estimation from within\-question activation differences preserves the cross\-domain separation: all eight adjustedΔCB\\Delta\_\{\\mathrm\{CB\}\}point estimates are positive \(0\.130–0\.748\), with joint 95% CIs excluding zero\. The complementary target\-side all\-cell analysis restores all four target cells while absorbing target\-question intercepts\. Across all eight transfers,WmetaW\_\{\\mathrm\{meta\}\}has a positive SJ coefficient, and allΔFE\\Delta\_\{\\mathrm\{FE\}\}estimates are positive \(0\.230–0\.530\), with 95% CIs excluding zero\. Within both transferred scores, the SJ coefficient exceeds the OC coefficient in every transfer, with all difference CIs excluding zero\. Thus,WmetaW\_\{\\mathrm\{meta\}\}retains SJ specificity, whereasWtruthW\_\{\\mathrm\{truth\}\}does not retain OC specificity\. Response\-level controls yield the same separation: residualizing projection scores against response log\-probability and response\-token count leaves all eight fixed\-windowΔCB\\Delta\_\{\\mathrm\{CB\}\}estimates positive \(0\.180–0\.751\), with CIs excluding zero\. Rendered\-sequence B/C token\-count matching yields a minimum estimate of 0\.198 \[0\.065, 0\.329\]\.
#### Cross\-domain transfer to MMLU and TruthfulQA\.
Directions estimated from Math and Movies are next evaluated on MMLU and binary TruthfulQA without fitting on either target dataset\. Each model–target conflict set contains 198–548 candidate responses from 160–483 questions\. In the fixed window, all 16Wmeta→SJW\_\{\\mathrm\{meta\}\}\\rightarrow\\mathrm\{SJ\}AUCs exceed 0\.5 \(0\.544–0\.823\), with 15/16 95% CIs above chance\. By contrast, all 16Wtruth→OCW\_\{\\mathrm\{truth\}\}\\rightarrow\\mathrm\{OC\}AUCs are below 0\.5 \(0\.206–0\.487\), with 11 of the 16 corresponding 95% CIs lying entirely below chance\. The resultingΔCB\\Delta\_\{\\mathrm\{CB\}\}values range from 0\.091 to 0\.581\. Figure[4](https://arxiv.org/html/2607.16799#Sx5.F4)shows the same component separation across depth: the SJ\-associated curves rise above chance in middle\-to\-late layers, whereas the OC\-associated curves remain near chance or reverse polarity\.
#### Forced\-choice and elicitation controls\.
The OOD separation is not explained by immediate preference for the forced answer option\. Residualization against answer\-option log\-probability and token count leaves 15/16 estimates positive, with 15/16 95% CIs above zero\. Adding answer\-letter fixed effects yields the same counts\. In both analyses, the sole nonpositive estimate is OLMo\-3\-7B Movies→\\rightarrowTruthfulQA, whose letter\-controlled estimate is−0\.030\-0\.030\[−0\.138\-0\.138, 0\.077\]\. With counterbalanced X/Y judgement labels, all 16 fixed\-window estimates remain positive \(0\.176–0\.682\), and 12/16 95% CIs lie above zero\. The four inconclusive CIs are the two source directions into Llama\-3\.1\-8B or OLMo\-3\-7B MMLU cells, where the balanced conflict sets contain only 13 or 4 B responses, respectively\.
#### Null and sensitivity analyses\.
The observed window means exceed the corresponding 95% null intervals in all 16 model–direction–control comparisons, both when source OC and SJ labels are shuffled and when fitted vectors are replaced with random directions of comparable norm \(Supplementary Fig\. S2\)\. The main pattern remains robust across confidence thresholds and a second Qwen2\.5\-7B strict\-pair draw\. A scoring\-rule audit likewise found no changes to SJ labels or the primary pair sample\.
## Discussion
#### Semantic validity of transferable readouts\.
A readout’s transferability and semantic validity are distinct\. When OC and SJ conflict, conventional correctness\-labelled contrasts often follow the model’s judgement rather than the ordering predicted by OC\. After factorial decomposition, the SJ\-associated contrast predicts SJ above chance in all eight held\-out within\-domain evaluations and preserves its expected polarity in all eight model\-by\-direction cross\-domain evaluations\. The OC\-associated contrast does not preserve its polarity under cross\-domain transfer and is often reversed\. Transferability alone therefore does not establish objective\-correctness semantics\.
#### Interpreting the transferable SJ signal\.
OC is externally defined by the relation between an answer and task\-specific ground truth\. SJ is generated from cues available to the model, which may include answer evidence, familiarity, fluency, or response policy\. Math requires multi\-step numerical reasoning whereas Movies requires person\-name recall, yet the SJ\-associated directions transfer between the domains despite their different knowledge demands and error structures\. At the final answer token, the separation is weak in early layers but sustained from middle to late layers, consistent with response\-level evaluation information becoming more separable at greater network depth\. This pattern may reflect endorsement or commitment as well as correctness monitoring; the judgement measure does not distinguish among these accounts\.
#### Implications for truth and metacognitive readouts\.
Aggregate correctness prediction does not identify what a probe reads out\. Semantic validation requires cases where external correctness and the model’s evaluation disagree, because aligned examples cannot separate them\. Before the separate judgement prompt, the final answer\-token representation predicts the model’s later behavioural judgement without establishing privileged introspective access\(Songet al\.[2025](https://arxiv.org/html/2607.16799#bib.bib40); Singhet al\.[2026](https://arxiv.org/html/2607.16799#bib.bib41)\)\.
#### Scope\.
The diagnostic sample comprises high\-confidence judgements for questions yielding usable correct and incorrect responses; it does not estimate ordinary\-output prevalence\. Evidence spans four instruction\-tuned models up to 14B and two free\-response domains; OOD uses externally specified answer\-letter candidates\. The factorial contrasts characterize associations between linear activation\-mean directions and subsequent behavioural judgements; they do not assess nonlinear or task\-adapted correctness predictors\. Broader source\-side elicitation, larger models, free\-response OOD tasks, and causal interventions along fitted directions would clarify generality and mechanism\.
Across tested models and diagnostic subsets, the most transferable component is more closely associated with subsequent self\-judgement than externally scored correctness\.
## References
- T\. Ashuach, S\. Gretz, Y\. Katz, Y\. Belinkov, and L\. Ein\-Dor \(2026\)Masked by consensus: disentangling privileged knowledge in LLM correctness\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),San Diego, California, United States,pp\. 10577–10596\.External Links:[Link](https://aclanthology.org/2026.acl-long.483/),[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.483)Cited by:[Truth\-direction transfer and semantic validation\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px1.p1.1)\.
- A\. Azaria and T\. Mitchell \(2023\)The internal state of an LLM knows when it’s lying\.InFindings of the Association for Computational Linguistics: EMNLP 2023,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 967–976\.External Links:[Link](https://aclanthology.org/2023.findings-emnlp.68/),[Document](https://dx.doi.org/10.18653/v1/2023.findings-emnlp.68)Cited by:[Introduction](https://arxiv.org/html/2607.16799#Sx1.p1.1),[Truth\-direction transfer and semantic validation\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px1.p1.1)\.
- W\. Azizian, M\. Kirchhof, E\. Ndiaye, L\. Béthune, M\. Klein, P\. Ablin, and M\. Cuturi \(2025\)The geometries of truth are orthogonal across tasks\.InICML 2025 Workshop on Reliable and Responsible Foundation Models,External Links:[Link](https://openreview.net/forum?id=FdfvGu5rM5)Cited by:[Introduction](https://arxiv.org/html/2607.16799#Sx1.p2.1),[Truth\-direction transfer and semantic validation\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px1.p1.1)\.
- Y\. Bao, X\. Zhang, T\. Du, X\. Zhao, Z\. Feng, H\. Peng, and J\. Yin \(2025\)Probing the geometry of truth: consistency and generalization of truth directions in LLMs across logical transformations and question answering tasks\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 682–700\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.38),[Link](https://aclanthology.org/2025.findings-acl.38/)Cited by:[Introduction](https://arxiv.org/html/2607.16799#Sx1.p1.1),[Truth\-direction transfer and semantic validation\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px1.p1.1)\.
- Y\. Belinkov \(2022\)Probing classifiers: promises, shortcomings, and advances\.Computational Linguistics48\(1\),pp\. 207–219\.External Links:[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00422),[Link](https://aclanthology.org/2022.cl-1.7/)Cited by:[Truth\-direction transfer and semantic validation\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px1.p1.1)\.
- C\. Burns, H\. Ye, D\. Klein, and J\. Steinhardt \(2023\)Discovering latent knowledge in language models without supervision\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=ETKGuby0hcs)Cited by:[Introduction](https://arxiv.org/html/2607.16799#Sx1.p1.1),[Truth\-direction transfer and semantic validation\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px1.p1.1)\.
- C\. S\. Cheang, H\. P\. Chan, W\. Zhang, and Y\. Deng \(2026\)Do LLMs really know what they don’t know? internal states mainly reflect knowledge recall rather than truthfulness\.InFindings of the Association for Computational Linguistics: ACL 2026,M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),San Diego, California, United States,pp\. 713–730\.External Links:[Link](https://aclanthology.org/2026.findings-acl.34/),[Document](https://dx.doi.org/10.18653/v1/2026.findings-acl.34),ISBN 979\-8\-89176\-395\-1Cited by:[Introduction](https://arxiv.org/html/2607.16799#Sx1.p2.1),[Truth\-direction transfer and semantic validation\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px1.p1.1)\.
- K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano, C\. Hesse, and J\. Schulman \(2021\)Training verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168\.External Links:[Link](https://arxiv.org/abs/2110.14168)Cited by:[§S2\.1](https://arxiv.org/html/2607.16799#S2.SS1.p1.1),[Datasets and models\.](https://arxiv.org/html/2607.16799#Sx4.SS0.SSS0.Px1.p1.1)\.
- A\. Grattafioriet al\.\(2024\)The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.External Links:[Link](https://arxiv.org/abs/2407.21783)Cited by:[Datasets and models\.](https://arxiv.org/html/2607.16799#Sx4.SS0.SSS0.Px1.p1.1)\.
- P\. Haller, M\. Ibrahim, P\. Kirichenko, L\. Sagun, and S\. J\. Bell \(2025\)LLM knowledge is brittle: truthfulness representations rely on superficial resemblance\.External Links:2510\.11905,[Document](https://dx.doi.org/10.48550/arXiv.2510.11905),[Link](https://arxiv.org/abs/2510.11905)Cited by:[Introduction](https://arxiv.org/html/2607.16799#Sx1.p2.1),[Truth\-direction transfer and semantic validation\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px1.p1.1)\.
- D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. Steinhardt \(2021a\)Measuring massive multitask language understanding\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=d7KBjmI3GmQ)Cited by:[§S2\.3](https://arxiv.org/html/2607.16799#S2.SS3.p1.1),[Datasets and models\.](https://arxiv.org/html/2607.16799#Sx4.SS0.SSS0.Px1.p1.1)\.
- D\. Hendrycks, C\. Burns, S\. Kadavath, A\. Arora, S\. Basart, E\. Tang, D\. Song, and J\. Steinhardt \(2021b\)Measuring mathematical problem solving with the math dataset\.InProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks,J\. Vanschoren and S\. Yeung \(Eds\.\),Vol\.1\.External Links:[Link](https://datasets-benchmarks-proceedings.neurips.cc/paper_files/paper/2021/file/be83ab3ecd0db773eb2dc1b0a17836a1-Paper-round2.pdf)Cited by:[§S2\.1](https://arxiv.org/html/2607.16799#S2.SS1.p1.1),[Datasets and models\.](https://arxiv.org/html/2607.16799#Sx4.SS0.SSS0.Px1.p1.1)\.
- J\. Hewitt and P\. Liang \(2019\)Designing and interpreting probes with control tasks\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing,pp\. 2733–2743\.External Links:[Document](https://dx.doi.org/10.18653/v1/D19-1275),[Link](https://aclanthology.org/D19-1275/)Cited by:[Truth\-direction transfer and semantic validation\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px1.p1.1)\.
- L\. Huang, W\. Yu, W\. Ma, W\. Zhong, Z\. Feng, H\. Wang, Q\. Chen, W\. Peng, X\. Feng, B\. Qin, and T\. Liu \(2025\)A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions\.ACM Transactions on Information Systems43\(2\),pp\. 1–55\.External Links:[Document](https://dx.doi.org/10.1145/3703155),[Link](https://dl.acm.org/doi/10.1145/3703155)Cited by:[Introduction](https://arxiv.org/html/2607.16799#Sx1.p1.1)\.
- S\. Kadavath, T\. Conerly, A\. Askell, T\. Henighan, D\. Drain, E\. Perez, N\. Schiefer, Z\. Hatfield\-Dodds, N\. DasSarma, E\. Tran\-Johnson, S\. Johnston, S\. El\-Showk, A\. Jones, N\. Elhage, T\. Hume, A\. Chen, Y\. Bai, S\. Bowman, S\. Fort, D\. Ganguli, D\. Hernandez, J\. Jacobson, J\. Kernion, S\. Kravec, L\. Lovitt, K\. Ndousse, C\. Olsson, S\. Ringer, D\. Amodei, T\. Brown, J\. Clark, N\. Joseph, B\. Mann, S\. McCandlish, C\. Olah, and J\. Kaplan \(2022\)Language models \(mostly\) know what they know\.External Links:2207\.05221,[Link](https://arxiv.org/abs/2207.05221)Cited by:[Introduction](https://arxiv.org/html/2607.16799#Sx1.p3.1),[Self\-evaluation, uncertainty, and introspection\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px2.p1.1)\.
- D\. Kumaran, A\. Conmy, F\. Barbero, S\. Osindero, V\. Patraucean, and P\. Veličković \(2026a\)How do llms compute verbal confidence\.External Links:2603\.17839,[Link](https://arxiv.org/abs/2603.17839)Cited by:[Introduction](https://arxiv.org/html/2607.16799#Sx1.p3.1),[Self\-evaluation, uncertainty, and introspection\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px2.p1.1)\.
- D\. Kumaran, V\. Patraucean, S\. Osindero, P\. Veličković, and N\. Daw \(2026b\)How LLMs detect and correct their own errors: the role of internal confidence signals\.External Links:2604\.22271,[Document](https://dx.doi.org/10.48550/arXiv.2604.22271),[Link](https://arxiv.org/abs/2604.22271)Cited by:[Self\-evaluation, uncertainty, and introspection\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px2.p1.1)\.
- J\. Li, H\. Xiong, R\. Wilson, M\. G\. Mattar, and M\. K\. Benna \(2025\)Language models are capable of metacognitive monitoring and control of their internal activations\.InAdvances in Neural Information Processing Systems,D\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen \(Eds\.\),Vol\.38,pp\. 60073–60108\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2025/file/56a225639da77e8f7c0409f6d5ba996b-Paper-Conference.pdf)Cited by:[Introduction](https://arxiv.org/html/2607.16799#Sx1.p3.1)\.
- K\. Li, O\. Patel, F\. Viégas, H\. Pfister, and M\. Wattenberg \(2023\)Inference\-time intervention: eliciting truthful answers from a language model\.InAdvances in Neural Information Processing Systems,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),Vol\.36,pp\. 41451–41530\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/file/81b8390039b7302c909cb769f8b6cd93-Paper-Conference.pdf)Cited by:[Introduction](https://arxiv.org/html/2607.16799#Sx1.p1.1),[Truth\-direction transfer and semantic validation\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px1.p1.1)\.
- H\. Lightman, V\. Kosaraju, Y\. Burda, H\. Edwards, B\. Baker, T\. Lee, J\. Leike, J\. Schulman, I\. Sutskever, and K\. Cobbe \(2024\)Let’s verify step by step\.InInternational Conference on Learning Representations,External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/hash/aca97732e30bcf1303bc22ac3924fd16-Abstract-Conference.html)Cited by:[§S2\.1](https://arxiv.org/html/2607.16799#S2.SS1.p1.1)\.
- S\. Lin, J\. Hilton, and O\. Evans \(2022a\)Teaching models to express their uncertainty in words\.Transactions on Machine Learning Research\.Note:External Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=8s8K2UZGTZ)Cited by:[Self\-evaluation, uncertainty, and introspection\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px2.p1.1)\.
- S\. Lin, J\. Hilton, and O\. Evans \(2022b\)TruthfulQA: measuring how models mimic human falsehoods\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),S\. Muresan, P\. Nakov, and A\. Villavicencio \(Eds\.\),Dublin, Ireland,pp\. 3214–3252\.External Links:[Link](https://aclanthology.org/2022.acl-long.229/),[Document](https://dx.doi.org/10.18653/v1/2022.acl-long.229)Cited by:[§S2\.3](https://arxiv.org/html/2607.16799#S2.SS3.p1.1),[Introduction](https://arxiv.org/html/2607.16799#Sx1.p1.1),[Datasets and models\.](https://arxiv.org/html/2607.16799#Sx4.SS0.SSS0.Px1.p1.1)\.
- S\. Lin, J\. Hilton, and O\. Evans \(2025\)TruthfulQA dataset\.Note:GitHub repositoryUpdated January 21, 2025External Links:[Link](https://github.com/sylinrl/TruthfulQA)Cited by:[§S2\.3](https://arxiv.org/html/2607.16799#S2.SS3.p1.1)\.
- K\. Liu, S\. Casper, D\. Hadfield\-Menell, and J\. Andreas \(2023\)Cognitive dissonance: why do language model outputs disagree with internal representations of truthfulness?\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,Singapore,pp\. 4791–4797\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.291/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.291)Cited by:[Truth\-direction transfer and semantic validation\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px1.p1.1)\.
- X\. Long, Y\. Fu, R\. Li, M\. Sheng, H\. Yu, X\. Han, and P\. Li \(2025\)When truthful representations flip under deceptive instructions?\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 16315–16335\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.826),[Link](https://aclanthology.org/2025.emnlp-main.826/)Cited by:[Truth\-direction transfer and semantic validation\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px1.p1.1)\.
- Y\. Lu, J\. Song, and W\. Wang \(2025\)A unified representation underlying the judgment of large language models\.External Links:2510\.27328,[Document](https://dx.doi.org/10.48550/arXiv.2510.27328),[Link](https://arxiv.org/abs/2510.27328)Cited by:[Introduction](https://arxiv.org/html/2607.16799#Sx1.p3.1)\.
- S\. Marks and M\. Tegmark \(2024\)The geometry of truth: emergent linear structure in large language model representations of true/false datasets\.InFirst Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=aajyHYjjsk)Cited by:[Introduction](https://arxiv.org/html/2607.16799#Sx1.p1.1),[Truth\-direction transfer and semantic validation\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px1.p1.1),[Experimental comparisons\.](https://arxiv.org/html/2607.16799#Sx4.SS0.SSS0.Px3.p1.2)\.
- S\. Miao, C\. Liang, and K\. Su \(2020\)A diverse corpus for evaluating and developing English math word problem solvers\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 975–984\.External Links:[Link](https://aclanthology.org/2020.acl-main.92/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.92)Cited by:[§S2\.1](https://arxiv.org/html/2607.16799#S2.SS1.p1.1)\.
- H\. Orgad, M\. Toker, Z\. Gekhman, R\. Reichart, I\. Szpektor, H\. Kotek, and Y\. Belinkov \(2025\)LLMs know more than they show: on the intrinsic representation of LLM hallucinations\.InInternational Conference on Learning Representations,External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/hash/a712d461e57201efe35d429a6f1731c1-Abstract-Conference.html)Cited by:[§S2\.2](https://arxiv.org/html/2607.16799#S2.SS2.p1.1),[Datasets and models\.](https://arxiv.org/html/2607.16799#Sx4.SS0.SSS0.Px1.p1.1)\.
- A\. Patel, S\. Bhattamishra, and N\. Goyal \(2021\)Are NLP models really able to solve simple math word problems?\.InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,K\. Toutanova, A\. Rumshisky, L\. Zettlemoyer, D\. Hakkani\-Tur, I\. Beltagy, S\. Bethard, R\. Cotterell, T\. Chakraborty, and Y\. Zhou \(Eds\.\),Online,pp\. 2080–2094\.External Links:[Link](https://aclanthology.org/2021.naacl-main.168/),[Document](https://dx.doi.org/10.18653/v1/2021.naacl-main.168)Cited by:[§S2\.1](https://arxiv.org/html/2607.16799#S2.SS1.p1.1)\.
- H\. Patel, T\. Chen, H\. Wei, E\. E\. Papalexakis, and J\. Chen \(2026\)Are LLM uncertainty and correctness encoded by the same features? a functional dissociation via sparse autoencoders\.External Links:2604\.19974,[Document](https://dx.doi.org/10.48550/arXiv.2604.19974),[Link](https://arxiv.org/abs/2604.19974)Cited by:[Introduction](https://arxiv.org/html/2607.16799#Sx1.p3.1),[Self\-evaluation, uncertainty, and introspection\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px2.p1.1)\.
- A\. Poulis, M\. Crovella, and E\. Terzi \(2026\)Testing the limits of truth directions in LLMs\.External Links:2604\.03754,[Document](https://dx.doi.org/10.48550/arXiv.2604.03754),[Link](https://arxiv.org/abs/2604.03754)Cited by:[Introduction](https://arxiv.org/html/2607.16799#Sx1.p2.1),[Truth\-direction transfer and semantic validation\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px1.p1.1)\.
- Qwen \(2025\)Qwen2\.5 technical report\.External Links:2412\.15115,[Link](https://arxiv.org/abs/2412.15115)Cited by:[Datasets and models\.](https://arxiv.org/html/2607.16799#Sx4.SS0.SSS0.Px1.p1.1)\.
- S\. Roy and D\. Roth \(2015\)Solving general arithmetic word problems\.InProceedings of the 2015 Conference on Empirical Methods in Natural Language Processing,L\. Màrquez, C\. Callison\-Burch, and J\. Su \(Eds\.\),Lisbon, Portugal,pp\. 1743–1752\.External Links:[Link](https://aclanthology.org/D15-1202/),[Document](https://dx.doi.org/10.18653/v1/D15-1202)Cited by:[§S2\.1](https://arxiv.org/html/2607.16799#S2.SS1.p1.1)\.
- S\. F\. Schouten, P\. Bloem, I\. Markov, and P\. Vossen \(2025\)Truth\-value judgment in language models: ‘truth directions’ are context sensitive\.InSecond Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=2H85485yAb)Cited by:[Truth\-direction transfer and semantic validation\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px1.p1.1)\.
- G\. Servedio, A\. De Bellis, D\. Di Palma, V\. W\. Anelli, and T\. Di Noia \(2025\)Are the hidden states hiding something? testing the limits of factuality\-encoding capabilities in LLMs\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 6089–6104\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.304),[Link](https://aclanthology.org/2025.acl-long.304/)Cited by:[Truth\-direction transfer and semantic validation\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px1.p1.1)\.
- S\. Singh, T\. Linzen, and S\. Ravfogel \(2026\)Can LLMs introspect? a reality check\.External Links:2605\.26242,[Document](https://dx.doi.org/10.48550/arXiv.2605.26242),[Link](https://arxiv.org/abs/2605.26242)Cited by:[Self\-evaluation, uncertainty, and introspection\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px2.p1.1),[Implications for truth and metacognitive readouts\.](https://arxiv.org/html/2607.16799#Sx6.SS0.SSS0.Px3.p1.1)\.
- S\. Song, J\. Hu, and K\. Mahowald \(2025\)Language models fail to introspect about their knowledge of language\.InProceedings of the 2nd Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=AivRDOFi5H)Cited by:[Self\-evaluation, uncertainty, and introspection\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px2.p1.1),[Implications for truth and metacognitive readouts\.](https://arxiv.org/html/2607.16799#Sx6.SS0.SSS0.Px3.p1.1)\.
- M\. Steyvers and M\. A\. K\. Peters \(2026\)Metacognition and uncertainty communication in humans and large language models\.Current Directions in Psychological Science35\(3\),pp\. 131–139\.Note:First published online November 18, 2025External Links:[Document](https://dx.doi.org/10.1177/09637214251391158),[Link](https://journals.sagepub.com/doi/10.1177/09637214251391158)Cited by:[Introduction](https://arxiv.org/html/2607.16799#Sx1.p4.1)\.
- Team Olmo \(2026\)Olmo 3\.External Links:2512\.13961,[Link](https://arxiv.org/abs/2512.13961)Cited by:[Datasets and models\.](https://arxiv.org/html/2607.16799#Sx4.SS0.SSS0.Px1.p1.1)\.
- K\. Tian, E\. Mitchell, A\. Zhou, A\. Sharma, R\. Rafailov, H\. Yao, C\. Finn, and C\. D\. Manning \(2023\)Just ask for calibration: strategies for eliciting calibrated confidence scores from language models fine\-tuned with human feedback\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,Singapore,pp\. 5433–5442\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.330),[Link](https://aclanthology.org/2023.emnlp-main.330/)Cited by:[Self\-evaluation, uncertainty, and introspection\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px2.p1.1)\.
- G\. Wang, W\. Wu, G\. Ye, Z\. Cheng, X\. Chen, and H\. Zheng \(2025a\)Decoupling metacognition from cognition: a framework for quantifying metacognitive ability in LLMs\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.39,pp\. 25353–25361\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v39i24.34723),[Link](https://ojs.aaai.org/index.php/AAAI/article/view/34723)Cited by:[Introduction](https://arxiv.org/html/2607.16799#Sx1.p3.1),[Self\-evaluation, uncertainty, and introspection\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px2.p1.1)\.
- K\. Wang, J\. Li, S\. Yang, Z\. Zhang, and D\. Wang \(2026\)When truth is overridden: uncovering the internal origins of sycophancy in large language models\.Proceedings of the AAAI Conference on Artificial Intelligence40\(39\),pp\. 33566–33574\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v40i39.40645),[Link](https://ojs.aaai.org/index.php/AAAI/article/view/40645)Cited by:[Truth\-direction transfer and semantic validation\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px1.p1.1)\.
- Y\. Wang, P\. Zhang, B\. Yang, D\. F\. Wong, and R\. Wang \(2025b\)Latent space chain\-of\-embedding enables output\-free LLM self\-evaluation\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=jxo70B9fQo)Cited by:[Self\-evaluation, uncertainty, and introspection\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px2.p1.1)\.
- Z\. Xiao, D\. Dou, B\. Xiong, Y\. Chen, and G\. Chen \(2026\)Enhancing uncertainty estimation in LLMs with expectation of aggregated internal belief\.Proceedings of the AAAI Conference on Artificial Intelligence40\(40\),pp\. 34043–34051\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v40i40.40698),[Link](https://ojs.aaai.org/index.php/AAAI/article/view/40698)Cited by:[Introduction](https://arxiv.org/html/2607.16799#Sx1.p3.1)\.
- M\. Xiong, Z\. Hu, X\. Lu, Y\. Li, J\. Fu, J\. He, and B\. Hooi \(2024\)Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs\.InInternational Conference on Learning Representations,External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/hash/6733cf15e10e2cd1d59af033c3bb8507-Abstract-Conference.html)Cited by:[Self\-evaluation, uncertainty, and introspection\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px2.p1.1)\.
- Z\. Yin, Q\. Sun, Q\. Guo, J\. Wu, X\. Qiu, and X\. Huang \(2023\)Do large language models know what they don’t know?\.InFindings of the Association for Computational Linguistics: ACL 2023,Toronto, Canada,pp\. 8653–8665\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.findings-acl.551),[Link](https://aclanthology.org/2023.findings-acl.551/)Cited by:[Self\-evaluation, uncertainty, and introspection\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px2.p1.1)\.
- Z\. J\. Ying, S\. Ravfogel, N\. Kriegeskorte, and P\. Hase \(2026\)The truthfulness spectrum hypothesis\.External Links:2602\.20273,[Document](https://dx.doi.org/10.48550/arXiv.2602.20273),[Link](https://arxiv.org/abs/2602.20273)Cited by:[Introduction](https://arxiv.org/html/2607.16799#Sx1.p2.1),[Truth\-direction transfer and semantic validation\.](https://arxiv.org/html/2607.16799#Sx2.SS0.SSS0.Px1.p1.1)\.
- A\. Zou, L\. Phan, S\. Chen, J\. Campbell, P\. Guo, R\. Ren, A\. Pan, X\. Yin, M\. Mazeika, A\. Dombrowski, S\. Goel, N\. Li, M\. J\. Byun, Z\. Wang, A\. Mallen, S\. Basart, S\. Koyejo, D\. Song, M\. Fredrikson, J\. Z\. Kolter, and D\. Hendrycks \(2025\)Representation engineering: a top\-down approach to AI transparency\.External Links:2310\.01405,[Link](https://arxiv.org/abs/2310.01405)Cited by:[Experimental comparisons\.](https://arxiv.org/html/2607.16799#Sx4.SS0.SSS0.Px3.p1.2)\.
Supplementary Material
## S1Supplementary Overview
This appendix documents the data construction, representation contrasts, statistical procedures, and robustness analyses underlying the main paper\. Table[S1](https://arxiv.org/html/2607.16799#S1.T1)links each experimental claim to its supporting analyses\. Unless stated otherwise, bootstrap intervals are obtained by resampling target questions\.
Table S1:Roadmap from the main experiments to supplementary evidence\.
## S2Datasets and Stimulus Construction
### S2\.1Math reasoning mixture
The Math domain contains 16,700 problems drawn from eight locally defined source tags\. The pool includes GSM8K\(Cobbeet al\.[2021](https://arxiv.org/html/2607.16799#bib.bib4)\), ASDiv\(Miaoet al\.[2020](https://arxiv.org/html/2607.16799#bib.bib5)\), and SVAMP\(Patelet al\.[2021](https://arxiv.org/html/2607.16799#bib.bib6)\), together with MATH\(Hendryckset al\.[2021b](https://arxiv.org/html/2607.16799#bib.bib7)\), MultiArith\(Roy and Roth[2015](https://arxiv.org/html/2607.16799#bib.bib8)\), a MATH\-500 subset associated with process\-supervision evaluation\(Lightmanet al\.[2024](https://arxiv.org/html/2607.16799#bib.bib9)\), and two AMC subsets\. The source tags indicate local provenance only and are not used as experimental strata\.
A Qwen2\.5\-1\.5B\-Instruct pilot generated eight solutions per problem with temperature 0\.6 and a maximum generation length of 4,096 tokens\. Its prompt ended with: “Please reason step by step, and put your final answer within\\boxed\{\}\.” We retained items satisfyingncorrect<8n\_\{\\mathrm\{correct\}\}<8and mean pilot output length<380<380tokens\. This removes questions solved on every pilot attempt and the longest approximately 10% of pilot responses while retaining all\-wrong pilot questions\. Model\-specific strict\-pair construction later requires both a correct and an incorrect usable answer\. The resulting manuscript\-facing pool contains 5,549 questions \(Table[S2](https://arxiv.org/html/2607.16799#S2.T2)\)\.
All 5,549 references are numeric\. Retained responses have oneAnswer:field and normalize to a nonnegative integer or decimal; expressions and multiple candidate answers were excluded from the usable response pool\. Correctness labels were obtained using normalized and symbolic\-equivalence checks\.
Table S2:Math source composition before and after pilot filtering\. Counts reflect the data files used in the reported analyses\.
### S2\.2Movies factual recall
The Movies domain contains 17,856 person\-name recall records from the Movies QA data ofOrgadet al\.\([2025](https://arxiv.org/html/2607.16799#bib.bib30)\): 10,000 in the distributed train file and 7,856 in the test file\. Questions follow the form “Who acted as \[character\] in the movie \[movie\]?” The answer instruction requests only the actor name\. Responses are lowercased, Unicode\-normalized, stripped of diacritics, and whitespace\-normalized; a response is marked correct when the normalized record\-level reference actor appears in the normalized response\. Unlike Math, this domain requires direct entity recall rather than multi\-step numerical reasoning\.
The 17,856 source records contain 17,518 unique question strings\. The remaining repetitions arise from duplicate source entries and from missing or non\-unique character fields, including blank roles and labels such as “Himself” or “Herself”\. Among the repeated prompts, 159 are associated with more than one reference actor\. Depending on the model, prompts from this multi\-reference set account for 1\.4%–2\.4% of retained rows\. Identical prompt strings are assigned to the same question cluster in grouped cross\-validation and bootstrap inference\. A complementary sensitivity analysis removes the complete multi\-reference set before direction fitting or evaluation \(Sec\. S8\.7\)\. The sample\-flow table accordingly reports both the number of retained response pairs and the number of unique question clusters\.
### S2\.3MMLU and binary TruthfulQA
The out\-of\-distribution evaluation uses 400 MMLU test questions\(Hendryckset al\.[2021a](https://arxiv.org/html/2607.16799#bib.bib3)\), sampled with a fixed seed as 100 questions from each of the humanities, other, social\-sciences, and STEM categories\. Exact subjects and question identifiers are retained in the released data artifact\. Each question is expanded into four candidate\-response records, one for each answer letter, yielding 1,600 forced\-choice responses\. Binary TruthfulQA\(Linet al\.[2022b](https://arxiv.org/html/2607.16799#bib.bib1),[2025](https://arxiv.org/html/2607.16799#bib.bib2)\)comprises 790 questions\. Each question is paired with its designated true and false candidate answers, yielding 1,580 candidate\-response records\.
In out\-of\-distribution tasks, each candidate answer letter is inserted as the assistant response and followed by the same Yes/No self\-judgement query used in the main domains\. This procedure measures the model’s judgement of correct and incorrect candidates at an identical, exactly controlled response site\. Neither MMLU nor TruthfulQA contributes data to direction fitting\.
## S3Complete Prompts and Activation Site
### S3\.1Free\-response prompts
The complete Math user template is:
> Question: \{question\} Strictly output your response in the following exact format: Reasoning: A concise reasoning focusing on the core steps needed to solve the problem\. Answer: <final numeric answer only\>\.
The complete Movies user template is:
> \{question\} Return only the name of the actor\.
After a free\-response answer was sampled, the judgement conversation was formed by retaining the original user message and sampled assistant answer and appending the following user turn:
> Do you believe the answer above is correct? Answer only with Yes or No\.
Next\-token Yes/No scores define the stored judgement probabilitypjudgep\_\{\\mathrm\{judge\}\}\. This judgement turn is used only to define self\-judgement \(SJ\) and is excluded from the sequence used for activation extraction\.
### S3\.2Forced\-choice prompts
For MMLU and binary TruthfulQA, the user message contains the question and all answer options, followed by:
> Answer only with a single letter\.
The assistant turn is set to one candidate answer letter\. The same judgement turn shown above is then appended to computepjudgep\_\{\\mathrm\{judge\}\}\. Candidate expansion represents every target question with both objectively correct and objectively incorrect responses, without target\-task answer generation or target\-task direction fitting\.
### S3\.3Answer\-token extraction
For each retained response, the model\-specific chat template is used to render the user prompt and assistant answer\. Before tokenization, preprocessing removes any terminal chat marker \(for example,<\|im\_end\|\>,<\|eot\_id\|\>, or<\|endoftext\|\>\), and removes a terminal period when present\. The resulting sequence is tokenized, and the residual\-stream hidden state at the output of each complete transformer block is recorded at the final non\-padding sequence position\. This is the block output rather than an intermediate attention or MLP activation\. Accordingly,xi,ℓx\_\{i,\\ell\}is aligned with the final numeric\-answer token in Math, the final actor\-name token in Movies, and the forced answer\-letter token in the OOD tasks\. Forward passes are performed in bfloat16, and the resulting activation arrays are stored in float16\. Layer indices refer to transformer blocks and range from zero toL−1L\-1\.
Table S3:Prompt roles and stored signals\. OC denotes objective correctness, and SJ denotes the thresholded high\-confidence self\-judgement label\.
## S4Strict\-Pair Construction and Sample Flow
For each model and main domain, strict pairs are produced as follows:
1. 1\.Generate eight stochastic answers per question\. Main\-model sampling uses temperature 1\.0 and maximum generation lengths of 384 tokens for Math and 256 for Movies\.
2. 2\.Apply the domain\-specific correctness parser and discard unusable responses\.
3. 3\.Retain questions with at least one correct and one incorrect usable response, then sample one response of each type\. Self\-judgement is elicited for this sampled pair, after which the paired confidence criterion is applied\.
4. 4\.Computepjudgep\_\{\\mathrm\{judge\}\}for both selected responses using the same model\. The one\-token call uses temperature 1\.0 and retains up to four next\-token log\-probability candidates\. If the literalYesandNocandidates are both present, their probabilities are normalized as pjudge=exp\(ℓYes\)exp\(ℓYes\)\+exp\(ℓNo\)\.p\_\{\\mathrm\{judge\}\}=\\frac\{\\exp\(\\ell\_\{\\mathrm\{Yes\}\}\)\}\{\\exp\(\\ell\_\{\\mathrm\{Yes\}\}\)\+\\exp\(\\ell\_\{\\mathrm\{No\}\}\)\}\.If onlyYesis present, its full\-vocabulary probability is used; if onlyNois present, its complement is used; and if neither is present,pjudgep\_\{\\mathrm\{judge\}\}is set to 0\.5\. Rows assigned 0\.5 by this final fallback are excluded by the primary symmetric threshold atτ=0\.7\\tau=0\.7\.
5. 5\.Retain the pair only if both responses satisfypjudge≤1−τp\_\{\\mathrm\{judge\}\}\\leq 1\-\\tauorpjudge≥τp\_\{\\mathrm\{judge\}\}\\geq\\tau, withτ=0\.7\\tau=0\.7in the primary analysis\.
6. 6\.AssignSJ=0\\mathrm\{SJ\}=0on the lower interval andSJ=1\\mathrm\{SJ\}=1on the upper interval\. With tuples ordered as \(OC,SJ\), the four cells are A=\(1,1\),B=\(1,0\),C=\(0,1\),D=\(0,0\)\.A=\(1,1\),\\quad B=\(1,0\),\\quad C=\(0,1\),\\quad D=\(0,0\)\.
Section[S8\.7](https://arxiv.org/html/2607.16799#S8.SS7)evaluates two upstream construction choices directly\. Full Yes/No normalization changes neither SJ labels nor paired retention in any of the eight model–domain cells, while a second Qwen2\.5\-7B draw from the frozen pass@8 pools preserves the cross\-domain layer profiles and positive fixed\-window effects\.
Each retained source item contributes one objectively correct and one objectively incorrect response\. Because SJ is measured rather than experimentally assigned, the B and C cell counts need not be equal\. Table[S4](https://arxiv.org/html/2607.16799#S4.T4)reports the model\-specific sample flow\. “Eligible pairs” contain one sampled correct response and one sampled incorrect response before judgement filtering\. “Strict pairs” additionally require both responses to satisfy the pairedτ=0\.7\\tau=0\.7criterion\. “Uniqueqq” denotes exact prompt\-string clusters and can therefore be smaller than the number of pairs when multiple source records contain the same prompt\.
Table S4:Main\-domain sample flow and strict\-sample cell counts\. Generated items are source records presented to each model\. The number of retained rows is twice the number of strict pairs\. Uniqueqqdenotes exact prompt\-string clusters used for grouped cross\-validation and bootstrap inference\.The resulting estimand is conditional on questions for which the model produces both a correct response and an incorrect response and assigns high\-confidence judgements to the sampled pair\. The analyses therefore characterize the relation between OC and SJ within this diagnostic subset rather than the prevalence of OC–SJ conflict cases in unconstrained model outputs\.
## S5Factorial Contrasts and Formal Derivations
### S5\.1Cell means and cancellation of the interaction
Leto=2OC−1o=2\\mathrm\{OC\}\-1ands=2SJ−1s=2\\mathrm\{SJ\}\-1\. For domaintt, consider the associative representation model
x=αtovtOC\+βtsvtSJ\+ηtosvtINT\+ϵ,x=\\alpha\_\{t\}ov^\{\\mathrm\{OC\}\}\_\{t\}\+\\beta\_\{t\}sv^\{\\mathrm\{SJ\}\}\_\{t\}\+\\eta\_\{t\}osv^\{\\mathrm\{INT\}\}\_\{t\}\+\\epsilon,\(S1\)wherevtOCv^\{\\mathrm\{OC\}\}\_\{t\},vtSJv^\{\\mathrm\{SJ\}\}\_\{t\}, andvtINTv^\{\\mathrm\{INT\}\}\_\{t\}are domain\-specific directions andαt,βt,ηt\\alpha\_\{t\},\\beta\_\{t\},\\eta\_\{t\}are their response\-level effect magnitudes\.ϵ\\epsilonis a zero\-mean residual term\. Suppressing the domain subscript for readability, the expected cell representations are
μA\\displaystyle\\mu\_\{A\}=αvOC\+βvSJ\+ηvINT,\\displaystyle=\\alpha v^\{\\mathrm\{OC\}\}\+\\beta v^\{\\mathrm\{SJ\}\}\+\\eta v^\{\\mathrm\{INT\}\},\(S2\)μB\\displaystyle\\mu\_\{B\}=αvOC−βvSJ−ηvINT,\\displaystyle=\\alpha v^\{\\mathrm\{OC\}\}\-\\beta v^\{\\mathrm\{SJ\}\}\-\\eta v^\{\\mathrm\{INT\}\},\(S3\)μC\\displaystyle\\mu\_\{C\}=−αvOC\+βvSJ−ηvINT,\\displaystyle=\-\\alpha v^\{\\mathrm\{OC\}\}\+\\beta v^\{\\mathrm\{SJ\}\}\-\\eta v^\{\\mathrm\{INT\}\},\(S4\)μD\\displaystyle\\mu\_\{D\}=−αvOC−βvSJ\+ηvINT\.\\displaystyle=\-\\alpha v^\{\\mathrm\{OC\}\}\-\\beta v^\{\\mathrm\{SJ\}\}\+\\eta v^\{\\mathrm\{INT\}\}\.\(S5\)The two factorial main effects are
Wmeta\\displaystyle W\_\{\\mathrm\{meta\}\}=\(μA−μB\)\+\(μC−μD\)2=2βvSJ,\\displaystyle=\\frac\{\(\\mu\_\{A\}\-\\mu\_\{B\}\)\+\(\\mu\_\{C\}\-\\mu\_\{D\}\)\}\{2\}=2\\beta v^\{\\mathrm\{SJ\}\},\(S6\)Wtruth\\displaystyle W\_\{\\mathrm\{truth\}\}=\(μA−μC\)\+\(μB−μD\)2=2αvOC\.\\displaystyle=\\frac\{\(\\mu\_\{A\}\-\\mu\_\{C\}\)\+\(\\mu\_\{B\}\-\\mu\_\{D\}\)\}\{2\}=2\\alpha v^\{\\mathrm\{OC\}\}\.\(S7\)The interaction cancels in both contrasts\. By comparison, the conventional aligned contrast is
Wmix=μA−μD=2αvOC\+2βvSJ=Wtruth\+Wmeta\.W\_\{\\mathrm\{mix\}\}=\\mu\_\{A\}\-\\mu\_\{D\}=2\\alpha v^\{\\mathrm\{OC\}\}\+2\\beta v^\{\\mathrm\{SJ\}\}=W\_\{\\mathrm\{truth\}\}\+W\_\{\\mathrm\{meta\}\}\.\(S8\)Thus, whenever both main\-effect components are present,WmixW\_\{\\mathrm\{mix\}\}combines OC\- and SJ\-associated variation rather than separating OC\-associated variation from SJ\-associated variation\.
### S5\.2Conflict\-set ordering
For a scorerW\(x\)=⟨x,W⟩r\_\{W\}\(x\)=\\langle x,W\\rangle, define
AUCC:B\(W\)=Pr\{rW\(XC\)\>rW\(XB\)\},\\mathrm\{AUC\}\_\{C:B\}\(W\)=\\Pr\\\{r\_\{W\}\(X\_\{C\}\)\>r\_\{W\}\(X\_\{B\}\)\\\},\(S9\)with ties assigned half weight\. Cell B contains objectively correct but self\-rejected responses, whereas cell C contains objectively incorrect but self\-endorsed responses\. An OC\-associated score should rank B above C, yieldingAUCC:B<0\.5\\mathrm\{AUC\}\_\{C:B\}<0\.5; an SJ\-associated score should rank C above B, yieldingAUCC:B\>0\.5\\mathrm\{AUC\}\_\{C:B\}\>0\.5\. Within one domain, Eq\.[S1](https://arxiv.org/html/2607.16799#S5.E1)gives
𝔼\[rWmix\(XC\)−rWmix\(XB\)\]=4β2‖vSJ‖2−4α2‖vOC‖2\.\\mathbb\{E\}\[r\_\{W\_\{\\mathrm\{mix\}\}\}\(X\_\{C\}\)\-r\_\{W\_\{\\mathrm\{mix\}\}\}\(X\_\{B\}\)\]=4\\beta^\{2\}\\\|v^\{\\mathrm\{SJ\}\}\\\|^\{2\}\-4\\alpha^\{2\}\\\|v^\{\\mathrm\{OC\}\}\\\|^\{2\}\.\(S10\)The ordering of B and C underWmixW\_\{\\mathrm\{mix\}\}therefore indicates whether the SJ\- or OC\-associated component contributes more strongly to the mixed contrast in Exp1\.
### S5\.3Cross\-domain transfer
LetWmetaaW^\{a\}\_\{\\mathrm\{meta\}\}andWtruthaW^\{a\}\_\{\\mathrm\{truth\}\}denote directions estimated in source domainaa, and letxbx^\{b\}denote an activation from target domainbb\. The corresponding transfer scores are
rmetaa→b\(xb\)=⟨xb,Wmetaa⟩,rtrutha→b\(xb\)=⟨xb,Wtrutha⟩,r^\{a\\rightarrow b\}\_\{\\mathrm\{meta\}\}\(x^\{b\}\)=\\langle x^\{b\},W^\{a\}\_\{\\mathrm\{meta\}\}\\rangle,\\qquad r^\{a\\rightarrow b\}\_\{\\mathrm\{truth\}\}\(x^\{b\}\)=\\langle x^\{b\},W^\{a\}\_\{\\mathrm\{truth\}\}\\rangle,\(S11\)where both directions are estimated from the source\-domain cell means\. Their scale does not affect the AUC rankings\. Transfer performance is summarized byAUC\(rmeta→SJ\)\\mathrm\{AUC\}\(r\_\{\\mathrm\{meta\}\}\\rightarrow\\mathrm\{SJ\}\)andAUC\(rtruth→OC\)\\mathrm\{AUC\}\(r\_\{\\mathrm\{truth\}\}\\rightarrow\\mathrm\{OC\}\)\.
Within the B/C conflict set, OC and SJ induce complementary class assignments\. Their relative transfer asymmetry is
ΔCB=AUC\(Wmeta→SJ\)−AUC\(Wtruth→OC\)\.\\Delta\_\{\\mathrm\{CB\}\}=\\mathrm\{AUC\}\(W\_\{\\mathrm\{meta\}\}\\rightarrow\\mathrm\{SJ\}\)\-\\mathrm\{AUC\}\(W\_\{\\mathrm\{truth\}\}\\rightarrow\\mathrm\{OC\}\)\.\(S12\)The two component AUCs indicate whether each source direction transfers above chance, whereasΔCB\\Delta\_\{\\mathrm\{CB\}\}compares their transfer performance on the same target responses\.
The target\-domain mean difference isμCb−μBb=−2αbvbOC\+2βbvbSJ\\mu\_\{C\}^\{b\}\-\\mu\_\{B\}^\{b\}=\-2\\alpha\_\{b\}v\_\{b\}^\{\\mathrm\{OC\}\}\+2\\beta\_\{b\}v\_\{b\}^\{\\mathrm\{SJ\}\}\. Its projection onto a source SJ direction increases with cross\-domain alignment betweenvaSJv\_\{a\}^\{\\mathrm\{SJ\}\}andvbSJv\_\{b\}^\{\\mathrm\{SJ\}\}and decreases with alignment between the source SJ direction and the target OC component\. Exp2B estimates these relations without imposing orthogonality between the latent components\.
### S5\.4Source\-question\-adjusted direction sensitivity
The pooled factorial directions weight the four source cells equally\. A complementary estimator removes additive source\-question variation before estimating the component directions\. Letxq\+x\_\{q\}^\{\+\}andxq−x\_\{q\}^\{\-\}denote the activations of the correct and incorrect responses retained for source questionqq, and letsq\+,sq−∈\{−1,\+1\}s\_\{q\}^\{\+\},s\_\{q\}^\{\-\}\\in\\\{\-1,\+1\\\}denote their signed SJ labels\. Under the response\-level model
xq±=hq\+bOCoq±\+bSJsq±\+bINToq±sq±\+ϵq±,x\_\{q\}^\{\\pm\}=h\_\{q\}\+b\_\{\\mathrm\{OC\}\}o\_\{q\}^\{\\pm\}\+b\_\{\\mathrm\{SJ\}\}s\_\{q\}^\{\\pm\}\+b\_\{\\mathrm\{INT\}\}o\_\{q\}^\{\\pm\}s\_\{q\}^\{\\pm\}\+\\epsilon\_\{q\}^\{\\pm\},\(S13\)whereoq\+=\+1o\_\{q\}^\{\+\}=\+1andoq−=−1o\_\{q\}^\{\-\}=\-1, the within\-question difference is
dq=xq\+−xq−=2bOC\+\(sq\+−sq−\)bSJ\+\(sq\+\+sq−\)bINT\+ϵ~q\.d\_\{q\}=x\_\{q\}^\{\+\}\-x\_\{q\}^\{\-\}=2b\_\{\\mathrm\{OC\}\}\+\(s\_\{q\}^\{\+\}\-s\_\{q\}^\{\-\}\)b\_\{\\mathrm\{SJ\}\}\+\(s\_\{q\}^\{\+\}\+s\_\{q\}^\{\-\}\)b\_\{\\mathrm\{INT\}\}\+\\widetilde\{\\epsilon\}\_\{q\}\.\(S14\)The additive question intercepthqh\_\{q\}cancels exactly\. At each layer, multivariate least squares across source pairs estimates the three coefficient vectors;b^SJ\\widehat\{b\}\_\{\\mathrm\{SJ\}\}andb^OC\\widehat\{b\}\_\{\\mathrm\{OC\}\}replaceWmetaW\_\{\\mathrm\{meta\}\}andWtruthW\_\{\\mathrm\{truth\}\}, respectively, in the unchanged cross\-domain B/C evaluation\. The four possible source\-pair patterns are A/C, A/D, B/C, and B/D, with design rows\[2,sq\+−sq−,sq\+\+sq−\]\[2,s\_\{q\}^\{\+\}\-s\_\{q\}^\{\-\},s\_\{q\}^\{\+\}\+s\_\{q\}^\{\-\}\]\. Identifiability is checked from the rank of this three\-column design for every source sample and bootstrap replicate\.
### S5\.5Target\-question\-controlled all\-cell model
The B/C analysis focuses on responses for which OC and SJ conflict\. A complementary analysis uses responses from all four cells\. For each componentk∈\{meta,truth\}k\\in\\\{\\mathrm\{meta\},\\mathrm\{truth\}\\\}, projection scores are standardized separately within each target\-domain and layer combination and then averaged over the fixed layer window\. The resulting scores are entered into
ziq\(k\)=γq\+βOC\(k\)oiq\+βSJ\(k\)siq\+βINT\(k\)oiqsiq\+ϵiq\.z^\{\(k\)\}\_\{iq\}=\\gamma\_\{q\}\+\\beta^\{\(k\)\}\_\{\\mathrm\{OC\}\}o\_\{iq\}\+\\beta^\{\(k\)\}\_\{\\mathrm\{SJ\}\}s\_\{iq\}\+\\beta^\{\(k\)\}\_\{\\mathrm\{INT\}\}o\_\{iq\}s\_\{iq\}\+\\epsilon\_\{iq\}\.\(S15\)The coefficients are therefore identified from variation among responses associated with the same question\. The cross\-component contrast is
ΔFE=βSJ\(meta\)−βOC\(truth\)\.\\Delta\_\{\\mathrm\{FE\}\}=\\beta^\{\(\\mathrm\{meta\}\)\}\_\{\\mathrm\{SJ\}\}\-\\beta^\{\(\\mathrm\{truth\}\)\}\_\{\\mathrm\{OC\}\}\.\(S16\)As a within\-component semantic diagnostic, we also compute
Δspec\(meta\)\\displaystyle\\Delta^\{\(\\mathrm\{meta\}\)\}\_\{\\mathrm\{spec\}\}=βSJ\(meta\)−βOC\(meta\),\\displaystyle=\\beta^\{\(\\mathrm\{meta\}\)\}\_\{\\mathrm\{SJ\}\}\-\\beta^\{\(\\mathrm\{meta\}\)\}\_\{\\mathrm\{OC\}\},\(S17\)Δspec\(truth\)\\displaystyle\\Delta^\{\(\\mathrm\{truth\}\)\}\_\{\\mathrm\{spec\}\}=βOC\(truth\)−βSJ\(truth\)\.\\displaystyle=\\beta^\{\(\\mathrm\{truth\}\)\}\_\{\\mathrm\{OC\}\}\-\\beta^\{\(\\mathrm\{truth\}\)\}\_\{\\mathrm\{SJ\}\}\.\(S18\)A positive value denotes the specificity implied by the source construction: SJ forWmetaW\_\{\\mathrm\{meta\}\}and OC forWtruthW\_\{\\mathrm\{truth\}\}\. Each question\-cluster bootstrap replicate repeats both score standardization and model estimation\.
### S5\.6Response\-likelihood residualization
For free responses, response log\-probability is the mean teacher\-forced log\-probability over all tokens in the generated assistant response\. Within each target dataset and layer, each projection score is residualized against mean response log\-probability and response\-token count\. OOD residual models additionally include answer\-letter indicators\. Component AUCs are computed from the layer\-specific residual scores and then averaged over the fixed window, matching the primary estimand\. Layer\-wise nuisance regressions are refitted in every bootstrap replicate\.
## S6Evaluation and Statistical Inference
#### Direction fitting and centering\.
At each layer, cell means and the source centering vector are estimated exclusively from the source\-domain rows or the training partition\. AUCs are computed from the resulting projection scores; their rankings are invariant to positive rescaling of the fitted directions\. For the all\-cell fixed\-effect analysis, the two projected components are standardized separately before their coefficients are compared\.
Exp2A uses five\-foldGroupKFold, with the question\-cluster identifier as the grouping variable, so that the correct and incorrect responses associated with the same question remain in the same fold\. Out\-of\-fold scores are concatenated before evaluation\. Exp2B estimates each direction once from all strict source\-domain rows and uses no target\-domain labels during direction estimation\.
#### Layer summaries\.
Layer\-wise curves describe how each effect varies across network depth\. Primary fixed\-window summaries average layers satisfyingℓ/\(L−1\)∈\[0\.40,0\.80\]\\ell/\(L\-1\)\\in\[0\.40,0\.80\]: layers 11–21 for the 28\-layer Qwen2\.5\-7B, 13–24 for the 32\-layer models, and 19–37 for Qwen2\.5\-14B\. Within a bootstrap replicate, AUCs are computed by layer and then averaged over this fixed window\.
The window covers a broad middle\-to\-late portion of each architecture, from 40% to 80% of normalized depth, rather than selecting a single layer from the observed results\. Complete layer\-wise curves are reported alongside the fixed\-window summaries\. For descriptive peak analyses, the model\-level peak is the layer that maximizes the meanΔCB\\Delta\_\{\\mathrm\{CB\}\}across the two Math–Movies transfer directions\. Peak intervals are descriptive and are not interpreted as selection\-adjusted hypothesis tests\.
#### Cluster bootstrap\.
Unless noted otherwise, confidence intervals are based on 1,000 question\-cluster bootstrap resamples\. Sampling a question retains all associated response rows, thereby preserving the B/C pairing and the repeated candidate responses in the OOD tasks\. For the primary Exp2B fixed\-window analysis, source and target question clusters are resampled independently, and both factorial directions are re\-estimated in every replicate\. For the all\-layer figures and secondary analyses, the fitted source direction is held fixed, so the resulting intervals quantify target\-population uncertainty conditional on that direction\. Percentile 95% intervals are computed from valid replicates\. All\-layer counts summarize rows of the layer\-by\-direction result grid and are descriptive because adjacent layers are not independent observations\.
Table S5:Descriptive all\-layer summaries\. “CI above” counts layer\-direction rows whose target\-question 95% interval is above the corresponding reference \(0\.5 for Exp1, zero for the remaining analyses\)\.
## S7In\-Domain Validation and Layer\-Wise Results
Figure[S1](https://arxiv.org/html/2607.16799#S7.F1)reports Exp2A with question\-grouped folds\. Seven of the eight fixed\-window estimates are positive, with 95% confidence intervals excluding zero \(Table[S6](https://arxiv.org/html/2607.16799#S7.T6)\)\. The only estimate whose interval includes zero is Llama\-3\.1\-8B on Math,ΔCB=−0\.042\\Delta\_\{\\mathrm\{CB\}\}=\-0\.042\[−0\.097\-0\.097,0\.0120\.012\], while the same model on Movies shows the largest in\-domain contrast\. Across models and domains, positive contrasts generally emerge after the earliest layers and remain present across much of the fixed middle\-to\-late layer window\.
Figure S1:Question\-grouped in\-domain validation \(Exp2A\)\. Each direction is fitted within the training questions of a five\-fold split and evaluated on held\-out questions\. Curves showΔCB\\Delta\_\{\\mathrm\{CB\}\}by layer; shaded bands are 95% target\-question cluster\-bootstrap intervals\. The pale vertical band marks the fixed normalized layer window\[0\.40,0\.80\]\[0\.40,0\.80\]\.Table S6:Exp2A fixed\-window component AUCs and relative transfer asymmetry\. Intervals use 1,000 question\-cluster bootstrap resamples of the out\-of\-fold predictions\.Table[S7](https://arxiv.org/html/2607.16799#S7.T7)summarizes the descriptive peak layers for Exp2B\. Within each model, a single layer is selected by maximizing the mean point estimate ofΔCB\\Delta\_\{\\mathrm\{CB\}\}across the two Math–Movies transfer directions\. The resulting peak summaries describe the location of the strongest observed cross\-domain effect but are not used for the primary fixed\-window inference\.
Table S7:Descriptive cross\-domain peak layers\. Confidence intervals condition on the selected source direction and are not adjusted for peak selection\.### S7\.1Joint source–target bootstrap
Table[S8](https://arxiv.org/html/2607.16799#S7.T8)reports the fixed\-window inference used in the main Exp2B table\. Each replicate independently resamples source and target question clusters, reconstructsWmetaW\_\{\\mathrm\{meta\}\}andWtruthW\_\{\\mathrm\{truth\}\}at every selected layer, and evaluates the refitted directions on the resampled target\-domain B/C conflicts\. All 1,000 replicates were valid in each of the eight transfer conditions, and every resampled source dataset retained observations from all four factorial cells\.
Table S8:Exp2B joint source–target question bootstrap\. Sourceqqand targetqqare the numbers of question clusters entering direction fitting and B/C evaluation, respectively\. Values are fixed\-window AUCs or differences with 95% percentile intervals\.
## S8Main\-Domain Controls
### S8\.1Source\-question\-adjusted direction sensitivity
Table[S9](https://arxiv.org/html/2607.16799#S8.T9)reports the cross\-domain evaluation after estimating both source directions from Eq\.[S14](https://arxiv.org/html/2607.16799#S5.E14)\. All eightΔCB\\Delta\_\{\\mathrm\{CB\}\}estimates remain positive, ranging from 0\.130 to 0\.748, and every joint source–target bootstrap 95% interval excludes zero\. The point estimates differ from the corresponding pooled factorial estimates by at most 0\.026\. All eight source designs have rank three; all 1,000 bootstrap replicates remain identifiable, including OLMo\-3\-7B Movies, where the A/C, A/D, and B/D patterns identify the coefficients without observed B/C source pairs\.
Table S9:Source\-question\-adjusted Exp2B sensitivity\. Directions are estimated from within\-question correct\-minus\-incorrect activation differences\. AC/AD/BC/BD gives the source\-pair support for the four SJ combinations; Rank is the three\-column design rank, and Valid is the number of identifiable joint source–target bootstrap replicates\. Entries are fixed\-window AUCs or differences with 95% percentile intervals\. The source\-adjusted directionsb^SJ\\widehat\{b\}\_\{\\mathrm\{SJ\}\}andb^OC\\widehat\{b\}\_\{\\mathrm\{OC\}\}occupy the evaluation roles of the pooledWmetaW\_\{\\mathrm\{meta\}\}andWtruthW\_\{\\mathrm\{truth\}\}, respectively\.
### S8\.2OC\-only mass\-mean control
The OC\-only direction isWOC−only=E\[x∣OC=1\]−E\[x∣OC=0\]W\_\{\\mathrm\{OC\-only\}\}=E\[x\\mid\\mathrm\{OC\}=1\]\-E\[x\\mid\\mathrm\{OC\}=0\], estimated from all source rows without SJ labels\. This is a canonical unwhitened mass\-mean direction; because each retained source item contributes one correct and one incorrect response, it is also equal to the average within\-question correct\-minus\-incorrect activation difference\.
The fitted direction is evaluated on B/C conflicts in the target domain\. In all eight transfer conditions, it assigns higher scores to C responses, which are incorrect but self\-endorsed, than to B responses, which are correct but self\-rejected \(Table[S10](https://arxiv.org/html/2607.16799#S8.T10)\)\. Thus, the conflict\-set reversal observed for the mixed contrast is also present when the source direction is estimated solely from objective\-correctness labels\.
Table S10:Fixed\-window OC\-only mass\-mean control\. The direction is the unwhitened source class\-mean contrastE\[x∣OC=1\]−E\[x∣OC=0\]E\[x\\mid\\mathrm\{OC\}=1\]\-E\[x\\mid\\mathrm\{OC\}=0\], fitted without SJ labels\. AUC\(C\>\>B\) above 0\.5 means that this correctness\-labelled source direction ranks incorrect/self\-endorsed target responses above correct/self\-rejected responses\. Intervals condition on the fitted source direction\.
### S8\.3Window\-level null directions
Two nulls test whether large fixed\-window contrasts arise from the target quadrant structure alone\. The source\-label\-shuffle null independently permutes source OC/SJ labels before direction construction\. The random\-direction null replaces each fitted direction with a normalized random vector of the same dimensionality\. Each null uses 100 repetitions and is evaluated on the unchanged target conflict set\. Figure[S2](https://arxiv.org/html/2607.16799#S8.F2)and Table[S11](https://arxiv.org/html/2607.16799#S8.T11)compare their distributions with the observed fixed\-window effects\.
Figure S2:Window\-level null controls for Exp2B\. Points show the observedΔCB\\Delta\_\{\\mathrm\{CB\}\}; null distributions are obtained from 100 source\-label shuffles or random direction pairs\. The same target B/C rows and fixed normalized layer window are used for observed and null estimates\.Table S11:Numerical window\-level null summaries\. The empirical probability is the plus\-one Monte Carlo estimate\(b\+1\)/\(R\+1\)\(b\+1\)/\(R\+1\), wherebbis the number of valid null repetitions at least as large as the observed effect\.
### S8\.4B/C token\-count matching
Token\-count matching evaluates whether differences in complete rendered question\-plus\-response sequence length account for the separation between B and C responses\. Matching is performed within the target domain before AUC computation and does not alter source\-direction estimation\. The balance\-optimized procedure evaluates alternative token\-count coarsenings and selects the specification whose token\-count\-only SJ AUC is closest to 0\.5, breaking ties in favor of retaining more responses\. Standardized mean difference \(SMD\) is reported as an additional balance diagnostic\. A second specification uses fixed pooled token\-count deciles\.
Table[S12](https://arxiv.org/html/2607.16799#S8.T12)reports the number of retained responses, post\-matching balance, and the resulting fixed\-window contrasts\. The estimated contrasts remain positive under both matching procedures, including conditions in which achieving token\-count balance requires substantial trimming\.
Table S12:Token\-count matching diagnostics\. SMD is the standardized B/C mean difference in target rendered\-sequence token count\. Matching is fixed before inference; intervals resample target questions within the matched set\.
### S8\.5Response likelihood
Mean response log\-probability itself often discriminates OC and SJ on the conflict set, but it does not explain the projection asymmetry\. Table[S13](https://arxiv.org/html/2607.16799#S8.T13)reports raw and residualized effects after controlling each projection score for mean response log\-probability and response\-token count at each layer\. All eight residualΔCB\\Delta\_\{\\mathrm\{CB\}\}estimates remain positive with 95% intervals above zero\.
Table S13:Free\-response likelihood control\. “logp→\\rightarrowOC/SJ” is the AUC of mean assistant\-response token log\-probability alone on the same B/C rows\. Residual effects control projection scores for mean response log\-probability and response\-token count separately at each layer, followed by the same layer\-AUC averaging as the primary analysis\. Nuisance models are refitted in every bootstrap replicate\.
### S8\.6Question fixed effects
The all\-cell model in Eq\.[S15](https://arxiv.org/html/2607.16799#S5.E15)compares responses within question and does not require restricting the fit to B/C rows\. Across all eight transfers,βSJ\(meta\)\\beta^\{\(\\mathrm\{meta\}\)\}\_\{\\mathrm\{SJ\}\}exceedsβOC\(truth\)\\beta^\{\(\\mathrm\{truth\}\)\}\_\{\\mathrm\{OC\}\}and the cluster interval forΔFE\\Delta\_\{\\mathrm\{FE\}\}remains above zero \(Table[S14](https://arxiv.org/html/2607.16799#S8.T14)\)\. This result links the conflict\-set AUC pattern to the full2×22\\times 2design\.
The complete coefficient matrix in Table[S15](https://arxiv.org/html/2607.16799#S8.T15)shows that both transferred components have positive OC and SJ cross\-loadings\. In every direction, the SJ coefficient exceeds the OC coefficient for bothWmetaW\_\{\\mathrm\{meta\}\}andWtruthW\_\{\\mathrm\{truth\}\}\. The latter pattern is consistent with the conflict\-set reversal: a source correctness\-associated contrast need not retain OC specificity after transfer\. The factorial names identify how the source directions are constructed, not pure target\-domain latent axes\. The within\-component contrasts in Table[S16](https://arxiv.org/html/2607.16799#S8.T16)quantify this asymmetry\. All eightΔspec\(meta\)\\Delta^\{\(\\mathrm\{meta\}\)\}\_\{\\mathrm\{spec\}\}estimates are positive with 95% intervals above zero \(0\.192–0\.619\), whereas all eightΔspec\(truth\)\\Delta^\{\(\\mathrm\{truth\}\)\}\_\{\\mathrm\{spec\}\}estimates are negative with intervals below zero \(−0\.490\-0\.490–−0\.214\-0\.214\)\. Thus the transferredWmetaW\_\{\\mathrm\{meta\}\}direction retains SJ specificity, while the transferredWtruthW\_\{\\mathrm\{truth\}\}direction loads more strongly on SJ than OC\.
Table S14:Question\-fixed\-effect cross\-domain projection analysis\. Scores are standardized by component and target layer, then averaged over the fixed window\.Table S15:Complete question\-fixed\-effect cross\-loading matrix\. Panel A projects the sourceWmetaW\_\{\\mathrm\{meta\}\}direction; Panel B projects the sourceWtruthW\_\{\\mathrm\{truth\}\}direction\. Entries are coefficients with target\-question cluster\-bootstrap 95% confidence intervals\.Panel A:WmetaW\_\{\\mathrm\{meta\}\}projection
Panel B:WtruthW\_\{\\mathrm\{truth\}\}projection
Table S16:Within\-component specificity diagnostics for the all\-cell model\. Positive values indicate the source\-implied target specificity: SJ forWmetaW\_\{\\mathrm\{meta\}\}and OC forWtruthW\_\{\\mathrm\{truth\}\}\. Entries are coefficient differences with target\-question cluster\-bootstrap 95% confidence intervals\.
### S8\.7Judgement scoring and data\-construction robustness
The stored judgement score was originally computed from the returned next\-token candidates as described in Sec\. S4\. We recomputed the exact literalYesandNologits at the first assistant\-generation position for every pre\-filter pair row\. Because inference backends need not reproduce token logits bit\-for\-bit, the scoring\-rule audit holds these recomputed logits fixed and compares the historical top\-four fallback with full binary normalization,pjudge=sigmoid\(zYes−zNo\)p\_\{\\mathrm\{judge\}\}=\\operatorname\{sigmoid\}\(z\_\{\\mathrm\{Yes\}\}\-z\_\{\\mathrm\{No\}\}\)\. Across all eight model–domain cells, the maximum row\-wise difference was1\.30×10−61\.30\\times 10^\{\-6\}\. The two scoring rules produced no SJ label changes at 0\.5 and no complete\-pair membership changes under the primary pairedτ=0\.7\\tau=0\.7criterion\. The candidate\-return truncation therefore did not determine the analyzed strict samples\.
#### Strict\-pair resampling\.
The main analysis samples one correct and one incorrect response from each question’s usable pass@8 responses\. To measure sensitivity to this draw, we constructed a second Qwen2\.5\-7B pair sample from the frozen response pools using seed 2027\. Each correct and incorrect response was drawn uniformly from its corresponding pool; the originally selected response remained eligible, and repeated generation slots retained their empirical multiplicity\. At least one response changed in 2,441/2,703 Math pairs \(90\.3%\) and 645/1,101 Movies pairs \(58\.6%\)\. Full Yes/No normalization followed by the pairedτ=0\.7\\tau=0\.7criterion retained 2,289 Math pairs and 907 Movies pairs, with observations in all four factorial cells\.
All\-layerΔCB\\Delta\_\{\\mathrm\{CB\}\}profiles from the resampled pairs correlated with the original profiles at 0\.996 for Math→\\rightarrowMovies and 0\.991 for Movies→\\rightarrowMath\. In each direction, 23/28 layer\-wise estimates were positive\. The fixed\-window effects remained positive under 1,000 joint source–target question bootstrap resamples with source\-direction refitting \(Table[S17](https://arxiv.org/html/2607.16799#S8.T17)\)\. Thus the cross\-domain pattern is preserved under an independently repeated draw from the finite pass@8 response pools\.
Table S17:Qwen2\.5\-7B strict\-pair resampling sensitivity\. “Original” gives the main\-analysis window effect; R2 columns are computed from the second pair draw\. Entries are fixed\-window AUCs or differences with 95% percentile intervals from 1,000 joint source–target question bootstrap resamples\.
#### Multi\-reference Movies prompts\.
The 159 Movies prompt strings linked to more than one reference actor define a separate label\-ambiguity sensitivity\. We removed every retained response whose exact prompt belonged to this set before recomputing Exp2B\. For Math→\\rightarrowMovies this changes the target conflict set; for Movies→\\rightarrowMath it changes the source sample, and both factorial directions are refitted\. The filter removed 1\.38%–2\.41% of retained Movies rows across models and 6–39 B/C responses\.
All eight filtered fixed\-window effects remained positive with 95% intervals above zero \(Table[S18](https://arxiv.org/html/2607.16799#S8.T18)\)\. Their all\-layer profiles correlated with the original profiles atr≥0\.999r\\geq 0\.999, and the largest absolute change in a window effect was 0\.030\. The weakest original condition, Llama\-3\.1\-8B Movies→\\rightarrowMath, changed from 0\.138 to 0\.140 \[0\.075, 0\.204\]\. The cross\-domain result is therefore unchanged when multi\-reference Movies prompts are excluded rather than handled only through question clustering\.
Table S18:Movies multi\-reference prompt sensitivity\. All retained responses from prompts associated with multiple reference actors are removed before direction fitting or target evaluation\. Filtered entries are fixed\-windowΔCB\\Delta\_\{\\mathrm\{CB\}\}estimates with 95% percentile intervals from 1,000 joint source–target question bootstrap resamples with direction refitting\.
## S9Self\-Judgement Threshold Sensitivity
The primary strict analysis usesτ=0\.7\\tau=0\.7at the pair level\. A complementary row\-level sensitivity analysis instead reconstructs the pre\-filter population by joining the original activation files with a supplementary activation patch for rows that had been removed atτ=0\.7\\tau=0\.7\. For eachτ∈\{0\.5,0\.6,0\.7,0\.8\}\\tau\\in\\\{0\.5,0\.6,0\.7,0\.8\\\}, it removes the uncertain band1−τ<pjudge<τ1\-\\tau<p\_\{\\mathrm\{judge\}\}<\\tauand assigns SJ from the retained side\. This is a row\-filter sensitivity population; the primary sample instead requires both members of a question pair to passτ=0\.7\\tau=0\.7\.
Figure[S3](https://arxiv.org/html/2607.16799#S9.F3)shows why threshold changes affect models differently\. Qwen distributions are highly polarized, whereas OLMo assigns much more mass near 0\.5\. Atτ=0\.8\\tau=0\.8, only 14\.0% of OLMo Math rows remain, so that setting is a support stress test rather than an equally powered analysis\. Table[S19](https://arxiv.org/html/2607.16799#S9.T19)gives the complete retained quadrant counts\.
Figure S3:Pre\-filter distributions ofpjudgep\_\{\\mathrm\{judge\}\}for every model and main domain\. Vertical references mark the lower and upper boundaries induced by the symmetric confidence thresholds\. These probabilities precede the main paired strict filter\.Table S19:Row\-level support under symmetric confidence filtering\. “Kept” is the number of pre\-filter judged rows outside the uncertain band\.The sensitivity run uses 100 question\-cluster bootstrap resamples\. Across all 280 Exp2B layer\-direction rows, positiveΔCB\\Delta\_\{\\mathrm\{CB\}\}occurs in 228, 239, 252, and 251 rows atτ=0\.5,0\.6,0\.7,\\tau=0\.5,0\.6,0\.7,and 0\.8, respectively; the corresponding intervals lie above zero in 178, 193, 196, and 205 rows\. The cross\-domain pattern therefore persists across alternative confidence filters\. Increasingτ\\tauremoves more rows near the decision boundary and generally strengthens separation, while reducing support most sharply for OLMo\.
## S10Zero\-Target\-Fitting OOD Analyses
Source directions are fitted only on Math or Movies strict samples and applied unchanged to MMLU and binary TruthfulQA\. OOD records retain only high\-confidence B/C candidates under the same symmetricτ=0\.7\\tau=0\.7rule\. Because several candidates can originate from one question, all intervals cluster by the original target question\.
Table[S20](https://arxiv.org/html/2607.16799#S10.T20)reports support and the full layer\-by\-source descriptive counts\. PositiveΔCB\\Delta\_\{\\mathrm\{CB\}\}appears in 458/560 \(81\.8%\) rows, with question\-cluster intervals above zero in 318/560 \(56\.8%\)\. The main OOD figure’s component curves show that this difference is driven byWmeta→SJW\_\{\\mathrm\{meta\}\}\\rightarrow\\mathrm\{SJ\}rising above chance whileWtruth→OCW\_\{\\mathrm\{truth\}\}\\rightarrow\\mathrm\{OC\}remains lower on the same target conflicts\. The result occurs for both a four\-option knowledge benchmark and a binary truthfulness benchmark, although effect magnitude varies by model, source, and target\.
Table S20:OOD support and all\-layer heterogeneity\. “Rows” in the final two columns are source\-by\-layer evaluations: two source directions per model\-target pair times the number of model layers\.### S10\.1Forced\-choice likelihood and answer\-letter controls
Forced\-choice answer likelihood is measured as the teacher\-forced probability of the inserted option letter\. Within the B/C conflict set, its OC and SJ AUCs are complementary by construction, but their direction and magnitude vary across models and target datasets\. To control for candidate likelihood and option identity, each layer\-specific projection score is residualized with respect to option log\-probability, token count, and answer\-letter indicators\. The nuisance regression is re\-estimated within each bootstrap replicate\. Table[S21](https://arxiv.org/html/2607.16799#S10.T21)reports the resulting contrast\. Fifteen of sixteen letter\-controlled estimates remain positive with intervals above zero\. The exception is OLMo Movies→\\rightarrowTruthfulQA,ΔCB=−0\.030\\Delta\_\{\\mathrm\{CB\}\}=\-0\.030\[\-0\.138, 0\.077\], whose 95% CI includes zero\.
Table S21:OOD answer\-likelihood and letter control\. RawΔCB\\Delta\_\{\\mathrm\{CB\}\}is the component difference before residualization; the final column controls answer\-option log\-probability, target token count, and answer letter\. Binary TruthfulQA has one non\-reference letter indicator; MMLU has three\.
### S10\.2Counterbalanced target\-side judgement labels
The primary SJ label uses a Yes/No response\. To test whether the OOD result is tied to assent or a fixed response\-token mapping, we re\-elicited target\-side judgements with two prompts:
> Evaluate whether the answer above is correct\. Respond with X if it is correct and Y if it is incorrect\. Evaluate whether the answer above is correct\. Respond with Y if it is correct and X if it is incorrect\.
X/Y avoids reusing the A–D labels already assigned to forced\-choice answers\. For each candidate, we compute correctness\-oriented log odds
zXY=logP\(X\)−logP\(Y\),zYX=logP\(Y\)−logP\(X\),z\_\{XY\}=\\log P\(X\)\-\\log P\(Y\),\\qquad z\_\{YX\}=\\log P\(Y\)\-\\log P\(X\),and define
zbal=12\(zXY\+zYX\),pbal=sigmoid\(zbal\)\.z\_\{\\mathrm\{bal\}\}=\\tfrac\{1\}\{2\}\(z\_\{XY\}\+z\_\{YX\}\),\\qquad p\_\{\\mathrm\{bal\}\}=\\operatorname\{sigmoid\}\(z\_\{\\mathrm\{bal\}\}\)\.The arithmetic mean of the two mapping\-corrected probabilities is stored only as a secondary diagnostic and is not used to define the reported labels or transfer results\. The primary analysis applies the same symmetricτ=0\.7\\tau=0\.7rule topbalp\_\{\\mathrm\{bal\}\}\. Because this changes which candidates belong to B and C, answer\-token activations were extracted for the complete new conflict set at the same forced\-answer final\-token site\. Math\- and Movies\-fitted directions, layer windows, and target\-question bootstrap clusters remain unchanged\.
Mapping\-corrected XY and YX scores correlate strongly in every model–target cell \(ρ=0\.917\\rho=0\.917–0\.9750\.975\), and their jointly high\-confidence binary labels agree in 92\.7%–100% of candidates \(Table[S22](https://arxiv.org/html/2607.16799#S10.T22)\)\. The balanced MMLU conflict sets are sparse on B for Llama\-3\.1\-8B \(nB=13n\_\{B\}=13\) and OLMo\-3\-7B \(nB=4n\_\{B\}=4\); those cells consequently provide weak raw inferential support despite having many C rows\.
Table S22:Counterbalanced target\-side judgement diagnostics\.ρXY,YX\\rho\_\{XY,YX\}is the Spearman correlation between the two mapping\-corrected log odds\. HC agreement is computed where both mappings yield a high\-confidence binary label\. Retained counts all candidates outside the symmetric uncertain band; B/C and conflict questions describe the balanced conflict set\.Across the 16 model–source–target rows, fixed\-windowΔCB\\Delta\_\{\\mathrm\{CB\}\}is positive in all cases \(mean 0\.383, range 0\.176–0\.682\), with 12 raw question\-cluster intervals above zero \(Table[S23](https://arxiv.org/html/2607.16799#S10.T23)\)\. The four intervals crossing zero are exactly the two source directions evaluated in each of the low\-B MMLU cells\. After layer\-wise adjustment for answer\-option log\-probability and token count, all 16 point estimates and intervals are above zero\. Across individual layers, 464/560 \(82\.9%\) estimates are positive and 282/560 \(50\.4%\) intervals are above zero\. This control supports target\-side judgement\-label robustness; the source directions were still constructed from the original elicitation\.
Table S23:Fixed\-window OOD transfer under counterbalanced X/Y judgement labels\. Component AUCs and rawΔCB\\Delta\_\{\\mathrm\{CB\}\}use the complete balanced B/C sets\. ResidualΔCB\\Delta\_\{\\mathrm\{CB\}\}controls answer\-option log\-probability and token count separately at each layer\. All intervals use 1,000 target\-question\-cluster bootstrap resamples\.Figure S4:All\-layer component AUCs under counterbalanced X/Y target judgement labels\. Directions are fitted only on Math or Movies\. Blue curves showWmeta→W\_\{\\mathrm\{meta\}\}\\\!\\rightarrowSJ and brown\-orange curves showWtruth→W\_\{\\mathrm\{truth\}\}\\\!\\rightarrowOC; shading denotes 95% target\-question cluster\-bootstrap intervals\. The broad intervals for Llama\-3\.1\-8B and OLMo\-3\-7B on MMLU reflect conflict sets with only 13 and 4 B rows, respectively\.
## S11Implementation and Reproducibility
### S11\.1Models and activation artifacts
The analyses use four publicly available instruction\-tuned language models:
- •Qwen2\.5\-7B\-Instruct \(Qwen/Qwen2\.5\-7B\-Instruct\);
- •Llama\-3\.1\-8B\-Instruct \(meta\-llama/Llama\-3\.1\-8B\-Instruct\);
- •Qwen2\.5\-14B\-Instruct \(Qwen/Qwen2\.5\-14B\-Instruct\); and
- •OLMo\-3\-7B\-Instruct \(allenai/Olmo\-3\-7B\-Instruct\)\.
Activation extraction was performed on a single NVIDIA A800 80GB using bfloat16 forward computation and float16 activation storage, with a maximum sequence length of 4,096 tokens\. The batch size was four for the 7–8B models in the main\-domain extraction and two for Qwen2\.5\-14B\. OOD activation extraction used a batch size of two\.
Table S24:Main\-domain activation arrays\. Shapes are samples by transformer blocks by hidden width and apply to each sample\-aligned artifact\.Each saved artifact contains metadata, a JSONL sample index, and one or more NPZ activation arrays\. The sample index records the activation\-row index, question identifier, response, reference answer, OC label, SJ label, factorial cell,pjudgep\_\{\\mathrm\{judge\}\}, rendered question\-plus\-response token count, and prompt hash\. Probe fitting, bootstrap inference, and figure generation are performed offline from these saved arrays and do not require additional model inference\.
## S12Interpretive Scope
The factorial contrasts identify response\-level associations with the observed OC and SJ cells; neither OC nor SJ is experimentally manipulated\. Confidence intervals for the primary fixed\-window cross\-domain analysis jointly incorporate variation in source\- and target\-domain question samples, conditional on the realized pass\-at\-eight generations and the sampled correct–incorrect response pair for each eligible question\. They do not integrate generation\-level or within\-question pair\-selection variability\. In contrast, the all\-layer figures and several secondary controls condition on the fitted source direction and therefore quantify only target\-sample uncertainty\.
The OOD tasks use inserted forced\-choice responses rather than free generation\. They consequently test whether the fitted directions transfer to a controlled answer site, not whether the complete free\-response generation pipeline replicates across task formats\. In addition, SJ operationalizes correctness\-directed self\-evaluation under the specified elicitation procedure\. It may contain variance associated with confidence, familiarity, response policy, or other processes beyond metacognitive accuracy\.
Subject to these constraints, the principal pattern is observed across conflict\-set ordering, grouped in\-domain validation, cross\-domain component AUCs, source\-side within\-question direction estimation, the all\-cell target\-question\-fixed\-effect analysis, token\-count matching, null directions, alternative confidence thresholds, answer\-likelihood controls, and two target tasks excluded from direction fitting\.Similar Articles
The Knowing-Saying Gap: When Probes See Errors that Confidence Misses
This paper investigates the 'knowing-saying gap' in language models, showing that linear probes can detect corrupted context with near-perfect accuracy yet fail to predict final answer errors, with implications for deployment monitoring and intervention strategies.
Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection
This paper proposes a nuisance-controlled residual-stream probing protocol to detect benchmark contamination in language models, showing that naive probing fails and their corrected method controls false positives while achieving reasonable power.
Making LLMs tell you how confident they really are through probe-targeted fine tuning.[R]
This research presents probe-targeted fine-tuning (LoRA) to make LLMs verbally express their internal confidence, achieving causal control over confidence outputs and demonstrating that models often know when they are right or wrong but fail to articulate it.
What Predicts Correctness in Text-to-SQL? A Selective-Prediction Study
This paper studies which signals best predict correctness in text-to-SQL for selective prediction. It finds that verification-based signals from LLM judges outperform black-box statistical signals like self-consistency, and that a two-provider ensemble achieves 0.82 AUROC with well-calibrated probabilities.
Linear Probes Detect Task Format, Not Reasoning Mode in Language Model Hidden States
This paper demonstrates that linear probes on LLM hidden states detect task format confounds (e.g., source identity, response length) rather than distinct reasoning modes, using residualization and causal steering to show that high probe accuracy is due to superficial features, not computational structure.