The Hallucination Signal Is a Mean Shift: Why Simple Probes Suffice
Summary
The paper shows that the hallucination signal in LLMs is dominated by a mean shift, making simple probes like L2-regularized logistic regression sufficient and outperforming complex alternatives.
View Cached Full Text
Cached at: 09/01/26, 12:11 PM
# The Hallucination Signal Is a Mean Shift: Why Simple Probes Suffice
Source: [https://arxiv.org/html/2608.28930](https://arxiv.org/html/2608.28930)
Jaehyung Seo††thanks:Corresponding authors\.Affiliation:Konkuk UniversityEmail:[limhseok@korea\.ac\.kr](mailto:)Heuiseok Lim11footnotemark:1Affiliation:Korea UniversityEmail:[seojae777@konkuk\.ac\.kr](mailto:)
###### Abstract
Hidden\-state probes effectively detect LLM hallucinations, but the geometry of the signal remains poorly characterized, driving increasingly complex probe architectures\. Across three 7B\-scale models and three datasets in a paired\-example paradigm, we find the signal overwhelmingly dominated by a single mean\-shift component, and removing this direction collapses detection to chance\. Shrinkage linear discriminant analysis closes about 73% of the gap between 1D and full\-dimensional classifiers, so apparent architectural complexity largely reflects high\-dimensional covariance estimation difficulty rather than exploitable non\-linearity\. A simple L2\-regularized logistic regression \(0\.952 AUROC\) bounds or outperforms twelve controlled architectural alternatives, and our multi\-layer aggregation exceeds CLAP cross\-layer attention probing under matched paradigm\. Because the signal spans a contiguous layer band, LayerMix aggregates it to match oracle\-layer performance without oracle access\. Our claims characterize the geometry within the controlled paired\-example paradigm\. Our code is available at[https://github\.com/js\-lee\-AI/LayerMix](https://github.com/js-lee-AI/LayerMix)\.
## 1Introduction
While hidden\-state probing has emerged as an effective mechanism for detecting hallucinations in large language models \(LLMs\)\([Bürger et al\., 2024](https://arxiv.org/html/2608.28930#bib.bib9);[Hernandez et al\., 2023](https://arxiv.org/html/2608.28930#bib.bib19);[Han et al\., 2025](https://arxiv.org/html/2608.28930#bib.bib18);[Azaria and Mitchell, 2023](https://arxiv.org/html/2608.28930#bib.bib2)\), recent work has pursued increasingly complex architectures to isolate the truthfulness signal\([Luo et al\., 2026](https://arxiv.org/html/2608.28930#bib.bib34);[Bürger et al\., 2024](https://arxiv.org/html/2608.28930#bib.bib9);[Bang et al\., 2025](https://arxiv.org/html/2608.28930#bib.bib3)\)\. Methods such as ICR Probe\([Zhang et al\., 2025](https://arxiv.org/html/2608.28930#bib.bib46)\), TSV\([Park et al\., 2025](https://arxiv.org/html/2608.28930#bib.bib39)\), and HaloScope\([Du et al\., 2024](https://arxiv.org/html/2608.28930#bib.bib13)\)deploy intricate cross\-layer tracking, optimal\-transport pseudo\-labeling, and attention\-gated subspace projections\. Yet, a fundamental question remains under\-explored:*what is the exact geometry of the hallucination signal, and does it necessitate this architectural complexity?*
To answer this, we conduct a systematic geometric analysis of the hallucination signal\. To rigorously isolate the underlying representational structure from generation\-induced distribution shifts, we operate within the standard paired\-example paradigm across three 7B\-scale models and three datasets\. Although prior work established that truth is encoded linearly\([Marks and Tegmark, 2023](https://arxiv.org/html/2608.28930#bib.bib36);[Bao et al\., 2025](https://arxiv.org/html/2608.28930#bib.bib4)\), the quantitative boundaries of this linearity for hallucination detection remain unmapped\. Our analysis reveals a remarkably simple underlying structure: the detection signal is overwhelmingly dominated by a single mean\-shift component\. By disentangling covariance estimation difficulty from genuine non\-linear structure, we demonstrate that a simple L2\-regularized logistic regression \(L2\-LR\) probe matches or outperforms 12 heavily engineered alternatives\. This suggests that complex architectures often overfit to estimation noise rather than exploiting hidden structural complexities\.
Crucially, we observe that this simple mean\-shift signal is not isolated to a single oracle layer, but rather distributed across a model\-specific, contiguous layer band\([Meng et al\., 2022](https://arxiv.org/html/2608.28930#bib.bib37);[Patrawala et al\., 2025](https://arxiv.org/html/2608.28930#bib.bib40);[Huang et al\., 2025](https://arxiv.org/html/2608.28930#bib.bib22)\)\. Motivated directly by this geometric insight, we operationalize our findings into LayerMix, a computationally lightweight, cross\-validation\-based multi\-layer aggregation approach\. LayerMix is not presented as a complex algorithmic innovation, but rather as a principled consequence of our analysis: it exploits the distributed nature of the signal to match oracle\-layer performance without requiring oracle access, offering a robust, highly effective drop\-in baseline for future geometric research\.
Our contributions are as follows:
- •Quantitative Geometric Decomposition:We provide the first quantitative characterization of the hallucination signal’s geometry, demonstrating that a dominant mean\-shift accounts for the vast majority of the signal\. We show that shrinkage linear discriminant analysis \(LDA\) closes about 73% of the Fisher LDA gap, indicating that apparent structural complexity largely stems from high\-dimensional covariance estimation difficulty rather than exploitable non\-linearities\.
- •Controlled Geometric Ablation:To rigorously validate our geometric framework, we evaluate L2\-LR against 12 controlled architectural alternatives designed to isolate specific geometric properties, alongside complex subspace methods\. As predicted by our decomposition, L2\-LR strictly upper\-bounds their performance, achieving 0\.952 AUROC and demonstrating that mean\-targeting methods consistently dominate variance\-targeting ones\.
- •LayerMix:We introduce a principled, oracle\-free multi\-layer aggregation strategy\. LayerMix seamlessly identifies and aggregates the informative layer band, matching oracle performance \(0\.954 AUROC\) while remaining exceptionally computationally efficient \(about 35 seconds overhead\)\.
The paired\-example paradigm is the controlled experimental setup we use to isolate the intrinsic representational geometry from generation\-induced distribution shifts, the methodological control that makes the geometric measurement possible\. The conclusions reported throughout \(mean\-shift dominance, decision\-boundary linearity, multi\-layer aggregation sufficiency\) are paradigm\-bounded statements about the controlled measurement setting, not universality claims across all detection scenarios\.
## 2Related Work
### 2\.1Detection Methods and Evaluation
#### Training\-free detection\.
Methods requiring no labeled data span logit contrasting\([Chuang et al\., 2023](https://arxiv.org/html/2608.28930#bib.bib12)\), covariance eigenvalues\([Chen et al\., 2024](https://arxiv.org/html/2608.28930#bib.bib11)\), semantic entropy\([Kuhn et al\., 2023](https://arxiv.org/html/2608.28930#bib.bib26)\), and sampling consistency\([Manakul et al\., 2023](https://arxiv.org/html/2608.28930#bib.bib35)\), with LLM\-Polygraph\([Fadeeva et al\., 2023](https://arxiv.org/html/2608.28930#bib.bib15)\)providing a comprehensive benchmark\. The limited accuracy of these approaches motivates our focus on supervised probing\.
#### Probing\-based detection\.
Following[Azaria and Mitchell \(2023\)](https://arxiv.org/html/2608.28930#bib.bib2)’s foundational logistic regression probe, recent methods have introduced increasingly complex architectures, including attention\-gated SVD\([Du et al\., 2024](https://arxiv.org/html/2608.28930#bib.bib13)\), cross\-layer tracking\([Zhang et al\., 2025](https://arxiv.org/html/2608.28930#bib.bib46)\), and optimal\-transport pseudo\-labeling\([Park et al\., 2025](https://arxiv.org/html/2608.28930#bib.bib39)\)\. Crucially, these single\-layer methods lack a geometric characterization of the signal they target\. Our analysis fills this gap, demonstrating that much of this architectural complexity is unwarranted when using L2\-regularized logistic regression\.
#### Evaluation paradigms\.
SEP\([Kossen et al\., 2024](https://arxiv.org/html/2608.28930#bib.bib25)\)trains probes to predict semantic entropy from multiple sampled generations\. This open\-ended QA paradigm differs fundamentally from the paired\-example setting evaluated here \(and in SAPLMA, ICR Probe, and HaloScope\), precluding direct head\-to\-head comparison\. We discuss SEP’s cross\-domain implications in §[7](https://arxiv.org/html/2608.28930#S7)\.
### 2\.2Geometry and Layer Choice
#### Geometry of LLM representations\.
Prior work has identified linear structures for truthfulness\([Marks and Tegmark, 2023](https://arxiv.org/html/2608.28930#bib.bib36);[Bao et al\., 2025](https://arxiv.org/html/2608.28930#bib.bib4);[Burns et al\., 2022](https://arxiv.org/html/2608.28930#bib.bib10)\)and world representations\([Li et al\., 2022](https://arxiv.org/html/2608.28930#bib.bib30)\)\. Representation engineering\([Zou et al\., 2023](https://arxiv.org/html/2608.28930#bib.bib47)\)shows that projecting along such directions steers model behavior, a principle our intervention experiment \(§[7](https://arxiv.org/html/2608.28930#S7)\) applies to hallucinations\. While HARP\([Hu et al\., 2025](https://arxiv.org/html/2608.28930#bib.bib21)\)derives hallucination subspaces from unembedding weights, we extend this trajectory by decomposing the detection signal into mean\-shift and residual components, quantifying the Fisher LDA gap, and linking geometry to probe performance\.
#### Layer selection\.
While probing accuracy heavily depends on the chosen hidden layer\([Azaria and Mitchell, 2023](https://arxiv.org/html/2608.28930#bib.bib2);[Han et al\., 2025](https://arxiv.org/html/2608.28930#bib.bib18);[Belinkov, 2022](https://arxiv.org/html/2608.28930#bib.bib6)\), the exact cost of relying on non\-oracle layer heuristics remains unquantified for hallucination detection\. We systematically measure this degradation and demonstrate that cross\-validation\-based aggregation eliminates it\.
### 2\.3Concurrent Work
#### Cross\-layer and subspace probing\.
The closest concurrent work to LayerMix is CLAP\([Suresh et al\., 2025](https://arxiv.org/html/2608.28930#bib.bib44)\), which learns input\-dependent cross\-layer attention, whereas LayerMix selects a layer band by cross\-validation and uniformly averages independent per\-layer probes\. We run CLAP from its official implementation on our cached multi\-layer hidden states, so the contrast isolates the combination rule, not the features\.[Bürger et al\. \(2024\)](https://arxiv.org/html/2608.28930#bib.bib9)argue instead that hallucination and truth occupy a low\-dimensional subspace, which we test with partial least squares discriminant analysis \(PLS\-DA\) atk=2k\{=\}2against full\-dimensional probes\. Our shrinkage\-LDA\-gap analysis predicts that the residual beyond the mean shift needs high\-dimensional shrinkage, not two\-dimensional truncation, and §[6](https://arxiv.org/html/2608.28930#S6)reports both comparisons under matched paradigm\.
#### Findings in other paradigms\.
[Bhatnagar et al\. \(2026\)](https://arxiv.org/html/2608.28930#bib.bib7)report SOTA on free\-form generation with cross\-dataset generalization, which is not in tension with our orthogonal\-mean\-shift finding, because DRIFT operates in the dynamic\-generation paradigm, orthogonal to our paired\-example setting, where reweighting across directions can succeed even when raw paired\-example𝜹^\\hat\{\\boldsymbol\{\\delta\}\}are orthogonal across datasets\.[Liang and Wang \(2025\)](https://arxiv.org/html/2608.28930#bib.bib31)report MLP probes outperforming linear probes in token\-level free\-form detection, which we likewise read as paradigm\-dependence rather than a contradiction\.
#### Generalization and controls\.
[Orgad et al\. \(2025\)](https://arxiv.org/html/2608.28930#bib.bib38)document that representation\-based detectors do not generalize zero\-shot across datasets\. Our within\-vs\-between cosine analysis \(Appendix[F](https://arxiv.org/html/2608.28930#A6.SSx1)\) suggests this reflects hallucination\-type\-specific geometry rather than pure dataset artifacts\. Our random\-label diagnostic follows probing\-classifier control tasks\([Hewitt and Liang, 2019](https://arxiv.org/html/2608.28930#bib.bib20)\), formalizing selectivity for hallucination detection\.
## 3Geometric Analysis
We characterize the hallucination signal’s geometry to motivate LayerMix\. Given an LLM with hidden dimensionddand labeled examples\{\(xi,yi\)\}\\\{\(x\_\{i\},y\_\{i\}\)\\\}whereyi∈\{0,1\}y\_\{i\}\\in\\\{0,1\\\}indicates factual or hallucinated, let𝝁0,𝝁1\\boldsymbol\{\\mu\}\_\{0\},\\boldsymbol\{\\mu\}\_\{1\}be class centroids and𝜹=𝝁1−𝝁0\\boldsymbol\{\\delta\}=\\boldsymbol\{\\mu\}\_\{1\}\-\\boldsymbol\{\\mu\}\_\{0\}the mean\-shift direction\.
### 3\.1The Signal is a Mean Shift
Table[1](https://arxiv.org/html/2608.28930#S3.T1)provides direct evidence from decomposition experiments averaged over 3 models×\\times3 datasets\. Removing𝜹\\boldsymbol\{\\delta\}collapses detection to chance \(0\.499\), confirming it is*necessary*\. Crucially,𝜹\\boldsymbol\{\\delta\}is estimated within each training fold of the CV protocol, preventing in\-sample estimation artifacts\.
The 1D projection onto𝜹\\boldsymbol\{\\delta\}achieves 0\.834 AUROC, recovering 78% of the above\-chance performance of unregularized LR\. The gap between this and the operational L2\-LR baseline \(0\.952\) highlights the regularization benefit: L2 shrinkage suppresses noise in the roughly 4,000 residual dimensions, extracting signal that unregularized LR cannot reliably exploit\. Note that PLS\-DA component 1 aligns with𝜹\\boldsymbol\{\\delta\}by mathematical construction in the binary case\([Barker and Rayens, 2003](https://arxiv.org/html/2608.28930#bib.bib5)\), confirming that PLS\-DA inherently targets the mean shift\.
Cohen’sddalong𝜹\\boldsymbol\{\\delta\}ranges 1\.2–1\.6 despite accounting for<3%\{<\}3\\%of total variance\. Figure[1](https://arxiv.org/html/2608.28930#S3.F1)visually corroborates this: the clear macro\-separation along the mean\-shift axis contrasts sharply with the completely overlapping distributions in the residual orthogonal subspace, explaining why linear probes successfully recover the vast majority of the signal\.
Figure 1:2D projection of factual and hallucinated hidden states from Qwen2\.5\-7B \(Layer 18\) on TruthfulQA\.Table 1:Mean\-shift decomposition \(avg\. 9 conditions\)\. “Unreg\. LR” denotes unregularized logistic regression, used to enable clean decomposition into mean\-shift and residual components without regularization\-induced shrinkage\.
### 3\.2Mean Shift is Necessary but Insufficient
The mean\-shift direction is essential but captures only part of the signal\.
#### Fisher LDA gap\.
In the binary case, Fisher LDA projects onto𝒘F=ΣW−1𝜹\\boldsymbol\{w\}\_\{F\}=\\Sigma\_\{W\}^\{\-1\}\\boldsymbol\{\\delta\}, which reduces to𝜹\\boldsymbol\{\\delta\}only when the within\-class covarianceΣW\\Sigma\_\{W\}is proportional to the identity\. In high dimensions \(d≈4,000d\\approx 4\{,\}000\),ΣW\\Sigma\_\{W\}is far from spherical, so𝒘F\\boldsymbol\{w\}\_\{F\}and𝜹\\boldsymbol\{\\delta\}differ\. We report projection onto the raw mean\-shift direction𝜹\\boldsymbol\{\\delta\}to provide a clean lower bound on the 1D signal without requiring covariance inversion\.
Table[2](https://arxiv.org/html/2608.28930#S3.T2)shows that this optimal 1D classifier underperforms L2\-regularized LR by 0\.09–0\.16 AUROC\. This consistent gap confirms the presence of discriminative structure beyond𝜹\\boldsymbol\{\\delta\}in every condition\. To disentangle covariance estimation difficulty from genuine distributional structure, we apply shrinkage LDA\([Ledoit and Wolf, 2004](https://arxiv.org/html/2608.28930#bib.bib27)\), which achieves 0\.920 mean AUROC, closing at least 73% of the gap\. The residual 0\.032 gap may reflect non\-Gaussian structure or L2\-regularization’s implicit feature selection\. Applying ZCA whitening \(Σ−1/2\\Sigma^\{\-1/2\}\) before L2\-LR degraded performance \(0\.790 vs\. 0\.952 raw\), as expected when amplifying low\-variance noise in finite high\-dimensional samples\.
#### Distributional evidence\.
D’Agostino–Pearson tests reject normality \(p<10−4p<10^\{\-4\}\) for both classes in all 9 conditions, with skewness\|γ1\|≤0\.55\|\\gamma\_\{1\}\|\\leq 0\.55and excess kurtosis−1\.1\-1\.1to\+0\.2\+0\.2\. This is consistent with the residual gap reflecting distributional structure beyond what LDA can capture\.
Table 2:Raw mean\-shift projection \(1D, onto𝜹\\boldsymbol\{\\delta\}\) vs\. L2\-regularized LR on the optimal layer across all 9 conditions\.
### 3\.3Signal Distributes Across Layers
Beyond characterizing the signal within a single layer, we examine how it distributes*across*layers to motivate multi\-layer probing:
#### Optimal layer varies\.
The best single layer differs by model: L14 for Llama\([Touvron et al\., 2023](https://arxiv.org/html/2608.28930#bib.bib45);[Grattafiori et al\., 2024](https://arxiv.org/html/2608.28930#bib.bib17)\), L16 for Mistral\([Jiang et al\., 2023](https://arxiv.org/html/2608.28930#bib.bib23)\), L18 for Qwen\([Qwen et al\., 2025](https://arxiv.org/html/2608.28930#bib.bib41)\)\(Figure[2](https://arxiv.org/html/2608.28930#S3.F2); Table[7](https://arxiv.org/html/2608.28930#A2.T7)\)\. This variation means a fixed\-layer heuristic cannot be optimal across conditions\. As Figure[2](https://arxiv.org/html/2608.28930#S3.F2)illustrates, the hallucination signal instead concentrates in a model\-specific contiguous band\.
#### Adjacent layers carry signal\.
Cross\-validation scores for layers near the optimum remain high\. For example, on Qwen–TQA \(optimal L18\), CV selects layers\{17,18,19,20,21\}\\\{17,18,19,20,21\\\}, each scoring≥\\geq0\.93\.
#### Signal distributes across neurons\.
L1\-regularized feature selection followed by L2\-LR \(Appendix[H](https://arxiv.org/html/2608.28930#A8)\) shows that 200 neurons \(about 5% ofdd\) recover 98\.3% of full AUROC, while 100 random neurons achieve only 0\.840\. The top\-20 neurons share only 4 across CV folds, confirming a distributed signal not localized to fixed "hallucination neurons\."
#### Combining layers improves detection\.
Averaging predictions from CV\-selected layers consistently outperforms any single layer \(by \+0\.0004 to \+0\.0086 AUROC across all 9 conditions\), indicating systematic complementarity\.
These findings establish that the hallucination signal distributes across both layers and neurons, motivating multi\-layer aggregation of full\-dimensional probes\. The geometry also correctly predicts the performance ordering of subspace methods \(Table[16](https://arxiv.org/html/2608.28930#A12.T16)in Appendix[L](https://arxiv.org/html/2608.28930#A12)\): methods targeting the mean shift \(PLS\-DA: 0\.926\) consistently outperform variance\-targeting ones \(SVD: 0\.795\)\.
Figure 2:Layer\-wise AUROC across 3 models and 3 datasets\. Shaded bands show±1\{\\pm\}1std across CV folds\. Dots mark the optimal layer per condition\.
## 4Proposed Method: LayerMix
Our geometric analysis reveals that the hallucination signal distributes across a contiguous band of layers \(§[3\.3](https://arxiv.org/html/2608.28930#S3.SS3)\), with the optimal layer varying by model and dataset\. This motivates LayerMix: a principled, computationally lightweight approach that identifies the informative layer band via cross\-validation \(CV\), trains independent probes on the top\-KKlayers, and averages their predictions\. The method operates in three stages\.
#### Stage 1: Layer scoring\.
For each layerℓ∈\{1,…,L\}\\ell\\in\\\{1,\\ldots,L\\\}, we extract hidden states𝐇\(ℓ\)∈ℝN×d\\mathbf\{H\}^\{\(\\ell\)\}\\in\\mathbb\{R\}^\{N\\times d\}and evaluate an L2\-regularized logistic regression probe via stratifiedmm\-fold CV \(m=5m\{=\}5\)\. The scoresℓs\_\{\\ell\}is the mean AUROC across folds\. This step replaces oracle selection: rather than requiring a held\-out set to identify the best layer post hoc, CV scores provide a data\-driven ranking using only training data\.
#### Stage 2: Layer selection\.
We select the top\-KKlayers:ℒ∗=top\-K\(\{sℓ\}ℓ=1L\)\\mathcal\{L\}^\{\*\}=\\text\{top\-\}K\(\\\{s\_\{\\ell\}\\\}\_\{\\ell=1\}^\{L\}\)\. We useK=5K\{=\}5as the default, finding it robust across conditions \(Appendix[M](https://arxiv.org/html/2608.28930#A13)\)\. In practice, the selected layers naturally form a contiguous block around the model’s most informative region \(e\.g\., layers 16–20 for Qwen2\.5\-7B\), as adjacent layers share representational structure while providing marginal complementary signal\.
#### Stage 3: Aggregation\.
For each selected layerℓ∈ℒ∗\\ell\\in\\mathcal\{L\}^\{\*\}, we train an L2\-regularized LR probefℓf\_\{\\ell\}on the full training set\. The final prediction for a test inputxxis:
p^\(x\)=1K∑ℓ∈ℒ∗fℓ\(𝐡\(ℓ\)\(x\)\)\\hat\{p\}\(x\)=\\frac\{1\}\{K\}\\sum\_\{\\ell\\in\\mathcal\{L\}^\{\*\}\}f\_\{\\ell\}\(\\mathbf\{h\}^\{\(\\ell\)\}\(x\)\)\(1\)
The simplicity of this formulation is deliberate: each per\-layer probe operates on the fulldd\-dimensional hidden state with strong L2 regularization \(C=0\.001C\{=\}0\.001\), and uniform averaging provides implicit variance reduction\. No learned aggregation weights are needed, as the selected layers are high\-quality by construction\.
Algorithm 1LayerMix0:LLM
ℳ\\mathcal\{M\}, labeled data
𝒟=\{\(xi,yi\)\}\\mathcal\{D\}=\\\{\(x\_\{i\},y\_\{i\}\)\\\},
K=5K\{=\}5
0:Detector
f:x→\[0,1\]f:x\\to\[0,1\]
1:// Stage 1: Layer Scoring \(replaces oracle selection\)
2:for
ℓ=1,…,L\\ell=1,\\ldots,Ldo
3:
𝐇\(ℓ\)←\\mathbf\{H\}^\{\(\\ell\)\}\\leftarrowhidden states from layer
ℓ\\ell
4:
sℓ←s\_\{\\ell\}\\leftarrowCV\-AUROC
\(LR\(𝐇\(ℓ\),𝐲,C=0\.001\)\)\(\\text\{LR\}\(\\mathbf\{H\}^\{\(\\ell\)\},\\mathbf\{y\};C\{=\}0\.001\)\)
5:endfor
6:// Stage 2: Layer Selection
7:
ℒ∗←top\-K\(\{sℓ\}ℓ=1L\)\\mathcal\{L\}^\{\*\}\\leftarrow\\text\{top\-\}K\(\\\{s\_\{\\ell\}\\\}\_\{\\ell=1\}^\{L\}\)
8:// Stage 3: Per\-Layer Probes \+ Aggregation
9:for
ℓ∈ℒ∗\\ell\\in\\mathcal\{L\}^\{\*\}do
10:Train
fℓ←LR\(𝐇\(ℓ\),𝐲,C=0\.001\)f\_\{\\ell\}\\leftarrow\\text\{LR\}\(\\mathbf\{H\}^\{\(\\ell\)\},\\mathbf\{y\};\\ C\{=\}0\.001\)
11:endfor
12:
f\(x\)←1K∑ℓ∈ℒ∗fℓ\(𝐡\(ℓ\)\(x\)\)f\(x\)\\leftarrow\\frac\{1\}\{K\}\\sum\_\{\\ell\\in\\mathcal\{L\}^\{\*\}\}f\_\{\\ell\}\(\\mathbf\{h\}^\{\(\\ell\)\}\(x\)\)
#### Design choices\.
\(1\)*CV\-based selection*identifies layers by their actual classification performance, unlike heuristics \(last, middle\) or statistics such as Cohen’sddranking, which we show fails systematically \(Appendix[M](https://arxiv.org/html/2608.28930#A13)\)\. \(2\)*Full\-dimensional probes*avoid the information loss of dimensionality reduction; our geometric analysis confirms L2\-regularized LR outperforms subspace methods on single layers\. \(3\)*Prediction averaging*prevents the dimensionality from scaling withKK\(unlike feature concatenation\) and provides ensemble\-style variance reduction\.
#### Computational cost\.
LayerMix requires extracting hidden states from allLLlayers during the scoring phase \(a single forward pass\) and trainingKKlogistic regression probes\. At inference, hidden states fromK=5K\{=\}5layers are extracted during the standard forward pass \(no additional passes required\)\. The total wall\-clock overhead is minimal: scoring all layers takes about 30 seconds on a single GPU forN=5,000N\{=\}5\{,\}000examples, and theKKprobe trainings add only about 5 seconds\. We note that an alternative held\-out layer sweep is*not cheaper*: it requires \(a\) sacrificing held\-out data and \(b\) the same per\-layer LR fits as our CV scoring\. CV\-based selection therefore matches the held\-out sweep in compute while preserving data efficiency\.
## 5Experimental Setup
#### Models & Datasets\.
We evaluate three 7B\-scale base LLMs:Llama\-3\.1\-8B\([Grattafiori et al\., 2024](https://arxiv.org/html/2608.28930#bib.bib17)\),Mistral\-7B\([Jiang et al\., 2023](https://arxiv.org/html/2608.28930#bib.bib23)\), andQwen2\.5\-7B\([Qwen et al\., 2025](https://arxiv.org/html/2608.28930#bib.bib41)\), scaling up to 70B in §[6\.5](https://arxiv.org/html/2608.28930#S6.SS5)\. We use three standard benchmarks:TruthfulQA\(TQA;[Lin et al\., 2022b](https://arxiv.org/html/2608.28930#bib.bib33),N=4,135N\{=\}4\{,\}135\),HaluEval\-Dialogue\(HE;[Li et al\., 2023](https://arxiv.org/html/2608.28930#bib.bib29),N=20,000N\{=\}20\{,\}000\), andFaithDial\(FD;[Dziri et al\., 2022](https://arxiv.org/html/2608.28930#bib.bib14),N=5,848N\{=\}5\{,\}848\)\.
#### Baselines\.
We compare against a broad spectrum of methods: \(1\)Training\-free: Perplexity\([Ren et al\., 2022](https://arxiv.org/html/2608.28930#bib.bib42)\), Entropy, Self\-eval P\(True\)\([Kadavath et al\., 2022](https://arxiv.org/html/2608.28930#bib.bib24)\), Verbalize\([Lin et al\., 2022a](https://arxiv.org/html/2608.28930#bib.bib32)\), LLM\-Check\([Sriramanan et al\., 2024](https://arxiv.org/html/2608.28930#bib.bib43)\), DoLa\([Chuang et al\., 2023](https://arxiv.org/html/2608.28930#bib.bib12)\), INSIDE\([Chen et al\., 2024](https://arxiv.org/html/2608.28930#bib.bib11)\); \(2\)Unsup\./Semi\-sup\.: HaloScope\([Du et al\., 2024](https://arxiv.org/html/2608.28930#bib.bib13)\), CCS\([Burns et al\., 2022](https://arxiv.org/html/2608.28930#bib.bib10)\), TSV\([Park et al\., 2025](https://arxiv.org/html/2608.28930#bib.bib39)\); \(3\)Cross\-layer: ICR Probe\([Zhang et al\., 2025](https://arxiv.org/html/2608.28930#bib.bib46)\); \(4\)Subspace: SVD, MeanDiff\([Zou et al\., 2023](https://arxiv.org/html/2608.28930#bib.bib47)\), PLS\-DA, gcPCA\([Abid et al\., 2018](https://arxiv.org/html/2608.28930#bib.bib1)\); \(5\)Supervised: SAPLMA\([Azaria and Mitchell, 2023](https://arxiv.org/html/2608.28930#bib.bib2)\), Oracle\-layer LR, and ourLayerMix\.
#### Protocol\.
We employ 5\-fold stratified CV\. To isolate the representational geometry from generation\-induced distribution shifts, we evaluate within the standard paired\-example paradigm\. Hidden states are extracted at the last token position of each input sequence\. LayerMix scoring uses nested 5\-fold CV within training folds, ensuring strictly held\-out selection\. All methods use StandardScaler on training data\.Length confound check\.Length\-only AUROC \(word\-length of the answer/response token sequence\) remains far below probe AUROC across all three datasets: TQA 0\.546, HE 0\.610, FD 0\.511\. Probes consistently exceed length\-only AUROC by≥0\.34\\geq 0\.34, ruling out length shortcuts\.
## 6Results
### 6\.1Main Comparison
Table[3](https://arxiv.org/html/2608.28930#S6.T3)presents detection AUROC averaged across three models\. LayerMix achieves 0\.954 mean AUROC across 9 conditions, fully matching the single\-layer oracle \(0\.952\)\. Crucially, LayerMix captures this distributed mean\-shift signal perfectly without requiring post\-hoc oracle layer access, eliminating a major bottleneck for practical geometric analysis and deployment\.
The best zero\-overhead heuristic, middle\-layer LR, achieves 0\.941\. The 0\.013 gap closed by LayerMix is statistically significant \(p<0\.001p<0\.001\) in 8/9 conditions\. Practically, at a fixed recall of 0\.90, this AUROC improvement translates to a 15–20% relative reduction in false positives\. For practitioners without labeled data, the middle\-layer heuristic provides a strong zero\-cost alternative; however, LayerMix’s principled aggregation adds only modest cost \(a single forward pass and about 35 seconds for training\) when maximum fidelity to the geometric signal is required\.
All\-layer averaging \(0\.944\) underperforms LayerMix, confirming that structurally informed selection adds value beyond naive ensembling\. Against the two concurrent designs discussed in §[2](https://arxiv.org/html/2608.28930#S2), LayerMix exceeds CLAP in 9/9 conditions \(0\.954 vs\. 0\.928 mean AUROC\), and the two\-component PLS\-DA subspace probe reaches 0\.910, under\-performing full\-dimensional probes by 0\.044\. Furthermore, among single\-layer subspace methods, PLS\-DA \(0\.941\) clearly outperforms SVD \(0\.929\)\. This strongly corroborates our theoretical geometric decomposition \(§[3](https://arxiv.org/html/2608.28930#S3)\), demonstrating that methods explicitly targeting the mean\-shift vector strictly dominate those capturing generic variance\.
#### Standardized evaluation context as an in\-vitro control\.
As shown in Table[3](https://arxiv.org/html/2608.28930#S6.T3), several baselines underperform their originally published numbers\. All methods there are evaluated on identical pre\-extracted hidden states under our paired\-example paradigm, so the differences reflect each method’s interaction with the intrinsic representational geometry rather than its degree of paradigm\-specific tuning\. Methods originally designed for dynamic generation \(HaloScope, ICR Probe, TSV, CLAP\) are therefore evaluated outside their native deployment regime, and their published numbers under that regime are not in question\.
To ensure a strictly fair comparison under these equated geometric conditions, we re\-evaluated these baselines on instruct models with matched L2 regularization \(C=0\.001C\{=\}0\.001; Table[4](https://arxiv.org/html/2608.28930#S6.T4)\)\. Proper regularization significantly improves SEP \(0\.889→\\to0\.921\) and the oracle LR \(0\.935→\\to0\.957\)\. Nevertheless, LayerMix \(0\.960\) retains a clear advantage over SEP and edges out the matched oracle LR, confirming its benefits derive from capturing the distributed multi\-layer structure rather than favorable hyperparameter tuning\.
Table 3:Hallucination detection AUROC averaged over three 7B\-scale base models\. Bold marks LayerMix and underline marks the single\-layer oracle\.†\\daggerOfficial implementation\.∗Adapted for dataset\-specific pseudo\-labeling\.a[Suresh et al\. \(2025\)](https://arxiv.org/html/2608.28930#bib.bib44)cross\-layer attention probe, trained from the official implementation on the same cached multi\-layer hidden states under 3\-fold cross\-validation\.Table 4:Fairness ablation on instruct models \(3\-model average AUROC\)\. Performance is compared across methods under matched L2 regularization \(C=0\.001C\{=\}0\.001\)\.
### 6\.2Ablation Studies
#### Heuristic baselines\.
Every fixed\-rule layer heuristic underperforms LayerMix \(Table[17](https://arxiv.org/html/2608.28930#A13.T17)c in Appendix[M](https://arxiv.org/html/2608.28930#A13)\)\. Notably, CV\-based selection is the only strategy that consistently matches or outperforms the single\-layer oracle, whereas metric\-based rankings like Cohen’sddselect scattered layers rather than the optimal contiguous band where the mean\-shift is most pronounced\.
#### Number of layers\.
Performance is highly stable acrossK∈\{3,5,7\}K\\in\\\{3,5,7\\\}\(differences≤0\.002\\leq 0\.002; Table[17](https://arxiv.org/html/2608.28930#A13.T17)b\)\. We adoptK=5K\{=\}5as the default\.
#### Label efficiency\.
AtN≤200N\{\\leq\}200, PLS\-DAk=3k\{=\}3outperforms full\-dimensional LR \(Appendix[N](https://arxiv.org/html/2608.28930#A14)\), as the low\-rank constraint acts as an implicit regularizer\. AtN≥500N\{\\geq\}500, full\-dimensional LR overtakes PLS\-DA\.
#### Statistical significance & Robustness\.
Paired bootstrap tests confirm LayerMix vs\. last layer reachesp<0\.001p<0\.001in 9/9 conditions\. A stricter nested CV protocol \(outer 5\-fold, inner 3\-fold\) confirms negligible optimistic bias \(≤\\leq0\.010\)\. Random\-label controls yield chance AUROC \(0\.500±0\.0030\.500\\pm 0\.003\)\.
### 6\.3The Decision Boundary is Linear in the Paired\-Example Setting
A controlled MLP experiment \(256 units, ReLU, dropout 0\.3\) on the oracle layer yields an absolute difference of\|Δ\|≤0\.002\|\\Delta\|\\leq 0\.002AUROC vs\. L2\-LR across all 9 conditions \(Appendix[I](https://arxiv.org/html/2608.28930#A9)\)\. This demonstrates that, within the paired\-example paradigm, the class\-discriminative boundary is overwhelmingly linear, explaining why SAPLMA’s 2\-layer MLP \(0\.919\) underperforms a properly regularized LR \(0\.952\)\. The complexity of the signal does not warrant non\-linear functional forms in this setting\. We note that[Liang and Wang \(2025\)](https://arxiv.org/html/2608.28930#bib.bib31)report MLP probes outperforming linear probes in token\-level free\-form detection; on our reading this reflects a paradigm difference \(paired\-example sequence\-level vs\. token\-level free\-form\) rather than a contradiction\.
### 6\.4Testing Geometric Hypotheses via Controlled Ablation
Our geometric framework posits that methods discarding raw hidden states or projecting onto generic variance\-maximizing subspaces will inherently lose the mean\-shift signal\. Because evaluating existing SOTA architectures often conflates multiple algorithmic design choices, we rigorously tested our structural predictions by evaluating 12 controlled alternative architectures \(Appendix[K](https://arxiv.org/html/2608.28930#A11)\)\. These were explicitly designed to isolate and ablate specific geometric properties \(e\.g\., cross\-layer dynamics tracking, domain\-adversarial transforms, discriminant subspaces\)\. By evaluating these isolated factors within our controlled paradigm, we empirically validate the theoretical findings of Section 5: none of the 12 complex alternatives improve upon the simple L2\-LR \(0\.952 in\-domain\)\. Their performance ordering strictly correlates with how explicitly they preserve the mean\-shift component, providing evidence that complex architectures often overfit to domain\-specific covariance noise rather than exploiting hidden non\-linear structures\.
### 6\.5Instruction\-Tuned Models and Scaling
To verify generalization, we map the scaling trajectory across 25 models spanning 5 families \(0\.5B to 70B\) in Figure[3](https://arxiv.org/html/2608.28930#S6.F3)\.
Figure 3:Scaling law of hallucination detection across 25 models \(0\.5B to 70B\), evaluated on 75 conditions \(25 models×\\times3 datasets\)\. The plot compares the performance of LayerMix \(solid markers and trend line\) against the single\-layer Oracle \(hollow markers and dashed trend line\)\.First,hallucination detection scales predictably with model capacity\. Across all model families, we observe a monotonic increase in AUROC as parameter count grows\. Small models \(<2\{<\}2B\) start in the 0\.86–0\.91 range, while massive models \(e\.g\., Llama\-3\.1\-70B\) approach 0\.98 AUROC\. This implies that the linear mean\-shift structure becomes increasingly prominent and separable as the model’s representational capacity expands\.
Second,instruction tuning crystallizes the hallucination signal\. On Llama\-3\.1\-8B\-Instruct, the oracle AUROC rises to 0\.959 \(from 0\.951 base\), and LayerMix achieves 0\.961\. This confirms that the mean\-shift geometry is not merely a pretraining artifact but is preserved and even sharpened during alignment\.
Third,LayerMix scales robustly\. As depicted, our method tightly tracks or exceeds the single\-layer oracle across 72 of 75 evaluated conditions \(96\.0%\)\. This demonstrates that multi\-layer aggregation remains highly effective from 0\.5B up to the 70B scale without necessitating computationally expensive held\-out layer selection\.
## 7Discussion and Conclusion
#### Scope and predictions\.
Our geometric analysis within the paired\-example paradigm yields a testable prediction: approaches discarding raw hidden states or projecting onto low\-dimensional subspaces should underperform properly regularized LR\. This holds across all 12 hypothesis\-driven alternatives tested \(§[6\.4](https://arxiv.org/html/2608.28930#S6.SS4)\)\. While recent work corroborates linear truth\-directions\([Bao et al\., 2025](https://arxiv.org/html/2608.28930#bib.bib4);[Marks and Tegmark, 2023](https://arxiv.org/html/2608.28930#bib.bib36)\), our contribution is quantifying that the Fisher LDA gap primarily reflects covariance estimation difficulty \(about 73%\) rather than exploitable residual structure\.
#### Implications for method design\.
LayerMix exploits the signal’s distribution across a contiguous layer band, precisely where prediction averaging succeeds\([Breiman, 1996](https://arxiv.org/html/2608.28930#bib.bib8)\)\(9/9 conditions\)\. Within this paradigm, L2\-regularized LR establishes a rigorous baseline that future architectural proposals must match\.
#### When to use probing\-based detection\.
While probing requires white\-box access and labeled data, it drastically outperforms training\-free methods \(0\.95\+ vs\. 0\.50–0\.66 AUROC\)\. For practitioners, middle\-layer LR at⌊L/2⌋\\lfloor L/2\\rfloorprovides a strong zero\-overhead default \(0\.941\), while LayerMix adds 0\.013 AUROC at modest cost \(about 35 seconds\) when maximal accuracy is required\.
#### Is the mean shift an artifact?
One might worry that the dominant mean shift reflects superficial features \(e\.g\., length or style\) rather than a genuine hallucination signal\([Levinstein and Herrmann, 2025](https://arxiv.org/html/2608.28930#bib.bib28)\)\. Five observations argue against this: \(1\) length\-only AUROC \(≤0\.610\\leq 0\.610\) is far below probe performance, and regressing out length changes AUROC by≤0\.003\\leq 0\.003; \(2\) reinforcing𝜹^\\hat\{\\boldsymbol\{\\delta\}\}degrades truthfulness while projecting away improves it \(Figure[4](https://arxiv.org/html/2608.28930#S7.F4)\), providing bidirectional causal evidence; \(3\) random\-label controls yield chance AUROC; \(4\) the mean\-shift dominance replicates exactly across Llama, Mistral, and Qwen; \(5\) removing𝜹\\boldsymbol\{\\delta\}collapses detection to chance across all 9 conditions\. While confounds within the paired\-example paradigm cannot be definitively excluded without free\-form generation tests, this convergent evidence strongly supports a genuine hallucination signal\.
Figure 4:Intervention alpha sweep on Qwen2\.5\-7B \(Layer 18, TruthfulQA MC1\)\. The plot shows the effect of projecting hidden states along the mean\-shift direction𝜹^\\hat\{\\boldsymbol\{\\delta\}\}across varyingα\\alphavalues, compared against 10 random control directions\.
#### Controlled comparison\.
Addressing fairness concerns over varying regularization strengths, matchingC=0\.001C\{=\}0\.001\(Table[4](https://arxiv.org/html/2608.28930#S6.T4)\) improves SEP to 0\.921, yet LayerMix \(0\.960\) maintains a clear advantage\. Even matched oracle\-layer LR \(0\.957\) falls short, confirming LayerMix’s gains stem from multi\-layer aggregation rather than hyperparameter tuning\.
#### Domain\-specific geometry and adaptation\.
As mean\-shift directions are nearly orthogonal across datasets \(cosine about 0\.12\), hallucination signals exhibit highly dataset\-specific geometry\. Rather than reflecting spurious artifacts, this orthogonality suggests that different hallucination types \(e\.g\., knowledge deficiency, contextual exaggeration\) activate distinct geometric subspaces\. This inherent orthogonality necessitates domain adaptation across varied distributions\. However, our linear approach guarantees exceptional sample efficiency \(Appendix[E](https://arxiv.org/html/2608.28930#A5)\), rapidly recalibrating to new geometric shifts with merely a modest target\-domain adaptation set\. A natural extension for training\-free generalization is constructing a detector overSpan\(𝜹1,…,𝜹k\)\\text\{Span\}\(\\boldsymbol\{\\delta\}\_\{1\},\\ldots,\\boldsymbol\{\\delta\}\_\{k\}\)fromkkreference datasets\.
## 8Conclusion
The hallucination detection signal in LLM hidden states is overwhelmingly dominated by a linear mean shift, with apparent structural complexities largely reflecting high\-dimensional covariance estimation difficulty \(about 73% of the Fisher LDA gap\)\. By establishing that complex, variance\-targeting architectures underperform properly regularized L2\-LR, we provide a rigorous geometric framework for future probe design\. Building on this, LayerMix offers a highly efficient, oracle\-free multi\-layer aggregation strategy that captures the distributed hallucination signal to perfectly match oracle performance\. While the orthogonal nature of this signal across domains necessitates adaptation, the extreme simplicity of our linear probes ensures robust, sample\-efficient recalibration across varied distributions\.
## Limitations
1. 1\.White\-box access required\.LayerMix operates on hidden states, restricting its use to open\-weight models\.
2. 2\.Paradigm scope \(paired\-example setting\)\.The paired\-example paradigm is a methodological choice: it isolates pure representational geometry from generation\-induced distribution shifts, and is the controlled setup our geometric question requires\. As empirical evidence that this is a genuine paradigm boundary rather than a methodological convenience, we ran a small in\-house pilot on Qwen2\.5\-7B\-Instruct/TruthfulQA \(817 examples; 447 factual, 370 hallucinated\)\. The paired\-example oracle reaches 0\.943 AUROC, but direct cross\-paradigm transfer of the paired\-example mean\-shift direction𝜹^\\hat\{\\boldsymbol\{\\delta\}\}to model\-generated hidden states collapses to AUROC 0\.477 \(chance\), withcos\(𝜹^paired,𝜹^gen\)=−0\.095\\cos\(\\hat\{\\boldsymbol\{\\delta\}\}\_\{\\text\{paired\}\},\\hat\{\\boldsymbol\{\\delta\}\}\_\{\\text\{gen\}\}\)=\-0\.095\(essentially orthogonal\)\. This indicates that the geometry characterized in this paper does not directly carry over to the dynamic\-generation paradigm; effective hallucination detection during free\-form generation is a separate research direction with its own active literature\([Liang and Wang, 2025](https://arxiv.org/html/2608.28930#bib.bib31)\)\.
3. 3\.Scope of Baseline Comparisons\.Adapting dynamic generation\-based methods \(e\.g\., ICR Probe\) to our offline setting was a methodological necessity to ensure a strictly equated comparison of intrinsic representational geometry\. Consequently, our results benchmark the static geometric properties of these methods rather than their full operational capabilities in dynamic settings\. Similarly, the 12 custom architectures in our extended analysis \(§[6\.4](https://arxiv.org/html/2608.28930#S6.SS4)\) are controlled hypothesis tests designed to ablate specific geometric properties, not standalone SOTA competitors\.
4. 4\.Model Scale\.Although our primary evaluation and scaling experiments robustly cover models from 0\.5B up to 70B parameters \(Appendix[A](https://arxiv.org/html/2608.28930#A1)\), it remains unverified whether the extreme simplicity of the hallucination signal persists in frontier\-class models \(e\.g\., 400B\+ parameters\)\.
5. 5\.Absolute improvement\.The mean gain of LayerMix over the single\-layer oracle is \+0\.002 AUROC \(0\.954 vs\. 0\.952\); its primary practical and methodological contribution is eliminating the need for held\-out oracle layer selection rather than delivering massive accuracy gains\.
## Ethics Statement
LayerMix and our geometric analysis of hidden\-state probes aim to enhance AI safety by providing an efficient, transparent mechanism for hallucination detection\. However, we acknowledge several potential risks associated with this work\.
#### Potential Risks\.
First, by demonstrating that the hallucination signal is overwhelmingly dominated by a simple linear mean\-shift component, we inadvertently highlight a potential vulnerability: malicious actors could theoretically use representation engineering to craft adversarial prompts that shift hidden states orthogonal to the detection direction, thereby generating highly convincing hallucinations that bypass linear probes\. Second, no automated detection system is perfect\. There is a risk of over\-reliance, where users might develop a false sense of security and blindly trust LLM outputs that the probe fails to flag \(false negatives\), which is particularly dangerous in high\-stakes domains like healthcare or law\. Therefore, we strongly recommend integrating LayerMix strictly as a supportive diagnostic tool within a broader, human\-in\-the\-loop verification pipeline, rather than an absolute guarantee of truthfulness\.
#### Research Integrity\.
All experiments use publicly available models and datasets; no human subjects were involved\.
#### AI Assistant Disclosure\.
We employed AI assistants strictly for structural formatting and language refinement\. The conceptualization, experimental design, and data analysis were conducted entirely by the human authors\.
## Acknowledgments
This research was supported by Basic Science Research Program through the National Research Foundation of Korea\(NRF\) funded by the Ministry of Education\(NRF\-2021R1A6A1A03045425\)\. This work was supported by Institute for Information & communications Technology Promotion\(IITP\) grant funded by the Korea government\(MSIT\) \(RS\-2024\-00398115, Research on the reliability and coherence of outcomes produced by Generative AI\)\. This work was partly supported by the Institute of Information & Communications Technology Planning & Evaluation\(IITP\)\-ICT Creative Consilience Program grant funded by the Korea government\(MSIT\)\(IITP\-2026\-RS\-2020\-II201819, 25%\)\. This work was supported by the Institute of Information & Communications Technology Planning & Evaluation \(IITP\) grant funded by the Korea government \(MSIT\) \(IITP\-2026\-RS\-2026\-25615817, AI Star Fellowship Support Program\)\.
## References
- Abid et al\. \(2018\)Abubakar Abid, Martin J Zhang, Vivek K Bagaria, and James Zou\. 2018\.Exploring patterns enriched in a dataset with contrastive principal component analysis\.*Nature communications*, 9\(1\):2134\.
- Azaria and Mitchell \(2023\)Amos Azaria and Tom Mitchell\. 2023\.The internal state of an llm knows when it’s lying\.In*Findings of the Association for Computational Linguistics: EMNLP 2023*, pages 967–976\.
- Bang et al\. \(2025\)Yejin Bang, Ziwei Ji, Alan Schelten, Anthony Hartshorn, Tara Fowler, Cheng Zhang, Nicola Cancedda, and Pascale Fung\. 2025\.Hallulens: Llm hallucination benchmark\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 24128–24156\.
- Bao et al\. \(2025\)Yuntai Bao, Xuhong Zhang, Tianyu Du, Xinkui Zhao, Zhengwen Feng, Hao Peng, and Jianwei Yin\. 2025\.Probing the geometry of truth: Consistency and generalization of truth directions in llms across logical transformations and question answering tasks\.In*Findings of the Association for Computational Linguistics: ACL 2025*, pages 682–700\.
- Barker and Rayens \(2003\)Matthew Barker and William Rayens\. 2003\.Partial least squares for discrimination\.*Journal of Chemometrics: A Journal of the Chemometrics Society*, 17\(3\):166–173\.
- Belinkov \(2022\)Yonatan Belinkov\. 2022\.Probing classifiers: Promises, shortcomings, and advances\.*Computational Linguistics*, 48\(1\):207–219\.
- Bhatnagar et al\. \(2026\)Rohan Bhatnagar, Youran Sun, Chi Andrew Zhang, Yixin Wen, and Haizhao Yang\. 2026\.DRIFT: Detecting representational inconsistencies for factual truthfulness\.*arXiv preprint arXiv:2601\.14210*\.
- Breiman \(1996\)Leo Breiman\. 1996\.Bagging predictors\.*Machine learning*, 24\(2\):123–140\.
- Bürger et al\. \(2024\)Lennart Bürger, Fred A Hamprecht, and Boaz Nadler\. 2024\.Truth is universal: Robust detection of lies in llms\.*Advances in Neural Information Processing Systems*, 37:138393–138431\.
- Burns et al\. \(2022\)Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt\. 2022\.Discovering latent knowledge in language models without supervision\.*arXiv preprint arXiv:2212\.03827*\.
- Chen et al\. \(2024\)Chao Chen, Kai Liu, Ze Chen, Yi Gu, Yue Wu, Mingyuan Tao, Zhihang Fu, and Jieping Ye\. 2024\.Inside: Llms’ internal states retain the power of hallucination detection\.*arXiv preprint arXiv:2402\.03744*\.
- Chuang et al\. \(2023\)Yung\-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James Glass, and Pengcheng He\. 2023\.Dola: Decoding by contrasting layers improves factuality in large language models\.*arXiv preprint arXiv:2309\.03883*\.
- Du et al\. \(2024\)Xuefeng Du, Chaowei Xiao, and Yixuan Li\. 2024\.Haloscope: Harnessing unlabeled llm generations for hallucination detection\.*Advances in Neural Information Processing Systems*, 37:102948–102972\.
- Dziri et al\. \(2022\)Nouha Dziri, Ehsan Kamalloo, Sivan Milton, Osmar R Zaïane, Mo Yu, Edoardo M Ponti, and Siva Reddy\. 2022\.Faithdial: A faithful benchmark for information\-seeking dialogue\.*Transactions of the Association for Computational Linguistics*, 10:1473–1490\.
- Fadeeva et al\. \(2023\)Ekaterina Fadeeva, Roman Vashurin, Akim Tsvigun, Artem Vazhentsev, Sergey Petrakov, Kirill Fedyanin, Daniil Vasilev, Elizaveta Goncharova, Alexander Panchenko, Maxim Panov, and 1 others\. 2023\.Lm\-polygraph: Uncertainty estimation for language models\.In*Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations*, pages 446–461\.
- Gao et al\. \(2025\)Cheng Gao, Huimin Chen, Chaojun Xiao, Zhiyi Chen, Zhiyuan Liu, and Maosong Sun\. 2025\.H\-neurons: On the existence, impact, and origin of hallucination\-associated neurons in LLMs\.*arXiv preprint arXiv:2512\.01797*\.
- Grattafiori et al\. \(2024\)Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al\-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and 1 others\. 2024\.The llama 3 herd of models\.*arXiv preprint arXiv:2407\.21783*\.
- Han et al\. \(2025\)Jiatong Han, Neil Band, Muhammed Razzak, Jannik Kossen, Tim GJ Rudner, and Yarin Gal\. 2025\.Simple factuality probes detect hallucinations in long\-form natural language generation\.*Findings of the Association for Computational Linguistics: EMNLP*, pages 16209–16226\.
- Hernandez et al\. \(2023\)Evan Hernandez, Arnab Sen Sharma, Tal Haklay, Kevin Meng, Martin Wattenberg, Jacob Andreas, Yonatan Belinkov, and David Bau\. 2023\.Linearity of relation decoding in transformer language models\.*arXiv preprint arXiv:2308\.09124*\.
- Hewitt and Liang \(2019\)John Hewitt and Percy Liang\. 2019\.Designing and interpreting probes with control tasks\.In*Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing \(emnlp\-ijcnlp\)*, pages 2733–2743\.
- Hu et al\. \(2025\)Junjie Hu, Gang Tu, ShengYu Cheng, Jinxin Li, Jinting Wang, Rui Chen, Zhilong Zhou, and Dongbo Shan\. 2025\.Harp: Hallucination detection via reasoning subspace projection\.*arXiv preprint arXiv:2509\.11536*\.
- Huang et al\. \(2025\)Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and 1 others\. 2025\.A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions\.*ACM Transactions on Information Systems*, 43\(2\):1–55\.
- Jiang et al\. \(2023\)Albert Q\. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie\-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed\. 2023\.[Mistral 7b](https://arxiv.org/abs/2310.06825)\.*Preprint*, arXiv:2310\.06825\.
- Kadavath et al\. \(2022\)Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield\-Dodds, Nova DasSarma, Eli Tran\-Johnson, and 1 others\. 2022\.Language models \(mostly\) know what they know\.*arXiv preprint arXiv:2207\.05221*\.
- Kossen et al\. \(2024\)Jannik Kossen, Jiatong Han, Muhammed Razzak, Lisa Schut, Shreshth Malik, and Yarin Gal\. 2024\.Semantic entropy probes: Robust and cheap hallucination detection in llms\.*arXiv preprint arXiv:2406\.15927*\.
- Kuhn et al\. \(2023\)Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar\. 2023\.Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation\.*arXiv preprint arXiv:2302\.09664*\.
- Ledoit and Wolf \(2004\)Olivier Ledoit and Michael Wolf\. 2004\.A well\-conditioned estimator for large\-dimensional covariance matrices\.*Journal of multivariate analysis*, 88\(2\):365–411\.
- Levinstein and Herrmann \(2025\)Benjamin A Levinstein and Daniel A Herrmann\. 2025\.Still no lie detector for language models: probing empirical and conceptual roadblocks\.*Philosophical Studies*, 182\(7\):1539–1565\.
- Li et al\. \(2023\)Junyi Li, Xiaoxue Cheng, Wayne Xin Zhao, Jian\-Yun Nie, and Ji\-Rong Wen\. 2023\.Halueval: A large\-scale hallucination evaluation benchmark for large language models\.In*Proceedings of the 2023 conference on empirical methods in natural language processing*, pages 6449–6464\.
- Li et al\. \(2022\)Kenneth Li, Aspen K Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg\. 2022\.Emergent world representations: Exploring a sequence model trained on a synthetic task\.*arXiv preprint arXiv:2210\.13382*\.
- Liang and Wang \(2025\)Shize Liang and Hongzhi Wang\. 2025\.Neural probe\-based hallucination detection for large language models\.*arXiv preprint arXiv:2512\.20949*\.
- Lin et al\. \(2022a\)Stephanie Lin, Jacob Hilton, and Owain Evans\. 2022a\.Teaching models to express their uncertainty in words\.*arXiv preprint arXiv:2205\.14334*\.
- Lin et al\. \(2022b\)Stephanie Lin, Jacob Hilton, and Owain Evans\. 2022b\.Truthfulqa: Measuring how models mimic human falsehoods\.In*Proceedings of the 60th annual meeting of the association for computational linguistics \(volume 1: long papers\)*, pages 3214–3252\.
- Luo et al\. \(2026\)Wen Luo, Guangyue Peng, Wei Li, Shaohang Wei, Feifan Song, Liang Wang, Nan Yang, Xingxing Zhang, Jing Jin, Furu Wei, and 1 others\. 2026\.Two pathways to truthfulness: On the intrinsic encoding of llm hallucinations\.*arXiv preprint arXiv:2601\.07422*\.
- Manakul et al\. \(2023\)Potsawee Manakul, Adian Liusie, and Mark Gales\. 2023\.Selfcheckgpt: Zero\-resource black\-box hallucination detection for generative large language models\.In*Proceedings of the 2023 conference on empirical methods in natural language processing*, pages 9004–9017\.
- Marks and Tegmark \(2023\)Samuel Marks and Max Tegmark\. 2023\.The geometry of truth: Emergent linear structure in large language model representations of true/false datasets\.*arXiv preprint arXiv:2310\.06824*\.
- Meng et al\. \(2022\)Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov\. 2022\.Locating and editing factual associations in gpt\.*Advances in neural information processing systems*, 35:17359–17372\.
- Orgad et al\. \(2025\)Hadas Orgad, Michael Toker, Zorik Gekhman, Roi Reichart, Idan Szpektor, Hadas Kotek, and Yonatan Belinkov\. 2025\.LLMs know more than they show: On the intrinsic representation of LLM hallucinations\.In*International Conference on Learning Representations \(ICLR\)*\.ArXiv:2410\.02707\.
- Park et al\. \(2025\)Seongheon Park, Xuefeng Du, Min\-Hsuan Yeh, Haobo Wang, and Yixuan Li\. 2025\.Steer llm latents for hallucination detection\.*arXiv preprint arXiv:2503\.01917*\.
- Patrawala et al\. \(2025\)Arjun Patrawala, Jiahai Feng, Erik Jones, and Jacob Steinhardt\. 2025\.[LLM layers immediately correct each other](https://openreview.net/forum?id=7DY7kB8wyZ)\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*\.
- Qwen et al\. \(2025\)Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, and 25 others\. 2025\.[Qwen2\.5 technical report](https://arxiv.org/abs/2412.15115)\.*Preprint*, arXiv:2412\.15115\.
- Ren et al\. \(2022\)Jie Ren, Jiaming Luo, Yao Zhao, Kundan Krishna, Mohammad Saleh, Balaji Lakshminarayanan, and Peter J Liu\. 2022\.Out\-of\-distribution detection and selective generation for conditional language models\.*arXiv preprint arXiv:2209\.15558*\.
- Sriramanan et al\. \(2024\)Gaurang Sriramanan, Siddhant Bharti, Vinu Sankar Sadasivan, Shoumik Saha, Priyatham Kattakinda, and Soheil Feizi\. 2024\.Llm\-check: Investigating detection of hallucinations in large language models\.*Advances in Neural Information Processing Systems*, 37:34188–34216\.
- Suresh et al\. \(2025\)Malavika Suresh, Rahaf Aljundi, Ikechukwu Nkisi\-Orji, and Nirmalie Wiratunga\. 2025\.Cross\-layer attention probing for fine\-grained hallucination detection\.*arXiv preprint arXiv:2509\.09700*\.Code:[https://github\.com/itsmemala/CLAP](https://github.com/itsmemala/CLAP)\.
- Touvron et al\. \(2023\)Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, and 49 others\. 2023\.[Llama 2: Open foundation and fine\-tuned chat models](https://arxiv.org/abs/2307.09288)\.*Preprint*, arXiv:2307\.09288\.
- Zhang et al\. \(2025\)Zhenliang Zhang, Xinyu Hu, Huixuan Zhang, Junzhe Zhang, and Xiaojun Wan\. 2025\.Icr probe: Tracking hidden state dynamics for reliable hallucination detection in llms\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 17986–18002\.
- Zou et al\. \(2023\)Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann\-Kathrin Dombrowski, and 1 others\. 2023\.Representation engineering: A top\-down approach to ai transparency\.*arXiv preprint arXiv:2310\.01405*\.
## Appendix APer\-Model Detailed Results across All Scales
To evaluate the impact of model scale and family on hallucination detection, Table[5](https://arxiv.org/html/2608.28930#A1.T5)provides the detailed dataset\-level AUROC breakdown for all 25 evaluated models spanning 5 families and ranging from 0\.5B to 70B parameters\.
Additionally, Table[6](https://arxiv.org/html/2608.28930#A1.T6)presents the per\-model AUROC breakdown for the SVD\+LR and SAPLMA baselines evaluated on the oracle layer, demonstrating that cross\-model variation remains minimal\.
Table 5:Breakdown of hallucination detection AUROC across 25 evaluated models \(0\.5B–70B\)\.Table 6:Per\-model AUROC for SVD\+LR and SAPLMA baselines \(oracle layer\)\. Cross\-model variation is small \(≤0\.011\{\\leq\}0\.011\) across the models averaged in Table[3](https://arxiv.org/html/2608.28930#S6.T3)\.
## Appendix BLayer Sensitivity
Table[7](https://arxiv.org/html/2608.28930#A2.T7)details the optimal layer selected for each model and dataset condition, along with its corresponding full\-dimensional LR AUROC\.
Table 7:Optimal layer and full\-dim LR AUROC per condition\.
## Appendix CDetailed Geometric Decomposition and Fisher Gap Analysis
This section provides the full breakdown of the geometric analysis across all 9 conditions \(3 models×\\times3 datasets\), merging the mean\-shift decomposition and the Fisher LDA gap analysis into a single comprehensive view, as detailed in Table[8](https://arxiv.org/html/2608.28930#A3.T8)\.
Table 8:Detailed geometric decomposition and Fisher gap analysis across all 9 conditions\. The “Mean\-Shift Decomposition” block \(left\) shows that the 1D mean\-shift direction \(MS\-only\) captures the vast majority of the signal, while removing it \(w/o MS\) collapses detection to chance\.
## Appendix DRegularization Analysis
Figure[5](https://arxiv.org/html/2608.28930#A4.F5)illustrates the effect of varying the L2 regularization parameterCCon hallucination detection performance, confirming thatC=0\.001C=0\.001consistently acts as the optimal choice across datasets in high dimensions\.
Figure 5:L2\-regularization sweep \(Qwen2\.5\-7B\)\.C=0\.001C=0\.001consistently provides the optimal regularization strength across all three datasets\.
## Appendix EDomain Generalization and Adaptation
While the in\-domain performance of our simple linear probes is exceptionally high, real\-world deployment often requires operating across different data distributions\. Because hallucination signals exhibit dataset\-specific geometric variations, direct transfer can be challenging\. However, the inherent simplicity of our linear geometry allows the probe to rapidly recalibrate to a new domain\.
Table[9](https://arxiv.org/html/2608.28930#A5.T9)presents the cross\-domain adaptation performance across all dataset pairs\. We report the results for Mistral\-7B as a representative model to illustrate the rapid adaptation capability of our approach\. The results demonstrate that our probe effectively aligns with the target domain’s geometry without the need for complex, computationally expensive retraining\.
Table 9:Cross\-domain adaptation performance across all dataset directions \(Mistral\-7B, L2\-LR on oracle layer\)\. Using only a modest adaptation set \(N=500N=500\) from the target domain, our simple linear probe rapidly recovers performance, achieving an overall mean AUROC of \>0\.70\.
## Appendix FCross\-Domain Mean\-Shift Cosine Similarity
To further investigate cross\-domain generalization, Table[10](https://arxiv.org/html/2608.28930#A6.T10)quantifies the absolute cosine similarity between mean\-shift directions across different dataset pairs\. Furthermore, Figure[6](https://arxiv.org/html/2608.28930#A6.F6)visualizes the correlation between this mean\-shift alignment and zero\-shot transfer capabilities\.
Table 10:Absolute cosine similarity between mean\-shift directions \(\|cos\(𝜹A,𝜹B\)\|\|\\cos\(\\boldsymbol\{\\delta\}\_\{A\},\\boldsymbol\{\\delta\}\_\{B\}\)\|\) across all dataset pairs\. While random unit vectors inℝd\\mathbb\{R\}^\{d\}\(d≈4,000d\{\\approx\}4\{,\}000\) yield𝔼\[\|cos\|\]≈0\.013\\mathbb\{E\}\[\|\\cos\|\]\\approx 0\.013, the observed values remain low \(≤0\.154\{\\leq\}0\.154\), indicating that hallucination signals are nearly orthogonal across domains\.Figure 6:Mean\-shift cosine similarity vs\. direct zero\-shot transfer AUROC across 6 directional transfer pairs\. Dashed line: OLS fit \(r=0\.86r\{=\}0\.86\)\.### Within\-dataset vs\. between\-dataset cosine: artifact or hallucination type?
The low between\-dataset cosines reported above admit two interpretations: \(i\) different datasets contain different*hallucination types*, each activating a distinct geometric subspace; \(ii\) dataset\-specific*artefacts*\(prompt templates, formatting, label conventions\) drive the orthogonality\. Recent work\([Orgad et al\., 2025](https://arxiv.org/html/2608.28930#bib.bib38);[Bhatnagar et al\., 2026](https://arxiv.org/html/2608.28930#bib.bib7)\)documents a similar failure mode and does not always distinguish the two\.
To distinguish the two empirically, we compute within\-dataset cosine on TruthfulQA: split TQA into two disjoint sub\-corpora by question category \(a partition that does not depend on prompt\-template artefacts\) and measure\|cos\(𝜹^A,𝜹^B\)\|\|\\cos\(\\hat\{\\boldsymbol\{\\delta\}\}\_\{A\},\\hat\{\\boldsymbol\{\\delta\}\}\_\{B\}\)\|on cached hidden states\. The results appear in Table[11](https://arxiv.org/html/2608.28930#A6.T11)\.
Table 11:Within\-dataset \(category\-disjoint TQA halves\) vs\. between\-dataset \(mean over the three pairs in Table[10](https://arxiv.org/html/2608.28930#A6.T10)\) cosine similarity of mean\-shift directions\. The within / between ratio is about7×7\\times\.The about7×7\\timesasymmetry favours the hallucination\-type interpretation: within a single dataset \(and therefore similar hallucination types\) the geometry is largely shared, while across datasets \(different hallucination types\) it is nearly orthogonal\. This does not exclude residual artefactual contributions, but it argues the orthogonality is not predominantly artefactual\.
## Appendix GCovariance Structure
Table[12](https://arxiv.org/html/2608.28930#A7.T12)summarizes the covariance analysis metrics on TruthfulQA, highlighting the near\-identical per\-class eigenvalue spectra that further validate the mean\-shift dominance theory over differential covariance\.
Table 12:Covariance analysis \(TQA\)\. Per\-class eigenvalue spectra are near\-identical \(r\>0\.96r\>0\.96\)\. The signal is a mean shift, not differential covariance\.
## Appendix HSparse Neuron Probing
To quantify signal distribution across neurons, we perform L1\-regularized logistic regression \(liblinear,C=0\.01C\{=\}0\.01\) on each layer’s hidden states to rank neurons by absolute weight, then retrain L2\-LR on the top\-ppneurons\. Table[13](https://arxiv.org/html/2608.28930#A8.T13)reports results averaged across 9 conditions\.
Table 13:Sparse neuron probing \(9\-condition avg\)\. The signal degrades gracefully as neurons are removed\.Cross\-fold analysis of the top\-20 neurons reveals only 4 shared across 5 folds, further confirming that no fixed neuron subset carries the hallucination signal\.
#### Causal vs\. discriminative feature selection\.
[Gao et al\. \(2025\)](https://arxiv.org/html/2608.28930#bib.bib16)report that<0\.1%\{<\}0\.1\\%of neurons*causally*drive hallucinations under intervention \(“H\-neurons”\)\. Our finding that 200 neurons \(about 5% ofdd\) carry the*discriminative*signal is not in tension with this: causal selection asks “which neurons, when intervened on, change model behaviour?” while discriminative selection asks “which neurons let a probe separate factual from hallucinated?”\. A small causal core can coexist with a wider discriminative correlate; downstream neurons that merely*reflect*the H\-neuron state still carry probe\-usable signal without being causally necessary\.
## Appendix IPer\-Condition MLP vs\. L2\-LR
Table[14](https://arxiv.org/html/2608.28930#A9.T14)compares the performance of a multi\-layer perceptron \(MLP\) against L2\-LR across all 9 conditions\. The marginal differences observed firmly support the linearity of the decision boundary\.
Table 14:Per\-condition MLP \(256 units, ReLU, dropout 0\.3\) vs\. L2\-LR on the oracle layer\.\|Δ\|≤0\.002\|\\Delta\|\\leq 0\.002across all 9 conditions\.
## Appendix JAlternative Method Descriptions
Table[15](https://arxiv.org/html/2608.28930#A11.T15)evaluates 12 alternative approaches\. Below we describe each of them, together with the two published\-method comparisons reported in Table[3](https://arxiv.org/html/2608.28930#S6.T3)\(CLAP and the 2D subspace probe\)\.
#### Published\-method adaptation\.
ICR\-Offline: Adapts the cross\-layer contribution ratio dynamics of[Zhang et al\. \(2025\)](https://arxiv.org/html/2608.28930#bib.bib46)\. For each layer pair\(l,l\+1\)\(l,l\{\+\}1\), we compute norm ratios‖𝐡\(l\+1\)‖/‖𝐡\(l\)‖\\\|\\mathbf\{h\}^\{\(l\+1\)\}\\\|/\\\|\\mathbf\{h\}^\{\(l\)\}\\\|and inter\-layer cosine similarities, yielding a 108–124d feature vector\. Unlike the original ICR Probe which operates during generation, our adaptation uses pre\-extracted hidden states\.
CLAP: The cross\-layer attention probe of[Suresh et al\. \(2025\)](https://arxiv.org/html/2608.28930#bib.bib44), trained from the official implementation \(itsmemala/CLAP\)\. It applies a per\-layer projection, a learnable CLS token with sinusoidal positional encoding over the layer sequence, a 2\-block Transformer encoder, and a supervised contrastive loss with a classifier head\. We train it on the same cached multi\-layer hidden states under 3\-fold cross\-validation\.
2D subspace probe: A 2\-component PLS\-DA on standardized hidden states\. It is the direct empirical analog of the 2D subspace claim of[Bürger et al\. \(2024\)](https://arxiv.org/html/2608.28930#bib.bib9), and it does not require the polarity labels that our paired\-example data does not provide\.
#### Layer\-dynamics features\.
Trajectory shape: Fits a 3rd\-degree polynomial to the layer\-wise norm curve, extracts curvature, inflection points, and residual statistics \(12d total\)\.Multi\-layer concat: Concatenates hidden states from 3 selected layers \(early/mid/late\) and trains LR on the concatenated vector\.
#### Discriminant / subspace methods\.
CLDP\(Cross\-Layer Discriminant Pursuit\): Finds Fisher discriminant directions independently atR=15R\{=\}15layers and combines 1D projections\.CovFisher: Uses per\-class covariance difference matrices to find discriminant subspaces\.LogitFlow: Projects hidden states through the unembedding matrix and uses vocabulary\-space features for classification\.NODE\-Probe: Neural ODE\-inspired continuous\-depth model treating layers as time steps\.MSTP\(Multi\-Scale Temporal Probing\): Extracts features at multiple layer\-stride scales and concatenates\.
#### Novel architectures\.
RESIDE: Uses residual stream differences between adjacent layers as features for classification\.DAFT\(Domain\-Adversarial Feature Transform\): Applies a domain\-adversarial MLP to learn features that are discriminative for hallucination but invariant across datasets\.
#### Cross\-domain alignment\.
Procrustes: Learns orthogonal alignment between source and target hidden\-state spaces \(k=100k\{=\}100dimensions\) for cross\-domain transfer\.VocabBridge: Maps hidden states to vocabulary space via the unembedding matrix \(k=256k\{=\}256dimensions\) as a domain\-invariant representation\.
## Appendix KFull Geometric Hypothesis Ablation Results
Table[15](https://arxiv.org/html/2608.28930#A11.T15)presents the complete results of our hypothesis\-driven ablations discussed in Section[6\.4](https://arxiv.org/html/2608.28930#S6.SS4)\.
Table 15:In\-domain geometric hypothesis ablation vs\. L2\-LR \(9\-condition avg\)\. These test specific geometric predictions, not the full landscape of published methods\.†\\daggerAdapts cross\-layer dynamics of[Zhang et al\. \(2025\)](https://arxiv.org/html/2608.28930#bib.bib46)to offline features\. The twelfth, Procrustes, aligns a source space onto a target space and therefore has no in\-domain setting\.
## Appendix LSubspace Method Taxonomy
Table[16](https://arxiv.org/html/2608.28930#A12.T16)provides a taxonomy of the evaluated subspace methods, categorized by whether they target the mean\-shift direction and capture multi\-dimensional structure\.
Table 16:Subspace method taxonomy\. “Mean”: targets the mean\-shift direction\. “k\>1k\{\>\}1”: captures multi\-dimensional structure\. Methods targeting the mean shift outperform those that do not\. 3\-model avg\. on TQA\.
## Appendix MAblation Studies on LayerMix Architecture
To thoroughly evaluate the design choices of LayerMix, Table[17](https://arxiv.org/html/2608.28930#A13.T17)presents a comprehensive ablation study covering layer selection strategies, the number of aggregated layers \(KK\), and a comparison against heuristic baselines\.
Table 17:Ablation studies on LayerMix architecture \(3\-model avg\)\. Performance is stable acrossK∈\{3,5,7\}K\\in\\\{3,5,7\\\}\.
## Appendix NLabel Efficiency
Figure[7](https://arxiv.org/html/2608.28930#A14.F7)depicts the label efficiency of full\-dimensional LR compared to PLS\-DA, illustrating the crossover point where low\-rank dimensionality reduction becomes beneficial in low\-data regimes\.
Figure 7:Label efficiency \(6\-condition avg\)\. PLS\-DA outperforms full\-dimensional LR atN≤200N\\leq 200, acting as an implicit regularizer\.
## Appendix OStatistical Significance
Table[18](https://arxiv.org/html/2608.28930#A15.T18)details the results of paired bootstrap significance tests, formally verifying the robust performance advantage of LayerMix over various baseline layer selection strategies\.
Table 18:Paired bootstrap significance tests \(B=1,000B\{=\}1\{,\}000\) across 9 conditions\.†Bonferroni\-corrected \(αeff=0\.05/9=0\.0056\\alpha\_\{\\text\{eff\}\}=0\.05/9=0\.0056\)\. \*\*\* =p<0\.001p<0\.001in all significant conditions\.
## Appendix PRobustness Checks
#### Nested cross\-validation\.
We re\-ran the full pipeline with a stricter nested CV protocol \(outer 5\-fold, inner 3\-fold\)\. Estimates differ by≤0\.010\{\\leq\}0\.010from the standard 5\-fold CV across all conditions, confirming negligible optimistic bias\.
#### Selectivity control\.
Random\-label controls\([Hewitt and Liang, 2019](https://arxiv.org/html/2608.28930#bib.bib20)\)yield AUROC0\.500±0\.0030\.500\\pm 0\.003, confirming zero selectivity under the null\.
## Appendix QInstruction\-Tuned Model Results
Table[19](https://arxiv.org/html/2608.28930#A17.T19)compares the performance of the base Llama\-3\.1\-8B model against its instruction\-tuned counterpart, demonstrating that instruction tuning actually strengthens the underlying hallucination signal\.
Table 19:Base vs\. instruction\-tuned model\. Instruction tuning*strengthens*the hallucination signal\.
## Appendix RIntervention Experiment Details
As discussed in Section[7](https://arxiv.org/html/2608.28930#S7)and shown in Figure[4](https://arxiv.org/html/2608.28930#S7.F4)of the main text, we conducted a causal intervention test by modifying the hidden state at layer 18 of Qwen2\.5\-7B during generation\. We addα𝜹^\\alpha\\hat\{\\boldsymbol\{\\delta\}\}\(the unit mean\-shift direction\):𝐡′=𝐡\+α𝜹^\\mathbf\{h\}^\{\\prime\}=\\mathbf\{h\}\+\\alpha\\hat\{\\boldsymbol\{\\delta\}\}\. Positiveα\\alphaprojects away from the hallucination direction; negativeα\\alphareinforces it\.
As a specificity control, the same sweep applied along 10 random unit vectors yielded no systematic trend \(mean MC1=0\.331±0\.008=0\.331\\pm 0\.008, range\[0\.321,0\.342\]\[0\.321,0\.342\]\)\. Additionally, repeating the sweep along the full\-dimensional LR weight direction produced a weaker, non\-monotonic effect \(MC1 range0\.3250\.325–0\.3300\.330\), confirming that the causal effect is specific to the mean\-shift component itself\.
The effect size across the full sweep is substantial:Δ=0\.058\\Delta\{=\}0\.058\(17\.4%17\.4\\%relative\), withα=\+2\\alpha\{=\}\{\+\}2exceeding all 10 random controls andα=−2\\alpha\{=\}\{\-\}2falling below all of them\. While this remains a single\-model, single\-dataset pilot, the bidirectional monotonic pattern and direction specificity provide stronger causal evidence than unidirectional projection alone\. A systematic intervention study across models and datasets is needed to confirm these findings\.
## Appendix SAlternative Method Hyperparameters
Table[20](https://arxiv.org/html/2608.28930#A19.T20)reports the hyperparameters used for each of the 12 alternative methods in Table[15](https://arxiv.org/html/2608.28930#A11.T15)\. All methods use the same CV protocol and training data as L2\-LR\.
Table 20:Hyperparameters for alternative methods in Table[15](https://arxiv.org/html/2608.28930#A11.T15)\. All methods use 5\-fold CV with StandardScaler\. Neural models use Adam\. LR denotes L2\-regularized logistic regression at the specifiedCC\.Similar Articles
Hallucination Is Linearly Decodable from Mid-Layer Hidden States in Quantized LLMs
This paper investigates whether open-source quantized LLMs encode a linearly separable truthfulness signal in their hidden states. Across three 7B-8B instruction-tuned models, a linear probe on a single mid-network layer achieves 0.904-1.000 AUROC on hallucination detection benchmarks, outperforming sampling-based methods.
From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMs
This paper proposes methods for detecting hallucinations in black-box LLMs by combining semantic entropy and token-level uncertainty signals, evaluating techniques like TopK, CoCoA, Gated, and Stacked across multiple benchmarks to find that no single method is universally strongest but Stacked often performs best.
PARALLAX: Separating Genuine Hallucination Detection from Benchmark Construction Artifacts
This paper reveals that much of the reported progress in LLM hallucination detection is due to benchmark construction artifacts, where ground-truth answers are embedded in prompts, allowing a simple text-similarity baseline to achieve near-perfect scores. Through a large-scale controlled evaluation, the authors show that most methods perform near chance under proper controls, except for supervised probes on upper-layer hidden states such as SAPLMA and their proposed DRIFT.
Mind the Unseen Mass: Unmasking LLM Hallucinations via Soft-Hybrid Alphabet Estimation
Researchers introduce SHADE, a hybrid estimator that combines Good-Turing coverage with graph-spectral cues to quantify semantic uncertainty and detect LLM hallucinations when only a few black-box samples are available.
Detecting Hallucination in LLMs: Tracing the Topological Signatures of Impaired Context Sharing
This paper proposes a method to detect hallucinations in LLMs by analyzing topological signatures in attention graphs, showing improvements over existing baselines across multiple benchmarks.