The Detectability Gap: Hidden Heterogeneity in Hallucination Detection Across Language Models

arXiv cs.CL Papers

Summary

This paper identifies a 'detectability gap' in hallucination detection: hallucinations split into high-agreement (Ghost) and low-agreement (Flickering) regimes with a 0.35–0.46 AUC gap, persisting across four models and three factual QA datasets even after freezing regime assignments and using stricter trajectory-based tests. The authors argue aggregate detection metrics hide model-dependent heterogeneity and call for regime-conditioned evaluation.

arXiv:2609.35860v1 Announce Type: new Abstract: Sampling based consistency is widely used for hallucination detection, yet aggregate performance can conceal systematic differences in which errors are detectable. This work studies that heterogeneity across four language models and three factual question answering datasets. Partitioning hallucinations by answer agreement reveals high agreement (Ghost) and low agreement (Flickering) regimes with an apparent detectability gap of $0.35$ to $0.46$ AUC. Because the statistics used to define the regimes and measure this gap are strongly coupled ($|\rho|\approx0.94$ to $1.00$), the raw result is treated as a property of agreement based detection rather than independent evidence. After freezing regime assignments, lexical and semantic response dispersion preserve the asymmetry, with bootstrap $95\%$ intervals excluding zero in all $12$ model and dataset settings. A stricter test using individual diffusion trajectories and no cross seed information preserves the asymmetry across all three LLaDA datasets ($p<0.005$) and directionally across all three Dream datasets, with one reaching significance. The hard regime varies substantially in prevalence across models ($16\%$ to $77\%$), and matched prompts frequently change regimes between models. These findings show that aggregate detection metrics conceal persistent, model dependent heterogeneity in language model failures and motivate regime conditioned evaluation.
Original Article
View Cached Full Text

Cached at: 09/30/26, 09:49 AM

# The Detectability Gap: Hidden Heterogeneity in Hallucination Detection Across Language Models
Source: [https://arxiv.org/html/2609.35860](https://arxiv.org/html/2609.35860)
\\workshoptitle

GlobalSouthAI @ NeurIPS 2026: Rethinking AI for and from the Global South

Pranav DarshanAffiliation:Department of Computer Science and Engineering, R\.V\. College of Engineering, IndiaEmail:[pranavdarshan\.cs22@rvce\.edu\.in](mailto:)Pranav AAffiliation:Department of Computer Science and Engineering, R\.V\. College of Engineering, IndiaEmail:[pranava\.cs21@rvce\.edu\.in](mailto:)Sravan Karthick TAffiliation:Department of Computer Science and Engineering, R\.V\. College of Engineering, IndiaEmail:[sravankt\.cs20@rvce\.edu\.in](mailto:)Minal MoharirAffiliation:Department of Computer Science and Engineering, R\.V\. College of Engineering, IndiaEmail:[minalmoharir@rvce\.edu\.in](mailto:)Ivan P\. YamshchikovAffiliation:CAIRO, Technical University of Applied Sciences Würzburg\-Schweinfurt, GermanyEmail:[ivan\.yamshchikov@thws\.de](mailto:)

###### Abstract

Sampling based consistency is widely used for hallucination detection, yet aggregate performance can conceal systematic differences in which errors are detectable\. This work studies that heterogeneity across four language models and three factual question answering datasets\. Partitioning hallucinations by answer agreement reveals high agreement \(Ghost\) and low agreement \(Flickering\) regimes with an apparent detectability gap of0\.350\.35to0\.460\.46AUC\. Because the statistics used to define the regimes and measure this gap are strongly coupled \(\|ρ\|≈0\.94\|\\rho\|\\approx 0\.94to1\.001\.00\), the raw result is treated as a property of agreement based detection rather than independent evidence\. After freezing regime assignments, lexical and semantic response dispersion preserve the asymmetry, with bootstrap95%95\\%intervals excluding zero in all1212model and dataset settings\. A stricter test using individual diffusion trajectories and no cross seed information preserves the asymmetry across all three LLaDA datasets \(p<0\.005p<0\.005\) and directionally across all three Dream datasets, with one reaching significance\. The hard regime varies substantially in prevalence across models \(16%16\\%to77%77\\%\), and matched prompts frequently change regimes between models\. These findings show that aggregate detection metrics conceal persistent, model dependent heterogeneity in language model failures and motivate regime conditioned evaluation\.

## 1Introduction

Hallucination detectors often use*stochastic consistency*: sample a model several times and flag answers that fail to reproduce\([Manakul et al\., 2023](https://arxiv.org/html/2609.35860#bib.bib14);[Wang et al\., 2023](https://arxiv.org/html/2609.35860#bib.bib6);[Farquhar et al\., 2024](https://arxiv.org/html/2609.35860#bib.bib5);[Zhang et al\., 2023](https://arxiv.org/html/2609.35860#bib.bib15)\)\. A related line of work detects hallucinations from the denoising trace of diffusion language models directly\([Chang et al\., 2025](https://arxiv.org/html/2609.35860#bib.bib10);[Hemmat et al\., 2026](https://arxiv.org/html/2609.35860#bib.bib12);[Qian et al\., 2026](https://arxiv.org/html/2609.35860#bib.bib13)\)\. These works establish that consistency signals can miss confidently repeated errors and that diffusion trajectories carry detection relevant information; the present study is complementary; it shows that such errors form a distinct, model dependent behavioral population whose detectability is quantitatively lower for both a cross seed agreement signal and a within seed trajectory signal, and it measures how large that gap is and how it survives several independent operationalizations\. Aggregate scores can conceal systematic differences in which errors are detectable\. We examine whether stochastic generations reveal distinct hallucination populations, and whether their detectability differs when evaluation is separated from the information defining those populations\.

Our motivation is to support the development and deployment of trustworthy and reliable AI systems globally\. Uneven infrastructure, skills, and governance capacity can amplify AI risks\([UNDP, 2025](https://arxiv.org/html/2609.35860#bib.bib16)\)\. Where independent verification is difficult, confidently repeated errors may invite misplaced trust\. The goal is to inform affordable, effective safeguards that flag hallucinations before users rely on them\. Rather than assuming uniformly higher trust or lower AI literacy across these diverse communities, we study a technical obstacle: errors that evade otherwise useful detection signals\.

We evaluate four instruction tuned diffusion and autoregressive models on TriviaQA, HotpotQA, and PopQA\. Generations differ only in sampling seed\. Hallucinations are grouped into high agreement \(*Ghost*\) and low agreement \(*Flickering*\) regimes\. The primary contributions are:

1. 1\.We expose persistent detectability differences between Ghost and Flickering hallucinations, motivating regime specific evaluation rather than aggregate scores alone\.
2. 2\.We examine the circularity of raw agreement and test the asymmetry with two decoupled response level measures across all1212settings, then with individual diffusion trajectories without cross seed information\.
3. 3\.We show that regime prevalence and the regime assigned to the same prompt depend on the model\.

## 2Experimental Setup

#### Models and datasets\.

The study evaluates four instruction tuned language models: two masked diffusion models, LLaDA 8B\([Nie et al\., 2025](https://arxiv.org/html/2609.35860#bib.bib1)\)and Dream 7B\([Ye et al\., 2025](https://arxiv.org/html/2609.35860#bib.bib11)\), and two autoregressive models, Qwen2\.5 7B\([Team, 2025](https://arxiv.org/html/2609.35860#bib.bib2)\)and Llama 3\.1 8B\([Grattafiori et al\., 2024](https://arxiv.org/html/2609.35860#bib.bib3)\)\. Evaluation is conducted on three open domain factual question answering datasets: TriviaQA\([Joshi et al\., 2017](https://arxiv.org/html/2609.35860#bib.bib7)\), HotpotQA\([Yang et al\., 2018](https://arxiv.org/html/2609.35860#bib.bib8)\), and PopQA\([Mallen et al\., 2023](https://arxiv.org/html/2609.35860#bib.bib9)\)\. For each prompt,K=3K=3generations are sampled with different random seeds while all other decoding conditions are held fixed\. The autoregressive models provide agreement level controls, while only the diffusion models are used for the trajectory analysis\.

#### Correctness and hallucination labels\.

A generation is considered correct when its answer contains a gold alias\. A prompt is labeled as a hallucination \(Y=1Y=1\) when none of its sampled generations contains the gold answer\. Alias matching is performed at the token level for the diffusion corpora and using a case insensitive substring match for the autoregressive corpora\. Gold answers are used only for labeling and evaluation, never as input to a detection signal\.

#### Behavioral regimes\.

For each generation, a gold free answer span is extracted and compared across seeds using token level Jaccard similarity at thresholdτ\\tau\. LetA1,…,AKA\_\{1\},\\ldots,A\_\{K\}denote the answer spans and let𝒞τ\\mathcal\{C\}\_\{\\tau\}denote the resulting single linkage clusters\. The plurality fraction is defined as

θ^K=1K​maxc∈𝒞τ​\|c\|\.\\hat\{\\theta\}\_\{K\}=\\frac\{1\}\{K\}\\max\_\{c\\in\\mathcal\{C\}\_\{\\tau\}\}\|c\|\.\(1\)Among hallucinated prompts, two operational behavioral regimes are defined:

Ghost:θ^K\>12,Flickering:θ^K≤12\.\\textbf\{Ghost\}:\\hat\{\\theta\}\_\{K\}\>\\frac\{1\}\{2\},\\qquad\\textbf\{Flickering\}:\\hat\{\\theta\}\_\{K\}\\leq\\frac\{1\}\{2\}\.\(2\)Ghost hallucinations therefore contain a dominant repeated answer, whereas Flickering hallucinations are more dispersed across seeds\. Table[1](https://arxiv.org/html/2609.35860#S2.T1)illustrates the distinction\. These are operational behavioral partitions and do not imply distinct latent mechanisms\.

Table 1:Ghost hallucinations repeat incorrect answers; Flickering hallucinations produce different incorrect answers\. No sampled answer contains the gold answer\.
#### Evaluation signals\.

For a signalss, detectability within regimeRRis measured by the AUC separating hallucinations inRRfrom correct answers\. The detectability gap is

Δ⁡\(s\)=AUCFlickering​\(s\)−AUCGhost​\(s\)\.\\Delta\(s\)=\\mathrm\{AUC\}\_\{\\mathrm\{Flickering\}\}\(s\)\-\\mathrm\{AUC\}\_\{\\mathrm\{Ghost\}\}\(s\)\.\(3\)The raw agreement analysis uses pairwise disagreement across sampled answers\. Because this statistic is closely coupled to the plurality fraction that defines the regimes, it is treated as descriptive\. The main analysis instead freezes the regime assignments and evaluates lexical and semantic response dispersion that do not reuse the regime defining computation\. Bootstrap95%95\\%confidence intervals are computed at the prompt level usingB=2000B=2000resamples\. A separate diffusion only analysis uses denoising trajectory features from individual seeds without information across seeds\.

## 3The Detectability Gap and Its Circularity

#### Agreement based detectability\.

The raw disagreement signal1−π^1\-\\hat\{\\pi\}separates hallucinations from correct answers much more effectively in the Flickering regime than in the Ghost regime\. Across the evaluated model and dataset combinations, Ghost AUC ranges from0\.340\.34to0\.550\.55, while Flickering AUC ranges from0\.800\.80to0\.960\.96\. The resulting detectability gap is consistently large, ranging from0\.350\.35to0\.460\.46AUC, with all bootstrap95%95\\%confidence intervals excluding zero\. Thus, an aggregate agreement signal can appear effective while performing poorly on a substantial population of hallucinations\.

#### The role of circularity\.

The raw gap cannot be interpreted as independent evidence of heterogeneous detectability because the same answer agreement structure defines the regimes and constructs the disagreement signal\. Ghost hallucinations are defined by largeθ^\\hat\{\\theta\}, while the evaluated signal is its agreement based complement,1−π^1\-\\hat\{\\pi\}\. Their correlation ranges from−0\.94\-0\.94to−1\.00\-1\.00across settings\.

The raw agreement result is therefore treated as descriptive\. The stronger test freezes the Ghost and Flickering assignments and changes the evaluation signal so that it does not reuse answer span clustering, plurality, or the agreement threshold\.

## 4Detectability Heterogeneity Beyond Agreement

The central test freezes the Ghost and Flickering assignments and replaces the agreement signal with measurements that do not reuse the regime defining computation\.

000\.10\.10\.20\.20\.30\.30\.40\.4LLaDA TLLaDA HLLaDA PDream TDream HDream PQwen TQwen HQwen PLlama TLlama HLlama P00Decoupled gapΔ\\Delta\(Flickering minus Ghost AUC\)Lexical dispersionSemantic dispersion

Figure 1:Decoupled gapΔ\\Delta, defined as Flickering AUC minus Ghost AUC\. Bootstrap95%95\\%confidence intervals are shown for lexical and semantic dispersion\. All twelve model and dataset settings have positive gaps\.#### Decoupled evaluation\.

Two response level measures are considered\. Lexical dispersion is the mean pairwise1−1\-Jaccard over content words in the whole responses\. It does not use the answer span heuristic, cluster sizes, plurality fraction, or agreement threshold\. Semantic dispersion is the mean pairwise cosine distance between all MiniLM L6 v2 embeddings of the wholeresponses\([Reimers and Gurevych, 2019](https://arxiv.org/html/2609.35860#bib.bib4)\)\. Although both measures capture variation across stochastic generations, neither reuses the computation that defines the regimes\.

The lexical gap is\+0\.13\+0\.13to\+0\.29\+0\.29and the semantic gap is\+0\.08\+0\.08to\+0\.26\+0\.26\. Both signals favor Flickering in every model and dataset combination, and all bootstrap95%95\\%confidence intervals exclude zero\. Figure[1](https://arxiv.org/html/2609.35860#S4.F1)summarizes the result\. The asymmetry is smaller than the raw agreement gap, but its direction does not change when the evaluation signal is changed\.

## 5Within Seed Evidence

The response level analysis shows that the detectability gap survives measurements separated from the regime definition\. Diffusion models provide a stricter test because each generation exposes a denoising trajectory\. The trajectory signal therefore uses information from a single seed without access to the other sampled seeds\.

#### Trajectory only evaluation\.

The trajectory features capture when answer tokens become committed, confidence before commitment, and confidence change at commitment\. No information about the other seeds, answer agreement, clustering, or plurality is provided to the detector\. A leakage free out of fold logistic detector separates hallucinations from correct answers, after which predictions are evaluated separately for the frozen Ghost and Flickering groups\.

On LLaDA, the trajectory only gap is positive on all three datasets, withΔ=0\.106\\Delta=0\.106,0\.1230\.123, and0\.0920\.092on TriviaQA, HotpotQA, and PopQA, respectively\. All three permutation tests reachp<0\.005p<0\.005\. Dream shows the same direction on all three datasets, with gaps of0\.0350\.035,0\.0720\.072, and0\.1630\.163, although only PopQA reaches significance \(p=0\.003p=0\.003\)\. Complete results and confidence intervals are reported in Table[12](https://arxiv.org/html/2609.35860#A9.T12)\.

Importantly, Ghost AUC rises to0\.600\.60to0\.790\.79under trajectory information\. Ghost hallucinations are therefore not inherently undetectable\. They remain harder to detect than Flickering hallucinations even when information across seeds is removed from the signal\.

## 6Model Dependence of the Hard Regime

The detectability asymmetry is consistent across models, but the composition of the hallucination population is not\. The proportion of Ghost hallucinations ranges from6565to77%77\\%for LLaDA,5959to77%77\\%for Qwen, and6262to67%67\\%for Llama, while Dream has only1616to30%30\\%Ghost hallucinations\. Thus, the model with the smallest hard regime still exhibits the same positive decoupled detectability gaps\.

The dependence on the model is also visible when the same prompts are evaluated by different models\. Among347347prompts hallucinated by both LLaDA and Dream,181181move from Ghost under LLaDA to Flickering under Dream, whereas only1717make the reverse transition\. A McNemar test givesp<10−35p<10^\{\-35\}\. Only6767Ghost cases remain Ghost under both models\. These transitions show that, within this matched comparison, regime membership is not determined solely by the question\. The generation model itself plays an important role in determining whether a hallucination is stable or variable across seeds\.

## 7Conclusion

These results show that hallucination detectability is heterogeneous across behavioral regimes and that regime membership itself depends on the generation model\. The raw agreement gap is strongly coupled to the regime definition, but the asymmetry persists under lexical and semantic response level measurements that do not reuse that computation\. A stricter trajectory only analysis also preserves the asymmetry for LLaDA and shows the same direction across all three Dream datasets\.

For accessible safeguards in the Global South and other resource constrained settings, detection must address consistently wrong answers, not only variable ones\. Regime specific evaluation and trajectory based warnings are steps toward this goal, rather than guarantees of safe use\. Deployment affordability and performance on locally relevant languages and tasks remain to be evaluated\.

## References

- S\. Chang, J\. Yu, W\. Wang, Y\. Chen, J\. Yu, P\. Torr, and J\. GuTraceDet: hallucination detection from the decoding trace of diffusion large language models\.arXiv preprint arXiv:2510\.01274\.Cited by:[§1](https://arxiv.org/html/2609.35860#S1.p1.1)\.
- Farquharet al\.\(2024\)S\. Farquhar, J\. Kossen, L\. Kuhn, and Y\. GalDetecting hallucinations in large language models using semantic entropy\.Nature630,pp\. 625–630\.External Links:[Document](https://dx.doi.org/10.1038/s41586-024-07421-0)Cited by:[Appendix E](https://arxiv.org/html/2609.35860#A5.p1.1),[§1](https://arxiv.org/html/2609.35860#S1.p1.1)\.
- Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri,et al\.The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.External Links:[Link](https://arxiv.org/abs/2407.21783)Cited by:[§2](https://arxiv.org/html/2609.35860#S2.SS0.SSS0.Px1.p1.1)\.
- Hemmatet al\.\(2026\)A\. Hemmat, P\. Torr, Y\. Chen, and J\. YuTDGNet: hallucination detection in diffusion language models via temporal dynamic graphs\.arXiv preprint arXiv:2602\.08048\.Cited by:[§1](https://arxiv.org/html/2609.35860#S1.p1.1)\.
- Joshiet al\.\(2017\)M\. Joshi, E\. Choi, D\. S\. Weld, and L\. ZettlemoyerTriviaQA: a large scale distantly supervised challenge dataset for reading comprehension\.InProceedings of the 55th Annual Meeting of the Association for Computational Linguistics \(ACL\),pp\. 1601–1611\.External Links:[Document](https://dx.doi.org/10.18653/v1/P17-1147)Cited by:[§2](https://arxiv.org/html/2609.35860#S2.SS0.SSS0.Px1.p1.1)\.
- Mallenet al\.\(2023\)A\. Mallen, A\. Asai, V\. Zhong, R\. Das, D\. Khashabi, and H\. HajishirziWhen not to trust language models: investigating the effectiveness of parametric and non parametric memories\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(ACL\),pp\. 9802–9822\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.546)Cited by:[§2](https://arxiv.org/html/2609.35860#S2.SS0.SSS0.Px1.p1.1)\.
- Manakulet al\.\(2023\)P\. Manakul, A\. Liusie, and M\. J\. F\. GalesSelfCheckGPT: zero resource black box hallucination detection for generative large language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),External Links:[Link](https://arxiv.org/abs/2303.08896)Cited by:[§1](https://arxiv.org/html/2609.35860#S1.p1.1)\.
- Nieet al\.\(2025\)S\. Nie, F\. Zhu, Z\. You, X\. Zhang, J\. Ou, J\. Hu, J\. Zhou, Y\. Lin, J\. R\. Wen, and C\. LiLarge language diffusion models\.InProceedings of the 42nd International Conference on Machine Learning \(ICML\),External Links:[Link](https://arxiv.org/abs/2502.09992)Cited by:[§2](https://arxiv.org/html/2609.35860#S2.SS0.SSS0.Px1.p1.1)\.
- Qianet al\.\(2026\)Y\. Qian, Y\. Tan, Y\. Liu, W\. Yu, and S\. PanDynHD: hallucination detection for diffusion large language models via denoising dynamics deviation learning\.arXiv preprint arXiv:2603\.16459\.Cited by:[§1](https://arxiv.org/html/2609.35860#S1.p1.1)\.
- Reimers and Gurevych \(2019\)N\. Reimers and I\. GurevychSentence BERT: sentence embeddings using siamese BERT networks\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 3982–3992\.External Links:[Link](https://arxiv.org/abs/1908.10084)Cited by:[§4](https://arxiv.org/html/2609.35860#S4.SS0.SSS0.Px1.p1.1.1)\.
- Team \(2025\)Q\. TeamQwen2\.5 technical report\.arXiv preprint arXiv:2412\.15115\.External Links:[Link](https://arxiv.org/abs/2412.15115)Cited by:[§2](https://arxiv.org/html/2609.35860#S2.SS0.SSS0.Px1.p1.1)\.
- UNDP \(2025\)UNDPThe next great divergence: why AI may widen inequality between countries\.Technical reportUnited Nations Development Programme\.External Links:[Link](https://www.undp.org/asia-pacific/publications/next-great-divergence)Cited by:[§1](https://arxiv.org/html/2609.35860#S1.p2.1)\.
- Wanget al\.\(2023\)X\. Wang, J\. Wei, D\. Schuurmans, Q\. Le, E\. Chi, S\. Narang, A\. Chowdhery, and D\. ZhouSelf consistency improves chain of thought reasoning in language models\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://arxiv.org/abs/2203.11171)Cited by:[§1](https://arxiv.org/html/2609.35860#S1.p1.1)\.
- Yanget al\.\(2018\)Z\. Yang, P\. Qi, S\. Zhang, Y\. Bengio, W\. W\. Cohen, R\. Salakhutdinov, and C\. D\. ManningHotpotQA: a dataset for diverse, explainable multi hop question answering\.InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 2369–2380\.External Links:[Document](https://dx.doi.org/10.18653/v1/D18-1259)Cited by:[§2](https://arxiv.org/html/2609.35860#S2.SS0.SSS0.Px1.p1.1)\.
- Yeet al\.\(2025\)J\. Ye, Z\. Xie, L\. Zheng, J\. Gao, Z\. Wu, X\. Jiang, Z\. Li, and L\. KongDream 7b: diffusion large language models\.arXiv preprint arXiv:2508\.15487\.Cited by:[§2](https://arxiv.org/html/2609.35860#S2.SS0.SSS0.Px1.p1.1)\.
- Zhanget al\.\(2023\)J\. Zhang, Z\. Li, K\. Das, B\. A\. Malin, and S\. KumarSAC3: reliable hallucination detection in black box language models via semantic aware cross check consistency\.InFindings of the Association for Computational Linguistics: EMNLP 2023,External Links:[Link](https://arxiv.org/abs/2311.01740)Cited by:[§1](https://arxiv.org/html/2609.35860#S1.p1.1)\.

## Appendix AExperimental Setup and Reproducibility

The evaluation covers four instruction tuned models and three open domain factual question answering datasets\. Table[2](https://arxiv.org/html/2609.35860#A1.T2)gives the prompt counts used for each model and dataset\. The main analysis usesK=3K=3seeds per prompt, with all decoding parameters fixed while only the sampling seed varies\.

Table 2:Prompt counts per model and dataset\. Each prompt usesK=3K=3seeds in the main analysis\.Autoregressive generations use temperature0\.80\.8, nucleus probabilityp=0\.95p=0\.95, and seeds\{42,123,456\}\\\{42,123,456\\\}\. Diffusion generations use the models’ standard denoising samplers over128128steps\. The diffusion scripts do not expose a scalar sampling temperature, so no temperature is reported for LLaDA or Dream\.

The answer span is extracted with a gold free first sentence heuristic\. The token level Jaccard threshold isτ=0\.5\\tau=0\.5unless otherwise stated\. A generation is correct when its answer contains a gold alias\. A prompt is hallucinated when none of its sampled generations contains a gold alias\. Gold information is used only for labels and scoring, never as an input feature\.

Learned detectors use five fold out of fold cross validation with standardization local to each fold and ten shuffles\. Fixed scalar signals are evaluated directly on their prompt level scores\. Bootstrap confidence intervals useB=2000B=2000prompt level resamples\. Permutation tests useB=2000B=2000shuffles\.

The high depth LLaDA corpus contains5050prompts and1818seeds per dataset, for150150prompts in total\. It is used only for robustness over seed depth\.

## Appendix BConfidence and Aggregate Agreement

Mean denoising confidence provides little separation between correct and wrong answers on LLaDA\. Table[3](https://arxiv.org/html/2609.35860#A2.T3)reports the complete confidence baseline\.

Table 3:LLaDA confidence baseline\. Mean confidence is similar for correct and wrong answers, while confidence to error AUC remains near chance\.Agreement across seeds is more informative\. Hallucinated prompts disagree more than correct prompts on all three datasets, with Mann Whitneyp<10−6p<10^\{\-6\}\. As a single scalar feature, disagreement gives AUC0\.610\.61,0\.610\.61, and0\.660\.66on TriviaQA, HotpotQA, and PopQA respectively \(Table[4](https://arxiv.org/html/2609.35860#A2.T4)\)\.

Table 4:Aggregate agreement across seeds on LLaDA\. The first sentence variant uses content word disagreement in the first sentence\.
## Appendix CFull Agreement Results and Circularity

Table[5](https://arxiv.org/html/2609.35860#A3.T5)reports the complete raw agreement result for all twelve model and dataset settings\.

Table 5:Raw agreement gap under disagreement score1−π^1\-\\hat\{\\pi\}\. Bootstrap confidence intervals useB=2000B=2000prompt level resamples\.The raw Ghost AUC ranges from0\.340\.34to0\.550\.55, while Flickering AUC ranges from0\.800\.80to0\.960\.96\. The raw gap ranges from0\.350\.35to0\.460\.46and every gap confidence interval excludes zero\. The result is partly mechanical because the disagreement signal is closely related to the statistic used to define the regimes\. Table[6](https://arxiv.org/html/2609.35860#A3.T6)reports the correlation between plurality and disagreement\.

Table 6:Mechanical coupling between plurality and disagreement\. The raw agreement gap is therefore treated as descriptive rather than as independent evidence\.
## Appendix DDecoupled Response Level Controls

The Ghost and Flickering assignments are fixed using the plurality fraction\. Two response level measurements are then evaluated without reusing the answer span, clustering procedure, plurality statistic, or Jaccard threshold\.

Lexical dispersion is the mean pairwise11minus Jaccard distance over content words in the full responses\. Semantic dispersion is the mean pairwise cosine distance between normalized all MiniLM L6 v2 embeddings of the full responses\.

Table 7:Detectability gap for frozen Ghost and Flickering regimes\. G and F denote Ghost and Flickering AUC\. All twelve lexical and all twelve semantic confidence intervals exclude zero\.The lexical gap is positive in all twelve settings, ranging from0\.130\.13to0\.290\.29\. The semantic gap is also positive in all twelve settings, ranging from0\.080\.08to0\.260\.26\. The asymmetry is smaller than the raw agreement gap, but remains positive after the mechanically coupled component is removed\.

## Appendix EEntailment Based Regime Robustness

The Ghost and Flickering partition uses a gold free answer span with token level Jaccard clustering, which is a lexical proxy for whether two samples express the same answer\. A stricter and widely used criterion instead clusters samples by bidirectional entailment, so that two answers are grouped only when each entails the other\[[Farquhar et al\., 2024](https://arxiv.org/html/2609.35860#bib.bib5)\]\. This appendix recomputes the regimes under that criterion and checks whether the decoupled detectability gap survives\.

For each hallucinated prompt, the first sentence of every seed response is taken as the answer, and two seeds share a cluster when a natural language inference model assigns entailment in both directions\. The model is the publicnli\-deberta\-v3\-basecross encoder, and clustering is single linkage\. The plurality fraction and the Ghost cut at one half are unchanged\. The decoupled lexical and semantic dispersion signals are then scored against the frozen entailment regimes exactly as in Appendix[D](https://arxiv.org/html/2609.35860#A4); under the token level plurality partition these columns reproduce Table[7](https://arxiv.org/html/2609.35860#A4.T7)\.

Table 8:Regime robustness under bidirectional entailment clustering\. Ghost % is the share of hallucinations assigned to Ghost under the token level plurality criterion \(plur\.\) and the entailment criterion \(ent\.\)\. Lexical and semanticΔ\\Deltaare the decoupled Flickering minus Ghost AUC gaps under each partition\. The entailment criterion reassigns a large fraction of prompts and lowers Ghost prevalence, yet both decoupled gaps remain positive in all twelve settings\.The entailment criterion is stricter than token clustering and reassigns1515to47%47\\%of hallucinated prompts, which lowers the Ghost prevalence in every setting \(Table[8](https://arxiv.org/html/2609.35860#A5.T8)\)\. Despite this large change in the partition, the decoupled lexical gap remains positive in all twelve settings and the semantic gap remains positive in all twelve settings\. The cross model ordering of Ghost prevalence is preserved\. The detectability asymmetry therefore does not depend on the token level equivalence rule, although the prevalence magnitude does\.

## Appendix FDisjoint Generation Control

The response level controls freeze the regimes and change the measurement, but the regime assignment and the dispersion score are computed from the same sampled responses\. This appendix removes that shared dependence by assigning the regime from one set of generations and scoring detectability on a disjoint set\.

The high depth LLaDA corpus provides eighteen seeds per prompt, which supports this split\. For each of fifty random partitions the eighteen seeds are divided into an assignment half and a scoring half of nine seeds each\. The hallucination label and the Ghost or Flickering assignment use only the assignment half, while the lexical and semantic dispersion signals use only the scoring half\. The two halves never share a generation\. Table[9](https://arxiv.org/html/2609.35860#A6.T9)reports the mean gap and the fraction of splits with a positive gap\.

Table 9:Disjoint generation control on the high depth LLaDA corpus \(5050prompts,1818seeds\), averaged over5050random nine by nine seed splits\. Regime assignment uses one half of the seeds and the dispersion score uses the disjoint other half\.Under this stricter separation the decoupled gap remains clearly positive on TriviaQA and HotpotQA, where both the lexical and semantic gaps are positive in every split\. On PopQA the gap is weak: the lexical gap is small and the semantic gap is close to zero, positive in76%76\\%of splits\. The asymmetry therefore persists when assignment and scoring use disjoint generations on two of the three datasets, while PopQA, which already has the smallest response level gap, is not robust to this separation at this sample size\. Only the high depth LLaDA corpus supports this analysis, so it is not available for the other models\.

Taken together, the raw agreement gap \(Table[5](https://arxiv.org/html/2609.35860#A3.T5)\), the decoupled lexical and semantic dispersion gaps \(Table[7](https://arxiv.org/html/2609.35860#A4.T7)\), the entailment based recomputation of the regimes \(Table[8](https://arxiv.org/html/2609.35860#A5.T8)\), and the disjoint generation split \(Table[9](https://arxiv.org/html/2609.35860#A6.T9)\) are four independent operationalizations of answer agreement, spanning three different ways of deciding when two sampled answers count as the same response: token level span clustering, bidirectional entailment, and a split that never lets the same generation contribute to both the regime label and the detectability score\. The gap is positive under every one of them, so it is not an artifact of any single choice of how Ghost and Flickering are defined\.

## Appendix GRobustness Over Seed Depth

The high depth LLaDA corpus contains5050prompts and1818seeds per dataset\. The plurality estimate from three seeds averaged over many subsets tracks the eighteen seed estimate closely, with Spearman correlation approximately0\.9990\.999and mean absolute error approximately0\.0090\.009\. A single draw of three seeds is noisier, with mean absolute error approximately0\.140\.14\(Table[10](https://arxiv.org/html/2609.35860#A7.T10)\)\.

Table 10:Seed depth robustness on the high depth LLaDA corpus\.Recomputing the Ghost and Flickering labels from eighteen seeds rather than three seeds agrees on approximately79%79\\%of hallucinated prompts\. The raw agreement gap remains positive on all three datasets\. On HotpotQA the gap changes from approximately0\.2440\.244at three seeds to0\.2500\.250at eighteen seeds\.

Holding the eighteen seed partition fixed and recomputing the signal fromK∈\{3,6,9,12,18\}K\\in\\\{3,6,9,12,18\\\}subsets produces a roughly constant gap\. This analysis does not establish that three seeds are universally sufficient\.

## Appendix HThreshold Sensitivity

The token level Jaccard threshold is swept overτ∈\{0\.34,0\.5,0\.67,1\.0\}\\tau\\in\\\{0\.34,0\.5,0\.67,1\.0\\\}while the Ghost cut remains fixed atθ^\>12\\hat\{\\theta\}\>\\tfrac\{1\}\{2\}\(Table[11](https://arxiv.org/html/2609.35860#A8.T11)\)\.

Table 11:Raw agreement gap under alternative token level Jaccard thresholds on LLaDA\.The raw gap changes by less than0\.0060\.006within each dataset\. The lexical decoupled gap changes by at most approximately0\.050\.05per setting\. Examples of the lexical gap ranges are\[0\.211,0\.219\]\[0\.211,0\.219\]for LLaDA on HotpotQA,\[0\.265,0\.287\]\[0\.265,0\.287\]for Llama on TriviaQA, and\[0\.207,0\.228\]\[0\.207,0\.228\]for Qwen on PopQA\. The qualitative conclusion does not depend on the selected threshold\.

## Appendix IStrict Control Using Diffusion Trajectories

The trajectory analysis provides a stronger independence check because it uses no information across seeds\. For each answer token, the analysis records normalized commit timing, the fraction of low confidence before commit, maximum and mean confidence before commit, confidence at commit, a one step confidence jump, and a short confidence ramp\.

These quantities are summarized into a2121dimensional prompt level vector\. Agreement, plurality, answer span clustering, and all other cross seed consistency features are excluded\. Trajectory features are available for8383to98%98\\%of prompts depending on dataset\. Prompts without a resolvable trajectory are excluded before cross validation\.

A singleℓ2\\ell\_\{2\}logistic detector is trained using repeated five fold out of fold cross validation\. The resulting scores are evaluated separately on frozen Ghost and Flickering groups\.

Table 12:Trajectory only control\. The detector uses dynamics within a single seed and no information across seeds\.On LLaDA the gap is positive on all three datasets, with confidence intervals excluding zero and permutationp<0\.005p<0\.005\. On Dream the gap is positive on all three datasets but reaches significance only on PopQA\.

Ghost AUC is above chance under trajectory information, ranging from0\.600\.60to0\.790\.79\. The result therefore does not indicate that Ghost hallucinations are impossible to detect\. It indicates that they remain harder to detect than Flickering hallucinations\.

## Appendix JModel Dependent Regime Prevalence

The proportion of hallucinations assigned to Ghost varies substantially across models \(Table[13](https://arxiv.org/html/2609.35860#A10.T13)\)\.

Table 13:Ghost prevalence across the three datasets\.Ghost is the larger share of hallucinations in LLaDA, Qwen, and Llama, but a minority in Dream\. These prevalence values are computed under the token level plurality criterion and are sensitive to how answer equivalence is judged\. Recomputing the regimes with bidirectional entailment clustering of the sampled answers reassigns a fraction of hallucinated prompts and reduces the Ghost share, yet the cross model ordering is preserved and the decoupled lexical gap remains positive in all twelve settings \(Appendix[E](https://arxiv.org/html/2609.35860#A5)\)\. The prevalence magnitude should therefore be read as criterion dependent, whereas the detectability asymmetry itself is robust to this choice\.

## Appendix KMatched Model Transition Analysis

The600600Dream prompts form an exact subset of the LLaDA corpus, allowing prompt matched comparison\.

The full three state transition matrix, with LLaDA as rows and Dream as columns and states ordered as Correct, Ghost, Flickering, is

\[194916276718171782\]\.\\begin\{bmatrix\}194&9&16\\\\ 27&67&181\\\\ 7&17&82\\end\{bmatrix\}\.
Among the347347prompts hallucinated by both models, the two state matrix is

\[671811782\]\.\\begin\{bmatrix\}67&181\\\\ 17&82\\end\{bmatrix\}\.
Table 14:Matched regime transitions among the347347prompts hallucinated by both models\.The regime changes on57\.1%57\.1\\%of jointly hallucinated prompts \(Table[14](https://arxiv.org/html/2609.35860#A11.T14)\)\. The Ghost to Flickering transition occurs on181181prompts, while the reverse transition occurs on1717prompts\. McNemar’s test givesp≈8×10−36p\\approx 8\\times 10^\{\-36\}\. Only6767prompts remain Ghost under both models\. The correct or hallucination label agrees on90\.2%90\.2\\%of the600600matched prompts\.

## Appendix LBehavioral Audit

Each model’s three seeds are classified by dominant behavior among matched hallucinated prompts \(Table[15](https://arxiv.org/html/2609.35860#A12.T15)\)\.

Table 15:Dominant behavior among matched hallucinated prompts\.Within the Ghost to Flickering cell containing181181prompts, LLaDA repeats one wrong answer on96\.7%96\.7\\%of prompts\. Among these,33\.7%33\.7\\%are verbatim repetitions and63\.0%63\.0\\%are paraphrased repetitions\. Dream is diverse on84\.0%84\.0\\%of these prompts and fragmented on14\.9%14\.9\\%\.

Repeated abstention is the dominant behavior on at most2\.2%2\.2\\%of prompts for either model and occurs in only1\.1%1\.1\\%of Dream generations in this transition cell\. The transition is therefore associated with diversification and fragmentation rather than refusal\.

## Appendix MQualitative Examples

Table[16](https://arxiv.org/html/2609.35860#A13.T16)gives one Ghost and one Flickering example for each model, together with a matched model transition\. In every example, none of the three seeds contains the gold answer\.

TypeDataQuestionGoldSeed answersθ^\\hat\{\\theta\}Ghost, LLaDAPopQAWhat sport does Masahito Noto play?footballbaseball, baseball, and baseball1\.001\.00Flickering, LLaDAPopQAWho composed “The Mission”?Ennio Morriconevarious, James Horner, and The Doors0\.330\.33Ghost, DreamTriviaQAIn which city are the Oscar statuettes made?ChicagoHollywood, Los Angeles, and Los Angeles0\.670\.67Flickering, DreamTriviaQAWhich composer wrote “The Dam Busters March”?Eric Coatesan English composer, John Moore, and Edward Elgar0\.330\.33Ghost, QwenHotpotQALanguage of the people whose principal town was Anhaica?ApalacheeTimucua, Timucua, and Timucua1\.001\.00Flickering, QwenPopQAWho was the composer of “Hello”?Masaharu FukuyamaBruno Mars, “a covered pop song”, and “various songs”0\.330\.33Ghost, LlamaPopQAReligion of St George’s Cathedral?Greek Orthodox“several cathedrals, cannot determine” three times1\.001\.00Flickering, LlamaTriviaQAFictional school in ‘Please Sir’?Fenn Street Schoolno verification, Fenn St\. Secondary Modern, and Fenn St\. Elementary0\.330\.33*Matched LLaDA Ghost to Dream Flickering*LLaDA to DreamTriviaQAFinal, unfinished novel by Charles Dickens?Edwin DroodLLaDA: Bleak House three times; Dream: “Tale of Our Time”, Bleak House, and Dombeyn/aTable 16:Representative hallucinations across models\. Ghost examples show repeated incorrect answers, while Flickering examples show divergent incorrect answers\. The final row shows a matched prompt whose regime changes between models\.The matched Dickens example provides a direct illustration of model dependent regime membership\. LLaDA produces Bleak House on all three seeds, whereas Dream produces three different incorrect answers\. The question is fixed while the generation model changes the stochastic structure of the error\.

## Appendix NLeakage and Label Robustness

Several audits test whether the reported results can be explained by information leakage or implementation artifacts\.

All learned detector evaluations use five fold cross validation with preprocessing local to each training fold\. No prompt is scored by a model trained on that prompt\.

Overwriting every denoising step after a checkpoint with noise and recomputing the features gives a maximum absolute feature difference of0\.00\.0at all3131checkpoints\. This verifies construction from the causal prefix\.

No online or offline feature reads a gold field\. Permuting the labels collapses detection to chance\. Gaussian random features remain within0\.060\.06of chance AUC\. Duplicate checks pass, and corrupting the gold field leaves features unchanged with maximum difference0\.00\.0\.

Across these tests there is no evidence of the tested leakage or implementation artifacts\.

### N\.1Semantic Reassessment

The entire subset of incorrect Ghosts is reassessed to bound the effect of exact match grading\. The reassessment covers387387,404404, and423423prompts for TriviaQA, HotpotQA, and PopQA respectively\.

A consensus answer is relabeled correct only when it passes an embedding retrieval filter using BGE large with cosine similarity at least0\.750\.75and bidirectional DeBERTa entailment with both directions at least0\.50\.5\. The reassessment is automated rather than based on blinded human annotation\.

Table 17:Semantic reassessment of the Ghost population\. Every quantity depending on Ghost moves by less than0\.0050\.005\.Only44,22, and00Ghosts are verified correct \(Table[17](https://arxiv.org/html/2609.35860#A14.T17)\)\. A generous lexical upper bound gives1111,77, and11possible relabelings\. Every quantity that depends on Ghost moves by less than0\.0050\.005\. The reassessment therefore indicates that exact match grading explains only a small fraction of the evaluated high agreement errors\.

## Appendix OSupporting Online Detector

The full behavioral detector is included as supporting evidence\. It uses3636causal features evaluated at3131checkpoints\. The four feature blocks are prefix information with55features, consistency information with1212features, diffusion commit dynamics with99features, and question priors with1010features\.

The detector uses five fold out of fold evaluation andℓ2\\ell\_\{2\}logistic regression withC=0\.2C=0\.2\. The operating threshold is selected to control the cumulative false positive rate at no more than30%30\\%\.

Table 18:Full online detector on the scaled LLaDA corpus\.At cumulative FPR no greater than30%30\\%, the detector achieves TPR of66%66\\%,69%69\\%, and88%88\\%on TriviaQA, HotpotQA, and PopQA \(Table[18](https://arxiv.org/html/2609.35860#A15.T18)\)\. Median detection occurs at steps2828,2424, and88of128128\. Recall remains higher for Flickering than Ghost on every dataset\.

These results motivate warnings before generation is complete, potentially allowing a system to request verification before users rely on an answer\. This is a candidate safeguard for resource constrained deployments, including in the Global South, not a demonstrated low\-cost solution\. A small logistic classifier does not establish low end\-to\-end cost: multiple generations, feature extraction, latency, and false alarms must also be evaluated\. The reported30%30\\%false positive ceiling further requires validation against local needs and the consequences of unnecessary warnings\.

## Appendix PAdditional Diagnostic Analyses

The following analyses are retained in compressed form because they provide supporting evidence without carrying the central claim\.

### P\.1Per Step Agreement

Recomputing answer agreement at every denoising step gives little improvement over final answer agreement on TriviaQA and HotpotQA, with AUC changing from0\.610\.61to0\.610\.61on both datasets\. PopQA changes from0\.660\.66to0\.790\.79\(Table[19](https://arxiv.org/html/2609.35860#A16.T19)\)\.

Table 19:Per step agreement compared with final answer agreement on LLaDA\.The analysis establishes that temporal resolution of agreement alone does not eliminate the detectability asymmetry\. It does not establish that transient early disagreement across seeds is absent because convergence timing across seeds is not directly measured\.

### P\.2Transfer Across Corpora

Training the online detector on two datasets and evaluating it on the held out third gives mean transfer AUC of0\.7480\.748, compared with0\.7930\.793in domain\. The signal therefore transfers partly across datasets\.

### P\.3Candidate Mechanism Probe

A correlational test of subject popularity as an explanation for Ghost membership on PopQA gives AUC0\.410\.41\. This result is below chance and does not support popularity as an explanation for the hard regime\.

Stability under stochastic perturbation, early commitment, and collapse onto a popularity prior remain hypotheses rather than demonstrated mechanisms\. The present analyses do not establish a causal explanation for why some hallucinations remain highly consistent\.

## Appendix QClaim and Evidence Summary

Table[20](https://arxiv.org/html/2609.35860#A17.T20)summarizes the evidence hierarchy\.

Table 20:Summary of the central evidence\. The trajectory analysis is the only analysis in which the evaluated signal uses no information across seeds\.
## Appendix RStatistical Procedures

Agreement and correctness contrasts use two sided Mann Whitney tests\. The reported values are approximately5×10−125\\times 10^\{\-12\}for TriviaQA,3×10−73\\times 10^\{\-7\}for HotpotQA, and1×10−201\\times 10^\{\-20\}for PopQA\.

Learned detectors use five folds with predictions held out from training and averaged over shuffles\. Bootstrap confidence intervals useB=2000B=2000prompt level resamples\. Each bootstrap resample redraws prompts and recomputes both subgroup AUCs, so the gap intervals are paired\.

Permutation tests useB=2000B=2000shuffles of Ghost and Flickering assignments among hallucinated prompts while detector scores remain fixed\. The finite sample convention is

p=b\+1B\+1\.p=\\frac\{b\+1\}\{B\+1\}\.\(4\)
The smallest reportable value is approximately5×10−45\\times 10^\{\-4\}, which is reported asp<0\.001p<0\.001when appropriate\.

The analyses beyond the primary decoupled result are exploratory robustness analyses and are not corrected for multiple comparisons\.

## Appendix SAggregate AUC Decomposition

When the hallucination population is a mixture of Ghost and Flickering regimes evaluated against the same correct population, aggregate AUC is a prevalence weighted combination of the two regime specific AUCs\.

LetπG\\pi\_\{G\}denote the Ghost prevalence among hallucinations\. Then

AUCagg\\displaystyle\\mathrm\{AUC\}\_\{\\mathrm\{agg\}\}=Pr⁡\(Sh\>Sc\)\+12​Pr⁡\(Sh=Sc\)\\displaystyle=\\Pr\(S\_\{h\}\>S\_\{c\}\)\+\\frac\{1\}\{2\}\\Pr\(S\_\{h\}=S\_\{c\}\)\(5\)=∑r∈\{G,F\}Pr⁡\(Rh=r∣Yh=1\)​AUCr\\displaystyle=\\sum\_\{r\\in\\\{G,F\\\}\}\\Pr\(R\_\{h\}=r\\mid Y\_\{h\}=1\)\\,\\mathrm\{AUC\}\_\{r\}=πG​AUCG\+\(1−πG\)​AUCF\.\\displaystyle=\\pi\_\{G\}\\mathrm\{AUC\}\_\{G\}\+\(1\-\\pi\_\{G\}\)\\mathrm\{AUC\}\_\{F\}\.
HereShS\_\{h\}andScS\_\{c\}denote detector scores for hallucinated and correct examples, whileRhR\_\{h\}denotes the regime of a hallucinated example\. Aggregate performance therefore depends jointly on regime prevalence and detectability conditioned on regime\.

Because Ghost prevalence varies substantially across models, aggregate AUC can change as the composition of the hallucination population changes, even when regime conditioned detector behavior is similar\.

## Appendix TScope and Limitations

The experiments cover four models, one size range, one sampling configuration per model, andK=3K=3seeds in the main analysis\. The high depth robustness study contains150150LLaDA prompts and does not provide a matched seed depth study for autoregressive models\.

Correctness is determined by gold alias containment\. The diffusion corpora use token level matching, while the autoregressive corpora use a case insensitive substring match\. These procedures can misjudge paraphrases and formatting variants\.

The assignment of hallucinations to the Ghost and Flickering regimes depends on how answer equivalence is judged\. The reported partition uses a gold free answer span with token level clustering, which can group answers that share surface tokens without sharing meaning\. Recomputing the regimes with bidirectional entailment clustering reassigns a fraction of hallucinated prompts and lowers the Ghost prevalence, so prevalence magnitudes are criterion dependent\. The decoupled detectability gap remains positive in every setting under this alternative criterion \(Appendix[E](https://arxiv.org/html/2609.35860#A5)\), so the central asymmetry does not depend on the specific equivalence rule\.

The semantic reassessment is automated and reduces the measured effect of possible exact match errors, but it does not constitute human adjudication\.

The trajectory analysis applies only to diffusion models\. Evidence is strong for LLaDA and directional but less precise for Dream\.

The Ghost and Flickering taxonomy is an operational behavioral partition\. It identifies a reproducible distinction in the evaluated settings but does not establish that every hallucination in other architectures or domains belongs to the same two regimes\.

The Global South framing is a deployment motivation, not a regional impact finding\. We do not measure users’ trust, AI literacy, or comparative exposure to hallucination harms, nor do these benchmarks establish performance across local languages and knowledge contexts\. Future evaluation should involve communities in defining relevant tasks and acceptable error rates, and measure end\-to\-end compute, latency, and cost before claiming affordable protection\.

Similar Articles

PARALLAX: Separating Genuine Hallucination Detection from Benchmark Construction Artifacts

arXiv cs.CL

This paper reveals that much of the reported progress in LLM hallucination detection is due to benchmark construction artifacts, where ground-truth answers are embedded in prompts, allowing a simple text-similarity baseline to achieve near-perfect scores. Through a large-scale controlled evaluation, the authors show that most methods perform near chance under proper controls, except for supervised probes on upper-layer hidden states such as SAPLMA and their proposed DRIFT.