TPA: Next Token Probability Attribution for Detecting Hallucinations in RAG
Summary
TPA proposes a novel method for detecting hallucinations in RAG systems by attributing next-token probabilities to seven distinct sources (Query, RAG Context, Past Token, Self Token, FFN, Final LayerNorm, Initial Embedding) and aggregating by Part-of-Speech tags. The approach achieves state-of-the-art performance across five LLMs including Llama2, Llama3, Mistral, and Qwen.
View Cached Full Text
Cached at: 04/20/26, 08:31 AM
# Next Token Probability Attribution for Detecting Hallucinations in RAG
Source: https://arxiv.org/html/2512.07515
Guangquan Zhang
Australian Artificial Intelligence Institute (AAII)
University of Technology Sydney
Ultimo, NSW 2007, Australia
{Pengqian.Lu@student., Jie.Lu@, Anjin.Liu@, Guangquan.Zhang@}uts.edu.au
###### Abstract
Detecting hallucinations in Retrieval-Augmented Generation (RAG) remains a critical reliability challenge, as ungrounded responses can have severe consequences in high-stakes applications such as clinical decision support, legal research assistants, and autonomous agents that act on retrieved evidence. Prior approaches attribute hallucinations to a binary conflict between internal knowledge stored in FFNs and the retrieved context. However, this perspective is incomplete, failing to account for the impact of other components of the LLM, such as the user query, previously generated tokens, the self token, and the Final LayerNorm adjustment. To comprehensively capture the impact of these components on hallucination detection, we propose TPA which mathematically attributes each token's probability to seven distinct sources: Query, RAG Context, Past Token, Self Token, FFN, Final LayerNorm, and Initial Embedding. This attribution quantifies how each source contributes to the generation of the next token. Specifically, we aggregate these attribution scores by Part-of-Speech (POS) tags to quantify the contribution of each model component to the generation of specific linguistic categories within a response. By leveraging these patterns, such as detecting anomalies where Nouns rely heavily on LayerNorm, TPA effectively identifies hallucinated responses. Extensive experiments on five LLMs (Llama2-7B/13B, Llama3-8B, Mistral-7B, and Qwen3-8B) demonstrate that TPA achieves state-of-the-art performance across diverse architectures.
## TPA: Next Token Probability Attribution for Detecting Hallucinations in RAG
Pengqian Lu, Jie Lu†
†Corresponding author.
, Anjin Liu, and Guangquan Zhang
Australian Artificial Intelligence Institute (AAII)
University of Technology Sydney
Ultimo, NSW 2007, Australia
{Pengqian.Lu@student., Jie.Lu@, Anjin.Liu@, Guangquan.Zhang@}uts.edu.au
[Figure 1a: TPA attributes each next-token probability to seven sources, then aggregates source attributions by POS tags to form the features for detecting hallucination.]
[Figure 1b: Feature-importance analysis (SHAP) shows that the detector leverages source contributions conditioned on POS tags. For example, responses are more likely to be hallucinated when RAG contributes little to NOUN tokens or when LN contributes too much to NUM tokens.]
**Figure 1:** Applying the TPA framework to a Llama2-7b response from RAGTruth dataset (Niu et al., 2024).
## 1 Introduction
Large Language Models (LLMs), despite their impressive capabilities, are prone to hallucinations (Huang et al., 2025). Consequently, Retrieval-Augmented Generation (RAG) (Lewis et al., 2020) is widely used to alleviate hallucinations by grounding models in external knowledge. However, RAG systems are not perfect. They can still hallucinate by ignoring or misinterpreting the retrieved information (Sun et al., 2025). Detecting such failures is therefore a critical challenge. Following Sun et al. (2025), we define a *hallucination* as a response containing content inconsistent with the retrieved RAG context (assuming the context is relevant and correct), excluding retrieval errors, outdated knowledge, and ambiguous evidence.
The cost rises with stakes: hallucinated medication dosages in clinical decision support can harm patients, fabricated case citations in legal research assistants (Lu et al., 2026b) have led to sanctioned court filings, and ungrounded intermediate responses in autonomous agents (Lu et al., 2026a) propagate into downstream actions.
The prevailing paradigm for hallucination detection typically relies on hand-crafted proxy signals. For example, common approaches detect hallucination through consistency checks (Manakul et al., 2023) or scalar uncertainty metrics such as semantic entropy (Han et al., 2024). However, these methods only measure the symptoms of hallucination, such as output variance or surface confidence, rather than the underlying architectural causes. Consequently, they often fail when a model is confidently incorrect (Simhi et al., 2025).
To address the root cause of hallucination, recent research has shifted focus to the model's internal representations. Pioneering works such as ReDeEP (Sun et al., 2025) explicitly assume the RAG context is correct. They reveal that hallucinations in RAG typically stem from a disproportionate dominance of internal parametric knowledge (stored in FFNs) over the retrieved external context. This insight inspires a fundamental question:
*Is the binary conflict between FFNs and RAG the only cause of hallucination? Critical components like LayerNorm and User Query are often overlooked. Do contributions from these sources also drive hallucinations?*
In this paper, we extend the analysis to cover all additive components along the transformer residual stream. This approach enables detection based on the model's full internal mechanics instead of relying on partial proxy signals. To achieve this, we also assume the RAG context contains relevant information and introduce TPA (Next Token Probability Attribution for Detecting Hallucinations in RAG). This framework mathematically attributes the final probability of each token to seven distinct sources: Query, RAG, Past, Self Token, FFN, Final LayerNorm, and Initial Embedding. The attribution scores of these seven parts sum to the token's final probability, ensuring we capture the complete generation process.
To compute these attributions, we propose a probe function similar to nostalgebraist (2020) that uses the model's unembedding matrix to read out the next-token probability directly from an intermediate residual-stream state. Concretely, for each component on the residual stream, we define its contribution as the change in the probed next-token probability *before* versus *after* applying that component. In this way, we can compute the contribution from Initial Embedding, attention block, FFN, and Final LayerNorm. For the attention block, we further distribute its contribution to Query, RAG, Past and Self Token according to their attention weights.
However, these attention scores alone are insufficient for detection. A high reliance on internal parametric knowledge (FFNs) does not necessarily imply a hallucination. This pattern is expected for function words like "the" or "of". Yet, it becomes highly suspicious when found in named entities. Therefore, treating all tokens equally fails to capture these critical distinctions.
To capture this distinction, we aggregate the attribution scores using Part-of-Speech (POS) tags. We employ POS tags to capture comprehensive syntactic patterns. Unlike Named Entity Recognition (NER), which is limited to specific entity types, POS tagging covers all tokens (including critical categories like Numerals and Adpositions) and maintains high computational efficiency.
Figure 1 illustrates how TPA turns a single response into detection features: we first compute token-level source attributions, then aggregate them by POS tags. The second step is critical since hallucination signals vary across distinct parts of speech. For example, low RAG contribution on nouns or high LN contribution on numerals is often indicative of hallucination. These patterns are harder to capture if we only use raw token-level attribution scores without POS information.
Our main contributions are:
1. We propose TPA, a novel framework that mathematically attributes each token's probability to seven distinct attribution sources. This provides a comprehensive mechanistic view of the token generation process.
2. We introduce a syntax-aware aggregation mechanism. By quantifying how attribution sources drive distinct parts of speech, this approach enables the detector to pinpoint anomalies in specific entities while ignoring benign grammatical patterns.
3. Extensive experiments demonstrate that TPA achieves state-of-the-art performance. Our framework also offers transparent interpretability, automatically uncovering novel mechanistic signatures, such as anomalous LayerNorm contributions, that extend beyond the traditional FFN-RAG binary conflict.
## 2 Related Work
##### Uncertainty and Proxy Metrics.
Approaches in this category estimate hallucination via output inconsistency or proxy signals. Some methods quantify uncertainty using model ensembles (Malinin and Gales, 2021) or by measuring self-consistency across multiple sampled generations from a single model (Manakul et al., 2023). Others utilize lightweight proxy scores computable from a single generation pass, such as energy-based OOD proxy scores (Liu et al., 2020), distributional prototype learning that models each in-distribution class with class-conditioned continuous distributions for OOD detection (Peng et al., 2025), embedding-based distance scores for conditional LMs (Ren et al., 2023), and token-level uncertainty heuristics for hallucination detection (Lee et al., 2024; Zhang et al., 2023). While efficient, these scores provide indirect signals (e.g., confidence or distribution shift) and therefore may be imperfect indicators of factual correctness.
##### Distribution Shift and Predictive-Distribution Modeling.
Hallucination can also be viewed as a failure under distribution shift. Most relevant, knowledge distillation with auxiliary variables unifies logit- and feature-level predictive-distribution matching (Peng et al., 2024). Related distribution-shift lines include bias-aware prediction under long-tailed regimes (Lu et al., 2025), online adaptation under concept drift (Yu et al., 2024, 2026a, 2026b), and structural representation learning such as deep subspace clustering (Peng and Zhu, 2021) and multiplex community detection (Zhou et al., 2025). In contrast, TPA decomposes the model's *internal* next-token distribution into explicit source contributions.
##### LLM-based Evaluation.
External LLMs are also employed as verifiers. In RAG settings, outputs can be checked against retrieved evidence (Friel and Sanyal, 2023) or through claim extraction and reference-based verification (Hu et al., 2024), and LLM-as-a-judge baselines are often instantiated using curated prompts (Niu et al., 2024). Automated evaluation suites (Es et al., 2024; TrueLens, 2024) have also been developed. Other strategies include cross-examination to expose inconsistencies (Cohen et al., 2023; Yehuda et al., 2024) or fine-tuning detectors for span-level localization (Su et al., 2025). Structured multi-agent frameworks have also been explored for domain-specific reasoning pipelines (e.g., legal consultation with statutory grounding) (Lu et al., 2026b), where verifiers coordinate evidence retrieval and response refinement. However, many of these approaches require extra LLM calls or multi-step verification.
[Figure 2: Overview of the TPA framework. (1) Coarse-Grained Decomposition: Complete decomposition of token probability into four components (Section 3.2). (2) Fine-Grained Attribution: Mapping attention contributions to four input sources via head-specific weights (Section 3.3). (3) Syntax-Aware Feature Engineering: Aggregating these attributions by POS tags to construct the final detection features (Section 3.3.4).]
##### Probing Internal Activations.
Recent work extracts factuality signals from internal representations, e.g., linear truthful directions or inference-time shifts (Burns et al., 2022; Li et al., 2023), and probe-based detectors trained on hidden states (Azaria and Mitchell, 2023; Han et al., 2024). Related studies show internal states remain predictive for hallucination detection (Chen et al., 2024). Beyond detection, mechanistic analyses reveal conflicts between FFN and RAG context (Sun et al., 2025), and lightweight indicators use attention-head norms (Ho et al., 2025). Active approaches steer or edit activations (Park et al., 2025; Li et al., 2023), or adjust decoding probabilities for diagnosis (Chen et al., 2025). In contrast, we decompose the final token probability into fine-grained sources.
## 3 Methodology
As illustrated in Figure 2, TPA operates in three stages and can be implemented with a fully parallel teacher-forced pass. Given the generated response sequence **y** of length T, we can feed the entire sequence into the model with standard causal masking to extract hidden states and attention maps for all T tokens in a single teacher-forced pass. This avoids autoregressive resampling while enabling efficient attribution computation.
We first derive a complete decomposition of token probabilities (Sec. 3.2), then attribute attention contributions to specific attribution sources (Sec. 3.3). Finally, we aggregate these scores to quantify how sources drive distinct parts of speech (Sec. 3.4). The pseudo-code and complexity analysis are provided in the Appendix. We report complexity instead of wall-clock time since the latter varies in different implementation hardware.
To provide the theoretical basis for our method, we first formalize the transformer's architecture.
### 3.1 Preliminaries: Transformer Architecture
#### 3.1.1 Notations
We consi[der]...Similar Articles
CORTEX: Token-Level Hallucination Detection in RAG via Comparative Internal Representations
Proposes CORTEX, a token-level hallucination detection method for RAG that compares LLM internal representations with and without retrieved documents to identify ungrounded spans. It improves fine-grained localization of hallucinations in long-form RAG outputs.
RAGognizer: Hallucination-Aware Fine-Tuning via Detection Head Integration
RAGognizer introduces a hallucination-aware fine-tuning approach that integrates a lightweight detection head into LLMs for joint optimization of language modeling and hallucination detection in RAG systems. The paper presents RAGognize, a dataset of naturally occurring closed-domain hallucinations with token-level annotations, and demonstrates state-of-the-art hallucination detection while reducing hallucination rates without degrading language quality.
From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMs
This paper proposes methods for detecting hallucinations in black-box LLMs by combining semantic entropy and token-level uncertainty signals, evaluating techniques like TopK, CoCoA, Gated, and Stacked across multiple benchmarks to find that no single method is universally strongest but Stacked often performs best.
Beyond Document Grounding: Span-Level Hallucination Detection over Code, Tool Output, and Documents
This paper introduces a unified benchmark for span-level hallucination detection in RAG systems that extends beyond natural language to code, tool output, and structured documents, and presents a fine-tuned Qwen3.5-2B detector that outperforms existing methods on these new domains while remaining competitive on standard NLP benchmarks.
Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic Critique
This paper introduces the Latent Critic, a lightweight LoRA adapter that detects hallucinated agent actions in real time by restructuring the transformer's residual stream into localized natural-language feedback, achieving 0.966 AUROC and enabling self-correction.