Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models

arXiv cs.LG Papers

Summary

This paper presents a function-centric framework using Transcoders to trace computational pathways in vision-language models, demonstrating stronger attribution of visual grounding and the ability to predict hallucinations via graph-based features.

arXiv:2605.22902v1 Announce Type: new Abstract: Generative Vision-Language Models (VLMs) perform well on multimodal reasoning, but how visual inputs are transformed to text remains poorly understood. Existing interpretability work on VLMs uses Sparse Autoencoders (SAEs), which decompose static residual representations and miss the functional updates that drive cross-modal interaction. We adopt a function-centric framework based on Transcoders, sparse approximations of MLP sublayers that act as a causal proxy for layer-wise computation. Applied to Gemma 3-4B-IT, the framework decomposes the model into interpretable computational pathways linking image patches to directions in token generation. Transcoder attributions produce stronger and more stable effects on visually grounded tokens under patch ablation than SAE attributions, and align better with semantically relevant image regions. A False Visual Grounding counterfactual analysis confirms that the recovered pathways are specific to vision-language interaction.Finally, we perform a structural analysis of hallucinated generations, by extracting graph-based indicators from circuit traces produced by the transcoders. A logistic classifier over these mechanistic graph features predicts hallucinations at AUC $0.68$. These results show that function-centric circuit decomposition yields interpretable and predictive accounts of multimodal computation in VLMs.
Original Article
View Cached Full Text

Cached at: 05/25/26, 08:56 AM

# Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models
Source: [https://arxiv.org/html/2605.22902](https://arxiv.org/html/2605.22902)
Dimitrios Damianos Leon Voukoutis Georgios Skyrianos Vassilis Katsouros Georgios Paraskevopoulos Institute of Language and Speech Processing, Athena Research Center Athens, Greece \{d\.damianos, leon\.voukoutis, george\.skyrianos, vsk, g\.paraskevopoulos\}@athenarc\.gr

###### Abstract

Generative Vision\-Language Models \(VLMs\) perform well on multimodal reasoning, but how visual inputs are transformed to text remains poorly understood\. Existing interpretability work on VLMs uses Sparse Autoencoders \(SAEs\), which decompose static residual representations and miss the functional updates that drive cross\-modal interaction\. We adopt a function\-centric framework based on Transcoders, sparse approximations of MLP sublayers that act as a causal proxy for layer\-wise computation\. Applied to Gemma 3\-4B\-IT, the framework decomposes the model into interpretable computational pathways linking image patches to directions in token generation\. Transcoder attributions produce stronger and more stable effects on visually grounded tokens under patch ablation than SAE attributions, and align better with semantically relevant image regions\. A False Visual Grounding counterfactual analysis confirms that the recovered pathways are specific to vision\-language interaction\. Finally, we perform a structural analysis of hallucinated generations, by extracting graph\-based indicators from circuit traces produced by the transcoders\. A logistic classifier over these mechanistic graph features predicts hallucinations at AUC0\.680\.68\. These results show that function\-centric circuit decomposition yields interpretable and predictive accounts of multimodal computation in VLMs\.

## 1Introduction

Visual Language Models \(VLMs\), such as Gemma 3\(Gemma Team,[2025](https://arxiv.org/html/2605.22902#bib.bib15)\), Qwen\-VL\(Baiet al\.,[2025](https://arxiv.org/html/2605.22902#bib.bib16)\), and LLaVA\(Liuet al\.,[2023](https://arxiv.org/html/2605.22902#bib.bib17)\), have achieved state\-of\-the\-art performance in complex visual reasoning and grounded question\-answering, significantly exceeding the capabilities of contrastive frameworks such as CLIP\(Radfordet al\.,[2021](https://arxiv.org/html/2605.22902#bib.bib18)\)and SigLIP\(Zhaiet al\.,[2023](https://arxiv.org/html/2605.22902#bib.bib19)\)\. This architectural leap is driven by the integration of a Large Language Model \(LLM\) backbone, which is responsible for processing visual embeddings with linguistic context\. However, the internal mechanisms of these generative backbones remain largely underexplored\. Mechanistic interpretability research has focused mainly on the semantic properties of LLMs or the visual encoders of contrastive VLMs\.

In the LLM domain, Sparse Autoencoders \(SAEs\)\(Cunninghamet al\.,[2023](https://arxiv.org/html/2605.22902#bib.bib11); Brickenet al\.,[2023](https://arxiv.org/html/2605.22902#bib.bib12)\)have been used to decompose hidden states into human\-interpretable feature directions\. These features have provided insight into model operations\(Bereska and Gavves,[2024](https://arxiv.org/html/2605.22902#bib.bib20); Zhaoet al\.,[2024](https://arxiv.org/html/2605.22902#bib.bib21)\)and enabled steering\(Aradet al\.,[2025](https://arxiv.org/html/2605.22902#bib.bib22); Sooet al\.,[2025](https://arxiv.org/html/2605.22902#bib.bib23)\)to affect behavior\. Based on this, the adoption of Transcoders\(Dunefskyet al\.,[2024](https://arxiv.org/html/2605.22902#bib.bib24)\)and Cross\-coders\(Lindseyet al\.,[2024](https://arxiv.org/html/2605.22902#bib.bib25)\)marked a shift from state\-based decomposition to functional circuit analysis\. These architectures have become a standard tool for tracing computational pathways in Transformers, as they isolate how individual sublayers transform information, rather than focusing on how the residual stream encodes it\.

Regarding VLMs, current interpretability work remains largely confined to contrastive encoders, focusing on the emergence of monosemantic visual features\(Pachet al\.,[2025](https://arxiv.org/html/2605.22902#bib.bib3); Zaigrajewet al\.,[2025](https://arxiv.org/html/2605.22902#bib.bib27); Stevenset al\.,[2025](https://arxiv.org/html/2605.22902#bib.bib4)\)\. Although recent studies have begun to apply SAEs to Large VLMs to disentangle multimodal information\(Ortuet al\.,[2025](https://arxiv.org/html/2605.22902#bib.bib6)\)or assess cross\-modal alignment\(Louet al\.,[2025](https://arxiv.org/html/2605.22902#bib.bib10)\), these approaches typically treat the LLM backbone as a sequence of static, independent states\. However, standard SAEs are fundamentally unable to isolate the specific functional updates that affect cross\-modal interaction between visual and language information\.

In this work, we extend the foundational circuit analysis ofYanget al\.\([2026](https://arxiv.org/html/2605.22902#bib.bib29)\)by presenting a comprehensive mechanistic decomposition of VLMs\. By shifting from state\-based to functional decomposition, we offer the following contributions:

Multimodal Functional Decomposition:We apply Transcoders to generative VLM and show that they produce more stable and more semantically faithful attributions for visually grounded tokens than SAEs\. A False Visual Grounding setting confirms that the recovered pathways are specific to vision\-language interaction rather than generic MLP behavior\.

Cross\-Modal Mechanistic Tracing:Using circuit tracing, we identify computational pathways that link visual embeddings to text tokens within the LLM backbone\. These pathways reveal how cross\-modal information propagates and provide evidence that Transcoders capture a more faithful functional decomposition of the model\.

Structural Analysis of Hallucination:Structural analysis of computational paths reveals consistent differences between grounded and hallucinated outputs\. Metrics such as contribution entropy highlight these patterns and may enable mechanistic analysis of multimodal hallucinations\.

These findings indicate that function\-centric circuit decomposition could serve as a foundation for more interpretable VLM computation and provide mechanistic tools for structural analysis of common failure modes, such as multimodal hallucinations\. Code, trained Transcoders, and datasets will be released upon acceptance under the Apache 2\.0 and CC\-BY\-4\.0 licenses\.

## 2Methodology

### 2\.1Preliminaries

Both SAEs and Transcoders decompose dense activationsx∈ℝdmodelx\\in\\mathbb\{R\}^\{d\_\{\\text\{model\}\}\}into a sparse feature vectorf​\(x\)∈ℝdfeatf\(x\)\\in\\mathbb\{R\}^\{d\_\{\\text\{feat\}\}\}\(dfeat≫dmodeld\_\{\\text\{feat\}\}\\gg d\_\{\\text\{model\}\}\)\. The forward pass is defined as:

f​\(x\)=ReLU​\(We​x\+be\),y^=Wd​f​\(x\)\+bdf\(x\)=\\text\{ReLU\}\(W\_\{e\}x\+b\_\{e\}\),\\quad\\hat\{y\}=W\_\{d\}f\(x\)\+b\_\{d\}\(1\)The models are trained to minimize a reconstruction loss with anL1L\_\{1\}sparsity penalty,ℒ=‖y−y^‖22\+λ​‖f​\(x\)‖1\\mathcal\{L\}=\\\|y\-\\hat\{y\}\\\|^\{2\}\_\{2\}\+\\lambda\\\|f\(x\)\\\|\_\{1\}, where the choice of targetyydetermines the functional objective:

- •SAEs \(y=xy=x\):Reconstruct a static representation, mapping the data manifold at a specific layer\.
- •Transcoders \(y=MLP​\(x\)y=\\text\{MLP\}\(x\)\):Reconstruct a computational transformation, acting as a sparse proxy for the layer’s internal logic\.

By mapping inputs directly to MLP outputs, transcoders transition from state\-based decomposition to functional circuit tracing, isolating the causal mechanisms of multimodal integration\.

### 2\.2Experimental Setup

We integrate Transcoders and SAEs with a16×16\\timesexpansion \(40,960 features\) into every layer of Gemma 3\-4B\-IT\. For a fair comparison, both methods are applied to the MLP outputs\. To enable spatial grounding, we map flattened vision token indices back to their corresponding pixel coordinates using the processor’s grid metadata\.

Training follows a two\-stage curriculum totaling 500M tokens: \(1\) a 200M\-token text\-only warm\-up phase using Natural Instructions\(Mishraet al\.,[2022](https://arxiv.org/html/2605.22902#bib.bib30)\), and \(2\) a 300M\-token multimodal phase using a balanced mixture of COCO\(Linet al\.,[2014](https://arxiv.org/html/2605.22902#bib.bib31)\), VQAv2\(Goyalet al\.,[2017](https://arxiv.org/html/2605.22902#bib.bib32)\), and CLEVR\(Johnsonet al\.,[2017](https://arxiv.org/html/2605.22902#bib.bib33)\)Since COCO provides only image captions, we synthesize prompts for each example \(e\.g\., “Describe the image”\) to align it with the instruction\-following format\. Each layer is trained independently for∼7\\sim 7hours on a single NVIDIA A100 64GB GPU\.

Table 1:Reconstruction comparison of top\-K and L1 training approaches across all layers\.Recent work\(Yanget al\.,[2026](https://arxiv.org/html/2605.22902#bib.bib29); Gaoet al\.,[2024](https://arxiv.org/html/2605.22902#bib.bib13)\)has moved from softL​1L1regularization towardt​o​p​\-​ktop\\text\{\-\}ksparsity to enable more targeted feature selection\. In this work, we adopt at​o​p​\-​64top\\text\{\-\}64sparsity setting, as it consistently yields strong reconstruction performance compared to alternativet​o​p​\-​ktop\\text\{\-\}kvariants andL​1L1regularization, as shown in Table[1](https://arxiv.org/html/2605.22902#S2.T1)\.

Reconstruction quality is assessed using the Fraction of Variance Unexplained \(FVU\) metric\(Lieberumet al\.,[2024](https://arxiv.org/html/2605.22902#bib.bib37)\)\.

## 3Visual Grounding in Image Captioning

We evaluate whether Transcoders provide a more faithful functional decomposition of cross\-modal computation than SAEs\. Specifically, we use each method’s attribution maps to identify image patches most relevant to a target token, and then assess the magnitude and stability of the resulting changes in model output when those patches are ablated\. Our experiments are conducted on Flickr30K\(Younget al\.,[2014](https://arxiv.org/html/2605.22902#bib.bib34)\), focusing on descriptive tokens, primarily nouns and adjectives\. For Transcoders, we additionally apply circuit tracing to examine the connections between image patches and visually grounded text tokens\.

### 3\.1Attribution Mapping

Letxt\(l\)∈ℝdmodelx\_\{t\}^\{\(l\)\}\\in\\mathbb\{R\}^\{d\_\{\\text\{model\}\}\}denote the hidden state at layerll\. For a target logityy, we define a decomposed directional attribution scoreSdecS\_\{\\text\{dec\}\}\(Sikdaret al\.,[2021](https://arxiv.org/html/2605.22902#bib.bib48); Simonyanet al\.,[2013](https://arxiv.org/html/2605.22902#bib.bib49)\)as:

Sdec=∑l,ifi\(l\)​\(xt\(l\)\)⋅\(di\(l\)⋅∇h\(l\)y\),S\_\{\\text\{dec\}\}=\\sum\_\{l,i\}f\_\{i\}^\{\(l\)\}\(x\_\{t\}^\{\(l\)\}\)\\cdot\\big\(d\_\{i\}^\{\(l\)\}\\cdot\\nabla\_\{h^\{\(l\)\}\}y\\big\),\(2\)wherefi\(l\)​\(xt\(l\)\)f\_\{i\}^\{\(l\)\}\(x\_\{t\}^\{\(l\)\}\)denotes the activation of featureiiat layerll,di\(l\)d\_\{i\}^\{\(l\)\}is the corresponding decoder direction \(e\.g\., from an SAE or Transcoder\), and∇h\(l\)y=∂y∂h\(l\)\\nabla\_\{h^\{\(l\)\}\}y=\\frac\{\\partial y\}\{\\partial h^\{\(l\)\}\}is the gradient of the target logit with respect to the MLP outputh\(l\)h^\{\(l\)\}at layerll\. The termdi\(l\)⋅∇h\(l\)yd\_\{i\}^\{\(l\)\}\\cdot\\nabla\_\{h^\{\(l\)\}\}ycaptures directional influence along the feature direction\. Since SAEs and Transcoders are trained on different points in the MLP computation, the variablext\(l\)x\_\{t\}^\{\(l\)\}denotes the MLP output when using SAEs, and the MLP input when using Transcoders\.

SAEs and Transcoders are trained to approximate MLP computations under sparsity constraints, creating a feature basis that provides a low\-dimensional decomposition of intermediate representations\. Within this framework, decoder directions define a structured basis over which local model behavior can be expressed\. We therefore interpretSdecS\_\{\\text\{dec\}\}as a decomposition of local behavior in this feature space, where each term combines feature activation with directional influence to quantify contribution to the target logit\.

To obtain patch\-level attributions, we aggregate feature\-level contributions across all features and layers associated with each image patchpp, yielding a scalar importance score per patch\.

### 3\.2Ablation\-Based Evaluation

To evaluate the quality of the resulting attributions, we measure changes in token probability and entropy under targeted patch ablations\. Specifically, we remove the top\-MMimage patches ranked by attribution score and compute the resulting changes in target token probability,Δ​p=po​r​i​g​i​n​a​l−pa​b​l​a​t​e​d\\Delta p=p\_\{original\}\-p\_\{ablated\}, and entropyΔ​H=Ho​r​i​g​i​n​a​l−Ha​b​l​a​t​e​d\\Delta H=H\_\{original\}\-H\_\{ablated\}\. An example of this procedure is shown in Fig\.[1](https://arxiv.org/html/2605.22902#S3.F1), where we report results forM=10M=10\.

![Refer to caption](https://arxiv.org/html/2605.22902v1/images/ranked_grid_top10_001.png)Figure 1:Comparison of thet​o​p−10top\-10most important image patches identified by SAEs and Transcoders\. Transcoders identify patches that are more aligned with visually grounded tokens, as reflected in both visual correspondence and their impact on token probability and entropy\.As shown in the examples, Transcoders tend to identify image patches that are more relevant to the target token compared to SAEs\. This is reflected in the effect of targeted patch ablations on token probability and entropy\. In particular, removing Transcoder\-identified patches leads to a larger decrease in token probability and a larger increase in entropy compared to SAE\-identified patches\. Overall, these results suggest that Transcoders more consistently highlight image regions associated with the generation of the target token\. We provide additional examples in Appendix[B\.1](https://arxiv.org/html/2605.22902#A2.SS1), including results for top\-1 and top\-5 patch ablations\.

### 3\.3Circuit tracing

We examine the computational pathways linking image patches to target tokens through circuit tracing\. As shown in Fig\.[2](https://arxiv.org/html/2605.22902#S3.F2), visually grounded tokens \(e\.g\., nouns and adjectives\) depend on a small subset of image regions\. These pathways are localized to specific regions and form structured computational chains that trace how information flows through the model to influence token generation\.

![Refer to caption](https://arxiv.org/html/2605.22902v1/captions/token_288_copy_2.png)Figure 2:Circuit analysis on captions: visually grounded tokens have clear semantic links to specific visual regions\.This analysis complements our results in Section[3\.1](https://arxiv.org/html/2605.22902#S3.SS1): while attribution maps identify relevant regions, circuit tracing reveals the underlying structure linking these regions to language outputs\. In particular, the traced paths indicate that the model conditions its predictions on distinct visual inputs with consistent semantic correspondence, providing a mechanistic view of how visual grounding is implemented\. Additional examples and visualizations are provided in Appendix[B\.2](https://arxiv.org/html/2605.22902#A2.SS2)\.

## 4Counterfactual Analysis: False Visual Grounding

To assess whether Transcoders selectively capture vision–language links rather than attributing importance to image inputs indiscriminately, we evaluate them in a False Visual Grounding \(FVG\) setting\. This setting serves as a counterfactual test in which the correct answer does not depend on the image, allowing us to examine whether Transcoders assign spurious visual relevance\.

We construct arithmetic and symbolic question–answer pairs \(e\.g\., “What is 20 \+ 30?” or “What is the capital of France?”\) and pair them with images from our captioning dataset\. Since these tasks can be solved without visual input, the image acts as a distractor, providing a controlled setting to evaluate whether attribution methods incorrectly rely on visual information\.

We follow the same procedure as in Section[3\.1](https://arxiv.org/html/2605.22902#S3.SS1), using Transcoder\-based attribution maps to identify and zero out thet​o​p−Mtop\-Mimage patches, and measuring the resulting changes in token probability \(Δ​p\\Delta p\) and entropy \(Δ​H\\Delta H\)\. Example attribution maps are shown in Fig\.[3](https://arxiv.org/html/2605.22902#S4.F3)\.

![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg_attr_maps_top10/entropy_ranked_top10_001.png)Figure 3:Attribution maps in the False Visual Grounding setting\.We observe that the patches identified by Transcoders in the FVG setting show no clear visual correspondence with the target token\. Consistently, ablating these patches results in negligible changes in both token probability and entropy \(Δ​p≈0\\Delta p\\approx 0,Δ​H≈0\\Delta H\\approx 0\)\. This suggests that the identified patches do not carry measurable information relevant to the prediction in this setting\.

Finally, circuit tracing \(Fig\.[4](https://arxiv.org/html/2605.22902#S4.F4)\) reveals minimal connectivity between image tokens and target logits, indicating that visual information plays a negligible role in the underlying computation\. This contrasts with the more structured and localized pathways observed in visually grounded settings\.

![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg/token_295.png)Figure 4:Circuit analysis on FVG setting: Transcoders reveal no correlation between the target and image tokens\.Overall, these results suggest that Transcoder\-based attributions selectively respond to meaningful visual signals in visually grounded settings, while not assigning systematic importance to image inputs in the absence of such signals\. We present additional visual examples in Appendix[B\.3](https://arxiv.org/html/2605.22902#A2.SS3)and[B\.4](https://arxiv.org/html/2605.22902#A2.SS4)\.

![Refer to caption](https://arxiv.org/html/2605.22902v1/images/graph_two.png)Figure 5:Example of computation path: We examine both per layer and per token paths\. We visualize the metrics we use in our structural analysis\.
## 5Structural Analysis of Hallucinations

Having established that Transcoders provide an interpretable functional decomposition of model activations, we now analyze computation graphs derived from circuit analysis and investigate whether their structural properties can help distinguish hallucinated from correct outputs\.

To this end, we use 150 samples from theHaloQuestbenchmark\(Wanget al\.,[2024](https://arxiv.org/html/2605.22902#bib.bib35)\), where responses are generated by Gemma 3\-4B\-IT and labeled as eitherCorrect\(n=50n=50\) orHallucinated\(n=100n=100\) based on GPT\-5\(Singhet al\.,[2025](https://arxiv.org/html/2605.22902#bib.bib46)\)\-generated annotations\.

### 5\.1Computational Graph Structure

We analyze computational graphs using five structural descriptors that characterize depth, span, and locality of computation in token\-level pathways\. An example computation graph is shown in Fig\.[5](https://arxiv.org/html/2605.22902#S4.F5)\.

Mean Path Depth:The average number of nodes along pathways, capturing the length of computation chains\.

Layer Span:The vertical extent of a pathway,Lroot−LminL\_\{\\text\{root\}\}\-L\_\{\\text\{min\}\}, measuring how many layers are traversed\.

Token Distance:The positional offset\|oleaf−oroot\|\|o\_\{\\text\{leaf\}\}\-o\_\{\\text\{root\}\}\|, capturing the extent of information propagation across the token sequence\.

Contribution Entropy:H=−∑ipi​log⁡piH=\-\\sum\_\{i\}p\_\{i\}\\log p\_\{i\}, wherepip\_\{i\}denotes the normalized activation of each feature across all features and layers\.

Top\-1 Feature Fraction:The maximum normalized feature activation, measuring how concentrated the representation is in the most dominant feature\.

Table 2:Comparative analysis of structural metrics: correct vs\. hallucinated computation paths\. Metrics marked with ‘\*’ are statistically significant atp<0\.01p<0\.01\.In Table[2](https://arxiv.org/html/2605.22902#S5.T2)we observe average differences between correct and hallucinated cases\. Hallucinations show slightly lower token distance and entropy, alongside a slightly higher top\-1 feature fraction, suggesting subtle shifts toward more localized token context and more concentrated feature usage\.

### 5\.2Predicting Hallucinations through Structure

To evaluate whether these structural differences carry predictive signal, we frame hallucination detection as a binary classification task\. We use a dataset of 150 examples \(100 hallucinations and 50 correct\) with five\-fold cross\-validation, and train a logistic regression model using the five structural features introduced above\.

Given the class imbalance, we report Balanced Accuracy, F1 Score and Area Under the ROC Curve \(AUC\)\. We compare against a Majority Class baseline \(always predicting hallucination\) and a Random Chance classifier\. Results are reported in Table[3](https://arxiv.org/html/2605.22902#S5.T3)\.

Table 3:Hallucination classification performance\. While the majority baseline is limited by the data imbalance, the logistic model shows a measurable gain in both AUC and Balanced Accuracy\.The model outperforms both baselines across all metrics\. In particular, an AUC of 0\.68 indicates that the structural features contain a measurable, though modest, predictive signal for hallucination detection\. While these results are preliminary, they suggest that internal circuit structure carries information relevant to distinguishing hallucinated from correct outputs\.

Table 4:Mean absolute SHAP values for the logistic regression hallucination classifier\. Values indicate the global importance of each structural feature\.To understand which structural characteristics drive the classifier’s predictions, we apply the SHAP framework\(Lundberg and Lee,[2017](https://arxiv.org/html/2605.22902#bib.bib50)\)to compute feature attributions for the logistic regression model\. Specifically, we report the mean absolute SHAP value for each feature across the evaluation set, where larger values indicate greater overall influence on the model’s predictions, independent of direction\. Since hallucinations are defined as the positive class, positive SHAP contributions push the prediction toward hallucination, while negative contributions push it toward correct outputs\.

As shown in Table[4](https://arxiv.org/html/2605.22902#S5.T4), Mean Token Distance, Mean Layer Span, and feature concentration measures exhibit the highest mean absolute SHAP values, indicating stronger contribution to the classifier’s decisions\. In particular, hallucinated outputs are associated with lower token distances and lower entropy, alongside higher top\-1 feature fraction\. This suggests that hallucinations are characterized by more localized interactions in token space and a more concentrated distribution of feature\-level contributions, relative to correct outputs\.

## 6Related Work

Mechanistic Interpretability on LLMs\.Sparse Autoencoders \(SAEs\) decompose dense hidden states into human\-interpretable directions\(Templetonet al\.,[2024](https://arxiv.org/html/2605.22902#bib.bib36); Lieberumet al\.,[2024](https://arxiv.org/html/2605.22902#bib.bib37)\)\. Beyond feature extraction, recent work has moved towards concept extraction\(Helffet al\.,[2026](https://arxiv.org/html/2605.22902#bib.bib2)\), and explanation verifications through falsification frameworks\(Billset al\.,[2025](https://arxiv.org/html/2605.22902#bib.bib38)\), applying these findings to model control\. This includes the use of “persona vectors” to monitor character traits\(Chenet al\.,[2025](https://arxiv.org/html/2605.22902#bib.bib39)\)and the deployment of SAE\-based steering to correct erroneous reasoning paths in real\-time\(Fanget al\.,[2026](https://arxiv.org/html/2605.22902#bib.bib42); Choet al\.,[2026](https://arxiv.org/html/2605.22902#bib.bib43)\)\. On the same time, the introduction of Transcoders\(Dunefskyet al\.,[2024](https://arxiv.org/html/2605.22902#bib.bib24)\)and Cross\-coders\(Team,[2025](https://arxiv.org/html/2605.22902#bib.bib40)\)has allowed a functional analysis of LLMs\. These architectures isolate the specific updates performed by sublayers rather than merely characterizing the residual stream\. This paradigm allows for the identification of "modular circuits"—reusable computational motifs across diverse tasks\(He and others,[2025](https://arxiv.org/html/2605.22902#bib.bib41)\)—providing a robust, verified toolkit for aligning LLM behavior with human intent\(Naseemet al\.,[2026](https://arxiv.org/html/2605.22902#bib.bib44)\)\.

Mechanistic Interpretability on VLMs\.Interpretability for VLMs has evolved from analyzing contrastive encoders like CLIP\(Radfordet al\.,[2021](https://arxiv.org/html/2605.22902#bib.bib18)\)to decomposing generative multimodal backbones\. Previous work established that SAEs can extract monosemantic visual features from ViTs\(Pachet al\.,[2025](https://arxiv.org/html/2605.22902#bib.bib3); Stevenset al\.,[2025](https://arxiv.org/html/2605.22902#bib.bib4)\)\. Recent efforts have extended this to Large VLMs, identifying specialized components that disentangle factual priors from visual grounding\(Ortuet al\.,[2025](https://arxiv.org/html/2605.22902#bib.bib6)\)and interpreting multimodal alignment\(Louet al\.,[2025](https://arxiv.org/html/2605.22902#bib.bib10)\)\. However, these approaches primarily focus on state\-based representations\. Our work bridges this gap by extending recent research into cross\-modal circuit tracing\(Yanget al\.,[2026](https://arxiv.org/html/2605.22902#bib.bib29)\)\. By utilizing Transcoders, we shift the focus toward the functional transformation of visual tokens, providing a causal account of how grounding behaves during inference\.

## 7Discussion and Future Work

Our central finding is that moving from state\-based to function\-based decomposition yields a more useful mechanistic account of generative Vision\-Language Models\. This shift enables the analysis in Sections[3\.3](https://arxiv.org/html/2605.22902#S3.SS3),[4](https://arxiv.org/html/2605.22902#S4),[5](https://arxiv.org/html/2605.22902#S5), by moving from mere feature attribution, towards a linked computation graph analysis across the network layers and input tokens\. The structural analysis of this graph shows that common failure modes of generative models, i\.e\., hallucination, can be explored through the lens of mechanistic interpretability by analyzing the properties of the traced computation circuit\.

Focusing on the hallucination detection result, a simple logistic classifier trained over five graph\-level metrics, without access to the generated text, output probabilities, or any external reference, distinguishes hallucinated from correct generations at AUC0\.680\.68versus0\.550\.55for a stratified\-prior baseline\. Although we make no claims of competitive performance against output\-level detectors, we identify a potential signal for hallucination in the mere structure of the LLM computation\.

Regarding future work, we plan to explore potential mechanistic interventions for hallucination mitigation based on our findings\. Furthermore, a common critique for mechanistic interpretability research regards the monosemanticity of extracted features, or lack thereof\. Indeed, in our analysis in Appendix[B\.5](https://arxiv.org/html/2605.22902#A2.SS5)we find that extracted features are not monosemantic, but we firmly believe that advancing mechanistic exploration of LLM behavior is a worthwhile pursuit, especially paired with more structural circuit trace analysis, made available through the use of functional decomposition frameworks, like Transcoders\. To this end, we plan to explore multimodal concept extraction, following approaches such as\(Helffet al\.,[2026](https://arxiv.org/html/2605.22902#bib.bib2)\)\. This could enable more interpretable ways of guiding model behavior, e\.g\., through visually grounded steering\. Finally, Transcoders provide further capabilities, such as de\-embeddings and virtual weights\(Dunefskyet al\.,[2024](https://arxiv.org/html/2605.22902#bib.bib24)\), which are not explored in this work\. Studying these mechanisms may offer additional insight into how VLMs represent and process multimodal information\.

## 8Limitations and Scope

While prior work has explored the use of Transcoders for circuit tracing in VLMs, our study extends this line of research with a more detailed analysis of cross\-modal representations in generative settings using attribution\-based and structural circuit analyses\. Nevertheless, several limitations should be considered\. First, our Transcoders are trained on a 500M\-token corpus\. Although this is sufficient to capture common visual–semantic patterns, it may not cover less frequent or more specialized forms of grounding\. Second, our analysis focuses on the Gemma 3\-4B\-IT model\. While this model provides a strong baseline, evaluating larger models \(e\.g\., 27B\+\) or alternative architectures such as Qwen\-VL would be important to assess the scalability and robustness of our findings\. Finally, our analysis is based on relatively small sample sizes\. Extending the evaluation to larger and more diverse datasets would help determine the generality of the observed patterns\.

## References

- Saes are good for steering–if you select the right features\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 10252–10270\.Cited by:[§1](https://arxiv.org/html/2605.22902#S1.p2.1)\.
- S\. Bai, Y\. Cai, R\. Chen, K\. Chen, X\. Chen, Z\. Cheng, L\. Deng, W\. Ding, C\. Gao, C\. Ge,et al\.\(2025\)Qwen3\-vl technical report\.arXiv preprint arXiv:2511\.21631\.Cited by:[§1](https://arxiv.org/html/2605.22902#S1.p1.1)\.
- L\. Bereska and E\. Gavves \(2024\)Mechanistic interpretability for ai safety–a review\.arXiv preprint arXiv:2404\.14082\.Cited by:[§1](https://arxiv.org/html/2605.22902#S1.p2.1)\.
- S\. Bills, N\. Cammarata, J\. Wu,et al\.\(2025\)Revising and falsifying sparse autoencoder feature explanations\.arXiv preprint arXiv:2502\.12345\.Cited by:[§6](https://arxiv.org/html/2605.22902#S6.p1.1)\.
- T\. Bricken, A\. Templeton, J\. Batson, B\. Chen, A\. Jermyn, T\. Conerly, N\. Turner, C\. Anil, C\. Denison, A\. Askell, R\. Lasenby, Y\. Wu, S\. Kravec, N\. Schiefer, T\. Maxwell, N\. Joseph, Z\. Hatfield\-Dodds, A\. Tamkin, K\. Nguyen, B\. McLean, J\. E\. Burke, T\. Hume, S\. Carter, T\. Henighan, and C\. Olah \(2023\)Towards monosemanticity: decomposing language models with dictionary learning\.Transformer Circuits Thread\.Note:Accessed: 2026\-04\-26External Links:[Link](https://transformer-circuits.pub/2023/monosemantic-features/index.html)Cited by:[§1](https://arxiv.org/html/2605.22902#S1.p2.1)\.
- R\. Chen, A\. Arditi, H\. Sleight, O\. Evans, and J\. Lindsey \(2025\)Persona vectors: monitoring and controlling character traits in language models\.arXiv preprint arXiv:2507\.21509\.Cited by:[§6](https://arxiv.org/html/2605.22902#S6.p1.1)\.
- M\. Cho, D\. Kim,et al\.\(2026\)CorrSteer: generation\-time llm steering via correlated sae features\.arXiv preprint arXiv:2601\.09876\.Cited by:[§6](https://arxiv.org/html/2605.22902#S6.p1.1)\.
- H\. Cunningham, A\. Ewart, L\. Riggs, R\. Huben, and L\. Sharkey \(2023\)Sparse autoencoders find highly interpretable features in language models\.arXiv preprint arXiv:2309\.08600\.Cited by:[§1](https://arxiv.org/html/2605.22902#S1.p2.1)\.
- J\. Dunefsky, P\. Chlenski, and N\. Nanda \(2024\)Transcoders find interpretable llm feature circuits\.Advances in Neural Information Processing Systems37,pp\. 24375–24410\.Cited by:[§1](https://arxiv.org/html/2605.22902#S1.p2.1),[§6](https://arxiv.org/html/2605.22902#S6.p1.1),[§7](https://arxiv.org/html/2605.22902#S7.p3.1)\.
- Y\. Fang, W\. Wang, M\. Xue, B\. Deng, F\. Xu, D\. Liu, and F\. Feng \(2026\)Controllable llm reasoning via sparse autoencoder\-based steering\.arXiv preprint arXiv:2601\.03595\.Cited by:[§6](https://arxiv.org/html/2605.22902#S6.p1.1)\.
- L\. Gao, T\. D\. la Tour, H\. Tillman, G\. Goh, R\. Troll, A\. Radford, I\. Sutskever, J\. Leike, and J\. Wu \(2024\)Scaling and evaluating sparse autoencoders\.arXiv preprint arXiv:2406\.04093\.Cited by:[§2\.2](https://arxiv.org/html/2605.22902#S2.SS2.p3.5)\.
- Gemma Team \(2025\)Gemma 3 technical report\.External Links:2503\.19786,[Link](https://arxiv.org/abs/2503.19786)Cited by:[§1](https://arxiv.org/html/2605.22902#S1.p1.1)\.
- Y\. Goyal, T\. Khot, D\. Summers\-Stay, D\. Batra, and D\. Parikh \(2017\)Making the v in vqa matter: elevating the role of image understanding in visual question answering\.InProceedings of the IEEE conference on computer vision and pattern recognition,pp\. 6904–6913\.Cited by:[§2\.2](https://arxiv.org/html/2605.22902#S2.SS2.p2.1)\.
- Y\. Heet al\.\(2025\)Towards global\-level mechanistic interpretability: a perspective of modular circuits of large language models\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§6](https://arxiv.org/html/2605.22902#S6.p1.1)\.
- L\. Helff, R\. Härle, W\. Stammer, F\. Friedrich, M\. Brack, A\. Wüst, H\. Shindo, P\. Schramowski, and K\. Kersting \(2026\)ActivationReasoning: logical reasoning in latent activation spaces\.External Links:2510\.18184,[Link](https://arxiv.org/abs/2510.18184)Cited by:[§6](https://arxiv.org/html/2605.22902#S6.p1.1),[§7](https://arxiv.org/html/2605.22902#S7.p3.1)\.
- J\. Johnson, B\. Hariharan, L\. Van Der Maaten, L\. Fei\-Fei, C\. Lawrence Zitnick, and R\. Girshick \(2017\)Clevr: a diagnostic dataset for compositional language and elementary visual reasoning\.InProceedings of the IEEE conference on computer vision and pattern recognition,pp\. 2901–2910\.Cited by:[§2\.2](https://arxiv.org/html/2605.22902#S2.SS2.p2.1)\.
- T\. Lieberum, V\. Veitch, S\. Ward\-Foxton,et al\.\(2024\)Gemma scope: open sparse autoencoders for gemma 2\.Google DeepMind Technical Report\.Cited by:[§2\.2](https://arxiv.org/html/2605.22902#S2.SS2.p4.1),[§6](https://arxiv.org/html/2605.22902#S6.p1.1)\.
- T\. Lin, M\. Maire, S\. Belongie, J\. Hays, P\. Perona, D\. Ramanan, P\. Dollár, and C\. L\. Zitnick \(2014\)Microsoft coco: common objects in context\.InEuropean Conference on Computer Vision,pp\. 740–755\.Cited by:[§2\.2](https://arxiv.org/html/2605.22902#S2.SS2.p2.1)\.
- J\. Lindsey, A\. Templeton, J\. Marcus, T\. Conerly, J\. Batson, and C\. Olah \(2024\)Sparse crosscoders for cross\-layer features and model diffing\.Transformer Circuits Thread\.Note:Accessed: 2026\-04\-26External Links:[Link](https://transformer-circuits.pub/2024/crosscoders/index.html)Cited by:[§1](https://arxiv.org/html/2605.22902#S1.p2.1)\.
- H\. Liu, C\. Li, Q\. Wu, and Y\. J\. Lee \(2023\)Visual instruction tuning\.Advances in neural information processing systems36,pp\. 34892–34916\.Cited by:[§1](https://arxiv.org/html/2605.22902#S1.p1.1)\.
- H\. Lou, C\. Li, J\. Ji, and Y\. Yang \(2025\)SAE\-v: interpreting multimodal models for enhanced alignment\.External Links:2502\.17514,[Link](https://arxiv.org/abs/2502.17514)Cited by:[§1](https://arxiv.org/html/2605.22902#S1.p3.1),[§6](https://arxiv.org/html/2605.22902#S6.p2.1)\.
- S\. M\. Lundberg and S\. Lee \(2017\)A unified approach to interpreting model predictions\.Advances in neural information processing systems30\.Cited by:[§5\.2](https://arxiv.org/html/2605.22902#S5.SS2.p4.1)\.
- S\. Mishra, D\. Khashabi, C\. Baral, and H\. Hajishirzi \(2022\)Cross\-task generalization via natural language crowdsourcing instructions\.InACL,Cited by:[§2\.2](https://arxiv.org/html/2605.22902#S2.SS2.p2.1)\.
- U\. Naseem, J\. Smith,et al\.\(2026\)Mechanistic interpretability for llm alignment: progress and challenges\.Journal of AI Research\.Cited by:[§6](https://arxiv.org/html/2605.22902#S6.p1.1)\.
- F\. Ortu, Z\. Jin, D\. Doimo, and A\. Cazzaniga \(2025\)When seeing overrides knowing: disentangling knowledge conflicts in vision\-language models\.External Links:2507\.13868,[Link](https://arxiv.org/abs/2507.13868)Cited by:[§1](https://arxiv.org/html/2605.22902#S1.p3.1),[§6](https://arxiv.org/html/2605.22902#S6.p2.1)\.
- M\. Pach, S\. Karthik, Q\. Bouniot, S\. Belongie, and Z\. Akata \(2025\)Sparse autoencoders learn monosemantic features in vision\-language models\.External Links:2504\.02821,[Link](https://arxiv.org/abs/2504.02821)Cited by:[§1](https://arxiv.org/html/2605.22902#S1.p3.1),[§6](https://arxiv.org/html/2605.22902#S6.p2.1)\.
- A\. Radford, J\. W\. Kim, C\. Hallacy, A\. Ramesh, G\. Goh, S\. Agarwal, G\. Sastry, A\. Askell, P\. Mishkin, J\. Clark,et al\.\(2021\)Learning transferable visual models from natural language supervision\.InInternational conference on machine learning,pp\. 8748–8763\.Cited by:[§1](https://arxiv.org/html/2605.22902#S1.p1.1),[§6](https://arxiv.org/html/2605.22902#S6.p2.1)\.
- S\. Sikdar, P\. Bhattacharya, and K\. Heese \(2021\)Integrated directional gradients: feature interaction attribution for neural NLP models\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),C\. Zong, F\. Xia, W\. Li, and R\. Navigli \(Eds\.\),Online,pp\. 865–878\.External Links:[Link](https://aclanthology.org/2021.acl-long.71/),[Document](https://dx.doi.org/10.18653/v1/2021.acl-long.71)Cited by:[§3\.1](https://arxiv.org/html/2605.22902#S3.SS1.p1.4)\.
- K\. Simonyan, A\. Vedaldi, and A\. Zisserman \(2013\)Deep inside convolutional networks: visualising image classification models and saliency maps\.arXiv preprint arXiv:1312\.6034\.Cited by:[§3\.1](https://arxiv.org/html/2605.22902#S3.SS1.p1.4)\.
- A\. Singh, A\. Fry, A\. Perelman, A\. Tart, A\. Ganesh, A\. El\-Kishky, A\. McLaughlin, A\. Low, A\. Ostrow, A\. Ananthram,et al\.\(2025\)Openai gpt\-5 system card\.arXiv preprint arXiv:2601\.03267\.Cited by:[§5](https://arxiv.org/html/2605.22902#S5.p2.2)\.
- S\. Soo, C\. Guang, W\. Teng, C\. Balaganesh, T\. Guoxian, and Y\. Ming \(2025\)Interpretable steering of large language models with feature guided activation additions\.arXiv preprint arXiv:2501\.09929\.Cited by:[§1](https://arxiv.org/html/2605.22902#S1.p2.1)\.
- S\. Stevens, W\. Chao, T\. Berger\-Wolf, and Y\. Su \(2025\)Interpretable and testable vision features via sparse autoencoders\.External Links:2502\.06755,[Link](https://arxiv.org/abs/2502.06755)Cited by:[§1](https://arxiv.org/html/2605.22902#S1.p3.1),[§6](https://arxiv.org/html/2605.22902#S6.p2.1)\.
- A\. I\. Team \(2025\)Insights on crosscoder model diffing\.Anthropic Technical Blog\.Cited by:[§6](https://arxiv.org/html/2605.22902#S6.p1.1)\.
- A\. Templeton, T\. Conerly, J\. Marcus, J\. Lindsay, T\. Bricken,et al\.\(2024\)Scaling monosemanticity: extracting interpretable features from claude 3 sonnet\.Anthropic Technical Report\.Cited by:[§6](https://arxiv.org/html/2605.22902#S6.p1.1)\.
- Z\. Wang, G\. Bingham, A\. W\. Yu, Q\. V\. Le, T\. Luong, and G\. Ghiasi \(2024\)HaloQuest: a visual hallucination dataset for advancing multimodal reasoning\.InEuropean Conference on Computer Vision,pp\. 288–304\.Cited by:[§5](https://arxiv.org/html/2605.22902#S5.p2.2)\.
- J\. Yang, T\. Xiong, S\. Qian, K\. Nahrstedt, and M\. Wu \(2026\)Circuit tracing in vision\-language models: understanding the internal mechanisms of multimodal thinking\.arXiv preprint arXiv:2602\.20330\.Cited by:[§1](https://arxiv.org/html/2605.22902#S1.p4.1),[§2\.2](https://arxiv.org/html/2605.22902#S2.SS2.p3.5),[§6](https://arxiv.org/html/2605.22902#S6.p2.1)\.
- P\. Young, A\. Lai, M\. Hodosh, and J\. Hockenmaier \(2014\)From image descriptions to visual denotations: new similarity metrics for semantic inference over event descriptions\.Transactions of the Association for Computational Linguistics2,pp\. 67–78\.Cited by:[§3](https://arxiv.org/html/2605.22902#S3.p1.1)\.
- V\. Zaigrajew, H\. Baniecki, and P\. Biecek \(2025\)Interpreting clip with hierarchical sparse autoencoders\.arXiv preprint arXiv:2502\.20578\.Cited by:[§1](https://arxiv.org/html/2605.22902#S1.p3.1)\.
- X\. Zhai, B\. Mustafa, A\. Kolesnikov, and L\. Beyer \(2023\)Sigmoid loss for language image pre\-training\.InProceedings of the IEEE/CVF international conference on computer vision,pp\. 11975–11986\.Cited by:[§1](https://arxiv.org/html/2605.22902#S1.p1.1)\.
- H\. Zhao, H\. Chen, F\. Yang, N\. Liu, H\. Deng, H\. Cai, S\. Wang, D\. Yin, and M\. Du \(2024\)Explainability for large language models: a survey\.ACM Transactions on Intelligent Systems and Technology15\(2\),pp\. 1–38\.Cited by:[§1](https://arxiv.org/html/2605.22902#S1.p2.1)\.

## Appendix ABroader Impact

This work advances mechanistic interpretability for Vision\-Language Models, with potential benefits for transparency and trustworthiness in high\-stakes deployments such as medical imaging or accessibility tools\. By tracing which image regions drive specific predictions, our framework provides a foundation for human\-auditable explanations of model behavior and offers a preliminary mechanistic signal for hallucination detection\. We release trained Transcoders, code, and datasets under permissive open licenses \(Apache 2\.0 and CC\-BY\-4\.0\) to lower the barrier to interpretability research for groups without access to large computational resources\. Practitioners should treat circuit traces and attribution maps as diagnostic approximations rather than ground\-truth explanations\. Misuse of interpretability tools to certify model safety on the basis of sparse circuit analysis alone could be harmful, and we encourage their use as one component among several in a broader evaluation pipeline\. We do not foresee significant dual\-use risks beyond those already present in existing gradient\-based saliency methods\.

## Appendix BAppendix and supplementary material

### B\.1Captioning visualizations \- Attribution maps

We present additional examples of patch ablation using thet​o​p−1top\-1,t​o​p−5top\-5, andt​o​p−10top\-10patches for both SAEs and Transcoders, and analyze their effects when zeroed out on probability and entropy\. Across all cases—\-particularly for thet​o​p−5top\-5andt​o​p−10top\-10selections—\-the patches chosen by the Transcoder show stronger alignment with the target token and produce a greater impact on both token probability and entropy\.

![Refer to caption](https://arxiv.org/html/2605.22902v1/ranked_grids_top1/ranked_grid_top1_001.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/ranked_grids_top1/ranked_grid_top1_002.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/ranked_grids_top1/ranked_grid_top1_004.png)

Figure 6:Top\-1 ablation comparison![Refer to caption](https://arxiv.org/html/2605.22902v1/ranked_grids_top5/ranked_grid_top5_001.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/ranked_grids_top5/ranked_grid_top5_002.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/ranked_grids_top5/ranked_grid_top5_005.png)

Figure 7:Top\-5 ablation comparison![Refer to caption](https://arxiv.org/html/2605.22902v1/ranked_grids_top5/ranked_grid_top5_006.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/ranked_grids_top5/ranked_grid_top5_007.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/ranked_grids_top5/ranked_grid_top5_008.png)

Figure 8:Top\-5 ablation comparison![Refer to caption](https://arxiv.org/html/2605.22902v1/ranked_grids_top10/ranked_grid_top10_002.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/ranked_grids_top10/ranked_grid_top10_005.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/ranked_grids_top10/ranked_grid_top10_004.png)

Figure 9:Top\-10 ablation comparison![Refer to caption](https://arxiv.org/html/2605.22902v1/ranked_grids_top10/ranked_grid_top10_006.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/ranked_grids_top10/ranked_grid_top10_007.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/ranked_grids_top10/ranked_grid_top10_008.png)

Figure 10:Top\-10 ablation comparison
### B\.2Captioning visualizations \- Circuit Analysis

In this section, we provide further examples illustrating how circuit tracing uncovers the dependency structure connecting visual regions to generated language\. Despite the limitations of our method, it reliably identifies interpretable links between visually grounded tokens and their corresponding image patches\.

![Refer to caption](https://arxiv.org/html/2605.22902v1/captions/token_285_copy.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/captions/token_285.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/captions/token_292_copy.png)

Figure 11:Circuit analysis results![Refer to caption](https://arxiv.org/html/2605.22902v1/captions/token_287_copy.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/captions/token_287.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/captions/token_292_copy_4.png)

Figure 12:Circuit analysis results![Refer to caption](https://arxiv.org/html/2605.22902v1/captions/token_291.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/captions/token_292_copy_2.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/captions/token_292_copy_3.png)

Figure 13:Circuit analysis results
### B\.3False Visual Grounding \- Attribution maps

We present additional examples of patch ablation using thet​o​p−1top\-1,t​o​p−5top\-5, andt​o​p−10top\-10patches for Transcoders in the FVG setting, and analyze the effects of zeroing them out on probability and entropy\. Across all cases, the regions selected by the Transcoder show little correlation with the target token, which is further supported by the minimal change in token probability observed when these patches are removed\.

![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg_attr_maps_top1/entropy_ranked_top1_001.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg_attr_maps_top1/entropy_ranked_top1_002.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg_attr_maps_top1/entropy_ranked_top1_003.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg_attr_maps_top1/entropy_ranked_top1_006.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg_attr_maps_top1/entropy_ranked_top1_005.png)

Figure 14:Top\-1 FVG ablation![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg_attr_maps_top5/entropy_ranked_top5_001.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg_attr_maps_top5/entropy_ranked_top5_002.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg_attr_maps_top5/entropy_ranked_top5_003.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg_attr_maps_top5/entropy_ranked_top5_004.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg_attr_maps_top5/entropy_ranked_top5_005.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg_attr_maps_top5/entropy_ranked_top5_006.png)

Figure 15:Top\-5 FVG comparison![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg_attr_maps_top10/entropy_ranked_top10_001.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg_attr_maps_top10/entropy_ranked_top10_002.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg_attr_maps_top10/entropy_ranked_top10_003.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg_attr_maps_top10/entropy_ranked_top10_005.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg_attr_maps_top10/entropy_ranked_top10_006.png)

Figure 16:Top\-10 FVG ablation
### B\.4False Visual Grounding \- Circuit Analysis

In this section, we provide additional circuit tracing examples in the False Visual Grounding setting\. These results further support the claim that Transcoders identify semantically meaningful connections between visually grounded tokens and image patches, rather than merely decomposing MLP activations\.

![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg/token_286.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg/token_287.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg/token_288.png)

Figure 17:Circuit analysis on FVG![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg/token_289.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg/token_290.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg/token_291.png)

Figure 18:Circuit analysis on FVG![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg/token_293.png)
![Refer to caption](https://arxiv.org/html/2605.22902v1/zvg/token_294.png)

Figure 19:Circuit analysis on FVG
### B\.5Monosemanticity Investigation

To investigate feature monosemanticity, we analyze how activated features are shared across tokens in different image contexts\. Specifically, we examine the overlap of features across tokens to determine whether they are uniquely tied to individual tokens or shared more broadly\. For this analysis, we use 10,000 samples from the Flickr30K dataset\.

Table 5:Distribution of feature sharing per layer across the corpus\.Table[5](https://arxiv.org/html/2605.22902#A2.T5)provides a breakdown of how frequently features are shared per layer across different tokens\. The data indicates that monosemanticity is rare, with only 4 out of 40,960 active features \(0\.01%\) activating for a single token\. Instead, a majority of features \(51\.4%\) are active for between 21 and 100 distinct tokens\. Additionally, 45 features appear to be "universal," firing for every token in the analyzed corpus\. These numbers suggest that the model relies on a distributed representational scheme rather than a one\-to\-one mapping between features and concepts\.

The analysis further shows an inverse relationship between token frequency and the number of features activated\. Rare tokens \(occurring fewer than 20 times\) activate an average of 49\.4 features per layer, while high\-frequency tokens activate only 35\.4 features\. This approximately 40% increase for rare tokens suggests that the model may use larger combinations of features to represent more specific or less common concepts\. These results indicate that the Transcoder represents information through overlapping sets of features rather than through dedicated, monosemantic units\.

Similar Articles

Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding

arXiv cs.AI

This paper reveals that hallucination in large vision-language models is caused by a dynamic structural misalignment where certain attention heads act as risky mediators, decoupling from visual evidence to lock onto language priors. The authors propose Fox, a training-free causal intervention framework that diagnoses and physically severs these pathological shortcuts, achieving state-of-the-art performance in faithful decoding.

From Architecture to Output: Structural Origins of Hallucination in Large Language Models and the Amplifying Role of Data

arXiv cs.AI

This paper analyzes hallucination in large language models as a structural consequence of three architectural decisions: self-attention's co-occurrence learning, maximum likelihood estimation training objective, and autoregressive decoding's left-to-right commitment. It maps each mechanism to specific hallucination types and argues that dataset pathologies amplify but do not cause these vulnerabilities.

Mechanisms of Prompt-Induced Hallucination in Vision-Language Models

arXiv cs.CL

This paper investigates prompt-induced hallucinations in vision-language models through mechanistic analysis, identifying specific attention heads responsible for the models' tendency to favor textual prompts over visual evidence. The authors demonstrate that ablating these PIH-heads reduces hallucinations by at least 40% without additional training, revealing model-specific mechanisms underlying this failure mode.

Hallucination as Trajectory Commitment: Causal Evidence for Asymmetric Attractor Dynamics in Transformer Generation

arXiv cs.CL

This paper presents causal evidence that hallucination in autoregressive language models results from early trajectory commitment governed by asymmetric attractor dynamics, using same-prompt bifurcation and activation patching experiments on Qwen2.5-1.5B to show that hallucinated trajectories diverge at the first token and exhibit strong causal asymmetry across model layers.