activation-patching

Tag

Cards List
#activation-patching

A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID

arXiv cs.CL ↗ · 3d ago Cached

The paper investigates neurons in frozen BERT that drive AI-text detection using sparse probing and activation patching on the RAID benchmark, identifying a small set of causally relevant neurons that generalize across generator families.

0 favorites 0 likes
#activation-patching

Are Stated Reasoning Steps Causally Load-Bearing?

arXiv cs.AI ↗ · 2026-09-24 Cached

This paper uses activation patching to causally measure if reasoning steps in chain-of-thought are load-bearing, finding that behavioral tests overestimate faithfulness and larger models like Qwen3-4B maintain better faithfulness across reasoning depths.

0 favorites 0 likes
#activation-patching

Inside VLM Chart Reading: Tracing Value Reading from Vertical Bar Charts Across Space and Depth

arXiv cs.CL ↗ · 2026-09-15 Cached

This paper investigates how vision-language models read exact values from vertical bar charts using controlled counterfactual activation patching, revealing insights into internal computations and differences between models like Qwen and InternVL.

0 favorites 0 likes
#activation-patching

Mechanistic Interpretability of Structure-Aware Numerical Reasoning in LLaMA 3.1 8B

arXiv cs.LG ↗ · 2026-08-20 Cached

This paper uses mechanistic interpretability to analyze how LLaMA 3.1 8B models numerical sequences, revealing that it internally computes and stores first differences to extrapolate patterns.

0 favorites 0 likes
#activation-patching

Shared Circuits for Shared Grammar: Tracing Subject-Verb Agreement Across Languages

arXiv cs.CL ↗ · 2026-08-20 Cached

This research investigates how multilingual large language models internally handle subject-verb agreement across languages, finding that models reuse shared computational structure for languages with overt inflection, indicating cross-lingual overlap in morphosyntactic processing.

0 favorites 0 likes
#activation-patching

LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning

arXiv cs.CL ↗ · 2026-08-14 Cached

This paper argues that LLM failures on hidden-constraint reasoning are routing problems, not knowledge problems, and introduces a quartet diagnostic to dissociate knowledge, symmetry, routing, and repair across 14 models, with activation probing and patching experiments.

0 favorites 0 likes
#activation-patching

Test, then Route: How Language Models Execute In-Context Conditional Rules Across Models and Languages

arXiv cs.CL ↗ · 2026-08-06 Cached

This paper investigates how language models execute in-context conditional rules by probing whether testing and routing are separable mechanisms. Using activation patching across three open models and six languages, the authors find that predicate testing is modular while route representations are token-bound and non-transferable.

0 favorites 0 likes
#activation-patching

Visualizing Graph-to-Answer Mechanism Recovery in Materials-Science Hypothesis Generation

arXiv cs.CL ↗ · 2026-08-06 Cached

This paper presents a graph-to-answer mechanism-tracing case study for Graph-PRefLexOR-8B, a materials-science hypothesis generation model, using visualization and activation-based diagnostics to localize where mechanism support is lost or recovered during generation.

0 favorites 0 likes
#activation-patching

Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction?

arXiv cs.CL ↗ · 2026-08-03 Cached

This paper tests whether increasing interventional data in pretraining improves LLMs' causal direction reasoning, using controlled Simpson's-paradox worlds. It finds that the training mixture does not govern interventional evidence use; instead the evidence type in the inference-time context is the decisive factor.

0 favorites 0 likes
#activation-patching

Cost Accounting for Reactive Computational Graphs: Exhaustive Sweeps, Sequential Mutation, and the Backward-Locality Gap

arXiv cs.LG ↗ · 2026-07-22 Cached

This paper presents a theoretical cost accounting for exhaustive sweeps and sequential mutations on reactive computational graphs, deriving speedup ratios and validating them on a Julia-based graph engine.

0 favorites 0 likes
#activation-patching

Are Arithmetic Heuristic Neurons Form-Invariant? A Mechanistic Analysis of Symbols, Text, and Code in LLMs

arXiv cs.CL ↗ · 2026-07-21 Cached

This paper investigates whether arithmetic heuristic neurons in LLMs are form-invariant across symbolic arithmetic, natural language word problems, and Python code. Using activation patching, they find a shared circuit of neurons that is necessary and sufficient for arithmetic computation, and that cross-format failures arise from activation states rather than distinct circuits.

0 favorites 0 likes
#activation-patching

Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models

arXiv cs.CL ↗ · 2026-06-29 Cached

This paper investigates how vision-language models resolve conflicts between visual evidence and world knowledge, revealing that visual grounding is the default while prior knowledge depends on a small set of late-layer attention heads. The authors perform causal analysis across three VLM families, demonstrating an asymmetric structure where ablating these heads shifts predictions from knowledge-grounded to visually grounded answers.

0 favorites 0 likes
#activation-patching

The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching

arXiv cs.LG ↗ · 2026-06-29 Cached

This paper re-derives activation patching from causal mediation analysis, revealing that the natural indirect effect (NIE) captures not only a component's causal effect but also interaction effects with other components. It demonstrates these hidden interactions in the GPT-2 IOI circuit and argues that they are a diagnostic tool rather than a nuisance.

0 favorites 0 likes
#activation-patching

Vernier: Probing Representational Misalignment Behind Lexical Gaps in Causal Reasoning

arXiv cs.CL ↗ · 2026-06-16 Cached

This paper investigates why instruction-tuned language models give different answers to causal reasoning questions when variable names are replaced with placeholders, finding that the issue stems from representational misalignment rather than information loss. The authors introduce Vernier, a method using paired-view weight updates and mechanism inspection to reveal that answer-relevant content is still present in the placeholder view but misaligned.

0 favorites 0 likes
#activation-patching

Measuring the Depth of LLM Unlearning via Activation Patching

arXiv cs.CL ↗ · 2026-05-26 Cached

The paper proposes the Unlearning Depth Score (UDS), a metric that uses activation patching to quantify how thoroughly target knowledge is erased from LLMs, achieving state-of-the-art faithfulness and robustness across multiple unlearning methods.

0 favorites 0 likes
#activation-patching

Interaction Locality in Hierarchical Recursive Reasoning

arXiv cs.AI ↗ · 2026-05-22 Cached

Proposes interaction locality, a task-geometry-aware framework for measuring whether information flow in spatial reasoning models stays within local cells or crosses into global structure, and applies it to HRM, TRM, and MTU3D models on grid benchmarks and embodied 3D grounding.

0 favorites 0 likes
#activation-patching

From Correlation to Cause: A Five-Stage Methodology for Feature Analysis in Transformer Language Models

arXiv cs.CL ↗ · 2026-05-22 Cached

This paper proposes a five-stage methodology for causal feature analysis in transformer language models, demonstrated on GPT-2 small for the IOI task. It finds that features are specifically causal but not necessary, and exposes a gap between detection and causal robustness.

0 favorites 0 likes
#activation-patching

Negative Before Positive: Asymmetric Valence Processing in Large Language Models

arXiv cs.CL ↗ · 2026-05-08 Cached

This paper investigates how large language models process emotional valence through mechanistic interpretability. Using activation patching and steering on three open-source LLMs, the authors find that negative valence is localized to early layers while positive valence peaks in mid-to-late layers, and they validate this through topic-controlled flip tests.

0 favorites 0 likes
#activation-patching

Hallucination as Trajectory Commitment: Causal Evidence for Asymmetric Attractor Dynamics in Transformer Generation

arXiv cs.CL ↗ · 2026-04-20 Cached

This paper presents causal evidence that hallucination in autoregressive language models results from early trajectory commitment governed by asymmetric attractor dynamics, using same-prompt bifurcation and activation patching experiments on Qwen2.5-1.5B to show that hallucinated trajectories diverge at the first token and exhibit strong causal asymmetry across model layers.

0 favorites 0 likes
← Back to home

Submit Feedback