activation-patching

Tag

Cards List
#activation-patching

Test, then Route: How Language Models Execute In-Context Conditional Rules Across Models and Languages

arXiv cs.CL · 5d ago Cached

This paper investigates how language models execute in-context conditional rules by probing whether testing and routing are separable mechanisms. Using activation patching across three open models and six languages, the authors find that predicate testing is modular while route representations are token-bound and non-transferable.

0 favorites 0 likes
#activation-patching

Visualizing Graph-to-Answer Mechanism Recovery in Materials-Science Hypothesis Generation

arXiv cs.CL · 5d ago Cached

This paper presents a graph-to-answer mechanism-tracing case study for Graph-PRefLexOR-8B, a materials-science hypothesis generation model, using visualization and activation-based diagnostics to localize where mechanism support is lost or recovered during generation.

0 favorites 0 likes
#activation-patching

Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction?

arXiv cs.CL · 2026-08-03 Cached

This paper tests whether increasing interventional data in pretraining improves LLMs' causal direction reasoning, using controlled Simpson's-paradox worlds. It finds that the training mixture does not govern interventional evidence use; instead the evidence type in the inference-time context is the decisive factor.

0 favorites 0 likes
#activation-patching

Cost Accounting for Reactive Computational Graphs: Exhaustive Sweeps, Sequential Mutation, and the Backward-Locality Gap

arXiv cs.LG · 2026-07-22 Cached

This paper presents a theoretical cost accounting for exhaustive sweeps and sequential mutations on reactive computational graphs, deriving speedup ratios and validating them on a Julia-based graph engine.

0 favorites 0 likes
#activation-patching

Are Arithmetic Heuristic Neurons Form-Invariant? A Mechanistic Analysis of Symbols, Text, and Code in LLMs

arXiv cs.CL · 2026-07-21 Cached

This paper investigates whether arithmetic heuristic neurons in LLMs are form-invariant across symbolic arithmetic, natural language word problems, and Python code. Using activation patching, they find a shared circuit of neurons that is necessary and sufficient for arithmetic computation, and that cross-format failures arise from activation states rather than distinct circuits.

0 favorites 0 likes
#activation-patching

Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models

arXiv cs.CL · 2026-06-29 Cached

This paper investigates how vision-language models resolve conflicts between visual evidence and world knowledge, revealing that visual grounding is the default while prior knowledge depends on a small set of late-layer attention heads. The authors perform causal analysis across three VLM families, demonstrating an asymmetric structure where ablating these heads shifts predictions from knowledge-grounded to visually grounded answers.

0 favorites 0 likes
#activation-patching

The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching

arXiv cs.LG · 2026-06-29 Cached

This paper re-derives activation patching from causal mediation analysis, revealing that the natural indirect effect (NIE) captures not only a component's causal effect but also interaction effects with other components. It demonstrates these hidden interactions in the GPT-2 IOI circuit and argues that they are a diagnostic tool rather than a nuisance.

0 favorites 0 likes
#activation-patching

Vernier: Probing Representational Misalignment Behind Lexical Gaps in Causal Reasoning

arXiv cs.CL · 2026-06-16 Cached

This paper investigates why instruction-tuned language models give different answers to causal reasoning questions when variable names are replaced with placeholders, finding that the issue stems from representational misalignment rather than information loss. The authors introduce Vernier, a method using paired-view weight updates and mechanism inspection to reveal that answer-relevant content is still present in the placeholder view but misaligned.

0 favorites 0 likes
#activation-patching

Measuring the Depth of LLM Unlearning via Activation Patching

arXiv cs.CL · 2026-05-26 Cached

The paper proposes the Unlearning Depth Score (UDS), a metric that uses activation patching to quantify how thoroughly target knowledge is erased from LLMs, achieving state-of-the-art faithfulness and robustness across multiple unlearning methods.

0 favorites 0 likes
#activation-patching

Interaction Locality in Hierarchical Recursive Reasoning

arXiv cs.AI · 2026-05-22 Cached

Proposes interaction locality, a task-geometry-aware framework for measuring whether information flow in spatial reasoning models stays within local cells or crosses into global structure, and applies it to HRM, TRM, and MTU3D models on grid benchmarks and embodied 3D grounding.

0 favorites 0 likes
#activation-patching

From Correlation to Cause: A Five-Stage Methodology for Feature Analysis in Transformer Language Models

arXiv cs.CL · 2026-05-22 Cached

This paper proposes a five-stage methodology for causal feature analysis in transformer language models, demonstrated on GPT-2 small for the IOI task. It finds that features are specifically causal but not necessary, and exposes a gap between detection and causal robustness.

0 favorites 0 likes
#activation-patching

Negative Before Positive: Asymmetric Valence Processing in Large Language Models

arXiv cs.CL · 2026-05-08 Cached

This paper investigates how large language models process emotional valence through mechanistic interpretability. Using activation patching and steering on three open-source LLMs, the authors find that negative valence is localized to early layers while positive valence peaks in mid-to-late layers, and they validate this through topic-controlled flip tests.

0 favorites 0 likes
#activation-patching

Hallucination as Trajectory Commitment: Causal Evidence for Asymmetric Attractor Dynamics in Transformer Generation

arXiv cs.CL · 2026-04-20 Cached

This paper presents causal evidence that hallucination in autoregressive language models results from early trajectory commitment governed by asymmetric attractor dynamics, using same-prompt bifurcation and activation patching experiments on Qwen2.5-1.5B to show that hallucinated trajectories diverge at the first token and exhibit strong causal asymmetry across model layers.

0 favorites 0 likes
← Back to home

Submit Feedback