mechanistic-interpretability

Tag

Cards List
#mechanistic-interpretability

Are Stated Reasoning Steps Causally Load-Bearing?

arXiv cs.AI ↗ · yesterday Cached

This paper uses activation patching to causally measure if reasoning steps in chain-of-thought are load-bearing, finding that behavioral tests overestimate faithfulness and larger models like Qwen3-4B maintain better faithfulness across reasoning depths.

0 favorites 0 likes
#mechanistic-interpretability

Rewired or Gated? How Instruction Tuning Shapes Knowledge-Conflict Circuits in LLMs

arXiv cs.LG ↗ · 2d ago Cached

This paper provides a mechanistic comparison of knowledge-conflict circuits in LLMs under instruction tuning, finding that tuning gates rather than rewires these circuits across multiple model families, with implications for interpretability.

0 favorites 0 likes
#mechanistic-interpretability

Matryoshka attribution: Learning to attribute language model outputs to representations and weights

arXiv cs.CL ↗ · 2d ago Cached

The paper introduces Matryoshka Attribution (MAttr), a mask-learning method for attributing language model outputs to internal components, achieving top performance on the Mechanistic Interpretability Benchmark and demonstrating practical use in modifying LLM behaviors by adjusting weights.

0 favorites 0 likes
#mechanistic-interpretability

Dissecting Hierarchical Reasoning Models: A Mechanistic Study

arXiv cs.LG ↗ · 3d ago Cached

This paper provides a mechanistic analysis of Hierarchical Reasoning Models (HRM) to understand their internal reasoning processes in latent space, using techniques like causal interventions and sparse autoencoders on tasks such as Sudoku and ARC-AGI-2.

0 favorites 0 likes
#mechanistic-interpretability

MechaTerp-TRACE: A Novel Approach for Component Ablation Analysis in Language Models

arXiv cs.CL ↗ · 3d ago Cached

The paper introduces MechaTerp-TRACE, a method for component ablation analysis in language models, finding that entity knowledge is largely attributable to generic generation machinery rather than localized components.

0 favorites 0 likes
#mechanistic-interpretability

Read-Best Is Not Steer-Best: A Probing--Steering Layer Dissociation in Omni-Modal Large Language Models

arXiv cs.CL ↗ · 3d ago Cached

This paper reveals that the optimal layer for linear probing to read concepts differs from the optimal layer for activation steering in omni-modal large language models, challenging common heuristics in representation engineering.

0 favorites 0 likes
#mechanistic-interpretability

If the AI Industry Followed Its Own Research, It Might Have Paused Already

Wired ↗ · 6d ago Cached

The article explores how the AI industry might have paused development if it followed its own research on catastrophic risks, highlighting Anthropic's findings on AI deception and the challenges of mechanistic interpretability.

0 favorites 0 likes
#mechanistic-interpretability

A Four-Stage Decomposition of Word-Problem Solving and Mechanistic Fragility in LLM Math Reasoning

arXiv cs.AI ↗ · 2026-09-17 Cached

This paper proposes a four-stage mechanistic decomposition of how large language models solve math word problems, localizing fragility from irrelevant clauses to the Operation Planning stage via specific attention heads.

0 favorites 0 likes
#mechanistic-interpretability

Relation Before Entity: Deferred Commitment in Language Model Factual Recall

arXiv cs.CL ↗ · 2026-09-17 Cached

The paper finds that in language models, relation-type information becomes generation-controlling before entity-specific information during factual recall, indicating a temporal asymmetry that is robust across various models and prompt families.

0 favorites 0 likes
#mechanistic-interpretability

@Jack_W_Lindsey: Here are the questions that currently seem most important to me: -- Better methods for "mind-reading" model activations…

X AI KOLs Timeline ↗ · 2026-09-16 Cached

The tweet discusses key questions in AI interpretability, specifically the advancement of methods for decoding neural network activations into human-readable language.

0 favorites 0 likes
#mechanistic-interpretability

Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration

arXiv cs.LG ↗ · 2026-09-16 Cached

Decoy Direction Optimization (DDO) is a fast, post-hoc defense method that protects open-weight LLMs from refusal feature ablation attacks by injecting decoy signals into the network, achieving high robustness at lower cost than trained defenses.

0 favorites 0 likes
#mechanistic-interpretability

Latent Undertow: How Ordinary Typos Break Probes

arXiv cs.CL ↗ · 2026-09-16 Cached

This paper demonstrates that ordinary typos in user inputs can significantly reduce the effectiveness of activation probes designed to detect prompt injection in LLMs, but introduces a KV-cache fork method that recovers most of the performance loss.

0 favorites 0 likes
#mechanistic-interpretability

Harmfulness Propagation Dynamics: Layer-wise Trajectories of Adversarial Intent in Large Language Models

arXiv cs.CL ↗ · 2026-09-15 Cached

The paper identifies Harmfulness Propagation Dynamics in large language models and introduces Herald, a lightweight input moderator that uses cross-layer activation patterns to detect harmful prompts efficiently.

0 favorites 0 likes
#mechanistic-interpretability

A Mathematical Framework for Transformer Circuits (2021)

Hacker News Top ↗ · 2026-09-12 Cached

This paper presents a mathematical framework for reverse-engineering transformer models to enhance mechanistic interpretability and address safety concerns in AI systems.

0 favorites 0 likes
#mechanistic-interpretability

Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations

arXiv cs.AI ↗ · 2026-09-11 Cached

This paper introduces methods to calibrate confidence in agentic systems using internal representations, demonstrating improved performance over baselines in multi-turn benchmarks.

0 favorites 0 likes
#mechanistic-interpretability

Capsule Lens: Locating and Tracking Concept Geometry in Model Representations

arXiv cs.LG ↗ · 2026-09-10 Cached

Capsule Lens is a framework for mechanistic interpretability that uses geometric capsules to locate and track how concepts are encoded in neural network representations, enabling analysis of both static and dynamic model internals.

0 favorites 0 likes
#mechanistic-interpretability

Contrastive Projection: Reading Transformer Internals by Differencing Logit Lenses

arXiv cs.CL ↗ · 2026-09-10 Cached

This paper introduces a contrastive projection technique to read transformer internal states by differencing logit lenses from paired prompts, enabling the identification of computation-specific differences across architectures like Phi-2.

0 favorites 0 likes
#mechanistic-interpretability

SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?

Hugging Face Daily Papers ↗ · 2026-09-08 Cached

This paper introduces SAEScientist-Bench, a benchmark to evaluate AI agents' ability to autonomously conduct mechanistic interpretability research using sparse autoencoders, revealing progress but significant gaps compared to expert baselines.

0 favorites 0 likes
#mechanistic-interpretability

ObserverBench: Testing Mechanistic Estimates for Intervention and Control

arXiv cs.LG ↗ · 2026-09-04 Cached

ObserverBench is a benchmark framework designed to evaluate whether internal estimators in AI models are adequate for guiding interventions, control, or safety tasks by reporting both estimation accuracy and downstream action performance.

0 favorites 0 likes
#mechanistic-interpretability

Interpretable Symptom Vectors for Depression in a Large Language Model

arXiv cs.CL ↗ · 2026-09-03 Cached

This paper uses mechanistic interpretability on Gemma-3-27B-PT to extract and align symptom vectors for depression with clinician judgments, demonstrating potential for interpretable clinical assessment tools.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback