mechanistic-interpretability

Tag

Cards List
#mechanistic-interpretability

ObserverBench: Testing Mechanistic Estimates for Intervention and Control

arXiv cs.LG ↗ · 2026-09-04 Cached

ObserverBench is a benchmark framework designed to evaluate whether internal estimators in AI models are adequate for guiding interventions, control, or safety tasks by reporting both estimation accuracy and downstream action performance.

0 favorites 0 likes
#mechanistic-interpretability

Interpretable Symptom Vectors for Depression in a Large Language Model

arXiv cs.CL ↗ · 2026-09-03 Cached

This paper uses mechanistic interpretability on Gemma-3-27B-PT to extract and align symptom vectors for depression with clinician judgments, demonstrating potential for interpretable clinical assessment tools.

0 favorites 0 likes
#mechanistic-interpretability

From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling

arXiv cs.CL ↗ · 2026-09-02 Cached

This paper introduces a mechanistic approach to improve LLM safety by characterizing a circuit for refusal behavior and using circuit-guided weight scaling, enhancing safety rates by 26.5% under attacks with minimal accuracy loss.

0 favorites 0 likes
#mechanistic-interpretability

How Language Models Choose Sides: Internal Representations of Instruction Hierarchy

arXiv cs.AI ↗ · 2026-09-01 Cached

This paper investigates how instruction-tuned LLMs arbitrate conflicts between system and user instructions, revealing that user-preferring behavior coexists with a readable internal arbitration signal, as demonstrated through mechanistic interpretability techniques and steering experiments.

0 favorites 0 likes
#mechanistic-interpretability

A Unifying Perspective on Language Model Representations: From Filler-Role Structure to Mechanistic Interpretability

arXiv cs.CL ↗ · 2026-09-01 Cached

This paper proposes Tensor Product Representations as a unifying framework for language model interpretability, showing mathematically and empirically that it integrates methods like additive analogies, linear probing, sparse autoencoders, and activation patching.

0 favorites 0 likes
#mechanistic-interpretability

Causal Interventions Reveal Typologically Organized Syntactic Mechanisms in Multilingual Language Models

arXiv cs.CL ↗ · 2026-09-01 Cached

This paper uses causal interventions to investigate syntactic mechanisms in multilingual language models, revealing cross-lingual transfer that is graded based on typological similarity.

0 favorites 0 likes
#mechanistic-interpretability

Mechanistic Circuit Identification for Controllable Data Generation

arXiv cs.LG ↗ · 2026-08-26 Cached

This paper proposes a circuit-grounded framework that leverages mechanistic interpretability for controllable data generation in language models, introducing SAMS for stage-aware data scheduling to improve training performance.

0 favorites 0 likes
#mechanistic-interpretability

Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders

arXiv cs.LG ↗ · 2026-08-26 Cached

This research investigates whether multilingual large language models develop shared internal representations for mathematical reasoning across languages, introducing a novel Geometry-Invariant Sparse Autoencoder (GI-SAE) method. It finds that cross-language feature sharing is model-dependent and that geometric similarity does not consistently imply functional interchangeability.

0 favorites 0 likes
#mechanistic-interpretability

Align, Unify, Suppress, Route: A Coherentist View of Transformer Computation

arXiv cs.CL ↗ · 2026-08-25 Cached

The paper introduces Coherentist Probabilistic Compositionalism (CPC), a framework for interpreting transformer computation through four operator roles, and validates it across 15 models from five architecture families.

0 favorites 0 likes
#mechanistic-interpretability

Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability

arXiv cs.LG ↗ · 2026-08-21 Cached

The paper proposes mechanistic tomography as a unified framework for designing measurements to recover internal mechanisms in AI models, improving control-oriented interpretability through interventions and calibration.

0 favorites 0 likes
#mechanistic-interpretability

Mechanistic Interpretability of Structure-Aware Numerical Reasoning in LLaMA 3.1 8B

arXiv cs.LG ↗ · 2026-08-20 Cached

This paper uses mechanistic interpretability to analyze how LLaMA 3.1 8B models numerical sequences, revealing that it internally computes and stores first differences to extrapolate patterns.

0 favorites 0 likes
#mechanistic-interpretability

Bidirectional representational alignment between biological and artificial neural networks

arXiv cs.LG ↗ · 2026-08-20 Cached

This paper presents a computational framework for steering representational geometry to improve bidirectional alignment between biological and artificial neural networks, showing a 55% relative enhancement in bidirectional predictivity.

0 favorites 0 likes
#mechanistic-interpretability

Abliteration Mitigation via Refusal Aliases

arXiv cs.CL ↗ · 2026-08-20 Cached

The paper introduces AMRA, a weight-editing method to mitigate abliteration in large language models by obscuring the refusal signal, improving post-abliteration refusal scores on Llama-3-8B and Gemma-2-9B with minimal utility degradation.

0 favorites 0 likes
#mechanistic-interpretability

J-Miner: Recovering Executable Decision Knowledge from Language-Model Classifiers

arXiv cs.LG ↗ · 2026-08-19 Cached

J-Miner recovers executable decision knowledge from fine-tuned language-model classifiers by mining named concepts and learning decision rules, enabling inspection and transfer to lightweight models with high fidelity.

0 favorites 0 likes
#mechanistic-interpretability

When Do Concepts Become Functionally Sufficient During Language-Model Training?

arXiv cs.CL ↗ · 2026-08-18 Cached

This paper introduces a checkpoint-wise activation-intervention framework to study how concepts become functionally sufficient during language model training, comparing preservation targets across layers and models.

0 favorites 0 likes
#mechanistic-interpretability

Explanation Multiplicity: Circuit-Level Interpretability Evidence Does Not Survive Defensible Analytic Variation

arXiv cs.AI ↗ · 2026-08-17 Cached

The study shows that circuit-level interpretability evidence for AI systems exhibits high variability across analytic settings, failing to meet consistency standards required by regulations like the EU AI Act.

0 favorites 0 likes
#mechanistic-interpretability

A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure

arXiv cs.AI ↗ · 2026-08-17 Cached

This paper introduces a calibrated test of internal action maps in language models, showing that state signals can be decodable and causally usable without global affine closure, using an evidence lattice framework validated on finite worlds and the Qwen3-4B model.

0 favorites 0 likes
#mechanistic-interpretability

Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia

arXiv cs.LG ↗ · 2026-08-14 Cached

This paper introduces PRISM, a perturbation-based method for spatially resolved interpretability of large language models, adapting neuroimaging subtraction analysis to transformers and applying it in parallel to post-stroke aphasia patients to recover shared phonemic-favoring dissociations.

0 favorites 0 likes
#mechanistic-interpretability

Geometric and Behavioral Stratification in Transformer Residual Streams

arXiv cs.LG ↗ · 2026-08-14 Cached

This paper investigates the residual stream geometry of trained transformers, showing that the prediction direction acts as a privileged anchor that stratifies residual-stream variation into narrow, readout-relevant and broad, computational regions across 18 models. The findings have implications for interpretability and evaluation.

0 favorites 0 likes
#mechanistic-interpretability

Falsehood and Impossibility Are Different Directions in an AI's Representation of Language

arXiv cs.CL ↗ · 2026-08-14 Cached

This paper reports an activation study of Gemma 3 4B IT showing that the model's internal representations distinguish necessary falsehoods (impossibilities) from contingent falsehoods, with impossibility directions orthogonal to truth directions and overlapping with semantic anomaly directions.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback