@noahdgoodman: how should we find representations in neural nets? how about the way we find everything else these days — gradient dece…
Summary
A new paper introduces Matryoshka Attribution, a method using gradient descent to identify responsible parts of neural networks, achieving top performance on the Mechanistic Interpretability Benchmark.
View Cached Full Text
Cached at: 09/27/26, 07:08 AM
how should we find representations in neural nets? how about the way we find everything else these days — gradient decent!
Aryaman Arora (@aryaman2020): New paper! 🫡
We introduce Matryoshka Attribution, a new attribution method which uses gradient descent to find which parts of a neural network are responsible for a behaviour.
MAttr is #1 on the Mechanistic Interpretability Benchmark by a wide margin (2.9× the runner up).
Similar Articles
Matryoshka attribution: Learning to attribute language model outputs to representations and weights
The paper introduces Matryoshka Attribution (MAttr), a mask-learning method for attributing language model outputs to internal components, achieving top performance on the Mechanistic Interpretability Benchmark and demonstrating practical use in modifying LLM behaviors by adjusting weights.
Mechanistic interpretability: a first paper on disentangling a convolutional neuron [R]
This paper introduces a technique to disentangle a single convolutional neuron in Inceptionv1 by analyzing Hadamard products, revealing clean monosemantic clusters (cars, cats, dogs) and also low-valued clusters (letters, faces) with distributed weights as evidence of gradient descent behavior.
Functional Gradient Descent with Adaptive Representations [R]
The paper introduces adaptive representations for functional gradient descent, provably ensuring convergence to global minimizers and outperforming neural networks by orders of magnitude in various settings.
From Graphs to Gradients: Physics-Inspired Structural Attribution for Cyber-Physical IoT Systems and Beyond
This paper introduces a physics-inspired framework for structural attribution in cyber-physical IoT systems, using an undirected energy-based representation to provide dependency-aware explanations without requiring a directed causal graph. Experiments on an industrial IoT testbed demonstrate higher attribution accuracy, robustness, and scalability compared to existing graph-based methods.
@NousResearch: Today we release Contrastive Neuron Attribution (CNA), a method for steering LLM behavior by identifying and ablating s…
NousResearch releases Contrastive Neuron Attribution (CNA), a method to steer LLM behavior by ablating sparse MLP circuits without training autoencoders or degrading benchmarks, validated on refusal circuits across models up to 70B parameters.