@noahdgoodman: how should we find representations in neural nets? how about the way we find everything else these days — gradient dece…

X AI KOLs Following Papers

Summary

A new paper introduces Matryoshka Attribution, a method using gradient descent to identify responsible parts of neural networks, achieving top performance on the Mechanistic Interpretability Benchmark.

how should we find representations in neural nets? how about the way we find everything else these days — gradient decent!
Original Article
View Cached Full Text

Cached at: 09/27/26, 07:08 AM

how should we find representations in neural nets? how about the way we find everything else these days — gradient decent!

Aryaman Arora (@aryaman2020): New paper! 🫡

We introduce Matryoshka Attribution, a new attribution method which uses gradient descent to find which parts of a neural network are responsible for a behaviour.

MAttr is #1 on the Mechanistic Interpretability Benchmark by a wide margin (2.9× the runner up).

Similar Articles

From Graphs to Gradients: Physics-Inspired Structural Attribution for Cyber-Physical IoT Systems and Beyond

arXiv cs.AI

This paper introduces a physics-inspired framework for structural attribution in cyber-physical IoT systems, using an undirected energy-based representation to provide dependency-aware explanations without requiring a directed causal graph. Experiments on an industrial IoT testbed demonstrate higher attribution accuracy, robustness, and scalability compared to existing graph-based methods.