mask-learning

Tag

Cards List
#mask-learning

Matryoshka attribution: Learning to attribute language model outputs to representations and weights

arXiv cs.CL ↗ · 2d ago Cached

The paper introduces Matryoshka Attribution (MAttr), a mask-learning method for attributing language model outputs to internal components, achieving top performance on the Mechanistic Interpretability Benchmark and demonstrating practical use in modifying LLM behaviors by adjusting weights.

0 favorites 0 likes
← Back to home

Submit Feedback