Tag
This paper introduces SAEScientist-Bench, a benchmark to evaluate AI agents' ability to autonomously conduct mechanistic interpretability research using sparse autoencoders, revealing progress but significant gaps compared to expert baselines.
The paper presents an empirical analysis showing that AI agents in post-training excel at execution but fail to spontaneously reevaluate their strategy, which is a bottleneck for autonomous AI R&D.