Tag
This paper introduces SAEScientist-Bench, a benchmark to evaluate AI agents' ability to autonomously conduct mechanistic interpretability research using sparse autoencoders, revealing progress but significant gaps compared to expert baselines.
HyperSAE is an open-source Python library that uses hyperbolic geometry to organize LLM learned concepts into browsable tree structures, improving on flat feature lists. It captures 99.8% of Gemma-2-2B's features and includes interactive demos.
This preregistered study tests whether holonomy (a geometric measure) concentrates on active SAE feature planes in the Gemma 2 2B language model. Contrary to the semantic-concentration prediction, active-feature planes carried less holonomy than matched mixed-feature controls, resulting in a narrow operational reversal with the underlying cause remaining open.
A researcher presents evidence that strong target text can induce a measurable latent-state shift in Gemma 3 12B IT before final output, distinct from lexical or content overlaps, and discusses implications for AI safety beyond output-only evaluation.