Tag
This paper introduces Sparse Readout Prism (SRP), a method to analyze language model predictions by decomposing the readout matrix into sparse features, enabling corpus-independent explanations of logit-lens scores and improving interpretability.
This paper investigates whether single-token sparse autoencoder features are causally necessary for model outputs, finding that causal roles vary across SAE families and are sensitive to training methodology rather than activation function or scale.