Tag
This paper introduces reward-informed sparse autoencoders (RI-SAEs) to use reinforcement learning rewards for interpretability, but finds that the separation between good and bad reasoning is largely driven by solution completeness rather than reasoning quality.