solution-completeness

Tag

Cards List
#solution-completeness

Reward-Informed Sparse Autoencoders and the Solution-Completeness Confound

arXiv cs.CL ↗ · 2026-08-28 Cached

This paper introduces reward-informed sparse autoencoders (RI-SAEs) to use reinforcement learning rewards for interpretability, but finds that the separation between good and bad reasoning is largely driven by solution completeness rather than reasoning quality.

0 favorites 0 likes
← Back to home

Submit Feedback