Reward-Informed Sparse Autoencoders and the Solution-Completeness Confound
Summary
This paper introduces reward-informed sparse autoencoders (RI-SAEs) to use reinforcement learning rewards for interpretability, but finds that the separation between good and bad reasoning is largely driven by solution completeness rather than reasoning quality.
View Cached Full Text
Cached at: 08/28/26, 09:19 AM
# Reward-Informed Sparse Autoencoders and the Solution-Completeness Confound
Source: [https://arxiv.org/html/2608.26136](https://arxiv.org/html/2608.26136)
###### Abstract
Sparse autoencoders \(SAEs\) decompose language\-model activations into sparse, interpretable features, and an appealing way to aim them at reasoning is to curate their data with a signal reinforcement learning already produces: the reward\. We build such a*reward\-informed*SAE \(RI\-SAE\): we split GRPO trajectories into high\-reward \(“good”\) and low\-reward \(“bad”\) reasoning continuations, train a standard JumpReLU SAE on their activations, and then ask what the resulting good/bad separation actually measures\. On Llama\-3\.1\-8B a sparse subset of the16,38416\{,\}384features does separate the classes \(silhouette0\.790\.79on the selected features versus0\.0050\.005for the full code\), but a control battery shows the separation is largely*solution completeness*rather than reasoning quality: a TF\-IDF text classifier already splits the classes \(AUC0\.750\.75–0\.830\.83\), and three structural cues alone \(length, a closed reasoning block, and a boxed answer\) reach AUC0\.700\.70\(99%99\\%of good versus69%69\\%of bad completions are boxed\)\. A generic SAE that never saw the reward does not separate the classes at all \(silhouette0\.010\.01, no discriminative features\), so the0\.790\.79is in\-sample fitting of this curated signal rather than structure that a reward\-blind dictionary recovers\. We therefore present the recipe and its control battery together: reward filtering is a cheap, label\-free way to reuse RL signals for interpretability, but most of what it surfaces is completion form\. Two discriminative features are still readable \(symbolic mathematics; procedural and evaluative language\), which we take as illustrative rather than as isolated reasoning\.
11footnotetext:Lead authors; equal contribution\.22footnotetext:Corresponding author\.## 1Introduction
Language models solve multi\-step reasoning problems, but the internal computations behind a correct derivation are hard to read off from activations\(Weiet al\.,[2022](https://arxiv.org/html/2608.26136#bib.bib8); Geipinget al\.,[2025](https://arxiv.org/html/2608.26136#bib.bib3)\)\. Sparse autoencoders address this by reconstructing a layer’s activations under a sparsity penalty, recovering dictionaries of features that are often monosemantic\(Brickenet al\.,[2023](https://arxiv.org/html/2608.26136#bib.bib6); Lieberumet al\.,[2024](https://arxiv.org/html/2608.26136#bib.bib9)\)\. But SAEs are typically trained on undifferentiated text: they learn whatever reconstructs the activation distribution, with no preference for the features that separate competent from incompetent reasoning\.
A signal that already encodes that distinction sits unused in RL pipelines: the reward\. Reinforcement learning with verifiable rewards, and GRPO in particular, produces many model rollouts, each tagged with a scalar reward\(Shaoet al\.,[2024](https://arxiv.org/html/2608.26136#bib.bib11)\)\. We treat this reward as a cheap form of supervision for interpretability and use it to decide which activations an SAE should look at\.
We study the simplest version of this idea, a*reward\-informed*SAE \(RI\-SAE\)\. The reward enters through the data, not the objective: we keep high\-reward GRPO continuations as “good reasoning” and low\-reward ones as “bad reasoning”, and train an otherwise standard SAE on their activations\. No new loss, no retraining of the base model; the method can be attached to an existing RL run\. The recipe works in the narrow sense that its features separate the classes; the harder question, and our main one, is what that separation actually measures\. Our contributions are: \(i\) a control battery that answers it, showing the good/bad split is largely*solution completeness*\(whether a complete, well\-formed answer was produced\) rather than reasoning quality, with a TF\-IDF text baseline at AUC0\.750\.75–0\.830\.83, structure\-only features at0\.700\.70, and a generic reward\-blind SAE that does not separate the classes at all \(so the effect is in\-sample fitting, not a generic artifact\); \(ii\) the reward\-filtering recipe itself, a cheap, label\-free way to turn an RL reward into SAE supervision; and \(iii\) two interpretable features and a Gemma\-2\-2B GSM8K fine\-tuning study that sketch where the method could go\. We see RI\-SAEs less as a finished diagnostic than as a reusable signal whose results must be read against the control battery we provide\.
## 2Method
##### Reward\-based data curation\.
We build the corpus from a public pool of GRPO continuations, each with a scalar reward in\[−1,3\]\[\-1,3\]and a<think\>…\\ldots</think\><answer\>…\\ldots</answer\>format\. The raw pool contains substantial non\-English \(largely Chinese\) text, which we remove with a non\-ASCII filter\. We then labelgood reasoning\(11\) as reward≥2\.0\\geq 2\.0andbad reasoning\(0\) as reward≤0\.5\\leq 0\.5; both require at least3030tokens and must pass a coherence check \(no excessivenn\-gram repetition or canned refusals\), so the bad class is genuinely low\-reward reasoning rather than broken text\. Our main set is balanced at1,0001\{,\}000good and1,0001\{,\}000bad \(mean reward2\.832\.83vs\.0\.060\.06\); good completions are longer than bad \(median1,2111\{,\}211vs\.453453words\)\. Two caveats are built in: the labels are a reward proxy \(“non\-reasoning” is shorthand for low\-reward reasoning\), and the reward enters only here, in data selection\.
##### Sparse autoencoder\.
For a residual\-stream activationx∈ℝdx\\in\\mathbb\{R\}^\{d\}, the SAE computesz=ReLU\(Wencx\+benc−θ\)z=\\mathrm\{ReLU\}\(W\_\{\\text\{enc\}\}x\+b\_\{\\text\{enc\}\}\-\\theta\)andx^=Wdecz\+bdec\\hat\{x\}=W\_\{\\text\{dec\}\}z\+b\_\{\\text\{dec\}\}, whereθ\\thetais a learned per\-feature threshold \(JumpReLU\) andWdecW\_\{\\text\{dec\}\}has unit\-norm rows\. We train with the standard objective
ℒ=∥x−x^∥22\+λ∥z∥0,λ=0\.1,\\mathcal\{L\}=\\lVert x\-\\hat\{x\}\\rVert\_\{2\}^\{2\}\+\\lambda\\lVert z\\rVert\_\{0\},\\qquad\\lambda=0\.1,\(1\)using Adam \(lr3×10−43\\times 10^\{\-4\}, batch1616\) for500500steps\. The reward does not appear in Eq\.[1](https://arxiv.org/html/2608.26136#S2.E1)\. Feature analyses use a wide, overcomplete16,38416\{,\}384\-feature dictionary \(roughly4×4\\timesexpansion for Llama\-3\.1\-8B\), the regime in which interpretable features have been reported\(Lieberumet al\.,[2024](https://arxiv.org/html/2608.26136#bib.bib9); Rajamanoharanet al\.,[2024](https://arxiv.org/html/2608.26136#bib.bib10)\); activations are taken from layer2222\. A reward\-weighted variant of Eq\.[1](https://arxiv.org/html/2608.26136#S2.E1)is a straightforward extension we have not implemented\.
## 3Results
### 3\.1What does the good/bad separation measure?
Reward\-filtered features do separate the two classes, but only after selection and only in a way a plain text classifier matches\. In the full16,38416\{,\}384\-dimensional code the classes do not separate \(silhouette0\.0050\.005, Davies–Bouldin6\.946\.94\)\. Keeping the features that individually discriminate them \(per\-feature silhouette\>0\.1\>0\.1\) and embedding that subspace with UMAP yields clean clusters \(silhouette0\.790\.79, Davies–Bouldin0\.280\.28; Figure[1](https://arxiv.org/html/2608.26136#S3.F1)a\)\. That jump is produced by the selection step \(we pick discriminative features and then report separation on them\), so it shows a*sparse subset*of features carries the distinction, not that the SAE separates reasoning globally\.
##### The separation is mostly solution completeness\.
That distinction is mostly structural\. High\- and low\-reward completions differ in vocabulary, formatting, and length, not only in reasoning, and a control battery \(Table[1](https://arxiv.org/html/2608.26136#S3.T1)\) quantifies it\. A TF\-IDFnn\-gram classifier on the raw text reaches AUC0\.830\.83cross\-validated and0\.750\.75held out \(Figure[1](https://arxiv.org/html/2608.26136#S3.F1)b\); removing digits or answer\-formatting tokens barely moves it \(it stays near0\.830\.83–0\.850\.85\); and three structural features alone \(length, whether the reasoning block is closed, and whether a`\\boxed\{\}`answer appears\) reach AUC0\.700\.70\. The cues are concrete: the top good\-classnn\-grams areboxed,final answer, andtherefore;99%99\\%of good versus69%69\\%of bad completions contain a boxed answer; and100%100\\%versus83%83\\%close the reasoning block\. Much of “good versus bad reasoning” is thus*solution completeness*, which bounds how much of any reward\-filtered SAE result can be read as reasoning rather than completion form\.
##### A generic SAE does not recover the separation\.
Is the discriminative subspace produced by reward filtering, or by the SAE alone? We ran the same pipeline on a generic, pretrained Llama Scope SAE for the matched Llama\-3\.1\-8B residual stream\(Heet al\.,[2024](https://arxiv.org/html/2608.26136#bib.bib13)\), one trained on ordinary text that never saw the reward\. On the same1,000/1,0001\{,\}000/1\{,\}000set it does not separate the classes: full\-code silhouette0\.0120\.012\(cf\.0\.0050\.005\),*no*feature exceeds our per\-feature selection threshold \(max single\-feature silhouette0\.0590\.059\), total\-activation AUC0\.610\.61, and selecting and UMAP\-embedding its5050most class\-different features still yields silhouette0\.020\.02\. The0\.790\.79is therefore neither intrinsic to these activations nor a generic UMAP artifact; it requires an SAE trained on the curated set itself\. Because that training and the feature selection are in\-sample, and the distinction is largely completeness, we read the0\.790\.79as in\-sample fitting of a completeness\-dominated signal rather than reward\-independent reasoning structure\. This control isolates the SAE from the data but not reward filtering from in\-domain training; a same\-recipe SAE on*unfiltered*in\-domain data would isolate it, and is left to future work\.
Table 1:Anatomy of the good/bad separation \(Llama\-3\.1\-8B; cross\-validated AUC of a logistic classifier on the named features\)\. Stripping lexical content barely changes the separation, and three structural cues alone \(length, a closed block, a boxed answer\) recover most of it\.\(a\)
\(b\)
Figure 1:Separating good\- from bad\-reasoning trajectories \(Llama\-3\.1\-8B, layer 22\)\. \(a\) UMAP of the SAE features that individually discriminate the classes; the selected subspace clusters cleanly \(silhouette0\.790\.79\), while the full code gives0\.0050\.005\. \(b\) A TF\-IDFnn\-gram classifier on the raw text reaches AUC0\.830\.83\(CV\) /0\.750\.75\(held out\), showing a strong lexical confound\.
### 3\.2Two discriminative features are interpretable
Ranking features by the difference in mean activation between the classes gives consistent but modest differences \(≈±0\.1\\approx\\pm 0\.1\), spread across many features\. We interpret two highly discriminative ones by their top\-activating tokens, computed by running the encoder over the tokens of5050examples and averaging each feature’s strongest tokens \(Figure[2](https://arxiv.org/html/2608.26136#S3.F2); in the byte\-level BPE tokenizer a leading space is written as a special prefix symbol\)\. Feature1596815968fires on symbolic\-mathematics tokens \(matrix,rows,theta,\(x,digits\), a clean example of a structured\-mathematics feature\. Feature42054205is associated with procedural and evaluative language \(think,values,formula,remember,must\), though it also picks up common function words, so its interpretation is suggestive rather than definitive\. Starting only from a reward signal, we arrive at named features a practitioner can inspect; we do not claim they are causally responsible for reasoning\.
\(a\)
\(b\)
Figure 2:Top\-activating tokens for two discriminative features \(Llama\-3\.1\-8B, layer 22, averaged over5050examples\)\. \(a\) Feature1596815968: symbolic mathematics\. \(b\) Feature42054205: procedural/evaluative language \(with some common function words\)\.
### 3\.3Toward monitoring fine\-tuning
The use we ultimately have in mind is monitoring reasoning features as a model trains\. As a first step we fine\-tuned Gemma\-2\-2B on GSM8K, saved five checkpoints \(steps600600–30003000\), and evaluated each \(Appendix[A](https://arxiv.org/html/2608.26136#A1)\)\. Exact match rises from0\.5%0\.5\\%to∼\\sim5% and MAE on numeric answers falls from∼\\sim4950 to∼\\sim100, with most of the gain between steps12001200and18001800\. This characterizes the model, not its features: we did not run an SAE across checkpoints\. It does, however, give a ready setup: training an RI\-SAE on these checkpoints to test whether reasoning features sharpen in that same window is the immediate next step\.
## 4Limitations
The labels are a reward proxy, not reasoning versus its absence\. The solution\-completeness confound \(Table[1](https://arxiv.org/html/2608.26136#S3.T1)\) is the paper’s main result rather than a caveat, and it bounds the rest: on these numbers alone we cannot read the SAE separation as reasoning structure\. The headline silhouette is computed on features selected for being discriminative and so reflects that selection, not global separation \(0\.0050\.005on the full code\)\. Evaluation sets are small and results come from single runs without seeds or error bars, so we present them as descriptive\. Our generic\-SAE control shows the SAE alone does not drive the separation, but it does not isolate reward filtering from in\-domain training; the cleanest remaining test \(a same\-recipe SAE on*unfiltered*in\-domain data\), along with causal interventions and a concept\-alignment metric\(Felet al\.,[2025](https://arxiv.org/html/2608.26136#bib.bib14)\), is future work\. The reward\-weighted objective is so far only a proposal, and the feature analyses \(Llama\-3\.1\-8B\) and fine\-tuning study \(Gemma\-2\-2B\) use different backbones\.
## 5Conclusion
Reward signals are an underused resource for interpretability, but they are not a free one\. Filtering RL trajectories by reward and training a standard SAE does surface a sparse set of features that separate good from bad reasoning on Llama\-3\.1\-8B, yet our control battery shows that most of that separation is solution completeness, recoverable from length, a closed reasoning block, and a boxed answer \(AUC0\.700\.70\) and already matched by a plain text classifier\. The honest reading is that reward filtering is a cheap, reusable way to point an SAE at reasoning\-adjacent data, but its good/bad signal must be read against a completeness baseline before any feature is called a reasoning feature; the battery we report is the tool for doing so\. The natural next steps sharpen the test rather than the claim: a same\-recipe SAE on unfiltered in\-domain data to isolate filtering from in\-domain training \(a generic SAE already fails to separate the classes\); a reward\-weighted objective; and SAE\-based monitoring across fine\-tuning checkpoints\.
#### Acknowledgments
We thank Kevin Zhu and Ryan Lagasse for their guidance and feedback throughout this project\.
## References
- T\. Bricken, A\. Templeton, J\. Batson, B\. Chen, A\. Jermyn, T\. Conerly, N\. Turner, C\. Anil, C\. Denison, A\. Askell, R\. Lasenby, Y\. Wu, S\. Kravec, N\. Schiefer, T\. Maxwell, N\. Joseph, Z\. Hatfield\-Dodds, A\. Tamkin, K\. Nguyen, B\. McLean, J\. E\. Burke, T\. Hume, S\. Carter, T\. Henighan, and C\. Olah \(2023\)Towards monosemanticity: decomposing language models with dictionary learning\.Transformer Circuits Thread\.Note:https://transformer\-circuits\.pub/2023/monosemantic\-features/index\.htmlCited by:[§1](https://arxiv.org/html/2608.26136#S1.p1.1)\.
- T\. Fel, E\. S\. Lubana, J\. S\. Prince, M\. Kowal, V\. Boutin, I\. Papadimitriou, B\. Wang, M\. Wattenberg, D\. Ba, and T\. Konkle \(2025\)Archetypal SAE: adaptive and stable dictionary learning for concept extraction in large vision models\.External Links:2502\.12892,[Link](https://arxiv.org/abs/2502.12892)Cited by:[§4](https://arxiv.org/html/2608.26136#S4.p1.1)\.
- J\. Geiping, S\. McLeish, N\. Jain, J\. Kirchenbauer, S\. Singh, B\. R\. Bartoldson, B\. Kailkhura, A\. Bhatele, and T\. Goldstein \(2025\)Scaling up test\-time compute with latent reasoning: a recurrent depth approach\.External Links:2502\.05171,[Link](https://arxiv.org/abs/2502.05171)Cited by:[§1](https://arxiv.org/html/2608.26136#S1.p1.1)\.
- Z\. He, W\. Shu, X\. Ge, L\. Chen, J\. Wang, Y\. Zhou, F\. Liu, Q\. Guo, X\. Huang, Z\. Wu, Y\. Jiang, and X\. Qiu \(2024\)Llama scope: extracting millions of features from llama\-3\.1\-8b with sparse autoencoders\.External Links:2410\.20526,[Link](https://arxiv.org/abs/2410.20526)Cited by:[§3\.1](https://arxiv.org/html/2608.26136#S3.SS1.SSS0.Px2.p1.9)\.
- T\. Lieberum, S\. Rajamanoharan, A\. Conmy, L\. Smith, N\. Sonnerat, V\. Varma, J\. Kramár, A\. Dragan, R\. Shah, and N\. Nanda \(2024\)Gemma scope: open sparse autoencoders everywhere all at once on gemma 2\.External Links:2408\.05147,[Link](https://arxiv.org/abs/2408.05147)Cited by:[§1](https://arxiv.org/html/2608.26136#S1.p1.1),[§2](https://arxiv.org/html/2608.26136#S2.SS0.SSS0.Px2.p1.11)\.
- Jumping ahead: improving reconstruction fidelity with jumprelu sparse autoencoders\.External Links:2407\.14435,[Link](https://arxiv.org/abs/2407.14435)Cited by:[§2](https://arxiv.org/html/2608.26136#S2.SS0.SSS0.Px2.p1.11)\.
- Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. K\. Li, Y\. Wu, and D\. Guo \(2024\)DeepSeekMath: pushing the limits of mathematical reasoning in open language models\.External Links:2402\.03300,[Link](https://arxiv.org/abs/2402.03300)Cited by:[§1](https://arxiv.org/html/2608.26136#S1.p2.1)\.
- J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, B\. Ichter, F\. Xia, E\. Chi, Q\. Le, and D\. Zhou \(2022\)Chain\-of\-thought prompting elicits reasoning in large language models\.External Links:2201\.11903,[Link](https://arxiv.org/abs/2201.11903)Cited by:[§1](https://arxiv.org/html/2608.26136#S1.p1.1)\.
## Appendix AFine\-tuning details and curves
We fine\-tuned Gemma\-2\-2B on GSM8K with LoRA and evaluated five checkpoints on the GSM8K test set, scoring exact match \(EM\) on the final answer and mean absolute error \(MAE\) on the numeric answer\. EM is low in absolute terms because of a strict last\-line string match; MAE shows the numeric answers nonetheless converge toward the correct values\.
\(a\)
\(b\)
\(c\)
Figure 3:Gemma\-2\-2B GSM8K fine\-tuning\. \(a\) MAE on numeric answers\. \(b\) Exact match\. \(c\) Training/validation loss\. These are model\-performance curves; SAE\-based feature tracking across the checkpoints is left to future work\.
## Appendix BSAE configuration
JumpReLU SAE \(learned per\-feature threshold initialized at0\.0010\.001, unit\-norm decoder rows\); objective MSE\+0\.1∥z∥0\+\\,0\.1\\lVert z\\rVert\_\{0\}; Adam, lr3×10−43\\times 10^\{\-4\}, batch1616,500500steps; dictionary width16,38416\{,\}384; Llama\-3\.1\-8B activations at layer 22 for the reported analyses\. Data curation as in Section[2](https://arxiv.org/html/2608.26136#S2): reward thresholds2\.02\.0/0\.50\.5, minimum3030tokens, trigram\-repetition and refusal\-phrase coherence filters\.Similar Articles
Rational Sparse Autoencoder
Introduces Rational Sparse Autoencoder (RSAE), which replaces fixed encoder activations with trainable rational functions, improving reconstruction and sparsity trade-offs on residual-stream activations of open-weight language models across multiple baseline families.
Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders
This paper investigates preference instability in reward models for LLMs, where subtle input variations cause contradictory preference assignments. The authors propose two SAE-based mitigation strategies—SAE Feature Steering and SAE Residual Correction—to reduce incorrect preference assignments without retraining.
Thinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse Autoencoders
This paper uses Top-K sparse autoencoders to analyze DeepSeek-R1-Distill-Qwen-7B's internal reasoning, contrasting Thinking (CoT) and NoThinking modes, and finds distinct feature activation patterns with causal intervention experiments.
Discovering Millions of Interpretable Features with Sparse Autoencoders
This paper introduces Qwen3-Instruct SAE, a suite of sparse autoencoders trained on Qwen3 instruction-tuned models, enabling the discovery of millions of interpretable features and demonstrating refusal steering capabilities.
Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models
SAE-FT introduces a novel fine-tuning method for CLIP models that uses sparse autoencoder constraints to regularize visual representations, improving robustness against distribution shifts while maintaining performance and enabling interpretability.