Thinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse Autoencoders
Summary
This paper uses Top-K sparse autoencoders to analyze DeepSeek-R1-Distill-Qwen-7B's internal reasoning, contrasting Thinking (CoT) and NoThinking modes, and finds distinct feature activation patterns with causal intervention experiments.
View Cached Full Text
Cached at: 08/11/26, 08:07 AM
# Thinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse Autoencoders
Source: [https://arxiv.org/html/2608.08168](https://arxiv.org/html/2608.08168)
Bo Cheng School of Artificial Intelligence, Jilin University chengbo9691@gmail\.com &Qiaolin Lu The Hong Kong Polytechnic University qiaolin\.lu@connect\.polyu\.hk &Yi Chang School of Artificial Intelligence, Jilin University Engineering Research Center of Knowledge\-Driven Human\-Machine Intelligence, MOE, China International Center of Future Science, Jilin University yichang@jlu\.edu\.cn &Yuan Wu School of Artificial Intelligence, Jilin University yuanwu@jlu\.edu\.cn
###### Abstract
While Large Language Models \(LLMs\) employing Chain\-of\-Thought \(CoT\) exhibit superior reasoning capabilities, the neural mechanisms distinguishing this explicitThinkingmode from direct answer generation \(NoThinkingmode\) remain poorly understood\. To deconstruct this cognitive process, we apply Top\-K Sparse Autoencoders \(SAEs\) to the intermediate representations of DeepSeek\-R1\-Distill\-Qwen\-7B and examine the model’s divergent behaviors across math\-solving tasks of three distinct difficulty levels\. Observationally, we identify a clear distinction in how the model functions under two reasoning modes:Thinkingmode relies on sparse and high\-intensity feature activations driving verbal deduction independent of problem complexity, whereasNoThinkingmode exhibits an adaptive and diffuse pattern prioritizing symbolic manipulation\. Causally, suppressing the three most active sparse features by Total Activation Volume reveals three principles: \(i\) reasoning and syntactic structure are tightly coupled, as interventions consistently degradeLaTeXand boxed\-solution formatting; \(ii\)Thinkingresponds to disruption with compensatory over\-generation marked by increased metacognitive cues and repetitive, low\-information continuations; and \(iii\) coherent CoT behavior depends on a fragile coordination among specialized features, yielding distinct failure modes under perturbation but a consistently impaired output structure\.
## 1Introduction
Recent advancements in Large Language Models \(LLMs\) have demonstrated that explicitly eliciting a Chain\-of\-Thought \(CoT\) significantly enhances performance on complex reasoning tasks\. Models such as DeepSeek\-R1\(Guoet al\.,[2025](https://arxiv.org/html/2608.08168#bib.bib11)\)exemplify this paradigm by generating extended thinking process involving exploration, backtracking, and self\-correction before producing a final answer\. While the behavioral benefits of CoT are well\-documented, the underlying neural mechanisms remain opaque\(Weiet al\.,[2022](https://arxiv.org/html/2608.08168#bib.bib27); Turpinet al\.,[2023](https://arxiv.org/html/2608.08168#bib.bib28); Lightmanet al\.,[2023](https://arxiv.org/html/2608.08168#bib.bib29)\)\. Specifically, the structural distinctions between the internal states governing explicit reasoning \(Thinkingmode\) and direct answer generation \(NoThinkingmode\) have yet to be fully elucidated\(Chenet al\.,[2025](https://arxiv.org/html/2608.08168#bib.bib24);[Theodoruset al\.,](https://arxiv.org/html/2608.08168#bib.bib26); Nandaet al\.,[2023](https://arxiv.org/html/2608.08168#bib.bib31); Liet al\.,[2022](https://arxiv.org/html/2608.08168#bib.bib32); Zouet al\.,[2023](https://arxiv.org/html/2608.08168#bib.bib33)\)\. A critical and unresolved issue is whether this thinking process constitutes a distinct computational regime or functions merely as a prolonged extension of standard sequence generation\.
To deconstruct this black box, Sparse Autoencoders \(SAEs\) have emerged as a powerful microscopic toolRajamanoharanet al\.\([2024](https://arxiv.org/html/2608.08168#bib.bib30)\); Cunninghamet al\.\([2023](https://arxiv.org/html/2608.08168#bib.bib18)\)\. Pioneering works have successfully addressed the polysemanticity of dense activations by decomposing them into interpretable and monosemantic features\(Brickenet al\.,[2023](https://arxiv.org/html/2608.08168#bib.bib8); Cunninghamet al\.,[2023](https://arxiv.org/html/2608.08168#bib.bib18); Liet al\.,[2025](https://arxiv.org/html/2608.08168#bib.bib25); Menget al\.,[2022](https://arxiv.org/html/2608.08168#bib.bib16)\)\. With recent architectural innovations such as the Top\-K activation mechanism\(Gaoet al\.,[2024](https://arxiv.org/html/2608.08168#bib.bib20)\)and JumpReLU\(Lieberumet al\.,[2024](https://arxiv.org/html/2608.08168#bib.bib22)\), SAEs have been scaled to analyze massive open\-weights models such as Gemma 2\. While existing research has made significant strides in dictionary learning, safety auditing\(Gallifantet al\.,[2025](https://arxiv.org/html/2608.08168#bib.bib19)\), and model steering\(Aradet al\.,[2025](https://arxiv.org/html/2608.08168#bib.bib17)\), these studies focus predominantly on identifying static semantic concepts or manipulating output logits\. Few studies have utilized SAEs to dynamically decode the temporal evolution of reasoning processes or causally disentangle the neural circuits that govern the initiation and regulation of CoT\.
Bridging this gap, we apply Top\-K SAEs to analyze the intermediate representations of DeepSeek\-R1\-Distill\-Qwen\-7B\. Unlike prior work that analyzes features in isolation, we establish a comparative framework to contrast the feature dynamics betweenThinkingandNoThinkingmodes processing identical mathematical problems\.
Our investigation into the latent feature space reveals a fundamental mechanistic divergence between the two modes\. We observe thatThinkingmode operates through a sparse yet high\-intensity activation regime, where a dedicated subset of features drives verbal deduction\. Crucially, this reasoning pathway remains stable and invariant to problem complexity\. In contrast,NoThinkingmode exhibits a diffuse and adaptive pattern, recruiting a broader and variable coalition of features to prioritize symbolic manipulation\. This strategy bypasses explicit reasoning in favor of difficulty\-dependent pattern matching and syntactic retrieval\.
Causally, to test whether the discovered sparse features are functionally necessary forThinking, we perform targeted suppression on the top\-3 features ranked by Total Activation Volume \(TAV\)\. These interventions reveal three governing principles\. \(i\) Coupling between reasoning and syntactic structure\. Suppressing high\-impact features consistently degrades the model’s ability to produce formal mathematical outputs, suggesting that logical computation and structural realization are supported by overlapping representations rather than separable reasoning and formatting modules\. \(ii\) Compensatory sequence extension under disruption\. When the core feature 28634 is suppressed, the model tends to avoid termination and instead expands the generation with increased metacognitive cueing, producing longer but less informative and more repetitive continuations\. \(iii\) Fragile coordination under feature suppression\. Suppressing different components induces opposite\-signed shifts in monitoring\-related signals, yet structural degradation remains consistent, implying thatThinkingdepends on a finely balanced coordination among a small set of specialized, high\-intensity features with limited redundancy and is prone to distinct failure modes such as uncontrolled verbosity\.
## 2Related Work
##### SAE Architecture and Scaling\.
Sparse Autoencoders address the polysemanticity of LLM activations by decomposing dense internal states into sparse, interpretable feature combinations\. While early work byCunninghamet al\.\([2023](https://arxiv.org/html/2608.08168#bib.bib18)\)successfully demonstrated this capability in small language models, scaling was historically hindered by training instability and the prevalence of dead latents\. To overcome these challenges, recent research has shifted fromL1L\_\{1\}regularization to direct sparsity enforcement\.Gaoet al\.\([2024](https://arxiv.org/html/2608.08168#bib.bib20)\)introduced the Top\-kkactivation mechanism, which simplifies hyperparameter tuning and significantly reduces dead latents by retaining only thekkhighest\-magnitude features\. Building on this,Lieberumet al\.\([2024](https://arxiv.org/html/2608.08168#bib.bib22)\)proposed JumpReLU to dynamically threshold low\-magnitude noise\. These innovations have enabled the training of SAEs on state\-of\-the\-art open\-weight models, such as the massive Gemma Scope suite spanning up to 27B parameters, establishing robust scaling laws for reconstruction fidelity\.
##### Semantic Validation and Interpretability\.
With robust architectures established, focus has shifted to validating the semantic alignment of learned features\. A primary method involves automated interpretability, where strong LLMs generate natural language explanations for features based on maximally activating contextsGallifantet al\.\([2025](https://arxiv.org/html/2608.08168#bib.bib19)\); Gaoet al\.\([2024](https://arxiv.org/html/2608.08168#bib.bib20)\)\. Beyond general semantics,Jinget al\.\([2025](https://arxiv.org/html/2608.08168#bib.bib21)\)introduced the LinguaLens framework to rigorously analyze linguistic mechanisms\. By leveraging "minimal pair" counterfactuals, they demonstrated that SAE features align with specific theoretical categories across morphology and syntax, confirming that LLMs encode precise linguistic attributes in distinguishable sparse directions\.
##### Mechanistic Analysis and Downstream Utility\.
Recent studies have further utilized SAEs to probe model behavior and enhance downstream tasks through causal intervention\. In the context of model steering,Aradet al\.\([2025](https://arxiv.org/html/2608.08168#bib.bib17)\)distinguished between input features \(pattern detection\) and output features \(generation influence\), showing that steering is most effective when targeting features with high causal scores on output logits\. Regarding application,Gallifantet al\.\([2025](https://arxiv.org/html/2608.08168#bib.bib19)\)found that binarized SAE features outperform dense states in safety\-critical tasks like toxicity detection due to better transferability\. Similarly,Parket al\.\([2025](https://arxiv.org/html/2608.08168#bib.bib23)\)applied SAEs to discretize dense retriever embeddings, enabling Concept\-Level Sparse Retrieval \(CL\-SR\) that combines semantic expressiveness with the efficiency of sparse representations\.
## 3Preliminaries
In this section, we provide the formal background for sparse autoencoders, with a specific focus on the Top\-K sparse autoencoders\.
### 3\.1Problem Setup
Let𝐱∈ℝd\\mathbf\{x\}\\in\\mathbb\{R\}^\{d\}denote the input vector originating from a specific layer of a pre\-trained language model\. While typically dense and polysemantic, our aim is to decompose𝐱\\mathbf\{x\}into a sparse linear combination of interpretable feature directions from an overcomplete dictionary\. Formally, we seek to learn a dictionary matrix𝐖dec∈ℝd×m\\mathbf\{W\}\_\{\\text\{dec\}\}\\in\\mathbb\{R\}^\{d\\times m\}and latent activations𝐳∈ℝm\\mathbf\{z\}\\in\\mathbb\{R\}^\{m\}withm≫dm\\gg dsuch that the input vector is approximately𝐱≈𝐖dec𝐳\+𝐛dec\\mathbf\{x\}\\approx\\mathbf\{W\}\_\{\\text\{dec\}\}\\mathbf\{z\}\+\\mathbf\{b\}\_\{\\text\{dec\}\}\.
### 3\.2Sparse Autoencoders
Sparse Autoencoders implement the decomposition using an encoder\-decoder architecture\. The encoder maps the input vector𝐱\\mathbf\{x\}to latent activations𝐳\\mathbf\{z\}via an affine transformation followed by a non\-linear activation functionσ\(⋅\)\\sigma\(\\cdot\):
𝐳=σ\(𝐖enc𝐱\+𝐛enc\)\\mathbf\{z\}=\\sigma\(\\mathbf\{W\}\_\{\\text\{enc\}\}\\mathbf\{x\}\+\\mathbf\{b\}\_\{\\text\{enc\}\}\)\(1\)where𝐖enc∈ℝm×d\\mathbf\{W\}\_\{\\text\{enc\}\}\\in\\mathbb\{R\}^\{m\\times d\}and𝐛enc∈ℝm\\mathbf\{b\}\_\{\\text\{enc\}\}\\in\\mathbb\{R\}^\{m\}denote the encoder parameters\. ReLU is typically employed asσ\(⋅\)\\sigma\(\\cdot\)to enforce non\-negativity\(Brickenet al\.,[2023](https://arxiv.org/html/2608.08168#bib.bib8)\)\. Then, the decoder reconstructs the input vector using the learned feature directions:
𝐱^=𝐖dec𝐳\+𝐛dec\\hat\{\\mathbf\{x\}\}=\\mathbf\{W\}\_\{\\text\{dec\}\}\\mathbf\{z\}\+\\mathbf\{b\}\_\{\\text\{dec\}\}\(2\)where𝐖dec∈ℝd×m\\mathbf\{W\}\_\{\\text\{dec\}\}\\in\\mathbb\{R\}^\{d\\times m\}and𝐛dec∈ℝd\\mathbf\{b\}\_\{\\text\{dec\}\}\\in\\mathbb\{R\}^\{d\}represent the decoder parameters\. The model is trained to minimize the following composite objective:
ℒ=\|𝐱−𝐱^\|22\+λ\|𝐳\|1\\mathcal\{L\}=\|\\mathbf\{x\}\-\\hat\{\\mathbf\{x\}\}\|\_\{2\}^\{2\}\+\\lambda\|\\mathbf\{z\}\|\_\{1\}\(3\)where\|𝐱−𝐱^\|22\|\\mathbf\{x\}\-\\hat\{\\mathbf\{x\}\}\|\_\{2\}^\{2\}quantifies the reconstruction error, while\|𝐳\|1\|\\mathbf\{z\}\|\_\{1\}imposes anL1L\_\{1\}penalty weighted by the hyperparameterλ\\lambdato enforce sparsity\. However,L1L\_\{1\}regularization induces a shrinkage bias where the model suppresses the magnitudes of active feature to minimize the total loss, thereby compromising the fidelity of recovered semantic concepts\.
### 3\.3Top\-K Sparse Autoencoders
To mitigate the shrinkage bias caused byL1L\_\{1\}regularization, we adoptkk\-sparse autoencoders in this work\(Makhzani and Frey,[2013](https://arxiv.org/html/2608.08168#bib.bib9)\)\. This architecture enforces sparsity directly through the activation mechanism, rather than a soft penalty in the loss\.
Specifically, the model imposes a hard constraint by retaining only thekkmost significant latents\. Given the pre\-activation𝐡=𝐖enc𝐱\+𝐛enc\\mathbf\{h\}=\\mathbf\{W\}\_\{\\text\{enc\}\}\\mathbf\{x\}\+\\mathbf\{b\}\_\{\\text\{enc\}\}, the latent activations are computed via a TopK operator:
𝐳=TopK\(𝐡\)\\mathbf\{z\}=\\text\{TopK\}\(\\mathbf\{h\}\)\(4\)where theii\-th element𝐳i\\mathbf\{z\}\_\{i\}retains the value of𝐡i\\mathbf\{h\}\_\{i\}if and only if\|𝐡i\|\|\\mathbf\{h\}\_\{i\}\|ranks among the topkkmagnitudes in𝐡\\mathbf\{h\}\. Otherwise, it is set to zero\. In addition, ReLU can be applied implicitly or explicitly to ensure positive feature activations\. The decoding process remains identical to that of the standard SAEs\. Since sparsity is strictly enforced bykk, the loss function simplifies to the reconstruction loss:
ℒ=\|𝐱−𝐱^\|22\\mathcal\{L\}=\|\\mathbf\{x\}\-\\hat\{\\mathbf\{x\}\}\|\_\{2\}^\{2\}\(5\)
This architecture effectively decouples sparsity from activation magnitude, allowing the model to learn precise feature strengths without the downward pressure exerted by regularization penalties\.
## 4Experiments
To investigate the behavior of modern reasoning models when solving mathematical problems across three distinct difficulty levels under two inference modes, we first detail our experimental setup, and then present a comprehensive analysis of the observed patterns\.
### 4\.1Experiment Setup
#### 4\.1\.1Model
We employ DeepSeek\-R1\-Distill\-Qwen\-7B\(Guoet al\.,[2025](https://arxiv.org/html/2608.08168#bib.bib11)\)as our primary subject of analysis\. Initialized with Qwen2\.5\-Math\-7B and fine\-tuned on the outputs generated by DeepSeek\-R1, DeepSeek\-R1\-Distill\-Qwen\-7B preserves strong reasoning capabilities while offering a computationally tractable scale for interpretability research\.
#### 4\.1\.2Dataset
To capture the sparse features underlying complex reasoning mechanisms, we employ DeepMath\-103K\(Heet al\.,[2025](https://arxiv.org/html/2608.08168#bib.bib12)\)as the training corpus for the sparse autoencoders\. The large\-scale dataset comprises 103,000 mathematical problems, primarily sourced from Math StackExchange111https://math\.stackexchange\.com, and is specifically curated to advance reasoning capabilities\.
#### 4\.1\.3Training Details
##### Inference Modes\.
Modern reasoning architectures, such as R1 and R1\-Distill\-Qwen, typically segregate internal cognition from final output using specific delimiters \(e\.g\.,<\|beginning\_of\_thinking\|\>and<\|end\_of\_thinking\|\>\)\. Based on this structure, we adapt two distinct inference modes as described in\(Maet al\.,[2025](https://arxiv.org/html/2608.08168#bib.bib10)\):
- •Thinking: the mode follows the standard generation trajectory, preserving the full chain of thought within the thinking box before producing the final solution and answer\.
- •NoThinking: the mode bypasses the explicit reasoning phase\. By constraining the decoding process to keep the thinking box empty, the model is guided to generate only the final solution and answer directly\.
##### Activation Extraction\.
Training Sparse Autoencoders requires dense activations derived from a target language model\. We employ the DeepSeek\-R1\-Distill\-Qwen\-7b model on the DeepMath\-103K corpus and extract activations from the residual stream of the 13th layer, chosen as a representative intermediate layer222More details are further provided in Appendix[A\.1](https://arxiv.org/html/2608.08168#A1.SS1)\. The collection process operates in two distinct modes: \(1\)Thinkingmode, which processes sequences of length 1,024 to capture CoT reasoning, yielding a total of 108M tokens; and \(2\)NoThinkingmode, which excludes reasoning traces, resulting in a total of 65M tokens\.
##### Hyperparameters\.
Our training methodology and hyperparameter settings follow primarily established protocols\(Gaoet al\.,[2024](https://arxiv.org/html/2608.08168#bib.bib20); Lieberumet al\.,[2024](https://arxiv.org/html/2608.08168#bib.bib22);[Wuet al\.,](https://arxiv.org/html/2608.08168#bib.bib13)\)\. Specifically, we train a Top\-K Sparse Autoencoder for each mode withC=216C=2^\{16\}feature vectors333More details can be found in Appendix[A\.1](https://arxiv.org/html/2608.08168#A1.SS1)\. We optimize the model using Adam\(Adam and others,[2014](https://arxiv.org/html/2608.08168#bib.bib14)\)with a constant learning rate of1×10−31\\times 10^\{\-3\},β1=0\.9\\beta\_\{1\}=0\.9,β2=0\.999\\beta\_\{2\}=0\.999, andϵ=6\.25×10−10\\epsilon=6\.25\\times 10^\{\-10\}\. To ensure training stability, we implement a dynamic sparsity schedule where the Top\-K constraint anneals fromK=200K=200toK=20K=20during the first 50% of the initial epoch\. Training configurations differ slightly by mode:Thinkingmode is trained with a batch size of 1,024 for 4 epochs, whereasNoThinkingmode uses a batch size of 128 for 3 epochs\.
#### 4\.1\.4Benchmarks
To systematically verify the behavioral differences and internal feature dynamics of the reasoning model under varying levels of problem complexity, we categorize our evaluation benchmarks into three difficulty levels:Easy\(AMC23\),Medium\(AIME 24 & AIME 25 \), andHard\(OlympiadBench\)\. More details of the benchmarks can be found in Appendix[A\.2](https://arxiv.org/html/2608.08168#A1.SS2)\.
#### 4\.1\.5Comparison of Two Reasoning Modes
To reveal the fundamental differences in the model’s internal reasoning mechanisms, we compare feature activation patterns across different difficulty levels\.
##### Feature Selection\.
Rather than selecting features randomly, we target the most dominant components of the model’s latent representation\. We quantify feature importance by computingTotal Activation Volume\(TAV\) for each featureii, defined as the sum of activation magnitudes across a validation corpus comprisingNNtokens:TAVi=∑n=1N𝐳i\(n\)\\text\{TAV\}\_\{i\}=\\sum\_\{n=1\}^\{N\}\\mathbf\{z\}\_\{i\}^\{\(n\)\}, where𝐳i\(n\)\\mathbf\{z\}\_\{i\}^\{\(n\)\}denotes the activation of theii\-th feature for thenn\-th token\. Leveraging the metric, we employ mode\-specific SAE to compute the average activation strength across three difficulty levels, identifying the global top\-20 features with the highest aggregate activity separately for each mode\.
#### 4\.1\.6Causal Intervention
To move beyond correlational analysis and establish the functional necessity of specific sparse features, we design a causal intervention framework targeting the model’s internal reasoning process\.
Figure 1:The mean, maximum, and standard deviation \(variability\) of activations for the top\-20 feature vectors across three difficulty levels\. Features are selected based on the TAV metric\.Figure 2:Heatmaps showing the normalized activations of the top\-20 features across difficulty levels forThinkingmode andNoThinkingmode\. We apply row\-wise normalization, where each feature’s activation is scaled by its maximum value across difficulty levels\.##### Feature Selection\.
Guided by the TAV metric and theThinking\-specific SAE, we select the top\-3 features with the highest aggregate activity on the easy task, specifically Feature 4416, Feature 8893, and Feature 28634, as our primary subjects for intervention\. This selection criterion ensures that our analysis targets the neural units that are most salient during the model’s reasoning phase\.
##### Intervention Protocol\.
We implement a dynamic suppression hook at the 13th layer, which allows for precise manipulation of feature activations during inference\. To isolate the impact on reasoning logic, the intervention is applied exclusively when the model is generating tokens within the thinking block\. For a target feature indexiiand a suppression strength coefficientα∈\[0,1\]\\alpha\\in\[0,1\], the modified latent activation𝐳^i\(t\)\\hat\{\\mathbf\{z\}\}\_\{i\}^\{\(t\)\}at time stepttis computed as:
𝐳^i\(t\)=\(1−α\)𝐳i\(t\)\\hat\{\\mathbf\{z\}\}\_\{i\}^\{\(t\)\}=\(1\-\\alpha\)\\mathbf\{z\}\_\{i\}^\{\(t\)\}\(6\)where𝐳i\(t\)\\mathbf\{z\}\_\{i\}^\{\(t\)\}represents the original activation value computed by the SAE encoder\. We systematically explore four suppression strengths,α∈\{0\.1,0\.3,0\.5,1\.0\}\\alpha\\in\\\{0\.1,0\.3,0\.5,1\.0\\\}, ranging from mild attenuation to complete ablation\. This graded intervention strategy enables us to characterize non\-linear behavioral shifts and identify sensitivity thresholds within the reasoning process\.
##### Evaluation Metrics\.
To quantify the effects of intervention while accounting for the high variance in generated sequence lengths, we define a comprehensive set of density\-based metrics capturing linguistic patterns, mathematical formalization and output characteristics, as detailed in Appendix[A\.4](https://arxiv.org/html/2608.08168#A1.SS4)\. All density metrics are normalized per 1,000 tokens to enable fair and robust comparison between baseline and intervened trajectories\.
### 4\.2Comparison ofThinkingandNoThinking
#### 4\.2\.1Analysis of Activation Patterns\.
##### Differences in Activation Patterns\.
The activation statistics presented in Figure[1](https://arxiv.org/html/2608.08168#S4.F1)reveal distinct numerical patterns between the two reasoning modes\. Specifically,Thinkingmode exhibits a relatively low mean activation of approximately 9\.0 across all difficulty levels\. However, its maximum activation consistently reaches high values around 75\.0, accompanied by a high standard deviation of approximately 19\.0\. This indicates a highly sparse distribution where a small number of features activate intensely and the majority remain suppressed\. In contrast,NoThinkingmode maintains a significantly higher mean activation of approximately 11\.7, with a lower maximum activation stabilizing around 60\.0\. The slightly lower standard deviation of approximately 17\.5 suggests a more uniform and diffuse feature activity\.
##### Consistency of Dominant Feature Intensity\.
The activation intensity of dominant features exhibits intrinsic stability regardless of problem complexity, as evidenced by both visual patterns and quantitative metrics in Figures[2](https://arxiv.org/html/2608.08168#S4.F2)and[3](https://arxiv.org/html/2608.08168#S4.F3)\. Visually, the continuous horizontal bands in Figure[2](https://arxiv.org/html/2608.08168#S4.F2)confirm that the features dominating easy tasks retain their prominence in hard scenarios\. Quantitatively, Figure[3](https://arxiv.org/html/2608.08168#S4.F3)substantiates this robustness\. InThinkingmode, the dominant F4416 maintains high\-intensity activation with values of 75\.2 on easy tasks, 72\.8 on medium tasks, and 73\.9 on hard tasks, significantly outpacing the second\-ranked F8893 \(≈\\approx53\.0\)\. This stability in peak magnitude extends toNoThinkingmode, where the primary F10770 shifts minimally from 60\.5 on easy tasks to 59\.5 on hard tasks\. Overall, while the broader feature composition may evolve, increased difficulty does not induce significant shifts in the activation magnitude of the primary features\.
Figure 3:The activation trajectories of the top\-20 SAE features across difficulty levels for bothThinkingandNoThinkingmodes, highlighting absolute differences in feature intensity\.ModeDifficultyFeature CompositionThinkingEasyF4416 \(82\.0%\), F8893 \(18\.0%\)MediumF4416 \(96\.0%\), F8893 \(4\.0%\)HardF4416 \(100\.0%\), F8893 \(0\.0%\)NoThinkingEasyF10770 \(37\.0%\), F63887 \(25\.0%\), F226 \(20\.0%\), F9911 \(18\.0%\)MediumF10770 \(83\.0%\), F63887 \(15\.0%\), F226 \(1\.0%\), F9911 \(1\.0%\)HardF10770 \(100\.0%\), F63887 \(0\.0%\), F226 \(0\.0%\), F9911 \(0\.0%\)Table 1:Feature source distribution for top\-100 highest\-activation tokens\. Tokens are ranked by activation strength across the top\-20 SAE features to identify dominant feature contributions for each mode\-difficulty pair\.
#### 4\.2\.2Token\-Level Analysis of SAE Feature Activations
While feature\-level analysis reveals macroscopic differences between reasoning modes, understanding which specific tokens trigger these features provides crucial interpretability insights\. This analysis investigates the lexical characteristics of tokens that maximally activate top\-20 SAE features, enabling fine\-grained comparison betweenThinkingandNoThinkingmodes across problem difficulties\.
##### Feature Source Distribution\.
We first analyze the feature source distribution to determine the extent to which the model relies on specialized latent features\. We calculate the composition of the top\-100 highest\-activation tokens for each mode\-difficulty pair and summarize the results in Table[1](https://arxiv.org/html/2608.08168#S4.T1)\.
As shown in Table[1](https://arxiv.org/html/2608.08168#S4.T1), both modes ultimately converge to single\-feature dominance at the hard level, but exhibit fundamentally different trajectories\.Thinkingmode demonstrates early concentration\. Even at the easy level, 82% of top tokens originate from a single feature F4416, increasing monotonically to 100% at the hard level\. In contrast,NoThinkingmode transitions from a distributed state\. Tokens at the easy level disperse across four features, with the top feature F10770 comprising only 37%\. This distribution progressively consolidates into single\-feature dominance\. These findings suggest thatThinkingmode employs a stable computational pathway regardless of complexity, whereasNoThinkingmode adaptively recruits different feature combinations based on task demands\.
ThinkingNoThinkingCategoryEasyMediumHardEasyMediumHardWord28\.229\.437\.721\.732\.436\.1Number17\.020\.58\.016\.012\.85\.9Math Symbol15\.614\.613\.122\.314\.514\.0Reasoning6\.15\.56\.92\.95\.34\.4Variable7\.17\.710\.86\.66\.89\.7
Table 2:Token category distribution comparison across difficulty levels in two reasoning modes \(%\), with all tokens activating top\-20 SAE features involved\.
##### Token Category Distribution\.
We further categorize all activated tokens into functional groups to elucidate their semantic roles, with statistics detailed in Table[2](https://arxiv.org/html/2608.08168#S4.T2)\. The complete categorization taxonomy is provided in Appendix[A\.3](https://arxiv.org/html/2608.08168#A1.SS3)\. Distinct patterns regarding strategic preference and difficulty adaptation are observed from these statistics\.
First, the modes exhibit a fundamental divergence between verbal deduction and symbolic manipulation\.Thinkingmode relies on explicit verbalized logic by activating a high proportion of reasoning tokens including logical connectives \(e\.g\., "therefore", "implies"\) and procedural markers\. In easy tasks, the proportion of these tokens inThinkingmode is nearly double that ofNoThinkingmode\. In contrast,NoThinkingmode employs a formula\-centric strategy characterized by a heavy reliance on math symbols\. Its usage rate of 22\.3% in simple tasks significantly exceeds the 15\.6% observed inThinkingmode\.
Second, both modes demonstrate a consistent transition from numerical processing to conceptual abstraction as task difficulty increases\. InThinkingmode, the frequency of number tokens declines from 17\.0% to 8\.0%, while word tokens increase from 28\.2% to 37\.7%\.NoThinkingmode mirrors this pattern with number tokens falling from 16\.0% to 5\.9% and word tokens rising from 21\.7% to 36\.1%\. These results indicate that high\-difficulty tasks necessitate linguistic reasoning rather than direct numerical computation regardless of the specific generation strategy\.
##### Qualitative Context Analysis\.
By examining specific activation contexts, we identify distinct functional roles for the dominant features, as detailed in Table[3](https://arxiv.org/html/2608.08168#S4.T3)\. F4416 serves as a semantic proxy forexploratory reasoningandself\-correction\. It exhibits strong activation during intermediate computational steps \(e\.g\.,"8\.97 squared is…"\) and aligns with epistemic markers such as"Let’s see"or"isn’t working", effectively capturing the iterative and trial\-and\-error nature of the cognitive process\. F10770 functions primarily as asyntactic formatter\. Its activations are densely concentrated on LaTeX syntax \(e\.g\.,\\infty\) and formal notation\. This indicates thatNoThinkingmode bypasses intermediate logical derivation, prioritizing the retrieval and structural formatting of the final solution\.
ThinkingMode \(F4416\)NoThinkingMode \(F10770\)Functional RoleFunctional RoleReasoning Monitor: Focuses on dynamic and trial\-and\-error solution finding processes\.Structural Formatter: Focuses on syntax retrieval and structural formatting of the final solution\.Representative ActivationsRepresentative Activations•Calculation Trace:
"…8\.97 squared is about 80\.4609, still less than80\.92\. 8\.98 squared is approximately 80\.640…"•Epistemic Marker:
"…Because other pairs might have the last letter as L or something else\. Wait, let’ssee…"•Error Detection:
"…Wait, now I’m really confused\. Maybe this approach isn’tworking\. Let’s try to write down the equations…"•LaTeX Syntax
…p9=∑m=1\\infty12m…\\dots p\_\{9\}=\\sum\_\{m=1\}^\{\\text\{\{\\textbackslash in\\hbox\{\\pagecolor\{nothinkingbg\}\{\\color\[rgb\]\{0\.1171875,0\.51953125,0\.28515625\}fty\}\}\}\}\}\\frac\{1\}\{2^\{m\}\}\\dots
•Structure Block
…\(306\+289\)=919\\end\{align\*\}…\\dots\(306\+289\)=919\\;\\text\{\{\\textbackslash end\\\{\\hbox\{\\pagecolor\{nothinkingbg\}\{\\color\[rgb\]\{0\.1171875,0\.51953125,0\.28515625\}align\}\}\*\\\}\}\}\\dots
•Formal Logic
…sin\(5x\)=π2\+πm\\impliessin\(5x\)=1\+2m14…\\dots\\sin\(5x\)=\\frac\{\\pi\}\{2\}\+\\pi m\\text\{\{\\textbackslash impl\\hbox\{\\pagecolor\{nothinkingbg\}\{\\color\[rgb\]\{0\.1171875,0\.51953125,0\.28515625\}ies\}\}\}\}\\;\\sin\(5x\)=\\frac\{1\+2m\}\{14\}\\dotsTable 3:Analysis of activation contexts for dominant features F4416 \(Thinking\) and F10770 \(NoThinking\)\.Examples illustrate typical activation scenarios, where highlighted substrings denote the tokens with the highest activation scores within the context window\.
### 4\.3Causal Analysis
To establish the functional necessity of the identified sparse features, we conducted causal interventions on the top\-3 features with the highest Total Activation Volume \(TAV\):F28634,F4416, andF8893\. As shown in Table[4](https://arxiv.org/html/2608.08168#S4.T4), our results uncover three fundamental mechanisms governing the model’s reasoning process\.
##### Coupling of Reasoning and Syntactic Structure\.
Our experiments indicate a functional link between the explicit reasoning process and the generation of mathematical syntax\. As shown in Table[4](https://arxiv.org/html/2608.08168#S4.T4), suppressing critical features for reasoning severely impairs the model’s ability to produce formal output regardless of their specific roles\. Specifically, we observe a consistent drop inLaTeXdensity of−29\.54\-29\.54to−40\.28\-40\.28per 1,000 tokens alongside a near\-total failure to format solutions where Boxed Answer Retention frequently fell to0%0\\%\. These findings suggest that the sparse features driving the Chain\-of\-Thought \(e\.g\., F4416, F28634\) simultaneously encode the structural representations required for formal output, indicating that reasoning and formatting are not processed by independent modules\.
##### Compensatory Sequence Extension\.
Thinkingmode exhibits a compensatory mechanism when disrupted\. When the core reasoning feature F28634 is suppressed, the model does not terminate generation but instead exhibits an expansion of the output sequence\. This is evidenced by a 454% increase in output length\. Notably, this expansion is inversely correlated with generation quality where lexical diversity \(Distinct\-1\) declines by 63%, indicating that the model produces repetitive and low\-information sequences\. Furthermore, suppressing F28634 leads to a significant increase in metacognitive density \(\+34\.17\)\. This suggests that when the primary reasoning vector is blocked, the model generates additional epistemic markers \(e\.g\., “Wait,” “Let me think”\) and extends the sequence length, attempting to maintain the generative state despite the absence of effective computational progress\.
Metric \(Change per 1k tokens\)F28634F4416F8893Metacognitive Density \(Δ\\Delta\)\+34\.17\-19\.38\-4\.33Uncertainty Density \(Δ\\Delta\)\+9\.25\-4\.16\-0\.99LaTeXDensity \(Δ\\Delta\)\-40\.28\-35\.22\-29\.54Boxed Answer Retention0%10%0%Output Length Change\+454%\+410%\+107%Lexical Diversity \(Distinct\-1\)\-63%\-37%\-42%
Table 4:Impact of Feature Suppression on Reasoning Metrics\.We report relative changes \(Δ\\Delta\) from baseline\. Key observations include: \(1\) a universal decline in mathematical formalism \(reducedLaTeXdensity and boxed answers\); \(2\) divergent shifts in metacognition, which increases for F28634 \(\+34\.17\) but decreases for F4416 \(\-19\.38\) ; and \(3\) significant output length expansion \(up to \+454%\) accompanied by reduced lexical diversity across all groups\.
##### Fragile Coordination under Feature Suppression\.
Finally, our results suggest thatThinkingmode relies on a fragile coordination among distinct high\-impact features rather than a uniformly robust mechanism\. Importantly, any significant change in metacognitive density reflects a deviation from a stable state rather than an improvement\. For instance, suppressing feature 28634 increases metacognitive density by 34\.17 while suppressing feature 4416 decreases the same metric by 19\.38\. This contrast shows that different features regulate the process in opposite directions\. Despite these divergent internal effects, the structural indicators collapsed consistently\. We observe thatLaTeXdensity drops sharply and boxed answer retention remains near zero across all features\. This pattern indicates a coupled control system where specific features jointly maintain both process monitoring and output structure\. Consequently, disrupting any single feature drives the model into distinct failure modes such as uncontrolled verbosity\.
## 5Conclusion
In this work, we use Top\-K Sparse Autoencoders to probe intermediate representations in DeepSeek\-R1\-Distill\-Qwen\-7B and mechanistically distinguishThinkingfromNoThinking\. Observationally,Thinkingmode relies on sparse and high\-intensity feature activations driving verbal deduction independent of problem complexity, whereasNoThinkingmode exhibits an adaptive and diffuse pattern prioritizing symbolic manipulation\. Causally, targeted suppression of the most active sparse features shows that reasoning and output structure are tightly coupled, since interventions consistently disruptLaTeXand boxed\-solution formatting; it also reveals a compensatory response inThinking, where disruption triggers longer and more metacognitively signaled but less informative continuations\. Together, these findings characterize Chain\-of\-Thought as a finely tuned, low\-redundancy control regime maintained by coordinated feature interactions rather than a standalone reasoning module, pointing toward feature\-level control as a path to more reliable and controllable reasoning behavior\.
## References
- K\. D\. B\. J\. Adamet al\.\(2014\)A method for stochastic optimization\.arXiv preprint arXiv:1412\.69801412\(6\)\.Cited by:[§4\.1\.3](https://arxiv.org/html/2608.08168#S4.SS1.SSS3.Px3.p1.7)\.
- D\. Arad, A\. Mueller, and Y\. Belinkov \(2025\)SAEs are good for steering–if you select the right features\.arXiv preprint arXiv:2505\.20063\.Cited by:[§1](https://arxiv.org/html/2608.08168#S1.p2.1),[§2](https://arxiv.org/html/2608.08168#S2.SS0.SSS0.Px3.p1.1)\.
- T\. Bricken, A\. Templeton, J\. Batson, B\. Chen, A\. Jermyn, T\. Conerly, N\. Turner, C\. Anil, C\. Denison, A\. Askell,et al\.\(2023\)Towards monosemanticity: decomposing language models with dictionary learning\.Transformer Circuits Thread2\.Cited by:[§1](https://arxiv.org/html/2608.08168#S1.p2.1),[§3\.2](https://arxiv.org/html/2608.08168#S3.SS2.p1.6)\.
- X\. Chen, A\. Plaat, and N\. van Stein \(2025\)How does chain of thought think? mechanistic interpretability of chain\-of\-thought reasoning with sparse autoencoding\.arXiv preprint arXiv:2507\.22928\.Cited by:[§1](https://arxiv.org/html/2608.08168#S1.p1.1)\.
- H\. Cunningham, A\. Ewart, L\. Riggs, R\. Huben, and L\. Sharkey \(2023\)Sparse autoencoders find highly interpretable features in language models\.arXiv preprint arXiv:2309\.08600\.Cited by:[§1](https://arxiv.org/html/2608.08168#S1.p2.1),[§2](https://arxiv.org/html/2608.08168#S2.SS0.SSS0.Px1.p1.3)\.
- J\. Gallifant, S\. Chen, K\. Sasse, H\. Aerts, T\. Hartvigsen, and D\. Bitterman \(2025\)Sparse autoencoder features for classifications and transferability\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 29927–29951\.Cited by:[§1](https://arxiv.org/html/2608.08168#S1.p2.1),[§2](https://arxiv.org/html/2608.08168#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2608.08168#S2.SS0.SSS0.Px3.p1.1)\.
- L\. Gao, T\. D\. la Tour, H\. Tillman, G\. Goh, R\. Troll, A\. Radford, I\. Sutskever, J\. Leike, and J\. Wu \(2024\)Scaling and evaluating sparse autoencoders\.arXiv preprint arXiv:2406\.04093\.Cited by:[§A\.1](https://arxiv.org/html/2608.08168#A1.SS1.SSS0.Px2.p1.9),[§1](https://arxiv.org/html/2608.08168#S1.p2.1),[§2](https://arxiv.org/html/2608.08168#S2.SS0.SSS0.Px1.p1.3),[§2](https://arxiv.org/html/2608.08168#S2.SS0.SSS0.Px2.p1.1),[§4\.1\.3](https://arxiv.org/html/2608.08168#S4.SS1.SSS3.Px3.p1.7)\.
- D\. Guo, D\. Yang, H\. Zhang, J\. Song, R\. Zhang, R\. Xu, Q\. Zhu, S\. Ma, P\. Wang, X\. Bi,et al\.\(2025\)Deepseek\-r1: incentivizing reasoning capability in llms via reinforcement learning\.arXiv preprint arXiv:2501\.12948\.Cited by:[§1](https://arxiv.org/html/2608.08168#S1.p1.1),[§4\.1\.1](https://arxiv.org/html/2608.08168#S4.SS1.SSS1.p1.1)\.
- Z\. He, T\. Liang, J\. Xu, Q\. Liu, X\. Chen, Y\. Wang, L\. Song, D\. Yu, Z\. Liang, W\. Wang,et al\.\(2025\)Deepmath\-103k: a large\-scale, challenging, decontaminated, and verifiable mathematical dataset for advancing reasoning\.arXiv preprint arXiv:2504\.11456\.Cited by:[§4\.1\.2](https://arxiv.org/html/2608.08168#S4.SS1.SSS2.p1.1)\.
- Y\. Jing, Z\. Yao, H\. Guo, L\. Ran, X\. Wang, L\. Hou, and J\. Li \(2025\)LinguaLens: towards interpreting linguistic mechanisms of large language models via sparse auto\-encoder\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 28220–28239\.Cited by:[§2](https://arxiv.org/html/2608.08168#S2.SS0.SSS0.Px2.p1.1)\.
- J\. Li, M\. Galley, C\. Brockett, J\. Gao, and W\. B\. Dolan \(2016\)A diversity\-promoting objective function for neural conversation models\.InProceedings of the 2016 conference of the North American chapter of the association for computational linguistics: human language technologies,pp\. 110–119\.Cited by:[Table 5](https://arxiv.org/html/2608.08168#A1.T5.1.7.2.1.1)\.
- K\. Li, A\. K\. Hopkins, D\. Bau, F\. Viégas, H\. Pfister, and M\. Wattenberg \(2022\)Emergent world representations: exploring a sequence model trained on a synthetic task\.arXiv preprint arXiv:2210\.13382\.Cited by:[§1](https://arxiv.org/html/2608.08168#S1.p1.1)\.
- Z\. Li, X\. Wang, Y\. Yang, Z\. Yao, H\. Xiong, and M\. Du \(2025\)Feature extraction and steering for enhanced chain\-of\-thought reasoning in language models\.arXiv preprint arXiv:2505\.15634\.Cited by:[§1](https://arxiv.org/html/2608.08168#S1.p2.1)\.
- T\. Lieberum, S\. Rajamanoharan, A\. Conmy, L\. Smith, N\. Sonnerat, V\. Varma, J\. Kramár, A\. Dragan, R\. Shah, and N\. Nanda \(2024\)Gemma scope: open sparse autoencoders everywhere all at once on gemma 2\.arXiv preprint arXiv:2408\.05147\.Cited by:[§A\.1](https://arxiv.org/html/2608.08168#A1.SS1.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2608.08168#S1.p2.1),[§2](https://arxiv.org/html/2608.08168#S2.SS0.SSS0.Px1.p1.3),[§4\.1\.3](https://arxiv.org/html/2608.08168#S4.SS1.SSS3.Px3.p1.7)\.
- H\. Lightman, V\. Kosaraju, Y\. Burda, H\. Edwards, B\. Baker, T\. Lee, J\. Leike, J\. Schulman, I\. Sutskever, and K\. Cobbe \(2023\)Let’s verify step by step\.InThe Twelfth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2608.08168#S1.p1.1)\.
- W\. Ma, J\. He, C\. Snell, T\. Griggs, S\. Min, and M\. Zaharia \(2025\)Reasoning models can be effective without thinking\.arXiv preprint arXiv:2504\.09858\.Cited by:[§4\.1\.3](https://arxiv.org/html/2608.08168#S4.SS1.SSS3.Px1.p1.1)\.
- A\. Makhzani and B\. Frey \(2013\)K\-sparse autoencoders\.arXiv preprint arXiv:1312\.5663\.Cited by:[§3\.3](https://arxiv.org/html/2608.08168#S3.SS3.p1.2)\.
- K\. Meng, D\. Bau, A\. Andonian, and Y\. Belinkov \(2022\)Locating and editing factual associations in gpt\.Advances in neural information processing systems35,pp\. 17359–17372\.Cited by:[§A\.1](https://arxiv.org/html/2608.08168#A1.SS1.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2608.08168#S1.p2.1)\.
- N\. Nanda, L\. Chan, T\. Lieberum, J\. Smith, and J\. Steinhardt \(2023\)Progress measures for grokking via mechanistic interpretability\.arXiv preprint arXiv:2301\.05217\.Cited by:[§1](https://arxiv.org/html/2608.08168#S1.p1.1)\.
- S\. Park, T\. Kim, and Y\. Ko \(2025\)Decoding dense embeddings: sparse autoencoders for interpreting and discretizing dense retrieval\.arXiv preprint arXiv:2506\.00041\.Cited by:[§2](https://arxiv.org/html/2608.08168#S2.SS0.SSS0.Px3.p1.1)\.
- S\. Rajamanoharan, A\. Conmy, L\. Smith, T\. Lieberum, V\. Varma, J\. Kramár, R\. Shah, and N\. Nanda \(2024\)Improving dictionary learning with gated sparse autoencoders\.arXiv preprint arXiv:2404\.16014\.Cited by:[§1](https://arxiv.org/html/2608.08168#S1.p2.1)\.
- A\. Templeton, T\. Conerly, J\. Marcus, J\. Lindsey, T\. Bricken, B\. Chen, A\. Pearce, C\. Citro, E\. Ameisen, A\. Jones, H\. Cunningham, N\. L\. Turner, C\. McDougall, M\. MacDiarmid, A\. Tamkin, E\. E, S\. Kapoor, J\. Kaplan, S\. Fort, N\. Nanda, and C\. Olah \(2024\)Scaling monosemanticity: extracting interpretable features from claude 3 sonnet\.Note:[https://transformer\-circuits\.pub/2024/scaling\-monosemanticity/index\.html](https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html)Transformer Circuits ThreadExternal Links:[Link](https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html)Cited by:[§A\.1](https://arxiv.org/html/2608.08168#A1.SS1.SSS0.Px1.p1.1)\.
- \[23\]J\. Theodorus, V\. Swaytha, S\. Gautam, A\. Ward, M\. Shah, C\. Blondin, and K\. ZhuFinding sparse autoencoder representations of errors in cot prompting\.InICLR 2025 Workshop on Building Trust in Language Models and Applications,Cited by:[§1](https://arxiv.org/html/2608.08168#S1.p1.1)\.
- M\. Turpin, J\. Michael, E\. Perez, and S\. Bowman \(2023\)Language models don’t always say what they think: unfaithful explanations in chain\-of\-thought prompting\.Advances in Neural Information Processing Systems36,pp\. 74952–74965\.Cited by:[§1](https://arxiv.org/html/2608.08168#S1.p1.1)\.
- J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, F\. Xia, E\. Chi, Q\. V\. Le, D\. Zhou,et al\.\(2022\)Chain\-of\-thought prompting elicits reasoning in large language models\.Advances in neural information processing systems35,pp\. 24824–24837\.Cited by:[§1](https://arxiv.org/html/2608.08168#S1.p1.1)\.
- \[26\]X\. Wu, J\. Yuan, W\. Yao, X\. Zhai, and N\. LiuInterpreting and steering llm representations with mutual information\-based explanations on sparse autoencoders\.Cited by:[§4\.1\.3](https://arxiv.org/html/2608.08168#S4.SS1.SSS3.Px3.p1.7)\.
- A\. Zou, L\. Phan, S\. Chen, J\. Campbell, P\. Guo, R\. Ren, A\. Pan, X\. Yin, M\. Mazeika, A\. Dombrowski,et al\.\(2023\)Representation engineering: a top\-down approach to ai transparency\.arXiv preprint arXiv:2310\.01405\.Cited by:[§1](https://arxiv.org/html/2608.08168#S1.p1.1)\.
## Appendix AExperiment Setup
### A\.1Hyperparameters
##### Selection of Targeted Layer\.
We specifically target the residual stream of the 13th layer\. This decision follows recent insights from sparse autoencoder research\[[22](https://arxiv.org/html/2608.08168#bib.bib15),[14](https://arxiv.org/html/2608.08168#bib.bib22)\]\. These studies identify intermediate layers as the core of high\-level reasoning\. Early layers mostly handle local syntax\. Late layers focus on predicting the next token\. In contrast, intermediate layers contain the richest semantic information\[[18](https://arxiv.org/html/2608.08168#bib.bib16)\]\. Therefore, this layer is optimal for capturing the reasoning traces in Chain\-of\-Thought processes\.
##### Selection of Feature DimensionsCC\.
This choice is governed by the scaling lawC∝ZγC\\propto Z^\{\\gamma\}\[[7](https://arxiv.org/html/2608.08168#bib.bib20)\], whereZZdenotes the number of training tokens, with empirical exponents ranging fromγ≈0\.60\\gamma\\approx 0\.60\(GPT\-2 Small\) toγ≈0\.65\\gamma\\approx 0\.65\(GPT\-4\)\. Given the fixed number of feature vectorsCC, the implied scaling exponents areγ≈0\.62\\gamma\\approx 0\.62for theNoThinkingmode \(Z≈65MZ\\approx 65\\text\{M\}\) andγ≈0\.60\\gamma\\approx 0\.60for theThinkingmode \(Z≈108MZ\\approx 108\\text\{M\}\)\. Both values fall consistently within the established range, validating the rationality of our architectural choices\.
### A\.2Details for Benchmarks
To systematically verify the behavioral differences and internal feature dynamics of the reasoning model under varying levels of problem complexity, we categorize our evaluation benchmarks into three difficulty levels:Easy\(AMC23\)444https://huggingface\.co/datasets/AI\-MO/aimo\-validation\-amc,Medium\(AIME 24 & AIME 25 \)555https://huggingface\.co/datasets/AI\-MO/aimo\-validation\-aime, andHard\(OlympiadBench\)666https://huggingface\.co/datasets/Hothan/OlympiadBench\.
- •Easy\(AMC23\)\. Comprising 40 problems from the 2023 American Mathematics Competition\. It focuses on high\-school foundational topics such as algebraic manipulations and geometric principles, with all solutions constrained to integers between 0 and 999\.
- •Medium\(AIME 24 & AIME 25\)\. Consisting of 30 problems from the 2024 American Invitational Mathematics Examinations and 30 novel problems curated in 2025, respectively\. Unlike the AMC, these tasks necessitate deep combinatorial and geometric insights involving multi\-step reasoning and significantly higher computational complexity\. The answers are strictly constrained to integers from 0 to 999\.
- •Hard\(OlympiadBench\)\. A comprehensive corpus of 8,476 Olympiad\-level problems sourced from elite competitions such as the IMO and the Chinese Gaokao, which is much more challenging than AIME and AMC\. It is characterized by multi\-modal inputs \(e\.g\., diagrams\) and expert solutions that involve complex, long\-horizon logical chains\. We selected a subset of 503 samples from this benchmark for our analysis\.
### A\.3Token Categorization
We define a hierarchical taxonomy of 5 token categories to characterize the semantic properties of highly\-activated tokens: \(1\) reasoning: logical connectives \(e\.g\.,"therefore","since","implies"\), procedural markers \(e\.g\.,"first","step","then"\), and problem\-solving directives \(e\.g\.,"let","solve","calculate"\); \(2\) number: pure digit sequences \(e\.g\.,"123","2024"\); \(3\) math symbol: mathematical operators and brackets \(e\.g\.,"\+","=","\(\)"\); \(4\) variable: single alphabetic characters typically representing mathematical variables \(e\.g\.,"x","n","A"\); \(5\) word: multi\-character alphabetic strings \(e\.g\.,"the","angle"\)\.
DimensionMetricKey Indicators & DescriptionCognitive AbilityMetacognitive Density16 Markers:"wait", "hmm", "actually", "let’s see", "perhaps", "maybe", "alternatively", "hold on", "thinking about", "let me", "i think", "seems like", "looks like", "but wait", "oh wait", "hang on"\.Uncertainty Density9 Markers:"might", "could", "possibly", "probably", "maybe", "perhaps", "seems", "appears", "likely"\.Math FormalizationLaTeX DensityCounts mathematical environments, including both display \(\\\[\.\.\.\\\]\) and inline \(\\\(\.\.\.\\\)\) syntax usage\.Boxed Answer RetentionA binary indicator measuring the successful generation of the strict final answer format:\\boxed\{\.\.\.\}\.Generative StateOutput LengthTracks the total token count of the generated chain\-of\-thought to monitor verbosity changes\.Lexical DiversityQuantified viaDistinct\-1\[[11](https://arxiv.org/html/2608.08168#bib.bib34)\]: the ratio of unique unigrams to total tokens\. Low values indicate repetitive looping or mode collapse\.Table 5:Summary of Evaluation Metrics for Causal Interventions\.The framework categorizes metrics into cognitive, structural, and generative dimensions\. Note that all density metrics are normalized per 1,000 tokens\.
### A\.4Casual Evaluation Metrics
To quantify the impact of SAE feature interventions on model reasoning behavior, we define a comprehensive set of metrics capturing linguistic patterns, mathematical formalization and output characteristics, as detailed in Table[5](https://arxiv.org/html/2608.08168#A1.T5)\. All density metrics are normalized per 1,000 tokens to enable fair comparison across responses of varying lengths\. Specifically, we assessed cognitive ability through Metacognitive Density and Uncertainty Density\. The former tracks the frequency of 16 specific markers, while the latter considers a set of 9 indicators\. In addition, mathematical formalization was evaluated via LaTeX Density, measuring the prevalence of symbolic notation, and Boxed Answer Retention, which serves as a binary indicator of the model’s capacity to formulate valid and well\-structured conclusions\. Furthermore, we analyzed generative characteristics using Output Length Change and Lexical Diversity to detect behavioral anomalies such as repetitive looping or verbose degeneration indicative of reasoning breakdown\.Similar Articles
Efficient Reasoning Distillation: Small Video-Language Models via Synthetic CoT and Difficulty-Aware Fine-Tuning
The paper presents a method to distill reasoning into compact video-language models using synthetic chain-of-thought rationales and difficulty-aware fine-tuning, enabling smaller models to outperform larger ones with minimal compute.
Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders
This paper applies TopK Sparse Autoencoders to three EEG foundation models (SleepFM, REVE, LaBraM) to extract interpretable feature dictionaries and introduces a framework for concept steering, revealing representational failures and clinical entanglements.
Rethinking Dense Sequential Chains: Reasoning Language Models Can Extract Answers from Sparse, Order-Shuffling Chain-of-Thoughts
This research paper from MediaTek and National Taiwan University challenges the assumption that reasoning chains must be dense and sequential, showing that models can extract answers from sparse, shuffled, and noisy reasoning traces. The findings suggest that answer extraction is robust and order-independent, potentially enabling more efficient, parallelized reasoning generation.
Can We Understand How Large Language Models Reason?
This article explores the ongoing efforts and challenges in understanding how large language models reason, focusing on interpretability research.
Revisiting Complete Reasoning Traces for Post-Training
This paper finds that large language models can gain reasoning improvements from truncated reasoning trajectories rather than full ones during post-training, reducing redundancy while benefiting methods like supervised fine-tuning and reinforcement learning.