Observable Patterns Are Not Explanations: A Causal-Geometric Analysis of Latent Reasoning Models

arXiv cs.CL Papers

Summary

This paper analyzes latent reasoning models (LRMs) and demonstrates that observable patterns in latent states are not causal explanations of reasoning; it advocates for matched controls and causal tests in interpretability research.

arXiv:2606.12689v1 Announce Type: new Abstract: Latent reasoning models (LRMs) replace explicit chain-of-thought with continuous thoughts. Recent work treats observable latent-state patterns, such as BFS-like frontiers and decodable arithmetic computation, as evidence for internal reasoning mechanisms. Evaluating two LRMs (Coconut and CODI) against controls lacking the proposed recurrence or curriculum, we find these patterns also appear in the controls and do not always causally affect behavior. Causal interventions reveal that latent-thought utilization is not binary but graded, scaling with a thought's causal effect on model behavior. Geometric analyses reveal this effect concentrates in low-rank directions whose step-to-step geometry grows more structured as their behavioral influence increases. Latent thoughts should therefore be treated as hidden computation, not hidden explanation: decodability, attention, or static structure alone cannot establish mechanism. LRM interpretability thus requires matched controls and causal tests.
Original Article
View Cached Full Text

Cached at: 06/12/26, 08:50 AM

# A Causal-Geometric Analysis of Latent Reasoning Models
Source: [https://arxiv.org/html/2606.12689](https://arxiv.org/html/2606.12689)
## Observable Patterns Are Not Explanations: A Causal\-Geometric Analysis of Latent Reasoning Models

Darpan Aswal1,2Thomas Palmeira Ferraz1,3Yongxin Zhou1Maxime Peyrard1 1Université Grenoble Alpes, CNRS, Grenoble INP, LIG 2Université Paris\-Saclay3NAVER LABS Europe darpan\.aswal@universite\-paris\-saclay\.fr

###### Abstract

Latent reasoning models \(LRMs\) replace explicit chain\-of\-thought with continuous thoughts\. Recent work treats observable latent\-state patterns, such as BFS\-like frontiers and decodable arithmetic computation, as evidence for internal reasoning mechanisms\. Evaluating two LRMs \(Coconut and CODI\) against controls lacking the proposed recurrence or curriculum, we find these patterns also appear in the controls and do not always causally affect behavior\. Causal interventions reveal that latent\-thought utilization is not binary but graded, scaling with a thought’s causal effect on model behavior\. Geometric analyses reveal this effect concentrates in low\-rank directions whose step\-to\-step geometry grows more structured as their behavioral influence increases\. Latent thoughts should therefore be treated as hidden computation, not hidden explanation: decodability, attention, or static structure alone cannot establish mechanism\. LRM interpretability thus requires matched controls and causal tests\.

Observable Patterns Are Not Explanations: A Causal\-Geometric Analysis of Latent Reasoning Models

Darpan Aswal1,2Thomas Palmeira Ferraz1,3Yongxin Zhou1Maxime Peyrard11Université Grenoble Alpes, CNRS, Grenoble INP, LIG2Université Paris\-Saclay3NAVER LABS Europedarpan\.aswal@universite\-paris\-saclay\.fr

## 1Introduction

Chain\-of\-thought \(CoT\) prompting has become a standard approach for eliciting reasoning through verbalized intermediate steps in natural language\(Weiet al\.,[2022](https://arxiv.org/html/2606.12689#bib.bib28); Liet al\.,[2025](https://arxiv.org/html/2606.12689#bib.bib44); Lyuet al\.,[2023](https://arxiv.org/html/2606.12689#bib.bib45)\)\. However, CoT is computationally costly and fundamentally constrained by the discreteness of natural\-language tokens\(Zhanget al\.,[2025b](https://arxiv.org/html/2606.12689#bib.bib29)\)\. To address these limitations, latent reasoning models \(LRMs\) perform intermediate computation in continuous hidden states\(Haoet al\.,[2025](https://arxiv.org/html/2606.12689#bib.bib1); Shenet al\.,[2025](https://arxiv.org/html/2606.12689#bib.bib9)\)\. LRMs promise improved efficiency, but they also remove the primary artifact available for monitoring and oversight: visible CoT traces\. Despite well\-known faithfulness concerns\(Chenet al\.,[2025](https://arxiv.org/html/2606.12689#bib.bib30); Turpinet al\.,[2023](https://arxiv.org/html/2606.12689#bib.bib32)\), CoT traces still provide human\-interpretable signals that can be monitored for harmful or misaligned behavior\(Bereska and Gavves,[2024](https://arxiv.org/html/2606.12689#bib.bib33); Korbaket al\.,[2025](https://arxiv.org/html/2606.12689#bib.bib46)\)\. In contrast, LRMs reason through continuous states, raising new safety and interpretability concerns as they scale and are deployed in agentic systems\(Chanet al\.,[2023](https://arxiv.org/html/2606.12689#bib.bib31)\)\. This calls for a better understanding of their internal dynamics\.

Early interpretability efforts have searched for latent\-state analogues of CoT traces, aiming to decode interpretable patterns from latent thoughts\. For example,Haoet al\.\([2025](https://arxiv.org/html/2606.12689#bib.bib1)\)andZhuet al\.\([2025](https://arxiv.org/html/2606.12689#bib.bib25)\)argue that continuous latent thoughts demonstrate breadth\-first\-search\-like reasoning over multiple superposed candidate reasoning paths, whileShenet al\.\([2025](https://arxiv.org/html/2606.12689#bib.bib9)\)andWeiet al\.\([2026](https://arxiv.org/html/2606.12689#bib.bib35)\)suggest that implicit supervision can yield interpretable intermediate reasoning steps\. Later work has also used local interventions to investigate these findings\(Cywinskiet al\.,[2025](https://arxiv.org/html/2606.12689#bib.bib20); Peterset al\.,[2025](https://arxiv.org/html/2606.12689#bib.bib21)\)\. However, several methodological concerns remain, leaving important questions unaddressed\. First,emergence: whether the observed pattern emerges specifically from latent\-reasoning mechanisms, or also appears in matched non\-LRM controls\. Second,generalizability: whether the pattern is characteristic of latent reasoning more broadly, rather than specific to one task, model, curriculum, or training recipe\. Third,causality: whether the pattern actually contributes to the model’s behavior, rather than merely correlating with successful outcomes\. Without addressing these questions, observable patterns may lead to superficial evidence for reasoning\.

In this work, instead of treating latent thoughts as hidden explanations, we argue in favor of treating them as extended hidden computation states whose influence strength must be measured by their causal effect on model behavior\. Our approach contrasts with prior work also taking a causal perspective by asking binarily whether a latent thought is “used” or “not used”\(Zhanget al\.,[2025a](https://arxiv.org/html/2606.12689#bib.bib23); Cuiet al\.,[2026](https://arxiv.org/html/2606.12689#bib.bib22)\), shifting the question to where causal influence is concentrated within the latent representation\. These causally effective regions then become the object of analysis: once isolated, their geometry and dynamics can be studied\. This framing leads to the following research questions:

RQ1\.Do observable patterns in LRM latent space uniquely emerge from, and explain the performance of, latent\-reasoning mechanisms?In §[4](https://arxiv.org/html/2606.12689#S4), we compare LRMs against matched controls and find that reasoning patterns previously attributed to latent recurrence can also emerge in curriculum\-matched non\-recurrent models, and even in untrained models with forced thought positions\. This shows that observable patterns alone do not sufficiently establish mechanism\.

RQ2\.When do latent thoughts influence model predictions, and where in the latent state does this influence come from?In §[5](https://arxiv.org/html/2606.12689#S5), we combine latent thought ablations, causal tracing, and gradient\-subspace interventions to show that latent thought influence lies on a continuum\. In particular, the effect of latent thoughts is often concentrated in low\-rank, loss\-sensitive directions \(a causal gradient subspace\) rather than distributed across the full latent representation\.

RQ3\.Do behavior\-influential latent thoughts differ from weakly influential ones in how their representations evolve across steps?In §[6](https://arxiv.org/html/2606.12689#S6), we study the geometry and dynamics of both full latent thought trajectories and causally influential subspaces\. We find that weakly influential latent thoughts are often nearly static across steps, while behavior\-influential thoughts exhibit more structured evolution\.

Overall, our work shifts the target of interpretability from observational proxies to the actual causal trajectories within latent computation\. Finally, we discuss practical implications in more detail in §[7](https://arxiv.org/html/2606.12689#S7)\.

## 2Background and Related Work

### 2\.1Latent Reasoning Models

Unlike standard Chain\-of\-Thought \(CoT\) models that generate textual rationales, latent reasoning models such as implicit CoT\(Denget al\.,[2023](https://arxiv.org/html/2606.12689#bib.bib16)\)and recurrent/looped transformers\(Giannouet al\.,[2023](https://arxiv.org/html/2606.12689#bib.bib18); Dehghaniet al\.,[2019](https://arxiv.org/html/2606.12689#bib.bib19)\)replace intermediate steps with hidden state computations to bypass vocabulary projection\. We focus on*autoregressive continuous\-thought*models which insertKKlatent positions between the promptxxand the answeryyby recursively feeding projected soft\-tokenset\+1late^\{\\mathrm\{lat\}\}\_\{t\+1\}from hidden\-states:

ht=Fθ​\(E​\(x\),e1:tlat\)last,et\+1lat=g​\(ht\),h\_\{t\}=F\_\{\\theta\}\\\!\\left\(E\(x\),e^\{\\mathrm\{lat\}\}\_\{1:t\}\\right\)\_\{\\mathrm\{last\}\},\\qquad e^\{\\mathrm\{lat\}\}\_\{t\+1\}=g\(h\_\{t\}\),withFθF\_\{\\theta\}a causal transformer,EEthe token embedding map, andggidentity or learned\. We callht∈ℝdh\_\{t\}\\in\\mathbb\{R\}^\{d\}the*latent thought*\. AfterKKlatent steps, autoregressive decoding resumes:

pθ​\(y∣x,e1:Klat\)=∏mpθ​\(ym∣x,e1:Klat,y<m\)\.p\_\{\\theta\}\(y\\mid x,e^\{\\mathrm\{lat\}\}\_\{1:K\}\)=\\prod\_\{m\}p\_\{\\theta\}\(y\_\{m\}\\mid x,e^\{\\mathrm\{lat\}\}\_\{1:K\},y\_\{<m\}\)\.
We study two instances:COCONUTgradually replaces CoT segments with latent states via a staged curriculum\(Haoet al\.,[2025](https://arxiv.org/html/2606.12689#bib.bib1)\), whileCODIcompresses textual rationales into latents via self\-distillation\(Shenet al\.,[2025](https://arxiv.org/html/2606.12689#bib.bib9)\)\.

### 2\.2Observational vs\. Causal Interpretability

Mechanistic interpretability seeks to reverse engineer neural network computations into human\- understandable algorithms\(Bereska and Gavves,[2024](https://arxiv.org/html/2606.12689#bib.bib33); Geigeret al\.,[2025](https://arxiv.org/html/2606.12689#bib.bib47)\)\. A key distinction separates information that is*observable*from activations from information that is*causally used*: probes and visualization methods can reveal correlated structure without establishing mechanism\(Belinkov,[2022](https://arxiv.org/html/2606.12689#bib.bib48); Jain and Wallace,[2019](https://arxiv.org/html/2606.12689#bib.bib50); Elazaret al\.,[2021](https://arxiv.org/html/2606.12689#bib.bib6); Lasriet al\.,[2022](https://arxiv.org/html/2606.12689#bib.bib7); Teneyet al\.,[2022](https://arxiv.org/html/2606.12689#bib.bib5)\)\. This issue is central for LRMs, where claims about latent reasoning often rely on observational readouts, without showing that these structures drive performance\. This risks conflatingdecodabilitywithmechanism, an analogue of the*dead salmon*effect\(Mélouxet al\.,[2025a](https://arxiv.org/html/2606.12689#bib.bib14)\), symptoms of a general tendency of interpretability research to produce false positive findings\(Hewitt and Liang,[2019](https://arxiv.org/html/2606.12689#bib.bib3); Ravichanderet al\.,[2021](https://arxiv.org/html/2606.12689#bib.bib2); Kantamneniet al\.,[2025](https://arxiv.org/html/2606.12689#bib.bib4); Mélouxet al\.,[2025a](https://arxiv.org/html/2606.12689#bib.bib14)\)\. We rely on intervention\-based methods \(e\.g\., causal tracing, activation patching\)\(Menget al\.,[2022](https://arxiv.org/html/2606.12689#bib.bib58); Wanget al\.,[2023](https://arxiv.org/html/2606.12689#bib.bib51); Chanet al\.,[2022](https://arxiv.org/html/2606.12689#bib.bib52); Heimersheim and Nanda,[2024](https://arxiv.org/html/2606.12689#bib.bib53); Moneaet al\.,[2024](https://arxiv.org/html/2606.12689#bib.bib63)\)to test whether latent thoughts affect model behavior\.111recognizing causal interventions can still admit multiple compatible explanations\(Mélouxet al\.,[2025b](https://arxiv.org/html/2606.12689#bib.bib54)\)\.

### 2\.3Mechanistic Analyses of Latent Reasoning

#### Observable Latent Patterns\.

Prior work finds structured intermediates in latent states\. In CODI, logit\-lens, attention\-mass, and activation\-patching analyses are used to argue that models use latent thoughts as acomputational scratchpadfor reasoning: at each step, these states store operands and intermediate values for subsequent steps\(Shenet al\.,[2025](https://arxiv.org/html/2606.12689#bib.bib9); Cywinskiet al\.,[2025](https://arxiv.org/html/2606.12689#bib.bib20); Peterset al\.,[2025](https://arxiv.org/html/2606.12689#bib.bib21)\); however, the evidence remains largely single\-task and observational\.Haoet al\.\([2025](https://arxiv.org/html/2606.12689#bib.bib1)\)interpret Coconut’s latent thoughts as superposed reasoning paths with BFS\-like dynamics, where successive latent steps represent multiple candidate nodes at increasing graph depths and progressively concentrate mass on target\-reaching nodes\. Later work attributes similar patterns to shortcut behavior and task\-specific heuristics\(Cuiet al\.,[2026](https://arxiv.org/html/2606.12689#bib.bib22); Rizvi\-Martelet al\.,[2026](https://arxiv.org/html/2606.12689#bib.bib24)\)\. In this work, we investigate whether such observable patterns reflect genuine reasoning mechanisms\.

Shortcut Behavior\.Continuous thoughts may be decodable without driving reasoning: models can solve tasks via prompt or KV\-cache shortcuts, or commit to answers before latent computation\. Prior work frames thought use as binary, showing performance persisting under perturbations and ablations of latent states\(Zhanget al\.,[2025a](https://arxiv.org/html/2606.12689#bib.bib23); Cuiet al\.,[2026](https://arxiv.org/html/2606.12689#bib.bib22)\), suggesting stronger supervision as the cure\. However, such findings cannot establish whether models are simply incapable of using their thoughts or if they just find other circuits in their absence\. We therefore treat latent\-thought influence as a graded and localizable causal quantity, motivating the question of whether this influence is concentrated in specific directions or distributed\(Gur\-Ariet al\.,[2018](https://arxiv.org/html/2606.12689#bib.bib40); Bernaset al\.,[2026](https://arxiv.org/html/2606.12689#bib.bib39)\)\.

Geometry, Stability, and Latent Dynamics\.A complementary perspective on LLM interpretability studies reasoning as movement through the representation spaceElhageet al\.\([2022](https://arxiv.org/html/2606.12689#bib.bib38)\); Parket al\.\([2023](https://arxiv.org/html/2606.12689#bib.bib41)\); Gurnee and Tegmark \([2023](https://arxiv.org/html/2606.12689#bib.bib42)\); Gevaet al\.\([2023](https://arxiv.org/html/2606.12689#bib.bib43)\); Bhatiaet al\.\([2025](https://arxiv.org/html/2606.12689#bib.bib62),[2026](https://arxiv.org/html/2606.12689#bib.bib61)\); Bernaset al\.\([2026](https://arxiv.org/html/2606.12689#bib.bib39)\); Zhouet al\.\([2026](https://arxiv.org/html/2606.12689#bib.bib65)\)\. In particular,Wanget al\.\([2026](https://arxiv.org/html/2606.12689#bib.bib37)\)model explicit CoT as a Markov chain, with CoT being useful only when transitions are stable and aligned\. However, such analyses remain limited in LRM interpretability\.Zhuet al\.\([2025](https://arxiv.org/html/2606.12689#bib.bib25)\)compare the geometry of latent thoughts with optimal node embeddings, suggesting a superpositional graph\-search representation\.Weiet al\.\([2026](https://arxiv.org/html/2606.12689#bib.bib35)\)report that scaled implicit reasoning can become unstable, with representations homogenizing and drifting from the vocabulary manifold\. These dynamical studies show how latent states evolve, not whether the changing directions causally affect the answer, motivating pairing geometry with intervention\.

## 3Experimental Setup

Datasets\.We evaluate graph\-hopping on ProsQAHaoet al\.\([2025](https://arxiv.org/html/2606.12689#bib.bib1)\)and arithmetic\-reasoning on GSM8kCobbeet al\.\([2021](https://arxiv.org/html/2606.12689#bib.bib8)\)\.Models and Controls\.All models are trained separately for each task from pretrained GPT\-2 smallRadfordet al\.\([2019](https://arxiv.org/html/2606.12689#bib.bib11)\)\. Our design compares target LRM models withdifferent curriculabut thesame recurrence mechanism\(Coconut \(C\)andCODI\), and the following controls:

- ∙\\bulletRecurrence control:Pause\-as\-thought \(PaT\)follows Coconut’s format and curriculum but replaces recurrence with learned thought tokens\.
- ∙\\bulletCurriculum control:Coconutu\(Cu\), a curriculum\-perturbed Coconut variant that samples other stages with probabilityu=0\.3u=0\.3\.
- ∙\\bulletObservational controls:Base GPT\-2 \(B\)andExplicit\-CoT GPT\-2 \(CoT\)are included in §[4](https://arxiv.org/html/2606.12689#S4)to test whether similar patterns can be recovered without latent\-reasoning training, either from the probe itself or from ordinary task\-solving\. They are excluded from causal and geometric analyses, which require dedicated latent thought positions\.

Details on models and training are in Appendix[C](https://arxiv.org/html/2606.12689#A3)\.Statistical Evaluation\.We report 95% bootstraped confidence intervals \(1,000 resamples\) and use McNemarpp\-values for paired comparisons\. Details and full significance results are in Appendix[B](https://arxiv.org/html/2606.12689#A2)\.

![Refer to caption](https://arxiv.org/html/2606.12689v1/x1.png)Figure 1:Probing for breadth\-first\-search \(BFS\) patterns on graph\-hopping \(top\) and scratchpad\-thinking on arithmetic\-reasoning \(bottom\)\.Takeaway:Matched controls reproduce or invert observable patterns, showing they are not specific to the proposed recurrence or curriculum mechanisms and, thus, are insufficient alone for mechanistic attribution\.
## 4Observable Structures Are Insufficient for Mechanistic Attribution

We revisit two LRM interpretability patterns from prior work, discussed in §[2\.3](https://arxiv.org/html/2606.12689#S2.SS3): Coconut’s apparent simultaneous encoding of multiple candidate reasoning paths \(superposition\) and progressive refinement toward the correct path \(breadth\-first\-search exploration\) on graph\-tasks, and CODI’s claimed reasoning through decodable intermediate computations \(scratchpad\) on arithmetics\. To address RQ1, we compare models to non\-LRM model and curriculum controls introduced in §[3](https://arxiv.org/html/2606.12689#S3)\.

### 4\.1Case Study 1: Superposition and BFS\-like search on Graph\-Hopping

FollowingHaoet al\.\([2025](https://arxiv.org/html/2606.12689#bib.bib1)\), we probe the*depth frontier*afterkklatent thoughts by measuring the unnormalized joint probabilityp​\(concept\)=∏ip​\(toki∣ctx,“Every”,tok<i\)p\(\\text\{concept\}\)=\\prod\_\{i\}p\(\\text\{tok\}\_\{i\}\\mid\\text\{ctx\},\\ \\text\{\`\`Every''\},\\ \\text\{tok\}\_\{<i\}\)over candidate nodes at depthmax⁡\(k,1\)\\max\(k,1\)from the root: children fork=1k\{=\}1; grandchildren fork=2k\{=\}2\.Haoet al\.\([2025](https://arxiv.org/html/2606.12689#bib.bib1)\)read this as two patterns\.Superposition: early latent thoughts hold multiple candidates simultaneously \(high entropy; candidate mass spread across nodes\)\.BFS: increasing mass on target\-reaching nodes with recurrence \(risingP​\(c​o​r​r​e​c​t\)P\(correct\); falling entropy\)\. If recurrent feedback drives this BFS\-like exploration, C should show this pattern more strongly than PaT\. Results in Figure[1](https://arxiv.org/html/2606.12689#S3.F1)\(top row\)\.

B and CoT, lacking latent training, never form the frontier \(B near\-zero mass; CoT concentrates mass early but degrades with depth\)\. While C reproduces the BFS signature \(risingP​\(c​o​r​r​e​c​t\)P\(correct\), falling entropy\), PaT matches it with no recurrence \(but same curriculum\)\. Moreover, Cuwhich perturbs C’s curriculum inverts this pattern: high mass at shallow depth thatdegradeswith depth\. CODI, which lacks the curriculum, weakly mirrors C and PaT\. Candidate mass tracksP​\(c​o​r​r​e​c​t\)P\(correct\)closely, indicating that frontier mass concentrates on target\-reaching nodes\. However, superposition requires competing incorrect candidates held simultaneously\. Thus, BFS signature is not recurrence\-specific: PaT matches it without recurrence, while Cuchanges it with recurrence preserved\.

### 4\.2Case Study 2: Scratchpad Thinking on Arithmetic\-Reasoning via Logit\-Lens

FollowingShenet al\.\([2025](https://arxiv.org/html/2606.12689#bib.bib9)\), we project final\-layer hidden states at thought positions through the LM head and compare decoded tokens against ground\-truth intermediate CoT annotations\.Shenet al\.\([2025](https://arxiv.org/html/2606.12689#bib.bib9)\)read two patterns as evidence of latent scratchpad\.Hit Rate: decoded tokens match a ground\-truth CoT intermediate\.Step Alignment\(ordered\-scratchpad\): thett\-th thought matches thett\-th CoT step\. We additionally test Superposition Rate \(fraction of steps decoding≥\\geq2 distinct intermediates\) to test generalizability ofHaoet al\.\([2025](https://arxiv.org/html/2606.12689#bib.bib1)\)claims222We also track top\-1 decoded token changes across steps to assess trajectory structure on both tasks \(Appendix[A](https://arxiv.org/html/2606.12689#A1)\)\.\. Results in Figure[1](https://arxiv.org/html/2606.12689#S3.F1)\(bottom row\)\.

PaT lacks recurrence but exceeds CODI on hit\-rate overall\. C, CoT \(no latent training\) and Cualso reproduce CODI\-like step\-wise alternation, although with different magnitudes and phase structure\. This suggests the pattern is not specific to CODI’s latent\-distillation objective: it could reflect CoT fine\-tuning, probe\-setup, or another source, but its appearance in controls already contradicts that attribution\. Moreover, step\-alignment decays sharply with depth for every model, contradicting an ordered scratchpad\. B is near zero throughout\. Superposition stays low across all models, failing to reproduce COCONUT’s frontier in arithmetic\. Overall, the logit\-lens readouts recover scratchpad\-like signatures without uniquely tracking performance or requiring CODI’s distillation mechanism\.

#### Summary\. Both case studies reproduce the original reported findings but challenge the conclusions\. Similar patterns in non\-LRM controls indicating that the observed patterns are not specific to LRMs\. Going beyond observable patterns requires a causal perspective\.

## 5When and How do LRMs use Latent Thoughts?

Having shown that observational readouts are not specific enough for mechanistic attribution, we now investigate how latent thoughts causally influence model behavior\.

![Refer to caption](https://arxiv.org/html/2606.12689v1/x2.png)Figure 2:Per\-layerIEKL\\mathrm\{IE\}\_\{\\mathrm\{KL\}\}on the residual stream under partner\-prompt corruption across buckets\{Pfull,Pmax,Pb,T1,…,TK,Tfull,Ab\}\\\{P\_\{\\text\{full\}\},P\_\{\\max\},P\_\{b\},T\_\{1\},\\dots,T\_\{K\},T\_\{\\text\{full\}\},A\_\{b\}\\\}\. Prompt positions recover best across models and tasks\. Thought positionsTtT\_\{t\}contribute near zero on graph\-hopping; on arithmetic\-reasoning,TfullT\_\{\\text\{full\}\}yields the strongest recovery for C,CuC\_\{u\}, and CODI\.Takeaway:Latent\-thought influence is task\-varying—when large, it can override corrupted prompt positions and steer models toward correct paths\.ModelGraph\-HoppingThoughtsRemovedArithmetic\-ReasoningThoughtsRemovedB2\.4±1\.32\.4\{\\scriptstyle\\pm 1\.3\}–1\.4±0\.61\.4\{\\scriptstyle\\pm 0\.6\}–CoT79\.0±3\.679\.0\{\\scriptstyle\\pm 3\.6\}–41\.9±2\.741\.9\{\\scriptstyle\\pm 2\.7\}–PaT95\.4±1\.895\.4\{\\scriptstyle\\pm 1\.8\}95\.6±1\.795\.6\{\\scriptstyle\\pm 1\.7\}26\.4±2\.426\.4\{\\scriptstyle\\pm 2\.4\}21\.4±2\.221\.4\{\\scriptstyle\\pm 2\.2\}C98\.0±1\.398\.0\{\\scriptstyle\\pm 1\.3\}97\.8±1\.397\.8\{\\scriptstyle\\pm 1\.3\}35\.7±2\.535\.7\{\\scriptstyle\\pm 2\.5\}7\.7±1\.47\.7\{\\scriptstyle\\pm 1\.4\}Cu96\.0±1\.796\.0\{\\scriptstyle\\pm 1\.7\}91\.8±2\.491\.8\{\\scriptstyle\\pm 2\.4\}30\.8±2\.530\.8\{\\scriptstyle\\pm 2\.5\}39\.1±2\.639\.1\{\\scriptstyle\\pm 2\.6\}CODI80\.0±3\.580\.0\{\\scriptstyle\\pm 3\.5\}80\.0±3\.580\.0\{\\scriptstyle\\pm 3\.5\}41\.8±2\.741\.8\{\\scriptstyle\\pm 2\.7\}25\.6±2\.425\.6\{\\scriptstyle\\pm 2\.4\}Table 1:Accuracy under thought\-ablation at test\-time\. Performance degrades on arithmetic\-reasoning; graph\-hopping is unaffected\.Takeaway:Influence of latent thoughts on model performance is task\-varying\.#### Latent Thought Ablation at Test\-Time\.

First, we evaluate whether models require latent thoughts at all by removing them at test\-time:K:=Kmax→0K\{:=\}K\_\{\\max\}\\\!\\to\\\!0\(recurrence skipped forC,Cu,CODIC,C\_\{u\},\\mathrm\{CODI\}; thekkparallel tokens dropped for PaT\)\. Results in Table[1](https://arxiv.org/html/2606.12689#S5.T1)\.

Ongraph\-hopping, onlyCuC\_\{u\}is affected meaningfully\. Onarithmetic\-reasoning,CuC\_\{u\}improvesbut other models suffer major drops\. Causalnecessityof latent thoughts is task\-dependent, but the shortcut\-learning dichotomy \(used or bypassed\) warrants further checks due to Cu’s behavior\.

#### Per\-Position Causal Tracing\.

If performance is maintained under removed thoughts, does it necessarily imply that, when present, thoughts are not causally influencing the computation? LLMs are known to be causally over\-determinedMcGrathet al\.\([2023](https://arxiv.org/html/2606.12689#bib.bib60)\); Mélouxet al\.\([2025a](https://arxiv.org/html/2606.12689#bib.bib14)\): multiple circuits computing the same behaviorLanet al\.\([2024](https://arxiv.org/html/2606.12689#bib.bib59)\); when present, thoughts are still part of the computational paths and we can expect them to have a causal impact on the output\. To test this, we extendMenget al\.\([2022](https://arxiv.org/html/2606.12689#bib.bib58)\)’s causal tracing methodology to latent thoughts\.

Define sites=\(ℓ,p\)s=\(\\ell,p\)as a layerℓ\\elland positionppin the residual streamElhageet al\.\([2021](https://arxiv.org/html/2606.12689#bib.bib68)\)333Per component \(attention, MLP\) decompositions appear in Appendix[A](https://arxiv.org/html/2606.12689#A1)\.\. We run a*clean*pass, a*corrupted*pass \(cp\) on a partner prompt \(dataset instance with different answer\), and a*patched*pass \(pp\) that injects sitess’s clean activation into the corrupted forward pass\. At the answer boundaryAbA\_\{b\}, we greedily decodeNNtokens, recording at each stepjjthe full next\-token distributionP∙\(j\)=softmax​\(zj\)P^\{\(j\)\}\_\{\\bullet\}=\\text\{softmax\}\(z\_\{j\}\)for each pass\. We compute the forward KL from the clean distributionKL​\(Pclean\(j\)∥P∙\(j\)\)\\mathrm\{KL\}\\big\(P^\{\(j\)\}\_\{\\mathrm\{clean\}\}\\\|P^\{\(j\)\}\_\{\\bullet\}\\big\), and average over the content windowW=\[nfmt,e\)W=\[n\_\{\\mathrm\{fmt\}\},e\)spanning the clean pass output, wherenfmtn\_\{\\mathrm\{fmt\}\}skips the fixed format prefix \(e\.g\.,\#\#\#,The answer is:\), andeeis the first end\-of\-text step \(orNN\)\. Patching thus measures recovery of the model’s clean behavior:

KL¯∙=1\|W\|​∑j∈WKL​\(Pclean\(j\)∥P∙\(j\)\),IE​\(s\)=1−KL¯pp/KL¯cp\\displaystyle\\overline\{\\mathrm\{KL\}\}\_\{\\bullet\}=\\tfrac\{1\}\{\|W\|\}\\sum\_\{j\\in W\}\\mathrm\{KL\}\\big\(P^\{\(j\)\}\_\{\\mathrm\{clean\}\}\\\|P^\{\(j\)\}\_\{\\bullet\}\\big\),\\quad\\mathrm\{IE\}\(s\)=1\-\\overline\{\\mathrm\{KL\}\}\_\{\\mathrm\{pp\}\}/\\overline\{\\mathrm\{KL\}\}\_\{\\mathrm\{cp\}\}

The indirect effectIE​\(s\)≤1\\mathrm\{IE\}\(s\)\\leq 1measures how far the patch restores the clean output:11is full recovery,0none, and<0<0is worsening\. We report thearg⁡maxℓ\\arg\\max\_\{\\ell\}\-IE\(s\)\(s\)over the position bucket\{Pfull,Pmax,Pb,T1,…,TK,Tfull,Ab\}\\\{P\_\{\\mathrm\{full\}\},P\_\{\\max\},P\_\{b\},T\_\{1\},\\dots,T\_\{K\},T\_\{\\mathrm\{full\}\},A\_\{b\}\\\}, whereTtT\_\{t\}is thoughttt;Pb,AbP\_\{b\},A\_\{b\}the latent\-block delimiters;PmaxP\_\{\\max\}the strongest single prompt token; andPfull/TfullP\_\{\\mathrm\{full\}\}/T\_\{\\mathrm\{full\}\}the jointly patched prompt & thought positions\. Results in Figure[2](https://arxiv.org/html/2606.12689#S5.F2)\.

![Refer to caption](https://arxiv.org/html/2606.12689v1/x3.png)Figure 3:Gradient\-subspace intervention flip rates \(%\) across amplification strengthα\\alphaon the gradient\-derived vs random control directions\.Graph\-Hopping:Robust to ablation \(α=0\\alpha=0\) but flips at high\-α\\alphaon GRAD\.Arithmetic\-Reasoning:Degradation under GRAD ablation; GRAD\>\>Rand flip\-rates at lowα\\alpha\.Takeaway:Thought utilization isgradedand depends on their influence power over model behavior\.Prompt positions \(especiallyPfullP\_\{\\text\{full\}\}\) strongly recover performance across models ongraph\-hopping\. The thoughts sit at zero recoverability across all models, even Cuwhich shows small thoughtnecessityin the thought ablation experiment\. This is possibly a result of its small \(4\.8%\) thought necessity being averaged over the full dataset\. Onarithmetic\-reasoning, prompt\-positions of PaT and CODI still show very high recoverability throughout layers\. Individual thought positions also show recoverability rates albeit small\. Interestingly, patching all thought positions while other positions are corrupted shows the best recoverability through all layers on C, Cuand CODI and through the later layers on PaT\.

#### Gradient\-Subspace Interventions\.

The previous analyses test whether or not latent thoughts influence model behavior at test\-time\. Prior work typically frames this influence as binary: unused thoughts are simply bypassed shortcutsZhanget al\.\([2025a](https://arxiv.org/html/2606.12689#bib.bib23)\); Cuiet al\.\([2026](https://arxiv.org/html/2606.12689#bib.bib22)\)\. Butwhy? Do these thoughts lack the capacity to influence behavior, or is their effect simply too small to measure globally? To test this, we isolate the representation directions most sensitive to the loss allowing us to geometrically concentrate thought influence to a subspace and intervene right there\.

For thoughthi,t∈ℝDh\_\{i,t\}\\in\\mathbb\{R\}^\{D\}\(instanceii, positiontt\), lossℒi\\mathcal\{L\}\_\{i\}and gradientsgi,t=∇hi,tℒig\_\{i,t\}=\\nabla\_\{h\_\{i,t\}\}\\mathcal\{L\}\_\{i\}, we define the gradient\-subspaceBt∈ℝD×ktB\_\{t\}\\in\\mathbb\{R\}^\{D\\times k\_\{t\}\}using the top\-ktk\_\{t\}right singular vectors of the gradient matrixGt=\[g1,t;…;gN,t\]G\_\{t\}=\[g\_\{1,t\};\\dots;g\_\{N,t\}\], capturing99%99\\%cumulative energy\. To intervene, we splithi,th\_\{i,t\}ashi,tB=Bt​Bt⊤​hi,th\_\{i,t\}^\{B\}=B\_\{t\}B\_\{t\}^\{\\top\}h\_\{i,t\}\(projection ontoBtB\_\{t\}\) andhi,t⟂=hi,t−hi,tBh\_\{i,t\}^\{\\perp\}=h\_\{i,t\}\-h\_\{i,t\}^\{B\}\. The projection is scaled byα∈\{0,0\.5,1,1\.5,2,5,10,25,50,100\}\\alpha\\in\\\{0,0\.5,1,1\.5,2,5,10,25,50,100\\\}via the updatehi,t←hi,t⟂\+α​hi,tB=hi,t\+\(α−1\)​Bt​Bt⊤​hi,th\_\{i,t\}\\leftarrow h\_\{i,t\}^\{\\perp\}\+\\alpha h\_\{i,t\}^\{B\}=h\_\{i,t\}\+\(\\alpha\-1\)B\_\{t\}B\_\{t\}^\{\\top\}h\_\{i,t\}which ablates \(α=0\\alpha=0\) or amplifies \(α\>1\\alpha\>1\) the targeted component\. A rank\-matched orthonormal basisBtrandB\_\{t\}^\{\\mathrm\{rand\}\}via QR decomposition on Gaussian noise serves as the control\. Except for PaT, models skipt=Kt=Kas the gold answer token sits at stepKK\(meaningGKG\_\{K\}in empty\)\. Results in Figure[3](https://arxiv.org/html/2606.12689#S5.F3)\.

Ongraph\-hopping, ablation leaves accuracy intact, but PaT and C show notable flip\-rates at high\-α\\alphavalues that significantly exceed the random control\. This suggests that thoughts are not strictlybypassed, but retaingradedcausal effect over model behavior capable of steering predictions\. Cuis an exception: despite degrading under thought ablation, the gradient directions fail to localize its influence to a concentrated subspace\. Onarithmetic\-reasoning, ablation degrades all models but the random control stays inert\. GRAD flips exceed Rand only at very smallα\\alpha, demonstrating a highly influential subspace444Appendix[A](https://arxiv.org/html/2606.12689#A1)\(Table[3](https://arxiv.org/html/2606.12689#A1.T3)\) reports the dimensionality of the gradient\-derived subspaces\.\.

#### Summary\. Thought utilization isgradedrather than binary, and depends on how much influence the latent thoughts carry on model behavior\. Together with §[4](https://arxiv.org/html/2606.12689#S4), this motivates moving from finding low\-level observable latent patterns to studying the dynamics of behavior\-influential directions across latent steps\.

## 6The Dynamics and Geometry of Latent Thoughts

Observable patterns such as BFS and scratchpad\-thinking \(§[4](https://arxiv.org/html/2606.12689#S4)\) are not sufficient, on their own, to establish causal contribution to model behavior\. In this section, we argue that the dynamic evolution of latent thoughts \(and the directions where their causal influence concentrates\) can offer more meaningful insights into the structure of latent reasoning\. Using the behavior\-influential directions identified in §[5](https://arxiv.org/html/2606.12689#S5), we now apply geometric tools to distinguish the dynamics of highly causally influential latent thoughts from less causally influential ones\.

![Refer to caption](https://arxiv.org/html/2606.12689v1/x4.png)Figure 4:Markovianity of latent thoughts\.Graph\-Hopping:Static dynamics \(except Cu\);Arithmetic\-Reasoning:Evolving dynamics\.Takeaway:Even static appearing full\-thoughts show evolving dynamics within the gradient\-subspace where influence concentrates\.#### Markovianity of Thought Trajectories\.

We first investigate the evolution of thought trajectories, testing whether future states can be predicted from preceding ones\. By design, the transformer architecture makes every thought a function of all previous thoughts and prompt tokens, but the trajectories may encode a simpler approximate structure\.

For the thoughtshi,1,…,hi,Kh\_\{i,1\},\\ldots,h\_\{i,K\}for every test instanceii, we form pairsXi,t=\[hi,t−1;…;hi,t−n\]X\_\{i,t\}=\[h\_\{i,t\-1\};\\ldots;h\_\{i,t\-n\}\]andYi,t=hi,tY\_\{i,t\}=h\_\{i,t\}for a Markov ordernn\. We fit a shared mapf:ℝn​D→ℝDf:\\mathbb\{R\}^\{nD\}\\\!\\to\\\!\\mathbb\{R\}^\{D\}on the train split and report uniform\-averageR2R^\{2\}on the test split under two function classes: ridge\-regularized linear regression and a two\-layer MLP :ht=W2​σ​\(W1​X\+b1\)\+b2h\_\{t\}=W\_\{2\}\\,\\sigma\(W\_\{1\}X\+b\_\{1\}\)\+b\_\{2\}withσ=GELU\\sigma\{=\}\\mathrm\{GELU\},W1∈ℝ256×n​DW\_\{1\}\\in\\mathbb\{R\}^\{256\\times nD\}\(early\-stopped on a 10% train slice,R2R^\{2\}averaged over 3 seeds\)\. We compare against the*mean baseline*h^t=h¯train\\hat\{h\}\_\{t\}=\\bar\{h\}\_\{\\text\{train\}\}\(R2=0R^\{2\}=0on train split\), and the*identity baseline*h^t=ht−1\\hat\{h\}\_\{t\}=h\_\{t\-1\}\. Ordern=1n=1is the strict first\-order claimht=f​\(ht−1\)h\_\{t\}=f\(h\_\{t\-1\}\); we sweepn∈\{1,…,5\}n\\in\\\{1,\\ldots,5\\\}\. The same analysis on the projected thoughtshi,tB=Bt​Bt⊤​hi,th^\{B\}\_\{i,t\}=B\_\{t\}B\_\{t\}^\{\\top\}h\_\{i,t\}tests how the gradient\-subspace \(§[5](https://arxiv.org/html/2606.12689#S5)\) dynamics differ from the full\-dimensional thoughts\. Results in Figure[4](https://arxiv.org/html/2606.12689#S6.F4)\.

Ongraph\-hopping, the identity baseline dominates for PaT, C, and CODI: no fitted transition meaningfully improves over simply copying the previous thought, but collapses onCu\\mathrm\{C\}\_\{u\}for which linear and MLP are the best approximates\. Onarithmetic\-reasoning, identity dominates on PaT but collapses on C, Cu, and CODI, all best characterized by the MLP while linear closely follows\. The dynamics change in the gradient\-subspace; the MLP map nearly captures the fullgraph\-hoppingdynamics and dominates onarithmetic\-reasoning\.

#### Geometric Stability of Gradient\-Subspaces\.

Next, we ask whether the gradient\-subspace \(§[5](https://arxiv.org/html/2606.12689#S5)\) itself remains stable throughout the reasoning trajectory\. For adjacent time\-steps, we measure the mean squared principal\-angle similarity between subspaces:st=mean​\(σ​\(Bt⊤​Bt\+1\)2\),s\_\{t\}=\\mathrm\{mean\}\\\!\\left\(\\sigma\(B\_\{t\}^\{\\top\}B\_\{t\+1\}\)^\{2\}\\right\),whereσ​\(⋅\)\\sigma\(\\cdot\)denotes the singular values of the inter\-basis projection matrix\. Values ofst→1s\_\{t\}\\to 1indicate stable causal subspaces, whilest→0s\_\{t\}\\to 0indicate orthogonality\.

![Refer to caption](https://arxiv.org/html/2606.12689v1/x5.png)Figure 5:Geometric stability of the gradient\-subspaces\.Graph\-Hopping:Stable subspaces \(except Cu\);Arithmetic\-Reasoning:Rotating subspaces\.Takeaway:Thoughts with weak influence over model behavior exhibit highly stable gradient\-subspace geometries\. Points plotted at half\-integers as each measures an adjacent\-step paircos2⁡\(Bt,Bt\+1\)\\cos^\{2\}\(B\_\{t\},B\_\{t\+1\}\)\.Figure[5](https://arxiv.org/html/2606.12689#S6.F5)further localizes this distinction to the geometry of the gradient\-subspace\. Ongraph\-hopping, PaT, C and CODI maintain highly stable subspaces whileCuC\_\{u\}shows near\-orthogonal early rotations before stabilizing\. Onarithmetic\-reasoning, stability is lower overall: PaT remains most stable, C andCuC\_\{u\}show moderate alignment, and CODI undergoes large step\-to\-step rotations\.

#### Summary\. Beyond their causal necessity and sufficiency, the causally influential directions in latent thoughts are consistently low\-rank, and theirdynamicstrack influence\. Where causal effect is weak \(e\.g, graph\-hopping\), full trajectories barely evolve past the first thought, though the causally active subspace supports non\-trivial computation; where stronger, trajectory dynamics are more structured, and the causally active subspace supports more diverse computation\.

## 7Discussions

A causal\-first view of latent thoughts\.LRM interpretability is especially vulnerable to false explanations: decodable intermediates, attention mass, and visually salient geometries can appear even when the proposed mechanism is absent\. Observational tools are most natural when computation is partly exposed in token space; in LRMs, they can recover structure from continuous states that need not coincide with what the model uses to answer\. This makes an observational\-first workflow risky\. We therefore reverse the order of analysis: rather than starting from a readable latent pattern, we first test whether latent thoughts causally affect behavior and where this influence is concentrated\. Only then do we analyze the geometry and dynamics of these behavior\-influential regions\. Geometry is informative when it describes the part of latent computation that affects predictions\.

From thought use to causal localization\.This causal\-first view helps reframe the shortcut\-vs\-reasoning dichotomy\. The relevant question is not whether latent thoughts are simply “used” or “bypassed”, but how much they influence behavior and in which directions of the representation\. This influence can vary across tasks with different complexity and structure\. Moreover, a model may solve a task without requiring latent thoughts under ablation while still containing latent directions that can steer its predictions when intervened on\. Such cases do not show that the model is incapable of using latent thoughts; they may instead reflect alternative circuits that solve the task when latent computation is removed\. Conversely, forcing stronger dependence on latent thoughts \(for instance by stronger supervision\) does not by itself show that these thoughts implement faithful reasoning\. The more informative unit of analysis is therefore the strength and localization of latent\-thought influence for a given model and task\.

Implications for LRM interpretation and design\.Causal localization also gives interpretability a constructive role\. Once behavior\-influential regions are identified, they provide concrete targets for geometric analysis, monitoring, and downstream interventions\. For example, if influence is consistently concentrated in low\-rank, loss\-sensitive subspaces, compression could focus on preserving the task\-relevant component of the latent state rather than the full representation\. Similarly, if behavior\-influential directions are shown to follow stable or predictable step\-to\-step dynamics, they could inform simplified transition models or more targeted recurrent computation, which is particularly relevant because the recurrent mechanism described in §[2\.1](https://arxiv.org/html/2606.12689#S2.SS1)can impose sequential dependencies that are difficult to parallelize during training\.

Relation to CoT faithfulness concerns\.Our findings extend documented concerns about explicit CoT: verbalized rationales can misrepresent a model’s internal reasonsTurpinet al\.\([2023](https://arxiv.org/html/2606.12689#bib.bib32)\); Arcuschinet al\.\([2025](https://arxiv.org/html/2606.12689#bib.bib55)\), and intermediate steps \(including “aha” or self\-verification moments\) may exert little causal influence on the final answerBoppanaet al\.\([2026](https://arxiv.org/html/2606.12689#bib.bib56)\); Zhaoet al\.\([2025](https://arxiv.org/html/2606.12689#bib.bib57)\)\. LRMs inherit this concern while removing the surface artifact that makes CoT at least partially inspectable\. The risk is therefore not only opacity, but false interpretability: latent states may contain decodable or geometrically structured information that appears explanatory without being behaviorally relevant, strengthening the case for causal\-first monitoring\.

Overall, latent thoughts should be treated as hidden computation, not hidden explanation\. Instead of readable latent patterns, LRM interpretability should target the regions of latent computation that causally shape behavior and analyze their dynamics\. This causal\-first view provides a more grounded and principled basis for designing and auditing latent reasoning systems\.

## Limitations

First, we estimate the gradient subspace linearly via SVD; a nonlinear estimator \(e\.g\., autoencoder\) could capture causal structure a linear basis flattens\. Second, our causal evidence relies on local intervention\-based methods with known caveats: causal tracing and gradient\-subspace interventions probe behavior under specific perturbations that can shift the model off\-distribution, and interventions admit multiple compatible mechanistic explanations, so “behavior\-influential” is a necessary\-but\-not\-sufficient signal of mechanism,notproof\. Third, our conclusions draw on relatively small models at modest latent steps \(K=6\), and our LRM coverage is partial \(PaT, COCONUT, CODI\); future work should extend to larger model families and more recent reasoning paradigms\. Lastly, the task\-varying nature of thought utilization makes LRMs a natural setting for model diffing—comparing checkpoints across curricula or training stages to localize where behavior\-influential structure first emerges—which our per\-instance, per\-timestep analyses do not address: we study trained checkpoints only, not how the causal subspace forms over training\.

## Ethical Considerations and Societal Impact

Latent reasoning models may have substantial societal impact if they make multi\-step reasoning more efficient, scalable, and easier to integrate into agentic systems\. Yet by replacing explicit chain\-of\-thought traces with continuous hidden\-state computation, LRMs remove a surface artifact that can be inspected, filtered, or audited\. This creates two related risks\. First, without reliable tools for interpreting LRMs, users, developers, and auditors may have fewer signals for detecting harmful, deceptive, or misaligned behavior, especially in high\-stakes domains such as education, healthcare triage, legal assistance, scientific automation, or autonomous agents\. Second, unreliable interpretability tools can create false interpretations, which may lead to monitor the wrong features and miss the computations that actually drive harmful or misaligned behavior\. Our work contributes towards addressing this risk by arguing that latent thoughts should be treated as hidden computational states, not hidden explanations, and that mechanistic claims require controls and causal tests before latent structure is interpreted\. This causal\-first view is not a bullet\-proof procedure that solves this completely, since interventions can still be local, distribution\-shifting, or compatible with multiple explanations\. Rather, it provides a stricter evidential standard that should be combined with stress testing, adversarial evaluation, uncertainty reporting, and domain\-specific oversight\.

## Acknowledgements

This work was conducted within the French research unit UMR 5217 and was supported by CNRS \(grant ANR\-22\-CPJ2\-0036\-01 and ANR\-25\-CE23\-2059\-01\) and by MIAI@Grenoble\-Alpes \(grant ANR\-19\-P3IA\-0003 and ANR\-23\-IACL\-0006\)\. Thomas’ research is also partially supported by the French Ministry of Higher Education, Research and Innovation under CIFRE PhD Convention“Reconcile Efficiency and Interpretability in Structured Planning and Reasoning”\(no\.2025/0487\)\.

## References

- Chain\-of\-thought reasoning in the wild is not always faithful\.arXiv preprint arXiv:2503\.08679\.Cited by:[§7](https://arxiv.org/html/2606.12689#S7.p4.1)\.
- Y\. Belinkov \(2022\)Probing classifiers: promises, shortcomings, and advances\.Computational Linguistics48\(1\),pp\. 207–219\.External Links:[Link](https://aclanthology.org/2022.cl-1.7/),[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00422)Cited by:[§2\.2](https://arxiv.org/html/2606.12689#S2.SS2.p1.1)\.
- L\. Bereska and E\. Gavves \(2024\)Mechanistic interpretability for ai safety–a review\.arXiv preprint arXiv:2404\.14082\.Cited by:[§1](https://arxiv.org/html/2606.12689#S1.p1.1),[§2\.2](https://arxiv.org/html/2606.12689#S2.SS2.p1.1)\.
- R\. Bernas, F\. Jourdan, A\. Poché, and C\. Hudelot \(2026\)Revisiting anisotropy in language transformers: the geometry of learning dynamics\.arXiv preprint arXiv:2604\.08764\.Cited by:[§2\.3](https://arxiv.org/html/2606.12689#S2.SS3.SSS0.Px1.p2.1),[§2\.3](https://arxiv.org/html/2606.12689#S2.SS3.SSS0.Px1.p3.1)\.
- G\. Bhatia, A\. M\. Isa, M\. Peyrard, and W\. Zhao \(2026\)What really controls temporal reasoning in large language models: tokenisation or representation of time?\.External Links:2603\.19017,[Link](https://arxiv.org/abs/2603.19017)Cited by:[§2\.3](https://arxiv.org/html/2606.12689#S2.SS3.SSS0.Px1.p3.1)\.
- G\. Bhatia, M\. Peyrard, and W\. Zhao \(2025\)Date fragments: a hidden bottleneck of tokenization for temporal reasoning\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 3201–3219\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.159/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.159),ISBN 979\-8\-89176\-332\-6Cited by:[§2\.3](https://arxiv.org/html/2606.12689#S2.SS3.SSS0.Px1.p3.1)\.
- S\. Boppana, A\. Ma, M\. Loeffler, R\. Sarfati, E\. Bigelow, A\. Geiger, O\. Lewis, and J\. Merullo \(2026\)Reasoning theater: disentangling model beliefs from chain\-of\-thought\.arXiv preprint arXiv:2603\.05488\.Cited by:[§7](https://arxiv.org/html/2606.12689#S7.p4.1)\.
- A\. Chan, R\. Salganik, A\. Markelius, C\. Pang, N\. Rajkumar, D\. Krasheninnikov, L\. Langosco, Z\. He, Y\. Duan, M\. Carroll,et al\.\(2023\)Harms from increasingly agentic algorithmic systems\.InProceedings of the 2023 ACM conference on fairness, accountability, and transparency,pp\. 651–666\.Cited by:[§1](https://arxiv.org/html/2606.12689#S1.p1.1)\.
- L\. Chan, A\. Garriga\-Alonso, N\. Goldwosky\-Dill, R\. Greenblatt, J\. Nitishinskaya, A\. Radhakrishnan, B\. Shlegeris, and N\. Thomas \(2022\)Causal scrubbing, a method for rigorously testing interpretability hypotheses\.AI Alignment Forum\.External Links:[Link](https://www.alignmentforum.org/posts/JvZhhzycHu2Yd57RN)Cited by:[§2\.2](https://arxiv.org/html/2606.12689#S2.SS2.p1.1)\.
- Y\. Chen, J\. Benton, A\. Radhakrishnan, J\. Uesato, C\. Denison, J\. Schulman, A\. Somani, P\. Hase, M\. Wagner, F\. Roger,et al\.\(2025\)Reasoning models don’t always say what they think\.arXiv preprint arXiv:2505\.05410\.Cited by:[§1](https://arxiv.org/html/2606.12689#S1.p1.1)\.
- K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano,et al\.\(2021\)Training verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168\.Cited by:[§C\.1](https://arxiv.org/html/2606.12689#A3.SS1.SSS0.Px1.p3.3),[§3](https://arxiv.org/html/2606.12689#S3.p1.2)\.
- Y\. Cui, Z\. Dai, B\. He, Z\. Shi, H\. Liu, R\. Sun, Z\. Liu, Y\. Xing, J\. Tang, and B\. Dumoulin \(2026\)How Do Latent Reasoning Methods Perform Under Weak and Strong Supervision?\.InWorkshop on Latent & Implicit Thinking – Going Beyond CoT Reasoning,External Links:[Link](https://openreview.net/forum?id=F2KA5IkONu)Cited by:[§C\.1](https://arxiv.org/html/2606.12689#A3.SS1.SSS0.Px1.p2.3),[§C\.1](https://arxiv.org/html/2606.12689#A3.SS1.SSS0.Px1.p3.3),[§C\.3](https://arxiv.org/html/2606.12689#A3.SS3.p3.1),[§1](https://arxiv.org/html/2606.12689#S1.p3.1),[§2\.3](https://arxiv.org/html/2606.12689#S2.SS3.SSS0.Px1.p1.1),[§2\.3](https://arxiv.org/html/2606.12689#S2.SS3.SSS0.Px1.p2.1),[§5](https://arxiv.org/html/2606.12689#S5.SS0.SSS0.Px3.p1.1)\.
- B\. Cywinski, B\. Bussmann, A\. Conmy, J\. Engels, N\. Nanda, and S\. Rajamanoharan \(2025\)Can we interpret latent reasoning using current mechanistic interpretability tools?\.LessWrong\.External Links:[Link](https://www.lesswrong.com/posts/YGAimivLxycZcqRFR/)Cited by:[§1](https://arxiv.org/html/2606.12689#S1.p2.1),[§2\.3](https://arxiv.org/html/2606.12689#S2.SS3.SSS0.Px1.p1.1)\.
- M\. Dehghani, S\. Gouws, O\. Vinyals, J\. Uszkoreit, and L\. Kaiser \(2019\)Universal transformers\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=HyzdRiR9Y7)Cited by:[§2\.1](https://arxiv.org/html/2606.12689#S2.SS1.p1.4)\.
- Y\. Deng, K\. Prasad, R\. Fernandez, P\. Smolensky, V\. Chaudhary, and S\. Shieber \(2023\)Implicit chain of thought reasoning via knowledge distillation\.External Links:2311\.01460,[Link](https://arxiv.org/abs/2311.01460)Cited by:[§C\.1](https://arxiv.org/html/2606.12689#A3.SS1.SSS0.Px1.p3.3),[§C\.1](https://arxiv.org/html/2606.12689#A3.SS1.SSS0.Px2.p1.2),[§2\.1](https://arxiv.org/html/2606.12689#S2.SS1.p1.4)\.
- C\. Dilgren and S\. Wiegreffe \(2026\)Are Latent Reasoning Models Easily Interpretable?\.InWorkshop on Latent & Implicit Thinking – Going Beyond CoT Reasoning,External Links:[Link](https://openreview.net/forum?id=L4k8rbmwrr)Cited by:[§C\.1](https://arxiv.org/html/2606.12689#A3.SS1.SSS0.Px1.p3.3)\.
- Y\. Elazar, S\. Ravfogel, A\. Jacovi, and Y\. Goldberg \(2021\)Amnesic probing: behavioral explanation with amnesic counterfactuals\.Transactions of the Association for Computational Linguistics9,pp\. 160–175\.External Links:[Link](https://aclanthology.org/2021.tacl-1.10/),[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00359)Cited by:[§2\.2](https://arxiv.org/html/2606.12689#S2.SS2.p1.1)\.
- N\. Elhage, T\. Hume, C\. Olsson, N\. Schiefer, T\. Henighan, S\. Kravec, Z\. Hatfield\-Dodds, R\. Lasenby, D\. Drain, C\. Chen,et al\.\(2022\)Toy models of superposition\.arXiv preprint arXiv:2209\.10652\.Cited by:[§2\.3](https://arxiv.org/html/2606.12689#S2.SS3.SSS0.Px1.p3.1)\.
- N\. Elhage, N\. Nanda, C\. Olsson, T\. Henighan, N\. Joseph, B\. Mann, A\. Askell, Y\. Bai, A\. Chen, T\. Conerly,et al\.\(2021\)A mathematical framework for transformer circuits\.Transformer Circuits Thread1\(1\),pp\. 12\.Cited by:[§5](https://arxiv.org/html/2606.12689#S5.SS0.SSS0.Px2.p2.13)\.
- A\. Geiger, D\. Ibeling, A\. Zur, M\. Chaudhary, S\. Chauhan, J\. Huang, A\. Arora, Z\. Wu, N\. Goodman, C\. Potts, and T\. Icard \(2025\)Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability\.Journal of Machine Learning Research26\(83\),pp\. 1–64\.External Links:[Link](http://jmlr.org/papers/v26/23-0058.html)Cited by:[§2\.2](https://arxiv.org/html/2606.12689#S2.SS2.p1.1)\.
- M\. Geva, J\. Bastings, K\. Filippova, and A\. Globerson \(2023\)Dissecting recall of factual associations in auto\-regressive language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 12216–12235\.Cited by:[§2\.3](https://arxiv.org/html/2606.12689#S2.SS3.SSS0.Px1.p3.1)\.
- A\. Giannou, S\. Rajput, J\. Sohn, K\. Lee, J\. D\. Lee, and D\. Papailiopoulos \(2023\)Looped transformers as programmable computers\.InProceedings of the 40th International Conference on Machine Learning,A\. Krause, E\. Brunskill, K\. Cho, B\. Engelhardt, S\. Sabato, and J\. Scarlett \(Eds\.\),Proceedings of Machine Learning Research, Vol\.202,pp\. 11398–11442\.External Links:[Link](https://proceedings.mlr.press/v202/giannou23a.html)Cited by:[§2\.1](https://arxiv.org/html/2606.12689#S2.SS1.p1.4)\.
- S\. Goyal, Z\. Ji, A\. S\. Rawat, A\. K\. Menon, S\. Kumar, and V\. Nagarajan \(2024\)Think before you speak: training language models with pause tokens\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=ph04CRkPdC)Cited by:[§C\.2](https://arxiv.org/html/2606.12689#A3.SS2.SSS0.Px2.p1.1)\.
- G\. Gur\-Ari, D\. A\. Roberts, and E\. Dyer \(2018\)Gradient descent happens in a tiny subspace\.arXiv preprint arXiv:1812\.04754\.Cited by:[§2\.3](https://arxiv.org/html/2606.12689#S2.SS3.SSS0.Px1.p2.1)\.
- W\. Gurnee and M\. Tegmark \(2023\)Language models represent space and time\.arXiv preprint arXiv:2310\.02207\.Cited by:[§2\.3](https://arxiv.org/html/2606.12689#S2.SS3.SSS0.Px1.p3.1)\.
- S\. Hao, S\. Sukhbaatar, D\. Su, X\. Li, Z\. Hu, J\. E\. Weston, and Y\. Tian \(2025\)Training Large Language Models to Reason in a Continuous Latent Space\.InSecond Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=Itxz7S4Ip3)Cited by:[Appendix A](https://arxiv.org/html/2606.12689#A1.SS0.SSS0.Px1.p1.1),[§C\.1](https://arxiv.org/html/2606.12689#A3.SS1.SSS0.Px1.p2.3),[§C\.1](https://arxiv.org/html/2606.12689#A3.SS1.SSS0.Px2.p1.2),[§C\.2](https://arxiv.org/html/2606.12689#A3.SS2.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2606.12689#S1.p1.1),[§1](https://arxiv.org/html/2606.12689#S1.p2.1),[§2\.1](https://arxiv.org/html/2606.12689#S2.SS1.p2.1),[§2\.3](https://arxiv.org/html/2606.12689#S2.SS3.SSS0.Px1.p1.1),[§3](https://arxiv.org/html/2606.12689#S3.p1.2),[§4\.1](https://arxiv.org/html/2606.12689#S4.SS1.p1.6),[§4\.2](https://arxiv.org/html/2606.12689#S4.SS2.p1.3)\.
- S\. Heimersheim and N\. Nanda \(2024\)How to use and interpret activation patching\.External Links:2404\.15255,[Link](https://arxiv.org/abs/2404.15255)Cited by:[§2\.2](https://arxiv.org/html/2606.12689#S2.SS2.p1.1)\.
- J\. Hewitt and P\. Liang \(2019\)Designing and interpreting probes with control tasks\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),K\. Inui, J\. Jiang, V\. Ng, and X\. Wan \(Eds\.\),Hong Kong, China,pp\. 2733–2743\.External Links:[Link](https://aclanthology.org/D19-1275/),[Document](https://dx.doi.org/10.18653/v1/D19-1275)Cited by:[§2\.2](https://arxiv.org/html/2606.12689#S2.SS2.p1.1)\.
- S\. Jain and B\. C\. Wallace \(2019\)Attention is not Explanation\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\),J\. Burstein, C\. Doran, and T\. Solorio \(Eds\.\),Minneapolis, Minnesota,pp\. 3543–3556\.External Links:[Link](https://aclanthology.org/N19-1357/),[Document](https://dx.doi.org/10.18653/v1/N19-1357)Cited by:[§2\.2](https://arxiv.org/html/2606.12689#S2.SS2.p1.1)\.
- S\. Kantamneni, J\. Engels, S\. Rajamanoharan, M\. Tegmark, and N\. Nanda \(2025\)Are sparse autoencoders useful? a case study in sparse probing\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=rNfzT8YkgO)Cited by:[§2\.2](https://arxiv.org/html/2606.12689#S2.SS2.p1.1)\.
- T\. Korbak, M\. Balesni, E\. Barnes, Y\. Bengio, J\. Benton, J\. Bloom, M\. Chen, A\. Cooney, A\. Dafoe, A\. Dragan, S\. Emmons, O\. Evans, D\. Farhi, R\. Greenblatt, D\. Hendrycks, M\. Hobbhahn, E\. Hubinger, G\. Irving, E\. Jenner, D\. Kokotajlo, V\. Krakovna, S\. Legg, D\. Lindner, D\. Luan, A\. Mądry, J\. Michael, N\. Nanda, D\. Orr, J\. Pachocki, E\. Perez, M\. Phuong, F\. Roger, J\. Saxe, B\. Shlegeris, M\. Soto, E\. Steinberger, J\. Wang, W\. Zaremba, B\. Baker, R\. Shah, and V\. Mikulik \(2025\)Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety\.External Links:2507\.11473,[Link](https://arxiv.org/abs/2507.11473)Cited by:[§1](https://arxiv.org/html/2606.12689#S1.p1.1)\.
- M\. Lan, P\. Torr, and F\. Barez \(2024\)Towards interpretable sequence continuation: analyzing shared circuits in large language models\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,pp\. 12576–12601\.Cited by:[§5](https://arxiv.org/html/2606.12689#S5.SS0.SSS0.Px2.p1.1)\.
- K\. Lasri, T\. Pimentel, A\. Lenci, T\. Poibeau, and R\. Cotterell \(2022\)Probing for the usage of grammatical number\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),S\. Muresan, P\. Nakov, and A\. Villavicencio \(Eds\.\),Dublin, Ireland,pp\. 8818–8831\.External Links:[Link](https://aclanthology.org/2022.acl-long.603/),[Document](https://dx.doi.org/10.18653/v1/2022.acl-long.603)Cited by:[§2\.2](https://arxiv.org/html/2606.12689#S2.SS2.p1.1)\.
- X\. Li, Z\. Yu, Z\. Zhang, X\. Chen, Z\. Zhang, Y\. Zhuang, N\. Sadagopan, and A\. Beniwal \(2025\)When thinking fails: the pitfalls of reasoning for instruction\-following in llms\.arXiv preprint arXiv:2505\.11423\.Cited by:[§1](https://arxiv.org/html/2606.12689#S1.p1.1)\.
- J\. Liang and L\. Pan \(2026\)Do latent\-cot models think step\-by\-step? a mechanistic study on sequential reasoning tasks\.External Links:2602\.00449,[Link](https://arxiv.org/abs/2602.00449)Cited by:[§C\.1](https://arxiv.org/html/2606.12689#A3.SS1.SSS0.Px1.p3.3)\.
- Q\. Lyu, S\. Havaldar, A\. Stein, L\. Zhang, D\. Rao, E\. Wong, M\. Apidianaki, and C\. Callison\-Burch \(2023\)Faithful chain\-of\-thought reasoning\.InProceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia\-Pacific Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 305–329\.Cited by:[§1](https://arxiv.org/html/2606.12689#S1.p1.1)\.
- T\. McGrath, M\. Rahtz, J\. Kramar, V\. Mikulik, and S\. Legg \(2023\)The hydra effect: emergent self\-repair in language model computations\.External Links:2307\.15771,[Link](https://arxiv.org/abs/2307.15771)Cited by:[§5](https://arxiv.org/html/2606.12689#S5.SS0.SSS0.Px2.p1.1)\.
- M\. Méloux, G\. Dirupo, F\. Portet, and M\. Peyrard \(2025a\)The dead salmons of ai interpretability\.arXiv preprint arXiv:2512\.18792\.Cited by:[§2\.2](https://arxiv.org/html/2606.12689#S2.SS2.p1.1),[§5](https://arxiv.org/html/2606.12689#S5.SS0.SSS0.Px2.p1.1)\.
- M\. Méloux, S\. Maniu, F\. Portet, and M\. Peyrard \(2025b\)Everything, everywhere, all at once: is mechanistic interpretability identifiable?\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=5IWJBStfU7)Cited by:[footnote 1](https://arxiv.org/html/2606.12689#footnote1)\.
- K\. Meng, D\. Bau, A\. Andonian, and Y\. Belinkov \(2022\)Locating and editing factual associations in gpt\.Advances in neural information processing systems35,pp\. 17359–17372\.Cited by:[§2\.2](https://arxiv.org/html/2606.12689#S2.SS2.p1.1),[§5](https://arxiv.org/html/2606.12689#S5.SS0.SSS0.Px2.p1.1)\.
- G\. Monea, M\. Peyrard, M\. Josifoski, V\. Chaudhary, J\. Eisner, E\. Kiciman, H\. Palangi, B\. Patra, and R\. West \(2024\)A glitch in the matrix? locating and detecting language model grounding with fakepedia\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 6828–6844\.External Links:[Link](https://aclanthology.org/2024.acl-long.369/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.369)Cited by:[§2\.2](https://arxiv.org/html/2606.12689#S2.SS2.p1.1)\.
- K\. Park, Y\. J\. Choe, and V\. Veitch \(2023\)The linear representation hypothesis and the geometry of large language models\.arXiv preprint arXiv:2311\.03658\.Cited by:[§2\.3](https://arxiv.org/html/2606.12689#S2.SS3.SSS0.Px1.p3.1)\.
- B\. Peters, S\. Goyal, M\. E\. Granda, A\. V\. Narmadha, D\. Yugeswardeenoo, C\. S\. McDougall, S\. O’Brien, A\. Panda, K\. Zhu, and C\. Blondin \(2025\)Scratchpad Thinking: Alternation Between Storage and Computation in Latent Reasoning Models\.InMechanistic Interpretability Workshop at NeurIPS 2025,External Links:[Link](https://openreview.net/forum?id=EV30qkZXrR)Cited by:[§1](https://arxiv.org/html/2606.12689#S1.p2.1),[§2\.3](https://arxiv.org/html/2606.12689#S2.SS3.SSS0.Px1.p1.1)\.
- A\. Radford, J\. Wu, R\. Child, D\. Luan, D\. Amodei, and I\. Sutskever \(2019\)Language Models are Unsupervised Multitask Learners\.OpenAI blog1\(8\),pp\. 9\.External Links:[Link](https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf)Cited by:[§C\.3](https://arxiv.org/html/2606.12689#A3.SS3.SSS0.Px1.p1.2),[§3](https://arxiv.org/html/2606.12689#S3.p1.2)\.
- A\. Ravichander, Y\. Belinkov, and E\. Hovy \(2021\)Probing the probing paradigm: does probing accuracy entail task relevance?\.InProceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume,P\. Merlo, J\. Tiedemann, and R\. Tsarfaty \(Eds\.\),Online,pp\. 3363–3377\.External Links:[Link](https://aclanthology.org/2021.eacl-main.295),[Document](https://dx.doi.org/10.18653/v1/2021.eacl-main.295)Cited by:[§2\.2](https://arxiv.org/html/2606.12689#S2.SS2.p1.1)\.
- M\. Rizvi\-Martel, G\. Rabusseau, and M\. Mosbach \(2026\)The illusion of superposition? a principled analysis of latent thinking in language models\.External Links:2604\.06374,[Link](https://arxiv.org/abs/2604.06374)Cited by:[§C\.1](https://arxiv.org/html/2606.12689#A3.SS1.SSS0.Px1.p2.3),[§2\.3](https://arxiv.org/html/2606.12689#S2.SS3.SSS0.Px1.p1.1)\.
- Z\. Shen, H\. Yan, L\. Zhang, Z\. Hu, Y\. Du, and Y\. He \(2025\)Codi: compressing chain\-of\-thought into continuous space via self\-distillation\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 677–693\.Cited by:[§C\.1](https://arxiv.org/html/2606.12689#A3.SS1.SSS0.Px1.p3.3),[§C\.1](https://arxiv.org/html/2606.12689#A3.SS1.SSS0.Px2.p1.2),[§C\.2](https://arxiv.org/html/2606.12689#A3.SS2.SSS0.Px1.p1.1),[§C\.3](https://arxiv.org/html/2606.12689#A3.SS3.p3.1),[§1](https://arxiv.org/html/2606.12689#S1.p1.1),[§1](https://arxiv.org/html/2606.12689#S1.p2.1),[§2\.1](https://arxiv.org/html/2606.12689#S2.SS1.p2.1),[§2\.3](https://arxiv.org/html/2606.12689#S2.SS3.SSS0.Px1.p1.1),[§4\.2](https://arxiv.org/html/2606.12689#S4.SS2.p1.3)\.
- D\. Teney, M\. Peyrard, and E\. Abbasnejad \(2022\)Predicting is not understanding: recognizing and addressing underspecification in machine learning\.InComputer Vision – ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXIII,Berlin, Heidelberg,pp\. 458–476\.External Links:ISBN 978\-3\-031\-20049\-6,[Link](https://doi.org/10.1007/978-3-031-20050-2_27),[Document](https://dx.doi.org/10.1007/978-3-031-20050-2%5F27)Cited by:[§2\.2](https://arxiv.org/html/2606.12689#S2.SS2.p1.1)\.
- M\. Turpin, J\. Michael, E\. Perez, and S\. Bowman \(2023\)Language models don’t always say what they think: unfaithful explanations in chain\-of\-thought prompting\.Advances in Neural Information Processing Systems36,pp\. 74952–74965\.Cited by:[§1](https://arxiv.org/html/2606.12689#S1.p1.1),[§7](https://arxiv.org/html/2606.12689#S7.p4.1)\.
- K\. R\. Wang, A\. Variengien, A\. Conmy, B\. Shlegeris, and J\. Steinhardt \(2023\)Interpretability in the wild: a circuit for indirect object identification in GPT\-2 small\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=NpsVSN6o4ul)Cited by:[§2\.2](https://arxiv.org/html/2606.12689#S2.SS2.p1.1)\.
- Z\. Wang, Y\. Dong, and Q\. Lei \(2026\)When does Chain\-of\-Thought Help: A Markovian Perspective\.InWorkshop on Latent & Implicit Thinking – Going Beyond CoT Reasoning,External Links:[Link](https://openreview.net/forum?id=fz5BC8VJ6X)Cited by:[§2\.3](https://arxiv.org/html/2606.12689#S2.SS3.SSS0.Px1.p3.1)\.
- J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, F\. Xia, E\. Chi, Q\. V\. Le, D\. Zhou,et al\.\(2022\)Chain\-of\-thought prompting elicits reasoning in large language models\.Advances in neural information processing systems35,pp\. 24824–24837\.Cited by:[§1](https://arxiv.org/html/2606.12689#S1.p1.1)\.
- X\. Wei, X\. Liu, Y\. Zang, X\. Dong, Y\. Cao, J\. Wang, X\. Qiu, and D\. Lin \(2026\)SIM\-CoT: Supervised Implicit Chain\-of\-Thought\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=6YRJ4jmVQl)Cited by:[Appendix A](https://arxiv.org/html/2606.12689#A1.SS0.SSS0.Px7.p2.3),[§C\.1](https://arxiv.org/html/2606.12689#A3.SS1.SSS0.Px1.p3.3),[§1](https://arxiv.org/html/2606.12689#S1.p2.1),[§2\.3](https://arxiv.org/html/2606.12689#S2.SS3.SSS0.Px1.p3.1)\.
- Y\. Zhang, B\. Tang, T\. Ju, S\. Duan, and G\. Liu \(2025a\)Do Latent Tokens Think? A Causal and Adversarial Analysis of Chain\-of\-Continuous\-Thought\.External Links:2512\.21711,[Link](https://arxiv.org/abs/2512.21711)Cited by:[§1](https://arxiv.org/html/2606.12689#S1.p3.1),[§2\.3](https://arxiv.org/html/2606.12689#S2.SS3.SSS0.Px1.p2.1),[§5](https://arxiv.org/html/2606.12689#S5.SS0.SSS0.Px3.p1.1)\.
- Z\. Zhang, X\. He, W\. Yan, A\. Shen, C\. Zhao, S\. Wang, Y\. Shen, and X\. E\. Wang \(2025b\)Soft thinking: unlocking the reasoning potential of llms in continuous concept space\.arXiv preprint arXiv:2505\.15778\.Cited by:[§1](https://arxiv.org/html/2606.12689#S1.p1.1)\.
- J\. Zhao, Y\. Sun, W\. Shi, and D\. Song \(2025\)Can aha moments be fake? identifying true and decorative thinking steps in chain\-of\-thought\.arXiv preprint arXiv:2510\.24941\.Cited by:[§7](https://arxiv.org/html/2606.12689#S7.p4.1)\.
- Y\. Zhou, Y\. Wang, X\. Yin, S\. Zhou, and A\. Zhang \(2026\)The geometry of reasoning: flowing logics in representation space\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=ixr5Pcabq7)Cited by:[§2\.3](https://arxiv.org/html/2606.12689#S2.SS3.SSS0.Px1.p3.1)\.
- H\. Zhu, S\. Hao, Z\. Hu, J\. Jiao, S\. J\. Russell, and Y\. Tian \(2025\)Reasoning by superposition: a theoretical perspective on chain of continuous thought\.InAdvances in Neural Information Processing Systems,D\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen \(Eds\.\),Vol\.38,pp\. 79931–79963\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2025/file/72c363c2a573ca2128bd176d3317696b-Paper-Conference.pdf)Cited by:[§C\.1](https://arxiv.org/html/2606.12689#A3.SS1.SSS0.Px1.p2.3),[§1](https://arxiv.org/html/2606.12689#S1.p2.1),[§2\.3](https://arxiv.org/html/2606.12689#S2.SS3.SSS0.Px1.p3.1)\.

## Appendix AExtended Analyses

This section reports additional results for experiments conducted in §[4](https://arxiv.org/html/2606.12689#S4), §[5](https://arxiv.org/html/2606.12689#S5)and §[6](https://arxiv.org/html/2606.12689#S6)\.

kk0123456Depth probed1123456nn\(instanceswith≥\\geq2 candidates\)461461486439182263Table 2:Sample size at each probing depth\. Results atk≥5k\\geq 5are computed over too few instances \(n<50n\{<\}50\) to support aggregate conclusions and are omitted from the main analysis\.#### Superposition and BFS\-like search on Graph\-Hopping\.

The main text reports aggregate results fork=0​…​4k\{=\}0\\ldots 4only\.Haoet al\.\([2025](https://arxiv.org/html/2606.12689#bib.bib1)\)illustrate the probing methodology on individual examples; Table[2](https://arxiv.org/html/2606.12689#A1.T2)shows that the aggregate version necessarily confronts the fact that most ProsQA graphs do not extend to depth 6\.

![Refer to caption](https://arxiv.org/html/2606.12689v1/x6.png)Figure 6:Attention mass distribution of the boundary token across prompt tokens, latent thoughts, and itself, averaged across all transformer layers and heads\.
#### Attention Mass Distribution\.

As an additional demonstration of epiphenomenal patterns in interpretability, we evaluate how much attention the answer\-generating boundary token directs at the latent thought tokens versus the prompt tokens\. Lette​n​dt\_\{end\}be the sequence position of this boundary token\. We extract attention distribution over all preceding key positionsjj, averaged across allLLtransformer layers andHHattention heads:

α¯j=1L⋅H​∑l=1L∑h=1Hαl,h​\(te​n​d,j\)\\bar\{\\alpha\}\_\{j\}=\\frac\{1\}\{L\\cdot H\}\\sum\_\{l=1\}^\{L\}\\sum\_\{h=1\}^\{H\}\\alpha\_\{l,h\}\(t\_\{end\},j\)whereαl,h\\alpha\_\{l,h\}is the attention weight at layerlland headhh\. The total attention mass for a given sequence partition𝒮\\mathcal\{S\}\(e\.g\., prompt or latent thoughts\) is calculated as∑j∈𝒮α¯j\\sum\_\{j\\in\\mathcal\{S\}\}\\bar\{\\alpha\}\_\{j\}\. Results in Figure[6](https://arxiv.org/html/2606.12689#A1.F6)\.

Ongraph\-hopping, PaT and Cuattend to their latent thoughts while C and CODI largely ignore them\. Despite no training, both controls B and CoT place significant mass on their \(forced\) thoughts\. Forarithmetic\-reasoning, all models place considerable mass on their thought tokens\.

![Refer to caption](https://arxiv.org/html/2606.12689v1/x7.png)Figure 7:A logit\-lens projection example for all models on both tasks \(graph\-hopping and arithmetic\-reasoning\) and model families\. The final\-layer hidden states at every thought position are decoded through the LM head to obtain the vocabulary projections\.
#### Logit\-Lens Decoded Thought Trajectories\.

Figure[7](https://arxiv.org/html/2606.12689#A1.F7)illustrates representative logit\-lens trajectories\.

Ongraph\-hopping, PaT and C collapse to the answer delimiter at every position despite strong task performance; only Cuproduces somewhat evolving projections, while CODI and the controls remain semantically incoherent\. Onarithmetic\-reasoning, B has no decodable structure while CoT shows moderate decodability despite no latent training—suggesting inherited structure fromCoTfine\-tuning\. PaT collapses to the answer delimiter at most positions again\. CODI exhibits the clearest alternation between intermediate results and formatting tokens; C has lower decodability but similar alternation while Cushows alternation with operators at non\-result positions\.

![Refer to caption](https://arxiv.org/html/2606.12689v1/x8.png)\(a\)Attention outputs\.
![Refer to caption](https://arxiv.org/html/2606.12689v1/x9.png)\(b\)MLP outputs\.

Figure 8:Per\-layerIEKL\\mathrm\{IE\_\{KL\}\}under partner\-prompt corruption across buckets\{Pfull,Pmax,Pb,T1,…,TK,Tfull,Ab\}\\\{P\_\{\\text\{full\}\},P\_\{\\text\{max\}\},P\_\{b\},T\_\{1\},\\ldots,T\_\{K\},T\_\{\\text\{full\}\},A\_\{b\}\\\}, decomposed into attention \(a\) and MLP \(b\) outputs\. Unlike the full residual stream \(Figure[2](https://arxiv.org/html/2606.12689#S5.F2)\), thought positionsTtT\_\{t\}ongraph\-hoppingcarry asmall but nonzerocausal effect, most visible in the attention decomposition\.Arithmetic\-reasoningmirrors the full\-residual pattern, withTfullT\_\{\\text\{full\}\}recovering most for C, Cu, and CODI\.Takeaway:Latent\-thought influence is graded rather than binary—even where it appears negligible, thought positions exert a measurable causal effect\.
#### Decomposition of Residual Stream Causal Tracing Recovery into Attention and MLP\-Outputs\.

Figures[8\(a\)](https://arxiv.org/html/2606.12689#A1.F8.sf1)&[8\(b\)](https://arxiv.org/html/2606.12689#A1.F8.sf2)show the results\. In contrast to the full residual stream \(Figure[2](https://arxiv.org/html/2606.12689#S5.F2)\) the thought positions ofgraph\-hoppingmodels showvery small but realcausal effect, specially on the attention decomposition\. This further supporting our claim that thought use isn’t necessarily binary but depends on how much influence the latent thoughts carry\. The arithmetic\-reasoning models show similar patterns as in the full residual stream \(Figure[2](https://arxiv.org/html/2606.12689#S5.F2)\)\.

#### Gradient\-Subspace Dimensionality\.

Table[3](https://arxiv.org/html/2606.12689#A1.T3)reports per\-timestep gradient\-subspace diagnostics: rankktk\_\{t\}atρ=0\.95\\rho\{=\}0\.95, mean rankk¯\\bar\{k\}\(excluding the degenerate final recurrent step\), adjacent and off\-diagonalcos2¯​\(Bt,Bt′\)\\overline\{\\cos^\{2\}\}\(B\_\{t\},B\_\{t^\{\\prime\}\}\), and the norm fraction‖hc‖/‖h‖\\\|h^\{c\}\\\|/\\\|h\\\|retained in the gradient\-subspace \(§[5](https://arxiv.org/html/2606.12689#S5)\)\.

Causal influence concentrates along few directions across both tasks\. Ongraph\-hopping, the subspaces are extremely low\-rank for all models; onarithmetic\-reasoning, ranks grow \(although still low\-rank compared to full thought vector dimensionality\) except for CODI whose mean rank stays comparable across tasks, plausibly a consequence of its LoRA training\.

TaskModelk0k\_\{0\}k1k\_\{1\}k2k\_\{2\}k3k\_\{3\}k4k\_\{4\}k5k\_\{5\}k6k\_\{6\}k¯\\bar\{k\}Adj\.cos2¯\\overline\{\\cos^\{2\}\}Off\-diag\.cos2¯\\overline\{\\cos^\{2\}\}‖hc‖/‖h‖\\\|h^\{c\}\\\|/\\\|h\\\|Graph\-HoppingPaT1516161617171716\.30\.8710\.8710\.697±0\.0260\.697\{\\scriptstyle\\pm 0\.026\}0\.3530\.353C131211111212–11\.80\.9720\.9720\.671±0\.0240\.671\{\\scriptstyle\\pm 0\.024\}0\.0260\.026Cu151317171714–15\.50\.4860\.4860\.197±0\.0120\.197\{\\scriptstyle\\pm 0\.012\}0\.0450\.045CODI203335363637–32\.80\.8980\.8980\.587±0\.0080\.587\{\\scriptstyle\\pm 0\.008\}0\.1000\.100Arithmetic\-ReasoningPaT160145146109122156153141\.60\.7110\.7110\.623±0\.0190\.623\{\\scriptstyle\\pm 0\.019\}0\.3780\.378C155100129135128115–127\.00\.5120\.5120\.394±0\.0210\.394\{\\scriptstyle\\pm 0\.021\}0\.3690\.369Cu126114115122119134–121\.70\.5740\.5740\.404±0\.0140\.404\{\\scriptstyle\\pm 0\.014\}0\.2970\.297CODI511464185713–36\.20\.2710\.2710\.351±0\.0160\.351\{\\scriptstyle\\pm 0\.016\}0\.3060\.306Table 3:Gradient\-subspace diagnostics across tasks\.
#### Variance Decomposition of Latent\-Thoughts\.

We report the static organization of latent\-thought by decomposing their total variancehi,t∈ℝDh\_\{i,t\}\\in\\mathbb\{R\}^\{D\}\(instanceii, timesteptt\) into timestep\-specific, instance\-specific, and residual components usingμ=𝔼i,t​\[hi,t\]\\mu=\\mathbb\{E\}\_\{i,t\}\[h\_\{i,t\}\]\(global mean\),μt=𝔼i​\[hi,t\]\\mu\_\{t\}=\\mathbb\{E\}\_\{i\}\[h\_\{i,t\}\]\(temporal\-mean\) andμi=𝔼t​\[hi,t\]\\mu\_\{i\}=\\mathbb\{E\}\_\{t\}\[h\_\{i,t\}\]\(instance\-mean\) as:

V​a​r​\(hi,t\)\\displaystyle Var\(h\_\{i,t\}\)=𝔼t​‖μt−μ‖2⏟V​a​rtime\+𝔼i​‖μi−μ‖2⏟V​a​rinst\\displaystyle=\\underbrace\{\\mathbb\{E\}\_\{t\}\\\|\\mu\_\{t\}\-\\mu\\\|^\{2\}\}\_\{Var\_\{\\mathrm\{time\}\}\}\+\\underbrace\{\\mathbb\{E\}\_\{i\}\\\|\\mu\_\{i\}\-\\mu\\\|^\{2\}\}\_\{Var\_\{\\mathrm\{inst\}\}\}\+𝔼i,t​‖hi,t−μi−μt\+μ‖2⏟V​a​rresidual\\displaystyle\\quad\+\\underbrace\{\\mathbb\{E\}\_\{i,t\}\\\|h\_\{i,t\}\-\\mu\_\{i\}\-\\mu\_\{t\}\+\\mu\\\|^\{2\}\}\_\{Var\_\{\\mathrm\{residual\}\}\}For PaT,ttindexes spatial tokens; for C and Cu, it indexes recurrence\-steps\. Results in Table[4](https://arxiv.org/html/2606.12689#A1.T4)\.

Graph\-HoppingArithmetic\-ReasoningModelVartime\{\}\_\{\\textbf\{time\}\}Varinst\{\}\_\{\\textbf\{inst\}\}Vartime\{\}\_\{\\textbf\{time\}\}Varinst\{\}\_\{\\textbf\{inst\}\}PaT0\.6±0\.10\.6\{\\scriptstyle\\pm 0\.1\}98\.7±0\.198\.7\{\\scriptstyle\\pm 0\.1\}1\.9±0\.21\.9\{\\scriptstyle\\pm 0\.2\}72\.6±0\.872\.6\{\\scriptstyle\\pm 0\.8\}C22\.6±1\.922\.6\{\\scriptstyle\\pm 1\.9\}68\.9±2\.168\.9\{\\scriptstyle\\pm 2\.1\}25\.0±0\.425\.0\{\\scriptstyle\\pm 0\.4\}26\.0±0\.626\.0\{\\scriptstyle\\pm 0\.6\}Cu99\.1±0\.099\.1\{\\scriptstyle\\pm 0\.0\}0\.4±0\.00\.4\{\\scriptstyle\\pm 0\.0\}10\.8±0\.610\.8\{\\scriptstyle\\pm 0\.6\}26\.1±0\.526\.1\{\\scriptstyle\\pm 0\.5\}CODI89\.8±0\.789\.8\{\\scriptstyle\\pm 0\.7\}7\.4±0\.67\.4\{\\scriptstyle\\pm 0\.6\}5\.6±0\.35\.6\{\\scriptstyle\\pm 0\.3\}46\.3±1\.146\.3\{\\scriptstyle\\pm 1\.1\}Table 4:Variance decomposition of latent thoughts into temporal, instance, and residual components\.Takeaway:Lack of generalizable structure suggests that observable static decomposition alone might not reliably differentiate active and inert thoughts\. PaT carries 1\-2 orders smallerVartime\\text\{Var\}\_\{\\text\{time\}\}than other models across tasks\.![Refer to caption](https://arxiv.org/html/2606.12689v1/x10.png)Figure 9:2D PCA projection of the latent space structure forCuC\_\{u\}on the graph\-hopping task\.Takeaway:Cuis dominated by timestep\-specific variance with negligible instance\-specific variance, exhibiting perfectly separable timestep corresponding clusters\.Ongraph\-hoppingshows PaT and C are Varinst\{\}\_\{\\text\{inst\}\}dominated \(98\.7%98\.7\\%and68\.9%68\.9\\%, respectively\), whereas Cuis Vartime\{\}\_\{\\text\{time\}\}dominated \(99\.1%99\.1\\%\), forming nearly perfectly separable clusters corresponding to recurrence depth \(Figure[9](https://arxiv.org/html/2606.12689#A1.F9)\)\. Forarithmetic\-reasoning, all recurrence\-based models are Varresidual\{\}\_\{\\text\{residual\}\}dominated \(implicit in Table[4](https://arxiv.org/html/2606.12689#A1.T4)as the remaining variance: Cuat63\.1%63\.1\\%, C at49\.0%49\.0\\%, and CODI at48\.1%48\.1\\%\)\. This is with the exception of PaT, which remains Varinst\{\}\_\{\\text\{inst\}\}dominated \(72\.6%72\.6\\%\) and whose temporal component is 1\-2 orders of magnitude smaller than the recurrence\-based models\.

#### Mean\-Ablations and Isolations\.

We apply six counterfactual interventions to the latent thoughtshi,th\_\{i,t\}at inference to identify component\-wise importance\.Ablationsremove one component: temporal \(hi,t′=hi,t−μt\+μh^\{\\prime\}\_\{i,t\}=h\_\{i,t\}\-\\mu\_\{t\}\+\\mu\), instance \(hi,t′=hi,t−μ^i\+μh^\{\\prime\}\_\{i,t\}=h\_\{i,t\}\-\\hat\{\\mu\}\_\{i\}\+\\mu, withμ^i=1K\+1​∑k=0Khi,k\\hat\{\\mu\}\_\{i\}=\\frac\{1\}\{K\+1\}\\sum\_\{k=0\}^\{K\}h\_\{i,k\}\), or residual \(hi,t′=μt\+μ^i−μh^\{\\prime\}\_\{i,t\}=\\mu\_\{t\}\+\\hat\{\\mu\}\_\{i\}\-\\mu\)\.Isolationsretain only one component:hi,t′=μth^\{\\prime\}\_\{i,t\}=\\mu\_\{t\},hi,t′=μ^ih^\{\\prime\}\_\{i,t\}=\\hat\{\\mu\}\_\{i\}, orhi,t′=hi,t−μt−μ^i\+2​μh^\{\\prime\}\_\{i,t\}=h\_\{i,t\}\-\\mu\_\{t\}\-\\hat\{\\mu\}\_\{i\}\+2\\mu, respectively\. A*matched\-norm random control*adds isotropic noise tohi,th\_\{i,t\}rescaled to the temporal\-deviation norm:hi,t′=hi,t\+ε⋅‖μt−μ‖/‖ε‖h^\{\\prime\}\_\{i,t\}=h\_\{i,t\}\+\\varepsilon\\cdot\\\|\\mu\_\{t\}\-\\mu\\\|/\\\|\\varepsilon\\\|\(ε∼𝒩​\(0,ID\)\\varepsilon\\sim\\mathcal\{N\}\(0,I\_\{D\}\)\)\. To prevent test\-set leakage, all reference means \(μ,μt\\mu,\\mu\_\{t\}\) are computed from the training split\. Results in Figure[10](https://arxiv.org/html/2606.12689#A1.F10)\.

![Refer to caption](https://arxiv.org/html/2606.12689v1/x11.png)Figure 10:Targeted interventions isolating and ablating specific variance components of the latent thoughts\.Takeaway:Inert thoughts are largely unaffected under perturbations\. Active thoughts show varied dependence on components, affected least by the temporal and most by the residual\. Temporal structure appears functionally inert across models and tasks\.Graph\-hoppingmodels are robust to perturbations except PaT, which requiresμi\\mu\_\{i\}for stability\. Notably, neither variance nor gradient\-subspace interventions recover Cu’s thought\-ablation drop, suggesting that its signal is distributed rather than localized\. Onarithmetic reasoning, PaT shows the same pattern\. Ablatingμt\\mu\_\{t\}yields the smallest drop while isolating it causes the largest, indicating temporal structure is largely functionally inert—contrastingWeiet al\.\([2026](https://arxiv.org/html/2606.12689#bib.bib35)\)’s account of vocabulary\-space drift as a key vulnerability\. Ablating the residual component produces the largest drop, but its isolation has minimal impact\. The matched\-norm control preserves performance across all tasks\.

## Appendix BStatistical Evaluation

This section presents the full statistical\-tests for every experiment conducted in §[4](https://arxiv.org/html/2606.12689#S4), §[5](https://arxiv.org/html/2606.12689#S5)and §[6](https://arxiv.org/html/2606.12689#S6)\. Unless stated otherwise, all confidence intervals are 95% percentile bootstrap intervals computed with 1,000 resamples over per\-instance outcome vectors, and paired comparisons report exact two\-sided McNemarpp\-values on discordant pairs\(b,c\)\(b,c\), wherebbcounts instances correct only under condition A andcccounts instances correct only under condition B\. Significance markers follow the conventionp∗<0\.05\{\}^\{\*\}p\{<\}0\.05,p∗∗<0\.01\{\}^\{\*\*\}p\{<\}0\.01,p∗⁣∗∗<0\.001\{\}^\{\*\*\*\}p\{<\}0\.001\.

#### Epiphenomenal Patterns in LRM Interpretability \(§[4](https://arxiv.org/html/2606.12689#S4)\)\.

H/log2⁡NH/\\log\_\{2\}NTop\-1 correct \(%\)kkBCoTPaTCCuCODIBCoTPaTCCuCODI00\.58±0\.030\.58\{\\scriptstyle\\pm 0\.03\}0\.16±0\.020\.16\{\\scriptstyle\\pm 0\.02\}0\.51±0\.030\.51\{\\scriptstyle\\pm 0\.03\}0\.43±0\.030\.43\{\\scriptstyle\\pm 0\.03\}0\.44±0\.030\.44\{\\scriptstyle\\pm 0\.03\}0\.34±0\.030\.34\{\\scriptstyle\\pm 0\.03\}81\.3±3\.781\.3\{\\scriptstyle\\pm 3\.7\}84\.2±3\.584\.2\{\\scriptstyle\\pm 3\.5\}65\.1±4\.265\.1\{\\scriptstyle\\pm 4\.2\}63\.8±4\.163\.8\{\\scriptstyle\\pm 4\.1\}90\.7±2\.590\.7\{\\scriptstyle\\pm 2\.5\}64\.4±4\.664\.4\{\\scriptstyle\\pm 4\.6\}10\.33±0\.030\.33\{\\scriptstyle\\pm 0\.03\}0\.15±0\.020\.15\{\\scriptstyle\\pm 0\.02\}0\.51±0\.030\.51\{\\scriptstyle\\pm 0\.03\}0\.47±0\.030\.47\{\\scriptstyle\\pm 0\.03\}0\.46±0\.030\.46\{\\scriptstyle\\pm 0\.03\}0\.33±0\.030\.33\{\\scriptstyle\\pm 0\.03\}68\.8±4\.168\.8\{\\scriptstyle\\pm 4\.1\}83\.9±3\.683\.9\{\\scriptstyle\\pm 3\.6\}64\.9±4\.464\.9\{\\scriptstyle\\pm 4\.4\}65\.1±4\.165\.1\{\\scriptstyle\\pm 4\.1\}90\.7±2\.590\.7\{\\scriptstyle\\pm 2\.5\}64\.4±4\.464\.4\{\\scriptstyle\\pm 4\.4\}20\.37±0\.020\.37\{\\scriptstyle\\pm 0\.02\}0\.24±0\.020\.24\{\\scriptstyle\\pm 0\.02\}0\.48±0\.020\.48\{\\scriptstyle\\pm 0\.02\}0\.48±0\.020\.48\{\\scriptstyle\\pm 0\.02\}0\.32±0\.020\.32\{\\scriptstyle\\pm 0\.02\}0\.32±0\.020\.32\{\\scriptstyle\\pm 0\.02\}55\.6±4\.155\.6\{\\scriptstyle\\pm 4\.1\}70\.8±4\.170\.8\{\\scriptstyle\\pm 4\.1\}71\.6±4\.071\.6\{\\scriptstyle\\pm 4\.0\}63\.2±4\.163\.2\{\\scriptstyle\\pm 4\.1\}89\.1±2\.889\.1\{\\scriptstyle\\pm 2\.8\}51\.0±4\.351\.0\{\\scriptstyle\\pm 4\.3\}30\.39±0\.030\.39\{\\scriptstyle\\pm 0\.03\}0\.26±0\.020\.26\{\\scriptstyle\\pm 0\.02\}0\.23±0\.030\.23\{\\scriptstyle\\pm 0\.03\}0\.26±0\.030\.26\{\\scriptstyle\\pm 0\.03\}0\.37±0\.030\.37\{\\scriptstyle\\pm 0\.03\}0\.22±0\.030\.22\{\\scriptstyle\\pm 0\.03\}51\.3±4\.751\.3\{\\scriptstyle\\pm 4\.7\}65\.1±4\.665\.1\{\\scriptstyle\\pm 4\.6\}88\.6±3\.288\.6\{\\scriptstyle\\pm 3\.2\}85\.2±3\.485\.2\{\\scriptstyle\\pm 3\.4\}86\.6±3\.386\.6\{\\scriptstyle\\pm 3\.3\}66\.7±4\.366\.7\{\\scriptstyle\\pm 4\.3\}40\.40±0\.050\.40\{\\scriptstyle\\pm 0\.05\}0\.27±0\.040\.27\{\\scriptstyle\\pm 0\.04\}0\.13±0\.030\.13\{\\scriptstyle\\pm 0\.03\}0\.14±0\.040\.14\{\\scriptstyle\\pm 0\.04\}0\.38±0\.050\.38\{\\scriptstyle\\pm 0\.05\}0\.19±0\.040\.19\{\\scriptstyle\\pm 0\.04\}57\.7±7\.157\.7\{\\scriptstyle\\pm 7\.1\}70\.3±6\.370\.3\{\\scriptstyle\\pm 6\.3\}97\.3±2\.297\.3\{\\scriptstyle\\pm 2\.2\}92\.9±3\.892\.9\{\\scriptstyle\\pm 3\.8\}70\.9±6\.970\.9\{\\scriptstyle\\pm 6\.9\}78\.6±6\.078\.6\{\\scriptstyle\\pm 6\.0\}P​\(correct\)P\(\\text\{correct\}\)Cand\. masskkBCoTPaTCCuCODIBCoTPaTCCuCODI00\.05±0\.010\.05\{\\scriptstyle\\pm 0\.01\}0\.36±0\.040\.36\{\\scriptstyle\\pm 0\.04\}0\.00±0\.000\.00\{\\scriptstyle\\pm 0\.00\}0\.00±0\.000\.00\{\\scriptstyle\\pm 0\.00\}0\.84±0\.020\.84\{\\scriptstyle\\pm 0\.02\}0\.02±0\.010\.02\{\\scriptstyle\\pm 0\.01\}0\.07±0\.010\.07\{\\scriptstyle\\pm 0\.01\}0\.40±0\.040\.40\{\\scriptstyle\\pm 0\.04\}0\.00±0\.000\.00\{\\scriptstyle\\pm 0\.00\}0\.00±0\.000\.00\{\\scriptstyle\\pm 0\.00\}0\.95±0\.010\.95\{\\scriptstyle\\pm 0\.01\}0\.03±0\.010\.03\{\\scriptstyle\\pm 0\.01\}10\.00±0\.000\.00\{\\scriptstyle\\pm 0\.00\}0\.39±0\.040\.39\{\\scriptstyle\\pm 0\.04\}0\.00±0\.000\.00\{\\scriptstyle\\pm 0\.00\}0\.00±0\.000\.00\{\\scriptstyle\\pm 0\.00\}0\.85±0\.020\.85\{\\scriptstyle\\pm 0\.02\}0\.01±0\.010\.01\{\\scriptstyle\\pm 0\.01\}0\.01±0\.000\.01\{\\scriptstyle\\pm 0\.00\}0\.44±0\.040\.44\{\\scriptstyle\\pm 0\.04\}0\.00±0\.000\.00\{\\scriptstyle\\pm 0\.00\}0\.00±0\.000\.00\{\\scriptstyle\\pm 0\.00\}0\.97±0\.010\.97\{\\scriptstyle\\pm 0\.01\}0\.03±0\.010\.03\{\\scriptstyle\\pm 0\.01\}20\.00±0\.000\.00\{\\scriptstyle\\pm 0\.00\}0\.09±0\.020\.09\{\\scriptstyle\\pm 0\.02\}0\.01±0\.010\.01\{\\scriptstyle\\pm 0\.01\}0\.00±0\.000\.00\{\\scriptstyle\\pm 0\.00\}0\.71±0\.030\.71\{\\scriptstyle\\pm 0\.03\}0\.01±0\.010\.01\{\\scriptstyle\\pm 0\.01\}0\.01±0\.000\.01\{\\scriptstyle\\pm 0\.00\}0\.10±0\.020\.10\{\\scriptstyle\\pm 0\.02\}0\.01±0\.010\.01\{\\scriptstyle\\pm 0\.01\}0\.00±0\.000\.00\{\\scriptstyle\\pm 0\.00\}0\.81±0\.020\.81\{\\scriptstyle\\pm 0\.02\}0\.03±0\.010\.03\{\\scriptstyle\\pm 0\.01\}30\.00±0\.000\.00\{\\scriptstyle\\pm 0\.00\}0\.04±0\.010\.04\{\\scriptstyle\\pm 0\.01\}0\.34±0\.040\.34\{\\scriptstyle\\pm 0\.04\}0\.37±0\.040\.37\{\\scriptstyle\\pm 0\.04\}0\.38±0\.030\.38\{\\scriptstyle\\pm 0\.03\}0\.15±0\.030\.15\{\\scriptstyle\\pm 0\.03\}0\.00±0\.000\.00\{\\scriptstyle\\pm 0\.00\}0\.04±0\.020\.04\{\\scriptstyle\\pm 0\.02\}0\.34±0\.040\.34\{\\scriptstyle\\pm 0\.04\}0\.37±0\.040\.37\{\\scriptstyle\\pm 0\.04\}0\.40±0\.030\.40\{\\scriptstyle\\pm 0\.03\}0\.16±0\.030\.16\{\\scriptstyle\\pm 0\.03\}40\.00±0\.000\.00\{\\scriptstyle\\pm 0\.00\}0\.02±0\.020\.02\{\\scriptstyle\\pm 0\.02\}0\.58±0\.060\.58\{\\scriptstyle\\pm 0\.06\}0\.63±0\.070\.63\{\\scriptstyle\\pm 0\.07\}0\.23±0\.050\.23\{\\scriptstyle\\pm 0\.05\}0\.16±0\.050\.16\{\\scriptstyle\\pm 0\.05\}0\.00±0\.000\.00\{\\scriptstyle\\pm 0\.00\}0\.03±0\.020\.03\{\\scriptstyle\\pm 0\.02\}0\.59±0\.070\.59\{\\scriptstyle\\pm 0\.07\}0\.63±0\.070\.63\{\\scriptstyle\\pm 0\.07\}0\.24±0\.050\.24\{\\scriptstyle\\pm 0\.05\}0\.16±0\.050\.16\{\\scriptstyle\\pm 0\.05\}Table 5:Full results for superposition and BFS\-probing across timesteps and models with statistical testing\.ModelStepHit RateSuperpos\.Step Align\.B00\.1\(0\.0, 0\.2\)0\.0\(0\.0, 0\.0\)0\.0\(0\.0, 0\.0\)11\.4\(0\.8, 2\.1\)0\.0\(0\.0, 0\.0\)0\.3\(0\.0, 0\.6\)20\.3\(0\.1, 0\.6\)0\.0\(0\.0, 0\.0\)0\.0\(0\.0, 0\.0\)30\.3\(0\.1, 0\.6\)0\.0\(0\.0, 0\.0\)0\.1\(0\.0, 0\.2\)40\.2\(0\.0, 0\.5\)0\.0\(0\.0, 0\.0\)0\.0\(0\.0, 0\.0\)50\.2\(0\.0, 0\.4\)0\.0\(0\.0, 0\.0\)0\.0\(0\.0, 0\.0\)60\.2\(0\.0, 0\.4\)0\.0\(0\.0, 0\.0\)0\.0\(0\.0, 0\.0\)CoT051\.6\(49\.0, 54\.3\)15\.4\(13\.4, 17\.3\)35\.2\(32\.4, 37\.8\)17\.1\(5\.8, 8\.5\)0\.6\(0\.2, 1\.1\)1\.7\(1\.0, 2\.5\)240\.6\(38\.0, 43\.4\)11\.7\(10\.0, 13\.5\)9\.7\(8\.1, 11\.4\)310\.3\(8\.8, 11\.8\)1\.3\(0\.8, 2\.0\)0\.6\(0\.2, 1\.1\)439\.4\(36\.8, 42\.0\)11\.5\(9\.8, 13\.2\)2\.0\(1\.3, 2\.8\)512\.1\(10\.5, 13\.9\)1\.7\(1\.0, 2\.5\)0\.2\(0\.0, 0\.4\)636\.2\(33\.4, 38\.8\)10\.5\(8\.9, 12\.1\)0\.0\(0\.0, 0\.0\)PaT065\.4\(62\.9, 67\.9\)22\.6\(20\.4, 24\.9\)48\.3\(45\.7, 51\.0\)165\.6\(63\.1, 68\.1\)24\.1\(22\.0, 26\.5\)40\.7\(37\.9, 43\.4\)264\.3\(61\.9, 66\.9\)23\.8\(21\.4, 26\.1\)18\.7\(16\.7, 20\.9\)362\.2\(59\.6, 64\.7\)23\.1\(20\.9, 25\.4\)7\.9\(6\.5, 9\.5\)460\.0\(57\.6, 62\.6\)23\.4\(21\.1, 25\.8\)4\.2\(3\.1, 5\.2\)558\.9\(56\.3, 61\.5\)22\.8\(20\.7, 25\.1\)1\.3\(0\.7, 1\.9\)658\.6\(55\.9, 61\.1\)21\.9\(19\.7, 24\.1\)0\.2\(0\.0, 0\.4\)C029\.4\(27\.0, 32\.1\)4\.2\(3\.1, 5\.4\)17\.7\(15\.7, 19\.7\)173\.0\(70\.6, 75\.5\)20\.8\(18\.7, 23\.2\)20\.8\(18\.7, 23\.0\)219\.7\(17\.6, 21\.9\)2\.9\(2\.1, 3\.8\)4\.5\(3\.5, 5\.6\)371\.6\(69\.3, 74\.1\)18\.8\(16\.5, 20\.8\)6\.3\(5\.0, 7\.7\)430\.4\(28\.0, 32\.9\)5\.5\(4\.3, 6\.8\)1\.5\(0\.8, 2\.2\)564\.8\(62\.3, 67\.4\)16\.6\(14\.5, 18\.7\)0\.7\(0\.2, 1\.2\)629\.2\(26\.7, 31\.7\)6\.1\(4\.8, 7\.5\)0\.2\(0\.0, 0\.4\)CODI075\.1\(72\.8, 77\.2\)21\.9\(19\.7, 24\.1\)63\.9\(61\.3, 66\.3\)10\.5\(0\.2, 0\.8\)0\.0\(0\.0, 0\.0\)0\.1\(0\.0, 0\.2\)273\.0\(70\.6, 75\.5\)22\.1\(19\.8, 24\.4\)15\.3\(13\.4, 17\.2\)30\.4\(0\.1, 0\.8\)0\.1\(0\.0, 0\.2\)0\.0\(0\.0, 0\.0\)470\.8\(68\.5, 73\.2\)21\.8\(19\.6, 23\.9\)3\.2\(2\.2, 4\.2\)50\.5\(0\.2, 0\.8\)0\.0\(0\.0, 0\.0\)0\.0\(0\.0, 0\.0\)670\.1\(67\.8, 72\.6\)22\.1\(20\.1, 24\.4\)0\.2\(0\.0, 0\.4\)Pooled \(all steps\)B–0\.4\(0\.2, 0\.5\)0\.0\(0\.0, 0\.0\)–CoT–28\.2\(27\.3, 29\.1\)7\.5\(7\.0, 8\.0\)–PaT–62\.1\(61\.2, 63\.2\)23\.1\(22\.3, 23\.9\)–C–45\.4\(44\.5, 46\.6\)10\.7\(10\.0, 11\.4\)–CODI–41\.5\(40\.4, 42\.5\)12\.6\(11\.8, 13\.2\)–Table 6:Full results for scratchpad\-thinking probing across timesteps and models with statistical testing\.Table[5](https://arxiv.org/html/2606.12689#A2.T5)reports the results for case study 1 \(superposition and BFS\-like search on graph\-hopping\); Table[6](https://arxiv.org/html/2606.12689#A2.T6)reports results for case study 2 \(scratchpad reasoning on arithmetic\-reasoning via logit\-lens\)\.

#### When and How do LRMs use Latent Thoughts \(§[5](https://arxiv.org/html/2606.12689#S5)\)?

ModelAccKmax\{\}\_\{K\_\{\\max\}\}\[% CI\]AccK=0\[% CI\]Δ\\Delta\[% CI\]McNemarppnnGraph\-HoppingB2\.4 \[1\.2, 3\.8\]–––500CoT79\.0 \[75\.4, 82\.6\]–––500PaT95\.4 \[93\.6, 97\.2\]95\.6 \[93\.6, 97\.2\]\-0\.2 \[\-0\.6, 0\.0\]1\.0000 \(b=0, c=1\)500C98\.0 \[96\.6, 99\.0\]97\.8 \[96\.4, 99\.0\]\+0\.2 \[0\.0, 0\.6\]1\.0000 \(b=1, c=0\)500Cu96\.0 \[94\.2, 97\.6\]91\.8 \[89\.4, 94\.0\]\+4\.2 \[1\.6, 6\.8\]0\.0019∗∗\(b=32, c=11\)500CODI79\.6 \[76\.0, 83\.2\]79\.8 \[76\.2, 83\.4\]\-0\.2 \[\-1\.2, 0\.8\]1\.0000 \(b=3, c=4\)500Arithmetic\-ReasoningB1\.4 \[0\.8, 2\.0\]–––1319CoT41\.9 \[39\.3, 44\.6\]–––1319PaT26\.4 \[24\.1, 28\.7\]21\.4 \[19\.4, 23\.7\]\+5\.0 \[3\.5, 6\.7\]0\.0000∗∗∗\(b=95, c=29\)1319C35\.7 \[33\.4, 38\.1\]7\.7 \[6\.4, 9\.2\]\+28\.1 \[25\.6, 30\.5\]0\.0000∗∗∗\(b=391, c=21\)1319Cu30\.8 \[28\.3, 33\.2\]39\.1 \[36\.4, 41\.6\]\-8\.3 \[\-10\.7, \-6\.0\]0\.0000∗∗∗\(b=74, c=184\)1319CODI41\.8 \[39\.2, 44\.7\]24\.8 \[22\.4, 27\.1\]\+17\.1 \[14\.4, 19\.9\]0\.0000∗∗∗\(b=297, c=72\)1319Table 7:Full results for latent thought ablation at test\-time with statistical testing\.ModelPfull\{\}\_\{\\text\{full\}\}PmaxPbAbKLc→c′¯\\overline\{\\text\{KL\}\_\{c\\to c^\{\\prime\}\}\}Wilc\.ppMcN\.pp\(τ=0\.5\\tau\{=\}0\.5\)nnGraph\-HoppingPaT1\.355 \[0\.977, 1\.817\]5\.705 \[5\.008, 6\.388\]14\.577 \[13\.697, 15\.488\]14\.657 \[13\.791, 15\.565\]15\.158 \[14\.280, 16\.085\]0\.0000∗∗∗0\.0000∗∗∗\(0,421\)500C1\.554 \[1\.145, 1\.992\]5\.315 \[4\.740, 5\.934\]14\.204 \[13\.248, 15\.077\]14\.143 \[13\.196, 15\.032\]15\.320 \[14\.382, 16\.198\]0\.0000∗∗∗0\.0000∗∗∗\(0,435\)500CuC\_\{u\}1\.694 \[1\.201, 2\.208\]4\.499 \[3\.953, 5\.085\]13\.776 \[12\.953, 14\.662\]13\.528 \[12\.741, 14\.351\]13\.960 \[13\.122, 14\.835\]0\.0000∗∗∗0\.0000∗∗∗\(0,464\)500CODI0\.861 \[0\.607, 1\.167\]5\.324 \[4\.755, 5\.911\]11\.135 \[10\.410, 11\.857\]11\.218 \[10\.509, 11\.945\]11\.540 \[10\.831, 12\.255\]0\.0000∗∗∗0\.0000∗∗∗\(0,455\)500Arithmetic\-ReasoningPaT5\.361 \[5\.162, 5\.572\]6\.496 \[6\.255, 6\.743\]10\.408 \[10\.089, 10\.730\]10\.517 \[10\.195, 10\.847\]11\.736 \[11\.390, 12\.085\]0\.0000∗∗∗0\.0000∗∗∗\(136,423\)1319C4\.877 \[4\.684, 5\.071\]5\.863 \[5\.646, 6\.070\]7\.981 \[7\.741, 8\.213\]8\.905 \[8\.615, 9\.198\]10\.157 \[9\.869, 10\.447\]0\.0000∗∗∗0\.0000∗∗∗\(372,136\)1319CuC\_\{u\}3\.921 \[3\.737, 4\.121\]4\.444 \[4\.259, 4\.644\]6\.248 \[6\.029, 6\.486\]6\.971 \[6\.722, 7\.224\]7\.860 \[7\.594, 8\.155\]0\.0000∗∗∗0\.0000∗∗∗\(340,172\)1319CODI4\.194 \[3\.997, 4\.382\]5\.919 \[5\.669, 6\.172\]8\.599 \[8\.318, 8\.892\]10\.424 \[10\.111, 10\.757\]10\.542 \[10\.223, 10\.877\]0\.0000∗∗∗0\.0000∗∗∗\(164,454\)1319Table 8:Full causal tracing results with statistical testing \(prompt\-positions\)\.ModelT1T2T3T4T5T6Tfull\{\}\_\{\\text\{full\}\}Graph\-HoppingPaT14\.610 \[13\.735, 15\.521\]14\.621 \[13\.747, 15\.532\]14\.636 \[13\.767, 15\.537\]14\.637 \[13\.768, 15\.537\]14\.650 \[13\.783, 15\.555\]14\.651 \[13\.786, 15\.560\]13\.612 \[12\.721, 14\.475\]C15\.305 \[14\.369, 16\.184\]15\.307 \[14\.371, 16\.187\]15\.287 \[14\.332, 16\.166\]15\.288 \[14\.332, 16\.166\]15\.310 \[14\.373, 16\.190\]14\.203 \[13\.252, 15\.091\]14\.205 \[13\.255, 15\.087\]CuC\_\{u\}13\.842 \[13\.021, 14\.730\]13\.846 \[13\.023, 14\.735\]13\.779 \[12\.973, 14\.640\]13\.808 \[12\.995, 14\.681\]13\.808 \[12\.995, 14\.681\]13\.761 \[12\.936, 14\.622\]13\.661 \[12\.854, 14\.504\]CODI11\.192 \[10\.474, 11\.918\]11\.234 \[10\.514, 11\.963\]11\.221 \[10\.510, 11\.940\]11\.241 \[10\.523, 11\.965\]11\.228 \[10\.505, 11\.964\]11\.249 \[10\.544, 11\.987\]11\.033 \[10\.301, 11\.750\]Arithmetic\-ReasoningPaT10\.278 \[9\.967, 10\.596\]10\.153 \[9\.838, 10\.472\]10\.011 \[9\.691, 10\.325\]9\.955 \[9\.642, 10\.270\]10\.019 \[9\.710, 10\.333\]10\.105 \[9\.794, 10\.422\]6\.745 \[6\.483, 6\.998\]C6\.276 \[6\.061, 6\.471\]6\.071 \[5\.871, 6\.256\]6\.092 \[5\.886, 6\.293\]6\.638 \[6\.413, 6\.852\]8\.012 \[7\.754, 8\.290\]9\.406 \[9\.112, 9\.707\]3\.165 \[3\.006, 3\.303\]CuC\_\{u\}5\.254 \[5\.046, 5\.469\]5\.019 \[4\.826, 5\.222\]5\.005 \[4\.806, 5\.212\]5\.049 \[4\.875, 5\.252\]5\.975 \[5\.765, 6\.210\]7\.655 \[7\.397, 7\.942\]2\.812 \[2\.658, 2\.986\]CODI8\.105 \[7\.841, 8\.376\]8\.324 \[8\.054, 8\.601\]8\.193 \[7\.923, 8\.478\]8\.241 \[7\.974, 8\.521\]8\.317 \[8\.051, 8\.615\]9\.895 \[9\.584, 10\.220\]5\.449 \[5\.232, 5\.664\]Table 9:Full causal tracing results with statistical testing \(thought\-positions\)\.ModelAccorig\{\}\_\{\\text\{orig\}\}\[% CI\]Accgrad\{\}\_\{\\text\{grad\}\}\[% CI\]Accrand\{\}\_\{\\text\{rand\}\}\[% CI\]Δgrad\\Delta\_\{\\text\{grad\}\}\[% CI\]McNemargrad\{\}\_\{\\text\{grad\}\}ppΔrand\\Delta\_\{\\text\{rand\}\}\[% CI\]McNemarrand\{\}\_\{\\text\{rand\}\}ppnnGraph\-HoppingPaT95\.4±1\.895\.4\{\\scriptstyle\\pm 1\.8\}95\.8±1\.795\.8\{\\scriptstyle\\pm 1\.7\}95\.4±1\.095\.4\{\\scriptstyle\\pm 1\.0\}−0\.4±0\.5\-0\.4\{\\scriptstyle\\pm 0\.5\}0\.5000 \(b=0, c=2\)\+0\.0±0\.0\+0\.0\{\\scriptstyle\\pm 0\.0\}1\.0000 \(b=0, c=0\)–C98\.0±1\.298\.0\{\\scriptstyle\\pm 1\.2\}98\.0±1\.298\.0\{\\scriptstyle\\pm 1\.2\}98\.0±0\.798\.0\{\\scriptstyle\\pm 0\.7\}\+0\.0±0\.0\+0\.0\{\\scriptstyle\\pm 0\.0\}1\.0000 \(b=0, c=0\)\+0\.0±0\.0\+0\.0\{\\scriptstyle\\pm 0\.0\}1\.0000 \(b=0, c=0\)–Cu96\.0±1\.796\.0\{\\scriptstyle\\pm 1\.7\}96\.0±1\.796\.0\{\\scriptstyle\\pm 1\.7\}96\.0±1\.096\.0\{\\scriptstyle\\pm 1\.0\}\+0\.0±0\.0\+0\.0\{\\scriptstyle\\pm 0\.0\}1\.0000 \(b=0, c=0\)\+0\.0±0\.0\+0\.0\{\\scriptstyle\\pm 0\.0\}1\.0000 \(b=0, c=0\)–CODI80\.4±3\.580\.4\{\\scriptstyle\\pm 3\.5\}79\.8±3\.679\.8\{\\scriptstyle\\pm 3\.6\}80\.0±2\.080\.0\{\\scriptstyle\\pm 2\.0\}\+0\.6±0\.9\+0\.6\{\\scriptstyle\\pm 0\.9\}0\.3750 \(b=4, c=1\)\+0\.4±0\.4\+0\.4\{\\scriptstyle\\pm 0\.4\}0\.1094 \(b=8, c=2\)–Arithmetic\-ReasoningPaT26\.4±2\.326\.4\{\\scriptstyle\\pm 2\.3\}23\.6±2\.123\.6\{\\scriptstyle\\pm 2\.1\}26\.4±1\.426\.4\{\\scriptstyle\\pm 1\.4\}\+2\.8±1\.6\+2\.8\{\\scriptstyle\\pm 1\.6\}0\.0003∗∗∗\(b=70, c=33\)\+0\.0±0\.2\+0\.0\{\\scriptstyle\\pm 0\.2\}1\.0000 \(b=12, c=12\)–C35\.7±2\.435\.7\{\\scriptstyle\\pm 2\.4\}9\.2±1\.69\.2\{\\scriptstyle\\pm 1\.6\}34\.6±1\.534\.6\{\\scriptstyle\\pm 1\.5\}\+26\.5±2\.4\+26\.5\{\\scriptstyle\\pm 2\.4\}0\.0000∗∗∗\(b=371, c=22\)\+1\.1±0\.7\+1\.1\{\\scriptstyle\\pm 0\.7\}0\.0017∗∗\(b=121, c=76\)–Cu30\.8±2\.530\.8\{\\scriptstyle\\pm 2\.5\}9\.9±1\.69\.9\{\\scriptstyle\\pm 1\.6\}28\.9±1\.428\.9\{\\scriptstyle\\pm 1\.4\}\+20\.8±2\.6\+20\.8\{\\scriptstyle\\pm 2\.6\}0\.0000∗∗∗\(b=325, c=50\)\+1\.9±0\.9\+1\.9\{\\scriptstyle\\pm 0\.9\}0\.0000∗∗∗\(b=171, c=95\)–CODI41\.8±2\.741\.8\{\\scriptstyle\\pm 2\.7\}35\.7±2\.535\.7\{\\scriptstyle\\pm 2\.5\}41\.6±1\.541\.6\{\\scriptstyle\\pm 1\.5\}\+6\.1±1\.7\+6\.1\{\\scriptstyle\\pm 1\.7\}0\.0000∗∗∗\(b=104, c=23\)\+0\.2±0\.4\+0\.2\{\\scriptstyle\\pm 0\.4\}0\.3557 \(b=42, c=33\)–Table 10:Full gradient\-subspace ablation results with statistical testing\.Graph\-HoppingArithmetic\-ReasoningModelα\\alphaFlipgrad\{\}\_\{\\text\{grad\}\}\[% CI\]Fliprand\{\}\_\{\\text\{rand\}\}\[% CI\]Δ\\Delta\[% CI\]McNemarppnnFlipgrad\{\}\_\{\\text\{grad\}\}\[% CI\]Fliprand\{\}\_\{\\text\{rand\}\}\[% CI\]Δ\\Delta\[% CI\]McNemarppnnPaT1\.50\.0±0\.00\.0\{\\scriptstyle\\pm 0\.0\}0\.0±0\.00\.0\{\\scriptstyle\\pm 0\.0\}\+0\.0±0\.0\+0\.0\{\\scriptstyle\\pm 0\.0\}1\.0000 \(b=0, c=0\)29\.4±1\.79\.4\{\\scriptstyle\\pm 1\.7\}1\.5±0\.41\.5\{\\scriptstyle\\pm 0\.4\}\+7\.9±0\.8\+7\.9\{\\scriptstyle\\pm 0\.8\}0\.0000∗∗∗\(b=312, c=1\)56020\.2±0\.30\.2\{\\scriptstyle\\pm 0\.3\}0\.0±0\.00\.0\{\\scriptstyle\\pm 0\.0\}\+0\.2±0\.2\+0\.2\{\\scriptstyle\\pm 0\.2\}0\.2500 \(b=3, c=0\)15\.6±2\.015\.6\{\\scriptstyle\\pm 2\.0\}3\.8±0\.63\.8\{\\scriptstyle\\pm 0\.6\}\+11\.8±1\.0\+11\.8\{\\scriptstyle\\pm 1\.0\}0\.0000∗∗∗\(b=479, c=13\)51\.0±0\.91\.0\{\\scriptstyle\\pm 0\.9\}0\.1±0\.10\.1\{\\scriptstyle\\pm 0\.1\}\+0\.9±0\.5\+0\.9\{\\scriptstyle\\pm 0\.5\}0\.0005∗∗∗\(b=15, c=1\)43\.2±2\.843\.2\{\\scriptstyle\\pm 2\.8\}14\.1±1\.114\.1\{\\scriptstyle\\pm 1\.1\}\+29\.2±1\.5\+29\.2\{\\scriptstyle\\pm 1\.5\}0\.0000∗∗∗\(b=1268, c=114\)102\.0±1\.32\.0\{\\scriptstyle\\pm 1\.3\}0\.2±0\.20\.2\{\\scriptstyle\\pm 0\.2\}\+1\.8±0\.8\+1\.8\{\\scriptstyle\\pm 0\.8\}0\.0000∗∗∗\(b=30, c=3\)52\.5±2\.752\.5\{\\scriptstyle\\pm 2\.7\}28\.2±1\.428\.2\{\\scriptstyle\\pm 1\.4\}\+24\.4±1\.6\+24\.4\{\\scriptstyle\\pm 1\.6\}0\.0000∗∗∗\(b=1121, c=157\)2513\.0±2\.813\.0\{\\scriptstyle\\pm 2\.8\}1\.0±0\.51\.0\{\\scriptstyle\\pm 0\.5\}\+12\.0±1\.7\+12\.0\{\\scriptstyle\\pm 1\.7\}0\.0000∗∗∗\(b=183, c=3\)55\.0±2\.755\.0\{\\scriptstyle\\pm 2\.7\}47\.7±1\.647\.7\{\\scriptstyle\\pm 1\.6\}\+7\.4±1\.2\+7\.4\{\\scriptstyle\\pm 1\.2\}0\.0000∗∗∗\(b=478, c=186\)5029\.6±4\.029\.6\{\\scriptstyle\\pm 4\.0\}1\.0±0\.51\.0\{\\scriptstyle\\pm 0\.5\}\+28\.6±2\.4\+28\.6\{\\scriptstyle\\pm 2\.4\}0\.0000∗∗∗\(b=432, c=3\)56\.3±2\.856\.3\{\\scriptstyle\\pm 2\.8\}48\.2±1\.648\.2\{\\scriptstyle\\pm 1\.6\}\+8\.1±1\.3\+8\.1\{\\scriptstyle\\pm 1\.3\}0\.0000∗∗∗\(b=520, c=200\)10049\.4±4\.249\.4\{\\scriptstyle\\pm 4\.2\}1\.1±0\.51\.1\{\\scriptstyle\\pm 0\.5\}\+48\.3±2\.7\+48\.3\{\\scriptstyle\\pm 2\.7\}0\.0000∗∗∗\(b=727, c=3\)57\.4±2\.857\.4\{\\scriptstyle\\pm 2\.8\}48\.1±1\.548\.1\{\\scriptstyle\\pm 1\.5\}\+9\.3±1\.3\+9\.3\{\\scriptstyle\\pm 1\.3\}0\.0000∗∗∗\(b=579, c=210\)C1\.50\.0±0\.00\.0\{\\scriptstyle\\pm 0\.0\}0\.0±0\.00\.0\{\\scriptstyle\\pm 0\.0\}\+0\.0±0\.0\+0\.0\{\\scriptstyle\\pm 0\.0\}1\.0000 \(b=0, c=0\)None16\.1±2\.016\.1\{\\scriptstyle\\pm 2\.0\}10\.7±1\.010\.7\{\\scriptstyle\\pm 1\.0\}\+5\.4±1\.1\+5\.4\{\\scriptstyle\\pm 1\.1\}0\.0000∗∗∗\(b=404, c=190\)105920\.0±0\.00\.0\{\\scriptstyle\\pm 0\.0\}0\.0±0\.00\.0\{\\scriptstyle\\pm 0\.0\}\+0\.0±0\.0\+0\.0\{\\scriptstyle\\pm 0\.0\}1\.0000 \(b=0, c=0\)25\.6±2\.425\.6\{\\scriptstyle\\pm 2\.4\}20\.4±1\.320\.4\{\\scriptstyle\\pm 1\.3\}\+5\.2±1\.4\+5\.2\{\\scriptstyle\\pm 1\.4\}0\.0000∗∗∗\(b=524, c=319\)50\.0±0\.00\.0\{\\scriptstyle\\pm 0\.0\}0\.0±0\.00\.0\{\\scriptstyle\\pm 0\.0\}\+0\.0±0\.0\+0\.0\{\\scriptstyle\\pm 0\.0\}1\.0000 \(b=0, c=0\)48\.2±2\.748\.2\{\\scriptstyle\\pm 2\.7\}60\.2±1\.560\.2\{\\scriptstyle\\pm 1\.5\}−12\.0±1\.6\-12\.0\{\\scriptstyle\\pm 1\.6\}0\.0000∗∗∗\(b=324, c=799\)100\.2±0\.30\.2\{\\scriptstyle\\pm 0\.3\}0\.0±0\.00\.0\{\\scriptstyle\\pm 0\.0\}\+0\.2±0\.2\+0\.2\{\\scriptstyle\\pm 0\.2\}0\.2500 \(b=3, c=0\)65\.7±2\.565\.7\{\\scriptstyle\\pm 2\.5\}79\.9±1\.279\.9\{\\scriptstyle\\pm 1\.2\}−14\.1±1\.5\-14\.1\{\\scriptstyle\\pm 1\.5\}0\.0000∗∗∗\(b=245, c=804\)252\.4±1\.32\.4\{\\scriptstyle\\pm 1\.3\}0\.3±0\.30\.3\{\\scriptstyle\\pm 0\.3\}\+2\.1±0\.7\+2\.1\{\\scriptstyle\\pm 0\.7\}0\.0000∗∗∗\(b=32, c=1\)82\.3±1\.982\.3\{\\scriptstyle\\pm 1\.9\}84\.7±1\.184\.7\{\\scriptstyle\\pm 1\.1\}−2\.4±1\.2\-2\.4\{\\scriptstyle\\pm 1\.2\}0\.0003∗∗∗\(b=297, c=392\)5017\.2±3\.217\.2\{\\scriptstyle\\pm 3\.2\}0\.3±0\.20\.3\{\\scriptstyle\\pm 0\.2\}\+16\.9±2\.0\+16\.9\{\\scriptstyle\\pm 2\.0\}0\.0000∗∗∗\(b=256, c=2\)84\.7±2\.084\.7\{\\scriptstyle\\pm 2\.0\}84\.8±1\.184\.8\{\\scriptstyle\\pm 1\.1\}−0\.2±1\.3\-0\.2\{\\scriptstyle\\pm 1\.3\}0\.8475 \(b=335, c=341\)10034\.4±4\.234\.4\{\\scriptstyle\\pm 4\.2\}1\.2±0\.51\.2\{\\scriptstyle\\pm 0\.5\}\+33\.2±2\.4\+33\.2\{\\scriptstyle\\pm 2\.4\}0\.0000∗∗∗\(b=503, c=5\)85\.7±1\.985\.7\{\\scriptstyle\\pm 1\.9\}85\.4±1\.185\.4\{\\scriptstyle\\pm 1\.1\}\+0\.4±1\.3\+0\.4\{\\scriptstyle\\pm 1\.3\}0\.5906 \(b=346, c=331\)Cu1\.50\.2±0\.30\.2\{\\scriptstyle\\pm 0\.3\}0\.0±0\.00\.0\{\\scriptstyle\\pm 0\.0\}\+0\.2±0\.2\+0\.2\{\\scriptstyle\\pm 0\.2\}0\.2500 \(b=3, c=0\)220\.7±2\.220\.7\{\\scriptstyle\\pm 2\.2\}16\.6±1\.216\.6\{\\scriptstyle\\pm 1\.2\}\+4\.1±1\.3\+4\.1\{\\scriptstyle\\pm 1\.3\}0\.0000∗∗∗\(b=439, c=278\)122220\.2±0\.30\.2\{\\scriptstyle\\pm 0\.3\}0\.0±0\.00\.0\{\\scriptstyle\\pm 0\.0\}\+0\.2±0\.2\+0\.2\{\\scriptstyle\\pm 0\.2\}0\.2500 \(b=3, c=0\)31\.2±2\.531\.2\{\\scriptstyle\\pm 2\.5\}28\.8±1\.528\.8\{\\scriptstyle\\pm 1\.5\}\+2\.4±1\.6\+2\.4\{\\scriptstyle\\pm 1\.6\}0\.0033∗∗\(b=560, c=465\)52\.0±1\.22\.0\{\\scriptstyle\\pm 1\.2\}39\.7±2\.539\.7\{\\scriptstyle\\pm 2\.5\}−37\.7±2\.6\-37\.7\{\\scriptstyle\\pm 2\.6\}0\.0000∗∗∗\(b=20, c=585\)57\.7±2\.657\.7\{\\scriptstyle\\pm 2\.6\}78\.1±1\.378\.1\{\\scriptstyle\\pm 1\.3\}−20\.4±1\.7\-20\.4\{\\scriptstyle\\pm 1\.7\}0\.0000∗∗∗\(b=251, c=1057\)1043\.6±4\.243\.6\{\\scriptstyle\\pm 4\.2\}100\.0±0\.0100\.0\{\\scriptstyle\\pm 0\.0\}−56\.4±2\.3\-56\.4\{\\scriptstyle\\pm 2\.3\}0\.0000∗∗∗\(b=0, c=846\)69\.2±2\.569\.2\{\\scriptstyle\\pm 2\.5\}93\.9±0\.793\.9\{\\scriptstyle\\pm 0\.7\}−24\.7±1\.5\-24\.7\{\\scriptstyle\\pm 1\.5\}0\.0000∗∗∗\(b=66, c=1042\)2598\.8±0\.998\.8\{\\scriptstyle\\pm 0\.9\}100\.0±0\.0100\.0\{\\scriptstyle\\pm 0\.0\}−1\.2±0\.5\-1\.2\{\\scriptstyle\\pm 0\.5\}0\.0000∗∗∗\(b=0, c=18\)79\.5±2\.279\.5\{\\scriptstyle\\pm 2\.2\}99\.2±0\.399\.2\{\\scriptstyle\\pm 0\.3\}−19\.8±1\.2\-19\.8\{\\scriptstyle\\pm 1\.2\}0\.0000∗∗∗\(b=14, c=797\)50100\.0±0\.0100\.0\{\\scriptstyle\\pm 0\.0\}100\.0±0\.0100\.0\{\\scriptstyle\\pm 0\.0\}\+0\.0±0\.0\+0\.0\{\\scriptstyle\\pm 0\.0\}1\.0000 \(b=0, c=0\)88\.3±1\.788\.3\{\\scriptstyle\\pm 1\.7\}99\.6±0\.299\.6\{\\scriptstyle\\pm 0\.2\}−11\.3±1\.0\-11\.3\{\\scriptstyle\\pm 1\.0\}0\.0000∗∗∗\(b=9, c=456\)100100\.0±0\.0100\.0\{\\scriptstyle\\pm 0\.0\}99\.7±0\.299\.7\{\\scriptstyle\\pm 0\.2\}\+0\.3±0\.2\+0\.3\{\\scriptstyle\\pm 0\.2\}0\.1250 \(b=4, c=0\)92\.5±1\.492\.5\{\\scriptstyle\\pm 1\.4\}99\.7±0\.299\.7\{\\scriptstyle\\pm 0\.2\}−7\.2±0\.8\-7\.2\{\\scriptstyle\\pm 0\.8\}0\.0000∗∗∗\(b=7, c=291\)CODI1\.50\.4±0\.50\.4\{\\scriptstyle\\pm 0\.5\}0\.7±0\.40\.7\{\\scriptstyle\\pm 0\.4\}−0\.3±0\.3\-0\.3\{\\scriptstyle\\pm 0\.3\}0\.1250 \(b=1, c=6\)417\.7±2\.217\.7\{\\scriptstyle\\pm 2\.2\}7\.6±0\.87\.6\{\\scriptstyle\\pm 0\.8\}\+10\.2±1\.1\+10\.2\{\\scriptstyle\\pm 1\.1\}0\.0000∗∗∗\(b=502, c=99\)46220\.6±0\.70\.6\{\\scriptstyle\\pm 0\.7\}0\.7±0\.40\.7\{\\scriptstyle\\pm 0\.4\}−0\.1±0\.4\-0\.1\{\\scriptstyle\\pm 0\.4\}1\.0000 \(b=5, c=6\)24\.6±2\.324\.6\{\\scriptstyle\\pm 2\.3\}10\.7±1\.010\.7\{\\scriptstyle\\pm 1\.0\}\+13\.8±1\.3\+13\.8\{\\scriptstyle\\pm 1\.3\}0\.0000∗∗∗\(b=673, c=125\)50\.6±0\.70\.6\{\\scriptstyle\\pm 0\.7\}0\.5±0\.40\.5\{\\scriptstyle\\pm 0\.4\}\+0\.1±0\.4\+0\.1\{\\scriptstyle\\pm 0\.4\}1\.0000 \(b=6, c=5\)34\.3±2\.534\.3\{\\scriptstyle\\pm 2\.5\}30\.5±1\.430\.5\{\\scriptstyle\\pm 1\.4\}\+3\.9±1\.6\+3\.9\{\\scriptstyle\\pm 1\.6\}0\.0000∗∗∗\(b=554, c=401\)101\.2±0\.91\.2\{\\scriptstyle\\pm 0\.9\}0\.7±0\.40\.7\{\\scriptstyle\\pm 0\.4\}\+0\.5±0\.4\+0\.5\{\\scriptstyle\\pm 0\.4\}0\.0386∗\(b=10, c=2\)37\.5±2\.637\.5\{\\scriptstyle\\pm 2\.6\}58\.2±1\.558\.2\{\\scriptstyle\\pm 1\.5\}−20\.7±1\.8\-20\.7\{\\scriptstyle\\pm 1\.8\}0\.0000∗∗∗\(b=253, c=1074\)250\.8±0\.70\.8\{\\scriptstyle\\pm 0\.7\}0\.7±0\.40\.7\{\\scriptstyle\\pm 0\.4\}\+0\.1±0\.3\+0\.1\{\\scriptstyle\\pm 0\.3\}1\.0000 \(b=3, c=2\)39\.2±2\.639\.2\{\\scriptstyle\\pm 2\.6\}90\.8±0\.990\.8\{\\scriptstyle\\pm 0\.9\}−51\.6±1\.6\-51\.6\{\\scriptstyle\\pm 1\.6\}0\.0000∗∗∗\(b=48, c=2091\)500\.8±0\.70\.8\{\\scriptstyle\\pm 0\.7\}1\.0±0\.51\.0\{\\scriptstyle\\pm 0\.5\}−0\.2±0\.4\-0\.2\{\\scriptstyle\\pm 0\.4\}0\.5078 \(b=3, c=6\)39\.9±2\.739\.9\{\\scriptstyle\\pm 2\.7\}94\.0±0\.794\.0\{\\scriptstyle\\pm 0\.7\}−54\.1±1\.6\-54\.1\{\\scriptstyle\\pm 1\.6\}0\.0000∗∗∗\(b=41, c=2182\)1001\.2±0\.91\.2\{\\scriptstyle\\pm 0\.9\}0\.9±0\.50\.9\{\\scriptstyle\\pm 0\.5\}\+0\.3±0\.4\+0\.3\{\\scriptstyle\\pm 0\.4\}0\.2891 \(b=6, c=2\)40\.0±2\.740\.0\{\\scriptstyle\\pm 2\.7\}94\.4±0\.794\.4\{\\scriptstyle\\pm 0\.7\}−54\.4±1\.6\-54\.4\{\\scriptstyle\\pm 1\.6\}0\.0000∗∗∗\(b=38, c=2192\)Table 11:Full gradient\-subspace amplification results with statistical testing\.Table[7](https://arxiv.org/html/2606.12689#A2.T7)reports results for the latent thought ablation at test\-time experiment; Tables[8](https://arxiv.org/html/2606.12689#A2.T8)&[9](https://arxiv.org/html/2606.12689#A2.T9)\(prompt and thought\-positions respectively\) for the causal\-tracing experiment; and Tables[10](https://arxiv.org/html/2606.12689#A2.T10)&[11](https://arxiv.org/html/2606.12689#A2.T11)\(ablation and amplification respectively\) for the gradient\-subspace interventions experiment\.

#### The Dynamics and Geometry of Latent Thoughts \(§[6](https://arxiv.org/html/2606.12689#S6)\)\.

Graph\-HoppingArithmetic\-ReasoningProj\.ModelIdentityMeanLinearMLPIdentityMeanLinearMLPFullPaT0\.999 \[0\.999, 0\.999\]\-0\.0030\.597 \[0\.595, 0\.599\]0\.515 \[0\.501, 0\.528\]0\.867 \[0\.861, 0\.873\]\-0\.0130\.716 \[0\.712, 0\.720\]0\.697 \[0\.695, 0\.699\]C0\.991 \[0\.990, 0\.991\]\-0\.0020\.649 \[0\.647, 0\.651\]0\.412 \[0\.398, 0\.424\]\-0\.663 \[\-0\.676, \-0\.653\]\-0\.0210\.253 \[0\.251, 0\.256\]0\.412 \[0\.410, 0\.414\]Cu\-0\.560 \[\-0\.588, \-0\.533\]\-0\.0000\.734 \[0\.730, 0\.738\]0\.783 \[0\.780, 0\.785\]\-0\.719 \[\-0\.732, \-0\.707\]\-0\.0310\.253 \[0\.250, 0\.257\]0\.427 \[0\.424, 0\.429\]CODI0\.810 \[0\.797, 0\.823\]\-0\.0020\.623 \[0\.619, 0\.625\]0\.889 \[0\.889, 0\.890\]\-0\.683 \[\-0\.693, \-0\.675\]\-0\.0230\.394 \[0\.389, 0\.399\]0\.481 \[0\.478, 0\.484\]SubspacePaT\-0\.466 \[\-0\.485, \-0\.452\]\-0\.0000\.971 \[0\.970, 0\.971\]0\.996 \[0\.996, 0\.996\]0\.264 \[0\.257, 0\.270\]\-0\.0090\.605 \[0\.602, 0\.609\]0\.780 \[0\.778, 0\.783\]C0\.775 \[0\.764, 0\.786\]\-0\.0010\.686 \[0\.681, 0\.690\]0\.976 \[0\.975, 0\.977\]\-0\.893 \[\-0\.906, \-0\.880\]\-0\.0180\.250 \[0\.247, 0\.253\]0\.441 \[0\.438, 0\.444\]Cu\-0\.758 \[\-0\.796, \-0\.728\]\-0\.0000\.921 \[0\.919, 0\.923\]0\.968 \[0\.968, 0\.969\]\-0\.932 \[\-0\.945, \-0\.920\]\-0\.0300\.281 \[0\.278, 0\.284\]0\.475 \[0\.472, 0\.478\]CODI0\.296 \[0\.270, 0\.322\]\-0\.0030\.771 \[0\.766, 0\.776\]0\.974 \[0\.973, 0\.975\]\-1\.076 \[\-1\.105, \-1\.050\]\-0\.0320\.236 \[0\.231, 0\.240\]0\.341 \[0\.337, 0\.346\]Table 12:Full results for Markovianity of thought trajectories with statistical testing\.Table[12](https://arxiv.org/html/2606.12689#A2.T12)reports results for the Markovianity of thought trajectories experiment; and Table[3](https://arxiv.org/html/2606.12689#A1.T3)for the geometric stability of gradient\-subspaces experiment\.

## Appendix CModels, Controls, and Training Details

In this appendix, we provide details on the dataset, models, and training details\.

### C\.1Task and Dataset

#### Tasks\.

In this work, we use two case studies corresponding to prominent mechanistic claims about latent reasoning models\.

Prior work has interpreted latent trajectories as breadth\-first\-search\-like reasoning over graph frontiers in the graph\-hopping task\(Haoet al\.,[2025](https://arxiv.org/html/2606.12689#bib.bib1); Zhuet al\.,[2025](https://arxiv.org/html/2606.12689#bib.bib25)\)\. More recent works have investigated this claim more carefully, questioning its validity\(Cuiet al\.,[2026](https://arxiv.org/html/2606.12689#bib.bib22); Rizvi\-Martelet al\.,[2026](https://arxiv.org/html/2606.12689#bib.bib24)\)\. All these works used ProsQA\(Haoet al\.,[2025](https://arxiv.org/html/2606.12689#bib.bib1)\)as the benchmark for this investigation, which makes it a natural choice for our study\. Each ProsQA instance is a directed acyclic graph of logical relations paired with a query and a gold target node; the dataset comprises17,88617\{,\}886training,300300validation, and500500test instances\.

Prior work has interpreted decodable intermediate values as evidence that latent thoughts act as a compressed thought scratchpad in arithmetic reasoning\(Shenet al\.,[2025](https://arxiv.org/html/2606.12689#bib.bib9)\)\. More recent works also question the validity of this claim, raising the possibility of shortcut behavior or task\-specific heuristics\(Cuiet al\.,[2026](https://arxiv.org/html/2606.12689#bib.bib22); Liang and Pan,[2026](https://arxiv.org/html/2606.12689#bib.bib26); Dilgren and Wiegreffe,[2026](https://arxiv.org/html/2606.12689#bib.bib36); Weiet al\.,[2026](https://arxiv.org/html/2606.12689#bib.bib35)\)\. All these works used GSM8k\(Cobbeet al\.,[2021](https://arxiv.org/html/2606.12689#bib.bib8)\)as one of the benchmarks for this investigation, which makes it a natural choice for our study\. Each GSM8k instance is a grade\-school arithmetic word problem paired with a step\-by\-step solution; we use the GSM8K\-Aug variant ofDenget al\.\([2023](https://arxiv.org/html/2606.12689#bib.bib16)\), which augments each problem with a compressed chain\-of\-thought decomposition, giving385,620385\{,\}620training,500500validation, and1,3191\{,\}319test instances\.

Together, the two tasks \(graph traversal and arithmetic reasoning\) allow us to evaluate whether observational patterns of latent reasoning are robust across qualitatively different reasoning settings\.

#### Training data\.

All models are trained separately for each task\. FollowingHaoet al\.\([2025](https://arxiv.org/html/2606.12689#bib.bib1)\)andShenet al\.\([2025](https://arxiv.org/html/2606.12689#bib.bib9)\), we train on the GSM8K\-Aug training data ofDenget al\.\([2023](https://arxiv.org/html/2606.12689#bib.bib16)\)for GSM8k and on the original ProsQA training split\(Haoet al\.,[2025](https://arxiv.org/html/2606.12689#bib.bib1)\)for ProsQA\. The test splits \(500500for ProsQA,1,3191\{,\}319for GSM8k\) are held out and used for all reported results\.

### C\.2Role of Each Model and Control

Table[13](https://arxiv.org/html/2606.12689#A3.T13)summarizes the role of each model in our design\. The controls are not used only as task\-performance baselines: each removes or perturbs one factor that could otherwise make an observable latent\-state pattern look like evidence for a latent reasoning mechanism\.

ModelRoleQuestion addressedCoconut \(C\)Target LRMDoes staged recurrent latent computation exhibit the BFS\-like graph\-frontier patterns attributed to Coconut?CODITarget LRMDo distilled latent rationale states contain decodable arithmetic intermediates that are also behaviorally relevant?PaTRecurrence controlDo the same patterns persist when Coconut’s format and curriculum are kept but recurrent latent\-state feedback is removed?Coconutu\(Cu\)Curriculum controlDo the patterns survive a controlled perturbation of Coconut’s staged curriculum while the recurrent architecture is preserved?Base GPT\-2 \(B\)Observational controlCan the probes recover apparent structure without task\-specific latent\-thought training?Explicit\-CoT GPT\-2 \(CoT\)Observational controlCan similar patterns arise from ordinary task solving or explicit\-rationale supervision rather than latent\-thought mechanisms?Table 13:Role of each model and control in the experimental design\.#### Target latent reasoning models\.

Coconut\(Haoet al\.,[2025](https://arxiv.org/html/2606.12689#bib.bib1)\)and CODI\(Shenet al\.,[2025](https://arxiv.org/html/2606.12689#bib.bib9)\)are the target LRMs because they instantiate the two mechanistic claims studied in this work: BFS\-like latent search on graph\-hopping and scratchpad\-like latent arithmetic\. Both insert dedicated latent thought positions, but train them differently\. Coconut gradually replaces explicit reasoning segments with recurrent latent states through a staged curriculum\. CODI compresses textual rationales into latent representations by distillation\. Studying both lets us test whether our conclusions are tied to one training recipe or hold across two representative ways of constructing latent thought states\.

#### Pause\-as\-thought as a recurrence control\.

PaT is designed to isolate recurrent latent\-state feedback\. It follows Coconut’s input format and staged curriculum, inspired by pause\-token training\(Goyalet al\.,[2024](https://arxiv.org/html/2606.12689#bib.bib10)\), but replaces feedback from previous hidden states with learned thought\-position embeddings\. Thus, PaT preserves the thought slots, task formatting, and curriculum pressures that may shape observable representations, while removing the mechanism by which Coconut feeds latent states back into subsequent latent computation\. If PaT reproduces Coconut\-like patterns, those patterns are not sufficient evidence that recurrence is the cause\.

#### Coconutuas a curriculum control\.

Coconutupreserves Coconut’s recurrent latent architecture and number of thought positions, but perturbs the order in which curriculum stages are sampled\. At each training step, the model follows the current stage with probability1−u1\-uand samples a non\-current stage with probabilityuu\. We setu=0\.3u=0\.3to create a moderate perturbation: the model still trains mostly on the intended stage \(70%70\\%of updates\), keeping the comparison close to Coconut, while the remaining updates are frequent enough to test whether the reported patterns depend on a nearly deterministic stage schedule\. This control separates effects of recurrence from effects of the curriculum trajectory\.

#### Base GPT\-2 and Explicit\-CoT GPT\-2 as observational controls\.

Base GPT\-2 and Explicit\-CoT GPT\-2 are included only in the observational analyses of §[4](https://arxiv.org/html/2606.12689#S4)\. Base GPT\-2 tests whether the readout methods themselves can recover apparently interpretable structure from a model without task\-specific latent\-thought training\. Explicit\-CoT GPT\-2 tests whether similar structure appears after ordinary explicit\-rationale supervision\. These controls assess the specificity of observational readouts\. They are excluded from causal and geometric analyses because those analyses intervene on dedicated latent thought positions, which these models are not trained to use\.

### C\.3Training Details

Table[14](https://arxiv.org/html/2606.12689#A3.T14)reports the training configuration for the explicit\-CoT baseline and the Coconut\-family models: PaT, Coconut, and Coconutu\. Table[15](https://arxiv.org/html/2606.12689#A3.T15)then reports the CODI configuration\. CODI uses a different training recipe and is therefore listed separately\.

For the baseline table,ccdenotes the number of thought tokens per reasoning step, and batch size is reported per GPU\. Coconut and Coconutushare the same hyperparameters for a given dataset; Coconutudiffers only by using stage\-mixing probabilityu=0\.3u=0\.3, while standard Coconut usesu=0\.0u=0\.0\. All runs in Table[14](https://arxiv.org/html/2606.12689#A3.T14)use GPT\-2 initialization, bf16 disabled, and weight decay0\.010\.01\.

The CODI hyperparameters in Table[15](https://arxiv.org/html/2606.12689#A3.T15)use LoRA fine\-tuning with a projection head, following the configuration ofShenet al\.\([2025](https://arxiv.org/html/2606.12689#bib.bib9)\)\. The ProsQA CODI settings follow the configuration ofCuiet al\.\([2026](https://arxiv.org/html/2606.12689#bib.bib22)\)\.

ModelDatasetccEps\./stageMax stageuuResumeBatchGrad\. acc\.EpochsLRCoTGSM8K\-Aug0100\.001281251×10−41\\times 10^\{\-4\}CoTProsQA0100\.001281501×10−41\\times 10^\{\-4\}PaTGSM8K\-Aug2330\.031281251×10−41\\times 10^\{\-4\}PaTProsQA1560\.001281501×10−41\\times 10^\{\-4\}CoconutGSM8K\-Aug2330\.0/0\.331281251×10−41\\times 10^\{\-4\}CoconutProsQA1560\.0/0\.301282501×10−41\\times 10^\{\-4\}Table 14:Hyperparameters used to train the GPT\-2 CoT, PaT, Coconut, and Coconutumodels\. Coconut rows report both the standard setting \(u=0\.0u=0\.0\) and the stage\-mixed Coconutusetting \(u=0\.3u=0\.3\)\.HyperparameterGSM8K\-AugProsQAModel max length10241024Precisionbf16bf16CODI loss weight1\.01\.0Include last CoTFalseFalseNumber of latents66Use projectionTrueTrueProjection dimension768768Projection dropout0\.00\.0Use LoRATrueTrueLoRA rank128128LoRA alpha3232Learning rate3×10−33\\times 10^\{\-3\}1×10−31\\times 10^\{\-3\}LR schedulerCosineCosineWarmup ratio0\.030\.03OptimizerAdamWAdamWTotal Batch size128128Weight decay0\.10\.01Gradient clipping2\.02\.0Epochs4020Table 15:Hyperparameters used to train the GPT\-2 CODI models\.#### Model size\.

All GPT\-2 models are built on the pretrained GPT\-2 small \(124124M parameters;Radfordet al\.,[2019](https://arxiv.org/html/2606.12689#bib.bib11)\)\. CODI uses LoRA fine\-tuning \(rank128128\) on these backbones; all other models are fully fine\-tuned\. Full per\-model, per\-task hyperparameters appear in Tables[14](https://arxiv.org/html/2606.12689#A3.T14)and[15](https://arxiv.org/html/2606.12689#A3.T15)\.

Similar Articles

Are Latent Reasoning Models Easily Interpretable?

Lobsters Hottest

The paper investigates the interpretability of latent reasoning models, finding that reasoning tokens are often unnecessary but can be decoded to reveal interpretable traces when needed, suggesting these models implement expected solutions.

Large Reasoning Models Are (Not Yet) Multilingual Latent Reasoners

arXiv cs.CL

This paper investigates multilingual latent reasoning in large reasoning models across 11 languages, revealing that while latent reasoning capabilities exist, they are unevenly distributed—stronger in resource-rich languages and weaker in low-resource ones. The study finds that despite surface-level differences, the internal reasoning mechanisms are largely aligned with an English-centered pathway.