Final Checkpoints Are Not Enough: Analyzing Latent Reasoning Faithfulness Along Training Trajectories

arXiv cs.LG Papers

Summary

This paper investigates how the faithfulness of latent reasoning steps evolves during training, finding that it depends on training stage and answer format, rather than just final checkpoint performance.

arXiv:2607.06648v1 Announce Type: new Abstract: Latent reasoning methods perform multi-step inference entirely in the model's continuous hidden states, promising more compact and efficient reasoning. However, these opaque hidden states raise a question of faithfulness: whether these latent reasoning steps causally drive the final answer. Prior work investigates this question at converged checkpoints and reports several unfaithful behaviors, such as latent reasoning steps that can be replaced without changing the answer, but leaves how these behaviors form during training unexamined. We instead track how faithfulness evolves across saved checkpoints for different latent reasoning paradigms, applying a verifiable counterfactual edit on the input and a noise-ablation activation patch on the latent reasoning steps. We find that (i) at the output level, latent reasoning methods can look similarly unfaithful at convergence under counterfactual edits while following qualitatively divergent trajectories; (ii) at the activation level, the causal contribution of latent reasoning steps to the final answer decays across training for both paradigms, with the examples that flip on the output side in (i) also being the examples on which this contribution decays; and (iii) the activation-level trajectory diverges by answer format, decaying on binary choice and rising on open-ended decoding. These findings highlight that latent reasoning faithfulness depends on training stage and answer format.
Original Article
View Cached Full Text

Cached at: 07/09/26, 07:42 AM

# Final Checkpoints Are Not Enough: Analyzing Latent Reasoning Faithfulness Along Training Trajectories
Source: [https://arxiv.org/html/2607.06648](https://arxiv.org/html/2607.06648)
Hengyu Jin1,2,3,Shu Yang2,3,Di Wang2,3 1Tongji University 2Provable Responsible AI and Data Analytics \(PRADA\) Lab 3King Abdullah University of Science and Technology

###### Abstract

Latent reasoning methods perform multi\-step inference entirely in the model’s continuous hidden states, promising more compact and efficient reasoning\. However, these opaque hidden states raise a question of faithfulness: whether these latent reasoning steps causally drive the final answer\. Prior work investigates this question at converged checkpoints and reports several unfaithful behaviors, such as latent reasoning steps that can be replaced without changing the answer, but leaves how these behaviors form during training unexamined\. We instead track how faithfulness evolves across saved checkpoints for different latent reasoning paradigms, applying a verifiable counterfactual edit on the input and a noise\-ablation activation patch on the latent reasoning steps\. We find that\(i\)at the output level, latent reasoning methods can look similarly unfaithful at convergence under counterfactual edits while following qualitatively divergent trajectories;\(ii\)at the activation level, the causal contribution of latent reasoning steps to the final answer decays across training for both paradigms, with the examples that flip on the output side in \(i\) also being the examples on which this contribution decays; and\(iii\)the activation\-level trajectory diverges by answer format, decaying on binary choice and rising on open\-ended decoding\. These findings highlight that latent reasoning faithfulness depends on training stage and answer format\.

Final Checkpoints Are Not Enough: Analyzing Latent Reasoning Faithfulness Along Training Trajectories

Hengyu Jin1,2,3, Shu Yang2,3, Di Wang2,3††thanks:Corresponding author\.1Tongji University2Provable Responsible AI and Data Analytics \(PRADA\) Lab3King Abdullah University of Science and Technology

## 1Introduction

![Refer to caption](https://arxiv.org/html/2607.06648v1/x1.png)Figure 1:Analysis along the training trajectory\.\(a\)Counterfactual edit on ProsQA and noise\-ablation patch on all latent reasoning steps, each evaluated across saved checkpoints\.\(b–d\)Output\-level \(RQ1\), activation\-level \(RQ2\), and cross\-format \(RQ3\) trajectories under these analyses\.Chain\-of\-thought reasoning represents intermediate computation as natural\-language tokens, improving multi\-step reasoning by exposing a step\-by\-step trace\(wei2022\_cot;nye2021\_scratchpad\)\. This representation comes at a cost: every reasoning step must be committed to discrete textual tokens, and the resulting traces lengthen inference\.*Latent reasoning*methods instead perform multi\-step inference entirely in the model’s continuous hidden states, promising more compact and efficient reasoning by removing the explicit trace that chain\-of\-thought requires\(hao2025\_COCONUT;shen2025\_codi\)\.

This opacity raises a question of faithfulness analogous to that studied in explicit CoT, where generated rationales need not reflect the underlying computation\(turpin2023\_unfaithful;lanham2023\_measuring\_faithfulness\)\. Latent reasoning produces no such rationale to inspect, so the question is reframed as whether the latent reasoning steps themselves causally drive the final answer:do interventions that change or erase these steps actually change the prediction?Prior work investigates this question at converged checkpoints and reports several unfaithful behaviors of trained latent reasoners, including latent reasoning steps whose hidden states can be replaced or perturbed without changing the final prediction, and evidence that the answer can be produced by alternative paths in the network that do not pass through these steps\(zhang2025\_do\_latent\_tokens\_think;lin2025\_implicit\_shortcut;cui2026\_supervision;li2026\_dynamics\)\.

These analyses share a methodological choice: each trained latent reasoner is treated as a fixed object and evaluated only at its final checkpoint\. This leaves how the reported unfaithful behaviors form during training unexamined, a question that is complex because latent reasoning training paradigms differ substantially in mechanism\. A final\-checkpoint observation also cannot tell whether different paradigms reach a similar endpoint through similar trajectories or through qualitatively different ones\. Treating a trained model as the endpoint of a training trajectory aligns with a broader line of work on model properties across training\(tirumala2022\_memorization\_dynamics;tigges2024\_circuits\_across\_training\), but has not yet been applied to latent reasoning faithfulness\.

We therefore measure faithfulness across saved checkpoints along training on a shared GPT\-2 backbone, applying the two interventions shown in Figure[1](https://arxiv.org/html/2607.06648#S1.F1)a\. The first, illustrated on the left, is a*verified counterfactual edit on the input*that we construct on ProsQA\(hao2025\_COCONUT\), a benchmark whose questions are graph reachability queries solvable by a deterministic BFS oracle: we edit a single edge of the underlying graph so that the BFS\-verified oracle answer flips from the original target to the alternative, yielding405405paired original/edited inputs that differ only in a single rewritten clause\. The second, on the right, is a noise\-ablation activation patch on the latent reasoning steps, adapted from causal\-mediation analyses of model components\(vig2020\_causal\_mediation;heimersheim2024\_patching\)\. We organise the analysis around three research questions that move from the output level to the activation level to the answer format\.

RQ1: Output\-level faithfulness behaviors along training\.Do latent reasoning methods with different training paradigms exhibit qualitatively different faithfulness behaviors at the output level along training, even when their final\-checkpoint unfaithfulness looks similar?On the405405verified ProsQA counterfactual pairs above, tracing the joint label composition of each pair across six saved checkpoints in Figure[1](https://arxiv.org/html/2607.06648#S1.F1)b shows that surface similarity at convergence can hide qualitatively divergent output\-level trajectories\.

RQ2: Activation\-level trajectory of latent reasoning steps\.Building on RQ1, we ask whether the output\-level trajectory has a corresponding activation\-level trajectory, that is,whether the latent reasoning steps lose their causal contribution to the answer on the same examples that drive the output\-level transition\.Following the patched\-versus\-clean output change rate across training in Figure[1](https://arxiv.org/html/2607.06648#S1.F1)c, we find that the causal contribution at the latent reasoning steps decays across training in both paradigms; a cohort analysis in §[4](https://arxiv.org/html/2607.06648#S4)further shows that the examples flipping on the output side in RQ1 are also those on which this contribution decays\.

RQ3: Cross\-format dependence of the activation\-level trajectory\.The trajectories in RQ1 and RQ2 are both observed on a binary\-choice setting\. RQ3 askswhether they reflect an intrinsic property of latent reasoning or depend on the answer format, by varying only the answer format under a shared content and training recipe\. Comparing the same trajectory across three answer formats in Figure[1](https://arxiv.org/html/2607.06648#S1.F1)d, we find that the activation\-level trajectory reverses direction when only the answer format changes: the causal contribution at the latent reasoning steps decays across training on binary choice but rises on open\-ended decoding\.

## 2Related Work

#### Latent reasoning methods\.

Latent reasoning replaces some or all discrete chain\-of\-thought steps\(wei2022\_cot;nye2021\_scratchpad\)with computation in continuous hidden\-state space\. COCONUT\(hao2025\_COCONUT\)feeds the model’s last hidden state back as the next input embedding under a staged curriculum, while CODI\(shen2025\_codi\)self\-distils an explicit\-CoT teacher into fixed\-length continuous thoughts; related variants add non\-verbal computation through pause tokens, recurrent depth, or implicit\-CoT training\(goyal2024\_pause;geiping2025\_recurrent;deng2024\_implicit\_cot;chen2025\_survey\_latentcot\)\. We focus on COCONUT and CODI as two representative training\-paradigm choices \(staged curriculum vs\. self\-distillation\) and ask how the faithfulness profile of each evolves across training rather than at a single checkpoint\.

#### Faithfulness for latent reasoning\.

Faithfulness diagnoses for explicit chain\-of\-thought assess whether the generated rationale reflects the model’s underlying computation\(jacovi\-goldberg\-2020\-towards;turpin2023\_unfaithful;lanham2023\_measuring\_faithfulness;lyu\-etal\-2024\-towards\)\. Without an explicit trace to inspect, latent reasoning faithfulness is reframed as whether the continuous thought steps causally drive the final prediction\. Recent studies at the converged checkpoint suggest these continuous steps are often replaceable: ablation and perturbation on continuous thoughts largely preserve predictions\(zhang2025\_do\_latent\_tokens\_think\), models exploit shortcuts over intermediate steps\(lin2025\_implicit\_shortcut;wu2025\_single\_threaded;liang2026\_latentcot\_step\), and latent positions exhibit shallow causal structure and limited interpretability\(li2026\_dynamics;cui2026\_supervision;dilgren2026\_lrm\_interp\)\. These studies, however, evaluate models only at convergence\. Inspired by perturbation tests for explicit reasoning\(atanasova2023\_counterfactual;kaushik2020\_counterfactual;lanham2023\_measuring\_faithfulness\)and causal\-mediation analyses\(vig2020\_causal\_mediation;heimersheim2024\_patching\), we track the development of these behaviors along the training trajectory\.

#### Training\-time dynamics\.

Neural\-network capabilities can change abruptly during training, as documented by work on grokking, induction\-head phase transitions, and emergent abilities at scale\(power2022\_grokking;nanda2023\_progress;olsson2022\_induction\_heads;wei2022\_emergent\)\. Closer in spirit to our setting,tirumala2022\_memorization\_dynamicstrack per\-example behaviour across the training trajectory rather than at convergence, andtigges2024\_circuits\_across\_trainingextend mechanistic circuit analyses across checkpoints and scales; both motivate looking at trained\-model behaviour as the endpoint of a trajectory rather than as a fixed object\. To our knowledge, no prior work jointly tracks a latent reasoning method’s counterfactual\-following behaviour and the latent tokens’ causal contribution along its training trajectory\.

## 3Counterfactual Response Across Training

Following the framing of §[1](https://arxiv.org/html/2607.06648#S1), we measure faithfulness by whether interventions on the model’s input or on its latent reasoning steps change the prediction\. This section takes up the input intervention and the resulting output behaviour along training; the intervention on the latent reasoning steps is studied in §[4](https://arxiv.org/html/2607.06648#S4)\. Prior work has reported similar unfaithful behaviors at the converged checkpoints of latent reasoning methods trained under different paradigms\(zhang2025\_do\_latent\_tokens\_think;lin2025\_implicit\_shortcut;cui2026\_supervision;li2026\_dynamics\)\. We ask whether this surface similarity hides qualitatively different faithfulness behaviors underneath:do latent reasoning methods with different training paradigms exhibit qualitatively different faithfulness behaviors along training, even when their final\-checkpoint unfaithfulness looks similar?

### 3\.1Experimental Setup

#### Counterfactual edit on ProsQA\.

We formalise the edit introduced in §[1](https://arxiv.org/html/2607.06648#S1)as follows\. For each test instance with inputxix\_\{i\}and oracle answerAiA\_\{i\}, we construct a minimally edited inputxi′x^\{\\prime\}\_\{i\}whose oracle answer is the designated alternativeBiB\_\{i\}, and treat the model as counterfactually faithful on this pair when its prediction follows the edit fromAiA\_\{i\}toBiB\_\{i\}\.

ProsQA is chosen for three reasons\.*\(a\)*Its BFS\-verifiable oracle lets the construction scale across the test set without human\-edited labels, unlike NLI\- or sentiment\-style counterfactual sets\(kaushik2020\_counterfactual;gardner\-etal\-2020\-evaluating\)\.*\(b\)*ProsQA is the benchmark on which the original COCONUT report claims latent reasoning performs Breadth\-First Search \(BFS\), so the test directly probes that capability\.*\(c\)*ProsQA is also where latent reasoning’s lead over explicit chain\-of\-thought is largest, so a low counterfactual\-following rate here is a disconfirming observation rather than a weak\-baseline case\.

#### Construction algorithm\.

We instantiate this minimal edit as a rewrite of one directed edge in the underlying graph: changing a single edge of𝒢\\mathcal\{G\}alters exactly the reasoning fact that a graph traversal would have to use, while leaving every other sentence in the question intact\. Algorithm[1](https://arxiv.org/html/2607.06648#alg1)details the procedure\. For each example with directed graph𝒢\\mathcal\{G\}, rootrir\_\{i\}, original targetAiA\_\{i\}, and alternative targetBiB\_\{i\}, the procedure first attempts the canonical last\-hop swap that replaces the final edge of the gold reasoning path with an outgoing edge toBiB\_\{i\}\. If this swap fails to disconnectAiA\_\{i\}because the graph contains a parallel root\-to\-target path, the procedure falls back to a search over the bridge edges of the\(ri,Ai\)\(r\_\{i\},A\_\{i\}\)query and accepts the first swap that succeeds\. Every candidate swap is accepted only when BFS on the edited graph confirms¬\[ri→Ai\]∧\[ri→Bi\]\\neg\[r\_\{i\}\\to A\_\{i\}\]\\land\[r\_\{i\}\\to B\_\{i\}\], soBiB\_\{i\}becomes the unique reachable answer under the edit\.

Algorithm 1Verified answer\-flipping perturbation for ProsQA\.1:example

\(𝒢,r​o​o​t,t​a​r​g​e​t,t−,π\)\(\\mathcal\{G\},root,target,t^\{\-\},\\pi\)with gold path

π\\pi
2:swap

\(u,v\)→\(u,t−\)\(u,v\)\\to\(u,t^\{\-\}\)that flips the oracle answer, or

⊥\\bot
3:Tier 1 \(canonical last hop\):

4:

\(u,v\)←\(π\|π\|−2,π\|π\|−1\)\(u,v\)\\leftarrow\(\\pi\_\{\|\\pi\|\-2\},\\pi\_\{\|\\pi\|\-1\}\)
5:

𝒢′←\(𝒢∖\{\(u,v\)\}\)∪\{\(u,t−\)\}\\mathcal\{G\}^\{\\prime\}\\leftarrow\(\\mathcal\{G\}\\setminus\\\{\(u,v\)\\\}\)\\cup\\\{\(u,t^\{\-\}\)\\\}
6:if

¬BFS​\(𝒢′,r​o​o​t,t​a​r​g​e​t\)\\neg\\mathrm\{BFS\}\(\\mathcal\{G\}^\{\\prime\},root,target\)and

BFS​\(𝒢′,r​o​o​t,t−\)\\mathrm\{BFS\}\(\\mathcal\{G\}^\{\\prime\},root,t^\{\-\}\)then

7:return

\(u,v,canonical\)\(u,v,\\texttt\{canonical\}\)
8:endif

9:Tier 2 \(bridge search\):

10:

ℬ←Bridges​\(𝒢,r​o​o​t,t​a​r​g​e​t\)\\mathcal\{B\}\\leftarrow\\mathrm\{Bridges\}\(\\mathcal\{G\},root,target\)
11:for

\(u,v\)∈ℬ\(u,v\)\\in\\mathcal\{B\}do

12:

𝒢′←\(𝒢∖\{\(u,v\)\}\)∪\{\(u,t−\)\}\\mathcal\{G\}^\{\\prime\}\\leftarrow\(\\mathcal\{G\}\\setminus\\\{\(u,v\)\\\}\)\\cup\\\{\(u,t^\{\-\}\)\\\}
13:if

¬BFS​\(𝒢′,r​o​o​t,t​a​r​g​e​t\)\\neg\\mathrm\{BFS\}\(\\mathcal\{G\}^\{\\prime\},root,target\)and

BFS​\(𝒢′,r​o​o​t,t−\)\\mathrm\{BFS\}\(\\mathcal\{G\}^\{\\prime\},root,t^\{\-\}\)then

14:return

\(u,v,bridge\)\(u,v,\\texttt\{bridge\}\)
15:endif

16:endfor

17:return

⊥\\bot

#### Coverage and reuse\.

The9595rejected examples are those whose graph has multiple redundant root\-to\-target paths, for which no single\-edge swap can disconnect the target while keeping the graph well\-formed\. The accepted405405pairs are reused across every saved checkpoint of every paradigm; full coverage breakdown and reasoning\-hop distribution are in Appendix[A\.2](https://arxiv.org/html/2607.06648#A1.SS2)\.

#### Training paradigms and baselines\.

The four training paradigms share a GPT\-2 backbone\(radford2019\_gpt2\)\. CoT and answer\-only \(NoCoT\) are text baselines that bracket faithfulness from above and below\. COCONUT\(hao2025\_COCONUT\)and CODI\(shen2025\_codi\)are two latent reasoning methods that share the same architectural mechanism, feeding the previous\-step last hidden state back as the next\-step input embedding, but differ sharply in training: COCONUT uses a staged curriculum that gradually replaces explicit reasoning tokens with continuous\-thought positions, while CODI uses self\-distillation to transfer an explicit\-chain\-of\-thought signal into the same latent positions\. With architecture, backbone, and dataset fixed, any trajectory difference is attributable to the training mechanism, not data or architecture\. The unit of analysis is the saved checkpoint trajectory\{Mt\}t=1T\\\{M\_\{t\}\\\}\_\{t=1\}^\{T\}, not only the final checkpoint or the one chosen by validation\.

#### Metrics\.

For each checkpoint, predictions on\(xi,xi′\)\(x\_\{i\},x^\{\\prime\}\_\{i\}\)are mapped to joint labels over\{A,B,O\}2\\\{A,B,O\\\}^\{2\}, whereAAis the original oracleAiA\_\{i\},BBthe perturbed oracleBiB\_\{i\}, andOOany other answer\. The two key cells areAB\\mathrm\{AB\}\(correct on the original and following the perturbed oracle\) andAA\\mathrm\{AA\}\(correct on the original but retaining the original answer after the edit\)\. We report two complementary metrics\. Contrast Consistency, in the sense ofgardner\-etal\-2020\-evaluating’s contrast sets,

CC=Pr⁡\(y^​\(xi\)=Ai∧y^​\(xi′\)=Bi\),\\mathrm\{CC\}=\\Pr\\bigl\(\\hat\{y\}\(x\_\{i\}\)=A\_\{i\}\\,\\land\\,\\hat\{y\}\(x^\{\\prime\}\_\{i\}\)=B\_\{i\}\\bigr\),is the joint probability of being correct on the original input and following the perturbed oracle\. Conditional Flip Rate is its capability\-conditioned form,

CFR\\displaystyle\\mathrm\{CFR\}=Pr⁡\(y^​\(xi′\)=Bi∣y^​\(xi\)=Ai\)\\displaystyle=\\Pr\\bigl\(\\hat\{y\}\(x^\{\\prime\}\_\{i\}\)=B\_\{i\}\\mid\\hat\{y\}\(x\_\{i\}\)=A\_\{i\}\\bigr\)=CC/acc,\\displaystyle=\\mathrm\{CC\}/\\mathrm\{acc\},whereacc=Pr⁡\(y^​\(xi\)=Ai\)\\mathrm\{acc\}=\\Pr\(\\hat\{y\}\(x\_\{i\}\)=A\_\{i\}\)\.CFR\\mathrm\{CFR\}isolates faithfulness by conditioning on capability, unlikeCC\\mathrm\{CC\}\.

### 3\.2Experimental Result

Final checkpoints show a low\-faithfulness endpoint for both latent reasoning methods\.Figure[2](https://arxiv.org/html/2607.06648#S3.F2)reports clean accuracy,CC\\mathrm\{CC\}, andCFR\\mathrm\{CFR\}at each model’s best checkpoint\. Both latent reasoning methods end withCC\\mathrm\{CC\}andCFR\\mathrm\{CFR\}at roughly the NoCoT answer\-only baseline, even though COCONUT’s clean accuracy is the highest of the four paradigms; both methods’CFR\\mathrm\{CFR\}is roughly an order of magnitude below the explicit\-CoT baseline\.CC\\mathrm\{CC\}andCFR\\mathrm\{CFR\}agree at this snapshot, so the low\-faithfulness endpoint is robust to whether faithfulness is measured jointly with capability or conditioned on it\. This reproduces the final\-checkpoint concern from prior work using a verified counterfactual perturbation, but leaves open how the two paradigms arrived at the same endpoint\.

![Refer to caption](https://arxiv.org/html/2607.06648v1/x2.png)Figure 2:Clean accuracy, Contrast Consistency \(CC\\mathrm\{CC\}\), and Conditional Flip Rate \(CFR\\mathrm\{CFR\}\) on the405405ProsQA oracle pairs at each model’s best checkpoint \(CoT ep3030, NoCoT ep1414, CODI ep3737, COCONUT ep4848\)\.COCONUT shifts abruptly at the curriculum boundary\.Tracking the output\-level trajectory of COCONUT across saved checkpoints, that is, the sequence of joint\-label distributions on the counterfactual pairs at each saved epoch \(Figure[3](https://arxiv.org/html/2607.06648#S3.F3)\), changes the interpretation\. Through the early curriculum stages, COCONUT is dominated by the AB cell, that is, pairs on which the model is correct on the original input and follows the perturbed oracle\. At the epoch1515–1616curriculum boundary, where the latent budget grows from two to three positions, the joint label distribution flips within a single epoch from counterfactual\-following\-dominated to retention\-dominated, with detailed counts provided in Table[6](https://arxiv.org/html/2607.06648#A1.T6)in Appendix[A\.4](https://arxiv.org/html/2607.06648#A1.SS4); the AA cell roughly quadruples while AB nearly halves, andCFR\\mathrm\{CFR\}is approximately halved in the same epoch\. Clean accuracy continues to rise smoothly across this boundary, so the faithfulness drop is invisible from original\-question accuracy alone\.CC\\mathrm\{CC\}declines in parallel withCFR\\mathrm\{CFR\}but more gradually, because the simultaneous rise in clean accuracy partly compresses its change\.

![Refer to caption](https://arxiv.org/html/2607.06648v1/x3.png)Figure 3:Per\-epoch joint label distribution on the405405ProsQA oracle pairs for the four training paradigms\. Each column is a stacked area over\{AB,AA,BA,BB,AO,BO,O​\-cells\}\\\{\\mathrm\{AB\},\\mathrm\{AA\},\\mathrm\{BA\},\\mathrm\{BB\},\\mathrm\{AO\},\\mathrm\{BO\},\\mathrm\{O\}\\textrm\{\-cells\}\\\}, whereO​\-cells\\mathrm\{O\}\\textrm\{\-cells\}groupsOO\\mathrm\{OO\},OA\\mathrm\{OA\}, andOB\\mathrm\{OB\}\. On the COCONUT panel,AB→AA\\mathrm\{AB\}\\to\\mathrm\{AA\}collapses at the epoch1515–1616stage boundary\.CODI is retention\-like from early checkpoints\.CODI does not show this shift during training\. BothCFR\\mathrm\{CFR\}andCC\\mathrm\{CC\}sit at the NoCoT baseline level from the earliest saved checkpoint and never rise above it, so the trajectory enters the retention regime within the first few epochs and stays there\. The contrast with COCONUT is clean: COCONUT first becomes more counterfactually adaptive and then loses that behaviour at a curriculum boundary, while CODI never visits the counterfactual\-following regime in the first place\. A single snapshot at the best checkpoint collapses these two qualitatively different training stories into the same low\-CC\\mathrm\{CC\}, low\-CFR\\mathrm\{CFR\}summary\.

Takeaway 1Latent reasoning methods can look similarly unfaithful at convergence while following qualitatively different output\-level trajectories\.

## 4Latent Mediation Across Training

The output\-level trajectories from §[3](https://arxiv.org/html/2607.06648#S3)characterise what the model outputs but not how the latent reasoning steps contribute to that output\. Prior work probes this at the final checkpoint by applying interventions to the latent reasoning steps and reports that they are often replaceable or weakly causal for the prediction\(zhang2025\_do\_latent\_tokens\_think;cui2026\_supervision\), building on causal\-mediation and activation\-patching techniques originally developed for analysing model components\(vig2020\_causal\_mediation;heimersheim2024\_patching\)\. Such a final\-checkpoint view leaves open whether the output\-level trajectory in §[3](https://arxiv.org/html/2607.06648#S3)has an activation\-level counterpart, so we ask:does the causal contribution of the latent reasoning steps to the answer follow its own trajectory across training, and does its timing align with the output\-level transition within each paradigm?

### 4\.1Experimental Setup

#### Trajectory\-level patching\.

Following causal\-mediation and activation\-patching analyses\(vig2020\_causal\_mediation;heimersheim2024\_patching\), we patch all latent reasoning steps in a single intervention using one of three content\-erasing values:zero, the dataset mean latent embedding \(mean\), or norm\-matched Gaussian noise \(norm\-noise\)\. We replace the latent states at every step, forcing all downstream steps and the answer to be produced under the perturbed state \(implementation details in Appendix[A\.3](https://arxiv.org/html/2607.06648#A1.SS3)\)\. The unit of analysis is the trajectory: we track how the intervention’s effect size evolves across saved checkpoints along training\.

The first metric is the output change rate, which measures the probability that predictions change under intervention:

OCR=Pr⁡\[y^patch​\(x\)≠y^clean​\(x\)\]\.\\mathrm\{OCR\}=\\Pr\\bigl\[\\hat\{y\}^\{\\text\{patch\}\}\(x\)\\neq\\hat\{y\}^\{\\text\{clean\}\}\(x\)\\bigr\]\.\(1\)The second metric is the preserved\-when\-correct rate, which isolates the subset of inputs that are correct in the clean pass to measure how often the intervention is destructive:

PWC=Pr⁡\[y^patch​\(x\)=y\|y^clean​\(x\)=y\]\.\\mathrm\{PWC\}=\\Pr\\bigl\[\\hat\{y\}^\{\\text\{patch\}\}\(x\)=y\\,\\bigm\|\\,\\hat\{y\}^\{\\text\{clean\}\}\(x\)=y\\bigr\]\.\(2\)The two metrics are complementary: a non\-zeroOCR\\mathrm\{OCR\}can arise from breaking correct answers or from changing already\-incorrect ones, andPWC\\mathrm\{PWC\}separates these\. HighOCR\\mathrm\{OCR\}with lowPWC\\mathrm\{PWC\}means ablation destroys previously correct predictions, whereas lowOCR\\mathrm\{OCR\}with highPWC\\mathrm\{PWC\}indicates a small causal mediation effect\. We report the mean across four independent sampling seeds±1\\pm 1standard deviation, with further details in Appendix[A\.3](https://arxiv.org/html/2607.06648#A1.SS3)\.

#### Cohort\-aware patching\.

The aggregateAB→AA\\mathrm\{AB\}\\to\\mathrm\{AA\}transition in §[3](https://arxiv.org/html/2607.06648#S3)is a property of the COCONUT average; we test whether the examples that flip on the output side also lose their causal contribution on the activation side\. We therefore partition the test set into three cohorts by their joint label at checkpoints around the curriculum boundary, as defined in Table[1](https://arxiv.org/html/2607.06648#S4.T1)\. For each cohort we report the contrast indirect effect\(vig2020\_causal\_mediation;wang2023\_ioi;heimersheim2024\_patching\),IEicontrast=mipatch−miclean\\mathrm\{IE\}^\{\\mathrm\{contrast\}\}\_\{i\}=m\_\{i\}^\{\\text\{patch\}\}\-m\_\{i\}^\{\\text\{clean\}\}withmi=log⁡p​\(Bi\)−log⁡p​\(Ai\)m\_\{i\}=\\log p\(B\_\{i\}\)\-\\log p\(A\_\{i\}\), on the perturbed input; positive values indicate that the intervention shifts probability mass toward the perturbed oracle answer\. Full formulas are in Appendix[A\.3](https://arxiv.org/html/2607.06648#A1.SS3)\.

Cohortep1515ep1616ep4848GstableG\_\{\\text\{stable\}\}AB\\mathrm\{AB\}−\-AB\\mathrm\{AB\}GdriftG\_\{\\text\{drift\}\}AB\\mathrm\{AB\}AA\\mathrm\{AA\}AA\\mathrm\{AA\}GretainG\_\{\\text\{retain\}\}AA\\mathrm\{AA\}−\-AA\\mathrm\{AA\}Table 1:Cohort criteria for the COCONUT ProsQA analysis \(ep==epoch\): ep1515is just before the COCONUT curriculum boundary, ep1616is immediately after it, and ep4848is the best checkpoint\. A−\-entry means that epoch1616is not part of that cohort definition\.

### 4\.2Experimental Result

Noise\-ablation patching loses effect along the ProsQA trajectory\.Figure[4](https://arxiv.org/html/2607.06648#S4.F4)reportsOCR\\mathrm\{OCR\}andPWC\\mathrm\{PWC\}across the ProsQA checkpoint grid\. Both COCONUT and CODI follow the same overall trajectory:OCR\\mathrm\{OCR\}falls toward zero andPWC\\mathrm\{PWC\}rises toward one across training, ending at a regime where noise ablation at the latent reasoning steps no longer produces detectable extracted\-answer changes\. This is the activation\-level counterpart of the output\-level pattern at convergence in RQ1: both models not only retain the original answer under counterfactual edits, but also become insensitive to noise ablation at the latent reasoning steps\.

![Refer to caption](https://arxiv.org/html/2607.06648v1/x4.png)Figure 4:Causal analyses along the ProsQA training trajectory\. Top: COCONUT,OCR\\mathrm\{OCR\}andPWC\\mathrm\{PWC\}across epochs for thezero,mean, andnorm\-noisecontent\-erasing interventions\. Bottom: CODI, same metrics on its1111\-checkpoint grid\. Lines are the mean across four independent256256\-question sampling seeds; shaded bands are±1\\pm 1standard deviation across seeds\.The two paradigms approach the low\-dependence regime through qualitatively different trajectories\.Figure[4](https://arxiv.org/html/2607.06648#S4.F4)shows that for COCONUT,OCR\\mathrm\{OCR\}falls andPWC\\mathrm\{PWC\}rises smoothly across the curriculum and stabilises near the low\-dependence regime by the best checkpoint\. CODI reaches the same endpoint within the first few epochs and then stays there, aside from a brief epoch\-1010transient\. CODI thus enters the low\-dependence regime early and abruptly, whereas COCONUT approaches it gradually; the boundary effect documented in §[3](https://arxiv.org/html/2607.06648#S3)for COCONUT is recovered at the activation level only in the cohort analysis below\.

For examples that drift from counterfactual following to retention of the original answer, the causal contribution of the latent steps also decays\.Table[2](https://arxiv.org/html/2607.06648#S4.T2)reports the effects per cohort: onGdriftG\_\{\\text\{drift\}\}under all\-stepnorm\-noise, the medianIEcontrast\\mathrm\{IE\}^\{\\mathrm\{contrast\}\}on the perturbed input drops across the same epoch1515–1616boundary that produces the aggregate AB→\\toAA transition in RQ1, and remains lower by the best checkpoint\. This weakening is specific to the drift cohort under the counterfactual edit: the same cohort on the original input retains a large shift across the boundary, and the stable and retention cohorts, which are defined by epochs1515and4848, do not show a comparable drop\.

CohortInputep 15ep 16ep 48GstableG\_\{\\text\{stable\}\}\(n=41n=41\)perturbed1\.931\.710\.13GdriftG\_\{\\text\{drift\}\}\(n=62n=62\)original1\.761\.780\.20GdriftG\_\{\\text\{drift\}\}\(n=62n=62\)perturbed1\.710\.700\.23GretainG\_\{\\text\{retain\}\}\(n=35n=35\)perturbed2\.323\.440\.20Table 2:Median contrast indirect effectIEcontrast\\mathrm\{IE\}^\{\\mathrm\{contrast\}\}in nats under all\-stepnorm\-noiseon the three cohorts of Table[1](https://arxiv.org/html/2607.06648#S4.T1); positive values indicate that probability mass shifts toward the perturbed answerBB\. The*Input*column distinguishes the original ProsQA questionxix\_\{i\}from its counterfactual editxi′x^\{\\prime\}\_\{i\}defined in §[3](https://arxiv.org/html/2607.06648#S3)\.Takeaway 2Across training on ProsQA, the latent reasoning steps’ causal contribution to the answer decays for both paradigms; the examples that flip on the output side in §[3](https://arxiv.org/html/2607.06648#S3)are also the examples on which this contribution decays\.

## 5Cross\-Format Analysis: Latent Mediation Across Answer Formats

The output\-level and activation\-level trajectories identified in §[3](https://arxiv.org/html/2607.06648#S3)and §[4](https://arxiv.org/html/2607.06648#S4)are both observed on ProsQA, whose answer format is binary choice\. Prior final\-checkpoint studies report that latent reasoning steps show different levels of causal involvement across datasets\(li2026\_dynamics;cui2026\_supervision\)\. How this variability forms along training has not been examined, and if the activation\-level trajectory itself depends on the answer format, ProsQA alone cannot speak for latent reasoning as a whole\. We therefore ask:do the output\-level and activation\-level trajectories observed in §[3](https://arxiv.org/html/2607.06648#S3)and §[4](https://arxiv.org/html/2607.06648#S4)depend on the answer format, or do they persist when only the format changes?

![Refer to caption](https://arxiv.org/html/2607.06648v1/x5.png)Figure 5:Latent\-position dependence along the trajectory under the three all\-step content\-erasing interventions \(zero,mean,norm\-noise\) for COCONUT on three tasks\. Left: ProsQA\. Middle: GSM\-Choice\. Right: GSM\-open\. Top row: output change rateOCR\\mathrm\{OCR\}\. Bottom row: preserved\-when\-correct ratePWC\\mathrm\{PWC\}\. Lines are the mean over four independent sampling seeds and shaded bands are±1\\pm 1standard deviation across seeds\.### 5\.1Experimental Setup

We apply the same all\-step noise\-ablation activation patching from §[4](https://arxiv.org/html/2607.06648#S4)to COCONUT on two answer formats of the same GSM8K problems\(cobbe2021\_gsm8k\)\.*GSM\-open*keeps the original open\-ended numerical answer format\.*GSM\-Choice*rewrites each item as a two\-option multiple\-choice question containing the gold numerical answer and a distractor, aligned with the binary\-choice format used by ProsQA\. Because the GSM pair shares problem content, backbone, model family, and COCONUT training recipe, the comparison isolates the answer format more cleanly than a cross\-task comparison can\. We focus on the within\-format trajectory direction ofOCR\\mathrm\{OCR\}undernorm\-noise, rather than comparing absoluteOCR\\mathrm\{OCR\}levels across formats with different output spaces\. Dataset construction and full intervention details are in Appendix[A\.1](https://arxiv.org/html/2607.06648#A1.SS1)and Appendix[A\.3](https://arxiv.org/html/2607.06648#A1.SS3)\.

### 5\.2Experimental Result

Changing only the answer format reverses the trajectory direction\.Figure[5](https://arxiv.org/html/2607.06648#S5.F5)shows the main result\. On GSM\-open,OCR\\mathrm\{OCR\}undernorm\-noiserises along training and remains high at the final evaluated epoch\. On GSM\-Choice, the same intervention moves in the opposite direction:OCR\\mathrm\{OCR\}falls andPWC\\mathrm\{PWC\}rises across training, matching the direction observed on ProsQA in §[4](https://arxiv.org/html/2607.06648#S4)\.

EpochPatchPWCc→\\toww→\\toc4zero0\.8010\.1990\.51013zero0\.8500\.1500\.28025zero0\.9190\.0810\.18025mean0\.9330\.0670\.15525norm\-noise0\.9220\.0780\.162Table 3:Flip decomposition on GSM\-Choice×\\timesCOCONUT \(mean across four sampling seeds\): preserved\-when\-correct rate vs\. correct\-to\-wrong and wrong\-to\-correct flip rates\.The flip decomposition shows that non\-zeroOCR\\mathrm\{OCR\}can reflect removable bias\.On GSM\-Choice, output changes are not mainly destructive changes from correct to wrong\. Table[3](https://arxiv.org/html/2607.06648#S5.T3)shows that at the best checkpoint underzero, the wrong→\\tocorrect flip rate exceeds the correct→\\towrong rate: the intervention preserves most predictions that are correct before patching, while a larger share of predictions that were wrong before the intervention move to the gold answer after patching\. This breakdown keeps the interpretation ofOCR\\mathrm\{OCR\}precise: in binary\-choice settings, latent reasoning steps may carry a removable answer bias rather than information required for the correct answer\.

GSM\-ChoiceGSM\-openCOCONUT \(best ep\.\)0\.8680\.356CoT \(best ep\.\)0\.8280\.446Δ\\Delta\(COCONUT−\-CoT\)\+0\.040\+0\.040−0\.090\-0\.090Table 4:Validation accuracy at each model’s best epoch on the same GSM8K problems in two answer formats\.The same format change also affects headline accuracy\.Table[4](https://arxiv.org/html/2607.06648#S5.T4)shows a related implication for benchmark design: the ranking of COCONUT versus CoT in clean accuracy flips between open\-ended GSM8K and its binary\-choice reformulation\. This ranking flip reinforces the broader point that answer format can change both apparent performance and the causal contribution measured at the latent reasoning steps\.

Takeaway 3Changing only the answer format reverses the activation\-level trajectory—decaying on binary choice, rising on open\-ended decoding—so latent reasoning faithfulness depends on both the training trajectory and the answer format\.

## 6Conclusion

We track latent reasoning faithfulness across saved checkpoints rather than only at convergence, combining a verified counterfactual edit on the input with activation patching on the latent reasoning steps\. Across COCONUT and CODI, the two paradigms reach similarly unfaithful endpoints through qualitatively divergent output\-level and activation\-level trajectories, and the activation\-level trajectory reverses direction when only the answer format changes from binary choice to open\-ended decoding\. Latent reasoning faithfulness is therefore a property of the training stage and the answer format, not of the architecture or the final checkpoint alone\.

## Limitations

Our findings are scoped to settings where the analyses can give verifiable answers\. The verified counterfactual edit requires a deterministic oracle on the perturbed input, which ProsQA provides through a BFS solver on the underlying graph; extending it to settings without such an oracle requires a new construction\. We run all four training paradigms on the GPT\-2 small backbone, both to match the scale at which the reference COCONUT and CODI recipes are reported\(hao2025\_COCONUT;shen2025\_codi\)and to keep the cost of reproducing full training trajectories tractable, following prior work that studies model behaviour across training at comparable scales\(tirumala2022\_memorization\_dynamics;tigges2024\_circuits\_across\_training\)\. The specific transition points we report should therefore be read as concrete instances of the methodological claim that final\-checkpoint evaluation can miss substantial changes along training, rather than as universal claims across architectures\.

## References

## Appendix AAppendix

### A\.1Training and Checkpoint Details

All four model families \(CoT, NoCoT, COCONUT, CODI\) are trained on a single RTX PRO 6000 GPU\. CoT, NoCoT, and COCONUT use the COCONUT reference codebase\(hao2025\_COCONUT\)unmodified at the optimisation level; we adapt only the I/O paths and the per\-device batch size so that, with gradient accumulation, the effective batch matches the original four\-GPU recipe \(32×4=12832\\times 4=128\)\. CODI uses its reference repository\(shen2025\_codi\)with only absolute paths changed\. The main training budgets are approximately3636GPU\-hours for ProsQA×\\timesCOCONUT,8383GPU\-hours each for GSM\-open×\\timesCOCONUT and GSM\-Choice×\\timesCOCONUT, and55GPU\-hours for ProsQA×\\timesCODI\. The GPT\-2 checkpoint is MIT\-licensed, and the other external artifacts used here are publicly released open\-source research artifacts\. We use these artifacts only for research training and evaluation, and the ProsQA perturbation pairs and GSM\-Choice reformulation introduced here are intended as research evaluation artifacts\. Every checkpoint that the original training loop saves is retained; checkpoint saving is not filtered by validation score\. When an analysis uses a best checkpoint, we select it afterward by validation score and state that choice explicitly\.

#### Backbone and tokeniser\.

All models useopenai\-community/gpt2\(GPT\-2 small, 124M parameters\)\(radford2019\_gpt2\)from a local HuggingFace snapshot\. The tokeniser is the GPT\-2 BPE tokeniser\. COCONUT and CODI extend the embedding table with the special tokens<bot\>,<eot\>, and<latent\>\.

#### Optimiser\.

For CoT, NoCoT, and COCONUT we use AdamW with learning rate10−410^\{\-4\}, weight decay0\.010\.01, andreset\_optimizer=Truewhen COCONUT enters its first latent stage\. Precision is FP32 \(bf16=False\)\. For CODI we follow the published recipe: AdamW with learning rate3⋅10−33\\cdot 10^\{\-3\}, weight decay0\.10\.1, cosine schedule,0\.030\.03warm\-up ratio, gradient clipping at2\.02\.0, BF16 mixed precision, and LoRA adapters \(r=128r=128,α=32\\alpha=32, dropout0\) on top of frozen GPT\-2 weights\. CODI’s projection head has dimension768768with no dropout, and the distillation loss is normalised by the running standard deviation \(distill\_loss\_div\_std=True\)\.

#### ProsQA training\.

Train/valid/test splits contain17,88617\{,\}886/300300/500500examples respectively\. CoT, NoCoT, and COCONUT are each trained for5050epochs on the same loader\. The COCONUT curriculum usescthought=1c\_\{\\text\{thought\}\}=1,epochs\_per\_stage=5\\texttt\{epochs\\\_per\\\_stage\}=5, andmax\_latent\_stage=6\\texttt\{max\\\_latent\\\_stage\}=6, so the latent token budget grows as0→1→2→⋯→60\\to 1\\to 2\\to\\cdots\\to 6over the first3030epochs and then remains at66\. CoT keeps the same scaffolding but withcot=True\(no latent positions\)\. NoCoT usesno\_cot=True,cthought=0c\_\{\\text\{thought\}\}=0, andmax\_latent\_stage=0\\texttt\{max\\\_latent\\\_stage\}=0\. CODI is trained for4040epochs withnum\_latent=6\\texttt\{num\\\_latent\}=6,model\_max\_length=512\\texttt\{model\\\_max\\\_length\}=512,max\_token\_num=1000\\texttt\{max\\\_token\\\_num\}=1000,include\_last\_cot=True, andremove\_eos=True, using per\-device batch size1616with88accumulation steps \(effective batch128128\)\.

#### GSM\-open training\.

Train/valid/test splits contain385,620385\{,\}620/500500/1,3191\{,\}319examples \(thecoconut/data/gsm\_train\.jsonaugmentation expansion of GSM8K\(cobbe2021\_gsm8k\)\)\. Following the original COCONUT recipe, training is two\-staged\.*Stage A \(CoT\)*: GPT\-2 is trained for2525epochs withcthought=0c\_\{\\text\{thought\}\}=0andmax\_latent\_stage=0\\texttt\{max\\\_latent\\\_stage\}=0, saving every epoch \(save\_only\_improve=False\); we usecheckpoint\_19of this run as the initialisation for Stage B \(this is the highest\-validation checkpoint of the CoT teacher\)\.*Stage B \(COCONUT\)*: training continues for another2525epochs withcthought=2c\_\{\\text\{thought\}\}=2,epochs\_per\_stage=3\\texttt\{epochs\\\_per\\\_stage\}=3,max\_latent\_stage=3\\texttt\{max\\\_latent\\\_stage\}=3, andresume=3\(the curriculum starts from stage11, skipping stage0which would just reproduce the loaded CoT checkpoint\)\. The realised schedule is therefore: epochs 1–3 – stage 0,nlatent=0n\_\{\\text\{latent\}\}=0; epochs 4–6 – stage 1,nlatent=2n\_\{\\text\{latent\}\}=2; epochs 7–9 – stage 2,nlatent=4n\_\{\\text\{latent\}\}=4; epochs 10–25 – stage 3,nlatent=6n\_\{\\text\{latent\}\}=6\. Both stages usebatch\_size\_training=32,gradient\_accumulation\_steps=4\\texttt\{gradient\\\_accumulation\\\_steps\}=4\(effective batch128128\)\.

#### GSM\-Choice training\.

The GSM\-Choice corpus uses the same385,620385\{,\}620/500500/1,3191\{,\}319split sizes, but each item is rewritten as a two\-option multiple choice question following the methodology ofzhang2024\_gsm\_mc\. The question text is appended with\\\\backslashnA\) v\\1\{\}\_\{1\}\\backslashnB\) v\\2\{\}\_\{2\}\\backslashnAnswer:, where one of\{v1,v2\}\\\{v\_\{1\},v\_\{2\}\\\}is the gold numerical answer and the other is a distractor\. Distractors are drawn, in order of preference, from \(i\) the public GSM\-MC option pool ofzhang2024\_gsm\_mcwhen available, \(ii\) intermediate values that appear inside the<<…=N\>\>step calculations, and \(iii\) magnitude\-perturbed variants of the gold answer \(±1,±10,×2,/2\\pm 1,\\pm 10,\\times 2,/2\)\. The gold letter is balanced across positions \(Pr⁡\[gold=A\]=0\.5\\Pr\[\\text\{gold\}=A\]=0\.5\)\. Training proceeds with the same two\-stage protocol and identical hyperparameters as GSM\-open; we usecot\-gsm\-mc/checkpoint\_12as the Stage A initialisation for Stage B\. All other hyperparameters \(lr, weight decay, accumulation,2525epochs per stage\) are unchanged\.

#### Checkpoint grids for trajectory analyses\.

Behavioural analyses consume every saved checkpoint of every run \(epochs11–5050on ProsQA,11–2525on GSM\-open and GSM\-Choice,11–4040on CODI\-ProsQA\)\. For the activation patching analyses we evaluate on a per\-setting grid of checkpoints chosen to span the curriculum, with extra density placed where the trajectory in §[4](https://arxiv.org/html/2607.06648#S4)changes most rapidly:

- •ProsQA×\\timesCOCONUT: epochs\{6,10,15,16,20,30,40,48\}\\\{6,10,15,16,20,30,40,48\\\}\(88checkpoints;epoch\_15andepoch\_16bracket thenlatent=2→3n\_\{\\text\{latent\}\}\{=\}2\\\!\\to\\\!3curriculum boundary highlighted in §[3](https://arxiv.org/html/2607.06648#S3)\);
- •ProsQA×\\timesCODI: epochs\{1,2,3,4,5,10,15,20,25,30,40\}\\\{1,2,3,4,5,10,15,20,25,30,40\\\}\(1111checkpoints; the first five epochs capture the early drop in theOCR\\mathrm\{OCR\}trajectory\);
- •GSM\-open×\\timesCOCONUT: epochs\{4,7,10,13,16,19,22,25\}\\\{4,7,10,13,16,19,22,25\\\}\(88checkpoints; one checkpoint per stage in stages 0–3, then four checkpoints inside the terminal stage\);
- •GSM\-Choice×\\timesCOCONUT: same epoch grid as GSM\-open \(88checkpoints\)\.

The activation patching sampling protocol, decoding settings, and uncertainty methodology are in §[A\.3](https://arxiv.org/html/2607.06648#A1.SS3)\.

### A\.2Counterfactual Edit Construction on ProsQA

This appendix collects the technical details supporting the verified counterfactual edit of §[3](https://arxiv.org/html/2607.06648#S3): solver primitives, the swap acceptance criteria, the strategy\-level coverage breakdown, the reasoning\-hop distribution of the accepted pairs, and the surface\-form rewriting rule\.

#### Solver primitives\.

Reachability is decided by breadth\-first search on the directed adjacency list\. An edge\(u,v\)\(u,v\)is a*bridge for the \(root, target\) query*if removing it disconnectstargetfromroot\(this is a path\-specific notion, not the classical undirected bridge\)\. Bridges are enumerated by trying each edge\(u,v\)\(u,v\)that lies on at least oneroot→\\totargetpath and re\-running BFS without it\.

#### Swap acceptance criteria\.

Beyond the BFS\-verified answer flip used in Algorithm[1](https://arxiv.org/html/2607.06648#alg1), a candidate swap must also \(i\) leave the graph well\-formed \(no orphan nodes, no schema violations\) and \(ii\) admit a single\-sentence surface rewrite of the form “Everyuuis avv”→\\to“Everyuuis at−t^\{\-\}”, wheret−t^\{\-\}is the alternative candidateBiB\_\{i\}\. Edits that fail either criterion are rejected before the BFS check\.

#### Strategy\-level coverage\.

The405405accepted pairs split into335335found by the canonical last\-hop swap \(82\.7%82\.7\\%of accepted\) and7070found by the bridge search \(17\.3%17\.3\\%\)\. The rejected examples are those for which every edge on everyr​o​o​t→t​a​r​g​e​troot\\to targetpath has a bypass, so no single\-edge swap can achieve the conjunction¬\[r​o​o​t→t​a​r​g​e​t\]∧\[r​o​o​t→t−\]\\neg\[root\\to target\]\\land\[root\\to t^\{\-\}\]\. Because the original ProsQA graph satisfies¬\[r​o​o​t→t−\]\\neg\[root\\to t^\{\-\}\]by construction \(onlyt​a​r​g​e​ttargetis reachable fromr​o​o​troot\), the binding constraint is the existence of a bridge edge\(u,v\)\(u,v\)whose removal disconnectst​a​r​g​e​ttarget; these rejected examples would require coordinated multi\-edge edits, which we exclude in order to keep the surface change to a single rewritten sentence\.

#### Reasoning\-hop distribution\.

The405405accepted pairs preserve the reasoning\-difficulty mixture of the full ProsQA test set rather than concentrating on shorter paths\. Table[5](https://arxiv.org/html/2607.06648#A1.T5)reports the distribution of the number of gold reasoning\-step sentences \(which determines the canonical\-path length\) on the500500\-example test set and on the405405accepted pairs\. The acceptance rate is in fact*higher*for longer\-path examples \(9797–100%100\\%for55–66steps vs\.67%67\\%for33steps\), because graphs with longer canonical paths are less likely to have a parallel alternative path; the resulting set is therefore mildly enriched in44–66\-step problems but still spans the full hop range present in the test set\.

TestPert\.Accept\.Stepsnn%nn%rate320240\.413533\.366\.8%421743\.419147\.288\.0%56913\.86716\.597\.1%6122\.4123\.0100%total50040581\.0%Table 5:Distribution of the number of gold reasoning\-step sentences \(a proxy for canonical\-path length\) on the ProsQA test set \(n=500n=500\) and on the405405accepted perturbation pairs\. The two distributions are within a few percentage points of each other; the perturbation set spans the full hop range and is mildly enriched in longer\-path examples\.
#### Question rewriting\.

After accepting a swap\(u,v\)→\(u,t−\)\(u,v\)\\to\(u,t^\{\-\}\), the perturbed natural\-language question is generated by replacing the unique sentence “Everyσ​\(u\)\\sigma\(u\)is aσ​\(v\)\\sigma\(v\)\.” with “Everyσ​\(u\)\\sigma\(u\)is aσ​\(t−\)\\sigma\(t^\{\-\}\)\.” \(whereσ\\sigmais the symbol\-name map of the example\), leaving every other sentence unchanged\. The query sentence \(“Isσ​\(r​o​o​t\)\\sigma\(root\)aσ​\(t​a​r​g​e​t\)\\sigma\(target\)or aσ​\(t−\)\\sigma\(t^\{\-\}\)?”\) is preserved verbatim, so the answer flips fromσ​\(t​a​r​g​e​t\)\\sigma\(target\)toσ​\(t−\)\\sigma\(t^\{\-\}\)but the candidate set, entity inventory, and surface form are otherwise identical\.

### A\.3Metric and Intervention Definitions

This appendix collects the explicit formulas, cohort usage, and intervention protocols used in the analyses of §[3](https://arxiv.org/html/2607.06648#S3)–§[5](https://arxiv.org/html/2607.06648#S5)\.

#### Joint\-label coding\.

For each original/perturbed pairii, we map the extracted answer on both inputs into\{A,B,O\}\\\{A,B,O\\\}, whereAAis the original oracle answerAiA\_\{i\},BBis the perturbed oracle answerBiB\_\{i\}, andOOis any other extracted answer\. The resulting two\-letter codeℓtM​\(i\)\\ell\_\{t\}^\{M\}\(i\)takes valuesAB, the counterfactually adaptive cell;AA, original\-answer retention;AO, originally correct but predicts neitherAiA\_\{i\}norBiB\_\{i\}on the perturbed input;BA, the reversed pattern; andBB, predicting the perturbed oracle on both inputs\. Cells withOOin either coordinate are retained in the full3×33\\times 3table and grouped as needed in figures\.

#### Contrast Consistency, Conditional Flip Rate, and complementary rates\.

For each model checkpointMtM\_\{t\}and oracle pairii, letAiA\_\{i\}be the original oracle answer andBiB\_\{i\}be the perturbed oracle answer\. We report original\-question accuracy

acctM=Pr⁡\[y^Mt​\(xi\)=Ai\],\\mathrm\{acc\}\_\{t\}^\{M\}=\\Pr\[\\hat\{y\}\_\{M\_\{t\}\}\(x\_\{i\}\)=A\_\{i\}\],Contrast Consistency in the sense ofgardner\-etal\-2020\-evaluating’s contrast sets,

CC​\(Mt\)=Pr⁡\[y^Mt​\(xi\)=Ai∧y^Mt​\(xi′\)=Bi\],\\mathrm\{CC\}\(M\_\{t\}\)=\\Pr\\\!\\left\[\\hat\{y\}\_\{M\_\{t\}\}\(x\_\{i\}\)=A\_\{i\}\\,\\land\\,\\hat\{y\}\_\{M\_\{t\}\}\(x^\{\\prime\}\_\{i\}\)=B\_\{i\}\\right\],and the Conditional Flip Rate \(the conditional form ofCC\\mathrm\{CC\}, in the lineage ofkaushik2020\_counterfactual;lanham2023\_measuring\_faithfulness\),

CFR​\(Mt\)\\displaystyle\\mathrm\{CFR\}\(M\_\{t\}\)=Pr⁡\[y^Mt​\(xi′\)=Bi\|y^Mt​\(xi\)=Ai\]\\displaystyle=\\Pr\\\!\\left\[\\hat\{y\}\_\{M\_\{t\}\}\(x^\{\\prime\}\_\{i\}\)=B\_\{i\}\\,\\middle\|\\,\\hat\{y\}\_\{M\_\{t\}\}\(x\_\{i\}\)=A\_\{i\}\\right\]=CC​\(Mt\)/acctM\.\\displaystyle=\\mathrm\{CC\}\(M\_\{t\}\)/\\mathrm\{acc\}\_\{t\}^\{M\}\.Equivalently,CC\\mathrm\{CC\}is the AB rate on the full405405\-pair denominator andCFR\\mathrm\{CFR\}is the AB rate among the originally\-correct pairs \(AB\+\+AA\+\+AO\); the complementary AA \(retention\) and AO fractions on that same denominator account for the rest\. We reportCC\\mathrm\{CC\}andCFR\\mathrm\{CFR\}jointly becauseCC\\mathrm\{CC\}mixes capability with faithfulness whileCFR\\mathrm\{CFR\}isolates faithfulness by conditioning on capability\.

#### Transition cohorts on ProsQA×\\timesCOCONUT\.

We use the three COCONUT ProsQA cohorts defined in Table[1](https://arxiv.org/html/2607.06648#S4.T1)\. Epoch1616is reported for every cohort, but forGstableG\_\{\\text\{stable\}\}andGretainG\_\{\\text\{retain\}\}it is not used to select examples; it is only an evaluation checkpoint for the intervention\. OnlyGdriftG\_\{\\text\{drift\}\}is conditioned on epoch1616, because it isolates examples that move fromAB\\mathrm\{AB\}at epoch1515toAA\\mathrm\{AA\}at the curriculum boundary and are alsoAA\\mathrm\{AA\}at epoch4848\. On each cohort we apply the all\-stepnorm\-noiseintervention at epochs1515,1616, and4848, and measure the answer margin shift

mi\\displaystyle m\_\{i\}=log⁡p​\(Bi\)−log⁡p​\(Ai\),\\displaystyle=\\log p\(B\_\{i\}\)\-\\log p\(A\_\{i\}\),Δ​mi\\displaystyle\\Delta m\_\{i\}=mipatch−miclean,\\displaystyle=m\_\{i\}^\{\\text\{patch\}\}\-m\_\{i\}^\{\\text\{clean\}\},whereΔ​mi\>0\\Delta m\_\{i\}\>0indicates that the intervention shifts probability mass toward the perturbed oracle answer\. The cohort design controls for item difficulty: inGdriftG\_\{\\text\{drift\}\}, the same example is solved correctly under perturbation at epoch1515but not at epochs1616and4848, so comparisons across these checkpoints are not due to changing item composition\.

#### Activation patching protocol\.

For each \(checkpoint, intervention\) cell, evaluation is repeated on four independently sampled256256\-question subsets of the test set \(sampling seeds\{0,1,2,3\}\\\{0,1,2,3\\\}\), and we compare a clean forward pass with a patched forward pass on each\. The all\-step intervention replaces every latent position in the current curriculum stage\. We usezeroreplacement, corresponding\-positionmeanreplacement from a held\-out calibration batch, andnorm\-noisereplacement with isotropic Gaussian noise rescaled to the original latent norm and repeated for33noise seeds per question\. For COCONUT, replacement is applied to the continuous\-thought embedding before it is fed back as the next input embedding; for CODI, replacement is applied to the post\-projection hidden state used as the next latent\-loop input\. Decoding uses greedy generation with the model\-specific maximum new\-token budget \(3232tokens for ProsQA,256256for GSM\-open,88for GSM\-Choice\), and the answer is extracted by the same regex used by the COCONUT and CODI evaluation scripts\.

#### Output\-level metrics for activation patching\.

The output change rate is

OCR=Pr⁡\[y^patch≠y^clean\],\\mathrm\{OCR\}=\\Pr\[\\hat\{y\}^\{\\text\{patch\}\}\\neq\\hat\{y\}^\{\\text\{clean\}\}\],the preserved\-when\-correct rate is

PWC=Pr⁡\[y^patch=y∣y^clean=y\],\\mathrm\{PWC\}=\\Pr\[\\hat\{y\}^\{\\text\{patch\}\}=y\\mid\\hat\{y\}^\{\\text\{clean\}\}=y\],whereyyis the task gold answer after the task\-specific answer normalisation\. The flip decomposition refines output changes by correctness direction:

Pr⁡\[c→w\]\\displaystyle\\Pr\[\\mathrm\{c\}\\to\\mathrm\{w\}\]=Pr⁡\[y^patch≠y∣y^clean=y\],\\displaystyle=\\Pr\[\\hat\{y\}^\{\\text\{patch\}\}\\neq y\\mid\\hat\{y\}^\{\\text\{clean\}\}=y\],Pr⁡\[w→c\]\\displaystyle\\Pr\[\\mathrm\{w\}\\to\\mathrm\{c\}\]=Pr⁡\[y^patch=y∣y^clean≠y\],\\displaystyle=\\Pr\[\\hat\{y\}^\{\\text\{patch\}\}=y\\mid\\hat\{y\}^\{\\text\{clean\}\}\\neq y\],which is necessary whenOCR\\mathrm\{OCR\}is nonzero butaccc−accp≈0\\mathrm\{acc\}\_\{c\}\-\\mathrm\{acc\}\_\{p\}\\approx 0, the regime of GSM\-Choice×\\timesCOCONUT at late checkpoints in Table[3](https://arxiv.org/html/2607.06648#S5.T3)\. For GSM\-open,OCR\\mathrm\{OCR\}is computed after the same numeric\-answer normalisation used for accuracy\.

#### Sampling and uncertainty\.

Each \(checkpoint, intervention\) cell is estimated from4×256=10244\\times 256=1024clean forward passes and, fornorm\-noise,4×256×3=30724\\times 256\\times 3=3072patched forward passes\. Reported metrics in §[4](https://arxiv.org/html/2607.06648#S4)–§[5](https://arxiv.org/html/2607.06648#S5)are the mean across the four sampling seeds; uncertainty is summarised as±1\\pm 1standard deviation across seeds\.

EpochAB\\mathrm\{AB\}AA\\mathrm\{AA\}BA\\mathrm\{BA\}BB\\mathrm\{BB\}Clean acc517581100\.52810191446180\.6811522038090\.7901612716115210\.857177426120130\.8862010024211100\.9013077311690\.96348 \(best\)65326590\.965Table 6:COCONUT on ProsQA: joint label counts on the405405oracle pairs across selected checkpoints\. The AB→\\toAA collapse between epoch1515and1616is invisible from clean accuracy\.
#### Robustness across seeds\.

For COCONUT at the first and last evaluated checkpoint undernorm\-noise, the four\-seedOCR\\mathrm\{OCR\}mean±\\pmstandard deviation is ProsQA0\.139±0\.0090\.139\\pm 0\.009/0\.000±0\.0000\.000\\pm 0\.000, GSM\-Choice0\.316±0\.0130\.316\\pm 0\.013/0\.101±0\.0070\.101\\pm 0\.007, and GSM\-open0\.641±0\.0110\.641\\pm 0\.011/0\.874±0\.0170\.874\\pm 0\.017\. All four seeds preserve the qualitative direction of each trajectory \(decaying for ProsQA and GSM\-Choice, rising for GSM\-open\)\. We therefore report the cohort pattern around the boundary in §[4](https://arxiv.org/html/2607.06648#S4)descriptively rather than as a standalone significance claim\.

### A\.4Additional Results

We report additional quantitative results to support the findings in the main text\. Table[6](https://arxiv.org/html/2607.06648#A1.T6)lists the core joint label counts for COCONUT on ProsQA across selected checkpoints, highlighting the rapidAB→AA\\mathrm\{AB\}\\to\\mathrm\{AA\}flip at the epoch1515–1616boundary\.

### A\.5The Use of Large Language Models

We used LLMs for limited writing support, including refining grammar and improving language clarity and fluency\. The scientific content, experimental design, analysis, claims, and citations were produced and verified by the authors\. The authors reviewed and edited all LLM\-generated content and take full responsibility for the final submission\.

### A\.6Notation and Terminology

Table[7](https://arxiv.org/html/2607.06648#A1.T7)collects the symbols and terms used across the paper\.

Symbol / TermMeaning*Training paradigms*CoTChain\-of\-thought baseline: generates explicit step\-by\-step text tokens before the answer\.NoCoTAnswer\-only baseline: trained to emit the answer directly with no intermediate text\.COCONUTLatent\-reasoning method that replaces explicit reasoning tokens with continuous thoughts via a curriculum\(hao2025\_COCONUT\)\.CODILatent\-reasoning method that compresses CoT into latent thoughts via self\-distillation\(shen2025\_codi\)\.*Datasets*ProsQAReachability QA over small directed graphs; each question has a binary candidate answer\.GSM\-openOpen\-ended numerical answer format of GSM8K\(cobbe2021\_gsm8k\)\.GSM\-ChoiceTwo\-option multiple\-choice reformulation of GSM8K aligned with the binary interface of ProsQA\.*Pairs, inputs, and oracle answers*xix\_\{i\},xi′x^\{\\prime\}\_\{i\}Original and perturbed input for pairii\.AiA\_\{i\},BiB\_\{i\}Original oracle answer and perturbed oracle answer\.t−t^\{\-\}Negative target node used to construct the perturbation\.y^Mt​\(⋅\)\\hat\{y\}\_\{M\_\{t\}\}\(\\cdot\)Predicted answer of checkpointMtM\_\{t\}on a given input\.*Joint labels \(per pair\)*ℓtM​\(i\)∈\{A,B,O\}2\\ell\_\{t\}^\{M\}\(i\)\\in\\\{A,B,O\\\}^\{2\}Joint label encoding \(original\-answer, perturbed\-answer\) for checkpointMtM\_\{t\}on pairii, whereOOis any answer outside\{Ai,Bi\}\\\{A\_\{i\},B\_\{i\}\\\}\.AB\\mathrm\{AB\}Original→A\\to A, perturbed→B\\to B; counterfactually adaptive\.AA\\mathrm\{AA\}Original→A\\to A, perturbed→A\\to A; original\-answer retention\.AO\\mathrm\{AO\}Original→A\\to A, perturbed→O\\to O; correct on the original, off\-target on the perturbed input\.*Behavioural metrics*acctM\\mathrm\{acc\}\_\{t\}^\{M\}Original\-question accuracy at checkpointMtM\_\{t\}\.CC​\(Mt\)\\mathrm\{CC\}\(M\_\{t\}\)Contrast Consistency,Pr⁡\[y^​\(xi\)=Ai∧y^​\(xi′\)=Bi\]\\Pr\[\\hat\{y\}\(x\_\{i\}\)=A\_\{i\}\\land\\hat\{y\}\(x^\{\\prime\}\_\{i\}\)=B\_\{i\}\]\(gardner\-etal\-2020\-evaluating\)\.CFR​\(Mt\)\\mathrm\{CFR\}\(M\_\{t\}\)Conditional flip rate,Pr⁡\[y^​\(xi′\)=Bi∣y^​\(xi\)=Ai\]\\Pr\[\\hat\{y\}\(x^\{\\prime\}\_\{i\}\)=B\_\{i\}\\mid\\hat\{y\}\(x\_\{i\}\)=A\_\{i\}\]\.*Mechanistic metrics*OCR\\mathrm\{OCR\}Output change rate,Pr⁡\[y^patch≠y^clean\]\\Pr\[\\hat\{y\}^\{\\text\{patch\}\}\\neq\\hat\{y\}^\{\\text\{clean\}\}\]\.PWC\\mathrm\{PWC\}Preserved\-when\-correct rate,Pr⁡\[y^patch=y∣y^clean=y\]\\Pr\[\\hat\{y\}^\{\\text\{patch\}\}=y\\mid\\hat\{y\}^\{\\text\{clean\}\}=y\]\.accc\\mathrm\{acc\}\_\{c\},accp\\mathrm\{acc\}\_\{p\}Clean and patched accuracy on the same evaluation set\.Pr⁡\[c→w\]\\Pr\[\\mathrm\{c\}\\to\\mathrm\{w\}\]Directional flip: items correct in the clean pass that become wrong after patching\.Pr⁡\[w→c\]\\Pr\[\\mathrm\{w\}\\to\\mathrm\{c\}\]Directional flip: items wrong in the clean pass that become correct after patching\.mim\_\{i\}Answer\-margin,mi=log⁡p​\(Bi\)−log⁡p​\(Ai\)m\_\{i\}=\\log p\(B\_\{i\}\)\-\\log p\(A\_\{i\}\)\.Δ​mi\\Delta m\_\{i\}Intervention\-induced margin shift,Δ​mi=mipatch−miclean\\Delta m\_\{i\}=m\_\{i\}^\{\\text\{patch\}\}\-m\_\{i\}^\{\\text\{clean\}\}\.*Content\-erasing interventions*zeroReplace the latent activation at the target position with the zero vector\.meanReplace with the corresponding\-position mean activation from a held\-out calibration batch\.norm\-noiseReplace with isotropic Gaussian noise rescaled to the original latent norm\.*Cohorts \(ProsQA×\\timesCOCONUT\)*GstableG\_\{\\text\{stable\}\}Pairs withℓ15=ℓ48=AB\\ell\_\{15\}=\\ell\_\{48\}=\\mathrm\{AB\}\(n=41n=41\); counterfactually adaptive at epochs1515and4848\.GdriftG\_\{\\text\{drift\}\}Pairs withℓ15=AB\\ell\_\{15\}=\\mathrm\{AB\},ℓ16=AA\\ell\_\{16\}=\\mathrm\{AA\},ℓ48=AA\\ell\_\{48\}=\\mathrm\{AA\}\(n=62n=62\); adaptive at epoch1515and retains the original answer at epochs1616and4848\.GretainG\_\{\\text\{retain\}\}Pairs withℓ15=ℓ48=AA\\ell\_\{15\}=\\ell\_\{48\}=\\mathrm\{AA\}\(n=35n=35\); retains the original answer at epochs1515and4848\.*Checkpoints and curriculum*MtM\_\{t\}Model at training checkpointtt\.TTIndex of the last saved checkpoint in a training trajectory, so\{Mt\}t=1T\\\{M\_\{t\}\\\}\_\{t=1\}^\{T\}denotes the full saved\-checkpoint trajectory of a run\.nlatentn\_\{\\text\{latent\}\}Number of latent positions in the current curriculum stage\.cthoughtc\_\{\\text\{thought\}\}COCONUT hyperparameter: continuous thoughts inserted per replaced reasoning step\.Table 7:Notation and terminology used throughout the paper\.

Similar Articles

Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability

Hugging Face Daily Papers

This paper investigates whether training LLMs for shorter, more efficient chain-of-thought reasoning harms the faithfulness and monitorability of that reasoning. It finds faithfulness generally declines due to reduced model consistency, while monitorability remains relatively robust even with much shorter CoT outputs.

Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning

arXiv cs.LG

This paper proposes a framework to evaluate and improve faithfulness of chain-of-thought reasoning by controlling information flow, using entropy-based, KL-divergence, and gradient-based diagnostics, and introduces training interventions (attention masking, gradient masking, adversarial perturbations) that make reasoning more transparent and reduce shortcut reliance.

Revisiting Complete Reasoning Traces for Post-Training

Hugging Face Daily Papers

This paper finds that large language models can gain reasoning improvements from truncated reasoning trajectories rather than full ones during post-training, reducing redundancy while benefiting methods like supervised fine-tuning and reinforcement learning.

Measuring AI Faithfulness-For Better or For Worse

Reddit r/AI_Agents

This article discusses the importance of faithfulness in LLM optimization, introducing a Structural Fidelity Score that measures drift across word overlap, constraint survival, and task-type match to ensure prompt optimization does not sacrifice intent.