将推理路径重构为相位结构化的轨迹

arXiv cs.AI 论文

摘要

该论文提出了 PAIR(Phase-Aligned Intra-question Reasoning,相位对齐的题内推理),将 LLM 的推理路径视为相位结构化的轨迹,并且只在同一道问题内部评估路径质量。研究发现,标准的正确性探测在一定程度上依赖于题目层面的差异,而逐相位的引导能够提供关于路径质量方向的因果证据。

arXiv:2609.36461v1 Announce Type: new Abstract: Large language models often improve problem-solving performance by generating multi-step reasoning paths, yet how to analyze the hidden states along these paths remains unclear. Existing approaches typically assign each intermediate state the final-answer correctness label and train probes across heterogeneous questions. We argue that this protocol obscures reasoning dynamics in two ways: (1) correctness prediction can exploit question-level variation rather than path quality, and (2) states aligned by absolute step indices may correspond to different functional phases of reasoning. In this work, we propose to view reasoning paths as phase-structured trajectories within fixed questions. We instantiate this view as PAIR, short for Phase-Aligned Intra-question Reasoning. PAIR samples multiple trajectories for each question, maps variable-length paths into shared relative phases based on normalized trajectory progress, and compares successful and unsuccessful trajectories only within the same question and phase. This yields phase-specific path-quality directions that better isolate path-quality signals from question-level variation. Empirically, we find that standard across-question correctness probes lose much of their predictive power under within-question evaluation, suggesting that these probes partly rely on question-level information. PAIR improves within-question trajectory ranking and Best-of-N trajectory selection across models and benchmarks. Phase-wise steering further shows that the learned directions can change generation outcomes, providing causal evidence that they capture trajectory-relevant information.
查看原文
查看缓存全文

缓存时间: 2026/09/30 09:45

# Rethinking Reasoning Paths as Phase-Structured Trajectories
Source: [https://arxiv.org/html/2609.36461](https://arxiv.org/html/2609.36461)
Zhenghao HeGuangzhi XiongSanchit SinhaBohan LiuWenqian YeAidong ZhangDepartment of Computer Science, University of VirginiaAffiliation:\{zhenghao, aidong\}@virginia\.edu

###### Abstract

Large language models often improve problem\-solving performance by generating multi\-step reasoning paths, yet how to analyze the hidden states along these paths remains unclear\. Existing approaches typically assign each intermediate state the final\-answer correctness label and train probes across heterogeneous questions\. We argue that this protocol obscures reasoning dynamics in two ways: \(1\) correctness prediction can exploit question\-level variation rather than path quality, and \(2\) states aligned by absolute step indices may correspond to different functional phases of reasoning\. In this work, we propose to view reasoning paths as phase\-structured trajectories within fixed questions\. We instantiate this view asPAIR, short forPhase\-AlignedIntra\-questionReasoning\. PAIR samples multiple trajectories for each question, maps variable\-length paths into shared relative phases based on normalized trajectory progress, and compares successful and unsuccessful trajectories only within the same question and phase\. This yields phase\-specific path\-quality directions that better isolate path\-quality signals from question\-level variation\. Empirically, we find that standard across\-question correctness probes lose much of their predictive power under within\-question evaluation, suggesting that these probes partly rely on question\-level information\. PAIR improves within\-question trajectory ranking and Best\-of\-NNtrajectory selection across models and benchmarks\. Phase\-wise steering further shows that the learned directions can change generation outcomes, providing causal evidence that they capture trajectory\-relevant information\.

## 1Introduction

Large language models can improve their problem\-solving performance by generating multi\-step reasoning paths before producing final answers\. This behavior is most explicit in chain\-of\-thought prompting\([Wei et al\., 2022](https://arxiv.org/html/2609.36461#bib.bib2)\), where models are encouraged to write intermediate reasoning steps, and has been shown to improve performance on mathematical reasoning\([Cobbe et al\., 2021](https://arxiv.org/html/2609.36461#bib.bib19)\), symbolic reasoning, and code generation tasks\([Chen et al\., 2021](https://arxiv.org/html/2609.36461#bib.bib3)\)\. The accuracy gains from this style of generation are substantial, but the internal mechanism by which it succeeds or fails remains opaque\. As a model generates a reasoning path,*what information about the eventual correctness is encoded in its hidden states, and how does this information evolve along the trajectory?*

Recent work has increasingly analyzed language model reasoning through internal representations, using hidden\-state probes, representation geometry, and activation directions to uncover task\-relevant or correctness\-related signals\([Burns et al\.,](https://arxiv.org/html/2609.36461#bib.bib10);[Li et al\., 2023](https://arxiv.org/html/2609.36461#bib.bib12);[He et al\., 2026](https://arxiv.org/html/2609.36461#bib.bib4)\)\. In many such analyses, intermediate hidden states are treated as pointwise representations labeled by the final outcome, e\.g\.,y=𝕀\[a^=a∗\]y=\\mathbb\{I\}\[\\hat\{a\}=a^\{\\ast\}\], and a probe or scorer is trained to predict final correctness from each state \(Fig[1](https://arxiv.org/html/2609.36461#S1.F1)a\)\.

However, such across\-question correctness prediction can conflate path quality with question\-level variation \(Fig[1](https://arxiv.org/html/2609.36461#S1.F1)b\)\. When examples are pooled across heterogeneous questions, the correctness label depends not only on the generated trajectory, but also on properties of the question being solved\. This makes early correctness signals hard to interpret: strong prediction before substantial reasoning has occurred may reflect question\-level information rather than path\-level progress\.

![Refer to caption](https://arxiv.org/html/2609.36461v1/intro.png)Figure 1:Overview of our motivation\. \(a\) Prior probing labels intermediate states by final\-answer correctness across different questions\. \(b\) Across\-question prediction can exploit question identity rather than path quality\. \(c\) Absolute step indices can misalign reasoning phases across questions\. \(d\) We instead align trajectories by phase and compare paths within the same question\.A second issue concerns how intermediate states are compared along the trajectory\. Final correctness is a trajectory\-level outcome, but pointwise probing assigns this same label to every intermediate state along the path\. Recent work has begun to study reasoning as a structured trajectory, showing that step\-specific activations can form separable subspaces and that correct and incorrect trajectories diverge at later stages\([Sun et al\., 2026](https://arxiv.org/html/2609.36461#bib.bib5)\)\. These results suggest that reasoning has meaningful temporal geometry\. However, existing trajectory analyses often align states by explicit step markers or absolute step indices\. This creates a structural mismatch: “Step 2” can mean setup in one solution and the main calculation in another \(Fig[1](https://arxiv.org/html/2609.36461#S1.F1)c\)\.

In this work, we view reasoning paths as phase\-structured trajectories within fixed questions \(Fig\.[1](https://arxiv.org/html/2609.36461#S1.F1)d\)\. A useful comparison should hold the question fixed, so that question\-level information cannot serve as the main shortcut, and should compare states that occupy comparable positions within their trajectories\. Under this view, the relevant question is not whether a hidden state predicts final correctness across a heterogeneous dataset, but whether one trajectory state is better than another for the same question and at the same relative phase\.

We instantiate this view as PAIR, short forPhase\-AlignedIntra\-questionReasoning\. PAIR samples multiple reasoning trajectories for each question, maps variable\-length paths to a shared set of phase slots, and constructs phase\-matched comparisons between successful and unsuccessful trajectories from the same question\. From these comparisons, PAIR learns phase\-specific directions that rank successful states above unsuccessful states\. We then use these directions for analysis and steering, testing whether they can actively redirect generation\.

Empirically, we find that standard across\-question correctness signals weaken substantially under within\-question evaluation, suggesting that part of their apparent strength comes from question\-level information\. We further show that phase alignment reveals trajectory\-quality structure that is obscured by absolute\-step comparisons\. The learned phase\-specific directions improve within\-question trajectory ranking and Best\-of\-NNtrajectory selection\. Finally, phase\-wise steering shows that these directions can change generation outcomes, providing causal evidence that they capture trajectory\-relevant information\.

Our contributions are as follows:

- •We diagnose key ambiguities in pointwise correctness analysis, showing that across\-question probes can capture question\-level variation and that absolute step indices can misalign reasoning stages\.
- •We introduce PAIR, a phase\-aligned intra\-question framework that compares successful and unsuccessful trajectories from the same question at matched relative phases to learn path\-quality directions\.
- •We demonstrate that PAIR improves trajectory ranking and Best\-of\-NNselection, and use phase\-wise steering to provide causal evidence that the learned directions can affect reasoning generation\.

## 2Related Work

##### Reasoning Paths and Multi\-trajectory Inference\.

Chain\-of\-thought prompting improves language\-model reasoning by eliciting intermediate steps before the final answer\([Wei et al\., 2022](https://arxiv.org/html/2609.36461#bib.bib2);[Kojima et al\., 2022](https://arxiv.org/html/2609.36461#bib.bib6)\)\. Beyond single\-path generation, self\-consistency samples multiple reasoning paths for the same question and aggregates their answers\([Wang et al\., 2022](https://arxiv.org/html/2609.36461#bib.bib7)\), while tree\- and graph\-based methods cast reasoning as search over candidate thoughts or states\([Yao et al\., 2023](https://arxiv.org/html/2609.36461#bib.bib8);[Besta et al\., 2024](https://arxiv.org/html/2609.36461#bib.bib9)\)\. Other work evaluates or supervises intermediate reasoning through process supervision and step\-level verification\([Lightman et al\., 2023](https://arxiv.org/html/2609.36461#bib.bib20)\)\. These methods exploit path structure to improve answer selection or reasoning performance; in contrast, we use multiple trajectories for the same question to study how reasoning quality is represented in hidden states\.

##### Representation Analysis of Reasoning Trajectories\.

Recent work analyzes model behavior through internal representations, including latent knowledge, truthfulness, confidence, and task\-relevant structure\([Burns et al\.,](https://arxiv.org/html/2609.36461#bib.bib10);[Azaria and Mitchell, 2023](https://arxiv.org/html/2609.36461#bib.bib11)\), as well as representation geometry and activation directions for characterizing or intervening on behavior\([Li et al\., 2023](https://arxiv.org/html/2609.36461#bib.bib12);[Park et al\., 2023](https://arxiv.org/html/2609.36461#bib.bib13)\)\. In reasoning tasks, intermediate hidden states have been shown to encode signals related to final\-answer correctness, self\-verification, and reasoning progress before the answer is generated\([Zhang et al\., 2025](https://arxiv.org/html/2609.36461#bib.bib1);[Liu et al\., 2025](https://arxiv.org/html/2609.36461#bib.bib14)\)\. Most closely related to our work,[Sun et al\. \(2026\)](https://arxiv.org/html/2609.36461#bib.bib5)view reasoning as a structured trajectory in representation space, showing that activations near explicit reasoning markers form step\-specific subspaces and that correct and incorrect trajectories may diverge over generation\. Our work shares this trajectory\-level view but addresses a different comparison problem: absolute step indices are not always comparable across trajectories, since the second step of one solution may be problem setup while the second step of another may already contain the main computation\. We therefore align variable\-length trajectories by relative phase and compare states within the same question and phase\.

##### Correctness Signals and Activation Steering\.

Although correctness\-related information can often be decoded from intermediate hidden states\([Zhang et al\., 2025](https://arxiv.org/html/2609.36461#bib.bib1)\), such signals are hard to interpret when examples are pooled across heterogeneous questions\. A predictor trained across questions may capture question\-level information, such as difficulty or prior solvability, rather than the quality of a particular reasoning trajectory\. PAIR instead compares successful and unsuccessful trajectories within the same question and aligned phase, avoiding across\-question correctness predictors\. For causal validation, we apply the resulting phase\-specific directions through activation steering\. Prior work shows that hidden\-state directions can steer model behavior at inference time\([Turner et al\., 2023](https://arxiv.org/html/2609.36461#bib.bib15)\), and contrastive activation methods further learn behavior\-changing directions from paired examples\([Panickssery et al\., 2023](https://arxiv.org/html/2609.36461#bib.bib16)\)\.

## 3PAIR: Phase\-Aligned Intra\-question Reasoning

In this section, we introduce PAIR, a phase\-aligned intra\-question framework for analyzing and steering reasoning trajectories\. PAIR is built on the premise that reasoning paths should not be compared as absolute\-step sequences across heterogeneous questions\. Instead, paths should be compared within the same question and at comparable relative phases of generation\.

As shown in Fig\.[2](https://arxiv.org/html/2609.36461#S3.F2), PAIR consists of three stages\. First, for each question, we sample multiple reasoning trajectories and align their hidden states into a fixed set of relative phase slots\. Second, within each phase, we construct successful–unsuccessful trajectory pairs from the same question and learn a phase\-specific path\-quality direction through a pairwise ranking objective\. Third, we use the learned directions to steer hidden states during generation, thereby testing whether phase\-specific path\-quality signals can causally redirect reasoning trajectories\.

![Refer to caption](https://arxiv.org/html/2609.36461v1/neurips_framework.png)Figure 2:Overview of PAIR\. \(a\) For each question, PAIR samples multiple reasoning trajectories and aligns variable\-length paths into a fixed set of relative phase slots\. \(b\) Within each phase, PAIR constructs successful–unsuccessful trajectory pairs from the same question and learns a phase\-specific path\-quality direction using a pairwise ranking objective\. \(c\) The learned directions are used for phase\-wise activation steering, where hidden states are shifted toward successful trajectory directions during generation\.### 3\.1Problem Setup

Letqqdenote a question with ground\-truth answeraq∗a\_\{q\}^\{\\ast\}\. For each question, we sample multiple reasoning trajectories from the model\. Theii\-th trajectory is written as

τq,i=\(xq,i,1:Tq,i,Hq,i,1:Tq,i,a^q,i\),\\tau\_\{q,i\}=\\left\(x\_\{q,i,1:T\_\{q,i\}\},H\_\{q,i,1:T\_\{q,i\}\},\\hat\{a\}\_\{q,i\}\\right\),\(1\)wherexq,i,1:Tq,ix\_\{q,i,1:T\_\{q,i\}\}are generated tokens,Hq,i,1:Tq,iH\_\{q,i,1:T\_\{q,i\}\}are token\-level hidden states, anda^q,i\\hat\{a\}\_\{q,i\}is the final answer\. The trajectory\-level outcome is

yq,i=𝕀\[a^q,i=aq∗\]\.y\_\{q,i\}=\\mathbb\{I\}\\left\[\\hat\{a\}\_\{q,i\}=a\_\{q\}^\{\\ast\}\\right\]\.\(2\)
As discussed in Sec\.[1](https://arxiv.org/html/2609.36461#S1), directly usingyq,iy\_\{q,i\}as a label for every intermediate state can conflate local reasoning quality with final\-answer correctness and question\-level variation\. Our goal is therefore not to predict final correctness from isolated hidden states\. Instead, we ask whether one trajectory state is better than another when the question and the reasoning phase are matched\.

After phase alignment, lethq,i,b\(ℓ\)h\_\{q,i,b\}^\{\(\\ell\)\}denote the layer\-ℓ\\ellrepresentation of trajectoryτq,i\\tau\_\{q,i\}at phasebb\. For two trajectories sampled from the same question, we define the phase\-wise path\-quality preference:

hq,i,b\(ℓ\)≻hq,j,b\(ℓ\)ifyq,i=1,yq,j=0\.h\_\{q,i,b\}^\{\(\\ell\)\}\\succ h\_\{q,j,b\}^\{\(\\ell\)\}\\quad\\text\{if\}\\quad y\_\{q,i\}=1,\\;y\_\{q,j\}=0\.\(3\)Thus, final correctness is used to form preferences between trajectories matched by question and phase, rather than to label individual hidden states\.

### 3\.2Phase Alignment of Reasoning Trajectories

PAIR first converts sampled reasoning trajectories into phase\-aligned step representations\. This stage has two components: extracting step\-level features from each trajectory and mapping variable\-length step sequences to shared relative phases\.

##### Step\-level trajectory representation\.

For each questionqq, we samplemmreasoning trajectories under the same prompting and decoding configuration\. The trajectory and outcome notation follows Eqs\.[1](https://arxiv.org/html/2609.36461#S3.E1)–[2](https://arxiv.org/html/2609.36461#S3.E2)\. Following[Sun et al\. \(2026\)](https://arxiv.org/html/2609.36461#bib.bib5), we extract step\-level representations using explicit reasoning markers\. We first detect step markers, such as “Step 1”, numbered\-list markers, and discourse markers such as “First”; the full marker set is given in Appendix[A\.4\.1](https://arxiv.org/html/2609.36461#A1.SS4.SSS1)\. These markers indicate the beginning of reasoning steps\.

Letpq,i,kp\_\{q,i,k\}denote the token position of thekk\-th detected step marker in trajectoryτq,i\\tau\_\{q,i\}, and letcq,ic\_\{q,i\}denote the position of the conclusion marker\. If a trajectory containsLq,iL\_\{q,i\}reasoning steps, we represent stepjjusing therrtokens immediately before the next step marker; the final step is represented using therrtokens before the conclusion marker:

hq,i,j\(ℓ\)=\{Pool⁡\(\{Hq,i,t\(ℓ\):pq,i,j\+1−r≤t<pq,i,j\+1\}\),j<Lq,i,Pool⁡\(\{Hq,i,t\(ℓ\):cq,i−r≤t<cq,i\}\),j=Lq,i,h\_\{q,i,j\}^\{\(\\ell\)\}=\\begin\{cases\}\\operatorname\{Pool\}\\\!\\left\(\\left\\\{H\_\{q,i,t\}^\{\(\\ell\)\}:p\_\{q,i,j\+1\}\-r\\leq t<p\_\{q,i,j\+1\}\\right\\\}\\right\),&j<L\_\{q,i\},\\\\\[4\.0pt\] \\operatorname\{Pool\}\\\!\\left\(\\left\\\{H\_\{q,i,t\}^\{\(\\ell\)\}:c\_\{q,i\}\-r\\leq t<c\_\{q,i\}\\right\\\}\\right\),&j=L\_\{q,i\},\\end\{cases\}\(4\)whereHq,i,t\(ℓ\)H\_\{q,i,t\}^\{\(\\ell\)\}is the token\-level hidden state at layerℓ\\ell, andPool⁡\(⋅\)\\operatorname\{Pool\}\(\\cdot\)denotes window aggregation\. This gives each trajectory a step\-level hidden\-state sequence\{hq,i,j\(ℓ\)\}j=1Lq,i\\\{h\_\{q,i,j\}^\{\(\\ell\)\}\\\}\_\{j=1\}^\{L\_\{q,i\}\}\.

##### Relative phase alignment\.

Since trajectories may contain different numbers of reasoning steps, absolute step indices are not directly comparable\. PAIR maps each trajectory toKKrelative phase slots\. We setKKto the mode of the step counts in the training trajectories:

K=mode\(\{Lq,i:q∈𝒟train,i=1,…,m\}\)\.K=\\operatorname\{mode\}\\left\(\\left\\\{L\_\{q,i\}:q\\in\\mathcal\{D\}\_\{\\mathrm\{train\}\},\\;i=1,\\ldots,m\\right\\\}\\right\)\.\(5\)
For a trajectory withLq,iL\_\{q,i\}steps, stepjjis assigned to its nearest relative phase:

bq,i,j=1\+round⁡\(\(j−1\)​\(K−1\)max⁡\(Lq,i−1,1\)\),bq,i,j∈\{1,…,K\}\.b\_\{q,i,j\}=1\+\\operatorname\{round\}\\left\(\\frac\{\(j\-1\)\(K\-1\)\}\{\\max\(L\_\{q,i\}\-1,1\)\}\\right\),\\quad b\_\{q,i,j\}\\in\\\{1,\\ldots,K\\\}\.\(6\)This maps the first step toP1P\_\{1\}and the final step toPKP\_\{K\}, while intermediate steps are placed according to normalized progress\. As illustrated in Fig\.[2](https://arxiv.org/html/2609.36461#S3.F2)\(a\), shorter trajectories may skip some phase slots, whereas longer trajectories may assign multiple steps to the same phase\.

For each trajectory and phase, we collect the step representations assigned to that phase:

ℋq,i,b\(ℓ\)=\{hq,i,j\(ℓ\):bq,i,j=b\}\.\\mathcal\{H\}\_\{q,i,b\}^\{\(\\ell\)\}=\\left\\\{h\_\{q,i,j\}^\{\(\\ell\)\}:b\_\{q,i,j\}=b\\right\\\}\.\(7\)The resulting phase\-indexed sets are used to construct within\-question comparisons in the next stage\.

### 3\.3Learning Phase\-specific Path\-quality Directions

After phase alignment, PAIR learns one path\-quality direction for each phase\.

For phasebb, letℋq,i,b\(ℓ\)\\mathcal\{H\}\_\{q,i,b\}^\{\(\\ell\)\}be the set of step representations from trajectoryτq,i\\tau\_\{q,i\}assigned to phasebb, as defined in Eq\.[7](https://arxiv.org/html/2609.36461#S3.E7)\. We collect successful and unsuccessful phase\-bbstates for questionqq:

ℋq,b\+=⋃i:yq,i=1ℋq,i,b\(ℓ\),ℋq,b−=⋃i:yq,i=0ℋq,i,b\(ℓ\)\.\\mathcal\{H\}\_\{q,b\}^\{\+\}=\\bigcup\_\{i:\\,y\_\{q,i\}=1\}\\mathcal\{H\}\_\{q,i,b\}^\{\(\\ell\)\},\\qquad\\mathcal\{H\}\_\{q,b\}^\{\-\}=\\bigcup\_\{i:\\,y\_\{q,i\}=0\}\\mathcal\{H\}\_\{q,i,b\}^\{\(\\ell\)\}\.\(8\)
For each questionqq, PAIR constructs within\-question preference pairs:

𝒫q,b=\{\(h\+,h−\):h\+∈ℋq,b\+,h−∈ℋq,b−\}\.\\mathcal\{P\}\_\{q,b\}=\\left\\\{\(h^\{\+\},h^\{\-\}\):h^\{\+\}\\in\\mathcal\{H\}\_\{q,b\}^\{\+\},\\;h^\{\-\}\\in\\mathcal\{H\}\_\{q,b\}^\{\-\}\\right\\\}\.\(9\)Each pair states that, for the same question and same phase, the state from a successful trajectory should receive a higher path\-quality score than the state from an unsuccessful trajectory\.

For each phasebb, we fit a linear scorer after phase\-specific preprocessing\. Letϕb​\(⋅\)\\phi\_\{b\}\(\\cdot\)denote the preprocessing map fitted on the training states for phasebb\. The phase\-bbscore is

sb​\(h\)=w~b⊤​ϕb​\(h\)\+βb,s\_\{b\}\(h\)=\\tilde\{w\}\_\{b\}^\{\\top\}\\phi\_\{b\}\(h\)\+\\beta\_\{b\},\(10\)wherew~b\\tilde\{w\}\_\{b\}is the learned direction in the preprocessed space\. PAIR optimizes the pairwise logistic ranking objective:

ℒb=−∑q∑\(h\+,h−\)∈𝒫q,blogσ\(sb\(h\+\)−sb\(h−\)\)\+λ∥w~b∥22\.\\mathcal\{L\}\_\{b\}=\-\\sum\_\{q\}\\sum\_\{\(h^\{\+\},h^\{\-\}\)\\in\\mathcal\{P\}\_\{q,b\}\}\\log\\sigma\\left\(s\_\{b\}\(h^\{\+\}\)\-s\_\{b\}\(h^\{\-\}\)\\right\)\+\\lambda\\\|\\tilde\{w\}\_\{b\}\\\|\_\{2\}^\{2\}\.\(11\)
Optimizing Eq\.[11](https://arxiv.org/html/2609.36461#S3.E11)yields a phase\-specific path\-quality direction, which we map back to the original hidden\-state space and normalize aswbw\_\{b\}for analysis and steering\.

### 3\.4Phase\-wise Trajectory Steering

The learned directions are used as causal interventions on the reasoning trajectory\. As illustrated in Fig\.[2](https://arxiv.org/html/2609.36461#S3.F2)\(c\), PAIR modifies the hidden state according to the direction associated with the current reasoning phase\.

Letwbw\_\{b\}be the normalized raw\-space path\-quality direction for phasebb\. During generation, letHt\(ℓ\)H\_\{t\}^\{\(\\ell\)\}denote the hidden state at layerℓ\\elland decoding positiontt\. Letbt∈\{1,…,K\}b\_\{t\}\\in\\\{1,\\ldots,K\\\}be the phase assigned to the current position by the steering schedule\. PAIR applies the intervention

H~t\(ℓ\)=Ht\(ℓ\)\+λt​wbt,\\widetilde\{H\}\_\{t\}^\{\(\\ell\)\}=H\_\{t\}^\{\(\\ell\)\}\+\\lambda\_\{t\}w\_\{b\_\{t\}\},\(12\)whereλt\\lambda\_\{t\}is the intervention strength at positiontt\. The modified stateH~t\(ℓ\)\\widetilde\{H\}\_\{t\}^\{\(\\ell\)\}is passed to the remaining layers, with all model parameters fixed\.

Rather than using a constant strength, we use the phase\-specific score to adapt the intervention magnitude\. Lets~b​\(h\)\\widetilde\{s\}\_\{b\}\(h\)denote the learned phase\-bbpath\-quality score, and letτb\\tau\_\{b\}be a reference score estimated from successful training trajectories at phasebb\. The intervention strength is

λt=min⁡\{λmax,α​\[τbt−s~bt​\(Ht\(ℓ\)\)\]\+\},\\lambda\_\{t\}=\\min\\left\\\{\\lambda\_\{\\max\},\\;\\alpha\\left\[\\tau\_\{b\_\{t\}\}\-\\widetilde\{s\}\_\{b\_\{t\}\}\\left\(H\_\{t\}^\{\(\\ell\)\}\\right\)\\right\]\_\{\+\}\\right\\\},\(13\)where\[x\]\+=max⁡\(x,0\)\[x\]\_\{\+\}=\\max\(x,0\),α\\alphacontrols the global steering scale, andλmax\\lambda\_\{\\max\}clips overly large updates\. Thus, PAIR pushes a state only when its phase\-specific score falls below the reference level, and the clipping term prevents unstable interventions\.

We evaluate two intervention schedules\. In single\-phase steering, Eq\.[12](https://arxiv.org/html/2609.36461#S3.E12)is applied only at one selected phase\. In progressive steering, the intervention follows the trajectory phase and uses the corresponding directionwbtw\_\{b\_\{t\}\}as generation advances\. The threshold construction, clipping value, and phase\-schedule implementation are given in Appendix[A\.4\.3](https://arxiv.org/html/2609.36461#A1.SS4.SSS3)\.

## 4Experiments and Results

We structure our experiments to test the main claim that reasoning trajectories require phase\-aligned within\-question comparison\. We first diagnose the limitations of standard across\-question correctness probing, showing that it can conflate path quality with question\-level variation and step misalignment \(Sec\.[4\.2](https://arxiv.org/html/2609.36461#S4.SS2)\)\. We then evaluate whether PAIR learns phase\-specific path\-quality directions that improve within\-question ranking of successful and unsuccessful trajectories \(Sec\.[4\.3](https://arxiv.org/html/2609.36461#S4.SS3)\)\. Finally, we use phase\-wise activation steering to test whether these directions can causally redirect generation rather than merely correlate with final correctness \(Sec\.[4\.4](https://arxiv.org/html/2609.36461#S4.SS4)\)\.

### 4\.1Experimental Setup

##### Models and benchmarks\.

We evaluate on three reasoning benchmarks: GSM8K[Cobbe et al\. \(2021\)](https://arxiv.org/html/2609.36461#bib.bib19), MATH500[Lightman et al\. \(2023\)](https://arxiv.org/html/2609.36461#bib.bib20), and the logical\_deduction\_three\_objects task from Big\-Bench Hard \(BBH\)[Suzgun et al\. \(2022\)](https://arxiv.org/html/2609.36461#bib.bib18)\. We use four open\-weight models: LLaMA\-3\.1\-8B\-Instruct[Meta \(2024\)](https://arxiv.org/html/2609.36461#bib.bib21), Qwen3\-4B[Team \(2025\)](https://arxiv.org/html/2609.36461#bib.bib23), Gemma3\-4B[Team et al\. \(2025\)](https://arxiv.org/html/2609.36461#bib.bib22), and DeepSeek\-R1\-Distill\-Llama\-8B[DeepSeek\-AI \(2025\)](https://arxiv.org/html/2609.36461#bib.bib24)\. For GSM8K, we train probes on 1,000 questions from the official training split and evaluate on 500 questions from the official test split\. For MATH and BBH, we use a fixed 70/30 question\-level split, sampled once and reused across all methods and models\. All evaluations are conducted on held\-out test questions, so trajectories from the same question never appear in both training and test sets\.

##### Trajectory sampling and labeling\.

For each question, we sample 32 reasoning trajectories from the evaluated model with temperature1\.01\.0and top\-p=0\.95p=0\.95\. A trajectory is labeled correct if its final extracted answer matches the ground truth, using dataset\-specific extraction rules; unparsable outputs are treated as incorrect\. Sampling multiple trajectories per question lets us compare successful and unsuccessful paths under a shared question context, controlling for question\-level difficulty\.

##### Representation extraction\.

Unless otherwise specified, we probe the final transformer layer, where prior work has found the strongest signals for truthfulness and task outcomes\([Azaria and Mitchell, 2023](https://arxiv.org/html/2609.36461#bib.bib11);[Burns et al\.,](https://arxiv.org/html/2609.36461#bib.bib10);[Li et al\., 2023](https://arxiv.org/html/2609.36461#bib.bib12)\)\. For each trajectory, we locate step markers and average the hidden states of the four tokens immediately preceding each marker to form a step\-level representation\. These step\-level representations are shared across all hidden\-state scoring methods we compare, including final\-token probing, pooled correctness probing, absolute\-step alignment, normalized\-step alignment, and our phase\-aligned PAIR method\.

### 4\.2Diagnosing Ambiguities in Pointwise Correctness Analysis

Table 1:Across\-question and within\-question evaluation of correctness probes\. For each benchmark, we report across\-question \(AQ\) and within\-question \(WQ\) AUROC and balanced accuracy, together withΔ=WQ−AQ\\Delta=\\mathrm\{WQ\}\-\\mathrm\{AQ\}\. AQ/WQ entries are formatted asAQ/WQ\.ModelGSM8K\([Cobbe et al\., 2021](https://arxiv.org/html/2609.36461#bib.bib19)\)MATH\([Lightman et al\., 2023](https://arxiv.org/html/2609.36461#bib.bib20)\)BBH\([Suzgun et al\., 2022](https://arxiv.org/html/2609.36461#bib.bib18)\)AUCΔ\\DeltaBalanced Acc\.Δ\\DeltaAUCΔ\\DeltaBalanced Acc\.Δ\\DeltaAUCΔ\\DeltaBalanced Acc\.Δ\\DeltaLLaMA\-3\.1\-8B\-Instruct0\.72/0\.56\-0\.160\.66/0\.51\-0\.150\.65/0\.49\-0\.160\.64/0\.50\-0\.140\.72/0\.53\-0\.190\.66/0\.51\-0\.15Qwen3\-4B0\.72/0\.53\-0\.190\.73/0\.64\-0\.090\.64/0\.51\-0\.130\.60/0\.57\-0\.030\.77/0\.53\-0\.240\.76/0\.51\-0\.25Gemma3\-4B0\.65/0\.50\-0\.150\.63/0\.52\-0\.110\.68/0\.54\-0\.140\.61/0\.58\-0\.030\.71/0\.51\-0\.200\.71/0\.41\-0\.30DeepSeek\-R1\-Distill\-Llama\-8B0\.63/0\.62\-0\.010\.62/0\.59\-0\.030\.65/0\.51\-0\.140\.64/0\.56\-0\.080\.68/0\.51\-0\.170\.66/0\.57\-0\.09Average0\.68/0\.55\-0\.130\.66/0\.57\-0\.100\.66/0\.51\-0\.140\.62/0\.55\-0\.070\.72/0\.52\-0\.200\.70/0\.50\-0\.20

We first test whether pointwise correctness prediction remains reliable after controlling for question identity\. For each sampled trajectory, we extract the hidden states from a four\-token window around the position closest to15%15\\%relative progress and use their mean as the representation\. We then train a linear predictor using the trajectory\-level correctness labelyq,iy\_\{q,i\}\. Additional training details are given in Appendix[A\.4\.3](https://arxiv.org/html/2609.36461#A1.SS4.SSS3.Px8)\.

We report the standard across\-question evaluation \(AQ\), which tests whether the predictor can predict final correctness on held\-out questions\. We contrast it with within\-question evaluation \(WQ\), which ranks correct and incorrect trajectories sampled for the same question:

WQ\-AUC=1\|𝒬±\|∑q∈𝒬±1\|𝒫q\|​\|𝒩q\|∑i∈𝒫q∑j∈𝒩q𝕀\[s\(xq,i\+\)\>s\(xq,j−\)\],\\mathrm\{WQ\\text\{\-\}AUC\}=\\frac\{1\}\{\|\\mathcal\{Q\}\_\{\\pm\}\|\}\\sum\_\{q\\in\\mathcal\{Q\}\_\{\\pm\}\}\\frac\{1\}\{\|\\mathcal\{P\}\_\{q\}\|\|\\mathcal\{N\}\_\{q\}\|\}\\sum\_\{i\\in\\mathcal\{P\}\_\{q\}\}\\sum\_\{j\\in\\mathcal\{N\}\_\{q\}\}\\mathbb\{I\}\\left\[s\(x\_\{q,i\}^\{\+\}\)\>s\(x\_\{q,j\}^\{\-\}\)\\right\],\(14\)where𝒫q\\mathcal\{P\}\_\{q\}and𝒩q\\mathcal\{N\}\_\{q\}denote the successful and unsuccessful trajectories for questionqq, respectively, and𝒬±\\mathcal\{Q\}\_\{\\pm\}contains questions with at least one trajectory from each group\. Balanced accuracy is computed under the same within\-question protocol\.

Figure 3:Balanced AUROC of correctness prediction and forced\-answer accuracy across relative trajectory positions\.Table[1](https://arxiv.org/html/2609.36461#S4.T1)shows that AQ performance consistently overestimates the strength of correctness signals\. Across models and benchmarks, AUROC and balanced accuracy drop sharply under WQ evaluation\. For example, the average BBH AUROC decreases from0\.720\.72to0\.520\.52, and balanced accuracy decreases from0\.700\.70to0\.500\.50\. This gap indicates that across\-question correctness predictors capture substantial question\-level information, such as difficulty or solvability, rather than only trajectory\-level progress\.

To test whether early correctness signals reflect answerability, we compare correctness prediction with forced\-answer accuracy along the trajectory\. Following prior work on early correctness signals\([Zhang et al\., 2025](https://arxiv.org/html/2609.36461#bib.bib1)\), we train a two\-layer MLP predictor on window\-aggregated hidden states at different relative positions, using final\-answer correctness as the label\. Following the forced\-answer extraction protocol of[Boppana et al\. \(2026\)](https://arxiv.org/html/2609.36461#bib.bib17), we estimate answerability by truncating each trajectory at the same relative position, appending “The answer is”, and samplingk=128k=128completions with temperature0\.80\.8\. We defineacc​@​k=c/k\\mathrm\{acc@\}k=c/k, whereccis the number of completions that produce the correct answer\.

Fig\.[3](https://arxiv.org/html/2609.36461#S4.F3)shows the results on GSM8K for Llama\-3\.1\-8B\-Instruct and Qwen3\-4B\. The MLP predictor achieves high balanced AUROC at very early positions, consistent with prior findings, whileacc​@​128\\mathrm\{acc@\}128remains close to zero\. This indicates that early representations can predict final correctness even when the prefix is not yet answerable\. Across relative positions, AUROC first decreases and then increases, whereasacc​@​128\\mathrm\{acc@\}128rises as the trajectory approaches the final answer\. The late increase in AUROC is expected because later prefixes contain more information needed to recover the answer\.*The early decrease may reflect a reduced influence of question\-level information as generation moves away from the prompt\.*Together with Table[1](https://arxiv.org/html/2609.36461#S4.T1), this result supports the need to control question identity when analyzing pointwise correctness signals\.

### 4\.3Effectiveness of Phase\-Aligned Path\-quality Directions

Table 2:Evaluation of trajectory\-quality estimation and Best\-of\-32 trajectory selection\. We highlight the best value ingreenwith bold textand the second best value inblue\.ModelMethodGSM8K\([Cobbe et al\., 2021](https://arxiv.org/html/2609.36461#bib.bib19)\)MATH\([Lightman et al\., 2023](https://arxiv.org/html/2609.36461#bib.bib20)\)BBH\([Suzgun et al\., 2022](https://arxiv.org/html/2609.36461#bib.bib18)\)WQ\-AUC↑\\uparrowBoN@32↑\\uparrowWQ\-AUC↑\\uparrowBoN@32↑\\uparrowWQ\-AUC↑\\uparrowBoN@32↑\\uparrowLLaMA\-3\.1\-8B\-Instruct[Meta \(2024\)](https://arxiv.org/html/2609.36461#bib.bib21)Pre\-Ans token0\.7183\.30\.8346\.60\.6584\.0Pooled correctness0\.6684\.60\.6936\.00\.5177\.3Absolute\-step0\.6381\.60\.6336\.60\.5781\.3Normalized\-step0\.6785\.30\.7140\.00\.5868\.0PAIR0\.7687\.30\.7746\.60\.7492\.0Qwen3\-4B\([Team, 2025](https://arxiv.org/html/2609.36461#bib.bib23)\)Pre\-Ans token0\.4858\.70\.7983\.30\.8493\.3Pooled correctness0\.6174\.70\.5474\.60\.6592\.0Absolute\-step0\.5955\.30\.5273\.30\.6185\.0Normalized\-step0\.6880\.70\.6076\.00\.6294\.6PAIR0\.9180\.70\.8984\.70\.8697\.3Gemma3\-4B\([Team et al\., 2025](https://arxiv.org/html/2609.36461#bib.bib22)\)Pre\-Ans token0\.6089\.30\.6068\.70\.6794\.6Pooled correctness0\.4889\.30\.5361\.30\.5192\.0Absolute\-step0\.5583\.30\.4955\.30\.5594\.6Normalized\-step0\.5788\.00\.5466\.70\.6296\.0PAIR0\.6195\.00\.7475\.30\.7596\.0DeepSeek\-R1\-Distill\-Llama\-8B[DeepSeek\-AI \(2025\)](https://arxiv.org/html/2609.36461#bib.bib24)Pre\-Ans token0\.6586\.00\.8784\.60\.92100\.0Pooled correctness0\.6088\.00\.5372\.00\.5681\.3Absolute\-step0\.5884\.30\.6374\.60\.5069\.3Normalized\-step0\.5786\.60\.6680\.00\.84100\.0PAIR0\.7589\.30\.8786\.70\.94100\.0

We first examine how path\-quality separability changes across aligned phases\. For each phase, we train a phase\-specific scorer and evaluate it by pairwise AUROC \(Eq\.[14](https://arxiv.org/html/2609.36461#S4.E14)\) on held\-out questions, usingK∈\{3,4,5,6\}K\\in\\\{3,4,5,6\\\}phase bins\. As shown in Fig\.[4\(a\)](https://arxiv.org/html/2609.36461#S4.F4.sf1), AUROC generally increases in later phases, indicating that later states better separate successful and unsuccessful trajectories\. This is consistent with prior findings that correct and incorrect trajectories diverge over generation\([Sun et al\., 2026](https://arxiv.org/html/2609.36461#bib.bib5)\)\. Because the final phase gives the strongest and most stable separation, Table[2](https://arxiv.org/html/2609.36461#S4.T2)uses the final\-phase scorer for trajectory selection\. We also find that raw pairwise LR is weaker, while PCA\-compressed LR approaches the MLP, suggesting that linear directions are effective after reducing high\-dimensional nuisance variation\. We therefore use the PCA\-based linear scorer as the main PAIR estimator\.

\(a\)WQ\-AUC across phases\.\(b\)Absolute\-step versus phase alignment\.
Table[2](https://arxiv.org/html/2609.36461#S4.T2)evaluates trajectory scoring for within\-question ranking and Best\-of\-NNselection\. In BoN@32, the model samples 32 reasoning trajectories for each question, and the scorer selects the highest\-scoring trajectory as the final output\. We compare PAIR with four pointwise correctness baselines trained on greedy trajectories with final\-answer correctness labels\. Pre\-Ans token uses the representation immediately before the answer as a late\-stage reference\. Pooled correctness uses the mean four\-token representation near25%25\\%relative progress, following prior work on early correctness signals\([Zhang et al\., 2025](https://arxiv.org/html/2609.36461#bib.bib1)\)\. Absolute\-step uses the third\-step representation, reflecting step\-index\-based comparison\([Sun et al\., 2026](https://arxiv.org/html/2609.36461#bib.bib5)\)\. Normalized\-step uses the step representation closest to80%80\\%relative progress\. PAIR achieves the best or competitive WQ\-AUC and BoN@32 across most settings, showing that phase\-aligned within\-question training provides a stronger trajectory\-ranking signal than baselines\. Pre\-Ans token is also strong in several cases, but it is taken very close to the final answer, where answer information is often already recoverable\. The relatively strong Normalized\-step baseline further suggests that relative trajectory position matters, even without full phase\-aligned training\.

Figure[4\(b\)](https://arxiv.org/html/2609.36461#S4.F4.sf2)visualizes the effect of phase alignment\. When hidden states are grouped by absolute step indices on greedy paths, only the first step forms a clear cluster, and later steps are mixed, with an average NMI of0\.360\.36\. This reflects that the same absolute step can correspond to different reasoning roles across questions\. After aligning trajectories within each fixed question by relative phase, phase groups become more separated, increasing the average NMI to0\.530\.53\.

### 4\.4Phase\-Specific Steering Redirects Reasoning Trajectories

Table 3:Phase\-wise steering results across benchmarks\. Baseline denotes generation without intervention\. SinglePbP\_\{b\}applies the phase\-bbdirection only at phasePbP\_\{b\}, while PAIR applies phase\-specific directions progressively along the trajectory\. We report final\-answer accuracy, wrong\-to\-correct and correct\-to\-wrong flip rates, and generated\-token length after steering\.InterventionGSM8KMATHBBHAcc\.W→CW\{\\rightarrow\}CC→WC\{\\rightarrow\}WTok\.Acc\.W→CW\{\\rightarrow\}CC→WC\{\\rightarrow\}WTok\.Acc\.W→CW\{\\rightarrow\}CC→WC\{\\rightarrow\}WTok\.Baseline79\.3––234±\\pm7042\.0––383±\\pm12974\.6––146±\\pm28SingleP1P\_\{1\}83\.332\.23\.3242±\\pm7147\.312\.64\.7399±\\pm18176\.057\.817\.9167±\\pm25SingleP2P\_\{2\}82\.625\.82\.5233±\\pm6744\.76\.93\.2386±\\pm12170\.65\.27\.1187±\\pm34SingleP3P\_\{3\}80\.03\.20234±\\pm7042\.02\.33\.2387±\\pm17069\.336\.819\.6205±\\pm66SingleP4P\_\{4\}79\.30\.00\.0234±\\pm7042\.71\.10386±\\pm12376\.010\.51\.8198±\\pm54SingleP5P\_\{5\}79\.30\.00\.0234±\\pm7042\.00\.00\.0383±\\pm12972\.00\.03\.6169±\\pm26PAIR \(Ours\)84\.638\.73\.3244±\\pm7447\.312\.64\.7398±\\pm18176\.057\.817\.9167±\\pm25

We use phase\-wise activation steering as a causal validation of the learned path\-quality directions\. Given a generated trajectory, we intervene at the corresponding reasoning phase by adding the learned direction to the hidden state, following Eq\.[12](https://arxiv.org/html/2609.36461#S3.E12)\. Table[3](https://arxiv.org/html/2609.36461#S4.T3)reports steering results on LLaMA\-3\.1\-8B\-Instruct across benchmarks, comparing no intervention, single\-phase steering, and progressive PAIR steering\. SinglePbP\_\{b\}applies only the direction learned for phasePbP\_\{b\}, while PAIR applies phase\-specific directions as the trajectory advances\. We report final\-answer accuracy, wrong\-to\-correct flipsW→CW\{\\to\}C, correct\-to\-wrong flipsC→WC\{\\to\}W, and generated\-token length after steering\.

The steering results show that PAIR can correct some originally unsuccessful trajectories\. On GSM8K, accuracy increases from79\.379\.3to84\.684\.6, with38\.7%38\.7\\%of originally wrong trajectories flipped to correct and only3\.3%3\.3\\%of originally correct trajectories flipped to wrong\. MATH shows a similar accuracy gain, from42\.042\.0to47\.347\.3, while BBH shows a smaller net gain because the largeW→CW\{\\to\}Crate is partly offset by a higherC→WC\{\\to\}Wrate\. The single\-phase results also suggest that early phases are more causally sensitive: steering atP1P\_\{1\}orP2P\_\{2\}is often more effective than steering at later phases\. This is consistent with Fig\.[4\(a\)](https://arxiv.org/html/2609.36461#S4.F4.sf1)because readout separability and causal influence measure different aspects of the trajectory\. Later phases are easier to classify because they are closer to the final outcome, whereas earlier interventions can affect a longer portion of the remaining reasoning process\.

## 5Conclusion

We presented a phase\-aligned view of reasoning trajectories in LLMs\. We showed that standard correctness probes can be ambiguous: across\-question evaluation may capture question\-level variation, and absolute step indices may compare different reasoning stages\. PAIR addresses this by comparing successful and unsuccessful trajectories from the same question at matched phases to learn path\-quality directions\. Phase\-wise steering further shows that these directions can change reasoning outcomes rather than only predict them\. These results suggest analyzing reasoning trajectories as structured processes, controlling comparisons by both question identity and relative phase\.

## References

- Azaria and Mitchell \(2023\)A\. Azaria and T\. MitchellThe internal state of an llm knows when it’s lying\.InFindings of the Association for Computational Linguistics: EMNLP 2023,pp\. 967–976\.Cited by:[§2](https://arxiv.org/html/2609.36461#S2.SS0.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2609.36461#S4.SS1.SSS0.Px3.p1.1)\.
- Bestaet al\.\(2024\)M\. Besta, N\. Blach, A\. Kubicek, R\. Gerstenberger, M\. Podstawski, L\. Gianinazzi, J\. Gajda, T\. Lehmann, H\. Niewiadomski, P\. Nyczyk,et al\.Graph of thoughts: solving elaborate problems with large language models\.InProceedings of the AAAI conference on artificial intelligence,Vol\.38,pp\. 17682–17690\.Cited by:[§2](https://arxiv.org/html/2609.36461#S2.SS0.SSS0.Px1.p1.1)\.
- Boppanaet al\.\(2026\)S\. Boppana, A\. Ma, M\. Loeffler, R\. Sarfati, E\. Bigelow, A\. Geiger, O\. Lewis, and J\. MerulloReasoning theater: disentangling model beliefs from chain\-of\-thought\.arXiv preprint arXiv:2603\.05488\.Cited by:[§4\.2](https://arxiv.org/html/2609.36461#S4.SS2.p4.1)\.
- \[4\]C\. Burns, H\. Ye, D\. Klein, and J\. SteinhardtDiscovering latent knowledge in language models without supervision, 2024\.URL https://arxiv\. org/abs/2212\.03827\.Cited by:[§1](https://arxiv.org/html/2609.36461#S1.p2.1),[§2](https://arxiv.org/html/2609.36461#S2.SS0.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2609.36461#S4.SS1.SSS0.Px3.p1.1)\.
- Chenet al\.\(2021\)M\. Chen, J\. Tworek, H\. Jun, Q\. Yuan, H\. P\. D\. O\. Pinto, J\. Kaplan, H\. Edwards, Y\. Burda, N\. Joseph, G\. Brockman,et al\.Evaluating large language models trained on code\.arXiv preprint arXiv:2107\.03374\.Cited by:[§1](https://arxiv.org/html/2609.36461#S1.p1.1)\.
- Cobbeet al\.\(2021\)K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano, C\. Hesse, and J\. SchulmanTraining verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168\.Cited by:[§A\.3\.1](https://arxiv.org/html/2609.36461#A1.SS3.SSS1.p1.1),[§1](https://arxiv.org/html/2609.36461#S1.p1.1),[§4\.1](https://arxiv.org/html/2609.36461#S4.SS1.SSS0.Px1.p1.1),[Table 1](https://arxiv.org/html/2609.36461#S4.T1.6.1.1.2),[Table 2](https://arxiv.org/html/2609.36461#S4.T2.8.1.1.3)\.
- DeepSeek\-AI \(2025\)DeepSeek\-AIDeepSeek\-r1: incentivizing reasoning capability in llms via reinforcement learning\.External Links:2501\.12948,[Link](https://arxiv.org/abs/2501.12948)Cited by:[§A\.2\.4](https://arxiv.org/html/2609.36461#A1.SS2.SSS4.p1.1),[§4\.1](https://arxiv.org/html/2609.36461#S4.SS1.SSS0.Px1.p1.1),[Table 2](https://arxiv.org/html/2609.36461#S4.T2.8.1.18.1.1.2.1.2.1)\.
- Heet al\.\(2026\)Z\. He, G\. Xiong, B\. Liu, S\. Sinha, and A\. ZhangReasoning beyond chain\-of\-thought: a latent computational mode in large language models\.External Links:2601\.08058,[Link](https://arxiv.org/abs/2601.08058)Cited by:[§1](https://arxiv.org/html/2609.36461#S1.p2.1)\.
- Kojimaet al\.\(2022\)T\. Kojima, S\. S\. Gu, M\. Reid, Y\. Matsuo, and Y\. IwasawaLarge language models are zero\-shot reasoners\.Advances in neural information processing systems35,pp\. 22199–22213\.Cited by:[§2](https://arxiv.org/html/2609.36461#S2.SS0.SSS0.Px1.p1.1)\.
- Liet al\.\(2023\)K\. Li, O\. Patel, F\. Viégas, H\. Pfister, and M\. WattenbergInference\-time intervention: eliciting truthful answers from a language model\.Advances in Neural Information Processing Systems36,pp\. 41451–41530\.Cited by:[§1](https://arxiv.org/html/2609.36461#S1.p2.1),[§2](https://arxiv.org/html/2609.36461#S2.SS0.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2609.36461#S4.SS1.SSS0.Px3.p1.1)\.
- Lightmanet al\.\(2023\)H\. Lightman, V\. Kosaraju, Y\. Burda, H\. Edwards, B\. Baker, T\. Lee, J\. Leike, J\. Schulman, I\. Sutskever, and K\. CobbeLet’s verify step by step\.arXiv preprint arXiv:2305\.20050\.Cited by:[§A\.3\.2](https://arxiv.org/html/2609.36461#A1.SS3.SSS2.p1.1),[§2](https://arxiv.org/html/2609.36461#S2.SS0.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2609.36461#S4.SS1.SSS0.Px1.p1.1),[Table 1](https://arxiv.org/html/2609.36461#S4.T1.6.1.1.3),[Table 2](https://arxiv.org/html/2609.36461#S4.T2.8.1.1.4)\.
- Liuet al\.\(2025\)J\. Liu, J\. Jain, M\. Diab, and N\. SubramaniLLM microscope: what model internals reveal about answer correctness and context utilization\.arXiv preprint arXiv:2510\.04013\.Cited by:[§2](https://arxiv.org/html/2609.36461#S2.SS0.SSS0.Px2.p1.1)\.
- Meta \(2024\)MetaLlama 3\.1: open foundation and instruction\-tuned large language models\.Note:Llama\-3\.1\-8B\-Instruct modelExternal Links:[Link](https://ai.meta.com/llama/)Cited by:[§A\.2\.1](https://arxiv.org/html/2609.36461#A1.SS2.SSS1.p1.1),[§4\.1](https://arxiv.org/html/2609.36461#S4.SS1.SSS0.Px1.p1.1),[Table 2](https://arxiv.org/html/2609.36461#S4.T2.8.1.3.1.1.2.1.2.1)\.
- Panicksseryet al\.\(2023\)N\. Panickssery, N\. Gabrieli, J\. Schulz, M\. Tong, E\. Hubinger, and A\. M\. TurnerSteering llama 2 via contrastive activation addition\.arXiv preprint arXiv:2312\.06681\.Cited by:[§2](https://arxiv.org/html/2609.36461#S2.SS0.SSS0.Px3.p1.1)\.
- Parket al\.\(2023\)K\. Park, Y\. J\. Choe, and V\. VeitchThe linear representation hypothesis and the geometry of large language models\.arXiv preprint arXiv:2311\.03658\.Cited by:[§2](https://arxiv.org/html/2609.36461#S2.SS0.SSS0.Px2.p1.1)\.
- Sunet al\.\(2026\)L\. Sun, H\. Dong, B\. Qiao, Q\. Lin, D\. Zhang, and S\. RajmohanLLM reasoning as trajectories: step\-specific representation geometry and correctness signals\.arXiv preprint arXiv:2604\.05655\.Cited by:[§1](https://arxiv.org/html/2609.36461#S1.p4.1),[§2](https://arxiv.org/html/2609.36461#S2.SS0.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2609.36461#S3.SS2.SSS0.Px1.p1.1),[§4\.3](https://arxiv.org/html/2609.36461#S4.SS3.p1.1),[§4\.3](https://arxiv.org/html/2609.36461#S4.SS3.p2.1)\.
- Suzgunet al\.\(2022\)M\. Suzgun, N\. Scales, N\. Schärli, S\. Gehrmann, Y\. Tay, H\. W\. Chung, A\. Chowdhery, Q\. V\. Le, E\. H\. Chi, D\. Zhou, and J\. WeiChallenging big\-bench tasks and whether chain\-of\-thought can solve them\.arXiv preprint arXiv:2210\.09261\.Cited by:[§A\.3\.3](https://arxiv.org/html/2609.36461#A1.SS3.SSS3.p1.1),[§4\.1](https://arxiv.org/html/2609.36461#S4.SS1.SSS0.Px1.p1.1),[Table 1](https://arxiv.org/html/2609.36461#S4.T1.6.1.1.4),[Table 2](https://arxiv.org/html/2609.36461#S4.T2.8.1.1.5)\.
- Teamet al\.\(2025\)G\. Team, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Rivière,et al\.Gemma 3 technical report\.arXiv preprint arXiv:2503\.19786\.Cited by:[§A\.2\.3](https://arxiv.org/html/2609.36461#A1.SS2.SSS3.p1.1),[§4\.1](https://arxiv.org/html/2609.36461#S4.SS1.SSS0.Px1.p1.1),[Table 2](https://arxiv.org/html/2609.36461#S4.T2.8.1.13.1.1)\.
- Team \(2025\)Q\. TeamQwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§A\.2\.2](https://arxiv.org/html/2609.36461#A1.SS2.SSS2.p1.1),[§4\.1](https://arxiv.org/html/2609.36461#S4.SS1.SSS0.Px1.p1.1),[Table 2](https://arxiv.org/html/2609.36461#S4.T2.8.1.8.1.1)\.
- Turneret al\.\(2023\)A\. M\. Turner, L\. Thiergart, G\. Leech, D\. Udell, J\. J\. Vazquez, U\. Mini, and M\. MacDiarmidActivation addition: steering language models without optimization\. arxiv eprints, pages arxiv–2308\.Cited by:[§2](https://arxiv.org/html/2609.36461#S2.SS0.SSS0.Px3.p1.1)\.
- Wanget al\.\(2022\)X\. Wang, J\. Wei, D\. Schuurmans, Q\. Le, E\. Chi, S\. Narang, A\. Chowdhery, and D\. ZhouSelf\-consistency improves chain of thought reasoning in language models\.arXiv preprint arXiv:2203\.11171\.Cited by:[§2](https://arxiv.org/html/2609.36461#S2.SS0.SSS0.Px1.p1.1)\.
- Weiet al\.\(2022\)J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, F\. Xia, E\. Chi, Q\. V\. Le, D\. Zhou,et al\.Chain\-of\-thought prompting elicits reasoning in large language models\.Advances in neural information processing systems35,pp\. 24824–24837\.Cited by:[§1](https://arxiv.org/html/2609.36461#S1.p1.1),[§2](https://arxiv.org/html/2609.36461#S2.SS0.SSS0.Px1.p1.1)\.
- Yaoet al\.\(2023\)S\. Yao, D\. Yu, J\. Zhao, I\. Shafran, T\. Griffiths, Y\. Cao, and K\. NarasimhanTree of thoughts: deliberate problem solving with large language models\.Advances in neural information processing systems36,pp\. 11809–11822\.Cited by:[§2](https://arxiv.org/html/2609.36461#S2.SS0.SSS0.Px1.p1.1)\.
- Zhanget al\.\(2025\)A\. Zhang, Y\. Chen, J\. Pan, C\. Zhao, A\. Panda, J\. Li, and H\. HeReasoning models know when they’re right: probing hidden states for self\-verification\.External Links:2504\.05419,[Link](https://arxiv.org/abs/2504.05419)Cited by:[§2](https://arxiv.org/html/2609.36461#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.36461#S2.SS0.SSS0.Px3.p1.1),[§4\.2](https://arxiv.org/html/2609.36461#S4.SS2.p4.1),[§4\.3](https://arxiv.org/html/2609.36461#S4.SS3.p2.1)\.

## Appendix AAppendix

### A\.1Prompt Design

The prompt for each model are listed as below:

Figure 5:Prompt for LLaMA3\.1\-8B\-InstructFigure 6:Prompt for Gemma3\-4BFigure 7:Prompt for Qwen3\-4BFigure 8:Prompt for DeepSeek\-R1\-Distilled\-LLaMA\-8B
### A\.2Model Details

We evaluate four open\-weight LLMs with different instruction\-tuning and reasoning\-training backgrounds\. All experiments use the publicly released checkpoints without additional finetuning\. Unless otherwise specified, hidden\-state features are extracted from the final transformer layer, and steering interventions are applied to the same layer\.

#### A\.2\.1Llama\-3\.1\-8B\-Instruct

Llama\-3\.1\-8B\-Instruct is an instruction\-tuned model from the Llama 3\.1 family\([Meta, 2024](https://arxiv.org/html/2609.36461#bib.bib21)\)\. We use it as the primary model for both diagnostic analysis and phase\-wise steering because it provides stable multi\-step reasoning trajectories under chain\-of\-thought prompting\. For all experiments, we use the same decoding and trajectory\-sampling settings described in Sec\.[4\.1](https://arxiv.org/html/2609.36461#S4.SS1)\.

#### A\.2\.2Qwen3\-4B

Qwen3\-4B is an open\-weight model from the Qwen3 family\([Team, 2025](https://arxiv.org/html/2609.36461#bib.bib23)\)\. We include it to test whether the observed trajectory\-level patterns hold beyond the Llama model family\. In our experiments, Qwen3\-4B is evaluated with the same prompting format, trajectory\-sampling configuration, and hidden\-state extraction procedure as the other models\.

#### A\.2\.3Gemma3\-4B

Gemma3\-4B is an open\-weight model from the Gemma 3 family\([Team et al\., 2025](https://arxiv.org/html/2609.36461#bib.bib22)\)\. We use this model as an additional architecture family for evaluating pointwise correctness analysis, phase\-aligned trajectory ranking, and Best\-of\-NNtrajectory selection\. No model\-specific finetuning or calibration is applied\.

#### A\.2\.4DeepSeek\-R1\-Distill\-Llama\-8B

DeepSeek\-R1\-Distill\-Llama\-8B is a distilled reasoning model from the DeepSeek\-R1 series\([DeepSeek\-AI, 2025](https://arxiv.org/html/2609.36461#bib.bib24)\)\. We include it to evaluate PAIR on a model with stronger explicit reasoning behavior\. The model is evaluated with the same trajectory\-sampling and representation\-extraction pipeline used for the other models\. Because distilled reasoning models tend to generate longer reasoning traces, we report generated\-token lengths when evaluating steering interventions\.

### A\.3Dataset Details

We evaluate on three reasoning benchmarks: GSM8K, MATH\-500, and BBH\. All splits are made at the question level\. For every benchmark, trajectories sampled from the same question are assigned to the same split, so no question appears in both training and test sets\.

#### A\.3\.1GSM8K

GSM8K is a grade\-school math word\-problem benchmark\([Cobbe et al\., 2021](https://arxiv.org/html/2609.36461#bib.bib19)\)\. We use GSM8K as the main benchmark for learning phase\-specific path\-quality directions because it provides a sufficiently large training split\. Specifically, we train on 1,000 questions sampled from the official training split and evaluate on 500 questions sampled from the official test split\. For each question, we sample 32 reasoning trajectories with temperature1\.01\.0and top\-p=0\.95p=0\.95\. A trajectory is labeled correct if the extracted final answer matches the ground\-truth answer; otherwise, it is labeled incorrect\.

#### A\.3\.2MATH\-500

MATH\-500 is a 500\-problem subset of the MATH benchmark\([Lightman et al\., 2023](https://arxiv.org/html/2609.36461#bib.bib20)\)\. It contains competition\-style mathematical reasoning problems and is substantially more difficult than GSM8K\. Since MATH\-500 does not provide a large official training split for our setting, we use a fixed 70/30 question\-level split\. The split is sampled once and reused across all models and methods to ensure a consistent comparison\. For each question, we sample 32 reasoning trajectories using the same decoding configuration as GSM8K, and correctness is determined by dataset\-specific final\-answer extraction\.

#### A\.3\.3BBH Logical Deduction

We use thelogical\_deduction\_three\_objectstask from Big\-Bench Hard\([Suzgun et al\., 2022](https://arxiv.org/html/2609.36461#bib.bib18)\)\. This task evaluates symbolic and logical reasoning over a small set of objects and constraints\. As with MATH\-500, we use a fixed 70/30 question\-level split and reuse the same split across all methods and models\. For each question, we sample 32 reasoning trajectories with temperature1\.01\.0and top\-p=0\.95p=0\.95\. A trajectory is labeled correct if its extracted final answer matches the ground\-truth option or answer string under the task\-specific extraction rule\.

### A\.4Implementation Details

#### A\.4\.1Step Marker Extraction

We identify intermediate reasoning steps from the model’s generated response using a rule\-based*smart marker*procedure\. The input is first normalized with the same text cleaning used elsewhere in our pipeline, and step markers are then detected on the cleaned response text\.

##### Candidate step markers\.

The smart marker extractor considers four classes of candidate boundaries:

1. 1\.Explicit numeric markers at line start, such asStep 3:,3\.,3\), or\(3\)\.
2. 2\.Section\-style numeric headersat line start, such asPart 2:,Phase 3,Round 4, etc\.
3. 3\.Discourse markers, such asFirst,Second,Next,Then,After that,Finally, andLastly\. These are only accepted when they occur at a line boundary or at a sentence boundary\.
4. 4\.Heading\-like linesending with a colon, provided that they look like short natural\-language headers rather than equations or formatting artifacts\.

##### Filtering heuristics\.

To avoid spurious boundaries, we apply several filters\. Numeric or discourse candidates are kept only if the following text looks like genuine prose rather than a formula, a bare number, or a final\-answer statement\. We explicitly reject separator/table lines \(e\.g\., rows of dashes\), equation\-only lines, and lines that contain only symbols or numbers\. We also strip common Markdown decoration before testing whether a line looks like a valid heading or step\.

##### Preference for explicit structure\.

If the response already contains a clear explicit step structure \(at least two numbered orStep\-nnmarkers\), we trust that structure and ignore weaker discourse\-style transitions inside those numbered steps\. Otherwise, we merge numeric, discourse, and heading candidates into a single ordered candidate list\.

##### Step numbering\.

Candidates are processed from left to right and assigned monotonically increasing step indices\. Explicit numeric markers keep their stated index\. Discourse ordinals such asFirst,Second, andThirdare mapped to their canonical indices\. Other accepted discourse or heading markers are assigned the next available step index\. We discard any candidate whose assigned index would be non\-monotonic or implausibly large\.

##### Mapping markers to hidden states\.

After converting accepted character positions to token indices, we represent stepkkusing the hidden state immediately before the marker of stepk\+1k\{\+\}1\. Thus, a marker for stepk\+1k\{\+\}1defines the feature for the preceding stepkk\. In addition, we extract:

- •afinal\-stepfeature from the hidden state immediately before the last conclusion marker \(e\.g\., phrases such astherefore,so the answer is, or equivalent conclusion patterns\), and
- •ananswer\-prefeature from the hidden state immediately before the predicted answer token\.

In our implementation, if an answer token can be located, we prefer the last conclusion marker that appears before that answer; otherwise we use the last detected conclusion marker in the response\.

#### A\.4\.2Implementation details for learning path\-quality directions\.

For each phase, we train an independent pairwise logistic ranking model\. The input feature is the mean hidden state over the four tokens immediately preceding the extracted boundary, corresponding to thewindow4\_meanfeature\. The preprocessing mapϕb\\phi\_\{b\}is fitted separately for each phase using only training questions\. It consists of standardization followed by PCA:

ϕb​\(h\)=Pb​\(h−μbσb\),\\phi\_\{b\}\(h\)=P\_\{b\}\\left\(\\frac\{h\-\\mu\_\{b\}\}\{\\sigma\_\{b\}\}\\right\),\(15\)whereμb\\mu\_\{b\}andσb\\sigma\_\{b\}are the training\-set mean and standard deviation for phasebb, andPbP\_\{b\}is the PCA projection matrix\. We use at most 128 PCA components, with the actual number of components set to

db=min⁡\(128,Nb−1,D\),d\_\{b\}=\\min\(128,N\_\{b\}\-1,D\),\(16\)whereNbN\_\{b\}is the number of training states for phasebbandDDis the original hidden dimension\. If PCA is not applicable,ϕb\\phi\_\{b\}reduces to standardization\.

For each question and phase, we form all positive–negative pairs fromℋq,b\+×ℋq,b−\\mathcal\{H\}\_\{q,b\}^\{\+\}\\times\\mathcal\{H\}\_\{q,b\}^\{\-\}\. To control the number of training examples, we keep at most 128 positive–negative pairs per question, sampled uniformly when more are available\. In the logistic\-regression implementation, we add both the forward difference and the reversed difference:

x=ϕb​\(h\+\)−ϕb​\(h−\),y=1,x′=ϕb​\(h−\)−ϕb​\(h\+\),y′=0\.x=\\phi\_\{b\}\(h^\{\+\}\)\-\\phi\_\{b\}\(h^\{\-\}\),\\quad y=1,\\qquad x^\{\\prime\}=\\phi\_\{b\}\(h^\{\-\}\)\-\\phi\_\{b\}\(h^\{\+\}\),\\quad y^\{\\prime\}=0\.\(17\)This is equivalent to optimizing the ranking objective in Eq\.[11](https://arxiv.org/html/2609.36461#S3.E11)\.

We usesklearn\.linear\_model\.LogisticRegressionwithsolver=lbfgs,class\_weight=balanced,C=1\.0C=1\.0, andmax\_iter=2000\. All splits are made at the question level: training, validation, and test questions are disjoint, and pairs are constructed only within the corresponding split\.

For steering, the learned coefficientw~b\\tilde\{w\}\_\{b\}is mapped back to the raw hidden\-state space\. When PCA is used as in Eq\.[15](https://arxiv.org/html/2609.36461#A1.E15), the raw\-space direction is

wb=diag⁡\(σb\)−1​Pb⊤​w~b,w¯b=wb‖wb‖2\.w\_\{b\}=\\operatorname\{diag\}\(\\sigma\_\{b\}\)^\{\-1\}P\_\{b\}^\{\\top\}\\tilde\{w\}\_\{b\},\\qquad\\bar\{w\}\_\{b\}=\\frac\{w\_\{b\}\}\{\\\|w\_\{b\}\\\|\_\{2\}\}\.\(18\)The normalized directionw¯b\\bar\{w\}\_\{b\}is used for phase\-wise steering\.

#### A\.4\.3Details for PAIR Direction Learning and Steering

##### Feature Preprocessing and Linear Scorer

For each phasebb, we train an independent linear path\-quality scorer using only training questions\. The raw featurehhis the mean hidden state over the four tokens immediately preceding the extracted reasoning boundary:

h=14​∑r=14Ht−r\(ℓ\),h=\\frac\{1\}\{4\}\\sum\_\{r=1\}^\{4\}H\_\{t\-r\}^\{\(\\ell\)\},\(19\)whereHt−r\(ℓ\)H\_\{t\-r\}^\{\(\\ell\)\}is the token\-level hidden state at layerℓ\\ell\.

Each phase has its own preprocessing map\. We first standardize raw features using the training\-set mean and standard deviation:

h¯=h−μbσb\.\\bar\{h\}=\\frac\{h\-\\mu\_\{b\}\}\{\\sigma\_\{b\}\}\.\(20\)We then apply PCA when the number of available training states is sufficient:

ϕb​\(h\)=Pb​h¯,\\phi\_\{b\}\(h\)=P\_\{b\}\\bar\{h\},\(21\)wherePbP\_\{b\}is the phase\-specific PCA projection matrix\. The PCA dimension is set to

db=min⁡\(128,Nb−1,D\),d\_\{b\}=\\min\(128,N\_\{b\}\-1,D\),\(22\)whereNbN\_\{b\}is the number of training states at phasebbandDDis the raw hidden dimension\. If PCA is not applicable,ϕb\\phi\_\{b\}reduces to standardization\.

The phase\-bbscorer is

s~b​\(h\)=w~b⊤​ϕb​\(h\)\+βb,\\widetilde\{s\}\_\{b\}\(h\)=\\tilde\{w\}\_\{b\}^\{\\top\}\\phi\_\{b\}\(h\)\+\\beta\_\{b\},\(23\)wherew~b\\tilde\{w\}\_\{b\}andβb\\beta\_\{b\}are learned by pairwise logistic regression\.

For each question and phase, we form all successful–unsuccessful pairs\(h\+,h−\)\(h^\{\+\},h^\{\-\}\)from the same question\. To implement the pairwise objective, we add both the forward and reversed differences:

x=ϕb​\(h\+\)−ϕb​\(h−\),y=1,x′=ϕb​\(h−\)−ϕb​\(h\+\),y′=0\.x=\\phi\_\{b\}\(h^\{\+\}\)\-\\phi\_\{b\}\(h^\{\-\}\),\\quad y=1,\\qquad x^\{\\prime\}=\\phi\_\{b\}\(h^\{\-\}\)\-\\phi\_\{b\}\(h^\{\+\}\),\\quad y^\{\\prime\}=0\.\(24\)We keep at most 128 positive–negative pairs per question and phase, sampled uniformly when more pairs are available\.

The logistic\-regression model useslbfgs,class\_weight=balanced,C=1\.0C=1\.0,max\_iter=2000, and random seed 42\. All train, validation, and test splits are made at the question level\.

##### Mapping Directions Back to Raw Hidden Space

The learned coefficientw~b\\tilde\{w\}\_\{b\}lies in the preprocessed feature space\. For steering, we map it back to the original hidden\-state space\. When preprocessing consists of standardization followed by PCA, the raw\-space direction is

vb=diag⁡\(σb\)−1​Pb⊤​w~b\.v\_\{b\}=\\operatorname\{diag\}\(\\sigma\_\{b\}\)^\{\-1\}P\_\{b\}^\{\\top\}\\tilde\{w\}\_\{b\}\.\(25\)We then normalize it:

wb=vb‖vb‖2\.w\_\{b\}=\\frac\{v\_\{b\}\}\{\\\|v\_\{b\}\\\|\_\{2\}\}\.\(26\)This normalized raw\-space directionwbw\_\{b\}is used for activation steering\.

##### Thresholded and Clipped Phase\-wise Steering

For each phasebb, we estimate a reference score from successful training trajectories\. In our implementation, we use the lower quartile of successful\-state scores:

τb=Quantile0\.25\(\{s~b\(h\):h∈ℋq,b\+,q∈𝒟train\}\)\.\\tau\_\{b\}=\\operatorname\{Quantile\}\_\{0\.25\}\\left\(\\left\\\{\\widetilde\{s\}\_\{b\}\(h\):h\\in\\mathcal\{H\}\_\{q,b\}^\{\+\},\\;q\\in\\mathcal\{D\}\_\{\\mathrm\{train\}\}\\right\\\}\\right\)\.\(27\)This threshold defines the lower reference range of successful trajectories at phasebb\.

During generation, letHt\(ℓ\)H\_\{t\}^\{\(\\ell\)\}be the hidden state at layerℓ\\elland decoding positiontt\. Letbtb\_\{t\}be the phase assigned to the current position by the intervention schedule\. The adaptive steering strength is

λt=min⁡\{λmax,α​\[τbt−s~bt​\(Ht\(ℓ\)\)\]\+\},\\lambda\_\{t\}=\\min\\left\\\{\\lambda\_\{\\max\},\\;\\alpha\\left\[\\tau\_\{b\_\{t\}\}\-\\widetilde\{s\}\_\{b\_\{t\}\}\\left\(H\_\{t\}^\{\(\\ell\)\}\\right\)\\right\]\_\{\+\}\\right\\\},\(28\)where\[x\]\+=max⁡\(x,0\)\[x\]\_\{\+\}=\\max\(x,0\),α\\alphais the global steering scale, andλmax\\lambda\_\{\\max\}is the maximum allowed intervention magnitude\.

The hidden\-state update is

H~t\(ℓ\)=Ht\(ℓ\)\+λt​wbt\.\\widetilde\{H\}\_\{t\}^\{\(\\ell\)\}=H\_\{t\}^\{\(\\ell\)\}\+\\lambda\_\{t\}w\_\{b\_\{t\}\}\.\(29\)The modified hidden state is passed to the remaining layers, while all model parameters remain fixed\.

##### Intervention Schedules

We evaluate two schedules\.

##### Single\-phase steering\.

For a selected phasebb, the intervention is applied only when the current position belongs to that phase:

bt=b⇒H~t\(ℓ\)=Ht\(ℓ\)\+λt​wb\.b\_\{t\}=b\\quad\\Rightarrow\\quad\\widetilde\{H\}\_\{t\}^\{\(\\ell\)\}=H\_\{t\}^\{\(\\ell\)\}\+\\lambda\_\{t\}w\_\{b\}\.\(30\)No intervention is applied outside the selected phase\.

##### Progressive steering\.

In progressive steering, the intervention follows the trajectory phase:

H~t\(ℓ\)=Ht\(ℓ\)\+λt​wbt,bt∈\{1,…,K\}\.\\widetilde\{H\}\_\{t\}^\{\(\\ell\)\}=H\_\{t\}^\{\(\\ell\)\}\+\\lambda\_\{t\}w\_\{b\_\{t\}\},\\qquad b\_\{t\}\\in\\\{1,\\ldots,K\\\}\.\(31\)Thus, different parts of the generation are steered by the direction learned for their corresponding phase\.

##### Steering Hyperparameters

For phase\-wise steering, we use the adaptive intervention strength defined in Eq\.[13](https://arxiv.org/html/2609.36461#S3.E13)\. The two main hyperparameters are the global scaleα\\alpha, which controls the size of the score\-dependent update, and the clipping thresholdλmax\\lambda\_\{\\max\}, which limits the maximum intervention magnitude\. We select these values by sweeping a small grid on the validation split and choosing the setting that gives the best validation accuracy while avoiding large correct\-to\-wrong flip rates\. After selection, the same values are fixed and used for all test\-set steering evaluations\. Table[4](https://arxiv.org/html/2609.36461#A1.T4)reports the hyperparameters used for the steering experiments in Table[3](https://arxiv.org/html/2609.36461#S4.T3)\.

Table 4:Steering hyperparameters selected by validation\-set sweep\.α\\alphacontrols the scale of the adaptive update, andλmax\\lambda\_\{\\max\}clips the maximum intervention magnitude\.Benchmarkα\\alphaλmax\\lambda\_\{\\max\}GSM8K55MATH\-50055BBH2\.52\.5
##### Pointwise Probe Training Details

In addition to the final pairwise probe, we train four*single\-path*probes\. Each of these probes is trained as a binary classifier ongreedytrajectories, where the label indicates whether the final answer of the trajectory is correct\. The learned probe is then evaluated onsampledtrajectories from the same dataset using within\-question ranking metrics and best\-of\-NNselection\. The four probes differ only in the hidden\-state feature used as input\.

##### Data sources\.

For each question, we use two kinds of trajectories:

- •a singlegreedytrajectory for probe training, and
- •a set ofsampledtrajectories \(typicallyN=32N=32\) for evaluation\.

Greedy trajectories provide one labeled example per question for training\. Sampled trajectories provide multiple correct and incorrect candidate paths for the same question, allowing us to evaluate whether a probe score can rank good reasoning paths above bad ones\.

##### Outer train/test split\.

All splits are defined at thequestion level\. Let𝒬\\mathcal\{Q\}denote the set of question IDs\. We construct a fixed outer train/test split using only greedy\-path correctness labels, with a default ratio of70%/30%70\\%/30\\%\. Thus, all greedy and sampled paths from the same question belong to the same outer split, preventing leakage across train and test\.

##### Inner validation for hyperparameter selection\.

Within the outer training questions, we create an inner validation split used only to select the logistic regression regularization strength\. When a dedicated validation split is not available, we create one by splitting the outer training questions again, typically using an inner validation ratio of0\.20\.2\. The split is also performed at the question level\. If stratified splitting is not possible due to small class counts, we fall back to an unstratified random split\.

##### Feature representation\.

Each of the four probes takes as input a single hidden\-state vector extracted from one location in the trajectory\. In all cases, the hidden state is represented as the mean of thefour tokens immediately precedingthe target position\. Letht−3,ht−2,ht−1,ht∈ℝdh\_\{t\-3\},h\_\{t\-2\},h\_\{t\-1\},h\_\{t\}\\in\\mathbb\{R\}^\{d\}denote the hidden states of the four tokens ending at the target token positiontt\. The feature vector is

x=14​∑i=03ht−i∈ℝd\.x=\\frac\{1\}\{4\}\\sum\_\{i=0\}^\{3\}h\_\{t\-i\}\\in\\mathbb\{R\}^\{d\}\.
The four probes are defined as follows:

1. 1\.Answer\-pre probe\.We locate the predicted answer token and use the mean hidden state of the four tokens immediately preceding that answer position\.
2. 2\.Relative\-position\-25% probe\.LetLLbe the token length of the cleaned response\. We compute the target token index as t0\.25=round⁡\(0\.25⋅\(L−1\)\),t\_\{0\.25\}=\\mathrm\{round\}\\\!\\left\(0\.25\\cdot\(L\-1\)\\right\),and use the mean hidden state of the four\-token window ending att0\.25t\_\{0\.25\}\.
3. 3\.Step\-3 probe\.We run the step marker extraction procedure described in Appendix[A\.4\.1](https://arxiv.org/html/2609.36461#A1.SS4.SSS1)\. By our convention, the hidden state of stepkkis taken from the hidden state immediately before the marker of stepk\+1k\+1\. The step\-3 probe therefore uses the four\-token mean immediately before the marker of step 4\. If a trajectory does not contain such a step boundary, it is skipped for this probe\.
4. 4\.80%\-step probe\.We first extract the ordered list of reasoning\-step features for the trajectory, including the final\-step feature defined by the hidden state before the conclusion marker\. If there aremmsuch reasoning\-step features, we select the step at index j0\.8=round⁡\(0\.8⋅\(m−1\)\),j\_\{0\.8\}=\\mathrm\{round\}\\\!\\left\(0\.8\\cdot\(m\-1\)\\right\),and use that step’s four\-token mean hidden state as the probe input\.

##### Training labels\.

For greedy training, each question contributes one trajectory with a binary label

wherey=1y=1indicates that the final predicted answer is correct andy=0y=0indicates that it is incorrect\.

##### Preprocessing\.

For each probe, we collect all training feature vectors from the greedy training split\. Let the raw feature matrix be

X∈ℝn×d\.X\\in\\mathbb\{R\}^\{n\\times d\}\.We first standardize features dimension\-wise using aStandardScaler:

X~i​j=Xi​j−μjσj\.\\tilde\{X\}\_\{ij\}=\\frac\{X\_\{ij\}\-\\mu\_\{j\}\}\{\\sigma\_\{j\}\}\.We then fit PCA on the standardized training features and project them to a low\-dimensional space\. If the requested PCA dimension iskk, the actual number of retained components is

k′=min⁡\(k,n−1,d\)\.k^\{\\prime\}=\\min\(k,n\-1,d\)\.In our experiments the default target dimension isk=128k=128, so typically the probe input after PCA is

z∈ℝ128\.z\\in\\mathbb\{R\}^\{128\}\.

##### Classifier\.

Each probe is a logistic regression classifier trained on the PCA\-transformed features\. For a transformed feature vectorzz, the probe predicts

s⁡\(z\)=w⊤​z\+b,s\(z\)=w^\{\\top\}z\+b,and the probability of correctness is

p⁡\(y=1∣z\)=σ⁡\(w⊤​z\+b\),p\(y=1\\mid z\)=\\sigma\(w^\{\\top\}z\+b\),whereσ⁡\(⋅\)\\sigma\(\\cdot\)is the logistic sigmoid\.

We usesklearn\.linear\_model\.LogisticRegressionwith:

- •solver:lbfgs,
- •class weighting:balanced,
- •maximum iterations:5000,
- •random seed fixed for reproducibility\.

##### Hyperparameter selection\.

The main tuned hyperparameter is the inverse regularization strengthCCof logistic regression\. We evaluate a candidate set such as

C∈\{0\.03,0\.1,0\.3,1\.0,3\.0\}\.C\\in\\\{0\.03,0\.1,0\.3,1\.0,3\.0\\\}\.For each candidateCC, we train on the greedy inner\-training questions and evaluate on sampled paths from the inner\-validation questions\. Although the probe is trained on single\-path correctness labels, model selection is based onwithin\-question ranking performanceon sampled validation paths, since ranking is the downstream behavior of interest\.

Concretely, for each candidateCC, we score sampled validation paths and compute:

- •pooled within\-question AUC,
- •path\-level AUROC across sampled paths,
- •and optionally best\-of\-NNaccuracy\.

We choose the candidate with the best validation pooled within\-question AUC; ties are broken by path\-level AUROC and then by preferring smallerCC\(stronger regularization\)\.

##### Refitting\.

After selectingC⋆C^\{\\star\}, we refit the probe on all greedy trajectories from the outer training questions using the same preprocessing pipeline:

1. 1\.fitStandardScaleron outer\-train greedy features,
2. 2\.fit PCA on standardized outer\-train features,
3. 3\.train logistic regression withC⋆C^\{\\star\}on the transformed features\.

The fitted scaler, PCA parameters, and logistic regression coefficients are saved for later analysis and reuse\.

### A\.5Compute Resources

The main computational cost of our experiments comes from trajectory sampling and hidden\-state caching, rather than from training or steering\. For each model and benchmark, we sample 32 trajectories per question and store the hidden states needed for later phase alignment and scoring\. On a single NVIDIA A100 GPU, one full trajectory\-sampling run for a model–benchmark setting takes approximately 20 GPU\-hours, depending on the average generation length of the benchmark and model\. Distilled reasoning models tend to require more time because they generate longer reasoning traces\. The total compute scales approximately linearly with the number of evaluated model–benchmark settings and the number of sampled trajectories per question\.

After trajectories and hidden states are cached, training the PAIR linear scorers is lightweight\. The phase\-specific logistic\-regression models are trained on cached features and typically require only CPU\-level computation or a small amount of CPU/GPU time compared with sampling\. Steering also adds little overhead during generation, since it only applies a vector update to the hidden state at selected positions and does not require updating model parameters\. The main resource bottlenecks are therefore generation time and storage for cached trajectories, token\-level hidden states, and extracted step\-level features\.

### A\.6Supplementary MMLU Candidate\-Set Analysis

This appendix\-only analysis measures answer availability in a separate subject\-balanced subset of 500 official MMLU test questions per model\. It does not train or evaluate a PAIR scorer on MMLU\. For each question, we retain 32 raw independent draws from Llama\-3\.1\-8B\-Instruct or Qwen3\-4B, sampled with temperature 1\.0 and top\-pp0\.95 under a zero\-shot multiple\-choice chain\-of\-thought prompt; Qwen uses its native thinking mode\. The firstNNdraws define each candidate set, with duplicate draws retained\. Unparsed outputs count towardNNbut cast no vote\. Self\-consistency \(SC@NN\) chooses the most frequent parsed answer, breaking ties by the earliest valid draw; Oracle@NNis correct if any of the sameNNcandidates contains the correct answer\. Thus Oracle is a candidate\-set upper bound, not the accuracy of an implemented selector\.

Table 5:Self\-consistency and oracle accuracy on the same fixed MMLU candidate sets\. Each model uses 500 official\-test questions;NNis the number of raw draws retained per question\. Gap is Oracle minus SC in percentage points\.ModelNNSC@NN\(%\)Oracle@NN\(%\)Gap \(pp\)Llama\-3\.1\-8B168\.268\.20\.0Llama\-3\.1\-8B472\.485\.012\.6Llama\-3\.1\-8B875\.090\.615\.6Llama\-3\.1\-8B1676\.495\.419\.0Llama\-3\.1\-8B3275\.097\.422\.4Qwen3\-4B183\.283\.20\.0Qwen3\-4B484\.490\.86\.4Qwen3\-4B884\.291\.87\.6Qwen3\-4B1684\.693\.08\.4Qwen3\-4B3283\.093\.810\.8AtN=32N=32, the Oracle–SC gap is 22\.4 percentage points for Llama and 10\.8 points for Qwen; 95% question\-bootstrap intervals from 2,000 resamples are \[19\.0, 26\.2\] and \[8\.0, 13\.8\] points, respectively\. These gaps quantify selection headroom within this MMLU candidate set; they do not imply that PAIR can close the gap\. The multiple\-choice prompt and dataset differ from the three main PAIR benchmarks, so these accuracies are not directly comparable to Table[2](https://arxiv.org/html/2609.36461#S4.T2)\. The SC curve need not increase monotonically withNN: adding draws can change the majority\-vote answer\.

相似文章

ReasoningFlow: 用于理解LLM推理轨迹的篇章结构

arXiv cs.CL

介绍 ReasoningFlow,一个将大语言模型推理轨迹的篇章结构捕获为有向无环图的框架,从而能够细粒度分析推理行为(如自我反思和回溯)。基于对数千条轨迹的手动和自动标注,揭示了模型之间的结构相似性,并且大多数错误步骤并不贡献于最终答案。

监控内部独白:探针轨迹揭示推理动态

Hugging Face Daily Papers

本文介绍了一种通过分析探针轨迹(即概念概率在生成token上的演变)来监控大型推理模型推理过程的方法。该方法利用隐藏表示中的时间特征和信号处理特征,更好地预测未来模型行为,通过最大池化达到了高达95%的AUROC。