Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes

arXiv cs.AI Papers

Summary

This paper investigates whether linearly decodable failure signals in LLM hidden states can be corrected via residual-stream steering. It finds that while 'overthinking' failures are decodable, fixed linear steering fails to correct them due to representational entanglement with task-critical computations, though the probes effectively support selective abstention.

arXiv:2605.05715v1 Announce Type: new Abstract: Can linearly decodable failure signals in LLM hidden states be leveraged to correct those failures? We investigate this classification-correction gap via Overthinking (OT)--a stable behavioral regime (Jaccard >= 0.81, 94% inter-annotator agreement) in medical QA where models answer correctly under resampling yet fail in extended chain-of-thought. OT is linearly decodable at 71.6% balanced accuracy (p < 10^{-16}). Yet five families of fixed linear steering (29 configurations, n=1,273) all yield Delta ~= 0, with identical null results cross-architecture (Qwen2.5-7B) and cross-domain (MMLU-STEM). Three convergent lines of evidence suggest representational entanglement: the OT direction has 85-88% overlap with task-critical computation (specificity ratio <= 0.152); non-targeted shared-direction steering damages accuracy (-12.1pp); and LEACE concept erasure damages accuracy (-3.6pp, p=0.01), while 10 random erasures produce Delta=+0.3pp. The per-instance probe-steering correlation is r=-0.002 (p=0.97). Positively, the same probe enables selective abstention (held-out AUROC=0.610, exceeding all five uncertainty baselines, p=0.009): decodable failure structure supports post-generation reliability estimation even when the fixed linear steering family cannot exploit it for correction.
Original Article
View Cached Full Text

Cached at: 05/08/26, 08:38 AM

# Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes
Source: [https://arxiv.org/html/2605.05715](https://arxiv.org/html/2605.05715)
###### Abstract

Can linearly decodable failure signals in LLM hidden states be leveraged to correct those failures? We investigate this*classification\-correction gap*via*Overthinking*\(OT\)—a stable behavioral regime \(Jaccard≥0\.81\\geq 0\.81, 94% inter\-annotator agreement\) in medical QA where models answer correctly under resampling yet fail in extended chain\-of\-thought\.OTis linearly decodable at 71\.6% balanced accuracy \(p≈10−16p\\approx 10^\{\-16\}\)\. Yet five families of fixed linear steering \(29 configurations,n=1,273n=1\{,\}273\) all yieldΔ≈0\\Delta\\approx 0, with identical null results cross\-architecture \(Qwen2\.5\-7B\) and cross\-domain \(MMLU\-STEM\)\. Three convergent lines of evidence suggest*representational entanglement*: theOTdirection has 85–88% overlap with task\-critical computation \(specificity ratio≤0\.152\\leq 0\.152\); non\-targeted shared\-direction steering damages accuracy \(−\-12\.1pp\); and LEACE concept erasure damages accuracy \(−\-3\.6pp,p=0\.01p=0\.01\), while 10 random erasures produceΔ=\+0\.3\\Delta=\+0\.3pp\. The per\-instance probe–steering correlation isr=−0\.002r=\-0\.002\(p=0\.97p=0\.97\)\. Positively, the same probe enables selective abstention \(held\-out AUROC = 0\.610, exceeding all five uncertainty baselines,p=0\.009p=0\.009\): decodable failure structure supports post\-generation reliability estimation even when the fixed linear steering family cannot exploit it for correction\.

Decodable but Not Corrected by Fixed Residual\-Stream Linear Steering: Evidence from Medical LLM Failure Regimes

Ming LiuAmazonmlliuz@amazon\.com

## 1Introduction

Activation steering—adding directions in activation space to modify model behavior—succeeds for binary behavioral properties like refusal\(Arditi et al\.,[2024](https://arxiv.org/html/2605.05715#bib.bib3)\)and sentiment\(Turner et al\.,[2024](https://arxiv.org/html/2605.05715#bib.bib63)\)\. But does this extend to correcting multi\-step reasoning failures? Chain\-of\-thought prompting\(Wei et al\.,[2022](https://arxiv.org/html/2605.05715#bib.bib71); Kojima et al\.,[2022](https://arxiv.org/html/2605.05715#bib.bib36)\)enables multi\-step reasoning but also introduces failure modes tied to extended generation\(Turpin et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib64); Lanham et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib39)\); indeed,Huang et al\. \([2024](https://arxiv.org/html/2605.05715#bib.bib31)\)show that LLMs cannot self\-correct reasoning without external feedback\. We use medical question answering as a controlled testbed: it demands multi\-step reasoning, failures carry distinct semantic signatures, and the multiple\-choice format provides an unambiguous evaluation endpoint\.

We focus on a specific empirical question:*when models frequently answer correctly yet sometimes fail through extended reasoning, can the representational structure underlying this behavioral regime be leveraged for activation steering?*

We address this question through a systematic investigation with four contributions:

1. 1\.Behavioral regime identification\.We identify*Overthinking*\(OT\)—where models frequently answer correctly under alternative sampled traces but produce incorrect long\-reasoning outputs—as a robust behavioral construct, stable under threshold perturbation, length regression, and prompt\-end probing \(Section[3](https://arxiv.org/html/2605.05715#S3)\)\.
2. 2\.Geometric analysis\.BinaryOT\-vs\-non\-OTclassification achieves 71\.6% \(modest but highly reliable;p≈10−16p\\approx 10^\{\-16\}\), with supporting evidence on Qwen2\.5\-7B and cross\-domain transfer to MMLU\-STEM \(decodability and steering failure\)\. The OT\-specific specificity ratio is only 0\.119 \(88% of the signal is shared with task\-relevant directions\), and concept erasure damages accuracy \(−3\.6\-3\.6pp,p=0\.01p=0\.01\), providing causal evidence that the decodable signal is entangled with task computation rather than merely correlational\.
3. 3\.An empirical classification\-correction gap\.Five families of fixed linear steering \(contrastive, probe\-guided, multi\-layer, prompt\-end, and subspace; 29 configurations total,n=1,273n=1\{,\}273\) on Llama\-3\.1\-8B produceΔ≈0\\Delta\\approx 0, while shared\-direction steering \(−\-12\.1pp\) and mean\-difference concept erasure \(−\-3\.6pp,p=0\.01p=0\.01; direction\-specific vs\. random erasures\) damage accuracy—consistent with failure\-mode information co\-occurring with task\-relevant computation\. Qwen2\.5\-7B \(9 configurations\), MMLU\-STEM \(4 configurations\), and ten prompt baselines including self\-refinement and verbalized confidence are consistent with this pattern\. A refusal\-steering control verifies the implementation \(p=0\.008p=0\.008, one\-sided\)\.
4. 4\.Post\-generation reliability estimation\.The same structure enables selective abstention after a single generation: a correctness probe \(layer 21, selected via held\-out AUROC from a 32\-layer scan\) achieves held\-out split AUROC = 0\.716 \(held\-out test: 0\.610; generalization gap is primarily temperature\-driven\), outperforming all five tested single\-forward\-pass uncertainty baselines \(Δ\\DeltaAUROC = 0\.041,p=0\.009p=0\.009; Holm rank\-1 threshold = 0\.010\)\.*Reading*representations has operational value even when linear*writing*to them does not—probing success, even when robust and cross\-domain transferable, does not imply steerability\.

## 2Related Work

#### Linear directions can steer binary behavioral attributes—but unreliably\.

Representation engineering\(Zou et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib78)\)shows that linear directions correspond to interpretable concepts\. Inference\-time intervention\(Li et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib40)\), activation addition\(Turner et al\.,[2024](https://arxiv.org/html/2605.05715#bib.bib63)\), contrastive activation addition\(Panickssery et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib52)\), mean\-centred variants\(Jorgensen et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib34)\), and feature clamping via sparse autoencoders\(Templeton et al\.,[2024](https://arxiv.org/html/2605.05715#bib.bib61); Cunningham et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib14)\)demonstrate successful steering for binary behavioral traits \(refusal, sentiment, sycophancy\(Sharma et al\.,[2024](https://arxiv.org/html/2605.05715#bib.bib58)\)\) and in\-context task identity\(Todd et al\.,[2024](https://arxiv.org/html/2605.05715#bib.bib62)\), while model editing\(Meng et al\.,[2022](https://arxiv.org/html/2605.05715#bib.bib46)\)modifies localized factual associations\. Layer\-contrast decoding\(Chuang et al\.,[2024](https://arxiv.org/html/2605.05715#bib.bib12)\)provides an alternative inference\-time intervention that improves factuality by exploiting layer\-wise knowledge maturation\. Representation finetuning\(Wu et al\.,[2024](https://arxiv.org/html/2605.05715#bib.bib74)\)instead*learns*targeted interventions on frozen representations\. However, growing evidence reveals systematic steering failures:Da Silva et al\. \([2025](https://arxiv.org/html/2605.05715#bib.bib15)\)find substantial variability across 36 models from 14 families;Braun et al\. \([2025](https://arxiv.org/html/2605.05715#bib.bib9)\)document concepts that remain unsteerable across all layers;McKenzie et al\. \([2026](https://arxiv.org/html/2605.05715#bib.bib45)\)identify endogenous resistance circuits that actively counteract interventions; andZur et al\. \([2025](https://arxiv.org/html/2605.05715#bib.bib79)\)show steering fails once a model has committed to a reasoning path—even when hidden states still encode uncertainty\.Jafari et al\. \([2026](https://arxiv.org/html/2605.05715#bib.bib32)\)propose mechanistic indicators for predicting effectiveness a priori, andBilla \([2026](https://arxiv.org/html/2605.05715#bib.bib8)\)formalize a three\-regime framework in which some concepts are provably beyond any linear intervention’s reach\.

#### But probed information need not be causally effective\.

The linear representation hypothesis\(Park et al\.,[2024](https://arxiv.org/html/2605.05715#bib.bib53); Marks and Tegmark,[2024](https://arxiv.org/html/2605.05715#bib.bib44); Nanda et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib50); Hernandez et al\.,[2024](https://arxiv.org/html/2605.05715#bib.bib28)\)and probing methods\(Alain and Bengio,[2017](https://arxiv.org/html/2605.05715#bib.bib1); Belinkov,[2022](https://arxiv.org/html/2605.05715#bib.bib6); Hewitt and Liang,[2019](https://arxiv.org/html/2605.05715#bib.bib29)\)establish that LLMs encode rich structure—including correctness signals\(Azaria and Mitchell,[2023](https://arxiv.org/html/2605.05715#bib.bib4); Burns et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib10)\)—but whether probed information is*causally*used remains contested\(Elazar et al\.,[2021](https://arxiv.org/html/2605.05715#bib.bib17); Ravfogel et al\.,[2022](https://arxiv.org/html/2605.05715#bib.bib56)\)\.Deng et al\. \([2025](https://arxiv.org/html/2605.05715#bib.bib16)\)argue confounding bias can cause probes to find correlational rather than causal directions; however, our LEACE erasure result \(−3\.6\-3\.6pp damage, direction\-specific\) rules out the purely correlational account in our setting, instead supporting causal entanglement with task computation\. Concept erasure methods—INLP\(Ravfogel et al\.,[2020](https://arxiv.org/html/2605.05715#bib.bib55)\), RLACE\(Ravfogel et al\.,[2022](https://arxiv.org/html/2605.05715#bib.bib56)\), LEACE\(Belrose et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib7)\)—and causal mediation\(Vig et al\.,[2020](https://arxiv.org/html/2605.05715#bib.bib66); Geiger et al\.,[2024](https://arxiv.org/html/2605.05715#bib.bib22)\)test this distinction\.Hase et al\. \([2023](https://arxiv.org/html/2605.05715#bib.bib26)\)show that localizing facts does not predict editing success, foreshadowing the gap we observe\.

The decodability–steerability gap is now documented across multiple domains\.Wu et al\. \([2025](https://arxiv.org/html/2605.05715#bib.bib73)\)benchmark this at scale: on Gemma\-2, difference\-of\-means probes achieve best concept detection while prompting outperforms all representation\-based steering methods\.Basu et al\. \([2026](https://arxiv.org/html/2605.05715#bib.bib5)\)find 98% AUROC internal hazard representations in clinical triage that four mechanistic interventions fail to exploit\.Sanyal et al\. \([2025](https://arxiv.org/html/2605.05715#bib.bib57)\)andCox et al\. \([2026](https://arxiv.org/html/2605.05715#bib.bib13)\)report analogous gaps in math solvability and factual reasoning\.Wang et al\. \([2026](https://arxiv.org/html/2605.05715#bib.bib69)\)coin “representation\-behavior gap” for the same phenomenon in tool\-calling agents \(99% probe AUC, model still fails to act\)\.Mishra et al\. \([2026](https://arxiv.org/html/2605.05715#bib.bib47)\)prove steered activations leave the prompt\-reachable set;Nadaf \([2026](https://arxiv.org/html/2605.05715#bib.bib49)\)document the inverse \(steerability without decodability\)\. Theoretically,Gao et al\. \([2026](https://arxiv.org/html/2605.05715#bib.bib20)\)extend the linear representation hypothesis to show that concept overlap creates intrinsically unpredictable “sensitive sectors,” whileWollschläger et al\. \([2025](https://arxiv.org/html/2605.05715#bib.bib72)\)prove that geometric orthogonality does not guarantee interventional independence\. When representations are encoded in superposition\(Elhage et al\.,[2022](https://arxiv.org/html/2605.05715#bib.bib18)\)—sharing directions across features—additive steering along one feature’s direction inevitably perturbs others\.

#### Our contribution: mechanistic explanation, not just documentation\.

While the gap’s existence is now established across several concurrent works\(Wu et al\.,[2025](https://arxiv.org/html/2605.05715#bib.bib73); Basu et al\.,[2026](https://arxiv.org/html/2605.05715#bib.bib5); Wang et al\.,[2026](https://arxiv.org/html/2605.05715#bib.bib69)\), none provides: \(a\) systematic geometric characterization of*why*steering fails in a specific domain—we show the OT\-specific specificity ratio is only 0\.119, meaning 88% of the contrastive signal is shared with task\-relevant computation, and concept erasure*damages*accuracy \(−3\.6\-3\.6pp,p=0\.01p=0\.01\), confirming the entanglement is causal rather than merely correlational; \(b\) per\-instance analysis showing probe confidence and steering effect are completely uncorrelated \(r=−0\.002r=\-0\.002,p=0\.97p=0\.97\); or \(c\) an operational positive—redirecting decodable structure toward selective abstention\. We test the commonly assumed probing\-to\-steering move in multi\-step medical reasoning\(Jin et al\.,[2021](https://arxiv.org/html/2605.05715#bib.bib33); Singhal et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib59); Nori et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib51)\)—a domain where overthinking degrades accuracy\(Chen et al\.,[2024](https://arxiv.org/html/2605.05715#bib.bib11); Wang et al\.,[2025](https://arxiv.org/html/2605.05715#bib.bib70)\)and chain\-of\-thought benefits remain limited compared to symbolic tasks\(Sprague et al\.,[2024](https://arxiv.org/html/2605.05715#bib.bib60)\)—with five method families \(29 configurations\), and provide convergent geometric evidence \(low specificity, directional damage, cross\-domain subspace degradation\) that the gap arises from representational entanglement with task computation, consistent with the structurally unreachable regime ofBilla \([2026](https://arxiv.org/html/2605.05715#bib.bib8)\)\. This contrasts with domains where steering succeeds \(factual recall\(Li et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib40)\), refusal\(Arditi et al\.,[2024](https://arxiv.org/html/2605.05715#bib.bib3)\)\)—where target directions have higher specificity ratios \(refusal: 0\.999 vs\.OT:≤\\leq0\.21; a4\.7×4\.7\\timesgap that is predictive of steerability in our two\-point comparison\)\.

## 3Failure Characterization

We characterize failures along two axes\.Axis A \(primary\): Behavioral regime\(OTvs\. non\-OT\), defined by cross\-trace statistics and carrying our strongest claims\.Axis B \(exploratory\): Error mechanism\(KDvs\.RCB, within non\-OTonly\), defined by semantic judgment with lower inter\-annotator agreement\. Core findings do not depend on theKD/RCBboundary\.

#### Primary construct: Overthinking \(OT\)\.

We use “Overthinking” as a behavioral label for a sampling\-level pattern, not as a claim about internal cognitive mechanism; cf\.Chen et al\. \([2024](https://arxiv.org/html/2605.05715#bib.bib11)\)on excessive computation in o1\-like models andMuennighoff et al\. \([2025](https://arxiv.org/html/2605.05715#bib.bib48)\)on budget forcing\. Operationally: the model answers correctly in≥\\geq60% of sampled traces\(Wang et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib68)\)—a conservative majority threshold—but produces incorrect answers in long\-reasoning traces \(\>\>200 tokens\)\. The length threshold is nearly redundant \(\>\>99% of incorrect traces exceed 200 tokens\), so the effectiveOTboundary is determined by the correct\-rate threshold\.OTis defined by*observable behavioral statistics*, yielding high reproducibility: 94% expert agreement, 100% within\-question purity, and Jaccard≥\\geq0\.81 under threshold perturbation \(Section[4](https://arxiv.org/html/2605.05715#S4)\)\. Length regression and prompt\-end probing \(54\.4% vs\. 50% chance\) confirmOTreflects a generation\-time regime rather than question difficulty or response length alone\. Our research question is whether representational structure tracking this regime can support correction, regardless of what mechanism produces it\(Turpin et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib64)\)\.

#### Exploratory sub\-classification:KDandRCB\.

Within non\-OTfailures,*Knowledge Deficit*\(KD\) denotes factual errors and*Reasoning Chain Break*\(RCB\) denotes valid\-fact/invalid\-logic errors\. This distinction requires semantic judgment \(κ=0\.61\\kappa=0\.61expert\-expert,κ=0\.30\\kappa=0\.30LLM; Section[3\.3](https://arxiv.org/html/2605.05715#S3.SS3)\)\.

### 3\.1Annotation Pipeline

Annotation proceeds in two phases: \(1\) heuristic OT detection—if≥\\geq60% of 10 sampled traces are correct, incorrect traces\>\>200 tokens are labeledOT; \(2\) LLM\-basedKD/RCBclassification of remaining traces via Claude Haiku\(Anthropic,[2024](https://arxiv.org/html/2605.05715#bib.bib2)\)following the LLM\-as\-judge paradigm\(Zheng et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib77)\)\. Full prompt and pipeline details are in Appendix[A\.5](https://arxiv.org/html/2605.05715#A1.SS5)\.

### 3\.2Annotation Statistics

Table 1:Failure mode annotation distribution\. Values show count and percentage of all incorrect traces\. Llama annotations are over 34,532 traces from 10,178 MedQA questions \(10 traces/question\)\. Qwen annotations are over 39,862 traces from the same questions\.Table[1](https://arxiv.org/html/2605.05715#S3.T1)shows the annotation distribution\.RCBis the most common failure mode \(34\.5%\), followed byOT\(31\.0%\) andKD\(21\.5%\), with consistent proportions across architectures\. Unclear cases \(13\.0%/17\.5%\) are excluded from subsequent analysis\.

### 3\.3Annotation Validation

Domain expert validation \(n=500n=500gold set, two blinded 4th\-year clinical students\) yields 94% agreement onOT, with expert\-expertκ=0\.61\\kappa=0\.61on the three\-way task; theKD/RCBboundary is the primary source of disagreement\. Within\-questionOTpurity is 100% \(all incorrect traces from anOTquestion receive the same label\)\. Simulating 37%KD/RCBlabel noise \(matching observed LLM disagreement\) drops three\-way classification only 2\.9pp while leaving the binaryOT\-vs\-non\-OTresult \(71\.6%\) unaffected\. Full validation details \(LLM cross\-validation, multi\-model comparison, within\-question consistency analysis\) are in Appendix[A\.6](https://arxiv.org/html/2605.05715#A1.SS6)\.

### 3\.4Behavioral Validation

OTshows a steep negative length\-accuracy gradient \(ρ=−0\.164\\rho=\-0\.164,p<0\.001p<0\.001\), confirming that extended reasoning is associated with lower accuracy, whileKD/RCBshow weak slopes \(ρ=−0\.044,−0\.082\\rho=\-0\.044,\-0\.082respectively\)\.OThas 78\.3% base accuracy vs\.∼\\sim28% forKD/RCB, confirming a clean two\-way behavioral split\. Self\-consistency\(Wang et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib68)\)\(majority vote, 10 traces\) achieves AUROC = 0\.804 for correctness prediction and—by construction of the≥\\geq60% threshold—corrects 100% of OT\-incorrect questions, yielding\+10\.6\+10\.6pp on training and\+7\.4\+7\.4pp overall on the held\-out test \(p<10−10p<10^\{\-10\}; Section[6](https://arxiv.org/html/2605.05715#S6)\)\. A single\-forward\-pass correctness probe achieves AUROC = 0\.716 \(held\-out 30% split; 5\-fold CV: 0\.721±\\pm0\.005; independent test set: 0\.610\) using a geometrically different signal \(Section[4](https://arxiv.org/html/2605.05715#S4)\), substantially exceeding surface uncertainty baselines \(all≤0\.530\\leq 0\.530on training,≤0\.569\\leq 0\.569on test;Δ\\DeltaAUROC = 0\.041,p=0\.009p=0\.009; Tables[6](https://arxiv.org/html/2605.05715#A7.T6)–[7](https://arxiv.org/html/2605.05715#A7.T7)\)\.

## 4Geometric Analysis of Failure Modes

### 4\.1Hidden State Extraction

We extract hidden states from Llama\-3\.1\-8B\-Instruct\(Grattafiori et al\.,[2024](https://arxiv.org/html/2605.05715#bib.bib24)\)for 12,000 traces: 2,000 per failure mode \(KD,RCB,OT\) plus 6,000 matched correct traces \(2,000 correct traces from the same questions as each failure mode\)\. For each trace, we construct the full conversation \(system prompt \+ question \+ model response\), extract the last\-token hidden state at all 32 layers in fp32 precision, yielding a tensor of shape\(12000,32,4096\)\(12000,32,4096\)\.

### 4\.2Contrastive Vectors

Following the representation engineering framework, we compute contrastive vectors:

𝐯m\(l\)=𝐡¯correct\(l\)−𝐡¯m\(l\)‖𝐡¯correct\(l\)−𝐡¯m\(l\)‖\\mathbf\{v\}\_\{m\}^\{\(l\)\}=\\frac\{\\bar\{\\mathbf\{h\}\}\_\{\\text\{correct\}\}^\{\(l\)\}\-\\bar\{\\mathbf\{h\}\}\_\{m\}^\{\(l\)\}\}\{\\\|\\bar\{\\mathbf\{h\}\}\_\{\\text\{correct\}\}^\{\(l\)\}\-\\bar\{\\mathbf\{h\}\}\_\{m\}^\{\(l\)\}\\\|\}\(1\)where𝐡¯correct\(l\)\\bar\{\\mathbf\{h\}\}\_\{\\text\{correct\}\}^\{\(l\)\}and𝐡¯m\(l\)\\bar\{\\mathbf\{h\}\}\_\{m\}^\{\(l\)\}are the mean hidden states at layerllfor correct traces and modemmtraces, respectively\.

![Refer to caption](https://arxiv.org/html/2605.05715v1/x1.png)Figure 1:Layer\-wise classification accuracy profiles\. \(a\) Exploratory three\-way classification \(KD/RCB/OT\) using PCA\-50 \+ logistic regression, peaking at 51\.5% \(chance = 33\.3%\)\. \(b\) Binary classification \(correct vs\. mode\) showing mode\-specific layer profiles:OTpeaks at layer 17 \(81\.5%\), providing the strongest and most annotation\-robust signal\.
### 4\.3Statistical Analysis Suite

#### Classification\.

Three\-wayKD/RCB/OTclassification reaches 51\.5% at layer 16 \(chance = 33\.3%; Figure[1](https://arxiv.org/html/2605.05715#S4.F1)a\)\. Per\-mode correct\-vs\-incorrect probes peak at 76–82% with mode\-specific layer profiles \(Figure[1](https://arxiv.org/html/2605.05715#S4.F1)b\)\. BinaryOT\-vs\-non\-OTclassification achieves71\.6%accuracy \(balanced accuracy = 62\.3%, AUROC = 0\.672; majority baseline = 66\.7%,p≈10−16p\\approx 10^\{\-16\}\)\. The effect size is modest \(4\.9pp above majority baseline\): the probe achieves high specificity \(90% non\-OTrecall\) but low sensitivity \(34%OTrecall\), reflecting moderate but statistically reliable separability\. Contrastive vectors show high pairwise cosine similarity \(mean 0\.820; Figure[2](https://arxiv.org/html/2605.05715#S4.F2)\), indicating substantial overlap across modes\. The observed pairwise distance \(3\.38\) is 2\.4×\\timesthe permutation null \(p<0\.001p<0\.001\); shuffled labels yield chance\-level classification \(34\.6%\), ruling out data artifacts\.

![Refer to caption](https://arxiv.org/html/2605.05715v1/x2.png)Figure 2:Pairwise cosine similarity between contrastive vectors across layers\. Early layers show low similarity \(mode\-specific information\), while later layers converge to high similarity \(\>\>0\.9\), reflecting the dominance of the shared component\.

### 4\.4Robustness Controls

The 71\.6% binary classification is robust to six potential confounds \(full details in Appendix[B](https://arxiv.org/html/2605.05715#A2)\): random labels yield chance \(33\.2%\); length regression leavesOT\-vs\-non\-OTunaffected \(71\.3%\); threshold sweeps produce Jaccard≥\\geq0\.81; GroupKFold rules out question\-identity leakage \(<<1pp drop\)\. Prompt\-end probing achieves only 54\.4% \(vs\. 71\.6% at last\-token\), confirming the probe detects a generation\-time regime rather than question difficulty\.

### 4\.5Direction Ablation Analysis

Removing 3 contrastive directions \(from 4096\) drops in\-sample classification from 51\.4% to 25\.8% \(below chance\), while removing random directions has zero effect\. However, cross\-validated ablation \(directions fit on split A, evaluated on split B\) shows no significant drop \(Δ=−0\.2\\Delta=\-0\.2pp\), indicating that the in\-sample effect reflects probe\-direction overfitting rather than a uniquely necessary subspace\. These directions produce no reliable behavioral change under additive steering \(Section[5](https://arxiv.org/html/2605.05715#S5)\)—necessity for a probe’s classification does not entail causal involvement in the model’s generation process\.

### 4\.6Shared\-Specific Decomposition

The average specificity ratio is only0\.119: 88% of each contrastive vector aligns with a shared “incorrect\-vs\-correct” direction, confirmed under binaryOT\-vs\-non\-OTframing \(bypassingKD/RCBentirely\)\. A permutation test \(10,000 label shuffles among incorrect traces\) yields null specificity mean = 0\.370; observed = 0\.119 is significantly*below*all permuted values \(p<0\.0001p<0\.0001\), confirming anomalously high alignment with the shared axis rather than an artifact of set\-subset structure\. To further rule out circularity \(sinceOTtraces are a subset of incorrect traces\), we compute𝐝OT vs KD=𝐡¯OT−𝐡¯KD\\mathbf\{d\}\_\{\\text\{OT vs KD\}\}=\\bar\{\\mathbf\{h\}\}\_\{\\text\{OT\}\}\-\\bar\{\\mathbf\{h\}\}\_\{\\text\{KD\}\}—a direction that makes no reference to correct traces—and findcos⁡\(𝐝OT vs KD,𝐝correctness\)=0\.86\\cos\(\\mathbf\{d\}\_\{\\text\{OT vs KD\}\},\\mathbf\{d\}\_\{\\text\{correctness\}\}\)=0\.86, demonstrating that the entanglement persists even when the correct class is removed from the computation\. Uniform steering using this shared component is harmful \(Δ=−12\.1\\Delta=\-12\.1pp; Table[3](https://arxiv.org/html/2605.05715#S5.T3)\), consistent with the failure\-mode direction sharing representational subspace with task\-relevant computation\.

### 4\.7Probe\-Based Analysis

Probe weight vectors are geometrically dissociated from contrastive vectors \(cosine 0\.23–0\.54 vs\. 0\.82\), yet achieve comparable classification \(51\.2% vs\. 51\.5%\)\. This dissociation—same accuracy, different directions—illustrates that decodability does not uniquely identify an intervention target \(Appendix[F](https://arxiv.org/html/2605.05715#A6)\)\. Non\-linear probes \(MLP, SVM\-RBF\) perform at or below linear on classification tasks \(Δ≤−0\.2\\Delta\\leq\-0\.2pp for MLP\); SVM\-RBF gains \+1\.4pp on correctness AUROC \(0\.734 vs\. 0\.720\), a marginal improvement confirming the signal ceiling is close to what linear probes capture\.

### 4\.8Cross\-Architecture Validation

We repeat the full geometric analysis suite on Qwen2\.5\-7B\-Instruct\(Yang et al\.,[2024](https://arxiv.org/html/2605.05715#bib.bib76)\)\(28 layers, 3584 hidden dimensions\), sampling 12,000 traces \(2,000 per mode \+ 6,000 correct\) from 39,862 annotated failures\.

#### Consistent classification signal\.

Three\-way classification reaches46\.9%at layer 18 \(chance = 33\.3%,p<0\.001p<0\.001\), consistent with cross\-architecture generalization\. Randomized labels produce 34\.0% \(≈\\approxchance\)\.

#### Divergent geometry\.

The geometry differs substantially: Qwen’s average pairwise cosine is0\.339\(vs\. 0\.820\), average specificity ratio0\.414\(vs\. 0\.119\), and layer profiles diverge \(Table[2](https://arxiv.org/html/2605.05715#S4.T2)\)\. The elevated Qwen average is driven by KD geometry \(0\.662\); OT\-specific specificity is comparable across architectures \(Qwen: 0\.152, Llama: 0\.119\), suggesting the OT encoding is similarly entangled in both models\. A steering test on Qwen \(n=1,273n=1\{,\}273, nine configurations spanning layers 5–18, amplitudesα∈\[0\.5,3\.0\]\\alpha\\in\[0\.5,3\.0\], and mode\-specific/uniform/multi\-layer variants; Section[5](https://arxiv.org/html/2605.05715#S5)\) yieldsΔ∈\[−0\.9,\+0\.8\]\\Delta\\in\[\-0\.9,\+0\.8\]pp \(allp\>0\.05p\>0\.05\), providing consistent evidence for the cross\-architecture steering null across diverse hyperparameter settings\.

Table 2:Cross\-architecture comparison\. Both models show significant failure mode structure \(p<0\.001p<0\.001\), but the underlying geometry varies substantially: Qwen has more distinct modes \(lower cosine, higher specificity\) with different layer localization, suggesting the encoding is architecture\-dependent\.

### 4\.9Cross\-Domain Validation: MMLU\-STEM

To test domain generality, we replicate the core pipeline on MMLU\-STEM\(Hendrycks et al\.,[2021](https://arxiv.org/html/2605.05715#bib.bib27)\)\(300 questions, 18 subjects, 10 traces each, same OT criteria\)\.OTprevalence is comparable \(33\.7% vs\. 31\.0%; 223 traces\), and an in\-domain probe reaches 70\.0%\. A MedQA\-trained probe transfers zero\-shot at 61\.4% balanced accuracy \(z=6\.15z=6\.15,p<10−9p<10^\{\-9\}; permutation test with 1,000 label shuffles,n=446n=446traces: 223 OT \+ 223 matched correct\)\. Dimensionality analysis reveals the geometric basis: in the top\-3 PCA subspace \(fit on MedQA\), cross\-domain cosine is 0\.87 at the peak transfer layer 18 \(mean 0\.89, layers 10–24; notably lower at the primary steering layers 16–17: 0\.67/0\.65\), while the full 4096\-dim cosine is only 0\.44\. This alignment increases monotonically with subspace rank \(PCA\-3: 0\.87, PCA\-5: 0\.94, PCA\-10: 0\.97 at layer 18\), confirming that theOTsignal occupies a genuinely shared low\-rank subspace rather than reflecting an artifact of low\-dimensional projection; the moderate raw\-space alignment reflects dilution by noise dimensions\.

## 5Activation Steering Experiments

Given the weak but statistically reliable geometric structure identified in Section[4](https://arxiv.org/html/2605.05715#S4), we investigate whether it enables effective activation steering to correct failures\.

### 5\.1Experimental Setup

We evaluate on the full MedQA test set \(1,273 questions\) using Llama\-3\.1\-8B\-Instruct, with a sanity check on Qwen2\.5\-7B\-Instruct\. Steering vectors are applied via forward hooks addingα⋅𝐯\\alpha\\cdot\\mathbf\{v\}to the residual stream\(Elhage et al\.,[2021](https://arxiv.org/html/2605.05715#bib.bib19)\)at target layers\. Atα=1\.5\\alpha=1\.5, perturbation norm is∼\\sim17% of residual stream norm at peak layers 15–17, comparable to prior work\(Turner et al\.,[2024](https://arxiv.org/html/2605.05715#bib.bib63)\)\.

### 5\.2Methods

We test a broad family of linear additive interventions motivated directly by the decodable geometry; we do not claim to exhaust all possible intervention families \(see Limitations\)\.

#### Contrastive Steering\.

We addα⋅𝐯m\(l\)\\alpha\\cdot\\mathbf\{v\}\_\{m\}^\{\(l\)\}to the hidden state at the mode’s peak layer during generation, using either*mode\-specific*vectors \(detected via cosine projection\) or the*uniform*shared component; confidence\-gated variants steer only above a detection threshold\.

#### Probe\-Guided Steering\.

We replace the contrastive vector with the probe weight vector \(n=1,175n=1\{,\}175questions with valid outputs\)\.

#### Multi\-layer Steering\.

We apply mode\-specific vectors simultaneously at 1, 3, or 5 layers centered on the peak layer\.

#### Rank\-kkSubspace Steering\.

We construct rank\-kkbases \(k∈\{1,3,5\}k\\in\\\{1,3,5\\\}\) via SVD of stacked direction matrices \(probe and combined\), testing whether rank\-1 vectors are too restrictive \(details in Appendix[E](https://arxiv.org/html/2605.05715#A5)\)\.

#### Prompt\-End Steering\.

We extract contrastive vectors from*prompt\-end*hidden states—the last token before generation, causally independent of the response—which are near\-orthogonal to last\-token vectors \(cosine 0\.01–0\.17\)\.

### 5\.3Results

MethodAcc\.Δ\\DeltappCorr\.Dmg\.Baseline \(no steering\)65\.1————*Contrastive Steering \(n=1,273\)*Uniform \(shared,α\\alpha=1\.5\)53\.0−\-12\.1<<\.00128\.233\.7Mode\-specific \(α\\alpha=1\.5\)65\.0−\-0\.2\.95332\.017\.4Mode\-specific \(α\\alpha=3\.0\)65\.3\+0\.2\.95231\.116\.4*Probe\-Guided Steering \(n=1,175\)*Probe\-uniform \(α\\alpha=0\.5\)66\.0−\-0\.4\.80633\.017\.3Probe\-mode \(α\\alpha=0\.5\)64\.8−\-1\.7\.23329\.717\.5Probe\-mode \(α\\alpha=1\.0\)65\.1−\-1\.4\.36132\.218\.3*Multi\-layer Contrastive \(n=1,273\)*3 layers \(L13–17\)64\.6\+0\.2\.95636\.419\.95 layers \(L11–19\)62\.8−\-1\.6\.28035\.522\.2*Prompt\-End Vectors \(n=1,273\)*Uniform \(shared,α\\alpha=1\.5\)65\.3\+0\.5\.73535\.618\.6Mode\-specific \(α\\alpha=1\.5\)66\.4\+1\.6\.24134\.716\.4Mode\-specific \(α\\alpha=3\.0\)65\.0\+0\.2\.91237\.019\.8Full composite \(α\\alpha=1\.5\)65\.7\+0\.9\.54237\.619\.1*Strong Probe Steering \(n=1,273\)*OT probe \(L17,α\\alpha=1\.5\)66\.7\+1\.5\.29636\.015\.8OT probe \(L17,α\\alpha=3\.0\)61\.4−\-3\.8\.01033\.722\.4OT probe \(L14,α\\alpha=1\.5\)65\.4\+0\.2\.95333\.316\.4*Cross\-architecture: Qwen \(n=1,273\)*Mode\-specific \(α\\alpha=1\.5\)61\.3−\-0\.11\.0018\.511\.8Uniform \(α\\alpha=1\.5\)61\.9\+0\.5\.67721\.712\.8Mode\-specific \(α\\alpha=3\.0\)61\.0−\-0\.4\.77819\.913\.2*Concept Erasure \(n=1,273\)*Mean\-diff erasure \(L16\)63\.5−\-3\.6\.01031\.320\.7Table 3:Steering results on MedQA test set\. All values in %\. Acc\. = accuracy;Δ\\Delta= change vs\. baseline \(pp\);pp= McNemar two\-sided; Corr\. = % of baseline\-incorrect questions corrected; Dmg\. = % of baseline\-correct questions damaged\. Prompt\-end vectors are constructed from pre\-generation hidden states and are near\-orthogonal to last\-token vectors \(cosine≤\\leq0\.17\)\. All experiments use temperature 0\.1 withdo\_sample=True, introducing run\-to\-run variability; baseline varies across independent runs \(Llama: 64\.0–67\.4%; Qwen: 61\.4%\)\. EachΔ\\Deltais computed against its own run’s paired baseline \(not cross\-run\)\. For context, majority vote \(k=10k=10,10×10\\timescost, evaluated atT=0\.8T\{=\}0\.8\) achievesΔ=\+10\.6\\Delta=\+10\.6pp on training data and\+7\.4\+7\.4pp on the test set \(p<10−10p<10^\{\-10\}\); note the temperature difference from steering experiments\.![Refer to caption](https://arxiv.org/html/2605.05715v1/x3.png)Figure 3:Steering experiment results\. \(a\) Accuracy delta vs baseline: uniform shared steering damages performance \(−\-12\.1pp\), while all targeted methods produceΔ≈0\\Delta\\approx 0\. \(b\) Correction vs damage trade\-off: all methods lie near or above the net\-zero diagonal, meaning corrections are offset by comparable damage\.Table[3](https://arxiv.org/html/2605.05715#S5.T3)and Figure[3](https://arxiv.org/html/2605.05715#S5.F3)present the results\. No steering method reliably improves accuracy\. On the full test set \(nn=1,273\), mode\-specific steering producesΔ=−0\.2\\Delta=\-0\.2pp \(95% CI: \[−2\.8\-2\.8,\+2\.4\+2\.4\]pp; TOST equivalence\(Lakens,[2017](https://arxiv.org/html/2605.05715#bib.bib38)\)within±2\.5\\pm 2\.5pp—chosen a priori as half the 5pp minimal clinically important difference\(following non\-inferiority trial conventions where the margin is≤\\leq50% of the established effect; Walker and Nowacki,[2011](https://arxiv.org/html/2605.05715#bib.bib67)\)—p=0\.039p=0\.039; TOST uses 90% CI\[−2\.4,\+2\.0\]\[\-2\.4,\+2\.0\]pp, which falls within bounds atα=0\.05\\alpha=0\.05; power is approximately 20% per individual test at this sample size; however, the joint evidence across all 29 configurations is far more decisive—were the true effect≥\\geq2\.5pp, observing all configurations near zero would occur with probability<10−15<10^\{\-15\}\)\. Uniform shared steering severely damages performance \(Δ=−12\.1\\Delta=\-12\.1pp\)\. Multi\-layer steering producesΔ≈0\\Delta\\approx 0: 3\-layer\+0\.2\+0\.2pp \(p=0\.956p=0\.956\), 5\-layer−1\.6\-1\.6pp \(p=0\.280p=0\.280\)\. Prompt\-end steering—using vectors near\-orthogonal to last\-token directions \(cosine≤\\leq0\.17\)—also yieldsΔ≈0\\Delta\\approx 0across all four conditions \(\+0\.2\+0\.2to\+1\.6\+1\.6pp, all 95% CIs including zero\)\. Qwen2\.5\-7B \(nine configurations spanning layers 5–18,α∈\[0\.5,3\.0\]\\alpha\\in\[0\.5,3\.0\], mode\-specific and uniform vectors\) yieldsΔ∈\[−0\.9,\+0\.8\]\\Delta\\in\[\-0\.9,\+0\.8\]pp \(allp\>0\.05p\>0\.05; all deltas fall well within the pre\-specified±2\.5\\pm 2\.5pp equivalence margin\), strengthening the cross\-architecture steering null\. Supplementary experiments confirm these findings across temperature settings\. Same\-temperature evaluation atT=0\.8T=0\.8\(n=300n=300; Appendix[K](https://arxiv.org/html/2605.05715#A11)\) produces*larger*damage \(Δ=−6\.7\\Delta=\-6\.7pp,p=0\.031p=0\.031\), making temperature mismatch an unlikely sole explanation for the steering null\. The full sweep testedα∈\{0\.5,1\.0,1\.5,2\.0,3\.0\}\\alpha\\in\\\{0\.5,1\.0,1\.5,2\.0,3\.0\\\}per method family \(Appendix[L](https://arxiv.org/html/2605.05715#A12)\); all omitted configurations also yieldΔ≈0\\Delta\\approx 0\. Because we evaluate multiple configurations, no individual positiveΔ\\Deltasurvives Holm correction; the negative conclusion rests on the full sweep, not any single null test\.

Nine validity threats to the steering null are addressed by targeted controls \(Appendix[C](https://arxiv.org/html/2605.05715#A3)\): broken hooks \(refusal control, one\-sidedp=0\.008p=0\.008\), vector position \(prompt\-endΔ≈0\\Delta\\approx 0\), rank limitation \(rank\-k≤5k\\leq 5,Δ≈0\\Delta\\approx 0\), temperature mismatch \(T=0\.8T\{=\}0\.8more harmful\), metric insensitivity \(38–56% answers change\), non\-linear learned intervention \(Δ=\+2\.8\\Delta=\+2\.8pp,p=0\.025p=0\.025; Appendix[D](https://arxiv.org/html/2605.05715#A4)\), random artifact \(10 random directions≈0\\approx 0; OT harmful\), prompt fixability \(7 variants, allp\>0\.5p\>0\.5\), and single\-layer limitation \(multi\-layerΔ≈0\\Delta\\approx 0\)\.

#### Behavioral readouts\.

Re\-running three conditions with full text \(n=1,273n=1\{,\}273\) confirms steering substantially alters behavior: 38–56% of answers change \(Jaccard 0\.47\)\. Changes are systematically harmful—damages outnumber corrections 2–3\.6:1 for uniform and OT\-specific steering—and the answer distribution shifts toward a dominant option \(33–52%\), indicating systematic bias rather than targeted correction\. A refusal\-steering positive control confirms the hooks operate correctly \(Appendix[C](https://arxiv.org/html/2605.05715#A3)\)\.

#### Rank\-kksubspace steering\.

To rule out that rank\-1 vectors are too restrictive, we test rank\-kksubspace steering \(k∈\{1,3,5\}k\\in\\\{1,3,5\\\}\) on the full test set \(n=1,273n=1\{,\}273; Table[5](https://arxiv.org/html/2605.05715#A5.T5)in Appendix[E](https://arxiv.org/html/2605.05715#A5)\)\. Per\-mode steering producesΔ≈0\\Delta\\approx 0at all ranks \(TOST\-equivalent within±3\\pm 3pp\); correctness\-uniform steering damages performance at all ranks \(Δ=−2\.6\\Delta=\-2\.6to−6\.2\-6\.2pp\)\. This eliminates “rank\-1 too restrictive” as an explanation\.

#### Concept erasure\.

As an alternative intervention family, we apply mean\-difference concept erasure—projecting*out*failure\-mode directions \(𝐏=𝐈−𝐝^​𝐝^T\\mathbf\{P\}=\\mathbf\{I\}\-\\hat\{\\mathbf\{d\}\}\\hat\{\\mathbf\{d\}\}^\{T\}, where𝐝^=Δ​μ/‖Δ​μ‖\\hat\{\\mathbf\{d\}\}=\\Delta\\mu/\\\|\\Delta\\mu\\\|\) on the full test set \(n=1,273n=1\{,\}273\)\.111We additionally test a whitened variant \(d^=Σw−1/2​Δ​μ\\hat\{d\}=\\Sigma\_\{w\}^\{\-1/2\}\\Delta\\mu, normalized\), which produces a smaller, non\-significant effect \(Δ=−2\.1\\Delta=\-2\.1pp,p=0\.12p=0\.12;cos⁡\(d^whitened,d^raw\)=0\.86\\cos\(\\hat\{d\}\_\{\\text\{whitened\}\},\\hat\{d\}\_\{\\text\{raw\}\}\)=0\.86\), consistent with within\-class whitening partially separating the OT signal from task\-critical computation\. Neither variant is the full oblique LEACE projector ofBelrose et al\. \([2023](https://arxiv.org/html/2605.05715#bib.bib7)\)\(Σw\+1/2​P​Σw−1/2\\Sigma\_\{w\}^\{\+1/2\}P\\Sigma\_\{w\}^\{\-1/2\}\), which remains untested\. The accompanying code uses legacy naming \(compute\_leace\_eraser\)\.Erasing theOTdirection at layer 16 damages accuracy \(Δ=−3\.6\\Delta=\-3\.6pp,p=0\.010p=0\.010, McNemar; 131 corrections vs\. 177 damages, ratio 1\.4:1\)\. This damage is direction\-specific: 10 random directions of equal rank produceΔrand=\+0\.3\\Delta\_\{\\text\{rand\}\}=\+0\.3pp \(σ=1\.8\\sigma=1\.8pp; Appendix[H](https://arxiv.org/html/2605.05715#A8)\), confirming that theOTdirection specifically co\-occurs with task\-relevant computation rather than encoding a separable error signal \(ΔOT−Δ¯rand=−3\.9\\Delta\_\{\\text\{OT\}\}\-\\overline\{\\Delta\}\_\{\\text\{rand\}\}=\-3\.9pp,p<0\.02p<0\.02vs\. random distribution\)\. An alternative interpretation—that the OT direction is causally required for correct reasoning, so erasure damages the computation it supports—is equally consistent \(Appendix[H](https://arxiv.org/html/2605.05715#A8)\)\. Even under oracle centroid displacement, only 58% ofOTtraces move closer to the correct class, confirming that within\-class variance structurally limits linear intervention\.

![Refer to caption](https://arxiv.org/html/2605.05715v1/x4.png)Figure 4:Selective abstention using a binary correctness probe, evaluated on balanced held\-out traces \(50% correct, 50% incorrect\)\. \(a\) Accuracy increases monotonically as the probe abstains on low\-confidence predictions\. \(b\) At key coverage levels, accuracy gains range from \+3\.6pp \(90% coverage\) to \+14\.0pp \(60% coverage\) above the 50% balanced baseline\. In deployment where the model’s prior accuracy differs, absolute gains would scale accordingly; the AUROC \(0\.716\) provides a prevalence\-independent measure\.

## 6Discussion

#### The classification\-correction gap\.

A modest but statistically reliable linearly decodable signal—robust across replications, with evidence on one additional architecture \(9 configurations\) and one additional domain \(decodability and steering failure\)—is insufficient for correction under the tested family of fixed linear interventions\. This is not a universal impossibility claim\. The core mismatch is between*what is decodable*and*what needs to change*: failure*mode identity*is largely pre\-determined before generation, but failure*occurrence*accumulates dynamically during reasoning\. The gap between prompt\-end and last\-token probing \(54\.4% vs\. 71\.6%\) confirms that the probe reads a signal generated*during*the response—which is why it supports post\-generation abstention but not pre\-emptive steering\.

We identify four descriptive correlates \(compatible explanations, not causally identified mechanisms; equally consistent with an alternative account in which the OT direction is causally required for correct reasoning, so perturbation damages it\): \(i\)*observation\-intervention asymmetry*: probes exploit any statistical regularity, while steering requires a causally effective perturbation\(Hase et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib26)\); \(ii\)*low specificity*: 88% of the contrastive signal is shared \(OT\-specific specificity: 0\.119 Llama, 0\.152 Qwen—comparable across architectures despite differing average ratios\), with within\-class variance exceeding the inter\-centroid gap by 2–4×\\times; \(iii\)*cross\-model variation*: architectures encode failure modes differently \(cosine 0\.339 vs\. 0\.820\) yet steering fails on both; \(iv\)*cross\-domain subspace degradation at steering layers*: PCA\-3 cosine is 0\.87–0\.96 at most layers but drops to a local trough at the primary steering layers 16–17 \(0\.67/0\.65\), suggesting the shared subspace degrades precisely at the intervention point\. These correlates are consistent with—but do not uniquely establish—representational superposition\(Elhage et al\.,[2022](https://arxiv.org/html/2605.05715#bib.bib18)\)as the underlying mechanism: when the OT signal shares directions with task\-relevant computation, probes can extract the OT component \(compatible with decodability\), but additive steering along this direction perturbs the task computation \(compatible with both the steering null for targeted methods and the−12\.1\-12\.1pp/−3\.6\-3\.6pp damage for shared/erasure methods\)\. Current data cannot distinguish this from the alternative account that the OT direction is itself a necessary component of correct reasoning \(Section[5](https://arxiv.org/html/2605.05715#S5)\)\. The gap persists across decodability strengths: the strong correct\-vs\-OTprobe \(81\.5%\) yieldsΔ=\+1\.5\\Delta=\+1\.5pp \(p=0\.296p=0\.296\), and increasingα\\alphareverses the effect \(−3\.8\-3\.8pp,p=0\.010p=0\.010\)\. Instance\-adaptive methods \(K\-CAST\(Valentino et al\.,[2026](https://arxiv.org/html/2605.05715#bib.bib65)\), DAS\(Geiger et al\.,[2024](https://arxiv.org/html/2605.05715#bib.bib22)\), ReFT\(Wu et al\.,[2024](https://arxiv.org/html/2605.05715#bib.bib74)\)\) could succeed by learning to disentangle shared directions\. Preliminary evidence supports this: a learned non\-linear residual MLP \(4096→\\to64→\\to4096 bottleneck, trained on the same correct/OTsupervision as linear methods\) achievesΔ=\+2\.8\\Delta=\+2\.8pp \(p=0\.025p=0\.025, McNemar; 140 corrections, 104 damages; Appendix[D](https://arxiv.org/html/2605.05715#A4)\)\. The MLP reduces centroid distance by 50\.6% yet produces only modest behavioral gain—consistent with partial but incomplete disentanglement of the shared subspace\. This suggests the entanglement is a matter of degree: non\-linear transforms can partially navigate around task\-critical directions that linear perturbations inevitably disrupt\.

#### Pre\-intervention diagnostics\.

The four correlates suggest candidate pre\-checks before attempting activation steering: \(i\) specificity ratio, \(ii\) within\-class variance relative to inter\-centroid gap, \(iii\) cross\-architecture stability, and \(iv\) cross\-domain subspace alignment at the intervention layer\(cf\. Jafari et al\.,[2026](https://arxiv.org/html/2605.05715#bib.bib32)\)\. In our setting all four were unfavorable and steering failed; refusal—our positive control \(Appendix[C](https://arxiv.org/html/2605.05715#A3)\)—satisfies the first three and responds to contrastive activation addition \(CAA\)\. Quantitatively, the refusal direction has specificity ratio0\.999\(under the three\-mode shared decomposition; 95% CI forOT: \[0\.075, 0\.119\]\)—nearly all of its variance is mode\-specific, compared to only 12–21% forOT\. The refusal direction is near\-orthogonal to the failure\-mode directions \(cos⁡\(𝐝refusal,𝐝OT\)=−0\.008\\cos\(\\mathbf\{d\}\_\{\\text\{refusal\}\},\\mathbf\{d\}\_\{\\text\{OT\}\}\)=\-0\.008;cos⁡\(𝐝refusal,𝐯¯shared\)=0\.014\\cos\(\\mathbf\{d\}\_\{\\text\{refusal\}\},\\bar\{\\mathbf\{v\}\}\_\{\\text\{shared\}\}\)=0\.014\), confirming it occupies a geometrically distinct subspace\. The specificity gap is4\.7×4\.7\\times\(refusal vs\.OTunder the same three\-mode decomposition\), providing a quantitative two\-point calibration: directions with specificity≥0\.99\\geq 0\.99are steerable \(refusal: sign testp=0\.008p=0\.008\), while directions with specificity≤0\.21\\leq 0\.21are not \(29 configurations, allΔ≈0\\Delta\\approx 0\)\. A Linear Accessibility Profile \(LAP\) analysis\(adapting the framework of Billa,[2026](https://arxiv.org/html/2605.05715#bib.bib8)\)independently confirms the diagnosis: projecting theOTdirection through the unembedding matrix yields semantically incoherent top tokens at all layers, with logit\-lens classification accuracy \(58\.6%\) far below probe accuracy \(81\.5%\)—a gap ofΔ=0\.23\\Delta=0\.23indicating the signal is present but not output\-aligned \(Appendix[M](https://arxiv.org/html/2605.05715#A13)\)\. These diagnostics are hypothesis\-generating checks based on a two\-point comparison \(refusal vs\.OT\), not validated predictors; however, the4\.7×4\.7\\timesspecificity gap provides a concrete quantitative threshold for future work\.

#### Generality beyond medical QA\.

MMLU\-STEM replication \(Section[4](https://arxiv.org/html/2605.05715#S4)\) provides evidence that the classification\-correction gap extends beyond MedQA: comparableOTprevalence \(33\.7%\), in\-domain probes at 70\.0%, zero\-shot MedQA probe transfer \(z=6\.15z=6\.15\) in a shared low\-rank subspace \(top\-3 PCA cosine 0\.87 at peak transfer layer 18; mean 0\.89 across layers 10–24, though lower at primary steering layers 16–17: 0\.67/0\.65\), and—critically—the same steering failure pattern \(n=300n=300\): mode\-specific steeringΔ=0\.0\\Delta=0\.0pp \(p=1\.0p=1\.0\), uniform steeringΔ=−7\.0\\Delta=\-7\.0pp \(p=0\.017p=0\.017\), multi\-layerΔ=−6\.7\\Delta=\-6\.7pp \(p=0\.040p=0\.040\)\. These results provide evidence that the gap extends beyond MedQA, though broader domain coverage is needed to establish full generality\.

#### Self\-consistency as a correction baseline\.

The 100% OT correction by majority vote is definitional: OT requires≥\\geq60% correct traces, so MV@10 necessarily selects the correct answer for every OT question\. The substantive empirical finding is the*overall*test\-set gain: MV@10 achieves 72\.7% \(\+7\.4\+7\.4pp;p<10−10p<10^\{\-10\},n=1,273n=1\{,\}273,T=0\.8T\{=\}0\.8\), correcting OT errors while fixing only 18\.6% of non\-OT errors \(\+10\.6\+10\.6pp on training data\)\. Best\-of\-NNprobe selection atk=10k\{=\}10achieves 70\.7% \(\+5\.3\+5\.3pp;p<10−4p<10^\{\-4\}\), competitive with MV@10 \(−2\.0\-2\.0pp,p=0\.053p=0\.053\)\. Notably, atk=3k\{=\}3\(same generation cost as MV@3\), the probe achieves 70\.6% vs\. MV@3’s 69\.2% \(\+1\.4\+1\.4pp;p=0\.17p=0\.17\), suggesting complementary signal at smallNN; atk=5k\{=\}5, BoN \(71\.3%\) trails MV@5 \(71\.8%\) by only 0\.5pp\. The resulting hierarchy—oracle \(94\.2%,T=0\.8T\{=\}0\.8, 10 passes\)\>\>MV@10 \(72\.7%\)≈\\approxbest\-of\-NN\(70\.7%\)≫\\ggsteering \(Δ≤\+1\.5\\Delta\\leq\+1\.5pp,T=0\.1T\{=\}0\.1, single\-pass\)≈\\approxprompt baselines \(Δ≤\+0\.6\\Delta\\leq\+0\.6pp\)—shows theOTproblem is solvable at10×10\\timescost but not via single\-pass interventions\.

#### Prompt\-based baselines\.

We test ten prompt variants \(n=1,273n=1\{,\}273,T=0\.1T\{=\}0\.1; Table[4](https://arxiv.org/html/2605.05715#S6.T4)\), including few\-shot prompting \(k∈\{1,3,5\}k\\in\\\{1,3,5\\\}; answer\-only format without reasoning chains\), self\-refinement\(iterative self\-critique within a single generation; Madaan et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib43)\), and verbalized confidence\. None reliably improves: removing explicit CoT instruction achieves\+0\.6\+0\.6pp \(p=0\.68p=0\.68\), 3\-shot achieves\+0\.5\+0\.5pp \(p=0\.74p=0\.74\), and self\-refinement*significantly degrades*accuracy to 63\.0% \(Δ=−5\.9\\Delta=\-5\.9pp vs\. its paired baseline of 68\.9%, McNemarp<0\.001p<0\.001; 208 questions damaged, 133 corrected\)\. Verbalized confidence similarly degrades to 63\.2% \(Δ=−5\.7\\Delta=\-5\.7pp,p<0\.001p<0\.001\), though the model’s self\-reported confidence retains discriminative value for abstention \(AUROC = 0\.635; Figure[4](https://arxiv.org/html/2605.05715#S5.F4)\)\. Short\-generation catastrophically fails \(2\.5%, 90% parse failure due to conflicting length constraint and CoT instruction\)\. These results show thatOTpersists across tested prompt variants, extending the classification\-correction gap beyond activation steering to input\-level modifications\. Untested strategies \(MedPrompt, budget forcing\) remain open\.

Table 4:Prompt baselines \(n=1,273n=1\{,\}273, temperature 0\.1\)\. All values in %\. Zero\-shot variants share a baseline \(67\.2%\); few\-shot variants use a separate independent run \(baseline 65\.4%\) due to different prompt construction\. Self\-refinement and verbalized confidence use a dedicated paired baseline \(68\.9%\) with identical tokenization \(add\_special\_tokens=False\)\. AllΔ\\Deltavalues are McNemar\-tested against the run\-matched baseline\.
#### Selective abstention\.

The same structure that fails to support correction enables post\-generation reliability estimation\(Geifman and El\-Yaniv,[2017](https://arxiv.org/html/2605.05715#bib.bib21)\)\. A correctness probe \(layer 21, selected via held\-out AUROC from a 32\-layer scan on 70/30 training split\) achieves held\-out split AUROC = 0\.716 on training traces \(Figure[4](https://arxiv.org/html/2605.05715#S5.F4)\), with held\-out test\-set AUROC = 0\.610 \(95% CI: \[0\.577, 0\.642\];n=1,273n=1\{,\}273\)\. The generalization gap \(0\.716→\\to0\.610\) primarily reflects distributional shift between training traces \(balanced, temperature 0\.8\) and test\-set evaluation \(natural prevalence, temperature 0\.1\); evidence for this includes: \(a\) 5\-fold GroupKFold CV on training data yields AUROC = 0\.727 \(higher than the held\-out estimate\), and \(b\) the gap aligns with the cross\-temperature vector cosine of 0\.41 \(Appendix[J](https://arxiv.org/html/2605.05715#A10)\)\. Despite this shift, the held\-out probe exceeds all five tested single\-forward\-pass uncertainty baselines on both training traces \(AUROC≤\\leq0\.530; Table[6](https://arxiv.org/html/2605.05715#A7.T6)\) and the test set \(best baseline = 0\.569;Δ\\DeltaAUROC = 0\.041,p=0\.009p=0\.009, paired bootstrap; Holm\-corrected rank\-1 threshold = 0\.010; Table[7](https://arxiv.org/html/2605.05715#A7.T7)\)\. Verbalized confidence—where the model self\-reports a confidence percentage via a modified prompt—achieves comparable AUROC \(0\.635, 97\.4% parseable\) but requires prompt modification that alters the generation itself; the probe operates post\-hoc on unmodified outputs\. At 60% coverage, selective abstention yields 72\.3% accuracy \(\+5\.5pp\) on the test set; at 70% coverage, gains are comparable \(\+5\.7pp with lower abstention cost\)\. On balanced training data, gains are \+14\.0pp at 60% coverage \(relative to the 50% balanced baseline\)\. Proper 5\-fold stratified CV on training data confirms probe quality \(best layer L14: AUROC = 0\.730±\\pm0\.006; L21: AUROC = 0\.721±\\pm0\.005\)\. Self\-consistency achieves higher AUROC \(0\.804\) but at10×10\\timescost; the probe provides a cost\-effective single\-pass alternative\. Unlike process reward models\(Lightman et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib41)\), it requires no per\-step annotation\. The contribution is demonstrating that decodable structure has operational value for reliability estimation even when the tested steering family cannot exploit it for correction\.

## 7Conclusion

We demonstrate an empirical classification\-correction gap:OTis linearly decodable \(p≈10−16p\\approx 10^\{\-16\}\), robust across confounds with evidence on one additional architecture \(Qwen2\.5\-7B, 9 configs, allΔ<1\\Delta<1pp\) and cross\-domain steering failure \(MMLU\-STEM,n=300n=300: targetedΔ=0\\Delta=0pp, uniformΔ=−7\.0\\Delta=\-7\.0pp\), yet five families of fixed linear steering \(29 configurations\) and ten prompt baselines produceΔ≤0\\Delta\\leq 0while non\-targeted shared\-direction steering and concept erasure damage accuracy \(consistent with direction overlap with task computation\)\. The same decodable structure supports selective abstention \(held\-out test AUROC = 0\.610, exceeding all five tested uncertainty baselines;p=0\.009p=0\.009\) and best\-of\-NNprobe selection competitive with majority vote \(BoN@3 exceeds MV@3 by \+1\.4pp at identical generation cost, thoughp=0\.17p=0\.17\), demonstrating operational value for reliability estimation even when the tested steering family cannot exploit it for correction\. Four candidate pre\-intervention diagnostics \(based on a two\-point comparison with a refusal control, not validated as predictors\) are consistent with this failure pattern\. A learned non\-linear MLP intervention provides preliminary evidence that the gap is partially closeable \(Δ=\+2\.8\\Delta=\+2\.8pp,p=0\.025p=0\.025; Appendix[D](https://arxiv.org/html/2605.05715#A4)\), suggesting the entanglement is partial rather than absolute\. Whether more sophisticated learned interventions—Distributed Alignment Search, representation finetuning, or instance\-adaptive methods like K\-CAST—can fully close the gap remains the central open question; our correct/incorrect supervision pairs provide the requisite training signal for these methods\.

## Limitations

#### Scale, scope, and generalizability\.

We evaluate two 7–8B models on MedQA; at larger scales, higher baseline accuracy likely reducesOTprevalence \(though superposition may intensify with more features competing for shared directions\)\. Cross\-architecture steering on Qwen uses nine configurations spanning layers 5–18 and amplitudesα∈\[0\.5,3\.0\]\\alpha\\in\[0\.5,3\.0\]; the absence of Qwen concept\-erasure experiments leaves open whether Qwen’s distinct geometry \(specificity 0\.414 vs\. Llama’s 0\.119\) reflects a qualitatively different encoding\. MMLU\-STEM confirms both decodability transfer and steering failure \(n=300n=300, same pattern as MedQA\)\. The 60% correct\-rate threshold is empirically motivated \(Jaccard≥\\geq0\.81 under 50–70% sweeps\) but the MV@10 correction rate of 100% is definitional\. Prompt baselines do not include MedPrompt or budget\-forcing strategies; the no\-CoT condition does not enforce short generation, so implicit reasoning may persist\.

#### Intervention scope\.

The primary negative result covers*fixed*\-direction residual\-stream linear interventions: five method families \(29 additive configurations\) plus mean\-difference concept erasure \(rank\-1 projection and a whitened variant; see footnote in Section[5](https://arxiv.org/html/2605.05715#S5)\)\. The full oblique LEACE projector ofBelrose et al\. \([2023](https://arxiv.org/html/2605.05715#bib.bib7)\)\(Σw\+1/2​P​Σw−1/2\\Sigma\_\{w\}^\{\+1/2\}P\\Sigma\_\{w\}^\{\-1/2\}\) remains untested\. A probe\-gated dynamic variant was tested onn=300n=300\(Appendix[I](https://arxiv.org/html/2605.05715#A9)\) but produced no significant improvement \(allp\>0\.3p\>0\.3\)\. A learned non\-linear MLP \(4096→\\to64→\\to4096 residual adapter; Appendix[D](https://arxiv.org/html/2605.05715#A4)\) achievesΔ=\+2\.8\\Delta=\+2\.8pp \(p=0\.025p=0\.025\), providing preliminary evidence that non\-linear learned interventions can partially close the gap—though the damage rate remains high \(104/1273 = 8\.2%\) and the gain is modest relative to the 50\.6% centroid\-distance reduction, consistent with only partial disentanglement of the shared subspace\. Full DAS\(Geiger et al\.,[2024](https://arxiv.org/html/2605.05715#bib.bib22)\), ReFT\(Wu et al\.,[2024](https://arxiv.org/html/2605.05715#bib.bib74)\), and instance\-adaptive methods \(K\-CAST\(Valentino et al\.,[2026](https://arxiv.org/html/2605.05715#bib.bib65)\)\) remain untested; our correct/incorrect supervision pairs provide the requisite training signal\. Other untested families include path patching\(Goldowsky\-Dill et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib23)\)and parameter\-efficient fine\-tuning\(Hu et al\.,[2022](https://arxiv.org/html/2605.05715#bib.bib30)\)\. Qwen concept erasure experiments were not conducted\.

#### Temperature mismatch and annotation validity\.

Annotation usesT=0\.8T\{=\}0\.8, steering evaluationT=0\.1T\{=\}0\.1\. Same\-temperature steering produces*larger*damage \(−6\.7\-6\.7pp,p=0\.031p=0\.031; Appendix[K](https://arxiv.org/html/2605.05715#A11)\), directly falsifying temperature mismatch as an explanation for the steering null; the probe’s generalization gap \(0\.716→\\to0\.610\) reflects this distributional shift \(cross\-temperature vector cosine = 0\.41; Appendix[J](https://arxiv.org/html/2605.05715#A10)\)\.KD/RCBannotation shows moderate agreement \(κ=0\.30\\kappa=0\.30\) from two 4th\-year clinical students; the core binaryOTfinding bypasses this boundary and is robust to 37% simulated label noise\.

## Ethics Statement

This research analyzes model failures on a public benchmark \(MedQA\) with no patient data or human subjects\. LLM annotations use Claude models via the Anthropic API\. Llama\-3\.1 and Qwen2\.5 are used under their open\-weight licenses\. The selective abstention result may inform research on post\-generation reliability estimation under benchmark conditions, though it is not sufficient as a standalone clinical safeguard\.

## Reproducibility

All experiments use publicly available models \(Llama\-3\.1\-8B\-Instruct, Qwen2\.5\-7B\-Instruct\) and data \(MedQA, MMLU\-STEM\)\. Steering and evaluation experiments were run on NVIDIA A10G GPUs \(24GB\); full test\-set evaluation \(n=1,273n=1\{,\}273\) takes approximately 4 hours per configuration\. Generation usesdo\_sample=TruewithT=0\.1T\{=\}0\.1, introducing run\-to\-run variability \(baseline range: 64\.0–67\.4%\); all deltas are computed within\-run against paired baselines\. Appendix[A](https://arxiv.org/html/2605.05715#A1)provides annotation prompts, probe training hyperparameters, PCA settings, decoding parameters, and data split details\. Code and processed annotations will be released upon publication\.

## References

- Alain and Bengio \(2017\)Guillaume Alain and Yoshua Bengio\. 2017\.[Understanding intermediate layers using linear classifier probes](https://arxiv.org/abs/1610.01644)\.*arXiv preprint arXiv:1610\.01644*\.
- Anthropic \(2024\)Anthropic\. 2024\.[The Claude 3 model family: Opus, Sonnet, Haiku](https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf)\.Technical report, Anthropic\.We use Claude Haiku 4\.5 \(claude\-haiku\-4\-5\-20251001\)\.
- Arditi et al\. \(2024\)Andy Arditi, Oscar Obber, Ajeya Shlegeris, and Nicholas Schiefer\. 2024\.[Refusal in language models is mediated by a single direction](https://arxiv.org/abs/2406.11717)\.*arXiv preprint arXiv:2406\.11717*\.
- Azaria and Mitchell \(2023\)Amos Azaria and Tom Mitchell\. 2023\.[The internal state of an LLM knows when it’s lying](https://arxiv.org/abs/2304.13734)\.In*Findings of the Association for Computational Linguistics: EMNLP 2023*, pages 967–976\.
- Basu et al\. \(2026\)Sanjay Basu, Sadiq Y\. Patel, Parth Sheth, Bhairavi Muralidharan, Namrata Elamaran, Aakriti Kinra, John Morgan, and Rajaie Batniji\. 2026\.[Interpretability without actionability: Mechanistic methods cannot correct language model errors despite near\-perfect internal representations](https://arxiv.org/abs/2603.18353)\.*arXiv preprint arXiv:2603\.18353*\.
- Belinkov \(2022\)Yonatan Belinkov\. 2022\.[Probing classifiers: Promises, shortcomings, and advances](https://doi.org/10.1162/coli_a_00422)\.*Computational Linguistics*, 48\(1\):207–219\.
- Belrose et al\. \(2023\)Nora Belrose, David Schneider\-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, and Stella Biderman\. 2023\.[LEACE: Perfect linear concept erasure in closed form](https://arxiv.org/abs/2306.03819)\.In*Advances in Neural Information Processing Systems*, volume 36\.
- Billa \(2026\)Jayadev Billa\. 2026\.[Predicting where steering vectors succeed](https://arxiv.org/abs/2604.15557)\.*arXiv preprint arXiv:2604\.15557*\.
- Braun et al\. \(2025\)Joschka Braun, Dmitrii Krasheninnikov, Usman Anwar, Robert Kirk, Daniel Tan, and David Scott Krueger\. 2025\.[A sober look at steering vectors for LLMs](https://arxiv.org/abs/2411.02827)\.*arXiv preprint*\.
- Burns et al\. \(2023\)Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt\. 2023\.[Discovering latent knowledge in language models without supervision](https://arxiv.org/abs/2212.03827)\.In*International Conference on Learning Representations*\.
- Chen et al\. \(2024\)Xingyu Chen, Jiahao Xu, Tian Liang, Zhiwei He, Jianhui Pang, Dian Yu, Linfeng Song, Qiuzhi Liu, Mengfei Zhou, Zhuosheng Zhang, Rui Wang, Zhaopeng Tu, Haitao Mi, and Dong Yu\. 2024\.[Do NOT think that much for 2\+3=? on the overthinking of o1\-like LLMs](https://arxiv.org/abs/2412.21187)\.*arXiv preprint arXiv:2412\.21187*\.
- Chuang et al\. \(2024\)Yung\-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James Glass, and Pengcheng He\. 2024\.[DoLa: Decoding by contrasting layers improves factuality in large language models](https://arxiv.org/abs/2309.03883)\.In*Proceedings of the 12th International Conference on Learning Representations \(ICLR\)*\.
- Cox et al\. \(2026\)Kyle Cox, Darius Kianersi, and Adria Garriga\-Alonso\. 2026\.[Decoding answers before chain\-of\-thought: Evidence from pre\-CoT probes and activation steering](https://arxiv.org/abs/2603.01437)\.*arXiv preprint arXiv:2603\.01437*\.
- Cunningham et al\. \(2023\)Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey\. 2023\.[Sparse autoencoders find highly interpretable features in language models](https://arxiv.org/abs/2309.08600)\.*arXiv preprint arXiv:2309\.08600*\.
- Da Silva et al\. \(2025\)Patrick Queiroz Da Silva, Hari Sethuraman, Dheeraj Rajagopal, Hannaneh Hajishirzi, and Sachin Kumar\. 2025\.[Steering off course: Reliability challenges in steering language models](https://arxiv.org/abs/2504.04635)\.*arXiv preprint arXiv:2504\.04635*\.
- Deng et al\. \(2025\)Zhihong Deng, Jing Jiang, Guodong Long, and Chengqi Zhang\. 2025\.[Rethinking the reliability of representation engineering in large language models](https://openreview.net/forum?id=sYJQEgkkaI)\.*arXiv preprint arXiv:2409\.15726*\.
- Elazar et al\. \(2021\)Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg\. 2021\.[Amnesic probing: Behavioral explanation with amnesic counterfactuals](https://doi.org/10.1162/tacl_a_00359)\.*Transactions of the Association for Computational Linguistics*, 9:160–175\.
- Elhage et al\. \(2022\)Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield\-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, and 1 others\. 2022\.[Toy models of superposition](https://transformer-circuits.pub/2022/toy_model/index.html)\.*Transformer Circuits Thread*\.
- Elhage et al\. \(2021\)Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, and 1 others\. 2021\.[A mathematical framework for transformer circuits](https://transformer-circuits.pub/2021/framework/index.html)\.*Transformer Circuits Thread*\.
- Gao et al\. \(2026\)Lang Gao, Jinghui Zhang, Wei Liu, Fengxian Ji, Chenxi Wang, Zirui Song, Akash Ghosh, Youssef Mohamed, Preslav Nakov, and Xiuying Chen\. 2026\.[The cylindrical representation hypothesis for language model steering](https://arxiv.org/abs/2605.01844)\.*arXiv preprint arXiv:2605\.01844*\.
- Geifman and El\-Yaniv \(2017\)Yonatan Geifman and Ran El\-Yaniv\. 2017\.[Selective classification for deep neural networks](https://arxiv.org/abs/1705.08500)\.In*Advances in Neural Information Processing Systems*, volume 30\.
- Geiger et al\. \(2024\)Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah D\. Goodman\. 2024\.[Finding alignments between interpretable causal variables and distributed neural representations](https://arxiv.org/abs/2303.02536)\.*arXiv preprint arXiv:2303\.02536*\.
- Goldowsky\-Dill et al\. \(2023\)Nicholas Goldowsky\-Dill, Chris MacLeod, Lucas Sato, and Aryaman Arora\. 2023\.[Localizing model behavior with path patching](https://arxiv.org/abs/2304.05969)\.*arXiv preprint arXiv:2304\.05969*\.
- Grattafiori et al\. \(2024\)Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, and 1 others\. 2024\.[The Llama 3 herd of models](https://arxiv.org/abs/2407.21783)\.*arXiv preprint arXiv:2407\.21783*\.
- Guo et al\. \(2017\)Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger\. 2017\.[On calibration of modern neural networks](https://arxiv.org/abs/1706.04599)\.In*International Conference on Machine Learning*, pages 1321–1330\.
- Hase et al\. \(2023\)Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun\. 2023\.[Does localization inform editing? surprising differences in causality\-based localization vs\. knowledge editing in language models](https://arxiv.org/abs/2301.04213)\.In*Advances in Neural Information Processing Systems*, volume 36\.
- Hendrycks et al\. \(2021\)Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt\. 2021\.[Measuring massive multitask language understanding](https://arxiv.org/abs/2009.03300)\.In*International Conference on Learning Representations*\.
- Hernandez et al\. \(2024\)Evan Hernandez, Arnab Sen Sharma, Tal Haklay, Kevin Meng, Martin Wattenberg, Jacob Andreas, Yonatan Belinkov, and David Bau\. 2024\.[Linearity of relation decoding in transformer language models](https://arxiv.org/abs/2308.09124)\.In*International Conference on Machine Learning*\.
- Hewitt and Liang \(2019\)John Hewitt and Percy Liang\. 2019\.[Designing and interpreting probes with control tasks](https://doi.org/10.18653/v1/D19-1275)\.In*Proceedings of EMNLP*, pages 2733–2743\.
- Hu et al\. \(2022\)Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen\-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen\. 2022\.[LoRA: Low\-rank adaptation of large language models](https://arxiv.org/abs/2106.09685)\.In*International Conference on Learning Representations*\.
- Huang et al\. \(2024\)Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou\. 2024\.[Large language models cannot self\-correct reasoning yet](https://arxiv.org/abs/2310.01798)\.In*Proceedings of the 12th International Conference on Learning Representations \(ICLR\)*\.
- Jafari et al\. \(2026\)Amir Jafari, Aryaman Gangopadhyay, Jeffrey Long, and Ekdeep Singh Lubana\. 2026\.[Mechanistic indicators of steering effectiveness in large language models](https://arxiv.org/abs/2602.01716)\.*arXiv preprint arXiv:2602\.01716*\.
- Jin et al\. \(2021\)Di Jin, Eileen Pan, Nassim Oufattole, Wei\-Hung Weng, Hanyi Fang, and Peter Szolovits\. 2021\.[What disease does this patient have? a large\-scale open domain question answering dataset from medical exams](https://doi.org/10.3390/app11146421)\.*Applied Sciences*, 11\(14\):6421\.
- Jorgensen et al\. \(2023\)Ole Jorgensen, Dylan Cope, Nandi Schoots, and Murray Shanahan\. 2023\.[Improving activation steering in language models with mean\-centring](https://arxiv.org/abs/2312.03813)\.*arXiv preprint arXiv:2312\.03813*\.
- Kadavath et al\. \(2022\)Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield\-Dodds, Nova DasSarma, Eli Tran\-Johnson, and 1 others\. 2022\.[Language models \(mostly\) know what they know](https://arxiv.org/abs/2207.05221)\.*arXiv preprint arXiv:2207\.05221*\.
- Kojima et al\. \(2022\)Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa\. 2022\.[Large language models are zero\-shot reasoners](https://arxiv.org/abs/2205.11916)\.In*Advances in Neural Information Processing Systems*, volume 35\.
- Kuhn et al\. \(2023\)Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar\. 2023\.[Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation](https://arxiv.org/abs/2302.09664)\.In*International Conference on Learning Representations*\.
- Lakens \(2017\)Daniel Lakens\. 2017\.[Equivalence tests: A practical primer for t tests, correlations, and meta\-analyses](https://doi.org/10.1177/1948550617697177)\.*Social Psychological and Personality Science*, 8\(4\):355–362\.
- Lanham et al\. \(2023\)Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, KamilL̇ukić, Karina Nguyen, Newton Schiefer, Catherine Olsson, Tom Henighan, Andy Jones, Karina Ndousse, Oliver Bloom, Nelson Elhage, and 6 others\. 2023\.[Measuring faithfulness in chain\-of\-thought reasoning](https://arxiv.org/abs/2307.13702)\.*arXiv preprint arXiv:2307\.13702*\.
- Li et al\. \(2023\)Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg\. 2023\.[Inference\-time intervention: Eliciting truthful answers from a language model](https://doi.org/10.48550/arXiv.2306.03341)\.*Advances in Neural Information Processing Systems*, 36\.
- Lightman et al\. \(2023\)Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe\. 2023\.[Let’s verify step by step](https://arxiv.org/abs/2305.20050)\.*arXiv preprint arXiv:2305\.20050*\.
- Lin et al\. \(2022\)Stephanie Lin, Jacob Hilton, and Owain Evans\. 2022\.[Truthfulqa: Measuring how models mimic human falsehoods](https://doi.org/10.18653/v1/2022.acl-long.229)\.In*Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 3214–3252\.
- Madaan et al\. \(2023\)Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shravya Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark\. 2023\.[Self\-refine: Iterative refinement with self\-feedback](https://arxiv.org/abs/2303.17651)\.In*Advances in Neural Information Processing Systems*, volume 36\.
- Marks and Tegmark \(2024\)Samuel Marks and Max Tegmark\. 2024\.[The geometry of truth: Emergent linear structure in large language model representations of true/false datasets](https://arxiv.org/abs/2310.06824)\.In*Conference on Language Modeling*\.
- McKenzie et al\. \(2026\)Alex McKenzie, Keenan Pepper, Stijn Servaes, Martin Leitgab, Murat Cubuktepe, Mike Vaiana, Diogo de Lucena, Judd Rosenblatt, and Michael S\. A\. Graziano\. 2026\.[Endogenous resistance to activation steering in language models](https://arxiv.org/abs/2602.06941)\.*arXiv preprint arXiv:2602\.06941*\.
- Meng et al\. \(2022\)Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov\. 2022\.[Locating and editing factual associations in GPT](https://proceedings.neurips.cc/paper_files/paper/2022/hash/6f1d43d5a82a37e89b0665b33bf3a182-Abstract-Conference.html)\.In*Advances in Neural Information Processing Systems*, volume 35, pages 17359–17372\.
- Mishra et al\. \(2026\)Aayush Mishra, Daniel Khashabi, and Anqi Liu\. 2026\.[Steered LLM activations are non\-surjective](https://arxiv.org/abs/2604.09839)\.*arXiv preprint arXiv:2604\.09839*\.
- Muennighoff et al\. \(2025\)Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei\-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candela, and Dirk Groeneveld\. 2025\.[s1: Simple test\-time scaling](https://arxiv.org/abs/2501.19393)\.*arXiv preprint arXiv:2501\.19393*\.
- Nadaf \(2026\)Mohammed Suhail B Nadaf\. 2026\.[Steerable but not decodable: Function vectors operate beyond the logit lens](https://arxiv.org/abs/2604.02608)\.*arXiv preprint arXiv:2604\.02608*\.
- Nanda et al\. \(2023\)Neel Nanda, Andrew Lee, and Martin Wattenberg\. 2023\.[Emergent linear representations in world models of self\-supervised sequence models](https://arxiv.org/abs/2309.00941)\.In*Proceedings of the Annual Meeting of the Association for Computational Linguistics: BlackboxNLP Workshop*\.
- Nori et al\. \(2023\)Harsha Nori, Yin Tat Lee, Sheng Zhang, Dean Carignan, Richard Edgar, Nicolo Fusi, Nicholas King, Jonathan Larson, Yuanzhi Li, Weishung Liu, Renqian Luo, Scott Mayer McKinney, Robert Osazuwa Morrow, Tristan Nguyen, Hoifung Poon, Qiufeng Wei, and 1 others\. 2023\.[Can generalist foundation models outcompete special\-purpose tuning? Case study in medicine](https://arxiv.org/abs/2311.16452)\.*arXiv preprint arXiv:2311\.16452*\.
- Panickssery et al\. \(2023\)Nina Panickssery, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner\. 2023\.[Steering Llama 2 via contrastive activation addition](https://arxiv.org/abs/2312.06681)\.*arXiv preprint arXiv:2312\.06681*\.
- Park et al\. \(2024\)Kiho Park, Yo Joong Choe, and Victor Veitch\. 2024\.[The linear representation hypothesis and the geometry of large language models](https://arxiv.org/abs/2311.03658)\.In*International Conference on Machine Learning*\.
- Pedregosa et al\. \(2011\)Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, and 1 others\. 2011\.Scikit\-learn: Machine learning in Python\.*Journal of Machine Learning Research*, 12:2825–2830\.
- Ravfogel et al\. \(2020\)Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg\. 2020\.[Null it out: Guarding protected attributes by iterative nullspace projection](https://doi.org/10.18653/v1/2020.acl-main.647)\.In*Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics*, pages 7237–7256\. Association for Computational Linguistics\.
- Ravfogel et al\. \(2022\)Shauli Ravfogel, Michael Twiton, Yoav Goldberg, and Ryan Cotterell\. 2022\.[Linear adversarial concept erasure](https://proceedings.mlr.press/v162/ravfogel22a.html)\.In*Proceedings of ICML*, pages 18400–18421\.
- Sanyal et al\. \(2025\)Debdeep Sanyal, Manya Pandey, Dhruv Kumar, Saurabh Deshpande, and Murari Mandal\. 2025\.[Confidence is not competence: Probing vs steering LLM solvability beliefs](https://arxiv.org/abs/2510.24772)\.*arXiv preprint arXiv:2510\.24772*\.
- Sharma et al\. \(2024\)Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R Bowman, Newton Cheng, Esin Durmus, Zac Hatfield\-Dodds, Scott R Johnston, and 1 others\. 2024\.[Towards understanding sycophancy in language models](https://arxiv.org/abs/2310.13548)\.*arXiv preprint arXiv:2310\.13548*\.
- Singhal et al\. \(2023\)Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole\-Lewis, Stephen Pfohl, and 1 others\. 2023\.[Large language models encode clinical knowledge](https://doi.org/10.1038/s41586-023-06291-2)\.*Nature*, 620\(7972\):172–180\.
- Sprague et al\. \(2024\)Zayne Sprague, Fangcong Yin, Juan Diego Rodriguez, Dongwei Jiang, Manya Wadhwa, Prasann Singhal, Xinyu Zhao, Xi Ye, Kyle Mahowald, and Greg Durrett\. 2024\.[To CoT or not to CoT? chain\-of\-thought helps mainly on math and symbolic reasoning](https://arxiv.org/abs/2409.12183)\.*arXiv preprint arXiv:2409\.12183*\.
- Templeton et al\. \(2024\)Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Amezcua, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C Daniel Freeman, Theodore R Sumers, Edward Rees, Joshua Batson, Adam Jermyn, and 3 others\. 2024\.[Scaling monosemanticity: Extracting interpretable features from Claude 3 Sonnet](https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html)\.*Anthropic Technical Report*\.
- Todd et al\. \(2024\)Eric Todd, Millicent L Li, Arnab Sen Sharma, Aaron Mueller, Byron C Wallace, and David Bau\. 2024\.[Function vectors in large language models](https://arxiv.org/abs/2310.15213)\.In*International Conference on Learning Representations*\.
- Turner et al\. \(2024\)Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J\. Vazquez, Ulisse Mini, and Monte MacDiarmid\. 2024\.[Steering language models with activation engineering](https://arxiv.org/abs/2308.10248)\.*arXiv preprint arXiv:2308\.10248*\.
- Turpin et al\. \(2023\)Miles Turpin, Julian Michael, Ethan Perez, and Samuel R\. Bowman\. 2023\.[Language models don’t always say what they think: Unfaithful explanations in chain\-of\-thought prompting](https://arxiv.org/abs/2305.04388)\.In*Advances in Neural Information Processing Systems*, volume 36\.
- Valentino et al\. \(2026\)Marco Valentino, Mokanarangan Thayaparan, Tom Sherborne, and André Freitas\. 2026\.[Mitigating content effects on reasoning in language models through fine\-grained activation steering](https://arxiv.org/abs/2505.12189)\.In*Proceedings of the AAAI Conference on Artificial Intelligence*\.
- Vig et al\. \(2020\)Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber\. 2020\.[Investigating gender bias in language models using causal mediation analysis](https://proceedings.neurips.cc/paper/2020/hash/92650b2e92217715fe312e6fa7b90d82-Abstract.html)\.In*Advances in Neural Information Processing Systems*, volume 33, pages 12388–12401\.
- Walker and Nowacki \(2011\)Esteban Walker and Amy S Nowacki\. 2011\.Understanding equivalence and noninferiority testing\.*Journal of General Internal Medicine*, 26\(2\):192–196\.
- Wang et al\. \(2023\)Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou\. 2023\.[Self\-consistency improves chain of thought reasoning in language models](https://arxiv.org/abs/2203.11171)\.In*International Conference on Learning Representations*\.
- Wang et al\. \(2026\)Youjin Wang, Run Zhou, Rong Fu, Shuaishuai Cao, Hongwei Zeng, Jiaxuan Lu, Sicheng Fan, Jiaqiao Zhao, and Liangming Pan\. 2026\.[ASA: Training\-free representation engineering for tool\-calling agents](https://arxiv.org/abs/2602.04935)\.*arXiv preprint arXiv:2602\.04935*\.
- Wang et al\. \(2025\)Yue Wang, Qiuzhi Liu, Jiahao Xu, Tian Liang, Xingyu Chen, Zhiwei He, Linfeng Song, Dian Yu, Juntao Li, Zhuosheng Zhang, Rui Wang, Zhaopeng Tu, Haitao Mi, and Dong Yu\. 2025\.[Thoughts are all over the place: On the underthinking of o1\-like LLMs](https://arxiv.org/abs/2501.18585)\.*arXiv preprint arXiv:2501\.18585*\.
- Wei et al\. \(2022\)Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou\. 2022\.[Chain\-of\-thought prompting elicits reasoning in large language models](https://arxiv.org/abs/2201.11903)\.In*Advances in Neural Information Processing Systems*, volume 35\.
- Wollschläger et al\. \(2025\)Tom Wollschläger, Jannes Elstner, Simon Geisler, Vincent Cohen\-Addad, Stephan Günnemann, and Johannes Gasteiger\. 2025\.[The geometry of refusal in large language models: Concept cones and representational independence](https://arxiv.org/abs/2502.17420)\.*arXiv preprint arXiv:2502\.17420*\.
- Wu et al\. \(2025\)Zhengxuan Wu, Aryaman Arora, Atticus Geiger, Zheng Wang, Jing Huang, Dan Jurafsky, Christopher D\. Manning, and Christopher Potts\. 2025\.[AxBench: Steering LLMs? even simple baselines outperform sparse autoencoders](https://arxiv.org/abs/2501.17148)\.In*International Conference on Machine Learning*\.
- Wu et al\. \(2024\)Zhengxuan Wu, Aryaman Arora, Zhitao Wang, Atticus Geiger, Dan Jurafsky, Christopher D\. Manning, and Christopher Potts\. 2024\.[ReFT: Representation finetuning for language models](https://arxiv.org/abs/2404.03592)\.In*International Conference on Machine Learning*\.
- Xiong et al\. \(2024\)Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi\. 2024\.[Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs](https://arxiv.org/abs/2306.13063)\.In*International Conference on Learning Representations*\.
- Yang et al\. \(2024\)An Yang, Baosong Yang, Beichen Zhang, and 1 others\. 2024\.[Qwen2\.5 technical report](https://arxiv.org/abs/2412.15115)\.*arXiv preprint arXiv:2412\.15115*\.
- Zheng et al\. \(2023\)Lianmin Zheng, Wei\-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P\. Xing, and 1 others\. 2023\.[Judging LLM\-as\-a\-judge with MT\-Bench and Chatbot Arena](https://arxiv.org/abs/2306.05685)\.In*Advances in Neural Information Processing Systems*, volume 36\.
- Zou et al\. \(2023\)Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann\-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, and 2 others\. 2023\.[Representation engineering: A top\-down approach to AI transparency](https://arxiv.org/abs/2310.01405)\.*arXiv preprint arXiv:2310\.01405*\.
- Zur et al\. \(2025\)Amir Zur, Atticus Geiger, Ekdeep Singh Lubana, and Eric Bigelow\. 2025\.[Are language models aware of the road not taken? token\-level uncertainty and hidden state dynamics](https://arxiv.org/abs/2511.04527)\.*arXiv preprint arXiv:2511\.04527*\.

## Appendix AImplementation Details

### A\.1Data and Splits

We use MedQA\(Jin et al\.,[2021](https://arxiv.org/html/2605.05715#bib.bib33)\)with its standard train/test split: 10,178 training questions and 1,273 test questions\. For each question, we generate 10 traces using temperature 0\.8 sampling\. Hidden state extraction uses a stratified sample of 2,000 traces per failure mode plus 6,000 matched correct traces \(2,000 per mode’s question set\) from the training split, yielding 12,000 total traces\. All classification and probe training results use 5\-fold stratified cross\-validation \(random seed 42\)\. Steering evaluation uses the full 1,273\-question test set \(except probe\-guided experiments; see Section[5](https://arxiv.org/html/2605.05715#S5)\)\.

### A\.2Generation Parameters

All steering evaluation uses temperature 0\.1 withdo\_sample=Trueandmax\_new\_tokens=600\. The system prompt is:*“You are a medical expert\. Answer the following medical question\. Think through the problem step by step, then provide your final answer\.”*Answers are extracted via regex matching of letter choices \(A–E\)\.

### A\.3Hidden State Extraction

We extract last\-token hidden states at all layers \(32 for Llama, 28 for Qwen\) in fp32 precision\. We use last\-token representations following prior work on probing autoregressive models\(Burns et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib10); Li et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib40)\), as the final position aggregates information from the full reasoning chain via causal attention\.

### A\.4PCA and Probe Training

Hidden states are standardized \(zero mean, unit variance\) before applying PCA with 50 components\. At 50 components, PCA retains\>\>95% of variance across all layers; we verified that 20 and 100 components yield comparable classification accuracy \(±\\pm0\.5pp\)\. Linear probes use logistic regression with L2 regularization \(C=1\.0C=1\.0\), the LBFGS solver, and a maximum of 1,000 iterations, implemented via scikit\-learn\(Pedregosa et al\.,[2011](https://arxiv.org/html/2605.05715#bib.bib54)\)\.

### A\.5Annotation Pipeline

#### Phase 1 \(OT detection\)\.

For each question, if≥\\geq60% of 10 traces are correct, incorrect traces exceeding 200 tokens are labeledOT\. OT threshold sensitivity is reported in Section[4](https://arxiv.org/html/2605.05715#S4)\.

#### Phase 2 \(KD/RCB classification\)\.

Remaining incorrect traces are classified using Claude Haiku \(model:claude\-haiku\-4\-5\-20251001, temperature 0\.1, max tokens 300\) with the following prompt:

> *Analyze this incorrect medical reasoning trace\. The model answered WRONG\.* *Question:\{question\}/ Correct Answer:\{answer\}/ Model’s Reasoning:\{trace\}* *Classify the PRIMARY failure mode: 1\.KD: The model stated incorrect medical FACTS\. 2\.RCB: The facts are mostly correct, but the LOGIC connecting them is broken\. 3\.unclear: Cannot clearly determine\.* *Respond with JSON:\{“mode”: …, “confidence”: 0–1, “evidence”: \[…\]\}*

### A\.6Annotation Validation Details

Domain expert validation\.Two clinical medicine graduate students \(4th\-year, with clinical rotation experience\) independently annotated a stratified 500\-trace gold set, balanced across modes and enriched with boundary cases where the pipeline had low confidence\. Annotators were blinded to automated labels and to each other; they received the question, correct answer, model response, and definitions of each mode\. Expert\-automated agreement:OT94%,KD82%,RCB71%; expert\-expertκ=0\.61\\kappa=0\.61on the three\-way task\. As expected, theKD/RCBboundary is the primary source of disagreement; our strongest claims avoid depending on it\.

LLM cross\-validation\(Claude Opus 4\.6\):KDagreement 88\.1%,RCBagreement 44\.1% \(Cohen’sκ=0\.30\\kappa=0\.30,n=500n=500\)\.100\-trace three\-way comparison\(Haiku vs\. Sonnet vs\. Opus\): all three agree on 40\.4%; the two strongest models \(Sonnet\-Opus\) agree most \(68\.7%,κ=0\.28\\kappa=0\.28\), both independently reclassifying∼\\sim56% of Haiku\-RCBasKD\. The consistent asymmetric pattern—highKDagreement, lowRCBagreement—shows that LLM annotators systematically diverge at theKD/RCBboundary\.

Within\-question label consistency\.Per\-question annotation purity is high: across 10 traces per question,OTquestions have 100% purity \(all incorrect traces labeledOT\),RCBquestions 87%, andKDquestions 78%\. Of 3,912 questions with failure traces, 89\.8% have a single failure mode label; only 10\.2% mixKDandRCBacross traces\. Theκ=0\.30\\kappa=0\.30above reflects cross\-question boundary ambiguity \(where to draw theKD/RCBline\), not random within\-question noise\.

Geometric robustness to label noise\.Simulating noise matching the observed LLM disagreement rate: randomly flipping 37% ofKD/RCBlabels drops binary classification only from 66\.3% to 60\.7% \(chance = 50%\), and three\-way classification from 50\.5% to 47\.6% \(chance = 33\.3%\)\. This simulation uses i\.i\.d\. random flips; real annotator disagreement may be more systematic, which our simulation does not fully capture\. The binaryOT\-vs\-non\-OTclassification \(71\.6%,p≈10−16p\\approx 10^\{\-16\}\) bypasses theKD/RCBboundary entirely\.

### A\.7Diagnostic Metrics

The*specificity ratio*for modemmquantifies how much of the contrastive vector is mode\-specific vs\. shared across all modes\. We define the shared direction as the mean of all mode contrastive vectors \(stacked across all layers\):𝐯¯=1\|ℳ\|​∑m∈ℳ𝐯m\\bar\{\\mathbf\{v\}\}=\\frac\{1\}\{\|\\mathcal\{M\}\|\}\\sum\_\{m\\in\\mathcal\{M\}\}\\mathbf\{v\}\_\{m\}\. Then:specificitym=‖𝐯m−proj𝐯¯​𝐯m‖2/‖𝐯m‖2\\text\{specificity\}\_\{m\}=\\\|\\mathbf\{v\}\_\{m\}\-\\text\{proj\}\_\{\\bar\{\\mathbf\{v\}\}\}\\mathbf\{v\}\_\{m\}\\\|^\{2\}/\\\|\\mathbf\{v\}\_\{m\}\\\|^\{2\}\. A value of 0\.119 means only 12% of vector variance is mode\-specific\. This is a custom metric introduced in this work\.

#### Permutation null for specificity\.

To confirm the low specificity is not an artifact of set\-subset structure \(all modes are subsets of incorrect traces\), we permute category labels among the 6,000 incorrect traces 10,000 times, recomputing the specificity ratio each time\. The null distribution has mean 0\.370 \(SD = 0\.155\); the observed value of 0\.119 falls below all 10,000 permutations \(p<0\.0001p<0\.0001\)\. This confirms thatOThas anomalously*high*alignment with the shared incorrect\-vs\-correct axis—not merely the expected behavior of any incorrect subset\. Bootstrap resampling \(2,000 draws\) yields a 95% CI of \[0\.015, 0\.055\] on the measurement itself \(under the binary framing; the multi\-mode framing yields 0\.119\)\.

#### Refusal direction specificity\.

As a positive\-control calibration, we compute the specificity ratio for the refusal direction \(extracted from 20 benign vs\. 20 harmful prompts; Appendix[C](https://arxiv.org/html/2605.05715#A3)\) under the same three\-mode shared decomposition\. The refusal specificity is 0\.999 \(1,000\-bootstrap 95% CI for OT: \[0\.075, 0\.119\]\), meaning\>\>99% of the refusal direction’s variance is orthogonal to the failure\-mode shared axis\. Cosine similarities confirm geometric independence:cos⁡\(𝐝refusal,𝐝OT\)=−0\.008\\cos\(\\mathbf\{d\}\_\{\\text\{refusal\}\},\\mathbf\{d\}\_\{\\text\{OT\}\}\)=\-0\.008,cos⁡\(𝐝refusal,𝐯¯shared\)=0\.014\\cos\(\\mathbf\{d\}\_\{\\text\{refusal\}\},\\bar\{\\mathbf\{v\}\}\_\{\\text\{shared\}\}\)=0\.014\. Under a four\-mode decomposition \(adding refusal as a fourth direction\), OT specificity rises to 0\.335 and refusal is 0\.851—the ordering is robust to decomposition choice\. The4\.7×4\.7\\timesspecificity gap between refusal \(steerable,p=0\.008p=0\.008\) andOT\(not steerable, 29 configs\) provides quantitative evidence that the specificity ratio is predictive of steerability in our setting\.

The*spread ratio*for modemmis the ratio of within\-class standard deviation to inter\-centroid distance:spreadm=σm/‖𝝁correct−𝝁m‖\\text\{spread\}\_\{m\}=\\sigma\_\{m\}/\\\|\\boldsymbol\{\\mu\}\_\{\\text\{correct\}\}\-\\boldsymbol\{\\mu\}\_\{m\}\\\|, computed in PCA\-50 space at the peak layer\. Values\>\>1 indicate that within\-class variance exceeds the centroid gap\. The*signal\-to\-noise ratio*\(SNR\) measures how much of the centroid gap a steering vector closes per unit perturbation:SNRm=\|⟨𝐯m,𝝁correct−𝝁m⟩\|/‖𝝁correct−𝝁m‖2\\text\{SNR\}\_\{m\}=\|\\langle\\mathbf\{v\}\_\{m\},\\boldsymbol\{\\mu\}\_\{\\text\{correct\}\}\-\\boldsymbol\{\\mu\}\_\{m\}\\rangle\|/\\\|\\boldsymbol\{\\mu\}\_\{\\text\{correct\}\}\-\\boldsymbol\{\\mu\}\_\{m\}\\\|^\{2\}\. Low SNR \(≪1\\ll 1\) means the steering vector is nearly orthogonal to the correction direction in activation space\.

### A\.8Steering Hyperparameters

Contrastive steering usesα∈\{0\.5,1\.0,1\.5,2\.0,3\.0\}\\alpha\\in\\\{0\.5,1\.0,1\.5,2\.0,3\.0\\\}; probe\-guided steering usesα∈\{0\.5,1\.0,1\.5\}\\alpha\\in\\\{0\.5,1\.0,1\.5\\\}\. Multi\-layer steering usesα=1\.5\\alpha=1\.5at 1, 3, or 5 layers centered on the peak layer\. Steering vectors are applied via forward hooks registered on the full decoder layer module \(model\.layers\[l\]\), firing on the post\-layer residual stream \(after both attention and MLP sub\-blocks, inclusive of residual connections\) and addingα⋅𝐯\\alpha\\cdot\\mathbf\{v\}during generation\. Confidence\-gated variants use a detection margin threshold of 0\.1\. Rank\-kksubspace steering usesα=1\.5\\alpha=1\.5at the peak layer \(layer 16\)\. Subspace bases are constructed via SVD of stacked direction matrices: the probe matrix is \(4, 4096\) \(3 mode\-specific \+ 1 correctness probe\), the combined matrix is \(7, 4096\) \(probe \+ contrastive\)\. For each rankk∈\{1,3,5\}k\\in\\\{1,3,5\\\}, we take the top\-kkright singular vectors as an orthonormal basis and project the correction vector onto this subspace\. Each projected vector is applied uniformly to all 1,273 test questions; no runtime mode detection is performed\.

## Appendix BTaxonomy Robustness Controls

We verify the 71\.6% binaryOT\-vs\-non\-OTclassification against six potential confounds\.Random labelsyield 33\.2% \(chance\)\.Length regression: regressing out response length drops three\-way only 1\.1pp;OT\-vs\-non\-OTis unaffected \(71\.3%\), consistent with\>\>99% of non\-OTincorrect traces also exceeding 200 tokens\.Binary collapse:OTvs\. \{KD\+RCB\} achieves 71\.6% \(p≈10−16p\\approx 10^\{\-16\}\), independent ofKD/RCBannotation quality\.OT threshold sensitivity: sweeping correct\-rate \(0\.5–0\.7\) and length \(100–300\) thresholds, Jaccard≥\\geq0\.81 vs\. default\. Excluding borderlineOTquestions \(6/10 correct, 16%\) reduces balanced accuracy by only 1\.2pp, matching a same\-proportion random\-removal control, confirming minority\-class sample reduction rather than borderline\-specific content; even strict exclusion \(6–7/10, 37%\) retains 58\.3%\.Question\-disjoint validation: GroupKFold yields identical results \(<<1pp drop\), ruling out question\-identity leakage\.Prompt\-end probing:OTdetection at the question representation \(before generation\) achieves only 54\.4% balanced accuracy \(95% CI: \[53\.5%, 55\.2%\]; chance = 50%\), confirming a generation\-time regime\.

## Appendix CPositive Control: Infrastructure Validation

To verify that our steering infrastructure is functional, we test whether the identical code path \(same model, same forward hooks, same generation parameters\) can produce directional behavioral effects on a task where contrastive activation addition is known to work\.

#### Refusal steering \(primary control\)\.

We extract contrastive refusal vectors from 20 benign prompts \(always complied with\) and 20 clearly harmful prompts \(always refused\), computing𝐯comply\(l\)=\(𝐡¯comply\(l\)−𝐡¯refuse\(l\)\)/∥⋅∥\\mathbf\{v\}\_\{\\text\{comply\}\}^\{\(l\)\}=\(\\bar\{\\mathbf\{h\}\}\_\{\\text\{comply\}\}^\{\(l\)\}\-\\bar\{\\mathbf\{h\}\}\_\{\\text\{refuse\}\}^\{\(l\)\}\)/\\\|\\cdot\\\|at each layer—the same mean\-difference, per\-layer L2 normalization used for MedQA\. We then evaluate on 50 borderline prompts \(ambiguous safety context\) using steering layers 26–28 \(top\-3 by separation, capped below layer 29\) withα∈\{2,4,6\}\\alpha\\in\\\{2,4,6\\\}in both the compliance \(\+𝐯\+\\mathbf\{v\}\) and anti\-compliance \(−𝐯\-\\mathbf\{v\}\) directions\.

#### Results\.

Baseline refusal rate is 4/50 \(8%\)\. Anti\-compliance steering consistently induces additional refusals \(best: 9/50 = 18% at L27/α\\alpha=6\), while compliance steering produces no change from baseline \(3–5/50\)\. Critically, the effect isdirectionally specific: all 5 prompts that flip from compliance to refusal under anti\-compliance steering remain compliant under compliance steering\. Across all 50 prompts, anti\-compliance produces strictly more refusal than compliance in 7 prompts, with 0 prompts showing the reverse pattern \(sign testp=0\.008p=0\.008, one\-sided\)\. This confirms that \(1\) forward hooks modify model behavior during generation, \(2\) contrastive vectors encode directionally meaningful signals, and \(3\) the effect is not a non\-directional perturbation artifact\.

#### TruthfulQA \(secondary evidence\)\.

We additionally apply contrastive truthfulness vectors on TruthfulQA MC1\(Lin et al\.,[2022](https://arxiv.org/html/2605.05715#bib.bib42)\)\(n=417n=417, baseline = 48\.9%\) using the identical code path\. Pro\-truthfulness steering at the best\-separation layer \(layer 31\) producesΔ∈\[\+0\.3,\+3\.1\]\\Delta\\in\[\+0\.3,\+3\.1\]pp acrossα∈\{1\.0,1\.5,3\.0\}\\alpha\\in\\\{1\.0,1\.5,3\.0\\\}, none reaching significance \(allp\>0\.34p\>0\.34\)\. This is consistent with the refusal result: mean\-difference CAA can shift binary behavioral gates \(refusal/compliance\) but does not improve multi\-choice reasoning accuracy—the same pattern observed on MedQA\.

#### Geometric comparison\.

Computing the specificity ratio for the refusal direction under the same three\-mode shared decomposition used for failure modes \(Section[4](https://arxiv.org/html/2605.05715#S4)\), we find specificityrefusal\{\}\_\{\\text\{refusal\}\}= 0\.999 \(bootstrap 95% CI for OT: \[0\.075, 0\.119\]\)\. The refusal direction is near\-orthogonal to both the OT direction \(cos=−0\.008\\cos=\-0\.008\) and the shared failure\-mode axis \(cos=0\.014\\cos=0\.014\), confirming geometric independence\. Under a four\-mode decomposition \(including refusal\), OT specificity rises to 0\.335 while refusal remains at 0\.851—the ordering is preserved regardless of decomposition choice\. This provides a quantitative geometric basis for the behavioral contrast: refusal occupies a dedicated, non\-overlapping direction in activation space \(specificity≈1\\approx 1\), while the OT signal is4\.7×4\.7\\timesless specific \(specificity≤0\.21\\leq 0\.21\), sharing 79–88% of its variance with task\-critical computation\.

#### Interpretation\.

The refusal control validates the infrastructure while illustrating why the classification\-correction gap arises\. Refusal is a binary behavioral gate decided at the first generation token; a single direction in activation space causally controls this gate\. MC accuracy requires sustained correct reasoning across many tokens, where a single additive perturbation is insufficient to redirect an incorrect reasoning chain\. The gap is thus between*behavioral steering*\(where CAA succeeds\) and*reasoning correction*\(where it does not\)\. The geometric comparison above adds a quantitative dimension: high\-specificity directions \(refusal: 0\.999\) respond to CAA because perturbation along them does not interfere with other computations, while low\-specificity directions \(OT:≤\\leq0\.21\) are entangled with task\-relevant signals, causing collateral damage that neutralizes any corrective effect\. We acknowledge this is a lower\-bound control that validates the code path, not the difficulty of the target task; to our knowledge, no multi\-step reasoning task has been established as a CAA positive control\.

## Appendix DLearned Non\-Linear MLP Steering

To test whether the classification\-correction gap is specific to the fixed linear intervention family, we train a non\-linear residual MLP to steer hidden states\. The MLP uses the*same*correct/OTsupervision available to linear methods \(contrastive vectors, probe\-guided steering, and LEACE all use these labels\)—the only difference is functional capacity\.

#### Architecture and training\.

A two\-layer MLP with GELU activation and bottleneck \(4096→\\to64→\\to4096\) is trained as a residual adapter:𝐡′=𝐡\+fθ​\(𝐡\)\\mathbf\{h\}^\{\\prime\}=\\mathbf\{h\}\+f\_\{\\theta\}\(\\mathbf\{h\}\)\. Training minimizes L2 distance between transformedOThidden states and the correct\-class centroid, withλ=0\.01\\lambda=0\.01regularization penalizing perturbation norm on correct\-class inputs \(encouraging near\-identity on already\-correct states\)\. Training uses 3,200 hidden states \(80/20 split from the training\-set annotation pool; no test\-question overlap\), Adam optimizer with cosine LR schedule, 50 epochs\. Best validation loss at epoch 49 \(1\.82\); train and validation losses track closely throughout \(no overfitting divergence\)\.

#### Inference\.

The trained MLP is registered as a forward hook on layer 16 \(matching all linear experiments\)\. A separate baseline is generated from a fresh model load without the hook\. Both conditions use identical generation parameters \(T=0\.1T=0\.1,do\_sample=True,max\_new\_tokens=600\)\. AtT=0\.1T=0\.1, generation is near\-deterministic \(\>\>99\.9% top\-token probability at typical logit gaps\), ensuring valid paired comparison\.

#### Results\.

On the full test set \(n=1,273n=1\{,\}273\):

- •Baseline accuracy: 66\.8% \(95% CI: \[64\.2, 69\.4\]%\)
- •MLP\-steered accuracy: 69\.7% \(95% CI: \[67\.1, 72\.1\]%\)
- •Δ=\+2\.8\\Delta=\+2\.8pp; McNemarp=0\.025p=0\.025\(two\-sided\)
- •Corrections: 140; Damages: 104; ratio 1\.35:1
- •TOST within±2\.5\\pm 2\.5pp:p=0\.57p=0\.57\(not equivalent to zero\)
- •Mean perturbation norm: 3\.07 \(comparable to linearα=1\.5\\alpha=1\.5\)
- •Centroid distance reduction: 50\.6% \(3\.64→\\to1\.80\)

#### Interpretation\.

The MLP achieves statistically significant improvement where all 29 fixed linear configurations produceΔ≈0\\Delta\\approx 0\. This is a separate hypothesis family \(learned non\-linear\) from the fixed linear sweep, so no multiple\-testing correction across families applies\. The result provides evidence that non\-linear transforms can partially disentangle the shared subspace: the 64\-dimensional bottleneck can implement state\-dependent, direction\-conditional corrections that avoid the task\-critical directions a fixed additive vector inevitably perturbs\. However, the gain is modest \(\+2\.8pp\) relative to the large geometric improvement \(50\.6% centroid\-distance reduction\), and the damage rate remains high \(8\.2%\), indicating only partial disentanglement\. The MLP changes 25% of answers \(vs\. 38–56% for linear steering\), with no systematic bias toward any answer letter—consistent with targeted rather than distributional perturbation\.

#### Relation to the gap\.

The MLP result refines the paper’s thesis: the classification\-correction gap is specific to*fixed linear*interventions;*learned non\-linear*methods can partially exploit the decodable signal, consistent with the entanglement being a matter of degree rather than absolute\. This supports the paper’s existing framing that “learned interventions could succeed by learning to disentangle shared directions” \(Section[6](https://arxiv.org/html/2605.05715#S6)\) and is consistent with the structurally unreachable regime ofBilla \([2026](https://arxiv.org/html/2605.05715#bib.bib8)\)applying specifically to the linear family\.

## Appendix ERank\-kkSubspace Steering

MethodkkΔ\\Delta95% CITOST*Correctness\-uniform*Probe sub\.1−\-4\.6\[−\-7\.4,−\-1\.9\]—Probe sub\.3−\-3\.8\[−\-6\.6,−\-1\.0\]—Combined sub\.1−\-6\.2\[−\-9\.2,−\-3\.2\]—Combined sub\.3−\-3\.5\[−\-6\.4,−\-0\.7\]—Combined sub\.5−\-2\.6\[−\-5\.4,\+\+0\.2\]—*Mode\-specific \(all questions\)*KDprobe1−\-0\.5\[−\-3\.1,\+\+2\.0\]\.030KDprobe3\+\+1\.8\[−\-0\.8,\+\+4\.4\]\.183RCBprobe1−\-0\.5\[−\-3\.2,\+\+2\.2\]\.033RCBprobe3−\-0\.5\[−\-3\.2,\+\+2\.1\]\.034OTprobe1−\-0\.8\[−\-3\.4,\+\+1\.9\]\.051OTprobe3−\-0\.2\[−\-2\.9,\+\+2\.6\]\.020Table 5:Rank\-kksubspace steering \(n=1,273n=1\{,\}273\)\.Δ\\Deltain pp\. Correctness\-uniform vectors damage performance at all ranks\. Per\-mode vectors produceΔ≈0\\Delta\\approx 0\(TOSTp<\.05p<\.05= equivalent within±\\pm3pp\)\. Higher rank does not help\.We construct subspace bases via SVD of stacked direction matrices: \(1\) a probe matrix of shape \(4, 4096\), stacking 3 mode\-specific and 1 correctness probe weight vector; \(2\) a combined matrix \(7, 4096\), stacking probe \+ contrastive vectors\. For each rankk∈\{1,3,5\}k\\in\\\{1,3,5\\\}, we take the top\-kkright singular vectors as an orthonormal basis, project the correction vector onto this subspace, and apply the projected vector uniformly to all 1,273 test questions\. Results \(Table[5](https://arxiv.org/html/2605.05715#A5.T5)\) show that neither expanding rank nor changing the basis source closes the gap\.

## Appendix FProbe vs\. Contrastive Vector Analysis

![Refer to caption](https://arxiv.org/html/2605.05715v1/x5.png)Figure 5:Probe vs contrastive vector comparison\. \(a\) Pairwise cosine similarity: probe vectors are much less correlated \(0\.23–0\.54\) than contrastive vectors \(0\.79–0\.86\), indicating the probe optimization finds more discriminative directions\. \(b\) Mode specificity: probes achieve 5–6×\\timeshigher specificity than contrastive decomposition\.Per\-layer logistic regression probes \(PCA\-50 features\) reveal a strikingly different geometry from mean\-difference contrastive vectors \(Figure[5](https://arxiv.org/html/2605.05715#A6.F5)\)\. Probe weight vector pairs have average cosine 0\.23–0\.54, showing the probe finds more discriminative, mode\-specific directions than the contrastive approach\. Yet despite this geometric dissociation, both methods achieve comparable classification accuracy \(probes: 51\.2% at layer 15; contrastive: 51\.5% at layer 16\)\. This means multiple distinct directions in activation space support the same classification—further evidence that decodability does not identify a unique causal intervention target\.

## Appendix GUncertainty Baseline Comparison

Table 6:Uncertainty baseline comparison for selective abstention \(n=3,600n=3\{,\}600held\-out traces, balanced 50% correct / 50% incorrect\)\. All baselines are single\-forward\-pass metrics computed from the same hidden states\. The correctness probe outperforms the best uncertainty baseline \(hidden state norm\) byΔ\\DeltaAUROC = 0\.186 \(p<10−4p<10^\{\-4\}, paired bootstrap, 5,000 resamples\)\. AUROC is prevalence\-independent; the “\+pp” column shows gains relative to the 50% balanced baseline at 60% coverage\.Table 7:Test\-set uncertainty baselines \(n=1,273n=1\{,\}273test questions, temperature 0\.1 generation\)\. The correctness probe—trained on temperature 0\.8 training traces—outperforms the best uncertainty baseline \(Δ\\DeltaAUROC = 0\.041,p=0\.009p=0\.009, paired bootstrap\)\. Baselines are computed directly on test\-set logits and hidden states; the probe faces an additional generalization challenge \(cross\-temperature, cross\-question transfer\)\. CIs are 95% bootstrap \(5,000 resamples\)\.We focus on single\-forward\-pass baselines computed from the same hidden states for fair comparison; modern neural networks are known to be poorly calibrated\(Guo et al\.,[2017](https://arxiv.org/html/2605.05715#bib.bib25)\), motivating hidden\-state probes over raw confidence scores\. Multi\-sample methods such as semantic entropy\(Kuhn et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib37)\)require multiple generations, comparable in cost to self\-consistency\(Wang et al\.,[2023](https://arxiv.org/html/2605.05715#bib.bib68)\)\(AUROC = 0\.804\), which we report as the multi\-sample performance ceiling\. Verbalized confidence\(Xiong et al\.,[2024](https://arxiv.org/html/2605.05715#bib.bib75)\)andP​\(True\)P\(\\text\{True\}\)probing\(Kadavath et al\.,[2022](https://arxiv.org/html/2605.05715#bib.bib35)\)remain untested single\-pass alternatives\.

## Appendix HConcept Erasure Random Direction Control

To test whether the mean\-difference erasure damage \(full test set:Δ=−3\.6\\Delta=\-3\.6pp,p=0\.010p=0\.010,n=1,273n=1\{,\}273\) is specific to the failure\-mode direction or a generic artifact of rank\-1 projection, we compare theOTdirection against 10 random directions of equal rank on a 300\-question subset under identical conditions\. For each random direction𝐪rand∼𝒩​\(0,𝐈\)\\mathbf\{q\}\_\{\\text\{rand\}\}\\sim\\mathcal\{N\}\(0,\\mathbf\{I\}\)\(normalized\), we construct𝐏rand=𝐈−𝐪rand​𝐪randT\\mathbf\{P\}\_\{\\text\{rand\}\}=\\mathbf\{I\}\-\\mathbf\{q\}\_\{\\text\{rand\}\}\\mathbf\{q\}\_\{\\text\{rand\}\}^\{T\}and evaluate accuracy under the same generation protocol\.

Onn=300n=300questions, the OT direction producesΔOT=−3\.9\\Delta\_\{\\text\{OT\}\}=\-3\.9pp, while 10 random directions produceΔrand=\+0\.3±1\.8\\Delta\_\{\\text\{rand\}\}=\+0\.3\\pm 1\.8pp \(mean±\\pmSD; range−2\.0\-2\.0to\+3\.3\+3\.3pp\)\. The OT\-specific excess \(ΔOT−Δ¯rand=−4\.2\\Delta\_\{\\text\{OT\}\}\-\\bar\{\\Delta\}\_\{\\text\{rand\}\}=\-4\.2pp; OT falls below all 10 random directions\) confirms that the erasure effect is direction\-specific, not a generic consequence of removing any rank\-1 subspace\. Two interpretations are equally consistent: \(1\) the failure\-mode direction co\-occurs with \(shares subspace with\) task\-relevant computation at layer 16; or \(2\) the OT direction is causally required for correct reasoning, and erasure removes information needed for both failure\-mode encoding and successful computation\. No random direction produces damage comparable to theOTdirection \(worst random:−2\.0\-2\.0pp vs\. OT:−3\.9\-3\.9pp\)\.

## Appendix IProbe\-Gated Dynamic Steering

We test whether dynamic, token\-conditional steering can succeed where static interventions fail \(n=300n=300\)\. At each generation step, a correctness probe \(trained on last\-token hidden states at layer 16\) evaluatesP​\(correct\)P\(\\text\{correct\}\)from the current token’s hidden state\. Steering is applied only when the probe signals low confidence\.

Table 8:Probe\-gated dynamic steering \(n=300n=300\)\. All values in %; Corr\. and Dmg\. show counts \(not %\)\. Static contrastive steering significantly damages accuracy \(p=0\.018p=0\.018, McNemar\)\. No dynamic gating condition produces reliable improvement \(all McNemarp\>0\.3p\>0\.3\)\. The probe reportsP​\(correct\)<0\.5P\(\\text\{correct\}\)<0\.5for\>\>93% of intermediate tokens, resulting in high steering rates; the strict threshold \(P<0\.3P<0\.3\) reduces this to 61% but still yieldsΔ≈0\\Delta\\approx 0\.No dynamic condition reliably improves accuracy \(Table[8](https://arxiv.org/html/2605.05715#A9.T8); all McNemarp\>0\.3p\>0\.3\)\. Static contrastive steering is significantly harmful \(Δ=−8\.0\\Delta=\-8\.0pp,p=0\.018p=0\.018; 63 damages vs\. 39 corrections\)\. Dynamic gating reduces this damage—the best gating variants achieveΔ≈0\\Delta\\approx 0—but cannot push it into positive territory\. A contributing factor is that the probe, trained on last\-token hidden states, is poorly calibrated on intermediate generation tokens: it reportsP​\(correct\)<0\.5P\(\\text\{correct\}\)<0\.5for 93% of tokens under binary gating\. Even the conservative threshold \(P<0\.3P<0\.3\) steers 61% of tokens\. This highlights an additional challenge for dynamic steering: the probe’s distributional assumptions break down during autoregressive generation, and the gating mechanism partially degenerates toward static application\.

## Appendix JCross\-Temperature Vector Comparison

To directly test whether contrastive directions are preserved across the temperature gap \(training at 0\.8, evaluation at 0\.1\), we generate 5 traces at temperature 0\.1 for 397 training questions \(100OT\+ 100KD\+ 100RCB\+ 97 other\), yielding 1,111 correct and 869 incorrect traces\. We compute binary correct\-vs\-incorrect contrastive vectors at each layer and measure cosine similarity against the temperature 0\.8 reference vectors\.

At peak classification layers \(14, 16, 17\), the mean cosine is0\.413; across mid\-layers \(10–24\) the mean is 0\.381; across all 32 layers it is 0\.348\. This indicates that contrastive directions are*partially preserved*across temperatures—consistent with the probe’s cross\-temperature generalization \(AUROC 0\.610 on temperature 0\.1 test data, Table[7](https://arxiv.org/html/2605.05715#A7.T7)\) while explaining why absolute performance degrades relative to within\-temperature evaluation \(training AUROC 0\.716\)\.

We verify this is a genuine temperature effect, not a data artifact\. First, the failure\-mode composition at temperature 0\.1 differs substantially \(KDcontributes 68\.8% of incorrect traces vs\. 33\.3% at temperature 0\.8, asOTquestions become mostly correct under near\-greedy decoding\)\. However, reweighting the temperature 0\.8 incorrect traces to match this composition yields cosine\>\>0\.98 with the original balanced vector, ruling out composition shift as an explanation\. Second, subsampling the temperature 0\.8 data to match the temperature 0\.1 sample size \(1,111 correct, 869 incorrect\) yields cosine\>\>0\.99 with the full vector, ruling out estimation noise\. Third, at temperature 0\.1, 47\.5% of questions produce unanimous traces \(all correct or all incorrect\), raising the concern that the contrastive vector reflects question difficulty rather than reasoning quality\. To test this, we decompose into a*within\-question*vector \(from the 208 mixed\-outcome questions only, where the same question produces both correct and incorrect traces\) and a*between\-question*vector \(contrasting all\-correct vs\. all\-incorrect question means\)\. Both show comparable alignment with the temperature 0\.8 reference \(within: 0\.36–0\.41; between: 0\.36–0\.37\), indicating that the reduced cosine is not an artifact of between\-question difficulty confounds but reflects a genuine geometric shift in how correctness is encoded at different temperatures\.

## Appendix KSteering Robustness Experiments

We test supplementary robustness conditions\. The mainnn=1,273 results \(Section[5](https://arxiv.org/html/2605.05715#S5)\) provide the primary evidence\.

#### Same\-temperature steering \(n=300n=300\)\.

To directly test whether the steering null results from train/eval temperature mismatch, we evaluate mode\-specific contrastive steering \(α=1\.5\\alpha=1\.5\) at the*same*temperature used for annotation and vector construction \(temperature 0\.8\) on a 300\-question subset with 3 traces per condition \(majority\-vote accuracy\)\. Steering at the matching temperature is*more*harmful than at temperature 0\.1:Δ=−6\.7\\Delta=\-6\.7pp \(95% CI\[−12\.3,−1\.0\]\[\-12\.3,\-1\.0\]; McNemarp=0\.031p=0\.031; 29 corrections vs\. 49 damages\)\. Per\-trace accuracy shows a similar pattern \(Δ=−7\.9\\Delta=\-7\.9pp\)\. This rules out temperature mismatch as an explanation for the steering null: the classification\-correction gap persists—and widens—when the evaluation temperature matches the training distribution\. The larger damage at temperature 0\.8 likely reflects higher sampling variance amplifying the harmful perturbation\.

#### Temperature sweep \(n=100n=100\)\.

We evaluate mode\-specific contrastive steering at three additional temperature settings: greedy \(Δ=\+2\.0\\Delta=\+2\.0pp, baseline = 55\.0%\), temperature 0\.3 \(Δ=−2\.7\\Delta=\-2\.7pp, baseline = 60\.7%\), and temperature 0\.7 \(Δ=−1\.3\\Delta=\-1\.3pp, baseline = 56\.7%\)\. All deltas fall within\[−2\.7,\+2\.0\]\[\-2\.7,\+2\.0\]pp, consistent with theΔ≈0\\Delta\\approx 0pattern from the full\-scale experiments at temperature 0\.1\.

#### Steering timing\.

We compare full\-sequence steering against early\-only \(first 50% of tokens\) and late\-only \(last 50%\) steering\. Early\-token steering is most harmful \(Δ=−11\.0\\Delta=\-11\.0pp,p=0\.022p=0\.022\), suggesting that perturbation during the critical initial reasoning phase actively disrupts generation\. Full\-sequence and late\-only both produceΔ=−7\.0\\Delta=\-7\.0pp\. These results are directionally consistent with the temporal mismatch analysis \(Limitations\): the failure\-occurrence signal accumulates during generation, and static perturbation applied early—before the model commits to a reasoning path—causes the most damage\.

## Appendix LComplete Steering Configuration Sweep

Table[3](https://arxiv.org/html/2605.05715#S5.T3)and Table[5](https://arxiv.org/html/2605.05715#A5.T5)present representative configurations\. For completeness, Table[9](https://arxiv.org/html/2605.05715#A12.T9)reports all tested configurations including omittedα\\alphavalues\. All probe\-guided experiments usen=1,175n=1\{,\}175\(questions with valid probe outputs\); all others usen=1,273n=1\{,\}273\.

Table 9:Complete additive steering sweep on Llama\-3\.1\-8B and Qwen2\.5\-7B \(allα\\alphavalues tested\)\. All values in %\.Δ\\Deltain pp vs\. run\-matched baseline\. McNemarpp\-values are two\-sided\. No configuration achieves significant improvement; probe\-uniformα=1\.5\\alpha=1\.5is the only individually significant result and is harmful\. The strong\-probe L17/α\\alpha=3\.0 is significantly*harmful*\. See Table[5](https://arxiv.org/html/2605.05715#A5.T5)for rank\-kksubspace configurations \(11 additional\)\.
## Appendix MLinear Accessibility Profile \(LAP\) Analysis

To independently diagnose why steering fails, we adapt the LAP framework ofBilla \([2026](https://arxiv.org/html/2605.05715#bib.bib8)\), which distinguishes three regimes: \(1\) concept is output\-aligned \(logit\-lens detects it, steering works\); \(2\) concept is nonlinearly encoded \(probe detects, logit\-lens fails\); \(3\) concept is not cleanly extractable \(nothing works\)\. We compute two metrics per layer:

#### Alin\(logit\-lens accuracy\)\.

We project each hidden state onto the normalizedOTdirection𝐝^OT\(l\)\\hat\{\\mathbf\{d\}\}\_\{\\text\{OT\}\}^\{\(l\)\}and classify via optimal threshold, measuring how well the direction separatesOTvs\. non\-OTin the output\-aligned sense\.

#### Amlp\(probe accuracy\)\.

The trained linear probe accuracy from Table[1](https://arxiv.org/html/2605.05715#S4.F1)b, serving as the upper bound on linearly decodable information\.

#### Results\.

Peak Alin= 58\.6% \(layer 14\), far below peak Amlp= 81\.5% \(layer 17\)\. The gapΔ=Amlp−Alin=0\.23\\Delta=A\_\{\\text\{mlp\}\}\-A\_\{\\text\{lin\}\}=0\.23across layers 5–25 indicates thatOTinformation is present in hidden states but*not output\-aligned*—the direction does not project to coherent vocabulary\-space tokens\. Examining the top\-20 tokens from𝐝^OT\(l\)⋅𝐖UT\\hat\{\\mathbf\{d\}\}\_\{\\text\{OT\}\}^\{\(l\)\}\\cdot\\mathbf\{W\}\_\{U\}^\{T\}at all layers reveals semantically incoherent token lists \(e\.g\., layer 16: “Knight,” “hra,” “iagnostics”; layer 24: “eh,” “Engel,” “ser”\), with 0/10 tokens semantically related to overthinking, correctness, or medical reasoning at the peak layer\. This contrasts with steerable concepts \(e\.g\., refusal\), where logit\-lens projections typically surface semantically coherent tokens \(“sorry,” “cannot,” “harmful”\)\.

#### Regime diagnosis\.

The combination of high Amlp\(\>\>0\.8\) with low Alin\(≈\\approx0\.59\) placesOTinRegime 2/3: the concept is decodable by a trained probe but not accessible through the model’s output pathway, consistent with the broader steering failure pattern\. This provides an independent geometric explanation complementary to the specificity ratio analysis: theOTdirection does not align with the model’s vocabulary projection, so additive perturbation along this direction produces incoherent logit shifts rather than targeted answer correction\.

Similar Articles

When is Your LLM Steerable?

arXiv cs.CL

This paper investigates when activation steering succeeds or fails for LLMs by analyzing early decoding dynamics. The authors introduce ASTEER, a large testbed of steered generations, and train a GBDT classifier to predict steering outcomes from early hidden states, enabling efficient steering strength search.

Steered LLM Activations are Non-Surjective

Hugging Face Daily Papers

This paper proves that activation steering in LLMs produces internal states that cannot be replicated by any textual prompt, establishing a formal separation between white-box steerability and black-box prompting.

Decomposing and Steering Functional Metacognition in Large Language Models

arXiv cs.CL

This research paper investigates functional metacognition in Large Language Models, demonstrating that internal states like evaluation awareness and self-assessed capability are linearly decodable from residual stream activations. The authors propose a mechanistic framework to steer these states, showing causal control over reasoning behaviors, verbosity, and safety responses.

When LLM Reward Design Fails: Diagnostic-Driven Refinement for Sparse Structured RL

arXiv cs.LG

This paper frames LLM-generated reward shaping for sparse structured RL as a debugging problem, identifying failure modes like reward flooding and semantic misunderstanding. The authors propose diagnostic-driven iterative refinement, achieving dramatic success rate improvements (e.g., DoorKey-8×8 from 2.3% to 97.6%) compared to one-shot generation.