纠正何时才算修复?工具使用型大语言模型内部干预的机制化审计

arXiv cs.CL 论文

摘要

本文提出了SAKIKO——一个机制化审计框架,揭示了工具使用型大语言模型中的激活引导干预可能只是将概率质量重新分配,而非真正修复问题;研究发现,其中一个净增益为+55的干预,会破坏其所触及的基准正确决策中超过一半的案例。

arXiv:2609.36138v1 Announce Type: new Abstract: Before invoking external tools, an agentic LLM must select among a K-way action space: executing a call, seeking clarification, answering directly, or declining. While internal activation steering can alter these pre-execution decisions, conventional aggregate metrics obscure where altered states land and what collateral damage they inflict. We present SAKIKO, an auditing framework that formalizes representation repair via directional error discovery, router-conditioned intervention, destination-resolved verification, and prospectively frozen statistical licensing. Across seven LLMs on When2Call and MetaTool, channel-keyed interventions induce direction-specific net gains in five models; across three sealed evaluations, none of 59 budget-matched random directions matches calibrated target gain. Crucially, destination auditing shows that behavioral movement does not equal repair: an intervention achieving +55 net gain corrupts over half of the baseline-correct decisions it touches, and promising point estimates on Qwen3-4B and Gemma-2-9B are formally declined due to finite-sample uncertainty. SAKIKO establishes the necessity of outcome-resolved adjudication before claiming internal repair. Code: https://github.com/ruizheliUOA/mechanistic-tool-use-llm.
查看原文
查看缓存全文

缓存时间: 2026/09/30 09:50

# When Does Correction Become Repair? Mechanistic Auditing of Internal Interventions in Tool-Using LLMs
Source: [https://arxiv.org/html/2609.36138](https://arxiv.org/html/2609.36138)
Ruizhe LiAffiliation:School of Computer Science, University of Birmingham, UK

###### Abstract

Before invoking external tools, an agentic LLM must select among aKK\-way action space: executing a call, seeking clarification, answering directly, or declining\. While internal activation steering can alter these pre\-execution decisions, conventional aggregate metrics obscure where altered states land and what collateral damage they inflict\. We presentSAKIKO, an auditing framework that formalizes representation repair via directional error discovery, router\-conditioned intervention, destination\-resolved verification, and prospectively frozen statistical licensing\. Across seven LLMs on When2Call and MetaTool, channel\-keyed interventions induce direction\-specific net gains in five models; across three sealed evaluations, none of 59 budget\-matched random directions matches calibrated target gain\. Crucially, destination auditing shows that behavioral movement does not equal repair: an intervention achieving\+55\+55net gain corrupts over half of the baseline\-correct decisions it touches, and promising point estimates on Qwen3\-4B and Gemma\-2\-9B are formally declined due to finite\-sample uncertainty\.SAKIKOestablishes the necessity of outcome\-resolved adjudication before claiming internal repair\.

## 1Introduction

The evolution of LLMs into autonomous agents fundamentally alters how model failures arise\. Beyond factual hallucination in generated text\([Ji et al\., 2023](https://arxiv.org/html/2609.36138#bib.bib1)\), tool\-augmented agents face a structurally distinct failure mode:pre\-execution tool\-decision errors\. Before an external API is invoked, evidence retrieved, or execution feedback received, an agent must decide upon an appropriate course of action\. It may invoke a tool when prerequisites are absent, answer prematurely when clarification is required, decline an answerable query, or omit an indispensable external tool call\. In these scenarios, the failure is situated not in the generated content, but in the upstream selection of action type\.

While extensive research has advanced downstream execution, such as API parameter synthesis, multi\-tool composition, and retrieval\([Schick et al\., 2023](https://arxiv.org/html/2609.36138#bib.bib4);[Tang et al\., 2023](https://arxiv.org/html/2609.36138#bib.bib6);[Patil et al\., 2024](https://arxiv.org/html/2609.36138#bib.bib7);[Qin et al\., 2024b](https://arxiv.org/html/2609.36138#bib.bib5)\), the upstream decision ofwhetherandwhento act\([Wei et al\., 2022](https://arxiv.org/html/2609.36138#bib.bib3);[Qin et al\., 2024a](https://arxiv.org/html/2609.36138#bib.bib2)\)remains underexplored as a target for representation\-level intervention\. Recent mechanistic studies, such as the Activation Steering Adapter \(ASA\)\([Wang et al\., 2026b](https://arxiv.org/html/2609.36138#bib.bib8)\)and CAST\([Lee et al\., 2025](https://arxiv.org/html/2609.36138#bib.bib13)\), demonstrate that pre\-execution states can be conditionally decoded and guided via internal activations\. However, existing interventions typically assume binary action spaces \(e\.g\., tool vs\. no\-tool\), where departing an incorrect state deterministically lands the model in the correct one\. In realistic agentic environments withKK\-way action topologies, leaving an erroneous state provides no guarantee of reaching the target, and an intervention may simply displace probability mass to a third, equally invalid action\. Consequently,behavioral movement does not constitute repair\.

This distinction motivates our central question:When an internal intervention alters a pre\-execution tool decision, when does that change constitute genuine repair rather than mere behavioral movement?We approach this problem through a scientific progression across three core dimensions: i\)Correction: Can directional error channels be discovered from a model’s internal representation, detected at inference, and selectively steered? ii\)Verification: Does the intervention transition the model to the valid target action, and does it preserve decisions that were already correct? iii\)Licensing: Do the observed destination outcomes and their quantified uncertainty provide sufficient empirical evidence to license a formal claim of repair?

Figure 1:Overview of SAKIKO\. The framework discovers directional pre\-execution tool\-decision error channels, detects channel\-associated states per input, applies the corresponding channel\-keyed intervention, verifies the post\-intervention destination and collateral on initially correct decisions, and licenses only the strongest repair claim supported by the evidence\.To operationalize this progression, we introduceSAKIKO\(StructuredAdjudication ofKeyedInterventions overK\-wayOutcomes\), a staged adjudication framework that treats representation\-level repair as an evidence process rather than an isolated steering operation \(Fig\.[1](https://arxiv.org/html/2609.36138#S1.F1)\)\. SAKIKO structures intervention and evaluation across five functional stages, i\.e\.,Discoveryof a model’s directional error channels,Detectionof channel\-associated states per input,Keyed Interventionapplied only where a channel fires,Destination\-Resolved Verificationover the completeKK\-way action space, andStatistical Licensingagainst a prospectively frozen evidence hierarchy\.

Our empirical evaluations demonstrate why each stage of this pipeline does distinct work\. At theCorrectionstage, channel\-keyed interventions succeed in altering behavior, producing marked net gains \(e\.g\.,\+79\+79on the locked test set for Qwen2\.5\-7B and consistent positive gains across all seeds for Phi\-3\.5\)\. Subjecting those corrections toVerificationandLicensinguncovers three dissociations that aggregate metrics conceal\. A Phi\-3\.5 intervention worth a net\+55\+55corrupts 52 of the 93 baseline\-correct decisions its Router fires on\. On Qwen3\-8B, activation steering and an output\-score baseline reach near\-identical net gains at target\-hit rates of0\.73080\.7308against0\.63790\.6379, so one lands on the intended target while the other redistributes mass across alternate errors\. And Qwen3\-4B and Gemma\-2\-9B carry promising point estimates \(0\.57810\.5781and0\.62960\.6296\) thatSAKIKOdeclines to license, because their confidence bounds at the available sample sizes remain inconclusive\. These empirical observations confirm thatbehavioral improvement, destination\-correct repair, preservation of correct behavior, and evidential sufficiency do not inherently coincide\.

Our main contributions are as follows:

1. 1\.Conceptual Framing: We formalize pre\-execution tool\-decision errors withinKK\-way action spaces, demonstrating that leaving an error state cannot be equated with target repair\.
2. 2\.SAKIKO Framework: We formulate an end\-to\-end pipeline that integrates directional error discovery, conditioned steering, destination\-resolved verification, and statistical licensing\.
3. 3\.Empirical Adjudication & Controls: Across seven LLMs spanning five model families and three benchmarks in three roles, with an eighth model stopped at pre\-intervention screening, we test channel\-keyed steering against structural controls to show that headline accuracy metrics can conceal collateral damage and destination redistribution\.
4. 4\.Mechanistic Decoupling: We show that linear decodability does not imply steerability, and optimal injection site, dosage, and intervention success diverge from probe accuracy, showing that verification cannot be reduced to internal detection alone\.

## 2Related Work

Tool Learning and Epistemic Selection\.Prior tool learning focuses on downstream execution: API adherence, parameter synthesis and environment generalization via fine\-tuning or retrieval\([Schick et al\., 2023](https://arxiv.org/html/2609.36138#bib.bib4);[Qin et al\., 2024b](https://arxiv.org/html/2609.36138#bib.bib5);[Tang et al\., 2023](https://arxiv.org/html/2609.36138#bib.bib6);[Patil et al\., 2024](https://arxiv.org/html/2609.36138#bib.bib7);[Qin et al\., 2024a](https://arxiv.org/html/2609.36138#bib.bib2)\)\. Before parameterization, however, an agent must resolve the epistemic decision ofwhetherexternal action is needed\.[Wang et al\. \(2026a\)](https://arxiv.org/html/2609.36138#bib.bib9)taxonomize failure modes such as tool bypass and over\-delegation, but treat the agent as a black box, omitting latent geometry and representation\-level intervention\. Closest to our setting,[Wu et al\. \(2026a\)](https://arxiv.org/html/2609.36138#bib.bib24)read tool need from hidden states to drive controllers, but score the outcome by task performance rather than by where redirected decisions land\.

Agent Observability and Activation Steering\.Pre\-execution states are linearly decodable, enabling real\-time detection of tool hallucinations by residual\-stream probing\([Healy et al\., 2026](https://arxiv.org/html/2609.36138#bib.bib10)\)and localized risk monitoring by sparse autoencoders\([Tatsat and Shater, 2026](https://arxiv.org/html/2609.36138#bib.bib11)\)\. Beyond passive detection, representation engineering\([Zou et al\., 2025](https://arxiv.org/html/2609.36138#bib.bib12);[Lee et al\., 2025](https://arxiv.org/html/2609.36138#bib.bib13)\)enables direct control: linear steering vectors invert tool choices in constrained menus\([Wu et al\., 2026b](https://arxiv.org/html/2609.36138#bib.bib14)\), and router\-conditioned vectors overcome inertia in multi\-turn settings\([Wang et al\., 2026b](https://arxiv.org/html/2609.36138#bib.bib8)\)\. Closest in apparatus,[Shi et al\. \(2026\)](https://arxiv.org/html/2609.36138#bib.bib26)recover a sparse\-autoencoder basis for the call/no\-call decision on the benchmark we also use, estimate an activation\-independent CALL offset, and cancel it with a closed\-form counter\-bias shift\.

SAKIKO Distinction: Audited Repair inKK\-Way Action Spaces\.Prior tool steering evaluates narrow binary exits or pairwise tool swaps\([Wu et al\., 2026b](https://arxiv.org/html/2609.36138#bib.bib14);[Shi et al\., 2026](https://arxiv.org/html/2609.36138#bib.bib26)\), while diagnostic multiclass work remains observational\([Zhao et al\., 2026](https://arxiv.org/html/2609.36138#bib.bib25)\)\. Crucially, behavioral readouts often diverge from internal representation effects\([Jiang et al\., 2026](https://arxiv.org/html/2609.36138#bib.bib27)\)\. InKK\-way action topologies, simply measuring state exits obscures destination divergence, collateral degradation, and lack of statistical support\. Moving beyond passive probes\([Healy et al\., 2026](https://arxiv.org/html/2609.36138#bib.bib10);[Zhao et al\., 2026](https://arxiv.org/html/2609.36138#bib.bib25)\)and uncalibrated steering\([Wu et al\., 2026b](https://arxiv.org/html/2609.36138#bib.bib14);[Wang et al\., 2026b](https://arxiv.org/html/2609.36138#bib.bib8)\),SAKIKOestablishes an end\-to\-end adjudication pipeline: discovering channel vectors, verifying multiclass destinations, auditing baseline preservation, and enforcing preregistered uncertainty bounds for certified repair \(App\.[C](https://arxiv.org/html/2609.36138#A3)for more discussion\)\.

## 3The SAKIKO Framework

SAKIKOformalizes the criteria under which an internal intervention on a pre\-execution tool decision constitutes a valid claim of repair\. Rather than proposing another ad\-hoc steering vector,SAKIKOframes evaluation as a non\-collapsing evidentiary pipeline:Correction→\\rightarrowDestination Correctness→\\rightarrowPreservation→\\rightarrowEvidential Sufficiency→\\rightarrowLicensed Repair\. In multiclass action spaces, no property guarantees the next \(Figs\.[4](https://arxiv.org/html/2609.36138#S5.F4)and[5](https://arxiv.org/html/2609.36138#S5.F5)\): vacating an error can divert mass to another invalid action, the same intervention can corrupt decisions that were already correct, and positive point estimates may lack statistical significance\. Consequently,SAKIKOdecouples representation steering\([Lee et al\., 2025](https://arxiv.org/html/2609.36138#bib.bib13);[Wang et al\., 2026b](https://arxiv.org/html/2609.36138#bib.bib8)\)from destination verification and formal statistical licensing\.

### 3\.1Problem Formulation and Outcome Topology

Let𝒴\\mathcal\{Y\}be the pre\-execution action space \(direct response, API call, clarification request, task declination\)\. For an inputxxa frozen agent produces a baseline actiony^0∈𝒴\\hat\{y\}\_\{0\}\\in\\mathcal\{Y\}, evaluated against a referencey⋆∈𝒴y^\{\\star\}\\in\\mathcal\{Y\}, which partitions inputs into baseline\-correct and baseline\-error populations\. Rather than treating errors monolithically we model failure as an ordereddirectional error channelc≡\(y⋆→y^0\)c\\equiv\(y^\{\\star\}\\rightarrow\\hat\{y\}\_\{0\}\), with gold targetg≡y⋆g\\equiv y^\{\\star\}and erroneous sources≡y^0s\\equiv\\hat\{y\}\_\{0\}\(g≠sg\\neq s\)\. The intervention emits an updated decisiony^1∈𝒴\\hat\{y\}\_\{1\}\\in\\mathcal\{Y\}\. This induces an exhaustive, mutually exclusive five\-way partition \(Table[1](https://arxiv.org/html/2609.36138#A2.T1), App\.[B](https://arxiv.org/html/2609.36138#A2)\): theNcN\_\{c\}channel errors split intosource retained\(ScS\_\{c\}\),gold arrival\(AcA\_\{c\}\) andother wrong\(OcO\_\{c\}\), with source exitsXc=Ac\+OcX\_\{c\}=A\_\{c\}\+O\_\{c\}, and the baseline\-correct decisions the Router touches \(CexpC\_\{\\mathrm\{exp\}\}\) split intocorrect retainedandbroken\(BB\)\. Conventional benchmarks assess steering via an aggregate net gain, and two conventions are in use\. We keep them apart throughout:Gwhole=Fixed−BG\_\{\\mathrm\{whole\}\}=\\mathrm\{Fixed\}\-Bcounts every baseline error in the evaluation population that becomes correct, whileGchannel=∑cAc−BG\_\{\\mathrm\{channel\}\}=\\sum\_\{c\}A\_\{c\}\-Bcounts only arrivals inside the adjudicated channel, so the first exceeds the second whenever correction reaches an error the channel does not key \(\+40\+40against\+38\+38on Qwen3\-8B\)\. Every net quantity we report carries one of the two labels\. However, the map from the five outcome classes to either aggregate is many\-to\-one: a given net gain does not determineTHc\\mathrm\{TH\}\_\{c\}or the collateral, as App\.[B](https://arxiv.org/html/2609.36138#A2)shows\.

### 3\.2Overview of the SAKIKO Pipeline

SAKIKOintroduces a five\-stage pipeline integrating representation\-level steering with statistical auditing \(Fig\.[2](https://arxiv.org/html/2609.36138#S3.F2)\): \(1\)Discoveryisolates directional error channels\(g→s\)\(g\\rightarrow s\)from baseline confusion matrices; \(2\)Detectionfits a linear Router over frozen mid\-layer representations, with a firing threshold selected on validation data, to gate interventions; \(3\)Correctionsteers routed inputs along channel\-specific latent vectors; \(4\)Destination Verificationtracks post\-intervention trajectories across the fullKK\-way action topology for both error and clean cohorts; and \(5\)Statistical Licensingbounds empirical gains via a pre\-registered evidence hierarchy, certifying repair only when confidence bounds rule out stochastic drift\. Therefore, while latent manipulation confirms steerability, only structured adjudication certifies genuine repair\.

![Refer to caption](https://arxiv.org/html/2609.36138v1/SAKIKO_ModelSide_Corrected_FullyEditable.png)Figure 2:Model\-side view of SAKIKO\.\(a\)Over a frozen model, a first pass reads hidden state at observation layerℓobs\\ell\_\{\\mathrm\{obs\}\}and a channel Router decides whether to fire; on a miss the baseline decisiony^0\\hat\{y\}\_\{0\}stands, and on a hit the input is rerun so the correction can be applied at the injection layerℓinj\\ell\_\{\\mathrm\{inj\}\}, which precedesℓobs\\ell\_\{\\mathrm\{obs\}\}in the stack\.\(b\)Each baseline error is keyed by its directionc=\(y⋆→y^0\)c=\(y^\{\\star\}\\rightarrow\\hat\{y\}\_\{0\}\)and corrected by addingqc​sc​dcq\_\{c\}s\_\{c\}d\_\{c\}at that site, withdcd\_\{c\}unit\-norm,scs\_\{c\}a local activation scale andqcq\_\{c\}a dose frozen before evaluation\.\(c\)The decision precedes any tool call: a frozen readout scores the action space𝒴\\mathcal\{Y\}and its argmax givesy^0\\hat\{y\}\_\{0\}, andy^1\\hat\{y\}\_\{1\}after intervention, with destination, collateral and licensing\.Channel Discovery and Latent Detection\.Rather than assuming error topologiesa priori,SAKIKOenumerates them from the training\-split confusion matrix and retains directional channelsc=\(y⋆→y^0\)c=\(y^\{\\star\}\\rightarrow\\hat\{y\}\_\{0\}\)that clear a support minimum on both the training and validation splits \(App\.[D\.1](https://arxiv.org/html/2609.36138#A4.SS1)and[D\.2](https://arxiv.org/html/2609.36138#A4.SS2)\)\. Because dominant error modes vary across architectures and share non\-trivial cosine similarities, e\.g\., on Qwen2\.5\-7B’s three When2Call channels, pairwise cosines of0\.4880\.488,0\.5690\.569and0\.7060\.706measured at a common layer, a single monolithic tool\-error vector conflates functionally disparate failures\. For each retained channelcc, we train a linear Routerr⁡\(hobs\)→\(c^,pc^\)r\(h\_\{\\mathrm\{obs\}\}\)\\rightarrow\(\\hat\{c\},p\_\{\\hat\{c\}\}\)on final\-token representationshobsh\_\{\\mathrm\{obs\}\}at layerℓobs\\ell\_\{\\mathrm\{obs\}\}, triggering steering iffpc^≥τc^p\_\{\\hat\{c\}\}\\geq\\tau\_\{\\hat\{c\}\}, and unrouted inputs remain unperturbed \(y^1=y^0\\hat\{y\}\_\{1\}=\\hat\{y\}\_\{0\}\)\. Crucially, probe decodability does not guarantee steerability: while Router accuracy plateaus across intermediate layers, downstream steering efficacy drops sharply \(§[M](https://arxiv.org/html/2609.36138#A13)\)\. The Router thus serves to bound exposure on baseline\-correct samples \(CexpC\_\{\\mathrm\{exp\}\}\), not to guarantee repair\.

Channel\-Keyed Intervention Protocol\.Upon Router activation for channelcc, we intervene at layerℓinj\\ell\_\{\\mathrm\{inj\}\}on the MLP output, before it rejoins the decoder residual stream, at every sequence position, parameterized by a fixed tupleθc=\{dc,ℓobs,ℓinj,qc,τc\}\\theta\_\{c\}=\\\{d\_\{c\},\\ell\_\{\\mathrm\{obs\}\},\\ell\_\{\\mathrm\{inj\}\},q\_\{c\},\\tau\_\{c\}\\\}:hℓinj′=hℓinj\+qc​sc​dch^\{\\prime\}\_\{\\ell\_\{\\mathrm\{inj\}\}\}=h\_\{\\ell\_\{\\mathrm\{inj\}\}\}\+q\_\{c\}\\,s\_\{c\}\\,d\_\{c\}, wheredcd\_\{c\}is a unit\-norm steering direction,scs\_\{c\}is a layer scale factor, andqcq\_\{c\}denotes normalized dosage\. Subsequent decoding remains unmodified\. Ifℓinj≤ℓobs\\ell\_\{\\mathrm\{inj\}\}\\leq\\ell\_\{\\mathrm\{obs\}\}, two\-pass inference avoids lookahead leakage\. Structural controls run as each protocol permits: budget\-matched random vectors everywhere, sign reversals and layer misallocations where the records permit, and cross\-channel substitution only in the historical battery, which the frozen sealed inventory excludes by construction \(App\.[K](https://arxiv.org/html/2609.36138#A11)\)\.

Destination\-Resolved Verification and Collateral Auditing\.Verification resolves exact transitions across𝒴\\mathcal\{Y\}\. On baseline errors, destination correctness is read over source exits asTHc=Ac/Xc\\mathrm\{TH\}\_\{c\}=A\_\{c\}/X\_\{c\}, separating gold arrivals from lateral redistribution intoOcO\_\{c\}, and asTGc=\(Ac−Oc\)/Nc\\mathrm\{TG\}\_\{c\}=\(A\_\{c\}\-O\_\{c\}\)/N\_\{c\}, which normalises over the whole channel and so carries coverage as well as composition\. Collateral is reported on two denominators,E1=B/\|C\|\\mathrm\{E1\}=B/\|C\|andE2=B/\|Cexp\|\\mathrm\{E2\}=B/\|C\_\{\\mathrm\{exp\}\}\|withE2≥E1\\mathrm\{E2\}\\geq\\mathrm\{E1\}: because a selective Router minimises\|Cexp\|\|C\_\{\\mathrm\{exp\}\}\|,E1\\mathrm\{E1\}alone understates the risk to a touched decision by\|C\|/\|Cexp\|\|C\|/\|C\_\{\\mathrm\{exp\}\}\|\. Appendix[B](https://arxiv.org/html/2609.36138#A2)defines every quantity against its named population\.

Statistical Licensing and Repair Adjudication\.Licensing evaluates empirical evidence against a pre\-registered evidence hierarchy \(Table[6](https://arxiv.org/html/2609.36138#A5.T6), App\.[E](https://arxiv.org/html/2609.36138#A5)\), assessing evidentiary rigor rather than raw model capability\. A configurationπ\\piearns a formal repair claim iff it satisfies the conjunction, in whichPRESERVINGpop\\textsc\{PRESERVING\}\_\{\\text\{pop\}\}is the population\-level form the frozen gate tests and not the stronger exposure\-conditional one \(App\.[E](https://arxiv.org/html/2609.36138#A5)\):REPAIR​\(π\)≡CORRECTABLE​\(π\)∧PRESERVINGpop​\(π\)∧LICENSABLE​\(π\)\\textsc\{REPAIR\}\(\\pi\)\\;\\equiv\\;\\textsc\{CORRECTABLE\}\(\\pi\)\\;\\wedge\\;\\textsc\{PRESERVING\}\_\{\\text\{pop\}\}\(\\pi\)\\;\\wedge\\;\\textsc\{LICENSABLE\}\(\\pi\)\. Under our protocol, adjudication is governed by a ten\-condition gate evaluating sample support, specificity over matched\-random baselines \(p≤0\.05p\\leq 0\.05\), threshold criteria forTHc\\mathrm\{TH\}\_\{c\}andE1\\mathrm\{E1\}, bootstrap confidence bounds, and exact baseline replication \(§[5\.3](https://arxiv.org/html/2609.36138#S5.SS3)\)\. The gate outputs an explicit verdict:ADMITlicenses the repair claim for that setting under the frozen conditions, whileDECLINErejects it and isolates the failed condition\. The licence is narrower than the property it is named for\. Preservation is adjudicated onE1\\mathrm\{E1\}, the denominator the preregistration fixed, so anADMITcertifies that collateral damage is bounded across the correct population and*not*that it is bounded on the decisions the Router actually touches; no setting we evaluate attains the latter \(§[5\.2](https://arxiv.org/html/2609.36138#S5.SS2)\)\. By separating unverified steering from a licensed claim,SAKIKOprevents aggregate net gains or optimistic point estimates from obscuring localized intervention failures\.

## 4Empirical Evaluation: Channel\-Keyed Correction

We evaluate the first stage of theSAKIKOprogression: whether channel\-keyed internal interventions produce genuine, direction\-specific corrections of pre\-execution tool decisions\. We address three sequential questions: \(1\) Does intervening along an extracted channel direction shift decisions toward the target action? \(2\) Is this displacement driven by the geometry of the identified direction rather than arbitrary representation noise? \(3\) How reliably does this steerability hold across diverse model architectures? We confine this analysis to theCorrectionstage, and destination\-resolved outcomes, collateral costs on baseline\-correct samples, and statistical licensing are evaluated in §[5](https://arxiv.org/html/2609.36138#S5)\.

### 4\.1Experimental Setup

Tasks and Readout Protocol\.We treat three benchmarks independently rather than pooling results \(App\.[F](https://arxiv.org/html/2609.36138#A6)\)\. We primarily evaluate on When2Call\([Ross et al\., 2025](https://arxiv.org/html/2609.36138#bib.bib15)\), scoring its four pre\-execution actions via teacher forcing to cleanly isolate steering effects\. We test binary bidirectional steering using MetaTool\([Huang et al\., 2024](https://arxiv.org/html/2609.36138#bib.bib16)\)on Qwen2\.5\-7B\. ACEBench provides an out\-of\-domain four\-action ontology evaluated via free generation and rule parsing\. Although the readout protocol certified successfully \(0\.7560\.756accuracy;0\.6870\.687macro\-F1\), ACEBench failed downstream channel discovery criteria, i\.e\., specifically lacking minimum transition support and yielding sub\-threshold paraphrase agreement \(0\.560\.56vs\.0\.800\.80\)\. Per our preregistered protocol, ACEBench was retired without intervention experiments\. This negative case demonstrates that our framework can rigorously validate readout stability and screen out unviable task topologies prior to intervention\.

Models, Protocols, and Metrics\.We evaluate seven LLMs \(3\.8B–9B\) across two protocol cohorts using a 2,556/548/548 train/val/test split\. Thehistoricalcohort \(Phi\-3\.5\-mini\([Abdin et al\., 2024](https://arxiv.org/html/2609.36138#bib.bib18)\), Qwen2\.5\-7B\([Qwen et al\., 2025](https://arxiv.org/html/2609.36138#bib.bib19)\), Llama\-3\.1\-8B\([Grattafiori et al\., 2024](https://arxiv.org/html/2609.36138#bib.bib20)\), and Mistral\-7B\([Jiang et al\., 2023](https://arxiv.org/html/2609.36138#bib.bib21)\)\) targets three channels, which on Qwen2\.5\-7B cover about78%78\\%of errors, evaluated via whole\-population net gainGwholeG\_\{\\mathrm\{whole\}\}\. The prospectivesealedcohort \(Qwen3\-4B, Qwen3\-8B\([Yang et al\., 2025](https://arxiv.org/html/2609.36138#bib.bib22)\), and Gemma\-2\-9B\([Team et al\., 2024](https://arxiv.org/html/2609.36138#bib.bib23)\)\) pre\-registers all vectors, thresholds, and doses∈\{0\.125,0\.25,1\.0\}\\in\\\{0\.125,0\.25,1\.0\\\}on thecannot\_answer→\\rightarrowtool\_callchannel prior to test unblinding, evaluated via target gainTGc=\(Ac−Oc\)/Nc\\mathrm\{TG\}\_\{c\}=\(A\_\{c\}\-O\_\{c\}\)/N\_\{c\}\. Each run is benchmarked against the structural controls its protocol permits: zero, budget\-matched random, reversed and wrong\-layer arms in the sealed inventory, with cross\-channel substitution excluded from that inventory by construction and run only on Phi\-3\.5\-mini \(App\.[K](https://arxiv.org/html/2609.36138#A11)\); each sealed setting executes6565arms in total, six named and5959random\. Sealed significance is adjudicated against those 59 directions by an add\-one Monte Carlo test over the empirical null,p=\(1\+∑𝕀\[TGrand≥TGreal\]\)/60p=\(1\+\\sum\\mathbb\{I\}\[\\mathrm\{TG\}\_\{\\mathrm\{rand\}\}\\geq\\mathrm\{TG\}\_\{\\mathrm\{real\}\}\]\)/60; they are isotropic unit vectors rescaled to the calibrated budgetqc​scq\_\{c\}s\_\{c\}, drawn under seeds disjoint from development and hash\-pinned before the split was opened \(App\.[K](https://arxiv.org/html/2609.36138#A11)\)\.

Figure 3:Correction is direction\-specific\.\(a\)On Phi\-3\.5 locked test, calibrated direction yields a net gain of\+55\+55, whereas random and sign\-reversed vectors yield\+14\+14, layer shifts yield\+3\+3, and mismatched channel vectors incur−17\-17\(arms colored by corrupted component; four arms show aggregate counts\)\.\(b\)Under sealed protocol, all 59 budget\-matched random vectors passing the Router gate trail calibrated target gain \(add\-onep=0\.017p=0\.017across all settings\)\.
### 4\.2Internal Interventions Induce Targeted Action Transitions

Quantitative Gains Across Architectures and Task Formulations\.Channel\-keyed interventions consistently drive intended behavioral shifts across model families and tasks\. In the historical cohort on When2Call, intervention yields net gains of\+79\+79decisions on Qwen2\.5\-7B \(92 corrections vs\. 13 regressions\) and\+55\+55on Phi\-3\.5\-mini\. In the sealed cohort, all models achieve positive target gains \(TGc\\mathrm\{TG\}\_\{c\}\) on the frozencannot\_answer→\\rightarrowtool\_callchannel: Qwen3\-8B reachesTGc=0\.276\\mathrm\{TG\}\_\{c\}=0\.276\(\+24\+24net arrivals /8787errors\), Qwen3\-4B achievesTGc=0\.081\\mathrm\{TG\}\_\{c\}=0\.081\(\+10/124\+10/124\), and Gemma\-2\-9B achievesTGc=0\.073\\mathrm\{TG\}\_\{c\}=0\.073\(\+7/96\+7/96\)\. On Qwen3\-8B, the Router fires on82\.8%82\.8\\%of errors, triggering59\.8%59\.8\\%source exits and43\.7%43\.7\\%target arrivals\. Finally, bidirectional steering on MetaTool with Qwen2\.5\-7B corrects both premature and omitted tool calls, yielding mean net gains of\+8\.6\+8\.6and\+13\.0\+13\.0across five independent runs\.

Localized Latent Steerability vs\. Global Propensity Drift\.These results confirm that pre\-execution tool decisions can be selectively steered in latent space at inference time\. Crucially, bidirectional corrections on opposing MetaTool channels demonstrate that interventions act via channel\-keyed rectification rather than indiscriminate shifts in global tool\-use propensity\. However, while target gains confirm departure from source errors across binary and 4\-way settings, aggregate displacement does not ensure ground\-truth arrival, which motivatesSAKIKO’s destination verification stage\.

### 4\.3Empirical Confirmation of Directional Specificity

Adjudication Against Structural Counterfactuals and Permutations\.Rigorous structural controls confirm that behavioral shifts stem from specific channel geometry rather than arbitrary perturbation\. On Phi\-3\.5\-mini \(Fig\.[3](https://arxiv.org/html/2609.36138#S4.F3)\), calibrated vector \(\+55\+55net gain\) dramatically outperforms norm\-matched random \(\+14\+14\), sign\-reversed \(\+14\+14\), layer\-misallocated \(\+3\+3\), and performance\-degrading cross\-channel baselines\. Likewise, on Qwen2\.5\-7B, calibrated direction \(\+79\+79\) substantially exceeds sign\-reversed controls \(\+16\+16\) and ten budget\-matched random vectors \(mean\+25\.2\+25\.2, max\+50\+50\)\. In sealed evaluations, calibrated intervention decisively beats the empirical null of 59 pre\-registered random vectors, achieving the minimum add\-one Monte Carlo significance floor \(p=1/60=0\.017p=1/60=0\.017\) across all architectures\. Calibrated vectors yield significantly more gold arrivals than random baselines on Qwen3\-8B \(3838vs\. max88\), Qwen3\-4B \(3737vs\.22\), and Gemma\-2\-9B \(1717vs\.00\), while zero\-dose, reversed, and wrong\-layer perturbations cause negligible displacement \(TGc≤0\.058\\mathrm\{TG\}\_\{c\}\\leq 0\.058\)\.

Geometric Alignment and Layer\-Specific Causal Control\.These counterfactual adjudications confirm that behavioral repair is causally governed by internal directional alignment\. The failure of layer\-mismatched and cross\-channel controls demonstrates that intervention requires dual specificity: the vector must encode the exact channel geometry and target the precise layer arbitrating tool selection\. Because uncalibrated noise and arbitrary perturbations fail to produce targeted gold arrivals, the intervention acts via site\- and direction\-specific control rather than a non\-specific activation shock, though it identifies locus of intervention rather than fine\-grained circuit components\.

### 4\.4Cross\-Model Analysis and Pre\-Intervention Attrition

Model Screening, Structural Failures, and Attrition Statistics\.Directional specificity is confirmed in five of the seven evaluated architectures, alongside two control failures and one pre\-intervention screening exit \(Table[5](https://arxiv.org/html/2609.36138#A5.T5)\); together with the retired third benchmark, this demonstrates that the protocol halts unviable settings as rigorously as it licenses valid ones\. In the historical cohort, Llama\-3\.1\-8B and Mistral\-7B fail structural specificity: for Llama\-3\.1\-8B, 5 of 20 random vectors match or exceed intervention gains, while on Mistral\-7B, the calibrated direction \(\+12\+12\) underperforms both sign\-reversed \(\+19\+19\) and mean random baselines \(\+29\+29\)\. In the prospective cohort, despite the highest baseline accuracy \(0\.45330\.4533\), Qwen3\.5\-9B was disqualified before intervention due to severe class skew, i\.e\., favoring information requests \(0\.3260\.326\) over refusals \(0\.0870\.087\)\. This produced ample training/validation errors \(300/170\) but only 29 baseline\-correct reference instances, violating the preregistered minimum support threshold of 30 and triggering an automatic protocol halt\.

Non\-Universality of Linear Steering and Role of Protocol Gates\.Linear steerability is neither universally shared across architectures nor guaranteed by high baseline accuracy\. In models like Mistral\-7B and Llama\-3\.1\-8B, latent steering behaves as non\-specific noise, inducing stochastic drift rather than targeted repair\. Furthermore, the pre\-intervention disqualification of Qwen3\.5\-9B demonstrates the critical utility ofSAKIKO’s preregistered gates: by filtering out under\-supported reference populations prior to unblinding, the protocol prevents underpowered point estimates from producing false claims of repair before verification begins\.

## 5Analysis: From Directional Correction to Licensed Repair

Figure 4:Correction is not repair\.\(a\)Outcome\-resolved steering on Phi\-3\.5\-mini\. Of 200 routed errors, 107 reach the target while 46 shift into alternative errors; concurrently, 52 of 93 exposed correct decisions break, despite a net\+55\+55gain\.\(b\)Target\-hit rates \(95%95\\%bootstrap CIs\) against the frozen0\.500\.50criterion\. Qwen3\-8B clears the gate; Qwen3\-4B and Gemma\-2\-9B yield favorable point estimates whose confidence intervals cross the threshold, warranting formalDECLINEverdicts due to evidentiary shortfall rather than uncorrectability\.While §[4](https://arxiv.org/html/2609.36138#S4)confirms genuine, direction\-specific behavioral steering, we assess whether such shifts warrant a formal claim of repair\. In multiclass action spaces, transitions do not collapse into binary shifts\. Hence, licensing a repair requires clearing three independent, non\-trivial hurdles beyond raw steerability: \(1\)destination correctness\(reaching gold target\), \(2\)preservation\(sparing baseline\-correct decisions\), \(3\)evidential sufficiency\(requiring the confidence interval, not only the point estimate, to clear the prospective threshold\)\. Crucially, each criterion can fail independently even when prior conditions are met\.

### 5\.1Decoupling Correction from Destination Correctness

Auditing Multiclass Exits and Lateral Error Mass Redistribution\.Row\-level destination audits reveal substantial lateral probability leakage into non\-target error classes across models\. On Phi\-3\.5\-mini, 153 of 200 routed errors successfully vacate the source state \(Xc=153X\_\{c\}=153\); however, while 107 reach the gold target action, 46 spill into alternative error classes, yielding a target\-hit rate ofTHc=0\.699\\mathrm\{TH\}\_\{c\}=0\.699\(Fig\.[4](https://arxiv.org/html/2609.36138#S5.F4)a\)\. On Qwen3\-8B, internal activation steering and external logit shifting achieve near\-identical top\-line gains \(\+38\+38vs\.\+35\+35net gain; 38 vs\. 37 gold arrivals; Fig\.[5](https://arxiv.org/html/2609.36138#S5.F5)a\) but diverge in destination topology: internal steering misdirects only 14 of 52 exits into alternative errors \(THc=0\.731\\mathrm\{TH\}\_\{c\}=0\.731\), whereas logit shifting diverts 21 of 58 exits into off\-target failures \(THc=0\.638\\mathrm\{TH\}\_\{c\}=0\.638\)\.

Vacating Error States Does Not Imply Target Convergence\.Exiting an error state does not guarantee ground\-truth convergence in multiclass settings\. On Phi\-3\.5\-mini, nearly a third of exits \(46 of 153\) drift into alternative errors rather than gold target, i\.e\., a defect entirely concealed by aggregate metrics\. On Qwen3\-8B, matching headline gains obscure diverging destination profiles \(a descriptive contrast asserting no formal ordering\)\. By auditing target\-hit rate,SAKIKOformally decouples source departure from verified target arrival, resolving what standard metrics conflate\.

Figure 5:Destination correctness and preservation are separate properties\.\(a\)On Qwen3\-8B, activation steering and logit shifting yield similar target arrivals but differ descriptively in lateral error spillage \(14 vs\. 21 exits;THc=0\.731\\mathrm\{TH\}\_\{c\}=0\.731vs\.0\.6380\.638\), without surviving multiplicity correction \(McNemarp=1\.0p=1\.0andp=0\.189p=0\.189\)\.\(b\)Collateral damage across denominators vs\. a5%5\\%tolerance: Phi\-3\.5\-mini incurs52/9352/93exposed \(E2\\mathrm\{E2\}\) and52/26452/264overall \(E1\\mathrm\{E1\}\) breaks; Qwen3\-8B incurs0/2110/211overall \(6 exposed\); Qwen3\-4B incurs1/501/50exposed; and Gemma\-2\-9B incurs1/111/11exposed \(E2=9\.1%\>5%\\mathrm\{E2\}=9\.1\\%\>5\\%\)\. By construction,E2≥E1\\mathrm\{E2\}\\geq\\mathrm\{E1\}\. \(See Figure[4](https://arxiv.org/html/2609.36138#S5.F4)b for licensing intervals\)\.
### 5\.2Decoupling Destination Correctness from Preservation

Collateral Breakdown Profiles and Denominator\-Dependent Degradation\.Audits of baseline\-correct instances expose severe collateral damage unmitigated by router gating\. On Phi\-3\.5\-mini, steering corrupts 52 of 93 exposed correct decisions \(Fig\.[4](https://arxiv.org/html/2609.36138#S5.F4)a\), yielding acute \(E2=55\.9%\\mathrm\{E2\}=55\.9\\%\) and population \(E1=19\.7%\\mathrm\{E1\}=19\.7\\%\) losses far exceeding tolerance \(ϵtol=0\.05\\epsilon\_\{\\mathrm\{tol\}\}=0\.05\)\. Router confidence fails to separate broken from intact instances \(Cliff’sδ=−0\.094\\delta=\-0\.094,p=0\.44p=0\.44\), leaving break rates\>50%\>50\\%even at high thresholds \(τc=0\.9\\tau\_\{c\}=0\.9\)\. Moreover, safety conclusions hinge critically on denominator selection: on Qwen3\-4B, a benign population loss \(E1=1\.3%\\mathrm\{E1\}=1\.3\\%\) conceals acute localized degradation \(E2=5\.7%\\mathrm\{E2\}=5\.7\\%\), whereas on Qwen3\-8B, zero breaks produce an uninformative exposed bound \(39\.3%39\.3\\%\) owing to sparse router exposure \(\|Cexp\|=6\|\{\}C\_\{\\mathrm\{exp\}\}\|\{\}=6\)\.

The Orthogonality of Target Steering and Behavioral Invariance\.Successful steering on error instances provides no inherent safeguard against corrupting clean inputs\. Latent overlap between under\- and fully specified prompts causes inference\-time routers to misallocate steering vectors, and raising confidence thresholds fails to mitigate break rates among exposed decisions\. Because breaks fall entirely within exposed instances, population lossE1\\mathrm\{E1\}understates localized risk by\|C\|/\|Cexp\|\|\{\}C\|\{\}/\|\{\}C\_\{\\mathrm\{exp\}\}\|\{\}, necessitating reporting bothE1\\mathrm\{E1\}and exposed lossE2\\mathrm\{E2\}alongside raw counts\. However, because preregistered protocols formally fixed preservation onE1\\mathrm\{E1\}\(clean\_collateral\_rateon Qwen3\-8B, with exposure\-conditional metrics added only in subsequent runs\), we reportE2\\mathrm\{E2\}without retrofitting frozen rules\. Consequently, the resulting repair license guarantees preservation strictly at the population level, not conditional on router exposure\.

### 5\.3Prospective Licensing and Evidential Sufficiency

Empirical Adjudication Under Pre\-Registered Statistical Gating\.SAKIKO’s frozen ten\-condition gate dissociates nominal point estimates from statistically robust repairs\. Qwen3\-8B clears all criteria for anADMIT, maintaining confidence bounds strictly above required thresholds \(THc=0\.731\\mathrm\{TH\}\_\{c\}=0\.731, CI\[0\.604,0\.846\]\[0\.604,0\.846\];TGc=0\.276\\mathrm\{TG\}\_\{c\}=0\.276, CI\[0\.115,0\.425\]\[0\.115,0\.425\]\) with zero breaks \(B=0B=0\)\. In contrast, Qwen3\-4B and Gemma\-2\-9B receiveDECLINEverdicts: despite positive point metrics, their bootstrap intervals cross sub\-threshold bounds on target\-hit rate and span zero on target gain \(TGc\\mathrm\{TG\}\_\{c\}CIs\[−0\.048,0\.202\]\[\-0\.048,0\.202\]and\[−0\.031,0\.177\]\[\-0\.031,0\.177\], respectively\)\. On Gemma\-2\-9B, wide intervals breach Conditions 3 and 6, while an acute exposed loss ofE2=9\.1%\\mathrm\{E2\}=9\.1\\%\(1/111/11breaks\) underscores that a nominal preservation pass does not guarantee deployment safety\.

Bounding Optimism via Finite\-Sample Statistical Power\.Finite\-sample point estimates cannot prove claims of mechanistic repair\. Due to limited statistical power, resolving a modest hit rate of0\.580\.58against a0\.500\.50null requires245245exits \(versus∼30\{\\sim\}30for0\.730\.73\), reflecting inadequate precision rather than guaranteed admission under larger samples\. Both model declines persist even if hit\-rate gates are relaxed, as target\-gain intervals cross zero and exposed collateral damage remains uncontained\. By mandating that confidence intervals, not isolated point estimates, satisfy preregistered safety criteria,SAKIKOensures small\-sample optimism is never mistaken for validated repair\.

Figure 6:One question under three input variants \(panels A–C\)\. Open circles and solid dots indicate pre\- and post\-intervention decisions, showing an identical shift across panels \(tool\_call→\\rightarrowrequest\_for\_info\)\. Filled labels denote the ground\-truth target; in panel B, the post\-intervention dot lands above the target, capturing an exit without target arrival\. The bottom bands contrast readout resolutions: transition is unsegmented; accuracy captures A and C but misses B \(consistently incorrect\); destination resolves all three distinct outcomes\. Visualizations are illustrative; counts represent frozen outcome\-class totals rather than specific transition frequencies\.
### 5\.4One Transition, Three Verdicts

Uniform Output Transitions Across Heterogeneous Task Demands\.On When2Call, an identical intervention shift, i\.e\., from invokingsearch\_flightsto requesting information, yields three distinct outcome classes depending on context, illustrated for Phi\-3\.5\-mini in Fig\.[6](https://arxiv.org/html/2609.36138#S5.F6)\. In Case A \(missing origin\), the model suppresses a hallucinated departure city to ask for clarification, achieving a target repair \(GOLD ARRIVAL, 107 instances\)\. In Case B \(unserviceable status inquiry\), it leaves the tool call but requests an unneeded flight number instead of refusing, shifting laterally into a new error \(OTHER WRONG, 46 instances\)\. In Case C \(fully specified query\), an already\-correct tool invocation is needlessly routed and broken into redundant clarification \(BROKEN, 52 of 93 exposed instances\)\. Standard accuracy tallies these identical shifts as\+1\+1,±0\\pm 0, and−1\-1, collapsing qualitatively divergent behaviors into an aggregate net gain of\+55\+55\(107−52107\-52\)\.

Representational Opacity of Behavioral Uniformity\.Uniform action\-level shifts mask fundamentally divergent epistemic outcomes, e\.g\., spanning true repair, lateral error drift, and collateral corruption\. Standard aggregate accuracy cannot separate these phenomena: it ignores lateral errors \(Case B\) while permitting true repairs \(Case A\) and broken decisions \(Case C\) to neutralize each other\. Because top\-line metrics conceal lateral probability redistribution and collateral damage, validating representation repair requires row\-level, destination\-resolved auditing across entireKK\-way action space\.

## 6Conclusion

In autonomous agents, pre\-execution action errors represent silent failures rooted in internal representations\. We introducedSAKIKO, a destination\-resolved framework that formally tests when representation interventions constitute genuine repair\. While channel\-keyed steering induces direction\-specific corrections across five architectures, conventional net gains conceal crucial failure modes: transitions frequently misroute into lateral errors, degrade baseline\-correct decisions, and fail formal confidence bounds, i\.e\., declining two of three sealed models on destination intervals alone despite favorable point estimates\. By pairing outcome\-resolved tracking with preregistered statistical gating, SAKIKO separates certified interventions from cosmetic behavioral drift\. Crucially, this license certifies only categorical action\-mode selection rather than downstream execution validity or exposed decision safety \(App\.[A](https://arxiv.org/html/2609.36138#A1)\), establishing that rigorous behavioral repair requires auditing destinations, collateral damage, and uncertainty alongside top\-line gains\.

## Reproducibility Statement

Appendices[G](https://arxiv.org/html/2609.36138#A7)and[D](https://arxiv.org/html/2609.36138#A4); Appendix[R](https://arxiv.org/html/2609.36138#A18)gives the artifact locations, the exact invocation, the determinism controls and the Git LFS requirement\. The sealed evaluations were each executed once under a preregistered gate, and the one\-shot ledger is recorded in Appendix[P](https://arxiv.org/html/2609.36138#A16)\. The code is available in[https://github\.com/ruizheliUOA/mechanistic\-tool\-use\-llm](https://github.com/ruizheliUOA/mechanistic-tool-use-llm)\.

## References

- Abdinet al\.\(2024\)M\. Abdin, J\. Aneja, H\. Awadalla, A\. Awadallah, A\. A\. Awan, N\. Bach, A\. Bahree, A\. Bakhtiari, J\. Bao, H\. Behl, A\. Benhaim, M\. Bilenko, J\. Bjorck, S\. Bubeck, M\. Cai, Q\. Cai, V\. Chaudhary, D\. Chen, D\. Chen, W\. Chen, Y\. Chen, Y\. Chen, H\. Cheng, P\. Chopra, X\. Dai, M\. Dixon, R\. Eldan, V\. Fragoso, J\. Gao, M\. Gao, M\. Gao, A\. Garg, A\. D\. Giorno, A\. Goswami, S\. Gunasekar, E\. Haider, J\. Hao, R\. J\. Hewett, W\. Hu, J\. Huynh, D\. Iter, S\. A\. Jacobs, M\. Javaheripi, X\. Jin, N\. Karampatziakis, P\. Kauffmann, M\. Khademi, D\. Kim, Y\. J\. Kim, L\. Kurilenko, J\. R\. Lee, Y\. T\. Lee, Y\. Li, Y\. Li, C\. Liang, L\. Liden, X\. Lin, Z\. Lin, C\. Liu, L\. Liu, M\. Liu, W\. Liu, X\. Liu, C\. Luo, P\. Madan, A\. Mahmoudzadeh, D\. Majercak, M\. Mazzola, C\. C\. T\. Mendes, A\. Mitra, H\. Modi, A\. Nguyen, B\. Norick, B\. Patra, D\. Perez\-Becker, T\. Portet, R\. Pryzant, H\. Qin, M\. Radmilac, L\. Ren, G\. de Rosa, C\. Rosset, S\. Roy, O\. Ruwase, O\. Saarikivi, A\. Saied, A\. Salim, M\. Santacroce, S\. Shah, N\. Shang, H\. Sharma, Y\. Shen, S\. Shukla, X\. Song, M\. Tanaka, A\. Tupini, P\. Vaddamanu, C\. Wang, G\. Wang, L\. Wang, S\. Wang, X\. Wang, Y\. Wang, R\. Ward, W\. Wen, P\. Witte, H\. Wu, X\. Wu, M\. Wyatt, B\. Xiao, C\. Xu, J\. Xu, W\. Xu, J\. Xue, S\. Yadav, F\. Yang, J\. Yang, Y\. Yang, Z\. Yang, D\. Yu, L\. Yuan, C\. Zhang, C\. Zhang, J\. Zhang, L\. L\. Zhang, Y\. Zhang, Y\. Zhang, Y\. Zhang, and X\. ZhouPhi\-3 technical report: a highly capable language model locally on your phone\.External Links:2404\.14219,[Link](https://arxiv.org/abs/2404.14219)Cited by:[§4\.1](https://arxiv.org/html/2609.36138#S4.SS1.p2.1)\.
- Agarwalet al\.\(2026\)M\. Agarwal, I\. Abdelaziz, K\. Basu, M\. Unuvar, L\. A\. Lastras, Y\. Rizk, and P\. KapanipathiToolRM: outcome reward models for tool\-calling large language models\.External Links:2509\.11963,[Link](https://arxiv.org/abs/2509.11963)Cited by:[§C\.4](https://arxiv.org/html/2609.36138#A3.SS4.SSS0.Px1.p1.1)\.
- Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. D\. Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. MaThe llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[§4\.1](https://arxiv.org/html/2609.36138#S4.SS1.p2.1)\.
- Healyet al\.\(2026\)K\. Healy, B\. Srinivasan, V\. Madathil, and J\. WuInternal representations as indicators of hallucinations in agent tool selection\.External Links:2601\.05214,[Link](https://arxiv.org/abs/2601.05214)Cited by:[§C\.2](https://arxiv.org/html/2609.36138#A3.SS2.p2.1),[§C\.4](https://arxiv.org/html/2609.36138#A3.SS4.p1.1),[Table 3](https://arxiv.org/html/2609.36138#A3.T3.10.3.1.1.1),[§2](https://arxiv.org/html/2609.36138#S2.p2.1),[§2](https://arxiv.org/html/2609.36138#S2.p3.1)\.
- Huanget al\.\(2024\)Y\. Huang, J\. Shi, Y\. Li, C\. Fan, S\. Wu, Q\. Zhang, Y\. Liu, P\. Zhou, Y\. Wan, N\. Z\. Gong, and L\. SunMetaTool benchmark for large language models: deciding whether to use tools and which to use\.External Links:2310\.03128,[Link](https://arxiv.org/abs/2310.03128)Cited by:[§4\.1](https://arxiv.org/html/2609.36138#S4.SS1.p1.1)\.
- Jiet al\.\(2023\)Z\. Ji, N\. Lee, R\. Frieske, T\. Yu, D\. Su, Y\. Xu, E\. Ishii, Y\. J\. Bang, A\. Madotto, and P\. FungSurvey of hallucination in natural language generation\.ACM Comput\. Surv\.55\(12\)\.External Links:ISSN 0360\-0300,[Link](https://doi.org/10.1145/3571730),[Document](https://dx.doi.org/10.1145/3571730)Cited by:[§C\.2](https://arxiv.org/html/2609.36138#A3.SS2.p1.1),[§1](https://arxiv.org/html/2609.36138#S1.p1.1)\.
- Jianget al\.\(2023\)A\. Q\. Jiang, A\. Sablayrolles, A\. Mensch, C\. Bamford, D\. S\. Chaplot, D\. de las Casas, F\. Bressand, G\. Lengyel, G\. Lample, L\. Saulnier, L\. R\. Lavaud, M\. Lachaux, P\. Stock, T\. L\. Scao, T\. Lavril, T\. Wang, T\. Lacroix, and W\. E\. SayedMistral 7b\.External Links:2310\.06825,[Link](https://arxiv.org/abs/2310.06825)Cited by:[§4\.1](https://arxiv.org/html/2609.36138#S4.SS1.p2.1)\.
- Jianget al\.\(2026\)E\. Jiang, A\. Gjølbye, Y\. J\. Zhang, and S\. KoyejoWhen behavioral safety evaluation fails: a representation\-level perspective\.External Links:2606\.08044,[Link](https://arxiv.org/abs/2606.08044)Cited by:[§C\.2](https://arxiv.org/html/2609.36138#A3.SS2.p3.1),[§2](https://arxiv.org/html/2609.36138#S2.p3.1)\.
- Laskaret al\.\(2026\)M\. T\. R\. Laskar, X\. Fu, S\. S\. Sarfjoo, Q\. McNamara, J\. Robertson, and S\. B\. TNFrom text to voice: a reproducible and verifiable framework for evaluating tool calling llm agents\.External Links:2605\.15104,[Link](https://arxiv.org/abs/2605.15104)Cited by:[§C\.4](https://arxiv.org/html/2609.36138#A3.SS4.SSS0.Px1.p1.1)\.
- Leeet al\.\(2025\)B\. W\. Lee, I\. Padhi, K\. N\. Ramamurthy, E\. Miehling, P\. Dognin, M\. Nagireddy, and A\. DhurandharProgramming refusal with conditional activation steering\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=Oi47wc10sm)Cited by:[§C\.3](https://arxiv.org/html/2609.36138#A3.SS3.p1.1),[Table 3](https://arxiv.org/html/2609.36138#A3.T3.10.6.1.1.1),[§1](https://arxiv.org/html/2609.36138#S1.p2.1),[§2](https://arxiv.org/html/2609.36138#S2.p2.1),[§3](https://arxiv.org/html/2609.36138#S3.p1.1)\.
- Liet al\.\(2026a\)D\. Li, Y\. Yao, Z\. Tan, H\. Liu, and R\. GuoToolPRMBench: evaluating and advancing process reward models for tool\-using agents\.InFindings of the Association for Computational Linguistics: ACL 2026,M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),San Diego, California, United States,pp\. 12378–12391\.External Links:[Link](https://aclanthology.org/2026.findings-acl.602/),[Document](https://dx.doi.org/10.18653/v1/2026.findings-acl.602),ISBN 979\-8\-89176\-395\-1Cited by:[§C\.4](https://arxiv.org/html/2609.36138#A3.SS4.SSS0.Px1.p1.1)\.
- Liet al\.\(2026b\)R\. Li, C\. Chen, Y\. Hu, Y\. Gao, X\. Wang, and E\. YilmazAttributing response to context: a jensen–shannon divergence driven mechanistic study of context attribution in retrieval\-augmented generation\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=7bHjHAOXH9)Cited by:[§C\.2](https://arxiv.org/html/2609.36138#A3.SS2.p1.1)\.
- Liuet al\.\(2024\)H\. Liu, Z\. Dou, Y\. Wang, N\. Peng, and Y\. YueUncertainty calibration for tool\-using language agents\.InFindings of the Association for Computational Linguistics: EMNLP 2024,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 16781–16805\.External Links:[Link](https://aclanthology.org/2024.findings-emnlp.978/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.978)Cited by:[§D\.3](https://arxiv.org/html/2609.36138#A4.SS3.p1.2)\.
- Liuet al\.\(2026\)X\. Liu, Y\. E\. Zhang, V\. Kasprova, P\. Rabbani, P\. S\. Zahraei, T\. Zhang, A\. Ebrahimpour\-Boroojeny, and V\. ChandrasekaranAgentAbstain: do llm agents know when not to act?\.External Links:2607\.10059,[Link](https://arxiv.org/abs/2607.10059)Cited by:[§C\.1](https://arxiv.org/html/2609.36138#A3.SS1.p3.1)\.
- Luoet al\.\(2026\)Z\. Luo, T\. P\. Kutralingam, O\. N\. Okoani, W\. Xu, H\. Wei, and X\. HuLost in execution: on the multilingual robustness of tool calling in large language models\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),San Diego, California, United States,pp\. 44059–44077\.External Links:[Link](https://aclanthology.org/2026.acl-long.2039/),[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.2039),ISBN 979\-8\-89176\-390\-6Cited by:[§C\.4](https://arxiv.org/html/2609.36138#A3.SS4.SSS0.Px1.p1.1)\.
- Patilet al\.\(2024\)S\. G\. Patil, T\. Zhang, X\. Wang, and J\. E\. GonzalezGorilla: large language model connected with massive apis\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 126544–126565\.External Links:[Document](https://dx.doi.org/10.52202/079017-4020),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/e4c61f578ff07830f5c37378dd3ecb0d-Paper-Conference.pdf)Cited by:[§C\.1](https://arxiv.org/html/2609.36138#A3.SS1.p1.1),[§1](https://arxiv.org/html/2609.36138#S1.p2.1),[§2](https://arxiv.org/html/2609.36138#S2.p1.1)\.
- Qinet al\.\(2024a\)Y\. Qin, S\. Hu, Y\. Lin, W\. Chen, N\. Ding, G\. Cui, Z\. Zeng, X\. Zhou, Y\. Huang, C\. Xiao, C\. Han, Y\. R\. Fung, Y\. Su, H\. Wang, C\. Qian, R\. Tian, K\. Zhu, S\. Liang, X\. Shen, B\. Xu, Z\. Zhang, Y\. Ye, B\. Li, Z\. Tang, J\. Yi, Y\. Zhu, Z\. Dai, L\. Yan, X\. Cong, Y\. Lu, W\. Zhao, Y\. Huang, J\. Yan, X\. Han, X\. Sun, D\. Li, J\. Phang, C\. Yang, T\. Wu, H\. Ji, G\. Li, Z\. Liu, and M\. SunTool learning with foundation models\.ACM Comput\. Surv\.57\(4\)\.External Links:ISSN 0360\-0300,[Link](https://doi.org/10.1145/3704435),[Document](https://dx.doi.org/10.1145/3704435)Cited by:[§C\.1](https://arxiv.org/html/2609.36138#A3.SS1.p1.1),[§1](https://arxiv.org/html/2609.36138#S1.p2.1),[§2](https://arxiv.org/html/2609.36138#S2.p1.1)\.
- Qinet al\.\(2024b\)Y\. Qin, S\. Liang, Y\. Ye, K\. Zhu, L\. Yan, Y\. Lu, Y\. Lin, X\. Cong, X\. Tang, B\. Qian, S\. Zhao, L\. Hong, R\. Tian, R\. Xie, J\. Zhou, M\. Gerstein, dahai li, Z\. Liu, and M\. SunToolLLM: facilitating large language models to master 16000\+ real\-world APIs\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=dHng2O0Jjr)Cited by:[§C\.1](https://arxiv.org/html/2609.36138#A3.SS1.p1.1),[§1](https://arxiv.org/html/2609.36138#S1.p2.1),[§2](https://arxiv.org/html/2609.36138#S2.p1.1)\.
- Qwenet al\.\(2025\)Qwen, :, A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, H\. Wei, H\. Lin, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Lin, K\. Dang, K\. Lu, K\. Bao, K\. Yang, L\. Yu, M\. Li, M\. Xue, P\. Zhang, Q\. Zhu, R\. Men, R\. Lin, T\. Li, T\. Tang, T\. Xia, X\. Ren, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Cui, Z\. Zhang, and Z\. QiuQwen2\.5 technical report\.External Links:2412\.15115,[Link](https://arxiv.org/abs/2412.15115)Cited by:[§4\.1](https://arxiv.org/html/2609.36138#S4.SS1.p2.1)\.
- Rosset al\.\(2025\)H\. Ross, A\. S\. Mahabaleshwarkar, and Y\. SuharaWhen2Call: when \(not\) to call tools\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),L\. Chiruzzo, A\. Ritter, and L\. Wang \(Eds\.\),Albuquerque, New Mexico,pp\. 3391–3409\.External Links:[Link](https://aclanthology.org/2025.naacl-long.174/),[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.174),ISBN 979\-8\-89176\-189\-6Cited by:[Appendix A](https://arxiv.org/html/2609.36138#A1.p2.1),[§4\.1](https://arxiv.org/html/2609.36138#S4.SS1.p1.1)\.
- Schicket al\.\(2023\)T\. Schick, J\. Dwivedi\-Yu, R\. Dessi, R\. Raileanu, M\. Lomeli, E\. Hambro, L\. Zettlemoyer, N\. Cancedda, and T\. ScialomToolformer: language models can teach themselves to use tools\.InThirty\-seventh Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=Yacmpz84TH)Cited by:[§C\.1](https://arxiv.org/html/2609.36138#A3.SS1.p1.1),[§1](https://arxiv.org/html/2609.36138#S1.p2.1),[§2](https://arxiv.org/html/2609.36138#S2.p1.1)\.
- Shiet al\.\(2026\)W\. Shi, Z\. Peng, S\. Li, X\. Wang, X\. Wang, M\. Du, and N\. ZouTo call or not to call: diagnosing intrinsic over\-calling bias in LLM agents\.External Links:2605\.18882,[Link](https://arxiv.org/abs/2605.18882)Cited by:[§C\.2](https://arxiv.org/html/2609.36138#A3.SS2.p3.1),[Table 3](https://arxiv.org/html/2609.36138#A3.T3.10.10.1.1.1),[§2](https://arxiv.org/html/2609.36138#S2.p2.1),[§2](https://arxiv.org/html/2609.36138#S2.p3.1)\.
- Suriet al\.\(2026\)M\. Suri, P\. Mathur, N\. Lipka, F\. Dernoncourt, R\. A\. Rossi, and D\. ManochaStructured uncertainty guided clarification for LLM agents\.InFindings of the Association for Computational Linguistics: ACL 2026,M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),San Diego, California, United States,pp\. 40811–40838\.External Links:[Link](https://aclanthology.org/2026.findings-acl.2028/),[Document](https://dx.doi.org/10.18653/v1/2026.findings-acl.2028),ISBN 979\-8\-89176\-395\-1Cited by:[§C\.1](https://arxiv.org/html/2609.36138#A3.SS1.p3.1)\.
- Tanget al\.\(2023\)Q\. Tang, Z\. Deng, H\. Lin, X\. Han, Q\. Liang, B\. Cao, and L\. SunToolAlpaca: generalized tool learning for language models with 3000 simulated cases\.External Links:2306\.05301,[Link](https://arxiv.org/abs/2306.05301)Cited by:[§C\.1](https://arxiv.org/html/2609.36138#A3.SS1.p1.1),[§1](https://arxiv.org/html/2609.36138#S1.p2.1),[§2](https://arxiv.org/html/2609.36138#S2.p1.1)\.
- Tatsat and Shater \(2026\)H\. Tatsat and A\. ShaterBeyond the black box: interpretability of agentic ai tool use\.External Links:2605\.06890,[Link](https://arxiv.org/abs/2605.06890)Cited by:[§C\.2](https://arxiv.org/html/2609.36138#A3.SS2.p2.1),[§C\.4](https://arxiv.org/html/2609.36138#A3.SS4.p1.1),[Table 3](https://arxiv.org/html/2609.36138#A3.T3.10.4.1.1.1),[§2](https://arxiv.org/html/2609.36138#S2.p2.1)\.
- Teamet al\.\(2024\)G\. Team, M\. Riviere, S\. Pathak, P\. G\. Sessa, C\. Hardin, S\. Bhupatiraju, L\. Hussenot, T\. Mesnard, B\. Shahriari, A\. Ramé, J\. Ferret, P\. Liu, P\. Tafti, A\. Friesen, M\. Casbon, S\. Ramos, R\. Kumar, C\. L\. Lan, S\. Jerome, A\. Tsitsulin, N\. Vieillard, P\. Stanczyk, S\. Girgin, N\. Momchev, M\. Hoffman, S\. Thakoor, J\. Grill, B\. Neyshabur, O\. Bachem, A\. Walton, A\. Severyn, A\. Parrish, A\. Ahmad, A\. Hutchison, A\. Abdagic, A\. Carl, A\. Shen, A\. Brock, A\. Coenen, A\. Laforge, A\. Paterson, B\. Bastian, B\. Piot, B\. Wu, B\. Royal, C\. Chen, C\. Kumar, C\. Perry, C\. Welty, C\. A\. Choquette\-Choo, D\. Sinopalnikov, D\. Weinberger, D\. Vijaykumar, D\. Rogozińska, D\. Herbison, E\. Bandy, E\. Wang, E\. Noland, E\. Moreira, E\. Senter, E\. Eltyshev, F\. Visin, G\. Rasskin, G\. Wei, G\. Cameron, G\. Martins, H\. Hashemi, H\. Klimczak\-Plucińska, H\. Batra, H\. Dhand, I\. Nardini, J\. Mein, J\. Zhou, J\. Svensson, J\. Stanway, J\. Chan, J\. P\. Zhou, J\. Carrasqueira, J\. Iljazi, J\. Becker, J\. Fernandez, J\. van Amersfoort, J\. Gordon, J\. Lipschultz, J\. Newlan, J\. Ji, K\. Mohamed, K\. Badola, K\. Black, K\. Millican, K\. McDonell, K\. Nguyen, K\. Sodhia, K\. Greene, L\. L\. Sjoesund, L\. Usui, L\. Sifre, L\. Heuermann, L\. Lago, L\. McNealus, L\. B\. Soares, L\. Kilpatrick, L\. Dixon, L\. Martins, M\. Reid, M\. Singh, M\. Iverson, M\. Görner, M\. Velloso, M\. Wirth, M\. Davidow, M\. Miller, M\. Rahtz, M\. Watson, M\. Risdal, M\. Kazemi, M\. Moynihan, M\. Zhang, M\. Kahng, M\. Park, M\. Rahman, M\. Khatwani, N\. Dao, N\. Bardoliwalla, N\. Devanathan, N\. Dumai, N\. Chauhan, O\. Wahltinez, P\. Botarda, P\. Barnes, P\. Barham, P\. Michel, P\. Jin, P\. Georgiev, P\. Culliton, P\. Kuppala, R\. Comanescu, R\. Merhej, R\. Jana, R\. A\. Rokni, R\. Agarwal, R\. Mullins, S\. Saadat, S\. M\. Carthy, S\. Cogan, S\. Perrin, S\. M\. R\. Arnold, S\. Krause, S\. Dai, S\. Garg, S\. Sheth, S\. Ronstrom, S\. Chan, T\. Jordan, T\. Yu, T\. Eccles, T\. Hennigan, T\. Kocisky, T\. Doshi, V\. Jain, V\. Yadav, V\. Meshram, V\. Dharmadhikari, W\. Barkley, W\. Wei, W\. Ye, W\. Han, W\. Kwon, X\. Xu, Z\. Shen, Z\. Gong, Z\. Wei, V\. Cotruta, P\. Kirk, A\. Rao, M\. Giang, L\. Peran, T\. Warkentin, E\. Collins, J\. Barral, Z\. Ghahramani, R\. Hadsell, D\. Sculley, J\. Banks, A\. Dragan, S\. Petrov, O\. Vinyals, J\. Dean, D\. Hassabis, K\. Kavukcuoglu, C\. Farabet, E\. Buchatskaya, S\. Borgeaud, N\. Fiedel, A\. Joulin, K\. Kenealy, R\. Dadashi, and A\. AndreevGemma 2: improving open language models at a practical size\.External Links:2408\.00118,[Link](https://arxiv.org/abs/2408.00118)Cited by:[§4\.1](https://arxiv.org/html/2609.36138#S4.SS1.p2.1)\.
- Wanget al\.\(2026a\)H\. Wang, C\. Qian, M\. Li, J\. Qiu, B\. Xue, M\. Wang, H\. Ji, A\. Storkey, and K\. WongPosition: agent should invoke external tools only when epistemically necessary\.External Links:2506\.00886,[Link](https://arxiv.org/abs/2506.00886)Cited by:[§C\.1](https://arxiv.org/html/2609.36138#A3.SS1.p2.1),[Table 3](https://arxiv.org/html/2609.36138#A3.T3.10.2.1.1.1),[§2](https://arxiv.org/html/2609.36138#S2.p1.1)\.
- Wanget al\.\(2026b\)Y\. Wang, R\. Zhou, Y\. Ma, R\. Fu, J\. Liang, S\. Cao, M\. Huang, T\. Fang, and L\. PanASA: backbone\-training\-free representation engineering for tool\-calling agents\.External Links:2602\.04935,[Link](https://arxiv.org/abs/2602.04935)Cited by:[§C\.3](https://arxiv.org/html/2609.36138#A3.SS3.p2.1),[§C\.4](https://arxiv.org/html/2609.36138#A3.SS4.p1.1),[Table 3](https://arxiv.org/html/2609.36138#A3.T3.10.8.1.1.1),[§1](https://arxiv.org/html/2609.36138#S1.p2.1),[§2](https://arxiv.org/html/2609.36138#S2.p2.1),[§2](https://arxiv.org/html/2609.36138#S2.p3.1),[§3](https://arxiv.org/html/2609.36138#S3.p1.1)\.
- Weiet al\.\(2022\)J\. Wei, Y\. Tay, R\. Bommasani, C\. Raffel, B\. Zoph, S\. Borgeaud, D\. Yogatama, M\. Bosma, D\. Zhou, D\. Metzler, E\. H\. Chi, T\. Hashimoto, O\. Vinyals, P\. Liang, J\. Dean, and W\. FedusEmergent abilities of large language models\.Transactions on Machine Learning Research\.Note:Survey CertificationExternal Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=yzkSU5zdwD)Cited by:[§C\.1](https://arxiv.org/html/2609.36138#A3.SS1.p1.1),[§1](https://arxiv.org/html/2609.36138#S1.p2.1)\.
- Wuet al\.\(2026a\)Q\. Wu, S\. Das, M\. Amani, A\. Nag, S\. Lee, K\. P\. Gummadi, A\. Ravichander, and M\. B\. ZafarTo call or not to call: a framework to assess and optimize llm tool calling\.External Links:2605\.00737,[Link](https://arxiv.org/abs/2605.00737)Cited by:[§C\.2](https://arxiv.org/html/2609.36138#A3.SS2.p3.1),[Table 3](https://arxiv.org/html/2609.36138#A3.T3.10.9.1.1.1),[§2](https://arxiv.org/html/2609.36138#S2.p1.1)\.
- Wuet al\.\(2026b\)Z\. Wu, Z\. Wang, S\. Cho, Y\. Yang, A\. Koshiyama, S\. Bulathwela, and M\. Perez\-OrtizTool calling is linearly readable and steerable in language models\.InWorkshop on Failure Modes of Agentic AI at ICML 2026,External Links:[Link](https://openreview.net/forum?id=FnXL1RKC06)Cited by:[§C\.3](https://arxiv.org/html/2609.36138#A3.SS3.p2.1),[§C\.3](https://arxiv.org/html/2609.36138#A3.SS3.p3.1),[§C\.4](https://arxiv.org/html/2609.36138#A3.SS4.p1.1),[Table 3](https://arxiv.org/html/2609.36138#A3.T3.10.7.1.1.1),[§2](https://arxiv.org/html/2609.36138#S2.p2.1),[§2](https://arxiv.org/html/2609.36138#S2.p3.1)\.
- Yanet al\.\(2026\)L\. Yan, R\. Li, G\. Chen, Q\. Li, J\. Geng, W\. Li, L\. Wang, and C\. LyuSpurious rewards paradox: mechanistically understanding how RLVR activates memorization shortcuts in LLMs\.InForty\-third International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=SGUSUm2491)Cited by:[§C\.2](https://arxiv.org/html/2609.36138#A3.SS2.p1.1)\.
- Yanget al\.\(2025\)A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv, C\. Zheng, D\. Liu, F\. Zhou, F\. Huang, F\. Hu, H\. Ge, H\. Wei, H\. Lin, J\. Tang, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Zhou, J\. Lin, K\. Dang, K\. Bao, K\. Yang, L\. Yu, L\. Deng, M\. Li, M\. Xue, M\. Li, P\. Zhang, P\. Wang, Q\. Zhu, R\. Men, R\. Gao, S\. Liu, S\. Luo, T\. Li, T\. Tang, W\. Yin, X\. Ren, X\. Wang, X\. Zhang, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Wang, Z\. Cui, Z\. Zhang, Z\. Zhou, and Z\. QiuQwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§4\.1](https://arxiv.org/html/2609.36138#S4.SS1.p2.1)\.
- Zhaoet al\.\(2026\)K\. Zhao, J\. Li, H\. Shen, W\. Chow, L\. Li, H\. Song, L\. Kong, C\. Zhi, T\. Zhao, S\. Liu, and J\. YinCalibration is the bottleneck: an action\-class diagnostic of multi\-turn tool\-calling\.External Links:2609\.00949,[Link](https://arxiv.org/abs/2609.00949)Cited by:[§C\.2](https://arxiv.org/html/2609.36138#A3.SS2.p3.1),[Table 3](https://arxiv.org/html/2609.36138#A3.T3.10.5.1.1.1),[§2](https://arxiv.org/html/2609.36138#S2.p3.1)\.
- Zhouet al\.\(2026\)Y\. Zhou, L\. Zeng, X\. Lu, W\. Xie, D\. Liu, J\. Yan, and J\. ShaoExploring agentic tool\-calling decisions via uncertainty\-aligned reinforcement learning\.External Links:2606\.06976,[Link](https://arxiv.org/abs/2606.06976)Cited by:[§C\.1](https://arxiv.org/html/2609.36138#A3.SS1.p3.1)\.
- Zouet al\.\(2025\)A\. Zou, L\. Phan, S\. Chen, J\. Campbell, P\. Guo, R\. Ren, A\. Pan, X\. Yin, M\. Mazeika, A\. Dombrowski, S\. Goel, N\. Li, M\. J\. Byun, Z\. Wang, A\. Mallen, S\. Basart, S\. Koyejo, D\. Song, M\. Fredrikson, J\. Z\. Kolter, and D\. HendrycksRepresentation engineering: a top\-down approach to ai transparency\.External Links:2310\.01405,[Link](https://arxiv.org/abs/2310.01405)Cited by:[§C\.3](https://arxiv.org/html/2609.36138#A3.SS3.p1.1),[§2](https://arxiv.org/html/2609.36138#S2.p2.1)\.

## Appendix ALimitations

We explicitly delineate the boundaries of our empirical findings across five primary areas:

Benchmark Scope and External Transfer\.All formal confirmatory evaluations are established exclusively on When2Call\([Ross et al\., 2025](https://arxiv.org/html/2609.36138#bib.bib15)\)\. While MetaTool provides historical evidence in a binary setting where destination analysis is structurally inapplicable, ACEBench failed pre\-intervention channel\-support criteria and was retired prior to steering experiments\. Furthermore, because ground\-truth labels in When2Call are synthetically generated by its creators without post\-hoc re\-annotation, unmodeled systematic labeling biases could shift destination distributions without triggering gate violations\. Cross\-benchmark, multi\-modal, or real\-world policy transfer remains unestablished\.

Specificity and Scope of Licensed Claims\.Our positive repair claim rests upon a single admitted configuration: Qwen3\-8B on thecannot\_answer→\\rightarrowtool\_callchannel under gradient activation steering\. This outcome does not support general claims regarding the model family, other error transitions, or activation steering broadly\. Furthermore, activation interventions are not established as superior to score\-space comparators: paired tests after multiplicity correction remain non\-significant \(exact McNemarp=1\.0p=1\.0andp=0\.189p=0\.189\), and performance rankings between the two alternate across architectures\. Existing steering methods \(e\.g\., ASA, CAST\) were not re\-implemented as empirical baselines because our evaluation protocol was strictly locked prior to sealed execution\.

Exposure\-Conditional Preservation Deficits\.Although our pre\-registered gate bounded population\-level collateral damage \(E​1≤0\.05E1\\leq 0\.05\), no setting in this study certifies exposure\-conditional preservation \(E​2E2\)\. On Qwen3\-8B, zero breaks were observed over only six router\-exposed correct decisions, yielding an uninformative one\-sided95%95\\%upper bound of0\.3930\.393\(whereas certifying a5%5\\%bound requires at least 59 exposed correct instances\)\. Similarly, Gemma\-2\-9B corrupted11of1111exposed correct decisions \(E​2=0\.091E2=0\.091\) despite passing the overall population threshold\. Demonstrating localized safety on touched clean decisions remains open across all evaluated models\.

Sensitivity to Sample Support and Finite\-Sample Uncertainty\.In sealed evaluations, formal declines for Qwen3\-4B and Gemma\-2\-9B were driven entirely by bootstrap interval conditions rather than point\-estimate failures\. The outcome\-resolved apparatus provides the necessary statistical resolution to compute these intervals, confirming that small\-sample optimism cannot substitute for verified repair\. Similarly, the pre\-intervention halt of Qwen3\.5\-9B \(at 29 reference rows versus a 30\-instance threshold\) was split\-sensitive, clearing eligibility in56\.7%56\.7\\%of retrospective split reallocations rather than reflecting intrinsic uncorrectability\.

Mechanistic Scope and Procedural Transparency\.While we confirm site\- and direction\-specific causal steerability, our framework audits the empirical evidence pipeline rather than isolating causal circuits or tracing detailed propagation dynamics through model components\. Finally, to ensure complete provenance, we disclose that three procedural gate adjustments, one mechanical restart on Qwen3\-4B, and historical variations in random null sizes \(1 to 10 directions on early Qwen2\.5\-7B runs versus 59 in sealed settings\) occurred\. All events were cryptographically recorded with pre\-access hashes prior to endpoint unblinding without altering any frozen verdict\.

## Appendix BProblem Formalisation and Notation

This appendix fixes the populations and denominators used throughout the paper\. Ambiguity about which population a rate is computed over has been the single largest source of misreading in this line of work, so every quantity below is defined against a named set rather than against “the evaluation set”\.

### B\.1Actions, channels and populations

Let𝒴\\mathcal\{Y\}be the set of pre\-execution action modes available to the model\. On the primary benchmark\|𝒴\|=4\|\\mathcal\{Y\}\|=4: calling a tool, requesting missing information, answering directly, and declining\. For an inputxxwe writey⋆​\(x\)∈𝒴y^\{\\star\}\(x\)\\in\\mathcal\{Y\}for the reference action andy^0​\(x\)\\hat\{y\}\_\{0\}\(x\)for the action the unmodified model produces\. An input is a*baseline error*wheny^0≠y⋆\\hat\{y\}\_\{0\}\\neq y^\{\\star\}and*baseline\-correct*wheny^0=y⋆\\hat\{y\}\_\{0\}=y^\{\\star\}\.

Adirectional error channelis the ordered pair

c=\(g→s\),g=y⋆,s=y^0,g≠s,c\\;=\\;\\bigl\(g\\rightarrow s\\bigr\),\\qquad g=y^\{\\star\},\\quad s=\\hat\{y\}\_\{0\},\\quad g\\neq s,\(1\)whereggis the required action andssthe source action actually produced\. Channels are properties of a model on a task: the transition carrying the most error mass differs between models, so the framework discovers them \(Appendix[D](https://arxiv.org/html/2609.36138#A4)\) rather than assuming a fixed set\.

Four populations matter, and they are not interchangeable \(Table[2](https://arxiv.org/html/2609.36138#A2.T2)collects every symbol with its denominator\):

Ec\\displaystyle E\_\{c\}=\{x:y^0\(x\)=s,y⋆\(x\)=g\}\\displaystyle=\\\{x:\\hat\{y\}\_\{0\}\(x\)=s,\\;y^\{\\star\}\(x\)=g\\\}channel errors,​Nc=\|Ec\|\\displaystyle\\text\{channel errors, \}N\_\{c\}=\|E\_\{c\}\|\(2\)Ecrt\\displaystyle E\_\{c\}^\{\\mathrm\{rt\}\}=\{x∈Ec:fire⁡\(x\)\}\\displaystyle=\\\{x\\in E\_\{c\}:\\mathrm\{fire\}\(x\)\\\}routed channel errors\(3\)C\\displaystyle C=\{x:y^0​\(x\)=y⋆​\(x\)\}\\displaystyle=\\\{x:\\hat\{y\}\_\{0\}\(x\)=y^\{\\star\}\(x\)\\\}all baseline\-correct decisions\(4\)Cexp\\displaystyle C\_\{\\mathrm\{exp\}\}=\{x∈C:fire⁡\(x\)\}\\displaystyle=\\\{x\\in C:\\mathrm\{fire\}\(x\)\\\}exposed \(at\-risk\) correct decisions\(5\)withEcrt⊆EcE\_\{c\}^\{\\mathrm\{rt\}\}\\subseteq E\_\{c\}andCexp⊆CC\_\{\\mathrm\{exp\}\}\\subseteq C\. The distinction between equation[2](https://arxiv.org/html/2609.36138#A2.E2)and equation[3](https://arxiv.org/html/2609.36138#A2.E3), and between equation[4](https://arxiv.org/html/2609.36138#A2.E4)and equation[5](https://arxiv.org/html/2609.36138#A2.E5), is what makes the rates below well defined\. On Qwen3\-8B, for instance,Nc=87N\_\{c\}=87while\|Ecrt\|=72\|E\_\{c\}^\{\\mathrm\{rt\}\}\|=72; a count reported “of 72” is not the same statement as the same count reported “of 87”\.

### B\.2Post\-intervention outcome classes

The intervention produces a second actiony^1\\hat\{y\}\_\{1\}\. Every decision in the two adjudicated cohorts — the errors of channelccand the baseline\-correct population — falls into exactly one of five classes, three on the error population and two on the correct one:

Nc=Sc\+Ac\+Oc⏟baseline errors,\|Cexp\|=R\+B⏟exposed correct,\\underbrace\{N\_\{c\}=S\_\{c\}\+A\_\{c\}\+O\_\{c\}\}\_\{\\text\{baseline errors\}\},\\qquad\\underbrace\{\|C\_\{\\mathrm\{exp\}\}\|=R\+B\}\_\{\\text\{exposed correct\}\},\(6\)whereScS\_\{c\}retain the source \(y^1=s\\hat\{y\}\_\{1\}=s\),AcA\_\{c\}arrive at the required action \(y^1=g\\hat\{y\}\_\{1\}=g\),OcO\_\{c\}exit to a third action \(y^1∉\{g,s\}\\hat\{y\}\_\{1\}\\notin\\\{g,s\\\}\),RRremain correct andBBare broken\. Decisions the Router declines are returned unchanged\. An unrouted channel error therefore retains its source and counts inScS\_\{c\}, which is defined over all ofNcN\_\{c\}; an unrouted correct decision is by definition outsideCexpC\_\{\\mathrm\{exp\}\}and counts in neitherRRnorBB, both of which are defined over the exposed cohort alone\. We writeXc=Ac\+OcX\_\{c\}=A\_\{c\}\+O\_\{c\}for the*source exits*\.

Table 1:The five mutually exclusive post\-intervention outcome classes\. Aggregate net gain \(G=∑cAc−BG=\\sum\_\{c\}A\_\{c\}\-B\) leaves exits to other incorrect actions \(OcO\_\{c\}\) unobserved and allows target arrivals \(AcA\_\{c\}\) and collateral breaks \(BB\) to artificially cancel\.
### B\.3Endpoints

Gwhole\\displaystyle G\_\{\\mathrm\{whole\}\}=Fixed−B,Fixed=\|\{x:y^0≠y⋆,y^1=y⋆\}\|\\displaystyle=\\text\{Fixed\}\-B,\\quad\\text\{Fixed\}=\\bigl\|\\\{x:\\hat\{y\}\_\{0\}\\neq y^\{\\star\},\\ \\hat\{y\}\_\{1\}=y^\{\\star\}\\\}\\bigr\|whole\-population net\(7\)Gchannel\\displaystyle G\_\{\\mathrm\{channel\}\}=∑cAc−B\\displaystyle=\\textstyle\\sum\_\{c\}A\_\{c\}\-Bchannel\-level net\(8\)THc\\displaystyle\\mathrm\{TH\}\_\{c\}=AcAc\+Oc=AcXc\\displaystyle=\\frac\{A\_\{c\}\}\{A\_\{c\}\+O\_\{c\}\}=\\frac\{A\_\{c\}\}\{X\_\{c\}\}target\-hit; undefined at​Xc=0\\displaystyle\\text\{target\-hit; undefined at \}X\_\{c\}=0\(9\)TGc\\displaystyle\\mathrm\{TG\}\_\{c\}=Ac−OcNc\\displaystyle=\\frac\{A\_\{c\}\-O\_\{c\}\}\{N\_\{c\}\}target gain, over*all*channel errors\(10\)E1\\displaystyle\\mathrm\{E1\}=B\|C\|,E2=B\|Cexp\|\\displaystyle=\\frac\{B\}\{\|C\|\},\\qquad\\mathrm\{E2\}=\\frac\{B\}\{\|C\_\{\\mathrm\{exp\}\}\|\}collateral, two denominators\(11\)
#### The two net conventions are distinct and are never substituted\.

Fixed\\mathrm\{Fixed\}counts every baseline error in the evaluation population that becomes correct;∑cAc\\sum\_\{c\}A\_\{c\}counts only arrivals inside the adjudicated channels\. The first exceeds the second whenever the intervention corrects an error the channel does not key, soGwhole≥GchannelG\_\{\\mathrm\{whole\}\}\\geq G\_\{\\mathrm\{channel\}\}, with equality only when no such error is corrected\. On Qwen3\-8B the two are\+40\+40and\+38\+38\(Table[10](https://arxiv.org/html/2609.36138#A8.T10)\)\. Every net quantity in this paper is labelledwholeorchannel\-leveland the two are never interchanged\.

Three consequences follow directly and are used throughout the paper\.

#### Net compresses five classes into two\.

OnlyAcA\_\{c\}andBBappear in equation[7](https://arxiv.org/html/2609.36138#A2.E7)and equation[8](https://arxiv.org/html/2609.36138#A2.E8)\.OcO\_\{c\}does not, because a move from one wrong action to another leaves accuracy unchanged\. A givenGGtherefore does not determineTHc\\mathrm\{TH\}\_\{c\}or the collateral: for anyGGand any rationalt∈\(0,1\]t\\in\(0,1\]whose denominator dividesAcA\_\{c\}, settingOc=Ac​\(1−t\)/tO\_\{c\}=A\_\{c\}\(1\-t\)/tgives target\-hit exactlyttat unchangedGG, and every such profile is admissible whenever the population is large enough to contain it\. The mapping from the five classes toGGis many\-to\-one; that, and not an absence of any constraint, is the formal content of the paper’s central claim\.

#### The two collateral denominators are ordered\.

Breaks occur only withinCexpC\_\{\\mathrm\{exp\}\}, andCexp⊆CC\_\{\\mathrm\{exp\}\}\\subseteq C, so

E2=B\|Cexp\|≥B\|C\|=E1,\\mathrm\{E2\}=\\frac\{B\}\{\|C\_\{\\mathrm\{exp\}\}\|\}\\;\\geq\\;\\frac\{B\}\{\|C\|\}=\\mathrm\{E1\},\(12\)with equality whenB=0B=0or when every correct decision is exposed, and strict inequality otherwise\. Reporting E1 alone understates the risk to a touched decision by the factor\|C\|/\|Cexp\|\|C\|/\|C\_\{\\mathrm\{exp\}\}\|\. The frozen licensing gate uses E1, the permissive denominator; we report both throughout, with their counts\.

#### Target\-hit is degenerate in a binary action space\.

If\|𝒴\|=2\|\\mathcal\{Y\}\|=2thenOc=0O\_\{c\}=0by construction andTHc≡1\\mathrm\{TH\}\_\{c\}\\equiv 1wheneverXc\>0X\_\{c\}\>0\. The destination question is therefore not posable on a binary benchmark, which is why the destination\-resolved results in this paper come from a multiclass action space \(Appendix[F](https://arxiv.org/html/2609.36138#A6)\)\.

Table 2:Notation and endpoint definitions\. The*population*column is the denominator; quoting a rate without it is ambiguous\. Populations are defined in equation[2](https://arxiv.org/html/2609.36138#A2.E2)–equation[5](https://arxiv.org/html/2609.36138#A2.E5)\.

## Appendix CExtended Related Work

We draw on three lines of work: tool\-augmented agents, interpretability of agent decisions, and representation\-level intervention\. In this section we survey each and delineate the structural distinctions that characterizeSAKIKO\.

### C\.1Tool\-Augmented Agents and the Epistemics of Action Selection

The integration of external tools into LLMs has significantly expanded their functional scope from static text generators to autonomous problem solvers\([Wei et al\., 2022](https://arxiv.org/html/2609.36138#bib.bib3);[Qin et al\., 2024a](https://arxiv.org/html/2609.36138#bib.bib2)\)\. Early and influential paradigms, such as Toolformer\([Schick et al\., 2023](https://arxiv.org/html/2609.36138#bib.bib4)\), Gorilla\([Patil et al\., 2024](https://arxiv.org/html/2609.36138#bib.bib7)\), ToolLLM\([Qin et al\., 2024b](https://arxiv.org/html/2609.36138#bib.bib5)\), and ToolAlpaca\([Tang et al\., 2023](https://arxiv.org/html/2609.36138#bib.bib6)\), focused predominantly on the syntactic and parametric execution of tool calls, i\.e\., improving how models format API queries, satisfy parameter schemas, and adapt to tool documentation via fine\-tuning or in\-context demonstration\.

However, before an agent can parameterize or execute a specific tool, it must resolve the fundamentally upstream decision ofwhetherexternal action is warranted at all\. A recent conceptual framework proposed by[Wang et al\. \(2026a\)](https://arxiv.org/html/2609.36138#bib.bib9)formalizes this criterion, positing that agents should invoke external tools only when epistemically necessary, i\.e\., when task uncertainty cannot be resolved through internal reasoning alone\. Deviations from this criterion give rise to failure regimes such as tool bypass \(unwarranted direct answering despite missing evidence\) or over\-delegation \(unnecessary external queries when internal knowledge suffices\)\. While[Wang et al\. \(2026a\)](https://arxiv.org/html/2609.36138#bib.bib9)conceptually formalize the normative decision boundary between internal reasoning and external interaction, they do not study the underlying internal representation geometry nor provide intervention mechanisms to correct miscalibrated tool decisions\.SAKIKOprovides an empirical and representation\-level operationalization of this boundary, identifying and intervening on the directional channels that mediate pre\-execution tool decisions\.

A parallel line treats the same decision as a policy to be trained rather than a representation to be edited\.[Zhou et al\. \(2026\)](https://arxiv.org/html/2609.36138#bib.bib30)optimise tool\-calling decisions by uncertainty\-aligned reinforcement learning over the same four next\-action choices, and[Suri et al\. \(2026\)](https://arxiv.org/html/2609.36138#bib.bib29)select clarification questions from a structured uncertainty estimate, reporting large When2Call gains from uncertainty\-weighted training\.[Liu et al\. \(2026\)](https://arxiv.org/html/2609.36138#bib.bib34)evaluate calibrated restraint directly, pairing should\-act with should\-abstain variants of the same task\. These are alternatives to post\-hoc intervention rather than comparators for it: they change the policy, whereSAKIKOleaves the weights frozen and asks what evidence an edit to a frozen model’s activations can support\. None of them resolves where a changed decision lands among the remaining actions, which is the accounting this paper adds\.

### C\.2Internal Observability and Detection of Tool\-Use Failures

Paralleling the rise of agent frameworks, recent work in mechanistic interpretability has begun investigating how internal representations reflect an agent’s operational state prior to execution\([Yan et al\., 2026](https://arxiv.org/html/2609.36138#bib.bib28)\)\. Traditional reliability taxonomies primarily address factual hallucinations in ungrounded text generation\([Ji et al\., 2023](https://arxiv.org/html/2609.36138#bib.bib1);[Li et al\., 2026b](https://arxiv.org/html/2609.36138#bib.bib17)\)\. In contrast, agentic settings introduce structured failures in action selection\.

To monitor these failures,[Healy et al\. \(2026\)](https://arxiv.org/html/2609.36138#bib.bib10)investigate internal representations during tool\-call generation, demonstrating that latent embeddings in transformer layers linearly encode tool\-selection and parameter hallucinations\. By training lightweight binary classifiers on contextualized hidden states, they show that impending tool errors can be flagged in real time before execution\. Concurrently,[Tatsat and Shater \(2026\)](https://arxiv.org/html/2609.36138#bib.bib11)construct an interpretability framework that couples sparse autoencoders \(SAEs\) with linear probes to monitor agent states before action\. Their architecture deploys a binary tool\-need probe to predict tool invocation and a ternary tool\-risk probe to estimate operational consequence, demonstrating that decision signals localize to specific sparse features and late transformer layers\.

While these approaches substantiate that an agent’s pre\-execution intent and validity are decodable from internal states, they remain strictly diagnostic: they focus on passive monitoring, risk scoring, and post\-hoc feature localization\. Furthermore, their formulation of error detection is typically collapsed into a binary classification \(e\.g\., correct vs\. hallucinated, or tool needed vs\. not needed\)\.[Zhao et al\. \(2026\)](https://arxiv.org/html/2609.36138#bib.bib25)are the exception, decomposing multi\-turn failures over a four\-class action space into action\-class miscalibration and execution error, but they diagnose the miscalibration without intervening on the states that produce it;[Wu et al\. \(2026a\)](https://arxiv.org/html/2609.36138#bib.bib24)do drive a controller from such readouts, but score the result by task performance rather than by where the redirected decision lands\.[Shi et al\. \(2026\)](https://arxiv.org/html/2609.36138#bib.bib26)go further and intervene, recovering a sparse\-autoencoder feature basis for the call/no\-call decision on When2Call, reducing it to a signed activation margin, and testing an activation\-independent CALL offset causally by a closed\-form counter\-bias shift along the decoder directions\. That study and this one share a benchmark and an intervention family and ask different questions: it diagnoses and cancels a scalar bias on a binary margin, where vacating the source leaves one destination, while the questions here — which of three or more remaining actions a corrected decision reaches, and what the correction costs on decisions that were already right — are not posable in that formulation\. Outside tool use,[Jiang et al\. \(2026\)](https://arxiv.org/html/2609.36138#bib.bib27)make the complementary point that a model can pass behavioural safety evaluation while remaining vulnerable to bounded latent perturbation, which is the same dissociation between an output\-level readout and an intervention\-level one that motivates our destination accounting\.SAKIKOdeparts from passive observability by developing active, targeted interventions\. More fundamentally,SAKIKOrecognizes that pre\-execution tool decisions operate in a multi\-way action space, where diagnostic flags alone cannot resolve which directional channel caused the failure or guide the model toward the correct target action\.

### C\.3Representation Engineering and Activation Steering in Agents

Representation engineering and activation steering have emerged as efficient, backbone\-training\-free techniques to manipulate model behaviour by perturbing intermediate residual states\([Zou et al\., 2025](https://arxiv.org/html/2609.36138#bib.bib12)\)\. In the context of alignment and behavioural control, Conditional Activation Steering \(CAST\)\([Lee et al\., 2025](https://arxiv.org/html/2609.36138#bib.bib13)\)demonstrates that refusal mechanisms can be gated conditionally using similarity metrics derived from prompt activations\.

Recently, representation steering has been extended directly to agent tool selection\.[Wu et al\. \(2026b\)](https://arxiv.org/html/2609.36138#bib.bib14)demonstrate that tool identity is linearly readable and steerable within the residual stream across a wide spectrum of open\-weight models\. By injecting the mean\-difference activation vector between two candidate tools, they show that an agent’s discrete tool choice can be flipped with high accuracy on single\-turn menus, with the autoregressively generated JSON arguments updating to match the schema of the new tool\. Similarly, the Activation Steering Adapter \(ASA\)\([Wang et al\., 2026b](https://arxiv.org/html/2609.36138#bib.bib8)\)applies router\-conditioned steering vectors at mid\-layers to counter the lazy agent problem, steering models out of inert states when tool invocation is warranted\.

Despite these empirical successes, the steering methodologies we survey operate under a crucial simplification: they evaluate steering as either a binary transition \(e\.g\., refusal vs\. compliance in CAST, or tool vs\. no\-tool in ASA\) or as an isolated pairwise swap between two designated tools\([Wu et al\., 2026b](https://arxiv.org/html/2609.36138#bib.bib14)\)\. In such constrained environments, vacating an incorrect state trivially coincides with arriving at the alternative candidate\. As we demonstrate, this assumption does not hold in the multiclass settings we evaluate\.

### C\.4Contrasting SAKIKO: From Behavioural Movement to Adjudicated Repair

Table 3:Comparison ofSAKIKOwith the agent interpretability and steering frameworks we survey\.Unlike the passive diagnostic tools and binary or pairwise steering methods listed here,SAKIKOaddresses multiclass action spaces, tracks post\-intervention destination distributions, verifies baseline preservation, and formally adjudicates repair claims under sample uncertainty\.*Dest\.*is destination verification and*Lic\.*is evidential licensing\. The four\-class row is diagnostic only: it decomposes failures without intervening, so no destination distribution is produced to verify\.Table[3](https://arxiv.org/html/2609.36138#A3.T3)highlights the structural position ofSAKIKOrelative to the works we survey\. While this research establishes that tool intent is linearly decodable\([Healy et al\., 2026](https://arxiv.org/html/2609.36138#bib.bib10);[Tatsat and Shater, 2026](https://arxiv.org/html/2609.36138#bib.bib11)\)and steerable\([Wang et al\., 2026b](https://arxiv.org/html/2609.36138#bib.bib8);[Wu et al\., 2026b](https://arxiv.org/html/2609.36138#bib.bib14)\), these studies evaluate intervention success through aggregate metric shifts or binary state exits\.SAKIKOaddresses three issues that the surveyed evaluations do not:

1. 1\.Multiclass Destination Divergence \(KK\-way Outcomes\):Real\-world pre\-execution agent decisions are not binary, and they span invoking specific tools, requesting user clarification, answering directly, or declining inappropriate prompts\. In aKK\-way action space, perturbing an internal activation away from an erroneous state does not guarantee arrival at the ground\-truth target\. An intervention can induce substantial behavioural movement while simply redirecting the model to an alternative, equally incorrect action \(e\.g\., misdirecting a tool\-bypass error into a refusal\)\.SAKIKOexplicitly tracks the full destination composition rather than treating state\-departure as repair\.
2. 2\.Decoupling Net Gain from Baseline Preservation:Prior steering works report aggregate improvements \(e\.g\., overall tool\-calling accuracy\) across evaluation sets\. However, as our experiments reveal \(e\.g\., on Phi\-3\.5\), an intervention can produce a net gain of\+55\+55while concurrently corrupting5252of the9393already\-correct decisions on which its detector fired\. By auditing both destination hit\-rates and collateral damage on the baseline\-correct inputs the intervention actually touches,SAKIKOensures that repair does not come at the expense of unmonitored capability degradation\.
3. 3\.Evidence\-Qualified Licensing under Uncertainty:In the works we survey, favourable point estimates are generally reported without an accompanying uncertainty criterion\. However, under finite evaluation budgets and high variance, point gains can be statistically illusory\.SAKIKOintroduces a structured adjudication protocol that combines destination verification with pre\-registered uncertainty criteria, formally declining repair claims when empirical sample support is insufficient\.

#### Adjacent evaluation settings\.

Three further lines bound what the present evaluation covers rather than compete with it\.[Agarwal et al\. \(2026\)](https://arxiv.org/html/2609.36138#bib.bib32)train outcome reward models for tool calling and introduce a reward benchmark for them, which scores whether a call was good rather than whether an action mode was the right one; linking pre\-execution repair to outcome quality is the natural next step and one this paper does not take\.[Li et al\. \(2026a\)](https://arxiv.org/html/2609.36138#bib.bib35)supply step\-level process labels over trajectories, a richer setting than the single pre\-execution decision audited here\.[Luo et al\. \(2026\)](https://arxiv.org/html/2609.36138#bib.bib36)show that tool\-calling failures shift systematically across languages, and[Laskar et al\. \(2026\)](https://arxiv.org/html/2609.36138#bib.bib33)that they shift again when the same benchmarks are rendered as speech — including When2Call, which carries every formal outcome in this paper\. Both mark a boundary on the generality of our results: nothing here establishes that the routing and collateral findings persist beyond English text prompts\.

Therefore, rather than viewing activation steering as a monolithic tool\-flipping switch,SAKIKOtreats it as one stage in anadjudicated repairprocess, progressing from directional error discovery and channel\-keyed correction to destination verification and evidence\-qualified licensing\.

## Appendix DThe SAKIKO Protocol in Full

The five stages named in Section[3](https://arxiv.org/html/2609.36138#S3)are the reader\-facing grouping; the protocol as executed runs eight, becauseDiscoverysplits into Discover and Support,Correctioninto Estimate and Intervene, andVerificationinto Verify and Control\. The six properties of the evidence ladder \(Appendix[E](https://arxiv.org/html/2609.36138#A5)\) are what those stages establish, not the stages themselves\.

In order,

Discover→Support→Detect→Estimate→Intervene→Verify→Control→Admit/Decline,\\text\{Discover\}\\rightarrow\\text\{Support\}\\rightarrow\\text\{Detect\}\\rightarrow\\text\{Estimate\}\\rightarrow\\text\{Intervene\}\\rightarrow\\text\{Verify\}\\rightarrow\\text\{Control\}\\rightarrow\\text\{Admit/Decline\},of which the first seven are offline and only Detect and Intervene have an inference\-time counterpart \(§[D\.6](https://arxiv.org/html/2609.36138#A4.SS6)\)\.

### D\.1Discover: channel construction

Baseline predictions are collected on the training split and grouped by the ordered pair\(y⋆,y^0\)\(y^\{\\star\},\\hat\{y\}\_\{0\}\), giving the full confusion topology rather than a preselected set of transitions\. Every off\-diagonal cell is a candidate channel\. The number of channels is a property of the model and task and is not fixed: the historical When2Call pipeline retains three channels on Phi\-3\.5 and Qwen2\.5\-7B, whereas each sealed evaluation adjudicates a single preselected channel,cannot\_answer→\\rightarrowtool\_call, disclosed before sealing\.

### D\.2Support: qualification before intervention

A candidate channel advances only if it can be estimated on training data and evaluated afterwards with enough rows to adjudicate\. The frozen sealed criterion is a minimum of3030channel errors in the evaluation population; the development\-stage gate applies a matching minimum of3030baseline\-correct reference rows per split, since a direction cannot be estimated without both sides of the channel\. Channels that fail are set aside as*unadjudicable*, not as uncorrectable: the stage reports a property of the available sample, not of the model\.

This gate is load\-bearing rather than decorative\. It stopped two settings in this paper before any intervention was run\. ACEBench produced a valid readout but no directional transition met the preregistered per\-split support minimum, and Qwen3\.5\-9B failed that reference\-side minimum with2929baseline\-correct rows against3030, on a channel that had ample errors \(Appendix[Q](https://arxiv.org/html/2609.36138#A17)\)\.

### D\.3Detect: observation site and Router

For a retained channelcc, a per\-channel Router is trained on the hidden statehobsh\_\{\\mathrm\{obs\}\}read at a single observation layerℓobs\\ell\_\{\\mathrm\{obs\}\}, taken at the last prompt position of the residual stream\. At inference it returns a channel, a confidence, and a firing decision:

r⁡\(hobs\)⟶\(c^,pc^,fire\),fire⇔pc^≥τc^,r\\\!\\left\(h\_\{\\mathrm\{obs\}\}\\right\)\\longrightarrow\\bigl\(\\hat\{c\},\\,p\_\{\\hat\{c\}\},\\,\\mathrm\{fire\}\\bigr\),\\qquad\\mathrm\{fire\}\\iff p\_\{\\hat\{c\}\}\\geq\\tau\_\{\\hat\{c\}\},\(13\)with the per\-channel thresholdτc\\tau\_\{c\}selected on validation data and frozen with the rest of the configuration\. The Router is a firing rule, not a calibrated probability: we do not recalibratepc^p\_\{\\hat\{c\}\}and make no claim that it is well calibrated in the sense of[Liu et al\. \(2024\)](https://arxiv.org/html/2609.36138#bib.bib31), whose recalibration of internal tool\-use probabilities addresses a different quantity\. When the rule does not fire, the forward pass is left untouched andy^1=y^0\\hat\{y\}\_\{1\}=\\hat\{y\}\_\{0\}\.

#### Router: objective, populations and threshold\.

The Router is a per\-channel linear discriminator fitted on training activations atℓobs\\ell\_\{\\mathrm\{obs\}\}: features are standardised to zero mean and unit variance, and the classifier isL2L\_\{2\}\-regularised logistic regression atC=1\.0C=1\.0with an intercept, solved byliblinearto a tolerance of10−410^\{\-4\}under a cap of2,0002\{,\}000iterations and a fixed seed, with non\-convergence at that cap treated as a hard failure rather than a fitted model\. The two populations are asymmetric by design\. The positive class is the channel’s own training errors, the rows whose reference isggand whose baseline prediction isss\. The negative class is*every*baseline\-correct training row,y^0=y⋆\\hat\{y\}\_\{0\}=y^\{\\star\}, and not only those whose reference isgg: the Router is asked to separate a channel\-error state from correct behaviour anywhere in the action space, which is the discrimination its firing decision actually makes at inference\. No class weighting is applied, so the fit is unbalanced in the ratio the training split supplies\.

The firing threshold is selected on a gridτ∈\{0\.4,0\.5,0\.6,0\.7,0\.8\}\\tau\\in\\\{0\.4,0\.5,0\.6,0\.7,0\.8\\\}evaluated on the development rows whose baseline prediction is the source action — the population the rule will face in operation, rather than the full development split — andτc\\tau\_\{c\}is the smallest grid value at which precision on that population reaches0\.500\.50\. A channel is Router\-eligible only if its development ROC AUC against all\-correct negatives is at least0\.750\.75and the selected threshold meets that precision floor; channels failing either are set aside before any intervention\. The fitted standardiser and coefficients are written to the freeze with the selectedτ\\tau, so the operating point is fixed before evaluation rather than tuned against it\. The Router is not recalibrated as a probability, andpc^p\_\{\\hat\{c\}\}is used only through this threshold\.

#### What Router performance does and does not establish\.

A Router that separates channel\-error states from a reference population establishesREADABLEand nothing further\. In the layer analysis reported in the main text, Router discrimination stays near\-constant across a range of observation layers while the correction direction estimated at those same layers degrades from usable to noise\. High Router AUC is therefore evidence that the state is decodable at that site, not that intervening there will help\.

### D\.4Estimate: correction directions

Three estimators appear in this work\. Only the first is used in the sealed evaluations; the others are historical or serve as comparator arms\.

#### Gradient / target\-axis estimator \(dgradd\_\{\\mathrm\{grad\}\}\)\.

The frozen sealed estimator\. In the Gemma configuration it is recorded as

dgrad=unit⁡\(1n​∑iunit⁡\(Gi,gold−Gi,source\)\),d\_\{\\mathrm\{grad\}\}\\;=\\;\\mathrm\{unit\}\\\!\\left\(\\frac\{1\}\{n\}\\sum\_\{i\}\\mathrm\{unit\}\\\!\\left\(G\_\{i,\\text\{gold\}\}\-G\_\{i,\\text\{source\}\}\\right\)\\right\),a normalised mean of per\-example normalised gradient differences between the gold and source action scores, computed on training rows only \(n=401n=401training channel errors for Gemma\)\. The direction is unit\-norm and is hashed into the freeze manifest before evaluation\.

For one input and one candidate mode the scored quantity is the mean over that candidate’s token positions of the log\-probability the model assigns to them\. Its gradient is taken with respect to the MLP forward output atℓinj\\ell\_\{\\mathrm\{inj\}\}— the same tensor the intervention later modifies — and reduced to a vector by summing over sequence positions, which givesGi,mG\_\{i,m\}\. The inner normalisation above makes the outer mean an average of directions rather than of magnitudes, so no single large\-gradient row dominates it\. Nothing inside the estimator is optimised, weighted, filtered or searched: there is no layer search, sign search, sample selection or dose tuning\. The difference\-in\-means comparator follows the opposite subtraction by construction, mean correct\-reference activation minus mean channel\-error activation, so that both directions point from the error state toward the required action\.

#### Difference\-in\-means \(DiffMean\)\.

The historical estimator, and in the sealed protocol a comparator arm rather than the primary\. It is the difference between the mean activation of correct reference rows and the mean activation of channel\-error rows, computed on training data at a declared layer\.

#### Principal component \(PCA\-1\)\.

Used in historical configuration studies where the difference\-in\-means direction was poorly aligned with the channel’s leading principal component\. It is not part of the sealed protocol and is excluded by construction from the formal arm inventory \(Appendix[K](https://arxiv.org/html/2609.36138#A11)\)\.

### D\.5Intervene

When the Router fires for channelcc, the hidden state at the injection layerℓinj\\ell\_\{\\mathrm\{inj\}\}is modified and the forward pass continues unchanged\. The site is the MLP forward output, after the MLP internals and before the decoder residual addition, and the perturbation is applied at every sequence position, matching the aggregation the direction was estimated under:

hℓinj′=hℓinj\+qc​sc​dc,h^\{\\prime\}\_\{\\ell\_\{\\mathrm\{inj\}\}\}\\;=\\;h\_\{\\ell\_\{\\mathrm\{inj\}\}\}\\;\+\\;q\_\{c\}\\,s\_\{c\}\\,d\_\{c\},\(14\)wheredcd\_\{c\}is the unit\-norm correction direction,scs\_\{c\}a scale statistic of the activation norms at that site, andqcq\_\{c\}a relative dose\. The productqc​scq\_\{c\}s\_\{c\}is the*absolute perturbation budget*, and it is this quantity, notqcq\_\{c\}, that is held identical across every directional arm so that comparisons are budget\-matched\. Thescs\_\{c\}above is the median Euclidean norm of the MLP output atℓinj\\ell\_\{\\mathrm\{inj\}\}, taken over the training rows on which the committed Router fires\. It is computed once from the committed Routers and the baseline predictions, without a model load or any development outcome, and enters the freeze as a number: the runner multiplies it by the dose and verifies the achieved perturbation norm against the declared one rather than recomputing the statistic at evaluation time\.

For Qwen3\-8B the budget isqc​sc=41\.5659848890q\_\{c\}s\_\{c\}=41\.5659848890atqc=1\.0q\_\{c\}=1\.0; for Gemma\-2\-9B it is0\.159511785159828830\.15951178515982883atqc=0\.125q\_\{c\}=0\.125withsc=1\.2760942812786307s\_\{c\}=1\.2760942812786307\. The dose rule ismin\(admissible\): the smallest dose satisfying the declared development criteria is taken, which makes the resulting placements conservative\.

#### Observation and injection sites are distinct\.

The layer at which the state is read need not be the layer at which it is modified\. In every evaluated configuration the injection site*precedes*the observation site \(ℓinj=21\\ell\_\{\\mathrm\{inj\}\}=21againstℓobs=26\\ell\_\{\\mathrm\{obs\}\}=26for Qwen3\-8B;2424against3030for Gemma\-2\-9B\), so a firing decision taken fromhobsh\_\{\\mathrm\{obs\}\}is realised by rerunning the input and intervening atℓinj\\ell\_\{\\mathrm\{inj\}\}on a second forward pass\. Model weights are frozen throughout; nothing downstream of the injection site is modified\.

### D\.6Offline construction versus online operation

Table 4:Which stages require labelled data and which run at inference time\. Only Detect and Intervene have an online counterpart; everything that consumes reference labels happens before deployment\.Table[4](https://arxiv.org/html/2609.36138#A4.T4)separates what is built offline from what runs at inference\. At inference the system holds a frozen backbone, one Router per retained channel, a direction bank\{dc\}\\\{d\_\{c\}\\\}with doses\{qc\}\\\{q\_\{c\}\\\}and thresholds\{τc\}\\\{\\tau\_\{c\}\\\}, and a hook atℓinj\\ell\_\{\\mathrm\{inj\}\}\. An input is scored, the Router either fires or does not, and the answer is produced with or without one additive modification\. No weights are updated and no reference label is consulted\.

## Appendix ELicensing Criteria

### E\.1The property hierarchy

Six properties are adjudicated, and they are ordered only in the sense that the later ones presuppose the earlier ones being posable:

ADJUDICABLEReference labels, an error population, a reference population and a defined exposure set exist, and at least one action lies outside\{g,s\}\\\{g,s\\\}so thatOcO\_\{c\}is not forced to zero by construction\.

READABLEChannel\-error states are discriminable from a reference population at the declared observation site\.

STEERABLEThe intervention has a direction\-specific effect beyond zero, reversed, matched\-random and wrong\-site controls\.

CORRECTABLEArrivals concentrate onggrather than on other wrong actions\. Undefined when\|𝒴\|=2\|\\mathcal\{Y\}\|=2\.

PRESERVINGCollateral disruption stays within a prespecified bound\. Two forms are distinguished throughout and only the first is adjudicated here\.*Population\-level preservation*boundsE1\\mathrm\{E1\}over all baseline\-correct decisions, and is the property the frozen gate tests and the one a licence in this paper asserts\.*Exposure\-conditional preservation*boundsE2\\mathrm\{E2\}over the decisions the Router actually touches; it is the stronger property, it is reported for every setting, and no setting in this paper attains it \(Appendix[E\.4](https://arxiv.org/html/2609.36138#A5.SS4)\)\.

LICENSABLEThe evidence meets a prespecified rule, at declared thresholds and confidence, for the claim being made\.

A*repair claim*is not a rung\. It is the conjunction

REPAIR​\(π\)≡CORRECTABLE​\(π\)∧PRESERVINGpop​\(π\)∧LICENSABLE​\(π\)\\textsc\{REPAIR\}\(\\pi\)\\;\\equiv\\;\\textsc\{CORRECTABLE\}\(\\pi\)\\wedge\\textsc\{PRESERVING\}\_\{\\text\{pop\}\}\(\\pi\)\\wedge\\textsc\{LICENSABLE\}\(\\pi\)\(15\)for a settingπ=\(model,dataset,channel,intervention,dose\)\\pi=\(\\text\{model\},\\text\{dataset\},\\text\{channel\},\\text\{intervention\},\\text\{dose\}\), and only a formal ADMIT licenses it\.

Table 5:Correction across seven evaluated settings\.Effectis net gain on the locked test \(historical\) or the target gain on the evaluation population \(sealed\)\.Controlssummarises the comparison that decides direction specificity\. Historical and sealed settings use different protocols and are not compared with each other\. Historical settings carry no formal verdict: the frozen gate was written afterwards and was never run on them, and aDECLINEdenotes insufficient evidentiary support under the protocol, not uncorrectability\. Figure[3](https://arxiv.org/html/2609.36138#S4.F3), Appendix[J](https://arxiv.org/html/2609.36138#A10)and Appendix[K](https://arxiv.org/html/2609.36138#A11)give each setting’s full control battery\.†Secondary audit records report only the comparison with random directions\.ModelProtocolEffectControlsSpecificVerdictPhi\-3\.5\-minihistoricalnet \+55random \+14, reversed \+14, wrong layer \+3, mismatched channel−17\-17; five seeds \+53 to \+67yesnot sealedQwen2\.5\-7Bhistoricalnet \+79reversed \+16; ten random directions, mean \+25\.2, largest \+50yesnot sealedLlama\-3\.1\-8Bhistorical—†as many as 5 of 20 random directions reach the real effectnonot sealedMistral\-7Bhistoricalnet \+12random mean \+29, 17 of 20 reach the real effect; reversed \+19nonot sealedQwen3\-8Bsealed0\.2760 of 59 random directions reach it \(largest 0\.023\)yesADMITQwen3\-4Bsealed0\.0810 of 59 \(largest 0\.000\)yesDECLINEGemma\-2\-9Bsealed0\.0730 of 59 \(largest 0\.000\)yesDECLINETable 6:The prospectively frozen SAKIKO evidence hierarchy\. A formal repair claim is not a standalone property, but the strict conjunctionREPAIR​\(π\)≡CORRECTABLE​\(π\)∧PRESERVINGpop​\(π\)∧LICENSABLE​\(π\)\\textsc\{REPAIR\}\(\\pi\)\\equiv\\textsc\{CORRECTABLE\}\(\\pi\)\\wedge\\textsc\{PRESERVING\}\_\{\\text\{pop\}\}\(\\pi\)\\wedge\\textsc\{LICENSABLE\}\(\\pi\), whose preservation term is the population\-level form\.
### E\.2The frozen ten\-condition gate

Each sealed setting is adjudicated by the ten conditions of Table[7](https://arxiv.org/html/2609.36138#A5.T7), written in the Qwen3\-8B preregistration before the first sealed evaluation and applied unchanged to all three settings\. The gate has a deliberate symmetry: every scientific quantity it uses is tested twice, once at its point estimate and once at the edge of its95%95\\%interval\. Destination therefore contributes four conditions and preservation two; the remaining four are integrity and specificity checks that are not about effect size at all\. Two of the four destination conditions are not independent at the point estimate: sinceTGc=\(Xc/Nc\)​\(2​THc−1\)\\mathrm\{TG\}\_\{c\}=\(X\_\{c\}/N\_\{c\}\)\(2\\,\\mathrm\{TH\}\_\{c\}\-1\), whenXc\>0X\_\{c\}\>0the conditionsTGc\>0\\mathrm\{TG\}\_\{c\}\>0andTHc\>0\.50\\mathrm\{TH\}\_\{c\}\>0\.50are algebraically equivalent\. They are retained as separate numbered conditions because the freeze numbered them so, and because the two quantities carry different information —TGc\\mathrm\{TG\}\_\{c\}retains coverage overNcN\_\{c\},THc\\mathrm\{TH\}\_\{c\}the destination composition overXcX\_\{c\}— but they are not two independent pieces of positive evidence, and their interval versions, Conditions 3 and 6, are not equivalent\.

Table 7:The canonical definition of a SAKIKO licence\. Conditions are numbered as in the frozen artifacts and grouped here by what they check\. An ADMIT requires all ten; the rule recorded in the freeze is10/10; 9/10 = DECLINE\. Destination intervals are bootstrap intervals over channel errors with10,00010\{,\}000draws; the collateral interval is the one\-sided upper bound from the same draws\. Apart from the zero boundary in Conditions 2 and 3,0\.500\.50is the only destination threshold the gate uses\.
### E\.3What the licence adds over simpler rules

A natural objection is that the ten conditions are more apparatus than the evidence requires: perhaps a simpler reporting rule would reach the same conclusions\. Figure[7](https://arxiv.org/html/2609.36138#A5.F7)answers this directly by replaying every evaluated unit — sealed, historical and development alike — against eight alternative reporting rules, from “Net gain above zero” to the full frozen conjunction\. The sequence is not a nested tightening: R4 drops the specificity requirement rather than adding to it\. The replay is a retrospective diagnostic over units gathered under different protocols, not nine independent confirmatory experiments, and a unit passing a rule here is not a formal verdict for that unit\.

A rule based on aggregate gain alone accepts all nine units\. Adding specificity removes three\. Adding destination correctness on point estimates removes one more\. Requiring collateral to be bounded removes another, and requiring the collateral bound to be non\-vacuous removes a further one\. The full conjunction accepts one\.

Two features of the sequence are worth stating plainly\. It is not monotone: R4 rises to six because it tests destination*without*requiring specificity, so it readmits a setting whose movement is not established as direction\-specific\. And the gap between R5b and R6 is the entire contribution of the interval layer — three units clear every point estimate and every non\-vacuous bound, and two of them are still declined\.

Figure 7:What the licence adds over simpler reporting rules\.\(a\)Every evaluated unit against eight alternative reporting rules, applied retrospectively; the sequence is not nested, since R4 drops specificity rather than adding a requirement, and passing a rule here is a diagnostic reading rather than a formal verdict\. Units span both protocol generations and include one development\-stage setting \(Gemma ca→\\rightarrowdirect DEV\), which is retained here because it is the clearest case of a vacuous preservation bound: zero exposed correct decisions\.\(b\)How many units each rule accepts\. A rule based on aggregate gain alone accepts all nine; the full frozen conjunction accepts one\. R4 rises because it drops the specificity requirement rather than because the evidence improves\.
### E\.4What a preservation\-certifying gate would require

The gate this paper ran adjudicates preservation onE1\\mathrm\{E1\}, and the reason is historical rather than principled: the Qwen3\-8B preregistration fixed that denominator and the exposure\-conditional field entered the record only with the later Qwen3\-4B run\. We did not retrofitE2\\mathrm\{E2\}into an already\-frozen conjunction, so the licence reported here is population\-level\. We state here what a gate that certified the property as defined would require, so that the gap is specified rather than merely disclosed\.

Two conditions would replace Conditions 7 and 8: a one\-sided95%95\\%upper bound onE2\\mathrm\{E2\}belowϵtol=0\.05\\epsilon\_\{\\mathrm\{tol\}\}=0\.05— one\-sided rather than bootstrap, because the bootstrap bound is degenerate whenever no break is observed \(Appendix[O](https://arxiv.org/html/2609.36138#A15)\) — and a minimum exposure\|Cexp\|\|C\_\{\\mathrm\{exp\}\}\|large enough for that bound to be attainable\. The second is not a free parameter: Appendix[O](https://arxiv.org/html/2609.36138#A15)fixes it at5959exposed decisions with zero observed breaks,9393with one,124124with two and153153with three\. Against that requirement Qwen3\-8B exposes66, Gemma\-2\-9B1111and Qwen3\-4B5050, so none of the three would clear the exposure minimum and the ADMIT would become a fourth decline\. No re\-analysis of the existing records closes a gap of that size; it requires a Router that exposes more decisions, or an evaluation population large enough to supply them\.

Two further conditions are absent from the frozen conjunction and belong in any successor to it\. The first is coverage:THc\\mathrm\{TH\}\_\{c\}is computed over source exits, so it says nothing about how much of the channel was reached, and a gate written on it alone rewards a Router that fires rarely on easy rows\. A successor should condition on Router recall overNcN\_\{c\}and on gold arrivals overNcN\_\{c\}rather than overXcX\_\{c\}; on the three sealed settings those arrival shares are0\.4370\.437,0\.2980\.298and0\.1770\.177\. The second is effect size: Condition 2 asks onlyTGc\>0\\mathrm\{TG\}\_\{c\}\>0, which licenses an arbitrarily small positive gain provided its interval clears zero, where a threshold tied to what a deployment would find worth the collateral risk would be a stronger rule than a sign test\. We add neither here: both would have to be fixed before evaluation to mean anything, and the evaluation is already spent\.

### E\.5What each verdict asserts

An ADMIT licenses the repair claim for that setting, under that protocol and dose, and nothing more general\. It is not a deployment\-safety claim: the preservation conditions the gate uses are computed on E1, the population denominator, and a setting may pass them while leaving the exposure\-conditional rate unbounded \(Appendix[H](https://arxiv.org/html/2609.36138#A8)\)\.

A DECLINE states that the evidence is insufficient for the stronger claim under the evaluated protocol and the available sample support\. It does not state that the intervention cannot work, that the setting cannot be corrected, or that one model ranks below another\. Because every property is conditional on the model, dataset, channel, protocol and dose, a licence attaches to evidence gathered under a declared protocol rather than to a model\.

## Appendix FDatasets and Action Spaces

Three benchmarks appear in this work, in three different roles \(Table[8](https://arxiv.org/html/2609.36138#A6.T8)\)\. Only one carries formal licence outcomes, and we keep the roles separate rather than pooling them into a single breadth claim\.

### F\.1When2Call

The primary benchmark\. Each input offers four candidate responses, one per pre\-execution action mode\. Three of the four occur as reference labels; answering directly appears only as a distractor\. Labels are synthetically generated and automatically assigned by the benchmark’s authors, and we did not re\-annotate them \(see Appendix[Q](https://arxiv.org/html/2609.36138#A17)for what this limits\)\. A decision is read out by teacher\-forced scoring of the four complete candidates, and the predicted action is the highest\-scoring one; the readout is frozen and deterministic, so a baseline run and an intervened run differ only through the intervention\. The3,6523\{,\}652items are split2,5562\{,\}556/548548/548548into training, validation and locked test\.

In the released artifacts the four modes appear astool\_call,request\_for\_info,cannot\_answeranddirect; the paper writes the last asdirect\_answerfor readability\.

### F\.2MetaTool

Used in its binary form for one model, where the only choice is whether to call a tool\. Because leaving the source mode is arriving at the target by construction,Oc=0O\_\{c\}=0and target\-hit is identically one: the destination question of Appendix[B](https://arxiv.org/html/2609.36138#A2)cannot be posed\. We use MetaTool only for the argument that the correction effect is not a single global shift in tool\-call propensity, since two opposite channels on disjoint populations both improve\.

### F\.3ACEBench

Instantiated as a second action ontology \(tool\_call,ask\_user,flag\_param\_error,cannot\_comply\), read out by free generation and a rule parser rather than by teacher\-forced scoring, and evaluated bilingually\. A criterion set locked before any result was seen first certified the readout, and it passed: over800800items, accuracy0\.7560\.756, macro\-F10\.6870\.687, no predicted class above a0\.700\.70share,7\.25%7\.25\\%unparseable,100%100\\%agreement on a replay check and0\.98%0\.98\\%parser mislabelling on an audited subset of102102rows\.

What failed was the channel structure, not the instrument\. No directional transition met the preregistered per\-split support minimum, and a paraphrase\-agreement check required before channel discovery returned0\.560\.56against a threshold of0\.800\.80\. Because136136of the137137native\-label errors trace to the official parse rules, the failure localises to the benchmark’s size and wording sensitivity\. Under the locked decision rule this retires ACEBench from the intervention line: we report no correction, no destination and no licence outcome for it\. Two qualifications belong with that verdict — the parser figure is a point estimate on102102rows whose stratification is undocumented, and the paraphrase reformulation is recorded in our own audit as carrying a weaker format directive than the original, so the0\.560\.56is not clean evidence of wording fragility\.

Table 8:Benchmark roles\. Only When2Call carries formal licence outcomes\.*Gold*is whether gold labels are recoverable,*Dest\.*whether the destination question is posable,*Interv\.*whether an intervention was run\. A dataset entering the table is not thereby evidence of breadth; the*status*column records what each one actually supports\.

## Appendix GModels and Implementation

Table 9:Evaluated settings and their frozen configuration\. The horizontal rule separates the two protocol generations, which are never pooled\. Entries markedN/Vhave no retained artifact to verify them against: the historical settings predate the configuration\-freeze policy that governs the sealed rows, and we leave them unstated rather than quote an unverified value\.ModelProtocolℓobs\\ell\_\{\\mathrm\{obs\}\}ℓinj\\ell\_\{\\mathrm\{inj\}\}EstimatorDoseqqChannelsPhi\-3\.5\-minihistorical1814 / 16‡DiffMean / PCA\-1α=10\.0\\alpha=10\.03Qwen2\.5\-7Bhistorical2016DiffMean / PCA\-1N/V3Llama\-3\.1\-8BhistoricalN/VN/VN/VN/VN/VMistral\-7B\-v0\.3historicalN/VN/VN/VN/VN/VQwen3\-8Bsealed2621dgradd\_\{\\mathrm\{grad\}\}1\.01Qwen3\-4Bsealed2621dgradd\_\{\\mathrm\{grad\}\}0\.251Gemma\-2\-9Bsealed3024dgradd\_\{\\mathrm\{grad\}\}0\.1251#### Layer mapping\.

Sealed observation and injection sites are derived from a single normalised\-depth rule anchored on the Qwen2\.5\-7B configuration,\(ℓobs,ℓinj\)=\(20,16\)\(\\ell\_\{\\mathrm\{obs\}\},\\ell\_\{\\mathrm\{inj\}\}\)=\(20,16\)on2828layers:ℓnew=round⁡\(ℓold/\(Lold−1\)×\(Lnew−1\)\)\\ell^\{\\mathrm\{new\}\}=\\mathrm\{round\}\\\!\\left\(\\ell^\{\\mathrm\{old\}\}/\(L^\{\\mathrm\{old\}\}\-1\)\\times\(L^\{\\mathrm\{new\}\}\-1\)\\right\)\. This reproduces the committed sites for Qwen3\-8B \(26/2126/21\) and yields30/2430/24for Gemma\-2\-9B\. The rule was fixed before the sealed settings were configured, so the sites are mechanically derived rather than searched per model\. Qwen3\-4B shares Qwen3\-8B’s layer count and therefore its sites,26/2126/21, confirmed in its formal run log\.

#### Dose grid\.

Relative doses are drawn from a frozen six\-point grid,q∈\{0\.0,0\.125,0\.25,0\.5,1\.0,2\.0\}q\\in\\\{0\.0,\\,0\.125,\\,0\.25,\\,0\.5,\\,1\.0,\\,2\.0\\\}, under the rulemin\(admissible\)\. The three sealed settings take1\.01\.0,0\.250\.25and0\.1250\.125from this grid; the absolute budgets areqc​sc=41\.5659848890q\_\{c\}s\_\{c\}=41\.5659848890,6\.17489054056\.1748905405and0\.159511785159828830\.15951178515982883respectively\. The historical Phi\-3\.5 configuration uses a different parameterisation, a Router thresholdτ=0\.4\\tau=0\.4withα=10\.0\\alpha=10\.0at seed4242; two of its five seeds share that configuration and three re\-select both values on validation data\.

‡Phi\-3\.5 records two injection variants,mlp\_allatL​14L14andmlp\_promptatL​16L16\. Its direction\-estimation layer is recorded inconsistently across two internal documents and is marked for verification below\.

#### Execution environment\.

The formal Qwen3\-8B run is frozen under protocol versionQWEN3\_STAGE2\_FORMAL\_V1withbfloat16weights, eager attention, batch size11, a single process and a single model load, fixed arm order and fixed sample order, no checkpointing between scientific arms and no resume between arms\. A zero\-magnitude gate runs before any endpoint is computed\. The recorded software stack for the ACEBench readout, which shares the environment, is PyTorch2\.1\.2\+2\.1\.2\{\+\}cu121, Transformers4\.49\.04\.49\.0, CUDA12\.112\.1on an NVIDIA RTX 4090 D\. Qwen3\-8B is pinned in the preregistration toQwen/Qwen3\-8Bat revisionb968826d9c46dd6066d109eabc6255188de91218, Qwen3\-4B toQwen/Qwen3\-4Bat revision1cfa9a7208912126459214e8b04321603b3df60c, and Llama\-3\.1\-8B tometa\-llama/Llama\-3\.1\-8B\-Instructat revision0e9e39f249a16976918f6564b8830bc894c89659,bfloat16, seed4242\. Table[9](https://arxiv.org/html/2609.36138#A7.T9)carries no revision column, and the remaining model revisions are not listed in this paper; they must be read from the repository manifests\.

## Appendix HFormal Evaluation: Qwen3\-8B

### H\.1Procedure in temporal order

Baseline predictions were collected and the confusion topology enumerated; the channelcannot\_answer→\\rightarrowtool\_callwas selected and the selection disclosed; the direction was estimated on training rows only and hashed; the Router and threshold were fitted on training and validation; the dose was calibrated on development data under themin\(admissible\)rule; the runner, configuration, direction hashes and random seeds were frozen and a gate\-only pass verified them without opening the evaluation payload; an access marker was written; the formal run executed once; endpoints were computed after the zero\-arm exactness gate passed\.

### H\.2Destination\-resolved outcome

Table[10](https://arxiv.org/html/2609.36138#A8.T10)resolves every decision in the two adjudicated cohorts, the8787channel errors and the211211baseline\-correct rows\. Of the8787channel errors,7272were routed and1515were returned unchanged; of the7272,2020retained the source,3838reached the required action and1414landed on a third\. The Router exposed66of the211211baseline\-correct decisions and broke none\.

Table 10:Qwen3\-8B formal run, every decision in the two adjudicated cohorts resolved; the remaining evaluation rows are neither channel errors nor baseline\-correct\. Of8787channel errors,7272were routed; the1515unrouted errors are returned unchanged by construction\. Note the two Net conventions:*whole*counts every fixed decision in the evaluation population,*channel*counts arrivals in the adjudicated channel net of breaks, which is why the score\-space comparator’s3737arrivals and22breaks appear as\+35\+35in Table[13](https://arxiv.org/html/2609.36138#A12.T13)\. The paper quotes both and never substitutes one for the other\.Baseline stateFinal destinationCountPopulationchannel errorsource retained \(routed\)20routed, 72channel errorGOLD ARRIVAL38routed, 72channel errorOTHER WRONG14routed, 72channel errorunrouted, returned unchanged15Nc=87N\_\{c\}=87baseline correctretained correct6 exposed, 0 brokenCexp=6C\_\{\\mathrm\{exp\}\}=6baseline correctnot exposed205\|C\|=211\|C\|=211source exitsXcX\_\{c\}52target\-hitTHc=38/52\\mathrm\{TH\}\_\{c\}=38/520\.7308CI\[0\.60416,0\.84615\]\[0\.60416,0\.84615\]target gainTGc=24/87\\mathrm\{TG\}\_\{c\}=24/870\.2759CI\[0\.115,0\.425\]\[0\.115,0\.425\]Fixed / Broke / Net \(whole\)40 / 0 /\+40\+40Net \(channel\-level\)\+38\+38collateral E10 / 211upper bound0\.01410\.0141collateral E20 / 6upper bound0\.39300\.3930
### H\.3Adjudication

All ten conditions passed and the frozen verdict isFORMAL\_CONFIRMATORY\_SUCCESS\. None of the5959budget\-matched random directions reached the real direction’s target gain; the largest random target gain was0\.02300\.0230against the real0\.27590\.2759, giving an add\-oneppof0\.01670\.0167\. The zero arm reproduced the baseline exactly, the reversed direction moved at most one channel error off its source, and the same direction injected atℓ=26\\ell=26reached a target gain of0\.05750\.0575\.

#### The licence carries a qualifier\.

Preservation passes on E1 — zero breaks among all211211baseline\-correct decisions, one\-sided upper bound0\.01410\.0141— but the Router exposed only six correct decisions, so the exposure\-conditional bound is0\.39300\.3930and E2 is*vacuous*as evidence of preservation\. The statement this setting supports is that no collateral damage was observed in the population, not that the intervention is safe on the decisions it touches\. This qualifier belongs to the claim and is not separable from it\.

## Appendix IFormal Evaluation: Qwen3\-4B

This setting is a formal adjudication, not a failed experiment\. The intervention produced directional movement that survived every specificity control, and every point estimate cleared its threshold\. The gate declined it on the precision of the destination evidence alone\.

### I\.1Outcome

Of124124channel errors,118118were routed\. The intervention produced6464source exits, of which3737reached the required action and2727landed on a third wrong action, givingTHc=0\.5781\\mathrm\{TH\}\_\{c\}=0\.5781over exits andTGc=10/124=0\.081\\mathrm\{TG\}\_\{c\}=10/124=0\.081over all channel errors\. Fixed and Broke were4040and11, a Net of\+39\+39on the whole population and\+36\+36at channel level\. The Router exposed5050correct decisions and broke one of them, so E2 is1/501/50with a one\-sided upper bound of0\.09140\.0914; the population rate E1 passed both its point and interval conditions\. The frozen collateral audit quantifies why that pass is weak evidence:164164of the214214rows in the frozen denominator,76\.6%76\.6\\%, were never perturbed at all, so the gate’s own metric is diluted by the rows the Router declined\. The audit records the exposure\-conditional rate beside it as a required diagnostic disclosure rather than as the frozen metric, and carries a claim lock stating that passing the formal collateral endpoint must not be read as demonstrated deployment safety\. That lock is the reason this paper reports both denominators\. None of the5959budget\-matched random directions reached the real target gain — the largest was0\.00000\.0000— for an add\-oneppof0\.01670\.0167\.

The named control arms separate cleanly here, and we report them because the same battery on Qwen3\-8B is quoted only in prose\. Against the real arm’sTGc=0\.081\\mathrm\{TG\}\_\{c\}=0\.081over6464exits, the zero arm reproduced the baseline exactly and the reversed direction moved*no*decision off its source at all \(00exits\)\. Both misallocations are negative rather than merely weaker: the frozen DiffMean estimator reachesTGc=−0\.016\\mathrm\{TG\}\_\{c\}=\-0\.016on44exits, and the wrong\-layer arm−0\.024\-0\.024on1515\. Removing the Router leaves arrivals unchanged at3737but raises breaks from11to66, so gating buys a six\-fold reduction in collateral at no cost in arrivals\. The score\-space comparator is the one arm not dominated, and it is the clearest illustration in the paper of why the two axes must be read together: it reaches a higher target gain \(0\.1050\.105,4444arrivals\) and breaks66exposed decisions against the activation arm’s11\. Neither arm is better on both axes at once, which is precisely the situation an aggregate would collapse and a destination\-resolved account keeps visible\. We report the contrast and claim no ordering between them\.

### I\.2Why the protocol declined

Table 11:The two sealed Qwen3 settings against the frozen conjunction\. Both clear every point estimate and both satisfy the frozen preservation conditions, which are written onE1\\mathrm\{E1\}; onE2\\mathrm\{E2\}the point estimates are below the bound but the interval evidence is not \(0\.39300\.3930and0\.09140\.0914\)\. The verdicts diverge on the two destination\-interval conditions and on nothing else\. Aggregate gain does not separate them:\+40\+40against\+39\+39\.\#ConditionLayerThresholdQwen3\-8BQwen3\-4B1channel supportintegrity≥30\\geq 30passpass9zero arm exactnessintegrityexactpasspass10no structural failureintegrity—passpass4specificity,K=59K=59specificityp≤0\.05p\\leq 0\.05passpass2target gain, pointdestination\>0\>0passpass5target\-hit, pointdestination\>0\.50\>0\.50passpass3target gain, intervaldestinationCI low\>0\>0passfail6target\-hit, intervaldestinationCI low\>0\.50\>0\.50passfail7collateral, pointpreservation≤0\.05\\leq 0\.05passpass8collateral, intervalpreservation≤0\.05\\leq 0\.05passpasstarget\-hit point estimate0\.73080\.5781target\-hit 95% interval\[0\.604,0\.846\]\[0\.604,0\.846\]\[0\.456,0\.697\]\[0\.456,0\.697\]source exits5264Net \(whole\)\+40\+40\+39\+39frozen verdictADMITDECLINETable[11](https://arxiv.org/html/2609.36138#A9.T11)sets the two sealed Qwen3 settings against the conjunction condition by condition\. The target\-hit interval\[0\.4559,0\.6970\]\[0\.4559,0\.6970\]reaches below one half and the target\-gain interval includes zero\. Under the frozen rule — ten of ten admits, nine of ten declines — the protocol returnedQWEN3\_4B\_FORMAL\_DECLINE\.

This outcome is the clearest demonstration in the paper that the licensing rule changes the scientific conclusion\. A reporting standard based on aggregate gain would not distinguish these two settings at all:\+40\+40against\+39\+39, a difference of one decision\. A standard that added specificity would still admit both, since both clear theK=59K=59null\. What separates them is the width of the destination evidence at the sample sizes available, and Appendix[O](https://arxiv.org/html/2609.36138#A15)shows that the declines are what those sample sizes predict\.

## Appendix JHistorical Experiments

The four historical settings developed the framework and are reported under the protocol generation that produced them; Figure[8](https://arxiv.org/html/2609.36138#A10.F8)shows the two whose control batteries survive in full\. They receive no formal verdict: the sealed gate was written afterwards and was never run on them\. Applying its criteria to their recorded counts is a retrospective reading, and we label it as such wherever it appears\.

Figure 8:The historical record behind Sections[4](https://arxiv.org/html/2609.36138#S4)and[5](https://arxiv.org/html/2609.36138#S5)\.\(a\)E1 collateral for each of the five artifact\-backed Phi\-3\.5 seeds against the5%5\\%bound, with each run’s Net gain: breakage ranges from1616to5252of264264baseline\-correct decisions while the Net gain stays between\+53\+53and\+67\+67, and every realisation exceeds the bound\.\(b\)The Qwen2\.5\-7B locked battery: the real direction \(\+79\+79\) against its reversal \(\+16\+16\) and the mean of ten budget\-matched random directions \(\+25\.2\+25\.2, the largest reaching\+50\+50\)\. Both panels are historical protocol and both survive as aggregate counts except where noted; they carry no formal verdict\.### J\.1Phi\-3\.5\-mini

The only historical setting with surviving row\-level records, and therefore the only one whose destination composition can be resolved\. Under the locked configuration at seed4242the Router fired on293293decisions, comprising200200channel errors and9393baseline\-correct decisions\. Of the200200errors,4747retained their source and153153left it; of those exits,107107reached the required action and4646landed on a third wrong action, forTH=107/153=0\.6993\\mathrm\{TH\}=107/153=0\.6993\. Of the9393exposed correct decisions,4141survived and5252were broken\.

The aggregate for the same run is Fixed107107, Broke5252, Net\+55\+55\. Collateral is52/93=55\.9%52/93=55\.9\\%on the exposure\-conditional denominator and52/264=19\.7%52/264=19\.7\\%on the population denominator; the one\-sided upper bound on the former is0\.64680\.6468\. Both exceed the5%5\\%bound the sealed protocol later prespecified\.

#### Seed variation\.

Repeating the full pipeline under five seeds gives Net gains of5555,5454,5353,6060and6767, positive in all five, with Broke ranging from1616to5252\. This is stability across the evaluated seeds rather than a general property: two of the five share the frozen configuration and three re\-select their threshold and dose on validation data, so the runs are configuration variants rather than repeated draws from one configuration\. An independent execution of the same code, seed and configuration fired on243243decisions rather than293293and broke4444, for reasons the surviving records cannot establish\. The magnitude of the damage depends on the execution; its presence does not\.

#### Controls\.

Against the real configuration’s\+55\+55: a random direction gives\+14\+14, the reversed direction\+14\+14, injection at a different layer\+3\+3, and one channel’s direction applied to another channel−17\-17, which is worse than leaving the model alone\. Four of these five arms survive only as aggregate counts, so the battery establishes an ordering rather than a decomposition of the effect\.

### J\.2Qwen2\.5\-7B

Applying the frozen channel configurations to the locked test set,9292baseline errors become correct while1313already\-correct decisions become wrong, a Net of\+79\+79\. Three channels are retained — a follow\-up question answered instead with a tool call, a refusal answered instead with a tool call, and a refusal answered instead with a direct answer — and together they carry about78%78\\%of all baseline errors\.

Its destination composition*cannot*be adjudicated: the When2Call records for this setting survive only as aggregate counts, soAcA\_\{c\}andOcO\_\{c\}cannot be separated\. This is why Qwen2\.5\-7B appears on the hierarchy as passing specificity and then stopping, rather than as a destination result\.

Against its own control battery the real direction gives\+79\+79, the reversed direction\+16\+16, and ten budget\-matched random directions a mean of\+25\.2\+25\.2with a largest value of\+50\+50\. Across five seed\-level runs the real direction again exceeds every random draw, although three of those runs drew a single random direction rather than ten — a heterogeneity in null strength that we record rather than average over\.

### J\.3MetaTool

The same machinery was applied to two opposite channels, unnecessary tool calls and missed ones, on disjoint sample populations\. Both improved, with mean Net gains of8\.68\.6and13\.013\.0across five runs and21\.621\.6when applied together, against a baseline accuracy of0\.7730\.773\.

The argument this supports is narrow and specific: because the two channels move tool\-calling behaviour in*opposite*directions on non\-overlapping inputs, a single global shift in tool\-call propensity cannot account for the result\. It supports nothing about destination, because in a binary action spaceOc=0O\_\{c\}=0by construction\.

### J\.4Mistral\-7B and Llama\-3\.1\-8B

Both settings fail the specificity comparison, and the framework stops there\. On Llama\-3\.1\-8B as many as55of2020random directions reach the real effect\. On Mistral\-7B the real direction’s Net gain of\+12\+12falls*below*the random mean of\+29\+29,1717of2020random directions reach it, and the reversed direction gives\+19\+19\.

Two qualifications are required\. Both results rest on secondary audit records rather than on primary row\-level artifacts, and both are flagged in those records as estimator\-suspect and compared against2020random directions rather than the5959of the sealed protocol\. The conclusion these settings support is therefore that, under the evaluated configuration, estimator and site, there is insufficient evidence that the correction is caused by the recovered direction\. They do not establish that no correctable direction exists in these models\.

Llama\-3\.1\-8B is, descriptively, the clearest case of movement without arrival in our records, with1919exits reaching the required action against4444landing elsewhere\. Because the setting fails specificity first, it cannot serve as evidence that a*specific*correction fails to arrive, and we do not use it as such\.

## Appendix KControl Experiments

Every directional arm in a sealed evaluation receives the identical absolute perturbation budgetqc​scq\_\{c\}s\_\{c\}and passes through the identical Router gate as the real arm; Table[12](https://arxiv.org/html/2609.36138#A11.T12)states what each arm exists to rule out\. Matching the budget rather than the relative dose is what makes the comparison meaningful across layers and models\.

Table 12:Control arms in the sealed protocol, transcribed from the frozen control specification\. The*alternative removed*column states what each arm exists to rule out\. Budget and gate are identical to the real arm unless noted\. Qwen3\-8B values are quoted; Gemma\-2\-9B uses the same inventory atℓinj=24\\ell\_\{\\mathrm\{inj\}\}=24with wrong\-layer atℓ=30\\ell=30\.#### The random null is near\-orthogonal to what it tests\.

The5959directions are verified unit\-norm to2\.7×10−92\.7\\times 10^\{\-9\}, mutually near\-orthogonal \(median\|cos\|=0\.011\|\\cos\|=0\.011, maximum0\.0610\.061\) and near\-orthogonal to the calibrated direction \(maximum\|cos\|=0\.050\|\\cos\|=0\.050\), so the null is a budget\-matched comparison against directions that do not overlap the one under test\. Seeds are disjoint from development and the direction matrix is hash\-pinned in the freeze manifest before the evaluation split is opened\.

The construction is recorded in full and is reproducible from the manifest alone: for eachi<Ki<Kthe seed is the first eight bytes ofSHA256\(salt∥i\)\\mathrm\{SHA256\}\(\\text\{salt\}\\,\\\|\\,i\), aPCG64DXSMgenerator drawsz∼𝒩⁡\(0,I\)z\\sim\\mathcal\{N\}\(0,I\)in float64, the vector is normalised in float64 and cast once to float32\. The null is therefore isotropic by construction rather than by inspection\. The manifest also records, per vector, the seed input, the seed digest and the vector digest, so any reader can regenerate the set and check it against the frozen hashes\. Four exclusions are asserted in the same record and are what make the null a null: no rejection sampling, no orthogonalisation, no cosine filtering and no outcome matching, with no regeneration and no post\-hoc selection\. The cosines quoted above were computed*after*the complete set was frozen, and no vector was removed on their basis\.

#### Arms excluded by construction\.

The frozen inventory explicitly excludes a same\-layer DiffMean at the injection site, another channel’s direction, an orthogonalised mismatch, any pooled or shared direction, PCA or rank\-2 directions, any new estimator, and any additional dose\. Excluding them before sealing is what prevents the control battery from becoming a search\.

#### Gating\.

Removing the gate changes the outcome, but not uniformly, and we describe it only by its measured effects\. On Phi\-3\.5 an ungated arm reverses the sign of the effect within an independent execution of the locked configuration, from\+62\+62to−32\-32; on both MetaTool channels the ungated arms fall to−39\-39and−34\-34; on Qwen3\-8B the ungated Net drops from\+38\+38to\+13\+13while its collateral rises from0/60/6to29/176=16\.5%29/176=16\.5\\%\. On Gemma\-2\-9B the ungated arm is a genuine counterexample: its Net is not lower but marginally higher \(\+16\+16gated against\+17\+17ungated\), and its collateral*rate*is lower rather than higher \(4/188=2\.1%4/188=2\.1\\%ungated against1/11=9\.1%1/11=9\.1\\%gated\), though a single break on eleven exposed rows makes that rate comparison unstable, which is why we state the counterexample on Net where it is unambiguous\. The five settings with both a gated and an ungated arm are Phi\-3\.5, the two MetaTool channels taken together, Qwen2\.5\-7B, Qwen3\-8B and Gemma\-2\-9B, and Net is the only endpoint recorded for every one of them\. Read on Net, gating raises the effect in three of the five — Phi\-3\.5, MetaTool and Qwen3\-8B — and lowers it in two, on Qwen2\.5\-7B by\+79\+79against\+94\+94and on Gemma\-2\-9B by\+16\+16against\+17\+17\. We therefore make no necessity claim on Net\. Arrivals per broken decision and collateral rate order the arms differently in places, and we do not substitute one for another\. On Qwen2\.5\-7B the ungated arm reaches a*larger*Net \(\+94\+94against\+79\+79\) by acting on293293decisions instead of202202and absorbing more damage\. What gating raises consistently, in the three settings where both arms are directly comparable, is the number of arrivals per broken decision: from5\.275\.27to7\.087\.08on Qwen2\.5\-7B, from5\.255\.25to17\.0017\.00on Gemma\-2\-9B, and from1\.451\.45on Qwen3\-8B to a gated arm that breaks nothing\.

#### The named arms on Gemma\-2\-9B\.

The third sealed setting separates in the same direction as Qwen3\-4B and we give it for completeness\. Against the real arm’sTGc=0\.073\\mathrm\{TG\}\_\{c\}=0\.073over2727exits \(1717arrivals,1010elsewhere,11break of1111exposed\), the zero arm reproduced the baseline exactly and the reversed direction again moved*no*decision off its source\. Both misallocations are negative: the frozen DiffMean estimator produced a single exit and no arrival at all \(THc=0\\mathrm\{TH\}\_\{c\}=0,TGc=−0\.010\\mathrm\{TG\}\_\{c\}=\-0\.010\), and the wrong\-layer arm−0\.010\-0\.010on33exits\. The ungated arm reaches2121arrivals against the gated1717but breaks44exposed decisions against11, which is the5\.255\.25to17\.0017\.00change in arrivals per break quoted above\.

#### The comparator does not order consistently\.

Across the three sealed settings the activation arm and the frozen score\-space comparator admit no stable ordering\. On Qwen3\-8B the activation arm leads on target gain \(0\.2760\.276against0\.1840\.184\) and on Gemma\-2\-9B it leads again \(0\.0730\.073against0\.0520\.052\), while on Qwen3\-4B the comparator leads \(0\.1050\.105against0\.0810\.081\)\. The comparator’sbcb\_\{c\}is calibrated on development data at the selected dose and never tuned on the sealed split\. That the sign of the difference changes with the setting, while no paired test survives multiplicity correction, is why the frozen interpretation rule for this contrast is honoured rather than set aside: we report it and claim no ordering\.

## Appendix LDestination\-Resolved Evaluation

Net is not a wrong statistic\. It is the right statistic for the question “did accuracy improve”, and on that question it is exact\. The claim of this paper is narrower: Net is a many\-to\-one compression of the five\-class flow in equation[6](https://arxiv.org/html/2609.36138#A2.E6), and the classes it discards are the ones a repair claim depends on\.

### L\.1Two interventions, equal aggregate, different destinations

On the sealed Qwen3\-8B population the preregistered activation intervention and the score\-space comparator produce nearly identical aggregates and materially different destination compositions\.

Table 13:Activation intervention against the score\-space comparator on the same channel and population\. The aggregates are close; the destinations are not\. The paired tests are not significant after multiplicity correction and the comparison was not prespecified, so this contrast is*descriptive*: it establishes neither that activation intervention is better nor that the two differ in general\.A study reporting only the aggregate would have treated these two interventions as interchangeable:\+40\+40against\+39\+39\. Resolved by destination they differ by seven decisions redistributed to a third wrong action\. The paired comparisons are exact McNemarp=1\.0p=1\.0for arrivals andp=0\.189p=0\.189for redistributions, and the archived artifact records that all\-548 paired correctness was not prespecified as a paired comparison\. We therefore report the contrast and draw no superiority claim in either direction\.

## Appendix MMechanistic Insights and Stage Decoupling

Decoupled Dynamics Across Probing, Geometry, Gating, and Dosage\.The analyses in this subsection are a preregistered follow\-up and are exploratory: they are not part of the frozen conjunction, and they locate no component\. Mechanistic analyses on Qwen2\.5\-7B and prospective cohorts demonstrate sharp dissociations between representation decoding and intervention efficacy\. Across layers 16–22 of Qwen2\.5\-7B, linear probe decodability plateaus atAUROC≈0\.96\\text\{AUROC\}\\approx 0\.96, yet steering vector norms scale non\-linearly \(2\.12→14\.262\.12\\rightarrow 14\.26\) and validation net gain jumps from\+13\+13at layer 16 to\+46\+46at layer 20, where layer 16 steering fails due to near\-zero PC1 alignment\. Across geometric formulations, PCA extraction outperforms difference\-in\-means \(\+50\+50vs\.\+36\+36locked net gain\), while cross\-channel injections produce severe negative performance \(−17\-17net gain on Phi\-3\.5\-mini\); the channels are related but not interchangeable, with pairwise cosines of0\.4880\.488,0\.5690\.569and0\.7060\.706between Qwen2\.5\-7B’s three When2Call directions measured at a common layer\. Ablating router gating triggers catastrophic degradation on unconstrained runs, causing net gains to flip negative on Phi\-3\.5\-mini \(\+62→−32\+62\\rightarrow\-32\) and MetaTool \(−39\-39and−34\-34\); conversely, gating sharpens precision, reducing breaks from 6 to 1 on Qwen3\-4B while holding arrivals constant at 37\. Read on Net, gating is not uniformly necessary: it raises Net on Phi\-3\.5, MetaTool and Qwen3\-8B, and lowers it on Qwen2\.5\-7B \(\+79\+79gated against\+94\+94ungated\) and marginally on Gemma\-2\-9B \(\+16\+16against\+17\+17\), so the effect on Net is setting\-dependent\. What gating changes consistently is exposure and the collateral that follows from it\. Finally, dose\-response curves on Qwen3\-4B reveal non\-monotonic scaling, where target\-hit rates span0\.5250\.525to0\.9110\.911across dosage regimes, peaking at intermediate injection magnitudes before regressing\.

The Multi\-Dimensional Architecture of Latent Steering\.These mechanistic investigations reveal that representation decodability does not entail causal steerability: the internal layers that maximally linearly separate error states are not necessarily the computational sites where interventions effectively shift downstream decisions\. Furthermore, the sensitivity of outcomes to extraction geometry and injection dosage demonstrates that causal steering acts within specific, multi\-dimensional subspaces rather than along arbitrary, monolithic rank\-one directions\. Finally, gating limits exposure but does not by itself establish preservation among exposed decisions: it shields the decisions the Router declines, and the evidence here says nothing stronger about the ones it touches\. Consequently, mechanistic repair cannot be treated as a static model property, but must be conceptualized as a delicate balance between extraction geometry, injection dosage, and exposure control\.

## Appendix NSynthesis: The Spectrum from Movement to Repair

Systematic Stage\-Wise Attrition Across Model Architectures\.Mapping the complete empirical evaluation acrossSAKIKO’s staged hierarchy demonstrates progressive attrition, wherein individual model configurations fail at distinct, predictable criteria \(Table[5](https://arxiv.org/html/2609.36138#A5.T5)\)\. In the historical cohort, Llama\-3\.1\-8B and Mistral\-7B fail the foundational specificity controls, failing to differentiate from budget\-matched random vectors\. Moving higher in the hierarchy, Phi\-3\.5\-mini clears destination correctness \(THc=0\.699\\mathrm\{TH\}\_\{c\}=0\.699\) but fails preservation by corrupting 52 baseline\-correct decisions, while Qwen2\.5\-7B exhibits high net gain \(\+79\+79\) but lacks row\-level destination logging\. In the sealed prospective cohort, Qwen3\-4B and Gemma\-2\-9B satisfy point thresholds for correctability and preservation, but fail the licensing stage due to inconclusive uncertainty intervals crossing null boundaries\. Ultimately, only Qwen3\-8B successfully navigates the entire five\-stage pipeline, clearing all structural, empirical, and statistical thresholds to earn anADMITverdict\. The historical stops carry no formal verdict: the gate was written afterwards and was never run on them\.

Enforcing the Boundary Between Behavioral Steering and Verified Repair\.This progressive attrition shows that behavioral movement, destination\-resolved correction, clean\-state preservation, and statistical certitude do not inherently coincide in the settings we evaluate\. Standard steering literature routinely conflates these levels, claiming repair whenever an internal intervention yields positive aggregate deltas or forces an agent out of an error state\. By establishing an evidence pipeline where each verification stage addresses independent failure modes,SAKIKOexposes how aggregate gains can conceal lateral error redistribution, collateral damage, and finite\-sample uncertainty\. Formalizing this staged adjudication transforms internal representation interventions from ad\-hoc activation steering into a rigorous, verifiable methodology for agentic repair\.

## Appendix OStatistical Procedures

#### Bootstrap intervals\.

Destination intervals are percentile bootstrap intervals resampling*channel errors*— not exits, and not evaluation rows — with10,00010\{,\}000draws\. Resampling the error population rather than the exits keeps the denominator ofTGc\\mathrm\{TG\}\_\{c\}fixed and propagates the uncertainty in how many decisions leave their source at all\. Collateral is bounded two ways and we keep them distinct\. Gate Condition 8 uses the bootstrap95%95\\%upper bound on E1, as written in the preregistration\. The bounds quoted in the tables are one\-sided Clopper–Pearson upper bounds at95%95\\%, which at these counts are the more conservative of the two —0/2110/211gives0\.00000\.0000under the bootstrap and0\.01410\.0141under Clopper–Pearson — so every collateral bound the paper prints is at least as strict as the one the gate applied\.

#### With no observed break the bootstrap bound is degenerate\.

This is not a cosmetic difference in the one case it decides\. WhenB=0B=0every nonparametric resample of the correct population also hasB=0B=0, so the bootstrap95%95\\%upper bound on E1 is exactly zero and carries no finite\-sample information: Condition 8 cannot fail on a setting with no observed break, however few decisions were exposed\. Qwen3\-8B is that setting\. We state the consequence rather than leave it to be inferred — its Condition 8 pass is uninformative as written, and the bound a reader should use is the one\-sided Clopper–Pearson value,0\.01410\.0141on0/2110/211\. Substituting Clopper–Pearson into the gate does not change the verdict, since0\.0141≤0\.050\.0141\\leq 0\.05, so no result here turns on the choice; but a gate written now should name the rare\-event\-informative bound prospectively, and Appendix[E\.4](https://arxiv.org/html/2609.36138#A5.SS4)states what else it would have to require\.

#### Undefined target\-hit under resampling\.

A resample in whichXc=0X\_\{c\}=0leavesTHc\\mathrm\{TH\}\_\{c\}undefined\. At the observed exit shares the probability of drawing one is below10−1310^\{\-13\}in every sealed setting —1\.7×10−141\.7\\times 10^\{\-14\}for Gemma\-2\-9B, the least favourable of the three — so the case does not arise in10,00010\{,\}000draws and no discard rule was exercised\.

#### Empirical random null\.

Specificity is judged by an add\-one Monte Carlopp\-value overKKbudget\-matched random directions passed through the same Router gate\. Here and belowKKcounts random directions; the cardinality of the action space, writtenKK\-way in the title, is\|𝒴\|\|\\mathcal\{Y\}\|throughout the formal text\. The null is an empirical distribution over isotropic directions at a matched budget, not a permutation of labels:

padd\-one=1\+\#\{Trand≥Treal\}K\+1,T=TGc\.p\_\{\\text\{add\-one\}\}\\;=\\;\\frac\{1\+\\\#\\\{\\,T\_\{\\mathrm\{rand\}\}\\geq T\_\{\\mathrm\{real\}\}\\,\\\}\}\{K\+1\},\\qquad T=\\mathrm\{TG\}\_\{c\}\.\(16\)WithK=59K=59and no random direction reaching the real target gain, this attains its floor of1/60=0\.01671/60=0\.0167in each sealed setting\. The valueK=59K=59was chosen on computational cost alone, and it bounds the resolution of the test:0\.01670\.0167is the smallestppthis design can report, not a measure of how far the real direction exceeds the null\. The observed margins are reported separately — the largest random target gain was0\.02300\.0230on Qwen3\-8B and0\.00000\.0000on both Qwen3\-4B and Gemma\-2\-9B\.

#### Design sensitivity\.

The main text reports how many source exits a true target\-hit requires before a one\-sided exact binomial test against0\.500\.50rejects atα=0\.05\\alpha=0\.05with80%80\\%power\. Table[14](https://arxiv.org/html/2609.36138#A15.T14)gives the full curve\.

Table 14:Source exits required to certify a true target\-hit above0\.500\.50\. This approximates the gate’s bootstrap criterion rather than replacing it, and we apply it only to the settings in this paper\.Read at the observed point estimates, Qwen3\-8B’s5252exits exceed its requirement of about3030, while Gemma\-2\-9B’s2727and Qwen3\-4B’s6464fall several times short of theirs\. The declines are what the sample sizes predict\.

#### Preservation power\.

Certifying a5%5\\%collateral bound atα=0\.05\\alpha=0\.05with zero observed breaks requires at least5959exposed decisions:0/580/58gives0\.05030\.0503and only0/590/59reaches0\.04950\.0495\. With one break it requiresn≥93n\\geq 93, with twon≥124n\\geq 124, with threen≥153n\\geq 153\. No setting in this paper attains the required exposure, and no re\-analysis, denominator substitution or pooling can close a gap of that size on the existing records\.

#### Boundary conventions\.

Target\-hit is undefined whenXc=0X\_\{c\}=0and is reported as such rather than as zero or one\. A collateral rate computed on zero exposed rows is recorded as*vacuous*and never as preservation\. Multiplicity: the paired activation\-versus\-score\-space tests are corrected and reported as non\-significant; no other comparison in the sealed protocol is multiple\.

## Appendix PSealing, Preregistration and Provenance

### P\.1Separation of development and confirmation

TRAIN and DEV data were used to select the channel, estimator, observation and injection sites, dose and Router threshold\. The evaluation population was sealed throughout that process\. Before any formal run, a gate\-only pass verified — without opening the evaluation payload — all artifact hashes, model and tokenizer identity, software versions, the selected channel and configuration, direction hashes, formal random count and hashes, seed disjointness, arm completeness, fixed arm order, batch size, the single\-load path, that the zero gate precedes endpoint computation, the raw record schema, the primary/secondary hierarchy, the bootstrap specification, the support\-outcome branch, the VOID rules, the sealed\-evaluation firewall, and an empty formal output namespace\.

#### Rows excluded for hardware, and where they came from\.

Six of the3,6523\{,\}652rows could not be run: the four\-mode gradient stack in bfloat16 with eager attention exceeds2424GB at their sequence lengths\. Three fall in train and three in development;*none*falls in the sealed evaluation split, and the exclusion manifest lists their raw indices\. The boundary is a natural gap in the data rather than a chosen threshold — the last feasible row has maximum sequence length3,1693\{,\}169and the first infeasible one4,5724\{,\}572— and it was fixed before any endpoint was observed\. A release scan over the committed record files reports no credentials, no activation caches and no dataset payload, and confirms the record schema: immutable sample identifiers, four\-mode log\-probability scores, predictions, Router probabilities, arm identity, dose and destination, with no prompt, candidate string or generated text retained\.

An access marker was then written, and the formal run executed*once*\. For Gemma\-2\-9B the one\-shot ledger recordsformal\_run\_count = 1,sealed\_rows\_read = 548and6565executed arms \(66named forward arms plus5959randoms\), against authorised repository head23d6da2a64639ca50d43d94de04a5b6d6275d896\.

#### What the freeze fixed, and what it forbade\.

The preregistration is not only a list of thresholds\. It fixes, before any evaluation access, the primary statistic, the ten\-condition conjunction, an*ordered*secondary hierarchy that is explicitly non\-promotable, the fourteen\-step execution order, the bootstrap method and its derived seed, the support rule, thirteen namedVOIDconditions, and a set of interpretation rules that map each possible outcome onto what may be claimed from it\. Three of those provisions do work that a post\-hoc reader cannot replicate\. The conjunction may not be replaced with a best\-looking subset after execution\. Secondary results may not rescue a failed primary\. And the interpretation rules bind in both directions: the frozen rule for the case in which the score\-space comparator matches or exceeds the activation arm states that no claim may then be made that activation intervention provides behaviour unavailable to a frozen score shift\.K=59K=59was fixed from a compute\-cost table overK∈\{39,59,79\}K\\in\\\{39,59,79\\\}and declared independent of development rankings, effect sizes, anticipatedpp\-value and random geometry\. The runner reads every frozen quantity from committed artifacts at run time rather than restating it in code, and the raw record schema retains the*achieved*perturbation norm alongside the intended one, so a drift between specification and execution would be visible in the records rather than silent\.

#### What selection conditions on\.

The channel, estimator and dose were chosen before sealing, on development data, and every sealed result conditions on that disclosed selection\. This is legitimate and it is not blind; we state it rather than present the sealed evaluation as if the configuration had been arbitrary\.

The frozen disclosure names the criteria and the rejected alternatives\. The adjudicated channel was retained because it cleared the support, Router, geometric and dose gates, and because its collateral bound is a real constraint rather than a formality:368368development rows are both routed\-eligible and baseline\-correct\. Two other eligible channels were excluded before sealing and neither is reported as confirmatory afterwards, which would be selection after the fact\. One,cannot\_answer→\\rightarrowdirect, was excluded because its collateral rate is*structurally*zero — the pinned split contains no golddirectrows, so no routed row can ever be baseline\-correct and its safety gate carries no information\. The other,request\_for\_info→\\rightarrowtool\_call, ranked the two interventions in the opposite order on development data and was set aside on that basis\. The single confirmatory channel is therefore chosen on development evidence that includes how the two compare, so the sealed run establishes destination\-correct repair on the channel it adjudicates and is not a general ranking of activation intervention against a frozen score shift\. That is the claim the paper makes, and the disclosure is what keeps it that size\.

### P\.2Corrections applied after the fact

Table[15](https://arxiv.org/html/2609.36138#A16.T15)lists every historical value a later audit corrected, with the artifact that is authoritative in each case\.

Table 15:Historical values that later audits corrected\. Each correction makes a derived index agree with the frozen artifact it cites; no frozen artifact was modified and no verdict changed\.#### A resolved attribution\.

A prose rendering of one mechanism table places the layer\-sweep observation — difference\-in\-means norms of2\.122\.12and14\.2614\.26with validation gains of\+13\+13and\+46\+46— in a column headed*Phi\-3\.5*\. The machine\-readable indices place it on*Qwen2\.5\-7B*’s refusal\-to\-direct\-answer channel, and name the source file accordingly; the recorded sweep is the four\-point series2\.122\.12,5\.885\.88,10\.4110\.41,14\.2614\.26over layers1616to2222, withcos⁡\(DiffMean,PC1\)=0\.0002\\cos\(\\text\{DiffMean\},\\mathrm\{PC\}\_\{1\}\)=0\.0002atL​16L16\. The tabular rendering is a formatting error and the indices are authoritative\. The main text and this appendix attribute the sweep to Qwen2\.5\-7B\.

The one\-shot ledgers record what a restart was permitted to be\. On Qwen3\-4B a mechanicalVOIDfor incomplete all\-arm execution was followed by exactly one authorised restart, and the ledger states its terms: every arm re\-executed from scratch, no resume from partial records, and the voided attempt retained by its identifier\. Its gate\-only pass is recorded as25/2525/25checks against the frozen specification before any sealed row was read\. One mechanical incident is recorded in the Gemma formal freeze and its pre\-repair runner is retained alongside the incident note\. Three pre\-access gate repairs and one mechanical VOID occurred, all before any endpoint was observed, and all are disclosed with before\-and\-after hashes\.

## Appendix QDeclined and Failed Settings

A decline is an output of the protocol, not a failure of it\. That framing does not convert a failure into a success, and each entry below states plainly what did not hold\.

#### Insufficient support: Qwen3\.5\-9B\.

An eighth model was evaluated and never reached the intervention stage\. On every naive indicator it was the most promising setting in the panel — the highest baseline accuracy \(0\.45330\.4533\), the least collapse onto a single predicted class, and errors lying closest to the decision boundary\. Its accuracy gain is unevenly distributed across gold classes: recall on a follow\-up question roughly doubles relative to the other models \(0\.3260\.326against0\.1700\.170and0\.1320\.132\) while recall on a refusal falls to the lowest in the panel \(0\.0870\.087\)\. Because reference and error rows within a gold class compete for a fixed number of items, that redistribution starves the reference side of every refusal channel — including the exact channel carrying the sealed evaluations — leaving2929reference rows against a frozen minimum of3030, and enriches the one channel that three models already agree is not linearly readable at the observation site\. The sealed population was never opened\. The stop is split\-sensitive: under3030retrospective group\-preserving reallocations the refusal channels were eligible in56\.7%56\.7\\%of them, against a median achievable support of3434\.

#### Dataset ineligibility: ACEBench\.

Described in Appendix[F](https://arxiv.org/html/2609.36138#A6)\. The readout passed every validity criterion; the channel structure did not meet the preregistered support minimum, and the paraphrase gate returned0\.560\.56against0\.800\.80\. Retired from the intervention line by the locked decision rule\.

#### Readable but not direction\-specific: Llama\-3\.1\-8B, Mistral\-7B\.

Described in Appendix[J\.4](https://arxiv.org/html/2609.36138#A10.SS4)\. Both fail the specificity comparison, both rest on secondary audit records, and both usedK=20K=20rather thanK=59K=59\.

#### Destination evidence insufficient: Qwen3\-4B, Gemma\-2\-9B\.

Described in Appendix[I](https://arxiv.org/html/2609.36138#A9)\. Both clear every point estimate and both satisfy preservation on the frozenE1\\mathrm\{E1\}denominator; both fail the two destination\-interval conditions and nothing else\.

#### Preservation not certified anywhere, including the ADMIT\.

No setting in this paper attains exposure\-conditional preservation certification\. Qwen3\-8B’s0/60/6gives an upper bound of0\.39300\.3930; Gemma\-2\-9B’s1/111/11gives0\.36440\.3644; Qwen3\-4B’s1/501/50gives0\.09140\.0914; Phi\-3\.5’s52/9352/93gives0\.64680\.6468\. The paper reports that none attains it, which is a true statement fully supported by the data, and does not claim preservation certification anywhere\.

## Appendix RReproducibility

Table 16:Where each result in the paper is recorded\. Paths are relative to the repository root\. Row\-level records exist only where indicated; the remaining historical settings survive as aggregate counts and cannot be re\-resolved by destination\.Table[16](https://arxiv.org/html/2609.36138#A18.T16)gives the artifact location for every result in the paper\. The formal Qwen3\-8B run is reproduced bypython scripts/qwen3\_stage2\_formal\.py \-\-run\-formalunder protocol versionQWEN3\_STAGE2\_FORMAL\_V1, guarded by an approval environment variable; the configuration check is\-\-gate\-onlyand runs without opening the evaluation payload\. Determinism controls are greedy decoding, batch size11, fixed arm order, fixed sample order, a single process and a single model load, with no checkpointing or resume between scientific arms\.

All\*\.jsonl,\*\.npzand\*\.npyartifacts are tracked with Git LFS\. A checkout withoutgit lfs pullyields pointer files of approximately132132bytes that read as empty rather than failing, so any reproduction attempt should verify artifact sizes before interpreting an empty result\.

相似文章

视觉工具使用的幻觉:对图像思考的因果审计

Hugging Face Daily Papers

本文通过因果干预审计多模态大语言模型中的视觉工具使用,揭示尽管总体准确率有所提升,返回的观察结果往往缺乏因果效应。它识别出诸如“不看就调用”和“不做规划就看”等失败模式,并引入了“视觉工具使用的幻觉”这一概念。