Present but Rescaled: Chat-to-Agent Transfer of Additive Activation Steering
Summary
This paper presents the first systematic study of how additive activation steering, calibrated in single-turn chat, transfers to tool-using ReAct agents. It finds that while the injected direction reaches late layers at near-full strength across settings, the behavioral coupling varies unpredictably between models and contexts, with amplification up to 2x or attenuation, posing immediate safety concerns.
View Cached Full Text
Cached at: 07/13/26, 07:59 AM
# Present but Rescaled: Chat-to-Agent Transfer of Additive Activation Steering
Source: [https://arxiv.org/html/2607.09156](https://arxiv.org/html/2607.09156)
###### Abstract
Additive activation steering \(injecting a scaled residual\-stream direction during generation\) is calibrated almost entirely in single\-turn chat, yet the models it targets are increasingly deployed as tool\-usingReActagents\. We present the first systematic chat\-to\-agent transfer study of additive steering, coupling behavioral measurement with a representation read\-out in a matched\-information design: the same items rendered as plain chat or as aReActtool\-use episode, with matched\-norm random\-direction controls and the transcript re\-encoded every turn to exclude KV\-cache contamination\. Transfer is real but rescaled, and the right description is a dissociation: the injected direction reaches the late layers at near\-full strength in every setting and model tested \(install\-site agent\-over\-chat ratios 0\.83–1\.16 across three families\), while the behavioral coupling is reset per model and context\. On Qwen2\.5\-7B a refusal bypass vector amplifies in the agent \(T=1\.45T\\\!=\\\!1\.45, CI \[1\.20, 1\.78\],N=300N\\\!=\\\!300\); across a powered uniform\-protocol distribution the coupling spans amplification \(Gemma\-2\-9BT=2\.00T\\\!=\\\!2\.00\) to attenuation \(Yi\-1\.5\-9BT=0\.43T\\\!=\\\!0\.43, CI \[0\.29, 0\.60\]\), with no universal constant and a single clean attenuator against a universal sign\. Directional ablation of the same axis does not amplify \(T=0\.93T\\\!=\\\!0\.93, CI including 1\) while additive injection amplifies \(T=1\.50T\\\!=\\\!1\.50\), a 20\.1\-point gain difference \(CI \[13\.4, 26\.8\]\) that identifies an additive\-specific mechanism\. Two pre\-registered instruments converge to localize the rescaling to theReActformat scaffold, before any tool observation, rather than to the observation boundary where a dilution account would predict it\. The safety implication is immediate and*unpredictable*: agentic deployment amplifies steering\-based refusal bypass by up to 2\.00×\\timeson some models while others attenuate, so a deployment cannot assume a given model is safe under additive steering\.
## Introduction
Additive activation steering extracts a direction in a model’s residual stream by contrasting activations on trait\-exhibiting and trait\-suppressing completions, then injects a scaled copy of that direction during generation to shift behavior in a graded, interpretable way\(Zouet al\.[2023](https://arxiv.org/html/2607.09156#bib.bib5); Turneret al\.[2023](https://arxiv.org/html/2607.09156#bib.bib7)\)\. The technique needs no weight updates, is cheap to apply, and has been demonstrated across refusal, sycophancy, honesty, and a wide range of persona traits\(Arditiet al\.[2024](https://arxiv.org/html/2607.09156#bib.bib4); Panicksseryet al\.[2024](https://arxiv.org/html/2607.09156#bib.bib6)\)\. It is increasingly proposed as a deployment\-time control and monitoring primitive for safety applications\.
Almost all of this evidence comes from single\-turn chat\. The models being steered are increasingly deployed as tool\-using ReAct agents: they plan over multiple turns, call external tools, read returned observations, and commit to actions across an extended context\. An agent episode differs from a chat reply in ways that are not obviously neutral for an additive control method: the context is longer and its residual\-stream norm grows each turn, the model is rendered into a structured ReAct format rather than a free\-form reply, real tool observations inject ground\-truth content that can contradict a steered disposition, and behavioral commitments are often made in a planning step rather than the final natural\-language token\.
Whether a chat\-calibrated additive control keeps its behavioral grip once the same model runs as a ReAct agent is open and practically important\.
Why three nearby results do not answer it\.First, refusal\-direction*ablation*\(a destructive weight edit\) is known to transfer to agents\(Lermenet al\.[2024](https://arxiv.org/html/2607.09156#bib.bib11)\): a refusal\-ablated model completes harmful agentic tasks\. Ablation removes the direction permanently and does not compete with a growing context norm; additive injection does, which is why the additive case cannot be read off the ablation result\. Second, persona steering degrades over multi\-turn*dialogue*through KV\-cache contamination\(Kanget al\.[2026](https://arxiv.org/html/2607.09156#bib.bib14)\)\. Our loop re\-encodes the full transcript from scratch every turn by construction, ruling that mechanism out\. Third, additive steering has been applied*inside*agent loops with directions extracted in the agent context\(Yap[2026](https://arxiv.org/html/2607.09156#bib.bib12)\); what is new here is the*chat\-to\-agent*transfer specifically, with a matched chat baseline and a representation\-level read\-out alongside the behavioral measurement\.
Our approach\.We use a matched\-information five\-rung ladder \(C0–C4\) that holds the harmful instruction byte\-identical while varying only the deployment wrapper: from plain single\-turn chat \(C0\) to multi\-turn ReAct with a real deterministic tool \(C3\)\. We instrument the representation leg with a read\-only projection hook at a late read layer, measuring how much of the injected direction survives into the agent’s residual stream at a behavior\-independent install site\. We instrument the behavioral leg with a setting\-invariant parser\-based metric applied identically to the natural\-language output in every rung\. And we apply a matched\-norm random\-direction band at every behavioral result: five random unit vectors at the same injection coefficient must not produce a comparable effect, making direction\-specificity a gate rather than an observation\.
Headline finding: a dissociation\.The representational and behavioral legs of the steering effect come apart\. The injected direction survives the chat\-to\-agent transfer at near\-full or above\-chat strength in every setting and model tested: install\-site agent\-over\-chat ratios 0\.83–1\.16 across three families \(Qwen2\.5\-7B, Llama\-3\.1\-8B, Gemma\-2\-9B\-IT\)\. The behavioral coupling the surviving direction buys, however, is reset by the deployment context and the model\. On Qwen2\.5\-7B, refusal bypass amplifies \(T=1\.45T\\\!=\\\!1\.45, CI \[1\.20, 1\.78\],N=300N\\\!=\\\!300\); a direction that induces refusal on harmless prompts amplifies by at least3\.68×3\.68\\\!\\times\(lower bound, agent\-arm ceiling\)\. On Llama\-3\.1\-8B, refusal amplification is absent\. A powered uniform\-protocol distribution sharpens this per\-model reset into a two\-sided result\. Running a family roster on one design \(bypass arm,N=200N\\\!=\\\!200items, AUC over a sub\-saturation dose grid, matched\-norm random band, KV\-recompute every turn\), the agentic coupling spans clean amplification \(Gemma\-2\-9BT=2\.00T\\\!=\\\!2\.00, CI \[1\.68, 2\.43\]; Qwen2\.5\-7BT=1\.41T\\\!=\\\!1\.41\) through a boundary cluster \(Qwen2\-7B, OLMo\-2\-7B, Starling\-7B, all CI spanning 1\) to clean attenuation \(Yi\-1\.5\-9BT=0\.43T\\\!=\\\!0\.43, CI \[0\.29, 0\.60\]\): the coupling is reset per model with no universal constant*or sign*\. Families whose baseline refusal is near\-ceiling or near\-floor saturate the dose\-response with no sub\-saturation window and are reported as gate\-fails: a characterization of alignment geometry, not a coupling result\.
We further show that the amplification is specific to the additive mechanism\. On the same refusal axis in the same items, directional ablation does not amplify while additive injection amplifies, a 20\.1\-point gain difference \(CI \[13\.4, 26\.8\]\) that replicated in sign on Llama\.
Two pre\-registered instruments localize the behavioral rescaling to the ReAct format scaffold the model reads before any tool observation, not to the observation boundary a dilution reading would predict\. The activation transplant shows the coupling cannot be committed to a static prefill state; it requires continuous re\-assertion at each generation step\.
Summary of contributions:\(1\) A matched\-information, cache\-excluded, random\-controlled chat\-to\-agent transfer protocol for additive steering; \(2\) a representation\-survival/behavioral\-coupling dissociation measured across three model families; \(3\) an in\-setup additive\-vs\-ablation asymmetry \(20\.1\-point gain gap\); \(4\) a powered uniform\-protocol coupling distribution spanning amplification \(Gemma\-2\-9BT=2\.00T\\\!=\\\!2\.00\) to attenuation \(Yi\-1\.5\-9BT=0\.43T\\\!=\\\!0\.43\), establishing no universal coupling constant and, via one clean attenuator, no universal sign; \(5\) a two\-instrument convergent localization of the rescaling to frame priming, not tool observation; \(6\) a pre\-registered sign\-mechanism test narrowing the coupling\-sign question; \(7\) a safety\-relevant finding: agentic deployment amplifies the dangerous direction by up to2\.00×2\.00\\timeson some models yet attenuates it on others, so deployment safety under additive steering cannot be assumed to transfer across models\.
## Related Work
Additive steering\.The representation\-engineering recipe ofZouet al\.\([2023](https://arxiv.org/html/2607.09156#bib.bib5)\)establishes difference\-of\-means extraction and additive injection;Panicksseryet al\.\([2024](https://arxiv.org/html/2607.09156#bib.bib6)\)apply it to persona traits andArditiet al\.\([2024](https://arxiv.org/html/2607.09156#bib.bib4)\)to refusal \(for both inference\-time control and ablation\)\. Persona\-vectors work\(Chenet al\.[2025](https://arxiv.org/html/2607.09156#bib.bib8); Moskvoretskiiet al\.[2026](https://arxiv.org/html/2607.09156#bib.bib9)\)proposes chat\-extracted vectors as deployment monitoring and control primitives; we test whether the control half survives chat\-to\-agent transfer and quantify the rescaling\.Tanet al\.\([2024](https://arxiv.org/html/2607.09156#bib.bib10)\)document out\-of\-distribution brittleness of steering in chat, which we extend into the agent deployment context\.
Agentic deployment of steering\.Lermenet al\.\([2024](https://arxiv.org/html/2607.09156#bib.bib11)\)showed that*ablation*of the refusal direction transfers to Llama\-3\.1\-8B agents; neither that work nor additive\-in\-agent steering with agent\-native vectors\(Yap[2026](https://arxiv.org/html/2607.09156#bib.bib12); Chenet al\.[2026](https://arxiv.org/html/2607.09156#bib.bib23)\)measures chat\-to\-agent transfer of a chat\-extracted vector against a matched chat baseline with a behavioral coupling ratio\. Concurrent AgentLens\(Luoet al\.[2026](https://arxiv.org/html/2607.09156#bib.bib13)\)builds in\-agent safety\-steering subspaces but has no chat comparison, and workflow\-level agentic jailbreaks reach the same refused\-in\-chat\-yet\-run\-as\-agent conclusion through prompt decomposition rather than steering\(Kumar and Maple[2026](https://arxiv.org/html/2607.09156#bib.bib25)\)\. Our contribution is the transfer ratio itself\.
Degradation over multi\-turn context\.Kanget al\.\([2026](https://arxiv.org/html/2607.09156#bib.bib14)\)show persona steering degrades over multi\-turn*dialogue*via KV\-cache contamination\. We exclude this channel by re\-encoding the full transcript each turn, and localize our rescaling to single\-turn priming rather than multi\-turn accumulation\.
Output\-level and mechanistic accounts\.The Belief Dynamics account\(Bigelowet al\.[2025](https://arxiv.org/html/2607.09156#bib.bib15)\)predicts context enters only through a baseline shift, implyingT=1T\\\!=\\\!1; we falsify this on the refusal induce arm \(67\-point gap\) while it stays compatible with sycophancy \(the formal result is a cross\-behavior interaction\)\.Zhong and Li \([2026](https://arxiv.org/html/2607.09156#bib.bib17)\)show refusal is gated downstream of persona, consistent with our frame\-priming localization;Galeoneet al\.\([2026](https://arxiv.org/html/2607.09156#bib.bib16)\)find detection and intervention directions orthogonal in chat, a within\-context gap to our deployment\-context one; andFominet al\.\([2026](https://arxiv.org/html/2607.09156#bib.bib19)\)andWalsh and Barkett \([2026](https://arxiv.org/html/2607.09156#bib.bib22)\)both find internal signals that decode but do not drive behavior, corroborating the decode side of our dissociation, which we localize to the chat\-to\-agent shift\. WhereCristofano \([2026](https://arxiv.org/html/2607.09156#bib.bib24)\)transfer a shared refusal circuit across models, our axis is instead the same model’s chat\-to\-agent coupling, which has no universal sign\. AndDeng \([2026](https://arxiv.org/html/2607.09156#bib.bib20)\)formalize norm\-accumulation limits on superposition, the competing\-norm source our capability\-matched control rules out as generic degradation\.
## Experimental Design
### Models
Primary:Qwen2\.5\-7B\-Instruct\(Team[2025](https://arxiv.org/html/2607.09156#bib.bib1)\), the model used in the persona\-vectors work\(Chenet al\.[2025](https://arxiv.org/html/2607.09156#bib.bib8)\), enabling direct replication\.Cross\-model replication:Llama\-3\.1\-8B\-Instruct\(Meta AI[2024](https://arxiv.org/html/2607.09156#bib.bib2)\)\.Cross\-family generalization:Gemma\-2\-9B\-IT\(Google DeepMind[2024](https://arxiv.org/html/2607.09156#bib.bib3)\)\.Uniform\-protocol distribution:an eight\-family roster plus a scale\-axis addition \(the per\-model distribution below\)\. All models run in\-process with forward hooks via PyTorch; vLLM is not used for any steered model \(hook access requires in\-process generation\)\. Primary\-family weights are pinned to fixed revision SHAs \(logged inconfigs/provenance\.yaml\)\.
### Steering Vectors
We extract difference\-of\-means residual\-stream directions on 128 harmful versus 128 harmless last\-token activations, followingArditiet al\.\([2024](https://arxiv.org/html/2607.09156#bib.bib4)\)for refusal andPanicksseryet al\.\([2024](https://arxiv.org/html/2607.09156#bib.bib6)\)for sycophancy\. The hook injects the scaled direction at every token position of the forward pass \(or a restricted phase thereof in phase\-localization experiments\)\. A hook sanity check \(max\|Δlogit\|\|\\Delta\\text\{logit\}\|at the target layer\) verifies injection before every behavioral run; a nonzero check at a nonzero coefficient confirms the hook is active\.
### The Matched\-Information Ladder
Figure 1:Setup schematic\. A chat\-extracted directionvv\(with a matched\-norm random controlrr\) is added at layerLLover the model’s generated positions\. The same item is rendered as plain chat \(C0\) and as a ReAct episode with a real tool \(C3\), the endpoints of the matched\-information ladder\. Each cell is read two ways: the read\-leg \(residual projection ontovv\) and the write\-leg \(judge\-scored transfer ratioTT\)\.To separate deployment\-context effects from surface\-format effects, we construct a five\-rung ladder \(Figure[1](https://arxiv.org/html/2607.09156#Sx3.F1)\) holding the harmful instruction byte\-identical while varying only the wrapper:
- •C0: Plain single\-turn chat\.
- •C1: Multi\-turn chat, no tools \(same semantic content, two turns\)\.
- •C2: ReAct format with a placeholder tool call; no real tool dispatch\.
- •C3: ReAct format with a real, deterministic tool: a lookup returning a constant value, so tool\-call success is fixed and does not confound the behavioral metric\.
- •C4: Held out for external benchmarks \(tau2\-bench\)\.
The primary comparison is C0 versus C3\. C2 isolates the format scaffold from the observation insertion\. C1 confirms that simple multi\-turn context is not the driver\.
### The Setting\-Invariant Metric
Chat replies and ReAct “Thought” tokens cannot be compared directly on task success\. We score the model’s natural\-language output identically in every rung: a validated parser\-based binary \(refuses/complies for refusal; agrees/disagrees for sycophancy\), validated at 83\.8% agreement with blind human labels on a stratified 100\-item slice\. The behavioral metric is the refusal rate \(bypass arm: rate of complying; induce arm: rate of refusing\) in each arm, and the transfer ratioT=Δagent/ΔchatT=\\Delta\_\{\\text\{agent\}\}/\\Delta\_\{\\text\{chat\}\}is the agent\-over\-chat steering\-effect ratio, normalized within matched items\. The representation metric is the induced projection of the residual stream onto the unit steering direction, measured by a read\-only hook at specified layers\.
### Controls
We apply six pre\-registered controls that rule out alternative explanations\. Together they form the methodological spine of the paper; we do not cut them to save compute\.
\(A\) Matched\-norm random direction\.At every behavioral run,nrand≥5n\_\{\\text\{rand\}\}\\\!\\geq\\\!5random unit vectors scaled to‖αv‖\\\|\\alpha v\\\|are injected at the same layer and positions\. The real effect must exceed max\|random effect\|\|\\text\{random effect\}\|for a result to count as direction\-specific\. We reportΔ\(real\)−max\|Δ\(random\)\|\\Delta\(\\text\{real\}\)\\\!\-\\\!\\max\|\\Delta\(\\text\{random\}\)\|as the direction\-specific margin\.
\(B\) KV\-cache recompute\.The full transcript is re\-encoded from scratch at each turn; no cross\-turn key\-value cache is reused\. This removes the KV\-contamination mechanism ofKanget al\.\([2026](https://arxiv.org/html/2607.09156#bib.bib14)\)by construction\.
\(C\) Capability\-matched perturbation\.A benign verbosity direction scaled to match the same cross\-entropy cost \(within 7%\) as the refusal vector is run as a control\. It does not reproduce the bypass effect \(chat:−6\.9\-6\.9pp, agent:\+0\.6\+0\.6pp\), confirming that the coupling rescaling we report is direction\-specific and not generic residual\-stream perturbation at this norm\(Nguyenet al\.[2026](https://arxiv.org/html/2607.09156#bib.bib21)\)\.
\(D\) Extraction asymmetry\.We extract directions in the agent context and evaluate in chat, testing whether transfer asymmetry is systematic across extraction contexts\.
\(E\) Phase\-restricted injection\.We fire injection at only specified loop phases \(pre\-observation vs\. post\-observation\) to localize where the forward pass must be touched for the effect to obtain\.
\(F\) Ablation vs\. addition\.On the same refusal axis, we compare additive injection to directional ablation \(rank\-1 projection subtracted from weight matrices, as inArditiet al\.[2024](https://arxiv.org/html/2607.09156#bib.bib4)\) in the same harness and the same items, isolating the mechanism\.
Primary transfer and distribution cells useN≥150N\\\!\\geq\\\!150paired items \(mechanism analyses state their ownNN; same items across rungs, so within\-item differences cancel item\-level variance\), and item\-paired bootstrap CIs atB=10,000B\\\!=\\\!10\{,\}000\.
## Representation Survives; Coupling Rescales
### The Behavioral Dissociation
Table[1](https://arxiv.org/html/2607.09156#Sx4.T1)summarizes the primary behavioral results\.
Table 1:Primary behavioral transfer results\.T=Δagent/ΔchatT=\\Delta\_\{\\text\{agent\}\}/\\Delta\_\{\\text\{chat\}\}; CI = 95% item\-paired bootstrap interval\. Direction\-specific margin = real agent\-cell effect minus the max absolute matched\-random effect, in points\. Sycophancy: CI spans 1; formal result is the cross\-behavior interaction\. Qwen2\.5\-3B: saturated \(both doses collapse in the C3 frame; AUC undefined\)\.†Qwen2\-7B: single eligible sub\-saturation dose \(c12; c16 and c20 agent\-saturated\), soTfamilyT\_\{\\text\{family\}\}is the point ratio at that dose; Gemma\-2\-9B, OLMo\-2\-7B, Starling\-7B, and Yi\-1\.5\-9B are≥\\geq2\-dose grid AUCs\.BehaviorModelArmProtocolNNTT95% CIDirection\-specific?Refusal bypassQwen2\.5\-7BbypassC0→\\toC3, c163001\.45\[1\.20, 1\.78\]Yes \(\+33\+33pt margin\)Refusal induceQwen2\.5\-7BinduceC0→\\toC3, c24159≥\\geq3\.68\[2\.88, 4\.97\]Yes \(agent ceiling\)SycophancyQwen2\.5\-7BinduceC0→\\toC32010\.78\[0\.55, 1\.06\]Point est\. onlyRefusal induceLlama\-3\.1\-8BinduceC0→\\toC3, c41570\.057\[0\.0, 0\.162\]No \(below gate\)Refusal bypassLlama\-3\.1\-8BbypassC0→\\toC3160n/a—Gate fail \(near\-ceiling C0\)Ablation \(refusal\)Qwen2\.5\-7BbypassC0→\\toC3, matched3000\.93incl\. 1FlatAdditive \(refusal\)Qwen2\.5\-7BbypassC0→\\toC3, effect\-matched3001\.50—Yes \(Φ=20\.1\\Phi\\\!=\\\!20\.1pt\)Exp\-A uniform\-protocol coupling distribution:AUCTfamilyT\_\{\\text\{family\}\}Gemma\-2\-9BbypassExp\-A uniform2002\.00\[1\.68, 2\.43\]AMPLIFYAUCTfamilyT\_\{\\text\{family\}\}Qwen2\.5\-7BbypassExp\-A uniform2001\.41\[1\.14, 1\.77\]AMPLIFYAUCTfamilyT\_\{\\text\{family\}\}Qwen2\-7B†bypassExp\-A uniform2001\.26\[0\.96, 1\.66\]BOUNDARYAUCTfamilyT\_\{\\text\{family\}\}OLMo\-2\-7BbypassExp\-A uniform2001\.24\[0\.99, 1\.55\]BOUNDARYAUCTfamilyT\_\{\\text\{family\}\}Starling\-7BbypassExp\-A uniform2001\.14\[0\.92, 1\.43\]BOUNDARYAUCTfamilyT\_\{\\text\{family\}\}Yi\-1\.5\-9BbypassExp\-A uniform2000\.43\[0\.29, 0\.60\]ATTENUATEAUCTfamilyT\_\{\\text\{family\}\}Qwen2\.5\-3BbypassExp\-A uniform200–\[–, –\]sat\.Behavioral leg: refusal bypass\.On Qwen2\.5\-7B, with the refusal direction injected at layer 16 at the primary coefficient \(c16\), the bypass transfer ratio isT=1\.45T\\\!=\\\!1\.45\(CI \[1\.20, 1\.78\],N=300N\\\!=\\\!300paired items\)\. The effect is amplified in the agent: removing refusal is 45% easier in ReAct deployment than in chat\. The direction\-specific margin is 33 points; the matched\-norm random band yields near\-zero transfer, establishing that the amplification is not a generic consequence of perturbing the residual stream at this coefficient\.
On Llama\-3\.1\-8B, the bypass arm cannot be cleanly measured: the unsteered C0 refusal rate is 97\.5%, leaving a 2\.5\-point ceiling on any bypass effect\. Llama’s primary quantitative contributions to the cross\-model picture are the install\-site survival ratio \(1\.16\) and the additive\-vs\-ablation gain gap \(detailed below\)\.
Behavioral leg: refusal induce\.Inducing refusal on harmless requests yields a lower boundT≥3\.68T\\\!\\geq\\\!3\.68\(CI \[2\.88, 4\.97\],N=159N\\\!=\\\!159\) on Qwen\. This is a lower bound: the agent arm saturates at near\-100% induced refusal, so the true ratio cannot be observed from below\. The ceiling is itself the stronger claim: the direction that induces refusal in chat induces it still more completely in the agent\.
Behavioral leg: sycophancy\.A sycophancy induction vector \(layer 20\) attenuates in point estimate:T=0\.78T\\\!=\\\!0\.78\(CI \[0\.55, 1\.06\],N=201N\\\!=\\\!201\)\. This interval spans 1; we do not claim significant sycophancy attenuation\. The formal result is the cross\-behavior interaction: refusal bypass at1\.451\.45versus sycophancy at0\.780\.78yields a difference of0\.570\.57\(CI \[0\.06, 0\.99\], excluding zero\)\. Context enters the coupling in a behavior\-specific way\.
Representational leg\.A read\-only hook at a late read layer measures the induced projection of the residual stream onto the unit steering direction at a behavior\-independent install site \(fixed\-length prefix in the system prompt, before any behavioral token\)\. The agent\-over\-chat install\-site ratio is0\.9750\.975\(CI \[0\.963, 0\.986\]\) on Qwen2\.5\-7B, at chat strength in the agent\. On Llama\-3\.1\-8B the ratio is1\.161\.16\(CI \[1\.16, 1\.17\]\), above chat strength\. On Gemma\-2\-9B\-IT the ratio across read layers \(up to layer 28\) is0\.830\.83–0\.900\.90, well above a 0\.5 retention floor\. The direction does not collapse in any family\. This is the survival half of the dissociation: the direction is present in the agent at near\-chat or above\-chat strength while the behavioral coupling is rescaled\.
### Additive Injection vs\. Directional Ablation
To confirm that the amplification is specific to the additive mechanism, we run the same refusal axis in the same harness and the same items with directional ablation: the refusal projection is subtracted from the weight matrices via rank\-1 update, as inArditiet al\.\([2024](https://arxiv.org/html/2607.09156#bib.bib4)\)\.
Ablation does not amplify:T=0\.93T\\\!=\\\!0\.93\(CI including 1\)\. Additive injection at its registered coefficient \(c16\), with ablation effect\-matched to the same chat\-side bypass swing, amplifies:T=1\.50T\\\!=\\\!1\.50\. The chat\-to\-agent gain difference isΦ=20\.1\\Phi\\\!=\\\!20\.1points \(CI \[13\.4, 26\.8\], excluding zero\)\. On Llama\-3\.1\-8B the difference replicates in sign:Φ=8\.0\\Phi\\\!=\\\!8\.0\(CI \[0\.5, 15\.6\]\)\.
The interpretation: ablation removes the direction outright and does not compete with a growing context norm; its chat\-to\-agent gain is not distinguishable from flat at thisNN\. Additive injection competes additively with the context norm at each generation step; the agentic context reshapes how much behavioral traction the injected direction obtains per unit of injection norm\. These are mechanistically different effects, and the prior literature’s conjecture that additive and ablation transfer must differ becomes a within\-harness quantified result\.
### Falsifying the Output\-Level Baseline\-Shift Account
An output\-level rival model\(Bigelowet al\.[2025](https://arxiv.org/html/2607.09156#bib.bib15)\)predicts that context enters only through a log\-baseline shift, implyingT=1T\\\!=\\\!1once the baseline is accounted for\. Fit on Qwen using the same base model, this account fails on the refusal induce arm with a predicted\-versus\-observed gap of 67 points\. It remains compatible with sycophancy \(where the CI does span 1\), making the cross\-behavior interaction the formal contrast: context enters through the coupling between the direction and the behavior, not through the baseline alone\.
## Per\-Model Distribution: No Universal Constant
### Uniform\-Protocol Experiment Design
A pre\-registered, uniform\-protocol Experiment A was run across a roster of eight distinct families \(Qwen2\.5\-7B\-Instruct, Llama\-3\.1\-8B\-Instruct, OLMo\-2\-1124\-7B\-Instruct, Starling\-7B\-beta, Yi\-1\.5\-9B\-Chat, Gemma\-2\-9B\-IT, InternLM2\.5\-7B\-Chat, and Qwen2\-7B\-Instruct\), plus Qwen2\.5\-3B\-Instruct as a scale\-axis addition\. The protocol is identical across families: same extraction procedure\(Arditiet al\.[2024](https://arxiv.org/html/2607.09156#bib.bib4)\), same sub\-saturation gate \(25–60 pp chat swing, coherence≥0\.85\\geq\\\!0\.85, effect exceeds the random band\), sameN=200N\\\!=\\\!200paired items, samenrand=5n\_\{\\text\{rand\}\}\\\!=\\\!5matched\-norm random directions, AUC\-over\-dosesTfamilyT\_\{\\text\{family\}\}as the primary statistic\. Each family’s operating layer and dose grid are fixed from its own chat pilot before any agent cell runs, so dose selection cannot be tuned to the coupling outcome\.
### Gate Structure and Dose Resolution
An initial coarse\-ladder pass measured coupling for Qwen2\.5\-7B, Gemma\-2\-9B, and Qwen2\-7B and gate\-failed the rest\. A pre\-registered follow\-up then probed each gate\-failed family once more, at finer \(single\-coefficient\) dose resolution and, where the pilot slope warranted, one adjacent injection layer\. This second pass recovered three families as measured multi\-dose AUCs \(OLMo\-2\-7B and Starling\-7B as boundary, and Yi\-1\.5\-9B a clean attenuator\), by locating the narrow sub\-saturation window the coarse ladder had stepped over\. We disclose this two\-pass structure explicitly: the finer ladders search for*measurability*\(two or more sub\-saturation doses at one layer\), not for a coupling sign; Yi’s attenuating sign was pre\-registered from a prior held\-out run before this pass; and a family that fails the finer pass is closed with no further probing\.
Three families remain intrinsic gate\-fails, a characterization of alignment geometry, not a coupling result:near\-ceilingC0C\_\{0\}refusal\(Llama\-3\.1\-8B, 97\.5%\), where any bypass coefficient large enough to move behavior immediately saturates the dose\-response;near\-floorC0C\_\{0\}refusal\(InternLM2\.5\-7B\-Chat,C0=12\.1%C\_\{0\}\\\!=\\\!12\.1\\%\), where the maximum achievable swing is bounded below the 25 pp window floor; andagent\-frame saturation\(Qwen2\.5\-3B\), where both tested doses collapse C3\-frame refusal to at or below the 5% margin so no AUC is defined even though the chat side moves\. That the sub\-saturation regime for refusal bypass is available only in a subset of models, and only within a narrow dose window when it is, is itself a finding about how alignment training shapes steerability\.
### Coupling Distribution Among Passing Families
Figure 2:The two\-sided per\-model coupling distribution \(TfamilyT\_\{\\text\{family\}\}, AUC\-over\-doses\) under the uniform protocol \(N=200N\\\!=\\\!200target; Yi189189scorable,nrand=5n\_\{\\text\{rand\}\}\\\!=\\\!5\)\. Forest plot with 95% bootstrap CI: two clean amplifiers \(Gemma\-2\-9B, Qwen2\.5\-7B\), a boundary cluster \(Qwen2\-7B, OLMo\-2\-7B, Starling\-7B\), and one clean attenuator \(Yi\-1\.5\-9B, CI below 1\)\. Dashed line atT=1T\\\!=\\\!1\(flat transfer\); shaded band = boundary zone; gate\-fail families are listed below the axis\.Figure[2](https://arxiv.org/html/2607.09156#Sx5.F2)and Table[1](https://arxiv.org/html/2607.09156#Sx4.T1)give the result across the six families with a coupling estimate\. Two amplify with CIs clear of the boundary band \(Gemma\-2\-9BT=2\.00T\\\!=\\\!2\.00, Qwen2\.5\-7BT=1\.41T\\\!=\\\!1\.41\); three sit at the boundary with CIs spanning 1 \(Qwen2\-7B, OLMo\-2\-7B, Starling\-7B\); and one attenuates with its CI entirely below 1 \(Yi\-1\.5\-9BT=0\.43T\\\!=\\\!0\.43, CI \[0\.29, 0\.60\], the real effect clearing the random band at all three doses\)\.
The distribution has no universal constant*and*no universal sign: the amplifiers and the attenuator fall on opposite sides of flat transfer on one protocol, spanning0\.430\.43to2\.002\.00\(roughly4\.5×4\.5\\times\) and crossing from amplification to attenuation\. We state the honest limit: the “no universal constant” half rests on five multi\-dose families, while the “no universal sign” half is anchored on the single attenuating family \(Yi\), recovered by the finer\-ladder pass with its sign pre\-registered\. The practical consequence: chat\-calibrated steering magnitudes, and even their sign, cannot be carried to agent deployment as safety constants; the coupling must be characterized per model\.
### The Room\-to\-Push Mechanism
The same unsteered baseline shiftdb=refusalC3−refusalC0\\text\{db\}=\\text\{refusal\}\_\{C3\}\-\\text\{refusal\}\_\{C0\}underwrites two accounts that make*opposite*predictions, and neither survives\. Our pre\-registered signed predictor maps a negative db \(agent refusal erodes\) to predicted*amplification*; it is refuted out\-of\-sample on Yi, whose strongly negative db \(db=−0\.519=\\\!\-0\.519\) instead attenuates, and it abstains on the other five measured families \(NULL: db intervals spanning zero, or coupling inside the boundary band\), so it confirms on none\. The competing room\-to\-push reading \(more erosion leaves less headroom, predicting*lower*coupling\) gets Yi right but is broken by Starling\-7B, which does not erode \(db=0\.000=\\\!0\.000\) yet couples at the boundary rather than amplifying, and by OLMo\-2\-7B and Qwen2\-7B, which share erosion \(db=−0\.317=\\\!\-0\.317,−0\.333\-0\.333\) without separating\. The coupling direction is real and two\-sided; which model lands where is left as an open per\-model question, the honest successor to the now\-closed search for a universal constant\.
Two predictor out\-of\-sample verdicts are null \(Gemma and Qwen2\-7B\); Qwen2\.5\-3B is excluded \(saturation, no validTfamilyT\_\{\\text\{family\}\}\)\. The db CI for Gemma and Qwen2\.5\-7B spans zero \(predictor abstains\); Qwen2\-7B’s significant negative db predicts AMPLIFY but the resultingTTis BOUNDARY \(predictor abstains\)\. The gradient is therefore an empirical regularity rather than a confirmed one\-step predictor result\. It is, however, a candidate mechanistic account: the agent frame’s own suppression of baseline refusal competes with the additive injection, and the amount of residual headroom may shape how amplified the injection can be\.
## Mechanism: Rescaling Is Set at Frame Priming
### Two Independent Localization Instruments
We identify*where*in the agent forward pass the coupling rescaling is set, using two independently pre\-registered instruments that make the same call\.
Instrument 1: Nested input\-frame ablation\.We build a ladder holding the harmful instruction byte\-identical and adding one frame ingredient at a time: \(i\) agent role header; \(ii\) ReAct format scaffold \(“Thought / Action / Observation” grammar\); \(iii\) tool schemas; \(iv\) a real multi\-turn loop with byte\-identical system prompt isolating the observation\-insertion machinery\. Each step is gated against a matched\-norm random\-direction band\.
The endpoint chat\-to\-C3 gain on the bypass arm is\+0\.36\+0\.36ratio units \(CI \[0\.087, 0\.726\]\)\. A single ingredient carries nearly all of it: adding the ReAct format scaffold lifts the coupling by\+0\.375\+0\.375\(CI \[0\.116, 0\.689\]\), 104% of the endpoint gain, clearing the random band of 0\.31\. The role header, tool schemas, and real observation loop each produce a marginal whose interval spans zero\. The observation loop’s own contribution is−0\.014\-0\.014\(CI\[−0\.104,0\.078\]\[\-0\.104,0\.078\]\), a null\.
Figure 3:Frame\-priming localization results\.Left: Nested input\-frame ablation \(Instrument 1\)\. Each point is the marginal coupling step from adding one frame ingredient; the ReAct format scaffold carries 104% of the endpoint gain; all other steps are within the random band\.Right: Forward\-pass phase restriction \(Instrument 2\)\. Pre\-observation block injection reproduces 0\.88 of the full agent gain; post\-observation block returns exactly zero\. Both instruments localize the rescaling to a single\-turn priming step before any tool dispatch\.A post\-hoc decomposition across two additional families indicates this is not a Qwen idiosyncrasy\.111Localization runs use each family’s own coefficient \(Gemma c96; Yi c8\), a different operating point from the Exp\-A grid \(Gemma c80/c88; Yi L30 c8–c11\): the magnitudes differ slightly but the sign matches \(both amplify for Gemma, both attenuate for Yi\), and the localization question is orthogonal to the Exp\-A AUC\.On Gemma\-2\-9B \(which amplifies,TC0\-to\-C3=1\.78T\_\{\\text\{C0\-to\-C3\}\}\\\!=\\\!1\.78, CI \[1\.54, 2\.09\]\) the format step commits the amplification \(formatT=1\.76T\\\!=\\\!1\.76, CI \[1\.53, 2\.08\]\) and the tool step adds nothing \(1\.01, CI \[0\.97, 1\.05\]\)\. On Yi\-1\.5\-9B\-Chat \(which attenuates,T=0\.31T\\\!=\\\!0\.31, CI \[0\.19, 0\.47\]\) the attenuation is likewise committed at the format step \(0\.14, CI \[0\.06, 0\.24\]\), with the tool step spanning one \(1\.49, CI \[0\.91, 2\.61\]\)\.
Instrument 2: Forward\-pass phase restriction\.Holding the full C3 agent frame fixed, we fire the injection only at one phase of the forward pass: either only the pre\-observation block \(turn\-0 prefill and first\-thought generation, before any tool dispatch\) or only the post\-observation block \(re\-encoded observation context and final\-answer generation\)\. Each phase\-restricted injection is gated against its own matched\-norm random band\.
Pre\-observation block injection reproduces0\.880\.88\(CI \[0\.810\.81,0\.950\.95\]\) of the full agent coupling gain\. Post\-observation block injection reproduces exactly0\.000\.00\(CI \[0, 0\]\)\. This phase localization replicates at an independent dose \(c8\): pre\-block share 0\.90 \(CI \[0\.81, 0\.97\]\), post\-block zero\. The two instruments agree: the rescaling is set at frame adoption, before the model conditions on any tool observation \(Figure[3](https://arxiv.org/html/2607.09156#Sx6.F3)\)\.
### Behavior Specificity of the Pre\-Observation Site
To confirm that the pre\-observation block is specific to the refusal\-bypass axis rather than a generic gain stage, we inject the sycophancy induction direction at the same block and coefficient\. Sycophancy’s pre\-block coupling gain is−0\.93\-0\.93\(CI\[−1\.12,−0\.74\]\[\-1\.12,\-0\.74\]\), not positive\. The refusal\-bypass pre\-block gain is\+1\.17\+1\.17; the cross\-behavior contrast is1\.101\.10\(CI \[0\.78, 1\.48\], excluding zero\)\. The pre\-observation site amplifies refusal and attenuates sycophancy: it is a behavior\-specific rescaling site, not a generic gain stage\.
### Layer Localization and the Continuous\-Injection Requirement
Two further checks confirm the mechanism\. Sweeping injection layers\{8,12,16,20,24\}\\\{8,12,16,20,24\\\}in chat and agent, the effective layer \(argmax bypass swing\) does not move under the agentic frame \(peak at layer 16 in both; bootstrapP\(Δlayer=0\)=0\.97P\(\\Delta\\text\{layer\}\\\!=\\\!0\)\\\!=\\\!0\.97\), so the frame rescales the coupling at the same site rather than relocating it, disposing of the objection that our layer and coefficient choices were tuned for chat\. Second, a 40\-item activation transplant that overwrites the*entire*steered turn\-0 residual state into an otherwise unsteered run \(faithful whole\-state overwriter,≥90%\\geq\\\!90\\%argmax reproduction\) and generates with no further injection recovers a null transplant gap \(2\.5 pt, CI\[−5\.0,10\.0\]\[\-5\.0,10\.0\], restoration fractionρ≤0\.21\\rho\\\!\\leq\\\!0\.21, against a 50\-point steered gap\); this whole\-state overwrite is the ceiling for any prefill\-state transplant, so no positional subset is committed regardless ofNN\. The amplified coupling is maintained by continuous injection, not a committed prefill state, the mechanistic form of the additive\-vs\-ablation asymmetry: ablation removes the direction once, additive injection must re\-assert it against the context norm at every step\.
### Sign Mechanism: Context\-Signal Test
Why does the same ReAct frame amplify refusal yet \(in point estimate\) attenuate sycophancy? A pre\-registered, powered context\-signal test \(N=101N\\\!=\\\!101sycophancy responders; three observation\-content arms: irrelevant, opposing, confirming\) finds a confirming observation significantly raises the sycophancy coupling \(awth=0\.146a\_\{\\text\{wth\}\}\\\!=\\\!0\.146, CI \[0\.013, 0\.276\], excluding zero\), while the opposing arm is a powered null \(aopp=0\.047a\_\{\\text\{opp\}\}\\\!=\\\!0\.047, CI \[\-0\.113, 0\.213\]\) and the refusal arm falls inside the random band\. Context content thus couples in a behavior\-specific way but not as a clean signed law: the sign mechanism is narrowed, not resolved\. What the test does establish is orthogonal and load\-bearing: the agent rescales the coupling through which the surviving direction acts, not the direction itself\.
## Safety Implications
Amplification of refusal bypass in agents\.Additive steering can be used offensively to override safety training, and on the models where the bypass coupling amplifies, agentic deployment makes the attack*more*effective, not less\. On Qwen2\.5\-7B the bypass transfer ratio is1\.45×1\.45\\times\(N=300N\\\!=\\\!300, CI \[1\.20, 1\.78\]\), on top of a separately documented baseline erosion \(unsteered agent refusal 61% versus 72% in chat on the same items\); on Gemma\-2\-9B it is2\.00×2\.00\\times\. The hazard is compounded by unpredictability: the coupling is model\-specific and can invert \(Yi\-1\.5\-9B attenuates\), so a chat safety evaluation neither bounds the agent effect nor reveals in advance which models amplify\. A white\-box jailbreak using additive residual\-stream injection is thus strictly more effective against agentic deployment on an amplifying model, and a chat\-only evaluation understates its effect on exactly the deployments that matter\.
Monitoring survives; control must be recalibrated\.The representational survival result is good news for activation\-level safety monitoring: a probe trained on chat activations continues to fire on the steered direction in the agent at near\-chat strength\. The behavioral control half does not: a controller calibrated in chat will over\- or under\-steer in deployment, so the coupling must be calibrated in the deployment context rather than in chat\.
## Limitations
Primary model family\.Most mechanism experiments are on Qwen2\.5\-7B\. The cross\-model replication \(Llama\-3\.1\-8B, Gemma\-2\-9B\-IT\) confirms the survival half generalizes and the additive\-vs\-ablation gain gap replicates in sign on Llama, and the two pre\-registered localization instruments run in full only on Qwen2\.5\-7B, with a post\-hoc format\-vs\-tool decomposition supporting the same conclusion on Gemma\-2\-9B and Yi\-1\.5\-9B at their own coefficients\.
Sycophancy attenuation is a point estimate\.The sycophancy transfer ratio CI spans 1 \(0\.78, CI \[0\.55, 1\.06\]\)\. The formal cross\-behavior interaction is significant; the sycophancy attenuation per se is not\. We report this honestly and use the interaction as the load\-bearing contrast\.
Coupling sign: two hypotheses remain live\.The context\-signal re\-test narrowed the question: a confirming tool output raises the sycophancy coupling, but the opposing arm is a powered null and the refusal arm falls inside the random band\. A single law governing the coupling sign across behaviors and contexts is not yet established\.
Coupling distribution: the two\-sided claim is asymmetrically anchored\.Five families yield multi\-dose AUC estimates and a sixth \(Qwen2\-7B\) a single\-dose estimate; three of these \(OLMo\-2\-7B, Starling\-7B, Yi\-1\.5\-9B\) were recovered by a pre\-registered finer\-ladder second pass after the coarse pass gate\-failed them, a two\-pass structure we disclose\. The “no universal constant” claim rests on the full multi\-dose spread \(4\.5×4\.5\\times\); the “no universal sign” claim is anchored on a single attenuating family \(Yi\), whose attenuating sign was pre\-registered but whose measurability required the finer pass\. Three roster families yield no coupling estimate \(Llama\-3\.1\-8B near\-ceiling, InternLM2\.5\-7B near\-floor, Qwen2\.5\-3B agent\-saturated\)\. The pre\-registered signed baseline\-shift predictor is refuted out\-of\-sample on Yi; which model lands where is left as an open per\-model question, not a confirmed law\.
No external benchmark\.The agent evaluation uses a custom deterministic\-tool harness\. External validity on tau2\-bench or SHADE\-Arena is reserved for future work\.
## Conclusion
We present the first systematic chat\-to\-agent transfer study of additive activation steering\. The headline result is a dissociation: the injected direction survives into the agent’s residual stream at near\-full or above\-chat strength in every setting and family measured, while the behavioral coupling it buys is reset per model and deployment context\. On Qwen2\.5\-7B, refusal vectors amplify in the agent; and a powered uniform\-protocol distribution spans clean amplification \(Gemma\-2\-9BT=2\.00T\\\!=\\\!2\.00, Qwen2\.5\-7BT=1\.41T\\\!=\\\!1\.41\) through a boundary cluster to clean attenuation \(Yi\-1\.5\-9BT=0\.43T\\\!=\\\!0\.43\), establishing that the agentic coupling has no universal constant and \(on a single clean attenuator\) no universal sign, only per\-model rescaling\.
The same refusal axis removed by directional ablation does not amplify, while additive injection amplifies, a gain difference of 20\.1 points \(CI \[13\.4, 26\.8\]\) that mechanistically separates the two control primitives\. Two pre\-registered instruments converge to localize the behavioral rescaling to the ReAct format scaffold before any tool observation\. The activation transplant shows the coupling cannot be locked into a prefill state but requires continuous re\-assertion at each generation step, the mechanistic form of the additive\-vs\-ablation asymmetry\.
The operational consequence: on Qwen2\.5\-7B, additive injection makes agentic deployment*more*exploitable than chat \(1\.45×1\.45\\times, atop baseline erosion\)\. Monitoring survives; control must be recalibrated per deployment, not read off the surviving representation\.
## Acknowledgments
The author thanks the maintainers of the open\-sourcerefusal\_directionandpersona\_vectorsrepositories for the extraction tooling this work builds on\.
## References
- A\. Arditi, O\. Obeso, A\. Syed, D\. Paleka, N\. Panickssery, W\. Gurnee, and N\. Nanda \(2024\)Refusal in language models is mediated by a single direction\.arXiv2406\.11717\.External Links:[Link](https://arxiv.org/abs/2406.11717)Cited by:[Appendix G](https://arxiv.org/html/2607.09156#A7.p1.9),[Introduction](https://arxiv.org/html/2607.09156#Sx1.p1.1),[Related Work](https://arxiv.org/html/2607.09156#Sx2.p1.1),[Steering Vectors](https://arxiv.org/html/2607.09156#Sx3.SSx2.p1.1),[Controls](https://arxiv.org/html/2607.09156#Sx3.SSx5.p7.1),[Additive Injection vs\. Directional Ablation](https://arxiv.org/html/2607.09156#Sx4.SSx2.p1.1),[Uniform\-Protocol Experiment Design](https://arxiv.org/html/2607.09156#Sx5.SSx1.p1.4)\.
- E\. Bigelow, D\. Wurgaft, Y\. Wang, N\. Goodman, T\. Ullman, H\. Tanaka, and E\. S\. Lubana \(2025\)Belief dynamics reveal the dual nature of in\-context learning and activation steering\.arXiv2511\.00617\.External Links:[Link](https://arxiv.org/abs/2511.00617)Cited by:[Appendix D](https://arxiv.org/html/2607.09156#A4.p1.4),[Related Work](https://arxiv.org/html/2607.09156#Sx2.p4.1),[Falsifying the Output\-Level Baseline\-Shift Account](https://arxiv.org/html/2607.09156#Sx4.SSx3.p1.1)\.
- R\. Chen, A\. Arditi, H\. Sleight, O\. Evans, and J\. Lindsey \(2025\)Persona vectors: monitoring and controlling character traits in language models\.arXiv2507\.21509\.External Links:[Link](https://arxiv.org/abs/2507.21509)Cited by:[Related Work](https://arxiv.org/html/2607.09156#Sx2.p1.1),[Models](https://arxiv.org/html/2607.09156#Sx3.SSx1.p1.1)\.
- Y\. Chen, V\. Siu, Y\. Liu, D\. Song, and C\. Wang \(2026\)Controlling tool use with heading\-specific activation steering\.arXiv2607\.05790\.External Links:[Link](https://arxiv.org/abs/2607.05790)Cited by:[Related Work](https://arxiv.org/html/2607.09156#Sx2.p2.1)\.
- T\. Cristofano \(2026\)Universal refusal circuits across LLMs: cross\-model transfer via trajectory replay and concept\-basis reconstruction\.arXiv2601\.16034\.External Links:[Link](https://arxiv.org/abs/2601.16034)Cited by:[Related Work](https://arxiv.org/html/2607.09156#Sx2.p4.1)\.
- Y\. Deng \(2026\)GEMS: geometric constraints enable multi\-semantic superposition in LLMs\.arXiv2606\.19946\.External Links:[Link](https://arxiv.org/abs/2606.19946)Cited by:[Related Work](https://arxiv.org/html/2607.09156#Sx2.p4.1)\.
- M\. Fomin, E\. David, and A\. LeVi \(2026\)Internal\-state probes read the situation, not the action: three negative results for pre\-action misalignment monitoring\.arXiv2606\.30449\.External Links:[Link](https://arxiv.org/abs/2606.30449)Cited by:[Related Work](https://arxiv.org/html/2607.09156#Sx2.p4.1)\.
- C\. Galeone, A\. Ettorre, M\. Park, G\. Ettorre, and D\. Ligorio \(2026\)Perfect detection, failed control: the geometry of knowing vs\. steering in language models\.arXiv2606\.24952\.External Links:[Link](https://arxiv.org/abs/2606.24952)Cited by:[Related Work](https://arxiv.org/html/2607.09156#Sx2.p4.1)\.
- Google DeepMind \(2024\)Gemma 2: improving open language models at a practical size\.arXiv2408\.00118\.External Links:[Link](https://arxiv.org/abs/2408.00118)Cited by:[Models](https://arxiv.org/html/2607.09156#Sx3.SSx1.p1.1)\.
- D\. Kang, Z\. Liu, N\. Ma, Y\. Huang, Z\. Tan, and M\. Jiang \(2026\)Prompt\-activation duality: improving activation steering via attention\-level interventions\.arXiv2605\.10664\.External Links:[Link](https://arxiv.org/abs/2605.10664)Cited by:[Introduction](https://arxiv.org/html/2607.09156#Sx1.p4.1),[Related Work](https://arxiv.org/html/2607.09156#Sx2.p3.1),[Controls](https://arxiv.org/html/2607.09156#Sx3.SSx5.p3.1)\.
- A\. Kumar and C\. Maple \(2026\)Refused in chat, written in code: workflow\-level jailbreak construction in IDE coding agents\.arXiv2607\.03968\.External Links:[Link](https://arxiv.org/abs/2607.03968)Cited by:[Related Work](https://arxiv.org/html/2607.09156#Sx2.p2.1)\.
- S\. Lermen, M\. Dziemian, and G\. Pimpale \(2024\)Applying refusal\-vector ablation to Llama 3\.1 70b agents\.arXiv2410\.10871\.External Links:[Link](https://arxiv.org/abs/2410.10871)Cited by:[Introduction](https://arxiv.org/html/2607.09156#Sx1.p4.1),[Related Work](https://arxiv.org/html/2607.09156#Sx2.p2.1)\.
- W\. Luo, Q\. Zhang, Y\. Quan, M\. Jin, J\. Cai, C\. Xiao, J\. Niu, and Z\. Xiang \(2026\)AgentLens: interpretable safety steering via mechanistic subspaces for multi\-turn coding agent\.arXiv2606\.22673\.External Links:[Link](https://arxiv.org/abs/2606.22673)Cited by:[Related Work](https://arxiv.org/html/2607.09156#Sx2.p2.1)\.
- Meta AI \(2024\)The Llama 3 herd of models\.arXiv2407\.21783\.External Links:[Link](https://arxiv.org/abs/2407.21783)Cited by:[Models](https://arxiv.org/html/2607.09156#Sx3.SSx1.p1.1)\.
- V\. Moskvoretskii, D\. Glandorf, J\. Medina Moreira, T\. Käser, and R\. West \(2026\)Tracing persona vectors through LLM pretraining\.arXiv2605\.13329\.External Links:[Link](https://arxiv.org/abs/2605.13329)Cited by:[Related Work](https://arxiv.org/html/2607.09156#Sx2.p1.1)\.
- T\. Nguyen, T\. A\. Nguyen, S\. Alemohammad, and R\. G\. Baraniuk \(2026\)Minimizing collateral damage in activation steering\.arXiv2605\.01167\.External Links:[Link](https://arxiv.org/abs/2605.01167)Cited by:[Controls](https://arxiv.org/html/2607.09156#Sx3.SSx5.p4.2)\.
- N\. Panickssery, N\. Gabrieli, J\. Schulz, M\. Tong, E\. Hubinger, and A\. M\. Turner \(2024\)Steering Llama 2 via contrastive activation addition\.arXiv2312\.06681\.External Links:[Link](https://arxiv.org/abs/2312.06681)Cited by:[Introduction](https://arxiv.org/html/2607.09156#Sx1.p1.1),[Related Work](https://arxiv.org/html/2607.09156#Sx2.p1.1),[Steering Vectors](https://arxiv.org/html/2607.09156#Sx3.SSx2.p1.1)\.
- D\. Tan, D\. Chanin, A\. Lynch, D\. Kanoulas, B\. Paige, A\. Garriga\-Alonso, and R\. Kirk \(2024\)Analyzing the generalization and reliability of steering vectors\.arXiv2407\.12404\.External Links:[Link](https://arxiv.org/abs/2407.12404)Cited by:[Related Work](https://arxiv.org/html/2607.09156#Sx2.p1.1)\.
- Q\. Team \(2025\)Qwen2\.5 technical report\.arXiv2412\.15115\.External Links:[Link](https://arxiv.org/abs/2412.15115)Cited by:[Models](https://arxiv.org/html/2607.09156#Sx3.SSx1.p1.1)\.
- A\. M\. Turner, L\. Thiergart, G\. Leech, D\. Udell, J\. J\. Vazquez, U\. Mini, and M\. MacDiarmid \(2023\)Steering language models with activation engineering\.arXiv2308\.10248\.External Links:[Link](https://arxiv.org/abs/2308.10248)Cited by:[Introduction](https://arxiv.org/html/2607.09156#Sx1.p1.1)\.
- C\. Walsh and E\. Barkett \(2026\)Representation without control: testing the realization effect in language models\.arXiv2605\.25151\.External Links:[Link](https://arxiv.org/abs/2605.25151)Cited by:[Related Work](https://arxiv.org/html/2607.09156#Sx2.p4.1)\.
- J\. Q\. Yap \(2026\)Behavioral steering in a 35b MoE language model via SAE\-decoded probe vectors: one agency axis, not five traits\.arXiv2603\.16335\.External Links:[Link](https://arxiv.org/abs/2603.16335)Cited by:[Introduction](https://arxiv.org/html/2607.09156#Sx1.p4.1),[Related Work](https://arxiv.org/html/2607.09156#Sx2.p2.1)\.
- V\. Zhong and Q\. Li \(2026\)Refusal lives downstream of persona in chat models\.ICML 2026 Mechanistic Interpretability Workshop / arXiv2606\.26161\.External Links:[Link](https://arxiv.org/abs/2606.26161)Cited by:[Related Work](https://arxiv.org/html/2607.09156#Sx2.p4.1)\.
- A\. Zou, L\. Phan, S\. Chen, J\. Campbell, P\. Guo, R\. Ren, A\. M\. Turner, B\. Robey, Z\. Kolter, M\. Fredrikson, and D\. Hendrycks \(2023\)Representation engineering: a top\-down approach to AI transparency\.arXiv2310\.01405\.External Links:[Link](https://arxiv.org/abs/2310.01405)Cited by:[Introduction](https://arxiv.org/html/2607.09156#Sx1.p1.1),[Related Work](https://arxiv.org/html/2607.09156#Sx2.p1.1)\.
## Appendix APreregistrations and Reproducibility
Every behavioral experiment was pre\-registered before pilot launch, with the live/die criteria and the verdict\-mapping engine committed to the public repository before any data were collected\. The registration documents indocs/includebypass\_transfer\_prereg\.md,sycophancy\_transfer\_prereg\.md,sign\_retest\_prereg\.md,phase\_specificity\_prereg\.md,induce\_spec\_phase\_prereg\.md,effective\_layer\_shift\_prereg\.md, the uniform\-protocol masterexp\_a\_coupling\_distribution\_prereg\.md, and the per\-family preregistrations\*\_exp\_a\_prereg\.md\. The two\-pass rescue structure of the per\-model distribution was itself pre\-registered \(exp\_a\_strong\_rescue\_prereg\.mdandexp\_a\_attenuator\_yi\_finegrid\_prereg\.md\), fixing the finer dose ladders and the per\-family operating layer from chat data alone, before any agent cell ran, so dose selection could not be tuned to the coupling outcome\. Yi\-1\.5\-9B’s attenuating sign was registered from a prior held\-out run before the finer\-ladder pass that established its measurability\.
The three primary families are pinned to exact revisions inconfigs/provenance\.yaml\(Qwen2\.5\-7Ba09a354, Llama\-3\.1\-8B0e9e39f2, Gemma\-2\-9B11c9b30\); the additional Exp\-A roster families were run on their default HuggingFace revision, recorded in the per\-run logs rather than pinned to a SHA\. Every rollout is logged as JSONL \(prompt, all read activations, action, observation, parse status, score, seed\)\. Scripts reproducing every figure and table are inscripts/; the coupling estimator isscripts/exp\_a\_coupling\_auc\.py\(unit\-tested to recover the reference Qwen factor\), and the mechanism analyses arescripts/w5\_analyze\.py,scripts/bd\_headtohead\.py, andscripts/toggle\_analyze\.py\. Decoding is greedy throughout, so each rollout is deterministic and the only randomness in the intervals is the item bootstrap\.
## Appendix BThe Uniform\-Protocol Coupling Distribution in Full
Table[2](https://arxiv.org/html/2607.09156#A2.T2)gives, per family, the pilot\-selected operating layer, the committed dose grid, the per\-dose chat swing, and the agent\-frame status that determined inclusion\. The coarse\-ladder pass measured Qwen2\.5\-7B, Gemma\-2\-9B, and Qwen2\-7B; the pre\-registered finer\-ladder pass recovered OLMo\-2\-7B and Starling\-7B as boundary families and Yi\-1\.5\-9B as the clean attenuator by locating the narrow sub\-saturation window the coarse grid had stepped over \(for Yi, layer 30 with a one\-coefficient ladder found four eligible doses where the original layers 24/28/32 had found at most one\)\.
Table 2:Per\-family dose grids and gate outcomes\. Chat swing is the C0 refusal\-rate change at the operating layer; a dose is sub\-saturation if the C3 real refusal rate stays above the5%5\\%margin\. “Gate\-fail” families never present two sub\-saturation doses at one layer\.The coupling factorTfamilyT\_\{\\text\{family\}\}is the trapezoid area of the agent dose\-response divided by that of the chat dose\-response over the sub\-saturation grid, with a single shared item\-paired bootstrap \(B=10,000B\\\!=\\\!10\{,\}000\) reused across every cell of the family so the interval propagates the C0/C3 and real/random pairing\. Each per\-dose effect is the real refusal\-rate change minus the matched\-norm random band \(nrand=5n\_\{\\text\{rand\}\}\\\!=\\\!5\); a dose enters the integral only if it clears the5%5\\%two\-frame saturation margin\. The signed BD\-band label uses the pre\-committedAMP\_MARGIN=1\.10\\text\{AMP\\\_MARGIN\}\\\!=\\\!1\.10: amplify if theTTinterval lower bound exceeds1\.101\.10, attenuate if the upper bound is below1\.01\.0, and boundary otherwise\. Qwen2\.5\-3B is the informative near\-miss: its chat side moves cleanly \(39 and 54 pp\) but both agent doses collapse refusal to at or below the saturation margin, so the AUC is undefined even though the direction is behaviorally live; the informational point ratio at c12 is1\.141\.14\.
## Appendix CRepresentation Survival: The Projection Half\-Life
The dissociation’s read half is measured directly by replaying the committed rollout under a read\-only hook and projecting the residual stream onto the unit steering direction at read layers\{16,20,24,27\}\\\{16,20,24,27\\\}\(Qwen2\.5\-7B,N=201N\\\!=\\\!201, 2814 instrumented rollouts;configs/w5\_mechanism\.json\)\. The decisive quantity is the induced alignment,ct\(\+v\)−ct\(unsteered\)c\_\{t\}\(\+v\)\\\!\-\\\!c\_\{t\}\(\\text\{unsteered\}\), the projection contributed by the injection, read in chat \(C0 reply\) versus agent \(C3 final\-answer tokens, the ones the behavioral metric scores\)\. If the agent context diluted the steering signal, this would be smaller in the agent\. It is not \(Table[3](https://arxiv.org/html/2607.09156#A3.T3)\)\.
Table 3:Induced alignment \(real minus unsteered\), chat C0 reply / agent C3 final answer \(N=201N\\\!=\\\!201source set; per\-condition scorablenn190–201\)\. The injected component decays with network depth \(layer 20→\\to27\) identically in chat and agent; at the last layer it is if anything slightly larger in the agent\.The controls are textbook: the matched\-norm random direction induces≈0\\approx 0alignment ontovv\(\+0\.001\+0\.001at\+25\+25\),−v\-vinduces the negative mirror \(−0\.073\-0\.073\), and\+40\+40induces more than\+25\+25\. There is a small, direction\-specific dip at each tool\-observation boundary \(−0\.009\-0\.009\[−0\.013,−0\.005\-0\.013,\-0\.005\] at\+25\+25, paired real minus baseline,P\(<0\)=1\.0P\(<0\)\\\!=\\\!1\.0\), but it does not accumulate: the induced component measured across successive turns is flat\-to\-rising \(retention1\.131\.13–1\.141\.14\), so the transient dip washes out within the next thought\. Items with a larger boundary dip do not lose more behavioral effect \(first\-dropr=−0\.18r\\\!=\\\!\-0\.18\[−0\.32,−0\.03\-0\.32,\-0\.03\] at\+25\+25\), the opposite of the dilution prediction\. The chat\-extracted direction is therefore present in the agent’s residual stream at the output layer at full chat strength while the behavior attenuates: the loss is downstream of the representation, a routing rather than a dilution effect\.
## Appendix DBelief\-Dynamics Head\-to\-Head
The strongest output\-level rival\(Bigelowet al\.[2025](https://arxiv.org/html/2607.09156#bib.bib15)\)holds that context enters only through the unsteered baseline and combines additively with steering in log\-odds, so one chat\-fitted slopekkmust predict the agent’s steered behavior once the baseline is measured\. We give it its most faithful output\-level rendering,logitp\(X,m\)=logitp0\(X\)\+km\\operatorname\{logit\}p\(X,m\)\\\!=\\\!\\operatorname\{logit\}p\_\{0\}\(X\)\\\!\+\\\!k\\,m, fitkkon chat alone per arm, and treat agent cells as out\-of\-sample predictions \(configs/bd\_headtohead\.json,B=20,000B\\\!=\\\!20\{,\}000\)\.
Table 4:Belief\-Dynamics predicted vs observed refusal transfer \(ε=0\.5\\varepsilon\\\!=\\\!0\.5\)\. Gap = observed agent effect minus predicted; every refusal cell is underpredicted\.The decisive cell is induce c24: both baselines are0/1590/159, so after identical smoothing the model predicts the agent curve equals the chat curve \(Tpred≈1T\_\{\\text\{pred\}\}\\\!\\approx\\\!1\), yet the observedT=3\.68T\\\!=\\\!3\.68and the observed count sits∼\\sim72 orders of magnitude outside the model’s binomial prediction, invariant acrossε∈\[0\.1,2\.0\]\\varepsilon\\\!\\in\\\!\[0\.1,2\.0\]\. Sigmoid geometry does not rescue the headline bypass c16 cell either: the chat logit shift applied to the agent baseline predictsT=1\.00T\\\!=\\\!1\.00against an observed1\.451\.45, so essentially none of that amplification is mechanical\. Refitting one slope per context, both arms \(disjoint items, opposite signs\) demand the same missing multiplier,kagent/kchat=1\.59k\_\{\\text\{agent\}\}/k\_\{\\text\{chat\}\}\\\!=\\\!1\.59\[1\.37, 1\.85\] \(bypass\) and1\.601\.60\[1\.49, 1\.76\] \(induce\)\. Sycophancy, by contrast, is compatible with the additive model at every coefficient \(k\-ratio1\.021\.02\[0\.70, 1\.48\]\), which is why the cross\-behavior interaction, not the sycophancy point estimate, is the formal contrast: the model needs exactly the context\-dependent coupling term whose refusal signature it cannot reproduce\.
## Appendix ESecond Model: Llama\-3\.1\-8B
The read leg replicates and the write\-leg coupling does not, which is the cross\-model spine \(configs/refusal\_transfer\_result\_llama\_core\_ci\.json\)\. On Llama the chat\-extracted refusal direction again survives into the last layer with the agent carrying*more*of the injected component \(final\-segment projection ratio1\.141\.14, CI \[1\.11, 1\.17\], excluding 1; Qwen1\.321\.32\[1\.19, 1\.46\], same direction, smaller magnitude\)\. But the∼\\sim1\.6×\\timesagent amplification that made Qwen the headline does not reappear: at Llama’s one clearly sub\-saturation operating point the induce arm flips hard the other way \(T=0\.057T\\\!=\\\!0\.057\[0\.0, 0\.16\], the agent barely responds where chat moves\+22\+22pp\), the bypass arm is inconclusive \(chat gate fails\), and higher doses are ceiling\-compressed \(T≈1\.0T\\\!\\approx\\\!1\.0–1\.081\.08\)\. Llama also shows no agent baseline erosion \(91%→98%91\\%\\\!\\to\\\!98\\%, versus Qwen’s72%→61%72\\%\\\!\\to\\\!61\\%\)\. This is a second, independent violation of output\-level additivity, in the*opposite*direction from Qwen’s: strictly additive models fail both ways, which strengthens the case for a per\-model coupling term while removing any temptation to assign it a universal sign\.
## Appendix FPhase Localization: Two Instruments
Two pre\-registered instruments locate the behavioral rescaling to the ReAct format scaffold the model reads before any tool observation, not the observation boundary a dilution account predicts \(Qwen2\.5\-7B, bypass, c16\)\.Instrument 1 \(nested input\-frame ablation\)builds a ladder holding the harmful instruction byte\-identical and adding one frame ingredient at a time \(agent role header; ReAct “Thought/Action/Observation” grammar; tool schemas; a real multi\-turn loop with byte\-identical observations\)\. The endpoint chat\-to\-agent gain is\+0\.36\+0\.36ratio units \(CI \[0\.087, 0\.726\]\); the ReAct*format step*alone carries\+0\.375\+0\.375\(CI \[0\.116, 0\.689\]\),104%104\\%of it, while the role header, tool schemas, and observation loop each contribute a marginal step whose interval spans zero \(the observation loop’s own contribution is−0\.014\-0\.014, CI \[−0\.104,0\.078\-0\.104,0\.078\], a null\)\.Instrument 2 \(forward\-pass phase restriction\)holds the full agent frame fixed and fires injection at only one phase: pre\-observation block injection reproduces0\.880\.88\(CI \[0\.81, 0\.95\]\) of the full agent gain, post\-observation block injection reproduces0\.000\.00\. Both instruments place the rescaling at a single\-turn priming step\. A behavior\-specificity check confirms the pre\-observation site is not a generic gain stage: injecting the sycophancy direction there yields a pre\-block coupling gain of−0\.93\-0\.93\(CI \[−1\.12,−0\.74\-1\.12,\-0\.74\]\), the opposite sign, so the site amplifies refusal and attenuates sycophancy\.
## Appendix GAdditive\-vs\-Ablation and Capability\-Matched Controls
Additive vs\. ablation\.On the same refusal axis, the same harness, and the same items, directional ablation \(rank\-1 projection subtracted from the weight matrices,Arditiet al\.[2024](https://arxiv.org/html/2607.09156#bib.bib4)\) does not amplify \(T=0\.93T\\\!=\\\!0\.93, CI including 1\) while additive injection \(ablation effect\-matched to its chat\-side bypass swing\) amplifies \(T=1\.50T\\\!=\\\!1\.50\), a chat\-to\-agent gain difference of20\.120\.1points \(CI \[13\.4, 26\.8\], excluding zero\)\. The difference replicates in sign on Llama\-3\.1\-8B \(Φ=8\.0\\Phi\\\!=\\\!8\.0, CI \[0\.5, 15\.6\]\), through a model\-specific mechanism \(Qwen drives the gap through additive amplification, Llama through ablation attenuation\), so the sign of the asymmetry is the cross\-model invariant, not its magnitude\.Capability\-matched control\.A benign verbosity direction scaled to match the refusal vector’s cross\-entropy cost \(within7%7\\%\) does not reproduce the bypass effect: its chat effect of−6\.9\-6\.9pp \(CI \[−11\.9,−2\.5\-11\.9,\-2\.5\]\) becomes a null\+0\.6\+0\.6pp \(CI \[−4\.4,5\.6\-4\.4,5\.6\]\) in the agent rather than amplifying, so the coupling gain is specific to the refusal direction’s behavioral channel and not a generic consequence of perturbing the residual stream at this norm\.
## Appendix HSign Mechanism: The Context\-Signal Re\-Test
To ask whether the deployment context’s own content sets the coupling, we hold the agent frame byte\-identical and vary only the bound tool observation across three content arms —irrelevant, opposing, and confirming—and measure the resulting coupling against a matched\-norm random\-direction band \(configs/sign\_retest\_result\_joint\_20260615\.json\)\. On sycophancy \(N=101N\\\!=\\\!101responders\), a confirming observation significantly raises the coupling \(awth=\+0\.146a\_\{\\text\{wth\}\}\\\!=\\\!\+0\.146\[0\.013, 0\.276\], clearing the random band\), while the opposing arm is a powered null \(aopp=\+0\.047a\_\{\\text\{opp\}\}\\\!=\\\!\+0\.047\[−0\.113,0\.213\-0\.113,0\.213\]\) and the direct contrastawth−aopp=\+0\.100a\_\{\\text\{wth\}\}\\\!\-\\\!a\_\{\\text\{opp\}\}\\\!=\\\!\+0\.100\[−0\.043,0\.249\-0\.043,0\.249\] does not exclude zero\. On refusal \(N=200N\\\!=\\\!200, c16\), the opposing\-content effect does*not*clear the matched\-norm random band \(aopp=−0\.0018a\_\{\\text\{opp\}\}\\\!=\\\!\-0\.0018,\|random\|=0\.0018\|\\text\{random\}\|\\\!=\\\!0\.0018; joint registered verdict BAND\_FAIL\), so an earlier apparent refusal trim does not survive the content\-matched control and is not interpreted\. The sign mechanism is therefore narrowed to a behavior\-specific, with\-axis content sensitivity on sycophancy, not resolved into a signed law across behaviors, and the sign of the rescale remains the open per\-model question\.Similar Articles
Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions
This paper identifies KV-cache contamination as a failure mode for activation steering in dialogue and proposes GCAD, a method that extracts steering signals from prompt contributions and applies token-level gating to improve long-horizon coherence, achieving substantial gains on multi-turn benchmarks.
Controlling Tool Use with Heading-Specific Activation Steering
This paper investigates whether tool-use decisions in large language models have stable internal representations that can be extracted and manipulated via activation steering, demonstrating that heading-specific steering vectors can suppress unnecessary tool use across five open-source models and three domains. The geometric analysis reveals that tool-invocation steps exhibit diffuse, bimodal alignment rather than the clean linear structure expected for parametrically grounded concepts.
A Geometric Account of Activation Steering through Angle-Norm Decomposition
This paper analyzes linear activation steering in language models by decomposing interventions into angular and radial components. It finds that concepts are primarily encoded in angular structure, but norm adjustments are crucial for stability, supporting spherical steering methods while showing that additive coefficients conflate geometry.
Beyond Steering Vector: Flow-based Activation Steering for Inference-Time Intervention
This paper introduces FLAS, a flow-based activation steering method that learns a concept-conditioned velocity field to steer language model activations at inference time. On the AxBench benchmark, FLAS is the first learned method to consistently outperform in-context prompting on held-out concepts without per-concept tuning.
Closed-Loop Neural Activation Control in Vision-Language-Action Models
Proposes CTRL-STEER, a closed-loop framework for adaptive steering of vision-language-action models using time-varying control signals, achieving better trade-off between concept regulation and task success without retraining.