Where Steering Signals Come From: Activation Source Selection in Activation Steering
Summary
This paper investigates how the choice of source activations influences activation steering in language models, finding that execution-boundary states (where the model is about to produce target behavior) yield stronger signals, and introduces tail subtraction to improve steering stability.
View Cached Full Text
Cached at: 07/29/26, 09:54 AM
# Activation Source Selection in Activation Steering
Source: [https://arxiv.org/html/2607.25270](https://arxiv.org/html/2607.25270)
## Where Steering Signals Come From: Activation Source Selection in Activation Steering
Jiaran Ye1,3,\*Lingxu Ran3,\*Zijun Yao3Chenpeng Wang6 Yong Jiang5Lei Hou3Juanzi Li3Liangming Pan1,2,4,† 1MOE Key Laboratory of Computational Linguistics, Peking University 2School of Computer Science, Peking University 3Department of Computer Science and Technology, Tsinghua University 4Beijing Academy of Artificial Intelligence, Beijing, China 5Tsinghua Shenzhen International Graduate School, Tsinghua University 6YiXin\-AILab, YIXIN, Beijing, China \{yejr23,rlx22\}@mails\.tsinghua\.edu\.cnliangmingpan@pku\.edu\.cn
###### Abstract
Activation steering controls language models by adding vectors or features to hidden states at inference time, but the upstream source of these steering signals is often treated as a secondary detail\. We study this source choice as*activation source selection*: the combination of source context and activation readout policy used to collect the hidden states from which a steering signal is built\. Holding the downstream intervention fixed, we show across three instruction\-tuned models and four steering task families that changing only the source activations substantially changes steering success\. We further find that effective steering is not explained simply by whether the desired behavior appears in the source text\. Instead, strong signals come from*execution\-boundary*states, where the model is about to produce or continue the target behavior\. This pre\-/post\-realization distinction explains why answer\-based sources sometimes work: their useful component aligns with execution\-boundary directions rather than target appearance alone\. Building on this view, we introduce tail subtraction, which removes shared prompt and continuation semantics from boundary states and yields cleaner, more stable steering signals\. Overall, our results suggest that steering depends on representations of what the model is*about to do*, not merely on what has already appeared\.
Where Steering Signals Come From: Activation Source Selection in Activation Steering
Jiaran Ye1,3,\*Lingxu Ran3,\*Zijun Yao3Chenpeng Wang6Yong Jiang5Lei Hou3Juanzi Li3Liangming Pan1,2,4,†1MOE Key Laboratory of Computational Linguistics, Peking University2School of Computer Science, Peking University3Department of Computer Science and Technology, Tsinghua University4Beijing Academy of Artificial Intelligence, Beijing, China5Tsinghua Shenzhen International Graduate School, Tsinghua University6YiXin\-AILab, YIXIN, Beijing, China\{yejr23,rlx22\}@mails\.tsinghua\.edu\.cnliangmingpan@pku\.edu\.cn
††footnotetext:\*Equal contribution\.†Corresponding author\.## 1Introduction
Large language models are usually controlled by changing their inputs or updating their parameters\.*Activation steering*offers a third option: at inference time, we add a small signal to the model’s hidden states so that the same model becomes more likely to behave in a desired way\(Subramani et al\.,[2022](https://arxiv.org/html/2607.25270#bib.bib25); Turner et al\.,[2023](https://arxiv.org/html/2607.25270#bib.bib31); Zou et al\.,[2023](https://arxiv.org/html/2607.25270#bib.bib37)\)\. Such signals can make answers more concise, more emoji\-heavy, more likely to mention a target entity, or more likely to refuse unsafe requests\.
A steering signal, however, must first be found somewhere\. In much of the activation\-steering pipeline, this is done by running the model on examples related to the desired behavior, reading hidden activations from those examples, and turning the activations into a vector\. A common recipe is to use source prompts that exhibit the target behavior, then average hidden states from those prompts\. For example, to make answers emoji\-heavy, one may collect activations from prompts and completions that contain emoji\-heavy responses and use their mean as a style direction\.
Figure 1:Overview of our activation source view\. With the downstream steering pipeline fixed, different source activations yield substantially different steering behavior\. Pre\-realization execution\-boundary states are more effective than post\-realization target\-appearance traces\.These upstream choices are common in practice, but often treated as setup details\. Most recent work starts after the source activations have been collected: it asks how to turn those activations into a better vector or feature, and how to inject the resulting signal into the model more effectively\(Rimsky et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib20); Konen et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib13); He et al\.,[2025](https://arxiv.org/html/2607.25270#bib.bib9); Postmus and Abreu,[2024](https://arxiv.org/html/2607.25270#bib.bib19); You et al\.,[2026](https://arxiv.org/html/2607.25270#bib.bib36)\)\. We instead ask an earlier question:where should the steering signal come from?This matters because a steering vector is estimated from the source activations themselves; if those activations contain different information, the same construction and injection method can yield different steering behavior\. We study this upstream choice as*activation source selection*: the combination ofsource context, the text used to elicit activations, andreadout policy, the rule that selects or aggregates hidden states\. Figure[1](https://arxiv.org/html/2607.25270#S1.F1)illustrates this view\.
Our main finding is that effective steering is not explained simply by whether the target behavior is visible in the source text\. A common target\-source recipe assumes that activations are useful when they are read from contexts that already realize the desired behavior\. We call this the*target\-appearance*intuition\. However, a state that records what the model has already said may not be the best handle for making it do so again\. Strong steering often comes instead from*execution\-boundary*states, where the model is about to produce or continue the target behavior\.
We test this view across three instruction\-tuned models and four families of steering tasks\. Holding the downstream intervention fixed, changing only the source activations produces large differences in steering success\. Prompt\-only final\-token sources are strongest on average, while answer\-only sources are among the weakest, even though they visibly contain the target behavior\. Relation completion further separates states before an answer is produced from states after it appears: pre\-realization boundary states steer reliably, whereas post\-realization answer states are much weaker\. When answer\-based sources do work, their useful signal partly aligns with an execution\-boundary direction\.
This interpretation also leads to a practical construction\. Execution\-boundary states contain useful target preparation, but also chat formatting, the user question, and generic continuation semantics\. We introduce*tail subtraction*, which subtracts a matched tail\-only state from each target\-conditioned boundary state\. This isolates a cleaner boundary signal and improves over positive\-only and ordinary contrastive baselines\.
#### Contributions\.
Our contributions are fourfold\. First, we formulate*activation source selection*as an explicit upstream design choice in activation steering, separating the source context from the readout policy, and operationally distinguish execution\-boundary states from post\-realization trace states\. Second, we show that this choice has a large empirical effect under a fixed downstream intervention\. Third, we provide evidence that effective steering is better explained by execution\-boundary states, where the model is about to produce the target behavior, than by target appearance alone\. Finally, we introduce tail subtraction, a simple source\-construction method that removes shared local continuation semantics and yields cleaner boundary signals\.
## 2Background and Related Work
### 2\.1Activation Steering
Activation steering modifies model behavior at inference time by intervening on hidden representations\. Prior work has shown that such interventions can be constructed from activation additions, contrastive directions, representation engineering directions, mean\-centered vectors, style or persona directions, and behavior\-specific directions such as refusal\(Subramani et al\.,[2022](https://arxiv.org/html/2607.25270#bib.bib25); Li et al\.,[2023](https://arxiv.org/html/2607.25270#bib.bib15); Turner et al\.,[2023](https://arxiv.org/html/2607.25270#bib.bib31); Zou et al\.,[2023](https://arxiv.org/html/2607.25270#bib.bib37); Jorgensen et al\.,[2023](https://arxiv.org/html/2607.25270#bib.bib12); Rimsky et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib20); Konen et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib13); Arditi et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib3); Chen et al\.,[2025](https://arxiv.org/html/2607.25270#bib.bib6)\)\. More recent work studies alternative intervention rules, including activation scaling, conceptor\-based steering, conditional refusal steering, dynamic steering, and geometry\-aware rotations\(Stoehr et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib22); Postmus and Abreu,[2024](https://arxiv.org/html/2607.25270#bib.bib19); Lee et al\.,[2025](https://arxiv.org/html/2607.25270#bib.bib14); Stolfo et al\.,[2025](https://arxiv.org/html/2607.25270#bib.bib23); Wang et al\.,[2025a](https://arxiv.org/html/2607.25270#bib.bib32); Li et al\.,[2025](https://arxiv.org/html/2607.25270#bib.bib16); Scialanga et al\.,[2025](https://arxiv.org/html/2607.25270#bib.bib21); Su et al\.,[2025](https://arxiv.org/html/2607.25270#bib.bib24); You et al\.,[2026](https://arxiv.org/html/2607.25270#bib.bib36)\)\. Sparse\-autoencoder methods provide a related feature\-level route for interpreting and steering model behavior\(Templeton et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib29); Huben et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib11); Chalnev et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib5); He et al\.,[2025](https://arxiv.org/html/2607.25270#bib.bib9)\)\.
This paper uses additive vector steering as its basic intervention setting\. Letfθf\_\{\\theta\}be an autoregressive language model with hidden statehℓ,th\_\{\\ell,t\}at layerℓ\\elland generation positiontt\. Given a steering vectorvv, additive steering modifies the hidden state as
h~ℓ,t=hℓ,t\+αv,\\tilde\{h\}\_\{\\ell,t\}=h\_\{\\ell,t\}\+\\alpha v,\(1\)whereα\\alphacontrols intervention strength\. Our focus is not to introduce a new form of Eq\.[1](https://arxiv.org/html/2607.25270#S2.E1), nor to optimize layer or strength as independent objects of study\. Instead, we hold the downstream intervention family fixed and ask which upstream hidden states should provide the evidence from whichvvis constructed\.
### 2\.2Activation Source Selection
A steering vector is computed from activations collected before intervention\. We call these activations*source activations*\. For a source examplezz, let a source\-context functionccproduce a token sequencex1:Tz\(z\)=c\(z\)x^\{\(z\)\}\_\{1:T\_\{z\}\}=c\(z\), such as a sentence containing the target concept, a prompt that makes the target behavior likely, or a matched control context\. Given this context and a source layerℓs\\ell\_\{s\}, a readout policyρ\\rhoselects or aggregates positions from the resulting activations:
sℓsρ,c\(z\)=ρ\(hℓs,1\(c\(z\)\),…,hℓs,Tz\(c\(z\)\)\)\.s\_\{\\ell\_\{s\}\}^\{\\rho,c\}\(z\)=\\rho\\\!\\left\(h\_\{\\ell\_\{s\},1\}\(c\(z\)\),\\ldots,h\_\{\\ell\_\{s\},T\_\{z\}\}\(c\(z\)\)\\right\)\.\(2\)For example,ρ\\rhomay read the final token state or average over source tokens\.
We define an*activation source selection*asσ=\(c,ρ\)\\sigma=\(c,\\rho\): the choice of source context and readout policy\. For a set of positive and optional negative source examples, this selection induces source activation setsSσ,ℓs\+S\_\{\\sigma,\\ell\_\{s\}\}^\{\+\}andSσ,ℓs−S\_\{\\sigma,\\ell\_\{s\}\}^\{\-\}, which a vector\-construction ruleggmaps into a steering vector:
v=g\(Sσ,ℓs\+,Sσ,ℓs−\)\.v=g\(S\_\{\\sigma,\\ell\_\{s\}\}^\{\+\},S\_\{\\sigma,\\ell\_\{s\}\}^\{\-\}\)\.\(3\)This notation separates two design choices that are often coupled in practice: how hidden states are converted into a direction, and which hidden states are made available to that conversion\. Prior work usually treats source text and readout policy as setup choices; a recent study discusses source\-text effects, but only for a small set of persona traits and reaches a different conclusion from ours\(Chen et al\.,[2025](https://arxiv.org/html/2607.25270#bib.bib6)\)\. In this paper, wedo not introduce a new vector\-construction rule; instead, we instantiateggwith standard methods such as mean directions, PCA directions and SAE\-based features, and study how changingσ\\sigmaaffects steering under the fixed downstream intervention in Eq\.[1](https://arxiv.org/html/2607.25270#S2.E1)\.
## 3Experimental Setup
Our experiments isolate the upstream source activations used to construct the steering signal\. We fix the intervention family, vector construction, layer–strength search, and evaluation protocol, and vary only the*activation source condition*: the source context and readout positions instantiatingσ=\(c,ρ\)\\sigma=\(c,\\rho\)from Section[2\.2](https://arxiv.org/html/2607.25270#S2.SS2)\.
### 3\.1Models
We evaluate on three open\-weight instruction\-tuned models:Gemma\-2\-9B\-IT\(Team,[2024a](https://arxiv.org/html/2607.25270#bib.bib26)\),Qwen2\.5\-7B\-Instruct\(Yang et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib35)\), andLlama\-3\.1\-8B\-Instruct\(Team,[2024b](https://arxiv.org/html/2607.25270#bib.bib28)\)\. This lets us test whether activation source effects persist across model families\.
### 3\.2Steering Tasks
We evaluate steering on four heterogeneous task families:Entitysteering makes generations mention ordinary concepts such ascat,coffee, ormusic\(11 targets\)\(Templeton et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib29); He et al\.,[2025](https://arxiv.org/html/2607.25270#bib.bib9); Wu et al\.,[2025](https://arxiv.org/html/2607.25270#bib.bib34)\);Persona / Styleinduces affective or stylistic responses such as happy, angry, flattering, rhetorical\-question, or emoji\-heavy outputs \(7 targets\)\(Konen et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib13); Chen et al\.,[2025](https://arxiv.org/html/2607.25270#bib.bib6)\);Rejectinduces refusal\-like behavior \(1 target\)\(Arditi et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib3); Wang et al\.,[2025b](https://arxiv.org/html/2607.25270#bib.bib33)\); andNonsenseinduces factually wrong or nonsensical but readable answers \(1 target\)\(Chen et al\.,[2025](https://arxiv.org/html/2607.25270#bib.bib6)\)\. Together, these tasks test content insertion, response manner, action policy, and semantic correctness\.
#### Evaluation\.
For each behavioral task, the steered model answers held\-out generic questions independent of the target attribute\. A fixed Qwen2\.5\-14B\-Instruct judge\(Yang et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib35)\)assigns a binary label for target expression and generation validity, where valid outputs must remain normal and readable rather than collapsing into flooding, repetition, or malformed text\. We validate the automatic judge on a stratified sample of behavioral outputs, finding high agreement with independent validation labels\. A full rescoring with Gemma\-3\-27B\(Team,[2025](https://arxiv.org/html/2607.25270#bib.bib27)\)also preserves the main source\-condition pattern; details are in Appendix[A](https://arxiv.org/html/2607.25270#A1.SS0.SSS0.Px7)\. We count a generation as successful only when both criteria hold, report the resultingsteering success rate, and retain the best score over the searched layer–strength grid\. Appendix Section[A](https://arxiv.org/html/2607.25270#A1)gives behavioral task data, judge prompts, and example outputs\.
### 3\.3Vector Construction
We use five standard ways to convert source activations into steering directions\(Turner et al\.,[2023](https://arxiv.org/html/2607.25270#bib.bib31); Jorgensen et al\.,[2023](https://arxiv.org/html/2607.25270#bib.bib12); Subramani et al\.,[2022](https://arxiv.org/html/2607.25270#bib.bib25); Zou et al\.,[2023](https://arxiv.org/html/2607.25270#bib.bib37); Rimsky et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib20); Templeton et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib29); Chalnev et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib5); He et al\.,[2025](https://arxiv.org/html/2607.25270#bib.bib9)\):Meanaverages positive source activations;Diff\-Meanaverages positive\-minus\-negative differences;PCAuses the first principal component of the centered positive source activations;Diff\-PCAuses the first principal component of the centered positive\-minus\-negative differences; andSAEselects sparse\-autoencoder features consistently activated by the source activations\. SAE details are provided in Appendix[C](https://arxiv.org/html/2607.25270#A3)\.
### 3\.4Intervention Protocol
We use the additive steering intervention defined in Section[2\.1](https://arxiv.org/html/2607.25270#S2.SS1)\. Vectors are injected only during generation, not during source\-prompt prefill\.
For each source condition and construction method, we sweep source layers, intervention layers, and steering strengths, with strength ranges chosen separately for each method and model\. We report the best accuracy over this grid as the source condition’s steering potential under the fixed downstream protocol\.
## 4Does Activation Source Selection Change Steering?
### 4\.1Coarse Controlled Activation Source Grid
We begin with a coarse grid of activation source choices mirroring common activation\-sourcing decisions\. The experiment askswhether changing only the upstream hidden states can change steering success\.
The grid first varies the source text used to elicit activations\.Prompt\-onlycontains the target instruction and query but no model response, matching prompt\- or query\-based refusal and safety sources\(Arditi et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib3); Wang et al\.,[2025b](https://arxiv.org/html/2607.25270#bib.bib33)\)\.Prompt\-and\-answeradditionally includes a response exhibiting the target behavior, as in representation engineering and contrastive activation addition\(Zou et al\.,[2023](https://arxiv.org/html/2607.25270#bib.bib37); Rimsky et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib20)\)\.Answer\-onlyuses only a target\-bearing response or completion, mirroring target\-text averaging choices used in mean\-centered and style steering\(Jorgensen et al\.,[2023](https://arxiv.org/html/2607.25270#bib.bib12); Konen et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib13)\)\. Table[1](https://arxiv.org/html/2607.25270#S4.T1)illustrates the three source\-text conditions\.
We also vary the hidden\-state readout\.Last tokentakes the final non\-padding token, a natural decoder\-only summary position used in representation\-engineering and refusal work\(Zou et al\.,[2023](https://arxiv.org/html/2607.25270#bib.bib37); Arditi et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib3); Wang et al\.,[2025b](https://arxiv.org/html/2607.25270#bib.bib33)\)\.Sequence meanaverages all non\-padding source tokens, following target\-text averaging choices in mean\-centered and style\-vector methods\(Jorgensen et al\.,[2023](https://arxiv.org/html/2607.25270#bib.bib12); Konen et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib13)\)\. Thus, answer\-only with sequence\-mean is closest to averaging activations from target\-bearing completions\.
Table 1:Example source texts for the emoji target\.
### 4\.2Main Result: Source Activations Strongly Affect Steering
For each model, source condition, readout, and vector\-construction method, we keep the best score over the same source\-layer, intervention\-layer, and strength grid\. We then macro\-average over the evaluated vector\-construction methods\. Within behavioral tasks, we average questions within each target, average targets within the entity and persona families, and treat reject and nonsense as single\-target families\. Overall scores are the unweighted average of the four family scores, preventing larger target families from dominating\. Unless otherwise stated, later sections use the same aggregation protocol\. Appendix[F](https://arxiv.org/html/2607.25270#A6)reports a more detailed Gemma breakdown by vector\-construction method and target family\.
Table 2:Overall activation source comparison\. Scores average steering success over persona, entity, nonsense, and reject\. Gemma and Llama use five vector\-construction methods; Qwen uses the four non\-SAE methods\.Source textHidden readoutSuccess ratePromptAnswerLastMeanLlamaQwenGemmaAvg✓✓0\.3360\.5060\.5860\.476✓✓0\.1640\.4000\.4350\.333✓✓✓0\.0600\.2100\.4730\.248✓✓✓0\.2310\.5270\.5210\.426✓✓0\.0780\.2410\.3300\.216✓✓0\.1190\.1930\.2810\.198
Activation source selection substantially changes steering success\. Prompt\-only with last\-token readout is strongest on average, while answer\-only sources are weakest despite visibly exhibiting the target behavior\. Thus, changing only the source activations can substantially shift steering success under the same downstream protocol\.
### 4\.3First Semantic Interpretation
The weak answer\-only result is notable because this condition is closest to the common averaging\-from\-target\-completions recipe\. If visible target text were sufficient, answer\-only sources should be reliable; instead, their weakness suggests that useful signal also comes from the computation that leads into or sustains the target behavior\.
The strongest conditionprompt\-only \+ lastalso supports this view\. The prompt\-only source already specifies the target behavior, but contains no realized answer; the last\-token readout then captures the state just before generation, where that instruction has been integrated into an imminent response\.
The next section tests more directly whether steerability comes from target appearance after the behavior appears, or from execution\-boundary semantics before it is produced\.
## 5What Semantics Make Source Activations Steerable?
### 5\.1Target Appearance vs\. Execution\-Boundary Semantics
The coarse comparison in Section[4\.3](https://arxiv.org/html/2607.25270#S4.SS3)suggests that activation source selection depends not only on whether the source text contains the target, but on where the activation is taken\. We distinguish two operational source\-state types\.
For a target behaviorBB, write a source instance asz=\(p,a\)z=\(p,a\), whereppis an open\-context token sequence andaais a continuation constructed to realize or sustainBB\. At layerℓ\\ell, we define the*execution\-boundary state*as
bℓ\(z\)=hℓ,\|p\|\(p\),b\_\{\\ell\}\(z\)=h\_\{\\ell,\|p\|\}\(p\),namely, the state read immediately before this target\-bearing continuation\. If a prefixa1:ka\_\{1:k\}has already realizedBBin the continuation, we define the corresponding*post\-realization trace state*as
rℓ,k\(z\)=hℓ,\|p\|\+k\(p⊕a1:k\)\.r\_\{\\ell,k\}\(z\)=h\_\{\\ell,\|p\|\+k\}\(p\\oplus a\_\{1:k\}\)\.Here,⊕\\oplusdenotes token\-sequence concatenation\. Different choices of the source\-context functionccand readout policyρ\\rhotherefore yield boundary states, post\-realization traces, or aggregations over multiple positions\.
Post\-realization states may make the target readable\(Alain and Bengio,[2016](https://arxiv.org/html/2607.25270#bib.bib2); Belinkov and Glass,[2019](https://arxiv.org/html/2607.25270#bib.bib4); Wu et al\.,[2025](https://arxiv.org/html/2607.25270#bib.bib34)\), but readability need not imply a causal handle for reproducing the behavior elsewhere\. We therefore examine whether steering efficacy is better explained by readout at an execution boundary than by target appearance alone\.
### 5\.2Refined Semantic Source Comparison
Table 3:Refined boundary prompts for the conceptdog\. The table abbreviates demonstrations for space; experiments use three demonstrations foriclandhybrid\. Additional target\-family examples are in Appendix[A](https://arxiv.org/html/2607.25270#A1.SS0.SSS0.Px5)\.To isolate execution\-boundary states, we add prompts that make the target continuation imminent while leaving the answer open\. For each target, we use three ingredients: a target instruction, complete target\-bearing question–answer examples, and open question–prefix pairs whose assistant prefix stops just before the target behavior would be realized or continued\. We combine these ingredients into three boundary constructions \(Table[3](https://arxiv.org/html/2607.25270#S5.T3)\):instruction, which prepends only the explicit target instruction;icl, which prepends three complete demonstrations; andhybrid, which prepends both the instruction and demonstrations\. For all refined sources, we use thelast\-tokenreadout from Section[4\.1](https://arxiv.org/html/2607.25270#S4.SS1), taking the final hidden state of the open source context before the target is produced\. The three constructions therefore vary the source\-context functionccwhile using the same last\-token readout policyρ\\rho, which extracts the boundary statebℓ\(z\)b\_\{\\ell\}\(z\)\. Construction details, negative\-source matching, and additional examples are provided in Appendix[A](https://arxiv.org/html/2607.25270#A1.SS0.SSS0.Px5); sequence\-mean readout ablations are reported in Appendix[E](https://arxiv.org/html/2607.25270#A5)\.
We rerun the activation source comparison with these sources and compare them with the coarse conditions from Section[4](https://arxiv.org/html/2607.25270#S4)\.
Instr\.\+LastICL\+LastHybrid\+LastPrompt\-only\+LastPrompt\+Ans\.\+MeanAnswer\-only\+Mean0202040406060808033\.833\.8161638\.938\.931\.631\.623\.123\.111\.911\.955\.555\.520\.120\.1646450\.650\.652\.752\.719\.319\.355\.755\.728\.128\.162\.562\.558\.658\.652\.152\.128\.128\.1Overall success rate \(%\)Llama\-3\.1\-8B\-InstructQwen2\.5\-7B\-InstructGemma\-2\-9B\-ITFigure 2:Refined semantic source comparison\. The first three conditions are boundary sources with last\-token readout; the final three are coarse baselines from Table[2](https://arxiv.org/html/2607.25270#S4.T2)\. Bars show overall success rate averaged over persona, entity, nonsense, and reject results\.Figure[2](https://arxiv.org/html/2607.25270#S5.F2)compares the refined boundary sources against the strongest coarse patterns from Section[4](https://arxiv.org/html/2607.25270#S4)\. Thehybrid \+ lastcondition is best for all three models, outperforming the post\-realizationanswer\-only \+ meansource and the stronger coarse baselinesprompt\-only \+ lastandprompt\-and\-answer \+ mean\. This supports the execution\-boundary interpretation: source activations are most reliable when read at the boundary before target production\. The competitiveinstruction \+ lastresult, together with the further gain fromhybrid \+ last, suggests that instruction following and example\-based continuation cues contribute to the same execution\-oriented signal\. A held\-out analysis over five random validation–test splits likewise finds that bothprompt\-only \+ lastandhybrid \+ lastoutperformanswer\-only \+ meanafter all three hyperparameters are selected exclusively on validation questions \(Appendix[D](https://arxiv.org/html/2607.25270#A4)\)\.
The weakericl \+ lastresult suggests that demonstrations alone may not reliably induce the intended boundary here, consistent with prior work on the sensitivity of in\-context learning to demonstration format and task structure\(Min et al\.,[2022](https://arxiv.org/html/2607.25270#bib.bib18); Garg et al\.,[2022](https://arxiv.org/html/2607.25270#bib.bib7)\)\.
## 6Why Do Target\-Appearance Sources Sometimes Work?
### 6\.1Mixed\-State Explanation
The execution\-boundary account raises an immediate question: if effective steering depends on states where the model is preparing or continuing a target behavior, why do answer\-based sources sometimes work\(Turner et al\.,[2023](https://arxiv.org/html/2607.25270#bib.bib31); Jorgensen et al\.,[2023](https://arxiv.org/html/2607.25270#bib.bib12); Konen et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib13)\)? The same pattern appears in our coarse grid: answer\-only sources are weak, butprompt\-and\-answer \+ meanis competitive\.
Our account is that many target\-appearance sources are*mixed states*, not pure post\-realization traces\. Their text visibly realizes the target, but some answer positions also reflect the model continuing that behavior\. This is especially natural for style and persona tasks: once an answer has begun in a particular style, later states help sustain that style in the following tokens\. A mean readout can therefore capture an execution\-continuation component even though the target is already visible; boundary\-oriented sources read this component more directly\. An answer\-level mean readout may aggregate states from multiple positions, including states involved in continuing the target behavior, rather than select a single trace staterℓ,k\(z\)r\_\{\\ell,k\}\(z\)\. We refer to such an aggregated source activation as a mixed state\.
### 6\.2Relation Tasks as an Unentangled Test
To test this cleanly, we use relation completion tasks, wheretarget appearance is not naturally entangled with continuing the same behavior\. The model must execute a mapping from inputxxto answeryybeforeyyis produced; afteryyappears, that local computation is largely complete\.
We use 16 relation\-completion tasks commonly used to study in\-context learning and task/function vectors\(Garg et al\.,[2022](https://arxiv.org/html/2607.25270#bib.bib7); Hendel et al\.,[2023](https://arxiv.org/html/2607.25270#bib.bib10); Todd et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib30)\)\. For each task, we compare three boundary sources from Section[5\.2](https://arxiv.org/html/2607.25270#S5.SS2)\(instruction,icl, andhybrid\), all read at the last token before the answer, against apost\-realizationsource that containsyyand uses sequence\-mean readout\. Table[4](https://arxiv.org/html/2607.25270#S6.T4)illustrates the separation: the boundary rows require producingsmall, while the post\-realization row already contains it\. Operationally, for a relation pair\(x,y\)\(x,y\), boundary sources end withUser:xx,Assistant:, optionally preceded by the task instruction and/or three complete demonstrations; post\-realization sources include the completed answeryy\. Full construction details and the relation\-task inventory are listed in Appendix[B](https://arxiv.org/html/2607.25270#A2)\.
Table 4:Illustrative source formats for the antonym relation task\. Boundary sources read the last token before the answer; the post\-realization source contains the answer and uses sequence\-mean readout\.For each source condition, task, and model, we construct steering vectors and measure whether open\-tail generation follows the intended relation\. Scores use the aggregation protocol from Section[4\.2](https://arxiv.org/html/2607.25270#S4.SS2)\.
Instr\.\+LastICL\+LastHybrid\+LastPost\-real\.\+Mean01010202030304040505060607070Success rate \(%\)LlamaQwenGemmaFigure 3:Relation\-task source comparison\. Boundary sources use the three source text types from Section[5\.2](https://arxiv.org/html/2607.25270#S5.SS2)with last\-token readout\. The post\-realization source contains the target answer and uses sequence\-mean readout\.Figure[3](https://arxiv.org/html/2607.25270#S6.F3)shows a consistent separation: all three boundary sources steer far better than the answer\-bearing mean source, withhybrid \+ lastmost stable across models\. This matches the mixed\-state account\. When answer appearance is no longer tied to continuing the same behavior, target appearance alone provides little signal; useful information is concentrated before the model executes the relation\. The nonzeroPost\-real\. \+ Meanscores suggest that target appearance may still have a small but real effect, just much weaker than boundary\-state information\.
### 6\.3Removing Execution\-Relevant Components
The previous experiment separates target appearance from execution\-boundary states by changing the task family\. We now test the same explanation within answer\-based sources: if an answer\-mean vector works as a mixed state, removing its execution\-aligned component should weaken steering\.
We useDiff\-PCAvectors from boundary source activations as a representative direction for the execution\-relevant component\. PCA\-based steering extracts a dominant latent direction from activation variation\(Subramani et al\.,[2022](https://arxiv.org/html/2607.25270#bib.bib25); Zou et al\.,[2023](https://arxiv.org/html/2607.25270#bib.bib37)\); applying PCA to positive\-minus\-negative differences makes Diff\-PCA target the main contrastive variation, analogous to contrastive PCA’s use of principal components to isolate enriched structure\(Abid et al\.,[2017](https://arxiv.org/html/2607.25270#bib.bib1)\)\. Diff\-PCA is also empirically strongest in our main comparison, with the same pattern visible in the Gemma method breakdown in Appendix[F](https://arxiv.org/html/2607.25270#A6)\. For each model and target, we construct an execution directionuu, take an answer\-based steering vectorvv, and remove its projection ontouu:
vablated=v−v⊤u‖u‖22u\.v\_\{\\mathrm\{ablated\}\}=v\-\\frac\{v^\{\\top\}u\}\{\\\|u\\\|\_\{2\}^\{2\}\}u\.We then rerun steering withvablatedv\_\{\\mathrm\{ablated\}\}for two answer\-based baselines:answer\-only \+ mean, our direct post\-realization baseline, andprompt\-and\-answer \+ mean, the stronger mixed source from Section[4\.2](https://arxiv.org/html/2607.25270#S4.SS2)\.
Table 5:Projection ablation of the Diff\-PCA execution component from answer\-based mean vectors\. Scores are average steering success\.Table[5](https://arxiv.org/html/2607.25270#S6.T5)shows that the ablation consistently reduces steering\. The drop is modest foranswer\-only \+ mean, whose baseline is already low, and larger forprompt\-and\-answer \+ mean, where it removes 8–22 percentage points across models\. Thus the stronger answer\-based source is effective partly because it contains the same execution\-relevant direction recovered from boundary states; after removing that component, the remaining post\-realization signal steers much less reliably\.
## 7Phase\-Aware Source Construction: Tail Subtraction
### 7\.1Motivation: Execution States Still Contain Noise
The previous sections argue that execution\-boundary source activations can provide more effective steering signals than pure target\-appearance states\. However, these states also include chat\-format, question, answer\-prefix, and generic continuation features unrelated to the target behavior\.
Consider thehybridexample in Table[3](https://arxiv.org/html/2607.25270#S5.T3)\. Its full source makes the continuationdoglikely, while the final tailUser: What helps you feel less stressed … Assistant: During the week, I feel less stressed with mycan be presented alone\. This tail\-only prompt preserves local continuation semantics while removing target conditioning\. High full–tail similarity would show that boundary activations contain shared local semantics,not only the target\-specific execution signal\.
For each task, we construct pairedhybrid/tail\-only prompts and read last\-token states from the best preceding source layer\. We average cosine similarities betweenhfullh\_\{\\mathrm\{full\}\}andhtailh\_\{\\mathrm\{tail\}\}, and betweenhfull−htailh\_\{\\mathrm\{full\}\}\-h\_\{\\mathrm\{tail\}\}andhfullh\_\{\\mathrm\{full\}\}\. Table[6](https://arxiv.org/html/2607.25270#S7.T6)reports the results\.
The pattern is clearest for Gemma and Qwen: full boundary states are close to their matched tails, while the residual is much less aligned with the full state\. Llama shows a weaker tail component, consistent with its weaker activation source steering\.
Table 6:Similarity diagnostic for fullhybridstates and matched tail\-only states\. Tail\-only keeps the final user question and assistant prefix but removes target conditioning\.Thus, execution\-boundary states contain local answer\-boundary semantics even without target conditioning\. The next subsection turns this into a source\-construction procedure\.
### 7\.2Tail Subtraction
The diagnostic above suggests a simple phase\-aware construction\. For each target\-conditioned boundary prompt, we subtract a matched tail\-only prompt that keeps the final user question and assistant prefix but removes target conditioning\. Letσfull=\(cfull,ρlast\)\\sigma\_\{\\mathrm\{full\}\}=\(c\_\{\\mathrm\{full\}\},\\rho\_\{\\mathrm\{last\}\}\)andσtail=\(ctail,ρlast\)\\sigma\_\{\\mathrm\{tail\}\}=\(c\_\{\\mathrm\{tail\}\},\\rho\_\{\\mathrm\{last\}\}\)denote the paired full and tail\-only boundary sources\. They share the same local query, assistant prefix, and readout position, whilectailc\_\{\\mathrm\{tail\}\}removes the target\-conditioning context\. The tail\-subtracted source activation is
s~ℓ\(i\)=sℓρlast,cfull\(zi\)−sℓρlast,ctail\(zi\)\.\\tilde\{s\}^\{\(i\)\}\_\{\\ell\}=s\_\{\\ell\}^\{\\rho\_\{\\mathrm\{last\}\},c\_\{\\mathrm\{full\}\}\}\(z\_\{i\}\)\-s\_\{\\ell\}^\{\\rho\_\{\\mathrm\{last\}\},c\_\{\\mathrm\{tail\}\}\}\(z\_\{i\}\)\.The downstream intervention is unchanged; only the source activations become residuals\.
We apply the same vector\-construction families from Section[3\.3](https://arxiv.org/html/2607.25270#S3.SS3):Tail\-Meanaverages residuals, andTail\-PCAtakes their leading component\. We compare against the corresponding positive\-only and negative\-subtracted baselines under the same evaluation protocol\. Table[7](https://arxiv.org/html/2607.25270#S7.T7)reports macro\-average success across the four task families\.
Table 7:Tail subtraction results\. Scores are macro\-average steering success rates across persona, entity, nonsense, and reject tasks\.Tail\-subtracted sources consistently perform best\. In both mean and PCA families, they improve over positive\-only and negative\-subtracted baselines\. The gain is largest for Gemma and Qwen, matching the diagnostic evidence for a large shared tail component; Llama improves less, but tail subtraction is still best in both families\. This supports isolating boundary\-state signal by subtracting the matched local continuation state\.
## 8Conclusion
We showed that activation steering depends strongly on the upstream hidden states used to build the steering signal\. Across tasks, effective sources are better explained by execution\-boundary semantics—what the model is about to do—than by target appearance alone\. This view explains the mixed behavior of answer\-based sources and motivates tail subtraction, which isolates cleaner boundary signals\. Source construction should therefore be treated as a first\-class design choice in activation steering\.
## Acknowledgements
This work was supported in part by the Beijing Major Science and Technology Project under Contract No\. Z251100008125054\. This work was supported by the Beijing Academy of Artificial Intelligence \(BAAI\)\.
## Limitations
Our experiments cover several standard vector\-construction methods, but not the full space of possible steering\-vector estimators\. More structured constructions, supervised probes, nonlinear directions, multi\-vector bases, or feature\-level decompositions may interact with activation source selection in different ways\.
We also fix the downstream intervention to simple additive activation steering\. Other intervention rules, such as token\- or layer\-dependent injection, subspace interventions, projection\-based methods, or steering through selected neurons and sparse features, are not evaluated here\. Our results therefore characterize how activation sources affect additive steering, not all possible forms of representation intervention\.
Finally, our empirical scope is limited to three instruction\-tuned open\-weight models around the 7B–9B scale and a finite set of behavioral and relation tasks\. We do not test base models, larger scales, multilingual or multimodal settings, long\-context generation, or multi\-turn steering\. Our behavioral scores use an automatic judge, although relation tasks provide complementary exact\-match evaluation\. Although we validate the automatic judge on a stratified sample, the full behavioral evaluation still relies on automatic labels and may miss subtle quality or safety issues\. We also report best scores over a common layer–strength grid; these scores are useful for comparing steering potential under a shared diagnostic protocol, but absolute numbers may overstate what would be obtained with a single preselected deployment setting\. Tail subtraction is likewise only a first phase\-aware construction, and better ways to isolate execution\-boundary signals remain open\.
## Ethical Considerations
This work studies activation steering, which can change model behavior at inference time\. While such methods can support controllable generation and model analysis, they may also be misused to induce unwanted styles, refusals, or incorrect outputs\. We therefore treat our experiments as controlled diagnostics of activation source selection rather than as a deployment recipe\.
Some targets, including refusal\-like and nonsensical\-output behaviors, are used only to measure whether source activations can control action policy and semantic correctness\. These are synthetic evaluation targets, not recommended user\-facing behaviors\. Practical systems using steering should include task\-specific safety checks and evaluation for unintended behavior changes\.
We use open\-weight models and automatic judges, which can introduce systematic evaluation errors\. Reported scores should therefore be read as controlled comparisons under a fixed judge, not as complete assessments of real\-world safety or truthfulness\. The study does not use private user data or human\-subject data\.
## References
- Abid et al\. \(2017\)Abubakar Abid, Vivek Kumar Bagaria, Martin J\. Zhang, and James Y\. Zou\. 2017\.[Contrastive principal component analysis](https://arxiv.org/abs/1709.06716)\.*CoRR*, abs/1709\.06716\.
- Alain and Bengio \(2016\)Guillaume Alain and Yoshua Bengio\. 2016\.[Understanding intermediate layers using linear classifier probes](https://arxiv.org/abs/1610.01644)\.*CoRR*, abs/1610\.01644\.
- Arditi et al\. \(2024\)Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda\. 2024\.[Refusal in language models is mediated by a single direction](https://doi.org/10.52202/079017-4322)\.*Advances in Neural Information Processing Systems*, 37\.
- Belinkov and Glass \(2019\)Yonatan Belinkov and James R\. Glass\. 2019\.[Analysis methods in neural language processing: A survey](https://doi.org/10.1162/TACL_A_00254)\.*Trans\. Assoc\. Comput\. Linguistics*, 7:49–72\.
- Chalnev et al\. \(2024\)Sviatoslav Chalnev, Matthew Siu, and Arthur Conmy\. 2024\.[Improving steering vectors by targeting sparse autoencoder features](https://doi.org/10.48550/ARXIV.2411.02193)\.*CoRR*, abs/2411\.02193\.
- Chen et al\. \(2025\)Runjin Chen, Andy Arditi, Henry Sleight, Owain Evans, and Jack Lindsey\. 2025\.[Persona vectors: Monitoring and controlling character traits in language models](https://doi.org/10.48550/ARXIV.2507.21509)\.*CoRR*, abs/2507\.21509\.
- Garg et al\. \(2022\)Shivam Garg, Dimitris Tsipras, Percy Liang, and Gregory Valiant\. 2022\.[What can transformers learn in\-context? A case study of simple function classes](https://proceedings.neurips.cc/paper_files/paper/2022/hash/c529dba08a146ea8d6cf715ae8930cbe-Abstract-Conference.html)\.In*Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022*\.
- He et al\. \(2024\)Zhengfu He, Wentao Shu, Xuyang Ge, Lingjie Chen, Junxuan Wang, Yunhua Zhou, Frances Liu, Qipeng Guo, Xuanjing Huang, Zuxuan Wu, Yu\-Gang Jiang, and Xipeng Qiu\. 2024\.[Llama scope: Extracting millions of features from llama\-3\.1\-8b with sparse autoencoders](https://doi.org/10.48550/ARXIV.2410.20526)\.*CoRR*, abs/2410\.20526\.
- He et al\. \(2025\)Zirui He, Haiyan Zhao, Yiran Qiao, Fan Yang, Ali Payani, Jing Ma, and Mengnan Du\. 2025\.[SAIF: A sparse autoencoder framework for interpreting and steering instruction following of language models](https://doi.org/10.48550/ARXIV.2502.11356)\.*CoRR*, abs/2502\.11356\.
- Hendel et al\. \(2023\)Roee Hendel, Mor Geva, and Amir Globerson\. 2023\.[In\-context learning creates task vectors](https://doi.org/10.18653/v1/2023.findings-emnlp.624)\.In*Findings of the Association for Computational Linguistics: EMNLP 2023*, pages 9318–9333\.
- Huben et al\. \(2024\)Robert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart, and Lee Sharkey\. 2024\.[Sparse autoencoders find highly interpretable features in language models](https://openreview.net/forum?id=F76bwRSLeK)\.In*The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7\-11, 2024*\. OpenReview\.net\.
- Jorgensen et al\. \(2023\)Ole Jorgensen, Dylan Cope, Nandi Schoots, and Murray Shanahan\. 2023\.[Improving activation steering in language models with mean\-centring](https://doi.org/10.48550/ARXIV.2312.03813)\.*CoRR*, abs/2312\.03813\.
- Konen et al\. \(2024\)Kai Konen, Sophie F\. Jentzsch, Diaoulé Diallo, Peer Schütt, Oliver Bensch, Roxanne El Baff, Dominik Opitz, and Tobias Hecking\. 2024\.[Style vectors for steering generative large language models](https://doi.org/10.18653/v1/2024.findings-eacl.52)\.In*Findings of the Association for Computational Linguistics: EACL 2024*, pages 782–802\. Association for Computational Linguistics\.
- Lee et al\. \(2025\)Bruce W\. Lee, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Erik Miehling, Pierre L\. Dognin, Manish Nagireddy, and Amit Dhurandhar\. 2025\.[Programming refusal with conditional activation steering](https://openreview.net/forum?id=Oi47wc10sm)\.In*The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24\-28, 2025*\. OpenReview\.net\.
- Li et al\. \(2023\)Kenneth Li, Oam Patel, Fernanda B\. Viégas, Hanspeter Pfister, and Martin Wattenberg\. 2023\.[Inference\-time intervention: Eliciting truthful answers from a language model](https://proceedings.neurips.cc/paper_files/paper/2023/hash/81b8390039b7302c909cb769f8b6cd93-Abstract-Conference.html)\.In*Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10\-16, 2023*\.
- Li et al\. \(2025\)Yichen Li, Zhiting Fan, Ruizhe Chen, Xiaotang Gai, Luqi Gong, Yan Zhang, and Zuozhu Liu\. 2025\.[FairSteer: inference time debiasing for LLMs with dynamic activation steering](https://doi.org/10.18653/v1/2025.findings-acl.589)\.In*Findings of the Association for Computational Linguistics: ACL 2025*, pages 11293–11312\. Association for Computational Linguistics\.
- Lieberum et al\. \(2024\)Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Nicolas Sonnerat, Vikrant Varma, János Kramár, Anca D\. Dragan, Rohin Shah, and Neel Nanda\. 2024\.[Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2](https://doi.org/10.18653/v1/2024.blackboxnlp-1.19)\.In*Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP*, pages 278–300\. Association for Computational Linguistics\.
- Min et al\. \(2022\)Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer\. 2022\.[Rethinking the role of demonstrations: What makes in\-context learning work?](https://doi.org/10.18653/v1/2022.emnlp-main.759)In*Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing*, pages 11048–11064\.
- Postmus and Abreu \(2024\)Joris Postmus and Steven Abreu\. 2024\.[Steering large language models using conceptors: Improving addition\-based activation engineering](https://doi.org/10.48550/ARXIV.2410.16314)\.*CoRR*, abs/2410\.16314\.
- Rimsky et al\. \(2024\)Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner\. 2024\.[Steering llama 2 via contrastive activation addition](https://doi.org/10.18653/V1/2024.ACL-LONG.828)\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 15504–15522\. Association for Computational Linguistics\.
- Scialanga et al\. \(2025\)Marco Scialanga, Thibault Laugel, Vincent Grari, and Marcin Detyniecki\. 2025\.[SAKE: steering activations for knowledge editing](https://doi.org/10.18653/v1/2025.acl-long.777)\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 15966–15978\. Association for Computational Linguistics\.
- Stoehr et al\. \(2024\)Niklas Stoehr, Kevin Du, Vésteinn Snæbjarnarson, Robert West, Ryan Cotterell, and Aaron Schein\. 2024\.[Activation scaling for steering and interpreting language models](https://doi.org/10.18653/V1/2024.FINDINGS-EMNLP.479)\.In*Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12\-16, 2024*, Findings of ACL, pages 8189–8200\. Association for Computational Linguistics\.
- Stolfo et al\. \(2025\)Alessandro Stolfo, Vidhisha Balachandran, Safoora Yousefi, Eric Horvitz, and Besmira Nushi\. 2025\.[Improving instruction\-following in language models through activation steering](https://openreview.net/forum?id=wozhdnRCtw)\.In*The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24\-28, 2025*\. OpenReview\.net\.
- Su et al\. \(2025\)Jingran Su, Jingfan Chen, Hongxin Li, Yuntao Chen, Li Qing, and Zhaoxiang Zhang\. 2025\.[Activation steering decoding: Mitigating hallucination in large vision\-language models through bidirectional hidden state intervention](https://doi.org/10.18653/v1/2025.acl-long.634)\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 12964–12974\. Association for Computational Linguistics\.
- Subramani et al\. \(2022\)Nishant Subramani, Nivedita Suresh, and Matthew E\. Peters\. 2022\.[Extracting latent steering vectors from pretrained language models](https://doi.org/10.18653/v1/2022.findings-acl.48)\.In*Findings of the Association for Computational Linguistics: ACL 2022*, pages 566–581\. Association for Computational Linguistics\.
- Team \(2024a\)Gemma Team\. 2024a\.[Gemma 2: Improving open language models at a practical size](https://doi.org/10.48550/ARXIV.2408.00118)\.*CoRR*, abs/2408\.00118\.
- Team \(2025\)Gemma Team\. 2025\.[Gemma 3 technical report](https://doi.org/10.48550/ARXIV.2503.19786)\.*CoRR*, abs/2503\.19786\.
- Team \(2024b\)Llama Team\. 2024b\.[The llama 3 herd of models](https://doi.org/10.48550/ARXIV.2407.21783)\.*CoRR*, abs/2407\.21783\.
- Templeton et al\. \(2024\)Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L\. Turner, Callum McDougall, Monte MacDiarmid, C\. Daniel Freeman, Theodore R\. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, and 3 others\. 2024\.[Scaling monosemanticity: Extracting interpretable features from Claude 3 sonnet](https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html)\.*Transformer Circuits Thread*\.
- Todd et al\. \(2024\)Eric Todd, Millicent L\. Li, Arnab Sen Sharma, Aaron Mueller, Byron C\. Wallace, and David Bau\. 2024\.[Function vectors in large language models](https://openreview.net/forum?id=AwyxtyMwaG)\.In*International Conference on Learning Representations*\.
- Turner et al\. \(2023\)Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J\. Vazquez, Ulisse Mini, and Monte MacDiarmid\. 2023\.[Steering language models with activation engineering](https://doi.org/10.48550/ARXIV.2308.10248)\.*CoRR*, abs/2308\.10248\.
- Wang et al\. \(2025a\)Weixuan Wang, Jingyuan Yang, and Wei Peng\. 2025a\.[Semantics\-adaptive activation intervention for LLMs via dynamic steering vectors](https://openreview.net/forum?id=8WQ7VTfPTl)\.In*The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24\-28, 2025*\. OpenReview\.net\.
- Wang et al\. \(2025b\)Xinpeng Wang, Mingyang Wang, Yihong Liu, Hinrich Schütze, and Barbara Plank\. 2025b\.[Refusal direction is universal across safety\-aligned languages](https://doi.org/10.48550/ARXIV.2505.17306)\.*CoRR*, abs/2505\.17306\.
- Wu et al\. \(2025\)Zhengxuan Wu, Aryaman Arora, Atticus Geiger, Zheng Wang, Jing Huang, Dan Jurafsky, Christopher D\. Manning, and Christopher Potts\. 2025\.[Axbench: Steering llms? even simple baselines outperform sparse autoencoders](https://proceedings.mlr.press/v267/wu25a.html)\.In*Forty\-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13\-19, 2025*, volume 267 of*Proceedings of Machine Learning Research*, pages 67035–67080\. PMLR / OpenReview\.net\.
- Yang et al\. \(2024\)An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, and 23 others\. 2024\.[Qwen2\.5 technical report](https://doi.org/10.48550/ARXIV.2412.15115)\.*CoRR*, abs/2412\.15115\.
- You et al\. \(2026\)Zejia You, Chunyuan Deng, and Hanjie Chen\. 2026\.[Spherical steering: Geometry\-aware activation rotation for language models](https://doi.org/10.48550/ARXIV.2602.08169)\.*CoRR*, abs/2602\.08169\.
- Zou et al\. \(2023\)Andy Zou, Long Phan, Sarah Li Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann\-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J\. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, and 2 others\. 2023\.[Representation engineering: A top\-down approach to AI transparency](https://doi.org/10.48550/ARXIV.2310.01405)\.*CoRR*, abs/2310\.01405\.
## Appendix ATask and Evaluation Details
#### Task inventory\.
The behavioral steering experiments use 20 targets across four task families: 11 entity targets \(cat, coffee, couch, dog, email, family, fruit, music, phone, tea, water\), 7 persona/style targets \(happy, angry, excited, flattering, rhetorical\-question, sad, emoji\), one refusal target, and one nonsense target\.
#### Dataset statistics\.
The 11 entity targets and 7 persona/style targets are each evaluated on 100 held\-out generic questions, while the reject and nonsense targets are evaluated on 80 and 70 questions, respectively\. This yields 1,950 target–question instances per model, or 5,850 instances across the three models for each evaluated configuration\. Each steering vector is constructed from 16 source prompts per target\. The source prompts and behavioral evaluation questions are disjoint, and the evaluation questions are not used for vector construction\.
Table 8:Behavioral evaluation\-set sizes\.
#### Behavioral boundary\-source construction\.
The refined behavioral boundary sources in Section[5\.2](https://arxiv.org/html/2607.25270#S5.SS2)are built directly from the task\-file fields\. For each target,instruction\_chatcontains a target instruction and a short assistant acknowledgement\. Thefull\_chatpool contains complete target\-bearing question–answer pairs used as in\-context demonstrations\. Theboundary\_chatpool contains question–prefix pairs whose assistant prefix stops immediately before the target behavior would be realized or continued\. For example, an entity prefix may end with “with my” before the entity name, a persona prefix may end with “I feel” before the intended affective continuation, and a factuality prefix may end before the incorrect answer\. For the reject task, the assistant prefix is empty, so the boundary state is read at the beginning of the assistant response\.
Given a boundary pair\(qb,pb\)\(q\_\{b\},p\_\{b\}\), theinstructionsource concatenatesinstruction\_chatwith the open pairUser:qbq\_\{b\},Assistant:pbp\_\{b\}\. Theiclsource concatenates three complete examples sampled fromfull\_chatwith the same open pair\. Thehybridsource concatenatesinstruction\_chat, three complete examples fromfull\_chat, and the open pair\. In all three cases, the main refined\-boundary experiments read the final hidden state of the resulting open context, before any omitted target token or target behavior is generated\.
For contrastive vector\-construction methods, negative sources use the matched negative file for the same family:plain\(neg\)for persona/style,no\-entity\(neg\)for entity,fact\(neg\)for nonsense, andnegfor reject\. These files have the sameinstruction\_chat/full\_chat/boundary\_chatstructure, so negative boundary sources are constructed with the same template while replacing the target\-conditioned instruction, demonstrations, and open prefixes with neutral, factual, helpful, or entity\-free counterparts\.
#### Held\-out evaluation questions\.
Behavioral steering is evaluated on generic open\-ended questions that do not name the target attribute, such as “What is a good way to begin the day?”, “What helps when feeling stuck?”, and “What supports creativity?” This tests whether the intervention changes the model’s general continuation tendency rather than merely completing a target\-specific prompt\. Table[9](https://arxiv.org/html/2607.25270#A1.T9)illustrates outputs that satisfy the success criterion, together with target\-expressing but invalid collapsed outputs that are scored as failures\.
Table 9:Illustrative successful outputs for each behavioral task family\. Outputs must both express the target attribute and remain readable\. Outputs that expose the target only through repetition, token flooding, or malformed text are counted as failures\.
#### Refined boundary prompt examples\.
Tables[10](https://arxiv.org/html/2607.25270#A1.T10)–[12](https://arxiv.org/html/2607.25270#A2.T12)give additional examples of the refined boundary prompt constructions used in Section[5\.2](https://arxiv.org/html/2607.25270#S5.SS2)\. The examples are drawn from the task files and illustrate how the same instruction, in\-context, and hybrid constructions are instantiated for persona, entity, nonsense, and reject targets\.
Table 10:Refined boundary prompt examples for two persona targets\. The table shows one representative demonstration where applicable; the experimentaliclandhybridsources sample three demonstrations from the corresponding task file before the final open\-boundary question\.Table 11:Refined boundary prompt examples for two entity targets\. As above, the experimental few\-shot conditions use three demonstrations sampled from the corresponding task file\.
#### LLM judge prompt\.
The main behavioral experiments use Qwen2\.5\-14B\-Instruct as a fixed binary judge\. For a target attributeATTRIBUTEand a generated texttext, the judge prompt asks whether the text shows a tendency towardATTRIBUTEthrough tone, wording, framing, or affective cues\. It returns 0 if the tendency is absent, or if the apparent target expression mainly comes from repetition, token flooding, severe corruption, gibberish, or malformed text\. The judge is asked to give a brief reason and place the final 0/1 label on the last line; the experiment script extracts this final binary label\.
#### Judge validation\.
We validate the automatic judge on a stratified sample of behavioral steering outputs\. Relation\-completion experiments are excluded from this validation because they are evaluated by normalized string matching rather than by the LLM judge\. The sample is stratified by model, task family, source condition, and automatic judge label, so that both positive and negative automatic decisions are represented across the main experimental settings\.
Each validation item contains the target attribute and the generated output\. Annotators are asked to apply the same success criterion used by the automatic judge: an output is positive only if it expresses the target attribute and remains readable, non\-collapsed, and well formed\. Annotators do not see the source condition, vector\-construction method, layer, strength, or model that produced the output\.
Table[13](https://arxiv.org/html/2607.25270#A2.T13)reports agreement between the Qwen2\.5\-14B automatic judge and the independent validation labels\. We report exact agreement, since the validation is intended as a sanity check for the automatic binary labels used in the main comparisons\. The overall agreement is 87\.0%, with agreement ranging from 82\.0% to 92\.0% across task families\. This suggests that the automatic judge is reliable enough for the controlled comparisons in the main experiments\.
As a complementary judge\-model robustness check, we rescore all behavioral outputs in the six source conditions compared in Figure[2](https://arxiv.org/html/2607.25270#S5.F2)with Gemma\-3\-27B, using the same binary target\-expression and generation\-validity criterion\. Table[14](https://arxiv.org/html/2607.25270#A2.T14)reports the resulting overall success rates under the alternative judge\. Gemma\-3\-27B is more permissive in absolute terms, but preserves the central source\-condition pattern:answer\-only \+ meanis weakest for every generator, while pre\-generation last\-token sources, especiallyinstructionandhybrid, remain strongest or among the strongest\.
## Appendix BRelation Task Details
#### Relation boundary\-source construction\.
Relation tasks use the same three boundary templates as the behavioral tasks, but the open boundary is defined by an input–output pair\(x,y\)\(x,y\)from the task’sdatapool\. Theinstructionsource concatenates the task instruction and acknowledgement frominstruction\_chatwith the open queryUser:xx,Assistant:\. Theiclsource concatenates three complete input–output demonstrations with the same open query\. Thehybridsource concatenates both the instruction pair and three demonstrations before the open query\. In all cases, the boundary readout is the final hidden state before the answeryyis generated\.
The post\-realization relation source uses the same task examples after the answer has appeared, i\.e\., contexts containingUser:xx,Assistant:yy, with sequence\-mean readout\. For contrastive methods, negative relation sources are constructed from therandommapping task with the corresponding template, so the negative examples preserve the input–output format but do not instantiate the evaluated relation\.
Table[15](https://arxiv.org/html/2607.25270#A2.T15)provides the complete inventory of the 16 relation\-completion tasks used in Section[6\.2](https://arxiv.org/html/2607.25270#S6.SS2)\. Each task file provides aninstruction\_chatpair and input–output examples;randomis used only as the negative source pool and is not counted as a relation task\.
#### Dataset statistics\.
The relation\-completion evaluation comprises 16 tasks and 958 input–output pairs in total, with 26–108 pairs per task\. Each source condition uses 16 source prompts, and performance is evaluated on all available pairs using normalized string matching\.
#### Relation scoring\.
Relation tasks are not evaluated with the LLM judge\. The model generates a short continuation and the output is normalized by lowercasing, whitespace normalization, and stripping punctuation\. A prediction is counted as correct if it matches the gold answer under the selected matching rule; the main setting uses prefix matching to allow harmless continuation after the answer token\.
Table 12:Refined boundary prompt examples for the nonsense and reject targets\. For reject, the boundary prefix is empty, so the last\-token readout is taken at the start of the assistant response\.Table 13:Judge validation on a stratified sample of behavioral steering outputs\. Agreement compares Qwen2\.5\-14B\-Instruct judge labels with independent validation labels\.Table 14:Overall success rates after rescoring all behavioral outputs with Gemma\-3\-27B\. Columns abbreviate prompt\-only \(P\-only\), prompt\-and\-answer \(P\+A\), answer\-only \(A\-only\), and instruction \(Instr\.\); the readout is shown beneath each condition\.
Table 15:Relation task inventory\. Instructions are shortened only by removing the repeated phrase “Output the answer only, no extra text\.”Table 17:Overall success rate \(%\) for refined boundary prompts under last\-token and sequence\-mean readout on the three base models\.
## Appendix CExperimental Details
#### SAE\-based steering\.
We use pretrained sparse autoencoders as feature dictionaries rather than as reconstruction modules\. A common SAE intervention amplifies the activation of a chosen feature, decodes the modified SAE code, and substitutes the reconstructed hidden state back into the residual stream\. In contrast, we attach the SAE at a target layer, encode source activations elicited by a batch of source examples, and identify*consensus*features that are active across that batch\. The decoder vector for each consensus feature is then treated as a candidate steering vector and evaluated with the same intervention protocol as the other vector\-construction methods\. We record the best\-performing feature vector for that condition\. This avoids routing the intervention through SAE reconstruction, which could introduce reconstruction error into the edited hidden state\.
We use GemmaScope SAEs for Gemma models\(Lieberum et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib17)\)and LlamaScope SAEs for Llama models\(He et al\.,[2024](https://arxiv.org/html/2607.25270#bib.bib8)\)\. We did not find a suitable open\-source SAE for Qwen2\.5\-7B\-Instruct, so SAE\-based experiments are not reported for that model\.
#### Layer and strength search\.
For full\-vector methods, candidate source and intervention layers are chosen at four roughly evenly spaced depths of each model\. For SAE\-based methods, the source layer is determined by the layer of the corresponding SAE, and the SAE\-derived feature vector is injected at candidate intervention layers\. Steering strengths are swept over method\- and model\-appropriate ranges\.
## Appendix DHeld\-Out Hyperparameter\-Selection Check
We repeat the central pairwise source comparisons over five random validation–test splits\. For each split, model, target, source condition, and vector\-construction method, we select the source layer, intervention layer, and steering strength exclusively on validation questions, then evaluate the selected configuration on disjoint test questions\. Held\-out scores follow the aggregation protocol in Section[4\.2](https://arxiv.org/html/2607.25270#S4.SS2)\.
Table 16:Held\-out source\-condition gains overanswer\-only \+ mean\. Entries are absolute differences in test steering success; brackets give paired hierarchical uncertainty intervals accounting for variation across splits, targets, and questions\.
Both boundary\-oriented source conditions retain large positive gains on held\-out questions for every model, and all paired intervals exclude zero\. Thus, the source effect persists without selecting hyperparameters on the test metric\.
Table 18:Gemma\-2\-9B\-IT source\-grid results across vector\-construction methods\.
## Appendix EReadout Ablations for Refined Boundary Sources
Table[17](https://arxiv.org/html/2607.25270#A2.T17)reports the overall success rates for last\-token and sequence\-mean readout on the refined boundary prompts\. Last\-token readout is strongest for the instruction and hybrid sources across all three models\. Sequence\-mean readout improvesiclfor Qwen2\.5 and Gemma, buticlremains weaker than the instruction and hybrid sources\.
## Appendix FGemma Source\-Grid Results by Vector\-Construction Method
Table[18](https://arxiv.org/html/2607.25270#A4.T18)reports the Gemma\-2\-9B\-IT source\-grid results for all five vector\-construction methods using the aggregation from Section[4\.2](https://arxiv.org/html/2607.25270#S4.SS2): persona/entity are concept averages, nonsense/reject are single\-target scores, and Avg is their unweighted mean\.Similar Articles
Forecasting Side Effects of Activation Steering
This paper investigates whether side effects of activation steering in language models can be predicted before intervention, constructing a cross-effect matrix across 67 behaviors and finding that side effects are systematic and forecastable from unsteered representations.
When is Your LLM Steerable?
This paper introduces a method to predict activation steering effectiveness in language models from early decoding states using a Gradient Boosting Decision Trees (GBDT) classifier, enabling efficient steering strength optimization without full rollouts.
Want Better Synthetic Data? Steer It: Activation Steering for Low-Resource Language Generation
This paper investigates activation steering as an alternative to few-shot prompting for generating synthetic data in low-resource languages. The authors propose LanguageSteering and QualitySteering strategies, showing that steering on early layers improves diversity and downstream model performance.
When is Your LLM Steerable?
This paper investigates when activation steering succeeds or fails for LLMs by analyzing early decoding dynamics. The authors introduce ASTEER, a large testbed of steered generations, and train a GBDT classifier to predict steering outcomes from early hidden states, enabling efficient steering strength search.
A Geometric Account of Activation Steering through Angle-Norm Decomposition
This paper analyzes linear activation steering in language models by decomposing interventions into angular and radial components. It finds that concepts are primarily encoded in angular structure, but norm adjustments are crucial for stability, supporting spherical steering methods while showing that additive coefficients conflate geometry.