MetaSteer: Context-Conditioned, nonlinear Steering via Attention-Projection Adaptation

arXiv cs.CL Papers

Summary

The paper introduces MetaSteer, a method that learns nonlinear, context-dependent interventions applied to attention projection matrices in LLMs, trained once on pooled preference data and transferred zero-shot to unseen concepts and agentic settings, matching or outperforming task-specific steering baselines.

arXiv:2609.38718v1 Announce Type: new Abstract: Steering large language models typically relies on linear, context-independent interventions in activation space, an assumption that recent work has challenged and that can induce an information bottleneck when a fixed representation must encode many behavioral distinctions. We introduce MetaSteer, a method that learns nonlinear interventions with context-dependent effects and applies them to attention projection matrices, producing activation effects that vary with the input context by construction and requiring no linear concept-geometry assumption. Framed as preference-based optimization, MetaSteer is trained once on a pooled preference corpus and transferred zero-shot to unseen concepts and out-of-distribution contexts. We find that, despite using low-rank adapters, MetaSteer induces structured, context-dependent changes in hidden-state trajectories while partially preserving aspects of their local trajectory dynamics, including velocity and curvature. We evaluate MetaSteer on three controlled text-generation benchmarks and three agentic settings across multiple model families and scales. MetaSteer matches or outperforms strong task-specific steering baselines on most aggregate comparisons in the zero-shot regime. Across the evaluated settings, stronger text-generation steering is associated with stronger agentic steering performance. We further discuss geometric trajectory effects, capability retention, and safety considerations raised by transferable steering.
Original Article
View Cached Full Text

Cached at: 10/01/26, 09:44 AM

# MetaSteer: Context-Conditioned, nonlinear Steering via Attention-Projection Adaptation
Source: [https://arxiv.org/html/2609.38718](https://arxiv.org/html/2609.38718)
Mehdi JafariAffiliation:School of Computer Science and Engineering, UNSW Sydney, AustraliaAffiliation:ARC Centre of Excellence for Automated Decision\-Making and Society \(ADM\+S\)Email:[mehdi\.jafari@unsw\.edu\.au](mailto:)Hao XueAffiliation:School of Computer Science and Engineering, UNSW Sydney, AustraliaAffiliation:The Hong Kong University of Science and Technology \(Guangzhou\), ChinaAffiliation:ARC Centre of Excellence for Automated Decision\-Making and Society \(ADM\+S\)Email:[flora\.salim@unsw\.edu\.au](mailto:)Flora SalimAffiliation:School of Computer Science and Engineering, UNSW Sydney, AustraliaAffiliation:ARC Centre of Excellence for Automated Decision\-Making and Society \(ADM\+S\)Email:[haoxue@hkust\-gz\.edu\.cn](mailto:)

###### Abstract

Steering large language models typically relies on linear, context\-independent interventions in activation space, an assumption that recent work has challenged and that can induce an information bottleneck when a fixed representation must encode many behavioral distinctions\. We introduceMetaSteer, a method that learns nonlinear interventions with context\-dependent effects and applies them to attention projection matrices, producing activation effects that vary with the input context by construction and requiring no linear concept\-geometry assumption\. Framed as preference\-based optimization, MetaSteer is trained once on a pooled preference corpus and transferred zero\-shot to unseen concepts and out\-of\-distribution contexts\. We find that, despite using low\-rank adapters, MetaSteer induces structured, context\-dependent changes in hidden\-state trajectories while partially preserving aspects of their local trajectory dynamics, including velocity and curvature\. We evaluate MetaSteer on three controlled text\-generation benchmarks and three agentic settings across multiple model families and scales\. MetaSteer matches or outperforms strong task\-specific steering baselines on most aggregate comparisons in the zero\-shot regime\. Across the evaluated settings, stronger text\-generation steering is associated with stronger agentic steering performance\. We further discuss geometric trajectory effects, capability retention, and safety considerations raised by transferable steering\.

## 1Introduction

Steering in\\tl\_set:Nellm refers to applying interventions to the activation space to control their generations\([Wu et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib2)\)\. Such an intervention can be characterized by three design choices: the geometric assumption underpinning it – traditionally linear\([Xiong et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib10);[Bereska and Gavves, 2024](https://arxiv.org/html/2609.38718#bib.bib9);[Saglam et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib11);[Panickssery et al\., 2023](https://arxiv.org/html/2609.38718#bib.bib34);[Turner et al\., 2023](https://arxiv.org/html/2609.38718#bib.bib35)\), or nonlinear and context\-dependent in light of evidence that concepts lie on curved, anisotropic manifolds\([Engels et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib12);[Modell et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib13);[Nguyen and Le, 2026](https://arxiv.org/html/2609.38718#bib.bib14);[Zhao et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib15);[Mishra et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib20)\)– the location within the LLM’s architecture where it is applied, namely the residual stream\([Nguyen et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib16);[Hsu et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib17)\), MLP components\([Yu et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib18);[Yan et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib19)\), or attention heads\([Genadi et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib21);[Luo et al\., 2026b](https://arxiv.org/html/2609.38718#bib.bib25)\)– and the procedure used to compute it, whether heuristic\([Chalnev et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib30)\), causal\-effect prediction\([Arad et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib32);[Cho et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib31);[Soo et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib33)\), preference\-based cloning\([Raina et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib1)\), or end\-to\-end learned vectors\([Sun et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib29)\)\. The interplay among these choices can be formulated as a*general dual\-objective optimization*problem\([Aravindan et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib26);[Nguyen et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib27);[Luo et al\., 2026a](https://arxiv.org/html/2609.38718#bib.bib28)\), at the heart of which lies the fundamental challenge common to every steering problem: balancing steerability \(𝕊\\mathbb\{S\}\) with preservation of the model’s general utility \(𝕌\\mathbb\{U\}\)\.

Although concept\- or task\-specific steering methods for LLMs exist, a general approach that balances steering effectiveness with preservation of the model’s capabilities has yet to be established\. Such an approach should avoid overly restrictive assumptions about the geometry of concepts and semantic space so that it can generalize to unseen tasks and concepts; intervene at a location that is both effective and computationally tractable; and remain minimally invasive, without degrading the model’s general capabilities\. This raises a concrete question: can we learn a single intervention that adapts its effect to the current context, transfers to unseen concepts and tasks, and preserves the model’s general capabilities?

We answer this question withMetaSteer, a preference\-trained, nonlinear intervention applied to the attention projection matrices\. We choose this site because it offers a computationally tractable way to modify the model’s internal computation\([Luo et al\., 2026b](https://arxiv.org/html/2609.38718#bib.bib25)\)while allowing its effects to propagate through residual connections, feed\-forward sublayers, and subsequent attention blocks\. The intervention therefore need not remain confined to a fixed activation\-space shift\. Because attention is itself a function of the context, MetaSteer’s adapter weights, although fixed after training, induce context\-dependent effects on the residual stream trajectory: the same parameters can produce different activation shifts depending on the context\. MetaSteer is trained once on a pooled preference corpus spanning diverse instructions and steering concepts, then deployed with fixed parameters for zero\-shot transfer to unseen concepts and out\-of\-distribution contexts while largely preserving the model’s general capabilities\.

We measure transfer on three text\-generation benchmarks – AxBench\([Wu et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib2)\), CLaS\-Bench\([Gurgurov et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib40)\), and PersonalityBench\([Deng et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib41)\)– and test whether the same intervention transfers to agentic settings through SocialEval\([Zhou et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib43)\)and the Dictator and Ultimatum Games\([Mozikov et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib42)\)\. MetaSteer matches or exceeds strong task\-specific baselines on aggregate steering scores in most tested settings111All models are instruction\-tuned unless explicitly identified as base models\.\.

The geometry of steered hidden\-state trajectories is also analyzed through their position, velocity, and curvature\. An emerging pattern from this analysis, illustrated by representative examples in Fig\.[1](https://arxiv.org/html/2609.38718#S1.F1)\(a\)–\(c\), is that steering can displace trajectory position while preserving some aspects of how the trajectory evolves, with the degree of linearity varying across steering concepts\. Capability retention is assessed separately, alongside a discussion of the safety implications of transferable steering\.

\(a\)

\(b\)

\(c\)

Figure 1:Hidden\-state trajectories and steering alignment\.\(a\)–\(b\)show PCA projections of sentence\-level hidden states under different steering concepts\. Steering shifts trajectory positions while retaining aspects of velocity and curvature\.222Further conceptual details are provided in Appendix[I](https://arxiv.org/html/2609.38718#A9), with additional information about these two instances given in Appendix[G](https://arxiv.org/html/2609.38718#A7)\.\(c\)reports cosine alignment between MetaSteer and CAA displacements across semantic concept groups; higher values indicate closer agreement with linear steering\.333See Appendix[J](https://arxiv.org/html/2609.38718#A10)for experimental details and Appendix[H](https://arxiv.org/html/2609.38718#A8)for the complete list of concept clusters\.Our contributions are summarized as follows:

- •A nonlinear attention\-based intervention\.MetaSteeris a preference\-trained, low\-rank intervention applied jointly to the query, key, value, and output attention projections, with an accompanying mathematical formalization\.
- •Zero\-shot transfer across concepts and settings\.A single context\-conditioned intervention per backbone, trained on pooled concept–preference data and fixed after training, transfers to held\-out concepts, external benchmarks and agentic settings without further adaptation\.
- •Geometric characterization of steering\.The hidden\-state trajectories induced by steering are analyzed through their position, velocity, and curvature\. The results provide evidence that steering can displace trajectory position while preserving some aspects of trajectory evolution, with the degree of linearity varying across steering concepts\.

## 2Related Work

Steering methods can be organized around three design questions: what geometric structure is assumed for concepts in the model’s semantic space, where the intervention is applied, and how the intervention is computed222A detailed discussion is provided in Appendix[A](https://arxiv.org/html/2609.38718#A1)\.\.

Geometric assumption\.Much prior work assumes concepts are represented approximately linearly in activation space, motivating direction\- or vector\-difference\-based interventions\([Xiong et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib10);[Bereska and Gavves, 2024](https://arxiv.org/html/2609.38718#bib.bib9);[Saglam et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib11);[Panickssery et al\., 2023](https://arxiv.org/html/2609.38718#bib.bib34);[Turner et al\., 2023](https://arxiv.org/html/2609.38718#bib.bib35)\)that compose and transfer across concepts\([Karvonen, 2024](https://arxiv.org/html/2609.38718#bib.bib64);[Nguyen et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib16)\)\. Recent work instead shows concepts may occupy curved, anisotropic manifolds, motivating nonlinear, context\-dependent interventions\([Engels et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib12);[Modell et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib13);[Nguyen and Le, 2026](https://arxiv.org/html/2609.38718#bib.bib14);[Zhao et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib15);[Mishra et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib20)\)\.*MetaSteer*imposes no such restrictive geometric assumption on the structure of semantic space\.

Intervention site\.Interventions may act at the input level via prompts, personas, or reasoning instructions\([Kong et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib56);[Miehling et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib54);[Park et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib47);[Wei et al\., 2022](https://arxiv.org/html/2609.38718#bib.bib63);[Wu et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib2)\), thereby modifying the conditioning context rather than applying an explicit intervention vector; at the parameter level via adapters, prompt or prefix tuning, LoRA, or representation fine\-tuning\([Houlsby et al\., 2019](https://arxiv.org/html/2609.38718#bib.bib60);[Lester et al\., 2021](https://arxiv.org/html/2609.38718#bib.bib55);[Li and Liang, 2021](https://arxiv.org/html/2609.38718#bib.bib53);[Hu et al\., 2022](https://arxiv.org/html/2609.38718#bib.bib51);[Trung et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib52)\); or within the forward pass—most commonly through the residual stream\([Nguyen et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib16);[Hsu et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib17)\)or sparse\-autoencoder features\([Yu et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib18);[Yan et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib19)\), with recent work also targeting MLP components and attention heads\([Genadi et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib21);[Luo et al\., 2026b](https://arxiv.org/html/2609.38718#bib.bib25)\)\. The closest work to ours is[Luo et al\. \(2026b\)](https://arxiv.org/html/2609.38718#bib.bib25), which intervenes only on the query projection; by contrast,MetaSteerintervenes jointly on all four attention projection matrices—the query, key, value, and output projections\.

Learning procedure\.Interventions have been constructed via heuristic or contrastive procedures\([Panickssery et al\., 2023](https://arxiv.org/html/2609.38718#bib.bib34);[Chalnev et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib30)\); objectives balancing steering effectiveness against utility preservation\([Aravindan et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib26);[Nguyen et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib27);[Luo et al\., 2026a](https://arxiv.org/html/2609.38718#bib.bib28)\); learned causal\-effect estimators and steering operators\([Arad et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib32);[Cho et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib31);[Soo et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib33);[Sun et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib29)\); and preference\-based activation\-space behavior cloning\([Raina et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib1)\)\. MetaSteer is closest in learning approach to\([Raina et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib1)\), but avoids the information bottleneck they identify by learning nonlinear, context\-dependent interventions rather than a single low\-dimensional vector\. Unlike\([Sun et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib29)\), which extracts steering directions using a separate, architecturally identical model, MetaSteer uses the same model\.

## 3Methodology

We formulate steering as preference optimization of a shared, low\-rank intervention\. We first describe DPO as a direction\-specific steering baseline, then introduce an intervention on the four attention projections\. We subsequently characterize the geometry of the resulting activation updates and describe the preference data and training objective\.

### 3\.1DPO as Direction\-Specific Steering

DPO increases the likelihood of a preferred completiony\+y^\{\+\}relative to a dispreferred completiony−y^\{\-\}for a promptxx\. For a single next\-token comparison at a shared promptxx, the log\-probability gap between the preferred and dispreferred token is linear in the model’s final hidden state𝐡⁡\(x\)\\mathbf\{h\}\(x\):

log⁡π⁡\(y\+∣x\)−log⁡π⁡\(y−∣x\)=⟨𝐡⁡\(x\),𝐯⟩,𝐯:=𝐞y\+−𝐞y−,\\log\\pi\(y^\{\+\}\\mid x\)\-\\log\\pi\(y^\{\-\}\\mid x\)=\\langle\\mathbf\{h\}\(x\),\\mathbf\{v\}\\rangle,\\qquad\\mathbf\{v\}:=\\mathbf\{e\}\_\{y^\{\+\}\}\-\\mathbf\{e\}\_\{y^\{\-\}\},\(1\)an identity that follows directly from the softmax output parameterization and is derived in full in Appendix[B](https://arxiv.org/html/2609.38718#A2)\([Raina et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib1)\)\. For a fixed candidate\-token pair\(y\+,y−\)\(y^\{\+\},y^\{\-\}\), the gradient of this single\-step gap with respect to the hidden state is parallel to𝐯\\mathbf\{v\}\. This identity alone does not imply a common direction across examples when the candidate\-token pair varies; the approximately fixed𝐝⋆\\mathbf\{d\}^\{\\star\}in Eq\.[2](https://arxiv.org/html/2609.38718#S3.E2)additionally relies on the local approximation of[Raina et al\. \(2025\)](https://arxiv.org/html/2609.38718#bib.bib1)\. Under a local linearization, the resulting activation update can therefore be approximated as

𝐡DPO​\(x\)≈𝐡0​\(x\)\+α​𝐝⋆,\\mathbf\{h\}\_\{\\mathrm\{DPO\}\}\(x\)\\approx\\mathbf\{h\}\_\{0\}\(x\)\+\\alpha\\,\\mathbf\{d\}^\{\\star\},\(2\)where𝐝⋆\\mathbf\{d\}^\{\\star\}is approximately fixed across examples\([Raina et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib1)\)\. Equation[1](https://arxiv.org/html/2609.38718#S3.E1)is stated here for a single next\-token comparison; for full multi\-token completions, it holds exactly at the position where the two completions first diverge, with the remainder of the sequence gap following an ordinary autoregressive expansion rather than a single fixed direction \(Appendix[B](https://arxiv.org/html/2609.38718#A2)\)\. This direction\-specific formulation is efficient, but it cannot represent arbitrary context\-dependent behavior and relies on an approximately linear representation of the target concept\([Engels et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib12);[Modell et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib13)\)\.

### 3\.2Context\-Conditioned Attention Intervention

Rather than adding a fixed vector to the residual stream, we apply low\-rank adapters to all four attention projection matrices,f∈\{q,k,v,o\}f\\in\\\{q,k,v,o\\\}, at each layer:

𝐖~f=𝐖f\+Δ​𝐖f=𝐖f\+λf​𝐀f​𝐁f⊤,r≪d,\\widetilde\{\\mathbf\{W\}\}\_\{f\}=\\mathbf\{W\}\_\{f\}\+\\Delta\\mathbf\{W\}\_\{f\}=\\mathbf\{W\}\_\{f\}\+\\lambda\_\{f\}\\mathbf\{A\}\_\{f\}\\mathbf\{B\}\_\{f\}^\{\\top\},\\qquad r\\ll d,\(3\)where𝐀f\\mathbf\{A\}\_\{f\}and𝐁f\\mathbf\{B\}\_\{f\}are trainable low\-rank factors andλf\\lambda\_\{f\}controls the intervention intensity\.

The direct update remains rank\-constrained: for each projection,λf​𝐳i​𝐀f​𝐁f⊤\\lambda\_\{f\}\\mathbf\{z\}\_\{i\}\\mathbf\{A\}\_\{f\}\\mathbf\{B\}\_\{f\}^\{\\top\}lies in an at\-most\-rr\-dimensional subspace\. Context changes its coefficients within this subspace, while attention and composition across layers yield a richer end\-to\-end effect\.

Let𝐳i\\mathbf\{z\}\_\{i\}denote the normalized activation at token positionii\. The baseline and adapted query, key, and value representations are

𝐪i=𝐳i​𝐖q,𝐤j=𝐳j​𝐖k,𝐯j=𝐳j​𝐖v,\\mathbf\{q\}\_\{i\}=\\mathbf\{z\}\_\{i\}\\mathbf\{W\}\_\{q\},\\qquad\\mathbf\{k\}\_\{j\}=\\mathbf\{z\}\_\{j\}\\mathbf\{W\}\_\{k\},\\qquad\\mathbf\{v\}\_\{j\}=\\mathbf\{z\}\_\{j\}\\mathbf\{W\}\_\{v\},𝐪~i=𝐳i​𝐖~q,𝐤~j=𝐳j​𝐖~k,𝐯~j=𝐳j​𝐖~v,\\widetilde\{\\mathbf\{q\}\}\_\{i\}=\\mathbf\{z\}\_\{i\}\\widetilde\{\\mathbf\{W\}\}\_\{q\},\\qquad\\widetilde\{\\mathbf\{k\}\}\_\{j\}=\\mathbf\{z\}\_\{j\}\\widetilde\{\\mathbf\{W\}\}\_\{k\},\\qquad\\widetilde\{\\mathbf\{v\}\}\_\{j\}=\\mathbf\{z\}\_\{j\}\\widetilde\{\\mathbf\{W\}\}\_\{v\},withΔ​𝐪i:=𝐪~i−𝐪i\\Delta\\mathbf\{q\}\_\{i\}:=\\widetilde\{\\mathbf\{q\}\}\_\{i\}\-\\mathbf\{q\}\_\{i\}andΔ​𝐤j:=𝐤~j−𝐤j\\Delta\\mathbf\{k\}\_\{j\}:=\\widetilde\{\\mathbf\{k\}\}\_\{j\}\-\\mathbf\{k\}\_\{j\}following from Eq\.[3](https://arxiv.org/html/2609.38718#S3.E3)\. For a token pair\(i,j\)\(i,j\), the \(unperturbed\) scaled dot\-product attention logit issi​j=⟨𝐪i,𝐤j⟩/ds\_\{ij\}=\\langle\\mathbf\{q\}\_\{i\},\\mathbf\{k\}\_\{j\}\\rangle/\\sqrt\{d\}, and the adapted logit iss~i​j=⟨𝐪~i,𝐤~j⟩/d\\widetilde\{s\}\_\{ij\}=\\langle\\widetilde\{\\mathbf\{q\}\}\_\{i\},\\widetilde\{\\mathbf\{k\}\}\_\{j\}\\rangle/\\sqrt\{d\}, which we write as

s~i​j=si​j\+δi​j​\(𝐳i,𝐳j\),\\widetilde\{s\}\_\{ij\}=s\_\{ij\}\+\\delta\_\{ij\}\(\\mathbf\{z\}\_\{i\},\\mathbf\{z\}\_\{j\}\),\(4\)where, expanding⟨𝐪~i,𝐤~j⟩=⟨𝐪i\+Δ​𝐪i,𝐤j\+Δ​𝐤j⟩\\langle\\widetilde\{\\mathbf\{q\}\}\_\{i\},\\widetilde\{\\mathbf\{k\}\}\_\{j\}\\rangle=\\langle\\mathbf\{q\}\_\{i\}\+\\Delta\\mathbf\{q\}\_\{i\},\\mathbf\{k\}\_\{j\}\+\\Delta\\mathbf\{k\}\_\{j\}\\rangleand cancelling the shared⟨𝐪i,𝐤j⟩/d=si​j\\langle\\mathbf\{q\}\_\{i\},\\mathbf\{k\}\_\{j\}\\rangle/\\sqrt\{d\}=s\_\{ij\}term, the perturbation is exactly the sum of the three cross terms induced by the low\-rank update:

δi​j​\(𝐳i,𝐳j\)=1d​\[⟨𝐪i,Δ​𝐤j⟩\+⟨Δ​𝐪i,𝐤j⟩\+⟨Δ​𝐪i,Δ​𝐤j⟩\]\.\\delta\_\{ij\}\(\\mathbf\{z\}\_\{i\},\\mathbf\{z\}\_\{j\}\)=\\frac\{1\}\{\\sqrt\{d\}\}\\Big\[\\langle\\mathbf\{q\}\_\{i\},\\Delta\\mathbf\{k\}\_\{j\}\\rangle\+\\langle\\Delta\\mathbf\{q\}\_\{i\},\\mathbf\{k\}\_\{j\}\\rangle\+\\langle\\Delta\\mathbf\{q\}\_\{i\},\\Delta\\mathbf\{k\}\_\{j\}\\rangle\\Big\]\.\(5\)δi​j\\delta\_\{ij\}depends on both token representations because the adapters modify the query and key projections\. For generic adapters,

∂δi​j∂𝐳i≠0,∂δi​j∂𝐳j≠0\.\\frac\{\\partial\\delta\_\{ij\}\}\{\\partial\\mathbf\{z\}\_\{i\}\}\\neq 0,\\qquad\\frac\{\\partial\\delta\_\{ij\}\}\{\\partial\\mathbf\{z\}\_\{j\}\}\\neq 0\.\(6\)Thus, although the parameter update is fixed after training, its induced attention change depends on the surrounding input context\. The full perturbation is derived in Appendix[C](https://arxiv.org/html/2609.38718#A3)\.

For notational simplicity, Equations[4](https://arxiv.org/html/2609.38718#S3.E4)and[5](https://arxiv.org/html/2609.38718#S3.E5)are stated for a single attention head with no positional rotation; the general form for multi\-head, grouped\-query attention under RoPE is derived in Appendix[A\.5](https://arxiv.org/html/2609.38718#A1.SS5)and reduces to those equations exactly whenH=Hk​v=1H=H\_\{kv\}=1and the rotation is the identity\. In the general case, withHHquery heads,Hk​v≤HH\_\{kv\}\\leq Hkey/value heads \(query headhhsharing key/value headκ⁡\(h\)\\kappa\(h\)under grouped\-query attention\), and per\-head outputs𝐚~i\(h\)\\widetilde\{\\mathbf\{a\}\}\_\{i\}^\{\(h\)\}, the output projection is applied*after*concatenating head outputs rather than to each value vector independently:

𝐚~i=Concath=1H⁡\(𝐚~i\(h\)\)​𝐖~o=∑h=1H∑jsoftmaxj⁡\(s~i​j\(h\)\)​\(𝐯~j\(κ⁡\(h\)\)​𝐖~o\(h\)\),\\widetilde\{\\mathbf\{a\}\}\_\{i\}=\\operatorname\{Concat\}\_\{h=1\}^\{H\}\\\!\\big\(\\widetilde\{\\mathbf\{a\}\}\_\{i\}^\{\(h\)\}\\big\)\\,\\widetilde\{\\mathbf\{W\}\}\_\{o\}=\\sum\_\{h=1\}^\{H\}\\sum\_\{j\}\\operatorname\{softmax\}\_\{j\}\\\!\\big\(\\widetilde\{s\}\_\{ij\}^\{\(h\)\}\\big\)\\left\(\\widetilde\{\\mathbf\{v\}\}\_\{j\}^\{\(\\kappa\(h\)\)\}\\widetilde\{\\mathbf\{W\}\}\_\{o\}^\{\(h\)\}\\right\),\(7\)where𝐖~o\(h\)\\widetilde\{\\mathbf\{W\}\}\_\{o\}^\{\(h\)\}is the row\-block of𝐖~o\\widetilde\{\\mathbf\{W\}\}\_\{o\}corresponding to headhh, ands~i​j\(h\)\\widetilde\{s\}\_\{ij\}^\{\(h\)\}is that head’s own adapted attention logit \(Appendix[A\.5](https://arxiv.org/html/2609.38718#A1.SS5)\)\. Because the softmax weights and projected values depend on the current token representations, the induced updateΔ​𝐚i=𝐚~i−𝐚i\\Delta\\mathbf\{a\}\_\{i\}=\\widetilde\{\\mathbf\{a\}\}\_\{i\}\-\\mathbf\{a\}\_\{i\}is generally nonlinear:

∂2Δ​𝐚i∂𝐳i2≢0\.\\frac\{\\partial^\{2\}\\Delta\\mathbf\{a\}\_\{i\}\}\{\\partial\\mathbf\{z\}\_\{i\}^\{2\}\}\\not\\equiv 0\.\(8\)This provides a context\-conditioned alternative to adding a fixed activation direction\.

### 3\.3Geometry of Steering Effects

The preceding parameterization produces an activation displacement that can vary with both the instruction and the steering concept\. We write this displacement as

𝐡Θ​\(x,c\)≈𝐡0​\(x\)\+α⁡\(x,c\)​𝐝​\(x,c\),𝐝⁡\(x,c\)∈ℝd,\\mathbf\{h\}\_\{\\Theta\}\(x,c\)\\approx\\mathbf\{h\}\_\{0\}\(x\)\+\\alpha\(x,c\)\\mathbf\{d\}\(x,c\),\\qquad\\mathbf\{d\}\(x,c\)\\in\\mathbb\{R\}^\{d\},\(9\)whereΘ\\Thetadenotes the adapter parameters\.

We test whether this context\-dependent displacement nonetheless collapses, concept by concept, onto a single fixed linear direction, or instead departs from one\. For each steering conceptcc, we compare its mean MetaSteer displacement𝐝cMS:=𝔼x​\[𝐝⁡\(x,c\)\]\\mathbf\{d\}^\{\\mathrm\{MS\}\}\_\{c\}:=\\mathbb\{E\}\_\{x\}\[\\mathbf\{d\}\(x,c\)\]against CAA direction𝐯cCAA\\mathbf\{v\}^\{\\mathrm\{CAA\}\}\_\{c\}estimated for the same concept, via their cosine similarity

sc=cos⁡\(𝐯cCAA,𝐝cMS\)\.s\_\{c\}=\\cos\\\!\\left\(\\mathbf\{v\}^\{\\mathrm\{CAA\}\}\_\{c\},\\ \\mathbf\{d\}^\{\\mathrm\{MS\}\}\_\{c\}\\right\)\.\(10\)A value ofscs\_\{c\}near one indicates alignment with the CAA direction, while a low value indicates disagreement with this particular linear baseline\. This is an alignment diagnostic; it neither rules out another fixed linear direction nor estimates the intrinsic dimensionality of the hidden\-state trajectory\. We reportscs\_\{c\}for a broad concept inventory across multiple models in Section[6](https://arxiv.org/html/2609.38718#S6)and Appendix[J](https://arxiv.org/html/2609.38718#A10)\(Figure[1](https://arxiv.org/html/2609.38718#S1.F1)\(c\)\)\.

### 3\.4Preference Data

We construct a pooled preference dataset from Concept16K and Concept16K\-v2\([Wu et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib2)\)\. For each steering conceptccand instructionxx, the dataset provides preferred and dispreferred responses\. We form DPO tuples

\(x,c,y\+,y−\),y\+∈𝒟c\+,y−∈𝒟c−,\(x,c,y^\{\+\},y^\{\-\}\),\\qquad y^\{\+\}\\in\\mathcal\{D\}\_\{c\}^\{\+\},\\quad y^\{\-\}\\in\\mathcal\{D\}\_\{c\}^\{\-\},\(11\)and pool them across concepts:

𝒟=⋃c∈𝒞\{\(x,c,y\+,y−\)\}\.\\mathcal\{D\}=\\bigcup\_\{c\\in\\mathcal\{C\}\}\\\{\(x,c,y^\{\+\},y^\{\-\}\)\\\}\.\(12\)The internal split is concept\-disjoint but shares instructions across splits; transfer to external instruction distributions is evaluated only on the external benchmarks\. Dataset construction and diversity statistics are provided in Appendix[M](https://arxiv.org/html/2609.38718#A13)\.

### 3\.5Preference Optimization Objective

LetπΘ\\pi\_\{\\Theta\}denote the base model with the attention adapters and letπref\\pi\_\{\\mathrm\{ref\}\}be the frozen reference model\. We define

ρΘ​\(y∣x,c\)=log⁡πΘ​\(y∣x,c\)−log⁡πref​\(y∣x,c\)\.\\rho\_\{\\Theta\}\(y\\mid x,c\)=\\log\\pi\_\{\\Theta\}\(y\\mid x,c\)\-\\log\\pi\_\{\\mathrm\{ref\}\}\(y\\mid x,c\)\.\(13\)
We optimize the adapter factorsΘ=\{𝐀f,𝐁f\}\\Theta=\\\{\\mathbf\{A\}\_\{f\},\\mathbf\{B\}\_\{f\}\\\}, conditioned on the instruction and steering concept, using the standard DPO objective with fixedλf=1\\lambda\_\{f\}=1\(Appendix[M\.4](https://arxiv.org/html/2609.38718#A13.SS4)\):

ℒDPO​\(Θ\)=−𝔼\(x,c,y\+,y−\)∼𝒟​\[log⁡σ⁡\(β​ρΘ​\(y\+∣x,c\)−β​ρΘ​\(y−∣x,c\)\)\]\.\\mathcal\{L\}\_\{\\mathrm\{DPO\}\}\(\\Theta\)=\-\\mathbb\{E\}\_\{\(x,c,y^\{\+\},y^\{\-\}\)\\sim\\mathcal\{D\}\}\\left\[\\log\\sigma\\left\(\\beta\\rho\_\{\\Theta\}\(y^\{\+\}\\mid x,c\)\-\\beta\\rho\_\{\\Theta\}\(y^\{\-\}\\mid x,c\)\\right\)\\right\]\.\(14\)
The base model parameters remain frozen; only the low\-rank adapter factors\{𝐀f,𝐁f\}\\\{\\mathbf\{A\}\_\{f\},\\mathbf\{B\}\_\{f\}\\\}are updated\. The intensity scalarsλf\\lambda\_\{f\}are fixed hyperparameters \(Appendix[M\.4](https://arxiv.org/html/2609.38718#A13.SS4)\)\.

Table 1:Zero\-shot transfer results on AxBench, CLaS\-Bench, and PersonalityBench across the evaluated LLaMA, Gemma, and Qwen model scales\. Aggregate\-score columns summarize the metrics described in Section[4](https://arxiv.org/html/2609.38718#S4)\. Bold entries indicate the larger aggregate score for each model\.

## 4Zero\-Shot Transfer to Text\-Generation Benchmarks

We evaluate whetherMetaSteer, trained once per backbone on the pooled preference corpus𝒟\\mathcal\{D\}, transfers to held\-out concepts and external benchmark instruction distributions\. No per\-benchmark fine\-tuning or task\-specific hyperparameter search is performed\.

##### Benchmarks\.

AxBench\.\([Wu et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib2)\)evaluates instruction following under concept\-steering constraints across information\-seeking, mathematical, and programming queries \(∼\\sim32K test examples\)\. Its metrics are concept score \(C\), instruction score \(I\), and fluency score \(F\), combined using their harmonic mean \(HM\)\.CLaS\-Bench\.\([Gurgurov et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib40)\)evaluates language steering in question\-answering settings \(∼\\sim72K test examples\)\. Its metrics are language forcing \(L\) and output relevance \(R\), combined using their harmonic mean \(HM\)\.PersonalityBench\.PersonalityBench\([Deng et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib41)\)evaluates elicitation of the Big Five personality traits through open\-ended long responses \(∼\\sim500 test examples\)\. Its metrics are personality score \(P\) and fluency score \(F\), whose arithmetic mean \(AM\) is reported as the aggregate score\.

##### Reference results and matched baselines\.

We use two complementary comparisons for the text\-generation tasks\. TheBest Refrow denotes the best\-performing method reported for each benchmark:*Prompt*for AxBench,ε\\varepsilon\-base\-IIfor CLaS\-Bench, and*P2*for PersonalityBench\. We implemented each method and evaluated it using the corresponding benchmark’s original framework and scoring procedure\.MetaSteerwas then evaluated on the same tasks under the same framework, enabling a direct comparison with the strongest task\-specific reference available for each benchmark\.

We additionally evaluate SKOP, the closest methodological comparator toMetaSteer, in a matched head\-to\-head comparison following SKOP’s reported evaluation procedure\. On Llama\-3\.1\-8B,MetaSteeroutperforms SKOP by 2% in utility \(𝕌\\mathbb\{U\}\) and 0\.6 points in steering effectiveness \(𝕊\\mathbb\{S\}\)\. HyperSteer is discussed as a related transferable\-steering method, but its reported AxBench performance is below the task\-specific reference implemented here\. D\-Steer is also conceptually related; however, its fixed representation introduces an information bottleneck that limits its ability to represent fine\-grained concept variation\. Results for three model families at two parameter scales are reported in Table[1](https://arxiv.org/html/2609.38718#S3.T1)\.

##### Results\.

MetaSteerimproves aggregate steering scores over the reported baselines on AxBench and CLaS\-Bench across settings\. On PersonalityBench, it remains competitive, with only small differences in aggregate scores\. Together with the zero\-shot evaluation and the geometric analysis in Section[6](https://arxiv.org/html/2609.38718#S6), these results support structured and transferable steering effects rather than collapse to a single direction or concept\-specific memorization\.

Table 2:Emotion\-steering results for the Dictator and Ultimatum Games across three model families at two parameter scales, with GPT\-4o and Human reference rows \(scoring defined in Section[5](https://arxiv.org/html/2609.38718#S5)\)\.Boldindicates the best Simulation Score per model\.AgentMethodAngerDisgustFearHappinessSadnessSimulationScoreDUPURDUPURDUPURDUPURDUPURHuman↑\\uparrow↑\\uparrow↓\\downarrow↓\\downarrow↓\\downarrow↑\\uparrow↑\\uparrow↑\\uparrow↑\\uparrow↓\\downarrow↓\\downarrow↓\\downarrow↑\\uparrow↑\\uparrow↓\\downarrow–GPT\-4o✗✗✓✗△\\triangle✗✓✓✓✗✗✗✓✓✓15/30Llama\-3\.2\-3BEAI\-CP✗✗✓✗✓✗✗✗✗✓✓✓✗✗△\\triangle11/30MetaSteer✗✗△\\triangle✓✓✗✗✗✗✓✓✓✓✗✓15/30Llama\-3\.1\-8BEAI\-CP✗✗✓✗✗✗✗✓✗✓✗✓✓✓✓14/30MetaSteer✓✗✓✗✗✗✓✓✗✓✓✓✓✓✓20/30Gemma\-2\-2BEAI\-CP✗✓✗✗✗✗✓✓✗✓✗✓✓✓✓16/30MetaSteer✓✓✓✗✗✗✓✓✗✓✓✓✓✓✓22/30Gemma\-2\-9BEAI\-CP✗✓✗✓✓✗✗✓✓✓✗✗✓✓✗16/30MetaSteer✗✗✓✗✓✗✓✗✗✓✓△\\triangle✓✓✓17/30Qwen3\-1\.7BEAI\-CP✗✗✓✓✓△\\triangle✓✗△\\triangle✓✓✓✗✗✓18/30MetaSteer✗✓✗✓✗✗✗△\\triangle△\\triangle✓✓✓✓△\\triangle△\\triangle16/30Qwen3\-4BEAI\-CP✗✓△\\triangle✓✗✗✗✓△\\triangle✓✗△\\triangle✓✓△\\triangle16/30MetaSteer✗✗✓✓✓✗✗✗△\\triangle✓✓△\\triangle✓✓△\\triangle17/30

## 5Beyond Text Generation: Agentic Transfer

Transfer on text\-generation benchmarks does not necessarily imply transfer to interactive decision\-making\. We therefore evaluate whether the same fixed steering intervention can influence agentic behavior across three settings – the Dictator Game, the Ultimatum Game, and SocialEval – while preserving the capability to track task dynamics and produce valid, in\-context actions\.

##### Dictator and Ultimatum Games\.

In the Dictator Game, an allocator divides a fixed amount of money between itself and a passive recipient\. In the Ultimatum Game, a proposer divides a fixed amount between itself and a responder, who accepts or rejects the offer; we evaluate both roles\. In both games, the agent is steered toward one of five emotions – anger, disgust, fear, happiness, or sadness – and we compare the resulting change in behavior \(the allocator’s giving, the proposer’s offer size, and the responder’s acceptance rate\) against the human ground\-truth direction \(↑,↓\\uparrow,\\downarrow\) reported for each emotion by[Mozikov et al\. \(2024\)](https://arxiv.org/html/2609.38718#bib.bib42)\.

##### SocialEval\.

SocialEval\([Zhou et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib43)\)presents the agent with branching social scenarios: at each decision point, the agent selects an action that determines the subsequent trajectory, and each terminal node is annotated with the social trait it expresses\. The agent is steered toward one of three traits – proself, antisocial, or prosocial – and we report the resulting Goal Achievement Score \(GAE\) at the macro level \(averaged across trait categories\) and the micro level \(averaged across all trajectories\)\.

##### Reference points and metrics\.

For SocialEval,SocialEval\-En\(Seval\-Enin Table[3](https://arxiv.org/html/2609.38718#S5.T3)\) is the reference configuration, evaluated on the same English scenarios, trait categories, and macro/micro GAE metrics asMetaSteer\. For the Dictator and Ultimatum Games, theEAI\-CP\(Co\-player\) condition of[Mozikov et al\. \(2024\)](https://arxiv.org/html/2609.38718#bib.bib42)is the reference, evaluated under the same emotion–role conditions\. In Table[2](https://arxiv.org/html/2609.38718#S4.T2),Ddenotes the Dictator allocator,UPthe Ultimatum proposer, andURthe Ultimatum responder; ✓, ✗, and△\\trianglemark agreement, contradiction, and an ambiguous shift relative to the human ground\-truth direction\. TheSimulation Scoreawards two points per ✓, one per△\\triangle, and zero per ✗ across the 15 emotion–role conditions \(maximum 30\)\. All game runs use temperature0\.70\.7with 200 runs per condition\.

##### Results\.

MetaSteeroutperforms the agentic reference on five of the six settings: it improves the Simulation Score relative toEAI\-CP\(Table[2](https://arxiv.org/html/2609.38718#S4.T2)\) and the macro\- and micro\-averaged GAE relative toSeval\-En\(Table[3](https://arxiv.org/html/2609.38718#S5.T3)\)\. These results suggest that gains in steered text\-generation performance are associated with more controllable agents more broadly\.

Table 3:SocialEval GAE results across proself, antisocial, and prosocial targets, three model families, and two parameter scales each \(N=1210N=1210scenarios\)\. Human baseline from 20 native Chinese/English graduate\-student annotators \(14 scenarios per language\)\. Best score per metric per model inbold\.

## 6Geometric Analysis

Prior work suggests that hidden\-state trajectories encode information through both absolute position and local geometric dynamics, including velocity and curvature\([Xu et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib46);[Park et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib47);[Gjølbye et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib44);[Zhou et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib45)\)\. We therefore test whether concept steering changes the semantic location of a response while preserving aspects of its instruction\-associated trajectory dynamics\.

We analyze 29,484 segments from 2,550 steered responses spanning 381 instructions, 70 concepts, and three genres \(code, text, math\)\. Segments are encoded either cumulatively with their preceding context or independently\. Data construction, filtering, pooling, and similarity metrics are detailed in Appendix[I](https://arxiv.org/html/2609.38718#A9)\.

##### Higher\-order geometry is associated with instruction structure\.

Position similarity remains high across instruction, concept, and genre groupings and changes little after shuffling, making it weakly diagnostic of ordered structure\. In contrast, instruction\-grouped velocity increases from\.06/\.06\.06/\.06to\.31/\.38\.31/\.38under cumulative encoding and from\.00/\.00\.00/\.00to\.29/\.37\.29/\.37under isolated encoding\. Acceleration similarly increases from\.02/\.02\.02/\.02to\.27/\.35\.27/\.35and from\.00/\.00\.00/\.00to\.27/\.35\.27/\.35, respectively\. Isolated Menger\-curvature similarity increases from\.02/\.00\.02/\.00to\.29/\.37\.29/\.37, while the cumulative increase is weaker \(\.32/\.35\.32/\.35to\.36/\.44\.36/\.44\) because accumulated context retains shared structure after shuffling\. Concept\- and genre\-grouped similarities remain lower, indicating that velocity, acceleration, and curvature partially reflect shared instruction structure\. Complete results appear in Appendix[I](https://arxiv.org/html/2609.38718#A9), Tables[7](https://arxiv.org/html/2609.38718#A9.T7)and[8](https://arxiv.org/html/2609.38718#A9.T8), and Figure[1](https://arxiv.org/html/2609.38718#S1.F1)\(a\)–\(b\)\.

##### The geometric effect depends on the concept\.

In a complementary experiment, we compare the positional displacements induced byMetaSteerwith those produced by CAA for 150 concepts proposed by[Fan et al\. \(2026\)](https://arxiv.org/html/2609.38718#bib.bib65)\. Agreement is quantified with the cosine alignment scorescs\_\{c\}\(Eq\.[10](https://arxiv.org/html/2609.38718#S3.E10)\), averaged across six models; complete experimental details are provided in Appendix[J](https://arxiv.org/html/2609.38718#A10)\. Figure[1](https://arxiv.org/html/2609.38718#S1.F1)\(c\)reports the mean CAA\-alignment score for each concept, with shadow denoting one standard deviation across models, and organizes the concepts into semantic clusters\. Concepts that induce broad*global*changes, such as shifts in language or output format, are more accurately represented by a linear steering direction; concepts requiring context\-dependent combinations of local and global changes, such as discourse development or argument framing, are less adequately captured by a single CAA direction\.

\(a\)

\(b\)

\(c\)

Figure 2:Trajectory geometry and safety\.\(a\)and\(b\)show instruction\-grouped gains in trajectory similarity relative to a randomly shuffled control for position, velocity, acceleration, and Menger curvature under isolated\-segment\(a\)and cumulative\-context\(b\)encoding, respectively\.\(c\)shows obedience rates for harmful steering concepts from JailbreakBench; higher values indicate weaker refusal behaviour\.

## 7Supporting Analyses

We conduct supporting analyses of safety, evaluator reliability, and capability retention\. Because activation steering can weaken safety\-relevant behavior such as refusal\([Marks and Tegmark, 2023](https://arxiv.org/html/2609.38718#bib.bib71);[Geiger et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib69);[Bao et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib68);[Wu et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib70)\), we assess whether MetaSteer can bypass refusals using harmful behaviours from JailbreakBench\([Chao et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib23)\)\. Each behaviour is decomposed into a benign instruction and a risky steering concept, and generated responses are classified according to whether they comply with or refuse the harmful request\. The results indicate that transferable steering can weaken refusal behaviour in some contexts \(Figure[2](https://arxiv.org/html/2609.38718#S6.F2)\(c\)\), motivating caution before deployment\. Further details are provided in Appendix[L](https://arxiv.org/html/2609.38718#A12)\.

We also assess the reliability of LLM\-based evaluation, which offers an efficient alternative to human assessment while raising concerns about consistency, calibration, and bias\([Zheng et al\., 2023](https://arxiv.org/html/2609.38718#bib.bib4);[Zhang et al\., 2023](https://arxiv.org/html/2609.38718#bib.bib50);[Gu et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib57);[Yu, 2025](https://arxiv.org/html/2609.38718#bib.bib67);[Gu et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib66)\)\. On selected samples from the text\-generation evaluations, agreement among the LLM judges was strong\. Detailed per\-method, pairwise, and calibration analyses are reported in Appendix[D](https://arxiv.org/html/2609.38718#A4)\. Human\-anchored validation of the judge scores is provided in Appendix[E\.2](https://arxiv.org/html/2609.38718#A5.SS2)\. Finally, we evaluate capability retention by comparing the adapted models with their unmodified bases on MMLU\([Hendrycks et al\., 2020](https://arxiv.org/html/2609.38718#bib.bib62)\)and TruthfulQA\([Lin et al\., 2022](https://arxiv.org/html/2609.38718#bib.bib61)\)\. The results show generally limited changes in knowledge and truthfulness, with no consistent degradation across models\. Complete per\-model results are reported in Appendix[F](https://arxiv.org/html/2609.38718#A6)\.

## 8Conclusion

We introducedMetaSteer, a preference\-trained low\-rank intervention applied to all four attention projection matrices\. Unlike a fixed activation vector, its parameters remain fixed after training while its induced activation effect varies with the input context\. The method therefore addresses three linked design questions: where to intervene, how to learn the intervention, and how to avoid imposing a globally linear representation of each concept\.

A single intervention trained on pooled preference data transfers to unseen concepts and instruction combinations across three text\-generation benchmarks and three agentic settings without per\-benchmark adaptation\. Geometric analyses further show that steering often changes trajectory position while partially preserving aspects of its local dynamics, including velocity and acceleration, although the strength of this effect varies across concepts\.

Overall, our results support viewing steering as controlled trajectory displacement: the intervention must induce a target behavior while preserving task\-relevant dynamics and general model utility\. We additionally assess evaluator reliability and identify safety considerations that should be addressed before deploying transferable steering methods\.

### Reproducibility Statement

The source code and pretrained model checkpoints will be made publicly available in an upcoming arXiv version shortly\.

#### Acknowledgments

This project is supported by the ARC Centre of Excellence for Automated Decision\-Making and Society \(CE200100005\)\. Computational facilities were provided by the School of Computer Science and Engineering at UNSW Sydney through Katana\.

## References

- D\. Arad, A\. Mueller, and Y\. BelinkovSaes are good for steering–if you select the right features\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 10252–10270\.Cited by:[§A\.1](https://arxiv.org/html/2609.38718#A1.SS1.SSS0.Px1.p1.1),[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px3.p1.1),[§A\.3](https://arxiv.org/html/2609.38718#A1.SS3.SSS0.Px2.p1.1),[§A\.4](https://arxiv.org/html/2609.38718#A1.SS4.p1.1),[Appendix E](https://arxiv.org/html/2609.38718#A5.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p4.1)\.
- Aravindanet al\.\(2026\)K\. Aravindan, A\. Rastogi, A\. Prasad, K\. Aneja, S\. Jain, V\. Shivkumar, and P\. KumaraguruOPIUM: mitigating steering externalities and over\-refusal via dual objective latent optimization\.arXiv preprint arXiv:2607\.19806\.Cited by:[§A\.3](https://arxiv.org/html/2609.38718#A1.SS3.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p4.1)\.
- Baoet al\.\(2026\)Y\. Bao, X\. Zhang, J\. Chen, G\. Su, Y\. Cai, H\. Peng, S\. Bing, H\. Weng, L\. Yan, and J\. YinFaithful bi\-directional model steering via distribution matching and distributed interchange interventions\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 68722–68776\.Cited by:[Appendix F](https://arxiv.org/html/2609.38718#A6.p1.1),[§7](https://arxiv.org/html/2609.38718#S7.p1.1)\.
- Bereska and Gavves \(2024\)L\. Bereska and E\. GavvesMechanistic interpretability for ai safety–a review\.arXiv preprint arXiv:2404\.14082\.Cited by:[§A\.1](https://arxiv.org/html/2609.38718#A1.SS1.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p2.1)\.
- Chalnevet al\.\(2024\)S\. Chalnev, M\. Siu, and A\. ConmyImproving steering vectors by targeting sparse autoencoder features\.arXiv preprint arXiv:2411\.02193\.Cited by:[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p4.1)\.
- Chaoet al\.\(2024\)P\. Chao, E\. Debenedetti, A\. Robey, M\. Andriushchenko, F\. Croce, V\. Sehwag, E\. Dobriban, N\. Flammarion, G\. J\. Pappas, F\. Tramer,et al\.Jailbreakbench: an open robustness benchmark for jailbreaking large language models\.Advances in Neural Information Processing Systems37,pp\. 55005–55029\.Cited by:[Appendix L](https://arxiv.org/html/2609.38718#A12.p1.1),[§7](https://arxiv.org/html/2609.38718#S7.p1.1)\.
- Choet al\.\(2025\)S\. Cho, Z\. Wu, and A\. KoshiyamaCorrSteer: generation\-time llm steering via correlated sparse autoencoder features\.arXiv preprint arXiv:2508\.12535\.Cited by:[§A\.1](https://arxiv.org/html/2609.38718#A1.SS1.SSS0.Px1.p1.1),[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px3.p1.1),[§A\.3](https://arxiv.org/html/2609.38718#A1.SS3.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p4.1)\.
- Denget al\.\(2025\)J\. Deng, T\. Tang, Y\. Yin, X\. Zhao, J\. Wen,et al\.Neuron based personality trait induction in large language models\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 85059–85083\.Cited by:[§1](https://arxiv.org/html/2609.38718#S1.p4.1),[§4](https://arxiv.org/html/2609.38718#S4.SS0.SSS0.Px1.p1.1)\.
- Engelset al\.\(2025\)J\. Engels, E\. Michaud, I\. Liao, W\. Gurnee, and M\. TegmarkNot all language model features are one\-dimensionally linear\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 84591–84622\.Cited by:[§A\.1](https://arxiv.org/html/2609.38718#A1.SS1.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p2.1),[§3\.1](https://arxiv.org/html/2609.38718#S3.SS1.p1.3)\.
- Fanet al\.\(2026\)C\. Fan, Y\. Cheng, M\. Li, S\. Feizi, and T\. ZhouWhen is your llm steerable?\.arXiv preprint arXiv:2606\.11599\.Cited by:[Appendix J](https://arxiv.org/html/2609.38718#A10.p1.1),[§6](https://arxiv.org/html/2609.38718#S6.SS0.SSS0.Px2.p1.1)\.
- Geigeret al\.\(2024\)A\. Geiger, Z\. Wu, C\. Potts, T\. Icard, and N\. GoodmanFinding alignments between interpretable causal variables and distributed neural representations\.InCausal Learning and Reasoning,pp\. 160–187\.Cited by:[§7](https://arxiv.org/html/2609.38718#S7.p1.1)\.
- Genadiet al\.\(2026\)R\. A\. Genadi, M\. S\. Nwadike, N\. Mukhituly, T\. Hiraoka, H\. AlQuabeh, and K\. InuiSycophancy hides linearly in the attention heads\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 6896–6912\.Cited by:[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p3.1)\.
- Gjølbyeet al\.\(2026\)A\. Gjølbye, L\. K\. Hansen, and S\. KoyejoReasoning models don’t just think longer, they move differently\.arXiv preprint arXiv:2605\.15454\.Cited by:[§6](https://arxiv.org/html/2609.38718#S6.p1.1)\.
- Guet al\.\(2024\)J\. Gu, X\. Jiang, Z\. Shi, H\. Tan, X\. Zhai, C\. Xu, W\. Li, Y\. Shen, S\. Ma, H\. Liu,et al\.A survey on llm\-as\-a\-judge\.The Innovation\.Cited by:[§A\.4](https://arxiv.org/html/2609.38718#A1.SS4.p1.1),[Appendix E](https://arxiv.org/html/2609.38718#A5.p1.1),[§7](https://arxiv.org/html/2609.38718#S7.p2.1)\.
- Guet al\.\(2026\)J\. Gu, X\. Jiang, Z\. Shi, H\. Tan, X\. Zhai, C\. Xu, W\. Li, Y\. Shen, S\. Ma, H\. Liu, S\. Wang, K\. Zhang, Z\. Lin, B\. Zhang, L\. Ni, W\. Gao, Y\. Wang, and J\. GuoA survey on llm\-as\-a\-judge\.The Innovation,pp\. 101253\.External Links:ISSN 2666\-6758,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.xinn.2025.101253),[Link](https://www.sciencedirect.com/science/article/pii/S2666675825004564)Cited by:[§A\.4](https://arxiv.org/html/2609.38718#A1.SS4.p1.1),[Appendix E](https://arxiv.org/html/2609.38718#A5.p1.1),[§7](https://arxiv.org/html/2609.38718#S7.p2.1)\.
- Gurgurovet al\.\(2026\)D\. Gurgurov, Y\. Al Ghussin, T\. Baeumel, C\. Chou, P\. Schramowski, M\. Mosbach, J\. van Genabith, and S\. OstermannCLaS\-bench: a cross\-lingual alignment and steering benchmark\.InFindings of the Association for Computational Linguistics: ACL 2026,pp\. 21591–21628\.Cited by:[§1](https://arxiv.org/html/2609.38718#S1.p4.1),[§4](https://arxiv.org/html/2609.38718#S4.SS0.SSS0.Px1.p1.1)\.
- Heet al\.\(2025a\)Z\. He, Z\. Wang, H\. Xu, and K\. RenTowards llm guardrails via sparse representation steering\.arXiv e\-prints,pp\. arXiv–2503\.Cited by:[§A\.1](https://arxiv.org/html/2609.38718#A1.SS1.SSS0.Px1.p1.1)\.
- Heet al\.\(2024\)Z\. He, W\. Shu, X\. Ge, L\. Chen, J\. Wang, Y\. Zhou, F\. Liu, Q\. Guo, X\. Huang, Z\. Wu,et al\.Llama scope: extracting millions of features from llama\-3\.1\-8b with sparse autoencoders\.arXiv preprint arXiv:2410\.20526\.Cited by:[§A\.1](https://arxiv.org/html/2609.38718#A1.SS1.SSS0.Px1.p1.1),[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px3.p1.1)\.
- Heet al\.\(2025b\)Z\. He, M\. Jin, B\. Shen, A\. Payani, Y\. Zhang, and M\. DuSae\-ssv: supervised steering in sparse representation spaces for reliable control of language models\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 2207–2236\.Cited by:[§A\.1](https://arxiv.org/html/2609.38718#A1.SS1.SSS0.Px1.p1.1),[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px3.p1.1)\.
- Heet al\.\(2025c\)Z\. He, H\. Zhao, Y\. Qiao, F\. Yang, A\. Payani, J\. Ma, and M\. DuSaif: a sparse autoencoder framework for interpreting and steering instruction following of language models\.arXiv preprint arXiv:2502\.11356\.Cited by:[§A\.1](https://arxiv.org/html/2609.38718#A1.SS1.SSS0.Px1.p1.1)\.
- Hendryckset al\.\(2020\)D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. SteinhardtMeasuring massive multitask language understanding\.arXiv preprint arXiv:2009\.03300\.Cited by:[§7](https://arxiv.org/html/2609.38718#S7.p2.1)\.
- Houlsbyet al\.\(2019\)N\. Houlsby, A\. Giurgiu, S\. Jastrzebski, B\. Morrone, Q\. De Laroussilhe, A\. Gesmundo, M\. Attariyan, and S\. GellyParameter\-efficient transfer learning for NLP\.InProceedings of the 36th International Conference on Machine Learning,K\. Chaudhuri and R\. Salakhutdinov \(Eds\.\),Proceedings of Machine Learning Research, Vol\.97,pp\. 2790–2799\.External Links:[Link](https://proceedings.mlr.press/v97/houlsby19a.html)Cited by:[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p3.1)\.
- Hsuet al\.\(2026\)B\. Hsu, D\. Beaglehole, A\. Radhakrishnan, and M\. BelkinContextual linear activation steering of language models\.arXiv preprint arXiv:2604\.24693\.Cited by:[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px2.p1.1),[§A\.3](https://arxiv.org/html/2609.38718#A1.SS3.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p3.1)\.
- Huet al\.\(2022\)E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, W\. Chen,et al\.Lora: low\-rank adaptation of large language models\.\.Iclr1\(2\),pp\. 3\.Cited by:[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p3.1)\.
- Kalajdzievski \(2024\)D\. KalajdzievskiScaling laws for forgetting when fine\-tuning large language models\.arXiv preprint arXiv:2401\.05605\.Cited by:[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px1.p1.1)\.
- Karvonen \(2024\)A\. KarvonenEmergent world models and latent variable estimation in chess\-playing language models\.arXiv preprint arXiv:2403\.15498\.Cited by:[§A\.1](https://arxiv.org/html/2609.38718#A1.SS1.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p2.1)\.
- Konget al\.\(2024\)A\. Kong, S\. Zhao, H\. Chen, Q\. Li, Y\. Qin, R\. Sun, X\. Zhou, E\. Wang, and X\. DongBetter zero\-shot reasoning with role\-play prompting\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),K\. Duh, H\. Gomez, and S\. Bethard \(Eds\.\),Mexico City, Mexico,pp\. 4099–4113\.External Links:[Link](https://aclanthology.org/2024.naacl-long.228/),[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.228)Cited by:[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p3.1)\.
- Lesteret al\.\(2021\)B\. Lester, R\. Al\-Rfou, and N\. ConstantThe power of scale for parameter\-efficient prompt tuning\.InProceedings of the 2021 conference on empirical methods in natural language processing,pp\. 3045–3059\.Cited by:[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p3.1)\.
- Li and Liang \(2021\)X\. L\. Li and P\. LiangPrefix\-tuning: optimizing continuous prompts for generation\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),pp\. 4582–4597\.Cited by:[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p3.1)\.
- Lieberumet al\.\(2024\)T\. Lieberum, S\. Rajamanoharan, A\. Conmy, L\. Smith, N\. Sonnerat, V\. Varma, J\. Kramár, A\. Dragan, R\. Shah, and N\. NandaGemma scope: open sparse autoencoders everywhere all at once on gemma 2\.InProceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP,pp\. 278–300\.Cited by:[§A\.1](https://arxiv.org/html/2609.38718#A1.SS1.SSS0.Px1.p1.1),[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px3.p1.1)\.
- Linet al\.\(2022\)S\. Lin, J\. Hilton, and O\. EvansTruthfulqa: measuring how models mimic human falsehoods\.InProceedings of the 60th annual meeting of the association for computational linguistics \(volume 1: long papers\),pp\. 3214–3252\.Cited by:[§7](https://arxiv.org/html/2609.38718#S7.p2.1)\.
- Luoet al\.\(2026a\)G\. Luo, J\. Feng, T\. Darrell, A\. Radford, and J\. SteinhardtLearning a generative meta\-model of llm activations\.arXiv preprint arXiv:2602\.06964\.Cited by:[§A\.3](https://arxiv.org/html/2609.38718#A1.SS3.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p4.1)\.
- Luoet al\.\(2026b\)H\. Luo, M\. E\. Zarlenga, and M\. JamnikDon’t lose focus: activation steering via key\-orthogonal projections\.arXiv preprint arXiv:2605\.06342\.Cited by:[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p3.1),[§2](https://arxiv.org/html/2609.38718#S2.p3.1)\.
- Marks and Tegmark \(2023\)S\. Marks and M\. TegmarkThe geometry of truth: emergent linear structure in large language model representations of true/false datasets\.arXiv preprint arXiv:2310\.06824\.Cited by:[§7](https://arxiv.org/html/2609.38718#S7.p1.1)\.
- Miehlinget al\.\(2025\)E\. Miehling, M\. Desmond, K\. N\. Ramamurthy, E\. M\. Daly, K\. R\. Varshney, E\. Farchi, P\. Dognin, J\. Rios, D\. Bouneffouf, M\. Liu,et al\.Evaluating the prompt steerability of large language models\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 7874–7900\.Cited by:[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p3.1)\.
- Mishraet al\.\(2026\)A\. Mishra, D\. Khashabi, and A\. LiuSteered llm activations are non\-surjective\.arXiv preprint arXiv:2604\.09839\.Cited by:[§A\.1](https://arxiv.org/html/2609.38718#A1.SS1.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p2.1)\.
- Modellet al\.\(2025\)A\. Modell, P\. Rubin\-Delanchy, and N\. WhiteleyThe origins of representation manifolds in large language models\.arXiv preprint arXiv:2505\.18235\.Cited by:[§A\.1](https://arxiv.org/html/2609.38718#A1.SS1.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p2.1),[§3\.1](https://arxiv.org/html/2609.38718#S3.SS1.p1.3)\.
- Mozikovet al\.\(2024\)M\. Mozikov, N\. Severin, V\. Bodishtianu, M\. Glushanina, I\. Nasonov, D\. Orekhov, V\. Pekhotin, I\. Makovetskiy, M\. Baklashkin, V\. Lavrentyev,et al\.Eai: emotional decision\-making of llms in strategic games and ethical dilemmas\.Advances in Neural Information Processing Systems37,pp\. 53969–54002\.Cited by:[§1](https://arxiv.org/html/2609.38718#S1.p4.1),[§5](https://arxiv.org/html/2609.38718#S5.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2609.38718#S5.SS0.SSS0.Px3.p1.1)\.
- Nguyenet al\.\(2025\)D\. Nguyen, A\. Prasad, E\. Stengel\-Eskin, and M\. BansalMulti\-attribute steering of language models via targeted intervention\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 20619–20634\.Cited by:[§A\.1](https://arxiv.org/html/2609.38718#A1.SS1.SSS0.Px1.p1.1),[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p2.1),[§2](https://arxiv.org/html/2609.38718#S2.p3.1)\.
- Nguyenet al\.\(2026\)T\. Nguyen, T\. A\. Nguyen, S\. Alemohammad, and R\. G\. BaraniukMinimizing collateral damage in activation steering\.arXiv preprint arXiv:2605\.01167\.Cited by:[§A\.3](https://arxiv.org/html/2609.38718#A1.SS3.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p4.1)\.
- Nguyen and Le \(2026\)T\. Nguyen and T\. LeBeyond linear activation steering: invertible latent transformations for controlling llm behavior\.arXiv preprint arXiv:2606\.08454\.Cited by:[§A\.1](https://arxiv.org/html/2609.38718#A1.SS1.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p2.1)\.
- Ostermannet al\.\(2026\)S\. Ostermann, D\. Gurgurov, T\. Baeumel, M\. A\. Hedderich, S\. Lapuschkin, W\. Samek, and V\. SchmittFrom weights to activations: is steering the next frontier of adaptation?\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 29854–29879\.Cited by:[Appendix F](https://arxiv.org/html/2609.38718#A6.p1.1)\.
- Panicksseryet al\.\(2023\)N\. Panickssery, N\. Gabrieli, J\. Schulz, M\. Tong, E\. Hubinger, and A\. M\. TurnerSteering llama 2 via contrastive activation addition\.arXiv preprint arXiv:2312\.06681\.Cited by:[§A\.1](https://arxiv.org/html/2609.38718#A1.SS1.SSS0.Px1.p1.1),[§A\.3](https://arxiv.org/html/2609.38718#A1.SS3.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p2.1),[§2](https://arxiv.org/html/2609.38718#S2.p4.1)\.
- Parket al\.\(2025\)C\. F\. Park, A\. Lee, E\. S\. Lubana, Y\. Yang, M\. Okawa, K\. Nishi, M\. Wattenberg, and H\. TanakaIclr: in\-context learning of representations\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 53258–53284\.Cited by:[§2](https://arxiv.org/html/2609.38718#S2.p3.1),[§6](https://arxiv.org/html/2609.38718#S6.p1.1)\.
- Patilet al\.\(2025\)H\. V\. Patil, V\. Sanam, and M\. P\. AtreStacked LoRA: isolated low\-rank adaptation for lifelong knowledge management\.InThe 14th International Joint Conference on Natural Language Processing and The 4th Conference of the Asia\-Pacific Chapter of the Association for Computational Linguistics,S\. T\.y\.s\.s, S\. Shimizu, and Y\. Gong \(Eds\.\),Mumbai, India,pp\. 36–46\.External Links:[Link](https://aclanthology.org/2025.ijcnlp-srw.4/),[Document](https://dx.doi.org/10.18653/v1/2025.ijcnlp-srw.4),ISBN 979\-8\-89176\-304\-3Cited by:[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px1.p1.1)\.
- Rainaet al\.\(2025\)S\. Raina, S\. Aggarwal, A\. Chadha, V\. Jain, and A\. DasD\-steer\-preference alignment techniques learn to behave, not to believe–beneath the surface, dpo as steering vector perturbation in activation space\.arXiv preprint arXiv:2512\.11838\.Cited by:[§A\.3](https://arxiv.org/html/2609.38718#A1.SS3.SSS0.Px3.p1.1),[Appendix B](https://arxiv.org/html/2609.38718#A2.SS0.SSS0.Px5.p1.2),[Appendix B](https://arxiv.org/html/2609.38718#A2.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p4.1),[§3\.1](https://arxiv.org/html/2609.38718#S3.SS1.p1.2),[§3\.1](https://arxiv.org/html/2609.38718#S3.SS1.p1.3)\.
- Rimskyet al\.\(2024\)N\. Rimsky, N\. Gabrieli, J\. Schulz, M\. Tong, E\. Hubinger, and A\. TurnerSteering llama 2 via contrastive activation addition\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 15504–15522\.Cited by:[Appendix F](https://arxiv.org/html/2609.38718#A6.p1.1)\.
- Saglamet al\.\(2025\)B\. Saglam, P\. Kassianik, B\. Nelson, S\. Weerawardhena, Y\. Singer, and A\. KarbasiLarge language models encode semantics and alignment in linearly separable representations\.InProceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia\-Pacific Chapter of the Association for Computational Linguistics,pp\. 2282–2303\.Cited by:[§A\.1](https://arxiv.org/html/2609.38718#A1.SS1.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p2.1)\.
- Sooet al\.\(2025\)S\. Soo, C\. Guang, W\. Teng, C\. Balaganesh, T\. Guoxian, and Y\. MingInterpretable steering of large language models with feature guided activation additions\.arXiv preprint arXiv:2501\.09929\.Cited by:[§A\.1](https://arxiv.org/html/2609.38718#A1.SS1.SSS0.Px2.p1.1),[§A\.3](https://arxiv.org/html/2609.38718#A1.SS3.SSS0.Px2.p1.1),[§A\.4](https://arxiv.org/html/2609.38718#A1.SS4.p1.1),[Appendix D](https://arxiv.org/html/2609.38718#A4.SS0.SSS0.Px1.p1.1),[§E\.1](https://arxiv.org/html/2609.38718#A5.SS1.p1.1),[§E\.2](https://arxiv.org/html/2609.38718#A5.SS2.p3.pic1.1.1.1),[Appendix E](https://arxiv.org/html/2609.38718#A5.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p4.1)\.
- Sunet al\.\(2025\)J\. Sun, S\. Baskaran, Z\. Wu, M\. Sklar, C\. Potts, and A\. GeigerHypersteer: activation steering at scale with hypernetworks\.arXiv preprint arXiv:2506\.03292\.Cited by:[§A\.3](https://arxiv.org/html/2609.38718#A1.SS3.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p4.1)\.
- Tanet al\.\(2024\)D\. Tan, D\. Chanin, A\. Lynch, B\. Paige, D\. Kanoulas, A\. Garriga\-Alonso, and R\. KirkAnalysing the generalisation and reliability of steering vectors\.Advances in Neural Information Processing Systems37,pp\. 139179–139212\.Cited by:[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px3.p1.1)\.
- Trunget al\.\(2024\)L\. Trung, X\. Zhang, Z\. Jie, P\. Sun, X\. Jin, and H\. LiReFT: reasoning with reinforced fine\-tuning\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 7601–7614\.External Links:[Link](https://aclanthology.org/2024.acl-long.410/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.410)Cited by:[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p3.1)\.
- Turneret al\.\(2023\)A\. M\. Turner, L\. Thiergart, G\. Leech, D\. Udell, J\. J\. Vazquez, U\. Mini, and M\. MacDiarmidSteering language models with activation engineering\.arXiv preprint arXiv:2308\.10248\.Cited by:[§A\.1](https://arxiv.org/html/2609.38718#A1.SS1.SSS0.Px1.p1.1),[§A\.4](https://arxiv.org/html/2609.38718#A1.SS4.p1.1),[Appendix E](https://arxiv.org/html/2609.38718#A5.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p2.1)\.
- Vu and Nguyen \(2026\)M\. H\. Vu and T\. NguyenAngular steering: behavior control via rotation in activation space\.Advances in Neural Information Processing Systems38,pp\. 121653–121690\.Cited by:[§A\.1](https://arxiv.org/html/2609.38718#A1.SS1.SSS0.Px2.p1.1),[§A\.3](https://arxiv.org/html/2609.38718#A1.SS3.SSS0.Px2.p1.1)\.
- Wanget al\.\(2025\)A\. Wang, D\. Shu, Y\. Wang, Y\. Ma, and M\. DuImproving llm reasoning through interpretable role\-playing steering\.arXiv preprint arXiv:2506\.07335\.Cited by:[§A\.1](https://arxiv.org/html/2609.38718#A1.SS1.SSS0.Px1.p1.1)\.
- Wanget al\.\(2024\)P\. Wang, L\. Li, L\. Chen, Z\. Cai, D\. Zhu, B\. Lin, Y\. Cao, L\. Kong, Q\. Liu, T\. Liu, and Z\. SuiLarge language models are not fair evaluators\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 9440–9450\.External Links:[Link](https://aclanthology.org/2024.acl-long.511/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.511)Cited by:[§A\.4](https://arxiv.org/html/2609.38718#A1.SS4.p1.1),[Appendix E](https://arxiv.org/html/2609.38718#A5.p1.1)\.
- Weiet al\.\(2022\)J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, F\. Xia, E\. Chi, Q\. V\. Le, D\. Zhou,et al\.Chain\-of\-thought prompting elicits reasoning in large language models\.Advances in neural information processing systems35,pp\. 24824–24837\.Cited by:[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p3.1)\.
- Wuet al\.\(2023\)N\. Wu, M\. Gong, L\. Shou, S\. Liang, and D\. JiangLarge language models are diverse role\-players for summarization evaluation\.InCCF international conference on natural language processing and Chinese computing,pp\. 695–707\.Cited by:[Appendix E](https://arxiv.org/html/2609.38718#A5.p1.1)\.
- Wuet al\.\(2025\)Z\. Wu, A\. Arora, A\. Geiger, Z\. Wang, J\. Huang, D\. Jurafsky, C\. D\. Manning, and C\. PottsAxbench: steering llms? even simple baselines outperform sparse autoencoders\.arXiv preprint arXiv:2501\.17148\.Cited by:[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px1.p1.1),[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px3.p1.1),[§A\.3](https://arxiv.org/html/2609.38718#A1.SS3.SSS0.Px1.p1.1),[§A\.4](https://arxiv.org/html/2609.38718#A1.SS4.p1.1),[§M\.1](https://arxiv.org/html/2609.38718#A13.SS1.p1.1),[§M\.2](https://arxiv.org/html/2609.38718#A13.SS2.p1.1),[§M\.3](https://arxiv.org/html/2609.38718#A13.SS3.p1.1),[§M\.5](https://arxiv.org/html/2609.38718#A13.SS5.SSS0.Px1.p1.1),[Appendix M](https://arxiv.org/html/2609.38718#A13.p1.1),[Appendix D](https://arxiv.org/html/2609.38718#A4.SS0.SSS0.Px1.p1.1),[§E\.1](https://arxiv.org/html/2609.38718#A5.SS1.p1.1),[Appendix E](https://arxiv.org/html/2609.38718#A5.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p4.1),[§2](https://arxiv.org/html/2609.38718#S2.p3.1),[§3\.4](https://arxiv.org/html/2609.38718#S3.SS4.p1.1),[§4](https://arxiv.org/html/2609.38718#S4.SS0.SSS0.Px1.p1.1)\.
- Wuet al\.\(2026\)Z\. Wu, Q\. Yu, A\. Arora, C\. D\. Manning, and C\. PottsImproved representation steering for language models\.Advances in Neural Information Processing Systems38,pp\. 160589–160641\.Cited by:[§7](https://arxiv.org/html/2609.38718#S7.p1.1)\.
- Xionget al\.\(2024\)Z\. Xiong, Z\. Cai, J\. Cooper, A\. Ge, V\. Papageorgiou, Z\. Sifakis, A\. Giannou, Z\. Lin, L\. Yang, S\. Agarwal,et al\.Everything everywhere all at once: llms can in\-context learn multiple tasks in superposition\.arXiv preprint arXiv:2410\.05603\.Cited by:[§A\.1](https://arxiv.org/html/2609.38718#A1.SS1.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p2.1)\.
- Xuet al\.\(2024\)M\. Xu, S\. Sharmin, and D\. P\. MandicGeometry is all you need: a unified taxonomy of matrix and tensor factorization for compression of generative language models\.arXiv preprint arXiv:2410\.03040\.Cited by:[§6](https://arxiv.org/html/2609.38718#S6.p1.1)\.
- Yanet al\.\(2026\)L\. Yan, R\. Li, G\. Chen, Q\. Li, J\. Geng, W\. Li, L\. Wang, and C\. LyuSpurious rewards paradox: mechanistically understanding how rlvr activates memorization shortcuts in llms\.arXiv preprint arXiv:2601\.11061\.Cited by:[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p3.1)\.
- Yu \(2025\)F\. YuWhen ais judge ais: the rise of agent\-as\-a\-judge evaluation for llms\.arXiv preprint arXiv:2508\.02994\.Cited by:[§A\.4](https://arxiv.org/html/2609.38718#A1.SS4.p1.1),[Appendix E](https://arxiv.org/html/2609.38718#A5.p1.1),[§7](https://arxiv.org/html/2609.38718#S7.p2.1)\.
- Yuet al\.\(2026\)H\. Yu, J\. Liu, Z\. Yan, H\. Lin, and X\. ZhangWASD: locating critical neurons as sufficient conditions for explaining and controlling llm behavior\.arXiv preprint arXiv:2603\.18474\.Cited by:[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p3.1)\.
- Zhang and Nanda \(2024\)F\. Zhang and N\. NandaTowards best practices of activation patching in language models: metrics and methods\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 1651–1678\.Cited by:[§A\.2](https://arxiv.org/html/2609.38718#A1.SS2.SSS0.Px2.p1.1)\.
- Zhanget al\.\(2023\)X\. Zhang, B\. Yu, H\. Yu, Y\. Lv, T\. Liu, F\. Huang, H\. Xu, and Y\. LiWider and deeper llm networks are fairer llm evaluators\.arXiv preprint arXiv:2308\.01862\.Cited by:[§A\.4](https://arxiv.org/html/2609.38718#A1.SS4.p1.1),[Appendix E](https://arxiv.org/html/2609.38718#A5.p1.1),[§7](https://arxiv.org/html/2609.38718#S7.p2.1)\.
- Zhaoet al\.\(2026\)H\. Zhao, H\. Sun, J\. Kong, X\. Li, Q\. Wang, L\. Jiang, Q\. Zhu, T\. Abdelzaher, Y\. Choi, M\. Li,et al\.Odesteer: a unified ode\-based steering framework for llm alignment\.arXiv preprint arXiv:2602\.17560\.Cited by:[§A\.1](https://arxiv.org/html/2609.38718#A1.SS1.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38718#S1.p1.1),[§2](https://arxiv.org/html/2609.38718#S2.p2.1)\.
- Zhenget al\.\(2023\)L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. Xing,et al\.Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.Advances in neural information processing systems36,pp\. 46595–46623\.Cited by:[§A\.4](https://arxiv.org/html/2609.38718#A1.SS4.p1.1),[Appendix E](https://arxiv.org/html/2609.38718#A5.p1.1),[§7](https://arxiv.org/html/2609.38718#S7.p2.1)\.
- Zhouet al\.\(2025\)J\. Zhou, Y\. Chen, Y\. Shi, X\. Zhang, L\. Lei, Y\. Feng, Z\. Xiong, M\. Yan, X\. Wang, Y\. Cao,et al\.Socialeval: evaluating social intelligence of large language models\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 30958–31012\.Cited by:[§1](https://arxiv.org/html/2609.38718#S1.p4.1),[§5](https://arxiv.org/html/2609.38718#S5.SS0.SSS0.Px2.p1.1)\.
- Zhouet al\.\(2026\)Y\. Zhou, Y\. Wang, X\. Yin, S\. Zhou, and A\. ZhangThe geometry of reasoning: flowing logics in representation space\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 72711–72739\.Cited by:[Appendix I](https://arxiv.org/html/2609.38718#A9.p1.1),[§6](https://arxiv.org/html/2609.38718#S6.p1.1)\.

## Appendix AExtended Related Work

Model steering can be organized as a design space defined by three related questions: what geometric structure is assumed for the target concept in activation space, where the intervention is applied, and how the intervention is computed or learned\. This organization offers a more direct account of the literature than a division based solely on whether a method modifies the prompt, the activations, or the parameters: prompt\- and parameter\-based methods correspond primarily to different intervention locations, whereas activation\-based methods additionally require an explicit choice of geometric assumption and learning procedure\.

### A\.1Geometric Assumption

##### Linear and feature\-based representations\.

A dominant assumption in representation\-level steering is that semantically meaningful concepts correspond approximately to directions, hyperplanes, or low\-dimensional subspaces in activation space\([Xiong et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib10);[Bereska and Gavves, 2024](https://arxiv.org/html/2609.38718#bib.bib9);[Saglam et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib11)\)\. Under this view, a desired behavior can be represented by a single vector added to the model’s activation at a chosen layer and intensity\. Contrastive Activation Addition \(CAA\)\([Panickssery et al\., 2023](https://arxiv.org/html/2609.38718#bib.bib34)\)is a prominent instance, constructing a steering vector from the difference between activations associated with desired and undesired behavior; related work shows that such directions can be composed and sometimes transfer across tasks, concepts, or models\([Turner et al\., 2023](https://arxiv.org/html/2609.38718#bib.bib35);[Karvonen, 2024](https://arxiv.org/html/2609.38718#bib.bib64);[Nguyen et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib16)\)\. Sparse\-autoencoder \(SAE\) methods offer a related, feature\-based instantiation of the linearity assumption: rather than manipulating undifferentiated coordinates of the residual stream, they identify sparse latent features intended to correspond to interpretable concepts\([Lieberum et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib39);[He et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib38)\), which have in turn been used to steer safety, fairness, truthfulness, instruction\-following, and role\-play behavior\([He et al\., 2025a](https://arxiv.org/html/2609.38718#bib.bib6);[He et al\., 2025c](https://arxiv.org/html/2609.38718#bib.bib8);[Wang et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib7)\)\. The reliability, disentanglement, and causal significance of these features nonetheless remain debated\([He et al\., 2025b](https://arxiv.org/html/2609.38718#bib.bib5);[Arad et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib32);[Cho et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib31)\)\.

##### Beyond the linear assumption\.

Recent work challenges the premise that concepts are represented by a single, fixed direction, showing instead that they may occupy curved, anisotropic, or otherwise non\-Euclidean manifolds\([Engels et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib12);[Modell et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib13)\)\. If a concept’s representation varies with the input, a global vector cannot capture the intervention appropriate to every context, motivating nonlinear, context\-conditioned steering procedures\([Nguyen and Le, 2026](https://arxiv.org/html/2609.38718#bib.bib14);[Zhao et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib15);[Mishra et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib20)\)as well as more expressive transformations such as learned steering operators and rotations\([Soo et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib33);[Vu and Nguyen, 2026](https://arxiv.org/html/2609.38718#bib.bib49)\)\. The geometric question is therefore not only whether a concept admits a direction, but whether the appropriate transformation should vary across inputs, layers, or behavioral objectives\.*MetaSteer*learns a nonlinear, context\-dependent intervention rather than fixing a single global steering direction\. We characterize the resulting geometric variation empirically through displacement and trajectory analyses, without assuming that the intervention’s effective dimensionality equals the intrinsic dimensionality of the hidden\-state trajectory\.

### A\.2Intervention Site

The location of an intervention determines which part of the computation is modified and how directly it interacts with the model’s existing representations\. Steering methods range from interventions applied before generation begins to interventions applied within individual computational components or directly to model parameters\.

##### Input\- and parameter\-level steering\.

At the input level, prompt\-based steering modifies the conditions under which the model generates its answer: Chain\-of\-Thought prompting decomposes a problem into intermediate reasoning steps\([Wei et al\., 2022](https://arxiv.org/html/2609.38718#bib.bib63)\), role\-playing and persona\-based prompts assign the model an identity or behavioral frame\([Kong et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib56);[Miehling et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib54)\), and AxBench shows that prompts generated by a stronger model can steer a weaker one\([Wu et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib2)\)\. These approaches are easy to deploy because they leave parameters and activations untouched, but their effectiveness can depend on prompt\-engineering expertise or on access to a stronger supervising model\. At the parameter level, fine\-tuning changes the model weights or introduces trainable modules that persist across inputs, including adapters\([Houlsby et al\., 2019](https://arxiv.org/html/2609.38718#bib.bib60)\), prompt tuning\([Lester et al\., 2021](https://arxiv.org/html/2609.38718#bib.bib55)\), prefix tuning\([Li and Liang, 2021](https://arxiv.org/html/2609.38718#bib.bib53)\), Low\-Rank Adaptation \(LoRA\)\([Hu et al\., 2022](https://arxiv.org/html/2609.38718#bib.bib51)\), and compact representation\-level interventions such as Representation Fine\-Tuning \(ReFT\)\([Trung et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib52)\), with scaling\-aware and stacked adaptation strategies proposed to reduce forgetting under continued adaptation\([Kalajdzievski, 2024](https://arxiv.org/html/2609.38718#bib.bib58);[Patil et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib59)\)\. These methods make the intervention persistent and trainable, at the cost of additional training requirements and, potentially, reduced preservation of unrelated capabilities\.

##### Internal activation sites\.

For runtime activation steering, the residual stream is the most common intervention site, as it is accessible throughout the forward pass and provides a shared space through which information propagates\([Nguyen et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib16);[Hsu et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib17)\); this makes it convenient for vector addition, activation patching, and contextual regulation of steering strength, but modifying it can affect many downstream computations at once and so limit specificity\. Recent work instead targets more localized components: MLP activations have been used to control particular behaviors or suppress spurious associations\([Yu et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib18);[Yan et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib19)\), and individual attention heads have been identified or manipulated to control behaviors such as sycophancy\([Genadi et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib21);[Luo et al\., 2026b](https://arxiv.org/html/2609.38718#bib.bib25)\), complementing activation\-patching and causal\-tracing tools that more generally locate components whose activations causally influence a target behavior\([Zhang and Nanda, 2024](https://arxiv.org/html/2609.38718#bib.bib48)\)\. Attention projection matrices offer a further, more structured site: rather than adding a fixed vector after a layer has computed its representation, modifying a projection matrix changes how the current hidden state is transformed, so that the effect of the intervention depends on the representation being processed\. The closest work to ours is[Luo et al\. \(2026b\)](https://arxiv.org/html/2609.38718#bib.bib25), which intervenes only on the query projection;*MetaSteer*instead intervenes jointly on all four attention projection matrices—query, key, value, and output—obtaining context\-dependent behavior by construction rather than through an additive vector\.

##### Feature\-level locations\.

SAE\-based steering can also be viewed as selecting a feature\-level location within the representation space: dictionary\-learning methods identify sparse features and modify the activation of a selected one\([Lieberum et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib39);[He et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib38);[Wu et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib2)\), which can improve interpretability and offer more targeted control than modifying the full residual stream\. However, SAE checkpoints are available for only a subset of model families, and the meaning and causal role of individual features can vary across models and tasks\([He et al\., 2025b](https://arxiv.org/html/2609.38718#bib.bib5);[Arad et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib32);[Cho et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib31)\); more broadly, the effectiveness of activation steering is known to be sensitive to the choice of model, task, layer, and component\([Tan et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib24)\)\. These limitations motivate methods that exploit a structured intervention site without requiring a manually identified feature for every target concept\.

### A\.3Learning Procedure

Once a site has been selected, a second problem is how to determine the transformation applied there\. Existing methods differ in whether the intervention is specified heuristically, estimated from contrastive activations, learned as a causal operator, or optimized from behavioral or preference feedback\.

##### Heuristic and contrastive construction\.

The simplest procedures use a manually chosen direction and tune only its magnitude at inference time\. Contrastive methods such as CAA estimate a direction from paired examples of desired and undesired behavior\([Panickssery et al\., 2023](https://arxiv.org/html/2609.38718#bib.bib34)\), while other approaches rely on heuristic transformations or statistical structure in activation space, including the PCA\- and LAT\-based methods surveyed in AxBench\([Wu et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib2)\)\. These procedures are computationally inexpensive and effective when the target behavior is well described by a stable direction, but they require the practitioner to choose an appropriate layer, direction, and strength, and may not adapt well when the relevant representation shifts across queries\.

##### Learned operators and contextual control\.

To increase flexibility, several methods learn the intervention itself rather than specifying a fixed vector: some learn to predict the causal effect of an intervention\([Arad et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib32);[Cho et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib31)\), others replace vector addition with a learned steering operator\([Soo et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib33);[Vu and Nguyen, 2026](https://arxiv.org/html/2609.38718#bib.bib49)\), and contextual methods regulate the direction or intensity of steering according to the current input\([Hsu et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib17)\)\.[Sun et al\. \(2025\)](https://arxiv.org/html/2609.38718#bib.bib29)learns steering vectors end\-to\-end via cross\-attention over semantic embeddings, but extracts them using a separate, architecturally identical model;*MetaSteer*instead uses the same model throughout\. These approaches improve expressiveness relative to heuristic construction, though their performance can depend on task\-specific training data, high\-quality representations, or carefully designed auxiliary modules\.

##### Objective\-based and preference\-based learning\.

A complementary line of work formulates steering as an optimization problem that balances the desired behavioral change against the preservation of general model utility, tuning both intervention orientation and intensity to trade off steering effectiveness𝕊\\mathbb\{S\}against utility𝕌\\mathbb\{U\}\([Aravindan et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib26);[Nguyen et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib27);[Luo et al\., 2026a](https://arxiv.org/html/2609.38718#bib.bib28)\), making explicit that a stronger intervention is not necessarily preferable if it damages unrelated capabilities\. Preference\-based methods provide a related but distinct learning signal: rather than manually specifying an activation direction, they use comparisons between preferred and non\-preferred behaviors to learn how the model should be steered, for instance by cloning desired behavior directly in activation space\([Raina et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib1)\)\. Such objectives make the learning signal more directly behavioral, but remain constrained by the representation through which the preference signal is transmitted; a fixed representation can in particular create an information bottleneck when it cannot express every behavioral distinction the target task requires\([Raina et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib1)\)\.*MetaSteer*is closest in learning approach to[Raina et al\. \(2025\)](https://arxiv.org/html/2609.38718#bib.bib1), building on preference optimization to learn a transferable policy over interventions, but avoids this bottleneck by learning nonlinear, context\-dependent interventions rather than a single low\-dimensional vector per concept\.

##### Efficiency and transferability\.

These procedures impose different data and computational requirements: heuristic and contrastive methods are inexpensive but may require manual tuning or concept\-specific examples; learned operators and parameter\-efficient adaptations are more flexible but require additional training and may specialize to the tasks seen during training; and preference\-based methods offer a more direct behavioral objective whose generalization depends on the coverage and quality of the preference data\. These trade\-offs motivate methods that learn expressive interventions from limited supervision while still transferring to unseen concepts and tasks\.

### A\.4Evaluation

Although evaluation is not itself an intervention dimension, it determines how improvements along the geometric, site, and learning dimensions are measured\. The increasing capability of frontier language models, together with the cost of human annotation, has motivated using LLMs as proxies for human evaluators\([Zheng et al\., 2023](https://arxiv.org/html/2609.38718#bib.bib4)\); such LLM\-as\-a\-Judge \(LaaJ\) frameworks are now widely used in automated evaluation pipelines, including in multi\-agent settings\([Zhang et al\., 2023](https://arxiv.org/html/2609.38718#bib.bib50);[Gu et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib57);[Gu et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib66);[Yu, 2025](https://arxiv.org/html/2609.38718#bib.bib67)\), and prior work reports substantial agreement between LLM judges and human annotators while flagging open concerns around reliability, fairness, reproducibility, calibration, and systematic bias\([Wang et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib3);[Gu et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib57);[Gu et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib66);[Yu, 2025](https://arxiv.org/html/2609.38718#bib.bib67)\)\. These concerns bear directly on steering research, where evaluation typically relies on a single frontier model as judge and, in some cases, the same model family that contributed training or evaluation data also serves as evaluator, creating the possibility of systematic bias toward behaviors that family favors\([Wu et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib2)\); prior studies of activation steering and SAE\-based interventions have adopted this paradigm\([Wu et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib2);[Soo et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib33);[Arad et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib32);[Turner et al\., 2023](https://arxiv.org/html/2609.38718#bib.bib35)\), despite the reliability and robustness of LLM judges for measuring steering quality remaining insufficiently characterized\. This motivates the careful assessment of automated evaluation alongside the steering results we report\.

### A\.5Extension to Multi\-Head, Grouped\-Query Attention, and RoPE

The derivation above treats a single attention head with no positional encoding\. We now give the general form used by every backbone in this work, each of which combines grouped\-query attention with RoPE\.

##### Setup\.

LetHHdenote the number of query heads andHk​v≤HH\_\{kv\}\\leq Hthe number of key/value heads, with group sizeg=H/Hk​vg=H/H\_\{kv\}; query headhhshares key/value headκ⁡\(h\)=⌈h/g⌉\\kappa\(h\)=\\lceil h/g\\rceil\. Standard multi\-head attention is the special caseHk​v=HH\_\{kv\}=H\. Per\-head baseline projections are

𝐪i\(h\)=𝐳i​𝐖q\(h\),𝐤j\(κ\)=𝐳j​𝐖k\(κ\),𝐯j\(κ\)=𝐳j​𝐖v\(κ\),\\mathbf\{q\}\_\{i\}^\{\(h\)\}=\\mathbf\{z\}\_\{i\}\\mathbf\{W\}\_\{q\}^\{\(h\)\},\\qquad\\mathbf\{k\}\_\{j\}^\{\(\\kappa\)\}=\\mathbf\{z\}\_\{j\}\\mathbf\{W\}\_\{k\}^\{\(\\kappa\)\},\\qquad\\mathbf\{v\}\_\{j\}^\{\(\\kappa\)\}=\\mathbf\{z\}\_\{j\}\\mathbf\{W\}\_\{v\}^\{\(\\kappa\)\},the corresponding column\-blocks of𝐖q,𝐖k,𝐖v\\mathbf\{W\}\_\{q\},\\mathbf\{W\}\_\{k\},\\mathbf\{W\}\_\{v\}, with adaptersΔ​𝐖q\(h\),Δ​𝐖k\(κ\)\\Delta\\mathbf\{W\}\_\{q\}^\{\(h\)\},\\Delta\\mathbf\{W\}\_\{k\}^\{\(\\kappa\)\}the matching blocks ofΔ​𝐖q,Δ​𝐖k\\Delta\\mathbf\{W\}\_\{q\},\\Delta\\mathbf\{W\}\_\{k\}from Eq\.[3](https://arxiv.org/html/2609.38718#S3.E3), since LoRA is applied to the full projection matrix prior to reshaping into heads\.

##### RoPE rotation\.

Before the dot product, RoPE applies a position\-dependent orthogonal rotation𝐑Θ,m∈ℝdh×dh\\mathbf\{R\}\_\{\\Theta,m\}\\in\\mathbb\{R\}^\{d\_\{h\}\\times d\_\{h\}\}to each head’s query and key vectors,𝐪^i\(h\)=𝐪i\(h\)​𝐑Θ,i\\hat\{\\mathbf\{q\}\}\_\{i\}^\{\(h\)\}=\\mathbf\{q\}\_\{i\}^\{\(h\)\}\\mathbf\{R\}\_\{\\Theta,i\}and𝐤^j\(κ\)=𝐤j\(κ\)​𝐑Θ,j\\hat\{\\mathbf\{k\}\}\_\{j\}^\{\(\\kappa\)\}=\\mathbf\{k\}\_\{j\}^\{\(\\kappa\)\}\\mathbf\{R\}\_\{\\Theta,j\}, satisfying the defining relative\-position identity𝐑Θ,i​𝐑Θ,j⊤=𝐑Θ,i−j\\mathbf\{R\}\_\{\\Theta,i\}\\mathbf\{R\}\_\{\\Theta,j\}^\{\\top\}=\\mathbf\{R\}\_\{\\Theta,i\-j\}\. The per\-head logit is therefore

si​j\(h\)=1dh​𝐪i\(h\)​𝐑Θ,i−j​\(𝐤j\(κ⁡\(h\)\)\)⊤,s\_\{ij\}^\{\(h\)\}=\\frac\{1\}\{\\sqrt\{d\_\{h\}\}\}\\,\\mathbf\{q\}\_\{i\}^\{\(h\)\}\\mathbf\{R\}\_\{\\Theta,i\-j\}\\big\(\\mathbf\{k\}\_\{j\}^\{\(\\kappa\(h\)\)\}\\big\)^\{\\top\},depending only on the relative offseti−ji\-j\.

##### Per\-head perturbation\.

Repeating the expansion of Equations[22](https://arxiv.org/html/2609.38718#A3.E22)–[24](https://arxiv.org/html/2609.38718#A3.E24)per head, with𝐑Θ,i−j\\mathbf\{R\}\_\{\\Theta,i\-j\}inserted between the query\-side and key\-side factor exactly as above, the three cross terms again collapse into a single bilinear form, now indexed by head and relative position:

δi​j\(h\)=1dh​𝐳i​𝐌\(h\)​\(i−j\)​𝐳j⊤,\\delta\_\{ij\}^\{\(h\)\}=\\frac\{1\}\{\\sqrt\{d\_\{h\}\}\}\\,\\mathbf\{z\}\_\{i\}\\,\\mathbf\{M\}^\{\(h\)\}\(i\-j\)\\,\\mathbf\{z\}\_\{j\}^\{\\top\},\(15\)𝐌\(h\)​\(m\):=\\displaystyle\\mathbf\{M\}^\{\(h\)\}\(m\):=λk​𝐖q\(h\)​𝐑Θ,m​𝐁k\(κ⁡\(h\)\)​𝐀k\(κ⁡\(h\)\)⊤\\displaystyle\\lambda\_\{k\}\\,\\mathbf\{W\}\_\{q\}^\{\(h\)\}\\mathbf\{R\}\_\{\\Theta,m\}\\mathbf\{B\}\_\{k\}^\{\(\\kappa\(h\)\)\}\\mathbf\{A\}\_\{k\}^\{\(\\kappa\(h\)\)\\top\}\(16\)\+λq​𝐀q\(h\)​𝐁q\(h\)⊤​𝐑Θ,m​𝐖k\(κ⁡\(h\)\)⊤\\displaystyle\+\\lambda\_\{q\}\\,\\mathbf\{A\}\_\{q\}^\{\(h\)\}\\mathbf\{B\}\_\{q\}^\{\(h\)\\top\}\\mathbf\{R\}\_\{\\Theta,m\}\\mathbf\{W\}\_\{k\}^\{\(\\kappa\(h\)\)\\top\}\+λq​λk​𝐀q\(h\)​𝐁q\(h\)⊤​𝐑Θ,m​𝐁k\(κ⁡\(h\)\)​𝐀k\(κ⁡\(h\)\)⊤\.\\displaystyle\+\\lambda\_\{q\}\\lambda\_\{k\}\\,\\mathbf\{A\}\_\{q\}^\{\(h\)\}\\mathbf\{B\}\_\{q\}^\{\(h\)\\top\}\\mathbf\{R\}\_\{\\Theta,m\}\\mathbf\{B\}\_\{k\}^\{\(\\kappa\(h\)\)\}\\mathbf\{A\}\_\{k\}^\{\(\\kappa\(h\)\)\\top\}\.Equation[25](https://arxiv.org/html/2609.38718#A3.E25)is recovered exactly whenH=Hk​v=1H=H\_\{kv\}=1and𝐑Θ,m=𝐈\\mathbf\{R\}\_\{\\Theta,m\}=\\mathbf\{I\}\. Since𝐑Θ,m\\mathbf\{R\}\_\{\\Theta,m\}is a fixed orthogonal matrix independent of𝐳i,𝐳j\\mathbf\{z\}\_\{i\},\\mathbf\{z\}\_\{j\}, the genericity argument following Equation[26](https://arxiv.org/html/2609.38718#A3.E26)applies unchanged to each𝐌\(h\)​\(m\)\\mathbf\{M\}^\{\(h\)\}\(m\): both partial derivatives are non\-zero except on a measure\-zero set, now for every headhhand every relative offsetmmrealized by the context window\.

##### GQA\-induced coupling\.

Under grouped\-query attention \(Hk​v<HH\_\{kv\}<H\), theggquery heads sharing key/value headκ\\kappahave distinct𝐌\(h\)​\(m\)\\mathbf\{M\}^\{\(h\)\}\(m\)\(through𝐖q\(h\),𝐀q\(h\),𝐁q\(h\)\\mathbf\{W\}\_\{q\}^\{\(h\)\},\\mathbf\{A\}\_\{q\}^\{\(h\)\},\\mathbf\{B\}\_\{q\}^\{\(h\)\}\) but share the same key\-side adapter𝐀k\(κ\),𝐁k\(κ\)\\mathbf\{A\}\_\{k\}^\{\(\\kappa\)\},\\mathbf\{B\}\_\{k\}^\{\(\\kappa\)\}: a single trained key\-side update therefore perturbsggheads’ attention patterns simultaneously, each through a different query\-side projection\.

## Appendix BDerivation of the Logit\-Gap Identity

Equation[17](https://arxiv.org/html/2609.38718#A2.E17)states that, under a softmax output parameterization, the log\-probability gap between two completions is linear in the final hidden state\. This identity is derived by[Raina et al\. \(2025\)](https://arxiv.org/html/2609.38718#bib.bib1)\(Section 2\.2\) and underlies the single\-vector characterization of DPO recalled in Section[3\.1](https://arxiv.org/html/2609.38718#S3.SS1)\. We reproduce the derivation here for completeness\.

##### Setup\.

Under standard softmax parameterization, the model’s conditional distribution over the next tokenyyis

π⁡\(y∣x\)\\displaystyle\\pi\(y\\mid x\)=exp⁡\(zy​\(x\)\)∑y′exp⁡\(zy′​\(x\)\),\\displaystyle=\\frac\{\\exp\\big\(z\_\{y\}\(x\)\\big\)\}\{\\sum\_\{y^\{\\prime\}\}\\exp\\big\(z\_\{y^\{\\prime\}\}\(x\)\\big\)\},\(17\)zy​\(x\)\\displaystyle z\_\{y\}\(x\)=⟨𝐡⁡\(x\),𝐞y⟩,\\displaystyle=\\langle\\mathbf\{h\}\(x\),\\mathbf\{e\}\_\{y\}\\rangle,where𝐡⁡\(x\)∈ℝd\\mathbf\{h\}\(x\)\\in\\mathbb\{R\}^\{d\}is the model’s final hidden state for promptxx, and𝐞y\\mathbf\{e\}\_\{y\}is the output \(unembedding\) vector for tokenyy, so that the logitzy​\(x\)z\_\{y\}\(x\)is their inner product\.

##### Log\-probability of a single completion\.

Taking the logarithm of Equation[17](https://arxiv.org/html/2609.38718#A2.E17),

log⁡π⁡\(y∣x\)\\displaystyle\\log\\pi\(y\\mid x\)=zy\(x\)−log∑y′exp\(zy′\(x\)\)\\displaystyle=z\_\{y\}\(x\)\-\\log\\sum\_\{y^\{\\prime\}\}\\exp\\big\(z\_\{y^\{\\prime\}\}\(x\)\\big\)\(18\)=⟨𝐡⁡\(x\),𝐞y⟩−log⁡Z⁡\(x\),\\displaystyle=\\langle\\mathbf\{h\}\(x\),\\mathbf\{e\}\_\{y\}\\rangle\-\\log Z\(x\),whereZ⁡\(x\)=∑y′exp⁡⟨𝐡⁡\(x\),𝐞y′⟩Z\(x\)=\\sum\_\{y^\{\\prime\}\}\\exp\\langle\\mathbf\{h\}\(x\),\\mathbf\{e\}\_\{y^\{\\prime\}\}\\rangleis the partition function\. Crucially,Z⁡\(x\)Z\(x\)depends only onxx\(through𝐡⁡\(x\)\\mathbf\{h\}\(x\)\) and on the full vocabulary,*not*on the particular tokenyywhose probability is being evaluated\.

##### Cancellation of the partition function\.

Consider two completions,y\+y^\{\+\}andy−y^\{\-\}, evaluated at the same promptxx\(and hence the same hidden state𝐡⁡\(x\)\\mathbf\{h\}\(x\)and the sameZ⁡\(x\)Z\(x\)\)\. Applying Equation[18](https://arxiv.org/html/2609.38718#A2.E18)to each and subtracting,

log⁡π⁡\(y\+∣x\)−log⁡π⁡\(y−∣x\)=\[⟨𝐡⁡\(x\),𝐞y\+⟩−log⁡Z⁡\(x\)\]−\[⟨𝐡⁡\(x\),𝐞y−⟩−log⁡Z⁡\(x\)\]=⟨𝐡⁡\(x\),𝐞y\+⟩−⟨𝐡⁡\(x\),𝐞y−⟩,\\begin\{gathered\}\\log\\pi\(y^\{\+\}\\mid x\)\-\\log\\pi\(y^\{\-\}\\mid x\)\\\\ =\\Big\[\\langle\\mathbf\{h\}\(x\),\\mathbf\{e\}\_\{y^\{\+\}\}\\rangle\-\\log Z\(x\)\\Big\]\-\\Big\[\\langle\\mathbf\{h\}\(x\),\\mathbf\{e\}\_\{y^\{\-\}\}\\rangle\-\\log Z\(x\)\\Big\]\\\\ =\\langle\\mathbf\{h\}\(x\),\\mathbf\{e\}\_\{y^\{\+\}\}\\rangle\-\\langle\\mathbf\{h\}\(x\),\\mathbf\{e\}\_\{y^\{\-\}\}\\rangle,\\end\{gathered\}\(19\)sincelog⁡Z⁡\(x\)\\log Z\(x\)is identical in both terms and cancels exactly\. This cancellation is the entire content of the result: it holds regardless of vocabulary size, temperature, or the specific values of𝐡⁡\(x\)\\mathbf\{h\}\(x\), because it follows purely from the algebraic structure of the softmax, not from any property of the trained model\.

##### Linearity in the hidden state\.

By bilinearity of the inner product, Equation[19](https://arxiv.org/html/2609.38718#A2.E19)simplifies to

log⁡π⁡\(y\+∣x\)−log⁡π⁡\(y−∣x\)=⟨𝐡⁡\(x\),𝐞y\+−𝐞y−⟩=⟨𝐡\(x\),𝐯⟩,𝐯:=𝐞y\+−𝐞y−,\\begin\{gathered\}\\log\\pi\(y^\{\+\}\\mid x\)\-\\log\\pi\(y^\{\-\}\\mid x\)\\\\ =\\big\\langle\\mathbf\{h\}\(x\),\\ \\mathbf\{e\}\_\{y^\{\+\}\}\-\\mathbf\{e\}\_\{y^\{\-\}\}\\big\\rangle\\\\ =\\langle\\mathbf\{h\}\(x\),\\mathbf\{v\}\\rangle,\\qquad\\mathbf\{v\}:=\\mathbf\{e\}\_\{y^\{\+\}\}\-\\mathbf\{e\}\_\{y^\{\-\}\},\\end\{gathered\}\(20\)which is exactly Equation[1](https://arxiv.org/html/2609.38718#S3.E1)\. The gap is therefore an inner product between the hidden state and a single, context\-independent vector𝐯\\mathbf\{v\}determined only by the output embeddings of the two completions being compared\.

##### Consequence for the DPO gradient\.

Since the DPO loss is a function ofρθ​\(y\+∣x\)−ρθ​\(y−∣x\)\\rho\_\{\\theta\}\(y^\{\+\}\\mid x\)\-\\rho\_\{\\theta\}\(y^\{\-\}\\mid x\), and by Equation[20](https://arxiv.org/html/2609.38718#A2.E20)this quantity is linear in𝐡⁡\(x\)\\mathbf\{h\}\(x\)with fixed slope𝐯\\mathbf\{v\}, its gradient with respect to the hidden state is

∇𝐡⁡\(x\)ℒDPO=−β​σ​\(−β⁡⟨𝐡⁡\(x\),𝐯⟩\+const\)​𝐯∝−𝐯,\\nabla\_\{\\mathbf\{h\}\(x\)\}\\mathcal\{L\}\_\{\\text\{DPO\}\}=\-\\beta\\,\\sigma\\\!\\big\(\-\\beta\\langle\\mathbf\{h\}\(x\),\\mathbf\{v\}\\rangle\+\\text\{const\}\\big\)\\,\\mathbf\{v\}\\ \\propto\\ \-\\mathbf\{v\},\(21\)where the scalar prefactor depends onxxbut the direction does not: for every promptxx, the gradient points along the same vector𝐯\\mathbf\{v\}, up to sign and magnitude\. This is the source of the rank\-one/single\-direction characterization of single\-concept DPO steering discussed in Section[3\.1](https://arxiv.org/html/2609.38718#S3.SS1)\([Raina et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib1)\)\.

## Appendix CFull Expansion of the Query\-Key Perturbation

This appendix expands Equation[5](https://arxiv.org/html/2609.38718#S3.E5)and derives the partial derivatives in Equation[6](https://arxiv.org/html/2609.38718#S3.E6)in full, showing that the query\-key perturbationδi​j\\delta\_\{ij\}reduces to a single bilinear form in𝐳i\\mathbf\{z\}\_\{i\}and𝐳j\\mathbf\{z\}\_\{j\}, and that its dependence on each is generically non\-zero\.

##### Setup\.

Recall𝐪i=𝐳i​𝐖q\\mathbf\{q\}\_\{i\}=\\mathbf\{z\}\_\{i\}\\mathbf\{W\}\_\{q\},𝐤j=𝐳j​𝐖k\\mathbf\{k\}\_\{j\}=\\mathbf\{z\}\_\{j\}\\mathbf\{W\}\_\{k\},Δ​𝐪i=λq​𝐳i​𝐀q​𝐁q⊤\\Delta\\mathbf\{q\}\_\{i\}=\\lambda\_\{q\}\\mathbf\{z\}\_\{i\}\\mathbf\{A\}\_\{q\}\\mathbf\{B\}\_\{q\}^\{\\top\},Δ​𝐤j=λk​𝐳j​𝐀k​𝐁k⊤\\Delta\\mathbf\{k\}\_\{j\}=\\lambda\_\{k\}\\mathbf\{z\}\_\{j\}\\mathbf\{A\}\_\{k\}\\mathbf\{B\}\_\{k\}^\{\\top\}, with𝐳i,𝐳j\\mathbf\{z\}\_\{i\},\\mathbf\{z\}\_\{j\}treated as row vectors and𝐖q,𝐖k,𝐀f​𝐁f⊤∈ℝd×d\\mathbf\{W\}\_\{q\},\\mathbf\{W\}\_\{k\},\\mathbf\{A\}\_\{f\}\\mathbf\{B\}\_\{f\}^\{\\top\}\\in\\mathbb\{R\}^\{d\\times d\}\(we taked=d′d=d^\{\\prime\}for notational simplicity; the argument is unchanged ford≠d′d\\neq d^\{\\prime\}\)\. The three terms of Equation[5](https://arxiv.org/html/2609.38718#S3.E5)are, expanding each inner product as a matrix product with a transpose,

⟨𝐪i,Δ​𝐤j⟩\\displaystyle\\langle\\mathbf\{q\}\_\{i\},\\Delta\\mathbf\{k\}\_\{j\}\\rangle=λk​𝐳i​𝐖q​\(𝐳j​𝐀k​𝐁k⊤\)⊤\\displaystyle=\\lambda\_\{k\}\\,\\mathbf\{z\}\_\{i\}\\mathbf\{W\}\_\{q\}\\big\(\\mathbf\{z\}\_\{j\}\\mathbf\{A\}\_\{k\}\\mathbf\{B\}\_\{k\}^\{\\top\}\\big\)^\{\\top\}\(22\)=λk​𝐳i​\(𝐖q​𝐁k​𝐀k⊤\)​𝐳j⊤\\displaystyle=\\lambda\_\{k\}\\,\\mathbf\{z\}\_\{i\}\\big\(\\mathbf\{W\}\_\{q\}\\mathbf\{B\}\_\{k\}\\mathbf\{A\}\_\{k\}^\{\\top\}\\big\)\\mathbf\{z\}\_\{j\}^\{\\top\}
⟨Δ​𝐪i,𝐤j⟩\\displaystyle\\langle\\Delta\\mathbf\{q\}\_\{i\},\\mathbf\{k\}\_\{j\}\\rangle=λq​𝐳i​𝐀q​𝐁q⊤​\(𝐳j​𝐖k\)⊤\\displaystyle=\\lambda\_\{q\}\\,\\mathbf\{z\}\_\{i\}\\mathbf\{A\}\_\{q\}\\mathbf\{B\}\_\{q\}^\{\\top\}\\big\(\\mathbf\{z\}\_\{j\}\\mathbf\{W\}\_\{k\}\\big\)^\{\\top\}\(23\)=λq​𝐳i​\(𝐀q​𝐁q⊤​𝐖k⊤\)​𝐳j⊤\\displaystyle=\\lambda\_\{q\}\\,\\mathbf\{z\}\_\{i\}\\big\(\\mathbf\{A\}\_\{q\}\\mathbf\{B\}\_\{q\}^\{\\top\}\\mathbf\{W\}\_\{k\}^\{\\top\}\\big\)\\mathbf\{z\}\_\{j\}^\{\\top\}
⟨Δ​𝐪i,Δ​𝐤j⟩\\displaystyle\\langle\\Delta\\mathbf\{q\}\_\{i\},\\Delta\\mathbf\{k\}\_\{j\}\\rangle=λq​λk​𝐳i​𝐀q​𝐁q⊤​\(𝐳j​𝐀k​𝐁k⊤\)⊤\\displaystyle=\\lambda\_\{q\}\\lambda\_\{k\}\\,\\mathbf\{z\}\_\{i\}\\mathbf\{A\}\_\{q\}\\mathbf\{B\}\_\{q\}^\{\\top\}\\big\(\\mathbf\{z\}\_\{j\}\\mathbf\{A\}\_\{k\}\\mathbf\{B\}\_\{k\}^\{\\top\}\\big\)^\{\\top\}\(24\)=λq​λk​𝐳i​\(𝐀q​𝐁q⊤​𝐁k​𝐀k⊤\)​𝐳j⊤\\displaystyle=\\lambda\_\{q\}\\lambda\_\{k\}\\,\\mathbf\{z\}\_\{i\}\\big\(\\mathbf\{A\}\_\{q\}\\mathbf\{B\}\_\{q\}^\{\\top\}\\mathbf\{B\}\_\{k\}\\mathbf\{A\}\_\{k\}^\{\\top\}\\big\)\\mathbf\{z\}\_\{j\}^\{\\top\}Each term is a scalar of the identical form𝐳i​\(⋅\)​𝐳j⊤\\mathbf\{z\}\_\{i\}\(\\cdot\)\\mathbf\{z\}\_\{j\}^\{\\top\}, differing only in thed×dd\\times dmatrix sandwiched between𝐳i\\mathbf\{z\}\_\{i\}and𝐳j⊤\\mathbf\{z\}\_\{j\}^\{\\top\}\.

##### Collapse to a single bilinear form\.

Summing Equations[22](https://arxiv.org/html/2609.38718#A3.E22)–[24](https://arxiv.org/html/2609.38718#A3.E24)per Equation[5](https://arxiv.org/html/2609.38718#S3.E5)and factoring out the common𝐳i​\(⋅\)​𝐳j⊤\\mathbf\{z\}\_\{i\}\(\\cdot\)\\mathbf\{z\}\_\{j\}^\{\\top\}structure,

δi​j=1d​𝐳i​𝐌​𝐳j⊤,𝐌:=λk​𝐖q​𝐁k​𝐀k⊤\+λq​𝐀q​𝐁q⊤​𝐖k⊤\+λq​λk​𝐀q​𝐁q⊤​𝐁k​𝐀k⊤,\\begin\{gathered\}\\delta\_\{ij\}=\\frac\{1\}\{\\sqrt\{d\}\}\\,\\mathbf\{z\}\_\{i\}\\,\\mathbf\{M\}\\,\\mathbf\{z\}\_\{j\}^\{\\top\},\\\\ \\mathbf\{M\}:=\\lambda\_\{k\}\\mathbf\{W\}\_\{q\}\\mathbf\{B\}\_\{k\}\\mathbf\{A\}\_\{k\}^\{\\top\}\+\\lambda\_\{q\}\\mathbf\{A\}\_\{q\}\\mathbf\{B\}\_\{q\}^\{\\top\}\\mathbf\{W\}\_\{k\}^\{\\top\}\+\\lambda\_\{q\}\\lambda\_\{k\}\\mathbf\{A\}\_\{q\}\\mathbf\{B\}\_\{q\}^\{\\top\}\\mathbf\{B\}\_\{k\}\\mathbf\{A\}\_\{k\}^\{\\top\},\\end\{gathered\}\(25\)
where𝐌∈ℝd×d\\mathbf\{M\}\\in\\mathbb\{R\}^\{d\\times d\}depends only on the frozen weights and the trained adapters\{𝐀q,𝐁q,𝐀k,𝐁k\}\\\{\\mathbf\{A\}\_\{q\},\\mathbf\{B\}\_\{q\},\\mathbf\{A\}\_\{k\},\\mathbf\{B\}\_\{k\}\\\}and intensitiesλq,λk\\lambda\_\{q\},\\lambda\_\{k\}— it does*not*depend on the token positionsi,ji,j\. The entire query\-key perturbation, despite being built from three separate inner products, is therefore exactly a single bilinear form in the two token activations𝐳i\\mathbf\{z\}\_\{i\}and𝐳j\\mathbf\{z\}\_\{j\}, mediated by the fixed matrix𝐌\\mathbf\{M\}\.

##### Partial derivatives\.

Equation[25](https://arxiv.org/html/2609.38718#A3.E25)is linear in each argument holding the other fixed, so its gradients follow directly from the bilinear\-form identity∂\(𝐳i​𝐌𝐳j⊤\)/∂𝐳i=𝐌𝐳j⊤\\partial\(\\mathbf\{z\}\_\{i\}\\mathbf\{M\}\\mathbf\{z\}\_\{j\}^\{\\top\}\)/\\partial\\mathbf\{z\}\_\{i\}=\\mathbf\{M\}\\mathbf\{z\}\_\{j\}^\{\\top\}and∂\(𝐳i​𝐌𝐳j⊤\)/∂𝐳j=𝐌⊤​𝐳i⊤\\partial\(\\mathbf\{z\}\_\{i\}\\mathbf\{M\}\\mathbf\{z\}\_\{j\}^\{\\top\}\)/\\partial\\mathbf\{z\}\_\{j\}=\\mathbf\{M\}^\{\\top\}\\mathbf\{z\}\_\{i\}^\{\\top\}:

∂δi​j∂𝐳i=1d​𝐌​𝐳j⊤,∂δi​j∂𝐳j=1d​𝐌⊤​𝐳i⊤\.\\frac\{\\partial\\delta\_\{ij\}\}\{\\partial\\mathbf\{z\}\_\{i\}\}=\\frac\{1\}\{\\sqrt\{d\}\}\\,\\mathbf\{M\}\\,\\mathbf\{z\}\_\{j\}^\{\\top\},\\qquad\\frac\{\\partial\\delta\_\{ij\}\}\{\\partial\\mathbf\{z\}\_\{j\}\}=\\frac\{1\}\{\\sqrt\{d\}\}\\,\\mathbf\{M\}^\{\\top\}\\mathbf\{z\}\_\{i\}^\{\\top\}\.\(26\)

##### Genericity\.

Equation[26](https://arxiv.org/html/2609.38718#A3.E26)vanishes only if𝐌𝐳j⊤=𝟎\\mathbf\{M\}\\mathbf\{z\}\_\{j\}^\{\\top\}=\\mathbf\{0\}\(respectively𝐌⊤​𝐳i⊤=𝟎\\mathbf\{M\}^\{\\top\}\\mathbf\{z\}\_\{i\}^\{\\top\}=\\mathbf\{0\}\), which requires either \(a\)𝐌=𝟎\\mathbf\{M\}=\\mathbf\{0\}— an exact, measure\-zero cancellation of the three adapter terms in Equation[25](https://arxiv.org/html/2609.38718#A3.E25)that is not enforced by, and would not be a stable fixed point of, gradient\-based training on a non\-trivial preference\-optimization objective — or \(b\)𝐳j\\mathbf\{z\}\_\{j\}\(resp\.𝐳i\\mathbf\{z\}\_\{i\}\) lying in the null space of𝐌\\mathbf\{M\}\(resp\.𝐌⊤\\mathbf\{M\}^\{\\top\}\), a codimension\-≥1\\geq 1subset of activation space that generic hidden states do not occupy\. Excluding these non\-generic cases, both partial derivatives in Equation[26](https://arxiv.org/html/2609.38718#A3.E26)are non\-zero, establishing Equation[6](https://arxiv.org/html/2609.38718#S3.E6)\.

##### Interpretation\.

Equation[26](https://arxiv.org/html/2609.38718#A3.E26)makes explicit*why*the intervention is context\-dependent rather than a fixed offset: the sensitivity ofδi​j\\delta\_\{ij\}to the query token𝐳i\\mathbf\{z\}\_\{i\}is not a constant vector, but𝐌𝐳j⊤\\mathbf\{M\}\\mathbf\{z\}\_\{j\}^\{\\top\}— a vector that itself changes with the key token𝐳j\\mathbf\{z\}\_\{j\}it is being compared against, and hence with whatever content \(including the concept specificationcc\) that key token encodes\. This is the precise sense in which the effective steering direction is a function of the surrounding sequence rather than a parameter fixed at the end of training: the same query activation𝐳i\\mathbf\{z\}\_\{i\}receives a different perturbation gradient depending on which key𝐳j\\mathbf\{z\}\_\{j\}it attends to, mediated entirely through the single learned matrix𝐌\\mathbf\{M\}\.

## Appendix DExtended Inter\-Rater Reliability Analysis

This appendix extends the reliability analysis summarized in the main text\. A separate, human\-anchored validation of judge scores is reported in Appendix[E\.1](https://arxiv.org/html/2609.38718#A5.SS1)\.

##### Setup\.

We evaluate two tasks: the steering\-evaluation protocol of[Wu et al\. \(2025\)](https://arxiv.org/html/2609.38718#bib.bib2)\(steering\-score and language\-quality metrics\) and the open\-ended generation protocol of[Soo et al\. \(2025\)](https://arxiv.org/html/2609.38718#bib.bib33)\. Each is scored independently by three judges – GPT\-4o\-mini, Gemini\-3\.1\-Flash\-Lite, and Claude Haiku 4\.5– on steering success and output quality\.

##### Overall and per\-method agreement\.

Pooled across all method×\\timesitem units, judges agree strongly \(Krippendorff’sα=0\.81\\alpha=0\.81; ICC\(3,1\)=0\.81=0\.81\)\. Table[4](https://arxiv.org/html/2609.38718#A4.T4)breaks this down by method on Gemma\-2\-2B \(layer 10\):LoRAandLoReFTshow the strongest, most consistent agreement;PromptSteering,LAT, andDiffMeanare intermediate;PCAandDPOare lowest333DPO outputs may convey concepts more subtly, making binary scoring harder\.\.

Table 4:Inter\-judge agreement across steering methods for the steering evaluation task on Gemma\-2\-2B \(layer 10\)\.![Refer to caption](https://arxiv.org/html/2609.38718v1/fig02_pairwise_agreement.png)Figure 3:Macro\-averaged pairwise judge agreement, steering evaluation task\.
##### Pairwise agreement and calibration\.

Weighted Cohen’sκ\\kappaand Spearman correlations \(Fig\.[3](https://arxiv.org/html/2609.38718#A4.F3)\) show moderate\-to\-good pairwise agreement: judges are consistent in relative ordering but not fully interchangeable at fine granularity, particularly forDPO\. This pattern holds on the open\-ended task as well\. Judges also differ systematically in scale \(Fig\.[4](https://arxiv.org/html/2609.38718#A4.F4)\): one is consistently more lenient \(e\.g\., onDPOandPromptSteering\), another more conservative, confirmed by a mixed\-effects model with significant fixed effects\. These differences are largely scale effects rather than disagreement in ordering – a variance decomposition on the open\-ended task attributes most variance to concept\-level differences, with smaller contributions from model and judge, and method rankings remain stable across judges on both tasks\.

##### Summary\.

Judges disagree modestly in absolute calibration but strongly agree in relative ordering, supporting LLM\-as\-a\-Judge as a reliable evaluation strategy in our setting\.

Figure 4:Judge score distributions, steering evaluation task\.

## Appendix EHuman Validation of LLM Judges

The increasing capability of frontier language models, together with the high cost of human annotation, has motivated using LLMs as proxies for human evaluators\([Zheng et al\., 2023](https://arxiv.org/html/2609.38718#bib.bib4);[Zhang et al\., 2023](https://arxiv.org/html/2609.38718#bib.bib50);[Gu et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib57);[Gu et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib66);[Yu, 2025](https://arxiv.org/html/2609.38718#bib.bib67)\)\. While prior work reports substantial LLM\-human agreement\([Zheng et al\., 2023](https://arxiv.org/html/2609.38718#bib.bib4)\), concerns remain about reliability, fairness, reproducibility, and systematic bias\([Wang et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib3);[Gu et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib57);[Gu et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib66);[Yu, 2025](https://arxiv.org/html/2609.38718#bib.bib67)\), particularly in model\-steering research, where a single frontier model often both generates and evaluates outputs\([Wu et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib2);[Soo et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib33);[Arad et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib32);[Turner et al\., 2023](https://arxiv.org/html/2609.38718#bib.bib35)\)\. We therefore cross\-check our own evaluation setup against human judgment below, and follow[Wu et al\. \(2023\)](https://arxiv.org/html/2609.38718#bib.bib22)in assigning the LLM judge an explicit evaluator role via prompting\.

### E\.1Quantitative Analysis

We sampled 50 generated outputs \(25 from[Wu et al\. \(2025\)](https://arxiv.org/html/2609.38718#bib.bib2), 25 from[Soo et al\. \(2025\)](https://arxiv.org/html/2609.38718#bib.bib33)\) and scored each, under an identical prompt, with both GPT\-4o\-mini and a human annotator \(a postgraduate student and second\-language English speaker\)\. Agreement between the two sets of ratings was moderately strong, with ICC\(3,1\)≈0\.70\\approx 0\.70, supporting LLM\-as\-a\-Judge as a reasonable proxy for human evaluation in our setting\. Representative side\-by\-side judgments are given in Appendix[E\.2](https://arxiv.org/html/2609.38718#A5.SS2)\.

### E\.2Qualitative Examples

We show a sample of human and LLM judgments for two identical inputs below\.

AxBench SampleSteering Concept:Scientific terms related to research findings and implications Instruction:What are different drawers I should have for clothes?Prompt Steering OutputWhile the question of optimal drawer organization isn’t directly related to scientific research, let’s view it through a lens of anthropological insights into human behavior and organizational psychology\. Understanding why we organize our belongings is fascinating\! One might draw on heuristic decision\-making models\. Our brains are naturally efficient and seek shortcuts to navigate everyday activities, and choosing a drawer organization system is an example of this\. A structured system like this can make accessing and putting away clothes more efficient – reducing cognitive load and increasing self\-efficacy\.•Fluency: LLM = 1, Human =2•Concept Relevance: LLM = 1, Human =1•Instruction Relevance: LLM = 1, Human =1DPO OutputThe scientific literature on organizational behaviors suggests that the strategic placement of clothing compartments, such as drawers, can significantly impact efficiency within an individual’s wardrobe system\. Recent studies indicate a correlation between drawer design and optimal apparel access, facilitating streamlined clothing selection based on contextual factors like occasion and activity\. This finding emphasizes the necessity for tailored drawer configurations that align with personal preferences, offering a personalized approach to optimizing storage and retrieval, ultimately enhancing overall time management in the context of daily dressing routines\. …•Fluency: LLM = 2, Human =2•Concept Relevance: LLM = 2, Human =2•Instruction Relevance: LLM = 2, Human =1

Free\-Text Generation Sample \(Base Model,[Soo et al\. \(2025\)](https://arxiv.org/html/2609.38718#bib.bib33)\)Steering Concept:Anger Prompt:I think …Prompt Steering Output…I was just so mad at this restaurant for going overboard on the lack of customer service earlier this week, so after getting out of therapy I lost my, …•Steering Score: LLM = 0\.67, Human =0\.75•Coherence: LLM = 0\.67, Human =0\.60•Product: LLM = 0\.45, Human =0\.45DPO Output…wrath is the driving force behind my actions, fueled by resentment and frustration\. Like a roaring inferno, it ignites my passion, shaping …•Steering Score: LLM = 0\.78, Human =0\.85•Coherence: LLM = 0\.78, Human =0\.70•Product: LLM = 0\.61, Human =0\.60

These examples illustrate that LLM\-as\-a\-Judge can serve as a reliable proxy for human evaluation in the text\-generation setting\.

## Appendix FPerformance Degradation Analysis

One of the main limitations of steering a large language model toward a particular behaviour is an unwanted side effect on the model’s general capabilities\. These side effects can range from degraded inherent safety guardrails[Bao et al\. \(2026\)](https://arxiv.org/html/2609.38718#bib.bib68)to drops of22–4%4\\%on general\-capability benchmarks\([Ostermann et al\., 2026](https://arxiv.org/html/2609.38718#bib.bib37);[Rimsky et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib36)\)\. We therefore run a side\-check experiment to verify that MetaSteer does not induce significant performance degradation\.

##### Testing scenario\.

For each backbone we compare two variants under an identical evaluation protocol: the original HuggingFace model \(*Base*\) and the same model with the trained MetaSteer LoRA adapter applied \(*MetaSteer*\)\. Evaluations use the EleutherAIlm\-evaluation\-harnesswith multiple\-choice log\-likelihood scoring \(not generative decoding\)\. We report accuracy on three held\-out capability / truthfulness suites, using the full official test sets:

- •MMLU: broad knowledge and reasoning;
- •TruthfulQA MC1: single\-true\-answer truthfulness;
- •TruthfulQA MC2: multi\-true\-answer truthfulness\.

Each task is evaluated atk∈\{0,2,4\}k\\in\\\{0,2,4\\\}few\-shot examples\. Chat templates are disabled, and Qwen3 thinking mode is turned off, so scores reflect standard multiple\-choice log\-likelihood rather than chain\-of\-thought generation\. The reported delta is

Δ=Acc⁡\(MetaSteer\)−Acc⁡\(Base\),\\Delta\\;=\\;\\mathrm\{Acc\}\(\\text\{MetaSteer\}\)\-\\mathrm\{Acc\}\(\\text\{Base\}\),so negative values indicate degradation\.

##### Results\.

Table[5](https://arxiv.org/html/2609.38718#A6.T5)summarises Base vs\. MetaSteer accuracy \(%\)\. Across the evaluated models, MMLU changes are generally small\. The largest TruthfulQA drops occur for Qwen3\-0\.6B, reaching 1\.51 percentage points; among the six primary backbones, the largest drop is 0\.37 points, while Gemma\-2\-2B and Qwen3\-4B improve on some TruthfulQA settings\.

Table 5:Capability retention after MetaSteer\. Accuracy \(%\) of the unsteered base model vs\. the MetaSteer LoRA on MMLU and TruthfulQA \(MC1/MC2\) atk=0,2,4k=0,2,4shots\.Δ\\Deltais MetaSteer−\-Base in percentage points \(negative==degradation\)\.

## Appendix GSteering as a Geometric Transformation

The analysis presented in Appendix[I](https://arxiv.org/html/2609.38718#A9)indicates that steering has a noisy influence on the generated text at the local level, particularly at the word level\. However, when the generated text is represented at a coarser granularity—for example, at the sentence level using pooling—the overall structure of the response remains comparatively stable\.

Figure[1](https://arxiv.org/html/2609.38718#S1.F1)\(a\)–\(b\)visualizes trajectories obtained from responses generated under different steering directions\. Each trajectory is constructed from the first six points extracted from a generated response\. To make the relationship between the trajectories and the generated language explicit, the instruction, steering direction, and complete recorded response are provided below for each example\.

The examples show that steering changes the wording, tone, formatting, and organizational structure of the generated text\. Nevertheless, the responses generally preserve the broader semantic objective specified by the instruction\. Thus, steering appears to influence the local realization of the response while leaving several global properties of the trajectory—including its general curvature, velocity profile, and semantic direction—comparatively stable\.

The following examples reproduce the recorded generated responses in full\. Some supplied responses terminate mid\-sentence; these endings are retained as recorded\.

### G\.1Instruction: “How to Lose Weight?”

#### G\.1\.1Generated Response 1

Instruction:How to lose weight?

Steering direction:Particular patterns related to connectivity and relationships in a structured format\(\#3\)\.

Final generated response:

Establish clear goals within a network of support, such as friends or family engaged in similar health journeys\.Create routines that interconnect meal planning and physical activity, ensuring a flow that promotes accountability\.Monitor progress while maintaining relationships with health professionals, integrating their insights into a structure that is adaptable\.Leverage community resources, joining group fitness or dieting programs to foster connections that reinforce commitment\.Reflect on emotional patterns surrounding food choices, identifying connections between feelings, habits, and health objectives to achieve a balanced approach\.

#### G\.1\.2Generated Response 2

Instruction:How to lose weight?

Steering direction:Entities, particularly names and classifications, as well as structured data elements\(\#4\)\.

Final generated response:

Nutritional categories: carbohydrates, proteins, fats\.Activity types: aerobic, anaerobic, flexibility\.Weight\-loss entities: caloric deficit, metabolism, body composition\.Structured data: weekly meal plan, daily exercise log, progress tracker\.Classification examples\.

#### G\.1\.3Generated Response 3

Instruction:How to lose weight?

Steering direction:Non\-textual elements or formatting indicators present in a document\(\#5\)\.

Final generated response:

How to Lose Weight: A Visual GuideImages:Incorporate before\-and\-after photos to illustrate progress\.Charts:Use bar graphs to show calorie intake versus calorie expenditure over time\.Icons:Utilize food icons to represent healthy meal choices such as fruits and vegetables\.Tables:Create a meal\-planning table that outlines daily calorie goals and meal options\.Color Coding:Apply different colors to categorize exercises, for example, cardio in blue and strength training in green\.\.

#### G\.1\.4Generated Response 4

Instruction:How to lose weight?

Steering direction:Elements related to lists and structured requests\(\#6\)\.

Final generated response:

Set clear goals\.Track food intake\.Incorporate regular exercise\.Prioritize hydration\.Focus on whole foods\.Avoid sugary drinks\.Plan meals in advance\.Monitor portion sizes\.Get enough sleep\.Stay consistent with habits\.

#### G\.1\.5Generated Response 5

Instruction:How to lose weight?

Steering direction:Procedural instructions related to software or game installation\(\#7\)\.

Final generated response:

Open the application interface for your weight\-management program\. Ensure that all prerequisites are met, including dietary preferences and exercise restrictions\.Download the meal\-plan module\. Once the download is complete, execute the program to initiate the dietary\-customization process\.Navigate to the settings panel and adjust the caloric\-intake parameters\. Confirm your selections to ensure that the new limits are activated\.Install the physical\-activity tracker add\-on\. After installation, link it with your device for real\-time updates on your exercises\.Monitor progress through the analytics dashboard\. Regularly check the updates and evaluate your metrics\.

### G\.2Instruction: “What Are Some Tools That Every Homeowner Should Have?”

Instruction:What are some tools that every homeowner should have?

#### G\.2\.1Generated Response 1

Steering direction:Specific structural markers or formatting indicators in a document\.

Final generated response:

Hammer:Essential for driving nails into or pulling them from materials, often marked by its head size and shape\.Screwdriver Set:Featuring various tips, such as flat and Phillips, categorized for different screw types and crucial for assembly and repairs\.Tape Measure:Typically sectioned in inches and centimeters, it provides accuracy in measuring spaces and objects and is indicated by clear markings\.Level:Equipped with horizontal and vertical indicators, it ensures that surfaces are even and properly aligned, making it vital for installations\.Pliers:Differentiated by type \(needle\-nose, slip\-joint\),\.

#### G\.2\.2Generated Response 2

Instruction:What are some tools that every homeowner should have?

Steering direction:Keywords and structural elements relevant to programming or technical documentation\.

Final generated response:

Hammer:Essential for executing tasks that require impact and precision; serves as the primary execution framework for fastening components\.Screwdriver:Integral for accessing or manipulating closures; supports various modules via interchangeable bits for different operations\.Pliers:Provides gripping capabilities, enabling multi\-use functionality similar to error handling; essential for enhancing torque and manipulating non\-digital components\.Level:A measurement tool essential for ensuring alignment; offers validation of horizontal and vertical parameters in project specifications\.Tape Measure:Fundamental for quantifying distances; functions as a scoring mechanism to determine the size and\.

#### G\.2\.3Generated Response 3

Instruction:What are some tools that every homeowner should have?

Steering direction:Components related to job listings and their associated properties in a structured format\.

Final generated response:

Hammer:Essential for various household tasks, akin to a key skill in job listings\.Screwdriver set:A fundamental tool for assembly, similar to the necessary qualifications for a position\.Tape measure:Vital for accurate measurements, paralleling the precision required in job descriptions\.Level:Ensures that projects are even, much like maintaining balance in job responsibilities\.Pliers:Useful for gripping and twisting, reflecting the adaptability sought in candidates\.Utility knife:Versatile for cutting tasks, resembling multi\-talented applicants in a competitive market\.Wrench set:Critical for\.

#### G\.2\.4Generated Response 4

Instruction:What are some tools that every homeowner should have?

Steering direction:Terms related to articles and forms in a structured format\.

Final generated response:

Hammer:Essential tool, typically categorized under fasteners and attachments\.Screwdriver:Commonly found in home\-maintenance checklists and classified as a specific tool\.Wrench:Integral for plumbing tasks and noted within repair documentation\.Pliers:Frequently referenced in do\-it\-yourself guides and serving various functions in toolboxes\.Tape Measure:Critical for measurements and usually documented in project\-planning forms\.Level:Important for ensuring alignment and often included in construction articles\.Utility Knife:Listed in safety protocols and widely used for cutting tasks\.Drill:Noted in inventory lists,\.

Table 6:Inter\-rater agreement across six model\-based raters and 150 evaluation cases\.MetricValueNotesIntraclass correlationICC\(2,1\)0\.396Single\-rater absolute agreementICC\(2,k\)0\.798Average of six raters under absolute agreementICC\(3,1\)0\.408Single\-rater consistencyICC\(3,k\)0\.805Average of six raters under consistencyRank\-based agreementKendall’sWW0\.441χ2​\(149\)=394\.3\\chi^\{2\}\(149\)=394\.3,p=4\.6×10−24p=4\.6\\times 10^\{\-24\}Mean pairwise Spearmanρ\\rho0\.329Mean Spearman correlation implied byWW

## Appendix HConcept Inventory and Inter\-Rater Agreement

The experimental details of the concept\-level geometric analysis—including concept selection, construction of the CAA and MetaSteer directions, cosine similarity computation, and model evaluation—are provided in Appendix[J](https://arxiv.org/html/2609.38718#A10)\. This appendix documents the concept inventory and reports inter\-rater agreement for the associated evaluations\.

For each concept, six model\-based raters assessed the corresponding steering behavior\. The raters were drawn from the Llama, Gemma, and Qwen model families\. The resulting inter\-rater agreement statistics are reported in Table[6](https://arxiv.org/html/2609.38718#A7.T6)\.

The results indicate moderate single\-rater reliability, with an intraclass correlation of approximately0\.400\.40\. Agreement increases substantially when ratings are averaged across the six raters, reaching approximately0\.800\.80\. Kendall’sWWand the mean pairwise Spearman correlation likewise indicate moderate but statistically significant rank agreement\.

The complete concept inventory and cluster assignments are provided below\. Cluster names are kept outside the shaded lists, while the concepts are grouped within compact colored boxes\.

### H\.1Concept Inventory by Cluster

#### Response language

#### Format markers

#### Persona & tone

#### Content devices

#### Expert framing

#### Explanatory structure

#### List & decision format

#### Opening/closing conventions

#### Answer templates

#### Discourse connectors

#### Argumentative framing

## Appendix ICross\-Model Similarity of Steering Trajectories

We examine whether responses to the same instruction retain common geometric structure across steering concepts, model families, and encoding modes\. Following[Zhou et al\. \(2026\)](https://arxiv.org/html/2609.38718#bib.bib45), we compare position, velocity, acceleration, and Menger\-curvature representations of hidden\-state trajectories\. We additionally use sentence\-order shuffling as a negative control for ordered trajectory structure\.

##### Data and models\.

We combine Concept16K\-v1 and Concept16K\-v2[M](https://arxiv.org/html/2609.38718#A13)into a dataset of 131,363 records, covering 2,134 instructions, 10,488 steering concepts, and three response genres:code,text, andmath\. To reduce the influence of sparsely represented groups, we retain only records whose instruction, concept, and genre each occur more than 50 times\. This produces 2,550 instruction–concept pairs spanning 381 instructions, 70 concepts, and all three genres\.

We evaluate the six primary backbones and additionally include Qwen3\-0\.6B as a small\-model capability\-retention stress test\. Qwen3\-1\.7B, Qwen3\-4B, Gemma\-2\-2B, Gemma\-2\-9B, Llama\-3\.2\-3B, and Llama\-3\.1\-8B\. Responses are represented using the steps supplied by the dataset; responses stored as single strings are segmented at newlines or sentence\-ending punctuation\. The retained responses contain 29,484 segments in total\.

##### Cumulative and isolated representations\.

Letx1,…,xTx\_\{1\},\\ldots,x\_\{T\}be the ordered segments of a response and let

St=x1​‖⋯‖​xtS\_\{t\}=x\_\{1\}\\\|\\cdots\\\|x\_\{t\}denote the cumulative prefix\. We extract representations under two encoding modes,m∈\{cum,iso\}m\\in\\\{\\mathrm\{cum\},\\mathrm\{iso\}\\\}\.

In cumulative mode, segmentxtx\_\{t\}is encoded with all preceding segments:

𝐬t\(cum\)=1\|It\|​∑j∈It𝐡j\(L\)​\(St\),\\mathbf\{s\}^\{\(\\mathrm\{cum\}\)\}\_\{t\}=\\frac\{1\}\{\|I\_\{t\}\|\}\\sum\_\{j\\in I\_\{t\}\}\\mathbf\{h\}^\{\(L\)\}\_\{j\}\(S\_\{t\}\),\(27\)whereItI\_\{t\}contains the token positions introduced byxtx\_\{t\}\.

In isolated mode, each segment is encoded independently:

𝐬t\(iso\)=1\|It\|​∑j∈It𝐡j\(L\)​\(xt\)\.\\mathbf\{s\}^\{\(\\mathrm\{iso\}\)\}\_\{t\}=\\frac\{1\}\{\|I\_\{t\}\|\}\\sum\_\{j\\in I\_\{t\}\}\\mathbf\{h\}^\{\(L\)\}\_\{j\}\(x\_\{t\}\)\.\(28\)Thus, both modes use the same segments and pooling rule, but only cumulative encoding allows preceding segments to contextualize the current segment\.

For either mode, we define position, velocity, and acceleration vectors as

𝐩t\(m\)\\displaystyle\\mathbf\{p\}^\{\(m\)\}\_\{t\}=𝐬t\(m\),\\displaystyle=\\mathbf\{s\}^\{\(m\)\}\_\{t\},\(29\)𝐯t\(m\)\\displaystyle\\mathbf\{v\}^\{\(m\)\}\_\{t\}=𝐬t\+1\(m\)−𝐬t\(m\),\\displaystyle=\\mathbf\{s\}^\{\(m\)\}\_\{t\+1\}\-\\mathbf\{s\}^\{\(m\)\}\_\{t\},\(30\)𝐚t\(m\)\\displaystyle\\mathbf\{a\}^\{\(m\)\}\_\{t\}=𝐯t\+1\(m\)−𝐯t\(m\)\.\\displaystyle=\\mathbf\{v\}^\{\(m\)\}\_\{t\+1\}\-\\mathbf\{v\}^\{\(m\)\}\_\{t\}\.\(31\)The scalar quantity used as speed in the main text is

vt\(m\)=‖𝐯t\(m\)‖2\.v^\{\(m\)\}\_\{t\}=\\\|\\mathbf\{v\}^\{\(m\)\}\_\{t\}\\\|\_\{2\}\.\(32\)Our similarity analysis retains the complete vector𝐯t\(m\)\\mathbf\{v\}^\{\(m\)\}\_\{t\}, rather than only its magnitude\.

Menger curvature is computed from three consecutive states:

κt\(m\)=4​Area⁡\(𝐬t−1\(m\),𝐬t\(m\),𝐬t\+1\(m\)\)‖𝐬t−1\(m\)−𝐬t\(m\)‖2​‖𝐬t\(m\)−𝐬t\+1\(m\)‖2​‖𝐬t\+1\(m\)−𝐬t−1\(m\)‖2\.\\kappa^\{\(m\)\}\_\{t\}=\\frac\{4\\,\\operatorname\{Area\}\\left\(\\mathbf\{s\}^\{\(m\)\}\_\{t\-1\},\\mathbf\{s\}^\{\(m\)\}\_\{t\},\\mathbf\{s\}^\{\(m\)\}\_\{t\+1\}\\right\)\}\{\\\|\\mathbf\{s\}^\{\(m\)\}\_\{t\-1\}\-\\mathbf\{s\}^\{\(m\)\}\_\{t\}\\\|\_\{2\}\\\|\\mathbf\{s\}^\{\(m\)\}\_\{t\}\-\\mathbf\{s\}^\{\(m\)\}\_\{t\+1\}\\\|\_\{2\}\\\|\\mathbf\{s\}^\{\(m\)\}\_\{t\+1\}\-\\mathbf\{s\}^\{\(m\)\}\_\{t\-1\}\\\|\_\{2\}\}\.\(33\)

##### Trajectory similarity\.

Forr∈\{𝐩,𝐯,𝐚\}r\\in\\\{\\mathbf\{p\},\\mathbf\{v\},\\mathbf\{a\}\\\}, the similarity between trajectoriesiiandjjis the mean cosine similarity between temporally aligned vectors:

si​j\(r,m\)=1ℓi​j\(r,m\)​∑t=1ℓi​j\(r,m\)⟨𝐫i,t\(m\),𝐫j,t\(m\)⟩‖𝐫i,t\(m\)‖2​‖𝐫j,t\(m\)‖2,s\_\{ij\}^\{\(r,m\)\}=\\frac\{1\}\{\\ell\_\{ij\}^\{\(r,m\)\}\}\\sum\_\{t=1\}^\{\\ell\_\{ij\}^\{\(r,m\)\}\}\\frac\{\\left\\langle\\mathbf\{r\}^\{\(m\)\}\_\{i,t\},\\mathbf\{r\}^\{\(m\)\}\_\{j,t\}\\right\\rangle\}\{\\\|\\mathbf\{r\}^\{\(m\)\}\_\{i,t\}\\\|\_\{2\}\\\|\\mathbf\{r\}^\{\(m\)\}\_\{j,t\}\\\|\_\{2\}\},\(34\)whereℓi​j\(r,m\)\\ell\_\{ij\}^\{\(r,m\)\}is the shorter sequence length\. Menger\-curvature similarity is the Pearson correlation between aligned scalar sequences:

si​j\(κ,m\)=Corr\(κi,1:ℓ\(m\),κj,1:ℓ\(m\)\)\.s\_\{ij\}^\{\(\\kappa,m\)\}=\\operatorname\{Corr\}\\left\(\\kappa^\{\(m\)\}\_\{i,1:\\ell\},\\kappa^\{\(m\)\}\_\{j,1:\\ell\}\\right\)\.\(35\)
We average similarities within instruction, steering\-concept, and genre groups\. In Tables[7](https://arxiv.org/html/2609.38718#A9.T7)and[8](https://arxiv.org/html/2609.38718#A9.T8),I,C, andGdenote these respective grouping criteria\. Each entry reports macro/micro similarity\. Macro averaging weights groups equally, whereas micro averaging pools all within\-group pairs\.

##### Shuffled control\.

For every response, we independently permute its segment order using a fixed random seed of 42\. The same permutations are used across models\. In cumulative mode, the shuffled order changes both segment adjacency and the context preceding each segment\. In isolated mode, individual segment embeddings remain context\-free, but shuffling changes which segments are adjacent when velocity, acceleration, and curvature are calculated\.

Table 7:Cumulative\-context trajectory similarity\. Each model is evaluated using its original and shuffled segment order\. Entries report macro/micro averages\. Position, velocity, and acceleration use mean cosine similarity; Menger curvature uses Pearson correlation\.Table 8:Isolated\-segment trajectory similarity\. Each segment is encoded without preceding context\. Shuffling therefore changes segment adjacency but not the context used to embed an individual segment\. Entries report macro/micro averages\.
##### Results\.

As shown in Tables[7](https://arxiv.org/html/2609.38718#A9.T7)and[8](https://arxiv.org/html/2609.38718#A9.T8), position similarity changes only modestly from the shuffled to the original ordering and remains high under instruction, concept, and genre grouping\. Position similarity is therefore weakly discriminative of ordered or instruction\-specific structure; this does not imply that positional representations contain no semantic information, but rather that similarity at this level cannot reliably isolate it\. In contrast, instruction\-grouped velocity exhibits substantial increases over the shuffled baseline, from\.057/\.059\.057/\.059to\.305/\.381\.305/\.381under cumulative encoding and from\.001/\.001\.001/\.001to\.289/\.365\.289/\.365under isolated encoding\. Discrete acceleration shows the same pattern, increasing from\.024/\.024\.024/\.024to\.269/\.347\.269/\.347in cumulative mode and from\.002/\.002\.002/\.002to\.269/\.347\.269/\.347in isolated mode\. These results indicate that responses sharing an instruction possess meaningfully aligned velocity and acceleration structure, and that this alignment depends on the original segment ordering\. Menger\-curvature similarity likewise increases strongly under isolated encoding, from\.018/\.004\.018/\.004to\.289/\.373\.289/\.373\. The corresponding cumulative increase is weaker, from\.324/\.347\.324/\.347to\.361/\.439\.361/\.439, consistent with cumulative context inducing shared curvature structure even after the segment order is shuffled\.

##### Limitations\.

Models from the same family are related and should not be interpreted as independent population samples\. Instruction\-grouped averages also include all within\-instruction pairs and do not require every pair to differ in both concept and genre\. The shuffled control destroys response coherence and is therefore a negative control rather than a realistic steering condition\. Finally, Menger curvature is a scalar correlation\-based measure and is more sensitive to short trajectories and local irregularities than the vector\-based cosine similarities\.

## Appendix JConcept\-Level Geometric Analysis

For each of the 150 concepts of[Fan et al\. \(2026\)](https://arxiv.org/html/2609.38718#bib.bib65), we compare two steering directions in the last\-layer hidden state\. CAA supplies a linear contrastive direction, the difference of mean last\-prompt states on positive and negative examples,

𝐯c,mCAA=𝔼⁡\[hm​\(x\+\)\]−𝔼⁡\[hm​\(x−\)\]\.\\mathbf\{v\}^\{\\mathrm\{CAA\}\}\_\{c,m\}=\\mathbb\{E\}\[h\_\{m\}\(x^\{\+\}\)\]\-\\mathbb\{E\}\[h\_\{m\}\(x^\{\-\}\)\]\.MetaSteer supplies a more efficient intervention based on the results\. Its effect can be partly described as displacement of the prompt state when the concept is applied,

𝐝c,mMS=𝔼i​\[hm​\(i,c\)−hm​\(i\)\]\.\\mathbf\{d\}^\{\\mathrm\{MS\}\}\_\{c,m\}=\\mathbb\{E\}\_\{i\}\\bigl\[h\_\{m\}\(i,c\)\-h\_\{m\}\(i\)\\bigr\]\.The share of that MetaSteer displacement captured by the CAA direction is their cosine,

sc,m=cos⁡\(𝐯c,mCAA,𝐝c,mMS\)\.s\_\{c,m\}=\\cos\\bigl\(\\mathbf\{v\}^\{\\mathrm\{CAA\}\}\_\{c,m\},\\,\\mathbf\{d\}^\{\\mathrm\{MS\}\}\_\{c,m\}\\bigr\)\.
This cosine measures directional agreement between the CAA displacement and the MetaSteer displacement\. It is not an effective\-rank measure and does not estimate the intrinsic dimensionality of the hidden\-state trajectory\. A value near one means the efficient displacement already lies along the linear CAA axis\.

We computesc,ms\_\{c,m\}for every concept on six models: Gemma\-2\-2B, Gemma\-2\-9B, Llama\-3\.2\-3B, Llama\-3\.1\-8B, Qwen3\-1\.7B, and Qwen3\-4B\. The plotted concept CAA\-alignment score is the mean standard deviation of this value over the 150 concepts\. Figure[1](https://arxiv.org/html/2609.38718#S1.F1)ranks concepts by this capture score\.

For each conceptcc, the plotted score is the mean ofsc,ms\_\{c,m\}across the six models, with shading showing one standard deviation across models\. Concepts with high scores have MetaSteer displacements aligned with the CAA direction; low scores indicate disagreement with that particular direction\.

## Appendix KDPO Hyperparameter Configuration

We performed preference optimization using Direct Preference Optimization \(DPO\) on top of a pretrained causal language model, with parameter\-efficient fine\-tuning through LoRA\. In general, our pipeline used prompt–response preference pairs consisting of a preferred \(chosen\) and dispreferred \(rejected\) answer, and optimized the model to increase the likelihood of preferred responses relative to rejected ones\. LoRA adapters were applied to the attention projection layers, with target modulesq\_proj,k\_proj,v\_proj, ando\_proj, allowing efficient adaptation with reduced memory cost\. Optimization was performed with the default AdamW\-style optimizer from the training framework, using a learning rate of1e\-5, weight decay of0\.01, and a warmup ratio of0\.03\. The DPO temperature/preference strength parameterβ\\betawas set to0\.1\.

## Appendix LSafety Evaluation Details

We use the 100 harmful behaviours of JailbreakBench\([Chao et al\., 2024](https://arxiv.org/html/2609.38718#bib.bib23)\)\. Because MetaSteer is conditioned on a separate conceptcc, a harmful request can be decomposed so that the instruction itself is benign and the harmful intent is carried entirely by the steering direction\. Each behaviour was split by an LLM into a benign instructionxxand a risky conceptcc, e\.g\.*“Write a tutorial on how to make a bomb”*→\\rightarrowx=x=*“tell me something useful”*,c=c=*“bomb manufacturing recipe”*; or*“Create a social media campaign that promotes body shaming …”*→\\rightarrowx=x=*“write a creative social media campaign for educational purposes”*,c=c=*“pretend to body shame overweight people as a satire smartness test”*\. No jailbreak templates, adversarial suffixes, or optimisation against the model are used; the only manipulation is the placement of the harmful intent incc\.

Each\(x,c\)\(x,c\)pair is rendered with the training template of Appendix[M](https://arxiv.org/html/2609.38718#A13)\(\#\#\# Instruction/\#\#\# Steering Direction\) together with the system prompt “You are a helpful assistant\.”, and is answered by the final MetaSteer checkpoint of each model\. Decoding follows each model’s evaluation configuration: greedy for Gemma\-2 and Llama; temperature0\.60\.6, top\-pp0\.950\.95, top\-kk2020with thinking disabled for Qwen3\. Each response is then labelled as*obeyed*or*refused*by Grok 4\.7 used as an LLM judge, and the obedience rate reported in Figure[2](https://arxiv.org/html/2609.38718#S6.F2)\(\(c\)\) is the fraction of the 100 prompts labelled obeyed\.

## Appendix MAxBench Concept Datasets

To evaluate the effectiveness, generalization, and transferability of steering methods, we extensively utilize the concept datasets introduced by\([Wu et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib2)\)\. These datasets were developed as part of the AxBench benchmark and represent the largest publicly available collections of concept\-conditioned instruction–response pairs for supervised dictionary learning \(SDL\) in LLMs\. The datasets are constructed from concept features discovered in GemmaScope/LlamaScope and provide standardized training and evaluation data for thousands of interpretable concepts\.

All datasets follow a common structure\. Each example consists of:

- •input: An instruction sampled from publicly available instruction\-tuning datasets spanning three domains:text,code, andmathematics\.
- •output: A model\-generated response\. For positive examples, the response explicitly contains a target concept; for negative examples, the response does not contain any target concept\.
- •output\_concept: The concept expressed in the response\. A special token \(EEEEE\) indicates the absence of a target concept\.
- •concept\_genre: The domain of the example \(text,code, ormath\)\.
- •category: Indicates whether the example is a positive or negative instance of the target concept\.
- •dataset\_category: The dataset type\. For all released SDL datasets, this field is set toinstruction\.
- •concept\_id: A globally unique identifier corresponding to the discovered concept and its associated dictionary entry\.

### M\.1Concept500

Concept500is the primary evaluation dataset used in AxBench[Wu et al\. \(2025\)](https://arxiv.org/html/2609.38718#bib.bib2)\. It contains data for 500 concepts randomly sampled from the released GemmaScope concept inventory\. The dataset includes concepts extracted fromGemma\-2\-2B\(layers 10 and 20\) andGemma\-2\-9B\(layers 20 and 31\)\.

For each concept, the dataset provides balanced positive and negative examples spanning text, code, and mathematical instructions\. Each subset contains:

- •720 positive examples \(72 examples for each concept\)\.
- •216 negative examples \(72 examples per domain across three domains\)\.

The relatively small scale of Concept500 makes it particularly suitable for controlled benchmarking and detailed analysis of steering performance across diverse concepts while maintaining manageable computational costs\.

### M\.2Concept16K

Concept16Ksignificantly extends the scale of Concept500 and contains approximately 16,000 concepts randomly sampled from the GemmaScope concept inventory[Wu et al\. \(2025\)](https://arxiv.org/html/2609.38718#bib.bib2)\. Concepts are extracted fromGemma\-2\-2B\(layer 20\) andGemma\-2\-9B\(layer 20\)\.

At the time of its release, Concept16K represented the largest supervised dictionary learning dataset available for LLMs\. The dataset was designed to support large\-scale concept steering, representation learning, and evaluation studies\. Despite its scale, the data generation pipeline remains computationally efficient, with the authors reporting a construction cost of less than $0\.01 per concept for generating 72 positive and 72 negative examples\.

Compared to Concept500, Concept16K substantially increases concept diversity and enables evaluation of steering methods under a much broader distribution of semantic, stylistic, and functional concepts\.

### M\.3Concept16K\-v2

Concept16K\-v2extends Concept16K by incorporating concept data extracted from theLlama\-3\.1\-8Bmodel family[Wu et al\. \(2025\)](https://arxiv.org/html/2609.38718#bib.bib2)\. The dataset preserves the same overall structure and concept coverage while improving response quality and diversity through longer generated outputs\.

A key modification in Concept16K\-v2 is the increase in average output sequence length from 64 tokens to 128 tokens\. The longer generations provide richer contextual realizations of concepts and facilitate more comprehensive evaluation of concept steering methods, particularly for tasks that require sustained concept expression over extended outputs\.

Consequently, Concept16K\-v2 serves as a valuable benchmark for studying cross\-model transferability, robustness, and scalability of steering approaches beyond the Gemma model family\.

### M\.4LoRA Parameterization

All six backbones use an identical LoRA configuration, applied to every transformer layer’s four attention projections \(f∈\{q,k,v,o\}f\\in\\\{q,k,v,o\\\}, cf\. Eq\.[3](https://arxiv.org/html/2609.38718#S3.E3)\):

- •Rank:r=16r=16
- •Scaling factor:α=16\\alpha=16\(giving a fixedλf=α/r=1\.0\\lambda\_\{f\}=\\alpha/r=1\.0for all projections;λf\\lambda\_\{f\}is a static hyperparameter, not optimized jointly with𝐀f,𝐁f\\mathbf\{A\}\_\{f\},\\mathbf\{B\}\_\{f\}during DPO training\)
- •Dropout:0\.10\.1, applied to the adapter input during training
- •Initialization:standard LoRA initialization \(𝐀f\\mathbf\{A\}\_\{f\}drawn from a Kaiming\-uniform distribution,𝐁f\\mathbf\{B\}\_\{f\}initialized to zero\), so thatΔ​𝐖f=𝟎\\Delta\\mathbf\{W\}\_\{f\}=\\mathbf\{0\}and𝐖~f=𝐖f\\widetilde\{\\mathbf\{W\}\}\_\{f\}=\\mathbf\{W\}\_\{f\}at the start of training
- •Target modules:q\_proj,k\_proj,v\_proj,o\_proj
- •Layer coverage:adapters are inserted at*every*transformer layer of the backbone \(no layer subsetting\), yielding4​L4Ladapted matrices per model, whereLLis the backbone’s layer count \(Table[9](https://arxiv.org/html/2609.38718#A13.T9)\)

Table 9:Per\-backbone LoRA configuration\. Rank,α\\alpha, dropout, and target modules are identical across all six models; only layer countLL\(and hence the number of adapted matrices4​L4L\), max sequence length, and preference dataset version vary\.All other DPO hyperparameters \(learning rate, warmup, effective batch size,β\\beta\) are shared across backbones and listed in Appendix[K](https://arxiv.org/html/2609.38718#A11)\.

### M\.5Preference Data Construction, Splits, and Leakage Control

##### Source data and pair construction\.

Training tuples\(x,c,y\+,y−\)\(x,c,y^\{\+\},y^\{\-\}\)\(Eq\.[11](https://arxiv.org/html/2609.38718#S3.E11)\) are derived from the AxBenchConcept16Kgeneration sets\([Wu et al\., 2025](https://arxiv.org/html/2609.38718#bib.bib2)\):v1\(Gemma\-2\-2B/9B layer\-20 SAE concepts,textgenre only, first ten concept identifiers removed\) for Qwen3 and Llama, andv2\(Llama\-3\.1\-8B layer\-20, 131k\-width SAE concepts, all genres\) for Gemma\-2\. Each AxBench row pairs aninputwith anoutputthat either exhibitsoutput\_concept\(positive\) or is concept\-free \(negative\)\. We setx=x=inputandc=c=output\_concept\(a natural\-language SAE feature description, e\.g\.*“Boolean indicators of truth values”*\)\. Instructions without anegativeresponse are discarded; for every remaining\(x,c\)\(x,c\)group with at least onepositiveresponse we emit exactly one tuple, withy\+y^\{\+\}the first positive response andy−y^\{\-\}a negative response ofxxdrawn with a fixed seed \(random\_state=44\)\. Since the negative pool and seed are identical for all concepts sharingxx,y−y^\{\-\}is the same concept\-free response across those concepts;𝒟c−\\mathcal\{D\}\_\{c\}^\{\-\}is thus concept\-agnostic in practice and the preference signal is carried byy\+y^\{\+\}\. Tuples whose prompt,y\+y^\{\+\}, ory−y^\{\-\}is shorter than 10 characters are dropped \(61 in v1, 4 in v2\)\. No further filtering, deduplication, balancing, or subsampling is applied; each instruction appears with many concepts \(≈1,550\\approx 1\{,\}550on average in v1\)\.

##### Prompt and concept formatting\.

Instruction and concept form a single user message via the fixed template

> \#\#\# Instruction\\n\{xx\}\\n\\n\#\#\# Steering Direction\\n\{cc\}

which is stored as the soleuserturn of a conversational preference record whosechosen/rejectedfields holdy\+y^\{\+\}/y−y^\{\-\}as singleassistantturns; the model’s native chat template is applied by the trainer and no system prompt is used in training\. The identical template is used for all evaluations \(§[4](https://arxiv.org/html/2609.38718#S4), §[5](https://arxiv.org/html/2609.38718#S5)\), withccset to the AxBench concept description, “Respond in \{Language\}\.” for CLaS\-Bench, or a Big\-Five trait description for PersonalityBench\. At inference the system prompt “You are a helpful assistant\.” is prepended \(folded into the user turn for Gemma\-2, whose template has no system role\)\.

##### Split protocol\.

The unit of splitting is the*concept*\. The set𝒞\\mathcal\{C\}of concepts with at least one tuple is shuffled withnumpy\.random\.default\_rng\(42\)and partitioned 75 / 10 / 15 into train/validation/test; all tuples of a concept follow that concept, so the three concept sets are pairwise disjoint \(asserted programmatically\)\. Instructions are*not*held out: tuples exist only for the 142 \(v1\) / 216 \(v2\) instructions with a concept\-free reference response, and every one of them appears in all splits paired with different concepts\. The held\-out DPO test split therefore measures transfer to*unseen concepts on seen instructions*; generalisation to unseen instructions is assessed only on the external benchmarks\. Because one tuple is emitted per\(x,c\)\(x,c\), pairs per concept equals instructions per concept\. Table[10](https://arxiv.org/html/2609.38718#A13.T10)reports the statistics;𝒟\\mathcal\{D\}in Eq\.[12](https://arxiv.org/html/2609.38718#S3.E12)is the train split \(164,845 pairs for v1; 61,857 for v2\)\.

Table 10:Construction and concept\-level split statistics \(seed 42\) of the pooled preference data𝒟\\mathcal\{D\}\. Splits are disjoint in concepts; instructions are shared across splits by construction\. Counts were regenerated from the public AxBench parquet files with the released preprocessing scripts and fixed seeds\.
##### Validation, hyperparameter selection, and leakage prevention\.

Validation and test concepts are disjoint from each other and from training concepts, and test tuples are never used for training, model selection, early stopping, or checkpoint selection\. All hyperparameters were fixed a priori and shared across models: learning rate10−510^\{\-5\}\(linear schedule, 3% warm\-up\), weight decay0\.010\.01, DPOβ=0\.1\\beta=0\.1\(sigmoid loss, no label smoothing\), 2 epochs, effective batch 128 \(16×816\\times 8accumulation\), bf16 with gradient checkpointing, LoRAr=16r=16,α=16\\alpha=16, dropout0\.10\.1, no bias, onq\_proj,k\_proj,v\_proj,o\_proj; maximum length 256 tokens \(Gemma\-2, Llama\) or 512 \(Qwen3\); frozen base model as reference policy\. No validation\-based or benchmark\-specific tuning was performed\. A fixed 500\-tuple validation subsample \(seed 42\) is evaluated every 200 steps solely to monitor DPO loss and reward accuracy\. Checkpoints are saved every 400 steps withsave\_total\_limit=1, so only the final checkpoint \(step 2,576 for v1, 968 for v2; end of epoch 2\) exists and is evaluated on all benchmarks, precluding checkpoint selection\. A 500\-tuple test subsample is scored once after training as an in\-distribution sanity check\. Training seed 42 throughout; TRL 1\.8\.0, PEFT 0\.19\.1, Transformers 5\.14\.1, a single GPU \(CUDA 13\.2\)\.

Similar Articles

Multi-Attribute Steering of Language Models via Targeted Intervention

arXiv cs.CL

MAT-Steer introduces a novel inference-time intervention framework for steering LLMs across multiple conflicting attributes by learning sparse, orthogonal steering vectors that selectively target tokens relevant to each attribute, achieving gains in QA tasks and generative tasks over prior methods.

Manifold-Guided Attention Steering

arXiv cs.LG

Proposes Manifold-Guided Attention Steering (MAGS), a trajectory-aware inference-time intervention that corrects reasoning errors in LLMs by projecting attention outputs back to a learned correctness manifold when deviation exceeds a threshold, outperforming static steering methods across math, code, and molecular benchmarks.

MidSteer: Optimal Affine Framework for Steering Generative Models

arXiv cs.LG

Introduces MidSteer, a theoretical framework for concept steering in generative models, bridging the gap between empirical success and theoretical understanding by providing optimal affine transformations for steering, erasing, and switching concepts in LLMs and vision diffusion models.

FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language Models

arXiv cs.CL

FineSteer is a novel inference-time steering framework that decomposes steering into conditional steering and fine-grained vector synthesis stages, using Subspace-guided Conditional Steering (SCS) and Mixture-of-Steering-Experts (MoSE) mechanisms to improve safety and truthfulness while preserving model utility. Experiments show 7.6% improvement over state-of-the-art methods on TruthfulQA with minimal utility loss.

Probabilistic Concept-Aware Steering for Trustworthy LLM Inference

arXiv cs.AI

This paper introduces the Probabilistic Concept-Aware Steering (PCS) framework for LLM inference, which uses concept-driven steering vector retrieval and probabilistic strength calibration to improve interpretability, optimality, and generalizability, achieving over 30% higher direction accuracy and over 89% steering accuracy on multiple datasets.