大语言模型中基于因果激活引导的自适应多值控制

arXiv cs.LG 论文

摘要

本文介绍了AIMES,一个用于大语言模型中自适应多值激活引导的框架,它使用在线观察者反馈进行状态感知适应,无需额外训练,显示出相比固定方法提高了可控性。

arXiv:2609.30405v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in settings where responses must reflect multiple, potentially interacting social norms and human values. Activation steering offers a lightweight alternative to training-based alignment by modifying internal activations at inference time. However, prior human-value steering methods have largely considered values in isolation, while direct composition of multiple directions relies on fixed intervention strengths that cannot respond to the model's evolving internal state. Motivated by this key observation, we introduce AIMES, a framework for adaptive multi-value activation steering. AIMES constructs layer-specific bipolar directions for moral-foundation values and uses intermediate-layer vocabulary readouts as online observers. An observer-guided controller then adapts the strength of each requested value intervention at every decoding step based on its current observed state, without training a separate value-state estimator. Across multiple instruction-tuned model families, value combinations, and intervention depths, we find that multi-value controllability varies across both value combinations and intervention locations. Compared with fixed joint steering and prompt-based steering, AIMES shows depth-dependent advantages that are broadly supported across two independent evaluators, with some variation in the precise depth at which specific control effects emerge. These advantages come with smaller realized activation-space interventions than fixed-joint steering and comparable response quality. Overall, our results suggest that online observer feedback can provide lightweight, state-aware adaptation for single-pass multi-value steering.
查看原文
查看缓存全文

缓存时间: 2026/09/29 09:36

# Adaptive Multi-Value Control in LLMs via Causal Activation Steering
Source: [https://arxiv.org/html/2609.30405](https://arxiv.org/html/2609.30405)
Payel Bhattacharjee Ravi TandonSchool of Electrical, Computing, and Software EngineeringUniversity of Arizona, Tucson, AZ, USAEmail:\{payelb,tandonr\}@arizona\.edu

###### Abstract

Large language models \(LLMs\) are increasingly deployed in settings where responses must reflect multiple, potentially interacting social norms and human values\. Activation steering offers a lightweight alternative to training\-based alignment by modifying internal activations at inference time\. However, prior human\-value steering methods have largely considered values in isolation, while direct composition of multiple directions relies on fixed intervention strengths that cannot respond to the model’s evolving internal state\. Motivated by this key observation, we introduceAIMES, a framework for adaptive multi\-value activation steering\.AIMESconstructs layer\-specific bipolar directions for moral\-foundation values and uses intermediate\-layer vocabulary readouts as online observers\. An observer\-guided controller then adapts the strength of each requested value intervention at every decoding step based on its current observed state, without training a separate value\-state estimator\. Across multiple instruction\-tuned model families, value combinations, and intervention depths, we find that multi\-value controllability varies across both value combinations and intervention locations\. Compared with fixed joint steering and prompt\-based steering,AIMESshows depth\-dependent advantages that are broadly supported across two independent evaluators, with some variation in the precise depth at which specific control effects emerge\. These advantages come with smaller realized activation\-space interventions than fixed\-joint steering and comparable response quality\. Overall, our results suggest that online observer feedback can provide lightweight, state\-aware adaptation for single\-pass multi\-value steering\.

## 1Introduction

Large Language Models \(LLMs\) are increasingly deployed across critical domains such as education\([Al Faraby and Romadhony, 2024](https://arxiv.org/html/2609.30405#bib.bib29);[Alhafni et al\., 2024](https://arxiv.org/html/2609.30405#bib.bib30)\), research\([Ren et al\., 2025](https://arxiv.org/html/2609.30405#bib.bib35);[Liao et al\., 2024](https://arxiv.org/html/2609.30405#bib.bib36)\), healthcare\([Yang et al\., 2023](https://arxiv.org/html/2609.30405#bib.bib32);[Cascella et al\., 2023](https://arxiv.org/html/2609.30405#bib.bib31)\), and finance\([Lakkaraju et al\., 2023](https://arxiv.org/html/2609.30405#bib.bib33);[Zhao et al\., 2024](https://arxiv.org/html/2609.30405#bib.bib34)\), where model behavior may need to reflect multiple, potentially interacting human values and social norms\. This challenge is closely related to pluralistic alignment, which seeks to accommodate diverse and potentially competing values and perspectives, including through steerable adaptation to specified value trade\-offs\([Sorensen et al\., 2024](https://arxiv.org/html/2609.30405#bib.bib44)\)\. Existing alignment methods include reinforcement\-learning\-based approaches such as RLHF\([Ouyang et al\., 2022](https://arxiv.org/html/2609.30405#bib.bib37)\), RLAIF\([Lee et al\., 2023](https://arxiv.org/html/2609.30405#bib.bib38)\), TRPO\([Schulman et al\., 2015](https://arxiv.org/html/2609.30405#bib.bib39)\), and PPO\([Schulman et al\., 2017](https://arxiv.org/html/2609.30405#bib.bib40)\), as well as reward\-model\-free methods such as DPO\([Rafailov et al\., 2023](https://arxiv.org/html/2609.30405#bib.bib41)\)\. However, these approaches generally require additional preference collection, optimization, or fine\-tuning, making post\-hoc adaptation to changing value requirements costly\.

Activation steering\([Subramani et al\., 2022](https://arxiv.org/html/2609.30405#bib.bib53);[Turner et al\., 2023](https://arxiv.org/html/2609.30405#bib.bib5);[Zou et al\., 2023](https://arxiv.org/html/2609.30405#bib.bib4);[Li et al\., 2023](https://arxiv.org/html/2609.30405#bib.bib6)\)provides a lightweight inference\-time alternative mechanism for controlling model behavior by modifying internal representations without updating model parameters\. Recent work has extended this idea to human values, showing that moral concepts can be captured by layer\-specific activation directions and causally influenced through intervention\([Yu et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib1)\), but primarily considers one value at a time\. Jointly steering multiple values is more challenging because their representations need not be independent: directions may be correlated or conflicting, and intervention along one value can affect the expression of others\. Similar interference has been observed in generic multi\-attribute activation steering, where composing several directions can reduce control reliability\([van der Weij et al\., 2024](https://arxiv.org/html/2609.30405#bib.bib20);[Nguyen et al\., 2025b](https://arxiv.org/html/2609.30405#bib.bib16);[Ghasemi et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib19);[Jiang et al\., 2025](https://arxiv.org/html/2609.30405#bib.bib17)\)\. Existing approaches to multi\-value alignment largely address this problem through other mechanisms, such as prompting\([Kim et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib25)\)or combining separately generated value\-conditioned outputs\([Zheng et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib26)\)\. As a result, joint activation\-level control of multiple human values remains comparatively underexplored\.

![Refer to caption](https://arxiv.org/html/2609.30405v1/AIMES_workflow.png)Figure 1:Overview ofAIMES\.At each decoding step, the hidden state activation\(hℓ,t\)\(h\_\{\\ell,t\}\)at the intervention layer is mapped to value\-specific observer scores, which determine adaptive steering weights\(𝐜ℓ,t\)\(\\mathbf\{c\}\_\{\\ell,t\}\)for the requested values\. These weights modulate precomputed model and layer\-specific value directions\(uk,ℓ\)\(u\_\{k,\\ell\}\)before generation continues, enabling closed\-loop multi\-value steering\.A natural extension to multiple values is to jointly apply several value directions with fixed strengths, but such open\-loop steering cannot adapt to changes in the model’s internal state during the generation process\. Recent work instead formulates activation steering as feedback control, using estimates of the current concept state to adjust intervention strength\([Nguyen et al\., 2025a](https://arxiv.org/html/2609.30405#bib.bib21);[Bharadwaj, 2025](https://arxiv.org/html/2609.30405#bib.bib22);[Prokopiou et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib23);[Skifstad et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib24)\)\. Nevertheless, these methods typically rely on separately trained probes, classifiers, or dynamical models\. Intermediate\-layer vocabulary readouts, includingLogit Lens\([nostalgebraist, 2020](https://arxiv.org/html/2609.30405#bib.bib9)\),Tuned Lens\([Belrose et al\., 2023](https://arxiv.org/html/2609.30405#bib.bib10)\), andJ\-Lens\([Gurnee et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib2)\), offer a lightweight alternative by exposing information directly from hidden representations\. However, representational accessibility does not necessarily imply causal steerability\([Billa, 2026](https://arxiv.org/html/2609.30405#bib.bib14);[Nadaf, 2026](https://arxiv.org/html/2609.30405#bib.bib15)\), motivating an empirical comparison of candidate readouts before using them as feedback signals for adaptive control\.

Motivated by these observations, we introduceAIMES\(AdaptiveIntervention forMulti\-ValueEvaluation andSteering\), a framework that turns joint multi\-value activation steering into a state\-aware feedback process \(Figure[1](https://arxiv.org/html/2609.30405#S1.F1)\)\. Rather than composing multiple value directions with fixed strengths,AIMESobserves their current expression in an intermediate representation and adjusts their relative strength during decoding\. This enables adaptive coordination of interacting value directions without training a separate value\-state estimator\. Across multiple models, value objectives, and intervention depths, our results show that multi\-value controllability depends on both value combination and intervention location, and the cross\-judge analysis reveals robust advantages in some depth regions\. Our contributions are:

- •Bipolar value directions and online observation\.Building on contrastive activation steering\([Rimsky et al\., 2024](https://arxiv.org/html/2609.30405#bib.bib28);[Yu et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib1)\), we construct model\- and layer\-specific bipolar human\-value directions for bidirectional value control\. Intermediate\-layer vocabulary readouts are used as online value\-state observers without training a separate state estimator\.
- •Adaptive multi\-value steering\.We formulate joint value steering as a closed\-loop control problem and introduce an observer\-guided controller that adapts the relative strength of multiple value\-specific interventions at every decoding step\.
- •Depth\-aware and cross\-judge empirical evaluation\.Across five models spanning three families and parameter scales, and multiple multi\-value objectives, we systematically evaluate adaptive control across intervention depth under two independent judges,GPT\-5\.6\-SolandClaude\-Opus\-4\.8, characterizing the robustness and evaluator sensitivity of the observed multi\-value control patterns\.

## 2Related Work

In this section, we review prior work on activation steering and vocabulary readout methods, and additional background is provided in Appendix[A](https://arxiv.org/html/2609.30405#A1)\.

#### Activation Steering for Human Values and Multiple Attributes\.

Activation steering\([Subramani et al\., 2022](https://arxiv.org/html/2609.30405#bib.bib53)\)modifies model behavior at inference time by intervening along directions identified in hidden representation space, without updating model parameters\([Turner et al\., 2023](https://arxiv.org/html/2609.30405#bib.bib5);[Zou et al\., 2023](https://arxiv.org/html/2609.30405#bib.bib4);[Li et al\., 2023](https://arxiv.org/html/2609.30405#bib.bib6)\)\. In the human\-value setting,\([Yu et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib1)\)construct layer\-specific contrastive directions for the five dimensions of Moral Foundations Theory\(MFT\)\([Graham et al\., 2013](https://arxiv.org/html/2609.30405#bib.bib3)\)and show that interventions along these directions can causally alter foundation\-relevant behavior\. Importantly, they also observe cross\-value effects, suggesting that value directions cannot generally be treated as independent\. This issue becomes more pronounced when multiple directions are applied jointly\. More broadly, multi\-attribute steering methods such asMAT\-Steer\([Nguyen et al\., 2025b](https://arxiv.org/html/2609.30405#bib.bib16)\),ORBIT\([Ghasemi et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib19)\),MSRS\([Jiang et al\., 2025](https://arxiv.org/html/2609.30405#bib.bib17)\), andK\-Steering\([Oozeer et al\., 2025](https://arxiv.org/html/2609.30405#bib.bib18)\)address interference among jointly controlled attributes using mechanisms including gating, orthogonalization, subspace decomposition, and nonlinear direction construction\. These methods are motivated in part by evidence that naive vector composition becomes less reliable as the number of controlled attributes increases\([van der Weij et al\., 2024](https://arxiv.org/html/2609.30405#bib.bib20)\)\. However, the intervention is typically fixed before generation and does not adapt to evolving attribute expression during decoding\.

#### Steering Multiple Human Values\.

Recent work has explored multi\-value control through prompting, activation\-level, and compositional approaches:VALUEFLOW\([Kim et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib25)\)composes natural\-language value specifications and shows that value interactions can be reinforcing or competing, whileVISPA\([Zheng et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib26)\)combines separately generated value\-conditioned outputs rather than jointly steering multiple directions within a single generation\.\([Yu et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib1)\)construct moral\-foundation activation directions but intervene on values individually\. Joint activation\-level control of multiple human values within a single generation therefore remains comparatively underexplored\. Inspired by these approaches, we compareAIMESwith two non\-adaptive alternatives:Fixed Multi\-Value Steering, which jointly applies multiple value directions with predetermined strengths, andprompt steering, which specifies the same multi\-value objective through natural\-language instructions without modifying internal activations\. Fixed joint steering is closely related to generic multi\-attribute activation steering, where vector composition can introduce interference and reduce controllability\([van der Weij et al\., 2024](https://arxiv.org/html/2609.30405#bib.bib20);[Nguyen et al\., 2025b](https://arxiv.org/html/2609.30405#bib.bib16);[Ghasemi et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib19);[Jiang et al\., 2025](https://arxiv.org/html/2609.30405#bib.bib17)\)\. Concurrent work by\([Pan et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib43)\)attributes such interference to geometric entanglement and proposes a static Gram\-matrix correction\. In contrast,AIMESuses online observer feedback to adapt value\-specific intervention strengths during generation\.

#### Closed\-Loop Activation Steering\.

Recent work formulates activation steering as a feedback\-control problem, adapting intervention strength to the model’s evolving internal state\.\([Nguyen et al\., 2025a](https://arxiv.org/html/2609.30405#bib.bib21)\)distinguish fixed\-strength open\-loop steering from feedback\-based control using the discrepancy between desired and observed concept states;\([Bharadwaj, 2025](https://arxiv.org/html/2609.30405#bib.bib22)\)use PID\-style feedback with a chunk\-level classifier during reasoning;\([Wang et al\., 2025](https://arxiv.org/html/2609.30405#bib.bib55)\)adapt truthfulness\-steering strength online using probe\-estimated activation states across different hallucination categories;\([Prokopiou et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib23)\)study feedback\-controlled multi\-attribute steering in symbolic music generation; and\([Skifstad et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib24)\)model layer\-wise computation as a locally linear dynamical system and derive interventions using linear\-quadratic regulation\. Relatedly,ODESteer\([Zhao et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib51)\)makes the steering direction state\-dependent through the gradient of a learned density\-ratio barrier, but focuses on single\-attribute control and requires fitting the barrier\. A key requirement in observer\-based closed\-loop steering is anobserverthat estimates the controlled state, typically through a separately trained classifier, probe, or fitted dynamical model\([Nguyen et al\., 2025a](https://arxiv.org/html/2609.30405#bib.bib21);[Bharadwaj, 2025](https://arxiv.org/html/2609.30405#bib.bib22);[Prokopiou et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib23);[Skifstad et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib24)\)\.AIMESinstead uses intermediate vocabulary readouts as online observers and adapts the strengths of multiple jointly applied value directions during generation\.

#### Intermediate\-Layer Vocabulary Readouts\.

To interpret the intermediate states, intermediate\-layer readout methods map transformer hidden states into vocabulary\-aligned coordinates: theLogit Lens\([nostalgebraist, 2020](https://arxiv.org/html/2609.30405#bib.bib9)\)directly applies the output projection, theTuned Lens\([Belrose et al\., 2023](https://arxiv.org/html/2609.30405#bib.bib10)\)learns layer\-specific transformations toward the final output space, andJ\-Lens\([Gurnee et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib2)\)uses layer\-specific Jacobian information to approximate downstream vocabulary\-aligned content\. Beyond interpretability,\([Billa, 2026](https://arxiv.org/html/2609.30405#bib.bib14)\)show that Logit\-Lens\-based accessibility can help predict concept steerability and effective intervention depth across models and concept families\. However, representational visibility does not necessarily imply causal controllability: effective steering directions can remain weakly exposed by intermediate readouts such as the Logit Lens\([Billa, 2026](https://arxiv.org/html/2609.30405#bib.bib14);[Nadaf, 2026](https://arxiv.org/html/2609.30405#bib.bib15)\)\. This distinction is especially important in closed\-loop steering, where the readout itself can potentially serve as the feedback signal for intervention\.

## 3AIMES: Observer\-Guided Adaptive Multi\-Value Steering

In this section, we present our main proposed framework,AIMES\(AdaptiveIntervention forMulti\-ValueEvaluation andSteering\), which consists of an offline value direction\-construction stage and an online adaptive\-steering stage\. Offline, we construct model and layer\-specific human\-value directionsuk,ℓu\_\{k,\\ell\}for each target valueVkV\_\{k\}\([Rimsky et al\., 2024](https://arxiv.org/html/2609.30405#bib.bib28);[Yu et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib1)\)\. During generation, intermediate\-layer observers\([Gurnee et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib2);[nostalgebraist, 2020](https://arxiv.org/html/2609.30405#bib.bib9)\)estimate the current expression of the controlled values and guide the allocation of value\-specific steering strengths\. Relative to Fixed Multi\-Value Steering,AIMESreplaces fixed coefficients with observer\-guided adaptive weights that are updated online\. TheAIMESframework has three major components: \(1\)value\-specific steering directions, which define the intervention axes; \(2\) anintermediate\-layer value observer, which estimates the current value state from a vocabulary\-space readout; and \(3\) anobserver\-guided controller, which adjusts the contribution of each requested value direction during generation\. The intervention layer is fixed within each generation run and varied across experiments to study the effect of intervention depth\.

### 3\.1Value\-Specific Steering Direction Estimation

![Refer to caption](https://arxiv.org/html/2609.30405v1/3model_cosine.png)Figure 2:Cross\-value alignment analysis\.The figure shows early\-layer cosine similarities for three representative models, highlighting positive Care\-Fairness alignment, weaker Loyalty\-Authority alignment, and positive Sanctity alignment with Care and Fairness\.Given an autoregressive language modelMMwithLLlayers, letℓ∈\{1,…,L\}\\ell\\in\\\{1,\\ldots,L\\\}denote the layer at which an intervention is applied during a given generation run\. For an input promptxix\_\{i\}, the model generates a responseyi=\(yi1,…,yiTi\)y\_\{i\}=\(y\_\{i\}^\{1\},\\ldots,y\_\{i\}^\{T\_\{i\}\}\), whereTiT\_\{i\}is the number of generated tokens\. We consider a set ofKKtarget human values\{V1,…,VK\}\\\{V\_\{1\},\\ldots,V\_\{K\}\\\}and construct each value\-specific steering direction from a matched contrastive corpus, following the mean\-difference principle commonly used in activation steering\([Rimsky et al\., 2024](https://arxiv.org/html/2609.30405#bib.bib28);[Yu et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib1)\)\. For each valueVkV\_\{k\}, we form paired sets𝒟k\+\\mathcal\{D\}\_\{k\}^\{\+\}and𝒟k−\\mathcal\{D\}\_\{k\}^\{\-\}, where each pair describes the same underlying scenario while expressing opposite poles of the target value dimension\. For the Moral Foundations Theory dimensions considered here\([Graham et al\., 2013](https://arxiv.org/html/2609.30405#bib.bib3)\), these contrasts areCare/Harm, Fairness/Cheating, Loyalty/Betrayal, Authority/Subversion, and Sanctity/Degradation\. The complete vignette\-generation, pairing, and validation procedure is described in Appendix[C\.2](https://arxiv.org/html/2609.30405#A3.SS2.SSS0.Px2)\. Formally, at layerℓ\\ell, we compute the positive and negative mean activations,μk,ℓ\+\\mu\_\{k,\\ell\}^\{\+\}andμk,ℓ−\\mu\_\{k,\\ell\}^\{\-\}, respectively, and define the normalized value directionuk,ℓu\_\{k,\\ell\}as:

μk,ℓ\+=1\|𝒟k\+\|​∑x∈𝒟k\+hℓ​\(x\),μk,ℓ−=1\|𝒟k−\|​∑x∈𝒟k−hℓ​\(x\),uk,ℓ=μk,ℓ\+−μk,ℓ−‖μk,ℓ\+−μk,ℓ−‖2\.\\mu\_\{k,\\ell\}^\{\+\}=\\frac\{1\}\{\|\\mathcal\{D\}\_\{k\}^\{\+\}\|\}\\sum\_\{x\\in\\mathcal\{D\}\_\{k\}^\{\+\}\}h\_\{\\ell\}\(x\),\\qquad\\mu\_\{k,\\ell\}^\{\-\}=\\frac\{1\}\{\|\\mathcal\{D\}\_\{k\}^\{\-\}\|\}\\sum\_\{x\\in\\mathcal\{D\}\_\{k\}^\{\-\}\}h\_\{\\ell\}\(x\),\\qquad u\_\{k,\\ell\}=\\frac\{\\mu\_\{k,\\ell\}^\{\+\}\-\\mu\_\{k,\\ell\}^\{\-\}\}\{\\left\\\|\\mu\_\{k,\\ell\}^\{\+\}\-\\mu\_\{k,\\ell\}^\{\-\}\\right\\\|\_\{2\}\}\.\(1\)
Since the positive and negative sets consist of one\-to\-one matched pairs of equal size, the mean\-difference direction is equivalent to the average pairwise activation difference\. The resulting directionuk,ℓu\_\{k,\\ell\}therefore defines a bipolar activation\-space axis, with its positive and negative orientations corresponding to the two poles of valueVkV\_\{k\}\. FollowingContrastive Activation Addition \(CAA\)\([Rimsky et al\., 2024](https://arxiv.org/html/2609.30405#bib.bib28)\), this matched construction provides a direct bidirectional intervention axis while reducing variation due to topic, actors, setting, and writing style\. This is particularly useful forAIMES, which supports both amplification and suppression along the same value dimension\. However, the existence of a separable direction does not imply that its behavior is identical across layers\. We therefore evaluate multi\-value steering across multiple intervention depths in Section[4](https://arxiv.org/html/2609.30405#S4)\.

#### Geometry\-Guided Value\-Objective Selection\.

For the multi\-value steering analysis, we select three objectives to represent distinct interaction regimes suggested by the learned value geometry\. As illustrated in Figure[2](https://arxiv.org/html/2609.30405#S3.F2),CareandFairnessare positively aligned, motivating\(↑Care↑Fairness\)\(\\uparrow\\mathrm\{Care\}\\uparrow\\mathrm\{Fairness\}\)as an aligned coordination setting\.LoyaltyandAuthorityshow substantially weaker and, in some models, negative alignment, motivating the\(↑Loyalty↑Authority\)\(\\uparrow\\mathrm\{Loyalty\}\\uparrow\\mathrm\{Authority\}\)objective as a weakly\-coupled setting\. These two pairings have also been considered in prior multi\-value MFT steering\([Kim et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib25)\)and are consistent with the broader individualizing\-binding distinction in MFT\([Graham et al\., 2011](https://arxiv.org/html/2609.30405#bib.bib27);[Graham et al\., 2013](https://arxiv.org/html/2609.30405#bib.bib3)\)\. Finally,Sanctityis positively aligned with bothCareandFairness; we therefore consider\(↑Care↑Fairness↓Sanctity\)\(\\uparrow\\mathrm\{Care\}\\uparrow\\mathrm\{Fairness\}\\downarrow\\mathrm\{Sanctity\}\)as a competing setting, where a value geometrically coupled to the promoted pair is intentionally steered in the opposite direction\. Full similarity matrices across five models and depth regions are provided in Appendix[C\.4](https://arxiv.org/html/2609.30405#A3.SS4)\.

### 3\.2Intermediate\-Layer Value Observer: Design and Implementation

To adapt steering strength during generation,AIMESrequires an estimate of how strongly each target valueVkV\_\{k\}is expressed in the model’s current intermediate state\. Rather than training a separate value classifier, we derive this estimate directly from intermediate\-layer vocabulary readouts\. We consider existing vocabulary readout methods\([Gurnee et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib2);[nostalgebraist, 2020](https://arxiv.org/html/2609.30405#bib.bib9)\)as candidate observers, and use J\-Lens as the default based on the ablation in Appendix[C\.3](https://arxiv.org/html/2609.30405#A3.SS3)\. No value\-specific estimator is trained: the readout is fixed for a given model and layer and is shared across all values and steering objectives\. Logit Lens uses the model’s unembedding matrixWUW\_\{U\}, while J\-Lens additionally uses a precomputed layer\-specific JacobianJℓJ\_\{\\ell\}; further details are provided in Appendix[A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px4)\. For a given readout, letzℓ,t∈ℝ\|𝒱\|z\_\{\\ell,t\}\\in\\mathbb\{R\}^\{\|\\mathcal\{V\}\|\}denote its vocabulary\-space output at layerℓ\\elland generation steptt\. For each valueVkV\_\{k\}, we define a fixed token set𝒯k⊂𝒱\\mathcal\{T\}\_\{k\}\\subset\\mathcal\{V\}containing surface forms associated with that value \(see Appendix[B](https://arxiv.org/html/2609.30405#A2)for details\), followingMoral Foundations Theory \(MFT\)terminology\([Graham et al\., 2013](https://arxiv.org/html/2609.30405#bib.bib3)\)\. These sets are fixed before evaluation and are independent of model activations\. Extending them with additional value\-specific terms from external lexicons\([Atari et al\., 2023](https://arxiv.org/html/2609.30405#bib.bib7)\)may improve observer sensitivity and is left for future work\.

Vocabulary\-space readouts are known to expose evolving intermediate predictions\([nostalgebraist, 2020](https://arxiv.org/html/2609.30405#bib.bib9);[Belrose et al\., 2023](https://arxiv.org/html/2609.30405#bib.bib10)\), and prior work has shown that such projections can reveal interpretable concept structure in vocabulary space\([Geva et al\., 2022](https://arxiv.org/html/2609.30405#bib.bib45)\)\. Because different readout methods may produce scores with different offsets and scales, as also observed when comparing logit distributions across sources\([Sun et al\., 2024](https://arxiv.org/html/2609.30405#bib.bib46)\), we standardize each readout with respect to its vocabulary distribution at the current decoding step:z^ℓ,t=zℓ,t−Mean​\(zℓ,t\)Std​\(zℓ,t\),\\hat\{z\}\_\{\\ell,t\}=\\frac\{z\_\{\\ell,t\}\-\\text\{Mean\}\\left\(z\_\{\\ell,t\}\\right\)\}\{\\text\{Std\}\\left\(z\_\{\\ell,t\}\\right\)\},whereMean⁡\(⋅\)\\operatorname\{Mean\}\(\\cdot\)andStd⁡\(⋅\)\\operatorname\{Std\}\(\\cdot\)are computed over the vocabulary dimension\. This places candidate readouts on a common scale without requiring learned normalization parameters or readout\-specific calibration\. Motivated by the bag\-of\-words attribute model of PPLM\([Dathathri et al\., 2020](https://arxiv.org/html/2609.30405#bib.bib47)\), which aggregates output\-distribution probability mass over a curated word set to score an attribute, we introduce the*value\-specific readout score*\(sk,ℓ,t\)\(s\_\{k,\\ell,t\}\), which aggregates standardized vocabulary evidence for valueVkV\_\{k\}at an arbitrary intermediate layerℓ\\ellrather than only at the final output distribution, with larger values indicating stronger value\-specific evidence\.

###### Definition 1\(Value\-Specific Readout Score\)\.

Given a fixed token vocabulary\(𝒯k\)\(\\mathcal\{T\}\_\{k\}\)for a targeted valueVkV\_\{k\}, layerℓ\\ell, and generation steptt, thevalue\-specific readout scoreis defined as

scoreℓ,t​\(Vk\):=sk,ℓ,t=1\|𝒯k\|​∑w∈𝒯kz^ℓ,t​\[w\]\.\\text\{score\}\_\{\\ell,t\}\(V\_\{k\}\):=s\_\{k,\\ell,t\}=\\frac\{1\}\{\|\\mathcal\{T\}\_\{k\}\|\}\\sum\_\{w\\in\\mathcal\{T\}\_\{k\}\}\\hat\{z\}\_\{\\ell,t\}\[w\]\.\(2\)

We map this unbounded score to a bounded value state via a sigmoid, following the use of sigmoid\-transformed linear scores as bounded confidence signals for gating or scaling activation interventions in prior work\([Hościłowicz et al\., 2024](https://arxiv.org/html/2609.30405#bib.bib48);[Nguyen et al\., 2025b](https://arxiv.org/html/2609.30405#bib.bib16)\):ck,ℓ,t=σ⁡\(sk,ℓ,t\)=11\+exp⁡\(−sk,ℓ,t\)\.c\_\{k,\\ell,t\}=\\sigma\\\!\\left\(s\_\{k,\\ell,t\}\\right\)=\\frac\{1\}\{1\+\\exp\\left\(\-s\_\{k,\\ell,t\}\\right\)\}\.Thus,ck,ℓ,t∈\(0,1\)c\_\{k,\\ell,t\}\\in\(0,1\), with larger values indicating stronger observed expression ofVkV\_\{k\}\.

For layerℓ\\elland generation steptt, the*intermediate\-layer value observer*maps the current intermediate representationhℓ,th\_\{\\ell,t\}to𝐜ℓ,t=\[c1,ℓ,t,…,cK,ℓ,t\]⊤∈\(0,1\)K\.\\mathbf\{c\}\_\{\\ell,t\}=\[c\_\{1,\\ell,t\},\\ldots,c\_\{K,\\ell,t\}\]^\{\\top\}\\in\(0,1\)^\{K\}\.The vector𝐜ℓ,t\\mathbf\{c\}\_\{\\ell,t\}provides a bounded estimate of the model’s current value orientation under a readout\. Unlike the trained per\-attribute gates used in prior sigmoid\-based interventions\([Hościłowicz et al\., 2024](https://arxiv.org/html/2609.30405#bib.bib48);[Nguyen et al\., 2025b](https://arxiv.org/html/2609.30405#bib.bib16)\), our observer requires no probe training:𝐜ℓ,t\\mathbf\{c\}\_\{\\ell,t\}is computed directly from the standardized vocabulary evidence in equation[2](https://arxiv.org/html/2609.30405#S3.E2), and depends only on the current intermediate representation, using no information from previous decoding steps\. This keeps the feedback mechanism lightweight, training\-free, and responsive to changes during generation\.

### 3\.3Observer\-Guided Causal Multi\-Value Steering

This final stage ofAIMESperforms adaptive causal multi\-value steering by directly intervening on the model’s hidden representations\. We use*causal*here in the interventionist sense; rather than only measuring associations in representation space,AIMESmodifies the hidden state along value\-specific directions and evaluates the resulting change in model behavior\. For a fixed readout method and intervention layerℓ\\ell, the observer provides the current value state𝐜ℓ,t=\[c1,ℓ,t,…,cK,ℓ,t\]⊤\\mathbf\{c\}\_\{\\ell,t\}=\[c\_\{1,\\ell,t\},\\ldots,c\_\{K,\\ell,t\}\]^\{\\top\}\. Let𝐠=\[g1,…,gK\]⊤\\mathbf\{g\}=\[g\_\{1\},\\ldots,g\_\{K\}\]^\{\\top\}, andgk∈\{−1,0,\+1\}g\_\{k\}\\in\\\{\-1,0,\+1\\\}encode the steering objective for valueVkV\_\{k\}:\+1\+1for amplification,−1\-1for suppression, and00for no intervention\. At each decoding step,AIMESuses the current observer state to determine the contribution of each requested value direction\. The resulting joint intervention is

hℓ,t′=hℓ,t\+γ\[∑k:gk=\+1\(1−ck,ℓ,t\)uk,ℓ−∑k:gk=−1ck,ℓ,tuk,ℓ\]\.h^\{\\prime\}\_\{\\ell,t\}=h\_\{\\ell,t\}\+\\gamma\\left\[\\sum\_\{k:g\_\{k\}=\+1\}\(1\-c\_\{k,\\ell,t\}\)\\,u\_\{k,\\ell\}\-\\sum\_\{k:g\_\{k\}=\-1\}c\_\{k,\\ell,t\}\\,u\_\{k,\\ell\}\\right\]\.\(3\)
Here,γ\>0\\gamma\>0is the shared base steering magnitude anduk,ℓu\_\{k,\\ell\}is the unit\-normalized direction for valueVkV\_\{k\}\. The observer stateck,ℓ,tc\_\{k,\\ell,t\}determines the value\-specific steering strength: amplification is weighted by\(1−ck,ℓ,t\)\(1\-c\_\{k,\\ell,t\}\), whereas suppression is weighted byck,ℓ,tc\_\{k,\\ell,t\}\. The intervention therefore adapts to the currently observed value state at each decoding step\. The controller is memoryless and depends only on the current readout, unlike PID\-style controllers that accumulate error over time\([Bharadwaj, 2025](https://arxiv.org/html/2609.30405#bib.bib22)\)\. Together, Sections[3\.1](https://arxiv.org/html/2609.30405#S3.SS1)\-[3\.3](https://arxiv.org/html/2609.30405#S3.SS3)define the completeAIMESframework: value directions specify the intervention axes, the intermediate\-layer readout estimates the current value state, and the controller maps these estimates to adaptive steering coefficients during generation\. The fullAIMESalgorithm and inference procedure are provided in Appendix[B](https://arxiv.org/html/2609.30405#A2)\.

## 4Experimental Analysis and Results

#### Experimental Setup and Baselines\.

We evaluate five instruction\-tuned models spanning three families and parameter scales:Gemma\-3\-4B\-ITandGemma\-3\-12B\-IT\([Gemma Team et al\., 2025](https://arxiv.org/html/2609.30405#bib.bib13)\),Qwen3\-4BandQwen3\-14B\([Yang et al\., 2025](https://arxiv.org/html/2609.30405#bib.bib12)\), andLlama\-3\.1\-8B\-Instruct\([Grattafiori et al\., 2024](https://arxiv.org/html/2609.30405#bib.bib11)\)\. Following\([Yu et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib1)\), we study the five MFT dimensions:Care, Fairness, Loyalty, Authority, andSanctity\([Graham et al\., 2013](https://arxiv.org/html/2609.30405#bib.bib3)\)\. For each foundation, we construct 200 matched positive\-negative vignette pairs to derive value\-specific steering directions\. Our multi\-value evaluation considers three objectives:\(↑Care↑Fairness\)\(\\uparrow\\mathrm\{Care\}\\uparrow\\mathrm\{Fairness\}\),\(↑Loyalty↑Authority\)\(\\uparrow\\mathrm\{Loyalty\}\\uparrow\\mathrm\{Authority\}\), and\(↑Care↑Fairness↓Sanctity\)\(\\uparrow\\mathrm\{Care\}\\uparrow\\mathrm\{Fairness\}\\downarrow\\mathrm\{Sanctity\}\), and each objective is evaluated across ten model\-specific intervention depths spanning the network\. We compareAIMESwith two non\-adaptive baselines:\(i\) Fixed Multi\-Value Steeringwhich uses the same directionsuk,ℓu\_\{k,\\ell\}, prompts, intervention layers, generation settings, and base magnitude asAIMES, but applies all requested directions simultaneously with fixed strength,Δhℓ,t=γ∑k:gk≠0gkuk,ℓ\\Delta h\_\{\\ell,t\}=\\gamma\\sum\_\{k:g\_\{k\}\\neq 0\}g\_\{k\}u\_\{k,\\ell\}, without intermediate\-state feedback\. This is the direct multi\-value extension of contrastive activation addition\([Rimsky et al\., 2024](https://arxiv.org/html/2609.30405#bib.bib28)\)and corresponds to naive vector composition in multi\-attribute steering\([van der Weij et al\., 2024](https://arxiv.org/html/2609.30405#bib.bib20)\)\.\(ii\) Intensity\-Anchor Prompt Steeringwhich expresses the same multi\-value objective directly in the prompt\([Kim et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib25)\), without modifying model activations\. Because it is independent of intervention depth, the same prompt\-steered response is compared withAIMESat each tested layer\. Behavioral control is evaluated using blind value\-specific judgments from two independent evaluators,GPT\-5\.6\-Sol\([OpenAI, 2026](https://arxiv.org/html/2609.30405#bib.bib52)\)andClaude\-Opus\-4\.8\([Anthropic, 2026](https://arxiv.org/html/2609.30405#bib.bib8)\), applied to the same generated responses\. Our main analysis reports matched prompt\-level interaction contrasts and prompt\-level consistency under both judges, while the full depth\-wise cross\-judge results are provided in Appendix[D\.4](https://arxiv.org/html/2609.30405#A4.SS4)\. The interaction contrasts isolate the effect of adding a new value constraint, whereas prompt\-level consistency measures how broadly the resulting advantage holds across examples\. Response quality is evaluated separately using the blindGPT\-5\.6\-Soljudge; complete metric definitions are provided in Appendix[D](https://arxiv.org/html/2609.30405#A4.SS0.SSS0.Px1)\. All experiments are implemented with PyTorch and Hugging Face Transformers and run on NVIDIA A100 GPUs with Google Colab Pro\+\. Further details are provided in Appendix[C](https://arxiv.org/html/2609.30405#A3)\. The full code is publicly available at[https://anonymous\.4open\.science/status/AIMES\-9A6C](https://anonymous.4open.science/status/AIMES-9A6C)\.

![Refer to caption](https://arxiv.org/html/2609.30405v1/aimes_fs_ps_heatmap.png)Figure 3:Depth\-wiseAIMESadvantage under GPT\-5\.6\-Sol\.Panels \(a,b\) show Care/Fairness retention \(DxD\_\{x\}\) and panels \(c,d\) Sanctity suppression \(DyD\_\{y\}\) relative to Fixed Multi\-Value Steering and Prompt Steering\. Positive values favorAIMES; black outlines mark depths with unanimous positive advantage across all five models\.Table 1:Prompt\-level consistency across judges\.Tie\-adjustedAIMESconsistency \(%\) at early \(30%30\\%\) and middle \(60%60\\%\) depths\.DxD\_\{x\}measures Care/Fairness retention andDyD\_\{y\}Sanctity suppression\.
#### \(1\) Overall Analysis of Adaptive Multi\-Value Steering\.

Figure[3](https://arxiv.org/html/2609.30405#S4.F3)shows that the advantage of adaptive steering depends on both intervention depth and the controlled objective\. Relative to Fixed Multi\-Value Steering,Dx\>0D\_\{x\}\>0for all five models at around30%30\\%depth underGPT\-5\.6\-Sol, a pattern that is analyzed under the independentClaude\-Opus\-4\.8evaluator \(Appendix[D\.4](https://arxiv.org/html/2609.30405#A4.SS4), Figure[12](https://arxiv.org/html/2609.30405#A4.F12)\)\. UnderGPT\-5\.6\-Sol,Dy\>0D\_\{y\}\>0for all five models at around60%60\\%depth; however, the depth at which this Sanctity\-suppression advantage concentrates shifts under Claude, indicating evaluator sensitivity\. Against Prompt Steering, the Care/Fairness advantage is more persistent underGPT\-5\.6\-Sol, withDx\>0D\_\{x\}\>0for all five models at approximately30%30\\%,50%50\\%,80%80\\%, and90%90\\%depth, while the strongestDyD\_\{y\}consensus reaches four of five models near50%50\\%\. Overall, the depth localization of the Care/Fairness\-retention advantage is more robust across evaluators than that of the Sanctity\-suppression advantage\.

#### \(2\) Analyzing Prompt\-Level Consistency\.

Table[1](https://arxiv.org/html/2609.30405#S4.T1)complements the mean interaction contrasts in Figure[3](https://arxiv.org/html/2609.30405#S4.F3)by measuring how broadly each advantage is distributed across prompts\. Against Prompt Steering, GPT\-5\.6\-Sol yields majorityDxD\_\{x\}consistency for all five models at both displayed depths, whereas the Claude results are more heterogeneous\. Relative to Fixed Multi\-Value Steering, consistency varies more strongly across models, depths, and evaluators\. Thus, the mean interaction contrasts characterize the average advantage, while prompt\-level consistency indicates how broadly that advantage holds across examples\. Full results across all ten depths are provided in Appendix[D\.1](https://arxiv.org/html/2609.30405#A4.SS1)\.

Figure 4:GeoGain\-based multi\-value control evaluation\.GeoGain advantage ofAIMESover Fixed Multi\-Value Steering and Prompt Steering across sampled intervention depths, evaluated using two independent judges: GPT\-5\.6\-Sol \(a,b\) and Claude\-Opus\-4\.8 \(c,d\)\. In the figures, positive values favorAIMES; GeoGain measures balanced movement in\[Δ​C,Δ​F,−Δ​S\]\[\\Delta C,\\Delta F,\-\\Delta S\]\.Figure 5:Variation in AIMES steering strength within a response at 30% depth\.For each model and objective, we measure the within\-response range of the steering coefficient,Rα,k=maxt⁡αk,t−mint⁡αk,tR\_\{\\alpha,k\}=\\max\_\{t\}\\alpha\_\{k,t\}\-\\min\_\{t\}\\alpha\_\{k,t\}\. Larger values indicate greater changes in steering strength during generation\. Bars show response means and error bars denote95%95\\%prompt\-bootstrap confidence intervals\.
#### \(3\) Evaluating Balanced Multi\-Value Control\.

Figure[4](https://arxiv.org/html/2609.30405#S4.F4)evaluates joint control of\(↑Care↑Fairness↓Sanctity\)\(\\uparrow\\mathrm\{Care\}\\uparrow\\mathrm\{Fairness\}\\downarrow\\mathrm\{Sanctity\}\)using GeoGain\([Ghasemi et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib19)\)\. Relative to Fixed Multi\-Value Steering, the advantage ofAIMESremains model\- and depth\-dependent, with positive GeoGain differences in 15/25 plotted settings under GPT\-5\.6\-Sol and 16/25 under Claude\-Opus\-4\.8\. In contrast,AIMESachieves higher GeoGain than Prompt Steering in all 25 plotted settings under both judges\. Thus, the balanced\-control advantage is consistent relative to prompt\-based steering, while the comparison with fixed joint steering remains dependent on model and intervention depth\. Full depth\-wise results are provided in Appendix[D](https://arxiv.org/html/2609.30405#A4)\(Figure[10](https://arxiv.org/html/2609.30405#A3.F10)and Figure[13](https://arxiv.org/html/2609.30405#A4.F13)\)\.

#### \(4\) Within\-Response Steering Variation\.

Compared to fixed activation steering,AIMESupdates its intervention coefficients during generation using the online value readout\. We measure this variation with the coefficient range\(Rα,k=maxt⁡αk,t−mint⁡αk,t\)\(R\_\{\\alpha,k\}=\\max\_\{t\}\\alpha\_\{k,t\}\-\\min\_\{t\}\\alpha\_\{k,t\}\)\. Figure[5](https://arxiv.org/html/2609.30405#S4.F5)shows the mean range at the representative30%30\\%depth\. Across all five models and three objectives, the ranges are consistently nonzero \(0\.450\.45\-0\.830\.83\), showing that the steering strength varies within responses rather than remaining fixed\. Full results across all ten sampled depths are provided in Appendix[D\.3](https://arxiv.org/html/2609.30405#A4.SS3)\.

## 5Conclusion

We introducedAIMES, a closed\-loop framework for adaptive activation\-level steering of multiple human values using intermediate\-layer vocabulary readouts as online observers\. Across five models from three families, we find that multi\-value controllability depends strongly on both the steering objective and intervention depth\.AIMESprovides depth\-dependent advantages over fixed joint steering and, at favorable depths, improves joint\-control performance relative to prompt\-based steering, including for objectives that simultaneously promote and suppress values\. These gains are achieved with smaller realized activation\-space interventions than Fixed Multi\-Value Steering while maintaining comparable response quality\. Within the Moral Foundations Theory values and model scales studied here,AIMESoffers a lightweight approach to adaptive multi\-value control and motivates future work on broader value spaces, richer observers, and geometry\-aware controllers\.

### AI Use Statement

Generative AI tools were used to assist with synthetic contrastive\-vignette generation, code debugging, and manuscript polishing and editing\. All code, analyses, experimental decisions, and manuscript content were reviewed and verified by the authors\. Synthetic data generated with AI tools were subjected to the validation and filtering procedures described in Appendix[C\.2](https://arxiv.org/html/2609.30405#A3.SS2.SSS0.Px2)\. The authors take full responsibility for the final content of the paper, including all reported results, claims, and artifacts produced with the aid of generative AI\.

### Ethics Statement

This work studies inference\-time control of value\-related behavior in language models\. While such methods may support more adaptable model behavior, they could also be used to impose particular value specifications or manipulate model outputs in undesirable ways\. Our experiments focus on the five Moral Foundations Theory \(MFT\) dimensions and should not be interpreted as providing a complete or universal representation of human values\. The study uses synthetic and publicly available evaluation data and does not involve direct human\-subject experimentation\.

### Reproducibility Statement

We provide the value\-direction construction, observer definition, adaptive controller, evaluation metrics, and experimental setup in the main paper\. The appendix further documents the contrastive\-vignette generation procedure, model\-specific intervention layers, complete inference algorithm, observer ablations, value\-direction geometry, and full layer\-wise results\. An anonymous code repository containing the implementation and supporting artifacts is linked in the experimental section\.

## References

- S\. Al Faraby and A\. RomadhonyAnalysis of llms for educational question classification and generation\.Computers and Education: Artificial Intelligence7,pp\. 100298\.Cited by:[§1](https://arxiv.org/html/2609.30405#S1.p1.1)\.
- Alhafniet al\.\(2024\)B\. Alhafni, S\. Vajjala, S\. Bannò, K\. K\. Maurya, and E\. KochmarLLMs in education: novel perspectives, challenges, and opportunities\.arXiv preprint arXiv:2409\.11917\.Cited by:[§1](https://arxiv.org/html/2609.30405#S1.p1.1)\.
- Angrist and Pischke \(2009\)J\. D\. Angrist and J\. PischkeMostly harmless econometrics: an empiricist’s companion\.Princeton University Press\.Cited by:[Appendix D](https://arxiv.org/html/2609.30405#A4.SS0.SSS0.Px2.p1.1)\.
- Anthropic \(2026\)AnthropicClaude Opus 4\.8 System Card\.Note:[https://www\.anthropic\.com/system\-cards](https://www.anthropic.com/system-cards)Cited by:[§4](https://arxiv.org/html/2609.30405#S4.SS0.SSS0.Px1.p1.1)\.
- Atariet al\.\(2023\)M\. Atari, J\. Haidt, J\. Graham, S\. Koleva, S\. T\. Stevens, and M\. DehghaniMorality beyond the weird: how the nomological network of morality varies across cultures\.\.Journal of personality and social psychology125\(5\),pp\. 1157\.Cited by:[§3\.2](https://arxiv.org/html/2609.30405#S3.SS2.p1.1)\.
- Belroseet al\.\(2023\)N\. Belrose, Z\. Furman, L\. Smith, D\. Halawi, I\. Ostrovsky, L\. McKinney, S\. Biderman, and J\. SteinhardtEliciting latent predictions from transformers with the tuned lens\.External Links:2303\.08112Cited by:[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px4.p3.1),[§1](https://arxiv.org/html/2609.30405#S1.p3.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px4.p1.1),[§3\.2](https://arxiv.org/html/2609.30405#S3.SS2.p2.1)\.
- Bharadwaj \(2025\)A\. R\. BharadwajSTU\-PID: steering token usage via PID controller for efficient large language model reasoning\.arXiv preprint arXiv:2506\.18831\.Cited by:[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2609.30405#S1.p3.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px3.p1.1),[§3\.3](https://arxiv.org/html/2609.30405#S3.SS3.p2.1)\.
- Billa \(2026\)J\. BillaPredicting where steering vectors succeed\.arXiv preprint arXiv:2604\.15557\.Cited by:[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px4.p6.1),[§1](https://arxiv.org/html/2609.30405#S1.p3.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px4.p1.1)\.
- Card and Krueger \(1994\)D\. Card and A\. B\. KruegerMinimum wages and employment: a case study of the fast\-food industry in new jersey and pennsylvania\.American Economic Review84\(4\),pp\. 772–793\.Cited by:[Appendix D](https://arxiv.org/html/2609.30405#A4.SS0.SSS0.Px2.p1.1)\.
- Cascellaet al\.\(2023\)M\. Cascella, J\. Montomoli, V\. Bellini, and E\. BignamiEvaluating the feasibility of chatgpt in healthcare: an analysis of multiple clinical and research scenarios\.Journal of medical systems47\(1\),pp\. 33\.Cited by:[§1](https://arxiv.org/html/2609.30405#S1.p1.1)\.
- Dathathriet al\.\(2020\)S\. Dathathri, A\. Madotto, J\. Lan, J\. Hung, E\. Frank, P\. Molino, J\. Yosinski, and R\. LiuPlug and play language models: a simple approach to controlled text generation\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§3\.2](https://arxiv.org/html/2609.30405#S3.SS2.p2.1)\.
- Gemma Teamet al\.\(2025\)Gemma Team, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Rivière,et al\.Gemma 3 technical report\.arXiv preprint arXiv:2503\.19786\.Cited by:[§4](https://arxiv.org/html/2609.30405#S4.SS0.SSS0.Px1.p1.1)\.
- Gevaet al\.\(2022\)M\. Geva, A\. Caciularu, K\. Wang, and Y\. GoldbergTransformer feed\-forward layers build predictions by promoting concepts in the vocabulary space\.InProceedings of the 2022 conference on empirical methods in natural language processing,pp\. 30–45\.Cited by:[§3\.2](https://arxiv.org/html/2609.30405#S3.SS2.p2.1)\.
- Ghasemiet al\.\(2026\)N\. Ghasemi, A\. Ziashahabi, S\. Avestimehr, and J\. MayORBIT: training\-free multi\-attribute behavioral steering via orthogonal subspace rotation\.arXiv preprint arXiv:2606\.22357\.Cited by:[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px1.p2.1),[§C\.1](https://arxiv.org/html/2609.30405#A3.SS1.SSS0.Px5.p1.1),[Appendix D](https://arxiv.org/html/2609.30405#A4.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2609.30405#S1.p2.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px2.p1.1),[§4](https://arxiv.org/html/2609.30405#S4.SS0.SSS0.Px4.p1.1)\.
- Grahamet al\.\(2013\)J\. Graham, J\. Haidt, S\. Koleva, M\. Motyl, R\. Iyer, S\. P\. Wojcik, and P\. H\. DittoMoral foundations theory: the pragmatic validity of moral pluralism\.InAdvances in experimental social psychology,Vol\.47,pp\. 55–130\.Cited by:[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px1.p1.1),[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px5.p1.1),[§C\.2](https://arxiv.org/html/2609.30405#A3.SS2.SSS0.Px2.p2.2),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2609.30405#S3.SS1.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2609.30405#S3.SS1.p1.1),[§3\.2](https://arxiv.org/html/2609.30405#S3.SS2.p1.1),[§4](https://arxiv.org/html/2609.30405#S4.SS0.SSS0.Px1.p1.1)\.
- Grahamet al\.\(2011\)J\. Graham, B\. A\. Nosek, J\. Haidt, R\. Iyer, S\. Koleva, and P\. H\. DittoMapping the moral domain\.Journal of Personality and Social Psychology101\(2\),pp\. 366–385\.External Links:[Document](https://dx.doi.org/10.1037/a0021847)Cited by:[§C\.2](https://arxiv.org/html/2609.30405#A3.SS2.SSS0.Px2.p2.2),[§3\.1](https://arxiv.org/html/2609.30405#S3.SS1.SSS0.Px1.p1.1)\.
- Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan,et al\.The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.Cited by:[§4](https://arxiv.org/html/2609.30405#S4.SS0.SSS0.Px1.p1.1)\.
- Gurneeet al\.\(2026\)W\. Gurnee, N\. Sofroniew, A\. Pearce, M\. Piotrowski, I\. Kauvar, R\. Chen, A\. Soligo, P\. Bogdan, E\. Ong, R\. Wang,et al\.Verbalizable representations form a global workspace in language models\. transformer circuits thread\.Cited by:[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px4.p5.1),[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px4.p7.1),[§1](https://arxiv.org/html/2609.30405#S1.p3.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px4.p1.1),[§3\.2](https://arxiv.org/html/2609.30405#S3.SS2.p1.1),[§3](https://arxiv.org/html/2609.30405#S3.p1.1)\.
- Hościłowiczet al\.\(2024\)J\. Hościłowicz, M\. Sowański, P\. Czubowski, and A\. JanickiNon\-linear inference time intervention: improving llm truthfulness\.arXiv preprint arXiv:2403\.18680\.Cited by:[§3\.2](https://arxiv.org/html/2609.30405#S3.SS2.p3.1),[§3\.2](https://arxiv.org/html/2609.30405#S3.SS2.p4.1)\.
- Jianget al\.\(2025\)X\. Jiang, L\. Zhang, J\. Zhang, Q\. Yang, G\. Hu, D\. Wang, and L\. HuMSRS: adaptive multi\-subspace representation steering for attribute alignment in large language models\.arXiv preprint arXiv:2508\.10599\.Cited by:[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px1.p2.1),[§1](https://arxiv.org/html/2609.30405#S1.p2.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px2.p1.1)\.
- Kimet al\.\(2026\)W\. Kim, S\. Hyeon, J\. Oh, and J\. DoVALUEFLOW: toward pluralistic and steerable value\-based alignment in large language models\.InProceedings of the 43rd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.306\.Cited by:[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.30405#S1.p2.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px2.p1.1),[§3\.1](https://arxiv.org/html/2609.30405#S3.SS1.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2609.30405#S4.SS0.SSS0.Px1.p1.1)\.
- Lakkarajuet al\.\(2023\)K\. Lakkaraju, S\. E\. Jones, S\. K\. R\. Vuruma, V\. Pallagani, B\. C\. Muppasani, and B\. SrivastavaLLMs for financial advisement: a fairness and efficacy study in personal decision making\.InProceedings of the Fourth ACM International Conference on AI in Finance,pp\. 100–107\.Cited by:[§1](https://arxiv.org/html/2609.30405#S1.p1.1)\.
- Leeet al\.\(2023\)H\. Lee, S\. Phatale, H\. Mansoor, K\. R\. Lu, T\. Mesnard, J\. Ferret, C\. Bishop, E\. Hall, V\. Carbune, and A\. RastogiRLAIF: scaling reinforcement learning from human feedback with ai feedback\.Cited by:[§1](https://arxiv.org/html/2609.30405#S1.p1.1)\.
- Liet al\.\(2023\)K\. Li, O\. Patel, F\. Viégas, H\. Pfister, and M\. WattenbergInference\-time intervention: eliciting truthful answers from a language model\.InAdvances in Neural Information Processing Systems,Cited by:[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.30405#S1.p2.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px1.p1.1)\.
- Liaoet al\.\(2024\)Z\. Liao, M\. Antoniak, I\. Cheong, E\. Y\. Cheng, A\. Lee, K\. Lo, J\. C\. Chang, and A\. X\. ZhangLLMs as research tools: a large scale survey of researchers’ usage and perceptions\.arXiv preprint arXiv:2411\.05025\.Cited by:[§1](https://arxiv.org/html/2609.30405#S1.p1.1)\.
- Nadaf \(2026\)M\. S\. B\. NadafSteerable but not decodable: function vectors operate beyond the logit lens\.arXiv preprint arXiv:2604\.02608\.Cited by:[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px4.p6.1),[§1](https://arxiv.org/html/2609.30405#S1.p3.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px4.p1.1)\.
- Nguyenet al\.\(2025a\)D\. V\. Nguyen, H\. M\. Vu, N\. Y\. Pham, L\. Zhang, and T\. M\. NguyenActivation steering with a feedback controller\.arXiv preprint arXiv:2510\.04309\.Cited by:[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2609.30405#S1.p3.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px3.p1.1)\.
- Nguyenet al\.\(2025b\)D\. Nguyen, A\. Prasad, E\. Stengel\-Eskin, and M\. BansalMulti\-attribute steering of language models via targeted intervention\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 20619–20634\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1007)Cited by:[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px1.p2.1),[§1](https://arxiv.org/html/2609.30405#S1.p2.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2609.30405#S3.SS2.p3.1),[§3\.2](https://arxiv.org/html/2609.30405#S3.SS2.p4.1)\.
- nostalgebraist \(2020\)nostalgebraistInterpreting gpt: the logit lens\.Note:LessWrong[https://www\.lesswrong\.com/posts/AcKRB8wDpdaN6v6ru/interpreting\-gpt\-the\-logit\-lens](https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens)Cited by:[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2609.30405#S1.p3.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px4.p1.1),[§3\.2](https://arxiv.org/html/2609.30405#S3.SS2.p1.1),[§3\.2](https://arxiv.org/html/2609.30405#S3.SS2.p2.1),[§3](https://arxiv.org/html/2609.30405#S3.p1.1)\.
- Oozeeret al\.\(2025\)N\. F\. Oozeer, L\. Marks, F\. Barez, and A\. AbdullahBeyond linear steering: unified multi\-attribute control for language models\.InFindings of the Association for Computational Linguistics: EMNLP 2025,pp\. 23513–23557\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.1278)Cited by:[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px1.p2.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px1.p1.1)\.
- OpenAI \(2026\)OpenAIGPT\-5\.6: frontier intelligence that scales with your ambition\.Note:Product releaseCited by:[Appendix D](https://arxiv.org/html/2609.30405#A4.SS0.SSS0.Px5.p1.1),[§4](https://arxiv.org/html/2609.30405#S4.SS0.SSS0.Px1.p1.1)\.
- Ouyanget al\.\(2022\)L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, and A\. RayTraining language models to follow instructions with human feedback\.Advances in Neural Information Processing Systems35,pp\. 27730–27744\.Cited by:[§1](https://arxiv.org/html/2609.30405#S1.p1.1)\.
- Panet al\.\(2026\)W\. Pan, X\. Barron, J\. Zhou, and Z\. LiuSpillover\-aware multi\-value steering for pluralistic llm alignment\.External Links:2609\.05800Cited by:[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px2.p1.1)\.
- Prokopiouet al\.\(2026\)I\. Prokopiou, P\. Vikatos, M\. Kaliakatsos\-Papakostas, T\. Giannakopoulos, and T\. StafylakisClosing the loop: PID feedback control for interpretable activation steering in symbolic music generation\.arXiv preprint arXiv:2606\.18790\.Cited by:[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2609.30405#S1.p3.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px3.p1.1)\.
- Rafailovet al\.\(2023\)R\. Rafailov, A\. Sharma, E\. Mitchell, C\. D\. Manning, S\. Ermon, and C\. FinnDirect preference optimization: your language model is secretly a reward model\.Advances in Neural Information Processing Systems36,pp\. 53728–53741\.Cited by:[§1](https://arxiv.org/html/2609.30405#S1.p1.1)\.
- Renet al\.\(2025\)S\. Ren, P\. Jian, Z\. Ren, C\. Leng, C\. Xie, and J\. ZhangTowards scientific intelligence: a survey of llm\-based scientific agents\.arXiv preprint arXiv:2503\.24047\.Cited by:[§1](https://arxiv.org/html/2609.30405#S1.p1.1)\.
- Rimskyet al\.\(2024\)N\. Rimsky, N\. Gabrieli, J\. Schulz, M\. Tong, E\. Hubinger, and A\. TurnerSteering llama 2 via contrastive activation addition\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 15504–15522\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.828)Cited by:[§C\.2](https://arxiv.org/html/2609.30405#A3.SS2.SSS0.Px2.p6.4),[1st item](https://arxiv.org/html/2609.30405#S1.I1.i1.p1.1),[§3\.1](https://arxiv.org/html/2609.30405#S3.SS1.p1.1),[§3\.1](https://arxiv.org/html/2609.30405#S3.SS1.p2.1),[§3](https://arxiv.org/html/2609.30405#S3.p1.1),[§4](https://arxiv.org/html/2609.30405#S4.SS0.SSS0.Px1.p1.1)\.
- Saltonet al\.\(1975\)G\. Salton, A\. Wong, and C\. YangA vector space model for automatic indexing\.Communications of the ACM18\(11\),pp\. 613–620\.Cited by:[Table 4](https://arxiv.org/html/2609.30405#A3.T4)\.
- Scalenaet al\.\(2024\)D\. Scalena, G\. Sarti, and M\. NissimMulti\-property steering of large language models with dynamic activation composition\.InProceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP,Miami, Florida, US,pp\. 577–603\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.blackboxnlp-1.34),[Link](https://aclanthology.org/2024.blackboxnlp-1.34/)Cited by:[Appendix D](https://arxiv.org/html/2609.30405#A4.SS0.SSS0.Px4.p1.1)\.
- Schulmanet al\.\(2015\)J\. Schulman, S\. Levine, P\. Abbeel, M\. Jordan, and P\. MoritzTrust region policy optimization\.InInternational conference on machine learning,pp\. 1889–1897\.Cited by:[§1](https://arxiv.org/html/2609.30405#S1.p1.1)\.
- Schulmanet al\.\(2017\)J\. Schulman, F\. Wolski, P\. Dhariwal, A\. Radford, and O\. KlimovProximal policy optimization algorithms\.arXiv preprint arXiv:1707\.06347\.Cited by:[§1](https://arxiv.org/html/2609.30405#S1.p1.1)\.
- Skifstadet al\.\(2026\)J\. Skifstad, X\. A\. Yang, and G\. ChouLocal linearity of LLMs enables activation steering via model\-based linear optimal control\.arXiv preprint arXiv:2604\.19018\.Cited by:[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2609.30405#S1.p3.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px3.p1.1)\.
- Sorensenet al\.\(2024\)T\. Sorensen, J\. Moore, J\. Fisher, M\. L\. Gordon, N\. Mireshghallah, C\. M\. Rytting, A\. Ye, L\. Jiang, X\. Lu, N\. Dziri, T\. Althoff, and Y\. ChoiPosition: a roadmap to pluralistic alignment\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 46280–46302\.Cited by:[§1](https://arxiv.org/html/2609.30405#S1.p1.1)\.
- Subramaniet al\.\(2022\)N\. Subramani, N\. Suresh, and M\. E\. PetersExtracting latent steering vectors from pretrained language models, 2022\.URL https://arxiv\. org/abs/2205\.051241\.Cited by:[§1](https://arxiv.org/html/2609.30405#S1.p2.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px1.p1.1)\.
- Sunet al\.\(2024\)S\. Sun, W\. Ren, J\. Li, R\. Wang, and X\. CaoLogit standardization in knowledge distillation\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 15731–15740\.Cited by:[§3\.2](https://arxiv.org/html/2609.30405#S3.SS2.p2.1)\.
- Turneret al\.\(2023\)A\. M\. Turner, L\. Thiergart, G\. Leech, D\. Udell, J\. J\. Vazquez, U\. Mini, and M\. MacDiarmidSteering language models with activation engineering\.arXiv preprint arXiv:2308\.10248\.Cited by:[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.30405#S1.p2.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px1.p1.1)\.
- van der Weijet al\.\(2024\)T\. van der Weij, M\. Poesio, and N\. SchootsExtending activation steering to broad skills and multiple behaviours\.arXiv preprint arXiv:2403\.05767\.Cited by:[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px1.p2.1),[§1](https://arxiv.org/html/2609.30405#S1.p2.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px2.p1.1),[§4](https://arxiv.org/html/2609.30405#S4.SS0.SSS0.Px1.p1.1)\.
- Wanget al\.\(2025\)T\. Wang, X\. Jiao, Y\. Zhu, Z\. Chen, Y\. He, X\. Chu, J\. Gao, Y\. Wang, and L\. MaAdaptive activation steering: a tuning\-free llm truthfulness improvement method for diverse hallucinations categories\.InProceedings of the ACM on Web Conference 2025,pp\. 2562–2578\.External Links:[Document](https://dx.doi.org/10.1145/3696410.3714640)Cited by:[Appendix D](https://arxiv.org/html/2609.30405#A4.SS0.SSS0.Px4.p1.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px3.p1.1)\.
- Yanget al\.\(2025\)A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§4](https://arxiv.org/html/2609.30405#S4.SS0.SSS0.Px1.p1.1)\.
- Yanget al\.\(2023\)R\. Yang, T\. F\. Tan, W\. Lu, A\. J\. Thirunavukarasu, D\. S\. W\. Ting, and N\. LiuLarge language models in health care: development, applications, and challenges\.Health Care Science2\(4\),pp\. 255–263\.Cited by:[§1](https://arxiv.org/html/2609.30405#S1.p1.1)\.
- Yuet al\.\(2026\)C\. Yu, B\. Yi, F\. Karimi\-Malekabadi, S\. Abdurahman, J\. Ye, S\. Narayanan, Y\. Zhao, and M\. DehghaniTracing moral foundations in large language models\.arXiv preprint arXiv:2601\.05437\.Cited by:[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px1.p1.1),[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px2.p1.1),[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px5.p1.1),[§C\.2](https://arxiv.org/html/2609.30405#A3.SS2.SSS0.Px1.p2.1),[1st item](https://arxiv.org/html/2609.30405#S1.I1.i1.p1.1),[§1](https://arxiv.org/html/2609.30405#S1.p2.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px2.p1.1),[§3\.1](https://arxiv.org/html/2609.30405#S3.SS1.p1.1),[§3](https://arxiv.org/html/2609.30405#S3.p1.1),[§4](https://arxiv.org/html/2609.30405#S4.SS0.SSS0.Px1.p1.1)\.
- Zhaoet al\.\(2026\)H\. Zhao, H\. Sun, J\. Kong, X\. Li, Q\. Wang, L\. Jiang, Q\. Zhu, T\. Abdelzaher, Y\. Choi, M\. Li, and H\. ShaoODESteer: a unified ODE\-based steering framework for LLM alignment\.InInternational Conference on Learning Representations,Note:arXiv:2602\.17560Cited by:[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px3.p1.1)\.
- Zhaoet al\.\(2024\)H\. Zhao, Z\. Liu, Z\. Wu, Y\. Li, T\. Yang, P\. Shu, S\. Xu, H\. Dai, L\. Zhao, and H\. JiangRevolutionizing finance with llms: an overview of applications and insights\.arXiv preprint arXiv:2401\.11641\.Cited by:[§1](https://arxiv.org/html/2609.30405#S1.p1.1)\.
- Zhaoet al\.\(2025\)W\. Zhao, J\. Guo, Y\. Hu, Y\. Deng, A\. Zhang, X\. Sui, X\. Han, Y\. Zhao, B\. Qin, T\. Chua, and T\. LiuAdaSteer: your aligned LLM is inherently an adaptive jailbreak defender\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,Suzhou, China,pp\. 24559–24577\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.1248),[Link](https://aclanthology.org/2025.emnlp-main.1248/)Cited by:[Appendix D](https://arxiv.org/html/2609.30405#A4.SS0.SSS0.Px4.p1.1)\.
- Zhenget al\.\(2026\)S\. Zheng, J\. Zhong, A\. Shetty, H\. Ji, P\. Nakov, and U\. NaseemVISPA: pluralistic alignment via automatic value selection and activation\.arXiv preprint arXiv:2601\.12758\.Cited by:[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.30405#S1.p2.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px2.p1.1)\.
- Zouet al\.\(2023\)A\. Zou, L\. Phan, S\. Chen, J\. Campbell, P\. Guo, R\. Ren, A\. Pan, X\. Yin, M\. Mazeika, A\. Dombrowski,et al\.Representation engineering: a top\-down approach to ai transparency\.arXiv preprint arXiv:2310\.01405\.Cited by:[Appendix A](https://arxiv.org/html/2609.30405#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.30405#S1.p2.1),[§2](https://arxiv.org/html/2609.30405#S2.SS0.SSS0.Px1.p1.1)\.

## Appendix

The Appendix is organized as follows:

- •[A](https://arxiv.org/html/2609.30405#A1)Background and Preliminaries
- •[B](https://arxiv.org/html/2609.30405#A2)AIMESAlgorithm
- •[C](https://arxiv.org/html/2609.30405#A3)Detailed Experimental Setup
- •[D](https://arxiv.org/html/2609.30405#A4)Additional Experimental Results

## Appendix ABackground and Preliminaries

#### Activation\-Space Control and Direction Composition\.

Activation steering modifies model behavior at inference time by perturbing hidden representations along directions associated with desired concepts or behaviors\([Turner et al\., 2023](https://arxiv.org/html/2609.30405#bib.bib5);[Zou et al\., 2023](https://arxiv.org/html/2609.30405#bib.bib4);[Li et al\., 2023](https://arxiv.org/html/2609.30405#bib.bib6)\)\. For human values,\([Yu et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib1)\)construct layer\-specific contrastive directions for the five Moral Foundations Theory dimensions\([Graham et al\., 2013](https://arxiv.org/html/2609.30405#bib.bib3)\)and show that interventions along these directions can shift foundation\-relevant behavior\. Their analysis also reveals cross\-value effects, indicating that moral\-value directions are not independent\.

This issue is closely related to the broader problem of composing multiple activation directions\. Methods such as MAT\-Steer\([Nguyen et al\., 2025b](https://arxiv.org/html/2609.30405#bib.bib16)\), ORBIT\([Ghasemi et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib19)\), MSRS\([Jiang et al\., 2025](https://arxiv.org/html/2609.30405#bib.bib17)\), and K\-Steering\([Oozeer et al\., 2025](https://arxiv.org/html/2609.30405#bib.bib18)\)address interference among jointly controlled attributes using mechanisms including gating, orthogonalization, subspace decomposition, rotation, and nonlinear direction construction\. These methods build on evidence that naive vector addition can become unreliable when several attributes are controlled simultaneously\([van der Weij et al\., 2024](https://arxiv.org/html/2609.30405#bib.bib20)\)\. In most cases, however, the resulting intervention rule is determined before generation and remains fixed during decoding\.

#### Multi\-Value Control\.

Recent work considers multiple human values through several complementary mechanisms\. VALUEFLOW\([Kim et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib25)\)uses natural\-language value specifications to study how multiple value intensities interact, showing that their effects need not compose additively\. VISPA\([Zheng et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib26)\)combines separately generated value\-conditioned outputs for pluralistic alignment, while\([Yu et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib1)\)intervene on moral\-foundation directions individually\. Concurrently,\([Pan et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib43)\)study joint activation\-level multi\-value steering and account for direction interference through a static geometry\-based correction\.

These works highlight two distinct challenges in multi\-value control:*geometric interference*among value directions and*state dependence*during generation\. The former concerns how several directions interact when applied jointly; the latter concerns whether their relative strengths should remain fixed as the model’s internal state evolves\.AIMESfocuses on the second problem, using online feedback to adapt the contributions of jointly applied value directions during decoding\.

#### Feedback\-Based Activation Steering\.

A separate line of work treats activation steering as a feedback\-control problem\.\([Nguyen et al\., 2025a](https://arxiv.org/html/2609.30405#bib.bib21)\)distinguish fixed\-strength open\-loop steering from controllers that adjust intervention magnitude according to the difference between desired and observed concept states\.\([Bharadwaj, 2025](https://arxiv.org/html/2609.30405#bib.bib22)\)use PID\-style feedback with a chunk\-level classifier during LLM reasoning, while\([Prokopiou et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib23)\)study feedback\-controlled multi\-attribute steering in symbolic music generation\.\([Skifstad et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib24)\)instead approximate layer\-wise computation as a locally linear dynamical system and derive interventions through linear\-quadratic regulation\.

These approaches illustrate that intervention strength can be adjusted online, but their feedback signals are typically obtained from a trained classifier, probe, or fitted dynamical model\. Our setting instead uses an intermediate\-layer vocabulary readout directly as the feedback signal, allowing value\-specific intervention strengths to be updated without training a separate value\-state estimator\.

#### Layer\-Wise Vocabulary Readouts\.

Several methods map intermediate transformer representations into vocabulary space to make hidden states more interpretable\. The*Logit Lens*\([nostalgebraist, 2020](https://arxiv.org/html/2609.30405#bib.bib9)\)directly applies the model’s unembedding matrix to an intermediate residual\-stream activation

zℓ,tLL=WU​norm​\(hℓ,t\),\\displaystyle z\_\{\\ell,t\}^\{\\mathrm\{LL\}\}=W\_\{U\}\\,\\mathrm\{norm\}\(h\_\{\\ell,t\}\),\(4\)
effectively treating the intermediate state as if it were already expressed in final\-layer output coordinates\. The*Tuned Lens*\([Belrose et al\., 2023](https://arxiv.org/html/2609.30405#bib.bib10)\)instead learns a layer\-specific affine transformation,

zℓ,tTL=WU​norm​\(Aℓ​hℓ,t\+bℓ\),\\displaystyle z\_\{\\ell,t\}^\{\\mathrm\{TL\}\}=W\_\{U\}\\,\\mathrm\{norm\}\\left\(A\_\{\\ell\}h\_\{\\ell,t\}\+b\_\{\\ell\}\\right\),\(5\)
to better align intermediate activations with the model’s final output space\. J\-Lens\([Gurnee et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib2)\)takes a different approach, using a layer\-specific average JacobianJℓ∈ℝd×dJ\_\{\\ell\}\\in\\mathbb\{R\}^\{d\\times d\}to approximate how information at layerℓ\\ellpropagates toward the final representation:

zℓ,tJ=WU​norm​\(Jℓ​hℓ,t\)\.\\displaystyle z\_\{\\ell,t\}^\{J\}=W\_\{U\}\\,\\mathrm\{norm\}\(J\_\{\\ell\}h\_\{\\ell,t\}\)\.\(6\)The corresponding token directions form the layer\-specific dictionaryDℓ=\(WU​Jℓ\)⊤∈ℝd×\|𝒱\|D\_\{\\ell\}=\(W\_\{U\}J\_\{\\ell\}\)^\{\\top\}\\in\\mathbb\{R\}^\{d\\times\|\\mathcal\{V\}\|\}, providing a vocabulary\-aligned view of the residual stream\.

Beyond interpretability, intermediate\-layer vocabulary readouts have recently been connected to activation steering\.\([Billa, 2026](https://arxiv.org/html/2609.30405#bib.bib14)\)show that Logit\-Lens\-based accessibility can help predict both concept steerability and effective intervention depth across models and concept families\. However, representational visibility does not necessarily coincide with steerability:\([Nadaf, 2026](https://arxiv.org/html/2609.30405#bib.bib15)\)show that effective steering directions can remain weakly exposed or absent under the Logit Lens\. This distinction is particularly relevant toAIMES, where the readout is used not only to inspect the intermediate representation, but also as the feedback signal that determines the strength of the subsequent intervention\. We therefore compare candidate readouts as online observers under the same steering setup in Appendix[C\.3](https://arxiv.org/html/2609.30405#A3.SS3)\.

We briefly review the two lines of work most closely related to our setting: J\-lens representations and moral\-value tracing in LLM activation space\. J\-Lens\([Gurnee et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib2)\)maps intermediate residual\-stream activations into vocabulary\-aligned coordinates while accounting for how information propagates through the remaining layers of the model\. Lethℓ,t∈ℝdh\_\{\\ell,t\}\\in\\mathbb\{R\}^\{d\}denote the residual\-stream activation at layerℓ\\elland token positiontt, and letWU∈ℝ\|𝒱\|×dW\_\{U\}\\in\\mathbb\{R\}^\{\|\\mathcal\{V\}\|\\times d\}denote the unembedding matrix\. J\-Lens introduces a layer\-specific average Jacobian

Jℓ=𝔼prompt,t,t′≥t​\[∂hL,t′∂hℓ,t\],\\displaystyle J\_\{\\ell\}=\\mathbb\{E\}\_\{\\mathrm\{prompt\},\\,t,\\,t^\{\\prime\}\\geq t\}\\left\[\\frac\{\\partial h\_\{L,t^\{\\prime\}\}\}\{\\partial h\_\{\\ell,t\}\}\\right\],\(7\)which provides a first\-order approximation of how an intermediate representation influences later final\-layer representations\. The corresponding vocabulary\-aligned readout is

zℓ,tJ=WU​norm​\(Jℓ​hℓ,t\)\.\\displaystyle z\_\{\\ell,t\}^\{J\}=W\_\{U\}\\,\\mathrm\{norm\}\(J\_\{\\ell\}h\_\{\\ell,t\}\)\.\(8\)AIMESuses this intermediate vocabulary\-space signal to construct the value\-specific observer scores defined in Section[3\.2](https://arxiv.org/html/2609.30405#S3.SS2)\.

#### Moral\-Value Representations in LLMs\.

\([Yu et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib1)\)study how the five Moral Foundations Theory dimensions:Care, Fairness, Loyalty, Authority, andSanctity\([Graham et al\., 2013](https://arxiv.org/html/2609.30405#bib.bib3)\)are represented in LLM residual streams and whether these representations can be behaviorally influenced through intervention\. For a target foundationVkV\_\{k\}, they construct a layer\-specific direction from the difference between mean activations of a positive setℰk\+\\mathcal\{E\}\_\{k\}^\{\+\}and a contrast setℰk−\\mathcal\{E\}\_\{k\}^\{\-\}:

uk,ℓraw=1\|ℰk\+\|​∑i∈ℰk\+h~ℓ\(i\)−1\|ℰk−\|​∑j∈ℰk−h~ℓ\(j\),\\displaystyle u\_\{k,\\ell\}^\{\\mathrm\{raw\}\}=\\frac\{1\}\{\|\\mathcal\{E\}\_\{k\}^\{\+\}\|\}\\sum\_\{i\\in\\mathcal\{E\}\_\{k\}^\{\+\}\}\\tilde\{h\}\_\{\\ell\}^\{\(i\)\}\-\\frac\{1\}\{\|\\mathcal\{E\}\_\{k\}^\{\-\}\|\}\\sum\_\{j\\in\\mathcal\{E\}\_\{k\}^\{\-\}\}\\tilde\{h\}\_\{\\ell\}^\{\(j\)\},\(9\)and normalizing that, we obtain:

uk,ℓ=uk,ℓraw‖uk,ℓraw‖2\\displaystyle u\_\{k,\\ell\}=\\frac\{u\_\{k,\\ell\}^\{\\mathrm\{raw\}\}\}\{\\\|u\_\{k,\\ell\}^\{\\mathrm\{raw\}\}\\\|\_\{2\}\}\(10\)whereh~ℓ\(i\)\\tilde\{h\}\_\{\\ell\}^\{\(i\)\}denotes the residual\-stream representation at the final input\-token position\. The resultinguk,ℓ∈ℝdu\_\{k,\\ell\}\\in\\mathbb\{R\}^\{d\}defines a layer\-specific axis associated with the target moral foundation\.

The authors further analyze these directions using pretrained sparse autoencoders \(SAEs\), identifying decoder features aligned withuk,ℓu\_\{k,\\ell\}and interpreting them through their top\-activating contexts\. They then test the behavioral relevance of the learned directions through inference\-time interventions of the form

hℓ,t′=hℓ,t\+αℓ​uk,ℓ\.\\displaystyle h^\{\\prime\}\_\{\\ell,t\}=h\_\{\\ell,t\}\+\\alpha\_\{\\ell\}u\_\{k,\\ell\}\.\(11\)These results provide evidence that moral\-foundation information is organized along distinguishable layer\-wise activation directions whose manipulation can shift foundation\-relevant behavior\.AIMESbuilds on this contrastive\-direction framework, but uses matched positive\-negative poles for each foundation and extends the setting from single\-value interventions to adaptive joint multi\-value control\.

## Appendix BAIMESAlgorithm

#### AIMES: Inference\-Time Adaptive Multi\-Value Steering\.

For each generation run,AIMESfixes an intervention layerℓ\\ell, readoutrr, and base steering magnitudeγ\\gamma\. The intervention layer is varied across experiments to study depth effects, but remains fixed within a response\. EveryNNdecoding steps, the selected readout estimates the current expression of each controlled value and updates the corresponding adaptive coefficients\. We useN=1N\{=\}1in all main experiments;N\>1N\>1reduces readout cost\. These coefficients determine the relative contribution of each requested value direction until the next controller update\.

We use a fixed value\-specific token set𝒯k\\mathcal\{T\}\_\{k\}for each moral foundation, with ten surface\-form tokens per value\. The complete token sets are listed in Table[2](https://arxiv.org/html/2609.30405#A2.T2)\. These sets are fixed before evaluation and used consistently across all models and steering objectives\.

Table 2:Value\-specific observer token sets\.Each set𝒯k\\mathcal\{T\}\_\{k\}contains ten fixed surface\-form tokens associated with the corresponding Moral Foundations Theory value and is used to compute the value\-specific readout score\.Because each coefficient is computed from the observed state of its corresponding value while all requested directions are applied jointly, the same mechanism naturally supports both single\-value and multi\-value objectives\. In our multi\-value experiments,AIMESis compared with Fixed Multi\-Value Steering, which uses the same value directions, intervention layers, generation settings, and base magnitudeγ\\gammabut fixed coefficients, and with Intensity\-Anchor Prompt Steering, which expresses the same objective through the input prompt without modifying model activations\.

The completeAIMESframework is summarized in Algorithm[1](https://arxiv.org/html/2609.30405#alg1)\.

Algorithm 1AIMES:AdaptiveIntervention forMulti\-ValueEvaluation andSteering1:Model

MM, prompt

xix\_\{i\}, intervention layer

ℓ\\ell, value directions

\{uk,ℓ\}k=1K\\\{u\_\{k,\\ell\}\\\}\_\{k=1\}^\{K\}, steering objective

𝐠\\mathbf\{g\}, token sets

\{𝒯k\}k=1K\\\{\\mathcal\{T\}\_\{k\}\\\}\_\{k=1\}^\{K\}, readout

rr, base steering magnitude

γ\\gamma
2:Generated response

yiy\_\{i\}
3:for

t=1,…,Tit=1,\\ldots,T\_\{i\}do

4:Compute the pre\-intervention hidden state

hℓ,th\_\{\\ell,t\}
5:Compute

zℓ,t←readout​\(hℓ,t\)z\_\{\\ell,t\}\\leftarrow\\text\{readout\}\(h\_\{\\ell,t\}\)
6:Standardize

zℓ,tz\_\{\\ell,t\}as

z^ℓ,t=zℓ,t−Mean​\(zℓ,t\)Std​\(zℓ,t\)\\hat\{z\}\_\{\\ell,t\}=\\frac\{z\_\{\\ell,t\}\-\\text\{Mean\}\\left\(z\_\{\\ell,t\}\\right\)\}\{\\text\{Std\}\\left\(z\_\{\\ell,t\}\\right\)\}\(highlighted in Section[3\.2](https://arxiv.org/html/2609.30405#S3.SS2)\)

7:foreach

kksuch that

gk≠0g\_\{k\}\\neq 0do

8:Compute value

VkV\_\{k\}specific score

sk,ℓ,ts\_\{k,\\ell,t\}using equation[2](https://arxiv.org/html/2609.30405#S3.E2)

9:Compute steering weights

ck,ℓ,t←σ⁡\(sk,ℓ,t\)c\_\{k,\\ell,t\}\\leftarrow\\sigma\(s\_\{k,\\ell,t\}\)
10:endfor

11:Apply the joint adaptive intervention in equation[3](https://arxiv.org/html/2609.30405#S3.E3)to obtain steered state

hℓ,t′h^\{\\prime\}\_\{\\ell,t\}
12:Continue the forward pass from

hℓ,t′h^\{\\prime\}\_\{\\ell,t\}and generate the next token

13:endfor

14:return

yiy\_\{i\}

## Appendix CDetailed Experimental Setup

### C\.1Construction and Overview

#### Value\-Direction Construction\.

We study the five Moral Foundations Theory dimensions: Care, Fairness, Loyalty, Authority, and Sanctity\. For each foundation, we construct200200matched positive\-negative vignette pairs corresponding to Care/Harm, Fairness/Cheating, Loyalty/Betrayal, Authority/Subversion, and Sanctity/Degradation, and use them to derive value\-specific steering directions\. Each pair describes the same underlying setting while reversing the relevant moral orientation, reducing variation due to topic, actors, and writing style\.

For each model and transformer layer, we extract the final\-token hidden representation of the positive and negative examples and compute the normalized contrastive mean\-difference direction described in Section[3\.1](https://arxiv.org/html/2609.30405#S3.SS1)\. Direction\-construction examples are kept disjoint from all prompts used for behavioral evaluation\. The same precomputed directions are used by Fixed Multi\-Value Steering andAIMES, ensuring that differences between methods arise from the intervention strategy rather than from different direction estimates\.

#### Multi\-Value Evaluation Data\.

For the main multi\-value experiments, we construct a separate evaluation set of100100held\-out MFRC prompts, balanced across the five foundations \(2020prompts per foundation\)\. For each steering objective, we retain only prompts whose source foundation is not one of the values being directly controlled\. This avoids evaluating a multi\-value intervention on prompts that are already explicitly associated with one of its target foundations\.

Under this filtering, the two\-value objectives\(↑Care↑Fairness\)\(\\uparrow\\mathrm\{Care\}\\uparrow\\mathrm\{Fairness\}\)and\(↑Loyalty↑Authority\)\(\\uparrow\\mathrm\{Loyalty\}\\uparrow\\mathrm\{Authority\}\)each retain6060prompts, corresponding to the three non\-target foundations\. The three\-value competing objective\(↑Care↑Fairness↓Sanctity\)\(\\uparrow\\mathrm\{Care\}\\uparrow\\mathrm\{Fairness\}\\downarrow\\mathrm\{Sanctity\}\)retains4040prompts from the two remaining foundations, Loyalty and Authority\. Each objective is evaluated at all ten model\-specific intervention depths using the same prompts forAIMES, Fixed Multi\-Value Steering, and Prompt Steering\.

For the matched interaction analysis comparing\(↑Care↑Fairness↓Sanctity\)\(\\uparrow\\mathrm\{Care\}\\uparrow\\mathrm\{Fairness\}\\downarrow\\mathrm\{Sanctity\}\)with\(↑Care↑Fairness\)\(\\uparrow\\mathrm\{Care\}\\uparrow\\mathrm\{Fairness\}\), we restrict both objectives to the same4040prompts\. Thus, theDxD\_\{x\},DyD\_\{y\}, and prompt\-level consistency analyses usen=40n=40matched prompts per model\-layer\-comparator setting\.

#### Intervention Layers\.

Since the evaluated architectures contain different numbers of transformer layers, we compare intervention locations using normalized depth

dℓ=ℓL,d\_\{\\ell\}=\\frac\{\\ell\}\{L\},whereLLis the number of transformer layers\. For each model, we evaluate ten approximately evenly spaced depths spanning the network\. The exact layer indices are reported in Table[3](https://arxiv.org/html/2609.30405#A3.T3)\. The intervention layers are selected a priori to approximately span the network depth; no layer is selected post hoc based on behavioral performance\. The complete depth sweep is used to characterize how steering behavior changes through the network\.

Table 3:Transformer layers used in the depth sweep\. The layer indices are chosen to approximately cover normalized depths from0\.10\.1to1\.01\.0for each architecture\.
#### Multi\-Value Steering\.

The multi\-value experiments compareAIMESwith two non\-adaptive baselines: Fixed Multi\-Value Steering and Intensity\-Anchor Prompt Steering\. Fixed Multi\-Value Steering jointly applies all requested value directions at the same intervention layer using fixed coefficients, whileAIMESuses the same prompts, value directions, intervention layers, generation settings, and base magnitude but replaces the fixed coefficients with the observer\-dependent weights defined in Section[3\.3](https://arxiv.org/html/2609.30405#S3.SS3)\. Prompt Steering specifies the same multi\-value objective directly in the input prompt without modifying model activations\. Unsteered responses are used as matched references when computing value\-score changes\. For eachAIMESgeneration, the intervention layer, observer, and base magnitude are fixed, while the value\-specific coefficients are recomputed at every decoding step \(N=1N=1in all main experiments\)\. The controller is memoryless and retains no accumulated state across decoding steps\.

#### Metrics and Statistical Analysis\.

All method comparisons are performed on matched prompts\. Our primary analysis uses matched interaction contrasts to measure how adding a value constraint changes control relative to a shared base objective, together with prompt\-level consistency to characterize how broadly the effect holds across examples\. We additionally use GeoGain\([Ghasemi et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib19)\)to evaluate balanced movement across all requested values, and separately report blind response quality and realized activation\-space perturbation\. Pointwise bootstrap confidence intervals are used where reported\. Complete metric definitions and statistical details are provided in Appendix[D](https://arxiv.org/html/2609.30405#A4.SS0.SSS0.Px1)\.

### C\.2Contrastive Vignette Generation

![Refer to caption](https://arxiv.org/html/2609.30405v1/value_direction_workflow.png)Figure 6:Value\-specific steering direction construction\.Value examples are selected, converted into contrastive pairs, verified, and used to extract model and value\-specific activation directions\.#### Independent semantic validity check\.

We have generated the contrastive value\-specific and model\-specific directions as shown in Figure[6](https://arxiv.org/html/2609.30405#A3.F6)\. As a preliminary sanity check, we evaluated whether each generated foundation\-specific vignette set was semantically distinguishable from the Social Norm reference set using three independent sentence\-embedding models: BGE\-large\-en\-v1\.5, E5\-large\-v2, and all\-MPNet\-base\-v2\. For foundationVkV\_\{k\}and embedding modelmm, we computed

Δk\(m\)=𝔼x∼𝒟k​\[cos⁡\(em​\(x\),em​\(dk\)\)\]−𝔼x∼𝒟SN​\[cos⁡\(em​\(x\),em​\(dk\)\)\],\\displaystyle\\Delta\_\{k\}^\{\(m\)\}=\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\}\_\{k\}\}\\left\[\\cos\\\!\\left\(e\_\{m\}\(x\),e\_\{m\}\(d\_\{k\}\)\\right\)\\right\]\-\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\}\_\{\\mathrm\{SN\}\}\}\\left\[\\cos\\\!\\left\(e\_\{m\}\(x\),e\_\{m\}\(d\_\{k\}\)\\right\)\\right\],\(12\)whereem​\(⋅\)e\_\{m\}\(\\cdot\)denotes the representation produced by embedding modelmm, anddkd\_\{k\}denotes the textual definition of foundationVkV\_\{k\}\. A positiveΔk\(m\)\\Delta\_\{k\}^\{\(m\)\}therefore indicates that the foundation\-specific vignette set is, on average, more semantically aligned with its intended foundation definition than the Social Norm reference set\. All five foundations yielded positive contrasts under all three embedding models, indicating consistent set\-level separation from the Social Norm distribution\.

Table 4:Independent semantic sanity check of the generated foundation\-specific vignette sets\.Δ¯k\\overline\{\\Delta\}\_\{k\}denotes the mean foundation\-versus\-Social\-Norm cosine\-similarity\([Salton et al\., 1975](https://arxiv.org/html/2609.30405#bib.bib42)\)difference averaged across BGE\-large\-en\-v1\.5, E5\-large\-v2, and all\-MPNet\-base\-v2\.Averaged across embedding models, the observed differences were0\.0630\.063for Care,0\.0520\.052for Fairness,0\.0600\.060for Loyalty,0\.0350\.035for Authority, and0\.0780\.078for Sanctity, with Authority showing the weakest separation and Sanctity the strongest\. Because these embedding models are general\-purpose semantic encoders rather than Moral Foundations Theory classifiers, we use this analysis only as a corpus\-level sanity check\. This foundation\-versus\-Social\-Norm comparison also provides a semantic counterpart to the contrastive construction used in prior moral\-tracing work\([Yu et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib1)\), while the matched positive–negative corpus described separately is used to construct the bidirectional steering directions employed byAIMES\. The primary validation of these directions is subsequently performed in the target models’ activation space\.

#### Matched Positive\-Negative Contrastive Vignette Generation\.

For the primaryAIMESsteering directions, we construct a second synthetic corpus consisting of matched positive–negative vignette pairs for each of the five Moral Foundations Theory dimensions\. In contrast to the foundation\-versus\-Social\-Norm corpus used for reproduction and representation\-level comparison, this corpus is designed specifically to recover a bidirectional moral axis suitable for activation steering\.

For each moral foundationVkV\_\{k\}, we define the corresponding positive and negative poles according to the standard Moral Foundations Theory formulation:

Care\\displaystyle\\mathrm\{Care\}↔Harm,\\displaystyle\\leftrightarrow\\mathrm\{Harm\},\(13\)Fairness\\displaystyle\\mathrm\{Fairness\}↔Cheating,\\displaystyle\\leftrightarrow\\mathrm\{Cheating\},\(14\)Loyalty\\displaystyle\\mathrm\{Loyalty\}↔Betrayal,\\displaystyle\\leftrightarrow\\mathrm\{Betrayal\},\(15\)Authority\\displaystyle\\mathrm\{Authority\}↔Subversion,\\displaystyle\\leftrightarrow\\mathrm\{Subversion\},\(16\)Sanctity\\displaystyle\\mathrm\{Sanctity\}↔Degradation\.\\displaystyle\\leftrightarrow\\mathrm\{Degradation\}\.\(17\)These paired dimensions follow the canonical virtue–vice structure of Moral Foundations Theory\([Graham et al\., 2011](https://arxiv.org/html/2609.30405#bib.bib27);[Graham et al\., 2013](https://arxiv.org/html/2609.30405#bib.bib3)\)\. For each foundation, we generateN=200N=200matched contrastive pairs,

𝒟k±=\{\(xk,i\+,xk,i−\)\}i=1200,\\displaystyle\\mathcal\{D\}\_\{k\}^\{\\pm\}=\\left\\\{\\left\(x\_\{k,i\}^\{\+\},x\_\{k,i\}^\{\-\}\\right\)\\right\\\}\_\{i=1\}^\{200\},\(18\)wherexk,i\+x\_\{k,i\}^\{\+\}expresses the positive pole of the foundation andxk,i−x\_\{k,i\}^\{\-\}expresses its corresponding negative pole\. Equivalently, these matched pairs induce positive and negative example sets

𝒟k\+\\displaystyle\\mathcal\{D\}\_\{k\}^\{\+\}=\{xk,i\+\}i=1200,𝒟k−=\{xk,i−\}i=1200\.\\displaystyle=\\left\\\{x\_\{k,i\}^\{\+\}\\right\\\}\_\{i=1\}^\{200\},\\qquad\\mathcal\{D\}\_\{k\}^\{\-\}=\\left\\\{x\_\{k,i\}^\{\-\}\\right\\\}\_\{i=1\}^\{200\}\.\(19\)Across the five foundations, this produces1,0001\{,\}000matched pairs, or2,0002\{,\}000individual vignettes\.

The two members of each pair are generated jointly rather than independently\. The generator is instructed to preserve, as closely as possible, the same actors, setting, underlying event, grammatical structure, level of detail, and approximate length across the positive and negative versions\. The principal difference between the two members should be the moral polarity associated with the target foundation\. This matched construction reduces variation due to topic, setting, lexical content, and writing style, making the resulting activation difference more directly attributable to the target moral dimension\.

For example, a Care/Harm pair may take the form

> Care:A student stopped to comfort a classmate who was crying after being mocked\. Harm:A student joined the others in mocking a classmate who was already crying\.

Similarly, a Fairness/Cheating pair may contrast returning money accidentally overpaid by a customer with intentionally keeping the excess amount\. The specific wording is not fixed; rather, the generation procedure enforces structural similarity within each pair while varying scenarios across the corpus\.

To promote contextual diversity, the same twelve generation\-time social settings used in the foundation\-versus\-Social\-Norm corpus are retained:

1. 1\.family and household,
2. 2\.friends and interpersonal relationships,
3. 3\.school and education,
4. 4\.workplace,
5. 5\.community and neighborhood,
6. 6\.groups, clubs, and organizations,
7. 7\.public and service interactions,
8. 8\.online and digital interactions,
9. 9\.sports and recreation,
10. 10\.health and caregiving,
11. 11\.travel and transportation, and
12. 12\.social gatherings\.

These settings are used only as diversity controls and are not treated as moral labels or theoretically defined context categories\. Generation is approximately balanced across the twelve contexts, yielding roughly1616–1717pairs per context for each foundation\. Each generated record stores the pair identifier, target foundation, diversity context, positive vignette, negative vignette, generator identifier, and generation metadata\. The resulting data therefore have the conceptual structure

\(pair\_id,foundation,context,positive\_text,negative\_text\)\.\\displaystyle\\left\(\\texttt\{pair\\\_id\},\\texttt\{foundation\},\\texttt\{context\},\\texttt\{positive\\\_text\},\\texttt\{negative\\\_text\}\\right\)\.\(20\)
To construct the layer\-wise activation direction, we first compute the mean hidden representation of the positive and negative sets at layerℓ\\ell:

μk,ℓ\+\\displaystyle\\mu\_\{k,\\ell\}^\{\+\}=1N​∑i=1Nhℓ​\(xk,i\+\),μk,ℓ−=1N​∑i=1Nhℓ​\(xk,i−\)\.\\displaystyle=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}h\_\{\\ell\}\(x\_\{k,i\}^\{\+\}\),\\qquad\\mu\_\{k,\\ell\}^\{\-\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}h\_\{\\ell\}\(x\_\{k,i\}^\{\-\}\)\.\(21\)We then define the normalized positive–negative direction as

uk,ℓPN=μk,ℓ\+−μk,ℓ−‖μk,ℓ\+−μk,ℓ−‖2\.\\displaystyle u\_\{k,\\ell\}^\{\\mathrm\{PN\}\}=\\frac\{\\mu\_\{k,\\ell\}^\{\+\}\-\\mu\_\{k,\\ell\}^\{\-\}\}\{\\left\\\|\\mu\_\{k,\\ell\}^\{\+\}\-\\mu\_\{k,\\ell\}^\{\-\}\\right\\\|\_\{2\}\}\.\(22\)Because the two sets consist of one\-to\-one matched pairs of equal size, the unnormalized mean difference is equivalently

μk,ℓ\+−μk,ℓ−=1N​∑i=1N\[hℓ​\(xk,i\+\)−hℓ​\(xk,i−\)\]\.\\displaystyle\\mu\_\{k,\\ell\}^\{\+\}\-\\mu\_\{k,\\ell\}^\{\-\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\left\[h\_\{\\ell\}\(x\_\{k,i\}^\{\+\}\)\-h\_\{\\ell\}\(x\_\{k,i\}^\{\-\}\)\\right\]\.\(23\)Thus, the mean\-difference and pairwise\-difference formulations yield the same direction\. Here,hℓ​\(⋅\)h\_\{\\ell\}\(\\cdot\)denotes the hidden representation extracted from layerℓ\\ellusing a fixed token\-selection rule\. This construction follows the general contrastive mean\-difference principle used in activation\-steering methods\([Rimsky et al\., 2024](https://arxiv.org/html/2609.30405#bib.bib28)\)\.

The positive–negative corpus is kept separate from the foundation\-versus\-Social\-Norm corpus\. The latter is used primarily for reproducing and comparing against the representation structure reported in prior moral\-tracing work, whereas the positive–negative corpus provides the primary bidirectional steering directions used by AIMES\.

### C\.3Observer Ablation: J\-Lens vs\. Logit Lens

![Refer to caption](https://arxiv.org/html/2609.30405v1/Figures/observer_ablation_4panel.png)Figure 7:Observer ablation: J\-Lens versus Logit Lens\.Positive values favor J\-Lens\. J\-Lens yields higher directional responsiveness in all five models and generally favorable downstream trends for worst\-target gain, non\-target interference, and response quality\. Error bars in panels \(b\)\-\(d\) denote95%95\\%bootstrap confidence intervals\.Figure[7](https://arxiv.org/html/2609.30405#A3.F7)compares J\-Lens and Logit Lens as the feedback observer while holding the steering directions, controller, intervention layers, prompts, and generation settings fixed\. J\-Lens exhibits higher directional responsiveness in all five models, indicating that its internal value estimates more consistently move in the requested direction under intervention\. Downstream, AIMES\-JL achieves higher worst\-target gain in four of five models, lower non\-target interference in three of five, and higher response quality in four of five\. Most behavioral confidence intervals overlap zero, indicating that the downstream advantage is not uniform across architectures; however, these results, together with the consistently stronger directional responsiveness, motivate our use of J\-Lens as the default observer in the main experiments\.

### C\.4Layer\-Wise Geometry of Value Directions

Figure[8](https://arxiv.org/html/2609.30405#A3.F8)shows the pairwise cosine similarity among the five value directions at representative early, middle, and late layers for all five models\. The geometry is strongly model and depth\-dependent\. In Llama\-3\.1\-8B\-Instruct, pairwise similarities are predominantly positive and relatively stable across depth, although Loyalty remains less aligned with the other foundations\. The two Qwen3 models exhibit a clearer depth\-dependent transition: Loyalty is nearly orthogonal, and in some cases weakly anti\-aligned, with several other directions at early layers, while the directions become more positively correlated in middle and late layers\.

The Gemma models show the largest changes across depth\. In Gemma\-3\-4B\-IT, several early\-layer pairs are negatively aligned, whereas the middle\-layer directions become highly correlated before separating again toward later layers\. Gemma\-3\-12B\-IT similarly exhibits strong positive and negative relationships at early depth, followed by broadly positive correlations in the middle and late network\. These results indicate that the interaction structure among value directions cannot be treated as fixed across either models or layers\. Consequently, simultaneously applying several directions may produce substantially different geometric interactions depending on the intervention depth, motivating our evaluation of multi\-value steering across multiple depths rather than selecting a single layer a priori\.

![Refer to caption](https://arxiv.org/html/2609.30405v1/Figures/gemma3-4b-valuedirection.png)

![Refer to caption](https://arxiv.org/html/2609.30405v1/Figures/qwen3-4b-valuedirections.png)

![Refer to caption](https://arxiv.org/html/2609.30405v1/Figures/llama-8b-valuedirections.png)

![Refer to caption](https://arxiv.org/html/2609.30405v1/Figures/gemma3-12b-valuedirections.png)

![Refer to caption](https://arxiv.org/html/2609.30405v1/Figures/qwen3-14b-valuedirections.png)

Figure 8:Pairwise cosine similarity among the five value directions across models and network depth\. Each model is shown at representative early, middle, and late layers, with entries reportingcos⁡\(ui,ℓ,uj,ℓ\)\\cos\(u\_\{i,\\ell\},u\_\{j,\\ell\}\)for Care, Fairness, Loyalty, Authority, and Sanctity\. Positive values indicate aligned directions, values near zero indicate approximate orthogonality, and negative values indicate opposing geometry\.![Refer to caption](https://arxiv.org/html/2609.30405v1/Figures/RPR.png)Figure 9:Realized activation\-space perturbation reduction relative to Fixed Multi\-Value Steering,Rℓ=100​\(1−‖Δ​hAIMES,ℓ‖2/‖Δ​hFixed,ℓ‖2\)R\_\{\\ell\}=100\\left\(1\-\\\|\\Delta h\_\{\\mathrm\{AIMES\},\\ell\}\\\|\_\{2\}/\\\|\\Delta h\_\{\\mathrm\{Fixed\},\\ell\}\\\|\_\{2\}\\right\), on shared prompt states\.Rℓ\>0R\_\{\\ell\}\>0in all150150model\-layer\-objective settings \(95%95\\%CIs excluding zero\), with magnitude varying substantially by model, depth, and objective rather than reflecting a constant rescaling\.#### AIMESachieves control with smaller realized interventions\.

Figure[9](https://arxiv.org/html/2609.30405#A3.F9)compares the realized activation\-space update ofAIMESwith Fixed Multi\-Value Steering across all 150 model\-layer\-objective configurations\.AIMESproduces a smaller perturbation in every setting, with reductions ranging from approximately8%8\\%to87%87\\%; the pointwise95%95\\%bootstrap confidence interval excludes zero throughout\. The magnitude of the reduction varies substantially across models, layers, and objectives, indicating that the adaptive controller does not simply rescale the fixed intervention by a constant factor\. Thus, the behavioral gains reported in the main text are obtained with smaller realized activation updates rather than stronger interventions\.

#### Full Layer\-Wise GeoGain Analysis\.

Figure[10](https://arxiv.org/html/2609.30405#A3.F10)extends the balanced\-control analysis across the full set of tested intervention depths\. The comparison with Fixed Multi\-Value Steering varies substantially across models and layers, confirming that the relative benefit of adaptive steering is depth\-dependent\. In contrast, the comparison with Prompt Steering is more consistently positive across models and depth\. These results reinforce the main\-text finding that balanced control under the competing objective depends strongly on intervention location, particularly when compared with fixed joint steering\.

![Refer to caption](https://arxiv.org/html/2609.30405v1/Figures/GeoGain.png)Figure 10:GPT\-5\.6\-Sol GeoGain advantage for balanced multi\-value control under\(↑Care↑Fairness↓Sanctity\)\(\\uparrow\\text\{Care\}\\uparrow\\text\{Fairness\}\\downarrow\\text\{Sanctity\}\)\.GeoGain uses the geometric mean of the aligned target changes\[Δ​C,Δ​F,−Δ​S\]\[\\Delta C,\\Delta F,\-\\Delta S\], with a positive sign only when all targets move in the requested directions\. Curves show the prompt\-paired mean \(AIMES–comparator\) difference across intervention depths; positive values favorAIMES\.

## Appendix DAdditional Experimental Results

#### Evaluation Metrics\.

We evaluate adaptive multi\-value control using four complementary measures\. We use a shared base steering magnitude ofγ=1\.0\\gamma=1\.0for all main multi\-value experiments\.

#### \(a\) Matched Interaction Contrast\.

We use a matched interaction contrast, adapting the difference\-in\-differences formulation\([Card and Krueger, 1994](https://arxiv.org/html/2609.30405#bib.bib49);[Angrist and Pischke, 2009](https://arxiv.org/html/2609.30405#bib.bib50)\), to measure how adding a new value constraint changes control relative to a shared base objective\. LetA=\(↑Care↑Fairness↓Sanctity\)A=\(\\uparrow\\mathrm\{Care\}\\uparrow\\mathrm\{Fairness\}\\downarrow\\mathrm\{Sanctity\}\)andB=\(↑Care↑Fairness\)B=\(\\uparrow\\mathrm\{Care\}\\uparrow\\mathrm\{Fairness\}\)\. For promptii, methodmm, objectiveoo, and layerℓ\\ell, we define

GCF,i\(m,o,ℓ\)\\displaystyle G\_\{\\mathrm\{CF\},i\}^\{\(m,o,\\ell\)\}=Δ​Ci\(m,o,ℓ\)\+Δ​Fi\(m,o,ℓ\)2,\\displaystyle=\\frac\{\\Delta C\_\{i\}^\{\(m,o,\\ell\)\}\+\\Delta F\_\{i\}^\{\(m,o,\\ell\)\}\}\{2\},GS,i\(m,o,ℓ\)\\displaystyle G\_\{\\mathrm\{S\},i\}^\{\(m,o,\\ell\)\}=−Δ​Si\(m,o,ℓ\),\\displaystyle=\-\\Delta S\_\{i\}^\{\(m,o,\\ell\)\},\(24\)where changes are measured relative to the corresponding unsteered response\. The incremental effect of adding↓Sanctity\\downarrow\\mathrm\{Sanctity\}isIx=GCFA−GCFBI\_\{x\}=G\_\{\\mathrm\{CF\}\}^\{A\}\-G\_\{\\mathrm\{CF\}\}^\{B\}for Care/Fairness retention andIy=GSA−GSBI\_\{y\}=G\_\{\\mathrm\{S\}\}^\{A\}\-G\_\{\\mathrm\{S\}\}^\{B\}for additional Sanctity suppression\. For comparatorq∈\{Fixed,Prompt\}q\\in\\\{\\mathrm\{Fixed\},\\mathrm\{Prompt\}\\\}, we report

Dx\(q\)​\(ℓ\)\\displaystyle D\_\{x\}^\{\(q\)\}\(\\ell\)=1n​∑i\[Ix,i\(AIMES,ℓ\)−Ix,i\(q,ℓ\)\],\\displaystyle=\\frac\{1\}\{n\}\\sum\_\{i\}\\left\[I\_\{x,i\}^\{\(\\mathrm\{AIMES\},\\ell\)\}\-I\_\{x,i\}^\{\(q,\\ell\)\}\\right\],Dy\(q\)​\(ℓ\)\\displaystyle D\_\{y\}^\{\(q\)\}\(\\ell\)=1n​∑i\[Iy,i\(AIMES,ℓ\)−Iy,i\(q,ℓ\)\]\.\\displaystyle=\\frac\{1\}\{n\}\\sum\_\{i\}\\left\[I\_\{y,i\}^\{\(\\mathrm\{AIMES\},\\ell\)\}\-I\_\{y,i\}^\{\(q,\\ell\)\}\\right\]\.\(25\)Positive values favorAIMES\. We additionally report*prompt\-level consistency*, the tie\-adjusted percentage of matched prompts for which the corresponding \(AIMES–comparator\) incremental effect is positive\.

#### \(b\) Geometric Gain \(GeoGain\)\.

To measure balanced progress across all components of a composite objective, we use ORBIT\-style Geometric Gain \(GeoGain\)\([Ghasemi et al\., 2026](https://arxiv.org/html/2609.30405#bib.bib19)\)\. For aligned target changesΔ~k=gk​Δk\\tilde\{\\Delta\}\_\{k\}=g\_\{k\}\\Delta\_\{k\}, we compute

Mϵ=exp⁡\(1K​∑k=1Klog⁡\(\|Δ~k\|\+ϵ\)\),M\_\{\\epsilon\}=\\exp\\\!\\left\(\\frac\{1\}\{K\}\\sum\_\{k=1\}^\{K\}\\log\(\|\\tilde\{\\Delta\}\_\{k\}\|\+\\epsilon\)\\right\),and assign GeoGain\+Mϵ\+M\_\{\\epsilon\}only when allΔ~k\>0\\tilde\{\\Delta\}\_\{k\}\>0, and−Mϵ\-M\_\{\\epsilon\}otherwise\. GeoGain therefore rewards simultaneous movement in all requested directions while penalizing imbalanced control\. Method comparisons use prompt\-paired differences before averaging within each model–layer setting\.

#### \(c\) Within\-Response Controller Adaptation\.

Prior work on adaptive activation steering has shown the utility of varying steering intensity during inference\([Scalena et al\., 2024](https://arxiv.org/html/2609.30405#bib.bib54);[Wang et al\., 2025](https://arxiv.org/html/2609.30405#bib.bib55);[Zhao et al\., 2025](https://arxiv.org/html/2609.30405#bib.bib56)\)\. To quantify the realized amount of token\-level adaptation in our controller, we use the range of each target\-value coefficient trajectory,

Rα,k=maxt⁡αk,t−mint⁡αk,t,R\_\{\\alpha,k\}=\\max\_\{t\}\\alpha\_\{k,t\}\-\\min\_\{t\}\\alpha\_\{k,t\},whereαk,t\\alpha\_\{k,t\}denotes the steering coefficient applied to valueVkV\_\{k\}at decoding steptt\. For each model, steering objective, and intervention depth, we averageRα,kR\_\{\\alpha,k\}across target values and responses\. Larger values indicate greater within\-response variation, whereas values near zero indicate that the corresponding coefficient remains nearly constant throughout generation\. This diagnostic characterizes coefficient variation only and does not, by itself, establish whether that variation is behaviorally beneficial\.

#### \(d\) Response Quality\.

Stronger behavioral control can potentially come at the cost of generation quality, we independently evaluate whether each steering method preserves the quality of the generated response\. We use a blind GPT\-5\.6\-Sol\([OpenAI, 2026](https://arxiv.org/html/2609.30405#bib.bib52)\)judge that observes only the original prompt and generated response; the steering method, intervention layer, steering objective, and source foundation are hidden from the judge\. The judge independently scores*coherence*\(CC\),*fluency*\(FF\), and*relevance*\(RR\) on a four\-point scale, and we summarize overall response quality as

Q=C\+F\+R3\.Q=\\frac\{C\+F\+R\}\{3\}\.We report both the raw quality scores and their changes relative to the corresponding unsteered response,Δ​Q=Qsteered−Qunsteered\\Delta Q=Q\_\{\\mathrm\{steered\}\}\-Q\_\{\\mathrm\{unsteered\}\}, with analogous definitions for coherence, fluency, and relevance\. Identical prompt–response pairs are judged only once and their scores are reused across duplicated conditions, such as the layer\-independent Prompt Steering baseline\. All quality comparisons are computed at the exact intervention layer, without averaging across layers\.

### D\.1Robustness Analysis

#### GPT\-5\.6\-Sol: Full depth\-wise prompt\-level consistency\.

Table 5:Prompt\-level consistency across all intervention depths\.Tie\-adjustedAIMESconsistency \(%\) with pointwise 95% prompt\-bootstrap confidence intervals in brackets \(n=40n=40matched prompts per cell\)\.vs\. Prompt Steeringvs\. Fixed MultiModelDepthℓ\\ellDxD\_\{x\}DyD\_\{y\}DxD\_\{x\}DyD\_\{y\}Llama\-3\.1\-8B10%342\.5 \[28\.7, 56\.2\]50\.0 \[43\.8, 57\.5\]41\.2 \[27\.5, 55\.0\]50\.0 \[43\.8, 56\.2\]20%651\.2 \[37\.5, 65\.0\]50\.0 \[41\.2, 58\.8\]38\.7 \[25\.0, 53\.8\]50\.0 \[41\.2, 58\.8\]30%1055\.0 \[40\.0, 70\.0\]51\.2 \[43\.8, 58\.8\]60\.0 \[46\.2, 73\.8\]52\.5 \[47\.5, 57\.5\]40%1350\.0 \[36\.3, 63\.7\]47\.5 \[40\.0, 53\.8\]36\.2 \[22\.5, 51\.2\]48\.8 \[42\.5, 55\.0\]50%1651\.2 \[36\.2, 66\.2\]55\.0 \[48\.8, 61\.3\]51\.2 \[38\.7, 65\.0\]56\.2 \[48\.8, 63\.7\]60%1953\.8 \[38\.8, 67\.5\]52\.5 \[43\.8, 61\.3\]36\.2 \[23\.8, 50\.0\]52\.5 \[45\.0, 60\.0\]70%2245\.0 \[31\.2, 58\.8\]47\.5 \[41\.2, 53\.8\]50\.0 \[37\.5, 62\.5\]43\.8 \[36\.3, 50\.0\]80%2650\.0 \[36\.2, 63\.8\]50\.0 \[41\.2, 58\.8\]57\.5 \[43\.8, 71\.2\]51\.2 \[43\.8, 58\.8\]90%2960\.0 \[47\.5, 72\.5\]50\.0 \[42\.5, 57\.5\]58\.8 \[46\.2, 71\.2\]47\.5 \[40\.0, 55\.0\]100%3248\.8 \[36\.2, 62\.5\]51\.2 \[42\.5, 58\.8\]65\.0 \[53\.8, 76\.2\]52\.5 \[46\.2, 58\.7\]Gemma\-3\-4B10%351\.2 \[37\.5, 65\.0\]48\.8 \[41\.2, 56\.2\]40\.0 \[27\.5, 52\.5\]46\.2 \[38\.8, 53\.8\]20%756\.2 \[43\.8, 68\.8\]53\.8 \[46\.2, 61\.3\]43\.8 \[30\.0, 57\.5\]55\.0 \[48\.8, 62\.5\]30%1067\.5 \[55\.0, 78\.8\]52\.5 \[43\.8, 61\.3\]62\.5 \[50\.0, 75\.0\]50\.0 \[43\.8, 56\.2\]40%1461\.3 \[48\.8, 73\.8\]48\.8 \[42\.5, 55\.0\]51\.2 \[38\.8, 63\.7\]48\.8 \[42\.5, 55\.0\]50%1767\.5 \[55\.0, 78\.8\]51\.2 \[43\.8, 58\.8\]53\.8 \[41\.2, 66\.2\]53\.8 \[48\.8, 60\.0\]60%2058\.8 \[46\.2, 71\.2\]53\.8 \[46\.2, 61\.3\]45\.0 \[32\.5, 57\.5\]50\.0 \[43\.8, 57\.5\]70%2457\.5 \[45\.0, 70\.0\]50\.0 \[42\.5, 57\.5\]50\.0 \[38\.7, 61\.3\]47\.5 \[42\.5, 52\.5\]80%2753\.8 \[41\.2, 67\.5\]51\.2 \[45\.0, 57\.5\]45\.0 \[33\.8, 56\.2\]53\.8 \[50\.0, 58\.7\]90%3156\.2 \[43\.8, 68\.8\]51\.2 \[43\.8, 58\.8\]48\.8 \[38\.8, 58\.7\]48\.8 \[45\.0, 52\.5\]100%3458\.8 \[46\.2, 71\.2\]51\.2 \[45\.0, 57\.5\]50\.0 \[50\.0, 50\.0\]50\.0 \[50\.0, 50\.0\]Gemma\-3\-12B10%553\.8 \[40\.0, 67\.5\]55\.0 \[47\.5, 63\.7\]41\.2 \[31\.2, 51\.2\]51\.2 \[46\.2, 56\.2\]20%1055\.0 \[41\.2, 68\.8\]55\.0 \[46\.2, 63\.7\]47\.5 \[36\.2, 58\.8\]48\.8 \[41\.2, 56\.2\]30%1460\.0 \[46\.2, 73\.8\]51\.2 \[43\.8, 60\.0\]48\.8 \[37\.5, 58\.8\]46\.2 \[40\.0, 52\.5\]40%1952\.5 \[38\.8, 66\.2\]50\.0 \[41\.2, 58\.8\]55\.0 \[43\.8, 66\.2\]48\.8 \[43\.8, 53\.8\]50%2456\.2 \[41\.2, 70\.0\]52\.5 \[45\.0, 60\.0\]48\.8 \[37\.5, 60\.0\]55\.0 \[48\.8, 62\.5\]60%2953\.8 \[40\.0, 67\.5\]57\.5 \[50\.0, 65\.0\]48\.8 \[37\.5, 60\.0\]51\.2 \[46\.2, 56\.2\]70%3456\.2 \[42\.5, 70\.0\]55\.0 \[47\.5, 62\.5\]51\.2 \[40\.0, 61\.3\]51\.2 \[45\.0, 57\.5\]80%3857\.5 \[42\.5, 72\.5\]52\.5 \[45\.0, 60\.0\]55\.0 \[46\.2, 63\.7\]48\.8 \[45\.0, 52\.5\]90%4358\.8 \[45\.0, 71\.2\]53\.8 \[46\.2, 61\.3\]53\.8 \[45\.0, 62\.5\]52\.5 \[47\.5, 57\.5\]100%4857\.5 \[43\.8, 71\.2\]52\.5 \[45\.0, 60\.0\]50\.0 \[50\.0, 50\.0\]50\.0 \[50\.0, 50\.0\]Qwen3\-4B10%453\.8 \[40\.0, 67\.5\]55\.0 \[47\.5, 62\.5\]47\.5 \[37\.5, 57\.5\]51\.3 \[47\.5, 55\.0\]20%761\.3 \[47\.5, 73\.8\]56\.2 \[50\.0, 62\.5\]46\.2 \[35\.0, 57\.5\]51\.2 \[45\.0, 57\.5\]30%1162\.5 \[48\.8, 76\.2\]52\.5 \[45\.0, 60\.0\]51\.2 \[40\.0, 62\.5\]46\.2 \[41\.2, 51\.3\]40%1461\.3 \[48\.8, 73\.8\]51\.2 \[43\.8, 60\.0\]62\.5 \[53\.8, 71\.2\]45\.0 \[38\.8, 50\.0\]50%1857\.5 \[43\.8, 71\.2\]52\.5 \[45\.0, 60\.0\]46\.2 \[35\.0, 57\.5\]51\.2 \[45\.0, 57\.5\]60%2265\.0 \[52\.5, 77\.5\]56\.2 \[48\.8, 63\.7\]60\.0 \[50\.0, 70\.0\]55\.0 \[50\.0, 61\.3\]70%2561\.3 \[48\.8, 73\.8\]53\.8 \[46\.2, 61\.3\]56\.2 \[48\.8, 65\.0\]48\.8 \[45\.0, 52\.5\]80%2961\.3 \[47\.5, 73\.8\]53\.8 \[46\.2, 61\.3\]52\.5 \[43\.8, 61\.3\]50\.0 \[46\.2, 53\.8\]90%3257\.5 \[43\.8, 71\.2\]52\.5 \[45\.0, 58\.8\]55\.0 \[50\.0, 61\.3\]48\.8 \[46\.2, 50\.0\]100%3660\.0 \[47\.5, 72\.5\]53\.8 \[46\.2, 61\.3\]47\.5 \[42\.5, 52\.5\]50\.0 \[50\.0, 50\.0\]Qwen3\-14B10%455\.0 \[41\.2, 67\.5\]48\.8 \[40\.0, 57\.5\]47\.5 \[36\.3, 58\.7\]52\.5 \[50\.0, 56\.2\]20%863\.8 \[50\.0, 76\.2\]48\.8 \[40\.0, 57\.5\]53\.8 \[41\.2, 66\.3\]50\.0 \[45\.0, 55\.0\]30%1251\.2 \[37\.5, 65\.0\]50\.0 \[41\.2, 58\.8\]48\.8 \[38\.7, 58\.7\]51\.2 \[46\.2, 56\.2\]40%1655\.0 \[41\.2, 68\.8\]50\.0 \[41\.2, 58\.8\]55\.0 \[43\.8, 66\.2\]47\.5 \[42\.5, 52\.5\]50%2050\.0 \[36\.3, 63\.7\]48\.8 \[40\.0, 57\.5\]45\.0 \[33\.8, 57\.5\]48\.8 \[43\.8, 53\.8\]60%2456\.2 \[43\.8, 67\.5\]51\.2 \[43\.8, 60\.0\]51\.2 \[40\.0, 62\.5\]51\.2 \[46\.2, 56\.2\]70%2856\.2 \[42\.5, 70\.0\]52\.5 \[43\.8, 61\.3\]52\.5 \[42\.5, 62\.5\]53\.8 \[50\.0, 58\.7\]80%3258\.8 \[45\.0, 71\.2\]48\.8 \[40\.0, 57\.5\]48\.8 \[40\.0, 58\.8\]51\.3 \[47\.5, 55\.0\]90%3655\.0 \[42\.5, 68\.8\]50\.0 \[41\.2, 58\.8\]40\.0 \[32\.5, 47\.5\]51\.2 \[50\.0, 53\.8\]100%4055\.0 \[42\.5, 67\.5\]50\.0 \[41\.2, 58\.8\]47\.5 \[43\.8, 50\.0\]48\.8 \[46\.2, 50\.0\]
#### Claude\-Opus\-4\.8: Full depth\-wise prompt\-level consistency\.

Table[6](https://arxiv.org/html/2609.30405#A4.T6)complements the Claude mean interaction contrasts in Figure[12](https://arxiv.org/html/2609.30405#A4.F12)by showing how consistently the corresponding advantage occurs across individual prompts at every tested intervention depth\. Each entry reports the tie\-adjusted percentage of matched prompts for whichAIMESobtains a largerDxD\_\{x\}orDyD\_\{y\}interaction effect than the corresponding comparator\. Values above50%50\\%therefore indicate that theAIMESadvantage occurs for a majority of prompts\. The results show that prompt\-level consistency varies across models, objectives, and intervention depths, while positive\-majority consistency appears across a broad range of settings for both comparison methods\.

Table 6:Claude\-Opus\-4\.8 prompt\-level consistency across all intervention depths\.Tie\-adjustedAIMESconsistency \(%\) under Claude\-Opus\-4\.8 with pointwise 95% prompt\-bootstrap confidence intervals in brackets \(n=40n=40matched prompts per cell\)\.DxD\_\{x\}measures Care/Fairness retention andDyD\_\{y\}additional Sanctity suppression\.vs\. Prompt Steeringvs\. Fixed MultiModelDepthℓ\\ellDxD\_\{x\}DyD\_\{y\}DxD\_\{x\}DyD\_\{y\}Llama\-3\.1\-8B10%10\\%331\.2 \[21\.2, 42\.5\]48\.8 \[41\.2, 56\.2\]45\.0 \[33\.8, 56\.2\]50\.0 \[43\.8, 57\.5\]20%20\\%643\.8 \[31\.2, 56\.2\]48\.8 \[41\.2, 56\.2\]45\.0 \[33\.8, 56\.2\]47\.5 \[41\.2, 53\.8\]30%30\\%1048\.8 \[36\.3, 61\.3\]55\.0 \[47\.5, 62\.5\]55\.0 \[43\.8, 66\.2\]51\.2 \[45\.0, 57\.5\]40%40\\%1342\.5 \[31\.2, 53\.8\]55\.0 \[48\.8, 62\.5\]52\.5 \[38\.8, 66\.2\]56\.2 \[51\.2, 61\.3\]50%50\\%1642\.5 \[31\.2, 53\.8\]51\.2 \[43\.8, 58\.8\]53\.8 \[42\.5, 65\.0\]52\.5 \[46\.2, 58\.8\]60%60\\%1945\.0 \[32\.5, 57\.5\]47\.5 \[40\.0, 55\.0\]48\.8 \[37\.5, 58\.8\]45\.0 \[37\.5, 51\.2\]70%70\\%2241\.2 \[28\.7, 53\.8\]52\.5 \[46\.2, 60\.0\]53\.8 \[42\.5, 65\.0\]50\.0 \[43\.8, 56\.2\]80%80\\%2643\.8 \[32\.5, 55\.0\]50\.0 \[42\.5, 57\.5\]41\.2 \[31\.2, 51\.2\]47\.5 \[41\.2, 53\.8\]90%90\\%2951\.2 \[38\.8, 63\.7\]53\.8 \[46\.2, 61\.3\]56\.2 \[46\.2, 65\.0\]51\.2 \[46\.2, 56\.2\]100%100\\%3250\.0 \[37\.5, 62\.5\]51\.2 \[45\.0, 57\.5\]55\.0 \[46\.2, 63\.7\]50\.0 \[45\.0, 55\.0\]Gemma\-3\-4B10%10\\%358\.8 \[48\.8, 68\.8\]48\.8 \[42\.5, 53\.8\]52\.5 \[42\.5, 62\.5\]51\.2 \[46\.2, 56\.2\]20%20\\%753\.8 \[43\.8, 63\.7\]48\.8 \[42\.5, 53\.8\]56\.2 \[48\.8, 63\.7\]48\.8 \[43\.8, 53\.8\]30%30\\%1053\.8 \[43\.8, 62\.5\]48\.8 \[42\.5, 55\.0\]55\.0 \[47\.5, 62\.5\]50\.0 \[45\.0, 55\.0\]40%40\\%1452\.5 \[42\.5, 62\.5\]48\.8 \[43\.8, 53\.8\]47\.5 \[38\.7, 56\.2\]48\.8 \[43\.8, 53\.8\]50%50\\%1752\.5 \[43\.8, 61\.3\]47\.5 \[41\.2, 53\.8\]48\.8 \[40\.0, 56\.2\]51\.2 \[46\.2, 56\.2\]60%60\\%2050\.0 \[41\.2, 58\.8\]51\.2 \[45\.0, 57\.5\]46\.2 \[38\.7, 53\.8\]52\.5 \[46\.2, 58\.8\]70%70\\%2451\.2 \[40\.0, 62\.5\]48\.8 \[43\.8, 53\.8\]51\.2 \[41\.2, 61\.3\]50\.0 \[46\.2, 53\.8\]80%80\\%2753\.8 \[45\.0, 62\.5\]51\.2 \[46\.2, 56\.2\]56\.2 \[48\.8, 63\.7\]46\.2 \[41\.2, 50\.0\]90%90\\%3153\.8 \[45\.0, 62\.5\]48\.8 \[42\.5, 53\.8\]53\.8 \[47\.5, 60\.0\]48\.8 \[46\.2, 50\.0\]100%100\\%3453\.8 \[45\.0, 62\.5\]48\.8 \[43\.8, 53\.8\]47\.5 \[41\.2, 53\.8\]50\.0 \[50\.0, 50\.0\]Gemma\-3\-12B10%10\\%557\.5 \[47\.5, 67\.5\]55\.0 \[48\.8, 62\.5\]50\.0 \[42\.5, 57\.5\]48\.8 \[45\.0, 52\.5\]20%20\\%1052\.5 \[42\.5, 62\.5\]55\.0 \[48\.8, 62\.5\]48\.8 \[40\.0, 57\.5\]51\.3 \[47\.5, 55\.0\]30%30\\%1455\.0 \[45\.0, 65\.0\]55\.0 \[50\.0, 61\.3\]50\.0 \[42\.5, 57\.5\]47\.5 \[43\.8, 50\.0\]40%40\\%1953\.8 \[43\.8, 63\.7\]55\.0 \[50\.0, 61\.3\]48\.8 \[40\.0, 56\.2\]50\.0 \[45\.0, 55\.0\]50%50\\%2455\.0 \[45\.0, 65\.0\]53\.8 \[47\.5, 60\.0\]51\.2 \[46\.2, 56\.2\]50\.0 \[46\.2, 53\.8\]60%60\\%2957\.5 \[46\.2, 67\.5\]58\.8 \[52\.5, 66\.2\]57\.5 \[50\.0, 65\.0\]53\.8 \[50\.0, 58\.7\]70%70\\%3453\.8 \[43\.8, 63\.7\]52\.5 \[46\.2, 60\.0\]51\.2 \[43\.8, 58\.8\]46\.2 \[41\.2, 50\.0\]80%80\\%3855\.0 \[45\.0, 65\.0\]56\.2 \[50\.0, 62\.5\]50\.0 \[43\.8, 56\.2\]52\.5 \[50\.0, 56\.2\]90%90\\%4355\.0 \[45\.0, 65\.0\]55\.0 \[48\.8, 62\.5\]48\.8 \[43\.8, 52\.5\]48\.8 \[45\.0, 52\.5\]100%100\\%4852\.5 \[43\.8, 61\.3\]57\.5 \[51\.2, 63\.7\]45\.0 \[40\.0, 48\.8\]52\.5 \[50\.0, 56\.2\]Qwen3\-4B10%10\\%441\.2 \[32\.5, 50\.0\]51\.2 \[45\.0, 57\.5\]42\.5 \[33\.8, 50\.0\]51\.2 \[47\.5, 55\.0\]20%20\\%752\.5 \[42\.5, 62\.5\]52\.5 \[46\.2, 58\.8\]53\.8 \[45\.0, 62\.5\]50\.0 \[46\.2, 53\.8\]30%30\\%1158\.8 \[48\.8, 68\.8\]51\.2 \[46\.2, 56\.2\]63\.7 \[56\.2, 71\.2\]46\.2 \[41\.2, 50\.0\]40%40\\%1446\.2 \[36\.2, 56\.2\]51\.2 \[45\.0, 57\.5\]55\.0 \[46\.2, 63\.7\]46\.2 \[41\.2, 50\.0\]50%50\\%1843\.8 \[35\.0, 52\.5\]55\.0 \[48\.8, 62\.5\]45\.0 \[36\.3, 53\.8\]52\.5 \[47\.5, 57\.5\]60%60\\%2251\.2 \[41\.2, 61\.3\]50\.0 \[43\.8, 56\.2\]55\.0 \[46\.2, 63\.7\]50\.0 \[43\.8, 56\.2\]70%70\\%2547\.5 \[37\.5, 57\.5\]53\.8 \[47\.5, 60\.0\]50\.0 \[42\.5, 57\.5\]52\.5 \[50\.0, 56\.2\]80%80\\%2948\.8 \[40\.0, 58\.7\]52\.5 \[46\.2, 58\.8\]51\.2 \[43\.8, 58\.8\]52\.5 \[50\.0, 56\.2\]90%90\\%3247\.5 \[38\.8, 56\.2\]52\.5 \[46\.2, 58\.7\]52\.5 \[50\.0, 56\.2\]48\.8 \[46\.2, 50\.0\]100%100\\%3650\.0 \[41\.2, 58\.8\]52\.5 \[46\.2, 58\.7\]51\.2 \[47\.5, 55\.0\]50\.0 \[50\.0, 50\.0\]Qwen3\-14B10%10\\%457\.5 \[45\.0, 70\.0\]52\.5 \[45\.0, 60\.0\]48\.8 \[41\.2, 56\.2\]51\.2 \[47\.5, 55\.0\]20%20\\%855\.0 \[43\.8, 66\.3\]53\.8 \[47\.5, 60\.0\]55\.0 \[46\.2, 63\.7\]51\.2 \[47\.5, 55\.0\]30%30\\%1256\.2 \[43\.8, 68\.8\]55\.0 \[48\.8, 62\.5\]55\.0 \[47\.5, 62\.5\]51\.3 \[47\.5, 56\.2\]40%40\\%1653\.8 \[42\.5, 65\.0\]52\.5 \[46\.2, 58\.7\]45\.0 \[37\.5, 51\.2\]50\.0 \[46\.2, 53\.8\]50%50\\%2055\.0 \[43\.8, 66\.2\]52\.5 \[46\.2, 58\.7\]50\.0 \[43\.8, 56\.2\]46\.2 \[41\.2, 50\.0\]60%60\\%2453\.8 \[42\.5, 65\.0\]53\.8 \[47\.5, 60\.0\]48\.8 \[43\.8, 53\.8\]50\.0 \[46\.2, 53\.8\]70%70\\%2858\.8 \[46\.2, 70\.0\]52\.5 \[46\.2, 58\.8\]52\.5 \[46\.2, 58\.8\]50\.0 \[46\.2, 53\.8\]80%80\\%3255\.0 \[45\.0, 65\.0\]52\.5 \[46\.2, 58\.8\]50\.0 \[43\.8, 56\.2\]51\.2 \[50\.0, 53\.8\]90%90\\%3657\.5 \[46\.2, 68\.8\]52\.5 \[46\.2, 58\.8\]55\.0 \[50\.0, 61\.3\]48\.8 \[46\.2, 50\.0\]100%100\\%4055\.0 \[43\.8, 66\.2\]52\.5 \[46\.2, 58\.7\]51\.2 \[45\.0, 57\.5\]50\.0 \[50\.0, 50\.0\]

### D\.2Detailed Response\-Quality Results

Table 7:Blind response\-quality comparison\.Number of model–depth settings in whichAIMESattains a higher mean score than each comparator\. Overall quality is defined asQ=\(C\+F\+R\)/3Q=\(C\+F\+R\)/3, whereCC,FF, andRRdenote coherence, fluency, and relevance\. Counts are descriptive rather than significance tests\.We evaluate whether the multi\-value control gains ofAIMESare associated with changes in generation quality using a blindGPT\-5\.6\-Soljudge\. Each response is scored for coherence \(CC\), fluency \(FF\), and relevance \(RR\) on a four\-point scale, with overall quality defined asQ=\(C\+F\+R\)/3Q=\(C\+F\+R\)/3\. Table[7](https://arxiv.org/html/2609.30405#A4.T7)summarizes the number of model–depth–objective settings in whichAIMESattains a higher mean score than each comparator\. Across the 150 settings,AIMESexceeds Fixed Multi\-Value Steering onQQin7676settings and Prompt Steering in8181, indicating broadly comparable overall response quality rather than a consistent quality advantage\. For relevance,AIMESexceeds Prompt Steering in106/150106/150settings \(70\.7%70\.7\\%\)\. Taken together, these results do not indicate a systematic loss in response quality under adaptive steering\.

### D\.3Layer\-wise Within\-Response Steering Variation

![Refer to caption](https://arxiv.org/html/2609.30405v1/Figures/controller_adaptation.png)Figure 11:Variation inAIMESsteering strength across intervention depths\.For each model and objective, we measure the within\-response range of the target\-value steering coefficients,Rα,k=maxt⁡αk,t−mint⁡αk,tR\_\{\\alpha,k\}=\\max\_\{t\}\\alpha\_\{k,t\}\-\\min\_\{t\}\\alpha\_\{k,t\}\. Each point shows the mean coefficient range across target\-disjoint responses at the corresponding model\-specific intervention layer, and shaded regions denote95%95\\%prompt\-bootstrap confidence intervals\. Larger values indicate greater changes in steering strength during generation\. Results are reported separately at each of the ten sampled depths; no averaging is performed across intervention layers\.A defining feature ofAIMESis that its steering coefficients are updated during generation using the online value readout, rather than remaining fixed for the entire response\. We therefore examine how much these coefficients vary within a response across the full intervention\-depth grid\.

#### Metric\.

For each controlled valueVkV\_\{k\}, we measure the within\-response coefficient range

Rα,k=maxt⁡αk,t−mint⁡αk,t,R\_\{\\alpha,k\}=\\max\_\{t\}\\alpha\_\{k,t\}\-\\min\_\{t\}\\alpha\_\{k,t\},\(26\)whereαk,t\\alpha\_\{k,t\}is the steering coefficient applied to valueVkV\_\{k\}at decoding steptt\. For objectives involving multiple target values, we first averageRα,kR\_\{\\alpha,k\}across the controlled values within each response and then average across target\-disjoint prompts\. Thus, larger values indicate greater changes in steering strength over the course of generation, while smaller values indicate more nearly constant steering\. The analysis is performed independently at each sampled intervention layer; no averaging is performed across layers\.

#### Depth\-wise analysis\.

Figure[11](https://arxiv.org/html/2609.30405#A4.F11)reports the resulting coefficient variation across all ten sampled intervention depths \(10%10\\%–100%100\\%\) for the three steering objectives\. Across all five models, the coefficient range remains clearly nonzero throughout the network, showing thatAIMEScontinues to adjust its steering strength during generation rather than reducing to a fixed\-coefficient intervention\. The magnitude of this variation depends on the model, objective, and intervention depth, as expected from a controller whose updates are driven by the model\-specific online value state\.

The same qualitative behavior is observed across both two\-value objectives,\(↑Care↑Fairness\)\(\\uparrow\\mathrm\{Care\}\\;\\uparrow\\mathrm\{Fairness\}\)and\(↑Loyalty↑Authority\)\(\\uparrow\\mathrm\{Loyalty\}\\;\\uparrow\\mathrm\{Authority\}\), as well as the competing three\-value objective\(↑Care↑Fairness↓Sanctity\)\(\\uparrow\\mathrm\{Care\}\\;\\uparrow\\mathrm\{Fairness\}\\;\\downarrow\\mathrm\{Sanctity\}\)\. Across the plotted conditions, mean coefficient ranges are typically substantial, with values approximately spanning0\.450\.45–0\.840\.84\. Importantly, this variation is not confined to a single model family or a narrow depth region: all five models exhibit clear within\-response changes across early, middle, and later intervention layers\.

The representative30%30\\%depth reported in the main text is therefore not an isolated case\. Rather, it is part of a broader depth\-wise pattern in which the adaptive coefficients remain responsive throughout generation\. At this depth, the coefficient range is consistently nonzero across all models and objectives, while the full analysis here shows that this behavior persists across the complete set of sampled intervention layers\.

Overall, the layer\-wise results confirm that the adaptive component ofAIMESremains active across the network: its steering coefficients change meaningfully within responses across models, objectives, and intervention depths, rather than behaving as approximately fixed steering weights\.

### D\.4Independent\-Judge\-based Analysis of Depth\-wise and GeoGain Results

#### Depth\-wise interaction contrasts\.

Figure[12](https://arxiv.org/html/2609.30405#A4.F12)reports the Claude\-based depth\-wise advantage ofAIMESusing the same interaction contrasts as Figure[3](https://arxiv.org/html/2609.30405#S4.F3)\. The Care/Fairness\-retention advantage over Fixed Multi\-Value Steering,DxD\_\{x\}, reproduces the strongest depth\-specific pattern observed under GPT\-5\.6\-Sol: at30%30\\%depth,Dx\>0D\_\{x\}\>0for all five models under both judges\. The corresponding comparison against Prompt Steering also exhibits 5/5 positive agreement at30%30\\%depth, although the additional depths identified as unanimous under GPT\-5\.6\-Sol do not all remain unanimous under Claude\. For example, at80%80\\%depth, Llama\-3\.1\-8B changes sign under Claude \(Dx=−0\.113D\_\{x\}=\-0\.113\)\.

![Refer to caption](https://arxiv.org/html/2609.30405v1/claude_Dx_Dy_depth_consistency_10depth.png)Figure 12:Claude\-Opus\-4\.8 judge\-based analysis of the depth\-wiseAIMESadvantage\.The analysis uses the same ten sampled intervention depths and the sameDx/DyD\_\{x\}/D\_\{y\}definitions as Figure[3](https://arxiv.org/html/2609.30405#S4.F3), but re\-scores the identical generations with Claude\-Opus\-4\.8\. Positive values indicate an advantage forAIMES; black outlines denote depths at which the advantage is positive for all five models\.The Sanctity\-suppression contrast,DyD\_\{y\}, shows greater evaluator sensitivity than the Care/Fairness\-retention contrast\. Against Fixed Multi\-Value Steering, the depth at which the advantage concentrates differs between evaluators: under GPT\-5\.6\-Sol, the strongest consensus occurs near60%60\\%depth, whereas under Claude, the strongest consensus for this comparator occurs near10%10\\%depth \(Figure[12](https://arxiv.org/html/2609.30405#A4.F12)\(c\)\)\. By contrast, the comparison with Prompt Steering is more consistent across evaluators, remaining broadly positive with four of five models positive at several depths under both judges, though without a fully unanimous column under either\. Overall, the full\-resolution analysis indicates that the30%30\\%Care/Fairness\-retention advantage over Fixed Multi\-Value Steering is the most evaluator\-robust depth\-specific finding in this analysis; the Sanctity\-suppression advantage over Fixed Multi\-Value Steering is depth\-dependent in a way that does not consistently localize to the same network region across evaluators, and we scope our main\-text claim about this comparison accordingly \(Section[4](https://arxiv.org/html/2609.30405#S4)\)\.

Figure 13:Claude\-Opus\-4\.8 judge\-based analysis of GeoGain for the competing three\-value objective\.Prompt\-paired GeoGain advantage ofAIMESover Fixed Multi\-Value Steering \(left\) and Prompt Steering \(right\) across all ten sampled intervention depths for\(↑Care↑Fairness↓Sanctity\)\(\\uparrow\\mathrm\{Care\}\\;\\uparrow\\mathrm\{Fairness\}\\;\\downarrow\\mathrm\{Sanctity\}\)\. Positive values indicate higher GeoGain metric forAIMESvs\. Fixed Multi\-Value Steering and Prompt Steering\. Under Claude\-Opus\-4\.8, 30/50 model\-depth settings favorAIMESover Fixed Multi\-Value Steering and 50/50 favorAIMESover Prompt Steering\.
#### GeoGain Analysis\.

We next recompute the prompt\-level GeoGain analysis using Claude scores over the complete ten\-depth grid\. Table[8](https://arxiv.org/html/2609.30405#A4.T8)summarizes the number of model–depth cells in whichAIMESattains higher GeoGain than each comparator\. Across the three objectives and two comparators, this yields3×2×50=3003\\times 2\\times 50=300model–depth comparisons\.

Relative to Fixed Multi\-Value Steering,AIMESattains higher GeoGain in 35/50 settings for\(↑Care↑Fairness\)\(\\uparrow\\mathrm\{Care\}\\;\\uparrow\\mathrm\{Fairness\}\), 29/50 for\(↑Loyalty↑Authority\)\(\\uparrow\\mathrm\{Loyalty\}\\;\\uparrow\\mathrm\{Authority\}\), and 30/50 for\(↑Care↑Fairness↓Sanctity\)\(\\uparrow\\mathrm\{Care\}\\;\\uparrow\\mathrm\{Fairness\}\\;\\downarrow\\mathrm\{Sanctity\}\)\. The latter exactly matches the 30/50 positive model–depth count obtained with GPT\-5\.6\-Sol, providing independent\-judge support for the model\- and depth\-dependent GeoGain advantage over Fixed Multi\-Value Steering\. Importantly, the 30/50 count understates the asymmetry in effect magnitude: on the three\-value objective, cells in whichAIMEStrails Fixed Multi\-Value Steering differ by at most approximately−0\.001\-0\.001in GeoGain, whereas favorable cells reach advantages as large as approximately\+0\.065\+0\.065\. Thus, the non\-positive cells are predominantly near ties, while several of the positive cells exhibit substantially larger margins\.

Table 8:Claude\-Opus\-4\.8 GeoGain Analysis\.Number of model–depth cells, out of 50 \(five models×\\timesten depths\), in whichAIMESattains higher prompt\-level GeoGain than each comparator\.The comparison with Prompt Steering is strongly objective\-dependent\. For\(↑Care↑Fairness\)\(\\uparrow\\mathrm\{Care\}\\;\\uparrow\\mathrm\{Fairness\}\),AIMEShas lower GeoGain in all 50 model–depth settings, whereas for\(↑Loyalty↑Authority\)\(\\uparrow\\mathrm\{Loyalty\}\\;\\uparrow\\mathrm\{Authority\}\), it is higher in 21/50 settings\. In contrast, for the competing three\-value objective\(↑Care↑Fairness↓Sanctity\)\(\\uparrow\\mathrm\{Care\}\\;\\uparrow\\mathrm\{Fairness\}\\;\\downarrow\\mathrm\{Sanctity\}\),AIMESattains higher GeoGain than Prompt Steering in50/50model–depth settings\. This strengthens the near\-universal advantage observed with GPT\-5\.6\-Sol, where 47/50 settings favoredAIMES, and shows that the principal balanced\-control result persists under an independent evaluator across the full depth grid\. Because GeoGain is positive only when all requested target directions improve simultaneously, the 50/50AIMESadvantage on the three\-value objective indicates more favorable joint satisfaction of the competing target directions under Claude, rather than simply larger movement along any individual value dimension\.

The model\-wise curves further illustrate this reversal\. Gemma\-3\-12B exhibits the largestAIMESadvantage over Prompt Steering on the three\-value objective, approximately\+0\.06\+0\.06to\+0\.085\+0\.085across depth, despiteAIMESbeing disadvantaged for this model on the simpler two\-value objectives\. This suggests that the relative advantage emerges specifically when an additional competing constraint must be satisfied jointly, rather than from uniformly stronger steering on individual target values\.

Taken together, the Claude judge\-based analysis supports the main GeoGain conclusion while revealing greater evaluator sensitivity in the finer\-grained depth\-wise decomposition\. In particular, the30%30\\%Care/Fairness\-retention effect and the balanced\-control advantage on the three\-value objective persist across evaluators, whereas the precise depth at which Sanctity\-suppression advantages emerge is less stable\.

相似文章

通过定向干预实现语言模型的多属性引导

arXiv cs.CL

MAT-Steer 提出了一种新颖的推理时干预框架,通过学习稀疏且正交的引导向量,选择性干预与每个属性相关的标记,从而在多属性冲突场景下引导大型语言模型,在问答任务和生成任务上的表现均优于现有方法。

驾驭语言轴:从线性可解码性到因果控制

arXiv cs.CL

本文研究LLM中的语言身份是否可以通过紧凑的激活方向进行线性解码和因果控制。通过在多个模型家族上进行引导(steering)和消融(ablation)实验,作者表明语言选择具有方向依赖性、层特异性,并且在语言信号被消融时会恢复为英语。