Read-Best Is Not Steer-Best: A Probing--Steering Layer Dissociation in Omni-Modal Large Language Models

arXiv cs.CL Papers

Summary

This paper reveals that the optimal layer for linear probing to read concepts differs from the optimal layer for activation steering in omni-modal large language models, challenging common heuristics in representation engineering.

arXiv:2609.22135v1 Announce Type: new Abstract: Omni-modal large language models integrate text, audio, and image signals into a shared residual stream, where concepts such as emotion can be linearly decoded and causally modified by activation steering. A common but rarely tested assumption is that the layer with the highest probing accuracy is also the best layer for steering, so injection layers are often selected by probe performance. We provide the first causal test of this assumption across three independently developed omni-modal models and find that it fails. Reading and intervention rely on different layers, a phenomenon we call the probing-steering layer dissociation. Using emotion as a controlled testbed, we measure layer-wise readability and steerability across text, audio, and image inputs. Probe-best layers vary widely across architectures, while steering-effective layers consistently fall within a narrow mid-to-late range of normalized depth. Paired random-direction controls show an approximately 26-fold causal gap, ruling out random perturbation and direction quality as explanations. Logit-lens analysis reveals a staged forward process: causal handle, probing saturation, and vocabulary commitment, and motivates a two-factor account in which steering effectiveness depends on both representational readability and downstream plasticity. These results show that probing accuracy is a poor heuristic for selecting intervention layers and suggest a cross-architecture mid-to-late selection criterion. We also identify a cross-modal emotion subspace organized by valence and arousal, with joy acting as a stable anchor across models. Code and data: https://github.com/YiboWang2002/Read-Best-Is-Not-Steer-Best.
Original Article
View Cached Full Text

Cached at: 09/22/26, 09:03 AM

# Read-Best Is Not Steer-Best: A Probing–Steering Layer Dissociation in Omni-Modal Large Language Models
Source: [https://arxiv.org/html/2609.22135](https://arxiv.org/html/2609.22135)
Jisheng DangBimei WangYitao WuWencan ZhangHong PengJizhao LiuBin HuQi Tianand Tat\-Seng ChuaThanks:Yibo Wang, Jisheng Dang, Bimei Wang, Hong Peng, Jizhao Liu, and Bin Hu are with Lanzhou University, Lanzhou, China\.Thanks:Yitao Wu is with Hainan University, Haikou, China\.Thanks:Qi Tian is with Cloud and AI BU, Huawei, Shenzhen, Guangdong 518129, China \(e\-mail: tian\.qi1@huawei\.com\)\.Thanks:Wencan Zhang and Tat\-Seng Chua are with the School of Computing, National University of Singapore, Singapore\.Thanks:Corresponding authors: Bin Hu, Jisheng Dang, and Hong Peng\.

###### Abstract

Omni\-modal large language models fold text, audio, and image signals into a single residual stream, where the emotion carried by an image or a voice can be linearly read out and causally rewritten by activation steering\. Practice, however, rests on a rarely tested assumption: the layer where a linear probe reads a concept most strongly is also the layer where injecting that concept steers behavior most effectively, so injection layers are chosen by probing accuracy\. We give the first causal test of this assumption across three independently built omni\-modal models\. The assumption fails\. Reading and intervention are distinct operations on the residual stream, carried by different layers, a phenomenon we call the probing–steering layer dissociation\. We use emotion as a controlled vehicle\. It is linearly readable across the three modalities and steerable by directional injection, so readability and steerability become two layer\-wise curves in one representation space\. The probe\-best layer scatters across nearly the full network depth and shifts with architecture, whereas the steering\-effective layer does not\. It is an architectural invariant, landing in the same narrow mid\-to\-late band of normalized depth in every model\. With paired random\-direction controls, the causal gap between the two is roughly twenty\-six\-fold, ruling out direction quality and random perturbation\. A logit\-lens analysis ties the dissociation to a staged forward pass \(a causal\-handle, probing\-saturation, vocabulary\-commitment ordering\) and motivates a two\-factor account that links steering effectiveness to representational readability and downstream plasticity\. Choosing the injection layer by probing accuracy is a mistaken heuristic, and the mid\-to\-late band gives a cross\-architecture selection criterion\. As a by\-product, we map a cross\-modal emotion subspace organized by valence–arousal and anchored across models by joy\. Code and data are available at[https://github\.com/YiboWang2002/Read\-Best\-Is\-Not\-Steer\-Best](https://github.com/YiboWang2002/Read-Best-Is-Not-Steer-Best)\.

###### Index Terms:

Affective Computing, Emotion Recognition, Multimodal Large Language Models, Mechanistic Interpretability, Activation Steering, Representation Engineering, Linear Probing\.

## IIntroduction

Omni\-modal large language models process text, audio, and image signals in one shared residual stream, and two operations dominate how we read and change the concepts these signals carry\. A linear probe\[[2](https://arxiv.org/html/2609.22135#bib.bib13)\]*reads*the concept encoded in a layer’s activations, while activation steering\[[67](https://arxiv.org/html/2609.22135#bib.bib1),[54](https://arxiv.org/html/2609.22135#bib.bib3)\]*edits*those activations to change behavior\. In practice the two are tied together by an assumption that is rarely stated and almost never tested: the layer where a probe reads a concept most strongly is taken to be the layer where injecting it steers behavior most effectively\. Injection layers are then chosen directly by probing accuracy\. If the assumption is wrong, much of the steering literature has been intervening at causally near\-inert layers, and underestimating what these methods can do\.

We show that this assumption fails in omni\-modal language models, and we measure how far\. Reading and intervention are distinct operations carried by different layers of the residual stream \([Figure1](https://arxiv.org/html/2609.22135#S1.F1)\)\. The separation is systematic, and the layer\-selection heuristic at the heart of representation engineering does not transfer to these models\.

![Refer to caption](https://arxiv.org/html/2609.22135v1/fig1.png)Fig\. 1:Read\-best is not steer\-best\. In an omni\-modal LLM, probe readability \(blue\) peaks at a late layer while the steering effect of an emotion direction \(red\) peaks earlier, in the mid\-to\-late zone\. Injecting at the read\-best layer is about26×26\\timesweaker than in the steering zone\. Curves are schematic\. Full results in[Figure3](https://arxiv.org/html/2609.22135#S1.F3)and §5\.![Refer to caption](https://arxiv.org/html/2609.22135v1/main1.png)Fig\. 2:Overview of the study\. Text, audio, and image inputs enter an omni\-modal LLM through a shared residual stream\. A linear probe reads a concept best at a late layer \(probe\-best\), while steering an emotion direction is most effective at a mid\-to\-late layer \(steer\-best\), a controlled gap of about26×26\\times\. A logit\-lens analysis orders the forward pass as causal handle, probing saturation, then vocabulary commitment, placing the steering window before commitment and motivating the two\-factor modelHthreshH\_\{\\mathrm\{thresh\}\}\. Bottom: \(A\) the steer\-best layer clusters in a shared band across three architectures while the probe\-best layer scatters; \(B\) the logit\-lens mechanism; \(C\) controls and external judges \(VAD\-BERT, GPT\-5\.5\)\.Fig\. 3:The probing–steering layer dissociation at a glance\.\(a\)On Qwen2\.5\-Omni\-7B the probe\-accuracy peak \(L27\) and the steering\-effect peak \(L20\) are 7 layers apart, with the band marking the mid\-to\-late steering zone \[L14, L22\]\.\(b\)Across three architectures the steer\-best layer stays in a narrow band of normalized depth\[0\.48,0\.74\]\[0\.48,0\.74\], while the probe\-best layer scatters across\[0\.29,1\.00\]\[0\.29,1\.00\]\(MiniCPM shows a reversed ordering\)\. Full evidence in §5–§6\.We study this probing–steering layer dissociation using emotion as a controlled vehicle, across three independently designed omni\-modal LLMs \(Qwen2\.5\-Omni\-7B, Phi\-4\-Multimodal, MiniCPM\-o\-4\.5\)\. Emotion is well suited to the task: it is shared across text, audio, and image, reads out linearly, and responds to directional injection, so we can measure readability and steerability as two layer\-wise curves in the same representation space and compare where they peak\.[Figure2](https://arxiv.org/html/2609.22135#S1.F2)gives an overview of the study, from the omni\-modal setup to the cross\-architecture findings, mechanism, and controls\.

The picture is consistent across the three architectures: the probe\-best layer scatters across nearly the full network depth while the steer\-best layer settles into a narrow mid\-to\-late band \([Figure3](https://arxiv.org/html/2609.22135#S1.F3)\), and on Qwen2\.5\-Omni single\-layer injection at the probe\-best layer is indistinguishable from a random direction\. A logit\-lens analysis locates the cause in the staged forward pass, in the window where a representation is already stable but not yet committed to a token\.

We make four contributions:

- •We provide the first causal characterization of the probing–steering dissociation in the omni\-modal setting, across three independently designed omni\-modal LLMs spanning text, audio, and image\. The steering\-effective zone is an architectural invariant\. It stays in the same mid\-to\-late band in every model, whereas the probe\-best layer varies across almost the full network depth\. We frame this stability as an architecture\-modulated regularity rather than a universal law \(§7\.2\)\.
- •We quantify the dissociation with a controlled causal gap of roughly 26×\\times, using paired random\-direction controls to rule out direction quality and random perturbation\.
- •We trace the dissociation to a logit\-lens forward\-pass account \(a causal handle, then probing saturation, then vocabulary commitment\) and propose a two\-factor account,HthreshH\_\{\\mathrm\{thresh\}\}, that relates steering effectiveness to representational readability and downstream plasticity, as a first step toward predicting the intervention\-optimal layer \(§5\.5\)\.
- •We draw the direct consequence for representation engineering\. Choosing the injection layer by probing accuracy is a mistaken heuristic, and the mid\-to\-late zone gives a cross\-architecture selection criterion\. As a substrate for these mechanisms, we also characterize a cross\-modal emotion geometry, a shared subspace organized by valence–arousal and anchored across models by joy\.

§3 describes the models, data, and protocol\. §4 characterizes the cross\-modal emotion geometry\. §5 presents the core dissociation evidence\. §6 extends it to tri\-modal, cross\-model causal control\. §7 discusses the implications and the geometric basis\.

## IIRelated Work

### II\-ARepresentation Engineering and Activation Steering

Representation engineering, introduced by Zou et al\.\[[67](https://arxiv.org/html/2609.22135#bib.bib1)\], treats high\-level concepts as linear directions in activation space and intervenes on behavior by editing the residual stream\. A family of methods in this vein extracts concept directions from contrastive activations and injects them additively, covering truthfulness, general behavior steering, and style or emotion control\[[27](https://arxiv.org/html/2609.22135#bib.bib2),[54](https://arxiv.org/html/2609.22135#bib.bib3),[41](https://arxiv.org/html/2609.22135#bib.bib4),[23](https://arxiv.org/html/2609.22135#bib.bib5)\], with roots in controllable text generation\[[10](https://arxiv.org/html/2609.22135#bib.bib50),[22](https://arxiv.org/html/2609.22135#bib.bib51),[24](https://arxiv.org/html/2609.22135#bib.bib52)\]and extensions to function vectors, in\-context vectors, and refusal or conditional steering\[[53](https://arxiv.org/html/2609.22135#bib.bib7),[33](https://arxiv.org/html/2609.22135#bib.bib8),[4](https://arxiv.org/html/2609.22135#bib.bib9),[25](https://arxiv.org/html/2609.22135#bib.bib11)\]\. Almost all of these treat steering as*single\-layer*injection with the layer chosen by probing signal or empirical search, tacitly assuming that where a concept can be read is where it can be intervened upon\. Closest to our work, Tan et al\.\[[51](https://arxiv.org/html/2609.22135#bib.bib10)\]note that steering is highly sensitive to the injection layer and setup, but do not characterize the systematic read\-best/steer\-best separation as a phenomenon in its own right; that separation is the entry point of this work\.

### II\-BMechanistic Interpretability of Layer\-Wise Function

Mechanistic\-interpretability research has characterized how information forms progressively along a transformer’s layers\. The logit lens\[[40](https://arxiv.org/html/2609.22135#bib.bib57)\]projects each layer’s hidden state onto the vocabulary and the tuned lens\[[6](https://arxiv.org/html/2609.22135#bib.bib12)\]calibrates this readout, while linear probes are the standard tool for reading intermediate layers\[[2](https://arxiv.org/html/2609.22135#bib.bib13),[19](https://arxiv.org/html/2609.22135#bib.bib14),[5](https://arxiv.org/html/2609.22135#bib.bib15)\]\. At the circuit level, FFN key\-value memories\[[16](https://arxiv.org/html/2609.22135#bib.bib19),[15](https://arxiv.org/html/2609.22135#bib.bib20)\], factual\-association editing\[[36](https://arxiv.org/html/2609.22135#bib.bib17)\], circuit analyses\[[57](https://arxiv.org/html/2609.22135#bib.bib16),[9](https://arxiv.org/html/2609.22135#bib.bib18)\], and sparse\-autoencoder studies support a picture in which concepts exist as decomposable directions, aligning with the linear representation hypothesis\[[43](https://arxiv.org/html/2609.22135#bib.bib21),[42](https://arxiv.org/html/2609.22135#bib.bib22)\]and superposition\[[14](https://arxiv.org/html/2609.22135#bib.bib23)\]\. These works focus on where information is*represented*; whether that locus coincides with the locus of causal*intervenability*is exactly what our probing–steering dissociation characterizes\.

### II\-CEmotion Representation in Multimodal LLMs

Omni\-modal LLMs process text, audio, and image in a unified residual stream\. The three models we study, Qwen2\.5\-Omni, Phi\-4\-Multimodal\[[1](https://arxiv.org/html/2609.22135#bib.bib24)\], and MiniCPM\-o\[[64](https://arxiv.org/html/2609.22135#bib.bib25)\], are of this kind, with a lineage traceable to multimodal architectures\[[7](https://arxiv.org/html/2609.22135#bib.bib26),[32](https://arxiv.org/html/2609.22135#bib.bib28),[3](https://arxiv.org/html/2609.22135#bib.bib29),[26](https://arxiv.org/html/2609.22135#bib.bib30)\]and modality encoders\[[45](https://arxiv.org/html/2609.22135#bib.bib31),[65](https://arxiv.org/html/2609.22135#bib.bib32),[46](https://arxiv.org/html/2609.22135#bib.bib34),[17](https://arxiv.org/html/2609.22135#bib.bib33)\]\. In the image processing literature, emotion and affect have long been read directly from visual signals, through visual emotion analysis and emotion distribution learning\[[60](https://arxiv.org/html/2609.22135#bib.bib59),[63](https://arxiv.org/html/2609.22135#bib.bib60),[62](https://arxiv.org/html/2609.22135#bib.bib61)\], facial expression and micro\-expression recognition\[[56](https://arxiv.org/html/2609.22135#bib.bib58),[58](https://arxiv.org/html/2609.22135#bib.bib62),[30](https://arxiv.org/html/2609.22135#bib.bib63)\], and personality\-aware image aesthetics\[[28](https://arxiv.org/html/2609.22135#bib.bib64)\], while audio\-visual correspondence modeling\[[38](https://arxiv.org/html/2609.22135#bib.bib65)\]and image captioning\[[21](https://arxiv.org/html/2609.22135#bib.bib66),[31](https://arxiv.org/html/2609.22135#bib.bib67)\]connect visual signals to attention and text generation; omni\-modal LLMs fold these capabilities into a single residual stream, which is exactly the setting we probe\. In psychology, the dimensional organization of emotion is described by Russell’s circumplex model\[[48](https://arxiv.org/html/2609.22135#bib.bib46),[44](https://arxiv.org/html/2609.22135#bib.bib47)\]and the valence\-arousal\-dominance framework\[[35](https://arxiv.org/html/2609.22135#bib.bib49),[39](https://arxiv.org/html/2609.22135#bib.bib48)\]\. A recent line of work examines emotion representation inside LLMs directly, characterizing the emotional latent space, localizing emotion\-inference mechanisms, and discovering or controlling emotion circuits in text\-only models\[[47](https://arxiv.org/html/2609.22135#bib.bib41),[50](https://arxiv.org/html/2609.22135#bib.bib43),[55](https://arxiv.org/html/2609.22135#bib.bib44),[13](https://arxiv.org/html/2609.22135#bib.bib42)\]\. Unlike this line, we use emotion as a controlled vehicle to reveal the reading–intervention layer separation and test its generality across three omni\-modal architectures, moving past whether emotion representations exist or can be controlled within a single text modality\.

## IIIMethod

This section describes the experimental setup shared by §4–§6: models, data, activation processing and probing, the steering injection protocol, and evaluation\. All experiments follow pre\-registered decision thresholds, multi\-seed bootstrap confidence intervals, and paired random\-direction controls\. Results that do not meet a pre\-registered threshold are reported faithfully as findings rather than retried\.

### III\-AModels

We study three independently designed omni\-modal LLMs\. Qwen2\.5\-Omni\-7B is the primary model\. Its LLM backbone \(based on Qwen2\.5\[[59](https://arxiv.org/html/2609.22135#bib.bib27)\], the Thinker\) has 28 layers and hidden dimension 3584, with hook pathmodel\.model\.layers\[0\.\.27\]\. For cross\-model validation we use Phi\-4\-Multimodal\-Instruct \(Phi\-4\-Mini 3\.8B LLM, 32 layers, hidden 3072, with a SigLIP\-400M vision encoder, a 24\-layer Conformer audio encoder, and a Mixture\-of\-LoRAs adapter\) and MiniCPM\-o\-4\.5 \(36 layers, hidden 4096\)\. The three models differ in LLM backbone, tokenizer, and modality\-encoder lineage, so conclusions that hold across them are not easily attributed to shared components\. All analyses are performed at inference time, with no model weights updated\.

### III\-BData

The main analysis uses a self\-constructed, fixed tri\-modal sample set, M3: 5 emotions \(anger / calm / fear / joy / sadness\)×\\times3 modalities \(text / audio / image\)×\\times25 samples per class, for 375 samples in total \(configuration inconfigs/multimodal\_m3\_samples\.json\)\. A fixed sample set ensures comparability across modalities, layers, and seeds\. Probe training and cross\-modal transfer are carried out on M3\.

To support external controls and scaled\-up evaluation, we additionally use three public datasets: GoEmotions \(text\)\[[11](https://arxiv.org/html/2609.22135#bib.bib35)\], RAVDESS speech\-only \(audio\)\[[34](https://arxiv.org/html/2609.22135#bib.bib36)\], and EmoSet \(image\)\[[61](https://arxiv.org/html/2609.22135#bib.bib40)\]\. The domain\-shift control in the Supplement additionally uses two text emotion datasets, ISEAR\[[49](https://arxiv.org/html/2609.22135#bib.bib39)\]and DailyDialog\[[29](https://arxiv.org/html/2609.22135#bib.bib37)\]\. The non\-text\-context steering in §6\.2 is conditioned on 100 RAVDESS utterances and 102 EmoSet images, respectively\.

### III\-CActivation Processing and Probing

For each input, we extract the layer\-ℓ\\ellhidden state with a forward hook and mean\-pool over the token dimension to obtain a per\-layer activation vectorh\(ℓ\)∈ℝdh^\{\(\\ell\)\}\\in\\mathbb\{R\}^\{d\}\. Probing is linear\. We first fit a whitening transformWsW\_\{s\}on the training activations of source modalityss\(per\-feature standardization followed by decorrelation, which removes the domination of class centroids by shared principal axes\), then train a linear emotion probefsf\_\{s\}in the whitened spaceh~=Ws​\(h−μs\)\\tilde\{h\}=W\_\{s\}\(h\-\\mu\_\{s\}\)\. Accuracy is reported with 5\-fold cross\-validation\. Cross\-modal transfer means applying the probe trained on source modalityssdirectly to classify activations of target modalitytt, with transfer accuracy defined as

Accs→t=1\|𝒟t\|∑\(h,y\)∈𝒟t\[fs\(Ws\(h−μs\)\)=y\],\\mathrm\{Acc\}\_\{s\\to t\}=\\frac\{1\}\{\|\\mathcal\{D\}\_\{t\}\|\}\\sum\_\{\(h,y\)\\in\\mathcal\{D\}\_\{t\}\}\\mathbb\{1\}\\\!\\left\[\\,f\_\{s\}\\\!\\left\(W\_\{s\}\(h\-\\mu\_\{s\}\)\\right\)=y\\,\\right\],\(1\)where𝒟t\\mathcal\{D\}\_\{t\}is the target\-modality sample set andyythe emotion label\. Whitening is necessary\. In the raw space the largest one or two principal components dominate all class centroids, and naive nearest\-class\-centroid accuracy is only 0\.24–0\.40, and after whitening the same readout exceeds 0\.90\. Unless stated otherwise, statistics are reported as 95% confidence intervals from 1000 bootstrap\[[52](https://arxiv.org/html/2609.22135#bib.bib56)\]resamples over 5 seeds\.

### III\-DSteering Protocol

Emotion directions are extracted independently at each injection layerℓ\\ell\(raw\_cmd\_pc\_orth\_k2\)\. On the layer\-ℓ\\ellactivations of M3 we take the means over target\-emotion and neutral samples to obtain the contrastive activation mean differencev0\(ℓ\)=h¯emo\(ℓ\)−h¯neutral\(ℓ\)v\_\{0\}^\{\(\\ell\)\}=\\bar\{h\}^\{\(\\ell\)\}\_\{\\text\{emo\}\}\-\\bar\{h\}^\{\(\\ell\)\}\_\{\\text\{neutral\}\}, then orthogonalize away the top two dataset\-level principal components\{u1\(ℓ\),u2\(ℓ\)\}\\\{u^\{\(\\ell\)\}\_\{1\},u^\{\(\\ell\)\}\_\{2\}\\\}of that layer to obtain the unit direction

v\(ℓ\)=v0\(ℓ\)−∑k=12\(v0\(ℓ\)⊤​uk\(ℓ\)\)​uk\(ℓ\)∥v0\(ℓ\)−∑k=12\(v0\(ℓ\)⊤​uk\(ℓ\)\)​uk\(ℓ\)∥,v^\{\(\\ell\)\}=\\frac\{v\_\{0\}^\{\(\\ell\)\}\-\\sum\_\{k=1\}^\{2\}\\big\(v\_\{0\}^\{\(\\ell\)\\top\}u^\{\(\\ell\)\}\_\{k\}\\big\)\\,u^\{\(\\ell\)\}\_\{k\}\}\{\\big\\lVert v\_\{0\}^\{\(\\ell\)\}\-\\sum\_\{k=1\}^\{2\}\\big\(v\_\{0\}^\{\(\\ell\)\\top\}u^\{\(\\ell\)\}\_\{k\}\\big\)\\,u^\{\(\\ell\)\}\_\{k\}\\big\\rVert\},\(2\)which reduces confounding with the shared principal axes\. Injection uses a multi\-layer normalized scheme\. Over the injection\-layer setℒ\\mathcal\{L\}\(L10/14/18/22 for Qwen\), the layer\-ℓ\\ellhidden state is updated as

h\(ℓ\)←h\(ℓ\)\+α⁡∥h\(ℓ\)∥​v\(ℓ\),ℓ∈ℒ,h^\{\(\\ell\)\}\\leftarrow h^\{\(\\ell\)\}\+\\alpha\\,\\lVert h^\{\(\\ell\)\}\\rVert\\,v^\{\(\\ell\)\},\\qquad\\ell\\in\\mathcal\{L\},\(3\)so that the norm of the injected vector at each layer is always\|α\|\\lvert\\alpha\\rverttimes the current hidden\-state norm∥h\(ℓ\)∥\\lVert h^\{\(\\ell\)\}\\rVertof that layer \(per\-layer normalized\), making the perturbation magnitude comparable across layers\. Cross\-model experiments mapℒ\\mathcal\{L\}by normalized depth to L12/16/20/24 for Phi\-4 and L14/18/22/27 for MiniCPM\. The main experiments useα=\+0\.12\\alpha=\+0\.12\(some experiments sweep over\{±0\.04,±0\.08,±0\.12\}\\\{\\pm 0\.04,\\pm 0\.08,\\pm 0\.12\\\}\)\. Generation uses prefill\-suffix \+ decode\-all, with temperature 0\.7, top\-pp0\.9, andmax\_new\_tokens80\. Text steering uses 30 neutral prompts by default \(scaled to 150 in §6\.2\)\. Audio\- and image\-context steering use 100 RAVDESS utterances and 102 EmoSet images, respectively \(§6\.2\)\. Every condition is paired with a same\-magnitude random\-direction injection as control, and we monitor the degradation \(OOD\) rate\. The single\-layer control experiments in §5\.1–§5\.2 use the same direction\-extraction and normalization protocol, varying only the number and position of injection layers\.

### III\-EEvaluation

Emotion intensity is read out by an external evaluator, decoupled from the subject model\. The primary evaluator isj\-hartmann/emotion\-english\-distilroberta\-base\[[18](https://arxiv.org/html/2609.22135#bib.bib45)\]\(hereafter VAD\-BERT\), whose native 7\-emotion output we fold into the 5\-emotion space of §3\.2 \(anger / calm / fear / joy / sadness\) as 5\-class probabilities\. Writingpc​\(x\)p\_\{c\}\(x\)for the evaluator’s probability of target emotionccon textxx, the steering effect is defined as the difference between steered and baseline mean target\-emotion probability overNNprompts,

Δtarget=1N​∑i=1N\[pc​\(gsteer​\(xi\)\)−pc​\(gbase​\(xi\)\)\],\\Delta\_\{\\text\{target\}\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\Big\[\\,p\_\{c\}\\big\(g\_\{\\text\{steer\}\}\(x\_\{i\}\)\\big\)\-p\_\{c\}\\big\(g\_\{\\text\{base\}\}\(x\_\{i\}\)\\big\)\\,\\Big\],\(4\)wheregbaseg\_\{\\text\{base\}\}andgsteerg\_\{\\text\{steer\}\}are generation before and after injection\. We additionally use the compound valence of NLTK VADER\[[20](https://arxiv.org/html/2609.22135#bib.bib38)\]as a cross\-check on the valence dimension \(its agreement with the VAD\-BERT readout is reported as a Spearman correlation\)\. To rule out potential coupling between the evaluator and the subject model, §6\.2 scales text steering ton=150n=150and introduces an independent GPT\-5\.5 as an external judge \(LLM\-as\-a\-judge\[[66](https://arxiv.org/html/2609.22135#bib.bib53)\]\), which outputs five\-dimensional emotion probabilities and a fluency score\. The per\-sample agreement between GPT\-5\.5 and VAD\-BERT is likewise reported as a Spearman correlation\. The mechanistic analysis in §5\.4 uses the logit lens\[[40](https://arxiv.org/html/2609.22135#bib.bib57)\], projecting each layer’s hidden state directly through the unembedding matrix to the vocabulary to track the layer\-wise rank of the target\-emotion token\. All steering experiments report the OOD rate\. §6\.5 further verifies, via WikiText perplexity and GSM8K accuracy, that steering does not harm basic language or reasoning ability\.

## IVCross\-Modal Emotion Geometry

Before turning to the dissociation, we characterize the representation space the steering experiments operate on, the geometry of emotion across text, audio, and image inside Qwen2\.5\-Omni\-7B\. A linear emotion probe trained on one modality transfers above the random baseline \(0\.20\) in all six cross\-modal directions, confirming a non\-trivial shared subspace\. Transfer is asymmetric and strongest toward text \(audio\-to\-text 0\.558, image\-to\-text 0\.568\), so text carries the most readily readable component of the shared emotion information and we target text generation for steering in §5; the aggregate text\-target probe peaks at L27\. A same\-modality cross\-domain control \(ISEAR versus DailyDialog\) rules out ordinary domain shift: the cross\-domain drop of 16\.49 percentage points is far smaller than the cross\-modal decay, with a pure\-modality increment of 26\.09 points \(p=0\.0001p=0\.0001\)\.

The shared structure is organized by the two continuous valence–arousal dimensions rather than discrete categories\. Single\-axis valence and arousal transfer \(0\.853 and 0\.696\) far exceed 5\-class transfer \(0\.558\), and orthogonalizing the two axes collapses 5\-class transfer to 0\.176, near chance\.[Figure4](https://arxiv.org/html/2609.22135#S4.F4)visualizes this: the five emotions occupy their expected valence–arousal circumplex positions and the three modalities of each emotion overlap\. This organization is Qwen\-strong but not universal: it only partially replicates on Phi\-4\. The joy direction is the cross\-model exception, occupying the unique high\-valence, high\-arousal corner of the circumplex\[[48](https://arxiv.org/html/2609.22135#bib.bib46)\]and replicating robustly in both models, which anticipates joy as the strongest and most stable steering anchor in §5–§6\. Full transfer tables, the depth anatomy, the domain\-shift control, and the valence–arousal decoupling are reported in the Supplementary Material\.

Fig\. 4:Cross\-modal emotion geometry in Qwen2\.5\-Omni\-7B\. Whitened M3 activations at L14 \(375 samples\) projected onto the supervised valence and arousal axes\. Color encodes emotion, marker shape encodes modality, and shaded regions are 95% confidence ellipses\. The five emotions fall at their expected circumplex positions and the three modalities of each emotion overlap, visualizing the dimensional sharing reported in the Supplementary Material\.
## VProbing–Steering Dissociation

### V\-ADissociation in a Single Controlled Comparison

A core assumption of representation engineering is that the layer at which a linear probe reads the strongest signal is also the layer at which activation steering is most effective\. We test this assumption directly and under control\. On Qwen2\.5\-Omni\-7B, the aggregate text\-target best layer for cross\-modal probe transfer is L27, the peak of the layer\-wise mean of the audio\-to\-text and image\-to\-text transfer curves\. We apply single\-layer activation steering at this layer, holding the injection magnitude \(α=\+0\.12\\alpha=\+0\.12\) and the direction\-extraction protocol \(raw\_cmd\_pc\_orth\_k2\) identical to the multi\-layer configuration\. The causal effect on generation is almost zero\. The VAD\-BERT external readout givesΔjoy=\+0\.02\\Delta\_\{\\text\{joy\}\}=\+0\.02with a 95% bootstrap CI of\[−0\.08,\+0\.13\]\[\-0\.08,\+0\.13\]that contains 0, and the paired random\-direction control gives the sameΔjoy=\+0\.02\\Delta\_\{\\text\{joy\}\}=\+0\.02, so the main direction is essentially indistinguishable from a random one \(n=30n=30prompts, VADER cross\-checkΔ=0\.000\\Delta=0\.000, OOD rate 0\)\. Injecting the same emotion direction in a per\-layer normalized way into the mid\-to\-late multi\-layer combination \(L10/14/18/22\) tells a different story\. It givesΔjoy=\+0\.55\\Delta\_\{\\text\{joy\}\}=\+0\.55with CI\[\+0\.39,\+0\.69\]\[\+0\.39,\+0\.69\], about 26 times the single\-layer effect at L27 \(full numbers in[TableI](https://arxiv.org/html/2609.22135#S5.T1)\)\.

TABLE I:Four\-condition joy steering on Qwen2\.5\-Omni\-7B\. All conditions useα=\+0\.12\\alpha=\+0\.12, the same prompts \(n=30n=30\) and direction protocol, differing only in layer selection\. Ours and Baseline\-B/C use theraw\_cmd\_pc\_orth\_k2direction, Baseline\-A the centroid\.Δjoy\\Delta\_\{\\text\{joy\}\}is the VAD\-BERT change of steered relative to baseline \(95% CI from 1000 bootstrap\), RandomΔ\\Deltais the same\-condition random control, and the OOD rate is 0% throughout\. Visual comparison in[Figure5](https://arxiv.org/html/2609.22135#S5.F5)\.Fig\. 5:Four\-condition joy steering effect\. Solid bars areΔjoy\\Delta\_\{\\text\{joy\}\}of the main direction, lighter bars the same\-condition random control, and error bars are 95% bootstrap CIs\. Multi\-layer injection \(Ours, Baseline\-A\) rises sharply, while single\-layer probe\-best L27 \(Baseline\-C\) sits near zero and matches its random control\. The roughly 26\-fold gap is the empirical signature of the dissociation\. Exact values in[TableI](https://arxiv.org/html/2609.22135#S5.T1)\.[Figure5](https://arxiv.org/html/2609.22135#S5.F5)makes the gap visually clear\. The main\-direction bar for Baseline\-C \(single\-layer probe\-best L27\) is almost as low as its random\-direction control bar, whereas the multi\-layer bars \(Ours / Baseline\-A\) rise sharply, and their 95% bootstrap error bars do not overlap at all\.

Layer position, not layer count, drives this gap\. Baseline\-B in[TableI](https://arxiv.org/html/2609.22135#S5.T1), a*single*\-layer L18 injection in the mid\-to\-late range, already reachesΔjoy=\+0\.15\\Delta\_\{\\text\{joy\}\}=\+0\.15, about 7 times the single\-layer probe\-best L27 \(\+0\.02\+0\.02\)\. With layer count held at one, moving the injection from L27 to the mid\-zone L18 alone produces this 7\-fold change, and the additional multi\-layer gain \(§5\.2\) only stacks on top\. The 26\-fold gap is dominated by where one injects\.

This*readability–intervenability*gap is the empirical signature of the probing–steering dissociation: the layer where a linear probe reads most easily is not the layer where activation steering is most effective\.

### V\-BLayer\-Wise Causal Map

§5\.1 established the dissociation through a single L27\-vs\-Ours comparison\. We now give the full 28\-layer causal map on Qwen2\.5\-Omni\-7B and pre\-register the dissociation criteria for the cross\-model extension \(§5\.3\)\.

Single\-layer steering sweep\.We measure the single\-layer activation steering effect independently for all 28 layers \(3 emotions×\\times2α\\alpha×\\times5 seeds×\\times20 prompts, 16,800 generations in total\)\. As shown in[Figure6](https://arxiv.org/html/2609.22135#S5.F6)and[TableII](https://arxiv.org/html/2609.22135#S5.T2), the steering effect is near zero at both ends \(L0–L5 and L25–L27\) and peaks in the mid\-to\-late range\. The single\-layer best for joy is stable at L18 \(α=\+0\.12\\alpha=\+0\.12,Δjoy=\+0\.194\\Delta\_\{\\text\{joy\}\}=\+0\.194\), the best for anger is L17 \(Δanger=\+0\.079\\Delta\_\{\\text\{anger\}\}=\+0\.079\), and the best for sadness is L20 \(Δsadness=\+0\.040\\Delta\_\{\\text\{sadness\}\}=\+0\.040\)\. Sadness is the least stable of the three, with a secondary peak at L5 under lowα\\alpha\.

TABLE II:Single\-layer steering best layer and effect over the 28 layers of Qwen2\.5\-Omni\-7B \(α=\+0\.12\\alpha=\+0\.12,n=20n=20prompts×\\times5 seeds, VAD\-BERT\)\. All three emotions peak in the mid\-to\-late range \(L17–L20\), while the end layers are near 0\. Multi\-layer effect in[TableI](https://arxiv.org/html/2609.22135#S5.T1)\.Fig\. 6:The 28\-layer dissociation curves of Qwen2\.5\-Omni\-7B\. Blue is layer\-wise probing accuracy \(left axis\), and the colored lines are the single\-layer steering effect for joy / anger / sadness \(right axis,α=\+0\.12\\alpha=\+0\.12\)\. The probing peak \(L27\) and steering peak \(L20\) are 7 layers apart, and the shading marks the mid\-to\-late steering zone \[L14, L22\]\.Pre\-registered dissociation criteria\.We define the dissociation as holding by either of two independent criteria: \(i\) the Pearsonrrbetween layer\-wise probing accuracy and steering effect over the 28 layers is<0\.30<0\.30, or \(ii\) the absolute layer distance between the two curves’ peaks is≥5\\geq 5layers\.

On the Qwen aggregate \(text\-target probing\), the steering peak is L18 atα=\+0\.08\\alpha=\+0\.08and L20 atα=\+0\.12\\alpha=\+0\.12, against a probing peak of L27 in both cases \(a gap of 9 and 7 layers, withr=0\.36r=0\.36and0\.450\.45, respectively\)\. The peak\-gap criterion passes at both values ofα\\alpha\. The Pearsonrr, though close, does not cross the 0\.30 threshold, so we adopt the peak gap as the primary criterion and carry it over to the cross\-model extension in §5\.3\.

Layer\-pair synergy\.After 12 selected layer\-pair injection experiments, super\-additivity appears in 8/12 pairs for joy and 7/12 for anger\. These layer pairs cover combinations of shallow\+shallow, shallow\+final, middle\+middle, middle\+final, and final\+final, and the strongest pairs consistently concentrate within the mid\-to\-late corridor, e\.g\., L14\_18, L18\_22, L14\_22\. For sadness, 6/12 pairs are weaker than its strongest single layer\. This shows that the mid\-to\-late steering zone holds at the single\-layer level and is also the dense locus of pair\-level synergy\.

[Figure6](https://arxiv.org/html/2609.22135#S5.F6)combines this evidence in one plot: the probing\-accuracy curve \(blue\) climbs monotonically with depth and peaks at L27, while the three steering\-effect curves \(red/orange/blue\) are single\-peaked within the mid\-to\-late corridor and approach zero at both ends, with the two sets of peaks clearly offset\. The shape of the curves alone displays the layer separation between reading and intervention\.

Combining the single\-layer curves, the peak gap, and pair synergy as three independent lines of evidence, the probing–steering dissociation on Qwen2\.5\-Omni\-7B is confirmed at the level of a*quantitative layer\-wise causal map*\. §5\.3 extends the same pre\-registered pair of criteria to Phi\-4\-Multimodal and MiniCPM\-o\-4\.5\.

### V\-CUniversality of the Steering\-Effective Zone Across Three Models

Does the dissociation quantified on Qwen2\.5\-Omni\-7B \(§5\.1, §5\.2\) hold across architectures? We replicate the measurement under an identical protocol on Phi\-4\-Multimodal\-Instruct \(32 layers, hidden 3072\) and MiniCPM\-o\-4\.5 \(36 layers, hidden 4096\), which share no lineage with Qwen in LLM backbone, tokenizer, or modality encoder\.

The cross\-model comparison reveals a striking asymmetry\. The probe\-best layer sits at normalized depths of 1\.00 \(Qwen L27\), 0\.81 \(Phi\-4 L25\), and 0\.29 \(MiniCPM L10\), spanning almost the full network depth\. The steer\-best layer, by contrast, stays in the mid\-to\-late residual stream, at normalized depths of 0\.74 \(Qwen L20\), 0\.48 \(Phi\-4 L15\), and 0\.54 \(MiniCPM L19\) \([TableIII](https://arxiv.org/html/2609.22135#S5.T3),[Figure7](https://arxiv.org/html/2609.22135#S5.F7)\)\. The absolute layer distances between probe peak and steer peak are 7, 10, and 9 layers, all exceeding the pre\-registered dissociation threshold from §5\.2 \(\|peak gap\|≥5\\lvert\\text\{peak gap\}\\rvert\\geq 5layers\), and the layer\-wise Pearson correlations are 0\.45, 0\.55, and 0\.36, short of the*highly correlated*range\.

The steer\-best layer of all three architectures lands in the same narrow band of normalized depth\[0\.48,0\.74\]\[0\.48,0\.74\], while the probe\-best layer scatters across\[0\.29,1\.00\]\[0\.29,1\.00\], nearly the entire network\. This contrast holds across the three models with no counterexample: the steering\-effective zone is stable across architectures, the probe\-best layer is not\. MiniCPM shows an order reversal, its probe\-best layer sitting in the early network \(L10\) ahead of its steer\-best \(L19\), opposite to Qwen and Phi\-4\. The reversal sits entirely on the volatile probing side\. MiniCPM’s steer\-best \(0\.54\) still falls inside the universal band\. Why MiniCPM’s probing–steering ordering differs is a question of mechanism, which we take up in §5\.5 and §7\.2\.

TABLE III:Cross\-architecture probing–steering dissociation across three models \(Qwen2\.5\-Omni, Phi\-4\-Multimodal, MiniCPM\-o\-4\.5\)\. Probe and steer peaks are measured per model by a layer\-wise scan \(5 seeds, VAD\-BERT\); each cell gives the layer and its normalized depth \(layer /LmaxL\_\{\\max\}\)\. Under the pre\-registered criterion \(\|peak gap\|≥5\\lvert\\text\{peak gap\}\\rvert\\geq 5layers\), the gap passes in all three \(7 / 10 / 9 layers\), and the steering peak \(0\.48–0\.74\) is far more concentrated than the probing peak \(0\.29–1\.00\)\.Fig\. 7:Cross\-architecture probing–steering dissociation \(signature figure\)\. Subplots are Qwen2\.5\-Omni, Phi\-4\-Multimodal, and MiniCPM\-o\-4\.5\. The probe\-best layer \(triangle\) drifts across normalized depth 0\.29–1\.00, nearly the full network, while the steer\-best layer \(star\) stays in the steering\-effective zone \[0\.48, 0\.74\] \(band\) across all three architectures\.The three subplots of[Figure7](https://arxiv.org/html/2609.22135#S5.F7)make this invariance immediate: the steer\-best stars of all three models fall into the same pale\-yellow band \(normalized depth \[0\.48, 0\.74\]\), while the probe\-best triangles scatter across 0\.29, 0\.81, and 1\.00\. This is the strongest cross\-architecture evidence that reading and intervention are functionally distinct operations on the residual stream\.

### V\-DA Logit\-Lens Mechanism

§5\.1–§5\.3 quantified the empirical strength of the probing–steering dissociation, but*why*it holds remains open\. We use the logit lens\[[40](https://arxiv.org/html/2609.22135#bib.bib57)\]to measure, after projecting each layer’s hidden state directly through the unembedding matrix, the timing at which the target\-emotion token enters the top\-50 vocabulary candidates\.

On Qwen2\.5\-Omni\-7B \(n=30n=30prompts×\\times3 emotions=90=90trajectories, emotion\-conditioned, with the logit lens run as a deterministic forward pass with no sampling\-seed dimension\), the median layer at which the target emotion token first enters the top\-50 vocabulary candidates is L23 \(the vocabulary\-commitment layer\)\. Combined with the probing\-saturation layer established in §5\.2 \(L20, where the linear probe reaches 95% of its peak\) and the steering\-peak layer \(L19, the aggregate peak of the §5\.2 layer\-wise sweep\), we obtain a clear three\-stage ordering, with the steering peak at L19 first, then probing saturation at L20, and finally vocabulary commitment at L23\.

The three pre\-registered hypotheses are adjudicated in[TableIV](https://arxiv.org/html/2609.22135#S5.T4)\. H2 \(the steering\-effective layer is adjacent to the commitment window, pre\-registered threshold\|gap\|≤4\\lvert\\text\{gap\}\\rvert\\leq 4\) holds, with\|L19−L23\|=4\\lvert\\text\{L19\}\-\\text\{L23\}\\rvert=4layers, exactly at the threshold\. H3 \(probing saturates before vocabulary commitment\) also holds, with L20 before L23\. H1 \(commitment precedes saturation\) does not hold: saturation is 3 layers earlier than commitment\.

TABLE IV:Adjudication of the three pre\-registered logit\-lens hypotheses \(Qwen2\.5\-Omni\-7B, 90 emotion\-conditioned trajectories; layer numbers are medians of the timing\)\.Fig\. 8:Cross\-layer logit\-lens trajectory of the emotion token \(joy / anger / sadness, median rank overn=30n=30prompts\)\. The vertical axis is the median vocabulary rank of the target emotion token \(log scale, inverted, so the top is rank 1\)\. Shaded bands mark the three landmark layers: steering peak \(L19\), probing saturation \(L20\), and vocabulary commitment \(L23\)\.[Figure8](https://arxiv.org/html/2609.22135#S5.F8)shows the layer\-wise rank trajectory of the three emotion tokens: over L0–L22 the median rank of the target emotion token stays at the10410^\{4\}scale \(buried deep inside the vocabulary\) and only at L23 jumps sharply to the top of the vocabulary \(rank≤2\\leq 2\)\. This abrupt change confirms the timing of vocabulary commitment and shows that at the layers before commitment \(including steering peak L19 and probing saturation L20\) the emotion information is already encoded but not yet linearized into a token choice\.

The fact that H1*does not hold*is itself informative\. It indicates that the representation stabilizes first and is then linearized into token logits\. The causal window for activation steering thus lies in the stable\-but\-uncommitted intermediate state, near L19\. The probe\-best layer L27, already well past the L23 commitment, is where the model has committed to its token choice, so an injected perturbation can no longer rewrite the already\-linearized logit structure\. This timing mechanistically explains the phenomenon in §5\.1 whereby single\-layer injection at L27 collapses to a level indistinguishable from a random direction\.

### V\-EA Predictive Functional Model

The three\-stage ordering in §5\.4 suggests a specific functional form\. The effectiveness of steering at layerLLshould depend on two factors at once: \(i\) the information at that layer is already sufficiently encoded, i\.e\., the probe accuracy exceeds a thresholdθ\\theta, and \(ii\) there is still enough downstream amplification room after that layer, i\.e\., the number of remaining layers exceeds a thresholdτ\\tau\. We write this hypothesis as a two\-factor model:

steer⁡\(L\)\\displaystyle\\mathrm\{steer\}\(L\)∝max⁡\(0,probe⁡\(L\)−θ\)\\displaystyle\\propto\\max\\\!\\left\(0,\\,\\mathrm\{probe\}\(L\)\-\\theta\\right\)×max⁡\(0,Lmax−L−τ\)\.\\displaystyle\\times\\max\\\!\\left\(0,\\,L\_\{\\max\}\-L\-\\tau\\right\)\.
We grid\-search\(θ,τ\)\(\\theta,\\tau\)independently for each model to predict the steering peak layer \([Figure9](https://arxiv.org/html/2609.22135#S5.F9)\)\. On Qwen2\.5\-Omni,\(θ=0\.54,τ=0\)\(\\theta=0\.54,\\tau=0\)predicts a steering peak of L20, exactly matching the measured L20\. On Phi\-4\-Multimodal,\(θ=0\.24,τ=14\)\(\\theta=0\.24,\\tau=14\)predicts L16 against a measured L15, an error of 1 layer or about 3% of network depth\. On MiniCPM\-o\-4\.5,\(θ=0\.24,τ=0\)\(\\theta=0\.24,\\tau=0\)predicts L10 but the measured peak is L19, an error of 9 layers and a clear failure \([Figure9](https://arxiv.org/html/2609.22135#S5.F9)\)\.

Fig\. 9:HthreshH\_\{\\mathrm\{thresh\}\}prediction versus the measured steering effect across models\. Solid lines are measured, dashed lines theHthreshH\_\{\\mathrm\{thresh\}\}prediction, stars the measured peak, and triangles the predicted peak\. Qwen is hit exactly \(error 0\), Phi\-4 is off by 1 layer, and MiniCPM by 9 layers \(red shading marks the gap\)\. Fit parameters in the Supplementary Material\.The exact/approximate success on Qwen and Phi\-4 quantitatively supports the causal reading of the*stable\-but\-uncommitted*timing in §5\.4, where steering effectiveness≈\\approxrepresentational readability×\\timesdownstream plasticity\. But MiniCPM’s failure is equally structural\. MiniCPM’s probe\-best layer is in the early network \(L10\), andHthreshH\_\{\\mathrm\{thresh\}\}predicts that the steer peak should likewise be early, yet the measured peak is at L19, 9 layers later than predicted\. This deviation cannot be explained by the two factors of*information encoding×\\timesdownstream room,*suggesting a third architecture\-related factor, a model\-specific lower bound of the actionable zone, which dictates that even when early\-layer information is already encoded, steering takes effect only once the representation enters this intervenable band\.

This residual deviation is the most concrete open problem for a mechanistic theory of steering\. Does a cross\-architecture universal actionable zone exist? If so, how can it be measured directly, independent of the two diagnostics of probing and the logit lens? We take this up as a central topic of the discussion and future work \(§7\.3\)\.

## VICross\-Modal and Cross\-Model Causal Control

Can multi\-layer injection in the mid\-to\-late zone located by the dissociation causally control the emotion of generation across three input modalities and three models? We test this, validate it with an independent LLM judge, characterize the cross\-model specificity of joy, and confirm that steering preserves the model’s basic abilities\.

### VI\-AMulti\-Layer Normalized Injection Protocol

Building on the dissociation result of §5, we use multi\-layer normalized injection across the mid\-to\-late effective zone rather than a single probe\-best layer: theraw\_cmd\_pc\_orth\_k2direction is injected at L10/14/18/22 of Qwen, per\-layer normalized, withα=\+0\.12\\alpha=\+0\.12\(full protocol in §3\.4\)\. Emotion intensity is read out by the external VAD\-BERT evaluator with a same\-condition random\-direction paired control, and the OOD rate is monitored throughout\.

### VI\-BTri\-Modal Controllability

![Refer to caption](https://arxiv.org/html/2609.22135v1/fig10.png)Fig\. 10:Qualitative cross\-modal joy steering\. The same multi\-layer joy injection \(L10/14/18/22,α=\+0\.12\\alpha=\+0\.12\) is applied to an image input \(EmoSet, anger\) and an audio input \(RAVDESS, sad speech\)\. The baseline generation conveys the original negative tone, while the joy\-steered generation shifts to a celebratory one, with the VAD\-BERT joy probability rising from0\.000\.00to0\.990\.99in both cases\. Quantitative results in[TableV](https://arxiv.org/html/2609.22135#S6.T5)\.We test the same protocol under three input contexts: text\-only \(30 prompts\), audio\-conditioned \(100 RAVDESS utterances\), and image\-conditioned \(102 EmoSet images\)\.[TableV](https://arxiv.org/html/2609.22135#S6.T5)gives the VAD\-BERTΔtarget\\Delta\_\{\\text\{target\}\}atα=\+0\.12\\alpha=\+0\.12for the joy / sadness / anger directions under the three modalities\. Across the nine cells of three modalities and three emotions, joy is strong under all three modalities \(\+0\.33\+0\.33to\+0\.67\+0\.67\), while sadness / anger show moderate, context\-dependent positive effects\. The OOD rate is 0 in all conditions, and paired random\-direction controls are near zero\. This shows that the multi\-layer effective zone located in §5 remains causally effective on text generation and in non\-text omni\-modal contexts\.[Figure10](https://arxiv.org/html/2609.22135#S6.F10)illustrates this qualitatively for an image and an audio input: the same injection turns a negative baseline description into a joyful one in both cases, with the external VAD\-BERT joy probability rising from0\.000\.00to0\.990\.99\.

TABLE V:Multi\-layer emotion steering under three input contexts \(VAD\-BERTΔtarget\\Delta\_\{\\text\{target\}\},α=\+0\.12\\alpha=\+0\.12; text 30 prompts, audio 100 RAVDESS, image 102 EmoSet\)\. joy is strongly positive under all three modalities, the OOD rate is 0, and random controls are near zero\.Independent LLM\-judge validation\.VAD\-BERT is an automatic evaluator homologous to the subject model\. To rule out evaluator coupling, we scale text steering ton=150n=150prompts and rescore with an independent GPT\-5\.5 as the external judge\. The joy direction givesΔjoy=\+0\.246\\Delta\_\{\\text\{joy\}\}=\+0\.246under GPT\-5\.5 \(95% CI\[\+0\.211,\+0\.283\]\[\+0\.211,\+0\.283\], paired random control−0\.007\-0\.007\), and the per\-sample Spearman correlation between GPT\-5\.5 and VAD\-BERT reachesρ=0\.809\\rho=0\.809\. This magnitude is comparable to the automatic\-versus\-human agreement reported by the human\-evaluation study of style\-vector steering\[[12](https://arxiv.org/html/2609.22135#bib.bib6)\]\(meanr=0\.776r=0\.776across over 7,000 crowdsourced ratings, ICC=0\.71=0\.71–0\.870\.87\), supporting the automatic evaluator as a credible proxy for emotion intensity and confirming that the joy effect is not an artifact of a single evaluator\. Sadness \(\+0\.023\+0\.023\) and anger \(\+0\.024\+0\.024\) show only weak effects under GPT\-5\.5, and anger’s VAD\-BERT–GPT correlation is low \(ρ=0\.19\\rho=0\.19\), consistent with the dimensional structure of §4\. Joy occupies the unique position of the circumplex and has the cleanest effect, while mixed\-valence emotions have weak effects and low evaluator agreement\.

### VI\-CCross\-Model Steering Replication

We replicate joy steering on Phi\-4\-Multimodal and MiniCPM\-o\-4\.5 under the same G15 protocol \(layer positions mapped to each model by normalized depth\)\.[TableVI](https://arxiv.org/html/2609.22135#S6.T6)shows that Phi\-4 replicates a strong effect \(Δjoy=\+0\.415\\Delta\_\{\\text\{joy\}\}=\+0\.415, CI excluding 0, all pre\-registered criteria passed\), about 25% weaker than Qwen, consistent with the model\-size difference \(3\.8B vs 7B\)\. MiniCPM\-o gives a null result \(Δjoy=\+0\.022\\Delta\_\{\\text\{joy\}\}=\+0\.022, CI containing 0\)\. MiniCPM’s failure is not random\. It corroborates the failure ofHthreshH\_\{\\mathrm\{thresh\}\}to predict MiniCPM in §5\.5 \(peak error 9 layers\) and the unstable geometric diagnostics of MiniCPM in §4, jointly marking MiniCPM as a boundary\-condition model rather than a simple replication failure\.

TABLE VI:Cross\-model replication of joy steering \(VAD\-BERTΔjoy\\Delta\_\{\\text\{joy\}\},α=\+0\.12\\alpha=\+0\.12, layers mapped by normalized depth\)\. Qwen and Phi\-4 replicate a strong effect \(CI excludes 0\); MiniCPM\-o is null \(CI contains 0\), matching its boundary\-condition diagnostics in §4 / §5\.5\.
### VI\-DJoy as a Cross\-Model Anchor

Why is joy the most stable across modalities and models? We test the directional specificity of each emotion with a5×55\\times 5cross\-injection matrix \(GPT\-5\.5 judge\), scoring each direction by its diagonal increment minus the largest off\-diagonal increment\. Joy is the only emotion specific \(above the 0\.15 threshold\) in both Qwen and Phi\-4 \(0\.363 and 0\.469\); calm passes only on Phi\-4, and the mixed\-valence emotions pass in neither\. Joy occupies the unique high\-valence, high\-arousal corner of the circumplex, so its direction overlaps least with neighbors, giving the strongest and most consistent specificity\. This joy\-anchored separability is a cross\-model invariant, consistent with the dimensional organization of §4\. The full specificity matrix is in the Supplementary Material\.

### VI\-EPreservation of Fluency and Reasoning Ability

Steering changes emotion without harming basic abilities\. Joy steering \(α=\+0\.12\\alpha=\+0\.12\) raises WikiText\[[37](https://arxiv.org/html/2609.22135#bib.bib55)\]perplexity only moderately \(5\.62 to 8\.39, far below the PPL\>50\>50of destructive steering\), leaves the GPT\-5\.5 fluency score essentially unchanged \(Δfluency=−0\.03\\Delta\_\{\\text\{fluency\}\}=\-0\.03\), and does not lower GSM8K\[[8](https://arxiv.org/html/2609.22135#bib.bib54)\]accuracy \(16% to 18%, within noise\)\. Further detail is in the Supplementary Material\.

## VIIDiscussion

### VII\-AImplications for Representation Engineering

The results of §5 refute the standard assumption that the layer where a probe reads a concept best is also where intervention works best\. On Qwen2\.5\-Omni, single\-layer injection at the probe\-best layer is indistinguishable from a random direction while mid\-to\-late multi\-layer injection is about 26 times stronger \([TableI](https://arxiv.org/html/2609.22135#S5.T1)\), and the gap holds across all three architectures, with the probe\-best layer spanning nearly the full network depth and the steer\-best layer confined to a narrow mid\-to\-late band \([TableIII](https://arxiv.org/html/2609.22135#S5.T3)\)\.

This yields a directly practical corollary\. Selecting the injection layer by probing accuracy is a mistaken heuristic\. Under the common practice of choosing the layer at the probing peak, it can place intervention at causally inert layers and thus severely underestimate a method’s true controllability\. By contrast, the mid\-to\-late steering\-effective zone \(normalized depth about0\.50\.5–0\.750\.75\) is stable across architectures and provides a layer\-selection criterion that does not depend on any single model’s probing curve\. The robust gain of multi\-layer normalized injection over single\-layer injection \(§5\.2\) further shows that spreading the injection within this zone is more reliable than betting on a single layer\. The boundary of this criterion is set by MiniCPM\. Although its steering\-effective zone still lies in the mid\-to\-late range,HthreshH\_\{\\mathrm\{thresh\}\}fails to predict its peak \([Figure9](https://arxiv.org/html/2609.22135#S5.F9)\), indicating that the criterion gives an effective*range*rather than a single point that can be extrapolated exactly, and that precise cross\-architecture localization still requires model\-specific calibration\.

### VII\-BThe Geometric Basis of the Dissociation

Why does the dissociation hold at the functional level? The logit\-lens ordering of §5\.4 \(steering peak, then probing saturation, then vocabulary commitment\) places*reading*and*intervention*at different stages of the forward pass\. A probe reads the strongest signal where the representation*saturates*and is fully linearly separable, whereas steering is most effective in the earlier window where the representation is*stable but not yet committed to a token*\. Readability measures whether the information*is already present*\. Intervenability measures whether it*can still be rewritten*\. The two are functionally distinct operations on the residual stream, with no a priori reason to share the same layer\.

TheHthreshH\_\{\\mathrm\{thresh\}\}two\-factor model of §5\.5 makes this falsifiable: steering effectiveness scales with the product of information already encoded and downstream amplification room, hitting the peak exactly on Qwen and within one layer on Phi\-4 \([Figure9](https://arxiv.org/html/2609.22135#S5.F9)\)\. MiniCPM exposes what the two factors miss: its steering peaks late despite early\-encoded information with ample downstream room, pointing to a model\-specific lower bound on the actionable zone that the representation must enter before intervention takes hold\. The gap is in the predictive formula, not the empirical finding, as MiniCPM’s steer\-best layer still lies in the cross\-architecture mid\-to\-late band\. The dissociation is thus general in substance and architecture\-dependent in form: the steering\-effective zone is stable across architectures, while the exact inter\-layer geometry, especially where the probing peak lands, is set by the architecture and remains to be formally characterized \(§7\.3\)\.

### VII\-CFuture Work

This work leaves several directly extensible directions\. The first is the*actionable zone*open problem raised in §5\.5\. Whether a cross\-architecture universal intervenable band exists, and whether it can be measured directly and independently of the probing and logit\-lens diagnostics, is the key to explaining the MiniCPM deviation and completing the third factor ofHthreshH\_\{\\mathrm\{thresh\}\}\. Second, the dissociation framework can in principle transfer to other intervention tasks that require layer selection\. Our preliminary exploration shows that on a safety steering task the mid\-to\-late layers located by the dissociation outperform the probe\-best layer, suggesting that the framework has applicability beyond emotion generation\. The directions used in that exploration come from safety contrastive activations, not the emotion subspace, and their cross\-model stability is unresolved, so we treat it as directional evidence of the framework’s transferability and leave it to future systematic study\.

## VIIILimitations

This work has several boundaries to state explicitly\. First, all three models are at the 7B scale, and whether the general mid\-to\-late steering\-effective zone is preserved as scale extends to 70B\+ has not been verified\. Second, the emotion labels and datasets are mainly in English, and whether the emotion structure replicates across languages and cultures remains open\. Third, we use linear probes and linear directions throughout, and non\-linear readouts \(such as a sparse autoencoder\) might reveal structure that linear methods miss, which this work does not cover\. Fourth, causal validation concentrates on injection toward text generation, and within\-modality steering \(e\.g\., injecting an audio direction to produce audio output\) is left to future work\. Fifth,HthreshH\_\{\\mathrm\{thresh\}\}is exact on Qwen, approximate on Phi\-4, and fails on MiniCPM \([Figure9](https://arxiv.org/html/2609.22135#S5.F9)\), indicating that the two\-factor model does not yet capture some architecture\-related actionable\-zone factor, and we have listed its formal characterization as a central open problem \(§7\.3\)\. Each of these boundaries points to a specific extension, not a confounder for the core conclusion\. The dissociation itself holds robustly across three independent architectures and does not depend on relaxing any of the above conditions\.

## IXConclusion

Reading and intervention are not the same operation on the residual stream of omni\-modal language models\. The layer where a linear probe reads a concept most strongly and the layer where activation steering is most effective are functionally separate: on Qwen2\.5\-Omni the causal gap is roughly 26\-fold \([TableI](https://arxiv.org/html/2609.22135#S5.T1)\), and across the three independent architectures the probe\-best layer scatters across nearly the full network depth while the steer\-best layer converges in a narrow mid\-to\-late band \([TableIII](https://arxiv.org/html/2609.22135#S5.T3)\)\. A logit\-lens account \(causal handle, then probing saturation, then vocabulary commitment\) and the falsifiableHthreshH\_\{\\mathrm\{thresh\}\}model locate the steering window in the stable\-but\-uncommitted stage, instantiated in a cross\-modal emotion geometry organized by valence–arousal and anchored across models by joy \(§4\)\. The long\-assumed*readable implies steerable*premise of representation engineering thus does not hold in omni\-LLMs: the mid\-to\-late zone is the locus of cross\-architecture\-stable causal intervention, while the probe\-best layer only reflects representational saturation\.

## Acknowledgments

This work was supported in part by the National Natural Science Foundation of China \(Grant No\. 62227807 and Grant No\. U24B20186\), in part by the Brain Science and Brain\-like Intelligence Technology—National Science and Technology Major Project \(No\. 2021ZD0200600, No\. 2021ZD0200408\), and in part by the WQ & UCAS Research Academy Intelligent Computing Center \(WRA\-ICC\) and the Supercomputing Center of Lanzhou University\.

## References

- \[1\]A\. Abouelenin, A\. Ashfaq, A\. Atkinson, H\. Awadalla, N\. Bach, J\. Bao, A\. Benhaim, M\. Cai, V\. Chaudhary, C\. Chen,et al\.\(2025\)Phi\-4\-mini technical report: compact yet powerful multimodal language models via mixture\-of\-loras\.arXiv preprint arXiv:2503\.01743\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[2\]G\. Alain and Y\. Bengio\(2016\)Understanding intermediate layers using linear classifier probes\.arXiv preprint arXiv:1610\.01644\.Cited by:[§I](https://arxiv.org/html/2609.22135#S1.p1.1),[§II\-B](https://arxiv.org/html/2609.22135#S2.SS2.p1.1)\.
- \[3\]J\. Alayrac, J\. Donahue, P\. Luc, A\. Miech, I\. Barr, Y\. Hasson, K\. Lenc, A\. Mensch, K\. Millican, M\. Reynolds,et al\.\(2022\)Flamingo: a visual language model for few\-shot learning\.Advances in neural information processing systems35,pp\. 23716–23736\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[4\]A\. Arditi, O\. Obeso, A\. Syed, D\. Paleka, N\. Panickssery, W\. Gurnee, and N\. Nanda\(2024\)Refusal in language models is mediated by a single direction\.Advances in Neural Information Processing Systems37,pp\. 136037–136083\.Cited by:[§II\-A](https://arxiv.org/html/2609.22135#S2.SS1.p1.1)\.
- \[5\]Y\. Belinkov\(2022\)Probing classifiers: promises, shortcomings, and advances\.Computational Linguistics48\(1\),pp\. 207–219\.Cited by:[§II\-B](https://arxiv.org/html/2609.22135#S2.SS2.p1.1)\.
- \[6\]N\. Belrose, I\. Ostrovsky, L\. McKinney, Z\. Furman, L\. Smith, D\. Halawi, S\. Biderman, and J\. Steinhardt\(2023\)Eliciting latent predictions from transformers with the tuned lens\.arXiv preprint arXiv:2303\.08112\.Cited by:[§II\-B](https://arxiv.org/html/2609.22135#S2.SS2.p1.1)\.
- \[7\]Y\. Chu, J\. Xu, Q\. Yang, H\. Wei, X\. Wei, Z\. Guo, Y\. Leng, Y\. Lv, J\. He, J\. Lin,et al\.\(2024\)Qwen2\-audio technical report\.arXiv preprint arXiv:2407\.10759\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[8\]K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano,et al\.\(2021\)Training verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168\.Cited by:[§VI\-E](https://arxiv.org/html/2609.22135#S6.SS5.p1.1)\.
- \[9\]A\. Conmy, A\. Mavor\-Parker, A\. Lynch, S\. Heimersheim, and A\. Garriga\-Alonso\(2023\)Towards automated circuit discovery for mechanistic interpretability\.Advances in Neural Information Processing Systems36,pp\. 16318–16352\.Cited by:[§II\-B](https://arxiv.org/html/2609.22135#S2.SS2.p1.1)\.
- \[10\]S\. Dathathri, A\. Madotto, J\. Lan, J\. Hung, E\. Frank, P\. Molino, J\. Yosinski, and R\. Liu\(2019\)Plug and play language models: a simple approach to controlled text generation\.arXiv preprint arXiv:1912\.02164\.Cited by:[§II\-A](https://arxiv.org/html/2609.22135#S2.SS1.p1.1)\.
- \[11\]D\. Demszky, D\. Movshovitz\-Attias, J\. Ko, A\. Cowen, G\. Nemade, and S\. Ravi\(2020\)GoEmotions: a dataset of fine\-grained emotions\.InProceedings of the 58th annual meeting of the association for computational linguistics,pp\. 4040–4054\.Cited by:[§III\-B](https://arxiv.org/html/2609.22135#S3.SS2.p2.1)\.
- \[12\]D\. Diallo, K\. Dworatzyk, S\. Jentzsch, P\. Schütt, S\. Theis, and T\. Hecking\(2025\)The effectiveness of style vectors for steering large language models: a human evaluation\.IEEE Access13,pp\. 191443–191457\.Cited by:[§VI\-B](https://arxiv.org/html/2609.22135#S6.SS2.p2.1)\.
- \[13\]Y\. Dong, L\. Jin, Y\. Yang, B\. Lu, J\. Yang, and Z\. Liu\(2025\)From rational answers to emotional resonance: the role of controllable emotion generation in language models\.arXiv preprint arXiv:2502\.04075\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[14\]N\. Elhage, T\. Hume, C\. Olsson, N\. Schiefer, T\. Henighan, S\. Kravec, Z\. Hatfield\-Dodds, R\. Lasenby, D\. Drain, C\. Chen,et al\.\(2022\)Toy models of superposition\.arXiv preprint arXiv:2209\.10652\.Cited by:[§II\-B](https://arxiv.org/html/2609.22135#S2.SS2.p1.1)\.
- \[15\]M\. Geva, A\. Caciularu, K\. Wang, and Y\. Goldberg\(2022\)Transformer feed\-forward layers build predictions by promoting concepts in the vocabulary space\.InProceedings of the 2022 conference on empirical methods in natural language processing,pp\. 30–45\.Cited by:[§II\-B](https://arxiv.org/html/2609.22135#S2.SS2.p1.1)\.
- \[16\]M\. Geva, R\. Schuster, J\. Berant, and O\. Levy\(2021\)Transformer feed\-forward layers are key\-value memories\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,pp\. 5484–5495\.Cited by:[§II\-B](https://arxiv.org/html/2609.22135#S2.SS2.p1.1)\.
- \[17\]R\. Girdhar, A\. El\-Nouby, Z\. Liu, M\. Singh, K\. V\. Alwala, A\. Joulin, and I\. Misra\(2023\)Imagebind: one embedding space to bind them all\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 15180–15190\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[18\]J\. Hartmann\(2022\)Emotion english distilroberta\-base\.Cited by:[§III\-E](https://arxiv.org/html/2609.22135#S3.SS5.p1.1)\.
- \[19\]J\. Hewitt and C\. D\. Manning\(2019\)A structural probe for finding syntax in word representations\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\),pp\. 4129–4138\.Cited by:[§II\-B](https://arxiv.org/html/2609.22135#S2.SS2.p1.1)\.
- \[20\]C\. Hutto and E\. Gilbert\(2014\)Vader: a parsimonious rule\-based model for sentiment analysis of social media text\.InProceedings of the international AAAI conference on web and social media,Vol\.8,pp\. 216–225\.Cited by:[§III\-E](https://arxiv.org/html/2609.22135#S3.SS5.p1.2)\.
- \[21\]J\. Ji, C\. Xu, X\. Zhang, B\. Wang, and X\. Song\(2020\)Spatio\-temporal memory attention for image captioning\.IEEE Transactions on Image Processing29,pp\. 7615–7628\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[22\]N\. S\. Keskar, B\. McCann, L\. R\. Varshney, C\. Xiong, and R\. Socher\(2019\)Ctrl: a conditional transformer language model for controllable generation\.arXiv preprint arXiv:1909\.05858\.Cited by:[§II\-A](https://arxiv.org/html/2609.22135#S2.SS1.p1.1)\.
- \[23\]K\. Konen, S\. Jentzsch, D\. Diallo, P\. Schütt, O\. Bensch, R\. El Baff, D\. Opitz, and T\. Hecking\(2024\)Style vectors for steering generative large language models\.InFindings of the Association for Computational Linguistics: EACL 2024,pp\. 782–802\.Cited by:[§II\-A](https://arxiv.org/html/2609.22135#S2.SS1.p1.1)\.
- \[24\]B\. Krause, A\. D\. Gotmare, B\. McCann, N\. S\. Keskar, S\. Joty, R\. Socher, and N\. F\. Rajani\(2021\)Gedi: generative discriminator guided sequence generation\.InFindings of the Association for Computational Linguistics: EMNLP 2021,pp\. 4929–4952\.Cited by:[§II\-A](https://arxiv.org/html/2609.22135#S2.SS1.p1.1)\.
- \[25\]B\. W\. Lee, I\. Padhi, K\. Natesan Ramamurthy, E\. Miehling, P\. Dognin, M\. Nagireddy, and A\. Dhurandhar\(2025\)Programming refusal with conditional activation steering\.InInternational conference on learning representations,Vol\.2025,pp\. 90960–90985\.Cited by:[§II\-A](https://arxiv.org/html/2609.22135#S2.SS1.p1.1)\.
- \[26\]J\. Li, D\. Li, S\. Savarese, and S\. Hoi\(2023\)Blip\-2: bootstrapping language\-image pre\-training with frozen image encoders and large language models\.InInternational conference on machine learning,pp\. 19730–19742\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[27\]K\. Li, O\. Patel, F\. Viégas, H\. Pfister, and M\. Wattenberg\(2023\)Inference\-time intervention: eliciting truthful answers from a language model\.Advances in Neural Information Processing Systems36,pp\. 41451–41530\.Cited by:[§II\-A](https://arxiv.org/html/2609.22135#S2.SS1.p1.1)\.
- \[28\]L\. Li, H\. Zhu, S\. Zhao, G\. Ding, and W\. Lin\(2020\)Personality\-assisted multi\-task learning for generic and personalized image aesthetics assessment\.IEEE Transactions on Image Processing29,pp\. 3898–3910\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[29\]Y\. Li, H\. Su, X\. Shen, W\. Li, Z\. Cao, and S\. Niu\(2017\)Dailydialog: a manually labelled multi\-turn dialogue dataset\.InProceedings of the Eighth International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),pp\. 986–995\.Cited by:[§III\-B](https://arxiv.org/html/2609.22135#S3.SS2.p2.1)\.
- \[30\]Y\. Li, X\. Huang, and G\. Zhao\(2020\)Joint local and global information learning with single apex frame detection for micro\-expression recognition\.IEEE Transactions on Image Processing30,pp\. 249–263\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[31\]H\. Liu, S\. Zhang, K\. Lin, J\. Wen, J\. Li, and X\. Hu\(2021\)Vocabulary\-wide credit assignment for training image captioning models\.IEEE Transactions on Image Processing30,pp\. 2450–2460\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[32\]H\. Liu, C\. Li, Q\. Wu, and Y\. J\. Lee\(2023\)Visual instruction tuning\.Advances in neural information processing systems36,pp\. 34892–34916\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[33\]S\. Liu, H\. Ye, L\. Xing, and J\. Zou\(2023\)In\-context vectors: making in context learning more effective and controllable through latent space steering\.arXiv preprint arXiv:2311\.06668\.Cited by:[§II\-A](https://arxiv.org/html/2609.22135#S2.SS1.p1.1)\.
- \[34\]S\. R\. Livingstone and F\. A\. Russo\(2018\)The ryerson audio\-visual database of emotional speech and song \(ravdess\): a dynamic, multimodal set of facial and vocal expressions in north american english\.PloS one13\(5\),pp\. e0196391\.Cited by:[§III\-B](https://arxiv.org/html/2609.22135#S3.SS2.p2.1)\.
- \[35\]A\. Mehrabian\(1996\)Pleasure\-arousal\-dominance: a general framework for describing and measuring individual differences in temperament\.Current psychology14\(4\),pp\. 261–292\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[36\]K\. Meng, D\. Bau, A\. Andonian, and Y\. Belinkov\(2022\)Locating and editing factual associations in gpt\.Advances in neural information processing systems35,pp\. 17359–17372\.Cited by:[§II\-B](https://arxiv.org/html/2609.22135#S2.SS2.p1.1)\.
- \[37\]S\. Merity, C\. Xiong, J\. Bradbury, and R\. Socher\(2016\)Pointer sentinel mixture models\.arXiv preprint arXiv:1609\.07843\.Cited by:[§VI\-E](https://arxiv.org/html/2609.22135#S6.SS5.p1.1)\.
- \[38\]X\. Min, G\. Zhai, J\. Zhou, X\. Zhang, X\. Yang, and X\. Guan\(2020\)A multimodal saliency model for videos with high audio\-visual correspondence\.IEEE Transactions on Image Processing29,pp\. 3805–3819\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[39\]S\. Mohammad\(2018\)Obtaining reliable human ratings of valence, arousal, and dominance for 20,000 english words\.InProceedings of the 56th annual meeting of the association for computational linguistics \(volume 1: Long papers\),pp\. 174–184\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[40\]nostalgebraist\(2020\)Interpreting GPT: the logit lens\.Note:LessWrong[https://www\.lesswrong\.com/posts/AcKRB8wDpdaN6v6ru/interpreting\-gpt\-the\-logit\-lens](https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens)Cited by:[§II\-B](https://arxiv.org/html/2609.22135#S2.SS2.p1.1),[§III\-E](https://arxiv.org/html/2609.22135#S3.SS5.p1.2),[§V\-D](https://arxiv.org/html/2609.22135#S5.SS4.p1.1)\.
- \[41\]N\. Panickssery, N\. Gabrieli, J\. Schulz, M\. Tong, E\. Hubinger, and A\. M\. Turner\(2023\)Steering llama 2 via contrastive activation addition\.arXiv preprint arXiv:2312\.06681\.Cited by:[§II\-A](https://arxiv.org/html/2609.22135#S2.SS1.p1.1)\.
- \[42\]K\. Park, Y\. J\. Choe, Y\. Jiang, and V\. Veitch\(2025\)The geometry of categorical and hierarchical concepts in large language models\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 76441–76463\.Cited by:[§II\-B](https://arxiv.org/html/2609.22135#S2.SS2.p1.1)\.
- \[43\]K\. Park, Y\. J\. Choe, and V\. Veitch\(2023\)The linear representation hypothesis and the geometry of large language models\.arXiv preprint arXiv:2311\.03658\.Cited by:[§II\-B](https://arxiv.org/html/2609.22135#S2.SS2.p1.1)\.
- \[44\]J\. Posner, J\. A\. Russell, and B\. S\. Peterson\(2005\)The circumplex model of affect: an integrative approach to affective neuroscience, cognitive development, and psychopathology\.Development and psychopathology17\(3\),pp\. 715–734\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[45\]A\. Radford, J\. W\. Kim, C\. Hallacy, A\. Ramesh, G\. Goh, S\. Agarwal, G\. Sastry, A\. Askell, P\. Mishkin, J\. Clark,et al\.\(2021\)Learning transferable visual models from natural language supervision\.InInternational conference on machine learning,pp\. 8748–8763\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[46\]A\. Radford, J\. W\. Kim, T\. Xu, G\. Brockman, C\. McLeavey, and I\. Sutskever\(2023\)Robust speech recognition via large\-scale weak supervision\.InInternational conference on machine learning,pp\. 28492–28518\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[47\]B\. Reichman, A\. Avsian, and L\. Heck\(2025\)Emotions where art thou: understanding and characterizing the emotional latent space of large language models\.arXiv preprint arXiv:2510\.22042\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[48\]J\. A\. Russell\(1980\)A circumplex model of affect\.\.Journal of personality and social psychology39\(6\),pp\. 1161\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1),[§IV](https://arxiv.org/html/2609.22135#S4.p2.1)\.
- \[49\]K\. R\. Scherer and H\. G\. Wallbott\(1994\)Evidence for universality and cultural variation of differential emotion response patterning\.\.Journal of personality and social psychology66\(2\),pp\. 310\.Cited by:[§III\-B](https://arxiv.org/html/2609.22135#S3.SS2.p2.1)\.
- \[50\]A\. N\. Tak, A\. Banayeeanzade, A\. Bolourani, M\. Kian, R\. Jia, and J\. Gratch\(2025\)Mechanistic interpretability of emotion inference in large language models\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 13090–13120\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[51\]D\. Tan, D\. Chanin, A\. Lynch, B\. Paige, D\. Kanoulas, A\. Garriga\-Alonso, and R\. Kirk\(2024\)Analysing the generalisation and reliability of steering vectors\.Advances in Neural Information Processing Systems37,pp\. 139179–139212\.Cited by:[§II\-A](https://arxiv.org/html/2609.22135#S2.SS1.p1.1)\.
- \[52\]R\. J\. Tibshirani and B\. Efron\(1993\)An introduction to the bootstrap\.Monographs on statistics and applied probability57\(1\),pp\. 1–436\.Cited by:[§III\-C](https://arxiv.org/html/2609.22135#S3.SS3.p1.2)\.
- \[53\]E\. Todd, M\. Li, A\. Sen Sharma, A\. Mueller, B\. Wallace, and D\. Bau\(2024\)Function vectors in large language models\.InInternational conference on learning representations,Vol\.2024,pp\. 17282–17333\.Cited by:[§II\-A](https://arxiv.org/html/2609.22135#S2.SS1.p1.1)\.
- \[54\]A\. M\. Turner, L\. Thiergart, G\. Leech, D\. Udell, J\. J\. Vazquez, U\. Mini, and M\. MacDiarmid\(2023\)Steering language models with activation engineering\.arXiv preprint arXiv:2308\.10248\.Cited by:[§I](https://arxiv.org/html/2609.22135#S1.p1.1),[§II\-A](https://arxiv.org/html/2609.22135#S2.SS1.p1.1)\.
- \[55\]C\. Wang, Y\. Zhang, R\. Yu, Y\. Zheng, L\. Gao, Z\. Song, Z\. Xu, G\. Xia, H\. Zhang, D\. Zhao,et al\.\(2025\)Do llms” feel”? emotion circuits discovery and control\.arXiv preprint arXiv:2510\.11328\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[56\]K\. Wang, X\. Peng, J\. Yang, D\. Meng, and Y\. Qiao\(2020\)Region attention networks for pose and occlusion robust facial expression recognition\.IEEE Transactions on Image Processing29,pp\. 4057–4069\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[57\]K\. Wang, A\. Variengien, A\. Conmy, B\. Shlegeris, and J\. Steinhardt\(2022\)Interpretability in the wild: a circuit for indirect object identification in gpt\-2 small\.arXiv preprint arXiv:2211\.00593\.Cited by:[§II\-B](https://arxiv.org/html/2609.22135#S2.SS2.p1.1)\.
- \[58\]Z\. Xia, W\. Peng, H\. Khor, X\. Feng, and G\. Zhao\(2020\)Revealing the invisible with model and data shrinking for composite\-database micro\-expression recognition\.IEEE Transactions on Image Processing29,pp\. 8590–8605\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[59\]A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, H\. Wei,et al\.\(2024\)Qwen2\. 5 technical report\.arXiv preprint arXiv:2412\.15115\.Cited by:[§III\-A](https://arxiv.org/html/2609.22135#S3.SS1.p1.1)\.
- \[60\]J\. Yang, X\. Gao, L\. Li, X\. Wang, and J\. Ding\(2021\)Solver: scene\-object interrelated visual emotion reasoning network\.IEEE Transactions on Image Processing30,pp\. 8686–8701\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[61\]J\. Yang, Q\. Huang, T\. Ding, D\. Lischinski, D\. Cohen\-Or, and H\. Huang\(2023\)Emoset: a large\-scale visual emotion dataset with rich attributes\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 20383–20394\.Cited by:[§III\-B](https://arxiv.org/html/2609.22135#S3.SS2.p2.1)\.
- \[62\]J\. Yang, J\. Li, L\. Li, X\. Wang, Y\. Ding, and X\. Gao\(2022\)Seeking subjectivity in visual emotion distribution learning\.IEEE Transactions on Image Processing31,pp\. 5189–5202\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[63\]J\. Yang, J\. Li, X\. Wang, Y\. Ding, and X\. Gao\(2021\)Stimuli\-aware visual emotion analysis\.IEEE Transactions on Image Processing30,pp\. 7432–7445\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[64\]Y\. Yao, T\. Yu, A\. Zhang, C\. Wang, J\. Cui, H\. Zhu, T\. Cai, H\. Li, W\. Zhao, Z\. He,et al\.\(2024\)Minicpm\-v: a gpt\-4v level mllm on your phone\.arXiv preprint arXiv:2408\.01800\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[65\]X\. Zhai, B\. Mustafa, A\. Kolesnikov, and L\. Beyer\(2023\)Sigmoid loss for language image pre\-training\.InProceedings of the IEEE/CVF international conference on computer vision,pp\. 11975–11986\.Cited by:[§II\-C](https://arxiv.org/html/2609.22135#S2.SS3.p1.1)\.
- \[66\]L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. Xing,et al\.\(2023\)Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.Advances in neural information processing systems36,pp\. 46595–46623\.Cited by:[§III\-E](https://arxiv.org/html/2609.22135#S3.SS5.p1.2)\.
- \[67\]A\. Zou, L\. Phan, S\. Chen, J\. Campbell, P\. Guo, R\. Ren, A\. Pan, X\. Yin, M\. Mazeika, A\. Dombrowski,et al\.\(2023\)Representation engineering: a top\-down approach to ai transparency\.arXiv preprint arXiv:2310\.01405\.Cited by:[§I](https://arxiv.org/html/2609.22135#S1.p1.1),[§II\-A](https://arxiv.org/html/2609.22135#S2.SS1.p1.1)\.

## XBiography Section

![[Uncaptioned image]](https://arxiv.org/html/2609.22135v1/figures/wyb.jpg)Yibo Wangis currently pursuing the M\.S\. degree at the Gansu Provincial Key Laboratory of Wearable Computing, School of Information Science and Engineering, Lanzhou University, Lanzhou, China\. He received the B\.S\. degree from Dalian University of Technology, Dalian, China\. His research interests include multimodal large language models, affective computing, and reinforcement learning\.![[Uncaptioned image]](https://arxiv.org/html/2609.22135v1/figures/djs1.jpg)Jisheng Dangreceived the Ph\.D\. degree from Sun Yat\-sen University, China, advised by Prof\. Jianhuang Lai and Prof\. Huicheng Zheng\. He worked as a research fellow at the NExT\+\+ laboratory of the National University of Singapore, advised by Prof\. Tat\-Seng Chua\. He is now a tenured associate professor at the School of Information Science and Engineering, Lanzhou University\. His research interests include multimodal learning, video understanding, and embodied intelligence\. He has published several papers as the first author or corresponding author in major journals and conferences including IEEE TPAMI/TIP/TNNLS/TITS/IJCAI/AAAI/ICLR/NeurIPS/PR\.![[Uncaptioned image]](https://arxiv.org/html/2609.22135v1/figures/wyt.jpg)Yitao Wureceived the B\.S\. degree in information system and information management from Hainan University\. His research interests include large language models, vision\-language\-action models, and vision\-language models\.![[Uncaptioned image]](https://arxiv.org/html/2609.22135v1/figures/ph.png)Hong Pengreceived the Ph\.D\. degree from Lanzhou University, Lanzhou, China\. From 2010 to 2011, he was a Visiting Scholar with the Institute of Computer System, ETH Zurich, Switzerland\. He is currently an Associate Professor with the School of Information Science and Engineering, Lanzhou University\. He is also in charge of three projects from the National Natural Science Foundation of China, the Central College Foundation Project of Lanzhou University, and the Youth Cross\-Project of Lanzhou University\. He has authored or coauthored more than 30 papers in peer\-reviewed journals, conferences, and book chapters\. His research areas include bioinformation processing and ubiquitous affective computing\.![[Uncaptioned image]](https://arxiv.org/html/2609.22135v1/figures/binhu.png)Bin Hu \(Fellow, IEEE\)received the Ph\.D\. degree in computer science from the Institute of Computing Technology, Chinese Academy of Sciences, China, in 1998\. Since 2008, he has been a Professor and Dean of the School of Information Science and Engineering, Lanzhou University\. He has also held a guest professorship at ETH Zurich\. He serves as Editor\-in\-Chief of IEEE Transactions on Computational Social Systems and is a Fellow of IET and AAIA\. His research interests include pervasive computing, computational psychophysiology, data modeling, and artificial intelligence\.![[Uncaptioned image]](https://arxiv.org/html/2609.22135v1/figures/tianqi.png)Qi Tian \(Fellow, IEEE\)received the Ph\.D\. degree in ECE from the University of Illinois at Urbana\-Champaign, Champaign, IL, USA, in 2002\. He is currently the Chief Scientist of Huawei Terminal BG\. He was the Chief Scientist in computer vision with Huawei Noah’s Ark Laboratory from 2018 to 2020\. Before he joined Huawei, he was a Full Professor with the Department of Computer Science, The University of Texas at San Antonio, San Antonio, TX, USA, from 2002 to 2019\. He was listed in the top ten of the 2016 Most Influential Scholars in Multimedia by Aminer\.org\. He is an IEEE Fellow \(Class 2016\) and an Academician of the International Eurasian Academy of Sciences \(elected in 2021\), with 113,457 citations on Google Scholar\.![[Uncaptioned image]](https://arxiv.org/html/2609.22135v1/figures/Tat-SengChua.png)Tat\-Seng Chuareceived the Ph\.D\. degree from the University of Leeds, U\.K\. He is the KITHCT chair professor with the School of Computing, National University of Singapore, where he was the acting and founding dean of the School from 1998 to 2000\. He is the co\-director of NExT, a joint center between NUS and Tsinghua University, to develop technologies for live social media search\. He is the 2015 winner of the prestigious ACM SIGMM Award\. He is the chair of the Steering Committee of the ACM International Conference on Multimedia Retrieval \(ICMR\) and the Multimedia Modeling \(MMM\) conference series\. He is also the general co\-chair of ACM Multimedia 2005, ACM CIVR \(now ACM ICMR\) 2005, ACM SIGIR 2008, and ACM Web Science 2015\. He serves on the editorial boards of four international journals\. He is the co\-founder of two technology startups in Singapore and a Fellow of the Singapore Academy of Sciences, with 114,163 citations on Google Scholar\.

Similar Articles

Decomposing and Steering Functional Metacognition in Large Language Models

arXiv cs.CL

This research paper investigates functional metacognition in Large Language Models, demonstrating that internal states like evaluation awareness and self-assessed capability are linearly decodable from residual stream activations. The authors propose a mechanistic framework to steer these states, showing causal control over reasoning behaviors, verbosity, and safety responses.