Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models
摘要
This preprint introduces a generation-aligned diagnostic ladder that separates decision-rule misalignment from readout-coverage limitations in speech language models, showing that state decoding far exceeds generated accuracy in emotion recognition tasks.
arXiv:2608.06409v1 Announce Type: new
Abstract: Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines failures at different stages of the audio-to-answer computation. We introduce a generation-aligned diagnostic ladder that compares the emitted answer, the option logits, an affine readout of those logits, and a linear readout of the hidden state at the same answer token. Successive differences separate endpoint, decision-rule, and readout-coverage gaps. Across five systems and two emotion corpora, state decoding exceeds generation by 27.8 accuracy points on average, and both the decision-rule and readout-coverage gaps are positive in all ten conditions. A label-free logit correction improves generated accuracy in every condition, showing that part of the decision-rule gap is actionable. In rank-matched comparisons, emotion information outside the native readout generalizes to held-out speakers and survives controls for measured acoustic descriptors, but replacing the selected readout-external directions usually has little effect on emitted answers. These results distinguish information availability from behavioral use and localize performance losses across the decision rule and the state-to-answer readout.
查看缓存全文
缓存时间: 2026/08/10 08:00
# Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models
Source: [https://arxiv.org/html/2608.06409](https://arxiv.org/html/2608.06409)
Linkai Peng and Baorian NuchgedLinkai Peng is with the Institute for the Brain and Cognitive Sciences, University of Connecticut, Storrs, CT 06269, USA \(e\-mail: linkai\.peng@uconn\.edu\)\. Baorian Nuchged is with the Department of Linguistics, The University of Texas at Austin, Austin, TX 78712, USA \(e\-mail: baorian@utexas\.edu\)\.Preprint\.
###### Abstract
Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines failures at different stages of the audio\-to\-answer computation\. We introduce a generation\-aligned diagnostic ladder that compares the emitted answer, the option logits, an affine readout of those logits, and a linear readout of the hidden state at the same answer token\. Successive differences separate endpoint, decision\-rule, and readout\-coverage gaps\. Across five systems and two emotion corpora, state decoding exceeds generation by 27\.8 accuracy points on average, and both the decision\-rule and readout\-coverage gaps are positive in all ten conditions\. A label\-free logit correction improves generated accuracy in every condition, showing that part of the decision\-rule gap is actionable\. In rank\-matched comparisons, emotion information outside the native readout generalizes to held\-out speakers and survives controls for measured acoustic descriptors, but replacing the selected readout\-external directions usually has little effect on emitted answers\. These results distinguish information availability from behavioral use and localize performance losses across the decision rule and the state\-to\-answer readout\.
## IIntroduction
Speech language models couple audio interfaces with generative language backbones, allowing a single system to process speech and respond through natural\-language generation\[[1](https://arxiv.org/html/2608.06409#bib.bib1),[2](https://arxiv.org/html/2608.06409#bib.bib2),[3](https://arxiv.org/html/2608.06409#bib.bib3),[4](https://arxiv.org/html/2608.06409#bib.bib4),[5](https://arxiv.org/html/2608.06409#bib.bib5)\]\. Their evaluation has expanded beyond transcription to emotion, prosody, and other paralinguistic judgments, typically through prompted multiple\-choice or free\-form responses\[[6](https://arxiv.org/html/2608.06409#bib.bib6),[7](https://arxiv.org/html/2608.06409#bib.bib7),[8](https://arxiv.org/html/2608.06409#bib.bib8)\]\. Generation accuracy provides a convenient summary of behavioral performance, but it conflates at least three distinct failure locations: the relevant evidence may never reach the language\-model component; it may be retained in the state without being exposed through the answer readout; or it may reach the answer logits yet be misread by a mismatched decision rule\. These possibilities demand different interventions, yet behavioral accuracy alone cannot distinguish among them\.
We ground this localization problem in four\-class speech emotion recognition, whose acoustic correlates are well characterized\[[9](https://arxiv.org/html/2608.06409#bib.bib9),[10](https://arxiv.org/html/2608.06409#bib.bib10)\]and whose answer can be elicited as a single token\. Across ten system–corpus conditions, generated accuracy averages0\.4740\.474, whereas linear decoding from the hidden state at the same answer position averages0\.7520\.752; this deficit of0\.2780\.278in absolute accuracy persists under speaker\-disjoint evaluation and appears in every condition\. It is most pronounced in relative terms for Phi\-4\-MM on CREMA\-D\[[11](https://arxiv.org/html/2608.06409#bib.bib11)\], where generation reaches0\.2790\.279against a full\-state probe at0\.7220\.722\. These results reveal a substantial state\-to\-answer loss even when emotion remains linearly decodable at the answer position\.
This discrepancy points to two distinct failure locations within the language model, which we introduce in the order encountered when tracing back from the emitted answer\. First, closest to the answer, emotion evidence may be present in the option logits while the default decision rule uses those scores inefficiently\. We call this*decision\-rule misalignment*\. Second, deeper in the state, additional evidence may remain accessible at the answer position without being expressed through the option\-specific coordinates that determine the generated label\. We call this a*readout\-coverage gap*\. The practical consequence is that decision\-rule misalignment is amenable to logit correction, whereas a coverage gap cannot be recovered by reweighting the same option logits and requires moving beyond the native readout\. Phi\-4\-MM on CREMA\-D makes this concrete\. The best affine rule over its option logits reaches an accuracy of only0\.3540\.354, far below the0\.7220\.722supported by the full state, so most of its deficit cannot be repaired at the option logits\.
To localize these failures, we construct a generation\-aligned diagnostic ladder that compares actual generation, the default option\-logit decision, an optimized affine decision over the same logits, and regularized linear decoding from the full state\. Anchoring all four levels to one verified answer token ensures that behavior, logits, and states measure the same event\. Their successive differences split the distance between generated accuracy and full\-state decodability into three terms that sum exactly: endpoint validity \(agreement between the emitted answer and the option favored by the model’s own logits\), the decision\-rule gap, and the readout\-coverage gap\. Because the full\-state reader has many more input dimensions than the contrast reader, we treat the readout\-coverage gap as a performance gap; claims about readout\-external information rest on rank\-matched comparisons\. We then connect diagnosis to behavior\. A label\-free logit correction tests the decision\-rule gap during generation, while held\-out decoding and minimal\-pair subspace interventions distinguish the availability of readout\-external information from its causal use\. Acoustic controls characterize how much of that information is explained by measured surface cues\.
These analyses yield four contributions\. First, we introduce a generation\-aligned framework that exactly decomposes the distance between emitted behavior and full\-state decodability into endpoint validity, a decision\-rule gap, and a readout\-coverage gap\. Second, we show that the decision\-rule diagnosis is behaviorally actionable\. A standard label\-free logit correction\[[12](https://arxiv.org/html/2608.06409#bib.bib12),[13](https://arxiv.org/html/2608.06409#bib.bib13)\]improves the emitted answer across all ten conditions with negligible format cost \(Section[V\-B](https://arxiv.org/html/2608.06409#S5.SS2)\)\. Third, in rank\-matched comparisons, we identify emotion information that remains linearly accessible outside the native readout and generalizes to held\-out speakers\. Fourth, matched minimal\-pair interventions show that the selected readout\-external directions have limited influence on the emitted answer, separating information availability from causal use\.
## IIRelated Work
#### Prosody\-sensitive evaluation of speech language models
Speech language models route speech through an encoder and projector into a text LLM\[[14](https://arxiv.org/html/2608.06409#bib.bib14),[2](https://arxiv.org/html/2608.06409#bib.bib2),[15](https://arxiv.org/html/2608.06409#bib.bib15),[16](https://arxiv.org/html/2608.06409#bib.bib16),[1](https://arxiv.org/html/2608.06409#bib.bib1)\]\. Benchmarks such as Dynamic\-SUPERB, AIR\-Bench, and SD\-Eval include emotion and paralinguistic tasks\[[6](https://arxiv.org/html/2608.06409#bib.bib6),[7](https://arxiv.org/html/2608.06409#bib.bib7),[8](https://arxiv.org/html/2608.06409#bib.bib8)\], and controlled studies show that models often rely more on lexical than acoustic cues\[[17](https://arxiv.org/html/2608.06409#bib.bib17)\]\. These works establish behavioral gaps but do not localize whether a cue is lost, attenuated, or retained but underused\. We study speech without engineered text–audio conflict and localize these downstream failures within the audio\-to\-answer computation\.
#### Label\-free correction of class\-dependent readout offsets
Prompted multiple\-choice answers exhibit systematic position and token biases\[[18](https://arxiv.org/html/2608.06409#bib.bib18)\]\. Prior label\-free methods estimate such preferences from content\-free inputs\[[12](https://arxiv.org/html/2608.06409#bib.bib12)\], answer\-side scoring statistics\[[19](https://arxiv.org/html/2608.06409#bib.bib19)\], or the mean predicted distribution over an unlabeled batch\[[13](https://arxiv.org/html/2608.06409#bib.bib13)\]\. We adopt the latter estimator but use it as a behavioral intervention rather than offline rescoring, applying the offset at the first answer step, resuming full\-vocabulary generation, and scoring the emitted string\. This measures both accuracy gain and format stability, making the correction a behavioral test of the decision\-rule gap\.
#### Representation probing in speech models
Beyond the option scores, linear probing shows that emotion and prosody are recoverable from frozen speech encoders\[[20](https://arxiv.org/html/2608.06409#bib.bib20),[21](https://arxiv.org/html/2608.06409#bib.bib21),[22](https://arxiv.org/html/2608.06409#bib.bib22),[23](https://arxiv.org/html/2608.06409#bib.bib23),[24](https://arxiv.org/html/2608.06409#bib.bib24),[25](https://arxiv.org/html/2608.06409#bib.bib25),[26](https://arxiv.org/html/2608.06409#bib.bib26)\]\. Layer\-wise probing inside speech language models further shows that such attributes remain recoverable deep in the language stack\[[27](https://arxiv.org/html/2608.06409#bib.bib27)\]\. However, decodability does not imply use\[[28](https://arxiv.org/html/2608.06409#bib.bib28)\]; we therefore treat probing as an availability diagnostic rather than evidence that the complete system recruits the cue\.
#### Subspace interventions and the selection trap
The logit lens and hidden\-state interventions provide tools for localizing computation\[[29](https://arxiv.org/html/2608.06409#bib.bib29),[30](https://arxiv.org/html/2608.06409#bib.bib30),[31](https://arxiv.org/html/2608.06409#bib.bib31),[32](https://arxiv.org/html/2608.06409#bib.bib32),[33](https://arxiv.org/html/2608.06409#bib.bib33)\]\. Supervised subspace selection, however, can confound decodability with mechanism\[[34](https://arxiv.org/html/2608.06409#bib.bib34)\]\. We address this problem by fixing the native readout space from the output head and treating held\-out decoding and matched replacement as separate measurements of availability and use\. In text\-only models, a related knowledge–prediction gap has been reported on multiple\-choice questions\[[35](https://arxiv.org/html/2608.06409#bib.bib35)\]; our framework additionally separates readout coverage from the decision rule over option scores and anchors both to the generated answer\. Concurrent work retrieves sparse audio concepts\[[36](https://arxiv.org/html/2608.06409#bib.bib36)\]and studies text–audio conflict\[[37](https://arxiv.org/html/2608.06409#bib.bib37)\]; our focus is the availability and causal use of readout\-external information in ordinary prompted speech\.
## IIIDiagnosing the State\-to\-Answer Interface
Figure[1](https://arxiv.org/html/2608.06409#S3.F1)gives an overview of this section\. It develops a diagnostic ladder that aligns the emitted token, the option scores, and the underlying hidden state at the same first\-token event, and a decomposition of that state around the answer readout\. Both concern the last mile; they do not by themselves localize losses earlier in the audio pathway\.
Figure 1:Overview of the generation\-aligned diagnostic ladder\.Left: audio and prompt tokens are processed by the encoder, projector, and language model, and the output headWWmaps the answer\-position statextx\_\{t\}to full\-vocabulary logits\. Four readouts score the same answer event: greedy generation over the full vocabulary \(AgenA\_\{\\mathrm\{gen\}\}\), the post\-hoc argmax over the four option logits \(AoptA\_\{\\mathrm\{opt\}\}\), a learned affine reader on the option\-logit contrasts \(AaffA\_\{\\mathrm\{aff\}\}\), and a learned affine reader on the full state \(AstateA\_\{\\mathrm\{state\}\}\); dashed boxes mark the learned readers, and successive differences among the four accuracies give the gaps in Eq\. \([3](https://arxiv.org/html/2608.06409#S3.E3)\)\. Right:xtx\_\{t\}decomposes into its component in the answer\-readout span \(PVxtP\_\{V\}x\_\{t\}\) and the readout\-external complementV⟂V^\{\\perp\}, within which a supervised\-selected subspaceSdecodingS\_\{\\mathrm\{decoding\}\}is compared against same\-rank random subspaces \(Sections[III\-D](https://arxiv.org/html/2608.06409#S3.SS4)and[III\-E](https://arxiv.org/html/2608.06409#S3.SS5)\)\.### III\-AThree Views for One Answer
For each audio–prompt pair, letxt∈ℝdx\_\{t\}\\in\\mathbb\{R\}^\{d\}be the model\-native, post\-normalization hidden state used to predict the first answer token,𝒱\\mathcal\{V\}the vocabulary, andW∈ℝ\|𝒱\|×dW\\in\\mathbb\{R\}^\{\|\\mathcal\{V\}\|\\times d\}the language\-model output head\. For option\-token identifiersv1,…,v4v\_\{1\},\\ldots,v\_\{4\}, define
zt\\displaystyle z\_\{t\}=Wxt,\\displaystyle=Wx\_\{t\},\(1\)zt,opt\\displaystyle z\_\{t,\\mathrm\{opt\}\}=\(zt,v1,…,zt,v4\),\\displaystyle=\\big\(z\_\{t,v\_\{1\}\},\\ldots,z\_\{t,v\_\{4\}\}\\big\),yt\\displaystyle y\_\{t\}=argmaxv∈𝒱zt,v,\\displaystyle=\\operatorname\*\{arg\\,max\}\_\{v\\in\\mathcal\{V\}\}z\_\{t,v\},
at the same inference step\. Thus,yty\_\{t\},zt,optz\_\{t,\\mathrm\{opt\}\}, andxtx\_\{t\}provide three views of one answer event: emitted behavior, native option preference, and the state available to the readout\.
### III\-BA Four\-Level Diagnostic Ladder
We operationalize these views as four levels of emotion\-classification performance \(Fig\.[1](https://arxiv.org/html/2608.06409#S3.F1), left\), all evaluated on the same rows and with the same scoring rule; learned readers are assessed on the same held\-out splits:
- •AgenA\_\{\\mathrm\{gen\}\}: accuracy of the model’s actual generated answer, with any response outside the required format scored as incorrect;
- •AoptA\_\{\\mathrm\{opt\}\}: accuracy of the post\-hoc argmax restricted to the four option logits;
- •AaffA\_\{\\mathrm\{aff\}\}: accuracy of a learned affine reader applied to three reference\-relative option\-logit contrasts;
- •AstateA\_\{\\mathrm\{state\}\}: accuracy of a regularized affine reader applied to the full answer\-position state\.
An off\-option emission or a mismatch in native tie\-breaking can makeAgenA\_\{\\mathrm\{gen\}\}differ fromAoptA\_\{\\mathrm\{opt\}\}\. For the two learned readers, letct=\(zt,v1−zt,v4,zt,v2−zt,v4,zt,v3−zt,v4\)⊤c\_\{t\}=\(z\_\{t,v\_\{1\}\}\-z\_\{t,v\_\{4\}\},\\,z\_\{t,v\_\{2\}\}\-z\_\{t,v\_\{4\}\},\\,z\_\{t,v\_\{3\}\}\-z\_\{t,v\_\{4\}\}\)^\{\\top\}, and sethaff=cth\_\{\\mathrm\{aff\}\}=c\_\{t\}andhstate=xth\_\{\\mathrm\{state\}\}=x\_\{t\}\. Both predictions take the form
y^r=argmaxk\(Urhr\+br\)k,r∈\{aff,state\}\.\\hat\{y\}\_\{r\}=\\operatorname\*\{arg\\,max\}\_\{k\}\(U\_\{r\}h\_\{r\}\+b\_\{r\}\)\_\{k\},\\qquad r\\in\\\{\\mathrm\{aff\},\\mathrm\{state\}\\\}\.\(2\)AaffA\_\{\\mathrm\{aff\}\}andAstateA\_\{\\mathrm\{state\}\}are the corresponding held\-out accuracies\. Both readers are affine and include a bias;UUdistinguishes their learned weights from the fixed output headWW, and the subscript identifies only whether the input isctc\_\{t\}orxtx\_\{t\}\.
Sharing a function class makes the comparison interpretable\. The contrast reader can relearn combinations and offsets of the existing option contrasts but cannot access information outside them, whereas the state reader applies the same rule type to the complete state\. Their successive differences telescope:
Astate−Agen=Aopt−Agen⏟Δendpoint\+Aaff−Aopt⏟Δdecision\+Astate−Aaff⏟Δcoverage\.A\_\{\\mathrm\{state\}\}\-A\_\{\\mathrm\{gen\}\}=\\underbrace\{A\_\{\\mathrm\{opt\}\}\-A\_\{\\mathrm\{gen\}\}\}\_\{\\Delta\_\{\\mathrm\{endpoint\}\}\}\+\\underbrace\{A\_\{\\mathrm\{aff\}\}\-A\_\{\\mathrm\{opt\}\}\}\_\{\\Delta\_\{\\mathrm\{decision\}\}\}\+\\underbrace\{A\_\{\\mathrm\{state\}\}\-A\_\{\\mathrm\{aff\}\}\}\_\{\\Delta\_\{\\mathrm\{coverage\}\}\}\.\(3\)
Δendpoint\\Delta\_\{\\mathrm\{endpoint\}\}is the accuracy difference between emitted generation and the option\-only decision\. It primarily captures emission and formatting behavior, including off\-option responses and native tie\-breaking mismatches, so we report it without a separate mechanistic analysis\. A near\-zeroΔendpoint\\Delta\_\{\\mathrm\{endpoint\}\}certifies that the option\-restricted view is a faithful anchor for the two explanatory gaps\.Δdecision\\Delta\_\{\\mathrm\{decision\}\}, the*decision\-rule gap*, measures the gain available from a better rule over the existing contrasts\.Δcoverage\\Delta\_\{\\mathrm\{coverage\}\}, the*readout\-coverage gap*, measures the additional performance supported by the full state\. Here,AstateA\_\{\\mathrm\{state\}\}is a decodability reference rather than attainable model performance\. The identity is exact by telescoping, and finite\-sample gap estimates are reported with uncertainty\. We analyze the two explanatory gaps in ladder order, beginning with the decision\-rule gap\.
### III\-CLogit Correction of the Decision\-Rule Gap
A positiveΔdecision\\Delta\_\{\\mathrm\{decision\}\}means that an affine rule improves on the native option\-only choice using the same contrasts\. One possible source is stable option bias\[[18](https://arxiv.org/html/2608.06409#bib.bib18)\]\. We test it with the unlabeled\-batch estimator, part of a broader family of label\-free corrections\[[13](https://arxiv.org/html/2608.06409#bib.bib13),[12](https://arxiv.org/html/2608.06409#bib.bib12),[19](https://arxiv.org/html/2608.06409#bib.bib19)\], and score its effect on subsequent generation\.
For each utterancexxunder a fixed prompt variant, letqk\(x\)q\_\{k\}\(x\)be the model’s probability for optionkk, normalized over the four prompted options\. Averaging this quantity over a condition\- and prompt\-matched unlabeled set𝒞\\mathcal\{C\}estimates how strongly the model favors each option overall\. We define
p^k=1\|𝒞\|∑xi∈𝒞qk\(xi\),bk=−logp^k\+14∑j=14logp^j,\\hat\{p\}\_\{k\}=\\frac\{1\}\{\|\\mathcal\{C\}\|\}\\sum\_\{x\_\{i\}\\in\\mathcal\{C\}\}q\_\{k\}\(x\_\{i\}\),\\qquad b\_\{k\}=\-\\log\\hat\{p\}\_\{k\}\+\\frac\{1\}\{4\}\\sum\_\{j=1\}^\{4\}\\log\\hat\{p\}\_\{j\},\(4\)
Here,p^k\\hat\{p\}\_\{k\}is the estimated average preference andbkb\_\{k\}reverses and centers it\. A frequently favored option receives a smaller offset, whereas an underpreferred option receives a larger one\. If all four options are favored equally, every offset is zero, so neither their relative ordering nor their competition with off\-option tokens changes\. Because a marginal option preference can also reflect the target class distribution, interpreting it as bias requires a target\-prior assumption; here the intended target prior is uniform, and the retained four\-class subsets are nearly balanced\.
We addbkb\_\{k\}to the corresponding option\-token logit only at the first answer step and then continue ordinary full\-vocabulary generation\. LetAoffsetA\_\{\\mathrm\{offset\}\}be the resulting accuracy under the same parser used forAgenA\_\{\\mathrm\{gen\}\}\. The gainAoffset−AgenA\_\{\\mathrm\{offset\}\}\-A\_\{\\mathrm\{gen\}\}therefore measures whether this simple correction improves the answers the model actually emits\. Because the gain is scored at the generation endpoint, where the offsets also shift the options’ competition with off\-option tokens, it tests the bias account ofΔdecision\\Delta\_\{\\mathrm\{decision\}\}rather than estimating that gap directly\. Ground\-truth labels are used only afterward to score the gain; they do not enter the correction\. In our experiments,𝒞\\mathcal\{C\}contains the target\-batch inputs themselves, making the procedure label\-free but transductive\.
### III\-DLocalizing the Readout\-Coverage Gap
A positiveΔcoverage\\Delta\_\{\\mathrm\{coverage\}\}shows that the full hidden state supports better linear emotion decoding than the three option\-logit contrasts\. It does not show where the additional decodable information lies, partly because the full\-state reader receives many more input dimensions\. This subsection therefore asks whether emotion information is concentrated in the model’s native answer readout or also remains available outside it\. We use rank\-matched comparisons here and report a matched\-budget refit in Supplementary Section[S4](https://arxiv.org/html/2608.06409#S4a)\.
We first identify the hidden\-state directions that can directly change the relative option logits\. Using the fourth option as reference, define
C=\[Wv1,:−Wv4,:Wv2,:−Wv4,:Wv3,:−Wv4,:\]∈ℝ3×d,V3=row\(C\)\.C=\\begin\{bmatrix\}W\_\{v\_\{1\},:\}\-W\_\{v\_\{4\},:\}\\\\ W\_\{v\_\{2\},:\}\-W\_\{v\_\{4\},:\}\\\\ W\_\{v\_\{3\},:\}\-W\_\{v\_\{4\},:\}\\end\{bmatrix\}\\in\\mathbb\{R\}^\{3\\times d\},\\qquad V\_\{3\}=\\operatorname\{row\}\(C\)\.\(5\)
We callV3V\_\{3\}the prompt\-specific*answer\-readout space*\(Fig\.[1](https://arxiv.org/html/2608.06409#S3.F1), right\)\. It has three dimensions because four option scores have three independent relative contrasts; the rows ofCCare linearly independent in the analyzed systems\. The output headWWis fixed; the prompt enters only by determining which four option\-token rows ofWWdefineCC\. The contrast reader observesct=Cxtc\_\{t\}=Cx\_\{t\}, so it can access only the component of the state inV3V\_\{3\}\. LetPV3P\_\{V\_\{3\}\}andPV3⟂P\_\{V\_\{3\}^\{\\perp\}\}denote the orthogonal projections ontoV3V\_\{3\}and its complement\. Then
xt=PV3xt\+PV3⟂xt,CPV3⟂xt=0\.x\_\{t\}=P\_\{V\_\{3\}\}x\_\{t\}\+P\_\{V\_\{3\}^\{\\perp\}\}x\_\{t\},\\qquad CP\_\{V\_\{3\}^\{\\perp\}\}x\_\{t\}=0\.\(6\)
Thus, the component inV3V\_\{3\}can directly change the relative logits of the four answer options\. Information inV3⟂V\_\{3\}^\{\\perp\}may still be present in the hidden state, but it cannot directly change these relative logits at the final answer position\. Throughout,*readout\-external*is used in this geometric sense, meaning outside the span of the option\-token rows of the output head, not unrelated to the task\.
This decomposition is exact at the final answer position\. To study how the same information is organized before the final readout, we move to an intermediate layerL∗L^\{\*\}\. We selectL∗L^\{\*\}on the training split as the layer with the highest logit\-lens accuracy, obtained by applying the model’s final normalization and output head to the answer\-position state \(Section[IV\-B](https://arxiv.org/html/2608.06409#S4.SS2)\)\. AtL∗L^\{\*\},V3V\_\{3\}is therefore a reference aligned with the final readout, not an exact decomposition of the final logits\.
For the external subspaces, we remove the full option\-row spanV4=row\(Wv1:v4\)V\_\{4\}=\\operatorname\{row\}\(W\_\{v\_\{1\}:v\_\{4\}\}\)\. BecauseV3⊂V4V\_\{3\}\\subset V\_\{4\}, their complements satisfyV4⟂⊂V3⟂V\_\{4\}^\{\\perp\}\\subset V\_\{3\}^\{\\perp\}\. A direction selected inV4⟂V\_\{4\}^\{\\perp\}is therefore also external to the relative option readout represented byV3V\_\{3\}\. The decoding comparisons below use rank\-three spaces matched toV3V\_\{3\}; the interventions in Section[III\-E](https://arxiv.org/html/2608.06409#S3.SS5)instead use rank\-four spaces matched toV4V\_\{4\}, which also carries the absolute option logits that matter during unconstrained generation\.
We conduct two rank\-three comparisons atL∗L^\{\*\}\. First, we compare held\-out emotion decoding fromV3V\_\{3\}with decoding from random three\-dimensional subspacesrand3⊂V4⟂⊂V3⟂\\mathrm\{rand\}\_\{3\}\\subset V\_\{4\}^\{\\perp\}\\subset V\_\{3\}^\{\\perp\}\. This tests whether the native answer\-readout directions carry more emotion information than a random readout\-external slice of the same dimension\. Neither space is selected using emotion labels, making this the appropriate comparison for evaluating the relative informativeness of the native readout\.
Second, we ask whether a generalizable emotion signal can be found outside the native readout\. For each prompt, we project the training states ontoV4⟂⊂V3⟂V\_\{4\}^\{\\perp\}\\subset V\_\{3\}^\{\\perp\}, fit a supervised multinomial logistic model, and defineSdecodingS\_\{\\mathrm\{decoding\}\}from the three leading right\-singular directions of its coefficient matrix\. We then fit a decoder inSdecodingS\_\{\\mathrm\{decoding\}\}and evaluate it on held\-out speakers\. An advantage overrand3\\mathrm\{rand\}\_\{3\}shows that selected readout\-external directions contain emotion information that generalizes beyond the training speakers\. BecauseSdecodingS\_\{\\mathrm\{decoding\}\}is selected using emotion labels whereasV3V\_\{3\}is not, their accuracies do not provide a direct ranking of the native and external spaces\. Construction details are given in Section[IV\-B](https://arxiv.org/html/2608.06409#S4.SS2)and the supplementary material\.
### III\-ECausal Interventions on Readout\-External Information
Held\-out decoding shows what information is available outside the readout, but not whether that information affects the model’s answer\. We therefore replace selected components of the answer\-position state atL∗L^\{\*\}and continue generation\. The decoding analysis usesV3V\_\{3\}, which represents the three relative contrasts among four options\. The intervention instead uses the full option\-row spanV4V\_\{4\}defined above, because unconstrained generation also depends on the absolute option logits and their competition with off\-option tokens\.
We compare three rank\-four spaces\.V4V\_\{4\}is the model’s native option\-readout space\.Sintervention⊂V4⟂S\_\{\\mathrm\{intervention\}\}\\subset V\_\{4\}^\{\\perp\}contains the four leading supervised readout\-external directions\. It extends the rank\-three decoding space, soSdecoding⊂SinterventionS\_\{\\mathrm\{decoding\}\}\\subset S\_\{\\mathrm\{intervention\}\}\. Finally,rand4⊂V4⟂\\mathrm\{rand\}\_\{4\}\\subset V\_\{4\}^\{\\perp\}is a same\-rank random control\. ReplacingV4V\_\{4\}tests whether the answer responds to information directly aligned with the native readout\. ReplacingSinterventionS\_\{\\mathrm\{intervention\}\}tests whether selected information outside that readout can influence the answer, whilerand4\\mathrm\{rand\}\_\{4\}controls for a generic state perturbation\.
For each held\-out minimal pair, the receiver and donor share the same speaker and transcript but express different emotions\. Lethrh\_\{r\}andhdh\_\{d\}be their answer\-position states atL∗L^\{\*\}under the same prompt\. ForS∈\{V4,Sintervention,rand4\}S\\in\\\{V\_\{4\},S\_\{\\mathrm\{intervention\}\},\\mathrm\{rand\}\_\{4\}\\\}, we construct
h~r\(S\)=hr\+PS\(hd−hr\)\.\\widetilde\{h\}\_\{r\}^\{\(S\)\}=h\_\{r\}\+P\_\{S\}\(h\_\{d\}\-h\_\{r\}\)\.\(7\)This operation replaces only the receiver’s component inSSwith the donor’s component\. We then continue generation from the edited state and measure two outcomes\. The answer\-change rateRchg\(S\)R\_\{\\mathrm\{chg\}\}\(S\)records any change from the receiver’s original answer\. The donor\-following rateRdon\(S\)R\_\{\\mathrm\{don\}\}\(S\)counts only the cases in which the answer changes to the donor’s emotion category, and therefore measures content\-specific transfer\. Each reported effect is the paired difference from therand4\\mathrm\{rand\}\_\{4\}arm\. A full\-state replacement checks that downstream generation can respond to a state change atL∗L^\{\*\}\. For the depth analysis, we repeat the same intervention at several layers\. Implementation and statistical details are given in Section[IV\-B](https://arxiv.org/html/2608.06409#S4.SS2)\.
### III\-FControls for Surface Acoustic Confounds
The rank\-three analysis tests whetherSdecodingS\_\{\\mathrm\{decoding\}\}supports held\-out emotion decoding outside the native readout\. One possible explanation is that this performance is driven mainly by simple surface acoustic cues that covary with the emotion labels\. Such cue–label relationships can occur in acted\-emotion corpora; for example, overall recording level can itself support decoding\[[9](https://arxiv.org/html/2608.06409#bib.bib9),[10](https://arxiv.org/html/2608.06409#bib.bib10)\]\. We test this explanation with three controls\. First, we decode emotion from clip\-level acoustic\-prosodic descriptors alone; the resulting accuracyAdescA\_\{\\mathrm\{desc\}\}measures their predictive strength\. Second, we regress theSdecodingS\_\{\\mathrm\{decoding\}\}coordinates on those descriptors and decode from the residuals; the accuracy drop measures how much decoding depends on the measured cues\. Third, we equalize the loudness of every clip, re\-extract the states, and repeat the decoding to test dependence on absolute level\. An extended descriptor panel and a nonlinear removal variant provide stronger versions of the same control; panel composition and protocols are given in Section[IV\-C](https://arxiv.org/html/2608.06409#S4.SS3)\.
## IVExperimental Setup
### IV\-AModels, Corpora, and Evaluation
We evaluate Qwen2\.5\-Omni\-7B, Qwen2\-Audio\-7B, Audio\-Flamingo\-3, Kimi\-Audio\-7B, and Phi\-4\-MM\[[1](https://arxiv.org/html/2608.06409#bib.bib1),[2](https://arxiv.org/html/2608.06409#bib.bib2),[3](https://arxiv.org/html/2608.06409#bib.bib3),[4](https://arxiv.org/html/2608.06409#bib.bib4),[5](https://arxiv.org/html/2608.06409#bib.bib5)\]\. The task is four\-way classification of happy, sad, angry, and neutral\. CREMA\-D\[[11](https://arxiv.org/html/2608.06409#bib.bib11)\]contributes 4,900 clips from 91 speakers, and VESUS\[[38](https://arxiv.org/html/2608.06409#bib.bib38)\]contributes 10,073 clips from 10 speakers, giving 10 model–corpus conditions\.
Four Latin\-square prompt variants rotate the emotions through the four option positions\. Under each model’s tokenizer, every selected option verbalizer is a single native vocabulary token\. Generation is greedy over the full vocabulary without masking non\-option tokens\. A strict prefix parser maps valid answer surfaces to emotion labels; refusals, ambiguous answers, and off\-format responses are incorrect\. All four ladder levels use the same option rows, with the generated answer, option logits, and hidden state recorded at the same first\-answer\-token event\.
AgenA\_\{\\mathrm\{gen\}\}andAoptA\_\{\\mathrm\{opt\}\}require no fitting;AaffA\_\{\\mathrm\{aff\}\}andAstateA\_\{\\mathrm\{state\}\}use five speaker\-disjoint outer folds\. Standardization, penalty selection, and reader fitting are confined to the training speakers in each fold\. The observational unit is the clip\. Speaker\-clustered resampling keeps all clips and prompt variants from one speaker together\. We report paired, per\-condition 95% intervals without family\-wise adjustment\.
### IV\-BSubspace Decoding and Causal Replacement
We split speakers into fixed training and held\-out sets\. Training speakers selectL∗L^\{\*\}by logit\-lens accuracy, constructSdecodingS\_\{\\mathrm\{decoding\}\}, and fit the decoders; held\-out speakers are reserved for evaluation\. The random\-space results averagerand3\\mathrm\{rand\}\_\{3\}over 20 independent draws and report the spread across draws\.
The matched replacements use the speaker split and subspaces defined above\. Within each corpus, every system and intervention arm uses the same receiver–donor pairs, and the random arm uses one fixedrand4\\mathrm\{rand\}\_\{4\}\. Effects are paired differences from the random arm, with uncertainty clustered by speaker\. We also report the one\-sided 95% upper bound for each readout\-external effect as a share of the corresponding readout\-aligned effect\. A full\-state replacement provides a perturbability control atL∗L^\{\*\}\.
### IV\-CControls for Surface Acoustic Confounds
Surface acoustic cues can covary with emotion labels, so we test whether they explain the held\-out decodability ofSdecodingS\_\{\\mathrm\{decoding\}\}\. We use a ten\-descriptor base panel and a twenty\-descriptor extended panel\.AdescA\_\{\\mathrm\{desc\}\}fits the same decoder family to the descriptors alone using the same speaker split\. We then regress theSdecodingS\_\{\\mathrm\{decoding\}\}coordinates on each descriptor panel and decode from the residuals, with all statistics estimated on training speakers only\. The nonlinear variant replaces linear regression with gradient\-boosted trees\. The input\-side control RMS\-equalizes each clip, re\-extracts the answer\-position states under the same prompts andL∗L^\{\*\}, and repeats the subspace analysis\.
The Supplementary Material provides the remaining experimental configuration details, including model checkpoints and prompt templates, tokenizer and endpoint audits, reader fitting and speaker splits, subspace construction and random\-space sampling, receiver–donor pairing, and acoustic descriptor definitions and control protocols\.
## VResults
We organize the results around the generation\-aligned performance ladder\. We first report its endpoint, decision\-rule, and readout\-coverage gaps, then test label\-free logit correction, the availability of readout\-external information, its causal use, and finally controls for measured surface acoustic cues\.
### V\-ABoth Gaps Are Systematic, but Their Relative Importance Varies
Figure 2:Which gap dominates differs by condition\.Each point is one model–corpus condition, with color denoting the system and shape the corpus\. The dashed line marks equal gaps and the shading separates the two regimes: points in the blue region above the line lose more at the readout \(coverage\-dominant\), and points in the tan region below it lose more at the decision rule \(decision\-dominant\)\. Exact values and intervals are in Table[I](https://arxiv.org/html/2608.06409#S5.T1)\.Table[I](https://arxiv.org/html/2608.06409#S5.T1)reports all three performance\-ladder terms across the ten conditions, whose meanAgenA\_\{\\mathrm\{gen\}\}is0\.4740\.474\.Δendpoint\\Delta\_\{\\mathrm\{endpoint\}\}is effectively zero throughout\. The emitted answer achieves the same accuracy as the option favored by the model’s logits, so no meaningful performance is lost at this interface\. Both explanatory gaps are positive\.Δdecision\\Delta\_\{\\mathrm\{decision\}\}ranges from\+0\.0264\+0\.0264to\+0\.2067\+0\.2067, showing that the default decision over the option logits falls short of a fitted rule on those same logits\.Δcoverage\\Delta\_\{\\mathrm\{coverage\}\}ranges from\+0\.0062\+0\.0062to\+0\.3679\+0\.3679, showing that the fitted logit rule in turn falls short of a reader of the full answer\-position state\. The confidence intervals for both gaps exclude zero in every condition\.
Figure[2](https://arxiv.org/html/2608.06409#S5.F2)compares the relative sizes of the two explanatory gaps\. Qwen2\-Audio×\\timesCREMA\-D is decision\-dominant, withΔdecision\\Delta\_\{\\mathrm\{decision\}\}at\+0\.197\+0\.197compared with a\+0\.060\+0\.060readout\-coverage gap\. Phi\-4\-MM×\\timesCREMA\-D is coverage\-dominant; its\+0\.368\+0\.368readout\-coverage gap is the largest in the study, whereasΔdecision\\Delta\_\{\\mathrm\{decision\}\}is\+0\.075\+0\.075\. Qwen2\.5\-Omni×\\timesCREMA\-D has substantial losses at both transitions, including the largestΔdecision\\Delta\_\{\\mathrm\{decision\}\}\(\+0\.207\+0\.207\) and a\+0\.138\+0\.138readout\-coverage gap\. The remaining conditions lie between these patterns, with Kimi\-Audio showing the widest overall separation between generated behavior and state decodability\.
TABLE I:The distance between generated behavior and linear decodability decomposes into an endpoint\-validity term and two explanatory gaps\. Rows are ordered byAgenA\_\{\\mathrm\{gen\}\}\.Δdecision=Aaff−Aopt\\Delta\_\{\\mathrm\{decision\}\}=A\_\{\\mathrm\{aff\}\}\-A\_\{\\mathrm\{opt\}\}andΔcoverage=Astate−Aaff\\Delta\_\{\\mathrm\{coverage\}\}=A\_\{\\mathrm\{state\}\}\-A\_\{\\mathrm\{aff\}\}\.Robustness analyses reproduce the gaps with nonlinear logit\-side decoding, matched regularization budgets, and delete\-one\-speaker resampling; complete results are in Supplementary Section[S4](https://arxiv.org/html/2608.06409#S4a)\.
The ladder reveals no single universal failure profile\. Some conditions have a larger decision\-rule gap, others have a larger readout\-coverage gap, and several show substantial gaps at both transitions\.
### V\-BLabel\-Free Logit Correction Recovers Part of the Decision\-Rule Gap
Figure 3:A label\-free logit correction, measured on the emitted answer\.Gains are grouped by system, with a solid CREMA\-D bar and a hatched VESUS bar\. Bars show only the generated\-accuracy gain over the uncorrected baseline; baseline and corrected accuracies, together with transmission and parseability diagnostics, are in Supplementary Table[S7](https://arxiv.org/html/2608.06409#S5.T7)\.Figure[3](https://arxiv.org/html/2608.06409#S5.F3)shows that label\-free logit correction improves the emitted answer in all ten conditions, with accuracy gains from\+0\.0098\+0\.0098to\+0\.1392\+0\.1392\(per\-condition values in Supplementary Table[S7](https://arxiv.org/html/2608.06409#S5.T7)\)\. These generation\-time gains recover part of the decision\-rule gap without ground\-truth labels\. The largest gain is\+0\.1392\+0\.1392on Qwen2\.5\-Omni×\\timesCREMA\-D\. The correction also helps across performance regimes\. Phi\-4\-MM gains\+0\.0399\+0\.0399and\+0\.0178\+0\.0178with its coverage\-dominant profile, while Audio\-Flamingo\-3 gains\+0\.0098\+0\.0098on CREMA\-D from a 0\.9060 baseline\. Across conditions, the offset realizes 0\.27 to 0\.87 of the supervised decision\-rule gap\. The label\-free correction is therefore effective across all tested model–corpus conditions, although the size of the gain varies\. Because the offsets are estimated from the unlabeled evaluation batch itself, the correction is transductive; applying it to a single isolated example would require other target\-domain data \(Section[VII](https://arxiv.org/html/2608.06409#S7)\)\.
Because the offsets modify logits during unconstrained generation, they could also make the model produce answers outside the required format\. We therefore check whether the corrected option is actually emitted and whether the answer remains parseable\. The corrected option is emitted on 95\.53% to 100% of rows, with parseability unchanged except for a 0\.0019 loss on Qwen2\.5\-Omni×\\timesVESUS\. The remaining difference from the supervised gap indicates that systematic option\-prior bias is one contributor rather than its complete explanation\.
### V\-CReadout\-External Emotion Information Generalizes
Figure 4:Availability does not imply effective use\.\(a\) Held\-out decoding accuracy atL∗L^\{\*\}for the native readout spaceV3V\_\{3\}, a label\-free random subspacerand3⊂V4⟂⊂V3⟂\\mathrm\{rand\}\_\{3\}\\subset V\_\{4\}^\{\\perp\}\\subset V\_\{3\}^\{\\perp\}averaged over 20 draws, and the supervised spaceSdecodingS\_\{\\mathrm\{decoding\}\}selected in the same complement; the dashed line marks chance\. Panels \(b\) and \(c\) score the rank\-four minimal\-pair replacements atL∗L^\{\*\}, shown as paired differences from therand4\\mathrm\{rand\}\_\{4\}arm forV4V\_\{4\}\(circles, dark blue\) andSinterventionS\_\{\\mathrm\{intervention\}\}\(squares, light blue\)\. Panel \(b\) counts any answer change, whereas panel \(c\) counts only changes to the donor emotion\. Bars are speaker\-clustered 95% intervals, and the horizontal axis is symmetric\-logarithmic\. All panels share their row order\. Exact values for \(a\), \(b\), and complete donor\-following values for \(c\) are in Supplementary Tables[S8](https://arxiv.org/html/2608.06409#S6.T8),[S11](https://arxiv.org/html/2608.06409#S7.T11), and[S13](https://arxiv.org/html/2608.06409#S7.T13)\.Having tested the decision\-rule gap at the output, we turn to the readout\-coverage gap and ask whether generalizable emotion information remains decodable outside the native option readout\. Fig\.[4](https://arxiv.org/html/2608.06409#S5.F4)\(a\) shows the held\-out decoding results for the native, random, and selected readout\-external subspaces, with exact values in Supplementary Table[S8](https://arxiv.org/html/2608.06409#S6.T8)\. At the training\-selectedL∗L^\{\*\}, held\-out decoding fromSdecoding⊂V4⟂⊂V3⟂S\_\{\\mathrm\{decoding\}\}\\subset V\_\{4\}^\{\\perp\}\\subset V\_\{3\}^\{\\perp\}reaches 0\.481 to 0\.949, and its improvement overrand3\\mathrm\{rand\}\_\{3\}has a speaker\-clustered interval excluding zero in every condition\. The answer\-position state therefore retains linearly accessible emotion information outside the native option contrasts\. Phi\-4\-MM illustrates the distinction\.SdecodingS\_\{\\mathrm\{decoding\}\}reaches 0\.668 on CREMA\-D and 0\.481 on VESUS, whereasV3V\_\{3\}reaches only 0\.342 and 0\.259\.
The selection\-matchedV3−rand3V\_\{3\}\-\\mathrm\{rand\}\_\{3\}contrast measures how informative the native readout is relative to a same\-rank random space\. It is positive with an interval excluding zero in nine of the ten conditions, ranging from\+0\.037\+0\.037to\+0\.272\+0\.272\. The exception is Phi\-4\-MM×\\timesVESUS at−0\.018\-0\.018\[−0\.033,\+0\.001\-0\.033,\+0\.001\], and its CREMA\-D condition is only\+0\.022\+0\.022\[\+0\.004,\+0\.040\+0\.004,\+0\.040\], so Phi\-4\-MM’s native option contrasts carry little more emotion information than a random subspace of the same rank\. The next smallest contrast is Kimi\-Audio×\\timesVESUS at\+0\.037\+0\.037\[\+0\.026,\+0\.047\+0\.026,\+0\.047\], which shows that a weakly informative native readout is not confined to one language\-model family\.
Raw accuracy contrasts compress near ceiling\. Audio\-Flamingo\-3×\\timesCREMA\-D’s\+0\.068\+0\.068difference, for example, sits on a 0\.869 random\-space baseline\. Individualrand3\\mathrm\{rand\}\_\{3\}draws vary by up to 0\.05; per\-draw results and theSdecoding−rand3S\_\{\\mathrm\{decoding\}\}\-\\mathrm\{rand\}\_\{3\}intervals are in Supplementary Section[S6](https://arxiv.org/html/2608.06409#S6a)\.
### V\-DMinimal\-Pair Interventions Reveal Limited Causal Use
Held\-out decodability establishes availability; the matched interventions now measure use\. Fig\.[4](https://arxiv.org/html/2608.06409#S5.F4)\(b,c\) shows both outcomes\. Replacing the readout\-alignedV4V\_\{4\}component increases answer change over the random arm in all ten conditions, by\+0\.0181\+0\.0181to\+0\.3631\+0\.3631\. It also significantly increases donor following in eight conditions; the two Phi\-4\-MM conditions are nonsignificant\. These results show that the model’s answer is causally sensitive to changes in the native readout space\.
By contrast, replacing the readout\-externalSinterventionS\_\{\\mathrm\{intervention\}\}component has much less influence on the answer\. It produces a significant answer\-change effect in five conditions, but no effect exceeds\+0\.0206\+0\.0206\. Only one condition shows a significant increase in donor following\. These results indicate that the selected readout\-external information has only limited influence on the emitted answer\.
We perform two additional tests to rule out the possibility that the weakSinterventionS\_\{\\mathrm\{intervention\}\}effects arise only from small edits or an unresponsive downstream pathway\. Exact values, per\-conditionL∗L^\{\*\}, and remaining diagnostics are in Supplementary Table[S11](https://arxiv.org/html/2608.06409#S7.T11)and Section[S7](https://arxiv.org/html/2608.06409#S7a)\.
Figure 5:Depth profile of causal access in Qwen2\-Audio×\\timesCREMA\-D\.Donor\-following effects of matched rank\-four minimal\-pair replacement at eight depths, with speaker\-clustered 95% intervals; hollow markers mark intervals containing zero, and the vertical axis is symmetric\-logarithmic\. Exact values are in Supplementary Section[S7\-E](https://arxiv.org/html/2608.06409#S7.SS5)\.Qwen2\-Audio×\\timesCREMA\-D is the only condition with a significant readout\-external donor\-following effect atL∗L^\{\*\}\(\+0\.0106\+0\.0106\[\+0\.0058,\+0\.0154\+0\.0058,\+0\.0154\]\)\. We therefore repeat the intervention at eight layers to ask where this causal influence is strongest; Fig\.[5](https://arxiv.org/html/2608.06409#S5.F5)shows the resulting donor\-following effects\. When the replacement is applied at layer 16, the external effect is near zero\. It peaks at layer 20, where it matches the readout\-aligned effect at the same layer, and becomes weaker when the replacement is applied closer to the output\. By contrast, the readout\-aligned effect grows toward the output and reaches\+0\.4983\+0\.4983\. In this condition, readout\-external information has its strongest causal influence in the middle of the model, whereas readout\-aligned information becomes increasingly influential near the final readout\.
We further track the particular component injected at layer 20 to determine whether it reaches the final answer\. Relative to the random control, this intervention shifts the option logits toward the donor emotion by\+0\.190\+0\.190\[\+0\.176,\+0\.204\+0\.176,\+0\.204\]\. However, the shift is usually too small to change which option has the highest score\. The mid\-stack external pathway therefore reaches the final logits but rarely changes the emitted answer\. This depth pattern is established only for Qwen2\-Audio×\\timesCREMA\-D \(protocols and exact values in Supplementary Sections[S7\-E](https://arxiv.org/html/2608.06409#S7.SS5)and[S7\-F](https://arxiv.org/html/2608.06409#S7.SS6)\)\.
These interventions show that readout\-external emotion information has limited causal access at the answer\.
### V\-EReadout\-External Decodability Persists under Controls for Measured Surface Cues
Figure 6:Readout\-external decodability under progressively stronger acoustic controls\.Each line is one model–corpus condition, with color denoting the system, solid circles CREMA\-D, and dashed squares VESUS\. Shown is held\-outSdecodingS\_\{\\mathrm\{decoding\}\}accuracy for the raw states, after input loudness equalization, and after residualizing ten or twenty acoustic\-prosodic descriptors linearly or with gradient\-boosted trees\. Dotted lines mark the descriptor\-only decoder for each corpus and the dashed gray line marks chance\. Exact values and intervals are in Supplementary Tables[S17](https://arxiv.org/html/2608.06409#S8.T17),[S19](https://arxiv.org/html/2608.06409#S8.T19), and[S20](https://arxiv.org/html/2608.06409#S8.T20)\.The causal interventions above test whether readout\-external information affects the answer\. Here we return to the rank\-threeSdecodingS\_\{\\mathrm\{decoding\}\}used in Fig\.[4](https://arxiv.org/html/2608.06409#S5.F4)\(a\) and ask whether its held\-out decodability can be explained by measured surface cues\. Recording level alone separates several class pairs in these corpora \(Supplementary Table[S16](https://arxiv.org/html/2608.06409#S8.T16)\), making shallow clip statistics a plausible source of readout\-external decodability\. Figure[6](https://arxiv.org/html/2608.06409#S5.F6)applies the controls defined in Section[IV\-C](https://arxiv.org/html/2608.06409#S4.SS3)\.
Linear residualization of the measured cues reduces held\-outSdecodingS\_\{\\mathrm\{decoding\}\}accuracy by 0\.039 to 0\.255\. After nonlinear residualization of the extended descriptor panel, accuracy remains at 0\.347 to 0\.752, above the 0\.25 chance level in every condition\. Input\-side loudness equalization changes accuracy by at most 0\.022\. Exact values are in Supplementary Tables[S17](https://arxiv.org/html/2608.06409#S8.T17),[S19](https://arxiv.org/html/2608.06409#S8.T19), and[S20](https://arxiv.org/html/2608.06409#S8.T20)\. Thus, the measured surface acoustic cues and absolute recording level do not fully explain the held\-out decodability ofSdecodingS\_\{\\mathrm\{decoding\}\}\.
## VIDiscussion
The central result is not simply that a hidden\-state decoder outperforms the generated answer\. The diagnostic ladder identifies two downstream gaps between information available at the answer position and the answer the model emits\. One arises when the model converts its option logits into a choice; the other arises because the option contrasts expose only part of the information available in the hidden state\. Both gaps are positive in every evaluated condition, so a system can suffer from both problems at once\. They are aggregate properties of a model–corpus condition rather than mutually exclusive explanations of individual errors\. More broadly, low generated accuracy need not mean that the relevant emotion evidence never reached the language model\.
### VI\-AWhy the Decision Rule Loses Available Evidence
Option logits combine evidence from the audio with option\-token priors, positional preferences, and prompt\-conditioned response habits\. A stable preference for one option can therefore shift the default argmax even when the relative scores still contain useful emotion evidence\[[18](https://arxiv.org/html/2608.06409#bib.bib18)\]\. This interpretation is consistent with the label\-free correction\. Estimating and removing marginal option preferences improves generation in all ten conditions and recovers 0\.27 to 0\.87 of the supervised decision\-rule gap\. The three largest decision\-rule gaps also produce the three largest correction gains\. Thus, the gap is not only diagnostic; it indicates when a logit\-side repair is likely to help\.
The correction recovers only part of the gap because a fixed offset captures only the stable component of the mismatch\. The remaining difference from the supervised affine reader may reflect class\-dependent scaling or example\-dependent boundaries that marginal option frequencies cannot estimate\. This account predicts that reducing stable option preferences, through prompts or training procedures, should reduce both the decision\-rule gap and the benefit of the offset correction\.
### VI\-BWhy Decodable Information Has Limited Influence
The readout\-coverage results reveal a different mismatch\. Emotion remains decodable from selected directions outside the native option readout and generalizes to held\-out speakers, yet replacing those directions usually has little effect on the answer \(no answer\-change effect exceeds\+0\.0206\+0\.0206; Section[V\-D](https://arxiv.org/html/2608.06409#S5.SS4)\)\. A plausible explanation is objective mismatch\. Audio front ends and intermediate language\-model states may preserve rich prosodic information, while next\-token and audio\-instruction training reward only the information needed to produce the target text\. They do not directly require all emotion\-discriminative directions to align with the few output directions that separate the prompted option tokens\. A supervised decoder is explicitly trained to find such directions; the native readout is not\.
On this account, readout\-external information can persist as a usable representation without being part of the model’s normal answer pathway\. This explains why decodability and behavioral influence diverge, a distinction that probing studies must preserve\[[28](https://arxiv.org/html/2608.06409#bib.bib28),[39](https://arxiv.org/html/2608.06409#bib.bib39)\]\. It is especially important here becauseSdecodingS\_\{\\mathrm\{decoding\}\}andSinterventionS\_\{\\mathrm\{intervention\}\}are selected with emotion labels\. Successful decoding shows that the information is available to a supervised linear reader, not that the model naturally uses the same directions\[[34](https://arxiv.org/html/2608.06409#bib.bib34)\]\. The acoustic controls further show that the measured surface cues do not fully explain this decodability, although unmeasured acoustic properties may still contribute\.
The depth profile offers a more specific hypothesis about routing\. In Qwen2\-Audio×\\timesCREMA\-D, readout\-external replacement has its strongest donor\-directed effect in the middle of the stack and becomes weaker near the output, while the readout\-aligned effect grows\. This pattern is consistent with progressive consolidation, in which intermediate layers can use emotion information in several directions while later layers increasingly concentrate behaviorally relevant information into token\-aligned coordinates\. Information that is not transferred into those coordinates may be overwritten or lose access to the answer\. Because this pattern is established in one condition, it is best treated as a mechanism to test across models rather than a universal depth profile\.
The weak low\-rank replacement effects do not imply that every direction outside the readout is functionless\. They show that the selected linear component has limited causal access under the tested intervention\. Emotion information could also be distributed across more directions or participate through nonlinear interactions that a rank\-four replacement does not capture\. The strong response to readout\-aligned and full\-state replacements nevertheless shows that the downstream pathway can respond to state changes at the intervention layer; the main limitation lies in how the selected external information is routed\.
### VI\-CImplications for Repair and Evaluation
The two gaps suggest different repairs\. A decision\-rule gap can be addressed by calibrating the existing option logits, without changing the hidden representation\. A readout\-coverage gap instead requires changing how hidden\-state information reaches the answer, for example through a learned readout adapter, targeted fine\-tuning of the output mapping, or auxiliary supervision that aligns prosodic evidence with the option contrasts\. If such training reducesΔcoverage\\Delta\_\{\\mathrm\{coverage\}\}and increases the causal effect of readout\-external directions, it would support the routing explanation above\.
These results also change how generative paralinguistic systems should be evaluated\. Generated answers, option logits, and hidden states should be measured at the same answer event; diagnostic readers should be evaluated on identity\-disjoint splits; and claims about information use should include dimension\-matched causal controls\. Reporting these views alongside accuracy separates a failure to represent emotion from a failure to expose or use information that is already present at the answer position\.
## VIILimitations
#### Empirical and statistical coverage
We evaluate only five models\. Although they use several audio front ends, they cover only two language\-model families, and four use Qwen\-family backbones\. Our experiments focus on four\-way English emotion classification on CREMA\-D and VESUS, using single\-token multiple\-choice answers\. Results may differ for other architectures, spontaneous speech, other languages, other paralinguistic tasks, or free\-form answers\. VESUS contains only 10 speakers, which limits the precision of speaker\-clustered estimates\.
#### Diagnostic scope
AaffA\_\{\\mathrm\{aff\}\}reads only three option\-logit contrasts, whereasAstateA\_\{\\mathrm\{state\}\}reads the fulldd\-dimensional hidden state\. Although both use linear classifiers, the state reader has access to many more input dimensions\. A largerΔcoverage\\Delta\_\{\\mathrm\{coverage\}\}can therefore arise for two reasons: the hidden state may contain information that the native readout does not expose, or the full\-state reader may benefit from its larger input space\. We therefore interpretΔcoverage\\Delta\_\{\\mathrm\{coverage\}\}as a performance gap rather than a pure measure of readout geometry\.
#### Acoustic interpretation
Our acoustic controls account for recording level and twenty measured acoustic descriptors, but they do not cover every property of the speech signal\. The remaining decodable information may therefore still include acoustic cues that we did not measure\. It should not, by itself, be interpreted as an abstract representation of emotion\.
#### Correction and intervention scope
The logit correction does not use emotion labels, but it estimates its offsets from the full unlabeled evaluation batch\. It therefore requires a batch of examples from the target domain and must be recalibrated for a new domain\. It cannot be applied to one isolated example without other target\-domain data\. The matched replacements are diagnostic tests rather than a trained repair\. They ask whether replacing a selected part of the hidden state can change the model’s answer; they do not teach the model a new readout or routing mechanism\. The one significant readout\-external donor\-following effect therefore shows only a small amount of causal influence\. It does not close the readout\-coverage gap\.
## VIIIConclusion
Behavioral errors alone do not reveal where task information stops influencing a generated answer\. By aligning behavior, option logits, and hidden states at the same first\-token event, our diagnostic ladder separates endpoint validity, the decision\-rule gap, and the readout\-coverage gap\. Both explanatory gaps are positive in all ten evaluated conditions\. A label\-free logit correction recovers part of the decision\-rule gap during generation\. Meanwhile, emotion information remains decodable outside the native readout under controls for measured acoustic cues, but replacing the selected external component seldom moves the answer toward the donor emotion; this bounds the causal influence of the directions we selected, not of all readout\-external information\. These results distinguish information availability from behavioral use and motivate different responses: correcting the rule over existing option logits or adapting how hidden\-state information is read out and routed\. The same three views can be captured at any prompted answer token; in speech, they show where paralinguistic information available at the answer position stops contributing to the emitted response\. Reported alongside accuracy, they turn a benchmark score into a diagnosis of the state\-to\-answer interface\.
## References
- \[1\]J\. Xu, Z\. Guo, J\. He, H\. Hu, T\. He, S\. Bai, K\. Chen, J\. Wang, Y\. Fan, K\. Dang, B\. Zhang, X\. Wang, Y\. Chu, and J\. Lin, “Qwen2\.5\-omni technical report,” 2025\. \[Online\]\. Available:https://arxiv\.org/abs/2503\.20215
- \[2\]Y\. Chu, J\. Xu, Q\. Yang*et al\.*, “Qwen2\-audio technical report,” arXiv preprint arXiv:2407\.10759, 2024\.
- \[3\]A\. Goel, S\. Ghosh, J\. Kim, S\. Kumar, Z\. Kong, S\.\-g\. Lee, C\.\-H\. H\. Yang, R\. Duraiswami, D\. Manocha, R\. Valle, and B\. Catanzaro, “Audio flamingo 3: Advancing audio intelligence with fully open large audio language models,”*arXiv preprint arXiv:2507\.08128*, 2025\.
- \[4\]Kimi Team, “Kimi\-audio technical report,” 2025\.
- \[5\]M\. Abdin, J\. Aneja, H\. Behl, S\. Bubeck, R\. Eldan, S\. Gunasekar, M\. Harrison, R\. J\. Hewett, M\. Javaheripi, P\. Kauffmann*et al\.*, “Phi\-4 technical report,”*arXiv preprint arXiv:2412\.08905*, 2024\.
- \[6\]C\.\-y\. Huang, K\.\-H\. Lu, S\.\-H\. Wang, C\.\-Y\. Hsiao, C\.\-Y\. Kuan, H\. Wu, S\. Arora, K\.\-W\. Chang, J\. Shi, Y\. Peng*et al\.*, “Dynamic\-superb: Towards a dynamic, collaborative, and comprehensive instruction\-tuning benchmark for speech,” in*ICASSP 2024\-2024 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\)*\. IEEE, 2024, pp\. 12 136–12 140\.
- \[7\]Q\. Yang, J\. Xu, W\. Liu, Y\. Chu, Z\. Jiang, X\. Zhou, Y\. Leng, Y\. Lv, Z\. Zhao, C\. Zhou*et al\.*, “Air\-bench: Benchmarking large audio\-language models via generative comprehension,” in*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, 2024, pp\. 1979–1998\.
- \[8\]J\. Ao, Y\. Wang, X\. Tian, D\. Chen, J\. Zhang, L\. Lu, Y\. Wang, H\. Li, and Z\. Wu, “Sd\-eval: A benchmark dataset for spoken dialogue understanding beyond words,”*Advances in Neural Information Processing Systems*, vol\. 37, pp\. 56 898–56 918, 2024\.
- \[9\]K\. R\. Scherer, “Vocal communication of emotion: A review of research paradigms,”*Speech Communication*, vol\. 40, no\. 1–2, pp\. 227–256, 2003\.
- \[10\]P\. N\. Juslin and P\. Laukka, “Communication of emotions in vocal expression and music performance: Different channels, same code?”*Psychological bulletin*, vol\. 129, no\. 5, p\. 770, 2003\.
- \[11\]H\. Cao, D\. G\. Cooper, M\. K\. Keutmann, R\. C\. Gur, A\. Nenkova, and R\. Verma, “Crema\-d: Crowd\-sourced emotional multimodal actors dataset,”*IEEE transactions on affective computing*, vol\. 5, no\. 4, pp\. 377–390, 2014\.
- \[12\]T\. Z\. Zhao, E\. Wallace, S\. Feng, D\. Klein, and S\. Singh, “Calibrate before use: Improving few\-shot performance of language models,” in*Proceedings of the 38th International Conference on Machine Learning*, ser\. Proceedings of Machine Learning Research, vol\. 139\. PMLR, 2021, pp\. 12 697–12 706\.
- \[13\]H\. Zhou, X\. Wan, L\. Proleev, D\. Mincu, J\. Chen, K\. Heller, and S\. Roy, “Batch calibration: Rethinking calibration for in\-context learning and prompt engineering,” in*International Conference on Learning Representations*, 2024, arXiv:2309\.17249\.
- \[14\]D\. Zhang, S\. Li, X\. Zhang, J\. Zhan, P\. Wang, Y\. Zhou, and X\. Qiu, “Speechgpt: Empowering large language models with intrinsic cross\-modal conversational abilities,” in*Findings of the Association for Computational Linguistics: EMNLP 2023*, 2023, pp\. 15 757–15 773\.
- \[15\]C\. Tang, W\. Yu, G\. Sun, X\. Chen, T\. Tan, W\. Li, L\. Lu, Z\. Ma, and C\. Zhang, “Salmonn: Towards generic hearing abilities for large language models,” in*International Conference on Learning Representations*, vol\. 2024, 2024, pp\. 16 607–16 629\.
- \[16\]Z\. Kong, A\. Goel, R\. Badlani, W\. Ping, R\. Valle, and B\. Catanzaro, “Audio flamingo: A novel audio language model with few\-shot learning and dialogue abilities,”*arXiv preprint arXiv:2402\.01831*, 2024\.
- \[17\]J\. Chen, Z\. Guo, J\. Chun, P\. Wang, A\. Perrault, and M\. Elsner, “Do audio LLMs really LISTEN, or just transcribe? measuring lexical vs\. acoustic emotion cues reliance,” in*Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics*, 2026, pp\. 5848–5877\.
- \[18\]C\. Zheng, H\. Zhou, F\. Meng, J\. Zhou, and M\. Huang, “Large language models are not robust multiple choice selectors,” in*International Conference on Learning Representations*, 2024, spotlight; arXiv:2309\.03882\.
- \[19\]S\. Kumar, “Answer\-level calibration for free\-form multiple choice question answering,” in*Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*\. Dublin, Ireland: Association for Computational Linguistics, 2022, pp\. 665–679\.
- \[20\]W\.\-N\. Hsu, B\. Bolte, Y\.\-H\. H\. Tsai, K\. Lakhotia, R\. Salakhutdinov, and A\. Mohamed, “Hubert: Self\-supervised speech representation learning by masked prediction of hidden units,”*IEEE/ACM transactions on audio, speech, and language processing*, vol\. 29, pp\. 3451–3460, 2021\.
- \[21\]S\. Chen, C\. Wang, Z\. Chen, Y\. Wu, S\. Liu, Z\. Chen, J\. Li, N\. Kanda, T\. Yoshioka, X\. Xiao*et al\.*, “Wavlm: Large\-scale self\-supervised pre\-training for full stack speech processing,”*IEEE Journal of Selected Topics in Signal Processing*, vol\. 16, no\. 6, pp\. 1505–1518, 2022\.
- \[22\]A\. Radford, J\. W\. Kim, T\. Xu, G\. Brockman, C\. McLeavey, and I\. Sutskever, “Robust speech recognition via large\-scale weak supervision,” in*International conference on machine learning*\. PMLR, 2023, pp\. 28 492–28 518\.
- \[23\]A\. Pasad, J\.\-C\. Chou, and K\. Livescu, “Layer\-wise analysis of a self\-supervised speech representation model,” in*2021 IEEE Automatic Speech Recognition and Understanding Workshop \(ASRU\)*\. IEEE, 2021, pp\. 914–921\.
- \[24\]M\. De Seyssel, M\. Lavechin, Y\. Adi, E\. Dupoux, and G\. Wisniewski, “Probing phoneme, language and speaker information in unsupervised speech representations,” in*Interspeech 2022*, 2022, pp\. 1402–1406\.
- \[25\]J\. Wagner, A\. Triantafyllopoulos, H\. Wierstorf, M\. Schmitt, F\. Burkhardt, F\. Eyben, and B\. W\. Schuller, “Dawn of the transformer era in speech emotion recognition: Closing the valence gap,”*IEEE Transactions on Pattern Analysis and Machine Intelligence*, 2023\.
- \[26\]Z\. Ma, Z\. Zheng, J\. Ye, J\. Li, Z\. Gao, S\. Zhang, and X\. Chen, “emotion2vec: Self\-supervised pre\-training for speech emotion representation,” in*Findings of the Association for Computational Linguistics: ACL 2024*, 2024, pp\. 15 747–15 760\.
- \[27\]C\.\-K\. Yang, N\. Ho, Y\.\-J\. Lee, and H\.\-y\. Lee, “AudioLens: A closer look at auditory attribute perception of large audio\-language models,”*arXiv preprint arXiv:2506\.05140*, 2025\.
- \[28\]Y\. Belinkov, “Probing classifiers: Promises, shortcomings, and advances,”*Computational Linguistics*, vol\. 48, no\. 1, pp\. 207–219, 2022\.
- \[29\]nostalgebraist, “Interpreting GPT: The logit lens,” 2020, lessWrong post;https://www\.lesswrong\.com/posts/AcKRB8wDpdaN6v6ru/interpreting\-gpt\-the\-logit\-lens\.
- \[30\]N\. Belrose, I\. Ostrovsky, L\. McKinney, Z\. Furman, L\. Smith, D\. Halawi, S\. Biderman, and J\. Steinhardt, “Eliciting latent predictions from transformers with the tuned lens,”*arXiv preprint arXiv:2303\.08112*, 2023\.
- \[31\]J\. Vig, S\. Gehrmann, Y\. Belinkov, S\. Qian, D\. Nevo, Y\. Singer, and S\. Shieber, “Investigating gender bias in language models using causal mediation analysis,”*Advances in neural information processing systems*, vol\. 33, pp\. 12 388–12 401, 2020\.
- \[32\]A\. Geiger, H\. Lu, T\. Icard, and C\. Potts, “Causal abstractions of neural networks,” in*Advances in Neural Information Processing Systems*, vol\. 34, 2021, pp\. 9574–9586\.
- \[33\]K\. Meng, D\. Bau, A\. Andonian, and Y\. Belinkov, “Locating and editing factual associations in gpt,”*Advances in neural information processing systems*, vol\. 35, pp\. 17 359–17 372, 2022\.
- \[34\]A\. Makelov, G\. Lange, A\. Geiger, and N\. Nanda, “Is this the subspace you are looking for? an interpretability illusion for subspace activation patching,” arXiv preprint arXiv:2311\.17030, 2023\.
- \[35\]Y\. Park, H\. Pyun, and Y\. Jo, “Bridging the knowledge\-prediction gap in LLMs on multiple\-choice questions,” in*Proc\. International Conference on Machine Learning*, 2026, arXiv:2509\.23782\.
- \[36\]T\. F\. Chowdhury, D\. H\. Ta, S\. Pan, J\. Stoddard, and Z\. Liao, “AR&D: A framework for retrieving and describing concepts for interpreting AudioLLMs,” in*Proc\. IEEE International Conference on Acoustics, Speech and Signal Processing*, 2026, arXiv:2602\.22253\.
- \[37\]H\. Cho, S\. Yoo, J\. Jang, C\. Kim, and J\. S\. Chung, “Who wins the conflict? mechanistic interpretability of text bias in audio LLMs,” 2026\.
- \[38\]J\. Sager, R\. Shankar, J\. Reinhold, and A\. Venkataraman, “VESUS: A crowd\-annotated database to study emotion production and perception in spoken English,” in*Interspeech*, 2019\.
- \[39\]Y\. Xu, S\. Zhao, J\. Song, R\. Stewart, and S\. Ermon, “A theory of usable information under computational constraints,” in*International Conference on Learning Representations \(ICLR\)*, 2020\.
## Supplementary Material
## S1Experimental Setup Details
### S1\-ASystems
Table[S1](https://arxiv.org/html/2608.06409#S1.T1)lists the five systems evaluated in the paper, with the checkpoint, architecture, and parameter count of each\. For compactness, subsequent supplementary tables abbreviate them as Omni, Q2A, AF3, Kimi, and Phi\-4\.
We selectively evaluate Qwen2\.5\-Omni\[[1](https://arxiv.org/html/2608.06409#bib.bib1)\], Phi\-4\-MM\[[5](https://arxiv.org/html/2608.06409#bib.bib5)\], Audio\-Flamingo\-3\[[3](https://arxiv.org/html/2608.06409#bib.bib3)\], Qwen2\-Audio\[[2](https://arxiv.org/html/2608.06409#bib.bib2)\], and Kimi\-Audio\[[4](https://arxiv.org/html/2608.06409#bib.bib4)\]\. These systems are frequently represented in related work and span several widely used speech\-language\-model architectures\. All five expose hidden states along a speech\-understanding, text\-response pathway and provide the discrete answer endpoint required by our analysis\.
TABLE S1:The five evaluated systems\. Parameter counts are taken from the public model cards and are total architecture sizes, including text/decoder branches where applicable\.
### S1\-BCorpus Composition
The clip counts in Table[S2](https://arxiv.org/html/2608.06409#S1.T2)are filtered subsets of the published corpora, not the full releases\.
- •CREMA\-D\.We use the four\-emotion subset\{\\\{angry, happy, sad, neutral\}\\\}of the published six\-emotion corpus, dropping*disgust*and*fear*, and keep every clip of the retained classes\. The source corpus contains fewer neutral clips than non\-neutral clips, so the resulting 4,900\-clip subset is not class\-balanced: its majority class accounts for 25\.94% of clips, against the 25% four\-way chance rate\.
- •VESUS\.We use the four\-emotion subset\{\\\{happy, sad, angry, neutral\}\\\}of the five\-emotion corpus, dropping*fearful*\. VESUS reads a phonetically balanced, semantically neutral script of more than 250 short phrases, each spoken by 10 actors in every emotion, giving 10,073 clips whose majority class accounts for 25\.01%\.
Because neither subset is exactly balanced, we report the empirical majority\-class rates rather than treating 25% as an exact baseline\.
TABLE S2:Per\-class clip counts of the two evaluation corpora\.
### S1\-CPrompt Protocol
#### System prompts\.
All five systems are queried with their official chat templates, unmodified\. Qwen2\.5\-Omni receives the canonical system message distributed with the model: “You are Qwen, a virtual human developed by the Qwen Team, Alibaba Group, capable of perceiving auditory and visual inputs, as well as generating text and speech\.” Phi\-4\-MM’s speech\-understanding template \(<\|user\|\><\|audio\_1\|\>…<\|end\|\><\|assistant\|\>\) does not include a system turn by design\. Audio\-Flamingo\-3, Qwen2\-Audio, and Kimi\-Audio are each queried with a single user turn carrying the audio and the text prompt; for these we do not add a system message, following each model’s recommended speech\-understanding format\.
#### Prompt bank\.
Every clip is presented under the four prompt variants of Table[S3](https://arxiv.org/html/2608.06409#S1.T3)\. They share one instruction template and differ only in the Latin\-square assignment of emotions to option letters, which places each emotion at each letter position exactly once across the four variants and so controls for positional bias\. The four variants of a clip are repeated measurements of the same audio and are kept together in the same split and the same bootstrap cluster, as detailed in Section[S1\-D](https://arxiv.org/html/2608.06409#S1.SS4)\.
TABLE S3:The four\-variant emotion prompt bank\. The Latin\-square design guarantees that each emotion appears at each letter position exactly once across B1–B4\.
#### Output parsing\.
Generation is greedy over the full vocabulary, and the processor never masks non\-option tokens\. The parser locates a standalone option marker in leading position, such asA,\(A\), orA\., and the corresponding patterns for B, C, and D\. Markers are bounded by non\-letter characters so that letters inside words such as “Answer” are not matched, and the selected option is mapped to the emotion label assigned under the active Latin\-square variant\. Any generation without a valid leading option marker, including refusals and free\-form prose, is counted as incorrect in every main\-text analysis; it is never re\-parsed into a class by scanning the rest of the response\.
#### Layer\-selection readout\.
The criterion that fixesL∗L^\{\*\}\(Section[IV\-B](https://arxiv.org/html/2608.06409#S4.SS2)\) uses a separate readout from the option letters that the analyses score\. At each layer, the model’s own final normalization and unembedding are applied to the answer\-position state, and the score of a class is the log\-sum\-exp over that class’s emotion\-word tokens:\{\\\{happy, Happy, joyful\}\\\}for happy,\{\\\{sad, Sad, upset\}\\\}for sad,\{\\\{angry, Angry\}\\\}for angry, and the corresponding set for neutral\. Each listed verbalizer is a single native vocabulary token under the corresponding model tokenizer\. The highest\-scoring class is the layerwise prediction, andL∗L^\{\*\}is the layer with the highest accuracy on the training split\. These emotion words are never scored as answers; they enter only the layer\-selection criterion\.
### S1\-DReaders, Splits, and Uncertainty
All ladder levels are evaluated on identical rows\. The affine and state readers use five speaker\-disjoint outer folds, with feature standardization and every supervised fitting decision confined to the training side\. The affine reader selects its penalty by nested cross\-validation, whereas the main state reader uses a fixed penalty\. Thus the tuned reader is the one subtracted inΔcoverage\\Delta\_\{\\mathrm\{coverage\}\}; Section[S4](https://arxiv.org/html/2608.06409#S4a)reports the refit in which both readers receive the same nested selection budget\.
The audio clip is the observational unit\. A resampled speaker carries all of that speaker’s clips and all four prompt variants, preserving their dependence\. Intervals are paired and reported per condition without family\-wise adjustment\. This construction yields 91 speaker clusters on CREMA\-D and 10 on VESUS; the latter necessarily produces less precise condition\-level intervals\.
## S2Endpoint Audit
Table[S4](https://arxiv.org/html/2608.06409#S2.T4)reports three implementation checks for the answer endpoint\.
#### Answer surfaces\.
For each system we enumerate candidate surfaces with its own tokenizer and select the one whose tokens carry the full\-vocabulary top\-1 mass\. Every selected option verbalizer is a single native vocabulary token, so its logit is obtained from one output\-head row rather than an aggregation across subtokens\. The choice of surface is not cosmetic\. Audio\-Flamingo\-3 puts all of its mass on parenthesis\-merged tokens and none on bare letters, so an analysis that assumed bare letters would have scored a token that system never emits\. Phi\-4 places 0\.9998 and 0\.9992 of its top\-1 mass on bare letters, with the small remainder on the parenthesis variant\. Spaced surfaces receive no mass in any condition\.
#### Off\-format generations\.
The strict parser accepts only a leading option marker\. Under the uncorrected generation protocol, all generated answers satisfy this requirement\.
#### State\-to\-logit reconstruction and conditioning\.
Passing each saved post\-normalization state back through the native language\-model head reproduces the stored option\-logit contrasts with a mean absolute error one to two orders of magnitude below the tolerance derived from each condition’s logit scale\. The option\-contrast matrix is well conditioned everywhere, with a condition number between 2\.51 and 2\.94 for four systems and 8\.35 for Phi\-4\.
TABLE S4:Endpoint audit\.*Surface*is the selected answer surface and its full\-vocabulary top\-1 in\-option rate; the alternative surface is given where the sweep was stored\.*Off\-fmt*is the fraction of generations the strict parser rejected\.*MAE/tol*is the state\-to\-logit reconstruction error against its tolerance, andκ\\kappais the condition number of the option\-contrast matrix\.For Audio\-Flamingo\-3, the parenthesized form is the valid answer surface; the bare and spaced forms are not selected\. The surface sweep was not stored for the four Qwen conditions, which were run on an earlier pass of the pipeline, but all of their generated answers pass the strict parser and use one of the requested option forms\.
## S3The Performance Ladder in Full
Table[S5](https://arxiv.org/html/2608.06409#S3.T5)gives every rung and both explanatory gaps with speaker\-clustered intervals for all 10 conditions\.AMLPA\_\{\\mathrm\{MLP\}\}is the nonlinear control of Section[S4](https://arxiv.org/html/2608.06409#S4a): a multilayer perceptron on the same three option contrasts\.
TABLE S5:The full ladder, ordered byAgenA\_\{\\mathrm\{gen\}\}\. Intervals are paired speaker\-clustered bootstraps and are not family\-wise adjusted\.
## S4Robustness of the Readout\-Coverage Gap
Three controls target three different alternatives to the coverage gap\. Table[S6](https://arxiv.org/html/2608.06409#S4.T6)reports the second and third; the first is theAMLPA\_\{\\mathrm\{MLP\}\}column of Table[S5](https://arxiv.org/html/2608.06409#S3.T5)\.
#### The affine function class is not the bottleneck\.
Replacing the affine reader on the three option contrasts with a multilayer perceptron on the same contrasts does not absorb the state advantage\.Astate−AMLPA\_\{\\mathrm\{state\}\}\-A\_\{\\mathrm\{MLP\}\}ranges from\+0\.0053\+0\.0053to\+0\.3464\+0\.3464, and its interval excludes zero in all 10 conditions\. The nonlinear reader beats the affine one on the same contrasts in nine conditions, so the option scores do carry some nonlinearly accessible structure; it is simply far smaller than what the full state supports\.
#### The selection budget is not the explanation\.
AaffA\_\{\\mathrm\{aff\}\}selects its penalty by nested cross\-validation whileAstateA\_\{\\mathrm\{state\}\}uses a fixed penalty, so the subtracted term is the tuned one\. Refitting both ends under a matched nested budget changes the coverage gap by at most0\.01640\.0164in absolute accuracy \(Qwen2\.5\-Omni×\\timesVESUS\) and leaves five conditions slightly lower than reported\. Every coverage interval still excludes zero\. The main text reports the fixed\-penalty specification; the matched refit confirms the result without systematically favoring either reader\.
#### No single speaker drives either gap\.
Recomputing both terms with each speaker’s rows removed in turn, without refitting, leaves both terms positive in every replicate of every condition\. This check does not use the bootstrap’s resampling assumptions at all, which matters most for the 10\-speaker corpus\.
TABLE S6:Capacity\-matched refit and delete\-one\-speaker jackknife\.Δ\\Deltais the change in the coverage gap under the matched budget relative to Table[S5](https://arxiv.org/html/2608.06409#S3.T5)\.
## S5Logit Correction Details
Offsets are estimated per condition and per prompt variant from the marginal option probabilities of the evaluation rows, using no emotion labels, and are centered before use\. For every condition, the offsets and the corrected accuracy they imply were written to disk before the corrected\-generation pass ran, so each row of Table[S7](https://arxiv.org/html/2608.06409#S5.T7)compares a generated result against a prediction fixed in advance\. The four Qwen conditions were re\-estimated on the full corpora for this table; their baselines reproduceAgenA\_\{\\mathrm\{gen\}\}to four decimals\.
Prediction and generation agree exactly in five conditions\. Qwen2\.5\-Omni×\\timesVESUS is the only condition in which correction reduces parseability: 0\.0019 of rows leave the option set\. The remaining disagreements change one valid option into another and therefore do not create a format cost\.
TABLE S7:Label\-free logit correction\.*Predicted*is the corrected accuracy registered before the generation pass;*generated*is the observed one\.*Flip*gives the predicted and observed fraction of rows whose answer changes\.*Recovery*is the gain divided byΔdecision\\Delta\_\{\\mathrm\{decision\}\}\.The recovery ratio is reported per condition rather than summarized, because it is unstable when its denominator is small and because the supervised affine reader is a diagnostic reference rather than a target the label\-free offset is expected to reach\. The main text quotes the 0\.27 to 0\.87 range across all 10 conditions\.
## S6Readout Subspaces and Held\-Out Decoding
TABLE S8:Held\-out rank\-three subspace decodability atL∗L^\{\*\}, plotted in main\-text Fig\.[4](https://arxiv.org/html/2608.06409#S5.F4)\(a\)\.rand3\\mathrm\{rand\}\_\{3\}is averaged over 20 independent draws\. Intervals forV3−rand3V\_\{3\}\-\\mathrm\{rand\}\_\{3\}are speaker\-clustered over the held\-out speakers and exclude zero except on Phi\-4\-MM×\\timesVESUS\.With seed 0, half of the speakers are assigned to the training split and the remainder to the held\-out split; construction is prompt\-specific\. LetWoptW\_\{\\mathrm\{opt\}\}stack the four option\-token output rows,V4=row\(Wopt\)V\_\{4\}=\\operatorname\{row\}\(W\_\{\\mathrm\{opt\}\}\), andPV4P\_\{V\_\{4\}\}project onto that space\. LetHHcontain training states atL∗L^\{\*\}, and let𝒩f\\mathcal\{N\}\_\{f\}be the final normalization\. We standardizeH⟂=𝒩f\(H\)\(I−PV4\)H\_\{\\perp\}=\\mathcal\{N\}\_\{f\}\(H\)\(I\-P\_\{V\_\{4\}\}\), fit an L2 multinomial logistic regression \(C=0\.01C=0\.01, 3,000 iterations\), and define
B⟂\\displaystyle B\_\{\\perp\}=BstdD−1\(I−PV4\)=UΣR⊤,\\displaystyle=B\_\{\\mathrm\{std\}\}D^\{\-1\}\(I\-P\_\{V\_\{4\}\}\)=U\\Sigma R^\{\\top\},\(S1\)Sdecoding\\displaystyle S\_\{\\mathrm\{decoding\}\}=span\{Ri,:\}i=13,\\displaystyle=\\operatorname\{span\}\\\{R\_\{i,:\}\\\}\_\{i=1\}^\{3\},whereBstdB\_\{\\mathrm\{std\}\}is the standardized coefficient matrix andDDcontains the fitted feature scales\. Reprojection and re\-orthonormalization reduce leakage intoV4V\_\{4\}below10−610^\{\-6\}\. Because the relative\-contrast spaceV3V\_\{3\}is contained inV4V\_\{4\},Sdecoding⊂V4⟂⊂V3⟂S\_\{\\mathrm\{decoding\}\}\\subset V\_\{4\}^\{\\perp\}\\subset V\_\{3\}^\{\\perp\}\.
The rank\-four spaceSinterventionS\_\{\\mathrm\{intervention\}\}used in the causal analysis is built from the leading four discriminant directions\. Because both spaces come from the same singular basis,SdecodingS\_\{\\mathrm\{decoding\}\}is contained inSinterventionS\_\{\\mathrm\{intervention\}\}by construction\. The rank\-four intervention therefore contains the decoded directions, although causal effects need not be monotone under subspace expansion\. The causal analysis uses the matched rank\-four geometry\.
Held\-out decodability uses a separate standardized L2 multinomial probe \(C=0\.5C=0\.5, 2,000 iterations\), fitted on the training split and scored on the held\-out split\. Correctness is averaged over four prompt variants, and 95% intervals use 2,000 speaker\-bootstrap resamples\. For example, Phi\-4×\\timesCREMA\-D uses 986 clips from 45 training speakers and 1,014 clips from 46 held\-out speakers\. Eachrand3\\mathrm\{rand\}\_\{3\}draw is a standard\-normal sample projected intoV4⟂V\_\{4\}^\{\\perp\}and re\-orthonormalized, with the raw seed reused across the four prompts so that only the per\-prompt projector differs\.
### S6\-AAveraging the random reference over draws
A single random subspace is a noisy reference\. Repeating the entire probe over 20 independent draws \(Table[S9](https://arxiv.org/html/2608.06409#S6.T9)\) shows that same\-rank draws vary by0\.0170\.017to0\.0530\.053in standard deviation, enough to move a small contrast across zero\. The decoding results in the main text therefore use the draw\-averaged reference: per\-clip correctness is averaged over the 20 draws before scoring and bootstrapping, exactly as it is averaged over the four prompts\.
The native contrast is positive on all 20 draws in seven conditions\. It is positive on 16 draws for Kimi\-Audio×\\timesVESUS, 17 for Phi\-4\-MM×\\timesCREMA\-D, and only 2 for Phi\-4\-MM×\\timesVESUS, confirming that small native contrasts can depend on the random draw\. By contrast,Sdecoding−rand3S\_\{\\mathrm\{decoding\}\}\-\\mathrm\{rand\}\_\{3\}is positive on all 20 draws in all ten conditions\.
TABLE S9:The random reference across 20 independent rank\-three draws\.*sd*and*range*are over draws, not over speakers\. The last two columns count the draws on which each contrast is positive\.Table[S10](https://arxiv.org/html/2608.06409#S6.T10)adds the intervals omitted from the main text\.Sdecoding−rand3S\_\{\\mathrm\{decoding\}\}\-\\mathrm\{rand\}\_\{3\}is positive with an interval excluding zero in all ten conditions\. AddingSdecodingS\_\{\\mathrm\{decoding\}\}toV3V\_\{3\}also improves decoding in eight conditions; the two nonsignificant Audio\-Flamingo\-3 increments reflect that its native readout is already highly informative\.
TABLE S10:Held\-out subspace decodability with intervals\.V3−rand3V\_\{3\}\-\\mathrm\{rand\}\_\{3\}compares two spaces that are both fixed without emotion labels;SdecodingS\_\{\\mathrm\{decoding\}\}is selected on the training split, so its margin overrand3\\mathrm\{rand\}\_\{3\}includes a supervised search advantage\.*n\.s\.*marks an interval containing zero\.
## S7Minimal\-Pair Activation Replacement
### S7\-APairing, per\-arm rates, and statistical units
The primary causal pass uses strict minimal pairs from the held\-out split: receiver and donor share speaker and transcript but differ in emotion\. The donor representation is the state of one real clip under the same prompt, not a training\-set or class\-average state, and the same clip pair is used across all arms and prompts\. The training split still fixesL∗L^\{\*\},SinterventionS\_\{\\mathrm\{intervention\}\}, and the random subspace before any held\-out intervention\.
Within the 2,000\-clip stratified sample, 866 of 1,014 held\-out CREMA\-D clips and 474 of 1,000 held\-out VESUS clips have an eligible minimal\-pair donor\. We sample 400 eligible receivers per condition and evaluate four prompts, giving 1,600 rows in each of the ten conditions\. Because receiver and donor share a speaker, intervals cluster the paired outcomes by speaker: 46 clusters on CREMA\-D and 5 on VESUS\.
Table[S11](https://arxiv.org/html/2608.06409#S7.T11)reports the primary per\-condition effects that main\-text Fig\.[4](https://arxiv.org/html/2608.06409#S5.F4)\(b\) plots, together with each condition’sL∗L^\{\*\}\.
TABLE S11:Matched rank\-four activation replacement atL∗L^\{\*\}using held\-out minimal pairs\. The in\-span arm replacesV4V\_\{4\}and the readout\-external arm replacesSinterventionS\_\{\\mathrm\{intervention\}\}; each effect is the paired difference in answer\-change rate from the samerand4\\mathrm\{rand\}\_\{4\}arm, with speaker\-clustered 95% intervals, and n\.s\. marks an interval containing zero\. The last column places the readout\-external effect on a common scale as its one\-sided 95% upper bound divided by the in\-span effect of the same condition\.TABLE S12:Minimal\-pair answer\-change rates by arm atL∗L^\{\*\}\.VV,SinterventionS\_\{\\mathrm\{intervention\}\}, and Random are rank matched; Full replaces the entire state\.Subtracting the random rate gives positive in\-span effects in all ten conditions, from\+0\.0181\+0\.0181to\+0\.3631\+0\.3631\. The readout\-external effect reaches\+0\.0206\+0\.0206at most, and its relative upper bound is at most 7\.9% in eight conditions\. The two larger ratios occur where the in\-span denominator is small; Kimi\-Audio×\\timesVESUS is the only one of them with a detected readout\-external answer\-change effect\.
### S7\-BDonor\-content outcome
Answer change asks whether an edit moves the answer; donor following asks whether it moves specifically toward the donor emotion\. Table[S13](https://arxiv.org/html/2608.06409#S7.T13)reports both matched effects on this stricter outcome\.
TABLE S13:Minimal\-pair donor\-following effects relative to the same random arm, with speaker\-clustered 95% intervals\.Only Qwen2\-Audio×\\timesCREMA\-D shows readout\-external donor following above random\. Its\+0\.0106\+0\.0106effect shows limited content\-specific causal influence under a compatible replacement\. The Phi\-4\-MM in\-span arm does not transfer donor content in either corpus, so its already small answer\-change effects provide a weak scale reference\.
### S7\-CPerturbation\-magnitude diagnostic
Minimal pairing makes donor and receiver states more similar, so a null effect can coincide with a smaller edit\. Table[S14](https://arxiv.org/html/2608.06409#S7.T14)reports the applied relative state change in all ten conditions and, where an earlier cross\-speaker pass is available, the minimal\-to\-cross\-speaker ratio\.
TABLE S14:Mean relative intervention magnitude‖h~r−hr‖/‖hr‖\\\|\\widetilde\{h\}\_\{r\}\-h\_\{r\}\\\|/\\\|h\_\{r\}\\\|\. Ratios compare minimal\-pair with cross\-speaker replacement where available; the last column compares the two arms within the minimal\-pair pass\.For Qwen2\-Audio×\\timesCREMA\-D, minimal pairing retains 95% of the cross\-speakerSinterventionS\_\{\\mathrm\{intervention\}\}edit magnitude\. Within the minimal\-pair pass, that edit is 2\.23 times theV4V\_\{4\}edit, yet its answer\-change effect is\+0\.0206\+0\.0206rather than\+0\.3344\+0\.3344, ruling out a weaker external edit as the explanation in this condition\. Across all ten conditions,SinterventionS\_\{\\mathrm\{intervention\}\}is at least as large asV4V\_\{4\}in five\. Kimi\-Audio×\\timesVESUS belongs to this set, but both edits are the smallest in the study, matching its status as a weak\-intervention boundary case\.
### S7\-DRouting capacity aboveL∗L^\{\*\}
The selectedL∗L^\{\*\}lies one to four blocks below the top of each stack \(Table[S11](https://arxiv.org/html/2608.06409#S7.T11)\)\. A weakSinterventionS\_\{\\mathrm\{intervention\}\}effect could therefore have a simple explanation: the remaining blocks might be unable to route anyV4⟂V\_\{4\}^\{\\perp\}component into the option logits\.
The full\-state arm tests this possibility\. Relative to the in\-span replacement, it adds the donor’s entireV4⟂V\_\{4\}^\{\\perp\}component\. If the remaining blocks were unresponsive to that component, the two arms would have the same effect\. Instead, the full arm changes the answer more often than the in\-span arm in every condition, from 0\.288 versus 0\.245 on Qwen2\-Audio×\\timesVESUS to 0\.883 versus 0\.118 on Audio\-Flamingo\-3×\\timesCREMA\-D \(Table[S12](https://arxiv.org/html/2608.06409#S7.T12)\)\.
The additional changes generally move toward the donor emotion\. Relative to the unpatched baseline, the full arm raises donor following in all ten conditions: by\+0\.161\+0\.161to\+0\.839\+0\.839in the eight non\-Phi\-4\-MM conditions, and by\+0\.046\+0\.046and\+0\.009\+0\.009in the two Phi\-4\-MM conditions\. The weakSinterventionS\_\{\\mathrm\{intervention\}\}effects therefore cannot be explained solely by a downstream pathway that is unresponsive to readout\-external content\.
### S7\-EMinimal\-pair depth scan: protocol and consistency
TABLE S15:Minimal\-pair replacement at eight depths of Qwen2\-Audio×\\timesCREMA\-D, on all 866 eligible receivers, with the readout\-external subspace refit at each depth\. Each cell is the paired difference from the same\-rank random arm, with speaker\-clustered 95% intervals; n\.s\. marks an interval containing zero, andL=28L=28is theL∗L^\{\*\}used in Table[S11](https://arxiv.org/html/2608.06409#S7.T11)\.Table[S15](https://arxiv.org/html/2608.06409#S7.T15)repeats the minimal\-pair intervention at layers 16, 18, 20, 22, 24, 26, 28, and 31 of Qwen2\-Audio×\\timesCREMA\-D\. Main\-text Fig\.[5](https://arxiv.org/html/2608.06409#S5.F5)plots the donor\-following columns\. The pairing, receiver split,V4V\_\{4\}, andrand4\\mathrm\{rand\}\_\{4\}are the same as in the primary pass\. At each depth,SinterventionS\_\{\\mathrm\{intervention\}\}is refit on the training speakers using the construction in Section[S6](https://arxiv.org/html/2608.06409#S6a)\.
All 866 held\-out receivers with an eligible same\-speaker, same\-transcript donor are included under four prompt variants, giving 3,464 rows per cell\. Intervals cluster on receiver speaker\. At layer 28, the selectedL∗L^\{\*\}, the readout\-external answer\-change effect is\+0\.0147\+0\.0147\[\+0\.0111,\+0\.0183\+0\.0111,\+0\.0183\], close to the primary\-pass estimate of\+0\.0206\+0\.0206\[\+0\.0136,\+0\.0276\+0\.0136,\+0\.0276\] obtained from a 400\-receiver sample\.
Beyond the peak at layer 20, the decline is not layer\-by\-layer monotone; the effects at layers 22 through 26 sit within one another’s intervals, and the ordering claim the table supports is that the external effect is largest at layer 20 and smallest at layers 16 and 31\.
### S7\-FPropagation of the injected readout\-external component
We repeat the layer\-20 and layer\-24SinterventionS\_\{\\mathrm\{intervention\}\}replacements on 400 receivers and record the induced answer\-position differenceδ\(L′\)=hpatched\(L′\)−hunpatched\(L′\)\\delta\(L^\{\\prime\}\)=h\_\{\\mathrm\{patched\}\}\(L^\{\\prime\}\)\-h\_\{\\mathrm\{unpatched\}\}\(L^\{\\prime\}\)at every later layer, together with the endpoint option logits\.
After a layer\-20 replacement, the component ofδ\\deltain the injected subspace retains 0\.88 of its original norm at layer 31\. The perturbation therefore persists\. Its overlap withV4V\_\{4\}grows from zero at injection to about 0\.07 of‖δ‖\\\|\\delta\\\|by layer 26 and then remains near that level, showing that part of the perturbation reaches the native readout span\.
The endpoint logits also move toward the donor emotion\. Relative to the random arm, the donor\-option logit minus the mean of the other three option logits shifts by\+0\.190\+0\.190\[\+0\.176,\+0\.204\+0\.176,\+0\.204\] after the layer\-20 replacement and by\+0\.312\+0\.312\[\+0\.289,\+0\.335\+0\.289,\+0\.335\] after the layer\-24 replacement\. Thus the injected component reaches the final option scores with the donor’s sign, but usually not strongly enough to change their ordering\. Because the last saved state may include the model’s final normalization, this endpoint statement uses the recorded option logits rather than the state\-norm decomposition\.
## S8Controls for Measured Surface Acoustic Cues
These analyses test whether the held\-out decodability ofSdecodingS\_\{\\mathrm\{decoding\}\}can be explained by measured surface acoustic cues\. We define the descriptor panels, quantify how predictive the cues are, remove them from the subspace coordinates, and separately remove absolute level from the input audio\.
### S8\-ADescriptor Panel
Ten clip\-level acoustic\-prosodic descriptors are computed once per recording and are model\-independent: duration; voiced\-frame fraction; five fundamental\-frequency statistics \(mean, standard deviation, range, terminal value, and slope of the F0 track\); and three RMS\-energy statistics \(mean, standard deviation, and max\-minus\-min spread\)\. RMS descriptors enter all emotion\-corpus analyses in log units\. In the residualization and decoding analyses, missing descriptor values, standardization statistics, ordinary\-least\-squares coefficients, and class means are all estimated on training\-speaker rows only and then applied to held\-out rows\. The descriptor\-only decoderAdescA\_\{\\mathrm\{desc\}\}uses the same probe family, regularization\-selection protocol, and speaker split as the subspace decoders\.
An extended panel used for robustness adds ten further descriptors: spectral tilt \(the regression slope of the long\-term average spectrum in dB over log\-frequency\), spectral centroid mean and standard deviation, local jitter, local shimmer, harmonics\-to\-noise ratio, and the first four DCT coefficients of the time\-interpolated log\-F0 contour\. Jitter, shimmer, and harmonics\-to\-noise ratio are computed with Praat via parselmouth; coverage is complete on both corpora\. The nonlinear removal variant replaces the ordinary\-least\-squares residualization with per\-dimension gradient\-boosted trees, again fit on training\-speaker rows only\.
### S8\-BStimulus\-Level Loudness Statistics
Table[S16](https://arxiv.org/html/2608.06409#S8.T16)reports, for every pair of emotion classes, the class\-mean RMS\-level difference in dB and how well dB level alone separates the pair\. Folded AUC ismax\(AUC,1−AUC\)\\max\(\\mathrm\{AUC\},1\-\\mathrm\{AUC\}\)and is therefore orientation\-free\. Uncertainty is a speaker\-clustered bootstrap with 2,000 resamples\. On CREMA\-D, level alone is a strong class separator for several pairs; on VESUS the gaps are smaller\. These statistics motivate treating surface loudness as an explicit alternative explanation rather than an afterthought\.
TABLE S16:Stimulus\-level RMS\-loudness differences between emotion classes\.Δ\\DeltadB is the class\-mean dB RMS gap \(first class minus second\); folded AUC measures how well dB RMS alone separates the pair\. Speaker\-clustered bootstrap 95% intervals\.
### S8\-CBase Descriptor Removal
TABLE S17:Linear controls for the ten measured acoustic\-prosodic descriptors atL∗L^\{\*\}\.AdescA\_\{\\mathrm\{desc\}\}is held\-out four\-class accuracy from the descriptors alone, with no hidden state\. The resid columns re\-decode each rank\-three projection after removing the descriptors, with the regression fit on training speakers only\. The last column is the pairedSdecodingS\_\{\\mathrm\{decoding\}\}drop with speaker\-clustered 95% intervals; n\.s\. marks an interval containing zero\. TheV3V\_\{3\}andSdecodingS\_\{\\mathrm\{decoding\}\}columns match Table[S8](https://arxiv.org/html/2608.06409#S6.T8)\.Table[S17](https://arxiv.org/html/2608.06409#S8.T17)reports the base control summarized in main\-text Fig\.[6](https://arxiv.org/html/2608.06409#S5.F6)\. It gives the descriptor\-only referenceAdescA\_\{\\mathrm\{desc\}\}, decoding after linear removal of the ten descriptors, and the paired drop inSdecodingS\_\{\\mathrm\{decoding\}\}accuracy\.
### S8\-DSubspace–Descriptor Associations
Table[S18](https://arxiv.org/html/2608.06409#S8.T18)reports, for each condition and subspace, the single descriptor best predicted from the rank\-three projection, as out\-of\-sampleR2R^\{2\}under an ordinary\-least\-squares fit on training speakers\. The pooled column predicts the raw descriptor; the within\-class column first centers both the descriptor and the projection by their training\-speaker class means, so it measures covariation that is not explained by class membership\. The dominant descriptors are energy statistics in nearly every condition, and the within\-class values are substantially smaller than the pooled ones, indicating that much of the pooled association reflects class structure\. These associations are reported descriptively, without a multiplicity correction across descriptors and conditions\.
TABLE S18:Strongest acoustic\-descriptor association per subspace: the top\-R2R^\{2\}descriptor, with held\-out OLSR2R^\{2\}in parentheses\. Pooled: descriptor predicted directly from the rank\-three projection\. Within\-class: both centered by training\-speaker class means first\.
### S8\-EExtended Panel and Nonlinear Removal
Table[S19](https://arxiv.org/html/2608.06409#S8.T19)repeats the residualized decoding of the main text under descriptor removal of increasing strength: the ten\-descriptor panel removed linearly, the extended twenty\-descriptor panel removed linearly, and the extended panel removed with gradient\-boosted trees\. Held\-outSdecodingS\_\{\\mathrm\{decoding\}\}decodability remains above the 0\.25 chance level in every condition under every variant\. The descriptor\-only reference also strengthens slightly with the extended panel, from 0\.640 to 0\.664 on CREMA\-D and from 0\.343 to 0\.378 on VESUS, confirming that the added features carry usable information thatSdecodingS\_\{\\mathrm\{decoding\}\}nevertheless exceeds\.
TABLE S19:Held\-outSdecodingS\_\{\\mathrm\{decoding\}\}decodability after descriptor removal of increasing strength\. resid\-10: ordinary least squares on the ten\-descriptor panel \(main text\)\. resid\-20: the same on the extended twenty\-descriptor panel\. GBRT\-20: gradient\-boosted\-tree removal of the extended panel\. Chance is 0\.25\.
### S8\-FInput\-Side Loudness Equalization
The residualization analyses remove surface cues from the state side\. The input\-side check removes absolute level from the audio itself: every clip is RMS\-equalized to a fixed target with peak limiting, answer\-position states are re\-extracted under the main protocol \(the same clips, the four multiple\-choice prompts, and the per\-conditionL∗L^\{\*\}\), and the subspace decoding is repeated\. Two arms are scored\. The replication arm rebuildsSdecodingS\_\{\\mathrm\{decoding\}\}and refits the probe on equalized training\-speaker rows with the split held fixed;V3V\_\{3\}is a function of the model weights and is reused unchanged\. The transfer arm applies the probes fit on raw states, without refitting, to the equalized held\-out rows\.
TABLE S20:Raw versus loudness\-equalized held\-out decodability under the main protocol\. eq:SdecodingS\_\{\\mathrm\{decoding\}\}and probe rebuilt on equalized audio with the split held fixed\. transfer: raw\-fit probe applied without refitting to equalized held\-out clips\. The last column is the paired raw\-minus\-equalizedSdecodingS\_\{\\mathrm\{decoding\}\}difference with speaker\-clustered bootstrap 95% intervals\.Table[S20](https://arxiv.org/html/2608.06409#S8.T20)shows both arms\. Held\-outSdecodingS\_\{\\mathrm\{decoding\}\}accuracy changes by at most 0\.022 under replication and 0\.030 under transfer across all ten conditions, and the paired raw\-minus\-equalized interval includes zero in eight of ten\. Absolute recording level therefore does not explain most of the readout\-external decodability\. This control removes only absolute level, not energy dynamics, fundamental frequency, or spectral cues; gain normalization in the audio front ends may also contribute to the observed robustness\.相似文章
探究语言模型的思维失调过程
本文提出通过将LLM的失调分解为细粒度的认知过程(失调指标),并利用线性探针检测内部激活中的这些指标,从而在分布外对话记录上实现了高AUROC。
语言模型决策中的修辞错位表征
本文提出了一个语言模型修辞错位的框架,其中呈现方式可能在人类决策中引发有害的认知偏差,并通过临床场景的实验进行了验证。
在Token空间中解读情感:SpeechLLMs的情感识别判别式适配
本文提出了针对SpeechLLMs的情感识别判别式适配方法,通过在线性分类头上使用最终提示词token的隐状态,提升性能和可解释性,消除幻觉并增强情感方向分析。
缓解流形偏离:面向可信MLLM解码的不确定性感知子空间矫正
本文介绍了MGAP,一种无需训练的解码方法,通过自适应地仅抑制语言先验中的有害部分,同时保留模型的语义流形,从而减少多模态大语言模型中的幻觉。该方法在POPE和CHAIR基准测试上优于先前的基线方法。
无需重新训练的跨方言泛化:面向MLIR的基于模式约束解码的基准与评估
本文介绍了跨多种方言的自然语言到MLIR代码生成的基准测试,以及一个基于模式约束的解码栈,该栈使得小型语言模型无需重新训练即可在结构验证器任务上匹配或超越大型代码语言模型。