Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models

arXiv cs.CL Papers

Summary

This research paper explores emotion-sensitive neurons in multimodal foundation models, revealing shared affective mechanisms between speech and facial emotion recognition through causal interventions and cross-modal analysis.

arXiv:2608.17102v1 Announce Type: new Abstract: Modern multimodal foundation models (MFMs) have made rapid progress on tasks requiring integrated perception across speech, vision, and language, including emotion recognition. However, it remains unclear whether they recognize speech and facial emotion through shared affective functional units or modality-specific pathways. We explore emotion-sensitive neurons (ESNs), sparse decoder neurons selectively associated with emotion categories, in three MFMs: Gemma-4-12B-it, MiniCPM-o-4.5, and Qwen2.5-Omni-7B. Using speech emotion recognition and facial expression recognition as complementary probes, we identify acoustic and visual ESNs. Visual ESNs are causally meaningful: deactivating them selectively impairs recognition of the associated facial emotion, whereas steering their activations selectively enhances recognition of that emotion relative to other emotion categories. Acoustic and visual ESNs further show emotion-matched overlap and similar layer-wise distributions, indicating partial structural alignment between affective representations across speech and faces. Finally, cross-modal interventions reveal bidirectional causal transfer: ESNs identified from one modality produce emotion-specific effects when applied to the other. Our findings provide one of the first cross-modality activation-level analyses of affective functional units in MFMs, suggesting that speech and facial emotion recognition partially converge onto sparse decoder-level components that can be localized and manipulated without training.
Original Article
View Cached Full Text

Cached at: 08/19/26, 09:49 AM

# Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models
Source: [https://arxiv.org/html/2608.17102](https://arxiv.org/html/2608.17102)
Xiutian ZhaoLuqi SunAffiliation:Center for Language and Speech Processing \(CLSP\), Johns Hopkins University, USABjörn SchullerAffiliation:Group on Language, Audio & Music \(GLAM\), Imperial College London, UKBerrak SismanAffiliation:Center for Language and Speech Processing \(CLSP\), Johns Hopkins University, USA

###### Abstract

Modern multimodal foundation models \(MFMs\) have made rapid progress on tasks requiring integrated perception across speech, vision, and language, including emotion recognition\. However, it remains unclear whether they recognize speech and facial emotion through shared affective functional units or modality\-specific pathways\. We explore emotion\-sensitive neurons \(ESNs\), sparse decoder neurons selectively associated with emotion categories, in three MFMs: Gemma\-4\-12B\-it, MiniCPM\-o\-4\.5, and Qwen2\.5\-Omni\-7B\. Using speech emotion recognition and facial expression recognition as complementary probes, we identify acoustic and visual ESNs\. Visual ESNs are causally meaningful: deactivating them selectively impairs recognition of the associated facial emotion, whereas steering their activations selectively enhances recognition of that emotion relative to other emotion categories\. Acoustic and visual ESNs further show emotion\-matched overlap and similar layer\-wise distributions, indicating partial structural alignment between affective representations across speech and faces\. Finally, cross\-modal interventions reveal bidirectional causal transfer: ESNs identified from one modality produce emotion\-specific effects when applied to the other\. Our findings provide one of the first cross\-modality activation\-level analyses of affective functional units in MFMs, suggesting that speech and facial emotion recognition partially converge onto sparse decoder\-level components that can be localized and manipulated without training\.

###### Index Terms:

multimodal foundation models, affective computing, speech emotion recognition, facial expression recognition, activation intervention

## IIntroduction

Emotion perception in natural communication is inherently multimodal\. In face\-to\-face interaction, people infer affect from facial expressions, speech prosody, lexical content, body movement, and context\[[14](https://arxiv.org/html/2608.17102#bib.bib44),[40](https://arxiv.org/html/2608.17102#bib.bib49),[7](https://arxiv.org/html/2608.17102#bib.bib50),[1](https://arxiv.org/html/2608.17102#bib.bib51),[38](https://arxiv.org/html/2608.17102#bib.bib55),[29](https://arxiv.org/html/2608.17102#bib.bib53),[53](https://arxiv.org/html/2608.17102#bib.bib26)\]\. Among these channels, speech and facial behavior are among the most commonly studied and practically important sources of affective information\. Facial expressions provide visible cues such as muscle configuration and gaze, while speech conveys affect through pitch, rhythm, intensity, and timing\[[45](https://arxiv.org/html/2608.17102#bib.bib31),[33](https://arxiv.org/html/2608.17102#bib.bib43),[3](https://arxiv.org/html/2608.17102#bib.bib45)\]\. Psychological and affective computing studies have long debated the extent to which emotion perception reflects modality\-general affective categories, modality\-specific cue patterns, and contextual or culturally learned regularities\[[31](https://arxiv.org/html/2608.17102#bib.bib32)\]\. This makes speech and facial emotion recognition a natural testbed for asking whether modern multimodal foundation models \(MFMs\) develop shared internal mechanisms for affect\.

Recent progress in MFMs has substantially expanded the scope of machine perception\. Models such as MiniCPM\-o\-4\.5\[[54](https://arxiv.org/html/2608.17102#bib.bib57)\], Qwen2\.5\-Omni\[[50](https://arxiv.org/html/2608.17102#bib.bib58)\], and Gemma\-4\[[17](https://arxiv.org/html/2608.17102#bib.bib56)\]can process images, speech, and text within a unified interface\. These systems have achieved strong performance on audio and vision understanding tasks\[[13](https://arxiv.org/html/2608.17102#bib.bib35),[43](https://arxiv.org/html/2608.17102#bib.bib36),[21](https://arxiv.org/html/2608.17102#bib.bib34),[32](https://arxiv.org/html/2608.17102#bib.bib33),[9](https://arxiv.org/html/2608.17102#bib.bib37),[19](https://arxiv.org/html/2608.17102#bib.bib41),[10](https://arxiv.org/html/2608.17102#bib.bib47)\], and have shown growing potential for affective applications, including speech emotion recognition \(SER\) and facial emotion recognition \(FER\)\[[11](https://arxiv.org/html/2608.17102#bib.bib38),[16](https://arxiv.org/html/2608.17102#bib.bib42),[51](https://arxiv.org/html/2608.17102#bib.bib30),[20](https://arxiv.org/html/2608.17102#bib.bib3),[28](https://arxiv.org/html/2608.17102#bib.bib48)\]\. However, behavioral success alone does not reveal whether a model recognizes anger in a face and anger in a voice through related internal components, or whether cross\-modal consistency emerges only near final answer selection\.

Mechanistic interpretability provides tools for localizing and manipulating internal components that support model behavior on affective tasks\. Prior work has shown that individual neurons or sparse groups of units can align with human\-interpretable concepts in vision models\[[5](https://arxiv.org/html/2608.17102#bib.bib11),[6](https://arxiv.org/html/2608.17102#bib.bib12),[12](https://arxiv.org/html/2608.17102#bib.bib14)\], language models\[[41](https://arxiv.org/html/2608.17102#bib.bib16),[55](https://arxiv.org/html/2608.17102#bib.bib20),[23](https://arxiv.org/html/2608.17102#bib.bib29)\], and multimodal systems\[[22](https://arxiv.org/html/2608.17102#bib.bib2),[15](https://arxiv.org/html/2608.17102#bib.bib19),[24](https://arxiv.org/html/2608.17102#bib.bib18),[49](https://arxiv.org/html/2608.17102#bib.bib10)\]\. In this work, a “neuron” refers to a scalar hidden unit inside a model component, and we focus on decoder MLP neurons because they provide a large set of nonlinear units through which multimodal representations are transformed before text generation\[[18](https://arxiv.org/html/2608.17102#bib.bib17),[4](https://arxiv.org/html/2608.17102#bib.bib15),[39](https://arxiv.org/html/2608.17102#bib.bib13)\]\. In parallel, affective modeling has studied how emotional style can be represented and controlled in speech generation and conversion systems\[[46](https://arxiv.org/html/2608.17102#bib.bib27),[48](https://arxiv.org/html/2608.17102#bib.bib5)\]\. More recent work connects affect and interpretability by identifying emotion\-related neurons or circuits in language and audio\-language models, and by showing that activation steering can causally modulate emotional behavior\[[27](https://arxiv.org/html/2608.17102#bib.bib22),[44](https://arxiv.org/html/2608.17102#bib.bib6),[37](https://arxiv.org/html/2608.17102#bib.bib40)\]\. However, these analyses remain largely confined to a single modality, especially text or speech, leaving open whether affect\-sensitive functional units are modality\-specific or partially shared across perceptual channels\.

This work investigates cross\-modal emotion\-sensitive neurons \(ESNs\) in MFMs\. Following prior work on speech emotion recognition in large audio\-language models\[[58](https://arxiv.org/html/2608.17102#bib.bib39)\], we define*ESNs*as sparse decoder MLP units whose activation patterns are selectively associated with affective input categories and whose behavioral role can be tested through intervention\. We distinguish*acoustic emotion\-sensitive neurons*\(A\-ESNs\), identified from SER activations, from*visual emotion\-sensitive neurons*\(V\-ESNs\), identified from FER activations\. Across Gemma\-4\-12B\-it\[[17](https://arxiv.org/html/2608.17102#bib.bib56)\], MiniCPM\-o\-4\.5\[[54](https://arxiv.org/html/2608.17102#bib.bib57)\], and Qwen2\.5\-Omni\-7B\[[50](https://arxiv.org/html/2608.17102#bib.bib58)\], we identify ESNs using a contrastive activation criterion, analyze their structural alignment through overlap and layer\-wise distributions, and evaluate their causal role through deactivation and steering\. Deactivation tests whether suppressing a selected neuron set disrupts recognition of the associated emotion, while steering tests whether amplifying that set selectively enhances recognition of the corresponding emotion\.

Our first question is whether visual ESNs exist and whether they are functionally meaningful\. Prior work has shown that SER\-derived A\-ESNs can be identified and causally validated in audio\-language models\[[58](https://arxiv.org/html/2608.17102#bib.bib39)\]\. However, it remains unclear whether neurons with analogous functions can be discovered from affective tasks in other modalities, particularly vision\. We therefore identify V\-ESNs from correctly recognized FER examples and intervene on them during FER\. The results show that V\-ESNs are causally influential for MFMs’ affective recognition behavior in a manner analogous to A\-ESNs: deactivation selectively reduces recognition of the associated facial emotion, whereas steering increases predictions of the targeted emotion\. These findings provide positive evidence that FER\-derived V\-ESNs constitute functional affective units, and show that ESNs can be independently discovered from both speech and facial emotion recognition tasks\.

These observations motivate our next question: whether speech and facial emotion recognition converge onto related affective representations\. To address this, we analyze the structural agreement between acoustic and visual ESNs inside the decoder\. Specifically, we compare A\-ESN and V\-ESN sets using emotion\-specific neuron\-level overlap and layer\-wise ESN distributions\. Across models, the overlap is sparse but consistently stronger for matched emotion categories than for mismatched ones, suggesting partial cross\-modal alignment\. The layer\-wise analysis further shows that both A\-ESNs and V\-ESNs are distributed across multiple decoder layers, with a tendency to concentrate more in middle and later layers than in early layers\. This indicates that the observed ESNs are not isolated artifacts of a single block\. Rather, affect\-sensitive units form sparse depth\-wise patterns that are broadly comparable across acoustic and visual inputs, although their precise localization remains model\-dependent\.

Structural overlap alone, however, does not establish functional sharing\. Overlap between A\-ESNs and V\-ESNs could arise from incidental co\-activation, shared answer\-token processing, or biases introduced by the selection criterion\. We therefore go beyond descriptive analysis and test whether the observed alignment is functional through cross\-modal causal interventions\. We transfer ESN masks across modalities: A\-ESNs identified from speech are applied during FER, and V\-ESNs identified from faces are applied during SER\. These transferred interventions produce emotion\-specific effects in both directions\. In aggregate, deactivating transferred ESNs impairs recognition of the matched emotion more than non\-matched emotions, while steering them increases the corresponding target\-emotion response relative to non\-target responses\. These effects are consistently stronger and more structured than those produced by random controls\. Thus, speech and facial emotion recognition in MFMs is not entirely modality\-isolated; decoder MLPs contain sparse affective components that are partially transferable across speech and faces\.

In summary, this study makes the following contributions\. First, we identify and causally validate visual ESNs in MFMs, extending neuron\-level affective analysis from speech to facial expression recognition\. Second, we show that acoustic and visual ESNs exhibit sparse emotion\-matched overlap and broadly similar layer\-wise distributions\. Third,we present, to our knowledge, one of the first cross\-modality analyses of affective functional units in MFMs, with a focus on speech and facial emotion recognition\.Our findings demonstrate bidirectional cross\-modal causal transfer between speech\- and face\-derived ESNs, providing evidence that affective processing in MFMs partially relies on shared decoder\-level components\.

## IIRelated Work

Neuron\-level specialization has been widely studied as a route to interpretable transformer\-based models, with evidence that individual units or sparse unit groups can encode human\-interpretable concepts in vision models\[[5](https://arxiv.org/html/2608.17102#bib.bib11),[6](https://arxiv.org/html/2608.17102#bib.bib12)\], language models\[[23](https://arxiv.org/html/2608.17102#bib.bib29),[41](https://arxiv.org/html/2608.17102#bib.bib16),[55](https://arxiv.org/html/2608.17102#bib.bib20)\], and multimodal systems\[[26](https://arxiv.org/html/2608.17102#bib.bib54),[22](https://arxiv.org/html/2608.17102#bib.bib2),[15](https://arxiv.org/html/2608.17102#bib.bib19),[49](https://arxiv.org/html/2608.17102#bib.bib10),[24](https://arxiv.org/html/2608.17102#bib.bib18)\]\. In affective modeling, related work has examined how emotional and paralinguistic attributes are represented or controlled in speech systems, including learned style embeddings\[[46](https://arxiv.org/html/2608.17102#bib.bib27)\], continuous affect\-control variables\[[48](https://arxiv.org/html/2608.17102#bib.bib5)\], and activation\-based emotion steering or circuit analysis in language models\[[25](https://arxiv.org/html/2608.17102#bib.bib46),[27](https://arxiv.org/html/2608.17102#bib.bib22),[44](https://arxiv.org/html/2608.17102#bib.bib6),[37](https://arxiv.org/html/2608.17102#bib.bib40)\]\. These studies establish that affective behavior can often be localized or modulated through internal representations, but they do not directly address whether such representations are shared across perceptual modalities\. Indeed, related evidence suggests that emotion\-related neurons may not always transfer across input conditions, such as across languages in multilingual encoders\[[35](https://arxiv.org/html/2608.17102#bib.bib28)\]\.

Beyond text, probing and dissection methods have analyzed what audio and multimodal models encode internally, including phonetic, speaker, prosodic, and broader acoustic concepts\[[36](https://arxiv.org/html/2608.17102#bib.bib25),[2](https://arxiv.org/html/2608.17102#bib.bib24),[52](https://arxiv.org/html/2608.17102#bib.bib4),[47](https://arxiv.org/html/2608.17102#bib.bib9)\]\. Most directly related to our work, recent studies have identified ESNs in large audio\-language models for SER\[[58](https://arxiv.org/html/2608.17102#bib.bib39),[57](https://arxiv.org/html/2608.17102#bib.bib8)\]and emotional voice conversion\[[59](https://arxiv.org/html/2608.17102#bib.bib7)\]\. However, existing analyses primarily focus on a single modality\. It remains unclear whether ESNs are modality\-specific or whether they form a partially shared substrate across different affective channels\. Our work narrows this gap by identifying visual ESNs from facial expression recognition and testing their overlap and transfer with acoustic ESNs from SER in MFMs\.

## IIIMethod

We apply a shared activation\-based probing pipeline to SER and FER: collect decoder MLP activations on correctly recognized examples, select sparse ESNs using a contrastive criterion, and evaluate their causal role through deactivation and steering\.

### III\-AEmotion\-Conditioned Activation Collection

Letℰ\\mathcal\{E\}denote the set of emotion categories\. For each task modalityq∈\{SER,FER\}q\\in\\\{\\mathrm\{SER\},\\mathrm\{FER\}\\\}and emotione∈ℰe\\in\\mathcal\{E\}, we run the original, unintervened model on the corresponding emotion recognition task\. We use a multiple\-choice question answering protocol and retain only correctly recognized examples for neuron identification, which reduces noise from model failure cases and yields cleaner emotion\-conditioned activation statistics\. Dataset\-specific sample sizes for ESN identification are reported in Section[IV](https://arxiv.org/html/2608.17102#S4)\.

We instrument the decoder MLP modules and record the activated SwiGLU gate outputs\[[34](https://arxiv.org/html/2608.17102#bib.bib1)\]\. For layerll, neuronnn, and valid token positiontt, letal,n,t\(q,e\)a^\{\(q,e\)\}\_\{l,n,t\}be the scalar gate activation under task modalityqqand emotionee\. We use a binary maskmt∈\{0,1\}m\_\{t\}\\in\\\{0,1\\\}to exclude irrelevant positions such as padding and instruction\-only tokens\. For each\(q,e\)\(q,e\), we compute the positive firing count and the number of valid positions asKl,n\(q,e\)=∑tmt​𝕀​\(al,n,t\(q,e\)\>0\)K^\{\(q,e\)\}\_\{l,n\}=\\sum\_\{t\}m\_\{t\}\\mathbb\{I\}\(a^\{\(q,e\)\}\_\{l,n,t\}\>0\)andT\(q,e\)=∑tmtT^\{\(q,e\)\}=\\sum\_\{t\}m\_\{t\}\. The normalized activation probability is thenPl,n\(q,e\)=Kl,n\(q,e\)/T\(q,e\)P^\{\(q,e\)\}\_\{l,n\}=K^\{\(q,e\)\}\_\{l,n\}/T^\{\(q,e\)\}, which measures how frequently a neuron fires for a given emotion and modality\.

### III\-BEmotion\-Sensitive Neuron Identification

We identify ESNs using Contrastive Activation Margin \(ConAct\)\[[56](https://arxiv.org/html/2608.17102#bib.bib21)\]\. For each modalityqq, layerll, and neuronnn, ConAct assigns the neuron to the emotion for which it fires most frequently and scores it by the margin between the highest and second\-highest activation probabilities\. Specifically, we defineel,n\(1\)​\(q\)=arg⁡maxe∈ℰ⁡Pl,n\(q,e\)e^\{\(1\)\}\_\{l,n\}\(q\)=\\arg\\max\_\{e\\in\\mathcal\{E\}\}P^\{\(q,e\)\}\_\{l,n\},Pl,n\(1\)​\(q\)=maxe∈ℰ⁡Pl,n\(q,e\)P^\{\(1\)\}\_\{l,n\}\(q\)=\\max\_\{e\\in\\mathcal\{E\}\}P^\{\(q,e\)\}\_\{l,n\}, andPl,n\(2\)​\(q\)=maxe≠el,n\(1\)​\(q\)⁡Pl,n\(q,e\)P^\{\(2\)\}\_\{l,n\}\(q\)=\\max\_\{e\\neq e^\{\(1\)\}\_\{l,n\}\(q\)\}P^\{\(q,e\)\}\_\{l,n\}\. The emotion\-specific ConAct score issl,n\(q,e\)=Pl,n\(1\)​\(q\)−Pl,n\(2\)​\(q\)s^\{\(q,e\)\}\_\{l,n\}=P^\{\(1\)\}\_\{l,n\}\(q\)\-P^\{\(2\)\}\_\{l,n\}\(q\)ife=el,n\(1\)​\(q\)e=e^\{\(1\)\}\_\{l,n\}\(q\), andsl,n\(q,e\)=0s^\{\(q,e\)\}\_\{l,n\}=0otherwise\.

For each modality\-emotion pair\(q,e\)\(q,e\), we rank all decoder MLP neurons bysl,n\(q,e\)s^\{\(q,e\)\}\_\{l,n\}and select the top fractionrras the corresponding ESN set\. In the FER experiments, we user=0\.5%r=0\.5\\%of all decoder MLP neurons to obtain V\-ESNs; the same selection strategy is used for A\-ESNs in SER\. We denote the selected neuron set for emotioneeasℐ\(q,e\)\\mathcal\{I\}^\{\(q,e\)\}, withℐ\(SER,e\)\\mathcal\{I\}^\{\(\\mathrm\{SER\},e\)\}corresponding to A\-ESNs andℐ\(FER,e\)\\mathcal\{I\}^\{\(\\mathrm\{FER\},e\)\}corresponding to V\-ESNs\. As a control, we also construct random masks by selecting the same number of neurons uniformly from the decoder MLPs, without using activation statistics\. We average results over five independent random masks\.

### III\-CCausal Interventions: Deactivation and Steering

Given an ESN setℐ\(q,e\)\\mathcal\{I\}^\{\(q,e\)\}, we evaluate its causal role through two interventions applied to decoder MLP gate activations: deactivation and steering\. Letgl,t∈ℝDlg\_\{l,t\}\\in\\mathbb\{R\}^\{D\_\{l\}\}be the activated gate vector at layerlland token positiontt, wheregl,t=act⁡\(gate​\_​proj​\(xl,t\)\)g\_\{l,t\}=\\mathrm\{act\}\(\\mathrm\{gate\\\_proj\}\(x\_\{l,t\}\)\)\. The modified gate vector is then used in the standard SwiGLU computation, i\.e\.,down​\_​proj​\(g~l,t⊙up​\_​proj​\(xl,t\)\)\\mathrm\{down\\\_proj\}\(\\tilde\{g\}\_\{l,t\}\\odot\\mathrm\{up\\\_proj\}\(x\_\{l,t\}\)\)\.

#### Deactivation

For deactivation, we suppress the selected neurons by setting their gate activations to zero\. For neuronnnin layerll, the mask isrl,n=0r\_\{l,n\}=0ifn∈ℐ\(q,e\)n\\in\\mathcal\{I\}^\{\(q,e\)\}andrl,n=1r\_\{l,n\}=1otherwise, yieldingg~l,tdeact=gl,t⊙rl\\tilde\{g\}^\{\\mathrm\{deact\}\}\_\{l,t\}=g\_\{l,t\}\\odot r\_\{l\}\.

#### Steering

For steering, we increase the selected neurons’ activations with a multiplicative gain\. The steering mask issl,n​\(α\)=1\+αs\_\{l,n\}\(\\alpha\)=1\+\\alphaifn∈ℐ\(q,e\)n\\in\\mathcal\{I\}^\{\(q,e\)\}andsl,n​\(α\)=1s\_\{l,n\}\(\\alpha\)=1otherwise, givingg~l,tsteer=gl,t⊙sl​\(α\)\\tilde\{g\}^\{\\mathrm\{steer\}\}\_\{l,t\}=g\_\{l,t\}\\odot s\_\{l\}\(\\alpha\)\. Unless otherwise stated, we setα=0\.5\\alpha=0\.5\.

These interventions provide complementary loss\- and gain\-of\-function tests\. We interpret an ESN set as causally emotion\-relevant when its effects are emotion\-selective and stronger than same\-size random\-mask controls\. The same protocol is used for both within\-modality validation and cross\-modal transfer\.

## IVExperiment Setup

### IV\-ADatasets and Models

We evaluate emotion recognition from both acoustic and visual inputs\. For SER, we use MSP\-Podcast\[[8](https://arxiv.org/html/2608.17102#bib.bib59)\]; for FER, we use AffectNet\[[30](https://arxiv.org/html/2608.17102#bib.bib60)\]\. We focus on five emotion categories shared by the two tasks: anger, fear, happiness, neutral, and sadness\. We sample 150 utterances per emotion for evaluation in SER task and 300 facial images per emotion for evaluation in FER\. The remaining correctly recognized SER/FER examples are used for A\-ESN/V\-ESN identification, from which we sample 100 successful cases per emotion for both tasks\. This separation ensures that neuron selection and causal evaluation are performed on disjoint instances\.

We evaluate three open\-source MFMs: Gemma\-4\-12B\-it\[[17](https://arxiv.org/html/2608.17102#bib.bib56)\], MiniCPM\-o\-4\.5\[[54](https://arxiv.org/html/2608.17102#bib.bib57)\], and Qwen2\.5\-Omni\-7B\[[50](https://arxiv.org/html/2608.17102#bib.bib58)\]\. These models support both acoustic and visual inputs and expose decoder MLP modules for activation\-level intervention\. All models are evaluated under the same multiple\-choice question answering protocol\. To reduce position and label bias\[[60](https://arxiv.org/html/2608.17102#bib.bib23)\], the order of emotion options is randomized across examples, and the model is instructed to output only the option index \(alphabetic letters\) rather than the emotion word\.

### IV\-BPrompting and Decoding

All inference is performed with deterministic decoding, using greedy search with temperature 0\. We set the maximum generation length to 20 tokens and apply lightweight post\-processing to extract the predicted option index from the model output\. Randomness enters through dataset sampling, option\-order randomization, and random\-mask construction; dataset sampling and option ordering are controlled by fixed seeds unless otherwise specified\. For random\-selection controls, we report averages over five independently sampled masks\.

### IV\-CNeuron Selection and Intervention Settings

Unless stated otherwise, ESNs are selected with ConAct from decoder MLP activations\. For each emotion, we select the topr=0\.5%r=0\.5\\%of all decoder MLP neurons\. In steering experiments, we use a multiplicative gain withα=0\.5\\alpha=0\.5\.

TABLE I:Mono\-modal causal effects of ESN interventions\.V\-ESN masks are identified and evaluated on FER; A\-ESN masks are identified and evaluated on SER\. Columns report the unmasked UAR, random\-mask controls, matched\-emotion performance, average non\-matched\-emotion performance, and the Self\-Cross Gap, as defined in Section[IV\-D](https://arxiv.org/html/2608.17102#S4.SS4)\.Δ\\Deltavalues indicate the changes relative to the unmasked UAR\. Arrows \(↓\\downarrow,↑\\uparrow\) indicate the expected direction of the matched\-emotion performance change\.![Refer to caption](https://arxiv.org/html/2608.17102v1/figures/mono/FER_AffectNet_to_AffectNet_gemma4_CAM_100_top0.005_ablate_Accuracy.png)\(a\)Deactivation,
Gemma\-4\-12B\-it
![Refer to caption](https://arxiv.org/html/2608.17102v1/figures/mono/FER_AffectNet_to_AffectNet_gemma4_CAM_100_top0.005_steer_alpha0.5_Accuracy.png)\(b\)Steering,
Gemma\-4\-12B\-it
![Refer to caption](https://arxiv.org/html/2608.17102v1/figures/mono/FER_AffectNet_to_AffectNet_mini_CAM_100_top0.005_ablate_Accuracy.png)\(c\)Deactivation,
MiniCPM\-o\-4\.5
![Refer to caption](https://arxiv.org/html/2608.17102v1/figures/mono/FER_AffectNet_to_AffectNet_mini_CAM_100_top0.005_steer_alpha0.5_Accuracy.png)\(d\)Steering,
MiniCPM\-o\-4\.5
![Refer to caption](https://arxiv.org/html/2608.17102v1/figures/mono/FER_AffectNet_to_AffectNet_qwen25_CAM_100_top0.005_ablate_Accuracy.png)\(e\)Deactivation,
Qwen2\.5\-Omni\-7B
![Refer to caption](https://arxiv.org/html/2608.17102v1/figures/mono/FER_AffectNet_to_AffectNet_qwen25_CAM_100_top0.005_steer_alpha0.5_Accuracy.png)\(f\)Steering,
Qwen2\.5\-Omni\-7B

Fig\. 1:Emotion\-specific FER effects of V\-ESN interventions\.Heatmaps report changes in FER accuracy after applying FER\-derived V\-ESN masks\. Rows denote the mask emotion and columns denote the FER test\-subset emotion; “Random” gives same\-size random\-mask controls\. Diagonal cells correspond to matched mask–test emotion pairs\.
### IV\-DEvaluation Metrics

We report baseline emotion\-recognition performance using unweighted average recall \(UAR\) over the five emotion categories\. For intervention analysis, all effects are measured relative to the original, unintervened model under the same evaluation protocol\. For an ESN mask associated with emotionee, the Self\-Emotion score is the post\-intervention accuracy on test examples whose ground\-truth label isee\. We report the average of this score over all target emotionse∈Ee\\in E\. The Avg\. Cross\-Emotion score is computed by first averaging, for each target maskee, the post\-intervention accuracies on all non\-target emotion subsetse′≠ee^\{\\prime\}\\neq e, and then averaging over target emotions\. TheΔ\\Deltavalues in the tables denote changes relative to the corresponding unmasked UAR under the same evaluation protocol\. Under deactivation, a negative gap indicates selective degradation of matched\-emotion recognition; under steering, a positive gap indicates selective improvement in matched\-emotion recognition relative to non\-target emotions\.

## VResults

### V\-AMono\-Modal Causal Validation of ESNs

We first test whether ESNs discovered from a modality are causally involved in emotion recognition within that same modality\. As shown in Table[I](https://arxiv.org/html/2608.17102#S4.T1), this validation succeeds for both FER\-derived V\-ESNs and SER\-derived A\-ESNs across all three MFMs\. Deactivation consistently reduces performance on the matched emotion more than on non\-matched emotions, yielding negative Self\-Cross Gaps, while steering produces the opposite pattern with positive gaps\. Same\-size random masks have only mild effects, indicating that the results are not explained by arbitrary perturbations of decoder MLP activations\.

Figure[1](https://arxiv.org/html/2608.17102#S4.F1)gives a finer\-grained view of the FER results\. The heatmaps show a broadly matched\-emotion structure: deactivation often produces the largest or among the largest decreases on the diagonal, and steering often increases the corresponding target\-emotion response more than most non\-target responses\. The pattern is clearest for several affective categories such as anger, fear, happiness, and sadness, whereas neutral shows a less uniformly targeted pattern\. This difference is plausible because neutral is closer to an absence of overt affect than to a strongly expressed emotion, and its recognition may depend more on suppressing evidence for other emotions than on activating a single positive affective pattern\.

This confirms that ESNs can be identified from facial\-expression activations, extending prior evidence that SER\-derived A\-ESNs are functional in SER\[[58](https://arxiv.org/html/2608.17102#bib.bib39)\]\. Moreover, the approximately opposite and emotion\-selective effects of deactivation and steering support the interpretation that these sparse neuron sets are emotion\-specific causal components, providing the mono\-modal basis for the cross\-modal analyses that follow\.

![Refer to caption](https://arxiv.org/html/2608.17102v1/figures/layers/layer_dist_emotions_gemma-4-12B-it_AffectNet_CAM_100_log.png)
![Refer to caption](https://arxiv.org/html/2608.17102v1/figures/layers/layer_dist_emotions_gemma-4-12B-it_MSP-PODCAST-Publish-1.12_CAM_100_log.png)
![Refer to caption](https://arxiv.org/html/2608.17102v1/figures/layers/layer_dist_emotions_MiniCPM-o-4_5_AffectNet_CAM_100_log.png)
![Refer to caption](https://arxiv.org/html/2608.17102v1/figures/layers/layer_dist_emotions_MiniCPM-o-4_5_MSP-PODCAST-Publish-1.12_CAM_100_log.png)
![Refer to caption](https://arxiv.org/html/2608.17102v1/figures/layers/layer_dist_emotions_Qwen2.5-Omni-7B_AffectNet_CAM_100_log.png)
![Refer to caption](https://arxiv.org/html/2608.17102v1/figures/layers/layer_dist_emotions_Qwen2.5-Omni-7B_MSP-PODCAST-Publish-1.12_CAM_100_log.png)

Fig\. 2:Layer\-wise distribution of selected ESNs \(ConAct\-selected,r=0\.5%r=0\.5\\%\)\. Panels show FER\-derived V\-ESNs and SER\-derived A\-ESNs for each model; rows denote decoder MLP layers and columns denote emotions\. Colors indicate neuron counts on a logarithmic scale\. ESNs are sparse but span multiple layers, with stronger concentration in middle and later layers\.![Refer to caption](https://arxiv.org/html/2608.17102v1/figures/overlap/CrossModal_gemma-4-12B-it_CAM_100_top0.005_JSC.png)\(a\)Gemma\-4\-12B\-it
![Refer to caption](https://arxiv.org/html/2608.17102v1/figures/overlap/CrossModal_MiniCPM-o-4_5_CAM_100_top0.005_JSC.png)\(b\)MiniCPM\-o\-4\.5
![Refer to caption](https://arxiv.org/html/2608.17102v1/figures/overlap/CrossModal_Qwen2.5-Omni-7B_CAM_100_top0.005_JSC.png)\(c\)Qwen2\.5\-Omni\-7B

Fig\. 3:Cross\-modal overlap between FER\-derived V\-ESNs and SER\-derived A\-ESNs\.Each heatmap reports Jaccard similarity coefficient across emotion\-specific neuron sets; rows denote FER emotions, columns denote SER emotions, and darker cells indicate larger overlap\.
### V\-BAcoustic and Visual ESNs Show Cross\-Modal Alignment

We next examine whether FER\-derived V\-ESNs and SER\-derived A\-ESNs are modality\-specific or partially shared\. We analyze both emotion\-level neuron\-set overlap and layer\-wise ESN localization\. Figure[3](https://arxiv.org/html/2608.17102#S5.F3)reports the Jaccard similarity between A\-ESNs and V\-ESNs for each pair of emotion categories\. Across all three models, the overlap is sparse but tends to be diagonal\-dominant: neuron sets for the same speech and facial emotion overlap more than mismatched emotion pairs, whose similarities are often near zero\. The strongest matched\-category overlaps vary by model: Gemma\-4\-12B\-it shows clearer alignment for anger, happiness, and sadness; MiniCPM\-o\-4\.5 for sadness, with additional alignment for anger, happiness, and neutral; and Qwen2\.5\-Omni\-7B for happiness and sadness\. The absolute similarity coefficients are small, as expected from selecting only the top 0\.5% of decoder MLP neurons, but the matched\-emotion structure suggests that the overlap is not merely due to generic task activation\.

Figure[2](https://arxiv.org/html/2608.17102#S5.F2)shows where the selected neurons occur across decoder depth\. In all models, both A\-ESNs and V\-ESNs are distributed over multiple layers rather than confined to a single block\. They are relatively rare in early layers and more frequent in middle and later layers, consistent with emotion information becoming more task\- or decision\-relevant deeper in the decoder\. The FER and SER profiles are broadly similar within each model, although their exact localization differs: MiniCPM\-o\-4\.5 has dense middle\-to\-late ESN selection for both modalities, Qwen2\.5\-Omni\-7B shows a more concentrated intermediate band, and Gemma\-4\-12B\-it is more diffuse across depth\. The late\-layer concentration also requires caution, because decoder states near the output may partly encode category\-level decision variables or answer\-selection processes\. However, this pattern is unlikely to be explained solely by lexical association\[[42](https://arxiv.org/html/2608.17102#bib.bib52)\]: the option order is randomized, models are instructed to output only the option index, and, more importantly, ESNs identified from different modalities and datasets exhibit emotion\-matched overlap and emotion\-specific causal effects, including under cross\-modal transfer\. We therefore interpret the late\-layer ESNs as decoder\-level affective functional units that may combine perceptual evidence with emotion\-category decision representations, rather than as purely lexical\-answer neurons\.

Notably, these similarities arise despite A\-ESNs and V\-ESNs being identified from different tasks, datasets, and input modalities\. This makes it less likely that the diagonal overlap and comparable layer\-wise profiles are explained solely by the selection procedure\. Instead, they suggest that speech and facial emotion recognition may converge onto a sparse set of decoder MLP units associated with emotion\-relevant processing\. These structural results suggest partial alignment between acoustic and visual ESNs, but they do not establish functional sharing\. We therefore test whether ESNs selected in one modality causally affect recognition in the other\.

### V\-CCross\-Modal ESN Interventions Reveal Bidirectional Transfer

TABLE II:Cross\-modal causal effects of ESN interventions\.SER\-derived A\-ESN masks are applied during FER, and FER\-derived V\-ESN masks are applied during SER\. Columns follow Table[I](https://arxiv.org/html/2608.17102#S4.T1)\.Having established that FER\-derived V\-ESNs are causally involved in FER, we next ask whether these neurons are modality\-specific or whether they also participate in affective processing across modalities\. To this end, we perform cross\-modal interventions: SER\-derived A\-ESN masks are applied during FER, and FER\-derived V\-ESN masks are applied during SER\. In both settings, the neuron sets are identified in one modality but intervened on in the other, providing a direct test of whether ESNs discovered from one perceptual channel causally influence recognition behavior in another\.

Table[II](https://arxiv.org/html/2608.17102#S5.T2)summarizes the cross\-modal intervention results\. Across all three models, ESNs identified in one modality produce structured effects when transferred to the other modality, whereas random interventions generally remain close to the unmasked baseline\. When applying A\-ESNs to FER, deactivation reduces matched self\-emotion performance for all models, with larger drops than the average changes on non\-target emotions\. Steering A\-ESNs during FER produces the complementary pattern, increasing the corresponding self\-emotion score and yielding positive Self\-Cross Gaps\. The magnitudes vary across models, with the strongest transfer observed for Gemma\-4\-12B\-it and more modest but consistent effects for MiniCPM\-o\-4\.5 and Qwen2\.5\-Omni\-7B\.

The reverse transfer shows the same qualitative pattern\. When applying V\-ESNs to SER, deactivation reduces recognition of the corresponding self\-emotion across models, while random deactivation remains close to zero\. Steering V\-ESNs improves SER performance on the corresponding self\-emotion category, producing positive Self\-Cross Gaps for all models\. These results suggest that FER\-derived V\-ESNs can participate in SER behavior, although the effect sizes indicate partial rather than complete sharing between the modalities\. The bidirectional transfer reduces the likelihood that the effect is specific to a single modality, and instead suggests a recurring organization in which decoder MLPs contain sparse emotion\-sensitive components accessible from both speech and facial cues\.

Figure[4](https://arxiv.org/html/2608.17102#S5.F4)illustrates the emotion\-specific intervention patterns for MiniCPM\-o\-4\.5\. The transferred interventions show a clearer matched\-emotion structure in aggregate than the random controls, although the heatmaps also contain non\-negligible off\-diagonal effects\. In particular, deactivation of transferred ESNs tends to reduce performance more for the corresponding self\-emotion than for the average non\-target emotions, while steering tends to increase the corresponding self\-emotion response\. These trends are consistent with the positive and negative Self\-Cross Gaps in Table[II](https://arxiv.org/html/2608.17102#S5.T2), but they also indicate that cross\-modal transfer is selective rather than perfectly emotion\-isolated\. Overall, these patterns strengthen the descriptive overlap analysis by showing that ESNs selected in one modality are not merely artifacts of the selection procedure, but can also causally modulate emotion\-recognition behavior in the other modality\.

![Refer to caption](https://arxiv.org/html/2608.17102v1/figures/cross/FER_MSP-PODCAST-Publish-1.12_to_AffectNet_mini_CAM_100_top0.005_ablate_Accuracy.png)\(a\)Deactivation, A\-ESN on FER
![Refer to caption](https://arxiv.org/html/2608.17102v1/figures/cross/FER_MSP-PODCAST-Publish-1.12_to_AffectNet_mini_CAM_100_top0.005_steer_alpha0.5_Accuracy.png)\(b\)Steering, A\-ESN on FER
![Refer to caption](https://arxiv.org/html/2608.17102v1/figures/cross/SER_AffectNet_to_MSP-PODCAST-Publish-1.12_mini_CAM_100_top0.005_ablate_Accuracy.png)\(c\)Deactivation, V\-ESN on SER
![Refer to caption](https://arxiv.org/html/2608.17102v1/figures/cross/SER_AffectNet_to_MSP-PODCAST-Publish-1.12_mini_CAM_100_top0.005_steer_alpha0.5_Accuracy.png)\(d\)Steering, V\-ESN on SER

Fig\. 4:Emotion\-specific effects of cross\-modal ESN transfer for MiniCPM\-o\-4\.5\.Heatmaps report changes in recognition accuracy when ESN masks identified from one modality are applied to the other modality\. Panels \(a,b\) apply SER\-derived A\-ESN masks during FER, and panels \(c,d\) apply FER\-derived V\-ESN masks during SER\.

## VIConclusion

We investigated whether speech and facial emotion recognition in MFMs rely on separate modality\-specific pathways or partially shared internal functional units\. Across Gemma\-4\-12B\-it, MiniCPM\-o\-4\.5, and Qwen2\.5\-Omni\-7B, we identified sparse ESNs from decoder MLP activations using SER and FER as complementary probes\. Both acoustic ESNs and visual ESNs are causally meaningful within their respective modalities: deactivation selectively impairs recognition of the matched emotion, while steering selectively improves recognition of that emotion\. Beyond mono\-modal validation, we found that A\-ESNs and V\-ESNs exhibit emotion\-matched overlap and broadly comparable layer\-wise distributions, with selected neurons tending to appear more in middle and later decoder layers\. Most importantly, cross\-modal interventions reveal partial bidirectional causal transfer: A\-ESNs identified from speech affect facial emotion recognition, and V\-ESNs identified from faces affect SER, with matched self\-emotion effects that are stronger than the corresponding random\-mask controls\. Together, these results suggest that speech and facial affect processing in MFMs partially converges onto sparse decoder\-level components that can be localized and manipulated without training\.

This work provides one of the first cross\-modality activation\-level analyses of affective functional units in MFMs, especially across speech and facial emotion recognition\. Our results offer a mechanistic view of how MFMs internally organize affective information across modalities\. Future work could move beyond post\-hoc localization toward building controllable affective interfaces for MFMs: identifying whether shared ESNs can be used to calibrate, debias, or personalize emotion perception across speech and vision\.

## Acknowledgment

This work is supported by the National Science Foundation \(NSF\) CAREER Award IIS\-2533652\.

## References

- \[1\]S\. M\. S\. A\. Abdullah, S\. Y\. A\. Ameen, M\. A\. Sadeeq, and S\. Zeebaree\(2021\)Multimodal emotion recognition using deep learning\.Journal of Applied Science and Technology Trends2\(01\),pp\. 73–79\.Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p1.1)\.
- \[2\]A\. Akman, Q\. Sun, and B\. W\. Schuller\(2025\)Improving audio explanations using audio language models\.IEEE Signal Processing Letters32\(\),pp\. 741–745\.External Links:[Document](https://dx.doi.org/10.1109/LSP.2025.3532218)Cited by:[§II](https://arxiv.org/html/2608.17102#S2.p2.1)\.
- \[3\]L\. F\. Barrett, R\. Adolphs, S\. Marsella, A\. M\. Martinez, and S\. D\. Pollak\(2019\)Emotional expressions reconsidered: challenges to inferring emotion from human facial movements\.Psychological science in the public interest20\(1\),pp\. 1–68\.Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p1.1)\.
- \[4\]A\. Bau, Y\. Belinkov, H\. Sajjad, N\. Durrani, F\. Dalvi, and J\. Glass\(2019\)Identifying and controlling important neurons in neural machine translation\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=H1z-PsR5KX)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p3.1)\.
- \[5\]D\. Bau, B\. Zhou, A\. Khosla, A\. Oliva, and A\. Torralba\(2017\)Network Dissection: Quantifying Interpretability of Deep Visual Representations\.In2017 IEEE Conference on Computer Vision and Pattern Recognition \(CVPR\),Vol\.,Los Alamitos, CA, USA,pp\. 3319–3327\.External Links:ISSN 1063\-6919,[Document](https://dx.doi.org/10.1109/CVPR.2017.354),[Link](https://doi.ieeecomputersociety.org/10.1109/CVPR.2017.354)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p3.1),[§II](https://arxiv.org/html/2608.17102#S2.p1.1)\.
- \[6\]D\. Bau, J\. Zhu, H\. Strobelt, A\. Lapedriza, B\. Zhou, and A\. Torralba\(2020\)Understanding the role of individual units in a deep neural network\.Proceedings of the National Academy of Sciences117\(48\),pp\. 30071–30078\.External Links:ISSN 1091\-6490,[Link](http://dx.doi.org/10.1073/pnas.1907375117),[Document](https://dx.doi.org/10.1073/pnas.1907375117)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p3.1),[§II](https://arxiv.org/html/2608.17102#S2.p1.1)\.
- \[7\]C\. Busso, Z\. Deng, S\. Yildirim, M\. Bulut, C\. M\. Lee, A\. Kazemzadeh, S\. Lee, U\. Neumann, and S\. Narayanan\(2004\)Analysis of emotion recognition using facial expressions, speech and multimodal information\.InProceedings of the 6th International Conference on Multimodal Interfaces,ICMI ’04,New York, NY, USA,pp\. 205–211\.External Links:ISBN 1581139950,[Link](https://doi.org/10.1145/1027933.1027968),[Document](https://dx.doi.org/10.1145/1027933.1027968)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p1.1)\.
- \[8\]C\. Busso, R\. Lotfian, K\. Sridhar, A\. N\. Salman, W\. Lin, L\. Goncalves, S\. Parthasarathy, A\. R\. Naini, S\. Leem, L\. Martinez\-Lucas, H\. Chou, and P\. Mote\(2026\)The MSP\-Podcast Corpus\.IEEE Transactions on Affective Computing1,pp\. 1–19\.External Links:ISSN 1949\-3045,[Document](https://dx.doi.org/10.1109/TAFFC.2026.3678489),[Link](https://doi.ieeecomputersociety.org/10.1109/TAFFC.2026.3678489)Cited by:[§IV\-A](https://arxiv.org/html/2608.17102#S4.SS1.p1.1)\.
- \[9\]U\. Cappellazzo, M\. Kim, H\. Chen, P\. Ma, S\. Petridis, D\. Falavigna, A\. Brutti, and M\. Pantic\(2025\)Large language models are strong audio\-visual speech recognition learners\.InICASSP 2025 \- 2025 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),Vol\.,pp\. 1–5\.External Links:[Document](https://dx.doi.org/10.1109/ICASSP49660.2025.10889251)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p2.1)\.
- \[10\]J\. Chen, Z\. Guo, J\. Chun, P\. Wang, A\. Perrault, and M\. Elsner\(2026\)Do audio LLMs really LISTEN, or just transcribe? measuring lexical vs\. acoustic emotion cues reliance\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 5848–5877\.External Links:[Link](https://aclanthology.org/2026.eacl-long.274/),[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.274),ISBN 979\-8\-89176\-380\-7Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p2.1)\.
- \[11\]Z\. Cheng, Z\. Cheng, J\. He, K\. Wang, Y\. Lin, Z\. Lian, X\. Peng, and A\. G\. Hauptmann\(2024\)Emotion\-LLaMA: multimodal emotion recognition and reasoning with instruction tuning\.InThe Thirty\-eighth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=qXZVSy9LFR)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p2.1)\.
- \[12\]F\. Dalvi, N\. Durrani, H\. Sajjad, Y\. Belinkov, A\. Bau, and J\. Glass\(2019\)What is one grain of sand in the desert? analyzing individual neurons in deep nlp models\.InProceedings of the Thirty\-Third AAAI Conference on Artificial Intelligence and Thirty\-First Innovative Applications of Artificial Intelligence Conference and Ninth AAAI Symposium on Educational Advances in Artificial Intelligence,AAAI’19/IAAI’19/EAAI’19\.External Links:ISBN 978\-1\-57735\-809\-1,[Link](https://doi.org/10.1609/aaai.v33i01.33016309),[Document](https://dx.doi.org/10.1609/aaai.v33i01.33016309)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p3.1)\.
- \[13\]S\. Deshmukh, B\. Elizalde, R\. Singh, and H\. Wang\(2023\)Pengi: an audio language model for audio tasks\.InAdvances in Neural Information Processing Systems,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),Vol\.36,pp\. 18090–18108\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/file/3a2e5889b4bbef997ddb13b55d5acf77-Paper-Conference.pdf)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p2.1)\.
- \[14\]Y\. Etesam, Ö\. N\. Yalçın, C\. Zhang, and A\. Lim\(2024\)Contextual emotion recognition using large vision language models\.In2024 IEEE/RSJ International Conference on Intelligent Robots and Systems \(IROS\),Vol\.,pp\. 4769–4776\.External Links:[Document](https://dx.doi.org/10.1109/IROS58592.2024.10802538)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p1.1)\.
- \[15\]J\. Fang, Z\. Bi, R\. Wang, H\. Jiang, Y\. Gao, K\. Wang, A\. Zhang, J\. Shi, X\. Wang, and T\. Chua\(2024\)Towards neuron attributions in multimodal large language models\.InProceedings of the 38th International Conference on Neural Information Processing Systems,NIPS ’24,Red Hook, NY, USA\.External Links:ISBN 9798331314385Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p3.1),[§II](https://arxiv.org/html/2608.17102#S2.p1.1)\.
- \[16\]N\. M\. Foteinopoulou and I\. Patras\(2024\)EmoCLIP: a vision\-language method for zero\-shot video facial expression recognition\.In2024 IEEE 18th International Conference on Automatic Face and Gesture Recognition \(FG\),Vol\.,pp\. 1–10\.External Links:[Document](https://dx.doi.org/10.1109/FG59268.2024.10581982)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p2.1)\.
- \[17\]Google\(2026\)Gemma\-4\-12b\-it\.Note:https://huggingface\.co/google/gemma\-4\-12B\-itCited by:[§I](https://arxiv.org/html/2608.17102#S1.p2.1),[§I](https://arxiv.org/html/2608.17102#S1.p4.1),[§IV\-A](https://arxiv.org/html/2608.17102#S4.SS1.p2.1)\.
- \[18\]W\. Gurnee, T\. Horsley, Z\. C\. Guo, T\. R\. Kheirkhah, Q\. Sun, W\. Hathaway, N\. Nanda, and D\. Bertsimas\(2024\)Universal neurons in GPT2 language models\.Transactions on Machine Learning Research\.Note:External Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=ZeI104QZ8I)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p3.1)\.
- \[19\]K\. Hacioglu, M\. K\. E, and A\. Stolcke\(2025\)SpeechLLMs for large\-scale contextualized zero\-shot slot filling\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track,S\. Potdar, L\. Rojas\-Barahona, and S\. Montella \(Eds\.\),Suzhou \(China\),pp\. 703–715\.External Links:[Link](https://aclanthology.org/2025.emnlp-industry.49/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-industry.49),ISBN 979\-8\-89176\-333\-3Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p2.1)\.
- \[20\]Y\. He, Z\. Liu, S\. Sun, B\. Wang, W\. Zhang, X\. Zou, N\. F\. Chen, and A\. T\. Aw\(2025\)MERaLiON\-audiollm: bridging audio and language with large language models\.External Links:2412\.09818,[Link](https://arxiv.org/abs/2412.09818)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p2.1)\.
- \[21\]C\. Huang, K\. Lu, S\. Wang, C\. Hsiao, C\. Kuan, H\. Wu, S\. Arora, K\. Chang, J\. Shi, Y\. Peng, R\. Sharma, S\. Watanabe, B\. Ramakrishnan, S\. Shehata, and H\. Lee\(2024\)Dynamic\-superb: towards a dynamic, collaborative, and comprehensive instruction\-tuning benchmark for speech\.InICASSP 2024 \- 2024 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),Vol\.,pp\. 12136–12140\.External Links:[Document](https://dx.doi.org/10.1109/ICASSP48485.2024.10448257)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p2.1)\.
- \[22\]K\. Huang, J\. Huo, Y\. Yan, K\. Wang, Y\. Yue, and X\. Hu\(2024\)MINER: mining the underlying pattern of modality\-specific neurons in multimodal large language models\.External Links:2410\.04819,[Link](https://arxiv.org/abs/2410.04819)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p3.1),[§II](https://arxiv.org/html/2608.17102#S2.p1.1)\.
- \[23\]R\. Huben, H\. Cunningham, L\. R\. Smith, A\. Ewart, and L\. Sharkey\(2024\)Sparse autoencoders find highly interpretable features in language models\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=F76bwRSLeK)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p3.1),[§II](https://arxiv.org/html/2608.17102#S2.p1.1)\.
- \[24\]J\. Huo, Y\. Yan, B\. Hu, Y\. Yue, and X\. Hu\(2024\)MMNeuron: discovering neuron\-level domain\-specific interpretation in multimodal large language model\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 6801–6816\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.387/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.387)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p3.1),[§II](https://arxiv.org/html/2608.17102#S2.p1.1)\.
- \[25\]K\. Konen, S\. Jentzsch, D\. Diallo, P\. Schütt, O\. Bensch, R\. El Baff, D\. Opitz, and T\. Hecking\(2024\)Style vectors for steering generative large language models\.InFindings of the Association for Computational Linguistics: EACL 2024,Y\. Graham and M\. Purver \(Eds\.\),St\. Julian’s, Malta,pp\. 782–802\.External Links:[Link](https://aclanthology.org/2024.findings-eacl.52/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-eacl.52)Cited by:[§II](https://arxiv.org/html/2608.17102#S2.p1.1)\.
- \[26\]P\. Kumar, V\. Kaushik, and B\. Raman\(2021\)Towards the Explainability of Multimodal Speech Emotion Recognition\.InInterspeech 2021,pp\. 1748–1752\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2021-1718),ISSN 2958\-1796Cited by:[§II](https://arxiv.org/html/2608.17102#S2.p1.1)\.
- \[27\]J\. Lee, W\. Lee, O\. Kwon, and H\. Kim\(2025\)Do large language models have “emotion neurons”? investigating the existence and role\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 15617–15639\.External Links:[Link](https://aclanthology.org/2025.findings-acl.806/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.806),ISBN 979\-8\-89176\-256\-5Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p3.1),[§II](https://arxiv.org/html/2608.17102#S2.p1.1)\.
- \[28\]P\. Li, B\. Zhao, Z\. Kang, J\. Peng, X\. Qu, Y\. He, and J\. Wang\(2025\)EMO\-RL: emotion\-rule\-based reinforcement learning enhanced audio\-language model for generalized speech emotion recognition\.InFindings of the Association for Computational Linguistics: EMNLP 2025,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 18744–18754\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.1018/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.1018),ISBN 979\-8\-89176\-335\-7Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p2.1)\.
- \[29\]H\. Lian, C\. Lu, S\. Li, Y\. Zhao, C\. Tang, and Y\. Zong\(2023\)A survey of deep learning\-based multimodal emotion recognition: speech, text, and face\.Entropy25\(10\)\.External Links:[Link](https://www.mdpi.com/1099-4300/25/10/1440),ISSN 1099\-4300,[Document](https://dx.doi.org/10.3390/e25101440)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p1.1)\.
- \[30\]A\. Mollahosseini, B\. Hasani, and M\. H\. Mahoor\(2019\)AffectNet: a database for facial expression, valence, and arousal computing in the wild\.IEEE Transactions on Affective Computing10\(1\),pp\. 18–31\.External Links:[Document](https://dx.doi.org/10.1109/TAFFC.2017.2740923)Cited by:[§IV\-A](https://arxiv.org/html/2608.17102#S4.SS1.p1.1)\.
- \[31\]J\. A\. Russell\(1991\)Culture and the categorization of emotions\.\.Psychological bulletin110\(3\),pp\. 426\.Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p1.1)\.
- \[32\]S\. Sakshi, U\. Tyagi, S\. Kumar, A\. Seth, R\. Selvakumar, O\. Nieto, R\. Duraiswami, S\. Ghosh, and D\. Manocha\(2025\)MMAU: a massive multi\-task audio understanding and reasoning benchmark\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=TeVAZXr3yv)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p2.1)\.
- \[33\]A\. Schirmer and R\. Adolphs\(2017\)Emotion perception from face, voice, and touch: comparisons and convergence\.Trends in cognitive sciences21\(3\),pp\. 216–228\.Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p1.1)\.
- \[34\]N\. Shazeer\(2020\)GLU variants improve transformer\.External Links:2002\.05202,[Link](https://arxiv.org/abs/2002.05202)Cited by:[§III\-A](https://arxiv.org/html/2608.17102#S3.SS1.p2.1)\.
- \[35\]P\. Singh, O\. De Clercq, and E\. Lefever\(2026\)Lost in activations: a neuron\-level analysis of encoders for cross\-lingual emotion detection\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 2: Short Papers\),V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 154–159\.External Links:[Link](https://aclanthology.org/2026.eacl-short.9/),[Document](https://dx.doi.org/10.18653/v1/2026.eacl-short.9),ISBN 979\-8\-89176\-381\-4Cited by:[§II](https://arxiv.org/html/2608.17102#S2.p1.1)\.
- \[36\]Y\. K\. Singla, J\. Shah, C\. Chen, and R\. R\. Shah\(2022\)What do audio transformers hear? probing their representations for language delivery & structure\.In2022 IEEE International Conference on Data Mining Workshops \(ICDMW\),pp\. 910–925\.Cited by:[§II](https://arxiv.org/html/2608.17102#S2.p2.1)\.
- \[37\]N\. Sofroniew, I\. Kauvar, W\. Saunders, R\. Chen, T\. Henighan, S\. Hydrie, C\. Citro, A\. Pearce, J\. Tarng, W\. Gurnee, J\. Batson, S\. Zimmerman, K\. Rivoire, K\. Fish, C\. Olah, and J\. Lindsey\(2026\)Emotion concepts and their function in a large language model\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2026/emotions/index.html)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p3.1),[§II](https://arxiv.org/html/2608.17102#S2.p1.1)\.
- \[38\]L\. Sun, B\. Liu, J\. Tao, and Z\. Lian\(2021\)Multimodal cross\- and self\-attention network for speech emotion recognition\.InICASSP 2021 \- 2021 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),Vol\.,pp\. 4275–4279\.External Links:[Document](https://dx.doi.org/10.1109/ICASSP39728.2021.9414654)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p1.1)\.
- \[39\]T\. Tang, W\. Luo, H\. Huang, D\. Zhang, X\. Wang, X\. Zhao, F\. Wei, and J\. Wen\(2024\)Language\-specific neurons: the key to multilingual capabilities in large language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 5701–5715\.External Links:[Link](https://aclanthology.org/2024.acl-long.309/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.309)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p3.1)\.
- \[40\]P\. Tzirakis, G\. Trigeorgis, M\. A\. Nicolaou, B\. W\. Schuller, and S\. Zafeiriou\(2017\)End\-to\-end multimodal emotion recognition using deep neural networks\.IEEE Journal of Selected Topics in Signal Processing11\(8\),pp\. 1301–1309\.External Links:[Document](https://dx.doi.org/10.1109/JSTSP.2017.2764438)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p1.1)\.
- \[41\]E\. Voita, J\. Ferrando, and C\. Nalmpantis\(2024\)Neurons in large language models: dead, n\-gram, positional\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 1288–1301\.External Links:[Link](https://aclanthology.org/2024.findings-acl.75/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.75)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p3.1),[§II](https://arxiv.org/html/2608.17102#S2.p1.1)\.
- \[42\]C\. Wang, G\. Deng, X\. Yang, H\. Qiu, and T\. Zhang\(2025\)When audio and text disagree: revealing text bias in large audio\-language models\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 4878–4888\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.246/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.246),ISBN 979\-8\-89176\-332\-6Cited by:[§V\-B](https://arxiv.org/html/2608.17102#S5.SS2.p2.1)\.
- \[43\]C\. Wang, S\. Dai, Y\. Wang, F\. Yang, M\. Qiu, K\. Chen, W\. Zhou, and J\. Huang\(2022\)ARoBERT: an asr robust pre\-trained language model for spoken language understanding\.IEEE/ACM Transactions on Audio, Speech, and Language Processing30\(\),pp\. 1207–1218\.External Links:[Document](https://dx.doi.org/10.1109/TASLP.2022.3153268)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p2.1)\.
- \[44\]C\. Wang, Y\. Zhang, R\. Yu, Y\. Zheng, L\. Gao, Z\. Song, Z\. Xu, G\. Xia, H\. Zhang, D\. Zhao, and X\. Chen\(2025\)Do llms ”feel”? emotion circuits discovery and control\.External Links:2510\.11328,[Link](https://arxiv.org/abs/2510.11328)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p3.1),[§II](https://arxiv.org/html/2608.17102#S2.p1.1)\.
- \[45\]J\. M\. Wilce\(2009\)Language and emotion\.Studies in the Social and Cultural Foundations of Language,Cambridge University Press\.Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p1.1)\.
- \[46\]P\. Wu, Z\. Ling, L\. Liu, Y\. Jiang, H\. Wu, and L\. Dai\(2019\)End\-to\-end emotional speech synthesis using style tokens and semi\-supervised training\.In2019 Asia\-Pacific Signal and Information Processing Association Annual Summit and Conference \(APSIPA ASC\),pp\. 623–627\.Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p3.1),[§II](https://arxiv.org/html/2608.17102#S2.p1.1)\.
- \[47\]T\. Wu, Y\. Lin, and T\. Weng\(2024\)AND: audio network dissection for interpreting deep acoustic models\.InProceedings of the 41st International Conference on Machine Learning,ICML’24\.Cited by:[§II](https://arxiv.org/html/2608.17102#S2.p2.1)\.
- \[48\]T\. Xie, S\. Yang, C\. Li, D\. Yu, and L\. Liu\(2025\)EmoSteer\-tts: fine\-grained and training\-free emotion\-controllable text\-to\-speech via activation steering\.External Links:2508\.03543,[Link](https://arxiv.org/abs/2508.03543)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p3.1),[§II](https://arxiv.org/html/2608.17102#S2.p1.1)\.
- \[49\]J\. Xu, C\. Lan, and Y\. Lu\(2025\)Deciphering functions of neurons in vision\-language models\.InProceedings of the 33rd ACM International Conference on Multimedia,pp\. 3173–3181\.Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p3.1),[§II](https://arxiv.org/html/2608.17102#S2.p1.1)\.
- \[50\]J\. Xu, Z\. Guo, J\. He, H\. Hu, T\. He, S\. Bai, K\. Chen, J\. Wang, Y\. Fan, K\. Dang, B\. Zhang, X\. Wang, Y\. Chu, and J\. Lin\(2025\)Qwen2\.5\-omni technical report\.External Links:2503\.20215,[Link](https://arxiv.org/abs/2503.20215)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p2.1),[§I](https://arxiv.org/html/2608.17102#S1.p4.1),[§IV\-A](https://arxiv.org/html/2608.17102#S4.SS1.p2.1)\.
- \[51\]Y\. Xu, H\. Chen, J\. Yu, Q\. Huang, Z\. Wu, S\. Zhang, G\. Li, Y\. Luo, and R\. Gu\(2024\)SECap: speech emotion captioning with large language model\.Proceedings of the AAAI Conference on Artificial Intelligence38\(17\),pp\. 19323–19331\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/29902),[Document](https://dx.doi.org/10.1609/aaai.v38i17.29902)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p2.1)\.
- \[52\]C\. Yang, N\. Ho, Y\. Lee, and H\. Lee\(2025\)AudioLens: a closer look at auditory attribute perception of large audio\-language models\.External Links:2506\.05140,[Link](https://arxiv.org/abs/2506.05140)Cited by:[§II](https://arxiv.org/html/2608.17102#S2.p2.1)\.
- \[53\]C\. Yang, N\. S\. Ho, and H\. Lee\(2025\)Towards holistic evaluation of large audio\-language models: a comprehensive survey\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 10144–10170\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.514/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.514),ISBN 979\-8\-89176\-332\-6Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p1.1)\.
- \[54\]Y\. Yao, T\. Yu, A\. Zhang, C\. Wang, J\. Cui, H\. Zhu, T\. Cai, H\. Li, W\. Zhao, Z\. He, Q\. Chen, H\. Zhou, Z\. Zou, H\. Zhang, S\. Hu, Z\. Zheng, J\. Zhou, J\. Cai, X\. Han, G\. Zeng, D\. Li, Z\. Liu, and M\. Sun\(2024\)MiniCPM\-v: a gpt\-4v level mllm on your phone\.External Links:2408\.01800,[Link](https://arxiv.org/abs/2408.01800)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p2.1),[§I](https://arxiv.org/html/2608.17102#S1.p4.1),[§IV\-A](https://arxiv.org/html/2608.17102#S4.SS1.p2.1)\.
- \[55\]Z\. Yu and S\. Ananiadou\(2024\)Neuron\-level knowledge attribution in large language models\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 3267–3280\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.191/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.191)Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p3.1),[§II](https://arxiv.org/html/2608.17102#S2.p1.1)\.
- \[56\]X\. Zhao, R\. Choenni, R\. Saxena, and I\. Titov\(2026\)Finding culture\-sensitive neurons in vision\-language models\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 3366–3381\.External Links:[Link](https://aclanthology.org/2026.eacl-long.155/),ISBN 979\-8\-89176\-380\-7Cited by:[§III\-B](https://arxiv.org/html/2608.17102#S3.SS2.p1.1)\.
- \[57\]X\. Zhao, P\. Koehn, B\. Schuller, and B\. Sisman\(2026\)Multilingual emotion neurons in large audio\-language models\.External Links:2608\.08772,[Link](https://arxiv.org/abs/2608.08772)Cited by:[§II](https://arxiv.org/html/2608.17102#S2.p2.1)\.
- \[58\]X\. Zhao, B\. Schuller, and B\. Sisman\(2026\)Discovering and causally validating emotion\-sensitive neurons in large audio\-language models\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),San Diego, California, United States,pp\. 15056–15071\.External Links:[Link](https://aclanthology.org/2026.acl-long.687/),[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.687),ISBN 979\-8\-89176\-390\-6Cited by:[§I](https://arxiv.org/html/2608.17102#S1.p4.1),[§I](https://arxiv.org/html/2608.17102#S1.p5.1),[§II](https://arxiv.org/html/2608.17102#S2.p2.1),[§V\-A](https://arxiv.org/html/2608.17102#S5.SS1.p3.1)\.
- \[59\]X\. Zhao, I\. R\. Ulgen, P\. Koehn, B\. Schuller, and B\. Sisman\(2026\)Neuron\-level emotion control in speech\-generative large audio\-language models\.External Links:2603\.17231,[Link](https://arxiv.org/abs/2603.17231)Cited by:[§II](https://arxiv.org/html/2608.17102#S2.p2.1)\.
- \[60\]X\. Zhao, K\. Wang, and W\. Peng\(2024\)Measuring the inconsistency of large language models in preferential ranking\.InProceedings of the 1st Workshop on Towards Knowledgeable Language Models \(KnowLLM 2024\),S\. Li, M\. Li, M\. J\. Zhang, E\. Choi, M\. Geva, P\. Hase, and H\. Ji \(Eds\.\),Bangkok, Thailand,pp\. 171–176\.External Links:[Link](https://aclanthology.org/2024.knowllm-1.14/),[Document](https://dx.doi.org/10.18653/v1/2024.knowllm-1.14)Cited by:[§IV\-A](https://arxiv.org/html/2608.17102#S4.SS1.p2.1)\.

## Appendix AReproducibility

### A\-ADatasets and Models

TABLE III:Dataset statistics showing utterance/images counts per emotion\.TABLE IV:Sources and licenses for the three evaluated MFMs\.

Similar Articles

Evaluating multimodal emotion recognition in proactive conversational agents: A user study

arXiv cs.AI

This paper presents a multimodal emotion recognition module for proactive conversational agents, using facial recognition and linguistic analysis. A user study with 20 participants reveals a 'poker face' effect where visual cues are unreliable, while linguistic analysis proves more accurate; the study also shows agents can elicit emotions through conversational adaptation.

Multilingual Emotion Neurons in Large Audio-Language Models

arXiv cs.CL

A first neuron-level interpretability study of how large audio-language models encode multilingual emotion, introducing Consistency-Regularized Fusion to identify Multilingual Emotion Neurons across 12 languages and showing cross-lingual transfer benefits.

Do Speech Emphasis Models Generalize across Languages and Emotions?

arXiv cs.CL

Introduces MMEE, a multilingual multi-emotion emphasis corpus of 10,000 utterances across 7 languages and 34 emotions, and benchmarks emphasis detection models under various transfer settings, finding that multilingual training improves robustness while monolingual models show limited zero-shot transfer.