Multilingual Emotion Neurons in Large Audio-Language Models

arXiv cs.CL Papers

Summary

A first neuron-level interpretability study of how large audio-language models encode multilingual emotion, introducing Consistency-Regularized Fusion to identify Multilingual Emotion Neurons across 12 languages and showing cross-lingual transfer benefits.

arXiv:2608.08772v1 Announce Type: new Abstract: Emotion is central to human communication, and its expression varies across languages. Large audio-language models (LALMs) achieve strong performance on multilingual speech tasks, yet it remains unclear whether they encode emotion through language-specific correlations or language-agnostic representations. We present the first neuron-level interpretability study of this question. We define Multilingual Emotion Neurons (MLENs) as functional units exhibiting stable emotional selectivity and aligned causal effects across languages, and introduce Consistency-Regularized Fusion (CR-Fusion) to identify them. Across four modern LALMs and 12 typologically diverse languages, emotion-sensitive neurons identified independently per language show minimal overlap, and additional monolingual identification data saturates quickly without isolating more transferable units, motivating identification from pooled cross-lingual evidence. Causal interventions demonstrate that MLENs identified by CR-Fusion provide more precise and transferable affective control than monolingual neuron sets in both zero-shot and low-resource settings. Leave-one-out ablations further reveal asymmetric transfer: individual identification languages, including low-resource ones, contribute non-redundant evidence, while several low-resource languages benefit most from the resulting cross-lingual transfer. Together, our findings provide the first causal, neuron-level account of how LALMs encode emotion across languages, and establish multilingual neuron identification as an effective mechanism for understanding cross-lingual affective behavior.
Original Article
View Cached Full Text

Cached at: 08/11/26, 08:09 AM

# Multilingual Emotion Neurons in Large Audio-Language Models
Source: [https://arxiv.org/html/2608.08772](https://arxiv.org/html/2608.08772)
Xiutian Zhao1,Philipp Koehn1,Björn Schuller2,Berrak Sisman1 1Center for Language and Speech Processing \(CLSP\), Johns Hopkins University, USA 2Group on Language, Audio & Music \(GLAM\), Imperial College London, UK

###### Abstract

Emotion is central to human communication, and its expression varies across languages\. Large audio\-language models \(LALMs\) achieve strong performance on multilingual speech tasks, yet it remains unclear whether they encode emotion through language\-specific correlations or language\-agnostic representations\. We present the first neuron\-level interpretability study of this question\. We define Multilingual Emotion Neurons \(MLENs\) as functional units exhibiting stable emotional selectivity and aligned causal effects across languages, and introduce Consistency\-Regularized Fusion \(CR\-Fusion\) to identify them\. Across four modern LALMs and 12 typologically diverse languages, emotion\-sensitive neurons identified independently per language show minimal overlap, and additional monolingual identification data saturates quickly without isolating more transferable units, motivating identification from pooled cross\-lingual evidence\. Causal interventions demonstrate that MLENs identified by CR\-Fusion provide more precise and transferable affective control than monolingual neuron sets in both zero\-shot and low\-resource settings\. Leave\-one\-out ablations further reveal asymmetric transfer: individual identification languages, including low\-resource ones, contribute non\-redundant evidence, while several low\-resource languages benefit most from the resulting cross\-lingual transfer\. Together, our findings provide the first causal, neuron\-level account of how LALMs encode emotion across languages, and establish multilingual neuron identification as an effective mechanism for understanding cross\-lingual affective behavior\.

Multilingual Emotion Neurons in Large Audio\-Language Models

## 1Introduction

Emotion is a core channel of human communication: beyond what we say, speech conveys how we feel through prosody, timing, intensity, and culturally shaped display rulesWilce \([2009](https://arxiv.org/html/2608.08772#bib.bib145)\); Lindquistet al\.\([2015](https://arxiv.org/html/2608.08772#bib.bib146)\)\. In a multilingual world, this affective layer must be robust to linguistic diversity: people routinely infer emotion from unfamiliar languages, yet recognition is modulated by language, culture, and learned acoustic conventionsRussell \([1991](https://arxiv.org/html/2608.08772#bib.bib148)\); Pellet al\.\([2009](https://arxiv.org/html/2608.08772#bib.bib147)\); Jacksonet al\.\([2019](https://arxiv.org/html/2608.08772#bib.bib149)\)\. This tension between pan\-cultural cues and language\-specific realizations makes multilingual emotion an ideal testbed for probing the generalizability of concept representations in modern foundation models\.

Decades of cross\-lingual SER research show that emotion cues partially generalize to unseen languages, yet transfer remains sensitive to language and corpus mismatch, and human and model transfer patterns diverge systematically\(Albornoz and Milone,[2017](https://arxiv.org/html/2608.08772#bib.bib143); Hanet al\.,[2025](https://arxiv.org/html/2608.08772#bib.bib138)\)\. It thus remains open to what extent current models rely on cross\-lingually shared affective representations versus dataset\-, corpus\-, and language\-local correlations\.

Parallel interpretability work shows that individual neurons in LLMs and multimodal models align with human\-interpretable concepts, including affective states\(Hubenet al\.,[2024](https://arxiv.org/html/2608.08772#bib.bib134); Bauet al\.,[2017](https://arxiv.org/html/2608.08772#bib.bib11);sofroniew2026emotionconceptsfunctionlarge\)\. However, emotion\-correlated neurons identified in one language often fail to generalize to others in multilingual encoders\(Singhet al\.,[2026](https://arxiv.org/html/2608.08772#bib.bib133)\), and cross\-lingual generalization of activation steering remains contested\(Maraiaet al\.,[2026](https://arxiv.org/html/2608.08772#bib.bib132)\)\. This tension between language\-specific and language\-agnostic emotion representations motivates our investigation \(§[2](https://arxiv.org/html/2608.08772#S2)\)\.

Modern multimodal foundation modelsyao2024minicpmvgpt4vlevelmllm;xu2025qwen25omnitechnicalreport, particularly LALMskimiteam2025kimiaudiotechnicalreport;goel2025audioflamingo3advancing, offer a new paradigm for probing this question\. By integrating vast acoustic and linguistic pre\-training, LALMs have achieved competitive performance on a broad range of spoken language understanding tasks\(Huanget al\.,[2024](https://arxiv.org/html/2608.08772#bib.bib158); Sakshiet al\.,[2025](https://arxiv.org/html/2608.08772#bib.bib157)\), including speech emotion recognition\(Chenget al\.,[2024](https://arxiv.org/html/2608.08772#bib.bib166);he2025meralionaudiollmbridgingaudiolanguage\)\. However, mechanistic accounts of how these models represent affective information remain lacking\. It is unclear whether LALMs utilize language\-specific functional units or possess a shared representational structure that generalizes emotion processing across broad linguistic diversity\.

To the best of our knowledge, this work represents the first systematic attempt to investigate and manipulate the intrinsic multilingual emotion representations within LALMs\. Using SER as a probe task, we identify affective neurons through activations across four open\-source LALMs, Audio\-Flamingo\-3goel2025audioflamingo3advancing, Qwen2\.5\-Omni\-7Bxu2025qwen25omnitechnicalreport, MiniCPM\-o\-4\.5yao2024minicpmvgpt4vlevelmllm, and Kimi\-Audiokimiteam2025kimiaudiotechnicalreport, and 12 typologically diverse languages, including high\-resource languages \(e\.g\., English, Mandarin\) and low\-resource languages such as Amharic, Bengali and Urdu, where emotional speech data is often sparseKoehn and Knowles \([2017](https://arxiv.org/html/2608.08772#bib.bib92)\); Mohmad Dar and Delhibabu \([2024](https://arxiv.org/html/2608.08772#bib.bib94)\)\. To isolate neurons with stable cross\-lingual emotional selectivity, we propose*Consistency\-Regularized Fusion*\(CR\-Fusion\) that operates atop any neuron selector \(e\.g\., Mean\-activation\-basedBauet al\.\([2019](https://arxiv.org/html/2608.08772#bib.bib16)\); Dalviet al\.\([2019](https://arxiv.org/html/2608.08772#bib.bib15)\)\) to leverage multilingual evidence for identifying cross\-lingually stable affective units\.

Our investigation yields five principal findings\. \(1\)Emotion\-sensitive neurons \(ESNs\) identified independently in different languages exhibit minimal set intersectionwith weak to moderate rank correlation in selectivity scores, revealing that monolingual identification procedures select substantially different neurons even when underlying ranking structures are similar\. \(2\)Monolingual evidence saturates rapidlybeyond 50 instances; additional within\-language data yields no further improvement in isolating transferable neurons\. \(3\) We define Multilingual Emotion Neurons \(MLENs\) as functional units exhibiting stable emotional selectivity across languages\. Causal interventions confirm thatMLENs play a functional role in emotion processing: deactivation selectively impairs recognition of target emotions, while steering selectively enhances it\. Our proposed CR\-Fusionmatches or outperforms the best monolingual identificationin zero\-shot and low\-resource settings across models, with the exception of steering on Qwen2\.5\-Omni\-7B, the model with the lowest language\-invariant share\. \(4\) Causal effects are heterogeneous across emotions: anger, happiness and sadness show the largest and most stable intervention effects, while fear and neutral are weaker and more variable\. This ordering tracks baseline recognition accuracy and class support rather than representational universality—per\-emotion invariant shares are in fact highest for neutral and fear\. \(5\) We uncoverasymmetric transfer patterns, with individual identification languages contributing non\-substitutable evidence\. These results demonstrate thatlow\-resource languages can contribute non\-redundant evidence for multilingual neuron identification while simultaneously benefiting from improved cross\-lingual transfer\.

These findings provide a mechanistic framework for understanding how multimodal foundation models process emotion across languages, with practical implications for training\-free cross\-lingual generalization and equitable affective computing for low\-resource language communities\.

## 2Related Work

#### Multilingual Speech Emotion Recognition\.

Research on multilingual SER has progressed from early cross\-lingual transfer using traditional deep neural networksAlbornoz and Milone \([2017](https://arxiv.org/html/2608.08772#bib.bib143)\); Neumann and Thang Vu \([2018](https://arxiv.org/html/2608.08772#bib.bib139)\); Zehraet al\.\([2021](https://arxiv.org/html/2608.08772#bib.bib144)\)to more scalable approaches leveraging self\-supervised encoders such as wav2vec 2\.0 and transformer\-based architecturesSharma \([2022](https://arxiv.org/html/2608.08772#bib.bib141)\); Al\-onaziet al\.\([2022](https://arxiv.org/html/2608.08772#bib.bib137)\)\. Recent benchmarks like EmoBoxMaet al\.\([2024](https://arxiv.org/html/2608.08772#bib.bib135)\)have standardized multilingual evaluation while highlighting persistent cross\-lingual robustness challenges, with studies revealing systematic divergences between human and model transfer patternssingh2023decodingemotionscomprehensivemultilingual; Hanet al\.\([2025](https://arxiv.org/html/2608.08772#bib.bib138)\)\. Current work explores LALMs for zero\-shot cross\-lingual recognition, emotion captioning, and joint audio\-text reasoningZouet al\.\([2025](https://arxiv.org/html/2608.08772#bib.bib140)\); Xuet al\.\([2024](https://arxiv.org/html/2608.08772#bib.bib142)\);he2025meralionaudiollmbridgingaudiolanguage, complemented by parameter\-efficient methods like LoRA for cross\-lingual alignmentGoncalveset al\.\([2024](https://arxiv.org/html/2608.08772#bib.bib136)\)\. However, there is no mechanistic study of emotion representation across languages in LALMsshou2025multimodallargelanguagemodels\.

#### Neuron\-Level Concept and Emotion Interpretability in Audio\-Language Models\.

Identifying functional units that respond selectively to human\-interpretable concepts is a long\-standing theme in interpretability, established in visionBauet al\.\([2017](https://arxiv.org/html/2608.08772#bib.bib11),[2020](https://arxiv.org/html/2608.08772#bib.bib12)\)and LLMsHubenet al\.\([2024](https://arxiv.org/html/2608.08772#bib.bib134)\); Voitaet al\.\([2024](https://arxiv.org/html/2608.08772#bib.bib17)\); Yu and Ananiadou \([2024](https://arxiv.org/html/2608.08772#bib.bib23)\); Tanget al\.\([2024](https://arxiv.org/html/2608.08772#bib.bib13)\)\. For affect specifically, recent work identifies clustered emotion neurons in LLMsLeeet al\.\([2025](https://arxiv.org/html/2608.08772#bib.bib26)\)that are causally controllable via activation steeringwang2025llmsfeelemotioncircuits, withsofroniew2026emotionconceptsfunctionlargeshowing how affective states decompose into modular circuits; yetSinghet al\.\([2026](https://arxiv.org/html/2608.08772#bib.bib133)\)find that emotion\-correlated neurons identified in one language fail to generalize in multilingual encoders, and cross\-lingual generalization of steering remains contestedMaraiaet al\.\([2026](https://arxiv.org/html/2608.08772#bib.bib132)\)\. In LALMs, neuron\-level analysis has so far targeted modality attribution and generic acoustic conceptsHuoet al\.\([2024](https://arxiv.org/html/2608.08772#bib.bib20)\); Wuet al\.\([2024](https://arxiv.org/html/2608.08772#bib.bib9)\), or probed phonetic and prosodic encoding layer\-wiseyang2025audiolenscloserlookauditory; affective units have been identified only monolingually, for SERZhaoet al\.\([2026b](https://arxiv.org/html/2608.08772#bib.bib167)\)and emotional voice conversionzhao2026neuronlevelemotioncontrolspeechgenerative\. Their stability across linguistic boundaries has not been explored, which motivates our investigation \(extended discussion in Appendix[C](https://arxiv.org/html/2608.08772#A3)\)\.

## 3Method

Our pipeline comprises three stages: \(1\) activation logging on correctly recognized SER instances, \(2\) neuron scoring and selection using principled fusion strategies that integrate language\-conditioned statistics into cross\-lingual neuron sets, and \(3\) causal analysis through targeted intervention\.

### 3\.1Multilingual Activation Logging

Letℒ\\mathcal\{L\}denote a set of languages andℰ\\mathcal\{E\}the set of emotions\. We run the unintervened model on an SER task of emotion\-labeled speech and restrict logging to correctly predicted items to reduce contamination from failure modes and to obtain cleaner emotion\-conditioned statistics\. Let𝒟\(ℓ,e\)\\mathcal\{D\}^\{\(\\ell,e\)\}denote the set of correctly predicted utterances of emotione∈ℰe\\in\\mathcal\{E\}in languageℓ∈ℒ\\ell\\in\\mathcal\{L\}\. For an utterancexxwith token positionst∈\{1,…,Tx\}t\\in\\\{1,\\dots,T\_\{x\}\\\}, letal,n,t​\(x\)a\_\{l,n,t\}\(x\)denote the scalar SwiGLU gate activationshazeer2020gluvariantsimprovetransformerof neuronnnin layerll, and letwx,t∈\{0,1\}w\_\{x,t\}\\in\\\{0,1\\\}mask out irrelevant positions \(e\.g\., padding and instruction\-prompt tokens\)\. We aggregate positive\-activation counts and valid token counts as

Kl,n\(ℓ,e\)=∑x∈𝒟\(ℓ,e\)∑t=1Txwx,t​𝕀​\[al,n,t​\(x\)\>0\],\\displaystyle K^\{\(\\ell,e\)\}\_\{l,n\}=\\\!\\\!\\sum\_\{x\\in\\mathcal\{D\}^\{\(\\ell,e\)\}\}\\sum\_\{t=1\}^\{T\_\{x\}\}w\_\{x,t\}\\,\\mathbb\{I\}\\\!\\left\[a\_\{l,n,t\}\(x\)\>0\\right\],\(1\)T\(ℓ,e\)=∑x∈𝒟\(ℓ,e\)∑t=1Txwx,t,\\displaystyle T^\{\(\\ell,e\)\}=\\\!\\\!\\sum\_\{x\\in\\mathcal\{D\}^\{\(\\ell,e\)\}\}\\sum\_\{t=1\}^\{T\_\{x\}\}w\_\{x,t\},and form the activation\-probability profilePl,n\(ℓ,e\)=Kl,n\(ℓ,e\)/T\(ℓ,e\)P^\{\(\\ell,e\)\}\_\{l,n\}=K^\{\(\\ell,e\)\}\_\{l,n\}\\big/T^\{\(\\ell,e\)\}\.

### 3\.2Monolingual Evidence and Neuron Identification Methods

We first obtain monolingual evidence by scoring neurons independently per language\. Letsl,n\(ℓ,e\)s^\{\(\\ell,e\)\}\_\{l,n\}denote a selector score derived fromPl,n\(ℓ,e\)P^\{\(\\ell,e\)\}\_\{l,n\}\. We use a margin\-based selector,Contrastive Activation Margin \(ConAct\)Zhaoet al\.\([2026a](https://arxiv.org/html/2608.08772#bib.bib25)\), as our default due to its consistently stronger intervention effectiveness in comparison with LAPHubenet al\.\([2024](https://arxiv.org/html/2608.08772#bib.bib134)\); Gurneeet al\.\([2024](https://arxiv.org/html/2608.08772#bib.bib19)\), LAPETanget al\.\([2024](https://arxiv.org/html/2608.08772#bib.bib13)\); Namazifard and Poech \([2025](https://arxiv.org/html/2608.08772#bib.bib29)\), and MADBauet al\.\([2019](https://arxiv.org/html/2608.08772#bib.bib16)\); Dalviet al\.\([2019](https://arxiv.org/html/2608.08772#bib.bib15)\)\. ConAct assigns each neuron to its most preferred emotion and measures a margin against the second most preferred\. UsingPl,n\(ℓ,e\)P^\{\(\\ell,e\)\}\_\{l,n\}, we define:Pl,n\(1\)​\(ℓ\)=maxe∈ℰ⁡Pl,n\(ℓ,e\)P^\{\(1\)\}\_\{l,n\}\(\\ell\)=\\max\_\{e\\in\\mathcal\{E\}\}P^\{\(\\ell,e\)\}\_\{l,n\}

el,n\(1\)​\(ℓ\)=arg⁡maxe⁡Pl,n\(ℓ,e\),Pl,n\(2\)​\(ℓ\)=maxe≠el,n\(1\)​\(ℓ\)⁡Pl,n\(ℓ,e\)e^\{\(1\)\}\_\{l,n\}\(\\ell\)=\\arg\\max\_\{e\}P^\{\(\\ell,e\)\}\_\{l,n\},P^\{\(2\)\}\_\{l,n\}\(\\ell\)=\\max\_\{e\\neq e^\{\(1\)\}\_\{l,n\}\(\\ell\)\}P^\{\(\\ell,e\)\}\_\{l,n\}sl,n\(ℓ,e\)=\{Pl,n\(1\)​\(ℓ\)−Pl,n\(2\)​\(ℓ\),e=el,n\(1\)​\(ℓ\)0,otherwise\.s^\{\(\\ell,e\)\}\_\{l,n\}=\\begin\{cases\}P^\{\(1\)\}\_\{l,n\}\(\\ell\)\-P^\{\(2\)\}\_\{l,n\}\(\\ell\),&e=e^\{\(1\)\}\_\{l,n\}\(\\ell\)\\\\ 0,&\\text\{otherwise\}\\end\{cases\}\.
For each\(ℓ,e\)\(\\ell,e\), we rank neurons by the emotion\-conditioned scoresl,n\(ℓ,e\)s^\{\(\\ell,e\)\}\_\{l,n\}and select a fixed fractionr=0\.5%r\{=\}0\.5\\%to obtain monolingual ESN setsℐl\(ℓ,e\)\\mathcal\{I\}\_\{l\}^\{\(\\ell,e\)\}\.

### 3\.3Fusion Strategies for Multilingual Emotion Neurons

We now construct MLENs by fusing monolingual evidence\{sl,n\(ℓ,e\)\}ℓ∈ℒ\\\{s^\{\(\\ell,e\)\}\_\{l,n\}\\\}\_\{\\ell\\in\\mathcal\{L\}\}\. The goal is to identify neurons that are not only emotion\-selective, but also consistent across languages\. Our designs are inspired by multi\-view learning and ensembling principles \(e\.g\., intersection/consensus, averaging, and variance\-regularized objectives\), adapted here to neuron selection\.

#### Overlap Fusion\.

This conservative strategy retains neurons supported by*all*languages\. We form a fixed\-budget MLEN set by re\-ranking candidates with a joint score that prefers agreement \(i\.e\. by its weakest per\-language evidence\),sl,noverlap,\(e\)=minℓ∈L⁡s~l,n\(ℓ,e\),s^\{\\mathrm\{overlap\},\(e\)\}\_\{l,n\}\\;=\\;\\min\_\{\\ell\\in L\}\\,\\tilde\{s\}^\{\(\\ell,e\)\}\_\{l,n\},wheres~\\tilde\{s\}denotes per\-language normalized scores \(z\-score or robust z\-score across neurons\)\. We then select the top\-rrneurons bysoverlaps^\{\\text\{overlap\}\}subject to a same\-emotion constraint \(the preferred emotion agrees across languages\)\.

#### Joint Fusion\.

This strategy treats all languages as a single unified dataset by pooling raw activation statistics before computing selector scores\. We first merge the activation counts and token counts across all languages:Kl,n\(joint,e\)=∑ℓ∈ℒKl,n\(ℓ,e\),T\(joint,e\)=∑ℓ∈ℒT\(ℓ,e\)\.K^\{\(\\text\{joint\},e\)\}\_\{l,n\}=\\sum\_\{\\ell\\in\\mathcal\{L\}\}K^\{\(\\ell,e\)\}\_\{l,n\},\\quad T^\{\(\\text\{joint\},e\)\}=\\sum\_\{\\ell\\in\\mathcal\{L\}\}T^\{\(\\ell,e\)\}\.Then, we compute the pooled activation probability:Pl,n\(joint,e\)=Kl,n\(joint,e\)T\(joint,e\)P^\{\(\\text\{joint\},e\)\}\_\{l,n\}=\\frac\{K^\{\(\\text\{joint\},e\)\}\_\{l,n\}\}\{T^\{\(\\text\{joint\},e\)\}\}\. We apply the selector directly to these pooled probabilities to obtain joint scoressl,n\(joint,e\)s^\{\(\\text\{joint\},e\)\}\_\{l,n\}, and select the top\-rrneurons\. This approach is data\-efficient and treats cross\-lingual evidence as natural augmentation, but may be dominated by high\-resource languages if case counts are imbalanced acrossℒ\\mathcal\{L\}\.

#### Consistency\-Regularized Fusion \(CR\-Fusion\)\.

To explicitly prefer cross\-lingual stability, we penalize dispersion of scores across languages\.s~l,n\(ℓ,e\)=\(sl,n\(ℓ,e\)−μ\(ℓ\)\)/\(σ\(ℓ\)\+ϵ0\)\\tilde\{s\}^\{\(\\ell,e\)\}\_\{l,n\}=\\bigl\(s^\{\(\\ell,e\)\}\_\{l,n\}\-\\mu^\{\(\\ell\)\}\\bigr\)/\\bigl\(\\sigma^\{\(\\ell\)\}\+\\epsilon\_\{0\}\\bigr\), whereμ\(ℓ\)\\mu^\{\(\\ell\)\}andσ\(ℓ\)\\sigma^\{\(\\ell\)\}are the mean and standard deviation computed over all \(layer, neuron, emotion\) triplets\.

μl,nCR,\(e\)\\displaystyle\\mu^\{\\mathrm\{CR\},\(e\)\}\_\{l,n\}=1\|ℒ\|​∑ℓ∈ℒs~l,n\(ℓ,e\),\\displaystyle=\\frac\{1\}\{\|\\mathcal\{L\}\|\}\\sum\_\{\\ell\\in\\mathcal\{L\}\}\\tilde\{s\}^\{\(\\ell,e\)\}\_\{l,n\},σl,nCR,\(e\)\\displaystyle\\sigma^\{\\mathrm\{CR\},\(e\)\}\_\{l,n\}=1\|ℒ\|​∑ℓ∈ℒ\(s~l,n\(ℓ,e\)−μl,nCR,\(e\)\)2\.\\displaystyle=\\sqrt\{\\frac\{1\}\{\|\\mathcal\{L\}\|\}\\sum\_\{\\ell\\in\\mathcal\{L\}\}\\big\(\\tilde\{s\}^\{\(\\ell,e\)\}\_\{l,n\}\-\\mu^\{\\mathrm\{CR\},\(e\)\}\_\{l,n\}\\big\)^\{2\}\}\.
We then score neurons by:sl,nCR,\(e\)=μl,nCR,\(e\)−λ⋅σl,nCR,\(e\),s^\{\\mathrm\{CR\},\(e\)\}\_\{l,n\}=\\mu^\{\\mathrm\{CR\},\(e\)\}\_\{l,n\}\-\\lambda\\cdot\\sigma^\{\\mathrm\{CR\},\(e\)\}\_\{l,n\},whereλ≥0\\lambda\{\\geq\}0acts as a regularization parameter that penalizes cross\-lingual variance\. Higher values ofλ\\lambdaprioritize neurons with high stability over those with high absolute monolingual scores, thereby filtering out language\-specific units\. Sensitivity toλ\\lambdais discussed in §[6\.2](https://arxiv.org/html/2608.08772#S6.SS2)\.

### 3\.4The Estimand of Cross\-Lingual Fusion

#### Fusion Strategies as a Robustness Family\.

Write the per\-language selector score for a fixed neuron ass\(ℓ,e\)=θ​\(e\)\+δ\(ℓ,e\)\+ξ\(ℓ,e\)s^\{\(\\ell,e\)\}=\\theta\(e\)\+\\delta^\{\(\\ell,e\)\}\+\\xi^\{\(\\ell,e\)\}, whereθ\(e\)\\theta^\{\(e\)\}is a language\-invariant emotion effect,δ\(ℓ,e\)\\delta^\{\(\\ell,e\)\}is genuine language\-specific deviation with𝔼ℓ​\[δ\(ℓ,e\)\]=0\\mathbb\{E\}\_\{\\ell\}\[\\delta^\{\(\\ell,e\)\}\]=0and varianceτ2\\tau^\{2\}, andξ\(ℓ,e\)\\xi^\{\(\\ell,e\)\}is finite\-sample noise\. Under a Gaussian approximation, the CR\-Fusion objectiveμCR−λ​σCR\\mu^\{\\mathrm\{CR\}\}\-\\lambda\\,\\sigma^\{\\mathrm\{CR\}\}is the\(1−q\)\(1\{\-\}q\)lower confidence bound of the neuron’s selectivity in a randomly drawn language, withλ=Φ−1​\(1−q\)\\lambda=\\Phi^\{\-1\}\(1\-q\)\. The three fusion strategies of §[3\.3](https://arxiv.org/html/2608.08772#S3.SS3)are therefore one family indexed by robustness: mean fusion \(λ=0\\lambda\{=\}0\), of which Joint Fusion is the token\-weighted counterpart, targets expected transfer to a new language, CR\-Fusion targets quantile transfer, and Overlap Fusion targets worst\-case transfer\.

#### The Predicted Optimal Penalty\.

Our evaluation averages ESS over held\-out languages, which constitutes an expected\-transfer objective\. If languages are exchangeable, the Bayes\-optimal selection criterion for this objective isθ\\thetaitself, whose minimum\-variance unbiased estimator is the cross\-language mean; the predicted optimum is thereforeλ\\lambdanear zero\. Moreover, the observed dispersion estimatesτ2\+v¯\\tau^\{2\}\+\\bar\{v\}, wherev¯\\bar\{v\}is within\-language sampling variance: sinceP^\\hat\{P\}is a binomial proportion,v≈P​\(1−P\)/T\(ℓ,e\)v\\approx P\(1\-P\)/T^\{\(\\ell,e\)\}, andTTvaries by more than an order of magnitude across our conditions \(Appendix[A\.5](https://arxiv.org/html/2608.08772#A1.SS5)\)\. A large penalty potentially removes neurons whose statistics are estimated from sparse data rather than neurons that are meaningfully language\-specific\. Moreover, because selection at a fixed top\-rrfraction operates in the upper tail, where score magnitude and dispersion are positively correlated, a large penalty drives the selected set toward uniformly weak units\. Both predictions are confirmed in §[6\.2](https://arxiv.org/html/2608.08772#S6.SS2)\.

#### The Recovered Estimand\.

Decompose the activation\-probability tensor additively over languages and emotions,

Pl,n\(ℓ,e\)=ml,n\+ul,n\(ℓ\)\+bl,n\(e\)\+γl,n\(ℓ,e\)\+εl,n\(ℓ,e\),P^\{\(\\ell,e\)\}\_\{l,n\}\\;=\\;m\_\{l,n\}\+u^\{\(\\ell\)\}\_\{l,n\}\+b^\{\(e\)\}\_\{l,n\}\+\\gamma^\{\(\\ell,e\)\}\_\{l,n\}\+\\varepsilon^\{\(\\ell,e\)\}\_\{l,n\},\(2\)wheremmis the neuron’s baseline activity,u\(ℓ\)u^\{\(\\ell\)\}a language main effect \(overall firing shift\),b\(e\)b^\{\(e\)\}the language\-invariant emotion effect, andγ\(ℓ,e\)\\gamma^\{\(\\ell,e\)\}the language×\\timesemotion interaction\. For a complete design with uniform weights, the least\-squares fit satisfiesm^\+b^\(e\)=1\|ℒ\|​∑ℓP\(ℓ,e\)\\hat\{m\}\+\\hat\{b\}^\{\(e\)\}=\\frac\{1\}\{\|\\mathcal\{L\}\|\}\\sum\_\{\\ell\}P^\{\(\\ell,e\)\}, the unweighted cross\-language mean profile\. Joint Fusion scores the token\-weighted pooled profile∑ℓK\(ℓ,e\)/∑ℓT\(ℓ,e\)\\sum\_\{\\ell\}K^\{\(\\ell,e\)\}/\\sum\_\{\\ell\}T^\{\(\\ell,e\)\}, and small\-λ\\lambdaCR\-Fusion approximates it\.

Empirically, applying ConAct tom\+b^m\+\\hat\{b\}selects neuron sets with mean Jaccard similarity0\.980\.98against Joint Fusion on the two models with complete language–emotion coverage, dropping to0\.410\.41–0\.630\.63on the two models with missing cells, where the weighted fit and naive pooling legitimately diverge \(Appendix[B](https://arxiv.org/html/2608.08772#A2)\)\. Cross\-lingual fusion thus succeeds not by consensus filtering but by estimating a specific, identifiable estimand: the invariant emotion effect, which no single\-corpus procedure can isolate\.

#### Magnitude of the Invariant Share\.

Noise\-corrected variance components of Eq\. \([2](https://arxiv.org/html/2608.08772#S3.E2)\) yield invariant shares between0\.300\.30\(Qwen2\.5\-Omni\-7B\) and0\.590\.59\(MiniCPM\-o\-4\.5\) among top\-ranked emotion\-selective neurons \(Appendix[B](https://arxiv.org/html/2608.08772#A2)\)\. In no model does the invariant share exceed0\.60\.6: much of the emotion\-conditioned activation structure, in two models the majority, is specific to individual languages or corpora \(see Limitations\)\.

### 3\.5Intervention

Given selected neuron indicesℐl\(e\)\\mathcal\{I\}\_\{l\}^\{\(e\)\}for emotioneeat layerll\(monolingual or fused\), we evaluate causal effects via two interventions: deactivation and steering\. Our interventions operate on the post\-activation SwiGLU gate outputs in decoder MLP modules\. Letgl,t∈ℝDlg\_\{l,t\}\\in\\mathbb\{R\}^\{D\_\{l\}\}denote the activated gate vector at layerlland token positiontt, computed asgl,t=act​\(gate\_proj​\(xl,t\)\)g\_\{l,t\}=\\text\{act\}\(\\text\{gate\\\_proj\}\(x\_\{l,t\}\)\), whereact​\(⋅\)\\text\{act\}\(\\cdot\)is the activation function\.

#### Deactivation\.

We zero selected neurons’ activations through element\-wise masking:

d​el,n\(e\)=\{0,n∈ℐl\(e\)1,otherwise,g~l,tdeact=gl,t⊙d​el\(e\)\.de^\{\(e\)\}\_\{l,n\}=\\begin\{cases\}0,&n\\in\\mathcal\{I\}^\{\(e\)\}\_\{l\}\\\\ 1,&\\text\{otherwise\}\\end\{cases\},\\qquad\\tilde\{g\}^\{\\mathrm\{deact\}\}\_\{l,t\}\{=\}g\_\{l,t\}\\odot de^\{\(e\)\}\_\{l\}\.

#### Steering\.

We amplify selected neurons by applying a multiplicative gain factorα≥0\\alpha\\geq 0:

s​tl,n\(e\)​\(α\)=\{1\+α,n∈ℐl\(e\)1,otherwise,g~l,tsteer=gl,t⊙s​tl\(e\)​\(α\)\.st^\{\(e\)\}\_\{l,n\}\(\\alpha\)=\\begin\{cases\}1\{\+\}\\alpha,n\\in\\mathcal\{I\}^\{\(e\)\}\_\{l\}\\\\ 1,\\text\{otherwise\}\\end\{cases\},\\tilde\{g\}^\{\\mathrm\{steer\}\}\_\{l,t\}\{=\}g\_\{l,t\}\\odot st^\{\(e\)\}\_\{l\}\(\\alpha\)\.
The modified gate activationsg~l,t\\tilde\{g\}\_\{l,t\}are then used in place ofgl,tg\_\{l,t\}\. We apply these interventions systematically across evaluated languages to assess whether MLENs induce aligned causal changes, thereby validating their cross\-lingual generalizability and functional specificity\.

## 4Experiment Setup

#### Datasets and Partition\.

Our test bed comprises 12 monolingual emotional speech datasets categorized by their role in the transferability pipeline\. TheIdentification Set \(ℒid\\mathcal\{L\}\_\{\\text\{id\}\}\)consists of languages used during activation logging to ensure cross\-lingual consistency: Amharic \(ASEDRettaet al\.\([2023](https://arxiv.org/html/2608.08772#bib.bib79)\)\), Moroccan Arabic \(MDERSoumiaa \([2024](https://arxiv.org/html/2608.08772#bib.bib88)\)\), Bengali \(SUBESCOSultanaet al\.\([2021](https://arxiv.org/html/2608.08772#bib.bib80)\)\), English \(MSP\-PodcastBussoet al\.\([2026](https://arxiv.org/html/2608.08772#bib.bib3)\)\), Italian \(EmozionalmenteCataniaet al\.\([2025](https://arxiv.org/html/2608.08772#bib.bib83)\)\), Mandarin \(BIICUpadhyayet al\.\([2023](https://arxiv.org/html/2608.08772#bib.bib82)\)\), Polish \(nEMOChristop \([2024](https://arxiv.org/html/2608.08772#bib.bib84)\)\), and Urdu \(UrduSERAkhtaret al\.\([2025](https://arxiv.org/html/2608.08772#bib.bib81)\)\)\. TheHeld\-out Set \(ℒheld\\mathcal\{L\}\_\{\\text\{held\}\}\)contains languages reserved strictly for zero\-shot evaluation: French \(CaFEGournayet al\.\([2018](https://arxiv.org/html/2608.08772#bib.bib91)\)\), German \(EmoDBBurkhardtet al\.\([2025](https://arxiv.org/html/2608.08772#bib.bib90)\)\), Persian \(ShEMOMohamad Nezamiet al\.\([2019](https://arxiv.org/html/2608.08772#bib.bib85)\)\) and Russian \(RESDAmenteset al\.\([2023](https://arxiv.org/html/2608.08772#bib.bib89)\)\)\. Notably, Amharic, Bengali, Urdu, and Moroccan Arabic are treated as low\-resource languages within this frameworkKoehn and Knowles \([2017](https://arxiv.org/html/2608.08772#bib.bib92)\)\. The dataset statistics are provided in Appendix[A](https://arxiv.org/html/2608.08772#A1)Table[3](https://arxiv.org/html/2608.08772#A1.T3)\. Evaluation partitions sample up to 150 utterances per emotion \(Appendix[A\.2](https://arxiv.org/html/2608.08772#A1.SS2)\)\.

#### Models, Prompting and Decoding\.

We evaluate four open\-sourced LALMs: Audio\-Flamingo\-3goel2025audioflamingo3advancing, Kimi\-audiokimiteam2025kimiaudiotechnicalreport, MiniCPM\-o\-4\.5yao2024minicpmvgpt4vlevelmllm, and Qwen2\.5\-Omni\-7Bxu2025qwen25omnitechnicalreport\. All selected models designedly support English and Mandarin, with varying, officially unclaimed capabilities in other languages\. All models are evaluated in a controlled multiple\-choice format using a single instruction template \(Appendix[A\.3](https://arxiv.org/html/2608.08772#A1.SS3)\)\. All inference uses deterministic decoding \(greedy, temperature=0\); reported variance therefore reflects cross\-language variability \(Appendix[A\.4](https://arxiv.org/html/2608.08772#A1.SS4)\)\.

#### Metrics\.

We assess baseline performances using unweighted average recall \(UAR\) across five emotions: anger, fear, happiness, neutral, and sadness\. To quantify causal effects, we defineSelf\-Effects \(SE\)as the accuracy change on a specific emotion subset when applying the matched neuron mask, relative to the unintervened baseline:SE​\(e\)=Acce→e−Accunintervened​\(e\)\\text\{SE\}\(e\)=\\mathrm\{Acc\}\_\{e\\to e\}\-\\mathrm\{Acc\}\_\{\\text\{unintervened\}\}\(e\)\.Average Cross\-Effects \(ACE\)measures, for each target emotionee, the mean accuracy change oneeunder masks matched to other emotions:ACE​\(e\)=1\|ℰ\|−1​∑e′≠e\(Acce′→e−Accunintervened​\(e\)\)\\mathrm\{ACE\}\(e\)=\\frac\{1\}\{\|\\mathcal\{E\}\|\-1\}\\sum\_\{e^\{\\prime\}\\neq e\}\\left\(\\mathrm\{Acc\}\_\{e^\{\\prime\}\\to e\}\-\\mathrm\{Acc\}\_\{\\text\{unintervened\}\}\(e\)\\right\)\.Emotion Selectivity Score \(ESS\)quantifies specificity by comparing self\-effect against cross\-effect\. The per\-emotion ESS is defined as:ESS​\(e\)=SE​\(e\)−ACE​\(e\)\.\\text\{ESS\}\(e\)=\\text\{SE\}\(e\)\-\\text\{ACE\}\(e\)\.The global ESS, which aggregates across all valid emotions, is computed as:ESS=1\|ℰvalid\|​∑e∈ℰvalidESS​\(e\),\\text\{ESS\}=\\frac\{1\}\{\|\\mathcal\{E\}\_\{\\text\{valid\}\}\|\}\\sum\_\{e\\in\\mathcal\{E\}\_\{\\text\{valid\}\}\}\\text\{ESS\}\(e\),whereℰvalid\\mathcal\{E\}\_\{\\text\{valid\}\}denotes emotions with at least one test instance\. Unless otherwise stated, all reported ESS values refer to this global metric\. Negative ESS under deactivation indicates selective impairment of target emotions while preserving recognition of other emotions; positive ESS under steering indicates selective enhancement of target emotions\.

## 5Results

### 5\.1Cross\-Lingual Agreement of Monolingual Emotion\-Sensitive Neurons

![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/descriptive/Combined_Descriptive_audio-flamingo-3-hf_CAM_50_top0.005_pairwise_additive.png)\(a\)Audio\-Flamingo\-3
![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/descriptive/Combined_Descriptive_Kimi-Audio-7B-Instruct_CAM_50_top0.005_pairwise_additive.png)\(b\)Kimi\-Audio

Figure 1:Cross\-lingual agreement of monolingually\-identified ESNs across two models \(see other two in Appendix[D\.1](https://arxiv.org/html/2608.08772#A4.SS1)\)\. Each heatmap displays JSC \(lower triangle\), measuring neuron set intersection, and Spearman’sρ\\rho\(upper triangle\), measuring agreement in selectivity score rankings\. All neurons selected using ConAct withr=0\.5%r\{=\}0\.5\\%\.We first examine whether ESNs identified independently per language exhibit cross\-lingual agreement in both ranking structure and set membership\. Figure[1](https://arxiv.org/html/2608.08772#S5.F1)present pairwise Jaccard Similarity Coefficient \(JSC\) and Spearman’s rank correlation \(ρ\\rho\) for all four LALMs across eight identification languages\. Across all models, we observe a consistent dissociation: the exact set intersection of top\-ranked neurons \(discrete identity overlap\) is minimal \(JSC predominantly below 0\.10, with occasional pairs reaching 0\.15–0\.16\), while the underlying selectivity scores retain a shared ranking structure \(ρ\\rhomostly between 0\.10 and 0\.25 for Audio\-Flamingo\-3 and Kimi\-Audio, and up to≈\\approx0\.6 for MiniCPM\-o\-4\.5\)\.

### 5\.2Saturation of Monolingual Evidence

![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/case/Case_Plot_MSP-PODCAST-Publish-1.12_CAM_ablate_top0.005.png)\(a\)MSP\-Podcast \(Deact\.\)
![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/case/Case_Plot_MSP-PODCAST-Publish-1.12_CAM_steer_top0.005_alpha0.5.png)\(b\)MSP\-Podcast \(Steering\)

Figure 2:Saturation of monolingual evidence: UAR \(%\) under intervention as identification sample size grows \(MSP\-Podcast; BIIC in Appendix[D\.2](https://arxiv.org/html/2608.08772#A4.SS2)shows the same plateau beyond∼50\{\\sim\}50instances\)\.We evaluate the causal properties of identified neurons through functional intervention\. To establish monolingual baselines before introducing fusion, Figure[2](https://arxiv.org/html/2608.08772#S5.F2)shows the causal efficacy of ESN masks identified using varying numbers of correctly predicted instances from the Identification Set \(ℒid\\mathcal\{L\}\_\{\\text\{id\}\}\)\. Across all evaluated models, intervention effects plateau rapidly beyond approximately 50 instances\.

The bottleneck for isolating transferable affective units is therefore the linguistic scope of the identification data rather than sample volume, making cross\-lingual evidence integration essential\. Accordingly, all subsequent experiments use a fixed budget ofc=50c\{=\}50instances per emotion to focus on the causal benefits of cross\-lingual consistency \(see Appendix[A\.5](https://arxiv.org/html/2608.08772#A1.SS5)for exact counts and justification\)\.

### 5\.3Causal Validation of MLENs Through Targeted Intervention

Table 1:Deactivation effects on ESS\. “Avg\.” rows denote means across all languages regardless of group; standard deviation in parentheses\. Negative values indicate reduced specificity, and color intensity indicates effect magnitude\. Per\-language results in Appendix[D\.5](https://arxiv.org/html/2608.08772#A4.SS5)Table[8](https://arxiv.org/html/2608.08772#A4.T8)\.Table 2:Steering effects on ESS withα=0\.5\\alpha\{=\}0\.5\. “Avg\.” rows denote means across all languages regardless of group; standard deviation in parentheses\. Positive values indicate enhanced specificity, and color intensity indicates effect magnitude\. Per\-language results in Appendix[D\.5](https://arxiv.org/html/2608.08772#A4.SS5)Table[9](https://arxiv.org/html/2608.08772#A4.T9)\.We evaluate the functional necessity and influence of the identified neurons through two targeted interventions: deactivation \(to assess necessity for recognition\) and activation steering \(to assess the neurons’ capacity to modulate model behavior\)\.

Deactivation\.To assess whether identified neurons contribute causally to emotion recognition, we measure the reduction in ESS under deactivation, which are summarized in Table[1](https://arxiv.org/html/2608.08772#S5.T1)across two language groups\. CR\-Fusion \(λ=0\.3\\lambda\{=\}0\.3\) exceeds the best monolingual mask for three models in both groups\. Joint Fusion does so on Audio\-Flamingo\-3 and Kimi\-Audio, but falls short of the strongest monolingual mask on MiniCPM\-o\-4\.5 \(−6\.26\-6\.26vs\.−7\.64\-7\.64\) and Qwen2\.5\-Omni\-7B \(−6\.05\-6\.05vs\.−6\.99\-6\.99\), consistent with its token\-weighted pooling being dominated by high\-resource conditions\. Overlap Fusion is the weakest fusion strategy on three of four models and reverses sign on Qwen2\.5\-Omni\-7B \(\+0\.58\+0\.58\); Audio\-Flamingo\-3 is the exception, where the consensus set remains strongly causal\.

Steering\.Table[2](https://arxiv.org/html/2608.08772#S5.T2)shows steering effects across two language groups\. CR\-Fusion \(λ\\lambda=0\.3\) yields superior steering efficacy inℒheld\\mathcal\{L\}\_\{\\text\{held\}\}for three of the four LALMs\. For MiniCPM\-o\-4\.5, CR\-Fusion achieves\+6\.48\+6\.48pp in held\-out languages, substantially outperforming English \(\+4\.70\+4\.70pp\) and Mandarin \(\+5\.48\+5\.48pp\) baselines\. The exception is Qwen2\.5\-Omni\-7B, where the Mandarin mask remains marginally stronger \(\+3\.97\+3\.97vs\.\+3\.64\+3\.64pp\); this is consistent with Qwen having the lowest language\-invariant share \(§3\.4, Appendix[B](https://arxiv.org/html/2608.08772#A2)\), so there is less invariant structure for fusion to recover\. Together, these results provide causal evidence that CR\-Fusion identifies functionally important, language\-shared affective units\.

## 6Discussion

### 6\.1Emotion\-Level Heterogeneity in Cross\-Lingual Generalization

Causal interventions reveal substantial heterogeneity across emotions\. Anger, happiness, and sadness exhibit the largest ESS magnitudes across all evaluated LALMs under both deactivation and steering, and this ordering is preserved across the identification and held\-out groups \(Appendix[A\.6](https://arxiv.org/html/2608.08772#A1.SS6)\)\. Fear and neutral show weaker and more variable effects; for fear, this plausibly reflects acoustic overlap with high\-arousal anger and a floor effect from near\-zero baseline accuracy in several model–language pairs \(Appendix[D\.4](https://arxiv.org/html/2608.08772#A4.SS4)\), which leaves few correctly predicted instances available for identification \(Appendix[A\.5](https://arxiv.org/html/2608.08772#A1.SS5)\) rather than a shortage of fear utterances in the corpora themselves \(Table[6](https://arxiv.org/html/2608.08772#A2.T6)\)\. Strong causal potency does not, however, imply more language\-invariant encoding: per\-emotion invariant shares are highest for neutral and fear, so the causal ordering is better explained by baseline accuracy and class support than by representational universality\.

![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/Emotion_Violin_ASMDSUMSEMBINEUR_consistency_L0.3_agg_CAM_ablate.png)\(a\)Deactivation
![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/Emotion_Violin_ASMDSUMSEMBINEUR_consistency_L0.3_agg_CAM_steer_alpha0.5.png)\(b\)Steering,α=0\.5\\alpha\{=\}0\.5

Figure 3:ESS distributions under \(a\) deactivation and \(b\) steering, comparing CR\-Fusion \(λ=0\.3\\lambda\{=\}0\.3\) against monolingual baselines across all evaluated emotions\. Per\-emotion magnitudes are provided in Appendix[A\.6](https://arxiv.org/html/2608.08772#A1.SS6)\.Independently of these differences, Figure[3](https://arxiv.org/html/2608.08772#S6.F3)shows that CR\-Fusion \(λ=0\.3\\lambda\{=\}0\.3\) matches or exceeds monolingual baselines for every emotion in both groups\. Consistent with §[3\.4](https://arxiv.org/html/2608.08772#S3.SS4), this advantage does not come from suppressing cross\-lingual variance \(the penalty is near zero at the operating point\) but from pooling multiple corpora to estimate the language\-invariant component\.

### 6\.2The Impact of Consistency Penalty Strength

![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/lambda/Consistency_ESS_Trends_MiniCPM-o-4_5_CAM_ablate.png)\(a\)MiniCPM \(Deact\.\)
![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/lambda/Consistency_ESS_Trends_MiniCPM-o-4_5_CAM_steer_alpha0.5.png)\(b\)MiniCPM \(Steering\)

Figure 4:Sensitivity of intervention effects to consistency penaltyλ\\lambda\. Average ESS \(±1\\pm 1standard error of the mean across 12 languages, shown as the shaded band\) under \(a\) deactivation and \(b\) steering for MiniCPM\. Results for other LALMs are provided in Appendix[D](https://arxiv.org/html/2608.08772#A4)\.We analyze the sensitivity of causal effects to the consistency penaltyλ\\lambda, which trades aggregate selector strength \(μ\\mu\) against cross\-lingual stability \(lowσ\\sigma\)\. Figure[4](https://arxiv.org/html/2608.08772#S6.F4)and Appendix[D\.3](https://arxiv.org/html/2608.08772#A4.SS3)Figure[9](https://arxiv.org/html/2608.08772#A4.F9)show that intervention effects are strongest at smallλ\\lambdaand degrade sharply beyondλ≈1\\lambda\\approx 1, approaching the random\-selection baseline byλ≈3\\lambda\\approx 3for all four LALMs; for Qwen2\.5\-Omni\-7B under deactivation the effect reverses sign, indicating that at largeλ\\lambdathe selected units are no longer emotion\-selective at all\. We therefore adoptλ=0\.3\\lambda\{=\}0\.3as the operating point, where the penalty is active but weak\.

This profile favors both predictions of §[3\.4](https://arxiv.org/html/2608.08772#S3.SS4): the optimal penalty is near zero under the expected\-transfer objective, and largeλ\\lambdapreferentially discards low\-resource conditions while shifting selection toward uniformly weak units\. Overlap Fusion is correspondingly the weakest fusion strategy in Table[2](https://arxiv.org/html/2608.08772#S5.T2)for all four models and in Table[1](https://arxiv.org/html/2608.08772#S5.T1)for three\.

### 6\.3Language Contributions and Asymmetric Transfer

![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/loo/SER_LOO_Aggregate_af3-kimi-mini-qwen25_CAM_50_top0.005_ablate_consistency_L0.3_fixed.png)\(a\)Deactivation
![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/loo/SER_LOO_Aggregate_af3-kimi-mini-qwen25_CAM_50_top0.005_steer_consistency_L0.3_alpha0.5_fixed.png)\(b\)Steering,α=0\.5\\alpha\{=\}0\.5

Figure 5:Leave\-one\-outΔ\\DeltaESS heatmaps averaged over four models: effect of excluding each identification language \(columns\) on each evaluation language \(rows\)\.We use leave\-one\-out to test whether the multilingual mask is driven by a few dominant identification languages: for each identification language, we remove it from the fusion pool, re\-identify the CR\-Fusion mask, and measure the change in ESS on each evaluation language \(Figure[5](https://arxiv.org/html/2608.08772#S6.F5)\)\. Under deactivation, no single removal explains the gains of CR\-Fusion; removals instead produce consistent but non\-uniform ESS reductions across evaluation languages, including when low\-resource languages are removed\. The fused mask thus aggregates partially complementary evidence, with low\-resource corpora supplying non\-redundant signal for selecting transferable units\. Steering shows the same qualitative pattern with smaller and more localized changes\. Finally, because each language is represented by a single corpus, language identity and recording condition \(elicitation style, channel, speakers\) are confounded; we therefore read it as evidence of non\-redundant cross\-corpus contributions rather than a typological transfer hierarchy\.

## 7Conclusion

This work presents the first systematic mechanistic investigation of multilingual emotion representations in LALMs\. Through causal interventions across 12 typologically diverse languages and four modern LALMs, we showed that monolingually identified emotion neurons share little set\-level overlap and that monolingual evidence saturates rapidly, motivating multilingual identification; that CR\-Fusion recovers a cross\-lingually shared emotion component inaccessible to any single\-corpus procedure, yielding more transferable causal control in zero\-shot and low\-resource settings; and that leave\-one\-out ablation reveals asymmetric, non\-substitutable language contributions, with low\-resource corpora both supplying non\-redundant identification evidence and being among those that benefit most from fusion\.

By demonstrating shared affective representations that generalize across diverse spoken languages, this work provides a mechanistic foundation for equitable affective computing, enabling speech emotion recognition systems that can serve low\-resource language communities without requiring extensive language\-specific training data\.

## Limitations

Our language stratification reflects functional roles in the experimental pipeline rather than strict claims about model exposure, as LALMs may have encountered these languages implicitly during web\-crawled pre\-training\. Additionally, our analysis addresses discrete emotion categories; extending to dimensional affect models or fine\-grained emotional states remains for future work\.

We also admit that each language in our study is represented by a single corpus, so language identity and recording condition \(elicitation style, channel, speaker population\) are confounded by construction\. A post\-hoc analysis of the language×\\timesemotion interaction component of Eq\. \([2](https://arxiv.org/html/2608.08772#S3.E2)\) indicates that similarity in language\-conditioned activation structure tracks corpus elicitation type \(acted vs\. naturalistic\) at least as strongly as typological relatedness; the asymmetric transfer patterns of §[6\.3](https://arxiv.org/html/2608.08772#S6.SS3)should therefore be read as corpus\-level rather than purely linguistic effects\. Disentangling the two requires at least one language represented by corpora of different elicitation types, which we leave to future work\. Third, all fusion strategies studied here recover only the language\-invariant component of the emotion code; the language\-specific component, which accounts for4141–70%70\\%of emotion\-conditioned activation variance across models, is discarded, and exploiting it for target\-aware adaptation remains open\.

## Ethical Considerations and Broader Societal Impact

While our work aims to advance mechanistic understanding of emotion processing in multimodal foundation models, we acknowledge potential risks associated with neuron\-level emotion manipulation techniques\. The intervention methods demonstrated in this study could be misused to artificially induce or suppress emotional signals in generated speech, enabling more persuasive synthetic media or deceptive voice\-based systems\. Furthermore, fine\-grained understanding of ESNs could facilitate targeted attacks on emotion recognition systems or be exploited to bypass affective content moderation\. We emphasize that our interventions are evaluated only in controlled research settings on pre\-trained models and are not packaged or evaluated as a deployment system; nevertheless, the underlying techniques could be employed in undesirable ways\. We hence encourage the research community to develop appropriate safeguards, such as detection mechanisms for artificially manipulated emotional content, alongside continued interpretability research\.

## Acknowledgments

This work was supported by the National Science Foundation \(NSF\) under CAREER Award IIS\-2533652, and in part by the Singapore Ministry of Digital Development and Information under its AIVP Programme \(Award Number: AIVP\-2026\-010\)\.

## References

- UrduSER: a comprehensive dataset for speech emotion recognition in urdu language\.60,pp\. 111627\.External Links:ISSN 2352\-3409,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.dib.2025.111627),[Link](https://www.sciencedirect.com/science/article/pii/S2352340925003580)Cited by:[§4](https://arxiv.org/html/2608.08772#S4.SS0.SSS0.Px1.p1.2)\.
- A\. Akman, Q\. Sun, and B\. W\. Schuller \(2025\)Improving audio explanations using audio language models\.IEEE Signal Processing Letters32\(\),pp\. 741–745\.External Links:[Document](https://dx.doi.org/10.1109/LSP.2025.3532218)Cited by:[Appendix C](https://arxiv.org/html/2608.08772#A3.SS0.SSS0.Px2.p1.1)\.
- B\. B\. Al\-onazi, M\. A\. Nauman, R\. Jahangir, M\. M\. Malik, E\. H\. Alkhammash, and A\. M\. Elshewey \(2022\)Transformer\-based multilingual speech emotion recognition using data augmentation and feature fusion\.12\(18\)\.External Links:[Link](https://www.mdpi.com/2076-3417/12/18/9188),ISSN 2076\-3417Cited by:[§2](https://arxiv.org/html/2608.08772#S2.SS0.SSS0.Px1.p1.1)\.
- E\.M\. Albornoz and D\.H\. Milone \(2017\)Emotion recognition in never\-seen languages using a novel ensemble method with emotion profiles\.8\(1\),pp\. 43–53\.External Links:[Document](https://dx.doi.org/10.1109/TAFFC.2015.2503757)Cited by:[§1](https://arxiv.org/html/2608.08772#S1.p2.1),[§2](https://arxiv.org/html/2608.08772#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Amentes, N\. Davidchuk, and I\. Lubenets \(2023\)RESD \(revision 75ed61a\)\.Hugging Face\.External Links:[Link](https://huggingface.co/datasets/Aniemore/resd),[Document](https://dx.doi.org/10.57967/hf/1273)Cited by:[§4](https://arxiv.org/html/2608.08772#S4.SS0.SSS0.Px1.p1.2)\.
- A\. Bau, Y\. Belinkov, H\. Sajjad, N\. Durrani, F\. Dalvi, and J\. Glass \(2019\)Identifying and controlling important neurons in neural machine translation\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=H1z-PsR5KX)Cited by:[§1](https://arxiv.org/html/2608.08772#S1.p5.1),[§3\.2](https://arxiv.org/html/2608.08772#S3.SS2.p1.4)\.
- D\. Bau, B\. Zhou, A\. Khosla, A\. Oliva, and A\. Torralba \(2017\)Network Dissection: Quantifying Interpretability of Deep Visual Representations\.In2017 IEEE Conference on Computer Vision and Pattern Recognition \(CVPR\),Vol\.,Los Alamitos, CA, USA,pp\. 3319–3327\.External Links:ISSN 1063\-6919,[Document](https://dx.doi.org/10.1109/CVPR.2017.354),[Link](https://doi.ieeecomputersociety.org/10.1109/CVPR.2017.354)Cited by:[§1](https://arxiv.org/html/2608.08772#S1.p3.1),[§2](https://arxiv.org/html/2608.08772#S2.SS0.SSS0.Px2.p1.1)\.
- D\. Bau, J\. Zhu, H\. Strobelt, A\. Lapedriza, B\. Zhou, and A\. Torralba \(2020\)Understanding the role of individual units in a deep neural network\.Proceedings of the National Academy of Sciences117\(48\),pp\. 30071–30078\.External Links:ISSN 1091\-6490,[Link](http://dx.doi.org/10.1073/pnas.1907375117),[Document](https://dx.doi.org/10.1073/pnas.1907375117)Cited by:[§2](https://arxiv.org/html/2608.08772#S2.SS0.SSS0.Px2.p1.1)\.
- F\. Burkhardt, O\. Schrüfer, U\. Reichel, H\. Wierstorf, A\. Derington, F\. Eyben, and B\. W\. Schuller \(2025\)EmoDB 2\.0: A Database of Emotional Speech in a World that is not Black or White but Grey\.InInterspeech 2025,pp\. 4488–4492\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2025-1951),ISSN 2958\-1796Cited by:[§4](https://arxiv.org/html/2608.08772#S4.SS0.SSS0.Px1.p1.2)\.
- C\. Busso, R\. Lotfian, K\. Sridhar, A\. N\. Salman, W\. Lin, L\. Goncalves, S\. Parthasarathy, A\. R\. Naini, S\. Leem, L\. Martinez\-Lucas, H\. Chou, and P\. Mote \(2026\)The MSP\-Podcast Corpus\.IEEE Transactions on Affective Computing1,pp\. 1–19\.External Links:ISSN 1949\-3045,[Document](https://dx.doi.org/10.1109/TAFFC.2026.3678489),[Link](https://doi.ieeecomputersociety.org/10.1109/TAFFC.2026.3678489)Cited by:[§4](https://arxiv.org/html/2608.08772#S4.SS0.SSS0.Px1.p1.2)\.
- F\. Catania, J\. W\. Wilke, and F\. Garzotto \(2025\)Emozionalmente: a crowdsourced corpus of simulated emotional speech in italian\.33\(\),pp\. 1142–1155\.External Links:[Document](https://dx.doi.org/10.1109/TASLPRO.2025.3540662)Cited by:[§4](https://arxiv.org/html/2608.08772#S4.SS0.SSS0.Px1.p1.2)\.
- Z\. Cheng, Z\. Cheng, J\. He, K\. Wang, Y\. Lin, Z\. Lian, X\. Peng, and A\. G\. Hauptmann \(2024\)Emotion\-LLaMA: multimodal emotion recognition and reasoning with instruction tuning\.InThe Thirty\-eighth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=qXZVSy9LFR)Cited by:[§1](https://arxiv.org/html/2608.08772#S1.p4.1)\.
- I\. Christop \(2024\)NEMO: dataset of emotional speech in Polish\.InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING 2024\),N\. Calzolari, M\. Kan, V\. Hoste, A\. Lenci, S\. Sakti, and N\. Xue \(Eds\.\),Torino, Italia,pp\. 12111–12116\.External Links:[Link](https://aclanthology.org/2024.lrec-main.1059/)Cited by:[§4](https://arxiv.org/html/2608.08772#S4.SS0.SSS0.Px1.p1.2)\.
- F\. Dalvi, N\. Durrani, H\. Sajjad, Y\. Belinkov, A\. Bau, and J\. Glass \(2019\)What is one grain of sand in the desert? analyzing individual neurons in deep nlp models\.InProceedings of the Thirty\-Third AAAI Conference on Artificial Intelligence and Thirty\-First Innovative Applications of Artificial Intelligence Conference and Ninth AAAI Symposium on Educational Advances in Artificial Intelligence,AAAI’19/IAAI’19/EAAI’19\.External Links:ISBN 978\-1\-57735\-809\-1,[Link](https://doi.org/10.1609/aaai.v33i01.33016309),[Document](https://dx.doi.org/10.1609/aaai.v33i01.33016309)Cited by:[§1](https://arxiv.org/html/2608.08772#S1.p5.1),[§3\.2](https://arxiv.org/html/2608.08772#S3.SS2.p1.4)\.
- J\. Fang, Z\. Bi, R\. Wang, H\. Jiang, Y\. Gao, K\. Wang, A\. Zhang, J\. Shi, X\. Wang, and T\. Chua \(2024\)Towards neuron attributions in multimodal large language models\.InProceedings of the 38th International Conference on Neural Information Processing Systems,NIPS ’24,Red Hook, NY, USA\.External Links:ISBN 9798331314385Cited by:[Appendix C](https://arxiv.org/html/2608.08772#A3.SS0.SSS0.Px2.p1.1)\.
- L\. Goncalves, D\. Robinson, E\. Richerson, and C\. Busso \(2024\)Bridging Emotions Across Languages: Low Rank Adaptation for Multilingual Speech Emotion Recognition\.InInterspeech 2024,pp\. 4688–4692\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2024-1226),ISSN 2958\-1796Cited by:[§2](https://arxiv.org/html/2608.08772#S2.SS0.SSS0.Px1.p1.1)\.
- P\. Gournay, O\. Lahaie, and R\. Lefebvre \(2018\)A canadian french emotional speech dataset\.InProceedings of the 9th ACM Multimedia Systems Conference,MMSys ’18,New York, NY, USA,pp\. 399–402\.External Links:ISBN 9781450351928,[Link](https://doi.org/10.1145/3204949.3208121),[Document](https://dx.doi.org/10.1145/3204949.3208121)Cited by:[§4](https://arxiv.org/html/2608.08772#S4.SS0.SSS0.Px1.p1.2)\.
- W\. Gurnee, T\. Horsley, Z\. C\. Guo, T\. R\. Kheirkhah, Q\. Sun, W\. Hathaway, N\. Nanda, and D\. Bertsimas \(2024\)Universal neurons in GPT2 language models\.Transactions on Machine Learning Research\.Note:External Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=ZeI104QZ8I)Cited by:[§3\.2](https://arxiv.org/html/2608.08772#S3.SS2.p1.4)\.
- Z\. Han, T\. Geng, H\. Feng, J\. Yuan, K\. Richmond, and Y\. Li \(2025\)Cross\-lingual speech emotion recognition: humans vs\. self\-supervised models\.InICASSP 2025 \- 2025 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),Vol\.,pp\. 1–5\.External Links:[Document](https://dx.doi.org/10.1109/ICASSP49660.2025.10889008)Cited by:[§1](https://arxiv.org/html/2608.08772#S1.p2.1),[§2](https://arxiv.org/html/2608.08772#S2.SS0.SSS0.Px1.p1.1)\.
- C\. Huang, K\. Lu, S\. Wang, C\. Hsiao, C\. Kuan, H\. Wu, S\. Arora, K\. Chang, J\. Shi, Y\. Peng, R\. Sharma, S\. Watanabe, B\. Ramakrishnan, S\. Shehata, and H\. Lee \(2024\)Dynamic\-superb: towards a dynamic, collaborative, and comprehensive instruction\-tuning benchmark for speech\.InICASSP 2024 \- 2024 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),Vol\.,pp\. 12136–12140\.External Links:[Document](https://dx.doi.org/10.1109/ICASSP48485.2024.10448257)Cited by:[§1](https://arxiv.org/html/2608.08772#S1.p4.1)\.
- R\. Huben, H\. Cunningham, L\. R\. Smith, A\. Ewart, and L\. Sharkey \(2024\)Sparse autoencoders find highly interpretable features in language models\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=F76bwRSLeK)Cited by:[§1](https://arxiv.org/html/2608.08772#S1.p3.1),[§2](https://arxiv.org/html/2608.08772#S2.SS0.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2608.08772#S3.SS2.p1.4)\.
- J\. Huo, Y\. Yan, B\. Hu, Y\. Yue, and X\. Hu \(2024\)MMNeuron: discovering neuron\-level domain\-specific interpretation in multimodal large language model\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 6801–6816\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.387/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.387)Cited by:[Appendix C](https://arxiv.org/html/2608.08772#A3.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2608.08772#S2.SS0.SSS0.Px2.p1.1)\.
- J\. C\. Jackson, J\. Watts, T\. R\. Henry, J\. List, R\. Forkel, P\. J\. Mucha, S\. J\. Greenhill, R\. D\. Gray, and K\. A\. Lindquist \(2019\)Emotion semantics show both cultural variation and universal structure\.366\(6472\),pp\. 1517–1522\.External Links:[Document](https://dx.doi.org/10.1126/science.aaw8160),[Link](https://www.science.org/doi/abs/10.1126/science.aaw8160),https://www\.science\.org/doi/pdf/10\.1126/science\.aaw8160Cited by:[§1](https://arxiv.org/html/2608.08772#S1.p1.1)\.
- P\. Koehn and R\. Knowles \(2017\)Six challenges for neural machine translation\.InProceedings of the First Workshop on Neural Machine Translation,T\. Luong, A\. Birch, G\. Neubig, and A\. Finch \(Eds\.\),Vancouver,pp\. 28–39\.External Links:[Link](https://aclanthology.org/W17-3204/),[Document](https://dx.doi.org/10.18653/v1/W17-3204)Cited by:[§1](https://arxiv.org/html/2608.08772#S1.p5.1),[§4](https://arxiv.org/html/2608.08772#S4.SS0.SSS0.Px1.p1.2)\.
- J\. Lee, W\. Lee, O\. Kwon, and H\. Kim \(2025\)Do large language models have “emotion neurons”? investigating the existence and role\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 15617–15639\.External Links:[Link](https://aclanthology.org/2025.findings-acl.806/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.806),ISBN 979\-8\-89176\-256\-5Cited by:[§2](https://arxiv.org/html/2608.08772#S2.SS0.SSS0.Px2.p1.1)\.
- Y\. Lei, S\. Yang, and L\. Xie \(2021\)Fine\-grained emotion strength transfer, control and prediction for emotional speech synthesis\.In2021 IEEE Spoken Language Technology Workshop \(SLT\),pp\. 423–430\.Cited by:[Appendix C](https://arxiv.org/html/2608.08772#A3.SS0.SSS0.Px1.p1.1)\.
- K\. A\. Lindquist, J\. K\. MacCormack, and H\. Shablack \(2015\)The role of language in emotion: predictions from psychological constructionism\.6,pp\. 121301\.Cited by:[§1](https://arxiv.org/html/2608.08772#S1.p1.1)\.
- R\. Liu, B\. Sisman, G\. Gao, and H\. Li \(2021\)Expressive tts training with frame and style reconstruction loss\.IEEE/ACM Transactions on Audio, Speech, and Language Processing29,pp\. 1806–1818\.Cited by:[Appendix C](https://arxiv.org/html/2608.08772#A3.SS0.SSS0.Px1.p1.1)\.
- J\. Lorenzo\-Trueba, G\. E\. Henter, S\. Takaki, J\. Yamagishi, Y\. Morino, and Y\. Ochiai \(2018\)Investigating different representations for modeling and controlling multiple emotions in dnn\-based speech synthesis\.Speech Communication99,pp\. 135–143\.Cited by:[Appendix C](https://arxiv.org/html/2608.08772#A3.SS0.SSS0.Px1.p1.1)\.
- Z\. Ma, M\. Chen, H\. Zhang, Z\. Zheng, W\. Chen, X\. Li, J\. Ye, X\. Chen, and T\. Hain \(2024\)EmoBox: Multilingual Multi\-corpus Speech Emotion Recognition Toolkit and Benchmark\.InInterspeech 2024,pp\. 1580–1584\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2024-788),ISSN 2958\-1796Cited by:[§2](https://arxiv.org/html/2608.08772#S2.SS0.SSS0.Px1.p1.1)\.
- G\. Maraia, L\. Ranaldi, M\. Valentino, and F\. M\. Zanzotto \(2026\)Can activation steering generalize across languages? a study on syllogistic reasoning in language models\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 2739–2753\.External Links:[Link](https://aclanthology.org/2026.eacl-long.125/),[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.125),ISBN 979\-8\-89176\-380\-7Cited by:[§1](https://arxiv.org/html/2608.08772#S1.p3.1),[§2](https://arxiv.org/html/2608.08772#S2.SS0.SSS0.Px2.p1.1)\.
- O\. Mohamad Nezami, P\. Jamshid Lou, and M\. Karami \(2019\)ShEMO: a large\-scale validated database for persian speech emotion detection\.53\(1\),pp\. 1–16\.Cited by:[§4](https://arxiv.org/html/2608.08772#S4.SS0.SSS0.Px1.p1.2)\.
- G\. H\. Mohmad Dar and R\. Delhibabu \(2024\)Speech databases, speech features, and classifiers in speech emotion recognition: a review\.12\(\),pp\. 151122–151152\.External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2024.3476960)Cited by:[§1](https://arxiv.org/html/2608.08772#S1.p5.1)\.
- D\. Namazifard and L\. G\. Poech \(2025\)Isolating culture neurons in multilingual large language models\.InProceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia\-Pacific Chapter of the Association for Computational Linguistics,K\. Inui, S\. Sakti, H\. Wang, D\. F\. Wong, P\. Bhattacharyya, B\. Banerjee, A\. Ekbal, T\. Chakraborty, and D\. P\. Singh \(Eds\.\),Mumbai, India,pp\. 768–785\.External Links:[Link](https://aclanthology.org/2025.findings-ijcnlp.45/),ISBN 979\-8\-89176\-303\-6Cited by:[§3\.2](https://arxiv.org/html/2608.08772#S3.SS2.p1.4)\.
- M\. Neumann and N\. g\. Thang Vu \(2018\)CRoss\-lingual and multilingual speech emotion recognition on english and french\.In2018 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),Vol\.,pp\. 5769–5773\.External Links:[Document](https://dx.doi.org/10.1109/ICASSP.2018.8462162)Cited by:[§2](https://arxiv.org/html/2608.08772#S2.SS0.SSS0.Px1.p1.1)\.
- M\. D\. Pell, L\. Monetta, S\. Paulmann, and S\. A\. Kotz \(2009\)Recognizing emotions in a foreign language\.33\(2\),pp\. 107–120\.Cited by:[§1](https://arxiv.org/html/2608.08772#S1.p1.1)\.
- E\. A\. Retta, E\. Almekhlafi, R\. Sutcliffe, M\. Mhamed, H\. Ali, and J\. Feng \(2023\)A new amharic speech emotion dataset and classification benchmark\.ACM Trans\. Asian Low\-Resour\. Lang\. Inf\. Process\.22\(1\)\.External Links:ISSN 2375\-4699,[Link](https://doi.org/10.1145/3529759),[Document](https://dx.doi.org/10.1145/3529759)Cited by:[§4](https://arxiv.org/html/2608.08772#S4.SS0.SSS0.Px1.p1.2)\.
- J\. A\. Russell \(1991\)Culture and the categorization of emotions\.\.110\(3\),pp\. 426\.Cited by:[§1](https://arxiv.org/html/2608.08772#S1.p1.1)\.
- S\. Sakshi, U\. Tyagi, S\. Kumar, A\. Seth, R\. Selvakumar, O\. Nieto, R\. Duraiswami, S\. Ghosh, and D\. Manocha \(2025\)MMAU: a massive multi\-task audio understanding and reasoning benchmark\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=TeVAZXr3yv)Cited by:[§1](https://arxiv.org/html/2608.08772#S1.p4.1)\.
- M\. Sharma \(2022\)Multi\-lingual multi\-task speech emotion recognition using wav2vec 2\.0\.InICASSP 2022 \- 2022 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),Vol\.,pp\. 6907–6911\.External Links:[Document](https://dx.doi.org/10.1109/ICASSP43922.2022.9747417)Cited by:[§2](https://arxiv.org/html/2608.08772#S2.SS0.SSS0.Px1.p1.1)\.
- P\. Singh, O\. De Clercq, and E\. Lefever \(2026\)Lost in activations: a neuron\-level analysis of encoders for cross\-lingual emotion detection\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 2: Short Papers\),V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 154–159\.External Links:[Link](https://aclanthology.org/2026.eacl-short.9/),[Document](https://dx.doi.org/10.18653/v1/2026.eacl-short.9),ISBN 979\-8\-89176\-381\-4Cited by:[§1](https://arxiv.org/html/2608.08772#S1.p3.1),[§2](https://arxiv.org/html/2608.08772#S2.SS0.SSS0.Px2.p1.1)\.
- Y\. K\. Singla, J\. Shah, C\. Chen, and R\. R\. Shah \(2022\)What do audio transformers hear? probing their representations for language delivery & structure\.In2022 IEEE International Conference on Data Mining Workshops \(ICDMW\),pp\. 910–925\.Cited by:[Appendix C](https://arxiv.org/html/2608.08772#A3.SS0.SSS0.Px2.p1.1)\.
- R\. Skerry\-Ryan, E\. Battenberg, Y\. Xiao, Y\. Wang, D\. Stanton, J\. Shor, R\. Weiss, R\. Clark, and R\. A\. Saurous \(2018\)Towards end\-to\-end prosody transfer for expressive speech synthesis with tacotron\.Ininternational conference on machine learning,pp\. 4693–4702\.Cited by:[Appendix C](https://arxiv.org/html/2608.08772#A3.SS0.SSS0.Px1.p1.1)\.
- M\. Soumiaa \(2024\)Moroccan dialect emotion recognition dataset\.IEEE Dataport\.External Links:[Document](https://dx.doi.org/10.21227/ev21-c430),[Link](https://dx.doi.org/10.21227/ev21-c430)Cited by:[§4](https://arxiv.org/html/2608.08772#S4.SS0.SSS0.Px1.p1.2)\.
- S\. Sultana, M\. S\. Rahman, M\. R\. Selim, and M\. Z\. Iqbal \(2021\)SUST bangla emotional speech corpus \(subesco\): an audio\-only emotional speech corpus for bangla\.PLOS ONEData in BriefIEEE Transactions on Audio, Speech and Language ProcessingLanguage Resources and EvaluationSpeech CommunicationIEEE AccessIEEE transactions on pattern analysis and machine intelligenceApplied SciencesApplied IntelligenceMachine Intelligence ResearchIEEE accessSpeech CommunicationProceedings of the IEEEIEEE Transactions on Affective ComputingIEEE/ACM Transactions on Audio, Speech, and Language ProcessingIEEE/ACM Transactions on Audio, Speech, and Language ProcessingSpeech CommunicationIEEE Trans\. Affect\. Comput\.ACM Trans\. Inf\. Syst\.Transformer Circuits ThreadApplied SciencesProceedings of the AAAI Conference on Artificial IntelligenceIEEE Transactions on Affective ComputingComplex & Intelligent SystemsFrontiers in psychologyJournal of Nonverbal BehaviorPsychological bulletinScienceACM Comput\. Surv\.The American Journal of PsychologyNew phytologistTheory and Practice in Language StudiesIEEE Transactions on Pattern Analysis and Machine IntelligenceIEEE/ACM Transactions on Audio, Speech, and Language Processing16,pp\. 1–27\.External Links:[Document](https://dx.doi.org/10.1371/journal.pone.0250173),[Link](https://doi.org/10.1371/journal.pone.0250173)Cited by:[§4](https://arxiv.org/html/2608.08772#S4.SS0.SSS0.Px1.p1.2)\.
- T\. Tang, W\. Luo, H\. Huang, D\. Zhang, X\. Wang, X\. Zhao, F\. Wei, and J\. Wen \(2024\)Language\-specific neurons: the key to multilingual capabilities in large language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 5701–5715\.External Links:[Link](https://aclanthology.org/2024.acl-long.309/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.309)Cited by:[§2](https://arxiv.org/html/2608.08772#S2.SS0.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2608.08772#S3.SS2.p1.4)\.
- I\. R\. Ulgen, Z\. Du, C\. Busso, and B\. Sisman \(2024\)Revealing emotional clusters in speaker embeddings: a contrastive learning strategy for speech emotion recognition\.InICASSP 2024 \- 2024 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),Vol\.,pp\. 12081–12085\.External Links:[Document](https://dx.doi.org/10.1109/ICASSP48485.2024.10447060)Cited by:[Appendix C](https://arxiv.org/html/2608.08772#A3.SS0.SSS0.Px2.p1.1)\.
- S\. G\. Upadhyay, W\. Chien, B\. Su, L\. Goncalves, Y\. Wu, A\. N\. Salman, C\. Busso, and C\. Lee \(2023\)An intelligent infrastructure toward large scale naturalistic affective speech corpora collection\.In2023 11th International Conference on Affective Computing and Intelligent Interaction \(ACII\),Vol\.,pp\. 1–8\.External Links:[Document](https://dx.doi.org/10.1109/ACII59096.2023.10388175)Cited by:[§4](https://arxiv.org/html/2608.08772#S4.SS0.SSS0.Px1.p1.2)\.
- E\. Voita, J\. Ferrando, and C\. Nalmpantis \(2024\)Neurons in large language models: dead, n\-gram, positional\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 1288–1301\.External Links:[Link](https://aclanthology.org/2024.findings-acl.75/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.75)Cited by:[§2](https://arxiv.org/html/2608.08772#S2.SS0.SSS0.Px2.p1.1)\.
- Y\. Wang, D\. Stanton, Y\. Zhang, R\. Ryan, E\. Battenberg, J\. Shor, Y\. Xiao, Y\. Jia, F\. Ren, and R\. A\. Saurous \(2018\)Style tokens: unsupervised style modeling, control and transfer in end\-to\-end speech synthesis\.InInternational conference on machine learning,pp\. 5180–5189\.Cited by:[Appendix C](https://arxiv.org/html/2608.08772#A3.SS0.SSS0.Px1.p1.1)\.
- J\. M\. Wilce \(2009\)Language and emotion\.Studies in the Social and Cultural Foundations of Language,Cambridge University Press\.Cited by:[§1](https://arxiv.org/html/2608.08772#S1.p1.1)\.
- P\. Wu, Z\. Ling, L\. Liu, Y\. Jiang, H\. Wu, and L\. Dai \(2019\)End\-to\-end emotional speech synthesis using style tokens and semi\-supervised training\.In2019 Asia\-Pacific Signal and Information Processing Association Annual Summit and Conference \(APSIPA ASC\),pp\. 623–627\.Cited by:[Appendix C](https://arxiv.org/html/2608.08772#A3.SS0.SSS0.Px1.p1.1)\.
- T\. Wu, Y\. Lin, and T\. Weng \(2024\)AND: audio network dissection for interpreting deep acoustic models\.InProceedings of the 41st International Conference on Machine Learning,ICML’24\.Cited by:[§2](https://arxiv.org/html/2608.08772#S2.SS0.SSS0.Px2.p1.1)\.
- J\. Xu, C\. Lan, and Y\. Lu \(2025\)Deciphering functions of neurons in vision\-language models\.InProceedings of the 33rd ACM International Conference on Multimedia,pp\. 3173–3181\.Cited by:[Appendix C](https://arxiv.org/html/2608.08772#A3.SS0.SSS0.Px2.p1.1)\.
- Y\. Xu, H\. Chen, J\. Yu, Q\. Huang, Z\. Wu, S\. Zhang, G\. Li, Y\. Luo, and R\. Gu \(2024\)SECap: speech emotion captioning with large language model\.38\(17\),pp\. 19323–19331\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/29902),[Document](https://dx.doi.org/10.1609/aaai.v38i17.29902)Cited by:[§2](https://arxiv.org/html/2608.08772#S2.SS0.SSS0.Px1.p1.1)\.
- Z\. Yu and S\. Ananiadou \(2024\)Neuron\-level knowledge attribution in large language models\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 3267–3280\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.191/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.191)Cited by:[§2](https://arxiv.org/html/2608.08772#S2.SS0.SSS0.Px2.p1.1)\.
- W\. Zehra, A\. R\. Javed, Z\. Jalil, H\. U\. Khan, and T\. R\. Gadekallu \(2021\)Cross corpus multi\-lingual speech emotion recognition using ensemble learning\.7\(4\),pp\. 1845–1854\.Cited by:[§2](https://arxiv.org/html/2608.08772#S2.SS0.SSS0.Px1.p1.1)\.
- X\. Zhao, R\. Choenni, R\. Saxena, and I\. Titov \(2026a\)Finding culture\-sensitive neurons in vision\-language models\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 3366–3381\.External Links:[Link](https://aclanthology.org/2026.eacl-long.155/),ISBN 979\-8\-89176\-380\-7Cited by:[§3\.2](https://arxiv.org/html/2608.08772#S3.SS2.p1.4)\.
- X\. Zhao, B\. Schuller, and B\. Sisman \(2026b\)Discovering and causally validating emotion\-sensitive neurons in large audio\-language models\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),San Diego, California, United States,pp\. 15056–15071\.External Links:[Link](https://aclanthology.org/2026.acl-long.687/),[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.687),ISBN 979\-8\-89176\-390\-6Cited by:[§2](https://arxiv.org/html/2608.08772#S2.SS0.SSS0.Px2.p1.1)\.
- H\. Zou, F\. Lv, D\. Zheng, E\. S\. Chng, and D\. Rajan \(2025\)Large language models meet contrastive learning: zero\-shot emotion recognition across languages\.In2025 IEEE International Conference on Multimedia and Expo \(ICME\),Vol\.,pp\. 1–6\.External Links:[Document](https://dx.doi.org/10.1109/ICME59968.2025.11209040)Cited by:[§2](https://arxiv.org/html/2608.08772#S2.SS0.SSS0.Px1.p1.1)\.

## Appendix AReproducibility

### A\.1Computational Resources

All experiments were conducted using NVIDIA A100 and H100 GPUs from a shared cluster budget with 8 GPUs of each type\. The total GPU consumption, including failed runs and debugging, was approximately 312 A100\-hours and 687 H100\-hours\. The majority of compute time was dedicated to intervention experiments across 12 languages, four LALMs and multiple parameter settings, with activation logging requiring comparatively fewer resources\. Each LALM required between 40\-80 GB of GPU memory depending on model size and batch configuration\. We estimate that reproducing only the final reported experiments would require approximately 400 A100\-hours or equivalent\.

### A\.2Datasets and Models

Table 3:Dataset statistics showing utterance counts per emotion across 12 multilingual speech emotion languages/datasets, grouped by experimental role: Identification and Held\.For each dataset, we sample up to 150 utterances per emotion for the evaluation partition using a fixed random seed to ensure reproducibility; when a dataset contains fewer than 150 utterances for an emotion \(e\.g\., CaFE, EmoDB, ShEMO; Table[3](https://arxiv.org/html/2608.08772#A1.T3)\), all available utterances are used\.

Table 4:Sources and licenses for the four evaluated LALMs\.
### A\.3SER Prompt Template

To reduce intrinsic positional and label preference biases within LALMs, we randomize the index↔\\leftrightarrowemotion mapping per instance: for each speech clip instance, each emotion is randomly assigned to an option letter \(e\.g\., “A”\)\. We implement the following prompt template for SER\.

Based on the provided speech clip, identify the emotion expressed in the speech\.Choose the option that best matches the perceived emotion from the audio:A: emotion 1B: emotion 2C: emotion 3D: emotion 4E: emotion 5Output exactly one option letter and no other text\.

### A\.4Decoding and Statistical Reporting

We decode deterministically \(greedy; temperature 0\) with a 20\-token generation limit and apply lightweight post\-processing to extract the option letter from model outputs\. Deterministic decoding ensures that model outputs are as reproducible as possible given fixed inputs and model weights\. Consequently, repeated runs on identical data produce identical results, and variance in reported metrics reflects cross\-language or cross\-condition variability rather than stochastic sampling noise\. In Tables[1](https://arxiv.org/html/2608.08772#S5.T1)and[2](https://arxiv.org/html/2608.08772#S5.T2), we report mean and standard deviation \(in parentheses\) computed across evaluation languages within each group, quantifying the consistency of intervention effects across typologically diverse linguistic contexts\. For sensitivity analyses \(Figure[4](https://arxiv.org/html/2608.08772#S6.F4)\), shaded bands represent±1\\pm 1SEM computed over all 12 evaluation languages \(n=12n\{=\}12\)\. Data partitioning uses fixed random seeds \(specified in our released code\) to ensure exact reproducibility of all reported results\.

We do not report seed\-based confidence intervals, as deterministic decoding renders repeated runs identical and seed variance is exactly zero\. Bootstrap resampling over test utterances would quantify sampling variability of the evaluation sets; we instead report cross\-language dispersion, which is the variability most relevant to our claims about multilingual generalization\.

All standard deviations and standard errors are computed using closed\-form formulas implemented via NumPy111https://numpy\.org/library functions\.

### A\.5Identification Instance Budget and Data Availability Constraints

Our neuron identification procedure targetsc=50c\{=\}50correctly predicted instances per emotion per language\. However, certain model\-language\-emotion triples exhibit low baseline accuracy, yielding fewer correct predictions available for activation logging\. In such cases, we use all available correctly predicted instances rather than discarding the emotion or language entirely\.

We retain low\-count conditions for two reasons\. First, excluding underperforming conditions would systematically bias our analysis toward high\-resource, well\-recognized settings, undermining our goal of investigating multilingual emotion representations\. Second, our saturation analysis \(§[5\.2](https://arxiv.org/html/2608.08772#S5.SS2)\) demonstrates that intervention effects plateau beyond approximately 50 instances, indicating that reduced counts can still provide informative activation statistics\.

For severely underrepresented cases such as MiniCPM\-o\-4\.5 Fear \(36 total instances, with zero for Arabic and Polish\), we acknowledge reduced reliability\. However, this primarily reflects MiniCPM\-o\-4\.5’s strong bias toward Neutral predictions \(98\.83% baseline accuracy\)\. The emotion\-level analysis \(§[6\.1](https://arxiv.org/html/2608.08772#S6.SS1)\) independently confirms that Fear exhibits weaker causal effects, consistent with data sparsity\. Our primary claims rest on emotions with robust identification data across all models, namely anger, happiness, and sadness\.

### A\.6Per\-Emotion Detail

![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/Emotion_Radar_ASMDSUMSEMBINEUR_consistency_L0.3_agg_CAM_ablate.png)\(a\)Deactivation
![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/Emotion_Radar_ASMDSUMSEMBINEUR_consistency_L0.3_agg_CAM_steer_alpha0.5.png)\(b\)Steering,α=0\.5\\alpha\{=\}0\.5

Figure 6:Per\-emotion ESS magnitudes under \(a\) deactivation and \(b\) steering, comparing CR\-Fusion \(λ\\lambda=0\.3\) against monolingual baselines, reported separately forℒid\\mathcal\{L\}\_\{\\text\{id\}\}andℒheld\\mathcal\{L\}\_\{\\text\{held\}\}\. Radial axes are not shared across panels\.#### Group consistency\.

Figure[6](https://arxiv.org/html/2608.08772#A1.F6)reports per\-emotion ESS magnitudes separately forℒid\\mathcal\{L\}\_\{\\text\{id\}\}andℒheld\\mathcal\{L\}\_\{\\text\{held\}\}\. The relative ordering of emotions is preserved across groups for all models, under both deactivation and steering: the emotions with the strongest effects on identification languages are also the strongest on languages never seen during identification\. Radial axes are not shared across panels\.

#### Causal potency versus representational invariance\.

It is tempting to interprete the ordering in §[6\.1](https://arxiv.org/html/2608.08772#S6.SS1)as evidence that anger, happiness, and sadness are encoded in more language\-invariant form than fear or neutral\. However, the variance decomposition of §[3\.4](https://arxiv.org/html/2608.08772#S3.SS4)does not support that interpretation: per\-emotion invariant shares \(Table[6](https://arxiv.org/html/2608.08772#A2.T6)\) are not highest for the three emotions with the strongest causal effects in any model, and averaged across models the ordering is inverted: neutral \(0\.520\.52\) and fear \(0\.500\.50\) exceed anger \(0\.430\.43\), sadness \(0\.430\.43\), and happiness \(0\.410\.41\)\.

The two quantities measure different things, and the dissociation is informative\. The invariant share is a property of the representation: how much of a neuron’s emotion tuning is shared across languages\. ESS is a property of model behavior under intervention, bounded by baseline recognition accuracy, per\-emotion test support, and confusability within the response set\. Fear is the least or second\-least accurately recognized emotion in every model, and neutral, although nominally the most accurate in three of four models, is inflated by a strong neutral response bias \(Appendix[D\.4](https://arxiv.org/html/2608.08772#A4.SS4)\): its high accuracy reflects a default prediction rather than discriminative use of emotion\-sensitive units, so suppressing neutral\-selective units cannot remove it\. Anger and sadness are recognized well above the fear floor in every model\. Baseline accuracy, response bias and per\-emotion test support therefore suffice to produce the observed causal ordering without any difference in representational invariance\.

## Appendix BVariance Decomposition of Multilingual Emotion Coding

#### Estimation\.

We fit Eq\. \([2](https://arxiv.org/html/2608.08772#S3.E2)\) per neuron by alternating weighted centering over the language and emotion axes of the\|L\|×\|E\|\|L\|\\times\|E\|activation\-probability table, after dividing each language’s tensor by its global standard deviation to remove differences in dynamic range\. For complete tables the fit converges in one iteration and coincides with the two\-way ANOVA solution; for tables with missing cells \(emotions for which a model produced too few correct predictions in a given language, Appendix[A\.5](https://arxiv.org/html/2608.08772#A1.SS5)\) it converges to the weighted least\-squares additive fit\. All quantities below are computed over the top\-r=0\.5%r\{=\}0\.5\\%of neurons ranked by the ConAct margin of the pooled profile, i\.e\. the population from which masks are drawn\.

#### Noise correction\.

SinceP^\\hat\{P\}is a binomial proportion, each cell carries sampling variancev=P​\(1−P\)/T\(ℓ,e\)v=P\(1\-P\)/T^\{\(\\ell,e\)\}, which inflates the apparent interaction variance far more than the invariant variance: noise enters the residualγ^\\hat\{\\gamma\}with weight\(\|L\|−1\)​\(\|E\|−1\)/\(\|L\|​\|E\|\)\(\|L\|\{\-\}1\)\(\|E\|\{\-\}1\)/\(\|L\|\|E\|\)but entersb^\\hat\{b\}with weight only\(\|E\|−1\)/\(\|L\|​\|E\|\)\(\|E\|\{\-\}1\)/\(\|L\|\|E\|\), a factor\|L\|−1\|L\|\{\-\}1smaller, because averaging over languages suppresses it\. We therefore report shares based onσ^γ2=ms​\(γ^\)/fγ−v¯\\hat\{\\sigma\}\_\{\\gamma\}^\{2\}=\\mathrm\{ms\}\(\\hat\{\\gamma\}\)/f\_\{\\gamma\}\-\\bar\{v\}andσ^b2=\(ms​\(b^\)−v¯​fb/\|L\|\)/fb\\hat\{\\sigma\}\_\{b\}^\{2\}=\\big\(\\mathrm\{ms\}\(\\hat\{b\}\)\-\\bar\{v\}f\_\{b\}/\|L\|\\big\)/f\_\{b\}, wherems​\(⋅\)\\mathrm\{ms\}\(\\cdot\)is the mean square,fγ=\(\|L\|−1\)​\(\|E\|−1\)/\(\|L\|​\|E\|\)f\_\{\\gamma\}=\(\|L\|\{\-\}1\)\(\|E\|\{\-\}1\)/\(\|L\|\|E\|\)andfb=\(\|E\|−1\)/\|E\|f\_\{b\}=\(\|E\|\{\-\}1\)/\|E\|are residual projection factors, andv¯\\bar\{v\}is the mean estimated binomial variance\. Uncorrected shares are within0\.030\.03of corrected values for every model, and estimated noise accounts for only1313–22%22\\%of the raw interaction energy, so the language\-specific component is not a sampling artifact\.

Table 5:Variance components of Eq\. \([2](https://arxiv.org/html/2608.08772#S3.E2)\) among top\-ranked emotion\-selective neurons \(r=0\.5%r\{=\}0\.5\\%, ConAct,c=50c\{=\}50\)\. “Share” is the invariant fractionσb2/\(σb2\+σγ2\)\\sigma\_\{b\}^\{2\}/\(\\sigma\_\{b\}^\{2\}\+\\sigma\_\{\\gamma\}^\{2\}\), raw and noise\-corrected\. “JSC” is the mean Jaccard similarity between masks selected on the fitted invariant componentm\+b^m\+\\hat\{b\}and Joint Fusion masks: near\-identity \(0\.980\.98\) for the two models with complete language–emotion tables, as predicted by the algebraic equivalence of §[3\.4](https://arxiv.org/html/2608.08772#S3.SS4), with divergence appearing only where missing cells make the weighted fit differ from naive pooling\.Table 6:Per\-emotion noise\-corrected invariant share\. The share is*not*highest for anger, happiness, and sadness in any model; averaged across models, neutral and fear show the highest invariant shares\. Representational universality therefore does not explain the causal\-potency ordering of §[6\.1](https://arxiv.org/html/2608.08772#S6.SS1)\(Appendix[A\.6](https://arxiv.org/html/2608.08772#A1.SS6)\)\.
#### Interpretation\.

From Table[5](https://arxiv.org/html/2608.08772#A2.T5)and[6](https://arxiv.org/html/2608.08772#A2.T6), we observe that: \(1\) the invariant share is below0\.60\.6across models; in two of four it is below one half\. The language\-invariant emotion code that fusion isolates is real and causally potent \(§[5](https://arxiv.org/html/2608.08772#S5)\), but it is a minority\-to\-half share of the emotion\-conditioned activation structure; \(2\) The masks selected on the fitted invariant component coincide with Joint Fusion masks wherever the design is complete \(JSC0\.980\.98\), confirming empirically that Joint Fusion and small\-λ\\lambdaCR\-Fusion estimate the invariant component rather than performing consensus filtering; \(3\) With only four models we do not treat the invariant share as a predictor of fusion gain, and the two orderings are not monotonically related: MiniCPM\-o\-4\.5 has the highest invariant share \(0\.590\.59\) but one of the smallest deactivation margins over the best monolingual mask \(0\.390\.39pp\), whereas Kimi\-Audio and Audio\-Flamingo\-3 have lower shares and larger margins\. The one systematic observation we draw is that Qwen2\.5\-Omni\-7B, with the lowest share \(0\.300\.30\), is the only model for which a monolingual mask \(Mandarin\) matches fusion under steering\. We report this as a consistency check on the decomposition, not as evidence of a dose–response relation between invariant share and fusion benefit\.

## Appendix CExtended Related Work

#### Controllable Affective Speech Synthesis\.

A complementary line of research controls paralinguistic affect in speech generation through learned style representationsSkerry\-Ryanet al\.\([2018](https://arxiv.org/html/2608.08772#bib.bib53)\); Wanget al\.\([2018](https://arxiv.org/html/2608.08772#bib.bib54)\); Wuet al\.\([2019](https://arxiv.org/html/2608.08772#bib.bib55)\); Liuet al\.\([2021](https://arxiv.org/html/2608.08772#bib.bib56)\)and continuous control variablesLorenzo\-Truebaet al\.\([2018](https://arxiv.org/html/2608.08772#bib.bib60)\); Leiet al\.\([2021](https://arxiv.org/html/2608.08772#bib.bib59)\);xie2025emosteerttsfinegrainedtrainingfreeemotioncontrollable\. These methods control affect via architectural conditioning learned during training, whereas we intervene training\-free on neurons of pre\-trained LALMs\.

#### Neuron Attribution in Multimodal Models\.

Beyond the audio modality, neuron\-level dissection in multimodal foundation models has focused on modality\-specific functional attributionshuang2024minerminingunderlyingpattern; Fanget al\.\([2024](https://arxiv.org/html/2608.08772#bib.bib21)\); Xuet al\.\([2025](https://arxiv.org/html/2608.08772#bib.bib10)\); Huoet al\.\([2024](https://arxiv.org/html/2608.08772#bib.bib20)\), and audio interpretability studies rely on layer\-wise probing of phonetic, speaker, and prosodic cuesSinglaet al\.\([2022](https://arxiv.org/html/2608.08772#bib.bib50)\); Ulgenet al\.\([2024](https://arxiv.org/html/2608.08772#bib.bib168)\); Akmanet al\.\([2025](https://arxiv.org/html/2608.08772#bib.bib42)\)\.

## Appendix DSupplementary Results

### D\.1Cross\-Lingual Agreement of Monolingual Emotion\-Sensitive Neurons

![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/descriptive/Combined_Descriptive_MiniCPM-o-4_5_CAM_50_top0.005_pairwise_additive.png)\(a\)MiniCPM\-o\-4\.5
![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/descriptive/Combined_Descriptive_Qwen2.5-Omni-7B_CAM_50_top0.005_pairwise_additive.png)\(b\)Qwen2\.5\-Omni\-7B

Figure 7:Cross\-lingual agreement of monolingually\-identified ESNs for MiniCPM\-o\-4\.5 and Qwen2\.5\-Omni\-7B\.
### D\.2Saturation of Monolingual Evidence \(BIIC\)

![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/case/Case_Plot_BIIC-Podcast-v1.01_CAM_ablate_top0.005.png)\(a\)BIIC \(Deact\.\)
![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/case/Case_Plot_BIIC-Podcast-v1.01_CAM_steer_top0.005_alpha0.5.png)\(b\)BIIC \(Steering\)

Figure 8:Saturation of monolingual evidence: UAR \(%\) under intervention as identification sample size grows \(BIIC\)\.
### D\.3Sensitivity Analysis for Consistency Penalty

![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/lambda/Consistency_ESS_Trends_audio-flamingo-3-hf_CAM_ablate.png)\(a\)Audio\-Flamingo\-3 \(Deact\.\)
![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/lambda/Consistency_ESS_Trends_audio-flamingo-3-hf_CAM_steer_alpha0.5.png)\(b\)Audio\-Flamingo\-3 \(Steer\.\)
![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/lambda/Consistency_ESS_Trends_Kimi-Audio-7B-Instruct_CAM_ablate.png)\(c\)Kimi\-Audio \(Deact\.\)
![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/lambda/Consistency_ESS_Trends_Kimi-Audio-7B-Instruct_CAM_steer_alpha0.5.png)\(d\)Kimi\-Audio \(Steering\)
![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/lambda/Consistency_ESS_Trends_Qwen2.5-Omni-7B_CAM_ablate.png)\(e\)Qwen2\.5\-Omni \(Deact\.\)
![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/lambda/Consistency_ESS_Trends_Qwen2.5-Omni-7B_CAM_steer_alpha0.5.png)\(f\)Qwen2\.5\-Omni \(Steering\)

Figure 9:Sensitivity to consistency penaltyλ\\lambdaunder deactivation \(left\) and steering \(α\\alpha=0\.5, right\), complementing Figure[4](https://arxiv.org/html/2608.08772#S6.F4): \(a,b\) Audio\-Flamingo\-3, \(c,d\) Kimi\-Audio, \(e,f\) Qwen2\.5\-Omni\-7B\. Shaded bands show±1\\pm 1SEM across 12 languages\.
### D\.4Unintervened Baseline Results

Table 7:Baseline SER accuracy \(%\) across 12 languages with deterministic decoding\. “UAR” rows report unweighted average recall computed over available emotions; “Avg\.” column reports UAR across languages\.Table[7](https://arxiv.org/html/2608.08772#A4.T7)shows the baseline SER accuracies for unintervened models\. Test set class distributions vary across datasets, with some emotions \(e\.g\., fear in Persian\) having few or zero instances\. Global ESS is computed only over emotions with available test cases; per\-emotion breakdowns in Appendix[D\.5](https://arxiv.org/html/2608.08772#A4.SS5)assess robustness across class frequencies\. We note that deactivation effects are bounded by baseline accuracy, creating potential floor effects for low\-accuracy emotions\. Fear’s weak deactivation effect in MiniCPM\-o\-4\.5 \(baseline≈1%\\approx 1\\%\) should be interpreted cautiously given this constraint\.

### D\.5Per\-Language Detailed Results

![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/loo/SER_ESS_Distribution_agg_CAM_50_top0.005_ablate.png)\(a\)ESS Distribution, Deactivation
![Refer to caption](https://arxiv.org/html/2608.08772v1/figures/loo/SER_ESS_Distribution_agg_CAM_50_top0.005_steer_alpha0.5.png)

Figure 10:ESS distributions across monolingual and fusion masks under \(a\) deactivation and \(b\) steering\.Table[8](https://arxiv.org/html/2608.08772#A4.T8)and[9](https://arxiv.org/html/2608.08772#A4.T9)provide per\-language intervention results\. Figure[10](https://arxiv.org/html/2608.08772#A4.F10)displays ESS distributions across mask sources and evaluation settings under deactivation, averaged over four LALMs\. Monolingual masks show substantial variability, with high\-resource anchors \(English, Mandarin\) and certain low\-resource languages \(Bengali, Urdu\) achieving stronger effects than others \(Amharic, Arabic\)\. Consistency fusion \(λ\\lambda=0\.3\) produces a distribution shifted toward the strongest effects, surpassing the best monolingual mask\.

Table 8:Per\-language deactivation effects\. More negative values indicate stronger causal necessity of deactivated neurons\.Table 9:Per\-language steering effects \(α=0\.5\\alpha\{=\}0\.5\)\. More positive values indicate stronger causal sufficiency of amplified neurons\.

Similar Articles

Do Speech Emphasis Models Generalize across Languages and Emotions?

arXiv cs.CL

Introduces MMEE, a multilingual multi-emotion emphasis corpus of 10,000 utterances across 7 languages and 34 emotions, and benchmarks emphasis detection models under various transfer settings, finding that multilingual training improves robustness while monolingual models show limited zero-shot transfer.

Negative Before Positive: Asymmetric Valence Processing in Large Language Models

arXiv cs.CL

This paper investigates how large language models process emotional valence through mechanistic interpretability. Using activation patching and steering on three open-source LLMs, the authors find that negative valence is localized to early layers while positive valence peaks in mid-to-late layers, and they validate this through topic-controlled flip tests.