面向鲁棒多模态情感分析的可靠性感知跨样本增强

arXiv cs.LG 论文

摘要

本文提出了一种可靠性感知跨样本增强(RCE)框架,用于鲁棒多模态情感分析。该框架通过自适应信息瓶颈和跨样本来处理噪声和模态缺失问题,并在多种设置中展示出优越性能。

arXiv:2609.30470v1 Announce Type: new Abstract: Multimodal Sentiment Analysis (MSA) aims to infer human emotions from multiple modalities such as text, audio, and vision. In practice, inputs are often corrupted by noise and missing modalities, which degrades performance. Existing methods typically address these challenges in isolation, limiting their effectiveness in realistic settings. To address this limitation, we propose a Reliability-aware Cross-sample Enhancement (RCE) framework. Specifically, RCE first introduces an adaptive variational information bottleneck to model modality-wise uncertainty and perform quality-aware information compression, thereby suppressing redundant noise in unreliable modalities. Furthermore, we design a reliability-aware cross-sample enhancement strategy that retrieves high-confidence, semantically consistent neighbors from a large candidate pool to enrich and calibrate current representations, effectively alleviating information deficiency caused by missing modalities. Building upon this, RCE integrates cross-modal interactions with a multilevel reliability-aware fusion mechanism to adaptively aggregate information across modalities and enhancement stages, leading to more robust multimodal representations. Extensive experiments demonstrate that RCE consistently outperforms state-of-the-art methods across full, noisy, and missing-modality settings.
查看原文
查看缓存全文

缓存时间: 2026/09/29 09:37

# Reliability-aware Cross-sample Enhancement for Robust Multimodal Sentiment Analysis
Source: [https://arxiv.org/html/2609.30470](https://arxiv.org/html/2609.30470)
Menghua JiangAffiliation:School of Computer Science and Engineering, Sun Yat\-sen UniversityEmail:[jiangmh26@mail2\.sysu\.edu\.cn](mailto:[email protected])Xiangui Kang††thanks:Corresponding authors\.Affiliation:School of Computer Science and Engineering, Sun Yat\-sen UniversityHaifeng HuAffiliation:School of Electronics and Information Technology, Sun Yat\-sen UniversitySijie Mai11footnotemark:1Affiliation:School of Computer Science, South China Normal University

###### Abstract

Multimodal Sentiment Analysis \(MSA\) aims to infer human emotions from multiple modalities such as text, audio, and vision\. In practice, inputs are often corrupted by noise and missing modalities, which degrades performance\. Existing methods typically address these challenges in isolation, limiting their effectiveness in realistic settings\. To address this limitation, we propose a Reliability\-aware Cross\-sample Enhancement \(RCE\) framework\. Specifically, RCE first introduces an adaptive variational information bottleneck to model modality\-wise uncertainty and perform quality\-aware information compression, thereby suppressing redundant noise in unreliable modalities\. Furthermore, we design a reliability\-aware cross\-sample enhancement strategy that retrieves high\-confidence, semantically consistent neighbors from a large candidate pool to enrich and calibrate current representations, effectively alleviating information deficiency caused by missing modalities\. Building upon this, RCE integrates cross\-modal interactions with a multilevel reliability\-aware fusion mechanism to adaptively aggregate information across modalities and enhancement stages, leading to more robust multimodal representations\. Extensive experiments demonstrate that RCE consistently outperforms state\-of\-the\-art methods across full, noisy, and missing\-modality settings\.

## 1Introduction

Multimodal Sentiment Analysis \(MSA\) integrates text, audio, and vision to better capture complex human emotional states[Zadeh et al\. \(2016\)](https://arxiv.org/html/2609.30470#bib.bib4)\. With advances in multimodal learning[Fang et al\. \(2025\)](https://arxiv.org/html/2609.30470#bib.bib50);[Li et al\. \(2024b\)](https://arxiv.org/html/2609.30470#bib.bib37);[He et al\. \(2026\)](https://arxiv.org/html/2609.30470#bib.bib57), MSA models leverage complementary cross\-modal information, narrowing the gap between human expression and machine understanding, and showing promise in applications such as mental health analysis[Ye et al\. \(2025\)](https://arxiv.org/html/2609.30470#bib.bib53)and human–robot interaction[Liu et al\. \(2017\)](https://arxiv.org/html/2609.30470#bib.bib64)\. However, practical deployment remains challenging: multimodal data are often incomplete, noisy, or uncertain due to imperfect collection[Zhang et al\. \(2024\)](https://arxiv.org/html/2609.30470#bib.bib40), sensor noise[Zhuang et al\. \(2026\)](https://arxiv.org/html/2609.30470#bib.bib54), and privacy constraints[Jaiswal and Provost \(2020\)](https://arxiv.org/html/2609.30470#bib.bib65), which significantly degrades model reliability and predictive performance\.

To achieve optimal multimodal performance, existing studies are often conducted under idealized or isolated settings, either implicitly assuming that all modalities available during training remain fully accessible at inference time, or treating modality missingness and noise contamination as separate problems\. Under the full\-modality setting, some approaches focus on modeling cross\-modal discrepancies\. For example, MODS[Yang et al\. \(2026\)](https://arxiv.org/html/2609.30470#bib.bib58)captures modality importance via dynamic dominant\-modality selection, while DecAlign[Qian et al\. \(2026\)](https://arxiv.org/html/2609.30470#bib.bib56)achieves cross\-modal alignment through prototype\-guided optimal transport and distribution matching\. To handle noisy inputs, several methods adopt an information\-theoretic perspective, where MIB[Mai et al\. \(2023b\)](https://arxiv.org/html/2609.30470#bib.bib30)introduces an information bottleneck[Tishby et al\. \(2000\)](https://arxiv.org/html/2609.30470#bib.bib1)to filter redundant information, while OMIB[Wu et al\. \(2025a\)](https://arxiv.org/html/2609.30470#bib.bib45)imposes theoretical constraints on the regularization weights to enable more stable bottleneck representation learning\. To address modality missingness, existing approaches typically rely on reconstruction or cross\-sample information augmentation\. For instance, CyIN[Lin et al\. \(2025\)](https://arxiv.org/html/2609.30470#bib.bib44)reconstructs missing information via cross\-modal cycle translation, while HME[Zhuang et al\. \(2025b\)](https://arxiv.org/html/2609.30470#bib.bib43)enhances representations by retrieving semantically relevant information from other samples\.

However, in real\-world scenarios, modality missingness and noise are often unknown and dynamically evolving, posing significant challenges to methods developed under idealized or isolated settings[Zhuang et al\. \(2026\)](https://arxiv.org/html/2609.30470#bib.bib54)\. Furthermore, many existing approaches improve robustness to incomplete inputs at the expense of performance in full\-modality settings, making it difficult to achieve consistent performance across diverse conditions within a unified framework[Lin et al\. \(2025\)](https://arxiv.org/html/2609.30470#bib.bib44)\. Finally, noisy data typically introduce aleatoric uncertainty[Hu et al\. \(2025\)](https://arxiv.org/html/2609.30470#bib.bib61), yet most existing denoising strategies overlook this factor and fail to explicitly model modality reliability during the denoising process\.

To address the above challenges, we propose a unified framework called Reliability\-aware Cross\-sample Enhancement \(RCE\), which learns robust multimodal representations under complex scenarios where noise corruption and modality missingness coexist\. RCE consists of three key components: ❶*Adaptive Variational Information Bottleneck*, which maps each modality into a von Mises–Fisher \(vMF\) distribution parameterized by directional and concentration variables, explicitly capturing modality\-wise uncertainty and enabling quality\-aware information compression; ❷*Reliability\-aware Modality Enhancement*, which mitigates information deficiency caused by missing or degraded modalities by retrieving semantically consistent and high\-confidence neighbors from a large candidate pool to enrich and calibrate modality representations; and ❸*Hyper\-Modality Generation and Multilevel Fusion*, which models cross\-modal interactions via hyper\-modality representation learning, integrates enhanced modality features with global contextual information, and employs a multilevel reliability\-aware fusion mechanism to adaptively assign weights based on modality confidence, enabling unified modeling from intra\-modality to cross\-modality interactions\. Through these designs, RCE achieves consistent and robust performance across diverse scenarios, including full\-modality, noisy, and missing\-modality settings\.

Our main contributions are summarized as follows:

- •We propose a unified framework, termed RCE, to address multimodal sentiment analysis under realistic scenarios encompassing full\-modality, noisy, and missing\-modality conditions, without requiring separate treatment of each condition\.
- •We develop a reliability\-aware paradigm that incorporates explicit uncertainty modeling, facilitating adaptive information compression, cross\-sample information augmentation, and multi\-level fusion for learning robust multimodal representations\.
- •Extensive experiments on four benchmark datasets and three evaluation settings demonstrate that RCE consistently outperforms existing methods with stronger robustness\.

## 2Related Works

### 2\.1Multimodal Sentiment Analysis

Multimodal Sentiment Analysis \(MSA\) aims to integrate text, audio, and visual signals to predict human emotional states[Zadeh et al\. \(2016\)](https://arxiv.org/html/2609.30470#bib.bib4)\. Most existing methods assume that all modalities are available at inference time and focus on designing sophisticated fusion strategies to learn discriminative multimodal representations[Zadeh et al\. \(2017\)](https://arxiv.org/html/2609.30470#bib.bib14);[Zhang et al\. \(2023a\)](https://arxiv.org/html/2609.30470#bib.bib29);[Wu et al\. \(2025b\)](https://arxiv.org/html/2609.30470#bib.bib48);[Mai et al\. \(2025\)](https://arxiv.org/html/2609.30470#bib.bib47);[Fang et al\. \(2025\)](https://arxiv.org/html/2609.30470#bib.bib50);[Wang et al\. \(2025\)](https://arxiv.org/html/2609.30470#bib.bib46);[Jiang et al\. \(2025b\)](https://arxiv.org/html/2609.30470#bib.bib69);[Jiang et al\. \(2025a\)](https://arxiv.org/html/2609.30470#bib.bib70);[Mai and Han \(2026\)](https://arxiv.org/html/2609.30470#bib.bib55)\. Some studies focus on cross\-modal alignment, such as DecAlign[Qian et al\. \(2026\)](https://arxiv.org/html/2609.30470#bib.bib56), which decouples shared and modality\-specific features and leverages prototype\-guided optimal transport\. Other methods model modality importance, such as MODS[Yang et al\. \(2026\)](https://arxiv.org/html/2609.30470#bib.bib58), which enhances cross\-modal interactions through dynamic primary modality selection\. In addition, some approaches improve representation capacity by incorporating additional semantic information\. For example, PSA\-MF[Xie et al\. \(2026\)](https://arxiv.org/html/2609.30470#bib.bib60)introduces personality\-aware embeddings for fine\-grained fusion\. Although these methods perform well under full\-modality settings, their performance degrades significantly in real\-world scenarios with missing modalities, thereby limiting their practical applicability\.

### 2\.2Multimodal Sentiment Analysis with Missing Modalities

To address missing modalities, existing MSA methods attempt to compensate for incomplete data from multiple perspectives\. One line of work focuses on reconstructing missing modalities from observed ones[Pham et al\. \(2019\)](https://arxiv.org/html/2609.30470#bib.bib16);[Zhao et al\. \(2021\)](https://arxiv.org/html/2609.30470#bib.bib18);[Wang et al\. \(2023a\)](https://arxiv.org/html/2609.30470#bib.bib35);[Lian et al\. \(2023\)](https://arxiv.org/html/2609.30470#bib.bib28);[Zhang et al\. \(2024\)](https://arxiv.org/html/2609.30470#bib.bib40)\. For example, IMDer[Wang et al\. \(2023b\)](https://arxiv.org/html/2609.30470#bib.bib34)employs a diffusion\-based model for reconstruction, while CyIN[Lin et al\. \(2025\)](https://arxiv.org/html/2609.30470#bib.bib44)restores missing information in the latent space via cross\-modal cyclic translation\. Another line of work adopts a teacher–student framework[Zhuang et al\. \(2025a\)](https://arxiv.org/html/2609.30470#bib.bib42)\. In this paradigm, methods such as HRLF[Li et al\. \(2024a\)](https://arxiv.org/html/2609.30470#bib.bib41)leverage a teacher model trained on fully observed multimodal data to guide student models in handling missing modalities through knowledge distillation\. In addition, some approaches enhance representations using cross\-sample information\. CorrKD[Li et al\. \(2024b\)](https://arxiv.org/html/2609.30470#bib.bib37)proposes a sample\-level contrastive distillation mechanism to capture global cross\-sample correlations, whereas HME[Zhuang et al\. \(2025b\)](https://arxiv.org/html/2609.30470#bib.bib43)enriches observed modalities by retrieving semantically relevant clues from other samples, thereby avoiding explicit reconstruction\.

Despite their effectiveness, these methods typically assume noise\-free data, which limits their applicability in real\-world scenarios with substantial noise\. Moreover, HME[Zhuang et al\. \(2025b\)](https://arxiv.org/html/2609.30470#bib.bib43)and CorrKD[Li et al\. \(2024b\)](https://arxiv.org/html/2609.30470#bib.bib37)rely on samples within the current mini\-batch, leading to unstable and limited neighborhood information\. Semantically similar samples with opposite sentiment polarity may further introduce “negative enhancement,” degrading model performance\. In contrast, RCE retrieves semantically consistent and high\-confidence neighbor representations from a large\-scale candidate pool, resulting in more robust multimodal representations\. We also provide a review of related work on multimodal sentiment analysis with noisy modalities\. Please refer to Appendix[B\.1](https://arxiv.org/html/2609.30470#A2.SS1)\.

![Refer to caption](https://arxiv.org/html/2609.30470v1/main.png)Figure 1:Overview of the RCE framework\. For clarity, we provide detailed descriptions and definitions of all notations used in the figure in Appendix[A](https://arxiv.org/html/2609.30470#A1)\.

## 3Methodology

The overall architecture of the proposed RCE framework is shown in Figure[1](https://arxiv.org/html/2609.30470#S2.F1)\. It comprises four modules\. The Pre\-Processing module handles multimodal inputs and extracts modality\-specific representations \(see Section[3\.2](https://arxiv.org/html/2609.30470#S3.SS2)\)\. The Adaptive Variational Information Bottleneck \(AVIB\) module performs reliability\-aware adaptive compression \(see Section[3\.3](https://arxiv.org/html/2609.30470#S3.SS3)\)\. The Reliability\-aware Modality Enhancement \(RME\) module leverages reliable samples to enhance representations \(see Section[3\.4](https://arxiv.org/html/2609.30470#S3.SS4)\)\. Finally, the Hyper\-Modality Generation and Multilevel Fusion module generates hyper\-modality representations and fuses them with original representations via reliability\-aware multilevel fusion \(see Section[3\.5](https://arxiv.org/html/2609.30470#S3.SS5)\)\.

### 3\.1Task Formulation

MSA aims to infer the speaker’s emotional state from video segments comprising textual \(XTX\_\{T\}\), acoustic \(XAX\_\{A\}\), and visual \(XVX\_\{V\}\) modalities\. Each modality is represented as a sequence of features, i\.e\.,Xm∈ℝLm×dmX\_\{m\}\\in\\mathbb\{R\}^\{L\_\{m\}\\times d\_\{m\}\}, wherem∈\{T,A,V\}m\\in\\\{T,A,V\\\}indexes the modality\. Here,LmL\_\{m\}anddmd\_\{m\}denote the sequence length and feature dimension of modalitymm, respectively\. Given a training dataset𝒟=\{\(XTi,XAi,XVi,yi\)\}i=1N\\mathcal\{D\}=\\\{\(X\_\{T\}^\{i\},X\_\{A\}^\{i\},X\_\{V\}^\{i\},y^\{i\}\)\\\}\_\{i=1\}^\{N\}, the objective is to learn a modelℳϕ\\mathcal\{M\}\_\{\\phi\}that predictsy^i=ℳϕ​\(XTi,XAi,XVi\)\\hat\{y\}^\{i\}=\\mathcal\{M\}\_\{\\phi\}\(X\_\{T\}^\{i\},X\_\{A\}^\{i\},X\_\{V\}^\{i\}\), whereϕ\\phidenotes the learnable parameters\. Depending on the task setting, the predictiony^i\\hat\{y\}^\{i\}can be either a continuous sentiment score \(regression\) or a discrete emotion category \(classification\)\.

### 3\.2Pre\-Processing

To map heterogeneous multimodal input features into a unified representation space, following prior work[Tsai et al\. \(2019\)](https://arxiv.org/html/2609.30470#bib.bib15);[Li et al\. \(2024b\)](https://arxiv.org/html/2609.30470#bib.bib37);[Zhuang et al\. \(2025b\)](https://arxiv.org/html/2609.30470#bib.bib43), we employ one\-dimensional convolutions to project features from each modality, yieldingX^mi=W1​Dm​\(Xmi\)\\hat\{X\}\_\{m\}^\{i\}=W\_\{1D\}^\{m\}\(X\_\{m\}^\{i\}\), whereW1​DmW\_\{1D\}^\{m\}denotes the modality\-specific one\-dimensional convolutional projection layer andX^mi∈ℝLm×d\\hat\{X\}\_\{m\}^\{i\}\\in\\mathbb\{R\}^\{L\_\{m\}\\times d\}, whereddis the shared embedding dimension across modalities\. Subsequently, the projected feature sequences are fed into a Transformer[Vaswani et al\. \(2017\)](https://arxiv.org/html/2609.30470#bib.bib8)encoder to model intra\-modal temporal dependencies\. Positional encoding is added to preserve sequential order information, yielding modality\-specific representationsSmi=Fϕm​\(X^mi\+P​E​\(Lm,d\)\)S\_\{m\}^\{i\}=F\_\{\\phi\_\{m\}\}\(\\hat\{X\}\_\{m\}^\{i\}\+PE\(L\_\{m\},d\)\), whereFϕmF\_\{\\phi\_\{m\}\}denotes the modality\-specific Transformer encoder with parametersϕm\\phi\_\{m\}, andP​E​\(Lm,d\)PE\(L\_\{m\},d\)denotes positional encoding of lengthLmL\_\{m\}and dimensiondd\. Finally, global max pooling is applied to obtain the initial global feature representation of each modality, i\.e\.,Hmi=G​M​P​\(Smi\)H\_\{m\}^\{i\}=GMP\(S\_\{m\}^\{i\}\), whereHmi∈ℝdH\_\{m\}^\{i\}\\in\\mathbb\{R\}^\{d\}andG​M​P​\(⋅\)GMP\(\\cdot\)denotes the global max pooling operation\.

### 3\.3Adaptive Variational Information Bottleneck

To achieve adaptive compression of information conditioned on modality quality, we propose the AVIB module, consisting of two key components: vMF\-based uncertainty estimation and quality\-aware adaptive compression\. We describe each component in detail below\.

#### vMF\-based Uncertainty Estimation\.

To explicitly model sample\-level reliability across modalities, we map each modality’s feature representation to the von Mises–Fisher \(vMF\) distribution on the unit hypersphere\. For theii\-th sample and modalitymm,HmiH\_\{m\}^\{i\}parameterizes a posterior distribution overzz, i\.e\.,zmi∼vMF⁡\(μmi,κmi\)z\_\{m\}^\{i\}\\sim\\mathrm\{vMF\}\(\\mu\_\{m\}^\{i\},\\kappa\_\{m\}^\{i\}\), whereμmi\\mu\_\{m\}^\{i\}andκmi\\kappa\_\{m\}^\{i\}denote the mean direction and concentration parameter, respectively\. Specifically, we define:

μmi=fm​\(Hmi\)‖fm​\(Hmi\)‖2∈𝕊d−1,κmi=softplus⁡\(gm​\(Hmi\)\)\+ϵ∈ℝ\+,\\mu\_\{m\}^\{i\}=\\frac\{f\_\{m\}\(H\_\{m\}^\{i\}\)\}\{\\\|f\_\{m\}\(H\_\{m\}^\{i\}\)\\\|\_\{2\}\}\\in\\mathbb\{S\}^\{d\-1\},\\quad\\kappa\_\{m\}^\{i\}=\\mathrm\{softplus\}\(g\_\{m\}\(H\_\{m\}^\{i\}\)\)\+\\epsilon\\in\\mathbb\{R\}^\{\+\},\(1\)wherefm​\(⋅\)f\_\{m\}\(\\cdot\)andgm​\(⋅\)g\_\{m\}\(\\cdot\)are two separate subnetworks used to estimate the directional and concentration parameters, respectively, andϵ\\epsilonis a small constant for numerical stability\. Under this formulation,μmi\\mu\_\{m\}^\{i\}represents the principal direction of the latent representation on the unit hypersphere, whileκmi\\kappa\_\{m\}^\{i\}characterizes the degree of concentration of the distribution around this direction\. A largerκmi\\kappa\_\{m\}^\{i\}corresponds to a more concentrated posterior distribution\. Following prior work[Ming et al\. \(2023\)](https://arxiv.org/html/2609.30470#bib.bib63);[Du et al\. \(2024\)](https://arxiv.org/html/2609.30470#bib.bib62);[Hu et al\. \(2025\)](https://arxiv.org/html/2609.30470#bib.bib61), we interpret it as a proxy for sample\-level reliability of the corresponding modality, reflecting how trustworthy the modality is for prediction\. Additional details are provided in Appendix[C\.1](https://arxiv.org/html/2609.30470#A3.SS1)\.

#### Quality\-aware Adaptive Compression\.

Building on the vMF representation, we model directional information and sample\-wise posterior uncertainty within a unified hyperspherical space, enabling modality\-quality\-aware adaptive compression\.

We define the variational posterior asq⁡\(zmi∣Hmi\)=vMF⁡\(μmi,κmi\)q\(z\_\{m\}^\{i\}\\mid H\_\{m\}^\{i\}\)=\\mathrm\{vMF\}\(\\mu\_\{m\}^\{i\},\\kappa\_\{m\}^\{i\}\), whereμmi\\mu\_\{m\}^\{i\}andκmi\\kappa\_\{m\}^\{i\}are estimated from modality\-specific featuresHmiH\_\{m\}^\{i\}\. Our goal is to learn a compressed latent representationzmiz\_\{m\}^\{i\}that preserves task\-relevant information while discarding redundant and noisy components in the original representation\. The corresponding information bottleneck objective is:

ℒVIBm=β​I​\(Hmi,zmi\)−I⁡\(zmi,Y\),\\mathcal\{L\}\_\{\\text\{VIB\}\}^\{m\}=\\beta I\(H\_\{m\}^\{i\};z\_\{m\}^\{i\}\)\-I\(z\_\{m\}^\{i\};Y\),\(2\)whereβ\\betais a hyperparameter controlling the trade\-off between compression and predictive power\.

Since the mutual informationI⁡\(Hmi,zmi\)I\(H\_\{m\}^\{i\};z\_\{m\}^\{i\}\)is intractable, we use variational inference to approximate it\. In the absence of additional prior knowledge and to preserve directional isotropy, we adopt a uniform distribution on the unit hypersphere as the prior:p⁡\(zmi\)=Uniform⁡\(𝕊d−1\)p\(z\_\{m\}^\{i\}\)=\\mathrm\{Uniform\}\(\\mathbb\{S\}^\{d\-1\}\)\. Accordingly, the original objective can be reformulated as the following tractable variational objective:

ℒAVIBm=β⋅KL\(q\(zmi∣Hmi\)∥p\(zmi\)\)−𝔼zmi∼q\[logp\(Y∣zmi\)\],\\mathcal\{L\}\_\{\\text\{AVIB\}\}^\{m\}=\\beta\\cdot\\mathrm\{KL\}\\Big\(q\(z\_\{m\}^\{i\}\\mid H\_\{m\}^\{i\}\)\\;\\\|\\;p\(z\_\{m\}^\{i\}\)\\Big\)\-\\mathbb\{E\}\_\{z\_\{m\}^\{i\}\\sim q\}\\big\[\\log p\(Y\\mid z\_\{m\}^\{i\}\)\\big\],\(3\)where the expectation term𝔼zmi∼q​\[log⁡p⁡\(Y∣zmi\)\]\\mathbb\{E\}\_\{z\_\{m\}^\{i\}\\sim q\}\\big\[\\log p\(Y\\mid z\_\{m\}^\{i\}\)\\big\]can be implemented via the task lossℒtask​\(y^m,y\)\\mathcal\{L\}\_\{\\text\{task\}\}\(\\hat\{y\}\_\{m\},y\)\.

To enable differentiable sampling on the unit hypersphere, we adopt an approximate sampling strategy based on direction–orthogonal decomposition\. Given vMF parameters\(μ,κ\)\(\\mu,\\kappa\), the latent variable is constructed as:

z=w​μ\+1−w2​v,z=w\\mu\+\\sqrt\{1\-w^\{2\}\}\\,v,\(4\)wherew∈\[−1,1\]w\\in\[\-1,1\]is a scalar random variable governed byκ\\kappa, andv∼Uniform⁡\(𝕊d−2\)v\\sim\\mathrm\{Uniform\}\(\\mathbb\{S\}^\{d\-2\}\)is a unit vector sampled from the subspace orthogonal toμ\\mu\. This construction ensures thatzzlies on the unit hypersphere𝕊d−1\\mathbb\{S\}^\{d\-1\}\. Under this formulation, the KL divergence admits a closed\-form expression:

KL\(vMF\(μ,κ\)∥Uniform\)=κId/2​\(κ\)Id/2−1​\(κ\)\+logCd\(κ\)−logCd\(0\),\\mathrm\{KL\}\(\\mathrm\{vMF\}\(\\mu,\\kappa\)\\;\\\|\\;\\mathrm\{Uniform\}\)=\\kappa\\frac\{I\_\{d/2\}\(\\kappa\)\}\{I\_\{d/2\-1\}\(\\kappa\)\}\+\\log C\_\{d\}\(\\kappa\)\-\\log C\_\{d\}\(0\),\(5\)whenκ=0\\kappa=0, the vMF distribution reduces to the uniform distribution on the unit hypersphere, with normalization constantCd​\(0\)=Γ⁡\(d/2\)2​πd/2C\_\{d\}\(0\)=\\frac\{\\Gamma\(d/2\)\}\{2\\pi^\{d/2\}\}\.

During training, we introduce auxiliary prediction branches based on modality\-specific latent samples and jointly optimize the task loss with the KL regularization\. This mechanism assigns a higher information cost to highly concentrated posteriors, encouraging the model to retain them only when the modality provides discriminative information\. Conversely, low\-quality or noisy modalities are driven toward more dispersed posteriors approaching the uniform prior, reducing their impact on downstream predictions\. In this way, AVIB achieves adaptive information compression conditioned on modality quality\.

### 3\.4Reliability\-aware Modality Enhancement

Inspired by HME[Zhuang et al\. \(2025b\)](https://arxiv.org/html/2609.30470#bib.bib43), we enrich each modality by retrieving semantically relevant cues from other samples\. However, HME is prone to “negative enhancement” and relies solely on mini\-batch samples, leading to insufficient and unstable neighborhood information\. To address these limitations, we propose the RME module\. Unlike methods that rely only on mini\-batch samples, RME maintains a modality\-specific memory bank of historical samples, enabling the retrieval of more informative and semantically similar neighbors from a larger candidate pool\.

Specifically, for modalitymm, letzmi∈ℝdz\_\{m\}^\{i\}\\in\\mathbb\{R\}^\{d\}denote the latent representation of theii\-th sample in the current mini\-batch obtained from the AVIB module, with corresponding vMF mean directionμmi\\mu\_\{m\}^\{i\}and concentration parameterκmi\\kappa\_\{m\}^\{i\}\. The batch\-level representations are denoted asZm\(b\)=\{zm1,…,zmB\}Z\_\{m\}^\{\(b\)\}=\\\{z\_\{m\}^\{1\},\\dots,z\_\{m\}^\{B\}\\\},Mm\(b\)=\{μm1,…,μmB\}M\_\{m\}^\{\(b\)\}=\\\{\\mu\_\{m\}^\{1\},\\dots,\\mu\_\{m\}^\{B\}\\\}, andKm\(b\)=\{κm1,…,κmB\}K\_\{m\}^\{\(b\)\}=\\\{\\kappa\_\{m\}^\{1\},\\dots,\\kappa\_\{m\}^\{B\}\\\}, whereBBis the batch size\. We maintain a modality\-specific memory bank with capacityNbN\_\{b\}:ℬm=\{\(zmr,μmr,κmr,yr\)\}r=1Nm\\mathcal\{B\}\_\{m\}=\\\{\(z\_\{m\}^\{r\},\\mu\_\{m\}^\{r\},\\kappa\_\{m\}^\{r\},y^\{r\}\)\\\}\_\{r=1\}^\{N\_\{m\}\}withNm≤NbN\_\{m\}\\leq N\_\{b\}, whereyry^\{r\}is the label\. For a given sampleii, its candidate neighbor set is defined as:

𝒞im=\{ℬm,if​Nm=Nb,\{\(zmj,μmj,κmj,yj\)∣j≠i\},o​t​h​e​r​w​i​s​e\.\\mathcal\{C\}\_\{i\}^\{m\}=\\begin\{cases\}\\mathcal\{B\}\_\{m\},&\\text\{if \}N\_\{m\}=N\_\{b\},\\\\ \\\{\(z\_\{m\}^\{j\},\\mu\_\{m\}^\{j\},\\kappa\_\{m\}^\{j\},y^\{j\}\)\\mid j\\neq i\\\},&otherwise\.\\end\{cases\}\(6\)In practice, the memory bank is rebuilt before each training epoch by a forward pass over the training set without gradient updates\. The model retrieves neighbors from the memory bank only when it is fully populated, providing a broader and more stable candidate space; otherwise, it falls back to intra\-batch retrieval\. This design avoids unreliable neighbors in early training while enabling more effective cross\-sample complementarity once sufficient historical samples are accumulated\.

Notably, to avoid recency bias and ensure balanced retention under limited capacity, we use reservoir sampling[Vitter \(1985\)](https://arxiv.org/html/2609.30470#bib.bib68)for memory updates, such that after observingttsamples, each sample is retained with probabilitymin⁡\(1,Nb/t\)\\min\(1,N\_\{b\}/t\)\. More details are provided in Appendix[C\.2](https://arxiv.org/html/2609.30470#A3.SS2)\.

To improve the stability of neighbor retrieval, we do not compute similarity directly on the sampled representationszmiz\_\{m\}^\{i\}, but instead perform matching in a more stable reference direction space\. Specifically, for a candidate samplejjin𝒞im\\mathcal\{C\}\_\{i\}^\{m\}, the similarity is defined asSm​\(i,j\)=μmi⊤​μmjS\_\{m\}\(i,j\)=\{\\mu\_\{m\}^\{i\}\}^\{\\top\}\\mu\_\{m\}^\{j\}, whereμmi,μmj\\mu\_\{m\}^\{i\},\\mu\_\{m\}^\{j\}are unit vectors on the hypersphere\. Thus, this is equivalent to cosine similarity[Li et al\. \(2024b\)](https://arxiv.org/html/2609.30470#bib.bib37)\. This design effectively reduces fluctuations caused by sampling noise, making the similarity measure more robust in reflecting semantic consistency between samples\. Based on this, we select only the top\-kkmost similar neighbors to form the neighborhood set𝒩im\\mathcal\{N\}\_\{i\}^\{m\}, i\.e\.,

𝒩im=Top\-​k​\(𝒞im\),\\mathcal\{N\}\_\{i\}^\{m\}=\\text\{Top\-\}k\(\\mathcal\{C\}\_\{i\}^\{m\}\),\(7\)wherek=⌊ρ​\|𝒞im\|⌋k=\\lfloor\\rho\|\\mathcal\{C\}\_\{i\}^\{m\}\|\\rfloor, andρ∈\(0,1\]\\rho\\in\(0,1\]is a proportion hyperparameter\. This design helps filter out weakly related candidates and reduces the introduction of irrelevant noise\.

To improve the reliability of neighbor aggregation, we incorporate both label distance and confidence constraints into the similarity weighting\. For any neighborj∈𝒩imj\\in\\mathcal\{N\}\_\{i\}^\{m\}, its weight is defined as

ei​j\(m\)=Sm​\(i,j\)−α​\|yi−yj\|\+log⁡\(1\+κmj\),e\_\{ij\}^\{\(m\)\}=S\_\{m\}\(i,j\)\-\\alpha\|y\_\{i\}\-y\_\{j\}\|\+\\log\(1\+\\kappa\_\{m\}^\{j\}\),\(8\)whereyiy\_\{i\}andyjy\_\{j\}are the labels of the current sample and its neighbor, respectively, andα\\alphacontrols the penalty for label discrepancy\. The similarity term captures semantic proximity, the label distance penalizes mismatched labels, andlog⁡\(1\+κmj\)\\log\(1\+\\kappa\_\{m\}^\{j\}\)favors high\-confidence neighbors\. The weights are normalized to obtain aggregation coefficients, and the neighbor\-enhanced representation is given by

z~mi=∑j∈𝒩imwi​j\(m\)​zmj,wi​j\(m\)=exp⁡\(ei​j\(m\)\)∑k∈𝒩imexp⁡\(ei​k\(m\)\)\.\\tilde\{z\}\_\{m\}^\{i\}=\\sum\_\{j\\in\\mathcal\{N\}\_\{i\}^\{m\}\}w\_\{ij\}^\{\(m\)\}z\_\{m\}^\{j\},\\quad w\_\{ij\}^\{\(m\)\}=\\frac\{\\exp\(e\_\{ij\}^\{\(m\)\}\)\}\{\\sum\_\{k\\in\\mathcal\{N\}\_\{i\}^\{m\}\}\\exp\(e\_\{ik\}^\{\(m\)\}\)\}\.\(9\)If no valid neighbors are available, we setz~mi=zmi\\tilde\{z\}\_\{m\}^\{i\}=z\_\{m\}^\{i\}\. This mechanism promotes aggregation from semantically similar, label\-consistent, and high\-confidence samples, leading to more robust modality enhancement\.

Notably, the RME module is used only during training and is not required at inference time\. It can also be applied to incomplete data by retrieving semantically relevant cues from available samples, avoiding explicit modality reconstruction and remaining effective under missing modalities\.

### 3\.5Hyper\-Modality Generation and Multilevel Fusion

#### Hyper\-Modality Representation Generation\.

Next, following HME[Zhuang et al\. \(2025b\)](https://arxiv.org/html/2609.30470#bib.bib43), we employ a Perceiver\-style latent Transformer[Zhang et al\. \(2023a\)](https://arxiv.org/html/2609.30470#bib.bib29);[Zhang et al\. \(2024\)](https://arxiv.org/html/2609.30470#bib.bib40)to encode the modality\-enhanced representations from the RME module, aiming to obtain compact and context\-aware modality\-level representations\. Specifically, for each modalitymm, we introduce a set of learnable latent promptsEm′∈ℝlp×dE\_\{m\}^\{\\prime\}\\in\\mathbb\{R\}^\{l\_\{p\}\\times d\}, wherelpl\_\{p\}is the prompt length andddis the feature dimension\. TakingEm′E\_\{m\}^\{\\prime\}as queries and the enhanced representationz~mi\\tilde\{z\}\_\{m\}^\{i\}as keys and values, we apply cross\-attention:

E¯mi=A​T​T​Nψm​\(Em′,z~mi\)=Softmax​\(Em′​WmQ​\(z~mi​WmK\)⊤d\)​z~mi​WmV,\\bar\{E\}\_\{m\}^\{i\}=ATTN\_\{\\psi\_\{m\}\}\(E\_\{m\}^\{\\prime\},\\tilde\{z\}\_\{m\}^\{i\}\)=\\text\{Softmax\}\\left\(\\frac\{E\_\{m\}^\{\\prime\}W\_\{m\}^\{Q\}\(\\tilde\{z\}\_\{m\}^\{i\}W\_\{m\}^\{K\}\)^\{\\top\}\}\{\\sqrt\{d\}\}\\right\)\\tilde\{z\}\_\{m\}^\{i\}W\_\{m\}^\{V\},\(10\)whereA​T​T​Nψm​\(⋅\)ATTN\_\{\\psi\_\{m\}\}\(\\cdot\)denotes the cross\-attention function parameterized byψm\\psi\_\{m\}, andWmQ,WmK,WmV∈ℝd×dW\_\{m\}^\{Q\},W\_\{m\}^\{K\},W\_\{m\}^\{V\}\\in\\mathbb\{R\}^\{d\\times d\}are learnable projection matrices\. After several standard Transformer blocks, we apply global average pooling over the latent prompt dimension to obtain the modality\-level representation:Emi=G​A​P​\(E¯mi\)E\_\{m\}^\{i\}=GAP\(\\bar\{E\}\_\{m\}^\{i\}\), whereEmi∈ℝdE\_\{m\}^\{i\}\\in\\mathbb\{R\}^\{d\}andG​A​P​\(⋅\)GAP\(\\cdot\)denotes the global average pooling operation\.

Furthermore, to model shared multimodal context, we introduce a shared Perceiver\-style Latent Transformer\. Unlike the modality\-specific branch, this branch takes the original unimodal representations of the current sample as input\. GivenHTi,HAi,HViH\_\{T\}^\{i\},H\_\{A\}^\{i\},H\_\{V\}^\{i\}, we constructHSi=\[HTi;HAi;HVi\]∈ℝ3×dH\_\{S\}^\{i\}=\[H\_\{T\}^\{i\};H\_\{A\}^\{i\};H\_\{V\}^\{i\}\]\\in\\mathbb\{R\}^\{3\\times d\}and apply the same latent cross\-attention mechanism to extract global contextual information, yielding the shared representation:ESi=LatentTransformerS​\(HSi\)E\_\{S\}^\{i\}=\\text\{LatentTransformer\}\_\{S\}\(H\_\{S\}^\{i\}\),ESi∈ℝdE\_\{S\}^\{i\}\\in\\mathbb\{R\}^\{d\}\.

Subsequently, we leverage cross\-modal complementarity to construct hyper\-modality representations\. For anym,m1,m2∈\{T,A,V\}m,m\_\{1\},m\_\{2\}\\in\\\{T,A,V\\\}withm≠m1m\\neq m\_\{1\},m≠m2m\\neq m\_\{2\}, andm1≠m2m\_\{1\}\\neq m\_\{2\}, we first model inter\-modal interactions via cross\-attention:

Cm→m1i=A​T​T​Nψm,m1​\(Emi,Em1i\),Cm→m2i=A​T​T​Nψm,m2​\(Emi,Em2i\)\.C\_\{m\\to m\_\{1\}\}^\{i\}=ATTN\_\{\\psi\_\{m,m\_\{1\}\}\}\(E\_\{m\}^\{i\},E\_\{m\_\{1\}\}^\{i\}\),\\quad C\_\{m\\to m\_\{2\}\}^\{i\}=ATTN\_\{\\psi\_\{m,m\_\{2\}\}\}\(E\_\{m\}^\{i\},E\_\{m\_\{2\}\}^\{i\}\)\.\(11\)Then, using the shared representationESiE\_\{S\}^\{i\}as the query and the interaction features from other modalities as the keys and values, we derive the hyper\-modality representation:

Rmi=A​T​T​NψS,m​\(ESi,\[Cm→m1i;Cm→m2i\]\),R\_\{m\}^\{i\}=ATTN\_\{\\psi\_\{S,m\}\}\(E\_\{S\}^\{i\},\[C\_\{m\\to m\_\{1\}\}^\{i\};C\_\{m\\to m\_\{2\}\}^\{i\}\]\),\(12\)whereRmi∈ℝdR\_\{m\}^\{i\}\\in\\mathbb\{R\}^\{d\}\. This allows each modality to absorb complementary information from others while incorporating global shared context, leading to more robust cross\-modal representations\.

#### Reliability\-aware Multilevel Fusion\.

To integrate multi\-granularity information across stages, we design a hierarchical aggregation mechanism from intra\-modality fusion to cross\-modality weighting\.

For each modalitym∈\{T,A,V\}m\\in\\\{T,A,V\\\}, we first fuse the original unimodal representationHmiH\_\{m\}^\{i\}with the hyper\-modality representationRmiR\_\{m\}^\{i\}under shared context\. Specifically, we concatenate them and pass them through a lightweight attention network to obtain fusion logits:ami=Fm\(\[Hmi∥Rmi\]\)a\_\{m\}^\{i\}=F\_\{m\}\(\[H\_\{m\}^\{i\}\\\|R\_\{m\}^\{i\}\]\), whereFm​\(⋅\)F\_\{m\}\(\\cdot\)denotes a lightweight mapping function composed of a linear layer, Dropout, and ReLU\. Since the concentration parameterκmi\\kappa\_\{m\}^\{i\}in the vMF distribution reflects representation uncertainty, we use it to adaptively modulate the fusion weights, yielding the intra\-modality fused representation:

umi=βm,1i​Hmi\+βm,2i​Rmi,βmi=Softmax​\(ami⊙\[1,κmi\]d\),u\_\{m\}^\{i\}=\\beta\_\{m,1\}^\{i\}H\_\{m\}^\{i\}\+\\beta\_\{m,2\}^\{i\}R\_\{m\}^\{i\},\\quad\\beta\_\{m\}^\{i\}=\\text\{Softmax\}\\left\(\\frac\{a\_\{m\}^\{i\}\\odot\[1,\\kappa\_\{m\}^\{i\}\]\}\{\\sqrt\{d\}\}\\right\),\(13\)where⊙\\odotdenotes element\-wise multiplication andddis the feature dimension\. This allows adaptive balancing between unimodal and enhanced representations based on modality reliability\.

To obtain a global representation, we further perform reliability\-aware fusion overRTi,RAi,RViR\_\{T\}^\{i\},R\_\{A\}^\{i\},R\_\{V\}^\{i\}\. We compute fusion logits via concatenation:aHi=FH​\(\[RTi​‖RAi‖​RVi\]\)∈ℝ3a\_\{H\}^\{i\}=F\_\{H\}\(\[R\_\{T\}^\{i\}\\\|R\_\{A\}^\{i\}\\\|R\_\{V\}^\{i\}\]\)\\in\\mathbb\{R\}^\{3\}\. The modality reliability weights are defined asγi=Softmax​\(\[κTi,κAi,κVi\]\)\\gamma^\{i\}=\\text\{Softmax\}\(\[\\kappa\_\{T\}^\{i\},\\kappa\_\{A\}^\{i\},\\kappa\_\{V\}^\{i\}\]\)\. The logits are then modulated and normalized:ηi=Softmax​\(aHi⊙γid\)\\eta^\{i\}=\\text\{Softmax\}\\left\(\\frac\{a\_\{H\}^\{i\}\\odot\\gamma^\{i\}\}\{\\sqrt\{d\}\}\\right\)\. The resulting global representation is

uHi=ηTi​RTi\+ηAi​RAi\+ηVi​RVi\.u\_\{H\}^\{i\}=\\eta\_\{T\}^\{i\}R\_\{T\}^\{i\}\+\\eta\_\{A\}^\{i\}R\_\{A\}^\{i\}\+\\eta\_\{V\}^\{i\}R\_\{V\}^\{i\}\.\(14\)
Finally, we form four tokens by combining the three intra\-modality representations and the global representation:Ui=\[uTi;uAi;uVi;uHi\]∈ℝ4×dU^\{i\}=\[u\_\{T\}^\{i\};u\_\{A\}^\{i\};u\_\{V\}^\{i\};u\_\{H\}^\{i\}\]\\in\\mathbb\{R\}^\{4\\times d\}, which are fed into a two\-layer Transformer encoder to model cross\-level interactions:U^i=TransformerEncoder​\(Ui\)\\hat\{U\}^\{i\}=\\text\{TransformerEncoder\}\(U^\{i\}\)\. The output tokens are concatenated to form the final joint representation:ri=\[u^Ti∥u^Ai∥u^Vi∥u^Hi\]∈ℝ4​dr^\{i\}=\[\\hat\{u\}\_\{T\}^\{i\}\\\|\\hat\{u\}\_\{A\}^\{i\}\\\|\\hat\{u\}\_\{V\}^\{i\}\\\|\\hat\{u\}\_\{H\}^\{i\}\]\\in\\mathbb\{R\}^\{4d\}\. The final representationrir^\{i\}is then passed to a two\-layer MLP to produce the prediction:y^i=MLP​\(ri\)\\hat\{y\}^\{i\}=\\text\{MLP\}\(r^\{i\}\)\.

#### Overall Objective\.

To jointly optimize task performance and latent representation compression, we define the overall objective as

ℒa​l​l=ℒt​a​s​k​\(y^,y\)\+λ⋅ℒAVIB,\\mathcal\{L\}\_\{all\}=\\mathcal\{L\}\_\{task\}\(\\hat\{y\},y\)\+\\lambda\\cdot\\mathcal\{L\}\_\{\\text\{AVIB\}\},\(15\)whereℒAVIB=13​∑m∈\{T,A,V\}ℒAVIBm\\mathcal\{L\}\_\{\\text\{AVIB\}\}=\\frac\{1\}\{3\}\\sum\_\{m\\in\\\{T,A,V\\\}\}\\mathcal\{L\}\_\{\\text\{AVIB\}\}^\{m\}, andλ\\lambdais a hyperparameter that balances the task loss and the information bottleneck regularization\. For classification and regression tasks,ℒt​a​s​k\\mathcal\{L\}\_\{task\}is defined as

ℒt​a​s​k=\{C​r​o​s​s​E​n​t​r​o​p​y​\(y^,y\),c​l​a​s​s​i​f​i​c​a​t​i​o​n\|y^−y\|,r​e​g​r​e​s​s​i​o​n\.\\mathcal\{L\}\_\{task\}=\\begin\{cases\}CrossEntropy\(\\hat\{y\},y\),&classification\\\\ \|\\hat\{y\}\-y\|,&regression\\end\{cases\}\.\(16\)

## 4Experiments

### 4\.1Experimental Setup

We evaluate RCE on MSA under full, noisy, and missing\-modality settings, where both noisy and missing modalities are evaluated under two protocols\. To further assess generalization, we also evaluate it on multimodal humor detection \(MHD\) and multimodal sarcasm detection \(MSD\) tasks\.

Datasets\.For the MSA task, we conduct experiments on two widely used datasets, CMU\-MOSI[Zadeh et al\. \(2016\)](https://arxiv.org/html/2609.30470#bib.bib4)and CMU\-MOSEI[Zadeh et al\. \(2018b\)](https://arxiv.org/html/2609.30470#bib.bib5)\. For the MHD and MSD tasks, we adopt the UR\-FUNNY[Hasan et al\. \(2019\)](https://arxiv.org/html/2609.30470#bib.bib3)and MUStARD[Castro et al\. \(2019\)](https://arxiv.org/html/2609.30470#bib.bib2)datasets, respectively\. Details are provided in Appendix[D\.1](https://arxiv.org/html/2609.30470#A4.SS1)\.

Baselines\.We compare RCE with several state\-of\-the\-art methods, including recent full\-modality approaches such as DecAlign[Qian et al\. \(2026\)](https://arxiv.org/html/2609.30470#bib.bib56)\[ICLR’26\], MODS[Yang et al\. \(2026\)](https://arxiv.org/html/2609.30470#bib.bib58)\[AAAI’26\], and PSA\-MF[Xie et al\. \(2026\)](https://arxiv.org/html/2609.30470#bib.bib60)\[AAAI’26\]\. In addition, we focus on two representative categories of methods that are most relevant to RCE\. The first category consists of information bottleneck\-based approaches, including C\-MIB[Mai et al\. \(2023b\)](https://arxiv.org/html/2609.30470#bib.bib30)\[TMM’23\], KAN\-MCP[Luo et al\. \(2025\)](https://arxiv.org/html/2609.30470#bib.bib51)\[ACM MM’25\], OMIB[Wu et al\. \(2025a\)](https://arxiv.org/html/2609.30470#bib.bib45)\[ICML’25\], and CyIN[Lin et al\. \(2025\)](https://arxiv.org/html/2609.30470#bib.bib44)\[NeurIPS’25\]\. The second category includes uncertainty\-aware fusion methods such as QMF[Zhang et al\. \(2023b\)](https://arxiv.org/html/2609.30470#bib.bib33)\[ICML’23\], PML[Hu et al\. \(2025\)](https://arxiv.org/html/2609.30470#bib.bib61)\[IJCAI’25\], and HME[Zhuang et al\. \(2025b\)](https://arxiv.org/html/2609.30470#bib.bib43)\[NeurIPS’25\]\. Detailed descriptions of these methods are provided in Appendix[D\.2](https://arxiv.org/html/2609.30470#A4.SS2)\.

Due to space limitations, details on feature extraction, implementation, and evaluation metrics are provided in Appendix[D\.3](https://arxiv.org/html/2609.30470#A4.SS3), Appendix[D\.4](https://arxiv.org/html/2609.30470#A4.SS4), and Appendix[D\.5](https://arxiv.org/html/2609.30470#A4.SS5), respectively\. In addition, extensive additional experimental results are provided in Appendix[E](https://arxiv.org/html/2609.30470#A5)for a more comprehensive analysis\.

### 4\.2Quantitative Results

Table 1:Comparison on CMU\-MOSI and CMU\-MOSEI\.BestandSecondresults are highlighted\.Full\-modality Setting\.The results under the fully multimodal setting are reported in Table[1](https://arxiv.org/html/2609.30470#S4.T1)\. This setting is adopted by most existing MSA methods\. It can be observed that, compared with state\-of\-the\-art approaches such as DecAlign[Qian et al\. \(2026\)](https://arxiv.org/html/2609.30470#bib.bib56)and MODS[Yang et al\. \(2026\)](https://arxiv.org/html/2609.30470#bib.bib58), RCE achieves the best performance on nearly all metrics\. In particular, RCE significantly outperforms existing methods onAcc7andMAE, highlighting its capability for fine\-grained sentiment prediction\. This advantage stems from its reliability\-aware compression, cross\-sample enhancement, and multilevel fusion, which collectively facilitate the learning of robust and discriminative multimodal representations\.

Missing\-modality Setting\.We compare RCE with state\-of\-the\-art methods specifically designed for missing modality scenarios, including HME[Zhuang et al\. \(2025b\)](https://arxiv.org/html/2609.30470#bib.bib43)and CyIN[Lin et al\. \(2025\)](https://arxiv.org/html/2609.30470#bib.bib44), under the random missing protocol\. The results \(see Figure[2](https://arxiv.org/html/2609.30470#S4.F2)\) show that RCE consistently outperforms these methods across most missing rates\. Notably, even without modality\-specific tuning, RCE demonstrates strong generalization ability\. We further provide comparisons under fixed missing protocols against a variety of advanced methods, with results reported in Appendix[E\.1](https://arxiv.org/html/2609.30470#A5.SS1)\. Overall, RCE significantly enhances robustness to diverse missing patterns by leveraging informative cues from other reliable and available samples\.

Figure 2:Comparison on CMU\-MOSEI under the random missing protocol\.MHD and MSD Tasks\.To further evaluate the generalization capability of RCE on other multimodal tasks, we conduct experiments on the MHD and MSD tasks using the UR\-FUNNY[Hasan et al\. \(2019\)](https://arxiv.org/html/2609.30470#bib.bib3)and MUStARD[Castro et al\. \(2019\)](https://arxiv.org/html/2609.30470#bib.bib2)datasets\. As shown in Figure[4](https://arxiv.org/html/2609.30470#S4.F4), RCE outperforms the best\-performing baselines, AtCAF[Huang et al\. \(2025\)](https://arxiv.org/html/2609.30470#bib.bib49)and MOAC[Mai et al\. \(2025\)](https://arxiv.org/html/2609.30470#bib.bib47), by over 3\.2% and 1\.5% in accuracy, respectively, achieving state\-of\-the\-art performance\. These results further demonstrate the strong generalization ability of RCE across different multimodal tasks\.

### 4\.3Ablation Study

To better understand the contribution of each component in RCE, we conduct an ablation study on several key modules of the proposed framework: \(i\) ‘w/o AVIB’ removes the AVIB module and directly uses the modality representations without adaptive compression; \(ii\) ‘w/o RME’ removes the RME module; \(iii\) ‘w/o MB’ disables the memory bank and only uses mini\-batch samples for neighbor retrieval; \(iv\) ‘w/o CP’ removes the confidence promotion term in neighbor weighting; \(v\) ‘w/o HMG’ removes the hyper\-modality generation component; and \(vi\) ‘w/o RF’ removes reliability\-aware fusion and replaces it with standard attention weighting\. The results are reported in Table[2](https://arxiv.org/html/2609.30470#S4.T2)\. Full results and additional variant analyses are provided in Appendix[E\.3](https://arxiv.org/html/2609.30470#A5.SS3)\.

Effects of the RCE Components\.Overall, removing most components degrades performance on both CMU\-MOSI and CMU\-MOSEI, demonstrating the effectiveness of the proposed modules\. Among them, ‘w/o HMG’ causes the largest drop on CMU\-MOSI, with Acc7 decreasing from 50\.22 to 46\.72 and MAE increasing from 0\.589 to 0\.637, highlighting the importance of hyper\-modality integration\. Removing AVIB also consistently harms performance, indicating that vMF\-based uncertainty estimation and adaptive compression effectively suppress noisy modality information\. In addition, ‘w/o MB’ yields lower Acc7 on both datasets, suggesting that retrieving reliable neighbors from the memory bank provides more stable enhancement than using only mini\-batch samples\. Some components exhibit dataset\- and metric\-dependent effects\. Specifically, ‘w/o CP’ causes only minor changes on CMU\-MOSEI, while ‘w/o RF’ slightly improves Acc7 but worsens MAE, suggesting that reliability\-aware fusion contributes more to prediction calibration than classification accuracy\. Although ‘w/o RME’ slightly improves Acc7 on CMU\-MOSI, it increases MAE and degrades performance on CMU\-MOSEI, indicating that RME generally improves robustness despite minor fluctuations on smaller datasets\. Overall, AVIB, MB, and HMG are the most influential components\.

Figure 3:Results on UR\-FUNNY and MUStARD\.Figure 4:Feature visualization\.
Table 2:Ablation results of RCE components on CMU\-MOSI and CMU\-MOSEI\. We reportAcc7andMAE\. Full results are provided in Appendix[E\.3](https://arxiv.org/html/2609.30470#A5.SS3)\.
### 4\.4Further Analysis

We provide additional analyses in the appendix, including parameter sensitivity analysis \(Appendix[E\.5](https://arxiv.org/html/2609.30470#A5.SS5)\), testing stability analysis \(Appendix[E\.6](https://arxiv.org/html/2609.30470#A5.SS6)\), model complexity analysis \(Appendix[E\.7](https://arxiv.org/html/2609.30470#A5.SS7)\), and theoretical complexity analysis \(Appendix[E\.8](https://arxiv.org/html/2609.30470#A5.SS8)\)\.

Feature Visualization\.To investigate the differences between the original modality representationsHmH\_\{m\}and the generated hyper\-modality representationsRmR\_\{m\}, we visualize them using t\-SNE[Van der Maaten and Hinton \(2008\)](https://arxiv.org/html/2609.30470#bib.bib13)\. As shown in Figure[4](https://arxiv.org/html/2609.30470#S4.F4), the hyper\-modality representations form distinct clusters, indicating that RCE captures complementary sentiment\-relevant information beyond unimodal features\.

Table 3:Training negative enhancement rates on CMU\-MOSI and CMU\-MOSEI\.Negative Enhancement Analysis\.As shown in Table[3](https://arxiv.org/html/2609.30470#S4.T3), we report the training\-stage negative enhancement rates of HME[Zhuang et al\. \(2025b\)](https://arxiv.org/html/2609.30470#bib.bib43)and RCE across text, audio, and vision modalities\. Compared with HME, RCE consistently reduces negative enhancement cases across all modalities, showing that the proposed reliability\-aware and label\-aware enhancement strategy effectively suppresses sentiment\-conflicting information from neighboring samples\. On CMU\-MOSI, the rates drop from around 46% to below 7%\. On CMU\-MOSEI, RCE also achieves clear reductions, especially in text, while the smaller improvement in audio may be due to the higher ambiguity of acoustic sentiment cues\. Detailed results and analysis are provided in Appendix[E\.4](https://arxiv.org/html/2609.30470#A5.SS4)\.

## 5Conclusion

In this paper, we propose Reliability\-aware Cross\-sample Enhancement \(RCE\), a unified framework for robust multimodal sentiment analysis under full, noisy, and missing\-modality settings\. RCE learns robust multimodal representations by estimating modality reliability with adaptive variational information bottleneck, retrieving reliable cross\-sample cues for modality enhancement, and performing reliability\-aware multilevel fusion for prediction\.

Limitations\.RCE introduces additional hyperparameters and a larger number of parameters\. Although it is generally stable within moderate parameter ranges, dataset\-specific tuning may still be needed\. Future work will explore adaptive and lightweight enhancement strategies\.

## References

- Castroet al\.\(2019\)S\. Castro, D\. Hazarika, V\. Pérez\-Rosas, R\. Zimmermann, R\. Mihalcea, and S\. PoriaTowards multimodal sarcasm detection \(an \_obviously\_ perfect paper\)\.InACL,pp\. 4619–4629\.Cited by:[§D\.1](https://arxiv.org/html/2609.30470#A4.SS1.p5.1),[§4\.1](https://arxiv.org/html/2609.30470#S4.SS1.p2.1),[§4\.2](https://arxiv.org/html/2609.30470#S4.SS2.p3.1)\.
- Degottexet al\.\(2014\)G\. Degottex, J\. Kane, T\. Drugman, T\. Raitio, and S\. SchererCOVAREP—a collaborative voice analysis repository for speech technologies\.InICASSP,pp\. 960–964\.Cited by:[§D\.3](https://arxiv.org/html/2609.30470#A4.SS3.p3.1)\.
- Devlinet al\.\(2019\)J\. Devlin, M\. Chang, K\. Lee, and K\. ToutanovaBert: pre\-training of deep bidirectional transformers for language understanding\.InNAACL,pp\. 4171–4186\.Cited by:[§D\.3](https://arxiv.org/html/2609.30470#A4.SS3.p1.1),[§D\.3](https://arxiv.org/html/2609.30470#A4.SS3.p2.1)\.
- Duet al\.\(2024\)C\. Du, Y\. Wang, S\. Song, and G\. HuangProbabilistic contrastive learning for long\-tailed visual recognition\.IEEE Transactions on Pattern Analysis and Machine Intelligence46\(9\),pp\. 5890–5904\.Cited by:[§C\.1](https://arxiv.org/html/2609.30470#A3.SS1.SSS0.Px1.p3.1),[§3\.3](https://arxiv.org/html/2609.30470#S3.SS3.SSS0.Px1.p1.2)\.
- Fanget al\.\(2025\)Y\. Fang, W\. Huang, G\. Wan, K\. Su, and M\. YeEMOE: modality\-specific enhanced dynamic emotion experts\.InCVPR,pp\. 14314–14324\.Cited by:[§1](https://arxiv.org/html/2609.30470#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.30470#S2.SS1.p1.1)\.
- Gaoet al\.\(2024\)Z\. Gao, X\. Jiang, X\. Xu, F\. Shen, Y\. Li, and H\. T\. ShenEmbracing unimodal aleatoric uncertainty for robust multimodal fusion\.InCVPR,pp\. 26876–26885\.Cited by:[§E\.2](https://arxiv.org/html/2609.30470#A5.SS2.SSS0.Px1.p1.1)\.
- Gonget al\.\(2024\)P\. Gong, J\. Liu, X\. Zhang, X\. Li, L\. Wei, and H\. HeAdaptive multimodal graph integration network for multimodal sentiment analysis\.IEEE Transactions on Audio, Speech and Language Processing33,pp\. 23–36\.Cited by:[§B\.1](https://arxiv.org/html/2609.30470#A2.SS1.p2.1)\.
- Hanet al\.\(2021\)W\. Han, H\. Chen, A\. Gelbukh, A\. Zadeh, L\. Morency, and S\. PoriaBi\-bimodal modality fusion for correlation\-controlled multimodal sentiment analysis\.InICMI,pp\. 6–15\.Cited by:[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p21.1),[§D\.4](https://arxiv.org/html/2609.30470#A4.SS4.p6.1)\.
- Hasanet al\.\(2021\)M\. K\. Hasan, S\. Lee, W\. Rahman, A\. Zadeh, R\. Mihalcea, L\. Morency, and E\. HoqueHumor knowledge enriched transformer for understanding multimodal humor\.InAAAI,pp\. 12972–12980\.Cited by:[§D\.1](https://arxiv.org/html/2609.30470#A4.SS1.p4.1),[§D\.3](https://arxiv.org/html/2609.30470#A4.SS3.p5.1)\.
- Hasanet al\.\(2019\)M\. K\. Hasan, W\. Rahman, A\. B\. Zadeh, J\. Zhong, M\. I\. Tanveer, L\. Morency, and M\. E\. HoqueUR\-funny: a multimodal language dataset for understanding humor\.InEMNLP\-IJCNLP,pp\. 2046–2056\.Cited by:[§D\.1](https://arxiv.org/html/2609.30470#A4.SS1.p4.1),[§4\.1](https://arxiv.org/html/2609.30470#S4.SS1.p2.1),[§4\.2](https://arxiv.org/html/2609.30470#S4.SS2.p3.1)\.
- Hazarikaet al\.\(2022\)D\. Hazarika, Y\. Li, B\. Cheng, S\. Zhao, R\. Zimmermann, and S\. PoriaAnalyzing modality robustness in multimodal sentiment analysis\.InACL,pp\. 685–696\.Cited by:[§B\.1](https://arxiv.org/html/2609.30470#A2.SS1.p1.1)\.
- Heet al\.\(2026\)K\. He, B\. Chen, Y\. Ding, F\. Li, C\. Teng, and D\. JiPase: prototype\-aligned calibration and shapley\-based equilibrium for multimodal sentiment analysis\.InAAAI,Vol\.40,pp\. 30960–30968\.Cited by:[§1](https://arxiv.org/html/2609.30470#S1.p1.1)\.
- Heet al\.\(2020\)P\. He, X\. Liu, J\. Gao, and W\. ChenDeberta: decoding\-enhanced bert with disentangled attention\.arXiv preprint arXiv:2006\.03654\.Cited by:[§D\.3](https://arxiv.org/html/2609.30470#A4.SS3.p1.1)\.
- Huet al\.\(2025\)P\. Hu, Y\. Qin, Y\. Gou, Y\. Li, M\. Yang, and X\. PengProbabilistic multimodal learning with von mises\-fisher distributions\.InIJCAI,pp\. 5390–5398\.Cited by:[§C\.1](https://arxiv.org/html/2609.30470#A3.SS1.SSS0.Px1.p3.1),[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p10.1),[§D\.4](https://arxiv.org/html/2609.30470#A4.SS4.p5.1),[§E\.1](https://arxiv.org/html/2609.30470#A5.SS1.p1.1),[Table 12](https://arxiv.org/html/2609.30470#A5.T12.5.1.7.1),[Table 6](https://arxiv.org/html/2609.30470#A5.T6.5.1.18.1),[Table 6](https://arxiv.org/html/2609.30470#A5.T6.5.1.27.1),[Table 6](https://arxiv.org/html/2609.30470#A5.T6.5.1.9.1),[Table 7](https://arxiv.org/html/2609.30470#A5.T7.5.1.18.1),[Table 7](https://arxiv.org/html/2609.30470#A5.T7.5.1.27.1),[Table 7](https://arxiv.org/html/2609.30470#A5.T7.5.1.9.1),[§1](https://arxiv.org/html/2609.30470#S1.p3.1),[§3\.3](https://arxiv.org/html/2609.30470#S3.SS3.SSS0.Px1.p1.2),[§4\.1](https://arxiv.org/html/2609.30470#S4.SS1.p3.1),[Table 1](https://arxiv.org/html/2609.30470#S4.T1.7.1.11.1)\.
- Huanget al\.\(2025\)C\. Huang, J\. Chen, Q\. Huang, S\. Wang, Y\. Tu, and X\. HuangAtCAF: attention\-based causality\-aware fusion network for multimodal sentiment analysis\.Information Fusion114,pp\. 102725\.Cited by:[§D\.1](https://arxiv.org/html/2609.30470#A4.SS1.p4.1),[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p17.1),[§D\.5](https://arxiv.org/html/2609.30470#A4.SS5.p7.1),[§4\.2](https://arxiv.org/html/2609.30470#S4.SS2.p3.1)\.
- Jaiswal and Provost \(2020\)M\. Jaiswal and E\. M\. ProvostPrivacy enhanced multimodal neural representations for emotion recognition\.InAAAI,Vol\.34,pp\. 7985–7993\.Cited by:[§1](https://arxiv.org/html/2609.30470#S1.p1.1)\.
- Jianget al\.\(2025a\)M\. Jiang, Y\. Jiang, H\. Hu, and S\. MaiTowards minimal causal representations for human multimodal language understanding\.arXiv preprint arXiv:2509\.21805\.Cited by:[§2\.1](https://arxiv.org/html/2609.30470#S2.SS1.p1.1)\.
- Jianget al\.\(2025b\)M\. Jiang, Y\. Lin, B\. Chen, H\. Hu, Y\. Jiang, and S\. MaiDisentangling bias by modeling intra\-and inter\-modal causal attention for multimodal sentiment analysis\.arXiv preprint arXiv:2508\.04999\.Cited by:[§2\.1](https://arxiv.org/html/2609.30470#S2.SS1.p1.1)\.
- Lanet al\.\(2020\)Z\. Lan, M\. Chen, S\. Goodman, K\. Gimpel, P\. Sharma, and R\. SoricutAlbert: a lite bert for self\-supervised learning of language representations\.InICLR,Cited by:[§D\.3](https://arxiv.org/html/2609.30470#A4.SS3.p1.1)\.
- Liet al\.\(2023a\)H\. Li, X\. Li, P\. Hu, Y\. Lei, C\. Li, and Y\. ZhouBoosting multi\-modal model performance with adaptive gradient modulation\.InICCV,pp\. 22214–22224\.Cited by:[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p16.1)\.
- Liet al\.\(2024a\)M\. Li, D\. Yang, Y\. Liu, S\. Wang, J\. Chen, S\. Wang, J\. Wei, Y\. Jiang, Q\. Xu, X\. Hou,et al\.Toward robust incomplete multimodal sentiment analysis via hierarchical representation learning\.InNeurIPS,Vol\.37,pp\. 28515–28536\.Cited by:[§2\.2](https://arxiv.org/html/2609.30470#S2.SS2.p1.1)\.
- Liet al\.\(2024b\)M\. Li, D\. Yang, X\. Zhao, S\. Wang, Y\. Wang, K\. Yang, M\. Sun, D\. Kou, Z\. Qian, and L\. ZhangCorrelation\-decoupled knowledge distillation for multimodal sentiment analysis with incomplete modalities\.InCVPR,pp\. 12458–12468\.Cited by:[§1](https://arxiv.org/html/2609.30470#S1.p1.1),[§2\.2](https://arxiv.org/html/2609.30470#S2.SS2.p1.1),[§2\.2](https://arxiv.org/html/2609.30470#S2.SS2.p2.1),[§3\.2](https://arxiv.org/html/2609.30470#S3.SS2.p1.1),[§3\.4](https://arxiv.org/html/2609.30470#S3.SS4.p4.1)\.
- Liet al\.\(2023b\)Y\. Li, Y\. Wang, and Z\. CuiDecoupled multimodal distilling for emotion recognition\.InCVPR,pp\. 6631–6640\.Cited by:[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p15.1),[§D\.4](https://arxiv.org/html/2609.30470#A4.SS4.p6.1)\.
- Liet al\.\(2026\)Y\. Li, Y\. Wang, Y\. Ding, S\. Zhang, K\. Lu, and C\. GuanDecoupled hierarchical distillation for multimodal emotion recognition\.IEEE Transactions on Pattern Analysis and Machine Intelligence\.Cited by:[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p19.1),[§D\.5](https://arxiv.org/html/2609.30470#A4.SS5.p7.1)\.
- Lianet al\.\(2023\)Z\. Lian, L\. Chen, L\. Sun, B\. Liu, and J\. TaoGcnet: graph completion network for incomplete multimodal learning in conversation\.IEEE Transactions on pattern analysis and machine intelligence45\(7\),pp\. 8419–8432\.Cited by:[§2\.2](https://arxiv.org/html/2609.30470#S2.SS2.p1.1)\.
- Lianget al\.\(2019\)P\. P\. Liang, Z\. Liu, Y\. H\. Tsai, Q\. Zhao, R\. Salakhutdinov, and L\. MorencyLearning representations from imperfect time series data via tensor rank regularization\.InACL,pp\. 1569–1576\.Cited by:[§B\.1](https://arxiv.org/html/2609.30470#A2.SS1.p3.1)\.
- Linet al\.\(2025\)R\. Lin, Q\. He, S\. Mai, Y\. Zeng, A\. Xiong, L\. Huang, Y\. Tan, and H\. HuCyIN: cyclic informative latent space for bridging complete and incomplete multimodal learning\.InNeurIPS,Cited by:[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p12.1),[§D\.3](https://arxiv.org/html/2609.30470#A4.SS3.p2.1),[§1](https://arxiv.org/html/2609.30470#S1.p2.1),[§1](https://arxiv.org/html/2609.30470#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.30470#S2.SS2.p1.1),[§4\.1](https://arxiv.org/html/2609.30470#S4.SS1.p3.1),[§4\.2](https://arxiv.org/html/2609.30470#S4.SS2.p2.1)\.
- Liuet al\.\(2017\)Z\. Liu, M\. Wu, W\. Cao, L\. Chen, J\. Xu, R\. Zhang, M\. Zhou, and J\. MaoA facial expression emotion recognition based human\-robot interaction system\.\.IEEE CAA J\. Autom\. Sinica4\(4\),pp\. 668–676\.Cited by:[§1](https://arxiv.org/html/2609.30470#S1.p1.1)\.
- Loshchilov and Hutter \(2017\)I\. Loshchilov and F\. HutterDecoupled weight decay regularization\.arXiv preprint arXiv:1711\.05101\.Cited by:[§D\.4](https://arxiv.org/html/2609.30470#A4.SS4.p1.1)\.
- Luoet al\.\(2025\)M\. Luo, Y\. Jiang, and S\. MaiTowards explainable fusion and balanced learning in multimodal sentiment analysis\.InACM MM,pp\. 1997–2006\.Cited by:[§B\.1](https://arxiv.org/html/2609.30470#A2.SS1.p3.1),[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p7.1),[§D\.4](https://arxiv.org/html/2609.30470#A4.SS4.p5.1),[§D\.5](https://arxiv.org/html/2609.30470#A4.SS5.p3.1),[§E\.1](https://arxiv.org/html/2609.30470#A5.SS1.p1.1),[§E\.2](https://arxiv.org/html/2609.30470#A5.SS2.SSS0.Px3.p2.1),[§E\.7](https://arxiv.org/html/2609.30470#A5.SS7.p1.1),[Table 12](https://arxiv.org/html/2609.30470#A5.T12.5.1.4.1),[Table 6](https://arxiv.org/html/2609.30470#A5.T6.5.1.15.1),[Table 6](https://arxiv.org/html/2609.30470#A5.T6.5.1.24.1),[Table 6](https://arxiv.org/html/2609.30470#A5.T6.5.1.6.1),[Table 7](https://arxiv.org/html/2609.30470#A5.T7.5.1.15.1),[Table 7](https://arxiv.org/html/2609.30470#A5.T7.5.1.24.1),[Table 7](https://arxiv.org/html/2609.30470#A5.T7.5.1.6.1),[§4\.1](https://arxiv.org/html/2609.30470#S4.SS1.p3.1),[Table 1](https://arxiv.org/html/2609.30470#S4.T1.7.1.8.1)\.
- Mai and Han \(2026\)S\. Mai and S\. HanCaReFlow: cyclic adaptive rectified flow for multimodal fusion\.arXiv preprint arXiv:2602\.19140\.Cited by:[§2\.1](https://arxiv.org/html/2609.30470#S2.SS1.p1.1)\.
- Maiet al\.\(2023a\)S\. Mai, Y\. Sun, A\. Xiong, Y\. Zeng, and H\. HuMultimodal boosting: addressing noisy modalities and identifying modality contribution\.IEEE Transactions on Multimedia26,pp\. 3018–3033\.Cited by:[§B\.1](https://arxiv.org/html/2609.30470#A2.SS1.p2.1)\.
- Maiet al\.\(2023b\)S\. Mai, Y\. Zeng, and H\. HuMultimodal information bottleneck: learning minimal sufficient unimodal and multimodal representations\.IEEE Transactions on Multimedia25,pp\. 4121–4134\.Cited by:[§B\.1](https://arxiv.org/html/2609.30470#A2.SS1.p3.1),[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p5.1),[§D\.4](https://arxiv.org/html/2609.30470#A4.SS4.p5.1),[§E\.1](https://arxiv.org/html/2609.30470#A5.SS1.p1.1),[Table 12](https://arxiv.org/html/2609.30470#A5.T12.5.1.2.1),[Table 6](https://arxiv.org/html/2609.30470#A5.T6.5.1.13.1),[Table 6](https://arxiv.org/html/2609.30470#A5.T6.5.1.22.1),[Table 6](https://arxiv.org/html/2609.30470#A5.T6.5.1.4.1),[Table 7](https://arxiv.org/html/2609.30470#A5.T7.5.1.13.1),[Table 7](https://arxiv.org/html/2609.30470#A5.T7.5.1.22.1),[Table 7](https://arxiv.org/html/2609.30470#A5.T7.5.1.4.1),[§1](https://arxiv.org/html/2609.30470#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.30470#S4.SS1.p3.1),[Table 1](https://arxiv.org/html/2609.30470#S4.T1.7.1.6.1)\.
- Maiet al\.\(2025\)S\. Mai, Y\. Zeng, and H\. HuLearning by comparing: boosting multimodal affective computing through ordinal learning\.InWWW,pp\. 2120–2134\.Cited by:[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p24.1),[§D\.5](https://arxiv.org/html/2609.30470#A4.SS5.p7.1),[§2\.1](https://arxiv.org/html/2609.30470#S2.SS1.p1.1),[§4\.2](https://arxiv.org/html/2609.30470#S4.SS2.p3.1)\.
- Maoet al\.\(2023\)H\. Mao, B\. Zhang, H\. Xu, Z\. Yuan, and Y\. LiuRobust\-msa: understanding the impact of modality noise on multimodal sentiment analysis\.InAAAI,Vol\.37,pp\. 16458–16460\.Cited by:[§B\.1](https://arxiv.org/html/2609.30470#A2.SS1.p1.1)\.
- Minget al\.\(2023\)Y\. Ming, Y\. Sun, O\. Dia, and Y\. LiHow to exploit hyperspherical embeddings for out\-of\-distribution detection?\.InICLR,Cited by:[§C\.1](https://arxiv.org/html/2609.30470#A3.SS1.SSS0.Px1.p3.1),[§3\.3](https://arxiv.org/html/2609.30470#S3.SS3.SSS0.Px1.p1.2)\.
- Phamet al\.\(2019\)H\. Pham, P\. P\. Liang, T\. Manzini, L\. Morency, and B\. PóczosFound in translation: learning robust joint representations by cyclic translations between modalities\.InAAAI,Vol\.33,pp\. 6892–6899\.Cited by:[§2\.2](https://arxiv.org/html/2609.30470#S2.SS2.p1.1)\.
- Pramanicket al\.\(2022\)S\. Pramanick, A\. Roy, and V\. M\. PatelMultimodal learning using optimal transport for sarcasm and humor detection\.InWACV,pp\. 3930–3940\.Cited by:[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p22.1),[§D\.4](https://arxiv.org/html/2609.30470#A4.SS4.p6.1)\.
- Qianet al\.\(2026\)C\. Qian, S\. Xing, S\. Li, Y\. Zhao, and Z\. TuDecalign: hierarchical cross\-modal alignment for decoupled multimodal representation learning\.InICLR,Cited by:[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p4.1),[§1](https://arxiv.org/html/2609.30470#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.30470#S2.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.30470#S4.SS1.p3.1),[§4\.2](https://arxiv.org/html/2609.30470#S4.SS2.p1.1),[Table 1](https://arxiv.org/html/2609.30470#S4.T1.7.1.5.1)\.
- Rahmanet al\.\(2020\)W\. Rahman, M\. K\. Hasan, S\. Lee, A\. Zadeh, C\. Mao, L\. Morency, and E\. HoqueIntegrating multimodal information in large pretrained transformers\.InACL,pp\. 2359\.Cited by:[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p20.1),[§D\.4](https://arxiv.org/html/2609.30470#A4.SS4.p6.1)\.
- Tishbyet al\.\(2000\)N\. Tishby, F\. C\. Pereira, and W\. BialekThe information bottleneck method\.arXiv preprint physics/0004057\.Cited by:[§B\.1](https://arxiv.org/html/2609.30470#A2.SS1.p3.1),[§1](https://arxiv.org/html/2609.30470#S1.p2.1)\.
- Tsaiet al\.\(2019\)Y\. H\. Tsai, S\. Bai, P\. P\. Liang, J\. Z\. Kolter, L\. Morency, and R\. SalakhutdinovMultimodal transformer for unaligned multimodal language sequences\.InACL,pp\. 6558\.Cited by:[§3\.2](https://arxiv.org/html/2609.30470#S3.SS2.p1.1)\.
- Van der Maaten and Hinton \(2008\)L\. Van der Maaten and G\. HintonVisualizing data using t\-sne\.\.Journal of machine learning research9\(11\)\.Cited by:[§4\.4](https://arxiv.org/html/2609.30470#S4.SS4.p2.1)\.
- Vaswaniet al\.\(2017\)A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, and I\. PolosukhinAttention is all you need\.InNeurIPS,Vol\.30\.Cited by:[§3\.2](https://arxiv.org/html/2609.30470#S3.SS2.p1.1)\.
- Vitter \(1985\)J\. S\. VitterRandom sampling with a reservoir\.ACM Transactions on Mathematical Software11\(1\),pp\. 37–57\.Cited by:[§C\.2](https://arxiv.org/html/2609.30470#A3.SS2.p1.1),[§3\.4](https://arxiv.org/html/2609.30470#S3.SS4.p3.1)\.
- Wanget al\.\(2025\)P\. Wang, Q\. Zhou, Y\. Wu, T\. Chen, and J\. HuDLF: disentangled\-language\-focused multimodal sentiment analysis\.InAAAI,pp\. 21180–21188\.Cited by:[§2\.1](https://arxiv.org/html/2609.30470#S2.SS1.p1.1)\.
- Wanget al\.\(2023a\)Y\. Wang, Z\. Cui, and Y\. LiDistribution\-consistent modal recovering for incomplete multimodal learning\.InCVPR,pp\. 22025–22034\.Cited by:[§2\.2](https://arxiv.org/html/2609.30470#S2.SS2.p1.1)\.
- Wanget al\.\(2023b\)Y\. Wang, Y\. Li, and Z\. CuiIncomplete multimodality\-diffused emotion recognition\.InNeurIPS,Vol\.36,pp\. 17117–17128\.Cited by:[§2\.2](https://arxiv.org/html/2609.30470#S2.SS2.p1.1)\.
- Wuet al\.\(2025a\)Q\. Wu, Y\. Shao, J\. Wang, and X\. SunLearning optimal multimodal information bottleneck representations\.InICML,Cited by:[§B\.1](https://arxiv.org/html/2609.30470#A2.SS1.p3.1),[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p8.1),[§D\.4](https://arxiv.org/html/2609.30470#A4.SS4.p5.1),[§E\.1](https://arxiv.org/html/2609.30470#A5.SS1.p1.1),[Table 12](https://arxiv.org/html/2609.30470#A5.T12.5.1.5.1),[Table 6](https://arxiv.org/html/2609.30470#A5.T6.5.1.16.1),[Table 6](https://arxiv.org/html/2609.30470#A5.T6.5.1.25.1),[Table 6](https://arxiv.org/html/2609.30470#A5.T6.5.1.7.1),[Table 7](https://arxiv.org/html/2609.30470#A5.T7.5.1.16.1),[Table 7](https://arxiv.org/html/2609.30470#A5.T7.5.1.25.1),[Table 7](https://arxiv.org/html/2609.30470#A5.T7.5.1.7.1),[§1](https://arxiv.org/html/2609.30470#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.30470#S4.SS1.p3.1),[Table 1](https://arxiv.org/html/2609.30470#S4.T1.7.1.9.1)\.
- Wuet al\.\(2025b\)S\. Wu, D\. He, X\. Wang, L\. Wang, and J\. DangEnriching multimodal sentiment analysis through textual emotional descriptions of visual\-audio content\.InAAAI,pp\. 1601–1609\.Cited by:[§2\.1](https://arxiv.org/html/2609.30470#S2.SS1.p1.1)\.
- Xiaoet al\.\(2024\)X\. Xiao, G\. Liu, G\. Gupta, D\. Cao, S\. Li, Y\. Li, T\. Fang, M\. Cheng, and P\. BogdanNeuro\-inspired information\-theoretic hierarchical perception for multimodal learning\.InICLR,Cited by:[§B\.1](https://arxiv.org/html/2609.30470#A2.SS1.p3.1),[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p6.1),[§D\.4](https://arxiv.org/html/2609.30470#A4.SS4.p5.1),[§D\.5](https://arxiv.org/html/2609.30470#A4.SS5.p3.1),[§E\.1](https://arxiv.org/html/2609.30470#A5.SS1.p1.1),[Table 12](https://arxiv.org/html/2609.30470#A5.T12.5.1.3.1),[Table 6](https://arxiv.org/html/2609.30470#A5.T6.5.1.14.1),[Table 6](https://arxiv.org/html/2609.30470#A5.T6.5.1.23.1),[Table 6](https://arxiv.org/html/2609.30470#A5.T6.5.1.5.1),[Table 7](https://arxiv.org/html/2609.30470#A5.T7.5.1.14.1),[Table 7](https://arxiv.org/html/2609.30470#A5.T7.5.1.23.1),[Table 7](https://arxiv.org/html/2609.30470#A5.T7.5.1.5.1),[Table 1](https://arxiv.org/html/2609.30470#S4.T1.7.1.7.1)\.
- Xieet al\.\(2026\)H\. Xie, K\. Zhu, Z\. Wen, J\. Tao, X\. Liu, R\. Fu, and C\. LiPSA\-mf: personality\-sentiment aligned multi\-level fusion for multimodal sentiment analysis\.InAAAI,Cited by:[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p3.1),[§2\.1](https://arxiv.org/html/2609.30470#S2.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.30470#S4.SS1.p3.1),[Table 1](https://arxiv.org/html/2609.30470#S4.T1.7.1.4.1)\.
- Xuet al\.\(2025\)Z\. Xu, D\. Yang, M\. Li, Y\. Wang, Z\. Chen, J\. Chen, J\. Wei, and L\. ZhangDebiased multimodal understanding for human language sequences\.InAAAI,pp\. 14450–14458\.Cited by:[§D\.1](https://arxiv.org/html/2609.30470#A4.SS1.p4.1),[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p18.1),[§D\.4](https://arxiv.org/html/2609.30470#A4.SS4.p6.1),[§D\.5](https://arxiv.org/html/2609.30470#A4.SS5.p7.1)\.
- Xue and Marculescu \(2023\)Z\. Xue and R\. MarculescuDynamic multimodal fusion\.InCVPR,pp\. 2575–2584\.Cited by:[§B\.1](https://arxiv.org/html/2609.30470#A2.SS1.p2.1)\.
- Yanget al\.\(2022\)D\. Yang, S\. Huang, H\. Kuang, Y\. Du, and L\. ZhangDisentangled representation learning for multimodal emotion recognition\.InACM MM,pp\. 1642–1651\.Cited by:[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p14.1),[§D\.4](https://arxiv.org/html/2609.30470#A4.SS4.p6.1)\.
- Yanget al\.\(2026\)D\. Yang, M\. Li, X\. Wu, Z\. Chen, K\. Jiang, K\. Liu, P\. Zhai, and L\. ZhangImproving multimodal sentiment analysis via modality optimization and dynamic primary modality selection\.InAAAI,Cited by:[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p2.1),[§1](https://arxiv.org/html/2609.30470#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.30470#S2.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.30470#S4.SS1.p3.1),[§4\.2](https://arxiv.org/html/2609.30470#S4.SS2.p1.1),[Table 1](https://arxiv.org/html/2609.30470#S4.T1.7.1.3.1)\.
- Yeet al\.\(2025\)J\. Ye, J\. Zhang, and H\. ShanDepmamba: progressive fusion mamba for multimodal depression detection\.InICASSP,pp\. 1–5\.Cited by:[§1](https://arxiv.org/html/2609.30470#S1.p1.1)\.
- Yuet al\.\(2021\)W\. Yu, H\. Xu, Z\. Yuan, and J\. WuLearning modality\-specific representations with self\-supervised multi\-task learning for multimodal sentiment analysis\.InAAAI,pp\. 10790–10797\.Cited by:[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p13.1),[§D\.4](https://arxiv.org/html/2609.30470#A4.SS4.p6.1)\.
- Zadehet al\.\(2017\)A\. Zadeh, M\. Chen, S\. Poria, E\. Cambria, and L\. MorencyTensor fusion network for multimodal sentiment analysis\.InEMNLP,pp\. 1103–1114\.Cited by:[§2\.1](https://arxiv.org/html/2609.30470#S2.SS1.p1.1)\.
- Zadehet al\.\(2018a\)A\. Zadeh, Y\. C\. Lim, and L\. MorencyOpenface 2\.0: facial behavior analysis toolkit tadas baltrušaitis\.InFG,pp\. 59–66\.Cited by:[§D\.3](https://arxiv.org/html/2609.30470#A4.SS3.p4.1)\.
- Zadehet al\.\(2016\)A\. Zadeh, R\. Zellers, E\. Pincus, and L\. MorencyMultimodal sentiment intensity analysis in videos: facial gestures and verbal messages\.IEEE Intelligent Systems31,pp\. 82–88\.Cited by:[§D\.1](https://arxiv.org/html/2609.30470#A4.SS1.p2.1),[§1](https://arxiv.org/html/2609.30470#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.30470#S2.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.30470#S4.SS1.p2.1)\.
- Zadehet al\.\(2018b\)A\. B\. Zadeh, P\. P\. Liang, S\. Poria, E\. Cambria, and L\. MorencyMultimodal language analysis in the wild: cmu\-mosei dataset and interpretable dynamic fusion graph\.InACL,pp\. 2236–2246\.Cited by:[§D\.1](https://arxiv.org/html/2609.30470#A4.SS1.p3.1),[§4\.1](https://arxiv.org/html/2609.30470#S4.SS1.p2.1)\.
- Zhanget al\.\(2024\)H\. Zhang, W\. Wang, and T\. YuTowards robust multimodal sentiment analysis with incomplete data\.InNeurIPS,Vol\.37,pp\. 55943–55974\.Cited by:[§1](https://arxiv.org/html/2609.30470#S1.p1.1),[§2\.2](https://arxiv.org/html/2609.30470#S2.SS2.p1.1),[§3\.5](https://arxiv.org/html/2609.30470#S3.SS5.SSS0.Px1.p1.1)\.
- Zhanget al\.\(2023a\)H\. Zhang, Y\. Wang, G\. Yin, K\. Liu, Y\. Liu, and T\. YuLearning language\-guided adaptive hyper\-modality representation for multimodal sentiment analysis\.InEMNLP,pp\. 756–767\.Cited by:[§2\.1](https://arxiv.org/html/2609.30470#S2.SS1.p1.1),[§3\.5](https://arxiv.org/html/2609.30470#S3.SS5.SSS0.Px1.p1.1)\.
- Zhanget al\.\(2023b\)Q\. Zhang, H\. Wu, C\. Zhang, Q\. Hu, H\. Fu, J\. T\. Zhou, and X\. PengProvable dynamic fusion for low\-quality multimodal data\.InICML,pp\. 41753–41769\.Cited by:[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p9.1),[§D\.4](https://arxiv.org/html/2609.30470#A4.SS4.p5.1),[§E\.1](https://arxiv.org/html/2609.30470#A5.SS1.p1.1),[§E\.2](https://arxiv.org/html/2609.30470#A5.SS2.SSS0.Px1.p1.1),[§E\.2](https://arxiv.org/html/2609.30470#A5.SS2.SSS0.Px3.p1.1),[Table 12](https://arxiv.org/html/2609.30470#A5.T12.5.1.6.1),[Table 6](https://arxiv.org/html/2609.30470#A5.T6.5.1.17.1),[Table 6](https://arxiv.org/html/2609.30470#A5.T6.5.1.26.1),[Table 6](https://arxiv.org/html/2609.30470#A5.T6.5.1.8.1),[Table 7](https://arxiv.org/html/2609.30470#A5.T7.5.1.17.1),[Table 7](https://arxiv.org/html/2609.30470#A5.T7.5.1.26.1),[Table 7](https://arxiv.org/html/2609.30470#A5.T7.5.1.8.1),[§4\.1](https://arxiv.org/html/2609.30470#S4.SS1.p3.1),[Table 1](https://arxiv.org/html/2609.30470#S4.T1.7.1.10.1)\.
- Zhanget al\.\(2023c\)Y\. Zhang, Y\. Yu, D\. Zhao, Z\. Li, B\. Wang, Y\. Hou, P\. Tiwari, and J\. QinLearning multitask commonness and uniqueness for multimodal sarcasm detection and sentiment analysis in conversation\.IEEE Transactions on Artificial Intelligence5,pp\. 1349–1361\.Cited by:[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p23.1)\.
- Zhaoet al\.\(2021\)J\. Zhao, R\. Li, and Q\. JinMissing modality imagination network for emotion recognition with uncertain missing modalities\.InACL,pp\. 2608–2618\.Cited by:[§2\.2](https://arxiv.org/html/2609.30470#S2.SS2.p1.1)\.
- Zhuanget al\.\(2025a\)Y\. Zhuang, M\. Liu, W\. Bai, Y\. Zhang, X\. Zhang, J\. Deng, and F\. RenCMAD: correlation\-aware and modalities\-aware distillation for multimodal sentiment analysis with missing modalities\.InICCV,pp\. 4626–4636\.Cited by:[§2\.2](https://arxiv.org/html/2609.30470#S2.SS2.p1.1)\.
- Zhuanget al\.\(2026\)Y\. Zhuang, M\. Liu, Y\. Zhang, J\. Deng, and F\. RenTMDC: a two\-stage modality denoising and complementation framework for multimodal sentiment analysis with missing and noisy modalities\.InAAAI,Vol\.40,pp\. 2281–2289\.Cited by:[§1](https://arxiv.org/html/2609.30470#S1.p1.1),[§1](https://arxiv.org/html/2609.30470#S1.p3.1)\.
- Zhuanget al\.\(2025b\)Y\. Zhuang, L\. Minhao, W\. Bai, Y\. Zhang, W\. Li, J\. Deng, and F\. RenHyper\-modality enhancement for multimodal sentiment analysis with missing modalities\.InNeurIPS,Cited by:[§D\.2](https://arxiv.org/html/2609.30470#A4.SS2.p11.1),[§D\.3](https://arxiv.org/html/2609.30470#A4.SS3.p2.1),[§D\.4](https://arxiv.org/html/2609.30470#A4.SS4.p5.1),[§E\.1](https://arxiv.org/html/2609.30470#A5.SS1.p1.1),[§E\.2](https://arxiv.org/html/2609.30470#A5.SS2.SSS0.Px2.p2.1),[§E\.2](https://arxiv.org/html/2609.30470#A5.SS2.SSS0.Px3.p2.1),[§E\.7](https://arxiv.org/html/2609.30470#A5.SS7.p1.1),[Table 12](https://arxiv.org/html/2609.30470#A5.T12.5.1.8.1),[Table 6](https://arxiv.org/html/2609.30470#A5.T6.5.1.10.1),[Table 6](https://arxiv.org/html/2609.30470#A5.T6.5.1.19.1),[Table 6](https://arxiv.org/html/2609.30470#A5.T6.5.1.28.1),[Table 7](https://arxiv.org/html/2609.30470#A5.T7.5.1.10.1),[Table 7](https://arxiv.org/html/2609.30470#A5.T7.5.1.19.1),[Table 7](https://arxiv.org/html/2609.30470#A5.T7.5.1.28.1),[§1](https://arxiv.org/html/2609.30470#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.30470#S2.SS2.p1.1),[§2\.2](https://arxiv.org/html/2609.30470#S2.SS2.p2.1),[§3\.2](https://arxiv.org/html/2609.30470#S3.SS2.p1.1),[§3\.4](https://arxiv.org/html/2609.30470#S3.SS4.p1.1),[§3\.5](https://arxiv.org/html/2609.30470#S3.SS5.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2609.30470#S4.SS1.p3.1),[§4\.2](https://arxiv.org/html/2609.30470#S4.SS2.p2.1),[§4\.4](https://arxiv.org/html/2609.30470#S4.SS4.p3.1),[Table 1](https://arxiv.org/html/2609.30470#S4.T1.7.1.12.1)\.

Reliability\-aware Cross\-sample Enhancement for Robust Multimodal Sentiment Analysis \(Appendix\)

## The Use of Large Language Models \(LLMs\)

In this work, large language models \(LLMs\) were used solely for language polishing to improve clarity and fluency\. All scientific content, including the experimental design, results, and conclusions, is the responsibility of the authors\. No LLM was involved in the ideation, analysis, or interpretation of the research, and all LLM\-assisted content was carefully reviewed and verified by the authors\.

## Appendix ANotations

We provide a comprehensive overview of the commonly used notations and their definitions in Table[4](https://arxiv.org/html/2609.30470#A1.T4)\.

Table 4:Notation and Definitions
## Appendix BMore Related Work

### B\.1Multimodal Sentiment Analysis with Noisy Modalities

Existing approaches for handling modality noise can be broadly categorized into three groups\. The first group, noise augmentation methods, simulates real\-world noise by injecting perturbations into features or raw inputs and developing corresponding defenses\. For example, Robust\-MSA\[[35](https://arxiv.org/html/2609.30470#bib.bib26)\]constructs an interactive analysis platform to visualize the impact of modality\-specific noise on model performance, and incorporates simple defense strategies to improve interpretability\. Similarly, MSA\-Robustness\[[11](https://arxiv.org/html/2609.30470#bib.bib22)\]introduces a suite of diagnostic tools to assess the sensitivity of multimodal models to noise in individual modalities, and systematically explores robust training strategies to enhance stability under noisy perturbations\. However, designing dedicated defenses for different noise types is complex and often lacks generalization\.

The second group focuses on noise identification and filtering mechanisms to suppress noisy signals and improve the robustness of multimodal representations\[[7](https://arxiv.org/html/2609.30470#bib.bib66)\]\. For instance, DynMM\[[54](https://arxiv.org/html/2609.30470#bib.bib25)\]introduces a dynamic gating mechanism that adaptively adjusts modality contributions or fusion pathways based on input features, thereby mitigating the influence of noisy modalities\. However, such implicit weighting strategies lack explicit reliability modeling and thus remain susceptible to noise\. To further improve modality contribution estimation, Multimodal Boosting\[[32](https://arxiv.org/html/2609.30470#bib.bib38)\]leverages unimodal prediction losses as supervisory signals to identify noisy modalities\. Nevertheless, this approach relies on the construction of absolute labels, which are difficult to define precisely and may introduce additional noise during training\.

The third group adopts representation regularization strategies, which have emerged as a promising direction\. These methods typically leverage techniques such as tensor rank minimization\[[26](https://arxiv.org/html/2609.30470#bib.bib67)\]or information\-theoretic principles\[[41](https://arxiv.org/html/2609.30470#bib.bib1)\]to extract discriminative representations from noisy features\. For example, MIB\[[33](https://arxiv.org/html/2609.30470#bib.bib30)\], based on the information bottleneck principle\[[41](https://arxiv.org/html/2609.30470#bib.bib1)\], suppresses redundancy and noise in both unimodal and multimodal representations by learning minimal sufficient task\-relevant features\. ITHP\[[51](https://arxiv.org/html/2609.30470#bib.bib39)\]designates a primary modality while treating the remaining modalities as auxiliary detectors to extract salient information\. OMIB\[[49](https://arxiv.org/html/2609.30470#bib.bib45)\]further extends the multimodal information bottleneck framework by constraining regularization weights to ensure an optimal bottleneck representation\. Similarly, KAN\-MCP\[[30](https://arxiv.org/html/2609.30470#bib.bib51)\]integrates cross\-modal modeling with Dimensionality Reduction and Denoising Modal Information Bottleneck \(DRD\-MIB\)\-based feature compression to learn more robust representations\.

Despite the success of representation regularization methods, particularly those based on the information bottleneck principle, such approaches typically compress all modalities equally and lack the ability to distinguish modality reliability, which makes them less effective in the presence of noise or heterogeneous modality quality\. In contrast, RCE introduces an Adaptive Variational Information Bottleneck \(AVIB\) in the von Mises–Fisher \(vMF\) representation space\. Specifically, AVIB jointly models directional information and sample\-wise posterior uncertainty within a unified hyperspherical space, enabling reliability\-aware adaptive compression that preserves informative signals from reliable modalities while suppressing noise from less reliable ones\.

## Appendix CMethodological Details

### C\.1von Mises–Fisher Distribution

In this section, we provide a detailed description of the von Mises–Fisher \(vMF\) distribution and its role in modeling sample\-level uncertainty in our RCE framework\.

#### vMF\-based Uncertainty Estimation\.

To explicitly model sample\-level reliability differences across modalities, we map the feature representation of each modality onto a probability distribution defined on the unit hypersphere\. Specifically, for theii\-th sample and modalitymm, the representationHmiH\_\{m\}^\{i\}is used to parameterize a von Mises–Fisher \(vMF\) posterior distribution:

zmi∼𝒱d​\(μmi,κmi\),z\_\{m\}^\{i\}\\sim\\mathcal\{V\}\_\{d\}\(\\mu\_\{m\}^\{i\},\\kappa\_\{m\}^\{i\}\),\(17\)whereμmi∈𝕊d−1\\mu\_\{m\}^\{i\}\\in\\mathbb\{S\}^\{d\-1\}denotes the mean direction andκmi∈ℝ\+\\kappa\_\{m\}^\{i\}\\in\\mathbb\{R\}^\{\+\}is the concentration parameter\.

The parametersμmi\\mu\_\{m\}^\{i\}andκmi\\kappa\_\{m\}^\{i\}are computed as:

μmi=fm​\(Hmi\)‖fm​\(Hmi\)‖2∈𝕊d−1,κmi=softplus⁡\(gm​\(Hmi\)\)\+ϵ∈ℝ\+,\\mu\_\{m\}^\{i\}=\\frac\{f\_\{m\}\(H\_\{m\}^\{i\}\)\}\{\\\|f\_\{m\}\(H\_\{m\}^\{i\}\)\\\|\_\{2\}\}\\in\\mathbb\{S\}^\{d\-1\},\\qquad\\kappa\_\{m\}^\{i\}=\\mathrm\{softplus\}\(g\_\{m\}\(H\_\{m\}^\{i\}\)\)\+\\epsilon\\in\\mathbb\{R\}^\{\+\},\(18\)wherefm​\(⋅\)f\_\{m\}\(\\cdot\)andgm​\(⋅\)g\_\{m\}\(\\cdot\)are two separate subnetworks used to estimate the directional and concentration parameters, respectively\. The normalization ensures thatμmi\\mu\_\{m\}^\{i\}lies on the unit hypersphere, while the softplus function guarantees thatκmi\\kappa\_\{m\}^\{i\}remains positive\. The constantϵ\\epsilonis a small value introduced to ensure numerical stability\.

Under this formulation,μmi\\mu\_\{m\}^\{i\}represents the principal direction of the latent representation on the hypersphere, whileκmi\\kappa\_\{m\}^\{i\}controls the dispersion of the distribution around this direction\. Intuitively, a largerκmi\\kappa\_\{m\}^\{i\}results in a more concentrated distribution, indicating higher certainty, whereas a smallerκmi\\kappa\_\{m\}^\{i\}corresponds to a more dispersed distribution, reflecting greater uncertainty\. Following prior work\[[36](https://arxiv.org/html/2609.30470#bib.bib63),[4](https://arxiv.org/html/2609.30470#bib.bib62),[14](https://arxiv.org/html/2609.30470#bib.bib61)\], we interpret the concentration parameterκmi\\kappa\_\{m\}^\{i\}as a proxy for sample\-level reliability of the corresponding modality\. This interpretation allows us to quantify the confidence of each modality and later leverage it for reliability\-aware fusion\.

#### Probability Density Function\.

The vMF distribution is defined on the unit hypersphere𝕊d−1⊂ℝd\\mathbb\{S\}^\{d\-1\}\\subset\\mathbb\{R\}^\{d\}\. The probability density function of a random variablez∈𝕊d−1z\\in\\mathbb\{S\}^\{d\-1\}is given by:

p⁡\(z∣μmi,κmi\)=Cd​\(κmi\)​exp⁡\(κmi​μmi⊤​z\),p\(z\\mid\\mu\_\{m\}^\{i\},\\kappa\_\{m\}^\{i\}\)=C\_\{d\}\(\\kappa\_\{m\}^\{i\}\)\\exp\\left\(\\kappa\_\{m\}^\{i\}\{\\mu\_\{m\}^\{i\}\}^\{\\top\}z\\right\),\(19\)whereCd​\(κmi\)C\_\{d\}\(\\kappa\_\{m\}^\{i\}\)is the normalization constant defined as:

Cd​\(κmi\)=\(κmi\)d2−1\(2​π\)d2​Id2−1​\(κmi\)\.C\_\{d\}\(\\kappa\_\{m\}^\{i\}\)=\\frac\{\(\\kappa\_\{m\}^\{i\}\)^\{\\frac\{d\}\{2\}\-1\}\}\{\(2\\pi\)^\{\\frac\{d\}\{2\}\}I\_\{\\frac\{d\}\{2\}\-1\}\(\\kappa\_\{m\}^\{i\}\)\}\.\(20\)
Here,Iν​\(⋅\)I\_\{\\nu\}\(\\cdot\)denotes the modified Bessel function of the first kind of orderν\\nu, which is defined as:

Iν​\(t\)=∑j=0∞\(t2\)2​j\+νj\!​Γ​\(j\+ν\+1\),I\_\{\\nu\}\(t\)=\\sum\_\{j=0\}^\{\\infty\}\\frac\{\\left\(\\frac\{t\}\{2\}\\right\)^\{2j\+\\nu\}\}\{j\!\\,\\Gamma\(j\+\\nu\+1\)\},\(21\)whereΓ⁡\(⋅\)\\Gamma\(\\cdot\)is the Gamma function\.

#### Discussion\.

The vMF distribution provides a natural way to model directional data on the hypersphere\. In our setting, it enables a unified representation that jointly captures both the semantic direction \(viaμmi\\mu\_\{m\}^\{i\}\) and the uncertainty \(viaκmi\\kappa\_\{m\}^\{i\}\)\. This property is important for multimodal learning, where different modalities may exhibit varying levels of reliability across samples\. By leveraging the concentration parameter as an uncertainty measure, our framework achieves adaptive, reliability\-aware representation learning\.

### C\.2Reservoir Sampling Strategy

We use reservoir sampling\[[45](https://arxiv.org/html/2609.30470#bib.bib68)\]to maintain a fixed\-size modality\-specific memory bank while avoiding bias toward recently observed samples\. For each modalitymm, the memory bank is denoted as

ℬm=\{\(zmr,μmr,κmr,yr\)\}r=1Nm,\\mathcal\{B\}\_\{m\}=\\\{\(z\_\{m\}^\{r\},\\mu\_\{m\}^\{r\},\\kappa\_\{m\}^\{r\},y^\{r\}\)\\\}\_\{r=1\}^\{N\_\{m\}\},\(22\)whereNm≤NbN\_\{m\}\\leq N\_\{b\}andNbN\_\{b\}is the memory bank capacity\. During memory construction, samples are sequentially processed by a forward pass without gradient updates\. Letttdenote the number of samples that have been observed so far for updating the memory bank of modalitymm\. For thett\-th observed sample, we obtain its latent representationzmtz\_\{m\}^\{t\}, vMF mean directionμmt\\mu\_\{m\}^\{t\}, concentration parameterκmt\\kappa\_\{m\}^\{t\}, and labelyty^\{t\}\.

If the memory bank is not full, i\.e\.,t≤Nbt\\leq N\_\{b\}, the sample is directly inserted intoℬm\\mathcal\{B\}\_\{m\}\. Once the memory bank reaches its capacity, i\.e\.,t\>Nbt\>N\_\{b\}, the new sample is retained with probabilityNb/tN\_\{b\}/t\. Specifically, we draw a random integer

q∼Uniform​\{1,2,…,t\}\.q\\sim\\mathrm\{Uniform\}\\\{1,2,\\dots,t\\\}\.\(23\)Ifq≤Nbq\\leq N\_\{b\}, we replace theqq\-th element inℬm\\mathcal\{B\}\_\{m\}with the current sample; otherwise, the current sample is discarded\. The update rule can be written as:

ℬm​\[q\]←\(zmt,μmt,κmt,yt\),if​q≤Nb\.\\mathcal\{B\}\_\{m\}\[q\]\\leftarrow\(z\_\{m\}^\{t\},\\mu\_\{m\}^\{t\},\\kappa\_\{m\}^\{t\},y^\{t\}\),\\quad\\text\{if \}q\\leq N\_\{b\}\.\(24\)Otherwise,ℬm\\mathcal\{B\}\_\{m\}remains unchanged\.

This strategy guarantees that after observingttsamples, each sample has equal probability of being retained in the memory bank:

Pr⁡\(\(zms,μms,κms,ys\)∈ℬm\)=min⁡\(1,Nbt\),∀s∈\{1,…,t\}\.\\Pr\\left\(\(z\_\{m\}^\{s\},\\mu\_\{m\}^\{s\},\\kappa\_\{m\}^\{s\},y^\{s\}\)\\in\\mathcal\{B\}\_\{m\}\\right\)=\\min\\left\(1,\\frac\{N\_\{b\}\}\{t\}\\right\),\\quad\\forall s\\in\\\{1,\\dots,t\\\}\.\(25\)Whent≤Nbt\\leq N\_\{b\}, all observed samples are stored\. Whent\>Nbt\>N\_\{b\}, each observed sample is retained with probabilityNb/tN\_\{b\}/t, regardless of its arrival order\.

For completeness, we briefly justify this property\. For a sample observed at steps≤Nbs\\leq N\_\{b\}, it is initially inserted into the memory bank\. At any later stepτ\>Nb\\tau\>N\_\{b\}, it is replaced only if the new sample is accepted and its stored position is selected\. This happens with probability

Nbτ⋅1Nb=1τ\.\\frac\{N\_\{b\}\}\{\\tau\}\\cdot\\frac\{1\}\{N\_\{b\}\}=\\frac\{1\}\{\\tau\}\.\(26\)Therefore, the probability that this sample survives from stepNb\+1N\_\{b\}\+1to stepttis

∏τ=Nb\+1t\(1−1τ\)=∏τ=Nb\+1tτ−1τ=Nbt\.\\prod\_\{\\tau=N\_\{b\}\+1\}^\{t\}\\left\(1\-\\frac\{1\}\{\\tau\}\\right\)=\\prod\_\{\\tau=N\_\{b\}\+1\}^\{t\}\\frac\{\\tau\-1\}\{\\tau\}=\\frac\{N\_\{b\}\}\{t\}\.\(27\)For a sample observed at steps\>Nbs\>N\_\{b\}, it is inserted with probabilityNb/sN\_\{b\}/s\. If inserted, it survives each subsequent stepτ=s\+1,…,t\\tau=s\+1,\\dots,twith probability1−1/τ1\-1/\\tau\. Hence its final retention probability is

Nbs​∏τ=s\+1t\(1−1τ\)=Nbs​∏τ=s\+1tτ−1τ=Nbs⋅st=Nbt\.\\frac\{N\_\{b\}\}\{s\}\\prod\_\{\\tau=s\+1\}^\{t\}\\left\(1\-\\frac\{1\}\{\\tau\}\\right\)=\\frac\{N\_\{b\}\}\{s\}\\prod\_\{\\tau=s\+1\}^\{t\}\\frac\{\\tau\-1\}\{\\tau\}=\\frac\{N\_\{b\}\}\{s\}\\cdot\\frac\{s\}\{t\}=\\frac\{N\_\{b\}\}\{t\}\.\(28\)Thus, reservoir sampling maintains a uniformly sampled subset of historical samples under the fixed capacity constraint\.

In our implementation, the above procedure is applied independently for each modality\-specific memory bank\. At the beginning of each training epoch, we rebuild the memory bank by forwarding training samples through the AVIB module without gradient updates and updatingℬm\\mathcal\{B\}\_\{m\}using the reservoir sampling rule\. This yields a balanced and order\-agnostic memory bank for neighbor retrieval\. When the memory bank is fully populated, RME retrieves candidate neighbors fromℬm\\mathcal\{B\}\_\{m\}; otherwise, it uses intra\-batch candidates as described in Section[3\.4](https://arxiv.org/html/2609.30470#S3.SS4)\.

## Appendix DExperimental Details

### D\.1Datasets Information

We evaluate the proposed RCE on four benchmark datasets covering three tasks: multimodal sentiment analysis \(MSA\), multimodal humor detection \(MHD\), and multimodal sarcasm detection \(MSD\)\. We next provide a brief description of each dataset\.

CMU\-MOSI\[[61](https://arxiv.org/html/2609.30470#bib.bib4)\]: A widely adopted benchmark for MSA, consisting of over 2,000 video segments collected from online sources\. Each segment is assigned a sentiment intensity score on a seven\-point Likert scale ranging from−3\-3\(most negative\) to33\(most positive\)\.

CMU\-MOSEI\[[62](https://arxiv.org/html/2609.30470#bib.bib5)\]: One of the largest and most diverse MSA datasets, comprising more than 22,000 video segments from over 1,000 YouTube speakers across approximately 250 topics\. Each segment is annotated with both categorical emotions \(six classes\) and sentiment scores on the same−3\-3to33scale as CMU\-MOSI\. In our experiments, we focus on the sentiment annotations for consistency\.

UR\-FUNNY\[[10](https://arxiv.org/html/2609.30470#bib.bib3)\]: A benchmark dataset for MHD, derived from TED talk videos featuring 1,741 speakers\. Each target segment, referred to as a punchline, is annotated across language, acoustic, and visual modalities, while preceding segments serve as contextual input\. Punchlines are identified using thelaughtertag in transcripts, which indicates audience laughter; negative samples are constructed when no such cue is present\. The dataset is split into 7,614 training, 980 validation, and 994 testing samples\. Following prior works\[[9](https://arxiv.org/html/2609.30470#bib.bib20),[15](https://arxiv.org/html/2609.30470#bib.bib49),[53](https://arxiv.org/html/2609.30470#bib.bib52)\], we adopt version 2 of UR\-FUNNY\.

MUStARD\[[1](https://arxiv.org/html/2609.30470#bib.bib2)\]: A dataset for MSD, collected from popular television series such asFriends,The Big Bang Theory,The Golden Girls, andSarcasmaholics\. It contains 690 video segments labeled as sarcastic or non\-sarcastic\. Each instance includes the target punchline along with preceding dialogue to provide context\.

### D\.2Baselines

This section offers a thorough overview of the baseline methods utilized in our study\.

MODS\[[56](https://arxiv.org/html/2609.30470#bib.bib58)\]\[AAAI’26\]: Modality Optimization and Dynamic Primary Modality Selection \(MODS\) tackles modality imbalance by adaptively identifying the most informative modality for each sample\. It incorporates a graph\-based dynamic sequence compression module to minimize redundancy in non\-linguistic modalities, and leverages a primary\-modality\-focused cross\-attention mechanism to strengthen cross\-modal integration\.

PSA\-MF\[[52](https://arxiv.org/html/2609.30470#bib.bib60)\]\[AAAI’26\]: Personality\-Sentiment Aligned Multi\-level Fusion \(PSA\-MF\) integrates personality traits into the feature extraction process to learn personalized sentiment representations, and employs a hierarchical fusion strategy to progressively combine multimodal information, thereby enhancing sentiment recognition performance;

DecAlign\[[39](https://arxiv.org/html/2609.30470#bib.bib56)\]\[ICLR’26\]: Decoupled Cross\-modal Alignment \(DecAlign\) tackles multimodal heterogeneity by decomposing representations into modality\-specific and shared components\. It employs a prototype\-guided optimal transport strategy to align modality\-specific features and leverages MMD\-based distribution matching to enforce cross\-modal consistency, further enhanced by a multimodal transformer for high\-level fusion\.

MIB\[[33](https://arxiv.org/html/2609.30470#bib.bib30)\]\[TMM’23\]: Multimodal Information Bottleneck \(MIB\) applies the information bottleneck principle to learn compact, task\-relevant representations by reducing redundancy and noise in both unimodal and multimodal features\. It provides three variants—E\-MIB, L\-MIB, and C\-MIB—that impose information constraints at different stages of multimodal fusion\.

ITHP\[[51](https://arxiv.org/html/2609.30470#bib.bib39)\]\[ICLR’24\]: Information\-Theoretic Hierarchical Perception \(ITHP\) leverages the information bottleneck principle by assigning a primary modality and utilizing auxiliary modalities as detectors to extract salient information\.

KAN\-MCP\[[30](https://arxiv.org/html/2609.30470#bib.bib51)\]\[ACM MM’25\]: Kolmogorov–Arnold Network with Multimodal Clean Pareto \(KAN\-MCP\) combines interpretable cross\-modal modeling with DRD\-MIB\-based feature compression and denoising, enabling discriminative representation learning while mitigating modality imbalance\.

OMIB\[[49](https://arxiv.org/html/2609.30470#bib.bib45)\]\[ICML’25\]: Optimal Multimodal Information Bottleneck \(OMIB\) extends the MIB framework by theoretically constraining regularization weights to ensure an optimal bottleneck representation, while dynamically adjusting modality\-specific regularization to alleviate modality imbalance\.

QMF\[[65](https://arxiv.org/html/2609.30470#bib.bib33)\]\[ICML’23\]: Quality\-aware Multimodal Fusion \(QMF\) provides a theoretically grounded framework for robust multimodal learning by analyzing multimodal fusion from a generalization perspective\. It leverages uncertainty estimation to dynamically assess modality quality, enabling adaptive cross\-modal interaction and mitigating the impact of low\-quality data, thereby promoting more reliable and quality\-aware feature integration\.

PML\[[14](https://arxiv.org/html/2609.30470#bib.bib61)\]\[IJCAI’25\]: Probabilistic Multimodal Learning \(PML\) models representations with von Mises–Fisher \(vMF\) distributions to capture intrinsic uncertainty and learn reliable directional features\. It employs a vMF\-based prototypical contrastive learning paradigm to enhance class discrimination, and introduces a reliability\-aware fusion mechanism to dynamically regulate modality contributions and resolve cross\-modal conflicts\.

HME\[[70](https://arxiv.org/html/2609.30470#bib.bib43)\]\[NeurIPS’25\]: Hyper\-Modality Enhancement \(HME\) tackles missing modality scenarios by enriching observed modalities with semantically relevant cues from other samples, avoiding explicit reconstruction\. It further adopts an uncertainty\-aware fusion strategy to adaptively balance original and enhanced representations, improving robustness to incomplete and noisy inputs\.

CyIN\[[27](https://arxiv.org/html/2609.30470#bib.bib44)\]\[NeurIPS’25\]: Cyclic INformative Learning \(CyIN\) addresses dynamic missing modality scenarios by constructing an informative latent space via token\- and label\-level Information Bottleneck \(IB\) applied cyclically across modalities\. It further introduces cross\-modal cyclic translation to reconstruct missing modalities through forward and reverse processes, enabling unified optimization for both complete and incomplete multimodal learning while enhancing cross\-modal interaction and fusion\.

Self\-MM\[[58](https://arxiv.org/html/2609.30470#bib.bib19)\]\[AAAI’21\]: Self\-Supervised Multi\-task Multimodal \(Self\-MM\) sentiment analysis framework generates pseudo labels for each modality based on annotated global sentiment labels, enabling the learning of more discriminative unimodal representations\.

FDMER\[[55](https://arxiv.org/html/2609.30470#bib.bib23)\]\[ACM MM’22\]: Feature\-Disentangled Multimodal Emotion Recognition \(FDMER\) alleviates distribution gaps and redundancy across modalities by decomposing features into modality\-invariant and modality\-specific subspaces\. It employs common and private encoders with adversarial training to enforce consistency and diversity, and utilizes a cross\-modal attention fusion module to learn adaptive multimodal representations\.

DMD\[[23](https://arxiv.org/html/2609.30470#bib.bib27)\]\[CVPR’23\]: Decoupled Multimodal Distillation \(DMD\) improves emotion recognition by separating each modality into modality\-relevant and modality\-exclusive components, and conducting adaptive cross\-modal knowledge distillation through a dynamic graph structure\.

AGM\[[20](https://arxiv.org/html/2609.30470#bib.bib31)\]\[ICCV’23\]: Adaptive Gradient Modulation \(AGM\) addresses modality competition in multimodal learning by adaptively adjusting gradients during training to balance modality contributions across various fusion strategies\. It further introduces a novel metric based on the mono\-modal concept to quantify competition strength, providing insights into modality dominance and improving overall model performance\.

AtCAF\[[15](https://arxiv.org/html/2609.30470#bib.bib49)\]\[Information Fusion’25\]: Attention\-based Causality\-Aware Fusion \(AtCAF\) learns causality\-aware multimodal representations through a text debiasing module and counterfactual cross\-modal attention for sentiment analysis\.

SuCI\[[53](https://arxiv.org/html/2609.30470#bib.bib52)\]\[AAAI’25\]: Subject Causal Intervention \(SuCI\) mitigates subject\-related spurious correlations by modeling subjects as confounders and applying causal intervention, leading to more robust multimodal language understanding and improved cross\-subject generalization\.

DHMD\[[24](https://arxiv.org/html/2609.30470#bib.bib59)\]\[TPAMI’26\]: Decoupled Hierarchical Multimodal Distillation \(DHMD\) separates multimodal features into modality\-invariant and modality\-specific components via a self\-regression mechanism, and adopts a hierarchical distillation framework—combining graph\-based coarse\-grained and dictionary\-based fine\-grained distillation—to improve cross\-modal alignment and emotion recognition\.

MAG\[[40](https://arxiv.org/html/2609.30470#bib.bib17)\]\[ACL’20\]: Multimodal Adaptation Gate \(MAG\) introduces an adaptation gate that enables large pre\-trained transformers to incorporate multimodal signals during the fine\-tuning stage\.

BBFN\[[8](https://arxiv.org/html/2609.30470#bib.bib21)\]\[ICMI’21\]: Bi\-Bimodal Fusion Network \(BBFN\) conducts both fusion and separation over pairwise modality representations, and utilizes a gated Transformer to alleviate modality imbalance\.

MuLOT\[[38](https://arxiv.org/html/2609.30470#bib.bib24)\]\[WACV’22\]: Multimodal Learning using Optimal Transport \(MuLOT\) addresses resource\-constrained sarcasm and humor detection by leveraging self\-attention to model intra\-modal relationships and optimal transport to capture cross\-modal alignment, followed by multimodal attention fusion to integrate complementary information across modalities\.

MIL\[[66](https://arxiv.org/html/2609.30470#bib.bib32)\]\[TAI’23\]: Multimodal Interaction Learning \(MIL\) is a multitask framework for joint sarcasm and sentiment detection that models both shared and task\-specific features\. It introduces a cross\-modal target attention mechanism to enhance multimodal fusion and facilitate interaction between related tasks\.

MOAC\[[34](https://arxiv.org/html/2609.30470#bib.bib47)\]\[WWW’25\]: Multimodal Ordinal Affective Computing \(MOAC\) improves affective modeling by incorporating ordinal learning at both the label level \(coarse\-grained\) and feature level \(fine\-grained\) within multimodal representations\.

### D\.3Feature Extraction

Textual Modality:For the CMU\-MOSI and CMU\-MOSEI datasets, we adopt DeBERTa\[[13](https://arxiv.org/html/2609.30470#bib.bib7)\]to extract high\-level textual representations, following prior work\. For the UR\-FUNNY dataset, ALBERT\[[19](https://arxiv.org/html/2609.30470#bib.bib11)\]is employed as the text encoder\. Specifically, the context and punchline token sequences are concatenated to form the input sequence:Ul=Cl⊕\[SEP\]⊕PlU\_\{l\}=C\_\{l\}\\oplus\\texttt\{\[SEP\]\}\\oplus P\_\{l\}, where\[SEP\]is used to separate the context tokensClC\_\{l\}and punchline tokensPlP\_\{l\}\. For the MUStARD dataset, contextual word representations are extracted using a pretrained BERT\[[3](https://arxiv.org/html/2609.30470#bib.bib6)\]model\.

For the experiments reported in Figure[2](https://arxiv.org/html/2609.30470#S4.F2)of the main paper, in order to ensure a fair comparison with CyIN\[[27](https://arxiv.org/html/2609.30470#bib.bib44)\], both RCE and HME\[[70](https://arxiv.org/html/2609.30470#bib.bib43)\]adopt BERT\[[3](https://arxiv.org/html/2609.30470#bib.bib6)\]to extract textual modality features\.

Acoustic Modality:Acoustic features for CMU\-MOSI, CMU\-MOSEI, UR\-FUNNY, and MUStARD are extracted using COVAREP\[[2](https://arxiv.org/html/2609.30470#bib.bib9)\], including 12\-dimensional MFCCs, pitch, speech polarity, glottal closure instants, and spectral envelope descriptors\. These features are computed over the full audio segment of each utterance, forming temporal sequences that capture dynamic vocal variations\.

Visual Modality:For CMU\-MOSI and CMU\-MOSEI, visual features are obtained using Facet \(iMotions 2017,[https://imotions\.com/](https://imotions.com/)\), including facial action units, landmarks, head pose, and other expression\-related cues, organized as temporal sequences to capture facial dynamics\. For UR\-FUNNY and MUStARD, OpenFace\[[60](https://arxiv.org/html/2609.30470#bib.bib10)\]is used to extract facial action units along with both rigid and non\-rigid facial shape parameters\.

For the CMU\-MOSI dataset, the feature dimensions of the textual, acoustic, and visual modalities are 768, 74, and 47, respectively\. For CMU\-MOSEI, the corresponding dimensions are 768, 74, and 35\. For both UR\-FUNNY and MUStARD, the dimensions of the textual, acoustic, and visual modalities are 768, 60, and 36, respectively\. In addition, UR\-FUNNY includes an extra humor centric feature \(HCF\) modality with a dimension of 4\. Details of the HCF feature extraction process can be found in\[[9](https://arxiv.org/html/2609.30470#bib.bib20)\]\.

### D\.4Implementation Details

We implement RCE using the PyTorch framework and conduct all experiments on a single NVIDIA GeForce RTX 4090 GPU, using CUDA 12\.8 and PyTorch 2\.8\.0\. The model is optimized using AdamW\[[29](https://arxiv.org/html/2609.30470#bib.bib12)\]with linear warm\-up\. Detailed hyper\-parameter settings for different datasets are summarized in Table[5](https://arxiv.org/html/2609.30470#A4.T5)\.

Several hyper\-parameters are shared across all datasets\. Specifically, warm\-up is applied in all experiments, and AdamW is consistently adopted as the optimizer\. The modality enhancement process uses a modality\-specific memory bank with capacityNbN\_\{b\}, where neighbor retrieval is performed based on the top\-kkratioρ\\rho\. The prompt length is denoted bylpl\_\{p\}, the label penalty weight byα\\alpha, the AVIB trade\-off coefficient byβ\\beta, and the overall loss balancing weight byλ\\lambda\.

For the MSA datasets CMU\-MOSI and CMU\-MOSEI, we use relatively larger feature dimensions and memory bank capacities to capture richer multimodal contextual information\. Specifically, the feature dimensions are set tod=256d=256andd=192d=192, respectively, while the memory bank capacities are set to10241024and20482048\. The batch sizes are set to4848for CMU\-MOSI and128128for CMU\-MOSEI, with initial learning rates of2×10−52\\times 10^\{\-5\}and3×10−53\\times 10^\{\-5\}, respectively\. Both datasets use a dropout rate of0\.40\.4and AVIB trade\-off coefficientβ=1​e−3\\beta=1\\mathrm\{e\}\{\-3\}\. For the MHD and MSD datasets UR\-FUNNY and MUStARD, we adopt comparatively smaller dropout rates to preserve discriminative multimodal cues\. The feature dimensions are set tod=192d=192andd=144d=144, respectively\. The prompt lengths differ across datasets, withlp=3l\_\{p\}=3for UR\-FUNNY andlp=8l\_\{p\}=8for MUStARD\. Both datasets use a memory bank capacity of10241024, while the top\-kkratios are set to0\.30\.3and0\.50\.5, respectively\. The AVIB trade\-off coefficients are set to1​e−41\\mathrm\{e\}\{\-4\}for UR\-FUNNY and1​e−21\\mathrm\{e\}\{\-2\}for MUStARD\.

To determine the optimal configuration, we perform a grid search with thirty random trials on the validation set\. The batch size is selected from32,48,64,128,256\{32,48,64,128,256\}, the initial learning rate is searched over1​e−6,3​e−6,8​e−6,1​e−5,2​e−5,3​e−5\{1\\mathrm\{e\}\{\-6\},3\\mathrm\{e\}\{\-6\},8\\mathrm\{e\}\{\-6\},1\\mathrm\{e\}\{\-5\},2\\mathrm\{e\}\{\-5\},3\\mathrm\{e\}\{\-5\}\}, and the feature dimensionddis selected from96,144,192,256\{96,144,192,256\}\. The dropout rate is tuned within0\.1,0\.2,0\.3,0\.4,0\.5\{0\.1,0\.2,0\.3,0\.4,0\.5\}, the prompt lengthlpl\_\{p\}is searched over2,3,4,6,8\{2,3,4,6,8\}, and the memory bank capacityNbN\_\{b\}is selected from256,512,1024,2048\{256,512,1024,2048\}\. In addition, the top\-kkratioρ\\rhois tuned within0\.1,0\.2,…,0\.7\{0\.1,0\.2,\\dots,0\.7\}, while the label penalty weightα\\alphaand loss balancing weightλ\\lambdaare tuned within0\.1,0\.2,…,1\.0\{0\.1,0\.2,\\dots,1\.0\}\. The AVIB trade\-off coefficientβ\\betais searched over1​e−2,1​e−3,1​e−4,1​e−5\{1\\mathrm\{e\}\{\-2\},1\\mathrm\{e\}\{\-3\},1\\mathrm\{e\}\{\-4\},1\\mathrm\{e\}\{\-5\}\}\. The final model is selected based on the configuration that achieves the lowestMAEon the validation set\.

Notably, for a fair comparison, we re\-implement C\-MIB\[[33](https://arxiv.org/html/2609.30470#bib.bib30)\], ITHP\[[51](https://arxiv.org/html/2609.30470#bib.bib39)\], KAN\-MCP\[[30](https://arxiv.org/html/2609.30470#bib.bib51)\], OMIB\[[49](https://arxiv.org/html/2609.30470#bib.bib45)\], QMF\[[65](https://arxiv.org/html/2609.30470#bib.bib33)\], PML\[[14](https://arxiv.org/html/2609.30470#bib.bib61)\], and HME\[[70](https://arxiv.org/html/2609.30470#bib.bib43)\]using the same feature extraction strategy as RCE\. The implementations are based on their publicly available official codebases, and the hyper\-parameter search is conducted according to the parameter ranges reported in their original papers using thirty random search trials on the validation set to determine the optimal configuration\. For the remaining baselines, unless otherwise specified, the reported results are directly taken from the corresponding original papers\.

For the MHD task, the results of Self\-MM\[[58](https://arxiv.org/html/2609.30470#bib.bib19)\], FDMER\[[55](https://arxiv.org/html/2609.30470#bib.bib23)\], and DMD\[[23](https://arxiv.org/html/2609.30470#bib.bib27)\]are taken from the SuCI\[[53](https://arxiv.org/html/2609.30470#bib.bib52)\]paper\. For the MSD task, the results of MAG\-XLNet\[[40](https://arxiv.org/html/2609.30470#bib.bib17)\]and BBFN\[[8](https://arxiv.org/html/2609.30470#bib.bib21)\]are taken from the MuLOT\[[38](https://arxiv.org/html/2609.30470#bib.bib24)\]paper\. Results for all other baselines are directly taken from their corresponding original papers\.

Table 5:Hyper\-parameter settings of RCE across different datasets\.Hyper\-parameterCMU\-MOSICMU\-MOSEIUR\-FUNNYMUStARDBatch Size4812825632Epochs50101020Warm\-up✓✓✓✓Initial Learning Rate2×10−52\\times 10^\{\-5\}3×10−53\\times 10^\{\-5\}1×10−51\\times 10^\{\-5\}3×10−63\\times 10^\{\-6\}OptimizerAdamWAdamWAdamWAdamWDropout Rate0\.40\.40\.10\.1Feature Dimensiondd256192192144Prompt Lengthlpl\_\{p\}4638Memory Bank CapacityNbN\_\{b\}1024204810241024Top\-kkRatioρ\\rho0\.60\.40\.30\.5Label Penalty Weightα\\alpha0\.40\.10\.30\.3AVIB Trade\-off Coefficientβ\\beta1e\-31e\-31e\-41e\-2Loss Balancing Weightλ\\lambda0\.71\.00\.60\.4

### D\.5Evaluation Metrics

We evaluate the model’s performance on the MSA task using a set of well\-established metrics, reported for both CMU\-MOSI and CMU\-MOSEI datasets\. For interpretability, classification results are presented as percentages\. These metrics are calculated as follows:

Seven\-category Classification Accuracy \(Acc7\):Measures the model’s ability to predict fine\-grained sentiment categories by dividing the sentiment score range \(−3\-3to33\) into seven equal intervals\. The metric is defined as:

Acc7=1n​∑i=1n𝟏​\(c^i=ci\),\\mathrm\{Acc7\}=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\mathbf\{1\}\\big\(\\hat\{c\}\_\{i\}=c\_\{i\}\\big\),\(29\)wherecic\_\{i\}andc^i\\hat\{c\}\_\{i\}denote the ground\-truth and predicted categories of sampleii, respectively, and𝟏​\(⋅\)\\mathbf\{1\}\(\\cdot\)is the indicator function\. Higher values indicate better fine\-grained sentiment classification\.

Binary Classification Accuracy \(Acc2\):Following prior work\[[51](https://arxiv.org/html/2609.30470#bib.bib39),[30](https://arxiv.org/html/2609.30470#bib.bib51)\], we report binary sentiment classification results by distinguishing negative \(<0<0\) and positive \(\>0\>0\) samples, excluding neutral cases\. The metric is formulated as:

Acc2=T​P\+T​NT​P\+T​N\+F​P\+F​N,\\mathrm\{Acc2\}=\\frac\{TP\+TN\}\{TP\+TN\+FP\+FN\},\(30\)whereT​PTP,T​NTN,F​PFP, andF​NFNdenote true positives, true negatives, false positives, and false negatives, respectively\.

Weighted F1\-score \(F1\):Computes the harmonic mean of precision and recall while considering class\-specific weights to mitigate imbalance\. It is formulated as:

F1=2⋅P​r​e​c​i​s​i​o​n⋅R​e​c​a​l​lP​r​e​c​i​s​i​o​n\+R​e​c​a​l​l,\\mathrm\{F1\}=2\\cdot\\frac\{Precision\\cdot Recall\}\{Precision\+Recall\},\(31\)wherePrecision=T​PT​P\+F​P\\mathrm\{Precision\}=\\tfrac\{TP\}\{TP\+FP\}andRecall=T​PT​P\+F​N\\mathrm\{Recall\}=\\tfrac\{TP\}\{TP\+FN\}\.

Mean Absolute Error \(MAE\):Represents the average magnitude of prediction errors with respect to the ground\-truth sentiment scores\. It directly corresponds to the original sentiment scale, making it both intuitive and informative:

MAE⁡\(y^,y\)=1n​∑i=1n\|y^i−yi\|,\\mathrm\{MAE\}\(\\hat\{y\},y\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\|\\hat\{y\}\_\{i\}\-y\_\{i\}\|,\(32\)whereyiy\_\{i\}is the true label,y^i\\hat\{y\}\_\{i\}is the predicted value, andnnis the total number of predictions\.

Pearson Correlation Coefficient \(Corr\):Quantifies the strength and direction of the linear relationship between predicted and true sentiment scores:

Corr⁡\(x,y\)=∑i=1n\(xi−x¯\)​\(yi−y¯\)∑i=1n\(xi−x¯\)2​∑i=1n\(yi−y¯\)2,\\mathrm\{Corr\}\(x,y\)=\\frac\{\\sum\_\{i=1\}^\{n\}\(x\_\{i\}\-\\bar\{x\}\)\(y\_\{i\}\-\\bar\{y\}\)\}\{\\sqrt\{\\sum\_\{i=1\}^\{n\}\(x\_\{i\}\-\\bar\{x\}\)^\{2\}\}\\sqrt\{\\sum\_\{i=1\}^\{n\}\(y\_\{i\}\-\\bar\{y\}\)^\{2\}\}\},\(33\)wherexix\_\{i\}andyiy\_\{i\}denote predicted and ground\-truth values, respectively, andx¯,y¯\\bar\{x\},\\bar\{y\}are their means\.

For the MHD and MSD tasks, following prior work\[[53](https://arxiv.org/html/2609.30470#bib.bib52),[15](https://arxiv.org/html/2609.30470#bib.bib49),[34](https://arxiv.org/html/2609.30470#bib.bib47),[24](https://arxiv.org/html/2609.30470#bib.bib59)\], we report binary accuracy only, which measures the model’s capability to differentiate between humorous and non\-humorous, as well as sarcastic and non\-sarcastic, instances\.

## Appendix EMore Experimental Results

Table 6:Performance comparison of RCE and baseline methods under fixed missing protocol\.### E\.1Results under Fixed Missing Protocols

Under the fixed missing protocol, we further evaluate RCE against a wide range of competitive missing\-modality methods, including C\-MIB\[[33](https://arxiv.org/html/2609.30470#bib.bib30)\], ITHP\[[51](https://arxiv.org/html/2609.30470#bib.bib39)\], KAN\-MCP\[[30](https://arxiv.org/html/2609.30470#bib.bib51)\], OMIB\[[49](https://arxiv.org/html/2609.30470#bib.bib45)\], QMF\[[65](https://arxiv.org/html/2609.30470#bib.bib33)\], PML\[[14](https://arxiv.org/html/2609.30470#bib.bib61)\], and HME\[[70](https://arxiv.org/html/2609.30470#bib.bib43)\]\. As reported in Table[6](https://arxiv.org/html/2609.30470#A5.T6), RCE achieves the best overall performance across most fixed modality\-availability settings\. Specifically, when only the textual modality is available, RCE obtains the best results on four out of five metrics on CMU\-MOSI and three out of five metrics on CMU\-MOSEI, showing that the proposed method can still produce reliable predictions under highly sparse modality conditions\. When both text and audio are available, RCE consistently outperforms all baselines on both datasets across all evaluation metrics, demonstrating its strong ability to exploit complementary acoustic cues in addition to the dominant textual modality\.

For the setting where text and vision are available, RCE also exhibits highly competitive and stable performance\. On CMU\-MOSEI, it achieves the best results on all five metrics, while on CMU\-MOSI it obtains the lowestM​A​EMAEand highest correlation, withA​c​c​2Acc2andF​1F1remaining very close to the best\-performing baseline\. These results suggest that RCE is not limited to a specific modality combination, but can adaptively benefit from different observed modalities\. Overall, under the fixed missing protocol, RCE achieves the best performance on 24 out of 30 metric comparisons, validating its robustness to diverse deterministic missing patterns\. The improvements can be attributed to the reliability\-aware modality enhancement mechanism, which retrieves informative cross\-sample cues from the memory bank, and the multilevel fusion strategy, which adaptively integrates unimodal, enhanced, and cross\-modal representations according to modality reliability\.

### E\.2Results under Noisy Modalities

In addition to missing\-modality evaluation, we conduct noise robustness experiments to further examine whether RCE can handle corrupted but still observable modalities\.

Table 7:Comparison under different Gaussian noise intensities on CMU\-MOSI and CMU\-MOSEI\.#### Noise setting\.

Following the noise evaluation protocols used in EAU\[[6](https://arxiv.org/html/2609.30470#bib.bib36)\]and QMF\[[65](https://arxiv.org/html/2609.30470#bib.bib33)\], we further examine the robustness of RCE under corrupted multimodal inputs\. Since the original noise operations are mainly designed for image\-like data, we adapt them to multimodal sentiment analysis with text tokens and continuous acoustic/visual feature sequences\. For the textual modality, noise is directly applied to the encodedinput\_idsby replacing selected token ids, which minimally changes the preprocessing pipeline\. Textual noise can be applied to the train, validation, and test splits\. For acoustic and visual modalities, we adapt Gaussian and Salt\-pepper noise to continuous feature sequences\. Specifically, Gaussian noise is not generated using image\-style amplitude scaling; instead, the perturbation scale is determined by the standard deviation of valid features in each sample multiplied bynoise/10\\texttt\{noise\}/10\. For Salt\-pepper noise, rather than replacing pixels with 0 or 255, we replace selected valid feature elements with the minimum or maximum value of the corresponding sample\. Acoustic and visual noise is applied only to the test split to evaluate robustness under corrupted inference\-time observations\. The noise intensity parameternoiseranges from 0 to 10, where larger values indicate stronger corruption\. For text, after a sample\-level trigger with probability 0\.5, each token is replaced with probabilitynoise/10\\texttt\{noise\}/10\. For acoustic and visual Gaussian noise, a sample\-level trigger with probability 0\.5 determines whether noise is added, and the perturbation magnitude is proportional tonoise/10\\texttt\{noise\}/10\. For acoustic and visual Salt\-pepper noise, the corruption process is first triggered with probability 0\.5, and the corruption strength is further controlled bynoise/10\\texttt\{noise\}/10\.

Table 8:Comparison under different Salt\-pepper noise intensities on CMU\-MOSI and CMU\-MOSEI\.
#### Results under Gaussian noise\.

Table[7](https://arxiv.org/html/2609.30470#A5.T7)reports the comparison under different Gaussian noise intensities\. Overall, RCE demonstrates strong robustness across both datasets, especially under moderate and severe Gaussian perturbations\. When the noise intensity is 1\.0, RCE achieves the best results on most metrics\. On CMU\-MOSI, RCE obtains the bestA​c​c​7Acc7,A​c​c​2Acc2,F​1F1, andM​A​EMAE, while itsC​o​r​rCorrremains very close to the best result\. On CMU\-MOSEI, RCE achieves the bestA​c​c​7Acc7,M​A​EMAE, andC​o​r​rCorr, and remains highly competitive onA​c​c​2Acc2andF​1F1\. These results indicate that when the perturbation is mild, RCE can effectively preserve useful sentiment cues while suppressing unreliable corrupted information\.

As the Gaussian noise intensity increases to 5\.0, all methods experience performance degradation, but RCE still maintains competitive performance\. On CMU\-MOSI, RCE achieves the bestA​c​c​2Acc2andF​1F1, showing that it preserves binary sentiment discrimination under stronger feature perturbation, although HME\[[70](https://arxiv.org/html/2609.30470#bib.bib43)\]obtains betterA​c​c​7Acc7,M​A​EMAE, andC​o​r​rCorrin this setting\. On CMU\-MOSEI, RCE achieves the bestA​c​c​2Acc2,F​1F1, andM​A​EMAE, demonstrating stable performance on the larger\-scale dataset\. Under the most severe Gaussian noise intensity of 10\.0, RCE shows clearer advantages\. On CMU\-MOSI, it achieves the best performance across all five metrics, improvingA​c​c​2Acc2andF​1F1notably compared with the strongest baselines\. On CMU\-MOSEI, RCE also obtains the bestA​c​c​7Acc7,A​c​c​2Acc2, andM​A​EMAE, while remaining close to the bestC​o​r​rCorr\. These results suggest that the reliability\-aware enhancement and multilevel fusion mechanisms help RCE maintain robust predictions when acoustic and visual observations are heavily corrupted\.

#### Results under Salt\-pepper noise\.

Table[8](https://arxiv.org/html/2609.30470#A5.T8)presents the results under Salt\-pepper noise with different corruption intensities\. Compared with Gaussian noise, Salt\-pepper noise introduces more abrupt feature\-level corruption by replacing valid acoustic and visual elements with extreme values from the same sample, making it a challenging test of robustness to outlier\-like perturbations\[[65](https://arxiv.org/html/2609.30470#bib.bib33)\]\. When the noise intensity is 1\.0, RCE achieves the bestA​c​c​7Acc7,A​c​c​2Acc2,F​1F1, andM​A​EMAEon CMU\-MOSI, and the bestA​c​c​7Acc7,A​c​c​2Acc2,F​1F1, andC​o​r​rCorron CMU\-MOSEI\. This shows that RCE is effective under mild sparse corruption and can still extract reliable complementary information from noisy continuous features\.

When the Salt\-pepper noise intensity increases to 5\.0, the performance of all methods decreases more noticeably\. On CMU\-MOSI, RCE achieves the bestA​c​c​2Acc2,F​1F1,M​A​EMAE, andC​o​r​rCorr, and remains close to the bestA​c​c​7Acc7\. This indicates that RCE is particularly effective at preserving overall sentiment discrimination and regression quality under medium\-level outlier corruption\. On CMU\-MOSEI, however, KAN\-MCP\[[30](https://arxiv.org/html/2609.30470#bib.bib51)\]and HME\[[70](https://arxiv.org/html/2609.30470#bib.bib43)\]obtain better results on several metrics, while RCE becomes less competitive\. This suggests that medium Salt\-pepper corruption may introduce feature\-level outliers that affect RCE’s retrieved or fused representations on the larger dataset\. Nevertheless, under the strongest Salt\-pepper noise intensity of 10\.0, RCE again shows robust overall performance\. It achieves the bestA​c​c​7Acc7,A​c​c​2Acc2,F​1F1, andM​A​EMAEon CMU\-MOSI, and the bestA​c​c​7Acc7,A​c​c​2Acc2,M​A​EMAE, andC​o​r​rCorron CMU\-MOSEI\. These results demonstrate that RCE is generally resilient to severe outlier\-style corruption\. The advantage is especially clear under high noise intensity, where reliability\-aware fusion can reduce the influence of unreliable modality information and better exploit stable cues from the remaining useful representations\.

### E\.3Detailed Ablation Study

Table 9:Ablation details\.To further examine the contribution of each component in RCE, we report detailed ablation results in Table[9](https://arxiv.org/html/2609.30470#A5.T9)\. Specifically, ‘w/o AVIB’ removes the Adaptive Variational Information Bottleneck, where modality representations are directly used without vMF\-based uncertainty estimation, latent sampling, and KL regularization\. ‘w/o RME’ removes the Reliability\-aware Modality Enhancement module by directly setting the enhanced representation as the original modality representation\. ‘w/o MB’ keeps the enhancement mechanism but disables the memory bank, using only mini\-batch samples for neighbor retrieval\. ‘w/o LP’ removes the label\-distance penalty in neighbor weighting, while ‘w/o CP’ removes the confidence promotion term based onκ\\kappa\. ‘w/o HMG’ removes the hyper\-modality generation module and directly uses modality\-level representations for fusion\. ‘w/o ES’ removes the shared sample\-level representation, and ‘w/o RF’ disables reliability\-aware fusion by replacingκ\\kappa\-modulated fusion with vanilla attention weights\.

Overall, the complete RCE achieves the best or competitive results on most metrics, especially on CMU\-MOSI, where it obtains the bestA​c​c​2Acc2,F​1F1,M​A​EMAE, andC​o​r​rCorr\. This indicates that the proposed components are complementary and jointly improve both classification accuracy and regression quality\. On CMU\-MOSI, removing ES leads to the largest degradation inA​c​c​7Acc7andM​A​EMAE, withA​c​c​7Acc7dropping from 50\.22 to 45\.69 andM​A​EMAEincreasing from 0\.589 to 0\.654, showing that the shared sample\-level representation is important for capturing global multimodal context\. Removing HMG also causes a clear performance drop, especially inA​c​c​7Acc7andM​A​EMAE, indicating that hyper\-modality generation contributes to robust cross\-modal representation learning\. In addition, removing AVIB consistently weakens all metrics on CMU\-MOSI, demonstrating that vMF\-based uncertainty estimation and adaptive compression help reduce noisy and redundant modality information\. The memory bank also plays an important role: compared with RCE, ‘w/o MB’ obtains lowerA​c​c​7Acc7,M​A​EMAE, andC​o​r​rCorr, suggesting that historical reliable samples provide more stable neighborhood information than mini\-batch\-only retrieval\.

For the internal design of RME, the effects of label penalty and confidence promotion show different tendencies\. On CMU\-MOSI, ‘w/o LP’ achieves relatively strong classification results but is still inferior to RCE inA​c​c​2Acc2,F​1F1,M​A​EMAE, andC​o​r​rCorr, suggesting that label consistency is beneficial for stable prediction\. In contrast, removing CP causes a larger decline inA​c​c​7Acc7, indicating that confidence\-aware neighbor weighting helps suppress unreliable enhancement samples\. On CMU\-MOSEI, several ablated variants obtain slightly better results on individual metrics, such as ‘w/o LP’ onA​c​c​7Acc7and ‘w/o ES’ or ‘w/o RF’ onA​c​c​2Acc2/F​1F1\. This may be because CMU\-MOSEI contains more training samples and provides more stable multimodal statistics, making some auxiliary constraints less critical for coarse\-grained classification\. Nevertheless, the complete RCE still achieves the bestC​o​r​rCorrand nearly the bestM​A​EMAE, showing better overall regression consistency and prediction calibration\. In particular, although ‘w/o RF’ slightly improvesA​c​c​2Acc2andF​1F1on CMU\-MOSEI, it performs worse than RCE inM​A​EMAEandC​o​r​rCorr, indicating that reliability\-aware fusion is more helpful for fine\-grained sentiment intensity estimation\. Overall, these results show that AVIB, RME with memory\-based retrieval, HMG, ES, and RF contribute from different aspects, and their combination yields the most balanced performance across datasets and evaluation metrics\.

### E\.4Negative Enhancement Analysis

Table 10:Detailed training\-stage negative enhancement statistics on CMU\-MOSI and CMU\-MOSEI\.We further analyze the negative enhancement phenomenon during training\. For each sample, we compare the sentiment direction of its ground\-truth label with the sentiment direction of the samples used for enhancement\. Specifically, for HME, the retrieved samples whose similarities exceed the threshold are averaged to obtain the enhancement source\. For RCE, we use the final weighted neighbors after top\-kkretrieval, label\-distance penalty, and reliability\-aware weighting\. If the weighted average label of the enhancement sources has the opposite sentiment polarity to the current sample, this case is counted as negative enhancement\.

The reported rate is computed as the number of negative enhancement cases divided by the number of valid enhanced samples in the training stage\. In our logs, CMU\-MOSI contains 61,400 valid training enhancement cases over all epochs, corresponding to 1,228 valid training samples per epoch, while CMU\-MOSEI contains 127,350 valid training enhancement cases, corresponding to 12,735 valid training samples per epoch\. Neutral samples are not counted since their sentiment polarity is ambiguous\.

As shown in Table[3](https://arxiv.org/html/2609.30470#S4.T3), HME suffers from severe negative enhancement\. On CMU\-MOSI, nearly 46% of training samples are negatively enhanced across all modalities\. In contrast, RCE reduces the negative enhancement rates to 2\.60%, 5\.90%, and 6\.48% for text, audio, and vision, respectively\. This demonstrates that the proposed label\-aware and reliability\-aware enhancement mechanism effectively suppresses sentiment\-conflicting information from neighboring samples\.

On CMU\-MOSEI, RCE also consistently reduces negative enhancement compared with HME\. The reduction is most significant in the text modality, from 38\.40% to 10\.57%\. However, the decrease for the audio modality is relatively smaller\. This may be because acoustic cues in CMU\-MOSEI are more ambiguous and diverse: samples with similar acoustic patterns may still express different sentiment polarities\. As a result, even after reliability\-aware weighting, some sentiment\-conflicting audio neighbors can still receive non\-negligible weights, leading to a higher remaining negative enhancement rate\.

\(a\)AVIB Trade\-off Coefficientβ\\beta\.\(b\)Memory Bank CapacityNbN\_\{b\}\.
Figure 5:Parameter sensitivity analysis ofβ\\betaandNbN\_\{b\}on the CMU\-MOSI dataset\.\(a\)Label Penalty Weightα\\alpha\.\(b\)Loss Balancing Weightλ\\lambda\.\(c\)Top\-kkRatioρ\\rho\.
Figure 6:Parameter sensitivity analysis ofα\\alpha,λ\\lambda, andρ\\rhoon the CMU\-MOSI dataset\.
### E\.5Parameter Sensitivity Analysis

We analyze the sensitivity of RCE to five important hyperparameters, including the AVIB trade\-off coefficientβ\\beta, the memory bank capacityNbN\_\{b\}, the label penalty weightα\\alpha, the loss balancing weightλ\\lambda, and the top\-kkratioρ\\rho\. The results on CMU\-MOSI and CMU\-MOSEI are shown in Fig\.[5](https://arxiv.org/html/2609.30470#A5.F5), Fig\.[6](https://arxiv.org/html/2609.30470#A5.F6), Fig\.[7](https://arxiv.org/html/2609.30470#A5.F7), and Fig\.[8](https://arxiv.org/html/2609.30470#A5.F8)\.

\(1\) AVIB trade\-off coefficientβ\\beta\.The coefficientβ\\betacontrols the strength of the information bottleneck regularization in AVIB\. On CMU\-MOSI, RCE obtains competitive performance under a wide range ofβ\\betavalues, with the bestA​c​c​7Acc7achieved aroundβ=1​e\\beta=1e\-11and the bestM​A​EMAEobtained atβ=1​e\\beta=1e\-33\. On CMU\-MOSEI, the model is also relatively stable, andβ=1​e\\beta=1e\-33achieves both strongA​c​c​7Acc7and the lowestM​A​EMAE\. These results indicate that a moderate bottleneck constraint helps remove noisy or redundant information, while overly large or overly small values may weaken either task\-relevant information preservation or regularization effectiveness\. Therefore, we adoptβ=1​e\\beta=1e\-33as the default setting\.

\(a\)AVIB Trade\-off Coefficientβ\\beta\.\(b\)Memory Bank CapacityNbN\_\{b\}\.
Figure 7:Parameter sensitivity analysis ofβ\\betaandNbN\_\{b\}on the CMU\-MOSEI dataset\.\(a\)Label Penalty Weightα\\alpha\.\(b\)Loss Balancing Weightλ\\lambda\.\(c\)Top\-kkRatioρ\\rho\.
Figure 8:Parameter sensitivity analysis ofα\\alpha,λ\\lambda, andρ\\rhoon the CMU\-MOSEI dataset\.\(2\) Memory bank capacityNbN\_\{b\}\.The memory bank capacity determines the number of historical samples available for reliability\-aware modality enhancement\. On both datasets, very small memory banks provide limited neighborhood information and lead to inferior performance\. AsNbN\_\{b\}increases, the performance generally improves, showing that a larger candidate pool can provide more stable and informative enhancement samples\. On CMU\-MOSI,Nb=512N\_\{b\}=512achieves the bestA​c​c​7Acc7, whileNb=1024N\_\{b\}=1024obtains the lowestM​A​EMAE\. On CMU\-MOSEI,Nb=512N\_\{b\}=512also achieves the bestA​c​c​7Acc7, andNb=1024N\_\{b\}=1024or20482048gives the bestM​A​EMAE\. This suggests that a moderate\-to\-large memory bank is beneficial, but continuously increasing the capacity does not always bring further gains\.

\(3\) Label penalty weightα\\alpha\.The label penalty weightα\\alphacontrols how strongly label discrepancy is penalized during neighbor weighting\. On CMU\-MOSI, the model performs best aroundα=0\.4\\alpha=0\.4, where bothA​c​c​7Acc7andM​A​EMAEachieve the best results\. Whenα\\alphabecomes too large, performance drops clearly, suggesting that overly strong label constraints may suppress useful complementary samples\. On CMU\-MOSEI, the performance is relatively stable for mostα\\alphavalues, althoughα=0\.1\\alpha=0\.1gives the best result\. This may be because CMU\-MOSEI contains more samples and more diverse sentiment expressions, making a mild label penalty sufficient for filtering conflicting neighbors\.

\(4\) Loss balancing weightλ\\lambda\.The weightλ\\lambdabalances the main prediction loss and the AVIB\-related auxiliary objective\. On CMU\-MOSI, RCE achieves the best performance whenλ=0\.7\\lambda=0\.7, indicating that a proper contribution from AVIB is important for robust sentiment prediction\. On CMU\-MOSEI, the model remains stable across differentλ\\lambdavalues, with the bestA​c​c​7Acc7aroundλ=0\.4\\lambda=0\.4and the lowestM​A​EMAEaroundλ=1\.0\\lambda=1\.0\. Overall, moderate values ofλ\\lambdaprovide a good balance between prediction accuracy and representation regularization\.

\(5\) Top\-kkratioρ\\rho\.The top\-kkratioρ\\rhocontrols the proportion of retrieved neighbors used for modality enhancement\. Ifρ\\rhois too small, the model may not obtain sufficient complementary information\. Ifρ\\rhois too large, less relevant or sentiment\-conflicting samples may be introduced\. On CMU\-MOSI, the bestA​c​c​7Acc7is obtained aroundρ=0\.5\\rho=0\.5, while the lowestM​A​EMAEis achieved atρ=0\.6\\rho=0\.6\. On CMU\-MOSEI,ρ=0\.4\\rho=0\.4achieves the best overall performance\. These results show that selecting a moderate number of reliable neighbors is more effective than using either too few or too many candidates\.

Overall, RCE is robust to these hyperparameters on both datasets\. Although different datasets prefer slightly different values, the performance remains stable within moderate ranges\. This indicates that the proposed reliability\-aware enhancement and multilevel fusion mechanisms are not overly sensitive to hyperparameter choices\.

### E\.6Testing Stability Analysis

Table[11](https://arxiv.org/html/2609.30470#A5.T11)reports the testing stability of RCE on the CMU\-MOSI and CMU\-MOSEI datasets\. For each dataset, we run RCE with five different random seeds and compute the mean performance together with the 95% confidence interval for each metric\. Specifically, the confidence interval is calculated asx¯±t0\.975,4​s5\\bar\{x\}\\pm t\_\{0\.975,4\}\\frac\{s\}\{\\sqrt\{5\}\}, wherex¯\\bar\{x\}andssdenote the mean and standard deviation over the five runs, respectively\. The narrow intervals on both CMU\-MOSI and CMU\-MOSEI show that RCE has small performance variation across different random initializations, indicating stable and reproducible results\.

Table 11:Testing stability with five random seeds under the complete\-modality setting \(95% confidence intervals\)\.
### E\.7Model Complexity Analysis

We compare the computational complexity of RCE with state\-of\-the\-art methods in terms of FLOPs, memory usage, parameter count, and training time\. As shown in Table[12](https://arxiv.org/html/2609.30470#A5.T12), RCE introduces moderate additional cost due to the adaptive variational information bottleneck, memory bank\-based reliability\-aware modality enhancement, and multilevel fusion modules\. Specifically, RCE requires 12\.660G FLOPs and 9\.49GB memory, which are comparable to most baselines and only slightly higher than HME\[[70](https://arxiv.org/html/2609.30470#bib.bib43)\]\. Although RCE has the largest parameter count, its training time is 8\.54s per epoch, lower than both HME\[[70](https://arxiv.org/html/2609.30470#bib.bib43)\]and KAN\-MCP\[[30](https://arxiv.org/html/2609.30470#bib.bib51)\]\. This indicates that the proposed enhancement and fusion mechanisms do not incur prohibitive overhead\. Overall, RCE improves robustness and prediction performance while maintaining acceptable computational complexity\.

Table 12:Comparison of model complexity on the CMU\-MOSI dataset with batch size 48\.
### E\.8Theoretical Complexity Analysis

We further analyze the computational complexity of the key components in RCE\. LetBBdenote the batch size,ddthe hidden dimension,MMthe number of modalities,NbN\_\{b\}the memory bank capacity, andρ\\rhothe top\-kkretrieval ratio\. The AVIB module mainly consists of vMF parameter estimation, latent sampling, and auxiliary prediction, with a complexity ofO⁡\(M​B​d2\)O\(MBd^\{2\}\)\. The RME module retrieves reliable neighbors from the modality\-specific memory bank\. Its similarity computation requiresO⁡\(M​B​Nb​d\)O\(MBN\_\{b\}d\), and weighted aggregation over selected neighbors requiresO⁡\(M​B​ρ​Nb​d\)O\(MB\\rho N\_\{b\}d\)\. The hyper\-modality generation and multilevel fusion modules are based on lightweight latent attention and fusion operations, with an approximate complexity ofO⁡\(M​B​d2\)O\(MBd^\{2\}\)since the number of latent prompts and fusion tokens is small and fixed\. Therefore, the overall additional complexity of RCE is mainly dominated by memory\-bank retrieval, i\.e\.,O⁡\(M​B​Nb​d\)O\(MBN\_\{b\}d\), while other components scale linearly with respect to the batch size and number of modalities\. SinceMMandNbN\_\{b\}are fixed hyperparameters in practice, RCE remains computationally tractable\. This is also consistent with Table[12](https://arxiv.org/html/2609.30470#A5.T12), where RCE achieves stronger robustness with only moderate additional cost compared to HME and other missing\-modality baselines\.

相似文章

RHEA:面向稳健多模态属性图聚类的可靠性协调重建与分配

arXiv cs.LG

本文提出RHEA,一种面向多模态属性图聚类的可靠性感知框架。该框架通过邻域一致性估计节点特定的模态可靠性,重建不可靠模态,并使用可靠性感知融合和最优传输聚类。在四个基准上的实验显示出一致的性能提升,尤其在属性含噪或缺失的情况下。