Tag
Proposes MSR-IVA, a state-aware framework for fusing structural MRI and dynamic functional network connectivity, improving matched source coupling by 6.5% and reducing unmatched dependence by 15.7% in an Alzheimer's disease cohort.
The paper proposes a framework for interpretable multimodal classification using Linear Discriminant Tree Ensembles, which balance accuracy and interpretability, outperforming Transformer models in F1-mod gains and human-annotator agreement scores.
This paper proposes an iterative proxy correction framework to enhance robustness in multimodal sentiment analysis when dealing with incomplete or corrupted inputs by refining a language proxy for better sentiment prediction.
This paper introduces TIER-MoE, a risk-guided subspace mixture-of-experts model for multimodal biomedical classification that estimates sample-specific modality reliability from out-of-fold predictions and routes modalities to experts, improving performance and calibration on four public datasets.
This paper presents a method that uses frozen medical large language model (LLM) representations as a shared embedding space to predict primary ICD diagnosis categories from both structured and unstructured electronic health record data, achieving improved accuracy over baseline methods on MIMIC-IV and showing transferability to MIMIC-III.
This paper proposes a leakage-safe diagnostic to test whether quality-aware multimodal fusion methods actually use reliability scores during inference, by permuting these scores across test examples. Experiments on StressID and CMU-MOSEI show that shuffled reliability scores leave performance unchanged, indicating that quality signals only influence decisions when they reliably predict unimodal correctness.
This paper evaluates deep learning models (LSTM, TCN, Transformer) on the WESAD dataset for multimodal emotion recognition from physiological signals, showing that an ensemble achieves 98.91% accuracy.
FusionSense introduces a tri-stage near-sensor learning framework for multimodal edge intelligence that jointly reduces compute and communication by using fusion-aware filtering, achieving up to 33× energy savings and significant data-reduction gains on RGB-Depth/LiDAR tasks.
MuteBench is a benchmark for evaluating multimodal fusion models under modality missing and within-modality missing conditions across clinical datasets. It provides insights into architecture robustness and suggests that diffusion-based imputation can help.