Tag
This paper investigates weaknesses in emotion recognition in conversations (ERC) by analyzing LLM performance in zero-shot settings, revealing systematic failures due to annotation ambiguity, and proposes an LLM-as-Judge framework for more robust evaluation.
The paper presents Chiaro, a new benchmark dataset for contrastive emotion recognition where two individuals experience opposing emotions from a shared event, grounded in appraisal theory. It evaluates seven LLMs and four emotion classifiers, revealing that current models fall short of human performance.
The paper introduces VoiceLongMemEval (VLME), a benchmark that evaluates AI assistants' ability to remember and reason over paralinguistic metadata like emotion and prosody from voice in long-term conversations, revealing an 'affect gap' in current models.
VocalAffectBench is introduced as a public benchmark for evaluating vocal emotion recognition in AI audio models, demonstrating that current baselines have limited accuracy, particularly for non-neutral emotions.
AffectOmni is a GRPO-trained framework for verifiable affective reasoning in multimodal large language models, introducing People Focus and Temporal Order rewards to enhance people-centric evidence selection and temporally structured reasoning, with experiments showing improvements over 7B scale baselines.
The research finds that emotion representation is mostly universal across seven languages in speech models, with a measurable 'accent' that mirrors human cross-cultural studies and affects cross-lingual transfer.
This research paper explores emotion-sensitive neurons in multimodal foundation models, revealing shared affective mechanisms between speech and facial emotion recognition through causal interventions and cross-modal analysis.
Introduces Rationale-Guided Learning (RGL), a framework that reframes multimodal emotion recognition in conversation as a cognitively-inspired reasoning task using dual-process theory and MLLM-generated rationales, achieving state-of-the-art results on IEMOCAP and MELD.
This paper proposes CONFER, a graph-based conflict-aware evidence negotiation framework for weakly supervised multimodal emotion recognition, addressing self-report unreliability and cross-modal conflict. It achieves competitive accuracy on AMIGOS, MAHNOB-HCI, and DEAP benchmarks.
This preprint introduces a generation-aligned diagnostic ladder that separates decision-rule misalignment from readout-coverage limitations in speech language models, showing that state decoding far exceeds generated accuracy in emotion recognition tasks.
The paper proposes C²MOE, a Consistency and Complementarity-guided Mixture of Experts framework for incomplete multimodal emotion recognition in conversations, using information-theoretic decomposition to improve robustness when modalities are missing.
This paper examines how evaluation protocols affect reported accuracy in EEG emotion recognition, using a DGCNN on SEED and SEED-IV datasets. It demonstrates that subject-dependent, subject-disjoint, and cross-session evaluations answer different questions, and that checkpoint selection and test-set reuse can inflate accuracy.
This paper introduces AtmosERC, a model that models dialogue-level affective atmosphere to enhance emotion recognition in conversations.
Proposes DCC, a dynamic commonsense coordination framework for empathetic response generation that integrates residual-based interaction, association-guided filtering, and iterative decoding, achieving improved emotion classification and response diversity over baselines.
Introduces SCoPE, a lightweight module for emotion recognition in conversations that models speaker-specific emotional priors and uses emotion shift prediction to dynamically fuse prior and multimodal evidence, achieving state-of-the-art on IEMOCAP.
The paper proposes Light-MER, a lightweight multimodal emotion recognition framework that uses knowledge distillation from an 8B teacher model to a sub-1B student, achieving state-of-the-art performance with significantly higher inference efficiency, challenging the necessity of models larger than 1B parameters.
The paper introduces a graph-regularized deep learning framework for EEG-based emotion recognition that incorporates psychologically-grounded emotion topology into the training objective, achieving up to +5.42% accuracy and 39% reduction in psychologically implausible misclassifications on SEED datasets.
This paper proposes SHAP-weighted cross-modal expert fusion (XGAF) for emotion and sentiment recognition, demonstrating that sum-abs SHAP aggregation achieves early-fusion-level performance on MELD and CMU-MOSEI datasets.
A study from OrukLabs shows that speech models trained solely on transcription or masked audio tasks spontaneously learn to represent emotions in their deeper layers, as revealed by mapping with real voice clips.
PRISM is a novel framework for cross-subject EEG emotion recognition that combines prioritized channel importance weighting via a lightweight expert ensemble with semi-supervised domain adaptation using confidence-filtered pseudo-labels, achieving state-of-the-art results on DEAP, DREAMER, and SEED datasets.