Tag
The paper studies zero-shot cross-subject continuous valence-arousal regression from synchronized EEG-fNIRS data, decomposing affect into a stimulus-shared component and an individual component calibrated via label-free alpha-band cross-channel synchrony, outperforming EEGNet and ASAC-Net baselines while reporting a systematic negative-result search over alternative architectures and features.
This paper reveals that the optimal layer for linear probing to read concepts differs from the optimal layer for activation steering in omni-modal large language models, challenging common heuristics in representation engineering.
This paper presents the YNU-HPCC team's approach for SemEval-2025 Task 11 on text-based emotion recognition, using a RoBERTa model with enhanced output headers and achieving a ranking score of 0.44 through English-translated datasets.
The paper introduces the Modality Discrepancy Transformer (MDT), a novel multimodal fusion framework that enhances cross-modal discrepancy modeling for recognizing ambivalence and hesitancy in clinical videos, outperforming baselines on the BAH dataset.
This paper presents a confidence-gated hybrid system for emotion recognition in conversational AI that routes most traffic through a low-cost ensemble and escalates uncertain cases to an LLM, achieving high accuracy while reducing costs and latency for CCaaS platforms.
This paper introduces MultiHuSE, a multimodal dataset comprising videos of actors with annotations for humour styles and emotions. Baseline experiments show that multimodal fusion improves humour style classification accuracy over unimodal approaches.
This paper investigates weaknesses in emotion recognition in conversations (ERC) by analyzing LLM performance in zero-shot settings, revealing systematic failures due to annotation ambiguity, and proposes an LLM-as-Judge framework for more robust evaluation.
The paper presents Chiaro, a new benchmark dataset for contrastive emotion recognition where two individuals experience opposing emotions from a shared event, grounded in appraisal theory. It evaluates seven LLMs and four emotion classifiers, revealing that current models fall short of human performance.
The paper introduces VoiceLongMemEval (VLME), a benchmark that evaluates AI assistants' ability to remember and reason over paralinguistic metadata like emotion and prosody from voice in long-term conversations, revealing an 'affect gap' in current models.
VocalAffectBench is introduced as a public benchmark for evaluating vocal emotion recognition in AI audio models, demonstrating that current baselines have limited accuracy, particularly for non-neutral emotions.
AffectOmni is a GRPO-trained framework for verifiable affective reasoning in multimodal large language models, introducing People Focus and Temporal Order rewards to enhance people-centric evidence selection and temporally structured reasoning, with experiments showing improvements over 7B scale baselines.
The research finds that emotion representation is mostly universal across seven languages in speech models, with a measurable 'accent' that mirrors human cross-cultural studies and affects cross-lingual transfer.
This research paper explores emotion-sensitive neurons in multimodal foundation models, revealing shared affective mechanisms between speech and facial emotion recognition through causal interventions and cross-modal analysis.
Introduces Rationale-Guided Learning (RGL), a framework that reframes multimodal emotion recognition in conversation as a cognitively-inspired reasoning task using dual-process theory and MLLM-generated rationales, achieving state-of-the-art results on IEMOCAP and MELD.
This paper proposes CONFER, a graph-based conflict-aware evidence negotiation framework for weakly supervised multimodal emotion recognition, addressing self-report unreliability and cross-modal conflict. It achieves competitive accuracy on AMIGOS, MAHNOB-HCI, and DEAP benchmarks.
This preprint introduces a generation-aligned diagnostic ladder that separates decision-rule misalignment from readout-coverage limitations in speech language models, showing that state decoding far exceeds generated accuracy in emotion recognition tasks.
The paper proposes C²MOE, a Consistency and Complementarity-guided Mixture of Experts framework for incomplete multimodal emotion recognition in conversations, using information-theoretic decomposition to improve robustness when modalities are missing.
This paper examines how evaluation protocols affect reported accuracy in EEG emotion recognition, using a DGCNN on SEED and SEED-IV datasets. It demonstrates that subject-dependent, subject-disjoint, and cross-session evaluations answer different questions, and that checkpoint selection and test-set reuse can inflate accuracy.
This paper introduces AtmosERC, a model that models dialogue-level affective atmosphere to enhance emotion recognition in conversations.
Proposes DCC, a dynamic commonsense coordination framework for empathetic response generation that integrates residual-based interaction, association-guided filtering, and iterative decoding, achieving improved emotion classification and response diversity over baselines.