OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction
Summary
Introduces OneEmo, a unified multimodal reasoning model for emotion perception, understanding, and interaction, along with the EmoWorld-130K dataset and Emo-Chord reinforcement learning strategy.
View Cached Full Text
Cached at: 08/10/26, 10:14 AM
Paper page - OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction
Source: https://huggingface.co/papers/2608.06013
Abstract
MultimodalLargeLanguageModels(MLLMs)havedemonstratedremarkablecapabilitiesinemotionalintelligence.However,prevailingresearchpredominantlyfocusesontask-specificspecialization,oftenneglectinginter-tasksynergyandleavinglatentreasoningpotentialunderexplored.Tobridgethisgap,weintroduceOneEmo,aunifiedaffectivegeneralistcapableofmasteringemotionperception,comprehension,andinteraction.Forthispurpose,wefirstconstructEmoWorld-130K,acomprehensivedatasetthatdistillsspecializedaffectiveknowledgeintoexplicitreasoningtrajectoriesviaahuman-in-the-loopworkflow.Supervisedfine-tuningonthiscorpusrevealssignificantmutualbenefitsderivedfrommulti-tasklearning.Second,tofullyunlockthelatentreasoningpotential,weproposeEmo-Chord,anovelreinforcementlearningstrategythatstabilizesoptimizationthroughunifiedmulti-taskrewardallocation.ExtensiveexperimentsdemonstratethatOneEmoachievesstate-of-the-artperformanceagainstsimilarlysizedbaselinesacrossmostbenchmarks.Notably,despitehavingsignificantlyfewerparametersthancommercialmodels,OneEmodelivershighlycompetitiveresults.Thispaperpavesthewayformorereliableandinterpretableaffectivecomputing.Thecodeisavailableathttps://github.com/waHAHJIAHAO/OneEmo.
View arXiv pageView PDFGitHub5Add to collection
Get this paper in your agent:
hf papers read 2608\.06013
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper2
#### Jiaha0Hu4ng/OneEmo Video-Text-to-Text• 5B• Updated2 days ago • 122 • 2
#### Jiaha0Hu4ng/OneEmo-Base Video-Text-to-Text• Updated3 days ago
Datasets citing this paper1
#### Jiaha0Hu4ng/EmoWorld-130K
Spaces citing this paper1
Collections including this paper1
Similar Articles
EmoS: A High-Fidelity Multimodal Benchmark for Fine-grained Streaming Emotional Understanding
This article introduces EmoS, a high-fidelity multimodal benchmark designed for fine-grained streaming emotional understanding, addressing limitations in ecological validity and labeling reliability found in existing datasets.
MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy
This paper introduces MER-R1, a reinforcement learning framework that synergizes fast and slow thinking for multimodal emotion recognition. It achieves state-of-the-art performance by jointly optimizing recall and precision through dual-objective disentanglement and slow-fast confidence calibration.
Rationale-Guided Learning for Multimodal Emotion Recognition
Introduces Rationale-Guided Learning (RGL), a framework that reframes multimodal emotion recognition in conversation as a cognitively-inspired reasoning task using dual-process theory and MLLM-generated rationales, achieving state-of-the-art results on IEMOCAP and MELD.
tencent/Hy-Embodied-RxBrain-1.0 · Hugging Face
Tencent releases Hy-Embodied-RxBrain-1.0, a unified multimodal foundation model for embodied cognition that combines language reasoning with visual imagination for understanding, world state prediction, and subgoal planning.
C$^2$MOE: Consistency and Complementarity-guided Mixture of Experts for Incomplete Multimodal Emotion Learning
The paper proposes C²MOE, a Consistency and Complementarity-guided Mixture of Experts framework for incomplete multimodal emotion recognition in conversations, using information-theoretic decomposition to improve robustness when modalities are missing.