Tag
SpeechSense is a novel dataset for fine-grained speech sentiment analysis, focusing on paralinguistic cues to address limitations in text-centric approaches and validate the importance of acoustic features.
This article introduces EmoS, a high-fidelity multimodal benchmark designed for fine-grained streaming emotional understanding, addressing limitations in ecological validity and labeling reliability found in existing datasets.