Real-time voice AI can hear the emotion — but does it actually use it?
Summary
The study tests real-time voice AI models on emotional cues in speech, finding that while they can detect emotions like distress or sarcasm, this information doesn't reliably influence decisions, highlighting an emotional intelligence gap in current systems.
Similar Articles
Real-Time Voice AI Hears but Does Not Listen (arXiv:2606.26083)
This paper evaluates four leading real-time voice AI systems (GPT Realtime 2, Gemini 3.1 Flash Live, Qwen3.5 Omni Plus, Omni Flash) and finds they consistently act on words rather than vocal tone, ignoring distress, fear, or sarcasm even when they can perceive them—termed the 'emotional intelligence gap' of voice AI.
@OrukLabs: None of these models was ever told what emotion is. They were trained to transcribe words, or to fill in masked audio. …
A study from OrukLabs shows that speech models trained solely on transcription or masked audio tasks spontaneously learn to represent emotions in their deeper layers, as revealed by mapping with real voice clips.
@BenjaminDEKR: Even if you doubt that AI can "feel" anything, this is valuable as a reflection of humans: "> They wrote about being wo…
This tweet highlights a study where researchers identified a 'pain' signal in AI systems and tested AI's ability to detect fake relief mechanisms, offering insights into AI cognition and parallels with human emotions.
VocalAffectBench: Evaluating Vocal Emotion Recognition in AI Audio Models
VocalAffectBench is introduced as a public benchmark for evaluating vocal emotion recognition in AI audio models, demonstrating that current baselines have limited accuracy, particularly for non-neutral emotions.
SpeechEQ: Benchmarking Emotional Intelligence Quotient in Socially Aware Voice Conversational Models
SpeechEQ introduces a benchmark and dataset for evaluating emotional intelligence in speech-language models, covering 15 EQ subscales across 2,265 dialogues. Experiments reveal current models struggle with paralinguistic cues, exhibiting text-reliant shortcuts and other limitations.