@OrukLabs: None of these models was ever told what emotion is. They were trained to transcribe words, or to fill in masked audio. …

X AI KOLs Following News

Summary

A study from OrukLabs shows that speech models trained solely on transcription or masked audio tasks spontaneously learn to represent emotions in their deeper layers, as revealed by mapping with real voice clips.

None of these models was ever told what emotion is. They were trained to transcribe words, or to fill in masked audio. That's it. Map their layers with 3,276 real voice clips and emotion sorts itself out anyway. The deeper you go, the clearer it gets. https://t.co/RyN92FlsKU
Original Article
View Cached Full Text

Cached at: 07/05/26, 10:36 PM

None of these models was ever told what emotion is. They were trained to transcribe words, or to fill in masked audio. That’s it. Map their layers with 3,276 real voice clips and emotion sorts itself out anyway. The deeper you go, the clearer it gets. https://t.co/RyN92FlsKU

Similar Articles

Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs

arXiv cs.CL

This paper replicates the finding of 'emotion vectors' in open-weight LLMs Apertus-8B and Gemma-4-E4B, showing that valence geometry is recoverable across models with differences in layer emergence. The study also finds that arousal encoding is sensitive to the story corpus used for extraction.

Do Speech Emphasis Models Generalize across Languages and Emotions?

arXiv cs.CL

Introduces MMEE, a multilingual multi-emotion emphasis corpus of 10,000 utterances across 7 languages and 34 emotions, and benchmarks emphasis detection models under various transfer settings, finding that multilingual training improves robustness while monolingual models show limited zero-shot transfer.

Real-Time Voice AI Hears but Does Not Listen (arXiv:2606.26083)

Reddit r/artificial

This paper evaluates four leading real-time voice AI systems (GPT Realtime 2, Gemini 3.1 Flash Live, Qwen3.5 Omni Plus, Omni Flash) and finds they consistently act on words rather than vocal tone, ignoring distress, fear, or sarcasm even when they can perceive them—termed the 'emotional intelligence gap' of voice AI.