@OrukLabs: None of these models was ever told what emotion is. They were trained to transcribe words, or to fill in masked audio. …
Summary
A study from OrukLabs shows that speech models trained solely on transcription or masked audio tasks spontaneously learn to represent emotions in their deeper layers, as revealed by mapping with real voice clips.
View Cached Full Text
Cached at: 07/05/26, 10:36 PM
None of these models was ever told what emotion is. They were trained to transcribe words, or to fill in masked audio. That’s it. Map their layers with 3,276 real voice clips and emotion sorts itself out anyway. The deeper you go, the clearer it gets. https://t.co/RyN92FlsKU
Similar Articles
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
This paper replicates the finding of 'emotion vectors' in open-weight LLMs Apertus-8B and Gemma-4-E4B, showing that valence geometry is recoverable across models with differences in layer emergence. The study also finds that arousal encoding is sensitive to the story corpus used for extraction.
Do Speech Emphasis Models Generalize across Languages and Emotions?
Introduces MMEE, a multilingual multi-emotion emphasis corpus of 10,000 utterances across 7 languages and 34 emotions, and benchmarks emphasis detection models under various transfer settings, finding that multilingual training improves robustness while monolingual models show limited zero-shot transfer.
The Echo Amplifies the Knowledge: Somatic Marker Analogues in Language Models via Emotion Vector Re-Injection
This preprint introduces a method to inject emotion vectors into language models to simulate somatic markers, aiming to bridge the gap between semantic and episodic memory. The authors demonstrate that combining emotional echoes with semantic knowledge improves decision-making capabilities, replicating findings from human cognitive science.
Real-Time Voice AI Hears but Does Not Listen (arXiv:2606.26083)
This paper evaluates four leading real-time voice AI systems (GPT Realtime 2, Gemini 3.1 Flash Live, Qwen3.5 Omni Plus, Omni Flash) and finds they consistently act on words rather than vocal tone, ignoring distress, fear, or sarcasm even when they can perceive them—termed the 'emotional intelligence gap' of voice AI.
@multimodalart: they extracted only the audio bit of LTX-2.3, fine-tuned for TTS task and achieved SOTA TTS emotional control??? try it…
A fine-tuned version of the LTX-2.3 model's audio component achieves state-of-the-art emotional control in text-to-speech, now available as a Hugging Face Space called DramaBox by ResembleAI.