speech-models

Tag

Cards List
#speech-models

@OrukLabs: None of these models was ever told what emotion is. They were trained to transcribe words, or to fill in masked audio. …

X AI KOLs Following · 2026-07-04 Cached

A study from OrukLabs shows that speech models trained solely on transcription or masked audio tasks spontaneously learn to represent emotions in their deeper layers, as revealed by mapping with real voice clips.

0 favorites 0 likes
#speech-models

RedVox: Safety and Fairness Gaps in Speech Models Across Languages

arXiv cs.CL · 2026-06-26 Cached

This paper introduces RedVox, a multilingual safety and fairness benchmark for speech models. Evaluating eight state-of-the-art models across five languages, it finds persistent vulnerabilities that worsen in non-English settings and with spoken input.

0 favorites 0 likes
#speech-models

SpeechEQ: Benchmarking Emotional Intelligence Quotient in Socially Aware Voice Conversational Models

arXiv cs.CL · 2026-06-25 Cached

SpeechEQ introduces a benchmark and dataset for evaluating emotional intelligence in speech-language models, covering 15 EQ subscales across 2,265 dialogues. Experiments reveal current models struggle with paralinguistic cues, exhibiting text-reliant shortcuts and other limitations.

0 favorites 0 likes
#speech-models

Perceptual compensation for tonal context in self-supervised speech models

arXiv cs.CL · 2026-06-17 Cached

This paper investigates whether the wav2vec2.0 architecture exhibits perceptual compensation for tonal context in Mandarin Chinese, finding limited evidence in the self-supervised model compared to human listeners and suggesting that supervised fine-tuning may be necessary for such phonological abstraction.

0 favorites 0 likes
#speech-models

@kyutai_labs: New paper: Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models We use RL to post-train speech models (Mo…

X AI KOLs Following · 2026-06-10 Cached

Kyutai Labs released a new paper on using reinforcement learning to post-train speech models (Moshi and PersonaPlex) for more human-like interaction, including when to respond, wait, or give listening cues.

0 favorites 0 likes
← Back to home

Submit Feedback