Tag
DialectS2S is an end-to-end speech dialogue model for low-resource Chinese dialects, introducing a scalable data synthesis pipeline and a two-stage post-training strategy with self-aligned speech supervision. Experiments show improvements in dialect consistency, response quality, and intelligibility, with fully open-sourced models, data, and code.
Bland launches Speech v3, claiming it's the world's first Human Speech Engine and top model in Design Arena's Audio Realism benchmark, surpassing ElevenLabs, Grok, Cartesia, and OpenAI.
Microsoft is testing a new native real-time voice model, MAI Realtime, in early access on its MAI Playground. The full-duplex system supports multiple languages, low latency, and configurable turn-taking, positioning it as a competitor to OpenAI's GPT Live and Sesame.
A significant breakthrough in Levantine Arabic speech synthesis and English code-switching, achieving a 76% improvement using a single RTX 3060 in an evening.
This paper presents a case study using unsupervised articulatory probing to examine how self-supervised speech models encode phonetic features across Mandarin sub-dialects, finding that salient features like labiality remain stable while finer spectral distinctions show dialect-dependent variation.