speech-model

Tag

Cards List
#speech-model

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects

arXiv cs.CL · 2026-08-11 Cached

DialectS2S is an end-to-end speech dialogue model for low-resource Chinese dialects, introducing a scalable data synthesis pipeline and a two-stage post-training strategy with self-aligned speech supervision. Experiments show improvements in dialect consistency, response quality, and intelligibility, with fully open-sourced models, data, and code.

0 favorites 0 likes
#speech-model

@LinusEkenstam: This is life-changing tech. We need less hype, and more stuff like this. Imagine how many people this can help

X AI KOLs Following · 2026-08-04 Cached

Bland launches Speech v3, claiming it's the world's first Human Speech Engine and top model in Design Arena's Audio Realism benchmark, surpassing ElevenLabs, Grok, Cartesia, and OpenAI.

0 favorites 0 likes
#speech-model

Microsoft tests new MAI Realtime voice model (2 minute read)

TLDR AI · 2026-08-03 Cached

Microsoft is testing a new native real-time voice model, MAI Realtime, in early access on its MAI Playground. The full-duplex system supports multiple languages, low latency, and configurable turn-taking, positioning it as a competitor to OpenAI's GPT Live and Sesame.

0 favorites 0 likes
#speech-model

@alamin_ai_: OMG, guys, this is unbelievable Please listen to the Levantine Arabic and the seamless code-switching with english, a 7…

X AI KOLs Following · 2026-07-06 Cached

A significant breakthrough in Levantine Arabic speech synthesis and English code-switching, achieving a 76% improvement using a single RTX 3060 in an evening.

0 favorites 0 likes
#speech-model

Probing in the Wild: A Case Study of Self-Supervised Speech Representations on Mandarin Sub-dialects with Unsupervised Articulatory Analysis

arXiv cs.CL · 2026-06-25 Cached

This paper presents a case study using unsupervised articulatory probing to examine how self-supervised speech models encode phonetic features across Mandarin sub-dialects, finding that salient features like labiality remain stable while finer spectral distinctions show dialect-dependent variation.

0 favorites 0 likes
← Back to home

Submit Feedback