streaming-tts

Tag

Cards List
#streaming-tts

X2Streaming-TTS: Causal Token-Level Text-to-Speech from Streaming Text with Speech-State Inheritance

arXiv cs.CL · 2d ago Cached

X2Streaming-TTS presents a causal token-level text-to-speech framework for true streaming synthesis, using causal commitment and speech-state inheritance to handle uncertain text prefixes and maintain acoustic continuity in low-latency spoken dialogue systems.

0 favorites 0 likes
#streaming-tts

VoiceChat-TTS: A Low-Latency Continuous Speech Synthesis Model for Interactive Agents

arXiv cs.CL · 5d ago Cached

VoiceChat-TTS is a low-latency, continuous text-to-speech model designed for interactive agents, enabling real-time streaming and interruption handling without compromising speech quality.

0 favorites 0 likes
#streaming-tts

Gepard : 0.6B streaming TTS built for real-time dialogue - 20× realtime factor, ~50ms time-to-first-audio, vLLM-native, Apache 2.0

Reddit r/LocalLLaMA · 2026-07-07 Cached

Gepard is a new streaming TTS model capable of real-time dialogue with ~50ms time-to-first-audio, supporting voice cloning and high parallelism, released under Apache 2.0.

0 favorites 0 likes
#streaming-tts

I can't believe text normalization is so underdiscussed in streaming text-to-speech [D]

Reddit r/MachineLearning · 2026-04-22

Author highlights under-discussed text normalization issues in streaming TTS and shares a vendor benchmark evaluating 1000+ sentences across 31 categories for dates, URLs, acronyms, etc.

0 favorites 0 likes
← Back to home

Submit Feedback