Tag
X2Streaming-TTS presents a causal token-level text-to-speech framework for true streaming synthesis, using causal commitment and speech-state inheritance to handle uncertain text prefixes and maintain acoustic continuity in low-latency spoken dialogue systems.
VoiceChat-TTS is a low-latency, continuous text-to-speech model designed for interactive agents, enabling real-time streaming and interruption handling without compromising speech quality.
Gepard is a new streaming TTS model capable of real-time dialogue with ~50ms time-to-first-audio, supporting voice cloning and high parallelism, released under Apache 2.0.
Author highlights under-discussed text normalization issues in streaming TTS and shares a vendor benchmark evaluating 1000+ sentences across 31 categories for dates, URLs, acronyms, etc.