@Ali_TongyiLab: Qwen-Audio-3.0-TTS is here. Our latest text-to-speech model, in two flavors: • Flash: real-time interaction • Plus: hig…

X AI KOLs Timeline Models

Summary

Alibaba's Tongyi Lab released Qwen-Audio-3.0-TTS, a new text-to-speech model with Flash (real-time) and Plus (high-quality) versions, supporting 16 languages, natural language style control, and robust voice cloning.

Qwen-Audio-3.0-TTS is here. Our latest text-to-speech model, in two flavors: • Flash: real-time interaction • Plus: high-quality generation What's new: • Multilingual coverage across 16 languages • Style control in natural language • Fine-grained tags for non-verbal details • More robust voice cloning from imperfect audio
Original Article
View Cached Full Text

Cached at: 07/20/26, 07:31 PM

Qwen-Audio-3.0-TTS is here.

Our latest text-to-speech model, in two flavors: • Flash: real-time interaction • Plus: high-quality generation

What’s new: • Multilingual coverage across 16 languages • Style control in natural language • Fine-grained tags for non-verbal details • More robust voice cloning from imperfect audio

Similar Articles

Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice

Hugging Face Models Trending

Alibaba's Qwen team releases Qwen3-TTS-12Hz-1.7B-CustomVoice, a powerful text-to-speech model supporting 10 languages with low-latency streaming, instruction-based voice control, and robust contextual understanding.

Qwen3-TTS Technical Report

Papers with Code Trending

The Qwen3-TTS technical report introduces a series of advanced multilingual text-to-speech models with voice cloning and controllable generation, featuring a dual-track LM architecture and specialized tokenizers for low-latency streaming.