@Ali_TongyiLab: Qwen-Audio-3.0-TTS is here. Our latest text-to-speech model, in two flavors: • Flash: real-time interaction • Plus: hig…
Summary
Alibaba's Tongyi Lab released Qwen-Audio-3.0-TTS, a new text-to-speech model with Flash (real-time) and Plus (high-quality) versions, supporting 16 languages, natural language style control, and robust voice cloning.
View Cached Full Text
Cached at: 07/20/26, 07:31 PM
Qwen-Audio-3.0-TTS is here.
Our latest text-to-speech model, in two flavors: • Flash: real-time interaction • Plus: high-quality generation
What’s new: • Multilingual coverage across 16 languages • Style control in natural language • Fine-grained tags for non-verbal details • More robust voice cloning from imperfect audio
Similar Articles
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
Alibaba's Qwen team releases Qwen3-TTS-12Hz-1.7B-CustomVoice, a powerful text-to-speech model supporting 10 languages with low-latency streaming, instruction-based voice control, and robust contextual understanding.
Qwen3-TTS Technical Report
The Qwen3-TTS technical report introduces a series of advanced multilingual text-to-speech models with voice cloning and controllable generation, featuring a dual-track LM architecture and specialized tokenizers for low-latency streaming.
Qwen3 TTS is seriously underrated - I got it running locally in real-time and it's one of the most expressive open TTS models I've tried
Developer shows how to run Qwen3 TTS locally in real-time with streaming, quantization, word-level alignment, and custom voice fine-tuning for an expressive open-source TTS pipeline.
Qwen3-TTS voice cloning is now in mainline llama.cpp — the old demo finally became real support
Qwen3-TTS voice cloning has been merged into mainline llama.cpp, enabling local text-to-speech with voice cloning from short reference audio via the llama-tts binary, supporting multiple languages. Limitations remain, including only the Base model and no server endpoint yet.
Qwen3.7-Plus: Multimodal Agent Intelligence (36 minute read)
Qwen3.7-Plus is a multimodal agent model that unifies vision and language for seamless GUI and CLI interactions, now available via Alibaba Cloud Model Studio.