@TeksEdge: Tencent open-sourced a small 1.5B AI model that can replace a whole stack of separate audio tools. Beats Qwen3-TTS! Thi…

X AI KOLs Timeline Models

Summary

Tencent has open-sourced a 1.5B parameter AI model called AuK that can replace multiple audio tools, handling tasks like TTS, voice cloning, and denoising via natural language instructions.

🤯 Tencent open-sourced a small 1.5B AI model that can replace a whole stack of separate audio tools. 🏆 Beats Qwen3-TTS! This is the purpose of the AuK family of models. Instead of having one model for TTS, another for voice cloning, another for denoising, another for speaker separation... AuK does all of this through natural-language instructions ... 🗣️ Zero-shot voice cloning / TTS ✂️ Replace, insert or remove spoken words 🎭 Change emotion 🎚️ Change pitch, speed or volume 🗣️ Remove an accent 🤫 Convert normal speech ↔ whisper 🎵 Rewrite lyrics while preserving melody + voice 🧹 Denoise + dereverberate speech 👥 Separate speakers 🎤 Extract vocals from music And it is: 🧠 1.5B parameters 🔓 MIT licensed 📦 Weights released 💻 Code released 🖥️ CLI + Gradio + ComfyUI support There’s also AuK-Flash, a distilled version requiring only 4 inference steps. Tencent reports 4.5× faster wall-clock inference than the full model under matched conditions. Now one open speech system can handle jobs that normally require an entire collection of separate audio models and tools locally on your own hardware. (PyTorch-CUDA)🔥 ⚠️ Caveat: The 1.5B AuK model is not the entire runtime footprint. The reference setup also downloads Qwen2.5-Omni-3B as the multimodal encoder and loads the VAE separately. This is PyTorch-style local inference, not llama.cpp/GGUF.
Original Article
View Cached Full Text

Cached at: 09/10/26, 02:32 PM

🤯 Tencent open-sourced a small 1.5B AI model that can replace a whole stack of separate audio tools.

🏆 Beats Qwen3-TTS!

This is the purpose of the AuK family of models.

Instead of having one model for TTS, another for voice cloning, another for denoising, another for speaker separation…

AuK does all of this through natural-language instructions … 🗣️ Zero-shot voice cloning / TTS ✂️ Replace, insert or remove spoken words 🎭 Change emotion 🎚️ Change pitch, speed or volume 🗣️ Remove an accent 🤫 Convert normal speech ↔ whisper 🎵 Rewrite lyrics while preserving melody + voice 🧹 Denoise + dereverberate speech 👥 Separate speakers 🎤 Extract vocals from music

And it is: 🧠 1.5B parameters 🔓 MIT licensed 📦 Weights released 💻 Code released 🖥️ CLI + Gradio + ComfyUI support

There’s also AuK-Flash, a distilled version requiring only 4 inference steps.

Tencent reports 4.5× faster wall-clock inference than the full model under matched conditions.

Now one open speech system can handle jobs that normally require an entire collection of separate audio models and tools locally on your own hardware. (PyTorch-CUDA)🔥

⚠️ Caveat: The 1.5B AuK model is not the entire runtime footprint. The reference setup also downloads Qwen2.5-Omni-3B as the multimodal encoder and loads the VAE separately.

This is PyTorch-style local inference, not llama.cpp/GGUF.

Similar Articles