AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing
Summary
AuK is an open-source foundational model that unifies speech generation and editing through natural-language instructions, achieving leading performance with efficient inference via distillation.
Similar Articles
@TencentHunyuan: AuK is officially here. Nano banana for audio An open-source foundation model for unified speech generation and editing…
AuK is an open-source foundational model for unified speech generation and editing, supporting tasks like zero-shot TTS and content editing via natural-language instructions.
X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation
X-AuT is a progressive framework for compressing audio-encoder layers in speech large language models, reducing inference cost while restoring accuracy via techniques like cross-scale distillation and LoRA adaptation.
@TeksEdge: Tencent open-sourced a small 1.5B AI model that can replace a whole stack of separate audio tools. Beats Qwen3-TTS! Thi…
Tencent has open-sourced a 1.5B parameter AI model called AuK that can replace multiple audio tools, handling tasks like TTS, voice cloning, and denoising via natural language instructions.
Open source : Turning vocal imitations into sound effects. (New UX for sound generation)
An open-source AI model that generates sound effects from vocal imitations and text descriptions, addressing the challenge of searching for specific sounds.
@paulabartabajo_: Advice for AI engineers If you're building voice agents, stop wiring up 3 separate models, for audio-to-text, text-to-a…
Announces liquid-audio, an open-source repository for Liquid AI's end-to-end speech-to-speech LFM models (LFM2-Audio-1.5B and LFM2.5-Audio-1.5B) with interleaved and sequential generation modes and fine-tuning support.