Tag
MiniMax releases Music 3, a high-performance music generation model that creates complete songs up to five minutes long using lyrics and detailed descriptions, with an 8B global LLM and 0.6B local LLM for long-range coherence and acoustic detail.
WanSong is a pure diffusion-based music generation model that directly produces high-fidelity, multilingual songs up to 5 minutes long with dual stems (vocals and background music) in a single run, addressing challenges in efficient generation, long-form audio, and controllability.
VibeVoice is a new model from Microsoft that synthesizes long-form multi-speaker speech using next-token diffusion and a highly efficient continuous speech tokenizer. It achieves superior fidelity and compression, supporting up to 90 minutes of audio with multiple speakers.