Tag
YuE2-3B is an open-source music generation model that creates complete songs from lyrics and style prompts, featuring editable musical scores and state-of-the-art performance rivaling commercial models like Suno v5.
MiniMax releases Music 3, a high-performance music generation model that creates complete songs up to five minutes long using lyrics and detailed descriptions, with an 8B global LLM and 0.6B local LLM for long-range coherence and acoustic detail.
This paper presents a dependency-free native runtime that enables efficient text-to-music generation on embedded devices through quantization and activation steering, achieving no measurable quality loss at 8-bit precision and enabling the 1.2B parameter model to run on a Raspberry Pi 5 at 4-bit precision.
This paper presents a text-to-music generation system that leverages reward conditioning, expert iteration, and preference tuning to improve audio quality within a 120M-parameter model, submitted to the ATTM Grand Challenge at ICME 2026.
TuneJury is an open-source pairwise reward model for text-to-music generation that provides calibrated preference scoring and generalizes across multiple downstream applications.
This paper introduces a dual-layer caption poisoning attack on retrieval-augmented text-to-music systems, showing that an attacker can inject malicious captions into the knowledge database to steer generated music toward attacker-chosen intent without modifying user prompts or models.
Khala 1.0 is an open-source music generation model for high-fidelity full-song generation from text and lyrics, using a unified acoustic-token pipeline. It was released by the Central Conservatory of Music in Beijing with paper, code, weights, and demo.