Tag
Suno announces new watermarking and fingerprinting tools to mark AI-generated tracks, limit downloads, and update community guidelines to prevent copycat songs, amid ongoing lawsuits from major labels and artist bodies.
The Verge reports that Treblo's new open-source AI music classifier indicates Fenix Flexin's song 'Rubberz' was almost entirely AI-generated, potentially making it the first AI-generated song to reach the Billboard Hot 100. Flexin denies the claims, but the evidence adds to the ongoing controversy.
This preprint introduces hierarchical self-supervised world models for music co-creation agents, with fast CPU-friendly models and a live demo for MIDI inpainting and generation.
Wired reports that an AI music app, Treblo, has flagged songs by rappers Fenix Flexin and Tyga as likely AI-generated, reigniting the debate over AI use in popular music after the song 'Rubberz' drew accusations.
Mubert announces an update to its API featuring a new engine that enables editing tracks and stems while delivering consistent music generation.
Google launched Lyria 3.5 in Flow Music, a music generation model with improvements in musicality, lyrics quality, vocal expressiveness, and creative control over tempo and duration.
MusiChat presents a conversational system for human-AI music co-creation that enables iterative refinement through natural language interaction, achieving high accuracy in multi-turn editing.
WanSong is a pure diffusion-based music generation model that directly produces high-fidelity, multilingual songs up to 5 minutes long with dual stems (vocals and background music) in a single run, addressing challenges in efficient generation, long-form audio, and controllability.
Qwen-Music is a new paper that splits music generation into two stages: planning with compact semantic tokens and CoT reasoning over melody, then rendering high-fidelity audio, achieving state-of-the-art results.
Qwen-Music is a music generation model that produces high-fidelity songs with vocals, supporting text-to-music and cover song generation. It uses a novel Melody-Chain-of-Thought mechanism and achieves state-of-the-art results on 13 of 16 objective metrics.
A discussion on whether an AI-generated song could top the charts and how public opinion might react if it were revealed to be completely AI-made.
audio.cpp releases a major update adding music/SFX generation and source separation with ACE-Step, HeartMuLa, Stable Audio 3, and HTDemucs, achieving up to 10x real-time speed for long music generation in native C++/GGML.
A user tested six frontier LLMs on generating music from a Bach MusicXML file in a one-shot setting, presenting unedited results.
Suno launches the Spark incubator program for independent artists, offering grants, mentorship, and marketing support to feed into its AI music platform.
May saw over $1.8 billion in voice AI funding, led by Sierra's $925M and Hark's $700M rounds, while ElevenLabs launched new models for music generation and dubbing with enhanced control. The newsletter also highlights healthcare deals and India's growing voice market.
TuneJury is an open-source pairwise reward model for text-to-music generation that provides calibrated preference scoring and generalizes across multiple downstream applications.
PianoKontext generates variable-length expressive piano performances from deadpan MIDI scores by aligning audio and MIDI in latent space using Dynamic Time Warping and flow matching with DiT blocks.
This paper introduces the Eisbach log-barrier, a parameter-free weight derived from the entropy of DiT output's spatial energy distribution, which when applied to LoRA fine-tuning of Stable Audio 3 improves musical diversity and thematic development without causing mode collapse.
Google released Magenta RealTime 2 on Hugging Face, an open-weights model for real-time continuous music generation on device with ~200ms latency, steerable by text, audio, or MIDI.
This paper introduces a dual-layer caption poisoning attack on retrieval-augmented text-to-music systems, showing that an attacker can inject malicious captions into the knowledge database to steer generated music toward attacker-chosen intent without modifying user prompts or models.