YuE2 · Frontier Music with Symbolic Planning

Hacker News Top Models

Summary

YuE2 is a new AI music generation model using symbolic planning that achieves competitive performance on benchmarks like WildSongBench, outperforming or matching proprietary systems such as Suno v5.

No content available
Original Article
View Cached Full Text

Cached at: 09/11/26, 02:17 AM

# YuE2 · Frontier Music with Symbolic Planning Source: [https://map-yue2.github.io/](https://map-yue2.github.io/) ## From score to song Listen to a song, then explore the melody, rhythm, and chords in its symbolic plan\. Selected score ## Loading selected song… The symbolic score Original score recording **Interactive ABC score**Red notes follow the score recording\. View original score pages All selected scores ## Cover & Editing A familiar song can take a different shape\. Listen to changes in melody, lyrics, tempo, and arrangement\. ### Agentic music editing A song takes shape through a conversation\. Loading the editing story… ## Genre Explorer The listening selection, gathered across genres and languages\. SearchLanguageGenreGeneration ## Model & Results YuE2 \(best\-of\-8\) reaches**6\.9632**on SongBench, the highest observed mean among 15 evaluated settings on WildSongBench \(192 prompts\)\. Suno v5 scores 6\.8721 in the same comparison\. [![Paper Figure 1: a WildSongBench comparison of song quality and text alignment. YuE2 and YuE2 best-of-8 are competitive with the evaluated proprietary systems. Bubble area represents AudioBox production quality.](https://map-yue2.github.io/static/figures/frontier-teaser.svg)](https://map-yue2.github.io/static/figures/frontier-teaser.pdf)Song quality and text alignment on WildSongBench\. Bubble area shows AudioBox PQ; black outlines mark Pareto optima on the two plotted axes\. Bo8 = best\-of\-8\.[How to read the indices](https://map-yue2.github.io/#evaluation-protocol)\.Model architecture ### Composing in symbols, performing in audio\. [Vector PDF](https://map-yue2.github.io/static/figures/yue2-overview.pdf) [![YuE2 architecture: an AR–NAR Mixture-of-Transformers turns an editable symbolic score into semantic tokens, acoustic latents, and full-song audio. The same generator supports creation, covering, and editing.](https://map-yue2.github.io/static/figures/yue2-overview.svg)](https://map-yue2.github.io/static/figures/yue2-overview.pdf)An editable score becomes semantic music tokens, acoustic latents, and full\-song audio\. The model has approximately 3\.59B parameters and 28 layers, and supports creation, covering, and editing\. The AR and NAR experts share an attention computation while using separate normalization, projections, and MLPs\.Explore benchmark scoresWildSongBench · 15 settings · 7 metrics**WildSongBench**192 prompts Compare by How to read these results**WildSongBench\.**192 prompts and 15 system settings\. The table reports automatic evaluation scores\. Best\-of\-8 selects one of eight generations by musicality, prompt control, and lyric accuracy\. **Figure 1\.**Song quality combines SongBench and SongEval; text alignment combines MuLan, AllMusicCaps, and prompt control\. Both axes show normalized comparison indices\. Bubble area represents AudioBox production quality\. MERT2 · Music representations ### Learning the structure behind the sound\. [Vector PDF](https://map-yue2.github.io/static/figures/mert2-ssl.pdf) [![MERT2 architecture: offline target synthesis combines MuQ and Qwen2-Audio features into four code streams. A ConvNeXt frontend and 24-layer Conformer learn by masked prediction, then branch into full-song representations for SheetSage2 and a causal tokenizer curriculum for YuE2.](https://map-yue2.github.io/static/figures/mert2-ssl.svg)](https://map-yue2.github.io/static/figures/mert2-ssl.pdf)MERT2 provides the music representations behind both analysis and generation\. A ConvNeXt frontend and 24\-layer Conformer learn to predict masked codes built from complementary MuQ and Qwen2\-Audio features\. Full\-song adaptation supplies SheetSage2 with musical context; a separate causal branch becomes YuE2’s 25\-Hz semantic tokenizer\.#### State of the art on MARBLE MERT2\-30s and MERT2\-FS \(full\-song\) achieve**SOTA on 14 of 15 MARBLE metrics**, leading across tagging, key, genre, and emotion recognition\. SOTA metricsMERT2\-30s & MERT2\-FS14/ 15Genre accuracy · GTZANMERT2\-30s · score × 10091\.72Key refined accuracy · GiantStepsMERT2\-FS · score × 10067\.05 Explore MERT2 benchmark scoresMARBLE · 15 metrics · 2 encodersSOTA counts use the best score across the two MERT2 encoders against the nine published baselines in this comparison\. Both encoders have 632M parameters\. MERT2\-30s uses a 30\-second training context; MERT2\-FS uses 300 seconds\. These are full\-context representation benchmarks\. MERT2 reports the best observed results across representations selected using test scores; each ROC\-AUC / AP pair uses the same representation\. SheetSage2 · Audio to score ### Hear a song\. Read its composition\. [Vector PDF](https://map-yue2.github.io/static/figures/sheetsage2-overview.pdf) [![SheetSage2 architecture: a full-song MERT2-FS encoder with trainable adapters feeds a six-layer autoregressive decoder. Task prompts select beat, section, key, chord, and melody events, which share a timeline and become ABC notation and a lead sheet.](https://map-yue2.github.io/static/figures/sheetsage2-overview.svg)](https://map-yue2.github.io/static/figures/sheetsage2-overview.pdf)SheetSage2 turns a recording into an editable lead sheet\. A full\-song MERT2\-FS encoder, adapted with LoRA, feeds a six\-layer autoregressive decoder\. Task prompts select beats, sections, keys, chords, and melodies; a shared event timeline becomes ABC notation with vocal and instrumental melody voices\. These scores supply symbolic training targets for YuE2\.#### Six transcription tasks, one model SheetSage2 achieves**SOTA on 10 of 13 benchmark metrics**with one model for beat, downbeat, key, chord, structure, and melody transcription\. SOTA metricsOne model · six transcription tasks10/ 13Vocal melody · RWC\-PopPitch\-class note F1 · score × 10082\.51Chord recognition · osu2017Maj/min · score × 10090\.08 Explore SheetSage2 benchmark scores6 tasks · 13 metricsSOTA counts refer to the leading scores against SheetSage1, Madmom, and the task\-specific systems in this comparison\. Results use one model selected by validation loss\. Melody F1 measures pitch\-class notes; structure F1 measures section boundaries at the stated tolerance\. On Chords1217, ChordFormer uses five\-fold cross\-validation, while SheetSage2 evaluates one fixed model on all 1,217 tracks\. ## Training data Our models are trained primarily on CC0 music and synthetic data\.[Tokenwave\.AI](https://www.tokenwave.us/)provides most of our synthetic training data under license\. We are committed to the ethical and responsible use of data\. MERT2700KhoursSheetSage228\.4KhoursYuE2346Khours

Similar Articles

multimodal-art-projection/YuE

GitHub Trending (daily)

YuE2 is an AI model that generates songs by first creating a symbolic score from lyrics and style, then rendering it into audio, achieving frontier quality competitive with Suno.

New Music Model YuE2-3B Released!

Reddit r/LocalLLaMA

YuE2-3B is an open-source music generation model that creates complete songs from lyrics and style prompts, featuring editable musical scores and state-of-the-art performance rivaling commercial models like Suno v5.

Comfy-Org/YuE2

Hugging Face Models Trending

This article details the YuE2 model repackaged for ComfyUI, enabling audio and music generation workflows. It provides model files based on MERT-v2-FullSong and SheetSage2 for easy integration into ComfyUI projects.