WanSong v1.0 Technical Report
Summary
WanSong is a pure diffusion-based music generation model that directly produces high-fidelity, multilingual songs up to 5 minutes long with dual stems (vocals and background music) in a single run, addressing challenges in efficient generation, long-form audio, and controllability.
View Cached Full Text
Cached at: 07/20/26, 09:43 AM
Paper page - WanSong v1.0 Technical Report
Source: https://huggingface.co/papers/2607.14749
Abstract
Musicgenerationfoundationmodelshaverecentlyattractedsignificantindustryattention.However,achievingefficientgenerationandhigh-fidelitylong-formaudiowhilesupportingcontrollabilityremainschallenging.Toaddresstheseneeds,wepresentWanSong,asimpleyetpowerfulapproachforlong-form,commercial-gradesonggeneration.Unlikeautoregressive(AR)andcascadedmulti-stagepipelines(\eg,ARfollowedbydiffusion),WanSongisapurediffusion-basedmodelthatdirectlygenerateshigh-fidelity,multilingualsongsupto5minutesandoutputsdualstems(vocalsandbackgroundmusic)inasinglerun.Inaddition,ourdiffusionframeworkenablesfasterinferencethroughstep-distillation,andoffersanefficientpathwayforfine-tuningandcustomizationtosupportdownstreameditingtasks.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2607\.14749
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.14749 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.14749 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.14749 in a Space README.md to link it from this page.
Collections including this paper3
Similar Articles
Qwen-Music Technical Report
Qwen-Music is a music generation model that produces high-fidelity songs with vocals, supporting text-to-music and cover song generation. It uses a novel Melody-Chain-of-Thought mechanism and achieves state-of-the-art results on 13 of 16 objective metrics.
@junmingong: Khala 1.0 just dropped — a music generation model from the Central Conservatory of Music in Beijing. Paper, code, weigh…
Khala 1.0 is an open-source music generation model for high-fidelity full-song generation from text and lyrics, using a unified acoustic-token pipeline. It was released by the Central Conservatory of Music in Beijing with paper, code, weights, and demo.
@askalphaxiv: “Qwen-Music Technical Report” AI music generation has to solve two problems at once, composing a coherent song and rend…
Qwen-Music is a new paper that splits music generation into two stages: planning with compact semantic tokens and CoT reasoning over melody, then rendering high-fidelity audio, achieving state-of-the-art results.
Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation
Wan-Dancer introduces a hierarchical framework for generating minute-scale coherent dances from music, addressing long-duration choreography generation.
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models
Wan-Streamer is a unified end-to-end multimodal model for real-time audio-visual interaction using causal attention and integrated processing of visual, audio, and text modalities, achieving sub-second latency.