Qwen-Music Technical Report
Summary
Qwen-Music is a music generation model that produces high-fidelity songs with vocals, supporting text-to-music and cover song generation. It uses a novel Melody-Chain-of-Thought mechanism and achieves state-of-the-art results on 13 of 16 objective metrics.
View Cached Full Text
Cached at: 07/20/26, 09:39 AM
Paper page - Qwen-Music Technical Report
Source: https://huggingface.co/papers/2607.11699 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Inthisreport,weintroduceQwen-Music,apowerfulmusicgenerationmodelcapableofproducinghighlymusicalandhigh-fidelitysongswithcompletevocalsinging.Qwen-Musicsupportstwocoretasks:TexttoMusicGeneration,whichcreateentirelynewsongsfromtextdescriptions,lyrics,andmusicalattributes,andCoverSongGeneration,whichreinterpretsexistingsongswithdifferentstylesandvocalcharacteristics.Architecturally,Qwen-Musicintegratesthreecorecomponents:Qwen-Music-Tokenizer,Qwen-Music-LLM,andQwen-Music-Render.Qwen-Music-Tokenizercompressesaudiointoa25Hzsingle-codebookstreamofMusicSemanticTokensthatpreservesemanticandmelodicinformationforLLMprediction.Basedonthesetokens,Qwen-Music-LLMperformsautoregressivemusicsemanticmodeling,withakeynoveltybeingamelody-token-basedchain-of-thought(Melody-CoT)mechanismthatplansmelodiesbeforefull-songgeneration,improvingcreativity,musicality,structuralcoherence,andreference-audio-basedmelodycloning.Toovercomethefidelitylimitationsofdiscretesemantictokens,Qwen-Music-Renderperformsgenerativestereorendering,enrichingacousticdetailsandproducinghigh-fidelitystereowaveforms.Finally,wetrainQwen-Music-LLMonmorethan5millionhoursofmultilingualmusicdatacoveringhundredsoflanguages.Wefirstapplyquality-awarepre-trainingcurriculum,thenuseprogressivepost-training,comprisingsupervisedinitialization,offlineDPO,andonlineGSPO,tofurtherimprovemusicalityandinstruction-followingability.Across600ChineseandEnglishprompts,Qwen-Musicachievesstate-of-the-artresultsin13of16objectivemusicalityandaudio-qualitymetrics.ProfessionalevaluatorsalsopreferQwen-Musicoverleadingproprietarysystems.Forcoversonggeneration,Qwen-Musicpreservesreferencemelodiesmoreaccuratelythanleadingproprietarysystems.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2607\.11699
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.11699 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.11699 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.11699 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
@askalphaxiv: “Qwen-Music Technical Report” AI music generation has to solve two problems at once, composing a coherent song and rend…
Qwen-Music is a new paper that splits music generation into two stages: planning with compact semantic tokens and CoT reasoning over melody, then rendering high-fidelity audio, achieving state-of-the-art results.
Qwen-Image-2.0 Technical Report
Qwen-Image-2.0 is a new image generation foundation model that unifies high-fidelity synthesis and precise editing using Qwen3-VL and a Multimodal Diffusion Transformer. It excels in text-rich content, multilingual typography, and photorealistic generation.
Qwen3-TTS Technical Report
The Qwen3-TTS technical report introduces a series of advanced multilingual text-to-speech models with voice cloning and controllable generation, featuring a dual-track LM architecture and specialized tokenizers for low-latency streaming.
Qwen3.5-Omni Technical Report
Qwen3.5-Omni is a hundreds-of-billions-parameter multimodal model with advanced audio-visual understanding and generation capabilities, featuring novel Audio-Visual Vibe Coding and achieving SOTA results across 215 benchmarks while matching Gemini-3.1 Pro.
Qwen-Image-2.0 Technical Report (57 minute read)
This technical report presents Qwen-Image-2.0, a new image generation model from Alibaba's Qwen team, detailing its architecture and capabilities.