Qwen-Image-2.0 Technical Report (57 minute read)
Summary
This technical report presents Qwen-Image-2.0, a new image generation model from Alibaba's Qwen team, detailing its architecture and capabilities.
View Cached Full Text
Cached at: 05/14/26, 12:11 AM
# Qwen-Image-2.0 Technical Report Source: [https://arxiv.org/abs/2605.10730](https://arxiv.org/abs/2605.10730) Authors:[Bing Zhao](https://arxiv.org/search/cs?searchtype=author&query=Zhao,+B),[Chenfei Wu](https://arxiv.org/search/cs?searchtype=author&query=Wu,+C),[Deqing Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+D),[Hao Meng](https://arxiv.org/search/cs?searchtype=author&query=Meng,+H),[Jiahao Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+J),[Jie Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+J),[Jingren Zhou](https://arxiv.org/search/cs?searchtype=author&query=Zhou,+J),[Junyang Lin](https://arxiv.org/search/cs?searchtype=author&query=Lin,+J),[Kaiyuan Gao](https://arxiv.org/search/cs?searchtype=author&query=Gao,+K),[Kuan Cao](https://arxiv.org/search/cs?searchtype=author&query=Cao,+K),[Kun Yan](https://arxiv.org/search/cs?searchtype=author&query=Yan,+K),[Liang Peng](https://arxiv.org/search/cs?searchtype=author&query=Peng,+L),[Lihan Jiang](https://arxiv.org/search/cs?searchtype=author&query=Jiang,+L),[Niantong Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+N),[Ningyuan Tang](https://arxiv.org/search/cs?searchtype=author&query=Tang,+N),[Shengming Yin](https://arxiv.org/search/cs?searchtype=author&query=Yin,+S),[Tianhe Wu](https://arxiv.org/search/cs?searchtype=author&query=Wu,+T),[Xiao Xu](https://arxiv.org/search/cs?searchtype=author&query=Xu,+X),[Xiaoyue Chen](https://arxiv.org/search/cs?searchtype=author&query=Chen,+X),[Xihua Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+X),[Yan Shu](https://arxiv.org/search/cs?searchtype=author&query=Shu,+Y),[Yanran Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+Y),[Yi Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+Y),[Yilei Chen](https://arxiv.org/search/cs?searchtype=author&query=Chen,+Y),[Ying Ba](https://arxiv.org/search/cs?searchtype=author&query=Ba,+Y),[Yixian Xu](https://arxiv.org/search/cs?searchtype=author&query=Xu,+Y),[Yujia Wu](https://arxiv.org/search/cs?searchtype=author&query=Wu,+Y),[Yuxiang Chen](https://arxiv.org/search/cs?searchtype=author&query=Chen,+Y),[Zecheng Tang](https://arxiv.org/search/cs?searchtype=author&query=Tang,+Z),[Zekai Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+Z),[Zhendong Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+Z),[Zihao Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+Z),[Zikai Zhou](https://arxiv.org/search/cs?searchtype=author&query=Zhou,+Z),[An Yang](https://arxiv.org/search/cs?searchtype=author&query=Yang,+A),[Chen Cheng](https://arxiv.org/search/cs?searchtype=author&query=Cheng,+C),[Chenxu Lv](https://arxiv.org/search/cs?searchtype=author&query=Lv,+C),[Dayiheng Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+D),[Fan Zhou](https://arxiv.org/search/cs?searchtype=author&query=Zhou,+F),[Hantian Xiong](https://arxiv.org/search/cs?searchtype=author&query=Xiong,+H),[Hongzhu Shi](https://arxiv.org/search/cs?searchtype=author&query=Shi,+H),[Hu Wei](https://arxiv.org/search/cs?searchtype=author&query=Wei,+H),[Huihong Zhao](https://arxiv.org/search/cs?searchtype=author&query=Zhao,+H),[Ivy Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+I),[Jianwei Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+J),[Jiawei Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+J),[Kai Chen](https://arxiv.org/search/cs?searchtype=author&query=Chen,+K),[Kang He](https://arxiv.org/search/cs?searchtype=author&query=He,+K),[Levon Xue](https://arxiv.org/search/cs?searchtype=author&query=Xue,+L),[Lin Qu](https://arxiv.org/search/cs?searchtype=author&query=Qu,+L),[Linhan Tang](https://arxiv.org/search/cs?searchtype=author&query=Tang,+L),[Luwen Feng](https://arxiv.org/search/cs?searchtype=author&query=Feng,+L),[Minggang Wu](https://arxiv.org/search/cs?searchtype=author&query=Wu,+M),[Minmin Sun](https://arxiv.org/search/cs?searchtype=author&query=Sun,+M),[Na Ni](https://arxiv.org/search/cs?searchtype=author&query=Ni,+N),[Rui Men](https://arxiv.org/search/cs?searchtype=author&query=Men,+R),[Shuai Bai](https://arxiv.org/search/cs?searchtype=author&query=Bai,+S),[Sishou Zheng](https://arxiv.org/search/cs?searchtype=author&query=Zheng,+S),[Tao Lan](https://arxiv.org/search/cs?searchtype=author&query=Lan,+T),[Tianqi Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+T),[Tingkun Wen](https://arxiv.org/search/cs?searchtype=author&query=Wen,+T),[Wei Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+W),[Weixu Qiao](https://arxiv.org/search/cs?searchtype=author&query=Qiao,+W),[Weiyi Lu](https://arxiv.org/search/cs?searchtype=author&query=Lu,+W),[Wenmeng Zhou](https://arxiv.org/search/cs?searchtype=author&query=Zhou,+W),[Xiaodong Deng](https://arxiv.org/search/cs?searchtype=author&query=Deng,+X),[Xiaoxiao Xu](https://arxiv.org/search/cs?searchtype=author&query=Xu,+X),[Xinlei Fang](https://arxiv.org/search/cs?searchtype=author&query=Fang,+X),[Xionghui Chen](https://arxiv.org/search/cs?searchtype=author&query=Chen,+X),[Yanan Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+Y),[Yang Fan](https://arxiv.org/search/cs?searchtype=author&query=Fan,+Y),[Yichang Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+Y),[Yixuan Xu](https://arxiv.org/search/cs?searchtype=author&query=Xu,+Y),[Yu Wu](https://arxiv.org/search/cs?searchtype=author&query=Wu,+Y),[Zhiyuan Ma](https://arxiv.org/search/cs?searchtype=author&query=Ma,+Z),[Zhizhi Cai](https://arxiv.org/search/cs?searchtype=author&query=Cai,+Z) [View PDF](https://arxiv.org/pdf/2605.10730) > Abstract:We present Qwen\-Image\-2\.0, an omni\-capable image generation foundation model that unifies high\-fidelity generation and precise image editing within a single framework\. Despite recent progress, existing models still struggle with ultra\-long text rendering, multilingual typography, high\-resolution photorealism, robust instruction following, and efficient deployment, especially in text\-rich and compositionally complex scenarios\. Qwen\-Image\-2\.0 addresses these challenges by coupling Qwen3\-VL as the condition encoder with a Multimodal Diffusion Transformer for joint condition\-target modeling, supported by large\-scale data curation and a customized multi\-stage training pipeline\. This enables strong multimodal understanding while preserving flexible generation and editing capabilities\. The model supports instructions of up to 1K tokens for generating text\-rich content such as slides, posters, infographics, and comics, while significantly improving multilingual text fidelity and typography\. It also enhances photorealistic generation with richer details, more realistic textures, and coherent lighting, and follows complex prompts more reliably across diverse styles\. Extensive human evaluations show that Qwen\-Image\-2\.0 substantially outperforms previous Qwen\-Image models in both generation and editing, marking a step toward more general, reliable, and practical image generation foundation models\. ## Submission history From: Shengming Yin \[[view email](https://arxiv.org/show-email/8ff770d2/2605.10730)\] **\[v1\]**Mon, 11 May 2026 15:34:56 UTC \(45,347 KB\)
Similar Articles
Qwen-Image-2.0 Technical Report
Qwen-Image-2.0 is a new image generation foundation model that unifies high-fidelity synthesis and precise editing using Qwen3-VL and a Multimodal Diffusion Transformer. It excels in text-rich content, multilingual typography, and photorealistic generation.
Qwen-Image-Flash (26 minute read)
This paper from Alibaba revisits few-step distillation for visual generative models, focusing on training recipe factors such as data composition, teacher guidance, and task mixture, using Qwen-Image-2.0 as a case study to develop Qwen-Image-Flash.
Qwen3.7 Preview lands on Arena (1 minute read)
Alibaba Qwen announces two major model releases: Qwen3-Omni, the first natively end-to-end omni-modal AI unifying text, image, audio and video, and Qwen3-Next-80B-A3B, an ultra-efficient MoE model with 3B activated parameters per token, achieving SOTA performance and 10x faster inference than Qwen3-32B.
@AdinaYakup: Qwen @Alibaba_Qwen just dropped a new Text to Image benchmark + a judge model https://huggingface.co/collections/Qwen/q…
Qwen released a new Text-to-Image benchmark with 56 fine-grained evaluation facets, measuring creativity beyond prompt alignment, and includes a human-aligned judge model.
Qwen-Image-VAE-2.0 Technical Report
Qwen-Image-VAE-2.0 is a high-compression Variational Autoencoder suite that improves reconstruction fidelity and diffusability through enhanced architecture, large-scale training, and semantic alignment strategies.