Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning

Hugging Face Daily Papers Papers

Summary

Uni-Edit proposes using intelligent image editing as a single general task to simultaneously improve unified multimodal models' understanding, generation, and editing capabilities, with an automated data synthesis pipeline creating complex editing instructions.

Currently, enhancing Unified Multimodal Models (UMMs) with image understanding, generation, and editing capabilities mainly relies on mixed multi-task training. Due to inherent task conflicts, such strategy requires complex multi-stage pipelines, massive data mixing, and balancing tricks, merely resulting in a performance trade-off rather than true mutual reinforcement. To break this paradigm, we propose Uni-Edit, an intelligent image editing task that serves as the first general task for UMM tuning. Unlike complex mixed pipelines, Uni-Edit improves performance across all three abilities at once using only one task, one training stage, and one dataset. Specifically, we first identify image editing as an inherently ideal general task, as it naturally demands both visual understanding and generation. However, existing editing data relies on simplistic instructions that severely underutilize a model's understanding capacity. To address this, we introduce the first automated and scalable data synthesis pipeline for intelligent editing, transforming diverse VQA data into complex and effective editing instructions with embedded questions and nested logic. This yields Uni-Edit-148k, pairing diverse reasoning-intensive instructions with high-quality edited images. Extensive experiments on BAGEL and Janus-Pro demonstrate that tuning solely on Uni-Edit achieves comprehensive enhancements across all three capabilities without any auxiliary operations.
Original Article
View Cached Full Text

Cached at: 05/21/26, 06:20 AM

Paper page - Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning

Source: https://huggingface.co/papers/2605.21487

Abstract

Uni-Edit introduces an intelligent image editing task that simultaneously enhances unified multimodal models’ understanding, generation, and editing capabilities through a single training stage and dataset, utilizing an automated data synthesis pipeline for complex editing instructions.

Currently, enhancingUnified Multimodal Models(UMMs) with image understanding, generation, and editing capabilities mainly relies on mixedmulti-task training. Due to inherent task conflicts, such strategy requires complex multi-stage pipelines, massive data mixing, and balancing tricks, merely resulting in a performance trade-off rather than true mutual reinforcement. To break this paradigm, we propose Uni-Edit, an intelligentimage editingtask that serves as the first general task for UMM tuning. Unlike complex mixed pipelines, Uni-Edit improves performance across all three abilities at once using only one task, one training stage, and one dataset. Specifically, we first identifyimage editingas an inherently ideal general task, as it naturally demands both visual understanding and generation. However, existing editing data relies on simplistic instructions that severely underutilize a model’s understanding capacity. To address this, we introduce the first automated and scalabledata synthesis pipelinefor intelligent editing, transforming diverseVQA datainto complex and effective editing instructions with embedded questions and nested logic. This yields Uni-Edit-148k, pairing diversereasoning-intensive instructionswith high-quality edited images. Extensive experiments onBAGELandJanus-Prodemonstrate that tuning solely on Uni-Edit achieves comprehensive enhancements across all three capabilities without any auxiliary operations.

View arXiv pageView PDFProject pageGitHub2Add to collection

Models citing this paper1

#### Uni-Edit/Uni-Edit-BAGEL Any-to-Any• 15B• Updatedabout 4 hours ago • 14 • 1

Datasets citing this paper1

#### Uni-Edit/Train-Data Viewer• Updatedabout 4 hours ago • 8.39k • 24

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2605.21487 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles