@ArtificialAnlys: We have updated the Artificial Analysis Image Editing Arena to expand the range of editing tasks we test for, from Enha…
Summary
The Artificial Analysis Image Editing Arena has been updated to expand editing tasks and use cases, providing a comprehensive leaderboard that ranks AI models on human preference for various image editing scenarios.
View Cached Full Text
Cached at: 09/03/26, 08:14 PM
We have updated the Artificial Analysis Image Editing Arena to expand the range of editing tasks we test for, from Enhancement & Restoration to Identity-Preserving Edits, and from Marketing & Advertising to UI/UX Design. Our updated evaluation measures model performance on complex edit tasks, and assesses not just which model is best overall, but which model is best for specific editing needs. The new leaderboard is live and voting is open.
Image Editing models are advancing fast. AI now fits into different stages of image creation workflows, from generation through to post-production, and editing is no longer a side feature. We treat Image Editing and Reference to Image as separate benchmarks because they sit at different points in that workflow: Image Editing covers post-production, changing an image you already have; while our upcoming Reference to Image benchmark covers generating novel images from reference images.
For Image Editing, we test a model’s ability to make specific changes to an image while keeping everything else unchanged. This includes complex edit instructions that chain multiple different asks, as frontier models have largely saturated single-instruction edits.
Different edit requests call for different models. Relighting a cinematic scene is a different problem from reworking the design of a marketing asset. We rank models on human preference across 7 editing actions, such as Object-Level Edit, Identity-Preserving Edit, and Enhancement & Restoration, and 10 real-world use cases, such as Marketing & Advertising, UI/UX Design, and Live-Action Film. The overall benchmark samples evenly across both.
Initial insights from an in-depth analysis of the 10 highest ranking models on the Artificial Analysis Image Editing Leaderboard:
➤ MAI-Image-2.6-Preview leads the overall leaderboard and 4 of the 7 editing action boards: Scene & Style Edit, Text or Symbol Edits, Reasoning-Based Edit, and Enhancement & Restoration, where it is tied #1 with MAI-Image-2.5. It excels at restyling, relighting, and retouching images.
➤ GPT Image 2 (high) ranks #2 overall but #1 on Object-Level Edit and Composition & Framing. It is the strongest at precise local edits and spatial reframing. It is weaker at whole-image transformations that must keep the image’s content intact, ranking #7 on Scene & Style Edit and #6 on Enhancement & Restoration.
➤ Seedream 5.0 Pro is the character and identity specialist, ranking #1 on Identity-Preserving Edit.
➤ MAI-Image-2.5-Flash is the value pick of the top 10, at $20 per 1,000 images against $211 for GPT Image 2 (high) and $90 for Seedream 5.0 Pro.
See below for the editing action and use case breakdowns 🧵
The best model depends on the editing task. MAI-Image-2.6-Preview leads the overall Image Editing Leaderboard and the most category boards, taking #1 on 7 of the 17. GPT Image 2 is next with 5, concentrated in precise object-level and spatial edits.
Example Editing Action: Enhancement & Restoration measures image cleanup, including denoising, deblurring, flare and artifact removal, and restoring old or degraded photos without inventing detail.
MAI-Image-2.5, MAI-Image-2.6-Preview, MAI-Image-2.5-Pro, and MAI-Image-2.5-Flash sweeps the top 4, with GPT Image 2 back at #6. The MAI models preserves minute details of the images and applies proper relighting while cleaning up noise, which is exactly what this editing action scores.
The pattern extends to Scene & Style Edit, where MAI-Image-2.6-Preview leads, and GPT Image 2 falls to #7, its weakest editing action.
Example Editing Action: Composition & Framing measures 2D and 3D reframing, including camera angle changes, zoom-outs that extend a scene, and rearranging graphic layouts while preserving design, spacing, and style.
GPT Image 2 leads MAI-Image-2.6-Preview here by 20 Elo, demonstrating better ability to preserve spatial coherency while changing camera angle, and to keep graphic elements intact while rearranging a 2D layout.
Example Editing Action: Identity-Preserving Edit measures whether a character survives the edit, including the same face, likeness, outfit, and accessories while the pose, movement, age, or expression changes.
Seedream 5.0 Pro ranks higher than #1 overall editing model MAI-Image-2.6-Preview here, preserving detailed character design from likeness to accessories, and it does so at $90 per 1,000 images, with GPT Image 2 one rank behind at $211.
Part of Identity-Preserving Edit is Human Anatomy generation, since the edit should yield clearly-rendered faces, hands, and bodies. We also tag each prompt by what the model must render as part of the editing task, and on that Human Anatomy axis GPT Image 2 and Seedream 5.0 Pro lead the top 10 while the MAI models sit mid-pack or lower.
Example Use Case: UI/UX editing includes changing UI component sizes, colors, styles, and in-UI text for app, web, in-car, and spatial UI mockups and flows, dashboards, screens, icons, and design system assets.
MAI-Image-2.6-Preview leads by 26 Elo, with MAI-Image-2.5 behind it. GPT Image 2 and Nano Banana 2 follow at #3 and #5.
UI edits are dense with in-image text, and the Text or Symbol Edits board tells the same story: MAI-Image-2.6-Preview leads that editing action by 30 Elo. MAI-Image-2.5 ranks #3 at $48.1 per 1,000 images.
Example Use Case: Productivity & Knowledge Work covers edits for charts, diagrams, infographics, data visualizations, slides, and document graphics.
Productivity & Knowledge Work is a three-way tie at the top: GPT Image 2, Nano Banana Pro and MAI-Image-2.6-Preview sit within the margin of error of each other.
Productivity edits combine two editing actions: Object-Level Edits, moving a legend or changing one series, and Reasoning-Based Edits, since the edit must respect the information the chart or diagram conveys. GPT Image 2 is the only model in the top 10 ranking top on both, #1 on Object-Level Edit and #2 on Reasoning-Based Edit.
Example Use Case: Marketing & Advertising covers edits on posters, campaign assets, product shots, and brand visuals.
MAI-Image-2.6-Preview leads GPT Image 2 here by 23 Elo. Marketing assets lean on the two editing actions it leads: Scene & Style Edit for restyling and relighting campaign visuals, and Text or Symbol Edits for logos and copy, which it leads by 36 Elo. GPT Image 2 at #2 runs $211 per 1,000 images. Qwen-Image-3.0-Pro at #3 is the budget alternative at $40.
How much does editing quality cost?
The price to quality frontier runs from MAI-Image-2.5-Flash at $20 per 1,000 images to GPT Image 2 (high) at 211, with Qwen-Image-3.0-Pro (40), MAI-Image-2.5 (48.1), and Nano Banana 2 (67) in between. MAI-Image-2.6-Preview sits above the frontier, with pricing not yet announced.
Vote in the Image Editing Arena now: https://artificialanalysis.ai/text-to-image/arena…
Full methodology, including how prompts are authored, refreshed, and retired, and how Elo is calculated: https://artificialanalysis.ai/image/methodology#image-editing…
Similar Articles
Artificial Analysis updates its Intelligence Index to version 4.3
Artificial Analysis has updated its Intelligence Index to version 4.3, incorporating new benchmarks like Terminal-Bench v4.0 and AutomationBench-AA to better evaluate AI model performance and cost-efficiency.
Artificial Analysis Capability Indices v1.1 (2 minute read)
Artificial Analysis has released Intelligence Index v4.2, an interim update with new evaluations like AA-Briefcase and GDP.pdf, increased private test sets to prevent gaming, and key results showing Anthropic and OpenAI leading.
@rohanpaul_ai: Arena just released a real-world agent leaderboard that ranks AI models by how well they complete actual user jobs, not…
Agent Arena is a new leaderboard that evaluates AI models on real-world agentic tasks such as coding, research, and file analysis, using signals like task success, steerability, and recovery, with GPT-5.5 High leading.
IA para imágenes artísticas
An introduction to AI tools for generating and editing artistic images, likely covering popular models or platforms.
MAI-Image-2.6 Reaches No. 2 on Arena (4 minute read)
Microsoft announces MAI-Image-2.6, an image generation model that ranks No. 2 on the Arena text-to-image leaderboard, improving on its predecessor with gains in text rendering, photorealism, and other categories.