Tag
Elon Musk highlights that Grok Imagine's image editing capabilities have been greatly improved, with precise segment editing tools that let users target specific parts of an image instead of regenerating the whole thing.
Grok announces Imagine Image 2.0, a next-generation image model with precision editing, crisp text rendering, and improved factuality for real-world use.
SenseNova U1.5-Lite-Preview is an open-source 8B MoT multimodal model that natively generates and edits ultra-wide panoramas, product posters, and commercial photography at 4K, with improved material rendering and fewer artifacts.
Adrian Sieber announces the 1.0 release of Perspec, a desktop app for correcting the perspective of images, useful for photos of documents and receipts.
UniWorld-Design is a framework that redefines image generation using semantic RGBA layers as atomic units, comprising T2RGBA for generating layered assets and I2L for decomposing images into editable layers, achieving state-of-the-art results on the Crello benchmark.
This paper introduces a Multi-dimensional Evaluation-Verification Reward (EVR) for reinforcement learning fine-tuning of multi-reference image editing models, improving visual consistency and harmony.
Introduces MPIE-Bench, a 2,500-sample benchmark for multi-person interaction image editing, along with MPIE-Eval, a mesh-based evaluation method that tracks human judgment more closely than VLM checklists across ten editors.
Microsoft releases Mage, a family of lightweight 4B-parameter multimodal models for visual understanding and generation, including Mage-VL for image/video understanding and Mage-Flow for text-to-image generation and editing, designed for research and deployment on modest hardware.
Microsoft releases Mage-Flow-Edit-Turbo, a compact 4B-scale generative model for efficient text-to-image generation and instruction-based image editing, achieving state-of-the-art competitive quality through co-designed tokenizer and backbone.
Microsoft releases Mage-Flow, a compact 4B-parameter foundation model for efficient native-resolution text-to-image generation and instruction-based image editing, achieving competitive quality against much larger models.
Mage-Flow is a compact 4B-parameter generative stack for efficient text-to-image generation and instruction-based image editing, featuring a co-designed lightweight tokenizer (Mage-VAE) and a native-resolution multimodal diffusion transformer trained with rectified flow matching. It achieves competitive performance while enabling high-resolution generation at 0.59s on a single A100 GPU.
GIMP 3.0, its first major release in seven years, fixes long-standing issues like the confusing floating selection mechanism and introduces non-destructive editing for most GEGL-based effects, significantly improving the user experience.
Mojave Paint is a tool for direct manipulation of static images on the Mac platform.
Boogu-Image-0.1 is an open-source family of unified multimodal understanding and generation models that achieves competitive performance in text-to-image generation, fast inference, instruction-based editing, and bilingual text rendering, with low training cost of approximately $400K.
This paper introduces RINO (RGB In and RGB Out), a unified framework that represents diverse visual information (masks, depth, etc.) as RGB images and converts visual tasks into RGB-to-RGB image editing, enabling a single model to perform zero-shot transfer across tasks.
Seedream 5.0 Pro is a highly controllable AI image editing model, accessible via BytePlusGlobal API and Lumina platform, enabling precise edits for humans, items, and animals.
This paper introduces VIP-SAM for instance-level garment segmentation and CtrlVTON, a controllable virtual try-on framework that treats try-on as an image editing problem, allowing precise control over garment layout, style, and placement. Both methods achieve state-of-the-art results on their respective tasks.
The user tested the frontend capabilities of GPT 5.6-Terra, creating an interactive Guan Yu card webpage from the Three Kingdoms. The results were outstanding, particularly in image matting and depth-of-field perspective.
A community fine-tune of Krea 2 Raw that enables instruction-based, identity-preserving image editing. It edits images while preserving details like faces, using a custom ComfyUI node pack.
Boogu-Image-0.1 is a new unified image generation and editing model family with 10B parameters, available under Apache 2.0 license. It features fast turbo inference in 4 steps, trained on 10x less data, and supports Chinese and English.