Tag
Perceptual Flow Matching supervises flow matching in perceptual feature space, enabling high-quality few-step generation with 4-8 sampling steps instead of 35-50, without needing teacher models.
SenseNova-U1-8b-MoT-Infographic-V2 is an open-source state-of-the-art model released by SenseTime for infographic design and image editing tasks.
MirrorPPR introduces an exemplar-based portrait retouching framework using Diffusion Transformer with LoRA adaptation and self-augmented training data, achieving superior quality and identity preservation.
A collection of 50 practical websites covering categories like bypassing paywalls, free paper downloads, image/video editing, code beautification, privacy & security, music & audio, research tools, and developer tools, offering convenience for daily work and life.
This technical report presents Qwen-Image-2.0-RL, a post-training pipeline using reinforcement learning from human feedback and on-policy distillation to enhance visual quality and instruction-following in image generation and editing tasks.
DanceOPD proposes an on-policy generative field distillation framework for flow-matching models that unifies text-to-image generation, local editing, and global editing via capability-specific routing and velocity-based training, improving multi-capability composition while preserving anchor generation quality.
Boogu has released a series of open-source unified image generation and editing models, including Base, Turbo, and Edit variants.
Demonstrates a workflow combining Codex and Excalidraw: open Excalidraw in Codex's built-in browser, paste an image and annotate, then take a screenshot for Codex to modify — enabling point-and-fix creative editing.
Photoroom introduces an API for transforming product images at scale, enabling automated image editing.
ImageWAM proposes replacing video generation with pretrained image editing models in world action models for robot control, achieving superior performance while reducing FLOPs to 1/6 and latency to 1/4 of video-based approaches.
Boogu-Image-0.1 is an Apache-2.0 open-source unified image generation and editing model family, including variants for text-to-image, fast generation, editing, and Chinese-English text rendering, released as a research project on Hugging Face.
UniAR presents a unified autoregressive framework that uses a single discrete visual tokenizer to bridge visual understanding and generation, achieving state-of-the-art results in image generation and editing.
A new framework called TV-Edit combines textual instructions and visual prompts for precise image editing, along with a benchmark TV-Edit-Bench for evaluation. The method achieves better spatial control and semantic faithfulness than existing approaches.
An open-source Android image editing toolbox built with Kotlin and Jetpack Compose, offering various image manipulation tools.
The comment acknowledges that the model is state-of-the-art for editing but not for generation.
HiLo-Token introduces an input-adaptive token compression framework for Diffusion Transformers that allocates more tokens to high-frequency regions, achieving up to 3.13x speedup in image editing tasks without quality loss.
Airbrush Studio is an AI-powered photo editor that delivers professional results without requiring manual editing.
Apple announced new AI-powered editing features for its Photos app at WWDC 2026, including Reframe, Extend, and an upgraded Cleanup tool, leveraging Apple Intelligence.
Pixel Snapper is an editor designed to clean up AI-generated pixel art.
This paper from Alibaba revisits few-step distillation for visual generative models, focusing on training recipe factors such as data composition, teacher guidance, and task mixture, using Qwen-Image-2.0 as a case study to develop Qwen-Image-Flash.