microsoft/Mage-Flow-Edit-Turbo
Summary
Microsoft releases Mage-Flow-Edit-Turbo, a compact 4B-scale generative model for efficient text-to-image generation and instruction-based image editing, achieving state-of-the-art competitive quality through co-designed tokenizer and backbone.
View Cached Full Text
Cached at: 07/27/26, 01:46 AM
microsoft/Mage-Flow-Edit-Turbo · Hugging Face
Source: https://huggingface.co/microsoft/Mage-Flow-Edit-Turbo
Mage-Flow An Efficient Native-Resolution Foundation Model for Image Generation and Editing

Mage-Flowis a compact4B-scale generative stackfor efficienttext-to-image generationandinstruction-based image editing. Instead of scaling to tens of billions of parameters, Mage-Flow reaches state-of-the-art-competitive quality through carefultokenizer–backbone–system co-design, so it stays fast, memory-light, and easy to fine-tune under realistic compute budgets.
The stack is built fromtwo shared, co-designed components:
- Mage-VAE— a lightweight, high-fidelity latent tokenizer (one-step diffusion encode/decode with anchor-latent KL regularization).
- NR-MMDiT— a shared 4BNative-Resolution Multimodal Diffusion Transformer, trained with rectified flow matching in the Mage-VAE latent space.
Together with native-resolution packing and a fused-kernel training infrastructure, this shared stack powerstwo model instantiations:Mage-Flowfor text-to-image generation andMage-Flow-Editfor instruction-based image editing. Each ships inBase,RL-aligned, and4-step Turbovariants.
https://huggingface.co/microsoft/Mage-Flow-Edit-Turbo#%E2%9C%A8-highlights✨ Highlights
- **Compact & competitive.**A single 4B family for generationandediting that matches or beats much larger open systems (Qwen-Image 20B, Z-Image 6B, FLUX.2 32B, FireRed-Image-Edit 20B).
- Efficient tokenizer.Mage-VAE matches FLUX.2-VAE reconstruction fidelity while using~12× / ~22× fewer encode / decode MACs per pixel, removing the VAE as the high-resolution bottleneck.
- Native resolution.One checkpoint generates from512 to 2048on any aspect ratio, including extreme4:1(e.g.
512×2048,2048×512). - System-level speed.Native-resolution packing (FlashAttention var-len + per-sample 2D RoPE) + fused CUDA kernels cut per-step training time from~1.93 s → ~0.78 s(~2.5× faster training); CFG’s conditional/unconditional branches run inonepacked forward.
- Full family.**Base,RL-aligned, and4-step Turbo**variants for both generation and editing.
- **Versatile editing.**Mage-Flow-Edit supports semantic content editing, appearance transformation, image restoration, and structure-aware outputs within a unified image-and-text-conditioned model. See the report’s editing galleries.
- Interactive latency.At
1024²on a single A100:Mage-Flow-Turbo 0.59 s/image,Mage-Flow-Edit-Turbo 1.02 s/edit, peak memory~18–20 GB(lowest among compared systems).
One-to-many editing diversity — Mage-Flow-Edit can generate diverse outputs from a single reference image.
https://huggingface.co/microsoft/Mage-Flow-Edit-Turbo#%F0%9F%93%A5-model-zoo📥 Model Zoo
Each checkpoint is a self-contained diffusers-style repo (transformer/+ sharedvae/,text\_encoder/,scheduler/).
https://huggingface.co/microsoft/Mage-Flow-Edit-Turbo#%F0%9F%96%BC%EF%B8%8F-showcase🖼️ Showcase
Text-to-image— prompt following, fine detail, and legible English/Chinese text rendering.(The first panel is open; click a title to expand the others.)
Showcase
General scenes
Portraits
Cuisine & still life
English text rendering
Chinese text rendering
Instruction-based editing— appearance, content, scene/subject, human-centered & creative, low-level, and restoration edits (source → result).(The first panel is open; click a title to expand the others.)
Various Editing I
Various Editing II
Localized content & object editing
Scene, subject & camera transformations
Appearance & artistic rendering
Human-centered & creative editing
Low-level vision & conditional reconstruction
Bidirectional degradation & restoration
https://huggingface.co/microsoft/Mage-Flow-Edit-Turbo#%F0%9F%93%8A-performance📊 Performance
Full benchmark tables (text-to-image & image editing) — click to expand****Text-to-image— full benchmark suite: GenEval, DPG-Bench, TIIF-Bench (short/long splits), CVTG-2K, OneIG (EN/CN), LongText (EN/CN). Higher is better; GenEval / CVTG-2K / OneIG / LongText on a 0–1 scale, DPG / TIIF on 0–100. TheTypecolumn marks closed- vs open-source;bold/underline= best / second-best amongopen-sourcemodels (closed-source shown for reference, not ranked);–= not reported; ★ = ours.
ModelType#ParamsStepsGenEvalDPGTIIF-ShortTIIF-LongCVTG-2KOneIG-ENOneIG-CNLongText-ENLongText-CNSeedream 3.0Closed––0.8488.2786.0284.310.5920.5300.5280.8960.878Seedream 4.0Closed––0.8488.63––0.8920.5730.5540.9360.946GPT-Image-1Closed––0.8485.1589.1588.290.8570.5330.4740.9560.619Nano-Banana-ProClosed––0.8387.16––0.7790.5800.5700.9810.949FLUX.1-devOpen12B500.6683.8471.0971.780.4960.4340.2450.6070.005FLUX.1-Krea-devOpen12B500.7286.5980.3681.670.4440.4430.2710.6930.002FLUX.2-devOpen32B500.8787.5788.8288.100.893****0.5510.5160.9630.757FLUX.2-Klein-Base-4BOpen4B500.7883.0279.9480.010.6560.4850.3660.5540.071FLUX.2-Klein-Base-9BOpen9B500.8385.2981.4784.520.6550.5440.4000.8720.227FLUX.2-Klein-4BOpen4B40.8385.5378.9179.040.6280.5000.3640.6490.068FLUX.2-Klein-9BOpen9B40.8686.2085.2284.130.4240.5380.4060.8720.226Qwen-ImageOpen20B500.8788.3286.1486.830.8290.5390.5480.9430.946JoyAI-ImageOpen16B50–88.05––0.8740.5420.5210.963****0.963HunyuanImage-3.0Open80B500.7286.10––0.765––––LongCat-ImageOpen6B500.8786.8080.9381.300.8660.5160.5180.8850.956Z-Image-BaseOpen6B500.8488.1480.2083.040.8670.5460.5350.9350.936Z-Image-TurboOpen6B80.8284.8677.7380.050.8590.5280.5070.9170.926Mage-Flow-Base★Open4B300.7986.2682.5083.190.8510.5420.5090.9040.792Mage-Flow★Open4B200.9086.4982.1984.700.8870.5360.5050.9440.823Mage-Flow-Turbo★Open4B40.8885.4883.5884.160.8730.5230.4910.9110.801
Image editing— ImgEdit-Bench (0–5), GEdit-Bench EN/CN (0–10), TextEdit-Bench synthetic/real (0–25). Higher is better; theTypecolumn marks closed- vs open-source;bold/underline= best / second-best amongopen-sourcemodels;–= not reported; ★ = ours.
ModelType#ParamsStepsImgEditGEdit-ENGEdit-CNTextEdit-SynTextEdit-RealNano-BananaClosed––4.297.2917.39916.5418.22Seedream 4.0Closed––4.307.7017.69214.9018.54Seedream 4.5Closed––4.327.8207.800––Nano-Banana-ProClosed––4.377.7387.799––Step1X-Edit-v1.2Open19B503.957.4807.4679.2612.02FLUX.1-Kontext-devOpen12B283.716.4621.85712.1414.31FLUX.2-devOpen32B504.357.4137.27811.8614.71FLUX.2-Klein-Base-4BOpen4B503.807.0817.10211.0113.79FLUX.2-Klein-4BOpen4B44.017.7177.75011.8414.46FLUX.2-Klein-Base-9BOpen9B504.057.7407.74512.7615.65FLUX.2-Klein-9BOpen9B44.188.0408.05512.7315.75Z-Image-EditOpen6B504.307.5707.540––Qwen-Image-Edit-2509Open20B504.317.4807.46713.4015.81Qwen-Image-Edit-2511Open20B504.517.8777.81913.5316.81LongCat-Image-EditOpen6B504.457.7487.73112.4614.89FireRed-Image-Edit-1.0Open20B504.567.9437.88715.19****17.23JoyAI-Image-EditOpen16B504.468.2768.12514.8017.23****Mage-Flow-Edit-Base★Open4B304.287.8607.97013.6315.57Mage-Flow-Edit★Open4B304.348.1278.12314.1416.26Mage-Flow-Edit-Turbo★Open4B44.388.2718.26412.7715.41
https://huggingface.co/microsoft/Mage-Flow-Edit-Turbo#%F0%9F%8F%97%EF%B8%8F-architecture🏗️ Architecture
Mage-VAE— a latent tokenizer built as asymmetricone-step diffusion codec: the decoder is a fully-convolutional one-step pixel-diffusion model (no global-attention blocks), and the encoder is its architectural dual (a one-step latent generator conditioned on pixels). A standard Gaussian-prior KL is replaced with ananchor-latent KLthat regularizes the posterior toward FLUX.2-VAE latents, giving a generation-ready128-channel,16×-downsampled latent space.
Mage-VAE — anchor VAE (FLUX.2-VAE), the symmetric one-step encoder/decoder architecture, and the three-stage training pipeline.
Mage-Flow— a 4B Multimodal DiT that encodes prompts withQwen3-VLand images with Mage-VAE, then processespackedvariable-length image+text sequences with per-sample 2D rotary embeddings and joint self-attention.Native-resolution packingremoves bucket quantization and padding, lets one checkpoint generalize to any output size, and fuses the CFG cond/uncond branches into a single forward.
Mage-Flow — native-resolution packing of variable-length image+text tokens through the Native-Resolution MMDiT (left), and the dual-stream MMDiT block (right).
Post-training— fromBase, generation is aligned withDiffusion-NFT(prompt following, aesthetics, text rendering, preference) to produce the RL model, and distilled withdecoupled-DMD + adversarial perceptual guidanceinto the 4-stepTurbo. Editing models reuse the recipe, trained on a mixture of generation and editing data to keep the generative prior.
https://huggingface.co/microsoft/Mage-Flow-Edit-Turbo#%F0%9F%9A%80-quick-start🚀 Quick Start
https://huggingface.co/microsoft/Mage-Flow-Edit-Turbo#installationInstallation
Install everythingexceptflash\-attnfirst, then installflash\-attnseparately with build isolationoff— it compiles a CUDA extension against your installed torch, so torch and a matching CUDA toolkit must already be present.
cd Mage/mage_flow
uv venv && source .venv/bin/activate
# 1) Pinned, tested dependency set (torch 2.13, transformers 5.5, diffusers 0.38, pillow 12.3, …).
# Recommended for reproducibility. `uv pip install -e .` also works, but its loose
# bounds may resolve to a newer torch/transformers than the code was tested against.
uv pip install -r requirements.txt
uv pip install -e . --no-deps # the mage-flow package itself
# 2) flash-attn — needs build tools present and a CUDA toolkit whose MAJOR version
# matches your torch build (e.g. torch cu12x ↔ nvcc 12.x). A cu13/nvcc-12 mix fails.
uv pip install setuptools wheel ninja
uv pip install --no-build-isolation flash-attn==2.8.3
Plainpipis equivalent (pip install \-r requirements\.txt,pip install \-e \. \-\-no\-deps, then the two flash-attn lines). This registers three commands:mage\-flow,mage\-flow\-edit,mage\-flow\-app.
**torch / CUDA:**the default PyPI torch wheel targets the newest CUDA (currently cu13x). If your machine’s CUDA toolkit is 12.x, install torch from the matching index first, e.g.
uv pip install torch==2\.13\.0 torchvision==0\.28\.0 \-\-index\-url https://download\.pytorch\.org/whl/cu126, otherwise the flash-attn build will fail with a CUDA-version mismatch. (torch 2.13.0 ships cu126/cu129/cu130 wheels — pick the one matching yournvcc.)
https://huggingface.co/microsoft/Mage-Flow-Edit-Turbo#python-apiPython API
pipe\.generate\(prompts, \*\*kw\)andpipe\.edit\(prompts, ref\_images, \*\*kw\)return alist\[PIL\.Image\]aligned withprompts. Apromptslistis batched into one packed forward per denoise step (each sample can have its own resolution/seed).
Text-to-image:
from mage_flow import MageFlowPipeline
pipe = MageFlowPipeline.from_pretrained("microsoft/Mage-Flow", device="cuda")
# 1) single image
img = pipe.generate(["A close-up portrait of an elderly African man with deep wrinkles, wearing a traditional hat, soft natural lighting, ultra realistic."],
steps=20, cfg=5.0, heights=[1024], widths=[1024])[0]
img.save("t2i.png")
# 2) batch: several prompts / resolutions / seeds in ONE packed forward per step
imgs = pipe.generate(
["the Salar de Uyuni mirror surface captured at high noon, with intimate stillness permeating the air. dew beads on every blade of grass. National Geographic editorial, cinematic depth, fine-grained natural texture.",
"A close-up portrait of an elderly African man with deep wrinkles, wearing a traditional hat, soft natural lighting, ultra realistic.",
"An immersive close-up of a steaming bowl of Sichuan mapo tofu over jasmine rice served on a hand-thrown ceramic plate, finished with a wedge of citrus. Surface oils catch a tiny specular highlight. Shot with a Hasselblad H6D-100c, ambient window light, the kind of image that makes the viewer hungry."],
heights=[512, 1024, 1792], widths=[2048, 1024, 1024], # per-sample; 4:1 is fine
seeds=[1, 2, 3], steps=20, cfg=5.0,
)
Image editing:
from mage_flow import MageFlowPipeline
pipe = MageFlowPipeline.from_pretrained("microsoft/Mage-Flow-Edit-Turbo", device="cuda")
# single reference (path or PIL image)
img = pipe.edit(["Replace the background with a field of sunflowers"], ["assets/dog.jpg"],
steps=4, cfg=1.0, max_size=1024)[0]
img.save("single_edit.png")
# multi-image edit — ref_images[i] is a LIST of source images
img = pipe.edit(["blend the object from image 2 into image 1"],
[["scene.png", "object.png"]], steps=4, cfg=1.0)[0]
img.save("multi_edit.png")
# explicit output size (overrides max_size); Turbo edit = 4 steps / cfg 1
img = pipe.edit(["Replace the background with a field of sunflowers"], ["assets/dog.jpg"],
heights=[1024], widths=[1024], steps=4, cfg=1.0)[0]
img.save("single_edit_1024x1024.png")
Parameters(shared bygenerate/edit):
ParameterDefaultDescriptionprompts—string or list of strings; a list is batched into one forward per stepref\_images(edit)—per prompt: one image/path, or alistof images for multi-image editsteps``30denoising steps — Base30, RL20, Turbo4``cfg``5\.0classifier-free guidance scale (Turbo:1\.0)heights,widths``\[1024\]per-sample output size, multiple of 16; native resolution512–2048``max\_size(edit)source sizelongest output side; short side follows the reference’s aspect ratiovl\_cond\_long\_edge(edit)384cap the long edge of the reference image fed to theVL text encoder(matches training preprocessing; the VAE/generation path keeps the full output resolution).0/Nonedisablesneg\_prompts``" "per-sample negative prompt (applied whencfg \> 1)seeds``42per-sample seed;\-1= randombatch\_cfg``Truefuse the CFG conditional + unconditional passes into one packed forwardrenormalization``Falserescale guided velocity per token (reduces over-saturation at high cfg)static\_shift``6\.0override the flow-matching sigma shiftprompt\_template``mage\-flow/mage\-flow\-edittext-encoder prompt template
https://huggingface.co/microsoft/Mage-Flow-Edit-Turbo#cliCLI
# text-to-image (two prompts in one batch)
mage-flow --prompt "A close-up portrait of an elderly African man with deep wrinkles, wearing a traditional hat, soft natural lighting, ultra realistic." "An immersive landscape of a Greenlandic icefjord at midnight sun, painted by early dawn light, crystal-clear skies adding drama. fine grains of sand carving sharp shadows. Peter Lik gallery print, moody atmosphere, museum-grade composition." \
--model_path microsoft/Mage-Flow --steps 20 --cfg 5.0 \
--height 1024 512 --width 1024 2048 --seed 42 --out ./outputs
# editing (one --ref per prompt; comma-separate sources for multi-image edit)
mage-flow-edit --prompt "Replace the background with a field of sunflowers" "blend these two images" \
--ref assets/dog.jpg "scene.png,object.png" \
--model_path microsoft/Mage-Flow-Edit --max_size 1024 --out ./outputs
FlagScopeMeaning\-\-promptbothone or more prompts, run as a batch (sampleiuses\-\-seed \+ i)\-\-model\_pathbothlocal repo dir or HF Hub repo id (auto-downloaded + cached)\-\-stepsbothnumber of denoising steps\-\-cfgbothclassifier-free guidance scale\-\-heightbothoutput height — one value, or one per prompt for mixed resolutions\-\-widthbothoutput width — one value, or one per prompt for mixed resolutions\-\-seedbothbase seed (sampleiuses\-\-seed \+ i)\-\-neg\_promptbothnegative prompt\-\-static\_shiftbothoverride the flow-matching sigma shift\-\-outbothoutput directory\-\-refeditreference image per prompt (comma-separate paths for a multi-image edit)\-\-max\_sizeeditmax size of the reference image\-\-vl\_cond\_long\_edgeeditVL-condition long edge (default384)
https://huggingface.co/microsoft/Mage-Flow-Edit-Turbo#gradio-appGradio app
mage-flow-app # serve on http://0.0.0.0:7860 (or: python -m mage_flow.app)
A web UI withText → ImageandImage Edittabs; models load lazily on first use and are cached. Presets default to themicrosoft/Mage\-Flow\*Hugging Face repos(downloaded + cached on first use); setMAGEFLOW\_HF\_DIRto load local checkpoint dirs instead.
Launch options:
FlagDefaultMeaning\-\-host``0\.0\.0\.0bind address\-\-port``7860port\-\-device``cudainference device\-\-shareoffcreate a public Gradio share link\-\-preload*(lazy)*comma-separated repo ids / paths to load at startup instead of on first use
https://huggingface.co/microsoft/Mage-Flow-Edit-Turbo#%F0%9F%93%9D-citation📝 Citation
@article{zhang2026mageflow,
title={Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing},
author={Zhang, Xinjie and Zhang, Peng and Zheng, Shicheng and Guo, Jinghao and Jia, Zhaoyang and Shen, Yifei and Guo, Xun and Luo, Yuxuan and Li, Jiahao and Xie, Wenxuan and Pu, Fanyi and Zhang, Xiaoyi and Zhang, Kaichen and Guo, Zongyu and Bi, Tianci and Gui, Dongnan and Liu, Zhening and Wen, Zimo and Zheng, Zihan and Yang, Senqiao and Li, Xiao and Wang, Jinglu and Li, Bin and Lu, Yan},
journal={arXiv preprint arXiv:2607.19064},
year={2026}
}
Similar Articles
microsoft/Mage-Flow
Microsoft releases Mage-Flow, a compact 4B-parameter foundation model for efficient native-resolution text-to-image generation and instruction-based image editing, achieving competitive quality against much larger models.
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
Mage-Flow is a compact 4B-parameter generative stack for efficient text-to-image generation and instruction-based image editing, featuring a co-designed lightweight tokenizer (Mage-VAE) and a native-resolution multimodal diffusion transformer trained with rectified flow matching. It achieves competitive performance while enabling high-resolution generation at 0.59s on a single A100 GPU.
Mage (GitHub Repo)
Microsoft releases Mage, a family of lightweight 4B-parameter multimodal models for visual understanding and generation, including Mage-VL for image/video understanding and Mage-Flow for text-to-image generation and editing, designed for research and deployment on modest hardware.
Mage (2 minute read)
Microsoft unveils Mage, a family of compact 4B-parameter multimodal models for visual understanding and generation, designed for easy training and deployment on modest hardware while staying competitive with larger open models.
microsoft/Lens-Turbo
Microsoft releases Lens, a 3.8B-parameter foundational text-to-image model with efficient training and fast high-resolution generation, featuring dense-caption pre-training and mixed-resolution learning.