Mage-Flow / README.md
Xinjie-Q's picture
add README
2a94cad verified
|
Raw
History Blame
33.2 kB
metadata
license: mit
library_name: diffusers
pipeline_tag: text-to-image
tags:
  - text-to-image
  - image-generation
  - image-editing
  - diffusion
  - rectified-flow
  - mage-flow

Mage-Flow
An Efficient Native-Resolution Foundation Model for Image Generation and Editing

arXiv Project Page GitHub Hugging Face Hugging Face Hugging Face Hugging Face Hugging Face Hugging Face License: MIT

gallery

Mage-Flow is a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. Instead of scaling to tens of billions of parameters, Mage-Flow reaches state-of-the-art-competitive quality through careful tokenizer–backbone–system co-design, so it stays fast, memory-light, and easy to fine-tune under realistic compute budgets.

The stack is built from two shared, co-designed components:

  • Mage-VAE β€” a lightweight, high-fidelity latent tokenizer (one-step diffusion encode/decode with anchor-latent KL regularization).
  • NR-MMDiT β€” a shared 4B Native-Resolution Multimodal Diffusion Transformer, trained with rectified flow matching in the Mage-VAE latent space.

Together with native-resolution packing and a fused-kernel training infrastructure, this shared stack powers two model instantiations: Mage-Flow for text-to-image generation and Mage-Flow-Edit for instruction-based image editing. Each ships in Base, RL-aligned, and 4-step Turbo variants.

✨ Highlights

  • Compact & competitive. A single 4B family for generation and editing that matches or beats much larger open systems (Qwen-Image 20B, Z-Image 6B, FLUX.2 32B, FireRed-Image-Edit 20B).
  • Efficient tokenizer. Mage-VAE matches FLUX.2-VAE reconstruction fidelity while using ~12Γ— / ~22Γ— fewer encode / decode MACs per pixel, removing the VAE as the high-resolution bottleneck.
  • Native resolution. One checkpoint generates from 512 to 2048 on any aspect ratio, including extreme 4:1 (e.g. 512Γ—2048, 2048Γ—512).
  • System-level speed. Native-resolution packing (FlashAttention var-len + per-sample 2D RoPE) + fused CUDA kernels raise MFU from ~33% β†’ ~77% (~2.5Γ— faster training); CFG's conditional/unconditional branches run in one packed forward.
  • Full family. Base, RL-aligned, and 4-step Turbo variants for both generation and editing.
  • Versatile editing. Mage-Flow-Edit supports semantic content editing, appearance transformation, image restoration, and structure-aware outputs within a unified image-and-text-conditioned model. See the report's editing galleries.
  • Interactive latency. At 1024Β² on a single A100: Mage-Flow-Turbo 0.59 s/image, Mage-Flow-Edit-Turbo 1.02 s/edit, peak memory ~18–20 GB (lowest among compared systems).
One-to-many editing diversity
One-to-many editing diversity β€” Mage-Flow-Edit can generate diverse outputs from a single reference image.

πŸ“₯ Model Zoo

Each checkpoint is a self-contained diffusers-style repo (transformer/ + shared vae/, text_encoder/, scheduler/).

Model Task Variant Steps Hugging Face
Mage-Flow-4B-Base textβ†’image Base 30 πŸ€— microsoft/Mage-Flow-Base
Mage-Flow-4B textβ†’image RL-aligned 20 πŸ€— microsoft/Mage-Flow
Mage-Flow-4B-Turbo textβ†’image Few-step distilled 4 πŸ€— microsoft/Mage-Flow-Turbo
Mage-Flow-Edit-4B-Base editing Base 30 πŸ€— microsoft/Mage-Flow-Edit-Base
Mage-Flow-Edit-4B editing RL-aligned 30 πŸ€— microsoft/Mage-Flow-Edit
Mage-Flow-Edit-4B-Turbo editing Few-step distilled 4 πŸ€— microsoft/Mage-Flow-Edit-Turbo

πŸ–ΌοΈ Showcase

Text-to-image β€” prompt following, fine detail, and legible English/Chinese text rendering. (The first panel is open; click a title to expand the others.)

Showcase
t2i teaser
General scenes
general scenes
Portraits
portraits
Cuisine & still life
cuisine
English text rendering
english text
Chinese text rendering
chinese text

Instruction-based editing β€” appearance, content, scene/subject, human-centered & creative, low-level, and restoration edits (source β†’ result). (The first panel is open; click a title to expand the others.)

Various Editing I
editing showcase 1
Various Editing II
editing showcase 2
Localized content & object editing
content editing
Scene, subject & camera transformations
scene and subject
Appearance & artistic rendering
appearance
Human-centered & creative editing
human-centered and creative
Low-level vision & conditional reconstruction
low-level vision
Bidirectional degradation & restoration
restoration

πŸ“Š Performance

Full benchmark tables (text-to-image & image editing) β€” click to expand

Text-to-image β€” full benchmark suite: GenEval, DPG-Bench, TIIF-Bench (short/long splits), CVTG-2K, OneIG (EN/CN), LongText (EN/CN). Higher is better; GenEval / CVTG-2K / OneIG / LongText on a 0–1 scale, DPG / TIIF on 0–100. The Type column marks closed- vs open-source; bold / underline = best / second-best among open-source models (closed-source shown for reference, not ranked); – = not reported; β˜… = ours.

Model Type #Params Steps GenEval DPG TIIF-Short TIIF-Long CVTG-2K OneIG-EN OneIG-CN LongText-EN LongText-CN
Seedream 3.0 Closed – – 0.84 88.27 86.02 84.31 0.592 0.530 0.528 0.896 0.878
Seedream 4.0 Closed – – 0.84 88.63 – – 0.892 0.573 0.554 0.936 0.946
GPT-Image-1 Closed – – 0.84 85.15 89.15 88.29 0.857 0.533 0.474 0.956 0.619
Nano-Banana-Pro Closed – – 0.83 87.16 – – 0.779 0.580 0.570 0.981 0.949
FLUX.1-dev Open 12B 50 0.66 83.84 71.09 71.78 0.496 0.434 0.245 0.607 0.005
FLUX.1-Krea-dev Open 12B 50 0.72 86.59 80.36 81.67 0.444 0.443 0.271 0.693 0.002
FLUX.2-dev Open 32B 50 0.87 87.57 88.82 88.10 0.893 0.551 0.516 0.963 0.757
FLUX.2-Klein-Base-4B Open 4B 50 0.78 83.02 79.94 80.01 0.656 0.485 0.366 0.554 0.071
FLUX.2-Klein-Base-9B Open 9B 50 0.83 85.29 81.47 84.52 0.655 0.544 0.400 0.872 0.227
FLUX.2-Klein-4B Open 4B 4 0.83 85.53 78.91 79.04 0.628 0.500 0.364 0.649 0.068
FLUX.2-Klein-9B Open 9B 4 0.86 86.20 85.22 84.13 0.424 0.538 0.406 0.872 0.226
Qwen-Image Open 20B 50 0.87 88.32 86.14 86.83 0.829 0.539 0.548 0.943 0.946
JoyAI-Image Open 16B 50 – 88.05 – – 0.874 0.542 0.521 0.963 0.963
HunyuanImage-3.0 Open 80B 50 0.72 86.10 – – 0.765 – – – –
LongCat-Image Open 6B 50 0.87 86.80 80.93 81.30 0.866 0.516 0.518 0.885 0.956
Z-Image-Base Open 6B 50 0.84 88.14 80.20 83.04 0.867 0.546 0.535 0.935 0.936
Z-Image-Turbo Open 6B 8 0.82 84.86 77.73 80.05 0.859 0.528 0.507 0.917 0.926
Mage-Flow-Base β˜… Open 4B 30 0.79 86.26 82.50 83.19 0.851 0.542 0.509 0.904 0.792
Mage-Flow β˜… Open 4B 20 0.90 86.49 82.19 84.70 0.887 0.536 0.505 0.944 0.823
Mage-Flow-Turbo β˜… Open 4B 4 0.88 85.48 83.58 84.16 0.873 0.523 0.491 0.911 0.801

Image editing β€” ImgEdit-Bench (0–5), GEdit-Bench EN/CN (0–10), TextEdit-Bench synthetic/real (0–25). Higher is better; the Type column marks closed- vs open-source; bold / underline = best / second-best among open-source models; – = not reported; β˜… = ours.

Model Type #Params Steps ImgEdit GEdit-EN GEdit-CN TextEdit-Syn TextEdit-Real
Nano-Banana Closed – – 4.29 7.291 7.399 16.54 18.22
Seedream 4.0 Closed – – 4.30 7.701 7.692 14.90 18.54
Seedream 4.5 Closed – – 4.32 7.820 7.800 – –
Nano-Banana-Pro Closed – – 4.37 7.738 7.799 – –
Step1X-Edit-v1.2 Open 19B 50 3.95 7.480 7.467 9.26 12.02
FLUX.1-Kontext-dev Open 12B 28 3.71 6.462 1.857 12.14 14.31
FLUX.2-dev Open 32B 50 4.35 7.413 7.278 11.86 14.71
FLUX.2-Klein-Base-4B Open 4B 50 3.80 7.081 7.102 11.01 13.79
FLUX.2-Klein-4B Open 4B 4 4.01 7.717 7.750 11.84 14.46
FLUX.2-Klein-Base-9B Open 9B 50 4.05 7.740 7.745 12.76 15.65
FLUX.2-Klein-9B Open 9B 4 4.18 8.040 8.055 12.73 15.75
Z-Image-Edit Open 6B 50 4.30 7.570 7.540 – –
Qwen-Image-Edit-2509 Open 20B 50 4.31 7.480 7.467 13.40 15.81
Qwen-Image-Edit-2511 Open 20B 50 4.51 7.877 7.819 13.53 16.81
LongCat-Image-Edit Open 6B 50 4.45 7.748 7.731 12.46 14.89
FireRed-Image-Edit-1.0 Open 20B 50 4.56 7.943 7.887 15.19 17.23
JoyAI-Image-Edit Open 16B 50 4.46 8.276 8.125 14.80 17.23
Mage-Flow-Edit-Base β˜… Open 4B 30 4.28 7.860 7.970 13.63 15.57
Mage-Flow-Edit β˜… Open 4B 30 4.34 8.127 8.123 14.14 16.26
Mage-Flow-Edit-Turbo β˜… Open 4B 4 4.38 8.271 8.264 12.77 15.41

πŸ—οΈ Architecture

Mage-VAE β€” a latent tokenizer built as a symmetric one-step diffusion codec: the decoder is a fully-convolutional one-step pixel-diffusion model (no global-attention blocks), and the encoder is its architectural dual (a one-step latent generator conditioned on pixels). A standard Gaussian-prior KL is replaced with an anchor-latent KL that regularizes the posterior toward FLUX.2-VAE latents, giving a generation-ready 128-channel, 16Γ—-downsampled latent space.

Mage-VAE architecture and training
Mage-VAE β€” anchor VAE (FLUX.2-VAE), the symmetric one-step encoder/decoder architecture, and the three-stage training pipeline.

Mage-Flow β€” a 4B Multimodal DiT that encodes prompts with Qwen3-VL and images with Mage-VAE, then processes packed variable-length image+text sequences with per-sample 2D rotary embeddings and joint self-attention. Native-resolution packing removes bucket quantization and padding, lets one checkpoint generalize to any output size, and fuses the CFG cond/uncond branches into a single forward.

Mage-Flow / Native-Resolution MMDiT architecture
Mage-Flow β€” native-resolution packing of variable-length image+text tokens through the Native-Resolution MMDiT (left), and the dual-stream MMDiT block (right).

Post-training β€” from Base, generation is aligned with Diffusion-NFT (prompt following, aesthetics, text rendering, preference) to produce the RL model, and distilled with decoupled-DMD + adversarial perceptual guidance into the 4-step Turbo. Editing models reuse the recipe, trained on a mixture of generation and editing data to keep the generative prior.

πŸš€ Quick Start

Installation

Install everything except flash-attn first, then install flash-attn separately with build isolation off β€” it compiles a CUDA extension against your installed torch, so torch and a matching CUDA toolkit must already be present.

cd Mage/mage_flow
uv venv && source .venv/bin/activate

# 1) Pinned, tested dependency set (torch 2.10, transformers 5.5, diffusers 0.38, pillow 12.3, …).
#    Recommended for reproducibility. `uv pip install -e .` also works, but its loose
#    bounds may resolve to a newer torch/transformers than the code was tested against.
uv pip install -r requirements.txt
uv pip install -e . --no-deps           # the mage-flow package itself

# 2) flash-attn β€” needs build tools present and a CUDA toolkit whose MAJOR version
#    matches your torch build (e.g. torch cu12x ↔ nvcc 12.x). A cu13/nvcc-12 mix fails.
uv pip install setuptools wheel ninja
uv pip install --no-build-isolation flash-attn==2.8.3

Plain pip is equivalent (pip install -r requirements.txt, pip install -e . --no-deps, then the two flash-attn lines). This registers three commands: mage-flow, mage-flow-edit, mage-flow-app.

torch / CUDA: the default PyPI torch wheel targets the newest CUDA (currently cu13x). If your machine's CUDA toolkit is 12.x, install torch from the matching index first, e.g. uv pip install torch==2.10.0 torchvision==0.25.0 --index-url https://download.pytorch.org/whl/cu128, otherwise the flash-attn build will fail with a CUDA-version mismatch.

Python API

pipe.generate(prompts, **kw) and pipe.edit(prompts, ref_images, **kw) return a list[PIL.Image] aligned with prompts. A prompts list is batched into one packed forward per denoise step (each sample can have its own resolution/seed).

Text-to-image:

from mage_flow import MageFlowPipeline

pipe = MageFlowPipeline.from_pretrained("microsoft/Mage-Flow", device="cuda")

# 1) single image
img = pipe.generate(["A close-up portrait of an elderly African man with deep wrinkles, wearing a traditional hat, soft natural lighting, ultra realistic."],
                    steps=20, cfg=5.0, heights=[1024], widths=[1024])[0]
img.save("t2i.png")

# 2) batch: several prompts / resolutions / seeds in ONE packed forward per step
imgs = pipe.generate(
    ["the Salar de Uyuni mirror surface captured at high noon, with intimate stillness permeating the air. dew beads on every blade of grass. National Geographic editorial, cinematic depth, fine-grained natural texture.", 
     "A close-up portrait of an elderly African man with deep wrinkles, wearing a traditional hat, soft natural lighting, ultra realistic.", 
     "An immersive close-up of a steaming bowl of Sichuan mapo tofu over jasmine rice served on a hand-thrown ceramic plate, finished with a wedge of citrus. Surface oils catch a tiny specular highlight. Shot with a Hasselblad H6D-100c, ambient window light, the kind of image that makes the viewer hungry."],
    heights=[512, 1024, 1792], widths=[2048, 1024, 1024],   # per-sample; 4:1 is fine
    seeds=[1, 2, 3], steps=20, cfg=5.0,
)

Image editing:

pipe = MageFlowPipeline.from_pretrained("microsoft/Mage-Flow-Edit", device="cuda")

# single reference (path or PIL image)
img = pipe.edit(["Replace the background with a field of sunflowers"], ["assets/dog.jpg"],
                steps=30, cfg=5.0, max_size=1024)[0]
img.save("single_edit.png")

# multi-image edit β€” ref_images[i] is a LIST of source images
img = pipe.edit(["blend the object from image 2 into image 1"],
                [["scene.png", "object.png"]], steps=30, cfg=5.0)[0]
img.save("multi_edit.png")

# explicit output size (overrides max_size); Turbo edit = 4 steps / cfg 1
img = pipe.edit(["Replace the background with a field of sunflowers"], ["assets/dog.jpg"],
                heights=[1024], widths=[1024], steps=4, cfg=1.0)[0]
img.save("single_edit_1024x1024.png")

Parameters (shared by generate / edit):

Parameter Default Description
prompts β€” string or list of strings; a list is batched into one forward per step
ref_images (edit) β€” per prompt: one image/path, or alist of images for multi-image edit
steps 30 denoising steps β€” Base30, RL 20, Turbo 4
cfg 5.0 classifier-free guidance scale (Turbo:1.0)
heights, widths [1024] per-sample output size, multiple of 16; native resolution512–2048
max_size (edit) source size longest output side; short side follows the reference's aspect ratio
vl_cond_long_edge (edit) 384 cap the long edge of the reference image fed to theVL text encoder (matches training preprocessing; the VAE/generation path keeps the full output resolution). 0/None disables
neg_prompts " " per-sample negative prompt (applied whencfg > 1)
seeds 42 per-sample seed;-1 = random
batch_cfg True fuse the CFG conditional + unconditional passes into one packed forward
renormalization False rescale guided velocity per token (reduces over-saturation at high cfg)
static_shift 6.0 override the flow-matching sigma shift
prompt_template mage-flow / mage-flow-edit text-encoder prompt template

CLI

# text-to-image (two prompts in one batch)
mage-flow --prompt "A close-up portrait of an elderly African man with deep wrinkles, wearing a traditional hat, soft natural lighting, ultra realistic." "An immersive landscape of a Greenlandic icefjord at midnight sun, painted by early dawn light, crystal-clear skies adding drama. fine grains of sand carving sharp shadows. Peter Lik gallery print, moody atmosphere, museum-grade composition." \
          --model_path microsoft/Mage-Flow --steps 20 --cfg 5.0 \
          --height 1024 512 --width 1024 2048 --seed 42 --out ./outputs


# editing (one --ref per prompt; comma-separate sources for multi-image edit)
mage-flow-edit --prompt "Replace the background with a field of sunflowers" "blend these two images" \
               --ref assets/dog.jpg "scene.png,object.png" \
               --model_path microsoft/Mage-Flow-Edit --max_size 1024 --out ./outputs
Flag Scope Meaning
--prompt both one or more prompts, run as a batch (sample i uses --seed + i)
--model_path both local repo dir or HF Hub repo id (auto-downloaded + cached)
--steps both number of denoising steps
--cfg both classifier-free guidance scale
--height both output height β€” one value, or one per prompt for mixed resolutions
--width both output width β€” one value, or one per prompt for mixed resolutions
--seed both base seed (sample i uses --seed + i)
--neg_prompt both negative prompt
--static_shift both override the flow-matching sigma shift
--out both output directory
--ref edit reference image per prompt (comma-separate paths for a multi-image edit)
--max_size edit max size of the reference image
--vl_cond_long_edge edit VL-condition long edge (default 384)

Gradio app

mage-flow-app                     # serve on http://0.0.0.0:7860  (or: python -m mage_flow.app)

A web UI with Text β†’ Image and Image Edit tabs; models load lazily on first use and are cached. Presets default to the microsoft/Mage-Flow* Hugging Face repos (downloaded + cached on first use); set MAGEFLOW_HF_DIR to load local checkpoint dirs instead.

Launch options:

Flag Default Meaning
--host 0.0.0.0 bind address
--port 7860 port
--device cuda inference device
--share off create a public Gradio share link
--preload (lazy) comma-separated repo ids / paths to load at startup instead of on first use

πŸ“ Citation

@techreport{mageflow2026,
  title  = {Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing},
  author = {Microsoft Mage Team},
  year   = {2026},
  institution = {Microsoft},
  url    = {https://github.com/microsoft/Mage}
}