Docs
AI ToolkitModels

Mage-Flow Edit Base

'The undistilled instruction-editing base of Microsoft’s Mage-Flow: the same 4.1B dual-stream NR-MMDiT and Mage-VAE as Mage-Flow Base, trained to edit. Reference images reach the model twice, through Qwen3-VL with the instruction and as clean latents packed after the target.'

org
Microsoft
modality
image
tasks
image-editing
license
MIT
native output
512 to 2048 px, any aspect ratio up to 4:1
total params
8.73B
model.arch
mageflow_edit

Components

rolemodelparamssizedtypetrained
Transformer
Mage-Flow Edit 4B Base (NR-MMDiT)
MageFlow (ai-toolkit mageflow/src/transformer.py)

Same sha256 as the file in microsoft/Mage-Flow-Edit-Base, which now returns 404 on the Hub. The community mirror is listed because it is the one that still downloads.

4.12B
4,115,745,408
8.23 GBbf16yes
Text encoder
Qwen3-VL-4B-Instruct
transformers.Qwen3VLForConditionalGeneration

Byte-identical to Qwen/Qwen3-VL-4B-Instruct (same sha256 for both shards). Includes the vision tower, which edit training uses to see the reference images.

4.44B
4,437,815,808
8.88 GBbf16no
Tokenizer
Qwen3-VL processor and tokenizer
transformers.AutoProcessor
————
VAE
Mage-VAE
MageVAE (ai-toolkit mageflow/src/vae.py)

A convolutional encoder plus a one-step diffusion decoder (Mage calls it a one-step diffusion codec). The same file ships in every Mage-Flow repo.

172.5M
172,478,148
345 MBbf16no
total8.73B17.45 GB

Latent space

spatial
16×
channels
128
patch
1×1
autoencoder
Mage-VAE
pixels per token
16×16
notes
No patchify: each latent pixel is one transformer token, so 16×16 pixels per token as with an 8× VAE and 2×2 patches, but with 128 channels per token. Latents are used raw, with no scale or shift. Reference latents add their own tokens on top of the counts below.
inputlatent (c×t×h×w)tokens
512×512128×1×32×321,024smallest native size
1024×1024128×1×64×644,096ai-toolkit sample default
2048×512128×1×32×1284,0964:1
2048×2048128×1×128×12816,384largest native size

Architecture

Type
Dual-stream MMDiT (separate text and image weights, joint attention)
Blocks
12 double-stream, no single-stream blocks
Hidden size
3072 (24 heads × 128)
FFN size
12288 (GELU, mlp_ratio 4)
Text conditioning
Final Qwen3-VL hidden states (2560-d) with the 64-token edit system template dropped; no pooled vector
Image conditioning
Reference images go through Qwen3-VL with the instruction, and as clean latents packed after the target tokens, sharing the sample’s timestep; each image gets its own RoPE frame index
Sequence packing
Variable-length samples packed into one sequence, varlen attention
Norm / position
QK RMSNorm, 2D multi-scale RoPE on image tokens only (16/56/56 dims, θ 10000); text tokens unrotated
Objective
Rectified flow (target noise − clean), static shift 6.0
Recommended sampling
30 steps (Mage-Flow README)

In AI Toolkit

model.arch
mageflow_edit
UI label
Mage-Flow Edit (instruction)
model.name_or_path
microsoft/Mage-Flow-Edit-Base
extra UI sections
datasets.multi_control_paths, sample.multi_ctrl_imgs, model.low_vram, model.layer_offloading

UI defaults

quantize / quantize_te
true / true (qfloat8)
low_vram
true
timestep_type
linear
sample guidance_scale / sample_steps
4 / 25
train.unload_text_encoder
false (section hidden)
network.conv
disabled (linear LoRA only)

Specifics

Default repo is gone
The UI default microsoft/Mage-Flow-Edit-Base returns 404 on the Hub (checked 2026-09-23), so it only loads from an existing local cache. Set name_or_path to mage-flow-community/Mage-Flow-Edit-Base (same sha256 for every weight file) or a local copy.
Training data
Each item is a target image, an edit instruction as its caption, and up to three reference images (control_path_1 to _3). Reference images keep their own size and aspect in the dataloader.
Reference path 1: Qwen3-VL
References are scaled so the long edge is at most 384 px (model_kwargs.vl_cond_long_edge, never upscaled) and placed as “Image N:” vision placeholders before the instruction. The UI keeps the text encoder loaded; cached text embeddings are keyed on the control paths.
Reference path 2: latents
Each reference is resized to the target’s pixel area keeping its own aspect ratio (Mage’s reference code stretches it to the exact target size instead), snapped to 16 px, VAE-encoded and packed after the target tokens. Only target tokens enter the loss.
Loading
Reads transformer/config.json and transformer/diffusion_pytorch_model.safetensors, the text encoder from text_encoder/ and the VAE from vae/, all inside name_or_path. model_kwargs.text_encoder_path and vae_path override the last two.
Text encoding
Each instruction is wrapped in Mage’s edit system template and encoded at its natural length, up to 2048 tokens (model_kwargs.max_text_length). No padding: samples are packed.
VAE sampling
Training latents are sampled from the posterior (model_kwargs.vae_sample_posterior, default true); the repo’s vae/config.json says sample_posterior false.
Resolution
Buckets snap to multiples of 16 (16× VAE, patch 1).
Timesteps
flowmatch with a static shift of 6.0, no resolution-dependent shift.
Attention
Uses flash-attn varlen when installed, otherwise a per-sample SDPA loop (same result, slower).
Quantization
img_in, txt_in, txt_norm, time_text_embed*, norm_out* and proj_out stay in full precision.
LoRA target
MageFlow
Saving
LoRA keys are saved with the diffusion_model. prefix (ComfyUI layout). Full fine-tunes save a single safetensors of the MageFlow state dict.
Sampling
Built-in Euler sampler over Mage’s shifted sigma schedule (model_kwargs.static_shift, default 6.0). CFG turns on above guidance_scale 1.
Metadata base version
mageflow_edit

Example config

not verified

job: extensionconfig:  name: "my_mageflow_edit_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: bf16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/targets"          control_path_1: "/path/to/references"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "linear"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "mage-flow-community/Mage-Flow-Edit-Base"        arch: "mageflow_edit"        quantize: true        quantize_te: true        low_vram: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 4        sample_steps: 25        samples:          - prompt: "turn the car red"            ctrl_img_1: "/path/to/reference.jpg"

On this page