Mage-Flow Base
'The undistilled text-to-image base of Microsoft’s Mage-Flow: a 4.1B dual-stream NR-MMDiT on Qwen3-VL-4B text features, working in the 128-channel, 16× latent space of Mage-VAE with one token per latent pixel. Built to generate at native resolution from 512 to 2048 px.'
- org
- Microsoft
- modality
- image
- tasks
- text-to-image
- license
- MIT
- native output
- 512 to 2048 px, any aspect ratio up to 4:1
- total params
- 8.73B
- model.arch
- mageflow
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Mage-Flow 4B Base (NR-MMDiT) MageFlow (ai-toolkit mageflow/src/transformer.py) Same sha256 as the file in microsoft/Mage-Flow-Base, which now returns 404 on the Hub. The community mirror is listed because it is the one that still downloads. | 4.12B 4,115,745,408 | 8.23 GB | bf16 | yes |
| Text encoder | Qwen3-VL-4B-Instruct transformers.Qwen3VLForConditionalGeneration Byte-identical to Qwen/Qwen3-VL-4B-Instruct (same sha256 for both shards). Includes the vision tower, which ai-toolkit drops for text-to-image. | 4.44B 4,437,815,808 | 8.88 GB | bf16 | no |
| Tokenizer | Qwen3-VL tokenizer transformers.AutoTokenizer | — | — | — | — |
| VAE | Mage-VAE MageVAE (ai-toolkit mageflow/src/vae.py) A convolutional encoder plus a one-step diffusion decoder (Mage calls it a one-step diffusion codec). The same file ships in every Mage-Flow repo. | 172.5M 172,478,148 | 345 MB | bf16 | no |
| total | 8.73B | 17.45 GB | |||
Latent space
- autoencoder
- Mage-VAE
- pixels per token
- 16×16
- notes
- No patchify: each latent pixel is one transformer token, so 16×16 pixels per token as with an 8× VAE and 2×2 patches, but with 128 channels per token. Latents are used raw, with no scale or shift.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 512×512 | 128×1×32×32 | 1,024 | smallest native size |
| 1024×1024 | 128×1×64×64 | 4,096 | ai-toolkit sample default |
| 2048×512 | 128×1×32×128 | 4,096 | 4:1 |
| 2048×2048 | 128×1×128×128 | 16,384 | largest native size |
Architecture
- Type
- Dual-stream MMDiT (separate text and image weights, joint attention)
- Blocks
- 12 double-stream, no single-stream blocks
- Hidden size
- 3072 (24 heads × 128)
- FFN size
- 12288 (GELU, mlp_ratio 4)
- Text conditioning
- Final Qwen3-VL hidden states (2560-d) with the 34-token system template dropped; no pooled vector
- Sequence packing
- Variable-length samples packed into one sequence, varlen attention
- Norm / position
- QK RMSNorm, 2D multi-scale RoPE on image tokens only (16/56/56 dims, θ 10000); text tokens unrotated
- Objective
- Rectified flow (target noise − clean), static shift 6.0
- Recommended sampling
- 30 steps (Mage-Flow README)
In AI Toolkit
- model.arch
- mageflow
- UI label
- Mage-Flow (image)
- model.name_or_path
- microsoft/Mage-Flow-Base
- source
- extensions_built_in/diffusion_models/mageflow/mageflow.py
- extensions_built_in/diffusion_models/mageflow/src/transformer.py
- extensions_built_in/diffusion_models/mageflow/src/pipeline.py
- extensions_built_in/diffusion_models/mageflow/src/text_encoder.py
- extensions_built_in/diffusion_models/mageflow/src/vae.py
- extensions_built_in/diffusion_models/mageflow/src/attn.py
- toolkit/models/v2/text_encoders/qwen3_vl.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- model.low_vram, model.layer_offloading
UI defaults
- quantize / quantize_te
- true / true (qfloat8)
- low_vram
- true
- timestep_type
- linear
- sample guidance_scale / sample_steps
- 4 / 25
- network.conv
- disabled (linear LoRA only)
Specifics
- Default repo is gone
- The UI default microsoft/Mage-Flow-Base returns 404 on the Hub (checked 2026-09-23), so it only loads from an existing local cache. Set name_or_path to mage-flow-community/Mage-Flow-Base (same sha256 for every weight file) or a local copy.
- Loading
- Reads transformer/config.json and transformer/diffusion_pytorch_model.safetensors, the text encoder from text_encoder/ and the VAE from vae/, all inside name_or_path. model_kwargs.text_encoder_path and vae_path override the last two.
- Text encoding
- Each prompt is wrapped in Mage’s system template and encoded at its natural length, up to 2048 tokens (model_kwargs.max_text_length). No padding: samples are packed.
- VAE sampling
- Training latents are sampled from the posterior (model_kwargs.vae_sample_posterior, default true); the repo’s vae/config.json says sample_posterior false.
- Resolution
- Buckets snap to multiples of 16 (16× VAE, patch 1).
- Timesteps
- flowmatch with a static shift of 6.0, no resolution-dependent shift.
- Attention
- Uses flash-attn varlen when installed, otherwise a per-sample SDPA loop (same result, slower).
- Quantization
- img_in, txt_in, txt_norm, time_text_embed*, norm_out* and proj_out stay in full precision.
- LoRA target
- MageFlow
- Saving
- LoRA keys are saved with the diffusion_model. prefix (ComfyUI layout). Full fine-tunes save a single safetensors of the MageFlow state dict.
- Sampling
- Built-in Euler sampler over Mage’s shifted sigma schedule (model_kwargs.static_shift, default 6.0). CFG turns on above guidance_scale 1.
- Metadata base version
- mageflow
Example config
not verified
job: extensionconfig: name: "my_mageflow_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: bf16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "linear" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "mage-flow-community/Mage-Flow-Base" arch: "mageflow" quantize: true quantize_te: true low_vram: true sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 4 sample_steps: 25 prompts: - "a red fox sitting in fresh snow, golden hour"Links
Krea 2 Turbo
'The post-trained, step-distilled release of Krea 2: the same 12.8B single-stream MMDiT as Raw, sampled in about 8 steps without CFG. ai-toolkit trains it through a de-distill training adapter so LoRAs keep the fast sampling.'
Boogu-Image 0.1 Base
'The undistilled text-to-image base of Boogu-Image 0.1: a 10.3B Lumina2-style DiT with 8 double-stream and 32 single-stream layers, conditioned on Qwen3-VL-8B and decoding through the FLUX.1 VAE. Its authors pitch it for fine-tuning and dense Chinese and English text rendering.'