Ming-Image 0.1 Design
A 6B Z-Image-style DiT for UI, infographics, posters and other text-heavy design, with RGBA output. It is conditioned by a Ling-mini-2.0 MoE multimodal LLM through two caption streams. ai-toolkit trains it from the int8 ComfyUI repack with a training adapter.
- org
- inclusionAI
- modality
- image
- tasks
- text-to-image · image-editing
- license
- MIT
- native output
- Up to 2048×2048, RGBA
- total params
- 24.86B
- model.arch
- ming_image
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Ming-Image DiT (Z-Image 6B layout) DiffusionTransformer (vendor) / MingImageTransformer2DModel (ai-toolkit) The vendor checkpoint. Not a stock diffusers class: ai-toolkit vendors a port of the diffusers Z-Image transformer with Ming’s zero-masked padding, second caption stream and reference frame. | 6.15B 6,154,901,056 | 12.31 GB | bf16 | yes |
| Transformer (loaded by default) | Ming-Image DiT, ComfyUI int8 convrot repackalternate file The UI default name_or_path. With qtype convrot8 the int8 weights attach as-is; any other qtype loads ming_image_0.1_design_bf16.safetensors and quantizes that. Copies in the local ComfyUI models folder are used before downloading. The param count includes the quantization scales. | 6.16B 6,156,757,185 | 6.18 GB | int8+bf16+fp32+uint8 | yes |
| Text encoder | Ling-mini-2.0 MoE LLM + Qwen2.5-VL vision tower BailingMM2NativeForConditionalGeneration BailingMoeV2 decoder (20 layers, 2048-d, 256 experts, 8 routed + 1 shared per token) plus a 32-layer Qwen2.5-VL vision tower and projector. The vision tower only matters for editing, where it encodes the reference image. | 17.00B 17,001,125,376 | 34.00 GB | bf16 | no |
| Text encoder | Query-token MLP (proj_in, proj_out, proj_directvlm) The 256 learnable query tokens and the projections into and out of the connector and from the direct LLM states. | 31.2M 31,209,216 | 125 MB | fp32 | no |
| Text encoder | Connector (Qwen2 1.5B layout, bidirectional) transformers.Qwen2Model Runs over the 256 query-token states only. ai-toolkit drops its token embedding table (never used, and absent from the ComfyUI repack). | 1.54B 1,543,714,304 | 6.17 GB | fp32 | no |
| Text encoder (loaded by default) | MLLM + MLP + connector, ComfyUI int8 convrot repackalternate file One file holding the three folders above. ai-toolkit loads it as one text encoder module. Its thinker.lm_head.* tensors are dropped on load, since only hidden states are used. If the file has no vision tower, the tower and its projector are read from the vendor mllm/ shards. The tokenizer and image processor always come from the vendor repo. | 18.36B 18,360,862,021 | 19.51 GB | int8+bf16+fp32+uint8 | no |
| Tokenizer | Ling tokenizer + Qwen2-VL image processor | — | — | — | — |
| VAE | Ming-Image RGBA VAE (Qwen-Image VAE layout, 4 channels) diffusers.AutoencoderKLQwenImage Takes and returns RGBA (input_channels 4). ai-toolkit loads the repack copy (vae/ming_image_vae_bf16.safetensors in Comfy-Org/Ming-Image, same param count) by default and converts its keys. | 126.9M 126,897,716 | 254 MB | bf16 | no |
| Training adapter | Ming-Image 0.1 Design training adapter v1adapter A LoRA on the DiT, set as assistant_lora_path by the UI. Active at 1.0 during training and switched off for samples, so samples show the base model plus your LoRA. Never merged: the int8 weights would round the delta away. | 85.0M 85,032,960 | 170 MB | bf16 | no |
| total | 24.86B | 52.87 GB | |||
Latent space
- autoencoder
- Ming-Image RGBA VAE
- pixels per token
- 16×16
- notes
- Three 2× downsampling stages (dim_mult [1, 2, 4, 4]); latents are scaled by scaling_factor 8.0064 with shift 0. Four channels in: RGB images get an opaque alpha on encode, and decode drops it unless RGBA is on. Encoding uses the posterior mode, not a sample. For editing, the reference latent joins as a second frame.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 | 16×1×128×128 | 4,096 | ai-toolkit sample default |
| 2048×2048 | 16×1×256×256 | 16,384 | recommended output |
Architecture
- Blocks
- 30, plus 2 noise-refiner and 2 context-refiner layers
- Hidden size
- 3840 (30 heads × 128)
- FFN size
- 10240 (SwiGLU)
- Text conditioning
- Single stream. Two caption streams: 256 query-token states through the connector (2560-d), and direct LLM states from layers 5 and 12 plus the final output for every prompt token (projected to 3840-d)
- Image conditioning
- Editing: reference image through the vision tower, and its latent as a second frame
- Padding
- Sequences pad to multiples of 32 with zero-masked slots (no learned pad tokens)
- Objective
- Rectified flow. Dynamic shift (0.5 at 256 tokens to 1.15 at 4096), pinned at mu 1.35 for 1024² and up
- Norm / position
- QK RMSNorm, 3D RoPE (axes 32/48/48, theta 256)
In AI Toolkit
- model.arch
- ming_image
- UI label
- Ming-Image 0.1 Design (w/ Training Adapter) (image)
- model.name_or_path
- Comfy-Org/Ming-Image
- source
- extensions_built_in/diffusion_models/ming_image/ming_image.py
- extensions_built_in/diffusion_models/ming_image/src/checkpoints.py
- extensions_built_in/diffusion_models/ming_image/src/pipeline.py
- extensions_built_in/diffusion_models/ming_image/src/transformer.py
- extensions_built_in/diffusion_models/ming_image/src/text_encoder.py
- extensions_built_in/diffusion_models/ming_image/src/bailing_moe_v2.py
- extensions_built_in/diffusion_models/ming_image/src/moe_kernels.py
- extensions_built_in/diffusion_models/ming_image/src/vae.py
- toolkit/models/v2/vae/qwen_image.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- model.low_vram, model.layer_offloading, model.assistant_lora_path
UI defaults
- quantize / quantize_te
- true / true
- qtype / qtype_te
- convrot8 / convrot8
- low_vram
- true
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- shift
- assistant_lora_path
- ostris/ming_image_training_adapter/ming_image_01_design_training_adapter_v1.safetensors
- sample guidance / steps
- 1.0 / 12
- model_kwargs.rgba
- false (Transparency checkbox)
- network.conv
- disabled (linear LoRA only)
Specifics
- Resolution
- Buckets and sample sizes snap to multiples of 16 (8× VAE × 2×2 patch).
- Weight sources
- name_or_path decides. The default ComfyUI repack loads its int8 files as-is under convrot8. The vendor repo, a local checkpoint or a fine-tune loads exactly what it names. A single .safetensors transformer takes the rest from the repack. Configs, tokenizer and image processor come from inclusionAI/Ming-Image-0.1-Design.
- Older configs
- name_or_path Kijai/Ming-Image-ComfyUI, the repack’s earlier home, is rewritten to Comfy-Org/Ming-Image (same files).
- Text encoder quantization
- The MoE expert banks are not nn.Linear, so they are handled separately: any quantize_te makes them int8 weight-only. With layer offloading they get their own stager. Caching text embeddings lets the 16B MoE unload before training.
- Training adapter
- Attached after quantization as a live LoRA (not merged), on for training and off for sampling. Layer offloading is attached after it.
- Transparency
- With model_kwargs.rgba on, images load, encode and decode with their alpha channel, and samples keep their alpha.
- Editing
- A dataset control path turns on editing: one reference image per sample goes to the LLM as vision tokens and to the DiT as a clean second latent frame. The UI does not expose this yet.
- Sampling
- The negative is all-zero conditioning whatever neg says. Guidance 1.0 (no CFG) and 12 steps are the recommended settings. Decodes above 1 MP, or with low_vram, are tiled.
- Quantization
- t_embedder*, cap_embedder*, all_x_embedder* and all_final_layer* stay in full precision.
- LoRA target
- MingImageTransformer2DModel
- Saving
- LoRAs use the ComfyUI key prefix. Full fine-tunes save one .safetensors in the ComfyUI key layout (fused qkv); pre-quantized layers keep their int8 storage. ComfyUI and ai-toolkit both load it.
- Metadata base version
- ming_image
Example config
not verified
job: extensionconfig: name: "my_ming_image_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "shift" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "Comfy-Org/Ming-Image" arch: "ming_image" quantize: true qtype: "convrot8" quantize_te: true qtype_te: "convrot8" low_vram: true assistant_lora_path: "ostris/ming_image_training_adapter/ming_image_01_design_training_adapter_v1.safetensors" model_kwargs: rgba: false sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 1.0 sample_steps: 12 prompts: - "a minimalist poster for a jazz festival, bold typography reading BLUE NOTES"Links
Qwen-Image-2.1
'One 7B single-stream DiT that does both text-to-image and editing with up to 10 reference images, conditioned on Qwen3-VL-8B. Its new 16×, 64-channel VAE is natively RGBA, so it can generate and edit transparent images.'
HiDream-I1 Full
'A 17B sparse diffusion transformer with a mixture-of-experts feed-forward in every block, conditioned on four text encoders: CLIP-L, CLIP-G, T5-XXL and Llama 3.1 8B. Full is the undistilled base, the only HiDream-I1 variant meant for training.'