Docs
AI ToolkitModels

Ming-Image 0.1 Design

A 6B Z-Image-style DiT for UI, infographics, posters and other text-heavy design, with RGBA output. It is conditioned by a Ling-mini-2.0 MoE multimodal LLM through two caption streams. ai-toolkit trains it from the int8 ComfyUI repack with a training adapter.

org
inclusionAI
modality
image
tasks
text-to-image · image-editing
license
MIT
native output
Up to 2048×2048, RGBA
total params
24.86B
model.arch
ming_image

Components

rolemodelparamssizedtypetrained
Transformer
Ming-Image DiT (Z-Image 6B layout)
DiffusionTransformer (vendor) / MingImageTransformer2DModel (ai-toolkit)

The vendor checkpoint. Not a stock diffusers class: ai-toolkit vendors a port of the diffusers Z-Image transformer with Ming’s zero-masked padding, second caption stream and reference frame.

6.15B
6,154,901,056
12.31 GBbf16yes
Transformer (loaded by default)
Ming-Image DiT, ComfyUI int8 convrot repackalternate file

The UI default name_or_path. With qtype convrot8 the int8 weights attach as-is; any other qtype loads ming_image_0.1_design_bf16.safetensors and quantizes that. Copies in the local ComfyUI models folder are used before downloading. The param count includes the quantization scales.

6.16B
6,156,757,185
6.18 GBint8+bf16+fp32+uint8yes
Text encoder
Ling-mini-2.0 MoE LLM + Qwen2.5-VL vision tower
BailingMM2NativeForConditionalGeneration

BailingMoeV2 decoder (20 layers, 2048-d, 256 experts, 8 routed + 1 shared per token) plus a 32-layer Qwen2.5-VL vision tower and projector. The vision tower only matters for editing, where it encodes the reference image.

17.00B
17,001,125,376
34.00 GBbf16no
Text encoder
Query-token MLP (proj_in, proj_out, proj_directvlm)

The 256 learnable query tokens and the projections into and out of the connector and from the direct LLM states.

31.2M
31,209,216
125 MBfp32no
Text encoder
Connector (Qwen2 1.5B layout, bidirectional)
transformers.Qwen2Model

Runs over the 256 query-token states only. ai-toolkit drops its token embedding table (never used, and absent from the ComfyUI repack).

1.54B
1,543,714,304
6.17 GBfp32no
Text encoder (loaded by default)
MLLM + MLP + connector, ComfyUI int8 convrot repackalternate file

One file holding the three folders above. ai-toolkit loads it as one text encoder module. Its thinker.lm_head.* tensors are dropped on load, since only hidden states are used. If the file has no vision tower, the tower and its projector are read from the vendor mllm/ shards. The tokenizer and image processor always come from the vendor repo.

18.36B
18,360,862,021
19.51 GBint8+bf16+fp32+uint8no
Tokenizer
Ling tokenizer + Qwen2-VL image processor
————
VAE
Ming-Image RGBA VAE (Qwen-Image VAE layout, 4 channels)
diffusers.AutoencoderKLQwenImage

Takes and returns RGBA (input_channels 4). ai-toolkit loads the repack copy (vae/ming_image_vae_bf16.safetensors in Comfy-Org/Ming-Image, same param count) by default and converts its keys.

126.9M
126,897,716
254 MBbf16no
Training adapter
Ming-Image 0.1 Design training adapter v1adapter

A LoRA on the DiT, set as assistant_lora_path by the UI. Active at 1.0 during training and switched off for samples, so samples show the base model plus your LoRA. Never merged: the int8 weights would round the delta away.

85.0M
85,032,960
170 MBbf16no
total24.86B52.87 GB

Latent space

spatial
8×
channels
16
patch
2×2
autoencoder
Ming-Image RGBA VAE
pixels per token
16×16
notes
Three 2× downsampling stages (dim_mult [1, 2, 4, 4]); latents are scaled by scaling_factor 8.0064 with shift 0. Four channels in: RGB images get an opaque alpha on encode, and decode drops it unless RGBA is on. Encoding uses the posterior mode, not a sample. For editing, the reference latent joins as a second frame.
inputlatent (c×t×h×w)tokens
1024×102416×1×128×1284,096ai-toolkit sample default
2048×204816×1×256×25616,384recommended output

Architecture

Blocks
30, plus 2 noise-refiner and 2 context-refiner layers
Hidden size
3840 (30 heads × 128)
FFN size
10240 (SwiGLU)
Text conditioning
Single stream. Two caption streams: 256 query-token states through the connector (2560-d), and direct LLM states from layers 5 and 12 plus the final output for every prompt token (projected to 3840-d)
Image conditioning
Editing: reference image through the vision tower, and its latent as a second frame
Padding
Sequences pad to multiples of 32 with zero-masked slots (no learned pad tokens)
Objective
Rectified flow. Dynamic shift (0.5 at 256 tokens to 1.15 at 4096), pinned at mu 1.35 for 1024² and up
Norm / position
QK RMSNorm, 3D RoPE (axes 32/48/48, theta 256)

In AI Toolkit

model.arch
ming_image
UI label
Ming-Image 0.1 Design (w/ Training Adapter) (image)
model.name_or_path
Comfy-Org/Ming-Image
extra UI sections
model.low_vram, model.layer_offloading, model.assistant_lora_path

UI defaults

quantize / quantize_te
true / true
qtype / qtype_te
convrot8 / convrot8
low_vram
true
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
shift
assistant_lora_path
ostris/ming_image_training_adapter/ming_image_01_design_training_adapter_v1.safetensors
sample guidance / steps
1.0 / 12
model_kwargs.rgba
false (Transparency checkbox)
network.conv
disabled (linear LoRA only)

Specifics

Resolution
Buckets and sample sizes snap to multiples of 16 (8× VAE × 2×2 patch).
Weight sources
name_or_path decides. The default ComfyUI repack loads its int8 files as-is under convrot8. The vendor repo, a local checkpoint or a fine-tune loads exactly what it names. A single .safetensors transformer takes the rest from the repack. Configs, tokenizer and image processor come from inclusionAI/Ming-Image-0.1-Design.
Older configs
name_or_path Kijai/Ming-Image-ComfyUI, the repack’s earlier home, is rewritten to Comfy-Org/Ming-Image (same files).
Text encoder quantization
The MoE expert banks are not nn.Linear, so they are handled separately: any quantize_te makes them int8 weight-only. With layer offloading they get their own stager. Caching text embeddings lets the 16B MoE unload before training.
Training adapter
Attached after quantization as a live LoRA (not merged), on for training and off for sampling. Layer offloading is attached after it.
Transparency
With model_kwargs.rgba on, images load, encode and decode with their alpha channel, and samples keep their alpha.
Editing
A dataset control path turns on editing: one reference image per sample goes to the LLM as vision tokens and to the DiT as a clean second latent frame. The UI does not expose this yet.
Sampling
The negative is all-zero conditioning whatever neg says. Guidance 1.0 (no CFG) and 12 steps are the recommended settings. Decodes above 1 MP, or with low_vram, are tiled.
Quantization
t_embedder*, cap_embedder*, all_x_embedder* and all_final_layer* stay in full precision.
LoRA target
MingImageTransformer2DModel
Saving
LoRAs use the ComfyUI key prefix. Full fine-tunes save one .safetensors in the ComfyUI key layout (fused qkv); pre-quantized layers keep their int8 storage. ComfyUI and ai-toolkit both load it.
Metadata base version
ming_image

Example config

not verified

job: extensionconfig:  name: "my_ming_image_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "shift"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "Comfy-Org/Ming-Image"        arch: "ming_image"        quantize: true        qtype: "convrot8"        quantize_te: true        qtype_te: "convrot8"        low_vram: true        assistant_lora_path: "ostris/ming_image_training_adapter/ming_image_01_design_training_adapter_v1.safetensors"        model_kwargs:          rgba: false      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 1.0        sample_steps: 12        prompts:          - "a minimalist poster for a jazz festival, bold typography reading BLUE NOTES"

On this page