Docs
AI ToolkitModels

Nucleus-Image

A 17B sparse mixture-of-experts DiT that activates about 2B parameters per step. Text from Qwen3-VL-8B enters only as keys and values, and images use the Qwen-Image VAE. Released as a pre-trained base model with no preference tuning.

org
Nucleus AI
modality
image
tasks
text-to-image
license
Apache 2.0
total params
25.82B
model.arch
nucleus_image

Components

rolemodelparamssizedtypetrained
Transformer
Nucleus-Image MoE DiT (17B total)
diffusers.NucleusMoEImageTransformer2DModel

64 routed experts plus one shared expert in each MoE block. The UI default LoRA leaves the routed experts alone.

16.92B
16,922,679,296
33.85 GBbf16yes
Text encoder
Qwen3-VL-8B-Instruct
transformers.Qwen3VLForConditionalGeneration

The full VLM (36-layer 4096-d text model, 27-layer vision tower). Same parameter count and byte size as Qwen/Qwen3-VL-8B-Instruct. Only the text path runs.

8.77B
8,767,123,696
17.53 GBbf16no
Tokenizer
Qwen3-VL processor
transformers.Qwen3VLProcessor
————
VAE
Qwen-Image VAE
diffusers.AutoencoderKLQwenImage

The same VAE as Qwen-Image. A Wan-style video VAE; images are encoded as one frame.

126.9M
126,892,531
254 MBbf16no
total25.82B51.63 GB

Latent space

spatial
8×
channels
16
patch
2×2
autoencoder
Qwen-Image VAE
pixels per token
16×16
notes
Three 2× downsampling stages (dim_mult [1, 2, 4, 4]). Latents are normalized with the per-channel latents_mean / latents_std from the VAE config. The pipeline packs 2×2 patches, so the transformer takes 64 channels in and returns 16.
inputlatent (c×t×h×w)tokens
1024×102416×1×128×1284,096ai-toolkit sample default
512×51216×1×64×641,024

Architecture

Blocks
32 (first 3 dense, 29 MoE)
Hidden size
2048 (16 query heads × 128, 4 KV heads)
Experts
64 routed + 1 shared per MoE block, expert FFN 1344, route scale 2.5
Routing
Expert-choice with per-layer capacity factors (4.0 in the first two MoE blocks, 2.0 after)
Text conditioning
Joint attention: image queries attend to image + text keys/values. Text tokens (4096-d) are not updated by the blocks
Objective
Rectified flow, shift 1.0
Position
3-axis RoPE (16/56/56)

In AI Toolkit

model.arch
nucleus_image
UI label
Nucleus-Image (image)
model.name_or_path
NucleusAI/Nucleus-Image
extra UI sections
model.low_vram

UI defaults

quantize / quantize_te
true / true
timestep_type
linear
network.linear / linear_alpha
128 / 128
network_kwargs.ignore_if_contains
img_mlp.experts, img_mlp.gate
network.conv
disabled (linear LoRA only)

Specifics

Resolution
Buckets and sample sizes snap to multiples of 32. One token covers 16×16 pixels, so 16 would be enough.
LoRA on MoE
The UI default keeps LoRA off the router gate and the routed experts (fused parameter banks, not linear layers). Attention, the dense FFNs and the shared experts get rank 128.
Prompt encoding
Prompts are wrapped in the pipeline’s system-prompt chat template, capped at 1024 tokens, padded to a multiple of 8, and the hidden state 8 layers from the end is used.
Component sources
The transformer comes from name_or_path; the text encoder, processor and VAE from extras_name_or_path (defaults to name_or_path). A local folder with a text_encoder/ subfolder is used for everything.
Prediction sign
The transformer output is negated to match ai-toolkit’s noise − clean target. Timesteps are passed as t / 1000.
Grouped GEMM
On PyTorch builds without torch.nn.functional.grouped_mm, the experts fall back to a per-expert loop.
LoRA target
NucleusMoEImageTransformer2DModel
Saving
LoRAs use the ComfyUI key prefix. Full fine-tunes save transformer/ in Diffusers format.
Metadata base version
nucleus_image

Example config

not verified

job: extensionconfig:  name: "my_nucleus_image_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 128        linear_alpha: 128        network_kwargs:          ignore_if_contains:            - "img_mlp.experts"            - "img_mlp.gate"      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "linear"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "NucleusAI/Nucleus-Image"        arch: "nucleus_image"        quantize: true        quantize_te: true        low_vram: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 4        sample_steps: 30        prompts:          - "woman with red hair, playing chess at the park, bomb going off in the background"

On this page