Nucleus-Image
A 17B sparse mixture-of-experts DiT that activates about 2B parameters per step. Text from Qwen3-VL-8B enters only as keys and values, and images use the Qwen-Image VAE. Released as a pre-trained base model with no preference tuning.
- weights
- NucleusAI/Nucleus-Image ↗
- org
- Nucleus AI
- modality
- image
- tasks
- text-to-image
- license
- Apache 2.0
- total params
- 25.82B
- model.arch
- nucleus_image
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Nucleus-Image MoE DiT (17B total) diffusers.NucleusMoEImageTransformer2DModel 64 routed experts plus one shared expert in each MoE block. The UI default LoRA leaves the routed experts alone. | 16.92B 16,922,679,296 | 33.85 GB | bf16 | yes |
| Text encoder | Qwen3-VL-8B-Instruct transformers.Qwen3VLForConditionalGeneration The full VLM (36-layer 4096-d text model, 27-layer vision tower). Same parameter count and byte size as Qwen/Qwen3-VL-8B-Instruct. Only the text path runs. | 8.77B 8,767,123,696 | 17.53 GB | bf16 | no |
| Tokenizer | Qwen3-VL processor transformers.Qwen3VLProcessor | — | — | — | — |
| VAE | Qwen-Image VAE diffusers.AutoencoderKLQwenImage The same VAE as Qwen-Image. A Wan-style video VAE; images are encoded as one frame. | 126.9M 126,892,531 | 254 MB | bf16 | no |
| total | 25.82B | 51.63 GB | |||
Latent space
- autoencoder
- Qwen-Image VAE
- pixels per token
- 16×16
- notes
- Three 2× downsampling stages (dim_mult [1, 2, 4, 4]). Latents are normalized with the per-channel latents_mean / latents_std from the VAE config. The pipeline packs 2×2 patches, so the transformer takes 64 channels in and returns 16.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 | 16×1×128×128 | 4,096 | ai-toolkit sample default |
| 512×512 | 16×1×64×64 | 1,024 |
Architecture
- Blocks
- 32 (first 3 dense, 29 MoE)
- Hidden size
- 2048 (16 query heads × 128, 4 KV heads)
- Experts
- 64 routed + 1 shared per MoE block, expert FFN 1344, route scale 2.5
- Routing
- Expert-choice with per-layer capacity factors (4.0 in the first two MoE blocks, 2.0 after)
- Text conditioning
- Joint attention: image queries attend to image + text keys/values. Text tokens (4096-d) are not updated by the blocks
- Objective
- Rectified flow, shift 1.0
- Position
- 3-axis RoPE (16/56/56)
In AI Toolkit
- model.arch
- nucleus_image
- UI label
- Nucleus-Image (image)
- model.name_or_path
- NucleusAI/Nucleus-Image
- source
- extra UI sections
- model.low_vram
UI defaults
- quantize / quantize_te
- true / true
- timestep_type
- linear
- network.linear / linear_alpha
- 128 / 128
- network_kwargs.ignore_if_contains
- img_mlp.experts, img_mlp.gate
- network.conv
- disabled (linear LoRA only)
Specifics
- Resolution
- Buckets and sample sizes snap to multiples of 32. One token covers 16×16 pixels, so 16 would be enough.
- LoRA on MoE
- The UI default keeps LoRA off the router gate and the routed experts (fused parameter banks, not linear layers). Attention, the dense FFNs and the shared experts get rank 128.
- Prompt encoding
- Prompts are wrapped in the pipeline’s system-prompt chat template, capped at 1024 tokens, padded to a multiple of 8, and the hidden state 8 layers from the end is used.
- Component sources
- The transformer comes from name_or_path; the text encoder, processor and VAE from extras_name_or_path (defaults to name_or_path). A local folder with a text_encoder/ subfolder is used for everything.
- Prediction sign
- The transformer output is negated to match ai-toolkit’s noise − clean target. Timesteps are passed as t / 1000.
- Grouped GEMM
- On PyTorch builds without torch.nn.functional.grouped_mm, the experts fall back to a per-expert loop.
- LoRA target
- NucleusMoEImageTransformer2DModel
- Saving
- LoRAs use the ComfyUI key prefix. Full fine-tunes save transformer/ in Diffusers format.
- Metadata base version
- nucleus_image
Example config
not verified
job: extensionconfig: name: "my_nucleus_image_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 128 linear_alpha: 128 network_kwargs: ignore_if_contains: - "img_mlp.experts" - "img_mlp.gate" save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "linear" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "NucleusAI/Nucleus-Image" arch: "nucleus_image" quantize: true quantize_te: true low_vram: true sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 4 sample_steps: 30 prompts: - "woman with red hair, playing chess at the park, bomb going off in the background"Links
FLUX.2 [klein] 9B Base
The larger FLUX.2 [klein]: a 9B transformer conditioned on Qwen3-8B, using the same 32-channel FLUX.2 VAE and reference-image editing as FLUX.2 [dev]. The Base release is undistilled (no step or guidance distillation) and meant for fine-tuning. Unlike the 4B, it is under the non-commercial FLUX license.
HiDream-O1-Image
An 8B pixel-space image model built inside a Qwen3-VL language model. Text, timestep and 32×32 pixel patches share one token sequence in the same transformer, with no VAE and no separate text encoder. Generates up to 2048×2048.