Docs
AI ToolkitModels

Zeta-Chroma

'A work-in-progress pixel-space model from lodestones (the Chroma team), built on the Z-Image transformer. It has no VAE: 32×32 RGB patches go straight into the trunk, and a per-token MLP decoder predicts the clean pixels.'

org
lodestones
modality
image
tasks
text-to-image
license
Apache 2.0
released
2025-12-31not verified
total params
10.52B
model.arch
zeta_chroma

Components

rolemodelparamssizedtypetrained
Transformer
Zeta-Chroma (x0, pixel, DINO distance)
ZImageDCT (ai-toolkit)

One file holding the Z-Image-shaped trunk, a 3072 → 3840 patch embedder (11,800,320 params) and the pixel decoder dec_net (333,614,592 params). The repo also has zeta-chroma-base-x0-pixel-no-dino.safetensors and a -no-dino-1024 variant with the same parameter count; pick one by putting its filename at the end of name_or_path.

6.50B
6,498,841,344
13.00 GBbf16yes
Text encoder
Qwen3-4B
transformers.Qwen3ForCausalLM

Not in the Zeta-Chroma repo. ai-toolkit loads it from extras_name_or_path, Tongyi-MAI/Z-Image-Turbo. Only the second-to-last hidden state is used.

4.02B
4,022,468,096
8.04 GBbf16no
Tokenizer
Qwen3 BPE tokenizer
transformers.Qwen2Tokenizer
————
total10.52B21.04 GB

Latent space

spatial
1×
channels
3
patch
32×32
pixels per token
32×32
notes
No autoencoder: the model works on RGB pixels (ai-toolkit uses an identity FakeVAE with scaling 1.0). Each token is one 32×32×3 patch flattened to 3072 values, so a 1024×1024 image is 1024 tokens, a quarter of latent Z-Image at the same size.
inputlatent (c×t×h×w)tokens
512×5123×1×512×512256
1024×10243×1×1024×10241,024ai-toolkit sample default
2048×20483×1×2048×20484,096

Architecture

Blocks
Z-Image layout: 30 single-stream blocks, 2 noise-refiner and 2 context-refiner blocks
Hidden size
3840 (30 heads × 128)
Input
Linear embedder on 32×32 RGB patches (3072 → 3840)
Output
Per-token MLP decoder: 4 adaLN residual blocks, 3840 wide, conditioned on that token’s transformer output, with 8×8 DCT position features on the input
Prediction
x0 (clean pixels), turned into a flow velocity as (noisy − x0) / t
Text conditioning
Qwen3-4B second-to-last hidden states (2560-d), 512 tokens with a padding mask, in one sequence with the image tokens
Position
3-axis RoPE [32, 48, 48]: text tokens count up on axis 0, image patches start after the prompt length
Objective
Rectified flow in pixel space (ai-toolkit shift 3.0)

In AI Toolkit

model.arch
zeta_chroma
UI label
Zeta Chroma (experimental)
model.name_or_path
lodestones/Zeta-Chroma/zeta-chroma-base-x0-pixel-dino-distance.safetensors
extra UI sections

UI defaults

extras_name_or_path
Tongyi-MAI/Z-Image-Turbo
quantize / quantize_te
true / true
noise_scheduler / sampler
flowmatch / flowmatch
network.conv
disabled (linear LoRA only)

Specifics

Checkpoint resolution
A local path is used as is. Otherwise name_or_path is a Hub repo, optionally ending in a .safetensors filename; with no filename it downloads zeta-chroma-base-x0-pixel-dino-distance.safetensors.
x0 detection
An empty __x0__ tensor in the checkpoint switches on the x0 → velocity conversion.
VAE
None. An identity FakeVAE (scaling 1.0) stands in, so cached “latents” are pixels.
Text encoder source
text_encoder/ and tokenizer/ from extras_name_or_path. Never trained.
Resolution
Buckets snap to multiples of 32 (the pixel patch size).
Timesteps and loss
Model gets t = timestep / 1000 (1 = noise). Target noise − pixels. Train scheduler shift 3.0.
Sampling
Own pipeline with a resolution-shifted schedule (0.5 to 1.15). With ≤ 8 steps and guidance ≤ 1 it switches to a uniform low-step schedule. CFG runs when guidance_scale > 1 (not shifted like Z-Image).
LoRA target
ZImageDCT
Saving
Full fine-tunes save the raw ZImageDCT state dict (the layout the Hub files use), dequantized and cast to the save dtype. LoRA keys use the ComfyUI diffusion_model prefix.
Metadata base version
zeta_chroma

Example config

not verified

job: extensionconfig:  name: "my_zeta_chroma_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: bf16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16      model:        name_or_path: "lodestones/Zeta-Chroma/zeta-chroma-base-x0-pixel-dino-distance.safetensors"        extras_name_or_path: "Tongyi-MAI/Z-Image-Turbo"        arch: "zeta_chroma"        quantize: true        quantize_te: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 4        sample_steps: 25        prompts:          - "a bear building a log cabin in the snow covered mountains"

On this page