Zeta-Chroma
'A work-in-progress pixel-space model from lodestones (the Chroma team), built on the Z-Image transformer. It has no VAE: 32×32 RGB patches go straight into the trunk, and a per-token MLP decoder predicts the clean pixels.'
- weights
- lodestones/Zeta-Chroma ↗
- org
- lodestones
- modality
- image
- tasks
- text-to-image
- license
- Apache 2.0
- released
- 2025-12-31not verified
- total params
- 10.52B
- model.arch
- zeta_chroma
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Zeta-Chroma (x0, pixel, DINO distance) ZImageDCT (ai-toolkit) One file holding the Z-Image-shaped trunk, a 3072 → 3840 patch embedder (11,800,320 params) and the pixel decoder dec_net (333,614,592 params). The repo also has zeta-chroma-base-x0-pixel-no-dino.safetensors and a -no-dino-1024 variant with the same parameter count; pick one by putting its filename at the end of name_or_path. | 6.50B 6,498,841,344 | 13.00 GB | bf16 | yes |
| Text encoder | Qwen3-4B transformers.Qwen3ForCausalLM Not in the Zeta-Chroma repo. ai-toolkit loads it from extras_name_or_path, Tongyi-MAI/Z-Image-Turbo. Only the second-to-last hidden state is used. | 4.02B 4,022,468,096 | 8.04 GB | bf16 | no |
| Tokenizer | Qwen3 BPE tokenizer transformers.Qwen2Tokenizer | — | — | — | — |
| total | 10.52B | 21.04 GB | |||
Latent space
- pixels per token
- 32×32
- notes
- No autoencoder: the model works on RGB pixels (ai-toolkit uses an identity FakeVAE with scaling 1.0). Each token is one 32×32×3 patch flattened to 3072 values, so a 1024×1024 image is 1024 tokens, a quarter of latent Z-Image at the same size.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 512×512 | 3×1×512×512 | 256 | |
| 1024×1024 | 3×1×1024×1024 | 1,024 | ai-toolkit sample default |
| 2048×2048 | 3×1×2048×2048 | 4,096 |
Architecture
- Blocks
- Z-Image layout: 30 single-stream blocks, 2 noise-refiner and 2 context-refiner blocks
- Hidden size
- 3840 (30 heads × 128)
- Input
- Linear embedder on 32×32 RGB patches (3072 → 3840)
- Output
- Per-token MLP decoder: 4 adaLN residual blocks, 3840 wide, conditioned on that token’s transformer output, with 8×8 DCT position features on the input
- Prediction
- x0 (clean pixels), turned into a flow velocity as (noisy − x0) / t
- Text conditioning
- Qwen3-4B second-to-last hidden states (2560-d), 512 tokens with a padding mask, in one sequence with the image tokens
- Position
- 3-axis RoPE [32, 48, 48]: text tokens count up on axis 0, image patches start after the prompt length
- Objective
- Rectified flow in pixel space (ai-toolkit shift 3.0)
In AI Toolkit
- model.arch
- zeta_chroma
- UI label
- Zeta Chroma (experimental)
- model.name_or_path
- lodestones/Zeta-Chroma/zeta-chroma-base-x0-pixel-dino-distance.safetensors
- source
- extensions_built_in/diffusion_models/zeta_chroma/zeta_chroma_model.py
- extensions_built_in/diffusion_models/zeta_chroma/zeta_chroma_transformer.py
- extensions_built_in/diffusion_models/zeta_chroma/zeta_chroma_pipeline.py
- toolkit/models/v2/text_encoders/qwen3.py
- toolkit/models/FakeVAE.py
- toolkit/models/registry.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
UI defaults
- extras_name_or_path
- Tongyi-MAI/Z-Image-Turbo
- quantize / quantize_te
- true / true
- noise_scheduler / sampler
- flowmatch / flowmatch
- network.conv
- disabled (linear LoRA only)
Specifics
- Checkpoint resolution
- A local path is used as is. Otherwise name_or_path is a Hub repo, optionally ending in a .safetensors filename; with no filename it downloads zeta-chroma-base-x0-pixel-dino-distance.safetensors.
- x0 detection
- An empty __x0__ tensor in the checkpoint switches on the x0 → velocity conversion.
- VAE
- None. An identity FakeVAE (scaling 1.0) stands in, so cached “latents” are pixels.
- Text encoder source
- text_encoder/ and tokenizer/ from extras_name_or_path. Never trained.
- Resolution
- Buckets snap to multiples of 32 (the pixel patch size).
- Timesteps and loss
- Model gets t = timestep / 1000 (1 = noise). Target noise − pixels. Train scheduler shift 3.0.
- Sampling
- Own pipeline with a resolution-shifted schedule (0.5 to 1.15). With ≤ 8 steps and guidance ≤ 1 it switches to a uniform low-step schedule. CFG runs when guidance_scale > 1 (not shifted like Z-Image).
- LoRA target
- ZImageDCT
- Saving
- Full fine-tunes save the raw ZImageDCT state dict (the layout the Hub files use), dequantized and cast to the save dtype. LoRA keys use the ComfyUI diffusion_model prefix.
- Metadata base version
- zeta_chroma
Example config
not verified
job: extensionconfig: name: "my_zeta_chroma_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: bf16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 model: name_or_path: "lodestones/Zeta-Chroma/zeta-chroma-base-x0-pixel-dino-distance.safetensors" extras_name_or_path: "Tongyi-MAI/Z-Image-Turbo" arch: "zeta_chroma" quantize: true quantize_te: true sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 4 sample_steps: 25 prompts: - "a bear building a log cabin in the snow covered mountains"Links
Qwen2.5-Omni 7B (thinker)
The thinker half of Qwen2.5-Omni 7B, a 7B Qwen2.5 language model with its own audio and vision encoders. AI Toolkit trains it as a captioner, with audio, image or video in and text out, using LoRA on the text stack only.
Ideogram 4
Ideogram’s first open-weight model: a 9.3B single-stream DiT trained from scratch on structured JSON captions, with strong text rendering and bounding-box layout control. It reads 13 layers of Qwen3-VL-8B hidden states and works in the FLUX.2 VAE latent space. Upstream CFG uses a separate unconditional transformer; ai-toolkit swaps it for a LoRA.