Chroma1-Base
'An 8.9B text-to-image model built from FLUX.1-schnell, with CLIP removed and the per-block modulation replaced by one small approximator network. Chroma1-Base is the neutral checkpoint meant for fine-tuning; Chroma1-HD shares its architecture and size.'
- weights
- lodestones/Chroma1-Base ↗
- org
- lodestones
- modality
- image
- tasks
- text-to-image
- license
- Apache 2.0
- released
- 2025-07-29not verified
- total params
- 13.75B
- model.arch
- chroma
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Chroma1-Base MMDiT (single file) Chroma (ai-toolkit, original layout) ai-toolkit loads this single file, not the Diffusers transformer/ folder in the same repo (same parameter count). Chroma1-HD.safetensors in lodestones/Chroma1-HD has the same size and layout. | 8.90B 8,899,983,424 | 17.80 GB | bf16 | yes |
| Text encoder | T5-XXL v1.1 (encoder only) transformers.T5EncoderModel ai-toolkit ignores this folder and loads text_encoder_2/ from ostris/Flex.1-alpha instead (same config, same parameter count and byte size). | 4.76B 4,762,310,656 | 9.52 GB | bf16 | no |
| Tokenizer | T5 SentencePiece (32k vocab) transformers.T5Tokenizer ai-toolkit loads tokenizer_2/ from ostris/Flex.1-alpha. | — | — | — | — |
| VAE | FLUX.1 VAE diffusers.AutoencoderKL ai-toolkit loads vae/ from ostris/Flex.1-alpha, which is the same FLUX.1 VAE (same size and config). | 83.8M 83,819,683 | 168 MB | bf16 | no |
| total | 13.75B | 27.49 GB | |||
Latent space
- autoencoder
- FLUX.1 VAE
- pixels per token
- 16×16
- notes
- Same latent space as FLUX.1 (scaling 0.3611, shift 0.1159). The transformer has patch_size 1 in its config; the 2×2 patching happens outside it, by packing latents into 64-channel tokens (16 × 2 × 2).
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 512×512 | 16×1×64×64 | 1,024 | |
| 1024×1024 | 16×1×128×128 | 4,096 | ai-toolkit sample default |
| 1536×1536 | 16×1×192×192 | 9,216 |
Architecture
- Blocks
- 19 double-stream + 38 single-stream (FLUX layout)
- Hidden size
- 3072 (24 heads × 128)
- MLP ratio
- 4.0
- Text conditioning
- Joint attention on T5-XXL hidden states (4096-d), 512 tokens, padding masked except one token
- Pooled / CLIP input
- None. FLUX’s CLIP-L and pooled-vector path are removed
- Modulation
- An approximator MLP (5 layers, 5120 hidden) maps timestep + a per-slot index to all 344 modulation vectors, replacing FLUX’s per-block modulation layers
- Guidance input
- Kept in the approximator input but always fed 0 (no distilled guidance)
- Objective
- Rectified flow
- Position
- 3-axis RoPE [16, 56, 56], theta 10000
In AI Toolkit
- model.arch
- chroma
- UI label
- Chroma (image)
- model.name_or_path
- lodestones/Chroma1-Base
- source
- extensions_built_in/diffusion_models/chroma/chroma_model.py
- extensions_built_in/diffusion_models/chroma/pipeline.py
- extensions_built_in/diffusion_models/chroma/src/model.py
- extensions_built_in/diffusion_models/chroma/src/layers.py
- toolkit/models/v2/text_encoders/t5.py
- toolkit/models/v2/vae/autoencoder_kl.py
- toolkit/models/registry.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
UI defaults
- quantize / quantize_te
- true / true
- noise_scheduler / sampler
- flowmatch / flowmatch
- network.conv
- disabled (linear LoRA only)
Specifics
- Checkpoint resolution
- name_or_path must resolve to one .safetensors file. lodestones/Chroma picks the newest chroma-unlocked-vNN file, lodestones/Chroma/vNN picks that version, lodestones/Chroma1-<name> downloads <name>.safetensors from that repo; anything else must be a local file.
- Text encoder and VAE source
- Always ostris/Flex.1-alpha (text_encoder_2/, tokenizer_2/, vae/), not the Chroma repo. The CLIP slot is filled with a dummy module.
- Resolution
- Buckets snap to multiples of 32.
- Approximator
- The modulation approximator (distilled_guidance_layer) runs under no_grad in the forward pass, so training only reaches the blocks, not the approximator.
- Text encoder training
- Not supported. T5 stays frozen.
- Loss
- Flow-matching target noise − latents. Train scheduler uses dynamic shift (0.5 to 1.15).
- Sampling
- Flowmatch Euler with resolution-dependent shift. True CFG with the negative prompt runs when guidance_scale > 1.
- LoRA target
- Chroma
- Saving
- Full fine-tunes save one file in the original Chroma layout. LoRA keys use the ComfyUI diffusion_model prefix.
- Metadata base version
- chroma
Example config
not verified
job: extensionconfig: name: "my_chroma_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 16 linear_alpha: 16 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 cache_latents_to_disk: true resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 train_unet: true train_text_encoder: false gradient_checkpointing: true noise_scheduler: "flowmatch" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 model: name_or_path: "lodestones/Chroma1-Base" arch: "chroma" quantize: true quantize_te: true sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 neg: "" guidance_scale: 4 sample_steps: 25 prompts: - "a bear building a log cabin in the snow covered mountains"Links
Flex.1-alpha
An 8B, Apache 2.0 descendant of FLUX.1 [schnell] with 8 double-stream blocks instead of 19. It uses the FLUX.1 text encoders and VAE unchanged, and has a guidance embedder that can be bypassed, which is how it is meant to be fine-tuned.
Lumina-Image 2.0
A 2.6B flow-matching diffusion transformer that reads Gemma 2 2B hidden states and works in the FLUX.1 latent space. Text and image tokens pass through separate refiner layers, then share one stack of single-stream blocks.