Docs
AI ToolkitModels

Chroma1-Base

'An 8.9B text-to-image model built from FLUX.1-schnell, with CLIP removed and the per-block modulation replaced by one small approximator network. Chroma1-Base is the neutral checkpoint meant for fine-tuning; Chroma1-HD shares its architecture and size.'

org
lodestones
modality
image
tasks
text-to-image
license
Apache 2.0
released
2025-07-29not verified
total params
13.75B
model.arch
chroma

Components

rolemodelparamssizedtypetrained
Transformer
Chroma1-Base MMDiT (single file)
Chroma (ai-toolkit, original layout)

ai-toolkit loads this single file, not the Diffusers transformer/ folder in the same repo (same parameter count). Chroma1-HD.safetensors in lodestones/Chroma1-HD has the same size and layout.

8.90B
8,899,983,424
17.80 GBbf16yes
Text encoder
T5-XXL v1.1 (encoder only)
transformers.T5EncoderModel

ai-toolkit ignores this folder and loads text_encoder_2/ from ostris/Flex.1-alpha instead (same config, same parameter count and byte size).

4.76B
4,762,310,656
9.52 GBbf16no
Tokenizer
T5 SentencePiece (32k vocab)
transformers.T5Tokenizer

ai-toolkit loads tokenizer_2/ from ostris/Flex.1-alpha.

————
VAE
FLUX.1 VAE
diffusers.AutoencoderKL

ai-toolkit loads vae/ from ostris/Flex.1-alpha, which is the same FLUX.1 VAE (same size and config).

83.8M
83,819,683
168 MBbf16no
total13.75B27.49 GB

Latent space

spatial
8×
channels
16
patch
2×2
autoencoder
FLUX.1 VAE
pixels per token
16×16
notes
Same latent space as FLUX.1 (scaling 0.3611, shift 0.1159). The transformer has patch_size 1 in its config; the 2×2 patching happens outside it, by packing latents into 64-channel tokens (16 × 2 × 2).
inputlatent (c×t×h×w)tokens
512×51216×1×64×641,024
1024×102416×1×128×1284,096ai-toolkit sample default
1536×153616×1×192×1929,216

Architecture

Blocks
19 double-stream + 38 single-stream (FLUX layout)
Hidden size
3072 (24 heads × 128)
MLP ratio
4.0
Text conditioning
Joint attention on T5-XXL hidden states (4096-d), 512 tokens, padding masked except one token
Pooled / CLIP input
None. FLUX’s CLIP-L and pooled-vector path are removed
Modulation
An approximator MLP (5 layers, 5120 hidden) maps timestep + a per-slot index to all 344 modulation vectors, replacing FLUX’s per-block modulation layers
Guidance input
Kept in the approximator input but always fed 0 (no distilled guidance)
Objective
Rectified flow
Position
3-axis RoPE [16, 56, 56], theta 10000

In AI Toolkit

model.arch
chroma
UI label
Chroma (image)
model.name_or_path
lodestones/Chroma1-Base
extra UI sections

UI defaults

quantize / quantize_te
true / true
noise_scheduler / sampler
flowmatch / flowmatch
network.conv
disabled (linear LoRA only)

Specifics

Checkpoint resolution
name_or_path must resolve to one .safetensors file. lodestones/Chroma picks the newest chroma-unlocked-vNN file, lodestones/Chroma/vNN picks that version, lodestones/Chroma1-<name> downloads <name>.safetensors from that repo; anything else must be a local file.
Text encoder and VAE source
Always ostris/Flex.1-alpha (text_encoder_2/, tokenizer_2/, vae/), not the Chroma repo. The CLIP slot is filled with a dummy module.
Resolution
Buckets snap to multiples of 32.
Approximator
The modulation approximator (distilled_guidance_layer) runs under no_grad in the forward pass, so training only reaches the blocks, not the approximator.
Text encoder training
Not supported. T5 stays frozen.
Loss
Flow-matching target noise − latents. Train scheduler uses dynamic shift (0.5 to 1.15).
Sampling
Flowmatch Euler with resolution-dependent shift. True CFG with the negative prompt runs when guidance_scale > 1.
LoRA target
Chroma
Saving
Full fine-tunes save one file in the original Chroma layout. LoRA keys use the ComfyUI diffusion_model prefix.
Metadata base version
chroma

Example config

not verified

job: extensionconfig:  name: "my_chroma_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 16        linear_alpha: 16      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          cache_latents_to_disk: true          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 2000        train_unet: true        train_text_encoder: false        gradient_checkpointing: true        noise_scheduler: "flowmatch"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16      model:        name_or_path: "lodestones/Chroma1-Base"        arch: "chroma"        quantize: true        quantize_te: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        neg: ""        guidance_scale: 4        sample_steps: 25        prompts:          - "a bear building a log cabin in the snow covered mountains"

On this page