Stable Diffusion XL 1.0 Base
The 2.6B UNet latent diffusion model with two CLIP text encoders (CLIP-L and OpenCLIP bigG) and size/crop micro-conditioning. Still the base of a large ecosystem of fine-tunes, LoRAs and ControlNets.
- org
- Stability AI
- modality
- image
- tasks
- text-to-image
- license
- CreativeML Open RAIL++-M
- released
- 2023-07-26not verified
- native output
- 1024×1024not verified
- total params
- 3.47B
- model.arch
- sdxl
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| UNet | SDXL UNet diffusers.UNet2DConditionModel The file ai-toolkit loads (from_pretrained with use_safetensors, no variant). The repo also has an fp16 variant (5.1 GB), Flax, ONNX and OpenVINO copies, and single-file checkpoints at the root. | 2.57B 2,567,463,684 | 10.27 GB | fp32 | yes |
| Text encoder 1 | CLIP ViT-L/14 text transformers.CLIPTextModel | 123.1M 123,060,480 | 492 MB | fp32 | no |
| Text encoder 2 | OpenCLIP ViT-bigG/14 text (with projection) transformers.CLIPTextModelWithProjection Supplies the pooled embedding as well as hidden states. | 694.7M 694,659,840 | 2.78 GB | fp32 | no |
| Tokenizers | CLIP BPE ×2 (77 tokens) transformers.CLIPTokenizer tokenizer/ and tokenizer_2/. | — | — | — | — |
| VAE | SDXL VAE diffusers.AutoencoderKL Same architecture as the SD 1.5 VAE with different weights; scaling factor 0.13025. Its config sets force_upcast, so pipelines decode it in fp32. | 83.7M 83,653,863 | 335 MB | fp32 | no |
| total | 3.47B | 13.88 GB | |||
Latent space
- autoencoder
- SDXL VAE
- pixels per token
- 8×8
- notes
- A UNet has no patch embedding; each latent pixel is one position. The token counts below are latent pixels at full resolution. Attention only runs at the 2× and 4× downsampled levels.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 | 4×1×128×128 | 16,384 | native |
| 832×1216 | 4×1×152×104 | 15,808 | portrait bucket |
Architecture
- Levels
- 3 (320, 640, 1280 channels), 2 res blocks each
- Transformer depth
- 0 / 2 / 10 blocks at the three levels, 10 in the mid block
- Heads
- 64-d heads: 10 at 640 channels, 20 at 1280
- Text conditioning
- Cross-attention on penultimate-layer hidden states of both encoders, concatenated to 2048-d, 77 tokens
- Micro-conditioning
- Pooled bigG embedding (1280) + original size, crop and target size (6 × 256), added to the timestep embedding
- Objective
- Epsilon prediction, scaled-linear betas (0.00085–0.012), 1000 steps
In AI Toolkit
- model.arch
- sdxl
- UI label
- SDXL (image)
- model.name_or_path
- stabilityai/stable-diffusion-xl-base-1.0
- source
- extra UI sections
UI defaults
- noise_scheduler / sampler
- ddpm / ddpm
- sample.guidance_scale
- 6
- quantize / quantize_te
- false / false (section hidden)
- train.timestep_type
- hidden in the UI
Specifics
- Model code
- Handled by the legacy StableDiffusion class, not a model plugin. arch sdxl sets is_xl and loads a StableDiffusionXLPipeline.
- Loading
- A Hub id or folder loads with from_pretrained (fp32 safetensors, no fp16 variant). A local .safetensors file loads with from_single_file, so CivitAI-style checkpoints work. vae_path swaps the VAE.
- Resolution
- Buckets snap to multiples of 16. Size and crop conditioning are set to the bucket size with a (0, 0) crop.
- Text encoders
- Both are frozen by default. Penultimate hidden states from each are concatenated; the pooled embedding always comes from text_encoder_2.
- Noise schedule
- ddpm uses the scaled-linear SD schedule with epsilon prediction. is_v_pred switches to v-prediction for v-pred fine-tunes.
- Network
- Conv (LoCon) layers are available alongside linear LoRA; LoRA targets UNet2DConditionModel.
- Saving
- LoRAs use kohya-style keys (lora_unet_, lora_te1_, lora_te2_) that ComfyUI and A1111 load. Full fine-tunes save a single-file LDM checkpoint by default (save_format diffusers for a folder).
- Metadata base version
- sdxl_1.0
Example config
not verified
job: extensionconfig: name: "my_sdxl_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "ddpm" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 model: name_or_path: "stabilityai/stable-diffusion-xl-base-1.0" arch: "sdxl" sample: sampler: "ddpm" sample_every: 250 width: 1024 height: 1024 guidance_scale: 6 sample_steps: 25 prompts: - "a bear building a log cabin in the snow covered mountains"Links
HiDream-I1 Full
'A 17B sparse diffusion transformer with a mixture-of-experts feed-forward in every block, conditioned on four text encoders: CLIP-L, CLIP-G, T5-XXL and Llama 3.1 8B. Full is the undistilled base, the only HiDream-I1 variant meant for training.'
Stable Diffusion 1.5
The original 860M UNet latent diffusion model with a single CLIP ViT-L text encoder, trained at 512×512. Small and fast to train, with the largest library of community fine-tunes of any model.