Docs
AI ToolkitModels

Stable Diffusion XL 1.0 Base

The 2.6B UNet latent diffusion model with two CLIP text encoders (CLIP-L and OpenCLIP bigG) and size/crop micro-conditioning. Still the base of a large ecosystem of fine-tunes, LoRAs and ControlNets.

org
Stability AI
modality
image
tasks
text-to-image
license
CreativeML Open RAIL++-M
released
2023-07-26not verified
native output
1024×1024not verified
total params
3.47B
model.arch
sdxl

Components

rolemodelparamssizedtypetrained
UNet
SDXL UNet
diffusers.UNet2DConditionModel

The file ai-toolkit loads (from_pretrained with use_safetensors, no variant). The repo also has an fp16 variant (5.1 GB), Flax, ONNX and OpenVINO copies, and single-file checkpoints at the root.

2.57B
2,567,463,684
10.27 GBfp32yes
Text encoder 1
CLIP ViT-L/14 text
transformers.CLIPTextModel
123.1M
123,060,480
492 MBfp32no
Text encoder 2
OpenCLIP ViT-bigG/14 text (with projection)
transformers.CLIPTextModelWithProjection

Supplies the pooled embedding as well as hidden states.

694.7M
694,659,840
2.78 GBfp32no
Tokenizers
CLIP BPE ×2 (77 tokens)
transformers.CLIPTokenizer

tokenizer/ and tokenizer_2/.

————
VAE
SDXL VAE
diffusers.AutoencoderKL

Same architecture as the SD 1.5 VAE with different weights; scaling factor 0.13025. Its config sets force_upcast, so pipelines decode it in fp32.

83.7M
83,653,863
335 MBfp32no
total3.47B13.88 GB

Latent space

spatial
8×
channels
4
patch
1×1
autoencoder
SDXL VAE
pixels per token
8×8
notes
A UNet has no patch embedding; each latent pixel is one position. The token counts below are latent pixels at full resolution. Attention only runs at the 2× and 4× downsampled levels.
inputlatent (c×t×h×w)tokens
1024×10244×1×128×12816,384native
832×12164×1×152×10415,808portrait bucket

Architecture

Levels
3 (320, 640, 1280 channels), 2 res blocks each
Transformer depth
0 / 2 / 10 blocks at the three levels, 10 in the mid block
Heads
64-d heads: 10 at 640 channels, 20 at 1280
Text conditioning
Cross-attention on penultimate-layer hidden states of both encoders, concatenated to 2048-d, 77 tokens
Micro-conditioning
Pooled bigG embedding (1280) + original size, crop and target size (6 × 256), added to the timestep embedding
Objective
Epsilon prediction, scaled-linear betas (0.00085–0.012), 1000 steps

In AI Toolkit

model.arch
sdxl
UI label
SDXL (image)
model.name_or_path
stabilityai/stable-diffusion-xl-base-1.0
extra UI sections

UI defaults

noise_scheduler / sampler
ddpm / ddpm
sample.guidance_scale
6
quantize / quantize_te
false / false (section hidden)
train.timestep_type
hidden in the UI

Specifics

Model code
Handled by the legacy StableDiffusion class, not a model plugin. arch sdxl sets is_xl and loads a StableDiffusionXLPipeline.
Loading
A Hub id or folder loads with from_pretrained (fp32 safetensors, no fp16 variant). A local .safetensors file loads with from_single_file, so CivitAI-style checkpoints work. vae_path swaps the VAE.
Resolution
Buckets snap to multiples of 16. Size and crop conditioning are set to the bucket size with a (0, 0) crop.
Text encoders
Both are frozen by default. Penultimate hidden states from each are concatenated; the pooled embedding always comes from text_encoder_2.
Noise schedule
ddpm uses the scaled-linear SD schedule with epsilon prediction. is_v_pred switches to v-prediction for v-pred fine-tunes.
Network
Conv (LoCon) layers are available alongside linear LoRA; LoRA targets UNet2DConditionModel.
Saving
LoRAs use kohya-style keys (lora_unet_, lora_te1_, lora_te2_) that ComfyUI and A1111 load. Full fine-tunes save a single-file LDM checkpoint by default (save_format diffusers for a folder).
Metadata base version
sdxl_1.0

Example config

not verified

job: extensionconfig:  name: "my_sdxl_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [768, 1024]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "ddpm"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16      model:        name_or_path: "stabilityai/stable-diffusion-xl-base-1.0"        arch: "sdxl"      sample:        sampler: "ddpm"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 6        sample_steps: 25        prompts:          - "a bear building a log cabin in the snow covered mountains"

On this page