Docs
AI ToolkitModels

Stable Diffusion 1.5

The original 860M UNet latent diffusion model with a single CLIP ViT-L text encoder, trained at 512×512. Small and fast to train, with the largest library of community fine-tunes of any model.

org
Runway / CompVis
modality
image
tasks
text-to-image
license
CreativeML Open RAIL-M
released
2022-10-20not verified
native output
512×512not verified
total params
1.07B
model.arch
sd15

Components

rolemodelparamssizedtypetrained
UNet
SD 1.5 UNet (EMA weights)
diffusers.UNet2DConditionModel

The file Diffusers loads by default. The repo also has fp16 and .bin copies, non-EMA weights (diffusion_pytorch_model.non_ema.safetensors) and single-file checkpoints at the root.

859.5M
859,520,964
3.44 GBfp32yes
Text encoder
CLIP ViT-L/14 text
transformers.CLIPTextModel

The count includes 77 int64 position ids stored as a tensor; 123,060,480 weights.

123.1M
123,060,557
492 MBfp32+int64no
Tokenizer
CLIP BPE (77 tokens)
transformers.CLIPTokenizer
————
VAE
SD VAE (kl-f8)
diffusers.AutoencoderKL

No scaling_factor in its config, so the Diffusers default 0.18215 applies.

83.7M
83,653,863
335 MBfp32no
Safety checker
CLIP ViT-L/14 image classifier
StableDiffusionSafetyChecker

Shipped in the repo but not loaded: ai-toolkit passes safety_checker=None.

———no
total1.07B4.27 GB

Latent space

spatial
8×
channels
4
patch
1×1
autoencoder
SD VAE (kl-f8)
pixels per token
8×8
notes
A UNet has no patch embedding; each latent pixel is one position. The token counts below are latent pixels at full resolution. Attention runs at the 1×, 2× and 4× downsampled levels.
inputlatent (c×t×h×w)tokens
512×5124×1×64×644,096native
768×7684×1×96×969,216

Architecture

Levels
4 (320, 640, 1280, 1280 channels), 2 res blocks each
Transformer depth
1 block at each of the first 3 levels and in the mid block; none at the lowest
Heads
8 per attention layer (head size 40 / 80 / 160)
Text conditioning
Cross-attention on the final CLIP-L hidden states (768-d), 77 tokens
Objective
Epsilon prediction, scaled-linear betas (0.00085–0.012), 1000 steps

In AI Toolkit

model.arch
sd15
UI label
SD 1.5 (image)
model.name_or_path
stable-diffusion-v1-5/stable-diffusion-v1-5
extra UI sections

UI defaults

noise_scheduler / sampler
ddpm / ddpm
sample size
512×512
sample.guidance_scale
6
quantize / train.timestep_type
hidden in the UI

Specifics

Model code
Handled by the legacy StableDiffusion class, not a model plugin. Loads a StableDiffusionPipeline.
Loading
A Hub id or folder loads with from_pretrained (fp32 safetensors, no fp16 variant) and no safety checker. A local .safetensors or .ckpt file loads with from_single_file. vae_path swaps the VAE.
Resolution
Buckets snap to multiples of 16.
Text encoder
Frozen by default. Uses the final hidden states (no clip skip); long prompts can be chunked in 77-token windows.
Noise schedule
ddpm uses the scaled-linear SD schedule with epsilon prediction. is_v_pred switches to v-prediction for v-pred fine-tunes.
Network
Conv (LoCon) layers are available alongside linear LoRA; LoRA targets UNet2DConditionModel.
Saving
LoRAs use kohya-style keys (lora_unet_, lora_te_) that ComfyUI and A1111 load. Full fine-tunes save a single-file LDM checkpoint by default (save_format diffusers for a folder).
Metadata base version
sd_1.5

Example config

not verified

job: extensionconfig:  name: "my_sd15_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [512, 768]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "ddpm"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16      model:        name_or_path: "stable-diffusion-v1-5/stable-diffusion-v1-5"        arch: "sd1"      sample:        sampler: "ddpm"        sample_every: 250        width: 512        height: 512        guidance_scale: 6        sample_steps: 25        prompts:          - "a bear building a log cabin in the snow covered mountains"

On this page