Stable Diffusion 1.5
The original 860M UNet latent diffusion model with a single CLIP ViT-L text encoder, trained at 512×512. Small and fast to train, with the largest library of community fine-tunes of any model.
- org
- Runway / CompVis
- modality
- image
- tasks
- text-to-image
- license
- CreativeML Open RAIL-M
- released
- 2022-10-20not verified
- native output
- 512×512not verified
- total params
- 1.07B
- model.arch
- sd15
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| UNet | SD 1.5 UNet (EMA weights) diffusers.UNet2DConditionModel The file Diffusers loads by default. The repo also has fp16 and .bin copies, non-EMA weights (diffusion_pytorch_model.non_ema.safetensors) and single-file checkpoints at the root. | 859.5M 859,520,964 | 3.44 GB | fp32 | yes |
| Text encoder | CLIP ViT-L/14 text transformers.CLIPTextModel The count includes 77 int64 position ids stored as a tensor; 123,060,480 weights. | 123.1M 123,060,557 | 492 MB | fp32+int64 | no |
| Tokenizer | CLIP BPE (77 tokens) transformers.CLIPTokenizer | — | — | — | — |
| VAE | SD VAE (kl-f8) diffusers.AutoencoderKL No scaling_factor in its config, so the Diffusers default 0.18215 applies. | 83.7M 83,653,863 | 335 MB | fp32 | no |
| Safety checker | CLIP ViT-L/14 image classifier StableDiffusionSafetyChecker Shipped in the repo but not loaded: ai-toolkit passes safety_checker=None. | — | — | — | no |
| total | 1.07B | 4.27 GB | |||
Latent space
- autoencoder
- SD VAE (kl-f8)
- pixels per token
- 8×8
- notes
- A UNet has no patch embedding; each latent pixel is one position. The token counts below are latent pixels at full resolution. Attention runs at the 1×, 2× and 4× downsampled levels.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 512×512 | 4×1×64×64 | 4,096 | native |
| 768×768 | 4×1×96×96 | 9,216 |
Architecture
- Levels
- 4 (320, 640, 1280, 1280 channels), 2 res blocks each
- Transformer depth
- 1 block at each of the first 3 levels and in the mid block; none at the lowest
- Heads
- 8 per attention layer (head size 40 / 80 / 160)
- Text conditioning
- Cross-attention on the final CLIP-L hidden states (768-d), 77 tokens
- Objective
- Epsilon prediction, scaled-linear betas (0.00085–0.012), 1000 steps
In AI Toolkit
- model.arch
- sd15
- UI label
- SD 1.5 (image)
- model.name_or_path
- stable-diffusion-v1-5/stable-diffusion-v1-5
- source
- extra UI sections
UI defaults
- noise_scheduler / sampler
- ddpm / ddpm
- sample size
- 512×512
- sample.guidance_scale
- 6
- quantize / train.timestep_type
- hidden in the UI
Specifics
- Model code
- Handled by the legacy StableDiffusion class, not a model plugin. Loads a StableDiffusionPipeline.
- Loading
- A Hub id or folder loads with from_pretrained (fp32 safetensors, no fp16 variant) and no safety checker. A local .safetensors or .ckpt file loads with from_single_file. vae_path swaps the VAE.
- Resolution
- Buckets snap to multiples of 16.
- Text encoder
- Frozen by default. Uses the final hidden states (no clip skip); long prompts can be chunked in 77-token windows.
- Noise schedule
- ddpm uses the scaled-linear SD schedule with epsilon prediction. is_v_pred switches to v-prediction for v-pred fine-tunes.
- Network
- Conv (LoCon) layers are available alongside linear LoRA; LoRA targets UNet2DConditionModel.
- Saving
- LoRAs use kohya-style keys (lora_unet_, lora_te_) that ComfyUI and A1111 load. Full fine-tunes save a single-file LDM checkpoint by default (save_format diffusers for a folder).
- Metadata base version
- sd_1.5
Example config
not verified
job: extensionconfig: name: "my_sd15_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [512, 768] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "ddpm" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 model: name_or_path: "stable-diffusion-v1-5/stable-diffusion-v1-5" arch: "sd1" sample: sampler: "ddpm" sample_every: 250 width: 512 height: 512 guidance_scale: 6 sample_steps: 25 prompts: - "a bear building a log cabin in the snow covered mountains"Links
Stable Diffusion XL 1.0 Base
The 2.6B UNet latent diffusion model with two CLIP text encoders (CLIP-L and OpenCLIP bigG) and size/crop micro-conditioning. Still the base of a large ecosystem of fine-tunes, LoRAs and ControlNets.
OmniGen2
A 4B Lumina-style diffusion transformer conditioned on Qwen2.5-VL 3B. Reference images enter as extra latent tokens through their own refiner, so one model does text-to-image, instruction editing and subject-driven generation.