ACE-Step 1.5
A 2B diffusion transformer that writes full songs from a style caption, lyrics and metadata (BPM, key, time signature, duration). It works in the 48 kHz stereo latent space of a 1920× Oobleck VAE, 25 latent frames per second.
- org
- ACE Studio & StepFun
- modality
- audio
- tasks
- text-to-music
- license
- MITnot verified
- released
- 2026-01-23not verified
- native output
- 48 kHz stereo
- total params
- 5.01B
- model.arch
- ace_step_15
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| All-in-one checkpoint | ACE-Step 1.5 base AIO One ComfyUI-layout file with every part below, split by key prefix. The rows below point at the same file and give their parameter counts in the notes. The two fp32 tensors are logit_scale scalars. | 5.01B 5,012,588,298 | 10.03 GB | bf16+fp16+fp32 | — |
| Transformer | ACE-Step 1.5 DiT decoder DiTModel (ai-toolkit) 1,575,458,880 params, bf16, under model.diffusion_model.decoder.*. The only part the LoRA targets. | — | — | — | yes |
| Condition encoder | ACE-Step 1.5 condition encoder (text projector, lyric encoder, timbre encoder) ConditionEncoder (ai-toolkit) 608,367,616 params, bf16, under model.diffusion_model.encoder.*, plus the 2,048-param null_condition_emb used for CFG. Runs while encoding prompts, so its output is the prompt embedding and it is never trained. | — | — | — | no |
| Audio tokenizer / detokenizer | ACE-Step 1.5 FSQ audio tokenizer and detokenizer 105,032,198 (tokenizer) + 105,011,776 (detokenizer) params, bf16, under model.diffusion_model.tokenizer.* and .detokenizer.*. Used by cover mode upstream. ai-toolkit skips these keys. | — | — | — | no |
| Text encoder | Qwen3-Embedding-0.6B transformers.Qwen3Model 595,776,512 params, fp16, under text_encoders.qwen3_06b.*. The full model encodes the caption; only its token embedding table is used for the lyrics. | — | — | — | no |
| Tokenizer | Qwen3 BPE transformers.AutoTokenizer Not in the AIO file. ai-toolkit downloads it from Qwen/Qwen3-Embedding-0.6B. | — | — | — | — |
| Language model (planner) | ACE-Step 5Hz LM 1.7B Qwen3 (28 layers, 2048 hidden, 217,204 vocab) 1,854,243,840 params, bf16, under text_encoders.qwen3_2b.*. Shapes match ACE-Step/Ace-Step1.5 acestep-5Hz-lm-1.7B. Upstream it is the planner that writes song metadata, lyrics and captions for the DiT. ai-toolkit never loads it. | — | — | — | no |
| VAE | ACE-Step 1.5 Oobleck VAE OobleckVAE (ai-toolkit) 168,695,426 params (84,281,344 encoder + 84,414,082 decoder), fp16, under vae.*. Upstream copy: ACE-Step/Ace-Step1.5 vae/ (diffusers.AutoencoderOobleck). | — | — | — | no |
| total | 5.01B | 10.03 GB | |||
Audio latent space
- autoencoder
- ACE-Step 1.5 Oobleck VAE
- audio channels
- stereo
- patch
- 2 latent steps per token
- tokens per second
- 12.50
- notes
- Downsampling strides 2 × 4 × 4 × 6 × 10 = 1920, so 25 latent frames per second. The DiT patchifies 2 frames per token with a Conv1d over 192 channels: 64 noisy latent + 64 source latent + 64 chunk mask.
Architecture
- DiT blocks
- 24
- Hidden size
- 2048 (16 heads × 128, 8 KV heads)
- FFN size
- 6144 (SwiGLU)
- Self-attention
- Alternating sliding window (128 tokens) and full attention
- Text conditioning
- Cross-attention in every block on the condition encoder output (2048-d)
- Condition encoder
- Qwen3-Embedding caption states (1024-d, 256 tokens max) projected to 2048; lyric encoder, 8 layers, on Qwen3 token embeddings (2048 tokens max); timbre encoder, 4 layers, on reference latents
- Timesteps
- Two embeddings (t and t − r), summed, into 6-way AdaLN per block
- Objective
- Rectified flow; ai-toolkit uses shift 3.0
- Norm / position
- RMSNorm, QK norm, 1D RoPE (theta 1e6)
In AI Toolkit
- model.arch
- ace_step_15
- UI label
- ACE-Step 1.5 (audio)
- model.name_or_path
- ostris/ace_step_1.5_ComfyUI_files/ace_step_1.5_base_aio.safetensors
- source
- extra UI sections
- sample.multi_ctrl_imgs, model.low_vram, model.layer_offloading
UI defaults
- quantize / quantize_te
- true / true (qfloat8)
- low_vram
- true
- train.unload_text_encoder
- false
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- linear
- sample
- tagged song prompt, 180 s, guidance 4, 30 steps
- network.conv
- disabled (linear LoRA only)
Specifics
- Loading
- name_or_path is a local file or exactly org/repo/file.safetensors. The AIO file is split by key prefix into DiT, VAE and the 0.6B text encoder; the 5Hz LM and the audio tokenizer/detokenizer keys are ignored. The tokenizer comes from Qwen/Qwen3-Embedding-0.6B.
- Dataset
- Audio files only. Each file is loaded whole, converted to stereo, resampled to 48 kHz and bucketed by its length in ms. There is no cropping, so clip length sets memory use.
- Captions
- Tagged text: <CAPTION>, <LYRICS>, <BPM>, <KEYSCALE>, <TIMESIGNATURE>, <DURATION>, <LANGUAGE>. BPM defaults to 120.
- Conditioning
- The prompt embedding is the condition encoder output, built with a silence latent as the timbre reference and as the source context (text-to-music mode only; no cover or repaint).
- Loss
- Flow target noise − latents. The timestep is passed as t/1000 for both t and r.
- LoRA target
- DiTModel
- Quantization
- DiT (with the condition encoder) and text encoder are both quantized to qfloat8 by default.
- low_vram
- Loads the weights on the CPU before quantizing and turns on tiled VAE decoding (10 s tiles, 1 s overlap).
- Saving
- LoRA keys use the ComfyUI prefix (transformer. → diffusion_model.). Full fine-tunes cannot be saved.
- Sampling
- Euler flow matching with shift 3.0. With guidance above 1, the unconditional pass uses the learned null_condition_emb. Length comes from the DURATION tag. Output is mp3 or wav.
- Not supported
- layer_offloading, full fine-tune saving
Example config
not verified
job: extensionconfig: name: "my_ace_step_15_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: bf16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/songs" caption_ext: "txt" caption_dropout_rate: 0.05 cache_latents_to_disk: true resolution: [512] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "linear" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 model: name_or_path: "ostris/ace_step_1.5_ComfyUI_files/ace_step_1.5_base_aio.safetensors" arch: "ace_step_15" quantize: true quantize_te: true qtype: "qfloat8" low_vram: true sample: sampler: "flowmatch" sample_every: 250 guidance_scale: 4 sample_steps: 30 prompts: - "<CAPTION>upbeat synth pop</CAPTION><LYRICS>[Verse 1]\nHello world</LYRICS><BPM>112</BPM><DURATION>60</DURATION><LANGUAGE>en</LANGUAGE>"Links
YuE2 3B
A song model built as two 28-layer experts on one Qwen3-shaped backbone. The AR expert reads a style line and lyrics and writes a lead sheet and semantic codec tokens (25 per second); the NAR expert renders those tokens into 48 kHz stereo VAE latents with flow matching.
Qwen2.5-Omni 7B (thinker)
The thinker half of Qwen2.5-Omni 7B, a 7B Qwen2.5 language model with its own audio and vision encoders. AI Toolkit trains it as a captioner, with audio, image or video in and text out, using LoRA on the text stack only.