Docs
AI ToolkitModels

ACE-Step 1.5 XL

The XL (4B) base DiT of ACE-Step 1.5. It writes full songs from a style caption, lyrics and metadata, using the same condition encoder, text encoder and 48 kHz stereo Oobleck VAE (1920×, 25 latent frames per second) as the 2B model.

org
ACE Studio & StepFun
modality
audio
tasks
text-to-music
license
MITnot verified
released
2026-04-02not verified
native output
48 kHz stereo
total params
7.61B
model.arch
ace_step_15_xl

Components

rolemodelparamssizedtypetrained
All-in-one checkpoint
ACE-Step 1.5 XL base AIO

One ComfyUI-layout file with every part below, split by key prefix. The rows below point at the same file and give their parameter counts in the notes. The two fp32 tensors are logit_scale scalars. The same repo also has ace_step_1.5_xl_sft_aio.safetensors, the SFT variant, with the same size and parameter count.

7.61B
7,606,026,506
15.21 GBbf16+fp16+fp32—
Transformer
ACE-Step 1.5 XL DiT decoder
DiTModel (ai-toolkit)

4,168,897,088 params, bf16, under model.diffusion_model.decoder.*. The only part the LoRA targets.

———yes
Condition encoder
ACE-Step 1.5 condition encoder (text projector, lyric encoder, timbre encoder)
ConditionEncoder (ai-toolkit)

608,367,616 params, bf16, under model.diffusion_model.encoder.*, plus the 2,048-param null_condition_emb used for CFG. Runs while encoding prompts, so its output is the prompt embedding and it is never trained.

———no
Audio tokenizer / detokenizer
ACE-Step 1.5 FSQ audio tokenizer and detokenizer

105,032,198 (tokenizer) + 105,011,776 (detokenizer) params, bf16, under model.diffusion_model.tokenizer.* and .detokenizer.*. Used by cover mode upstream. ai-toolkit skips these keys.

———no
Text encoder
Qwen3-Embedding-0.6B
transformers.Qwen3Model

595,776,512 params, fp16, under text_encoders.qwen3_06b.*. The full model encodes the caption; only its token embedding table is used for the lyrics.

———no
Tokenizer
Qwen3 BPE
transformers.AutoTokenizer

Not in the AIO file. ai-toolkit downloads it from Qwen/Qwen3-Embedding-0.6B.

————
Language model (planner)
ACE-Step 5Hz LM 1.7B
Qwen3 (28 layers, 2048 hidden, 217,204 vocab)

1,854,243,840 params, bf16, under text_encoders.qwen3_2b.*. Shapes match ACE-Step/Ace-Step1.5 acestep-5Hz-lm-1.7B. Upstream it is the planner that writes song metadata, lyrics and captions for the DiT. ai-toolkit never loads it.

———no
VAE
ACE-Step 1.5 Oobleck VAE
OobleckVAE (ai-toolkit)

168,695,426 params (84,281,344 encoder + 84,414,082 decoder), fp16, under vae.*. Upstream copy: ACE-Step/Ace-Step1.5 vae/ (diffusers.AutoencoderOobleck).

———no
total7.61B15.21 GB

Audio latent space

sample rate
48 kHz
temporal
1,920×
channels
64
latent rate
25/s
autoencoder
ACE-Step 1.5 Oobleck VAE
audio channels
stereo
patch
2 latent steps per token
tokens per second
12.50
notes
Downsampling strides 2 × 4 × 4 × 6 × 10 = 1920, so 25 latent frames per second. The DiT patchifies 2 frames per token with a Conv1d over 192 channels: 64 noisy latent + 64 source latent + 64 chunk mask.

Architecture

DiT blocks
32
Hidden size
2560 (32 heads × 128, 8 KV heads)
FFN size
9728 (SwiGLU)
Self-attention
Alternating sliding window (128 tokens) and full attention
Text conditioning
Cross-attention in every block on the condition encoder output (2048-d, projected to 2560)
Condition encoder
Qwen3-Embedding caption states (1024-d, 256 tokens max) projected to 2048; lyric encoder, 8 layers, on Qwen3 token embeddings (2048 tokens max); timbre encoder, 4 layers, on reference latents
Timesteps
Two embeddings (t and t − r), summed, into 6-way AdaLN per block
Objective
Rectified flow; ai-toolkit uses shift 3.0
Norm / position
RMSNorm, QK norm, 1D RoPE (theta 1e6)

In AI Toolkit

model.arch
ace_step_15_xl
UI label
ACE-Step 1.5 XL (audio)
model.name_or_path
ostris/ace_step_1.5_ComfyUI_files/ace_step_1.5_xl_base_aio.safetensors
extra UI sections
sample.multi_ctrl_imgs, model.low_vram, model.layer_offloading

UI defaults

quantize / quantize_te
true / true (qfloat8)
low_vram
true
train.unload_text_encoder
false
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
linear
sample
tagged song prompt, 180 s, guidance 4, 30 steps
network.conv
disabled (linear LoRA only)

Specifics

Loading
name_or_path is a local file or exactly org/repo/file.safetensors. The AIO file is split by key prefix into DiT, VAE and the 0.6B text encoder; the 5Hz LM and the audio tokenizer/detokenizer keys are ignored. The tokenizer comes from Qwen/Qwen3-Embedding-0.6B.
Dataset
Audio files only. Each file is loaded whole, converted to stereo, resampled to 48 kHz and bucketed by its length in ms. There is no cropping, so clip length sets memory use.
Captions
Tagged text: <CAPTION>, <LYRICS>, <BPM>, <KEYSCALE>, <TIMESIGNATURE>, <DURATION>, <LANGUAGE>. BPM defaults to 120.
Conditioning
The prompt embedding is the condition encoder output, built with a silence latent as the timbre reference and as the source context (text-to-music mode only; no cover or repaint).
Loss
Flow target noise − latents. The timestep is passed as t/1000 for both t and r.
LoRA target
DiTModel
Quantization
DiT (with the condition encoder) and text encoder are both quantized to qfloat8 by default.
low_vram
Loads the weights on the CPU before quantizing and turns on tiled VAE decoding (10 s tiles, 1 s overlap).
Saving
LoRA keys use the ComfyUI prefix (transformer. → diffusion_model.). Full fine-tunes cannot be saved.
Sampling
Euler flow matching with shift 3.0. With guidance above 1, the unconditional pass uses the learned null_condition_emb. Length comes from the DURATION tag. Output is mp3 or wav.
Not supported
layer_offloading, full fine-tune saving

Example config

not verified

job: extensionconfig:  name: "my_ace_step_15_xl_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: bf16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/songs"          caption_ext: "txt"          caption_dropout_rate: 0.05          cache_latents_to_disk: true          resolution: [512]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "linear"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16      model:        name_or_path: "ostris/ace_step_1.5_ComfyUI_files/ace_step_1.5_xl_base_aio.safetensors"        arch: "ace_step_15_xl"        quantize: true        quantize_te: true        qtype: "qfloat8"        low_vram: true      sample:        sampler: "flowmatch"        sample_every: 250        guidance_scale: 4        sample_steps: 30        prompts:          - "<CAPTION>upbeat synth pop</CAPTION><LYRICS>[Verse 1]\nHello world</LYRICS><BPM>112</BPM><DURATION>60</DURATION><LANGUAGE>en</LANGUAGE>"

On this page