YuE2 3B
A song model built as two 28-layer experts on one Qwen3-shaped backbone. The AR expert reads a style line and lyrics and writes a lead sheet and semantic codec tokens (25 per second); the NAR expert renders those tokens into 48 kHz stereo VAE latents with flow matching.
- weights
- Comfy-Org/YuE2 ↗
- org
- M-A-P (Multimodal Art Projection)not verified
- modality
- audio
- tasks
- text-to-music · lyrics-to-song
- license
- CC BY-NC 4.0
- released
- 2026-09-09not verified
- native output
- 48 kHz stereo
- total params
- 5.10B
- model.arch
- yue2
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| All-in-one checkpoint (int8) | YuE2 3B, convrot int8 repackalternate file ai-toolkit default. Linear layers of both experts are int8 with convrot rotation; the count includes quantization scales and the embedded tokenizer (8,034,652 uint8 bytes stored as a tensor). The VAE is fp16 here. The rows below point at parts of this file. | 3.77B 3,772,873,990 | 3.96 GB | int8+fp16+bf16+uint8+fp32 | — |
| All-in-one checkpoint (bf16) | YuE2 3B, bf16 repack Same layout without quantization; the VAE is fp32 here. Also loads as name_or_path. | 3.77B 3,771,337,054 | 7.80 GB | bf16+fp32+uint8 | — |
| AR expert | YuE2 AR (composition) YuE2AR (ai-toolkit) 2,165,957,632 params in the bf16 file, under text_encoders.model.*. Token embedding, 28 causal layers, final norm and a 184,704-token lm head. Writes the ABC sheet and the codec tokens. Doubles as the text side: the prompt embedding is its own token embeddings. | — | — | — | yes |
| NAR expert | YuE2 NAR (rendering) YuE2NAR (ai-toolkit) 1,464,728,640 params in the bf16 file, under model.diffusion_model.*. The same 28-layer stack plus latent in/out projections (64 ↔ 2048) and a timestep embedder. Predicts flow velocity on VAE latents while attending into the AR key/value cache. | — | — | — | yes |
| Tokenizer | Qwen BPE with ABC and codec tokens Embedded in the checkpoint as text_encoders.yue2_tokenizer_json (uint8 bytes of a tokenizer.json). | — | — | — | — |
| VAE | YuE2 Oobleck-style VAE YuE2VAE (ai-toolkit) 132,616,130 params (66,241,664 encoder + 66,374,466 decoder), under vae.*. ai-toolkit runs it in fp32. Encoding takes the mean half of the bottleneck, in 60 s chunks. The checkpoint metadata names its source as YuE2-Vae (m-a-p/YuE2-Vae). | — | — | — | no |
| Audio tokenizer backbone | MERT-v2-FullSong Used at latent-cache time only. Layer-20 features of 24 kHz mono audio, resampled to 25 Hz, feed the community semantic head. The official audio-to-token encoder is unreleased. | 632.4M 632,429,312 | 2.53 GB | fp32 | no |
| Semantic head | Mothersuperior realaudio tokenizer v4 (community) 8-layer transformer classifier over 512-frame windows that maps MERT features to YuE2 codec tokens. A .pt file, so no parameter count. The repo also has matching NAR adapters (nar_lora_joint_v4 and later v5 / v8 / v9) that ai-toolkit can merge on load. | — | — | — | no |
| Lead-sheet transcriber | SheetSage2 Audio → ABC sheet, MERT-v2-based encoder with its adapter already merged (source m-a-p/SheetSage2). Used at latent-cache time only, when cot is full or melody. | 693.2M 693,240,089 | 1.39 GB | bf16+fp32 | no |
| total | 5.10B | 11.72 GB | |||
Audio latent space
- autoencoder
- YuE2 Oobleck-style VAE
- audio channels
- stereo
- patch
- 1 latent steps per token
- tokens per second
- 25.00
- notes
- Strides 2 × 2 × 4 × 4 × 5 × 6 = 1920, so 25 latent frames per second, the same rate as the codec tokens (one token per latent frame). The NAR adds a START and an END slot around the window.
Architecture
- Experts
- AR + NAR, each 28 layers, same shapes
- Hidden size
- 2048 (16 heads × 128, 8 KV heads)
- FFN size
- 6144 (SwiGLU, fused gate_up)
- Vocabulary
- 184,704 (Qwen text + ABC markers + 32,768 codec tokens)
- Context
- 24,576 tokens
- Coupling
- NAR layer i attends over its own tokens plus AR layer i cached keys/values for prefix + ABC + codec tokens + MUSIC_END (mixture of transformers)
- Generation
- AR writes an ABC lead sheet (cot full or melody, or none with off), then codec tokens; NAR renders chunks with flow matching
- Objective
- AR: next-token CE. NAR: rectified flow, timestep shift 1.0
- Norm / position
- RMSNorm, QK norm, RoPE (theta 1e6); NAR adds a frame position table
In AI Toolkit
- model.arch
- yue2
- UI label
- YuE2 (audio)
- model.name_or_path
- Comfy-Org/YuE2/checkpoints/yue2_3b_int8_convrot.safetensors
- source
- extensions_built_in/audio_models/yue2/yue2_model.py
- extensions_built_in/audio_models/yue2/src/model.py
- extensions_built_in/audio_models/yue2/src/pipeline.py
- extensions_built_in/audio_models/yue2/src/tokenizer.py
- extensions_built_in/audio_models/yue2/src/vae.py
- extensions_built_in/audio_models/yue2/src/sheetsage_generate.py
- extensions_built_in/audio_models/yue2/src/sheetsage_decoder.py
- extensions_built_in/audio_models/yue2/src/sheetsage_repair.py
- extensions_built_in/audio_models/base_audio_model.py
- extensions_built_in/audio_models/ui.tsx
- extra UI sections
- model.low_vram, sample.duration
UI defaults
- quantize
- true (convrot8)
- low_vram
- false
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- sigmoid
- datasets.cache_latents_to_disk
- true
- datasets.resolution
- [512] (one bucket)
- datasets.caption_dropout_rate
- 0
- model_kwargs
- cot: full, abc_dropout: 0.5, sample_ar_repetition_penalty: 1.2, ar_kl_weight: 0.2
- sample
- style + [Lyrics] prompt, up to 120 s, guidance 1, 32 steps
- network.conv / model.quantize_te
- disabled
Specifics
- Status
- Experimental. The AR memorizes small datasets within a few hundred steps (loss/ar_ce falls toward 0); it needs a large, varied dataset to learn a style.
- Loading
- Local file or org/repo/path, downloaded into the ComfyUI models folder. qtype convrot8 keeps the shipped int8 layers as they are. Embeddings, lm head, NAR in/out projections and the timestep embedder are never quantized.
- Captions
- A style line, then [Lyrics], then lyrics with [Verse 1] / [Chorus] headers (normalized to title case). Legacy <CAPTION>/<LYRICS> tags also parse. Do not use caption dropout: a blank prompt breaks lyric following.
- Latent cache (required)
- Per song it stores VAE latents, codec tokens from MERT + the semantic head, and (cot full/melody) a SheetSage2 ABC sheet, about 12 s per song. The cache records the cot mode; changing it means deleting _latent_cache. A non-default semantic head gets its own cache key. The encoders are dropped once training starts.
- Training step
- One LoRA covers both experts. The NAR gets the flow loss on a random window (train_window_frames 1500 = 60 s). The AR gets next-token CE over the codec tokens from the song start (ar_loss_weight 1.0), plus KL to the base model when ar_kl_weight > 0. The flow loss does not reach the AR through its cache.
- Sheet dropout
- abc_dropout (0.5) trains that fraction of items without the sheet, as an off-mode prompt, so one LoRA works with and without a lead sheet.
- Separation
- do_separation splits songs with MelBandRoformer at cache time and adds lyrics-only → vocals and tags-only → instrumental AR loss terms (weights 0.5 and 1.0).
- AR learning rate
- ar_lr_multiplier puts the AR expert LoRA in its own optimizer group.
- LoRA target
- YuE2AR, YuE2NAR
- Saving
- LoRA keys are rewritten for ComfyUI: NAR → diffusion_model.*, AR → text_encoders.*. Full fine-tunes cannot be saved.
- Sampling
- AR writes the sheet (up to 8192 tokens), then codec tokens up to the sample duration (repetition penalty from model_kwargs); NAR renders each context-sized chunk with sample_steps flow steps (32); tiled VAE decode. Output mp3, wav or flac.
- Not supported
- layer_offloading, full fine-tune saving
Example config
not verified
job: extensionconfig: name: "my_yue2_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: bf16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/songs" caption_ext: "txt" caption_dropout_rate: 0 cache_latents_to_disk: true resolution: [512] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "sigmoid" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 model: name_or_path: "Comfy-Org/YuE2/checkpoints/yue2_3b_int8_convrot.safetensors" arch: "yue2" quantize: true qtype: "convrot8" model_kwargs: cot: "full" abc_dropout: 0.5 sample_ar_repetition_penalty: 1.2 ar_kl_weight: 0.2 sample: sampler: "flowmatch" sample_every: 250 guidance_scale: 1 sample_steps: 32 duration: 120 prompts: - "upbeat synth pop, female vocals\n[Lyrics]\n[Verse 1]\nHello world\n"Links
ACE-Step 1.5 XL
The XL (4B) base DiT of ACE-Step 1.5. It writes full songs from a style caption, lyrics and metadata, using the same condition encoder, text encoder and 48 kHz stereo Oobleck VAE (1920×, 25 latent frames per second) as the 2B model.
ACE-Step 1.5
A 2B diffusion transformer that writes full songs from a style caption, lyrics and metadata (BPM, key, time signature, duration). It works in the 48 kHz stereo latent space of a 1920× Oobleck VAE, 25 latent frames per second.