Docs
AI ToolkitModels

YuE2 3B

A song model built as two 28-layer experts on one Qwen3-shaped backbone. The AR expert reads a style line and lyrics and writes a lead sheet and semantic codec tokens (25 per second); the NAR expert renders those tokens into 48 kHz stereo VAE latents with flow matching.

org
M-A-P (Multimodal Art Projection)not verified
modality
audio
tasks
text-to-music · lyrics-to-song
license
CC BY-NC 4.0
released
2026-09-09not verified
native output
48 kHz stereo
total params
5.10B
model.arch
yue2

Components

rolemodelparamssizedtypetrained
All-in-one checkpoint (int8)
YuE2 3B, convrot int8 repackalternate file

ai-toolkit default. Linear layers of both experts are int8 with convrot rotation; the count includes quantization scales and the embedded tokenizer (8,034,652 uint8 bytes stored as a tensor). The VAE is fp16 here. The rows below point at parts of this file.

3.77B
3,772,873,990
3.96 GBint8+fp16+bf16+uint8+fp32—
All-in-one checkpoint (bf16)
YuE2 3B, bf16 repack

Same layout without quantization; the VAE is fp32 here. Also loads as name_or_path.

3.77B
3,771,337,054
7.80 GBbf16+fp32+uint8—
AR expert
YuE2 AR (composition)
YuE2AR (ai-toolkit)

2,165,957,632 params in the bf16 file, under text_encoders.model.*. Token embedding, 28 causal layers, final norm and a 184,704-token lm head. Writes the ABC sheet and the codec tokens. Doubles as the text side: the prompt embedding is its own token embeddings.

———yes
NAR expert
YuE2 NAR (rendering)
YuE2NAR (ai-toolkit)

1,464,728,640 params in the bf16 file, under model.diffusion_model.*. The same 28-layer stack plus latent in/out projections (64 ↔ 2048) and a timestep embedder. Predicts flow velocity on VAE latents while attending into the AR key/value cache.

———yes
Tokenizer
Qwen BPE with ABC and codec tokens

Embedded in the checkpoint as text_encoders.yue2_tokenizer_json (uint8 bytes of a tokenizer.json).

————
VAE
YuE2 Oobleck-style VAE
YuE2VAE (ai-toolkit)

132,616,130 params (66,241,664 encoder + 66,374,466 decoder), under vae.*. ai-toolkit runs it in fp32. Encoding takes the mean half of the bottleneck, in 60 s chunks. The checkpoint metadata names its source as YuE2-Vae (m-a-p/YuE2-Vae).

———no
Audio tokenizer backbone
MERT-v2-FullSong

Used at latent-cache time only. Layer-20 features of 24 kHz mono audio, resampled to 25 Hz, feed the community semantic head. The official audio-to-token encoder is unreleased.

632.4M
632,429,312
2.53 GBfp32no
Semantic head
Mothersuperior realaudio tokenizer v4 (community)

8-layer transformer classifier over 512-frame windows that maps MERT features to YuE2 codec tokens. A .pt file, so no parameter count. The repo also has matching NAR adapters (nar_lora_joint_v4 and later v5 / v8 / v9) that ai-toolkit can merge on load.

———no
Lead-sheet transcriber
SheetSage2

Audio → ABC sheet, MERT-v2-based encoder with its adapter already merged (source m-a-p/SheetSage2). Used at latent-cache time only, when cot is full or melody.

693.2M
693,240,089
1.39 GBbf16+fp32no
total5.10B11.72 GB

Audio latent space

sample rate
48 kHz
temporal
1,920×
channels
64
latent rate
25/s
autoencoder
YuE2 Oobleck-style VAE
audio channels
stereo
patch
1 latent steps per token
tokens per second
25.00
notes
Strides 2 × 2 × 4 × 4 × 5 × 6 = 1920, so 25 latent frames per second, the same rate as the codec tokens (one token per latent frame). The NAR adds a START and an END slot around the window.

Architecture

Experts
AR + NAR, each 28 layers, same shapes
Hidden size
2048 (16 heads × 128, 8 KV heads)
FFN size
6144 (SwiGLU, fused gate_up)
Vocabulary
184,704 (Qwen text + ABC markers + 32,768 codec tokens)
Context
24,576 tokens
Coupling
NAR layer i attends over its own tokens plus AR layer i cached keys/values for prefix + ABC + codec tokens + MUSIC_END (mixture of transformers)
Generation
AR writes an ABC lead sheet (cot full or melody, or none with off), then codec tokens; NAR renders chunks with flow matching
Objective
AR: next-token CE. NAR: rectified flow, timestep shift 1.0
Norm / position
RMSNorm, QK norm, RoPE (theta 1e6); NAR adds a frame position table

In AI Toolkit

model.arch
yue2
UI label
YuE2 (audio)
model.name_or_path
Comfy-Org/YuE2/checkpoints/yue2_3b_int8_convrot.safetensors
extra UI sections
model.low_vram, sample.duration

UI defaults

quantize
true (convrot8)
low_vram
false
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
sigmoid
datasets.cache_latents_to_disk
true
datasets.resolution
[512] (one bucket)
datasets.caption_dropout_rate
0
model_kwargs
cot: full, abc_dropout: 0.5, sample_ar_repetition_penalty: 1.2, ar_kl_weight: 0.2
sample
style + [Lyrics] prompt, up to 120 s, guidance 1, 32 steps
network.conv / model.quantize_te
disabled

Specifics

Status
Experimental. The AR memorizes small datasets within a few hundred steps (loss/ar_ce falls toward 0); it needs a large, varied dataset to learn a style.
Loading
Local file or org/repo/path, downloaded into the ComfyUI models folder. qtype convrot8 keeps the shipped int8 layers as they are. Embeddings, lm head, NAR in/out projections and the timestep embedder are never quantized.
Captions
A style line, then [Lyrics], then lyrics with [Verse 1] / [Chorus] headers (normalized to title case). Legacy <CAPTION>/<LYRICS> tags also parse. Do not use caption dropout: a blank prompt breaks lyric following.
Latent cache (required)
Per song it stores VAE latents, codec tokens from MERT + the semantic head, and (cot full/melody) a SheetSage2 ABC sheet, about 12 s per song. The cache records the cot mode; changing it means deleting _latent_cache. A non-default semantic head gets its own cache key. The encoders are dropped once training starts.
Training step
One LoRA covers both experts. The NAR gets the flow loss on a random window (train_window_frames 1500 = 60 s). The AR gets next-token CE over the codec tokens from the song start (ar_loss_weight 1.0), plus KL to the base model when ar_kl_weight > 0. The flow loss does not reach the AR through its cache.
Sheet dropout
abc_dropout (0.5) trains that fraction of items without the sheet, as an off-mode prompt, so one LoRA works with and without a lead sheet.
Separation
do_separation splits songs with MelBandRoformer at cache time and adds lyrics-only → vocals and tags-only → instrumental AR loss terms (weights 0.5 and 1.0).
AR learning rate
ar_lr_multiplier puts the AR expert LoRA in its own optimizer group.
LoRA target
YuE2AR, YuE2NAR
Saving
LoRA keys are rewritten for ComfyUI: NAR → diffusion_model.*, AR → text_encoders.*. Full fine-tunes cannot be saved.
Sampling
AR writes the sheet (up to 8192 tokens), then codec tokens up to the sample duration (repetition penalty from model_kwargs); NAR renders each context-sized chunk with sample_steps flow steps (32); tiled VAE decode. Output mp3, wav or flac.
Not supported
layer_offloading, full fine-tune saving

Example config

not verified

job: extensionconfig:  name: "my_yue2_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: bf16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/songs"          caption_ext: "txt"          caption_dropout_rate: 0          cache_latents_to_disk: true          resolution: [512]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "sigmoid"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16      model:        name_or_path: "Comfy-Org/YuE2/checkpoints/yue2_3b_int8_convrot.safetensors"        arch: "yue2"        quantize: true        qtype: "convrot8"        model_kwargs:          cot: "full"          abc_dropout: 0.5          sample_ar_repetition_penalty: 1.2          ar_kl_weight: 0.2      sample:        sampler: "flowmatch"        sample_every: 250        guidance_scale: 1        sample_steps: 32        duration: 120        prompts:          - "upbeat synth pop, female vocals\n[Lyrics]\n[Verse 1]\nHello world\n"

On this page