Qwen2.5-Omni 7B (thinker)
The thinker half of Qwen2.5-Omni 7B, a 7B Qwen2.5 language model with its own audio and vision encoders. AI Toolkit trains it as a captioner, with audio, image or video in and text out, using LoRA on the text stack only.
- weights
- ai-toolkit/Qwen2.5-Omni-7B ↗
- org
- Qwen (Alibaba)
- modality
- llm
- tasks
- audio-to-text · image-to-text · video-to-text
- license
- Apache 2.0
- released
- 2025-03-22not verified
- native output
- Text (the talker, which is not loaded, adds speech upstream)
- total params
- 8.93B
- model.arch
- qwen25_omni
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Thinker checkpoint | Qwen2.5-Omni 7B thinker, convrot int8 ai-toolkit default. A single file with only the thinker, in ComfyUI layout. Linear layers of the text stack and both encoders are int8 with convrot rotation; the count includes their quantization scales. The rows below are parts of this file, with bf16 parameter counts from the upstream repo. | 8.93B 8,933,786,616 | 10.46 GB | int8+bf16+fp32+uint8 | — |
| Language model | Qwen2.5-7B text stack transformers.Qwen2_5OmniThinkerTextModel 7,615,616,512 params: 7,070,619,136 for embeddings, 28 layers and the final norm, plus a 544,997,376-param untied lm head. LoRA goes on the layers; the lm head stays frozen. | — | — | — | yes |
| Audio encoder | Qwen2.5-Omni audio tower (Whisper-style) transformers.Qwen2_5OmniAudioEncoder 639,647,232 params. 128-bin log-mel input at 16 kHz, 32 layers, 1280 wide, projected to 3584. | — | — | — | no |
| Vision encoder | Qwen2.5-Omni vision tower (ViT) transformers.Qwen2_5OmniVisionEncoder 676,550,144 params. 32 blocks, 1280 wide, 14 px patches, 2×2 patch merger projecting to 3584. | — | — | — | no |
| Processor / tokenizer | Qwen2.5-Omni processor (Qwen2 BPE, Whisper feature extractor, image/video processor) transformers.AutoProcessor Not in the single file. ai-toolkit loads config and processor from Qwen/Qwen2.5-Omni-7B, picked by the checkpoint hidden size (2048 maps to Qwen/Qwen2.5-Omni-3B). | — | — | — | — |
| Full upstream model | Qwen2.5-Omni-7B (thinker + talker + token2wav)alternate file transformers.Qwen2_5OmniForConditionalGeneration Thinker 8,931,813,888, talker 1,351,360,256 and token2wav 449,051,296 params. Also loads as name_or_path (a folder or repo id), in which case ai-toolkit keeps only the thinker and quantizes it itself. The talker and token2wav (speech output) are never used. | 10.73B 10,732,225,440 | 22.36 GB | bf16+fp32 | no |
| total | 8.93B | 10.46 GB | |||
Architecture
- Layers
- 28
- Hidden size
- 3584 (28 heads × 128, 4 KV heads)
- FFN size
- 18944
- Vocabulary
- 152,064
- Context
- 32,768 tokens
- Position
- TMRoPE: 3-axis multimodal RoPE (sections 16 / 24 / 24), theta 1e6
- Modalities in
- Text, audio, image, video
- Modalities out
- Text (thinker); speech needs the talker and token2wav
- Audio tokens
- 25 per second (100 mel frames/s, conv stride 2, pooled 2×)
- Vision tokens
- One per 28×28 px after the 2×2 merge; video also pairs frames (temporal patch 2)
In AI Toolkit
- model.arch
- qwen25_omni
- UI label
- Qwen2.5-Omni (llm)
- model.name_or_path
- ai-toolkit/Qwen2.5-Omni-7B/qwen2_5_omni_7b_convrot8.safetensors
- source
- extra UI sections
- model.model_kwargs.instruction, sample.ctrl_img, datasets.num_frames
UI defaults
- quantize / quantize_te
- true (convrot8) / false
- low_vram
- false
- train.batch_size
- 1
- datasets.cache_latents_to_disk
- false
- datasets.resolution
- [512]
- datasets.caption_dropout_rate
- 0
- model_kwargs.instruction
- "Describe this in detail."
- sample
- instruction prompts, 1 step, guidance 1
- disabled sections
- network.conv, trigger_word, diff_output_preservation, blank_prompt_preservation, unload_text_encoder, slider
Specifics
- Training loop
- is_llm routes the trainer to train_llm_accumulation, which calls get_llm_loss: no noise, scheduler, VAE or prompt encoding. The loss is next-token cross-entropy on the caption only (loss/ce).
- Dataset
- Each item is a media file with a caption file; one dataset can mix audio (mp3, wav, flac, ogg, ...), images and video. Audio is resampled to 16 kHz and mixed to mono. Video is used when num_frames > 1, as frames only (no audio track).
- Sequence
- Chat template: system prompt, user turn with the media span and model_kwargs.instruction, assistant header, then the caption + <|im_end|> as the target.
- Media encoding
- The frozen audio and vision towers run on the GPU each step (log-mel is computed on the GPU; audio is cut at the processor chunk length). The latent cache is optional and stores the tower outputs (about 54 MB per 300 s of audio).
- Image budget
- max_pixels 512×512 and min_pixels 64×28×28 by default (model_kwargs). Buckets snap to multiples of 16.
- Loading
- A .safetensors name_or_path loads the thinker alone; a pre-quantized convrot8 file is kept as is whatever qtype says. Otherwise the full model loads and is cut down to the thinker, then quantized with quanto.
- LoRA target
- Qwen2_5OmniThinkerTextModel
- Saving
- LoRA keys are thinker-relative (model.layers.N...), with the transformer. prefix stripped. Full fine-tunes cannot be saved.
- Sampling
- sample.ctrl_img is the media file and the prompt is the instruction. Greedy decoding, up to 512 new tokens (max_new_tokens); the text is written as a .txt file.
- Not supported
- layer_offloading, batch_size > 1, full fine-tune saving
Example config
not verified
job: extensionconfig: name: "my_qwen25_omni_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: bf16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/media_and_captions" caption_ext: "txt" caption_dropout_rate: 0 cache_latents_to_disk: false resolution: [512] train: batch_size: 1 steps: 2000 gradient_checkpointing: true optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 model: name_or_path: "ai-toolkit/Qwen2.5-Omni-7B/qwen2_5_omni_7b_convrot8.safetensors" arch: "qwen25_omni" quantize: true qtype: "convrot8" model_kwargs: instruction: "Describe this in detail." sample: sample_every: 250 sample_steps: 1 guidance_scale: 1 samples: - prompt: "Describe this in detail." ctrl_img: "/path/to/test_song.mp3"Links
ACE-Step 1.5
A 2B diffusion transformer that writes full songs from a style caption, lyrics and metadata (BPM, key, time signature, duration). It works in the 48 kHz stereo latent space of a 1920× Oobleck VAE, 25 latent frames per second.
Zeta-Chroma
'A work-in-progress pixel-space model from lodestones (the Chroma team), built on the Z-Image transformer. It has no VAE: 32×32 RGB patches go straight into the trunk, and a per-token MLP decoder predicts the clean pixels.'