Docs
AI ToolkitModels

Qwen2.5-Omni 7B (thinker)

The thinker half of Qwen2.5-Omni 7B, a 7B Qwen2.5 language model with its own audio and vision encoders. AI Toolkit trains it as a captioner, with audio, image or video in and text out, using LoRA on the text stack only.

org
Qwen (Alibaba)
modality
llm
tasks
audio-to-text · image-to-text · video-to-text
license
Apache 2.0
released
2025-03-22not verified
native output
Text (the talker, which is not loaded, adds speech upstream)
total params
8.93B
model.arch
qwen25_omni

Components

rolemodelparamssizedtypetrained
Thinker checkpoint
Qwen2.5-Omni 7B thinker, convrot int8

ai-toolkit default. A single file with only the thinker, in ComfyUI layout. Linear layers of the text stack and both encoders are int8 with convrot rotation; the count includes their quantization scales. The rows below are parts of this file, with bf16 parameter counts from the upstream repo.

8.93B
8,933,786,616
10.46 GBint8+bf16+fp32+uint8—
Language model
Qwen2.5-7B text stack
transformers.Qwen2_5OmniThinkerTextModel

7,615,616,512 params: 7,070,619,136 for embeddings, 28 layers and the final norm, plus a 544,997,376-param untied lm head. LoRA goes on the layers; the lm head stays frozen.

———yes
Audio encoder
Qwen2.5-Omni audio tower (Whisper-style)
transformers.Qwen2_5OmniAudioEncoder

639,647,232 params. 128-bin log-mel input at 16 kHz, 32 layers, 1280 wide, projected to 3584.

———no
Vision encoder
Qwen2.5-Omni vision tower (ViT)
transformers.Qwen2_5OmniVisionEncoder

676,550,144 params. 32 blocks, 1280 wide, 14 px patches, 2×2 patch merger projecting to 3584.

———no
Processor / tokenizer
Qwen2.5-Omni processor (Qwen2 BPE, Whisper feature extractor, image/video processor)
transformers.AutoProcessor

Not in the single file. ai-toolkit loads config and processor from Qwen/Qwen2.5-Omni-7B, picked by the checkpoint hidden size (2048 maps to Qwen/Qwen2.5-Omni-3B).

————
Full upstream model
Qwen2.5-Omni-7B (thinker + talker + token2wav)alternate file
transformers.Qwen2_5OmniForConditionalGeneration

Thinker 8,931,813,888, talker 1,351,360,256 and token2wav 449,051,296 params. Also loads as name_or_path (a folder or repo id), in which case ai-toolkit keeps only the thinker and quantizes it itself. The talker and token2wav (speech output) are never used.

10.73B
10,732,225,440
22.36 GBbf16+fp32no
total8.93B10.46 GB

Architecture

Layers
28
Hidden size
3584 (28 heads × 128, 4 KV heads)
FFN size
18944
Vocabulary
152,064
Context
32,768 tokens
Position
TMRoPE: 3-axis multimodal RoPE (sections 16 / 24 / 24), theta 1e6
Modalities in
Text, audio, image, video
Modalities out
Text (thinker); speech needs the talker and token2wav
Audio tokens
25 per second (100 mel frames/s, conv stride 2, pooled 2×)
Vision tokens
One per 28×28 px after the 2×2 merge; video also pairs frames (temporal patch 2)

In AI Toolkit

model.arch
qwen25_omni
UI label
Qwen2.5-Omni (llm)
model.name_or_path
ai-toolkit/Qwen2.5-Omni-7B/qwen2_5_omni_7b_convrot8.safetensors
extra UI sections
model.model_kwargs.instruction, sample.ctrl_img, datasets.num_frames

UI defaults

quantize / quantize_te
true (convrot8) / false
low_vram
false
train.batch_size
1
datasets.cache_latents_to_disk
false
datasets.resolution
[512]
datasets.caption_dropout_rate
0
model_kwargs.instruction
"Describe this in detail."
sample
instruction prompts, 1 step, guidance 1
disabled sections
network.conv, trigger_word, diff_output_preservation, blank_prompt_preservation, unload_text_encoder, slider

Specifics

Training loop
is_llm routes the trainer to train_llm_accumulation, which calls get_llm_loss: no noise, scheduler, VAE or prompt encoding. The loss is next-token cross-entropy on the caption only (loss/ce).
Dataset
Each item is a media file with a caption file; one dataset can mix audio (mp3, wav, flac, ogg, ...), images and video. Audio is resampled to 16 kHz and mixed to mono. Video is used when num_frames > 1, as frames only (no audio track).
Sequence
Chat template: system prompt, user turn with the media span and model_kwargs.instruction, assistant header, then the caption + <|im_end|> as the target.
Media encoding
The frozen audio and vision towers run on the GPU each step (log-mel is computed on the GPU; audio is cut at the processor chunk length). The latent cache is optional and stores the tower outputs (about 54 MB per 300 s of audio).
Image budget
max_pixels 512×512 and min_pixels 64×28×28 by default (model_kwargs). Buckets snap to multiples of 16.
Loading
A .safetensors name_or_path loads the thinker alone; a pre-quantized convrot8 file is kept as is whatever qtype says. Otherwise the full model loads and is cut down to the thinker, then quantized with quanto.
LoRA target
Qwen2_5OmniThinkerTextModel
Saving
LoRA keys are thinker-relative (model.layers.N...), with the transformer. prefix stripped. Full fine-tunes cannot be saved.
Sampling
sample.ctrl_img is the media file and the prompt is the instruction. Greedy decoding, up to 512 new tokens (max_new_tokens); the text is written as a .txt file.
Not supported
layer_offloading, batch_size > 1, full fine-tune saving

Example config

not verified

job: extensionconfig:  name: "my_qwen25_omni_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: bf16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/media_and_captions"          caption_ext: "txt"          caption_dropout_rate: 0          cache_latents_to_disk: false          resolution: [512]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16      model:        name_or_path: "ai-toolkit/Qwen2.5-Omni-7B/qwen2_5_omni_7b_convrot8.safetensors"        arch: "qwen25_omni"        quantize: true        qtype: "convrot8"        model_kwargs:          instruction: "Describe this in detail."      sample:        sample_every: 250        sample_steps: 1        guidance_scale: 1        samples:          - prompt: "Describe this in detail."            ctrl_img: "/path/to/test_song.mp3"

On this page