Docs
AI ToolkitModels

Boogu-Image 0.1 Base

'The undistilled text-to-image base of Boogu-Image 0.1: a 10.3B Lumina2-style DiT with 8 double-stream and 32 single-stream layers, conditioned on Qwen3-VL-8B and decoding through the FLUX.1 VAE. Its authors pitch it for fine-tuning and dense Chinese and English text rendering.'

org
Boogu
modality
image
tasks
text-to-image
license
Apache 2.0
released
2026-06-16
native output
1K, 1.5K or 2K (Boogu README)
total params
19.14B
model.arch
boogu_image

Components

rolemodelparamssizedtypetrained
Transformer
Boogu-Image 0.1 Base DiT
BooguImageTransformer2DModel (ai-toolkit boogu_image/src/transformer.py)

Includes the reference-image refiner used by the Edit model. The -fp8 sibling repo ships torchao float8 .bin weights, which ai-toolkit cannot load.

10.29B
10,292,556,288
20.59 GBbf16yes
Text encoder
Qwen3-VL-8B-Instruct
transformers.Qwen3VLModel

Byte-identical to Qwen/Qwen3-VL-8B-Instruct (same sha256 for all four shards). ai-toolkit loads the inner Qwen3VLModel from mllm/ and uses its last hidden state. The vision tower is loaded but never run for text-to-image.

8.77B
8,767,123,696
17.53 GBbf16no
Processor
Qwen3-VL processor and tokenizer
transformers.AutoProcessor
————
VAE
FLUX.1 VAE
diffusers.AutoencoderKL

Its config names FLUX.1-dev as the source. Scaling factor 0.3611, shift factor 0.1159.

83.8M
83,819,683
335 MBfp32no
total19.14B38.45 GB

Latent space

spatial
8×
channels
16
patch
2×2
autoencoder
FLUX.1 VAE
pixels per token
16×16
notes
Latents are shifted by 0.1159 and scaled by 0.3611, both read from the VAE config.
inputlatent (c×t×h×w)tokens
1024×102416×1×128×1284,0961K, ai-toolkit sample default
1536×153616×1×192×1929,2161.5K
2048×204816×1×256×25616,3842K, suggested for dense text

Architecture

Type
Lumina2-style DiT: double-stream layers, then single-stream layers on the joint sequence
Layers
40 (8 double-stream + 32 single-stream)
Refiners
2 blocks each for text, noisy image and reference images, before the main stack
Hidden size
3360 (28 query heads × 120, 7 KV heads)
FFN size
13568 (SwiGLU)
Text conditioning
Last hidden state of Qwen3-VL-8B (4096-d) over a chat template with Boogu’s system prompt
Norm / position
3-axis RoPE (40/40/40 dims, θ 10000); text positions run along all axes, images start after the text
Objective
Rectified flow. Native time runs 0 = noise to 1 = clean; the model predicts clean − noise
Recommended sampling
25–50 steps, CFG 2–5, e.g. 4.0 (Boogu README)

In AI Toolkit

model.arch
boogu_image
UI label
Boogu Image (image)
model.name_or_path
Boogu/Boogu-Image-0.1-Base
extra UI sections
model.low_vram, model.layer_offloading

UI defaults

quantize / quantize_te
true / true (qfloat8)
low_vram
true
timestep_type
linear
network.conv
disabled (linear LoRA only)

Specifics

Loading
transformer/, vae/, processor/ and mllm/ all come from name_or_path. model_kwargs.text_encoder_path and text_encoder_subfolder override the text encoder location. Use the bf16 repo, not -fp8; set quantize for fp8.
Time convention
Boogu’s time runs the other way from ai-toolkit’s, so the timestep is flipped (t = 1 − timestep / 1000) and the prediction negated to give the usual noise − clean target.
Text encoding
Each caption is put in a system + user chat template and encoded at its natural length, capped at 1024 tokens (model_kwargs.max_text_length). Padding to the batch max happens at the model call.
Resolution
Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
Timesteps
Training uses flowmatch with a static shift of 3.0. Boogu’s own resolution-dependent shift (μ 0.5 at 256 tokens to 1.15 at 4096) is only used by the preview sampler.
Attention
PyTorch SDPA by default. model_kwargs.attention_backend: "flash" switches to Flash Attention 2.
LoRA target
BooguImageTransformer2DModel
Saving
LoRA keys are saved with the diffusion_model. prefix (ComfyUI layout). Full fine-tunes save a Diffusers transformer/ folder with its config.json, plus aitk_meta.yaml.
Sampling
Built-in Euler sampler; CFG turns on above guidance_scale 1.
Metadata base version
boogu_image.0.1

Example config

not verified

job: extensionconfig:  name: "my_boogu_image_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: bf16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "linear"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "Boogu/Boogu-Image-0.1-Base"        arch: "boogu_image"        quantize: true        quantize_te: true        low_vram: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 4        sample_steps: 30        prompts:          - "a travel poster of a mountain lake with the title BOOGU in bold letters"

On this page