Boogu-Image 0.1 Base
'The undistilled text-to-image base of Boogu-Image 0.1: a 10.3B Lumina2-style DiT with 8 double-stream and 32 single-stream layers, conditioned on Qwen3-VL-8B and decoding through the FLUX.1 VAE. Its authors pitch it for fine-tuning and dense Chinese and English text rendering.'
- weights
- Boogu/Boogu-Image-0.1-Base ↗
- org
- Boogu
- modality
- image
- tasks
- text-to-image
- license
- Apache 2.0
- released
- 2026-06-16
- native output
- 1K, 1.5K or 2K (Boogu README)
- total params
- 19.14B
- model.arch
- boogu_image
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Boogu-Image 0.1 Base DiT BooguImageTransformer2DModel (ai-toolkit boogu_image/src/transformer.py) Includes the reference-image refiner used by the Edit model. The -fp8 sibling repo ships torchao float8 .bin weights, which ai-toolkit cannot load. | 10.29B 10,292,556,288 | 20.59 GB | bf16 | yes |
| Text encoder | Qwen3-VL-8B-Instruct transformers.Qwen3VLModel Byte-identical to Qwen/Qwen3-VL-8B-Instruct (same sha256 for all four shards). ai-toolkit loads the inner Qwen3VLModel from mllm/ and uses its last hidden state. The vision tower is loaded but never run for text-to-image. | 8.77B 8,767,123,696 | 17.53 GB | bf16 | no |
| Processor | Qwen3-VL processor and tokenizer transformers.AutoProcessor | — | — | — | — |
| VAE | FLUX.1 VAE diffusers.AutoencoderKL Its config names FLUX.1-dev as the source. Scaling factor 0.3611, shift factor 0.1159. | 83.8M 83,819,683 | 335 MB | fp32 | no |
| total | 19.14B | 38.45 GB | |||
Latent space
- autoencoder
- FLUX.1 VAE
- pixels per token
- 16×16
- notes
- Latents are shifted by 0.1159 and scaled by 0.3611, both read from the VAE config.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 | 16×1×128×128 | 4,096 | 1K, ai-toolkit sample default |
| 1536×1536 | 16×1×192×192 | 9,216 | 1.5K |
| 2048×2048 | 16×1×256×256 | 16,384 | 2K, suggested for dense text |
Architecture
- Type
- Lumina2-style DiT: double-stream layers, then single-stream layers on the joint sequence
- Layers
- 40 (8 double-stream + 32 single-stream)
- Refiners
- 2 blocks each for text, noisy image and reference images, before the main stack
- Hidden size
- 3360 (28 query heads × 120, 7 KV heads)
- FFN size
- 13568 (SwiGLU)
- Text conditioning
- Last hidden state of Qwen3-VL-8B (4096-d) over a chat template with Boogu’s system prompt
- Norm / position
- 3-axis RoPE (40/40/40 dims, θ 10000); text positions run along all axes, images start after the text
- Objective
- Rectified flow. Native time runs 0 = noise to 1 = clean; the model predicts clean − noise
- Recommended sampling
- 25–50 steps, CFG 2–5, e.g. 4.0 (Boogu README)
In AI Toolkit
- model.arch
- boogu_image
- UI label
- Boogu Image (image)
- model.name_or_path
- Boogu/Boogu-Image-0.1-Base
- source
- extensions_built_in/diffusion_models/boogu_image/boogu_image.py
- extensions_built_in/diffusion_models/boogu_image/src/transformer.py
- extensions_built_in/diffusion_models/boogu_image/src/pipeline.py
- extensions_built_in/diffusion_models/boogu_image/src/rope.py
- extensions_built_in/diffusion_models/boogu_image/src/block_lumina2.py
- extensions_built_in/diffusion_models/boogu_image/src/attention_processor.py
- extensions_built_in/diffusion_models/boogu_image/src/embeddings.py
- toolkit/models/v2/text_encoders/qwen3_vl.py
- toolkit/models/v2/vae/autoencoder_kl.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- model.low_vram, model.layer_offloading
UI defaults
- quantize / quantize_te
- true / true (qfloat8)
- low_vram
- true
- timestep_type
- linear
- network.conv
- disabled (linear LoRA only)
Specifics
- Loading
- transformer/, vae/, processor/ and mllm/ all come from name_or_path. model_kwargs.text_encoder_path and text_encoder_subfolder override the text encoder location. Use the bf16 repo, not -fp8; set quantize for fp8.
- Time convention
- Boogu’s time runs the other way from ai-toolkit’s, so the timestep is flipped (t = 1 − timestep / 1000) and the prediction negated to give the usual noise − clean target.
- Text encoding
- Each caption is put in a system + user chat template and encoded at its natural length, capped at 1024 tokens (model_kwargs.max_text_length). Padding to the batch max happens at the model call.
- Resolution
- Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
- Timesteps
- Training uses flowmatch with a static shift of 3.0. Boogu’s own resolution-dependent shift (μ 0.5 at 256 tokens to 1.15 at 4096) is only used by the preview sampler.
- Attention
- PyTorch SDPA by default. model_kwargs.attention_backend: "flash" switches to Flash Attention 2.
- LoRA target
- BooguImageTransformer2DModel
- Saving
- LoRA keys are saved with the diffusion_model. prefix (ComfyUI layout). Full fine-tunes save a Diffusers transformer/ folder with its config.json, plus aitk_meta.yaml.
- Sampling
- Built-in Euler sampler; CFG turns on above guidance_scale 1.
- Metadata base version
- boogu_image.0.1
Example config
not verified
job: extensionconfig: name: "my_boogu_image_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: bf16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "linear" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "Boogu/Boogu-Image-0.1-Base" arch: "boogu_image" quantize: true quantize_te: true low_vram: true sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 4 sample_steps: 30 prompts: - "a travel poster of a mountain lake with the title BOOGU in bold letters"Links
Mage-Flow Base
'The undistilled text-to-image base of Microsoft’s Mage-Flow: a 4.1B dual-stream NR-MMDiT on Qwen3-VL-4B text features, working in the 128-channel, 16× latent space of Mage-VAE with one token per latent pixel. Built to generate at native resolution from 512 to 2048 px.'
Flex.2-preview
The follow-up to Flex.1-alpha: the same 8B FLUX-style transformer with inpainting and a universal control input (line, pose, depth) trained into the base model. The extra inputs are concatenated on the latent channels, so the transformer takes 196 input channels per token instead of 64.