Docs
AI ToolkitModels

Boogu-Image 0.1 Edit

'The instruction-editing model of Boogu-Image 0.1: the same 10.3B Lumina2-style DiT, Qwen3-VL-8B encoder and FLUX.1 VAE as the Base, trained to edit. The reference image is read twice, by Qwen3-VL with the instruction and as VAE latents through a dedicated refiner.'

org
Boogu
modality
image
tasks
image-editing
license
Apache 2.0
released
2026-06-16
native output
1K, 1.5K or 2K; most stable at 1K (Boogu README)
total params
19.14B
model.arch
boogu_image_edit

Components

rolemodelparamssizedtypetrained
Transformer
Boogu-Image 0.1 Edit DiT
BooguImageTransformer2DModel (ai-toolkit boogu_image/src/transformer.py)

Same architecture and parameter count as the Base, different weights. The -fp8 sibling repo ships torchao float8 .bin weights, which ai-toolkit cannot load.

10.29B
10,292,556,288
20.59 GBbf16yes
Text encoder
Qwen3-VL-8B-Instruct
transformers.Qwen3VLModel

Byte-identical to Qwen/Qwen3-VL-8B-Instruct (same sha256 for all four shards). ai-toolkit loads the inner Qwen3VLModel from mllm/ and uses its last hidden state. The vision tower encodes the reference images.

8.77B
8,767,123,696
17.53 GBbf16no
Processor
Qwen3-VL processor and tokenizer
transformers.AutoProcessor
————
VAE
FLUX.1 VAE
diffusers.AutoencoderKL

Its config names FLUX.1-dev as the source. Scaling factor 0.3611, shift factor 0.1159.

83.8M
83,819,683
335 MBfp32no
total19.14B38.45 GB

Latent space

spatial
8×
channels
16
patch
2×2
autoencoder
FLUX.1 VAE
pixels per token
16×16
notes
Latents are shifted by 0.1159 and scaled by 0.3611, both read from the VAE config. Reference latents add their own tokens on top of the counts below.
inputlatent (c×t×h×w)tokens
1024×102416×1×128×1284,0961K, ai-toolkit sample default
1536×153616×1×192×1929,2161.5K
2048×204816×1×256×25616,3842K

Architecture

Type
Lumina2-style DiT: double-stream layers, then single-stream layers on the joint sequence
Layers
40 (8 double-stream + 32 single-stream)
Refiners
2 blocks each for text, noisy image and reference images, before the main stack
Image conditioning
Reference images go into Qwen3-VL with the instruction, and as VAE latents through the reference refiner into the image stream; each reference sits after the text on RoPE axis 0
Hidden size
3360 (28 query heads × 120, 7 KV heads)
FFN size
13568 (SwiGLU)
Text conditioning
Last hidden state of Qwen3-VL-8B (4096-d) over a chat template with Boogu’s edit system prompt and the reference images
Norm / position
3-axis RoPE (40/40/40 dims, θ 10000); text positions run along all axes, images start after the text
Objective
Rectified flow. Native time runs 0 = noise to 1 = clean; the model predicts clean − noise
Recommended sampling
25–50 steps, CFG 2–5, e.g. 5.0 (Boogu README)

In AI Toolkit

model.arch
boogu_image_edit
UI label
Boogu Image Edit (instruction)
model.name_or_path
Boogu/Boogu-Image-0.1-Edit
extra UI sections
datasets.multi_control_paths, sample.multi_ctrl_imgs, model.low_vram, model.layer_offloading, model.qie.match_target_res

UI defaults

quantize / quantize_te
true / true (qfloat8)
low_vram
true
timestep_type
linear
model_kwargs
{ match_target_res: false }
train.unload_text_encoder
false (section hidden)
network.conv
disabled (linear LoRA only)

Specifics

Loading
transformer/, vae/, processor/ and mllm/ all come from name_or_path. model_kwargs.text_encoder_path and text_encoder_subfolder override the text encoder location. Use the bf16 repo, not -fp8; set quantize for fp8.
Training data
Each item is a target image, an edit instruction as its caption, and reference images (control_path_1 to _3). A reference is required: prompt encoding fails without one. Boogu’s README says the released model supports one reference image for now.
Reference path 1: Qwen3-VL
References are downscaled (never up) to fit 384² pixels and a 768 px long side (model_kwargs.vlm_max_pixels, vlm_max_side_length), snapped down to 16 px, and placed before the instruction in the user message. The UI keeps the text encoder loaded; cached text embeddings are keyed on the control paths and use a separate cache version (boogu_image_edit_v1).
Reference path 2: latents
Each reference keeps its aspect ratio and is capped at 1 MP (model_kwargs.control_image_max_pixels), or resized to the target’s pixel area with match_target_res (off by default here), then snapped to 16 px and VAE-encoded. It goes through the reference refiner, not the loss.
Time convention
Boogu’s time runs the other way from ai-toolkit’s, so the timestep is flipped (t = 1 − timestep / 1000) and the prediction negated to give the usual noise − clean target.
Text encoding
Each instruction is put in a system + user chat template and encoded at its natural length, capped at 1024 tokens (model_kwargs.max_text_length). Padding to the batch max happens at the model call.
Resolution
Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
Timesteps
Training uses flowmatch with a static shift of 3.0. Boogu’s own resolution-dependent shift (μ 0.5 at 256 tokens to 1.15 at 4096) is only used by the preview sampler.
Attention
PyTorch SDPA by default. model_kwargs.attention_backend: "flash" switches to Flash Attention 2.
LoRA target
BooguImageTransformer2DModel
Saving
LoRA keys are saved with the diffusion_model. prefix (ComfyUI layout). Full fine-tunes save a Diffusers transformer/ folder with its config.json, plus aitk_meta.yaml.
Sampling
Built-in Euler sampler; CFG turns on above guidance_scale 1.
Metadata base version
boogu_image_edit.0.1

Example config

not verified

job: extensionconfig:  name: "my_boogu_image_edit_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: bf16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/targets"          control_path_1: "/path/to/references"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "linear"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "Boogu/Boogu-Image-0.1-Edit"        arch: "boogu_image_edit"        quantize: true        quantize_te: true        low_vram: true        model_kwargs:          match_target_res: false      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 5        sample_steps: 30        samples:          - prompt: "replace the sky with a starry night"            ctrl_img_1: "/path/to/reference.jpg"

On this page