Boogu-Image 0.1 Edit
'The instruction-editing model of Boogu-Image 0.1: the same 10.3B Lumina2-style DiT, Qwen3-VL-8B encoder and FLUX.1 VAE as the Base, trained to edit. The reference image is read twice, by Qwen3-VL with the instruction and as VAE latents through a dedicated refiner.'
- weights
- Boogu/Boogu-Image-0.1-Edit ↗
- org
- Boogu
- modality
- image
- tasks
- image-editing
- license
- Apache 2.0
- released
- 2026-06-16
- native output
- 1K, 1.5K or 2K; most stable at 1K (Boogu README)
- total params
- 19.14B
- model.arch
- boogu_image_edit
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Boogu-Image 0.1 Edit DiT BooguImageTransformer2DModel (ai-toolkit boogu_image/src/transformer.py) Same architecture and parameter count as the Base, different weights. The -fp8 sibling repo ships torchao float8 .bin weights, which ai-toolkit cannot load. | 10.29B 10,292,556,288 | 20.59 GB | bf16 | yes |
| Text encoder | Qwen3-VL-8B-Instruct transformers.Qwen3VLModel Byte-identical to Qwen/Qwen3-VL-8B-Instruct (same sha256 for all four shards). ai-toolkit loads the inner Qwen3VLModel from mllm/ and uses its last hidden state. The vision tower encodes the reference images. | 8.77B 8,767,123,696 | 17.53 GB | bf16 | no |
| Processor | Qwen3-VL processor and tokenizer transformers.AutoProcessor | — | — | — | — |
| VAE | FLUX.1 VAE diffusers.AutoencoderKL Its config names FLUX.1-dev as the source. Scaling factor 0.3611, shift factor 0.1159. | 83.8M 83,819,683 | 335 MB | fp32 | no |
| total | 19.14B | 38.45 GB | |||
Latent space
- autoencoder
- FLUX.1 VAE
- pixels per token
- 16×16
- notes
- Latents are shifted by 0.1159 and scaled by 0.3611, both read from the VAE config. Reference latents add their own tokens on top of the counts below.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 | 16×1×128×128 | 4,096 | 1K, ai-toolkit sample default |
| 1536×1536 | 16×1×192×192 | 9,216 | 1.5K |
| 2048×2048 | 16×1×256×256 | 16,384 | 2K |
Architecture
- Type
- Lumina2-style DiT: double-stream layers, then single-stream layers on the joint sequence
- Layers
- 40 (8 double-stream + 32 single-stream)
- Refiners
- 2 blocks each for text, noisy image and reference images, before the main stack
- Image conditioning
- Reference images go into Qwen3-VL with the instruction, and as VAE latents through the reference refiner into the image stream; each reference sits after the text on RoPE axis 0
- Hidden size
- 3360 (28 query heads × 120, 7 KV heads)
- FFN size
- 13568 (SwiGLU)
- Text conditioning
- Last hidden state of Qwen3-VL-8B (4096-d) over a chat template with Boogu’s edit system prompt and the reference images
- Norm / position
- 3-axis RoPE (40/40/40 dims, θ 10000); text positions run along all axes, images start after the text
- Objective
- Rectified flow. Native time runs 0 = noise to 1 = clean; the model predicts clean − noise
- Recommended sampling
- 25–50 steps, CFG 2–5, e.g. 5.0 (Boogu README)
In AI Toolkit
- model.arch
- boogu_image_edit
- UI label
- Boogu Image Edit (instruction)
- model.name_or_path
- Boogu/Boogu-Image-0.1-Edit
- source
- extensions_built_in/diffusion_models/boogu_image/boogu_image_edit.py
- extensions_built_in/diffusion_models/boogu_image/boogu_image.py
- extensions_built_in/diffusion_models/boogu_image/src/transformer.py
- extensions_built_in/diffusion_models/boogu_image/src/pipeline.py
- extensions_built_in/diffusion_models/boogu_image/src/rope.py
- extensions_built_in/diffusion_models/boogu_image/src/block_lumina2.py
- extensions_built_in/diffusion_models/boogu_image/src/attention_processor.py
- extensions_built_in/diffusion_models/boogu_image/src/embeddings.py
- toolkit/models/v2/text_encoders/qwen3_vl.py
- toolkit/models/v2/vae/autoencoder_kl.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- datasets.multi_control_paths, sample.multi_ctrl_imgs, model.low_vram, model.layer_offloading, model.qie.match_target_res
UI defaults
- quantize / quantize_te
- true / true (qfloat8)
- low_vram
- true
- timestep_type
- linear
- model_kwargs
- { match_target_res: false }
- train.unload_text_encoder
- false (section hidden)
- network.conv
- disabled (linear LoRA only)
Specifics
- Loading
- transformer/, vae/, processor/ and mllm/ all come from name_or_path. model_kwargs.text_encoder_path and text_encoder_subfolder override the text encoder location. Use the bf16 repo, not -fp8; set quantize for fp8.
- Training data
- Each item is a target image, an edit instruction as its caption, and reference images (control_path_1 to _3). A reference is required: prompt encoding fails without one. Boogu’s README says the released model supports one reference image for now.
- Reference path 1: Qwen3-VL
- References are downscaled (never up) to fit 384² pixels and a 768 px long side (model_kwargs.vlm_max_pixels, vlm_max_side_length), snapped down to 16 px, and placed before the instruction in the user message. The UI keeps the text encoder loaded; cached text embeddings are keyed on the control paths and use a separate cache version (boogu_image_edit_v1).
- Reference path 2: latents
- Each reference keeps its aspect ratio and is capped at 1 MP (model_kwargs.control_image_max_pixels), or resized to the target’s pixel area with match_target_res (off by default here), then snapped to 16 px and VAE-encoded. It goes through the reference refiner, not the loss.
- Time convention
- Boogu’s time runs the other way from ai-toolkit’s, so the timestep is flipped (t = 1 − timestep / 1000) and the prediction negated to give the usual noise − clean target.
- Text encoding
- Each instruction is put in a system + user chat template and encoded at its natural length, capped at 1024 tokens (model_kwargs.max_text_length). Padding to the batch max happens at the model call.
- Resolution
- Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
- Timesteps
- Training uses flowmatch with a static shift of 3.0. Boogu’s own resolution-dependent shift (μ 0.5 at 256 tokens to 1.15 at 4096) is only used by the preview sampler.
- Attention
- PyTorch SDPA by default. model_kwargs.attention_backend: "flash" switches to Flash Attention 2.
- LoRA target
- BooguImageTransformer2DModel
- Saving
- LoRA keys are saved with the diffusion_model. prefix (ComfyUI layout). Full fine-tunes save a Diffusers transformer/ folder with its config.json, plus aitk_meta.yaml.
- Sampling
- Built-in Euler sampler; CFG turns on above guidance_scale 1.
- Metadata base version
- boogu_image_edit.0.1
Example config
not verified
job: extensionconfig: name: "my_boogu_image_edit_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: bf16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/targets" control_path_1: "/path/to/references" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "linear" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "Boogu/Boogu-Image-0.1-Edit" arch: "boogu_image_edit" quantize: true quantize_te: true low_vram: true model_kwargs: match_target_res: false sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 5 sample_steps: 30 samples: - prompt: "replace the sky with a starry night" ctrl_img_1: "/path/to/reference.jpg"Links
Mage-Flow Edit Base
'The undistilled instruction-editing base of Microsoft’s Mage-Flow: the same 4.1B dual-stream NR-MMDiT and Mage-VAE as Mage-Flow Base, trained to edit. Reference images reach the model twice, through Qwen3-VL with the instruction and as clean latents packed after the target.'
Wan 2.1 T2V 1.3B
'The small text-to-video member of Wan 2.1. A 1.4B-parameter DiT in the 8× spatial / 4× temporal latent space of the Wan 2.1 VAE, with UMT5-XXL text conditioning. The cheapest Wan to train: it fits on a consumer GPU without quantizing the transformer.'