Z-Image L2P (pixel space)
'A pixel-space version of Z-Image made with the Latent-to-Pixel (L2P) transfer method: the VAE is gone, 16×16 RGB patches go straight into the Z-Image trunk, and a small U-Net decoder turns the transformer features back into pixels.'
- weights
- zhen-nan/L2P ↗
- org
- zhen-nan (NJU PCALab)not verified
- modality
- image
- tasks
- text-to-image
- license
- Apache 2.0
- released
- 2026-05-03not verified
- total params
- 10.19B
- model.arch
- zimage_l2p
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Z-Image L2P (trunk + pixel decoder, single file) ZImageTransformer2DModel (ai-toolkit L2P subclass) One file: the 30-layer Z-Image trunk and refiners, a 16×16 pixel patch embedder (768 → 3840) and a 10,085,123-parameter U-Net pixel decoder (local_decoder). The Z-Image final layer is not in the file. Main layers 5–24 are stored in fp32 (3,618,206,720 params), everything else in bf16; ai-toolkit casts all of it to the training dtype on load. | 6.17B 6,166,464,515 | 19.57 GB | fp32+bf16 | yes |
| Text encoder | Qwen3-4B transformers.Qwen3ForCausalLM Not in the L2P repo. ai-toolkit loads it from extras_name_or_path, Tongyi-MAI/Z-Image-Turbo. Only the second-to-last hidden state is used. | 4.02B 4,022,468,096 | 8.04 GB | bf16 | no |
| Tokenizer | Qwen3 BPE tokenizer transformers.Qwen2Tokenizer | — | — | — | — |
| total | 10.19B | 27.61 GB | |||
Latent space
- pixels per token
- 16×16
- notes
- No autoencoder: the model denoises RGB pixels directly (ai-toolkit uses an identity FakeVAE with scaling 1.0). Each token is one 16×16×3 pixel patch (768 values), so the token grid matches latent Z-Image at the same resolution (8× VAE × 2×2 patch).
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 512×512 | 3×1×512×512 | 1,024 | |
| 1024×1024 | 3×1×1024×1024 | 4,096 | ai-toolkit sample default |
| 2048×2048 | 3×1×2048×2048 | 16,384 |
Architecture
- Blocks
- Z-Image trunk: 30 single-stream blocks, 2 noise-refiner and 2 context-refiner blocks
- Hidden size
- 3840 (30 heads × 128)
- Input
- Linear embedder on 16×16 RGB patches (768 → 3840) replaces the latent 2×2 patch embedder
- Output
- 4-stage conv U-Net (64/128/256/512 channels, 16× down) on the noisy image, with the transformer feature map (one 3840-d vector per patch) fused at the bottleneck. Replaces the Z-Image final layer
- Text conditioning
- Same as Z-Image: Qwen3-4B second-to-last hidden states, up to 512 tokens, in one sequence with the image tokens
- Objective
- Rectified flow in pixel space (ai-toolkit shift 3.0)
- Norm / position
- RMSNorm, QK RMSNorm, 3-axis RoPE [32, 48, 48], theta 256
In AI Toolkit
- model.arch
- zimage_l2p
- UI label
- Z-Image L2P (pixel space) (image)
- model.name_or_path
- zhen-nan/L2P/model-1k-merge.safetensors
- source
- extra UI sections
- model.low_vram, model.layer_offloading
UI defaults
- extras_name_or_path
- Tongyi-MAI/Z-Image-Turbo
- quantize / quantize_te
- true / true
- low_vram
- true
- timestep_type
- linear
- network.conv
- disabled (linear LoRA only)
Specifics
- Checkpoint download
- An org/repo/file.safetensors name_or_path is downloaded once to MODELS_PATH/diffusion_models and reused from there.
- Single-file loading
- No config ships with the file. ai-toolkit builds the model from a built-in Z-Image config, detects pixel space from local_decoder.* or all_x_embedder.16-1 keys, and loads the weights non-strictly.
- Latent-to-pixel conversion
- Point name_or_path at a latent Z-Image checkpoint (in_channels 16) and it is converted on load: a new 16×16 patch embedder (random × 0.001), the final layer dropped, and a randomly initialised U-Net decoder. That starts an L2P transfer, not a finished model.
- VAE
- None. An identity FakeVAE (scaling 1.0) stands in, so cached “latents” are pixels.
- Text encoder source
- text_encoder/ and tokenizer/ from extras_name_or_path. Never trained.
- Resolution
- Buckets snap to multiples of 16 (inherited from Z-Image; also the pixel patch size).
- Timesteps and loss
- Inherited from Z-Image: model t = (1000 − timestep) / 1000, output negated, target noise − pixels, scheduler shift 3.0.
- Sampling
- Z-Image pipeline with the identity VAE. guidance_scale is shifted down by 1. The inference engine registry defaults to 8 steps, guidance 1.
- LoRA target
- ZImageTransformer2DModel
- Saving
- Full fine-tunes save the raw state dict (Diffusers key names, not the ComfyUI layout) minus all_final_layer.*, cast to the save dtype.
- Metadata base version
- zimage
Example config
not verified
job: extensionconfig: name: "my_zimage_l2p_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: bf16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "linear" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 model: name_or_path: "zhen-nan/L2P/model-1k-merge.safetensors" extras_name_or_path: "Tongyi-MAI/Z-Image-Turbo" arch: "zimage_l2p" quantize: true quantize_te: true low_vram: true sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 1 sample_steps: 8 prompts: - "a bear building a log cabin in the snow covered mountains"Links
HiDream-O1-Image
An 8B pixel-space image model built inside a Qwen3-VL language model. Text, timestep and 32×32 pixel patches share one token sequence in the same transformer, with no VAE and no separate text encoder. Generates up to 2048×2048.
PRXPixel
Photoroom’s 7B pixel-space variant of PRX. There is no VAE: the transformer denoises raw RGB in 16×16 pixel patches, predicts the clean image rather than a velocity, and starts from noise with std 2.0. Text comes from a Qwen3-VL text tower.