Docs
AI ToolkitModels

Z-Image L2P (pixel space)

'A pixel-space version of Z-Image made with the Latent-to-Pixel (L2P) transfer method: the VAE is gone, 16×16 RGB patches go straight into the Z-Image trunk, and a small U-Net decoder turns the transformer features back into pixels.'

org
zhen-nan (NJU PCALab)not verified
modality
image
tasks
text-to-image
license
Apache 2.0
released
2026-05-03not verified
total params
10.19B
model.arch
zimage_l2p

Components

rolemodelparamssizedtypetrained
Transformer
Z-Image L2P (trunk + pixel decoder, single file)
ZImageTransformer2DModel (ai-toolkit L2P subclass)

One file: the 30-layer Z-Image trunk and refiners, a 16×16 pixel patch embedder (768 → 3840) and a 10,085,123-parameter U-Net pixel decoder (local_decoder). The Z-Image final layer is not in the file. Main layers 5–24 are stored in fp32 (3,618,206,720 params), everything else in bf16; ai-toolkit casts all of it to the training dtype on load.

6.17B
6,166,464,515
19.57 GBfp32+bf16yes
Text encoder
Qwen3-4B
transformers.Qwen3ForCausalLM

Not in the L2P repo. ai-toolkit loads it from extras_name_or_path, Tongyi-MAI/Z-Image-Turbo. Only the second-to-last hidden state is used.

4.02B
4,022,468,096
8.04 GBbf16no
Tokenizer
Qwen3 BPE tokenizer
transformers.Qwen2Tokenizer
————
total10.19B27.61 GB

Latent space

spatial
1×
channels
3
patch
16×16
pixels per token
16×16
notes
No autoencoder: the model denoises RGB pixels directly (ai-toolkit uses an identity FakeVAE with scaling 1.0). Each token is one 16×16×3 pixel patch (768 values), so the token grid matches latent Z-Image at the same resolution (8× VAE × 2×2 patch).
inputlatent (c×t×h×w)tokens
512×5123×1×512×5121,024
1024×10243×1×1024×10244,096ai-toolkit sample default
2048×20483×1×2048×204816,384

Architecture

Blocks
Z-Image trunk: 30 single-stream blocks, 2 noise-refiner and 2 context-refiner blocks
Hidden size
3840 (30 heads × 128)
Input
Linear embedder on 16×16 RGB patches (768 → 3840) replaces the latent 2×2 patch embedder
Output
4-stage conv U-Net (64/128/256/512 channels, 16× down) on the noisy image, with the transformer feature map (one 3840-d vector per patch) fused at the bottleneck. Replaces the Z-Image final layer
Text conditioning
Same as Z-Image: Qwen3-4B second-to-last hidden states, up to 512 tokens, in one sequence with the image tokens
Objective
Rectified flow in pixel space (ai-toolkit shift 3.0)
Norm / position
RMSNorm, QK RMSNorm, 3-axis RoPE [32, 48, 48], theta 256

In AI Toolkit

model.arch
zimage_l2p
UI label
Z-Image L2P (pixel space) (image)
model.name_or_path
zhen-nan/L2P/model-1k-merge.safetensors
extra UI sections
model.low_vram, model.layer_offloading

UI defaults

extras_name_or_path
Tongyi-MAI/Z-Image-Turbo
quantize / quantize_te
true / true
low_vram
true
timestep_type
linear
network.conv
disabled (linear LoRA only)

Specifics

Checkpoint download
An org/repo/file.safetensors name_or_path is downloaded once to MODELS_PATH/diffusion_models and reused from there.
Single-file loading
No config ships with the file. ai-toolkit builds the model from a built-in Z-Image config, detects pixel space from local_decoder.* or all_x_embedder.16-1 keys, and loads the weights non-strictly.
Latent-to-pixel conversion
Point name_or_path at a latent Z-Image checkpoint (in_channels 16) and it is converted on load: a new 16×16 patch embedder (random × 0.001), the final layer dropped, and a randomly initialised U-Net decoder. That starts an L2P transfer, not a finished model.
VAE
None. An identity FakeVAE (scaling 1.0) stands in, so cached “latents” are pixels.
Text encoder source
text_encoder/ and tokenizer/ from extras_name_or_path. Never trained.
Resolution
Buckets snap to multiples of 16 (inherited from Z-Image; also the pixel patch size).
Timesteps and loss
Inherited from Z-Image: model t = (1000 − timestep) / 1000, output negated, target noise − pixels, scheduler shift 3.0.
Sampling
Z-Image pipeline with the identity VAE. guidance_scale is shifted down by 1. The inference engine registry defaults to 8 steps, guidance 1.
LoRA target
ZImageTransformer2DModel
Saving
Full fine-tunes save the raw state dict (Diffusers key names, not the ComfyUI layout) minus all_final_layer.*, cast to the save dtype.
Metadata base version
zimage

Example config

not verified

job: extensionconfig:  name: "my_zimage_l2p_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: bf16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "linear"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16      model:        name_or_path: "zhen-nan/L2P/model-1k-merge.safetensors"        extras_name_or_path: "Tongyi-MAI/Z-Image-Turbo"        arch: "zimage_l2p"        quantize: true        quantize_te: true        low_vram: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 1        sample_steps: 8        prompts:          - "a bear building a log cabin in the snow covered mountains"

On this page