PRXPixel
Photoroom’s 7B pixel-space variant of PRX. There is no VAE: the transformer denoises raw RGB in 16×16 pixel patches, predicts the clean image rather than a velocity, and starts from noise with std 2.0. Text comes from a Qwen3-VL text tower.
- weights
- Photoroom/prxpixel-t2i ↗
- org
- Photoroom
- modality
- image
- tasks
- text-to-image
- license
- Apache 2.0
- native output
- 1024×1024
- total params
- 8.72B
- model.arch
- prx_pixel
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | PRX-7B pixel DiT PRXTransformer2DModel The diffusers class is still an unmerged PR, so ai-toolkit vendors its own copy of the transformer and a small preview sampler. | 7.00B 7,003,535,872 | 14.01 GB | bf16 | yes |
| Text encoder | Qwen3-VL text tower (28 layers, 2048-d) transformers.Qwen3VLTextModel Text model only, no vision tower. Last hidden state. | 1.72B 1,720,574,976 | 3.44 GB | bf16 | no |
| Tokenizer | Qwen2 tokenizer transformers.Qwen2TokenizerFast | — | — | — | — |
| total | 8.72B | 17.45 GB | |||
Latent space
- pixels per token
- 16×16
- notes
- Pixel space: no VAE. The "latent" is the RGB image in [-1, 1]; ai-toolkit uses an identity FakeVAE so encode and decode do nothing. Each 16×16×3 patch (768 values) goes through a two-layer bottleneck projection (768 → 768 → 3584) into one token.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 | 3×1×1024×1024 | 4,096 | native, ai-toolkit sample default |
| 512×512 | 3×1×512×512 | 1,024 |
Architecture
- Blocks
- 24
- Hidden size
- 3584 (28 heads × 128)
- FFN size
- 12544 (mlp_ratio 3.5)
- Text conditioning
- Image queries attend to text + image keys/values in every block. Text states (2048-d, 256 tokens) are projected but not updated
- Resolution conditioning
- The output resolution is embedded into the timestep modulation
- Objective
- Flow matching with x-prediction (the model outputs the clean image), shift 3.0, noise std 2.0
- Norm / position
- QK norm, 2D RoPE on image tokens (axes 64/64)
In AI Toolkit
- model.arch
- prx_pixel
- UI label
- PRXPixel (pixel space) (image)
- model.name_or_path
- Photoroom/prxpixel-t2i
- source
- extensions_built_in/diffusion_models/prx_pixel_t2i/prx_pixel_t2i.py
- extensions_built_in/diffusion_models/prx_pixel_t2i/src/transformer_prx.py
- extensions_built_in/diffusion_models/prx_pixel_t2i/src/pipeline.py
- toolkit/models/v2/text_encoders/qwen3_vl.py
- toolkit/models/FakeVAE.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- model.low_vram, model.layer_offloading
UI defaults
- quantize / quantize_te
- true / true
- low_vram
- true
- timestep_type
- linear
- network.conv
- disabled (linear LoRA only)
Specifics
- Resolution
- Buckets and sample sizes snap to multiples of 16 (the patch size; there is no VAE).
- Loss
- x-prediction: the model output is compared with the clean image, not with noise − clean. The x0 → velocity conversion only happens while sampling.
- Noise
- Training noise and the starting noise for samples are randn × 2.0. noise_offset is applied on top.
- Prompt encoding
- Prompts are padded or truncated to exactly 256 tokens, with an attention mask.
- Sampling
- Flowmatch Euler, shift 3.0. CFG is applied on the x0 prediction, then converted to velocity with t clamped at 0.05. The vendor example uses 28 steps at guidance 5.0.
- LoRA target
- PRXTransformer2DModel
- Saving
- LoRAs use the ComfyUI key prefix. Full fine-tunes save transformer/ in Diffusers format.
- Metadata base version
- prx_pixel
Example config
not verified
job: extensionconfig: name: "my_prx_pixel_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "linear" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 model: name_or_path: "Photoroom/prxpixel-t2i" arch: "prx_pixel" quantize: true quantize_te: true low_vram: true sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 5.0 sample_steps: 28 prompts: - "A front-facing portrait of a lion in the golden savanna at sunset."Links
Z-Image L2P (pixel space)
'A pixel-space version of Z-Image made with the Latent-to-Pixel (L2P) transfer method: the VAE is gone, 16×16 RGB patches go straight into the Z-Image trunk, and a small U-Net decoder turns the transformer features back into pixels.'
Krea 2 Raw
'The undistilled base checkpoint of Krea 2, a 12.8B single-stream MMDiT conditioned on stacked Qwen3-VL-4B hidden states and working in the Qwen-Image VAE latent space. Krea ships it as the checkpoint to fine-tune, not to sample from.'