Docs
AI ToolkitModels

PRXPixel

Photoroom’s 7B pixel-space variant of PRX. There is no VAE: the transformer denoises raw RGB in 16×16 pixel patches, predicts the clean image rather than a velocity, and starts from noise with std 2.0. Text comes from a Qwen3-VL text tower.

org
Photoroom
modality
image
tasks
text-to-image
license
Apache 2.0
native output
1024×1024
total params
8.72B
model.arch
prx_pixel

Components

rolemodelparamssizedtypetrained
Transformer
PRX-7B pixel DiT
PRXTransformer2DModel

The diffusers class is still an unmerged PR, so ai-toolkit vendors its own copy of the transformer and a small preview sampler.

7.00B
7,003,535,872
14.01 GBbf16yes
Text encoder
Qwen3-VL text tower (28 layers, 2048-d)
transformers.Qwen3VLTextModel

Text model only, no vision tower. Last hidden state.

1.72B
1,720,574,976
3.44 GBbf16no
Tokenizer
Qwen2 tokenizer
transformers.Qwen2TokenizerFast
————
total8.72B17.45 GB

Latent space

spatial
1×
channels
3
patch
16×16
pixels per token
16×16
notes
Pixel space: no VAE. The "latent" is the RGB image in [-1, 1]; ai-toolkit uses an identity FakeVAE so encode and decode do nothing. Each 16×16×3 patch (768 values) goes through a two-layer bottleneck projection (768 → 768 → 3584) into one token.
inputlatent (c×t×h×w)tokens
1024×10243×1×1024×10244,096native, ai-toolkit sample default
512×5123×1×512×5121,024

Architecture

Blocks
24
Hidden size
3584 (28 heads × 128)
FFN size
12544 (mlp_ratio 3.5)
Text conditioning
Image queries attend to text + image keys/values in every block. Text states (2048-d, 256 tokens) are projected but not updated
Resolution conditioning
The output resolution is embedded into the timestep modulation
Objective
Flow matching with x-prediction (the model outputs the clean image), shift 3.0, noise std 2.0
Norm / position
QK norm, 2D RoPE on image tokens (axes 64/64)

In AI Toolkit

model.arch
prx_pixel
UI label
PRXPixel (pixel space) (image)
model.name_or_path
Photoroom/prxpixel-t2i
extra UI sections
model.low_vram, model.layer_offloading

UI defaults

quantize / quantize_te
true / true
low_vram
true
timestep_type
linear
network.conv
disabled (linear LoRA only)

Specifics

Resolution
Buckets and sample sizes snap to multiples of 16 (the patch size; there is no VAE).
Loss
x-prediction: the model output is compared with the clean image, not with noise − clean. The x0 → velocity conversion only happens while sampling.
Noise
Training noise and the starting noise for samples are randn × 2.0. noise_offset is applied on top.
Prompt encoding
Prompts are padded or truncated to exactly 256 tokens, with an attention mask.
Sampling
Flowmatch Euler, shift 3.0. CFG is applied on the x0 prediction, then converted to velocity with t clamped at 0.05. The vendor example uses 28 steps at guidance 5.0.
LoRA target
PRXTransformer2DModel
Saving
LoRAs use the ComfyUI key prefix. Full fine-tunes save transformer/ in Diffusers format.
Metadata base version
prx_pixel

Example config

not verified

job: extensionconfig:  name: "my_prx_pixel_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "linear"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16      model:        name_or_path: "Photoroom/prxpixel-t2i"        arch: "prx_pixel"        quantize: true        quantize_te: true        low_vram: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 5.0        sample_steps: 28        prompts:          - "A front-facing portrait of a lion in the golden savanna at sunset."

On this page