PixelDiT2: Representation-Grounded Pixel Diffusion Transformers

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决像素空间扩散模型收敛慢和图像质量低的问题,提出PixelDiT2,通过引入预训练视觉模型提供表示指导,使模型更专注于像素生成。
📝 Abstract
Recent advances in pixel-space diffusion models have narrowed the image quality gap with latent-space diffusion, but still converge more slowly and lag behind in final image quality. We argue that a key reason is the lack of an explicit representation prior: unlike latent diffusion, which usually denoises in a compact and structured latent space, pixel diffusion needs to learn denoising-friendly representations and pixel generation simultaneously from raw RGB space. To address this problem, we propose PixelDiT2, an end-to-end pixel-space diffusion model designed to decouple representation learning from pixel generation without introducing an autoencoder or latent reconstruction bottleneck. We propose representation grounding that uses a frozen pretrained vision foundation model to provide explicit per-patch representation guidance throughout denoising, allowing the pixel diffusion transformer to focus more on pixel generation. On ImageNet-256x256, PixelDiT2 achieves an FID of 1.46 after 600 epochs; at 512x512 resolution, PixelDiT2 achieves an FID of 1.48 after 680 epochs.
Problem

Research questions and friction points this paper is trying to address.

pixel diffusion
image quality
representation learning
latent space
Innovation

Methods, ideas, or system contributions that make the work stand out.

PixelDiT2
representation grounding
pixel-space diffusion model
frozen pretrained vision model
decoupling representation learning
Yongsheng Yu
Yongsheng Yu
University of Rochester
image generation
W
Wei Xiong
1NVIDIA 2University of Rochester
Yichen Sheng
Yichen Sheng
NVIDIA
computer graphicscomputer vision
S
Shiqiu Liu
1NVIDIA
J
Jiebo Luo
2University of Rochester