UDT: Reconciling U-Nets and Diffusion Transformers with Data-Adaptive Token Reduction

📅 2026-08-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the representational degradation and encoder–decoder imbalance in Diffusion Transformers (DiT) caused by fixed downsampling operations incompatible with the Transformer architecture. To resolve this, the authors propose the UDT framework, which introduces, for the first time in DiT, a data-adaptive token merging mechanism that replaces conventional downsampling, enabling efficient up- and down-sampling while preserving consistent token dimensions. UDT synergistically combines the multi-scale encoder–decoder strengths of U-Net with the powerful representational capacity of DiT, further enhanced by REPA regularization and a VA-VAE decoder. On ImageNet at 256×256 resolution, the XL variant achieves an FID of 7.9 after only 40 training epochs—matching the performance of SiT trained for 1,400 epochs, yielding a ~40× speedup. With classifier-free guidance and the VA-VAE decoder, UDT attains a state-of-the-art FID of 1.35 within 500 epochs.
📝 Abstract
Diffusion Transformers (DiTs) have emerged as a core architecture in generative modeling due to their scalability and adaptability to multimodal tasks. DiTs comprise isotropic transformer blocks, and learn representations progressively across depth, where the denoising objective drives later layers to focus on fine-detail reconstruction. This results in degraded representation quality and an imbalanced encoder-decoder behavior. Prior approaches such as representation alignment (REPA) mitigate this by encouraging stronger early representations via training regularization. Alternatively, U-Net-style DiT architectures introduce explicit multi-scale encoder-decoder structures for improved convergence. But they build on standard U-Net wisdom via learnable operators for spatial downsampling, which are not well-suited to transformer architectures, introducing inefficiencies and compatibility issues with components such as cross-attention and representation regularization. In this work, we propose UDT, a U-Net diffusion transformer that combines the representation power of DiTs with the encoding-decoding benefits of U-Nets, through data-adaptive token merging for downsampling and upsampling, while preserving the DiT token dimension. Our baseline UDT architecture outperforms existing U-Net DiTs and achieves performance comparable to REPA across all model sizes. Furthermore, using architectural optimization and REPA, UDT outperforms SiT's 7.9 FID at 1400 epochs (w/o CFG) within 40 epochs (~ 40x faster convergence) for XL model size on 256x256 ImageNet. Finally, it achieves strong image generation performance with CFG, reaching FID of 1.38 (320 epochs) with SD-VAE and 1.35 (500 epochs) with VA-VAE, providing a new backbone for DiTs with strong empirical benefits.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Transformers
U-Net
token reduction
representation imbalance
architectural compatibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion Transformer
U-Net
data-adaptive token reduction
multi-scale architecture
fast convergence
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Junno Yun
University of Minnesota
Y
Yaşar Utku Alçalar
University of Minnesota
M
Mehmet Akçakaya
University of Minnesota