Structured Residual Connectivity Matters for Diffusion Transformers

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited feature reuse efficiency in existing Diffusion Transformers (DiTs) caused by inflexible residual connections. To overcome this, we propose a structured adaptive connectivity mechanism that upgrades passive summation to active retrieval. By integrating local residuals with long-range pathways through differentiable cross-depth attention, our method enables dynamic cross-layer feature selection, thereby optimizing the image denoising process. Notably, this approach achieves these improvements with less than 0.1% additional parameters, reducing training iterations by 1.73× and lowering the FID score to 4.34. These results demonstrate significant enhancements in both generation quality and computational efficiency for diffusion-based models.
📝 Abstract
Diffusion Transformers (DiTs) have established themselves as a scalable backbone for high-fidelity image synthesis. However, unlike U-Net based diffusion models that rely on rigid, hand-crafted skip connections, DiTs predominantly use a uniform residual stream that integrates all preceding layers as a monolithic state. In this work, we rethink residual connections in diffusion transformers and propose to transform them from passive summation into an active retrieval mechanism optimized for image denoising. First, we conduct a systematic analysis of DiT's internal representation, revealing a latent preference for early-layer feature reuse and symmetric layer guidance. Motivated by this, we introduce a structured connectivity design that explicitly integrates local residual connections with long-range pathways. Instead of static skip connections or dense all-layer routing, our method enables each transformer block to selectively ``attend''to critical earlier representations, dynamically retrieving spatial and semantic cues through direct, differentiable cross-depth paths. Experiments show that our adaptive connectivity leads to faster convergence, with up to $1.73\times$ fewer training iterations, and significant gains in FID and visual quality with less than $0.1\%$ additional parameters, further improving a strong REPA-XL/2 model from $5.9$ to $4.34$ FID without guidance and reaching $1.39$ FID with classifier-free guidance. Our findings suggest that adaptive cross-layer connectivity is a critical yet underexplored factor in diffusion transformers, and that incorporating structured information pathways provides a simple and effective direction for improving scalable generative models.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Transformers
Residual Connectivity
Image Synthesis
Skip Connections
Cross-layer Connectivity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion Transformers
Structured Residual Connectivity
Cross-layer Attention
Image Denoising
Adaptive Skip Connections