Residual-Stream Burden Shapes Representation Learning in Diffusion Transformers

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the representational learning asymmetry in diffusion Transformers, where noise prediction targets underperform compared to clean data prediction. To elucidate this mechanism, we propose the theory of "residual stream burden," which reveals that noise targets compel deeper network layers to process highly noisy inputs, thereby constraining the degrees of freedom for representation organization. Motivated by this insight, we introduce the Spatially-indexed Hyper-Connections (SiHC) architecture, which directly expands residual bandwidth to alleviate such burden. Extensive experiments validate the critical role of residual bandwidth in noise prediction tasks. Notably, SiHC achieves a superior FID of 1.71 on ImageNet 256×256, demonstrating exceptional generative quality.
📝 Abstract
In diffusion-based generation, a neural network can be trained to predict the clean data, the noise, or the velocity from a noisy input. These prediction targets are interconvertible and describe the same generative process, yet plain Diffusion Transformers operating on large pixel patches succeed with clean prediction and fail with noise or velocity prediction. We argue that this asymmetry arises because noisy targets require the residual stream to preserve noise-dependent input variation through depth for the final readout, forcing subsequent layers to compute on noisy representations. A spectrally concentrated clean target imposes a lighter demand, leaving greater freedom to organize hidden representations for subsequent computation. We call this preservation requirement *residual-stream burden* and show how it shapes representation learning in Diffusion Transformers. Controlled experiments indicate that the exploitable structure is spectral concentration in patch space and that the bandwidth of the persistent residual state is a key resource for noisy prediction. We further show that this account is consistent with recent decoupled pixel-space architectures, whose diverse designs all reduce the residual-stream burden on the main pathway. To examine this understanding from a complementary direction, we expand and reorganize the residual-stream bandwidth directly, introducing Spatially Indexed Hyper-Connections (SiHC) that reach FID 1.71 on ImageNet $256^2$. Together, these results identify residual-stream burden as a mechanism through which prediction targets and architecture jointly shape representation learning in Diffusion Transformers.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Transformers
residual-stream burden
representation learning
prediction targets
spectral concentration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Residual-Stream Burden
Diffusion Transformers
Representation Learning
Spatially Indexed Hyper-Connections
Spectral Concentration