🤖 AI Summary
This study addresses the significant degradation in generation quality of Diffusion Transformers (DiTs) under out-of-distribution free-text prompts. By identifying the deficiency in image-text attention during early denoising stages as the root cause, this work proposes a timestep-aware gated attention mechanism. This method adaptively optimizes image-text attention allocation across all denoising phases through the dynamic injection of conditional biases, introducing a novel gating strategy that synergizes with the denoising process. Experimental results demonstrate that the proposed approach substantially outperforms existing baselines across multiple benchmarks, achieving a 9.5% improvement in DPG scores under raw prompts. These findings indicate that our method effectively enhances both the robustness and generative capability of DiTs when conditioned on open-domain text inputs.
📝 Abstract
Diffusion Transformers (DiTs) have emerged as the dominant architecture for high-fidelity image and video generation. Recent DiT systems increasingly use structured prompts for training, improving caption quality and prompt adherence. However, their generation quality can degrade severely under out-of-domain (OOD) prompts, including the free-form descriptions supplied by users at inference time. Although LLM-based rewriting can convert these prompts into structured formats, it does not guarantee that the rewritten prompts align with the training distribution. Our analysis links this degradation to attention sinks and reduced early-step image-to-text attention and shows that sink suppression alone is insufficient to restore generation quality. Despite effective sink suppression, models trained with standard gated attention exhibit reduced early-step image-to-text attention and suboptimal generation quality. Based on these insights, we propose Timestep-Aware Gated Attention (TSGate), which injects a timestep-conditioned bias into the gate signal so that gating behavior adapts across denoising steps. Extensive experiments show that TSGate consistently outperforms both the baseline and standard gated attention across multiple benchmarks, improving the raw-prompt DPG score by 9.5% over the baseline.