🤖 AI Summary
This study addresses the prohibitive computational overhead of bidirectional joint attention in conditional generation with diffusion Transformers. To this end, we propose a training-free adaptive acceleration framework that dynamically regulates reference-stream computation by leveraging target query variations and attention quality. Furthermore, by incorporating reference drift and exposure features to overcome the limitations of static reuse, the framework achieves fine-grained, block-level adaptive control. Experimental results demonstrate that the proposed method yields 2.1× and 3.5× speedups for 4-step and 8-step models, respectively, while preserving comparable visual quality.
📝 Abstract
Diffusion Transformers (DiTs) have become the standard backbone for high-quality generative modeling, yet deploying them in conditional generation tasks remains computationally prohibitive because bidirectional joint attention repeatedly processes large reference streams. While existing optimization schemes mitigate generic temporal redundancy, they typically rely on coarse-grained static reuse and overlook the distinct dynamics of references and targets. Specifically, we observe that reference representations often evolve slowly along the generation trajectory, while the target often assigns little attention mass to them; reference drift and this target-to-reference exposure jointly shape how strongly stale reference states affect the target. To exploit these patterns, we introduce \RefAdapt, a training-free framework for adaptive control of joint attention between references and targets. Instead of rigid static strategies, \RefAdapt combines consecutive target-Q change with previously observed target-to-reference attention mass to control reference computation adaptively at block granularity. Under ultra-few-step settings, \RefAdapt enables speedups of up to $2.097\times$ on 4-step MiniMax H3 and $3.54\times$ on 8-step Qwen Image Edit, while maintaining comparable visual quality.