RefAdapt-DiT: Adaptive Joint Attention for Reference-Conditioned Diffusion Transformers

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive computational overhead of bidirectional joint attention in conditional generation with diffusion Transformers. To this end, we propose a training-free adaptive acceleration framework that dynamically regulates reference-stream computation by leveraging target query variations and attention quality. Furthermore, by incorporating reference drift and exposure features to overcome the limitations of static reuse, the framework achieves fine-grained, block-level adaptive control. Experimental results demonstrate that the proposed method yields 2.1× and 3.5× speedups for 4-step and 8-step models, respectively, while preserving comparable visual quality.
📝 Abstract
Diffusion Transformers (DiTs) have become the standard backbone for high-quality generative modeling, yet deploying them in conditional generation tasks remains computationally prohibitive because bidirectional joint attention repeatedly processes large reference streams. While existing optimization schemes mitigate generic temporal redundancy, they typically rely on coarse-grained static reuse and overlook the distinct dynamics of references and targets. Specifically, we observe that reference representations often evolve slowly along the generation trajectory, while the target often assigns little attention mass to them; reference drift and this target-to-reference exposure jointly shape how strongly stale reference states affect the target. To exploit these patterns, we introduce \RefAdapt, a training-free framework for adaptive control of joint attention between references and targets. Instead of rigid static strategies, \RefAdapt combines consecutive target-Q change with previously observed target-to-reference attention mass to control reference computation adaptively at block granularity. Under ultra-few-step settings, \RefAdapt enables speedups of up to $2.097\times$ on 4-step MiniMax H3 and $3.54\times$ on 8-step Qwen Image Edit, while maintaining comparable visual quality.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Transformers
conditional generation
joint attention
computational efficiency
reference-conditioned
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion Transformers
Training-free Acceleration
Adaptive Joint Attention
Reference-Conditioned Generation
Block-granularity Control
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jian Tang
Platform and Content Group, Tencent
Jiawei Fan
Jiawei Fan
Intel Labs China
Deep LearningComputer VisionModel CompressionDeep Learning Acceleration
Q
Qiannan Zhou
Platform and Content Group, Tencent
Q
Qingbin Liu
Platform and Content Group, Tencent
J
Jiang Bian
Platform and Content Group, Tencent
Z
Zang Li
Platform and Content Group, Tencent