🤖 AI Summary
Stable diffusion models suffer from low inference efficiency due to the quadratic computational complexity of self-attention. Existing token merging methods fail to adequately model the locality and semantic importance of cross-modal attention in text-to-image generation, thus struggling to balance efficiency and generation quality. To address this, we propose a local representative token-guided token merging method: we introduce the novel concept of *local representative tokens*, integrating dynamic window partitioning with similarity-based adaptive token selection to identify the most representative tokens within context-aware local regions. This strategy is model-agnostic and requires no architectural modifications. Experiments demonstrate that our method achieves a 6.2% reduction in FID while significantly improving CLIP Score, all without compromising inference speed—effectively reconciling high-fidelity generation with computational efficiency.
📝 Abstract
Stable diffusion is an outstanding image generation model for text-to-image, but its time-consuming generation process remains a challenge due to the quadratic complexity of attention operations. Recent token merging methods improve efficiency by reducing the number of tokens during attention operations, but often overlook the characteristics of attention-based image generation models, limiting their effectiveness. In this paper, we propose local representative token guided merging (ReToM), a novel token merging strategy applicable to any attention mechanism in image generation. To merge tokens based on various contextual information, ReToM defines local boundaries as windows within attention inputs and adjusts window sizes. Furthermore, we introduce a representative token, which represents the most representative token per window by computing similarity at a specific timestep and selecting the token with the highest average similarity. This approach preserves the most salient local features while minimizing computational overhead. Experimental results show that ReToM achieves a 6.2% improvement in FID and higher CLIP scores compared to the baseline, while maintaining comparable inference time. We empirically demonstrate that ReToM is effective in balancing visual quality and computational efficiency.