Local Representative Token Guided Merging for Text-to-Image Generation

📅 2025-07-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Stable diffusion models suffer from low inference efficiency due to the quadratic computational complexity of self-attention. Existing token merging methods fail to adequately model the locality and semantic importance of cross-modal attention in text-to-image generation, thus struggling to balance efficiency and generation quality. To address this, we propose a local representative token-guided token merging method: we introduce the novel concept of *local representative tokens*, integrating dynamic window partitioning with similarity-based adaptive token selection to identify the most representative tokens within context-aware local regions. This strategy is model-agnostic and requires no architectural modifications. Experiments demonstrate that our method achieves a 6.2% reduction in FID while significantly improving CLIP Score, all without compromising inference speed—effectively reconciling high-fidelity generation with computational efficiency.

Technology Category

Computer Vision: Diffusion Models for VisionMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Generation

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Web query analysis, representation and understandingUser Modeling, Personalization and Recommendation: Federated recommendation systems and personalization
📝 Abstract
Stable diffusion is an outstanding image generation model for text-to-image, but its time-consuming generation process remains a challenge due to the quadratic complexity of attention operations. Recent token merging methods improve efficiency by reducing the number of tokens during attention operations, but often overlook the characteristics of attention-based image generation models, limiting their effectiveness. In this paper, we propose local representative token guided merging (ReToM), a novel token merging strategy applicable to any attention mechanism in image generation. To merge tokens based on various contextual information, ReToM defines local boundaries as windows within attention inputs and adjusts window sizes. Furthermore, we introduce a representative token, which represents the most representative token per window by computing similarity at a specific timestep and selecting the token with the highest average similarity. This approach preserves the most salient local features while minimizing computational overhead. Experimental results show that ReToM achieves a 6.2% improvement in FID and higher CLIP scores compared to the baseline, while maintaining comparable inference time. We empirically demonstrate that ReToM is effective in balancing visual quality and computational efficiency.
Problem

Research questions and friction points this paper is trying to address.

Reduces quadratic complexity in stable diffusion attention operations
Improves token merging for attention-based image generation models
Balances visual quality and computational efficiency in text-to-image
Innovation

Methods, ideas, or system contributions that make the work stand out.

Local representative token guided merging strategy
Adjustable window sizes for contextual merging
Highest similarity token selection per window
M
Min-Jeong Lee
Department of Artificial Intelligence, Korea University, Anam-dong, Seongbuk-ku, Seoul 02841, Korea
H
Hee-Dong Kim
Department of Artificial Intelligence, Korea University, Anam-dong, Seongbuk-ku, Seoul 02841, Korea
S
Seong-Whan Lee
Department of Artificial Intelligence, Korea University, Anam-dong, Seongbuk-ku, Seoul 02841, Korea