Importance-Based Token Merging for Efficient Image and Video Generation

📅 2024-11-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing vision generation models commonly employ token merging strategies that neglect semantic importance, leading to information loss, degradation of fine details, and generation artifacts. To address this, we propose a semantic-aware dynamic token merging method that— for the first time—systematically leverages importance scores derived from classifier-free guidance to inform merging decisions, prioritizing retention of high-information tokens. Our approach integrates seamlessly into mainstream diffusion models—including Stable Diffusion, Zero123++, AnimateDiff, and PixArt-α—without architectural modification. Extensive experiments across text-to-image, multi-view, and video generation tasks demonstrate substantial improvements: PSNR increases by 3.2 dB, FID decreases by 18.7%, and both detail fidelity and spatiotemporal coherence are significantly enhanced. Moreover, inference speed is accelerated by up to 2.1×. This work establishes a principled, guidance-driven paradigm for adaptive token compression in diffusion-based generative modeling.

Technology Category

Computer Vision: Diffusion Models for VisionNatural Language Processing: GenerationMachine Learning: Deep Generative Models & Autoencoders

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
Token merging can effectively accelerate various vision systems by processing groups of similar tokens only once and sharing the results across them. However, existing token grouping methods are often ad hoc and random, disregarding the actual content of the samples. We show that preserving high-information tokens during merging - those essential for semantic fidelity and structural details - significantly improves sample quality, producing finer details and more coherent, realistic generations. Despite being simple and intuitive, this approach remains underexplored. To do so, we propose an importance-based token merging method that prioritizes the most critical tokens in computational resource allocation, leveraging readily available importance scores, such as those from classifier-free guidance in diffusion models. Experiments show that our approach significantly outperforms baseline methods across multiple applications, including text-to-image synthesis, multi-view image generation, and video generation with various model architectures such as Stable Diffusion, Zero123++, AnimateDiff, or PixArt-$alpha$.
Problem

Research questions and friction points this paper is trying to address.

Improving token merging for efficient image and video generation
Preserving high-information tokens to enhance sample quality
Allocating computational resources to critical tokens using importance scores
Innovation

Methods, ideas, or system contributions that make the work stand out.

Importance-based token merging prioritizes critical tokens
Uses classifier-free guidance scores for token importance
Improves quality in image and video generation tasks
🔎 Similar Papers