🤖 AI Summary
Existing vision generation models commonly employ token merging strategies that neglect semantic importance, leading to information loss, degradation of fine details, and generation artifacts. To address this, we propose a semantic-aware dynamic token merging method that— for the first time—systematically leverages importance scores derived from classifier-free guidance to inform merging decisions, prioritizing retention of high-information tokens. Our approach integrates seamlessly into mainstream diffusion models—including Stable Diffusion, Zero123++, AnimateDiff, and PixArt-α—without architectural modification. Extensive experiments across text-to-image, multi-view, and video generation tasks demonstrate substantial improvements: PSNR increases by 3.2 dB, FID decreases by 18.7%, and both detail fidelity and spatiotemporal coherence are significantly enhanced. Moreover, inference speed is accelerated by up to 2.1×. This work establishes a principled, guidance-driven paradigm for adaptive token compression in diffusion-based generative modeling.
📝 Abstract
Token merging can effectively accelerate various vision systems by processing groups of similar tokens only once and sharing the results across them. However, existing token grouping methods are often ad hoc and random, disregarding the actual content of the samples. We show that preserving high-information tokens during merging - those essential for semantic fidelity and structural details - significantly improves sample quality, producing finer details and more coherent, realistic generations. Despite being simple and intuitive, this approach remains underexplored. To do so, we propose an importance-based token merging method that prioritizes the most critical tokens in computational resource allocation, leveraging readily available importance scores, such as those from classifier-free guidance in diffusion models. Experiments show that our approach significantly outperforms baseline methods across multiple applications, including text-to-image synthesis, multi-view image generation, and video generation with various model architectures such as Stable Diffusion, Zero123++, AnimateDiff, or PixArt-$alpha$.