ASAP: Attention Sink Anchored Pruning

📅 2026-05-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the computational bottleneck of high-resolution Vision Transformers caused by the quadratic complexity of self-attention. Existing token pruning methods based on local attention scores are prone to interference from “attention sinks”—dominant tokens that disproportionately attract attention—leading to erroneous removal of critical foreground information. The paper introduces, for the first time, a training-free, single-pass pruning framework that treats attention sinks as pivotal conduits of global information flow. It models token interactions via Lazy Random Walks, computes diffusion distances from each token to attention sinks using cumulative transition matrices, and achieves efficient compression through radial diffusion clustering combined with transition-weighted pooling. Evaluated across image, video, and vision-language tasks, the method outperforms state-of-the-art pruning strategies, yielding up to a 48% increase in inference throughput while preserving or even improving model accuracy.
📝 Abstract
Vision Transformers (ViTs) face severe computational bottlenecks due to the quadratic complexity of self-attention at high resolutions. Existing token reduction methods rely on local metrics - such as single-layer attention scores - that are inherently vulnerable to the attention sink phenomenon, where uninformative tokens are paradoxically preserved over salient foreground objects. We propose ASAP (Attention Sink Anchored Pruning), a training-free framework that recasts this sink as a feature. Modeling ViT information flow as a Lazy Random Walk, ASAP identifies the sink as a dominant accumulator of probability mass. By computing the diffusion distance to the sink within the cumulative transition matrix, ASAP partitions tokens via Radial Diffusion Clustering and compresses background redundancy through Transition Weight Pooling in a single shot. Extensive experiments across image, video, and vision-language tasks demonstrate ASAP outperforms state-of-the-art methods, accelerating throughput by up to 48% while maintaining - or even exceeding - baseline accuracy.
Problem

Research questions and friction points this paper is trying to address.

Vision Transformers
computational bottleneck
attention sink
token reduction
self-attention complexity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Attention Sink
Token Pruning
Vision Transformers
Diffusion Distance
Training-Free
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jaehyuk Lee
Department of Mathematics, Korea University, Seoul, Republic of Korea
H
Hanyoung Kim
Department of Mathematics, Korea University, Seoul, Republic of Korea
Y
Yanggee Kim
Korea University, Seoul, Republic of Korea
Donghun Lee
Donghun Lee
Seoul National University, ETRI
Machine Learning