Interpretability-Guided Soft Pruning of Attention Heads in Vision Transformers

📅 2026-07-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high computational cost, redundant attention heads, and structural opacity of vision foundation models such as DINOv2. The authors propose the SAPER framework, which leverages spectral analysis and visualization of Laplacian eigenvectors to uncover the semantic roles of individual attention heads and employs semantic clustering to identify redundancy among them. Building on these insights, they introduce an end-to-end differentiable pruning mechanism driven by LapSum Soft Top-K, enabling efficient head selection during training. Evaluated on ImageNet-1K, the method substantially outperforms the RAPTOR baseline, achieving significantly reduced FLOPs while preserving competitive classification accuracy. This approach establishes a new paradigm for interpretable and efficient compression of vision models.
📝 Abstract
Vision foundation models, such as DINOv2, learn highly expressive representations but rely on massive, opaque architectures that demand substantial computational power and memory. To provide an interpretable-guided and efficient solution to this issue, we first propose a spectral analysis and new visualization technique for individual attention heads based on the Laplacian eigenvectors of their attention maps. Building upon recent observations regarding the block structure of Vision Transformers, we perform semantic clustering of attention heads and identify functional redundancies. Leveraging these insights, we introduce SAPER (Soft Attention PrunER), an end-to-end differentiable pruning framework based on the LapSum Soft Top-K approach. Extensive experiments on ImageNet-1K demonstrate that SAPER achieves a highly favorable accuracy-efficiency trade-off, outperforming the competitive RAPTOR baseline in FLOPs reduction while preserving strong classification performance.
Problem

Research questions and friction points this paper is trying to address.

Vision Transformers
interpretability
attention pruning
model efficiency
computational cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Interpretability
Attention Head Pruning
Vision Transformers
Spectral Analysis
Soft Top-K
🔎 Similar Papers
Kamil Książek
Kamil Książek
Jagiellonian University
machine learningdeep learninghyperspectral imagingcontinual learningqueueing theory
P
Piotr Suszyński
Centre for Credible Artificial Intelligence, Warsaw University of Technology, Warsaw, Poland
M
Michał Jan Włodarczyk
Centre for Credible Artificial Intelligence, Warsaw University of Technology, Warsaw, Poland
Jacek Tabor
Jacek Tabor
Profesor informatyki, Uniwersytet Jagielloński
mathematicscomputer science
P
Przemysław Biecek
Centre for Credible Artificial Intelligence, Warsaw University of Technology, Warsaw, Poland; University of Warsaw, Warsaw, Poland