DIPrune: Task-Aware Token Pruning with Dual Importance for Efficient Multimodal Language Models

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the semantic degradation and unreliable attention estimation caused by task-agnostic designs in training-free pruning for multimodal large language models. To this end, we propose DIPrune, a framework that reveals, for the first time, the suppressive effect of shallow-layer numerical inertia on deep-layer semantics. By deriving a tractable upper bound based on final task loss distortion and introducing cross-layer gradient terms, DIPrune establishes a dual importance scoring mechanism. This mechanism jointly optimizes intra-layer static feature saliency and inter-layer dynamic semantic evolution to achieve efficient visual token pruning. Experimental results demonstrate that DIPrune attains state-of-the-art performance on both LLaVA and Qwen-VL benchmarks, significantly enhancing inference efficiency while preserving semantic integrity.
📝 Abstract
Recent training-free pruning approaches for Multimodal Large Language Models (MLLMs) effectively cut computational overhead by exploiting visual redundancy or text-vision attention. However, they frequently suffer from semantic degradation due to their task-agnostic design or unreliable attention estimates. Based on our empirical analysis, we have found that this issue arises because salient tokens in shallow layers persistently suppress emerging semantic ones through numerical inertia, leading to premature discarding of signals crucial for deep reasoning. To address the aforementioned issue, from the task-oriented aspects, we first reformulate training-free pruning as a minimization of the distortion in the final task loss and derive a tractable, token-wise upper bound to serve as a surrogate objective. Specifically, this formulation inherently reveals a previously neglected inter-layer term that accounts for gradients across layers. Accordingly, for the implementation, we propose DIPrune, a rank-based framework that employs a dual importance scoring mechanism to jointly optimize intra-layer static feature saliency and inter-layer dynamic semantic evolution. Extensive experiments on LLaVA and Qwen-VL demonstrate that DIPrune consistently achieves state-of-the-art results.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Token Pruning
Semantic Degradation
Training-free
Visual Redundancy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-free Token Pruning
Multimodal Large Language Models
Task-Aware Pruning
Dual Importance Scoring
Inter-layer Semantic Evolution
💼 Related Jobs
No related jobs found.
S
Shuo Yang
Beihang University
C
Changbai Li
Beihang University
Linlin Yang
Linlin Yang
Communication University of China
Computer VisionMachine Learning
H
Huobin Tan
Beihang University
R
Rongyu Chen
National University of Singapore
Tongfei Chen
Tongfei Chen
Microsoft AI
Natural language processingProgramming languages
T
Tian Wang
Beihang University
S
Sheng Xu
Communication University of China
Baochang Zhang
Baochang Zhang
Technische Universität München
Computer assisted interventionMedical image analysisDeep learning