Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?

๐Ÿ“… 2025-02-17
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Token pruning in multimodal large language models (MLLMs) aims to reduce inference overhead, yet existing methods suffer from fundamental flawsโ€”including unreliable attention-based scoring, limited linguistic contribution, suboptimal redundancy-repetition trade-offs, and biased evaluation protocols. Method: This work is the first to systematically challenge both pruning design principles and evaluation paradigms, proposing a novel assessment framework that jointly ensures interpretability and fairness. We develop a diagnostic evaluation protocol grounded in theoretical analysis, controlled ablation experiments, and cross-model validation (ViT-L/LLaMA-2/3). Contribution/Results: Key findings reveal that most state-of-the-art pruning methods perform no better than random pruning; visual tokens constitute the primary source of redundancy; and linguistic information exhibits sharply diminishing marginal utility under pruning. These insights establish a rigorous theoretical foundation and practical benchmark for efficient MLLM inference.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Language Grounding & Multi-modal NLPComputer Vision: Multi-modal Vision

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Web evaluation methodologies and metrics
๐Ÿ“ Abstract
Multimodal large language models (MLLMs) have shown remarkable performance for cross-modal understanding and generation, yet still suffer from severe inference costs. Recently, abundant works have been proposed to solve this problem with token pruning, which identifies the redundant tokens in MLLMs and then prunes them to reduce the computation and KV storage costs, leading to significant acceleration without training. While these methods claim efficiency gains, critical questions about their fundamental design and evaluation remain unanswered: Why do many existing approaches underperform even compared to naive random token selection? Are attention-based scoring sufficient for reliably identifying redundant tokens? Is language information really helpful during token pruning? What makes a good trade-off between token importance and duplication? Are current evaluation protocols comprehensive and unbiased? The ignorance of previous research on these problems hinders the long-term development of token pruning. In this paper, we answer these questions one by one, providing insights into the design of future token pruning methods.
Problem

Research questions and friction points this paper is trying to address.

Evaluate token pruning efficiency
Assess attention-based scoring adequacy
Analyze language information impact
Innovation

Methods, ideas, or system contributions that make the work stand out.

Token pruning reduces MLLM computation
Attention-based scoring identifies redundant tokens
Evaluates token importance and duplication trade-off