Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing training-free visual token pruning methods rely on static heuristic strategies that irreversibly discard tokens in early layers, overlooking the dynamic evolution of token importance across layers in multimodal large language models and consequently causing premature loss of critical visual information. This work proposes Trend-aware Pruning, a novel framework that formulates pruning as an attention flow trend prediction problem. By dynamically tracking the evolving importance of tokens, the method enables selective reactivation of “late-rising” tokens. It introduces a training-free, reversible, and layer-aware pruning mechanism coupled with cross-layer semantic momentum tracking. The approach achieves over 77.8% visual token compression—retaining only about 23 tokens in the final layer—while maintaining competitive performance, thereby significantly improving the efficiency–accuracy trade-off in multimodal tasks.
📝 Abstract
While visual token pruning is essential for efficient Multimodal Large Language Models (MLLMs), existing training-free methods suffer from a critical limitation: they rely on static, instantaneous heuristics to perform irreversible filtering. This approach ignores the hierarchical nature of MLLMs, where token importance often evolves dynamically rather than remaining fixed across layers. Consequently, tokens essential for deep-layer reasoning are often prematurely discarded by shallow-layer estimates. To address this, we propose Trend-aware Pruning, a novel framework that elevates pruning from a local snapshot decision to a temporal trajectory modeling problem. Instead of relying on isolated scores, our method captures the momentum of attention flow. This enables a dynamic rectification mechanism that selectively reactivates "late-blooming" tokens, those initially undervalued but exhibiting rising semantic importance, thereby preventing the loss of critical visual cues. Extensive experiments demonstrate that our approach achieves a superior efficiency-performance trade-off across diverse multimodal tasks. Notably, it reduces visual tokens by over 77.8%, retaining only approximately 23 tokens in the final layer while maintaining competitive performance, offering a robust and reversible solution for high-efficiency multimodal inference.
Problem

Research questions and friction points this paper is trying to address.

token pruning
multimodal large language models
training-free
dynamic token importance
visual token filtering
Innovation

Methods, ideas, or system contributions that make the work stand out.

Trend-aware Pruning
training-free token pruning
attention flow momentum
late-blooming tokens
multimodal efficiency
Jie Ma
Jie Ma
China University of Petroleum-Beijing
Environmental engineeringenvironmental microbiologymicrobial ecologyvapor intrusion.
Z
Zhike Qiu
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University
J
Jie Gao
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University
Jiayi Ji
Jiayi Ji
Rutgers University
Q
Qian Chen
School of Information Engineering, Xiamen Ocean Vocational College
X
Xiaoshuai Sun
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University; Sino-Russian Research Center for Digital Economy
R
Rongrong Ji
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University