TrimTokenator-LC: Towards Adaptive Visual Token Pruning for Large Multimodal Models with Long Contexts

📅 2025-12-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the prohibitive inference cost of large vision-language models (VLMs) under long-context, multi-image inputs, this paper proposes an adaptive visual token pruning method. Unlike prior approaches, it uniquely decouples redundancy modeling into two orthogonal dimensions: intra-image diversity and inter-image dissimilarity, and introduces a Pareto-optimal selection mechanism to jointly optimize both against text alignment. The proposed two-stage framework—comprising global candidate pool construction, diversity quantification, and greedy subset selection—dynamically allocates token budgets and identifies the most representative visual tokens without fine-tuning or supervision. Evaluated on multi-image long-context benchmarks, our method reduces visual tokens by up to 67% while maintaining or even improving performance on VQA and image captioning tasks. This achieves content-aware, cross-image cooperative inference with significant efficiency gains.

Technology Category

Computer Vision: Large Vision ModelsMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSearch and Retrieval-Augmented AI: Large language models for searchWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Large Multimodal Models (LMMs) have proven effective on various tasks. They typically encode visual inputs into Original Model sequences of tokens, which are then concatenated with textual tokens and jointly processed by the language model. However, the growing number of visual tokens greatly increases inference cost. Visual token pruning has emerged as a promising solution. However, existing methods often overlook scenarios involving long context inputs with multiple images. In this paper, we analyze the challenges of visual token pruning in long context, multi-image settings and introduce an adaptive pruning method tailored for such scenarios. We decompose redundancy into intra-image and inter-image components and quantify them through intra-image diversity and inter-image variation, which jointly guide dynamic budget allocation. Our approach consists of two stages. The intra-image stage allocates each image a content-aware token budget and greedily selects its most representative tokens. The inter-image stage performs global diversity filtering to form a candidate pool and then applies a Pareto selection procedure that balances diversity with text alignment. Extensive experiments show that our approach maintains strong performance in long context settings while significantly cutting down the number of visual tokens.
Problem

Research questions and friction points this paper is trying to address.

Adaptive visual token pruning for long multimodal contexts
Reducing inference cost in multi-image LMM scenarios
Balancing token diversity with text alignment efficiently
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive pruning for long multi-image contexts
Intra-image diversity and inter-image variation guide allocation
Two-stage selection balancing diversity and text alignment
H
Hao Zhang
Beijing Academy of Artificial Intelligence (BAAI)
M
Mengsi Lyu
Beijing Academy of Artificial Intelligence (BAAI)
B
Bo Huang
Beijing Academy of Artificial Intelligence (BAAI)
Y
Yulong Ao
Beijing Academy of Artificial Intelligence (BAAI)
Y
Yonghua Lin
Beijing Academy of Artificial Intelligence (BAAI)