MiCo: Mutual Information Coverage Optimization through Semantic Erasure Modeling for Efficient MLLM Inference

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the substantial computational overhead in multimodal large language models caused by visual token redundancy, where existing pruning methods often discard critical information. To this end, we propose MiCo, a training-free two-stage pruning framework. This method introduces a novel mutual information coverage objective derived from task loss, and leverages semantic erasure modeling to reformulate token selection as a monotone submodular function optimization problem, efficiently solved via a greedy algorithm to achieve theoretically grounded preservation of core visual information. Experiments demonstrate that MiCo attains state-of-the-art performance across multiple benchmarks. Notably, LLaVA-NEXT-13B retains 97.5% of its original performance using only 5.6% of tokens, achieving a 3.8× inference speedup.
📝 Abstract
Multimodal large language models (MLLMs) have demonstrated impressive performance in multimodal understanding, but processing large numbers of visual tokens results in high computational costs. While many methods have been proposed to reduce the number of visual tokens, most of them rely on heuristics and are prone to discarding substantial visual information during pruning, leading to degradation in model performance. In this work, by using a semantic erasure model, we derive a general mutual information coverage objective from task log-loss and propose MiCo, a training-free two-stage pruning method. MiCo first uses visual signals to select a representative candidate pool before visual tokens enter the language model, then performs task-aware subset selection within it. At each stage, suitable observable proxies instantiate the derived objective as a monotone submodular coverage function, which MiCo greedily optimizes under the token budget. MiCo is evaluated on diverse MLLMs ranging from 7B to 13B parameters across a broad range of image and video benchmarks spanning general visual reasoning, fine-grained OCR and grounding, hallucination detection, and long-video understanding. MiCo consistently achieves the best performance across nearly all evaluated models under all pruning ratios. On LLaVA-NEXT-13B, MiCo uses only 5.6% visual tokens, retains 97.5% of baseline performance, and achieves a 3.8-fold inference speedup. Our experiments demonstrate the effectiveness of MiCo and our mutual information coverage objective for visual token pruning.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Visual Token Pruning
Computational Efficiency
Information Loss
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mutual Information Coverage
Semantic Erasure Modeling
Training-free Pruning
Submodular Optimization
Multimodal Large Language Models
🔎 Similar Papers
No similar papers found.
T
Tinghao Wang
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University; University of Electronic Science and Technology of China
Yichen Guo
Yichen Guo
Master student in Nanyang Technological University
Qizhe Zhang
Qizhe Zhang
School of Computer Science, Peking University
Vision Language ModelComputer VisionMachine Learning
Y
Yuan Zhang
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
W
Weimin Ouyang
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
Rui Huang
Rui Huang
The Chinese University of Hong Kong, Shenzhen
Computer VisionPattern RecognitionImage ProcessingMachine Learning
Jiajun Cao
Jiajun Cao
Ph.D. Student, Peking University
MLLMComputer Vision
Sixiang Chen
Sixiang Chen
The Hong Kong University of Science and Technology (Guangzhou)
Computer VisionImage RestorationAIGCMLLM
H
Hao Jiang
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University; University of Electronic Science and Technology of China
J
Jixian Wu
New York University
Zheng Lu
Zheng Lu
University of Nottingham Ningbo China
Computer VisionNatural Language ProcessingMachine Learning
B
Bofan Zhu
University of Electronic Science and Technology of China
R
Renyuan Li
University of Electronic Science and Technology of China
Shanghang Zhang
Shanghang Zhang
Peking University
Embodied AIFoundation Models