SPIDER: Multi-Layer Semantic Token Pruning and Adaptive Sub-Layer Skipping in Multimodal Large Language Models.

📅 2026-09-17
🏛️ IEEE Transactions on Pattern Analysis and Machine Intelligence
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Multimodal large language models face inference efficiency bottlenecks arising from dual redundancy in data and computation. This work proposes SPIDER, a training-free acceleration framework that achieves fine-grained optimization through multi-layer semantic visual token pruning and an adaptive sub-layer skipping mechanism. Specifically, the method uncovers the shifting patterns of semantic focus across intermediate layers and quantifies the differential contributions of attention and feed-forward network modules to guide dynamic computational allocation. Evaluated on LLaVA-NeXT-7B, SPIDER reduces FLOPs by 79% while retaining 96% of baseline performance. Furthermore, it demonstrates broad compatibility with various mainstream architectures without requiring any fine-tuning.
📝 Abstract
Multimodal Large Language Models face significant efficiency challenges that stem from two distinct yet coupled sources: data redundancy and computational redundancy. While most methods focus on data redundancy by pruning visual tokens from the output of the visual encoder or computing redundancy in LLM decoders using blockwise importance, the finer-grained inter-layer representation shifts and the distribution differences within the layers themselves have not been fully explored. In this work, we comprehensively investigate this dual-level inefficiency. We posit that intermediate layer tokens from vision encoders should be considered for effective visual token pruning, as semantic focus shifts across layers, with middle-layer tokens capturing more detailed object-centric information that deeper layers may abstract away. Furthermore, we reveal the differential contributions of Attention and FFNs across distinct LLM decoder layers. Building upon these discoveries, we propose SPIDER, a training-free framework that integrates multi-layer Semantic visual token PrunIng with an aDaptive sub-layER skipping mechanism. Experimental evaluations demonstrate that SPIDER consistently maintains strong performance across various MLLM architectures and reduction ratios. For instance, on LLaVA-NeXT-7B, SPIDER reduces FLOPs by $80\%$ while maintaining 96$\%$ of the baseline performance.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Visual Token Pruning
Data Redundancy
Computational Redundancy
Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Large Language Models
Visual Token Pruning
Adaptive Sub-Layer Skipping
Training-Free
Computational Efficiency
T
Tianxiang Chen
End-User Intelligent Computing BU, Alibaba Cloud
Z
Zhentao Tan
Alibaba Group
Z
Zi Ye
Department of Computer Science, Maynooth University
Y
Yue Wu
Alibaba Group
X
Xiaobing Tu
End-User Intelligent Computing BU, Alibaba Cloud
J
Jinkui Ren
End-User Intelligent Computing BU, Alibaba Cloud
Xiantao Zhang
Xiantao Zhang
Beihang University
Large Language ModelsNatural Language ProcessingArtificial IntelligenceData Curation
Tao Gong
Tao Gong
University of Science and Technology of China
Computer VisionMachine Learning
Qi Chu
Qi Chu
University of Science and Technology of China
Computer visionArtificial intelligence security
Nenghai Yu
Nenghai Yu
University of Science and Technology of China
Computer VisionArtificial IntelligenceInformation Hiding
X
Xipeng Qiu
Fudan University
J
Jieping Ye
Alibaba Group