MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high inference cost of long sequences in multimodal large language models and the underexploited fine-grained redundancy therein, proposing a modality-aware width pruning method. This approach is the first to reveal and decouple modality interaction redundancy within attention pathways and FFN channels, pruning them independently based on first-order Taylor criteria. It achieves efficient acceleration by integrating LoRA-based recovery training, Triton sparse kernels, and compact FFN execution, while remaining compatible with existing token compression techniques. Evaluated on LLaVA-OneVision-7B, the proposed method delivers a 1.6× prefill speedup while retaining 99.7% of the original performance. When combined with token compression, the acceleration ratio further reaches 2.9×.
📝 Abstract
Multimodal large language models (MLLMs) incur substantial inference costs when processing long visual-textual sequences. While existing operation compression methods exploit modality-level redundancy, they largely treat computation within attention heads and shared feed-forward network (FFN) channels as unified units, leaving finer-grained redundancy underexplored. We find that redundancy varies both across modality-interaction paths within the same attention head and across visual and textual executions of the same FFN channel. Based on these findings, we propose Modality-aware Width-wise Operation Pruning (MWOP), which independently prunes visual-to-visual (V2V), text-to-visual (T2V), and text-to-text (T2T) attention paths within each layer, and separately selects FFN channels for visual and textual inputs. A first-order Taylor criterion guides the pruning process, with FFN importance re-evaluated after attention pruning and LoRA-based recovery training. To translate the resulting fine-grained sparsity into practical acceleration, we further develop path-sparse Triton attention kernels and compact visual-side FFN execution. MWOP preserves the token sequence while reducing attention and FFN computation, making it complementary to token compression and enabling simultaneous reduction of sequence length and per-token computation. On LLaVA-OneVision-7B, MWOP alone achieves a $1.6\times$ prefill speedup with 99.7\% average performance retention across 12 benchmarks. Combined with two representative token compression methods, it further increases their prefill speedups from $2.0\times$ and $1.9\times$ to $2.9\times$ and $2.7\times$, respectively. Results on Qwen2.5-VL-7B further demonstrate its applicability across architectures. The code is available at https://github.com/EIT-NLP/MWOP.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Inference Efficiency
Operation Pruning
Fine-grained Redundancy
Computational Cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Modality-aware Pruning
Operation Compression
Multimodal Large Language Models
Sparse Attention Kernels
Fine-grained Redundancy
🔎 Similar Papers
2024-08-29arXiv.orgCitations: 7
💼 Related Jobs
No related jobs found.
X
Xudong Wang
EIT-NLP Lab, Eastern Institute of Technology, Ningbo
H
Hao Wu
EIT-NLP Lab, Eastern Institute of Technology, Ningbo; Shanghai Jiao Tong University
H
Haozhe Hu
EIT-NLP Lab, Eastern Institute of Technology, Ningbo; Shanghai Jiao Tong University
P
Peiran Yin
EIT-NLP Lab, Eastern Institute of Technology, Ningbo; The Hong Kong Polytechnic University
X
Xinghao Chen
EIT-NLP Lab, Eastern Institute of Technology, Ningbo; The Hong Kong Polytechnic University
Yunpu Ma
Yunpu Ma
Ludwig Maximilian University of Munich
Foundation ModelsAgentic AITemporal Knowledge GraphQuantum AI
W
Wei Zhang
EIT-NLP Lab, Eastern Institute of Technology, Ningbo
Xiaoyu Shen
Xiaoyu Shen
Eastern Institute of Technology, Ningbo
language modelmulti-modal learningreasoning