DecoMoE: Decoupling Visual Propagation and Expert Computation for Efficient Multimodal MoE Inference

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive inference costs of multimodal Mixture-of-Experts (MoE) models caused by long visual sequences. We propose a two-dimensional structured compression framework that decouples visual propagation from expert computation to eliminate redundancy across both token and expert dimensions. Methodologically, we introduce a Sample-Adaptive Visual Boundary (SAVB) prediction module to dynamically determine the visual exit layer, alongside a Route-Calibrated Expert Prefix (RCEP) retention mechanism that offline-reorders experts and preserves the shortest prefix based on routing quality, enabling input-dependent dynamic pruning. Evaluated on Qwen3-VL-MoE, our approach retains 97.91% of baseline performance with only 16.73 TFLOPs of computation, halving latency and achieving a 1.69× speedup.
📝 Abstract
Multimodal mixture-of-experts (MoE) models combine sparse expert activation with visual-language capabilities, yet their inference remains costly because long visual-token sequences repeatedly incur attention, routing, dispatch, and expert-MLP computation. Existing methods typically compress either the token or expert dimension, leaving redundancy along the other. Our analysis reveals two complementary regularities: the depth required for visual propagation varies across inputs, while text-token routing exhibits concentrated and recurrent expert-importance patterns. Based on these observations, we propose DecoMoE, a two-dimensional structured compression framework that decouples visual propagation from expert computation. The Sample-Adaptive Visual Boundary (SAVB) predicts an input-dependent visual-exit layer at which the visual-token block is removed. The Routing-Calibrated Expert Prefix (RCEP) reorders experts offline using text-token routed mass and, from this predicted exit layer onward, retains at each MoE layer the shortest contiguous prefix covering a target routed-mass fraction. We evaluate DecoMoE on Qwen3-VL-MoE and InternVL3.5-30B-A3B across six benchmarks. On Qwen3-VL-MoE, DecoMoE retains 97.91% of dense-baseline performance while reducing computation from 27.06 to 16.73 TFLOPs and latency from 0.44 to 0.26 seconds, yielding a 1.69x speedup. Code will be available at https://github.com/ShawnTan86/DecoMoE.
Problem

Research questions and friction points this paper is trying to address.

Multimodal MoE
Inference Efficiency
Visual Token Redundancy
Expert Computation
Structured Compression
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal MoE
Structured Compression
Visual-Expert Decoupling
Sample-Adaptive Visual Boundary
Routing-Calibrated Expert Prefix
🔎 Similar Papers
No similar papers found.