🤖 AI Summary
This study addresses the prohibitive inference costs of multimodal Mixture-of-Experts (MoE) models caused by long visual sequences. We propose a two-dimensional structured compression framework that decouples visual propagation from expert computation to eliminate redundancy across both token and expert dimensions. Methodologically, we introduce a Sample-Adaptive Visual Boundary (SAVB) prediction module to dynamically determine the visual exit layer, alongside a Route-Calibrated Expert Prefix (RCEP) retention mechanism that offline-reorders experts and preserves the shortest prefix based on routing quality, enabling input-dependent dynamic pruning. Evaluated on Qwen3-VL-MoE, our approach retains 97.91% of baseline performance with only 16.73 TFLOPs of computation, halving latency and achieving a 1.69× speedup.
📝 Abstract
Multimodal mixture-of-experts (MoE) models combine sparse expert activation with visual-language capabilities, yet their inference remains costly because long visual-token sequences repeatedly incur attention, routing, dispatch, and expert-MLP computation. Existing methods typically compress either the token or expert dimension, leaving redundancy along the other. Our analysis reveals two complementary regularities: the depth required for visual propagation varies across inputs, while text-token routing exhibits concentrated and recurrent expert-importance patterns. Based on these observations, we propose DecoMoE, a two-dimensional structured compression framework that decouples visual propagation from expert computation. The Sample-Adaptive Visual Boundary (SAVB) predicts an input-dependent visual-exit layer at which the visual-token block is removed. The Routing-Calibrated Expert Prefix (RCEP) reorders experts offline using text-token routed mass and, from this predicted exit layer onward, retains at each MoE layer the shortest contiguous prefix covering a target routed-mass fraction. We evaluate DecoMoE on Qwen3-VL-MoE and InternVL3.5-30B-A3B across six benchmarks. On Qwen3-VL-MoE, DecoMoE retains 97.91% of dense-baseline performance while reducing computation from 27.06 to 16.73 TFLOPs and latency from 0.44 to 0.26 seconds, yielding a 1.69x speedup. Code will be available at https://github.com/ShawnTan86/DecoMoE.