🤖 AI Summary
This study addresses the unclear discriminative mechanisms and high computational costs of multimodal large models in video anomaly detection. By leveraging routing signals from a sparse Mixture-of-Experts (MoE) architecture, this work reveals the distribution of anomalous evidence to construct an efficient detection framework. Methodologically, it identifies critical features by discovering dynamically routed expert subnetworks, and introduces routing-modulated fusion alongside a temporal-aware network to enhance spatiotemporal modeling capabilities. Experimental results demonstrate that the proposed framework achieves state-of-the-art performance using only 5% weakly annotated data. Furthermore, its inference speed significantly surpasses that of dense backbone networks, successfully unifying high detection accuracy with low computational overhead.
📝 Abstract
Intermediate-layer features from multimodal large language models have shown strong potential for video anomaly detection (VAD), yet the origin of their discriminative power remains unclear. We study this question using sparse mixture-of-experts (MoE) models, whose explicit expert structure and sparse activation make their internal computation easier to inspect. With a fully frozen backbone and no additional training, we find that anomaly-related evidence is concentrated in a small set of experts. These experts recur across layers, spontaneously specialize in different anomaly types, and together form a dynamic routing subnetwork. We further show that the output channels most strongly influenced by these experts are also the hidden dimensions that contain the most anomaly-relevant information. Routing statistics can therefore serve as an internal anomaly cue that complements semantic features.Based on these findings, we propose RoMod, an efficient VAD framework trained with only \(5\%\) of weakly labeled videos. RoMod includes a Routing-Modulated Fusion module, RoMF, and a Routing-aware Temporal Network, RoTN. RoMF uses routing signals to adaptively recalibrate hidden semantic channels. Its design also prevents the routing branch from bypassing semantic features and making predictions on its own. RoTN captures the temporal evolution of anomalies from onset to persistence and termination. Experiments on three benchmarks show that RoMod achieves state-of-the-art performance while running substantially faster than dense backbones of comparable size.