🤖 AI Summary
Existing video anomaly detection methods rely on black-box pre-trained models, suffering from poor interpretability and limited capacity for explicit knowledge integration. To address this, we propose RuleVM, a weakly supervised, interpretable framework tailored for violent incident monitoring. RuleVM adopts a dual-branch paradigm: one branch leverages YOLO-World and vision-language alignment to extract visual-semantic representations; the other performs scene/action dual-channel feature disentanglement and data-driven association rule mining to enable rule-guided reasoning. The framework jointly supports coarse-grained anomaly classification and fine-grained attribution (e.g., “increased crowd size → elevated violence risk”). Evaluated on UCF-Crime and XD145, RuleVM achieves new state-of-the-art performance while providing human-verifiable decision rationales—thereby balancing detection accuracy and model transparency.
📝 Abstract
Recent advances in pre-trained models have demonstrated exceptional performance in video anomaly detection (VAD). However, most systems remain black boxes, lacking explainability during training and inference. A key challenge is integrating explicit knowledge into implicit models to create expert-driven, interpretable VAD systems. This paper introduces Rule-based Violence Monitoring (RuleVM), a novel weakly supervised video anomaly detection (WVAD) paradigm. RuleVM employs a dual-branch architecture: an implicit branch using visual features for coarse-grained binary classification, with feature extraction split into scene frames and action channels, and an explicit branch leveraging language-image alignment for fine-grained classification. The explicit branch utilizes the state-of-the-art YOLO-World model for object detection in video frames, with association rules mined from data as video descriptors. This design enables interpretable coarse- and fine-grained violence monitoring. Extensive experiments on two standard benchmarks show RuleVM outperforms state-of-the-art methods in both granularities. Notably, it reveals rules like increased violence risk with crowd size. Demo content is provided in the appendix.