From Explicit Rules to Implicit Reasoning in Weakly Supervised Video Anomaly Detection

📅 2024-10-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing video anomaly detection methods rely on black-box pre-trained models, suffering from poor interpretability and limited capacity for explicit knowledge integration. To address this, we propose RuleVM, a weakly supervised, interpretable framework tailored for violent incident monitoring. RuleVM adopts a dual-branch paradigm: one branch leverages YOLO-World and vision-language alignment to extract visual-semantic representations; the other performs scene/action dual-channel feature disentanglement and data-driven association rule mining to enable rule-guided reasoning. The framework jointly supports coarse-grained anomaly classification and fine-grained attribution (e.g., “increased crowd size → elevated violence risk”). Evaluated on UCF-Crime and XD145, RuleVM achieves new state-of-the-art performance while providing human-verifiable decision rationales—thereby balancing detection accuracy and model transparency.

Technology Category

Computer Vision: Interpretability, Explainability, and TransparencyData Mining & Knowledge Management: Anomaly/Outlier DetectionKnowledge Representation and Reasoning: Diagnosis and Abductive Reasoning

Application Category

Responsible Web: Machine-in-the-loop, human agency and autonomySemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalization
📝 Abstract
Recent advances in pre-trained models have demonstrated exceptional performance in video anomaly detection (VAD). However, most systems remain black boxes, lacking explainability during training and inference. A key challenge is integrating explicit knowledge into implicit models to create expert-driven, interpretable VAD systems. This paper introduces Rule-based Violence Monitoring (RuleVM), a novel weakly supervised video anomaly detection (WVAD) paradigm. RuleVM employs a dual-branch architecture: an implicit branch using visual features for coarse-grained binary classification, with feature extraction split into scene frames and action channels, and an explicit branch leveraging language-image alignment for fine-grained classification. The explicit branch utilizes the state-of-the-art YOLO-World model for object detection in video frames, with association rules mined from data as video descriptors. This design enables interpretable coarse- and fine-grained violence monitoring. Extensive experiments on two standard benchmarks show RuleVM outperforms state-of-the-art methods in both granularities. Notably, it reveals rules like increased violence risk with crowd size. Demo content is provided in the appendix.
Problem

Research questions and friction points this paper is trying to address.

Integrating explicit knowledge into implicit video anomaly detection models
Developing interpretable weakly supervised video anomaly detection systems
Combining coarse and fine-grained classification for violence monitoring
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-branch architecture for interpretable VAD
YOLO-World model for object detection
Language-image alignment for fine-grained classification
🔎 Similar Papers
No similar papers found.
Tamkang University | National Taipei University of Business | National Institute of Technology
Wen-Dong Jiang
Wen-Dong Jiang
Tamkang University
Interpretable Machine LeaningSmart CityMultimodal Explanation
C
Chih-Yung Chang
Department of Computer Science and Information Engineering, Tamkang University, New Taipei 25137, Taiwan
S
Ssu-Chi Kuai
Department of Information Management, National Taipei University of Business, Taipei 100025, Taiwan
Diptendu Sinha Roy
Diptendu Sinha Roy
Professor, Computer Science & Engineering, National Institute of Technology Meghalaya. India
IoTNext generation CommunicationsCloud ComputingMachine Learning