Detection, Retrieval, and Explanation Unified: A Violence Detection System Based on Knowledge Graphs and GAT

πŸ“… 2025-01-07
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing violence detection systems suffer from poor interpretability, functional limitations (e.g., classification or retrieval only), and low accessibility for non-expert users (e.g., middle school students). To address these issues, we propose the first end-to-end, multi-task unified framework for surveillance video analysis, integrating knowledge graphs (KGs) and graph attention networks (GATs) to jointly support violence detection, cross-modal retrieval, and interpretable reasoning. Our approach introduces a verifiable KG-based reasoning mechanism, synergistically combining ImageBind’s multimodal embeddings, lightweight temporal modeling, and video-LLM-driven natural language explanation generation. Evaluated on XD-Violence and UCF-Crime, our method achieves state-of-the-art performance. It significantly enhances both explanatory transparency and reasoning efficiency. Moreover, our framework uncovers, for the first time, a previously unreported social behavioral pattern: a negative correlation between bystander count and violence incidence.

Technology Category

Computer Vision: Multi-modal VisionKnowledge Representation and Reasoning: Geometric, Spatial, and Temporal ReasoningIntelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
πŸ“ Abstract
Recently, violence detection systems developed using unified multimodal models have achieved significant success and attracted widespread attention. However, most of these systems face two critical challenges: the lack of interpretability as black-box models and limited functionality, offering only classification or retrieval capabilities. To address these challenges, this paper proposes a novel interpretable violence detection system, termed the Three-in-One (TIO) System. The TIO system integrates knowledge graphs (KG) and graph attention networks (GAT) to provide three core functionalities: detection, retrieval, and explanation. Specifically, the system processes each video frame along with text descriptions generated by a large language model (LLM) for videos containing potential violent behavior. It employs ImageBind to generate high-dimensional embeddings for constructing a knowledge graph, uses GAT for reasoning, and applies lightweight time series modules to extract video embedding features. The final step connects a classifier and retriever for multi-functional outputs. The interpretability of KG enables the system to verify the reasoning process behind each output. Additionally, the paper introduces several lightweight methods to reduce the resource consumption of the TIO system and enhance its efficiency. Extensive experiments conducted on the XD-Violence and UCF-Crime datasets validate the effectiveness of the proposed system. A case study further reveals an intriguing phenomenon: as the number of bystanders increases, the occurrence of violent behavior tends to decrease.
Problem

Research questions and friction points this paper is trying to address.

Violence Detection
Transparency
Multifunctionality
Innovation

Methods, ideas, or system contributions that make the work stand out.

Knowledge Graph
Graph Attention Network (GAT)
Efficient Resource Utilization