🤖 AI Summary
This study addresses the challenge of post-hoc signal interference in multimodal risk recognition during the brief pre-violence window. Based on XD-Violence, we construct a pre-event dataset and introduce a novel variable-length slicing paradigm that exclusively leverages pre-incident behavioral cues. Methodologically, visual Transformer models such as DeiT-Tiny are employed alongside keypoint tracking and audio alignment techniques to deeply fuse facial, bodily, and acoustic features for short-term risk assessment. Ablation studies confirm the complementary advantages of these multimodal features. The optimal configuration achieves 91.21% accuracy, 88.96% balanced accuracy, and 96.38% ROC-AUC on the test set, demonstrating precise early warning capabilities for violence risk.
📝 Abstract
Detecting violence after it begins is important from recognizing behavioral cues that appear immediately beforehand. This work studies short-horizon pre-incident risk recognition from multimodal video signals. We construct a binary Normal-versus-Risky setting from temporally annotated XD-Violence clips, using 443 samples with source-level separation across training, validation, and test sets. Each sample consists of a variable-length pre-incident clip, with its duration determined by the observable behavioral context preceding the incident. The inci- dent itself is excluded from all input clips. We evaluate three complementary information sources: facial-region appearance, temporally aligned audio, and body-motion features derived from tracked keypoints. Controlled ablations are performed with Swin-Tiny, ViT-Tiny, and DeiT-Tiny to measure the contribution of each modality under the same split. Results show that combining all modalities is more effective than using any other combination alone. The best configuration, Deit-Tiny with audio, facial appearance, and motion, achieves 91.21% accuracy, 88.96% balanced accuracy, 93.65% F1-score, and 96.38% ROC-AUC on the held-out test set. These results suggest that complementary appearance, acoustic, and kinematic cues provide useful evidence for recognizing elevated pre-incident risk.