๐ค AI Summary
Addressing the challenges of modeling spatial context, reliance on pre-trained models, and poor interpretability in street-scene video anomaly detection, this paper proposes an unsupervised spatial consistency modeling framework. First, Gaussian Mixture Modeling (GMM) is applied to high-resolution feature maps for unsupervised clustering, jointly discovering object-level spatial attributes and spatially consistent regions. Subsequently, an inter-object spatial relation graph is constructed to generate pixel-level normality heatmaps. The method requires neither pre-trained segmentation models nor human annotations. Evaluated on the Street Scene dataset, it achieves state-of-the-art performance while reducing parameter count by one to two orders of magnitude. Moreover, it produces high-resolution, semantically interpretable anomaly localization mapsโenhancing model transparency and enabling efficient deployment.
๐ Abstract
We describe a method for modeling spatial context to enable video anomaly detection. The main idea is to discover regions that share similar object-level activities by clustering joint object attributes using Gaussian mixture models. We demonstrate that this straightforward approach, using orders of magnitude fewer parameters than competing models, achieves state-of-the-art performance in the challenging spatial-context-dependent Street Scene dataset. As a side benefit, the high-resolution discovered regions learned by the model also provide explainable normalcy maps for human operators without the need for any pre-trained segmentation model.