VESSI - VLM-Enhanced Support for Surveillance and Investigations

πŸ“… 2026-10-08
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitations of traditional surveillance systems, which lack semantic analysis capabilities and thereby constrain situational awareness and investigative efficiency. We propose a video semantic transformation framework based on vision-language models (VLMs) that leverages prompt engineering to map visual features into textual descriptions, enabling automated surveillance analysis. Furthermore, we introduce a Composite Model Utility Score (CMUS), a reference-free metric designed for the systematic evaluation of VLM performance. Experimental results demonstrate that the proposed approach accurately identifies target activities in over 66% of videos while reducing review time by more than 85%. These findings indicate that our framework significantly enhances both the automation level and operational flexibility of surveillance analysis.
πŸ“ Abstract
Automated video surveillance analysis has become a critical component of intelligence infrastructures and Law Enforcement agencies. Traditional systems lack the semantic module for comprehensive situational awareness and forensic tasks, limiting their ability to interpret events meaningfully or support post-incident investigations. This slows operational insight and increases the burden on human analysts. Recent advances in Vision-Language Models (VLMs) offer promising pathways to bridge this gap. To address this, we propose VLM-Enhanced Support for Surveillance and Investigations (VESSI), a VLM-based framework designed to enhance automated video surveillance analysis through prompt-driven interrogation of video sequences where salient visual features are converted into textual descriptions. We test our framework with four state-of-the-art models. Since most datasets for this task are unlabeled, we also propose the Composite Model Utility Score (CMUS) to assess VLM performance. Experimental results show that our solution substantially improves the analysis capabilities of human operators and enhances the flexibility of automated surveillance systems. In our evaluation, the most reliable model flagged potentially relevant activity in more than 66% of the videos while reducing review time by more than 85%, offering a practical balance between selectivity and efficiency. The model ordering produced by the reference-free CMUS evaluation was reproduced by the normal-video CMUS evaluation and matched the false-positive-rate ordering obtained from 5,909 manually referenced frames. This agreement supports the operational use of the score within the evaluated setting.
Problem

Research questions and friction points this paper is trying to address.

Video Surveillance
Situational Awareness
Forensic Analysis
Vision-Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Video Surveillance
Composite Model Utility Score
Reference-free Evaluation
Prompt-driven Interrogation