From Skeletons to Semantics: Design and Deployment of a Hybrid Edge-Based Action Detection System for Public Safety

📅 2026-03-31
📈 Citations: 0
Influential: 0
📄 PDF

career value

245K/year
🤖 AI Summary
This work addresses the challenges of high latency, privacy risks, and computational bottlenecks that hinder efficient violence detection in existing edge-based video analytics systems for public safety. To overcome these limitations, the authors propose a hybrid edge action detection architecture that innovatively integrates skeleton-based pose analysis with the semantic reasoning capabilities of vision-language foundation models. Deployed on GPU-enabled edge devices, the system enables low-latency, low-overhead real-time inference while supporting context-aware and zero-shot detection. It dynamically orchestrates motion- and semantics-driven paradigms according to scene requirements. Experimental results demonstrate that the approach achieves a favorable trade-off among latency, resource consumption, and accuracy, highlighting the complementary strengths of the two modalities and offering a practical solution for real-world deployment.

Technology Category

Application Category

📝 Abstract
Public spaces such as transport hubs, city centres, and event venues require timely and reliable detection of potentially violent behaviour to support public safety. While automated video analysis has made significant progress, practical deployment remains constrained by latency, privacy, and resource limitations, particularly under edge-computing conditions. This paper presents the design and demonstrator-based deployment of a hybrid edge-based action detection system that combines skeleton-based motion analysis with vision-language models for semantic scene interpretation. Skeleton-based processing enables continuous, privacy-aware monitoring with low computational overhead, while vision-language models provide contextual understanding and zero-shot reasoning capabilities for complex and previously unseen situations. Rather than proposing new recognition models, the contribution focuses on a system-level comparison of both paradigms under realistic edge constraints. The system is implemented on a GPU-enabled edge device and evaluated with respect to latency, resource usage, and operational trade-offs using a demonstrator-based setup. The results highlight the complementary strengths and limitations of motioncentric and semantic approaches and motivate a hybrid architecture that selectively augments fast skeletonbased detection with higher-level semantic reasoning. The presented system provides a practical foundation for privacy-aware, real-time video analysis in public safety applications.
Problem

Research questions and friction points this paper is trying to address.

action detection
public safety
edge computing
privacy-aware monitoring
real-time video analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

hybrid edge-based system
skeleton-based action detection
vision-language models
zero-shot reasoning
privacy-aware video analysis
🔎 Similar Papers
No similar papers found.