Analytical and Convolutional Neural Network-Based Motion-Vector Propagation for Efficient Video Object Detection

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent trade-off between high accuracy and low latency in continuous video analysis on embedded devices by proposing a hardware-software co-optimization approach. By reusing codec motion vectors for inter-frame object tracking, we design two propagation models: an analytical method and a parallel-friendly CNN-based approach. Furthermore, the temporal algorithms are jointly optimized with the execution pipeline of edge GPU/DLA accelerators. Experimental results demonstrate that the analytical model reduces latency to 9.03 ms and energy consumption by 36.4%, while the CNN-based model improves recall rate and further lowers power consumption. Ultimately, this work achieves synergistic enhancements in both accuracy and efficiency for edge-based video analytics.
📝 Abstract
Continuous video analytics requires accurate localization at low latency within embedded power budgets. This paper presents a hardware-software design methodology that reuses codec motion vectors (MVs) between detector invocations. Two alternative models support translation and scale changes: analytical motion-vector propagation (Analytical-MV) and learned propagation using a convolutional neural network (CNN) (CNN-MV). The learned model uses convolutional operations and independent object updates suited to parallel execution on an edge graphics processing unit (GPU). Analytical-MV combines a harmonic-mean precision-recall score (F1) of 0.909 with a mean end-to-end latency of 9.03 ms and an energy consumption of 0.177 J per frame, yielding the lowest latency and energy among the evaluated configurations. Relative to detection on every frame, it reduces mean latency by 25.9% and energy per frame by 36.4%. CNN-MV offers a different trade-off: its fastest configuration raises recall from 0.871 for Analytical-MV to 0.890 and lowers mean power from 19.64 to 17.32 W, while achieving a latency of 18.42 ms and an energy consumption of 0.319 J per frame. It is therefore useful when recall or operating power is more important than minimum latency and energy. Execution on a deep learning accelerator (DLA) further reduces time-averaged GPU utilization relative to GPU execution. Host-processing optimization substantially improves both latency and energy, demonstrating the value of jointly designing temporal models and their execution pipelines.
Problem

Research questions and friction points this paper is trying to address.

Video Object Detection
Motion Vector Propagation
Low Latency
Embedded Systems
Energy Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Motion-Vector Propagation
Video Object Detection
Edge Computing
Hardware-Software Co-design
Convolutional Neural Network
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Ashiyana Abdul Majeed
Ashiyana Abdul Majeed
PhD Student, Khalifa University
Mahmoud Meribout
Mahmoud Meribout
Khalifa University of Science & Technology
Embedded Systems and Instrumentation
N
Neethu Joseph
Department of Computer and Information Engineering, Khalifa University, Abu Dhabi, UAE