Score
Designs and implements systems that detect, associate, and maintain identities of objects across sequences of video frames, including algorithms for motion modeling, temporal feature extraction, and data association for multi-frame tracking. Builds and analyzes end-to-end video processing and analytics pipelines — including video preprocessing, feature extraction, super-resolution, temporal video modeling, real-time media pipelines, and integration with video generation/synthesis tools and pipelines — to support long-video and real-time tracking applications.
本文针对多目标跟踪中的评估不一致问题,通过系统回顾基于检测的跟踪方法,并从最小基线跟踪器出发公平评估各方法贡献。
Video Object Segmentation and Tracking (VOST) suffers from poor temporal consistency, limited generalization, and low computational efficiency. To address these challenges, this paper presents a systematic survey of SAM- and SAM2-based VOST methods and proposes a foundation-model-driven paradigm: (1) a motion-aware memory selection mechanism to mitigate error accumulation; (2) trajectory-guided prompting to enhance temporal robustness; and (3) integration of streaming memory architecture, dynamic feature extraction, and motion prediction for efficient inference. Experiments demonstrate that the framework achieves superior trade-offs between accuracy and real-time performance. Furthermore, the study identifies critical bottlenecks—including memory redundancy, suboptimal prompt efficiency, and long-term error propagation—offering, for the first time, a structured technical roadmap and concrete future research directions for adapting SAM to VOST.
To address the bandwidth–accuracy trade-off in multi-view video analytics for distributed IoT camera networks, this paper proposes STAC, a lightweight cross-camera surveillance system. Methodologically, STAC introduces the first ReID algorithm featuring full-scale spatiotemporal feature learning, integrating frame-level dynamic filtering with FFmpeg libx264-based adaptive video compression to significantly reduce transmission and computational overhead while preserving detection, tracking, and re-identification accuracy. It further incorporates omni-scale feature extraction and explicit spatiotemporal correlation modeling to enhance cross-camera target consistency representation. Evaluated on the AICity 2023 multi-camera dataset, STAC achieves state-of-the-art cross-camera pedestrian re-identification performance (mAP improved by 12.3%), compresses video stream volume by 78%, and maintains end-to-end inference latency below 200 ms—satisfying real-time operational requirements.
To address storage redundancy and inefficient retrieval in video surveillance, this paper proposes an activity-driven intelligent dynamic scene analysis system. Methodologically, it introduces a novel hybrid motion segmentation strategy integrating adaptive background modeling, Lucas-Kanade optical flow, and a deep temporal model (LSTM); combines multi-scale context-aware object detection (based on YOLO/SSD) with illumination-invariant feature optimization; and enhances tracking robustness via Kalman filtering and Siamese network-based re-identification. Evaluated on real-world CCTV footage, the system achieves significant improvements in critical event detection—e.g., person appearance and anomalous behavior—with average precision and recall gains of 12.3%. It operates at real-time speed (≥25 FPS), reduces video storage overhead by 63.7%, and enables efficient content-based retrieval and long-term archival.
This work addresses the challenges of 3D multi-object tracking (MOT) with small targets, high occlusion density, frequent entry/exit, and unknown camera poses in multi-view RGB setups. We propose a lightweight, modular 3D MOT framework integrating multi-view geometry, feature matching, PnP-based pose estimation, and extended Kalman filtering (EKF). Our key methodological contribution is the first automatic instance-aware EKF architecture supporting dynamic object creation and deletion—eliminating the need for prior camera calibration while enabling real-time, covariance-aware 3D trajectory estimation. Evaluated on the Table Setting Dataset comprising over ten million frames, our approach achieves centimeter-level average localization accuracy across hundreds of trials; the estimated covariance matrices effectively quantify positional uncertainty. The framework significantly enhances robustness and deployability of 3D MOT in complex, unstructured environments.
To address the high cost and low efficiency of manual annotation in video object detection and segmentation, this paper proposes a lightweight, end-to-end automated annotation framework. Methodologically, it introduces the first deep integration of YOLOv8-based video tracking and SAM-based interactive segmentation, augmented by an adaptive frame-sampling strategy and a Gradio-powered visual interface, resulting in a modular, open-source, and extensible prototype system. Key contributions include: (i) a lightweight co-design of tracking and segmentation models; (ii) efficient, interactive annotation support for multi-object, long-duration videos; and (iii) rapid generation of annotated datasets enabling a closed-loop annotation–training pipeline. Experiments on multiple public video benchmarks demonstrate over 10× higher annotation throughput compared to manual labeling, strong inter-annotator consistency, and substantial performance gains for downstream detection and segmentation models trained on the generated data.
This work addresses the high computational cost of existing RGB-based video object tracking methods, which hinders their large-scale deployment. The authors propose MVTrack, the first approach to achieve efficient tracking solely using motion vectors extracted from H.264 compressed bitstreams, entirely bypassing pixel-domain processing. MVTrack integrates a lightweight motion vector field detector (MVDet) with a minimalistic motion association module (MVLink) to enable accurate tracking without video decoding. Evaluated on the VIRAT dataset, MVTrack outperforms YOLOv2tiny in tracking accuracy while using 60× fewer parameters, requiring 40× lower FLOPs, and achieving 8.6× faster CPU inference speed, thereby significantly advancing the practicality and efficiency of compressed-domain tracking.
This work addresses the degradation in tracking accuracy caused by conventional frame-level sampling in static videos, which often discards fine-grained motion information. To overcome this limitation, the authors propose a tile-based polyomino modeling approach that enables efficient trajectory extraction through spatiotemporal joint pruning. The method employs a three-stage pipeline—tile classification, integer linear programming (ILP)-based pruning, and canvas packing—and supports user-specified detectors while adaptively adjusting the sampling strategy under given accuracy constraints. Experiments across seven static video datasets demonstrate that, with trajectory accuracy loss capped at 5%, the system achieves up to 17.4× higher throughput than state-of-the-art methods and 68.8× improvement over the full-frame, frame-by-frame baseline.
This work addresses the challenges of identity preservation and frequent identity switches in multi-object tracking, which are exacerbated by highly similar appearances and dense object distributions. To tackle these issues, the paper proposes a video-level association re-identification framework that formulates re-identification as a global trajectory matching task. By aggregating historical trajectory features with current detections for video-level association, the method enhances individual discriminability without requiring additional annotations. It further introduces two novel components: Frame-common Appearance Estimation (FCAE) and a Corresponding Appearance Suppression (CAS) mechanism. Evaluated on the BEE24 dataset, the approach achieves notable improvements, including a 1.1-point gain in HOTA, a 2.6-point increase in AssR, and a 28% reduction in identity switches, significantly outperforming existing state-of-the-art methods.
This work addresses the challenge of controllably editing the motion trajectory of a target object in videos while preserving the original scene content. To this end, the authors propose a two-stage framework: first, a cross-view motion transformation module maps a user-specified trajectory—provided only in the initial frame—into per-frame bounding boxes that account for camera motion; second, a motion-conditioned video resynthesis module generates the object along this trajectory while maintaining background consistency. By eliminating the need for complex point-trajectory inputs, the method significantly enhances user-friendliness and temporal coherence. Experiments demonstrate that the approach produces more realistic, temporally consistent, and controllable motion edits on diverse real-world videos compared to existing image-to-video or video-to-video methods.
研究提出TRACE框架,通过预训练视频模型的速度响应提取表示并建模帧间一致性,有效检测AI生成的视频,超越现有方法。
This study addresses the challenge faced by low-budget football teams in accessing professional-grade player and match analytics due to the high cost of multi-camera setups or GPS tracking systems. To overcome this barrier, the authors propose an end-to-end single-camera visual system that, for the first time, enables joint detection and tracking of players, referees, goalkeepers, and the ball directly from a single broadcast video stream. The system leverages a YOLO-based object detector integrated with the ByteTrack multi-object tracking algorithm to form a lightweight computer vision pipeline. Experimental results demonstrate high precision, recall, and strong mAP50 performance in tracking players and referees, confirming the feasibility of efficiently extracting spatial motion data from standard broadcast footage and substantially lowering the entry threshold for football performance analysis.