multi-frame tracking

Designs and implements systems that detect, associate, and maintain identities of objects across sequences of video frames, including algorithms for motion modeling, temporal feature extraction, and data association for multi-frame tracking. Builds and analyzes end-to-end video processing and analytics pipelines — including video preprocessing, feature extraction, super-resolution, temporal video modeling, real-time media pipelines, and integration with video generation/synthesis tools and pipelines — to support long-video and real-time tracking applications.

multi-frametracking

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.94
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$215K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Segment Anything for Video: A Comprehensive Review of Video Object Segmentation and Tracking from Past to Future

Jul 30, 2025
GX
Guoping Xu
🏛️ University of Texas Southwestern Medical Center | University of Pennsylvania | Mayo Clinic

Video Object Segmentation and Tracking (VOST) suffers from poor temporal consistency, limited generalization, and low computational efficiency. To address these challenges, this paper presents a systematic survey of SAM- and SAM2-based VOST methods and proposes a foundation-model-driven paradigm: (1) a motion-aware memory selection mechanism to mitigate error accumulation; (2) trajectory-guided prompting to enhance temporal robustness; and (3) integration of streaming memory architecture, dynamic feature extraction, and motion prediction for efficient inference. Experiments demonstrate that the framework achieves superior trade-offs between accuracy and real-time performance. Furthermore, the study identifies critical bottlenecks—including memory redundancy, suboptimal prompt efficiency, and long-term error propagation—offering, for the first time, a structured technical roadmap and concrete future research directions for adapting SAM to VOST.

Addressing domain generalization and temporal consistency in video segmentationEnhancing accuracy with motion-aware memory and trajectory-guided promptingImproving computational efficiency in video object tracking

Must-Read Papers

Most classic and influential ideas
View more

STAC: Leveraging Spatio-Temporal Data Associations For Efficient Cross-Camera Streaming and Analytics

Jan 27, 2024
VV
Volodymyr Vakhniuk
🏛️ University of Illinois at Urbana-Champaign

To address the bandwidth–accuracy trade-off in multi-view video analytics for distributed IoT camera networks, this paper proposes STAC, a lightweight cross-camera surveillance system. Methodologically, STAC introduces the first ReID algorithm featuring full-scale spatiotemporal feature learning, integrating frame-level dynamic filtering with FFmpeg libx264-based adaptive video compression to significantly reduce transmission and computational overhead while preserving detection, tracking, and re-identification accuracy. It further incorporates omni-scale feature extraction and explicit spatiotemporal correlation modeling to enhance cross-camera target consistency representation. Evaluated on the AICity 2023 multi-camera dataset, STAC achieves state-of-the-art cross-camera pedestrian re-identification performance (mAP improved by 12.3%), compresses video stream volume by 78%, and maintains end-to-end inference latency below 200 ms—satisfying real-time operational requirements.

Improving object tracking accuracy under network constraintsMinimizing redundant visual data without degrading model performanceReducing bandwidth demands in multi-camera video analytics

To address storage redundancy and inefficient retrieval in video surveillance, this paper proposes an activity-driven intelligent dynamic scene analysis system. Methodologically, it introduces a novel hybrid motion segmentation strategy integrating adaptive background modeling, Lucas-Kanade optical flow, and a deep temporal model (LSTM); combines multi-scale context-aware object detection (based on YOLO/SSD) with illumination-invariant feature optimization; and enhances tracking robustness via Kalman filtering and Siamese network-based re-identification. Evaluated on real-world CCTV footage, the system achieves significant improvements in critical event detection—e.g., person appearance and anomalous behavior—with average precision and recall gains of 12.3%. It operates at real-time speed (≥25 FPS), reduces video storage overhead by 63.7%, and enables efficient content-based retrieval and long-term archival.

Develop robust video surveillance systemImprove object detection and tracking accuracyOptimize storage and enhance digital searches

A Modular Pipeline for 3D Object Tracking Using RGB Cameras

Mar 06, 2025
LB
Lars Bredereke
🏛️ University Bremen

This work addresses the challenges of 3D multi-object tracking (MOT) with small targets, high occlusion density, frequent entry/exit, and unknown camera poses in multi-view RGB setups. We propose a lightweight, modular 3D MOT framework integrating multi-view geometry, feature matching, PnP-based pose estimation, and extended Kalman filtering (EKF). Our key methodological contribution is the first automatic instance-aware EKF architecture supporting dynamic object creation and deletion—eliminating the need for prior camera calibration while enabling real-time, covariance-aware 3D trajectory estimation. Evaluated on the Table Setting Dataset comprising over ten million frames, our approach achieves centimeter-level average localization accuracy across hundreds of trials; the estimated covariance matrices effectively quantify positional uncertainty. The framework significantly enhances robustness and deployability of 3D MOT in complex, unstructured environments.

Addresses challenges in tracking multiple objects with temporary occlusions.Develops a modular pipeline for 3D object tracking using RGB cameras.Enables scalable, accurate 3D trajectory calculation with minimal human input.

To address the high cost and low efficiency of manual annotation in video object detection and segmentation, this paper proposes a lightweight, end-to-end automated annotation framework. Methodologically, it introduces the first deep integration of YOLOv8-based video tracking and SAM-based interactive segmentation, augmented by an adaptive frame-sampling strategy and a Gradio-powered visual interface, resulting in a modular, open-source, and extensible prototype system. Key contributions include: (i) a lightweight co-design of tracking and segmentation models; (ii) efficient, interactive annotation support for multi-object, long-duration videos; and (iii) rapid generation of annotated datasets enabling a closed-loop annotation–training pipeline. Experiments on multiple public video benchmarks demonstrate over 10× higher annotation throughput compared to manual labeling, strong inter-annotator consistency, and substantial performance gains for downstream detection and segmentation models trained on the generated data.

Automate video annotation for computer vision modelsDevelop efficient video tracking and segmentation toolReduce time and resources for labeled data generation

This work addresses the high computational cost of existing RGB-based video object tracking methods, which hinders their large-scale deployment. The authors propose MVTrack, the first approach to achieve efficient tracking solely using motion vectors extracted from H.264 compressed bitstreams, entirely bypassing pixel-domain processing. MVTrack integrates a lightweight motion vector field detector (MVDet) with a minimalistic motion association module (MVLink) to enable accurate tracking without video decoding. Evaluated on the VIRAT dataset, MVTrack outperforms YOLOv2tiny in tracking accuracy while using 60× fewer parameters, requiring 40× lower FLOPs, and achieving 8.6× faster CPU inference speed, thereby significantly advancing the practicality and efficiency of compressed-domain tracking.

compressed bitstreamscomputational costmoving object detection

Latest Papers

What's happening recently
View more

This work addresses the degradation in tracking accuracy caused by conventional frame-level sampling in static videos, which often discards fine-grained motion information. To overcome this limitation, the authors propose a tile-based polyomino modeling approach that enables efficient trajectory extraction through spatiotemporal joint pruning. The method employs a three-stage pipeline—tile classification, integer linear programming (ILP)-based pruning, and canvas packing—and supports user-specified detectors while adaptively adjusting the sampling strategy under given accuracy constraints. Experiments across seven static video datasets demonstrate that, with trajectory accuracy loss capped at 5%, the system achieves up to 17.4× higher throughput than state-of-the-art methods and 68.8× improvement over the full-frame, frame-by-frame baseline.

efficient video processinghigh-fidelity trackingspatiotemporal sampling

This work addresses the challenges of identity preservation and frequent identity switches in multi-object tracking, which are exacerbated by highly similar appearances and dense object distributions. To tackle these issues, the paper proposes a video-level association re-identification framework that formulates re-identification as a global trajectory matching task. By aggregating historical trajectory features with current detections for video-level association, the method enhances individual discriminability without requiring additional annotations. It further introduces two novel components: Frame-common Appearance Estimation (FCAE) and a Corresponding Appearance Suppression (CAS) mechanism. Evaluated on the BEE24 dataset, the approach achieves notable improvements, including a 1.1-point gain in HOTA, a 2.6-point increase in AssR, and a 28% reduction in identity switches, significantly outperforming existing state-of-the-art methods.

highly similar objectsidentity preservationmulti-object tracking

This work addresses the challenge of controllably editing the motion trajectory of a target object in videos while preserving the original scene content. To this end, the authors propose a two-stage framework: first, a cross-view motion transformation module maps a user-specified trajectory—provided only in the initial frame—into per-frame bounding boxes that account for camera motion; second, a motion-conditioned video resynthesis module generates the object along this trajectory while maintaining background consistency. By eliminating the need for complex point-trajectory inputs, the method significantly enhances user-friendliness and temporal coherence. Experiments demonstrate that the approach produces more realistic, temporally consistent, and controllable motion edits on diverse real-world videos compared to existing image-to-video or video-to-video methods.

camera motionmotion pathobject motion editing

This study addresses the challenge faced by low-budget football teams in accessing professional-grade player and match analytics due to the high cost of multi-camera setups or GPS tracking systems. To overcome this barrier, the authors propose an end-to-end single-camera visual system that, for the first time, enables joint detection and tracking of players, referees, goalkeepers, and the ball directly from a single broadcast video stream. The system leverages a YOLO-based object detector integrated with the ByteTrack multi-object tracking algorithm to form a lightweight computer vision pipeline. Experimental results demonstrate high precision, recall, and strong mAP50 performance in tracking players and referees, confirming the feasibility of efficiently extracting spatial motion data from standard broadcast footage and substantially lowering the entry threshold for football performance analysis.

computer visionmulti-class detectionobject tracking

Hot Scholars

PW

Pengfei Wan

Head of Kling Video Generation Models, Kuaishou Technology
Generative ModelsComputer VisionMultimodal AIComputer Graphics
ZL

Ziwei Liu

Associate Professor, Nanyang Technological University
Computer VisionMachine LearningComputer Graphics
YS

Yujun Shen

Ant Group
Generative ModelingComputer VisionDeep Learning
QC

Qifeng Chen

HKUST
Computational PhotographyImage SynthesisGenerative AIAutonomous Driving
GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation