Score
Designs and implements models and processing pipelines that analyze video—both per-frame and across time—to detect and track people and assign visual body-type categories (e.g., shape or build) from appearance, pose, and silhouette cues. This work covers feature extraction, temporal aggregation over tracks, annotation and labeling schemes, training and evaluating classifiers on tracked segments, and engineering for robustness to occlusion, viewpoint and motion variability as well as dataset biases.
To address the high cost and low efficiency of manual annotation in video object detection and segmentation, this paper proposes a lightweight, end-to-end automated annotation framework. Methodologically, it introduces the first deep integration of YOLOv8-based video tracking and SAM-based interactive segmentation, augmented by an adaptive frame-sampling strategy and a Gradio-powered visual interface, resulting in a modular, open-source, and extensible prototype system. Key contributions include: (i) a lightweight co-design of tracking and segmentation models; (ii) efficient, interactive annotation support for multi-object, long-duration videos; and (iii) rapid generation of annotated datasets enabling a closed-loop annotation–training pipeline. Experiments on multiple public video benchmarks demonstrate over 10× higher annotation throughput compared to manual labeling, strong inter-annotator consistency, and substantial performance gains for downstream detection and segmentation models trained on the generated data.
To address storage redundancy and inefficient retrieval in video surveillance, this paper proposes an activity-driven intelligent dynamic scene analysis system. Methodologically, it introduces a novel hybrid motion segmentation strategy integrating adaptive background modeling, Lucas-Kanade optical flow, and a deep temporal model (LSTM); combines multi-scale context-aware object detection (based on YOLO/SSD) with illumination-invariant feature optimization; and enhances tracking robustness via Kalman filtering and Siamese network-based re-identification. Evaluated on real-world CCTV footage, the system achieves significant improvements in critical event detection—e.g., person appearance and anomalous behavior—with average precision and recall gains of 12.3%. It operates at real-time speed (≥25 FPS), reduces video storage overhead by 63.7%, and enables efficient content-based retrieval and long-term archival.
该研究通过融合生物识别、外观和3D身体特征,提出了一种基于多目标跟踪和个人再识别的视频摘要算法,以解决视频中身份识别与跟踪的问题。
该研究通过系统实验评估了多目标跟踪算法中检测和关联组件的贡献,揭示了检测质量对整体性能的影响远大于关联策略,并提供了设计优化MOT系统的实用指导。
To address frequent identity switches and trajectory discontinuities caused by occlusion in multi-object tracking, this paper proposes a lightweight online tracker. Methodologically: (1) it introduces the first multi-granularity appearance feature disentanglement framework, jointly modeling IoU, motion direction, clothing color/style, and ReID embeddings via an adaptive weighted composite distance metric; (2) it integrates a high-accuracy detector with a customized post-processing module to enhance detection robustness and trajectory smoothness. Experiments on mainstream benchmarks demonstrate significant improvements in MOTA (+2.1%) and IDF1 (+3.4%), a 32% reduction in ID switches, and markedly enhanced trajectory continuity under occlusion—while maintaining real-time inference speed (≥30 FPS), thus satisfying practical online deployment requirements.
Existing point tracking methods face significant challenges in real-world scenarios—including high motion complexity, frequent occlusions, and large object diversity—yet lack a systematic benchmark for evaluating robustness and failure modes. To address this, we introduce ITTO, the first high-challenge dynamic point tracking benchmark, comprising first-person real-world videos and multi-source data. We propose a multi-stage manual annotation protocol to precisely characterize motion patterns, occlusion events, and appearance variations. ITTO introduces a novel performance analysis protocol stratified by motion complexity and, for the first time, systematically exposes critical failure points of mainstream trackers—particularly in post-occlusion re-identification. Experiments reveal substantial performance degradation of state-of-the-art methods on ITTO, highlighting deficiencies in long-term occlusion handling and complex dynamic modeling. These findings provide concrete diagnostic insights and quantitative evaluation standards to guide algorithmic improvement.
This work proposes a novel end-to-end multi-object tracking approach that circumvents the limitations of traditional methods, which rely on separate detection and association modules and struggle to maintain identity consistency under occlusions due to complex external identity management. For the first time, the authors directly employ a large-scale text-to-video diffusion model (22B-parameter LTX-2.3) as a tracker, applying lightweight in-context LoRA fine-tuning to generate ID-map videos with temporally consistent identity colors. Identities are encoded implicitly at the pixel level, eliminating the need for explicit detection or association components. A windowed chain-generation strategy conditions each subsequent video window on the last frame of the previous one to preserve identity continuity. The method achieves 40.3 HOTA and 44.1 AssA on DanceTrack, substantially outperforming existing approaches, and successfully recovers identities in 42% of 383 occlusion events—dramatically surpassing appearance-embedding baselines, which exhibit near-zero recovery rates.
本文针对多目标跟踪中的评估不一致问题,通过系统回顾基于检测的跟踪方法,并从最小基线跟踪器出发公平评估各方法贡献。
Existing audio-visual speaker tracking methods suffer from limited performance in complex dynamic scenes, primarily due to their reliance on static co-occurrence assumptions and the absence of fine-grained, realistically challenging datasets. To address this gap, this work introduces AVTrack—the first human-centric Audio-Visual Instance Segmentation (AVIS) benchmark tailored for dynamic and complex environments, incorporating realistic challenges such as camera motion, occlusion, and speaker position changes. The dataset features multimodal fusion, instance-level annotations, and spatiotemporal alignment mechanisms. An accompanying evaluation protocol and baseline models reveal a significant performance drop compared to existing approaches, thereby validating the dataset’s difficulty and establishing a foundation for advancing high-order cross-modal spatiotemporal modeling research.
This study addresses the challenge of automated AfroBeats dance motion analysis by proposing a marker-free, equipment-agnostic end-to-end video analytics framework. Methodologically, it integrates YOLOv8/v11 with the Segment Anything Model (SAM) to achieve high-precision dancer detection and pixel-level instance segmentation—overcoming the limitations of bounding-box-based approaches. Motion quantification is performed via trajectory tracking and inter-frame displacement analysis, enabling step counting, spatial coverage measurement, and rhythm consistency assessment. Evaluated on a 49-second real-world AfroBeats video, the system achieves 94% detection precision, 89% recall, and an 83% IoU for SAM-based segmentation. Quantitative analysis reveals statistically significant disparities between lead and supporting dancers in step count (+23%), motion intensity (+37%), and spatial occupancy (+42%). To our knowledge, this is the first work to apply SAM for fine-grained African dance motion analysis, establishing a scalable, high-fidelity visual computing paradigm for markerless dance evaluation.