visual/body-type classification in video

Designs and implements models and processing pipelines that analyze video—both per-frame and across time—to detect and track people and assign visual body-type categories (e.g., shape or build) from appearance, pose, and silhouette cues. This work covers feature extraction, temporal aggregation over tracks, annotation and labeling schemes, training and evaluating classifiers on tracked segments, and engineering for robustness to occlusion, viewpoint and motion variability as well as dataset biases.

visualbody-typeclassificationin

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.09
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

To address the high cost and low efficiency of manual annotation in video object detection and segmentation, this paper proposes a lightweight, end-to-end automated annotation framework. Methodologically, it introduces the first deep integration of YOLOv8-based video tracking and SAM-based interactive segmentation, augmented by an adaptive frame-sampling strategy and a Gradio-powered visual interface, resulting in a modular, open-source, and extensible prototype system. Key contributions include: (i) a lightweight co-design of tracking and segmentation models; (ii) efficient, interactive annotation support for multi-object, long-duration videos; and (iii) rapid generation of annotated datasets enabling a closed-loop annotation–training pipeline. Experiments on multiple public video benchmarks demonstrate over 10× higher annotation throughput compared to manual labeling, strong inter-annotator consistency, and substantial performance gains for downstream detection and segmentation models trained on the generated data.

Automate video annotation for computer vision modelsDevelop efficient video tracking and segmentation toolReduce time and resources for labeled data generation

To address storage redundancy and inefficient retrieval in video surveillance, this paper proposes an activity-driven intelligent dynamic scene analysis system. Methodologically, it introduces a novel hybrid motion segmentation strategy integrating adaptive background modeling, Lucas-Kanade optical flow, and a deep temporal model (LSTM); combines multi-scale context-aware object detection (based on YOLO/SSD) with illumination-invariant feature optimization; and enhances tracking robustness via Kalman filtering and Siamese network-based re-identification. Evaluated on real-world CCTV footage, the system achieves significant improvements in critical event detection—e.g., person appearance and anomalous behavior—with average precision and recall gains of 12.3%. It operates at real-time speed (≥25 FPS), reduces video storage overhead by 63.7%, and enables efficient content-based retrieval and long-term archival.

Develop robust video surveillance systemImprove object detection and tracking accuracyOptimize storage and enhance digital searches

FeatureSORT: Essential Features for Effective Tracking

Jul 05, 2024
HH
Hamidreza Hashempoor
🏛️ Pintel Co. Ltd.

To address frequent identity switches and trajectory discontinuities caused by occlusion in multi-object tracking, this paper proposes a lightweight online tracker. Methodologically: (1) it introduces the first multi-granularity appearance feature disentanglement framework, jointly modeling IoU, motion direction, clothing color/style, and ReID embeddings via an adaptive weighted composite distance metric; (2) it integrates a high-accuracy detector with a customized post-processing module to enhance detection robustness and trajectory smoothness. Experiments on mainstream benchmarks demonstrate significant improvements in MOTA (+2.1%) and IDF1 (+3.4%), a 32% reduction in ID switches, and markedly enhanced trajectory continuity under occlusion—while maintaining real-time inference speed (≥30 FPS), thus satisfying practical online deployment requirements.

Enhancing multi-object tracking with enriched appearance featuresImproving association accuracy through complementary feature embeddingsReducing identity switches during occlusions in object tracking

Latest Papers

What's happening recently
View more

Is This Tracker On? A Benchmark Protocol for Dynamic Tracking

Oct 22, 2025
ID
Ilona Demler
🏛️ California Institute of Technology

Existing point tracking methods face significant challenges in real-world scenarios—including high motion complexity, frequent occlusions, and large object diversity—yet lack a systematic benchmark for evaluating robustness and failure modes. To address this, we introduce ITTO, the first high-challenge dynamic point tracking benchmark, comprising first-person real-world videos and multi-source data. We propose a multi-stage manual annotation protocol to precisely characterize motion patterns, occlusion events, and appearance variations. ITTO introduces a novel performance analysis protocol stratified by motion complexity and, for the first time, systematically exposes critical failure points of mainstream trackers—particularly in post-occlusion re-identification. Experiments reveal substantial performance degradation of state-of-the-art methods on ITTO, highlighting deficiencies in long-term occlusion handling and complex dynamic modeling. These findings provide concrete diagnostic insights and quantitative evaluation standards to guide algorithmic improvement.

Addressing motion complexity and occlusion in real scenesBenchmarking tracker performance after object occlusionEvaluating point tracking methods' capabilities and limitations

This work proposes a novel end-to-end multi-object tracking approach that circumvents the limitations of traditional methods, which rely on separate detection and association modules and struggle to maintain identity consistency under occlusions due to complex external identity management. For the first time, the authors directly employ a large-scale text-to-video diffusion model (22B-parameter LTX-2.3) as a tracker, applying lightweight in-context LoRA fine-tuning to generate ID-map videos with temporally consistent identity colors. Identities are encoded implicitly at the pixel level, eliminating the need for explicit detection or association components. A windowed chain-generation strategy conditions each subsequent video window on the last frame of the previous one to preserve identity continuity. The method achieves 40.3 HOTA and 44.1 AssA on DanceTrack, substantially outperforming existing approaches, and successfully recovers identities in 42% of 383 occlusion events—dramatically surpassing appearance-embedding baselines, which exhibit near-zero recovery rates.

Diffusion modelsIdentity preservationMulti-object tracking

Existing audio-visual speaker tracking methods suffer from limited performance in complex dynamic scenes, primarily due to their reliance on static co-occurrence assumptions and the absence of fine-grained, realistically challenging datasets. To address this gap, this work introduces AVTrack—the first human-centric Audio-Visual Instance Segmentation (AVIS) benchmark tailored for dynamic and complex environments, incorporating realistic challenges such as camera motion, occlusion, and speaker position changes. The dataset features multimodal fusion, instance-level annotations, and spatiotemporal alignment mechanisms. An accompanying evaluation protocol and baseline models reveal a significant performance drop compared to existing approaches, thereby validating the dataset’s difficulty and establishing a foundation for advancing high-order cross-modal spatiotemporal modeling research.

audio-visual trackingcomplex scenescross-modal reasoning

AfroBeats Dance Movement Analysis Using Computer Vision: A Proof-of-Concept Framework Combining YOLO and Segment Anything Model

Dec 03, 2025
KO
Kwaku Opoku-Ware
🏛️ University of Idaho | Kwame Nkrumah University of Science and Technology

This study addresses the challenge of automated AfroBeats dance motion analysis by proposing a marker-free, equipment-agnostic end-to-end video analytics framework. Methodologically, it integrates YOLOv8/v11 with the Segment Anything Model (SAM) to achieve high-precision dancer detection and pixel-level instance segmentation—overcoming the limitations of bounding-box-based approaches. Motion quantification is performed via trajectory tracking and inter-frame displacement analysis, enabling step counting, spatial coverage measurement, and rhythm consistency assessment. Evaluated on a 49-second real-world AfroBeats video, the system achieves 94% detection precision, 89% recall, and an 83% IoU for SAM-based segmentation. Quantitative analysis reveals statistically significant disparities between lead and supporting dancers in step count (+23%), motion intensity (+37%), and spatial occupancy (+42%). To our knowledge, this is the first work to apply SAM for fine-grained African dance motion analysis, establishing a scalable, high-fidelity visual computing paradigm for markerless dance evaluation.

Automated analysis of AfroBeats dance movements using computer visionEvaluating technical feasibility of YOLO and SAM integration for motion analysisTracking and quantifying dancer steps, spatial coverage, and rhythm consistency

Hot Scholars

AG

Anhong Guo

Assistant Professor, University of Michigan
human-computer interactionaccessibilityhuman-AI interactionaugmented reality
LL

Luca Luceri

Research Assistant Professor @University of Southern California - Information Sciences Institute
Computational Social ScienceNetwork ScienceMachine LearningSocial Media Manipulation