🤖 AI Summary
This study addresses the challenge of automated AfroBeats dance motion analysis by proposing a marker-free, equipment-agnostic end-to-end video analytics framework. Methodologically, it integrates YOLOv8/v11 with the Segment Anything Model (SAM) to achieve high-precision dancer detection and pixel-level instance segmentation—overcoming the limitations of bounding-box-based approaches. Motion quantification is performed via trajectory tracking and inter-frame displacement analysis, enabling step counting, spatial coverage measurement, and rhythm consistency assessment. Evaluated on a 49-second real-world AfroBeats video, the system achieves 94% detection precision, 89% recall, and an 83% IoU for SAM-based segmentation. Quantitative analysis reveals statistically significant disparities between lead and supporting dancers in step count (+23%), motion intensity (+37%), and spatial occupancy (+42%). To our knowledge, this is the first work to apply SAM for fine-grained African dance motion analysis, establishing a scalable, high-fidelity visual computing paradigm for markerless dance evaluation.
📝 Abstract
This paper presents a preliminary investigation into automated dance movement analysis using contemporary computer vision techniques. We propose a proof-of-concept framework that integrates YOLOv8 and v11 for dancer detection with the Segment Anything Model (SAM) for precise segmentation, enabling the tracking and quantification of dancer movements in video recordings without specialized equipment or markers. Our approach identifies dancers within video frames, counts discrete dance steps, calculates spatial coverage patterns, and measures rhythm consistency across performance sequences. Testing this framework on a single 49-second recording of Ghanaian AfroBeats dance demonstrates technical feasibility, with the system achieving approximately 94% detection precision and 89% recall on manually inspected samples. The pixel-level segmentation provided by SAM, achieving approximately 83% intersection-over-union with visual inspection, enables motion quantification that captures body configuration changes beyond what bounding-box approaches can represent. Analysis of this preliminary case study indicates that the dancer classified as primary by our system executed 23% more steps with 37% higher motion intensity and utilized 42% more performance space compared to dancers classified as secondary. However, this work represents an early-stage investigation with substantial limitations including single-video validation, absence of systematic ground truth annotations, and lack of comparison with existing pose estimation methods. We present this framework to demonstrate technical feasibility, identify promising directions for quantitative dance metrics, and establish a foundation for future systematic validation studies.