🤖 AI Summary
This study addresses the challenge of recognizing heading events in football broadcast videos, which are inherently brief and subtle. To this end, we propose BMASH, a multimodal action recognition framework that employs Video Swin Transformer for action feature extraction. The framework innovatively incorporates frame-level football detection, leveraging ball presence and dynamic trajectories as critical contextual cues. A multimodal feature fusion algorithm is further introduced to enhance the model’s discriminative capability for heading actions. Experimental results demonstrate that the proposed method significantly improves Average Precision (AP) and ROC-AUC metrics in clip-level classification while maintaining high F1 performance in continuous full-video detection tasks. These findings confirm the effectiveness of BMASH in tackling fine-grained sports action recognition challenges.
📝 Abstract
Recent advances in computer vision have made broadcast sports videos increasingly useful for event analysis, performance assessment, and player-safety applications. In soccer, however, header spotting remains a challenging problem due to the subtle and short-lived nature of header events. This paper focuses on soccer header spotting: identifying moments in broadcast videos where the ball contacts a player's head. We first adapt and evaluate Video Swin as a strong action-recognition baseline for this task, and then introduce BMASH, a ball-motion-aware fusion framework that integrates detector-derived ball features. BMASH combines Video Swin action representations with ball-presence and motion features from frame-level soccer-ball detection, integrating player-action context with ball dynamics to distinguish headers from visually similar events. We evaluate BMASH using game-level splits with separate test matches and rotating validation folds, considering both centered-window classification and continuous full-video spotting. Results show that Video Swin provides a strong baseline for header spotting, while BMASH improves clip-level AP and ROC-AUC over the corresponding Video Swin baseline. In continuous full-video spotting, BMASH achieves a comparable event-level F1-performance with a different precision--recall trade-off.