Precise Event Spotting in Sports Videos: Solving Long-Range Dependency and Class Imbalance

๐Ÿ“… 2025-02-28
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Precise Event Localization (PES) in long sports videos faces two key challenges: modeling long-range temporal dependencies and severe class imbalance among events. To address these, we propose an end-to-end framework featuring three core components: (1) an Adaptive Spatio-Temporal Refinement Module (ASTRM) to enhance local temporal discriminability; (2) a Long-Range Temporal Modeling Module explicitly designed to capture millisecond-level event boundaries; and (3) a novel Soft Instance Contrastive (SoftIC) loss function that dynamically balances feature compactness and inter-class separability. Evaluated on multiple sports event localization benchmarks, our method achieves state-of-the-art performance. Notably, it delivers substantial improvements in localization accuracy and robustness for rare-event detection and long-interval scenariosโ€”where prior methods suffer from degraded boundary estimation and poor generalization across imbalanced classes. The framework demonstrates strong generalizability while maintaining computational efficiency.

Technology Category

Computer Vision: Motion & TrackingMachine Learning: Time-Series/Data StreamsKnowledge Representation and Reasoning: Geometric, Spatial, and Temporal Reasoning

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Algorithms and analysis for heterogeneous, signed, attributed, multi-relational, temporal, higher-order, and annotated Web-related graphsUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systems
๐Ÿ“ Abstract
Precise Event Spotting (PES) aims to identify events and their class from long, untrimmed videos, particularly in sports. The main objective of PES is to detect the event at the exact moment it occurs. Existing methods mainly rely on features from a large pre-trained network, which may not be ideal for the task. Furthermore, these methods overlook the issue of imbalanced event class distribution present in the data, negatively impacting performance in challenging scenarios. This paper demonstrates that an appropriately designed network, trained end-to-end, can outperform state-of-the-art (SOTA) methods. Particularly, we propose a network with a convolutional spatial-temporal feature extractor enhanced with our proposed Adaptive Spatio-Temporal Refinement Module (ASTRM) and a long-range temporal module. The ASTRM enhances the features with spatio-temporal information. Meanwhile, the long-range temporal module helps extract global context from the data by modeling long-range dependencies. To address the class imbalance issue, we introduce the Soft Instance Contrastive (SoftIC) loss that promotes feature compactness and class separation. Extensive experiments show that the proposed method is efficient and outperforms the SOTA methods, specifically in more challenging settings.
Problem

Research questions and friction points this paper is trying to address.

Detect precise events in long sports videos
Address class imbalance in event distribution
Model long-range dependencies for better accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Spatio-Temporal Refinement Module (ASTRM)
Long-range temporal module for global context
Soft Instance Contrastive (SoftIC) loss for class imbalance
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.