🤖 AI Summary
To address the high annotation cost of fully supervised video action detection, this paper proposes an action-agnostic frame-level weak supervision paradigm (AAPL), requiring only sparse, unsupervised keyframe annotations—without exhaustive video scanning or instance-level localization—thereby drastically reducing labeling effort. Methodologically, we design an end-to-end temporal detection model coupled with a weakly supervised learning strategy to accurately localize action segments from sparse, non-instance-aligned frame labels. A contrastive pseudo-label distillation mechanism is further introduced to enhance the temporal convolutional network’s capacity for modeling action boundaries. Evaluated on five standard benchmarks—including THUMOS’14 and ActivityNet 1.3—our approach achieves state-of-the-art or competitive performance using significantly fewer annotations than video-level or point-level supervision, establishing, for the first time, a new paradigm that simultaneously delivers high detection accuracy and low annotation cost.
📝 Abstract
We propose action-agnostic point-level (AAPL) supervision for temporal action detection to achieve accurate action instance detection with a lightly annotated dataset. In the proposed scheme, a small portion of video frames is sampled in an unsupervised manner and presented to human annotators, who then label the frames with action categories. Unlike point-level supervision, which requires annotators to search for every action instance in an untrimmed video, frames to annotate are selected without human intervention in AAPL supervision. We also propose a detection model and learning method to effectively utilize the AAPL labels. Extensive experiments on the variety of datasets (THUMOS '14, FineAction, GTEA, BEOID, and ActivityNet 1.3) demonstrate that the proposed approach is competitive with or outperforms prior methods for video-level and point-level supervision in terms of the trade-off between the annotation cost and detection performance.