Action-Agnostic Point-Level Supervision for Temporal Action Detection

📅 2024-12-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the high annotation cost of fully supervised video action detection, this paper proposes an action-agnostic frame-level weak supervision paradigm (AAPL), requiring only sparse, unsupervised keyframe annotations—without exhaustive video scanning or instance-level localization—thereby drastically reducing labeling effort. Methodologically, we design an end-to-end temporal detection model coupled with a weakly supervised learning strategy to accurately localize action segments from sparse, non-instance-aligned frame labels. A contrastive pseudo-label distillation mechanism is further introduced to enhance the temporal convolutional network’s capacity for modeling action boundaries. Evaluated on five standard benchmarks—including THUMOS’14 and ActivityNet 1.3—our approach achieves state-of-the-art or competitive performance using significantly fewer annotations than video-level or point-level supervision, establishing, for the first time, a new paradigm that simultaneously delivers high detection accuracy and low annotation cost.

Technology Category

Computer Vision: Video Understanding & Activity AnalysisMachine Learning: Active LearningHumans and AI: Human-Aware Planning and Behavior Prediction

Application Category

Economics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applications
📝 Abstract
We propose action-agnostic point-level (AAPL) supervision for temporal action detection to achieve accurate action instance detection with a lightly annotated dataset. In the proposed scheme, a small portion of video frames is sampled in an unsupervised manner and presented to human annotators, who then label the frames with action categories. Unlike point-level supervision, which requires annotators to search for every action instance in an untrimmed video, frames to annotate are selected without human intervention in AAPL supervision. We also propose a detection model and learning method to effectively utilize the AAPL labels. Extensive experiments on the variety of datasets (THUMOS '14, FineAction, GTEA, BEOID, and ActivityNet 1.3) demonstrate that the proposed approach is competitive with or outperforms prior methods for video-level and point-level supervision in terms of the trade-off between the annotation cost and detection performance.
Problem

Research questions and friction points this paper is trying to address.

Video Action Recognition
Limited Annotation Data
Effective Action Identification
Innovation

Methods, ideas, or system contributions that make the work stand out.

AAPL Supervision
Key Frame Selection
Efficient Learning Strategy
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
NEC Corporation | Tohoku University | RIKEN Center for Advanced Intelligence Project | The University of Tokyo
S
Shuhei M. Yoshida
Visual Intelligence Research Laboratories, NEC Corporation, Kanagawa 211-8666, Japan
T
Takashi Shibata
Visual Intelligence Research Laboratories, NEC Corporation, Kanagawa 211-8666, Japan
M
Makoto Terao
Visual Intelligence Research Laboratories, NEC Corporation, Kanagawa 211-8666, Japan
T
Takayuki Okatani
Graduate School of Information Sciences, Tohoku University, Miyagi 980-8579, Japan; RIKEN Center for Advanced Intelligence Project, Tokyo 103-0027, Japan
Masashi Sugiyama
Masashi Sugiyama
Director, RIKEN Center for Advanced Intelligence Project / Professor, The University of Tokyo
Machine LearningData MiningArtificial Intelligence