HuC-VideoMAE: Human-Centric Video Masked Autoencoding from synthetic data

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the ethical concerns of action recognition models relying on unauthorized web videos and the suboptimal pre-training performance of synthetic data. To this end, it proposes a human-centric masked autoencoding strategy based on synthetic motion data. By leveraging keypoints and bounding boxes to design a novel masking mechanism, the method guides video Transformers to focus on human structural dynamics during self-supervised learning, enabling efficient pre-training without real images. Experimental results demonstrate that the proposed model significantly outperforms standard approaches on benchmarks such as NTU RGB+D, narrowing the performance gap between synthetic data and Kinetics real-data pre-training by 49%. All models and code are publicly released.
📝 Abstract
Modern action recognition models rely on video transformers pretrained on massive collections of web-crawled videos, such as Kinetics-700. However, the use of such data raises ethical concerns, as subjects' consent is typically not obtained. Recent high-quality synthetic video datasets generated from motion-capture data, such as BEDLAM2.0, offer a promising ethical alternative. In this work, we investigate self-supervised pretraining of video transformers on synthetic human-motion datasets. We first show that directly applying the standard VideoMAE masking strategy leads to substantially worse performance than pretraining on Kinetics. To address this limitation, we propose a human-centric masking scheme that leverages body keypoints and person bounding box regions. Our approach encourages the model to focus on the structure and dynamics of human motion during pretraining. Experiments on NTU RGB+D and Toyota-Smarthome demonstrate that our method significantly outperforms standard VideoMAE pretraining on synthetic data, closing 49% of the gap to Kinetics pretraining on NTU RGB+D cross-view-subject without using a single real frame during pretraining. To promote the use of ethical action recognition models, we will publicly release our pretrained models.
Problem

Research questions and friction points this paper is trying to address.

Action Recognition
Synthetic Data
Self-supervised Pretraining
Video Transformers
Ethical AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

Human-Centric Masking
Video Masked Autoencoding
Synthetic Data
Self-Supervised Pretraining
Action Recognition
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ricardo Pizarro
Universidad de Alcalá, Alcalá de Henares, Spain
R
Roberto Valle
Universidad Politécnica de Madrid, Madrid, Spain
J
José M. Buenaposada
Universidad Rey Juan Carlos, Móstoles, Spain
L
Luis M. Bergasa
Universidad de Alcalá, Alcalá de Henares, Spain
Luis Baumela
Luis Baumela
Departamento de Inteligencia Artificial, Universidad Politécnica de Madrid
computer vision