DiMoP: Diffusion-Driven Motion Representation Learning With Frame-Level Pseudo-Classification for Skeleton-Based Action Recognition

📅 2026-09-28
🏛️ IEEE Transactions on Biometrics Behavior and Identity Science
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge in existing skeleton-based action recognition methods of balancing the capture of motion patterns across varying intensities. To this end, this work proposes DiMoP, a framework that achieves joint generative and discriminative modeling through mask-diffusion-driven motion representation learning. Its core innovation lies in leveraging a controllable noise denoising process to uniformly learn dynamic features ranging from subtle to vigorous movements, while introducing an unlabeled frame-level pseudo-classifier to enhance temporal consistency. The proposed method attains state-of-the-art performance on both the NTU RGB+D and PKU-MMD datasets, achieving a 1.1% accuracy improvement under the cross-subject protocol.
📝 Abstract
Robust skeleton-based action recognition requires representations that capture a wide spectrum of motions, from subtle to moderate and strong ones. Existing methods often focus on strong motions. This paper introduces DiMoP, a masking- and diffusion-driven motion representation learning method with frame-level pseudo-classification to explicitly learn the distribution of joint motions rather than regressing deterministic coordinates, as existing methods often do. By diffusing masked joints with progressive noise and denoising them conditioned on visible joints, DiMoP learns through controllable noising and denoising processes, enabling uniform learning of weak, moderate, and strong dynamics. To enable the masking-based generative diffusion learning with a discriminative capability, a pseudo-frame classifier is proposed that enforces the learning towards sequence-consistent and temporally coherent pseudo-labels without manual annotations. Together, these strategies provide a principled mechanism for joint generative and discriminative motion modeling. DiMoP achieves state-of-the-art performance across NTU RGB+D 60/120, and PKUMMD, including a 1.1 percentage point gain over prior works on NTU RGB+D 120 with the cross-subject protocol.
Problem

Research questions and friction points this paper is trying to address.

Skeleton-based action recognition
Motion representation learning
Joint motion distribution
Weak and strong dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion-driven representation learning
Masked generative modeling
Frame-level pseudo-classification
Skeleton-based action recognition
Motion distribution modeling
💼 Related Jobs
No related jobs found.
S
Shanaka Ramesh Gunasekara
Advanced Multimedia Research Lab, University of Wollongong, Australia
Wanqing Li
Wanqing Li
Professor, University of Wollongong
Multimedia UnderstandingComputer VisionMachine Learning
N
Nikalal Kaldera
Advanced Multimedia Research Lab, University of Wollongong, Australia
P
Philip Ogunbona
Advanced Multimedia Research Lab, University of Wollongong, Australia
Jack Yang
Jack Yang
Senior Lecturer, University of New South Wales
Computational Material Science