Paving the Way Towards Kinematic Assessment Using Monocular Video: A Preclinical Benchmark of State-of-the-Art Deep-Learning-Based 3D Human Pose Estimators Against Inertial Sensors in Daily Living Activities

📅 2025-10-02
📈 Citations: 0
Influential: 0
📄 PDF

career value

216K/year
🤖 AI Summary
This study addresses the challenge of performing accurate, lab-free 3D human kinematic assessment during Activities of Daily Living (ADL) for telemedicine, sports science, and rehabilitation. We systematically benchmark monocular video-based and IMU-based approaches using state-of-the-art models—including MotionAGFormer, MotionBERT, MMPose (2D-to-3D), and NVIDIA BodyTrack—and unify evaluation via OpenSim inverse dynamics and Human3.6M joint-angle metrics. Results demonstrate that MotionAGFormer achieves the highest accuracy (RMSE = 9.27° ± 4.80°, MAE = 7.86° ± 4.18°, correlation r = 0.86 ± 0.15, R² = 0.67 ± 0.28), confirming the clinical feasibility of monocular video in real-world settings. We introduce a new benchmark for in-the-wild human motion capture, explicitly characterizing the trade-offs among accuracy, cost, and deployment practicality between video and IMU modalities. This work provides empirical validation and methodological guidance for scalable, low-cost remote motion monitoring.

Technology Category

Application Category

📝 Abstract
Advances in machine learning and wearable sensors offer new opportunities for capturing and analyzing human movement outside specialized laboratories. Accurate assessment of human movement under real-world conditions is essential for telemedicine, sports science, and rehabilitation. This preclinical benchmark compares monocular video-based 3D human pose estimation models with inertial measurement units (IMUs), leveraging the VIDIMU dataset containing a total of 13 clinically relevant daily activities which were captured using both commodity video cameras and five IMUs. During this initial study only healthy subjects were recorded, so results cannot be generalized to pathological cohorts. Joint angles derived from state-of-the-art deep learning frameworks (MotionAGFormer, MotionBERT, MMPose 2D-to-3D pose lifting, and NVIDIA BodyTrack) were evaluated against joint angles computed from IMU data using OpenSim inverse kinematics following the Human3.6M dataset format with 17 keypoints. Among them, MotionAGFormer demonstrated superior performance, achieving the lowest overall RMSE ($9.27deg pm 4.80deg$) and MAE ($7.86deg pm 4.18deg$), as well as the highest Pearson correlation ($0.86 pm 0.15$) and the highest coefficient of determination $R^{2}$ ($0.67 pm 0.28$). The results reveal that both technologies are viable for out-of-the-lab kinematic assessment. However, they also highlight key trade-offs between video- and sensor-based approaches including costs, accessibility, and precision. This study clarifies where off-the-shelf video models already provide clinically promising kinematics in healthy adults and where they lag behind IMU-based estimates while establishing valuable guidelines for researchers and clinicians seeking to develop robust, cost-effective, and user-friendly solutions for telehealth and remote patient monitoring.
Problem

Research questions and friction points this paper is trying to address.

Benchmarking monocular video 3D pose estimation against IMUs for kinematic assessment
Evaluating joint angle accuracy in daily living activities using deep learning
Comparing video and sensor trade-offs for telehealth movement analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Monocular video-based 3D human pose estimation
Benchmarking deep learning models against IMU sensors
Evaluating joint angles using VIDIMU dataset activities
M
Mario Medrano-Paredes
Department of Signal Theory, Communications and Telematics Engineering. University of Valladolid, 47011 Valladolid, Spain
C
Carmen Fernández-González
Department of Signal Theory, Communications and Telematics Engineering. University of Valladolid, 47011 Valladolid, Spain
F
Francisco-Javier Díaz-Pernas
Department of Signal Theory, Communications and Telematics Engineering. University of Valladolid, 47011 Valladolid, Spain
H
Hichem Saoudi
Department of Signal Theory, Communications and Telematics Engineering. University of Valladolid, 47011 Valladolid, Spain
J
Javier González-Alonso
Department of Signal Theory, Communications and Telematics Engineering. University of Valladolid, 47011 Valladolid, Spain
M
Mario Martínez-Zarzuela
Department of Signal Theory, Communications and Telematics Engineering. University of Valladolid, 47011 Valladolid, Spain