🤖 AI Summary
This study addresses the unclear impact of camera viewing angles on pose estimation accuracy in AI-based rehabilitation movement quality assessment. To this end, it introduces REHAB26-ViewAngles, the first multi-view dataset for this domain, and proposes a novel separability metric to quantify algorithmic discriminative capability. By integrating 2D and 3D pose estimation, triangulation, weighted fusion, and Transformer models, the work systematically evaluates single- and multi-view RGB strategies for distinguishing correct from incorrect movements. Experimental results demonstrate that the optimal 2D viewpoint improves separability by 16.9% over the frontal view, while dual-view configurations further increase accuracy by 13.1%. These findings establish an effective framework for multi-view rehabilitation assessment.
📝 Abstract
Automated quality assessment of rehabilitation exercises relies heavily on accurate human pose estimation from video data. Although numerous RGB-based pose estimation methods have been proposed, the impact of camera placement on detecting clinically relevant movement errors remains insufficiently explored. To address this gap, we introduce REHAB26-ViewAngles, a dataset comprising correct and incorrect rehabilitation exercise executions captured from a wide range of camera angles. Furthermore, we propose a novel separability metric to quantify an algorithm's ability to distinguish between valid and faulty exercise repetitions. Using these tools, we analyze how various RGB-based pose-estimation strategies are suitable for exercise quality assessment under varying camera placements. In particular, we analyze single-camera 2D and 3D pose estimation and four multi-camera strategies: a combination of two orthogonal 2D views, 3D triangulation, weighted 3D fusion, and an AI-based pose-estimation transformer model specifically trained from two synchronized cameras. Our findings reveal that an optimally placed 2D camera can improve the separability by 16.9\,\% over the commonly used $0^\circ$ frontal view and frequently outperforms single-camera 3D estimation, while combining two views can further improve accuracy by up to 13.1\,\%. These results offer practical guidance for deploying rehabilitation monitoring in both home and clinical settings.