🤖 AI Summary
Existing 3D human pose estimation methods suffer from severe overfitting to single datasets, exhibiting poor generalization across viewpoints, environments, and camera configurations.
Method: We introduce the first standardized cross-dataset evaluation framework—covering four major benchmarks with dynamic extensibility—and establish a fair, unified-interface evaluation paradigm. The framework features standardized data loading and preprocessing pipelines, dual-metric assessment (MPJPE and PA-MPJPE), and modular design for seamless integration of diverse models.
Contribution/Results: Through systematic re-evaluation of 18 state-of-the-art methods (>100 new results), we quantitatively reveal that model performance degrades by 30–60% under cross-domain settings; moreover, preprocessing choices and data configuration critically govern generalization capability. This work shifts the field’s focus from single-dataset optimization toward robust modeling for real-world deployment.
📝 Abstract
Reliable three-dimensional human pose estimation is becoming increasingly important for real-world applications, yet much of prior work has focused solely on the performance within a single dataset. In practice, however, systems must adapt to diverse viewpoints, environments, and camera setups -- conditions that differ significantly from those encountered during training, which is often the case in real-world scenarios. To address these challenges, we present a standardized testing environment in which each method is evaluated on a variety of datasets, ensuring consistent and fair cross-dataset comparisons -- allowing for the analysis of methods on previously unseen data. Therefore, we propose PoseBench3D, a unified framework designed to systematically re-evaluate prior and future models across four of the most widely used datasets for human pose estimation -- with the framework able to support novel and future datasets as the field progresses. Through a unified interface, our framework provides datasets in a pre-configured yet easily modifiable format, ensuring compatibility with diverse model architectures. We re-evaluated the work of 18 methods, either trained or gathered from existing literature, and reported results using both Mean Per Joint Position Error (MPJPE) and Procrustes Aligned Mean Per Joint Position Error (PA-MPJPE) metrics, yielding more than 100 novel cross-dataset evaluation results. Additionally, we analyze performance differences resulting from various pre-processing techniques and dataset preparation parameters -- offering further insight into model generalization capabilities.