🤖 AI Summary
This study addresses the significant performance degradation of existing pose estimation algorithms in challenging scenarios such as trampoline gymnastics, which involve extreme body postures and unconventional camera viewpoints. To tackle this issue, the authors introduce STP, the first synthetic pose dataset specifically designed for trampoline gymnastics, created by fitting a parametric human body model to noisy motion capture data and rendering photorealistic multi-view images. Building upon this dataset, they fine-tune the ViTPose model and integrate multi-view 3D triangulation to achieve high-accuracy 2D and 3D pose estimation. Experimental results demonstrate that the proposed approach achieves state-of-the-art 2D pose accuracy on real-world trampoline sequences and reduces the 3D mean per-joint position error (MPJPE) by 12.5 mm, representing a 19.6% improvement over the original ViTPose model.
📝 Abstract
Trampoline gymnastics involves extreme human poses and uncommon viewpoints, on which state-of-the art pose estimation models tend to under-perform. We demonstrate that this problem can be addressed by fine-tuning a pose estimation model on a dataset of synthetic trampoline poses (STP). STP is generated from motion capture recordings of trampoline routines. We develop a pipeline to fit noisy motion capture data to a parametric human model, then generate multiview realistic images. We use this data to fine-tune a ViTPose model, and test it on real multi-view trampoline images. The resulting model exhibits accuracy improvements in 2D which translates to improved 3D triangulation. In 2D, we obtain state-of-the-art results on such challenging data, bridging the performance gap between common and extreme poses. In 3D, we reduce the MPJPE by 12.5 mm with our best model, which represents an improvement of 19.6% compared to the pretrained ViTPose model.