🤖 AI Summary
This work addresses the limitations of existing SMPL-based 3D Gaussian Splatting methods for human avatars, demonstrating that the primary bottleneck lies in the representational capacity of the body model rather than architectural complexity. The authors propose replacing SMPL with the more expressive Momentum Human Rig (MHR), integrated with poses estimated by SAM-3D-Body, to achieve high-quality reconstructions through an exceptionally streamlined pipeline that avoids learning any deformation or pose-dependent corrections. Controlled ablation studies are conducted to isolate and validate, for the first time, the critical role of body model expressiveness in reconstruction fidelity. The method achieves state-of-the-art PSNR on both PeopleSnapshot and ZJU-MoCap datasets, while attaining leading or comparable performance in LPIPS and SSIM metrics.
📝 Abstract
Recent 3D Gaussian splatting methods built atop SMPL achieve remarkable visual fidelity while continually increasing the complexity of the overall training architecture. We demonstrate that much of this complexity is unnecessary: by replacing SMPL with the Momentum Human Rig (MHR), estimated via SAM-3D-Body, a minimal pipeline with no learned deformations or pose-dependent corrections achieves the highest reported PSNR and competitive or superior LPIPS and SSIM on PeopleSnapshot and ZJU-MoCap. To disentangle pose estimation quality from body model representational capacity, we perform two controlled ablations: translating SAM-3D-Body meshes to SMPL-X, and translating the original dataset's SMPL poses into MHR both retrained under identical conditions. These ablations confirm that body model expressiveness has been a primary bottleneck in avatar reconstruction, with both mesh representational capacity and pose estimation quality contributing meaningfully to the full pipeline's gains.