PACT: End-to-End Learning of Human Pose, Contacts, and Forces from Video

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitations of existing approaches that decouple visual pose reconstruction from contact force estimation, which suffer from restricted joint reasoning and error accumulation. We propose PACT, an end-to-end model built upon a human reconstruction foundation model that incorporates learnable contact force tokens, a temporal Transformer, and physics-constrained supervision to jointly estimate human pose, contact states, and contact forces from monocular video. Key contributions include enabling end-to-end joint learning of these three tasks, developing a data annotation pipeline integrating physics-based optimization, and releasing ForceWall, a real-world climbing benchmark. Experimental results demonstrate that PACT achieves state-of-the-art performance on contact and force estimation tasks, significantly outperforming staged methods while exhibiting strong generalization capabilities to out-of-distribution interactions.
πŸ“ Abstract
Human motion, environmental contacts, and interaction forces are governed by common physical laws, yet existing approaches typically separate visual pose reconstruction from contact and force estimation. This separation limits joint reasoning and can propagate errors between stages. We introduce PACT, an end-to-end model that jointly learns to estimate human pose, contacts and contact forces from monocular video. Our approach augments a human reconstruction foundation model with learnable contact-force tokens and a temporal transformer that integrates visual features with world-space motion. Joint prediction heads refine human poses and estimate contacts and forces, while physics-based supervision encourages consistency between the reconstructed motion and interaction forces. To address the scarcity of force annotations, we develop a data annotation pipeline that combines contact labeling with physics-based motion and force optimization, producing training supervision from synthetic and real-world videos. We also introduce a real-world climbing benchmark ForceWall with climbing videos and corresponding ground-truth contact forces obtained from the force sensors. Experiments demonstrate state-of-the-art contact and force estimation, outperforming staged reconstruction approaches and generalizing to interactions beyond the training distribution. These results support end-to-end joint learning as an effective approach to recovering human motion and physical interactions from video.
Problem

Research questions and friction points this paper is trying to address.

human pose estimation
contact force estimation
monocular video
end-to-end learning
physical interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

End-to-End Learning
Contact Force Estimation
Physics-based Supervision
Monocular Video
Human Pose Reconstruction
πŸ”Ž Similar Papers
No similar papers found.