PointWAM: 3D World Action Modeling for Dexterous Robotic Manipulation

πŸ“… 2026-10-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation of existing methods in capturing the 3D spatial structures and contact geometries essential for dexterous manipulation. We propose an explicitly decoupled 3D point trajectory representation that separates scene and hand dynamics, jointly predicting point cloud trajectories within a spatiotemporal coordinate system to enable self-supervised pretraining from human videos and subsequent robot action retargeting. This approach effectively leverages large-scale human video data without requiring object-specific or keypoint annotations, thereby establishing a unified 3D world action model. On the DexJoCo benchmark, our method achieves a 56.9% improvement in success rate, surpassing state-of-the-art approaches by 11.7 points. Furthermore, it significantly outperforms mainstream Vision-Language-Action (VLA) models in real-world robotic tasks.
πŸ“ Abstract
World action models jointly learn to forecast world dynamics and predict robot actions, such that the learned internal world dynamics guide accurate actions. Existing approaches typically represent the world as RGB frames or latent counterparts while predicting actions as end-effector poses or joint angles, but they often struggle to capture the 3D spatial structure and contact geometry central to dexterous manipulation. We introduce Point World Action Model (PointWAM), a 3D world action model that decomposes the world into a scene (i.e., environment) and hands (i.e., actor), and jointly forecasts both as 3D point trajectories within a shared space-time coordinate frame. This explicit, disentangled representation enables effective pre-training on large-scale human demonstration videos without requiring any task-specific object or keypoint selection. Given a colored point cloud and a language instruction, PointWAM predicts how the scene and hands co-evolve in 3D space over time, then retargets the forecast hand motion to robot actions. Pre-training on human videos improves average DexJoCo success by 56.9 percentage points, and scene-trajectory supervision adds 10.9 points over forecasting the hands alone. With both, PointWAM surpasses the prior state of the art on ten DexJoCo tasks by 11.7 points and outperforms strong VLAs on a real robot.
Problem

Research questions and friction points this paper is trying to address.

dexterous manipulation
world action model
3D spatial structure
contact geometry
Innovation

Methods, ideas, or system contributions that make the work stand out.

3D World Action Model
Dexterous Manipulation
Point Cloud Trajectories
Human Video Pre-training
Motion Retargeting