🤖 AI Summary
This study addresses the challenges of robotic data scarcity and insufficient precision in complex tool manipulation by proposing P2P-T, a framework that enables robots to learn tool-use skills from human videos. The method employs a two-stage strategy: first pre-training an object-centric world model to extract pose priors, and subsequently integrating them into an efficient low-level control policy. A core innovation lies in completely eliminating the reliance on human-robot alignment data while significantly reducing training overhead through foundation model-based data processing and automated pipelines. Experimental results demonstrate that P2P-T outperforms existing state-of-the-art methods by 73% in complex real-world tool manipulation tasks.
📝 Abstract
Scaling up robotic manipulation is primarily bottlenecked by the scarcity of real-world robot data. While recent approaches leverage human video demonstrations to mitigate this shortage, they remain computationally expensive and still rely on paired human-robot data for domain alignment. Although current state-of-the-arts excel at long-horizon tasks, they struggle with the delicate and precise control required for complex tool manipulation. To overcome these limitations, we introduce P2P-T, from Pixel to Poses for Tool Manipulation, a data-efficient, object-centric framework that learns tool use directly from human demonstrations. P2P-T bridges the cognitive and physical execution gap through a two-stage approach. First, pretraining an object-centric world model to extract stable pose priors; second, integrating these priors into an efficient, pose-aware low-level policy. By utilizing a robust automated data processing pipeline powered by modern foundation models, P2P-T completely bypasses the need for human-robot aligned data. This reduces overall training overhead drastically. With minimal per-task fine-tuning, our framework achieves a 73% improvement over the previous state of the art in execution performance on complex, real-world tool manipulation tasks that currently remain out of reach for standard large-scale pretrained models.