From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of robotic data scarcity and insufficient precision in complex tool manipulation by proposing P2P-T, a framework that enables robots to learn tool-use skills from human videos. The method employs a two-stage strategy: first pre-training an object-centric world model to extract pose priors, and subsequently integrating them into an efficient low-level control policy. A core innovation lies in completely eliminating the reliance on human-robot alignment data while significantly reducing training overhead through foundation model-based data processing and automated pipelines. Experimental results demonstrate that P2P-T outperforms existing state-of-the-art methods by 73% in complex real-world tool manipulation tasks.
📝 Abstract
Scaling up robotic manipulation is primarily bottlenecked by the scarcity of real-world robot data. While recent approaches leverage human video demonstrations to mitigate this shortage, they remain computationally expensive and still rely on paired human-robot data for domain alignment. Although current state-of-the-arts excel at long-horizon tasks, they struggle with the delicate and precise control required for complex tool manipulation. To overcome these limitations, we introduce P2P-T, from Pixel to Poses for Tool Manipulation, a data-efficient, object-centric framework that learns tool use directly from human demonstrations. P2P-T bridges the cognitive and physical execution gap through a two-stage approach. First, pretraining an object-centric world model to extract stable pose priors; second, integrating these priors into an efficient, pose-aware low-level policy. By utilizing a robust automated data processing pipeline powered by modern foundation models, P2P-T completely bypasses the need for human-robot aligned data. This reduces overall training overhead drastically. With minimal per-task fine-tuning, our framework achieves a 73% improvement over the previous state of the art in execution performance on complex, real-world tool manipulation tasks that currently remain out of reach for standard large-scale pretrained models.
Problem

Research questions and friction points this paper is trying to address.

robotic manipulation
tool manipulation
human demonstrations
data scarcity
domain alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Object-centric World Model
Tool Manipulation
Human Demonstrations
Pose-aware Policy
Foundation Models
💼 Related Jobs
No related jobs found.
Bangjun Wang
Bangjun Wang
University of Hong Kong
Artificial IntelligenceRobot LearningComputer Vision
L
Longyan Wu
Shanghai Innovation Institute
Y
Yukun Wei
The University of Hong Kong
S
Shenghe Shao
The University of Hong Kong
C
Chaoyi Huang
The University of Hong Kong
W
Wenze Cui
The University of Hong Kong
Z
Zetong Xu
The University of Hong Kong
Hanlin Wu
Hanlin Wu
Tsinghua University
Generative ModelsAI for Science
L
Long Chen
Xiaomi EV
Y
Yi Ma
The University of Hong Kong
Hongyang Li
Hongyang Li
Assistant Professor, University of Hong Kong
Computer VisionAutonomous DrivingRobotics