Robots Acquire Manipulation Skills in Seconds from a Single Human Video

📅 2026-07-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of enabling robots to rapidly acquire new manipulation skills from minimal human demonstrations while avoiding catastrophic forgetting. The authors propose HOST, a framework that leverages task-progress manifold-guided cross-modal alignment, future observation prediction, and self-supervised action generation within a cascaded skill transfer architecture. This approach allows robots to learn novel skills from a single human video in seconds during inference. Experimental results demonstrate that HOST achieves an average learning time of 29 seconds with a success rate of 62%, outperforming zero-shot baselines by 45% and surpassing models fine-tuned with 50 robot demonstrations. The method improves data efficiency by a factor of 507 and represents the first demonstration of highly efficient, catastrophic-forgetting-free robotic skill transfer from a single video.
📝 Abstract
The ability to acquire skills rapidly and effortlessly while retaining those already mastered is essential for robots. However, current methods still rely on a cumbersome training-time loop that is costly and slow, while eroding skills already mastered. In this paper, we introduce HOST (Human-to-robot One-Shot Skill AcquisiTion), a framework that enables a robot to acquire skills in seconds from a single human video while retaining previously mastered skills. HOST resolves skill acquisition through a cascade of self-grounded prediction. It first estimates the robot's progress within the demonstrated task, then translates the upcoming progression into the robot's own future observations, and finally derives actions from these predicted observations. This cascade is trained on targets coupled to the video demonstration, obtained by mapping the robot trajectory and the video demonstration onto a shared task progress manifold, then redefining each target to align with the future progression of the video. HOST thereby enables the robot to actively follow the demonstrated procedure and adapt it to the robot's embodiment. HOST acquires novel skills at inference time from a single human video in an average of 29 seconds and achieves a 62% average success rate. It exceeds the zero-shot baseline by 45% while retaining previously mastered skills. HOST even exceeds the baseline fine-tuned on 50 robot demonstrations per task while requiring 50 times fewer demonstrations and acquiring each skill 507 times faster. Additional information about HOST is available on the project website.
Problem

Research questions and friction points this paper is trying to address.

robotic manipulation
one-shot learning
skill acquisition
human-to-robot transfer
lifelong learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

one-shot learning
skill acquisition
human-to-robot transfer
self-grounded prediction
task progress manifold
Guangyan Chen
Guangyan Chen
Beijing Institute of Technology
M
Meiling Wang
Beijing Institute of Technology
Te Cui
Te Cui
Beijing Institute of Technology
Embodied AI
Z
Zichen Zhou
Beijing Institute of Technology
Q
Qi Shao
Beijing Institute of Technology
S
Shalfun Li
X SQUARE ROBOT
Hang Su
Hang Su
Associated Professor, Tsinghua University
Adversarial LearningAdversarial Attacks and DefenseInterpretable Learning
R
Roy Gan
X SQUARE ROBOT
H
Hao Wang
X SQUARE ROBOT
M
Mengyin Fu
Beijing Institute of Technology
Y
Yi Yang
Beijing Institute of Technology
Y
Yufeng Yue
Beijing Institute of Technology