🤖 AI Summary
This study addresses the overly conservative execution of imitation learning policies in robotics and the lack of task-adaptive references in existing acceleration methods. We propose a dynamic acceleration framework that, for the first time, leverages natural human manipulation tempo as a velocity reference. By performing cross-modal action alignment to match manipulation phases, our approach integrates multi-source velocity estimation with demonstration retiming to transfer relative human velocities to robotic policies, which are subsequently trained via behavior cloning. This enables phase-aware dynamic policy acceleration. Experimental results demonstrate that the proposed method increases success rates by an average of 25 percentage points and reduces execution time by 36.5% across real-world tasks.
📝 Abstract
Robot manipulation policies trained through imitation learning inherit not only the demonstrated behavior but also the conservative execution tempo of robot demonstrations. Existing acceleration approaches can execute faster than the original demonstrations, but determine the appropriate acceleration primarily from robot-side information or a predefined set of tempo factors, leaving open how to obtain a task-appropriate reference for how fast each manipulation phase should progress. We introduce EgoSpeedUp, a framework that uses human manipulation as temporal supervision for robot imitation learning. Our key insight is that human demonstrations naturally reveal task-appropriate, phase-wise manipulation tempo. Given slow robot demonstrations and human demonstrations of the same task, EgoSpeedUp aligns corresponding manipulation phases, estimates their relative execution tempos from multiple human demonstrations, and transfers the resulting phase-wise tempo by retiming the robot demonstrations. The retimed demonstrations are then used for standard behavior cloning, allowing the robot to retain its executable manipulation behavior while learning to perform it at a human-informed tempo. Across two real-world manipulation tasks, EgoSpeedUp improves the task success rate by an average of 25 percentage points (pp) while reducing successful execution time by 36.5%. These results demonstrate that human manipulation tempo provides an effective temporal reference for learning faster and more reliable robot policies.