🤖 AI Summary
This study addresses the poor generalization in robot imitation learning caused by the coupling of actions and object effects during demonstrations. To this end, it proposes a "programmable effect-to-execution" paradigm that decouples the executor from the intended effect by compiling demonstrations into effect programs, such as 3D keypoint trajectories. Methodologically, the approach leverages the PEWAM model and flow matching techniques to generate independent action flows, enabling closed-loop replanning and allowing users to intuitively modify task logic by editing coordinates. Experimental results demonstrate that this method significantly outperforms baselines on the LIBERO and Meta-World benchmarks. Furthermore, in real-world settings, it achieves a 90% success rate from single-video demonstrations and enables zero-shot transfer across different robotic arms, exhibiting strong robustness.
📝 Abstract
A robot demonstration records two things in the same frames: what happened to the objects, and how one particular arm made it happen. We condition on the first. A demonstration is compiled into an effect program: the 3D keypoint trajectories of the objects that moved, two points marking where each was held, and the configuration the scene ends in, with the demonstrator removed. PEWAM, a 71.5M-parameter world-action model, generates effect, robot execution, action and terminal state as four streams with independent flow-matching times, so clamping a program and sampling the execution turns inference into programming, re-solved closed loop from the live scene. On held-out LIBERO-Goal tasks, one demonstration's program completes 40 of 90 episodes, where the same backbone given a goal image or language, and published demonstration-conditioned methods, complete at most 19; on three of Meta-World's held-out classes it exceeds the best published results, though not on the five-class mean. Because a program is a set of coordinates, a person can edit it: the placement follows a shifted terminal state and the grasp turns with rotated contact points. The same program runs on four robot arms without retraining, and after a push, re-solving completes 23 of 60 episodes where replaying the demonstration completes 6. On a Franka arm, fine-tuned on real demonstrations of other tasks, programs compiled from single human videos complete 36 of 40 trials, against 22 for the same backbone conditioned on the video's last frame as a goal image.