🤖 AI Summary
This study addresses the challenges of high data costs and unsafe exploration in policy learning for robotic physical systems by proposing a vision-action policy learning framework that eliminates the need for human demonstrations. The method leverages a Model Predictive Path Integral (MPPI) controller to autonomously collect safe data, establishing a dual-mode training paradigm: simulation-based optimization combining Implicit Q-Learning with terminal value function iteration, and real-world Action Chunking Transformer training on offline data. The proposed framework successfully transfers locomotion policies to a Unitree Go2 quadruped robot and accomplishes grasping and force-aware wiping tasks on a Flexiv manipulator. These results validate the effectiveness and generalization capability of the approach across morphologically diverse robotic platforms.
📝 Abstract
Learning robot policies directly on physical systems remains difficult because data collection is costly and policy exploration can be unsafe. We introduce Visual and Motor Policies from Sampling-Based Planning (VAMPS), a framework that uses Model Predictive Path Integral (MPPI) control to train reusable policies without human demonstrations. VAMPS supports two training modes. For one-step proprioceptive policies, it operates iteratively in simulation: the policy warm-starts MPPI, and the refined trajectories provide new supervision as the policy changes. A learned terminal value improves short-horizon planning, while an Implicit Q-Learning (IQL) critic guides the policy update. Iterative refinement outperforms training once on frozen MPPI data, and we transfer the learned locomotion policy to a Unitree Go2. For visuomotor policies, VAMPS operates directly from real-robot data. MPPI uses task-specific state estimates to plan and execute trajectories while recording RGB and sensor observations on a Flexiv Rizon 10S. Action Chunking with Transformers predicts action chunks, reducing the effective prediction horizon, and is trained offline on this fixed dataset. We demonstrate visuomotor pick-and-place and force-aware whiteboard erasing. In the latter task, the policy additionally observes the measured $6$-D wrench and desired normal force. These results show that VAMPS can learn policies either in simulation followed by hardware transfer or directly from autonomously collected real-robot data.