🤖 AI Summary
To address the high cost and limited scale of high-quality real-world demonstration data in robotic imitation learning, this paper proposes a closed-loop learning framework integrating few-shot learning, simulation augmentation, and human-in-the-loop error correction. The method leverages a small set of real demonstrations (3–5 trials) augmented with synthetic data generated via GANs or diffusion models, followed by online policy fine-tuning to enable zero-shot cross-task transfer. Its key contributions are: (1) a novel vision-guided real-time human correction mechanism supporting natural multimodal interaction; (2) an end-to-end control pipeline built on ROS and PyTorch; and (3) empirical validation across manipulation tasks—including bottle collection, stacking, and hammering—achieving >92% success rates. Crucially, the framework transfers zero-shot to an unseen task—beverage tray arrangement—with 86% success, significantly outperforming pure-simulation baselines.
📝 Abstract
Imitation Learning (IL) has emerged as a powerful approach in robotics, allowing robots to acquire new skills by mimicking human actions. Despite its potential, the data collection process for IL remains a significant challenge due to the logistical difficulties and high costs associated with obtaining high-quality demonstrations. To address these issues, we propose a large-scale data generation from a handful of demonstrations through data augmentation in simulation. Our approach leverages affordable hardware and visual processing techniques to collect demonstrations, which are then augmented to create extensive training datasets for imitation learning. By utilizing both real and simulated environments, along with human-in-the-loop corrections, we enhance the generalizability and robustness of the learned policies. We evaluated our method through several rounds of experiments in both simulated and real-robot settings, focusing on tasks of varying complexity, including bottle collecting, stacking objects, and hammering. Our experimental results validate the effectiveness of our approach in learning robust robot policies from simulated data, significantly improved by human-in-the-loop corrections and real-world data integration. Additionally, we demonstrate the framework's capability to generalize to new tasks, such as setting a drink tray, showcasing its adaptability and potential for handling a wide range of real-world manipulation tasks. A video of the experiments can be found at: https://youtu.be/YeVAMRqRe64?si=R179xDlEGc7nPu8i