🤖 AI Summary
This study addresses the challenges of manual reward engineering and the difficulty of integrating perception with control in robotic skill development by proposing the RPG framework. This method leverages offline data to construct simulated practice tasks and introduces a novel self-improvement mechanism that requires no model weight updates. Specifically, it employs multimodal large language models, combining privileged state information with video feedback to diagnose failures, dynamically reconstruct a reusable symbolic skill library, and iteratively optimize prompts. Experimental results demonstrate that this framework increases the success rate from 28.6% to 95.0% across 22 manipulation tasks, significantly outperforming existing baselines. Furthermore, it achieves a 100% success rate over 30 real-world physical trials, validating its superior generalization capability and practical utility.
📝 Abstract
Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reuse. At test time, a multimodal LLM uses the resulting system prompt and skill library to coordinate perception and robot control. On held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks. Project Website: https://rpg-robot.github.io/