🤖 AI Summary
This study addresses the high data demands and reliance on failure modes when adapting robotic manipulation policies to novel tasks. We propose a recursive self-improvement (RSI) framework operating through a real-to-sim-to-real cycle. The method leverages an agent-based system for collaborative error correction and adaptive data collection, utilizing execution feedback loops to guide demonstration generation and policy updates, while employing deployment-specific simulation for efficient iterative optimization. Experimental results demonstrate that the proposed framework increases simulation success rates from 50.4% to 83.5% and achieves an 83.1% real-world success rate with only a few trajectories. This performance significantly surpasses conventional methods, enabling highly efficient task adaptation at minimal data cost.
📝 Abstract
Adapting robot manipulation policies to new tasks and environments remains highly data-intensive, while the data needed for further improvement depends on the policy's current capabilities and failure modes. We introduce EmbodiRSI, an agentic system for recursive self-improvement (RSI) in a real-to-sim-to-real setting, where task-specific simulations are constructed from target deployment scenarios and used as low-cost environments for iterative policy improvement before transfer back to the physical world. EmbodiRSI uses policy execution feedback to guide subsequent experience acquisition and policy updates. Two complementary mechanisms close this loop: Collaborative Error Correction generates agent-assisted corrective trajectories from policy-reached states, while Adaptive Data Collection directs expert demonstration generation toward the current policy's weaknesses. The task-specific simulation serves as a reusable workspace for policy warm-up, repeatable evaluation, failure diagnosis, and targeted data generation across successive RSI rounds. Across three tabletop environments and 14 subtasks, EmbodiRSI increases scene-balanced autonomous simulation success from 50.4% to 83.5% over two RSI updates. With 400 adaptive simulated trajectories and only ten real-world refinement trajectories per subtask, EmbodiRSI achieves 83.1% scene-balanced autonomous real-world success, compared with 75.0% for adaptation using 200 real-world demonstrations per subtask. These results demonstrate that feedback-driven recursive improvement in deployment-specific simulations can enable data-efficient adaptation of embodied policies to physical environments.