🤖 AI Summary
This study addresses the challenge of structurally attributing, verifying, and reusing robotic experience. To this end, it proposes a multi-agent collaboration framework based on top-down skill refinement, introducing a novel hierarchical skill contract mechanism. This approach enables precise code-level fault localization, isolated revision, and self-repair, thereby ensuring the stability of capability evolution. The proposed method is validated through 100 rounds of real-world household cleaning tasks as well as in simulated environments. Experimental results demonstrate that it outperforms the strongest baseline by 2.7 to 11 percentage points in success rate, effectively achieving reliable consolidation and reuse of robotic skills.
📝 Abstract
A generalist robot should not only perform diverse tasks but also improve through experience, turning what it learns during execution into capabilities that later tasks can reuse. Robot agents that act through code can already repair programs from execution feedback, yet it remains a central challenge to organize this experience around the task structure that gives it meaning, so that each repair is attributed to the responsible capability, supported by execution evidence, and validated before it is reused. We introduce RoboRSI, a robot self-improvement system built on Top-Down Skill Refinement (TSR). TSR decomposes tasks into compound, atomic, and base skills with scoped responsibilities and explicit input--output contracts, attributes each execution outcome to the responsible branch, and confines revision to that branch. Building upon this structure, a Manager, Planner, Engineer, and Reviewer coordinate planning, execution, diagnosis, and the validated release of new skills, while people steer the process through objectives and corrections; stable skill sequences are further consolidated into reusable compound skills. On a mobile manipulator, RoboRSI develops multi-object household cleanup over 104 rounds. In simulation, it achieves the highest success rate on LIBERO, LIBERO-PRO, LIBERO-Plus, and RoboTwin, exceeding the strongest baseline by 2.7 to 11.0 percentage points.