🤖 AI Summary
This study addresses the challenges faced by robotic agents in associating intentions with outcomes across repeated attempts, as well as the tendency of existing interfaces to cause context bloat or over-reliance on predefined tools. To this end, we propose RobotUse, a framework that innovatively decouples computation, context, and decision-making. Specifically, the frontend enables visual selection of targets and poses, while the backend handles geometric planning and control. Furthermore, a sub-agent mechanism is introduced to preserve interaction details and update a persistent action manual, facilitating continuous learning from execution feedback and cross-task knowledge transfer. Experimental results demonstrate that the proposed framework achieves a 45% success rate on the RoboLab benchmark, outperforming CaP-X by 6.7 percentage points. These findings confirm its effectiveness in reducing dependence on predefined actions while maintaining a compact context.
📝 Abstract
Robot agents must connect their intended actions to observed outcomes while retaining the context needed to revise their choices over repeated attempts. Existing interfaces often leave these choices inside predefined tools or require agents to manage detailed execution code and its growing history. We introduce RobotUse, a robot agent harness that organizes computation, context, and decisions around specifying and revising physical actions. Agents visually select targets and poses, while the backend handles geometry, motion planning, and control. Subagents retain detailed interactions within each subgoal and return the information needed for subsequent decisions. Continual harnessing lets agents learn from execution by updating a persistent playbook. On RoboLab, RobotUse achieves 45% task success, outperforming CaP-X by 6.7 percentage points while maintaining compact decision contexts and reducing reliance on predefined action abstractions. Furthermore, we show that RobotUse learns from real-world execution despite imperfect feedback and transfers what it learns to subsequent tasks. Project page is available at https://robotuse-team.github.io/.