🤖 AI Summary
This study addresses the inference inefficiency and excessive token consumption caused by frontier models controlling robots through action-by-action execution. To mitigate these issues, this work proposes the URAI framework, which decouples decision-making from control. Specifically, a programming agent constructs a reusable tool library, while an execution agent invokes these tools and provides feedback for iterative optimization. This design preserves model-level decision-making authority without requiring updates to the base model weights. Experimental results demonstrate that URAI improves the success rate on the RoboDojo benchmark from 18% to 53% and accelerates execution speed by 1.3 to 1.5 times. Furthermore, its effectiveness is validated on real-world dual-arm robotic tasks.
📝 Abstract
Frontier models can control robots, but reasoning through every reach, grasp, and retreat makes manipulation slow and token-intensive. We revisit code as policy with a different division of labor: models build executable tools, code handles multi-phase motions, and models decide what to do next. We introduce URAI (Universal Robot-Agent Interface), which couples a programming agent that constructs robot tools with an execution agent that uses them in a feedback loop. The programming agent writes reusable and task-specific tools from task intent and refines them through execution feedback and human guidance. The execution agent selects and parameterizes these tools from current observations; each call runs a complete motion locally before returning control to the agent. Unlike delegating subsequent decisions to a generated program, this design retains model-level decision-making between tool executions. Validated tool revisions persist across episodes without updating foundation-model weights, and a shared GUI and API make the same tools available to humans and agents. Across five RoboDojo tasks and four frozen execution agents, URAI raises aggregate success from 18.0% to 53.0% relative to direct fingertip control, with the largest gain on Swap Blocks; with the same tools, a program written in advance reaches only 24% against 56% for two agents deciding after each call. Three of the four agents also finish episodes 1.3-1.5 times faster with 1.5-1.7 times fewer execution-agent output tokens; DeepSeek-V4-Flash's cost barely changes. We further evaluate URAI on seven real-world AgileX dual-arm tasks, spanning object manipulation, cloth folding, and human-interactive tic-tac-toe. URAI connects the coding and decision-making capabilities of frontier agents, organizing robot control around reusable tools that agents can both invoke and revise.