RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant reliability gap that emerges when general-purpose agents transfer digital capabilities to physical-world robotic tasks. To investigate this, we construct a simulation testbed encompassing 84 tasks across manipulation, locomotion, and driving, and propose an execution-trajectory-based analytical framework. By integrating multimodal agents, image segmentation, and dynamics computation, this work systematically evaluates agent performance within perception and control workflows. Our analysis quantifies capability disparities among different models on spatially constrained and dynamically balanced tasks, revealing behavioral deficiencies such as state loss and error-correction failures. Crucially, we identify a key bottleneck: while these agents can construct complex control pipelines, they struggle to reliably compose behaviors. These findings provide empirical evidence and clear directions for enhancing the reliability of physical-world agents.
📝 Abstract
General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities carry into the physical world. To investigate this, we introduce RobotWorld, a challenging simulation testbed for robot use: turning instructions and observations into physical task execution through robot interfaces. Its 84 tasks span manipulation, mobile manipulation, locomotion, driving, and aerial control, with explicit interaction budgets and executable success checks. By analysing task outcomes alongside execution traces, we identify both the capabilities that transfer and the gaps that prevent reliable completion. Furthermore, we find that current agents can construct sophisticated perception and control workflows, including image segmentation, camera calibration, spatial estimation, and dynamics-based computation. These capabilities, however, do not consistently compose into successful behaviour: agents lose task-relevant object states despite reaching commanded poses, fail to correct ineffective actions, recover too late, or mistake unfinished tasks for completion. This uneven transfer also differs across models: Astra succeeds more often on spatial and constrained-contact goals, whereas Opus 5.5 succeeds more often on continuous-balance and timed-interaction goals. By linking these outcomes to execution behaviour, RobotWorld provides both a rigorous proving ground and an empirical account of the remaining capability gaps, thereby establishing concrete targets for training and designing more reliable physical-world agents.
Problem

Research questions and friction points this paper is trying to address.

multimodal agents
robot use
physical world transfer
capability gaps
benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Agents
Simulation Benchmark
Robot Manipulation
Execution Traces
Embodied AI