Are Frontier VLM Agents Ready to Be Robot Generalists? An Empirical Study with the Embodied Agent Arena

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant gap between the localized perception capabilities of frontier vision-language models (VLMs) and their ability to execute complete tasks as general-purpose robots. To bridge this divide, we construct Embodied Agent Arena, a benchmark comprising thousands of cases, and introduce the GeoProbe evaluation framework. By employing hybrid validation through Blender rendering and real-world scenes, our approach decouples perception accuracy from task completion via minimal testing, enabling a systematic assessment of seven prominent VLMs in embodied scenarios. This work explicitly delineates the boundary between perceptual precision and functional deployment, revealing fundamental shortcomings of current models in coordinated, goal-directed actions. Ultimately, these findings identify critical challenges that must be addressed to advance toward general-purpose robotic agents.
📝 Abstract
Frontier vision-language models (VLMs) combine scene estimation, interaction grounding, and executable actions. Understanding how these abilities support complete robotic tasks is central to evaluating their readiness as robot generalists. We introduce Embodied Agent Arena to examine where local competence supports, or falls short of, complete task success across Geometry, Spatial Reasoning, Affordance, Task Planning, and Manipulation. The arena contains 1,000 cases drawn from 32 established sources and GeoProbe, our new benchmark for geometric estimation on Blender renders and real-scene images. A minimal harness preserves source observations and operations while separating metric precision, functional grounding, and native goal completion. We evaluate seven VLMs, analyze Astra's task-specific advantages, and compare richer-observation execution protocols and multi-round review. Across the arena, Astra's advantage is strongest in precise estimation and usable-contact localization; completing coordinated, goal-directed actions remains the key gap to robot generalism.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Robot Generalists
Embodied Agents
Task Completion
Robotic Manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Embodied Agent Arena
Geometric Estimation
Robot Generalists
Task Planning
🔎 Similar Papers
No similar papers found.