RoboQuest: Generalist Physical Agents that Search, Inspect and Test

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inability of agents to execute actions effectively in unknown environments due to missing information by proposing a goal-directed embodied exploration paradigm. We construct an embodied exploration benchmark that requires agents to actively acquire information through physical interaction, defining three categories of uncertainty tasks: searching, inspecting, and testing. Methodologically, our approach integrates multimodal foundation models, visuomotor interfaces, and full-trajectory demonstration fine-tuning strategies to drive agent exploration. Experimental results reveal that even the best-performing agent achieves only a 23% success rate, exposing critical bottlenecks such as premature termination of exploration and insufficient disturbance recovery capabilities. These findings establish an important benchmark and highlight promising new directions for future research in embodied exploration.
📝 Abstract
Recent advances in multimodal foundation models have made them capable generalist physical agents for a range of manipulation tasks. However, successful operation in an unfamiliar environment may require an agent to seek task-relevant information through interaction when it is absent from the observations: it may need to determine where a relevant object is, inspect an unobserved property, or discover the effect of an unfamiliar tool. We thus introduce RoboQuest, a benchmark for goal-directed embodied exploration, where agents must actively acquire task-relevant information through physical interaction, use the resulting evidence to adapt subsequent actions, and autonomously decide when to commit to task completion. RoboQuest comprises ten mobile manipulation tasks centered on three forms of uncertainty: search, manipulation-based inspection, and interactive testing. We evaluate five frontier multimodal agents through a common visuomotor interface, as well as a $π_{0.5}$ policy fine-tuned on the full-episode demonstrations we release. The best agent succeeds in only 23\% of the episodes, and the fine-tuned policy almost never succeeds. Isolated tests of the execution skills the tasks are built from, with the hidden information supplied, show that the agents can carry out most of the required actions, and our failure analysis attributes only a minority of the failures to execution. Our failure analysis further finds that the agents often stop exploring too early as they make decisions before observing the required evidence for task completion. We also find that agents rarely prevent or repair the disturbances caused by their exploration. Moreover, learning by trial and error remains difficult for most models.
Problem

Research questions and friction points this paper is trying to address.

embodied exploration
physical agents
information seeking
mobile manipulation
multimodal foundation models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Embodied Exploration
Multimodal Foundation Models
Mobile Manipulation
Goal-directed Interaction
Benchmark
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Liu Renhang
Nanyang Technological University
Navonil Majumder
Navonil Majumder
Singapore University of Technology and Design
Natural Language ProcessingMachine LearningNeural NetworksDeep Learning
T
Tej Deep Pala
Nanyang Technological University
S
Soujanya Poria
Nanyang Technological University