Structured World-State Reasoning for Agentic Robotic Search

📅 2026-09-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出WORLDS框架,通过持续更新的图结构和多模态感知解决长期机器人搜索中自然语言与不完整、异质信息匹配的问题。
📝 Abstract
Long-horizon robotic search must resolve natural language against heterogeneous, incomplete, and often ambiguous evidence: textual information, prior maps, and observations arriving over time. The core challenge is to contextualize these streams and decide where to gather evidence before selecting a target. We present WORLDS: World-state Observation and Reasoning for Language-guided Discovery and Search, a framework that grounds reasoning in a persistent graph initialized from geospatial priors and updated by perception. Parallel Reasoners maintain competing candidate interpretations and request evidence to distinguish between them. We collect and process the requested observations with a multimodal Examiner, after which a Judge selects a grounded target or requests another pass. WORLDS achieves 51.8% navigation success across all 5,311 CityNav test episodes, the highest reported success rate, exceeding the previous published best by 15.7 percentage points under an OSM-only, high-resolution orthographic protocol. On 1,000 shared episodes, it achieves 50.0% versus 27.9% for the strongest adapted baseline using the same model, prior, sensing stack, and movement budget. Observation-based verification by the Examiner contributes 5.9 points of this success, and at a reduced reasoning-effort setting WORLDS still exceeds the adapted GeoNav baseline by 18.8 points while generating fewer tokens. We also demonstrate WORLDS on a quadrotor, which flies the generated sensing waypoints and grounds three language targets, including a vehicle absent from the map, from its onboard imagery.
Problem

Research questions and friction points this paper is trying to address.

long-horizon robotic search
natural language
heterogeneous evidence
Innovation

Methods, ideas, or system contributions that make the work stand out.

World-state Reasoning
Parallel Reasoners
Multimodal Examiner
Geospatial Priors
Long-horizon Search
🔎 Similar Papers
F
Finley R. Holt
Department of Aeronautics and Astronautics, Stanford University
L
Luis A. Pabon
Department of Aeronautics and Astronautics, Stanford University
John Irvin Alora
John Irvin Alora
Stanford University
RoboticsControl TheoryOptimizationMachine Learning
J
Jonas Frey
Department of Aeronautics and Astronautics, Stanford University
Marco Pavone
Marco Pavone
Stanford University and NVIDIA
RoboticsControl TheoryDistributed ControlIntelligent Transportation systems