SearchWorld: Spatial Value-Grounded Imagination for UAV Object Search via World Models

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the planning challenges faced by UAVs in partially observable urban environments, where limited fields of view, geometric complexity, and open-ended instructions hinder navigation. To this end, it proposes a world model that integrates explicit spatial memory with value-guided imagination. Specifically, a recurrent state-space world model is constructed to perceive spatial value maps through bird’s-eye-view (BEV) exploration and obstacle memory decoding tasks. A novel cognitive-action network anchors the imagination process to explicit spatial representations, enabling efficient lookahead reasoning via spatial value priors without requiring an independently trained scalar critic. Experimental results demonstrate that the proposed method achieves a 23.8% success rate on the UAV-ON benchmark, significantly outperforming existing state-of-the-art agents, while maintaining robust generalization with a 19.9% success rate in unseen scenarios.
📝 Abstract
Autonomous unmanned aerial vehicle (UAV) object search involves a closed loop of perception, decision-making, and action under partial observability. Urban environments pose several challenges: large search areas and narrow egocentric views limit coverage, dense 3D geometry constrains safe motion, and open-world instructions require identifying a specific target among distractors. Many existing methods mitigate partial observability through explicit maps or memory representations, yet remain largely reactive, reasoning over past observations without explicitly predicting future states. World models enable prospective reasoning through imagined rollouts. However, image-generating world models can incur high inference latency, while spatially grounded planning remains challenging for latent world models. We propose SearchWorld, a recurrent state-space world model that connects explicit spatial memory with value-guided imagination. The model maintains BEV exploration and obstacle memory and decodes a task-aware spatial value layer to guide search. A cognition-action network uses this learned spatial value prior to improve the policy through imagined rollouts, without training a separate scalar critic. Training progresses from world-model learning to expert imitation and imagination-based exploration refinement. On UAV-ON, SearchWorld improves the success rate to 23.8% (19.5% for the strongest published agent) and raises oracle success to 35.5%, while remaining robust on unseen scenes (19.9% success rate). By grounding imagination in explicit spatial representations, SearchWorld enables UAV agents to plan prospectively rather than react.
Problem

Research questions and friction points this paper is trying to address.

UAV object search
partial observability
world models
spatial grounding
prospective reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Models
Spatial Value Grounding
UAV Object Search
Imagined Rollouts
Bird's Eye View Memory
Y
Yatai Ji
National Key Laboratory of Digital Intelligent Modeling and Simulation, National University of Defense Technology; College of Systems Engineering, National University of Defense Technology
Z
Zhengqiu Zhu
National Key Laboratory of Digital Intelligent Modeling and Simulation, National University of Defense Technology; College of Systems Engineering, National University of Defense Technology
Y
Yong Zhao
National Key Laboratory of Digital Intelligent Modeling and Simulation, National University of Defense Technology; College of Systems Engineering, National University of Defense Technology
Y
Yue Hu
National Key Laboratory of Digital Intelligent Modeling and Simulation, National University of Defense Technology; College of Systems Engineering, National University of Defense Technology
F
Fanglong Yao
Aerospace Information Research Institute, Chinese Academy of Sciences
Chen Gao
Chen Gao
BNRist, Tsinghua University
Data MiningLLM AgentEmbodied AI
P
Pengfei Zhu
Southeast University
Q
Quanjun Yin
National Key Laboratory of Digital Intelligent Modeling and Simulation, National University of Defense Technology; College of Systems Engineering, National University of Defense Technology