AerialDojo-200K: A Large-Scale Benchmark Suite for Open-World Aerial Object-Goal Search

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of a unified benchmark for open-world aerial target search, which constrains agent generalization and scalability, by constructing a large-scale simulation-based evaluation suite. The proposed benchmark comprises 42 scenarios and 200,000 task instances—representing threefold and 18.7-fold expansions, respectively—spanning four environment categories, including urban and natural settings, alongside long-horizon configurations. It standardizes action spaces and data formats while providing multi-view videos and reference trajectories to facilitate both semantic and image-based search evaluation. Extensive experiments across nine mainstream models reveal significant performance deficiencies in existing methods, indicating that research on general-purpose aerial agents remains in its nascent stage.
📝 Abstract
Open-world aerial object-goal search is a foundational yet challenging task, requiring aerial agents to autonomously explore large-scale, unstructured three-dimensional environments and reach target objects specified by semantic descriptions or reference images, rather than following route-specific instructions. However, research in this task remains at a nascent stage and relies on small, environment-specific benchmarks with heterogeneous action spaces and data formats. These limitations hinder large-scale training and cross-benchmark evaluation, constraining the scalability and generalizability of aerial agents. To address this problem, we propose AerialDojo-200K, a large-scale benchmark suite for open-world aerial object-goal search, with 3 times as many scenes and 18.7 times as many task instances as the largest existing benchmark for this task. Specifically, we construct 42 simulation scenes spanning four scene families and 21 scene types, including 18 urban, 12 natural, six infrastructure, and six disaster scenes. To ensure data quality, 12 annotators spent two months manually annotating 109 landmarks, 2099 target objects, and 2099 object anchors across these scenes. We further construct 205,732 task instances, comprising over 100K semantic-goal and over 100K image-goal instances across Base, Standard, and Long-Horizon settings. Each task instance includes a collision-free reference trajectory and corresponding multi-view video recordings. We also develop a unified evaluation framework with a scene partition comprising 21 in-distribution scenes and 21 out-of-distribution scenes. Finally, our evaluation of five open-source and four closed-source multimodal large language models reveals that there is still a long way to go toward achieving general-purpose aerial agents. All can be found at https://fengtt42.github.io/AerialDojo/.
Problem

Research questions and friction points this paper is trying to address.

open-world aerial object-goal search
benchmark suite
scalability
generalizability
aerial agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Aerial Object-Goal Search
Large-Scale Benchmark
Open-World Navigation
Multimodal Large Language Models
Unified Evaluation Framework
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Tongtong Feng
Tongtong Feng
Tsinghua University
Environment LearningAutonomous Embodied AIMultimedia Intelligence
X
Xin Wang
Department of Computer Science and Technology, BNRist, Tsinghua University
H
Haoran Hou
Department of Computer Science and Technology, BNRist, Tsinghua University
R
Ren Wang
Department of Computer Science and Technology, BNRist, Tsinghua University
Weiran Wang
Weiran Wang
University of Iowa
Machine learningspeech processing
S
Shaokai Zhu
School of Electronics Engineering and Computer Science, Peking University
Z
Ziqi Jia
Department of Computer Science and Technology, BNRist, Tsinghua University
Hao Wang
Hao Wang
Tsinghua University
AI for Software EngineeringAI for SecuritySoftware and System Security
Yu-Wei Zhan
Yu-Wei Zhan
Tsinghua University
Z
Zongyuan Wu
Department of Computer Science and Technology, BNRist, Tsinghua University
J
Jinghao Cui
Department of Computer Science and Technology, BNRist, Tsinghua University
Wenwu Zhu
Wenwu Zhu
Professor, Computer Science, Tsinghua Univerisity
Multimedia ComputingNetwork Representation Learning