Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents

📅 2026-07-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work evaluates the capability of general-purpose multimodal large language models (MLLMs) to solve aerial, long-horizon, high-level instruction-driven embodied tasks without fine-tuning. To this end, we introduce MissionBench, a benchmark comprising 120 tasks across five 3D simulated environments, and establish the first task-level zero-shot evaluation framework tailored for aerial embodied agents. This framework requires agents to autonomously plan, navigate, and report outcomes using only egocentric observations and action history. Our experiments encompass 22 prominent MLLMs, revealing that even the best-performing model achieves a success rate below 35%, substantially lagging behind human performance (84.4%). The results underscore the critical role of multi-step planning and adaptive reasoning in task success and demonstrate that scaling model size effectively enhances zero-shot embodied capabilities.
📝 Abstract
Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, yet it remains unclear how well general-purpose models can solve long-horizon embodied tasks from a single high-level instruction. We introduce MissionBench, a benchmark for mission-level evaluation of MLLMs in aerial 3D environments. It comprises 120 missions across five simulated 3D environments and four task families. Agents must autonomously plan, navigate, and report outcomes using only egocentric observations and its action history, without aerial-specific fine-tuning. Across 22 open- and closed-source MLLMs, the strongest model succeeds on fewer than 35% of missions compared to 84.4% human performance, highlighting the difficulty of multi-step embodied tasks. Despite large variations between model families, we observe gains from scaling, indicating that larger general-purpose models possess stronger zero-shot embodied capabilities. Our analysis shows that mission-level competence requires coordinating multiple capabilities beyond spatial perception, including multi-step planning and adaptive reasoning. This motivates closed-loop evaluation and highlights both the promise and risk of scaling-driven improvements for embodied AI.
Problem

Research questions and friction points this paper is trying to address.

Zero-Shot
Embodied AI
Mission-Level Evaluation
Aerial Agents
Multimodal Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zero-Shot Evaluation
Mission-Level Benchmarking
Aerial Embodied Agents
Multimodal Large Language Models
Closed-Loop Embodied Reasoning