AeroEval: Staged Program and Execution Validation for AI-Generated Drone Missions

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that drone control code generated by large language models (LLMs), while syntactically correct, frequently violates task intentions and physical constraints. To overcome this limitation, this work proposes an agent-assisted middleware framework featuring a novel dual-layer verification architecture that integrates program-level and execution-level agents. By combining static analysis with simulated trajectory evaluation, the framework precisely localizes errors through staged verification and provides structured feedback to guide iterative code refinement. Experimental results demonstrate that the proposed approach increases navigation success rates from 55% to 95% and overall mission success rates from 34% to 88%, substantially enhancing the reliability and safety of LLM-generated drone control code.
📝 Abstract
Large Language Models (LLMs) can generate drone programs from natural-language mission descriptions, but syntactically valid programs may still violate user intent, environmental constraints, and mission-level behavior. This problem is pronounced in cyber-physical applications, where correctness depends on the interaction among generated code, mobile sensing, environmental geometry, event-driven analytics, and physical execution. Existing drone code-generation systems primarily use prompt guardrails or simulator outcomes and provide limited failure localization. We present AeroEval, an agent-assisted middleware for staged validation of AI-generated drone missions. AeroEval combines deterministic program analysis with context-grounded LLM agents. It first validates program syntax, platform API usage, and mission intent, and then evaluates the realized behavior using execution trajectories, mission requirements, and environmental context. Each stage returns structured failure information for iterative regeneration. In our evaluation using 20 navigation tasks and five analytical mission types over AirSim and Gazebo simulators, AeroEval improves navigation success from 55% to 95%. In a stagewise ablation study, our Code and Trajectory Validators by themselves achieve mean run-level success rates of 44% and 56%, respectively, while the full AeroEval pipeline achieves 88%; the stages detect complementary failures in program structure, API usage, mission intent, obstacle avoidance, altitude, coverage, and event-driven transitions and the guided regeneration corrects for them. Across the main analytics missions, AeroEval increases aggregate run-level success from 34% for one-shot AeroGen to 88% within the regeneration budget. These results demonstrate the benefit of combining program-level and execution-grounded agentic validation for AI-generated drone applications in the evaluated environment.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Drone mission generation
Program validation
Cyber-physical systems
Failure localization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Staged Validation
LLM Agents
Drone Missions
Iterative Regeneration
Deterministic Program Analysis