AeroManip-VLA: Scalable Vision-Language-Action Learning for Aerial Manipulation with RL-Generated Demonstrations

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive costs of data acquisition and the inherent evaluation challenges associated with Vision-Language-Action (VLA) models for aerial manipulators. To this end, we construct a GPU-accelerated simulation benchmark that integrates reinforcement learning with expert rules to automatically synthesize diverse demonstration data. This pipeline spans fundamental skills to long-horizon tasks without requiring manual teleoperation. Furthermore, by introducing automated event annotation and trajectory classification mechanisms, our framework enables fine-grained performance analysis and facilitates a systematic evaluation of various baseline models alongside their failure modes. Ultimately, this work achieves scalable data generation and structured analysis for aerial manipulation, establishing a solid foundation for real-world deployment.
📝 Abstract
Aerial manipulators extend robotic manipulation into 3D workspaces that are difficult for ground-based robots to access, creating new opportunities for general-purpose manipulation. However, extending Vision-Language-Action (VLA) models to aerial robots introduces distinct challenges due to the tight coupling between manipulation and flight, continuously changing observations, and safety-critical physical interactions. These challenges demand diverse training data and systematic policy evaluation, yet collecting demonstrations and evaluating policies directly on physical aerial platforms are costly, difficult to scale, and hard to repeat under controlled conditions. We present AeroManip-VLA, a scalable benchmark for aerial VLA data generation and policy evaluation. AeroManip-VLA provides a GPU-accelerated simulation framework with low-level payload-aware flight and manipulation control in massively parallel environments. Building on this framework, we combine reusable reinforcement learning policies with expert task rules to automatically generate demonstrations without human teleoperation across diverse objects, environments, and randomized initial conditions. The generated data include basic skills such as grasping and placing, as well as long-horizon tasks that require both navigation and manipulation. We further introduce automated event labeling and trajectory categorization to filter demonstrations. These mechanisms enable fine-grained analysis of task progress, behavioral outcomes, and safety-related failures. Finally, we evaluate a range of imitation learning and VLA baselines across different task settings, revealing their performance characteristics and failure modes. Together, AeroManip-VLA enables scalable aerial manipulation data generation, structured trajectory analysis, and systematic VLA evaluation in simulation prior to real-world deployment.
Problem

Research questions and friction points this paper is trying to address.

Aerial Manipulation
Vision-Language-Action
Data Generation
Policy Evaluation
Simulation Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action
Aerial Manipulation
Reinforcement Learning
Scalable Simulation
Automated Demonstration Generation
R
Rui Huang
National University of Singapore
Y
Yanlin Mu
Beijing Institute of Technology, National University of Singapore
L
Lidong Li
National University of Singapore
Y
Yucong Wang
National University of Singapore
Z
Zichen Yan
National University of Singapore
Lin Zhao
Lin Zhao
Assistant Professor, National University of Singapore
control theoryreinforcement learningroboticsautonomous vehiclespower system