From Local Whole-Body VLA Behaviors to Scene-Scale Aerial Manipulation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of high demonstration costs, latency-induced state misalignment, and difficult cross-site composition in vision-language-action models for aerial manipulation. To this end, we propose a unified framework that leverages a reconfigurable scene pipeline to generate synthetic data, designs a measurement-progress alignment mechanism to eliminate latency errors, and utilizes relational scene graphs with topological guidance to enable local skill transfer. Experimental results demonstrate that the proposed approach achieves a 65% success rate for local skills in simulated environments and a 42% completion rate for multi-site tasks within the full system. Furthermore, the effectiveness of the framework is validated through real-world deployment on physical unmanned aerial vehicles.
📝 Abstract
Vision-language-action (VLA) models enable task-conditioned interaction, but extending them to scene-scale aerial manipulation remains challenging due to costly whole-body demonstrations, latency-induced action-state misalignment, and cross-site behavior composition. We present a unified framework for synthetic policy training and scene-scale execution on articulated uncrewed aerial manipulators (UAMs). A scene-reconfigurable pipeline synthesizes task-conditioned, kinodynamically feasible trajectories and synchronized multiview observations for VLA training without physical-platform demonstrations. Measured-progress-aligned realization (MPAR) aligns asynchronously returned action chunks with measured execution progress and realizes them as continuous, dynamically feasible trajectories. A relational Scene Graph grounds language goals to object instances and feasible interaction regions, while topology-guided transfer connects local behaviors across sites. Local VLA skills achieve 39/60 successes (65.0%) in simulation under oracle target and feasible-handoff conditions. Under 500-ms added latency, with and without a transient command-update stall, MPAR reduces median takeover phase error by 0.212 s over nominal-time alignment. The complete system completes 21/50 simulated multi-site missions (42.0%) and is further validated on a physical articulated UAM.
Problem

Research questions and friction points this paper is trying to address.

Vision-language-action models
Aerial manipulation
Uncrewed aerial manipulators
Action-state misalignment
Cross-site behavior composition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action (VLA)
Aerial Manipulation
Synthetic Policy Training
Measured-Progress-Aligned Realization (MPAR)
Scene Graph
🔎 Similar Papers
W
Weixiang Guo
School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore 639798
Rui Jin
Rui Jin
Cold and Arid Regions Environment and Engineering Research Institute, Chinese Academy of Sciences
surface freeze/thaw cyclessoil moisturemicrowave remote snesingWSNdata assimilation
H
Haotian Jin
School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore 639798
Xinhang Xu
Xinhang Xu
NTU
R
Ruiyang Liu
School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore 639798
H
Haoran Zhao
School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore 639798
Y
Yi Wang
College of Electronics and Information Engineering, Shanghai Institute of Intelligent Science and Technology, Tongji University, Shanghai 201804, China
W
Weiqi Gai
School of Aeronautic Science and Engineering, Beihang University, Beijing 100191, China
Kun Cao
Kun Cao
Tongji University
multi-robot systemsinverse RLsoft robotics
Lihua Xie
Lihua Xie
Professor of Electrical Engineering, Nanyang Technological University
Robust controlNetworked ControlMult-agent Systems