🤖 AI Summary
This study addresses the challenges of high demonstration costs, latency-induced state misalignment, and difficult cross-site composition in vision-language-action models for aerial manipulation. To this end, we propose a unified framework that leverages a reconfigurable scene pipeline to generate synthetic data, designs a measurement-progress alignment mechanism to eliminate latency errors, and utilizes relational scene graphs with topological guidance to enable local skill transfer. Experimental results demonstrate that the proposed approach achieves a 65% success rate for local skills in simulated environments and a 42% completion rate for multi-site tasks within the full system. Furthermore, the effectiveness of the framework is validated through real-world deployment on physical unmanned aerial vehicles.
📝 Abstract
Vision-language-action (VLA) models enable task-conditioned interaction, but extending them to scene-scale aerial manipulation remains challenging due to costly whole-body demonstrations, latency-induced action-state misalignment, and cross-site behavior composition. We present a unified framework for synthetic policy training and scene-scale execution on articulated uncrewed aerial manipulators (UAMs). A scene-reconfigurable pipeline synthesizes task-conditioned, kinodynamically feasible trajectories and synchronized multiview observations for VLA training without physical-platform demonstrations. Measured-progress-aligned realization (MPAR) aligns asynchronously returned action chunks with measured execution progress and realizes them as continuous, dynamically feasible trajectories. A relational Scene Graph grounds language goals to object instances and feasible interaction regions, while topology-guided transfer connects local behaviors across sites. Local VLA skills achieve 39/60 successes (65.0%) in simulation under oracle target and feasible-handoff conditions. Under 500-ms added latency, with and without a transient command-update stall, MPAR reduces median takeover phase error by 0.212 s over nominal-time alignment. The complete system completes 21/50 simulated multi-site missions (42.0%) and is further validated on a physical articulated UAM.