AeroAct: Action-Centered World-Action Models for Language-Conditioned Quadrotor Flight

πŸ“… 2026-07-16
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge of language-guided quadrotor flight, which requires mapping high-level semantic instructions to dynamically feasible control actions. Existing approaches lack fine-grained modeling of future visual states, limiting their performance in real-world deployment. To overcome this, we propose AeroActβ€”the first action-centric world model (WAM) tailored for real quadrotor platforms. Built upon a pretrained video diffusion Transformer, AeroAct integrates first-person visual history, proprioceptive data, and language commands to predict short-horizon trajectory-action segments, leveraging synthetically generated future frames for dense supervision. A novel combination of low-cost handheld data collection and a self-guidance mechanism enhances temporal consistency across predicted trajectory segments. Experiments demonstrate that incorporating temporal visual context significantly improves target tracking and object search capabilities, with the learned policy deployable directly on physical quadrotors without further fine-tuning.
πŸ“ Abstract
Language-conditioned quadrotor flight requires a policy to ground semantic goals, anticipate the visual consequences of ego-motion, and output control references that remain smooth and dynamically executable under rapidly changing first-person views. Existing aerial vision-language navigation and vision-language-action methods commonly use discrete actions, high-level waypoints, or instantaneous velocity commands, which provide limited supervision about how flight actions change future observations. We present AeroAct, an action-centered world-action model (WAM) for quadrotor navigation. To the best of our knowledge, AeroAct is the first WAM instantiated and demonstrated for real-world aerial flight. The model adapts a pretrained video diffusion Transformer to predict local trajectory-action chunks from egocentric visual history, proprioception, and language. Future first-person frames are used during training as dense consequence supervision, while deployment directly decodes actions without generating future video. To obtain aligned visual, state, language, and dynamically feasible action data, we build a DiffAero-based pipeline with complementary Isaac Lab and 3D Gaussian splatting renderers. We further introduce a low-cost handheld collection device that couples camera observations with motion estimates to recreate flight-like egocentric trajectories, and a self-guidance procedure that improves temporal consistency across overlapping trajectory chunks. Closed-loop simulation and real-world experiments show that temporal visual context improves target tracking and object-search performance, and that WAM-based policies can be executed on a physical quadrotor.
Problem

Research questions and friction points this paper is trying to address.

language-conditioned quadrotor flight
vision-language navigation
world-action model
egocentric trajectory prediction
dynamically feasible control
Innovation

Methods, ideas, or system contributions that make the work stand out.

world-action model
video diffusion Transformer
language-conditioned flight
egocentric trajectory prediction
dense consequence supervision
X
Xinhong Zhang
School of Automation, Beijing Institute of Technology
Q
Qiyuan Zhu
School of Automation, Beijing Institute of Technology
Y
Yubo Huang
School of Automation, Beijing Institute of Technology
H
Haolin Chen
School of Mechanical Engineering, Beijing Institute of Technology
R
Runqing Wang
School of Automation, Beijing Institute of Technology
Y
Yuhao Mo
School of Automation, Beijing Institute of Technology
Z
Zhongxin Chen
School of Automation, Beijing Institute of Technology
Y
Yu Hu
Independent researcher
X
Xinjiang Wang
Independent researcher
Jian Sun
Jian Sun
Beijing Institute of Technology
Networked control systemstime-delay systemsSecurity of CPS
Gang Wang
Gang Wang
Beijing Institute of Technology
Distributed learningnon-convex optimizationreinforcement learningdata-driven control