WareFly-VLA: A Vision-Language-Action Framework for UAV Navigation and Human Tracking in Smart Warehouses

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of benchmarks and the challenge of continuous control for language-guided UAV navigation in intelligent warehousing. We construct a simulation framework and dataset based on NVIDIA Isaac Sim, supporting target search, localization, and tracking under natural language instructions. This work proposes the first joint Vision-Language-Action (VLA) benchmark for UAVs that integrates continuous low-level actions, fine-grained descriptions, and realistic industrial environments, employing a four-degree-of-freedom control architecture. Experimental results reveal that existing VLA models exhibit insufficient generalization capabilities in aerial scenarios. Furthermore, our findings demonstrate that continuous action modeling significantly outperforms discrete tokenization schemes, highlighting its necessity for effective UAV control in complex warehouse environments.
📝 Abstract
Vision-Language-Action (VLA) models have achieved impressive results in robotic manipulation and ground-mobile navigation, yet language-conditioned control of unmanned aerial vehicles (UAVs) in smart warehouses remains largely unexplored, hindered by the lack of benchmarks that jointly provide continuous low-level flight actions, fine-grained natural-language target descriptions, and realistic industrial environments. This paper introduces WareFly-VLA, a photorealistic UAV VLA framework and dataset for language-guided human search, localization, and tracking in warehouse environments. It contains 507 human-teleoperated flight episodes and 8,504 high-resolution RGB transitions collected in NVIDIA Isaac Sim, each paired with a human-written appearance description of the target worker and a synchronized four-degree-of-freedom control command. Two aerial tasks are covered: target approach and person following, under occlusion, long-range search, altitude variation, and clutter. A unified benchmark of four open-source VLA architectures (SmolVLA, GR00T N1.7, pi_0 and OpenVLA) is established under a leakage-free episode-level protocol at two control rates. The results show that language-conditioned aerial control in warehouses is far from solved: performance drops substantially under strict generalization settings, continuous action modeling consistently outperforms discrete action tokenization, only the forward channel is reliably learnable from a single frame, and current foundation-model interfaces transfer poorly from ground and humanoid embodiments to aerial platforms. The synchronized video, language, action, pose, and difficulty annotations further support world-model research. The dataset, baselines, and evaluation protocol are released to support language-grounded aerial autonomy in smart warehouses.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
UAV navigation
human tracking
smart warehouses
language-conditioned control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action
UAV Navigation
Smart Warehouses
Human Tracking
Continuous Action Modeling
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Thinh D. Le
Center for AI Research, VinUniversity, Ho Chi Minh City, Vietnam; College of Engineering and Computer Science, VinUniversity, Hanoi, Vietnam
S
Son T. Nguyen
Center for AI Research, VinUniversity, Ho Chi Minh City, Vietnam; College of Engineering and Computer Science, VinUniversity, Hanoi, Vietnam
D
Duong Q. Nguyen
Center for AI Research, VinUniversity, Ho Chi Minh City, Vietnam; College of Engineering and Computer Science, VinUniversity, Hanoi, Vietnam
D
Dung D. Le
Center for AI Research, VinUniversity, Ho Chi Minh City, Vietnam; College of Engineering and Computer Science, VinUniversity, Hanoi, Vietnam
Ngo Anh Vien
Ngo Anh Vien
VinRobotics & VinUni, ex-BCAI
machine learningrobotics
H. Nguyen-Xuan
H. Nguyen-Xuan
CIRTech Institute, HUTECH University, Vietnam
Computational EngineeringBiomorphic IntelligenceBioinspired Material Design3D Printing