LightVLN: Efficient Aerial Vision-and-Language Navigation with Compact Memory and History-Guided Local Aggregation

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational overhead and onboard deployment challenges in aerial vision-language navigation caused by large models and dense historical data. We propose a lightweight framework employing a 0.5B-parameter language backbone, featuring a novel single-token compression for historical frames and an instruction-conditioned local aggregation mechanism to substantially reduce observation tokens. By integrating visual feature reuse with history-aware memory, the framework significantly lowers inference burden. Evaluated on the OpenFly dataset, our approach outperforms 7B-parameter baselines while achieving real-time inference at 14.61 Hz and end-to-end decision updates at 11.13 Hz on Jetson Orin NX edge hardware, demonstrating both its efficiency and feasibility for onboard deployment.
📝 Abstract
Aerial vision-and-language navigation (VLN) enables unmanned aerial vehicles to execute long-horizon natural-language instructions from visual observations in complex three-dimensional environments. However, recent aerial VLN models often rely on large-scale vision-language backbones and dense visual histories, imposing substantial computation and memory costs that hinder onboard deployment. We propose LightVLN, a lightweight history-aware aerial VLN framework that combines a compact 0.5B language backbone with compact representations of both historical and current observations. LightVLN compresses each historical frame into a single token using visual features already computed by the policy. It further introduces history- and instruction-conditioned local aggregation to reduce the current observation from 256 to 32 visual tokens while preserving navigation-relevant spatial information. With up to 16 historical frames, the policy uses at most 48 observation-derived tokens. On the public OpenFly dataset, LightVLN achieves 50.93% Test-Seen and 36.14% Test-Unseen success rates (SR), outperforming the evaluated 7B language-backbone baselines on most reported metrics. It also achieves 25.83% SR on AerialVLN-S Val-Seen. In a reconstructed unseen campus, we deploy LightVLN on a DJI M350 RTK with an external Jetson Orin NX 16 GB for closed-loop onboard-compute real-to-sim hardware-in-the-loop (HIL) evaluation, achieving 14.61 Hz model inference and 11.13 Hz end-to-end decision updates. These results demonstrate the effectiveness and efficiency of LightVLN for aerial navigation.
Problem

Research questions and friction points this paper is trying to address.

Aerial Vision-and-Language Navigation
Computational Cost
Memory Overhead
Onboard Deployment
Unmanned Aerial Vehicles
Innovation

Methods, ideas, or system contributions that make the work stand out.

Aerial Vision-and-Language Navigation
Lightweight Framework
Compact Memory
Local Aggregation
Onboard Deployment
💼 Related Jobs
No related jobs found.
Y
Yiming Zhao
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China
T
Tianshun Li
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China
J
Jingle He
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China
R
Ruonan Chai
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China
Xinhu Zheng
Xinhu Zheng
Assistant Professor, The Hong Kong University of Science and Technology (Guangzhou)