NavJev: Efficient Vision-Language Navigation via Action-Centric Visual Compression and Discriminative Action-Semantic Memory

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high inference latency in zero-shot vision-language navigation caused by repeated invocations of multimodal large language models. To this end, we propose an action-centric online navigation framework. Methodologically, an action-centric visual compression mechanism integrates geometric features, captions, and semantic labels to construct compact representations, while a discriminative action-semantic memory module filters redundant shared semantics and maintains action-specific evidence. This design eliminates repetitive autoregressive generation, enabling efficient structured probabilistic decision-making. Experimental results on the R2R-CE dataset demonstrate that the proposed method achieves a 27.0% success rate and 22.4% SPL, requiring only 0.65 seconds per inference step, thereby significantly reducing system latency and computational overhead.
📝 Abstract
Recent zero-shot Vision-and-Language Navigation (VLN) methods increasingly rely on multimodal large language models (MLLMs) to reason over visual observations, navigation instructions, and candidate actions. Although effective, repeatedly invoking autoregressive multimodal reasoning at every navigation step introduces substantial inference latency, limiting the responsiveness of embodied agents. We propose NavJev, an efficient VLN framework that reformulates online navigation from repeated multimodal generation into compact visual compression followed by lightweight typed action selection. Specifically, Action-Centric Visual Compression (ACVC) integrates waypoint geometry, BLIP captions, and RAM semantic tags into compact representations of candidate actions, while Discriminative Action-Semantic Memory (DASM) filters shared semantics and maintains discriminative action-specific evidence across navigation steps. Based on these representations, Jev directly performs structured probabilistic decisions over the available action set. Experiments on R2R-CE show that NavJev achieves 27.0% SR and 22.4% SPL with only 0.65 s per navigation step, while substantially reducing inference latency and cost compared with MLLM-based VLN methods. The project page is available at https://kai-sheng-caesar.github.io/NavJev/.
Problem

Research questions and friction points this paper is trying to address.

Vision-and-Language Navigation
Inference Latency
Multimodal Large Language Models
Embodied Agents
Zero-shot VLN
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Navigation
Visual Compression
Action-Semantic Memory
Zero-shot Reasoning
Efficient Inference
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kai Sheng
College of Electronic and Information Engineering, Tongji University, Shanghai, China
Liuyi Wang
Liuyi Wang
Tongji University
computer visionnatural language processingartificial intelligence
Jinlong Li
Jinlong Li
University of Science and Technology of China
Data miningmachine learningdeep learningbig data
H
Haojie Dai
College of Electronic and Information Engineering, Tongji University, Shanghai, China
C
Chengju Liu
College of Electronic and Information Engineering, Tongji University, Shanghai, China
Q
Qijun Chen
College of Electronic and Information Engineering, Tongji University, Shanghai, China