AVERT-VLN: Abstention-aware Visual Error Recovery and Training for Vision-and-Language Navigation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of execution deviation in vision-language navigation within unseen environments, where unfamiliar layouts and the absence of human supervision hinder performance. To this end, this work proposes a closed-loop human-AI collaborative framework. Methodologically, it introduces an asynchronous sidecar monitoring mechanism that leverages vision-language models to independently evaluate instruction consistency and trigger selective human intervention. Furthermore, a counterfactual risk trajectory dataset is constructed to enable offline policy optimization through trajectory-anchored preference learning. Experimental results demonstrate that the proposed approach achieves success rates of 76.2% and 66.3% on the R2R-CE and RxR-CE benchmarks in unseen environments, respectively, significantly enhancing generalization performance across diverse architectures.
📝 Abstract
Deploying vision-and-language navigation (VLN) agents in unseen environments remains challenging because unfamiliar layouts and visual conditions can cause execution to go off track. Rather than relying on continuous human supervision, a practical strategy is to selectively request corrective guidance, recover the ongoing task, and reuse corrective interactions to improve subsequent navigation. We propose Abstention-aware Visual Error Recovery and Training for Vision-and-Language Navigation (AVERT-VLN), a closed-loop framework that uses a plug-in vision-language Monitor for online human-assisted recovery and offline preference learning. The Monitor operates separately from navigation decision generation and assesses instruction-execution consistency from the instruction, visual history, and current observation. To train the Monitor for deviation recognition, we construct LOSTNAV DATASET with 20K counterfactual risk trajectories and rule-based deviation labels. The Monitor is first fine-tuned on 40K normal trajectories to assess instruction progress and then jointly fine-tuned on normal and risk trajectories to recognize semantic deviations. At runtime, Asynchronous Sidecar Monitoring evaluates execution alongside the navigation model. When the controller accepts a LOST verdict, it suspends autonomous execution and requests human guidance for recovery. For offline policy improvement, Trajectory-Anchored Preference Learning converts deviation-associated failures into decision-level preference pairs under shared decision contexts, restricting supervision to the decisions targeted for correction. Under human-assisted evaluation, the full AVERT-VLN system achieves success rates of 76.2% and 66.3% on the val-unseen splits of R2R-CE and RxR-CE, respectively. The same monitoring and human-assisted recovery interface also improves success rates across the three evaluated navigation architectures.
Problem

Research questions and friction points this paper is trying to address.

Vision-and-Language Navigation
Error Recovery
Deviation Recognition
Human-Assisted Correction
Unseen Environments
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-and-Language Navigation
Abstention-aware Monitoring
Error Recovery
Preference Learning
Human-in-the-loop
🔎 Similar Papers
M
Minrui Liu
Institute of Cyber-Systems and Control, Zhejiang University, Hangzhou, China
J
Jingke Wang
Institute of Cyber-Systems and Control, Zhejiang University, Hangzhou, China
Y
Yuehao Huang
Institute of Cyber-Systems and Control, Zhejiang University, Hangzhou, China
H
Hao Su
Institute of Cyber-Systems and Control, Zhejiang University, Hangzhou, China
Jiajun Lv
Jiajun Lv
Zhejiang University
SLAM
Yukai Ma
Yukai Ma
Zhejiang University
Yong Liu
Yong Liu
Institute of Cyber-Systems and Control, Zhejiang University
Robotic Vision and PerceptionGraphicsInformation Fusion