Affordance-Conditioned Decision Making: Bridging the Semantic-Spatial Gap in Zero-Shot Cross-Floor Vision-and-Language Navigation

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge in zero-shot vision-and-language navigation where high-level semantic intents are difficult to translate into reliable physical execution, particularly in cross-floor scenarios constrained by spatial limitations and error accumulation. To this end, this work proposes the PACE module, which introduces a novel passability-aware pose anchoring mechanism that maps transitional semantics into traversable poses to condition short-horizon action generation. Furthermore, it incorporates failure-aware preference fine-tuning to enhance closed-loop error correction, achieving precise alignment between semantic planning and physical execution. When integrated into six open-source navigators, the proposed method improves cross-floor success rates on R2R-CE and RxR-CE by 27.65% and 12.06%, respectively, while demonstrating robust generalization and reliability in real-world unseen environments.
📝 Abstract
Vision-and-language navigation increasingly relies on general-purpose semantic planners, yet translating correct high-level intent into reliable physical execution remains difficult in spatially constrained transitions. Reaching a staircase, doorway, or narrow passage does not ensure traversal; the agent must identify an executable affordance pose and recover from accumulated action errors. We propose PACE (Preference-refined Affordance-Conditioned Execution), a supervised local execution module that augments frozen zero-shot semantic planners for reliable cross-floor navigation. PACE grounds transition-related semantics into a long-horizon, agent-centric traversable affordance pose and conditions short-horizon action generation on this spatial target, thereby aligning semantic goals with physical execution. We further post-train PACE through failure-aware preference refinement using rollout-derived pairs that contrast normal or recovery behaviors with deviation-amplifying behaviors, thereby improving closed-loop correction. We integrate PACE into six open-source zero-shot VLN navigators and demonstrate consistent improvements on the cross-floor subsets of R2R-CE and RxR-CE, increasing the average success rate from 16.35% to 27.65% and from 4.76% to 12.06%, respectively. Real-world experiments further demonstrate PACE's applicability in unseen environments, highlighting the potential of traversable affordances to bridge semantic intent and reliable embodied behavior.
Problem

Research questions and friction points this paper is trying to address.

Vision-and-Language Navigation
Zero-Shot Cross-Floor Navigation
Affordance
Semantic-Spatial Gap
Embodied AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-and-Language Navigation
Affordance-Conditioned Execution
Zero-Shot Planning
Preference Refinement
Cross-Floor Navigation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xuekang Yang
School of Computer Science, Shanghai Jiao Tong University
Lu Chen
Lu Chen
School of Computer Science, Shanghai Jiao Tong University
Large Language ModelsDialogue SystemsAI for Science
S
Shuang Luo
School of Computer Science, Shanghai Jiao Tong University
J
Jialing Zhu
School of Automation and Intelligent Sensing, Shanghai Jiao Tong University
Q
Qi Zhang
Defense Innovation Institute, Academy of Military Sciences
Y
Yue Gao
MoE Key Laboratory of Artificial Intelligence and AI Institute, Shanghai Jiao Tong University
X
Xiang Zhang
Defense Innovation Institute, Academy of Military Sciences