Affordance-Conditioned Decision Making: Bridging the Semantic-Spatial Gap in Zero-Shot Cross-Floor Vision-and-Language Navigation
This study addresses the challenge in zero-shot vision-and-language navigation where high-level semantic intents are difficult to translate into reliable physical execution, particularly in cross-floor scenarios constrained by spatial limitations and error accumulation. To this end, this work proposes the PACE module, which introduces a novel passability-aware pose anchoring mechanism that maps transitional semantics into traversable poses to condition short-horizon action generation. Furthermore, it incorporates failure-aware preference fine-tuning to enhance closed-loop error correction, achieving precise alignment between semantic planning and physical execution. When integrated into six open-source navigators, the proposed method improves cross-floor success rates on R2R-CE and RxR-CE by 27.65% and 12.06%, respectively, while demonstrating robust generalization and reliability in real-world unseen environments.