🤖 AI Summary
This study addresses the challenge in zero-shot vision-and-language navigation where high-level semantic intents are difficult to translate into reliable physical execution, particularly in cross-floor scenarios constrained by spatial limitations and error accumulation. To this end, this work proposes the PACE module, which introduces a novel passability-aware pose anchoring mechanism that maps transitional semantics into traversable poses to condition short-horizon action generation. Furthermore, it incorporates failure-aware preference fine-tuning to enhance closed-loop error correction, achieving precise alignment between semantic planning and physical execution. When integrated into six open-source navigators, the proposed method improves cross-floor success rates on R2R-CE and RxR-CE by 27.65% and 12.06%, respectively, while demonstrating robust generalization and reliability in real-world unseen environments.
📝 Abstract
Vision-and-language navigation increasingly relies on general-purpose semantic planners, yet translating correct high-level intent into reliable physical execution remains difficult in spatially constrained transitions. Reaching a staircase, doorway, or narrow passage does not ensure traversal; the agent must identify an executable affordance pose and recover from accumulated action errors. We propose PACE (Preference-refined Affordance-Conditioned Execution), a supervised local execution module that augments frozen zero-shot semantic planners for reliable cross-floor navigation. PACE grounds transition-related semantics into a long-horizon, agent-centric traversable affordance pose and conditions short-horizon action generation on this spatial target, thereby aligning semantic goals with physical execution. We further post-train PACE through failure-aware preference refinement using rollout-derived pairs that contrast normal or recovery behaviors with deviation-amplifying behaviors, thereby improving closed-loop correction. We integrate PACE into six open-source zero-shot VLN navigators and demonstrate consistent improvements on the cross-floor subsets of R2R-CE and RxR-CE, increasing the average success rate from 16.35% to 27.65% and from 4.76% to 12.06%, respectively. Real-world experiments further demonstrate PACE's applicability in unseen environments, highlighting the potential of traversable affordances to bridge semantic intent and reliable embodied behavior.