🤖 AI Summary
This study addresses the decision bias and unreliable updates caused by distribution shifts during test-time adaptation in vision-language navigation. To this end, we propose CGPI, a method that pioneers the extraction of signed credits from action-driven local observation transitions to guide lightweight online updates of a frozen pretrained policy. Furthermore, it introduces a history-based verification mechanism that enables reliable policy improvement and safe rollback without external feedback. Experimental results demonstrate that CGPI achieves consistent performance gains across multiple benchmarks and backbone networks, and successfully validates the feasibility of zero-shot sim-to-real transfer.
📝 Abstract
Test-time adaptation for vision-language navigation (TTA-VLN) enables pretrained policies to adapt online to unseen environments using only test-time observations and interaction history. However, distribution shifts can distort local action preferences and lead to off-course decisions. Existing methods rely on predictive uncertainty, trajectory-level feedback, or accumulated adaptation experience to correct such deviations. These signals, however, do not directly reveal whether an executed action supports instruction-guided progress toward the goal. Moreover, a plausible corrective signal does not guarantee a reliable policy update. The key challenge is thus twofold: identifying interactions that support goal-directed improvement and determining whether the resulting updates are worth retaining. We observe that each executed action induces an immediate observation transition, providing evidence of its local consequences. Based on this insight, we propose Credit-Guided Policy Improvement (CGPI), which recovers signed, reference-relative decision credit from action-induced observation transitions without external outcome feedback. With the pretrained navigation policy frozen, CGPI uses this credit to propose lightweight adaptation updates and verifies them against prior credit-supported interactions. Updates are retained only when supported and rolled back otherwise. CGPI achieves consistent gains across the evaluated VLN benchmarks and navigation backbones, while qualitative robot trials further illustrate the feasibility of zero-shot sim-to-real transfer.