🤖 AI Summary
This study addresses the accumulation of local decision errors in long-horizon urban navigation and the inherent limitation of imitation learning in leveraging failure experiences. To overcome these challenges, this work proposes a Recursive World-Action Model that establishes a self-improving closed loop through alternating updates between a world model and a policy. By utilizing imagined feedback to optimize the policy and generate an adaptive curriculum, the approach technically integrates a conditional world model, Group Relative Policy Optimization (GRPO), and a behavior-novelty-based self-curriculum learning mechanism. Experimental results demonstrate that the proposed method significantly outperforms existing baselines. Furthermore, real-world trials validate its practical applicability for long-horizon urban navigation tasks.
📝 Abstract
Long-horizon urban navigation requires sequential local decisions whose errors can compound over time. Imitation learning (IL) rarely learns from failures, while physical trial-and-error reinforcement learning (RL) is costly. Action-conditioned world models can provide imagined feedback by predicting visual consequences for candidate actions. However, a frozen world model may become less reliable as the policy evolves. In this paper, we introduce RIWANAV, a post-training framework that casts the coupled adaptation of a world model and an action model (policy) as task-specific recursive self-improvement (RSI). Each cycle alternates two updates. The world model evaluates policy actions through imagined outcomes, providing comparative feedback for group-relative policy optimization (GRPO). The improved policy then constructs a grounded self-curriculum, selecting expert-consistent action-video pairs by behavioral novelty and prediction error. The refined world model supplies feedback for the next policy update, closing the recursive self-improvement loop. Experiments show that RIWANAV outperforms training baselines and prior methods, validating the proposed recursive self-improvement loop between the policy and world model. Real-world trials further demonstrate its practical applicability.