🤖 AI Summary
This study addresses the limitation that long-horizon language agents rely solely on terminal rewards, lacking effective progress evaluation signals for intermediate steps. To overcome this, it proposes the World Potential Model (WPM), which leverages the inherent world knowledge of pretrained models to identify task-relative progress without fine-tuning, utilizing goal-conditioned evaluation and world potential anchoring techniques. By integrating temporal difference learning, WPM enables step-level credit assignment and guides GRPO policy optimization. This work is the first to demonstrate that pretrained world knowledge can be repurposed for real-time progress perception, serving as a viable alternative to independent value functions or process reward models. Evaluated on benchmarks such as ALFWorld, WPM significantly outperforms purely outcome-supervised baselines, effectively improving task success rates.
📝 Abstract
Long-horizon language agents often receive supervision only from terminal task outcomes, leaving little signal for distinguishing productive intermediate behavior from stagnation or even regression. Rather than learning a separate value function or process reward model for every task, we ask whether pretrained models can recognize task progress from their existing world knowledge. We formalize this capability with a World Potential Model (WPM), a goal-conditioned evaluator of task-relative realized progress in agent contexts. In ALFWorld and ScienceWorld, off-the-shelf pretrained models substantially outperform chance at recovering realized-progress structure without task-specific evaluator fine-tuning. We further anchor these progress judgments to task-specific milestones to obtain scalar world potentials, whose temporal differences provide process-sensitive step-level credit for policy optimization. Under matched comparisons, WPM-guided optimization improves success over outcome-only GRPO across all evaluated configurations. Together, these results provide initial evidence that pretrained world knowledge can support reusable realized-progress evaluation and provide useful supervision for long-horizon agents.