π€ AI Summary
This study investigates whether the performance gains from world model post-training stem primarily from enhanced predictive capabilities or from incidental effects of the optimization process. To disentangle predictive learning from optimization dynamics, we propose an experimental paradigm employing mismatched objectives and random rewards. Specifically, agent training is audited by substituting authentic feedback with erroneous observational targets and stochastic signals within interactive text environments and VisualWebArena. Our findings reveal that post-training substantially improves task coverage even in the absence of accurate predictions or environmental information. Notably, while mismatched objectives degrade predictive accuracy, they preserve task-level improvements. Furthermore, training with random rewards yields a 14.3% increase in pass@64 on VisualWebArena, confirming that task exploration can be effectively expanded without genuine environmental feedback. These results highlight the critical role of optimization side effects in driving post-training performance enhancements.
π Abstract
Predicting how an environment will change before acting is a natural route to better decision making for agents. Recent post-training methods therefore require agents to predict the next observation and turn that prediction into a reward or a direct supervision signal, which is called world model. Existing next-observation training methods help the agent to learn the environmental content. However, they additionally involve an optimization process, which may introduce several effects other than learning to predict the world. Consequently, where the performance gain comes from during the training process remains an open question. We answer this research question through replacing true next-observation targets with in-distribution mismatched observations during the training process. Across two interactive text environments, mismatched targets lower prediction accuracy by 15.3-61.6% relative to ground-truth targets, yet retain substantial task gains over the base model. Compared with the base model, trained models consider more candidate actions and exhibit less looping. We also introduce a setting that replaces prediction-based rewards with independent random signals. This training expands task coverage (pass@64) even when the reward carries no environment information. We also generalize this finding to VisualWebArena, where random-reward training raises pass@64 by 14.3% relative to the base model, without observation-matching rewards or an external multimodal teacher for reward construction.