Does Learning to Predict the World Help Agents Act? Auditing World-Model Post-Training

πŸ“… 2026-09-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study investigates whether the performance gains from world model post-training stem primarily from enhanced predictive capabilities or from incidental effects of the optimization process. To disentangle predictive learning from optimization dynamics, we propose an experimental paradigm employing mismatched objectives and random rewards. Specifically, agent training is audited by substituting authentic feedback with erroneous observational targets and stochastic signals within interactive text environments and VisualWebArena. Our findings reveal that post-training substantially improves task coverage even in the absence of accurate predictions or environmental information. Notably, while mismatched objectives degrade predictive accuracy, they preserve task-level improvements. Furthermore, training with random rewards yields a 14.3% increase in pass@64 on VisualWebArena, confirming that task exploration can be effectively expanded without genuine environmental feedback. These results highlight the critical role of optimization side effects in driving post-training performance enhancements.
πŸ“ Abstract
Predicting how an environment will change before acting is a natural route to better decision making for agents. Recent post-training methods therefore require agents to predict the next observation and turn that prediction into a reward or a direct supervision signal, which is called world model. Existing next-observation training methods help the agent to learn the environmental content. However, they additionally involve an optimization process, which may introduce several effects other than learning to predict the world. Consequently, where the performance gain comes from during the training process remains an open question. We answer this research question through replacing true next-observation targets with in-distribution mismatched observations during the training process. Across two interactive text environments, mismatched targets lower prediction accuracy by 15.3-61.6% relative to ground-truth targets, yet retain substantial task gains over the base model. Compared with the base model, trained models consider more candidate actions and exhibit less looping. We also introduce a setting that replaces prediction-based rewards with independent random signals. This training expands task coverage (pass@64) even when the reward carries no environment information. We also generalize this finding to VisualWebArena, where random-reward training raises pass@64 by 14.3% relative to the base model, without observation-matching rewards or an external multimodal teacher for reward construction.
Problem

Research questions and friction points this paper is trying to address.

world model
post-training
next-observation prediction
agent decision making
performance gain attribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Model Post-Training
Mismatched Observation Targets
Random Reward Signals
Task Coverage
VisualWebArena
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
X
Xinyu Che
Xi’an Jiaotong University
Hang Yan
Hang Yan
Xi'an Jiaotong University
LLM reasoningAgentKnowledge Graph
Y
Yanchen Liu
University of Southern California
Haochen Liu
Haochen Liu
University of the Chinese Academy of Sciences
R
Ruifeng Li
University of Southern California
A
Anran Shi
East China Normal University
H
Heng Wang
Xi’an Jiaotong University
J
Jun Liu
Xi’an Jiaotong University