Long-Horizon Textual World Modeling through Structured Reasoning

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of recursive error accumulation, multi-step state tracking difficulties, and sparse supervision signals in long-horizon world models by proposing a novel paradigm that formulates multi-step transitions as structured textual reasoning. The proposed method leverages sparsely changing states to reduce tracking overhead and optimizes training through a predictive gain objective function alongside intermediate rewards. Furthermore, it introduces intermediate trajectory supervision to yield interpretable state representations. Experimental results demonstrate that the model achieves state-of-the-art long-horizon performance on benchmarks such as ScienceWorld. Notably, it is the only approach exhibiting statistically significant counterfactual action sensitivity, effectively enhancing the reasoning capabilities of large language models within complex dynamic environments.
📝 Abstract
World models must predict how an environment evolves under sequences of actions, enabling agents to compare possible futures and reason about counterfactual actions before acting. Long-horizon prediction is commonly obtained by recursively applying a one-step transition model, but intermediate errors can compound over time. Multi-step dynamics models instead condition on a sequence of future actions and predict their consequences directly, but become harder to learn as horizon grows: the model must track interacting state changes across the trajectory, endpoint supervision provides weak credit assignment, and intermediate predictions can remain plausible while losing information needed for later states. We show that these challenges can be addressed by casting the internal evolution of a multi-step transition as structured reasoning over textual world states: reasoning over sparse state changes reduces the burden of state tracking, a predictive-gain objective rewards the learned state for improving over a matched predictor that conditions on raw history instead, and intermediate predictive rewards supervise each state along the trajectory. Because these intermediate states are explicit textual representations of the world, they provide semantically meaningful targets that can be inspected, scored, and corrected during training. Across ScienceWorld, Jericho, and CEO-Bench, our approach achieves the strongest average long-horizon performance against recursive and non-recursive baselines that condition directly on raw history, with gains increasing at longer horizons. In a controlled counterfactual study, our model is also the only one with statistically significant sensitivity to future actions.
Problem

Research questions and friction points this paper is trying to address.

world models
long-horizon prediction
multi-step dynamics
error compounding
credit assignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Models
Structured Reasoning
Long-Horizon Prediction
Predictive-Gain Objective
Textual State Representation