World Agent: Can Language Models Keep a World Running?

πŸ“… 2026-09-26
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation of existing world model evaluations, which are confined to static generation or single-step transitions and thus fail to verify a model’s capacity to sustain continuous world operation and organize causal event flows. To this end, this work proposes the World Agent task and benchmark, comprising two tracks: maintenance, which coordinates event states, and deduction, which predicts and acts upon partial observations. This framework shifts the evaluation paradigm from outcome delivery to continuous operation while providing an auditable mechanism for checking causal constraints. The assessment integrates LLM-based semantic judgment with programmatic verification techniques. Experiments across eight models demonstrate that performance degrades significantly when pre-built structures are removed, revealing causal relationship modeling as a pervasive bottleneck and effectively quantifying models’ capabilities in sustaining continuous world dynamics.
πŸ“ Abstract
World models are moving from generating realistic frames to generating playable worlds, yet whether a delivered world can keep running is not tested anywhere. Existing evaluations stop at generation, at delivery, or at single-step transitions, and each stops at a different point along the way. Correct local state transitions or intermediate outcomes do not guarantee a correctly organized causal event flow. We propose the world agent task, which moves the evaluation point of world generation from the moment of delivery to the continued operation that follows. In this task, a model is not asked to generate a world. It is held responsible for keeping the world running, which requires coordinating events and carrying forward their consequences to constrain subsequent evolution. We instantiate the task in WorldAgent-Benchmark with two complementary tracks. In the maintenance track, the model must ground the events of a continuous narrative into correct transitions of the explicit world state while respecting causal, temporal, and concurrency constraints. In the deduction track, the model must predict how the world will evolve under partial observations and act toward a goal. The maintenance track combines LLM-assisted semantic judgments with programmatic validation and scoring, while the deduction track is evaluated entirely programmatically. Individual judgments are auditable against world states and execution logs, and scores can be recomputed from the saved judgments and execution records. Across 8 models, scores decline steadily as pre-built structure is removed from the world, and causal-relation checking is the weakest component for every model. The benchmark makes the continued operation of a world measurable and distinguishes local completion from failures in event organization. Code and dataset will be released on https://github.com/HCPLab-SYSU/WorldAgent-Benchmark.
Problem

Research questions and friction points this paper is trying to address.

World Models
Language Models
Causal Reasoning
Evaluation Benchmark
World Agent
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Agent
World Model Evaluation
Causal Reasoning
Benchmark
State Transition
πŸ”Ž Similar Papers
2023-08-22Frontiers Comput. Sci.Citations: 866
Weixing Chen
Weixing Chen
Sun Yat-sen University
CausalityEmbodied AIMedical Image Analysis
W
Weipeng Zhang
South China University of Technology
N
Nan An
Sun Yat-sen University
Y
Yang Liu
Sun Yat-sen University, Guangdong Key Laboratory of Big Data Analysis and Processing, X-Era AI Lab
Liang Lin
Liang Lin
Fellow of IEEE/IAPR, Professor of Computer Science, Sun Yat-sen University
Embodied AICausal Inference and LearningMultimodal Data Analysis