Why Do Conventional World Models Fail to Learn Cellular Automata?

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of conventional world models in learning precise dynamics from observational histories, particularly their deficiencies in spatial locality, temporal consistency, and stability. Using cellular automata as a testbed, this work proposes a lightweight improvement strategy that modulates information flow rather than restructuring backbone networks. By introducing two-dimensional rotary position embeddings, neighborhood token mechanisms, and causal freezing techniques, the proposed approach effectively resolves structural bottlenecks in Transformers, CNNs, and diffusion models. Experimental results demonstrate that on tasks such as Conway's Game of Life, this method increases the success rate of complete trajectory rollouts from below 20% to nearly 100%, while significantly enhancing generalization to unseen rules.
📝 Abstract
Although conventional world models - auto-regressive or diffusion models based on transformers or convolutional networks - may learn surface statistics of world dynamics, can they learn the exact world dynamics from its observed history? Leveraging cellular automata as a simple testbed, we find the answer to be no in many cases. Conventional architectures predict most pixels correctly yet rarely complete a rollout: a CNN predicts 96.3% of cells but completes 18.9% of rollouts; a joint diffusion model completes none. We trace the gap to three failure modes of these world models - namely, they fail to exactly capture spatial locality, temporal locality or temporal stability. Simple changes repair each: (1) for spatial locality, two-dimensional rotary positions lift a transformer from 39.1% to 100% on the Game of Life; (2) for temporal locality, handing each token its cell's previous-frame neighbourhood lifts the same transformer from 25.8% to 99.9% on unseen rules; (3) for temporal stability, causal freezing lifts the same diffusion weights from 42.2% to 99.9%. None of the three changes touches the architectural backbone; each only modifies the information flow within it. We also compare joint and ordered sampling on billiards and, in an exploratory study, on a simulated Burgers equation.
Problem

Research questions and friction points this paper is trying to address.

world models
cellular automata
spatial locality
temporal locality
temporal stability
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Models
Cellular Automata
2D Rotary Position Encoding
Causal Freezing
Information Flow
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Shaoyang Guo
Shaoyang Guo
Peking University
PhysicsAI
Z
Ziming Liu
1Meta Circle(ᇆି) 3Tsinghua University