Training Object Permanence in World Models

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of physical cognitive priors, such as object permanence and entityness, in existing video generation models. We propose WROP, a data infrastructure that leverages Blender's procedural generation to randomize parameters while preserving underlying cognitive structures. Furthermore, we introduce the first cognitive science-inspired dataset and evaluation benchmark comprising 150 tasks, establishing an object permanence training paradigm specifically designed for world models. Built upon a large-scale synthetic data pipeline and an AWS Trainium2-native PyTorch training stack, this project releases a 1.5M-sample corpus and trains the PWM-WROP model. In blind Elo ranking evaluations, the proposed model achieves first place among video continuation models and third place overall.
📝 Abstract
Object permanence and solidity are hallmarks of human cognitive priors. Recent studies show that video generation models, a paradigmatic class of current world models, have begun to show emerged reasoning abilities, making them ideal candidates for building human-like physical intelligence. Do video models have emerged object permanence in them? If not, could we train them with a core-cognition inspired dataset? We introduce WROP (World Reasoning with Object Permanence), a data infrastructure of 150 hand-designed cognitive science inspired tasks, divided into six cognitive categories. We build Blender generators that randomize speed, lighting, camera angle, and other nuisance parameters while preserving each task's cognitive structure, yielding 10,000+ samples per task. We release a 1.5M-sample training corpus and a 300-question exam. On this exam we evaluate 14 video models: 3 reference-to-video, 7 edit, and 4 continuation, among which PWM-WROP, our 16B world model. In a blind pairwise Elo study, PWM-WROP ranks first among continuation models and third overall, behind only a statistical tie between two reference-to-video models. We release the data, exam, model answers, scores, weights, and PWM, our native-PyTorch training stack on AWS Trainium2.
Problem

Research questions and friction points this paper is trying to address.

Object Permanence
World Models
Video Generation Models
Physical Intelligence
Core Cognition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Object Permanence
World Models
Video Generation
Cognitive Science
Data Infrastructure
H
Haotian Zhang
University of Southern California
F
Fengyuan Yu
Carnegie Mellon University
Dezhi Luo
Dezhi Luo
University of Michigan
cognitive sciencephilosophyAI
H
Haoran Sun
Johns Hopkins University
Z
Zehong Zhao
University of California, San Diego
Q
Qingying Gao
Johns Hopkins University
Y
Yihan Li
Carnegie Mellon University
S
Siyuan An
Carnegie Mellon University
H
Huayi Qin
Carnegie Mellon University
Yilan Zhang
Yilan Zhang
King Abdullah University of Science and Technology
Computer VisionMedical Image Analysis
Z
Zhengze Jiang
Columbia University
P
Pinyuan Feng
Columbia University
Renrui Zhang
Renrui Zhang
Seed ByteDance & MMLab & PKU
Large Multimodal ModelGenerative ModelEmbodied AI
Ziyu Guo
Ziyu Guo
The Chinese University of Hong Kong
Multi-modality LearningLLM/VLMs3D Vision
Letian Wang
Letian Wang
University of Toronto | Carnegie Mellon University | UC Berkeley
MultiModal LearningReinforcement Learning3D/4D VisionHuman Robot Interaction
Mengyue Yang
Mengyue Yang
Lecturer, University of Bristol
CausalityTrustworthiness
Kangfu Mei
Kangfu Mei
Google DeepMind
Computational PhotographyGenerative Models
Maijunxian Wang
Maijunxian Wang
University of California, Berkeley
Social JusticeArtificial General IntelligenceAI AlignmentAI EthicsMachine Cognition
R
Ran Ji
University of California, San Diego
V
Vikash Kumar
Columbia University
Freda Shi
Freda Shi
Assistant Professor of Computer Science, University of Waterloo
Language GroundingMultilingualismComputational LinguisticsArtificial Intelligence
Chandra Sripada
Chandra Sripada
Professor, Psychiatry and Philosophy, University of Michigan
PsychiatryPhilosophySelf-ControlADHDNeuroimaging
V
Vincent C. Muller
Friedrich-Alexander-Universität Erlangen
Philip Torr
Philip Torr
Professor, University of Oxford
Department of Engineering
Alan Yuille
Alan Yuille
Professor of Cognitive Science and Computer Science, Johns Hopkins University
Computer VisionComputational Models of Mind and BrainMachine Learning