WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
WorldCrafter通过学习一种可由摄像机查询的隐式3D感知记忆,解决了视频世界模型在长时间和跨视角下难以保持与先前观察一致的问题。
📝 Abstract
Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory for this purpose. The key insight is to let the requested viewpoint shape how multi-view evidence is compressed into the video generator's limited token budget. Trained jointly with the video generator, a memory encoder and pose-conditioned readout module integrate historical observations into a fixed set of target view-specific tokens before denoising, without explicit depth-based correspondences. By combining this memory with recent temporal context and few-step distillation, WorldCrafter enables streaming scene exploration from a single input image or text prompt. Experiments across static and dynamic scenes show substantial gains in long-horizon consistency and camera-control accuracy while preserving visual quality during minute-scale exploration.
Problem

Research questions and friction points this paper is trying to address.

video world model
dynamic environments
long horizons
viewpoints
consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

implicit 3D-aware memory
camera-queryable
viewpoint-specific tokens
temporal context
few-step distillation
🔎 Similar Papers