Learning World Models for Interactive Video Generation

📅 2025-05-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Current long-video generation models suffer from compounding errors and weak memory mechanisms, hindering the construction of foundational world models that simultaneously ensure interactivity and spatiotemporal consistency. To address this, we propose Video Retrieval-Augmented Generation (VRAG), the first framework to introduce explicit global state conditioning—overcoming the inherent, irreducible error accumulation bottleneck in autoregressive modeling. VRAG integrates action-conditioned generation with efficient video retrieval to substantially suppress long-horizon generation errors. Furthermore, we establish the first comprehensive benchmark specifically designed to evaluate world-modeling capabilities in video generation. Extensive experiments demonstrate significant improvements in spatiotemporal consistency, interactive controllability, and long-sequence fidelity. Our approach establishes a novel paradigm for interactive, video-based world models.

Technology Category

Computer Vision: Image and Video RetrievalMachine Learning: Deep Generative Models & AutoencodersNatural Language Processing: Generation

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
Foundational world models must be both interactive and preserve spatiotemporal coherence for effective future planning with action choices. However, present models for long video generation have limited inherent world modeling capabilities due to two main challenges: compounding errors and insufficient memory mechanisms. We enhance image-to-video models with interactive capabilities through additional action conditioning and autoregressive framework, and reveal that compounding error is inherently irreducible in autoregressive video generation, while insufficient memory mechanism leads to incoherence of world models. We propose video retrieval augmented generation (VRAG) with explicit global state conditioning, which significantly reduces long-term compounding errors and increases spatiotemporal consistency of world models. In contrast, naive autoregressive generation with extended context windows and retrieval-augmented generation prove less effective for video generation, primarily due to the limited in-context learning capabilities of current video models. Our work illuminates the fundamental challenges in video world models and establishes a comprehensive benchmark for improving video generation models with internal world modeling capabilities.
Problem

Research questions and friction points this paper is trying to address.

Enhancing video models with interactive action conditioning
Addressing compounding errors in autoregressive video generation
Improving spatiotemporal coherence with retrieval-augmented generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Action conditioning enhances interactive video generation
VRAG reduces errors with global state conditioning
Autoregressive framework improves spatiotemporal coherence
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Taiye Chen
School of EECS, Peking University
X
Xun Hu
Department of Engineering Science, University of Oxford
Z
Zihan Ding
Department of Electrical and Computer Engineering, Princeton University
Chi Jin
Chi Jin
Assistant Professor, Princeton University
Machine LearningOptimization