Oneira: From Open-Ended Generation to Open-World Interaction in Video World Models

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the disconnect between open-ended generation and coherent interaction in video world models by proposing an interactive video world model. The method introduces a code agent-based explicit world state management mechanism that closes the loop between generation and interaction through dynamic reading and writing of an extensible state table. By integrating entity-grounded planning with conditional rendering, updated states are translated into videos while a generator supplements fine-grained details. This framework enables the immediate operability of newly generated entities and the persistent retention of interactive states, ensuring that objects discovered during exploration are effectively integrated and that the influence of prior interactions on world evolution endures over long horizons.
📝 Abstract
Generative video world models can now synthesize open-ended environments that agents can navigate and interact with in simple ways. Yet open-ended generation does not imply full interaction: as a generated world expands, newly created content through navigation should expand what the agent can act upon, and as the agent changes the world, those changes should become persistent parts of the environment rather than transient visual effects. We characterize these two requirements as Open-World Interactivity, where newly generated or encountered entities are incorporated into the actionable world, and Persistent State, where interaction outcomes are committed to the world state and continue to influence subsequent observations and interactions. We present Oneira, an interactive video world model that closes the loop between generation and interaction through an explicit, extensible world state managed by a coding agent. Given the current observation and an action or high-level goal, the agent reads the world state, grounds the relevant entities, plans the interaction, and writes its outcome back into a world state table. When exploration reveals new objects, the agent incorporates them from generated observations, allowing the interaction space to expand with the generated world. Meanwhile, previously induced state changes are carried across video segments, making the consequences of interaction persistent parts of subsequent world evolution. The updated world state is rendered along the camera action trajectory into a coarse conditioning video, from which a video generator fills in the appearance, motion, and interaction details not represented in the state. Experiments show that Oneira enables direct and consistent interaction with newly generated objects, while preserving the effects of prior interactions over long horizons. Project page: https://madaoer.github.io/projects/oneira
Problem

Research questions and friction points this paper is trying to address.

Video World Models
Open-World Interactivity
Persistent State
Open-Ended Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video World Models
Open-World Interactivity
Persistent State
Coding Agent
Interactive Generation
🔎 Similar Papers
No similar papers found.