🤖 AI Summary
This work addresses the limitation of existing world models in maintaining geometric consistency during scene expansion and cross-view revisitation due to the absence of explicit, persistent structural memory. We propose a novel framework that decouples world state maintenance from visual rendering to construct scalable 3D worlds for video generation. Specifically, our approach integrates elevation map generation, hierarchical semantic planning, diffusion-based image outpainting, agent-guided scene stitching, and depth-sequence-guided video synthesis. By establishing persistent geometric memory independent of short video generation, the method provides a consistent structural foundation for cross-trajectory navigation and repeated scene revisitation. Consequently, it enables unbounded maps to be incrementally expanded without compromising previously established memories, advancing the scalability and coherence of generative world models.
📝 Abstract
Despite rapid progress, world models still lack explicit, persistent structural memory, making it difficult to preserve consistent world structure during continual scene expansion and cross-view revisits. To address this limitation, we present WorldWeave, a world generation framework that decouples world-state maintenance from visual rendering. Specifically, WorldWeave combines continual elevation-map generation with agent-guided scene organization and stitching to build an expandable explicit 3D world that incrementally extends structural memory while preserving existing structure. First, its terrain module uses diffusion-based image outpainting to generate continuous metric elevation maps under neighborhood conditioning and boundary constraints. Next, an agent integrates user intent, terrain evidence, and cross-region connectivity constraints to construct scenes through hierarchical semantic planning, deterministic geometry compilation, and local revision. Finally, during visual generation, planned camera trajectories query world geometry through a read-only interface, producing depth sequences that guide video synthesis without writing the generated results back into the world state. As a result, structural memory remains independent of short-window video generation, enabling continual expansion without predefined map boundaries and providing a consistent geometric basis for observations across trajectories and repeated visits.