🤖 AI Summary
This study addresses the inherent difficulty of existing camera-controllable video generation models in maintaining four-dimensional spatiotemporal consistency. To overcome this limitation, we propose an “Observe–State–Reflect” framework that enforces multi-view constraints through a spatiotemporal epipolar causal attention mechanism. Furthermore, the method incorporates a geometry reflection pipeline driven by 4D retrieval and reconstruction to achieve dynamic self-correction. By integrating these components, our approach attains state-of-the-art performance in free-viewpoint 4D scene generation, delivering high-fidelity geometric accuracy, strong generalization capabilities, and cinematic visual quality.
📝 Abstract
While existing camera-controllable video generation models can produce visually compelling sequences, preserving intrinsic 4D spatiotemporal coherence remains challenging. To address this limitation, we propose ChronoWorld, an"Observation--State--Reflection"framework that leverages spatiotemporal causal cues and reconstruction priors to generate globally consistent, free-view 4D scenes. Given a context video, we introduce a Spatiotemporal Epipolar Causal Attention mechanism that enforces multi-view epipolar constraints and temporal causality throughout the generation process. In addition, we develop a reconstruction-driven geometric reflection pipeline with a 4D retrieval strategy to enable dynamic self-assessment and correction of generated outputs, improving consistency and accuracy. Extensive experiments show that ChronoWorld achieves state-of-the-art performance in spatiotemporally consistent, cinematic-quality 4D scene generation, with strong generalization and high-fidelity geometry across diverse scenarios.