Kepler4D: Controllable Future Video Generation via 4D Scene State Evolution

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of preserving scene structure and predicting dynamic object evolution in video world models by proposing a video generation framework grounded in explicit 4D scene state evolution. Methodologically, it constructs a shared 3D representation that decouples motion from appearance synthesis. A novel Chain-of-Motion mechanism is introduced, leveraging vision-language models to enable editable, deterministic trajectory inference. The technical pipeline integrates monocular 3D reconstruction, multimodal large model decision-making, and pretrained video rendering. Experimental results demonstrate that the proposed approach achieves controllable object motion, plausible future prediction, and high scene consistency on real-world videos.
📝 Abstract
Video world models aim to preserve scene structure and predict how dynamic objects evolve beyond visual observations. We present Kepler4D, a framework for future video generation through explicit 4D scene state evolution. Given a monocular video, Kepler4D constructs a shared 3D representation of background geometry, object motion histories, coarse spatial supports, and semantic context. Chain-of-Motion summarizes observed motion and uses a vision-language model to select structured speed and heading decisions and decide whether to bound object-center height from below. A deterministic rollout converts these decisions into future object trajectories for inspection and editing before synthesis. We render the evolving proxies into geometric controls for a pretrained video generator, separating coarse object motion from the synthesis of appearance and articulation. Experiments on real-world videos demonstrate that Kepler4D enables controllable object motion and plausible future rollout while preserving scene consistency.
Problem

Research questions and friction points this paper is trying to address.

video world models
future video generation
4D scene evolution
controllable object motion
scene consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

4D Scene State Evolution
Controllable Video Generation
Chain-of-Motion
Vision-Language Model
Geometric Control