π€ AI Summary
Existing world models for autonomous driving struggle to jointly optimize high-fidelity 3D scene reconstruction and temporally coherent video generation. This work proposes a Joint World Model (JWM) that integrates a feedforward 3D Gaussian reconstruction module, WorldRec, driven by structured 3D sparse queries, with a causal video generation module, WorldGen, which employs bidirectional pretraining followed by a three-stage causal fine-tuning strategy incorporating Teacher Forcing, ODE distillation, and DMDβenabling efficient video synthesis in just four denoising steps. By deeply fusing both components within a unified representation space, JWM significantly enhances cross-frame consistency and generation stability, achieving for the first time simultaneous high-fidelity, spatiotemporally coherent online 3D reconstruction and video generation, thereby establishing a new paradigm for closed-loop simulation and end-to-end autonomous driving training.
π Abstract
This report presents a unified technical system addressing the two core capabilities of world models for autonomous driving: world representation and world generation. For world representation, we propose WorldRec, a feed-forward reconstruction architecture driven by sparse scene queries. WorldRec initializes structured queries in 3D space, leveraging them to aggregate cross-view, cross-temporal features, thereby naturally enforcing spatial consistency across frames and yielding compact yet high-fidelity 3D Gaussian scene representations. For world generation, we propose WorldGen, a two-stage training framework of bidirectional pretraining followed by causal fine-tuning through three progressive stages (Teacher Forcing, ODE distillation, and DMD), enabling high-quality online causal video generation in as few as 4 denoising steps. Building on both modules, we further introduce the JWM, which deeply integrates WorldRec and WorldGen to achieve synergistic gains in generation stability, cross-frame consistency, and visual fidelity, providing a solid foundation for closed-loop simulation, data synthesis, and end-to-end training in autonomous driving.