Xiaomi EV World Model: A Joint World Model Integrating Reconstruction and Generation for Autonomous Driving

πŸ“… Unknown Date
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing world models for autonomous driving struggle to jointly optimize high-fidelity 3D scene reconstruction and temporally coherent video generation. This work proposes a Joint World Model (JWM) that integrates a feedforward 3D Gaussian reconstruction module, WorldRec, driven by structured 3D sparse queries, with a causal video generation module, WorldGen, which employs bidirectional pretraining followed by a three-stage causal fine-tuning strategy incorporating Teacher Forcing, ODE distillation, and DMDβ€”enabling efficient video synthesis in just four denoising steps. By deeply fusing both components within a unified representation space, JWM significantly enhances cross-frame consistency and generation stability, achieving for the first time simultaneous high-fidelity, spatiotemporally coherent online 3D reconstruction and video generation, thereby establishing a new paradigm for closed-loop simulation and end-to-end autonomous driving training.
πŸ“ Abstract
This report presents a unified technical system addressing the two core capabilities of world models for autonomous driving: world representation and world generation. For world representation, we propose WorldRec, a feed-forward reconstruction architecture driven by sparse scene queries. WorldRec initializes structured queries in 3D space, leveraging them to aggregate cross-view, cross-temporal features, thereby naturally enforcing spatial consistency across frames and yielding compact yet high-fidelity 3D Gaussian scene representations. For world generation, we propose WorldGen, a two-stage training framework of bidirectional pretraining followed by causal fine-tuning through three progressive stages (Teacher Forcing, ODE distillation, and DMD), enabling high-quality online causal video generation in as few as 4 denoising steps. Building on both modules, we further introduce the JWM, which deeply integrates WorldRec and WorldGen to achieve synergistic gains in generation stability, cross-frame consistency, and visual fidelity, providing a solid foundation for closed-loop simulation, data synthesis, and end-to-end training in autonomous driving.
Problem

Research questions and friction points this paper is trying to address.

world model
autonomous driving
3D scene representation
causal video generation
spatiotemporal consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Model
3D Gaussian Reconstruction
Causal Video Generation
Autonomous Driving
Joint Modeling