SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation

📅 2026-05-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high computational cost and redundancy inherent in dense scene representations used by existing autonomous driving world models, which hinder efficient prediction and planning. The authors propose SparseWorld, a lightweight world model that leverages a sparse scene representation to model only critical layout elements. It introduces a Sparse Dreamer mechanism that performs autoregressive rollout prediction in latent space via spatiotemporal joint attention, simultaneously generating future map features and agent states. This approach enables end-to-end trajectory planning, achieving a collision rate of just 0.05% on nuScenes—setting a new state of the art in open-loop evaluation—and significantly outperforming existing methods on the closed-loop Bench2Drive benchmark.
📝 Abstract
Recently, world models have made significant progress in enhancing end-to-end driving systems through both future situation forecasting and improved scene understanding. However, existing driving world models are typically built upon dense scene representations, causing high computational costs and redundant information. In this paper, we present SparseWorld, a lightweight world model that focuses on predicting only the critical layout of the scene, enabling efficient future forecasting for end-to-end driving systems. SparseWorld first performs autoregressive rollout to forecast future map elements and surrounding agents, enabling the model to learn how driving scenarios evolve over time. It then leverages these predicted futures to refine downstream motion prediction and trajectory planning. Specifically, we propose a Sparse Dreamer that anticipates future instances in the latent space through joint temporal and spatial attention. By interacting with predicted future instances, the motion planner captures more accurate motion patterns and generates more informed and safety-aware trajectories. Extensive experiments demonstrate that SparseWorld significantly reduces collision risk and achieves state-of-the-art performance on the open-loop planning metrics of the nuScenes dataset with a collision rate of 0.05\%. Moreover, it substantially outperforms the baseline method in closed-loop planning metrics on the Bench2Drive benchmark. Supplementary material is available at the project page: https://wryzju.github.io/SparseWorld/.
Problem

Research questions and friction points this paper is trying to address.

world models
sparse scene representation
end-to-end autonomous driving
computational efficiency
redundant information
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Scene Representation
World Models
Autonomous Driving
Future Prediction
Sparse Dreamer
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Ruoyu Wang
Ruoyu Wang
University of Science and Technology of China
Speech signal processingSpeech recognition
J
Jingke Wang
Institute of Cyber-Systems and Control, Zhejiang University, Hangzhou, China
Yukai Ma
Yukai Ma
Zhejiang University
Y
Yuehao Huang
Institute of Cyber-Systems and Control, Zhejiang University, Hangzhou, China
S
Shuangming Lei
Institute of Cyber-Systems and Control, Zhejiang University, Hangzhou, China
Guanglin Xu
Guanglin Xu
Assistant Professor, Systems Engineering and Engineering Management, UNC Charlotte
Operations ResearchData AnalyticsEnergy SystemsHealth Care
A
Aixue Ye
2012 Labs, Huawei
Yong Liu
Yong Liu
Institute of Cyber-Systems and Control, Zhejiang University
Robotic Vision and PerceptionGraphicsInformation Fusion