ForeDrive: Foresight-Guided End-to-End Autonomous Driving with a Planning-Relevant Latent World Model

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决自动驾驶中规划相关表示学习问题,提出ForeDrive模型,结合JEPA风格世界模型与Diffusion Transformer,通过门控视觉融合等方法提高规划准确性。
📝 Abstract
Existing latent world models are typically optimized for future predictability, yet the resulting representations are not necessarily useful for planning in autonomous driving. Predictions are commonly used for pretraining or auxiliary supervision rather than as direct conditioning signals for trajectory generation. We propose ForeDrive, which learns a planning-relevant latent representation and couples it asymmetrically to a Diffusion Transformer (DiT) planner. The planner consumes multi-horizon latent future representations learned with a JEPA-style world model; planning gradients update the shared online encoder, while stop-gradient routing trains the latent predictor with forecasting losses only. Because predicted futures have varying reliability across horizons and BEV trajectories are misaligned with image tokens, we use gated visual fusion, future-status injection, and Trajectory-Adaptive Bias (TAB) to inject future latents as guidance without overriding the current observation. Trained with pure imitation learning and using only the current front-view image as visual input at inference, ForeDrive attains 89.9 PDMS on NAVSIM v1 and 90.0 one-stage EPDMS on NAVSIM v2, without reinforcement learning or an external trajectory scorer.
Problem

Research questions and friction points this paper is trying to address.

latent world model
autonomous driving
planning
trajectory generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

planning-relevant latent representation
Diffusion Transformer (DiT) planner
gated visual fusion
future-status injection
Trajectory-Adaptive Bias (TAB)
💼 Related Jobs
No related jobs found.
Sinuo Wang
Sinuo Wang
PhD Candidate, The University of Adelaide
Vision-Language Machine Learning
Z
Zichong Gu
Shanghai Zaofu Intelligent Technology Co., Ltd.
Yuhan Huang
Yuhan Huang
Harbin Institute of Technology
transfer learning diagnostic methods for sparse feature
W
Wenxin Wen
Shanghai Zaofu Intelligent Technology Co., Ltd.
X
Xun Yang
Shanghai Zaofu Intelligent Technology Co., Ltd.
Yiqing Zhang
Yiqing Zhang
Worcester Polytechnic Institute
Xingyu Zhang
Xingyu Zhang
Horizon Robotics Inc
NLP&VLM&AD
N
Ningyu Che
Shanghai Zaofu Intelligent Technology Co., Ltd.
J
Jie Ling
Shanghai Zaofu Intelligent Technology Co., Ltd.
Q
Qiankun Yu
Shanghai Zaofu Intelligent Technology Co., Ltd.
W
Wei Liu
Huazhong University of Science and Technology
Jing Xu
Jing Xu
Hong Kong University of Science and Technology (Guangzhou)
Computer VisionAI applicationRepresentation learning
Xinggang Wang
Xinggang Wang
Professor, Huazhong University of Science and Technology
Artificial IntelligenceComputer VisionAutonomous DrivingObject DetectionObject Segmentation