ReDrive: Shaping Representations with World Modeling for End-to-End Driving

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the architectural complexity of existing end-to-end driving models and their reliance on auxiliary modules during inference. To overcome these limitations, this work proposes a framework that leverages world models to shape visual representations and reinforce planning-oriented features. A three-stage training pipeline—comprising video pre-training, joint modeling with planning, and planner adaptation—is introduced to demonstrate that robust visual representations combined with world models suffice for effective planning, thereby eliminating the need for auxiliary perception and prediction modules at inference time. Evaluated on the NAVSIM v1 and v2 benchmarks, the proposed method achieves 91.0 PDMS and 90.8 EPDMS, respectively. These results establish that a minimalist encoder-planner architecture can deliver high-performance end-to-end autonomous driving.
📝 Abstract
Driving policies require capabilities of scene understanding and future evolution prediction. To achieve this goal, current end-to-end models typically construct complex perception-planning pipelines or introduce world models that explicitly predict future states, resulting in a complex system architecture. Inspired by the transferability of general-purpose visual representations, we argue that combining sufficiently strong visual representations with representation world modeling can support effective planning without relying on complex inference-time auxiliary modules. Based on this insight, we present ReDrive, an end-to-end driving framework that strengthens planning-oriented visual features via future representation prediction. To achieve this, ReDrive adopts a three-stage training pipeline consisting of driving video pretraining, joint world-modeling and planning training, and planner adaptation. This yields a strong planning-oriented representation and a high-performance planner, while requiring neither auxiliary perception modules nor future prediction at inference time. Experiments on NAVSIM demonstrate strong performance, achieving 91.0 PDMS on NAVSIM v1 and 90.8 EPDMS on NAVSIM v2. These results show that shaping representations with world modeling is sufficient to enable high-performance end-to-end planning while retaining a simple encoder-planner inference pipeline.
Problem

Research questions and friction points this paper is trying to address.

End-to-End Driving
World Modeling
Representation Learning
Autonomous Planning
Innovation

Methods, ideas, or system contributions that make the work stand out.

End-to-End Driving
World Modeling
Representation Learning
Future Representation Prediction
NAVSIM
🔎 Similar Papers
No similar papers found.