🤖 AI Summary
This study addresses the challenges of low computational efficiency, inadequate safety, and limited preference annotation in end-to-end planning for autonomous driving by proposing EMPlan. This method introduces a novel hybrid architecture combining sparse anchors with offset refinement to achieve efficient multimodal trajectory prediction. Furthermore, it adopts a two-stage paradigm comprising pre-training followed by reward-guided fine-tuning, integrated with a rule-based unpaired preference supervision strategy that enhances safety without incurring additional inference overhead. Experimental results demonstrate that EMPlan achieves an excellent balance between planning accuracy and computational efficiency under real-time constraints on the NAVSIM benchmark, significantly outperforming existing baselines.
📝 Abstract
Safe and efficient trajectory planning is essential in autonomous driving. However, existing end-to-end approaches often fall short in both computational efficiency and safety guarantees. Methods based on imitation learning suffer from causal confusion, while rule-based scoring approaches often incur heavy computational overhead and suffer from objective misalignment. Additionally, preference-based methods rely on strict pairwise annotations, limiting data utilization. To overcome these limitations, we propose EMPlan, an efficient multi-modal trajectory planning method powered by reward-guided fine-tuning. We design a hybrid architecture that combines sparse anchors with an offset refinement module for efficient multi-modal trajectory prediction. Sparse anchors provide coarse trajectory candidates with low latency, which are subsequently refined by the offset module for higher prediction accuracy. To enhance safety without incurring additional inference costs, we adopt a two-stage training paradigm consisting of pretraining and reward-guided fine-tuning. During fine-tuning, we leverage rule-based reward signals and unpaired preference supervision to refine the pretrained policy toward safer trajectory selection. We evaluate EMPlan on the non-reactive NAVSIM benchmark, where it strikes a favorable balance between planning accuracy and efficiency, demonstrating superior performance under real-time constraints.