🤖 AI Summary
This study addresses the failure of optimal trajectory selection in multimodal planning caused by the asymmetry between generation and evaluation, proposing the iDriveVLA framework. This method introduces a unified evaluation mechanism that integrates a safety-aware scorer with a VLM-guided modulator. Through progressive imitation pre-training, candidate space refinement, and semantic ranking alignment techniques, it achieves semantically consistent trajectory ranking and scene-adaptive evaluation. The proposed framework attains state-of-the-art performance on the NAVSIM v1 leaderboard with a PDMS score of 94.95, surpassing human expert reference levels for the first time.
📝 Abstract
Multi-modal planning is promising for autonomous driving by representing multiple plausible behaviors in ambiguous and long-tail scenarios. Existing methods mainly focus on improving trajectory multi-modality, enhancing trajectory representations, or reshaping the candidate distribution. Nevertheless, we identify a pronounced generation-evaluation asymmetry in multi-modal planning: despite strong oracle performance, existing planners often fail to reliably select the best available candidate, leaving substantial planning potential unrealized. To address this challenge, we propose iDriveVLA, a multi-modal planning framework that improves the candidate trajectory space while enabling more reliable and context-aware trajectory evaluation. Specifically, iDriveVLA introduces a unified trajectory evaluator comprising a Safety-aware Scorer for quality and risk estimation, together with a VLM-guided Modulator for scene-adaptive criterion weighting. We further develop an oracle-aligned progressive training strategy consisting of candidate imitation pretraining, candidate space refinement, and semantic ranking alignment. On the public NAVSIM v1 leaderboard, iDriveVLA achieves a new state-of-the-art performance of 94.95 PDMS, surpassing the human-expert reference.