🤖 AI Summary
This study addresses the challenge in autonomous driving where unexecuted trajectories lack supervision, impeding scorers from accurately comparing candidate plans. To overcome this, we propose a JEPA-based trajectory-conditioned world model that jointly anchors candidate states using simulator labels and real future observations to generate reliable scores. Furthermore, a scene-matched low-score sample bank and an inertia-based re-ranking mechanism are introduced to enforce parameter-sharing constraints and ensure consistency in sequential decision-making. The proposed method achieves state-of-the-art performance on the NAVSIM-v2 and Bench2Drive benchmarks, while demonstrating effective generalization to robotic manipulation planning tasks.
📝 Abstract
Autonomous driving requires choosing a safe and efficient plan as surrounding traffic evolves. Generate-and-select planners propose multiple trajectories and score them for execution, and they have outperformed representative direct-prediction baselines on NAVSIM. Their scorer must compare plans that were never executed. Driving logs record the future of only the executed trajectory, so matching the logged future can leave predictions for the alternatives unconstrained; a simulator, in contrast, can label the outcome of every candidate. We introduce World4Scorer, which builds the scorer as a trajectory-conditioned JEPA-style predictor: it predicts a state for each candidate and reads the candidate's scores from that state. Simulator outcome labels supervise the states of all candidates, and the observed future of the executed trajectory anchors the predictor to real scene evolution. Because one predictor produces every candidate's state, the anchor can constrain shared parameters used to score unexecuted plans, while the future itself is needed only during training. Generated candidates mostly score well, so a scene-matched bank adds low-scoring plans to the outcome supervision; framewise choices can conflict, so inertial re-ranking keeps consecutive selections consistent. World4Scorer achieves state-of-the-art NAVSIM-v2 performance and a strong adapted-system result on closed-loop Bench2Drive. With the LeWM world model and planning budget fixed, outcome-based scoring also improves manipulation planning on the OGBench-Cube benchmark.