🤖 AI Summary
This study addresses the causal confusion induced by imitation learning in end-to-end autonomous driving and the action-visual misalignment prevalent in existing world models. To this end, it proposes a plug-and-play closed-loop reinforcement learning framework. The core innovation lies in designing an action-visual fidelity evaluator that, combined with geometry-aware auxiliary supervision, filters for reliable trajectories. By exclusively leveraging high-fidelity data for long-horizon closed-loop reinforcement learning post-training, the framework achieves safe policy optimization. Evaluated on benchmarks such as nuScenes, the proposed method reduces the infraction rates of DiffusionDrive and Qwen3-VL by 27.6% and 33.7%, respectively, demonstrating substantial improvements in driving safety.
📝 Abstract
End-to-end autonomous driving policies are commonly trained via imitation learning on logged demonstrations without observing the consequences of their own actions, leading to causal confusion in closed-loop real-world deployment. To address this issue, reinforcement learning (RL) post-training offers a promising alternative by leveraging world models as interactive training environments to enable future scene generation for policy improvement. Nevertheless, existing approaches either rely on reconstruction-based simulators, offering limited counterfactual interaction, or adopt synthetic simulators to enable long-horizon closed-loop interaction at the cost of a substantial sim-to-real gap. Recently, video world models have exhibited the ability to generate realistic multi-step future rollouts but may not faithfully reflect action conditions, resulting in action-vision mismatch. In this paper, we introduce RoXDrive, a plug-and-play closed-loop RL framework that enables reliable policy optimization by identifying action-faithful world-model rollouts, consisting of two stages: 1) Model pre-training: In addition to imitation-based policy pre-training, we devise an Action-Vision Faithfulness Evaluator for inverse dynamics estimation with our geometry-aware auxiliary trajectory supervision, enabling long-horizon assessment of whether visual dynamics faithfully reflect the conditioning ego actions. 2) Action-faithful RL post-training: Agents iteratively interact with world models to form long-horizon scene rollouts, retaining only action-faithful ones for dense safety-aware scoring and scene-level closed-loop RL post-training. Extensive experiments on nuScenes and an in-house dataset with over 130K training scenarios demonstrate consistent gains across planners, reducing safety violations by 27.6% with DiffusionDrive on nuScenes and 33.7% with Qwen3-VL on the internal data.