RefineDrive: Reliable Failure-Guided Learning for Vision-Language-Action Driving

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of Vision-Language-Action (VLA) driving models, which typically rely on successful demonstrations, neglect failure cases, and suffer from unreliable diagnostics. To overcome these issues, we propose a fault-guided post-training framework. Methodologically, we introduce a structured and verifiable fault diagnosis mechanism, combined with minimal-correction retrieval that preserves motion patterns to generate supervision signals. Furthermore, we design a hierarchical reward function prioritizing hard safety constraints, achieving policy optimization through GRPO-based reinforcement learning and conditional supervised fine-tuning (SFT). Evaluated on the NAVSIM benchmarks, our approach improves the PDMS score from 87.7 to 91.7 on v1 and achieves 89.4 EPMS on v2, significantly enhancing both the safety and planning capabilities of autonomous driving systems.
📝 Abstract
Vision-Language-Action (VLA) models for autonomous driving rely heavily on successful expert demonstrations, leaving model-specific failures underexploited. Learning from these failures is hindered by unreliable diagnoses, poorly matched correction targets, and coarse rewards. We propose RefineDrive, a failure-guided post-training framework that learns from self-generated failures through targeted supervision and safety-aware reinforcement learning. Reliable Diagnosis derives structured, verifiable feedback on collisions and drivable-area violations directly from simulator states. Minimum-Correction Target Retrieval searches a clustered human trajectory bank for nearby corrections that satisfy hard-safety constraints in the current scene, prioritizing preservation of the failed prediction's motion pattern. Conditioned on the driving context and failed trajectory, Correction SFT learns to generate the diagnosis followed by the retrieved correction as a training-only auxiliary task. We then apply GRPO with a Safety-Layered Reward that strictly prioritizes hard-safe trajectories, retains continuous safety feedback for both unsafe and hard-safe trajectories, and rewards driving progress only after hard safety is satisfied. At inference, the policy directly predicts trajectories from the driving context without an explicit diagnosis or repair stage. On NAVSIM v1, RefineDrive improves the 4B base SFT policy from 87.7 to 91.7 PDMS. Using the same checkpoint without additional training, RefineDrive achieves 89.4 EPDMS on the original NAVTEST scenes evaluated with NAVSIM v2 extended metrics. Controlled ablations support the benefits of structured diagnosis supervision, retrieved corrections, and safety-layered optimization for direct planning.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
autonomous driving
failure-guided learning
reliable diagnosis
safety-aware reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action
Failure-Guided Learning
Safety-Layered Reward
Minimum-Correction Retrieval
Autonomous Driving
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Zhe Sun
Zhe Sun
Swinburne University of Technology
Nonlinear ControlSliding Mode ControlMechatronics
Z
Ziyi Luo
College of Computer Science and Technology, Zhejiang University, Hangzhou, China
Yehao Lu
Yehao Lu
Zhejiang University
Autonomous Driving3D ReconstructionSwarm Robot
L
Lei Zhou
Yinwang Intelligent Technology Co., Ltd.
X
Xi Li
College of Computer Science and Technology, Zhejiang University, Hangzhou, China