🤖 AI Summary
This study addresses the inadequacy of existing driving benchmarks in evaluating planner reliability under rare, safety-critical scenarios. To this end, we construct a VLM-based counterfactual planning benchmark that injects hazardous elements into real-world driving scenes to assess planner robustness. We introduce a reference-free protocol and a reminder agent mechanism, enabling structured hazard logging and decision guidance without manual annotations. Benchmark quality is further ensured through multi-view editing, quality auditing, and zero-shot experiments. Experimental results demonstrate that mainstream planners frequently intrude into hazardous regions, whereas the proposed reminder agent significantly improves policy accuracy and reduces false negative rates.
📝 Abstract
Average performance on routine driving benchmarks does not establish planner reliability under rare, safety-critical hazards. We proposed ExceptionDrive, a counterfactual planning benchmark that uses VLM-assisted screening, localized multi-view editing, and quality auditing to insert hazards into real nuScenes scenes while preserving their context. Its 21 tasks span six safety families and define hazard or conflict regions, local safety constraints, and acceptable responses. Because hazard insertion can invalidate the recorded human trajectory, our reference-free protocol evaluates edited predictions using Unsafe Rate (UR), Hazard Clearance Compliance (HCC), Hazard Proximity Response (HPR), and Counterfactual Trajectory Shift (CTS), which measure core-region intrusion, clearance compliance, clearance relative to a prescribed margin, and counterfactual trajectory change. Seven representative planners frequently intrude into hazard regions or provide insufficient clearance. We also develop a Reminder Agent that, without sample-specific task labels, converts visual evidence and the shared taxonomy into structured records of hazard presence, type, and a recommended high-level strategy. The agent neither predicts trajectories nor controls the vehicle; its records guide a VLM-based decision agent. In zero-shot experiments, the reminders improve strategy accuracy and reduce under-warning.