When Shared Rollouts Fail in Defensive Driving Evaluation: A NAVSIM Score Basis Audit

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the confounding effect of numerical instability in shared replay trajectories on defensive driving scores, which obscures the performance distinction between environment-aware policies and blind-driving strategies due to error propagation from invalid reference trajectories. For the first time, it systematically reveals how the permissiveness of reference conditions in NAVSIM v2.2 is compromised by such instability. The work proposes a standardized auditing protocol comprising blind probes (Ignore-All and actor-blind), dependency stack control, solver replacement, and replay stability testing. Validation on 12,146 samples demonstrates that blind-driving policies anomalously outperform human replays and PDM-Closed; however, replacing the solver eliminates trajectory divergence and restores a rational score ranking, confirming numerical instability as the root cause.
📝 Abstract
Defensive driving scores are useful only when they preserve distinctions between policies that observe surrounding actors and those that do not. Re-simulation benchmarks may use reference-conditioned forgiveness, under which an agent receives credit when the logged human reference fails a compliance channel. When agent and reference share an unstable rollout transformation, this rule can propagate shared reference failures into broad compliance credit. We audit this risk in NAVSIM v2.2 original scene single-stage scoring. Under the affected documented-stack condition on the audited numerical backend, the route-blind Ignore-All probe and a route-aware actor-blind probe outrank human replay and PDM-Closed over the complete 12,146-token navtest split. A fresh installation following the public specification reproduces rollout divergence on a fixed 32-token diagnostic set. A same-source dependency stack control and an exact-input diagnostic isolate dependency-sensitive numerical behavior in the shared velocity refit. On a 450-token control pool, replacing only the solver eliminates rollout divergence and restores blind-last ordering while keeping forgiveness enabled. Thus, the numerical instability is the direct trigger. Reference-conditioned forgiveness propagates the resulting shared reference failures into compliance credit. We contribute an audit protocol requiring score basis and stack disclosure, blind probes, overwrite reporting, and rollout stability tests before using such scores for defensive driving claims.
Problem

Research questions and friction points this paper is trying to address.

defensive driving evaluation
shared rollouts
reference-conditioned forgiveness
numerical instability
compliance scoring
Innovation

Methods, ideas, or system contributions that make the work stand out.

shared rollout instability
reference-conditioned forgiveness
blind probe auditing
numerical solver sensitivity
defensive driving evaluation
🔎 Similar Papers
No similar papers found.