WM-Cov: Test Adequacy for Interactive World-Model-Style Autonomous Driving Simulation

๐Ÿ“… 2026-07-31
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the lack of effective adequacy metrics in existing world-model-based interactive simulation testing for autonomous driving, which hinders the assessment of whether generated hazardous scenarios constitute valid, non-redundant closed-loop test evidence. To this end, the paper introduces WM-Covโ€”an evaluation layer agnostic to simulation platformsโ€”that formally defines, for the first time, adequacy criteria for interactive world-model testing. It establishes a multidimensional evaluation framework encompassing effective failure discovery, diversity of failure modes, and realism, leveraging techniques such as coverage growth analysis, extraction of valid evidence, redundancy detection, and artifact suppression. Experiments on platforms like DriveArena demonstrate that the method identifies 304 high-quality, valid evidence instances from hundreds of simulations, supporting the key insight that test adequacy should be judged by the convergence of effective interactive evidence rather than raw failure counts.
๐Ÿ“ Abstract
World models and generative simulators are emerging as interactive testing infrastructure for autonomous driving because they can react to the ego planner and produce counterfactual, rare, and safety-critical rollouts. This changes a test scenario from a fixed replayed trajectory into an interactive scenario family whose realized evolution depends on the planner under test. The unresolved question is therefore not only whether dangerous rollouts can be generated, but what valid closed-loop evidence is enough to support a specified testing intent and stopping decision. This paper formulates interactive world-model-style testing adequacy and introduces WM-Cov, a provider-agnostic evaluation layer that converts raw provider outputs into requested, realized, and valid evidence. WM-Cov reports adequacy through coverage growth, valid-failure discovery, failure-mode diversity, realism, artifact suppression, duplicate accounting, and valid-evidence precision. Studies on executed TeraSim/SUMO events, WM-like mixed trace pools, and a real DriveArena TrafficManager--WorldDreamer matrix show that dangerous-looking events can include valid ADS failures, duplicates, partial realizations, and artifacts. The DriveArena matrix evaluates two planners, two horizons, six prompt conditions, and 360 ego-route requests; 304 attempts become fully realized evidence and 56 remain partial. A disjoint 80-request route-slice check yields 74 fully realized and 6 partial attempts. The results support evaluating world-model-style testing by convergence of valid interactive evidence under budget, rather than by raw generated failures or prompt coverage alone.
Problem

Research questions and friction points this paper is trying to address.

test adequacy
world models
autonomous driving simulation
interactive testing
evidence validity
Innovation

Methods, ideas, or system contributions that make the work stand out.

world models
test adequacy
interactive simulation
WM-Cov
autonomous driving
๐Ÿ”Ž Similar Papers
No similar papers found.