Score
Designs and implements procedural scenario generators that produce safety-critical test environments by defining parameterized environment templates, stochastic processes, and constraints. Builds controls for scenario difficulty, variability, and diversity and analyzes generated scenario distributions to ensure coverage, systematic evaluation, and appropriate challenge levels.
Safety verification of decision-making agents in dynamic environments faces challenges including susceptibility to local optima in high-dimensional scenario spaces and difficulty balancing scenario diversity with criticality. Method: This paper proposes a dual-space guided testing framework that jointly optimizes the scenario parameter space and agent behavioral space. It introduces a novel parameter–behavior closed-loop feedback mechanism, integrating hierarchical representation, dimensionality reduction modeling, multi-dimensional subspace evaluation, behavioral criticality quantification, and adaptive mode switching to dynamically balance local perturbation and global exploration. Results: Experiments on five decision-making agents show that the framework increases critical scenario generation by 56.23% on average. It significantly outperforms state-of-the-art methods under a joint parameter–behavior driving metric, achieving superior scenario diversity, coverage, and verification effectiveness.
To address the challenge that existing autonomous driving simulation validation tools require programming expertise—thus hindering adoption by non-technical users—this paper proposes a no-code, interactive scenario generation framework. Methodologically, it introduces a graph-structured scenario representation model that unifies manual editing with parameterized stochastic sampling, enabling intuitive construction of diverse, high-fidelity traffic scenarios; a graphical user interface supports configuration, editing, and execution, with seamless integration into the CARLA simulator and deep learning models. Key contributions include: (1) the first application of graph-based modeling to no-code scenario generation, enhancing semantic expressiveness and manageability; (2) substantial reduction in usability barriers, empowering domain experts without coding skills to efficiently construct large-scale, variable, and photorealistic test cases; and (3) empirical evaluation demonstrating superior performance over state-of-the-art tools in scenario diversity, generation efficiency, and simulation utility.
This work addresses the inefficiency in autonomous driving testing caused by cumbersome workflows and redundant code during complex scenario construction. To overcome these limitations, the authors propose Modular2Simple, a novel tool that introduces a modular-composition paradigm for scenario generation. By combining simple or modular OpenSCENARIO scenes, Modular2Simple enables flexible and efficient creation of diverse, complex test scenarios while strictly adhering to the OpenSCENARIO standard. The approach seamlessly integrates with mainstream simulation platforms such as CARLA, significantly enhancing scenario reusability and customizability while reducing development complexity. Experimental results demonstrate that, compared to conventional methods, the proposed solution substantially decreases both development time and labor costs, markedly improving the efficiency and diversity of test scenario construction.
This work addresses the limitations of existing autonomous driving testing methods, which rely on handcrafted templates or fixed models and struggle to efficiently reproduce complex real-world failure scenarios. The authors propose a modular synthesis framework powered by large language models (LLMs) that, for the first time, incorporates natural language descriptions and contextual information from real accident reports into the scenario generation process. Operating under predefined testing constraints, the framework automatically constructs diverse driving scenarios. Integrated with the MetaDrive platform and validated using NHTSA crash data, the approach successfully generates test cases covering four road types, three non-ego vehicle motion patterns, and construction-zone anomalies. Remarkably, only 20 generated scenarios effectively expose latent system failures, significantly outperforming conventional testing methodologies.
Autonomous driving systems (ADS) face a dual challenge in testing: naturally occurring traffic scenarios often lack sufficient risk, while synthetically generated critical scenarios suffer from low fidelity. This paper proposes the On-Demand Scenario Generation (OSG) framework, enabling controllable synthesis of diverse, naturalistic, and safety-critical traffic scenarios across a continuous risk spectrum. Our approach integrates real-world traffic data, multi-objective optimization, and heuristic search within a collaborative generation mechanism. Key innovations include: (i) the first quantifiable, tunable risk intensity controller; (ii) a synergistic generation pipeline unifying data-driven modeling, optimization, and search; and (iii) an end-to-end CARLA-based simulation infrastructure. Experiments demonstrate robust scenario generation across risk levels, uncover distinct ADS failure patterns under progressive risk exposure, and establish a new paradigm for objective, comparable, and systematic ADS safety and reliability evaluation.
This work addresses the challenge that traditional model-based testing is ill-suited for distributed robotic systems due to their high nondeterminism, dynamic reconfiguration, and inherent complexity. To overcome this limitation, the paper proposes the Scenario Specification Language (SCSL), which enables the construction of system-level tests by composing basic scenarios. The approach integrates runtime online test generation and execution with mechanisms for dynamic component joining/leaving and interface reconnection, thereby supporting automated testing and dynamic reconfiguration. The syntax and semantics of SCSL are validated through a robotic salvage mission case study, where automatically generated tests effectively demonstrate the feasibility and advantages of the proposed method.
This work addresses the limitations of existing approaches in autonomous driving simulation testing, where manually authored scenarios lack scalability and statistical models struggle to precisely control interactive behaviors in out-of-distribution settings. The authors propose modeling traffic scenario orchestration as a constraint satisfaction problem, uniquely integrating large language models (LLMs) with closed-loop simulation. Specifically, natural language descriptions are automatically translated by an LLM into formal constraints, which are then processed by off-the-shelf solvers to generate closed-loop scenarios aligned with test intents. Evaluated on diverse benchmarks, the method significantly outperforms baseline approaches, achieving markedly higher scenario generation success rates. The results underscore the critical role of closed-loop mechanisms in handling ego-vehicle-responsive complex scenarios and demonstrate a unified framework that reconciles high controllability with scalability.
Existing autonomous driving simulation frameworks lack native support for OpenSCENARIO 2.x, leading to spatiotemporal drift, event latency, and motion discontinuities during scenario execution. This work proposes the first CARLA-based simulation orchestration framework with native OpenSCENARIO 2.x support. It employs a multi-pass translator to compile the domain-specific language into a type-safe abstract syntax tree and dynamically generates deterministic behavior trees that invoke CARLA’s atomic APIs. This approach achieves, for the first time, an exact mapping from OpenSCENARIO 2.x to CARLA, elevating scenario testing from approximate interpretation to a mathematically rigorous and deterministic execution paradigm. The framework ensures frame-accurate determinism, precise spatial trigger evaluation, and 100-millisecond cross-agent blackboard synchronization under high-concurrency adversarial conditions, strictly adhering to continuous environmental boundary constraints.
This study addresses the limitations of current robotic system validation, which relies heavily on manual selection of test scenarios, thereby hindering scalability and compromising reproducibility and reliability of conclusions. To overcome these challenges, this work proposes a compositional, scenario-based modeling approach that integrates declarative test specifications, plugin-driven scenario generation, containerized parallel simulation, and unified result analysis to establish the first modular and scalable automated verification framework. The framework enables systematic parameter variation across multiple dimensions and facilitates robust identification of systemic faults versus stochastic anomalies. Evaluated across 5,480 distinct scenario configurations with over 100,000 simulation runs, the approach accumulated 1,800 hours of simulated operation and 1,873 virtual kilometers, demonstrating its efficacy in discerning consistent system deficiencies from random irregularities.
Existing automated test generation approaches heavily rely on code coverage metrics, often failing to capture diverse test scenarios driven by implicit requirements and thereby risking the omission of critical defects. To address this limitation, this work proposes TestGeneralizer, a novel framework that treats an initial test as an executable specification. By leveraging large language models to infer its underlying requirements, TestGeneralizer constructs reusable test scenario templates and systematically instantiates them, enabling generalization from a single test to a comprehensive set of scenarios. This paradigm shifts beyond traditional coverage-driven methods, and empirical evaluation on twelve open-source Java projects demonstrates that TestGeneralizer improves mutation detection rate by 31.66% and LLM-assessed scenario coverage by 23.08% compared to ChatTester.