Score
Designs and implements systems and artifacts that produce, curate, and evaluate simulation or test scenarios—including generative and search-based generators, scenario composition and taxonomies, scripting and annotation pipelines, and validation/formatting tools—so that scenario ensembles meet constraints on diversity, difficulty, coverage, and simulator compatibility. This competence covers conditioning and sampling mechanisms, scenario-based evaluation and testing, and methods for balancing, constructing, and validating scenario collections for downstream simulation or testing workflows.
This work addresses the challenge that traditional model-based testing is ill-suited for distributed robotic systems due to their high nondeterminism, dynamic reconfiguration, and inherent complexity. To overcome this limitation, the paper proposes the Scenario Specification Language (SCSL), which enables the construction of system-level tests by composing basic scenarios. The approach integrates runtime online test generation and execution with mechanisms for dynamic component joining/leaving and interface reconnection, thereby supporting automated testing and dynamic reconfiguration. The syntax and semantics of SCSL are validated through a robotic salvage mission case study, where automatically generated tests effectively demonstrate the feasibility and advantages of the proposed method.
Safety verification of decision-making agents in dynamic environments faces challenges including susceptibility to local optima in high-dimensional scenario spaces and difficulty balancing scenario diversity with criticality. Method: This paper proposes a dual-space guided testing framework that jointly optimizes the scenario parameter space and agent behavioral space. It introduces a novel parameter–behavior closed-loop feedback mechanism, integrating hierarchical representation, dimensionality reduction modeling, multi-dimensional subspace evaluation, behavioral criticality quantification, and adaptive mode switching to dynamically balance local perturbation and global exploration. Results: Experiments on five decision-making agents show that the framework increases critical scenario generation by 56.23% on average. It significantly outperforms state-of-the-art methods under a joint parameter–behavior driving metric, achieving superior scenario diversity, coverage, and verification effectiveness.
In model-based systems engineering, low experimental data reuse efficiency and excessive redundant experiments hinder digital engineering agility. To address this, this paper proposes a case-based reasoning (CBR)-driven experimental management framework that explicitly integrates domain knowledge. The framework features structured experimental metadata modeling, digital twin–enabled scenario semantic alignment, and an interpretable similarity assessment mechanism to intelligently determine whether historical experiments can be transferred to address new verification queries. Its key innovation lies in embedding domain knowledge explicitly into both the CBR retrieval and adaptation stages, thereby enabling trustworthy cross-operating-condition and cross-configuration experimental data reuse. Evaluated on an industrial-scale vehicle energy system design case, the framework reduces redundant experiments by 37% and shortens early verification cycles by 42% on average, significantly enhancing iterative efficiency in digital engineering and advancing intelligent experimental management.
This paper systematically examines the structural role and evolutionary trajectory of simulation methods across the statistical lifecycle. Addressing the current fragmentation and conceptual ambiguity in simulation practice, the study introduces, for the first time, a comprehensive functional taxonomy—spanning model specification, diagnostic checking, validation, and inference—and proposes a “simulation-driven” paradigm for statistical practice, prioritizing computational scalability. Methodologically, it integrates Monte Carlo simulation, approximate Bayesian computation (ABC), simulation-based calibration, and posterior predictive checking, implemented via high-performance computing frameworks to enable large-scale empirical analysis. Key contributions are: (1) establishing simulation as foundational statistical infrastructure; (2) providing an actionable roadmap for algorithm design, statistical software development, and pedagogical reform; and (3) advancing a paradigm shift in statistical practice—from model-centric to simulation-augmented inference.
This work proposes a novel approach to black-box testing of Functional Mock-up Units (FMUs) by integrating large language models (LLMs) with a human-in-the-loop mechanism. Addressing the inefficiency and poor interpretability of traditional FMU-based dynamic simulation testing—which relies on manually crafted scenarios—the method automatically generates structured Given-When-Then test objectives from FMU interface and functional specifications, and constructs complete test plans comprising input sequences and assertion oracles. Upon simulation execution, the framework produces visualizable logs and statistical evaluation metrics. The approach significantly enhances test design efficiency and result interpretability, facilitates test asset reuse, and demonstrates effectiveness on a lubricating oil cooling system by autonomously generating executable test scenarios and delivering objective-level pass-rate analysis.
Current verification workflows for autonomous systems suffer from a lack of coordination among scenario design, simulation execution, and telemetry analysis, leading to poor traceability between requirements, tests, and evidence, which undermines reproducibility and debugging efficiency. This work proposes a unified verification framework powered by large language models (LLMs) that bridges this gap through task-level structured scenario representations. The framework automatically translates high-level verification intents into temporally evolving scenarios, enabling automated simulation execution and context-aligned telemetry analysis. Furthermore, it incorporates a counterfactual scenario generation mechanism driven by failure cases to establish a closed-loop, self-evolving testing process. The approach substantially enhances traceability, reproducibility, and scalability of verification, accelerates test iteration cycles, and deepens insight into system behavior.
This work addresses the inefficiency in autonomous driving testing caused by cumbersome workflows and redundant code during complex scenario construction. To overcome these limitations, the authors propose Modular2Simple, a novel tool that introduces a modular-composition paradigm for scenario generation. By combining simple or modular OpenSCENARIO scenes, Modular2Simple enables flexible and efficient creation of diverse, complex test scenarios while strictly adhering to the OpenSCENARIO standard. The approach seamlessly integrates with mainstream simulation platforms such as CARLA, significantly enhancing scenario reusability and customizability while reducing development complexity. Experimental results demonstrate that, compared to conventional methods, the proposed solution substantially decreases both development time and labor costs, markedly improving the efficiency and diversity of test scenario construction.
Large language models often produce semantically homogeneous outputs in open-ended generation tasks, failing to meet diversity requirements. This work proposes a unified framework that, for the first time, systematically characterizes the design space of test-time diversity methods by automatically injecting controllable diversity into an intermediate latent representation and conditioning final response generation on this diverse representation. The approach integrates representation-level guidance, conditional language modeling, and a transfer score—quantifying the influence of source diversity on model outputs—for joint optimization. Experimental results across five open-ended tasks and four backbone architectures demonstrate that the proposed framework substantially enhances output diversity while maintaining generation quality on par with baseline models.
This study addresses the limitations of current robotic system validation, which relies heavily on manual selection of test scenarios, thereby hindering scalability and compromising reproducibility and reliability of conclusions. To overcome these challenges, this work proposes a compositional, scenario-based modeling approach that integrates declarative test specifications, plugin-driven scenario generation, containerized parallel simulation, and unified result analysis to establish the first modular and scalable automated verification framework. The framework enables systematic parameter variation across multiple dimensions and facilitates robust identification of systemic faults versus stochastic anomalies. Evaluated across 5,480 distinct scenario configurations with over 100,000 simulation runs, the approach accumulated 1,800 hours of simulated operation and 1,873 virtual kilometers, demonstrating its efficacy in discerning consistent system deficiencies from random irregularities.