Score
Designs and executes tests, analyses, and engineering plans that measure how a system or model’s performance, safety, robustness, and costs change under deployment choices and environmental factors; this includes evaluating effects of numeric precision, sampling temperature, quantization, attack vectors, and other axes of variability to detect decision instability and failure modes. Builds and assesses deployment artifacts and processes — experimental/field rollouts, operational validation, feasibility, risk and cost assessments, and deployment engineering — to inform safe, robust real‑world operation.
Contemporary AI systems are rapidly advancing toward transformative capabilities, necessitating safety evaluation methodologies that transcend conventional static benchmarks. Method: We propose a novel three-dimensional safety assessment framework—“Capability–Propensity–Control”—that systematically integrates measurement targets (e.g., deception capability, power-seeking propensity, adversarial robustness), measurement modalities (behavioral testing, internal analysis), and governance mapping. This framework overcomes limitations of static benchmarking by enabling dynamic, multi-layered evaluation. Contribution/Results: We introduce the first unified taxonomy covering the full stack of AI safety assessment; identify critical evaluation pitfalls—including “safety washing” and “model sandbagging”; formally define safety-critical capabilities and hazardous propensities; provide practitioners with actionable assessment guidelines; establish decision-support interfaces for regulators; and uncover fundamental research gaps in scalable, interpretable, and governance-aligned safety evaluation.
This study addresses the lack of a systematic framework for identifying critical input variables and conducting sensitivity analysis under uncertainty in complex simulations, particularly in military decision-making contexts. The authors propose a unified sensitivity analysis framework that integrates local and global methods—including variance-based, derivative-based, screening, and uncertainty quantification techniques—and strategically maps these approaches to specific decision objectives such as factor prioritization, fixing, variance reduction, and mapping. Innovatively, the framework introduces a “sensitivity audit” mechanism to enhance traceability of model assumptions and promote responsible model usage. By providing a structured guide for high-dimensional, complex simulation systems, this work significantly improves model interpretability, transparency, and the credibility of decisions derived from such models.
This study addresses a critical limitation of existing DORA metrics, which rely solely on first-order statistics and thus fail to capture the distributional characteristics of software release cadence or distinguish teams with markedly different release regularity. To overcome this, the work introduces second-order statistics into the DORA framework for the first time, proposing a novel Delivery Consistency (DC) metric based on the coefficient of variation of inter-release intervals. It further constructs an eight-prototype Delivery Health Matrix to enable multidimensional diagnosis and targeted intervention for software delivery rhythms across platforms. Validation using real-world data spanning 120 weeks from four platforms—including Jira, GitHub, and Firebase—demonstrates that the approach effectively identifies teams sharing identical DORA ratings yet exhibiting divergent release patterns, uncovering underlying organizational or process constraints common to such teams.
This study addresses safety risks arising from software faults in cyber-physical systems (CPS) for electric bicycles. We propose a simulation-driven functional Failure Mode and Effects Analysis (FMEA) method, leveraging Simulink Fault Analyzer to construct fault models, integrated with expert review and a systematic FMEA process to close the loop among fault modeling, simulation-based analysis, and impact assessment. Experimental evaluation identified 13 real-world faults with 100% model accuracy; among them, five revealed previously unrecognized safety implications, and 38.4% induced anomalous system behavior. The study distills ten reusable engineering practice guidelines, significantly enhancing the effectiveness and practicality of FMEA in industrial-scale CPS. It provides empirical validation and methodological contributions toward the operational deployment of simulation-driven safety analysis.
In the pre-prototype phase of complex novel systems, the absence of empirical data impedes rigorous assessment of simulation model credibility. Method: This paper proposes a physics-fidelity-based model trust evaluation method that bypasses reliance on real-world measurements. Instead, it quantifies model applicability by systematically analyzing the completeness of represented physical phenomena, the mathematical complexity of their formulation, and the fidelity of emergent behavior modeling. Contribution/Results: The approach enables objective, quantitative ranking of multiple candidate models under data-scarce conditions—thereby significantly enhancing the reliability of simulation-driven decisions during early-stage design. It establishes both theoretical foundations and practical tools for model-based design in high-uncertainty scenarios, advancing trustworthy digital twin development and physics-informed simulation validation.
Safety-critical small Unmanned Aircraft Systems (sUAS) lack systematic, standardized testing processes that are tightly integrated with safety analysis. Method: This paper proposes a requirement-driven coupled testing framework, introducing the novel triadic paradigm of “requirements–simulation testing–safety analysis.” It employs formal requirement modeling with bidirectional traceability, a simulation–hardware-in-the-loop cooperative testing architecture, scenario-driven test case generation, and deep integration of safety analysis methods (e.g., Fault Tree Analysis and System-Theoretic Process Analysis). Contribution/Results: Evaluated on an sUAS case study, the framework significantly improves simulation fidelity coverage and requirement coverage, enables end-to-end safety evidence generation, fills the gap in standardized sUAS testing procedures, and delivers reproducible, verifiable testing assets to support airworthiness certification.
This work addresses the challenge of silent updates to large language models (LLMs) by service providers, which often occur without version changes and can lead to behavioral drift and functional regressions, while existing mechanisms lack deployment-side control over compatibility governance. Framing LLM updates as a software supply chain governance problem, this study proposes a deployment-side control framework that defines rule-based production contracts, constructs risk-category-oriented test suites, and enforces compatibility gates to validate model safety and performance prior to updates. Experimental results demonstrate that the approach effectively uncovers fine-grained regressions missed by aggregate metrics, while also highlighting critical challenges in test design, threshold calibration, and drift attribution.
Current verification workflows for autonomous systems suffer from a lack of coordination among scenario design, simulation execution, and telemetry analysis, leading to poor traceability between requirements, tests, and evidence, which undermines reproducibility and debugging efficiency. This work proposes a unified verification framework powered by large language models (LLMs) that bridges this gap through task-level structured scenario representations. The framework automatically translates high-level verification intents into temporally evolving scenarios, enabling automated simulation execution and context-aligned telemetry analysis. Furthermore, it incorporates a counterfactual scenario generation mechanism driven by failure cases to establish a closed-loop, self-evolving testing process. The approach substantially enhances traceability, reproducibility, and scalability of verification, accelerates test iteration cycles, and deepens insight into system behavior.
This work addresses the problem of implementation drift in evolving distributed systems, where runtime behavior gradually deviates from the original design. To tackle this issue, the paper proposes a design conformance assessment method based on distributed tracing data. It introduces, for the first time in the domain of distributed systems, conformance checking techniques from process mining, leveraging runtime traces collected via the OpenTelemetry standard and automatically comparing them against behavioral models defined at design time to quantify their alignment. The key contribution lies in establishing persistent, monitorable conformance metrics that enable continuous, automated evaluation of deviations between system implementation and design. This approach is readily applicable to modern distributed systems widely adopting OpenTelemetry for observability.