Score
Designs, implements, and analyzes quantitative and qualitative metrics, tests, protocols, and policies that measure and enforce system safety—covering auditability, controllability, requirements, validation, and safety-state evaluation—and builds evaluation pipelines for safety testing, safety metric analysis, and safety assessment. Performs behavioral and comparative analyses such as isolating and measuring refusal behavior, comparing aligned versus ablated/abliterated systems, attributing performance differences to safety state, and measuring downstream impact of safety interventions.
Contemporary AI systems are rapidly advancing toward transformative capabilities, necessitating safety evaluation methodologies that transcend conventional static benchmarks. Method: We propose a novel three-dimensional safety assessment framework—“Capability–Propensity–Control”—that systematically integrates measurement targets (e.g., deception capability, power-seeking propensity, adversarial robustness), measurement modalities (behavioral testing, internal analysis), and governance mapping. This framework overcomes limitations of static benchmarking by enabling dynamic, multi-layered evaluation. Contribution/Results: We introduce the first unified taxonomy covering the full stack of AI safety assessment; identify critical evaluation pitfalls—including “safety washing” and “model sandbagging”; formally define safety-critical capabilities and hazardous propensities; provide practitioners with actionable assessment guidelines; establish decision-support interfaces for regulators; and uncover fundamental research gaps in scalable, interpretable, and governance-aligned safety evaluation.
This study addresses a critical “evaluation–safety gap” (EvalSafetyGap) in large language models (LLMs), wherein apparent performance gains do not necessarily reflect genuine safety capabilities. Through a systematic literature review, gray literature analysis, and a multidimensional audit of ten models across eight evidence streams, the work proposes the EvalSafetyGap hypothesis and introduces two novel constructs—“instability decomposition” and the “alignment trilemma”—to establish a unified terminology and evidence map supporting dynamic evaluation and auditable alignment. Empirical findings reveal no significant correlation between model capability and adversarial robustness (r = 0.232, p = 0.520). Moreover, safety differences between open- and closed-source models stem primarily from governance transparency rather than behavioral robustness, with results highly sensitive to model categorization and evaluation protocols.
In early-stage collaborative robot task design, safety experts struggle to comprehend task logic, and risk assessment outcomes often lack practical implementability. Method: This paper proposes a model-driven risk assessment approach based on Behavior Trees (BTs)—the first application of BTs in risk assessment—enabling early risk identification, formal verification, and end-to-end traceability via visual modeling. Integrating Model-Driven Engineering (MDE) with Human Factors evaluation, the method was empirically validated by cross-functional practitioners from five industrial enterprises. Contribution/Results: The approach significantly improves risk identification completeness (+32%) and enhances collaboration efficiency between safety experts and development teams, reducing communication overhead by 41%. It establishes a novel, industrial-grade paradigm for trustworthy robotic systems that unifies modeling, analysis, and implementation within a single coherent framework.
Current LLM security evaluations predominantly focus on base models, neglecting the critical impact of application-layer components—such as system prompts, retrieval processes, and safety mitigations—on end-to-end security. Method: We propose the first security assessment framework tailored to real-world LLM applications, integrating a domain-customized risk taxonomy, system prompt analysis, retrieval pipeline auditing, and mitigation mechanism evaluation to enable holistic, multi-component risk identification. The framework supports cross-use-case scalability and organization-level customization. Contribution/Results: Validated across multiple internal production scenarios, it significantly improves security test coverage and high-severity risk detection rates. This work bridges the gap between AI safety theory and engineering practice, delivering a reusable methodology and actionable guidelines for secure, large-scale LLM deployment.
Leading AI organizations lack systematic methodologies for identifying safety risks, particularly those arising from feedback-driven and interactive failure modes; manual assessment further suffers from limited causal traceability and incompleteness. Method: This work pioneers the adaptation of System Safety Engineering’s STPA (System-Theoretic Process Analysis)—a rigorous hazard analysis framework from aviation and industrial safety—to AI systems. We construct control structure models, integrate loss scenario simulation with an AI-specific safety case framework (Korbak et al., 2025), and automatically identify Unsafe Control Actions and associated Loss Scenarios overlooked by existing threat models. Contribution/Results: The approach enables LLM-augmented, scalable analysis, significantly improving coverage, causal traceability, and robustness in hazard identification. Experimental validation demonstrates that STPA provides a verifiable, extensible, and complementary assurance mechanism for AI safety governance.
This study addresses the limitations of current safety evaluations, which predominantly rely on isolated multiple-choice setups and overlook the real-world impact of agent scaffolding on model safety. Through a large-scale controlled experiment (N = 62,808), we systematically assess four scaffolding architectures—including Map-Reduce—across evaluation formats (open-ended vs. multiple-choice) on state-of-the-art language models. We find that evaluation format exerts a far stronger influence on safety scores than scaffolding effects themselves. Critically, strong model–scaffolding interactions lead to complete reversals in safety rankings across benchmarks (G = 0.000). Employing preregistration, evaluator blinding, TOST equivalence testing, and generalizability analyses, we demonstrate that safety must be evaluated for each specific model–deployment configuration: Map-Reduce significantly reduces safety (NNH = 14), whereas other architectures remain equivalent within ±2 percentage points.
This study addresses the limitations of existing System-Theoretic Process Analysis (STPA) methods, which rely heavily on manual effort for identifying Unsafe Control Actions (UCAs), resulting in low efficiency and susceptibility to human error. To overcome these challenges, this work proposes the first integration of robustness analysis into the STPA framework, combining model-driven engineering with formal methods to enable automated and complete UCA identification. The approach was validated through a case study on an aircraft braking control system, demonstrating significant improvements in both analytical efficiency and accuracy. Furthermore, user studies indicate that the accompanying tool offers practical utility for the majority of safety analysts, supporting its real-world applicability in safety-critical systems engineering.
This study addresses the limitations of existing SysML verification approaches, which are often tool-dependent and restricted to performance properties, lacking support for automated validation of behavioral and interface requirements. To overcome these shortcomings, this work proposes a tool-agnostic, automated verification workflow driven by SysML test cases, integrating UML Testing Profile and behavioral diagram constructs to enable unified validation of multidimensional attributes—including behavior, timing, and state responses. The methodology was developed through a mixed-methods research strategy combining literature review and stakeholder interviews, and its efficacy was empirically validated across two independent SysML toolchains. The approach not only transcends the constraints of conventional parametric methods but also enables automatic traceability of verification results back to the original model elements.
Current LLM-assisted security analysis tools lack rigorous self-validation and are vulnerable to hallucinations, unverifiable constraints, and insufficient auditability. This work addresses this gap by introducing Constitutional Meta-STPA, the first framework to apply Systems-Theoretic Process Analysis (STPA) at the meta-level of such tools. It systematically derives 29 actionable governance principles—including eight meta-safety principles—from hazard analysis and formally binds them to specific code execution points. By integrating Claude Opus/Sonnet, formal verification, and a constitution-driven constraint mechanism, the study successfully instantiates 18 tool-level principles alongside all meta-principles, thereby validating both the efficacy of the meta-analytic approach and its dependence on underlying model capabilities.