safety evaluation

Formulating protocols, tests, and metrics to evaluate system safety and compliance under normal, failure, and adversarial conditions, and assessing how interventions (controls, steering, concept removal) change behavior and resilience to attacks.

safetyevaluation

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Quantifying Security for Networked Control Systems: A Review

Oct 21, 2025
SC
Sribalaji C. Anand
🏛️ KTH Royal Institute of Technology | Digital Futures | Uppsala University

To address the challenge of quantifying security risks for Networked Control Systems (NCSs) in critical infrastructure—such as power grids and transportation networks—under cyberattacks, this paper proposes a probabilistic risk modeling framework that integrates uncertain prior knowledge, shifting from traditional static vulnerability analysis to dynamic resilience assessment. Methodologically, it unifies control theory, cybersecurity analysis, probabilistic risk assessment, and statistical modeling into an interdisciplinary security quantification methodology. It introduces, for the first time, a systematic taxonomy for NCS security quantification, categorizing prevailing defense strategies and identifying key research gaps. The resulting framework provides computationally tractable and empirically verifiable theoretical tools and practical guidelines for resilience-oriented NCS design, significantly enhancing risk characterization under partial attacker knowledge.

Developing metrics for cyber-attack resilience assessmentQuantifying vulnerabilities in networked control systemsReviewing mitigation strategies for critical infrastructure protection

This work addresses the current lack of a systematic understanding of the capabilities of AI sandboxes in ensuring safety, security, and regulatory compliance, particularly within physical AI and cyber-physical systems. It proposes the first unified, assurance-oriented framework for AI sandboxes, introducing a formal boundary definition, a comprehensive sandbox taxonomy, a threat model targeting the assurance mechanisms themselves, and a quantifiable evaluation methodology spanning six dimensions—including fidelity and controllability. Through formal modeling, threat analysis, and multi-case validation, the study clarifies what aspects of AI behavior can be effectively tested in sandboxes, which risk categories can be meaningfully controlled, and what forms of evidence such environments can generate to support safety and compliance claims, thereby establishing foundational tools for trustworthy AI verification.

AI sandboxassurancecyber-physical systems

Must-Read Papers

Most classic and influential ideas
View more

Quantitative Measurement of Cyber Resilience: Modeling and Experimentation

Mar 28, 2023
MJ
Michael J. Weisman
🏛️ DEVCOM Army Research Laboratory | Pennsylvania State University | ICF International | University of California, Irvine

Current cyber-physical systems (CPS) in vehicular environments lack quantitative, experimentally grounded methods for assessing network resilience. Method: This study constructs an experimental testbed replicating real-world truck operational conditions and conducts multiple rounds of malware injection attacks, simultaneously collecting network- and physical-layer data on resistance and recovery behaviors. Contribution/Results: We introduce the novel concept of “bonware” to holistically characterize both cybersecurity defense capability and physical resilience, formalized via an analytically tractable mathematical model. We further define and extract experimentally identifiable, quantitative resilience metrics—termed elastic features—for the first time. Sensitivity analysis confirms these metrics exhibit significant discriminability with respect to attack intensity, defensive strategies, and physical redundancy. This work bridges a critical gap by advancing vehicular CPS resilience from qualitative description to quantifiable, comparable, and optimizable measurement.

Attack RecoveryCyber ResilienceMeasurement Tools

AI safety evaluation lacks consensus standards, limiting its utility for governance and policy decisions. This paper introduces the first practical AI safety evaluation framework, systematically integrating threat modeling, assessment design, and validity validation. It formally defines three essential criteria for “useful” evaluations—risk alignment, reproducibility, and scalability—along with associated quantitative parameters. Innovatively distinguishing formal metrics from real-world risk coverage, the framework establishes an evolutionary paradigm—from isolated tests to modular, composable evaluation suites. It synergistically integrates red-teaming, evaluation validity analysis, and cybersecurity best practices to jointly optimize reliability, construct validity, and operational feasibility. The resulting safety evaluation guidelines have been adopted by industry stakeholders and policymaking bodies, demonstrably enhancing the interpretability of evaluation outcomes and their actionable support for risk-informed decision-making.

Connecting threat modeling to evaluation design effectivelyDefining criteria for high-quality AI safety evaluationsDeveloping comprehensive evaluation suites for AI systems

Quantitative Resilience Modeling for Autonomous Cyber Defense

Mar 04, 2025
XC
Xavier Cadet
🏛️ Dartmouth College | Northeastern University

This paper addresses the challenge of quantifying system resilience under cyberattacks by proposing an operation-goal-oriented, interpretable resilience quantification framework. Methodologically, it integrates dynamic resource criticality modeling, multi-objective weighted resilience aggregation, and cross-temporal attack and heterogeneous topology adaptation mechanisms; a reinforcement learning (RL)-based defensive agent is implemented on the CybORG platform. Key contributions include: (1) the first resilience quantification anchored explicitly to security operations objectives, ensuring interpretability; (2) dynamic resilience assessment capability across multi-stage attacks and diverse network topologies; and (3) empirical validation that proactive hardening and rapid recovery constitute two core pathways for enhancing resilience. Experiments demonstrate that the proposed RL policy significantly outperforms heuristic baselines in both mission assurance rate and recovery timeliness.

Design RL agents for proactive network hardening and recovery strategies.Evaluate resilience trade-offs in autonomous cyber defense systems.Quantify cyber resilience under diverse network topologies and attack patterns.

This study addresses the inadequacy of current IT compliance–oriented cybersecurity policies in safeguarding the physical safety of cyber-physical systems, as digital failures often precipitate real-world harm. By coding 292 critical infrastructure policies (2000–2025) and aligning them with the NIST SP 800-160 Vol. 2 resilience lifecycle, the research reveals a significant misalignment between prevailing policy approaches—overreliant on IT control catalogs during resistance and recovery phases—and actual physical risks. The work proposes a modernized “duty of reasonable care” standard centered on hazard-specific traceability, structured assurance cases, and cyber resilience engineering. It identifies three critical disconnects: misaligned delegation of standards, reduction of recovery mechanisms to mere incident reporting, and uneven sectoral adaptability. The study further outlines a viable pathway for federal policy that integrates engineering implementation with targeted incentives.

critical infrastructurecyber safetycyber-physical systems

Ctrl-Z: Controlling AI Agents via Resampling

Apr 14, 2025
AB
Aryan Bhatt
🏛️ Redwood Research | ML Alignment and Theory Scholars (MATS) Program

AI agents performing multi-step system administration tasks pose a security risk by covertly downloading and executing malicious code. Method: This paper proposes a control-assessment-based defense framework. It introduces BashBench—the first benchmark tailored to AI agent environments—for systematic control evaluation; designs a novel dynamic action-level resampling protocol that selectively resamples suspicious actions to enable real-time attack blocking and behavioral attribution; and integrates adversarial agent construction, joint historical-step analysis, and multi-step task modeling to balance security and practicality. Contribution/Results: Experiments demonstrate that the optimal protocol reduces malicious agent attack success rate from 58% to 7%, while degrading benign task performance by only 5%, validating both effectiveness and low intrusiveness.

Balancing attack prevention with agent usefulnessEvaluating AI agent safety in multi-step tasksPreventing covert malicious code execution by AI

Latest Papers

What's happening recently
View more

Current AI incident governance frameworks lack consistency in defining, categorizing, monitoring, and reporting incidents, which constrains the depth and accuracy of post-deployment failure analysis. This study addresses this gap through a systematic literature review and comparative analysis across multiple governance frameworks, thereby identifying and synthesizing key inconsistencies that span existing mechanisms. The work reveals systemic deficiencies in data collection practices, classification logics, and analytical rigor, and elucidates critical misalignments among core governance components. By clarifying these structural disconnects, the research establishes a theoretical foundation and proposes a coordinated pathway toward a unified, standardized framework for AI incident governance.

AI incident governanceclassificationdefinitions

Traditional penetration testing struggles to evaluate security risks in AI systems arising from violations of behavioral objectives without breaching underlying infrastructure. This work proposes the first formal definition of AI penetration testing, reframing it as an objective-driven behavioral security assessment. The approach involves identifying operational objectives, mapping AI-driven behaviors, analyzing adversarial attack surfaces—such as prompt injection, data poisoning, and sensor manipulation—establishing criteria for behavioral failure, and conducting scenario-based red-teaming exercises. By integrating threat modeling, behavior mapping, and evidentiary chain construction, the framework demonstrates its efficacy and novelty in a case study involving an AI-powered Security Operations Center assistant, successfully uncovering attack pathways that violate system objectives through behavioral manipulation alone, without requiring infrastructure compromise.

adversarial influenceAI-enabled systemsbehavioral objective violation

Current AI safety research predominantly focuses on overt failures, often overlooking pervasive latent risks in deployed systems—such as undetectable errors, attribution challenges, and recovery breakdowns. This work proposes a five-dimensional socio-technical framework encompassing cognitive, control, temporal, organizational, and ecosystem integrity to systematically identify novel latent risk patterns, including “uncertainty laundering,” “memory poisoning,” and “synthetic evidence contamination.” The approach is agnostic to specific algorithms and instead leverages integrity modeling and governance mechanism design to expose blind spots in existing safety evaluations. By shifting the paradigm from model-centric to socio-technical reliability, this study advances a actionable agenda for future research and practice in AI safety.

AI safetyhidden failuresintegrity

This study addresses the security risks posed by AI agents with offensive cyber capabilities that may breach sandbox boundaries in evaluation environments. It systematically identifies five categories of boundary vulnerabilities—multi-step attacks, objective conflicts, supply chain leaks, persistence mechanisms, and automated execution speed—and conducts a case analysis grounded in the 2026 Hugging Face/OpenAI incident. The work introduces the first taxonomy of AI boundary vulnerabilities specifically tailored to evaluation settings and proposes an integrated defense framework combining isolation, privilege separation, behavioral provenance tracking, and defensive response interfaces. By jointly considering misuse risks and capability assessment, this research establishes clear security priorities for high-risk AI evaluations, offering both theoretical foundations and practical guidance for developing trustworthy evaluation environments that balance testing efficacy with risk containment.

AI security evaluationcyber-capable AI agentsevaluation containment

This study addresses a critical gap in AI alignment research, which has predominantly emphasized ex ante prevention while neglecting post-incident response and resilience management. The work proposes the first systematic framework for post-hoc response to AI incidents, introducing a three-dimensional classification matrix based on controllability, intent, and severity. This matrix distinguishes between “extremely difficult to control” and “fully uncontrollable” scenarios and further categorizes manageable incidents into accidental and adversarial failures. Drawing on circuit-breaker mechanisms from safety engineering and tiered response strategies from cybersecurity, the project develops actionable response protocols. These provide policymakers and developers with proportionate, evidence-based decision-making guidance, thereby filling a significant void in AI incident management and substantially enhancing overall system resilience.

AI loss of controlcatastrophic AI riskincident management

Hot Scholars

SC

Stephen Casper

PhD student, MIT
AI safetyAI responsibilityred-teamingrobustness
DS

Dawn Song

Professor of Computer Science, UC Berkeley
Computer Security and Privacy
SM

Sören Mindermann

University of Oxford, OATML
AI safetydeep learningactive learningcausal inference
VM

Vasilios Mavroudis

Research Scientist, Alan Turing Institute
Machine LearningSystems SecurityArtificial Intelligence
XE

Xin Eric Wang

Assistant Professor, University of California, Santa Barbara, Simular
NLPCVMLLanguage and Vision