ai safety

Designs, builds, and evaluates methods, tools, and processes to reduce harmful or unintended behavior in AI systems, including robustness, verification, interpretability, alignment, monitoring, and incident response. Analyzes failure modes, threat models, and governance controls across model development, deployment, and operation to identify risks and implement mitigations.

aisafety

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.12
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$221K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Current AI incident governance frameworks lack consistency in defining, categorizing, monitoring, and reporting incidents, which constrains the depth and accuracy of post-deployment failure analysis. This study addresses this gap through a systematic literature review and comparative analysis across multiple governance frameworks, thereby identifying and synthesizing key inconsistencies that span existing mechanisms. The work reveals systemic deficiencies in data collection practices, classification logics, and analytical rigor, and elucidates critical misalignments among core governance components. By clarifying these structural disconnects, the research establishes a theoretical foundation and proposes a coordinated pathway toward a unified, standardized framework for AI incident governance.

AI incident governanceclassificationdefinitions

AI safety evaluation lacks consensus standards, limiting its utility for governance and policy decisions. This paper introduces the first practical AI safety evaluation framework, systematically integrating threat modeling, assessment design, and validity validation. It formally defines three essential criteria for “useful” evaluations—risk alignment, reproducibility, and scalability—along with associated quantitative parameters. Innovatively distinguishing formal metrics from real-world risk coverage, the framework establishes an evolutionary paradigm—from isolated tests to modular, composable evaluation suites. It synergistically integrates red-teaming, evaluation validity analysis, and cybersecurity best practices to jointly optimize reliability, construct validity, and operational feasibility. The resulting safety evaluation guidelines have been adopted by industry stakeholders and policymaking bodies, demonstrably enhancing the interpretability of evaluation outcomes and their actionable support for risk-informed decision-making.

Connecting threat modeling to evaluation design effectivelyDefining criteria for high-quality AI safety evaluationsDeveloping comprehensive evaluation suites for AI systems

Usage Governance Advisor: from Intent to AI Governance

Dec 02, 2024
EM
Elizabeth M. Daly
🏛️ IBM Research

To address governance challenges concerning safety, fairness, compliance, and privacy preservation in AI system deployment, this paper proposes the first intent-driven, end-to-end AI risk governance framework. The method automatically constructs semi-structured governance knowledge by parsing user intents and integrating heterogeneous, multi-source information; it employs knowledge graph modeling alongside rule- and model-based collaborative reasoning to enable dynamic risk identification, context-aware risk prioritization, and interpretable risk mitigation recommendations. Its key innovation lies in grounding governance on user intent, thereby closing the loop across “use scenario → risk → benchmark → assessment → mitigation.” Evaluated on real-world deployments, the framework achieves 92.3% risk identification accuracy and over 85% adoption rate of mitigation recommendations, significantly enhancing the automation, operationality, and interpretability of AI governance.

Artificial Intelligence SafetyEthical AIRisk Management in AI

Establishing Minimum Elements for Effective Vulnerability Management in AI Software

Nov 18, 2024
MF
Mohamad Fazelnia
🏛️ University of Hawaii at Manoa

The absence of a unified framework for identifying, assessing, and mitigating vulnerabilities in AI systems hinders systematic AI security governance. Method: This paper proposes the first minimal-element framework for AI software vulnerability management and designs a standardized Artificial Intelligence Vulnerability Database (AIVD). It formally defines four core vulnerability management phases—disclosure, analysis, cataloging, and documentation—and develops an AI-adapted severity scoring model, a weakness enumeration taxonomy, and multi-dimensional mitigation strategies. To support heterogeneous AI models, it introduces a standardized description language, an AI-specific classification ontology, and cross-modal representation techniques. Contribution/Results: The work yields a draft AIVD construction specification, identifies critical capability gaps, and provides a technical foundation for international standards bodies—including NIST—to institutionalize and scale AI security governance.

Addressing unique AI system vulnerabilities beyond traditional security approachesCreating standardized protocols for AI vulnerability documentation and disclosureEstablishing minimum elements for AI vulnerability management systems

Latest Papers

What's happening recently
View more

This study addresses the multidimensional risks—operational, security, and governance-related—that enterprises face when deploying large language models, noting that existing open-source tools are fragmented and fail to comprehensively cover authoritative risk taxonomies. To bridge this gap, the work proposes a structured mapping protocol that automatically aligns the capabilities of 21 prominent open-source tools with the 32 subcategories of the MIT AI Risk Framework, leveraging retrieval-augmented generation (RAG) and LLM-based parsing. The protocol’s validity is substantiated through source code and documentation analysis, majority voting, and inter-rater reliability assessment using Fleiss’ Kappa (κ = 0.509, F1 = 75.5%). Findings reveal a pronounced overconcentration of current tools on technical controls, with significant gaps in governance, legal, and market risk domains, thereby providing an empirical foundation for developing layered AI risk mitigation architectures.

AI risk mitigationgovernancelarge language models

Current AI safety research predominantly focuses on overt failures, often overlooking pervasive latent risks in deployed systems—such as undetectable errors, attribution challenges, and recovery breakdowns. This work proposes a five-dimensional socio-technical framework encompassing cognitive, control, temporal, organizational, and ecosystem integrity to systematically identify novel latent risk patterns, including “uncertainty laundering,” “memory poisoning,” and “synthetic evidence contamination.” The approach is agnostic to specific algorithms and instead leverages integrity modeling and governance mechanism design to expose blind spots in existing safety evaluations. By shifting the paradigm from model-centric to socio-technical reliability, this study advances a actionable agenda for future research and practice in AI safety.

AI safetyhidden failuresintegrity

This study addresses the growing global demand for risk-based AI regulation by systematically identifying, analyzing, and integrating the multifaceted risks of artificial intelligence across technical, ethical, and societal dimensions. Through a comprehensive literature review and framework analysis, it establishes the first systematic alignment between major international regulatory frameworks and AI risk typologies from academic research, thereby constructing a structured risk taxonomy. The work clarifies key dimensions and limitations of existing risk assessment methodologies, distills best practices, and identifies critical research gaps. By doing so, it provides a robust theoretical foundation and strategic guidance for the development of standardized, actionable AI risk management tools aligned with evolving regulatory expectations.

AI risk assessmentAI safetyregulatory frameworks