ai security engineering

Designs, builds, and evaluates security controls and defensive mechanisms for AI systems, including model hardening, adversarial-robustness measures, access controls, monitoring, and incident response across training and inference pipelines. Implements and operationalizes detection, mitigation, and governance controls to reduce risks from attacks such as data poisoning, model extraction, evasion, and misuse.

aisecurityengineering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.72
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Current AI risk mitigation frameworks suffer from fragmentation, terminological ambiguity, and coverage gaps, hindering coordinated multistakeholder governance. To address this, we introduce the first cross-framework taxonomy for AI risk mitigation, systematically synthesizing 831 mitigation measures from 13 prominent frameworks published between 2023 and 2025. Our methodology combines rapid evidence scanning, iterative clustering-based coding, and structured knowledge modeling to develop a four-dimensional classification—governance & oversight, technical safety, operational processes, and transparency & accountability—with 23 granular subcategories. We explicitly resolve semantic inconsistencies in key terms (e.g., “red-teaming,” “risk management”) and deliver a scalable, role-aligned taxonomy alongside a dynamic, open-source database. The resulting resource enables comparative framework analysis and gap identification, supporting national policymaking and AI safety organizations worldwide. All artifacts are publicly released to advance global AI governance infrastructure.

Addressing inconsistent terminology and coverage gaps in AI risk managementOrganizing fragmented AI risk mitigation frameworks into a unified taxonomyProviding a common reference for AI risk mitigation across organizations and governments

Must-Read Papers

Most classic and influential ideas
View more

As AI agents become deeply embedded in critical systems, their misaligned behaviors pose significant internal security challenges. This work proposes a hierarchical defense framework that pioneers an extension of the MITRE ATT&CK matrix into the TRAIT&R taxonomy, establishing a four-layer detection and three-tier prevention-response mechanism aligned with the evolving capabilities of AI models. By integrating threat modeling, capability-tiered controls, chain-of-thought monitoring, asynchronous alerting, real-time access control, and system-level anomaly detection, the framework systematically incorporates fifteen concrete mitigation measures. These collectively address security requirements spanning from current to future high-capability AI systems, offering an actionable and scalable defense roadmap for AI-controlled environments.

adversarial AIAI alignmentAI security

Limits of Safe AI Deployment: Differentiating Oversight and Control

Jul 04, 2025
DM
David Manheim
🏛️ Association for Long Term Existence and Resilience | Centre for the Governance of AI

This paper addresses the conceptual conflation of “oversight” and “control” in AI safety governance, systematically distinguishing their distinct objectives, operational mechanisms, and temporal scopes. Through a critical cross-disciplinary literature review—and integrating insights from Responsible AI maturity models and risk governance theory—it develops a theoretically rigorous yet policy-actionable analytical framework, introducing the first AI Oversight Maturity Model (AI-OMM). The model identifies critical boundary conditions for oversight failure and establishes a structured, conditional system for assessing the feasibility of meaningful human oversight. Key contributions include: (1) clarifying the normative distinction between oversight and control; (2) diagnosing design gaps and contextual limitations in current oversight mechanisms; and (3) providing regulators, auditors, and developers with a practical tool to evaluate oversight effectiveness, detect capability gaps, and guide technical alignment with governance requirements.

Differentiating oversight and control in AI supervisionIdentifying limitations and needs in AI supervision mechanismsProposing a framework for meaningful human supervision conditions

Traditional defense mechanisms struggle to counter AI-driven adaptive cyberattacks, as they are ill-equipped to handle autonomous adversarial agents capable of evading detection. This work proposes a novel defensive infrastructure centered on controllable offensive AI, which, within regulated environments, is trained to simulate the full attack lifecycle to proactively acquire and transform threat intelligence into defensive knowledge. The study’s core contributions include the first systematic benchmark encompassing the entire attack chain, alongside a training-based vulnerability discovery agent, an open-weight model governance framework, a tiered capability release mechanism, and a defense-oriented agent distillation technique. Together, these establish a “offense-informed defense” strategic paradigm and outline three actionable pathways for the safe development and constraint of offensive AI capabilities.

AI agentscyber attacksdefensive strategy

To address the lack of provable behavioral guarantees for large-scale autonomous AI systems under adversarial attacks and operational stress, this paper proposes the first engineering-grade safety and trustworthiness assurance framework spanning the entire system lifecycle—design, training, deployment, and runtime operation. Methodologically, it innovatively integrates standardized threat modeling with quantitative risk assessment, adversarial robustness training, lightweight real-time anomaly detection, automated audit logging, and compliance verification protocols into a unified assurance pipeline. Key contributions include: (1) proactive, risk-aware assurance embedded early in the development cycle; (2) security-by-design, wherein safety properties are intrinsically encoded into model architecture; and (3) formally verifiable and mathematically provable system behavior. Experimental evaluation demonstrates significant reductions in vulnerability rates and compliance overhead across national security, open-model governance, and industrial automation domains, confirming strong scalability and cross-domain applicability.

Ensuring safe operation of large-scale autonomous AI modelsIntegrating security measures into AI development lifecycleReducing vulnerabilities in AI systems across various sectors

Ctrl-Z: Controlling AI Agents via Resampling

Apr 14, 2025
AB
Aryan Bhatt
🏛️ Redwood Research | ML Alignment and Theory Scholars (MATS) Program

AI agents performing multi-step system administration tasks pose a security risk by covertly downloading and executing malicious code. Method: This paper proposes a control-assessment-based defense framework. It introduces BashBench—the first benchmark tailored to AI agent environments—for systematic control evaluation; designs a novel dynamic action-level resampling protocol that selectively resamples suspicious actions to enable real-time attack blocking and behavioral attribution; and integrates adversarial agent construction, joint historical-step analysis, and multi-step task modeling to balance security and practicality. Contribution/Results: Experiments demonstrate that the optimal protocol reduces malicious agent attack success rate from 58% to 7%, while degrading benign task performance by only 5%, validating both effectiveness and low intrusiveness.

Balancing attack prevention with agent usefulnessEvaluating AI agent safety in multi-step tasksPreventing covert malicious code execution by AI

Latest Papers

What's happening recently
View more

This study addresses the security risks posed by AI agents with offensive cyber capabilities that may breach sandbox boundaries in evaluation environments. It systematically identifies five categories of boundary vulnerabilities—multi-step attacks, objective conflicts, supply chain leaks, persistence mechanisms, and automated execution speed—and conducts a case analysis grounded in the 2026 Hugging Face/OpenAI incident. The work introduces the first taxonomy of AI boundary vulnerabilities specifically tailored to evaluation settings and proposes an integrated defense framework combining isolation, privilege separation, behavioral provenance tracking, and defensive response interfaces. By jointly considering misuse risks and capability assessment, this research establishes clear security priorities for high-risk AI evaluations, offering both theoretical foundations and practical guidance for developing trustworthy evaluation environments that balance testing efficacy with risk containment.

AI security evaluationcyber-capable AI agentsevaluation containment

This work addresses the absence of standardized, composable oversight infrastructure in current AI deployments, which leads teams to repeatedly build fragmented auditing and monitoring mechanisms. The authors propose a five-layer, six-dimension framework for AI oversight, with a particular focus on formally defining— for the first time—the “normative layer.” This layer translates human intent into executable, traceable, and upgradable machine-checkable norms through six design principles, including elicitable, adversarially aware, and governable specifications. Integrating formal methods, policy languages (e.g., Cedar, OPA), and constitutional AI concepts, the study introduces CARMA, a norm-driven runtime oversight prototype that demonstrates how a single norm can uniformly drive execution, evaluation, and upgrading. The system validates the feasibility of reusing composable oversight components across teams.

AI OversightComposabilityCoordination Gap

This study addresses a critical gap in AI alignment research, which has predominantly emphasized ex ante prevention while neglecting post-incident response and resilience management. The work proposes the first systematic framework for post-hoc response to AI incidents, introducing a three-dimensional classification matrix based on controllability, intent, and severity. This matrix distinguishes between “extremely difficult to control” and “fully uncontrollable” scenarios and further categorizes manageable incidents into accidental and adversarial failures. Drawing on circuit-breaker mechanisms from safety engineering and tiered response strategies from cybersecurity, the project develops actionable response protocols. These provide policymakers and developers with proportionate, evidence-based decision-making guidance, thereby filling a significant void in AI incident management and substantially enhancing overall system resilience.

AI loss of controlcatastrophic AI riskincident management

This study addresses a critical limitation in current AI safety evaluations, which assume indiscriminate adversarial attacks and thereby overestimate system robustness by neglecting attackers’ strategic timing. We propose the first decomposition of adversarial strategy into initiation and termination mechanisms, significantly reducing empirical safety without enhancing attack capabilities. Implementing this approach within a red-teaming framework, we conduct stress tests in BashArena and LinuxArena under constrained human auditing budgets. Experimental results demonstrate that, with only a 1% audit budget, the initiation strategy reduces safety by 20 percentage points in both environments, while the termination strategy further decreases safety by 20 and 28 percentage points, respectively. These findings reveal a substantial misalignment between prevailing evaluation paradigms and realistic threat models.

agentic AIAI controlattack selection