agent threat modeling

Designs and builds threat models that identify and characterize harmful capabilities, failure modes, and misuse pathways of autonomous or agentic systems, and maps those threats to applicable legal and regulatory obligations. Produces prioritized control sets and concrete, actionable mitigation requirements by analyzing how regulations amplify or constrain risks and translating regulatory obligations into technical, operational, and governance controls.

agentthreatmodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the lack of systematic analysis and actionable controls linking large language model (LLM) agent security threats to real-world financial regulations. It presents the first mapping of six categories of LLM agent risks to regulatory obligations in the U.S. and EU, and proposes four scalable compliance architectures centered on auditability, authorization, and boundary enforcement. Key technical contributions include agent-to-agent (A2A) compliance orchestration, audit-driven Grounded-RAG, case ID propagation, and reasoning-boundary de-identification proxies. Empirical evaluation demonstrates that the framework reduces manual processing from multiple days to same-day handling, automates approximately 80% of use cases, and uncovers two types of control failures detectable only through internal audit, as well as one category of legitimate applicants erroneously rejected.

agent securityauditabilitycompliance

To address the lack of provable behavioral guarantees for large-scale autonomous AI systems under adversarial attacks and operational stress, this paper proposes the first engineering-grade safety and trustworthiness assurance framework spanning the entire system lifecycle—design, training, deployment, and runtime operation. Methodologically, it innovatively integrates standardized threat modeling with quantitative risk assessment, adversarial robustness training, lightweight real-time anomaly detection, automated audit logging, and compliance verification protocols into a unified assurance pipeline. Key contributions include: (1) proactive, risk-aware assurance embedded early in the development cycle; (2) security-by-design, wherein safety properties are intrinsically encoded into model architecture; and (3) formally verifiable and mathematically provable system behavior. Experimental evaluation demonstrates significant reductions in vulnerability rates and compliance overhead across national security, open-model governance, and industrial automation domains, confirming strong scalability and cross-domain applicability.

Ensuring safe operation of large-scale autonomous AI modelsIntegrating security measures into AI development lifecycleReducing vulnerabilities in AI systems across various sectors

A Systematic Approach to Estimate the Security Posture of a Cyber Infrastructure: A Technical Report

Aug 29, 2025
QS
Qishen Sam Liang
🏛️ USC Information Sciences Institute

Scientific research cyberinfrastructure (CI) faces unique challenges—including high collaboration requirements, component heterogeneity, and the absence of adaptable security assessment frameworks. To address these, we propose a mission-centric security posture assessment method: first, top-down identification of critical assets and unacceptable losses; second, construction of a security knowledge graph integrating system components, dependencies, and threat behaviors; and third, integration with directed attack graphs to quantify multi-hop attack paths from entry points to critical assets—enabling visualization of attacker-defender relationships and identification of security blind spots. Unlike conventional generic standards, our approach is the first to deeply couple mission-driven assessment, knowledge graphs, and attack graphs. It supports risk prioritization and generation of actionable defensive strategies, significantly enhancing the precision and effectiveness of CI security defense.

Addressing lack of practical security assessment frameworksEstimating security posture of collaborative cyber infrastructuresSystematically mapping adversary attack paths to critical assets

An Approach to Technical AGI Safety and Security

Apr 02, 2025
RS
Rohin Shah
🏛️ Google DeepMind

Prior to large-scale AGI deployment, misuse and goal misalignment represent two critical safety risks. Method: We propose a dual-track collaborative defense framework: (1) at the model level, integrating amplified supervision, robust training, interpretability analysis, and uncertainty modeling; and (2) at the system level, implementing multi-tiered access control and real-time behavioral monitoring. Contribution/Results: This work is the first to systematically categorize and prioritize four risk types—misuse, misalignment, mistakes, and structural flaws—focusing explicitly on the former two. It introduces the first verifiable safety case framework for AGI, treating interpretability and uncertainty estimation as proactive enablers of safety assurance. The resulting end-to-end methodology comprehensively covers capability identification, safety hardening, dynamic monitoring, and failure containment—thereby enabling high-assurance, auditable, and formally verifiable AGI safety engineering.

Addressing misuse risks in AGI through security and access controlCombining techniques for robust AGI safety and security casesMitigating misalignment risks via model-level and system-level defenses

Dimensional Characterization and Pathway Modeling for Catastrophic AI Risks

Aug 08, 2025
ZS
Ze Shen Chin
🏛️ Oxford Martin AI Governance Initiative | AI Standards Lab

Current AI risk research lacks a systematic framework and rigorous causal modeling of pathways from hazardous AI capabilities to real-world harms. This paper addresses six categories of catastrophic AI risks by proposing the first seven-dimensional risk characterization framework—spanning intent, capability, agent type, and other critical dimensions—and introducing a stepwise causal pathway model (“Hazard → Harm → Consequence”). Methodologically, it integrates multidimensional feature analysis with formal causal path modeling to enable computationally tractable representation of risk evolution. The contributions are threefold: (1) a scalable, structured analytical framework for AI catastrophe risk assessment; (2) a decision-support tool that jointly enables general-purpose mitigation strategies and scenario-specific interventions; and (3) a theoretical and operational foundation for end-to-end AI risk governance across the full value chain.

Identify systematic risk mitigation strategiesLack comprehensive framework for AI catastrophic risksNeed mapping hazard to harm causal pathways

Latest Papers

What's happening recently
View more

This work addresses the inadequacy of existing large language model (LLM) lifecycle frameworks, which predominantly emphasize operational efficiency while lacking explicit support for security-critical activities—such as data provenance, component signing, and access control—and failing to align governance requirements with specific lifecycle phases. The paper proposes the first security-oriented LLM system lifecycle model, structured not by workflow but by security boundaries, organizing 32 phases into four layered pipelines: data, model, distribution, and application, while integrating LLMOps and governance pillars. It uniquely identifies 13 distinct security-critical phases and exposes a structural imbalance wherein regulatory evidence is concentrated at deployment despite pivotal decisions occurring during development. By mapping key standards—including NIST AI RMF, the EU AI Act, and ISO/IEC 42001—the study establishes a phase-to-governance correspondence mechanism, yielding a comprehensive, lifecycle-spanning security analysis framework that offers structured guidance for compliance and secure design.

governance frameworklarge language modelsLLM systems

This study addresses the multidimensional risks—operational, security, and governance-related—that enterprises face when deploying large language models, noting that existing open-source tools are fragmented and fail to comprehensively cover authoritative risk taxonomies. To bridge this gap, the work proposes a structured mapping protocol that automatically aligns the capabilities of 21 prominent open-source tools with the 32 subcategories of the MIT AI Risk Framework, leveraging retrieval-augmented generation (RAG) and LLM-based parsing. The protocol’s validity is substantiated through source code and documentation analysis, majority voting, and inter-rater reliability assessment using Fleiss’ Kappa (κ = 0.509, F1 = 75.5%). Findings reveal a pronounced overconcentration of current tools on technical controls, with significant gaps in governance, legal, and market risk domains, thereby providing an empirical foundation for developing layered AI risk mitigation architectures.

AI risk mitigationgovernancelarge language models

This study addresses the systemic risks posed by on-premises AI coding agents, whose autonomous modifications to code and infrastructure may lead to severe organizational or societal harm due to challenges in timely constraint, auditing, or reversal. For the first time, it systematically applies three systems safety methodologies—STECA, STPA, and FRAM—to model risks in cutting-edge laboratory settings from multiple perspectives. The analysis reveals critical blind spots in current AI governance frameworks, particularly concerning unverifiable accountability, control failure caused by intervention delays, and weakened safeguards due to operational drift. The work underscores the necessity of integrating model-level evaluations with system-level hazard analysis, offering a crucial complementary pathway for robust AI risk management.

Agentic AIAI Risk ManagementLoss of Control

Hot Scholars

WX

Wenrui Xu

University of Minnesota
Machine LearningHyperdimensional Computing
AS

Abhinav Sinha

Guidance, Autonomy, Learning, and Control for Intelligent Systems Lab; University of Cincinnati
Guidance and ControlReinforcement LearningMultiagent systemsNetworked Control Systems
BC

Benjamin C. M. Fung

Canada Research Chair & Professor, School of Information Studies, McGill University
Data miningmachine learningdata privacyauthorship analysis
JH

Ji He

Guangzhou Medical University
CT Image ReconstructionDeep Learning
YC

Yongcan Cao

UT San Antonio
autonomous systemsroboticscyber-physical systemshuman-robot interaction