Score
Designs and builds threat models that identify and characterize harmful capabilities, failure modes, and misuse pathways of autonomous or agentic systems, and maps those threats to applicable legal and regulatory obligations. Produces prioritized control sets and concrete, actionable mitigation requirements by analyzing how regulations amplify or constrain risks and translating regulatory obligations into technical, operational, and governance controls.
Rapid deployment of AI systems in regulated domains exposes a critical gap between technical security and legal compliance, as algorithmic vulnerabilities (e.g., those cataloged in MITRE ATLAS) lack systematic mapping to quantifiable financial impacts—undermining evidence-based decisions on contingency reserves and cyber-insurance pricing. Method: We propose the first cross-domain AI threat vector classification framework that directly links technical threats to five business loss dimensions: confidentiality, integrity, availability, legal liability, and reputational harm. Leveraging structured ontology modeling, we integrate MITRE ATLAS, the EU AI Act, NIST AI Risk Management Framework, and ISO/IEC 42001 to define 53 actionable sub-threats. Contribution/Results: The framework achieves 100% coverage across 133 real-world AI incidents reported in 2025, demonstrating both conceptual completeness and audit readiness for regulatory and economic risk assessment.
This study addresses the lack of systematic analysis and actionable controls linking large language model (LLM) agent security threats to real-world financial regulations. It presents the first mapping of six categories of LLM agent risks to regulatory obligations in the U.S. and EU, and proposes four scalable compliance architectures centered on auditability, authorization, and boundary enforcement. Key technical contributions include agent-to-agent (A2A) compliance orchestration, audit-driven Grounded-RAG, case ID propagation, and reasoning-boundary de-identification proxies. Empirical evaluation demonstrates that the framework reduces manual processing from multiple days to same-day handling, automates approximately 80% of use cases, and uncovers two types of control failures detectable only through internal audit, as well as one category of legitimate applicants erroneously rejected.
To address the lack of provable behavioral guarantees for large-scale autonomous AI systems under adversarial attacks and operational stress, this paper proposes the first engineering-grade safety and trustworthiness assurance framework spanning the entire system lifecycle—design, training, deployment, and runtime operation. Methodologically, it innovatively integrates standardized threat modeling with quantitative risk assessment, adversarial robustness training, lightweight real-time anomaly detection, automated audit logging, and compliance verification protocols into a unified assurance pipeline. Key contributions include: (1) proactive, risk-aware assurance embedded early in the development cycle; (2) security-by-design, wherein safety properties are intrinsically encoded into model architecture; and (3) formally verifiable and mathematically provable system behavior. Experimental evaluation demonstrates significant reductions in vulnerability rates and compliance overhead across national security, open-model governance, and industrial automation domains, confirming strong scalability and cross-domain applicability.
Scientific research cyberinfrastructure (CI) faces unique challenges—including high collaboration requirements, component heterogeneity, and the absence of adaptable security assessment frameworks. To address these, we propose a mission-centric security posture assessment method: first, top-down identification of critical assets and unacceptable losses; second, construction of a security knowledge graph integrating system components, dependencies, and threat behaviors; and third, integration with directed attack graphs to quantify multi-hop attack paths from entry points to critical assets—enabling visualization of attacker-defender relationships and identification of security blind spots. Unlike conventional generic standards, our approach is the first to deeply couple mission-driven assessment, knowledge graphs, and attack graphs. It supports risk prioritization and generation of actionable defensive strategies, significantly enhancing the precision and effectiveness of CI security defense.
Prior to large-scale AGI deployment, misuse and goal misalignment represent two critical safety risks. Method: We propose a dual-track collaborative defense framework: (1) at the model level, integrating amplified supervision, robust training, interpretability analysis, and uncertainty modeling; and (2) at the system level, implementing multi-tiered access control and real-time behavioral monitoring. Contribution/Results: This work is the first to systematically categorize and prioritize four risk types—misuse, misalignment, mistakes, and structural flaws—focusing explicitly on the former two. It introduces the first verifiable safety case framework for AGI, treating interpretability and uncertainty estimation as proactive enablers of safety assurance. The resulting end-to-end methodology comprehensively covers capability identification, safety hardening, dynamic monitoring, and failure containment—thereby enabling high-assurance, auditable, and formally verifiable AGI safety engineering.
Current AI risk research lacks a systematic framework and rigorous causal modeling of pathways from hazardous AI capabilities to real-world harms. This paper addresses six categories of catastrophic AI risks by proposing the first seven-dimensional risk characterization framework—spanning intent, capability, agent type, and other critical dimensions—and introducing a stepwise causal pathway model (“Hazard → Harm → Consequence”). Methodologically, it integrates multidimensional feature analysis with formal causal path modeling to enable computationally tractable representation of risk evolution. The contributions are threefold: (1) a scalable, structured analytical framework for AI catastrophe risk assessment; (2) a decision-support tool that jointly enables general-purpose mitigation strategies and scenario-specific interventions; and (3) a theoretical and operational foundation for end-to-end AI risk governance across the full value chain.
This work addresses the inadequacy of existing large language model (LLM) lifecycle frameworks, which predominantly emphasize operational efficiency while lacking explicit support for security-critical activities—such as data provenance, component signing, and access control—and failing to align governance requirements with specific lifecycle phases. The paper proposes the first security-oriented LLM system lifecycle model, structured not by workflow but by security boundaries, organizing 32 phases into four layered pipelines: data, model, distribution, and application, while integrating LLMOps and governance pillars. It uniquely identifies 13 distinct security-critical phases and exposes a structural imbalance wherein regulatory evidence is concentrated at deployment despite pivotal decisions occurring during development. By mapping key standards—including NIST AI RMF, the EU AI Act, and ISO/IEC 42001—the study establishes a phase-to-governance correspondence mechanism, yielding a comprehensive, lifecycle-spanning security analysis framework that offers structured guidance for compliance and secure design.
This study addresses the multidimensional risks—operational, security, and governance-related—that enterprises face when deploying large language models, noting that existing open-source tools are fragmented and fail to comprehensively cover authoritative risk taxonomies. To bridge this gap, the work proposes a structured mapping protocol that automatically aligns the capabilities of 21 prominent open-source tools with the 32 subcategories of the MIT AI Risk Framework, leveraging retrieval-augmented generation (RAG) and LLM-based parsing. The protocol’s validity is substantiated through source code and documentation analysis, majority voting, and inter-rater reliability assessment using Fleiss’ Kappa (κ = 0.509, F1 = 75.5%). Findings reveal a pronounced overconcentration of current tools on technical controls, with significant gaps in governance, legal, and market risk domains, thereby providing an empirical foundation for developing layered AI risk mitigation architectures.
This study addresses the systemic risks posed by on-premises AI coding agents, whose autonomous modifications to code and infrastructure may lead to severe organizational or societal harm due to challenges in timely constraint, auditing, or reversal. For the first time, it systematically applies three systems safety methodologies—STECA, STPA, and FRAM—to model risks in cutting-edge laboratory settings from multiple perspectives. The analysis reveals critical blind spots in current AI governance frameworks, particularly concerning unverifiable accountability, control failure caused by intervention delays, and weakened safeguards due to operational drift. The work underscores the necessity of integrating model-level evaluations with system-level hazard analysis, offering a crucial complementary pathway for robust AI risk management.