Score
Designs and documents the formal policies that govern autonomous agent behavior, specifying rules, constraints, permissions, priorities, failure modes, and safety or ethical guardrails. Implements and iterates policy representations (for example rule sets, decision trees, reward/penalty structures, access-control lists, and fallback procedures) that agent controllers and orchestration systems use to make decisions and handle exceptions.
To address runtime safety failures in large language model (LLM)-based agents—arising from autonomy and non-determinism in multi-step decision-making, goal planning, and tool invocation—this paper proposes a dynamic, hierarchical, cross-phase safety assurance framework. Methodologically, it innovatively adapts the Swiss Cheese Model to establish an AI safety reference architecture; introduces the first runtime protection taxonomy spanning quality attributes, pipeline stages, and architectural components; and formalizes an AI safety-by-design software architecture paradigm. Through systematic literature review, architectural modeling, and multi-layered defense design, the framework enables real-time monitoring and intervention over critical artifacts—including goals, plans, and tools. The contributions include a reusable classification system and a structured design guideline, collectively supporting the development of robust, verifiable, and evolvable safety-critical systems for foundation model agents. (149 words)
Autonomous AI agents deployed in industrial settings face significant challenges in ensuring runtime safety and regulatory compliance. Method: This paper introduces the “Policy-as-Prompt” paradigm, which automatically transforms unstructured design documentation into verifiable, auditable, real-time safety guardrails. Leveraging large language models, the approach parses technical documents to extract security policy semantics, enforces least-privilege constraints, constructs structured policy trees, and compiles them into lightweight, prompt-driven classifiers for low-overhead behavioral auditing. Results: Experiments demonstrate scalability and auditability across diverse industrial scenarios, effectively bridging the gap between policy formulation and enforcement. The key contribution is the first end-to-end automated translation of natural-language security policies into formally verifiable, runtime-enforceable guardrails—establishing a novel AI governance framework that jointly ensures security, regulatory compliance, and interpretability.
Enterprise-scale general-purpose agents lack built-in, reusable governance mechanisms for autonomous cross-tool operation, making it difficult to satisfy requirements for compliance, auditability, and behavioral controllability. This work proposes the CUGA policy system, which embeds runtime governance capabilities into five critical checkpoints of the agent execution pipeline—intent protection, playbook guidance, tool invocation control, human approval gating, and output formatting—through a modular “policy-as-code” architecture. Without requiring model fine-tuning, CUGA enables proactive, continuous, and structured behavior control. By integrating typed governance primitives, dynamic playbook injection, and human-in-the-loop approval, the system effectively blocks malicious requests, enforces structured tool sequences, and triggers manual review for high-risk operations in healthcare scenarios, significantly enhancing policy adherence, execution consistency, and deployment safety.
Existing requirements engineering approaches fail to explicitly define the scope of delegated decision-making, authorization hierarchies, oversight mechanisms, and control-relinquishment protocols in AI agent systems, often leaving critical requirements implicit in prompts or runtime policies. This work proposes the first requirements engineering framework tailored for autonomous agents, introducing the novel concept of “delegation autonomy boundaries” and modeling authority as a hierarchical structure. The framework specifies delegation behavior across six dimensions—purpose, authority, information, coordination, assurance, and evolution—through two complementary artifacts: Agent Justification Records (AJRs) and Agent Delegation Policies (ADPs). Empirical validation in hospital discharge coordination and automated code review scenarios demonstrates the framework’s effectiveness in enabling clear definition and management of delegation boundaries for both safety-critical and routine tasks.
This work addresses the lack of comprehensive governance mechanisms—such as permissions, prohibitions, obligations, exemptions, and policy conflict resolution—in existing autonomous agent systems for cross-organizational collaboration. To bridge this gap, the authors propose AgenticRei, a novel framework that, for the first time, integrates obligation lifecycle management, context-aware exemptions, and policy conflict resolution within a unified runtime governance architecture. The framework formalizes deontic policies using an OWL ontology and leverages the Rei policy language coupled with a high-performance logical reasoning engine to enable dynamic policy inference. AgenticRei seamlessly aligns with industry standards like A2AS and successfully enforces security and privacy constraints that current production-grade policy engines struggle to handle, thereby significantly enhancing both the expressiveness and practical feasibility of agent governance in complex, real-world scenarios.
This paper addresses the challenge policymakers face in predicting agent behavior under varying degrees of norm compliance. We propose a norm-aware agent architecture capable of dynamically switching behavioral modes. Methodologically, we are the first to embed the AOPL normative language into a multimodal planning simulation framework, enabling human operators to adjust agent compliance levels online and transparently—e.g., shifting in real time from strict norm adherence to high-risk operation. By integrating logic programming with authorization/obligation modeling, we achieve semantically grounded, causally traceable behavior policies. Our contributions are threefold: (1) the first online, interpretable mechanism for behavioral mode switching; (2) a prototype system tailored to time-critical domains such as emergency response; and (3) a norm-semantics–enabled simulation tool for policy design, significantly enhancing decision predictability and accountability.
本文提出Brain API,一种意图感知控制平面,通过决策工件解决现有系统中缺乏意图级决策治理的问题。
This work addresses a critical gap in current AI systems, which often conflate technical capability with operational authority, resulting in inadequate governance of authorized autonomy. The paper proposes a structured governance framework that systematically distinguishes between an AI system’s Autonomous Capability Level (ACL) and its Authorized Autonomy Level (AAL). By integrating risk exposure, action reversibility, and accountability, the framework introduces a dynamic authorization mechanism that decouples capability from permission. Validated in enterprise-grade data engineering agents, the approach enables high-capability systems to be safely constrained to lower authorization levels aligned with organizational risk tolerance. Through layered autonomy modeling and risk-aware decision protocols, the framework ensures that AI autonomy remains both effective and responsibly governed.
This study reframes multi-agent AI safety as a problem of institutional design, focusing on security mechanisms in task delegation, information flow, action execution, and resource sharing. Through structured workflow experiments, it systematically links institutional components—such as rules, authority states, and fallback pathways—to emergent safety behaviors in multi-agent systems, revealing distinct failure modes under identical violation rates. The methodology integrates constitutional prompting, provenance-aware protection, and local state guards, evaluated across 5,280 high-conflict episodes involving multiple model families. Results show that both constitutional prompting and provenance-aware protection achieve zero violations out of 384 trials; the latter successfully blocks 51 attempted violations and ensures 44 subsequent safe completions. In contrast, local state guards exhibit significant failure under policy-change scenarios.
This work addresses the limitation of existing autonomous agents, which typically ensure compliance only at the level of individual actions and struggle to maintain system-wide safety constraints throughout an entire task. To overcome this, the paper proposes a verifiable safety architecture spanning the agent’s full stack, embedding safety into its design, interaction protocols, and runtime mechanisms. The framework establishes a multi-layered assurance system—from input sanitization and multi-agent trust establishment to behavioral trajectory verification—by integrating formal verification, runtime monitoring, identity and capability control, model provenance, and end-to-end observability. This approach enables behavior containment that extends from pointwise safeguards to trajectory-level constraints, offering a scalable and verifiable deployment pathway for large language model–driven autonomous agents and effectively addressing complex security challenges in cross-organizational settings.
本文针对企业部署自主AI代理时遇到的控制模型不匹配问题,提出了五种运行时治理原语,并通过实现这些原语来解决身份、治理等问题。