Score
Designs, builds, and evaluates automated and programmatic constraints, configurations, and policies that prevent unsafe, insecure, or non‑compliant behaviors in AI models and systems, including model- and LLM-specific limits and runtime enforcement. Work includes specifying and implementing behavioral, security, and governance rules; runtime monitors, checks, and fallback actions; access and output controls and sanitization; deployment and configuration of guardrail infrastructure; and logging, alerting, audit, and incident‑response mechanisms.
To address runtime safety failures in large language model (LLM)-based agents—arising from autonomy and non-determinism in multi-step decision-making, goal planning, and tool invocation—this paper proposes a dynamic, hierarchical, cross-phase safety assurance framework. Methodologically, it innovatively adapts the Swiss Cheese Model to establish an AI safety reference architecture; introduces the first runtime protection taxonomy spanning quality attributes, pipeline stages, and architectural components; and formalizes an AI safety-by-design software architecture paradigm. Through systematic literature review, architectural modeling, and multi-layered defense design, the framework enables real-time monitoring and intervention over critical artifacts—including goals, plans, and tools. The contributions include a reusable classification system and a structured design guideline, collectively supporting the development of robust, verifiable, and evolvable safety-critical systems for foundation model agents. (149 words)
Autonomous AI agents deployed in industrial settings face significant challenges in ensuring runtime safety and regulatory compliance. Method: This paper introduces the “Policy-as-Prompt” paradigm, which automatically transforms unstructured design documentation into verifiable, auditable, real-time safety guardrails. Leveraging large language models, the approach parses technical documents to extract security policy semantics, enforces least-privilege constraints, constructs structured policy trees, and compiles them into lightweight, prompt-driven classifiers for low-overhead behavioral auditing. Results: Experiments demonstrate scalability and auditability across diverse industrial scenarios, effectively bridging the gap between policy formulation and enforcement. The key contribution is the first end-to-end automated translation of natural-language security policies into formally verifiable, runtime-enforceable guardrails—establishing a novel AI governance framework that jointly ensures security, regulatory compliance, and interpretability.
This study addresses the challenge of translating high-level governance standards into enforceable runtime safeguards for agentic AI systems, whose multi-step external actions complicate direct application of conventional norms. To bridge this gap, the authors propose a hierarchical translation framework that systematically maps governance objectives from standards such as ISO/IEC and NIST across four levels: governance goals, design-time constraints, runtime mediators, and assurance feedback. The framework explicitly distinguishes governance intent, technical controls, runtime protections, and evidentiary assurance, introducing control tuples and runtime executability criteria to guide the appropriate architectural placement of safeguards. Validation through a procurement agent case study demonstrates that only observable, deterministic, and time-sensitive controls are suitable for runtime enforcement, effectively reconciling normative requirements with practical system implementation.
This work proposes a modular and configurable safety framework to address critical security and ethical risks associated with large language models (LLMs), including privacy leakage, generation of misinformation, and malicious misuse. The framework employs an adaptive sequence scheduling mechanism to dynamically integrate trustworthy components—such as content filtering, behavioral constraints, and context-aware controls—into a flexible guardrail architecture. This approach enables real-time, context-sensitive ethical and safety oversight of model outputs, effectively mitigating the generation of harmful, factually inaccurate, or privacy-sensitive content. By doing so, it significantly enhances the safety, regulatory compliance, and deployment adaptability of LLMs in real-world applications.
This work addresses the security risks posed by large language model (LLM) agents during tool invocation, such as inadvertent leakage of sensitive data or overwriting of critical records—hazards for which existing approaches lack verifiable guarantees. To bridge this gap, the paper introduces a novel integration of System-Theoretic Process Analysis (STPA) with formal specifications to systematically identify hazards in agent workflows and derive enforceable safety requirements. These requirements are then translated into executable constraints on data flows and tool invocation sequences. Building upon an enhanced Model Context Protocol (MCP) framework, the approach incorporates structured capability control and trust-labeling mechanisms to enable proactive, verifiable protection of tool interactions. By significantly reducing reliance on manual verification, this method advances LLM agent design from empirical reliability toward a paradigm grounded in formal security assurances.
Traditional text-level safety mechanisms fail to adequately address security risks arising from LLM agent behaviors. To bridge this gap, this paper proposes GuardAgent—the first dynamic, agent-level safety guard framework. Its core comprises knowledge-enhanced, two-stage LLM reasoning (security requirement parsing → plan-to-code mapping) integrated with memory-augmented contextual retrieval, enabling real-time behavioral verification of target agents and generation of lightweight, executable code-based safeguards. We introduce a novel agent-level safety guarding paradigm and establish two domain-specific, rigorously designed benchmarks: EICU-AC (for healthcare access control) and Mind2Web-SC (for secure web interaction). Experimental results demonstrate that GuardAgent achieves 98% and 83% guard accuracy on these benchmarks, respectively, effectively suppressing policy violations while maintaining high flexibility, low computational overhead, and strong generalization across diverse agent tasks and environments.
Enterprise-scale general-purpose agents lack built-in, reusable governance mechanisms for autonomous cross-tool operation, making it difficult to satisfy requirements for compliance, auditability, and behavioral controllability. This work proposes the CUGA policy system, which embeds runtime governance capabilities into five critical checkpoints of the agent execution pipeline—intent protection, playbook guidance, tool invocation control, human approval gating, and output formatting—through a modular “policy-as-code” architecture. Without requiring model fine-tuning, CUGA enables proactive, continuous, and structured behavior control. By integrating typed governance primitives, dynamic playbook injection, and human-in-the-loop approval, the system effectively blocks malicious requests, enforces structured tool sequences, and triggers manual review for high-risk operations in healthcare scenarios, significantly enhancing policy adherence, execution consistency, and deployment safety.
This work addresses the challenge of enforcing temporal safety constraints throughout the lifecycle of black-box AI systems, such as large language models (LLMs), which are inherently difficult to verify. The paper proposes the first offline auditing and online monitoring framework that integrates Linear Temporal Logic (LTL) with machine learning. This framework enables formal verification of complex temporal behavioral specifications by introducing a sampling-driven predictive monitor and an intervenable runtime monitor, effectively preventing policy violations. Experimental results demonstrate that the proposed approach significantly outperforms existing LLM-based evaluators in detecting temporal violations. Notably, it achieves performance on par with or superior to state-of-the-art large models using only a small labeled model, while its intervention mechanism substantially reduces violation rates without compromising task performance.
This work addresses the semantic gap faced by current AI agents in enforcing natural language policies: the intended policy semantics are difficult to enforce precisely and interpretably at the system level. To bridge this gap, the authors propose a novel approach that integrates agent-side context with kernel-level enforcement mechanisms. For the first time, policy context is preserved on the agent side, while a domain-specific language (DSL) for information flow control (IFC), implemented via eBPF, enables comprehensive, action-level policy enforcement within the operating system kernel. This framework supports cross-event data-flow and ordering constraints, significantly improving policy compliance rates by covering indirect execution paths invisible to conventional tool-call interception. The system incurs only 1.9%–8.4% runtime overhead and provides semantically clear feedback instead of ambiguous errors.