Score
Designs and implements formal representations of human constraints and exception behaviors for automated systems, including explicit error‑handling rules and execution policies. Builds or analyzes mechanisms that translate workflows and human rules into enforceable policies for agents and training processes so systems follow prioritized safe behaviors.
To address runtime safety failures in large language model (LLM)-based agents—arising from autonomy and non-determinism in multi-step decision-making, goal planning, and tool invocation—this paper proposes a dynamic, hierarchical, cross-phase safety assurance framework. Methodologically, it innovatively adapts the Swiss Cheese Model to establish an AI safety reference architecture; introduces the first runtime protection taxonomy spanning quality attributes, pipeline stages, and architectural components; and formalizes an AI safety-by-design software architecture paradigm. Through systematic literature review, architectural modeling, and multi-layered defense design, the framework enables real-time monitoring and intervention over critical artifacts—including goals, plans, and tools. The contributions include a reusable classification system and a structured design guideline, collectively supporting the development of robust, verifiable, and evolvable safety-critical systems for foundation model agents. (149 words)
Enterprise-scale general-purpose agents lack built-in, reusable governance mechanisms for autonomous cross-tool operation, making it difficult to satisfy requirements for compliance, auditability, and behavioral controllability. This work proposes the CUGA policy system, which embeds runtime governance capabilities into five critical checkpoints of the agent execution pipeline—intent protection, playbook guidance, tool invocation control, human approval gating, and output formatting—through a modular “policy-as-code” architecture. Without requiring model fine-tuning, CUGA enables proactive, continuous, and structured behavior control. By integrating typed governance primitives, dynamic playbook injection, and human-in-the-loop approval, the system effectively blocks malicious requests, enforces structured tool sequences, and triggers manual review for high-risk operations in healthcare scenarios, significantly enhancing policy adherence, execution consistency, and deployment safety.
This work addresses the challenge of control-flow violations that arise when large language model (LLM) agents automate high-judgment quality management processes in regulated industries, often due to insufficient integration of symbolic structures such as regulatory rules and typed process models. To overcome this limitation, the paper introduces a “compliance-by-construction” paradigm, which internalizes compliance constraints as core components of the agent architecture rather than relying solely on external guardrails. By synergistically combining LLMs with symbolic systems—integrating typed process models, formal compliance constraints, and neuro-symbolic reasoning—the approach structurally prevents violations while preserving the ability to detect semantic errors. The study also systematically delineates the foundational and capability-level challenges required to realize this paradigm, offering a viable neuro-symbolic pathway for automation in regulation-intensive domains.
Large language model (LLM) agents exhibit unreliable adherence to corporate policies in business process automation. To address this, we propose a deterministic, transparent, and modular policy compliance framework comprising two phases: (1) an offline phase that compiles natural-language policy documents into verifiable guard code, and (2) a runtime phase that inserts lightweight, policy-agnostic guards before tool invocation—thereby decoupling policy enforcement from agent logic. This design enhances interpretability, maintainability, and agility in policy updates. Experiments on the τ-bench Airlines testbed demonstrate the framework’s effectiveness in intercepting policy-violating actions, validating its feasibility. However, empirical evaluation also uncovers critical deployment challenges, including incompleteness in policy coverage and difficulties in dynamically adapting guards to contextual changes. The framework thus advances policy-aware LLM agent deployment while surfacing key open issues for future work.
Existing human–robot collaboration systems struggle to interpret complex human behaviors—such as task deviation and deception—in real time due to the absence of rigorous human behavioral modeling and the inability of descriptive ontologies to support efficient, runtime reasoning. Method: This paper introduces the first behavior interpretation ontology framework that integrates cognitive science principles with modular engineering design. It overcomes dual limitations: (i) technology-centric robot architectures lacking human behavior modeling capabilities, and (ii) static ontologies unsuitable for real-time semantic inference. The framework establishes a machine-processable, semantically tractable, and architecturally extensible behavior taxonomy, enabled by formal ontology modeling, lightweight semantic reasoning, and modular system design. Contribution/Results: Evaluated in heterogeneous domains—industrial manufacturing and interactive gaming—the framework achieves significant improvements in behavior recognition accuracy and collaborative safety. It provides a deployable theoretical foundation and technical infrastructure for trustworthy human–machine coexistence in Industry 5.0.
This work addresses the absence of a formal specification language that clearly delineates responsibility boundaries, approval checkpoints, and governance constraints between humans and AI agents throughout the software development lifecycle (SDLC). To this end, we propose a domain-specific protocol language tailored for AI-augmented SDLCs, which decouples policy intent from mechanistic implementation through a formal grammar, well-formedness conditions, operational semantics, and execution invariants to mitigate collaborative uncertainty. Our approach innovatively formalizes the principle of separation of duties as a 2+N team model and natively integrates Kleene closure with protocol self-consistency verification. Theoretical analysis demonstrates that structured execution reduces system failure rates to the weighted product of agent and verifier failure probabilities. A prototype implementation confirms the feasibility and effectiveness of the proposed method.
This study addresses the challenge of deploying agentic AI in regulated environments, where existing approaches lack a systematic design framework that jointly accounts for autonomy and agency, often failing to balance compliance, auditability, and error correction. The work introduces the first unified model of these two dimensions, defining a two-dimensional hierarchical design space with five operational levels each. It proposes six architectural strategies—checkpoints, escalation mechanisms, multi-agent delegation, tool provisioning, tool sandboxing, and write staging—to enable flexible system configuration under real-world regulatory constraints. Validated through public-sector case studies, the framework establishes a shared terminology and actionable design guidelines, facilitating interpretable, controllable, and compliant AI deployment amid evolving model capabilities and tool fidelity.
This work addresses the lack of runtime-enforceable mechanisms for natural language behavioral policies in AI agents when invoking third-party skills, which hinders fine-grained monitoring and intervention against policy violations across actions. The authors propose VIGIL, an end-to-end runtime enforcement framework that translates skill specifications, operator constraints, and global rules into executable policies through a context-aware policy language and symbolic evaluation. VIGIL dynamically detects violations along agent execution traces by enabling fine-grained reasoning over event orderings, parameter relationships, and cross-call data flows, thereby overcoming limitations of fixed event models. Leveraging SMT-based symbolic execution and bounded trajectory constraints, VIGIL achieves efficient runtime monitoring and intervention. Empirical evaluation demonstrates that the approach attains over 95% violation detection accuracy with a false positive rate below 10% in real-world LLM agent tasks.