Score
Designs, builds, and analyzes systems, frameworks, and tooling that implement and automate policy rules—including rule-based, referential-integrity, safety, and security policies—via enforcement engines, mechanisms, workflows, and coordination to provide scalable oversight. Implements monitoring and observability for those enforcement systems by defining and collecting logs, metrics, traces, alerts, and audit trails and by designing dashboards and diagnostics to measure enforcement effectiveness, detect failures or policy violations, and support incident response and continuous improvement.
Enterprise-scale general-purpose agents lack built-in, reusable governance mechanisms for autonomous cross-tool operation, making it difficult to satisfy requirements for compliance, auditability, and behavioral controllability. This work proposes the CUGA policy system, which embeds runtime governance capabilities into five critical checkpoints of the agent execution pipeline—intent protection, playbook guidance, tool invocation control, human approval gating, and output formatting—through a modular “policy-as-code” architecture. Without requiring model fine-tuning, CUGA enables proactive, continuous, and structured behavior control. By integrating typed governance primitives, dynamic playbook injection, and human-in-the-loop approval, the system effectively blocks malicious requests, enforces structured tool sequences, and triggers manual review for high-risk operations in healthcare scenarios, significantly enhancing policy adherence, execution consistency, and deployment safety.
This work addresses the semantic gap faced by current AI agents in enforcing natural language policies: the intended policy semantics are difficult to enforce precisely and interpretably at the system level. To bridge this gap, the authors propose a novel approach that integrates agent-side context with kernel-level enforcement mechanisms. For the first time, policy context is preserved on the agent side, while a domain-specific language (DSL) for information flow control (IFC), implemented via eBPF, enables comprehensive, action-level policy enforcement within the operating system kernel. This framework supports cross-event data-flow and ordering constraints, significantly improving policy compliance rates by covering indirect execution paths invisible to conventional tool-call interception. The system incurs only 1.9%–8.4% runtime overhead and provides semantically clear feedback instead of ambiguous errors.
Large language model (LLM) agents exhibit unreliable adherence to corporate policies in business process automation. To address this, we propose a deterministic, transparent, and modular policy compliance framework comprising two phases: (1) an offline phase that compiles natural-language policy documents into verifiable guard code, and (2) a runtime phase that inserts lightweight, policy-agnostic guards before tool invocation—thereby decoupling policy enforcement from agent logic. This design enhances interpretability, maintainability, and agility in policy updates. Experiments on the τ-bench Airlines testbed demonstrate the framework’s effectiveness in intercepting policy-violating actions, validating its feasibility. However, empirical evaluation also uncovers critical deployment challenges, including incompleteness in policy coverage and difficulties in dynamically adapting guards to contextual changes. The framework thus advances policy-aware LLM agent deployment while surfacing key open issues for future work.
This study addresses the lack of systematic comparative analysis in business process compliance monitoring, particularly for non-conformance checking techniques. Through a systematic literature review (SLR), process mining, compliance modeling, and qualitative comparative analysis, it maps real-world applications across domains, operational workflows, technical foundations, and result representations. The analysis identifies key implementation barriers—especially pervasive human dependence and the absence of standardized evaluation criteria. As the first structured survey framework dedicated to non-conformance checking, the study introduces a standardized, multi-dimensional evaluation framework that clarifies commonalities and distinctions across the technical landscape. It further proposes an extensible theoretical pathway and practical guidelines for automated compliance monitoring. This work provides a methodological foundation and strategic direction for both academic research and industrial deployment. (149 words)
This work addresses the lack of runtime-enforceable mechanisms for natural language behavioral policies in AI agents when invoking third-party skills, which hinders fine-grained monitoring and intervention against policy violations across actions. The authors propose VIGIL, an end-to-end runtime enforcement framework that translates skill specifications, operator constraints, and global rules into executable policies through a context-aware policy language and symbolic evaluation. VIGIL dynamically detects violations along agent execution traces by enabling fine-grained reasoning over event orderings, parameter relationships, and cross-call data flows, thereby overcoming limitations of fixed event models. Leveraging SMT-based symbolic execution and bounded trajectory constraints, VIGIL achieves efficient runtime monitoring and intervention. Empirical evaluation demonstrates that the approach attains over 95% violation detection accuracy with a false positive rate below 10% in real-world LLM agent tasks.
This work addresses the challenge of control-flow violations that arise when large language model (LLM) agents automate high-judgment quality management processes in regulated industries, often due to insufficient integration of symbolic structures such as regulatory rules and typed process models. To overcome this limitation, the paper introduces a “compliance-by-construction” paradigm, which internalizes compliance constraints as core components of the agent architecture rather than relying solely on external guardrails. By synergistically combining LLMs with symbolic systems—integrating typed process models, formal compliance constraints, and neuro-symbolic reasoning—the approach structurally prevents violations while preserving the ability to detect semantic errors. The study also systematically delineates the foundational and capability-level challenges required to realize this paradigm, offering a viable neuro-symbolic pathway for automation in regulation-intensive domains.
Existing runtime enforcement techniques struggle to handle reactive systems with complex continuous dynamics and lack effective mechanisms for intervening in hybrid behaviors. This work proposes the first framework that integrates hybrid automata into runtime enforcement, enabling coordinated discrete event editing and continuous-time monitoring to correct system behavior at any instant by suppressing, delaying, or inserting events. The paper establishes formal enforceability conditions and devises an online strategy synthesis algorithm based on reachability analysis. Evaluation on an adaptive cruise control case study demonstrates that the approach ensures safety properties even when the underlying controller is unsafe, all while incurring minimal computational overhead.