Score
Designs and implements runtime enforcement systems that apply unified policy rules to LLM model calls—intercepting requests, monitoring model compliance, logging policy decisions, and blocking or modifying outputs that violate rules. Builds adaptive enforcement components that operate under real-time and latency/deadline constraints, including scheduler-weight adaptation and lightweight xapp / xapp-controller enforcement executed at scheduling or control intervals.
This work addresses the challenge of ensuring that tool-augmented large language models (TaLLMs) adhere to domain-specific operational policies in sensitive scenarios. To this end, the authors propose a runtime policy compliance verification framework that, for the first time, translates natural-language policies into formal SMT-LIB constraints through human-in-the-loop collaboration. The framework integrates the Z3 solver to perform real-time validation of tool invocation parameters against environmental states, immediately blocking any action that violates prescribed policies. Experimental evaluation on the TauBench benchmark demonstrates that this approach substantially reduces policy violation rates while preserving task accuracy, thereby providing TaLLMs with strong formal guarantees of compliant behavior.
Large language model (LLM) agents exhibit unreliable adherence to corporate policies in business process automation. To address this, we propose a deterministic, transparent, and modular policy compliance framework comprising two phases: (1) an offline phase that compiles natural-language policy documents into verifiable guard code, and (2) a runtime phase that inserts lightweight, policy-agnostic guards before tool invocation—thereby decoupling policy enforcement from agent logic. This design enhances interpretability, maintainability, and agility in policy updates. Experiments on the τ-bench Airlines testbed demonstrate the framework’s effectiveness in intercepting policy-violating actions, validating its feasibility. However, empirical evaluation also uncovers critical deployment challenges, including incompleteness in policy coverage and difficulties in dynamically adapting guards to contextual changes. The framework thus advances policy-aware LLM agent deployment while surfacing key open issues for future work.
Existing runtime enforcement techniques struggle to handle reactive systems with complex continuous dynamics and lack effective mechanisms for intervening in hybrid behaviors. This work proposes the first framework that integrates hybrid automata into runtime enforcement, enabling coordinated discrete event editing and continuous-time monitoring to correct system behavior at any instant by suppressing, delaying, or inserting events. The paper establishes formal enforceability conditions and devises an online strategy synthesis algorithm based on reachability analysis. Evaluation on an adaptive cruise control case study demonstrates that the approach ensures safety properties even when the underlying controller is unsafe, all while incurring minimal computational overhead.
This work addresses “policy-invisible violations” in large language model agents—instances where outputs appear compliant but substantively breach organizational policies due to missing critical information such as entity attributes, contextual cues, or dialogue history. To tackle this, the authors propose Sentinel, a novel policy enforcement framework that introduces a world-state–based paradigm. Sentinel models agent actions as counterfactual update proposals over an organizational knowledge graph and validates structural invariants through speculative execution to precisely detect violations. The study also introduces PhantomPolicy, the first clean benchmark devoid of policy metadata, covering eight representative violation scenarios. Experiments demonstrate that Sentinel achieves 93.0% accuracy against human-audited ground-truth labels, substantially outperforming content-only DLP baselines (68.8%), thereby validating the efficacy of world-state modeling for accurate policy enforcement.
This work addresses the lack of system-level support in existing large language model (LLM) agents for long-running execution, state persistence, access control, and auditability, as well as the common misconception of treating tool invocations as trust boundaries. Inspired by library operating systems, the paper proposes a novel runtime environment that models LLM agents as AgentProcesses endowed with identity, lifecycle management, and capability-based access control. Authorization and isolation are enforced through runtime primitives rather than tool dispatching. Key innovations include a capability-based security model, a human approval queue, checkpointing mechanisms, and comprehensive audit logging. The system features asynchronous scheduling, namespace-isolated object memory, just-in-time tool registration, and an injectable resource layer. A prototype implementation passes 123 regression tests and real-model evaluations, supporting working directories, shell primitives, and one-time permission grants to ensure secure and controllable long-term operation.
This work addresses the lack of runtime-enforceable mechanisms for natural language behavioral policies in AI agents when invoking third-party skills, which hinders fine-grained monitoring and intervention against policy violations across actions. The authors propose VIGIL, an end-to-end runtime enforcement framework that translates skill specifications, operator constraints, and global rules into executable policies through a context-aware policy language and symbolic evaluation. VIGIL dynamically detects violations along agent execution traces by enabling fine-grained reasoning over event orderings, parameter relationships, and cross-call data flows, thereby overcoming limitations of fixed event models. Leveraging SMT-based symbolic execution and bounded trajectory constraints, VIGIL achieves efficient runtime monitoring and intervention. Empirical evaluation demonstrates that the approach attains over 95% violation detection accuracy with a false positive rate below 10% in real-world LLM agent tasks.
This work addresses the security threat posed by external tools that, while appearing benign at their interfaces, may embed malicious behaviors when invoked by large language model agents. To mitigate this risk, the authors propose ToolGuardian, a novel framework featuring a declarative policy layer based on Answer Set Programming. This layer explicitly models tool capabilities, effects, task context, and compositional relationships to enable auditable, deterministic pre-execution screening and runtime task-aware authorization. By integrating progressive tool representations—including natural-language descriptions, system call traces, simulated executions, and source code analysis—with logical reasoning, ToolGuardian achieves an F1 score of 0.86 and 88% accuracy in admission control across 16 MCP tools (including eight malicious variants) and 20 diverse scenarios, and attains 100% classification accuracy during runtime authorization.