Score
Design and implement machine-executable policy artifacts—rules, constraints, and decision logic—expressed as source code and accompanied by tests, linters, and version control; and build the tooling and integrations (policy engines, CI/CD checks, runtime/admission controls, and audit mechanisms) that validate, enforce, and monitor those policies across development and deployment workflows.
Enterprise-scale general-purpose agents lack built-in, reusable governance mechanisms for autonomous cross-tool operation, making it difficult to satisfy requirements for compliance, auditability, and behavioral controllability. This work proposes the CUGA policy system, which embeds runtime governance capabilities into five critical checkpoints of the agent execution pipeline—intent protection, playbook guidance, tool invocation control, human approval gating, and output formatting—through a modular “policy-as-code” architecture. Without requiring model fine-tuning, CUGA enables proactive, continuous, and structured behavior control. By integrating typed governance primitives, dynamic playbook injection, and human-in-the-loop approval, the system effectively blocks malicious requests, enforces structured tool sequences, and triggers manual review for high-risk operations in healthcare scenarios, significantly enhancing policy adherence, execution consistency, and deployment safety.
AI governance policies—typically authored in natural language—are manually translated into executable rules, resulting in low efficiency, high error rates, and poor scalability, thereby impeding the deployment of safety mechanisms. Method: We propose Policy-to-Tests (P2T), the first systematic framework to automatically convert multi-source AI policy documents into standardized, machine-executable rules. P2T introduces a compact domain-specific language (DSL) to structurally encode policy elements—including risk categories, scope, conditions, exceptions, and evidentiary requirements—and employs an LLM-driven parsing and adjudication pipeline for end-to-end policy interpretation and behavioral compliance assessment. Contribution/Results: Experiments show P2T-generated rules match human-authored baselines in coverage and granularity (high inter-annotator agreement); when integrated with HIPAA-aligned safeguards, P2T significantly reduces agent policy violations. All artifacts—including source code, DSL specification, prompt templates, and rule sets—are publicly released to ensure reproducibility.
Large language model (LLM) agents exhibit unreliable adherence to corporate policies in business process automation. To address this, we propose a deterministic, transparent, and modular policy compliance framework comprising two phases: (1) an offline phase that compiles natural-language policy documents into verifiable guard code, and (2) a runtime phase that inserts lightweight, policy-agnostic guards before tool invocation—thereby decoupling policy enforcement from agent logic. This design enhances interpretability, maintainability, and agility in policy updates. Experiments on the τ-bench Airlines testbed demonstrate the framework’s effectiveness in intercepting policy-violating actions, validating its feasibility. However, empirical evaluation also uncovers critical deployment challenges, including incompleteness in policy coverage and difficulties in dynamically adapting guards to contextual changes. The framework thus advances policy-aware LLM agent deployment while surfacing key open issues for future work.
This work addresses the semantic gap faced by current AI agents in enforcing natural language policies: the intended policy semantics are difficult to enforce precisely and interpretably at the system level. To bridge this gap, the authors propose a novel approach that integrates agent-side context with kernel-level enforcement mechanisms. For the first time, policy context is preserved on the agent side, while a domain-specific language (DSL) for information flow control (IFC), implemented via eBPF, enables comprehensive, action-level policy enforcement within the operating system kernel. This framework supports cross-event data-flow and ordering constraints, significantly improving policy compliance rates by covering indirect execution paths invisible to conventional tool-call interception. The system incurs only 1.9%–8.4% runtime overhead and provides semantically clear feedback instead of ambiguous errors.
Existing automated tool-calling systems often suffer from insufficient generalization due to model-centric designs and heavy reliance on prompting, leading to recurrent failures such as unsafe side effects, invalid parameters, uncontrolled retries, and sensitive data leakage. This work proposes a model-agnostic, policy-first framework for tool orchestration that enforces permission control prior to invocation, enhancing safety through explicit constraints, risk-aware gating, recovery mechanisms, and auditable explanations. Key contributions include a policy-first paradigm for tool workflows, a lightweight domain-specific language (DSL) for policies, a runtime execution engine, and a reproducible safety benchmark based on trajectory replay. In 225 controlled experiments, the strictest policy configuration achieved a violation prevention rate of 0.681, reduced retry amplification to 1.378, and attained a sensitive information leakage recall of 0.875, effectively quantifying the trade-off between safety and utility.
本文提出Brain API,一种意图感知控制平面,通过决策工件解决现有系统中缺乏意图级决策治理的问题。
本文提出Agile-V Assurance Spine,通过权威源配置文件、工件绑定、风险适当独立性和时效性等方法解决工程生命周期中对代理输出的正当行动问题。
为了解决生成式AI应用中内容约束难以实施的问题,提出了可操作策略模式和合成数据生成管道,以实现从模型对齐到运行时监控的全生命周期政策执行。
This study addresses the unclear relationships among information sources, generation strategies, and quality evidence in test case generation using large language models (LLMs). Through a systematic literature review of 95 studies, this work constructs a multidimensional taxonomy and a benchmark analysis framework. Specifically, it proposes a four-dimensional classification system that elucidates how execution feedback influences oracle independence. Furthermore, it establishes a unified theoretical framework connecting the generation process with quality assessment. By identifying independent oracle evaluation as a critical yet underexplored dimension, this research formulates a future agenda centered on rigorous, oracle-independent quality measurement for LLM-generated test cases.
本文提出一种自主生成业务需求文档的代理框架,通过逆向工程处理未记录的企业软件,解决系统迁移、扩展或审计时缺乏正式需求文档的问题。