Score
Designs and builds systems that translate high-level policy specifications or rule descriptions (including automatically generated ones) into deterministic, executable artifacts such as code, rule-engine configurations, or control-flow graphs. Implements correctness-preserving compilation and separation of deterministic control flow from reasoning, plus runtime enforcement and dense diagnostic telemetry to ensure faithful, debuggable rule execution.
Current approaches to automated program synthesis lack effective governance mechanisms to ensure the compliance of generated code. This work proposes Protocol-Driven Development (PDD), a model that treats machine-executable protocols as primary artifacts and delineates the space of valid implementations through structural, behavioral, and operational invariants. PDD mandates that every implementation be accompanied by a verifiable chain of compliance evidence. By integrating formal methods, property-based testing, policy-as-code, and software provenance techniques, PDD establishes a unified framework for protocol specification and verification. This framework enables trustworthy admission control over automatically synthesized code, guaranteeing that all adopted implementations strictly adhere to protocol constraints and are backed by complete, auditable proofs of compliance.
AI governance policies—typically authored in natural language—are manually translated into executable rules, resulting in low efficiency, high error rates, and poor scalability, thereby impeding the deployment of safety mechanisms. Method: We propose Policy-to-Tests (P2T), the first systematic framework to automatically convert multi-source AI policy documents into standardized, machine-executable rules. P2T introduces a compact domain-specific language (DSL) to structurally encode policy elements—including risk categories, scope, conditions, exceptions, and evidentiary requirements—and employs an LLM-driven parsing and adjudication pipeline for end-to-end policy interpretation and behavioral compliance assessment. Contribution/Results: Experiments show P2T-generated rules match human-authored baselines in coverage and granularity (high inter-annotator agreement); when integrated with HIPAA-aligned safeguards, P2T significantly reduces agent policy violations. All artifacts—including source code, DSL specification, prompt templates, and rule sets—are publicly released to ensure reproducibility.
This work addresses the challenge of ensuring deterministic policy enforcement in large language model (LLM) agents operating within complex authorization scenarios—such as customer service, approval workflows, and data access control—where conventional approaches often fail to guarantee compliance. To this end, we propose PCAS, the first policy compiler for agent systems, which models system state via dependency graphs, expresses policies declaratively using Datalog rules, and integrates a reference monitor to intercept policy-violating actions prior to execution. By decoupling policy enforcement from model inference and embedding compliance directly into system construction, PCAS automatically synthesizes policy-compliant agent systems without requiring security-oriented redesign. Evaluation across three real-world scenarios demonstrates that our approach increases policy compliance from 48% to 93% and achieves zero runtime violations.
This work addresses the challenge that large language model (LLM) agents often produce redundant, exploratory, and non-deterministic execution trajectories that are difficult to reuse. To overcome this, the authors propose a skill-guided framework that extracts reusable structures from noisy trajectories and compiles them into near-deterministic workflows. The core innovations include a dependency inference mechanism based on evidence tuples—establishing strong dependencies only when parameters are uniquely traceable and flagging ambiguous relations as suspect—along with fine-grained binding-type categorization. The method integrates trajectory clustering, dependency rule mining, deterministic replay, and leave-one-out validation into a unified pipeline. Experiments demonstrate high precision (0.928) and recall (0.943) in dependency identification on the T1 dataset; for Venmo tasks, API calls are reduced from 34 to 11 while passing 15 of 21 test cases, and the system correctly rejects ill-posed or irreversible intents in Spotify and Todoist scenarios.
This study addresses the challenge of providing certifiable runtime safety guarantees prior to tool invocation, focusing on three core issues: the representability of policy states, the observability of monitoring evidence, and the impact of interventions on future behavior. To this end, we propose the first formal theoretical framework for runtime safety-executable boundaries, distinguishing among static policy executability, statistical calibration under exogenous legal constraints, and closed-loop intervention effects. Building upon finitely controlled models, we develop a method for closed-loop safety certification that integrates register model identification, Neyman–Pearson hypothesis testing, conformal calibration, and occupancy planning. Empirical validation through static diagnosis, model enumeration, representation rewriting, and closed-loop re-execution experiments demonstrates the efficacy of our approach and exposes the fundamental limitations of static calibration under representation attacks.
This work addresses critical challenges in configuration management for large language model (LLM)-based coding agents, including configuration reuse ambiguities, unclear permission boundaries, and inadequate versioning. To tackle these issues, the authors propose Rel(AI)Build—the first deterministic, tool-agnostic configuration governance framework specifically designed for LLM coding agents. Treating agent definitions as managed supply chains, Rel(AI)Build enforces configuration integrity through SHA-256 content addressing, HMAC-signed lockfiles, hash-chain audit logs, hierarchical access controls, and a state-machine-driven development workflow. The framework also supports multi-IDE target compilation. Empirical evaluation demonstrates that Rel(AI)Build effectively preserves configuration immutability under adversarial compliance tests, thereby validating its reliability and security guarantees.
This work addresses the high computational cost, slow inference speed, and lack of determinism in large language model (LLM) agents stemming from token-by-token reasoning. The authors propose an AGI compilation paradigm that records and analyzes agent behavior to identify deterministic segments, which are then extracted as verifiable programs or distilled into expert models. These components are compiled into WebAssembly cognitive binaries equipped with capability declarations and performance guarantees, and executed within a sandboxed environment. This approach enables, for the first time, the automatic conversion of agent experiences into permanent skills with near-zero marginal cost and supports quantifiable uncertainty estimation. Experiments on AUTO-BENCH show that 87.1% of behavioral segments exhibit observational determinism; under distribution shift, per-query inference cost drops from 59 to 2 micro-dollars (a 6.4× speedup), achieving 96.9% accuracy with zero errors.
Enterprise-scale general-purpose agents lack built-in, reusable governance mechanisms for autonomous cross-tool operation, making it difficult to satisfy requirements for compliance, auditability, and behavioral controllability. This work proposes the CUGA policy system, which embeds runtime governance capabilities into five critical checkpoints of the agent execution pipeline—intent protection, playbook guidance, tool invocation control, human approval gating, and output formatting—through a modular “policy-as-code” architecture. Without requiring model fine-tuning, CUGA enables proactive, continuous, and structured behavior control. By integrating typed governance primitives, dynamic playbook injection, and human-in-the-loop approval, the system effectively blocks malicious requests, enforces structured tool sequences, and triggers manual review for high-risk operations in healthcare scenarios, significantly enhancing policy adherence, execution consistency, and deployment safety.