🤖 AI Summary
This work addresses the critical need for policy enforcement mechanisms in high-stakes domains that simultaneously offer scalability and formal safety guarantees. Existing approaches either lack formal verification or rely on handcrafted rules that do not scale. To bridge this gap, the authors propose a novel generate-and-critic loop framework grounded in large language models, which for the first time enables fully automated translation from natural language policy documents, agent prompts, and MCP tool descriptions into the formally verifiable Cedar policy language. The method preserves rigorous formal guarantees while substantially expanding policy coverage and scalability. Evaluated on the MedAgentBench benchmark, the automatically generated policies capture a more comprehensive subset of the original specifications than manually encoded counterparts.
📝 Abstract
Agent safety in high-stakes domains requires formal policy enforcement, but most existing approaches either rely on probabilistic guardrails (fine-tuned classifiers, prompt-based steering) that offer no formal guarantees, or on hand-coded symbolic enforcement that does not scale to the breadth of real policy specifications. We present an autoformalization pipeline that translates agent prompts, MCP tool descriptions, and natural language policy documents into formally verified policies using an LLM-based generator-critic loop. The resulting policies are written in the Cedar Policy Language. On the MedAgentBench benchmark, our autoformalized policies cover substantially more of the source natural-language specification than the hand-coded symbolic enforcement in prior work.