RAISE: Reinforcing Access Control Policy Synthesis in LLMs via Symbolic Evaluation

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the frequent violation of authorization semantics when large language models generate access control policies. To this end, it introduces CedarInstruct, the first policy synthesis dataset supporting formal verification, and proposes RAISE, a two-stage training framework. This framework integrates supervised fine-tuning with a reinforcement learning algorithm guided by symbolic counterexample exploration (RAISE-OC), leveraging LoRA and GRPO to optimize policy synthesis. Experimental results demonstrate that the trained Qwen3.5-9B model surpasses GPT-6 Astra and Claude Opus 5 in semantic success rate while exhibiting strong generalization capabilities on independent benchmarks.
📝 Abstract
Translating natural-language access-control requirements into policies requires careful reasoning about permissions, constraints, and exceptions, and even frontier LLMs often produce policies that violate the intended authorization semantics. We construct CedarInstruct, to our knowledge the first dataset that supports both training and semantic evaluation for formally verifiable Cedar policy synthesis. It contains 5,800 scenarios across 44 domains and 1,408 representing a single synthetic organization, each with a verified target policy and an executable verification plan. On this data we introduce RAISE, which trains policy synthesizers from formal verification in two stages, verified supervised fine-tuning (SFT) followed by a reinforcement learning (RL) stage that learns from verifier signal. We find that SFT succeeds largely by letting models express authorization logic they already have, since untrained models rarely write valid Cedar but often reason correctly when they do. After SFT, how the verifier's information is used matters more than how much of it is used. Of six RL instantiations that consume progressively richer verifier signal, only RAISE-OC improves meaningfully on SFT; it turns failed checks and symbolic counterexamples into guided exploration and learns from the result with off-context GRPO. With about 5.4K verified scenarios and LoRA fine-tuning, RAISE-OC trains Qwen3.5-9B to surpass zero-shot GPT-6 Astra and Claude Opus 5 by 13.33 and 16.26 percentage points in semantic success on held-out scenarios, and training transfers to the independently constructed CedarBench.
Problem

Research questions and friction points this paper is trying to address.

access control policy synthesis
large language models
formal verification
natural language requirements
Cedar policy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Formal Verification
Symbolic Evaluation
Reinforcement Learning
Access Control Policy Synthesis
Off-context GRPO
🔎 Similar Papers