AdaGuard: An Adaptive Guard Model with User-defined Policies

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge that static risk taxonomies struggle to accommodate dynamic user policies in compliance evaluation by proposing AdaGuard, an adaptive guardrail model. We construct the AdaptiveSafety dataset incorporating counterfactual samples and introduce a novel dynamic detection paradigm supporting 1 to 100 rules. The SafePO reinforcement learning algorithm is employed to balance reasoning and judgment rewards, while structured augmentation ensures consistency during rule reordering. Among the trained 0.6B–8B parameter model series, the 4B variant achieves accuracies of 89.30% and 71.82% on the AdaptiveSafety and DynaBench benchmarks, respectively. Ultimately, this approach enables real-time policy-driven dynamic evaluation of agent trajectories.
📝 Abstract
Guard models support the safe deployment of language model agents, but fixed risk taxonomies limit their ability to accommodate requirements that vary across applications and tasks. Under user-defined policies, detecting violations requires interpreting both the applicable rules and the agent's behavior, since identical actions can receive different judgments under different policies. To support learning this capability, we introduce AdaptiveSafety, a dataset of 10,939 training examples and 1,000 test examples covering policies with 1--100 rules. The dataset combines trajectories from multiple sources with policy and behavioral counterfactuals, pairing each example with an explanation and the complete set of violated rules. These counterfactuals expose changes that alter compliance, while structural augmentations provide supervision for consistency under rule reordering and identifier remapping. Building on this supervision, we propose SafePO, a reinforcement learning algorithm for refining violation identification while balancing explanatory reasoning and final verdicts. SafePO uses structured rewards to assess prediction correctness, retains group-relative advantages at the response level, and employs a separately trained value model to modulate token weights within explanation and verdict regions. Separate normalization controls their relative contribution to training despite differences in length. Through supervised initialization followed by SafePO, we develop AdaGuard, a family of 0.6B, 4B, and 8B guard models that assess agent trajectories under policies supplied at inference time. Our 4B model achieves binary accuracies of 89.30\% on AdaptiveSafety and 71.82\% on DynaBench. The project repository is available at https://github.com/Yunhao-Feng/AdaGuard
Problem

Research questions and friction points this paper is trying to address.

guard models
user-defined policies
violation detection
language model agents
risk taxonomies
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Guard Model
User-defined Policies
Counterfactual Augmentation
SafePO
Reinforcement Learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yunhao Feng
National University of Defense Technology
Y
Yifan Ding
Fudan University
Y
Yuxiang Xie
National University of Defense Technology
Zheng Li
Zheng Li
Chengdu University of Technology
Micro- and nano-scale flow in porous mediaShale oilShale gasCCUSUnderground hydrogen storage
M
Mingrui Lao
National University of Defense Technology
Zeyuan Wang
Zeyuan Wang
PhD, The University of Sydney
NLPMedical Informatics
Yanming Guo
Yanming Guo
National University of Defense Technology
deep learningcomputer vision