🤖 AI Summary
This study addresses the rigidity, opacity, and difficulty of balancing dynamic policies with latency constraints in enterprise-grade generative AI safety guardrails by proposing an adaptive LLM-as-a-Judge framework. Trained via supervised fine-tuning (SFT) and Group Relative Policy Optimization (GRPO), the core innovation lies in dynamically inferring input-policy pair complexity to automatically switch between high-speed black-box reasoning and interpretable auditing. This enables adaptive inference budget allocation and generalizes user-defined compliance policies without frequent model updates. Experimental results demonstrate that the proposed method achieves performance comparable to frontier models with several times more parameters, while recovering always-on reasoning accuracy at minimal latency, thereby substantially enhancing the efficiency of safety auditing.
📝 Abstract
Enterprise generative AI applications require robust safety mechanisms that can accommodate diverse risk postures, evolving policies, and varying latency constraints. Current guardrail solutions often suffer from rigidity, relying on fixed policy sets and offering limited transparency or reasoning flexibility. We present Adaguard, an adaptive LLM-as-a-Judge framework designed to address these challenges through dynamic policy enforcement and adaptive reasoning-budget allocation. Built using supervised fine-tuning (SFT) and reinforcement learning (GRPO), AdaGuard generalizes to user-defined safety and compliance policies at runtime without requiring frequent model updates. A core innovation of our approach is the ability to dynamically infer the complexity of input-policy pairs, allowing the model to switch between high-speed black-box inference and explainable, reasoning-enabled moderation. This flexibility enables developers to balance stringent latency requirements with the need for actionable transparency. This adaptive capability allows AdaGuard to rival other guardrail and frontier models several times its size, while its auto-reasoning mode recovers the accuracy of always-on reasoning at a fraction of the latency