🤖 AI Summary
This work addresses the limitation of existing defenses in distinguishing malicious requests from out-of-domain queries, which fails to ensure large language model (LLM) agents operate within prescribed functional boundaries. We propose a production-grade runtime defense framework whose core innovation lies in semantic specification-based boundary enforcement. By synchronously validating inputs and outputs through semantic whitelists and blacklists, the framework ensures LLM agents strictly adhere to explicit functional constraints. Additionally, we introduce the PAGE benchmark for systematically evaluating function-specific defense capabilities. Experimental results demonstrate that the proposed framework achieves an overall accuracy of 95.9% and an out-of-domain detection rate of 93.5%, while reducing the false pass-through rate to 4.7%. Furthermore, it satisfies the low-latency requirements essential for production deployment.
📝 Abstract
Large language model-based agents are increasingly deployed to perform domain-specific tasks by interacting with enterprise knowledge, tools, and external services. Existing runtime guardrails primarily target prompt injection and other attack-specific behaviors under a black-box threat model, but provide limited guarantees that agents operate within their intended functionality. As a result, production agents remain vulnerable to malicious requests and out-of-domain queries that existing defenses often fail to distinguish. We present Persona Guardrail, a production-grade runtime defense framework that enforces explicit functional boundaries for customer-facing agentic AI systems through synchronous input and output validation driven by semantic allowlist and blocklist specifications. We also introduce PAGE (Persona-Aware Guardrail Evaluation), a benchmark for evaluating function-specific guardrails across benign, adversarial, and out-of-domain interactions on both user and agent turns. Compared with a generic LLM-based guardrail, Persona Guardrail improves overall accuracy from 85.7% to 95.9%, increases out-of-domain detection from 57.3% to 93.5%, and reduces the false-approved rate from 25.0% to 4.7%. Currently deployed in production, Persona Guardrail meets its latency budget while sustaining a very low false-block and false-allow rate under realistic production workloads. These results demonstrate that Persona Guardrail provides a practical, scalable, and production-ready foundation for securing agentic AI systems.