🤖 AI Summary
To address the dual challenges of insufficient contextual awareness and over-defensiveness in large language model (LLM) prompt injection detection—manifesting as high false-negative rates under adversarial conditions and high false-positive rates on benign inputs—this paper introduces the first lightweight benchmark jointly evaluating attack detection capability and over-defense propensity. Methodologically, we propose a context-aware evaluation framework integrating dynamic context modeling, adversarial sample generation, and fine-grained defense behavior analysis, augmented by a minimally supervised bias calibration mechanism enabling robust assessment with only a few domain-specific examples. Our contributions are threefold: (1) the first simultaneous quantification of both false-negative and false-positive rates, overcoming limitations of static testing paradigms; (2) empirical identification of critical vulnerabilities in state-of-the-art protective models; and (3) a reproducible, generalizable, and trustworthy evaluation standard for LLM security.
📝 Abstract
Prompt injection remains a major security risk for large language models. However, the efficacy of existing guardrail models in context-aware settings remains underexplored, as they often rely on static attack benchmarks. Additionally, they have over-defense tendencies. We introduce CAPTURE, a novel context-aware benchmark assessing both attack detection and over-defense tendencies with minimal in-domain examples. Our experiments reveal that current prompt injection guardrail models suffer from high false negatives in adversarial cases and excessive false positives in benign scenarios, highlighting critical limitations.