🤖 AI Summary
Current evaluations of prompt injection and jailbreak detectors for large language models often suffer from dataset-specific threshold tuning and opaque operating points. This work proposes a unified evaluation framework that enforces a single global operating point—maximizing F1 score under a false positive rate (FPR) constraint of ≤1%—across 16 public benchmarks. The framework employs a dual-channel cross-validation strategy combining StratifiedKFold with a StratifiedGroupKFold variant based on parent prompt IDs and MinHash+LSH clustering to mitigate data leakage. To systematically assess detector robustness and generalization, it integrates multidimensional diagnostic mechanisms, including adversarial validation, permutation feature importance, and paraphrase invariance probes. This approach enables consistent, reproducible cross-dataset performance comparisons and establishes quantifiable criteria for evaluating detector generalization.
📝 Abstract
Published evaluations of prompt-injection and jailbreak detectors for Large Language Models often suffer from two systematic weaknesses: per-dataset threshold tuning and undisclosed operating points. We describe an evaluation harness that addresses both. The detector under evaluation is scored across 16 public benchmarks (12,111 samples) using 5-fold cross-validation. StratifiedKFold (by row) is the headline pass; a parallel StratifiedGroupKFold pass over a composite key (parent-prompt id plus MinHash + LSH near-duplicate clusters at Jaccard $\gtrsim 0.8$) runs alongside it as a leakage-premium diagnostic. A single global operating point is selected on the held-out folds (max F1 subject to FPR $\leq 1\%$) and applied uniformly to every dataset, so per-dataset results reflect one threshold rather than per-benchmark optimisation. Generalisation is examined through a battery of diagnostics (leave-one-dataset-out cross-validation, a random-label control, adversarial validation, permutation feature importance, length-bias correlation, classifier-head agreement, cross-source near-duplicate detection, threshold transferability, train-vs-OOF agreement, and a paraphrase-invariance probe), most with a quantitative pass threshold and the remainder with a stated failure mode. For every external comparison, the detector's threshold is re-tuned to the competitor's published false-positive rate so head-to-head values are evaluated at matched operating points.