🤖 AI Summary
This study addresses the challenge of detecting prompt injection attacks embedded within retrieved content in Retrieval-Augmented Generation (RAG) systems by proposing a leakage-aware benchmark construction pipeline and a rigorous evaluation protocol. The authors construct a dedicated benchmark dataset comprising 4,876 samples and systematically compare the performance of multiple detectors, including keyword matching, TF-IDF, and DistilBERT. Experimental results demonstrate that DistilBERT achieves the highest F1 score of 0.896, while also validating the effectiveness of sparse baseline methods. Overall, this work provides a reliable evaluation benchmark and methodological reference for enhancing the security defenses of RAG systems.
📝 Abstract
Retrieval-Augmented Generation (RAG) systems are vulnerable to prompt-injection attacks embedded in retrieved content. We introduce RAG-PIBench, a benchmark for RAG-style prompt-injection detection containing 4,876 contextual examples across frozen train, validation, and protected-test splits. Using a leakage-aware construction pipeline and strict evaluation protocol, we compare keyword-based, semantic-reference, TF-IDF, and transformer-based detectors. DistilBERT achieves the best protected-test performance (F1 = 0.896, PR-AUC = 0.968), while TF-IDF SVM and logistic regression remain competitive. Our results demonstrate the value of leakage-aware benchmark design and strong sparse baselines for reliable prompt-injection detection in RAG systems.