RAG-PIBench: A Leakage-Aware Benchmark for Prompt-Injection Detection in Trustworthy RAG Systems

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of detecting prompt injection attacks embedded within retrieved content in Retrieval-Augmented Generation (RAG) systems by proposing a leakage-aware benchmark construction pipeline and a rigorous evaluation protocol. The authors construct a dedicated benchmark dataset comprising 4,876 samples and systematically compare the performance of multiple detectors, including keyword matching, TF-IDF, and DistilBERT. Experimental results demonstrate that DistilBERT achieves the highest F1 score of 0.896, while also validating the effectiveness of sparse baseline methods. Overall, this work provides a reliable evaluation benchmark and methodological reference for enhancing the security defenses of RAG systems.
📝 Abstract
Retrieval-Augmented Generation (RAG) systems are vulnerable to prompt-injection attacks embedded in retrieved content. We introduce RAG-PIBench, a benchmark for RAG-style prompt-injection detection containing 4,876 contextual examples across frozen train, validation, and protected-test splits. Using a leakage-aware construction pipeline and strict evaluation protocol, we compare keyword-based, semantic-reference, TF-IDF, and transformer-based detectors. DistilBERT achieves the best protected-test performance (F1 = 0.896, PR-AUC = 0.968), while TF-IDF SVM and logistic regression remain competitive. Our results demonstrate the value of leakage-aware benchmark design and strong sparse baselines for reliable prompt-injection detection in RAG systems.
Problem

Research questions and friction points this paper is trying to address.

Retrieval-Augmented Generation
Prompt Injection
Benchmark
Trustworthy AI
Data Leakage
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prompt-Injection Detection
Retrieval-Augmented Generation
Leakage-Aware Benchmark
DistilBERT
Trustworthy RAG