Towards More Robust Retrieval-Augmented Generation: Evaluating RAG Under Adversarial Poisoning Attacks

📅 2024-12-21
🏛️ arXiv.org
📈 Citations: 2
✨ Influential: 0
📄 PDF
🤖 AI Summary
Retrieval-augmented generation (RAG) systems mitigate large language model hallucinations but remain vulnerable to adversarial corpus poisoning attacks, which induce factual errors in generated outputs. This paper presents the first systematic analysis of RAG’s two-stage failure mechanism, identifying retrieval ranking bias as the primary driver of successful attacks. To address this, we propose “skeptical prompting”—a lightweight, model-agnostic self-validation framework that operates at the generation stage without fine-tuning. It integrates multi-round consistency verification with knowledge activation assessment to enhance output robustness. Through retrieval quality attribution analysis and targeted adversarial sample construction, we conduct empirical validation across diverse benchmarks. Experimental results demonstrate that our approach reduces erroneous response rates by up to 47%, offering a practical, deployable defense for secure RAG systems.

Technology Category

Machine Learning: Adversarial Learning & RobustnessComputer Vision: Adversarial Attacks & RobustnessNatural Language Processing: Safety and Robustness

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Retrieval-Augmented Generation (RAG) systems have emerged as a promising solution to mitigate LLM hallucinations and enhance their performance in knowledge-intensive domains. However, these systems are vulnerable to adversarial poisoning attacks, where malicious passages injected into retrieval databases can mislead the model into generating factually incorrect outputs. In this paper, we investigate both the retrieval and the generation components of RAG systems to understand how to enhance their robustness against such attacks. From the retrieval perspective, we analyze why and how the adversarial contexts are retrieved and assess how the quality of the retrieved passages impacts downstream generation. From a generation perspective, we evaluate whether LLMs' advanced critical thinking and internal knowledge capabilities can be leveraged to mitigate the impact of adversarial contexts, i.e., using skeptical prompting as a self-defense mechanism. Our experiments and findings provide actionable insights into designing safer and more resilient retrieval-augmented frameworks, paving the way for their reliable deployment in real-world applications.
Problem

Research questions and friction points this paper is trying to address.

Evaluating RAG vulnerability to adversarial poisoning attacks
Improving robustness of RAG systems against malicious inputs
Assessing skeptical prompting for self-defense in LLMs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Structured taxonomy for context types analysis
Evaluating retrievers under adversarial poisoning attacks
Skeptical prompting to activate LLMs self-defense
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Cornell University | Mohamed bin Zayed University of Artificial Intelligence