Provably Secure Retrieval-Augmented Generation

📅 2025-08-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
RAG systems face critical security threats—including data leakage and poisoning—yet existing defenses lack formal security guarantees, suffer from poor interpretability, and are vulnerable to adaptive attacks. To address this, we propose SAG, the first provably secure RAG framework. SAG employs end-to-end encryption prior to storage, simultaneously safeguarding both raw documents and their vector embeddings. We introduce the first cryptographically grounded formal security model for RAG, rigorously proving confidentiality and integrity under standard assumptions. By integrating ciphertext-based retrieval with protected embedding representations, SAG effectively mitigates state-of-the-art attacks across multiple benchmarks. Experiments demonstrate that SAG maintains competitive retrieval accuracy and generation quality while substantially enhancing security—achieving up to 98% attack mitigation without compromising latency or utility. Our work establishes both theoretical foundations and practical mechanisms for verifiably secure RAG deployment.

Technology Category

Natural Language Processing: Safety and RobustnessComputer Vision: Adversarial Attacks & RobustnessPhilosophy and Ethics of AI: Privacy & Security

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSecurity and Privacy: Data transparency and provenanceSemantics and Knowledge: Provenance, trust, security and privacy, and ethical issues in managing semantic data
📝 Abstract
Although Retrieval-Augmented Generation (RAG) systems have been widely applied, the privacy and security risks they face, such as data leakage and data poisoning, have not been systematically addressed yet. Existing defense strategies primarily rely on heuristic filtering or enhancing retriever robustness, which suffer from limited interpretability, lack of formal security guarantees, and vulnerability to adaptive attacks. To address these challenges, this paper proposes the first provably secure framework for RAG systems(SAG). Our framework employs a pre-storage full-encryption scheme to ensure dual protection of both retrieved content and vector embeddings, guaranteeing that only authorized entities can access the data. Through formal security proofs, we rigorously verify the scheme's confidentiality and integrity under a computational security model. Extensive experiments across multiple benchmark datasets demonstrate that our framework effectively resists a range of state-of-the-art attacks. This work establishes a theoretical foundation and practical paradigm for verifiably secure RAG systems, advancing AI-powered services toward formally guaranteed security.
Problem

Research questions and friction points this paper is trying to address.

Addressing privacy and security risks in RAG systems
Providing formal security guarantees for RAG frameworks
Ensuring confidentiality and integrity against advanced attacks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pre-storage full-encryption for dual protection
Formal security proofs for confidentiality and integrity
Resists state-of-the-art attacks in experiments
P
Pengcheng Zhou
Beijing University of Posts and Telecommunications
Y
Yinglun Feng
Beijing University of Posts and Telecommunications
Zhongliang Yang
Zhongliang Yang
Associate Professor, Beijing University of Posts and Telecommunications
AI SecurityFinTech