🤖 AI Summary
This study addresses privacy leakage arising from user queries and access patterns in retrieval-augmented generation (RAG), as well as the high latency and degraded retrieval quality of existing solutions. To this end, it proposes the first privacy-preserving RAG system built upon private information retrieval (PIR). The system introduces a two-stage sparse-dense hybrid retrieval architecture that leverages BM25 precomputed indices to filter candidate sets, thereby circumventing multi-round PIR overhead. It further integrates vector embedding reranking, hash binning, and oblivious tree traversal to preserve retrieval accuracy. Experimental results demonstrate that, while strictly guaranteeing query-content obliviousness, the proposed approach achieves lower latency and superior retrieval quality compared to state-of-the-art baselines.
📝 Abstract
Retrieval-augmented generation (RAG) hands the user's query to whoever hosts the corpus. We propose PILLAR, a Privacy-Preserving RAG (PPRAG) system based on Private Information Retrieval (PIR) in which a client utilizes the k documents most similar to their query from a server-held and publicly known corpus to respond to their query, while the server learns nothing about the query, either its terms or its access pattern. Prior PPRAG constructions rely on dense retrieval alone, translating approximate nearest-neighbor search into many query-dependent rounds of PIR, and pay for it in both latency and retrieval quality. PILLAR instead performs private hybrid retrieval in two stages. A sparse stage issues a small, fixed number of PIR queries against a carefully designed index of precomputed BM25 scores, filtering the corpus down to candidates that share terms with the query without the server ever seeing which terms these are. A dense stage then fetches only those candidates' document embeddings and re-ranks them locally, avoiding the many costly PIR queries that private dense retrieval typically requires. We instantiate PILLAR with two protocols that trade latency against retrieval quality, each built on a different private rendering of lexical search. PILLAR-Bin bins posting lists into a hash table and is a single-round design that achieves lower latency than state-of-the-art private retrieval schemes. PILLAR-Tree turns block-max pruning into an oblivious tree traversal combined with cuckoo hash tables and achieves the highest retrieval quality at lower latency than state-of-the-art schemes.