🤖 AI Summary
Existing black-box privacy scores lack interpretability because they cannot pinpoint which component within a retrieval-augmented generation (RAG) system a privacy defense actually affects. This work proposes an active path auditing method that inserts source-level hooks into the retrieval, retrieved content, and generation stages to map privacy metrics to specific leakage channels. By integrating exact-match canary testing with the NEL_strict metric to evaluate named entity leakage, the study reveals for the first time that certain differential privacy (DP)-based defenses only adjust retrieval scores without influencing generation, thereby failing to suppress named entity leakage. In contrast, end-to-end LPRAG completely blocks all 150 canary leaks in the email channel, significantly outperforming existing approaches.
📝 Abstract
Black-box privacy scores for retrieval-augmented generation (RAG) are difficult to interpret unless the audited defense's active pipeline hook is known. We propose an active-path audit: inventory source-level hooks over retrieval, retrieved content, and generation; map each metric to the leakage channel it observes; and validate generated-text effects with exact-match canaries. In our benchmark reimplementations, the DP-style defenses modify retrieval scores only: their generation hooks are TODO-flagged stubs that return responses unchanged. This active path explains why they affect membership-inference behavior but track No-Defense on generated-text named-entity leakage, measured by NEL_strict. By contrast, the end-to-end LPRAG path is canary-validated on the email channel, recovering 53/150 canaries under No-Defense and 0/150 under LPRAG. These findings concern our reimplementations on our stack, not released defenses or defense families; the contribution is a methodology and case study, not a universal ranking