🤖 AI Summary
This study addresses a set-level poisoning threat in Retrieval-Augmented Generation (RAG) systems, wherein individually benign documents collectively mislead model outputs upon retrieval. To this end, we propose the LENS framework, which formalizes this vulnerability through a nested double-loop workflow. By employing a generator-agnostic multi-agent architecture to collaboratively synthesize adversarial evidence sets, and integrating query-conditioned interpretation with counterexample-guided refinement, LENS achieves precise semantic manipulation that remains ineffective at the subset level yet succeeds when the full set is retrieved. Experimental results demonstrate that the attack success rate reaches 0.852 for the complete set compared to merely 0.069 for subsets, yielding a 48.8% improvement in strict end-to-end success rates. These findings confirm that LENS effectively circumvents existing defenses, establishing a new paradigm for evaluating the security boundaries of RAG systems.
📝 Abstract
Retrieval-augmented generation (RAG) aggregates evidence from multiple external documents, yet this joint integration creates an underexamined vulnerability: attack effects absent in individual documents can emerge through set-level composition. Existing coordinated attacks do not explicitly enforce that every proper subset remains insufficient in frozen single-round RAG. We formalize set-level compositional poisoning, where documents designed to remain individually plausible jointly redirect RAG outputs to a target answer, while proper subsets fail to induce the target on their own. To construct such attacks, we propose LENS, a generator-black-box multi-agent framework that casts construction as constrained evidence composition. LENS factorizes target inference into a query-conditioned interpretation lens and complementary facts, then uses a nested dual-loop workflow to concentrate steering in the full set while suppressing subset leakage. The outer loop plans the interpretation lens and semantic roles; the inner loop synthesizes documents and applies counterexample-guided repair. Across four benchmarks and three generators, returned packets achieve 0.852 full-set ASR and 0.784 post-retrieval ASR@5, while their strongest proper subsets reach only 0.069. Against construction baselines evaluated on the same frozen manifest, LENS improves all-attempt E2E-Strict@5 from 0.244 to 0.363, a 48.8% relative gain. A blinded human audit finds that 68.3% of returned packets combine an incorrect target, a definite answer-criterion shift, and no target entailment under the original semantics. Across four published defenses, LENS attains the highest defended all-attempt ASR@5, exceeding the strongest baseline by 0.141 on average. Together, these results establish evidence composition as a distinct RAG security boundary and position LENS as a stress test for defenses that reason over document sets.