🤖 AI Summary
This work addresses the redundant computation caused by duplicate texts in retrieval-augmented generation (RAG) by proposing a byte-level exact block deduplication method that significantly compresses context while preserving generation quality. For the first time, the compression efficacy of this approach is quantified across three real-world RAG scenarios—academic, enterprise, and conversational—achieving compression ratios of 0.16%, 24.03%, and 80.34%, respectively. Rigorous evaluation via multi-vendor large language model APIs, a five-category human-in-the-loop noise filtering protocol, and statistical validation using Wilson confidence intervals consistently demonstrates that all tested cases remain within a <5% quality degradation threshold. These results establish a deterministic optimization pathway that guarantees zero quality regression.
📝 Abstract
This preprint presents an empirical analysis of byte-exact chunk-level deduplication in Retrieval-Augmented Generation (RAG) pipelines. We measure context reduction across three distinct operating regimes: clean academic retrieval (0.16% byte reduction on 22.2M BeIR passages), constructed enterprise patterns (24.03% reduction), and multi-turn conversational AI (80.34% reduction). To validate quality preservation, we conducted a cross-vendor 5-judge calibrated panel evaluation across four production APIs (Google Gemini 2.5 Flash, Anthropic Claude Sonnet 4.6, Meta Llama 3.3 70B, and OpenAI GPT-5.1). Applying a five-category human-in-the-loop noise-removal protocol to panel-majority materially different (MAT) pairs, we establish that byte-exact deduplication introduces zero measurable quality regression. Post-audit, all four vendors clear the strict <5% Wilson 95% upper-bound MAT threshold in both the clean and high-redundancy RAG regimes. This work demonstrates that substantial inference compute savings can be achieved deterministically without compromising evaluation-grade model quality.