ScoreGate: Adaptive Chunk Selection for Retrieval-Augmented Generation via Dual-Score Statistical Fusion

πŸ“… 2026-06-12
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Fixed-size retrieval struggles to accommodate varying query complexity, often leading to either excessive or insufficient document retrieval. This work proposes ScoreGate, a lightweight adaptive mechanism that dynamically determines the number of retrieved passages by fusing similarity scores from a bi-encoder with reranking scores from a cross-encoderβ€”without requiring additional model invocations. ScoreGate is the first approach to jointly leverage both scoring signals to identify relevant documents previously underestimated due to lexical mismatch, thereby overcoming the limitations of fixed top-K retrieval or single-threshold strategies. Experiments show that on MS MARCO, ScoreGate achieves an MRR@10 of 0.401 while reducing retrieved passages by 35%. In internal evaluations, it attains near-perfect recall with zero false positives, cuts per-query token usage by 34.8%, and introduces only 31ms of latency.
πŸ“ Abstract
Fixed-cardinality retrieval injects a constant top-K chunks into the generator regardless of query complexity, causing over-retrieval for narrow queries and under-retrieval for compositional ones. We describe ScoreGate, a lightweight score-space decision mechanism that controls retrieval cardinality at inference time using two scores already produced by the standard pipeline: bi-encoder similarity s_i and cross-encoder reranker score r_i, with no additional model inference calls required. Its core insight is that cross-encoder affirmation can rescue semantically relevant chunks that bi-encoder retrieval ranks poorly due to vocabulary mismatch -- a failure mode unaddressed by fixed-K or single-score thresholding. On MS MARCO (200 dev queries), ScoreGate achieves MRR@10 = 0.401 with 35% fewer retained chunks than Standard Top-K. On an internal benchmark (n=300, Fleiss' kappa=0.87), ScoreGate observed zero false positives (95% CI [96.4%, 100%]) at 97.77-99.34% recall, with 34.8% fewer tokens per query and only 31ms added latency. Results on both MS MARCO and real-world production traffic suggest that adaptive retrieval cardinality can improve retrieval efficiency without degrading retrieval quality.
Problem

Research questions and friction points this paper is trying to address.

retrieval-augmented generation
adaptive retrieval
fixed-cardinality retrieval
query complexity
retrieval efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

adaptive retrieval
retrieval-augmented generation
dual-score fusion
cross-encoder reranking
retrieval efficiency
πŸ”Ž Similar Papers
2024-06-01International Conference on Computational LinguisticsCitations: 4