π€ AI Summary
Fixed-size retrieval struggles to accommodate varying query complexity, often leading to either excessive or insufficient document retrieval. This work proposes ScoreGate, a lightweight adaptive mechanism that dynamically determines the number of retrieved passages by fusing similarity scores from a bi-encoder with reranking scores from a cross-encoderβwithout requiring additional model invocations. ScoreGate is the first approach to jointly leverage both scoring signals to identify relevant documents previously underestimated due to lexical mismatch, thereby overcoming the limitations of fixed top-K retrieval or single-threshold strategies. Experiments show that on MS MARCO, ScoreGate achieves an MRR@10 of 0.401 while reducing retrieved passages by 35%. In internal evaluations, it attains near-perfect recall with zero false positives, cuts per-query token usage by 34.8%, and introduces only 31ms of latency.
π Abstract
Fixed-cardinality retrieval injects a constant top-K chunks into the generator regardless of query complexity, causing over-retrieval for narrow queries and under-retrieval for compositional ones. We describe ScoreGate, a lightweight score-space decision mechanism that controls retrieval cardinality at inference time using two scores already produced by the standard pipeline: bi-encoder similarity s_i and cross-encoder reranker score r_i, with no additional model inference calls required. Its core insight is that cross-encoder affirmation can rescue semantically relevant chunks that bi-encoder retrieval ranks poorly due to vocabulary mismatch -- a failure mode unaddressed by fixed-K or single-score thresholding. On MS MARCO (200 dev queries), ScoreGate achieves MRR@10 = 0.401 with 35% fewer retained chunks than Standard Top-K. On an internal benchmark (n=300, Fleiss' kappa=0.87), ScoreGate observed zero false positives (95% CI [96.4%, 100%]) at 97.77-99.34% recall, with 34.8% fewer tokens per query and only 31ms added latency. Results on both MS MARCO and real-world production traffic suggest that adaptive retrieval cardinality can improve retrieval efficiency without degrading retrieval quality.