๐ค AI Summary
In large-scale, dynamically evolving knowledge base retrieval, the total number of relevant documents is unknown, rendering conventional recall incomputable and hindering rigorous evaluation of retrieval quality. To address this, we propose NR-Metricโa novel retrieval effectiveness measure that does not require ground-truth relevance counts. Instead, it quantifies how well retrieval results predict the quality of downstream large language model (LLM) responses. Our method integrates multi-dataset comparative experiments, analysis against traditional metrics, LLM-based response quality assessment, and statistical correlation validation. Evaluated across multiple real-world datasets with 2โ15 relevant documents per query, NR-Metric demonstrates strong correlation with LLM response quality (average Spearmanโs ฯ > 0.82) and consistently outperforms recall-dependent mainstream metrics. It offers a scalable, empirically verifiable evaluation paradigm for open-domain, dynamic retrieval settings.
๐ Abstract
In realistic retrieval settings with large and evolving knowledge bases, the total number of documents relevant to a query is typically unknown, and recall cannot be computed. In this paper, we evaluate several established strategies for handling this limitation by measuring the correlation between retrieval quality metrics and LLM-based judgments of response quality, where responses are generated from the retrieved documents. We conduct experiments across multiple datasets with a relatively low number of relevant documents (2-15). We also introduce a simple retrieval quality measure that performs well without requiring knowledge of the total number of relevant documents.