How important is Recall for Measuring Retrieval Quality?

๐Ÿ“… 2025-12-23
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
In large-scale, dynamically evolving knowledge base retrieval, the total number of relevant documents is unknown, rendering conventional recall incomputable and hindering rigorous evaluation of retrieval quality. To address this, we propose NR-Metricโ€”a novel retrieval effectiveness measure that does not require ground-truth relevance counts. Instead, it quantifies how well retrieval results predict the quality of downstream large language model (LLM) responses. Our method integrates multi-dataset comparative experiments, analysis against traditional metrics, LLM-based response quality assessment, and statistical correlation validation. Evaluated across multiple real-world datasets with 2โ€“15 relevant documents per query, NR-Metric demonstrates strong correlation with LLM response quality (average Spearmanโ€™s ฯ > 0.82) and consistently outperforms recall-dependent mainstream metrics. It offers a scalable, empirically verifiable evaluation paradigm for open-domain, dynamic retrieval settings.

Technology Category

Data Mining & Knowledge Management: Conversational Systems for Recommendation & RetrievalMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Question Answering

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
๐Ÿ“ Abstract
In realistic retrieval settings with large and evolving knowledge bases, the total number of documents relevant to a query is typically unknown, and recall cannot be computed. In this paper, we evaluate several established strategies for handling this limitation by measuring the correlation between retrieval quality metrics and LLM-based judgments of response quality, where responses are generated from the retrieved documents. We conduct experiments across multiple datasets with a relatively low number of relevant documents (2-15). We also introduce a simple retrieval quality measure that performs well without requiring knowledge of the total number of relevant documents.
Problem

Research questions and friction points this paper is trying to address.

Evaluates recall alternatives for large knowledge bases
Measures correlation between retrieval metrics and LLM judgments
Proposes new metric without needing total relevant documents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evaluates retrieval metrics via LLM-based quality judgments
Tests across datasets with few relevant documents
Introduces recall-free retrieval quality measure
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
Primer Technologies Inc.
S
Shelly Schwartz
Primer Technologies Inc., San Francisco, California
O
Oleg Vasilyev
Primer Technologies Inc., San Francisco, California
R
Randy Sawaya
Primer Technologies Inc., San Francisco, California