Score
Design and implement evaluation frameworks, benchmarks, and metric computations to measure retrieval system performance, accuracy, recall, and impact on downstream generation across retrieval architectures and modalities. Build diagnostic and stress tests (e.g., should-change/should-not-change cases), curate controlled datasets, isolate retrieval from generation, and analyze retrieval error modes and metrics (hallucination rates, robustness, embedding similarity, confusions) to compare approaches and guide retriever engineering.
In large-scale, dynamically evolving knowledge base retrieval, the total number of relevant documents is unknown, rendering conventional recall incomputable and hindering rigorous evaluation of retrieval quality. To address this, we propose NR-Metric—a novel retrieval effectiveness measure that does not require ground-truth relevance counts. Instead, it quantifies how well retrieval results predict the quality of downstream large language model (LLM) responses. Our method integrates multi-dataset comparative experiments, analysis against traditional metrics, LLM-based response quality assessment, and statistical correlation validation. Evaluated across multiple real-world datasets with 2–15 relevant documents per query, NR-Metric demonstrates strong correlation with LLM response quality (average Spearman’s ρ > 0.82) and consistently outperforms recall-dependent mainstream metrics. It offers a scalable, empirically verifiable evaluation paradigm for open-domain, dynamic retrieval settings.
This study investigates whether retrieval quality can serve as a reliable early indicator of information coverage in responses generated by Retrieval-Augmented Generation (RAG) systems. Through systematic experiments across three benchmarks—TREC NeuCLIR 2024, TREC RAG 2024, and WikiVideo—the authors evaluate 15 text-based and 10 multimodal retrieval systems using the Auto-ARGUE and MiRAGE assessment frameworks. The work provides the first empirical evidence of a strong correlation between coverage-oriented retrieval metrics and the informational coverage of generated outputs. Findings reveal that, at both topic and system levels, such metrics effectively predict RAG output coverage when retrieval and generation objectives are aligned, underscoring the critical role of goal consistency in optimizing RAG performance.
This study addresses the quantification of discriminative power in query-document relevance judgments (qrels) for information retrieval (IR) evaluation, with particular emphasis on the historically underexamined Type II error (false negatives) and its joint analysis with Type I error (false positives). Method: We systematically introduce Type II error modeling into IR evaluation for the first time and propose replacing conventional significance testing with balanced classification metrics—such as balanced accuracy—as a more principled basis for assessing qrels’ discriminative capability. A unified, comparable measurement framework is thereby established. Results: Empirical hypothesis testing and statistical significance analysis across multiple qrels generation strategies demonstrate that jointly evaluating both error types exposes latent quality deficiencies in qrels more comprehensively than traditional approaches. Balanced classification metrics robustly aggregate discriminative performance, substantially enhancing the reliability and interpretability of IR system evaluation.
This work addresses the challenge of deploying a shared retrieval backbone in industrial systems, where balancing performance and deployment flexibility across multiple downstream tasks remains difficult. To overcome the limitations of conventional approaches that rely on a single optimal checkpoint, the authors propose a multi-stage optimization framework that tailors component-level and hybrid-stage configuration strategies to the distinct performance characteristics of dense retrievers and rerankers throughout training. This approach significantly enhances the adaptability of the shared backbone and improves overall retrieval effectiveness. End-to-end evaluation demonstrates that the resulting shared retrieval service has been successfully deployed across multiple industrial applications, delivering substantial gains in both system performance and scalability.
This study addresses the lack of standardized validation criteria for user simulators in information retrieval evaluation, which undermines the reliability of simulation outcomes. Through a systematic literature review, it proposes the first structured taxonomy of metrics specifically designed for validating simulated search queries. The work empirically analyzes the interrelationships among these metrics across four diverse datasets and, based on the findings, offers tailored validation recommendations for different application scenarios. To foster standardization and reproducibility in simulation-based evaluation, the authors also release an open-source toolkit implementing commonly used validation metrics, thereby supporting future research extension and benchmarking.
Current RAG system evaluations overly rely on end-to-end accuracy, failing to capture enterprise-level requirements across dimensions such as reasoning complexity, retrieval difficulty, document structural diversity, and interpretability. Consequently, models achieving high scores often exhibit insufficient reliability in real-world deployments. To address this gap, this work proposes the first difficulty taxonomy integrating these four dimensions and introduces a multidimensional diagnostic framework and benchmark tailored for enterprise applications. The framework systematically identifies weaknesses of RAG systems in complex, realistic settings and effectively exposes performance bottlenecks that hinder practical deployment, thereby offering actionable pathways for evaluation and optimization to enhance real-world reliability.
This work addresses the absence of standardized benchmarks for end-to-end performance and quality evaluation in retrieval-augmented generation (RAG) systems. The authors propose a modular and configurable RAG benchmarking framework that, for the first time, enables decoupled assessment of individual pipeline stages—including embedding, indexing, retrieval, reranking, and generation. The framework supports multimodal data, multiple vector databases (e.g., Milvus, Qdrant), and diverse large language models, while accommodating realistic query loads and update patterns. It automatically collects key metrics such as throughput, resource utilization, and accuracy. Experimental results demonstrate that the framework introduces negligible performance overhead while providing comprehensive evaluation of RAG system effectiveness. The implementation is publicly released as open-source software.
This study addresses the confounding effects of metadata, structured representations, and retrieval mechanisms in current RAG systems, which often combine multiple context-augmentation strategies, obscuring their individual contributions to answer quality. Through controlled experiments across six benchmarks, four models, and five augmentation levels—totaling over 24,000 evaluations—the work reveals that increased contextual richness does not necessarily improve accuracy. It introduces the “tractability hierarchy” theory, emphasizing that context must align with model capacity. The findings demonstrate that most augmentation strategies actually degrade performance; however, when metadata and retrieval strategies are carefully matched to a model’s capabilities, smaller models can outperform state-of-the-art large models by up to 19 F1 points on specific tasks, challenging the prevailing RAG design paradigm centered on stacking metadata.
Traditional keyword-based code retrieval struggles to meet the demands of natural language queries, intent understanding, and code quality assessment. This work proposes a hybrid retrieval system that integrates semantic search with large language model (LLM)-generated quality metadata, supporting four query modes: semantic, quality-filtered, hybrid, and automatic routing. The approach innovatively incorporates function-level code slicing, text-code embeddings, and ChromaDB vector storage, and—novelly—leverages LLM-generated quality scores for dynamic query routing. Experiments on a C-language educational code corpus demonstrate strong performance: semantic retrieval achieves nDCG@5 of 0.820 and Success@5 of 0.800; automatic routing attains 100% accuracy; and in 9 out of 12 cases, LLM-predicted quality scores deviate by no more than one point from human evaluations.