Evaluating Multi-Hop Reasoning in RAG Systems: A Comparison of LLM-Based Retriever Evaluation Strategies

📅 2026-04-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Current evaluation methods for retrieval-augmented generation (RAG) systems predominantly focus on single-hop queries, failing to accurately assess retriever performance in multi-hop reasoning scenarios. To address this limitation, this work proposes Context-Aware Retriever Evaluation (CARE), the first framework to systematically incorporate contextual relevance into multi-hop retrieval evaluation. Leveraging RAG simulation environments built on HotPotQA, MuSiQue, and SQuAD, and employing large language models from OpenAI, Meta, and Google as judges, CARE establishes an LLM-as-judge automated evaluation paradigm. Experimental results demonstrate that CARE significantly outperforms existing evaluation approaches on multi-hop queries, with particularly pronounced gains when using large-parameter models with extended context windows, whereas single-hop settings exhibit minimal sensitivity to contextual awareness.

Technology Category

Natural Language Processing: Question AnsweringKnowledge Representation and Reasoning: Common-Sense ReasoningData Mining & Knowledge Management: Conversational Systems for Recommendation & Retrieval

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Retrieval-augmented generation (RAG) enhances large language models (LLMs) with external knowledge to answer questions more accurately. However, research on evaluating RAG systems-particularly the retriever component-remains limited, as most existing work focuses on single-context retrieval rather than multi-hop queries, where individual contexts may appear irrelevant in isolation but are essential when combined. In this research, we use the HotPotQA, MuSiQue, and SQuAD datasets to simulate a RAG system and compare three LLM-as-judge evaluation strategies, including our proposed Context-Aware Retriever Evaluation (CARE). Our goal is to better understand how multi-hop reasoning can be most effectively evaluated in RAG systems. Experiments with LLMs from OpenAI, Meta, and Google demonstrate that CARE consistently outperforms existing methods for evaluating multi-hop reasoning in RAG systems. The performance gains are most pronounced in models with larger parameter counts and longer context windows, while single-hop queries show minimal sensitivity to context-aware evaluation. Overall, the results highlight the critical role of context-aware evaluation in improving the reliability and accuracy of retrieval-augmented generation systems, particularly in complex query scenarios. To ensure reproducibility, we provide the complete data of our experiments at https://github.com/lorenzbrehme/CARE.
Problem

Research questions and friction points this paper is trying to address.

multi-hop reasoning
RAG systems
retriever evaluation
context-aware evaluation
LLM-as-judge
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-hop reasoning
retrieval-augmented generation
context-aware evaluation
retriever evaluation
LLM-as-judge
🔎 Similar Papers
2024-02-26Annual Meeting of the Association for Computational LinguisticsCitations: 97
L
Lorenz Brehme
Universität Innsbruck, Technikerstraße 21a, 6020 Innsbruck, Austria
T
Thomas Ströhle
Universität Innsbruck, Technikerstraße 21a, 6020 Innsbruck, Austria
R
Ruth Breu
Universität Innsbruck, Technikerstraße 21a, 6020 Innsbruck, Austria