🤖 AI Summary
Current evaluation methods for retrieval-augmented generation (RAG) systems predominantly focus on single-hop queries, failing to accurately assess retriever performance in multi-hop reasoning scenarios. To address this limitation, this work proposes Context-Aware Retriever Evaluation (CARE), the first framework to systematically incorporate contextual relevance into multi-hop retrieval evaluation. Leveraging RAG simulation environments built on HotPotQA, MuSiQue, and SQuAD, and employing large language models from OpenAI, Meta, and Google as judges, CARE establishes an LLM-as-judge automated evaluation paradigm. Experimental results demonstrate that CARE significantly outperforms existing evaluation approaches on multi-hop queries, with particularly pronounced gains when using large-parameter models with extended context windows, whereas single-hop settings exhibit minimal sensitivity to contextual awareness.
📝 Abstract
Retrieval-augmented generation (RAG) enhances large language models (LLMs) with external knowledge to answer questions more accurately.
However, research on evaluating RAG systems-particularly the retriever component-remains limited, as most existing work focuses on single-context retrieval rather than multi-hop queries, where individual contexts may appear irrelevant in isolation but are essential when combined. In this research, we use the HotPotQA, MuSiQue, and SQuAD datasets to simulate a RAG system and compare three LLM-as-judge evaluation strategies, including our proposed Context-Aware Retriever Evaluation (CARE). Our goal is to better understand how multi-hop reasoning can be most effectively evaluated in RAG systems.
Experiments with LLMs from OpenAI, Meta, and Google demonstrate that CARE consistently outperforms existing methods for evaluating multi-hop reasoning in RAG systems. The performance gains are most pronounced in models with larger parameter counts and longer context windows, while single-hop queries show minimal sensitivity to context-aware evaluation. Overall, the results highlight the critical role of context-aware evaluation in improving the reliability and accuracy of retrieval-augmented generation systems, particularly in complex query scenarios. To ensure reproducibility, we provide the complete data of our experiments at https://github.com/lorenzbrehme/CARE.