🤖 AI Summary
This study addresses the performance degradation of Retrieval-Augmented Generation (RAG) systems in multi-turn dialogues caused by evaluation bias. Through large-scale simulation experiments, it systematically diagnoses the failure mechanisms of both RAG and GraphRAG in multi-hop question-answering scenarios. The research innovatively identifies two distinct failure modes—“translation loss” and “dialogue loss”—effectively addressing the blind spots inherent in conventional single-turn evaluations. Experimental results demonstrate that multi-turn interactions reduce system accuracy by 21% while increasing unreliability by 47%, pinpointing the core sources of performance deterioration. These findings provide critical evidence for the robustness evaluation and optimization of multi-turn retrieval-augmented generation systems.
📝 Abstract
When conversing with large language models (LLMs), users often begin with a simple question and build towards a multi-hop question through follow-up turns. Retrieval-augmented generation (RAG) and its graph-based variant (GraphRAG) have become the dominant approaches for grounding LLM responses in external evidence, yet both are evaluated almost exclusively on single-turn, fully specified queries. We systematically investigate this evaluation mismatch through a large-scale simulation study. Building on prior work on multi-turn LLM evaluation, we transform questions from multi-hop question answering (QA) benchmarks into underspecified conversations and evaluate ten LLM assistants with eight retrieval systems across 1.5 million simulated conversations. Our findings reveal that multi-turn interaction causes widespread performance degradation, incurring relative performance drops of up to 21% and increasing unreliability by 47%, making RAG systems simultaneously less accurate and less reliable. We identify two distinct failure modes behind this degradation. Systems are either lost in translation, where conversational rephrasing distorts the retrieval query, or lost in conversation, where retrieval succeeds but the LLM fails to synthesize evidence distributed across turns.