Lost in Conversation or Lost in Translation? Diagnosing Multi-Turn Degradation in RAG

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance degradation of Retrieval-Augmented Generation (RAG) systems in multi-turn dialogues caused by evaluation bias. Through large-scale simulation experiments, it systematically diagnoses the failure mechanisms of both RAG and GraphRAG in multi-hop question-answering scenarios. The research innovatively identifies two distinct failure modes—“translation loss” and “dialogue loss”—effectively addressing the blind spots inherent in conventional single-turn evaluations. Experimental results demonstrate that multi-turn interactions reduce system accuracy by 21% while increasing unreliability by 47%, pinpointing the core sources of performance deterioration. These findings provide critical evidence for the robustness evaluation and optimization of multi-turn retrieval-augmented generation systems.
📝 Abstract
When conversing with large language models (LLMs), users often begin with a simple question and build towards a multi-hop question through follow-up turns. Retrieval-augmented generation (RAG) and its graph-based variant (GraphRAG) have become the dominant approaches for grounding LLM responses in external evidence, yet both are evaluated almost exclusively on single-turn, fully specified queries. We systematically investigate this evaluation mismatch through a large-scale simulation study. Building on prior work on multi-turn LLM evaluation, we transform questions from multi-hop question answering (QA) benchmarks into underspecified conversations and evaluate ten LLM assistants with eight retrieval systems across 1.5 million simulated conversations. Our findings reveal that multi-turn interaction causes widespread performance degradation, incurring relative performance drops of up to 21% and increasing unreliability by 47%, making RAG systems simultaneously less accurate and less reliable. We identify two distinct failure modes behind this degradation. Systems are either lost in translation, where conversational rephrasing distorts the retrieval query, or lost in conversation, where retrieval succeeds but the LLM fails to synthesize evidence distributed across turns.
Problem

Research questions and friction points this paper is trying to address.

Retrieval-Augmented Generation
Multi-turn Dialogue
Performance Degradation
Large Language Models
Evaluation Mismatch
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-turn RAG
Evaluation Mismatch
Failure Mode Diagnosis
Large-scale Simulation
Multi-hop QA
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.