🤖 AI Summary
Multi-source retrieval-augmented generation (RAG) suffers from data sparsity and cross-source information conflicts, significantly exacerbating hallucination. To address these challenges, we propose a knowledge-guided multi-source graph augmentation framework. First, we construct a multi-source line graph to explicitly model inter-document logical relationships, mitigating sparsity. Second, we design a dual-level confidence mechanism—operating at both graph and node levels—to dynamically identify and suppress inconsistent information. Our approach integrates graph-structured modeling, fine-grained knowledge fusion, and multi-level confidence-driven retrieval optimization. Evaluated on four multi-domain query datasets and two multi-hop question answering benchmarks, our method consistently outperforms state-of-the-art RAG approaches: hallucination rates decrease by 23.6% on average, while answer accuracy improves by 19.4%. The framework substantially enhances the reliability and robustness of multi-source knowledge retrieval in complex, real-world scenarios.
📝 Abstract
Retrieval Augmented Generation (RAG) has emerged as a promising solution to address hallucination issues in Large Language Models (LLMs). However, the integration of multiple retrieval sources, while potentially more informative, introduces new challenges that can paradoxically exacerbate hallucination problems. These challenges manifest primarily in two aspects: the sparse distribution of multi-source data that hinders the capture of logical relationships and the inherent inconsistencies among different sources that lead to information conflicts. To address these challenges, we propose MultiRAG, a novel framework designed to mitigate hallucination in multi-source retrieval-augmented generation through knowledge-guided approaches. Our framework introduces two key innovations: (1) a knowledge construction module that employs multi-source line graphs to efficiently aggregate logical relationships across different knowledge sources, effectively addressing the sparse data distribution issue; and (2) a sophisticated retrieval module that implements a multi-level confidence calculation mechanism, performing both graph-level and node-level assessments to identify and eliminate unreliable information nodes, thereby reducing hallucinations caused by inter-source inconsistencies. Extensive experiments on four multi-domain query datasets and two multi-hop QA datasets demonstrate that MultiRAG significantly enhances the reliability and efficiency of knowledge retrieval in complex multi-source scenarios. extcolor{blue}{Our code is available in https://github.com/wuwenlong123/MultiRAG.