๐ค AI Summary
Existing RAG methods overlook the heterogeneity of source reliability in multi-source retrieval, relying solely on relevance-based rankingโleading to hallucinations and erroneous information propagation. To address this, we propose RA-RAG, a reliability-aware RAG framework that, for the first time, jointly models and iteratively optimizes source reliability and answer truthfulness in a fully unsupervised manner. RA-RAG introduces three core components: (i) a reliability estimation mechanism, (ii) a reliability-constrained scalable retrieval strategy, and (iii) a reliability-weighted majority voting aggregation method. To rigorously evaluate source reliability heterogeneity, we construct the first benchmark evaluation framework specifically designed for this challenge. Extensive experiments demonstrate that RA-RAG significantly outperforms state-of-the-art RAG baselines in both answer accuracy and retrieval efficiency, while effectively mitigating erroneous information propagation.
๐ Abstract
Retrieval-augmented generation (RAG) addresses key limitations of large language models (LLMs), such as hallucinations and outdated knowledge, by incorporating external databases. These databases typically consult multiple sources to encompass up-to-date and various information. However, standard RAG methods often overlook the heterogeneous source reliability in the multi-source database and retrieve documents solely based on relevance, making them prone to propagating misinformation. To address this, we propose Reliability-Aware RAG (RA-RAG) which estimates the reliability of multiple sources and incorporates this information into both retrieval and aggregation processes. Specifically, it iteratively estimates source reliability and true answers for a set of queries with no labelling. Then, it selectively retrieves relevant documents from a few of reliable sources and aggregates them using weighted majority voting, where the selective retrieval ensures scalability while not compromising the performance. We also introduce a benchmark designed to reflect real-world scenarios with heterogeneous source reliability and demonstrate the effectiveness of RA-RAG compared to a set of baselines.