Retrieval-Augmented Generation with Estimation of Source Reliability

๐Ÿ“… 2024-10-30
๐Ÿ›๏ธ arXiv.org
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Existing RAG methods overlook the heterogeneity of source reliability in multi-source retrieval, relying solely on relevance-based rankingโ€”leading to hallucinations and erroneous information propagation. To address this, we propose RA-RAG, a reliability-aware RAG framework that, for the first time, jointly models and iteratively optimizes source reliability and answer truthfulness in a fully unsupervised manner. RA-RAG introduces three core components: (i) a reliability estimation mechanism, (ii) a reliability-constrained scalable retrieval strategy, and (iii) a reliability-weighted majority voting aggregation method. To rigorously evaluate source reliability heterogeneity, we construct the first benchmark evaluation framework specifically designed for this challenge. Extensive experiments demonstrate that RA-RAG significantly outperforms state-of-the-art RAG baselines in both answer accuracy and retrieval efficiency, while effectively mitigating erroneous information propagation.

Technology Category

Reasoning under Uncertainty: Relational Probabilistic ModelsMultiagent Systems: Multiagent Systems under UncertaintyKnowledge Representation and Reasoning: Reasoning with Beliefs

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGWeb Mining and Content Analysis: Web data provenance, reliability, and authenticityUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
๐Ÿ“ Abstract
Retrieval-augmented generation (RAG) addresses key limitations of large language models (LLMs), such as hallucinations and outdated knowledge, by incorporating external databases. These databases typically consult multiple sources to encompass up-to-date and various information. However, standard RAG methods often overlook the heterogeneous source reliability in the multi-source database and retrieve documents solely based on relevance, making them prone to propagating misinformation. To address this, we propose Reliability-Aware RAG (RA-RAG) which estimates the reliability of multiple sources and incorporates this information into both retrieval and aggregation processes. Specifically, it iteratively estimates source reliability and true answers for a set of queries with no labelling. Then, it selectively retrieves relevant documents from a few of reliable sources and aggregates them using weighted majority voting, where the selective retrieval ensures scalability while not compromising the performance. We also introduce a benchmark designed to reflect real-world scenarios with heterogeneous source reliability and demonstrate the effectiveness of RA-RAG compared to a set of baselines.
Problem

Research questions and friction points this paper is trying to address.

Estimates source reliability in multi-source databases
Enhances retrieval-augmented generation accuracy
Reduces misinformation propagation in language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Estimates source reliability iteratively
Selectively retrieves from reliable sources
Uses weighted majority voting aggregation
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
Pohang University of Science and Technology (POSTECH)
Jeongyeon Hwang
Jeongyeon Hwang
Ph.D student at POSTECH
Machine Learning
J
Junyoung Park
Pohang University of Science and Technology (POSTECH), South Korea
H
Hyejin Park
Pohang University of Science and Technology (POSTECH), South Korea
S
Sangdon Park
Pohang University of Science and Technology (POSTECH), South Korea
Jungseul Ok
Jungseul Ok
Associate Professor, CSE/AI, POSTECH
Reinforcement LearningMachine Learning