When Evidence Contradicts: Toward Safer Retrieval-Augmented Generation in Healthcare

📅 2025-11-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the declining factual accuracy of Retrieval-Augmented Generation (RAG) in high-stakes medical applications, where outdated or contradictory source documents undermine reliability. We introduce a novel medical question-answering benchmark built upon Australian Therapeutic Goods Administration (TGA) product labeling and propose time-stratified PubMed abstract retrieval to enable controlled evaluation of outdated evidence. Experiments compare five state-of-the-art LLMs on integrating temporally dispersed, semantically similar yet contradictory medical literature. Results demonstrate that lexical or semantic similarity alone does not ensure RAG reliability in medicine; contradictory evidence significantly degrades both answer accuracy and inter-response consistency. Our key contribution is the empirical revelation that “similarity ≠ reliability” — a critical limitation in medical RAG — and the first systematic validation of the necessity of contradiction-aware filtering mechanisms. This work provides empirical grounding and methodological guidance for enhancing RAG safety in high-risk domains.

Technology Category

Reasoning under Uncertainty: Other Foundations of Reasoning under UncertaintyKnowledge Representation and Reasoning: Reasoning with BeliefsData Mining & Knowledge Management: Conversational Systems for Recommendation & Retrieval

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
In high-stakes information domains such as healthcare, where large language models (LLMs) can produce hallucinations or misinformation, retrieval-augmented generation (RAG) has been proposed as a mitigation strategy, grounding model outputs in external, domain-specific documents. Yet, this approach can introduce errors when source documents contain outdated or contradictory information. This work investigates the performance of five LLMs in generating RAG-based responses to medicine-related queries. Our contributions are three-fold: i) the creation of a benchmark dataset using consumer medicine information documents from the Australian Therapeutic Goods Administration (TGA), where headings are repurposed as natural language questions, ii) the retrieval of PubMed abstracts using TGA headings, stratified across multiple publication years, to enable controlled temporal evaluation of outdated evidence, and iii) a comparative analysis of the frequency and impact of outdated or contradictory content on model-generated responses, assessing how LLMs integrate and reconcile temporally inconsistent information. Our findings show that contradictions between highly similar abstracts do, in fact, degrade performance, leading to inconsistencies and reduced factual accuracy in model answers. These results highlight that retrieval similarity alone is insufficient for reliable medical RAG and underscore the need for contradiction-aware filtering strategies to ensure trustworthy responses in high-stakes domains.
Problem

Research questions and friction points this paper is trying to address.

Detect outdated medical information in retrieval-augmented generation
Evaluate LLM performance with contradictory healthcare evidence
Develop contradiction-aware filtering for reliable medical responses
Innovation

Methods, ideas, or system contributions that make the work stand out.

Benchmark dataset from TGA medicine documents
Retrieval of PubMed abstracts across publication years
Contradiction-aware filtering for trustworthy medical responses
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Saeedeh Javadi
Saeedeh Javadi
PhD Student, RMIT University
AIGNNLLMKnowledge graphMachine Learning
S
Sara Mirabi
Deakin University, Melbourne, Australia
M
Manan Gangar
Deakin University, Melbourne, Australia
B
B. Ofoghi
Deakin University, Melbourne, Australia