Introducing Answered with Evidence -- a framework for evaluating whether LLM responses to biomedical questions are founded in evidence

📅 2025-06-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of scientifically grounded evidence supporting large language model (LLM) answers in biomedical question answering. Methodologically, we propose a multi-source retrieval-augmented generation (RAG) framework that integrates heterogeneous biomedical literature sources—including the novel observational evidence repository Alexandria (formerly Atropos Evidence Library), PubMed, and Perplexity—to conduct empirical analysis on real-world physician questions. Our key contributions are: (1) the first incorporation of dynamic observational study evidence into LLM evaluation, substantially broadening the scope of evidence coverage; and (2) cross-source validation revealing that PubMed alone supports answers to ~44% of questions, Alexandria supports ~50%, and their combination enables evidence-based, reliable answers for over 70% of questions. This establishes a reproducible benchmark and technical pathway for trustworthy biomedical AI.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Reasoning under Uncertainty: Other Foundations of Reasoning under UncertaintyNatural Language Processing: Question Answering

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
The growing use of large language models (LLMs) for biomedical question answering raises concerns about the accuracy and evidentiary support of their responses. To address this, we present Answered with Evidence, a framework for evaluating whether LLM-generated answers are grounded in scientific literature. We analyzed thousands of physician-submitted questions using a comparative pipeline that included: (1) Alexandria, fka the Atropos Evidence Library, a retrieval-augmented generation (RAG) system based on novel observational studies, and (2) two PubMed-based retrieval-augmented systems (System and Perplexity). We found that PubMed-based systems provided evidence-supported answers for approximately 44% of questions, while the novel evidence source did so for about 50%. Combined, these sources enabled reliable answers to over 70% of biomedical queries. As LLMs become increasingly capable of summarizing scientific content, maximizing their value will require systems that can accurately retrieve both published and custom-generated evidence or generate such evidence in real time.
Problem

Research questions and friction points this paper is trying to address.

Evaluating evidence support in LLM biomedical answers
Comparing retrieval systems for scientific literature grounding
Improving reliability of LLM responses to medical questions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Framework evaluates LLM answers with evidence
Uses retrieval-augmented generation (RAG) systems
Combines PubMed and novel evidence sources
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Julian D Baldwin
Atropos Health; New York, NY, USA
C
Christina Dinh
Atropos Health; New York, NY, USA
A
Arjun Mukerji
Atropos Health; New York, NY, USA
N
Neil Sanghavi
Atropos Health; New York, NY, USA
S
Saurabh Gombar
Stanford School of Medicine - Department of Pathology, Stanford University, Stanford, California, USA