Evaluating Biomedical Reranking for LLM-Based Question Answering over Longitudinal Clinical Notes

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of degraded patient question-answering quality caused by fragmented evidence and heterogeneous terminology in longitudinal clinical records by constructing a localized retrieval-augmented generation (RAG) pipeline. Methodologically, it integrates PubMedBERT-based dense retrieval with BM25 lexical retrieval via weighted reciprocal rank fusion, and incorporates a MedCPT cross-encoder for biomedical reranking to optimize evidence selection. Results demonstrate that the top-10 retrieval hit rate increases from 46.6% to 60.6%, while the answer accuracy of Qwen3-8B improves from 44.8% to 48.6%. This work validates the effectiveness of reranking for context optimization and reveals a nonlinear relationship between retrieval performance gains and downstream answer improvements.
📝 Abstract
Patient-specific clinical question answering requires locating the right evidence within long, heterogeneous longitudinal clinical records in which relevant facts may be scattered across encounters, repeated in copied-forward notes, or expressed using different clinical terminology. We evaluated whether biomedical reranking can improve evidence selection and downstream answer quality in a locally deployed retrieval-augmented generation pipeline for longitudinal clinical notes. The pipeline combines PubMedBERT dense retrieval, BM25 lexical retrieval, weighted reciprocal-rank fusion, and MedCPT cross-encoder reranking. Across 1,000 open- and closed-ended question-answer pairs from a cohort of 200 bariatric surgery patients, reranking increased exact source-chunk retrieval within the top 10 items, Hit@10 from 46.6% to 60.6% and mean reciprocal rank from 0.2371 to 0.3252. With Qwen3-8B generation, local judge-assessed answer correctness increased from 44.8% to 48.6%. These results show that biomedical reranking can improve the placement of relevant clinical evidence within a limited context window, although gains in retrieval do not translate proportionally into gains in answer correctness.
Problem

Research questions and friction points this paper is trying to address.

clinical question answering
longitudinal clinical notes
biomedical reranking
retrieval-augmented generation
evidence selection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Biomedical Reranking
Retrieval-Augmented Generation
Longitudinal Clinical Notes
Reciprocal-Rank Fusion
Cross-Encoder
💼 Related Jobs
No related jobs found.
M
Maryam Shahbaz Ali
Children’s National Hospital, Washington, DC, USA
L
Laura B. Strachan
Children’s National Hospital, Washington, DC, USA; University of Florida, Gainesville, FL, USA
C
Caitlin Sherman
Children’s National Hospital, Washington, DC, USA
M
Mark Kovler
Children’s National Hospital, Washington, DC, USA
E
Eleanor Mackey
Children’s National Hospital, Washington, DC, USA; George Washington University, Washington DC, USA
Syed Muhammad Anwar
Syed Muhammad Anwar
Childrens National Hospital/George Washington University
Biomedical Signal processingmedical image analysisgraph learningself-supervised learning