🤖 AI Summary
This study addresses the challenge of degraded patient question-answering quality caused by fragmented evidence and heterogeneous terminology in longitudinal clinical records by constructing a localized retrieval-augmented generation (RAG) pipeline. Methodologically, it integrates PubMedBERT-based dense retrieval with BM25 lexical retrieval via weighted reciprocal rank fusion, and incorporates a MedCPT cross-encoder for biomedical reranking to optimize evidence selection. Results demonstrate that the top-10 retrieval hit rate increases from 46.6% to 60.6%, while the answer accuracy of Qwen3-8B improves from 44.8% to 48.6%. This work validates the effectiveness of reranking for context optimization and reveals a nonlinear relationship between retrieval performance gains and downstream answer improvements.
📝 Abstract
Patient-specific clinical question answering requires locating the right evidence within long, heterogeneous longitudinal clinical records in which relevant facts may be scattered across encounters, repeated in copied-forward notes, or expressed using different clinical terminology. We evaluated whether biomedical reranking can improve evidence selection and downstream answer quality in a locally deployed retrieval-augmented generation pipeline for longitudinal clinical notes. The pipeline combines PubMedBERT dense retrieval, BM25 lexical retrieval, weighted reciprocal-rank fusion, and MedCPT cross-encoder reranking. Across 1,000 open- and closed-ended question-answer pairs from a cohort of 200 bariatric surgery patients, reranking increased exact source-chunk retrieval within the top 10 items, Hit@10 from 46.6% to 60.6% and mean reciprocal rank from 0.2371 to 0.3252. With Qwen3-8B generation, local judge-assessed answer correctness increased from 44.8% to 48.6%. These results show that biomedical reranking can improve the placement of relevant clinical evidence within a limited context window, although gains in retrieval do not translate proportionally into gains in answer correctness.