A Living Benchmark for Information Retrieval from Electronic Health Records

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing Electronic Health Record (EHR) retrieval benchmarks, which rely on manual annotation and rapidly become outdated, thereby hindering effective evaluation of clinical large language models (LLMs). To overcome these challenges, this work proposes BRIE, a benchmarking framework that automatically generates question-answer pairs from longitudinal EHR notes. By integrating LLMs with expert validation mechanisms, BRIE enables automated benchmark generation, supports multiple reasoning variants per answer, and facilitates continuous dynamic refreshing to effectively prevent data leakage. Experimental evaluations reveal that mainstream models frequently overlook critical information when processing complex multi-document queries. These findings demonstrate that the proposed dynamic benchmark provides a more rigorous and reliable assessment of clinical LLM performance compared to conventional static approaches.
📝 Abstract
Large language model (LLM)-based clinical assistants are increasingly being integrated into electronic health record (EHR) systems, transforming how clinicians retrieve and synthesize information from patient records. Their safety and utility depend on rigorous evaluation, yet existing benchmarks are manually curated, costly to update, and rapidly become obsolete with evolving technological advancements. We present a scalable framework that automatically generates question--answer pairs from longitudinal EHR notes. Nineteen clinicians validate the benchmark generator, producing the Benchmark for Retrieving Information in EHRs (BRIE), a continuously maintainable evaluation dataset. Across nine LLMs and five inference strategies, state-of-the-art systems frequently omit clinically important information, particularly for questions requiring synthesis across multiple documents and encounters. Because the generator itself is validated, BRIE supports evaluations that static benchmarks cannot, including the generation of multiple answers that reflect variation in clinician reasoning for robust performance assessment and continuously refreshing benchmark content to guard against leakage. Our results demonstrate that scalable benchmark generation enables rigorous, up-to-date evaluation of clinical LLMs as they are deployed in rapidly evolving healthcare settings.
Problem

Research questions and friction points this paper is trying to address.

Electronic Health Records
Large Language Models
Clinical Benchmark
Information Retrieval
Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Electronic Health Records
Benchmark Generation
Large Language Models
Information Retrieval
Clinical Evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jordan L. Cahoon
Department of Biomedical Data Science, Stanford University, Stanford, CA; Department of Pathology, Stanford University, Stanford, CA
C
Chloe O. Stanwyck
Department of Biomedical Data Science, Stanford University, Stanford, CA; Department of Anesthesiology, Perioperative and Pain Medicine, Stanford University, Stanford, CA
S
Sulaiman Somani
Department of Medicine, Stanford University, Stanford, CA
Philip Chung
Philip Chung
Department of Anesthesiology, Perioperative and Pain Medicine, Stanford University, Stanford, CA
K
Kevin R Keet
Department of Medicine, Stanford University, Stanford, CA
K
Kameron C. Black
Department of Medicine, Stanford University, Stanford, CA
A
Andrea T. Fisher
Department of Surgery, Stanford University, Stanford, CA
S
Sarita Khemani
Department of Medicine, Stanford University, Stanford, CA
Jerry Liu
Jerry Liu
Stanford University
S
Stephen Ma
Department of Medicine, Stanford University, Stanford, CA
S
Saloni K. Maharaj
Department of Medicine, Stanford University, Stanford, CA
R
Rita M. Pandya
Department of Medicine, Stanford University, Stanford, CA
E
Eduardo Perez-Guerrero
Department of Medicine, Stanford University, Stanford, CA
P
Priyanka Pillai
Department of Medicine, Stanford University, Stanford, CA
L
Lisa Shieh
Department of Medicine, Stanford University, Stanford, CA
D
David J. H. Wu
Department of Radiation Oncology, Stanford Cancer Center, Palo Alto, CA, USA
J
James Xie
Department of Anesthesiology, Perioperative and Pain Medicine, Stanford University, Stanford, CA
J
James C. McAvoy
Department of Anesthesiology, Perioperative and Pain Medicine, Stanford University, Stanford, CA
T
Teresa Nguyen
Department of Anesthesiology, Perioperative and Pain Medicine, Stanford University, Stanford, CA
J
Jessica Tran
Department of Medicine, Stanford University, Stanford, CA
L
Lucy Yin
Department of Biomedical Data Science, Stanford University, Stanford, CA
B
Bridget Lin
Department of Biomedical Data Science, Stanford University, Stanford, CA
Alison Callahan
Alison Callahan
Stanford University
electronic health recordsclinical decision supportdata miningcausal inferencemachine learning
J
Jason A. Fries
Department of Biomedical Data Science, Stanford University, Stanford, CA; Department of Medicine, Stanford University, Stanford, CA; Weill Cancer Hub West
N
Nigam H. Shah
Department of Biomedical Data Science, Stanford University, Stanford, CA; Department of Medicine, Stanford University, Stanford, CA; Center for Clinical Excellence Research, Stanford School of Medicine, Stanford, CA, USA