Synthetic Hospital: An Open, Verifiable, Physician-Validated Longitudinal EHR Benchmark

πŸ“… 2026-09-24
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the lack of publicly shareable longitudinal electronic health record (EHR) benchmarks with verifiable ground truth for clinical AI. We construct the first fully synthetic longitudinal EHR benchmark derived from public medical education materials. By anchoring diagnoses to ICD-10/SNOMED standard ontologies and preserving complete provenance chains, the benchmark enables privacy-compliant open sharing alongside verifiable ground truth, while high-fidelity simulation of real-world hospital systems is achieved through API and access control technologies. Physician blinding tests yield near-random distinguishability, confirming exceptional data realism. The best-performing model achieves an F1 score of 0.73, matching the physician average yet remaining substantially below the expert upper bound, thereby underscoring the benchmark’s considerable challenge and evaluation value.
πŸ“ Abstract
Frontier language models are rarely used in clinical workflows because the realistic, longitudinal benchmarks needed to develop them are scarce. Real electronic health record (EHR) data cannot be openly shared due to privacy, ethics or data use issues and it does not contain verifiable ground truth since the chart records only reflect what clinicians documented. We introduce Synthetic Hospital, an open, fully synthetic, fact-grounded longitudinal EHR benchmark that resolves the open sharing and verifiable ground truth barriers. Built entirely from public medical-education material with no protected health information, it comprises 1,268 longitudinal patients and 5,602 encounters, where every diagnosis, finding, and temporal relation is grounded in standard ontologies (ICD-10-CM, SNOMED CT, LOINC) and with a complete provenance chain back to its source medical education material. Synthetic Hospital is served through a simulated hospital record system that mirrors real EHR infrastructure (standard interoperability APIs, role-based access and function-calling interface). In a blinded review, physicians distinguished its records from real patient charts at near-chance rates (53\%). Across 10 frontier and open models, none approaches ceiling: the best model reconstructs a patient's longitudinal problem list with a severity-weighted F1 of 0.73, level with the mean of seven physicians on a matched subset but well below the best of them (0.89), and misses roughly half of clinically relevant findings when summarizing a chart. Overall, these results highlight that Synthetic Hospital is a difficult and realistic test of clinical AI performance.
Problem

Research questions and friction points this paper is trying to address.

Electronic Health Records
Longitudinal Benchmark
Clinical AI
Synthetic Data
Ground Truth
Innovation

Methods, ideas, or system contributions that make the work stand out.

Synthetic EHR Benchmark
Longitudinal Patient Records
Verifiable Ground Truth
Standard Medical Ontologies
Clinical AI Evaluation
πŸ”Ž Similar Papers
No similar papers found.