MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clinically Realistic EHR Settings
Existing clinical reasoning evaluation benchmarks predominantly rely on unstructured or static data, failing to capture the structured and interoperable nature of real-world electronic health records (EHRs). To address this gap, this work proposes a novel pipeline that integrates staged large language model (LLM) generation with terminology-anchored validation and repair, yielding MedCase-Structured—the first HL7 FHIR R4–compliant structured dataset for clinical reasoning assessment. Built upon the MedCaseReasoning benchmark, the pipeline successfully generates valid FHIR bundles for 82.5% of cases. Experimental results demonstrate that LLMs exhibit significantly lower diagnostic accuracy when provided with structured FHIR inputs compared to plain text, underscoring both the necessity of aligning evaluations with authentic clinical workflows and the innovative contribution of this dataset.