๐ค AI Summary
This study addresses the automated extraction of Medical History Entities (MHEs)โincluding Chief Complaint (CC), History of Present Illness (HPI), and Past/Family/Social History (PFSH)โfrom clinical narratives to improve the conversion of unstructured electronic health records (EHRs) into standardized formats. Methodologically, it conducts the first systematic evaluation of seven clinical large language models (cLLMs), including fine-tuned GatorTron/GatorTronS and zero-shot GPT-4o, using a fine-grained manually annotated MTSamples dataset; error analysis and ablation studies examine impacts of text segmentation, entity length, and other linguistic features. A novel contribution is the integration of foundational medical entities (BMEs) as auxiliary signals. Results show that fine-tuned cLLMs reduce MHE extraction latency by over 20%; GatorTron variants achieve the highest performance; BME augmentation improves F1 scores for certain MHE types by up to 5.3%; and explicitly structured, title-annotated paragraphs significantly enhance extraction accuracy.
๐ Abstract
Extracting medical history entities (MHEs) related to a patient's chief complaint (CC), history of present illness (HPI), and past, family, and social history (PFSH) helps structure free-text clinical notes into standardized EHRs, streamlining downstream tasks like continuity of care, medical coding, and quality metrics. Fine-tuned clinical large language models (cLLMs) can assist in this process while ensuring the protection of sensitive data via on-premises deployment. This study evaluates the performance of cLLMs in recognizing CC/HPI/PFSH-related MHEs and examines how note characteristics impact model accuracy. We annotated 1,449 MHEs across 61 outpatient-related clinical notes from the MTSamples repository. To recognize these entities, we fine-tuned seven state-of-the-art cLLMs. Additionally, we assessed the models' performance when enhanced by integrating, problems, tests, treatments, and other basic medical entities (BMEs). We compared the performance of these models against GPT-4o in a zero-shot setting. To further understand the textual characteristics affecting model accuracy, we conducted an error analysis focused on note length, entity length, and segmentation. The cLLMs showed potential in reducing the time required for extracting MHEs by over 20%. However, detecting many types of MHEs remained challenging due to their polysemous nature and the frequent involvement of non-medical vocabulary. Fine-tuned GatorTron and GatorTronS, two of the most extensively trained cLLMs, demonstrated the highest performance. Integrating pre-identified BME information improved model performance for certain entities. Regarding the impact of textual characteristics on model performance, we found that longer entities were harder to identify, note length did not correlate with a higher error rate, and well-organized segments with headings are beneficial for the extraction.