🤖 AI Summary
Clinical notes exported from electronic medical records (EMRs) are predominantly unstructured or semi-structured, posing significant challenges for paragraph boundary detection—thereby hindering downstream tasks such as information extraction, cohort construction, and automatic summarization. This work presents the first systematic benchmark on a unified, curated subset of 1,000 MIMIC-IV cases, evaluating rule-based methods, domain-finetuned Transformers (e.g., BioClinicalBERT), and closed-source large language model APIs (e.g., GPT-5-mini). We introduce a dual-granularity evaluation framework incorporating both sentence-level and free-text segmentation annotations. Results show that GPT-5-mini achieves a mean F1-score of 72.4 on free-text paragraph segmentation—substantially outperforming lightweight models—while smaller models remain competitive on structured segmentation tasks. Our study demonstrates the paradigmatic advantage of LLMs in fine-grained clinical text segmentation and establishes a reproducible benchmark with methodological guidance for clinical NLP.
📝 Abstract
Clinical notes are often stored in unstructured or semi-structured formats after extraction from electronic medical record (EMR) systems, which complicates their use for secondary analysis and downstream clinical applications. Reliable identification of section boundaries is a key step toward structuring these notes, as sections such as history of present illness, medications, and discharge instructions each provide distinct clinical contexts. In this work, we evaluate rule-based baselines, domain-specific transformer models, and large language models for clinical note segmentation using a curated dataset of 1,000 notes from MIMIC-IV. Our experiments show that large API-based models achieve the best overall performance, with GPT-5-mini reaching a best average F1 of 72.4 across sentence-level and freetext segmentation. Lightweight baselines remain competitive on structured sentence-level tasks but falter on unstructured freetext. Our results provide guidance for method selection and lay the groundwork for downstream tasks such as information extraction, cohort identification, and automated summarization.