CNSight: Evaluation of Clinical Note Segmentation Tools

📅 2025-12-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Clinical notes exported from electronic medical records (EMRs) are predominantly unstructured or semi-structured, posing significant challenges for paragraph boundary detection—thereby hindering downstream tasks such as information extraction, cohort construction, and automatic summarization. This work presents the first systematic benchmark on a unified, curated subset of 1,000 MIMIC-IV cases, evaluating rule-based methods, domain-finetuned Transformers (e.g., BioClinicalBERT), and closed-source large language model APIs (e.g., GPT-5-mini). We introduce a dual-granularity evaluation framework incorporating both sentence-level and free-text segmentation annotations. Results show that GPT-5-mini achieves a mean F1-score of 72.4 on free-text paragraph segmentation—substantially outperforming lightweight models—while smaller models remain competitive on structured segmentation tasks. Our study demonstrates the paradigmatic advantage of LLMs in fine-grained clinical text segmentation and establishes a reproducible benchmark with methodological guidance for clinical NLP.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: SummarizationComputer Vision: Segmentation

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labeling
📝 Abstract
Clinical notes are often stored in unstructured or semi-structured formats after extraction from electronic medical record (EMR) systems, which complicates their use for secondary analysis and downstream clinical applications. Reliable identification of section boundaries is a key step toward structuring these notes, as sections such as history of present illness, medications, and discharge instructions each provide distinct clinical contexts. In this work, we evaluate rule-based baselines, domain-specific transformer models, and large language models for clinical note segmentation using a curated dataset of 1,000 notes from MIMIC-IV. Our experiments show that large API-based models achieve the best overall performance, with GPT-5-mini reaching a best average F1 of 72.4 across sentence-level and freetext segmentation. Lightweight baselines remain competitive on structured sentence-level tasks but falter on unstructured freetext. Our results provide guidance for method selection and lay the groundwork for downstream tasks such as information extraction, cohort identification, and automated summarization.
Problem

Research questions and friction points this paper is trying to address.

Evaluates segmentation tools for structuring clinical notes
Compares rule-based, transformer, and large language models
Guides method selection for downstream clinical applications
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evaluated rule-based, transformer, and large language models
Large API-based models achieved best performance for segmentation
Lightweight baselines competitive on structured sentence-level tasks
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Risha Surana
University of Southern California
A
Adrian Law
University of Southern California
S
Sunwoo Kim
University of Southern California
R
Rishab Sridhar
University of Southern California
A
Angxiao Han
University of Southern California
P
Peiyu Hong
University of Southern California