Clinical Note Bloat Reduction for Efficient LLM Use

📅 2026-03-21
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges posed by templated and duplicated content in clinical notes, which dilutes informative signals, constrains context windows, and escalates inference costs for large language models (LLMs). This work presents the first systematic quantification of electronic health record (EHR) metadata for text deduplication, leveraging it to precisely identify and remove redundant information to extend longitudinal context. A comprehensive evaluation is conducted using the TRACE algorithm, frequency-based deduplication, zero-shot LLMs, and embedding classifiers. Results demonstrate that the proposed approach reduces text volume by 47.3% without compromising model performance, with projected computational savings exceeding one million dollars over three years. Ultimately, this research establishes an efficient data preprocessing paradigm for deploying large-scale models in healthcare applications.
📝 Abstract
Background: Clinical notes contain extensive duplicated text from templates, copy-paste, and auto-populated fields ("note bloat"), diluting clinical signal, limiting longitudinal context, and increasing large language model (LLM) costs. Methods: TRACE removes note bloat using note-level EHR metadata to identify templated and copied content, with frequency-based de-duplication when metadata are unavailable. We evaluated TRACE using blinded physician span review and gold-standard templated-text annotations across four cohorts spanning liver transplant, obstetrics, and inpatient populations at multiple health systems (5.3M notes). We compared zero-shot LLMs and embedding-based classifiers using original and TRACE-processed notes for 20 information extraction tasks and prediction of 5-year survival, postpartum hemorrhage, and 30-day readmission. Results: Only 0.3-6.6% of removed text was flagged as author-generated; TRACE captured 86% of annotated templated characters. Information extraction F1 differences averaged by cohort ranged from -0.009 to +0.004; task-specific prediction F1 differences ranged from -0.011 to +0.018. Among 1,000 randomly sampled Stanford Health Care patients, TRACE reduced chart text by 47.3% (742.7M characters), averaging 220,167 fewer tokens per patient. Using 2024 encounter volumes at a large tertiary academic center and one query per encounter, projected three-year net savings ranged from $1.00M to $13.58M across evaluated model pricing schemes, including initial and annual TRACE processing costs. Conclusion: TRACE substantially reduces clinical note redundancy while preserving information extraction and prediction performance. Underused EHR metadata can reduce LLM inference costs, expand usable longitudinal context, and support scalable clinical AI.
Problem

Research questions and friction points this paper is trying to address.

clinical note bloat
text redundancy
large language model costs
electronic health records
Innovation

Methods, ideas, or system contributions that make the work stand out.

Clinical Note Bloat Reduction
EHR Metadata
De-duplication
Large Language Model Efficiency
Information Extraction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jordan L. Cahoon
Department of Biomedical Data Science, Stanford University, Stanford, CA
C
Chloe Stanwyck
Department of Anesthesiology, Perioperative and Pain Medicine, Stanford University, Stanford, CA
A
Asad Aali
Department of Radiology, Stanford University, Stanford, CA
R
Rachel Madding
Department of Obstetrics and Gynecology, Stanford University, Stanford, CA
S
Sulaiman S. Somani
Department of Medicine, Stanford University, Stanford, CA
E
Emma Sun
Department of Computer Science, Stanford University, Stanford, CA
Yixing Jiang
Yixing Jiang
Stanford
R
Renumathy Dhanasekaran
Division of Gastroenterology and Hepatology, Stanford University, Stanford, CA
Emily Alsentzer
Emily Alsentzer
Assistant Professor, Stanford University
machine learning for healthcare