clinical concept extraction

Designs and builds models, pipelines, and tools that identify and extract clinical concepts and named entities (e.g., problems, medications, procedures, laboratory results, findings) from clinical text such as EHR notes. Analyzes and evaluates methods for candidate phrase detection, mention boundary identification, entity typing and normalization, and highlighting or retrieving relevant EHR details for downstream clinical NLP tasks.

clinicalconceptextraction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.21
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Extracting Patient History from Clinical Text: A Comparative Study of Clinical Large Language Models

Mar 30, 2025
HN
Hieu Nghiem
🏛️ Oklahoma State University | Moffitt Cancer Center and Research Institute | South Dakota School of Mines and Technology | University of South Florida Morsani College of Medicine | South University School of Pharmacy | Istinye University | The State University of New York Upstate Medical University | The State University of New York at New Paltz

This study addresses the automated extraction of Medical History Entities (MHEs)—including Chief Complaint (CC), History of Present Illness (HPI), and Past/Family/Social History (PFSH)—from clinical narratives to improve the conversion of unstructured electronic health records (EHRs) into standardized formats. Methodologically, it conducts the first systematic evaluation of seven clinical large language models (cLLMs), including fine-tuned GatorTron/GatorTronS and zero-shot GPT-4o, using a fine-grained manually annotated MTSamples dataset; error analysis and ablation studies examine impacts of text segmentation, entity length, and other linguistic features. A novel contribution is the integration of foundational medical entities (BMEs) as auxiliary signals. Results show that fine-tuned cLLMs reduce MHE extraction latency by over 20%; GatorTron variants achieve the highest performance; BME augmentation improves F1 scores for certain MHE types by up to 5.3%; and explicitly structured, title-annotated paragraphs significantly enhance extraction accuracy.

Assessing impact of note characteristics on entity recognition accuracyComparing fine-tuned cLLMs against GPT-4o in zero-shot settingsEvaluating clinical LLMs for extracting medical history entities from notes

Structured Semantics from Unstructured Notes: Language Model Approaches to EHR-Based Decision Support

Jun 01, 2025
WH
Wu Hao Ran
🏛️ Southern China University | University of the Chinese Academy of Sciences | Columbia University

This study addresses three critical challenges in electronic health record (EHR) analytics: (1) the limited utility of unstructured clinical text for high-quality clinical decision support; (2) cross-institutional semantic heterogeneity among EHR data; and (3) insufficient generalizability and fairness of medical AI models. To tackle these, we propose the first systematic, large language model (LLM)-driven framework that integrates heterogeneous EHR modalities—including free-text notes, structured laboratory values, and clinical codes. Our method introduces an ontology-guided, cross-institutional semantic alignment mechanism, coupled with interpretable fine-tuning and bias-correction strategies, to enable text-augmented multimodal representation learning. Evaluated on multicenter clinical prediction tasks, our framework achieves a mean AUC improvement of 5.2%, demonstrating enhanced model robustness. Furthermore, it exhibits superior predictive fairness across diverse demographic subgroups, validating its equitable performance in real-world heterogeneous healthcare settings.

Enhancing clinical decision support using language modelsEnsuring generalizability and fairness of healthcare AI modelsExtracting structured semantics from unstructured EHR notes

This work addresses the challenge of constructing interoperable patient digital twins from unstructured electronic health records (EHRs), which is hindered by clinical text heterogeneity and the lack of standardized mappings. The authors propose the first end-to-end semantic natural language processing (NLP) pipeline that tightly integrates with the Fast Healthcare Interoperability Resources (FHIR) standard. By combining named entity recognition, concept normalization to SNOMED-CT and ICD-10 terminologies, and relation extraction, the pipeline automatically transforms free-text clinical notes into structured FHIR resources. Evaluated on the MIMIC-IV Clinical Database Demo, the approach significantly improves F1 scores for both entity and relation extraction, outperforms baseline methods in schema completeness and system interoperability, and enables the automated construction of patient digital twins with high semantic consistency.

digital twinsFHIRinteroperability

Automated Detection of Clinical Entities in Lung and Breast Cancer Reports Using NLP Techniques

May 14, 2025
JM
J. Moreno-Casanova
🏛️ GMV | Health Research Institute Hospital La Fe

This study addresses the time-consuming, error-prone manual extraction of key clinical information from lung and breast cancer reports, which hinders the full utilization of healthcare data. To this end, we propose an end-to-end clinical natural language processing (NLP) system. Methodologically, we introduce the first joint application of the uQuery context-aware parsing engine and a fine-tuned RoBERTa model—bsc-bio-ehr-en3—on Spanish electronic health records (EHRs), enabling negation detection, temporal modeling, and patient-level semantic association. The system integrates named entity recognition (NER), standardized mapping to SNOMED CT and OMOP ontologies, and annotation via the Doccano platform to generate structured outputs. Evaluated on 600 real-world clinical reports, our approach achieves F1-scores exceeding 92% for both MET (metastasis) and PAT (pathology) critical entities, demonstrating strong cross-cancer generalizability and robustness in automated clinical information structuring.

Automating clinical data extraction from cancer reports using NLPEnhancing efficiency of EHR data processing for patient outcomesImproving accuracy in identifying lung and breast cancer entities

This work addresses three key challenges in zero-shot clinical named entity recognition (NER): fine-grained entity omission, class imbalance, and low recall for rare/long-tail entities. To this end, we propose the Entity Decomposition and Filtering (EDF) framework—the first of its kind—decoupling open-domain NER into two stages: subtype-aware retrieval and collaborative result filtering. EDF leverages open-source, NER-specialized large language models and integrates task decomposition with type-aware mechanisms. Extensive experiments across multiple clinical benchmarks demonstrate that EDF consistently outperforms all baseline methods across all evaluation metrics, model configurations, and entity types. Notably, it achieves substantial improvements in the recognition accuracy of rare and long-tail clinical entities, significantly enhancing zero-shot generalization capability.

Evaluating open NER LLMs in clinical entity recognitionImproving recognition of missed clinical entitiesIntroducing entity decomposition with filtering framework

Latest Papers

What's happening recently
View more

Leveraging LLMs for Structured Data Extraction from Unstructured Patient Records

Dec 03, 2025
MA
Mitchell A. Klusty
🏛️ University of Kentucky | University of Kentucky College of Medicine

In clinical research, manual extraction of structured clinical features from unstructured electronic health records (EHRs) is time-consuming, inefficient, and error-prone. To address this, we propose a privacy-preserving, modular, on-premises large language model (LLM) framework that integrates retrieval-augmented generation (RAG) with structured-output prompt engineering, enabling secure, scalable, containerized deployment in HIPAA-compliant environments. Our approach uniquely synergizes RAG with deterministic structured-response mechanisms for clinical text parsing—balancing domain adaptability and strict data privacy. Evaluated across multiple medical feature extraction tasks, the framework achieves high accuracy, substantially reducing manual annotation effort and improving data consistency. Notably, its systematic evaluation uncovered previously undetected systematic errors in prior human annotations, thereby validating both its reliability and its capacity for quality assurance and error discovery.

Automates structured data extraction from unstructured patient recordsEnhances data consistency and accuracy using LLMsReduces manual chart review burden in clinical research

This study addresses the challenges of accurately querying structured data and extracting information from unstructured clinical text in electronic health records (EHRs). To this end, the authors propose a unified framework that integrates large language models (LLMs) with retrieval-augmented generation (RAG): LLMs are employed to execute structured queries (e.g., Pandas operations), while RAG enhances information extraction from unstructured clinical narratives. The work introduces an innovative automatic evaluation pipeline based on synthetically generated question-answer pairs, combining exact match metrics, semantic similarity scores, and human assessments. Evaluated on a subset of MIMIC-III, the approach demonstrates improved semantic accuracy and task adaptability, offering clinical data science a flexible and reliable tool for automated reasoning and evaluation.

Clinical Data ScienceElectronic Health RecordsInformation Extraction

CNSight: Evaluation of Clinical Note Segmentation Tools

Dec 28, 2025
RS
Risha Surana
🏛️ University of Southern California

Clinical notes exported from electronic medical records (EMRs) are predominantly unstructured or semi-structured, posing significant challenges for paragraph boundary detection—thereby hindering downstream tasks such as information extraction, cohort construction, and automatic summarization. This work presents the first systematic benchmark on a unified, curated subset of 1,000 MIMIC-IV cases, evaluating rule-based methods, domain-finetuned Transformers (e.g., BioClinicalBERT), and closed-source large language model APIs (e.g., GPT-5-mini). We introduce a dual-granularity evaluation framework incorporating both sentence-level and free-text segmentation annotations. Results show that GPT-5-mini achieves a mean F1-score of 72.4 on free-text paragraph segmentation—substantially outperforming lightweight models—while smaller models remain competitive on structured segmentation tasks. Our study demonstrates the paradigmatic advantage of LLMs in fine-grained clinical text segmentation and establishes a reproducible benchmark with methodological guidance for clinical NLP.

Compares rule-based, transformer, and large language modelsEvaluates segmentation tools for structuring clinical notesGuides method selection for downstream clinical applications

Enhancing Clinical Note Generation with ICD-10, Clinical Ontology Knowledge Graphs, and Chain-of-Thought Prompting Using GPT-4

Dec 04, 2025
IM
Ivan Makohon
🏛️ Old Dominion University | University of Arkansas for Medical Sciences

Clinicians’ manual documentation of clinical notes is time-consuming, impeding diagnostic efficiency and patient experience. To address this, we propose a domain-informed, prompt-engineered method for automated clinical note generation: leveraging ICD-10 codes as input, we integrate a clinical ontology knowledge graph to enhance semantic understanding, and combine semantic retrieval with Chain-of-Thought prompting to guide GPT-4 in producing high-quality, professional, and structured notes. Our approach significantly improves medical accuracy, logical coherence, and clinical plausibility. Evaluations on six real-world clinical cases from the CodiEsp test set demonstrate substantial improvements over standard single-shot prompting—particularly in clinical professionalism, information completeness, and consistency with clinical practice. The method effectively reduces clinicians’ documentation burden and shows strong potential for clinical deployment.

Automating clinical note generation to reduce physician documentation timeEnhancing note quality through chain-of-thought prompting with domain knowledgeImproving AI-generated medical notes using structured clinical codes and ontologies

Balancing Natural Language Processing Accuracy and Normalisation in Extracting Medical Insights

Nov 19, 2025
PT
Paulina Tworek
🏛️ Sano - Centre for Computational Personalized Medicine | Jagiellonian University | AGH University of Krakow | Voivodeship Rehabilitation Hospital for Children in Ameryka | University of Warmia and Mazury

This study addresses the trade-off among accuracy, standardization, and computational efficiency in NLP for low-resource, non-English medical settings, using unstructured electronic health records from a Polish pediatric rehabilitation hospital. We propose a hybrid approach integrating rule-based systems—offering high precision and low computational cost—with multilingual large language models (LLMs)—providing strong generalization and adaptability. We systematically compare performance on demographic, clinical finding, and medication information extraction tasks using both original Polish text and machine-translated English text. Results show rule-based methods outperform LLMs in age and gender identification, while LLMs significantly improve drug name recognition accuracy. Critically, machine translation introduces non-negligible information loss, degrading downstream performance. This work establishes a new paradigm for resource-constrained, multilingual clinical NLP that balances accuracy, robustness, and practical deployability.

Assessing trade-offs between accuracy, normalization, and computational costsComparing rule-based methods and LLMs for EHR information extractionExtracting structured medical insights from unstructured clinical text

Hot Scholars

WL

Weixin Liu

Baidu Inc.
Natural Language ProcessingMachine LearningDeep Learning
ZL

Zhiyong Lu

Senior Investigator, NLM; Adjunct Professor of CS, UIUC
BioNLPBiomedical InformaticsMedical AIArtificial Intelligence
SY

Siyuan Yan

Research Fellow@Monash University
AI for MedicineFoundation Model
MK

Murat Kantarcioglu

Professor of Computer Science, Virginia Tech
Security and Privacy in AIDatabasesData ScienceComputer Security
WN

Weizhi Nie

Tianjin University
Medical Image ProcessingComputer VisionLLMs