clinical terminology curation

Designs, builds, and maintains clinical interface terminologies and curated concept sets for electronic health records, including hierarchical groupings, synonyms, abbreviations, and mappings to standard vocabularies. Uses semi-automated extraction and human review workflows to refine candidate concepts, validate mappings, and produce deployment- and training-ready clinical vocabularies.

clinicalterminologycuration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of information overload and complex terminology in electronic health records (EHRs), which can lead clinicians to overlook critical details. To mitigate this, the authors propose a three-stage, minimally supervised approach for constructing a Cardiology Interface Terminology (CIT). The method first integrates SNOMED CT subhierarchies, EHR-derived concepts, and terminological components to generate an initial CIT. It then iteratively extracts phrases and employs semi-automated review to build a training set (TCIT). Finally, a supervised machine learning model is trained on TCIT to expand the CIT and highlight relevant terms in EHRs. This novel integration of semi-automated term curation with machine learning achieves 74.21% coverage and a breadth of 1.68 on the test set, with manual evaluation of 20 random clinical notes showing 98.2% average completeness and 84.2% conciseness.

Cardiology Interface TerminologyClinical TerminologyDetail Extraction

Efficient Standardization of Clinical Notes using Large Language Models

Dec 31, 2024
DB
D. B. Hier
🏛️ Missouri University of Science & Technology | University of Illinois at Chicago | Missouri State University

Clinical handwritten notes suffer from inconsistent handwriting, pervasive abbreviations, nonstandard terminology, grammatical and spelling errors, and disorganized formatting—severely hindering information extraction and interoperability in electronic health records (EHRs). To address this, we propose the first end-to-end LLM-based framework for clinical note standardization. It integrates terminology mapping, medical ontology alignment, and rule-augmented generative text reconstruction to jointly perform grammar/spelling correction, normalization of nonstandard clinical terms, abbreviation expansion, and structural reformatting—natively supporting interoperability standards such as FHIR. Evaluated on 1,618 real-world clinical notes, our method corrects on average 4.9 grammatical errors, 3.3 spelling errors, 3.1 nonstandard terms, and 15.8 abbreviations per note. Expert evaluation confirms high semantic fidelity and negligible information loss, significantly improving readability and downstream task performance.

Electronic Health Records AnalysisHandwritten Notes RecognitionMedical Information Extraction

Structured Semantics from Unstructured Notes: Language Model Approaches to EHR-Based Decision Support

Jun 01, 2025
WH
Wu Hao Ran
🏛️ Southern China University | University of the Chinese Academy of Sciences | Columbia University

This study addresses three critical challenges in electronic health record (EHR) analytics: (1) the limited utility of unstructured clinical text for high-quality clinical decision support; (2) cross-institutional semantic heterogeneity among EHR data; and (3) insufficient generalizability and fairness of medical AI models. To tackle these, we propose the first systematic, large language model (LLM)-driven framework that integrates heterogeneous EHR modalities—including free-text notes, structured laboratory values, and clinical codes. Our method introduces an ontology-guided, cross-institutional semantic alignment mechanism, coupled with interpretable fine-tuning and bias-correction strategies, to enable text-augmented multimodal representation learning. Evaluated on multicenter clinical prediction tasks, our framework achieves a mean AUC improvement of 5.2%, demonstrating enhanced model robustness. Furthermore, it exhibits superior predictive fairness across diverse demographic subgroups, validating its equitable performance in real-world heterogeneous healthcare settings.

Enhancing clinical decision support using language modelsEnsuring generalizability and fairness of healthcare AI modelsExtracting structured semantics from unstructured EHR notes

Enhancing Clinical Note Generation with ICD-10, Clinical Ontology Knowledge Graphs, and Chain-of-Thought Prompting Using GPT-4

Dec 04, 2025
IM
Ivan Makohon
🏛️ Old Dominion University | University of Arkansas for Medical Sciences

Clinicians’ manual documentation of clinical notes is time-consuming, impeding diagnostic efficiency and patient experience. To address this, we propose a domain-informed, prompt-engineered method for automated clinical note generation: leveraging ICD-10 codes as input, we integrate a clinical ontology knowledge graph to enhance semantic understanding, and combine semantic retrieval with Chain-of-Thought prompting to guide GPT-4 in producing high-quality, professional, and structured notes. Our approach significantly improves medical accuracy, logical coherence, and clinical plausibility. Evaluations on six real-world clinical cases from the CodiEsp test set demonstrate substantial improvements over standard single-shot prompting—particularly in clinical professionalism, information completeness, and consistency with clinical practice. The method effectively reduces clinicians’ documentation burden and shows strong potential for clinical deployment.

Automating clinical note generation to reduce physician documentation timeEnhancing note quality through chain-of-thought prompting with domain knowledgeImproving AI-generated medical notes using structured clinical codes and ontologies

This work addresses the challenge of constructing interoperable patient digital twins from unstructured electronic health records (EHRs), which is hindered by clinical text heterogeneity and the lack of standardized mappings. The authors propose the first end-to-end semantic natural language processing (NLP) pipeline that tightly integrates with the Fast Healthcare Interoperability Resources (FHIR) standard. By combining named entity recognition, concept normalization to SNOMED-CT and ICD-10 terminologies, and relation extraction, the pipeline automatically transforms free-text clinical notes into structured FHIR resources. Evaluated on the MIMIC-IV Clinical Database Demo, the approach significantly improves F1 scores for both entity and relation extraction, outperforms baseline methods in schema completeness and system interoperability, and enables the automated construction of patient digital twins with high semantic consistency.

digital twinsFHIRinteroperability

Latest Papers

What's happening recently
View more

This study addresses the critical issue of documentation inconsistencies in electronic health records (EHRs), which can compromise clinical decision-making and patient safety. The authors propose the first hierarchical ontology framework specifically designed for EHR inconsistencies, capturing a spectrum ranging from strict contradictions to ambiguous discrepancies. They introduce a fine-grained annotation schema structured along four axes: category, section, clinical domain, and inconsistency type. Leveraging this framework, they develop a two-stage large language model pipeline: Gemini 2.5 Pro first identifies candidate inconsistencies, followed by context-anchored validation using Gemini 2.5 Flash. Applied to 3,000 discharge summaries from MIMIC-IV-Note, the pipeline detected 3,460 inconsistencies across 69.7% of records, predominantly involving demographics, allergies, diagnoses, and medications, while also exposing systematic model limitations in temporal reasoning and outpatient medication knowledge.

clinical safetydischarge summariesdocumentation inconsistencies

This study addresses the inefficiency and error-proneness of information extraction in medical form filling by proposing CLAIRE, a hybrid workflow. The method adopts a "schema-grounded, verification-first" architecture that automates form completion through field state discovery, source-to-field mapping, deterministic validation, and audit trails. Large language models from the Qwen series are restricted to assisting with mapping tasks without authorization for critical operations, while a bounded error-correction mechanism ensures data rigor. Benchmark evaluations demonstrate that CLAIRE achieves both a success rate and an accuracy of 1.000, substantially reducing the burden of manual review and enhancing processing throughput.

data qualityelectronic health recordsform completion

Existing clinical reasoning evaluation benchmarks predominantly rely on unstructured or static data, failing to capture the structured and interoperable nature of real-world electronic health records (EHRs). To address this gap, this work proposes a novel pipeline that integrates staged large language model (LLM) generation with terminology-anchored validation and repair, yielding MedCase-Structured—the first HL7 FHIR R4–compliant structured dataset for clinical reasoning assessment. Built upon the MedCaseReasoning benchmark, the pipeline successfully generates valid FHIR bundles for 82.5% of cases. Experimental results demonstrate that LLMs exhibit significantly lower diagnostic accuracy when provided with structured FHIR inputs compared to plain text, underscoring both the necessity of aligning evaluations with authentic clinical workflows and the innovative contribution of this dataset.

benchmarkingclinical reasoningelectronic health records

This study addresses inconsistencies between medication information in clinical notes and structured electronic health records (EHRs), which primarily arise from terminological discrepancies, temporal misalignment, and divergent documentation practices. To systematically evaluate concordance between these data sources, the authors propose a hybrid approach integrating large language models (LLMs) with human expert review to construct a high-quality reference standard. Their methodology combines deterministic medication normalization, mapping to the OMOP ontology, and semantic–temporal alignment analysis. This framework improves the exact match rate after standardization from 0.7226 to 0.8429, achieves a semantic overlap of 55.17%, and reduces the proportion of strictly non-overlapping entries to 3.97%. The findings highlight terminology variation and normalization challenges as key drivers of discordance, offering a scalable foundation for multimodal healthcare data integration.

clinical notesdocumentation timingmedication reconciliation

Hot Scholars

BS

Bolan Su

Senior Researcher, Tencent
Image ProcessingComputer VisionReinforcement LearningRecommender System
JG

Jatin Gupta

Sharda University
AIMLDeep LearningLLMs
ZA

Zahra Atf

Interdisciplinary Researcher, PhD in Business.
Digital MarketingExplainable AILLMsMoral Philosophy
PR

Peter Rohloff

Brigham and Women's Hospital, Maya Health Alliance
global healthGuatemalaindigenous healthdiabetes
WL

Wenda Li

University of Edinburgh
theorem provingformal verificationmachine learning