nlp model development and adaptation

Designs and implements natural language processing models and adaptation pipelines that process and normalize clinical text (e.g., notes and reports) to parse and extract structured information such as diagnoses, temporal events, and treatment-related entities. Builds preprocessing and annotation-correction workflows, adapts models to new clinical subdomains or annotation schemes, and maps clinical language to standardized representations or translations for use in downstream predictive models.

nlpmodeldevelopmentand

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.35
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$195K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

ELMTEX: Fine-Tuning Large Language Models for Structured Clinical Information Extraction. A Case Study on Clinical Reports

Feb 08, 2025
AG
Aynur Guluzade
🏛️ Fraunhofer Institute for Applied Information Technology

This study addresses interoperability challenges in European healthcare systems by proposing a large language model (LLM)-based approach for multilingual clinical text information extraction. To process unstructured clinical reports, we design an end-to-end extraction pipeline with an interactive interface and jointly apply prompt engineering and supervised fine-tuning to adapt medium- and small-scale LLMs. Our key contributions are threefold: (1) We release the first bilingual clinical summarization dataset—comprising 60,000 English and 24,000 German samples—validated via multi-dimensional automatic and human annotation; (2) We propose a multi-granularity evaluation framework integrating ROUGE, BERTScore, and entity-level metrics; (3) Our fine-tuned model achieves 89.3% F1 on critical entity extraction—outperforming zero-shot LLMs by 22.7%—while accelerating inference by 3.8× and reducing GPU memory consumption by 64%, significantly enhancing clinical deployability.

Enhance healthcare system interoperabilityExtract structured clinical informationFine-tune Large Language Models

Balancing Natural Language Processing Accuracy and Normalisation in Extracting Medical Insights

Nov 19, 2025
PT
Paulina Tworek
🏛️ Sano - Centre for Computational Personalized Medicine | Jagiellonian University | AGH University of Krakow | Voivodeship Rehabilitation Hospital for Children in Ameryka | University of Warmia and Mazury

This study addresses the trade-off among accuracy, standardization, and computational efficiency in NLP for low-resource, non-English medical settings, using unstructured electronic health records from a Polish pediatric rehabilitation hospital. We propose a hybrid approach integrating rule-based systems—offering high precision and low computational cost—with multilingual large language models (LLMs)—providing strong generalization and adaptability. We systematically compare performance on demographic, clinical finding, and medication information extraction tasks using both original Polish text and machine-translated English text. Results show rule-based methods outperform LLMs in age and gender identification, while LLMs significantly improve drug name recognition accuracy. Critically, machine translation introduces non-negligible information loss, degrading downstream performance. This work establishes a new paradigm for resource-constrained, multilingual clinical NLP that balances accuracy, robustness, and practical deployability.

Assessing trade-offs between accuracy, normalization, and computational costsComparing rule-based methods and LLMs for EHR information extractionExtracting structured medical insights from unstructured clinical text

Natural Language Processing for Electronic Health Records in Scandinavian Languages: Norwegian, Swedish, and Danish

Mar 24, 2025
AZ
Ashenafi Zebene Woldaregay
🏛️ University Hospital of Northern Norway | UiT The Arctic University of Norway | Norwegian Centre for E-health Research | Copenhagen University Hospital | Stockholm University | University Hospital of North Norway

Clinical natural language processing (NLP) for Scandinavian languages remains underexplored, with no systematic, cross-lingual assessment of progress, resource availability, or methodological trends. Method: We conducted a systematic review of 113 peer-reviewed studies (2010–2024) from PubMed, ACL Anthology, IEEE Xplore, Scopus, and Web of Science, focusing on Norwegian, Swedish, and Danish clinical text processing. We quantitatively analyzed model adoption, task coverage, and resource sharing across the three languages. Results: Swedish dominates the field (72% of studies), while Norwegian (18%) and Danish (10%) lag significantly—especially in critical tasks like de-identification and in adopting Transformer-based models. Data, code, and pretrained model sharing rates are extremely low, hindering regional reproducibility and collaboration. We further evaluated rule-based systems, classical machine learning, and BERT-family models on EHR text, identifying persistent adaptation bottlenecks and limited cross-lingual transferability. This study provides the first empirical evidence of structural imbalance in Scandinavian clinical NLP and offers actionable insights for equitable, multilingual health AI resource development.

Assessing NLP methods for Scandinavian clinical textsEvaluating resource sharing in clinical NLP researchIdentifying disparities in NLP adoption across languages

Excessive clinical documentation burden impedes healthcare efficiency. Method: We propose a large language model (LLM)-based automation framework comprising (1) structured table generation from nurse verbal notes and (2) precise medical instruction extraction from physician–patient consultation transcripts. To address data scarcity and privacy constraints, we design an intelligent agent pipeline that synthesizes high-fidelity, de-identified, non-sensitive spoken clinical data. We release SYNUR—the first open-source dataset for nurse observation summarization—and SIMORD—the first dedicated benchmark for medical instruction extraction. We systematically evaluate both open-weight (e.g., Qwen, Llama) and proprietary (e.g., GPT-4o, o1) LLMs on real-world clinical tasks. Contribution/Results: Experiments demonstrate LLMs’ effectiveness on these high-value clinical NLP tasks, establishing a reproducible, scalable pathway toward structured electronic health record generation. Our work fills critical gaps in high-quality annotated clinical datasets and open evaluation benchmarks.

Extracting medical orders from doctor-patient consultationsGenerating synthetic clinical data for NLP researchStructuring nurse dictations into tabular reports

Will Large Language Models Transform Clinical Prediction?

May 23, 2025
YY
Yusuf Yildiz
🏛️ University of Manchester

Large language models (LLMs) face critical limitations in clinical prediction—including handling heterogeneous data, ensuring interpretable risk stratification, and integrating with real-world clinical decision workflows—hindering their regulatory and practical adoption. Method: We propose the first LLM-oriented validation framework tailored to clinical prediction, integrating fairness quantification, survival analysis–aware modeling, structured alignment of multi-source clinical text, and an ethics-technology co-assessment pathway compliant with healthcare regulations. Our approach synergizes clinical natural language processing, interpretable statistical learning, and medical ethics analysis. Contribution/Results: The study identifies key translational gaps between LLM capabilities and clinical operational requirements. It delivers an actionable methodology guide and a prioritized development roadmap, establishing both theoretical foundations and implementation paradigms for evidence-based, trustworthy, and regulation-compliant LLM-driven clinical prediction tools.

Address fairness and bias in LLM-based clinical predictionsDevelop regulations for LLM integration in healthcare workflowsEvaluate LLMs for clinical prediction accuracy and reliability

Latest Papers

What's happening recently
View more

This study addresses the challenges of accurately querying structured data and extracting information from unstructured clinical text in electronic health records (EHRs). To this end, the authors propose a unified framework that integrates large language models (LLMs) with retrieval-augmented generation (RAG): LLMs are employed to execute structured queries (e.g., Pandas operations), while RAG enhances information extraction from unstructured clinical narratives. The work introduces an innovative automatic evaluation pipeline based on synthetically generated question-answer pairs, combining exact match metrics, semantic similarity scores, and human assessments. Evaluated on a subset of MIMIC-III, the approach demonstrates improved semantic accuracy and task adaptability, offering clinical data science a flexible and reliable tool for automated reasoning and evaluation.

Clinical Data ScienceElectronic Health RecordsInformation Extraction

Clinical communication data contain critical medical information, yet their unstructured nature and the scarcity of authentically annotated corpora hinder the training of conventional NLP models. This work proposes the first systematic framework for generating synthetic clinical communication data using large language models, encompassing 13 underexplored scenarios—such as emergency dispatch, nurse handoffs, and patient triage—that lack real-world annotated examples. The approach produces diverse synthetic texts and incorporates fine-tuned encoder models, further enhanced by deliberate degradation of communication fidelity to improve robustness. Experimental results demonstrate that the proposed method significantly outperforms zero-shot baselines across multiple downstream tasks, thereby validating the efficacy and potential of synthetic data in advancing practical clinical NLP systems.

annotated corporaclinical communicationhealthcare NLP

Leveraging LLMs for Structured Data Extraction from Unstructured Patient Records

Dec 03, 2025
MA
Mitchell A. Klusty
🏛️ University of Kentucky | University of Kentucky College of Medicine

In clinical research, manual extraction of structured clinical features from unstructured electronic health records (EHRs) is time-consuming, inefficient, and error-prone. To address this, we propose a privacy-preserving, modular, on-premises large language model (LLM) framework that integrates retrieval-augmented generation (RAG) with structured-output prompt engineering, enabling secure, scalable, containerized deployment in HIPAA-compliant environments. Our approach uniquely synergizes RAG with deterministic structured-response mechanisms for clinical text parsing—balancing domain adaptability and strict data privacy. Evaluated across multiple medical feature extraction tasks, the framework achieves high accuracy, substantially reducing manual annotation effort and improving data consistency. Notably, its systematic evaluation uncovered previously undetected systematic errors in prior human annotations, thereby validating both its reliability and its capacity for quality assurance and error discovery.

Automates structured data extraction from unstructured patient recordsEnhances data consistency and accuracy using LLMsReduces manual chart review burden in clinical research

Integrating Text and Time-Series into (Large) Language Models to Predict Medical Outcomes

Sep 17, 2025
IB
Iyadh Ben Cheikh Larbi
🏛️ German Research Center for Artificial Intelligence (DFKI)

Existing research inadequately explores the capability of large language models (LLMs) to jointly process clinical text and time-series data for predictive tasks. This paper proposes a lightweight, prompt-driven multimodal modeling approach: leveraging the DSPy framework to construct an instruction-tuned prompt optimization pipeline, enabling off-the-shelf LLMs—without architectural modification—to jointly reason over unstructured clinical narratives and structured temporal data (e.g., vital signs, lab results). The method achieves performance on par with specialized multimodal models across multiple clinical outcome classification tasks, while substantially reducing system complexity and improving cross-task generalization. Its core innovation lies in “injecting” temporal modeling capacity into the LLM’s prompt layer—enabling unified representation and reasoning over both textual and sequential modalities. This offers an efficient, scalable, and general-purpose solution for clinical AI.

Adapting LLMs to handle clinical classification with structured EHRCombining text and time-series data for medical outcome predictionOptimizing multimodal processing for clinical tasks with reduced complexity

This work addresses the inflated performance of clinical NLP models caused by temporal and lexical leakage, which poses serious risks to real-world deployment safety. To mitigate this, the authors propose a lightweight auditing framework that integrates interpretability mechanisms early into the model development pipeline, systematically ensuring temporal validity, probability calibration, and behavioral robustness. By jointly leveraging temporal leakage detection and interpretability analysis, the framework effectively curbs the model’s reliance on spurious cues—such as discharge-related vocabulary—that do not reflect genuine clinical signals. Experimental results demonstrate that audited models produce more conservative and well-calibrated prediction probabilities, significantly enhancing clinical reliability and safety without compromising overall performance.

clinical NLPlexical leakagemodel deployment

Hot Scholars

CY

Carl Yang

Waymo LLC, PhD at University of California, Davis
GPU ComputingParallel ComputingGraph Processing
JS

Jimeng Sun

Professor at University of Illinois Urbana-Champaign
AI for healthcareMachine learning for healthcaredeep learning for healthcare
EC

Edward Choi

KAIST
Machine LearningArtificial IntelligenceHealthcare
YW

Yonghui Wu

Associate Professor, University of Florida
Natural Language ProcessingMachine LearningMedical InformaticsPharmacovigilance