Score
Designs and implements natural language processing models and adaptation pipelines that process and normalize clinical text (e.g., notes and reports) to parse and extract structured information such as diagnoses, temporal events, and treatment-related entities. Builds preprocessing and annotation-correction workflows, adapts models to new clinical subdomains or annotation schemes, and maps clinical language to standardized representations or translations for use in downstream predictive models.
This paper systematically reviews the current state, challenges, and deployment pathways of large language models (LLMs) in healthcare. It addresses four core barriers to clinical adoption: poor deployability, high privacy risks, weak task-specific adaptation, and absence of standardized evaluation frameworks. To tackle these, the study proposes: (1) a hierarchical ethical framework tailored to healthcare, integrating data security, algorithmic fairness, and clinical accountability; (2) a structured, multidimensional taxonomy of clinical LLM tasks—encompassing text generation, information extraction, multimodal understanding, and conversational interaction; and (3) an integrated technical pathway combining localized inference, in-context learning, and multimodal modeling. The resulting comprehensive guide bridges theoretical rigor and practical implementation, offering a methodological foundation and actionable roadmap for developing trustworthy, evaluable, and deployable clinical AI systems.
This study addresses interoperability challenges in European healthcare systems by proposing a large language model (LLM)-based approach for multilingual clinical text information extraction. To process unstructured clinical reports, we design an end-to-end extraction pipeline with an interactive interface and jointly apply prompt engineering and supervised fine-tuning to adapt medium- and small-scale LLMs. Our key contributions are threefold: (1) We release the first bilingual clinical summarization dataset—comprising 60,000 English and 24,000 German samples—validated via multi-dimensional automatic and human annotation; (2) We propose a multi-granularity evaluation framework integrating ROUGE, BERTScore, and entity-level metrics; (3) Our fine-tuned model achieves 89.3% F1 on critical entity extraction—outperforming zero-shot LLMs by 22.7%—while accelerating inference by 3.8× and reducing GPU memory consumption by 64%, significantly enhancing clinical deployability.
This study addresses the trade-off among accuracy, standardization, and computational efficiency in NLP for low-resource, non-English medical settings, using unstructured electronic health records from a Polish pediatric rehabilitation hospital. We propose a hybrid approach integrating rule-based systems—offering high precision and low computational cost—with multilingual large language models (LLMs)—providing strong generalization and adaptability. We systematically compare performance on demographic, clinical finding, and medication information extraction tasks using both original Polish text and machine-translated English text. Results show rule-based methods outperform LLMs in age and gender identification, while LLMs significantly improve drug name recognition accuracy. Critically, machine translation introduces non-negligible information loss, degrading downstream performance. This work establishes a new paradigm for resource-constrained, multilingual clinical NLP that balances accuracy, robustness, and practical deployability.
Clinical natural language processing (NLP) for Scandinavian languages remains underexplored, with no systematic, cross-lingual assessment of progress, resource availability, or methodological trends. Method: We conducted a systematic review of 113 peer-reviewed studies (2010–2024) from PubMed, ACL Anthology, IEEE Xplore, Scopus, and Web of Science, focusing on Norwegian, Swedish, and Danish clinical text processing. We quantitatively analyzed model adoption, task coverage, and resource sharing across the three languages. Results: Swedish dominates the field (72% of studies), while Norwegian (18%) and Danish (10%) lag significantly—especially in critical tasks like de-identification and in adopting Transformer-based models. Data, code, and pretrained model sharing rates are extremely low, hindering regional reproducibility and collaboration. We further evaluated rule-based systems, classical machine learning, and BERT-family models on EHR text, identifying persistent adaptation bottlenecks and limited cross-lingual transferability. This study provides the first empirical evidence of structural imbalance in Scandinavian clinical NLP and offers actionable insights for equitable, multilingual health AI resource development.
Excessive clinical documentation burden impedes healthcare efficiency. Method: We propose a large language model (LLM)-based automation framework comprising (1) structured table generation from nurse verbal notes and (2) precise medical instruction extraction from physician–patient consultation transcripts. To address data scarcity and privacy constraints, we design an intelligent agent pipeline that synthesizes high-fidelity, de-identified, non-sensitive spoken clinical data. We release SYNUR—the first open-source dataset for nurse observation summarization—and SIMORD—the first dedicated benchmark for medical instruction extraction. We systematically evaluate both open-weight (e.g., Qwen, Llama) and proprietary (e.g., GPT-4o, o1) LLMs on real-world clinical tasks. Contribution/Results: Experiments demonstrate LLMs’ effectiveness on these high-value clinical NLP tasks, establishing a reproducible, scalable pathway toward structured electronic health record generation. Our work fills critical gaps in high-quality annotated clinical datasets and open evaluation benchmarks.
Large language models (LLMs) face critical limitations in clinical prediction—including handling heterogeneous data, ensuring interpretable risk stratification, and integrating with real-world clinical decision workflows—hindering their regulatory and practical adoption. Method: We propose the first LLM-oriented validation framework tailored to clinical prediction, integrating fairness quantification, survival analysis–aware modeling, structured alignment of multi-source clinical text, and an ethics-technology co-assessment pathway compliant with healthcare regulations. Our approach synergizes clinical natural language processing, interpretable statistical learning, and medical ethics analysis. Contribution/Results: The study identifies key translational gaps between LLM capabilities and clinical operational requirements. It delivers an actionable methodology guide and a prioritized development roadmap, establishing both theoretical foundations and implementation paradigms for evidence-based, trustworthy, and regulation-compliant LLM-driven clinical prediction tools.
This study addresses the challenges of accurately querying structured data and extracting information from unstructured clinical text in electronic health records (EHRs). To this end, the authors propose a unified framework that integrates large language models (LLMs) with retrieval-augmented generation (RAG): LLMs are employed to execute structured queries (e.g., Pandas operations), while RAG enhances information extraction from unstructured clinical narratives. The work introduces an innovative automatic evaluation pipeline based on synthetically generated question-answer pairs, combining exact match metrics, semantic similarity scores, and human assessments. Evaluated on a subset of MIMIC-III, the approach demonstrates improved semantic accuracy and task adaptability, offering clinical data science a flexible and reliable tool for automated reasoning and evaluation.
Clinical communication data contain critical medical information, yet their unstructured nature and the scarcity of authentically annotated corpora hinder the training of conventional NLP models. This work proposes the first systematic framework for generating synthetic clinical communication data using large language models, encompassing 13 underexplored scenarios—such as emergency dispatch, nurse handoffs, and patient triage—that lack real-world annotated examples. The approach produces diverse synthetic texts and incorporates fine-tuned encoder models, further enhanced by deliberate degradation of communication fidelity to improve robustness. Experimental results demonstrate that the proposed method significantly outperforms zero-shot baselines across multiple downstream tasks, thereby validating the efficacy and potential of synthetic data in advancing practical clinical NLP systems.
In clinical research, manual extraction of structured clinical features from unstructured electronic health records (EHRs) is time-consuming, inefficient, and error-prone. To address this, we propose a privacy-preserving, modular, on-premises large language model (LLM) framework that integrates retrieval-augmented generation (RAG) with structured-output prompt engineering, enabling secure, scalable, containerized deployment in HIPAA-compliant environments. Our approach uniquely synergizes RAG with deterministic structured-response mechanisms for clinical text parsing—balancing domain adaptability and strict data privacy. Evaluated across multiple medical feature extraction tasks, the framework achieves high accuracy, substantially reducing manual annotation effort and improving data consistency. Notably, its systematic evaluation uncovered previously undetected systematic errors in prior human annotations, thereby validating both its reliability and its capacity for quality assurance and error discovery.
Existing research inadequately explores the capability of large language models (LLMs) to jointly process clinical text and time-series data for predictive tasks. This paper proposes a lightweight, prompt-driven multimodal modeling approach: leveraging the DSPy framework to construct an instruction-tuned prompt optimization pipeline, enabling off-the-shelf LLMs—without architectural modification—to jointly reason over unstructured clinical narratives and structured temporal data (e.g., vital signs, lab results). The method achieves performance on par with specialized multimodal models across multiple clinical outcome classification tasks, while substantially reducing system complexity and improving cross-task generalization. Its core innovation lies in “injecting” temporal modeling capacity into the LLM’s prompt layer—enabling unified representation and reasoning over both textual and sequential modalities. This offers an efficient, scalable, and general-purpose solution for clinical AI.
This work addresses the inflated performance of clinical NLP models caused by temporal and lexical leakage, which poses serious risks to real-world deployment safety. To mitigate this, the authors propose a lightweight auditing framework that integrates interpretability mechanisms early into the model development pipeline, systematically ensuring temporal validity, probability calibration, and behavioral robustness. By jointly leveraging temporal leakage detection and interpretability analysis, the framework effectively curbs the model’s reliance on spurious cues—such as discharge-related vocabulary—that do not reflect genuine clinical signals. Experimental results demonstrate that audited models produce more conservative and well-calibrated prediction probabilities, significantly enhancing clinical reliability and safety without compromising overall performance.