Score
Designs, implements, or analyzes systems that encode clinical text into representations usable by language models and fuse language-model outputs with other patient data modalities to produce interpretable, clinically actionable predictions or decision support. Work includes developing clinical-text encoding pipelines, multimodal fusion architectures and integration, calibration and explainability of LM-based components, and evaluation of outputs at region- or patient-level.
This paper systematically reviews the current state, challenges, and deployment pathways of large language models (LLMs) in healthcare. It addresses four core barriers to clinical adoption: poor deployability, high privacy risks, weak task-specific adaptation, and absence of standardized evaluation frameworks. To tackle these, the study proposes: (1) a hierarchical ethical framework tailored to healthcare, integrating data security, algorithmic fairness, and clinical accountability; (2) a structured, multidimensional taxonomy of clinical LLM tasks—encompassing text generation, information extraction, multimodal understanding, and conversational interaction; and (3) an integrated technical pathway combining localized inference, in-context learning, and multimodal modeling. The resulting comprehensive guide bridges theoretical rigor and practical implementation, offering a methodological foundation and actionable roadmap for developing trustworthy, evaluable, and deployable clinical AI systems.
This work addresses the limitations of conventional clinical prediction systems, which rely on task-specific multimodal fusion architectures that suffer from poor generalizability and high complexity. The authors propose a unified text serialization paradigm that converts both structured and unstructured clinical data into natural language sequences, enabling end-to-end prediction by directly fine-tuning large language models—including ModernBERT, Llama 3.1, Gemma, DeepSeek-R1-Qwen, and Qwen3—without requiring specialized fusion modules. This approach substantially reduces system complexity while achieving performance on par with or superior to existing task-specific multimodal baselines across three clinical prediction tasks. Notably, it outperforms the gradient-boosting model currently used in clinical practice for graft failure prediction, demonstrating strong cross-task and cross-scenario generalization capabilities.
Despite rapid advances in multimodal large language models (MLLMs), their clinical deployment remains hindered by domain-specific bottlenecks—including scarce annotated medical data, modality bias, and limited interpretability. Method: This work systematically reviews the evolution from large language models (LLMs) to MLLMs and empirically analyzes their integration of text, medical imaging, and audio modalities for clinical decision support, radiology/pathology interpretation, patient interaction, and biomedical research. Contribution/Results: We identify three critical research directions: (1) construction of medical-domain multimodal datasets, (2) novel modality alignment techniques, and (3) an ethics-aware governance framework. Empirical evaluation demonstrates that MLLMs improve diagnostic assistance accuracy and accelerate structured reporting generation; however, performance is constrained by data scarcity, cross-modal misalignment, and opaque reasoning. This study provides both theoretical foundations and actionable pathways toward trustworthy, clinically viable MLLMs.
This study addresses critical challenges hindering the clinical deployment of large language models (LLMs): patient privacy, algorithmic bias, regulatory compliance, and operational sustainability. Methodologically, we propose the first healthcare-specific, four-dimensional adaptation framework comprising: (1) domain-adaptive fine-tuning, (2) clinically informed prompt engineering, (3) multimodal electronic health record (EHR) integration—unifying unstructured text and structured data—and (4) a novel evaluation paradigm centered on clinical accuracy, fairness, robustness, and outcome-oriented metrics. Crucially, privacy-preserving mechanisms, bias mitigation strategies, and regulatory requirements (e.g., HIPAA, FDA guidelines) are systematically embedded throughout the technical design lifecycle. The work yields a reproducible implementation roadmap with clearly defined interdisciplinary collaboration protocols. It provides both theoretical foundations and actionable guidance for the safe, effective, and compliant integration of LLMs into clinical decision support, patient-facing applications, and healthcare administrative automation.
Large language models (LLMs) face critical limitations in clinical prediction—including handling heterogeneous data, ensuring interpretable risk stratification, and integrating with real-world clinical decision workflows—hindering their regulatory and practical adoption. Method: We propose the first LLM-oriented validation framework tailored to clinical prediction, integrating fairness quantification, survival analysis–aware modeling, structured alignment of multi-source clinical text, and an ethics-technology co-assessment pathway compliant with healthcare regulations. Our approach synergizes clinical natural language processing, interpretable statistical learning, and medical ethics analysis. Contribution/Results: The study identifies key translational gaps between LLM capabilities and clinical operational requirements. It delivers an actionable methodology guide and a prioritized development roadmap, establishing both theoretical foundations and implementation paradigms for evidence-based, trustworthy, and regulation-compliant LLM-driven clinical prediction tools.
Clinical NLP faces challenges including scarcity of real-world data, stringent privacy requirements, and the complexity of medical terminology. Method: We propose ClinGen, a resource-efficient knowledge-injection framework for synthetic clinical text generation. ClinGen introduces a novel prompting mechanism that synergistically integrates external biomedical knowledge graphs with large language models (LLMs) to jointly model clinical topics and writing styles, ensuring privacy preservation and regulatory compliance while enhancing fidelity and lexical/semantic diversity. The approach comprises knowledge extraction, knowledge graph embedding, and context-aware prompt engineering, complemented by a multi-task evaluation framework. Results: Extensive experiments across seven clinical NLP tasks and sixteen benchmark datasets demonstrate significant performance gains. Generated data exhibit improved distributional alignment with real clinical corpora, and training sample diversity increases by an average of 42%.
Excessive clinical documentation burden impedes healthcare efficiency. Method: We propose a large language model (LLM)-based automation framework comprising (1) structured table generation from nurse verbal notes and (2) precise medical instruction extraction from physician–patient consultation transcripts. To address data scarcity and privacy constraints, we design an intelligent agent pipeline that synthesizes high-fidelity, de-identified, non-sensitive spoken clinical data. We release SYNUR—the first open-source dataset for nurse observation summarization—and SIMORD—the first dedicated benchmark for medical instruction extraction. We systematically evaluate both open-weight (e.g., Qwen, Llama) and proprietary (e.g., GPT-4o, o1) LLMs on real-world clinical tasks. Contribution/Results: Experiments demonstrate LLMs’ effectiveness on these high-value clinical NLP tasks, establishing a reproducible, scalable pathway toward structured electronic health record generation. Our work fills critical gaps in high-quality annotated clinical datasets and open evaluation benchmarks.
Clinical communication data contain critical medical information, yet their unstructured nature and the scarcity of authentically annotated corpora hinder the training of conventional NLP models. This work proposes the first systematic framework for generating synthetic clinical communication data using large language models, encompassing 13 underexplored scenarios—such as emergency dispatch, nurse handoffs, and patient triage—that lack real-world annotated examples. The approach produces diverse synthetic texts and incorporates fine-tuned encoder models, further enhanced by deliberate degradation of communication fidelity to improve robustness. Experimental results demonstrate that the proposed method significantly outperforms zero-shot baselines across multiple downstream tasks, thereby validating the efficacy and potential of synthetic data in advancing practical clinical NLP systems.
This study addresses the challenge of accurately predicting short-term mortality risk in heart failure patients when relying solely on structured electronic health records. Leveraging a French heart failure cohort, the authors systematically evaluate various Transformer-based modeling strategies and propose a supervised multimodal fusion approach that effectively integrates entity-level representations from clinical text with structured variables. This method significantly outperforms conventional CLS embeddings and existing large language model (LLM) prompting techniques. Experimental results demonstrate that the proposed entity-aware multimodal Transformer achieves the best performance for short-term mortality prediction. In contrast, LLMs exhibit inconsistent performance across different input modalities and decoding strategies, with plain-text prompting yielding better results than structured or multimodal inputs.
Medical multimodal AI is hindered by the scarcity of high-quality heterogeneous data, particularly in dermatology, where image datasets lack rich clinical textual annotations—limiting model robustness and generalization. To address this, we propose a fine-tuning-free prompt engineering framework that leverages structured medical metadata (e.g., lesion location, age, sex) to guide large language models in generating high-fidelity, low-hallucination synthetic clinical notes. This approach is the first to enable image-to-text cross-modal retrieval solely through prompt design. Evaluated across multiple dermatological benchmarks, the synthetic notes significantly improve multimodal classification accuracy (+3.2–7.8%), with even greater gains under domain shift. Our core innovation lies in embedding clinical priors directly into the prompting mechanism—ensuring both clinical plausibility and modeling efficacy—thereby bridging the modality gap without architectural modification or parameter updates.
This work addresses the inflated performance of clinical NLP models caused by temporal and lexical leakage, which poses serious risks to real-world deployment safety. To mitigate this, the authors propose a lightweight auditing framework that integrates interpretability mechanisms early into the model development pipeline, systematically ensuring temporal validity, probability calibration, and behavioral robustness. By jointly leveraging temporal leakage detection and interpretability analysis, the framework effectively curbs the model’s reliance on spurious cues—such as discharge-related vocabulary—that do not reflect genuine clinical signals. Experimental results demonstrate that audited models produce more conservative and well-calibrated prediction probabilities, significantly enhancing clinical reliability and safety without compromising overall performance.