Score
Designs and implements systems that ingest multiple input modalities (e.g., images, text, signals) and produce structured diagnostic outputs such as syndrome labels, differential diagnoses, suggested interventions or prescriptions, and explanatory rationale. Builds and fine-tunes multimodal models or pipelines (for example multimodal LLMs), maps visual and other modality-specific features into text or structured fields, and analyzes the fidelity, format compliance, and interpretability of the generated diagnoses.
This study systematically reviews medical multimodal large language models (MLLMs), focusing on three core clinical applications: medical report generation, diagnostic support, and therapeutic assistance. Drawing on 330 state-of-the-art publications, it unifies technical methodologies, multimodal data compositions, and evaluation benchmarks for the first time. A taxonomy classifying six principal medical modalities—including radiological imaging, histopathological slides, and electronic health records—is established, along with corresponding evaluation frameworks. Key challenges—data privacy, hallucination, and lack of interpretability—are distilled, and actionable solutions are proposed, including medical vision–language alignment, domain-adaptive fine-tuning, and structured medical knowledge injection. The work culminates in a structured knowledge graph demonstrating that MLLMs have achieved clinically deployable performance in tasks such as radiology report generation and pathology-assisted diagnosis.
Despite rapid advances in multimodal large language models (MLLMs), their clinical deployment remains hindered by domain-specific bottlenecks—including scarce annotated medical data, modality bias, and limited interpretability. Method: This work systematically reviews the evolution from large language models (LLMs) to MLLMs and empirically analyzes their integration of text, medical imaging, and audio modalities for clinical decision support, radiology/pathology interpretation, patient interaction, and biomedical research. Contribution/Results: We identify three critical research directions: (1) construction of medical-domain multimodal datasets, (2) novel modality alignment techniques, and (3) an ethics-aware governance framework. Empirical evaluation demonstrates that MLLMs improve diagnostic assistance accuracy and accelerate structured reporting generation; however, performance is constrained by data scarcity, cross-modal misalignment, and opaque reasoning. This study provides both theoretical foundations and actionable pathways toward trustworthy, clinically viable MLLMs.
Medical multimodal AI is hindered by the scarcity of high-quality heterogeneous data, particularly in dermatology, where image datasets lack rich clinical textual annotations—limiting model robustness and generalization. To address this, we propose a fine-tuning-free prompt engineering framework that leverages structured medical metadata (e.g., lesion location, age, sex) to guide large language models in generating high-fidelity, low-hallucination synthetic clinical notes. This approach is the first to enable image-to-text cross-modal retrieval solely through prompt design. Evaluated across multiple dermatological benchmarks, the synthetic notes significantly improve multimodal classification accuracy (+3.2–7.8%), with even greater gains under domain shift. Our core innovation lies in embedding clinical priors directly into the prompting mechanism—ensuring both clinical plausibility and modeling efficacy—thereby bridging the modality gap without architectural modification or parameter updates.
This work addresses three critical limitations of medical large language models (MedLMs) in clinical diagnosis: poor interpretability, inadequate multimodal fusion, and difficulty integrating domain-specific knowledge. To tackle these challenges, we propose a unified diagnostic AI evolution framework comprising three synergistic pillars: domain adaptation, structured knowledge injection, and cross-modal alignment. For the first time, we systematically integrate four paradigmatic approaches—BioBERT, Med-PaLM, DR.KNOWS, and multimodal foundation models (MMFMs)—by unifying biomedical pretraining, instruction tuning, medical knowledge graph embedding, and vision-language joint modeling. Evaluated on benchmarks including MultiMedQA, our approach achieves significant improvements in clinical question answering accuracy, named entity recognition F1-score, and lesion detection performance, while simultaneously reducing harmful outputs. Moreover, it enhances model interpretability and clinical utility, delivering a deployable technical framework for personalized medicine.
Current conversational diagnostic systems predominantly rely on text-only interaction, failing to support real-time multimodal clinical analysis—such as medical images, ECGs, and PDF reports—essential for telemedicine. To address this, we propose a state-driven multimodal conversational diagnostic framework built upon Gemini 2.0 Flash, featuring a dynamic state-aware mechanism that jointly enables multimodal understanding, uncertainty modeling, and structured clinical questioning. Crucially, we introduce the first method for autonomously generating follow-up questions based on patient-state uncertainty, emulating expert clinicians’ diagnostic reasoning. Evaluated on 105 OSCE cases, our system significantly outperformed general practitioners across 7 of 9 multimodal and 29 of 32 non-multimodal clinical dimensions—including diagnostic accuracy—demonstrating synergistic enhancement between multimodal capability and diagnostic efficacy.
This study investigates the capability of large language models (LLMs) for zero-shot multimodal (text + audio) diagnosis of depression and PTSD. Leveraging the E-DAIC dataset, we systematically evaluate single-modal and cross-modal fusion performance of models including Gemini 1.5 Pro and GPT-4o mini. Methodologically, we propose two novel metrics—Modal Superiority Score and Disagreement Resolution Score—to quantify multimodal synergy, and employ zero-shot prompt engineering—without fine-tuning—to enable end-to-end cross-modal reasoning. Results demonstrate that Gemini 1.5 Pro achieves an F1-score of 0.67 and balanced accuracy of 77.4% under multimodal fusion, outperforming the best single-modal baseline by 2.7–3.1%. This work constitutes the first empirical validation that LLMs can effectively and robustly perform early mental health screening across unseen diagnostic categories via zero-shot cross-modal integration.
Existing clinical diagnostic evaluation methods struggle to capture the true complexity of progressive multimodal information disclosure and dynamic reasoning. To address this gap, this work proposes ClinMM-Bench—the first large-scale, multi-turn, multimodal diagnostic benchmark grounded in real-world clinical scenarios—encompassing 1,089 challenging cases and 3,760 medical images, along with a dual-layer evaluation framework that systematically assesses both diagnostic accuracy and reasoning quality. Evaluation of 15 state-of-the-art multimodal large language models reveals that, although closed-source models achieve the best performance, their fully correct diagnosis rate remains limited; while models can generate plausible hypotheses, they exhibit significant deficiencies in constructing reliable and coherent reasoning chains. The study further identifies five representative failure patterns, offering concrete directions for future model improvement.
This study addresses the challenge that multimodal models struggle to reliably bind textual prompts (e.g., “image”) to their corresponding input modalities, resulting in inadequate source modality tracking. The work formally defines and empirically investigates the “source modality monitoring” problem for the first time, introducing an evaluation paradigm grounded in target-modality information retrieval. By integrating syntactic manipulation with semantic perturbation, the authors systematically assess binding mechanisms across eleven prominent vision-language models. Their findings reveal that semantic cues dominate the binding process when modality distributions exhibit significant divergence, consistently outweighing syntactic signals. These insights offer critical evidence and a novel perspective for enhancing the reliability and robustness of multimodal agents.
This work addresses the limitations of handcrafted input representations in supervised learning on complex heterogeneous data—such as time series, text, and structured records—by proposing a large language model (LLM)-based agent pipeline. The approach automatically derives global normalization rules from a small set of diverse textual serialization examples and combines them with task-conditioned local rules to construct efficient, auditable, and low-cost tabular input representations. Evaluated on 15 clinical tasks in the EHRSHOT benchmark, the method significantly outperforms conventional count-based feature models, naive LLM text serialization baselines, and large-scale pretrained clinical foundation models, demonstrating strong effectiveness and generalization capability in few-shot medical settings.
This work addresses the challenge of leveraging unstructured clinical notes in electronic health records (EHR) during training to enhance the performance of models that rely solely on structured data at deployment. The authors propose a multimodal learning framework that integrates clinical notes—encoded via BioClinicalBERT—with structured features such as demographics and medical codes during training. By combining a teacher–student architecture, contrastive learning, and contrastive knowledge distillation, the approach enables the final deployed model to achieve high inference efficiency using only structured inputs, without requiring access to textual data at test time. To the best of the authors’ knowledge, this is the first method to effectively augment structured EHR representations with unstructured clinical notes while maintaining deployment constraints. Evaluated on a cohort of 3,466 children with late language emergence, the model achieves an AUROC of 0.705, significantly outperforming baseline methods (AUROC = 0.656).