Progress in Medical AI: Reviewing Large Language Models and Multimodal Systems for Diagonosis

📅 2025-02-10
🏛️ AI Med
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses three critical limitations of medical large language models (MedLMs) in clinical diagnosis: poor interpretability, inadequate multimodal fusion, and difficulty integrating domain-specific knowledge. To tackle these challenges, we propose a unified diagnostic AI evolution framework comprising three synergistic pillars: domain adaptation, structured knowledge injection, and cross-modal alignment. For the first time, we systematically integrate four paradigmatic approaches—BioBERT, Med-PaLM, DR.KNOWS, and multimodal foundation models (MMFMs)—by unifying biomedical pretraining, instruction tuning, medical knowledge graph embedding, and vision-language joint modeling. Evaluated on benchmarks including MultiMedQA, our approach achieves significant improvements in clinical question answering accuracy, named entity recognition F1-score, and lesion detection performance, while simultaneously reducing harmful outputs. Moreover, it enhances model interpretability and clinical utility, delivering a deployable technical framework for personalized medicine.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Knowledge Representation and Reasoning: Diagnosis and Abductive ReasoningComputer Vision: Multi-modal Vision

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAG
📝 Abstract
The rapid advancement of artificial intelligence (AI) in healthcare has significantly enhanced diagnostic accuracy and clinical decision-making processes. This review examines four pivotal studies that highlight the integration of large language models (LLMs) and multimodal systems in medical diagnostics. BioBERT demonstrates the efficacy of domain-specific pretraining on biomedical texts, improving performance in tasks such as named entity recognition, relation extraction, and question answering. Med-PaLM, a large-scale language model tailored for clinical question answering, leverages instruction prompt tuning to enhance accuracy and reduce harmful outputs, validated through the MultiMedQA benchmark. DR.KNOWS integrates medical knowledge graphs with LLMs, enhancing diagnostic reasoning and interpretability by grounding model predictions in structured medical knowledge. Medical Multimodal Foundation Models (MMFMs) combine textual and imaging data to improve tasks like segmentation, lesion detection, and automated report generation. These studies demonstrate the importance of domain adaptation, structured knowledge integration, and multimodal data fusion in developing robust and interpretable AI-driven diagnostic tools.
Problem

Research questions and friction points this paper is trying to address.

Enhancing disease prediction and diagnosis
Improving personalized treatment and drug discovery
Advancing medical image analysis and anatomical modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models enhance diagnostics
Graph neural networks boost drug discovery
Vision-Language Models improve medical imaging
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ran Tong
T
Ting Xu
X
Xinxin Ju
L
Lanruo Wang