Multimodal Large Language Models for Medicine: A Comprehensive Survey

📅 2025-04-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study systematically reviews medical multimodal large language models (MLLMs), focusing on three core clinical applications: medical report generation, diagnostic support, and therapeutic assistance. Drawing on 330 state-of-the-art publications, it unifies technical methodologies, multimodal data compositions, and evaluation benchmarks for the first time. A taxonomy classifying six principal medical modalities—including radiological imaging, histopathological slides, and electronic health records—is established, along with corresponding evaluation frameworks. Key challenges—data privacy, hallucination, and lack of interpretability—are distilled, and actionable solutions are proposed, including medical vision–language alignment, domain-adaptive fine-tuning, and structured medical knowledge injection. The work culminates in a structured knowledge graph demonstrating that MLLMs have achieved clinically deployable performance in tasks such as radiology report generation and pathology-assisted diagnosis.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Computer Vision: Multi-modal VisionNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
MLLMs have recently become a focal point in the field of artificial intelligence research. Building on the strong capabilities of LLMs, MLLMs are adept at addressing complex multi-modal tasks. With the release of GPT-4, MLLMs have gained substantial attention from different domains. Researchers have begun to explore the potential of MLLMs in the medical and healthcare domain. In this paper, we first introduce the background and fundamental concepts related to LLMs and MLLMs, while emphasizing the working principles of MLLMs. Subsequently, we summarize three main directions of application within healthcare: medical reporting, medical diagnosis, and medical treatment. Our findings are based on a comprehensive review of 330 recent papers in this area. We illustrate the remarkable capabilities of MLLMs in these domains by providing specific examples. For data, we present six mainstream modes of data along with their corresponding evaluation benchmarks. At the end of the survey, we discuss the challenges faced by MLLMs in the medical and healthcare domain and propose feasible methods to mitigate or overcome these issues.
Problem

Research questions and friction points this paper is trying to address.

Exploring MLLMs' potential in medical reporting, diagnosis, and treatment
Reviewing 330 papers to assess MLLMs' healthcare capabilities and limitations
Addressing challenges of MLLMs in medical applications with solutions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Utilizes Multimodal Large Language Models (MLLMs)
Applies MLLMs to medical reporting and diagnosis
Reviews 330 papers for comprehensive healthcare insights
J
Jiarui Ye
School of Computer Science and Engineering, Nanjing University of Science and Technology, Nanjing 210094, China
H
Hao Tang
School of Computer Science, Peking University, Beijing 100871, China