🤖 AI Summary
This study systematically reviews medical multimodal large language models (MLLMs), focusing on three core clinical applications: medical report generation, diagnostic support, and therapeutic assistance. Drawing on 330 state-of-the-art publications, it unifies technical methodologies, multimodal data compositions, and evaluation benchmarks for the first time. A taxonomy classifying six principal medical modalities—including radiological imaging, histopathological slides, and electronic health records—is established, along with corresponding evaluation frameworks. Key challenges—data privacy, hallucination, and lack of interpretability—are distilled, and actionable solutions are proposed, including medical vision–language alignment, domain-adaptive fine-tuning, and structured medical knowledge injection. The work culminates in a structured knowledge graph demonstrating that MLLMs have achieved clinically deployable performance in tasks such as radiology report generation and pathology-assisted diagnosis.
📝 Abstract
MLLMs have recently become a focal point in the field of artificial intelligence research. Building on the strong capabilities of LLMs, MLLMs are adept at addressing complex multi-modal tasks. With the release of GPT-4, MLLMs have gained substantial attention from different domains. Researchers have begun to explore the potential of MLLMs in the medical and healthcare domain. In this paper, we first introduce the background and fundamental concepts related to LLMs and MLLMs, while emphasizing the working principles of MLLMs. Subsequently, we summarize three main directions of application within healthcare: medical reporting, medical diagnosis, and medical treatment. Our findings are based on a comprehensive review of 330 recent papers in this area. We illustrate the remarkable capabilities of MLLMs in these domains by providing specific examples. For data, we present six mainstream modes of data along with their corresponding evaluation benchmarks. At the end of the survey, we discuss the challenges faced by MLLMs in the medical and healthcare domain and propose feasible methods to mitigate or overcome these issues.