multimodal diagnosis generation

Designs and implements systems that ingest multiple input modalities (e.g., images, text, signals) and produce structured diagnostic outputs such as syndrome labels, differential diagnoses, suggested interventions or prescriptions, and explanatory rationale. Builds and fine-tunes multimodal models or pipelines (for example multimodal LLMs), maps visual and other modality-specific features into text or structured fields, and analyzes the fidelity, format compliance, and interpretability of the generated diagnoses.

multimodaldiagnosisgeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.18
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Despite rapid advances in multimodal large language models (MLLMs), their clinical deployment remains hindered by domain-specific bottlenecks—including scarce annotated medical data, modality bias, and limited interpretability. Method: This work systematically reviews the evolution from large language models (LLMs) to MLLMs and empirically analyzes their integration of text, medical imaging, and audio modalities for clinical decision support, radiology/pathology interpretation, patient interaction, and biomedical research. Contribution/Results: We identify three critical research directions: (1) construction of medical-domain multimodal datasets, (2) novel modality alignment techniques, and (3) an ethics-aware governance framework. Empirical evaluation demonstrates that MLLMs improve diagnostic assistance accuracy and accelerate structured reporting generation; however, performance is constrained by data scarcity, cross-modal misalignment, and opaque reasoning. This study provides both theoretical foundations and actionable pathways toward trustworthy, clinically viable MLLMs.

Addresses implementation challenges including data limitations and ethical concernsExamines MLLM applications in clinical decision support and medical imagingExplores the evolution of text-based LLMs to multimodal systems in healthcare

Evaluating Strategies for Synthesizing Clinical Notes for Medical Multimodal AI

Nov 26, 2025
NM
Niccolo Marini
🏛️ National Library of Medicine | National Institutes of Health

Medical multimodal AI is hindered by the scarcity of high-quality heterogeneous data, particularly in dermatology, where image datasets lack rich clinical textual annotations—limiting model robustness and generalization. To address this, we propose a fine-tuning-free prompt engineering framework that leverages structured medical metadata (e.g., lesion location, age, sex) to guide large language models in generating high-fidelity, low-hallucination synthetic clinical notes. This approach is the first to enable image-to-text cross-modal retrieval solely through prompt design. Evaluated across multiple dermatological benchmarks, the synthetic notes significantly improve multimodal classification accuracy (+3.2–7.8%), with even greater gains under domain shift. Our core innovation lies in embedding clinical priors directly into the prompting mechanism—ensuring both clinical plausibility and modeling efficacy—thereby bridging the modality gap without architectural modification or parameter updates.

Enhancing classification and cross-modal retrieval in dermatology datasetsEvaluating strategies to reduce hallucinations in synthetic medical textGenerating synthetic clinical notes using LLMs for multimodal medical AI

This work addresses three critical limitations of medical large language models (MedLMs) in clinical diagnosis: poor interpretability, inadequate multimodal fusion, and difficulty integrating domain-specific knowledge. To tackle these challenges, we propose a unified diagnostic AI evolution framework comprising three synergistic pillars: domain adaptation, structured knowledge injection, and cross-modal alignment. For the first time, we systematically integrate four paradigmatic approaches—BioBERT, Med-PaLM, DR.KNOWS, and multimodal foundation models (MMFMs)—by unifying biomedical pretraining, instruction tuning, medical knowledge graph embedding, and vision-language joint modeling. Evaluated on benchmarks including MultiMedQA, our approach achieves significant improvements in clinical question answering accuracy, named entity recognition F1-score, and lesion detection performance, while simultaneously reducing harmful outputs. Moreover, it enhances model interpretability and clinical utility, delivering a deployable technical framework for personalized medicine.

Advancing medical image analysis and anatomical modelingEnhancing disease prediction and diagnosisImproving personalized treatment and drug discovery

Advancing Conversational Diagnostic AI with Multimodal Reasoning

May 06, 2025
KS
Khaled Saab
🏛️ Google DeepMind | Google Research

Current conversational diagnostic systems predominantly rely on text-only interaction, failing to support real-time multimodal clinical analysis—such as medical images, ECGs, and PDF reports—essential for telemedicine. To address this, we propose a state-driven multimodal conversational diagnostic framework built upon Gemini 2.0 Flash, featuring a dynamic state-aware mechanism that jointly enables multimodal understanding, uncertainty modeling, and structured clinical questioning. Crucially, we introduce the first method for autonomously generating follow-up questions based on patient-state uncertainty, emulating expert clinicians’ diagnostic reasoning. Evaluated on 105 OSCE cases, our system significantly outperformed general practitioners across 7 of 9 multimodal and 29 of 32 non-multimodal clinical dimensions—including diagnostic accuracy—demonstrating synergistic enhancement between multimodal capability and diagnostic efficacy.

Comparing AI diagnostic performance with physicians in structured clinical scenariosEnhancing AI's ability to interpret diverse medical data during consultationsEvaluating LLMs for multimodal medical diagnosis beyond text-only interactions

Leveraging Audio and Text Modalities in Mental Health: A Study of LLMs Performance

Dec 09, 2024
AA
Abdelrahman A. Ali
🏛️ Compumacy for Artificial Intelligence solutions

This study investigates the capability of large language models (LLMs) for zero-shot multimodal (text + audio) diagnosis of depression and PTSD. Leveraging the E-DAIC dataset, we systematically evaluate single-modal and cross-modal fusion performance of models including Gemini 1.5 Pro and GPT-4o mini. Methodologically, we propose two novel metrics—Modal Superiority Score and Disagreement Resolution Score—to quantify multimodal synergy, and employ zero-shot prompt engineering—without fine-tuning—to enable end-to-end cross-modal reasoning. Results demonstrate that Gemini 1.5 Pro achieves an F1-score of 0.67 and balanced accuracy of 77.4% under multimodal fusion, outperforming the best single-modal baseline by 2.7–3.1%. This work constitutes the first empirical validation that LLMs can effectively and robustly perform early mental health screening across unseen diagnostic categories via zero-shot cross-modal integration.

Detecting depression and PTSD using text and audio modalities.Enhancing diagnostic accuracy by integrating text and audio inputs.Evaluating LLMs performance in multimodal mental health diagnostics.

Latest Papers

What's happening recently
View more

Existing clinical diagnostic evaluation methods struggle to capture the true complexity of progressive multimodal information disclosure and dynamic reasoning. To address this gap, this work proposes ClinMM-Bench—the first large-scale, multi-turn, multimodal diagnostic benchmark grounded in real-world clinical scenarios—encompassing 1,089 challenging cases and 3,760 medical images, along with a dual-layer evaluation framework that systematically assesses both diagnostic accuracy and reasoning quality. Evaluation of 15 state-of-the-art multimodal large language models reveals that, although closed-source models achieve the best performance, their fully correct diagnosis rate remains limited; while models can generate plausible hypotheses, they exhibit significant deficiencies in constructing reliable and coherent reasoning chains. The study further identifies five representative failure patterns, offering concrete directions for future model improvement.

clinical diagnosisdiagnostic reasoningmulti-turn evaluation

This study addresses the challenge that multimodal models struggle to reliably bind textual prompts (e.g., “image”) to their corresponding input modalities, resulting in inadequate source modality tracking. The work formally defines and empirically investigates the “source modality monitoring” problem for the first time, introducing an evaluation paradigm grounded in target-modality information retrieval. By integrating syntactic manipulation with semantic perturbation, the authors systematically assess binding mechanisms across eleven prominent vision-language models. Their findings reveal that semantic cues dominate the binding process when modality distributions exhibit significant divergence, consistently outweighing syntactic signals. These insights offer critical evidence and a novel perspective for enhancing the reliability and robustness of multimodal agents.

binding probleminformation originmultimodal models

This work addresses the limitations of handcrafted input representations in supervised learning on complex heterogeneous data—such as time series, text, and structured records—by proposing a large language model (LLM)-based agent pipeline. The approach automatically derives global normalization rules from a small set of diverse textual serialization examples and combines them with task-conditioned local rules to construct efficient, auditable, and low-cost tabular input representations. Evaluated on 15 clinical tasks in the EHRSHOT benchmark, the method significantly outperforms conventional count-based feature models, naive LLM text serialization baselines, and large-scale pretrained clinical foundation models, demonstrating strong effectiveness and generalization capability in few-shot medical settings.

domain-specific engineeringheterogeneous datasetsinput representation

This work addresses the challenge of leveraging unstructured clinical notes in electronic health records (EHR) during training to enhance the performance of models that rely solely on structured data at deployment. The authors propose a multimodal learning framework that integrates clinical notes—encoded via BioClinicalBERT—with structured features such as demographics and medical codes during training. By combining a teacher–student architecture, contrastive learning, and contrastive knowledge distillation, the approach enables the final deployed model to achieve high inference efficiency using only structured inputs, without requiring access to textual data at test time. To the best of the authors’ knowledge, this is the first method to effectively augment structured EHR representations with unstructured clinical notes while maintaining deployment constraints. Evaluated on a cohort of 3,466 children with late language emergence, the model achieves an AUROC of 0.705, significantly outperforming baseline methods (AUROC = 0.656).

model deploymentmultimodal trainingstructured EHR data

Hot Scholars

ZW

Zhe Wang

Tsinghua University
Computer VisionAutonomous Driving
GA

Geo Ahn

Kyung Hee University, Republic of Korea
Computer visionVideo understanding
XL

Xiaoxiao Li

Assistant Professor, UBC; Vector Institute; CIFAR AI Chair; Canada Research Chair
Deep LearningTrustworthy AIAI for Healthcare
WD

Wenlong Deng

University of British Columbia, Vector Institute
Machine LearningLarge Language Model
YL

Yushu Li

University of British Columbia
Natural Language ProcessingComputer VisionMachine Learning