Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases

📅 2026-07-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing clinical diagnostic evaluation methods struggle to capture the true complexity of progressive multimodal information disclosure and dynamic reasoning. To address this gap, this work proposes ClinMM-Bench—the first large-scale, multi-turn, multimodal diagnostic benchmark grounded in real-world clinical scenarios—encompassing 1,089 challenging cases and 3,760 medical images, along with a dual-layer evaluation framework that systematically assesses both diagnostic accuracy and reasoning quality. Evaluation of 15 state-of-the-art multimodal large language models reveals that, although closed-source models achieve the best performance, their fully correct diagnosis rate remains limited; while models can generate plausible hypotheses, they exhibit significant deficiencies in constructing reliable and coherent reasoning chains. The study further identifies five representative failure patterns, offering concrete directions for future model improvement.
📝 Abstract
Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive disclosure of multimodal information, dynamic updating of diagnostic hypotheses, and continuous refinement of clinical reasoning. However, existing evaluations of multimodal large language models (MLLMs) typically rely on single-turn or isolated tasks, making it difficult to fully capture the complexity of real-world clinical diagnosis. To bridge this gap, we developed ClinMM-Bench, the largest multi-turn multimodal clinical diagnostic evaluation benchmark to date. ClinMM-Bench contains 1,089 challenging real-world clinical cases and 3,760 medical images across eight specialties. We systematically evaluated 15 representative MLLMs using a two-level evaluation framework that assessed both diagnostic accuracy and diagnostic reasoning quality. Results showed that proprietary models achieved the highest overall diagnostic accuracy, but the proportion of completely correct diagnoses remained limited across all models. In terms of diagnostic reasoning quality, current models can identify plausible diagnostic directions but still have considerable limitations in generating reliable diagnostic reasoning. Error analysis further identified five representative failure modes: information synthesis failure, knowledge mapping error, perception error, premature closure, and visual hallucination.
Problem

Research questions and friction points this paper is trying to address.

multimodal large language models
clinical diagnosis
multi-turn evaluation
diagnostic reasoning
real-world clinical cases
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-turn evaluation
multimodal diagnostic reasoning
ClinMM-Bench
clinical reasoning quality
failure mode analysis
🔎 Similar Papers
Rui Yang
Rui Yang
Duke-NUS Medical School
Medical InformaticsMedical Text MiningMedical Knowledge Graph
Weihao Xuan
Weihao Xuan
The University of Tokyo, RIKEN
Natural Language ProcessingComputer VisionMultimodal AIGenerative AILLM Agent
Y
Yi Lin
Department of Population Health Sciences, Weill Cornell Medicine, New York, NY, USA
Z
Zhuhan Bao
Department of Biostatistics and Bioinformatics, Duke University, Durham, NC, USA
J
Jonathan Chong Kai Liew
Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA, USA; Hospital of the University of Pennsylvania, Philadelphia, PA, USA
M
Matthew Yu Heng Wong
School of Clinical Medicine, University of Cambridge, Cambridge, UK
N
Nicolás Lescano
Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA, USA; Hospital of the University of Pennsylvania, Philadelphia, PA, USA
N
Nikita R. Paripati
Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA, USA; Hospital of the University of Pennsylvania, Philadelphia, PA, USA; Children’s Hospital of Philadelphia (CHOP), Philadelphia, PA, USA
E
Emily Ling-Lin Pai
Department of Anatomic Pathology and Laboratory Medicine, Hospital of the University of Pennsylvania, PA, USA; Department of Pathology and Laboratory Medicine, University of California, San Francisco, CA, USA
Jiarui Liu
Jiarui Liu
Carnegie Mellon University
Natural Language Processing
Heli Qi
Heli Qi
Waseda University, RIKEN
Multi-Modal Learning
Heng-Jui Chang
Heng-Jui Chang
Massachusetts Institute of Technology
Speech ProcessingDeep Learning
B
Benny Kai Guo Loo
Sport and Exercise Medicine Service, KK Women’s and Children’s Hospital, Singapore, Singapore; Paediatrics Academic Clinical Programme, SingHealth Duke-NUS Academic Medical center, Singapore, Singapore
Huitao Li
Huitao Li
Duke-Nus Medical School
Medical Informatics
K
Kunyu Yu
Center for Biomedical Data Science, Duke-NUS Medical School, Singapore, Singapore; Duke-NUS AI + Medical Sciences Initiative, Duke-NUS Medical School, Singapore, Singapore
Y
Yufan Wang
Department of Population Health Sciences, Weill Cornell Medicine, New York, NY, USA
C
Chuan Hong
Department of Biostatistics and Bioinformatics, Duke University, Durham, NC, USA
Shijian Lu
Shijian Lu
College of Computing and Data Science, NTU
Image and video analyticscomputer visionmachine learning
Douglas Teodoro
Douglas Teodoro
Professor, University of Geneva
biomedical NLPmachine learning for healthcaremedical informatics
Naoto Yokoya
Naoto Yokoya
The University of Tokyo, RIKEN
Remote SensingComputer VisionMachine LearningData Fusion
Ross Koppel
Ross Koppel
University of Pennsylvania
patient safetyusabilityworkflowcybersecurity
Mona Diab
Mona Diab
Professor & Director of Language Technologies Institute, Carnegie Mellon University, ACL Fellow
Responsible AINLP/CLArabic NLPCross lingual/multilingual & Low Resource Lang Processing
Hua Xu
Hua Xu
Robert T. McCluskey Professor, Section of Biomedical Informatics and Data Science, Yale University
natural language processingtext mining
D
David W. Bates
Division of General Internal Medicine and Primary Care, Brigham and Women’s Hospital, Boston, MA, USA; Department of Medicine, Harvard Medical School, Boston, MA, USA; Department of Health Care Policy and Management, Harvard T. H. Chan School of Public Health, Boston, MA, USA
N
Nan Liu
Center for Biomedical Data Science, Duke-NUS Medical School, Singapore, Singapore; Duke-NUS AI + Medical Sciences Initiative, Duke-NUS Medical School, Singapore, Singapore; Department of Biostatistics and Bioinformatics, Duke University, Durham, NC, USA; Pre-hospital and Emergency Research Center, Health Services Research and Population Health, Duke-NUS Medical School, Singapore, Singapore; NUS Artificial Intelligence Institute, National University of Singapore, Singapore, Singapore