Evaluation of Medical Vision Language Models HuluMed and MedGemma, and general purpose chatbots Gemma 3, ChatGPT Plus, and Claude Pro on real previously unseen wound images

📅 2026-06-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant clinical challenge of comprehensively assessing chronic wounds, which requires integrated evaluation of morphology, tissue characteristics, vascular status, and infection risk. The authors construct a real-world clinical dataset comprising 20 multi-etiology wound cases and, for the first time, systematically evaluate the clinical reasoning capabilities of six vision-language models (VLMs) on unseen authentic wound images using a structured twelve-question clinical framework. Performance is assessed across wound classification, infection risk identification, and treatment recommendation tasks. Results demonstrate that general-purpose models substantially outperform medical-specific counterparts: ChatGPT achieves the highest accuracy at 72.5%, followed by Claude at 62.08%, while HuluMed emerges as the top-performing medical model at 40.00%. This work highlights the untapped potential of current general multimodal models in real-world clinical settings and establishes a new benchmark for intelligent wound assessment.
📝 Abstract
Chronic wound assessment remains a clinically challenging task that requires accurate interpretation of wound morphology, tissue composition, vascular characteristics, and infection risk. Recent advances in Vision-Language Models (VLMs) have introduced the possibility of automated multimodal wound analysis through image understanding combined with clinical reasoning. This study evaluates the performance of several general-purpose and medically specialized open-source and proprietary VLMs for clinical wound assessment using an expanded, curated dataset of 20 clinically diverse wounds spanning vascular, surgical, ischemic, venous, lymphedema, and amputation-related etiologies. Six VLMs were evaluated using a structured twelve-question clinical framework covering wound classification, infection risk, vascular intervention recommendations, debridement urgency, wound therapy selection, and advanced management planning. Across 20 wound cases and 240 clinician-graded wound-analysis decisions, ChatGPT achieved the highest overall performance with 174/240 correct responses (72.50%), followed by Claude with 149/240 (62.08%). Among the open-source and medically specialized models, HuluMed achieved the strongest performance with 96/240 correct responses (40.00%), followed by Gemma 3 (81/240, 33.75%), MedGemma 4B (62/240, 25.83%), and MedGemma 27B (42/240, 17.50%). The findings suggest that frontier general-purpose multimodal systems currently demonstrate substantially stronger wound-analysis performance than medically specialized alternatives, highlighting the continued importance of broad multimodal reasoning capabilities alongside domain-specific medical knowledge. Although current VLMs demonstrate promising potential for clinical decision support, substantial limitations remain in advanced wound-management reasoning, procedural planning, and autonomous clinical reliability.
Problem

Research questions and friction points this paper is trying to address.

chronic wound assessment
wound morphology
infection risk
vascular characteristics
clinical decision support
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
chronic wound assessment
multimodal clinical reasoning
medical AI evaluation
wound image analysis
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yunzhe Xue
Department of Data Science, New Jersey Institute of Technology, Newark, NJ, USA
M
Mohammed Saim Ahmed Quadri
Department of Computer Science, New Jersey Institute of Technology, Newark, NJ, USA
N
Neal Panse
Vascular and Endovascular Surgery, Robert Wood Johnson Hospital, New Brunswick, NJ, USA
J
Justin W. Ady
Vascular and Endovascular Surgery, Robert Wood Johnson Hospital, New Brunswick, NJ, USA
Usman Roshan
Usman Roshan
Associate Professor, Department of Computer Science, New Jersey Institute of Technology
Machine learningDeep learningMedical AIBioinformatics