🤖 AI Summary
This study addresses the significant clinical challenge of comprehensively assessing chronic wounds, which requires integrated evaluation of morphology, tissue characteristics, vascular status, and infection risk. The authors construct a real-world clinical dataset comprising 20 multi-etiology wound cases and, for the first time, systematically evaluate the clinical reasoning capabilities of six vision-language models (VLMs) on unseen authentic wound images using a structured twelve-question clinical framework. Performance is assessed across wound classification, infection risk identification, and treatment recommendation tasks. Results demonstrate that general-purpose models substantially outperform medical-specific counterparts: ChatGPT achieves the highest accuracy at 72.5%, followed by Claude at 62.08%, while HuluMed emerges as the top-performing medical model at 40.00%. This work highlights the untapped potential of current general multimodal models in real-world clinical settings and establishes a new benchmark for intelligent wound assessment.
📝 Abstract
Chronic wound assessment remains a clinically challenging task that requires accurate interpretation of wound morphology, tissue composition, vascular characteristics, and infection risk. Recent advances in Vision-Language Models (VLMs) have introduced the possibility of automated multimodal wound analysis through image understanding combined with clinical reasoning. This study evaluates the performance of several general-purpose and medically specialized open-source and proprietary VLMs for clinical wound assessment using an expanded, curated dataset of 20 clinically diverse wounds spanning vascular, surgical, ischemic, venous, lymphedema, and amputation-related etiologies. Six VLMs were evaluated using a structured twelve-question clinical framework covering wound classification, infection risk, vascular intervention recommendations, debridement urgency, wound therapy selection, and advanced management planning. Across 20 wound cases and 240 clinician-graded wound-analysis decisions, ChatGPT achieved the highest overall performance with 174/240 correct responses (72.50%), followed by Claude with 149/240 (62.08%). Among the open-source and medically specialized models, HuluMed achieved the strongest performance with 96/240 correct responses (40.00%), followed by Gemma 3 (81/240, 33.75%), MedGemma 4B (62/240, 25.83%), and MedGemma 27B (42/240, 17.50%). The findings suggest that frontier general-purpose multimodal systems currently demonstrate substantially stronger wound-analysis performance than medically specialized alternatives, highlighting the continued importance of broad multimodal reasoning capabilities alongside domain-specific medical knowledge. Although current VLMs demonstrate promising potential for clinical decision support, substantial limitations remain in advanced wound-management reasoning, procedural planning, and autonomous clinical reliability.