Bridging Research and Practice: A Systematic Evaluation of Generalist and Dermatology-Specific Models in Clinical Skin Lesion Classification

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient model robustness and clinical deployment challenges in skin lesion classification arising from data heterogeneity. We present the first systematic comparative analysis between cross-modal foundation models and task-specific classifiers. Methodologically, by integrating vision-language models, convolutional neural networks, and feature embedding techniques, this work comprehensively evaluates the performance discrepancies between general-purpose and specialized models under distribution shifts while quantifying their generalization gaps. Our findings reveal the inherent limitations of existing state-of-the-art models, providing a critical benchmark and empirical evidence for developing safe, equitable, and trustworthy AI systems in dermatology.
📝 Abstract
The application of machine learning to dermatology has grown substantially in recent years, moving beyond proof-of-concept studies toward potential applications. However, clinical dermatology remains a challenging and still open problem. Diagnostic assessment is often ambiguous, and skin lesions exhibit high variability, compounded by differences in acquisition modality, device quality, and patient demographics. These factors hinder the development of robust models suitable for safe and equitable clinical use. To support translation into practice, it is essential to systematically evaluate how contemporary models generalize across heterogeneous data sources. In this work, we benchmark a diverse set of architectures on recent dermatology datasets, spanning dermoscopic images and smartphone-based clinical photographs. We assess the robustness of recent general-purpose and medical vision-language models, as well as foundation models, and compare them against task-specific dermatology classifiers, including embedding-based approaches and convolutional neural networks. Our study provides an evaluation of model performance under distribution shifts, modality changes, and demographic variability. By quantifying the gap between current state-of-the-art models and the requirements of clinical deployment, we aim to contribute to the development of reliable, accessible, and clinically applicable AI systems for dermatology.
Problem

Research questions and friction points this paper is trying to address.

clinical dermatology
skin lesion classification
model robustness
distribution shift
vision-language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Foundation Models
Dermatology Classification
Distribution Shifts
Systematic Benchmarking
🔎 Similar Papers
No similar papers found.
E
Emanoel dos Santos
Centro de Informática, Universidade Federal de Pernambuco (UFPE), Brazil
K
Kelvin Cunha
Centro de Informática, Universidade Federal de Pernambuco (UFPE), Brazil
R
Rodrigo Mota
Centro de Informática, Universidade Federal de Pernambuco (UFPE), Brazil
F
Fabio Papais
Centro de Informática, Universidade Federal de Pernambuco (UFPE), Brazil
T
Thales Bezerra
Centro de Informática, Universidade Federal de Pernambuco (UFPE), Brazil
N
Natalia Lopes
Centro de Informática, Universidade Federal de Pernambuco (UFPE), Brazil
E
Erico Medeiros
Centro de Informática, Universidade Federal de Pernambuco (UFPE), Brazil
S
Shirley Cruz
Hospital das Clinicas, Ebserh, UFPE, Pernambuco, Brazil
J
Jessica Araujo
Hospital das Clinicas, Ebserh, UFPE, Pernambuco, Brazil
Paulo Borba
Paulo Borba
Federal University of Pernambuco
Software EngineeringProgramming Languages
Tsang Ing Ren
Tsang Ing Ren
Center for Informatics - CIn, Federal University of Pernambuco - UFPE
Image ProcessingComputer VisionPattern RecognitionMachine Learning