🤖 AI Summary
Current visual explanations for skin disease classification models lack systematic, clinically relevant evaluation. This work proposes an explainability assessment framework grounded in large language models (LLMs), integrating Grad-CAM heatmaps with state-of-the-art LLMs—including GPT-5.5, Gemini 3.5 Flash, and Claude Sonnet 4.6—guided by progressive, structured prompt engineering to evaluate both lesion localization accuracy and explanation credibility. For the first time, a domain-adapted LLM-as-a-Judge paradigm is introduced into dermatological AI explainability assessment. Validated on models such as EfficientNet-B0, MobileNetV3, and ResNet18, the approach significantly enhances alignment between automated evaluations and clinical judgments, enabling quantifiable and standardized measurement of explanation quality.
📝 Abstract
This study proposes a domain-specific LLM-based Visual Explanation Evaluation Framework for assessing Grad-CAM explanations in facial skin disease diagnosis models. While previous studies have primarily focused on improving classification performance through data augmentation techniques, relatively few studies have systematically examined whether model explanations are grounded in clinically relevant lesion regions.
In this study, geometric augmentation, color-based augmentation, and mixed augmentation strategies were applied to facial skin disease classification models based on EfficientNet-B0, MobileNetV3, and ResNet18. Grad-CAM was employed to generate visual explanations representing the models' decision-making processes. Furthermore, an LLM-as-a-Judge evaluation framework was designed using GPT-5.5, Gemini 3.5 Flash, and Claude Sonnet 4.6 to assess Grad-CAM explanations from the perspectives of lesion localization and explanation trustworthiness. To improve evaluation consistency and clinical grounding, a progressive prompt engineering strategy was introduced, incorporating evaluation rubrics, clinical knowledge, penalty rules, and structured output formats.