🤖 AI Summary
This study addresses the subjectivity of conventional acne assessment, its neglect of lesion size, and the lack of spatial priors in existing segmentation models by proposing a multimodal acne analysis framework based on vision-language models. The method leverages CLIP with regional text prompts to incorporate spatial prior knowledge, enabling precise localization and segmentation of facial lesions. Furthermore, it introduces a novel single global prompt protocol that operates without positional information, combined with area-based severity quantification to support cross-domain generalization on unconstrained smartphone images. Experimental results demonstrate that the proposed model achieves a Dice coefficient of 0.5082, outperforming existing methods. Its area measurements exhibit correlation with expert annotations comparable to established standards, while maintaining robust performance across diverse datasets.
📝 Abstract
Acne assessment is crucial for clinical decision-making, yet traditional grading and counting are subjective and fail to account for lesion size. While area-based assessment has emerged as a promising alternative, acne segmentation has continued to rely on general-purpose architectures. To address this gap, we propose VL-AcneSeg, a multimodal framework for acne lesion segmentation that leverages CLIP and region-level text prompts to incorporate spatial priors, enabling lesions to be localized across the whole face. Because region-level prompts indicate which facial areas contain lesions, we report a single global prompt, which requires no such information, as our primary setting. On our internal clinical dataset, VL-AcneSeg achieves a Dice score of 0.5082 and an IoU of 0.3407 under this protocol, the highest among all compared methods, including recent vision-language segmentation methods that are themselves given region-level prompts; region-level prompting raises these to 0.5296 and 0.3602. Moreover, lesion area measurements derived from our segmentation correlate with IGA scores at a level comparable to expert annotations (Pearson r = 0.719 versus 0.658). Notably, our framework maintains consistent performance across external validation datasets, performing reliably even on uncontrolled smartphone images without requiring additional training or fine-tuning. By pairing a protocol that requires no lesion-location information with area-based severity estimation, this work provides a foundation for objective acne assessment outside the clinic. Our implementation is publicly available at: https://github.com/sukjuoh/VL-AcneSeg