🤖 AI Summary
This work addresses the performance bottlenecks in semi-supervised semantic segmentation caused by noisy pseudo-labels and domain discrepancies between labeled and unlabeled data. To mitigate these issues, the authors propose a novel approach that integrates ClassMix augmentation with supervised–unsupervised feature alignment. Specifically, ground-truth class regions from labeled images are pasted onto unlabeled images and their corresponding pseudo-labels, while a feature discriminator enforces alignment of the model’s predictions with the feature distribution of the labeled data. This dual strategy effectively reduces the adverse effects of inaccurate pseudo-labels and domain shift. Experimental results on the CHASE and COVID-19 datasets demonstrate a consistent improvement, with an average mIoU gain of 2.07%, substantially outperforming existing semi-supervised segmentation methods.
📝 Abstract
In semantic segmentation, the creation of pixel-level labels for training data incurs significant costs. To address this problem, semi-supervised learning, which utilizes a small number of labeled images alongside unlabeled images to enhance the performance, has gained attention. A conventional semi-supervised learning method, ClassMix, pastes class labels predicted from unlabeled images onto other images. However, since ClassMix performs operations using pseudo-labels obtained from unlabeled images, there is a risk of handling inaccurate labels. Additionally, there is a gap in data quality between labeled and unlabeled images, which can impact the feature maps. This study addresses these two issues. First, we propose a method where class labels from labeled images, along with the corresponding image regions, are pasted onto unlabeled images and their pseudo-labeled images. Second, we introduce a method that trains the model to make predictions on unlabeled images more similar to those on labeled images. Experiments on the Chase and COVID-19 datasets demonstrated an average improvement of 2.07% in mIoU compared to conventional semi-supervised learning methods.