🤖 AI Summary
This study addresses the pseudo-label bias caused by foreground-background imbalance and the leakage of such imbalance during optimization in semi-supervised remote sensing image extraction. To tackle these issues, this work proposes a dual-stage class rebalancing framework that reveals the inherent limitations of single-stage balancing. It pioneers a joint mechanism regulating both pseudo-label generation and unsupervised optimization through adaptive class thresholds, confidence-aware reweighting, distribution alignment, and self-training modules, effectively suppressing the resurgence of background bias during optimization. Experimental results demonstrate that the proposed framework achieves state-of-the-art IoU and F1 scores across multiple datasets under low annotation rates, surpassing fully supervised baselines using only 1% labeled data.
📝 Abstract
Accurate building footprint extraction from high-resolution remote sensing imagery is essential for urban planning, disaster response, and environmental monitoring. However, obtaining dense pixel-level annotations is costly, motivating the use of semi-supervised learning (SSL) to leverage unlabeled imagery. In remote sensing, severe foreground--background imbalance poses a particular challenge for self-training, as it can bias pseudo-label generation and the resulting unsupervised optimization toward the majority background class. We show that addressing this imbalance at only one stage is insufficient: balancing pseudo-label selection alone does not prevent background bias from re-emerging during unsupervised loss optimization, a failure mode we term \emph{imbalance leak}. To address this issue, we propose \textbf{RBMatch}, a dual-level class-rebalancing framework that jointly regulates pseudo-label generation and unsupervised optimization. RBMatch combines a supervised learning pathway with a self-training module comprising three components: adaptive class-specific thresholding (ACT) for balanced pseudo-label selection, confidence-aware class-balanced reweighting (CACBR) for mitigating class bias in the unsupervised loss, and distribution alignment (DAL) for matching the predicted unlabeled-data distribution to the labeled-data prior. Experiments on the WHU, INRIA, and Massachusetts building footprint datasets across labeled ratios of 1%--10% show that RBMatch consistently achieves the best building IoU and F1-score among the evaluated methods. The improvement is most pronounced on the highly imbalanced Massachusetts dataset, where RBMatch improves IoU by 1.37 points over the strongest baseline at a 1% labeling ratio and is the only method to outperform the fully supervised baseline across all twelve dataset--ratio settings.