🤖 AI Summary
This study addresses the challenges of confirmation bias and the complexity of multi-stage training pipelines in unsupervised domain adaptation for panoptic segmentation by proposing a concise and robust single-stage training framework. Methodologically, it introduces self-supervised visual encoder initialization to enhance feature transferability and designs a class-level adaptive mask loss scaling strategy to effectively suppress pseudo-label noise. Compared with existing approaches, this work unifies cumbersome multi-stage procedures into end-to-end training. By maintaining architectural simplicity, the proposed framework significantly improves cross-domain generalization performance while substantially reducing reliance on manual annotations.
📝 Abstract
Unsupervised domain adaptation (UDA) reduces the annotation burden in panoptic segmentation by leveraging a cost-effectively labeled source domain (e.g., synthetic) and an unlabeled target domain to bridge the distribution gap. Existing panoptic UDA methods rely on teacher-student consistency learning built upon suboptimal per-pixel segmentation architectures. In contrast, state-of-the-art mask transformers are rarely adopted due to their pronounced vulnerability to confirmation bias in consistency learning, where erroneous teacher predictions are reinforced during training. Our earlier approach, MC-PanDA, mitigates this issue through fine-grained confidence estimation, which suppresses gradients from unreliable masks while sampling informative yet reliable locations for loss computation. However, this method entails a complex multi-stage training and requires careful hyperparameter tuning. This work presents MC-PanDA++, which addresses these limitations by introducing: (i) self-supervised vision encoders that provide a stronger and more robust initialization, further reducing the reliance on human annotations, (ii) per-class, self-adapting mask-wide loss scaling that stabilizes training and enables the usage of a single set of hyperparameters across domains, and (iii) a single-stage training pipeline that decreases overall conceptual complexity. Together, these improvements result in a conceptually simpler, better-performing, and more robust method for domain-adaptive panoptics. Source code: https://github.com/martinovicivan/MC-PanDA