TC-ADA: One-Shot Active Domain Adaptation for Semantic Segmentation

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive annotation costs and inefficiency of multi-round active domain adaptation in cross-domain semantic segmentation by proposing a single-round, image-level active domain adaptation framework. The method leverages vision foundation models to identify representative target images and acquire dense annotations in a single pass. By integrating these representations with unsupervised domain adaptation predictions and semi-supervised learning, it establishes a joint calibration mechanism that aligns source-domain knowledge with limited target supervision, enabling uninterrupted optimization. Evaluated across multiple synthetic-to-real and real-to-real driving scenario transfers, the proposed approach surpasses existing baselines using minimal annotations, achieving performance closely approaching fully supervised upper bounds with a gap of less than 1.9 mIoU.
📝 Abstract
Manual dense annotation remains a major obstacle to deploying semantic segmentation models in new driving environments. Active domain adaptation (ADA) seeks label-efficient transfer by annotating only a selected portion of the target domain. Existing ADA methods commonly implement this process through multiple rounds of acquisition, annotation, and retraining. We study a practical one-shot image-level setting that selects and densely annotates a fixed target subset in a single round, followed by uninterrupted adaptation. Within this setting, we develop Target-Calibrated Active Domain Adaptation (TC-ADA) as a joint design of complete-image acquisition and target-calibrated adaptation. Stage~1 uses visual representations from a vision foundation model (VFM) together with semantic predictions from a fixed unsupervised domain adaptation model to select representative and informative target images without target annotations. Stage~2 jointly uses labeled source data, labeled target data, and the remaining unlabeled target data, while calibrating source and target supervision under limited target labels. Extensive experiments across five synthetic-to-real and real-to-real driving transfers show consistent improvements over representative ADA baselines. With only 23 to 46 labeled target images on four transfers and 140 on Mapillary, TC-ADA stays within 1.9 mean intersection over union (mIoU) points of target-only full supervision. Code will be available at https://github.com/ywher/TC-ADA.
Problem

Research questions and friction points this paper is trying to address.

Semantic Segmentation
Active Domain Adaptation
One-Shot Learning
Driving Environments
Label Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

One-Shot Active Domain Adaptation
Semantic Segmentation
Vision Foundation Model
Target-Calibrated Adaptation
Domain Adaptation
🔎 Similar Papers
No similar papers found.