$\oslash$ Source Models Leak What They Shouldn't $\nrightarrow$: Unlearning Zero-Shot Transfer in Domain Adaptation Through Adversarial Optimization

📅 2026-04-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the privacy risks in source-free domain adaptation, where models trained on the source domain may inadvertently leak information about source-private classes into the target domain. To tackle this issue, we formally introduce and solve the problem of source-private class forgetting under a novel setting termed SCADA-UL, extending it to scenarios involving continual forgetting and unknown forgotten classes. By leveraging adversarially generated forgetting samples, a label rescaling strategy, and adversarial optimization, our method achieves effective machine unlearning under distribution shift. Experimental results demonstrate that the proposed approach attains forgetting performance comparable to full retraining across multiple benchmark datasets, significantly outperforming existing baselines.

Technology Category

Machine Learning: Adversarial Learning & RobustnessComputer Vision: Adversarial Attacks & RobustnessGame Theory and Economic Paradigms: Adversarial Learning

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsSecurity and Privacy: Security and privacy of machine learning and AI applications
📝 Abstract
The increasing adaptation of vision models across domains, such as satellite imagery and medical scans, has raised an emerging privacy risk: models may inadvertently retain and leak sensitive source-domain specific information in the target domain. This creates a compelling use case for machine unlearning to protect the privacy of sensitive source-domain data. Among adaptation techniques, source-free domain adaptation (SFDA) calls for an urgent need for machine unlearning (MU), where the source data itself is protected, yet the source model exposed during adaptation encodes its influence. Our experiments reveal that existing SFDA methods exhibit strong zero-shot performance on source-exclusive classes in the target domain, indicating they inadvertently leak knowledge of these classes into the target domain, even when they are not represented in the target data. We identify and address this risk by proposing an MU setting called SCADA-UL: Unlearning Source-exclusive ClAsses in Domain Adaptation. Existing MU methods do not address this setting as they are not designed to handle data distribution shifts. We propose a new unlearning method, where an adversarially generated forget class sample is unlearned by the model during the domain adaptation process using a novel rescaled labeling strategy and adversarial optimization. We also extend our study to two variants: a continual version of this problem setting and to one where the specific source classes to be forgotten may be unknown. Alongside theoretical interpretations, our comprehensive empirical results show that our method consistently outperforms baselines in the proposed setting while achieving retraining-level unlearning performance on benchmark datasets. Our code is available at https://github.com/D-Arnav/SCADA
Problem

Research questions and friction points this paper is trying to address.

machine unlearning
source-free domain adaptation
privacy leakage
zero-shot transfer
domain adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

machine unlearning
source-free domain adaptation
adversarial optimization
zero-shot transfer leakage
privacy-preserving AI
A
Arnav Devalapally
Indian Institute of Technology, Hyderabad; University of Michigan
P
Poornima Jain
Indian Institute of Technology, Hyderabad
K
Kartik Srinivas
Indian Institute of Technology, Hyderabad; Carnegie Mellon University
V
Vineeth N. Balasubramanian
Indian Institute of Technology, Hyderabad; Microsoft Research