Partial Identifiability for Domain Adaptation

📅 2023-06-10
🏛️ arXiv.org
📈 Citations: 4
✨ Influential: 0
📄 PDF
🤖 AI Summary
Unsupervised domain adaptation (UDA) faces a fundamental identifiability challenge: the target-domain joint distribution (P(Y,X)) is unidentifiable under arbitrary domain shifts, especially when cross-domain causal mechanisms differ only slightly. Method: This paper proposes a generative modeling framework grounded in the causal mechanism stability assumption. We establish the first “partial identifiability” theory, rigorously proving—under mild conditions—that both the latent representation and the target joint distribution are partially recoverable. Building on this, we introduce iMSDA, which explicitly decomposes latent variables into domain-invariant and sparsely varying components, and enforces causal constraints within a variational inference framework to achieve disentangled learning. Contribution/Results: iMSDA achieves significant improvements over state-of-the-art methods across multiple standard UDA benchmarks, empirically validating the effectiveness and generalization robustness of our theory-driven design.
📝 Abstract
Unsupervised domain adaptation is critical to many real-world applications where label information is unavailable in the target domain. In general, without further assumptions, the joint distribution of the features and the label is not identifiable in the target domain. To address this issue, we rely on the property of minimal changes of causal mechanisms across domains to minimize unnecessary influences of distribution shifts. To encode this property, we first formulate the data-generating process using a latent variable model with two partitioned latent subspaces: invariant components whose distributions stay the same across domains and sparse changing components that vary across domains. We further constrain the domain shift to have a restrictive influence on the changing components. Under mild conditions, we show that the latent variables are partially identifiable, from which it follows that the joint distribution of data and labels in the target domain is also identifiable. Given the theoretical insights, we propose a practical domain adaptation framework called iMSDA. Extensive experimental results reveal that iMSDA outperforms state-of-the-art domain adaptation algorithms on benchmark datasets, demonstrating the effectiveness of our framework.
Problem

Research questions and friction points this paper is trying to address.

Unsupervised Domain Adaptation
Data-Label Correlation
Slight Variations in Learning Patterns
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unsupervised Domain Adaptation
iMSDA
Invariant Learning
🔎 Similar Papers
No similar papers found.
Carnegie Mellon University | Mohamed bin Zayed University of Artificial Intelligence | Broad Institute of MIT and Harvard
Lingjing Kong
Lingjing Kong
Carnegie Mellon University
Machine Learning
Shaoan Xie
Shaoan Xie
Carnegie Mellon University
Representation LearningGenerative ModelCausality
W
Weiran Yao
Carnegie Mellon University, USA
Yujia Zheng
Yujia Zheng
Carnegie Mellon University
Machine LearningCausal Discovery and InferenceLatent Variable ModelsGenerative Models
G
Guan-Hong Chen
Mohamed bin Zayed University of Artificial Intelligence, UAE
P
P. Stojanov
Broad Institute of MIT and Harvard, USA
V
Victor Akinwande
Carnegie Mellon University, USA
K
Kun Zhang
Carnegie Mellon University, USA