🤖 AI Summary
This work addresses unsupervised domain adaptation under covariate shift by proposing Target-Induced Loss Tilting (TILT). The method decomposes the source-domain predictor into a shared backbone \(f\) and an auxiliary component \(b\), jointly training \(f + b\) on the source domain while penalizing \(b\) on the target domain, ultimately deploying \(f\) as the target predictor. TILT implicitly implements importance weighting on the target side through this decomposition, yielding an estimator that is locally adaptive and uniformly bounded for any source–target distribution pair—even when their supports are disjoint—thus ensuring stability. The theoretical analysis leverages a novel objective function, sparse ReLU networks, and finite-sample oracle inequalities. Empirical results demonstrate that TILT significantly outperforms source-only training, exact importance weighting, and density ratio baselines on regression tasks and shifted CIFAR-100 distillation, while exhibiting robustness to regularization hyperparameters.
📝 Abstract
We introduce and analyze Target-Induced Loss Tilting (TILT) for unsupervised domain adaptation under covariate shift. It is based on a novel objective function that decomposes the source predictor as $f+b$, fits $f+b$ on labeled source data while simultaneously penalizing the auxiliary component $b$ on unlabeled target inputs. The resulting fit $f$ is deployed as the final target predictor. At the population level, we show that this target-side penalty implicitly induces relative importance weighting at the population level, but in terms of an estimand $b^*_f$ that is self-localized to the current error, and remains uniformly bounded for any source-target pair (even those with disjoint supports). We prove a general finite-sample oracle inequality on the excess risk, and use it to give an end-to-end guarantee for training with sparse ReLU networks. Experiments on controlled regression problems and shifted CIFAR-100 distillation show that TILT improves target-domain performance over source-only training, exact importance weighting, and relative density-ratio baselines, with a stable dependence on the regularization parameter.