🤖 AI Summary
This paper addresses density ratio estimation under covariate shift without assuming boundedness—a departure from conventional bounded-density-ratio assumptions. Methodologically, it establishes novel upper bounds on estimation error for unbounded domains and ranges, leveraging least-squares and logistic regression losses, nonparametric estimation, and minimax analysis to derive optimal convergence rates up to logarithmic factors. Theoretically, it is the first to reveal that tail decay behavior of the density ratio fundamentally governs cross-domain generalization error, and identifies a sufficient condition—requiring no loss correction—for controlled generalization error. Moreover, it proves that, on suitably chosen target domains, source-domain estimators can outperform correction-based methods relying on the true density ratio. Extensive simulations validate the theoretical findings. The results provide a rigorous, error-controlled generalization theory for nonparametric regression and conditional flow models under covariate shift.
📝 Abstract
The density ratio is an important metric for evaluating the relative likelihood of two probability distributions, with extensive applications in statistics and machine learning. However, existing estimation theories for density ratios often depend on stringent regularity conditions, mainly focusing on density ratio functions with bounded domains and ranges. In this paper, we study density ratio estimators using loss functions based on least squares and logistic regression. We establish upper bounds on estimation errors with standard minimax optimal rates, up to logarithmic factors. Our results accommodate density ratio functions with unbounded domains and ranges. We apply our results to nonparametric regression and conditional flow models under covariate shift and identify the tail properties of the density ratio as crucial for error control across domains affected by covariate shift. We provide sufficient conditions under which loss correction is unnecessary and demonstrate effective generalization capabilities of a source estimator to any suitable target domain. Our simulation experiments support these theoretical findings, indicating that the source estimator can outperform those derived from loss correction methods, even when the true density ratio is known.