🤖 AI Summary
This study addresses the absence of a unified theoretical framework for uncertainty estimation and generalization analysis in distributionally robust linear regression. By leveraging the Wasserstein distance, the proposed approach employs quadratic reformulation to unify square-root Lasso and adversarial training, revealing the equivalence between large and small ambiguity set solutions while establishing noise-insensitive pivotal properties. The primary contribution lies in deriving non-asymptotic error bounds that achieve convergence rates of O(n^{-1/2}) under general conditions and O(n^{-1}) under sparsity assumptions. Numerical experiments further demonstrate that the method successfully combines theoretical rigor with computationally efficient solving capabilities.
📝 Abstract
Distributionally robust optimization (DRO) studies parameter estimation under uncertainty in the underlying probability distribution and has emerged as a principled framework for analyzing robustness and generalization. In particular, Wasserstein DRO, with distributional uncertainty induced by the Wasserstein distance, generalizes several popular regularizers. This paper studies Wasserstein DRO linear regression, unifying square-root Lasso and adversarial linear regression as important special cases. We prove that many properties of these two special cases carry over to this general method. In particular, we show (i) deterministic and non-asymptotic in-sample error bounds $O(n^{-1/2})$ in general and $O(n^{-1})$ under design matrix and sparsity conditions; (ii) insensitivity to the noise level, also known as the pivotal property; and (iii) solution equivalences for small and large ambiguity sets. The key proof step is to recast the method into a quadratic form, mimicking adversarial linear regression. We also show that the method can be solved efficiently, and we validate our findings through numerical simulations.