Sinkhorn Distributionally Robust Optimization

πŸ“… 2021-09-24
πŸ›οΈ arXiv.org
πŸ“ˆ Citations: 33
✨ Influential: 9
πŸ“„ PDF
πŸ€– AI Summary
This work addresses key limitations of Wasserstein distributionally robust optimization (DRO), including high computational cost in computing the Wasserstein distance and the unrealistic nature of the worst-case distribution. We propose a general DRO framework based on the Sinkhorn distanceβ€”an entropy-regularized variant of the Wasserstein distance. First, we establish a convex dual characterization of Sinkhorn-DRO applicable to arbitrary nominal distributions, transportation costs, and loss functions, ensuring both tractability and modeling flexibility while yielding more realistic worst-case distributions. Second, we design a stochastic mirror descent algorithm with bias-corrected gradient estimation, achieving near-optimal sample complexity and supporting both smooth and nonsmooth losses. Experiments on synthetic and real-world datasets demonstrate that our method significantly improves computational efficiency, generalization performance, and robustness compared to classical Wasserstein DRO.
πŸ“ Abstract
We study distributionally robust optimization (DRO) with Sinkhorn distance -- a variant of Wasserstein distance based on entropic regularization. We derive convex programming dual reformulation for general nominal distributions, transport costs, and loss functions. Compared with Wasserstein DRO, our proposed approach offers enhanced computational tractability for a broader class of loss functions, and the worst-case distribution exhibits greater plausibility in practical scenarios. To solve the dual reformulation, we develop a stochastic mirror descent algorithm with biased gradient oracles. Remarkably, this algorithm achieves near-optimal sample complexity for both smooth and nonsmooth loss functions, nearly matching the sample complexity of the Empirical Risk Minimization counterpart. Finally, we provide numerical examples using synthetic and real data to demonstrate its superior performance.
Problem

Research questions and friction points this paper is trying to address.

Study distributionally robust optimization with Sinkhorn distance
Derive convex dual reformulation for general distributions
Develop stochastic mirror descent algorithm for solution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses Sinkhorn distance for robust optimization
Develops convex dual reformulation method
Proposes stochastic mirror descent algorithm
πŸ’Ό Related Jobs
No related jobs found.
Georgia Institute of Technology | University of Texas at Austin
J
Jie Wang
School of Industrial and Systems Engineering, Georgia Institute of Technology, Atlanta, GA 30332
R
Rui Gao
Department of Information, Risk, and Operations Management, University of Texas at Austin, Austin, TX 78712
Yao Xie
Yao Xie
Coca-Cola Foundation Chair and Professor, Georgia Institute of Technology
statisticsmachine learningoptimizationsignal processing