π€ AI Summary
This work addresses key limitations of Wasserstein distributionally robust optimization (DRO), including high computational cost in computing the Wasserstein distance and the unrealistic nature of the worst-case distribution. We propose a general DRO framework based on the Sinkhorn distanceβan entropy-regularized variant of the Wasserstein distance. First, we establish a convex dual characterization of Sinkhorn-DRO applicable to arbitrary nominal distributions, transportation costs, and loss functions, ensuring both tractability and modeling flexibility while yielding more realistic worst-case distributions. Second, we design a stochastic mirror descent algorithm with bias-corrected gradient estimation, achieving near-optimal sample complexity and supporting both smooth and nonsmooth losses. Experiments on synthetic and real-world datasets demonstrate that our method significantly improves computational efficiency, generalization performance, and robustness compared to classical Wasserstein DRO.
π Abstract
We study distributionally robust optimization (DRO) with Sinkhorn distance -- a variant of Wasserstein distance based on entropic regularization. We derive convex programming dual reformulation for general nominal distributions, transport costs, and loss functions. Compared with Wasserstein DRO, our proposed approach offers enhanced computational tractability for a broader class of loss functions, and the worst-case distribution exhibits greater plausibility in practical scenarios. To solve the dual reformulation, we develop a stochastic mirror descent algorithm with biased gradient oracles. Remarkably, this algorithm achieves near-optimal sample complexity for both smooth and nonsmooth loss functions, nearly matching the sample complexity of the Empirical Risk Minimization counterpart. Finally, we provide numerical examples using synthetic and real data to demonstrate its superior performance.