Brenier Meets Adversarial Training: Optimal Transport Geometry for Robust Learning

πŸ“… 2026-09-25
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of evaluating worst-case risk in distributionally robust optimization (DRO) under non-convex losses by reconstructing the adversarial training framework through the lens of optimal transport geometry. It reveals that standard approaches incur computational waste by violating cyclic monotonicity, and accordingly proposes a multi-start particle ascent algorithm alongside an input convex neural network-based parameterization for adversarial mappings. Experimental results demonstrate that the proposed method significantly enhances model robustness and generalization under distribution shifts across regression, classification, and control tasks.
πŸ“ Abstract
Distributionally robust optimization (DRO) provides a principled framework for learning under distribution shift, but its practical use is hindered by the difficulty of evaluating worst-case risks for nonconvex loss functions. We study a penalized DRO formulation in which the adversary may choose any distribution but incurs a Wasserstein penalty for deviating from the empirical distribution. We show that the adversary's problem can be reformulated as an optimization problem over transport maps that push empirical samples to adversarial ones, and we prove that optimal maps are cyclically monotone. We also show that standard adversarial training---based on per-sample local optimization---violates cyclical monotonicity and wastes transport costs unless the adversary is severely restricted. We propose two remedies. First, we introduce multi-start particle ascent, which alternates parallel gradient ascent with reassignment to enforce cyclical monotonicity across samples. Second, we parameterize adversarial maps as gradients of input-convex neural networks, which guarantees cyclical monotonicity by construction. Experiments on robust regression, image classification, and robust control show that our methods consistently outperform standard adversarial training and state-of-the-art baselines, achieving improved robustness and better generalization under distribution shift.
Problem

Research questions and friction points this paper is trying to address.

Distributionally Robust Optimization
Adversarial Training
Optimal Transport
Wasserstein Penalty
Cyclical Monotonicity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Distributionally Robust Optimization
Optimal Transport
Cyclical Monotonicity
Adversarial Training
Input-Convex Neural Networks
πŸ”Ž Similar Papers
No similar papers found.