curriculum distributionally robust optimization

Designs and implements training algorithms and objective schedules that progressively impose distributionally robust optimization (DRO) criteria—e.g., gradually increasing emphasis on worst-case or rare subpopulations—so models are optimized for worst-group performance while stabilizing learning under stricter data splits. Analyzes convergence, sensitivity to rare subpopulations, and the trade-offs between robustness and generalization when applying curriculum-style schedules to DRO.

curriculumdistributionallyrobustoptimization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.35
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Near-Optimal Algorithms for Group Distributionally Robust Optimization and Beyond

Dec 28, 2022
TS
Tasuku Soma
🏛️ Institute of Statistical Mathematics | MIT

To address performance degradation and algorithmic inefficiency in fairness-sensitive tasks—such as group distributionally robust optimization (Group DRO)—caused by subgroup distributional shifts, this paper proposes a novel stochastic optimization algorithm. Methodologically, it unifies Group DRO, subgroup fairness, and empirical conditional value-at-risk (CVaR) optimization via a framework integrating stochastic gradient updates, dual-variable coupling, and information-theoretic entropy regularization. Theoretically, it achieves the first near-optimal convergence rate for Group DRO and establishes a tight information-theoretic lower bound, rigorously proving algorithmic optimality. Empirically, the algorithm demonstrates faster convergence and superior robustness across multiple DRO benchmarks, consistently outperforming state-of-the-art methods. It thus bridges theoretical rigor with practical efficacy, offering both provable guarantees and strong empirical performance.

Efficient AlgorithmsFairnessOptimization in Learning Algorithms

Distributionally Robust Optimization

Nov 04, 2024
DK
Daniel Kuhn
🏛️ École Polytechnique Fédérale de Lausanne | Cornell University | Imperial College London

This paper addresses distributionally robust optimization (DRO): seeking decisions that remain optimal under the worst-case distribution within an ambiguity set—defined either by Wasserstein distance or φ-divergence—when the true data-generating distribution is unknown. Methodologically, it establishes, for the first time, systematic equivalences between DRO and key machine learning paradigms, including regularization and adversarial training, thereby unifying statistical learning, operations research, and control theory into a coherent theoretical framework. The approach integrates ambiguity set construction, min-max expected loss optimization, duality analysis, and rigorous robustness verification, balancing theoretical interpretability with computational tractability. The resulting methodology significantly enhances model generalization and decision robustness under distributional shifts. It has been successfully deployed in high-stakes domains including financial risk management, medical diagnosis, and AI safety.

Connects to regularization and adversarial training in MLFocuses on worst-case performance within ambiguity setsStudies decision-making under uncertain probability distributions

This work addresses the lack of finite-sample theoretical guarantees and systematic comparisons for existing robust learning methods under distribution shift between training and deployment environments. Focusing on Distributionally Robust Optimization (DRO) and Robust Satisficing (RS), the paper establishes, for the first time, dimension-free finite-sample generalization error bounds for the target domain and introduces an information-guided hyperparameter calibration strategy that leverages partial knowledge of the distributional shift. Theoretical analysis reveals a complementary relationship between DRO and RS under partial shift information, while empirical studies in inventory network planning demonstrate their distinct response mechanisms to positively shifted demand, thereby providing principled guidance for method selection in practice.

distributional shiftsfinite-sample guaranteesgeneralization error

Distributionally Robust Self Paced Curriculum Reinforcement Learning

Nov 07, 2025
AS
Anirudh Satheesh
🏛️ University of Maryland, College Park | Purdue University

Reinforcement learning (RL) policies often fail when deployed in real-world environments due to train-test distributional shift. Conventional approaches fix the robustness budget ε, yet this static choice inherently trades off nominal performance against robustness: overly small ε yields insufficient robustness, while excessively large ε induces over-conservatism or instability. Method: We propose an adaptive robustness budget curriculum learning framework, modeling ε as a continuous, learnable curriculum variable. The uncertainty set is dynamically expanded during training to progressively increase robustness requirements. Our method integrates distributionally robust optimization, self-paced learning, and adversarial worst-case training. Contribution/Results: Experiments across diverse tasks show an average 11.8% improvement in episode return—reaching 1.9× that of baseline algorithms. The approach significantly alleviates the robustness–performance trade-off and, for the first time, enables end-to-end curriculum scheduling of the robustness budget.

Addresses reinforcement learning policy failure under environmental distribution shiftsBalances nominal performance and robustness against environmental perturbationsOvercomes fixed robustness budget limitations via adaptive curriculum scheduling

Federated Distributionally Robust Optimization with Non-Convex Objectives: Algorithm and Analysis

Jul 25, 2023
YJ
Yang Jiao
🏛️ Tongji University | University of Connecticut

To address three key challenges in federated distributionally robust optimization (FDRO) for non-convex settings—difficulty in achieving convergence under asynchronous updates, insufficient exploitation of prior distributional knowledge, and lack of adaptive control over robustness levels—this paper proposes ASPIRE-EASE. The algorithm integrates asynchronous single-loop optimization, alternating gradient projection, and the iterative active-set method (EASE), coupled with a constraint-based D-norm uncertainty set. ASPIRE-EASE establishes the first theoretical convergence guarantee for non-convex FDRO and enables tunable trade-offs between robustness and model performance. Extensive experiments on real-world datasets demonstrate its rapid convergence, strong robustness against data heterogeneity and adversarial attacks, and superior generalization across diverse federated learning scenarios.

Address asynchronous updates in distributed DRO environmentsAdjust robustness degree flexibly across scenariosEffectively leverage prior distribution in optimization

Latest Papers

What's happening recently
View more

This work addresses the challenge of insufficient generalization in few-shot learning caused by distributional shifts. The authors propose a Prototype-Guided Distributionally Robust Optimization (PG-DRO) framework that integrates class-adaptive priors with optimal transport for the first time. By leveraging hierarchical optimal transport, PG-DRO learns structure-aware prototype priors from base-class data and embeds them into a Sinkhorn-based distributionally robust optimization formulation. This enables the dynamic construction of uncertainty sets aligned with transferable semantic structures. Extensive experiments demonstrate that PG-DRO significantly outperforms standard learners and existing DRO methods across multiple few-shot benchmarks, effectively enhancing model robustness and generalization under distributional shift.

distribution shiftsdistributional robustnessfew-shot learning

Traditional robust optimization is often overly conservative due to its exclusive focus on worst-case scenarios, limiting its ability to leverage predictive information for improved scheduling performance. This work proposes the first framework that explicitly incorporates predictions as an independent benchmark in robust scheduling, achieving a principled trade-off between consistency—near-optimality under predicted scenarios—and robustness—guaranteed performance under worst-case uncertainty. By developing a consistency–robustness trade-off mechanism and employing duality theory, upper-envelope reductions, and support-function blocks, the paper systematically analyzes scheduling problems under interval, budgeted, and general uncertainty sets. Smooth $(1+1/\lambda, 1+\lambda)$ trade-offs are established for restricted assignment and related machine models, while the impossibility of constant-factor trade-offs is proven for unrelated machines; constant performance guarantees are provided for identical machines.

consistencyrobust optimizationrobustness

This work addresses the lack of robustness in multi-objective optimization under distributional shifts, as existing methods fail to explicitly model worst-case performance across objectives. The paper introduces, for the first time, a Distributionally Robust Multi-Objective Optimization (DR-MOO) framework that minimizes each objective’s loss under its worst-case perturbed distribution. It proposes the ε-Pareto stationary point as the solution concept and develops a single-loop, double-clipped Multiple Gradient Descent Algorithm (MGDA). By integrating Lagrangian duality reformulation with gradient clipping, the algorithm handles non-convex, unbounded, and biased gradient settings. Theoretically, it achieves a sample complexity of 𝒪(ε⁻⁴), significantly improving upon existing approaches, while experiments demonstrate performance on par with state-of-the-art MGDA baselines.

distributional shiftdistributionally robustmulti-objective optimization

This work addresses the sensitivity of hyperparameters in Learning-to-Optimize (L2O) to perturbations in data distribution by proposing the first hyperparameter learning framework based on Wasserstein distributionally robust optimization (DRO). The approach unifies empirical-performance-driven L2O with worst-case performance estimation (PEP) paradigms. It employs a stochastic gradient algorithm capable of differentiably solving an inner semidefinite program to learn hyperparameters of first-order optimizers over a given set of problem instances. The method provides provable generalization bounds, smoothly interpolating between empirical and worst-case optimality as the sample size grows. Experiments on unconstrained quadratic optimization, LASSO, and linear programming tasks demonstrate that the learned algorithms significantly outperform conventional L2O and worst-case optimal baselines while maintaining certifiable robustness.

distributionally robust optimizationhyperparameter learninglearning to optimize

This work addresses the challenge of robust decision-making in online learning under distributional uncertainty by modeling distributionally robust online learning as a stochastic dynamic game between a decision-maker and a worst-case adversary, where ambiguity is characterized via a Wasserstein ambiguity set. The study establishes the first convergence analysis framework for this setting, proving that the proposed algorithm converges to a robust Nash equilibrium. Furthermore, it reveals an equivalence between worst-case expected optimization and the classical budget allocation problem, enabling the design of an efficient customized algorithm tailored for piecewise-concave loss functions. Theoretical guarantees ensure convergence, while empirical results demonstrate that the method significantly outperforms general-purpose solvers such as Gurobi, effectively alleviating computational bottlenecks in online settings.

computational bottleneckdistributionally robust optimizationonline learning

Hot Scholars

LQ

Lianhui Qin

UC San Diego, Computer Science and Engineering
Natural Language ProcessingMachine Learning
JC

Jixuan Chen

UC San Diego
Multimodal agentsNatural language processingMachine learning
ES

Elias Stengel-Eskin

Assistant Professor, University of Texas at Austin
Natural language processingcomputational semanticscomputational linguistics