distributionally robust listwise optimization

Designs, implements, and analyzes distributionally robust training objectives and algorithms for listwise ranking models—particularly robust Plackett–Luce likelihoods—by formulating and solving worst-case (e.g., total-variation) perturbation problems; this includes deriving DRO formulations and decompositions of the robust loss into nominal plus correction terms and producing procedures that preserve ranking performance under clean labels and label noise.

distributionallyrobustlistwiseoptimization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the uncertainty in listwise preference ranking for language model alignment—arising from inconsistent annotations, near-ties, or reward noise—by proposing a robust Plackett–Luce objective grounded in total variation distributionally robust optimization. The method directly optimizes ranking labels given a prompt and a candidate list, and reduces the worst-case enumeration complexity of rankings from \(K!\) to \(O(K \log K)\), thereby providing the first theoretical guarantees for both offline and online listwise preference learning. Experiments demonstrate that the approach preserves performance on clean labels in offline settings while significantly improving robustness to noise; in online settings, it enhances the reliability of candidate expansion. Consistent gains are observed across both reward models and GPT-4-based evaluation metrics.

distributionally robust optimizationlanguage model alignmentlistwise preference

Learning Optimal Classification Trees Robust to Distribution Shifts

Oct 26, 2023
NJ
Nathan Justin
🏛️ University of Southern California

In high-stakes domains such as public health, distributional shifts between training and test data—arising from questionnaire design variations, heterogeneous data collection environments, and differing respondent trust levels—severely degrade classifier reliability. Method: This paper proposes the first robust optimal classification tree learning framework. It formulates robust tree learning as a single-stage nonlinear mixed-integer robust optimization problem and equivalently recasts it as a tractable two-stage linear robust optimization model. A customized constraint-generation algorithm is then developed to solve it efficiently. Results: Evaluated on multiple public datasets, the method improves worst-case accuracy by up to 12.48% and average accuracy by up to 4.85% over non-robust optimal trees. Crucially, it establishes, for the first time, provable robustness guarantees for optimal decision trees under distribution shift while maintaining computational tractability—unifying theoretical robustness and practical solvability.

Addressing data sensitivity in high-stakes settings like public healthLearning classification trees robust to training-testing distribution shiftsProposing mixed-integer optimization for optimal robust tree learning

Existing robust learning methods are often designed in isolation and struggle to handle unknown dominant failure modes, such as distribution shifts or label noise. This work proposes the first unified framework for robust learning, decomposing approaches into four modular stages: reference distribution augmentation, input perturbation, label perturbation, and sample aggregation—each configurable with pessimistic, neutral, or optimistic strategies. The framework establishes a joint design space encompassing diverse robustness techniques, including distributionally robust optimization, label smoothing, neighborhood risk minimization, and Mixup, enabling end-to-end hyperparameter optimization to adaptively compose optimal strategies. Experiments demonstrate that this approach matches or surpasses the best specialized method across tabular data, image classification, and reward modeling benchmarks, offering a general and reliable default solution for unseen failure modes.

distribution shiftempirical risk minimisationfinite-sample degeneracies

Addressing the fundamental trade-off between robustness and accuracy in high-dimensional linear regression, this paper proposes an automated radius selection method grounded in Wasserstein distributionally robust optimization. For the first time under the high-dimensional asymptotic regime, we derive a convex–concave error characterization formula involving only four scalar variables, precisely capturing how estimation error varies with the robustness radius. This analytical expression eliminates the need for computationally expensive cross-validation, enabling theory-driven hyperparameter tuning. The theoretically predicted error closely matches empirical results, and the optimal radius selected by our method aligns with cross-validation outcomes—while reducing computational cost by over two orders of magnitude. Our core contribution is the establishment of an analytically tractable and numerically computable explicit model for robust estimation error, yielding the first automatic tuning framework for high-dimensional robust regression that simultaneously satisfies theoretical rigor and engineering practicality.

Characterizing estimation error via convex-concave optimizationEfficient hyperparameter tuning without cross-validation costOptimal robustness radius selection in high-dimensional DRO

Drago: Primal-Dual Coupled Variance Reduction for Faster Distributionally Robust Optimization

Mar 16, 2024
RM
Ronak Mehta
🏛️ University of Washington | University of Wisconsin

This paper studies penalty-based distributionally robust optimization (DRO) with a closed convex uncertainty set, encompassing canonical settings such as $f$-DRO and spectral/$L$-risk minimization. Exploiting the problem’s strongly convex–strongly concave structure, we propose a cyclic–stochastic hybrid sampling scheme, coupled with regularized primal updates and dual variance reduction. This yields the first linearly convergent algorithm whose convergence rate depends *finely* on both primal and dual condition numbers. Theoretical analysis establishes that our method achieves the current state-of-the-art linear convergence rate. Numerical experiments on regression and classification tasks demonstrate significant improvements over existing baseline methods. Our core contributions lie in the synergistic integration of hybrid sampling design, variance reduction, and condition-number-sensitive analysis—establishing a new paradigm for high-accuracy, high-efficiency DRO optimization.

Faster distributionally robust optimizationLinear convergence for convex-concave problemsPrimal-dual variance reduction algorithm

Latest Papers

What's happening recently
View more

This work addresses the sensitivity of Bayesian optimization to model misspecification, which can lead to fragile out-of-sample decisions. The authors propose a distributionally robust optimization framework that formalizes and quantifies model robustness under perturbations to both parameters and likelihood through novel measures: posterior sensitivity and likelihood sensitivity. Theoretical analysis reveals that posterior sensitivity vanishes as variance decreases, whereas likelihood sensitivity persists; parameter learning mitigates the former but cannot eliminate the latter. By constructing an uncertainty set based on a bias-aware divergence measure, the method achieves a near-Pareto-optimal trade-off between expected performance and dual robustness. Empirical experiments validate the effectiveness of the proposed approach.

Bayesian distributionally robust optimizationlikelihood sensitivityposterior sensitivity

This work addresses the lack of finite-sample theoretical guarantees and systematic comparisons for existing robust learning methods under distribution shift between training and deployment environments. Focusing on Distributionally Robust Optimization (DRO) and Robust Satisficing (RS), the paper establishes, for the first time, dimension-free finite-sample generalization error bounds for the target domain and introduces an information-guided hyperparameter calibration strategy that leverages partial knowledge of the distributional shift. Theoretical analysis reveals a complementary relationship between DRO and RS under partial shift information, while empirical studies in inventory network planning demonstrate their distinct response mechanisms to positively shifted demand, thereby providing principled guidance for method selection in practice.

distributional shiftsfinite-sample guaranteesgeneralization error

This work addresses the vulnerability of probabilistic circuits to overfitting and poor generalization under data noise, limited samples, or distribution shifts. To mitigate this, the authors propose PeTeR, the first data-free post-training framework that enhances the robustness of pretrained probabilistic circuits to distributional shifts without requiring retraining. PeTeR leverages distributionally robust optimization by modeling worst-case distributions within a Wasserstein ball and introduces a data-agnostic parameter adjustment mechanism grounded in this principle. Empirical evaluations across multiple density estimation benchmarks demonstrate that PeTeR significantly improves model robustness against both random and adversarial perturbations, matching or outperforming existing data-dependent robust learning baselines.

distribution shiftgeneralizationoverfitting

This work addresses the grouped distributionally robust (GDR) least squares problem, which seeks to minimize the worst-case loss across multiple data groups. The authors introduce block Lewis weights—a novel geometric tool—to reformulate the problem as a specially weighted least squares instance. By integrating an accelerated proximal algorithm with a structured linear system solver tailored for systems of the form \(A^\top B A\), they achieve an efficient solution method that unifies optimization frameworks for both average and robust losses. The proposed approach outperforms interior-point methods at moderate accuracy levels. Theoretically, it attains a \((1+\varepsilon)\)-approximate solution using only \(\widetilde{O}(\min\{\mathrm{rank}(A), m\}^{1/3} \varepsilon^{-2/3})\) linear system solves, yielding the current best-known guarantee for the special case of \(\ell_\infty\) regression.

block Lewis weightsdistributionally robust optimizationgroup robustness

Hot Scholars

JZ

Jizhi Zhang

USTC
RecommendationTrustworthy AILarge Personalized Model
MH

Matthias Hagen

Friedrich-Schiller-Universität Jena
Information RetrievalNatural Language Processing
MP

Martin Potthast

University of Kassel, hessian.AI, and ScaDS.AI
Information RetrievalNatural Language Processing
YZ

Yi Zhang

Huawei Co., Ltd
CVAITrustworthy AI