Score
Designs and implements optimization models, algorithms, and decision policies that explicitly account for uncertainty in model parameters, data, or outcomes so solutions minimize expected loss, control risk, or limit worst‑case cost. This work includes formulating uncertainty‑aware objectives and constraints, propagating parameter uncertainty, using gradient estimators or other stochastic/robust optimization techniques, and deriving conservative, robust, or adaptive policies that trade off competing costs under parameter uncertainty.
Traditional robust control theory fails in reinforcement learning settings due to gradient uncertainty introduced when approximating value function gradients, undermining stability and performance guarantees. Method: We propose the first robust control framework that explicitly models gradient uncertainty. By formulating a zero-sum dynamic game integrating system dynamics and gradient perturbations, we expose the mismatch mechanism of quadratic value function assumptions under nonzero gradient uncertainty. This yields a generalized Hamilton–Jacobi–Bellman–Isaacs (GU-HJBI) equation featuring non-polynomial correction terms and an associated nonlinear optimal control law. Using viscosity solution theory and uniform ellipticity, we establish a comparison principle; further, we integrate perturbation analysis with an Actor–Critic architecture to design the Gradient-Uncertainty-Robust Actor–Critic (GURAC) algorithm. Results: We theoretically guarantee well-posedness of the GU-HJBI equation, and empirical evaluations demonstrate significantly improved training stability and robustness.
In robust optimization, the robustness level is often set empirically, leading to either insufficient protection or excessive conservatism. While existing data-driven approaches provide finite-sample coverage guarantees, they still require pre-specified coverage targets and lack theoretical guidance for balancing robustness against cost-risk trade-offs. Method: We propose the first distribution-free, finite-sample certifiable robust optimization framework, unifying conformal prediction with robust optimization to construct a Pareto frontier between miscoverage rate and regret upper bound, thereby jointly calibrating robustness and conservatism. Contribution/Results: Our method enables decision-makers to reliably assess and tune robustness levels according to preferences. It significantly improves finite-sample performance on classical optimization problems and delivers interpretable, verifiable robust decision support for high-stakes applications.
This paper identifies an “overfitting” phenomenon in adaptive robust optimization (ARO): when adaptive decisions depend on uncertainty realizations, policies often fail outside the nominal uncertainty set—mirroring generalization failure in machine learning. To address this, we propose a regularization-inspired approach: a hierarchical uncertainty set structure that dynamically allocates set sizes according to constraint importance, thereby providing stronger probabilistic feasibility guarantees for critical constraints. Our method integrates adaptive robust modeling, probabilistic feasibility analysis, and constraint decoupling techniques. Theoretically, we establish a novel analytical framework for mitigating adaptive overfitting. Empirically, the proposed mechanism significantly improves out-of-distribution feasibility and solution stability, achieving a superior trade-off between robustness and adaptivity.
Classical dynamic programming fails for robust infinite-horizon MDPs under non-rectangular uncertainty sets, as it cannot accommodate their coupled, non-separable structure. Method: This paper introduces the first policy-gradient framework for such problems that simultaneously ensures global optimality and computational tractability. We propose a deterministic policy gradient method with provable error bounds quantifying deviation from rectangularity; design a robust Actor-Critic algorithm with controllable convergence rate; and introduce a novel metric characterizing the degree of non-rectangularity of uncertainty sets. Contribution/Results: We establish an $O(1/varepsilon^4)$ iteration complexity bound for convergence to an $varepsilon$-optimal policy, with rigorous global optimality guarantees. Numerical experiments on multiple benchmark tasks demonstrate that our approach significantly outperforms existing state-of-the-art methods.
Directly incorporating machine learning predictions into optimization constraints often yields high constraint violation probabilities. Method: This paper proposes a robust optimization framework that models the model’s loss function as a theoretically certified compact uncertainty set—bypassing conventional distributional assumptions or sampling-based approximations. The framework derives a rigorous upper bound on the constraint violation probability and proves that the resulting uncertainty set radius is up to ten times smaller than those of existing methods. Results: In synthetic experiments, the proposed approach significantly reduces constraint violation probability while compressing the uncertainty set size by an order of magnitude, achieving a favorable trade-off between solution feasibility and conservatism. The core contribution lies in a geometric transformation from the loss function to an uncertainty set, coupled with a probabilistic guarantee mechanism that ensures theoretical soundness and practical efficacy.
This work addresses the challenges of model inaccuracy and ambiguous state distributions in nonlinear systems by proposing a distributionally robust optimization-based chance-constrained control framework. The approach constructs an ambiguity set using relative entropy constraints and derives an upper bound on risk expectations via the variational representation of the exponential integral. It further integrates nonlinear covariance propagation with adaptive determination of the ambiguity set radius based on second-order dynamic truncation error. Notably, the method recovers nominal risk in the zero-divergence limit, thereby overcoming the restrictive assumptions of Gaussianity prevalent in conventional approaches. Validation on spacecraft stochastic guidance tasks demonstrates that the proposed framework effectively enforces probabilistic safety constraints under distributional uncertainty.
Traditional robust optimization is often overly conservative due to its exclusive focus on worst-case scenarios, limiting its ability to leverage predictive information for improved scheduling performance. This work proposes the first framework that explicitly incorporates predictions as an independent benchmark in robust scheduling, achieving a principled trade-off between consistency—near-optimality under predicted scenarios—and robustness—guaranteed performance under worst-case uncertainty. By developing a consistency–robustness trade-off mechanism and employing duality theory, upper-envelope reductions, and support-function blocks, the paper systematically analyzes scheduling problems under interval, budgeted, and general uncertainty sets. Smooth $(1+1/\lambda, 1+\lambda)$ trade-offs are established for restricted assignment and related machine models, while the impossibility of constant-factor trade-offs is proven for unrelated machines; constant performance guarantees are provided for identical machines.
This work addresses the challenge in safety-sensitive stochastic optimization where tail risk exhibits high sensitivity to decision variables, making it difficult for conventional methods to simultaneously optimize average performance and control extreme risks. The authors propose a novel framework that integrates a “safe-start” mechanism with variance reduction techniques. By employing simulation-guided safe initialization, the approach overcomes convergence failures in stochastic gradient descent caused by step-size selection, achieving provably improved sample complexity. Empirical evaluations on portfolio optimization and robust neural network classification demonstrate that the method substantially enhances decision safety under extreme risk scenarios while improving algorithmic efficiency.
This study addresses a central challenge in data-driven optimization: determining the optimal stopping time for data acquisition under parameter uncertainty by balancing sampling costs against information gains. The authors propose a Bayesian learning–based sequential data collection framework that explicitly models the trade-off between information gain and sampling cost, enabling a reward-driven adaptive stopping mechanism that jointly optimizes data acquisition and decision-making. Integrating Bayesian parameter updating, sequential decision theory, and stochastic programming, the approach formulates multiple stopping strategies within the newsvendor model. Numerical experiments demonstrate that the proposed strategies significantly reduce redundant sampling compared to fixed-budget and ex post optimal benchmarks while maintaining near-optimal decision performance.
This work addresses the challenge of selecting uncertainty representations that align with decision objectives to achieve optimal and trustworthy decisions under state-variable uncertainty. Drawing on decision theory, it systematically analyzes the optimal forms of uncertainty representation for both risk-neutral and risk-averse agents in known and unknown environments, revealing the minimal uncertainty information required under distinct risk preferences. The study innovatively unifies three approaches to epistemic uncertainty—calibrated prediction, confidence-set robust optimization, and Bayesian inference—establishing a theoretical link between uncertainty representation and decision goals. This integration yields a reliable decision-making framework that provides agents with verifiable utility guarantees.