Score
Designs and trains decision-making policies that select and schedule interventions (e.g., control, treatment, or mitigation actions) to achieve specified objectives under uncertainty; this includes formulating optimization or reinforcement-learning approaches, building or using simulators and controllers, coping with partial observability and stochastic dynamics, and evaluating trade-offs between effectiveness and costs or constraints.
This paper addresses the challenge of deeply integrating model predictive control (MPC) and reinforcement learning (RL), stemming from their fundamentally divergent model usage paradigms. To resolve this, we propose the first unified taxonomy for MPC–RL fusion, centered on *how models are used*, categorizing approaches into three paradigms: MPC-augmented RL, RL-augmented MPC, and co-designed architectures. Leveraging a unified Actor–Critic modeling framework, we systematically analyze how MPC’s online optimization enhances RL’s closed-loop performance and establish a performance-gain-oriented evaluation perspective grounded in closed-loop metrics. The survey comprehensively covers six application domains—including robotics, energy systems, and autonomous driving—and synthesizes cross-cutting modeling techniques bridging control theory and RL. Our work provides a scalable methodology and principled design guidelines for hybrid intelligent control systems.
This paper addresses optimal policy assignment under partial treatment coverage in heterogeneous populations, where experiments only implement a subset of possible treatment values, limiting generalizability to untested interventions. Method: We propose the first framework integrating shape-constrained partial identification of treatment effects—incorporating monotonicity or convexity constraints—with a minimax regret criterion, formulated as a tractable mixed-integer linear program (MILP). The method combines nonparametric conditional average treatment effect (CATE) estimation, shape restrictions, minimax optimization, and efficient linear/integer programming solvers. Contribution/Results: Applied to a Kenyan rural electricity subsidy experiment, our framework recommends novel, experimentally untested treatment levels—covering nearly the entire population—while reducing maximum regret by over 60%. It substantially enhances policy extrapolation capability and robustness beyond conventional methods constrained to observed treatment supports.
Policy optimization algorithms suffer from poor interpretability and error-prone implementation due to the complexity of Markov decision process (MDP) modeling and inconsistent use of discounted versus average-reward settings. Method: This paper introduces a unified analytical framework that, for the first time, systematically integrates generalized ergodicity theory with perturbation analysis to characterize the steady-state behavior of diverse policy optimization algorithms under both discounted and average-reward criteria. Contribution/Results: The framework clarifies fundamental algorithmic principles, identifies and corrects common implementation pitfalls, and significantly enhances interpretability and robustness. Empirical validation on MDP modeling and linear quadratic regulator (LQR) benchmarks confirms the framework’s ability to capture algorithmic consistency. Quantitative analysis further demonstrates that minor adjustments to key design parameters exert decisive influence on convergence properties and performance.
This paper addresses welfare loss in policy decision-making arising from estimation uncertainty in treatment effect evaluation. We propose the first framework that embeds statistical confidence guarantees directly into policy learning. Methodologically, we construct confidence sets for heterogeneous treatment effects via robust causal inference and integrate them into a risk-constrained optimization formulation to define and solve the “risk-controlled efficient decision frontier,” jointly optimizing policy assignment and budget allocation for both experimental and observational data. Our key contribution is the first verifiable lower bound guarantee for policy selection: at a prespecified confidence level, the actual welfare is guaranteed to be no less than the reported estimated welfare. This framework substantially enhances the reliability, interpretability, and verifiability of policy deployment in high-stakes domains such as healthcare and public policy.
This paper addresses the problem of optimizing initial expenditure allocation across multiple policies under statistical uncertainty to maximize social welfare. Methodologically, it develops an analytically tractable Bayesian risk minimization framework and proposes the first empirical Bayes decision rule for multi-policy comparison—integrating empirical Bayes estimation with statistical decision theory while retaining theoretical optimality even when the prior is unknown. Its contributions are threefold: (1) it derives the first closed-form Bayesian decision rule for multi-policy settings; (2) it provides rigorous theoretical guarantees that the rule strictly dominates conventional sample interpolation methods under small-sample and heteroscedastic noise conditions; and (3) empirical evaluation demonstrates substantial improvements in both accuracy and robustness of policy selection. The framework delivers a computationally feasible, theoretically grounded, and broadly applicable decision paradigm for evidence-based policy design under uncertainty.
Classical dynamic programming fails for robust infinite-horizon MDPs under non-rectangular uncertainty sets, as it cannot accommodate their coupled, non-separable structure. Method: This paper introduces the first policy-gradient framework for such problems that simultaneously ensures global optimality and computational tractability. We propose a deterministic policy gradient method with provable error bounds quantifying deviation from rectangularity; design a robust Actor-Critic algorithm with controllable convergence rate; and introduce a novel metric characterizing the degree of non-rectangularity of uncertainty sets. Contribution/Results: We establish an $O(1/varepsilon^4)$ iteration complexity bound for convergence to an $varepsilon$-optimal policy, with rigorous global optimality guarantees. Numerical experiments on multiple benchmark tasks demonstrate that our approach significantly outperforms existing state-of-the-art methods.
Existing hierarchical decision-making approaches often struggle to simultaneously satisfy constraints and maintain computational efficiency due to misalignment between low-level policies and high-level objectives. This work proposes a principled inverse optimization–based hierarchical framework that, for the first time, systematically constructs structured low-level optimization problems from expert demonstrations, thereby aligning high-level task abstractions with low-level decision-making. By integrating inverse optimization, hierarchical reinforcement learning, and optimal control, the method achieves both interpretability and computational efficiency. Empirical evaluations on resource allocation and obstacle avoidance tasks demonstrate that the approach significantly outperforms end-to-end reinforcement learning, learning-augmented optimal control, and existing hierarchical methods, achieving state-of-the-art performance in both decision quality and computational speed.
This study investigates the robustness of agent policies under uncertain perturbations—such as action corruption caused by actuator failures—specifically examining the magnitude of disturbances under which a policy can still guarantee reachability or safety objectives. For both Markov decision processes and stochastic games, the work formally introduces the notion of policy resilience for the first time and establishes a comprehensive theoretical framework encompassing perturbation modeling, impact aggregation (under expected or worst-case semantics), and quantitative analysis (e.g., via frequency-based metrics). The paper delineates solvability boundaries under different aggregation mechanisms and extends the framework to stochastic games, thereby providing a rigorous foundation for designing highly reliable autonomous systems.
This study addresses the challenges of deploying Actor-Critic algorithms in real-world control systems, where poor reliability and high sensitivity to hyperparameters often hinder practical application. Focusing on a real-world water treatment plant control task, the authors conduct over 33,000 large-scale ablation experiments to systematically evaluate how key algorithmic components—such as policy update schemes, action distribution representations, gradient estimation methods, and update frequencies—affect performance stability and hyperparameter robustness. Their empirical analysis reveals, for the first time, that commonly adopted default configurations (e.g., Gaussian action distributions with pathwise derivatives) exhibit low reliability, whereas bounded action distributions combined with adaptive update strategies substantially enhance robustness. The work identifies high-stability algorithmic configurations that significantly reduce performance variance under limited tuning budgets, offering actionable, component-level design guidelines for industrial deployment.