🤖 AI Summary
This work addresses the challenge of optimizing root-node risk-sensitive objectives in finite-horizon Markov decision processes when using rank-dependent utility functionals, which violate Bellman optimality and are notoriously difficult to optimize. The authors propose the ERQDP method, which constructs a rank-quantile surrogate model to enable exact dynamic programming over a discretized return grid—without requiring scenario enumeration or sampling—and integrates an anytime optimization loop to provide rigorous upper and lower bounds on the objective. This approach yields, for the first time, certifiably exact solutions for both risk-averse and risk-seeking behaviors and supports policy evaluation with suboptimality gap guarantees. Experiments demonstrate that ERQDP efficiently computes certified solutions across multiple benchmarks, significantly accelerates risk-parameter sweeps, and substantially improves computational efficiency.
📝 Abstract
We study finite-horizon MDP planning under \emph{root-based} (resolute) risk objectives that apply a rank-dependent functional to the distribution of total returns. Such objectives are non-linear in the return distribution and generally break Bellman optimality, so direct optimization by scenario-tree enumeration is intractable. We propose \textbf{ERQDP}, an enumeration-free and sampling-free method that solves a rank--quantile surrogate via exact DP (Dynamic Programming), evaluates candidate policies exactly by DP over return Probability Mass Functions (PMFs) on a discretized return grid (with an explicit rounding bound), and refines the surrogate in an anytime loop that reports an explicit upper--lower gap (certificate) for the target objective up to discretization budgets. Across tested benchmarks, ERQDP returns certified solutions or explicit residual gaps, enables fast risk-parameter sweeps with substantial runtime gains, and supports both risk-averse and risk-seeking behaviors.