Score
Designs and implements deterministic decision sequences by reducing randomized strategies to static optimization problems and constructing optimized finite supports that replace randomization with single or small sets of deterministic actions. This includes formulating static-reduction and support-extension models, extracting implementable de‑randomized policies from a static solver, and building efficient algorithms to compute and represent those deterministic sequences while preserving the randomized policy’s performance guarantees.
This work addresses the problem of efficiently and exactly solving discounted Markov decision processes (DMDPs) for the optimal value function and policy. We propose a novel reduction framework that decomposes the exact solution into two subproblems: policy evaluation and computation of an approximately optimal value function. Leveraging state-of-the-art techniques in approximate dynamic programming, we design both deterministic and randomized algorithms tailored to these subtasks. Our approach achieves significantly improved computational efficiency, yielding the fastest known exact DMDP solver to date. The resulting algorithms demonstrate clear advantages over existing methods, both theoretically—through tighter complexity bounds—and empirically—via superior practical performance.
This work investigates whether computationally efficient reinforcement learning algorithms exist for Markov decision processes (MDPs) with deterministic dynamics, large action spaces, stochastic initial states, and stochastic rewards, under the linear Bellman completeness framework. To address error amplification—a key challenge in value estimation—we propose the first computationally efficient (polynomial-time) optimistic value iteration algorithm: it injects structured random noise *only* into the null space of the training data during least-squares regression, yielding strictly optimistic value estimates without excessive conservatism. Our method integrates linear function approximation with optimistic value iteration. Theoretically, it achieves a regret bound of $ ilde{O}(sqrt{d^3 H^3 T})$, where $d$ is the feature dimension, $H$ the horizon, and $T$ the total time steps. This bound unifies classical settings—including linear MDPs and linear quadratic regulators (LQR)—and breaks computational bottlenecks in large-action-space regimes under statistical learnability assumptions.
Discrete partially observable Markov decision processes (POMDPs) lack deterministic performance guarantees for online planning solutions. Method: This paper proposes an online planning framework that, for the first time, establishes a deterministic error bound between any time-bounded approximate solution and the optimal value function. The approach integrates POMDP modeling, online Monte Carlo tree search (MCTS), upper-confidence-bound propagation, and rigorous error bound derivation—designed as a plug-and-play enhancement to existing MCTS-based planners. Contribution/Results: Evaluated on standard benchmarks, the method achieves significantly improved solution quality with negligible increase in computational overhead, while providing verifiable theoretical guarantees. It is the first POMDP planning framework to simultaneously ensure real-time execution, deterministic error bounds, and plug-and-play compatibility with state-of-the-art tree search algorithms.
This work addresses the challenge of computing deterministic optimal policies for constrained Markov decision processes (CMDPs) with continuous state-action spaces. Existing policy gradient methods struggle in this setting due to their reliance on stochastic policies and discrete action enumeration, rendering them ill-suited for continuous domains and hard constraints. To overcome this bottleneck, we propose the first provably convergent deterministic policy gradient primal-dual algorithm (D-PGPD). D-PGPD establishes a non-asymptotic, function-approximation-compatible framework for deterministic policy search and introduces a quadratically regularized Lagrangian primal-dual update to handle hard constraints in continuous spaces. We prove that D-PGPD converges sublinearly to the regularized optimal primal-dual solution and provide explicit bounds on approximation error induced by function approximation. Empirical evaluation on robotic navigation and fluid control tasks demonstrates substantial improvements over state-of-the-art baselines.
This paper studies monotone submodular maximization under matroid constraints. Addressing a long-standing bottleneck in the approximation ratio of deterministic algorithms—previously capped at 0.5008—it introduces the first deterministic non-blind local search algorithm achieving an approximation ratio of $1 - 1/e - varepsilon$. This bridges the theoretical gap between deterministic and randomized algorithms. The method fully exploits matroid structure to attain nearly linear query complexity $ ilde{O}_varepsilon(nr)$. By incorporating lightweight randomization, the complexity improves to $ ilde{O}_varepsilon(n + rsqrt{n})$. Notably, this is the first deterministic framework—retaining full determinism in its core design—to achieve the $1 - 1/e - varepsilon$ guarantee, significantly surpassing all prior deterministic approaches. The result advances the state-of-the-art both in approximation quality and computational efficiency for constrained submodular optimization.
This work addresses the problem of multi-objective expected reward optimization in infinite-state Markov decision processes, aiming to synthesize policies that approximate the Pareto front. To this end, it introduces the first deductive program-level reasoning framework that integrates multi-objective optimization with weak expectation semantics. The approach features a novel multi-objective expectation transformer and employs a convex hull power domain to symbolically represent post-expectation tuples. By combining hybrid determinization rules for policy synthesis with operational semantics modeling, the method enables symbolic policy synthesis over infinite state spaces. Experimental evaluation demonstrates its effectiveness in solving multi-objective optimization problems across several case studies.
To address the scalability bottleneck in MDP policy synthesis under LTL objectives—caused by state-space explosion during generalized Rabin automaton (GFM) construction—this paper proposes a novel state-space reduction framework. Methodologically, it introduces (1) a game-theoretically optimal “good-for-games minimization” technique, integrating formal translation with specialized GFM construction, and (2) for the key LTL fragment $mathsf{G}mathsf{F}varphi$, a direct GFM construction algorithm achieving single-exponential time complexity, breaking the classical double-exponential barrier. Experimental evaluation on standard benchmarks demonstrates that the approach reduces automaton size by one to two orders of magnitude, significantly improving policy synthesis efficiency and overall scalability. These advances provide a practical pathway for large-scale LTL-constrained MDP planning.
This work addresses the minimax regret optimization problem in Markov decision processes (MDPs) with model uncertainty under a strict constraint on the number of deployable policies. It formally introduces, for the first time, the k-adaptable policy synthesis framework: at most k policies are precomputed before uncertainty is revealed, and during execution, the best among them is selected to minimize worst-case regret. The problem is shown to be NP-hard, prompting the development of KAPS, an exact algorithm that jointly optimizes MDP clustering and policy selection via nested branch-and-bound, enhanced with problem-specific upper and lower bounds and heuristic strategies for computational efficiency. Experiments demonstrate that increasing the policy budget from one to two yields substantial regret reduction; furthermore, under the single-policy setting, KAPS consistently matches or outperforms existing methods in solution quality and more frequently certifies optimality.
This work addresses the maximization of non-monotone submodular functions under both matroid and knapsack constraints. The authors propose a deterministic approximation algorithm based on an extended multilinear extension, which integrates deterministic continuous optimization with efficient discretization techniques. This approach achieves improved approximation ratios for both constraint types simultaneously while maintaining polynomial query complexity—the first such result in the deterministic setting. Specifically, the algorithm attains a (0.385 − ε)-approximation under matroid constraints and a (0.367 − ε)-approximation under knapsack constraints, surpassing the previous best deterministic guarantees of 0.367 and 0.25, respectively, and establishing the current state-of-the-art deterministic performance for these problems.
This work addresses the challenge of verifying Markov decision processes (MDPs) with unknown transition probabilities that exhibit both nondeterminism and probabilistic uncertainty. To tackle this problem, the authors propose an online statistical model checking method grounded in confidence sequences. By integrating dynamic sampling, statistical hypothesis testing, and a novel online confidence sequence construction, the approach effectively mitigates the conservativeness and inefficiency inherent in traditional union-bound techniques. The resulting specialized verification tool maintains rigorous reliability guarantees while achieving a dramatic reduction in sample complexity—requiring on average approximately 50 times fewer samples than the current state-of-the-art methods—thereby substantially enhancing both efficiency and practical applicability.