mdp formulation

Design and build Markov Decision Process formulations by specifying state and action spaces, reward functions, transition dynamics, and encoding constraints or objectives into the model. Synthesize policies for these MDPs and perform theoretical analysis of policy optimality, robustness, feasibility under constraints, and related performance guarantees.

mdpformulation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.47
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

MDP modeling for multi-stage stochastic programs

Sep 26, 2025
DP
David P. Morton
🏛️ Northwestern University | Dowson Farms | SKEMA Business School | Université Côte d’Azur

This paper addresses a class of multistage stochastic programming problems characterized by continuous state and action spaces, decision-dependent uncertainty, and limited statistical learning capability. To overcome the expressive limitations of conventional models, we propose an extended policy graph framework that explicitly captures the feedback effect of decisions on uncertainty and incorporates online learning mechanisms. Building upon this, we design a novel stochastic dual dynamic programming (SDDP) algorithm and its nonconvex approximation variant, tailored for efficiently solving such structured Markov decision processes. Experimental results on a suite of benchmark instances—increasing in complexity—demonstrate that our approach significantly improves policy quality and computational scalability. The work establishes a new paradigm for stochastic optimization that jointly integrates statistical learning with sequential decision-making, offering both enhanced expressiveness and tractability.

Develops new stochastic dual dynamic programming variantsExtends MDP modeling for multi-stage stochastic programsIncorporates decision-dependent uncertainty in transition probabilities

Constrained and Robust Policy Synthesis with Satisfiability-Modulo-Probabilistic-Model-Checking

Nov 11, 2025
LH
Linus Heck
🏛️ Radboud University | Brno University of Technology

This work addresses the synthesis of reward-optimal policies for finite Markov decision processes (MDPs) that simultaneously satisfy structural constraints—such as policy conciseness or implementation cost—and exhibit robustness against model perturbations. We propose the first robust policy synthesis framework supporting arbitrary first-order logic–expressible structural constraints. Our approach tightly integrates Satisfiability Modulo Theories (SMT) solving with probabilistic model checking to jointly optimize constraint satisfaction and robust performance. Crucially, it retains computational tractability while significantly enhancing the reliability of synthesized policies in practical deployment. Experimental evaluation on hundreds of benchmark instances demonstrates competitive performance across diverse constraint classes, substantially advancing both the solvability and practical applicability of structured robust policy synthesis.

Computing robust policies for perturbed Markov decision processesIntegrating satisfiability solvers with probabilistic model checking algorithmsSatisfying structural constraints on policy representation and cost

Efficient Strategy Synthesis for MDPs via Hierarchical Block Decomposition

Jun 21, 2025
AE
Alexandros Evangelidis
🏛️ University of York

To address the poor scalability of conventional policy synthesis methods for large-scale Markov decision processes (MDPs), this paper proposes a vulnerability-driven hierarchical block decomposition approach. The method iteratively refines the model dynamically and selects regions based on uncertainty awareness, focusing computational effort exclusively on the currently most vulnerable state subsets for fine-grained modeling and optimization—thereby jointly improving accuracy and efficiency. Its core innovation lies in recasting policy synthesis as an incremental refinement process targeted at critical uncertain regions, circumventing prohibitively expensive global computations. Experiments on MDP benchmarks with over one million states demonstrate that our approach achieves up to a 2× speedup over the state-of-the-art tool PRISM, significantly enhancing the feasibility and practicality of policy synthesis for large-scale systems.

Balancing accuracy and speed in iterative refinementImproving efficiency via dynamic MDP refinementScaling policy synthesis for large MDP state spaces

Tackling Decision Processes with Non-Cumulative Objectives using Reinforcement Learning

May 22, 2024
MN
Maximilian Nägele
🏛️ Max Planck Institute for the Science of Light | Friedrich-Alexander-Universität Erlangen-Nürnberg

This paper addresses non-cumulative Markov decision processes (NCMDPs), where the objective is to optimize the expectation of an arbitrary function—e.g., maximum reward, Sharpe ratio—of the reward sequence, rather than the conventional discounted cumulative reward. We propose the first general, theoretically rigorous state-augmentation mapping that equivalently transforms any NCMDP into a standard MDP. This reduction enables direct application of classical reinforcement learning algorithms (e.g., DQN, policy gradients) and dynamic programming methods. Empirical evaluation across diverse domains—including control, finance (portfolio optimization), and combinatorial optimization—demonstrates substantial improvements in final performance and training efficiency. Our core contribution is the establishment of a formal theoretical equivalence between NCMDPs and standard MDPs, accompanied by a scalable algorithmic framework for practical implementation. The approach unifies treatment of non-cumulative objectives within the standard RL paradigm while preserving computational tractability and theoretical soundness.

Enabling RL techniques to optimize arbitrary reward functions in NCMDPsImproving performance and training efficiency in diverse NCMDP applicationsMapping non-cumulative MDPs to standard MDPs for broader applicability

A safe exploration approach to constrained Markov decision processes

Dec 01, 2023
TN
Tingting Ni
🏛️ SYCAMORE | EPFL

This paper addresses policy optimization for discounted infinite-horizon constrained Markov decision processes (CMDPs) in online learning for safety-critical systems, aiming to maximize expected cumulative reward while *strictly satisfying* cumulative constraints throughout the entire learning process. We propose the first model-free and simulation-free interior-point method framework, which guarantees policy feasibility at *every training iteration*—not merely asymptotically. Our approach constructs an interior-point regularized objective using a log-barrier function and integrates it with policy gradient updates under a Fisher non-degeneracy assumption on policy parameterization. Theoretically, we establish a sample complexity of $ ilde{mathcal{O}}(varepsilon^{-6})$ for converging to an $varepsilon$-optimal feasible policy. This incurs only an $mathcal{O}(varepsilon^{-2})$ overhead relative to the unconstrained C-NPG-PDA algorithm, significantly improving both learning efficiency and reliability under safety constraints.

Develop model-free, simulator-free algorithm for safety-critical systems.Ensure constraint satisfaction during learning process.Maximize reward in constrained Markov decision processes.

Latest Papers

What's happening recently
View more

This work addresses the problem of policy synthesis in Markov decision processes (MDPs) under entropy-based constraints that enforce concentration of state visitation distributions. It formalizes entropy maximization as a policy synthesis objective for the first time, establishes its computational complexity, and introduces a novel method combining convex duality theory with invariant synthesis to handle nonlinear entropy constraints in a conditionally complete manner. By systematically analyzing the roles of memory and randomization in policies, the approach effectively synthesizes and verifies entropy-constrained policies across multiple benchmark instances, substantially extending the expressiveness and applicability of existing policy synthesis frameworks.

concentration propertycontrol policy synthesisentropy objectives

This work addresses the lack of theoretical foundations and efficient algorithms for policy optimization in Markov decision processes (MDPs) with unbounded costs and general state and action spaces. By formulating the MDP as an optimization problem over linear operators in a function space, the paper systematically introduces perturbation theory from functional analysis to derive gradients of the objective function, thereby establishing a policy gradient framework applicable to general MDPs. Building on this foundation, the authors propose a low-complexity proximal policy optimization (PPO)-style algorithm that overcomes the limitations of prior methods, which are typically confined to finite spaces or specific function approximators. This approach successfully extends classical reinforcement learning theory to general MDPs and enables efficient policy optimization in continuous or large-scale state-action spaces.

general state and action spacesMarkov decision processesoperator-theoretic foundations

This work investigates how to achieve optimal policies in Markov decision processes (MDPs) that incorporate future information—such as reference trajectories or predictions—by leveraging model predictive control (MPC). The authors formulate MPC as a class of parameterized policies and train them end-to-end via reinforcement learning. Their key contribution lies in establishing, for the first time, the precise structural conditions under which MPC can exactly represent the optimal value function and policy, thereby providing a theoretical foundation for MPC as a structured function approximator with formal guarantees. Empirical validation on a point-mass racing task with future reference trajectories demonstrates that the proposed approach learns policies approaching optimality, confirming its effectiveness.

future informationMarkov Decision ProcessesModel Predictive Control

This work addresses the minimax regret optimization problem in Markov decision processes (MDPs) with model uncertainty under a strict constraint on the number of deployable policies. It formally introduces, for the first time, the k-adaptable policy synthesis framework: at most k policies are precomputed before uncertainty is revealed, and during execution, the best among them is selected to minimize worst-case regret. The problem is shown to be NP-hard, prompting the development of KAPS, an exact algorithm that jointly optimizes MDP clustering and policy selection via nested branch-and-bound, enhanced with problem-specific upper and lower bounds and heuristic strategies for computational efficiency. Experiments demonstrate that increasing the policy budget from one to two yields substantial regret reduction; furthermore, under the single-policy setting, KAPS consistently matches or outperforms existing methods in solution quality and more frequently certifies optimality.

k-adaptable PoliciesMinimax RegretPolicy Selection

Reinforcement learning policies often fail in simulation-to-reality (sim-to-real) transfer due to modeling inaccuracies. This work systematically investigates how key components of the Markov Decision Process (MDP)—including state representation, objectives, reward function, termination conditions, and dynamics model—affect transfer performance in industrial control tasks. It reveals, for the first time, the underlying mechanisms by which these design choices critically influence sim-to-real success and formulates actionable MDP design principles. Through physically accurate modeling and ablation studies on a color-mixing task, the proposed approach achieves up to 50% success rate in the real world, whereas simplified models completely fail, thereby demonstrating that rigorous MDP design is essential for effective sim-to-real transfer.

industrial process controlMarkov Decision Processpolicy transfer

Hot Scholars

SJ

Sebastian Junges

Assistant Professor, Radboud University, Nijmegen
Formal methodsMarkov Decision ProcessesController SynthesisProbabilistic Inference
AG

Arnob Ghosh

Assistant Professor of ECE at New Jersey Institute of Technology
Reinforcement LearningGame thoeryIntelligent Transportation SystemComputer Networks
KC

Krishnendu Chatterjee

Professor, IST Austria
Game theoryLogic and automata theoryAlgorithmsEvolutionary Game theory
AA

Alessandro Abate

Professor of Verification and Control, University of Oxford, UK
Formal VerificationControl TheoryStochastic Hybrid SystemsCyber-Physical Systems
GQ

Guannan Qu

Carnegie Mellon University
Machine LearningGenerative AIReinforcement LearningControl Theory