design decision policies

Design, build, and analyze formal decision-making artifacts—policies, planners, and decision procedures—that select actions under uncertainty and constraints. This includes constructing and evaluating probabilistic and decision-theoretic models (e.g., MDPs), hierarchical, modular, online, and model-based reinforcement-learning policies, and information-theoretic or budgeted probing operators to optimize trade-offs among objectives and resources.

designdecisionpolicies

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.5
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$197K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Efficient Strategy Synthesis for MDPs via Hierarchical Block Decomposition

Jun 21, 2025
AE
Alexandros Evangelidis
🏛️ University of York

To address the poor scalability of conventional policy synthesis methods for large-scale Markov decision processes (MDPs), this paper proposes a vulnerability-driven hierarchical block decomposition approach. The method iteratively refines the model dynamically and selects regions based on uncertainty awareness, focusing computational effort exclusively on the currently most vulnerable state subsets for fine-grained modeling and optimization—thereby jointly improving accuracy and efficiency. Its core innovation lies in recasting policy synthesis as an incremental refinement process targeted at critical uncertain regions, circumventing prohibitively expensive global computations. Experiments on MDP benchmarks with over one million states demonstrate that our approach achieves up to a 2× speedup over the state-of-the-art tool PRISM, significantly enhancing the feasibility and practicality of policy synthesis for large-scale systems.

Balancing accuracy and speed in iterative refinementImproving efficiency via dynamic MDP refinementScaling policy synthesis for large MDP state spaces

MDP modeling for multi-stage stochastic programs

Sep 26, 2025
DP
David P. Morton
🏛️ Northwestern University | Dowson Farms | SKEMA Business School | Université Côte d’Azur

This paper addresses a class of multistage stochastic programming problems characterized by continuous state and action spaces, decision-dependent uncertainty, and limited statistical learning capability. To overcome the expressive limitations of conventional models, we propose an extended policy graph framework that explicitly captures the feedback effect of decisions on uncertainty and incorporates online learning mechanisms. Building upon this, we design a novel stochastic dual dynamic programming (SDDP) algorithm and its nonconvex approximation variant, tailored for efficiently solving such structured Markov decision processes. Experimental results on a suite of benchmark instances—increasing in complexity—demonstrate that our approach significantly improves policy quality and computational scalability. The work establishes a new paradigm for stochastic optimization that jointly integrates statistical learning with sequential decision-making, offering both enhanced expressiveness and tractability.

Develops new stochastic dual dynamic programming variantsExtends MDP modeling for multi-stage stochastic programsIncorporates decision-dependent uncertainty in transition probabilities

Small Decision Trees for MDPs with Deductive Synthesis

Jan 17, 2025
RA
Roman Andriushchenko
🏛️ Brno University of Technology | Radboud University

Optimal policies for Markov decision processes (MDPs) are often large and unintelligible, hindering interpretability and deployment. Method: This paper proposes an SMT-based policy compression framework that synthesizes compact, interpretable decision trees while guaranteeing optimality. It integrates abstraction-refinement with family-based Markov chain verification in a synergistic search scheme, explicitly constraining tree size. Policy synthesis is encoded as a semantic-constrained SMT problem, augmented by policy-space pruning and model-family verification to ensure efficiency and correctness. Contribution/Results: Evaluated on benchmark MDPs with up to 10,000 states and 19-dimensional feature spaces, our method achieves up to 20× reduction in decision tree size while preserving near-optimal performance. The resulting trees significantly outperform those produced by state-of-the-art approaches in both compactness and fidelity to the optimal policy.

Decision Tree SimplificationMarkov Decision ProcessesOptimal Policy Representation

Existing hierarchical decision-making approaches often struggle to simultaneously satisfy constraints and maintain computational efficiency due to misalignment between low-level policies and high-level objectives. This work proposes a principled inverse optimization–based hierarchical framework that, for the first time, systematically constructs structured low-level optimization problems from expert demonstrations, thereby aligning high-level task abstractions with low-level decision-making. By integrating inverse optimization, hierarchical reinforcement learning, and optimal control, the method achieves both interpretability and computational efficiency. Empirical evaluations on resource allocation and obstacle avoidance tasks demonstrate that the approach significantly outperforms end-to-end reinforcement learning, learning-augmented optimal control, and existing hierarchical methods, achieving state-of-the-art performance in both decision quality and computational speed.

Hierarchical Decision MakingInverse OptimizationOptimal Control

Tackling Decision Processes with Non-Cumulative Objectives using Reinforcement Learning

May 22, 2024
MN
Maximilian Nägele
🏛️ Max Planck Institute for the Science of Light | Friedrich-Alexander-Universität Erlangen-Nürnberg

This paper addresses non-cumulative Markov decision processes (NCMDPs), where the objective is to optimize the expectation of an arbitrary function—e.g., maximum reward, Sharpe ratio—of the reward sequence, rather than the conventional discounted cumulative reward. We propose the first general, theoretically rigorous state-augmentation mapping that equivalently transforms any NCMDP into a standard MDP. This reduction enables direct application of classical reinforcement learning algorithms (e.g., DQN, policy gradients) and dynamic programming methods. Empirical evaluation across diverse domains—including control, finance (portfolio optimization), and combinatorial optimization—demonstrates substantial improvements in final performance and training efficiency. Our core contribution is the establishment of a formal theoretical equivalence between NCMDPs and standard MDPs, accompanied by a scalable algorithmic framework for practical implementation. The approach unifies treatment of non-cumulative objectives within the standard RL paradigm while preserving computational tractability and theoretical soundness.

Enabling RL techniques to optimize arbitrary reward functions in NCMDPsImproving performance and training efficiency in diverse NCMDP applicationsMapping non-cumulative MDPs to standard MDPs for broader applicability

Latest Papers

What's happening recently
View more

This work addresses the lack of theoretical foundations and efficient algorithms for policy optimization in Markov decision processes (MDPs) with unbounded costs and general state and action spaces. By formulating the MDP as an optimization problem over linear operators in a function space, the paper systematically introduces perturbation theory from functional analysis to derive gradients of the objective function, thereby establishing a policy gradient framework applicable to general MDPs. Building on this foundation, the authors propose a low-complexity proximal policy optimization (PPO)-style algorithm that overcomes the limitations of prior methods, which are typically confined to finite spaces or specific function approximators. This approach successfully extends classical reinforcement learning theory to general MDPs and enables efficient policy optimization in continuous or large-scale state-action spaces.

general state and action spacesMarkov decision processesoperator-theoretic foundations

This study addresses the challenge of tracing causes and processes when automated decision systems fail, a task inadequately handled by existing compliance governance mechanisms. The authors propose an operational governance evidence framework that integrates structural accountability diagnostics, decision trajectory tracing, sufficiency metrics for evidentiary support, and label-free monitoring. The framework’s applicability is validated across four canonical system architectures. The research uncovers a “governance coverage gradient” phenomenon and introduces an uncertainty cascade model to identify three types of structural discontinuities in agent-based AI systems, along with methods for their analytical extension. By formalizing four propositions that delineate the framework’s boundaries, the work demonstrates full fillability in rule-based engines while simultaneously revealing inherent structural governance gaps in agent-centric AI systems.

agentic AI systemsauditable decisioningevidence fragmentation

This work addresses the problem of multi-objective expected reward optimization in infinite-state Markov decision processes, aiming to synthesize policies that approximate the Pareto front. To this end, it introduces the first deductive program-level reasoning framework that integrates multi-objective optimization with weak expectation semantics. The approach features a novel multi-objective expectation transformer and employs a convex hull power domain to symbolically represent post-expectation tuples. By combining hybrid determinization rules for policy synthesis with operational semantics modeling, the method enables symbolic policy synthesis over infinite state spaces. Experimental evaluation demonstrates its effectiveness in solving multi-objective optimization problems across several case studies.

multiobjective optimizationnondeterminismPareto front

This work addresses the minimax regret optimization problem in Markov decision processes (MDPs) with model uncertainty under a strict constraint on the number of deployable policies. It formally introduces, for the first time, the k-adaptable policy synthesis framework: at most k policies are precomputed before uncertainty is revealed, and during execution, the best among them is selected to minimize worst-case regret. The problem is shown to be NP-hard, prompting the development of KAPS, an exact algorithm that jointly optimizes MDP clustering and policy selection via nested branch-and-bound, enhanced with problem-specific upper and lower bounds and heuristic strategies for computational efficiency. Experiments demonstrate that increasing the policy budget from one to two yields substantial regret reduction; furthermore, under the single-policy setting, KAPS consistently matches or outperforms existing methods in solution quality and more frequently certifies optimality.

k-adaptable PoliciesMinimax RegretPolicy Selection

This work addresses the asymmetry in sequential decision-making where information continually accumulates while the set of feasible actions progressively shrinks due to deadlines, commitments, or resource constraints. To tackle this challenge, the paper introduces the Mature Markov Decision Process (MMDP) framework, which formally characterizes the information–action asymmetry. The framework incorporates stage-aware policies and a “prioritize expiring actions” principle, and integrates search-augmented reinforcement learning with knowledge distillation to develop a structure-aware learning approach. Evaluated on multi-supplier replenishment, cash management, and production-scale simulation environments, the proposed method demonstrates significantly improved learning efficiency, with performance gains amplifying as problem scale increases.

decision urgencyexpiring actionsinformation-action asymmetry

Hot Scholars

ZG

Ziyang Guo

Northwestern Univeristy
Human-AI complementarityDecision making
BW

Bryan Wilder

Assistant Professor of Machine Learning, Carnegie Mellon University
Artificial intelligenceoptimizationmachine learningsocial networks
HH

Hamed Hassani

University of Pennsylvania (EE, CS, Stat) & Google Research
Machine LearningAI SafetyTrustworthy MLInformation Theory
JH

Jessica Hullman

Ginni Rometty Professor of Computer Science, Northwestern University
uncertainty quantificationAI for decision-makingmetasciencevisualization
MW

Michael Wooldridge

University of Oxford
artificial intelligencemulti-agent systemsmultiagent systemsknowledge representation