Maturing Markov Decision Processes: Decision Making under Increasing Information and Shrinking Action Sets

📅 2026-06-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the asymmetry in sequential decision-making where information continually accumulates while the set of feasible actions progressively shrinks due to deadlines, commitments, or resource constraints. To tackle this challenge, the paper introduces the Mature Markov Decision Process (MMDP) framework, which formally characterizes the information–action asymmetry. The framework incorporates stage-aware policies and a “prioritize expiring actions” principle, and integrates search-augmented reinforcement learning with knowledge distillation to develop a structure-aware learning approach. Evaluated on multi-supplier replenishment, cash management, and production-scale simulation environments, the proposed method demonstrates significantly improved learning efficiency, with performance gains amplifying as problem scale increases.
📝 Abstract
Sequential decision problems often exhibit an asymmetric evolution of information and decision flexibility: as a decision cycle unfolds, the agent receives richer information while feasible actions expire due to operational cutoffs, commitments, or resource constraints. Standard MDP formulations typically flatten this structure into stage-dependent state descriptions and action masks, thereby obscuring the nested information--action asymmetry that determines which decisions are urgent and which can be deferred. We introduce Maturing Markov Decision Processes (MMDPs), a formulation built around this information--action asymmetry. We characterize one of its key consequences through an expiring-action priority principle, which identifies the actions that must be resolved before the next stage. Motivated by this structure, we develop a structure-aware reinforcement learning framework with stage-aware policy design, expiring-action abstraction, and search-augmented learning with distillation. Experiments on a controlled multi-supplier replenishment problem, simplified cash-management environments of increasing complexity, and a production-scale simulator show that explicitly modeling this asymmetry improves learning efficiency and becomes increasingly valuable as decision problems scale.
Problem

Research questions and friction points this paper is trying to address.

Markov Decision Processes
sequential decision making
information-action asymmetry
expiring actions
decision urgency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Maturing MDPs
information-action asymmetry
expiring-action priority
structure-aware reinforcement learning
action abstraction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jiaxi Liu
Ant International, School of Economics, Sichuan University
A
Aiping Yang
Ant International
Y
Yuhang Yang
Ant International
S
Shuqi Zhang
Ant International
Z
Zewei Dong
Ant International
J
Jiangming Yang
Ant International
X
Xuebin Chen
School of Economics, Sichuan University, School of Economics, Fudan University