Policy Iteration Is Not Strongly Polynomial for Deterministic Markov Decision Processes: The Price of Algorithmic Anarchy

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the open question of whether Howard policy iteration admits strongly polynomial complexity in deterministic discounted Markov decision processes (MDPs), alongside the price of anarchy incurred by decentralized improvement algorithms. By integrating combinatorial optimization theory with lower-bound construction techniques, this work analyzes the iteration complexity under bounded bit-length reward constraints and establishes an exponential separation from Dantzig’s simplex method. The authors prove both exponential and stretched-exponential iteration lower bounds for Howard policy iteration, definitively ruling out its strongly polynomial nature. Furthermore, they precisely quantify the algorithmic price of anarchy arising from decentralized improvements. These findings provide critical theoretical foundations for understanding the fundamental limits of policy iteration in deterministic MDPs.
πŸ“ Abstract
We establish an exponential iteration lower bound in the number of states for Howard's policy iteration on deterministic discounted Markov decision processes, with at most two actions per state. This rules out strong polynomiality of Howard's policy iteration when the discount factor is part of the input and yields an exponential separation from the simplex method with Dantzig's pivoting rule, which is proved to be strongly polynomial on this class. Even when each reward is restricted to logarithmic bit length, we obtain a stretched-exponential iteration lower bound. The gap between Howard's decentralized and simultaneous selfish improvements and Dantzig's coordinated selection of a single action with the largest gain across all states reveals a ``price'' of algorithmic anarchy.
Problem

Research questions and friction points this paper is trying to address.

Policy Iteration
Markov Decision Processes
Strong Polynomiality
Iteration Lower Bound
Algorithmic Anarchy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Policy Iteration
Markov Decision Processes
Strong Polynomiality
Exponential Lower Bound
Price of Anarchy
πŸ”Ž Similar Papers