🤖 AI Summary
This study addresses the challenge of policy optimization in multichain Markov decision processes, where the optimal gain depends on the initial state and exhibits a complex recursive structure. We propose the first general average-reward reinforcement learning algorithm for multichain MDPs that requires no prior model knowledge and avoids reduction to discounted formulations. Based on Bather decomposition, the state space is hierarchically partitioned into communicating classes and transient states. Asynchronous value iteration is then employed to reformulate global decision-making as structured subproblems, approximately solving the average optimality equation. The proposed algorithm converges almost surely to the optimal gain, offering finite-time convergence guarantees and near bias-optimality while significantly improving transient performance. Experiments further validate the trade-offs among different algorithmic variants.
📝 Abstract
We study learning optimal policies in average-reward multichain Markov decision processes (MDPs), where the optimal gain may depend on the initial state and recurrence structures vary across policies, creating challenges for reinforcement learning (RL) methods. We propose an asynchronous value-iteration-based RL algorithm that requires no model knowledge beyond the MDP's transition graph and leverages Bather's decomposition to hierarchically partition the state space into communicating subsystems and transient states. This decomposition induces a recasting of the global decision problem into structured subproblems, which our algorithm exploits. We show that the algorithm converges to the optimal gain and produces gain-optimal policies after finite time. Building on this base algorithm, we develop two further algorithms: one approximately solves the multichain average optimality equations to obtain near gain-optimal policies, and another targets near bias-optimality by approximating the optimal bias function and solving an induced average-reward multichain MDP using the base algorithm. We provide almost-sure convergence guarantees for all three algorithms and empirically compare their tradeoffs, showing that the latter two also consistently improve transient performance relative to the base algorithm. To our knowledge, these are the first essentially model-free average-reward RL algorithms for general multichain MDPs without reductions to discounted problems.