Single or Multiple Policies for Phase-Structured Reinforcement Learning?

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear applicability boundaries between single-policy and multi-policy approaches in non-stationary reinforcement learning by proposing a phase-structure-based duration decomposition method. By uncovering the root causes underlying the divergence between theoretical equivalence and practical performance, we establish a policy selection criterion centered on the ratio of transient to quasi-stationary durations, validated through state augmentation techniques and numerical simulations. The research confirms key hypotheses, demonstrating that prolonged phases favor multi-policy methods while increased environmental heterogeneity imposes greater burdens on single-policy approaches. These findings provide a rigorous theoretical foundation for determining optimal policy types in non-stationary environments, thereby effectively guiding efficient policy design.
📝 Abstract
Many reinforcement-learning (RL) problems are non-stationary yet structured and can be decomposed into phases, each with its own transition probabilities and reward functions. When the phase sequence is known, the common solution augments the state with information to satisfy the Markovian property and applies standard RL techniques. However, prior work finds that the multi-policy approach for different phases can outperform a single state-augmented policy shared among the phases, for reasons that remain unclear. In this work, we first show that the shared policy can theoretically achieve performance of any multi-policy solution. However, whether a multi-policy solution can perform better than the corresponding single shared policy in practice depends on function approximation, learning and optimization processes, as well as, for multi-policy solutions, the sample efficiency and loss of continuity from one policy to another. We propose a regime-based phase decomposition method to identify which policy can provide better performance. The method is based on consideration of the duration of transient system dynamics relative to the duration of the quasi-stationary period. Numerical experiments are conducted with different non-stationary RL problems to validate our four major hypotheses: (a) longer phase durations favor multi-policies, (b) the heterogeneity between phases increases the burden on single policy, (c) multi-policies need sufficient data for each phase, and (d) environment-specific transition dynamics between phases can affect which policy is preferable.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Non-stationary
Phase-structured
Single vs Multiple Policies
Function Approximation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Phase-Structured Reinforcement Learning
Single vs Multiple Policies
Regime-Based Phase Decomposition
Non-Stationary RL
Function Approximation
🔎 Similar Papers
No similar papers found.