๐ค AI Summary
This study addresses the challenge of achieving theoretically optimal rates for regret and constraint violation in adversarial linear constrained Markov decision processes by proposing a novel primal-dual algorithm. The method integrates adaptive Follow-the-Regularized-Leader (FTRL), shrinkage value estimation, and exponential Lyapunov function techniques. By introducing an adaptive dual regularizer, it eliminates reliance on Slaterโs condition and policy mixing while maintaining compatibility with the policy class and uniform concentration inequalities. This work bridges the theoretical gap between existing algorithms and optimal rates, attaining $\tilde{O}(\sqrt{K})$ performance without assuming Slaterโs condition. Furthermore, its computational complexity remains independent of the state space size.
๐ Abstract
We study episodic adversarial linear constrained Markov decision processes (CMDPs) with unknown transitions, where both the loss and constraint functions may vary adversarially across episodes. The best previous algorithm achieves $\widetilde{\mathcal{O}}(K^{3/4})$ regret and cumulative constraint violation, leaving a gap to the optimal $\widetilde{\mathcal{O}}(\sqrt{K})$ dependence on the number of episodes $K$. We close this gap by proposing a new primal dual algorithm that achieves $\widetilde{\mathcal{O}}(\sqrt{K})$ regret and cumulative constraint violation without assuming Slater's condition. The main challenge is that learning linear CMDPs requires uniform concentration over a value function class with a controlled covering number, whereas standard techniques in constrained online learning, such as policy mixing, can make this class more complex. Our algorithm combines adaptive Follow the Regularized Leader (FTRL), contracted value estimation, and an exponential Lyapunov function. An adaptive dual regularizer offsets the dependence on the dual weights in the primal regret bound, removing the need for policy mixing. We further show that the normalization in the FTRL update bounds the policy parameters independently of the magnitudes of the dual weights, which explains why the resulting policy class remains compatible with uniform concentration. Under feature access, the computational complexity is independent of the size of the state space.