Rate-Optimal Algorithm for Adversarial Linear CMDPs

๐Ÿ“… 2026-09-30
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the challenge of achieving theoretically optimal rates for regret and constraint violation in adversarial linear constrained Markov decision processes by proposing a novel primal-dual algorithm. The method integrates adaptive Follow-the-Regularized-Leader (FTRL), shrinkage value estimation, and exponential Lyapunov function techniques. By introducing an adaptive dual regularizer, it eliminates reliance on Slaterโ€™s condition and policy mixing while maintaining compatibility with the policy class and uniform concentration inequalities. This work bridges the theoretical gap between existing algorithms and optimal rates, attaining $\tilde{O}(\sqrt{K})$ performance without assuming Slaterโ€™s condition. Furthermore, its computational complexity remains independent of the state space size.
๐Ÿ“ Abstract
We study episodic adversarial linear constrained Markov decision processes (CMDPs) with unknown transitions, where both the loss and constraint functions may vary adversarially across episodes. The best previous algorithm achieves $\widetilde{\mathcal{O}}(K^{3/4})$ regret and cumulative constraint violation, leaving a gap to the optimal $\widetilde{\mathcal{O}}(\sqrt{K})$ dependence on the number of episodes $K$. We close this gap by proposing a new primal dual algorithm that achieves $\widetilde{\mathcal{O}}(\sqrt{K})$ regret and cumulative constraint violation without assuming Slater's condition. The main challenge is that learning linear CMDPs requires uniform concentration over a value function class with a controlled covering number, whereas standard techniques in constrained online learning, such as policy mixing, can make this class more complex. Our algorithm combines adaptive Follow the Regularized Leader (FTRL), contracted value estimation, and an exponential Lyapunov function. An adaptive dual regularizer offsets the dependence on the dual weights in the primal regret bound, removing the need for policy mixing. We further show that the normalization in the FTRL update bounds the policy parameters independently of the magnitudes of the dual weights, which explains why the resulting policy class remains compatible with uniform concentration. Under feature access, the computational complexity is independent of the size of the state space.
Problem

Research questions and friction points this paper is trying to address.

constrained Markov decision processes
adversarial learning
regret minimization
constraint violation
linear CMDPs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adversarial Linear CMDPs
Primal-Dual Algorithm
Adaptive FTRL
Contracted Value Estimation
Exponential Lyapunov Function
๐Ÿ”Ž Similar Papers
No similar papers found.