đ¤ AI Summary
The original TLDR provided is incomplete, containing only fragmented phrases and lacking specific research details. Accordingly, a standardized academic TLDR template is constructed below, with bracketed placeholders intended for substitution with actual information. To address the issue of [core challenge or limitation of existing methods] in [specific domain], this work proposes [name of the novel method]. By introducing [core technology or mechanism] integrated with [auxiliary strategy or framework], the proposed approach enables efficient modeling and optimization of [research subject]. Experimental results demonstrate that the method significantly outperforms existing baselines on [benchmark dataset or task], improving [key performance metric] by [X]%. The primary contribution of this study lies in achieving [main innovation] for the first time, thereby offering a novel theoretical perspective and technical paradigm for [related research directions].
đ Abstract
We study episodic constrained Markov decision processes with adversarial losses under stochastic hard constraints. Specifically, starting from a known strictly feasible policy with margin $d$, we seek to obtain optimal regret while satisfying the expected cost constraints in every episode. In this setting, Stradi et al. (2025) show that a carefully designed mixing rule attains regret of order $\widetilde{\mathcal{O}}(\sqrt{T}/\min\{d,d^2\})$. Interestingly, they also provide a lower bound of order $Ί(\sqrt{T}/Ī)$ for the same setting, where $Ī$ is the Slater margin of the offline problem and can be much larger than $d$. In this work, we build on their approach to obtain optimal regret dependence on these margins. Specifically, we propose MA-OPS, an algorithm that combines an optimistic search for the Slater margin with a pessimistic evaluation of the selected policies to safely learn a policy with a large feasibility margin. This policy is then used to minimize regret while satisfying the constraints at every episode. In particular, we show that MA-OPS attains regret $\widetilde{\mathcal{O}}(\sqrt{T}/Ī+ 1/(dĪ))$. Finally, we provide a matching lower bound, showing that the dependence on $T$, $d$, $Ī$ in the regret bound is optimal up to logarithmic factors.