Toward Optimal Regret in Adversarial MDPs with Stochastic Hard Constraints

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
The original TLDR provided is incomplete, containing only fragmented phrases and lacking specific research details. Accordingly, a standardized academic TLDR template is constructed below, with bracketed placeholders intended for substitution with actual information. To address the issue of [core challenge or limitation of existing methods] in [specific domain], this work proposes [name of the novel method]. By introducing [core technology or mechanism] integrated with [auxiliary strategy or framework], the proposed approach enables efficient modeling and optimization of [research subject]. Experimental results demonstrate that the method significantly outperforms existing baselines on [benchmark dataset or task], improving [key performance metric] by [X]%. The primary contribution of this study lies in achieving [main innovation] for the first time, thereby offering a novel theoretical perspective and technical paradigm for [related research directions].
📝 Abstract
We study episodic constrained Markov decision processes with adversarial losses under stochastic hard constraints. Specifically, starting from a known strictly feasible policy with margin $d$, we seek to obtain optimal regret while satisfying the expected cost constraints in every episode. In this setting, Stradi et al. (2025) show that a carefully designed mixing rule attains regret of order $\widetilde{\mathcal{O}}(\sqrt{T}/\min\{d,d^2\})$. Interestingly, they also provide a lower bound of order $Ί(\sqrt{T}/΁)$ for the same setting, where $΁$ is the Slater margin of the offline problem and can be much larger than $d$. In this work, we build on their approach to obtain optimal regret dependence on these margins. Specifically, we propose MA-OPS, an algorithm that combines an optimistic search for the Slater margin with a pessimistic evaluation of the selected policies to safely learn a policy with a large feasibility margin. This policy is then used to minimize regret while satisfying the constraints at every episode. In particular, we show that MA-OPS attains regret $\widetilde{\mathcal{O}}(\sqrt{T}/΁+ 1/(d΁))$. Finally, we provide a matching lower bound, showing that the dependence on $T$, $d$, $΁$ in the regret bound is optimal up to logarithmic factors.
Problem

Research questions and friction points this paper is trying to address.

Constrained MDPs
Adversarial losses
Stochastic hard constraints
Optimal regret
Slater margin
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adversarial MDPs
Stochastic Hard Constraints
Optimal Regret
Slater Margin
MA-OPS
🔎 Similar Papers
đŸ’ŧ Related Jobs
No related jobs found.