$\tilde{O}(\sqrt{T})$ Regret and Polylogarithmic Constraint Violation for COCO

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of excessive cumulative constraint violation in online convex optimization with constraints under adversarial convex losses. To this end, it proposes a coupled framework integrating the continuous Hedge algorithm with a shrinking feasible set elimination strategy. The approach controls regret via adaptive learning rate scheduling and potential function techniques, while leveraging Grünbaum's inequality to precisely analyze probability mass reduction for suppressing constraint violations. This work makes the first contribution to reducing cumulative constraint violation from polynomial to poly-logarithmic order. Theoretically, the proposed algorithm simultaneously achieves a near-optimal regret bound of $O(\sqrt{T \log T})$ and a cumulative constraint violation of $O(\log^2 T)$, representing a significant breakthrough in balancing optimization performance with constraint satisfaction.
📝 Abstract
We study constrained online convex optimization with adversarial convex losses and constraints ($\mathsf{COCO}$). At each round \(t\in[T]\), a learner selects \(x_t\) from a \(d\)-dimensional convex decision set \(\mathcal X\), after which an adaptive adversary reveals a convex cost function \(f_t\) and constraint function \(g_t\). Consequently, the learner incurs cost \(f_t(x_t)\) and constraint violation \(\max\{0,g_t(x_t)\}\), and aims to simultaneously minimize regret and cumulative constraint violation ($\mathsf{CCV}$) over the entire horizon. Existing algorithms achieve \(O(\sqrt{T})\) regret and \(\widetilde O(\sqrt{T})\) $\mathsf{CCV}$. We show that an online policy can achieve \(O(\sqrt{T\log T})\) regret and \(O(\log^2 T)\) $\mathsf{CCV}$, reducing the $\mathsf{CCV}$ from polynomial to polylogarithmic while retaining near-optimal regret. Our approach combines continuous Hedge with elimination on shrinking feasible sets. The key observation is that whenever the mean of the Hedge distribution violates a constraint, Gr\"unbaum's inequality guarantees that a constant fraction of the Hedge probability mass is eliminated. We use an adaptive learning-rate schedule and a potential function coupling the surviving volume with the learning rate to convert this probability-mass reduction into a bound of \(O(\log^2 T)\) on the $\mathsf{CCV}$.
Problem

Research questions and friction points this paper is trying to address.

Constrained Online Convex Optimization
Regret Minimization
Cumulative Constraint Violation
Adversarial Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Constrained Online Convex Optimization
Cumulative Constraint Violation
Continuous Hedge
Grünbaum's Inequality
Adaptive Learning Rate
🔎 Similar Papers
No similar papers found.