The Price of Hidden Curvature: An $\widetildeΩ (d^{5/4} \sqrt{T})$ Lower Bound for Bandit Convex Optimization

📅 2026-07-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work investigates the minimax regret lower bound for stochastic constrained convex optimization under bandit feedback, aiming to determine whether it surpasses the well-known $d\sqrt{T}$ rate established for linear bandits. By constructing a novel class of convex functions that embed hidden curvature—built from combinations of soft-max and distance functions—and leveraging posterior diffusion analysis of the Fisher information matrix together with adaptive action sequence techniques, the study establishes the first regret lower bound of $\widetilde{\Omega}(d^{5/4} \sqrt{T})$. This result breaks the long-standing $d\sqrt{T}$ barrier and further yields a corresponding sample complexity lower bound of $\widetilde{\Omega}(d^{5/2}/\varepsilon^2)$. The findings demonstrate that bandit convex optimization is inherently more challenging than its linear counterpart and extend naturally to the unconstrained setting.
📝 Abstract
We establish a $\widetildeΩ(d^{5/4}\sqrt T)$ lower bound on the minimax expected regret of stochastic bandit convex optimization of $1$-Lipschitz functions on the Euclidean ball. This presents the first nontrivial regret lower bound that grows faster than $d\sqrt{T}$ for this problem, establishing that stochastic bandit convex optimization is fundamentally harder than linear bandits. The hard class of convex functions we construct takes the following form in dimension $2d$: for an action $a = (a^1,a^2) \in \mathbb{B}^{2d}_2$, each function is the scaled soft maximum of a "tube", $r^{-1} \| W^\star a^1 - \frac{r}{8\varepsilon} a^2 \|_2$ (hyperparameterized by $\varepsilon,r$), and a squared distance function, $\frac12 \| a^1 - u^\star \|_2^2 - \frac12 \| u^\star \|_2^2$. Here, $W^\star \in \mathbb{R}^{d \times d}$ is an unknown linear transformation, and $u^\star \in \mathbb{R}^{d}$ is an unknown vector which must be learned to minimize the function. Observations are informative about $u^\star$ only when the learner's action lies near the tube determined by $W^\star$, satisfying $a^2 \approx \frac{8\varepsilon}{r} W^\star a^1$: thus the learner must either find this tube without knowing $W^\star$, or spend observations learning useful directions of $W^\star$. Formally, our regret analysis exploits this tradeoff by bounding the posterior spread of Fisher information matrices obtained under an adaptive sequence of actions. Together, these ingredients give a sample complexity lower bound of $\widetildeΩ(d^{5/2}/\varepsilon^2)$ to find an $\varepsilon$-optimal action, which translates to an $\widetildeΩ (d^{5/4} \sqrt{T})$ regret lower bound. We also extend this lower bound to the unconstrained setting where the action space is $\mathbb{R}^d$.
Problem

Research questions and friction points this paper is trying to address.

bandit convex optimization
minimax regret
lower bound
stochastic optimization
Lipschitz functions
Innovation

Methods, ideas, or system contributions that make the work stand out.

bandit convex optimization
minimax lower bound
hidden curvature
Fisher information
adaptive exploration
🔎 Similar Papers
2024-02-09arXiv.orgCitations: 10