Parameter-Free Heavy-Tailed Bandits

📅 2026-07-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a fundamental limitation of existing bandit algorithms under heavy-tailed reward distributions: their reliance on prior knowledge of tail parameters—such as the tail index $\varepsilon$ and moment bound $u$—which are typically unknown in practice. We propose the first fully adaptive bandit algorithm that requires no prior information about these parameters. By employing a scheduled exploration mechanism to adapt to the unknown moment bound $u$ and incorporating a calibrated design tailored to the critical endpoint case $\varepsilon = 1$, our method achieves sublinear regret under the minimal assumption $\varepsilon > 0$. When $u$ is unknown, it matches the optimal regret rate attainable with known parameters up to logarithmic factors. Furthermore, we resolve an open problem posed at COLT 2025 by proving that uniformly sublinear regret over the entire range $\varepsilon \in (0,1]$ is unattainable, thereby establishing the fundamental limits of adaptation and the optimal regret lower bound without any parametric assumptions.
📝 Abstract
Heavy-tailed distributions arise naturally in sequential decision-making problems such as financial investment, online advertising, and network management, where rare but extreme outcomes can dominate performance. Heavy-tailed bandits model online decision-making in these settings by assuming only that rewards $X$ satisfy $\mathbb{E}[|X|^{1+ε}]\leq u$, for some tail exponent $ε\in(0,1]$ and moment bound $u<+\infty$. However, most existing regret minimization algorithms require these parameters to be known. This assumption is particularly restrictive in practice: $ε$ and $u$ govern the frequency and magnitude of rare events and are therefore precisely the quantities that are hardest to infer reliably from limited observations. Motivated by an open problem posed by Genalti and Metelli at COLT 2025, we resolve the assumption-free adaptation problem for heavy-tailed bandits and characterize the price in the regret of not knowing the tail parameters. We first study adaptation to the moment bound $u$ for a fixed tail exponent $ε$. We prove that every algorithm unaware of $u$, or of any upper bound on it, must obey a sharp trade-off between its distribution-dependent and distribution-free regret guarantees. We then introduce a scheduled-exploration algorithm that requires no knowledge of $u$ and matches the resulting adaptation frontier up to logarithmic factors. Finally, we show that the same algorithm can be instanced without knowing $ε$ by calibrating its exploration schedule to the endpoint $ε=1$. It achieves sublinear regret for every fixed $ε>0$, while no algorithm can guarantee sublinear regret uniformly over all $ε\in(0,1]$. Altogether, our results resolve the COLT open problem without additional distributional assumptions and provide a sharp characterization of the statistical cost of adapting to unknown heavy tails.
Problem

Research questions and friction points this paper is trying to address.

heavy-tailed bandits
parameter-free
regret minimization
tail exponent
moment bound
Innovation

Methods, ideas, or system contributions that make the work stand out.

parameter-free
heavy-tailed bandits
regret minimization
adaptation frontier
scheduled exploration
🔎 Similar Papers