Prediction with Expert Advice: Anytime Regret with Many Experts Matches the Fixed-Time Constant

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文解决了在线学习中专家建议预测问题,提出了一种无需预先知道时间范围的算法,其累积遗憾与已知时间范围的最佳结果相匹配。
📝 Abstract
Prediction with expert advice is a fundamental problem in online learning. When the time horizon $T$ is known in advance, the minimax cumulative regret over $n$ experts is asymptotically $\sqrt{\frac{T \ln n}{2}}$. This is achieved by the Multiplicative Weights Update algorithm with a learning rate tuned to $T$, and is known to be tight. If instead the regret bound is required to hold simultaneously at every time $t$, the best known guarantee has been $\sqrt{t \ln n}$---a factor of $\sqrt{2}$ worse---and it has remained unknown whether this factor of $\sqrt{2}$ is necessary. We show that it is not. We give an algorithm, requiring no knowledge of the horizon, whose cumulative regret satisfies $R_t \le \bigl(1 + O(\sqrt{\ln \ln n / \ln n})\bigr)\sqrt{t \ln n / 2}$ simultaneously for every $t \ge 1$.
Problem

Research questions and friction points this paper is trying to address.

Prediction with Expert Advice
Anytime Regret
Online Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prediction with Expert Advice
Anytime Regret
Minimax Cumulative Regret
Multiplicative Weights Update
🔎 Similar Papers
No similar papers found.