On the Dynamic Regret of Following the Regularized Leader: Optimism with History Pruning

📅 2025-05-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the loose dynamic regret bounds of Follow-the-Regularized-Leader (FTRL) in dynamic online convex optimization (OCO), identifying the root cause as the decoupling between state updates and iterates—not the projection mechanism, as conventionally assumed. To overcome this, we propose a novel analytical framework integrating optimistic prediction of future costs with linearized gradient pruning over historical gradients. Our approach employs recursive regularization to tightly couple states and iterates, enabling loop-free optimistic design and continuous interpolation between greediness and agility. The framework recovers classical dynamic regret upper bounds as special cases, yields finer-grained control over regret terms, and achieves the optimal $O(sqrt{T})$ dynamic regret over compact domains—without increasing gradient queries or memory overhead.

Technology Category

Search and Optimization: Learning to SearchMachine Learning: Online Learning & BanditsReasoning under Uncertainty: Stochastic Optimization

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
We revisit the Follow the Regularized Leader (FTRL) framework for Online Convex Optimization (OCO) over compact sets, focusing on achieving dynamic regret guarantees. Prior work has highlighted the framework's limitations in dynamic environments due to its tendency to produce"lazy"iterates. However, building on insights showing FTRL's ability to produce"agile"iterates, we show that it can indeed recover known dynamic regret bounds through optimistic composition of future costs and careful linearization of past costs, which can lead to pruning some of them. This new analysis of FTRL against dynamic comparators yields a principled way to interpolate between greedy and agile updates and offers several benefits, including refined control over regret terms, optimism without cyclic dependence, and the application of minimal recursive regularization akin to AdaFTRL. More broadly, we show that it is not the lazy projection style of FTRL that hinders (optimistic) dynamic regret, but the decoupling of the algorithm's state (linearized history) from its iterates, allowing the state to grow arbitrarily. Instead, pruning synchronizes these two when necessary.
Problem

Research questions and friction points this paper is trying to address.

Achieving dynamic regret guarantees in FTRL for OCO
Overcoming lazy iterates limitation in dynamic environments
Synchronizing algorithm state and iterates via history pruning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Optimistic composition of future costs
Careful linearization of past costs
Pruning to synchronize state and iterates
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
N
N. Mhaisen
Faculty of Electrical Engineering, Mathematics and Computer Science, TU Delft, Netherlands
G
G. Iosifidis
Faculty of Electrical Engineering, Mathematics and Computer Science, TU Delft, Netherlands