Vanishing L2 regularization for the softmax Multi Armed Bandit

📅 2026-05-05
📈 Citations: 0
Influential: 0
📄 PDF

career value

219K/year
🤖 AI Summary
This work addresses the long-standing challenge of characterizing the convergence behavior of L2-regularized softmax policy gradient methods as the regularization coefficient vanishes. Focusing on the multi-armed bandit setting, the paper establishes the first rigorous convergence analysis framework for softmax policy gradient under vanishing L2 regularization, circumventing the conventional reliance on convexity assumptions. Through a combination of theoretical analysis and numerical experiments, the study formally proves that the algorithm converges even as the regularization parameter approaches zero. Furthermore, extensive evaluations on standard benchmarks demonstrate that this vanishing-regularization regime consistently outperforms both unregularized and fixed-regularization strategies, highlighting its practical efficacy and theoretical significance.
📝 Abstract
Multi Armed Bandit (MAB) algorithms are a cornerstone of reinforcement learning and have been studied both theoretically and numerically. One of the most commonly used implementation uses a softmax mapping to prescribe the optimal policy and served as the foundation for downstream algorithms, including REINFORCE. Distinct from vanilla approaches, we consider here the L2 regularized softmax policy gradient where a quadratic term is subtracted from the mean reward. Previous studies exploiting convexity failed to identify a suitable theoretical framework to analyze its convergence when the regularization parameter vanishes. We prove here theoretical convergence results and confirm empirically that this regime makes the L2 regularization numerically advantageous on standard benchmarks.
Problem

Research questions and friction points this paper is trying to address.

L2 regularization
softmax
Multi Armed Bandit
convergence
vanishing regularization
Innovation

Methods, ideas, or system contributions that make the work stand out.

vanishing regularization
softmax policy gradient
Multi-Armed Bandit
L2 regularization
convergence analysis
🔎 Similar Papers
No similar papers found.