Sample complexity of variance-reduced policy gradient: weaker assumptions and lower bounds

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the theoretical deficiency in sample complexity of existing variance-reduced policy gradient methods, which rely on impractical importance sampling assumptions. To overcome this limitation, we propose a defensive policy gradient algorithm that leverages defensive importance sampling techniques, enabling rigorous sample complexity analysis within a black-box policy optimization framework without requiring additional assumptions such as bounded importance weights. This work is the first to achieve the optimal O(ε⁻³) sample complexity under no extra assumptions, significantly outperforming standard REINFORCE. Furthermore, we establish a matching theoretical lower bound, demonstrating that this acceleration rate is both separated and unimprovable, thereby defining a fundamental theoretical limit for algorithmic efficiency in this setting.
📝 Abstract
Several variance-reduced versions of REINFORCE based on importance sampling achieve an improved $O(ε^{-3})$ sample complexity to find an $ε$-stationary point, under an unrealistic assumption on the variance of the importance weights. In this paper, we propose the \algo (Defensive Policy Gradient) algorithm, based on defensive importance sampling, which achieves the same rate without any assumption on the variance of ordinary importance weights. We also establish lower bounds in a generalized black-box policy-optimization model that hides states and actions and permits parameter-dependent rewards. In this model, the optimal rates are $Θ(ε^{-4})$ with bounded-variance one-policy feedback and $Θ(ε^{-3})$ with mean-square-smooth coupled two-policy feedback. Under standard policy-regularity conditions, REINFORCE and \algo realize the corresponding oracle conditions and attain the $O(ε^{-4})$ and $O(ε^{-3})$ upper bounds, respectively. Although the lower bounds do not apply directly to the classical MDP interaction model in which these algorithms operate, this correspondence provides oracle-level evidence that the faster rate of \algo is optimal and genuinely separated from that of vanilla policy gradient.
Problem

Research questions and friction points this paper is trying to address.

sample complexity
variance-reduced policy gradient
importance sampling
lower bounds
black-box policy optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

variance-reduced policy gradient
defensive importance sampling
sample complexity
lower bounds
black-box policy optimization
🔎 Similar Papers
No similar papers found.