Vanilla Policy Optimization Is Both Optimal and Differentially Private for Stochastic Contextual Bandits

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the suboptimal regret bottleneck encountered in policy optimization for stochastic contextual bandits when explicit exploration bonuses are absent. It reveals an implicit exploration mechanism inherent in standard exponential policy updates under realizability conditions, proving that near-optimal regret can be achieved without explicit exploration and extending this result to batched settings. Furthermore, by incorporating a private regression oracle, the authors design a differentially private algorithm that effectively circumvents the accumulation of compositional errors. Theoretically, this work establishes high-probability near-optimal regret bounds, while empirically validating the effectiveness of the proposed approach under privacy constraints. Ultimately, it introduces a novel paradigm for policy optimization that simultaneously reconciles rigorous privacy preservation with efficient learning.
📝 Abstract
Can vanilla policy optimization explore enough to achieve near-optimal regret in stochastic contextual bandits? We show that standard exponential policy updates driven by offline regression do so under realizability, without exploration bonuses or importance weighting. For $A$ actions, $T$ rounds, and a finite prediction class $F$, vanilla PO achieves $\widetilde O(\sqrt{AT\log(|F|)})$ regret with high probability. Our analysis reveals an implicit exploration mechanism of independent interest: gradual policy updates prevent actions from losing probability too quickly, allowing the regression oracle to learn their expected losses. We further develop a batched version using only $O(\log T)$ regression calls and policy switches, and show how private regression oracles yield differentially private contextual bandit algorithms without composition across batches. For a finite class, this gives pure $\varepsilon_{\rm priv}$-DP and regret $\widetilde O\left( \sqrt{AT \log(|F|/\delta)}(1+\varepsilon_{\rm priv}^{-1/2}) \right)$. Finally, experiments across oracle-based contextual bandit algorithms, with and without privacy, demonstrate the practical effectiveness of policy optimization and the value of explicit exploration under stronger privacy constraints.
Problem

Research questions and friction points this paper is trying to address.

stochastic contextual bandits
policy optimization
regret minimization
differential privacy
batched learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vanilla Policy Optimization
Stochastic Contextual Bandits
Implicit Exploration
Batched Algorithm
Differential Privacy
🔎 Similar Papers
No similar papers found.