Reusing Past Samples in Proximal Policy Optimization: When and How Does It Help?

πŸ“… 2026-10-01
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the low sample efficiency of Proximal Policy Optimization (PPO) and the lack of systematic analysis in existing data reuse research. Within a multiple importance weighting framework, we propose two algorithms, wPPO-U and wPPO-BH, which reuse recent experience samples and derive a lower bound on policy improvement to provide theoretical support for data reuse. To our knowledge, this work is the first to systematically delineate the applicable scenarios and extent of data reuse in PPO, incorporating a balancing heuristic correction to optimize policy updates. Experimental results demonstrate that the proposed methods significantly enhance both sample efficiency and final performance on continuous control tasks.
πŸ“ Abstract
Among on-policy deep reinforcement learning methods, Proximal Policy Optimization (PPO) has become the de facto standard, due to its consistently strong empirical performance across diverse application domains. However, on-policy methods are inherently sample inefficient: fresh data collected under the current policy is used for just a few updates before being discarded. Off-policy methods avoid this inefficiency via experience replay, achieving notable sample efficiency gains, but at the cost of training instabilities or extensive tuning. This motivated the rise of hybrid strategies that augment PPO with off-policy data reuse. Existing sample-reuse variants of PPO demonstrated improved sample efficiency over vanilla PPO, yet a systematic study of when reuse helps, in which scenarios, and to what extent remains missing. In this work, we study the effectiveness of sample reuse in PPO by instantiating two variants within a multiple importance weighting framework. Both retain the core PPO mechanics, reusing only samples from a window of recent iterations, thereby isolating the effect of data reuse from other factors. The variants, termed wPPO-U and wPPO-BH, employ vanilla importance weights or balance-heuristic-corrected ones, respectively. For both, we derive policy improvement lower bounds providing theoretical grounding for their respective losses. We use them to empirically study when and how data reuse improves sample efficiency or final performance of PPO across continuous control tasks.
Problem

Research questions and friction points this paper is trying to address.

Proximal Policy Optimization
sample reuse
sample efficiency
deep reinforcement learning
on-policy methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

Proximal Policy Optimization
Sample Reuse
Multiple Importance Weighting
Policy Improvement Lower Bounds
Deep Reinforcement Learning