Score
Designs, implements, and operates experience replay buffers and their management policies—specifying storage formats, capacity, sampling strategies, insertion/eviction rules, and APIs for integration with continual-learning or reinforcement-learning systems. Builds tests and audit procedures to detect and quantify initialization- and manipulation-induced biases (e.g., empty vs prefilled starts), tunes buffer size and sampling to preserve class or representational support, and analyzes the buffer's contribution to overall learning suboptimality.
This work systematically investigates the underutilized potential of experience replay in reinforcement learning-based post-training of large language models, challenging the prevailing assumption that fresh online data generation is indispensable. By carefully balancing data staleness, sample diversity, and computational cost, the authors design an efficient replay buffer mechanism that effectively substitutes strict online sampling. Their approach demonstrates, for the first time, that experience replay can significantly reduce inference-time computational overhead while maintaining or even improving model performance and effectively preserving policy entropy. This finding offers a compelling alternative to costly online data collection, suggesting that strategic reuse of historical interactions can sustain training efficacy without compromising behavioral diversity or learning stability.
Catastrophic forgetting remains a fundamental challenge in continual learning. Method: This paper systematically investigates memory sampling strategies for experience replay, challenging the conventional default assumption of uniform sampling. We propose 50 non-uniform sampling strategies based on randomized weighted probability distributions and evaluate them within standard replay frameworks, integrated with mainstream continual learning methods (e.g., EWC, LwF) and benchmarks (Split-CIFAR100, PermutedMNIST). Contribution/Results: Experiments across diverse buffer sizes, model architectures, and task sequences consistently identify at least one non-uniform strategy that significantly outperforms uniform sampling—demonstrating its robust and general effectiveness. Crucially, we provide the first empirical evidence that adaptive replay strategies play a pivotal role in mitigating forgetting. Our work establishes a new paradigm for experience replay design and delivers a reproducible, optimization-guided pathway for improving replay-based continual learning systems.
This work addresses the low sample efficiency and training instability of experience replay in reinforcement learning. We systematically introduce the Random Reshuffling (RR) mechanism—previously shown to yield superior convergence properties in supervised learning—into RL experience replay for the first time. We propose RR-based extensions applicable to both uniform and prioritized replay buffers, overcoming statistical redundancy and convergence limitations inherent in traditional independent, with-replacement sampling. Theoretical analysis demonstrates accelerated convergence under RR. Empirical evaluation within the DQN framework on the Atari benchmark shows that, compared to standard prioritized sampling, our approach significantly improves sample efficiency, accelerates convergence, and simultaneously enhances policy performance and training stability. This work establishes a novel paradigm for experience replay in reinforcement learning.
This work addresses the limitation of conventional experience replay buffers in reinforcement learning, which lack semantic-aware prioritization and thereby constrain sample efficiency and policy performance. The authors propose a novel approach that integrates a frozen pre-trained vision-language model (VLM) into experience replay to automatically assess the semantic value of agent interaction sub-trajectories, enabling semantic-driven prioritized sampling without fine-tuning. This method maintains high interpretability while significantly enhancing generalization across diverse tasks. Evaluated on a range of discrete and continuous control benchmarks—including gaming and robotic manipulation—it achieves an average success rate improvement of 11–52% and boosts sample efficiency by 19–45% compared to standard baselines.
In online reinforcement learning, uniform experience replay is inefficient, while conventional prioritized replay risks overfitting due to excessive sampling of sparse high-value transitions. To address this, we propose a generative experience replay buffer grounded in conditional diffusion models, guided by a differentiable relevance function jointly driven by curiosity and value estimation—replacing explicit priority-based sampling with controllable synthesis of high-value, diverse experiences. This work represents the first deep integration of prioritized replay principles with parametric generative modeling, enabling explicit diversity control while preserving experience density. Experiments demonstrate significant improvements in sample efficiency and policy performance across both state- and pixel-based environments, support higher update-to-data ratios during training, and effectively mitigate overfitting.
This work addresses the instability and low sample efficiency in Generalized Reinforcement Learning with Policy Optimization (GRPO) for large language model post-training, which stem from rapid policy drift causing experience replay samples to become outdated. To mitigate this issue, the authors propose a rollout-level experience replay mechanism tailored for GRPO, integrating rollout-based storage and sampling, advantage-weighted prioritized replay, a maximum step-age (τ_max) eviction strategy, and the construction of fresh anchor batches. Experiments on three Qwen3-Base model scales across five mathematical benchmarks demonstrate that the proposed method significantly outperforms both standard GRPO and naive replay baselines, achieving an average accuracy gain of 4.35 percentage points and an improvement of 0.579 in the AES efficiency metric for the 4B model.
This work addresses the high memory overhead and redundancy in experience replay buffers commonly used in deep reinforcement learning, where conventional compression techniques often introduce bias. The authors propose a novel experience compression method based on n-step sequence endpoints, which constructs a compact buffer by retaining only the initial and terminal states of n-step trajectories. This approach significantly reduces storage requirements while preserving the ability to propagate long-horizon value estimates. By integrating n-step temporal difference learning with an endpoint sampling strategy, the method effectively mitigates systematic bias. Empirical results demonstrate that on the Pinball and Atari 2600 benchmarks, the proposed technique achieves performance comparable to that of standard large-scale replay buffers using only one-tenth of the buffer capacity.
This work addresses a critical vulnerability in existing experience replay mechanisms for continual learning: their susceptibility to stealthy attacks that manipulate replay sample selection without altering the data itself. We demonstrate for the first time that an adversary with access only to replay indices can severely degrade model performance by skewing the class distribution of replayed samples. To this end, we propose an audit-aware stealthy attack framework that balances invisibility and impact. Our method estimates per-class utility via lightweight metrics—either EMA loss or confidence—and optimizes the replay distribution using projection techniques based on KL divergence (via exponential tilting) or total variation distance (via mass redistribution). A sliding-window scheduler further adapts the attack to rolling audits. Experiments across multiple continual learning benchmarks show significant drops in final accuracy and backward transfer, with the KL-based variant achieving high destructiveness while remaining highly evasive under diverse auditing mechanisms.