implement experience replay

Designs and implements the data and software infrastructure (buffers, storage, sampling and prioritization policies, roll‑in mechanisms, and tooling) that capture, retain, and serve past experience or simulation traces for training, evaluation, and replay. Builds and analyzes replay procedures and infrastructure (mixing strategies, offline replay testing, management and debugging interfaces) to integrate replay with update pipelines and preserve prior knowledge / reduce catastrophic forgetting.

implementexperiencereplay

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$201K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work systematically investigates the underutilized potential of experience replay in reinforcement learning-based post-training of large language models, challenging the prevailing assumption that fresh online data generation is indispensable. By carefully balancing data staleness, sample diversity, and computational cost, the authors design an efficient replay buffer mechanism that effectively substitutes strict online sampling. Their approach demonstrates, for the first time, that experience replay can significantly reduce inference-time computational overhead while maintaining or even improving model performance and effectively preserving policy entropy. This finding offers a compelling alternative to costly online data collection, suggesting that strategic reuse of historical interactions can sustain training efficacy without compromising behavioral diversity or learning stability.

computational costExperience ReplayLLM post-training

Get Experience from Practice: LLM Agents with Record&Replay

May 23, 2025
EF
Erhu Feng
🏛️ Shanghai Jiao Tong University

To address systemic challenges in LLM-based agents—namely reliability, privacy, cost efficiency, and performance—this paper proposes AgentRR, a novel paradigm that introduces a “record-and-replay” mechanism into LLM agent frameworks. It systematically records task execution traces and abstracts them into structured, reusable “experiences.” The method features multi-level experience abstraction, safety- and generalization-aware check functions, and an experience repository supporting user demonstration recording, collaborative inference between large and small models, and privacy-sensitive execution. Experiments demonstrate that AgentRR significantly improves execution consistency and reliability, reduces redundant computation and API call costs, ensures compliance with privacy regulations, and enables cross-task and cross-user experience reuse. By promoting efficient knowledge retention and transfer, AgentRR advances the deployment of lightweight, trustworthy LLM agents.

Address LLM uncertainty and resource challenges in AI agentsImprove agent reliability, privacy, cost, and performancePropose record-and-replay mechanism for efficient task execution

Experience Replay with Random Reshuffling

Mar 04, 2025
YF
Yasuhiro Fujita
🏛️ Preferred Networks, Inc.

This work addresses the low sample efficiency and training instability of experience replay in reinforcement learning. We systematically introduce the Random Reshuffling (RR) mechanism—previously shown to yield superior convergence properties in supervised learning—into RL experience replay for the first time. We propose RR-based extensions applicable to both uniform and prioritized replay buffers, overcoming statistical redundancy and convergence limitations inherent in traditional independent, with-replacement sampling. Theoretical analysis demonstrates accelerated convergence under RR. Empirical evaluation within the DQN framework on the Atari benchmark shows that, compared to standard prioritized sampling, our approach significantly improves sample efficiency, accelerates convergence, and simultaneously enhances policy performance and training stability. This work establishes a novel paradigm for experience replay in reinforcement learning.

Evaluates methods on Atari benchmarks for effectiveness.Extends random reshuffling to reinforcement learning experience replay.Improves convergence and sample efficiency in deep reinforcement learning.

This work addresses the challenge of catastrophic forgetting and misalignment in continual instruction tuning, where fixed replay ratios fail to adapt to dynamic task distributions. The authors propose PROXYMIX, a novel framework that leverages the “forgetting mirror” hypothesis—empirically validated for the first time—which posits that the relative forgetting sensitivity across tasks remains consistent across model scales. By training a dynamic replay controller on a small proxy model, PROXYMIX transfers this policy to large models without requiring knowledge of future tasks. The controller constructs its state from normalized validation loss and its temporal dynamics, then adaptively blends old and new data via a mask-based mixing mechanism. Evaluated on five sequential instruction-tuning benchmarks with LLaMA-3-8B, PROXYMIX improves average accuracy by 3.4 points, reduces final forgetting by 3.5 points, enhances safety by 5.8 points, and achieves these gains at only 1/50th the policy learning cost of Oracle Target RL.

catastrophic forgettingcontinual instruction tuningdynamic replay

Prioritized Trajectory Replay: A Replay Memory for Data-driven Reinforcement Learning

Jun 27, 2023
JL
Jinyi Liu
🏛️ Tianjin University | NetEase Fuxi AI Lab

In offline reinforcement learning, conventional single-step transition sampling fails to improve policy performance and often introduces out-of-distribution actions, causing training instability. To address this, we propose Trajectory-level Replay (TR), the first framework to extend prioritized sampling to complete trajectories. TR introduces a reverse-trajectory sampling strategy and a trajectory-level priority metric grounded in both TD error and cumulative return, effectively avoiding out-of-distribution action selection. Furthermore, we incorporate a weighted critic objective to mitigate distributional shift inherent in trajectory-level sampling. Evaluated on the D4RL benchmark, TR consistently enhances state-of-the-art algorithms—including BCQ and CQL—achieving average normalized score improvements of 12%–28%. These results empirically validate the effectiveness and generalizability of trajectory-level data utilization as a novel paradigm for offline RL.

Enhancing data sampling from limited offline datasetsImproving offline reinforcement learning with trajectory replayPrioritizing trajectories to boost RL algorithm efficiency

Latest Papers

What's happening recently
View more

This work addresses significant biases in existing trace-based evaluations of caching strategies for Mixture-of-Experts (MoE) models, which suffer from ignoring replay semantics, workload contamination, and runtime mechanisms—leading to distorted rankings and performance assessments. We identify and quantify these three sources of bias for the first time, demonstrating their disruptive impact on cache policy evaluation, and advocate reporting the per-step union of experts and per-layer capacity ratios to ensure comparability. Through event-level atomic simulation, matched-pair interventions, a compulsory-admission oracle, and a causal next-use predictor, we find that lightweight causal eviction policies recover only −11.4% of the offline optimal gap, with an optimal victim selection rate of merely 3.4%, substantially lower than LRU/LFRU (20.6–22.1%), revealing a severe overestimation of practical gains in current mechanisms.

Expert CachingMixture-of-ExpertsReplay Semantics

Hot Scholars

JM

Jianbiao Mei

Zhejiang University
computer visiondeep learning
DT

Dacheng Tao

Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processing
JH

Jianye Hao

Huawei Noah's Ark Lab/Tianjin University
Multiagent SystemsEmbodied AI
LW

Licheng Wen

Shanghai AI Laboratory
AI AgentsAutonomous DrivingRobotics