hierarchical experience replay

Designs and implements memory and replay systems that store, organize, and represent past experiences at multiple granularities—from fine-grained transitions to aggregated episodes and abstract summaries—using structures such as hierarchical pools, banks, or indexed memories. Creates aggregation, indexing, and retrieval policies that enable coarse-to-fine replay for planning and learning, including matching stored items to degradation or applicability patterns, prioritizing or ordering removals, and producing representations suitable for efficient lookup and replay.

hierarchicalexperiencereplay

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.13
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$185K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Mitigating Catastrophic Forgetting in Streaming Generative and Predictive Learning via Stateful Replay

Nov 22, 2025
WD
Wenzhang Du
🏛️ Mahanakorn University of Technology

To address catastrophic forgetting in memory-constrained streaming learning, this paper proposes a unified continual learning framework based on state replay, applicable to generative (autoencoding), time-series forecasting, and classification tasks. Unlike naive sequential fine-tuning or black-box replay, we formulate state replay as a joint optimization objective, enabling cooperative parameter updates over new and replayed samples via stochastic gradient methods; theoretical analysis grounded in gradient alignment reveals necessary conditions for effective forgetting mitigation. Evaluated across six heterogeneous and stationary streaming settings—constructed from Rotated MNIST, Electricity, and Airlines datasets—the method reduces average forgetting by 2–3× under heterogeneous multi-task streams, while matching fine-tuning performance on stationary streams. Our key contributions are: (i) the first theoretical analysis framework for state replay that unifies generative and discriminative tasks, and (ii) empirical validation of its effectiveness and robustness as a strong baseline for streaming continual learning.

Evaluating stateful replay on heterogeneous multitask and time-based data streamsMitigating catastrophic forgetting in streaming learning under memory constraintsStudying replay mechanisms across generative and predictive learning objectives

This work addresses the challenge of catastrophic forgetting in large language models during continual fine-tuning, a problem inadequately mitigated by existing replay methods that either rely on heuristic rules or incur high computational costs. Inspired by human memory mechanisms, the authors propose the Memory Strength–aware Sample Replay (MSSR) framework, which dynamically optimizes both the content and timing of replayed samples through sample-level memory strength estimation and adaptive replay scheduling. By integrating experience replay, memory strength modeling, and a lightweight scheduling algorithm, MSSR effectively balances the mitigation of forgetting with rapid adaptation to new tasks. Extensive experiments across three backbone models and eleven sequential tasks demonstrate that MSSR significantly outperforms current replay strategies, with particularly notable gains on reasoning-intensive and multiple-choice tasks.

catastrophic forgettingcontinual fine-tuningexperience replay

This work addresses the stability–plasticity dilemma in continual learning for large language model agents, which manifests as competition between old and new experiences during external memory retrieval. The authors propose a (k,v) memory framework that decouples experience representation from organization, enabling systematic investigation of memory reuse mechanisms in ALFWorld and BabyAI. Their findings reveal that external memory does not eliminate continual learning challenges but reframes them as issues of representation and retrieval design. Specifically, abstract procedural memories outperform detailed trajectories, while fine-grained memory organization exacerbates forgetting. Abstract memories also facilitate better transfer, though negative transfer disproportionately affects more difficult tasks. Notably, certain design choices that enhance forward transfer simultaneously induce severe catastrophic forgetting.

continual learningexperience reusememory retrieval

Towards LifeSpan Cognitive Systems

Sep 20, 2024
YW
Yu Wang
🏛️ UCSD | UIUC | Monash | NUS | AIWaves | MIT-IBM | UCLA

This paper addresses the lack of lifespan cognitive capabilities—such as rapid learning, long-term precise recall, and experience transfer—in large language models (LLMs). To this end, it proposes the first theoretical framework for a Lifespan Cognitive System (LSCS) tailored to high-frequency incremental interaction. Methodologically, it introduces a novel four-component synergistic paradigm integrating memory-augmented networks, dynamic sparse coding, neuro-symbolic representation, and online meta-learning, jointly modeling experience streams according to storage complexity. It further establishes dual core mechanisms—“experience absorption” and “response generation”—to overcome continuous learning bottlenecks. Contributions include: (1) formally characterizing the solvability boundaries of LSCS’s two fundamental challenges; (2) empirically validating, for the first time, both the necessity and feasibility of multi-technique coupling; and (3) providing a foundational theoretical framework and concrete technical pathway toward brain-inspired, lifelong-learning cognitive systems.

Continuous LearningLanguage ModelLife-Span Cognitive System

Latest Papers

What's happening recently
View more

This work addresses a key limitation in existing memory-augmented agents, which typically replay static experiences without accounting for semantic mismatches between retrieved memories and the current context, often leading to negative transfer. To overcome this, the paper proposes MemHarness, a novel framework that abandons the conventional replay paradigm in favor of a human-like memory reconstruction mechanism. Within this framework, a large language model agent critically reinterprets retrieved memories in light of the current state during decision-making, generating dynamic, context-adapted guidance signals. MemHarness integrates memory retrieval, critique, and reconstruction into a unified policy model trained end-to-end via GRPO. Experiments on ALFWorld and WebShop demonstrate substantial improvements over pure reinforcement learning and static memory baselines, with strong robustness in out-of-distribution scenarios and effective mitigation of negative transfer.

context alignmentexperience replaylarge language model agents

This work addresses the limitations of large language model agents in long-term interactions, where statelessness and constrained context windows hinder effective memory utilization. Existing memory retrieval approaches struggle to efficiently integrate heterogeneous memories and suffer from poor sample efficiency. To overcome these challenges, the paper proposes the Explore-Assimilate Reflection (EAR) framework, which first conducts exploratory reflection to iteratively retrieve initial memories and accumulate experiences, followed by assimilative reflection that replays an experience buffer to optimize a global reranker. EAR uniquely unifies exploration and assimilation mechanisms, achieving both high retrieval performance and strong sample adaptability. Experiments demonstrate that EAR improves memory retrieval accuracy by up to 17.9% over baselines on two long-term dialogue benchmarks, while exhibiting superior sample efficiency and robustness to noise.

autonomous agentsheterogeneous memorylong-term memory

Existing long-horizon multimodal memory agents lack the ability to diagnose retrieval failures and adapt their strategies accordingly. This work proposes a reflective retrieval memory framework that dynamically guides current query-based memory retrieval by distilling transferable retrieval strategies from historical task trajectories, ensuring that generated answers rely solely on newly retrieved video evidence. The framework integrates an entity-centric multimodal memory graph, a reflective experience memory mechanism, query-level guided generation, and a lifecycle management strategy that jointly considers usage frequency, feedback signals, and temporal decay to effectively reduce redundancy and noise. Experimental results demonstrate consistent performance gains over state-of-the-art methods across the M3-Bench-Robot, M3-Bench-Web, and Video-MME-Long benchmarks, validating the framework’s effectiveness and generalizability.

long-horizon multimodal reasoningmemory retrievalmultimodal memory agents

Hot Scholars

XD

Xiangyu Dong

Staff Software Engineer, Google
Computer architecture
ML

Meng Li

China University of Mining and Technology
Mining Engineering
TN

Tien N. Nguyen

Professor, School of Engineering and Computer Science - The University of Texas at Dallas
AI4SEAutomated Software EngineeringArtificial IntelligenceMining Software Repositories
YY

Yang You

Postdoc, Stanford University
3D visioncomputer graphicscomputational geometry
TZ

Tieying Zhang

Research Scientist at Bytedance
AI for SystemsSystems for AI