Score
Designs and implements memory and replay systems that store, organize, and represent past experiences at multiple granularities—from fine-grained transitions to aggregated episodes and abstract summaries—using structures such as hierarchical pools, banks, or indexed memories. Creates aggregation, indexing, and retrieval policies that enable coarse-to-fine replay for planning and learning, including matching stored items to degradation or applicability patterns, prioritizing or ordering removals, and producing representations suitable for efficient lookup and replay.
To address catastrophic forgetting in memory-constrained streaming learning, this paper proposes a unified continual learning framework based on state replay, applicable to generative (autoencoding), time-series forecasting, and classification tasks. Unlike naive sequential fine-tuning or black-box replay, we formulate state replay as a joint optimization objective, enabling cooperative parameter updates over new and replayed samples via stochastic gradient methods; theoretical analysis grounded in gradient alignment reveals necessary conditions for effective forgetting mitigation. Evaluated across six heterogeneous and stationary streaming settings—constructed from Rotated MNIST, Electricity, and Airlines datasets—the method reduces average forgetting by 2–3× under heterogeneous multi-task streams, while matching fine-tuning performance on stationary streams. Our key contributions are: (i) the first theoretical analysis framework for state replay that unifies generative and discriminative tasks, and (ii) empirical validation of its effectiveness and robustness as a strong baseline for streaming continual learning.
This work addresses the challenge of catastrophic forgetting in large language models during continual fine-tuning, a problem inadequately mitigated by existing replay methods that either rely on heuristic rules or incur high computational costs. Inspired by human memory mechanisms, the authors propose the Memory Strength–aware Sample Replay (MSSR) framework, which dynamically optimizes both the content and timing of replayed samples through sample-level memory strength estimation and adaptive replay scheduling. By integrating experience replay, memory strength modeling, and a lightweight scheduling algorithm, MSSR effectively balances the mitigation of forgetting with rapid adaptation to new tasks. Extensive experiments across three backbone models and eleven sequential tasks demonstrate that MSSR significantly outperforms current replay strategies, with particularly notable gains on reasoning-intensive and multiple-choice tasks.
This work addresses the stability–plasticity dilemma in continual learning for large language model agents, which manifests as competition between old and new experiences during external memory retrieval. The authors propose a (k,v) memory framework that decouples experience representation from organization, enabling systematic investigation of memory reuse mechanisms in ALFWorld and BabyAI. Their findings reveal that external memory does not eliminate continual learning challenges but reframes them as issues of representation and retrieval design. Specifically, abstract procedural memories outperform detailed trajectories, while fine-grained memory organization exacerbates forgetting. Abstract memories also facilitate better transfer, though negative transfer disproportionately affects more difficult tasks. Notably, certain design choices that enhance forward transfer simultaneously induce severe catastrophic forgetting.
This paper addresses the lack of lifespan cognitive capabilities—such as rapid learning, long-term precise recall, and experience transfer—in large language models (LLMs). To this end, it proposes the first theoretical framework for a Lifespan Cognitive System (LSCS) tailored to high-frequency incremental interaction. Methodologically, it introduces a novel four-component synergistic paradigm integrating memory-augmented networks, dynamic sparse coding, neuro-symbolic representation, and online meta-learning, jointly modeling experience streams according to storage complexity. It further establishes dual core mechanisms—“experience absorption” and “response generation”—to overcome continuous learning bottlenecks. Contributions include: (1) formally characterizing the solvability boundaries of LSCS’s two fundamental challenges; (2) empirically validating, for the first time, both the necessity and feasibility of multi-technique coupling; and (3) providing a foundational theoretical framework and concrete technical pathway toward brain-inspired, lifelong-learning cognitive systems.
This work addresses a key limitation in existing memory-augmented agents, which typically replay static experiences without accounting for semantic mismatches between retrieved memories and the current context, often leading to negative transfer. To overcome this, the paper proposes MemHarness, a novel framework that abandons the conventional replay paradigm in favor of a human-like memory reconstruction mechanism. Within this framework, a large language model agent critically reinterprets retrieved memories in light of the current state during decision-making, generating dynamic, context-adapted guidance signals. MemHarness integrates memory retrieval, critique, and reconstruction into a unified policy model trained end-to-end via GRPO. Experiments on ALFWorld and WebShop demonstrate substantial improvements over pure reinforcement learning and static memory baselines, with strong robustness in out-of-distribution scenarios and effective mitigation of negative transfer.
This work addresses the limitations of large language model agents in long-term interactions, where statelessness and constrained context windows hinder effective memory utilization. Existing memory retrieval approaches struggle to efficiently integrate heterogeneous memories and suffer from poor sample efficiency. To overcome these challenges, the paper proposes the Explore-Assimilate Reflection (EAR) framework, which first conducts exploratory reflection to iteratively retrieve initial memories and accumulate experiences, followed by assimilative reflection that replays an experience buffer to optimize a global reranker. EAR uniquely unifies exploration and assimilation mechanisms, achieving both high retrieval performance and strong sample adaptability. Experiments demonstrate that EAR improves memory retrieval accuracy by up to 17.9% over baselines on two long-term dialogue benchmarks, while exhibiting superior sample efficiency and robustness to noise.
Existing long-horizon multimodal memory agents lack the ability to diagnose retrieval failures and adapt their strategies accordingly. This work proposes a reflective retrieval memory framework that dynamically guides current query-based memory retrieval by distilling transferable retrieval strategies from historical task trajectories, ensuring that generated answers rely solely on newly retrieved video evidence. The framework integrates an entity-centric multimodal memory graph, a reflective experience memory mechanism, query-level guided generation, and a lifecycle management strategy that jointly considers usage frequency, feedback signals, and temporal decay to effectively reduce redundancy and noise. Experimental results demonstrate consistent performance gains over state-of-the-art methods across the M3-Bench-Robot, M3-Bench-Web, and Video-MME-Long benchmarks, validating the framework’s effectiveness and generalizability.