memory distillation

Design, build, and evaluate methods that compress, summarize, and transfer sequences of interactions or procedural experiences into compact, fixed‑size memory representations usable by downstream models. This work covers teacher–student distillation, procedural memory learning, cross‑executor and third‑party summarization to align student memories to teacher representations, synthesize or compare multiple trajectories into executor‑independent candidate memories, and reduce noise or bias while training recurrent models with fixed‑size targets.

memorydistillation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Current AI agents exhibit significant degradation in procedural memory retrieval when encountering novel tasks containing unseen vocabulary, revealing a fundamental bottleneck in cross-instance recognition of functionally equivalent programs. Method: We introduce the first diagnostic benchmark for procedural memory retrieval in language agents, decoupling memory retrieval from task execution to isolate and rigorously evaluate agents’ understanding of functional equivalence. Leveraging ALFWorld, we construct a dual-source trajectory corpus—comprising expert-annotated and LLM-generated demonstrations—and design a hierarchical query protocol to systematically assess six retrieval methods via controlled ablation studies. Results: State-of-the-art embedding methods suffer sharp performance drops in novel scenarios, exposing inherent limitations in modeling temporal structure and cross-context generalization. In contrast, LLM-generated program abstractions demonstrate superior generalizability: their transfer stability and scalability with corpus size substantially outperform representation-learning enhancements, highlighting abstraction—not just embedding—as a critical axis for advancing procedural memory.

Evaluates agents' ability to recognize functionally equivalent procedures across different object instantiations.Exposes generalization cliff in embedding-based methods when handling novel tasks and vocabularies.Separates genuine procedural understanding from surface-level memorization in language agents.

This work addresses the construction of human-like memory mechanisms in large language models and multimodal foundation models to support continual learning, personalized reasoning, and cross-modal consistency. It proposes the first unified taxonomic framework that systematically integrates three major memory paradigms: implicit memory (e.g., parameterized memory), explicit memory (e.g., external retrieval and graph-structured knowledge bases), and agent memory. The framework is further extended to multimodal settings, elucidating the critical role of memory in cross-modal alignment and agent collaboration. Through a comprehensive literature review and taxonomic analysis, the study surveys existing technical approaches, evaluation benchmarks, and core challenges, thereby establishing a theoretical foundation and offering clear directions for future research on human-like memory systems in artificial intelligence.

autonomous agentscontinual learningLarge Language Models

This work addresses the limited tool-use capability of small-scale language model agents, which stems from their difficulty in generating sufficient successful execution trajectories. To overcome this, the authors propose a training-free hierarchical memory distillation framework that systematically transfers structured knowledge from a large-model teacher to a small-model student. The framework encodes successful experiences into three types of memory: task-level workflows, subtask behavior examples, and function-calling specifications, and integrates both proactive injection prior to task execution and passive retrieval during inference. This approach achieves the first structured knowledge transfer tailored for small-agent models, yielding significant performance gains—average accuracy improvements of 27.2%, 11.2%, and 3.4% across three tool-use benchmarks—outperforming existing memory-augmentation methods, with the most pronounced gains observed in 4B-scale models.

agent memoryknowledge transfersmall language models

Large language model (LLM) agents often struggle to reuse past experiences in recurring scenarios, leading to computational redundancy and behavioral instability. To address this, this work proposes the Skill-MDP framework, which converts interaction histories into executable skills and introduces a non-parametric variant of Proximal Policy Optimization (PPO). By integrating semantic gradient-guided skill generation with a PPO-based gating mechanism, the approach enables procedural memory learning and validation without requiring parameter updates. Coupled with a score-driven memory maintenance strategy, the method significantly enhances skill reuse and task performance across in-domain, cross-task, and cross-agent settings, while achieving extremely high compression ratios in storing procedural memories.

computational redundancyexecution instabilityexperience reuse

Recite, Reconstruct, Recollect: Memorization in LMs as a Multifaceted Phenomenon

Jun 25, 2024
US
USVSN Sai Prashanth
🏛️ EleutherAI | Microsoft | DatologyAI | New York University | Indraprastha Institute of Information Technology Delhi | Northeastern University | Google DeepMind | University of Illinois at Urbana-Champaign | Harvard University | Kempner Institute

Language model memory is often oversimplified as a homogeneous phenomenon, neglecting sample-specific characteristics and heterogeneous model–corpus interactions. Method: We propose the *memory heterogeneity hypothesis*, decomposing memory into three distinct types: *parroting* (highly repetitive sequences), *reconstructing* (highly predictable sequences), and *reminiscing* (low-repetition, low-predictability sequences). Leveraging causal-inspired feature engineering, we develop a category-aware, multi-factor logistic regression model for interpretable, cross-category attribution. Contribution/Results: This work establishes the first memory taxonomy driven jointly by sample attributes and model–corpus co-adaptation. We identify distinct dominant factors per type—e.g., repetition rate for parroting, local entropy for reminiscing—and achieve an AUC of 0.89, significantly outperforming homogeneous baselines. Our framework enables fine-grained, mechanistic understanding of memorization behavior in large language models.

Classify memorization into recitation, reconstruction, recollectionModel memorization as multifaceted, not homogeneousPredict memorization using taxonomic category factors

Latest Papers

What's happening recently
View more

This work addresses the challenge of deploying large language model (LLM) agents on resource-constrained devices, where complex procedural tasks are hindered by reliance on massive models, long contexts, and iterative reasoning. To overcome this, the authors propose DuoMem, a novel framework that synergistically combines context-space and parameter-space distillation. By precomputing high-quality teacher memories and integrating them with LoRA-based fine-tuning, DuoMem efficiently transfers the problem-solving capabilities of a large teacher model to a lightweight student model. Experiments in embodied decision-making environments such as ALFWorld demonstrate that the method boosts the task success rate of a 4B-parameter student model from 4.3% to 77.9%—approaching the 87.1% performance of a 72B-parameter teacher—while introducing fewer than 10 million trainable parameters and achieving over 3× inference speedup.

large language modelsmemory-augmented agentson-device deployment

This work addresses the challenges large language models face in long-horizon tasks regarding knowledge retention, organization, and reuse. Existing approaches are often limited to static fact storage or replay of successful experiences, struggling to incorporate failure cases and lacking online extensibility. To overcome these limitations, the paper proposes a unified memory framework that, for the first time, integrates semantic, episodic, and procedural memory within a dual-layer short- and long-term storage architecture. A multi-agent design—comprising actor, memory, and critic components—enables automatic memory generation, reward-based annotation, and adaptive retrieval. A reward-driven long-term memory management strategy, including evaluation, consolidation, and pruning, facilitates continual learning and online evolution. Experiments demonstrate that the proposed method significantly improves success rates and robustness across diverse complex long-horizon tasks, outperforming current baselines.

LLM-based agentslong-horizon tasksmemory framework

Existing large language model agents lack systematic evaluation of procedural memory transfer across tasks, roles, and models. This work proposes the AFTER benchmark, comprising 382 real-world enterprise tasks, to establish the first procedural memory evaluation framework tailored for enterprise-level workflows. The study reveals a fundamental trade-off between general-purpose and specialized skills and introduces a multi-model execution trajectory fusion strategy. Experimental results demonstrate that single-round skill optimization improves performance by 3.7–6.7 points, while the fused multi-model approach achieves 73.1% accuracy in cross-model evaluations, significantly outperforming individual models.

cross-task generalizationenterprise tasksLLM agents

Hot Scholars

YZ

Yue Zhao

Assistant Professor of Computer Science, University of Southern California
Anomaly DetectionOut-of-Distribution DetectionTrustworthy AIAI for Science
WH

Wen Huang

Tsinghua University
Generative model
WS

Weiyan Shi

Northeastern University
Natural Language ProcessingPersuasionDialogue systemsAI Safety
KL

Kaiwen Liu

University of Michigan
Control TheoryRoboticsMachine LearningHuman-Robot Interactions
TH

Thorsten Holz

Max Planck Institute for Security and Privacy (MPI-SP)
Computer Security