Towards In-Parameter Memory Augmentation for Large Language Models

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited context length and redundant encoding overhead of large language models (LLMs) by systematically surveying parameterized memory augmentation methods that embed reusable knowledge into model parameters or adapters to extend inference-time capabilities. It innovatively proposes a dual-axis taxonomy based on "parameter placement location" and "acquisition timing (online/offline)," comprehensively reviewing parameter injection techniques across embedding, attention, and feed-forward network layers, alongside adapter fine-tuning strategies. Ultimately, this work constructs a complete technical landscape of parameterized memory, providing systematic guidance for LLM knowledge updating, safety compliance, and synergistic design with in-context learning.
📝 Abstract
Recently Large Language Models (LLMs) and LLM-based agents increasingly need to incorporate knowledge acquired after pretraining, e.g., domain facts, user preferences, documents, and interaction experience. In-context learning (ICL) and ICL-based agent harness remain flexible, but they consume context capacity and incur repeated discretized encoding cost that grows with context length. \textbf{In-parameter memory} offers a complementary substrate: reusable memory information is represented in model parameters, adapters, or other parameter-like objects that are composed into the forward pass at inference time. This survey focuses on methods that augment LLMs with such parametric memory at deployment: a memory-bearing parameter object is plugged into the forward pass during inference, whether it is acquired before or during deployment. We organize the landscape with two orthogonal axes: \textbf{Parameter Placement}, which includes Embedding, Attention, FFN layers, or Hybrid when two or more layers are used; and \textbf{Parameter Acquisition Time}, which distinguishes methods whose memory object is acquired during deployment (online) from those acquired before it (offline). We clarify boundaries, conduct comparisons, and discuss open directions in interference, safety, co-design with ICL, and recursive self-improvement.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
In-Parameter Memory
Knowledge Augmentation
In-Context Learning
Deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-Parameter Memory
Large Language Models
Memory Augmentation
Parameter Placement
In-Context Learning
🔎 Similar Papers