🤖 AI Summary
This work addresses the limitation of conventional language models, which tightly couple long-term memory and reasoning within a single parameter set, hindering independent scaling of memory capacity. The authors propose a decoupled architecture featuring a parameterized long-term memory module that can be scaled independently. By integrating distributed Faiss indexing with sparse batch-loaded kNN retrieval, they demonstrate—for the first time—the feasibility of independently scaling pretraining memory at scales of billions of parameters and 300 billion tokens. Experiments show that adding only a 6.9B-parameter memory module boosts the average score of Pythia-410M from 29.86 to 37.34, surpassing Pythia-12B while reducing total parameters by 39%. Similar gains are observed on the Qwen3 series, achieving an average improvement of over 9 points across domains and substantially enhancing parameter efficiency.
📝 Abstract
Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independently. Memory Decoder introduces a parametric long-term memory module but only studies it at a relatively small scale. In this work, we present Memory Decoder at Scale, scaling memory models up to 6.9B parameters and pretraining them on 300B tokens. At this data scale, the combined cost of indexing and search makes a standard Faiss pipeline infeasible. We address this bottleneck with a distributed pipeline for Faiss indexing and retrieval, together with sparse, batch-wise loading of kNN distributions. Across model scales, we find that allocating more parameters to memory yields a better parameter-performance tradeoff than scaling the base model alone. On 17 benchmarks, pairing a 6.9B general memory with Pythia-410M raises its average score from 29.86 to 37.34, surpassing Pythia-12B (37.24) with 39% fewer total parameters. For Qwen3 Base models ranging from 0.6B to 14B, 1.7B domain memories improve the average score across the three domains by more than 9 points at every scale. Overall, our results demonstrate that independently scaling pretrained memory offers a more parameter efficient path to improving language model performance.