π€ AI Summary
Current AI agents rely on external modules for memory, lacking an intrinsic, native memory mechanism within foundational models. This work proposes Metis, a memory-augmented foundation model that embeds persistent and dynamically evolving memory states directly into the model backbone, thereby introducing the first end-to-end trainable native memory capability in a foundation model. By integrating a novel memory attention mechanism, specialized large-scale training data, and a multi-objective intermediate training strategy, Metis enables memory updates during inference through a single forward passβwithout requiring gradient computation or weight adjustments. This design significantly enhances computational efficiency and architectural unity while effectively supporting long-term information retention and utilization.
π Abstract
Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal foundation models and large reasoning models. However, agent memory is still primarily implemented through external modules, leaving the native memory capability largely unexplored. In this paper, we take a first step toward this direction by introducing memory foundation models, which empower foundation models with native memory capabilities. We formalize native memory from two perspectives: a persistent and dynamically evolving memory state within the backbone, and native memory procedures that autonomously store and utilize information through model computation. We show that native memory offers advantages in architecture, end-to-end optimization, and efficiency. Based on this formulation, we propose Metis, the first prototype of memory foundation models. Metis introduces a new architecture that equips a foundation model with a native memory state, allowing historical information to be compressed into the model and accessed through memory attention. We construct large-scale memory-specific training data and introduce multiple optimization objectives to acquire these native memory procedures through mid-training. The online memory maintenance of Metis is gradient-free, and the memory update requires only a forward pass. At inference time, all learned model weights remain frozen, while the native memory states are autonomously transformed through standard forward computation. Through extensive experiments, we show that Metis exhibits native memory capabilities and further provide a detailed analysis of its strengths, limitations, and behaviors. To facilitate future research on memory foundation models, we release our project and model checkpoints.