LSTMem: Hierarchical Long Short-Term Online Memory for Large Language Models

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing online memory mechanisms, which fail to decouple historical accumulation from current retrieval, resulting in uncontrollable information access during long-horizon interactions with large language models. To overcome this, we propose LSTMem, a hierarchical online memory architecture that equips each layer of a frozen LLM with cell and hidden states. It employs input, forget, and output gates to independently regulate information storage and exposure. Furthermore, cross-layer forward propagation and block-end feedback mechanisms are designed to leverage high-level gradients for reverse optimization of lower-level states. Evaluations on Qwen3-4B demonstrate significant performance improvements on benchmarks such as MemoryAgentBench. These results validate that hierarchically decoupled memory outperforms associative memory and confirm that cross-layer propagation is essential for effective long-range modeling.
📝 Abstract
Large language models increasingly serve as long-horizon assistants and agents, where they must both accumulate information across interactions and make the relevant parts available when later requests depend on them. Existing compact online memories typically use a single persistent state both to accumulate history and to serve readout, so what the memory stores cannot be controlled separately from what it exposes to the current computation. We propose LSTMem, an LSTM-inspired online memory that instead equips each layer of a frozen LLM with two matrix-valued states: a cell state that accumulates history and a hidden state whose readouts correct the backbone's attention. Input and forget gates control what the cell stores, while an output gate separately controls what the cell exposes through the hidden state. LSTMem further connects memory across depth through forward hidden-state propagation and block-end feedback, and uses higher-layer reconstruction gradients to refine lower-layer cell states before rebuilding hidden states from shallow to deep layers. Across memory benchmarks on Qwen3-4B-Instruct, LSTMem consistently improves MemoryAgentBench, LoCoMo, and HotpotQA over the plain backbone. Comparisons further show that the LSTM-based memory formulation outperforms an associative-memory counterpart, while removing cross-layer hidden-memory propagation degrades performance. These results demonstrate the benefits of separating memory accumulation from memory expression and organizing memory hierarchically across model depth. The code is available at https://github.com/Longchentong/LSTMem.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Online Memory
Long-horizon Agents
Memory Accumulation
Memory Readout
Innovation

Methods, ideas, or system contributions that make the work stand out.

Online Memory
LSTM
Hierarchical Memory
Large Language Models
State Decoupling
🔎 Similar Papers
No similar papers found.