HasMem: Hard-Origin Adaptively Softened Memory for Long-Term LLM Agents

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance degradation in long-horizon LLM agents caused by the coupling of memory capacity adjustment with frozen model readout. To mitigate this, we propose an Adaptive Softened Memory mechanism that optimizes cross-turn adaptation through controller-driven dynamic memory width regulation, Writer re-encoding, and global state-aware readout. Furthermore, hard-source initial state verification is introduced to decouple memory compression from readout operations. Evaluated on the MSC and LongMemEval-S benchmarks, our approach improves lexical F1 scores by 4.4 and 5.5 percentage points, respectively, while significantly reducing negative log-likelihood (NLL). By outperforming rule-based baselines, the proposed method effectively enhances long-term question-answering performance.
📝 Abstract
Text-based memory and context compression support reuse of past interactions. Resizing continuous memory changes the input to a frozen LLM, coupling capacity allocation with readout. We propose Hard-Origin Adaptively Softened Memory (HasMem). Frozen hard-prompt embeddings provide a verifiable initial state. A controller adjusts memory widths, a Writer re-encodes resized entries, and Reader and Global provide readout adaptation and cross-turn state. On all $535$ questions in a reconstruction probe derived from the Multi-Session Chat (MSC) development split, the main configuration achieves lexical F1 of $95.3$ ($+4.4$ percentage points) at $93.6\%$ of the hard reference's framed memory positions. With approximately matched per-question target body budgets, six configurations at mean per-entry retention around $0.83$--$0.91$ exceed rule-based re-encoding by $8.0$--$23.6$ exact-match (EM) percentage points. With fixed model parameters and rule target width ratio $0.75$, Global's EM gain passes a user-level exact paired test with Bonferroni correction over eight comparisons. On all $500$ LongMemEval-S questions, local lexical F1 rises from the hard reference's $3.4$ to $8.9$, and answer negative log-likelihood (NLL) falls from $12.257$ to $5.274$. F1 gains accompany lower EM on both evaluations.
Problem

Research questions and friction points this paper is trying to address.

Long-Term Memory
LLM Agents
Context Compression
Memory Resizing
Multi-Session Chat
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Memory
Context Compression
Long-Term Agents
Memory Resizing
Hard-Prompt Embeddings
Z
Zihong He
The Hong Kong University of Science and Technology (Guangzhou)
J
Junxiao Shen
University of Bristol
C
Chen Liang
The Hong Kong University of Science and Technology (Guangzhou)
Hai-Ning Liang
Hai-Ning Liang
The Hong Kong University of Science and Technology (Guangzhou)
VR/AR/MRGamesHCI