MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入MemCalib基准和MemCalib-RL算法,解决了大语言模型在使用记忆时过度或不足的问题,优化了模型响应的质量。
📝 Abstract
The effectiveness of agent memory ultimately depends on whether the underlying LLM gives each memory in context an appropriate degree of influence over its response. Yet this capability has remained largely overlooked. To assess this capability, we introduce MemCalib, a benchmark grounded in realistic memory-system scenarios for evaluating memory use and advancing optimization algorithms. Results on the MemCalib test set reveal that frontier open- and closed-source models struggle to use memory appropriately. They frequently over-use or under-use memory rather than matching each proposition's actual use to its target level, leading to biased, low-quality responses. Experiments with common post-training algorithms, including group relative policy optimization and on-policy self-distillation, further reveal a clear directional skew: trained models improve in one direction while deteriorating in the other. We therefore propose MemCalib-RL, an ordered bidirectional counterfactual credit-assignment algorithm that separates over- and under-use signals and localizes their credit to response tokens through exact atom ablation. Results across model families and scales (Qwen3-8B, Ministral-3-8B-Instruct, and Qwen3.5-35B-A3B) show that MemCalib-RL achieves the best overall performance while better balancing over-use and under-use, with gains generalizing beyond MemCalib in external benchmark evaluation. Further experiments support its design choices and robustness and provide insight into its training dynamics.
Problem

Research questions and friction points this paper is trying to address.

memory use
LLM agents
over-use
under-use
response quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

MemCalib
memory use optimization
bidirectional counterfactual credit-assignment
atom ablation
MemCalib-RL
🔎 Similar Papers
No similar papers found.
R
Ruike Cao
University of Science and Technology of China
F
Fanyu Zhao
Fudan University
F
Fugen Yao
Qwen Applications Business Group, Alibaba Group
L
Liang Dong
Qwen Applications Business Group, Alibaba Group
Jian Xu
Jian Xu
Senior Director, Ad Platform, Alibaba Group
Computational AdvertisingMachine LearningData MiningData Privacy
G
Guanjun Jiang
Qwen Applications Business Group, Alibaba Group
Yifei Zhao
Yifei Zhao
上海科技大学
H
Han Zhang
Fudan University
Li Xiao
Li Xiao
University of Science and Technology of China