VibeMemBench: Evaluating Memory Systems for Coding Agents on Real Repository Coding Tasks

📅 2026-09-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过VibeMemBench评估编码代理在真实代码库任务中使用持久内存系统的效果,揭示了现有内存系统与实际需求之间的差距。
📝 Abstract
Coding agents operate on real repository coding tasks, and persistent memory systems promise to reuse experience across tasks. Yet existing evaluations do not show whether those systems improve executable repository work. Repository benchmarks test code changes but do not isolate memory, while memory benchmarks score recall without measuring downstream coding outcomes. We introduce VibeMemBench, a benchmark for evaluating memory systems on 111 coding targets from 90 SWE-rebench V2 repositories and 3,634 history trajectories from the target repositories. The targets follow the SWE benchmark style and cover bug fixes, feature requests, interface changes, and configuration work. An agent edits each target codebase under a declared memory condition. Executable tests decide task resolution. Each target is retained only when injected history experience improves its executable outcome in a reference setting, so every target carries a prior experience whose usefulness is verified by execution in that setting. The frozen verified experience is then transferred to five held-out solvers. Direct injection raises observed task resolution on four of them by 1.1 to 4.5 percentage points while lowering agent steps on all five. Yet when four existing memory systems must construct and retrieve experience from the same history, eleven of twelve solver and system pairings fail to exceed the matched memory-off baseline. VibeMemBench exposes the gap between the useful experience that repository history holds and the experience existing memory systems deliver for repository coding tasks.
Problem

Research questions and friction points this paper is trying to address.

persistent memory systems
executable repository work
memory benchmarks
Innovation

Methods, ideas, or system contributions that make the work stand out.

VibeMemBench
memory systems
executable repository work
coding agents
history trajectories
🔎 Similar Papers
No similar papers found.
L
Liyang Fan
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences
Y
Yingcheng Shi
Alibaba Token Hub, Alibaba Group
Y
Yongbin Li
Alibaba Token Hub, Alibaba Group
C
Chenghao Sun
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences
X
Xin Chen
Alibaba Group
X
Xander Xu
Alibaba Group
H
Hu Wei
Alibaba Group
S
Shiwen Ni
SUAT
Min Yang
Min Yang
Bytedance
Vision Language ModelComputer VisionVideo Understanding
J
Jieping Ye
Alibaba Token Hub, Alibaba Group