Bookkeeping, Composition, or Unreachable Gold? Reading MemoryAgentBench's Conflict-Resolution Scores Against a Frozen Last-Write Resolver

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a critical flaw in the "selective forgetting" metric of MemoryAgentBench, where unreachable samples artificially inflate memory scores for long-context models. To rectify this, we propose an evaluation paradigm that replaces aggregated scoring with item-wise reachability partitioning. We further construct a zero-learning baseline by freezing the last write operation and employ BM25 reimplementation, dual-model comparison, and preregistered statistical methods to attribute evaluation errors. Our analysis reveals that the benchmark contains logically unsolvable noise. Moreover, we demonstrate that simple heuristic rules can resolve over 80% of tasks yet consistently fail on the unreachable subset. These findings necessitate a fundamental reassessment of current memory agent evaluations in complex reasoning scenarios.
📝 Abstract
MemoryAgentBench's Conflict Resolution split is read as measuring "selective forgetting". We execute the benchmark's own rule - the newest statement about a fact wins - as a zero-learning resolver frozen on one of the four fact lists. Under the official metric the rule answers 80.25% of the questions (74.5% on the three held-out lists). Of the rest, 67 items have a released gold that the last-write graph cannot reach but overwritten statements would ("The capital of India is New Delhi." superseded by "The capital of India is Grosseto."; gold New Delhi); such items are a third of the multi-hop questions at 262K. Two long-context models and our pre-registered approximate re-implementation of the benchmark's BM25 agent, one retained run per item and outcomes only, score 84.7%, 82.6% and 41.6% on the items the rule solves against 10.4%, 11.9% and 6.0% on those 67. The failures are a reachability split plus a small parser-scope residual; the per-item split, not the aggregate, is the unit at which a score here can be read.
Problem

Research questions and friction points this paper is trying to address.

MemoryAgentBench
conflict resolution
selective forgetting
benchmark evaluation
unreachable gold
Innovation

Methods, ideas, or system contributions that make the work stand out.

MemoryAgentBench
Conflict Resolution
Last-Write Resolver
Reachability Analysis
Benchmark Evaluation
💼 Related Jobs
No related jobs found.
E
Egor Pakhomov
Salesforce AI Research
Erik Nijkamp
Erik Nijkamp
Salesforce Research
Generative modeling