MemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM Agents

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the "memory-action gap" in long-horizon large language model agents, which arises from suboptimal memory organization and presentation. To bridge this gap, the authors propose MemArbiter, a framework that introduces a function-aware dynamic memory arbitration mechanism at decision time. It decomposes interaction history into atomic memory items and organizes them into five functional memory banks. Memory salience is then modulated through a multi-dimensional process incorporating demand assessment, relevance scoring, focus-context representation, and temporal gating. Experimental results on ALFWorld demonstrate that MemArbiter achieves success rates of 82.8% and 92.5% under token budgets of 500 and 750, respectively—surpassing the strongest baseline by 20.9 and 25.4 percentage points—and effectively mitigates issues of repetitive failures and state-action loops.
📝 Abstract
Large language model (LLM) agents must retain and use cross-step information to act coherently in long-horizon tasks. Existing methods improve memory accessibility, yet action-relevant information may still fail to guide the current decision because it is poorly formed, organized, prioritized, or presented. We call this post-access failure the Memory-Action Gap. We propose MemArbiter, a function-aware memory arbitration framework that addresses the memory-management-induced component of this gap. MemArbiter decomposes interaction histories into atomic items, organizes them into five functional Memory Banks, and combines bank-level demand, item-level relevance, focal-ambient representations, and a temporal presentation gate to dynamically control memory salience. We evaluate MemArbiter on ALFWorld against Flat Retrieval and Flat Recency under unified per-step memory budgets. With an open-weight action-generation model, MemArbiter achieves success rates of 82.8% and 92.5% under 500- and 750-token budgets, outperforming the strongest baseline by 20.9 and 25.4 percentage points, respectively. It also improves post-failure recovery and reduces failed-action repetition and state-action recurrence. These results show that function-aware memory arbitration enables accessible information to guide actions more effectively.
Problem

Research questions and friction points this paper is trying to address.

Memory-Action Gap
long-horizon tasks
memory arbitration
LLM agents
decision-time memory
Innovation

Methods, ideas, or system contributions that make the work stand out.

Memory Arbitration
Function-Aware Memory
Memory-Action Gap
Long-Horizon LLM Agents
Dynamic Memory Salience