AMBER: Training Long-Horizon Web Agents through Append-Only Memory

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of critical information loss in long-horizon web agents due to context overflow, where existing overwrite-based memory mechanisms struggle to retain facts and feedback under sparse rewards. We propose AMBER, a framework introducing a novel append-only free-form memory mechanism that enables agents to jointly learn reasoning, acting, and memory writing. This architecture ensures the permanent preservation of key evidence and facilitates end-to-end reinforcement training without costly supervised data. Experiments demonstrate that AMBER improves success rates by 4.09% over baselines on WebArena Lite and increases multi-turn task completion by 4.8%, outperforming methods reliant on extensive annotated data at significantly lower cost.
📝 Abstract
Modern language-model agents increasingly interact with external environments over long-horizon, multi-step trajectories, where the accumulated interaction history can quickly exceed practical context budgets. To ensure reliability, agents must maintain factual information over long horizons, remember execution errors and corrective feedback, and track progress across actions. Several approaches have been proposed to achieve this without the need for maintaining the entire execution history in context, such as using the reasoning and action history, learning to maintain a fixed-size memory through an overwrite mechanism, and periodic summarization. Although overwrite memory can in principle retain anything an append-only memory can, it must learn to carry each fact through every subsequent rewrite, which is difficult to learn from sparse outcome rewards; for interactive applications like web agents, we find that trained overwrite memories delete key information required by the trajectory, as well as corrective feedback received from the environment. We introduce AMBER (Append-only Memory Bank for Evidence Retention) - a simple and scalable framework where an agent jointly learns to reason, act, and write free-form memory, while an append-only rule guarantees retention by construction. This allows AMBER to be trained end-to-end with reinforcement learning from outcome rewards without the need for extensive curated SFT data. On WebArena Lite, AMBER improves average success over overwrite-based memory by 4.09 percentage points, increases the fraction of tasks solved in five repeated runs by 4.8 percentage points, and matches an overwrite baseline trained on substantially more expensive curated supervision. AMBER achieves these improvements while maintaining a practical token budget, providing a strong balance between context efficiency, task performance, and reliable long-horizon execution.
Problem

Research questions and friction points this paper is trying to address.

long-horizon web agents
context budget
memory retention
overwrite memory
Innovation

Methods, ideas, or system contributions that make the work stand out.

Append-Only Memory
Web Agents
Reinforcement Learning
Long-Horizon Execution
End-to-End Training
🔎 Similar Papers