AMBER: Training Long-Horizon Web Agents through Append-Only Memory
This study addresses the challenge of critical information loss in long-horizon web agents due to context overflow, where existing overwrite-based memory mechanisms struggle to retain facts and feedback under sparse rewards. We propose AMBER, a framework introducing a novel append-only free-form memory mechanism that enables agents to jointly learn reasoning, acting, and memory writing. This architecture ensures the permanent preservation of key evidence and facilitates end-to-end reinforcement training without costly supervised data. Experiments demonstrate that AMBER improves success rates by 4.09% over baselines on WebArena Lite and increases multi-turn task completion by 4.8%, outperforming methods reliant on extensive annotated data at significantly lower cost.