Coding Agent Memory Post-training: Unlocking the Memory Potential of Pre-trained File Operations for Long-Horizon Tasks via Reinforcement Learning

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenges of memory failure due to context limitations and poor generalization caused by reliance on task-specific interfaces in long-horizon agent tasks. To overcome these issues, this work proposes a file-based persistent memory control strategy. Methodologically, it introduces CAMG, a novel environment suite that directly leverages the native file operations of pre-trained models as a memory mechanism, eliminating the need for additional tool interface definitions. Furthermore, an asynchronous Proximal Policy Optimization (PPO) algorithm is employed to jointly train Qwen3.5 series models across multiple environments. Experimental results demonstrate that smaller 4B and 9B models achieve performance comparable to 35B–122B large models on benchmarks such as SWE-bench, thereby realizing efficient memory augmentation and robust cross-task generalization.
πŸ“ Abstract
Language-model agents increasingly tackle long-horizon tasks whose interaction histories exceed the model's active context. Recent work has begun to use reinforcement learning to make memory control part of the policy, often relying on predefined memory tools within domain-specific training environments of relatively short horizons. This setup ties learned memory behavior to environment-specific interfaces that lie outside the base model's pre-training and must be learned from scratch, so even after post-training, agents struggle to use memory in long-horizon tasks. To address these limitations, we introduce Coding Agent Memory Gym (CAMG), a suite of long-horizon agentic-RL environments spanning Shop, Coding, DeepResearch, and AutoResearch. Alongside each environment's native task interface, CAMG provides executable shell access and an episode-persistent workspace, enabling agents to create, revise, search, and reuse files as memory throughout an episode. We also introduce CAMG-RL, which trains a single policy jointly across all four environments with fully asynchronous PPO, learning this file-based memory behavior directly from downstream task reward, and we train CAMG-RL-4B and CAMG-RL-9B from Qwen3.5 models of matching size. On SWE-bench Verified and MLE-bench Lite, CAMG-RL-4B and CAMG-RL-9B are competitive with Qwen3.5-35B-A3B and Qwen3.5-122B-A10B, respectively.
Problem

Research questions and friction points this paper is trying to address.

Coding Agent
Long-Horizon Tasks
Memory
Reinforcement Learning
Pre-trained File Operations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Long-Horizon Tasks
Agent Memory
File Operations
Asynchronous PPO
πŸ”Ž Similar Papers