MemTrace: State-Consistent Memory for Long-Horizon Coding Agents

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of constrained context budgets in long-horizon coding agents and the invalidation of historical evidence caused by repository modifications. To tackle these issues, this work proposes a provenance-aware memory system that organizes dependencies through an immutable trajectory graph and memory anchors. By incorporating a dynamic validity verification mechanism, the system enables state-consistent memory reconstruction during iterative cross-file repairs. Evaluated on benchmarks such as DeepSWE, the proposed approach improves pass rates by up to 21.2 percentage points over baselines, significantly enhancing the reliability and efficiency of agents in complex software engineering tasks.
📝 Abstract
As coding agents take on long-horizon software evolution tasks spanning multiple files and stages, longer execution trajectories introduce two coupled challenges: (1) accumulated histories strain context budgets, and (2) repository changes can invalidate earlier execution evidence. Existing approaches address these challenges through techniques like larger context windows, compression, retrieval, or repository representations, but often fail to reconstruct a consistent task state after a context refresh or verify whether recalled evidence remains valid. Thus, we introduce MemTrace, a provenance-aware memory system that preserves execution history and aligns its reuse with the evolving task (e.g., iterative cross-file repair) and repository state. MemTrace stores history as immutable Memory Traces anchored to key information (e.g., files, symbols, tests), and organizes their execution order and dependencies in a Memory Trace Graph. When context is constrained, working memory retains only compact Memory Anchors, from which the agent can reconstruct the latest execution state and locate evidence relevant to its next action. Before restoring historical evidence, MemTrace checks its validity against the current repository state and retrieves only what the next action requires. Across three complementary long-horizon coding benchmarks, MemTrace consistently outperforms all fully evaluated baselines under the same backbone and harness, improving DeepSWE pass@1 by 21.2 points, SWE-EVO Resolved Rate by 4.4 points, and SWE-Milestone Score by 17.8 points under Codex CLI.
Problem

Research questions and friction points this paper is trying to address.

coding agents
long-horizon tasks
context budget
state consistency
evidence validity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Provenance-Aware Memory
Memory Trace Graph
Memory Anchors
Long-Horizon Coding Agents
State Consistency
H
Hongming Xu
Shanghai Jiao Tong University; MemTensor (Shanghai) Technology Co., Ltd.; Theseus Lab
L
Le Zhou
Shanghai Jiao Tong University; MemTensor (Shanghai) Technology Co., Ltd.; Theseus Lab
Z
ZhongHe Jin
Shanghai Jiao Tong University; Theseus Lab
X
Xiang Zhang
Shanghai Jiao Tong University; Theseus Lab
B
Bo Tang
MemTensor (Shanghai) Technology Co., Ltd.
Zhiyu Li
Zhiyu Li
Tianjin University
Robust controlattitude control
Xuanhe Zhou
Xuanhe Zhou
Assistant Professor, Shanghai Jiao Tong University
Data ManagementArtificial Intelligence
J
Juncheng Zhang
MemTensor (Shanghai) Technology Co., Ltd.