Why Git Is the Memory Solution for the Agentic Development Lifecycle

📅 2026-07-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the frequent loss of reasoning behind code changes in agent development—such as trade-offs, constraints, and discarded alternatives—due to inadequate memory mechanisms. To remedy this, the paper proposes a novel paradigm that embeds reasoning traces directly into Git as a memory substrate, leveraging native version control operations like commits, merges, and code reviews to ensure authenticity, timeliness, and verifiability. Through a Git-anchored architecture featuring multi-path routing retrieval, structured mapping, confidence-gated snippet injection, and decision synthesis, the approach enables reproduction of results without additional annotations and achieves, for the first time, cross-session reconstruction of “why” reasoning chains. Evaluated on approximately 50,000 lines of production code, it attains an answer adequacy of 0.83 with outputs compressed by three orders of magnitude (382–980 tokens) and a seed retrieval MRR of 0.31, significantly outperforming baseline methods.
📝 Abstract
Coding agents now produce a growing share of a team's code, while the reasoning behind each change -- the alternatives weighed, the constraints discovered, the approaches rejected -- is trapped in assistant transcripts that vanish with the session. Memory for this setting, the agentic development lifecycle (ADLC), is usually posed as one retrieval problem and built as machinery: tiered stores, memory graphs, compiled wikis, model-judged admission. We argue memory should instead be git-bound -- built into the repository's version control, inheriting the guarantees the machinery struggles to construct: ground truth from commits, freshness from rebuild, verification from the merge, containment from review. On this ledger we solve two problems separately, then combine them. Seed supply is closed as an eight-corpus retrieval study under a pre-registered ship discipline: five imported ranking mechanisms rejected, two kept, and a best configuration of ~0.31 pooled MRR -- ~60x the raw-transcript grep floor, ~15x an honest parsed-turn floor. Answer assembly is where ranking stops helping: single-shot retrieval scores only 0.07-0.20 answer-sufficiency on real developer questions, and ungated episode injection measurably degrades good answers. A router dispatches breadth to a git-anchored structural map, pointed lookups to confidence-gated episodes, and rationale to decision synthesis, which reconstructs why-arcs no single session contains (0.83 sufficiency on a young ~50k-LOC production system). Routed, the system answers at 382-980 tokens per question -- three orders of magnitude below the recorded history. Because ground truth is mined from commit-session links rather than annotated, every result is replicable on any user's own history at zero labeling cost. The remaining constraint is capture. Code, benchmark, and paper source: github.com/rekal-dev/rekal-cli.
Problem

Research questions and friction points this paper is trying to address.

agentic development lifecycle
development memory
reasoning trace
version control
commit-session linkage
Innovation

Methods, ideas, or system contributions that make the work stand out.

git-bound memory
agentic development lifecycle
commit-session linking
memory routing
zero-label replication
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
F
Frank Guo
Rekal