🤖 AI Summary
This study investigates whether persistent context files—such as AGENTS.md—enhance the task correctness of AI coding agents in real-world codebases. Through controlled ablation experiments on 17 authentic repository tasks using Claude Code and Codex, the work employs gold-standard testing, failure mode categorization, and equivalence testing to rigorously evaluate context injection strategies in a multi-agent, realistic setting for the first time. Results indicate that context files do not significantly improve correctness for either agent (with an upper bound of ≤15 percentage points) and fail to convert near-correct outputs into passing solutions. The primary cause of failure stems from insufficient implementation capability rather than missing knowledge. Furthermore, task difficulty exhibits agent-specific characteristics (Spearman ρ = 0.75), clarifying the source of contradictory findings in prior literature.
📝 Abstract
Persistent context files (AGENTS.md, CLAUDE.md) are standard practice for guiding AI coding agents, yet evidence for their effectiveness is contradictory. We present a controlled ablation of context-injection strategy across two frontier agents (Claude Code and Codex), 17 real tasks from 3 repositories (15 shared + 2 Codex-only), and 288 evaluated runs with gold-test evaluation. Context strategy does not measurably move correctness on either agent (bounded to <=10-15pp via equivalence testing). A failure-mode triage reveals why: agents fail on implementation skill---feature design, pattern selection, exact wiring---not missing repository knowledge that a context file could supply; a manipulation probe confirms the real AGENTS.md never converts a near-miss to a pass on either agent. We further show that borderline task difficulty is agent-specific (Spearman rho=0.75), offering a candidate explanation for prior contradictions: single-agent studies draw tasks from different agents' informative bands. We release all code, data, and analysis.