🤖 AI Summary
This study addresses the unreliability of environment construction for code agents caused by scattered requirements and implicit dependencies in execution environments. To this end, it proposes Graph2Env, a framework that introduces a typed dependency graph (DepGraph) to explicitly model environment requirements and states, thereby decoupling dependencies from interaction histories. By integrating a graph-based agent architecture with an execution feedback loop, Graph2Env enables the persistent storage and replay verification of successful repair experiences, generating reproducible build processes. Evaluated on a benchmark of 200 Python repositories, the framework achieves an environment construction success rate of 81.0% and a setup success rate of 59.3%, significantly outperforming state-of-the-art baselines.
📝 Abstract
Coding agents now increasingly rely on execution to validate their solutions, making the construction of reliable execution environments a critical enabling capability. However, repository environment construction is challenging because execution requirements are fragmented across repository artifacts and may only become apparent during execution. Existing agent-based approaches address this problem through iterative interaction, but information about the current construction state, including discovered requirements, satisfied and unresolved prerequisites, and their dependencies, can remain distributed across the interaction history. We present Graph2Env, an agent-based approach centered on DepGraph, a typed dependency graph that explicitly represents the environment requirements needed for repository execution, their dependency relations, and their states. Graph2Env uses DepGraph to guide environment construction and continuously refines it with execution feedback, while persisting successful repairs into a replayable construction procedure. The resulting artifacts are then applied in a fresh environment to verify that the constructed environment can be reproduced. We evaluate Graph2Env on a benchmark of 200 Python repositories drawn from RATBench and EnvBench, against a static dependency-inference baseline (pipreqs), three specialized environment-construction systems (Repo2Run, RAT, and SetupX), and two general-purpose coding agents (SWE-agent and Claude Code). Graph2Env achieves an 81.0% Environment Build Success Rate (EBSR) and a 59.3% Environment Setup Success Rate (ESSR), outperforming the strongest baseline by 9.5 and 9.0 percentage points, respectively.