π€ AI Summary
This work addresses the limitation of existing code agent evaluations, which focus solely on functional correctness and fail to reveal whether agents effectively reuse existing code, leading to latent accumulation of structural redundancy. We propose RepoReuse, a multi-turn evaluation benchmark that constructs an automated task synthesis pipeline based on AST dependency graphs and guided evidence collection, introducing for the first time a cross-turn structural redundancy metric to quantify code reuse rates. Experiments demonstrate that agents progressively lose repository exploration capabilities across multiple interaction turns, with 50.8% of task chains exhibiting duplicated logicβa deficiency obscured by conventionally high pass rates. These findings underscore the necessity of incorporating code reuse into agent evaluation frameworks.
π Abstract
Coding agents are increasingly deployed for iterative development on real repositories, yet existing evaluation barely answers a basic question: \emph{do coding agents reuse existing code or reinvent the wheel?} The question matters: every duplicated implementation is a fix applied twice and agents produce code far faster than humans can audit, so redundancy accumulates unsupervised. Thus, we present \textbf{RepoReuse}, a multi-turn benchmark for auditing code reuse in real repositories, where requirements are revealed turn by turn and the workspace accumulates across turns. It is built by a fully automated pipeline combining AST-based dependency graphs, guided evidence collection, and execution-verified task synthesis, and scales readily to new repositories. Beyond pass rates, we measure the reuse rate together with recall and cross-turn structural redundancy. An audit over 3{,}000 turns shows that agents progressively stop exploring relevant repository code, reuse their own history less even when it is fully in the workspace, and leave duplicated logic in 50.8\% of task chains by turn~5---all while pass rates barely move. Such deficiencies are invisible to pass rates, underscoring the need to evaluate code generation beyond functional correctness.