🤖 AI Summary
This study addresses the limitations in correctness and misleading test feedback inherent to repository-level code generation by large language models. We propose a test-first reasoning framework coupled with a dual-track iterative refinement mechanism for code and tests. By integrating test-driven development into the generation process, this approach transforms tests from static validators into dynamic reasoning carriers, enabling co-evolution through an execution feedback loop. Evaluations on the RepoEval benchmark demonstrate that our method significantly outperforms retrieval-augmented and agent-based baselines. It achieves comprehensive improvements across pass rates, coverage metrics, and mutation scores, effectively enhancing both code reliability and test quality in complex software engineering tasks.
📝 Abstract
Large Language Models (LLMs) have achieved remarkable progress in code generation, yet ensuring correctness in complex, repository-level tasks remains challenging. Existing approaches often use generated tests as static post-hoc validators, which limits their ability to guide implementation and may introduce misleading feedback when the tests themselves are incomplete or incorrect. In this paper, we introduce TDD-Agent, which operationalizes the test-driven development paradigm for code generation. TDD-Agent first prompts the model to generate executable tests, encouraging it to clarify expected behaviors before implementation, and then performs iterative dual-track refinement over both the generated code and tests using execution feedback. We first isolate the effect of test-first reasoning through a prompt variant TDD-prompt on LiveCodeBench, where it consistently improves upon reasoning-based prompting baselines. Building on this finding, we evaluate the full TDD-Agent framework on RepoEval, a repository-level benchmark, and show that it consistently outperforms retrieval-based and agent-based baselines. Additional analyses show that iterative refinement improves not only code correctness but also the effectiveness of the generated tests, yielding higher pass rates, coverage, and mutation scores, suggesting that tests can serve as evolving reasoning artifacts rather than fixed validators. Our source code is available at https://anonymous.4open.science/r/TDD-Agent-Framework-6370/.