Zero2Repo: Can Coding Agents Build Repositories from Scratch?

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the reliance on manual annotation in existing multilingual code generation benchmarks and their inadequacy in evaluating full-repository construction capabilities by proposing an automated evaluation framework. The framework leverages a language-agnostic task pipeline and open-source project conversion techniques to automatically generate multilingual tasks comprising requirement documents, interface contracts, and hidden tests. Adversarial validation, sandboxed container execution, and binary reward mechanisms are integrated to ensure objective and reliable assessment. Experimental results demonstrate that state-of-the-art models remain significantly limited in delivering complete repositories within native ecosystems, with failures primarily attributable to omitted specification details. This work provides clear directions for improvement and a robust benchmark for enhancing the precise implementation capabilities of AI agents.
πŸ“ Abstract
Coding agents are increasingly asked to build software rather than patch it, yet benchmarks for from-scratch repository construction are mostly limited to a single language and depend on manually curated tasks. We introduce Zero2Repo, a benchmark in which an agent receives a product requirements document, an interface contract, and an empty workspace, and must deliver a complete repository in the project's native ecosystem. Tasks are produced by a language-agnostic authoring pipeline that converts real, version-pinned open-source projects into behavioral specifications, reproducible environments, and hidden acceptance tests. Each task is validated by execution: a reference implementation derived from the upstream project must pass, and adversarial validation must show that the tests reject incorrect implementations. Evaluation runs production coding agents in isolated containers, withholds the acceptance tests until an explicit submission, and assigns a binary reward only when every test passes, with no LLM judge. The pipeline and harness make no language-specific assumptions and apply to mainstream programming ecosystems; the current release contains Python, TypeScript, Go, and C++ tasks. Even on 11 tasks drawn from repositories that frontier models have very likely seen during training, the strongest agent solves only 10, and every failing submission passes 90-99% of the hidden tests; for the two strongest agents, 67-100% of failed tests trace to a single omission or a low-frequency rule stated in the specification rather than to a missing subsystem, so each failure is a concrete target for improvement.
Problem

Research questions and friction points this paper is trying to address.

coding agents
repository construction
benchmark evaluation
software engineering
from-scratch generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Coding Agents
Benchmark
Language-Agnostic Pipeline
Execution-Based Evaluation
Repository Construction
πŸ”Ž Similar Papers
No similar papers found.
P
Pei Yang
Gradient Data
Tianyu Shi
Tianyu Shi
University of Toronto
Reinforcement learningIntelligent Transportation SystemLarge Language ModelsAILLM agent
Yuhang Yao
Yuhang Yao
Carnegie Mellon University
Federated Graph LearningSuper Agent System
W
Wanyi Chen
Independent Researcher
Tongyun Yang
Tongyun Yang
Marie-Curie PhD Fellow @ IMDEA Networks
Applied Machine LearningIntegrated Sensing & CommunicationComputer Vision
D
Dun Pei
Independent Researcher
H
Haonan Wang
Independent Researcher
Pengbin Feng
Pengbin Feng
Xidian University
Malware detectionVulnerability detectionBinary analysis
G
Guanxu Yu
Independent Researcher
J
Jingchun Huang
Independent Researcher
Z
Zeyu Zhang
Independent Researcher
S
Shuhan Sun
University of The Cumberlands
H
Hao Li
Queen’s University
Alex Gu
Alex Gu
MIT
program synthesismachine learninglarge language modelscode generation
X
Xiang Li
University College London, University of London
Jie Xiao
Jie Xiao
University of Science and Technology of China
low level visiongenerative modelmachine learning
Xinyu Wang
Xinyu Wang
PhD student, McGill University
Large Language ModelRetrieval Augmented GenerationQuantization
H
Hankxin Chen
University of California, San Diego
D
Daqi Li
Independent Researcher
Q
Qi Jia
Independent Researcher
H
Hongshan Lin
Independent Researcher
Z
Zhizhou Gu
Independent Researcher
Z
Zijun Tian
Independent Researcher
W
Weizhi Du
Independent Researcher
L
Lynn Ai
Gradient Data
Eric Yang
Eric Yang
AI Scientist, Verily Life Sciences