🤖 AI Summary
This study addresses the engineering challenges encountered when large language models construct complete code repositories, including module coupling, interface conflicts, and complex dependency interactions. To overcome these limitations, this work proposes LEGO, a framework that introduces an agent-native reusable code primitive system. It employs persistent LLMs for relevance assessment and adaptive modification, combined with dependency closure resolution and cross-component constraint solving to activate, integrate, and debug primitives. Furthermore, the authors present LEGO-REPO, an end-to-end refactoring benchmark. Experimental results demonstrate an average performance improvement of 0.1474 across various models, with GPT-5.6-terra achieving a 61.4% gain. Notably, the framework maintains 95.1% homogenized evaluation scores at reduced computational cost, significantly enhancing the controllability of automated large-scale repository construction.
📝 Abstract
Large language models equipped with development environments have moved code generation toward repository-scale construction, yet building complete repositories remains difficult because interacting modules, interfaces, configurations, tests, and dependencies must work together. We introduce Code Primitives, agent-native reusable executable components with interface contracts, dependency closures, validation tests, and provenance. Each primitive uses a resident LLM to assess relevance and adapt its implementation, interfaces, and dependencies to the target repository, and we organize 1,424 validated primitives in CodeFace, a searchable library for repository construction. We introduce LEGO (Large-scale repository Engineering via aGent-native reusable cOde primitives), which activates task-relevant primitives, integrates their adapted implementations with task-specific code while resolving cross-component constraints, and revises the result against executed tests. To measure construction end to end, we build LEGO-REPO, a benchmark of 522 executable reconstruction tasks spanning seven software domains, 22 capability tracks, and five difficulty levels, scored against native test suites between an empty-package floor and original-source ceiling. The strongest of 13 evaluated backbones reaches a delivery score of 0.318 and scores zero on 41.0% of tasks; LEGO improves all 13 by 0.1474 on average and raises GPT-5.6-terra from 0.3180 to 0.5134 (+61.4%). In controlled comparisons, adapted primitives outperform retrieved code supplied as context or vendored unchanged. The effect persists against independent repository agents, across three external benchmarks, and with a disjointly re-mined CodeFace; GPT-OSS-20B for adaptation and diagnosis retains 95.1% of the homogeneous score at 24.0% lower cost.