🤖 AI Summary
This study evaluates the end-to-end capabilities of coding agents in translating high-level designs into code implementations within large-scale software systems. To this end, we construct a multilingual benchmark spanning 29 real-world repositories and comprising 100 long-horizon proposal tasks. This benchmark enables the first joint assessment of agents' intent comprehension and code implementation proficiency, incorporating large-scale change analysis and automated pull request verification. Experimental results reveal that the strongest agent resolves only 14% of tasks; however, augmenting agents with file-tree context substantially increases the success rate to 34%, indicating that code localization remains the primary bottleneck for current systems. By addressing the gap in full-pipeline evaluation, this work provides critical insights for advancing the engineering capabilities of coding agents.
📝 Abstract
Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents' implementation capability to produce correct code edits from detailed specifications. However, practical modular development tasks also require the perception capability of grounding user intent and high-level design to derive a specification. We introduce LoLBench to evaluate both capabilities through the entire proposal-to-implementation process on large software systems. It is a multilingual benchmark of 100 tasks across 29 software systems in five domains. Each task provides a human-written enhancement proposal with user intent and high-level design. On average, proposals contain about 5,000 words, software systems contain 2.4 million source lines of code (LoC), and implementation pull requests (PRs) change approximately 5,500 LoC. Across 28 agents we evaluated, the best agent resolves only 14% of tasks and achieves a 52.7% Fail-to-Pass (F2P) pass rate. Failure analysis identifies incomplete code localization as a major bottleneck, while providing reference-derived file trees alongside API specifications improves resolved rates by 16--22 percentage points (2.4--17$\times$), reaching at most 34%. These results show that both perception and implementation remain central challenges for coding agents in practical modular development on large software systems. LoLBench is available at https://huggingface.co/datasets/lolbench26/LoLBench.