LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study evaluates the end-to-end capabilities of coding agents in translating high-level designs into code implementations within large-scale software systems. To this end, we construct a multilingual benchmark spanning 29 real-world repositories and comprising 100 long-horizon proposal tasks. This benchmark enables the first joint assessment of agents' intent comprehension and code implementation proficiency, incorporating large-scale change analysis and automated pull request verification. Experimental results reveal that the strongest agent resolves only 14% of tasks; however, augmenting agents with file-tree context substantially increases the success rate to 34%, indicating that code localization remains the primary bottleneck for current systems. By addressing the gap in full-pipeline evaluation, this work provides critical insights for advancing the engineering capabilities of coding agents.
📝 Abstract
Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents' implementation capability to produce correct code edits from detailed specifications. However, practical modular development tasks also require the perception capability of grounding user intent and high-level design to derive a specification. We introduce LoLBench to evaluate both capabilities through the entire proposal-to-implementation process on large software systems. It is a multilingual benchmark of 100 tasks across 29 software systems in five domains. Each task provides a human-written enhancement proposal with user intent and high-level design. On average, proposals contain about 5,000 words, software systems contain 2.4 million source lines of code (LoC), and implementation pull requests (PRs) change approximately 5,500 LoC. Across 28 agents we evaluated, the best agent resolves only 14% of tasks and achieves a 52.7% Fail-to-Pass (F2P) pass rate. Failure analysis identifies incomplete code localization as a major bottleneck, while providing reference-derived file trees alongside API specifications improves resolved rates by 16--22 percentage points (2.4--17$\times$), reaching at most 34%. These results show that both perception and implementation remain central challenges for coding agents in practical modular development on large software systems. LoLBench is available at https://huggingface.co/datasets/lolbench26/LoLBench.
Innovation

Methods, ideas, or system contributions that make the work stand out.

coding agents
benchmark
long-horizon tasks
perception capability
code localization
Y
Yun Peng
Fudan University, China
Z
Zihan Wu
City University of Hong Kong, Hong Kong
Z
Zeyang Zhuang
Chinese University of Hong Kong, Hong Kong
X
Xin Zhou
Singapore Management University, Singapore
Rui Shu
Rui Shu
OpenAI
Machine LearningComputer VisionArtificial IntelligenceGenerative Models
X
Xu Han
HKUST (GZ), China
Chun Yong Chong
Chun Yong Chong
Monash University
Software Engineering
Y
Yuan Wang
Independent Researcher, Hong Kong
Jiakun Liu
Jiakun Liu
Harbin Institute of Technology
Empirical Software EngineeringIntelligent Software Engineering