🤖 AI Summary
Existing evaluations of code agents primarily focus on isolated tasks or final outcomes, failing to assess their capability in long-term, iterative software development. This work proposes the first long-horizon benchmark framework centered on cyclical engineering, modeling development tasks as directed acyclic graphs (DAGs) composed of independently testable units linked by source-evidence dependency edges. It introduces a flow-aware runtime that dynamically schedules tests and manages regression obligations. The benchmark encompasses 112 real-world tasks spanning eight programming languages and nine domains, comprising over 5,300 structured development units and associated test code. Even under the strongest configuration (Opus-4.7 + Claude Code), only a 25% task completion rate is achieved, underscoring the challenge and marking a significant departure from conventional static, endpoint-based evaluation paradigms.
📝 Abstract
Coding agent infrastructure is shifting from harness engineering toward loop engineering as coding agents are deployed for sustained long-horizon software development. Existing benchmarks often center on localized tasks or end-state outcomes, offering limited insight into sustained execution. We introduce LOOPSBENCH, a long-horizon benchmark for loop engineering in coding agent evaluation. Each task is a dependency DAG over separately testable development units with source-evidenced prerequisite edges. LOOPSBENCH comprises 112 tasks from authentic sources spanning 8 programming languages and 9 domains. Its flow-aware runtime releases tests along the ready frontier and retains completed nodes as regression obligations. We evaluate frontier coding agents paired with widely used loop implementations. The strongest configuration, Opus-4.7 with Claude Code and outer continuation, resolves 25.00% of tasks. Recorded plans recover only part of the source-recovered prerequisite DAG, and regression events remain visible across the evaluated loop profiles. We open source the benchmark data and code, including all tasks, more than 5,300 development units, and executable tests, at microsoft/Loopsbench.