🤖 AI Summary
This study investigates the potential and limitations of dynamic concurrent execution strategies for long-horizon coding tasks. Through end-to-end comparative experiments and trajectory analysis on 354 tasks involving mainstream code agents—including Codex, Claude Code, and Kimi Code—the authors demonstrate that task orchestration mechanisms substantially outperform underlying model capabilities in complex development scenarios. Furthermore, the research systematically identifies thirteen concurrency-specific failure modes and delineates the precise boundary conditions under which concurrent strategies yield performance gains. These findings provide empirical foundations for optimizing agent scheduling and informing engineering practices in code generation systems.
📝 Abstract
As coding agents advance from bounded software engineering tasks toward long horizon development, dynamic concurrency offers a promising way to scale complex development tasks. Under this policy, agents decide during execution whether and how to spawn concurrent sub-agents. Model capability largely determines outcomes on shorter tasks, whereas long horizon development makes orchestration central to task completion. Existing work, focused on coding agent failures on shorter tasks or collaboration in predefined multiagent workflows, offers little insight into dynamic concurrency in frontier agents across task complexity. We study dynamic concurrency as an execution policy through controlled comparisons of matched Codex, Claude Code, and Kimi Code executions with the policy enabled or disabled. Across 354 tasks and 2,124 executions spanning a range of task complexities and execution horizons, we evaluate its end to end effects and scheduling behavior, and analyze matched trajectories to characterize 13 concurrency specific failure modes, 28 observable patterns, and the conditions under which it provides an advantage.