AsynCodeBench: Benchmarking Collaboration of Asynchronous Multi-Agent Systems in Software Engineering

๐Ÿ“… 2026-09-26
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the absence of direct metrics for collaborative capability in asynchronous multi-agent systems, where collaboration is often confounded with individual coding proficiency. To this end, we propose a dependency graph-based benchmarking framework. By constructing an executable dependency checker and employing asynchronous trajectory analysis techniques, we introduce two quantitative metricsโ€”ADPR and DRSโ€”to enable explicit evaluation of cross-agent collaborative efficacy. Our investigation reveals a "jump window" phenomenon in collaboration, demonstrating that improvements in coding performance do not necessarily enhance collaborative effectiveness. Furthermore, successful coordination manifests as concentrated bursts rather than gradual progression. These findings effectively decouple coding from collaborative capabilities, providing a rigorous methodology for assessing true multi-agent cooperation independent of individual task competence.
๐Ÿ“ Abstract
Multi-agent coding has emerged as an increasingly active direction in software engineering, where complex development tasks are decomposed across multiple specialized agents working on different parts of the problem. Despite the shift from individual problem solving to distributed collaboration, multi-agent systems still lack a direct measure of collaboration and are largely evaluated through task-level outcomes inherited from single-agent coding, conflating individual coding capability with cross-agent coordination. We introduce AsynCodeBench, a dependency-centric benchmark for asynchronous multi-agent software engineering that represents each task with an explicit dependency graph and executable Dependency Checkers. Through this dependency-tracking process, we propose two complementary measures: Asynchronous Dependency Pass Rate (ADPR), which measures how many cross-agent dependencies are ultimately satisfied, and Dependency Resolution Step (DRS), which measures when each dependency first becomes satisfied during execution. AsynCodeBench comprises 19 tasks from real-world repositories, exposing 52 directed dependencies as explicit units for evaluating cross-agent collaboration. Experiments across model families, scales, and generations reveal a clear gap between coding and collaboration capability: improvements in coding performance do not necessarily translate into stronger collaboration, and task-level metrics can diverge substantially from dependency-level collaboration measures. Dependency-trajectory analysis further reveals that successful coordination often emerges not gradually, but through concentrated bursts in which many dependencies become resolved over a short portion of the execution trajectory, a pattern we term a hopping window.
Problem

Research questions and friction points this paper is trying to address.

Multi-agent systems
Software engineering
Collaboration evaluation
Asynchronous coordination
Benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Asynchronous Multi-Agent Systems
Dependency Graph
Collaboration Benchmark
Dependency Checkers
Hopping Window
๐Ÿ”Ž Similar Papers
No similar papers found.