🤖 AI Summary
This study addresses the evaluation bias arising from deviations between actual agent behaviors and predefined organizational structures in multi-agent coding systems. We propose a unified runtime architecture coupled with a fine-grained event tracing mechanism to enable programmable collaboration and execution control. Furthermore, we introduce an Adherence metric to quantify organizational compliance, revealing that single-dimensional variations significantly impact collaborative efficacy and establishing a controllable causal evaluation paradigm. Experimental results demonstrate that our dual-agent configuration achieves new state-of-the-art performance on mainstream benchmarks, while the single-agent configuration surpasses existing strong baselines with minimal token consumption.
📝 Abstract
Multi-agent coding systems are designed to tackle complex software engineering tasks through collaboration. However, existing evaluations typically assume configured organizations are followed faithfully, whereas reality differs. This behavioral gap, combined with differences in underlying system components, prevents clear attribution of observed gains. To this end, we introduce OpenCollab, a multi-agent coding framework that provides a unified infrastructure for programmable collaboration and controllable runtime. Specifically, OpenCollab unifies organization design, enforces experimental control on a shared runtime, and tracks execution through fine-grained event streams. On this basis, we define Adherence to quantify whether the declared organization is actually realized. Our experiments reveal that agents collaborate very differently across configurations: changing any single dimension shifts Adherence, from 47.2% to as high as 97.2%. Furthermore, extensive agentic coding benchmarks show that a two-coder workflow built on OpenCollab establishes new SOTA performance compared to the mainstream harnesses such as Mini-SWE-agent, Codex CLI, and Claude Code, showing that a well-designed organization can outperform strong existing harnesses, while OpenCollab's single-agent configuration uses the fewest tokens across all evaluated suites. OpenCollab establishes a unified multi-agent infrastructure for easy programmable collaboration and controlled causal evaluation.