WideSWE: Can Coding Agents Coordinate Changes Across Repositories?

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing coding agent benchmarks, which are largely confined to single repositories and fail to reflect real-world cross-repository collaborative development. We propose the first cross-repository evaluation framework for coding agents, constructing a benchmark of 120 real-world tasks by mining software ecosystem changes and adapting hidden tests for case generation. This work systematically compares independent and joint execution strategies, revealing the critical role of information sharing in achieving implementation completeness. Experimental results demonstrate that, using model configurations such as Codex CLI, the highest task success rate reaches 42.50%. Furthermore, joint execution more effectively leverages correlated repository information to guide code generation, significantly outperforming the independent mode. These findings establish a new paradigm for evaluating and enhancing the multi-repository collaboration capabilities of coding agents.
📝 Abstract
Coding-agent evaluation has progressed from resolving individual issues to carrying out long-horizon development, yet task completion is still largely assessed within a single codebase. In software ecosystems, many features and bug fixes require coordinated changes across multiple repositories. We introduce WideSWE to evaluate coding agents on such cross-repository tasks. Mining and reviewing changes across 103 software ecosystems yields 120 real-world tasks, balanced between 60 bug fixes and 60 features. We derive prompts from related issues and pull requests. We systematically review and adapt hidden tests to support diverse correct implementations while preserving required behavior and regression checks. Across seven agent configurations, full task success ranges from 10.83% to 42.50%, with the configuration pairing Codex CLI with GPT-5.6-sol achieving the highest rate. Trajectories show agents failing to identify necessary changes, recognizing changes but leaving them unfinished, or modifying the required repositories without fully satisfying the request. To examine whether working on one repository at a time can alleviate these difficulties, we compare it with joint execution under identical prompts. Independent execution mainly recovers omitted work and is less effective at correcting previously attempted but unsuccessful implementations. Joint execution can use information from related repositories to guide implementation and verification. Code is available at https://github.com/ZJU-ACES-ISE/WideSWE.
Problem

Research questions and friction points this paper is trying to address.

coding agents
cross-repository coordination
software ecosystems
agent evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-repository coordination
Coding agents evaluation
Software ecosystems
Benchmark dataset
Joint execution