SWE-Prometheus: Measuring Engineering Governance Improvements in Real-World Repositories

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing benchmarks evaluate only functional patches, making it difficult to measure the capacity of code agents to improve engineering governance in real-world repositories. To address this limitation, this work proposes SWE-Prometheus, a benchmark that enables open-ended goal-driven agents to autonomously identify risks, prioritize interventions, and validate changes. Methodologically, it establishes a six-dimensional governance evaluation framework incorporating pairwise evidence, behavioral gating, and a dual-teacher independent scoring mechanism to discern substantive improvements, alongside a comprehensive quantitative metric termed NGI. Experiments across ten models on sixty repositories yield an average NGI of 0.576, demonstrating that accurate assessment necessitates jointly reporting improvement magnitude, behavioral preservation, and evidence quality.
📝 Abstract
Large language model based coding agents have made substantial progress on repository-level software engineering tasks. Existing repository benchmarks, however, usually start from a human-identified issue and evaluate whether a patch satisfies a functional signal. We present SWE-Prometheus, a benchmark for the broader task of improving repository engineering governance. Each task provides a fixed snapshot and an open-ended objective, requiring the agent to identify risks, prioritize interventions, and verify the resulting changes. SWE-Prometheus evaluates six governance dimensions through paired evidence, clean-environment probes, behavior gates, and two independent teacher ratings of the same evidence. The benchmark contains 60 repositories; ten models are evaluated on a shared 22-repository public subset, where mean Normalized Governance Improvement ranges from 0.0568 to 0.5760 and observed behavior-breakage rates range from 0% to 23%. On a frozen ten-repository batch, a repository-blind template obtains mean NGI 0.272, but its gains concentrate in Tests & CI, Quality Gates, and Documentation; it improves Reproducible Environment and Dependency & Security on none of the repositories. This baseline makes the distinction between adding governance artifacts and producing execution-backed improvements measurable. The no-op condition has median NGI zero and standard deviation 0.073; two teachers agree exactly on 57 of 60 dimension scores for the same no-op evidence. For the two highest conditional-mean systems, common-valid NGI is similar, while full-pool comparisons that include behavior failures favor Kimi-K3. These results show why repository-governance evaluation should report improvement, behavior preservation, evidence quality, and coverage together.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Engineering Governance
Repository Benchmark
LLM Coding Agents
Multi-dimensional Evaluation
Behavior Preservation