🤖 AI Summary
This study addresses the issue that coding agents, despite passing functional tests, frequently violate repository governance standards. To this end, we introduce SWE-CC, a benchmark for systematically evaluating both code and process compliance. Methodologically, we construct 823 machine-verifiable policies, design an auditing mechanism that integrates runtime behavior with final deliverables, and propose an evaluation framework combining semi-automated document conversion, deterministic checking, and LLM agent workflows. Experimental results reveal that modern agents exhibit a policy violation rate of 43.1%, with nearly half occurring during intermediate execution steps rather than in final outputs. These findings underscore the critical necessity of process-level compliance auditing to ensure that autonomous coding agents adhere not only to functional requirements but also to established software engineering practices and repository governance norms.
📝 Abstract
Autonomous coding agents now resolve a substantial share of real-world GitHub issues. However, passing functional tests differs fundamentally from producing a high-quality contribution acceptable for merging. Mature open-source projects publish repository-specific contribution policies, spanning style, git, testing workflows, to ensure code quality and long-term maintainability. Because existing benchmarks evaluate patches solely on unit tests, agent compliance with repository governance remains unknown. In this paper, we introduce SWE-CC, a benchmark evaluating code and process compliance in autonomous software engineering. We develop a semi-automated pipeline that converts developer documentation across 12 open-source repositories into 823 machine-checkable atomic policies. SWE-CC introduces two features: 1) lightweight, deterministic checker functions that represent each policy, 2) a comprehensive auditing mechanism that inspects both agent runtime behaviors and final deliverables. We evaluate the compliance of agent workflows in 500 end-to-end software contribution tasks extended from SWE-bench Verified. Our evaluation of four LLMs under two agent scaffolds shows that modern agents suffer from coding compliance issues: although agents produce functionally correct patches, they still violate 43.1 percent of applicable project policies, with nearly half of all violations occurring during intermediate execution steps. These results show that functional correctness does not guarantee real-world readiness, highlighting that future software engineering agents must reliably conform to repository governance to enable safe and trustworthy deployment.