🤖 AI Summary
This study addresses the challenge that patches generated by coding agents often appear plausible yet omit critical behaviors, thereby complicating code review. Advocating a "verifiability over scale" principle, this work demonstrates that model parameter size is not the sole determinant of review quality. Methodologically, it introduces a cascaded review framework integrating execution tag tracing, static error detection, and failing test generation, enabling weaker reviewers to effectively audit stronger coding agents through structured evidence and automated tests. Experimental results indicate that the proposed framework achieves high coverage and defect detection rates even in the absence of official test suites, substantially enhancing the reliability of code review.
📝 Abstract
Coding agents can return plausible patches that omit required behavior. These failures are hard to review because long traces and confident summaries often hide what was missed. We ask when a nominally weaker reviewer can reliably decide whether a patch solves its issue. We study 411 execution-labeled traces from three agents and 101 controlled cases. On 154 GPT-5.4 traces, structured but unchecked evidence raises both defect catch and over-rejection. We then provide official execution evidence as an upper-bound diagnostic. After choosing and freezing one of two formats per reviewer, five of six reviewers improve both rates on 122 held-out traces; two classify every trace correctly. Reviewer size is not a consistent predictor of quality. Because official tests are unavailable in deployment, we also evaluate a frozen cascade with patch-caused static errors and generated tests that first fail on the unpatched repository. On 121 scored held-out GPT-5.4 traces and 59 Gemini traces, its coverage is 0.89 and 0.86, risk is 0.33 and 0.26, catch is 0.76 and 0.80, and over-rejection is 0.66 and 0.67. Most false rejections occur when unresolved cases reach the reviewer. Official execution evidence shows the potential of weak review when decisive checks are available. Producing equally reliable checks without official tests remains the main bottleneck.