CRJudgeBench: Can AI Detect Plausible but Invalid Code Reviews?

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of AI systems identifying plausible yet technically erroneous code review comments by proposing CRJudgeBench, a benchmark integrating real pull requests with expert-perturbed data to fill the gap in comment credibility assessment, and Sentinel, an agent framework built upon Qwen3-Coder. Sentinel incorporates a repository-context grounding architecture and enhances evidence-gathering capabilities through privileged teacher-guided iterative action-level reinforcement learning. Experimental results demonstrate that Sentinel achieves an accuracy of 76.60%, significantly outperforming strong baselines such as GLM-5.3 and effectively improving the verification of technical correctness in code review comments.
📝 Abstract
Large language models can generate plausible code-review comments, but such comments may contain technically incorrect claims that mislead developers. We study technical trustworthiness judgment: determining whether a review comment's core technical claims are correct and applicable to the reviewed code in its repository context. Existing code-review benchmarks primarily evaluate review generation, issue discovery, or general comment quality, but do not directly assess whether an agent can determine the technical trustworthy of an individual review comment. To fill this gap, we introduce CRJudgeBench, a benchmark of 1199 instances constructed from real pull requests and expert-verified perturbations, covering both trustworthy and plausible but untrustworthy comments. We further present Sentinel, a repository-grounded agentic judge that actively gathers code evidence to verify review comments before making judgments. Starting from Qwen3-Coder-30B-A3B-Instruct, Sentinel is trained on the CRJudgeBench training split through iterative action-level learning from a privileged teacher. On the 359-instance CRJudgeBench test set, Sentinel achieves 76.60\% accuracy, outperforming GLM-5.3 by 6.13 percentage points and its base model by 19.78 points. These results show that even state-of-the-art general-purpose LLMs struggle to identify untrustworthy comments, while iterative action-level learning substantially improves the accuracy of repository-grounded trustworthiness judgments. Our dataset is available at https://huggingface.co/datasets/dcloud347/CRJudgeBenchmark
Problem

Research questions and friction points this paper is trying to address.

code review
trustworthiness judgment
large language models
benchmark
hallucination
Innovation

Methods, ideas, or system contributions that make the work stand out.

Code Review Trustworthiness
CRJudgeBench
Agentic Judge
Repository-grounded Verification
Iterative Action-level Learning