🤖 AI Summary
Synthetic verification methods—such as self-generated test cases and reward modeling—lack systematic evaluation for code correctness assessment. Method: We introduce four novel benchmarks—HE-R, HE-R+, MBPP-R, and MBPP-R+—unifying programming evaluation datasets into dual-modality tasks: scoring and ranking. We propose the first standardized evaluation framework tailored to synthetic verification capability, enabling multi-dimensional accuracy assessment and attribution analysis. Contribution/Results: Experiments reveal a synergistic gain between reasoning model scale and test case quantity; mainstream LLMs significantly improve test generation quality; and increasing test cases consistently enhances verification accuracy. Cross-model evaluation across multiple LLMs identifies key determinants of verification performance, including test diversity, oracle reliability, and model calibration. This work establishes foundational infrastructure and empirical insights for advancing trustworthy code verification via synthetic methods.
📝 Abstract
Code verification has recently found great success as a critical component in training large scale reasoning models for coding. Synthetic techniques such as self-generated test cases and reward models provide a way to enhance code capabilities beyond predefined tests. Building on these advancements, we propose new benchmarks designed to systematically evaluate the impact of synthetic verification methods on assessing solution correctness. We introduce HE-R, HE-R+, MBPP-R, and MBPP-R+, which transform existing coding benchmarks into scoring and ranking datasets to evaluate the effectiveness of synthetic verifiers. Using these benchmarks, we analyze synthetic verification methods in standard, reasoning-based, and reward-based LLMs. Our results show that recent reasoning models significantly improve test case generation and that scaling test cases enhances verification accuracy.