Scoring Verifiers: Evaluating Synthetic Verification in Code and Reasoning

📅 2025-02-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Synthetic verification methods—such as self-generated test cases and reward modeling—lack systematic evaluation for code correctness assessment. Method: We introduce four novel benchmarks—HE-R, HE-R+, MBPP-R, and MBPP-R+—unifying programming evaluation datasets into dual-modality tasks: scoring and ranking. We propose the first standardized evaluation framework tailored to synthetic verification capability, enabling multi-dimensional accuracy assessment and attribution analysis. Contribution/Results: Experiments reveal a synergistic gain between reasoning model scale and test case quantity; mainstream LLMs significantly improve test generation quality; and increasing test cases consistently enhances verification accuracy. Cross-model evaluation across multiple LLMs identifies key determinants of verification performance, including test diversity, oracle reliability, and model calibration. This work establishes foundational infrastructure and empirical insights for advancing trustworthy code verification via synthetic methods.

Technology Category

Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageMachine Learning: Calibration & Uncertainty QuantificationConstraint Satisfaction and Optimization: Satisfiability Modulo Theories

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Code verification has recently found great success as a critical component in training large scale reasoning models for coding. Synthetic techniques such as self-generated test cases and reward models provide a way to enhance code capabilities beyond predefined tests. Building on these advancements, we propose new benchmarks designed to systematically evaluate the impact of synthetic verification methods on assessing solution correctness. We introduce HE-R, HE-R+, MBPP-R, and MBPP-R+, which transform existing coding benchmarks into scoring and ranking datasets to evaluate the effectiveness of synthetic verifiers. Using these benchmarks, we analyze synthetic verification methods in standard, reasoning-based, and reward-based LLMs. Our results show that recent reasoning models significantly improve test case generation and that scaling test cases enhances verification accuracy.
Problem

Research questions and friction points this paper is trying to address.

Evaluate synthetic verification in coding
Transform benchmarks into scoring datasets
Analyze verification methods in LLMs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Synthetic test case generation
Reward models integration
Enhanced benchmark transformation
🔎 Similar Papers
No similar papers found.