A Cheap Verifier is Good Enough: LLM Post-training is Robust to Erroneous Rewards

๐Ÿ“… 2026-09-27
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study investigates whether expensive verifiers are truly necessary for large language model (LLM) post-training, which currently relies heavily on costly verification. Through large-scale H100 experiments based on Qwen3 and Gemma models, we systematically examine the relationship between verifier accuracy and post-training performance. Our findings across medical and legal domains demonstrate that open-source, lightweight verifiers can effectively substitute high-precision alternatives while exhibiting significant robustness to evaluation errors. This work challenges the conventional reliance on expensive verifiers by achieving approximately 99% reduction in scoring overhead with only a marginal 1โ€“3 point decline in post-training performance. Ultimately, it establishes a cost-effective paradigm for efficient reinforcement learning from human feedback (RLHF).
๐Ÿ“ Abstract
When post-training large language models on tasks with semi-verifiable rewards, there are many factors (training steps, base model size, training order, data quality, verifier accuracy, etc.) that practitioners must contend with to maximize model performance. Yet, it remains unclear how well verifier agreement predicts post-training performance on such tasks. In this paper, we explore this question with over 11k H100 GPU-hours, across HealthBench and PRBench tasks in medical, legal, and finance domains. Across the tested domains, Qwen3 trainees (1.7B-8B on HealthBench; 8B on PRBench), evaluation splits, and frontier LLM reference judges (which we call golden verifiers), higher verifier agreement does not consistently identify the best training verifier. Expensive verifiers need not outperform inexpensive ones, and open-weight Gemma verifiers produce strong training outcomes. We compare two low-cost choices retrospectively -- a cost-reducing choice and a balanced choice -- with estimated grading cost reductions of 98.8%-99.7% relative to the golden grading protocols and average post-training score gaps of 1-3 points from the best evaluated training verifier. These averages include larger losses in individual settings; they do not establish that verifier choices are interchangeable.
Problem

Research questions and friction points this paper is trying to address.

LLM post-training
semi-verifiable rewards
verifier agreement
reward verification
cost reduction
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Post-training
Reward Verification
Cost Reduction
Verifier Agreement
Semi-verifiable Rewards
๐Ÿ”Ž Similar Papers
No similar papers found.