Verifier Errors in RLVR: Reward Hacking, Limits of Feedback, and Selective Control

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of reward hacking and insufficient feedback caused by imperfect verifiers in Reinforcement Learning with Verifiable Rewards (RLVR). We propose a selective control mechanism based on audited feedback. Through gradient flow analysis, we characterize the conditions under which verifier errors occur, revealing that such errors cannot be identified solely from observations. To overcome this limitation, our method introduces auxiliary audit signals to enable selective intervention. We validate the proposed approach through experiments involving linear and neural contextual bandits as well as language models. The results demonstrate that, even under partial auditing, our mechanism significantly reduces the false acceptance rate and improves response correctness. Overall, this work provides an effective solution for mitigating verification bias in RLVR.
📝 Abstract
In reinforcement learning with verifiable rewards (RLVR), imperfect verifiers can reward incorrect responses, creating opportunities for reward hacking. Using gradient flow with a fixed verifier, we characterize the conditions under which reward rises while correctness falls. We then show that the observations available during RLVR are, in general, insufficient to detect or identify accepted errors, or to guarantee their reduction without sacrificing correct responses. To address this limit, we construct a correction using additional feedback about correctness from audits. This correction achieves \emph{selective control}: at the current policy, it lowers the probability of accepted errors and raises that of correct responses, provided it outweighs the pressure toward errors from verifier reward. Experiments with log linear and neural contextual bandits and with a language model support the analysis and show that selective control under partial auditing reduces accepted errors while increasing correctness.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning with Verifiable Rewards
Reward Hacking
Verifier Errors
Selective Control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reward Hacking
Selective Control
Verifiable Rewards
Partial Auditing
Gradient Flow