Audit-First VAPO: Risk-Certified Selective Updates under Imperfect Verification

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the risk of harmful updates introduced by imperfect verifiers in reinforcement learning by proposing an "audit-before-execute" safety framework. Methodologically, it decouples discrete direction admission from continuous magnitude control, employing an observation-based accept-appeal-abstain strategy for audit selection under a limited budget. Risk is controlled via conditional Hoeffding-Azuma bounds and a bounded-trust clipped-KL executor. This work provides the first theoretical certification of risk, coverage, and cost under finite-sample conditions. Experiments demonstrate that the model achieves 74.1% accuracy on the GSM8K benchmark while effectively suppressing the risk of harmful updates and ensuring certificate satisfaction rates across diverse noise scenarios.
📝 Abstract
Imperfect verifiers can assign a harmful update direction even when clipping and regularization bound its magnitude. We introduce Audit-First VAPO, which separates discrete directional admission from continuous magnitude control. An observation-only accept-appeal-abstain policy uses a finite secondary-verification budget; its action trace is frozen before clean labels are joined. Simultaneous finite-sample bounds then certify selected harmful risk, coverage, and verifier-call rate over a predeclared policy family. Conditional Hoeffding-Azuma bounds account for the dependence induced by shared budgets, and rollout or verifier changes initiate a new certification stage. After admission, a bounded trust-clip-KL actuator controls magnitude. We evaluate two models on two reasoning benchmarks against static RLVR, matched-random selection, confidence thresholding, noise correction, and verifier augmentation. On Qwen3.5-0.8B and GSM8K at target risk $\rho=0.08$, RC-VAPO achieves 74.1% accuracy, selected harmful risk 0.0697, coverage 0.4125, and relative verifier cost $1.16\times$. At matched coverage and update magnitude, its selected-risk difference from matched random is -0.0260 with paired 95% interval $[-0.0364,-0.0157]$. Across asymmetric, confidence-dependent, and correlated-verifier noise, the certificate is satisfied on 57 of 60 independent runs. These comparisons isolate informative directional selection from proposal suppression, update shrinkage, and additional verifier computation.
Problem

Research questions and friction points this paper is trying to address.

imperfect verifiers
harmful risk
selective updates
risk certification
directional admission
Innovation

Methods, ideas, or system contributions that make the work stand out.

Audit-First VAPO
Directional Admission
Conditional Hoeffding-Azuma Bounds
Risk Certification
Trust-Clip-KL Actuator
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Miaobo Hu
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
S
Shuhao Hu
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
Xiaobo Guo
Xiaobo Guo
Dartmouth College
machine learningdeep learningnatural language processingsocia mediapropagantion
Xin Wang
Xin Wang
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences
Biomedical Engineering
Bokun Wang
Bokun Wang
Texas A&M University
Machine LearningArtificial IntelligenceMultimodal Machine Learning
R
Rui Chen
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
D
Daren Zha
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
J
Jun Xiao
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China