🤖 AI Summary
This study addresses the "reward hacking without quality improvement" problem in Group Relative Policy Optimization (GRPO) caused by reward model fragility, proposing a reliability-first advantage estimator. Methodologically, it introduces a novel reliability-based self-tuning normalization mechanism that disentangles quality signals from learning effects via pairwise evaluation, while employing robust fitting to suppress interference from unreliable rewards on baselines and updates, thereby theoretically constraining the influence of unsupported rewards. Experiments demonstrate that this approach effectively curtails over-optimization in token-level interface utilization and medical reasoning tasks. It significantly improves independent semantic evaluation scores, narrows the proxy-judge discrepancy, and reduces hallucinated claims.
📝 Abstract
Reward hacking occurs when policy optimization exploits a brittle reward interface or an overly permissive proxy objective, improving the training score without improving the underlying response quality. This phenomenon is amplified in group-relative policy optimization: an unsupported reward can shift the group baseline and alter the updates of other rollouts, while post-hoc or purely relative weighting cannot represent group-wide uncertainty. We propose \emph{Self-Tuned Anchored Reliability Group-Relative Policy Optimization} (STAR-GRPO), a reliability-first advantage estimator based on paired assessments of the same rollout. STAR separates the quality signal from its learning influence: score disagreement determines rollout reliability, relative reliability enters a self-tuned robust location--scale fit before group normalization, and absolute group reliability attenuates the resulting bounded advantage. The analysis establishes coordinate and second-moment bounds, characterizes exact centering through the weighted location equation, and gives reliability-dependent attenuation guarantees for outlying rewards. We evaluate STAR-GRPO in two complementary reward-hacking regimes. In token-interface exploitation, STAR prevents runaway optimization of the deployed-interface score while improving the canonical quality signal. In rubric-proxy overoptimization for medical reasoning, STAR improves independent semantic evaluation, narrows the proxy--judge discrepancy, and reduces overclaim while optimizing the same task proxy. Together, these results show that reliability-first normalization offers a principled way to limit unsupported reward influence on both group baselines and policy updates, while retaining the task reward as the optimization target.