STAR-GRPO: Canonical Anchoring and Reliability-First Advantages against Representation-Dependent Reward Hacking

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the "reward hacking without quality improvement" problem in Group Relative Policy Optimization (GRPO) caused by reward model fragility, proposing a reliability-first advantage estimator. Methodologically, it introduces a novel reliability-based self-tuning normalization mechanism that disentangles quality signals from learning effects via pairwise evaluation, while employing robust fitting to suppress interference from unreliable rewards on baselines and updates, thereby theoretically constraining the influence of unsupported rewards. Experiments demonstrate that this approach effectively curtails over-optimization in token-level interface utilization and medical reasoning tasks. It significantly improves independent semantic evaluation scores, narrows the proxy-judge discrepancy, and reduces hallucinated claims.
📝 Abstract
Reward hacking occurs when policy optimization exploits a brittle reward interface or an overly permissive proxy objective, improving the training score without improving the underlying response quality. This phenomenon is amplified in group-relative policy optimization: an unsupported reward can shift the group baseline and alter the updates of other rollouts, while post-hoc or purely relative weighting cannot represent group-wide uncertainty. We propose \emph{Self-Tuned Anchored Reliability Group-Relative Policy Optimization} (STAR-GRPO), a reliability-first advantage estimator based on paired assessments of the same rollout. STAR separates the quality signal from its learning influence: score disagreement determines rollout reliability, relative reliability enters a self-tuned robust location--scale fit before group normalization, and absolute group reliability attenuates the resulting bounded advantage. The analysis establishes coordinate and second-moment bounds, characterizes exact centering through the weighted location equation, and gives reliability-dependent attenuation guarantees for outlying rewards. We evaluate STAR-GRPO in two complementary reward-hacking regimes. In token-interface exploitation, STAR prevents runaway optimization of the deployed-interface score while improving the canonical quality signal. In rubric-proxy overoptimization for medical reasoning, STAR improves independent semantic evaluation, narrows the proxy--judge discrepancy, and reduces overclaim while optimizing the same task proxy. Together, these results show that reliability-first normalization offers a principled way to limit unsupported reward influence on both group baselines and policy updates, while retaining the task reward as the optimization target.
Problem

Research questions and friction points this paper is trying to address.

Reward Hacking
Group-Relative Policy Optimization
Policy Optimization
Proxy Overoptimization
Representation-Dependent Reward
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reward Hacking
Group-Relative Policy Optimization
Reliability-First Advantage Estimation
Robust Location-Scale Normalization
Self-Tuned Anchoring
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
W
Wan Tian
Peking University
Z
Zhongyi Li
Beihang University
X
Xiang Xu
Beihang University
Minhao Zou
Minhao Zou
Peking University
Yijie Peng
Yijie Peng
Peking University
SimulationBayesian LearningArtificial IntelligenceHealthcareFinancial Engineering
F
Fuzhen Zhuang
Beihang University