RMB: Reward Model Boosting Mitigates Reward Hacking

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reward hacking problem in Reinforcement Learning from Human Feedback (RLHF) caused by imperfections in proxy reward models. To this end, it introduces the boosting principle into reward modeling for the first time, proposing a novel reward model enhancement method. Specifically, the approach trains an ensemble of complementary reward models via diversity regularization and employs a lightweight neural network aggregator to integrate their outputs, thereby generating robust reward signals that guide policy optimization. Experimental results demonstrate that the proposed method significantly improves reward accuracy on both in-distribution and out-of-distribution data. Furthermore, it effectively mitigates reward hacking and comprehensively enhances the RLHF alignment performance of large language models.
📝 Abstract
Reinforcement Learning from Human Feedback (RLHF) is a powerful technique for aligning large language models (LLMs) with human preference. However, it often suffers from the reward hacking issue, where policy optimization improves the proxy reward model while actually degrading performance with respect to the true human preference, due to the imperfection of the proxy. To address this, we propose Reward Model Boosting (RMB), a novel approach that enhances the robustness and reliability of the reward signal for RLHF. RMB first trains a set of reward models with a diversity-promoting regularizer. This encourages each model to learn complementary aspects of the reward landscape. Then, RMB learns a lightweight aggregator in the principle of boosting to aggregate the outputs of the diverse reward models into a more accurate and robust reward signal. Our extensive experiments demonstrate that RMB significantly improves reward accuracy on both in-distribution and out-of-distribution datasets, substantially mitigating the reward hacking issue and ultimately improving RLHF performance.
Problem

Research questions and friction points this paper is trying to address.

Reward Hacking
RLHF
Proxy Reward Model
Large Language Models
Human Preference Alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reward Model Boosting
Reward Hacking
RLHF
Diversity Regularization
Lightweight Aggregator
🔎 Similar Papers