RMB: Comprehensively Benchmarking Reward Models in LLM Alignment

📅 2024-10-13
🏛️ arXiv.org
📈 Citations: 7
✨ Influential: 2
📄 PDF
🤖 AI Summary
Existing reward modeling (RM) evaluation suffers from two critical limitations: narrow data distribution and misalignment between evaluation protocols and the ultimate human preference alignment objective, resulting in poor correlation with actual alignment performance. To address this, we introduce RMB—the first comprehensive benchmark covering 49 realistic application scenarios—designed to support both pairwise comparison and Best-of-N evaluation modes, directly aligned with LLM human preference alignment goals. Methodologically, RMB is constructed via multi-scenario human annotation, augmented by correlation analysis and attribution studies to systematically expose generalization failures of mainstream RMs. We further provide the first empirical validation of generative RM capabilities and pioneer analyses of majority voting mechanisms and instruction design effects on evaluation fidelity. Experiments demonstrate strong positive correlation between RMB scores and downstream alignment performance, uncovering RM weaknesses missed by prior benchmarks. We open-source the dataset, code, and toolchain to enable reproducible, extensible, and standardized RM evaluation.

Technology Category

Machine Learning: Learning Preferences or RankingsHumans and AI: Learning Human Values and PreferencesKnowledge Representation and Reasoning: Preferences

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Reward models (RMs) guide the alignment of large language models (LLMs), steering them toward behaviors preferred by humans. Evaluating RMs is the key to better aligning LLMs. However, the current evaluation of RMs may not directly correspond to their alignment performance due to the limited distribution of evaluation data and evaluation methods that are not closely related to alignment objectives. To address these limitations, we propose RMB, a comprehensive RM benchmark that covers over 49 real-world scenarios and includes both pairwise and Best-of-N (BoN) evaluations to better reflect the effectiveness of RMs in guiding alignment optimization. We demonstrate a positive correlation between our benchmark and the downstream alignment task performance. Based on our benchmark, we conduct extensive analysis on the state-of-the-art RMs, revealing their generalization defects that were not discovered by previous benchmarks, and highlighting the potential of generative RMs. Furthermore, we delve into open questions in reward models, specifically examining the effectiveness of majority voting for the evaluation of reward models and analyzing the impact factors of generative RMs, including the influence of evaluation criteria and instructing methods. Our evaluation code and datasets are available at https://github.com/Zhou-Zoey/RMB-Reward-Model-Benchmark.
Problem

Research questions and friction points this paper is trying to address.

Evaluating reward models' alignment performance with limited data distribution
Assessing RM effectiveness in guiding LLM alignment across diverse scenarios
Investigating generalization defects and potential of generative reward models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Comprehensive RM benchmark with 49 scenarios
Pairwise and Best-of-N evaluation methods
Analyzes generalization defects of reward models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Fudan University | University of North Carolina at Chapel Hill
E
Enyu Zhou
NLP Group, Fudan University
G
Guodong Zheng
NLP Group, Fudan University
B
Bing Wang
NLP Group, Fudan University
Zhiheng Xi
Zhiheng Xi
Fudan University
LLM ReasoningLLM-based Agents
Shihan Dou
Shihan Dou
Fudan University
LLMsCode LMsRLAlignment
Rong Bao
Rong Bao
PhD student, Fudan University
AlignmentGenerative AIReinforcement Learning
W
Wei Shen
NLP Group, Fudan University
L
Limao Xiong
NLP Group, Fudan University
J
Jessica Fan
UNC Chapel Hill, USA
Y
Yurong Mou
NLP Group, Fudan University
R
Rui Zheng
NLP Group, Fudan University
T
Tao Gui
NLP Group, Fudan University
Q
Qi Zhang
NLP Group, Fudan University
X
Xuanjing Huang
NLP Group, Fudan University