RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge that generative reward models, due to their comparative outputs, are incompatible with the scalar rewards required by reinforcement learning, thereby hindering effective training of large language models. To overcome this limitation, the paper proposes a Ranking-based Reward Construction (RRC) method, which innovatively introduces self-competitive ranking and anchor-guided ranking strategies to transform relative preference rankings into scalar reward signals suitable for reinforcement learning. By circumventing the constraints of conventional scalar reward formulation, RRC achieves substantially improved training performance on open-ended dialogue and reasoning benchmarks, consistently outperforming existing reward modeling approaches.
📝 Abstract
Recent advances in reward modeling show a paradigm shift from discriminative reward models to generative reward models. However, despite their strong capabilities in response ranking, generative reward models have not realized their potential in reinforcement learning (RL). Our analysis reveals that this limitation arises from a mismatch between the comparative nature of generative reward modeling and the scalar scoring paradigm adopted by existing RL algorithms. To bridge this gap, we propose a Ranking-based Reward Construction (RRC) approach, which enables generative reward models to provide more effective RL learning signals by deriving rewards from relative preference rankings. RRC introduces two complementary strategies: self-competitive ranking, which exploits comparisons among sampled responses, and anchor-guided ranking, which enables scalable ranking-based reward construction with a small set of reference responses. Experiments across open-ended chat and reasoning benchmarks demonstrate that RRC substantially improves RL training with generative reward models, achieving consistent gains over existing reward construction approaches. Our code can be found at https://github.com/wangclnlp/RRC.
Problem

Research questions and friction points this paper is trying to address.

generative reward models
reinforcement learning
reward modeling
ranking
scalar scoring
Innovation

Methods, ideas, or system contributions that make the work stand out.

generative reward models
ranking-based reward construction
reinforcement learning
preference ranking
LLM alignment
🔎 Similar Papers
No similar papers found.
Chenglong Wang
Chenglong Wang
Northeastern University (Shenyang, China)
Natural Language ProcessingLanguage Model Alignment
Z
Ziming Zhu
School of Computer Science and Engineering, Northeastern University, Shenyang, China
Yifu Huo
Yifu Huo
Northeastern University
Bei Li
Bei Li
Meituan LLM Team
Machine TranslationDeep LearningLarge Language Models
Qiaozhi He
Qiaozhi He
ByteDance
LLMNatural Language Processing
Y
Yan Ding
School of Computer Science and Engineering, Northeastern University, Shenyang, China
Xiaoyang Hao
Xiaoyang Hao
Tencent
speech synthesis
Y
Yuxin Gao
School of Computer Science and Engineering, Northeastern University, Shenyang, China
Tianhua Zhou
Tianhua Zhou
Fujian Institute of Research on the Structure of Matter, Chinese Academy of Sciences
Photo- and Electro-catalysisCrystalline Porous MaterialsCO2RRSurfactant
X
Xiaojia Chang
Independent Researcher, Beijing, China
T
Tongran Liu
CAS Key Laboratory of Behavioral Science, Institute of Psychology, CAS, Beijing, China
Jingbo Zhu
Jingbo Zhu
Northeastern University, China
Machine TranslationLanguage ParsingNatural Language Processing