Fast-Slow Thinking RM: Efficient Integration of Scalar and Generative Reward Models

📅 2026-03-02
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent tension between the high computational cost of generative reward models and the limited performance of scalar reward models. Inspired by dual-process theory in cognitive science, this work proposes a hybrid reward architecture that integrates fast scalar judgments with slow chain-of-thought reasoning. The core innovation lies in a dual-confidence activation mechanism that dynamically schedules between fast and slow thinking modes, achieving an adaptive balance between efficiency and accuracy. Experimental results demonstrate that the proposed method attains state-of-the-art performance on RLHF benchmarks with an average score of 84.3, while simultaneously reducing token consumption by 22.5%. These findings establish a new paradigm for reward modeling that effectively reconciles performance and computational efficiency.
📝 Abstract
Reward models (RMs) are critical for aligning Large Language Models via Reinforcement Learning from Human Feedback (RLHF). While Generative Reward Models (GRMs) achieve superior accuracy through chain-of-thought (CoT) reasoning, they incur substantial computational costs. Conversely, Scalar Reward Models (SRMs) offer efficiency but suffer from limited performance and adaptability in complex scenarios. We introduce Fast-Slow Thinking Reward Models (F/S-RM), a hybrid RM architecture inspired by Dual Process Theory. It trains a single model to integrate two distinct reward paradigms: scalar-style first-token pairwise judgment (fast thinking) and CoT-based judgment (slow thinking), regulated by a dual-confidence activation mechanism that determines when to activate slow thinking. Under hybrid inference, F/S-RM achieves state-of-the-art accuracy with an average score of 84.3 across benchmarks, while reducing token consumption by 22.5%.
Problem

Research questions and friction points this paper is trying to address.

Reward Models
RLHF
Generative Reward Models
Scalar Reward Models
Computational Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fast-Slow Thinking Reward Model
Dual Process Theory
Hybrid Inference
Dual-Confidence Activation Mechanism
Chain-of-Thought Reasoning
Jiayun Wu
Jiayun Wu
Carnegie Mellon University
Machine Learning
P
Peng Zhang
Fudan University
Y
Yuanyuan Lu
Meituan
S
Shan Qu
Meituan
N
Ning Gu
Fudan University
T
Tun Lu
Fudan University