🤖 AI Summary
This study addresses the inherent tension between the high computational cost of generative reward models and the limited performance of scalar reward models. Inspired by dual-process theory in cognitive science, this work proposes a hybrid reward architecture that integrates fast scalar judgments with slow chain-of-thought reasoning. The core innovation lies in a dual-confidence activation mechanism that dynamically schedules between fast and slow thinking modes, achieving an adaptive balance between efficiency and accuracy. Experimental results demonstrate that the proposed method attains state-of-the-art performance on RLHF benchmarks with an average score of 84.3, while simultaneously reducing token consumption by 22.5%. These findings establish a new paradigm for reward modeling that effectively reconciles performance and computational efficiency.
📝 Abstract
Reward models (RMs) are critical for aligning Large Language Models via Reinforcement Learning from Human Feedback (RLHF). While Generative Reward Models (GRMs) achieve superior accuracy through chain-of-thought (CoT) reasoning, they incur substantial computational costs. Conversely, Scalar Reward Models (SRMs) offer efficiency but suffer from limited performance and adaptability in complex scenarios. We introduce Fast-Slow Thinking Reward Models (F/S-RM), a hybrid RM architecture inspired by Dual Process Theory. It trains a single model to integrate two distinct reward paradigms: scalar-style first-token pairwise judgment (fast thinking) and CoT-based judgment (slow thinking), regulated by a dual-confidence activation mechanism that determines when to activate slow thinking. Under hybrid inference, F/S-RM achieves state-of-the-art accuracy with an average score of 84.3 across benchmarks, while reducing token consumption by 22.5%.