🤖 AI Summary
This study investigates what discriminative reward models memorize from human preference data and the origins of their biases. Through counterfactual memory metrics, pairwise preference analysis, ablation studies, and out-of-distribution evaluation, it systematically reveals—for the first time—the models’ reliance on non-semantic cues such as model identity, user sampling strategies, and response length. The findings indicate that reward models tend to over-memorize high-margin, easily separable samples and dataset-specific shortcuts, and they generalize simplistic heuristic features to unseen preference pairs, resulting in insufficient context-sensitive judgment capabilities. This work provides critical empirical evidence for understanding the generalization behavior and bias mechanisms inherent in reward models trained on human preferences.
📝 Abstract
This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets. We show that RMs 1) misallocate memorization to easy, high margin preference pairs, 2) memorize dataset-specific shortcuts (e.g., model identity, user sampling strategy), and 3) overgeneralize simple heuristic correlates of human preference (e.g., length, compliance) when confronted with unseen preference pairs. Overall, our findings indicate that discriminative training of RMs from human preference data results in biased RMs not yet capable of judging response quality in context-dependent scenarios.