EnGRICH: Enhancing Generative Reward Modeling with Critiques from Humans

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unreliability of critiques in generative reward models due to their dependence on outcome supervision, alongside the underutilization of scarce human critique data. To this end, we propose EnGRICH, a framework that introduces a MetaCritic module to generalize limited human critiques across large-scale preference datasets. By constructing response-specific scoring criteria, this approach pioneers the transfer of evaluation standards from sparse human critiques to extensive outcome data, enabling fine-grained process supervision. The methodology integrates meta-critic learning, evidence coverage assessment, and reinforcement learning-based process rewards. Extensive experiments across seven benchmarks demonstrate that EnGRICH consistently outperforms competitive baselines, validating its effectiveness in enhancing both critique reliability and overall model performance.
📝 Abstract
Generative reward models (GRMs) are important for LLM optimization. Unlike scalar reward models, GRMs generate natural-language critiques alongside preference judgments, providing finer-grained evaluation signals. Their effectiveness depends heavily on critique reliability. However, existing GRM training typically uses final preference correctness as outcome supervision. Because the preference outcome space is highly constrained, unreliable critiques can still yield correct outcomes and thus be reinforced. Recent work leverages human critiques for process supervision, but such critiques are scarce and are often reduced to scalar rewards, leaving their fine-grained evaluative information underutilized. We argue that evaluative criteria learned from human critiques can be generalized to broader outcome-only preference data. To this end, we propose \textbf{EnGRICH}, a GRM training framework that pairs the GRM with a training-time MetaCritic learned from a small set of human critiques. MetaCritic constructs response-specific rubrics and uses them to evaluate the evidence coverage and correctness of generated critiques. The resulting signals provide both process rewards for fine-grained credit assignment and structured guidance for exploring better critiques. During GRM training, MetaCritic is further optimized to generalize human-grounded evaluative criteria to outcome-only data. At inference, the trained GRM operates independently. Experiments across seven reward-model benchmarks show that EnGRICH consistently improves over competitive baselines, while further analyses validate the effectiveness of its core mechanisms.
Problem

Research questions and friction points this paper is trying to address.

Generative Reward Models
Critique Reliability
Process Supervision
Outcome Supervision
Human Critiques
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative Reward Modeling
MetaCritic
Process Supervision
Human Critiques
Outcome Generalization
X
Xuancheng Li
Department of Computer Science and Technology, Tsinghua University, Beijing, China
B
Beining Wang
Department of Computer Science and Technology, Tsinghua University, Beijing, China
Haitao Li
Haitao Li
TsingHua University
Information Retrieval
H
Heng Wang
Tencent, Beijing, China
Yujia Zhou
Yujia Zhou
Tsinghua University
Information retrievalSocial Simulation
Q
Qingyi Pan
Department of Computer Science and Technology, Tsinghua University, Beijing, China
B
Blaze Chen
Department of Computer Science and Technology, Tsinghua University, Beijing, China
Y
Yiqun Liu
Department of Computer Science and Technology, Tsinghua University, Beijing, China
M
Min Zhang
Department of Computer Science and Technology, Tsinghua University, Beijing, China
Qingyao Ai
Qingyao Ai
Associate Professor, Dept. of CS&T, Tsinghua University
Information RetrievalMachine Learning