MatrixReward: Reward from Rubric Matrix for Open-Ended Generation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of quantifying rewards in open-ended generation tasks due to the absence of standard answers. We propose an adaptive reward mechanism based on pairwise comparison matrices. Specifically, this method constructs a scoring-rubric-driven win-rate matrix to capture fine-grained inter-sample differences, dynamically fuses weights via column distributions and correlations, and computes rewards using ideal silhouette distance, thereby effectively mitigating the information loss inherent in single scalar scores. Reinforcement learning experiments conducted on the Qwen3-8B model demonstrate that our approach achieves an average score of 63.02 across four benchmarks, outperforming the strongest baseline by approximately 2.0%. These results validate the effectiveness of constructing rewards through relative comparison matrices for open-ended text generation.
📝 Abstract
Open-ended query generation lacks standard answers, thus necessitating an effective reward mechanism. Pointwise scoring rubrics provide limited information about the relative quality of sample answers under the same prompt; merging multiple rubric judgments into a single score may also mask the differences between these answers. We propose MatrixReward, which constructs rewards from a rollout-by-rubric win-rate matrix obtained by comparing every pair of sampled responses under each rubric. The spread of each matrix column captures how strongly that rubric distinguishes the current rollouts, while correlations between columns reveal rubric repetition; together, these statistics yield data-dependent rubric weights. We combine these weights with the prior weights of rubrics. After column normalization and weighting, the observed per-rubric maxima and minima define positive and negative ideal profiles. Each rollout's distances to these two ideals determine its relative-closeness quality reward. Evaluated using Qwen3-8B on four open-ended query-answering benchmarks, MatrixReward achieves an average score of 63.02, outperforming the strongest baseline by approximately 2.0%. These results support the idea that matrices derived from relative comparisons can be used to construct rewards more reasonably for open-ended generative reinforcement learning.
Problem

Research questions and friction points this paper is trying to address.

open-ended generation
reward mechanism
scoring rubrics
reinforcement learning
relative quality assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

MatrixReward
win-rate matrix
data-dependent rubric weights
relative-closeness reward
open-ended generation
🔎 Similar Papers
No similar papers found.