Scoring Higher, Answering Worse: Mitigating Reward Hacking in Rubric-Based RL via Protocol-Level Rubrics

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses compensatory reward hacking in rubric-based reinforcement learning, where models exploit irrelevant content to offset critical omissions and artificially inflate scores. To mitigate this, we propose Protocol-level Rubrics (ProRubric), which groups evaluation criteria by dimension and replaces conventional linear weighted aggregation with an all-or-nothing decision mechanism. This approach suppresses reward cheating at the protocol level without altering the underlying rubric content. Experimental results demonstrate that ProRubric maintains evaluation coverage while improving clinical counseling appropriateness by 10.8 points and achieving the highest average score across seven benchmarks. These findings indicate that ProRubric effectively enhances the reliability of large language model alignment.
📝 Abstract
Rubric-based reinforcement learning (Rubric-RL) trains language models where no verifier exists. A judge checks each criterion of a rubric, and the verdicts are aggregated into a reward, most often by a weighted sum. We show that this additive aggregation is the weak point. Under a sum, criteria compensate for one another: a policy that misses the one decision that matters can buy the points back with advice nobody asked for. On clinical consultation, such a policy scores higher and answers worse. Rubric coverage rises while appropriateness on held-out physician criteria falls below the untrained model. The medical criteria are not to blame. Grouped so that they must hold together, the same criteria, unchanged to the word, recover a third of the loss; shorter answers recover almost none. We therefore propose Protocol-level Rubrics (ProRubric), which keeps what the criteria ask for and changes how they are aggregated. It groups a checklist into a few protocol-level dimensions. A dimension counts only when all of its criteria hold and its failure clause does not fire. The grouping is done once, offline, and leaves the optimizer unchanged. ProRubric raises appropriateness by 10.8 points without losing coverage and has the best seven-benchmark average at both scales. Reward validity is set not only by what a rubric verifies, but by how it aggregates. Code is available at https://github.com/Estrellajer/ProRubric
Problem

Research questions and friction points this paper is trying to address.

Reward Hacking
Rubric-based Reinforcement Learning
Additive Aggregation
Large Language Models
Clinical Consultation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reward Hacking
Rubric-based Reinforcement Learning
Protocol-level Rubrics
Reward Aggregation
Large Language Models
🔎 Similar Papers
No similar papers found.