RH-Detect: A Unified Benchmark for Reward Hacking Detection

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of standardized datasets and the difficulty of reproducible comparisons in reward hacking detection by constructing the first cross-source benchmark with a unified schema, integrating eleven public datasets. Methodologically, it evaluates language models as training-free detectors and explores the training benefits of single-token SFT and GRPO, employing AUROC for systematic evaluation. Experiments demonstrate that the optimal model achieves an AUROC of 0.962 and an accuracy exceeding 93%. Furthermore, the analysis reveals that a single aggregate score obscures significant performance disparities across data sources, and that detection capability degrades sharply in multi-turn tool-calling scenarios. Overall, this work provides a standardized evaluation framework and key empirical insights to advance research in reward hacking detection.
📝 Abstract
Reward hacking, where a model exploits an evaluation signal without completing the intended task, threatens the reliability of deployed language model systems. Existing datasets use different labels, response formats, and metadata conventions, making detector results difficult to compare. We present RH-Detect, a benchmark that combines reward-hacking-relevant subsets from eleven public datasets, comprising 92,761 rows and six behavior categories, into a common schema. On 5,021 open-ended evaluation units, each comprising a task prompt and a free-form model continuation, including multi-turn tool-use trajectories, we evaluate six off-the-shelf language models from five families as reward hacking detectors without additional training. The best model achieves a pooled AUROC of 0.962, with accuracy above 93%. For the four strongest models, however, accuracy on the two multi-turn tool-use datasets, MALT and TRACE, is 10.7-15.9 percentage points lower than on the other sources at a common decision threshold, highlighting a key gap for deployment-time monitoring. We find that different input formats have different effects across models. Removing thinking raises Qwen3.5-4B AUROC from 0.779 to 0.849, but lowers Qwen Flash from 0.977 to 0.950. We also evaluate the benchmark as a training dataset for detectors. Holding out each source in turn, single-token SFT improves average AUROC on five of six held-out sources. A GRPO follow-up on that failure case yields a slight improvement in detection performance. Our results show that a single pooled score can conceal variation across data sources, detector inputs, and training procedures.
Problem

Research questions and friction points this paper is trying to address.

reward hacking
benchmark
detection
language models
evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reward Hacking Detection
Unified Benchmark
LLM-as-a-Judge
Multi-turn Tool Use
Detector Training
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Junwei Quan
University of Toronto
E
Evgenii Opryshko
University of Toronto, Vector Institute, Trajectory Labs
Rohan Subramani
Rohan Subramani
University of Toronto
AI safetyLLM agents
Igor Gilitschenski
Igor Gilitschenski
Assistant Professor, University of Toronto
RoboticsMachine LearningComputer Vision