Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking

📅 2026-04-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Real-world reward signals—such as human preferences—are often uncertain, context-dependent, and inconsistent, leading to reward hacking and over-optimization. This work proposes a dual-source uncertainty-aware reward framework that treats uncertainty as a first-class component of the reward signal, explicitly modeling both epistemic uncertainty in value estimation (via ensemble prediction disagreement) and annotator variability in human preference labels. A confidence-adaptive reliability filter dynamically balances exploration and cautious decision-making based on uncertainty estimates. Experiments demonstrate that the method achieves more stable training in both discrete and continuous environments, reduces trap visits by 93.7%, maintains robust performance under 30% supervisory noise, and significantly mitigates reward hacking behaviors.
📝 Abstract
Reinforcement learning (RL) systems typically optimize scalar reward functions that assume precise and reliable evaluation of outcomes. However, real-world objectives--especially those derived from human preferences--are often uncertain, context-dependent, and internally inconsistent. This mismatch can lead to alignment failures such as reward hacking, over-optimization, and overconfident behavior. We introduce a dual-source uncertainty-aware reward framework that explicitly models both epistemic uncertainty in value estimation and uncertainty in human preferences. Model uncertainty is captured via ensemble disagreement over value predictions, while preference uncertainty is derived from variability in reward annotations. We combine these signals through a confidence-adjusted Reliability Filter that adaptively modulates action selection, encouraging a balance between exploitation and caution. Empirical results across multiple discrete grid configurations (6x6, 8x8, 10x10) and high-dimensional continuous control environments (Hopper-v4, Walker2d-v4) demonstrate that our approach yields more stable training dynamics and reduces exploitative behaviors under reward ambiguity, achieving a 93.7% reduction in reward-hacking behavior as measured by trap visitation frequency. We demonstrate statistical significance of these improvements and robustness under up to 30% supervisory noise, albeit with a trade-off in peak observed reward compared to unconstrained baselines. By treating uncertainty as a first-class component of the reward signal, this work offers a principled approach toward more reliable and aligned reinforcement learning systems.
Problem

Research questions and friction points this paper is trying to address.

reward hacking
uncertainty
reinforcement learning
human preferences
alignment failure
Innovation

Methods, ideas, or system contributions that make the work stand out.

uncertainty-aware reward
reward hacking
epistemic uncertainty
preference uncertainty
reliability filter
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Disha Singha