How to post-train on a surrogate: Envelope sampling mitigates reward hacking

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of large language models to reward hacking during post-training with cheap proxy rewards, a problem exacerbated by calibration bias. To mitigate this, the authors propose Envelope Sampling, a theoretically grounded method that recalibrates LLM judges using a small number of ground-truth annotations. This approach integrates rejection sampling with a regret minimization algorithm under an L2-ball constraint to optimize the modified reward function, thereby overcoming the limitations of conventional heuristics and effectively handling misjudgments on rare outputs. Empirical evaluations on clinical note generation and sycophancy tasks demonstrate that the proposed method significantly alleviates reward hacking, achieving superior performance compared to baseline sampling-based recalibration schemes.
📝 Abstract
Large language models (LLMs) are commonly post-trained against LLM judges and other cheap surrogates because the true reward, such as human preference, is too expensive to query at scale. This practice often leads to reward hacking, where reinforcement learning against a miscalibrated surrogate leads to undesirable side effects. In this work, we study a setting in which a small number $n$ of model outputs are annotated with ground-truth labels (e.g., from expert review) and used to recalibrate the LLM judge before optimizing against it. Prior approaches to judge recalibration are costly or heuristic, and it is known that on-policy sampling fails when the surrogate is miscalibrated on a rare set of outputs. In this work, we propose envelope sampling, a theoretically-grounded method for judge recalibration that seeks to minimize an upper bound on the regret of the post-trained model under the assumption that the human reward and re-calibrated reward lie in an $L^2$ ball around the judge. We give practical algorithms to sample from the envelope by rejection or by fine-tuning against a modified reward, and experiments on clinical note generation and on a controlled sycophancy task show that recalibrating on envelope samples mitigates reward hacking where recalibrating on base-model samples does not.
Innovation

Methods, ideas, or system contributions that make the work stand out.

envelope sampling
reward hacking
surrogate recalibration
regret minimization
post-training
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Sanjit Dandapanthula
Carnegie Mellon University Department of Statistics and Machine Learning Department
S
Shuvom Sadhuka
MIT EECS
S
Samir Khan
Abridge AI
M
Michael Oberst
Johns Hopkins University Department of Computer Science
Aaditya Ramdas
Aaditya Ramdas
Associate Professor (with tenure), Carnegie Mellon University
Machine LearningStatistics
Alexandra Chouldechova
Alexandra Chouldechova
Researcher @ MSR NYC FATE