How to post-train on a surrogate: Envelope sampling mitigates reward hacking
This study addresses the vulnerability of large language models to reward hacking during post-training with cheap proxy rewards, a problem exacerbated by calibration bias. To mitigate this, the authors propose Envelope Sampling, a theoretically grounded method that recalibrates LLM judges using a small number of ground-truth annotations. This approach integrates rejection sampling with a regret minimization algorithm under an L2-ball constraint to optimize the modified reward function, thereby overcoming the limitations of conventional heuristics and effectively handling misjudgments on rare outputs. Empirical evaluations on clinical note generation and sycophancy tasks demonstrate that the proposed method significantly alleviates reward hacking, achieving superior performance compared to baseline sampling-based recalibration schemes.