🤖 AI Summary
This study addresses the ongoing debate regarding why random rewards can enhance large language model (LLM) performance. We propose using reinforcement learning with random rewards as a probing mechanism to investigate model reachability, thereby distinguishing whether observed improvements reflect intrinsic representational capacity or task overfitting. Through analyses of OLMo checkpoints and comparative experiments with supervised fine-tuning, we link the spurious reward paradox to model reachability, establishing a novel evaluation perspective free from label leakage. Furthermore, this work reveals three distinct mechanisms governing reinforcement learning responses during pretraining and intermediate training stages, validating the generalizability of our approach. Ultimately, these findings provide a new theoretical framework for understanding the latent training upper bounds of large models.
📝 Abstract
We connect the spurious-reward paradox to a model's reachability and propose random-reward reinforcement learning (RL) as a useful tool for the probing enterprise, addressing a decade-long debate over what probing performance actually reveals about a model. There are two prevailing explanations for the surprising finding that even random rewards can improve the performance of large language models (LLMs): one attributes the gains to particular mechanisms within RL training; the other to data contamination. Our results motivate a different view: spurious-reward RL can probe a model's reachability, or what further training can attain from its current state under specified constraints, beyond what is reflected in its current performance. Two OLMo checkpoints with the same accuracy on synthetic arithmetic (3.5%), for example, reach 8.5% and 55% in their best runs under the same correctness-rewarded RL. Examining OLMo checkpoints across pre-training and mid-training reveals three distinct regimes of training response: early on, RL produces little improvement even when correct answers are rewarded; later in pre-training, rewarding correct answers becomes effective while random rewards remain weak; and, upon entering mid-training, even random rewards can produce large gains. A similar ordering appears in a number-masked supervised fine-tuning (SFT) analysis of these checkpoints, suggesting that the pattern is not specific to a particular RL mechanism. Moreover, RL with random rewards offers a distinctive perspective on what training can make an LLM do, since its reward signal supplies no information about which answers are correct. By asking what training can attain without correctness feedback, it addresses the label-leakage side of a central problem in decodability-based probing: whether a successful probe reveals the model's capabilities or learns the task itself.