🤖 AI Summary
This study investigates how pretraining configurations—specifically model scale and data volume—influence the effectiveness of subsequent reinforcement learning (RL) in enhancing reasoning capabilities. Using chess as a controlled testbed, the authors establish a comprehensive pipeline spanning pretraining, supervised fine-tuning (SFT), and RL to systematically analyze the evolution of reasoning ability. Key findings reveal that pretraining loss strongly predicts final post-RL performance, and the amount of pretraining data exhibits an approximately linear relationship with RL gains. Moreover, RL not only reinforces correct strategies acquired during SFT but also activates novel solution pathways absent in SFT, particularly on challenging problems. These patterns demonstrate robust generalization to mathematical reasoning tasks, offering the first controlled-environment characterization of the trajectory through which reasoning abilities evolve from pretraining to RL.
📝 Abstract
Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic questions remain open: (1) how do pretraining choices (model size, data) shape the returns to RL compute, and (2) what does RL actually do to the model? These questions are difficult to study in the standard LLM setting: pretraining corpora are vast and uncontrolled, making it hard to attribute behaviors to pretraining versus RL, and systematic compute sweeps across both stages are prohibitively expensive. To address these challenges, we use chess as a controlled testbed for studying reasoning across the full pretraining-to-post-training pipeline. We follow the standard LLM training pipeline by pretraining language models from 5M to 1B parameters on human chess games, supervised fine-tuning on synthetic reasoning traces, and running RL on chess puzzles with verifiable rewards. Using this framework, we find that the post-RL performance at given RL compute level is well-predicted from the pretraining loss, and slope of the RL reward curves improves approximately linearly with the pretraining tokens. Beyond scaling, we find that RL does not simply sharpen the SFT policy: on easy puzzles it amplifies correct moves the SFT policy already preferred, while on hard puzzles it surfaces correct moves that were nearly absent under SFT. We further test whether our findings transfer beyond chess by training a 1B language model on math-domain text, where the same predictive pattern emerges: longer-pretrained checkpoints reach higher post-RL performance and improve faster under RL. In sum, we provide a quantitative account of the pretraining-to-RL interface and a controlled testbed for studying the science of reasoning across the full pretraining-to-post-training pipeline.