🤖 AI Summary
This work addresses the challenge of balancing reference-guided and free-form generation in reinforcement learning for reasoning by proposing Adaptive Reference Guidance (ARG). By deriving a closed-form KL divergence for prefix continuation, ARG dynamically selects the optimal shared prefix length to efficiently construct correct trajectories under a fixed computational budget for GRPO training. Notably, this method eliminates the need for success rate estimation or additional generation overhead. Experiments on Qwen3 models across five mathematical benchmarks demonstrate that ARG achieves the highest aggregated pass@12 performance and competitive average sampling accuracy, effectively enhancing the sample efficiency of reasoning policy optimization.
📝 Abstract
A verified reference solution provides a correct trajectory for training a reasoning model. Alternatively, a prefix of the reference can guide the model in generating a trajectory of its own. How much reference guidance should we provide? We study this question through prefix continuation, where the model continues from a reference prefix and keeps the resulting trajectory if it passes verification, falling back to the reference otherwise. Since both procedures produce correct trajectories, we compare their distributions with the ideal distribution, the model's own distribution conditioned on successful verification. For one continuation, we derive the KL divergence in closed form, which, up to a bounded term, decreases with the product of the probability of generating a different correct trajectory and the reference surprisal, the negative log probability of the reference suffix given the prefix. Since a longer prefix tends to raise the former but lowers the latter, continuation success alone does not determine the preferred amount of guidance. From this analysis, we learn a prefix selector shared across training questions from continuation outcomes, without estimating success probabilities or additional generation. The resulting Adaptive Reference Guidance (ARG) constructs correct trajectories within a fixed generation budget, and we apply it to all-failure groups in Group Relative Policy Optimization (GRPO). Experiments on Qwen3-4B and Qwen3-8B across five mathematical reasoning benchmarks show that ARG achieves the highest aggregate pass@12 among the evaluated methods with competitive average sampled accuracy.