Bootstrapped Mixed Rewards for RL Post-Training: Injecting Canonical Action Order

📅 2025-12-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Standard reinforcement learning (RL) post-training for reasoning tasks often neglects the structural constraints of solution processes, leading to suboptimal sequence generation. Method: We propose incorporating *canonical action-order prompts* into scalar rewards to guide models toward solver-like behavior. Our approach employs a hybrid reward function combining cell-level accuracy with coarse-grained ranking signals, optimized via Group Relative Policy Optimization (GRPO). A bootstrapped scaling mechanism balances multi-objective reward components without altering supervision data or model architecture. Results: Evaluated on structured reasoning tasks (e.g., Sudoku), our method significantly improves generalization—achieving test accuracy surpassing pure accuracy-optimized baselines and approaching the upper bound of full supervised fine-tuning on canonical-order data. Contribution: This work is the first to implicitly model solution-order structure as an optimizable scalar prompt within RL-based post-training, enabling efficient, structure-aware policy refinement without architectural or data modifications.

Technology Category

Search and Optimization: Learning to SearchMachine Learning: Structured LearningReasoning under Uncertainty: Stochastic Optimization

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Post-training with reinforcement learning (RL) typically optimizes a single scalar objective and ignores structure in how solutions are produced. We ask whether a scalar hint toward a canonical solver ordering, used only during RL post-training, improves performance even when fine-tuned on randomized solution sequences. On Sudoku, we train a Transformer with standard fine-tuning on randomized solving orders, then post-train it with Group Relative Policy Optimization (GRPO) with two rewards: cell accuracy and an ordering reward that increases when the model's emission order aligns with the solver order. To compare signals cleanly, we combine them via fixed mixtures and use a simple bootstrapped scaling to equalize component magnitudes at initialization. Mixed rewards generally outperform cell-only optimization--the best mixture yields substantially higher test accuracy than the fine-tuned-only model trained on random-order and approaches the fine-tuned-only model trained on solver-order sequences in accuracy. These results suggest that coarse ordering signals can steer RL post-training toward solver-order trajectories without modifying supervised data or architecture.
Problem

Research questions and friction points this paper is trying to address.

Improves RL post-training with canonical action ordering hints
Combines accuracy and ordering rewards for better solution trajectories
Enhances model performance without modifying supervised data or architecture
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixed rewards combine accuracy and ordering signals
Bootstrapped scaling equalizes reward component magnitudes
Post-training steers model toward canonical solver trajectories
🔎 Similar Papers