🤖 AI Summary
This study addresses the inefficiency of token-level credit assignment and update drift in reinforcement middle training by proposing a dual-critic method. The approach leverages single-trajectory calibrated advantage estimation, employing a conditional moment saddle-point objective to eliminate bias while preserving learning signals. Furthermore, it introduces an action-dependent weight mixing mechanism, for which the optimal mixing strategy and the upper bound of residual drift are theoretically derived, thereby avoiding the overhead of repeated sampling. By integrating PPO clipping corrections with shared positional information, the proposed method achieves a 7.8% performance improvement on standard benchmarks and reduces training step time by up to 63.4%.
📝 Abstract
Reinforcement mid-training lets language models learn internal thoughts from unlabeled text, but efficient token-level credit assignment remains challenging. Existing group-relative methods require costly repeated generation. Learned critics offer single-rollout feedback, but accurate return prediction alone does not ensure reliable policy updates. Our analysis shows how training--inference mismatch and PPO clipping prevent a common offset in advantage estimates from cancelling out, introducing additional update drift. We propose \tfour{}, a twin-critic method that calibrates token-level advantages from a single generated trajectory. After warmup and held-out qualification, the critics provide two advantage estimates, combined using action-dependent weights learned through a conditional-moment saddle-point objective. This objective brings the average advantage at each prefix toward zero, while a signal-retention constraint prevents the correction from erasing the learning signal. Sharing information across text positions avoids repeated sampling of each prefix. Theoretically, we characterize optimal mixing under the signal-retention constraint and establish an upper bound on residual mean-induced drift. Experiments show that, compared with the state-of-the-art critic-free method, \tfour{} improves mean benchmark performance by 7.8\% and reduces mean training-step time by up to 63.4\%.