Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR
This study addresses the absence of gradient signals in completely failed groups caused by limited sampling in reinforcement learning by proposing the GRAFT framework. Without requiring a designated teacher model, this approach leverages the complementarity of multiple models to exchange trajectories across them, enhancing policy learning via a gated replacement mechanism. Furthermore, it introduces sequence-level compatibility weighting and token-level importance ratio clipping to effectively mitigate distribution shift. Experimental results demonstrate that GRAFT achieves an average improvement of 2.1 points on mathematical reasoning benchmarks, with gains reaching up to 4.5 points, thereby enabling efficient teacher-free mutual learning.