Cross-Benchmark Transfer from RL on Agentic Coding Tasks
This study addresses the persistent failure of coding agents at the "last mile" of task completion, where they frequently overlook requirements or disrupt existing logic. To tackle this, we apply reinforcement learning post-training to the Kimi K2.7 Mixture-of-Experts (MoE) model, employing the GSPO algorithm with Rank-32 LoRA adapters and a composite reward mechanism grounded in hidden test cases. Our work demonstrates that reinforcement learning alone can substantially enhance the coding capabilities of large MoE models. The resulting model achieves statistically significant improvements in Pass@1 on unseen benchmarks such as SWE-Bench Pro, generates more concise reasoning trajectories, effectively circumvents common failure modes, and exhibits strong cross-framework generalization.