Cross-Benchmark Transfer from RL on Agentic Coding Tasks

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the persistent failure of coding agents at the "last mile" of task completion, where they frequently overlook requirements or disrupt existing logic. To tackle this, we apply reinforcement learning post-training to the Kimi K2.7 Mixture-of-Experts (MoE) model, employing the GSPO algorithm with Rank-32 LoRA adapters and a composite reward mechanism grounded in hidden test cases. Our work demonstrates that reinforcement learning alone can substantially enhance the coding capabilities of large MoE models. The resulting model achieves statistically significant improvements in Pass@1 on unseen benchmarks such as SWE-Bench Pro, generates more concise reasoning trajectories, effectively circumvents common failure modes, and exhibits strong cross-framework generalization.
📝 Abstract
Coding agents often fail in the last mile: they build most of a feature but drop a requirement, test only the cases their implementation already handles, break behavior that was supposed to stay intact, or validate against an unchecked assumption. We ask whether reinforcement learning (RL) on expert-built agentic coding tasks closes this gap, and whether what the agent learns transfers beyond the training distribution. We post-train Kimi K2.7 Code, a 1T-parameter (32B active) open-weight mixture-of-experts model, with RL alone on 1,700 tasks: 1,000 repository tasks graded by hidden fail-to-pass tests and by pass-to-pass tests of existing behavior, and 700 terminal tasks graded by expert-written hidden verifiers. The reward is the fraction of target checks passed and drops to zero if any pass-to-pass test fails. One epoch of GSPO on a rank-32 LoRA adapter improves pass@1 on each of the six external benchmarks we evaluated, across three agent harnesses: SWE-Bench Pro (60.1 to 64.8), DeepSWE (31.0 to 43.4), Terminal-Bench 2.1 (67.4 to 82.0), Terminal-Bench 3 (1.4 to 12.1), Terminal-Bench 4 (0.0 to 7.6), and SWE-Marathon (5.0 to 25.0). Pooled over the five independent task sets (Terminal-Bench 4 revises Terminal-Bench 3), the improvement is significant (p < 0.001), and it remains significant on the three sets released after the training data was collected (p = 0.004); the model also improves under both harnesses never used in training. Median trajectories on DeepSWE and Terminal-Bench 3 are 24-35% shorter in agent steps. The base model's failed DeepSWE runs are mostly near-misses, and on the tasks the trained model newly solves, paired trajectories show it avoiding each of the four failure modes above.
Problem

Research questions and friction points this paper is trying to address.

coding agents
reinforcement learning
cross-benchmark transfer
agentic coding
last-mile failure
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Agentic Coding
Cross-Benchmark Transfer
Mixture-of-Experts
GSPO
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Sushant Mehta
Surge AI
L
Logan Ritchie
Surge AI
E
Edwin Chen
Surge AI