Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In reinforcement learning for code agents, binary test rewards assign identical advantages to all passing trajectories within a group, overlooking differences in implementation quality. This work proposes GAGAR, a quality-aware credit reallocation framework that uniquely combines dynamic sampling with intra-group joint scoring. Specifically, an SFT-trained agent scorer ranks passing trajectories by quality, and a sum-preserving weighted advantage redistribution mechanism refines the policy learning signal in GRPO. Experiments on the MiMo-V2.6 model series demonstrate that this approach effectively addresses the limitation of conventional RL in neglecting code cleanliness. It significantly enhances code agent performance while suppressing trajectory length inflation and improving training stability.
📝 Abstract
Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in implementation quality and adherence to task requirements. This leaves the policy without a learning signal that favors clean, targeted implementations over those containing unnecessary or out-of-scope changes. We introduce GAGAR, a framework for quality-aware credit redistribution in code agent RL. Built on dynamic sampling that retains groups containing both passing and failing trajectories, GAGAR places all trajectories from each group in a shared workspace, where an SFT-trained agentic grader jointly inspects them and ranks the test-passing candidates. Based on this ranking, we downweight lower-ranked trajectories and proportionally rescale the advantages of all test-passing trajectories to restore their original sum. This sum-preserving redistribution retains the relative weights established by quality-based downweighting while shifting credit toward higher-quality implementations. We evaluate GAGAR at industrial scale using pre-RL SFT checkpoints of MiMo-V2.6-Flash (310B total parameters) and MiMo-V2.6-Pro (1.02T total parameters). Controlled code-only Flash experiments show improved code agent performance, reduced trajectory-length growth, and more stable training. We further apply GAGAR in large-scale mixed-task RL with both Flash and Pro. Our results support combining test-based verification with groupwise agentic grading to improve the quality and stability of code agent RL.
Problem

Research questions and friction points this paper is trying to address.

Code Agent
Reinforcement Learning
GRPO
Reward Signal
Implementation Quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

Groupwise Agentic Grading
Advantage Redistribution
Code Agent RL
Quality-aware Credit Assignment
GRPO
🔎 Similar Papers
No similar papers found.
Jinhao Dong
Jinhao Dong
Peking University
SE Augments AITrustworthy Software DevelopmentPre-trainingCode Generation
L
Liang Zhao
LLM Core, Xiaomi
Zihao Yue
Zihao Yue
Renmin University of China
Multimodal AILanguage Modeling
W
Wenhan Ma
LLM Core, Xiaomi; Peking University
L
Linghao Zhang
LLM Core, Xiaomi
L
Lei Li
LLM Core, Xiaomi; University of Hong Kong
S
Shicheng Li
LLM Core, Xiaomi
Y
Yifan Song
LLM Core, Xiaomi
B
Bowen Ye
LLM Core, Xiaomi; Peking University
F
Fuli Luo
LLM Core, Xiaomi