🤖 AI Summary
This study addresses the signal degradation and unreliable guidance in self-rewarding reinforcement learning, where response rewards are distorted by stochastic group sampling contexts. To mitigate this, we propose Group Marginal Advantage Estimation (GMAE), which transcends single-group context limitations by constructing response-level distributions and aggregating multi-context rewards to estimate expected advantages. This approach enables zero-label self-evolution of large language models without human annotation. Experiments across eight benchmarks and four base models demonstrate that GMAE achieves superior performance and robust cross-domain generalization while significantly enhancing policy optimization stability with minimal computational overhead.
📝 Abstract
Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels. Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly. However, a response's reward representation also depends on its randomly sampled group context, i.e., the other responses in its group. Using only one group-context realization may miss desired reward signals and provide unreliable guidance for policy optimization. To address this issue, we propose Group-Marginalized Advantage Estimation (GMAE), which aggregates reward realizations across possible contexts into a response-level distribution and estimates expected advantages. Experiments across eight benchmarks and four base models demonstrate strong performance and cross-domain generalization. GMAE also exhibits stable learning, low extra cost, and good applicability across training datasets and RL backbones.