Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the signal degradation and unreliable guidance in self-rewarding reinforcement learning, where response rewards are distorted by stochastic group sampling contexts. To mitigate this, we propose Group Marginal Advantage Estimation (GMAE), which transcends single-group context limitations by constructing response-level distributions and aggregating multi-context rewards to estimate expected advantages. This approach enables zero-label self-evolution of large language models without human annotation. Experiments across eight benchmarks and four base models demonstrate that GMAE achieves superior performance and robust cross-domain generalization while significantly enhancing policy optimization stability with minimal computational overhead.
📝 Abstract
Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels. Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly. However, a response's reward representation also depends on its randomly sampled group context, i.e., the other responses in its group. Using only one group-context realization may miss desired reward signals and provide unreliable guidance for policy optimization. To address this issue, we propose Group-Marginalized Advantage Estimation (GMAE), which aggregates reward realizations across possible contexts into a response-level distribution and estimates expected advantages. Experiments across eight benchmarks and four base models demonstrate strong performance and cross-domain generalization. GMAE also exhibits stable learning, low extra cost, and good applicability across training datasets and RL backbones.
Problem

Research questions and friction points this paper is trying to address.

self-rewarding reinforcement learning
group context
reward signal
policy optimization
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Rewarding Reinforcement Learning
Group-Marginalized Advantage Estimation
Zero-Label Self-Evolving
Large Language Models
🔎 Similar Papers
2022-08-09IEEE Transactions on Evolutionary ComputationCitations: 14