Range-GRPO: Policy Optimization via Pairwise Relations among Reward Intervals

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the uncertainty of pseudo-rewards and the high supervision costs in large language model (LLM) post-training. It proposes a semi-supervised framework that, for the first time, incorporates conformally calibrated reward interval uncertainty into pairwise comparisons as a replacement for pointwise scores, enabling efficient training through intra-group relative policy optimization. Theoretically, this work demonstrates that the proposed approach generalizes the advantages of Group Relative Policy Optimization (GRPO). Empirically, experiments show that it achieves state-of-the-art performance on both in-distribution and out-of-distribution data while requiring significantly fewer resources. Overall, by integrating conformal prediction with group-based optimization, this research offers a principled and resource-efficient solution to mitigating reward uncertainty during LLM alignment.
📝 Abstract
As the use of large language models (LLMs) expands, post-training has become increasingly important for adapting them to downstream tasks. However, obtaining reliable supervision remains costly, especially in domains without reference answers or executable verifiers. LLM-as-a-Judge provides scalable pseudo-rewards for unlabeled responses, but a single point score does not explicitly represent reward uncertainty. This motivates representing pseudo-rewards as conformally calibrated reward ranges. We propose Range-GRPO, a semi-supervised post-training framework that combines limited labeled data with unlabeled prompts. In Group Relative Policy Optimization (GRPO), learning signals depend on relative reward comparisons within each rollout group. The proposed objective compares reward ranges pairwise rather than reducing them to point rewards, allowing interval uncertainty to affect both the magnitude and direction of these signals. Our theoretical analysis characterizes this distinction and shows that the proposed objective recovers the Dr.GRPO advantage when all reward ranges collapse to points. Empirically, Range-GRPO achieves the highest in-distribution and out-of-distribution average performance among the evaluated semi-supervised methods while requiring fewer training resources.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Post-training
Reward Uncertainty
LLM-as-a-Judge
Policy Optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Range-GRPO
Conformal Calibration
Reward Uncertainty
Semi-supervised Post-training
Pairwise Interval Comparison
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ryunyi Lee
Yonsei University
K
Kangjun Noh
Yonsei University
S
Somin Kim
Yonsei University
H
Heedong Kim
Yonsei University
Kyungwoo Song
Kyungwoo Song
Yonsei University
Machine LearningDeep LearningNeural Networks