RankBuffer: Efficient Ranking-Based Rewards for Open-Ended Generation

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge that open-ended generation lacks standard answers, rendering pointwise rewards in group-based reinforcement learning difficult to calibrate, while existing pairwise ranking methods incur prohibitive evaluation costs. We propose a relative reward construction method based on query-specific ordered response buffers. Specifically, our approach maintains reusable quality scales and generates rewards through coarse-grained anchoring followed by local fine-grained ranking. To adapt to policy evolution, the buffer is dynamically managed via mechanisms such as boundary expansion and inactive anchor pruning. Experimental results demonstrate that the proposed method outperforms all pointwise baselines across four benchmarks. Furthermore, it achieves performance comparable to the strongest ranking baseline while substantially reducing evaluation costs.
πŸ“ Abstract
Open-ended generation lacks canonical answers, making pointwise rewards difficult to calibrate for group-based reinforcement learning. Directly ranking same-query rollouts provides a more suitable relative reward signal, but existing ranking-based reward methods can incur substantial judging cost. We introduce RankBuffer, which maintains an ordered, query-specific buffer of previously judged responses as a reusable quality scale. Each rollout is first inserted into an anchor interval through an independent coarse judgment, after which only rollouts assigned to the same interval undergo local fine ranking. The resulting complete order is converted into bounded rank rewards, while boundary expansion, local refinement, and inactive-anchor pruning adapt the buffer as the policy evolves. Across four open-ended benchmarks, RankBuffer consistently outperforms all pointwise baselines. It also achieves nearly on-par performance with the strongest ranking-based reward baseline while substantially reducing judging cost. Ablations demonstrate the importance of both local fine ranking and anchor response content, while buffer analyses show that rollout-derived anchors progressively extend and refine the covered quality scale. These results establish response reuse as an effective approach to efficient relative reward construction.
Problem

Research questions and friction points this paper is trying to address.

Open-ended generation
Ranking-based rewards
Judging cost
Reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

RankBuffer
Open-ended generation
Ranking-based rewards
Response reuse
Reinforcement learning
πŸ”Ž Similar Papers
No similar papers found.