Score
Design and implement policy-optimization algorithms that train sequence-generating policies by comparing groups of candidate outputs—using group-relative, anchor-based, or group-filtered objectives and sequence-level/sufficiency-driven criteria to form update signals. Analyze and evaluate these methods to produce stable RL-style updates that improve fidelity of structured outputs and reduce erroneous or hallucinated elements.
Existing RLHF algorithms suffer from training instability, low efficiency, and poor compatibility with Mixture-of-Experts (MoE) architectures when employing token-level importance ratios. To address these issues, this paper proposes Sequence-level Importance Sampling (SIS): it defines importance weights via the full-sequence likelihood ratio, integrates sequence-level reward modeling, and employs an adaptive clipping strategy to enable end-to-end sequence-level policy optimization. By operating at the sequence level, SIS avoids cumulative token-level bias, significantly enhancing training stability—particularly for MoE-based large language models—while simplifying RL system design and reducing engineering complexity. Experiments on the Qwen3 series demonstrate that SIS improves training efficiency by 23% over GRPO, increases average win rate by 4.8%, and exhibits superior generalization across multi-task alignment and long-text generation benchmarks.
Existing LLM instruction-tuning algorithms—such as supervised fine-tuning (SFT), proximal policy optimization (PPO), and direct preference optimization (DPO)—are often explained with heavy reliance on prior knowledge, omit critical derivations, or remain overly abstract, resulting in high cognitive barriers and poor interpretability. Method: This paper systematically unifies mainstream reinforcement learning and preference optimization approaches under a concise, symbolically grounded derivation framework explicitly tailored to practical LLM training scenarios. Contribution/Results: We introduce GRAPE (Generalized Relative Advantage Policy Evolution), a novel paradigm for future preference learning designed to overcome fundamental limitations of current methods in objective design, training stability, and generalization. The framework provides a coherent, step-by-step exposition—from SFT through DPO—enhancing algorithmic intuition and theoretical transparency. It establishes a rigorous foundation for advancing preference-based LLM alignment and offers principled directions for subsequent research.
This work addresses the challenge of policy optimization in non-verifiable tasks, where explicit correctness signals are absent and existing group-wise comparison methods struggle to apply. The authors propose Reference Relative Policy Optimization (RRPO), a framework that constructs positive and negative anchor sets through hierarchical conditional trajectories and employs a set-based contrastive learning objective to train a metric projection head. This approach generates contrastive advantage scores without requiring a ground-truth verifier and integrates a group-wise centered policy update mechanism. RRPO effectively extends relative policy optimization to non-verifiable settings, achieving performance on par with or superior to verifier-based methods across verifiable reasoning, open-ended generation, and post-supervised fine-tuning scenarios, while significantly outperforming weakly supervised baselines.
This work uncovers a structural discrepancy between reward optimization and training objectives in group-based reinforcement learning methods such as GRPO. By establishing a unified surrogate objective framework and integrating optimization dynamics modeling with theoretical analysis, the study identifies three systematic flaws: gradient bias on prefix tokens induced by non-uniform group weighting, AdamW’s insensitivity to reward scaling, and momentum-driven excursions beyond clipping boundaries. These findings provide a rigorous theoretical foundation and concrete directions for designing more robust and consistent post-training algorithms.
This paper addresses the limitation of the GRPO algorithm—its restriction to on-policy training—by systematically proposing and validating its first off-policy variant. Methodologically, we introduce a clipped surrogate objective function, adapt GRPO to the off-policy setting within the PPO framework, employ offline advantage estimation, and incorporate verifiable reward evaluation. We theoretically prove that this objective guarantees monotonic improvement in expected reward. Our key contributions are threefold: (1) the first successful extension of GRPO to the off-policy paradigm; (2) theoretical analysis demonstrating superior training stability and higher sample and memory efficiency compared to on-policy GRPO; and (3) empirical validation showing that off-policy GRPO matches or significantly outperforms the original on-policy version across multiple benchmark tasks.
Policy updates in reinforcement learning are highly sensitive to distributional shifts, a problem exacerbated in large-scale settings where discrepancies in numerical precision and sampling between training and inference introduce further instability. Existing approaches often rely on fixed hyperparameters, limiting their adaptability to variations in tasks, model scales, or data distributions. This work proposes a batch-adaptive policy optimization objective that dynamically modulates update intensity based on the effective sample size of policy ratios within each batch. By replacing fixed clipping with an adaptive mechanism grounded in the empirical distribution of ratios, the method jointly addresses trust-region constraints and off-policy data reliability without introducing additional hyperparameters. Empirical results demonstrate that the proposed approach matches or surpasses carefully tuned baselines across diverse settings, significantly enhancing algorithmic robustness and generalization.
This study addresses the lack of theoretical foundations for off-policy self-generated data fine-tuning in large model post-training, where convergence is difficult to guarantee when the sampling distribution is updated infrequently. To this end, this work proposes RE(S), a unified framework that formulates the optimization as a staged KL-divergence minimization process and provides rigorous analysis by integrating multi-armed bandits with a generalized REINFORCE algorithm under a Softmax policy. The authors prove that global convergence is achievable for any fixed S, establishing a tight O(1/T) convergence rate. Furthermore, they reveal a distinct advantage of off-policy learning: under weak initialization, appropriately increasing S helps escape local traps and significantly accelerates convergence, thereby breaking the conventional reliance on strict on-policy training.
This work addresses the credit assignment trap in long-horizon agent tasks, where sparse rewards lead to undersampling of critical state-changing actions and over-reinforcement of ineffective behaviors due to sampling imbalance. To mitigate this, the authors propose ProGPO, a method that introduces a progress signal based on first-visit observation coverage within entirely failed trajectories. This signal enables a progress-conditioned advantage estimation mechanism that treats state coverage as an intrinsic reward, thereby guiding the policy to prioritize exploration of novel states. Implemented within a population-based policy optimization framework, ProGPO is evaluated using Qwen2.5-1.5B/7B-Instruct models in ALFWorld and WebShop environments, demonstrating significant performance gains over existing baselines on both overall and challenging tasks, effectively alleviating the credit assignment trap and enhancing exploration efficiency.
This work addresses the instability and frequent training collapse of traditional REINFORCE algorithms in neural combinatorial optimization, which stem from reliance on rollout baselines that degrade in quality on complex instances. For the first time, it introduces baseline-free policy optimization methods from large model alignment—such as GRPO and P3O—into this domain. Built upon the RL4CO framework, the approach replaces explicit baselines with intra-batch advantage normalization, thereby eliminating the structural fragility associated with baseline maintenance. Evaluated on routing problems including TSP and CVRP, the method substantially enhances training robustness: it successfully avoids collapse on TSP-100 while achieving solution quality within 2% of the strong POMO baseline, all without requiring any external or rollout baselines.
This work addresses the challenge in group-based reinforcement learning where step-level advantage estimation struggles to balance fairness and coverage in long-horizon tasks. The authors propose ProGPO, a method that maintains action history consistency while leveraging state potentials to derive transferable credit signals that augment sparse intra-group comparisons, enabling context-consistent step-level policy optimization. Its key innovation lies in reliably estimating state potentials without requiring a learned critic, instead integrating exact prefix-action contrasts, semantic expansion, and inverse-variance fusion across historical depths. Experimental results demonstrate that ProGPO significantly outperforms existing RL baselines on ALFWorld and WebShop benchmarks and exhibits strong scalability when applied to the Qwen2.5-3B-Instruct model.