Score
Designs, implements, and evaluates algorithms, training pipelines, and system components that learn, improve, compress, or transfer decision-making policies for agents — including proximal policy optimization (PPO/GRPO), policy gradient and policy-iteration methods, on-policy and off-policy algorithms — and builds policy engines, policy abstraction/localization, policy modeling, on-policy distillation, and policy-driven automation while measuring convergence, stability, and sample efficiency.
This paper addresses the limitation of the GRPO algorithm—its restriction to on-policy training—by systematically proposing and validating its first off-policy variant. Methodologically, we introduce a clipped surrogate objective function, adapt GRPO to the off-policy setting within the PPO framework, employ offline advantage estimation, and incorporate verifiable reward evaluation. We theoretically prove that this objective guarantees monotonic improvement in expected reward. Our key contributions are threefold: (1) the first successful extension of GRPO to the off-policy paradigm; (2) theoretical analysis demonstrating superior training stability and higher sample and memory efficiency compared to on-policy GRPO; and (3) empirical validation showing that off-policy GRPO matches or significantly outperforms the original on-policy version across multiple benchmark tasks.
Existing LLM instruction-tuning algorithms—such as supervised fine-tuning (SFT), proximal policy optimization (PPO), and direct preference optimization (DPO)—are often explained with heavy reliance on prior knowledge, omit critical derivations, or remain overly abstract, resulting in high cognitive barriers and poor interpretability. Method: This paper systematically unifies mainstream reinforcement learning and preference optimization approaches under a concise, symbolically grounded derivation framework explicitly tailored to practical LLM training scenarios. Contribution/Results: We introduce GRAPE (Generalized Relative Advantage Policy Evolution), a novel paradigm for future preference learning designed to overcome fundamental limitations of current methods in objective design, training stability, and generalization. The framework provides a coherent, step-by-step exposition—from SFT through DPO—enhancing algorithmic intuition and theoretical transparency. It establishes a rigorous foundation for advancing preference-based LLM alignment and offers principled directions for subsequent research.
To address the low sample efficiency and weak exploration capability of online reinforcement learning in resource-constrained settings, this paper proposes DiffPPO—the first integration of denoising diffusion probabilistic models (DDPMs) into the PPO framework. DiffPPO synthesizes high-fidelity synthetic trajectories to augment offline datasets, enabling an offline-online hybrid training paradigm. Methodologically, it jointly leverages trajectory generation and importance reweighting to facilitate effective policy transfer from online learning to computationally limited offline environments. Experiments on challenging continuous control benchmarks demonstrate that DiffPPO significantly improves cumulative reward (+23.6%), accelerates convergence (reducing training steps by 37%), and enhances policy stability. The complete implementation is open-sourced, establishing a reproducible, diffusion-augmented PPO paradigm for sample-efficient RL.
Proximal Policy Optimization (PPO) suffers from high variance and high sample complexity, undermining training stability and efficiency. To address this, we propose HP3O—a novel PPO variant that introduces *trajectory recency modeling* into the PPO framework for the first time. Specifically, HP3O employs a FIFO trajectory replay buffer to reuse recent high-return trajectories and designs a mixed policy update mechanism under distributional shift constraints, jointly leveraging optimal and randomly sampled trajectories. We theoretically establish its monotonic policy improvement guarantee. Empirical evaluation on multiple continuous-control benchmark tasks demonstrates that HP3O significantly improves sample efficiency and training stability over standard PPO, A2C, and PPO-RND: it reduces policy gradient variance by 32% and consistently achieves superior final performance.
Existing policy gradient methods (e.g., PPO, TRPO) perform parameter updates solely along a single stochastic gradient direction, neglecting local geometric structure in parameter space and thus often converging to suboptimal policies. To address this, we propose ExploRLer—a plug-and-play local exploration enhancement framework that, without increasing the number of gradient updates, models the local geometry around policy checkpoints and systematically explores high-return regions within the current update neighborhood. ExploRLer is fully compatible with mainstream on-policy algorithms and requires no modification to existing training pipelines. Empirically, it significantly improves both convergence speed and final performance across multiple challenging continuous-control benchmarks. These results demonstrate that explicitly modeling and leveraging local parameter-space geometry is both effective and essential for optimizing reinforcement learning policies.
研究提出多步近端策略改进(MPI)方法,解决离线强化学习中策略更新需保持价值估计可靠同时超越行为分布的问题。
This work addresses the low sample efficiency of traditional model-free reinforcement learning methods—such as Proximal Policy Optimization (PPO)—which rely on high-variance advantage estimates. The authors propose Analytic Policy Gradients (APG), a method that leverages differentiable environment dynamics to compute exact, end-to-end gradients of policy returns with respect to policy parameters. To mitigate gradient degradation in long-horizon tasks, APG incorporates a segment-wise backpropagation mechanism and combines Monte Carlo estimation with critic-guided bootstrapping for effective gradient guidance. Evaluated on four continuous control benchmarks under identical network architectures and training protocols, APG consistently outperforms PPO, demonstrating substantially higher sample efficiency and faster convergence.
This work addresses the challenges of high inference cost, deployment difficulty, and opaque decision-making in deep reinforcement learning for power grid topology control. The authors propose a stress-focused data collection strategy to train a Proximal Policy Optimization (PPO) teacher model and, for the first time, distill it into interpretable, lightweight agents—specifically decision trees and random forests—targeting high-load critical states. The distilled models not only surpass the original PPO policy in average reward and survival duration while significantly reducing inference overhead, but also maintain highly consistent action outputs, enabling human auditability. Furthermore, the study reveals fundamental differences in feature dependencies between neural policies and tree-based models, achieving a balanced trade-off among performance, real-time responsiveness, and interpretability.
This study addresses the challenge of balancing expert guidance with autonomous exploration in on-policy reinforcement learning, where existing methods often converge to suboptimal solutions due to fragmented optimization objectives. To overcome this, we propose an adaptive expert-guidance mechanism that treats the expert intervention weight as a learnable parameter, jointly optimized with the policy under a unified on-policy objective. Built upon Proximal Policy Optimization (PPO) and an alternating control framework, our approach enables automatic decay of expert influence without requiring auxiliary components or complex scheduling heuristics. Extensive evaluations across 34 tasks demonstrate significant improvements in sample efficiency over strong baselines, with notably low hyperparameter sensitivity. Crucially, the expert weight naturally diminishes to zero as performance improves, facilitating a smooth transition from reliance on expert demonstrations to independent policy execution, ultimately yielding policies that surpass the guiding expert.
This work addresses the slow convergence and low sample efficiency of Generative Flow Networks (GFlowNets) in structured discrete sampling tasks by leveraging their theoretical connection to entropy-regularized reinforcement learning. For the first time, Proximal Policy Optimization (PPO) is integrated into the GFlowNet training framework. Through a systematic design of policy gradient updates, advantage estimation, and baseline mechanisms, the proposed approach substantially enhances training stability and data efficiency. Experimental results on benchmark tasks—including synthetic energy functions and molecular graph generation—demonstrate that the method achieves faster convergence and higher sampling efficiency compared to standard GFlowNet objectives.