Score
Design and implement policy-optimization algorithms and training pipelines that incorporate explicit intermediate "milestone" signals or reward shaping (including milestone-based GRPO variants) to guide agent behavior. Build and analyze objective functions, credit-assignment mechanisms, and reward schedules that stabilize long-horizon optimization, encourage completion of intermediate milestones, and improve learning under sparse or delayed rewards.
This work addresses the challenges of credit misassignment and low sample efficiency in long-horizon language agents trained via reinforcement learning by introducing the BEACON framework. Leveraging the compositional structure of tasks, BEACON segments trajectories using milestones as anchors and employs segmented temporal reward shaping together with dual-scale advantage estimation to achieve precise credit assignment. This approach effectively decouples local action evaluation from interference caused by distant failures. Experimental results demonstrate that BEACON substantially outperforms GRPO and GiGPO across ALFWorld, WebShop, and ScienceWorld benchmarks. Notably, it achieves a 92.9% success rate on long-horizon ALFWorld tasks and improves sample utilization efficiency from 23.7% to 82.0%.
This paper addresses the Pass@K optimization objective in reinforcement learning with verifiable rewards (RLVR), where reward signals are sparse and only available upon successful task completion. Method: The authors establish the fundamental equivalence between direct policy gradient methods (e.g., REINFORCE) and advantage shaping techniques by reinterpreting advantage shaping as implicit maximization of a surrogate reward. Through reverse-engineering of existing algorithms—including GRPO and reward-regularized variants—they show that all implicitly optimize the same class of surrogate rewards. Building on this insight, they develop a unified framework that derives policy gradient algorithms systematically from surrogate reward specifications. Contribution/Results: This work provides the first theoretical unification of Pass@K policy gradient methods under RLVR, yielding a general analytical paradigm and principled design guidelines for algorithm development in verifiable-reward settings.
This work addresses the absence of a unified principled framework in existing large language model policy optimization methods, which obscures the mechanistic roles and design motivations of diverse algorithms within their objective functions. Starting from the expected reward objective, the paper constructs a diagnostic unified framework structured around two axes: trajectory-side and reward-side factors—centered on trajectory probabilities and rewards, respectively. Through first-principles derivation, this framework systematically integrates the evolutionary logic of REINFORCE, PPO, GRPO, and their variants (e.g., Agentic RL, GRPO-OPD), revealing compound failure modes that cannot be resolved by improvements on either axis alone and delineating the failure boundaries of current approaches. The proposed framework provides a principled foundation for joint policy optimization, offering both extensibility and diagnostic capability.
Addressing practical challenges in reinforcement learning—such as sparse and delayed rewards and training instability—this paper presents a systematic survey of reward engineering and reward shaping. We propose the first fine-grained taxonomy of reward design techniques, explicitly exposing their implicit assumptions and failure boundaries. Furthermore, we introduce an evaluation framework for reward shaping that jointly balances interpretability and empirical effectiveness. Our analysis integrates theoretical foundations of RL, deep RL practice, formal modeling of reward functions, and cross-domain applications—including robotics and autonomous driving. This work fills a critical gap by providing the first comprehensive, methodology-driven survey of reward design. It establishes a unified tripartite research framework comprising methodology, taxonomic classification, and application boundaries. The resulting synthesis delivers a reproducible, transferable engineering guide for algorithm designers, significantly enhancing the robustness and real-world deployability of RL systems. (149 words)
Sparse, delayed, and weakly informative reward signals severely hinder the efficiency of reinforcement learning, and existing reward shaping methods often fail to adapt to dynamic environments. This work proposes the first unified analytical framework encompassing temporal, informational, and theoretical dimensions to systematically categorize and compare twelve classes of dynamic reward shaping and related adaptive mechanisms. It clearly distinguishes between parameter corrections and state-dependent modifications, and precisely delineates the boundaries among additive shaping, reward replacement, and correlated guidance. Through integrated theoretical analysis and taxonomic synthesis, the study identifies conditions under which optimality is preserved in the presence of modern RL components such as experience replay, bootstrapped critics, and reward normalization, and for the first time elucidates the intrinsic relationship between adaptation rate and learning stability.
This study addresses the tendency of time-inconsistent agents to abandon long-term tasks prematurely by developing a continuous-time dynamic model that characterizes their behavior under deadlines and investigates optimal goal-setting and reward-scheduling mechanisms. It provides, for the first time in a continuous-time framework, analytical trajectories for the generalized hyperbolic discounting class, delineating precise conditions under which agents complete tasks, quit immediately, or partially disengage, while clarifying the fundamental differences between continuous and discrete interventions. Leveraging variational methods and optimal control theory, the work derives optimal goals both when exploitative rewards are permitted and prohibited. It further proves that, for a fixed number of stages, equal-length intervals paired with uniform rewards are optimal, and that progressively finer reward segmentation monotonically enhances final progress until reaching a discounting-independent limit.
This work addresses the challenges of sparse rewards and cross-turn credit assignment in multi-turn tool-use tasks, which hinder the effectiveness of reinforcement learning (RL). The authors propose a hybrid training framework integrating MT-GRPO and GTPO, trained on a realistic customer-service user simulator grounded in large language models (LLMs). They introduce an iterative reward calibration mechanism to refine per-turn reward design, enhancing reward discriminability. Notably, this is the first study to successfully apply RL training on the Tau-Bench benchmark. The proposed GTPO mixed advantage estimation effectively mitigates the misalignment between reward discriminability and advantage direction. Experimental results show performance gains of 2.9 and 11.5 percentage points for Qwen3.5-4B and Qwen3-30B-A3B, achieving 66.7% and 69.5% success rates, respectively—surpassing GPT-4.1/GPT-4o with the smaller model and approaching Claude Sonnet 4.5 with the larger one.