Score
Designs and implements reinforcement-learning training procedures that generate multiple candidate outputs or trajectories per input and use their relative comparisons to construct adaptive baselines, winner-guidance, or training signals. This includes building two-stage or memory-shaping RL pipelines that perform multi-candidate self-comparison, apply winner-candidate guidance, and schedule (e.g., anneal) entropy regularization to stabilize learning.
Addressing practical challenges in reinforcement learning—such as sparse and delayed rewards and training instability—this paper presents a systematic survey of reward engineering and reward shaping. We propose the first fine-grained taxonomy of reward design techniques, explicitly exposing their implicit assumptions and failure boundaries. Furthermore, we introduce an evaluation framework for reward shaping that jointly balances interpretability and empirical effectiveness. Our analysis integrates theoretical foundations of RL, deep RL practice, formal modeling of reward functions, and cross-domain applications—including robotics and autonomous driving. This work fills a critical gap by providing the first comprehensive, methodology-driven survey of reward design. It establishes a unified tripartite research framework comprising methodology, taxonomic classification, and application boundaries. The resulting synthesis delivers a reproducible, transferable engineering guide for algorithm designers, significantly enhancing the robustness and real-world deployability of RL systems. (149 words)
Self-supervised goal-conditioned reinforcement learning (GCRL) has long suffered from slow simulation data acquisition and unstable training, hindering its broad adoption. This paper introduces JaxGCRL: the first high-performance JAX library and benchmark suite specifically designed for GCRL. It integrates GPU-accelerated replay buffers, parallel vectorized environments, and a stable contrastive RL framework. We systematically evaluate key design choices—including contrastive learning, vector-quantized goal representations, and self-supervised goal sampling—under unified experimental conditions. Empirical results demonstrate up to 22× speedup over prior implementations, enabling million-step training within minutes on a single GPU. JaxGCRL facilitates rapid iteration and reproducible evaluation across diverse, challenging environments. By significantly lowering the barrier to GCRL research, it provides empirically grounded guidance for algorithmic development and fosters standardized, scalable experimentation.
This paper addresses the fragmentation and weak theoretical foundations of Reinforcement Learning from Human Feedback (RLHF) in large language model alignment. We propose the first multi-stage collaborative optimization framework integrating economic incentive mechanisms, philosophical value reasoning, and optimal control theory. Methodologically, we systematically unify instruction tuning, Bradley–Terry reward modeling, Proximal Policy Optimization (PPO), Direct Preference Optimization (DPO), rejection sampling, and a structured human feedback protocol. Our contributions are threefold: (1) a modular, reproducible end-to-end RLHF practice guide; (2) clarification of key open challenges—including synthetic data generation and multi-dimensional alignment evaluation; and (3) enhanced model safety, controllability, and value consistency. The framework bridges rigorous theoretical grounding with practical engineering applicability, providing a principled methodology for deploying trustworthy large language models.
In reinforcement learning, Linear Temporal Logic (LTL) tasks suffer from sparse rewards, hindering subgoal guidance, slowing policy convergence, and degrading robustness. To address this, we propose a progress-aware adaptive reward shaping method: for the first time, we quantify LTL satisfaction progress as a continuous, differentiable reward signal and design an online mechanism to dynamically update the reward function according to the agent’s current learning state. Our approach integrates LTL task compilation, formal progress modeling, and deep RL frameworks (PPO/SAC). Evaluated across multiple benchmark environments, the method significantly accelerates convergence, increases average expected return by 23%, improves task completion rate by 31%, and outperforms both traditional sparse-reward baselines and handcrafted reward-shaping approaches in terms of robustness and generalization.
This work addresses reward-free, offline, image-driven goal-oriented robotic manipulation—enabling end-to-end vision-based grasping and relocation of real-world objects using only a single target image. Methodologically, we propose a contrastive learning–based self-supervised offline reinforcement learning framework, integrating an image encoder–action decoder architecture with goal-conditioned policy learning; crucial architectural designs and hyperparameter configurations are introduced to ensure stable training for real-hardware deployment. To the best of our knowledge, this is the first demonstration of contrastive self-supervised RL on a physical robotic arm. Our approach achieves a twofold improvement in task success rate over baseline methods. Critically, it requires no handcrafted reward functions, online environment interaction, or pixel-level annotations—significantly lowering the barrier to real-world deployment.
This work addresses the instability commonly observed in reinforcement learning training with large language models, which arises from architectural mismatches and precision discrepancies between training and inference—such as FP8 inference versus higher-precision training. To mitigate this issue, the paper proposes Adaptive Control Reinforcement Learning (ACRL), a method that dynamically regulates the divergence between training and inference within a bounded range, thereby stabilizing the learning process. Notably, this mechanism naturally enhances policy entropy without requiring explicit entropy regularization, leading to improved exploration efficiency. Experimental results demonstrate that under FP8 low-precision inference, ACRL achieves significantly better training stability compared to importance sampling correction methods and attains final performance on par with BF16 high-precision baselines.
This work addresses the challenge of poor generalization in autonomous racing when encountering novel track and surface conditions due to insufficient simulation coverage. To overcome this limitation, the authors propose a continual reinforcement learning framework based on continual backpropagation that trains a universal driving policy exclusively from real-world data. This approach represents the first implementation of purely real-data-driven continual reinforcement learning on the RoboRacer platform, further enhanced by offline reinforcement learning for policy fine-tuning and plasticity analysis. Experimental results demonstrate that the learned policy rapidly adapts to new scenarios within 15 minutes, significantly outperforming classical controllers and thereby validating its strong generalization capability and rapid adaptability in real-world environments.
This work addresses the strong dataset dependency in post-training for reinforcement learning, where conventional fixed scheduling strategies struggle to dynamically balance exploration and exploitation and fail to adaptively adjust hyperparameters such as regularization. To overcome this limitation, the authors propose a tree-search framework powered by large language model (LLM) agents that automatically diagnoses trajectory pathologies and jointly optimizes multiple hyperparameters across multi-stage training. The study uncovers, for the first time, structural patterns wherein capacity-related parameters exhibit monotonic accumulation while regularization parameters oscillate, leading to transferable adaptive scheduling principles that uniformly explain the common dynamics of policy behavior across diverse tasks. Evaluated on four GRPO benchmarks, the method achieves performance gains of 9%–140% over baselines, substantially outperforming grid search (+6%–15%), random search, and skill-based agents.
This work addresses the limitations of large language models in generating high-quality BPMN process models, which are constrained by supervised fine-tuning data and the absence of well-defined multidimensional reward functions. The authors propose a reinforcement learning–based optimization approach that systematically explores a reward function encompassing 38 syntactic, pragmatic, and semantic metrics. They train Llama-3.1-8B and Qwen2.5-14B models across 48 configurations and find that uniformly weighted rewards outperform targeted weighting schemes, with significant interaction effects observed between reward composition and model architecture. Leveraging Group Relative Policy Optimization and an automated evaluation framework, the method substantially improves pragmatic and syntactic quality while preserving semantic fidelity and reducing output variability by over sixfold. All code is publicly released.
Sparse, delayed, and weakly informative reward signals severely hinder the efficiency of reinforcement learning, and existing reward shaping methods often fail to adapt to dynamic environments. This work proposes the first unified analytical framework encompassing temporal, informational, and theoretical dimensions to systematically categorize and compare twelve classes of dynamic reward shaping and related adaptive mechanisms. It clearly distinguishes between parameter corrections and state-dependent modifications, and precisely delineates the boundaries among additive shaping, reward replacement, and correlated guidance. Through integrated theoretical analysis and taxonomic synthesis, the study identifies conditions under which optimality is preserved in the presence of modern RL components such as experience replay, bootstrapped critics, and reward normalization, and for the first time elucidates the intrinsic relationship between adaptation rate and learning stability.