Score
Designs and implements reinforcement-learning procedures that run after initial training to fine‑tune a pre‑trained model or policy using explicit reward signals or online policy optimization. This includes building reward models and optimization loops, selecting or engineering reward functions, and measuring post‑training changes in attributes such as fidelity and reliability to analyze and validate improvement.
This work addresses the unclear efficacy of offline pretraining for the Q-function in online reinforcement learning fine-tuning under pretrained policy strategies. It reveals a fundamental mismatch between the objectives of Q-function pretraining and online fine-tuning, which limits performance gains. To overcome this issue, the paper proposes Initialization via Policy Ensemble (IPE), a method that leverages rollout data from an ensemble of diverse policies to guide the initialization of the Q-function, thereby enabling more effective knowledge transfer. Evaluated across multiple continuous control benchmarks, IPE achieves an average 1.26× improvement in fine-tuning performance over naive Q-function pretraining, substantially enhancing online learning efficiency.
Addressing practical challenges in reinforcement learning—such as sparse and delayed rewards and training instability—this paper presents a systematic survey of reward engineering and reward shaping. We propose the first fine-grained taxonomy of reward design techniques, explicitly exposing their implicit assumptions and failure boundaries. Furthermore, we introduce an evaluation framework for reward shaping that jointly balances interpretability and empirical effectiveness. Our analysis integrates theoretical foundations of RL, deep RL practice, formal modeling of reward functions, and cross-domain applications—including robotics and autonomous driving. This work fills a critical gap by providing the first comprehensive, methodology-driven survey of reward design. It establishes a unified tripartite research framework comprising methodology, taxonomic classification, and application boundaries. The resulting synthesis delivers a reproducible, transferable engineering guide for algorithm designers, significantly enhancing the robustness and real-world deployability of RL systems. (149 words)
To address the high cost of human feedback and low sample efficiency in reward function learning for human-in-the-loop reinforcement learning, this paper proposes the Suboptimal Data Pretraining (SDP) framework. SDP enables cold-start training of reward models without human annotations by leveraging unlabeled, low-quality trajectory data—augmented with pseudo-labels derived from environment-minimum rewards. The method integrates pseudo-labeling, scalar reward modeling, and preference-based learning within a human-in-the-loop RL architecture. Evaluated across diverse simulated robotic tasks, SDP achieves significant improvements over state-of-the-art methods: it attains comparable or superior performance while reducing human interaction counts by over 50%. Crucially, SDP is compatible with both simulated and real human teachers and, for the first time, enables efficient reward modeling without any manual annotation.
Online fine-tuning of offline pre-trained RL models typically requires continuous access to large-scale offline datasets, incurring high computational overhead, slow convergence, and risks of Q-function divergence and catastrophic forgetting due to distributional shift. Method: We theoretically establish, for the first time, that offline data are unnecessary during online fine-tuning, and propose Warm-start RL (WSRL)—a novel paradigm that initiates online adaptation using only a small number of rollouts generated by the pre-trained policy. WSRL integrates policy warmup, distribution-matching analysis, and an offline-to-online policy bridging mechanism, eliminating the need to store or revisit any offline data. Contribution/Results: Evaluated across multiple standard benchmarks, WSRL consistently outperforms state-of-the-art methods—both those retaining and discarding offline data—in final performance and sample efficiency. It accelerates convergence by 30–50%, achieves higher asymptotic returns, and reduces training cost by an order of magnitude.
Offline pretraining often suffers from rapid degradation and poor exploration during early online reinforcement learning. To address this, we propose a policy expansion mechanism that treats the frozen offline policy as a fixed behavioral prior, dynamically coordinating it with a learnable online policy. Our key contribution is the first adaptive dual-policy architecture, integrating a policy-ensemble-based gating mechanism with behavioral distribution matching constraints. This ensures the offline policy remains unupdated while continuously guiding exploration, while enabling the online policy to incrementally acquire novel behaviors. Evaluated on multiple continuous-control benchmarks, our method significantly improves sample efficiency and final performance, avoids initial performance collapse, and achieves more stable convergence—outperforming standard fine-tuning and policy distillation baselines across all metrics.
This work addresses the limitations of large language models in generating high-quality BPMN process models, which are constrained by supervised fine-tuning data and the absence of well-defined multidimensional reward functions. The authors propose a reinforcement learning–based optimization approach that systematically explores a reward function encompassing 38 syntactic, pragmatic, and semantic metrics. They train Llama-3.1-8B and Qwen2.5-14B models across 48 configurations and find that uniformly weighted rewards outperform targeted weighting schemes, with significant interaction effects observed between reward composition and model architecture. Leveraging Group Relative Policy Optimization and an automated evaluation framework, the method substantially improves pragmatic and syntactic quality while preserving semantic fidelity and reducing output variability by over sixfold. All code is publicly released.
This work addresses the misalignment between supervised learning training objectives and reinforcement learning decision goals in financial time series forecasting by proposing an end-to-end fine-tuning framework. The approach first pretrains a predictive model using supervised learning and then fine-tunes it via reinforcement learning, propagating policy gradients back through the original model. This method establishes an effective linkage between supervised pretraining and reinforcement-based fine-tuning, validated across three mainstream reinforcement learning algorithms. Experimental results demonstrate that the fine-tuned models achieve significantly improved performance on trading tasks, while exhibiting strong generalization and cross-market transfer capabilities, thereby offering a practical solution for real-world deployment of financial forecasting systems.
Current large language model training typically introduces reinforcement learning (RL) only after pretraining and supervised fine-tuning (SFT), which constrains its full potential. This work proposes a novel paradigm that integrates RL and SFT directly during multiple stages of pretraining, exploring their concurrent optimization. By intervening at pretraining checkpoints, designing a target objective averaging mechanism, and carefully controlling data composition, the study demonstrates that introducing RL early can match or even surpass the performance of the conventional SFT→RL pipeline—particularly on challenging tasks—without compromising general capabilities. Moreover, strategic design of data composition proves more effective for performance gains than merely scaling up model size. These findings offer a new, efficient, and flexible pathway for aligning language models with desired behaviors.
This work investigates the impact of dynamics and reward model errors on policy performance in imagination-based trajectory learning. By extending MDP error analysis to settings with learned reward models, the authors reformulate policy optimization under reward noise as a one-dimensional problem, targeting representations with low Lipschitz constants. Integrating theoretical error bounds, REINFORCE gradient estimation, and sample complexity analysis, they derive the optimal allocation ratio between dynamics and reward samples under a fixed sampling budget. The analysis theoretically establishes that zero-mean reward noise introduces no bias and that its variance decays with the number of imagined trajectories, yielding clear practical guidelines for sample-efficient policy training.
This work addresses the limitation of static constraints in reinforcement learning fine-tuning, which often suppress a model’s ability to explore superior solutions while preventing degenerate outputs. To overcome this trade-off, the authors propose a dynamic constraint mechanism that employs a reference model as an online corrector, applying minimal intervention only when degenerate outputs are detected. This approach is combined with supervised fine-tuning loss to guide the model toward high-quality responses, allowing the constraint strength to adaptively scale with output quality. Evaluated on dialogue and code generation tasks, the method significantly outperforms both KL-regularized and unconstrained baselines, achieving higher task rewards without compromising training stability—thus effectively balancing exploration capability with constraint efficacy.