post-training rl

Designs and implements reinforcement-learning procedures that run after initial training to fine‑tune a pre‑trained model or policy using explicit reward signals or online policy optimization. This includes building reward models and optimization loops, selecting or engineering reward functions, and measuring post‑training changes in attributes such as fidelity and reliability to analyze and validate improvement.

post-trainingrl

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.6
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the unclear efficacy of offline pretraining for the Q-function in online reinforcement learning fine-tuning under pretrained policy strategies. It reveals a fundamental mismatch between the objectives of Q-function pretraining and online fine-tuning, which limits performance gains. To overcome this issue, the paper proposes Initialization via Policy Ensemble (IPE), a method that leverages rollout data from an ensemble of diverse policies to guide the initialization of the Q-function, thereby enabling more effective knowledge transfer. Evaluated across multiple continuous control benchmarks, IPE achieves an average 1.26× improvement in fine-tuning performance over naive Q-function pretraining, substantially enhancing online learning efficiency.

offline-to-online transferonline RL fine-tuningpolicy pretraining

Comprehensive Overview of Reward Engineering and Shaping in Advancing Reinforcement Learning Applications

Jul 22, 2024
SI
Sinan Ibrahim
🏛️ Skolkovo Institute of Science and Technology | Innopolis University

Addressing practical challenges in reinforcement learning—such as sparse and delayed rewards and training instability—this paper presents a systematic survey of reward engineering and reward shaping. We propose the first fine-grained taxonomy of reward design techniques, explicitly exposing their implicit assumptions and failure boundaries. Furthermore, we introduce an evaluation framework for reward shaping that jointly balances interpretability and empirical effectiveness. Our analysis integrates theoretical foundations of RL, deep RL practice, formal modeling of reward functions, and cross-domain applications—including robotics and autonomous driving. This work fills a critical gap by providing the first comprehensive, methodology-driven survey of reward design. It establishes a unified tripartite research framework comprising methodology, taxonomic classification, and application boundaries. The resulting synthesis delivers a reproducible, transferable engineering guide for algorithm designers, significantly enhancing the robustness and real-world deployability of RL systems. (149 words)

Complex Real-World ProblemsReinforcement LearningReward Mechanism

Leveraging Sub-Optimal Data for Human-in-the-Loop Reinforcement Learning

Apr 30, 2024
CM
Calarina Muslimani
🏛️ University of Alberta

To address the high cost of human feedback and low sample efficiency in reward function learning for human-in-the-loop reinforcement learning, this paper proposes the Suboptimal Data Pretraining (SDP) framework. SDP enables cold-start training of reward models without human annotations by leveraging unlabeled, low-quality trajectory data—augmented with pseudo-labels derived from environment-minimum rewards. The method integrates pseudo-labeling, scalar reward modeling, and preference-based learning within a human-in-the-loop RL architecture. Evaluated across diverse simulated robotic tasks, SDP achieves significant improvements over state-of-the-art methods: it attains comparable or superior performance while reducing human interaction counts by over 50%. Crucially, SDP is compatible with both simulated and real human teachers and, for the first time, enables efficient reward modeling without any manual annotation.

Improving feedback efficiency in human-in-the-loop RLLeveraging sub-optimal data to pre-train reward modelsReducing human interactions for reward function learning

Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Dec 10, 2024
ZZ
Zhiyuan Zhou
🏛️ UC Berkeley | Carnegie Mellon University

Online fine-tuning of offline pre-trained RL models typically requires continuous access to large-scale offline datasets, incurring high computational overhead, slow convergence, and risks of Q-function divergence and catastrophic forgetting due to distributional shift. Method: We theoretically establish, for the first time, that offline data are unnecessary during online fine-tuning, and propose Warm-start RL (WSRL)—a novel paradigm that initiates online adaptation using only a small number of rollouts generated by the pre-trained policy. WSRL integrates policy warmup, distribution-matching analysis, and an offline-to-online policy bridging mechanism, eliminating the need to store or revisit any offline data. Contribution/Results: Evaluated across multiple standard benchmarks, WSRL consistently outperforms state-of-the-art methods—both those retaining and discarding offline data—in final performance and sample efficiency. It accelerates convergence by 30–50%, achieves higher asymptotic returns, and reduces training cost by an order of magnitude.

Eliminates need for retaining offline data in RL fine-tuningEnhances performance without offline data constraintsPrevents value function divergence during online fine-tuning

Policy Expansion for Bridging Offline-to-Online Reinforcement Learning

Feb 02, 2023
HZ
Haichao Zhang
🏛️ Horizon Robotics

Offline pretraining often suffers from rapid degradation and poor exploration during early online reinforcement learning. To address this, we propose a policy expansion mechanism that treats the frozen offline policy as a fixed behavioral prior, dynamically coordinating it with a learnable online policy. Our key contribution is the first adaptive dual-policy architecture, integrating a policy-ensemble-based gating mechanism with behavioral distribution matching constraints. This ensures the offline policy remains unupdated while continuously guiding exploration, while enabling the online policy to incrementally acquire novel behaviors. Evaluated on multiple continuous-control benchmarks, our method significantly improves sample efficiency and final performance, avoids initial performance collapse, and achieves more stable convergence—outperforming standard fine-tuning and policy distillation baselines across all metrics.

Bridging offline pre-training and online fine-tuning in reinforcement learningEnabling adaptive policy composition for improved exploration and performanceRetaining useful offline policy behaviors during online learning

Latest Papers

What's happening recently
View more

This work addresses the limitations of large language models in generating high-quality BPMN process models, which are constrained by supervised fine-tuning data and the absence of well-defined multidimensional reward functions. The authors propose a reinforcement learning–based optimization approach that systematically explores a reward function encompassing 38 syntactic, pragmatic, and semantic metrics. They train Llama-3.1-8B and Qwen2.5-14B models across 48 configurations and find that uniformly weighted rewards outperform targeted weighting schemes, with significant interaction effects observed between reward composition and model architecture. Leveraging Group Relative Policy Optimization and an automated evaluation framework, the method substantially improves pragmatic and syntactic quality while preserving semantic fidelity and reducing output variability by over sixfold. All code is publicly released.

LLMmulti-dimensional qualityprocess model generation

This work addresses the misalignment between supervised learning training objectives and reinforcement learning decision goals in financial time series forecasting by proposing an end-to-end fine-tuning framework. The approach first pretrains a predictive model using supervised learning and then fine-tunes it via reinforcement learning, propagating policy gradients back through the original model. This method establishes an effective linkage between supervised pretraining and reinforcement-based fine-tuning, validated across three mainstream reinforcement learning algorithms. Experimental results demonstrate that the fine-tuned models achieve significantly improved performance on trading tasks, while exhibiting strong generalization and cross-market transfer capabilities, thereby offering a practical solution for real-world deployment of financial forecasting systems.

financial forecastingfine-tuningreinforcement learning

Current large language model training typically introduces reinforcement learning (RL) only after pretraining and supervised fine-tuning (SFT), which constrains its full potential. This work proposes a novel paradigm that integrates RL and SFT directly during multiple stages of pretraining, exploring their concurrent optimization. By intervening at pretraining checkpoints, designing a target objective averaging mechanism, and carefully controlling data composition, the study demonstrates that introducing RL early can match or even surpass the performance of the conventional SFT→RL pipeline—particularly on challenging tasks—without compromising general capabilities. Moreover, strategic design of data composition proves more effective for performance gains than merely scaling up model size. These findings offer a new, efficient, and flexible pathway for aligning language models with desired behaviors.

Large Language ModelsPolicy OptimizationPre-training

This work investigates the impact of dynamics and reward model errors on policy performance in imagination-based trajectory learning. By extending MDP error analysis to settings with learned reward models, the authors reformulate policy optimization under reward noise as a one-dimensional problem, targeting representations with low Lipschitz constants. Integrating theoretical error bounds, REINFORCE gradient estimation, and sample complexity analysis, they derive the optimal allocation ratio between dynamics and reward samples under a fixed sampling budget. The analysis theoretically establishes that zero-mean reward noise introduces no bias and that its variance decays with the number of imagined trajectories, yielding clear practical guidelines for sample-efficient policy training.

dynamics model errorimagined rolloutsmodel-based reinforcement learning

This work addresses the limitation of static constraints in reinforcement learning fine-tuning, which often suppress a model’s ability to explore superior solutions while preventing degenerate outputs. To overcome this trade-off, the authors propose a dynamic constraint mechanism that employs a reference model as an online corrector, applying minimal intervention only when degenerate outputs are detected. This approach is combined with supervised fine-tuning loss to guide the model toward high-quality responses, allowing the constraint strength to adaptively scale with output quality. Evaluated on dialogue and code generation tasks, the method significantly outperforms both KL-regularized and unconstrained baselines, achieving higher task rewards without compromising training stability—thus effectively balancing exploration capability with constraint efficacy.

constraintsdegenerate outputsoptimization conflict

Hot Scholars

JC

Juan C. Pérez

AI Research Scientist, Meta
computer visionartificial intelligence
HX

Hongyu Xu

Research Scientist, Meta Reality Labs
Spatial PerceptionGenAIMultimodalRoomPlan
TF

Tao Feng

Department of Mathematics, Zhejiang University
information theorydiscrete mathematicscombinatorics
WZ

Wenhong Zhu

Shanghai Jiao Tong Unviersity
Natural Language Processing
JY

Junchi Yan

FIAPR & ICML Board Member, SJTU (2018-), SII (2024-), AWS (2019-2022), IBM (2011-2018)
Computational IntelligenceAI4ScienceMachine LearningAutonomous Driving