rl fine-tuning

Designs and implements training pipelines that fine-tune pretrained policies or sequence models by optimizing reinforcement-learning objectives derived from scalar or composite reward signals, in online or offline/post-training settings. Builds reward functions, policy-update rules and regularizers (including sequence-level, stepwise, layer-specific, and depth-normalized variants), and joint training schemes (e.g., simultaneous ranking/retrieval and generation) to improve task-level metrics and model behavior without relying on direct supervised labels.

rlfine-tuning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.85
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$237K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the problem of automatically setting the KL regularization coefficient in reinforcement learning fine-tuning of language models, aiming to balance improvement in task reward against deviation from a reference policy. The authors propose a game-theoretic framework that formulates fine-tuning as a sequential game between an agent maximizing reward and a monitor detecting significant policy deviations. They provide the first statistically interpretable characterization of the KL coefficient in terms of detectability and derive a Pareto-optimal regularization parameter using concave-convex fractional programming theory. This approach transforms equilibrium computation into a tractable optimization problem compatible with standard fine-tuning pipelines. Experiments on Qwen3-8B and Llama-3.2-1B demonstrate superior reward–retention trade-offs in continual learning and enable auditing of model modifications by API providers.

KL regularizationreference policyregularization coefficient

Current large language model training typically introduces reinforcement learning (RL) only after pretraining and supervised fine-tuning (SFT), which constrains its full potential. This work proposes a novel paradigm that integrates RL and SFT directly during multiple stages of pretraining, exploring their concurrent optimization. By intervening at pretraining checkpoints, designing a target objective averaging mechanism, and carefully controlling data composition, the study demonstrates that introducing RL early can match or even surpass the performance of the conventional SFT→RL pipeline—particularly on challenging tasks—without compromising general capabilities. Moreover, strategic design of data composition proves more effective for performance gains than merely scaling up model size. These findings offer a new, efficient, and flexible pathway for aligning language models with desired behaviors.

Large Language ModelsPolicy OptimizationPre-training

This work addresses the unclear efficacy of offline pretraining for the Q-function in online reinforcement learning fine-tuning under pretrained policy strategies. It reveals a fundamental mismatch between the objectives of Q-function pretraining and online fine-tuning, which limits performance gains. To overcome this issue, the paper proposes Initialization via Policy Ensemble (IPE), a method that leverages rollout data from an ensemble of diverse policies to guide the initialization of the Q-function, thereby enabling more effective knowledge transfer. Evaluated across multiple continuous control benchmarks, IPE achieves an average 1.26× improvement in fine-tuning performance over naive Q-function pretraining, substantially enhancing online learning efficiency.

offline-to-online transferonline RL fine-tuningpolicy pretraining

Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Dec 10, 2024
ZZ
Zhiyuan Zhou
🏛️ UC Berkeley | Carnegie Mellon University

Online fine-tuning of offline pre-trained RL models typically requires continuous access to large-scale offline datasets, incurring high computational overhead, slow convergence, and risks of Q-function divergence and catastrophic forgetting due to distributional shift. Method: We theoretically establish, for the first time, that offline data are unnecessary during online fine-tuning, and propose Warm-start RL (WSRL)—a novel paradigm that initiates online adaptation using only a small number of rollouts generated by the pre-trained policy. WSRL integrates policy warmup, distribution-matching analysis, and an offline-to-online policy bridging mechanism, eliminating the need to store or revisit any offline data. Contribution/Results: Evaluated across multiple standard benchmarks, WSRL consistently outperforms state-of-the-art methods—both those retaining and discarding offline data—in final performance and sample efficiency. It accelerates convergence by 30–50%, achieves higher asymptotic returns, and reduces training cost by an order of magnitude.

Eliminates need for retaining offline data in RL fine-tuningEnhances performance without offline data constraintsPrevents value function divergence during online fine-tuning

Latest Papers

What's happening recently
View more

This work addresses the limitations of large language models in generating high-quality BPMN process models, which are constrained by supervised fine-tuning data and the absence of well-defined multidimensional reward functions. The authors propose a reinforcement learning–based optimization approach that systematically explores a reward function encompassing 38 syntactic, pragmatic, and semantic metrics. They train Llama-3.1-8B and Qwen2.5-14B models across 48 configurations and find that uniformly weighted rewards outperform targeted weighting schemes, with significant interaction effects observed between reward composition and model architecture. Leveraging Group Relative Policy Optimization and an automated evaluation framework, the method substantially improves pragmatic and syntactic quality while preserving semantic fidelity and reducing output variability by over sixfold. All code is publicly released.

LLMmulti-dimensional qualityprocess model generation

This work addresses the need for post-training to enhance the accuracy and reasoning reliability of large language models on specific tasks, while the boundaries and synergies between supervised fine-tuning (SFT) and reinforcement learning (RL) remain unclear. The study proposes a unified analytical framework to systematically compare SFT and RL in terms of objective formulation, algorithmic structure, and data requirements, revealing their intrinsic connections. Building on this analysis, the authors design an integrated strategy to establish an efficient hybrid post-training paradigm. Through empirical evaluation across representative applications from 2023 to 2025, the research identifies a clear trend toward hybrid post-training approaches and distills key practical guidelines, offering both theoretical grounding and methodological guidance for scalable, effective, and generalizable post-training of large language models.

Large Language ModelsPost-TrainingReinforcement Learning

Bootstrapped Mixed Rewards for RL Post-Training: Injecting Canonical Action Order

Dec 03, 2025
PG
Prakhar Gupta
🏛️ University of Michigan | Independent

Standard reinforcement learning (RL) post-training for reasoning tasks often neglects the structural constraints of solution processes, leading to suboptimal sequence generation. Method: We propose incorporating *canonical action-order prompts* into scalar rewards to guide models toward solver-like behavior. Our approach employs a hybrid reward function combining cell-level accuracy with coarse-grained ranking signals, optimized via Group Relative Policy Optimization (GRPO). A bootstrapped scaling mechanism balances multi-objective reward components without altering supervision data or model architecture. Results: Evaluated on structured reasoning tasks (e.g., Sudoku), our method significantly improves generalization—achieving test accuracy surpassing pure accuracy-optimized baselines and approaching the upper bound of full supervised fine-tuning on canonical-order data. Contribution: This work is the first to implicitly model solution-order structure as an optimizable scalar prompt within RL-based post-training, enabling efficient, structure-aware policy refinement without architectural or data modifications.

Combines accuracy and ordering rewards for better solution trajectoriesEnhances model performance without modifying supervised data or architectureImproves RL post-training with canonical action ordering hints

Rethinking Expert Trajectory Utilization in LLM Post-training

Dec 12, 2025
BD
Bowen Ding
🏛️ Zhejiang University | Westlake University | Huawei Noah's Ark Lab

This study addresses the efficient utilization of expert trajectories in large language model (LLM) post-training, proposing the Plasticity–Ceiling theoretical framework to systematically characterize the synergy between supervised fine-tuning (SFT) and reinforcement learning (RL). Methodologically, it analyzes trajectory selection, scheduling, and scaling through empirical and analytical lenses. Key contributions include: (1) refuting the empirical heuristic that “fewer expert trajectories are always better”; (2) establishing SFT-first followed by RL as the stable optimal paradigm, with a precise switching criterion based on the inflection point of validation loss; and (3) quantifying the complementary interplay between data scale (governing latent capacity) and trajectory difficulty (providing multiplicative gain), identifying minimal validation loss as a robust proxy for trajectory quality. The framework yields significant, reproducible performance gains across multiple LLM benchmarks, offering both theoretical foundations and actionable guidelines for post-training data strategy.

Deriving scaling guidelines for data, difficulty, and loss indicatorsEstablishing superior SFT-then-RL pipeline over synchronized methodsOptimizing expert trajectory use in LLM post-training

Existing process reward models (PRMs) rely on costly step-level human annotations or ground-truth reference solutions, limiting their applicability to domains like mathematical reasoning where gold-standard process annotations are unavailable. Method: We propose SPARK, the first framework for ground-truth-free process-level reward modeling. It employs a generator-verifier collaborative paradigm to produce diverse solution paths, integrates parallel self-consistency scoring, sequence-level meta-critique, and chain-of-thought verification (PRM-CoT) to construct synthetic verification data for fine-tuning a generative PRM, and incorporates format constraints to mitigate reward hacking. Contribution/Results: On ProcessBench, SPARK achieves 67.5 F1—surpassing the ground-truth-supervised baseline (66.4). Across six mathematical reasoning benchmarks, it attains a mean accuracy of 47.4%, significantly outperforming RLVR (43.9%) and establishing the first effective process-supervised reinforcement learning method without reference answers.

Addresses the need for expensive step-level annotations in process reward models.Enhances mathematical reasoning accuracy by aggregating multiple step-level verifications.Proposes a reference-free reinforcement learning framework using synthetic verification data.

Hot Scholars

DT

Dacheng Tao

Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processing
SL

Sergey Levine

UC Berkeley, Physical Intelligence
Machine LearningRoboticsReinforcement Learning
YG

Yu-Gang Jiang

Professor, Fudan University. IEEE & IAPR Fellow
Video AnalysisEmbodied AITrustworthy AI
GC

Ganqu Cui

Shanghai AI Lab
LLM AlignmentReinforcement Learning
LS

Lifeng Shang

Huawei Noah's Ark Lab
Machine LearningComputer VisionPattern ReconitionNatural Language Processing