Score
Designs and implements training pipelines that fine-tune pretrained policies or sequence models by optimizing reinforcement-learning objectives derived from scalar or composite reward signals, in online or offline/post-training settings. Builds reward functions, policy-update rules and regularizers (including sequence-level, stepwise, layer-specific, and depth-normalized variants), and joint training schemes (e.g., simultaneous ranking/retrieval and generation) to improve task-level metrics and model behavior without relying on direct supervised labels.
This work systematically investigates reinforcement learning (RL)-driven alignment and capability enhancement of large language models (LLMs), addressing three core challenges: instruction following, ethical compliance, and complex reasoning. Methodologically, it introduces a two-dimensional classification framework grounded in reward modeling and policy optimization to unify the analysis of prominent paradigms—including RLHF, DPO, RLAIF, GRPO, and RLVR. Empirical analysis reveals emerging trends: RLHF excels at foundational alignment, while RLVR significantly improves stepwise reasoning. The study identifies critical bottlenecks—reward gaming, multi-objective trade-offs, and computational overhead—and proposes novel directions: hybrid RL architectures and verifier-guided training. Collectively, this work delivers a principled technical roadmap and methodological foundation for developing safe, reliable, and scalable RL-augmented LLMs.
This work addresses the problem of automatically setting the KL regularization coefficient in reinforcement learning fine-tuning of language models, aiming to balance improvement in task reward against deviation from a reference policy. The authors propose a game-theoretic framework that formulates fine-tuning as a sequential game between an agent maximizing reward and a monitor detecting significant policy deviations. They provide the first statistically interpretable characterization of the KL coefficient in terms of detectability and derive a Pareto-optimal regularization parameter using concave-convex fractional programming theory. This approach transforms equilibrium computation into a tractable optimization problem compatible with standard fine-tuning pipelines. Experiments on Qwen3-8B and Llama-3.2-1B demonstrate superior reward–retention trade-offs in continual learning and enable auditing of model modifications by API providers.
Current large language model training typically introduces reinforcement learning (RL) only after pretraining and supervised fine-tuning (SFT), which constrains its full potential. This work proposes a novel paradigm that integrates RL and SFT directly during multiple stages of pretraining, exploring their concurrent optimization. By intervening at pretraining checkpoints, designing a target objective averaging mechanism, and carefully controlling data composition, the study demonstrates that introducing RL early can match or even surpass the performance of the conventional SFT→RL pipeline—particularly on challenging tasks—without compromising general capabilities. Moreover, strategic design of data composition proves more effective for performance gains than merely scaling up model size. These findings offer a new, efficient, and flexible pathway for aligning language models with desired behaviors.
This work addresses the unclear efficacy of offline pretraining for the Q-function in online reinforcement learning fine-tuning under pretrained policy strategies. It reveals a fundamental mismatch between the objectives of Q-function pretraining and online fine-tuning, which limits performance gains. To overcome this issue, the paper proposes Initialization via Policy Ensemble (IPE), a method that leverages rollout data from an ensemble of diverse policies to guide the initialization of the Q-function, thereby enabling more effective knowledge transfer. Evaluated across multiple continuous control benchmarks, IPE achieves an average 1.26× improvement in fine-tuning performance over naive Q-function pretraining, substantially enhancing online learning efficiency.
Online fine-tuning of offline pre-trained RL models typically requires continuous access to large-scale offline datasets, incurring high computational overhead, slow convergence, and risks of Q-function divergence and catastrophic forgetting due to distributional shift. Method: We theoretically establish, for the first time, that offline data are unnecessary during online fine-tuning, and propose Warm-start RL (WSRL)—a novel paradigm that initiates online adaptation using only a small number of rollouts generated by the pre-trained policy. WSRL integrates policy warmup, distribution-matching analysis, and an offline-to-online policy bridging mechanism, eliminating the need to store or revisit any offline data. Contribution/Results: Evaluated across multiple standard benchmarks, WSRL consistently outperforms state-of-the-art methods—both those retaining and discarding offline data—in final performance and sample efficiency. It accelerates convergence by 30–50%, achieves higher asymptotic returns, and reduces training cost by an order of magnitude.
This work addresses the limitations of large language models in generating high-quality BPMN process models, which are constrained by supervised fine-tuning data and the absence of well-defined multidimensional reward functions. The authors propose a reinforcement learning–based optimization approach that systematically explores a reward function encompassing 38 syntactic, pragmatic, and semantic metrics. They train Llama-3.1-8B and Qwen2.5-14B models across 48 configurations and find that uniformly weighted rewards outperform targeted weighting schemes, with significant interaction effects observed between reward composition and model architecture. Leveraging Group Relative Policy Optimization and an automated evaluation framework, the method substantially improves pragmatic and syntactic quality while preserving semantic fidelity and reducing output variability by over sixfold. All code is publicly released.
This work addresses the need for post-training to enhance the accuracy and reasoning reliability of large language models on specific tasks, while the boundaries and synergies between supervised fine-tuning (SFT) and reinforcement learning (RL) remain unclear. The study proposes a unified analytical framework to systematically compare SFT and RL in terms of objective formulation, algorithmic structure, and data requirements, revealing their intrinsic connections. Building on this analysis, the authors design an integrated strategy to establish an efficient hybrid post-training paradigm. Through empirical evaluation across representative applications from 2023 to 2025, the research identifies a clear trend toward hybrid post-training approaches and distills key practical guidelines, offering both theoretical grounding and methodological guidance for scalable, effective, and generalizable post-training of large language models.
Standard reinforcement learning (RL) post-training for reasoning tasks often neglects the structural constraints of solution processes, leading to suboptimal sequence generation. Method: We propose incorporating *canonical action-order prompts* into scalar rewards to guide models toward solver-like behavior. Our approach employs a hybrid reward function combining cell-level accuracy with coarse-grained ranking signals, optimized via Group Relative Policy Optimization (GRPO). A bootstrapped scaling mechanism balances multi-objective reward components without altering supervision data or model architecture. Results: Evaluated on structured reasoning tasks (e.g., Sudoku), our method significantly improves generalization—achieving test accuracy surpassing pure accuracy-optimized baselines and approaching the upper bound of full supervised fine-tuning on canonical-order data. Contribution: This work is the first to implicitly model solution-order structure as an optimizable scalar prompt within RL-based post-training, enabling efficient, structure-aware policy refinement without architectural or data modifications.
This study addresses the efficient utilization of expert trajectories in large language model (LLM) post-training, proposing the Plasticity–Ceiling theoretical framework to systematically characterize the synergy between supervised fine-tuning (SFT) and reinforcement learning (RL). Methodologically, it analyzes trajectory selection, scheduling, and scaling through empirical and analytical lenses. Key contributions include: (1) refuting the empirical heuristic that “fewer expert trajectories are always better”; (2) establishing SFT-first followed by RL as the stable optimal paradigm, with a precise switching criterion based on the inflection point of validation loss; and (3) quantifying the complementary interplay between data scale (governing latent capacity) and trajectory difficulty (providing multiplicative gain), identifying minimal validation loss as a robust proxy for trajectory quality. The framework yields significant, reproducible performance gains across multiple LLM benchmarks, offering both theoretical foundations and actionable guidelines for post-training data strategy.
Existing process reward models (PRMs) rely on costly step-level human annotations or ground-truth reference solutions, limiting their applicability to domains like mathematical reasoning where gold-standard process annotations are unavailable. Method: We propose SPARK, the first framework for ground-truth-free process-level reward modeling. It employs a generator-verifier collaborative paradigm to produce diverse solution paths, integrates parallel self-consistency scoring, sequence-level meta-critique, and chain-of-thought verification (PRM-CoT) to construct synthetic verification data for fine-tuning a generative PRM, and incorporates format constraints to mitigate reward hacking. Contribution/Results: On ProcessBench, SPARK achieves 67.5 F1—surpassing the ground-truth-supervised baseline (66.4). Across six mathematical reasoning benchmarks, it attains a mean accuracy of 47.4%, significantly outperforming RLVR (43.9%) and establishing the first effective process-supervised reinforcement learning method without reference answers.