Score
Design, implement, and analyze mechanisms that dynamically adjust the relative weighting between forward and reverse Kullback–Leibler divergence terms during model training—e.g., learned controllers or policies that set fKL/rKL weights based on observed distributional characteristics and reward feedback. Evaluate how these adaptive weight schedules affect teacher–student distributional alignment, optimization stability, and downstream objective trade-offs.
This work addresses inconsistent implementations of KL regularization in Reinforcement Learning from Human Feedback (RLHF), where existing methods (e.g., GRPO) conflate the distinct functional roles of the KL term—as a reward correction versus an explicit loss. Through gradient analysis and equivalence derivation, we first prove that, under on-policy settings, optimizing “KL as a loss” is strictly gradient-equivalent to incorporating KL into the reward function. In contrast, we show that common off-policy implementations—such as adding KL as a separate loss term (e.g., $k_3$)—yield only a biased first-order approximation. To resolve this, we propose a unified framework grounded in reverse KL divergence modeling and importance sampling–based bias correction. Our theoretical analysis establishes a rigorous gradient-based foundation for KL regularization in RLHF, eliminating implementation-induced bias. Empirically, the proposed framework significantly improves training stability and sample efficiency.
This paper systematically identifies a gradient bias issue arising from KL regularization estimators in reinforcement learning (RL) training of large language models (LLMs): existing KL divergence estimation methods lack theoretical guarantees when integrated into objective functions, causing misalignment between the optimization target and the actual computed gradients. To address this, the authors propose an unbiased gradient configuration principle, empirically validated across both on-policy (PPO) and off-policy (DPO, GRPO) RL frameworks using Qwen2.5-7B, Llama-3.1-8B-Instruct, and Qwen3-4B-Instruct-2507. Their analysis provides the first empirical evidence that unbiased KL estimation significantly improves both in-domain and out-of-domain generalization. Moreover, in asynchronous offline RL settings, KL regularization is shown to play a previously unrecognized role—stabilizing training dynamics, suppressing gradient oscillations, and enhancing convergence robustness. These findings offer foundational insights for designing theoretically sound and empirically effective RLHF algorithms.
This work addresses mode collapse in KL-regularized reinforcement learning (RL). We systematically analyze how forward and reverse KL divergences affect multimodal coverage of the target distribution, identifying regularization strength and reward scaling as key determinants of mode coverage. We propose a theoretically grounded, scalable algorithm that adaptively optimizes the target distribution solely by adjusting reward magnitude—without requiring auxiliary diversity signals. We validate our method on post-training tasks for both large language models (LLMs) and chemical language models (CLMs). Empirical results demonstrate significant improvements in generation quality and diversity under both forward- and reverse-KL settings. Crucially, our approach remains robust even under strong KL regularization or low reward scales—regimes where conventional methods fail—thereby overcoming the inherent diversity limitation of existing KL-regularized RL frameworks.
In reinforcement learning, adaptive interaction data—where the behavior policy is nonstationary—invalidates standard estimators, undermining asymptotic normality for off-policy counterfactual policy evaluation and dynamic treatment effect (DTE) inference. To address this, we propose a weighted Z-estimation framework that constructs time-varying adaptive weights to stabilize heteroskedasticity, achieving, for the first time in the RL off-policy setting, both consistent and asymptotically normal DTE estimation. Our approach integrates dynamic causal inference with asymptotic statistical theory, enabling rigorous hypothesis testing and construction of uniformly valid confidence regions. Simulation studies and real-world RL experiments demonstrate substantial improvements in confidence interval coverage and statistical power. The method provides the first solution for structural parameter inference under adaptive experimentation that simultaneously offers theoretical guarantees—namely consistency, asymptotic normality, and uniform validity—and empirical robustness.
This work addresses the limitations of existing self-training methods for fine-tuning large language models, which are highly sensitive to synthetic data quality and suffer from diminishing margins between positive and negative samples during iterative optimization. To overcome these challenges, the authors propose the TPAW algorithm, which operates in a fully self-supervised setting by constructing a cooperative-competitive ensemble composed of the current policy model and historical checkpoints to engage in self-play. TPAW incorporates a dual adaptive weighting mechanism—comprising response reweighting and participant dynamic weighting—to enhance training stability and alignment efficacy. Requiring no human supervision and initialized solely from a supervised fine-tuned (SFT) model, TPAW iteratively refines model performance and consistently outperforms state-of-the-art baselines across multiple base models and LLM benchmarks, significantly improving alignment outcomes.
This work investigates the impact of exponential reward weighting in KL-regularized policy optimization on the performance of neural reward models, revealing the downstream policy’s sensitivity to errors in reward-skewed regions and its feedback interaction with feature learning. Under a Gaussian single-index model, the authors propose a two-stage analytical framework: first recovering the latent direction from reward-weighted samples, then fitting the output layer via weighted ridge regression. For the first time, the coupling between reward modeling and policy optimization is integrated into single-index theoretical analysis. By combining Hermite expansions with neural feature learning theory, they prove that when the feature-learning temperature exceeds a critical threshold, a constant fraction of neurons accurately recovers the latent direction. They further establish an upper bound on the value gap of the tilted policy, characterizing the trade-off governed by the deployment temperature β₂ between performance gain and learning cost.
This work addresses the instability and collapse of output diversity commonly encountered in post-training with reinforcement learning, as well as the lack of a unified design principle in existing advantage function methods. The authors propose FADE, a novel framework that systematically decouples the gradient weighting structure of the advantage function by decomposing it along the sign and difficulty axes into positive and negative gradient quality components. This decomposition reveals the dynamic trade-off between exploration and exploitation, enabling an adaptive scheduling mechanism that dynamically adjusts gradient weights to balance accuracy and diversity. Evaluated on 7B and 32B models, FADE achieves peak pass@1 performance 20k and 2k training steps earlier, respectively, and demonstrates state-of-the-art accuracy–diversity trade-offs on the LiveCodeBench and AIME benchmarks.
This work addresses the problem of automatically setting the KL regularization coefficient in reinforcement learning fine-tuning of language models, aiming to balance improvement in task reward against deviation from a reference policy. The authors propose a game-theoretic framework that formulates fine-tuning as a sequential game between an agent maximizing reward and a monitor detecting significant policy deviations. They provide the first statistically interpretable characterization of the KL coefficient in terms of detectability and derive a Pareto-optimal regularization parameter using concave-convex fractional programming theory. This approach transforms equilibrium computation into a tractable optimization problem compatible with standard fine-tuning pipelines. Experiments on Qwen3-8B and Llama-3.2-1B demonstrate superior reward–retention trade-offs in continual learning and enable auditing of model modifications by API providers.