Score
Designs and trains reinforcement-learning policies that automatically modify or correct noisy labels for supervised datasets, applied either on-the-fly during training/inference or as a post‑processing step per sample. This involves defining state and action representations for label adjustments, reward functions tied to downstream evaluation metrics, training and deployment procedures for the corrective agent, and methods to evaluate the quality of the rectified labels.
Noisy labels severely degrade model performance. To address this, this paper pioneers a reinforcement learning (RL) formulation for label correction, establishing an end-to-end closed-loop framework comprising states (joint sample-label representations), actions (label revision operations), and rewards (model performance gain after correction). We propose a deep feature-based Actor-Critic policy network that autonomously and adaptively rectifies noisy labels without requiring clean validation data or prior noise assumptions. Extensive experiments on multiple benchmark datasets demonstrate that our method consistently outperforms state-of-the-art robust learning approaches, achieving significant improvements in both classification accuracy and generalization across diverse noise settings. This work introduces a novel paradigm for learning with noisy labels, shifting from static noise modeling to dynamic, performance-driven label refinement.
This paper addresses the fragmentation and weak theoretical foundations of Reinforcement Learning from Human Feedback (RLHF) in large language model alignment. We propose the first multi-stage collaborative optimization framework integrating economic incentive mechanisms, philosophical value reasoning, and optimal control theory. Methodologically, we systematically unify instruction tuning, Bradley–Terry reward modeling, Proximal Policy Optimization (PPO), Direct Preference Optimization (DPO), rejection sampling, and a structured human feedback protocol. Our contributions are threefold: (1) a modular, reproducible end-to-end RLHF practice guide; (2) clarification of key open challenges—including synthetic data generation and multi-dimensional alignment evaluation; and (3) enhanced model safety, controllability, and value consistency. The framework bridges rigorous theoretical grounding with practical engineering applicability, providing a principled methodology for deploying trustworthy large language models.
In reinforcement learning, adaptive interaction data—where the behavior policy is nonstationary—invalidates standard estimators, undermining asymptotic normality for off-policy counterfactual policy evaluation and dynamic treatment effect (DTE) inference. To address this, we propose a weighted Z-estimation framework that constructs time-varying adaptive weights to stabilize heteroskedasticity, achieving, for the first time in the RL off-policy setting, both consistent and asymptotically normal DTE estimation. Our approach integrates dynamic causal inference with asymptotic statistical theory, enabling rigorous hypothesis testing and construction of uniformly valid confidence regions. Simulation studies and real-world RL experiments demonstrate substantial improvements in confidence interval coverage and statistical power. The method provides the first solution for structural parameter inference under adaptive experimentation that simultaneously offers theoretical guarantees—namely consistency, asymptotic normality, and uniform validity—and empirical robustness.
In reinforcement learning, reward modeling suffers from “error-regret mismatch”: low test error of the reward model does not guarantee low regret of the optimized policy, primarily due to distributional shift induced by policy optimization. Method: We provide the first theoretical proof that, for any arbitrarily small expected test error, there exist underlying data distributions yielding arbitrarily large regret. We construct explicit counterexamples, derive tight quantitative bounds linking reward estimation error and policy regret, and analyze the robustness of regularization techniques—including RLHF—against this mismatch. Contribution/Results: We show that low test error only ensures a worst-case regret upper bound, not actual policy performance; moreover, standard regularizers fail to eliminate the mismatch. Our analysis establishes a new theoretical benchmark for assessing reward model reliability and safety alignment in preference-based RL, with implications for trustworthy reward learning and deployment-critical applications.
Learning robust reward machines (RMs) from noisy execution traces remains challenging in reinforcement learning. Method: We propose a closed-loop co-learning framework featuring: (1) Bayesian posterior belief modeling to explicitly quantify trajectory uncertainty and noise tolerance; (2) an alternating online update mechanism jointly optimizing the RM and policy; (3) the first posterior-belief-based probabilistic reward shaping, enabling stable extraction of transferable RMs under high noise; and (4) integration of inductive logic programming (ILP), finite-state machine modeling, and online RM relearning. Results: Experiments show that the learned RMs closely approximate ground-truth structures under noise, and agents guided by them achieve performance on par with those using handcrafted RM baselines. The approach demonstrates significant improvements in robustness, transferability, and practical applicability.
Existing neural network editing methods rely on task-specific handcrafted algorithms, which are costly and exhibit poor generalization. This work proposes a unified, learnable framework by formulating model editing as a reinforcement learning problem for the first time. An agent learns to edit model parameters autonomously within two environments—MaskWorld (multiplicative mask scaling) and ShiftWorld (additive weight shifting)—guided by a multi-objective reward function that balances task-specific objectives with overall model performance preservation. Experiments demonstrate that the approach effectively reduces accuracy on forget sets to nearly 0% while maintaining over 90% accuracy on retain sets in machine unlearning tasks. In bias mitigation scenarios, it improves fairness metrics by more than 5% without compromising classification utility.
This work addresses a critical issue in existing process reward models based on Monte Carlo estimation (MCE), wherein policy-dependent labeling introduces noise that leads to incorrect rewards for reasoning steps. The study is the first to identify this problem and proposes a two-stage denoising framework to mitigate it. Initially, a large language model (LLM) is leveraged to detect reflective and self-correction behaviors to rectify noisy labels. Subsequently, a noise-aware iterative training mechanism dynamically refines these labels based on model confidence. By integrating LLM-based adjudication with iterative denoising, the approach substantially outperforms baseline methods in step-level correctness evaluation, achieving an absolute F1 score improvement of up to 27% and significantly enhancing the robustness of process reward modeling.
Policy updates in reinforcement learning are highly sensitive to distributional shifts, a problem exacerbated in large-scale settings where discrepancies in numerical precision and sampling between training and inference introduce further instability. Existing approaches often rely on fixed hyperparameters, limiting their adaptability to variations in tasks, model scales, or data distributions. This work proposes a batch-adaptive policy optimization objective that dynamically modulates update intensity based on the effective sample size of policy ratios within each batch. By replacing fixed clipping with an adaptive mechanism grounded in the empirical distribution of ratios, the method jointly addresses trust-region constraints and off-policy data reliability without introducing additional hyperparameters. Empirical results demonstrate that the proposed approach matches or surpasses carefully tuned baselines across diverse settings, significantly enhancing algorithmic robustness and generalization.
该研究通过解构强化学习后训练算法,探讨了其在提升大型语言模型能力时的机制和影响因素,如奖励信号、提示分布等,并分析了这些选择如何相互作用以影响后训练的成功。
This work addresses the vulnerability of large language models in reinforcement learning to noisy labels caused by scarce expert annotations, a challenge inadequately handled by existing methods. The study is the first to identify two distinct mechanisms through which label noise operates in reinforcement learning: inactive and active. To mitigate this issue, the authors propose an online label refinement strategy, OLR, which dynamically corrects labels by integrating rollout pass-rate slope and historical prediction consistency, further enhanced with majority voting and policy stability checks. Evaluated across six mathematical reasoning benchmarks and three out-of-distribution tasks under label noise ratios ranging from 0.1 to 0.9, OLR consistently improves average performance by 3.3%–4.6%, demonstrating significantly enhanced model robustness.