rl label rectification

Designs and trains reinforcement-learning policies that automatically modify or correct noisy labels for supervised datasets, applied either on-the-fly during training/inference or as a post‑processing step per sample. This involves defining state and action representations for label adjustments, reward functions tied to downstream evaluation metrics, training and deployment procedures for the corrective agent, and methods to evaluate the quality of the rectified labels.

rllabelrectification

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Learning to Clean: Reinforcement Learning for Noisy Label Correction

Nov 24, 2025
MH
Marzi Heidari
🏛️ Carleton University | Amii

Noisy labels severely degrade model performance. To address this, this paper pioneers a reinforcement learning (RL) formulation for label correction, establishing an end-to-end closed-loop framework comprising states (joint sample-label representations), actions (label revision operations), and rewards (model performance gain after correction). We propose a deep feature-based Actor-Critic policy network that autonomously and adaptively rectifies noisy labels without requiring clean validation data or prior noise assumptions. Extensive experiments on multiple benchmark datasets demonstrate that our method consistently outperforms state-of-the-art robust learning approaches, achieving significant improvements in both classification accuracy and generalization across diverse noise settings. This work introduces a novel paradigm for learning with noisy labels, shifting from static noise modeling to dynamic, performance-driven label refinement.

Correcting noisy labels in datasets using reinforcement learning frameworkDeveloping policy network for iterative label correction through actor-critic methodImproving prediction model performance by addressing noisy label degradation

Reinforcement Learning from Human Feedback

Apr 16, 2025
NL
Nathan Lambert

This paper addresses the fragmentation and weak theoretical foundations of Reinforcement Learning from Human Feedback (RLHF) in large language model alignment. We propose the first multi-stage collaborative optimization framework integrating economic incentive mechanisms, philosophical value reasoning, and optimal control theory. Methodologically, we systematically unify instruction tuning, Bradley–Terry reward modeling, Proximal Policy Optimization (PPO), Direct Preference Optimization (DPO), rejection sampling, and a structured human feedback protocol. Our contributions are threefold: (1) a modular, reproducible end-to-end RLHF practice guide; (2) clarification of key open challenges—including synthetic data generation and multi-dimensional alignment evaluation; and (3) enhanced model safety, controllability, and value consistency. The framework bridges rigorous theoretical grounding with practical engineering applicability, providing a principled methodology for deploying trustworthy large language models.

Detail optimization stages from tuning to alignmentExplore understudied topics in synthetic dataIntroduce core RLHF methods for quantitative backgrounds

Post Reinforcement Learning Inference

Feb 17, 2023
VS
Vasilis Syrgkanis
🏛️ Stanford University | Hong Kong University of Science and Technology

In reinforcement learning, adaptive interaction data—where the behavior policy is nonstationary—invalidates standard estimators, undermining asymptotic normality for off-policy counterfactual policy evaluation and dynamic treatment effect (DTE) inference. To address this, we propose a weighted Z-estimation framework that constructs time-varying adaptive weights to stabilize heteroskedasticity, achieving, for the first time in the RL off-policy setting, both consistent and asymptotically normal DTE estimation. Our approach integrates dynamic causal inference with asymptotic statistical theory, enabling rigorous hypothesis testing and construction of uniformly valid confidence regions. Simulation studies and real-world RL experiments demonstrate substantial improvements in confidence interval coverage and statistical power. The method provides the first solution for structural parameter inference under adaptive experimentation that simultaneously offers theoretical guarantees—namely consistency, asymptotic normality, and uniform validity—and empirical robustness.

Address nonstationary variance in adaptive reinforcement learning environmentsDevelop weighted Z-estimation for dynamic treatment effect analysisEstimate counterfactual policies post reinforcement learning data collection

The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret

Jun 22, 2024
LF
Lukas Fluri
🏛️ University of Amsterdam | Oxford University | University of Cambridge

In reinforcement learning, reward modeling suffers from “error-regret mismatch”: low test error of the reward model does not guarantee low regret of the optimized policy, primarily due to distributional shift induced by policy optimization. Method: We provide the first theoretical proof that, for any arbitrarily small expected test error, there exist underlying data distributions yielding arbitrarily large regret. We construct explicit counterexamples, derive tight quantitative bounds linking reward estimation error and policy regret, and analyze the robustness of regularization techniques—including RLHF—against this mismatch. Contribution/Results: We show that low test error only ensures a worst-case regret upper bound, not actual policy performance; moreover, standard regularizers fail to eliminate the mismatch. Our analysis establishes a new theoretical benchmark for assessing reward model reliability and safety alignment in preference-based RL, with implications for trustworthy reward learning and deployment-critical applications.

Distributional shift during policy optimization causes error-regret mismatch.Learned reward functions may have low training error but high regret.Policy regularization techniques do not fully resolve error-regret mismatch.

Learning Robust Reward Machines from Noisy Labels

Aug 27, 2024
RP
Roko Parac
🏛️ Imperial College London | University of Brescia | Cardiff University

Learning robust reward machines (RMs) from noisy execution traces remains challenging in reinforcement learning. Method: We propose a closed-loop co-learning framework featuring: (1) Bayesian posterior belief modeling to explicitly quantify trajectory uncertainty and noise tolerance; (2) an alternating online update mechanism jointly optimizing the RM and policy; (3) the first posterior-belief-based probabilistic reward shaping, enabling stable extraction of transferable RMs under high noise; and (4) integration of inductive logic programming (ILP), finite-state machine modeling, and online RM relearning. Results: Experiments show that the learned RMs closely approximate ground-truth structures under noise, and agents guided by them achieve performance on par with those using handcrafted RM baselines. The approach demonstrates significant improvements in robustness, transferability, and practical applicability.

Ensuring robustness against noisy label inconsistenciesInterleaving reward machine and policy learningLearning robust reward machines from noisy traces

Latest Papers

What's happening recently
View more

Existing neural network editing methods rely on task-specific handcrafted algorithms, which are costly and exhibit poor generalization. This work proposes a unified, learnable framework by formulating model editing as a reinforcement learning problem for the first time. An agent learns to edit model parameters autonomously within two environments—MaskWorld (multiplicative mask scaling) and ShiftWorld (additive weight shifting)—guided by a multi-objective reward function that balances task-specific objectives with overall model performance preservation. Experiments demonstrate that the approach effectively reduces accuracy on forget sets to nearly 0% while maintaining over 90% accuracy on retain sets in machine unlearning tasks. In bias mitigation scenarios, it improves fairness metrics by more than 5% without compromising classification utility.

bias mitigationmachine unlearningneural model editing

This work addresses a critical issue in existing process reward models based on Monte Carlo estimation (MCE), wherein policy-dependent labeling introduces noise that leads to incorrect rewards for reasoning steps. The study is the first to identify this problem and proposes a two-stage denoising framework to mitigate it. Initially, a large language model (LLM) is leveraged to detect reflective and self-correction behaviors to rectify noisy labels. Subsequently, a noise-aware iterative training mechanism dynamically refines these labels based on model confidence. By integrating LLM-based adjudication with iterative denoising, the approach substantially outperforms baseline methods in step-level correctness evaluation, achieving an absolute F1 score improvement of up to 27% and significantly enhancing the robustness of process reward modeling.

Label NoiseMonte Carlo EstimationPolicy-dependent Rewards

Policy updates in reinforcement learning are highly sensitive to distributional shifts, a problem exacerbated in large-scale settings where discrepancies in numerical precision and sampling between training and inference introduce further instability. Existing approaches often rely on fixed hyperparameters, limiting their adaptability to variations in tasks, model scales, or data distributions. This work proposes a batch-adaptive policy optimization objective that dynamically modulates update intensity based on the effective sample size of policy ratios within each batch. By replacing fixed clipping with an adaptive mechanism grounded in the empirical distribution of ratios, the method jointly addresses trust-region constraints and off-policy data reliability without introducing additional hyperparameters. Empirical results demonstrate that the proposed approach matches or surpasses carefully tuned baselines across diverse settings, significantly enhancing algorithmic robustness and generalization.

distribution mismatchoff-policy learningpolicy optimization

This work addresses the vulnerability of large language models in reinforcement learning to noisy labels caused by scarce expert annotations, a challenge inadequately handled by existing methods. The study is the first to identify two distinct mechanisms through which label noise operates in reinforcement learning: inactive and active. To mitigate this issue, the authors propose an online label refinement strategy, OLR, which dynamically corrects labels by integrating rollout pass-rate slope and historical prediction consistency, further enhanced with majority voting and policy stability checks. Evaluated across six mathematical reasoning benchmarks and three out-of-distribution tasks under label noise ratios ranging from 0.1 to 0.9, OLR consistently improves average performance by 3.3%–4.6%, demonstrating significantly enhanced model robustness.

label noiseLLMsnoisy supervision

Hot Scholars

GC

Georgios Chochlakis

CS PhD fellow, University of Southern California
Machine LearningNLPSubjectivity
WZ

Weizhong Zhang

Fudan University
Machine LearningDeep LearningOptimization
AA

Anastasia Antsiferova

MSU AI Institute, ISP RAS, Innopolis University
machine learningcomputer visionvideo compressionadversarial robustness
PW

Peter Wu

School of Computer Science, Carnegie Mellon University
machine learning