feedback-driven policy refinement

Designs and implements algorithms and training pipelines that refine existing policies using feedback signals—such as learned critics, step-level or dense rewards, or structured human/model feedback—by applying on-policy policy-gradient reinforcement learning and related techniques. Builds and analyzes methods for post-training fine-tuning (including KL-regularized updates) to improve properties like faithfulness, coherence, or process alignment while preserving prior behavior and handling cases without ground-truth targets.

feedback-drivenpolicyrefinement

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.53
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the problem of automatically setting the KL regularization coefficient in reinforcement learning fine-tuning of language models, aiming to balance improvement in task reward against deviation from a reference policy. The authors propose a game-theoretic framework that formulates fine-tuning as a sequential game between an agent maximizing reward and a monitor detecting significant policy deviations. They provide the first statistically interpretable characterization of the KL coefficient in terms of detectability and derive a Pareto-optimal regularization parameter using concave-convex fractional programming theory. This approach transforms equilibrium computation into a tractable optimization problem compatible with standard fine-tuning pipelines. Experiments on Qwen3-8B and Llama-3.2-1B demonstrate superior reward–retention trade-offs in continual learning and enable auditing of model modifications by API providers.

KL regularizationreference policyregularization coefficient

Residual Policy Gradient: A Reward View of KL-regularized Objective

Mar 14, 2025
PW
Pengcheng Wang
🏛️ University of California, Berkeley

Existing reinforcement learning and imitation learning approaches struggle to simultaneously satisfy novel task requirements and preserve desirable properties of prior policies, particularly lacking residual learning methods tailored for policy gradient frameworks. To address this, we propose Residual Policy Gradients (RPG), the first method to integrate the residual Q-learning paradigm into the policy gradient framework. RPG explicitly models a residual action distribution during policy updates, enabling controlled refinement of prior policies. Theoretically, we reinterpret the KL-regularized objective, uncovering its implicit maximum-entropy trade-off mechanism. Algorithmically, RPG unifies soft policy gradients with residual Q-value estimation to yield a differentiable, stable, and constraint-aware optimization objective. Evaluated on the MuJoCo benchmark, RPG significantly improves both stability and task adaptability in policy customization, establishing a new paradigm for gradient-based policy transfer.

Enables policy customization in gradient-based RL settings.Extends Residual Q-Learning to policy gradient methods.Reinterprets KL-regularized objective for reward-level balance.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology

Sep 04, 2025
YJ
Yuchen Jiao
🏛️ Chinese University of Hong Kong | University of Pennsylvania

Existing alignment methods—such as RLHF, human/internal feedback integration, test-time scaling, and diffusion guidance—lack a unified theoretical foundation, leading to instability (e.g., reward hacking, policy optimization divergence) and inflexibility in behavior control. Method: We establish the intrinsic equivalence among these paradigms, showing that soft-optimal N-sampling, resampling-based guidance, and reward modeling all instantiate implicit policy optimization. Building on this insight, we propose a novel resampling-based alignment framework that operates entirely at inference time: it dynamically reweights or resamples diffusion trajectories using heterogeneous feedback signals—explicit human ratings and implicit model self-assessments—without explicit RL training. Contribution/Results: Our approach eliminates reliance on unstable policy network updates and reward modeling pitfalls while preserving generation quality. It achieves superior controllability, robustness, and adaptability across diverse alignment objectives. Crucially, this work provides the first systematic theoretical unification of mainstream post-training alignment techniques under a coherent implicit optimization lens.

Connecting reinforcement learning feedback and test-time scalingExploring equivalences between human and internal feedback methodsIntroducing resampling for diffusion models without reinforcement learning

Existing LLM instruction-tuning algorithms—such as supervised fine-tuning (SFT), proximal policy optimization (PPO), and direct preference optimization (DPO)—are often explained with heavy reliance on prior knowledge, omit critical derivations, or remain overly abstract, resulting in high cognitive barriers and poor interpretability. Method: This paper systematically unifies mainstream reinforcement learning and preference optimization approaches under a concise, symbolically grounded derivation framework explicitly tailored to practical LLM training scenarios. Contribution/Results: We introduce GRAPE (Generalized Relative Advantage Policy Evolution), a novel paradigm for future preference learning designed to overcome fundamental limitations of current methods in objective design, training stability, and generalization. The framework provides a coherent, step-by-step exposition—from SFT through DPO—enhancing algorithmic intuition and theoretical transparency. It establishes a rigorous foundation for advancing preference-based LLM alignment and offers principled directions for subsequent research.

Explaining reinforcement learning algorithms for instruction tuningIntroducing new research directions with GRAPE frameworkProviding clear intuitive understanding of complex RL methods

Reinforcement Learning from Human Feedback

Apr 16, 2025
NL
Nathan Lambert

This paper addresses the fragmentation and weak theoretical foundations of Reinforcement Learning from Human Feedback (RLHF) in large language model alignment. We propose the first multi-stage collaborative optimization framework integrating economic incentive mechanisms, philosophical value reasoning, and optimal control theory. Methodologically, we systematically unify instruction tuning, Bradley–Terry reward modeling, Proximal Policy Optimization (PPO), Direct Preference Optimization (DPO), rejection sampling, and a structured human feedback protocol. Our contributions are threefold: (1) a modular, reproducible end-to-end RLHF practice guide; (2) clarification of key open challenges—including synthetic data generation and multi-dimensional alignment evaluation; and (3) enhanced model safety, controllability, and value consistency. The framework bridges rigorous theoretical grounding with practical engineering applicability, providing a principled methodology for deploying trustworthy large language models.

Detail optimization stages from tuning to alignmentExplore understudied topics in synthetic dataIntroduce core RLHF methods for quantitative backgrounds

Latest Papers

What's happening recently
View more

Policy updates in reinforcement learning are highly sensitive to distributional shifts, a problem exacerbated in large-scale settings where discrepancies in numerical precision and sampling between training and inference introduce further instability. Existing approaches often rely on fixed hyperparameters, limiting their adaptability to variations in tasks, model scales, or data distributions. This work proposes a batch-adaptive policy optimization objective that dynamically modulates update intensity based on the effective sample size of policy ratios within each batch. By replacing fixed clipping with an adaptive mechanism grounded in the empirical distribution of ratios, the method jointly addresses trust-region constraints and off-policy data reliability without introducing additional hyperparameters. Empirical results demonstrate that the proposed approach matches or surpasses carefully tuned baselines across diverse settings, significantly enhancing algorithmic robustness and generalization.

distribution mismatchoff-policy learningpolicy optimization

Current large language model training typically introduces reinforcement learning (RL) only after pretraining and supervised fine-tuning (SFT), which constrains its full potential. This work proposes a novel paradigm that integrates RL and SFT directly during multiple stages of pretraining, exploring their concurrent optimization. By intervening at pretraining checkpoints, designing a target objective averaging mechanism, and carefully controlling data composition, the study demonstrates that introducing RL early can match or even surpass the performance of the conventional SFT→RL pipeline—particularly on challenging tasks—without compromising general capabilities. Moreover, strategic design of data composition proves more effective for performance gains than merely scaling up model size. These findings offer a new, efficient, and flexible pathway for aligning language models with desired behaviors.

Large Language ModelsPolicy OptimizationPre-training

This work addresses the optimal allocation of post-training compute resources for reinforcement learning under a fixed FLOP budget. It introduces the first accounting framework that explicitly decomposes post-training computation into rollout/search, policy updates, and reward model evaluation, systematically quantifying the trade-offs among model scale, search intensity, number of learning steps, and feedback quality. Using GRPO with LoRA fine-tuning on the Qwen2.5 model family and combining rule-based and PRM rewards, the authors conduct large-scale ablation studies under a unified compute budget. Their findings reveal that the optimal allocation is highly sensitive to model size, total budget, reward type, and evaluation objective; notably, larger models incur higher per-inference costs, yielding fewer updates or rollouts within the same FLOP budget, thereby uncovering nonlinear coupling in compute allocation.

compute allocationFLOP budgetfoundation models

This work addresses a critical limitation in existing reinforcement learning post-training methods—the absence of mechanisms to validate the efficacy of policy updates, which often leads to optimization drift or collapse. To mitigate this, the authors propose the PIRL framework, which reframes the optimization objective from immediate reward maximization to cumulative policy improvement across training iterations. Central to this framework is the PIPO algorithm, which introduces, for the first time, a policy improvement feedback mechanism. By employing a sliding window to retrospectively validate historical baselines, PIPO establishes a self-correcting closed-loop optimization process that guarantees each policy update positively contributes to final performance. Empirical evaluations on mathematical reasoning benchmarks demonstrate that PIPO achieves superior stability and performance compared to GRPO and its variants.

Large Language ModelsOpen-loop OptimizationPolicy Improvement

This work addresses the limitation of static constraints in reinforcement learning fine-tuning, which often suppress a model’s ability to explore superior solutions while preventing degenerate outputs. To overcome this trade-off, the authors propose a dynamic constraint mechanism that employs a reference model as an online corrector, applying minimal intervention only when degenerate outputs are detected. This approach is combined with supervised fine-tuning loss to guide the model toward high-quality responses, allowing the constraint strength to adaptively scale with output quality. Evaluated on dialogue and code generation tasks, the method significantly outperforms both KL-regularized and unconstrained baselines, achieving higher task rewards without compromising training stability—thus effectively balancing exploration capability with constraint efficacy.

constraintsdegenerate outputsoptimization conflict

Hot Scholars

JR

Ji-Rong Wen

Gaoling School of Artificial Intelligence, Renmin University of China
Large Language ModelWeb SearchInformation RetrievalMachine Learning
MG

Maani Ghaffari

Assistant Professor, University of Michigan
RoboticsMachine LearningRobot PerceptionAutonomous Navigation
JZ

Jun Zhao

School of Marine Sciences, Sun Yat-sen University
ocean opticsremote sensingnumerical modeling
DC

Daphne Cornelisse

Graduate student, NYU
multi-agent systemsreinforcement learningimitation learning