adaptive policy optimization

Designs, builds, and evaluates policy-optimization algorithms and update procedures for reinforcement learning that adaptively adjust learning progress, per-step or horizon selection, weighting schemes, and update rules to account for robustness, group/role structure, reflection or introspection, and stitching of policy fragments; this includes mutual-information–weighted updates, process-reward formulations, and robust/group variants (e.g., APO, RAPO, MRPO, R-GRPO) to improve stability under noisy supervision, sparse feedback, or heterogeneous agent roles.

adaptivepolicyoptimization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.25
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$225K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing LLM instruction-tuning algorithms—such as supervised fine-tuning (SFT), proximal policy optimization (PPO), and direct preference optimization (DPO)—are often explained with heavy reliance on prior knowledge, omit critical derivations, or remain overly abstract, resulting in high cognitive barriers and poor interpretability. Method: This paper systematically unifies mainstream reinforcement learning and preference optimization approaches under a concise, symbolically grounded derivation framework explicitly tailored to practical LLM training scenarios. Contribution/Results: We introduce GRAPE (Generalized Relative Advantage Policy Evolution), a novel paradigm for future preference learning designed to overcome fundamental limitations of current methods in objective design, training stability, and generalization. The framework provides a coherent, step-by-step exposition—from SFT through DPO—enhancing algorithmic intuition and theoretical transparency. It establishes a rigorous foundation for advancing preference-based LLM alignment and offers principled directions for subsequent research.

Explaining reinforcement learning algorithms for instruction tuningIntroducing new research directions with GRAPE frameworkProviding clear intuitive understanding of complex RL methods

This study addresses the optimization imbalance in offline reinforcement learning arising from the coupling of policy execution and value guidance roles. We propose RAPO (Role-based Adaptive Policy Optimization), a framework that adaptively modulates update coefficients by distinguishing between execution and bootstrapping roles, effectively decoupling TD3+BC and optimizing the temperature parameter in IQL. To our knowledge, this work is the first to enable independent regulation of policy coefficients serving distinct purposes, integrating deep reinforcement learning with advantage-weighted extraction techniques to enhance existing algorithms. Extensive evaluations on the D4RL benchmark demonstrate that RAPO significantly outperforms mainstream baselines, achieving particularly notable performance improvements among TD3+BC variants.

Critic BootstrappingOffline Reinforcement LearningPolicy Regularization

Adaptive Policy Learning to Additional Tasks

May 24, 2023
WH
Wenjian Hao
🏛️ Purdue University

This work addresses the problem of incremental adaptation of pre-trained policies to new tasks while preserving performance on original tasks. To this end, we propose Adaptive Policy Gradient (APG), the first method to integrate the Bellman optimality principle into the policy gradient framework, establishing a hybrid optimization paradigm that synergistically combines policy gradient updates with dynamic programming principles. We provide theoretical guarantees showing that APG achieves a convergence rate of O(1/T) and sample complexity of O(1/ε). Empirical evaluation on benchmark control tasks—including CartPole, LunarLander, and Robot Arm—demonstrates that APG attains performance comparable to deterministic policy gradient methods, yet with significantly fewer environment samples and faster convergence. These results highlight substantial improvements in both incremental learning efficiency and generalization stability.

Adapting pre-trained policies to new tasks without affecting original performanceImproving convergence rate and sample efficiency in policy learningValidating method effectiveness through challenging robotic simulation environments

This work addresses the training instability and heavy hyperparameter tuning burden in large language model reinforcement learning with PPO/GRPO, which stem from fixed clipping thresholds and static decoding temperatures. To overcome these limitations, the authors propose Adaptive Group-based Policy Optimization (AGPO), a value-network-free approach that leverages multidimensional statistics—such as reward distributions, entropy, and KL divergence—from a population of policies to construct a shared probing state. This state drives an adaptive clipping mechanism and a bidirectional temperature controller, dynamically modulating policy update magnitudes and exploration intensity. Evaluated across nine Chinese and English mathematical and STEM benchmarks, AGPO substantially outperforms PPO and GRPO, achieving 67.3% on GSM8K and 40.5% on MATH with Qwen2.5-14B, with consistent gains transferable to Llama-3-8B and Gemma-2-9B.

decoding temperaturefixed clippingLLM reasoning

This study addresses the lack of theoretical foundations for off-policy self-generated data fine-tuning in large model post-training, where convergence is difficult to guarantee when the sampling distribution is updated infrequently. To this end, this work proposes RE(S), a unified framework that formulates the optimization as a staged KL-divergence minimization process and provides rigorous analysis by integrating multi-armed bandits with a generalized REINFORCE algorithm under a Softmax policy. The authors prove that global convergence is achievable for any fixed S, establishing a tight O(1/T) convergence rate. Furthermore, they reveal a distinct advantage of off-policy learning: under weak initialization, appropriately increasing S helps escape local traps and significantly accelerates convergence, thereby breaking the conventional reliance on strict on-policy training.

convergence ratefine-tuninglearning dynamics

Latest Papers

What's happening recently
View more

Policy updates in reinforcement learning are highly sensitive to distributional shifts, a problem exacerbated in large-scale settings where discrepancies in numerical precision and sampling between training and inference introduce further instability. Existing approaches often rely on fixed hyperparameters, limiting their adaptability to variations in tasks, model scales, or data distributions. This work proposes a batch-adaptive policy optimization objective that dynamically modulates update intensity based on the effective sample size of policy ratios within each batch. By replacing fixed clipping with an adaptive mechanism grounded in the empirical distribution of ratios, the method jointly addresses trust-region constraints and off-policy data reliability without introducing additional hyperparameters. Empirical results demonstrate that the proposed approach matches or surpasses carefully tuned baselines across diverse settings, significantly enhancing algorithmic robustness and generalization.

distribution mismatchoff-policy learningpolicy optimization

This work addresses the lack of a unified theoretical framework for reinforcement learning, which has hindered systematic analysis of its convergence, sample complexity, and generalization. Building upon Markov decision processes and Bellman operators, the paper introduces a cohesive analytical framework that integrates tools from operator theory, stochastic approximation, convex duality, and function approximation. This framework encompasses a broad range of algorithms, including value iteration, policy iteration, temporal difference methods, off-policy learning, and constrained MDPs. By leveraging contraction mappings, monotone operators, martingale techniques, mirror/proximal optimization, concentration inequalities, and mixing process theory, the study establishes finite-sample performance bounds and asymptotic convergence guarantees for diverse reinforcement learning algorithms, thereby forging a rigorous theoretical bridge between probability theory, optimization, and statistics.

function approximationMarkov decision processesmathematical foundations

This study addresses the high sensitivity of large language model (LLM) agent policies to perturbations during reinforcement learning by proposing a stable perturbation-robust policy optimization method. The approach introduces an adaptive sensitivity-aware perturbation mechanism that dynamically quantifies the policy's responsiveness to perturbations, thereby theoretically guaranteeing monotonic policy improvement and training stability. Experimental evaluations on the ALFWorld and WebShop benchmarks demonstrate that the proposed method significantly enhances the resilience of LLM agents against interference. By effectively improving policy robustness while ensuring stable convergence throughout the optimization process, this work provides a principled solution for deploying reliable LLM-based agents in complex interactive environments.

Large Language Model AgentsPerturbation RobustnessPolicy Optimization

Current reinforcement learning (RL) post-training of large language models (LLMs) is overly focused on policy gradient methods such as PPO and GRPO, largely neglecting the broader RL algorithmic landscape. This work proposes a modular analytical framework centered on three core dimensions—MDP formulation, exploration strategies, and learning mechanisms—and systematically maps classical RL techniques—including value functions, off-policy learning, bootstrapped credit assignment, intrinsic motivation, tree search, and curriculum learning—onto the LLM training context for the first time. The study reveals a predominant reliance in existing approaches on actor-only, Monte Carlo–style policy optimization and explicitly identifies underexplored yet promising directions, thereby offering a clear roadmap for future algorithmic innovation in LLM alignment and training.

Credit AssignmentExplorationLarge Language Models

This work addresses the challenge that traditional reinforcement learning methods, such as GRPO, struggle to generate high-quality rollouts when tasks exceed the model’s current capabilities, resulting in weak gradient signals and training stagnation. To overcome this limitation, the authors propose a feedback-driven dual-objective cooperative reinforcement learning framework that, for the first time, leverages environmental feedback to guide exploration while jointly optimizing two complementary objectives: exploitation-oriented policy alignment (EPA) and exploration-oriented capability cultivation (ECC). This mechanism dynamically balances exploration and exploitation, effectively breaking through training bottlenecks. Experimental results demonstrate that, under the same number of rollouts, the proposed method achieves significantly faster training convergence and superior final performance compared to GRPO and feedback-based baselines, while maintaining higher policy entropy and lower gradient norms.

gradient directionpolicy updatereinforcement learning

Hot Scholars

MH

Marco Hutter

Professor of Robotics, ETH Zurich
Legged RoboticsRoboticsControl
YZ

Yuheng Zhang

University of Illinois Urbana-Champaign
Machine LearningReinforcement LearningOnline LearningBandits
FW

Furu Wei

Distinguished Scientist, Microsoft Research
Natural Language ProcessingArtificial IntelligenceGeneral AIGenerative AI
YY

Yang Yu

Professor, Nanjing University
Artificial IntelligenceReinforcement LearningEvolutionary Algorithms