discriminator-guided rl

Designs and implements reinforcement learning training methods that use a learned discriminator's outputs (e.g., logits or scores) as the reward or shaping signal to guide a generator or policy. This includes building KL-regularized policy/objective updates and procedures that optimize the model on its own samples to improve perceptual or semantic realism.

discriminator-guidedrl

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.31
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Comprehensive Overview of Reward Engineering and Shaping in Advancing Reinforcement Learning Applications

Jul 22, 2024
SI
Sinan Ibrahim
🏛️ Skolkovo Institute of Science and Technology | Innopolis University

Addressing practical challenges in reinforcement learning—such as sparse and delayed rewards and training instability—this paper presents a systematic survey of reward engineering and reward shaping. We propose the first fine-grained taxonomy of reward design techniques, explicitly exposing their implicit assumptions and failure boundaries. Furthermore, we introduce an evaluation framework for reward shaping that jointly balances interpretability and empirical effectiveness. Our analysis integrates theoretical foundations of RL, deep RL practice, formal modeling of reward functions, and cross-domain applications—including robotics and autonomous driving. This work fills a critical gap by providing the first comprehensive, methodology-driven survey of reward design. It establishes a unified tripartite research framework comprising methodology, taxonomic classification, and application boundaries. The resulting synthesis delivers a reproducible, transferable engineering guide for algorithm designers, significantly enhancing the robustness and real-world deployability of RL systems. (149 words)

Complex Real-World ProblemsReinforcement LearningReward Mechanism

Bootstrapped Reward Shaping

Jan 02, 2025
JA
Jacob Adamczyk
🏛️ University of Massachusetts Boston | San José State University | Texas Tech University

To address slow convergence in reinforcement learning under sparse rewards, this paper proposes a bootstrapped potential-based reward shaping (V-PBRS) method grounded in state-value function estimation. Unlike conventional approaches, V-PBRS eliminates the need for handcrafted potential functions by dynamically constructing them directly from the current estimate of the state-value function. It rigorously preserves the optimality of the original policy and provides theoretical convergence guarantees. By integrating the potential function into both Q-learning and DQN frameworks—leveraging Monte Carlo or temporal-difference online estimation mechanisms—V-PBRS achieves stable and adaptive reward shaping. Empirical evaluation on the Atari benchmark demonstrates significantly accelerated training convergence, validating its effectiveness, robustness, and generalization capability in high-dimensional visual environments.

Reinforcement LearningSlow LearningSparse Rewards

Reward Design for Reinforcement Learning Agents

Mar 27, 2025
RD
Rati Devidze
🏛️ Saarland University

In reinforcement learning, reward function design faces critical challenges including delayed signals, ambiguity, misalignment with task objectives, and induction of undesirable behaviors. To address these, this paper proposes three novel reward mechanisms: teacher-driven, adaptive explainable, and agent-autonomous reward generation. Our core contributions are the first-ever adaptive explainable reward design method and a meta-learning–driven autonomous reward generation framework—enabling a paradigm shift from expert-guided reward specification to online inverse reward modeling by the agent. Technically, we integrate reward shaping, eXplainable AI (XAI)-informed reward modeling, policy-value alignment, and online inverse reward design. Experiments across multiple sparse-reward benchmarks demonstrate over 40% faster training convergence, significantly improved policy robustness, and high reward interpretability—validated by domain experts with 92% inter-rater agreement.

Creating adaptive interpretable rewards based on learner's policyDesigning informative reward signals for RL agentsDeveloping self-driven reward design via meta-learning

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games

Jun 30, 2025
AA
António Afonso
🏛️ SEED - Electronic Arts (EA) | KTH Royal Institute of Technology

Designing reward functions for reinforcement learning (RL) agents in games traditionally relies heavily on domain expertise and struggles to adapt to dynamic content changes. Method: This paper proposes an LLM-based automated iterative reward weight optimization method that takes user-specified behavioral objectives as input and leverages agent training feedback—such as success rate and episode length—to perform closed-loop, multi-round LLM reasoning for reward weight self-calibration, eliminating manual intervention. Contribution/Results: To our knowledge, this is the first work to integrate LLMs into online adaptive optimization of RL reward functions, substantially reducing dependence on human experts. Evaluated on a racing task, the approach improves agent success rate from 9% to 80% and reduces average lap steps to 855—performance approaching that achieved by expert manual tuning.

Adapts reward weights to game content changes automaticallyAutomates reward function tuning for RL agents in gamesUses language models to align behavior with goals

Highly Efficient Self-Adaptive Reward Shaping for Reinforcement Learning

Aug 06, 2024
HM
Haozhe Ma
🏛️ National University of Singapore | Nanyang Technological University

To address low sample efficiency and unstable convergence in reinforcement learning caused by sparse rewards, this paper proposes an adaptive reward shaping method grounded in historical success rates. The method models state-dependent success probability as a time-varying Beta distribution—explicitly capturing epistemic uncertainty for the first time in this context. It further introduces an uncertainty-driven stochastic annealing strategy that naturally balances exploration and exploitation. For scalable, model-free, nonparametric success-rate estimation in high-dimensional continuous state spaces, the approach integrates kernel density estimation (KDE) with Random Fourier Features. Experiments demonstrate substantial improvements in sample efficiency and convergence stability on extremely sparse-reward tasks, consistently outperforming state-of-the-art reward shaping and intrinsic motivation baselines across diverse benchmarks.

Addresses sparse-reward problem in reinforcement learningBalances exploration and exploitation with evolving Beta distributionsIntroduces self-adaptive reward shaping using historical success rates

Latest Papers

What's happening recently
View more

This study addresses the computational expense and instability of retraining from scratch when reward functions change during reinforcement learning (RL) post-training of foundation models. To this end, we propose a zero-shot policy prediction method that eliminates the need for additional RL training. Leveraging the low-rank subspace property of log-policies across different rewards, our approach directly estimates target policy weights under new reward functions through linear combinations of log-probabilities from existing policies. Experiments demonstrate that this method efficiently and accurately predicts RL post-training outcomes across both synthetic and real-world reward scenarios in text and image modalities. Consequently, it significantly reduces the computational overhead associated with multi-reward alignment.

Foundation ModelsPolicy PredictionPost-training

Sparse, delayed, and weakly informative reward signals severely hinder the efficiency of reinforcement learning, and existing reward shaping methods often fail to adapt to dynamic environments. This work proposes the first unified analytical framework encompassing temporal, informational, and theoretical dimensions to systematically categorize and compare twelve classes of dynamic reward shaping and related adaptive mechanisms. It clearly distinguishes between parameter corrections and state-dependent modifications, and precisely delineates the boundaries among additive shaping, reward replacement, and correlated guidance. Through integrated theoretical analysis and taxonomic synthesis, the study identifies conditions under which optimality is preserved in the presence of modern RL components such as experience replay, bootstrapped critics, and reward normalization, and for the first time elucidates the intrinsic relationship between adaptation rate and learning stability.

adaptive rewardsdynamic rewardreinforcement learning

This study addresses the issues of policy staleness bias in asynchronous reinforcement learning for large language models and the high sampling costs associated with GRPO by proposing the KLPO framework. This method anchors the KL regularization term to the sampling policy, enabling importance-weight-free updates through least-squares fitting of log-ratios. Furthermore, it derives a closed-form Gibbs solution that eliminates the need for partition function learning and group response dependencies, demonstrating that various existing algorithms constitute special cases thereof. Notably, KLPO supports single-trajectory, critic-free updates while strictly guaranteeing unbiased gradients, thereby substantially reducing computational overhead.

Asynchronous reinforcement learningCritic-free updateImportance weight bias

Hot Scholars

MR

Marcello Restelli

Full Professor, Politecnico di Milano
Machine LearningReinforcement Learning
JS

Jianzhun Shao

Alibaba Inc.
reinforcement learning in LLMmulti-agent reinforcement learningoffline reinforcement learning
RZ

Riccardo Zamboni

Politecnico di Milano
Reinforcement LearningMachine LearningArtificial Intelligence
SL

Shiyi Lan

NVIDIA
VisionLLM AgentVisual Gen
JZ

Julian Zimmert

Google Research
Bandit theoryReinforcement LearningMachine Learning