stagewise reinforcement learning

Designs and implements reinforcement-learning training regimes that partition learning into sequential stages or curricula, including stage definitions, transition criteria, reward shaping, and scheduling mechanisms to produce policies that evolve over stages. Builds and evaluates the training pipelines, algorithms, and analyses needed to refine behavior and safety, measure transfer and generalization between stages, and mitigate stage-to-stage overgeneralization or failure modes.

stagewisereinforcementlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Curriculum Reinforcement Learning for Complex Reward Functions

Oct 22, 2024
KF
Kilian Freitag
🏛️ Chalmers University of Technology | University of Gothenburg

Training deep reinforcement learning agents under multi-objective conflicting rewards often suffers from instability and difficulty balancing task performance against constraint satisfaction. To address this, we propose a two-stage reward curriculum learning framework: an initial phase optimizes a simplified reward to accelerate convergence, followed by a smooth transition to the full, complex reward. We introduce a novel Actor-Critic fidelity criterion for automatic, dynamic stage switching and design a flexible replay buffer enabling cross-phase sample reuse. Our approach integrates curriculum learning, dynamic reward shaping, and adaptive experience replay. Evaluated on the DeepMind Control Suite—including tasks with explicit constraints—and real-world mobile robot navigation, our method significantly outperforms non-curriculum baselines, achieving a more robust trade-off between task success rate and constraint violation rate.

Automates transition in reward curriculum stagesBalances task completion and constraint satisfactionHandles complex multi-term reward functions

Reinforcement Teaching

Apr 25, 2022
AL
Alex Lewandowski
🏛️ University of Alberta | Huawei Technologies Canada Co., Ltd. | Google Brain

Existing meta-learning methods suffer from limited generalizability, often being confined to specific algorithms or requiring differentiability assumptions. This paper proposes a general reinforcement learning–driven meta-learning framework that trains a teacher policy to dynamically guide arbitrary student algorithms—without imposing structural or differentiability constraints on the student. Key contributions include: (i) the first unified pedagogical paradigm for meta-learning; (ii) a parameter-behavior encoder that implicitly infers the student’s internal parameter state from its input-output behavior; and (iii) a reward function grounded in learning progress. Experiments across supervised and reinforcement learning tasks demonstrate that our framework significantly outperforms baselines relying on heuristic rewards and handcrafted state representations, validating its broad generalizability and empirical effectiveness.

AdaptabilityMachine Learning EfficiencyMeta-Learning

Comprehensive Overview of Reward Engineering and Shaping in Advancing Reinforcement Learning Applications

Jul 22, 2024
SI
Sinan Ibrahim
🏛️ Skolkovo Institute of Science and Technology | Innopolis University

Addressing practical challenges in reinforcement learning—such as sparse and delayed rewards and training instability—this paper presents a systematic survey of reward engineering and reward shaping. We propose the first fine-grained taxonomy of reward design techniques, explicitly exposing their implicit assumptions and failure boundaries. Furthermore, we introduce an evaluation framework for reward shaping that jointly balances interpretability and empirical effectiveness. Our analysis integrates theoretical foundations of RL, deep RL practice, formal modeling of reward functions, and cross-domain applications—including robotics and autonomous driving. This work fills a critical gap by providing the first comprehensive, methodology-driven survey of reward design. It establishes a unified tripartite research framework comprising methodology, taxonomic classification, and application boundaries. The resulting synthesis delivers a reproducible, transferable engineering guide for algorithm designers, significantly enhancing the robustness and real-world deployability of RL systems. (149 words)

Complex Real-World ProblemsReinforcement LearningReward Mechanism

To address the human dependency and poor generalizability in reinforcement learning (RL) task curriculum design, this paper proposes the first fully automated RL curriculum generation framework powered by large language models (LLMs). Our method operates via a three-stage pipeline: (1) natural-language-based subtask decomposition; (2) end-to-end compilation of executable reward functions and goal-conditioned code; and (3) curriculum refinement driven by policy rollout trajectory evaluation. Crucially, we systematically integrate LLMs’ planning and code-generation capabilities into RL curriculum design—eliminating manual intervention while enabling cross-domain curriculum synthesis across manipulation, navigation, and humanoid locomotion. Experiments in diverse robotic simulation environments demonstrate substantial improvements in complex skill acquisition efficiency. Furthermore, policies generated by our framework successfully transfer to a real-world humanoid robot, validating effective sim-to-real deployment. This work establishes a scalable, domain-agnostic paradigm for automated curriculum learning in RL.

Automates curriculum design for complex robot skills using LLMsEnhances learning efficiency across diverse robotics environmentsReduces need for human intervention in task decomposition

Adaptive teachers for amortized samplers

Oct 02, 2024
MK
Minsu Kim
🏛️ KAIST | Recursion | Université de Montréal | POSTECH | University of Edinburgh

To address the low sample efficiency and inadequate multimodal coverage in approximate inference for hard-to-sample unnormalized density distributions, this paper reformulates sampling as a sequential decision-making process and introduces a novel adaptive teacher-guided framework: dynamically identifying high-loss regions where the student sampler underperforms and actively constructing a progressive training curriculum. The method integrates reinforcement learning–based normalizing flows, off-policy training, auxiliary behavioral modeling, and amortized inference. Evaluated across synthetic exploration environments, two diffusion-based sampling tasks, and four biochemical discovery benchmarks, it achieves substantial improvements—averaging +37% in sample efficiency and +52% in mode coverage—while notably enhancing discovery of low-probability, high-reward modes.

Enhancing mode coverage through adaptive training distributionImproving exploration efficiency in reinforcement learning methodsTraining parametric models for intractable distribution approximation

Latest Papers

What's happening recently
View more

Current large language model training typically introduces reinforcement learning (RL) only after pretraining and supervised fine-tuning (SFT), which constrains its full potential. This work proposes a novel paradigm that integrates RL and SFT directly during multiple stages of pretraining, exploring their concurrent optimization. By intervening at pretraining checkpoints, designing a target objective averaging mechanism, and carefully controlling data composition, the study demonstrates that introducing RL early can match or even surpass the performance of the conventional SFT→RL pipeline—particularly on challenging tasks—without compromising general capabilities. Moreover, strategic design of data composition proves more effective for performance gains than merely scaling up model size. These findings offer a new, efficient, and flexible pathway for aligning language models with desired behaviors.

Large Language ModelsPolicy OptimizationPre-training

This work addresses the limitations of conventional training pipelines, which struggle to dynamically mitigate issues such as overfitting, loss imbalance, and unsafe exploration due to reliance on fixed policies or single-axis schedulers. The authors propose a large language model–based bounded supervisory controller that leverages structured telemetry snapshots to monitor training dynamics in real time and generates verifiable multi-parameter adjustment commands within a constrained action space. This enables closed-loop regulation of learning rate, regularization strength, loss weighting, and exploration strategy. Notably, it introduces pattern-constrained large language models into training supervision for the first time, supporting asynchronous, auditable multi-axis interventions applicable to both supervised and reinforcement learning. Experiments demonstrate a 60% loss reduction on TinyStories with effective overfitting correction, marked alleviation of overly conservative or unsafe exploration in robotic manipulation tasks, and generation of traceable intervention logs.

adaptive trainingexploration collapseloss imbalance

This work proposes the LLM-as-Environment-Engineer framework, which leverages a large language model (Qwen3-4B) as an “environment engineer” to dynamically optimize reinforcement learning training environments. Addressing the common reliance on manually tuned settings and the absence of automated, performance-driven environment adaptation, the framework analyzes policy failure trajectories, behavioral summaries, and environmental statistics to automatically reconstruct a multi-dimensional, configurable MAPF-FrozenLake environment. Experimental results demonstrate that reinforcement learning agents fine-tuned within this closed-loop, LLM-driven optimization process not only develop enhanced self-diagnostic capabilities to guide environmental refinements but also achieve superior overall performance—outperforming both larger closed-source models such as GPT and Gemini and fixed-environment baselines—thereby validating the efficacy and novelty of dynamic environment design in reinforcement learning.

Environment DesignLarge Language ModelsMulti-Agent Reasoning

This work addresses the challenge of sparse trajectory information in long-horizon reinforcement learning for large language model (LLM) agents, where weak policies often fail repeatedly, hindering effective policy optimization. The authors propose a policy-centric training paradigm that dynamically models skills as evolving scaffolds aligned with policy development. Specifically, during inference, the framework adaptively provides guidance through evidence card generation, task-specific evaluation, and context-aware adjustment mechanisms, gradually reducing reliance on external support as the agent’s capabilities improve—thus balancing guided assistance with growing autonomy. Integrated with standard RLVR optimization, this approach outperforms strong baselines by up to 18.6% on ALFWorld and WebShop benchmarks, achieves competitive performance across seven retrieval-augmented question-answering tasks, and reduces prompt usage by 32.1%.

LLM agentslong-horizon reinforcement learningpolicy optimization

Hot Scholars

ZX

Zhiheng Xi

Fudan University
LLM ReasoningLLM-based Agents
SL

Sicong Leng

Nanyang Technological University
Multi-modal Learning
HD

Haodong Duan

Shanghai AI Lab | CUHK | PKU
Computer VisionVideo UnderstandingMultimodal LearningGenerative AI
SF

Sicheng Feng

Nankai University
Efficient Deep LearningReasoning Models
LZ

Lanyun Zhu

NTU, CityUHK, SUTD, BUAA
Multimodal LearningComputer VisionResource-efficient LearningLarge Vision-Language Model