curriculum rl fine-tuning

Designs and implements stage-wise training pipelines that fine-tune models through curriculum learning followed by reinforcement-learning calibration, sequencing easier-to-harder tasks or small labeled subsets into multiple stages. Builds the scheduling, sample-selection and reward/policy optimization components needed to calibrate behavior in a second RL stage and improve task performance while minimizing labeled data and annotations.

curriculumrlfine-tuning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.43
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning

Oct 06, 2025
JH
Jonas Hübotter
🏛️ ETH Zürich | Max Planck Institute for Intelligent Systems

Humans excel at “learning in the job”—dynamically optimizing policies during task execution. This paper introduces Test-Time Curriculum Reinforcement Learning (TTC-RL), a framework enabling models to autonomously construct task-specific curricula during inference, select high-value samples from large-scale unlabeled data, and continuously fine-tune themselves to improve performance on target tasks. Its core innovation extends test-time learning into a goal-directed, online reinforcement training process spanning thousands of steps—fully unsupervised and annotation-free. TTC-RL integrates automatic curriculum selection with sparse-reward-driven policy optimization. Evaluated on mathematical reasoning (AIME25) and competitive programming (CodeElo) benchmarks, it significantly enhances Qwen3-8B: pass@1 improves by 1.8× and 2.1×, respectively, while pass@8 rises from 40% to 62% on AIME25 and from 28% to 43% on CodeElo.

Automatically creates task-specific curricula for reinforcement learningImproves model performance on math and coding benchmarksSelects relevant training data without human curation

Current large language model training typically introduces reinforcement learning (RL) only after pretraining and supervised fine-tuning (SFT), which constrains its full potential. This work proposes a novel paradigm that integrates RL and SFT directly during multiple stages of pretraining, exploring their concurrent optimization. By intervening at pretraining checkpoints, designing a target objective averaging mechanism, and carefully controlling data composition, the study demonstrates that introducing RL early can match or even surpass the performance of the conventional SFT→RL pipeline—particularly on challenging tasks—without compromising general capabilities. Moreover, strategic design of data composition proves more effective for performance gains than merely scaling up model size. These findings offer a new, efficient, and flexible pathway for aligning language models with desired behaviors.

Large Language ModelsPolicy OptimizationPre-training

ToolSample: Dual Dynamic Sampling Methods with Curriculum Learning for RL-based Tool Learning

Sep 18, 2025
ZF
Zihao Feng
🏛️ Harbin Institute of Technology | Tencent

In RL-based LLM tool learning, the pedagogical value of simple samples diminishes over time, and existing dynamic sampling methods struggle to accommodate multi-task architectures and fine-grained reward signals. To address these challenges, we propose a synergistic framework integrating dual dynamic sampling and curriculum learning. Specifically, we jointly model sampling dynamics along two orthogonal dimensions: (i) reward-aware sampling—weighting trajectories dynamically based on per-step reward mean and variance; and (ii) task-aware curriculum progression—sequencing subtasks according to mastery estimates. This work is the first to deeply integrate curriculum learning with multi-dimensional dynamic sampling in tool-use RL, while explicitly coupling fine-grained reward modeling. Evaluated on the BFCLv3 benchmark, our method achieves a +3.29% absolute performance gain over strong baselines, with concurrent improvements in training efficiency and cross-task generalization.

Addresses inefficient RL training from excessive simple samplesImproves training efficiency and performance through adaptive sampling strategiesTargets multi-task structure and fine-grained rewards in tool learning

Curriculum Reinforcement Learning for Complex Reward Functions

Oct 22, 2024
KF
Kilian Freitag
🏛️ Chalmers University of Technology | University of Gothenburg

Training deep reinforcement learning agents under multi-objective conflicting rewards often suffers from instability and difficulty balancing task performance against constraint satisfaction. To address this, we propose a two-stage reward curriculum learning framework: an initial phase optimizes a simplified reward to accelerate convergence, followed by a smooth transition to the full, complex reward. We introduce a novel Actor-Critic fidelity criterion for automatic, dynamic stage switching and design a flexible replay buffer enabling cross-phase sample reuse. Our approach integrates curriculum learning, dynamic reward shaping, and adaptive experience replay. Evaluated on the DeepMind Control Suite—including tasks with explicit constraints—and real-world mobile robot navigation, our method significantly outperforms non-curriculum baselines, achieving a more robust trade-off between task success rate and constraint violation rate.

Automates transition in reward curriculum stagesBalances task completion and constraint satisfactionHandles complex multi-term reward functions

Syllabus: Portable Curricula for Reinforcement Learning Agents

Nov 18, 2024
RS
Ryan Sullivan
🏛️ University of Maryland, College Park | University College London | Jamia Hamdard University

Current reinforcement learning (RL) frameworks lack native, non-intrusive support for curriculum learning (CL), requiring invasive code modifications to implement. To address this, we propose CLib—the first lightweight, general-purpose curriculum learning library. CLib features a unified API and modular architecture comprising: (i) an environment-agnostic curriculum scheduler, (ii) a distributed sampling adapter, and (iii) a cross-framework bridging layer supporting both PyTorch and TensorFlow backends. It integrates seamlessly with five+ mainstream RL libraries—including Ray RLlib and CleanRL—without altering underlying training logic. We demonstrate the first successful application of CL in complex environments NetHack and Neural MMO, and validate CLib across nine benchmark tasks, consistently outperforming state-of-the-art baselines. By eliminating implementation barriers, CLib lowers the entry threshold for CL adoption, promotes standardization, and enhances reproducibility in RL research.

Complex code changes needed for curriculum learning methodsDifficulty in adapting curriculum learning to new environmentsLack of direct support for curriculum learning in major RL libraries

Latest Papers

What's happening recently
View more

This work addresses the challenges of low training efficiency and suboptimal performance commonly faced by reinforcement learning agents in high-dimensional action spaces. It proposes, for the first time, a method that leverages large language models to dynamically generate action-level curricula, constructing multi-stage training trajectories for both Tabular Q-Learning and Deep Q-Network (DQN) agents in the game of Blackjack. By progressively introducing more complex actions, the approach integrates large language models, curriculum learning, and deep reinforcement learning to enhance learning efficacy. Evaluated in an eight-deck Blackjack environment, the method significantly improves agent performance: the DQN agent’s win rate increases from 43.97% to 47.41%, its bust rate decreases from 32.9% to 28.0%, and training converges over 74% faster—requiring less total training time than the evaluation phase of baseline methods.

Complex EnvironmentsCurriculum LearningEfficiency

This work addresses the limitations of large language models in generating high-quality BPMN process models, which are constrained by supervised fine-tuning data and the absence of well-defined multidimensional reward functions. The authors propose a reinforcement learning–based optimization approach that systematically explores a reward function encompassing 38 syntactic, pragmatic, and semantic metrics. They train Llama-3.1-8B and Qwen2.5-14B models across 48 configurations and find that uniformly weighted rewards outperform targeted weighting schemes, with significant interaction effects observed between reward composition and model architecture. Leveraging Group Relative Policy Optimization and an automated evaluation framework, the method substantially improves pragmatic and syntactic quality while preserving semantic fidelity and reducing output variability by over sixfold. All code is publicly released.

LLMmulti-dimensional qualityprocess model generation

This work addresses the challenge in Flow Matching models where fixed time-step sampling strategies, such as midpoint biasing, struggle to balance training efficiency and sample quality. The study reframes time-step sampling as a dynamic curriculum and reveals that the loss landscape exhibits a U-shaped difficulty distribution across time steps. To exploit this insight, the authors propose a two-stage curriculum sampling strategy: initially employing midpoint-biased sampling to accelerate structural learning, followed by a switch to uniform sampling to refine boundary details. Evaluated on CIFAR-10, the method improves the Fréchet Inception Distance (FID) from 3.85 to 3.22 and achieves peak performance within 100,000 training steps—significantly outpacing the 150,000 steps required by uniform sampling.

boundary regimesFlow Matchingsampling distribution

Hot Scholars

YZ

Yefeng Zheng

Professor, Westlake University, Hangzhou, China, IEEE Fellow, AIMBE Fellow
AI in HealthMedical ImagingComputer VisionNatural Language Processing
ZL

Zichen Liu

Sea AI Lab; National University of Singapore
reinforcement learningartificial intelligence
XD

Xiaoliang Dai

Research Scientist, Meta GenAI
Generative AIComputer vision
TH

Tingbo Hou

Google DeepMind
Computer VisionGenerative AI
ZH

Zecheng He

Meta GenAI
Generative AIEfficient ModelAI Security and Privacy