Score
Designs, implements, and evaluates training curricula and scheduling mechanisms that sequence and adapt tasks, data distributions, and objective phases over time to shape model learning dynamics. This work builds and analyzes curriculum strategies such as progressive difficulty (including burial-conditioned and progressive burial curricula), two-stage or multi-stage pretraining and synthetic warm-up then real-data training, recursive and online curricula with revisitation loops, curriculum alignment across tasks, and measures for curriculum scheduling and effectiveness.
This work addresses the lack of systematic understanding of difficulty metrics and scheduling strategies in curriculum learning for natural language processing, which has hindered method comparison and reproducibility. The authors propose a fine-grained taxonomy that decouples curriculum learning into two orthogonal components: difficulty assessment and training scheduling. They formally define the scheduler for the first time, explicitly distinguishing between sources of difficulty and task dependencies. Through conceptual decoupling, formal modeling, and systematic literature analysis, they introduce retention mechanisms and monotonicity properties to characterize scheduling behavior, thereby uncovering the root causes of conceptual conflation in prior work. This framework enables unified design, rigorous analysis, and fair comparison of curriculum strategies, paving the way toward a reproducible and comparable evaluation paradigm for curriculum learning.
This work proposes a reverse curriculum learning framework tailored for structurally complex and conceptually deep tasks such as advanced mathematical problem solving and code generation. The approach introduces a difficulty scoring mechanism based on structural complexity and conceptual depth, and employs a teacher–student architecture to recursively decompose challenging problems. The teacher model generates progressively simplified examples through step-by-step reasoning, thereby constructing an easy-to-hard curriculum that guides the student model in incremental learning. Experimental results on benchmarks including MATH and AIME demonstrate that this method significantly outperforms standard training strategies, effectively enhancing the model’s capacity to solve complex problems.
This study investigates whether curriculum learning in large language model pretraining alters the learning trajectory or merely reshapes the order of data exposure. We implement three linguistically motivated curricula—based on age of acquisition, word frequency, and verb variability—on the Pythia model series, comparing them against random data ordering. Through gradient variance analysis, spectral saturation measurements, and idealized optimization theory, we demonstrate that curriculum learning primarily enhances performance by stabilizing the optimization dynamics within existing learning phases rather than introducing new stages. Empirical results show that curricula significantly reduce gradient noise and output head spectral saturation in smaller models, yielding higher accuracy; however, these benefits diminish with increasing model scale, revealing a scaling-dependent efficacy of curriculum strategies and establishing a theoretical link between difficulty scheduling and optimization stability.
This work addresses the limitations of existing curriculum learning approaches, which rely on static or computationally expensive dynamic difficulty assessments and struggle to generate efficient, learner-specific training sequences. The authors propose a novel problem difficulty evaluation mechanism based on a relative measure of model capability, introducing and formally defining “transitional problems”—critical instances that shift from difficult to easy as the model’s competence improves. Leveraging this insight, they construct an adaptive curriculum that aligns dynamically with the learner’s evolving capacity, yielding a personalized, interpretable, and computationally efficient training trajectory. Experiments on chess and mathematical reasoning tasks demonstrate that the proposed strategy significantly outperforms current methods, effectively facilitating transitions to higher levels of model performance.
This study addresses the challenge in high-dimensional motor skill acquisition where true skill states are unobservable and task performance poorly reflects underlying learning progress, thereby hindering effective practice design. To overcome this limitation, the authors propose an automated curriculum generation framework that integrates a human motor learning model, personalized real-time skill estimation, and stochastic nonlinear model predictive control (SNMPC), enabling model-based dynamic curriculum optimization for the first time. Evaluated in both simulation and a hand exoskeleton experiment involving 36 participants, the proposed approach significantly accelerates skill acquisition—improving learning speed by approximately 23% over random curricula and 17% over performance-heuristic baselines—demonstrating its efficacy and novelty.
The empirical benefits of curriculum learning in post-training inference for large language models (LLMs) lack principled theoretical justification. Method: We propose curriculum strategies based on incrementally increasing reasoning-chain depth or decrementally shortening prompt length, and introduce a state-conditioned autoregressive reasoning tree model. This framework enables curriculum-aware fine-tuning via reinforcement learning under outcome-only reward signals. Contribution: We provide the first theoretical proof that curriculum learning can overcome the exponential sample complexity barrier inherent in tree-structured reasoning—reducing it to polynomial order—and establish polynomial-cost scaling guarantees for test-time inference. Experiments demonstrate substantial improvements in reasoning accuracy, alongside significant reductions in sampling overhead and API query costs.
This work addresses the challenge of dynamically varying sample complexity in multimodal learning, which existing curriculum learning approaches struggle to accommodate as models evolve. It introduces, for the first time, Partial Information Decomposition (PID) theory into multimodal curriculum learning, proposing a progressive curriculum framework that decomposes multimodal interactions into redundant, unique, and synergistic information components. This decomposition enables a dynamic characterization of sample complexity, which in turn drives an adaptive scheduling of training samples aligned with the model’s learning progress. The resulting approach yields an interpretable, training-aware sample selection mechanism that significantly outperforms conventional training strategies and state-of-the-art baselines across multiple multimodal benchmarks, demonstrating both its effectiveness and generalizability.
This work addresses the scalability bottleneck in AI tutoring systems caused by the labor-intensive, manual construction of structured procedural skill models. To overcome this limitation, the authors propose a human-in-the-loop text-to-model generation approach that leverages large language models to automatically transform instructional texts into procedural skill models conforming to the Task-Method-Knowledge (TMK) ontology. The method integrates ontology-constrained prompting with template-driven generation and incorporates expert validation of causal logic and failure conditions. This framework preserves model structural integrity and semantic alignment while substantially reducing expert modeling effort. Evaluated in a graduate-level AI course, the approach produced 23 skill models with 50–70% less expert time investment, and the generated models demonstrated high reproducibility under fixed inputs.
In sparse-reward reinforcement learning, heterogeneous tasks exhibit divergent learning progress, rendering uniform reset curricula ineffective across all contexts. To address this challenge, this work proposes SCOUT—an online, learner-agnostic reset controller that dynamically adjusts the curriculum at the granularity of individual contexts for the first time. Leveraging only binary episode success signals, SCOUT adaptively reduces assistance upon sustained success, restores it after failure, and cautiously attempts harder starting states during stagnation. Requiring no predefined context groupings, reward function modifications, or optimizer changes, SCOUT significantly improves performance across six navigation and manipulation tasks, successfully solving three that fail entirely without auxiliary resets. Notably, under deliberately induced pacing conflicts, SCOUT concurrently learns two distinct task sets, whereas all global scheduling baselines fail.