Disentangling Curriculum Learning in NLP: Towards a Unifying Taxonomy

📅 2026-07-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of systematic understanding of difficulty metrics and scheduling strategies in curriculum learning for natural language processing, which has hindered method comparison and reproducibility. The authors propose a fine-grained taxonomy that decouples curriculum learning into two orthogonal components: difficulty assessment and training scheduling. They formally define the scheduler for the first time, explicitly distinguishing between sources of difficulty and task dependencies. Through conceptual decoupling, formal modeling, and systematic literature analysis, they introduce retention mechanisms and monotonicity properties to characterize scheduling behavior, thereby uncovering the root causes of conceptual conflation in prior work. This framework enables unified design, rigorous analysis, and fair comparison of curriculum strategies, paving the way toward a reproducible and comparable evaluation paradigm for curriculum learning.
📝 Abstract
Despite more than a decade of curriculum learning (CL) research in NLP, the field lacks a principled account of which difficulty function or scheduler to use for a given problem. To understand what has hindered progress towards this account, we propose a fine-grained taxonomy separating difficulty evaluation from training scheduling to enable systematic analysis of CL strategies. For difficulty evaluation, we distinguish attribution source and task dependence, revealing difficulty as a perspectival concept encoding different assumptions about what makes an instance hard to learn. For scheduling, we provide the first formalisation of CL schedulers in terms of expected training contribution, enabling comparison across implementations by introducing retention regimes and monotonicity properties. Applied in a dedicated analysis of CL works in NLP, our taxonomy reveals a systematic incomparability problem: prior works conflate distinct notions of difficulty and scheduling, often pursuing different objectives under the same CL label -- hindering comparison and the accumulation of a coherent evidence base. Beyond diagnosis, the taxonomy supports the design, analysis, and comparison of CL strategies, and motivates evaluation practices that disentangle the sources of observed improvement.
Problem

Research questions and friction points this paper is trying to address.

Curriculum Learning
Difficulty Evaluation
Training Scheduling
Taxonomy
NLP
Innovation

Methods, ideas, or system contributions that make the work stand out.

curriculum learning
taxonomy
difficulty evaluation
training scheduling
NLP
V
Vanessa Toborek
University of Bonn, Germany; Lamarr Institute, Germany; Fraunhofer IAIS, Germany
F
Florian Seiffarth
University of Bonn, Germany; Lamarr Institute, Germany; Fraunhofer IAIS, Germany
S
Sebastian Müller
University of Bonn, Germany; Lamarr Institute, Germany; Fraunhofer IAIS, Germany
T
Tamás Horváth
University of Bonn, Germany; Lamarr Institute, Germany; Fraunhofer IAIS, Germany