Score
Designs and implements finetuning training schedules and dataset curricula for machine learning models, including learning-rate and batch-size schedules, epoch allocation, task/task-domain ordering, and strategies for data selection, weighting, mixing, filtering, and progression. Builds the dataset pipelines and curriculum policies and analyzes their effects on model convergence, generalization, robustness, and downstream behavior.
This work addresses the lack of systematic understanding of difficulty metrics and scheduling strategies in curriculum learning for natural language processing, which has hindered method comparison and reproducibility. The authors propose a fine-grained taxonomy that decouples curriculum learning into two orthogonal components: difficulty assessment and training scheduling. They formally define the scheduler for the first time, explicitly distinguishing between sources of difficulty and task dependencies. Through conceptual decoupling, formal modeling, and systematic literature analysis, they introduce retention mechanisms and monotonicity properties to characterize scheduling behavior, thereby uncovering the root causes of conceptual conflation in prior work. This framework enables unified design, rigorous analysis, and fair comparison of curriculum strategies, paving the way toward a reproducible and comparable evaluation paradigm for curriculum learning.
Existing data evaluation and selection methods for instruction tuning lack systematicity, and commonly used metrics exhibit misalignment with downstream task performance. Method: We propose the first three-dimensional evaluation taxonomy for instruction tuning of large language models—encompassing quality, diversity, and importance—establishing a unified classification framework and enabling cross-method comparative analysis. Through a systematic review of 120+ evaluation approaches, we identify seven recurrent limitations and expose structural bottlenecks between metric design and task adaptability. Empirical comparisons further reveal deficiencies in generalizability, reliability, and interpretability of current methods. Contribution/Results: We articulate five key open challenges and outline corresponding research directions. All implementations are open-sourced on GitHub, providing both theoretical foundations and practical guidelines for efficient, reliable, and interpretable instruction data engineering.
Existing curriculum learning methods rely on static difficulty metrics, failing to accommodate the dynamic capability evolution of large language models (LLMs) during instruction tuning—resulting in rigid, suboptimal learning trajectories. To address this, we propose CAMPUS, the first capability-aware dynamic curriculum learning framework for instruction tuning. CAMPUS continuously monitors the model’s multi-dimensional capabilities, integrates multi-perspective difficulty estimation, and introduces dynamic sub-curriculum selection alongside adaptive difficulty scheduling—thereby enabling personalized, evolution-aware instruction fine-tuning. Extensive experiments demonstrate that CAMPUS consistently outperforms state-of-the-art curriculum learning baselines on major benchmarks—including AlpacaEval and MT-Bench—with average improvements of +2.1–3.8 points—validating its effectiveness and generalizability across diverse LLMs and instruction datasets.
Traditional AI education suffers from a pedagogical disconnect between classical machine learning (ML) and large language models (LLMs), hindering students’ ability to grasp the historical and conceptual evolution of AI techniques. Method: This paper proposes a two-stage, progressive curriculum: Stage I establishes foundational ML concepts—including supervised learning and model evaluation—while Stage II advances to LLM-specific competencies such as prompt engineering, fine-tuning, and application development. The design integrates both paradigms via comparative technical analysis, historical contextualization, and cross-paradigm project-based learning. Contribution/Results: Empirical evaluation demonstrates that the approach significantly enhances students’ holistic understanding of the AI technology ecosystem, increasing conceptual linkage by 32%. Moreover, graduates exhibit stronger practical competencies aligned with industry demand for “ML+LLM” hybrid professionals. The framework provides a reusable, evidence-based pedagogical model for AI curriculum reform.
This work proposes a reverse curriculum learning framework tailored for structurally complex and conceptually deep tasks such as advanced mathematical problem solving and code generation. The approach introduces a difficulty scoring mechanism based on structural complexity and conceptual depth, and employs a teacher–student architecture to recursively decompose challenging problems. The teacher model generates progressively simplified examples through step-by-step reasoning, thereby constructing an easy-to-hard curriculum that guides the student model in incremental learning. Experimental results on benchmarks including MATH and AIME demonstrate that this method significantly outperforms standard training strategies, effectively enhancing the model’s capacity to solve complex problems.
Learning rate scheduling strategies significantly influence neural network performance, yet their selection typically relies on extensive manual trial and error. This work presents the first large-scale, systematic quantification of scheduler effectiveness across heterogeneous architectures. Leveraging the LEMUR dataset, we evaluate 25 configurations from nine PyTorch learning rate schedulers on 30 diverse convolutional and Transformer architectures, training a total of 3,938 model variants through automated source-code instrumentation. Our results reveal strong dependencies between scheduler choice and model architecture, with CosineAnnealingWarmRestarts and CyclicLR substantially outperforming conventional decay strategies. The best-performing configuration achieves a Top-1 accuracy of 86.45%, and 237 variants exceed 80% accuracy. All results are publicly released to support community research.
This work addresses the prevailing reliance on heuristic-based learning rate scheduling strategies, which lack a systematic understanding of optimal schedule shapes. The authors propose a method that decouples the base learning rate from the schedule shape, enabling automatic discovery of near-optimal learning rate schedules within a parameterized family tailored to specific tasks. Experiments across image classification, language modeling, and linear regression reveal that warmup followed by decay constitutes a robust characteristic of high-performing schedules, that commonly used schedule families are often suboptimal, and that weight decay significantly influences the optimal schedule shape. This study provides the first systematic characterization of universal properties underlying near-optimal learning rate schedules and establishes a clear connection between schedule morphology and optimization hyperparameters.
This work addresses the challenge in Flow Matching models where fixed time-step sampling strategies, such as midpoint biasing, struggle to balance training efficiency and sample quality. The study reframes time-step sampling as a dynamic curriculum and reveals that the loss landscape exhibits a U-shaped difficulty distribution across time steps. To exploit this insight, the authors propose a two-stage curriculum sampling strategy: initially employing midpoint-biased sampling to accelerate structural learning, followed by a switch to uniform sampling to refine boundary details. Evaluated on CIFAR-10, the method improves the Fréchet Inception Distance (FID) from 3.85 to 3.22 and achieves peak performance within 100,000 training steps—significantly outpacing the 150,000 steps required by uniform sampling.
This study addresses the widespread absence of systematic instruction on building, testing, deploying, and maintaining AI/ML systems in current undergraduate software engineering (SE) curricula. It presents the first comprehensive delineation of core AI/ML topics essential for SE practice, integrating curriculum mapping analysis, instructor needs surveys, and structured modeling to identify critical content gaps in existing programs. Grounded in empirical evidence, the work proposes actionable pathways for embedding high-priority AI/ML themes into established SE courses. The resulting framework offers a practical, implementable guide to enhance SE education’s capacity to support the development of intelligent software systems.
This work proposes a principled, trial-and-error-free method for determining optimal learning rate schedules in deep learning. By leveraging a solvable power-law random feature model and optimal control theory, the authors analytically derive a stage-wise optimal learning rate schedule: it exhibits polynomial decay during an “easy” learning phase and transitions to a warmup-stable-decay–like profile in a subsequent “hard” phase. This approach uncovers an intrinsic connection between the optimal schedule and the underlying task structure and naturally extends to the joint optimization of momentum and batch size. Experimental results demonstrate that the derived schedule significantly outperforms constant or power-law baselines, achieving faster convergence and accurately predicting the empirically observed compute-optimal scaling laws.