Score
Designs and implements algorithms, schedules, and controllers that set or adapt per-objective or per-class loss scales during model training (e.g., static schedules, dynamic reweighting, inverse-share softmax) to balance multiple objectives, mitigate class or task imbalance, and prevent any single loss from dominating optimization. Analyzes training trajectories and optimization stability to tune loss-reweighting strategies and ensure robust multi-objective convergence and fair emphasis across tasks or classes.
This work addresses the challenges of jointly optimizing numerous loss terms and managing high memory and computational overhead in multi-objective deep learning. Methodologically, we propose a hierarchical output-feedback control framework that eliminates explicit Lagrange multipliers by introducing time-varying multipliers, dynamically reshaping the loss landscape at the epoch level. We further introduce a novel hypervolume-based likelihood probabilistic graphical model that jointly captures the co-evolution of model parameters and multipliers, decomposing multi-objective optimization into a sequence of Pareto-adaptive constrained hierarchical optimal control subproblems. Evaluated on the PACS domain generalization benchmark—featuring a six-loss-term variational autoencoder—we demonstrate that our approach significantly outperforms existing multiplier-scheduling methods in both accuracy and robustness, while substantially reducing memory footprint and computational cost. Moreover, the framework supports modular extension for diverse multi-objective architectures.
This work addresses the challenge of simultaneously achieving calibration, low regret, and multi-accuracy in online learning under arbitrarily time-varying data distributions—a setting where existing methods struggle to balance these competing objectives. The authors propose a novel local adaptive mechanism that integrates a multi-objective optimization framework with adaptive online learning algorithms. Without requiring explicit definitions of local targets, their approach dynamically optimizes performance over contiguous subintervals, thereby circumventing the limitations of traditional global worst-case analyses. Empirical evaluations on energy forecasting and algorithmic fairness benchmarks demonstrate that the method significantly outperforms current state-of-the-art techniques, delivering unbiased predictions for subpopulations while maintaining robust multi-objective performance under distributional shifts.
Learning rate scheduling in large language model training lacks rigorous theoretical foundations, leading to heuristic designs and suboptimal convergence. Method: This paper establishes, for the first time, a quantitative alignment between practical schedulers (e.g., linear decay) and tight non-smooth convex optimization lower bounds—eliminating spurious logarithmic factors in prior analyses and enabling principled cross-scheduler optimal learning rate transfer. We integrate convex optimization theory, scheduler modeling, and empirical validation, conducting systematic evaluations on 124M- and 210M-parameter Llama models. Results: Theory-guided scheduler design yields faster convergence and improved stability, empirically validating optimization theory’s practical relevance for large-model training. Core contribution: bridging the gap between theoretical performance bounds and engineering schedulers by providing a transferable, interpretable, and theoretically grounded framework for learning rate tuning.
Existing multi-objective optimization research predominantly focuses on conflicting objectives and Pareto fronts, overlooking the prevalent “aligned objectives” scenario in machine learning—where objectives are non-conflicting and mutually reinforcing. This work formally defines the aligned multi-objective optimization problem and breaks from the traditional Pareto paradigm by proposing the first gradient-based optimization framework tailored to this setting. Methodologically, it introduces a dynamic weight allocation and gradient normalization fusion algorithm grounded in gradient direction alignment analysis, accompanied by theoretical convergence guarantees. Compared to naive strategies such as weighted sum, the approach achieves significantly improved optimization efficiency and stability. Empirical evaluation on multi-task learning and large language model training demonstrates synchronous performance gains across all objectives, faster convergence, enhanced robustness, and scalability to large-scale, highly correlated objective sets.
To address the inefficiency and poor generalizability of manual hyperparameter tuning—particularly for learning rates—this paper proposes a dynamic online meta-optimization framework that formulates learning rate adaptation as a discounted cumulative regret minimization problem over time. The method employs a gradient-based meta-update mechanism, enabling plug-and-play integration with any first-order optimizer (e.g., SGD, Adam) to achieve decoupled, real-time, adaptive step-size optimization. Key contributions include: (i) the first formalization of meta-optimization as discounted regret minimization; and (ii) a low-complexity variant that preserves theoretical rigor while ensuring computational efficiency and strong generalization. Experiments across diverse tasks demonstrate faster convergence, enhanced robustness to initialization and task heterogeneity, competitive performance against hand-tuned optimal schedulers, and significantly lower computational overhead compared to conventional hyperparameter search methods.
This study investigates whether the substantial training cost of deep reinforcement learning (DRL) in carbon-aware flow shop scheduling can be justified by long-term performance gains through policy generalization. To this end, the authors propose a framework that integrates DRL with dynamic algorithm configuration, training policies on small, simple instances and transferring them to unseen, complex instances for online parameter adaptation. Experimental results demonstrate that the proposed approach significantly outperforms baseline methods—such as static parameter tuning—on out-of-distribution, complex scenarios. These findings validate the strong generalization capability of DRL policies and confirm that the initial investment in training yields sustained performance benefits in practical applications.
This study investigates how to assign optimal layer-wise learning rates in the early phase of training deep linear neural networks to minimize test loss. By deriving an exact closed-form solution after two steps of gradient descent, the authors characterize the initial training dynamics and propose a gradient-based inter-layer learning rate scaling strategy, constructing an analytically tractable surrogate loss function. Theoretical analysis reveals that employing unequal learning rates across layers at initialization enhances performance, whereas uniform learning rates become preferable in subsequent steps. This work presents the first precise theoretical characterization of the dynamics during the first two training iterations, and numerical experiments confirm that the proposed strategy significantly reduces test loss compared to baseline approaches.
This work addresses the instability and performance degradation commonly observed during fine-tuning of pre-trained models, which often stems from gradient cancellation leading to optimization collapse. To mitigate this issue, the paper introduces, for the first time in the context of fine-tuning, a dynamic gradient scaling mechanism, proposing the Dynamic Scaled Gradient Descent (DSGD) algorithm. DSGD adaptively attenuates the gradient magnitudes of correctly classified samples, thereby effectively alleviating gradient cancellation. The method substantially enhances fine-tuning stability and robustness, consistently reducing performance variance and achieving higher accuracy than existing approaches across multiple benchmark datasets and large-scale models.