adaptive loss weighting

Designs and implements algorithms, schedules, and controllers that set or adapt per-objective or per-class loss scales during model training (e.g., static schedules, dynamic reweighting, inverse-share softmax) to balance multiple objectives, mitigate class or task imbalance, and prevent any single loss from dominating optimization. Analyzes training trajectories and optimization stability to tune loss-reweighting strategies and ensure robust multi-objective convergence and fair emphasis across tasks or classes.

adaptivelossweighting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.42
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

M-HOF-Opt: Multi-Objective Hierarchical Output Feedback Optimization via Multiplier Induced Loss Landscape Scheduling

Mar 20, 2024
XS
Xudong Sun
🏛️ Helmholtz Munich | Volkswagen Group | U.S. FDA | KTH Royal Institute of Technology

This work addresses the challenges of jointly optimizing numerous loss terms and managing high memory and computational overhead in multi-objective deep learning. Methodologically, we propose a hierarchical output-feedback control framework that eliminates explicit Lagrange multipliers by introducing time-varying multipliers, dynamically reshaping the loss landscape at the epoch level. We further introduce a novel hypervolume-based likelihood probabilistic graphical model that jointly captures the co-evolution of model parameters and multipliers, decomposing multi-objective optimization into a sequence of Pareto-adaptive constrained hierarchical optimal control subproblems. Evaluated on the PACS domain generalization benchmark—featuring a six-loss-term variational autoencoder—we demonstrate that our approach significantly outperforms existing multiplier-scheduling methods in both accuracy and robustness, while substantially reducing memory footprint and computational cost. Moreover, the framework supports modular extension for diverse multi-objective architectures.

Hierarchically dispatches multi-objective descent into constraint sub-problems.Optimizes multi-objective model parameters using time-varying multipliers.Reduces memory and computational burden in multi-objective deep learning.

This work addresses the challenge of simultaneously achieving calibration, low regret, and multi-accuracy in online learning under arbitrarily time-varying data distributions—a setting where existing methods struggle to balance these competing objectives. The authors propose a novel local adaptive mechanism that integrates a multi-objective optimization framework with adaptive online learning algorithms. Without requiring explicit definitions of local targets, their approach dynamically optimizes performance over contiguous subintervals, thereby circumventing the limitations of traditional global worst-case analyses. Empirical evaluations on energy forecasting and algorithmic fairness benchmarks demonstrate that the method significantly outperforms current state-of-the-art techniques, delivering unbiased predictions for subpopulations while maintaining robust multi-objective performance under distributional shifts.

distribution shiftfairnesslocal adaptivity

Learning rate scheduling in large language model training lacks rigorous theoretical foundations, leading to heuristic designs and suboptimal convergence. Method: This paper establishes, for the first time, a quantitative alignment between practical schedulers (e.g., linear decay) and tight non-smooth convex optimization lower bounds—eliminating spurious logarithmic factors in prior analyses and enabling principled cross-scheduler optimal learning rate transfer. We integrate convex optimization theory, scheduler modeling, and empirical validation, conducting systematic evaluations on 124M- and 210M-parameter Llama models. Results: Theory-guided scheduler design yields faster convergence and improved stability, empirically validating optimization theory’s practical relevance for large-model training. Core contribution: bridging the gap between theoretical performance bounds and engineering schedulers by providing a transferable, interpretable, and theoretically grounded framework for learning rate tuning.

Large Model TrainingLearning Rate AdjustmentTraining Efficiency

Aligned Multi Objective Optimization

Feb 19, 2025
YE
Yonathan Efroni
🏛️ Meta AI | Technion

Existing multi-objective optimization research predominantly focuses on conflicting objectives and Pareto fronts, overlooking the prevalent “aligned objectives” scenario in machine learning—where objectives are non-conflicting and mutually reinforcing. This work formally defines the aligned multi-objective optimization problem and breaks from the traditional Pareto paradigm by proposing the first gradient-based optimization framework tailored to this setting. Methodologically, it introduces a dynamic weight allocation and gradient normalization fusion algorithm grounded in gradient direction alignment analysis, accompanied by theoretical convergence guarantees. Compared to naive strategies such as weighted sum, the approach achieves significantly improved optimization efficiency and stability. Empirical evaluation on multi-task learning and large language model training demonstrates synchronous performance gains across all objectives, faster convergence, enhanced robustness, and scalability to large-scale, highly correlated objective sets.

Address lack of gradient-based methodsEnhance performance across related tasksExplore non-conflicting objectives optimization

MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters

Feb 04, 2024
AS
Arsalan Sharifnassab
🏛️ University of Alberta | Leiden University

To address the inefficiency and poor generalizability of manual hyperparameter tuning—particularly for learning rates—this paper proposes a dynamic online meta-optimization framework that formulates learning rate adaptation as a discounted cumulative regret minimization problem over time. The method employs a gradient-based meta-update mechanism, enabling plug-and-play integration with any first-order optimizer (e.g., SGD, Adam) to achieve decoupled, real-time, adaptive step-size optimization. Key contributions include: (i) the first formalization of meta-optimization as discounted regret minimization; and (ii) a low-complexity variant that preserves theoretical rigor while ensuring computational efficiency and strong generalization. Experiments across diverse tasks demonstrate faster convergence, enhanced robustness to initialization and task heterogeneity, competitive performance against hand-tuned optimal schedulers, and significantly lower computational overhead compared to conventional hyperparameter search methods.

Dynamically adjusting step sizes during model optimizationOptimizing meta-parameters for efficient machine learning trainingReducing regret by considering long-term impact of learning rates

Latest Papers

What's happening recently
View more

This study investigates whether the substantial training cost of deep reinforcement learning (DRL) in carbon-aware flow shop scheduling can be justified by long-term performance gains through policy generalization. To this end, the authors propose a framework that integrates DRL with dynamic algorithm configuration, training policies on small, simple instances and transferring them to unseen, complex instances for online parameter adaptation. Experimental results demonstrate that the proposed approach significantly outperforms baseline methods—such as static parameter tuning—on out-of-distribution, complex scenarios. These findings validate the strong generalization capability of DRL policies and confirm that the initial investment in training yields sustained performance benefits in practical applications.

carbon-aware schedulingcomputational costdeep reinforcement learning

This study investigates how to assign optimal layer-wise learning rates in the early phase of training deep linear neural networks to minimize test loss. By deriving an exact closed-form solution after two steps of gradient descent, the authors characterize the initial training dynamics and propose a gradient-based inter-layer learning rate scaling strategy, constructing an analytically tractable surrogate loss function. Theoretical analysis reveals that employing unequal learning rates across layers at initialization enhances performance, whereas uniform learning rates become preferable in subsequent steps. This work presents the first precise theoretical characterization of the dynamics during the first two training iterations, and numerical experiments confirm that the proposed strategy significantly reduces test loss compared to baseline approaches.

early training dynamicslayer-wise optimizationlearning rate

This work addresses the instability and performance degradation commonly observed during fine-tuning of pre-trained models, which often stems from gradient cancellation leading to optimization collapse. To mitigate this issue, the paper introduces, for the first time in the context of fine-tuning, a dynamic gradient scaling mechanism, proposing the Dynamic Scaled Gradient Descent (DSGD) algorithm. DSGD adaptively attenuates the gradient magnitudes of correctly classified samples, thereby effectively alleviating gradient cancellation. The method substantially enhances fine-tuning stability and robustness, consistently reducing performance variance and achieving higher accuracy than existing approaches across multiple benchmark datasets and large-scale models.

class imbalancefine-tuninggradient collapse

Hot Scholars

SG

Song Guo

Chair Professor of CSE, HKUST
Large Language ModelEdge AIMachine Learning Systems
QZ

Qichao Zhang

中国科学院自动化研究所
人工智能 强化学习 博弈论 自适应动态规划
ST

Songjun Tu

Institute of Automation, Chinese Academy of Sciences; Pengcheng Laboratory
Large Language ModelsReinforecement Learning
DZ

Dongbin Zhao

Institute of Automation, Chinese Academy of Sciences
Deep Reinforcement LearningAdaptive Dynamic ProgrammingGame AISmart driving