loss weighting

Designing and applying schemes to weight, reweight, or balance loss terms during training so multiple objectives (global vs local, multi-scale) are aligned, preventing overfitting and encoding desired properties with limited labels.

lossweighting

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Aligned Multi Objective Optimization

Feb 19, 2025
YE
Yonathan Efroni
🏛️ Meta AI | Technion

Existing multi-objective optimization research predominantly focuses on conflicting objectives and Pareto fronts, overlooking the prevalent “aligned objectives” scenario in machine learning—where objectives are non-conflicting and mutually reinforcing. This work formally defines the aligned multi-objective optimization problem and breaks from the traditional Pareto paradigm by proposing the first gradient-based optimization framework tailored to this setting. Methodologically, it introduces a dynamic weight allocation and gradient normalization fusion algorithm grounded in gradient direction alignment analysis, accompanied by theoretical convergence guarantees. Compared to naive strategies such as weighted sum, the approach achieves significantly improved optimization efficiency and stability. Empirical evaluation on multi-task learning and large language model training demonstrates synchronous performance gains across all objectives, faster convergence, enhanced robustness, and scalability to large-scale, highly correlated objective sets.

Address lack of gradient-based methodsEnhance performance across related tasksExplore non-conflicting objectives optimization

This work addresses the challenge of simultaneously achieving calibration, low regret, and multi-accuracy in online learning under arbitrarily time-varying data distributions—a setting where existing methods struggle to balance these competing objectives. The authors propose a novel local adaptive mechanism that integrates a multi-objective optimization framework with adaptive online learning algorithms. Without requiring explicit definitions of local targets, their approach dynamically optimizes performance over contiguous subintervals, thereby circumventing the limitations of traditional global worst-case analyses. Empirical evaluations on energy forecasting and algorithmic fairness benchmarks demonstrate that the method significantly outperforms current state-of-the-art techniques, delivering unbiased predictions for subpopulations while maintaining robust multi-objective performance under distributional shifts.

distribution shiftfairnesslocal adaptivity

Rethinking Multi-Objective Learning through Goal-Conditioned Supervised Learning

Dec 12, 2024
SL
Shijun Li
🏛️ The University of Texas at Austin | Intuit

To address task conflict, poor generalization, and high computational complexity in multi-objective learning, this paper proposes a general framework based on Goal-Conditioned Supervised Learning (GCSL). The method decouples scalar objectives into interpretable multidimensional goal vectors and enables end-to-end joint optimization of multiple objectives directly from offline sequential data. It formally characterizes objective feasibility and introduces a novel goal generation mechanism that requires no specialized model architecture, explicit constraints, or complex optimization procedures. Experiments on real-world recommendation datasets demonstrate that the approach significantly improves multi-objective trade-off performance while maintaining high efficiency, scalability, and strong generalization across diverse tasks and domains.

Multi-task LearningRecommendation SystemsTask Conflict Resolution

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards

Oct 01, 2025
YS
Yiran Shen
🏛️ UC San Diego | Databricks

This paper addresses the multi-objective alignment challenge for large language models—simultaneously optimizing verifiable rewards (e.g., mathematical correctness), unverifiable subjective preferences (e.g., human values), and complex interactive settings (e.g., multi-turn AI tutoring). We propose MAH-DPO, a unified framework integrating multi-action-head direct preference optimization (DPO), vectorized reward modeling, and process reward model (PRM) training to enable joint cross-objective optimization. Crucially, MAH-DPO supports fine-grained, user-controllable trade-offs among objectives during inference. Experiments demonstrate significant improvements in mathematical reasoning, value alignment, and multi-turn dialogue tasks, with reduced inter-objective compromise, enhanced alignment flexibility, and improved controllability compared to prior methods.

Aligning models across verifiable and non-verifiable reward domainsProviding fine-grained user control during model inferenceResolving conflicts between competing objectives in multi-objective training

M-HOF-Opt: Multi-Objective Hierarchical Output Feedback Optimization via Multiplier Induced Loss Landscape Scheduling

Mar 20, 2024
XS
Xudong Sun
🏛️ Helmholtz Munich | Volkswagen Group | U.S. FDA | KTH Royal Institute of Technology

This work addresses the challenges of jointly optimizing numerous loss terms and managing high memory and computational overhead in multi-objective deep learning. Methodologically, we propose a hierarchical output-feedback control framework that eliminates explicit Lagrange multipliers by introducing time-varying multipliers, dynamically reshaping the loss landscape at the epoch level. We further introduce a novel hypervolume-based likelihood probabilistic graphical model that jointly captures the co-evolution of model parameters and multipliers, decomposing multi-objective optimization into a sequence of Pareto-adaptive constrained hierarchical optimal control subproblems. Evaluated on the PACS domain generalization benchmark—featuring a six-loss-term variational autoencoder—we demonstrate that our approach significantly outperforms existing multiplier-scheduling methods in both accuracy and robustness, while substantially reducing memory footprint and computational cost. Moreover, the framework supports modular extension for diverse multi-objective architectures.

Hierarchically dispatches multi-objective descent into constraint sub-problems.Optimizes multi-objective model parameters using time-varying multipliers.Reduces memory and computational burden in multi-objective deep learning.

Latest Papers

What's happening recently
View more

This work addresses the challenge of modality imbalance in multimodal learning, where disparities in convergence rates often cause faster-converging modalities to dominate optimization while slower ones remain under-trained. To mitigate this issue, the paper introduces Balanced Multimodal Learning via Label-space Reshaping (BMLR), a novel strategy that aligns inter-modal learning difficulties by deliberately adjusting the mapping complexity from each modality to the label space. By reshaping the label space, BMLR facilitates more effective cross-modal interaction and enhances inter-class discriminability without altering the underlying model architecture. Theoretical analysis and extensive experiments demonstrate that BMLR consistently improves performance across diverse mainstream multimodal architectures, highlighting its strong compatibility and effectiveness as a plug-and-play solution for balanced multimodal representation learning.

feature-to-label mappinglabel spacelearning pace discrepancy

This work addresses the pervasive issue of cross-objective interference in large language models during multi-objective alignment, where improving performance on one objective often degrades others. The study formally characterizes this interference and introduces an analytical framework based on the covariance between reward signals and scalarized scores. It reveals a local covariance law and establishes global convergence conditions under non-convex optimization, incorporating Polyak–Łojasiewicz assumptions and clipped surrogate objectives. Building on these insights, the authors propose CTWA, a plug-and-play method that preserves positive covariance to mitigate interference. Extensive experiments across multiple scalarization algorithms demonstrate the ubiquity of interference and show that CTWA consistently enhances overall multi-objective alignment performance.

cross-objective interferencelarge language modelsmulti-objective alignment

This work addresses the growing risks of misuse and loss of control associated with the broad applicability of foundation models, which existing alignment methods struggle to mitigate through hard behavioral constraints. It establishes capability control as a core objective distinct from alignment and introduces a defense-in-depth framework spanning data, learning, and system layers to enforce multi-granular behavioral constraints throughout the model lifecycle. By integrating techniques such as data distribution shaping, representational intervention, and runtime input/output/action-level safeguards, the paper systematically constructs pathways for capability control. It further identifies critical challenges—including the dual-use nature of knowledge and combinatorial generalization—offering a new paradigm for developing safe and controllable AI systems.

adversarial elicitationcapability controlfoundation models

This work addresses the challenge of gradient conflict in multi-task learning, which often degrades performance on certain tasks. While existing dynamic weighting methods such as MGDA aim to mitigate this issue, they suffer from high computational overhead and poor scalability. To overcome these limitations, this paper formulates gradient balancing as a bilevel optimization problem for the first time and introduces a zeroth-order optimization approach to efficiently decouple model training from weight adjustment. The proposed method substantially reduces computational cost while maintaining or even improving multi-task performance across both public benchmarks and industrial-scale datasets, achieving a favorable trade-off between efficiency and effectiveness.

bi-level optimizationcomputational efficiencygradient balancing

This study addresses the vulnerability of large language models (LLMs) to malicious fine-tuning that can induce misalignment and compromise safety. The authors systematically evaluate four supervised fine-tuning (SFT) and two preference-based fine-tuning (PFT) methods in both inducing misalignment and restoring alignment, conducting experiments across four widely used safety-aligned LLMs. They find, for the first time, that ORPO is most susceptible to causing misalignment, whereas DPO demonstrates superior efficacy in realignment. The work further uncovers key asymmetries between attack and defense dynamics, persistent residual effects from multi-round adversarial interactions, and model-specific resistance patterns. These findings underscore the necessity of tailoring alignment strategies to individual model characteristics and strengthening pre-deployment safeguards—particularly for open-source models—to mitigate alignment erosion risks.

adversarial fine-tuninglarge language modelsmisalignment

Hot Scholars

FL

Fan Li

Department of Statistical Science, Duke University
statisticscausal inferencecomparative effectiveness researchmissing data
YQ

Yumou Qiu

Iowa State University
Statistics
HG

Hanxue Gu

Duke University
Medical imagingDeep learningMachine learning
MB

Michelle Blom

School of Computing and Information Systems, The University of Melbourne
OptimisationArtificial Intelligence