🤖 AI Summary
This study addresses the high computational cost and difficulty of approximating the complete Pareto front in multi-objective optimization for deep neural networks by proposing the MOSEL framework. MOSEL reformulates the problem as a Stackelberg game, leveraging network modularity to decouple representation learning from preference alignment. Through bilevel optimization and a single forward-backward pass, it achieves posteriori multi-objective optimization. Experiments demonstrate that MOSEL enables scalable Pareto front learning at the computational cost of single-objective training. Specifically, it effectively reveals diverse optimal solutions in strongly conflicting scenarios such as fairness-aware learning, while closely approximating the utopia point and significantly improving generalization performance in multi-task learning settings.
📝 Abstract
We propose MOSEL (Multi-Objective Stackelberg Efficient Learning), a framework for a posteriori multi-objective optimization (MOO) in deep neural networks that recovers a full front of Pareto stationary solutions at the computational cost of standard single-objective training. MOSEL reformulates the problem as a bilevel optimization problem that leverages network modularity to decouple representation learning from objective-preference alignment. Casting the bilevel optimization problem as a Stackelberg game enables solving the original a posteriori MOO problem in a single forward-backward pass. As a result, MOSEL matches the time and memory efficiency of standard single-objective training while enabling scalable Pareto stationary front learning. Empirically, MOSEL uncovers diverse and optimal Pareto frontiers in strongly conflicting settings (e.g., fairness-accuracy). Remarkably, even in weakly conflicting regimes such as multi-task learning, it consistently converges to solutions closer to the utopia point, outperforming both standard single-objective training and specialized multi-task learning methods. These results highlight the broader potential of a posteriori MOO learning as a pathway to efficiently learn more diverse and robust representations, ultimately improving generalization.