🤖 AI Summary
This work addresses the common oversight in multi-objective deep learning of ignoring the hierarchical importance among objectives, which often compromises the prioritization of primary goals while satisfying secondary constraints. To resolve this, the paper proposes Priority-Constrained Descent (PCD), the first method to explicitly model objective hierarchies by ensuring continuous descent of the primary objective while minimally perturbing the optimization trajectory to meet progress requirements of secondary objectives. A single, interpretable parameter τ governs the trade-off strength, and the approach yields scale-invariant closed-form solutions for two- to three-objective problems. Experiments on model compression and sparsification demonstrate that PCD achieves Pareto superiority, outperforming existing methods on both primary and secondary objectives, with τ effectively modulating optimization behavior.
📝 Abstract
Deep learning problems rarely involve objectives that are equal in importance. A primary objective defines the goal, whilst secondary objectives, such as sparsity, compression, or robustness constrain the solution. While existing multi-objective methods have proven effective in practice, they have a clear symmetry problem and neglect the inherent objective hierarchy built into these objective spaces. We introduce Priority-Constrained Descent (PCD), a gradient-based optimization framework designed to explicitly exploit hierarchical objective structures. PCD preserves the direction of primary descent whilst allowing for the minimal distortion necessary to guarantee progress on secondary objectives, controlled by a single $τ\in [0, 1]$ that dictates the strength of the distortion. The resulting formulation is invariant to objective scaling and admits exact closed-form solutions for problems with two and three objectives. We evaluate PCD within structured network compression settings, unstructured sparsity and low-rankness, and across a variety of synthetic experiments, showing Pareto dominance and better per-objective performance with secondary progress guarantees over existing methods, further exhibiting the interpretable trade-off that $τ$ provides.