๐ค AI Summary
To address the failure of gradient descent in multi-objective optimization caused by objective conflicts, this paper proposes the Jacobian Descent (JD) algorithm, which iteratively updates parameters directly in the Jacobian matrix space of the vector-valued objective function. Key contributions include: (1) a novel gradient projection mechanism that rigorously eliminates inter-objective conflicts while preserving the relative influence of each objective proportional to its gradient norm; (2) the first formalization of the instance-wise risk minimization (IWRM) paradigm, treating the loss of each training sample as an independent optimization objective; and (3) an efficient Gramian matrix approximation to substantially reduce memory and computational overhead in Jacobian computation. Theoretical analysis establishes JDโs improved convergence guarantees over standard methods. Empirical evaluation on image classification tasks demonstrates that IWRM consistently outperforms conventional average-loss minimization, validating both the efficacy and scalability of JD.
๐ Abstract
Many optimization problems require balancing multiple conflicting objectives. As gradient descent is limited to single-objective optimization, we introduce its direct generalization: Jacobian descent (JD). This algorithm iteratively updates parameters using the Jacobian matrix of a vector-valued objective function, in which each row is the gradient of an individual objective. While several methods to combine gradients already exist in the literature, they are generally hindered when the objectives conflict. In contrast, we propose projecting gradients to fully resolve conflict while ensuring that they preserve an influence proportional to their norm. We prove significantly stronger convergence guarantees with this approach, supported by our empirical results. Our method also enables instance-wise risk minimization (IWRM), a novel learning paradigm in which the loss of each training example is considered a separate objective. Applied to simple image classification tasks, IWRM exhibits promising results compared to the direct minimization of the average loss. Additionally, we outline an efficient implementation of JD using the Gramian of the Jacobian matrix to reduce time and memory requirements.