🤖 AI Summary
This study addresses the limitation in model merging where implicit regularization introduced during coefficient search constrains weights to a restricted subspace, thereby hindering multi-task performance improvements. To overcome this, we re-examine the implicit regularization mechanism in task arithmetic and propose directly searching the pre-trained weight space via unconstrained optimization, breaking through the subspace constraints inherent in traditional linear combinations. Our work reveals and eliminates these implicit regularization effects, demonstrating that superior solutions reside outside conventional subspaces. Across diverse architectures and extremely data-scarce scenarios, the proposed method significantly outperforms existing model merging techniques, establishing a new optimization paradigm for multi-task learning.
📝 Abstract
Model merging aims to build a multi-task model cheaply by combining the weights of individual task-specific models. To perform well across multiple tasks, most existing merging methods use an additional dataset to find the coefficients for the best linear combination of task-specific weight updates. However, we identify an implicit regularization in this standard practice: searching over coefficients restricts the candidate models to a subspace spanned by task-specific weight updates. In this work, we investigate whether this regularization is actually useful. Surprisingly, empirical results show that optimizing merged-model weights without this regularization significantly boosts the performance of common merging methods across multiple architectures, domains, and even in an extremely data-limited scenario where only one instance is available per class. Moreover, directly optimizing the pretrained model weights even outperforms some existing merging methods. Analysis shows that better multi-task weights exist outside the subspace and can be found using multiple methods. We study different strategies for using the additional dataset, discussing their practical use and implications for model merging. Overall, this work calls for revisiting the existing model-merging pipeline, motivating a broader exploration of the weight space and a reconsideration of the implicit regularization induced by task arithmetic.