🤖 AI Summary
Traditional model ensembling relies on averaging numerous fine-tuned models, incurring high computational cost and low efficiency. This paper proposes Model Stock: an efficient ensemble method requiring only two fine-tuned models. Its key insight is that fine-tuned weights closer to the layer-wise weight center exhibit superior in-distribution (ID) and out-of-distribution (OOD) generalization. Leveraging this, Model Stock introduces a layer-wise dual-model weighted averaging strategy designed to approximate the center of the weight space. Built upon the CLIP architecture, it jointly optimizes layer-wise center approximation and OOD robustness. On standard ID/OOD benchmarks, Model Stock consistently outperforms state-of-the-art methods—including Model Soup—achieving higher ID accuracy and stronger OOD robustness, while introducing negligible inference overhead.
📝 Abstract
This paper introduces an efficient fine-tuning method for large pre-trained models, offering strong in-distribution (ID) and out-of-distribution (OOD) performance. Breaking away from traditional practices that need a multitude of fine-tuned models for averaging, our approach employs significantly fewer models to achieve final weights yet yield superior accuracy. Drawing from key insights in the weight space of fine-tuned weights, we uncover a strong link between the performance and proximity to the center of weight space. Based on this, we introduce a method that approximates a center-close weight using only two fine-tuned models, applicable during or after training. Our innovative layer-wise weight averaging technique surpasses state-of-the-art model methods such as Model Soup, utilizing only two fine-tuned models. This strategy can be aptly coined Model Stock, highlighting its reliance on selecting a minimal number of models to draw a more optimized-averaged model. We demonstrate the efficacy of Model Stock with fine-tuned models based upon pre-trained CLIP architectures, achieving remarkable performance on both ID and OOD tasks on the standard benchmarks, all while barely bringing extra computational demands. Our code and pre-trained models are available at https://github.com/naver-ai/model-stock.