🤖 AI Summary
This study addresses the prohibitive computational and memory overhead incurred by coefficient search as model merging scales up. To this end, we propose $\alpha$Transfer, a paradigm grounded in the assumption that coefficient distributions remain consistent within a model family. This method decouples the coefficient search from large models to small proxy models, achieving efficient merging through parameter arithmetic, performance distribution analysis, and cross-model transfer techniques. Experimental results demonstrate that $\alpha$Transfer yields a 6× speedup with 70% memory reduction on Vision Transformers, and a 20× speedup with 85% memory savings on large language models, all while maintaining performance comparable to the original methods.
📝 Abstract
Model merging offers a promising solution for combining multiple fine-tuned checkpoints into a single model through parameter arithmetic. However, finding optimal merging coefficients requires an extensive search that becomes prohibitively expensive as models scale in both size and number, due to high memory requirements and combinatorial growth in the search space. We show that, within the same model family, models exhibit highly congruent performance distributions over merging coefficients across different model sizes. This distributional similarity enables a practical paradigm we call \textit{$α$Transfer}: searching for optimal coefficients on a small proxy model, then directly transfer them to larger target models. We verify $α$Transfer across multiple merging methods, model families, and tasks. Experimental results demonstrate a 6$\times$ speedup and 70\% memory reduction on vision transformers, and a 20$\times$ speedup and 85\% memory reduction on large language models, while maintaining comparable performance. Our findings establish $α$Transfer as an efficient and generalizable approach to scaling model merging.