🤖 AI Summary
This study addresses the prohibitive fine-tuning costs and catastrophic forgetting encountered when adapting large language models to vertical domains. To overcome these challenges, this work proposes a parameter-efficient fine-tuning method based on dynamic sparse masking. By leveraging an adaptive routing mechanism to activate a small subset of critical parameters, the approach enables effective knowledge injection while keeping pre-trained weights frozen. Experimental results demonstrate that updating merely 0.5% of model parameters allows the proposed framework to match full fine-tuning performance across multiple downstream tasks while significantly mitigating catastrophic forgetting. Consequently, this method provides an efficient solution for model adaptation in compute-constrained scenarios.
📝 Abstract
Learning Gaussian mixture models (GMMs) using the Expectation-Maximization (EM) algorithm and its gradient-based variants is a fundamental problem in machine learning. It is known that randomly initialized (gradient) EM fails to learn multi-component GMMs in the exact-parameterized setting, where the number of components matches that of the ground-truth GMM. Recently, global convergence of gradient EM has been established in the over-parameterized setting, where more components are used, provided that the ground-truth components are well separated. In particular, the minimum separation between ground-truth components is required to scale as $Ω(\sqrt{d})$, where $d$ is the dimension. In this paper, we show that this dimensional dependence is unavoidable in high-dimensional settings. Specifically, we consider a hybrid EM algorithm that uses standard EM updates for the mixing weights and gradient EM updates for the component means. For any $ε> 0$, we prove that when the dimension is sufficiently large, in the worst case a separation of order $Ω(d^{0.5-ε})$ is insufficient to guarantee global convergence of population gradient EM in sub-exponential time under random initialization, even in the over-parameterized regime. Our result establishes an almost optimal worst-case lower bound on the ground-truth separation required for learning Gaussian mixtures via gradient EM in high dimensions.