Is $\sqrt{d}$ Separation Necessary for Gradient EM to Learn Gaussian Mixtures in High Dimensions?

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive fine-tuning costs and catastrophic forgetting encountered when adapting large language models to vertical domains. To overcome these challenges, this work proposes a parameter-efficient fine-tuning method based on dynamic sparse masking. By leveraging an adaptive routing mechanism to activate a small subset of critical parameters, the approach enables effective knowledge injection while keeping pre-trained weights frozen. Experimental results demonstrate that updating merely 0.5% of model parameters allows the proposed framework to match full fine-tuning performance across multiple downstream tasks while significantly mitigating catastrophic forgetting. Consequently, this method provides an efficient solution for model adaptation in compute-constrained scenarios.
📝 Abstract
Learning Gaussian mixture models (GMMs) using the Expectation-Maximization (EM) algorithm and its gradient-based variants is a fundamental problem in machine learning. It is known that randomly initialized (gradient) EM fails to learn multi-component GMMs in the exact-parameterized setting, where the number of components matches that of the ground-truth GMM. Recently, global convergence of gradient EM has been established in the over-parameterized setting, where more components are used, provided that the ground-truth components are well separated. In particular, the minimum separation between ground-truth components is required to scale as $Ω(\sqrt{d})$, where $d$ is the dimension. In this paper, we show that this dimensional dependence is unavoidable in high-dimensional settings. Specifically, we consider a hybrid EM algorithm that uses standard EM updates for the mixing weights and gradient EM updates for the component means. For any $ε> 0$, we prove that when the dimension is sufficiently large, in the worst case a separation of order $Ω(d^{0.5-ε})$ is insufficient to guarantee global convergence of population gradient EM in sub-exponential time under random initialization, even in the over-parameterized regime. Our result establishes an almost optimal worst-case lower bound on the ground-truth separation required for learning Gaussian mixtures via gradient EM in high dimensions.
Problem

Research questions and friction points this paper is trying to address.

Gaussian Mixture Models
Gradient EM
High Dimensions
Separation Condition
Global Convergence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gaussian Mixture Models
Gradient EM
Over-parameterization
Separation Condition
High-dimensional Learning
💼 Related Jobs
No related jobs found.