๐ค AI Summary
This paper investigates the classification error behavior of Mixture Discriminant Analysis (MDA) under model over-specificationโi.e., when the number of mixture components exceeds the true number underlying the data distribution. Focusing on unimodal Gaussian generative data, we model each class as a two-component Gaussian mixture and employ the EM algorithm for parameter estimation and convergence analysis, with applications to remote sensing image classification. We establish, for the first time, a rigorous theoretical framework for over-specified MDA: under reasonable initialization, the EM algorithm converges exponentially fast to the Bayes-optimal risk; moreover, in finite-sample settings, the classification error converges to the Bayes risk at rate $n^{-1/2}$. Both theoretical analysis and remote sensing experiments demonstrate that over-specified MDA remains capable of approaching the Bayes error bound, thereby substantially enhancing robustness and practicality under model misspecification.
๐ Abstract
This study explores the classification error of Mixture Discriminant Analysis (MDA) in scenarios where the number of mixture components exceeds those present in the actual data distribution, a condition known as overspecification. We use a two-component Gaussian mixture model within each class to fit data generated from a single Gaussian, analyzing both the algorithmic convergence of the Expectation-Maximization (EM) algorithm and the statistical classification error. We demonstrate that, with suitable initialization, the EM algorithm converges exponentially fast to the Bayes risk at the population level. Further, we extend our results to finite samples, showing that the classification error converges to Bayes risk with a rate $n^{-1/2}$ under mild conditions on the initial parameter estimates and sample size. This work provides a rigorous theoretical framework for understanding the performance of overspecified MDA, which is often used empirically in complex data settings, such as image and text classification. To validate our theory, we conduct experiments on remote sensing datasets.