๐ค AI Summary
To address the slow convergence and poor stability of the traditional EM algorithm for Mixtures of Linear Regressions (MLR) in high-dimensional, noisy, and multi-cluster settings, this paper proposes an incremental seeded EM algorithm that integrates seed-based initialization with a parameter-driven dynamic update mechanism, significantly improving convergence speed and robustness. We innovatively introduce two unsupervised metricsโ*Resolvability* and *X-predictability*โenabling, for the first time, label-free quantitative evaluation of model quality and reliability assessment. Extensive experiments demonstrate that the proposed method consistently outperforms baseline approaches across diverse challenging scenarios. Moreover, the Resolvability index exhibits strong correlation with both regression and clustering performance, providing theoretical foundations and practical tools for MLR model selection, diagnostic analysis, and trustworthy interpretation.
๐ Abstract
This paper proposes Incremental Seeded Expectation Maximization, an algorithm that improves upon the traditional Expectation Maximization computational flow for clusterwise or finite mixture linear regression tasks. The proposed method shows significantly better performance, particularly in scenarios involving high-dimensional input, noisy data, or a large number of clusters. Alongside the new algorithm, this paper introduces the concepts of $ extit{Resolvability}$ and $ extit{X-predictability}$, which enable more rigorous discussions of clusterwise regression problems. The resolvability index is quantified using parameters derived from the model, and results demonstrate its strong connection to model quality without requiring knowledge of the ground truth. This makes the $ extit{Resolvability}$ especially useful for assessing the quality of clusterwise regression models, and by extension, the conclusions drawn from them.