Some Simplifications for the Expectation-Maximization (EM) Algorithm: The Linear Regression Model Case

๐Ÿ“… 2025-09-23
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This paper addresses linear regression under missing-at-random (MAR) data by proposing an ANCOVA-equivalent reformulation of the EM algorithm that simplifies computation. The method recasts EM iterations as standard linear or nonlinear regression procedures, enabling constrained prediction and asymptotic variance estimation. Through rigorous theoretical derivation, we establish six theorems thatโ€” for the first timeโ€”unify the maximum likelihood estimation (MLE) consistency foundations of diverse imputation strategies, thereby enhancing interpretability and implementation flexibility. The approach is validated within the SAS PROC MI framework and applied to reanalyze 14 canonical datasets; imputation results match the gold-standard reference exactly, confirming its accuracy, robustness, and broad applicability.

Technology Category

Machine Learning: Matrix & Tensor MethodsConstraint Satisfaction and Optimization: Satisfiability Modulo TheoriesReasoning under Uncertainty: Relational Probabilistic Models

Application Category

Graph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasets
๐Ÿ“ Abstract
The EM algorithm is a generic tool that offers maximum likelihood solutions when datasets are incomplete with data values missing at random or completely at random. At least for its simplest form, the algorithm can be rewritten in terms of an ANCOVA regression specification. This formulation allows several analytical results to be derived that permit the EM algorithm solution to be expressed in terms of new observation predictions and their variances. Implementations can be made with a linear regression or a nonlinear regression model routine, allowing missing value imputations, even when they must satisfy constraints. Fourteen example datasets gleaned from the EM algorithm literature are reanalyzed. Imputation results have been verified with SAS PROC MI. Six theorems are proved that broadly contextualize imputation findings in terms of the theory, methodology, and practice of statistical science.
Problem

Research questions and friction points this paper is trying to address.

Simplifying EM algorithm for maximum likelihood with missing data
Expressing EM solutions via regression models for imputation
Deriving analytical results for prediction and variance estimation
Innovation

Methods, ideas, or system contributions that make the work stand out.

EM algorithm reformulated as ANCOVA regression
Uses linear/nonlinear regression for constrained imputation
Derives analytical solutions for prediction variances
D
Daniel A. Griffith
School of Economic, Political and Policy Sciences, University of Texas at Dallas