A New Look at Gaussian Mixtures in the Presence of Missing-at-Random Responses and Covariates

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses multivariate linear regression and clustering when both the response variable and covariates are subject to missingness at random (MAR). The authors propose a general framework that employs a conditional–marginal decomposition to perform role-aware reparameterization of the joint Gaussian distribution, preserving the structural asymmetry between responses and covariates. By embedding variable roles directly into the distributional modeling, the approach avoids biases introduced by conventional symmetric treatments of variables. Missing values are imputed via an expectation–maximization (EM) algorithm, and the framework naturally extends to Gaussian mixture models to enable soft clustering. Simulation studies and experiments on the Automobile dataset demonstrate that the proposed reparameterization strategy substantially improves both parameter estimation accuracy and clustering performance.
📝 Abstract
Missing values present a common challenge in statistical modeling, so handling them properly is an important research direction. Among the various mechanisms that can generate missing values, the most common is the missing-at-random (MAR) mechanism, in which the probability of missingness depends only on observed data and not on unobserved data. This paper addresses the problem of estimating a multivariate linear regression model with multiple random covariates in the presence of MAR values in both the response and covariate spaces using a maximum likelihood (ML) framework. The proposed methodology models the joint distribution of responses and covariates through a conditional-marginal factorization of a multivariate Gaussian distribution. This formulation can be interpreted as a reparameterization of the multivariate normal distribution when the variables can be naturally partitioned into responses and covariates. Parameter estimation is performed using the expectation-maximization (EM) algorithm, which facilitates the imputation of missing values while preserving the distinct roles of responses and covariates. We extend this framework to the model-based clustering setting by considering a mixture of multivariate linear regressions with multiple random covariates. This extension enables soft clustering under incomplete data and accommodates MAR values in both the multivariate responses and covariates. Hence, it represents one of the most general model-based clustering solutions for regression data currently available in the literature. The effectiveness of the methodology is demonstrated through a simulation study, and the advantages of the proposed reparameterization are illustrated using the Automobile dataset, which contains missing values.
Problem

Research questions and friction points this paper is trying to address.

missing-at-random
multivariate linear regression
model-based clustering
Gaussian mixtures
incomplete data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gaussian mixture model
missing-at-random (MAR)
reparameterization
EM algorithm
model-based clustering
🔎 Similar Papers
No similar papers found.