🤖 AI Summary
This study addresses the challenge of jointly modeling missing and truncated values in compositional data by proposing a likelihood-based Expectation–Maximization (EM) algorithm for parameter estimation and imputation within a Dirichlet mixture model framework. The approach respects the simplex constraint and preserves data interpretability, circumventing the information loss inherent in conventional strategies such as case deletion or variable transformation. It represents the first unified methodology capable of simultaneously handling both missingness and truncation in compositional datasets while enabling clustering and model selection. Empirical evaluations demonstrate superior clustering performance and more accurate model selection on simulated data. Applied to real-world datasets—specifically rock-type compositions and PM2.5 source profiles—the method successfully identifies four distinct compositional patterns with clear geological and environmental interpretations.
📝 Abstract
Incomplete compositional data analysis faces a fundamental limitation: likelihood-based methods for compositional models generally require fully observed compositions, making it difficult to accommodate missing or censored proportions directly on the simplex. Consequently, analysts often discard partially observed compositions or transform the data into unconstrained spaces, potentially sacrificing interpretability and coherence. This paper proposes a likelihood-based method for incomplete compositional data without leaving the simplex. Specifically, we develop an Expectation-Maximisation (EM) type algorithm for fitting finite mixtures of Dirichlet distributions in the presence of missing and censored components. The proposed approach performs parameter estimation and model-based imputation simultaneously while preserving the compositional structure and interpretability of the original variables.
A simulation experiment evaluates the performance of the proposed estimators and imputations under increasingly complex coarsening mechanisms. Particular attention is paid to clustering performance, and model selection outcomes. The results showed beneficial clustering performance despite observations being incomplete, and a higher probability of model selection metrics identifying the correct number of clusters compared to current alternative of case-deletion.
The practical utility of the method is illustrated using two real datasets with distinct coarsened patterns. Analysis of the xenolith dataset identifies a four-component Dirichlet mixture that reveals interpretable profiles of rock types and speciation methods. Application to PM$_{2.5}$ speciation data from the Air Quality System, containing both left-censored and missing-at-random values, supports a four-component mixture model that characterises compositional parts of particulate matter across the United States.