Handling Missingness and Censoring in Dirichlet Models

📅 2026-07-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge posed by the coexistence of missingness and censoring in compositional data on the simplex, which renders conventional likelihood-based methods invalid. The work proposes the first unified framework under the Dirichlet model to jointly handle these two types of incomplete observations. Building upon coarsened data theory, a structure-preserving EM algorithm is developed to simultaneously estimate distributional parameters and perform model-driven imputation. The approach rigorously respects the unit-sum constraint inherent to compositional data, thereby preventing information distortion. Experimental results on both simulated and real-world mercury speciation datasets demonstrate that, under complex coarsening mechanisms, the proposed method significantly outperforms existing parametric and nonparametric imputation strategies in accurately recovering the original compositional structure.
📝 Abstract
Likelihood-based inference for compositional data generally requires fully observed compositions, hindering the direct treatment of missing or censored components on the simplex. In this paper, we develop an expectation-maximisation (EM)-type algorithm for maximum likelihood estimation of the Dirichlet parameters in the presence of missing and censored components under a unified coarsening framework. The Dirichlet distribution---the canonical probability model for compositional data, which plays a role analogous to that of the multivariate normal distribution for unconstrained multivariate data---provides the foundation for our methodology. Our methodology preserves the compositional structure of the data while simultaneously performing parameter estimation and model-based imputation. We evaluate the performance of our estimators and imputations through a simulation study under increasingly complex coarsening mechanisms, including both missing and censored data. We compare our method with an existing model-based approach and a nonparametric alternative. Finally, we illustrate the practical utility of our methodology using mercury speciation data, in which compositions are only partially observed because of detection limits and incomplete speciation. Our results indicate that the Dirichlet distribution provides a suitable model for these data and that our method yields imputations that better preserve the observed compositional structure than competing approaches.
Problem

Research questions and friction points this paper is trying to address.

compositional data
missingness
censoring
Dirichlet distribution
simplex
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dirichlet distribution
compositional data
missing data
censoring
EM algorithm
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
J. Pillay
Department of Statistics, University of Pretoria, Pretoria, South Africa
A
A. Bekker
Department of Statistics, University of Pretoria; National Institute for Theoretical and Computational Sciences (NITheCS), Pretoria Node, Pretoria, South Africa
C
C. Tortora
Department of Mathematics and Statistics, San José State University, San José, USA
A
A. Punzo
Department of Economics and Business, University of Catania, Catania, Italy