ZINBGT: Exploratory Data Analysis of Single-Cell Transcriptomic Expression Using Mixture Models

📅 2026-04-10
📈 Citations: 0
Influential: 0
📄 PDF

career value

219K/year
🤖 AI Summary
Single-cell RNA-seq data are inherently noisy and sparse, and existing visualization methods often lack a solid statistical foundation, making it difficult to distinguish technical artifacts from genuine biological signals. This work proposes ZINBGT—a zero-inflated negative binomial mixture model with a geometric tail—that uniquely integrates statistical rigor with biological interpretability. The model provides interpretable visualizations of gene expression at the individual gene level and employs Wasserstein distance to diagnose aberrant genes. Applying ZINBGT, the authors successfully identify outlier-expressing genes in *T. brucei*, uncover intrinsic relationships among sparsity, mean expression, and dispersion in human immune cells, and highlight limitations of current simulated datasets in accurately recapitulating key characteristics of real single-cell data.

Technology Category

Application Category

📝 Abstract
Single-cell transcriptomic data approximates the abundance of proteins at a high resolution, but its noisiness necessitates transformation by a pipeline of methods before analysis and inference. In the absence of robust validation of these pipelines and methods, it remains unclear how best to process any particular dataset. To compensate for this, popular visualisation methods, e.g., t-SNE and UMAP, are commonly used to produce descriptions of datasets. Such visualisations are incomplete and provide subjective descriptions of samples rather than statistically meaningful statements about technical noise or biology. In this paper, we introduce the Zero-Inflated Negative-Binomial with Geometric Tail (ZINBGT), a mixture-model-based strategy for producing interpretable visualisations of each gene's expression across cells, along with diagnostic summaries that use Wasserstein distance to highlight outlier genes. These diagnostics are used to reveal an outlier gene within a T. brucei sample. This method is applied to a human immune-cell dataset, highlighting the relationship between sparsity, mean, and spread across genes, as well as revealing an issue with the use of zero-inflated negative-binomial distributions to model single-cell RNA data. An investigation of simulated datasets intended to replicate the immune-cell data revealed discrepancies with the ground truth, establishing purposes for which these simulated datasets are unsuitable. Finally, we list a number of different domains to which this method can be applied.
Problem

Research questions and friction points this paper is trying to address.

single-cell transcriptomics
data noise
visualization
statistical inference
zero-inflation
Innovation

Methods, ideas, or system contributions that make the work stand out.

ZINBGT
mixture model
single-cell RNA-seq
Wasserstein distance
zero-inflation
🔎 Similar Papers
No similar papers found.
T
Toby Kettlewell
School of Mathematics and Statistics, University of Glasgow, University Place, G12 8QQ, Glasgow, United Kingdom
Y
Yiyi Cheng
School of Infection and Immunity, University of Glasgow, University Place, G12 8QQ, Glasgow, United Kingdom
T
Thomas D. Otto
School of Infection and Immunity, University of Glasgow, University Place, G12 8QQ, Glasgow, United Kingdom; Bernhard Nocht Institute for Tropical Medicine, Hamburg, Germany; University of Hamburg, Hamburg, Germany
V
Vincent Macaulay
School of Mathematics and Statistics, University of Glasgow, University Place, G12 8QQ, Glasgow, United Kingdom
M
Mayetri Gupta
School of Mathematics and Statistics, University of Glasgow, University Place, G12 8QQ, Glasgow, United Kingdom