dirichlet-tree residualization

Designs and implements algorithms that compute deviance-based residuals from Dirichlet‑tree multinomial (DTM) models to transform or test tree‑structured compositional count features, extending fixed‑dispersion residuals to both internal nodes and leaves while strictly respecting hierarchical compositional constraints. Builds scalable procedures that unify joint and feature‑wise null hypotheses and operate efficiently on sparse input representations.

dirichlet-treeresidualization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.33
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of normalizing sparse and overdispersed compositional count matrices arising from high-throughput sequencing data by proposing a unified residualization framework based on the Dirichlet–multinomial (DM) distribution and its generalizations, such as the Dirichlet-tree multinomial. The approach models within-sample overdispersion via a single concentration parameter, treats each sample as compositional data with a fixed total count, and naturally extends to features with ordinal or tree-structured relationships. The resulting family of residuals encompasses both joint and feature-wise null models, preserves the original sparsity, integrates seamlessly with existing sparse computational pipelines, and allows each residual to be computed in constant time. Theoretically and empirically, the transformation adaptively shrinks residuals according to the degree of overdispersion under repeated sampling, reduces to multinomial residuals for singleton observations, and converges exactly to the classical multinomial case as the concentration parameter tends to infinity.

compositional datacount normalizationDirichlet-multinomial

The application of the Nested Dirichlet distribution has been limited by the need to pre-specify a tree structure and the absence of effective diagnostic tools. This work proposes a data-driven greedy algorithm that, for the first time, enables automatic inference of the underlying tree structure. Furthermore, it introduces saddlepoint approximation–based pseudo-residuals and likelihood displacement measures, providing efficient model-fitting diagnostics even when marginal distributions are not analytically tractable. The method successfully identifies interpretable tree structures in both simulated data and real-world Morris water maze behavioral data, substantially improving model assessment accuracy. To facilitate reproducibility and broader adoption, the authors release an open-source R package implementing the proposed approach.

compositional datamodel diagnosticsNested Dirichlet Distribution

This work addresses the challenge of modeling conditional dependencies among multivariate compositional data under probability simplex constraints, which existing graphical models struggle to handle. The authors propose the first directed tree learning framework tailored for compositional variables. Their approach employs Kullback–Leibler divergence as a scoring function and models child compositions as mixtures of parent-driven and baseline components via column-stochastic transition matrices. This formulation preserves geometric consistency while ensuring edge identifiability through non-degeneracy conditions. Theoretical analysis establishes finite-sample consistency of the method, and it naturally accommodates zero-inflated data. Experiments on both synthetic and real-world microbiome and single-cell datasets successfully recover interpretable directed structures that align with established biological mechanisms.

compositional dataconditional dependencedirected tree

Microbiome multivariate count data often contain outlying observations that induce bias in parameter estimation under the conventional Dirichlet-multinomial (DM) model. This work proposes a contaminated Dirichlet-multinomial (CDM) distribution—the first of its kind—to construct a two-component Bayesian mixture model that explicitly distinguishes between typical low-dispersion and anomalous high-dispersion observations. By retaining all data points, the CDM enables robust parameter inference and simultaneous outlier detection, naturally down-weighting aberrant samples through posterior probabilities while yielding an interpretable estimate of the contamination proportion. Applied to colorectal cancer microbiome data, the CDM consistently outperforms the standard DM model across multiple information criteria and reveals biologically plausible between-group patterns of anomalies.

anomalous observationscontaminationmicrobiome

Latest Papers

What's happening recently
View more

This study addresses the limitations of existing membership inference benchmarks for large language models, including incomplete coverage, distribution misalignment, and insufficient filtering. Building upon OLMo 2, we construct a multi-stage membership inference benchmark spanning the entire pipeline from pre-training to post-training. Methodologically, we explicitly align the distributions of member and non-member data, introduce Shifted variants to evaluate robustness against distribution shifts, and employ infini-gram to rigorously filter non-member samples, thereby controlling confounding factors. Experimental results demonstrate that the optimal attack achieves an AUC of only 0.68, revealing that intermediate training stages exhibit the highest detectability and that unsupervised attacks are highly sensitive to distribution shifts.

Benchmark EvaluationConfounder ControlDistribution Alignment

This study addresses the challenge posed by the coexistence of missingness and censoring in compositional data on the simplex, which renders conventional likelihood-based methods invalid. The work proposes the first unified framework under the Dirichlet model to jointly handle these two types of incomplete observations. Building upon coarsened data theory, a structure-preserving EM algorithm is developed to simultaneously estimate distributional parameters and perform model-driven imputation. The approach rigorously respects the unit-sum constraint inherent to compositional data, thereby preventing information distortion. Experimental results on both simulated and real-world mercury speciation datasets demonstrate that, under complex coarsening mechanisms, the proposed method significantly outperforms existing parametric and nonparametric imputation strategies in accurately recovering the original compositional structure.

censoringcompositional dataDirichlet distribution

Hot Scholars

LK

Lawrence K. Saul

Senior Research Scientist, Flatiron Institute
Machine learning
KR

Kellin Rumsey

Los Alamos National Laboratory
Uncertainty QuantificationBayesian statistics
JO

Josue Obregon

Assistant Professor, Seoul National University of Science and Technology
Machine Learning ExplainabilityMachine LearningFault DetectionProcess Mining
AT

Armando Teixeira-Pinto

Professor of Biostatistics; School of Public Health - University of Sydney
Biostatistics