🤖 AI Summary
This work addresses the inaccuracy and lack of robustness in estimating the Gini–Simpson diversity index within nonparametric Bayesian frameworks. To overcome these limitations, we propose an improved approach that incorporates a species-count-dependent symmetric Dirichlet prior and provides a thorough analysis of the Poisson–Dirichlet process. We show that its posterior mean admits an analytically tractable and interpretable form as a convex combination of the classical unbiased estimator and the prior expectation. The proposed method effectively remedies deficiencies of conventional models in both expectation and dispersion, substantially outperforming the standard Dirichlet process in settings with infinitely many species. Theoretical advantages are rigorously confirmed through asymptotic analysis, yielding systematically enhanced accuracy and robustness in diversity estimation.
📝 Abstract
Many statistical problems concern the analysis of species distributions or, more generally, of discrete labeled quantities. Assessing species diversity constitutes a key step toward understanding population structure, and the Gini-Simpson index is among the most widely adopted diversity measures. In this manuscript, we examine several well-established nonparametric prior models for species frequencies and compare them with a newly proposed distribution within this framework. Specifically, we demonstrate that the conventional symmetric Dirichlet distribution leads to certain undesirable properties in terms of expectation and dispersion. These limitations can be mitigated by adopting an alternative symmetric Dirichlet specification, in which the parameter depends on the number of species. This modified formulation is characterized by analytical tractability and interpretability of its main summaries. Furthermore, when the number of distinct species is infinite, the classical Ferguson Dirichlet process exhibits unsatisfactory behavior compared to the more general Poisson-Dirichlet model. Notably, within this latter model, the posterior mean of the diversity index can be expressed as a convex combination of the optimal classical unbiased estimator and the prior expectation. Theoretical results are further supported by asymptotic analyses and systematically compared with their classical counterparts.