🤖 AI Summary
This study addresses estimation bias and hypothesis test failure arising from frequent zero counts in categorical data analysis. Methodologically: (1) it develops a robust Bayesian estimator for the binomial distribution; (2) it designs a regularized maximum likelihood–based hypothesis testing framework, covering sign tests, homogeneity tests, and symmetry tests; and (3) it introduces analytically tractable regularized measures of association for contingency tables—including regularized mutual information—that explicitly accommodate zero frequencies. The contributions are threefold: first, substantially improved stability and interpretability of probability estimation, hypothesis testing, and association assessment under data sparsity; second, provision of a theoretically rigorous yet computationally feasible foundational toolkit for high-dimensional sparse categorical data; and third, enhanced reliability of practical statistical inference in real-world applications involving sparse contingency structures.
📝 Abstract
This paper explores Bayesian estimation for categorical data, focusing on simple yet effective models that provide a foundation for applying more advanced methods accurately and reliably in real-world applications. We begin by revisiting Bayesian estimators for the binomial distribution and investigating their properties. Next, we develop hypothesis tests for categorical data (sign test, homogeneity test, symmetry test) based on regularized maximum likelihood estimates of the probabilities. Finally, we formulate regularized versions of common association measures for contingency tables and study the regularized version of mutual information, particular for the situation where the regularized version can effectively handle zero counts.