🤖 AI Summary
This paper addresses the challenge of unknown moment properties for self-normalized statistics—such as ratios and the Gini coefficient—under nonnegative data. We propose a general framework for computing exact moments based solely on sample means. For the first time, we derive closed-form moment expressions whose computational complexity remains constant with increasing sample size, overcoming scalability limitations of conventional methods in high dimensions or large samples. Theoretical analysis, supported by extensive numerical experiments, reveals novel patterns in bias and variance dependence on distributional skewness and sample size—enabling the design of a broadly applicable debiasing algorithm. The derived closed-form solutions apply to diverse ratio-based metrics across economics, ecology, and machine learning. Empirical evaluation on both synthetic and real-world datasets demonstrates substantial improvements in estimation accuracy, with average bias reduction exceeding 60%, while maintaining computational scalability and practical utility.
📝 Abstract
Following the student t-statistic, normalization has been a widely used method in statistic and other disciplines including economics, ecology and machine learning. We focus on statistics taking the form of a ratio over (some power of) the sample mean, the probabilistic features of which remain unknown. We develop a unified formula for the moments of these self-normalized statistics with non-negative observations, yielding closed-form expressions for several important cases. Moreover, the complexity of our formula doesn't scale with the sample size $n$. Our theoretical findings, supported by extensive numerical experiments, reveal novel insights into their bias and variance, and we propose a debiasing method illustrated with applications such as the odds ratio, Gini coefficient and squared coefficient of variation.