🤖 AI Summary
Existing methods for quantifying the association between real-valued and categorical variables often suffer from instability due to parametric assumptions or arbitrary encoding schemes. This work proposes ξ′, the first label-invariant, normalized measure of association that is robust to monotonic transformations of the real-valued variable, along with a sample estimator ξ′ₙ computable in O(n log n) time. The estimator enables Wald-type independence tests and asymptotic confidence intervals without requiring permutation procedures, while simultaneously offering normalization, discriminative power for independence, and the ability to detect functional relationships. Strong theoretical guarantees are provided, including strong consistency and asymptotic normality. Empirical evaluations on both simulated and TCGA data demonstrate that the method achieves encoding stability, high statistical power, and substantial computational efficiency.
📝 Abstract
Quantifying the association between a real-valued variable and a categorical variable is a fundamental task in data analysis. Existing methods often rely on parametric assumptions or arbitrary integer encoding, which may lead to unstable results. We propose a label-invariant population measure of association, $ξ'$, specifically designed for the mixed real-valued-categorical setting. The proposed measure is normalized between 0 and 1; it equals 0 if and only if the variables are independent and 1 if and only if the categorical variable is a measurable function of the real-valued one. We also introduce a corresponding sample estimator, $ξ_n'$, computable in $O(n \log n)$ time. These measures are invariant to permutations of category labels and strictly monotone transformations of the real-valued variable. We establish the strong consistency and asymptotic normality of the estimator $ξ_n'$, enabling a computationally efficient, permutation-free Wald test for independence, and an asymptotic confidence interval for the population measure $ξ'$. Extensive simulations and an application to The Cancer Genome Atlas (TCGA) data demonstrate that the proposed method provides coding stability, competitive power, and substantial computational advantages in nominal mixed-type settings.