A likelihood-based coefficient for biomedical independence testing: the binomial-cut composite likelihood ratio

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inability of conventional correlation coefficients to distinguish independence from non-monotonic dependence, particularly their failure to capture threshold effects and heteroscedasticity in biomarkers. To this end, we propose xi_cut, an independence test coefficient ranging from 0 to 1 based on the composite likelihood ratio of binomial cuts. This coefficient provides a distribution-free lower bound for mutual information and unifies multiple dependence measures. Methodologically, it employs Nadaraya-Watson plug-in estimation, maximum grid bandwidth selection, and exact permutation inference. Experimental results demonstrate that xi_cut significantly outperforms rank-based correlation methods under non-monotonic simulation settings. Furthermore, when applied to proteomics data, the proposed approach successfully identifies a greater number of age-dependent proteins compared to existing techniques.
📝 Abstract
The standard dependence summaries used in biomarker studies -- Pearson's r, Spearman's rho, Kendall's tau -- take values in [-1, 1] with 0 indicating no linear or monotone association. Zero does not distinguish independence from non-monotone dependence, so the scale cannot represent threshold effects, heteroscedasticity, and tail shifts common in biomarker practice. We formulate independence testing as a composite Bernoulli likelihood ratio: at each threshold t, comparing the Bernoulli laws of 1(Y <= t) conditionally on X versus marginally, aggregated over cut points. The resulting coefficient xi_cut lies on [0, 1] with 0 iff X and Y are independent (under continuity of Y) and 1 iff Y is a measurable function of X. Fisher weighting arises at second order from the Bernoulli likelihood, and xi_cut equals twice the threshold-averaged mutual information between X and 1(Y <= t), giving a distribution-free lower bound on I(X; Y). A second-order expansion recovers the Fisher-weighted Dette-Siburg-Stoimenov measure, which coincides under continuity with Chatterjee's rank correlation. Estimation uses a Nadaraya-Watson plug-in with a max-over-grid bandwidth; inference is by exact permutation. In biomarker-motivated simulations T_cut substantially outperforms rank-based coefficients on W-shaped non-monotone and heteroscedastic alternatives. We illustrate on the Seattle cohort (n=70, ages 21-88) of the aging plasma proteome dataset, screening all 1,305 proteins for age dependence: under Benjamini-Hochberg control at q<0.05, T_cut rejects on 70 proteins, six of which are missed by Pearson, Spearman, and Chatterjee at the same FDR level.
Problem

Research questions and friction points this paper is trying to address.

independence testing
biomarker studies
non-monotone dependence
composite likelihood ratio
heteroscedasticity
Innovation

Methods, ideas, or system contributions that make the work stand out.

composite likelihood ratio
independence testing
mutual information lower bound
non-monotone dependence
Nadaraya-Watson estimation
💼 Related Jobs
No related jobs found.