🤖 AI Summary
Assessing the statistical reliability of linear classifiers in high-dimensional biomedical data—such as ER-positive breast cancer gene expression profiles—remains challenging, particularly in distinguishing true discriminative power from spurious performance due to random correlations.
Method: We propose a novel homogeneity test grounded in linear separability, introducing the first analytically derived upper bound on the linear separability *p*-value. We rigorously prove its high accuracy under bivariate normality and extend it to high dimensions, enabling strict statistical inference on classifier significance—not merely random efficacy. The method integrates linear separability analysis, derivation of *p*-value upper bounds, and paired gene expression testing.
Contribution/Results: Applied to ER-positive breast cancer recurrence prediction, our approach identifies the IGFBP6–ELOVL5 gene pair as exhibiting statistically significant synergistic discriminative capacity (*p* < 0.01), offering a new paradigm for interpretable biomarker discovery.
📝 Abstract
We propose a homogeneity test closely related to the concept of linear separability between two samples. Using the test one can answer the question whether a linear classifier is merely ``random'' or effectively captures differences between two classes. We focus on establishing upper bounds for the test's emph{p}-value when applied to two-dimensional samples. Specifically, for normally distributed samples we experimentally demonstrate that the upper bound is highly accurate. Using this bound, we evaluate classifiers designed to detect ER-positive breast cancer recurrence based on gene pair expression. Our findings confirm significance of IGFBP6 and ELOVL5 genes in this process.