correlation analysis

Computing and interpreting correlation-based diagnostics (rank, Spearman, spatial, cross-correlation) to quantify relationships between features/metrics and empirical outcomes and to identify robust predictive signals across datasets or model families.

correlationanalysis

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of effectively integrating radiologists’ assessments with AI predictions in mammographic screening to optimize rule-out and rule-in diagnostic strategies. It introduces, for the first time, a unified joint ROC theoretical framework tailored to both clinical scenarios. By modeling the dependence between physician and AI diagnostic outputs using bivariate copulas, the work theoretically derives—and empirically validates—the impact of their correlation on AUC performance: higher correlation improves rule-out efficacy in diseased populations, whereas lower correlation is preferable in non-diseased populations; conversely, for rule-in tasks, the opposite pattern holds. This framework provides a rigorous theoretical foundation and practical guidance for designing collaborative diagnostic systems that strategically leverage human–AI synergy.

diagnostic testsmammographyROC analysis

Quantifying uncertainty and stability among highly correlated predictors: a subspace perspective

May 10, 2025
XZ
Xiaozhu Zhang
🏛️ University of Washington | University of Southern California

To address the ambiguity in false-positive definition, instability, and poor interpretability of linear feature selection under high feature correlation, this paper proposes a novel subspace-level feature selection framework. It elevates error rate and stability definitions from individual features to feature subspaces and establishes subspace-stable selection theory. We introduce an interchangeable model identification and surrogate structure discovery paradigm, designing a stability-generalization algorithm based on subspace similarity and a surrogate structure detection method, implemented as the R package `substab`. Experiments on synthetic and real gene expression datasets demonstrate that our approach significantly improves cross-sampling stability and model interpretability, while explicitly identifying equivalent model sets under multicollinearity.

Defining false positives in highly correlated feature selectionIdentifying interchangeable features due to multicollinearityMeasuring stability of feature subspaces, not individual features

Correlation vs causation in Alzheimer's disease: an interpretability-driven study

Jun 11, 2025
HD
Hamzah Dabool
🏛️ United Arab Emirates University | Damascus University

This study addresses the critical challenge of distinguishing correlation from causation in Alzheimer’s disease (AD) research. We propose an integrative framework combining interpretable machine learning with multimodal statistical inference. Using clinical, cognitive, genetic (e.g., APOE), and biomarker (e.g., Aβ) data, we employ XGBoost modeling augmented by SHAP values to quantify feature contributions, complemented by Pearson/Spearman correlation analyses to identify stage-specific drivers. To our knowledge, this is the first AD study to synergistically integrate SHAP-based interpretability with association analysis across heterogeneous, multimodal features. Results confirm that amyloid biomarkers exhibit strong association—but not necessary causation—with early cognitive decline, whereas cognitive scores and APOE status serve as robust discriminative features. The framework substantially enhances model transparency and clinical interpretability, establishing a novel paradigm and methodological foundation for causal inference in AD.

Distinguish causation vs correlation in Alzheimer's disease featuresIdentify key diagnostic factors using interpretable machine learningImprove early diagnosis by analyzing true pathological mechanisms

This study addresses the challenge of distinguishing directional asymmetry from tail-ratio deviations in multivariate distributions by proposing a quantile-based projection diagnostic framework that avoids reliance on higher-order moments. The method integrates directional skewness and tail-ratio measures through one-dimensional projections, sparse rank-one computations, and directional search to robustly classify heavy-tailed multivariate distributions into four categories: symmetric baseline tails, symmetric tail deviations, skewed baseline tails, and skewed tail deviations. Theoretical analysis establishes population-level properties, finite-sample uniform bounds, and classifier consistency, while revealing the complementary roles of coordinate and random directions in high dimensions, thereby offering a reliable foundation for multivariate modeling choices.

central symmetrydirectional asymmetrymultivariate data

High-Dimensional Independence Testing via Maximum and Average Distance Correlations

Jan 04, 2020
CS
Cencheng Shen
🏛️ University of Delaware | Temple University

This paper addresses the challenging problem of high-dimensional multivariate independence testing. We propose a novel nonparametric test based on maximum distance correlation (MaxDCOR) and average distance correlation (AvgDCOR). The method unifies Euclidean distance and Gaussian kernel metrics, and—crucially—systematically constructs their corresponding test statistics in high dimensions for the first time. We rigorously establish statistical consistency and derive a fast chi-square approximation for the null distribution, thereby circumventing the high computational cost and limited asymptotic theory inherent in classical distance correlation. Experiments demonstrate that the proposed method achieves over 30% higher detection power than standard distance correlation under sparse strong dependence. It also significantly outperforms existing approaches in both synthetic multivariate dependency settings and real-world cancer–peptide plasma data, effectively capturing complex, high-order multivariate dependence structures.

Comparing maximum and average distance correlations for effectivenessDeveloping non-parametric tests for Euclidean and Gaussian kernel metricsTesting multivariate independence in high-dimensional settings

Latest Papers

What's happening recently
View more

In causal machine learning, it remains unclear whether standard cross-fitting can effectively eliminate bias introduced by black-box algorithms when observational units exhibit spatial, clustered, or time-series dependencies. This study systematically evaluates the performance of conventional cross-fitting that ignores such dependence structures through theoretical analysis and simulation experiments. The findings reveal that, even without explicitly modeling inter-unit dependencies, standard cross-fitting successfully removes the dominant bias term and yields estimation bias and precision comparable to—or sometimes better than—specialized decorrelated folding strategies across a range of correlated data-generating mechanisms. These results challenge the prevailing assumption in the literature that customized cross-fitting procedures are necessary for dependent data, offering theoretical justification for simplifying causal inference pipelines.

bias reductioncausal machine learningcorrelated units

This study investigates how correlations among biomarkers influence the discriminative performance of predictive models, elucidating the mechanism by which adding new biomarkers does not necessarily improve model accuracy. Through theoretical derivations under multivariate normal and skewed distributions, simulation experiments—including log-folded bivariate normal and Gamma distributions—and validation using serum metabolomic data from pancreatic ductal adenocarcinoma patients, the work establishes, for the first time, an analytical relationship between biomarker correlation structures and the area under the ROC curve (AUC). The findings demonstrate that negative correlation most substantially enhances the joint AUC when individual biomarkers exhibit comparable predictive power, and real-world metabolomic data confirm that inter-biomarker correlation plays a decisive role in the performance of disease detection models.

biomarker correlationdiscrimination improvementmultivariate normality

This paper investigates the detectability of shared signals between two high-dimensional variable sets under undersampling regimes. Leveraging random matrix theory, we systematically analyze the signal-resolving capabilities—under additive noise—of three covariance-based estimators: individual autocovariances, cross-covariances, and the joint covariance of concatenated variables. We establish that both cross-covariance and joint covariance estimators surpass the Baik–Bénayoud–Péché phase transition threshold earlier than autocovariance, enabling reliable detection of shared signals; their relative advantage depends critically on the dimensional alignment between the two variable sets. This work provides the first unified characterization of the statistical gain from multivariate collaborative analysis in low signal-to-noise ratio and small-sample settings. It yields theoretical criteria for high-dimensional association inference and principled guidance for estimator selection.

Comparing covariance matrices for signal detectabilityDetecting shared signals between high-dimensional variablesDetermining optimal methods for correlation detection

Weight Space Correlation Analysis: Quantifying Feature Utilization in Deep Learning Models

Dec 15, 2025
CK
Chun Kit Wong
🏛️ Technical University of Denmark | University of Copenhagen

Medical imaging deep learning models are prone to shortcut learning, inadvertently or deliberately exploiting confounding metadata—such as scanner manufacturer—to degrade predictive reliability. To address this, we propose a weight-space correlation analysis method that quantifies, for the first time, a model’s *actual reliance* on confounders embedded in its representations—using cosine similarity among classification head weight vectors as an interpretable, post-hoc metric—rather than merely detecting their presence. Our approach leverages a multi-task projection framework to assess a model’s intrinsic ability to disentangle acquisition-invariant features under unbiased training conditions. Empirical evaluation on the SA-SonoNet architecture for spontaneous preterm birth (sPTB) prediction demonstrates that learned weights exhibit significant correlations with clinically meaningful variables (e.g., birth weight) while remaining decoupled from scanner-related metadata. This work introduces the first quantitative, interpretable diagnostic tool for shortcut learning in medical AI, advancing model trustworthiness and clinical deployability.

Detects shortcut learning in medical imaging modelsQuantifies feature utilization via weight space correlationVerifies model trustworthiness by analyzing feature selectivity

This study addresses the common misuse of Pearson correlation for mixed variable types—such as binary, ordinal, and nominal—which often introduces bias in traditional correlation analyses. To resolve this, the authors introduce smartcor (for R) and pysmartcor (for Python), the first toolkits to systematically support all ten possible combinations of variable types. These packages employ automatic variable-type detection and a rule-based engine to intelligently select the optimal correlation or association method from a repertoire of fourteen, while also providing interpretable justifications for each choice. Monte Carlo simulations demonstrate substantially improved selection accuracy, and real-world case studies reveal meaningful discrepancies between type-aware analyses and naive Pearson correlations, thereby enhancing the reliability of statistical inference.

correlationmethod selectionmixed data

Hot Scholars

LF

Long Feng

Professor of Nankai University
High Dimensional DataHigh Frequency Data
HL

Han Lin Shang

Department of Actuarial Studies and Business Analytics, Macquarie University
Functional data analysisnonparametric smoothingnonparametric statisticsmachine learning
CG

Chenjuan Guo

Professor, East China Normal University
Data AnalyticsMachine Learning
MC

Matteo Cinelli

Assistant Professor @Sapienza University of Rome
Data ScienceNetwork ScienceSocial MediaComputational Social Science
VD

Vince D. Calhoun

Director-Translational Research in Neuroimaging and Data Science (TReNDS;GSU/GAtech/Emory)
brain imaging/MRI/EEG/MEGdata fusiondata scienceimage analysis