Correlation vs causation in Alzheimer's disease: an interpretability-driven study

📅 2025-06-11
📈 Citations: 0
Influential: 0
📄 PDF

career value

217K/year
🤖 AI Summary
This study addresses the critical challenge of distinguishing correlation from causation in Alzheimer’s disease (AD) research. We propose an integrative framework combining interpretable machine learning with multimodal statistical inference. Using clinical, cognitive, genetic (e.g., APOE), and biomarker (e.g., Aβ) data, we employ XGBoost modeling augmented by SHAP values to quantify feature contributions, complemented by Pearson/Spearman correlation analyses to identify stage-specific drivers. To our knowledge, this is the first AD study to synergistically integrate SHAP-based interpretability with association analysis across heterogeneous, multimodal features. Results confirm that amyloid biomarkers exhibit strong association—but not necessary causation—with early cognitive decline, whereas cognitive scores and APOE status serve as robust discriminative features. The framework substantially enhances model transparency and clinical interpretability, establishing a novel paradigm and methodological foundation for causal inference in AD.

Technology Category

Application Category

📝 Abstract
Understanding the distinction between causation and correlation is critical in Alzheimer's disease (AD) research, as it impacts diagnosis, treatment, and the identification of true disease drivers. This experiment investigates the relationships among clinical, cognitive, genetic, and biomarker features using a combination of correlation analysis, machine learning classification, and model interpretability techniques. Employing the XGBoost algorithm, we identified key features influencing AD classification, including cognitive scores and genetic risk factors. Correlation matrices revealed clusters of interrelated variables, while SHAP (SHapley Additive exPlanations) values provided detailed insights into feature contributions across disease stages. Our results highlight that strong correlations do not necessarily imply causation, emphasizing the need for careful interpretation of associative data. By integrating feature importance and interpretability with classical statistical analysis, this work lays groundwork for future causal inference studies aimed at uncovering true pathological mechanisms. Ultimately, distinguishing causal factors from correlated markers can lead to improved early diagnosis and targeted interventions for Alzheimer's disease.
Problem

Research questions and friction points this paper is trying to address.

Distinguish causation vs correlation in Alzheimer's disease features
Identify key diagnostic factors using interpretable machine learning
Improve early diagnosis by analyzing true pathological mechanisms
Innovation

Methods, ideas, or system contributions that make the work stand out.

XGBoost algorithm for AD feature identification
SHAP values for interpretable feature contributions
Correlation matrices to reveal variable clusters