Causal Explainability of Machine Learning in Heart Failure Prediction from Electronic Health Records

📅 2025-06-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses causal inference for clinical variables in heart failure (HF) prediction—specifically, whether statistical correlations or machine learning (ML) feature importances reflect true causal relationships. Method: We propose the first causal structure discovery framework supporting mixed variable types (continuous, categorical, binary), overcoming the limitation of conventional methods that assume continuous variables. It integrates causal discovery algorithms, gradient-boosted trees, hybrid encoding, and a novel causal scoring model to quantify nonlinear causal strength for binary disease outcomes. Contribution/Results: Nonlinear causal modeling significantly outperforms linear alternatives. Key clinical variables—including BNP and LVEF—exhibit high causal strength and strongly align with top ML features (Spearman ρ > 0.82), empirically validating the intrinsic link between causal strength and predictive importance. This alignment enhances model interpretability and clinical credibility, establishing a principled foundation for causally informed HF risk prediction.

Technology Category

Application Category

📝 Abstract
The importance of clinical variables in the prognosis of the disease is explained using statistical correlation or machine learning (ML). However, the predictive importance of these variables may not represent their causal relationships with diseases. This paper uses clinical variables from a heart failure (HF) patient cohort to investigate the causal explainability of important variables obtained in statistical and ML contexts. Due to inherent regression modeling, popular causal discovery methods strictly assume that the cause and effect variables are numerical and continuous. This paper proposes a new computational framework to enable causal structure discovery (CSD) and score the causal strength of mixed-type (categorical, numerical, binary) clinical variables for binary disease outcomes. In HF classification, we investigate the association between the importance rank order of three feature types: correlated features, features important for ML predictions, and causal features. Our results demonstrate that CSD modeling for nonlinear causal relationships is more meaningful than its linear counterparts. Feature importance obtained from nonlinear classifiers (e.g., gradient-boosting trees) strongly correlates with the causal strength of variables without differentiating cause and effect variables. Correlated variables can be causal for HF, but they are rarely identified as effect variables. These results can be used to add the causal explanation of variables important for ML-based prediction modeling.
Problem

Research questions and friction points this paper is trying to address.

Investigating causal explainability of clinical variables in heart failure prediction
Proposing a framework for causal discovery with mixed-type clinical variables
Comparing importance of correlated, ML-prediction, and causal features in HF classification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Proposes causal structure discovery for mixed-type variables
Uses nonlinear CSD modeling for meaningful causal relationships
Links ML feature importance with causal strength directly
🔎 Similar Papers
No similar papers found.