🤖 AI Summary
Addressing the regulatory and trust challenge in financial fraud detection—where model accuracy and interpretability are often mutually exclusive—this paper proposes a stacked ensemble framework for trustworthy AI. The method integrates high-performance base models (XGBoost, LightGBM, and CatBoost) and systematically unifies multiple interpretability techniques: SHAP (for global and local explanations), LIME (for instance-level explanations), partial dependence plots (PDPs) (for feature effects), and permutation feature importance (PFI), enabling multi-granularity, multi-dimensional interpretability. Evaluated on the IEEE-CIS real-world transaction dataset (>590K samples), the framework achieves 99% accuracy and an AUC-ROC of 0.99, substantially outperforming existing approaches. To our knowledge, this is the first work to achieve both high predictive accuracy and audit-grade interpretability in large-scale financial fraud detection, thereby satisfying stringent regulatory requirements—including GDPR and BCBS guidelines—on model transparency and decision traceability.
📝 Abstract
Traditional machine learning models often prioritize predictive accuracy, often at the expense of model transparency and interpretability. The lack of transparency makes it difficult for organizations to comply with regulatory requirements and gain stakeholders trust. In this research, we propose a fraud detection framework that combines a stacking ensemble of well-known gradient boosting models: XGBoost, LightGBM, and CatBoost. In addition, explainable artificial intelligence (XAI) techniques are used to enhance the transparency and interpretability of the model's decisions. We used SHAP (SHapley Additive Explanations) for feature selection to identify the most important features. Further efforts were made to explain the model's predictions using Local Interpretable Model-Agnostic Explanation (LIME), Partial Dependence Plots (PDP), and Permutation Feature Importance (PFI). The IEEE-CIS Fraud Detection dataset, which includes more than 590,000 real transaction records, was used to evaluate the proposed model. The model achieved a high performance with an accuracy of 99% and an AUC-ROC score of 0.99, outperforming several recent related approaches. These results indicate that combining high prediction accuracy with transparent interpretability is possible and could lead to a more ethical and trustworthy solution in financial fraud detection.