🤖 AI Summary
This study addresses the challenge of reduced predictive reliability in credit card fraud detection caused by severe class imbalance. The authors propose an optimized modeling approach based on the Explainable Boosting Machine (EBM), which systematically integrates Taguchi experimental design to fine-tune both data preprocessing sequences and model hyperparameters. By leveraging feature importance analysis and modeling interaction effects, the method enhances performance without resorting to conventional sampling techniques that may introduce bias. Evaluated on a standard credit card transaction dataset, the proposed approach achieves a ROC-AUC of 0.983, outperforming the baseline EBM (0.975) as well as established models including logistic regression, random forest, XGBoost, and decision trees. The framework successfully balances high predictive accuracy with strong interpretability, maintaining transparency without compromising performance.
📝 Abstract
Addressing class imbalance is a central challenge in credit card fraud detection, as it directly impacts predictive reliability in real-world financial systems. To overcome this, the study proposes an enhanced workflow based on the Explainable Boosting Machine (EBM)-a transparent, state-of-the-art implementation of the GA2M algorithm-optimized through systematic hyperparameter tuning, feature selection, and preprocessing refinement. Rather than relying on conventional sampling techniques that may introduce bias or cause information loss, the optimized EBM achieves an effective balance between accuracy and interpretability, enabling precise detection of fraudulent transactions while providing actionable insights into feature importance and interaction effects. Furthermore, the Taguchi method is employed to optimize both the sequence of data scalers and model hyperparameters, ensuring robust, reproducible, and systematically validated performance improvements. Experimental evaluation on benchmark credit card data yields an ROC-AUC of 0.983, surpassing prior EBM baselines (0.975) and outperforming Logistic Regression, Random Forest, XGBoost, and Decision Tree models. These results highlight the potential of interpretable machine learning and data-driven optimization for advancing trustworthy fraud analytics in financial systems.