Evaluating the stability of model explanations in instance-dependent cost-sensitive credit scoring

📅 2025-09-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Interpretability stability remains underexplored in instance-dependent cost-sensitive (IDCS) credit scoring, particularly regarding how IDCS loss functions affect explanation consistency—despite growing regulatory demands for model transparency. Method: We introduce a novel stability metric and conduct the first systematic evaluation of feature importance ranking stability for LIME and SHAP across four public credit datasets, incorporating controlled resampling to quantify the impact of class imbalance. Contribution/Results: While IDCS models substantially improve cost-effectiveness, they exhibit significantly lower interpretability stability than conventional models—and this degradation intensifies sharply with increasing class imbalance. Our findings uncover a fundamental trade-off between cost optimization and explanation stability, providing both theoretical grounding and practical assessment tools for developing regulation-compliant, trustworthy credit scoring systems.

Technology Category

Machine Learning: Learning Preferences or RankingsComputer Vision: Interpretability, Explainability, and TransparencyPhilosophy and Ethics of AI: Accountability, Interpretability & Explainability

Application Category

User Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalizationSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Instance-dependent cost-sensitive (IDCS) classifiers offer a promising approach to improving cost-efficiency in credit scoring by tailoring loss functions to instance-specific costs. However, the impact of such loss functions on the stability of model explanations remains unexplored in literature, despite increasing regulatory demands for transparency. This study addresses this gap by evaluating the stability of Local Interpretable Model-agnostic Explanations (LIME) and SHapley Additive exPlanations (SHAP) when applied to IDCS models. Using four publicly available credit scoring datasets, we first assess the discriminatory power and cost-efficiency of IDCS classifiers, introducing a novel metric to enhance cross-dataset comparability. We then investigate the stability of SHAP and LIME feature importance rankings under varying degrees of class imbalance through controlled resampling. Our results reveal that while IDCS classifiers improve cost-efficiency, they produce significantly less stable explanations compared to traditional models, particularly as class imbalance increases, highlighting a critical trade-off between cost optimization and interpretability in credit scoring. Amid increasing regulatory scrutiny on explainability, this research underscores the pressing need to address stability issues in IDCS classifiers to ensure that their cost advantages are not undermined by unstable or untrustworthy explanations.
Problem

Research questions and friction points this paper is trying to address.

Evaluating stability of model explanations in cost-sensitive credit scoring
Assessing impact of instance-dependent loss functions on explanation stability
Investigating trade-off between cost optimization and interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evaluating LIME and SHAP explanation stability
Introducing novel cross-dataset comparability metric
Investigating feature importance under class imbalance
🔎 Similar Papers
No similar papers found.