Score
Design, build, and evaluate methods and tools that explain how machine learning models produce outputs, including computing feature attributions (e.g., SHAP), producing global and local explanations, and generating visualizations such as partial dependence plots. Analyze feature interactions, trends, and decision drivers to make model behavior interpretable and to communicate which inputs influence specific predictions.
This study addresses the strong model dependency of SHAP value interpretations, which lacks a standardized analytical framework and thereby limits reliable explanations of black-box model decisions in high-stakes applications. For the first time, this work systematically evaluates SHAP explanations across multiple mainstream machine learning models on diverse datasets, uncovering consistent patterns of model dependence. Furthermore, it proposes a generalized waterfall plot visualization method tailored for multi-class classification problems. Experimental results demonstrate the effectiveness and practical utility of the proposed approach, offering both theoretical grounding and actionable guidance for practitioners in the field of explainable artificial intelligence.
This work addresses the theoretical shortcomings of prevailing post-hoc feature attribution methods—such as SHAP—which often lack formal guarantees and can mislead human decision-making in high-stakes settings. To overcome this limitation, the paper introduces a novel paradigm grounded in symbolic explainable artificial intelligence (XAI), leveraging formal, verifiable symbolic reasoning to reconstruct feature importance assignments. By integrating formal verification with feature attribution analysis, the proposed framework establishes a theoretically sound and certifiable approach to interpretability. This integration significantly enhances the reliability and trustworthiness of XAI in safety-critical applications, offering a rigorous and verifiable pathway for generating explanations in high-risk machine learning systems.
This work addresses the longstanding challenge of jointly optimizing model interpretability and predictive performance in machine learning. We propose a SHAP-driven regularized training framework, whose core innovation is the first direct incorporation of TreeSHAP attribution values into the loss function via a novel joint regularization term based on the entropy of the attribution distribution. This term simultaneously enforces sparsity, concentration, and cross-sample stability of feature importances. Unlike post-hoc methods, our approach is end-to-end trainable, applicable to mainstream tree-based models (e.g., XGBoost, LightGBM), and supports both regression and classification tasks. Extensive experiments across multiple benchmark datasets demonstrate that the proposed method improves model generalization, yields more robust and interpretable SHAP attributions, and maintains or exceeds baseline accuracy—without sacrificing predictive performance.
This work addresses the fundamental trade-off in machine learning between high predictive performance and low interpretability inherent in “black-box” models (e.g., deep neural networks, ensemble methods). It rigorously distinguishes post-hoc explanation—applied after model training—from inherently interpretable modeling—designed for transparency from inception. To reconcile accuracy and interpretability, we propose a hybrid modeling paradigm centered on symbolic knowledge embedding, integrating differentiable symbolic modules, knowledge distillation, and symbolic reasoning into the model architecture itself. This enables joint optimization of fidelity and interpretability at the design stage. Extensive experiments across diverse domains demonstrate that our approach matches the predictive accuracy of state-of-the-art black-box models while generating human-understandable, logically grounded decision rules. As a result, it substantially enhances model trustworthiness and deployment viability in safety- and accountability-critical applications.
This study addresses the troubling inconsistency in feature attribution methods such as SHAP, which can yield substantially divergent explanations even for identical inputs and models, thereby undermining trustworthiness and auditability in high-stakes applications. The work formally defines and quantifies the phenomenon of “explanation multiplicity,” distinguishing its origins in model training or selection from inherent randomness in the explanation procedure itself. To assess stability, the authors introduce a dual-perspective metric incorporating both feature magnitude and ranking, establish a randomized null model as an interpretable baseline, and develop a comprehensive empirical evaluation framework spanning diverse datasets and model classes. Experiments demonstrate that explanation multiplicity is pervasive; relying solely on SHAP value magnitudes can lead to misleading conclusions, necessitating rank-sensitive metrics and principled baselines for reliable interpretability assessment.
This work addresses the limitations of existing explainable AI methods, which predominantly focus on associative predictions and fall short in supporting decision-making that requires causal reasoning and counterfactual analysis. To bridge this gap, the paper proposes a novel framework that integrates causal machine learning with intrinsically interpretable models—such as additive models and symbolic regression—by explicitly embedding causal inference mechanisms within the model architecture. This approach enables the explicit recovery of causal structures and functional forms among variables directly from cross-sectional data. While maintaining high predictive accuracy, the method achieves comprehensive transparency in system structure, causal relationships, and response mechanisms, thereby substantially enhancing both interpretability and causal reliability for trustworthy “What-if” analyses.
This study identifies a critical disconnect between generalizability and interpretability in machine learning–based agricultural yield forecasting: while models—including XGBoost, Random Forest, LSTM, and TCN—achieve high accuracy under spatially partitioned test sets at the German NUTS-3 level, their performance deteriorates substantially when evaluated on temporally independent validation years. Crucially, even under temporal generalization failure, SHAP-based attributions retain high confidence scores, revealing a fundamental limitation of post-hoc interpretability methods—the illusion of reliability. To address this, we propose “validation-aware interpretability,” a novel paradigm that grounds model explanation in robust spatiotemporal cross-validation. We rigorously demonstrate that high test-set accuracy does not imply reliable generalization, and that trustworthy interpretation must be predicated on temporal robustness. This work establishes validation-aware interpretability as a necessary condition for credible, actionable insights in time-sensitive agroecological modeling.
Current machine learning evaluation practices predominantly rely on surface-level performance metrics, often neglecting the internal mechanisms of models. This work proposes trustworthy interpretability as a central evaluation paradigm and, for the first time, systematically demonstrates that it satisfies core criteria from the philosophy of science—namely falsifiability, reproducibility, and predictive power. By constructing an evaluation framework that integrates causal analysis with mechanistic probing, the study delineates three functional pathways through which interpretability enables the identification of behavioral origins, detection of latent flaws, and prediction of potential failure modes. This approach advances model assessment beyond performance-oriented benchmarks toward a deeper understanding of underlying mechanisms.
SHAP-based visualizations of feature importance lack semantically meaningful explanations for non-technical users, undermining interpretability and trust. To address this, we propose LLM-SHAP—a framework that tightly integrates large language models (e.g., GPT) with SHAP, dynamically injecting user-provided feature aliases, domain-specific descriptions, and contextual background into the explanation generation process. This enables context-aware, natural-language interpretations without model retraining, synergistically enhancing both visual and textual explanations. We evaluate LLM-SHAP in a medical diagnosis use case via a user study. Results demonstrate statistically significant improvements over standard SHAP visualizations: +38.2% higher perceived interpretability and +41.5% greater contextual relevance among end users. Our approach establishes a novel paradigm for deploying explainable AI in real-world, non-expert settings—bridging the gap between technical model outputs and human-understandable, actionable insights.