model explainability

Design, build, and evaluate methods and tools that explain how machine learning models produce outputs, including computing feature attributions (e.g., SHAP), producing global and local explanations, and generating visualizations such as partial dependence plots. Analyze feature interactions, trends, and decision drivers to make model behavior interpretable and to communicate which inputs influence specific predictions.

modelexplainability

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$205K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the strong model dependency of SHAP value interpretations, which lacks a standardized analytical framework and thereby limits reliable explanations of black-box model decisions in high-stakes applications. For the first time, this work systematically evaluates SHAP explanations across multiple mainstream machine learning models on diverse datasets, uncovering consistent patterns of model dependence. Furthermore, it proposes a generalized waterfall plot visualization method tailored for multi-class classification problems. Experimental results demonstrate the effectiveness and practical utility of the proposed approach, offering both theoretical grounding and actionable guidance for practitioners in the field of explainable artificial intelligence.

explainable AIfeature contributionmachine learning models

This work addresses the theoretical shortcomings of prevailing post-hoc feature attribution methods—such as SHAP—which often lack formal guarantees and can mislead human decision-making in high-stakes settings. To overcome this limitation, the paper introduces a novel paradigm grounded in symbolic explainable artificial intelligence (XAI), leveraging formal, verifiable symbolic reasoning to reconstruct feature importance assignments. By integrating formal verification with feature attribution analysis, the proposed framework establishes a theoretically sound and certifiable approach to interpretability. This integration significantly enhances the reliability and trustworthiness of XAI in safety-critical applications, offering a rigorous and verifiable pathway for generating explanations in high-risk machine learning systems.

explainable artificial intelligencefeature attributionrigor

SHAP-Guided Regularization in Machine Learning Models

Jul 31, 2025
AS
Amal Saadallah
🏛️ Lamarr Institute for Machine Learning and AI

This work addresses the longstanding challenge of jointly optimizing model interpretability and predictive performance in machine learning. We propose a SHAP-driven regularized training framework, whose core innovation is the first direct incorporation of TreeSHAP attribution values into the loss function via a novel joint regularization term based on the entropy of the attribution distribution. This term simultaneously enforces sparsity, concentration, and cross-sample stability of feature importances. Unlike post-hoc methods, our approach is end-to-end trainable, applicable to mainstream tree-based models (e.g., XGBoost, LightGBM), and supports both regression and classification tasks. Extensive experiments across multiple benchmark datasets demonstrate that the proposed method improves model generalization, yields more robust and interpretable SHAP attributions, and maintains or exceeds baseline accuracy—without sacrificing predictive performance.

Applies entropy-based penalties for sparse, stable feature attributionsImproves generalization performance with robust, explainable modelsIncorporates SHAP-guided regularization to enhance predictive performance and interpretability

Investigating the Duality of Interpretability and Explainability in Machine Learning

Oct 28, 2024
MG
Moncef Garouani
🏛️ Université Toulouse Capitole | Université de Toulouse | Aix-Marseille University

This work addresses the fundamental trade-off in machine learning between high predictive performance and low interpretability inherent in “black-box” models (e.g., deep neural networks, ensemble methods). It rigorously distinguishes post-hoc explanation—applied after model training—from inherently interpretable modeling—designed for transparency from inception. To reconcile accuracy and interpretability, we propose a hybrid modeling paradigm centered on symbolic knowledge embedding, integrating differentiable symbolic modules, knowledge distillation, and symbolic reasoning into the model architecture itself. This enables joint optimization of fidelity and interpretability at the design stage. Extensive experiments across diverse domains demonstrate that our approach matches the predictive accuracy of state-of-the-art black-box models while generating human-understandable, logically grounded decision rules. As a result, it substantially enhances model trustworthiness and deployment viability in safety- and accountability-critical applications.

Addressing the need for transparent and trustworthy machine learning modelsClarifying the difference between explaining black box models and using inherently interpretable onesEvaluating hybrid methods combining symbolic knowledge with neural networks for interpretability

This study addresses the troubling inconsistency in feature attribution methods such as SHAP, which can yield substantially divergent explanations even for identical inputs and models, thereby undermining trustworthiness and auditability in high-stakes applications. The work formally defines and quantifies the phenomenon of “explanation multiplicity,” distinguishing its origins in model training or selection from inherent randomness in the explanation procedure itself. To assess stability, the authors introduce a dual-perspective metric incorporating both feature magnitude and ranking, establish a randomized null model as an interpretable baseline, and develop a comprehensive empirical evaluation framework spanning diverse datasets and model classes. Experiments demonstrate that explanation multiplicity is pervasive; relying solely on SHAP value magnitudes can lead to misleading conclusions, necessitating rank-sensitive metrics and principled baselines for reliable interpretability assessment.

explanation multiplicityexplanation stabilityfeature attribution

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing explainable AI methods, which predominantly focus on associative predictions and fall short in supporting decision-making that requires causal reasoning and counterfactual analysis. To bridge this gap, the paper proposes a novel framework that integrates causal machine learning with intrinsically interpretable models—such as additive models and symbolic regression—by explicitly embedding causal inference mechanisms within the model architecture. This approach enables the explicit recovery of causal structures and functional forms among variables directly from cross-sectional data. While maintaining high predictive accuracy, the method achieves comprehensive transparency in system structure, causal relationships, and response mechanisms, thereby substantially enhancing both interpretability and causal reliability for trustworthy “What-if” analyses.

causal machine learningcausal relationshipsdecision support

Generalization and Feature Attribution in Machine Learning Models for Crop Yield and Anomaly Prediction in Germany

Dec 17, 2025
RB
Roland Baatz
🏛️ Leibniz Centre for Agricultural Landscape Research (ZALF)

This study identifies a critical disconnect between generalizability and interpretability in machine learning–based agricultural yield forecasting: while models—including XGBoost, Random Forest, LSTM, and TCN—achieve high accuracy under spatially partitioned test sets at the German NUTS-3 level, their performance deteriorates substantially when evaluated on temporally independent validation years. Crucially, even under temporal generalization failure, SHAP-based attributions retain high confidence scores, revealing a fundamental limitation of post-hoc interpretability methods—the illusion of reliability. To address this, we propose “validation-aware interpretability,” a novel paradigm that grounds model explanation in robust spatiotemporal cross-validation. We rigorously demonstrate that high test-set accuracy does not imply reliable generalization, and that trustworthy interpretation must be predicated on temporal robustness. This work establishes validation-aware interpretability as a necessary condition for credible, actionable insights in time-sensitive agroecological modeling.

Addresses trust in model explanations under temporal validationEvaluates generalization of ML models for crop yield predictionExamines reliability of feature attribution in non-generalizing models

Current machine learning evaluation practices predominantly rely on surface-level performance metrics, often neglecting the internal mechanisms of models. This work proposes trustworthy interpretability as a central evaluation paradigm and, for the first time, systematically demonstrates that it satisfies core criteria from the philosophy of science—namely falsifiability, reproducibility, and predictive power. By constructing an evaluation framework that integrates causal analysis with mechanistic probing, the study delineates three functional pathways through which interpretability enables the identification of behavioral origins, detection of latent flaws, and prediction of potential failure modes. This approach advances model assessment beyond performance-oriented benchmarks toward a deeper understanding of underlying mechanisms.

behavioral metricsinterpretabilitymachine learning

ContextualSHAP : Enhancing SHAP Explanations Through Contextual Language Generation

Dec 08, 2025
LD
Latifa Dwiyanti
🏛️ Kanazawa University | Institut Teknologi Bandung

SHAP-based visualizations of feature importance lack semantically meaningful explanations for non-technical users, undermining interpretability and trust. To address this, we propose LLM-SHAP—a framework that tightly integrates large language models (e.g., GPT) with SHAP, dynamically injecting user-provided feature aliases, domain-specific descriptions, and contextual background into the explanation generation process. This enables context-aware, natural-language interpretations without model retraining, synergistically enhancing both visual and textual explanations. We evaluate LLM-SHAP in a medical diagnosis use case via a user study. Results demonstrate statistically significant improvements over standard SHAP visualizations: +38.2% higher perceived interpretability and +41.5% greater contextual relevance among end users. Our approach establishes a novel paradigm for deploying explainable AI in real-world, non-expert settings—bridging the gap between technical model outputs and human-understandable, actionable insights.

Addresses lack of meaningful contextual explanations for non-technical usersEnhances SHAP with contextual language generation for better explanationsIntegrates SHAP with LLMs to improve perceived understandability of outputs

Hot Scholars

OA

Omran Ayoub

Lecturer-Researcher at SUPSI
Communication NetworksNetwork OptimizationMachine LearningExplainable AI
RG

Riccardo Guidotti

Associate Professor @ University of Pisa
Explainable AIData MiningClustering AlgorithmsPersonal Data Analytics
GH

Griffin Higgins

PhD Student, Canadian Institute for Cybersecurity (CIC), University of New Brunswick
privacy preserving deep learningcyber threat intelligenceapplied cybersecurity
HS

Hossein Shokouhinejad

Research Scientist
Graph LearningMalware DetectionCybersecurity in Smart GridsFault Detection