explanation robustness testing

Designs and implements evaluation protocols, metrics, and benchmarks to test the robustness, stability, and reliability of model explanations such as saliency maps, feature attributions, and other XAI outputs. Builds perturbation and attack strategies, evaluation pipelines, statistical analyses, and visualizations to quantify fidelity, sensitivity, and susceptibility to manipulation or randomness across models, inputs, and explanation methods.

explanationrobustnesstesting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.13
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

When Can You Trust Your Explanations? A Robustness Analysis on Feature Importances

Jun 20, 2024
IV
Ilaria Vascotto
🏛️ University of Trieste | The Abdus Salam International Center for Theoretical Physics | Assicurazioni Generali Spa

This work addresses the robustness evaluation of feature importance explanations in eXplainable Artificial Intelligence (XAI) under non-adversarial perturbations. Existing methods fail to model natural perturbations intrinsic to the data manifold; we thus present the first systematic analysis of explanation fragility under such realistic, non-adversarial distortions. To ensure perturbations remain faithful to the underlying data distribution, we propose a manifold-aware perturbation generation strategy grounded in the manifold hypothesis. Further, we introduce a multi-explainer ensemble framework that aggregates explanations via consistency-based fusion to jointly enhance robustness and interpretability. Extensive experiments on multiple tabular datasets reveal that mainstream explanation methods—including Grad-CAM and SHAP variants—exhibit substantial robustness deficiencies. In contrast, our ensemble framework significantly improves explanation stability and decision trustworthiness. We publicly release an open-source evaluation framework enabling reproducible, quantitative robustness assessment across diverse XAI methods.

Analyzing robustness of neural network explanations to perturbationsProposing ensemble method to aggregate and evaluate explanationsProviding framework for assessing trustworthiness of model explanations

The Effect of Similarity Measures on Accurate Stability Estimates for Local Surrogate Models in Text-based Explainable AI

Jun 22, 2024
CB
Christopher Burger
🏛️ University of Mississippi | Indiana University

This work addresses the over-sensitivity of ranking-based similarity metrics—such as Kendall’s Tau, Spearman’s Footrule, and Rank-biased Overlap—to minor perturbations in evaluating the stability of local surrogate models (e.g., LIME) for text-based eXplainable AI (XAI). We systematically analyze their behavior under adversarial perturbations and find that uncalibrated metrics (e.g., vanilla Kendall’s Tau) significantly overestimate model fragility, leading to erroneous robustness assessments; this stems from excessive responsiveness to trivial rank shifts rather than semantic relevance. To remedy this, we propose the first task-aware stability analysis framework, which jointly adapts metric selection and thresholding to downstream interpretability objectives. Empirical validation demonstrates that principled metric configuration substantially improves the reliability and cross-method comparability of XAI evaluations. Our results underscore that similarity metric choice must be semantically grounded in the specific explanation task—not defaulting to generic, off-the-shelf measures.

Adversarial PerturbationsSimilarity MeasurementStability Assessment

A Systematic Literature Review on Explainability for Machine/Deep Learning-based Software Engineering Research

Jan 26, 2024
SC
Sicong Cao
🏛️ Yangzhou University | Singapore Management University | University of Southern Queensland | Washington University in St. Louis

Insufficient interpretability of machine learning models in software engineering (SE), particularly undermining decision transparency in critical tasks such as vulnerability detection. Method: We conduct a systematic literature review (SLR) analyzing 108 peer-reviewed papers spanning 23 SE tasks. Contribution/Results: This study establishes the first comprehensive XAI (Explainable AI) landscape for SE, identifying six high-value application scenarios and seven mainstream technical approaches. It reveals critical gaps in evaluation methodologies—especially the absence of structured, empirically grounded benchmarks for SE-XAI. Furthermore, we synthesize a prioritized list of industrial deployment challenges and actionable best practices. By bridging the gap between theoretical XAI research and SE practice, this work provides a foundational, evidence-based reference for both academic investigation and real-world engineering adoption of interpretable ML in software development and assurance.

Challenges in AI-driven SE modelsExplainability of AI in Software EngineeringSystematic review of XAI techniques

X Hacking: The Threat of Misguided AutoML

Jan 16, 2024
RS
Rahul Sharma
🏛️ Deutsches Forschungszentrum für Künstliche Intelligenz GmbH (DFKI)

This paper introduces “X-hacking”—a novel form of methodological bias wherein researchers systematically search the Rashomon set (i.e., the set of models with comparable predictive performance but divergent explanations) to manipulate XAI attribution metrics (e.g., SHAP values) in support of preconceived conclusions, analogous to p-hacking in statistics. To empirically validate X-hacking, the authors propose a multi-objective optimization framework that identifies models satisfying both high predictive accuracy and *a priori* desired explanation patterns, leveraging AutoML on UCI/tabular benchmarks. Results demonstrate that X-hacking is prevalent across mainstream XAI practices, critically undermining the trustworthiness and reproducibility of explainable AI. As a key contribution, the paper presents the first dedicated detection and mitigation mechanism for X-hacking, advocating a paradigm shift in XAI evaluation—from isolated explanation fidelity toward joint verification of explanation validity and predictive performance.

Analyzes vulnerability to X-hacking via feature information redundancyDemonstrates automated exploitation of model multiplicity for desired explanationsExposes manipulation of XAI metrics to support biased conclusions

A Unified Framework for Evaluating the Effectiveness and Enhancing the Transparency of Explainable AI Methods in Real-World Applications

Dec 05, 2024
MA
Md. Ariful Islam
🏛️ American International University-Bangladesh | Okinawa Institute of Science and Technology Graduate University (OIST) | Techno International New Town

Existing XAI methods lack a standardized, real-world-scenario-oriented evaluation framework, hindering systematic validation of their correctness, intelligibility, robustness, fairness, and completeness. To address this, we propose the first unified, multidimensional XAI evaluation framework integrating quantitative and qualitative metrics. Our approach encompasses: (1) construction of standardized benchmarks; (2) domain-adaptive interface design; (3) end-to-end integration of explanation generation and evaluation; (4) human-in-the-loop verification mechanisms; and (5) cross-domain, case-driven testing methodologies. We empirically validate the framework across four high-stakes domains—healthcare, finance, agriculture, and autonomous driving. Results demonstrate significant improvements in evaluation consistency, credibility, and practical deployability. The framework establishes a reusable, generalizable operational standard for AI transparency and algorithmic accountability.

Absence of unified framework for real-world XAI evaluationLack of standard metrics to evaluate XAI method effectivenessNeed for transparent and trustworthy AI decision explanations

Latest Papers

What's happening recently
View more

This work addresses a critical limitation of existing explainable AI methods, which predominantly offer passive attribution and thus fail to support practitioners in effectively intervening on model behavior. To bridge this gap, the authors propose an interactive analysis workflow that integrates sparse autoencoder (SAE)-based attribution with activation intervention, introducing activation steering into human-in-the-loop debugging for the first time and enabling a paradigm shift from observation to active intervention. Through semi-structured interviews with eight domain experts, the study reveals that users commonly engage in intervention-based hypothesis testing, primarily employing component suppression strategies and grounding their trust in model responses rather than the plausibility of explanations. The research also uncovers key risks—including ripple effects and limited instance-level generalizability of corrections—thereby charting a new path toward trustworthy AI debugging.

actionable explanationsactivation steeringExplainable AI

This work proposes an end-to-end trustworthy detection framework to address three major challenges in cybersecurity threat detection: large-scale data volume, high risk of feature leakage, and opaque model decisions. The framework uniquely integrates strategic sampling—preserving class distribution to enhance training efficiency—an automated data leakage prevention mechanism, and model-agnostic SHAP-based interpretability analysis. Experimental evaluation on the CIC-IDS2017 dataset demonstrates that the proposed approach significantly reduces computational overhead while maintaining high detection performance. Furthermore, it delivers actionable explanations for security analysts, thereby facilitating the practical deployment of trustworthy AI in Security Operations Centers (SOCs).

Cybersecurity Threat DetectionData Leakage PreventionExplainable AI

The evaluation of explainable AI (XAI) methods is affected by a lack of standardization. Metrics are inconsistently defined, incompletely reported, and rarely validated against common baselines. In this paper, we identify transparency of evaluation reporting as a central, under-addressed problem. We propose the XAI Evaluation Card, a documentation template analogous to model cards, designed to accompany any study that introduces an XAI evaluation metric. The card covers explicit declaration of target properties, grounding levels, metric assumptions, validation evidence, gaming risks, and known failure cases. We argue that adopting this template as a community norm would reduce evaluation fragmentation, support meta-analysis, and improve accountability in XAI research.

Evaluation MetricsExplainable AIReporting

This work addresses the lack of a unified, multidimensional standard for evaluating explainability, which hinders reliable comparisons across models, datasets, and explainable AI (XAI) methods. The authors propose a comprehensive evaluation framework encompassing fidelity, conciseness, and stability, integrating diverse metrics into a unified explainability scoring system for the first time. Through systematic benchmarking of mainstream XAI approaches—including LIME and SHAP—on multiple open-source datasets, they construct an explainability meta-knowledge base. This knowledge base enables predictive scoring of explainability for new models and data contexts, significantly enhancing contextual adaptability. Experimental results validate the framework’s effectiveness and uncover patterns in how explainability varies with model architecture, data characteristics, and user backgrounds, offering a practical and scalable tool for trustworthy AI development.

evaluation metricexplainabilitymultidimensional

Existing explanation methods struggle to effectively distinguish model behaviors within Rashomon sets of comparable performance and are vulnerable to spurious explanations. To address this, this work proposes AXE, the first evaluation framework that does not rely on ground-truth explanations. AXE is grounded in three principled criteria and assesses explanation quality by analyzing the consistency of feature importance under unlabeled settings. It reliably identifies adversarial fairness-washing attacks with 100% accuracy, detects whether models implicitly leverage protected attributes, and overcomes limitations of conventional sensitivity-based or ground-truth-comparison approaches. By revealing behavioral differences among models in Rashomon sets, AXE facilitates the selection of more trustworthy models.

explainable AIexplanation evaluationfairwashing

Hot Scholars

ML

Michael Lognoul

Legal researcher, CRIDS/Nadi, University of Namur
EU LawICTAIXAI
GV

Giulia Vilone

Analog Devices International
Artificial IntelligenceeXplainable Artificial Intelligence
FS

Francesco Sovrano

ETH Zurich, Collegium Helveticum
AI for Software EngineeringResponsible AIAI and LawXAI