A Synthetic Ground-Truth Framework for the Evaluation of Explainable AI Methods

๐Ÿ“… 2026-09-24
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the lack of reliable benchmarks and ground-truth explanations in evaluating explainable AI (XAI) by proposing a multimodal XAI evaluation framework based on synthetic ground truth. Moving beyond conventional fidelity-only metrics, this framework generates input importance labels through controlled data interventions to construct synthetic ground-truth benchmarks aligned with model behavior across image, tabular, and time-series domains. Leveraging this framework, nine mainstream XAI methods are systematically evaluated, revealing significant limitations in existing techniques. The results validate the critical role of synthetic benchmarks in reliably quantifying explanation quality, thereby providing rigorous methodological support for XAI evaluation.
๐Ÿ“ Abstract
Evaluating explainable Artificial Intelligence (XAI) methods is a challenging task due to the lack of reliable evaluation procedures and, in particular, the absence of ground truth explanations. In the literature, existing evaluation approaches typically assess explanations by measuring their fidelity with respect to the predictions of a black-box model. However, such evaluation strategies only quantify the degree to which an explanation reproduces the model's output, without ensuring that the explanation correctly reflects the underlying decision process. As a consequence, different explanations may achieve similar fidelity scores while providing inconsistent or misleading interpretations of the model behavior. In this paper, we propose a framework for the evaluation of XAI methods based on synthetic ground truth. The proposed approach relies on controlled interventions to generate synthetic datasets in which the importance of input components can be determined by design. This enables the construction of ground truth explanations that are directly aligned with the behavior of the model under analysis. The framework is instantiated across three data domains, namely binary images, tabular data, and time series, allowing a comprehensive assessment of explanation methods in heterogeneous settings. Experimental results obtained by evaluating nine widely used XAI methods show significant limitations in current techniques and highlight the importance of synthetic, intervention-based benchmarks for a reliable assessment of explanation quality.
Problem

Research questions and friction points this paper is trying to address.

Explainable AI
Evaluation framework
Ground truth
Fidelity
Black-box model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Explainable AI (XAI)
Synthetic Ground Truth
Controlled Interventions
Evaluation Framework
Benchmark