Score
Designs and implements quantitative metrics, scoring functions, and evaluation pipelines that measure how faithfully model explanations reflect model predictions and decision mechanisms, covering perturbation-driven fidelity tests, composite scores (e.g., entityscore, evidencescore, EQMs), and specialized few-class fidelity metrics. Also develops procedures to assess explanation trustworthiness and overall quality — including scoring, statistical validation, self-explanation assessment, and comparisons to human localization — and analyzes how these metrics relate to interpretability and downstream use.
This study investigates whether existing algorithmic evaluation metrics for counterfactual explanations align with users’ perceptions of explanation quality. Through user studies conducted on three datasets, the authors systematically compare widely used algorithmic metrics against multidimensional human subjective ratings of counterfactual explanations. Employing correlation analyses and multivariate regression models, they assess the consistency and predictive power of these metrics. The findings reveal that algorithmic metrics generally exhibit weak correlations with human judgments and are highly dataset-dependent. Moreover, increasing the number of metrics yields only marginal improvements in predictive performance. These results expose structural limitations in current evaluation practices, underscoring their inability to capture key aspects of explanation quality that matter to users, and provide empirical support for advancing human-centered evaluation paradigms in explainable AI.
Existing XAI evaluation lacks reliable, principled reference baselines for quantifying explanation quality. Method: This paper proposes Quality Gap Estimation (QGE), the first method to introduce “inverse explanations” as a conceptual quality reference—approximated via counterfactual perturbations and latent-space inverse mapping—to enable relative, comparable quantification of individual explanations across dimensions including faithfulness, localization, and robustness. QGE abandons conventional random baselines, instead grounding evaluation in semantically meaningful counterfactuals. Contribution/Results: QGE significantly improves statistical robustness and cross-model/cross-dataset transferability of explanation assessment. Experiments across diverse architectures and datasets show that QGE increases ranking consistency of explanation quality by 32% and reduces evaluation variance by 41% compared to random baselines, thereby enhancing the reliability of model behavior diagnosis.
This study addresses the uncontrolled quality of quality engineering (QE) artifacts—such as requirements specifications, test cases, and Behavior-Driven Development (BDD) scenarios—automatically generated by large language models (LLMs). We propose an iterative optimization framework integrating forward generation, backward generation, and rubric-guided scoring to enhance artifact quality along four dimensions: clarity, completeness, consistency, and testability. Our approach enables automated, quantitative, and reproducible quality assessment and improvement. Evaluated across 12 real-world projects, the method significantly improves output stability: it preserves high quality under high-quality inputs and substantially outperforms baselines under low-quality inputs. The core contribution is the first integration of backward generation with structured rubric-based guidance, establishing a closed-loop, artifact-centric quality enhancement paradigm for QE.
Current machine learning evaluation practices predominantly rely on surface-level performance metrics, often neglecting the internal mechanisms of models. This work proposes trustworthy interpretability as a central evaluation paradigm and, for the first time, systematically demonstrates that it satisfies core criteria from the philosophy of science—namely falsifiability, reproducibility, and predictive power. By constructing an evaluation framework that integrates causal analysis with mechanistic probing, the study delineates three functional pathways through which interpretability enables the identification of behavioral origins, detection of latent flaws, and prediction of potential failure modes. This approach advances model assessment beyond performance-oriented benchmarks toward a deeper understanding of underlying mechanisms.
Existing evaluation of explainable recommender systems over-relies on recommendation performance and subjective user feedback, lacking objective, content-oriented metrics for explanation veracity—the intrinsic informational quality of explanations. This paper introduces signal detection theory to explanation evaluation for the first time, proposing a dual-dimensional decomposition framework grounded in fidelity (explanation’s faithfulness to the underlying model) and attunement (explanation’s alignment with user expectations). Based on this, we develop a quantifiable Veracity scoring model that integrates decision-sensitivity analysis. Through multi-scenario simulation experiments, we demonstrate that the model effectively discriminates among explanations of varying informational quality, exhibiting strong discriminative power and robustness. Our work bridges a critical gap in objective, content-centric explanation assessment and establishes a novel, principled benchmark for developing and evaluating explainable recommender systems.
This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.
This work addresses a central challenge in explainable artificial intelligence (XAI): objectively evaluating the fidelity of explanations and aligning them with human understanding. The authors propose EPC, a model-agnostic scoring metric that quantifies explanation quality by jointly optimizing feature sparsity and preservation of model performance. For the first time, the method validates automatically generated explanations against human-provided semantic sentiment judgments and visual spatial annotations across multimodal data—including tabular, textual, and image domains—demonstrating strong alignment. Experiments show that the EPC score not only effectively reveals dependencies between explainer performance and factors such as network activations and data dimensionality but also exhibits high correlation with human-centered evaluations, thereby establishing a reliable and generalizable paradigm for XAI assessment.
This study addresses the prevalent ambiguity, inconsistency, and incompleteness in articulating explainability requirements for AI systems due to a lack of standardized specifications. Through a structured literature review and interviews with developers, the authors identify a set of explainability quality attributes, which are then refined via a large-scale survey of practitioners into ten core attributes. For the first time, these attributes are translated into a prioritized, actionable guideline for writing explainability requirements. Building on this foundation, the authors design a lightweight, iterative requirements engineering workflow augmented by a large language model to assist in requirement generation. An accompanying web-based tool reduces average requirement drafting time by 23.5%, and user evaluations indicate that the generated requirements match or slightly exceed manually written ones in terms of implementability and textual quality.