Score
Designs, implements, and integrates methods and interfaces that produce and present explanations for individual model decisions at decision time (case-level), such as feature attributions, counterfactuals, or example-based explanations. Evaluates and analyzes these explanations for fidelity to the model, robustness to input changes, timeliness, and human interpretability/usability within decision workflows.
This study addresses the human-centered explainability challenge in interactive information systems, aiming to enhance users’ comprehension, interpretation, and critical evaluation of AI outputs to support informed decision-making. Following the PRISMA guidelines, we systematically screened and structurally coded 100 peer-reviewed articles. Based on this analysis, we propose a novel five-dimensional conceptual model of explainability—comprising transparency, causality, relevance, controllability, and actionability—and introduce the first user-needs-driven taxonomy for explanation design. Furthermore, we identify six user-centered dimensions for explainability evaluation—the first such classification in the literature. These contributions advance explainability research from ad hoc, experience-based practice toward systematic, theory-grounded inquiry. The resulting integrative framework provides both theoretical rigor and practical guidance for designing transparent, trustworthy, and responsible interactive information systems.
Existing interpretability research in machine learning predominantly focuses on input-output mapping mechanisms, neglecting the functional role of explanations within concrete application contexts—such as clinical decision support, model debugging, or remedial intervention—leading to potential misuse. Method: This paper proposes a “use-case-driven” paradigm grounded in statistical decision theory, establishing a quantifiable analytical framework for explanation utility: it formally defines the maximum performance gain an explanation can yield in a given task and characterizes its theoretical utility upper bound, while unifying evaluation criteria across diverse application scenarios. Contribution/Results: The framework shifts interpretability research from qualitative description toward a rigorous, analyzable, verifiable, and reproducible scientific paradigm. It significantly enhances the reliability and practical utility of explanation methods in real-world decision-making environments.
Contemporary machine learning models’ complexity poses significant trust, regulatory, and ethical risks, yet existing explainability guidelines lack operational specificity. To address this gap, we conducted a controlled experiment with 124 developers, integrating cognitive process theory and sociological imagination to investigate how policy frameworks influence the design of end-user–oriented explanations for diabetic retinopathy screening models. Results reveal that all participants struggled to generate high-quality, policy-compliant, and empirically verifiable explanations; over 70% failed to accurately anticipate users’ comprehension barriers; and widely adopted technical methods (e.g., SHAP, Anchors) exhibit fundamental misalignment with real-world stakeholder needs. Our core contribution is identifying *developers’ inability to empathize with non-technical stakeholders* as the central mechanism underlying explanation failure—and proposing, for the first time, an empathy-centered educational intervention framework to bridge this gap.
Existing business process modeling practices separate process flows from business rules, leading to fragmented mental models and impaired comprehension among expert process workers. Method: This study employs a mixed-methods design integrating eye-tracking and concurrent verbal protocol analysis, coupled with cognitive-behavioral coding, to investigate how domain experts perform sensemaking when interpreting integrated process-rule models. Contribution/Results: We identify fine-grained visual search patterns and cognitive bottlenecks that critically affect comprehension efficiency during information foraging and cognitive processing stages. Based on these findings, we propose empirically grounded design principles for personalized cognitive support targeting knowledge workers. The results provide actionable evidence to enhance integrated modeling languages, tool interfaces, and training strategies—advancing business process modeling from syntactic formalism toward cognitive alignment and human-centered design.
This study addresses the lack of a domain-agnostic, human-centered explainable artificial intelligence (XAI) framework by investigating user preferences across healthcare, retail, and energy domains. Through expert interviews and multi-stakeholder structured surveys, it empirically identifies “interpretability over accuracy” as a cross-domain preference and establishes feature importance and counterfactual explanations as the two foundational pillars of a universal XAI framework. Method: The approach integrates qualitative transcription analysis, questionnaire-driven requirement modeling, and genetic programming (GP) to construct inherently interpretable models. Contribution/Results: We propose the first empirically validated, unified XAI framework spanning multiple domains; release an open-source, standardized XAI questionnaire toolkit; demonstrate its feasibility across three core machine learning tasks—prediction, diagnosis, and prescription—and advance the XAI paradigm from technology-centric design toward human consensus–driven development.
This study addresses the prevalent ambiguity, inconsistency, and incompleteness in articulating explainability requirements for AI systems due to a lack of standardized specifications. Through a structured literature review and interviews with developers, the authors identify a set of explainability quality attributes, which are then refined via a large-scale survey of practitioners into ten core attributes. For the first time, these attributes are translated into a prioritized, actionable guideline for writing explainability requirements. Building on this foundation, the authors design a lightweight, iterative requirements engineering workflow augmented by a large language model to assist in requirement generation. An accompanying web-based tool reduces average requirement drafting time by 23.5%, and user evaluations indicate that the generated requirements match or slightly exceed manually written ones in terms of implementability and textual quality.
This work addresses the limitations of existing explainable AI methods, which predominantly focus on associative predictions and fall short in supporting decision-making that requires causal reasoning and counterfactual analysis. To bridge this gap, the paper proposes a novel framework that integrates causal machine learning with intrinsically interpretable models—such as additive models and symbolic regression—by explicitly embedding causal inference mechanisms within the model architecture. This approach enables the explicit recovery of causal structures and functional forms among variables directly from cross-sectional data. While maintaining high predictive accuracy, the method achieves comprehensive transparency in system structure, causal relationships, and response mechanisms, thereby substantially enhancing both interpretability and causal reliability for trustworthy “What-if” analyses.
This study addresses the persistent challenge that explainable AI (XAI) often fails to effectively support human decision-making due to poor user comprehension. To bridge this gap, the work integrates cognitive modeling with user studies to formally represent—within a computationally tractable framework—the reasoning strategies humans employ when interacting with different XAI methods in structured data tasks. Through formative and summative user experiments, feature attribution analyses, and behavioral alignment evaluations, the resulting cognitive model demonstrates significantly greater accuracy than conventional machine learning surrogates in capturing human forward-simulation decision behavior. Beyond elucidating which XAI mechanisms genuinely aid human judgment, the model offers an empirically grounded foundation for designing more effective XAI systems and serves as a high-fidelity, efficient proxy for human-subject experimentation in XAI research.
Rapid AI model evolution has led to ad hoc, non-reproducible model selection in scientific software engineering, severely undermining reproducibility and transparency. To address this, we propose ModelSelect—the first evidence-driven framework that formalizes AI model selection as a multi-criteria decision-making (MCDM) problem, integrating automated metadata harvesting, a structured knowledge graph, and research-context-aware decision modeling. Its key contributions are: (1) establishing the first MCDM modeling paradigm tailored to scientific practice; (2) end-to-end integration of technical metrics and domain semantics, significantly enhancing recommendation interpretability and consistency; and (3) empirical validation across 50 real-world research scenarios, achieving 96.2% coverage, substantially higher rationale alignment than baselines, and superior traceability, cross-scenario robustness, and transparency.