Score
Designs, implements, and evaluates methods and tools that generate post-hoc explanations for ensemble model predictions, producing per-instance attributions or summaries that identify which features, base models, or components drove each prediction. Communicates and formats these explanations so users can inspect, compare, and reason about ensemble outputs to support transparent decision-making and model analysis.
This study addresses the challenge that post-hoc interpretability tools for credit risk models—such as SHAP—produce numerical outputs that are difficult for non-technical stakeholders to comprehend. It presents the first systematic evaluation of large language models (LLMs) in two distinct roles: as “translators” that convert attribution-based explanations into natural language narratives, and as “autonomous explainers” that generate explanations directly from input data. Using LendingClub data with logistic regression and XGBoost models, the authors experiment with few-shot prompting on GPT-4-turbo, Claude Sonnet 4, and Gemini-2.0-Flash. Results show that LLMs excel as translators, producing intelligible and auditable narrative explanations, but perform poorly as autonomous explainers, exhibiting low fidelity to model attributions—especially for nonlinear models. The work thus advocates integrating LLMs as complementary narrative interfaces to, rather than replacements for, established interpretability methods.
This study addresses the lack of a domain-agnostic, human-centered explainable artificial intelligence (XAI) framework by investigating user preferences across healthcare, retail, and energy domains. Through expert interviews and multi-stakeholder structured surveys, it empirically identifies “interpretability over accuracy” as a cross-domain preference and establishes feature importance and counterfactual explanations as the two foundational pillars of a universal XAI framework. Method: The approach integrates qualitative transcription analysis, questionnaire-driven requirement modeling, and genetic programming (GP) to construct inherently interpretable models. Contribution/Results: We propose the first empirically validated, unified XAI framework spanning multiple domains; release an open-source, standardized XAI questionnaire toolkit; demonstrate its feasibility across three core machine learning tasks—prediction, diagnosis, and prescription—and advance the XAI paradigm from technology-centric design toward human consensus–driven development.
Communication barriers between data scientists and domain experts arise from oversimplified, accuracy-centric model performance reporting, hindering shared understanding of model limitations and contextual applicability. Method: We propose a visualization-mediated model explanation framework grounded in human-computer interaction principles, participatory design, and visual narrative techniques. This yields the first domain-expert-oriented model communication guideline—emphasizing risk, trade-offs, and situational appropriateness rather than isolated metrics like accuracy. An iterative empirical study was conducted using regression models, incorporating structured expert feedback for evaluation. Contribution/Results: The framework significantly improves domain experts’ ability to identify model limitations, recognize inherent trade-offs, and proactively make context-driven adoption decisions. Its core innovation lies in repositioning visualization as an interdisciplinary consensus-building medium—shifting the paradigm from “metric reporting” to “collaborative understanding.”
This work addresses the fundamental trade-off in machine learning between high predictive performance and low interpretability inherent in “black-box” models (e.g., deep neural networks, ensemble methods). It rigorously distinguishes post-hoc explanation—applied after model training—from inherently interpretable modeling—designed for transparency from inception. To reconcile accuracy and interpretability, we propose a hybrid modeling paradigm centered on symbolic knowledge embedding, integrating differentiable symbolic modules, knowledge distillation, and symbolic reasoning into the model architecture itself. This enables joint optimization of fidelity and interpretability at the design stage. Extensive experiments across diverse domains demonstrate that our approach matches the predictive accuracy of state-of-the-art black-box models while generating human-understandable, logically grounded decision rules. As a result, it substantially enhances model trustworthiness and deployment viability in safety- and accountability-critical applications.
Quantifying the contribution of individual submodels to the overall predictive performance of ensemble systems is crucial for enhancing interpretability and construction efficiency. This work proposes and implements a unified R package that, for the first time in the R ecosystem, systematically supports model importance assessment across diverse ensemble methods under both point and probabilistic forecasting frameworks, with full compatibility with the hubverse infrastructure. The package offers flexible importance metrics and robust handling of missing values, substantially improving the understanding of submodel roles. It thereby empowers researchers to efficiently construct, diagnose, and optimize ensemble forecasting systems.
Current machine learning evaluation practices predominantly rely on surface-level performance metrics, often neglecting the internal mechanisms of models. This work proposes trustworthy interpretability as a central evaluation paradigm and, for the first time, systematically demonstrates that it satisfies core criteria from the philosophy of science—namely falsifiability, reproducibility, and predictive power. By constructing an evaluation framework that integrates causal analysis with mechanistic probing, the study delineates three functional pathways through which interpretability enables the identification of behavioral origins, detection of latent flaws, and prediction of potential failure modes. This approach advances model assessment beyond performance-oriented benchmarks toward a deeper understanding of underlying mechanisms.
Existing time series models lack effective methods for global interpretability, often limited to local explanations. This work introduces example-based global explanations to time series forecasting for the first time, proposing a model-agnostic, user-centered approach that automatically constructs representative subsets of time series by balancing sample importance and diversity through a domain-specific utility function, thereby generating concise and transparent summaries of global model behavior. Experiments demonstrate that the method efficiently produces high-quality summaries that are readily amenable to human evaluation. User studies further reveal that domain experts significantly prefer the model understanding provided by this approach, leading to a marked improvement in their comprehension of overall model behavior.
This study addresses a critical epistemological gap in scientific machine learning: while post-hoc interpretability methods—such as feature importance and counterfactual explanations—enhance model transparency, they remain insufficient for substantiating claims about the true causal structures or mechanistic underpinnings of natural phenomena. Through an integrated philosophical analysis and critique of scientific modeling practices, the paper systematically demonstrates that even when a model’s predictions are reliable and its explanations faithfully reflect its internal logic, this does not warrant inferences about the actual mechanisms governing the target phenomenon. The work exposes fundamental cognitive limitations of explainable AI in scientific discovery, arguing that external validation is indispensable for formulating credible scientific hypotheses. On this basis, it proposes a novel epistemological framework to delineate the appropriate boundaries for employing model-based explanations within scientific reasoning.
This work addresses the limitations of existing explainable AI methods, which predominantly focus on associative predictions and fall short in supporting decision-making that requires causal reasoning and counterfactual analysis. To bridge this gap, the paper proposes a novel framework that integrates causal machine learning with intrinsically interpretable models—such as additive models and symbolic regression—by explicitly embedding causal inference mechanisms within the model architecture. This approach enables the explicit recovery of causal structures and functional forms among variables directly from cross-sectional data. While maintaining high predictive accuracy, the method achieves comprehensive transparency in system structure, causal relationships, and response mechanisms, thereby substantially enhancing both interpretability and causal reliability for trustworthy “What-if” analyses.