Score
Designs and implements methods and models that produce human-readable, interpretable rationales for automated outputs or decisions, including chain-of-thought decompositions and multi-factor explanations. Builds components to elicit, generate, and stabilize repeatable textual explanations that support recommendations or decisions and to analyze their fidelity and interpretability.
This paper addresses the lack of systematic design and evaluation frameworks for prompt-based natural language explanations (NLEs) in transparent AI governance. We propose the first taxonomy specifically tailored to prompt-driven NLEs, structured along three dimensions: context, generation & presentation, and evaluation. The taxonomy integrates task objectives, data characteristics, stakeholder needs, generation mechanisms, interaction modalities, output formats, and user-centered evaluation criteria into a multidimensional, structured classification. Its key innovation lies in adapting eXplainable AI (XAI) taxonomic principles to the prompt-based NLE paradigm—moving beyond conventional model-internal explanations. The resulting framework provides researchers, auditors, and policymakers with actionable design guidelines and standardized evaluation benchmarks, thereby enhancing the practical utility of NLEs in transparency, auditability, and regulatory compliance. (149 words)
AI deployment in high-stakes domains demands explainability to foster trust and accountability, yet prevailing XAI methods often neglect foundational human cognitive mechanisms. This paper introduces the first XAI framework systematically integrating Malle’s five-category model of human explanatory reasoning—namely, knowledge structure, simulation/prediction, covariation, direct recall, and rationalization—by unifying attribution analysis, feature importance, attention visualization, and large language model–generated reasoning chains. The resulting multimodal explanation framework aligns technical outputs with empirically grounded human explanation preferences. Empirical evaluation in real-world credit risk assessment and regulatory compliance tasks demonstrates significant improvements in users’ depth of understanding and perceived trustworthiness of AI decisions. The core contribution lies in pioneering a cognition-informed design paradigm for explainability—shifting XAI from mere *technical interpretability* toward *human comprehensibility*, *acceptability*, and *reliability*.
This study addresses the opacity and non-auditable nature of large language model (LLM) decision-making. Methodologically, it introduces a modular, interpretable LLM agent framework that integrates deterministic analyzers—including Vester’s sensitivity analysis, normal-form and sequential game modeling, matrix classification, and backward induction—with an LLM (default: GPT-5) to jointly generate explicit, traceable intermediate reasoning artifacts. The framework supports dynamic switching among analytical paradigms and role-conditioned agency. Its key contribution is a dual-track “LLM + deterministic analyzer” architecture that preserves reasoning flexibility while enabling end-to-end auditability. Evaluated on a real-world logistics decision-making case, the approach achieves a factor alignment rate of 55.5% (full dataset) and 62.9% (core subset), a role-matching accuracy of 57%, and LLM-generated assessments statistically comparable to human expert baselines.
This study addresses controllable readability generation: guiding large language models to produce accurate, faithful natural language explanations for reasoning tasks tailored to diverse cognitive levels (e.g., sixth graders, college students), balancing comprehensibility and factual consistency. We conduct the first systematic investigation into how readability level affects free-text reasoning generation, proposing a prompt-engineering–based controllable generation framework. Evaluation integrates traditional readability metrics (e.g., Flesch-Kincaid), multidimensional automated assessment (BERTScore, QA-based faithfulness), and human annotation. Results demonstrate strong model controllability over readability; explanations of medium complexity achieve optimal performance across both automatic metrics and human evaluation—exhibiting the lowest hallucination and misinterpretation rates—while high-school–level explanations receive the highest human preference. Our core contribution is the empirical identification of the readability–faithfulness trade-off and the validation of feasible, effective controllable explanation generation.
This study addresses the lack of a domain-agnostic, human-centered explainable artificial intelligence (XAI) framework by investigating user preferences across healthcare, retail, and energy domains. Through expert interviews and multi-stakeholder structured surveys, it empirically identifies “interpretability over accuracy” as a cross-domain preference and establishes feature importance and counterfactual explanations as the two foundational pillars of a universal XAI framework. Method: The approach integrates qualitative transcription analysis, questionnaire-driven requirement modeling, and genetic programming (GP) to construct inherently interpretable models. Contribution/Results: We propose the first empirically validated, unified XAI framework spanning multiple domains; release an open-source, standardized XAI questionnaire toolkit; demonstrate its feasibility across three core machine learning tasks—prediction, diagnosis, and prescription—and advance the XAI paradigm from technology-centric design toward human consensus–driven development.
Current automated methods for evaluating interpretability struggle to keep pace with increasingly complex autonomous explanatory agents, particularly due to the subjectivity, opacity, and memory biases inherent in paradigms that aim to replicate human expert explanations. This work proposes a novel unsupervised intrinsic evaluation approach grounded in the functional interchangeability of model components and implements it within a large language model–driven autonomous research agent system that iteratively designs experiments and tests hypotheses in circuit analysis tasks. Experiments across six benchmark tasks show that while the system’s performance appears comparable to that of human experts, deeper analysis exposes fundamental flaws in replication-based evaluation. The proposed method effectively overcomes key limitations of traditional paradigms, offering a more reliable pathway for interpretability assessment.
This study addresses the question of what user-relevant content should be included in local, post-hoc explanations for industrial AI systems and how such content should be organized. Through a hybrid inductive–deductive qualitative content analysis, the authors integrate user research data from six domains with explainable AI theory to develop the first user-oriented model comprising fourteen categories of local explanation content. The model encompasses rule-based and causal dimensions alongside two cognitive dimensions. It demonstrates high content validity and clear conceptual boundaries, as evidenced by expert review, an item-level content validity index (I-CVI ≥ 0.82), and strong intercoder reliability (Krippendorff’s α = Cohen’s κ = 0.920). This framework provides both theoretical grounding and practical guidance for the elicitation, standardization, and evaluation of explanation content in industrial AI applications.
This work addresses the limitations of existing software documentation, which often suffers from poor consistency, weak relevance, and unclear expression, while purely automated generation methods struggle with insufficient reliability and controllability. To overcome these challenges, the authors propose a human-AI collaborative approach that integrates empirically derived quality guidelines with large language model (LLM) assistance through a generate–evaluate–iterate workflow. This method preserves developers’ domain control while enhancing explanation quality by embedding experience-driven quality criteria directly into the LLM-assisted writing process, enabling controllable, efficient, and high-quality output. Preliminary experiments demonstrate a 24.4% average improvement in authoring efficiency with tool support, and user studies reveal significantly higher satisfaction with the generated explanations compared to purely manual writing (p = 0.003, effect size = 0.86).
This study addresses the persistent challenge that explainable AI (XAI) often fails to effectively support human decision-making due to poor user comprehension. To bridge this gap, the work integrates cognitive modeling with user studies to formally represent—within a computationally tractable framework—the reasoning strategies humans employ when interacting with different XAI methods in structured data tasks. Through formative and summative user experiments, feature attribution analyses, and behavioral alignment evaluations, the resulting cognitive model demonstrates significantly greater accuracy than conventional machine learning surrogates in capturing human forward-simulation decision behavior. Beyond elucidating which XAI mechanisms genuinely aid human judgment, the model offers an empirically grounded foundation for designing more effective XAI systems and serves as a high-fidelity, efficient proxy for human-subject experimentation in XAI research.
This study addresses the misalignment between humans and AI in data-driven cognition, particularly regarding semantic understanding and reasoning intent. To bridge this gap, the authors propose a bidirectional alignment framework grounded in complementary Explorer-Guide roles. The framework enables both parties to explicitly articulate goals, hypotheses, evidence, and inferences within a shared reasoning space, distinguishing mechanisms for “AI aligning to humans” and “humans aligning to AI.” It integrates theories of human–machine collaboration, models of critical thinking, and the semantic reasoning capabilities of large language models. The work contributes an interpretable and collaborative paradigm for cognitive alignment and advocates for fine-grained, role-specific evaluation metrics, thereby advancing a new agenda for the design and assessment of human–AI collaborative alignment.